Επιστροφή στις Ειδήσεις
ΠροϊόνAI Understanding ενημέρωση

Το llama.cpp 0.4.0 προσθέτει υποστήριξη για νεότερα μοντέλα AI και είσοδο βίντεο

Η έκδοση llama.cpp 0.4.0 προσθέτει αρχική υποστήριξη για Qwen3.8-Flash-Next και NVIDIA Nemotron-3-Puzzle, επιλογές εισόδου βίντεο, όρια περιβάλλοντος διακομιστή ανά υποδοχή, ανάγνωση τεμπέλης τανυστήρα και αλλαγές ggml 0.23.0.

4 min readRead the primary source
Source-page capture accompanying llama.cpp 0.4.0 adds support for newer AI models and video input
Έγγραφο κύριας πηγήςΗ πηγή καταγράφηκε
Εκδότης
github.com
Σύνδεσμος πηγής
github.comhttps://github.com/ggml-org/llama.cpp/releases/tag/v0.4.0
Τύπος πηγής
Κύριο έγγραφο — μια επίσημη ανακοίνωση, χαρτί, αρχειοθέτηση ή σελίδα πρώτου μέρους που διαβάζουμε απευθείας.
Αναφέρεται επίσης

Η ιστορία αναθεωρήθηκε τελευταία

ΠλαίσιοΚαταλάβετε αυτό σε 60 δευτερόλεπτα

Ξεκινήστε εδώ

Βασικοί όροι

API (Διεπαφή προγραμματισμού εφαρμογών)
Ένας δομημένος τρόπος για ένα σύστημα λογισμικού να στέλνει αιτήματα και να λαμβάνει απαντήσεις από ένα άλλο σύστημα.
Μνήμη (μνήμη πράκτορα)
Το αποθηκευμένο πλαίσιο που χρησιμοποιεί ένας πράκτορας τεχνητής νοημοσύνης σε βήματα ή περιόδους σύνδεσης για να βελτιώσει τη συνέχεια.
Πολυτροπικό μοντέλο
Ένα μοντέλο που μπορεί να επεξεργαστεί ή να δημιουργήσει πολλούς τύπους δεδομένων όπως κείμενο, εικόνα και ήχος.
Δοκιμάστε τον εαυτό σαςΕξηγημένο Κουίζ Μοντέλων AI

Τι άλλαξε από τη δημοσίευση

  1. Πρωτοδημοσιεύτηκε
  2. The earlier llama.cpp entry covered pre-release support for NVIDIA’s Nemotron-3-Puzzle. Version 0.4.0 now includes Nemotron-3-Puzzle-75B-A9B support in a broader tagged release that also adds newer model architectures, video input, server context controls, lazy tensor reading, and ggml 0.23.0 backend work.

Τι έγινε

llama.cpp released version 0.4.0 on GitHub on September 4. The release adds initial support for Qwen3.8-Flash-Next, NVIDIA Nemotron-3-Puzzle-75B-A9B, DSpark for Nemotron 3.5, nanbeige4.2-3B, and DeepSeek-V4-Flash-Vision-Exp. It also adds video-input parameters, per-slot server context limits, lazy tensor reading, expert-routing changes, KV-cache improvements, and multiple backend optimizations. The release updates ggml from 0.22.0 to 0.23.0. The project says that version adds sparse flash attention, asynchronous execution and allocation-dependency APIs, RPC event and async APIs, and Apple RDMA transport support. The page lists a nightly build identified as b10809. It does not document packaged-binary availability, installation requirements, model-weight distribution, or pricing.

The GitHub release page identifies v0.4.0 as the latest release and records a September 4 publication time of 19:56, without specifying a timezone in the supplied text. The release includes API changes such as llama_lazy_mode, quantizer buffer-size controls, updated session and state versions, and new multimodal tokenization helpers.

Model and core changes include initial Qwen3.8-Flash-Next architecture support, NVIDIA Nemotron-3-Puzzle-75B-A9B support, DSpark support for Nemotron 3.5, nanbeige4.2-3B support, lazy tensor reading, per-layer expert routing, n-gram history lookup, KV-cache restoration improvements, and safeguards against RAM peaks during model loading.

Multimodal and server changes include DeepSeek-V4-Flash-Vision-Exp support, video command-line options, per-slot context limits, data URLs for media, default preservation of reasoning output, and rejection of prefilled assistant tool calls. The UI also changes tool-policy and settings behavior, but the source provides no user adoption or performance data.

Στοιχεία πηγής: github.com ↗

Γιατί έχει σημασία

This is a substantial update to an open-source runtime used by developers building AI inference and multimodal applications. Its expanded model coverage can make more recently released models testable within llama.cpp-based systems, while video-input support broadens the kinds of media those systems can process. The source describes performance and memory-related work, but provides no independent benchmarks, so practical gains will depend on hardware, model formats, and deployment configuration.

The release connects several newer model architectures to a widely used inference codebase, including initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support. That may reduce integration work for developers experimenting with those models, although the source does not establish compatibility across every platform or configuration.

The ggml update adds infrastructure for sparse attention, asynchronous backends, remote procedure call events, and Apple RDMA. These changes could matter for deployments that can use those specific backends, but the release page does not quantify latency, throughput, memory use, or reliability improvements.

Video parameters, data-URL support for media, and multimodal preprocessing changes provide a more concrete path for applications handling video and other media. The source does not state whether these capabilities are available in packaged builds or under what model-specific limitations.

Interactive Mechanism

Διαδραστικός Μηχανισμός: Πώς λειτουργεί στην πραγματικότητα

Εξερευνήστε την υποκείμενη τεχνολογία πίσω από αυτήν την εξέλιξη διαδραστικά.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Διαδραστικός Έλεγχος Έννοιας+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Τι να παρακολουθήσετε στη συνέχεια

Follow-up releases, backend-specific tests, and documentation should clarify how the new model and video features behave across supported hardware. The most immediate uncertainty is whether the initial Qwen3.8-Flash-Next implementation receives the promised optimization work and how much the sparse-attention and memory changes improve real workloads.

Watch for optimization updates to Qwen3.8-Flash-Next and additional fixes for the newly added paths.

Watch for benchmark results that separate the release page’s implementation claims from measured gains in speed, memory use, and concurrency.

Watch for documentation on packaged access, supported model files, hardware requirements, and production-readiness. Those details are not supplied by the release page.

Σχετικοί οδηγοί και κουίζ

Επεξήγηση μοντέλων AIChatGPT και LLMΜετασχηματιστέςΔοκιμάστε τι γνωρίζετε — δοκιμάστε ένα δωρεάν κουίζ AIΑναζητήστε έναν όρο AI στο γλωσσάρι μαςΑκολουθήστε τον ιχνηλάτη έκδοσης μοντέλου AI

Ενημερώσεις και διορθώσεις

Αυτή η κανονική ιστορία ενημερώνεται όταν αλλάζει ουσιαστικά το αναπτυσσόμενο γεγονός. Το URL και η αρχική ημερομηνία δημοσίευσής του δεν αλλάζουν ποτέ.

  • The earlier llama.cpp entry covered pre-release support for NVIDIA’s Nemotron-3-Puzzle. Version 0.4.0 now includes Nemotron-3-Puzzle-75B-A9B support in a broader tagged release that also adds newer model architectures, video input, server context controls, lazy tensor reading, and ggml 0.23.0 backend work.
Δείτε το δημόσιο αρχείο διορθώσεων
Βρήκατε αυτό χρήσιμο;