Powrót do Wiadomości
ProduktAI Understanding odprawa

llama.cpp 0.4.0 dodaje obsługę nowszych modeli AI i wejścia wideo

Wersja llama.cpp 0.4.0 dodaje wstępną obsługę Qwen3.8-Flash-Next i NVIDIA Nemotron-3-Puzzle, opcje wejścia wideo, limity kontekstu serwera na szczelinę, leniwy odczyt tensora i zmiany w ggml 0.23.0.

4 min readRead the primary source
Source-page capture accompanying llama.cpp 0.4.0 adds support for newer AI models and video input
Dokument źródłowyŹródło zapisane
Wydawca
github.com
Link źródłowy
github.comhttps://github.com/ggml-org/llama.cpp/releases/tag/v0.4.0
Typ źródła
Dokument podstawowy — oficjalne ogłoszenie, dokument, zgłoszenie lub strona własna, którą czytamy bezpośrednio.
Również cytowane

Historia ostatnio poprawiona

KontekstZrozum to w 60 sekund

Zacznij tutaj

Kluczowe terminy

API (interfejs programowania aplikacji)
Ustrukturyzowany sposób wysyłania żądań przez jeden system oprogramowania i otrzymywania odpowiedzi z innego systemu.
Pamięć (pamięć agenta)
Przechowywany kontekst, którego agent AI używa na różnych etapach lub sesjach, aby poprawić ciągłość.
Model multimodalny
Model, który może przetwarzać lub generować wiele typów danych, takich jak tekst, obraz i dźwięk.
Sprawdź sięQuiz objaśniający modele AI

Co się zmieniło od czasu publikacji

  1. Po raz pierwszy opublikowany
  2. The earlier llama.cpp entry covered pre-release support for NVIDIA’s Nemotron-3-Puzzle. Version 0.4.0 now includes Nemotron-3-Puzzle-75B-A9B support in a broader tagged release that also adds newer model architectures, video input, server context controls, lazy tensor reading, and ggml 0.23.0 backend work.

Co się stało

llama.cpp released version 0.4.0 on GitHub on September 4. The release adds initial support for Qwen3.8-Flash-Next, NVIDIA Nemotron-3-Puzzle-75B-A9B, DSpark for Nemotron 3.5, nanbeige4.2-3B, and DeepSeek-V4-Flash-Vision-Exp. It also adds video-input parameters, per-slot server context limits, lazy tensor reading, expert-routing changes, KV-cache improvements, and multiple backend optimizations. The release updates ggml from 0.22.0 to 0.23.0. The project says that version adds sparse flash attention, asynchronous execution and allocation-dependency APIs, RPC event and async APIs, and Apple RDMA transport support. The page lists a nightly build identified as b10809. It does not document packaged-binary availability, installation requirements, model-weight distribution, or pricing.

The GitHub release page identifies v0.4.0 as the latest release and records a September 4 publication time of 19:56, without specifying a timezone in the supplied text. The release includes API changes such as llama_lazy_mode, quantizer buffer-size controls, updated session and state versions, and new multimodal tokenization helpers.

Model and core changes include initial Qwen3.8-Flash-Next architecture support, NVIDIA Nemotron-3-Puzzle-75B-A9B support, DSpark support for Nemotron 3.5, nanbeige4.2-3B support, lazy tensor reading, per-layer expert routing, n-gram history lookup, KV-cache restoration improvements, and safeguards against RAM peaks during model loading.

Multimodal and server changes include DeepSeek-V4-Flash-Vision-Exp support, video command-line options, per-slot context limits, data URLs for media, default preservation of reasoning output, and rejection of prefilled assistant tool calls. The UI also changes tool-policy and settings behavior, but the source provides no user adoption or performance data.

Szczegóły źródła: github.com ↗

Dlaczego to ma znaczenie

This is a substantial update to an open-source runtime used by developers building AI inference and multimodal applications. Its expanded model coverage can make more recently released models testable within llama.cpp-based systems, while video-input support broadens the kinds of media those systems can process. The source describes performance and memory-related work, but provides no independent benchmarks, so practical gains will depend on hardware, model formats, and deployment configuration.

The release connects several newer model architectures to a widely used inference codebase, including initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support. That may reduce integration work for developers experimenting with those models, although the source does not establish compatibility across every platform or configuration.

The ggml update adds infrastructure for sparse attention, asynchronous backends, remote procedure call events, and Apple RDMA. These changes could matter for deployments that can use those specific backends, but the release page does not quantify latency, throughput, memory use, or reliability improvements.

Video parameters, data-URL support for media, and multimodal preprocessing changes provide a more concrete path for applications handling video and other media. The source does not state whether these capabilities are available in packaged builds or under what model-specific limitations.

Interactive Mechanism

Mechanizm interaktywny: jak to faktycznie działa

Poznaj interaktywnie technologię leżącą u podstaw tego rozwoju.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interaktywna kontrola koncepcji+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Co obejrzeć dalej

Follow-up releases, backend-specific tests, and documentation should clarify how the new model and video features behave across supported hardware. The most immediate uncertainty is whether the initial Qwen3.8-Flash-Next implementation receives the promised optimization work and how much the sparse-attention and memory changes improve real workloads.

Watch for optimization updates to Qwen3.8-Flash-Next and additional fixes for the newly added paths.

Watch for benchmark results that separate the release page’s implementation claims from measured gains in speed, memory use, and concurrency.

Watch for documentation on packaged access, supported model files, hardware requirements, and production-readiness. Those details are not supplied by the release page.

Powiązane przewodniki i quizy

Wyjaśnienie modeli AIChatGPT i LLMTransformatorySprawdź swoją wiedzę — wypróbuj darmowy quiz dotyczący sztucznej inteligencjiWyszukaj termin związany ze sztuczną inteligencją w naszym glosariuszuPostępuj zgodnie z modułem śledzenia wydań modeli AI

Aktualizacje i poprawki

Ta kanoniczna historia jest aktualizowana na miejscu, gdy rozwijające się wydarzenie ulegnie istotnej zmianie. Jego adres URL i pierwotna data publikacji nigdy się nie zmieniają.

  • The earlier llama.cpp entry covered pre-release support for NVIDIA’s Nemotron-3-Puzzle. Version 0.4.0 now includes Nemotron-3-Puzzle-75B-A9B support in a broader tagged release that also adds newer model architectures, video input, server context controls, lazy tensor reading, and ggml 0.23.0 backend work.
Zobacz publiczny dziennik poprawek
Uznałeś to za przydatne?