Subira ku makuru
IbicuruzwaAI Understanding ibisobanuro

llama.cpp 0.4.0 yongeraho inkunga kubintu bishya bya AI no kwinjiza amashusho

Isohora rya llama.cpp 0.4.0 ryongeramo inkunga yambere kuri Qwen3.8-Flash-Ibikurikira na NVIDIA Nemotron-3-Puzzle, amashusho-yinjiza, uburyo bwa seriveri igarukira, imipaka ya tensor, kandi gusoma ggml 0.23.0.

4 min readRead the primary source
Source-page capture accompanying llama.cpp 0.4.0 adds support for newer AI models and video input
Inyandiko y'ibanzeInkomoko yanditse
Umwanditsi
github.com
Ihuza ry'inkomoko
github.comhttps://github.com/ggml-org/llama.cpp/releases/tag/v0.4.0
Ubwoko bw'inkomoko
Inyandiko y'ibanze - itangazo ryemewe, impapuro, dosiye, cyangwa urupapuro rwambere-dusoma mu buryo butaziguye.
Byatanzwe kandi

Inkuru iheruka gusubirwamo

ImirongoSobanukirwa ibi mumasegonda 60

Tangira hano

Amagambo y'ingenzi

API (Imigaragarire ya Porogaramu)
Inzira yuburyo bwa sisitemu imwe yohereza ibyifuzo no kwakira ibisubizo bivuye murindi sisitemu.
Kwibuka (Memory Memory)
Imiterere yabitswe umukozi wa AI akoresha intambwe cyangwa amasomo kugirango atezimbere.
Icyitegererezo
Icyitegererezo gishobora gutunganya cyangwa kubyara amakuru menshi nkinyandiko, ishusho, n'amajwi.
IsuzumeModeri ya AI Yasobanuwe Ikibazo

Niki cyahindutse kuva cyatangazwa

  1. Byatangajwe bwa mbere
  2. The earlier llama.cpp entry covered pre-release support for NVIDIA’s Nemotron-3-Puzzle. Version 0.4.0 now includes Nemotron-3-Puzzle-75B-A9B support in a broader tagged release that also adds newer model architectures, video input, server context controls, lazy tensor reading, and ggml 0.23.0 backend work.

Byagenze bite

llama.cpp released version 0.4.0 on GitHub on September 4. The release adds initial support for Qwen3.8-Flash-Next, NVIDIA Nemotron-3-Puzzle-75B-A9B, DSpark for Nemotron 3.5, nanbeige4.2-3B, and DeepSeek-V4-Flash-Vision-Exp. It also adds video-input parameters, per-slot server context limits, lazy tensor reading, expert-routing changes, KV-cache improvements, and multiple backend optimizations. The release updates ggml from 0.22.0 to 0.23.0. The project says that version adds sparse flash attention, asynchronous execution and allocation-dependency APIs, RPC event and async APIs, and Apple RDMA transport support. The page lists a nightly build identified as b10809. It does not document packaged-binary availability, installation requirements, model-weight distribution, or pricing.

The GitHub release page identifies v0.4.0 as the latest release and records a September 4 publication time of 19:56, without specifying a timezone in the supplied text. The release includes API changes such as llama_lazy_mode, quantizer buffer-size controls, updated session and state versions, and new multimodal tokenization helpers.

Model and core changes include initial Qwen3.8-Flash-Next architecture support, NVIDIA Nemotron-3-Puzzle-75B-A9B support, DSpark support for Nemotron 3.5, nanbeige4.2-3B support, lazy tensor reading, per-layer expert routing, n-gram history lookup, KV-cache restoration improvements, and safeguards against RAM peaks during model loading.

Multimodal and server changes include DeepSeek-V4-Flash-Vision-Exp support, video command-line options, per-slot context limits, data URLs for media, default preservation of reasoning output, and rejection of prefilled assistant tool calls. The UI also changes tool-policy and settings behavior, but the source provides no user adoption or performance data.

Ibisobanuro birambuye: github.com ↗

Impamvu ari ngombwa

This is a substantial update to an open-source runtime used by developers building AI inference and multimodal applications. Its expanded model coverage can make more recently released models testable within llama.cpp-based systems, while video-input support broadens the kinds of media those systems can process. The source describes performance and memory-related work, but provides no independent benchmarks, so practical gains will depend on hardware, model formats, and deployment configuration.

The release connects several newer model architectures to a widely used inference codebase, including initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support. That may reduce integration work for developers experimenting with those models, although the source does not establish compatibility across every platform or configuration.

The ggml update adds infrastructure for sparse attention, asynchronous backends, remote procedure call events, and Apple RDMA. These changes could matter for deployments that can use those specific backends, but the release page does not quantify latency, throughput, memory use, or reliability improvements.

Video parameters, data-URL support for media, and multimodal preprocessing changes provide a more concrete path for applications handling video and other media. The source does not state whether these capabilities are available in packaged builds or under what model-specific limitations.

Interactive Mechanism

Uburyo bukoreshwa: Uburyo bukora

Shakisha ikoranabuhanga ryihishe inyuma yiri terambere.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Kugenzura Ibitekerezo Byagenzuwe+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Ibyo kureba

Follow-up releases, backend-specific tests, and documentation should clarify how the new model and video features behave across supported hardware. The most immediate uncertainty is whether the initial Qwen3.8-Flash-Next implementation receives the promised optimization work and how much the sparse-attention and memory changes improve real workloads.

Watch for optimization updates to Qwen3.8-Flash-Next and additional fixes for the newly added paths.

Watch for benchmark results that separate the release page’s implementation claims from measured gains in speed, memory use, and concurrency.

Watch for documentation on packaged access, supported model files, hardware requirements, and production-readiness. Those details are not supplied by the release page.

Ibijyanye nuyobora & ibibazo

Moderi ya AI YasobanuweChatGPT na LLMsAbahinduraGerageza ibyo uzi - gerageza ikibazo cya AI kubuntuReba ijambo AI mumagambo yacuKurikiza icyerekezo cya AI cyo kurekura

Kuvugurura no gukosora

Iyi nkuru yemewe ivugururwa mugihe iyo iterambere ryiterambere rihindutse mubintu. URL yayo nitariki yo gusohora itariki ntizigera ihinduka.

  • The earlier llama.cpp entry covered pre-release support for NVIDIA’s Nemotron-3-Puzzle. Version 0.4.0 now includes Nemotron-3-Puzzle-75B-A9B support in a broader tagged release that also adds newer model architectures, video input, server context controls, lazy tensor reading, and ggml 0.23.0 backend work.
Reba igitabo gikosora rusange
Basanze ari ingirakamaro?