Voltar às notícias
ProdutoInstruções AI Understanding

llama.cpp 0.4.0 adiciona suporte para modelos de IA mais recentes e entrada de vídeo

A versão llama.cpp 0.4.0 adiciona suporte inicial para Qwen3.8-Flash-Next e NVIDIA Nemotron-3-Puzzle, opções de entrada de vídeo, limites de contexto de servidor por slot, leitura lenta de tensor e alterações ggml 0.23.0.

4 min readRead the primary source
Source-page capture accompanying llama.cpp 0.4.0 adds support for newer AI models and video input
Documento de origem primáriaFonte registrada
Editora
github.com
Link da fonte
github.comhttps://github.com/ggml-org/llama.cpp/releases/tag/v0.4.0
Tipo de fonte
Documento primário - um anúncio oficial, papel, arquivamento ou página original que lemos diretamente.
Também citado

História revisada pela última vez

ContextoEntenda isso em 60 segundos

Comece aqui

Termos-chave

API (Interface de Programação de Aplicativo)
Uma maneira estruturada de um sistema de software enviar solicitações e receber respostas de outro sistema.
Memória (memória do agente)
Contexto armazenado que um agente de IA usa em etapas ou sessões para melhorar a continuidade.
Modelo Multimodal
Um modelo que pode processar ou gerar vários tipos de dados, como texto, imagem e áudio.
Teste você mesmoQuestionário explicado sobre modelos de IA

O que mudou desde a publicação

  1. Publicado pela primeira vez
  2. The earlier llama.cpp entry covered pre-release support for NVIDIA’s Nemotron-3-Puzzle. Version 0.4.0 now includes Nemotron-3-Puzzle-75B-A9B support in a broader tagged release that also adds newer model architectures, video input, server context controls, lazy tensor reading, and ggml 0.23.0 backend work.

O que aconteceu

llama.cpp released version 0.4.0 on GitHub on September 4. The release adds initial support for Qwen3.8-Flash-Next, NVIDIA Nemotron-3-Puzzle-75B-A9B, DSpark for Nemotron 3.5, nanbeige4.2-3B, and DeepSeek-V4-Flash-Vision-Exp. It also adds video-input parameters, per-slot server context limits, lazy tensor reading, expert-routing changes, KV-cache improvements, and multiple backend optimizations. The release updates ggml from 0.22.0 to 0.23.0. The project says that version adds sparse flash attention, asynchronous execution and allocation-dependency APIs, RPC event and async APIs, and Apple RDMA transport support. The page lists a nightly build identified as b10809. It does not document packaged-binary availability, installation requirements, model-weight distribution, or pricing.

The GitHub release page identifies v0.4.0 as the latest release and records a September 4 publication time of 19:56, without specifying a timezone in the supplied text. The release includes API changes such as llama_lazy_mode, quantizer buffer-size controls, updated session and state versions, and new multimodal tokenization helpers.

Model and core changes include initial Qwen3.8-Flash-Next architecture support, NVIDIA Nemotron-3-Puzzle-75B-A9B support, DSpark support for Nemotron 3.5, nanbeige4.2-3B support, lazy tensor reading, per-layer expert routing, n-gram history lookup, KV-cache restoration improvements, and safeguards against RAM peaks during model loading.

Multimodal and server changes include DeepSeek-V4-Flash-Vision-Exp support, video command-line options, per-slot context limits, data URLs for media, default preservation of reasoning output, and rejection of prefilled assistant tool calls. The UI also changes tool-policy and settings behavior, but the source provides no user adoption or performance data.

Detalhes da fonte: github.com ↗

Por que isso importa

This is a substantial update to an open-source runtime used by developers building AI inference and multimodal applications. Its expanded model coverage can make more recently released models testable within llama.cpp-based systems, while video-input support broadens the kinds of media those systems can process. The source describes performance and memory-related work, but provides no independent benchmarks, so practical gains will depend on hardware, model formats, and deployment configuration.

The release connects several newer model architectures to a widely used inference codebase, including initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support. That may reduce integration work for developers experimenting with those models, although the source does not establish compatibility across every platform or configuration.

The ggml update adds infrastructure for sparse attention, asynchronous backends, remote procedure call events, and Apple RDMA. These changes could matter for deployments that can use those specific backends, but the release page does not quantify latency, throughput, memory use, or reliability improvements.

Video parameters, data-URL support for media, and multimodal preprocessing changes provide a more concrete path for applications handling video and other media. The source does not state whether these capabilities are available in packaged builds or under what model-specific limitations.

Interactive Mechanism

Mecanismo interativo: como realmente funciona

Explore a tecnologia subjacente a este desenvolvimento de forma interativa.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Verificação de conceito interativo+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

O que assistir a seguir

Follow-up releases, backend-specific tests, and documentation should clarify how the new model and video features behave across supported hardware. The most immediate uncertainty is whether the initial Qwen3.8-Flash-Next implementation receives the promised optimization work and how much the sparse-attention and memory changes improve real workloads.

Watch for optimization updates to Qwen3.8-Flash-Next and additional fixes for the newly added paths.

Watch for benchmark results that separate the release page’s implementation claims from measured gains in speed, memory use, and concurrency.

Watch for documentation on packaged access, supported model files, hardware requirements, and production-readiness. Those details are not supplied by the release page.

Guias e questionários relacionados

Modelos de IA explicadosChatGPT e LLMTransformadoresTeste o que você sabe – experimente um teste gratuito de IAProcure um termo de IA em nosso glossárioSiga o rastreador de lançamento de modelo de IA

Atualizações e correções

Esta história canônica é atualizada quando o evento em desenvolvimento muda materialmente. Seu URL e a data de publicação original nunca mudam.

  • The earlier llama.cpp entry covered pre-release support for NVIDIA’s Nemotron-3-Puzzle. Version 0.4.0 now includes Nemotron-3-Puzzle-75B-A9B support in a broader tagged release that also adds newer model architectures, video input, server context controls, lazy tensor reading, and ggml 0.23.0 backend work.
Veja o registro de correções públicas
Achou isso útil?