Вернуться к новостям
ПродуктAI Understanding брифинг

В llama.cpp 0.4.0 добавлена поддержка новых моделей искусственного интеллекта и видеовхода.

В выпуске llama.cpp 0.4.0 добавлена первоначальная поддержка Qwen3.8-Flash-Next и NVIDIA Nemotron-3-Puzzle, параметры видеовхода, ограничения контекста сервера для каждого слота, отложенное чтение тензоров и изменения ggml 0.23.0.

4 min readRead the primary source
Source-page capture accompanying llama.cpp 0.4.0 adds support for newer AI models and video input
ПервоисточникИсточник записан
Издатель
github.com
Ссылка на источник
github.comhttps://github.com/ggml-org/llama.cpp/releases/tag/v0.4.0
Тип источника
Первичный документ — официальное объявление, документ, файл или собственная страница, которую мы читаем напрямую.
Также цитируется

Последняя редакция истории

КонтекстПоймите это за 60 секунд

Начните здесь

Ключевые термины

API (интерфейс прикладного программирования)
Структурированный способ отправки одной программной системой запросов и получения ответов от другой системы.
Память (Память агента)
Сохраненный контекст, который агент ИИ использует на этапах или сеансах для улучшения непрерывности.
Мультимодальная модель
Модель, которая может обрабатывать или генерировать несколько типов данных, таких как текст, изображение и аудио.
Проверьте себяВикторина с объяснением моделей искусственного интеллекта

Что изменилось с момента публикации

  1. Впервые опубликовано
  2. The earlier llama.cpp entry covered pre-release support for NVIDIA’s Nemotron-3-Puzzle. Version 0.4.0 now includes Nemotron-3-Puzzle-75B-A9B support in a broader tagged release that also adds newer model architectures, video input, server context controls, lazy tensor reading, and ggml 0.23.0 backend work.

Что случилось

llama.cpp released version 0.4.0 on GitHub on September 4. The release adds initial support for Qwen3.8-Flash-Next, NVIDIA Nemotron-3-Puzzle-75B-A9B, DSpark for Nemotron 3.5, nanbeige4.2-3B, and DeepSeek-V4-Flash-Vision-Exp. It also adds video-input parameters, per-slot server context limits, lazy tensor reading, expert-routing changes, KV-cache improvements, and multiple backend optimizations. The release updates ggml from 0.22.0 to 0.23.0. The project says that version adds sparse flash attention, asynchronous execution and allocation-dependency APIs, RPC event and async APIs, and Apple RDMA transport support. The page lists a nightly build identified as b10809. It does not document packaged-binary availability, installation requirements, model-weight distribution, or pricing.

The GitHub release page identifies v0.4.0 as the latest release and records a September 4 publication time of 19:56, without specifying a timezone in the supplied text. The release includes API changes such as llama_lazy_mode, quantizer buffer-size controls, updated session and state versions, and new multimodal tokenization helpers.

Model and core changes include initial Qwen3.8-Flash-Next architecture support, NVIDIA Nemotron-3-Puzzle-75B-A9B support, DSpark support for Nemotron 3.5, nanbeige4.2-3B support, lazy tensor reading, per-layer expert routing, n-gram history lookup, KV-cache restoration improvements, and safeguards against RAM peaks during model loading.

Multimodal and server changes include DeepSeek-V4-Flash-Vision-Exp support, video command-line options, per-slot context limits, data URLs for media, default preservation of reasoning output, and rejection of prefilled assistant tool calls. The UI also changes tool-policy and settings behavior, but the source provides no user adoption or performance data.

Подробности об источнике: github.com ↗

Почему это важно

This is a substantial update to an open-source runtime used by developers building AI inference and multimodal applications. Its expanded model coverage can make more recently released models testable within llama.cpp-based systems, while video-input support broadens the kinds of media those systems can process. The source describes performance and memory-related work, but provides no independent benchmarks, so practical gains will depend on hardware, model formats, and deployment configuration.

The release connects several newer model architectures to a widely used inference codebase, including initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support. That may reduce integration work for developers experimenting with those models, although the source does not establish compatibility across every platform or configuration.

The ggml update adds infrastructure for sparse attention, asynchronous backends, remote procedure call events, and Apple RDMA. These changes could matter for deployments that can use those specific backends, but the release page does not quantify latency, throughput, memory use, or reliability improvements.

Video parameters, data-URL support for media, and multimodal preprocessing changes provide a more concrete path for applications handling video and other media. The source does not state whether these capabilities are available in packaged builds or under what model-specific limitations.

Interactive Mechanism

Интерактивный механизм: как он на самом деле работает

Изучите технологию, лежащую в основе этой разработки, в интерактивном режиме.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Интерактивная проверка концепции+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Что посмотреть дальше

Follow-up releases, backend-specific tests, and documentation should clarify how the new model and video features behave across supported hardware. The most immediate uncertainty is whether the initial Qwen3.8-Flash-Next implementation receives the promised optimization work and how much the sparse-attention and memory changes improve real workloads.

Watch for optimization updates to Qwen3.8-Flash-Next and additional fixes for the newly added paths.

Watch for benchmark results that separate the release page’s implementation claims from measured gains in speed, memory use, and concurrency.

Watch for documentation on packaged access, supported model files, hardware requirements, and production-readiness. Those details are not supplied by the release page.

Сопутствующие руководства и викторины

Объяснение моделей искусственного интеллектаChatGPT и LLMТрансформерыПроверьте свои знания — пройдите бесплатную викторину по искусственному интеллектуНайдите термин ИИ в нашем глоссарии.Следите за трекером выпуска моделей AI

Обновления и исправления

Эта каноническая история обновляется при существенных изменениях развивающегося события. Его URL-адрес и первоначальная дата публикации никогда не меняются.

  • The earlier llama.cpp entry covered pre-release support for NVIDIA’s Nemotron-3-Puzzle. Version 0.4.0 now includes Nemotron-3-Puzzle-75B-A9B support in a broader tagged release that also adds newer model architectures, video input, server context controls, lazy tensor reading, and ggml 0.23.0 backend work.
Посмотреть журнал общедоступных исправлений
Нашли это полезным?