Dellu ci xibaar yi
ProduitAI Understanding

llama.cpp 0.4.0 yokk na jàppale ci xeeti IA yu bees ak dugal wideo

llama.cpp 0.4.0 génne yokk na ndimmbal bu njëkk ngir Qwen3.8-Flash-Next ak NVIDIA Nemotron-3-Puzzle, tànneefi dugal wideo, àppu serwër bu nekk, jàngu tensor bu tayeel, ak coppite ggml 0.23.0.

4 min readRead the primary source
Source-page capture accompanying llama.cpp 0.4.0 adds support for newer AI models and video input
Këyitu xët bu njëkkSource biñ enregistre
Siiwalkat
github.com
Lëkkalekaayu cosaan
github.comhttps://github.com/ggml-org/llama.cpp/releases/tag/v0.4.0
Xeetu balluwaay
Këyitu njëkk - ab yëgle ofisel, këyit, dosiye, wala xëtu pàrti bu njëkk bi ñuy jàng ci saasi.
ñu wax itam

Dañu mujjee soppali jaar-jaar bi

KontekstXam lii ci 60 seconde

Tambalil fii

Term yu am solo

API (interfaasu prograam aplikaasioŋ)
Benn anam buñ tëral ngir benn sistem losisel mëna yónnee ay laaj ak jot tontu ci beneen sistem.
Memoire (Memoire agent)
Kontekst buñ denc bi ab ndawu IA di jëfandikoo ci jéego yi wala sesioŋ yi ngir gëna mëna wéy.
Modèle multimodal
Modèle bu mëna defar wala defar ay xeeti done yu bari lu ci melni bind, nataal wala audio.
Nattal sa boppModèlu IA leeral quiz

Lu soppeeku ginaaw biñu ko siiwalee

  1. Ñu ngi ko njëkka siiwal
  2. Duggug llama.cpp bi njëkk dafa muur ndimmbalu génne bu njëkk ngir Nemotron-3-Puzzle bu NVIDIA. Version 0.4.0 leegi amna ndimmbalu Nemotron-3-Puzzle-75B-A9B ci benn génne bu gëna yaatu buñ etikete, loolu dafay yokk yeneen xeetu architecture yu bees, dugal wideo, seytu serwër, jàngu tensor bu taye, ak liggéeyu backend ggml 0.23.0.

Lu xew

llama.cpp genne na 0.4.0 ci GitHub ci 4 septembre. DeepSeek-V4-Flash-Vision-Exp. Dafay yokk itam ay parametru dugal wideo, ay yamaleg ëmbiitu serwër bu nekk, jàngu tensor bu taye, coppite ci yoonu kàngam, yokkute ci cache KV, ak gëna xéewale backend yu bari. ggml bi dafay yeesal 0.22.0 ba 0.23.0. Projet bi dafa wax ni versioŋ bi dafay yokk ay flash yu néew, jëfandikoo asynchrone ak API yu aju ci joxe, xew-xew RPC ak API async, ak ndimbalu dem ak dikk Apple RDMA. Xët wi dafay lim tabax guddi gu nekk buñu xamme ni b10809. Du bind disponibilite binaire buñ defar, li ñuy laaj ci instalaasioŋ bi, séddaleb poid model bi, wala njëg yi.

Xëtu génne bu GitHub dafay wane v0.4.0 ni mooy génne bu mujj ba noppi bind waxtu génne bu 4 septembre ci 19:56, te leeralul zone zone ci bind biñ joxe. Li ñuy génne dafa amaale coppite ci API yu melni llama_lazy_mode, seytu dayo tampon kwantizer, sesioŋ yu bees ak xeetu etaa yi, ak ndimbalu tokenization multimodal yu bees.

Coppite yiñ amal ci model bi bokkuna ci ndimbalu architecture Qwen3.8-Flash-Next, ndimbalu Nemotron-3-75B-A9B, ndimbalu DSpark ngir Nemotron 3.5, ndimmbalu nanbeige4.2-3B, jàngat taarixu tensor bu taye, xool bu nekk yokkute, ak kaaraange ci RAM peaks ci diiru model loading.

Coppite yi am ci anam yu bari ak ci serwër yi bokkuna ci jàppale DeepSeek-V4-Flash-Vision-Exp, tànneefi liiñ komand wideo, àppu ëmbiit bu nekk, URL done ngir media yi, sàmm bu njëkk ci génnug xalaat, ak bañ woote jumtukaayi ndimbal yu njëkk. UI bi itam dafay soppi jumtukaay-politigu ak jekkal jekkal, waaye source bi du joxe benn done jëfandikukat wala performance.

Ay leeral ci cosaan: github.com ↗

Lu tax mu am solo

Lii ab yeesal bu am solo la ci runtime open-source bi developpeur yi di jëfandikoo ngir tabax IA inference ak aplikaasioŋu multimodal. Modèle bimu yaatal mën na tax model yu bees yi ñu mëna natt ci sistem yu llama.cpp, ci noonu la jàppale wideo-input di yaatal xeetu media yi sistem yooyu mëna jëfandikoo. Source bi dafay leeral performance ak liggéey bi jëm ci mémoire, waaye du joxe benn benchmark bu moom boppam, kon njariñ yi ci jëfandikoo dañuy aju ci hardware bi, formaa model yi, ak configuration deployment bi.

Li ñuy génne dafay boole yeneen xeeti tabax yu bees ak benn kodu jëfandikoo bu bari, lu ci melni Qwen3.8-Flash-Next ak ndimbalu Nemotron-3-Puzzle. Loolu mën na wàññi liggéeyu boole bi ci developpeur yiy jàngat model yooyu, ndigam source bi du taxawal deggoo ci bépp platform wala configuration.

Coppite ggml bi dafa yokk jumtukaay ngir bàyyi xel ci lu néew, backend yu asynchrone, xew-xewu wooteb doxalin bu sori, ak Apple RDMA. Coppite yooyu mën nañu am solo ci jëfandikoo yi mëna jëfandikoo backend yooyu, waaye xëtu génne gi du xayma latency, throughput, jëfandikoo mémoire, wala yokkute ci wóor.

Parametru wideo, ndimbalu URL done ngir media yi, ak coppite yi am ci preprocessing multimodal dañuy joxe yoon wu gëna fëgër ngir aplikaasioŋ yiy jëfandikoo wideo ak yeneen media. Source bi waxul ndax mën-mën yooyu am nañu ci tabax yuñ defar wala ci ban model-specific limitations.

Interactive Mechanism

Mekanism buy weccoo xalaat: naka lay doxee

Saytu xarala yu bees yi ci ginaaw yokkute bii ci anam wu weccoo xalaat.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Saytu konsept buy weccoo xalaat+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Li nga wara seetaan ci topp

Li ñuy topp ci génne yi, test yiñ jagleel backend bi, ak këyitu liggéey yi dañu wara leeral ni model bu bees bi ak man-mani wideo yi di doxee ci hardware biñ jàppale. Li gëna wóorul ci saasi mooy ndax jëfandikoo gi njëkk ci Qwen3.8-Flash-Next jotna liggéeyu optimisation biñ digoon, ak ni sparse-attention ak coppite ci mémoire di gëna baaxal liggéey bi.

Seetal yeesali gëna xéewale ci Qwen3.8-Flash-Next ak yeneen defar ngir yooni model multimodal yu bees yiñ yokk.

Seetal njariñu référence yiy tàqale li ñuy wax ci jëfandikoo xëtu génne gi ak njariñ yiñ natt ci gaawaay, jëfandikoo mémoire, ak concurrence.

Saytu dokimaa ci wàllu jëfandikoo gi, fichier model yiñ jàppale, li ñuy laaj ci hardware bi, ak waajal liggéey bi. Xëtu génne gi joxeewul leeral yooyu.

Gid ak quiz yu ci méngoo

Model IA leeral nañu koChatGPT & LLMsTransformatërNatt li nga xam — natt quiz IA bu amul faydaSeetal benn baat IA ci sunu glossaireToppal toppukaayu génne xeetu IA

Yeesal ak jubbanti

Jaar-jaar canonical bii dañu koy yeesal ci barab bi su xew-xew bi di màgg soppeekoo ci anam wu amul benn werante. URL bi ak bisu siiwal bi duñu musa soppeeku.

  • Duggug llama.cpp bi njëkk dafa muur ndimmbalu génne bu njëkk ngir Nemotron-3-Puzzle bu NVIDIA. Version 0.4.0 leegi amna ndimmbalu Nemotron-3-Puzzle-75B-A9B ci benn génne bu gëna yaatu buñ etikete, loolu dafay yokk yeneen xeetu architecture yu bees, dugal wideo, seytu serwër, jàngu tensor bu taye, ak liggéeyu backend ggml 0.23.0.
Xoolal jubluwaayu njuumte yi ñépp bokk
Gis nga lii am njariñ?