What happened
Mistral announced the public‑preview API for its Large 4 model (Le Chonk) on Oct 6 2026 and detailed specifications, benchmark scores, and token pricing. The model is a 1.05 trillion‑ mixture‑of‑experts LLM with 49 billion active parameters, a 1 million‑token context window (though independent evaluators see 512 K–524 K), and a 1.6 billion‑parameter vision encoder. Pricing shown on the model card is $0.68 per M input tokens, $0.07 per M cached‑input tokens, and $2.09 per M output tokens, with crossed‑out higher rates indicating a 50 % preview discount. Independent benchmarks place Large 4 at 48.05 % on Vals’ broad industry index (32nd of 44) but higher on specialist tests: 15.83 % on the Harvey Legal Agent Benchmark and a modest edge on finance‑agent workloads. The open‑weight checkpoint is promised for the end of October, but no download link or license terms were available at the time of writing.
On Oct 6 2026 Mistral published a model card and API documentation for Large 4 (Le Chonk). The card lists 1.05 trillion total parameters, 49 billion active parameters (≈4.7 % activation per token), a 1.6 billion‑ vision encoder, and a 1 million‑token context window. The model uses a mixture‑of‑experts architecture that routes each token through a subset of experts rather than the full parameter set.
Pricing shown on the model card is $0.68 per M input tokens, $0.07 per M cached‑input tokens, and $2.09 per M output tokens, with crossed‑out rates of $1.36, $0.14, and $4.18 indicating a 50 % preview discount. The rates differ from those used by early independent evaluators, who applied the higher prices in their cost calculations.
Independent benchmark providers report mixed results. Vals’ broad industry index scores Large 4 at 48.05 % (32nd of 44), while the Harvey Legal Agent Benchmark gives it 15.83 %—ahead of several larger models on that specific task. Finance‑agent tests show a slight advantage over Astra, but latency is more than double. Vision tests place it at 42 % on Dense200, a one‑point lead over GPT‑6 Astra.
Mistral promises to release the model weights and a final license by the end of October 2026. At the time of the article, no download link or license details were available, meaning private deployment remains speculative.
Why it matters
Mistral Large 4 is one of the first trillion‑ mixture‑of‑experts models made available via a public API, offering a potential cost advantage for workloads that can exploit its specialist strengths in legal and financial document processing. The aggressive preview pricing could make high‑context, agentic applications financially viable for enterprises that need to keep data within the EU, as Mistral provides regional inference routes. However, the model’s open‑weight release is still pending, meaning private‑deployment or fine‑tuning plans must wait. The discrepancy between the advertised 1 M token context and evaluator‑reported limits also introduces uncertainty for developers planning long‑document use cases. Finally, benchmark results show that while Large 4 can outperform some frontier models on niche tasks, it lags behind leading generalist LLMs on broader indices, underscoring the need for workload‑specific testing before committing to it as a universal model.
The mixture‑of‑experts design allows a trillion‑ model to be served at a lower compute cost per token than a dense model of comparable size, potentially reducing inference expenses for high‑throughput workloads.
Specialist benchmark strengths suggest Large 4 could be a cost‑effective choice for legal document analysis and finance research, where the model’s reasoning depth and long‑context handling are valuable.
The preview pricing, if sustained, offers a significant token‑cost advantage over competing frontier models such as GPT‑6 Astra or Claude Opus, which could shift procurement decisions for budget‑constrained enterprises.
The pending open‑weight release is crucial for organizations that require on‑premise deployment for data‑sovereignty or custom fine‑tuning. Until the weights are available, users are limited to Mistral’s hosted API, which may not meet all regulatory or security requirements.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
Which component of an AI application is the machine-learning model itself?
What to watch next
Key follow‑ups include the official release of the downloadable weights and license terms at the end of October, confirmation of the true context‑window size across API endpoints, and independent replication of the specialist benchmark scores. Observers should also track any changes to the preview token pricing, especially whether the 50 % discount persists after the preview period. Finally, performance and cost data from real‑world deployments in legal and finance teams will clarify whether the model’s specialist edge translates into measurable productivity gains.
Release of downloadable weights and the accompanying license terms (expected end‑October 2026).
Verification of the actual context‑window size across API endpoints—whether the advertised 1 M tokens is fully supported or limited to ~512 K tokens.
Potential adjustments to the preview token pricing after the discount period ends.
Independent replication of legal and finance benchmark scores on real‑world corpora, and measurement of latency and throughput in production settings.
Adoption signals from EU‑based enterprises that prioritize regional inference routes and data residency.