What happened
Artificial Analysis released a benchmark on October 6, 2026 evaluating Mistral Large 4 preview (nicknamed Le Chonk). The model achieved a 38 score on the firm’s Intelligence Index version 4.3.2, matching GPT‑6 Luna (max) and just below DeepSeek V4.1 Flash (max) at 39. In the ranking of 225 models, Mistral Large 4 sits at 64, well above the class median of 26. The Cyber Index score is 50, equal to GLM‑5.3‑Flash and ahead of several regional rivals. Its strongest cyber benchmark, CyberGym‑E2E‑AA, recorded an 82 % success rate, surpassing MiMo‑V2.6‑Pro (79 %) and GPT‑6 Luna (max) (78 %). The analysis also reports operational metrics: token‑generation speed of 116.1 tokens / s (median 86.1) and a first‑token latency of 1.46 s (median 3.81). Over the Intelligence Index run the model produced 200 million output tokens, more than double the class median of 81 million. Cost calculations show $1.13 per task at standard pricing, reduced to $0.57 during a 50 % launch discount, still higher than comparable open‑weight models (GLM‑5.3‑Flash $0.25, DeepSeek V4.1 Flash $0.27). The model’s is 512 k tokens, with a trillion‑parameter architecture and 49 billion active parameters. Availability is limited to a research public preview via Mistral’s API; open weights are slated for release by the end of October 2026.
Artificial Analysis’s Intelligence Index combines ten sub‑evaluations (AA‑Briefcase, GDPval‑AA, AutomationBench‑AA, Terminal‑Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA‑Omniscience, AA‑LCR v1.1) weighted across agents (30 %), coding (20 %), scientific reasoning (20 %), and general (30 %). Mistral Large 4’s 38 score reflects balanced strength, particularly in agentic reasoning and general tasks.
The Cyber Index, a separate suite measuring security‑related reasoning, placed the model at 50, tying with GLM‑5.3‑Flash and ahead of several regional models. Its top result on CyberGym‑E2E‑AA (82 %) demonstrates notable competence in end‑to‑end cyber‑scenario reasoning.
Operational metrics were gathered via Mistral’s API during the benchmark run. The model generated 200 million output tokens, a token‑generation speed of 116.1 tokens / s, and a first‑token latency of 1.46 seconds, indicating efficient inference compared with the class median.
Cost analysis used Mistral’s published token pricing ($1.36 / M input tokens, $4.18 / M output tokens) and a 50 % launch discount for the first two weeks. Even with the discount, the per‑task cost ($0.57) remains higher than comparable open‑weight models, highlighting a pricing gap that could affect early‑stage adoption.
Why it matters
The benchmark positions Mistral Large 4 as the most intelligent open‑weight‑eligible model outside the United States and China, a claim that could shift competitive dynamics in the global LLM market. A 38 Intelligence Index score signals strong performance across agents, coding, scientific reasoning and general tasks, suggesting the model can handle a wide range of enterprise and research workloads. However, the higher per‑task cost—more than four times that of comparable open‑weight models—may limit early adoption unless pricing drops after the preview period. The speed advantage and large also make the model attractive for applications requiring long‑form reasoning or real‑time interaction. If the promised open‑weight release lives up to the benchmark results, developers worldwide could gain a high‑performing, Europe‑trained alternative to US‑dominated models, potentially diversifying the AI ecosystem and reducing reliance on a few dominant providers.
The score validates Mistral’s claim of delivering a world‑class, Europe‑trained LLM that can compete with top US and Chinese offerings, potentially encouraging more sovereign AI development in the region.
Higher per‑task costs may deter cost‑sensitive developers, especially startups and academic groups, unless the performance gains translate into tangible productivity improvements.
Speed and large are critical for applications such as long‑form document analysis, code generation, and interactive agents, giving Mistral a competitive edge in latency‑sensitive use cases.
If the open‑weight release matches the preview’s benchmark performance, it could broaden access to a high‑intelligence model without the licensing constraints of proprietary APIs, influencing the open‑weight ecosystem.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
Which component of an AI application is the machine-learning model itself?
What to watch next
Key signals to monitor include the official release of Mistral Large 4’s weights at the end of October 2026, subsequent benchmark updates from independent labs, and any pricing adjustments after the launch discount expires. Adoption trends among European research institutions and enterprises will indicate whether the cost premium is justified. Additionally, follow‑up evaluations on multilingual, speech and image tasks—currently benchmarked separately—will clarify the model’s true multimodal capabilities. Finally, competitor responses, especially from Chinese and US open‑weight projects, could reshape pricing and performance races in the coming months.
Weight release timeline: confirmation of open‑weight availability by end‑October 2026 will determine how quickly the broader community can test and adopt the model.
Pricing evolution: monitoring whether Mistral adjusts token rates after the launch discount period ends, and how that aligns with competitor pricing.
Multimodal benchmarks: upcoming results on speech and multilingual tasks will reveal whether the model’s claimed 100‑image‑per‑request capability translates into superior performance.
Competitive response: tracking benchmark updates from US and Chinese open‑weight models (e.g., GPT‑6, DeepSeek) to see if they close the intelligence gap.