返回新闻
产品展示AI Understanding 简报

Mistral Large 4预览版人工智能分析智能指数得分38

人工分析基准测试为 Mistral Large 4 预览版提供了 38 分的智能指数分数,使其跻身非美国/中国得分最高的型号之列,并突出了其速度、代币成本和多模式功能。

4 min readRead the linked source
Source-provided image accompanying Mistral Large 4 preview scores 38 on Artificial Analysis Intelligence Index
来源参考来源记录
出版商
unite.ai
来源类型
链接来源——主要来源状态尚未确定。
背景60 秒内了解这一点

关键术语

API(应用程序编程接口)
一种软件系统向另一个系统发送请求并接收响应的结构化方式。
大语言模型(LLM)
在海量文本语料库上训练来生成和分析文本的语言模型。
上下文窗口
语言模型一次可以处理的输入标记的最大数量。
测试一下自己AI 模型解释测验

发生了什么

Artificial Analysis released a benchmark on October 6, 2026 evaluating Mistral Large 4 preview (nicknamed Le Chonk). The model achieved a 38 score on the firm’s Intelligence Index version 4.3.2, matching GPT‑6 Luna (max) and just below DeepSeek V4.1 Flash (max) at 39. In the ranking of 225 models, Mistral Large 4 sits at 64, well above the class median of 26. The Cyber Index score is 50, equal to GLM‑5.3‑Flash and ahead of several regional rivals. Its strongest cyber benchmark, CyberGym‑E2E‑AA, recorded an 82 % success rate, surpassing MiMo‑V2.6‑Pro (79 %) and GPT‑6 Luna (max) (78 %). The analysis also reports operational metrics: token‑generation speed of 116.1 tokens / s (median 86.1) and a first‑token latency of 1.46 s (median 3.81). Over the Intelligence Index run the model produced 200 million output tokens, more than double the class median of 81 million. Cost calculations show $1.13 per task at standard pricing, reduced to $0.57 during a 50 % launch discount, still higher than comparable open‑weight models (GLM‑5.3‑Flash $0.25, DeepSeek V4.1 Flash $0.27). The model’s is 512 k tokens, with a trillion‑parameter architecture and 49 billion active parameters. Availability is limited to a research public preview via Mistral’s API; open weights are slated for release by the end of October 2026.

Artificial Analysis’s Intelligence Index combines ten sub‑evaluations (AA‑Briefcase, GDPval‑AA, AutomationBench‑AA, Terminal‑Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA‑Omniscience, AA‑LCR v1.1) weighted across agents (30 %), coding (20 %), scientific reasoning (20 %), and general (30 %). Mistral Large 4’s 38 score reflects balanced strength, particularly in agentic reasoning and general tasks.

The Cyber Index, a separate suite measuring security‑related reasoning, placed the model at 50, tying with GLM‑5.3‑Flash and ahead of several regional models. Its top result on CyberGym‑E2E‑AA (82 %) demonstrates notable competence in end‑to‑end cyber‑scenario reasoning.

Operational metrics were gathered via Mistral’s API during the benchmark run. The model generated 200 million output tokens, a token‑generation speed of 116.1 tokens / s, and a first‑token latency of 1.46 seconds, indicating efficient inference compared with the class median.

Cost analysis used Mistral’s published token pricing ($1.36 / M input tokens, $4.18 / M output tokens) and a 50 % launch discount for the first two weeks. Even with the discount, the per‑task cost ($0.57) remains higher than comparable open‑weight models, highlighting a pricing gap that could affect early‑stage adoption.

来源详情: unite.ai ↗

为什么这很重要

The benchmark positions Mistral Large 4 as the most intelligent open‑weight‑eligible model outside the United States and China, a claim that could shift competitive dynamics in the global LLM market. A 38 Intelligence Index score signals strong performance across agents, coding, scientific reasoning and general tasks, suggesting the model can handle a wide range of enterprise and research workloads. However, the higher per‑task cost—more than four times that of comparable open‑weight models—may limit early adoption unless pricing drops after the preview period. The speed advantage and large also make the model attractive for applications requiring long‑form reasoning or real‑time interaction. If the promised open‑weight release lives up to the benchmark results, developers worldwide could gain a high‑performing, Europe‑trained alternative to US‑dominated models, potentially diversifying the AI ecosystem and reducing reliance on a few dominant providers.

The score validates Mistral’s claim of delivering a world‑class, Europe‑trained LLM that can compete with top US and Chinese offerings, potentially encouraging more sovereign AI development in the region.

Higher per‑task costs may deter cost‑sensitive developers, especially startups and academic groups, unless the performance gains translate into tangible productivity improvements.

Speed and large are critical for applications such as long‑form document analysis, code generation, and interactive agents, giving Mistral a competitive edge in latency‑sensitive use cases.

If the open‑weight release matches the preview’s benchmark performance, it could broaden access to a high‑intelligence model without the licensing constraints of proprietary APIs, influencing the open‑weight ecosystem.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下来看什么

Key signals to monitor include the official release of Mistral Large 4’s weights at the end of October 2026, subsequent benchmark updates from independent labs, and any pricing adjustments after the launch discount expires. Adoption trends among European research institutions and enterprises will indicate whether the cost premium is justified. Additionally, follow‑up evaluations on multilingual, speech and image tasks—currently benchmarked separately—will clarify the model’s true multimodal capabilities. Finally, competitor responses, especially from Chinese and US open‑weight projects, could reshape pricing and performance races in the coming months.

Weight release timeline: confirmation of open‑weight availability by end‑October 2026 will determine how quickly the broader community can test and adopt the model.

Pricing evolution: monitoring whether Mistral adjusts token rates after the launch discount period ends, and how that aligns with competitor pricing.

Multimodal benchmarks: upcoming results on speech and multilingual tasks will reveal whether the model’s claimed 100‑image‑per‑request capability translates into superior performance.

Competitive response: tracking benchmark updates from US and Chinese open‑weight models (e.g., GPT‑6, DeepSeek) to see if they close the intelligence gap.

相关指南和测验

觉得这有用吗?