返回新闻
创新AI Understanding 简报

麻省理工学院人工智能系统 Ataraxos 击败顶级人类 Stratego 玩家

麻省理工学院和合作大学的研究人员推出了 Ataraxos,这是一种人工智能,它的表现远远超过了世界上最好的 Stratego 玩家,同时使用的训练资源却少得多。

4 min readRead the primary source
Source-provided image accompanying MIT AI system Ataraxos beats top human Stratego players
主要来源文件来源记录
出版商
news.mit.edu
来源链接
news.mit.eduhttps://news.mit.edu/2026/game-playing-ai-stratego-new-champ-0930
来源类型
主要文件——我们直接阅读的官方公告、文件、文件或第一方页面。
背景60 秒内了解这一点

从这里开始

关键术语

强化学习
通过奖励信号进行训练,代理学习能够最大化长期回报的行动。
基准测试
用于测量和比较模型性能的标准化测试或数据集。
推理
经过训练的模型生成预测或输出的运行时阶段。
测试一下自己什么是人工智能?测验

发生了什么

MIT, Carnegie Mellon, NYU, and Stanford researchers announced a new AI system called Ataraxos that achieved superhuman performance in the hidden‑information board game Stratego. Using self‑play combined with decision‑time planning and a generative model to infer hidden piece identities, Ataraxos defeated the strongest human player 15‑1‑4 and posted a 39‑2 record at the Stratego world championship. The system also topped human experts in several related imperfect‑information games, including Barrage Stratego, Hanabi, and Dou dizhu. Crucially, Ataraxos reached this level while using less than one‑hundredth of the training examples and one‑thirtieth of the self‑play games required by DeepMind’s prior DeepNash system, indicating a dramatic efficiency gain.

The team trained Ataraxos via self‑play , allowing the AI to generate a "blueprint strategy" for Stratego without human supervision. Efficient training algorithms reduced the number of required self‑play games dramatically compared with earlier systems.

During play, Ataraxos employs decision‑time planning: a generative model predicts the likely identities of hidden opponent pieces, then evaluates possible moves under those hypotheses before selecting an action. This approach replaces exhaustive enumeration with probabilistic , enabling rapid, informed decisions.

Performance metrics show Ataraxos beating the world’s top Stratego player with a 15‑1‑4 win‑loss‑draw record and achieving a 39‑2 record against other elite competitors. The system also surpassed human experts in Barrage Stratego, Hanabi, and Dou dizhu, confirming the method’s generality across imperfect‑information games.

Compared with DeepMind’s DeepNash, Ataraxos required less than 1 % of the training examples and roughly 3 % of the self‑play games, translating to a massive reduction in computational expense and energy consumption.

来源详情: news.mit.edu ↗

为什么这很重要

The breakthrough demonstrates that AI can now handle games with astronomically large hidden‑state spaces at a fraction of the computational cost previously thought necessary. This opens the door for practical applications in domains where information is incomplete, such as business negotiations, cybersecurity threat assessment, and military planning. By showing that decision‑time planning with a generative model can reliably infer hidden variables, the research provides a template for building more cost‑effective, real‑world decision‑support tools. Moreover, the efficiency gains lower the barrier for institutions without massive compute budgets to develop comparable AI capabilities, potentially democratizing access to advanced strategic reasoning.

Stratego’s hidden‑information complexity (over 10^66 possible piece configurations) has long been a for AI strategic reasoning. Ataraxos’ success indicates that AI can now navigate such vast uncertainty spaces efficiently, a capability previously limited to simpler games like chess or Go.

The ability to infer hidden information and plan under uncertainty is directly relevant to real‑world scenarios where decision makers lack full visibility, such as market negotiations, cyber‑threat detection, and tactical military operations.

By achieving these results with far lower training costs, the research lowers the entry barrier for organizations to develop comparable AI systems, potentially accelerating innovation across sectors that cannot afford multi‑million‑dollar compute budgets.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
What is AI? Quiz

A route planner searches possible journeys using explicit rules. What does this illustrate about AI?

接下来看什么

Future work will focus on adding interpretability layers so Ataraxos can explain its reasoning to human users, a prerequisite for adoption in high‑stakes settings. Researchers will also explore scaling the approach to larger, more complex real‑world problems and measuring its performance against human experts in those domains. Funding from the Office of Naval Research and other agencies suggests possible defense‑related deployments, raising questions about transparency, oversight, and ethical use. Monitoring how the system is integrated into decision‑making pipelines and whether it maintains its efficiency advantage outside the controlled game environment will be essential.

Interpretability: The team plans to embed mechanisms that let Ataraxos explain its inferred hidden states and chosen actions, which is critical for trust and accountability in high‑risk applications.

Real‑world deployment: Funding from the Office of Naval Research hints at possible defense uses, prompting scrutiny over ethical guidelines, oversight, and potential dual‑use concerns.

Scalability: Researchers will test whether the efficiency gains hold when scaling to larger, more complex domains beyond board games, such as multi‑agent simulations or real‑time strategic planning.

Adoption barriers: Even with improved efficiency, practical integration will require user‑friendly interfaces, validation against domain experts, and robust safety testing before the technology can be trusted in operational settings.

相关指南和测验

什么是人工智能?人工智能模型解释AI 的未来AI 伦理测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注 AI 模型发布跟踪器
觉得这有用吗?