뉴스로 돌아가기
혁신AI Understanding 브리핑

MIT AI 시스템 Ataraxos가 최고의 인간 Stratego 플레이어를 이겼습니다.

MIT와 파트너 대학의 연구원들은 훨씬 적은 훈련 자원을 사용하면서 세계 최고의 Stratego 플레이어보다 훨씬 더 뛰어난 성능을 발휘하는 AI인 Ataraxos를 공개했습니다.

4 min readRead the primary source
Source-provided image accompanying MIT AI system Ataraxos beats top human Stratego players
기본 소스 문서녹음된 소스
출판사
news.mit.edu
소스 링크
news.mit.eduhttps://news.mit.edu/2026/game-playing-ai-stratego-new-champ-0930
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

강화 학습
에이전트가 장기적인 수익을 극대화하는 행동을 학습하는 보상 신호를 통한 교육입니다.
벤치마크
모델 성능을 측정하고 비교하는 데 사용되는 표준화된 테스트 또는 데이터 세트입니다.
추론
훈련된 모델이 예측 또는 출력을 생성하는 런타임 단계입니다.
자신을 테스트해 보세요AI란 무엇인가? 퀴즈

무슨 일이 일어났나요?

MIT, Carnegie Mellon, NYU, and Stanford researchers announced a new AI system called Ataraxos that achieved superhuman performance in the hidden‑information board game Stratego. Using self‑play combined with decision‑time planning and a generative model to infer hidden piece identities, Ataraxos defeated the strongest human player 15‑1‑4 and posted a 39‑2 record at the Stratego world championship. The system also topped human experts in several related imperfect‑information games, including Barrage Stratego, Hanabi, and Dou dizhu. Crucially, Ataraxos reached this level while using less than one‑hundredth of the training examples and one‑thirtieth of the self‑play games required by DeepMind’s prior DeepNash system, indicating a dramatic efficiency gain.

The team trained Ataraxos via self‑play , allowing the AI to generate a "blueprint strategy" for Stratego without human supervision. Efficient training algorithms reduced the number of required self‑play games dramatically compared with earlier systems.

During play, Ataraxos employs decision‑time planning: a generative model predicts the likely identities of hidden opponent pieces, then evaluates possible moves under those hypotheses before selecting an action. This approach replaces exhaustive enumeration with probabilistic , enabling rapid, informed decisions.

Performance metrics show Ataraxos beating the world’s top Stratego player with a 15‑1‑4 win‑loss‑draw record and achieving a 39‑2 record against other elite competitors. The system also surpassed human experts in Barrage Stratego, Hanabi, and Dou dizhu, confirming the method’s generality across imperfect‑information games.

Compared with DeepMind’s DeepNash, Ataraxos required less than 1 % of the training examples and roughly 3 % of the self‑play games, translating to a massive reduction in computational expense and energy consumption.

소스 세부정보: news.mit.edu ↗

왜 중요한가요?

The breakthrough demonstrates that AI can now handle games with astronomically large hidden‑state spaces at a fraction of the computational cost previously thought necessary. This opens the door for practical applications in domains where information is incomplete, such as business negotiations, cybersecurity threat assessment, and military planning. By showing that decision‑time planning with a generative model can reliably infer hidden variables, the research provides a template for building more cost‑effective, real‑world decision‑support tools. Moreover, the efficiency gains lower the barrier for institutions without massive compute budgets to develop comparable AI capabilities, potentially democratizing access to advanced strategic reasoning.

Stratego’s hidden‑information complexity (over 10^66 possible piece configurations) has long been a for AI strategic reasoning. Ataraxos’ success indicates that AI can now navigate such vast uncertainty spaces efficiently, a capability previously limited to simpler games like chess or Go.

The ability to infer hidden information and plan under uncertainty is directly relevant to real‑world scenarios where decision makers lack full visibility, such as market negotiations, cyber‑threat detection, and tactical military operations.

By achieving these results with far lower training costs, the research lowers the entry barrier for organizations to develop comparable AI systems, potentially accelerating innovation across sectors that cannot afford multi‑million‑dollar compute budgets.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
What is AI? Quiz

A route planner searches possible journeys using explicit rules. What does this illustrate about AI?

다음에 무엇을 볼 것인가

Future work will focus on adding interpretability layers so Ataraxos can explain its reasoning to human users, a prerequisite for adoption in high‑stakes settings. Researchers will also explore scaling the approach to larger, more complex real‑world problems and measuring its performance against human experts in those domains. Funding from the Office of Naval Research and other agencies suggests possible defense‑related deployments, raising questions about transparency, oversight, and ethical use. Monitoring how the system is integrated into decision‑making pipelines and whether it maintains its efficiency advantage outside the controlled game environment will be essential.

Interpretability: The team plans to embed mechanisms that let Ataraxos explain its inferred hidden states and chosen actions, which is critical for trust and accountability in high‑risk applications.

Real‑world deployment: Funding from the Office of Naval Research hints at possible defense uses, prompting scrutiny over ethical guidelines, oversight, and potential dual‑use concerns.

Scalability: Researchers will test whether the efficiency gains hold when scaling to larger, more complex domains beyond board games, such as multi‑agent simulations or real‑time strategic planning.

Adoption barriers: Even with improved efficiency, practical integration will require user‑friendly interfaces, validation against domain experts, and robust safety testing before the technology can be trusted in operational settings.

관련 가이드 및 퀴즈

AI란 무엇인가?AI 모델 설명AI의 미래AI 윤리알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?