Volver a Noticias
InnovaciónAI Understanding sesión informativa

El sistema de inteligencia artificial del MIT, Ataraxos, supera a los mejores jugadores humanos de Stratego

Investigadores del MIT y universidades asociadas presentaron Ataraxos, una IA que superó a los mejores jugadores de Stratego del mundo por un amplio margen y utilizó muchos menos recursos de capacitación.

4 min readRead the primary source
Source-provided image accompanying MIT AI system Ataraxos beats top human Stratego players
Documento de fuente primariaFuente registrada
Editor
news.mit.edu
Enlace fuente
news.mit.eduhttps://news.mit.edu/2026/game-playing-ai-stratego-new-champ-0930
Tipo de fuente
Documento principal: un anuncio oficial, documento, archivo o página propia que leemos directamente.
ContextoEntiende esto en 60 segundos

Empieza aquí

Términos clave

Aprendizaje por refuerzo
Entrenamiento mediante señales de recompensa donde un agente aprende acciones que maximizan el retorno a largo plazo.
Punto de referencia
Una prueba o conjunto de datos estandarizado que se utiliza para medir y comparar el rendimiento del modelo.
Inferencia
La fase de tiempo de ejecución donde un modelo entrenado genera predicciones o resultados.
Ponte a prueba¿Qué es la IA? cuestionario

que paso

MIT, Carnegie Mellon, NYU, and Stanford researchers announced a new AI system called Ataraxos that achieved superhuman performance in the hidden‑information board game Stratego. Using self‑play combined with decision‑time planning and a generative model to infer hidden piece identities, Ataraxos defeated the strongest human player 15‑1‑4 and posted a 39‑2 record at the Stratego world championship. The system also topped human experts in several related imperfect‑information games, including Barrage Stratego, Hanabi, and Dou dizhu. Crucially, Ataraxos reached this level while using less than one‑hundredth of the training examples and one‑thirtieth of the self‑play games required by DeepMind’s prior DeepNash system, indicating a dramatic efficiency gain.

The team trained Ataraxos via self‑play , allowing the AI to generate a "blueprint strategy" for Stratego without human supervision. Efficient training algorithms reduced the number of required self‑play games dramatically compared with earlier systems.

During play, Ataraxos employs decision‑time planning: a generative model predicts the likely identities of hidden opponent pieces, then evaluates possible moves under those hypotheses before selecting an action. This approach replaces exhaustive enumeration with probabilistic , enabling rapid, informed decisions.

Performance metrics show Ataraxos beating the world’s top Stratego player with a 15‑1‑4 win‑loss‑draw record and achieving a 39‑2 record against other elite competitors. The system also surpassed human experts in Barrage Stratego, Hanabi, and Dou dizhu, confirming the method’s generality across imperfect‑information games.

Compared with DeepMind’s DeepNash, Ataraxos required less than 1 % of the training examples and roughly 3 % of the self‑play games, translating to a massive reduction in computational expense and energy consumption.

Detalles de la fuente: news.mit.edu ↗

Por qué es importante

The breakthrough demonstrates that AI can now handle games with astronomically large hidden‑state spaces at a fraction of the computational cost previously thought necessary. This opens the door for practical applications in domains where information is incomplete, such as business negotiations, cybersecurity threat assessment, and military planning. By showing that decision‑time planning with a generative model can reliably infer hidden variables, the research provides a template for building more cost‑effective, real‑world decision‑support tools. Moreover, the efficiency gains lower the barrier for institutions without massive compute budgets to develop comparable AI capabilities, potentially democratizing access to advanced strategic reasoning.

Stratego’s hidden‑information complexity (over 10^66 possible piece configurations) has long been a for AI strategic reasoning. Ataraxos’ success indicates that AI can now navigate such vast uncertainty spaces efficiently, a capability previously limited to simpler games like chess or Go.

The ability to infer hidden information and plan under uncertainty is directly relevant to real‑world scenarios where decision makers lack full visibility, such as market negotiations, cyber‑threat detection, and tactical military operations.

By achieving these results with far lower training costs, the research lowers the entry barrier for organizations to develop comparable AI systems, potentially accelerating innovation across sectors that cannot afford multi‑million‑dollar compute budgets.

Interactive Mechanism

Mecanismo interactivo: cómo funciona realmente

Explore la tecnología subyacente detrás de este desarrollo de forma interactiva.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Verificación interactiva del concepto+10 Points
What is AI? Quiz

A route planner searches possible journeys using explicit rules. What does this illustrate about AI?

Qué ver a continuación

Future work will focus on adding interpretability layers so Ataraxos can explain its reasoning to human users, a prerequisite for adoption in high‑stakes settings. Researchers will also explore scaling the approach to larger, more complex real‑world problems and measuring its performance against human experts in those domains. Funding from the Office of Naval Research and other agencies suggests possible defense‑related deployments, raising questions about transparency, oversight, and ethical use. Monitoring how the system is integrated into decision‑making pipelines and whether it maintains its efficiency advantage outside the controlled game environment will be essential.

Interpretability: The team plans to embed mechanisms that let Ataraxos explain its inferred hidden states and chosen actions, which is critical for trust and accountability in high‑risk applications.

Real‑world deployment: Funding from the Office of Naval Research hints at possible defense uses, prompting scrutiny over ethical guidelines, oversight, and potential dual‑use concerns.

Scalability: Researchers will test whether the efficiency gains hold when scaling to larger, more complex domains beyond board games, such as multi‑agent simulations or real‑time strategic planning.

Adoption barriers: Even with improved efficiency, practical integration will require user‑friendly interfaces, validation against domain experts, and robust safety testing before the technology can be trusted in operational settings.

Guías y cuestionarios relacionados

¿Qué es la IA?Modelos de IA explicadosFuturo de la IAÉtica de la IAPon a prueba lo que sabes: prueba un cuestionario gratuito sobre IABusque un término de IA en nuestro glosarioSiga el rastreador de lanzamientos de modelos de IA
¿Encontró esto útil?