A new AI system developed by MIT, Carnegie Mellon University, New York University, and Stanford University has mastered Stratego, outperforming top human players and previous models, with potential real-world applications in decision-making under uncertainty.
The Breakthrough in Imperfect Information Games
The AI system, named Ataraxos, achieved superhuman performance in Stratego, a complex board game with hidden information, by defeating top-ranked human players by a large margin. This marks a significant advancement in AI’s ability to handle real-world scenarios where information is incomplete, such as business negotiations or cybersecurity. The research, published in Nature, highlights how the system’s approach could generalize to other strategic challenges.
In the kind of imperfect information tasks you would face in reality, you often don’t have the luxury of enumerating through all the possibilities,
said Gabriele Farina, an assistant professor in the Department of Electrical Engineering and Computer Science (EECS), principal investigator at the Laboratory for Information and Decision Systems (LIDS), and senior author of a paper on this AI system. Having AI algorithms that are general purpose and can provably perform this challenging task so well is a big step forward,

Technical Innovations Behind Ataraxos
Ataraxos combines self-play reinforcement learning with specialized algorithms tailored for hidden information settings. The system uses two interdependent self-play processes: one for setting up game positions and another for making moves. This dual approach allows it to learn a blueprint strategy
efficiently, reducing computational costs compared to previous methods. The research team emphasized that the system’s design enables it to scale to problems with exponentially large decision spaces, such as Stratego’s 10 to the 66th power possible configurations.
With Stratego, there is an explosion of possible universes you might have to deal with. AI techniques that were developed for games like poker definitely could not scale in this setting,
Farina says. The system’s efficiency stems from its use of transformer networks and dynamic damping in self-play training, which optimizes learning without excessive computational demands.
Implications for Real-World Applications
The researchers suggest that Ataraxos’ framework could revolutionize fields requiring strategic decision-making under uncertainty. By simulating complex scenarios, the AI could assist in areas like financial trading, military logistics, or diplomatic negotiations. For instance, the system’s ability to handle hidden information mirrors challenges in cybersecurity, where adversaries may obscure their intentions.
The more you bluff, the more your opponent expects it, and the less each bluff is worth. It’s not obvious how to reason about that,
Sokota explains. It’s very different from a setting like chess, where the best move is still the best move no matter how often you’ve played it.
This distinction highlights the AI’s ability to adapt to dynamic, information-scarce environments.

Collaborative Research and Future Directions
The project involved a multidisciplinary team from MIT, Carnegie Mellon University, New York University, and Stanford University. The researchers plan to explore applications beyond gaming, focusing on scenarios where human-AI collaboration could enhance outcomes.
The AI system could be adapted to help humans tackle many real-world problems with hidden information, such as business negotiations or cybersecurity.