2 Sources
[1]
Scalable decision-making for games of imperfect information
The combination of reinforcement learning and search has proven to be a powerful recipe for strategic decision-making25,26,27,28,29,30,31. But in the past, this recipe has had limited applicability to imperfect-information settings. The success of the AIs presented herein shows that the presence of
[2]
This game-playing AI is the new champ at Stratego
"With Stratego, there is an explosion of possible universes you might have to deal with. AI techniques that were developed for games like poker definitely could not scale in this setting," Gabriele Farina says. A new AI system that excels at challenging games with hidden information could someday
Share
Copy Link
Researchers from MIT, Carnegie Mellon University, Stanford, and NYU developed Ataraxos, a game-playing AI that defeated top-ranked human Stratego players by a large margin. The system uses transformer-based networks and efficient reinforcement learning to handle games with hidden information, achieving superhuman performance at lower computational costs than previous models like DeepNash.
Researchers from MIT, Carnegie Mellon University, Stanford, and NYU have developed Ataraxos, a game-playing AI system that achieved superhuman performance in Stratego, defeating top-ranked human players by a significant margin
2
. Published in Nature, this breakthrough marks the first time an AI has successfully mastered the complex board wargame, which features more than 10^66 possible piece configurations—exponentially greater than chess2
.Stratego presents unique challenges for AI systems because players arrange 40 pieces on their side of the board with piece identities remaining secret until they collide. This creates what researchers call imperfect information scenarios, where decisions become intertwined in ways that make determining optimal strategies extremely difficult. "With Stratego, there is an explosion of possible universes you might have to deal with. AI techniques that were developed for games like poker definitely could not scale in this setting," explains Gabriele Farina, assistant professor in MIT's Department of Electrical Engineering and Computer Science and senior author of the research
2
.
Source: MIT
Ataraxos employs two interdependent self-play reinforcement learning processes realized by transformer-based networks
1
. The first handles set-up selection, where players privately determine starting positions for their pieces, while the second manages move selection during gameplay. The researchers chose this decomposition because an end-to-end approach would require a single model to accommodate two unrelated tasks1
.The set-up network uses a decoder-only transformer architecture, allowing training on entire set-ups with single forward-backward passes. Meanwhile, the move network employs an encoder-only architecture with a key-query matrix product to parameterize the policy over moves, which learned faster than less sophisticated approaches
1
. Both networks utilize learned absolute positional embeddings, with the move network sized specifically to balance sample efficiency against iteration speed.The system achieved greater performance at Stratego than competing models while being far cheaper and less computationally demanding to train
2
. Previous efforts, including Google's DeepNash, relied on sophisticated operations requiring millions of dollars in training costs but still couldn't beat top human Stratego players2
.Ataraxos generates self-play games by directly sampling from its networks, with moves computed using λ returns for estimating expected cumulants and advantages. The system trains only on moves with large estimated advantage magnitudes, a filtering approach that reduced overall wall-clock time per reinforcement learning iteration by a factor of approximately 2.5 while simultaneously increasing both sample efficiency and asymptotic performance
1
.For set-ups, the system uses Monte Carlo returns to estimate quantities based on final game outcomes. "Our system reaches strictly higher playing strength than DeepNash," the researchers noted
2
.Related Stories
The AI system demonstrated generalization capabilities by outperforming top human players in other strategic games with different rules and designs
2
. This versatility positions Ataraxos for adaptation to real-world problems involving imperfect information, including business negotiations, cybersecurity, and military strategy."In the kind of imperfect information tasks you would face in reality, you often don't have the luxury of enumerating through all the possibilities. There are just too many," Farina explains
2
. The combination of reinforcement learning and search has proven powerful for strategic decision-making, and the success of Ataraxos demonstrates that large amounts of hidden information are no longer prohibitive for practical AI applications1
.The research team includes lead author Samuel Sokota from Carnegie Mellon University, Eugene Vinitsky and Zico Kolter from NYU, Hengyuan Hu from Stanford, and Zhiyuan Fan, a graduate student at MIT
2
. Their work puts practical AI within reach for many strategic decision-making problems where fast, accurate simulators can be constructed, addressing scenarios from financial markets where traders lack knowledge of others' rationale to military forces operating without full knowledge of enemy positions1
2
.Summarized by
Navi
05 Jun 2026•Science and Research

21 Feb 2025•Technology

09 Jun 2025•Technology

1
Technology

2
Policy and Regulation

3
Policy and Regulation
