Researchers from MIT, Carnegie Mellon University, Stanford, and NYU developed Ataraxos, a game-playing AI that defeated top-ranked human Stratego players by a large margin. The system uses transformer-based networks and efficient reinforcement learning to handle games with hidden information, achieving superhuman performance at lower computational costs than previous models like DeepNash.

Ataraxos Defeats Top Players in Complex Board Wargame

Researchers from MIT, Carnegie Mellon University, Stanford, and NYU have developed Ataraxos, a game-playing AI system that achieved superhuman performance in Stratego, defeating top-ranked human players by a significant margin

2

. Published in Nature, this breakthrough marks the first time an AI has successfully mastered the complex board wargame, which features more than 10^66 possible piece configurations—exponentially greater than chess

2

.

Stratego presents unique challenges for AI systems because players arrange 40 pieces on their side of the board with piece identities remaining secret until they collide. This creates what researchers call imperfect information scenarios, where decisions become intertwined in ways that make determining optimal strategies extremely difficult. "With Stratego, there is an explosion of possible universes you might have to deal with. AI techniques that were developed for games like poker definitely could not scale in this setting," explains Gabriele Farina, assistant professor in MIT's Department of Electrical Engineering and Computer Science and senior author of the research

2

.

Source: MIT

Source: MIT

Scalable Decision-Making Through Dual Transformer Architecture

Ataraxos employs two interdependent self-play reinforcement learning processes realized by transformer-based networks

1

. The first handles set-up selection, where players privately determine starting positions for their pieces, while the second manages move selection during gameplay. The researchers chose this decomposition because an end-to-end approach would require a single model to accommodate two unrelated tasks

1

.

The set-up network uses a decoder-only transformer architecture, allowing training on entire set-ups with single forward-backward passes. Meanwhile, the move network employs an encoder-only architecture with a key-query matrix product to parameterize the policy over moves, which learned faster than less sophisticated approaches

1

. Both networks utilize learned absolute positional embeddings, with the move network sized specifically to balance sample efficiency against iteration speed.

Efficient Training Cuts Costs While Boosting Performance

The system achieved greater performance at Stratego than competing models while being far cheaper and less computationally demanding to train

2

. Previous efforts, including Google's DeepNash, relied on sophisticated operations requiring millions of dollars in training costs but still couldn't beat top human Stratego players

2

.

Ataraxos generates self-play games by directly sampling from its networks, with moves computed using λ returns for estimating expected cumulants and advantages. The system trains only on moves with large estimated advantage magnitudes, a filtering approach that reduced overall wall-clock time per reinforcement learning iteration by a factor of approximately 2.5 while simultaneously increasing both sample efficiency and asymptotic performance

1

.

For set-ups, the system uses Monte Carlo returns to estimate quantities based on final game outcomes. "Our system reaches strictly higher playing strength than DeepNash," the researchers noted

2

.

Real-World Applications Beyond Gaming

The AI system demonstrated generalization capabilities by outperforming top human players in other strategic games with different rules and designs

2

. This versatility positions Ataraxos for adaptation to real-world problems involving imperfect information, including business negotiations, cybersecurity, and military strategy.

"In the kind of imperfect information tasks you would face in reality, you often don't have the luxury of enumerating through all the possibilities. There are just too many," Farina explains

2

. The combination of reinforcement learning and search has proven powerful for strategic decision-making, and the success of Ataraxos demonstrates that large amounts of hidden information are no longer prohibitive for practical AI applications

1

.

The research team includes lead author Samuel Sokota from Carnegie Mellon University, Eugene Vinitsky and Zico Kolter from NYU, Hengyuan Hu from Stanford, and Zhiyuan Fan, a graduate student at MIT

2

. Their work puts practical AI within reach for many strategic decision-making problems where fast, accurate simulators can be constructed, addressing scenarios from financial markets where traders lack knowledge of others' rationale to military forces operating without full knowledge of enemy positions

1

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved