3 Sources
[1]
Move over math and reasoning, it's time to benchmark AI using Super Mario Bros.
The big picture: Benchmarking AI remains a thorny issue, with companies often accused of cherry-picking flattering results while burying less favorable ones. Instead of fixating on math and logic trials, perhaps it's time for a more unconventional test - one that challenges AI in a way humans
[2]
People are using Super Mario to benchmark AI now | TechCrunch
Thought Pokémon was a tough benchmark for AI? One group of researchers argues that Super Mario Bros. is even tougher. Hao AI Lab, a research org at the University of California San Diego, on Friday threw AI into live Super Mario Bros. games. Anthropic's Claude 3.7 performed the best, followed by
[3]
AI Models Tested in Super Mario Bros. Reveal Speed vs. Reasoning Trade-Off
Super Mario Bros. AI Test Highlights Strengths and Weaknesses of Modern Models A research team from Hao AI Lab at the University of California San Diego has uniquely tested artificial intelligence models -- by making them play Super Mario Bros. Unlike traditional benchmarks, this real-time gaming
Share
Copy Link
Researchers at UC San Diego's Hao AI Lab use Super Mario Bros. to test AI models, revealing unexpected strengths and weaknesses in different AI approaches.

In an innovative approach to AI evaluation, researchers at the Hao AI Lab at the University of California San Diego have employed an unexpected tool: the classic video game Super Mario Bros. This unconventional benchmark aims to test AI models' ability to navigate complex, real-time environments, offering a fresh perspective on their capabilities beyond traditional reasoning and mathematical tasks
1
.The experiment utilized an emulated version of Super Mario Bros. integrated with a custom framework called GamingAgent, developed by the Hao Lab. This system allowed AI models to control Mario by generating Python code based on basic instructions and screenshot visualizations of the game state
2
.The outcomes of this unique test revealed unexpected strengths and weaknesses among different AI models:
Top Performers: Anthropic's Claude 3.7 emerged as the leader, showcasing impressive reflexes and skillful gameplay. Its predecessor, Claude 3.5, also performed well
1
.Unexpected Struggles: Reasoning-heavy models like OpenAI's GPT-4o and Google's Gemini 1.5 Pro, despite their reputation for strong reasoning abilities, lagged behind in performance
3
.Related Stories
Researchers discovered that success in Super Mario Bros. hinged more on timing than logical reasoning. The game's fast-paced nature requires quick decision-making, with even slight delays potentially resulting in failure. This revelation suggests that more deliberative models may have taken too long to calculate their next moves, leading to frequent in-game deaths
1
.While using retro video games to benchmark AI is largely a playful experiment, it raises important questions about AI evaluation methods:
Real-world Applicability: The study highlights the need for diverse testing environments that challenge AI in ways that mirror complex, dynamic real-world scenarios
2
.Speed vs. Reasoning Trade-off: The results underscore a potential trade-off between quick decision-making and deep reasoning capabilities in AI models, prompting discussions about balancing these attributes for various applications
3
.Evaluation Crisis: Some experts, like Andrej Karpathy from OpenAI, point to an "evaluation crisis" in AI, questioning the reliability of current metrics in assessing AI capabilities
2
.Summarized by
Navi
[3]
1
Policy and Regulation

2
Technology

3
Technology
