5 Sources
[1]
A new, challenging AGI test stumps most AI models | TechCrunch
The Arc Prize Foundation, a nonprofit co-founded by prominent AI researcher François Chollet, announced in a blog post on Tuesday that it has created a new, challenging test to measure the general intelligence of leading AI models. So far, the new test, called ARC-AGI-2, has stumped most
[2]
Leading AI models fail new test of artificial general intelligence
The most sophisticated AI models in existence today have scored poorly on a new benchmark designed to measure their progress towards artificial general intelligence (AGI) - and brute-force computing power won't be enough to improve, as evaluators are now taking into account the cost of running the
[3]
ChatGPT, Gemini and Claude all failed to solve a simple test that humans are acing
As artificial intelligence continues to build on its reputation as the smartest thing in the room, it will be oddly therapeutic to hear that one test has it stumped. In fact, this new AI examination system is causing issues for even the most advanced models. ARC-AG2, or to use its more glamorous
[4]
A new AI test is outwitting OpenAI, Google models, among others
Humans are still way smarter than AI according to this new AGI benchmark. Credit: karetoria / Getty Images Google, OpenAI, DeepSeek, et al. are nowhere near achieving AGI (Artificial General Intelligence), according to a new benchmark. The Arc Prize Foundation, a nonprofit that measures AGI
[5]
LLMs Hit a New Low on ARC-AGI-2 Benchmark, Pure LLMs Score 0%
ARC Prize, a non-profit organisation that evaluates the effectiveness of AI models to demonstrate human-like intelligence, has announced the ARC-AGI-2 benchmark. The new benchmark is a successor to the ARC-AGI benchmark released a few years ago. Like its predecessor, the benchmark tests AI models
Share
Copy Link
The Arc Prize Foundation introduces ARC-AGI-2, a challenging new test for artificial general intelligence that current AI models, including those from OpenAI and Google, are struggling to solve. The benchmark emphasizes efficiency and adaptability, revealing limitations in current AI capabilities.

The Arc Prize Foundation, a nonprofit co-founded by prominent AI researcher François Chollet, has unveiled a new benchmark test called ARC-AGI-2, designed to measure the general intelligence of leading AI models
1
. This test has proven to be significantly more challenging than its predecessor, with most current AI models struggling to achieve even single-digit scores.The results of the ARC-AGI-2 test have been eye-opening:
2
.1
.1
3
.5
.In stark contrast, a human panel achieved an average score of 60% on the test, with some individuals solving all tasks perfectly
1
5
.The new benchmark introduces several important changes:
1
2
.3
.1
.5
.The poor performance of leading AI models on ARC-AGI-2 highlights the significant gap between current AI capabilities and human-level general intelligence. Greg Kamradt, co-founder of the Arc Prize Foundation, emphasized that intelligence is not solely about problem-solving ability but also about the efficiency of acquiring and deploying new skills
1
.This benchmark challenges the notion that brute-force computing power alone can lead to AGI. It suggests that fundamental advancements in AI architecture and learning approaches may be necessary to achieve human-like adaptability and efficiency
2
4
.Related Stories
While many in the tech industry welcome new benchmarks to measure AI progress, some experts question the framing of these tests. Catherine Flick from the University of Staffordshire argues that performing well on such benchmarks should not be seen as a major step towards AGI, as they only assess an AI's ability to complete specific tasks rather than demonstrate true general intelligence
2
.The introduction of ARC-AGI-2 raises questions about the future of AGI evaluation. Joseph Imperial from the University of Bath suggests that future iterations might incorporate additional metrics, such as the minimum number of humans required to solve tasks, alongside performance and efficiency measures
2
.As the debate over AGI continues, the Arc Prize Foundation has announced a new contest challenging developers to reach 85% accuracy on the ARC-AGI-2 test while spending only $0.42 per task
1
. This competition aims to drive innovation in both AI performance and efficiency, potentially bringing us closer to the elusive goal of artificial general intelligence.Summarized by
Navi
[2]
24 Jan 2025•Science and Research

11 Oct 2024•Science and Research

05 Apr 2025•Science and Research

1
Technology

2
Technology

3
Technology
