7 Sources
[1]
'Humanity's Last Exam' benchmark is stumping top AI models - can you do any better?
A new academic benchmark aims to 'test the limits of AI knowledge at the frontiers of human expertise.' So far, these LLMs are stumped. Are artificial intelligence (AI) models really surpassing human ability? Or are current tests just too easy for them? On Thursday, Scale AI and the Center for AI
[2]
A new AI benchmark called 'Humanity's Last Exam' stumped top models -- for now, at least
Despite facing increasingly harder tests, artificial intelligence models have been advancing quickly and passing even PhD-level exams with high scores, making it somewhat difficult to track just how good they're getting. But it seems the AI models have met their match -- at least for
[3]
Could you pass 'Humanity's Last Exam'? Probably not, but neither can AI
Did you know some of the smartest people on the planet create benchmarks to test AI's capabilities at replicating human intelligence? Well, scarily enough most AI benchmarks are easily completed by artificial intelligence models, showcasing just how smart the likes of ChatGPT's GPT-4o, Google
[4]
Humanity's Last Exam is the New MultiAgent AI Benchmark
The AI field welcomes a new benchmark: Humanity's Last Exam (HLE), introduced by the Center for AI Safety (CAIS) and Scale AI for testing AI systems on expert-level knowledge. The dataset includes 3,000 questions crowdsourced from 1,000 contributors across 500 institutions in 50 countries,
[5]
Even some of the best AI can't beat this new benchmark
The nonprofit Center for AI Safety (CAIS) and Scale AI, a company that provides a number of data labeling and AI development services, have released a challenging new benchmark for frontier AI systems. The benchmark, called Humanity's Last Exam, includes thousands of crowdsourced questions
[6]
When AI passes this test, look out
AI systems are surpassing traditional tests, prompting the creation of "Humanity's Last Exam", a collection of extremely difficult questions across various fields. This new benchmark aims to measure AI's ability to tackle complex problems, though initial results show AI models still struggle. The
[7]
A Test So Hard No AI System Can Pass It -- Yet
If you're looking for a new reason to be nervous about artificial intelligence, try this: Some of the smartest humans in the world are struggling to create tests that A.I. systems can't pass. For years, A.I. systems were measured by giving new models a variety of standardized benchmark tests. Many
Share
Copy Link
Scale AI and the Center for AI Safety have introduced a challenging new AI benchmark called 'Humanity's Last Exam', which has proven difficult for even the most advanced AI models, highlighting the current limitations of artificial intelligence.

Scale AI and the Center for AI Safety (CAIS) have introduced a groundbreaking new AI benchmark called "Humanity's Last Exam" (HLE), designed to test the limits of AI knowledge at the frontiers of human expertise
1
2
. This benchmark aims to address the issue of "benchmark saturation," where AI models have been rapidly excelling on standard tests, making it difficult to accurately gauge their capabilities3
.The HLE consists of 3,000 questions covering over 100 subjects in mathematics, science, and humanities
1
. These questions were carefully selected from an initial pool of 70,000, with input from nearly 1,000 subject expert contributors across 500 institutions in 50 countries2
4
. The benchmark includes multiple-choice and short-answer questions, as well as multi-modal elements incorporating text, diagrams, and images4
.In initial testing, current AI models struggled significantly with the HLE:
3
5
These results stand in stark contrast to the high scores (often over 90%) that many of these models achieve on other popular benchmarks like MMLU, MATH, and GPQA
1
2
.The poor performance of top AI models on the HLE reveals that there are still significant gaps in AI capabilities when it comes to expert-level knowledge and complex reasoning
2
. Dan Hendrycks, co-founder and executive director of CAIS, noted that while it's uncertain how quickly models will advance, the HLE currently demonstrates that there are still expert-level questions that AI models cannot answer1
2
.Related Stories
While the current results show a clear limitation in AI capabilities, researchers are cautious about making long-term predictions. Given the rapid pace of AI advancement, it's considered plausible that models could reach over 50% accuracy on the HLE by the end of the year
2
. However, the benchmark's creators emphasize that such an achievement would not necessarily indicate autonomous research capabilities or artificial general intelligence2
.CAIS and Scale AI plan to release the HLE dataset to researchers for further study of AI systems and their limitations
1
. The benchmark remains open for additional test questions, though cash prizes are no longer being awarded1
. This initiative represents an important step in creating more challenging and comprehensive evaluations of AI capabilities as the field continues to evolve rapidly.Summarized by
Navi
[5]
27 Feb 2026•Science and Research

04 Feb 2025•Technology

17 Sept 2024

1
Technology

2
Technology

3
Technology
