2 Sources
[1]
OpenAI's Deep Research smashes records for the world's hardest AI exam, with ChatGPT o3-mini and DeepSeek left in its wake
The world's hardest AI exam, Humanity's Last Exam, was launched less than two weeks ago, and we've already seen a huge jump in accuracy, with ChatGPT o3-mini and now OpenAI's Deep Reasoning topping the leaderboard. The AI benchmark created by experts from around the world contains some of the
[2]
Humanity's Last Exam Explained - The ultimate AI benchmark that sets the tone of our AI future
Artificial intelligence (AI) has been evolving at breakneck speed, with models like OpenAI's GPT-4 and DeepSeek's R1 pushing the boundaries of what machines can do. We are in an era where artificial intelligence (AI) systems can write poetry, diagnose diseases, and even drive cars. At this moment,
Share
Copy Link
OpenAI's Deep Research achieves a record-breaking 26.6% accuracy on Humanity's Last Exam, a new benchmark designed to test the limits of AI reasoning and problem-solving abilities across diverse fields.

In a significant leap forward for artificial intelligence, OpenAI's Deep Research has achieved a groundbreaking score of 26.6% accuracy on Humanity's Last Exam (HLE), a newly established benchmark designed to push AI systems to their limits
1
. This result represents a staggering 183% increase in accuracy compared to previous top performers, setting a new standard for AI capabilities in complex reasoning and problem-solving.HLE, developed by the Center for AI Safety (CAIS) and Scale AI, is considered the world's hardest AI exam. It comprises 3,000 challenging questions spanning over 100 subjects, including mathematics, physics, law, medicine, and philosophy
2
. Unlike previous benchmarks, HLE incorporates both text and image-based questions, with 10% of the exam requiring visual processing alongside written context.The AI community has witnessed remarkable progress in a short span of time. Just days before Deep Research's achievement, other models had set impressive benchmarks:
1
It's worth noting that Deep Research's exceptional performance is partly attributed to its web search capabilities, which are not available to other AI models. This feature provides an advantage in addressing general knowledge questions included in the exam
1
.Related Stories
HLE represents a critical shift in how AI progress is measured and evaluated:
Exposing AI Weaknesses: The exam reveals areas where AI still struggles, such as deep reasoning and multi-modal understanding
2
.Setting New Standards: HLE challenges AI companies to focus on meaningful advancements rather than superficial improvements
2
.Increasing Accountability: The benchmark introduces transparency and forces AI models to perform under pressure, mimicking real-world scenarios
2
.While Deep Research's 26.6% accuracy on HLE is impressive, it still falls short of what would be considered a passing grade in human terms. This underscores the significant challenges that remain in developing AI systems capable of human-level reasoning across diverse fields
1
.As AI continues to evolve rapidly, HLE will likely play a crucial role in gauging progress and directing research efforts. The AI community now faces the exciting challenge of pushing beyond current limitations, with many wondering how long it will take for an AI model to surpass the 50% mark on this rigorous exam
1
2
.Summarized by
Navi
24 Jan 2025•Science and Research

27 Feb 2026•Science and Research

17 Sept 2024

1
Technology

2
Technology

3
Technology
