8 Sources
[1]
A new math benchmark just dropped and leading AI models can solve 'less than 2%' of its problems... oh dear
Sometimes I forget there's a whole other world out there where AI models aren't just used for basic tasks such as simple research and quick content summaries. Out in the land of bigwigs, they're instead being used to help with everything from financial analysis to scientific research. That's why
[2]
New secret math benchmark stumps AI models and PhDs alike
On Friday, research organization Epoch AI released FrontierMath, a new mathematics benchmark that has been turning heads in the AI world because it contains hundreds of expert-level problems that leading AI models solve less than 2 percent of the time, according to Epoch AI. The benchmark tests AI
[3]
Testing AI systems on hard math problems shows they still perform very poorly
A team of AI researchers and mathematicians affiliated with several institutions in the U.S. and the U.K. has developed a math benchmark that allows scientists to test the ability of AI systems to solve exceptionally difficult math problems. Their paper is posted on the arXiv preprint server. Over
[4]
AI's math problem: FrontierMath benchmark shows how far technology still has to go
Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Artificial intelligence systems may be good at generating text, recognizing images, and even solving basic math problems -- but when it comes to advanced mathematical
[5]
GPT-4 and Gemini Scored Less Than 2 Percent on This New AI Benchmark
The company said older benchmarks do not truly test AI capabilities Epoch AI, a California-based research institute launched a new artificial intelligence (AI) benchmark last week. Dubbed FrontierMath, the new AI benchmark tests large language models (LLMs) on their capability of reseasoning and
[6]
Never Mind Coding -- o1 is Downright Awful at Maths!
It's not just OpenAI's o1 -- no LLM in the world is anywhere close to cracking the toughest problems in mathematics (yet). A few days ago, Epoch AI released FrontierMath, a new benchmark to evaluate the mathematical capabilities of large language models. The results revealed a startling low for
[7]
OpenAI o1 Can't Do Maths, But Excels at Making Excuses
It's not just OpenAI's o1 -- no LLM in the world is anywhere close to cracking the toughest problems in mathematics (yet). A few days ago, Epoch AI released FrontierMath, a new benchmark to evaluate the mathematical capabilities of large language models. The results revealed a startling low for
[8]
OpenAI is So Doomed if Inference Time Scaling for o1 Fails
But Sam Altman and his team are taking their biggest risk ever to bring AGI next year. OpenAI's progress from GPT-4 to Orion has slowed, The information reported recently. According to the report, although OpenAI has completed only 20% of Orion's training, it is already on par with GPT-4 in
Share
Copy Link
Epoch AI's FrontierMath, a new mathematics benchmark, reveals that leading AI models struggle with complex mathematical problems, solving less than 2% of the challenges.

Epoch AI, a California-based research institute, has introduced FrontierMath, a groundbreaking benchmark designed to test the advanced mathematical reasoning capabilities of large language models (LLMs). This new benchmark has exposed significant limitations in current AI systems, with even leading models solving less than 2% of the problems
1
.Existing mathematical benchmarks like GSM-8k and MATH have become less effective in evaluating AI capabilities, with top models scoring over 90% on these tests
2
. Epoch AI argues that these high scores are partly due to data contamination, where AI models have been trained on similar problems, leading to artificially inflated performance4
.FrontierMath consists of hundreds of original, expert-crafted mathematics problems that are:
1
The benchmark was developed in collaboration with over 60 mathematicians from leading institutions. The problems underwent peer review to ensure correctness and check for ambiguities, with about 1 in 20 problems requiring corrections during the review process
2
.Despite their high performance on simpler math benchmarks, top AI models like Claude 3.5 Sonnet, GPT-4o, o1-preview, and Gemini 1.5 Pro scored extremely poorly on FrontierMath, even with access to Python environments for testing and verification
2
.Related Stories
Fields Medalist Terence Tao commented on the difficulty of the problems, stating that solving them would likely require a combination of a semi-expert (like a graduate student in a related field), modern AI, and various algebra packages
4
.FrontierMath's results highlight the current limitations of AI in complex reasoning tasks. The benchmark serves as a crucial tool for evaluating genuine mathematical understanding and creativity in AI systems, rather than simple pattern matching or brute-force approaches
4
.While AI models have made significant strides in various domains, FrontierMath demonstrates that there is still a substantial gap between current AI capabilities and human-level mathematical reasoning. This benchmark sets a new standard for evaluating AI progress in advanced problem-solving and may guide future developments in AI research and applications
3
.Summarized by
Navi
[1]
[2]
13 Jan 2025•Science and Research

25 Mar 2025•Science and Research

24 Jan 2025•Science and Research

1
Technology

2
Policy and Regulation

3
Technology
