4 Sources
[1]
Did OpenAI Cheat on Its Big Math Test? - Decrypt
How intelligent is a model that memorizes the answers before an exam? That's the question facing OpenAI after it unveiled o3 in December, and touted its model's impressive benchmarks. At the time, some pundits hailed it as being almost as powerful as AGI, the level at which artificial intelligence
[2]
OpenAI Faces Scrutiny Over o3 Model's FrontierMath Benchmarking Transparency
AI researchers have put OpenAI in the spotlight, especially by the newest released AI model, o3 following its unprecedented performance on the FrontierMath benchmarking test it passed. While OpenAI recently reported 25% accuracy on this particular and difficult mathematician benchmark, issues of
[3]
OpenAI Just Pulled a Theranos With o3
OpenAI's o3 benchmark controversy is starting to look like a Theranos moment -- claiming record-breaking performance on EpochAI's FrontierMath benchmark while having access to much of the test data, and funding the same. Epoch AI's associate director, Tamay Besiroglu admitted they were
[4]
AI benchmarking organization criticized for waiting to disclose funding from OpenAI | TechCrunch
An organization developing math benchmarks for AI didn't disclose that it had received funding from OpenAI until relatively recently, drawing allegations of impropriety from some in the AI community. Epoch AI, a nonprofit primarily funded by Open Philanthropy, a research and grantmaking
Share
Copy Link
OpenAI's impressive performance on the FrontierMath benchmark with its o3 model is under scrutiny due to the company's involvement in creating the test and having access to problem sets, raising questions about the validity of the results and the transparency of AI benchmarking.

OpenAI recently unveiled its o3 model, claiming an impressive 25.2% accuracy on the FrontierMath benchmark, a challenging mathematical test developed by Epoch AI
1
. This score far surpassed previous high scores of just 2% from other powerful models, marking a significant leap in AI capabilities3
.The celebration of o3's performance was short-lived as it came to light that OpenAI had played a significant role in the creation of the FrontierMath benchmark. Epoch AI, the nonprofit behind FrontierMath, revealed that OpenAI had funded the benchmark's development and had access to a large portion of the problems and solutions
1
.This disclosure raised concerns about the validity of OpenAI's results and the transparency of the benchmarking process. Tamay Besiroglu, associate director at Epoch AI, admitted that they were contractually restricted from disclosing OpenAI's involvement until the o3 model was launched
3
.The lack of transparency extended to the mathematicians who contributed to FrontierMath. Six mathematicians confirmed they were unaware that OpenAI would have exclusive access to the benchmark
3
. This revelation led to regret among some contributors who might not have participated had they known about OpenAI's involvement2
.OpenAI maintains that it didn't directly train o3 on the benchmark and that some problems were "strongly held out"
1
. Epoch AI acknowledged the mistake in not being more transparent about OpenAI's involvement and committed to implementing a "hold out set" of 50 randomly selected problems to be withheld from OpenAI for future testing1
4
.Related Stories
This controversy highlights the challenges in creating truly independent evaluations for AI models. Experts argue that ideal testing would require a neutral sandbox, which is difficult to realize
1
. The incident has drawn comparisons to the Theranos scandal, with some AI experts questioning the legitimacy of OpenAI's claims3
.The FrontierMath controversy underscores the complexities of AI benchmarking and the need for greater transparency in the development and testing of AI models. It raises important questions about how to balance the need for resources in benchmark development with maintaining the integrity and objectivity of the evaluation process
4
.Summarized by
Navi
[2]
21 Apr 2025•Technology

12 Nov 2024•Technology

11 Apr 2025•Technology

1
Science and Research

2
Technology
3
Policy and Regulation
