7 Sources
[1]
OpenAI's most capable models hallucinate more than earlier ones
Researchers say the hallucinations make o3 "less useful" than it would be. OpenAI says its latest models, o3 and o4-mini, are its most powerful yet. However, research shows the models also hallucinate more -- at least twice as much as earlier models. Also: How to use ChatGPT: A beginner's guide
[2]
OpenAI's newest o3 and o4-mini models excel at coding and math - but hallucinate more often
A hot potato: OpenAI's latest artificial intelligence models, o3 and o4-mini, have set new benchmarks in coding, math, and multimodal reasoning. Yet, despite these advancements, the models are drawing concern for an unexpected and troubling trait: they hallucinate, or fabricate information, at
[3]
OpenAI's newest AI models hallucinate way more, for reasons unknown
This does not bode well if you're using the new o3 and o4-mini reasoning models for factual answers. Last week, OpenAI released its new o3 and o4-mini reasoning models, which perform significantly better than their o1 and o3-mini predecessors and have new capabilities like "thinking with images"
[4]
OpenAI's leading models keep making things up -- here's why
OpenAI's newly released o3 and o4-mini are some of the smartest AI models to ever be released, but they seem to be suffering from one major problem. Both models are hallucinating. This in itself isn't out of the ordinary, as most AI models still tend to do this. But these two new versions seem to
[5]
OpenAI's Hot New AI Has an Embarrassing Problem
OpenAI launched its latest AI reasoning models, dubbed o3 and o4-mini, last week. According to the Sam Altman-led company, the new models outperform their predecessors and "excel at solving complex math, coding, and scientific challenges while demonstrating strong visual perception and
[6]
It's not your imagination -- ChatGPT models actually do hallucinate more now
Table of Contents Table of Contents What do the tests say? What are AI "hallucinations" and why do they happen? What's the fix? OpenAI released a paper last week detailing various internal tests and findings about its o3 and o4-mini models. The main differences between these newer models and the
[7]
New OpenAI Models Hallucinating More Than their Predecessor
OpenAI's new artificial intelligence (AI) reasoning models, o3 and o4-mini, are hallucinating more than their predecessor, the company's internal testing report reveals. OpenAI launched the two new reasoning models, designed to pause and work through questions before responding, earlier this month,
Share
Copy Link
OpenAI's new o3 and o4-mini models show improved performance in various tasks but face a significant increase in hallucination rates, raising concerns about their reliability and usefulness.

OpenAI has released its latest AI models, o3 and o4-mini, touting significant improvements in coding, math, and multimodal reasoning capabilities
2
. These new "reasoning models" are designed to handle more complex tasks and provide more thorough, higher-quality answers1
. According to OpenAI, the models excel at solving complex math, coding, and scientific challenges while demonstrating strong visual perception and analysis5
.Despite their advanced capabilities, o3 and o4-mini have shown a concerning trend: they hallucinate, or fabricate information, at higher rates than their predecessors
1
2
3
. This development breaks the historical pattern of decreasing hallucination rates with each new model release2
.OpenAI's internal testing using the PersonQA benchmark revealed:
1
2
3
1
2
3
The exact reasons for this increase in hallucinations remain unclear, even to OpenAI's researchers
1
2
. Some hypotheses include:1
.2
.These hallucinations pose significant risks for industries where accuracy is crucial, such as law and finance
2
. Sarah Schwettmann, co-founder of Transluce, warns that the higher hallucination rate could limit o3's usefulness in real-world applications2
.Related Stories
Researchers have observed concerning behaviors in the new models:
1
.2
5
.5
.OpenAI acknowledges the challenge, stating that addressing hallucinations "across all our models is an ongoing area of research"
2
5
. The company is exploring potential solutions, including:2
.1
3
.As the AI industry shifts focus towards reasoning models, the experience with o3 and o4-mini highlights the need for balanced progress in both capabilities and reliability
2
. For now, users are advised to remain cautious and fact-check AI-generated information, especially when using these latest-generation reasoning models1
.Summarized by
Navi
[2]
[4]
[5]
05 May 2025•Technology

08 Sept 2025•Science and Research

22 Mar 2025•Technology

1
Science and Research

2
Policy and Regulation

3
Technology