6 Sources
[1]
OpenAI breaks barriers with the o3 model: is general artificial intelligence near? - Softonic
o3 could mark the beginning of a new era in artificial intelligence On December 20th, OpenAI revealed that its model o3 achieved an 88% score on the demanding ARC-AGI benchmark, far surpassing the previous record of 55% and reaching the level of the human average. The company's major breakthrough
[2]
OpenAI's o3 system has reached human level on a test for 'general intelligence'
A new artificial intelligence (AI) model has just achieved human-level results on a test designed to measure "general intelligence". On December 20, OpenAI's o3 system scored 85% on the ARC-AGI benchmark, well above the previous AI best score of 55% and on par with the average human score. It also
[3]
OpenAI Claims Its New Model Reached Human Level on a Test for â€~General Intelligence.’ What Does That Mean?
A new artificial intelligence (AI) model has just achieved human-level results on a test designed to measure “general intelligenceâ€. On December 20, OpenAI’s o3 system scored 85% on the ARC-AGI benchmark, well above the previous AI best score of 55% and on par with the average human score. It
[4]
An AI system has reached human level on a test for 'general intelligence' -- here's what that means
by Michael Timothy Bennett and Elija Perrier, The Conversation A new artificial intelligence (AI) model has just achieved human-level results on a test designed to measure "general intelligence." On December 20, OpenAI's o3 system scored 85% on the ARC-AGI benchmark, well above the previous AI
[5]
An AI system has reached human level on a test for 'general intelligence': here's what that means
While scepticism remains, many AI researchers and developers feel something just changed. For many, the prospect of AGI now seems more real, urgent and closer than anticipated. Are they right?A new artificial intelligence (AI) model has just achieved human-level results on a test designed to
[6]
An AI system has reached human level on a test for 'general intelligence'. Here's what that means
Australian National University provides funding as a member of The Conversation AU. A new artificial intelligence (AI) model has just achieved human-level results on a test designed to measure "general intelligence". On December 20, OpenAI's o3 system scored 85% on the ARC-AGI benchmark, well
Share
Copy Link
OpenAI's o3 model scores 85-88% on the ARC-AGI benchmark, matching human-level performance and surpassing previous AI systems, raising questions about progress towards artificial general intelligence (AGI).

On December 20th, OpenAI announced that its new o3 model had achieved a remarkable score of 85-88% on the ARC-AGI benchmark, a test designed to measure "general intelligence"
1
2
3
. This result not only surpasses the previous AI best score of 55% but also reaches the level of average human performance on the test1
2
.The ARC-AGI benchmark, created by French AI researcher Francois Chollet, evaluates an AI system's "sample efficiency" in adapting to new situations
2
3
. It presents visual problems in the form of grid patterns, where the AI must deduce transformation rules from just three examples and apply them correctly to a fourth case1
4
. This test is considered crucial for measuring an AI's ability to generalize and solve novel problems with limited data, a key aspect of intelligence2
3
.The o3 model's achievement is noteworthy for several reasons:
Sample Efficiency: Unlike models like GPT-4, which rely on vast amounts of training data, o3 demonstrates remarkable adaptability with very few examples
1
2
.Generalization Capability: The model's performance suggests it can find and apply "weak rules" - simple, general norms that maximize adaptability to new situations
1
4
.Potential AGI Implications: This breakthrough has reignited discussions about the proximity of artificial general intelligence (AGI), with some researchers viewing it as a significant step towards this goal
2
3
.While OpenAI has not disclosed detailed information about o3's architecture, experts have theories about its approach:
Chain of Thought Searching: Chollet believes o3 might search through different "chains of thought" to solve tasks, similar to how Google's AlphaGo system analyzed move sequences in the game of Go
2
3
.Heuristic Optimization: The model may use a heuristic to choose the best solution from multiple valid options, possibly favoring simpler or more generalizable rules
2
3
.Related Stories
Despite the excitement, several caveats remain:
Limited Disclosure: OpenAI has only shared initial test results and conducted private presentations, leaving many aspects of o3 unknown
2
3
5
.Specialized Training: The model was specifically trained for the ARC-AGI test, raising questions about its general applicability
2
3
.Need for Further Evaluation: Comprehensive testing is required to understand o3's full capabilities, limitations, and failure rates
2
3
5
.If o3 proves to be as adaptable as an average human across various tasks, it could have far-reaching implications:
Economic Impact: The technology could potentially revolutionize numerous industries and accelerate AI-driven innovation
2
3
5
.AGI Benchmarks: New standards may be needed to evaluate and define artificial general intelligence
2
3
.Governance Considerations: The rapid progress may necessitate serious discussions about AI governance and safety measures
2
3
5
.As the AI community awaits more information and the eventual release of o3, the debate continues on whether this breakthrough truly brings us closer to AGI or represents a more limited, albeit impressive, advancement in AI capabilities
2
3
5
.Summarized by
Navi
[1]
[4]
1
Technology

2
Technology

3
Policy and Regulation
