4 Sources
[1]
These AI models reason better than their open-source peers - but still can't rival humans
A study tested AI's ability to complete visual puzzles like those found on human IQ tests. It went poorly. Can artificial intelligence (AI) pass cognitive puzzles designed for human IQ tests? The results were mixed. Researchers from the USC Viterbi School of Engineering Information Sciences
[2]
Can advanced AI can solve visual puzzles and perform abstract reasoning?
Artificial Intelligence has learned to master language, generate art, and even beat grandmasters at chess. But can it crack the code of abstract reasoning -- those tricky visual puzzles that leave humans scratching their heads? Researchers at USC Viterbi School of Engineering Information Sciences
[3]
Can advanced AI can solve visual puzzles and perform abstract reasoning?
Artificial Intelligence has learned to master language, generate art, and even beat grandmasters at chess. But can it crack the code of abstract reasoning -- those tricky visual puzzles that leave humans scratching their heads? Researchers at USC Viterbi School of Engineering Information Sciences
[4]
Can AI Tackle Abstract Reasoning? Study Tests Cognitive Limits - Neuroscience News
Summary: Researchers tested artificial intelligence's ability to solve abstract visual puzzles similar to human IQ tests, revealing gaps in AI's reasoning skills. Open-source AI models struggled, while closed-source models like GPT-4V performed better, but far from perfectly. Detailed analyses
Share
Copy Link
A study by USC researchers reveals that AI models, particularly open-source ones, struggle with abstract visual reasoning tasks similar to human IQ tests. While closed-source models like GPT-4V perform better, they still fall short of human cognitive abilities.

Researchers from the USC Viterbi School of Engineering Information Sciences Institute (ISI) have conducted a groundbreaking study to assess the capabilities of artificial intelligence in solving abstract visual puzzles similar to those found in human IQ tests. The study, presented at the Conference on Language Modeling (COLM 2024) in Philadelphia, reveals significant limitations in AI's ability to perform nonverbal abstract reasoning tasks
1
.The research team, led by Kian Ahrabian and Zhivar Sourati, tested 24 different multi-modal large language models (MLLMs) using puzzles based on Raven's Progressive Matrices, a standard test of abstract reasoning. The results showed a stark contrast between open-source and closed-source AI models
2
.Open-source models performed poorly, with Ahrabian stating, "They were really bad. They couldn't get anything out of it." In contrast, closed-source models like GPT-4V demonstrated better performance, though still far from matching human cognitive abilities
3
.The researchers delved deeper to understand where the AI models were failing. They discovered that the issue was not limited to visual processing but extended to the reasoning process itself. Even when provided with detailed textual descriptions of the images, many models struggled to reason effectively
4
.Related Stories
To enhance AI performance, the team explored a technique called "Chain of Thought prompting." This method guides the AI through step-by-step reasoning tasks and led to significant improvements in some cases. Ahrabian noted, "By guiding the models with hints, we were able to see up to 100% improvement in performance"
2
.Jay Pujara, research associate professor and author of the study, emphasized the importance of understanding AI's limitations: "We still have such a limited understanding of what new AI models can do, and until we understand these limitations, we can't make AI better, safer, and more useful"
1
.The study's findings highlight both the current limitations of AI and the potential for future advancements. As AI models continue to evolve, this research could pave the way for developing AI systems that can not only understand but also reason in ways more comparable to human cognition
4
.Summarized by
Navi
[4]
25 Mar 2025•Science and Research

13 Oct 2024•Science and Research

22 Feb 2025•Science and Research

1
Technology

2
Technology

3
Technology
