17 Sources
[1]
Apple Says AI's Math Skills Fall Short | PYMNTS.com
Recent findings from Apple researchers have cast doubt on the mathematical prowess of large language models (LLMs), challenging the notion that artificial intelligence (AI) is on the brink of human-like reasoning. In a test of 20 state-of-the-art LLMs, performance on grade-school math problems
[2]
Apple Engineers Show How Flimsy AI 'Reasoning' Can Be
The new frontier in large language models is the ability to "reason" their way through problems. New research from Apple says it's not quite what it's cracked up to be. For a while now, companies like OpenAI and Google have been touting advanced "reasoning" capabilities as the next big step in
[3]
Top "Reasoning" AI Models Can be Brought to Their Knees With an Extremely Simple Trick
A team of Apple researchers has found that advanced AI models' alleged ability to "reason" isn't all it's cracked up to be. "Reasoning" is a word that's thrown around a lot in the AI industry these days, especially when it comes to marketing the advancements of frontier AI language models. OpenAI,
[4]
Apple's latest study proves that AI can't even solve basic grade-school math problems
Several Apple researchers have confirmed what had been previously thought to be the case regarding AI -- that there are serious logical faults in its reasoning, especially when it comes to basic grade school math. According to a recently published paper from six Apple researchers, 'GSM-Symbolic:
[5]
LLMs can't perform "genuine logical reasoning," Apple researchers suggest
What is going on inside that anthropomorphized digital brain? Credit: Getty Images For a while now, companies like OpenAI and Google have been touting advanced "reasoning" capabilities as the next big step in their latest artificial intelligence models. Now, though, a new study from six Apple
[6]
Apple Study Reveals Critical Flaws in AI's Logical Reasoning Abilities
Apple's AI research team has uncovered significant weaknesses in the reasoning abilities of large language models, according to a newly published study. The study, published on arXiv, outlines Apple's evaluation of a range of leading language models, including those from OpenAI, Meta, and other
[7]
Apple Researchers Suggest 'Fragile' AI Reasoning Capabilities Are Overstated
This suggests that AI models rely on "sophisticated pattern matching more than true logical reasoning" they concluded. According to commonly used benchmarks, frontier large language models (LLMs) have now surpassed the average human's ability to solve mathematical problems and perform complex
[8]
Reasoning failures highlighted by Apple research on LLMs
Apple plans to introduce its own version of AI starting with iOS 18.1 - image credit Apple A new paper from Apple's artificial intelligence scientists has found that engines based on large language models, such as those from Meta and OpenAI, still lack basic reasoning skills. The group has
[9]
Apple study reveals major AI flaw in OpenAI, Google, and Meta LLMs
Large Language Models (LLMs) may not be as smart as they seem, according to a study from Apple researchers. LLMs from OpenAI, Google, Meta, and others have been touted for their impressive reasoning skills. But research suggests their purported intelligence may be closer to "sophisticated pattern
[10]
Researchers question AI's 'reasoning' ability as models stumble on math problems with trivial changes
How do machine learning models do what they do? And are they really "thinking" or "reasoning" the way we understand those things? This is a philosophical question as much as a practical one, but a new paper making the rounds Friday suggests that the answer is, at least for now, a pretty clear
[11]
A New Apple Study Shows AI Reasoning Has Critical Flaws
It's no surprise that AI doesn't always get things right. Occasionally, it even hallucinates. However, a recent study by Apple researchers has shown even more significant flaws within the mathematical models used by AI for formal reasoning. As part of the study, Apple scientists asked an AI Large
[12]
Apple's Shocking AI Revelation: Are Language Models Just Pattern Machines?
Apple's recent research paper, "GSM Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models," challenges the perceived reasoning capabilities of current large language models (LLMs). The study suggests that these models primarily rely on pattern recognition rather
[13]
Artificial intelligence does not reason, according to Apple, but is there a solution? - Softonic
Apple's artificial intelligence research team has published an interesting paper on the weaknesses in the reasoning capabilities of language models. In the paper, available on arXiv (via Macrumors), the team explains how it evaluated a series of language models from different leading developers,
[14]
Apple says a high score on GSM8K dataset does not mean your AI is smarter
Recent research from Apple suggests that models that got a high score on the GSM8K dataset may not be as intelligent as they seem. Large Language Models (LLMs) have been widely praised for their seemingly impressive reasoning abilities. Models from companies like OpenAI, Google, and Meta are often
[15]
Apple agrees with Sam Altman: AI is incredibly dumb
Is ChatGPT getting smarter, or is it getting better at seeming smart? According to Apple, it's the latter. A team of AI researchers at Apple published a paper this weekend claiming that most leading large language AI models aren't actually capable of advanced reasoning, despite how intelligent
[16]
Apple researchers suggest artificial intelligence is still mostly an illusion
Researchers at Apple Computer Company have found evidence, via testing, showing that the seemingly intelligent responses given by AI-based LLMs are little more than an illusion. In their paper posted on the arXiv preprint server, the researchers argue that after testing several LLMs, they found
[17]
Apple Proves OpenAI o1 is Actually Good at Reasoning
While some say LLMs are our ticket to AGI, others think they're just glorified text-producing algorithms with a fancy name. Apple has gotten better at gaslighting AI companies that are spending all they have on making LLMs better at reasoning. A research team of six people at Apple recently
Share
Copy Link
A recent study by Apple researchers exposes significant flaws in the mathematical reasoning capabilities of large language models (LLMs), challenging the notion of AI's advanced reasoning skills and raising questions about their real-world applications.

A team of six Apple researchers has cast doubt on the mathematical prowess of large language models (LLMs), challenging the notion that artificial intelligence (AI) is approaching human-like reasoning capabilities. The study, titled "GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models," reveals significant weaknesses in AI systems when faced with tasks requiring robust logical reasoning
1
.The researchers utilized the GSM8K benchmark, a set of over 8,000 grade-school level mathematical word problems, to evaluate the performance of more than 20 state-of-the-art LLMs. They introduced two key modifications to the original benchmark:
The results were striking:
2
.3
.These findings suggest that current LLMs may not be capable of genuine logical reasoning. Instead, they appear to rely on pattern matching and replication of reasoning steps observed in their training data
4
.Dr. Selmer Bringsjord, professor at Rensselaer Polytechnic Institute, commented, "Any real-world application that requires reasoning of the sort that can be definitively verified (or not) is basically impossible for an LLM to get right with any degree of consistency"
1
.The implications of these limitations for AI applications in commerce and decision-making are significant. Financial institutions and other sectors relying on AI for complex calculations may need to reassess their use of these technologies
1
.However, not all experts view these limitations as equally problematic. Aravind Chandramouli, head of AI at Tredence, suggests that the impact on real-world applications may be minimal, as most do not require advanced mathematical reasoning
1
.Related Stories
Researchers and industry professionals are exploring several approaches to address these limitations:
1
.Eric Bravick, CEO of The Lifted Initiative, suggests that emerging technologies like retrieval-augmented generation (RAG) systems and multimodal AI could help address current limitations in AI reasoning
1
.This study emphasizes the need for more robust and adaptable evaluation methods for AI models. Lead study author Mehrdad Farajtabar stressed the importance of understanding LLMs' true reasoning capabilities for deploying them in real-world scenarios where accuracy and consistency are crucial
3
.As the field of AI continues to evolve, these findings highlight the significant work still needed to achieve artificial general intelligence (AGI) and underscore the importance of careful evaluation and testing of AI systems, particularly for high-stakes applications requiring reliable reasoning
5
.Summarized by
Navi
02 Nov 2024•Science and Research

09 Jun 2025•Science and Research

23 Jan 2026•Technology

1
Technology

2
Technology

3
Technology
