Large language models achieve under 1% accuracy at basic multiplication, new study reveals
Despite handling complex coding tasks, large language models fail at four-digit multiplication, achieving less than 1% accuracy with standard training. Researchers from University of Chicago, MIT, Harvard, and Google DeepMind discovered the culprit: models can't store and retrieve intermediate computations. But a specialized Implicit Chain of Thought method achieved 100% accuracy by teaching models to internalize reasoning processes.