2 Sources
[1]
AI can learn to show its workings through trial and error
You have full access to this article via Jozef Stefan Institute. When a student encounters a challenging mathematics problem or a programmer needs to write a complex algorithm, they will rarely solve it all in one go. Instead, they will reason through the task, jotting down notes and intermediate
[2]
DeepSeek bolsters AI 'reasoning' using trial-and-error
Chinese AI company DeepSeek has shown it can improve the reasoning of its LLM DeepSeek-R1 through trial-and-error based reinforcement learning, and even be made to explain its reasoning on math and coding problems, even though explanations might sometimes be unintelligible. The release of
Share
Copy Link
Chinese AI company DeepSeek has developed a novel approach to improve AI reasoning using reinforcement learning. Their model, DeepSeek-R1, demonstrates enhanced performance in math and coding tasks without relying on human examples.

In a groundbreaking development, Chinese AI company DeepSeek has introduced a novel method to enhance AI reasoning capabilities using reinforcement learning. The research, published in Nature, demonstrates how their large language model (LLM) DeepSeek-R1 can learn to reason and explain its thought process without relying on human examples
1
.DeepSeek's approach leverages reinforcement learning, a technique akin to how children learn through trial and error. This method contrasts with traditional prompting-based or supervised learning approaches, which rely heavily on human input or examples
2
.The model is rewarded for correct answers and penalized for incorrect ones, particularly in mathematics and programming tasks where answers are easily verifiable. This process naturally encourages the AI to develop its own reasoning strategies and output its thought process
1
.During training, DeepSeek-R1 exhibited interesting behaviors:
1
.However, the approach has limitations. The model occasionally produces extremely long reasoning traces and struggles with nuanced or subjective questions
2
.Related Stories
Despite these challenges, DeepSeek-R1 has achieved state-of-the-art accuracy in tasks assessing mathematics, coding skills, factual knowledge, and language understanding in both Chinese and English
2
.The release of DeepSeek-R1 in January 2025 had a significant impact on the AI market, causing a $589 billion decrease in Nvidia's market value. Investors viewed it as a potential cheaper alternative to systems like OpenAI's ChatGPT
2
.This research opens new avenues for AI development, potentially reducing the need for extensive human input in training advanced language models. As AI continues to evolve, DeepSeek's approach could lead to more efficient and capable AI systems, particularly in fields requiring complex reasoning and problem-solving skills.
Summarized by
Navi
[2]
1
Technology

2
Policy and Regulation

3
Technology
