7 Sources
[1]
Secrets of DeepSeek AI model revealed in landmark paper
The success of DeepSeek's powerful artificial intelligence (AI) model R1 -- that made the US stock market plummet when it was released in January -- did not hinge on being trained on the output of its rivals, researchers at the Chinese firm have said. The statement came in documents released
[2]
Secrets of Chinese AI Model DeepSeek Revealed in Landmark Paper
The first peer-reviewed study of the DeepSeek AI model shows how a Chinese start-up firm made the market-shaking LLM for $300,000 The success of DeepSeek's powerful artificial intelligence (AI) model R1 -- that made the US stock market plummet when it was released in January -- did not hinge on
[3]
DeepSeek didn't really train its flagship model for $294,000
Training costs detailed in R1 training report don't include 2.79 million GPU hours that laid its foundation Chinese AI darling DeepSeek's now infamous R1 research report was published in the Journal Nature this week, alongside new information on the compute resources required to train the model.
[4]
In rare disclosure, DeepSeek claims R1 model training cost just $294K
Serving tech enthusiasts for over 25 years. TechSpot means tech analysis and advice you can trust. Bottom line: China's DeepSeek has released detailed cost figures for training its R1 artificial intelligence model, providing rare insight into its development and drawing renewed scrutiny of the
[5]
We Finally Know How Much It Cost to Train China's Astonishing DeepSeek Model
Remember when DeepSeek briefly shook up the entire artificial intelligence industry by launching its large language model, R1, that was trained for a fraction of the money that OpenAI and other big players were pouring into their models? Thanks to a new paper published by the DeepSeek AI team in
[6]
DeepSeek releases R1 model trained for $294,000 on 512 H800 GPUs
The Chinese company DeepSeek AI has released its large language model, R1, which was trained for only $294,000 using 512 Nvidia H800 GPUs. In a paper published in the journal Nature, the company detailed how it achieved this low cost by using a trial-and-error reinforcement learning method,
[7]
China's DeepSeek says its hit AI model cost just $294,000 to train - The Economic Times
Chinese AI firm DeepSeek revealed its R1 model was trained for just $294,000 using 512 Nvidia H800 chips, far below US rivals' costs. The disclosure revives debates on China's AI progress, export restrictions, and transparency, with skepticism over DeepSeek's true access to banned Nvidia
Share
Copy Link
Chinese AI startup DeepSeek reveals groundbreaking training methods and costs for its R1 model in a peer-reviewed Nature paper, sparking debates over efficiency and transparency in AI development.
Chinese AI startup DeepSeek has made waves in the artificial intelligence community with the publication of a peer-reviewed paper in Nature, detailing the development of their R1 model. This landmark study marks the first major large language model (LLM) to undergo the rigorous peer-review process, setting a new precedent for transparency in AI research
1
2
.
Source: Nature
DeepSeek's primary innovation lies in its use of pure reinforcement learning to create R1. This automated trial-and-error approach rewards the model for reaching correct answers, rather than following human-selected reasoning examples. The process allowed R1 to develop its own reasoning-like strategies, including self-verification methods
1
2
.One of the most striking claims in the paper is the reported training cost of just $294,000 for R1. This figure, based on 512 Nvidia H800 GPUs running for 198 hours, is substantially lower than the tens of millions of dollars typically associated with training competitive AI models
3
4
.However, this claim has been met with skepticism. Critics argue that the $294,000 figure only accounts for the final reinforcement learning phase, not the entire training process. When including the development of the base V3 model, which required 2.79 million GPU hours, the total cost rises to approximately $5.87 million
3
.Related Stories
Despite the cost controversy, R1's performance has been impressive. It has become the most popular open-weight model on the AI community platform Hugging Face, with 10.9 million downloads. In scientific task challenges, R1 has proven to be highly competitive, particularly in balancing ability with cost
1
.
Source: ET
The paper also addresses concerns about DeepSeek's training data sources. While acknowledging that R1's base model was trained on web data, which may have included AI-generated content, the researchers deny deliberately using outputs from rival models like OpenAI's
1
2
.The publication of this paper in Nature has been widely welcomed as a step towards greater transparency in AI development. It sets a precedent that other firms may be encouraged to follow, potentially leading to more open evaluation of AI systems and their associated risks
1
2
.
Source: Gizmodo
As researchers continue to explore and apply DeepSeek's methods, the R1 model's influence is likely to grow, potentially revolutionizing how reasoning capabilities are developed in future AI systems
5
.Summarized by
Navi
[2]
[3]
1
Technology

2
Policy and Regulation

3
Technology
