7 Sources
[1]
How DeepSeek's new way to train advanced AI models could disrupt everything - again
Just before the start of the new year, the AI world was introduced to a potential game-changing new method for training advanced models. A team of researchers from Chinese AI firm DeepSeek released a paper on Wednesday outlining what it called Manifold-Constrained Hyper-Connections, or mHC for
[2]
DeepSeek Touts New Training Method as China Pushes AI Efficiency
DeepSeek published a paper outlining a more efficient approach to developing AI, illustrating the Chinese artificial intelligence industry's effort to compete with the likes of OpenAI despite a lack of free access to Nvidia Corp. chips. The document, co-authored by founder Liang Wenfeng,
[3]
DeepSeek develops mHC AI architecture to boost model performance - SiliconANGLE
DeepSeek develops mHC AI architecture to boost model performance DeepSeek researchers have developed a technology called Manifold-Constrained Hyper-Connections, or mHC, that can improve the performance of artificial intelligence models. The Chinese AI lab debuted the software in a paper published
[4]
New DeepSeek Research Shows Architectural Fix Can Boost Reasoning at Scale
DeepSeek's findings highlight a path to stronger AI reasoning without relying on ever-larger models. DeepSeek has released new research showing that a promising but fragile neural network design can be stabilised at scale, delivering measurable performance gains in large language models without
[5]
DeepSeek introduces Manifold-Constrained Hyper-Connections for R2
Just before the start of the new year, the artificial intelligence community was introduced to a potential breakthrough in model training. A team of researchers from the Chinese AI firm DeepSeek released a paper outlining a novel architectural approach called Manifold-Constrained Hyper-Connections,
[6]
DeepSeek's New Architecture Can Make AI Model Training More Efficient
DeepSeek, the Chinese artificial intelligence (AI) startup, that took the Silicon Valley by storm in November 2024 with its R1 AI model has now revealed a new architecture that can help bring down the cost and time taken to train large language models (LLMs). A new research paper has been published
[7]
DeepSeek touts new training method as China pushes AI efficiency
DeepSeek published a paper outlining a more efficient approach to developing AI, illustrating the Chinese artificial intelligence industry's effort to compete with the likes of OpenAI despite a lack of free access to Nvidia chips. The document, co-authored by founder Liang Wenfeng, introduces a
Share
Copy Link
Chinese AI firm DeepSeek has released research on Manifold-Constrained Hyper-Connections (mHC), a new AI training method that addresses signal degradation in neural networks. The framework could enable smaller developers to build frontier models without massive computational costs, potentially previewing the architecture behind the anticipated DeepSeek R2 model expected around February's Spring Festival.
DeepSeek has published research outlining Manifold-Constrained Hyper-Connections, or mHC, a new AI architecture that could fundamentally alter how engineers train advanced AI models
1
. The paper, released on arXiv and co-authored by founder Liang Wenfeng, introduces a framework designed to improve scalability while reducing the computational and energy demands of training advanced AI systems2
. This development comes as Chinese AI startups continue operating under significant constraints, with US restrictions preventing access to the most advanced semiconductors essential for developing and running AI2
.
Source: Gadgets 360
The Chinese AI startup gained prominence one year ago with its R1 model, which rivaled OpenAI's o1 capabilities but was reportedly trained at a fraction of the cost
1
. The new mHC paper could serve as the technological framework for DeepSeek R2, the company's next flagship system expected around the Spring Festival in February2
. The R2 model was originally anticipated in mid-2025 but was postponed due to China's limited access to advanced AI chips and concerns from CEO Liang Wenfeng about the model's performance1
.The mHC architecture addresses a critical technical challenge that has hindered the development of large language models. As neural networks grow deeper with additional layers, signals can become attenuated or degraded, increasing the risk they turn into noise. DeepSeek researchers describe this as optimizing the trade-off between plasticity and stability across many layers
1
.DeepSeek built upon Hyper-Connections, a framework introduced in 2024 by ByteDance researchers that diversifies channels through which neural network layers share information
1
. However, standard Hyper-Connections introduce risks of signal loss and come with high memory costs that make them difficult to implement at scale3
. The mHC approach constrains hyperconnectivity within a model, preserving informational complexity while sidestepping memory issues1
.
Source: SiliconANGLE
The architecture uses a manifold—a mathematical object—to maintain the stability of gradients while they travel between an AI model's layers
3
. This innovation restores training stability while preserving the benefits of richer internal routing, enabling models to train reliably up to 27 billion parameters4
.DeepSeek tested the new AI training method by training three large language models with 3 billion, 9 billion, and 27 billion parameters using mHC, then compared them against models trained with standard Hyper-Connections
3
. The mHC-powered models performed better across eight different AI benchmarks3
.On BIG-Bench Hard, a benchmark focused on complex, multi-step reasoning, accuracy rose from 43.8% to 51.0%
4
. AI model performance also improved on DROP, testing numerical and logical reasoning over long passages, and on GSM8K, a standard test of mathematical reasoning performance4
.Crucially, these gains came with only a 6.27% increase in hardware overhead during training
3
. This level of hardware efficiency makes the approach viable for production-scale models, particularly for developers with limited resources4
.Related Stories
The debut of the mHC framework suggests a shift in AI evolution. Industry wisdom has held that only the wealthiest companies can afford to build frontier models, but DeepSeek continues demonstrating that breakthroughs can be achieved through clever engineering rather than massive capital reserves
1
. By publishing this research openly through arXiv and Hugging Face, DeepSeek has made the method available to smaller developers, potentially democratizing access to advanced AI capabilities5
.Bloomberg Intelligence analysts Robert Lea and Jasmine Lyu note that DeepSeek's forthcoming R2 model has potential to upend the global AI sector again, despite recent gains by competitors
2
. The paper lists 19 authors, with Liang Wenfeng's name appearing last, reflecting his consistent role in steering DeepSeek's research agenda2
. The research incorporates "rigorous infrastructure optimization to ensure efficiency" and addresses challenges such as training instability and limited scalability2
. The technique holds promise "for the evolution of foundational models," the authors stated2
.Summarized by
Navi
1
Technology

2
Technology

3
Science and Research
