10 Sources
[1]
DeepSeek tests "sparse attention" to slash AI processing costs
Ever wonder why ChatGPT slows down during long conversations? The culprit is a fundamental mathematical challenge: processing long sequences of text requires massive computational resources, even with the efficiency tricks that companies have already deployed. While US tech giants can afford to
[2]
DeepSeek releases 'sparse attention' model that cuts API costs in half | TechCrunch
Researchers at DeepSeek on Monday released a new experimental model called V3.2-exp, designed to have dramatically lower inference costs when used in long-context operations. DeepSeek announced the model with a post on Hugging Face, also posting a linked academic paper on GitHub. The most
[3]
DeepSeek Debuts 'Sparse Attention' Method in Next-Gen AI Model
DeepSeek updated an experimental AI model Monday in what it called a step toward next-generation artificial intelligence. The secretive Chinese startup outlined the DeepSeek-V3.1-Exp platform, explaining it uses a new technique it calls DeepSeek Sparse Attention or DSA, according to a post on its
[4]
DeepSeek releases model it calls 'intermediate step' towards 'next-generation architecture'
BEIJING, Sept 29 (Reuters) - Chinese AI developer DeepSeek has released its latest model which it said was an "experimental release" that was more efficient to train and better at processing long sequences of text than previous iterations. The Hangzhou-based company called DeepSeek-V3.2-Exp an
[5]
China's DeepSeek launches next-gen AI model. Here's what makes it different
"This is DeepSeek's value prop all over: efficiency is becoming as important as raw power," according to Nick Patience, VP and Practice Lead for AI at The Futurum Group. Chinese startup DeepSeek's latest experimental model promises to increase efficiency and improve AI's ability to handle a lot of
[6]
DeepSeek's new V3.2-Exp model cuts API pricing in half to less than 3 cents per 1M input tokens
DeepSeek continues to push the frontier of generative AI...in this case, in terms of affordability. The company has unveiled its latest experimental large language model (LLM), DeepSeek-V3.2-Exp, that mostly matches or slightly improves the benchmarks of its predecessor DeepSeek-3.1-Terminus, but
[7]
Here's what we know about DeepSeek's latest AI offering
The announcement comes as China pressures its tech companies to break their reliance on foreign chip makers so it can compete in the AI race. The Chinese artificial intelligence (AI) company DeepSeek has launched its latest experimental model, which claims to handle a large amount of data and
[8]
DeepSeek Has 'Cracked' Cheap Long Context for LLMs With Its New Model | AIM
DeepSeek-V3.2-Exp is claimed to achieve 'significant efficiency improvements in both training and inference'. DeepSeek, the China-based AI lab, has released DeepSeek-V3.2-Exp, an experimental AI model on September 29. The model is claimed to achieve 'significant efficiency improvements in both
[9]
Deepseek 3.2 : New AI Model is Faster, Cheaper and Smarter
What if artificial intelligence could process information faster, cost less, and still deliver unparalleled accuracy? With the release of Deepseek 3.2 Experimental, that vision is no longer hypothetical. Building on the foundation of its predecessor, Deepseek 3.1 Terminus, this innovative iteration
[10]
China's DeepSeek releases 'intermediate' AI model on route to next generation
BEIJING -- Chinese AI developer DeepSeek has released its "experimental" latest model, which it said was more efficient to train and better at processing long sequences of text than previous iterations of its large language models. The Hangzhou-based company called DeepSeek-V3.2-Exp an
Share
Copy Link
Chinese AI company DeepSeek has released an experimental version of its latest language model, DeepSeek-V3.2-Exp, featuring a new 'sparse attention' technique that promises to significantly reduce processing costs for long-context AI operations.
Chinese AI company DeepSeek has made waves in the artificial intelligence community with the release of its experimental model, DeepSeek-V3.2-Exp. This latest iteration introduces a novel technique called 'DeepSeek Sparse Attention' (DSA), which promises to dramatically reduce processing costs for long-context AI operations
1
.
Source: Reuters
AI language models have long grappled with the computational challenges of processing extensive sequences of text. The traditional 'attention' mechanism, which helps models understand context by relating each word to every other word in a sequence, becomes increasingly resource-intensive as the text length grows. This quadratic growth in computational requirements has been been a significant bottleneck for AI performance in long conversations
1
.DeepSeek's sparse attention approach tackles this issue by selectively processing only the most relevant word relationships, rather than examining every possible connection. The model employs a 'lightning indexer' to identify the top 2,048 most important connections for each word, significantly reducing the computational load without compromising understanding
1
2
.
Source: VentureBeat
The efficiency gains from this new architecture are substantial. DeepSeek claims that its sparse attention technique has enabled them to cut API prices by 50% for long-context operations
2
. This dramatic reduction in processing costs could make powerful AI more accessible to developers, researchers, and smaller companies, potentially spurring a new wave of innovative applications5
.Related Stories
This latest release builds on DeepSeek's reputation for efficiency-focused AI development. The company previously garnered attention with its R1 model, which reportedly matched OpenAI's performance while costing only $6 million to train
1
. DeepSeek's approach to AI development, emphasizing cost-effectiveness and efficiency, has positioned it as a notable player in the global AI landscape5
.
Source: AIM
The release of DeepSeek-V3.2-Exp is seen as an intermediate step towards the company's next-generation AI architecture
3
4
. As an open-weight model available on Hugging Face, it invites third-party testing and validation of its performance claims2
.While the full impact of DeepSeek's sparse attention technique remains to be seen, it represents a significant step forward in addressing the crucial challenge of inference costs in AI. As the industry continues to evolve, innovations like these could play a pivotal role in shaping the future of AI technology and its applications across various sectors.
Summarized by
Navi
[1]
[4]
1
Science and Research

2
Policy and Regulation

3
Technology