12 Sources
[1]
DeepSeek's new model sets a template for powerful LLMs that run lean
Chinese AI darling DeepSeek unveiled an updated version of its cost-and-latency-optimized Flash model on Thursday, with a new version 4.1 that includes architectural improvements more significant than you would expect in a point release because the changes might open the door to larger, smarter,
[2]
DeepSeek's V4.1-Flash takes on Opus 5, and retires V4-Pro
DeepSeek V4.1-Flash is a 552bn-parameter open model that DeepSeek says beats its own flagship at a lower price. From Monday, it takes over every V4-Pro request. Its benchmarks rival Opus 5 on coding, but trail on hard reasoning, and none are independently verified. DeepSeek released V4.1-Flash on
[3]
DeepSeek-V4.1-Flash debuts with $0.003/1M off-peak cached-input rate and benchmarks eclipsing GPT-5.6 Sol, Claude Opus 5
DeepSeek launched DeepSeek-V4.1-Flash last night with a 552-billion-parameter mixture-of-experts backbone, native vision, a 1-million-token context window and an architecture built to make repeatedly reading large contexts cheaper. For developers evaluating the model for coding agents and other
[4]
DeepSeek's New Model Nearly Matches GPT-6 Astra on Design -- at 1.4% of the Cost
DeepSeek's technical paper for V4.1 Flash shows the model activates just 8 billion of its 552 billion parameters to read a prompt, the design choice behind its low price. OpenDesign, the company behind the benchmark site OpenDesign Arena, ran 13 AI models through the same batch of design tasks
[5]
DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro
DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro Chinese artificial intelligence startup Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co. Ltd. today released DeepSeek-V4.1-Flash, the smallest model in a new architecture family. The company said tests by
[6]
DeepSeek launches V4.1-Flash with 1M-token context
DeepSeek has launched DeepSeek-V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native vision, a 1 million-token context window and pricing built around low cached-input costs. During off-peak hours, DeepSeek prices the model at $0.003 per million input tokens on a cache hit,
[7]
DeepSeek: China's DeepSeek launches V4.1-Flash model
DeepSeek said the new model is designed for greater capability, faster inference, higher throughput and scaling to larger models. Chinese artificial intelligence startup DeepSeek on Thursday launched DeepSeek-V4.1-Flash, which the company said is the smallest model in its new architecture
[8]
How DeepSeek V4.1 Achieves a 437x Memory Footprint Drop
DeepSeek V4.1 introduces a significant breakthrough in AI memory efficiency, achieving a 437-fold reduction in memory usage compared to its previous version. As detailed by The Stack, this model operates with just 890 bytes per token, allowing it to handle a million-token context while maintaining
[9]
DeepSeek V4.1 Flash Beats OpenAI's GPT-5.6 Sol And Anthropic's Opus 5 On Coding And Cybersecurity At An ~86x Lower Cost, While Reducing HBM Requirements By 3.8x And SSD Ones By 8x
DeepSeek's engineers are a marvel, excelling in extracting every ounce of efficiency from the architectural constraints that characterize contemporary LLMs, as they appear to have done with the just-released V4.1 Flash, which is phenomenally competitive with the likes of OpenAI's GPT-5.6 Sol and
[10]
DeepSeek Escalates AI Price War With Cheaper V4.1 Flash Model, Puts Pressure on OpenAI and Anthropic
The aggressive pricing could put additional pressure on major US players such as OpenAI and Anthropic, while also intensifying competition among China's growing field of AI developers. Chinese AI startup DeepSeek has intensified the global artificial intelligence price war with the launch of its
[11]
DeepSeek V4.1 Flash Outperforms Opus 5 in New AI Benchmarks
DeepSeek V4.1 Flash represents a significant step forward in AI development, emphasizing both efficiency and accessibility. As outlined by Universe of AI, the model features a 552-billion-parameter Mixture of Experts architecture, which dynamically adjusts computational resources to optimize
[12]
9 Things to Know About DeepSeek's New 552B AI Model
DeepSeek V4.1 Flash: DeepSeek's new 552-billion-parameter model uses a Mixture-of-Experts architecture designed to balance capability with efficient computing. . 552B Parameters: The model contains 552 billion parameters, placing it among the largest AI models built for advanced general-purpose
Share
Copy Link
DeepSeek launched V4.1 Flash, a 552-billion-parameter AI model that beats its own flagship on coding and agent tasks while slashing operational costs by up to 70%. The mixture-of-experts model introduces causal encoder-decoder architecture that activates just 8 billion parameters for input processing, achieving performance comparable to GPT-6 Astra at 1.4% of the cost.
DeepSeek unveiled V4.1 Flash on Thursday, a 552-billion-parameter AI model that the Hangzhou-based lab claims outperforms its own flagship V4-Pro on coding, agent tasks, and cost efficiency
1
. Starting September 14 at 04:00 UTC, every request made to V4-Pro through DeepSeek's API will automatically route to V4.1 Flash and bill at the cheaper rates until a V4.1-Pro arrives2
. The open-weight AI model is available on Hugging Face under the MIT license, allowing developers to download, modify, and deploy it freely2
.
Source: VentureBeat
The mixture-of-experts AI model represents more than a point release. At 763 billion total parameters including 196 billion N-gram parameters, V4.1 Flash is 2.5 times larger than the V4 Flash it replaces and exceeds the size of the V3 and R1 models that put DeepSeek on the map in early 2025
1
. Despite this massive parameter count, the model activates only 8 billion parameters when reading input and 16 billion when writing output2
.The causal encoder-decoder architecture enables V4.1 Flash to handle up to 1 million tokens of context while dramatically reducing resource consumption
2
. DeepSeek achieved reduced KV cache consumption of just 890 bytes per token, roughly 25% of what V4 Flash required and approximately 437 times less than its first model in 20232
. This reduction means the model can support four to eight times as many users in the same memory footprint1
.The technical breakthrough centers on N-gram parameters that form a conditional memory module. These 196 billion parameters decouple memory from computation, allowing the AI model to provide smarter responses without the performance penalty typically associated with additional parameters
1
. The N-gram weights function as enormous lookup tables that quickly surface relevant information through cheap lookups rather than forcing the model to calculate token probabilities from scratch1
.DeepSeek's benchmark table positions V4.1 Flash competitively against leading closed models from Anthropic and OpenAI. On DeepSWE v1.1, a software engineering benchmark, the model scored 74.2% compared to Claude Opus 5 at 74.0% and GPT-5.6 Sol at 73.0%
2
. On CyberGym cybersecurity testing, V4.1 Flash reached 88.1, ahead of every rival with a listed score2
. The model also scored 90.6 on Terminal-Bench 2.1, narrowly ahead of Opus 5 at 89.1 and GPT-5.6 Sol at 88.85
.
Source: The Next Web
OpenDesign Arena testing revealed V4.1 Flash achieved 98% of GPT-6 Astra's design score while charging approximately 1.4% of the cost
4
. On everyday design tasks including web apps, dashboards, and mobile screens, GPT-6 Astra averaged 82.7 points at $1.61 per design, while V4.1 Flash scored 81.2 at $0.023 per design4
. The model completed designs in 5.3 minutes compared to Astra's 11.1 minutes4
.Gaps remain on harder reasoning tasks. On Humanity's Last Exam, V4.1 Flash scored 36.8 compared to Claude Opus 5's 56.3
2
. Both U.S. models still lead on GPQA Diamond science reasoning benchmarks5
.Related Stories
DeepSeek announced pricing cuts of up to 32% from its August rates, with off-peak API pricing set at $0.15 per million uncached input tokens and $0.60 per million output tokens
3
. Peak rates, which apply weekdays from 01:00-04:00 UTC and 06:00-10:00 UTC, double those figures3
.The cached-input rates transform long-running AI agent workflows economics. Off-peak cached input costs just $0.003 per million tokens, compared to $0.40 for GPT-5.6 Sol, $0.50 for Claude Opus 5, and $0.30 for Kimi K3
3
. For an agent retaining a 500,000-token reusable prefix across 100 requests, cache reads would cost approximately $0.15 on V4.1 Flash off-peak versus $15 on Kimi K3, $20 on GPT-5.6 Sol, and $25 on Claude Opus 53
.Developers still calling V4-Pro will see output costs drop from $3.96 per million tokens at peak to $1.20 for V4.1 Flash, representing a roughly 70% reduction
5
. This aggressive pricing pressured Chinese competitors, with MiniMax and Z.ai shares falling more than 8% in Hong Kong and Alibaba sliding over 2% following the announcement2
.DeepSeek is positioning V4.1 Flash to compete in the AI agentic stack beyond just model provision. The company invited operators planning deployments of 2,000 GPUs with storage clusters to collaborate directly
2
. Coding tools WorkBuddy and OpenCode already support the model2
. The company is also recruiting engineers in Beijing to build its own Code Harness, aiming to own the full agentic infrastructure rather than simply supplying the underlying model4
.
Source: The Register
The launch arrives as DeepSeek prepares for a listing on Shanghai's STAR Market, following IPO preparations reported in July
2
. Founder Liang Wenfeng reportedly contributed $3 billion to a funding round exceeding $7.4 billion in June that valued the company above $50 billion5
. The same day as the V4.1 Flash launch, Anthropic named DeepSeek in its threat intelligence report as one of seven China-based labs running distillation campaigns against Claude, attributing more than 12.1 million exchanges over 14 days in July to DeepSeek5
.Watch whether enterprises adjust procurement strategies around cache-hit economics and whether DeepSeek's architectural innovations force U.S. labs to rethink their own mixture-of-experts designs. The model's ability to deliver frontier-class performance at dramatically lower operational costs may accelerate adoption of open-weight AI models in production environments where cost per completed task matters more than benchmark leaderboard position.
Summarized by
Navi
[2]
[3]
24 Apr 2026•Technology

04 Aug 2026•Technology

29 Sept 2025•Technology
