DeepSeek V4 Pro launches with fourfold price increase as AI demand strains infrastructure capacity

Reviewed byNidhi Govil

5 Sources

Share

DeepSeek officially released its V4 Pro AI model with significant price increases taking effect August 16. The V4 Pro will cost $3.96 per million output tokens at peak hours, up from $0.87, while introducing half-price off-peak rates. Despite the fourfold increase, DeepSeek remains cheaper than competitors like OpenAI's GPT-5.6 Sol at $30 per million tokens.

DeepSeek V4 Pro Launch Marks End of Ultra-Low Pricing Era

DeepSeek has officially launched DeepSeek-V4-Pro-0813, ending the promotional pricing that made the Chinese AI startup a disruptive force in the market

1

. The general availability version supersedes the preview model released in April and introduces substantial changes to the company's pricing structure

3

. Starting at 16:00 UTC on August 16, API pricing for the V4 model family will increase by notable margins, with some rates rising more than 1,100%

1

. The V4 Pro will cost $3.96 for 1 million output tokens at peak hours, more than four times the current rate of $0.87

2

. Input tokens are priced at $0.435 per million

4

. DeepSeek is introducing peak and off-peak rates to allocate resources more reasonably as AI demand strains infrastructure capacity

2

.

Peak and Off-Peak Rates Offer Pricing Flexibility

The new pricing structure includes off-peak rates at half the peak-hour price, encouraging developers toward more flexible workload scheduling

1

. During off-peak hours, V4 Pro output tokens will cost $1.98 per million, while the V4-Flash model will be available for $0.66 per million output tokens compared to $1.32 during peak hours

2

. The current API pricing was originally supposed to be a promotion ending on May 31, but DeepSeek announced that month it would make the discounted prices permanent before ultimately deciding to proceed with the price hikes

2

. DeepSeek had offered a 75% promotional discount on V4 Pro through May 5, making the pricing shift a notable reversal of the company's earlier strategy

3

.

Cost-Effective AI Models Still Undercut Premium Competitors

Source: Geeky Gadgets

Source: Geeky Gadgets

Despite the fourfold increase, DeepSeek remains significantly cheaper than many competitors in the AI model pricing landscape

2

. Kimi K3, developed by Chinese company Moonshot, costs $15 per 1 million output tokens, while Anthropic's Fable 5 charges $50 per million output tokens

3

. OpenAI's most advanced model, GPT-5.6 Sol, costs $30 for 1 million output tokens

2

. DeepSeek V4 Pro is up to 57 times cheaper than Fable 5 while delivering nearly identical results in key benchmark scores

4

. OpenAI's low-cost offering, GPT-5.6 Luna, costs $1.20, which is less than the DeepSeek V4-Flash at peak hours but more expensive than the off-peak rate

2

.

Agentic AI Performance Shows Significant Improvements

The V4-Pro-0813 focuses on agent capabilities, where AI systems use tools, execute code, and complete multi-step workflows without human intervention

3

. Benchmark scores demonstrate substantial improvements over the preview version across agentic AI performance metrics

5

. The model scored 87.9 on Terminal Bench 2.1 compared to 72.1 for V4 Pro Preview, while achieving 62.7 on DeepSWE and 61.5 on NL2Repo

3

5

. On the CyberGym benchmark, it achieved 83.3, narrowly surpassing Fable 5's score of 83.1

4

. The model also scored 74.1 on Toolathlon-Verified, 25.7 on Agents' Last Exam, and 31.8 on AutomationBench, outperforming Fable 5's 29.1

4

5

.

Enhanced Features Include OpenAI Responses API Support and Codex Integration

Source: Engadget

Source: Engadget

The V4 Pro API has been updated to work with the OpenAI Responses API format out of the box, providing native compatibility for developers

3

. Built-in Codex support allows for one-click setup, optimizing the model for coding workflows

5

. Thinking effort levels for both V4 Pro and V4-Flash have been expanded to three settings: low for simple tasks, high for daily agent workflows, and max for complex tasks requiring more deliberation

3

5

. The model is capable of handling a context window of up to 1 million tokens and can produce outputs as long as 384,000 tokens, with the option to run in either thinking or non-thinking mode

3

.

DSpark Speculative Decoding Enhances Performance

DeepSeek-V4-Pro-0813 includes DSpark speculative decoding, a module built on the V4 Pro Preview architecture that improves inference efficiency

5

. Developers can enable DSpark with vLLM using seven speculative tokens with greedy draft sampling, or through SGLang without specifying a separate draft model path since target and draft weights come from the same checkpoint

5

. For local deployment, DeepSeek recommends temperature settings of 1.0 with top-p at 0.95 for agentic scenarios and top-p at 1.0 for other scenarios, with maximum output length of 384K tokens for high and max reasoning modes

5

.

Market Competition Intensifies as Rivals Close Gap

Source: InfoWorld

Source: InfoWorld

DeepSeek's R1 model went viral in early 2025 and briefly gave the company a commanding position in the AI race, but competitors including Moonshot AI, Alibaba, and ByteDance have since closed the gap with their own releases

3

. The launch comes after the cheaper V4-Flash model performed above expectations in independent tests, in some cases outpacing the April preview of V4 Pro, drawing attention given that Pro is positioned as the more capable product

3

. Models like SpaceX's Grock 4.6, priced at $2 per million input tokens and $6 per million output tokens, offer competitive performance and match GPT-5.6 Sol with a score of 61 on the Artificial Analysis Index

4

. The company secured roughly $7.4 billion in its first round of outside capital in June, and was reportedly in discussions for an additional round that would value it at around $74 billion

3

. The model weights are released under the MIT License, maintaining accessibility for developers

5

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved