3 Sources
[1]
DeepSeek unveils vision model challenging Anthropic's Opus 4.8 performance
DeepSeek has launched an experimental multimodal model called DeepSeek-V4-Flash-Vision-Exp, which it claims approaches the performance of Anthropic's Claude Opus 4.8 on various benchmarks. This release signifies DeepSeek's continued efforts to enhance visual AI capabilities while maintaining lower
[2]
DeepSeek Version 4 Flash Vision Excels on Apex Bench Tests
The emergence of Aux Alpha and DeepSeek Version 4 Flash Vision has sparked significant interest within the AI community, with both models showcasing impressive advancements in multimodal capabilities. Aux Alpha, in particular, stands out for its remarkable 1-million-token context window and its
[3]
DeepSeek says its new multimodal model nears Anthropic's Opus 4.8: Here's how
DeepSeek just landed yet another punch in the arms race for AI's price/performance ratio - and this punch has eyes. Hangzhou-based company DeepSeek announced the release of the DeepSeek-V4-Flash-Vision-Exp - an experimental multimodal variation of its star performer, the V4 Flash. While
Share
Copy Link
DeepSeek unveiled DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal AI model with image understanding capabilities that approaches Anthropic's Opus 4.8 performance. The model wins on three benchmarks including Agents' Last Exam and ZeroBench while maintaining 99% lower operational costs, intensifying competition in the global AI landscape.
DeepSeek has released DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal AI model that extends its V4-Flash text model with image understanding capabilities
1
. The Hangzhou-based company positions this release as approaching the performance of Anthropic Opus 4.8, one of the most advanced models currently available3
. This launch marks a strategic move by Chinese AI labs to compete in visual AI performance while maintaining significant cost advantages over American competitors.The model preserves the original text functionalities of DeepSeek-V4-Flash, including reasoning and general knowledge capabilities, while adding multimodal agent benchmarks support
1
. DeepSeek-V4-Flash-Vision-Exp supports JPEG, PNG, GIF and WebP formats, allowing developers to upload up to 600 images within a single request3
. This capability positions the model as particularly valuable for document and screenshot-based agentic system applications.DeepSeek conducted eleven benchmark comparisons with Anthropic Opus 4.8, demonstrating competitive visual AI performance across multiple tests
3
. On Apex Bench, the DeepSeek-V4-Flash-Vision-Exp model scored 36.5 at Pass@1, while Opus 4.8 achieved 39.41
2
. The model outperformed Anthropic's offering on Agents' Last Exam with a score of 27.3 compared to Opus 4.8's 25.71
.
Source: Digit
On ZeroBench, DeepSeek's multimodal AI model achieved 35.0, surpassing Opus 4.8's score of 34.0
1
. The model also demonstrated benchmark dominance on the Deep Software Engineering test, though detailed scores varied across different evaluations2
. However, the model lost to Opus 4.8 in eight of the eleven comparisons, with some differences reaching double digits on the hardest tests3
.DeepSeek evaluated its models using its internal Harness Minimal Mode, meaning the performance figures have not been independently verified
1
. The company also released version 0.1.1 of its Harness framework with built-in support for the new model3
.The operational cost of DeepSeek-V4-Flash-Vision-Exp is approximately 99% lower than Anthropic Opus 4.8, representing the most significant advantage of this release
3
. This cost efficiency makes the model particularly attractive to developers seeking affordable alternatives with competitive image understanding capabilities. The pricing structure bills images by token count, capped at 384 tokens per image1
.DeepSeek introduced a Files API that enables developers to upload images once and reference them multiple times using a file ID without incurring additional costs
1
3
. Integration with OpenRouter enhances accessibility for developers and researchers, reflecting DeepSeek's strategy to prioritize both performance and usability2
.Related Stories
This launch occurs amid heightened competition in the global AI landscape from domestic rivals like Alibaba and Moonshot AI, as well as U.S. firms such as Anthropic and OpenAI
1
. Chinese AI labs are leveraging extensive resources and technical expertise to narrow the gap with established players, driving rapid advancements in multimodal tasks and realistic simulations2
.
Source: Geeky Gadgets
The emergence of Aux Alpha, another high-performance multimodal AI model with a 1-million-token context window, has sparked additional interest within the AI community
2
. Speculation links Aux Alpha to GLM 5.3 or Xiaomi's rumored Mimo 3, with leaks from Chinese forums suggesting involvement of Chinese AI labs. The internal rivalry among Chinese AI labs is fueling innovation and challenging the dominance of global leaders.Developers can now access DeepSeek-V4-Flash-Vision-Exp through the DeepSeek API platform
3
. The model's future depends on developer adoption and real-world testing to determine whether its performance translates into reliable operations for purpose-built agents. Watch for independent benchmark verification and broader deployment across multimodal agent benchmarks to assess the model's true competitive position against Anthropic Opus 4.8, which remains fully supported until May 20273
.Summarized by
Navi
[2]
1
Science and Research

2
Technology

3
Policy and Regulation
