5 Sources
[1]
New Deepseek model drastically reduces resource usage by converting text and documents into images -- 'vision-text compression' uses up to 20 times fewer tokens
Could help cut costs and improve the efficiency of the latest AI models. Chinese developers of Deepseek AI have released a new model that leverages its multi-modal capabilities to improve the efficiency of its handling of complex documents and large blocks of text, by converting them into images
[2]
DeepSeek drops open-source model that compresses text 10x through images, defying conventions
DeepSeek, the Chinese artificial intelligence research company that has repeatedly challenged assumptions about AI development costs, has released a new model that fundamentally reimagines how large language models process information -- and the implications extend far beyond its modest branding as
[3]
DeepSeek's New OCR Model Can Process Over 2 Lakh Pages Daily on a Single GPU | AIM
The technology introduces a vision-based approach to context compression, converting text into compact visual tokens. DeepSeek AI has announced DeepSeek-OCR, a new optical character recognition (OCR) system designed to improve how large language models handle long text contexts through optical 2D
[4]
DeepSeek-OCR: New open-source AI model goes viral on GitHub
DeepSeek-OCR's power lies in its ability to compress information. According to its creators, the model can take a 1,000-word article and compress it into just 100 visual tokens. A new open-source model named DeepSeek-OCR has been released, disrupting the traditional paradigm of large models. The
[5]
DeepSeek-OCR Could Change How AI Reads Text From Images
The model turns text into pixels to improve its context memory DeepSeek, on Monday, released a new open-source artificial intelligence (AI) model that changes how these machines analyse and process plain text. Dubbed DeepSeek-OCR, it uses 2D mapping to convert text into pixels to compress long
Share
Copy Link
DeepSeek's new open-source AI model, DeepSeek-OCR, introduces a groundbreaking approach to text processing by converting text into images. This method achieves up to 20x compression while maintaining high accuracy, potentially revolutionizing AI language models' efficiency and capabilities.
Chinese AI company DeepSeek has unveiled a groundbreaking open-source model called DeepSeek-OCR, which challenges conventional approaches to text processing in large language models (LLMs). The model's innovative technique converts text into images, achieving significant compression while maintaining high accuracy
1
2
.
Source: Gadgets 360
DeepSeek-OCR's core innovation lies in its ability to compress textual information through visual representation. The model can achieve a compression ratio of up to 20 times, with a 97% accuracy rate at 10x compression
1
. This approach inverts the traditional hierarchy where text tokens were considered more efficient than vision tokens2
.DeepSeek-OCR comprises two main components:
The model outperforms existing OCR systems on benchmarks like OmniDocBench while using fewer vision tokens
2
3
.DeepSeek-OCR's efficiency translates to impressive real-world performance. A single Nvidia A100-40G GPU can process more than 200,000 pages per day, scaling up to 33 million pages daily with a cluster of 20 servers
2
3
. This efficiency makes it suitable for large-scale document digitization and AI training data generation3
.The compression breakthrough could potentially unlock 10 million token context windows for language models, a significant leap from current state-of-the-art models that typically handle context windows measured in hundreds of thousands of tokens
2
.Related Stories
The AI community has responded enthusiastically to DeepSeek-OCR. Andrej Karpathy, co-founder of OpenAI and former director of AI at Tesla, suggested that this approach could fundamentally change how AI systems process information
4
5
.DeepSeek has made both the code and model weights for DeepSeek-OCR available as an open-source project on GitHub and Hugging Face
3
5
. This release aims to support broader research into combining vision and language for more efficient AI systems, potentially leading to a paradigm shift in how language models process and understand information3
4
.Summarized by
Navi
[1]
[2]
[5]
10 Sept 2026•Technology

29 Sept 2025•Technology

25 Mar 2025•Technology

1
Technology

2
Policy and Regulation

3
Technology
