DeepSeek Ships Image Recognition to All Users at 10x Fewer Tokens Than Rivals, V4.1 Due in June
DeepSeek has made its image recognition mode generally available to all users as of May 9, completing a beta period that began last month. The feature marks DeepSeek’s formal entry into multimodal AI — a capability gap the lab has lagged behind Gemini, GPT-5.5, and Claude Opus 4.7 on since V4’s text-only launch.
The Token Efficiency Claim
The headline number is the compute cost. DeepSeek’s visual architecture processes an 800x800 resolution image in approximately 90 tokens. Mainstream frontier models — GPT-4o, Gemini 3 Flash, Claude Sonnet — consume 870 to 1,100 tokens for the same image size. At 9x to 12x lower token cost per visual input, DeepSeek’s approach has structural implications for any high-volume image-processing deployment: document analysis, product catalog search, satellite imagery, medical imaging pipelines.
The efficiency comes from a lightweight visual tokenizer that compresses spatial information aggressively before handing off to the language model. At API pricing of $0.14/M for DeepSeek V4 Flash, the cost delta per image versus a comparable GPT-5.5 call is significant enough to matter in production workloads.
V4.1 in June
DeepSeek has told investors to expect a V4.1 model update in June. No benchmark numbers have been released for V4.1, but the briefing is consistent with DeepSeek’s pattern of releasing model updates on roughly 6-to-8 week cycles since V3 launched in late 2025. The V4.1 designation implies an incremental improvement over V4 Pro and Flash rather than an architectural departure.
Fundraise Context
The capability launch coincides with DeepSeek’s reported fundraising. A planned round of up to 50 billion yuan ($7.35 billion) has been in progress, targeting a valuation of approximately $45 billion. The vision rollout gives the company a capability narrative that extends beyond code and math benchmarks — the areas where V4 established its reputation.
V4 Pro’s 80.6% SWE-bench Verified score and MIT-licensed release last month put it at the frontier of open-weight coding performance. V4.1 now has to extend that into multimodal territory while defending V4 Pro’s coding position against DeepSeek R2’s reasoning competition and the accelerating pace of Qwen3.6 updates from Alibaba.
Numbers
- Image token cost: ~90 tokens per 800x800 (vs 870–1,100 for mainstream models)
- Efficiency ratio: 9x–12x fewer tokens vs frontier competitors
- V4.1 timeline: June 2026 (per investor briefings)
- V4 Flash pricing: $0.14/M input (reference rate)
- Fundraise: Up to $7.35B at ~$45B valuation (in progress)
- V4 Pro SWE-bench: 80.6% Verified (current open-weight frontier)