GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%
← Back to feed

DeepSeek Ships Image Recognition to All Users at 10x Fewer Tokens Than Rivals, V4.1 Due in June

DeepSeek has made its image recognition mode generally available to all users as of May 9, completing a beta period that began last month. The feature marks DeepSeek’s formal entry into multimodal AI — a capability gap the lab has lagged behind Gemini, GPT-5.5, and Claude Opus 4.7 on since V4’s text-only launch.

The Token Efficiency Claim

The headline number is the compute cost. DeepSeek’s visual architecture processes an 800x800 resolution image in approximately 90 tokens. Mainstream frontier models — GPT-4o, Gemini 3 Flash, Claude Sonnet — consume 870 to 1,100 tokens for the same image size. At 9x to 12x lower token cost per visual input, DeepSeek’s approach has structural implications for any high-volume image-processing deployment: document analysis, product catalog search, satellite imagery, medical imaging pipelines.

The efficiency comes from a lightweight visual tokenizer that compresses spatial information aggressively before handing off to the language model. At API pricing of $0.14/M for DeepSeek V4 Flash, the cost delta per image versus a comparable GPT-5.5 call is significant enough to matter in production workloads.

V4.1 in June

DeepSeek has told investors to expect a V4.1 model update in June. No benchmark numbers have been released for V4.1, but the briefing is consistent with DeepSeek’s pattern of releasing model updates on roughly 6-to-8 week cycles since V3 launched in late 2025. The V4.1 designation implies an incremental improvement over V4 Pro and Flash rather than an architectural departure.

Fundraise Context

The capability launch coincides with DeepSeek’s reported fundraising. A planned round of up to 50 billion yuan ($7.35 billion) has been in progress, targeting a valuation of approximately $45 billion. The vision rollout gives the company a capability narrative that extends beyond code and math benchmarks — the areas where V4 established its reputation.

V4 Pro’s 80.6% SWE-bench Verified score and MIT-licensed release last month put it at the frontier of open-weight coding performance. V4.1 now has to extend that into multimodal territory while defending V4 Pro’s coding position against DeepSeek R2’s reasoning competition and the accelerating pace of Qwen3.6 updates from Alibaba.

Numbers

  • Image token cost: ~90 tokens per 800x800 (vs 870–1,100 for mainstream models)
  • Efficiency ratio: 9x–12x fewer tokens vs frontier competitors
  • V4.1 timeline: June 2026 (per investor briefings)
  • V4 Flash pricing: $0.14/M input (reference rate)
  • Fundraise: Up to $7.35B at ~$45B valuation (in progress)
  • V4 Pro SWE-bench: 80.6% Verified (current open-weight frontier)