Tencent's Hy3 Is Beating Claude by 50% on OpenRouter — At $0.066/M, Nobody Can Explain It
Something unusual is happening on OpenRouter’s model usage rankings. Tencent’s Hy3 preview — a 295 billion parameter MoE model with 21 billion active params, released April 22 and open-sourced the next day — is processing 3.47 trillion tokens per week. That puts it ahead of Claude Opus 4.7 by more than 50% in aggregate token volume. It also leads coding and tool-calling usage categories.
The model’s Artificial Analysis Intelligence Index score is 41.9, placing it in the 77th percentile — well below frontier models like Opus 4.7 or GPT-5.5. SWE-bench Verified: 74.4%. tau2-bench Telecom: 92.7%. Competent, not exceptional.
Pricing Doesn’t Fully Explain It Either
Hy3 preview is priced at $0.066 per million input tokens, $0.26 output. At 79.7% cache hit rate via its sole provider, SiliconFlow (Singapore), the effective input price drops to $0.036 per million. That is cheap — but not the cheapest.
DeepSeek V4 Flash, served directly through DeepSeek’s own API on OpenRouter, has a stated price of $0.10/M input — but a 2% cache read cost. With DeepSeek’s aggressive KV cache architecture baked into V4, the effective input price is $0.018 per million tokens. That is half Hy3’s effective rate.
Yet Hy3 is beating V4 Flash in total volume too.
What the Usage Pattern Reveals
The model launched free on OpenRouter around May 8, then moved to paid. Usage did not drop when pricing hit — a signal that whoever is consuming Hy3 tokens found the model worth paying for. Top five public apps account for less than 1% of Hy3’s total traffic. The usage is not explained by a single viral developer tool.
The input-to-output token ratio across all OpenRouter API calls is now running at 98:2 — 98% input tokens, 2% output. That ratio is the fingerprint of agentic workflows: long contexts reprocessed turn-by-turn, with agents re-ingesting prior output at each step. Someone is running Hy3 in a high-volume, long-context agentic loop. The best hypothesis is a single large platform not listed among public apps, using Hy3 as its data-processing backbone.
One complication: Hy3’s license is restrictive in ways that could limit commercial redistribution. The model is routed exclusively through SiliconFlow, which itself had minimal OpenRouter traffic until Hy3 arrived.
The Broader Signal
The usage data surfaces something more structurally important than Hy3 itself. LLM stated prices are becoming misleading in a direction that favours buyers. Cache hit rates of 44% to 80%+ mean actual per-token costs are a fraction of published rates — especially for providers that deeply integrate their own models’ caching architectures. DeepSeek’s 2% cache read cost (compared to the typical 10%) is not accidental; it is a structural advantage built into V4’s KV cache design.
For enterprise buyers evaluating AI spend, the relevant benchmark is not the headline price per million tokens. It is the effective price after cache, multiplied by the input:output ratio for a given workflow. At 98% input-heavy agentic tasks, cache arithmetic dominates everything else.
Hy3 may be the beneficiary of one large customer running on SiliconFlow who found it good enough, cheap enough, and geopolitically easier than using DeepSeek’s own infrastructure directly. Or it is something else entirely. OpenRouter’s data does not go deep enough to say which.
Key Numbers
- Weekly tokens: 3.47T (Hy3 preview), vs ~2.3T (Claude Opus 4.7, estimated rank)
- Stated input price: $0.066/M
- Effective input price after cache: $0.036/M (79.7% hit rate, SiliconFlow)
- DeepSeek V4 Flash effective via DeepSeek: $0.018/M
- Intelligence Index (AA): 41.9 — 77th percentile
- SWE-bench Verified: 74.4%
- tau2-bench Telecom: 92.7%
- Context: 262K tokens
- Top-5 app share of total traffic: less than 1%
- API input:output ratio across all OpenRouter calls: 98:2