Ant Group's Ling-2.6-Flash Was the Mystery Model Topping OpenRouter Charts
On April 16, 2026, a model called “Elephant Alpha” appeared on OpenRouter with no attribution and quickly climbed to the top of the trending leaderboard, hitting 100 billion daily token calls within a week. Initial coverage — including here — attributed it to OpenRouter. That attribution was wrong.
Ant Group officially confirmed on April 22 that Elephant Alpha is Ling-2.6-Flash, a production large language model from the research arm of China’s largest fintech company.
The Model
Ling-2.6-Flash is a sparse Mixture-of-Experts model with 104 billion total parameters and 7.4 billion active per forward pass. Ant Group designed it specifically for AI agent applications, claiming state-of-the-art performance for its size class on BFCL-V4, TAU2-bench, SWE-bench Verified, Claw-Eval, and PinchBench.
The standout number from Artificial Analysis: Ling-2.6-Flash scored Intelligence Index 26 while generating only 15 million output tokens to complete the full eval suite. Nemotron-3-Super needed 110 million tokens to reach a comparable score. Ant Group claims this translates to an 86% reduction in inference cost for equivalent intelligence, which is the actual story here: the model is not trying to top raw capability rankings, it is optimized to minimize token spend while staying competitive.
Inference speeds hit 340 tokens per second on a 4-card H20 setup, with Prefill throughput 2.2x that of Nemotron-3-Super.
Pricing
- Input: $0.10 per million tokens
- Output: $0.30 per million tokens
- One-week free trial via OpenRouter and Alipay Tbox
A commercial deployment called LingDT is available through Ant Digital Technologies for enterprise and SME use.
Why the Anonymous Launch
Ant Group did not explain why the model was soft-launched under a codename. The playbook is not unusual for Chinese labs — test real-world adoption and load characteristics before making a formal claim. What is unusual is how fast it moved: reaching the trending top before most observers had time to speculate about the source.
The actual use case Ant Group is pitching is agent pipelines where token cost compounds at scale. At $0.30/M output, Ling-2.6-Flash is 25x cheaper than Claude Opus 4.6 on output and roughly 10x cheaper than GPT-5.4 at standard pricing. For high-volume agentic workloads where the model does not need frontier reasoning but does need to execute thousands of tool calls per day, the cost delta is material.
What Changed
The earlier Stack Futures item attributed Elephant Alpha to OpenRouter directly. That was incorrect. OpenRouter was the distribution platform, not the lab. Ant Group built the model, tested it anonymously, and revealed the name once usage data validated the deployment.
Ling-2.6-Flash is now listed on Artificial Analysis under its official name, where it sits in the mid-tier of the intelligence rankings but near the top of the efficiency tier — which appears to be exactly where Ant Group intended it to land.