Anthropic's Best Model Is Struggling to Win Users as Cheaper Rivals Close the Gap
Anthropic’s Claude Opus 5 sits at the top of SWE-bench Verified at 97.0%, leads multiple agentic coding benchmarks, and ranks first on Artificial Analysis’s intelligence index. It is also struggling to attract users.
The Financial Times reported Monday that Anthropic’s best AI model is seeing weak uptake while cheaper tools are winning the market. Analyst Gary Marcus, who cited the story, called it “more bad news for frontier AI companies” and flagged direct implications for Anthropic’s upcoming IPO.
The Competitive Squeeze
The adoption pattern matches what leaderboard cost data shows. Claude Fable 5 Max Effort — Anthropic’s primary developer-facing frontier model — costs $1.439 per successful LiveBench task. Gemini 3.7 Flash High delivers a comparable overall LiveBench score at $0.157 per task. Kimi K3, an open-weight model from Moonshot AI, matches Fable 5’s agentic coding score of 62.2 while costing roughly one-quarter as much per task.
Production AI buyers are doing this math. A model at the top of every benchmark chart is only a compelling purchase when the performance premium justifies the price multiple. For most workloads, it is not.
| Model | LiveBench Overall | Agentic Coding | Cost/Task |
|---|---|---|---|
| Claude Fable 5 Max Effort | 83.0 | 62.2 | $1.439 |
| Claude Opus 5 Thinking Max | 80.1 | 65.2 | $0.699 |
| Kimi K3 open | 79.2 | 62.2 | $0.348 |
| Gemini 3.7 Flash High | 78.8 | 58.3 | $0.157 |
Gemini 3.7 Flash High delivers agentic coding performance within four points of Claude Fable 5 at roughly 11 cents on the dollar. Kimi K3 matches Fable 5 outright in that category at less than a quarter the price. The case for paying frontier-tier prices is harder to make than it was a year ago.
The IPO Question
Anthropic’s IPO, targeting a reported $2 trillion valuation in October, was premised on market leadership at the frontier, not just capability leadership. Those are different things. If enterprise and developer buyers are routing to Gemini Flash and Kimi K3 for the majority of workloads, the revenue base supporting that valuation is under pressure.
Marcus specifically linked the FT story to the IPO question, noting that “cheaper tools are thriving” at precisely the moment Anthropic needs to show a growing premium user base. The IPO narrative assumes frontier buyers pay for the top tier at scale. The market is showing they often will not.
What Anthropic Can Do
The problem is not capability. Claude Opus 5 at 97.0% on SWE-bench Verified and Claude Fable 5’s coding performance are genuine technical achievements. The issue is price elasticity and the speed at which open and mid-tier models are compressing the effective performance gap.
Anthropic has two obvious responses. One is pricing — OpenAI’s 20% cut to GPT-5.6 Sol last week shows the playbook, though Sol’s cut came with a three-month expiry clause that introduces its own uncertainty. The other is repositioning the value proposition around safety, reliability, and enterprise compliance, areas where Anthropic has built more credible infrastructure than any competitor.
Neither solves the core problem: the benchmark gap that justified premium pricing two years ago has largely closed at the application layer for the workloads most buyers actually run. Kimi K3 tying Fable 5 in agentic coding is the kind of result that changes procurement conversations, regardless of what the overall LiveBench score says.
The FT story names the trend. The leaderboard data explains why.