StepFun's Step 5 Preview: 600B MoE Enters the Agentic Race, Leads Chinese Peers on Most Benchmarks
StepFun (阶跃星辰) has released Step 5 Preview, its new flagship model built for agentic and professional knowledge work. The architecture is a sparse Mixture-of-Experts at 600B total parameters with 27B active per token, a 1M-token context window, and vision input. The lab positions it as a new Pareto point between intelligence and cost, though the numbers tell a more specific story.
Where It Lands
On the benchmarks StepFun chose to publish, Step 5 Preview occupies a consistent second tier behind GPT-6 Astra and Claude Opus 5, while beating Kimi K3 and GLM-5.3 across most tasks.
DeepSWE v1.1 (software engineering):
- Step 5 Preview (High): 67.7
- Kimi K3 (Max): 67.5
- GLM-5.3 (Max): 66.9
- Claude Opus 5 (Max): 74.0
- GPT-6 Astra (Max): 74.1
Step 5 edges Kimi K3 by 0.2 points and GLM-5.3 by 0.8, but sits 6.4 points behind GPT-6 Astra and 6.3 points behind Claude Opus 5.
ProgramBench (competitive programming):
- Step 5 Preview (High): 80.5
- Kimi K3 (Max): 77.8
- GLM-5.3 (Max): 72.0
- Claude Opus 5 (Max): 82.3
- GPT-6 Astra (Max): 85.4
Step 5 leads Kimi K3 by 2.7 points and GLM-5.3 by 8.5, but remains 4.9 points behind GPT-6 Astra and 1.8 points behind Opus 5.
Terminal-Bench v4 (end-to-end agent tasks):
- GLM-5.3 (Max): 41.9
- Step 5 Preview (High): 33.3
- Kimi K3 (Max): 12.6
- GPT-6 Astra (Max): 57.9
- Claude Opus 5 (Max): 52.3
Terminal-Bench v4 is the clearest gap: Step 5 at 33.3 is 24.6 points behind GPT-6 Astra and 19.0 points behind Opus 5. Kimi K3 at 12.6 is substantially weaker on this benchmark.
Agents’ Last Exam (ALE-CLI) (frontier agentic capability):
- Step 5 Preview (High): 29.5
- Claude Opus 5 (Max): 28.6
- GLM-5.3 (Max): 28.6
- Kimi K3 (Max): 27.6
- GPT-6 Astra (Max): 33.3
ALE-CLI is the benchmark where Step 5 comes closest to the frontier, beating Opus 5 by 0.9 points. GPT-6 Astra still leads at 33.3.
GDPval-AA v2 (aggregate capability index):
- GLM-5.3 (Max): 1634
- Step 5 Preview (High): 1571
- Kimi K3 (Max): 1548
What It Means
Step 5 Preview is the strongest model StepFun has shipped. On DeepSWE, ProgramBench, and ALE-CLI it consistently outperforms Kimi K3 and GLM-5.3, establishing it as a credible option in the Chinese-domestic agentic tier.
The “Pareto frontier” framing in the launch is marketing shorthand. The actual frontier on Terminal-Bench v4 sits at 57.9 (GPT-6 Astra) and 52.3 (Opus 5). Step 5 Preview at 33.3 is not on that boundary — it’s competitive at the sub-frontier tier where the cost-to-performance tradeoff does shift outward compared with smaller Chinese models.
StepFun has published pricing: $1.00 per 1M input tokens ($0.05 cached) and $2.70 per 1M output tokens. The model is available for API access through the StepFun platform.