GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

StepFun's Step 5 Preview: 600B MoE Enters the Agentic Race, Leads Chinese Peers on Most Benchmarks

StepFun (阶跃星辰) has released Step 5 Preview, its new flagship model built for agentic and professional knowledge work. The architecture is a sparse Mixture-of-Experts at 600B total parameters with 27B active per token, a 1M-token context window, and vision input. The lab positions it as a new Pareto point between intelligence and cost, though the numbers tell a more specific story.

Where It Lands

On the benchmarks StepFun chose to publish, Step 5 Preview occupies a consistent second tier behind GPT-6 Astra and Claude Opus 5, while beating Kimi K3 and GLM-5.3 across most tasks.

DeepSWE v1.1 (software engineering):

  • Step 5 Preview (High): 67.7
  • Kimi K3 (Max): 67.5
  • GLM-5.3 (Max): 66.9
  • Claude Opus 5 (Max): 74.0
  • GPT-6 Astra (Max): 74.1

Step 5 edges Kimi K3 by 0.2 points and GLM-5.3 by 0.8, but sits 6.4 points behind GPT-6 Astra and 6.3 points behind Claude Opus 5.

ProgramBench (competitive programming):

  • Step 5 Preview (High): 80.5
  • Kimi K3 (Max): 77.8
  • GLM-5.3 (Max): 72.0
  • Claude Opus 5 (Max): 82.3
  • GPT-6 Astra (Max): 85.4

Step 5 leads Kimi K3 by 2.7 points and GLM-5.3 by 8.5, but remains 4.9 points behind GPT-6 Astra and 1.8 points behind Opus 5.

Terminal-Bench v4 (end-to-end agent tasks):

  • GLM-5.3 (Max): 41.9
  • Step 5 Preview (High): 33.3
  • Kimi K3 (Max): 12.6
  • GPT-6 Astra (Max): 57.9
  • Claude Opus 5 (Max): 52.3

Terminal-Bench v4 is the clearest gap: Step 5 at 33.3 is 24.6 points behind GPT-6 Astra and 19.0 points behind Opus 5. Kimi K3 at 12.6 is substantially weaker on this benchmark.

Agents’ Last Exam (ALE-CLI) (frontier agentic capability):

  • Step 5 Preview (High): 29.5
  • Claude Opus 5 (Max): 28.6
  • GLM-5.3 (Max): 28.6
  • Kimi K3 (Max): 27.6
  • GPT-6 Astra (Max): 33.3

ALE-CLI is the benchmark where Step 5 comes closest to the frontier, beating Opus 5 by 0.9 points. GPT-6 Astra still leads at 33.3.

GDPval-AA v2 (aggregate capability index):

  • GLM-5.3 (Max): 1634
  • Step 5 Preview (High): 1571
  • Kimi K3 (Max): 1548

What It Means

Step 5 Preview is the strongest model StepFun has shipped. On DeepSWE, ProgramBench, and ALE-CLI it consistently outperforms Kimi K3 and GLM-5.3, establishing it as a credible option in the Chinese-domestic agentic tier.

The “Pareto frontier” framing in the launch is marketing shorthand. The actual frontier on Terminal-Bench v4 sits at 57.9 (GPT-6 Astra) and 52.3 (Opus 5). Step 5 Preview at 33.3 is not on that boundary — it’s competitive at the sub-frontier tier where the cost-to-performance tradeoff does shift outward compared with smaller Chinese models.

StepFun has published pricing: $1.00 per 1M input tokens ($0.05 cached) and $2.70 per 1M output tokens. The model is available for API access through the StepFun platform.