GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

StepFun Ships Five-Model StepAudio 3 Suite, Takes First on Conversational Dynamics and ASR

StepFun launched all five models in its StepAudio 3 lineup on September 15, the same day — an unusual simultaneous release covering every layer of an audio AI stack: real-time speech, speech recognition, text-to-speech, audio generation, and music creation. Third-party benchmarks from Artificial Analysis placed two of them at the top of their respective leaderboards immediately on release.

Five Models, One Day

The StepAudio 3 lineup:

  • StepAudio 3 Realtime — live speech-to-speech for agent and conversational use cases
  • StepAudio 3 ASR — speech recognition with non-streaming and streaming modes
  • StepAudio 3 TTS — text-to-speech synthesis
  • StepAudio 3 Gen — general audio generation
  • StepAudio 3 Music — music generation

All five went live on StepFun’s open platform simultaneously.

Benchmark Results

Third-party evaluator Artificial Analysis ran the models against its speech leaderboards:

ModelBenchmarkScoreRank
StepAudio 3 RealtimeConversational Dynamics98.9% composite#1
StepAudio 3 RealtimeSpeech Reasoning accuracy99.7%#1
StepAudio 3 ASRWER (non-streaming)1.7%#1 (tied)

The Conversational Dynamics leaderboard measures the full real-time speech-to-speech loop, including turn detection, latency, and response coherence. A 98.9% composite on that benchmark puts Realtime clear of the prior leader.

The 1.7% word error rate in non-streaming ASR is at the low end of what major commercial services report. StepAudio 3 ASR Max is a one-shot submission model; StepFun’s separate StepAudio 2.5 ASR Stream handles real-time audio input.

Context

StepFun has positioned itself in the efficiency tier of the AI market: strong on throughput and cost, competitive on capability. The StepAudio 3 results suggest that positioning has extended into audio, where latency and real-time coherence matter as much as raw accuracy.

The audio market is also a differentiated target. Voice AI is increasingly embedded in agent frameworks, customer service, and consumer hardware. Conversational Dynamics measures the full interaction loop — not just recognition accuracy — which is the dimension that determines user experience in deployed products.

StepAudio 3 is available on StepFun’s platform. Pricing published at launch: TTS at $0.36 per 10,000 characters (2.5 RMB), ASR Max at $0.24 per hour (2.8 RMB). Realtime, Gen, and Music launched under limited free-access preview pricing.