StepFun Ships Five-Model StepAudio 3 Suite, Takes First on Conversational Dynamics and ASR
StepFun launched all five models in its StepAudio 3 lineup on September 15, the same day — an unusual simultaneous release covering every layer of an audio AI stack: real-time speech, speech recognition, text-to-speech, audio generation, and music creation. Third-party benchmarks from Artificial Analysis placed two of them at the top of their respective leaderboards immediately on release.
Five Models, One Day
The StepAudio 3 lineup:
- StepAudio 3 Realtime — live speech-to-speech for agent and conversational use cases
- StepAudio 3 ASR — speech recognition with non-streaming and streaming modes
- StepAudio 3 TTS — text-to-speech synthesis
- StepAudio 3 Gen — general audio generation
- StepAudio 3 Music — music generation
All five went live on StepFun’s open platform simultaneously.
Benchmark Results
Third-party evaluator Artificial Analysis ran the models against its speech leaderboards:
| Model | Benchmark | Score | Rank |
|---|---|---|---|
| StepAudio 3 Realtime | Conversational Dynamics | 98.9% composite | #1 |
| StepAudio 3 Realtime | Speech Reasoning accuracy | 99.7% | #1 |
| StepAudio 3 ASR | WER (non-streaming) | 1.7% | #1 (tied) |
The Conversational Dynamics leaderboard measures the full real-time speech-to-speech loop, including turn detection, latency, and response coherence. A 98.9% composite on that benchmark puts Realtime clear of the prior leader.
The 1.7% word error rate in non-streaming ASR is at the low end of what major commercial services report. StepAudio 3 ASR Max is a one-shot submission model; StepFun’s separate StepAudio 2.5 ASR Stream handles real-time audio input.
Context
StepFun has positioned itself in the efficiency tier of the AI market: strong on throughput and cost, competitive on capability. The StepAudio 3 results suggest that positioning has extended into audio, where latency and real-time coherence matter as much as raw accuracy.
The audio market is also a differentiated target. Voice AI is increasingly embedded in agent frameworks, customer service, and consumer hardware. Conversational Dynamics measures the full interaction loop — not just recognition accuracy — which is the dimension that determines user experience in deployed products.
StepAudio 3 is available on StepFun’s platform. Pricing published at launch: TTS at $0.36 per 10,000 characters (2.5 RMB), ASR Max at $0.24 per hour (2.8 RMB). Realtime, Gen, and Music launched under limited free-access preview pricing.