StepAudio 3 Realtime
Chinese and English. Preview ids will be retired for a paid GA version when the trial ends. Predecessor stepaudio-2.5-realtime (2026-05-26; persona role-play, project page https://stepaudiollm.github.io/step-audio-2.5-realtime/ with self-reported 86.36 general dialogue / 79.80 spoken QA / 82.18 paralinguistics, claimed to beat GPT-Realtime-1.5 on StepFun's evals). Technical report arXiv 2609.14005 (56.0% task success on tau-Voice).
- Input
- audio, text
- Output
- audio, text
- License
- proprietary
- Pricing
- input: $0 · output: $0 (free during limited-time preview (successor stepaudio-2.5-realtime: $1.50 in / $0.30 cached / $10.00 out per 1M tokens)) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| StepFun API (Realtime WebSocket) | stepaudio-3-realtime-preview | — | docs |
| StepFun API (Chat Completions) | stepaudio-3-chat-preview | — | docs |
Notable capabilities (2)
- Think-while-speaking full duplex: Runs private chain-of-thought in parallel with spoken output; distinguishes real interruptions from backchannels; asynchronous tool execution (web search, knowledge retrieval). source
- #1 on Artificial Analysis conversational dynamics: 98.9 on Artificial Analysis Full-Duplex Bench (Conversational Dynamics) and 99.7% Speech Reasoning at launch, ahead of Qwen Audio 3.0 Realtime Plus and GPT-Live-1 per StepFun. source
StepFun's full-duplex voice agent model; part of the five-model StepAudio 3 family (Realtime, ASR Max, TTS, Gen, Music).
Sources: https://platform.stepfun.ai/docs/en/guides/models/audio · https://platform.stepfun.ai/docs/en/pricing/details · https://arxiv.org/abs/2609.14005
Timeline entry
- StepFun releases StepAudio 3 family; its Realtime model tops Artificial Analysis full-duplex rankings ★★★
Chinese lab StepFun launched StepAudio 3, five audio models (Realtime, ASR Max, TTS, Gen, Music). StepAudio 3 Realtime, a "think-while-speaking" full-duplex voice model, ranked #1 on Artificial Analysis for Conversational Dynamics (98.9%) and Speech Reasoning (99.7%), and StepAudio 3 ASR ranked #1…
Other StepFun models
StepAudio 3 ASR Max / StepAudio 3 TTS · StepFun Step-Audio-EditX · Step 5 Preview