StepFun Step-Audio-EditX
Official changelog lists a new model release on 2026-01-29 (overall ~4% improvement; new paralinguistic tags such as exhale, inhale, chuckle, clears throat, giggle; SFT/DPO/GRPO training code released); HF weights updated 2026-01-23/24, README edits to 2026-02-14. No March 2026 release appears in the official GitHub/HF changelog, so Artificial Analysis's 'Step Audio EditX (Mar 2026)' label (#3 open weights, ~1095 Elo, Sept 2026) probably refers to the Jan 2026 weights or a hosted snapshot (unverified).
- Input
- text, audio
- Output
- audio
- License
- apache-2.0 (code; check model card for weights)
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | stepfun-ai/Step-Audio-EditX | huggingface.co/stepfun-ai/Step-Audio-EditX | — |
| Hugging Face (4-bit) | stepfun-ai/Step-Audio-EditX-AWQ-4bit | huggingface.co/stepfun-ai/Step-Audio-EditX-AWQ-4bit | — |
| GitHub | — | github.com/stepfun-ai/Step-Audio-EditX | — |
Notable capabilities (1)
- Iterative LLM-based audio editing: 3B RL-trained audio LLM that edits emotion, speaking style and paralinguistics of existing speech step by step, plus zero-shot TTS cloning (Mandarin, English, Sichuanese, Cantonese; Japanese/Korean added 2025-11-28). source
Sources: https://github.com/stepfun-ai/Step-Audio-EditX , https://huggingface.co/stepfun-ai/Step-Audio-EditX
Other StepFun models
StepAudio 3 ASR Max / StepAudio 3 TTS · Step 5 Preview · StepAudio 3 Realtime