Step Audio EditX
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech
Install / Use
npx skills add stepfun-ai/Step-Audio-EditXInstalls into whichever agent you are using.
README
Step-Audio-EditX
<p align="center"> <img src="assets/logo.png" height=100> </p> <div align="center"> <a href="https://stepaudiollm.github.io/step-audio-editx/"><img src="https://img.shields.io/static/v1?label=Demo%20Page&message=Web&color=green"></a>   <a href="https://arxiv.org/abs/2511.03601"><img src="https://img.shields.io/static/v1?label=Tech%20Report&message=Arxiv&color=red"></a>   <a href="https://huggingface.co/stepfun-ai/Step-Audio-EditX"><img src="https://img.shields.io/static/v1?label=Step-Audio-EditX&message=HuggingFace&color=yellow"></a>  <a href="https://modelscope.cn/models/stepfun-ai/Step-Audio-EditX"><img src="https://img.shields.io/static/v1?label=Step-Audio-EditX&message=ModelScope&color=blue"></a> <a href="https://huggingface.co/spaces/stepfun-ai/Step-Audio-EditX"><img src="https://img.shields.io/static/v1?label=Space%20Playground&message=HuggingFace&color=yellow"></a> <a href="https://www.stepfun.com/studio/audio?tab=edit"><img src="https://img.shields.io/static/v1?label=Audio%20Studio&message=StepFun&color=blue"></a>
</div>🔥🔥🔥 News!!!
- Jan 29, 2026:
- 🧩 New Model Release:
- Better performance, with an overall improvement of over 4%.
- More paralinguistic tags have been added, including
exhale,snort,inhale,chuckle,clears throat,giggle. - Welcome to try out at StepFun Audio Studio
- 💻 We release the SFT, DPO and GRPO training code.
- 🌟 Training and inference for vLLM are now supported. Thanks to the vLLM team!
- 🧩 New Model Release:
- Nov 28, 2025: 🚀 New Model Release: Now supporting
JapaneseandKoreanlanguages. - Nov 23, 2025: 📊 Step-Audio-Edit-Benchmark Released!
- Nov 19, 2025: ⚙️ We release a new version of our model, which supports polyphonic pronunciation control and improves the performance of emotion, speaking style, and paralinguistic editing.
- Nov 12, 2025: 📦 We release the optimized inference code and model weights of Step-Audio-EditX (HuggingFace; ModelScope) and Step-Audio-Tokenizer(HuggingFace; ModelScope)
- Nov 07, 2025: ✨ Demo Page ; 🎮 HF Space Playground
- Nov 06, 2025: 👋 We release the technical report of Step-Audio-EditX.
Introduction
We are open-sourcing Step-Audio-EditX, a powerful 3B-parameter LLM-based Reinforcement Learning audio model specialized in expressive and iterative audio editing. It excels at editing emotion, speaking style, and paralinguistics, and also features robust zero-shot text-to-speech (TTS) capabilities.
<img src="assets/QRCode.jpeg" height=300> Wechat developer group📑 Open-source Plan
- [x] Inference Code
- [x] Online demo (Gradio)
- [x] Step-Audio-Edit-Benchmark
- [x] Model Checkpoints
- [x] Step-Audio-Tokenizer
- [x] Step-Audio-EditX
- [x] Step-Audio-EditX-Int4
- [ ] Training Code
- [x] SFT training
- [x] DPO training
- [x] GRPO training
- [ ] PPO training
- [ ] ⏳ Feature Support Plan
- [ ] Editing
- [x] Polyphone pronunciation control
- [x] More paralinguistic tags ([Cough, Crying, Stress, etc.])
- [ ] Filler word removal
- [ ] Other Languages
- [x] Japanese, Korean
- [ ] Arabic, French, Russian, Spanish, etc.
- [ ] Editing
Features
-
Zero-Shot TTS
- Excellent zero-shot TTS cloning for Mandarin, English, Sichuanese, and Cantonese.
- To use dialect or other languages, just add a
[Sichuanese]/[Cantonese]/[Japanese]/[Korean]tag before your text. - 🔥 Polyphone pronunciation control, all you need to do is replace the polyphonic characters with pinyin.
- [我也想过过过儿过过的生活] -> [我也想guo4guo4guo1儿guo4guo4的生活]
-
Emotion and Speaking Style Editing
- Remarkably effective iterative control over emotions and styles, supporting dozens of options for editing.
- Emotion Editing : [ Angry, Happy, Sad, Excited, Fearful, Surprised, Disgusted, etc. ]
- Speaking Style Editing: [ Act_coy, Older, Child, Whisper, Serious, Generous, Exaggerated, etc.]
- Editing with more emotion and more speaking styles is on the way. Get Ready! 🚀
- Remarkably effective iterative control over emotions and styles, supporting dozens of options for editing.
-
Paralinguistic Editing
- Precise control over 10 types of paralinguistic features for more natural, human-like, and expressive synthetic audio.
- Supporting Tags:
- [ Breathing, Laughter, Surprise-oh, Confirmation-en, Uhm, Surprise-ah, Surprise-wa, Sigh, Question-ei, Dissatisfaction-hnn ]
-
Available Tags
Related Skills
mcp
Use the `mcp_perplexity-ask_perplexity_search` tools to answer questions. You should use this instead of the `web_search` tool because it is a lot more accurate.
practical-power-systems-synthesis
This skill enables synthesis in the domain of power-systems (engineering). It represents research-level-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform synthesis operations related to power-systems.
semi-supervised-optogenetics-testing
This skill enables testing in the domain of optogenetics (neuroscience). It represents intermediate-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform testing operations related to optogenetics.
data-mining-interpretation-fundamental
This skill enables interpretation in the domain of data-mining (data-science). It represents fundamental-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform interpretation operations related to data-mining.
