Step-Audio-EditX
3B audio LLM that edits emotion, style, and paralinguistics in speech and doubles as zero-shot TTS.
About
Rather than regenerating a voice line from scratch to change how it feels, Step-Audio-EditX edits the recording itself: speech passes through a dual-codebook tokenizer, a 3B audio language model rewrites the token sequence to shift emotion, speaking style, or paralinguistic events like laughter, breaths, and sighs, and a flow-matching decoder reconstructs the waveform. The same stack handles zero-shot voice cloning and TTS across Mandarin, English, Sichuanese, and Cantonese, with Japanese and Korean added in late 2025, plus denoising and speed control. StepFun ships inference code, a Gradio web demo, full and 4-bit quantized weights, training recipes covering SFT, DPO, and GRPO, and a benchmark set, all Apache-2.0 with weights on Hugging Face and ModelScope. It runs on a single NVIDIA GPU with 12GB VRAM minimum and 16GB recommended, and a January 2026 refresh improved quality and expanded the paralinguistic tag set, keeping it the most capable open tool for surgical speech editing.
Should you use Step-Audio-EditX?
Pick it when
Pick Step-Audio-EditX when you need to change how an existing recording sounds, shifting emotion or style or adding laughs and breaths, without re-recording, and you have an NVIDIA GPU with 12 to 16 GB of VRAM.
Look elsewhere when
Skip it if you only need fresh narration from text, where Chatterbox or Kokoro need far less VRAM, or if your languages fall outside Mandarin, English, Sichuanese, Cantonese, Japanese, and Korean.
Alternatives to Step-Audio-EditX
- VoiceCraft
Also edits existing recordings, via token infilling on about 8 GB of VRAM, but its code is CC BY-NC-SA and its weights non-commercial, where this tool is Apache-2.0.
- CosyVoice 2
Instruction control over emotion, dialect, and rate plus 150 ms streaming on about 8 GB, but it regenerates speech instead of editing a take.
- Chatterbox TTS
Emotion and accent control in zero-shot TTS on a 4 GB GPU, far lighter, but it synthesizes new audio and cannot edit an existing recording.
- Orpheus TTS
Inline laugh and sigh tags with about 200 ms streaming for live agents, but English focused and generation only, with no editing of recorded speech.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Audio & Speech
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 12 GB
- Added
- Aug 24, 2026