Step-Audio-EditX
3B audio LLM that edits emotion, style, and paralinguistics in speech and doubles as zero-shot TTS.
About
Rather than regenerating a voice line from scratch to change how it feels, Step-Audio-EditX edits the recording itself: speech passes through a dual-codebook tokenizer, a 3B audio language model rewrites the token sequence to shift emotion, speaking style, or paralinguistic events like laughter, breaths, and sighs, and a flow-matching decoder reconstructs the waveform. The same stack handles zero-shot voice cloning and TTS across Mandarin, English, Sichuanese, and Cantonese, with Japanese and Korean added in late 2025, plus denoising and speed control. StepFun ships inference code, a Gradio web demo, full and 4-bit quantized weights, training recipes covering SFT, DPO, and GRPO, and a benchmark set, all Apache-2.0 with weights on Hugging Face and ModelScope. It runs on a single NVIDIA GPU with 12GB VRAM minimum and 16GB recommended, and a January 2026 refresh improved quality and expanded the paralinguistic tag set, keeping it the most capable open tool for surgical speech editing.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Audio & Speech
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 12 GB
- Added
- Aug 24, 2026
Related Tools
Free text-to-speech generator with multiple voices, accents, and languages. No signup required.
CTranslate2-based Whisper with 4x faster transcription
Universal neural vocoder from NVIDIA that converts mel spectrograms into waveforms up to 44 kHz.
End-to-end Chinese and English spoken dialogue model from Zhipu AI with streaming speech output.
Audio foundation model unifying speech recognition, understanding, and conversation in one 7B model.
Deep learning toolkit for text-to-speech synthesis