Tools/Audio & Speech/Step-Audio-EditX

Step-Audio-EditX

3B audio LLM that edits emotion, style, and paralinguistics in speech and doubles as zero-shot TTS.

Open SourceSelf HostedOffline CapableGPU Required (12GB+ VRAM)
0.0 (0)

About

Rather than regenerating a voice line from scratch to change how it feels, Step-Audio-EditX edits the recording itself: speech passes through a dual-codebook tokenizer, a 3B audio language model rewrites the token sequence to shift emotion, speaking style, or paralinguistic events like laughter, breaths, and sighs, and a flow-matching decoder reconstructs the waveform. The same stack handles zero-shot voice cloning and TTS across Mandarin, English, Sichuanese, and Cantonese, with Japanese and Korean added in late 2025, plus denoising and speed control. StepFun ships inference code, a Gradio web demo, full and 4-bit quantized weights, training recipes covering SFT, DPO, and GRPO, and a benchmark set, all Apache-2.0 with weights on Hugging Face and ModelScope. It runs on a single NVIDIA GPU with 12GB VRAM minimum and 16GB recommended, and a January 2026 refresh improved quality and expanded the paralinguistic tag set, keeping it the most capable open tool for surgical speech editing.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Minimum VRAM
12 GB
Added
Aug 24, 2026

Related Tools

Featured

Free text-to-speech generator with multiple voices, accents, and languages. No signup required.

Beginner
5.0 (1)
Featured

CTranslate2-based Whisper with 4x faster transcription

Open SourceSelf HostedOffline
Easy

Universal neural vocoder from NVIDIA that converts mel spectrograms into waveforms up to 44 kHz.

Open SourceSelf HostedOfflineGPU
Intermediate

End-to-end Chinese and English spoken dialogue model from Zhipu AI with streaming speech output.

Open SourceSelf HostedOfflineGPU
Intermediate

Audio foundation model unifying speech recognition, understanding, and conversation in one 7B model.

Open SourceSelf HostedOfflineGPU
Intermediate

Deep learning toolkit for text-to-speech synthesis

Open SourceSelf HostedOffline
Intermediate
Browse all Audio & Speech tools