Whisper Audio (Transcription+)
Audio processing toolkit building on Whisper for diarization and subtitling.
About
Stable-ts extends OpenAI Whisper to produce more reliable word and segment timestamps and adds tools for transcription workflows. It refines Whisper's native timing, supports regrouping and editing of segments, suppresses silent-region hallucinations, and can output subtitle formats. It works with any Whisper model size on CPU or GPU and is useful for subtitling, captioning, and audio editing. Open-source Python package.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Music & Audio Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Easy (2/5)
- License
- MIT
- Minimum VRAM
- 4 GB
- Added
- Apr 3, 2026
Related Tools
Latent diffusion model for text-to-audio, music, and speech generation.
Audio super-resolution model for upsampling audio to higher sample rates.
State-of-the-art music source separation model by Meta for splitting tracks.
Fast music generation model producing full songs with lyrics in seconds.
Audio diffusion model by Harmonai for generating music samples.
PyTorch library for deep learning research on audio generation including MusicGen and AudioGen.