MOSS-Transcribe
End-to-end model transcribing long multi-speaker audio with speaker labels and timestamps in one pass.
About
Meeting transcription normally chains ASR, voice activity detection, and a separate diarization model; MOSS-Transcribe-Diarize folds the whole job into one end-to-end model that ingests long multi-speaker recordings and emits a single stream of speaker-labeled, timestamped text with acoustic events included. The 0.9B open release pairs a Whisper-Medium-shaped audio encoder with a Qwen3-0.6B-style decoder, covers more than 50 languages, and took first place in the second MLC-SLM challenge at Interspeech 2026. OpenMOSS ships a Python package, serving recipes for vLLM and SGLang, a local subtitle web app exporting JSON, SRT, and ASS files, and a fine-tuning framework, all Apache-2.0 with weights on Hugging Face; a larger Pro model sits behind the same interface as a preview. Audio is processed at 16kHz in 30-second chunks and inference is GPU-oriented, with throughput quoted on H100 but a footprint small enough for consumer cards. Released July 2026, it is the strongest open answer to who-said-what transcription.
Should you use MOSS-Transcribe?
Pick it when
Pick MOSS-Transcribe when you transcribe long meetings, interviews, or podcasts and need who-said-what with timestamps in one pass, in any of 50+ languages, on a consumer GPU, with export to SRT or ASS subtitles.
Look elsewhere when
Skip it for single-speaker dictation or live captioning, where its chunked, GPU-oriented design adds little, and note it is a July 2026 release with a young ecosystem. WhisperX is the more established pipeline.
Alternatives to MOSS-Transcribe
- WhisperX
Whisper pipeline with forced word alignment and external diarization that also runs on CPU, more moving parts than one model but a longer track record than a July 2026 release.
- Reverb
ASR plus diarization tuned on English long-form audio with verbatim control, but its weights are limited to non-production use.
- Pyannote Audio
Diarization only, to pair with a recognizer you already trust, instead of replacing your ASR with one combined model.
- Granite Speech
Stronger fit for single-speaker transcription and translation in six languages, but it produces no speaker labels.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026