Reverb

Rev's ASR and diarization models trained on 200,000 hours of human-transcribed English audio.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

Reverb is Rev's release of the speech models behind its transcription business: an ASR model built on the WeNet framework plus speaker diarization built on Pyannote, trained on 200,000 hours of human-transcribed English audio, which Rev describes as the largest such corpus ever used for an open model. A distinctive verbatimicity control slides output between strict verbatim, keeping every false start and filler, and a cleaned-up reading style suited to captions. On the long-form Earnings21 benchmark the ASR scores 9.68 percent word error rate against 14.26 percent for Whisper large-v3, and the diarization component reaches a 0.046 error rate on the same data. The pipeline installs with pip on Python 3.10 or newer or runs through Docker, and pulling the weights requires Hugging Face authentication plus Git LFS; a GPU is the practical choice for long-form workloads. The inference code is Apache-2.0, but the weights ship under the Rev Model Non-Production License, which limits them to testing, research, and evaluation, with commercial terms available from Rev, whose paper is at arXiv 2410.03930.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Freemium
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Rev Model Non-Production License
Added
Jul 29, 2026

Related Tools

Convolution-augmented transformer for speech recognition in ESPnet toolkit.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

End-to-end speech processing toolkit covering ASR, TTS, and speech translation.

Open SourceSelf HostedOfflineGPU 8GB+
Expert
0.0 (0)

CLI tool that transcribes audio 10x faster using pipeline optimizations.

Open SourceSelf HostedOfflineGPU 6GB+
Easy
0.0 (0)

Established speech recognition toolkit used in research and production systems.

Open SourceSelf HostedOffline
Expert
0.0 (0)

Self-supervised speech representation model by Meta for ASR.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Multilingual ASR model by NVIDIA supporting 4 languages with translation.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)
Browse all Speech-to-Text / Speech Recognition tools