SenseVoice

Multilingual model covering speech recognition, emotion recognition, and audio event detection.

Open SourceSelf HostedOffline Capable
0.0 (0)

About

SenseVoice is the speech understanding member of Alibaba's FunAudioLLM family, sharing a lineage with the CosyVoice synthesis models. The released SenseVoice-Small checkpoint handles automatic speech recognition, spoken language identification, speech emotion recognition, and audio event detection in one non-autoregressive model covering Mandarin, Cantonese, English, Japanese, and Korean, drawn from a training program of more than 400,000 hours of audio. Because decoding is non-autoregressive, the authors report inference more than 5 times faster than Whisper-Small and 15 times faster than Whisper-Large at comparable parameter counts, with an accuracy edge on Chinese and Cantonese test sets such as AISHELL-1, AISHELL-2, and WenetSpeech. It runs on CPU or GPU through pip install funasr, and exports to ONNX, LibTorch, and GGUF for llama.cpp-style deployment. The repository code is MIT licensed while the weights use the FunASR model license, which permits commercial use with attribution; the project has around 9,000 GitHub stars and is widely used for Chinese, Japanese, and Korean transcription.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Easy (2/5)
License
MIT
Added
Jul 29, 2026

Related Tools

Convolution-augmented transformer for speech recognition in ESPnet toolkit.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

End-to-end speech processing toolkit covering ASR, TTS, and speech translation.

Open SourceSelf HostedOfflineGPU 8GB+
Expert
0.0 (0)

CLI tool that transcribes audio 10x faster using pipeline optimizations.

Open SourceSelf HostedOfflineGPU 6GB+
Easy
0.0 (0)

Established speech recognition toolkit used in research and production systems.

Open SourceSelf HostedOffline
Expert
0.0 (0)

Self-supervised speech representation model by Meta for ASR.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Multilingual ASR model by NVIDIA supporting 4 languages with translation.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)
Browse all Speech-to-Text / Speech Recognition tools