TEN Framework
Open source framework for building real-time multimodal conversational voice agents.
About
Voice agents that listen, think, and speak in real time are the target workload of the TEN Framework, an open source runtime for multimodal conversational AI backed by Agora. Pipelines are assembled from extensions written in Python, Go, C++, or TypeScript and connected through a low-latency runtime that handles full-duplex audio and video. The ecosystem ships more than the core: TEN VAD for voice activity detection, TEN Turn Detection for natural interruption handling, and agent examples covering voice assistants, speaker diarization, and lip-synced avatars. Deployment is Docker Compose based, using an Agora App ID for real-time transport, and the bundled examples integrate hosted services such as OpenAI, Deepgram, and ElevenLabs, so typical agents pair local components with cloud speech and LLM APIs. The license is Apache-2.0 with additional restrictions defined in the project's license file, worth reviewing before commercial deployment, though the framework is free to use and self-host. With roughly 11,000 GitHub stars it is among the most visible open foundations for production voice agent development.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Audio & Speech
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0 (with additional restrictions)
- Added
- Jul 29, 2026
Related Tools
Free text-to-speech generator with multiple voices, accents, and languages. No signup required.
Deep learning toolkit for text-to-speech synthesis
CTranslate2-based Whisper with 4x faster transcription
Universal neural vocoder from NVIDIA that converts mel spectrograms into waveforms up to 44 kHz.
End-to-end Chinese and English spoken dialogue model from Zhipu AI with streaming speech output.
Transformer-based text-to-audio model from Suno