Maya1
Expressive 3B TTS that designs voices from plain-English descriptions with inline emotion tags.
About
Voice design in Maya1 happens in prose: describe a speaker as a raspy 40-year-old baritone or a bright teenage narrator and the 3B model synthesizes that voice directly, no reference audio or cloning session required. Built as a Llama-style transformer, it predicts tokens for the SNAC neural codec, seven per audio frame at about 0.98 kbps, which makes low-latency streaming synthesis natural, and inline tags such as laugh and cry inject emotion mid-sentence for game characters, podcasts, and voice assistants. Maya Research released the weights, tokenizer, and inference scripts on Hugging Face under Apache-2.0 with commercial use allowed, plus a vLLM script for real-time streaming deployments. English is the focus at launch with multi-accent coverage and broader language support on the roadmap. A single GPU with 16GB of VRAM, anything from an RTX 4090 upward, runs it comfortably, and the November 2025 release quickly became the reference point for open description-driven voice generation.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Easy (2/5)
- License
- Apache-2.0
- Minimum VRAM
- 16 GB
- Added
- Aug 24, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.