Maya1

Expressive 3B TTS that designs voices from plain-English descriptions with inline emotion tags.

Open SourceSelf HostedOffline CapableGPU Required (16GB+ VRAM)
0.0 (0)

About

Voice design in Maya1 happens in prose: describe a speaker as a raspy 40-year-old baritone or a bright teenage narrator and the 3B model synthesizes that voice directly, no reference audio or cloning session required. Built as a Llama-style transformer, it predicts tokens for the SNAC neural codec, seven per audio frame at about 0.98 kbps, which makes low-latency streaming synthesis natural, and inline tags such as laugh and cry inject emotion mid-sentence for game characters, podcasts, and voice assistants. Maya Research released the weights, tokenizer, and inference scripts on Hugging Face under Apache-2.0 with commercial use allowed, plus a vLLM script for real-time streaming deployments. English is the focus at launch with multi-accent coverage and broader language support on the roadmap. A single GPU with 16GB of VRAM, anything from an RTX 4090 upward, runs it comfortably, and the November 2025 release quickly became the reference point for open description-driven voice generation.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Easy (2/5)
License
Apache-2.0
Minimum VRAM
16 GB
Added
Aug 24, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
Browse all Text-to-Speech (TTS) tools