TangoFlux

Text-to-audio model using flow matching that makes up to 30 seconds of 44.1 kHz audio in 3.7 seconds.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

Text prompts become 44.1 kHz audio in roughly 3.7 seconds with TangoFlux, a 515M parameter text-to-audio model from the DeCLaRe Lab at Singapore University of Technology and Design, produced with NVIDIA and Lambda and appearing at ICLR 2026. The architecture stacks FluxTransformer blocks, Diffusion Transformer and Multimodal Diffusion Transformer layers conditioned on text and duration embeddings, trained with flow matching over rectified flow trajectories; a final CRPO stage, CLAP-Ranked Preference Optimization, builds preference pairs ranked by the CLAP model. Generating up to 30 seconds of audio takes about 3.7 seconds on a single NVIDIA A40, with 25 sampling steps the default and 50 recommended for higher quality. The repository installs via pip from GitHub and provides a CLI, Python API, Gradio demo, and a ComfyUI integration, with weights hosted on Hugging Face. Licensing is the main caveat: the model falls under the Stability AI Community License with WavCaps-derived data restricted to academic use, making it research-only, so commercial deployment needs separate licensing. A GPU is effectively required.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Stability AI Community License
Added
Jul 29, 2026

Related Tools

Featured

Audio generation framework by Meta including MusicGen for text-to-music.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)

Latent diffusion model for text-to-audio, music, and speech generation.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)

Audio super-resolution model for upsampling audio to higher sample rates.

Open SourceSelf HostedOfflineGPU 6GB+
Intermediate
0.0 (0)
Featured

State-of-the-art music source separation model by Meta for splitting tracks.

Open SourceSelf HostedOffline
Easy
0.0 (0)

Fast music generation model producing full songs with lyrics in seconds.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)

PyTorch library for deep learning research on audio generation including MusicGen and AudioGen.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)
Browse all Music & Audio Generation tools