NVIDIA Cosmos
NVIDIA's open world foundation models that generate physics-aware video for robotics and autonomy.
About
Robotics and autonomous-vehicle teams use Cosmos as a source of physics-aware synthetic video and embodied reasoning. The platform spans three open model families plus tooling: Cosmos-Predict, now at version 2.5, simulates the future state of a scene as video, Cosmos-Transfer 2.5 renders high-quality world simulations conditioned on spatial control inputs such as depth or segmentation maps, and Cosmos-Reason 2 handles physical common sense, producing embodied decisions in natural language. Around the models sit Cosmos-RL, a reinforcement learning framework built for physical AI, and the Cosmos-Cookbook of post-training scripts for adapting the checkpoints to a specific robot or driving domain. Code across the primary repositories is Apache 2.0, with model weights distributed under the permissive NVIDIA Open Model License, so the whole stack self-hosts on your own GPUs; earlier Predict1, Transfer1, and Reason1 generations remain available alongside the current releases. NVIDIA positions the family as trained infrastructure for physical AI that developers post-train on their own data rather than a consumer video toy.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Video Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026
Related Tools
Open-source video generation model by Tencent with text and image conditioning.
Image-to-video generation model by Alibaba DAMO Academy.
Updated CogVideo model by Zhipu AI with improved video quality.
Infinite-length music-driven video generation with visual conditioning.
Text-to-video generation framework with cascaded latent diffusion.
Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.