Tools/Video Generation/NVIDIA Cosmos

NVIDIA Cosmos

NVIDIA's open world foundation models that generate physics-aware video for robotics and autonomy.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

Robotics and autonomous-vehicle teams use Cosmos as a source of physics-aware synthetic video and embodied reasoning. The platform spans three open model families plus tooling: Cosmos-Predict, now at version 2.5, simulates the future state of a scene as video, Cosmos-Transfer 2.5 renders high-quality world simulations conditioned on spatial control inputs such as depth or segmentation maps, and Cosmos-Reason 2 handles physical common sense, producing embodied decisions in natural language. Around the models sit Cosmos-RL, a reinforcement learning framework built for physical AI, and the Cosmos-Cookbook of post-training scripts for adapting the checkpoints to a specific robot or driving domain. Code across the primary repositories is Apache 2.0, with model weights distributed under the permissive NVIDIA Open Model License, so the whole stack self-hosts on your own GPUs; earlier Predict1, Transfer1, and Reason1 generations remain available alongside the current releases. NVIDIA positions the family as trained infrastructure for physical AI that developers post-train on their own data rather than a consumer video toy.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Advanced (4/5)
License
Apache-2.0
Added
Aug 24, 2026

Related Tools

Featured

Open-source video generation model by Tencent with text and image conditioning.

Open SourceSelf HostedOfflineGPU 24GB+
Advanced

Image-to-video generation model by Alibaba DAMO Academy.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced

Updated CogVideo model by Zhipu AI with improved video quality.

Open SourceSelf HostedOfflineGPU 16GB+
Advanced

Infinite-length music-driven video generation with visual conditioning.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced

Text-to-video generation framework with cascaded latent diffusion.

Open SourceSelf HostedOfflineGPU 16GB+
Advanced

Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced
Browse all Video Generation tools