Evoke
Interactive autoregressive world model that keeps generated scenes persistent under live camera control.
About
World models usually forget what is behind the camera; Evoke tackles that with an external world state bank that stores scene geometry indexed by camera pose, letting rollouts revisit earlier viewpoints without drift. The AlayaLab model is a 14B autoregressive diffusion transformer over latent chunks with multi-term memory carrying recent and distant context, distilled down to 3-step CFG-free sampling so each step costs one forward pass instead of two, and it leads the WBench benchmark against 50-step baselines. Users steer the camera and swap prompts mid-generation across text-to-video, image-to-video, and video-to-video modes. On a single H200 it produces 1.5 seconds of 384x640, 24fps video every 2.11 seconds, approaching real-time interaction. The August 2026 release ships training and inference code, four staged checkpoints plus teacher weights, and example data under Apache-2.0, though some vendored depth components carry CC-BY-NC terms. Setup is research-grade Python with datacenter GPU expectations.
Should you use Evoke?
Pick it when
Pick Evoke if you research interactive world models and need scenes that stay consistent when the camera returns to earlier viewpoints, with live camera steering and prompt swaps, and you have datacenter GPUs such as an H200.
Look elsewhere when
Skip it for ordinary clip production: output is 384x640 and setup is research-grade. Check the vendored depth components, which carry CC-BY-NC terms, before commercial use. For robotics simulation, NVIDIA Cosmos has a broader toolchain.
Alternatives to Evoke
- NVIDIA Cosmos
World models for robotics and driving with RL and post-training tooling under permissive NVIDIA terms, but not built around camera-steered scene revisits.
- MAGI-1
Also autoregressive chunk-by-chunk generation with streaming, aimed at text and image-to-video rather than interactive camera control and scene memory.
- SkyReels-V2
Diffusion forcing gives unlimited-length 540p or 720p clips with camera control, higher resolution but no world-state memory for revisiting viewpoints.
- LongCat-Video
MIT-licensed minutes-long 720p, 30 fps generation, better for finished long-form footage but not interactive or steerable in real time.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Video Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026