Phantom
ByteDance framework that keeps reference people and objects consistent in generated video.
About
Keeping the same face across every frame is the problem Phantom tackles: ByteDance's subject-to-video framework generates clips in which reference people and objects stay consistent, trained through cross-modal alignment on text-image-video triplet data. It builds on the Wan2.1 architecture, shipping Phantom-Wan checkpoints at 1.3B and 14B parameters that render at 480p or 720p, with the base Wan2.1-T2V-1.3B weights downloaded separately. Single and multi-subject generation are both supported, so a person and a product from separate reference images can appear together in one clip. Inference runs on a single GPU or scales out with FSDP and xDiT parallelism, and installation is a git clone plus requirements.txt with PyTorch 2.4 or newer. The work was accepted to ICCV 2025, the team released the companion Phantom-Data training dataset in June 2025, and a 14B Pro variant is planned. Code is Apache-2.0 with about 1,500 GitHub stars, and the paper is arXiv 2502.11079. Advertising, e-commerce, and character-driven content are the obvious use cases.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Video Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Added
- Jul 29, 2026
Related Tools
Open-source video generation model by Tencent with text and image conditioning.
Image-to-video generation model by Alibaba DAMO Academy.
Updated CogVideo model by Zhipu AI with improved video quality.
Infinite-length music-driven video generation with visual conditioning.
Text-to-video generation framework with cascaded latent diffusion.
Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.