Phantom

ByteDance framework that keeps reference people and objects consistent in generated video.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

Keeping the same face across every frame is the problem Phantom tackles: ByteDance's subject-to-video framework generates clips in which reference people and objects stay consistent, trained through cross-modal alignment on text-image-video triplet data. It builds on the Wan2.1 architecture, shipping Phantom-Wan checkpoints at 1.3B and 14B parameters that render at 480p or 720p, with the base Wan2.1-T2V-1.3B weights downloaded separately. Single and multi-subject generation are both supported, so a person and a product from separate reference images can appear together in one clip. Inference runs on a single GPU or scales out with FSDP and xDiT parallelism, and installation is a git clone plus requirements.txt with PyTorch 2.4 or newer. The work was accepted to ICCV 2025, the team released the companion Phantom-Data training dataset in June 2025, and a 14B Pro variant is planned. Code is Apache-2.0 with about 1,500 GitHub stars, and the paper is arXiv 2502.11079. Advertising, e-commerce, and character-driven content are the obvious use cases.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Added
Jul 29, 2026

Related Tools

Featured

Open-source video generation model by Tencent with text and image conditioning.

Open SourceSelf HostedOfflineGPU 24GB+
Advanced
0.0 (0)

Image-to-video generation model by Alibaba DAMO Academy.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced
0.0 (0)

Updated CogVideo model by Zhipu AI with improved video quality.

Open SourceSelf HostedOfflineGPU 16GB+
Advanced
0.0 (0)

Infinite-length music-driven video generation with visual conditioning.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced
0.0 (0)

Text-to-video generation framework with cascaded latent diffusion.

Open SourceSelf HostedOfflineGPU 16GB+
Advanced
0.0 (0)

Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced
0.0 (0)
Browse all Video Generation tools