LongCat-Video
View on GitHubLongCat-Video is Meituan's 13.6B video generation foundation model, unifying text-to-video, image-to-video and video-continuation plus audio-driven avatar animation. Ships PyTorch inference demos (Streamlit, multi-GPU/context-parallel, FlashAttention-2/3, INT8) and open weights.
Use Cases
Text-to-video generationImage-to-video animationVideo continuation / extensionMinutes-long video generation without driftAudio-driven talking-head avatar videoMulti-character dialogue video from dual audio streamsStylized (anime/animal) avatar animationInteractive video generation720p 30fps fast video inference on single or multi-GPU
Built With
- Language
- Python
- Frameworks
- PyTorch · Diffusers · Transformers · Streamlit · FlashAttention · xformers
Tags
text-to-video · image-to-video · video-generation · diffusion · avatar · audio-driven · lip-sync · long-video · dit · pytorch · streamlit · flash-attention · distillation · int8-quantization · world-models · inference