ConsistI2V
View on GitHubConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation [TMLR 2024]
Official code for ConsistI2V (TMLR 2024), a diffusion method that turns a still image plus text prompt into a consistent video using first-frame spatiotemporal attention and low-frequency noise initialization. Includes inference and training scripts, a Gradio demo, and Hugging Face weights.
Use Cases
Animate a single still image into a short video clip from a text promptGenerate visually consistent image-to-video results for creative and film previsualizationRun local Gradio demo for I2V generationDeploy I2V generation on Replicate or Hugging Face SpacesFine-tune the diffusion I2V model on a custom video-caption datasetBenchmark I2V consistency with the I2V-Bench datasetResearch spatiotemporal attention and noise initialization techniques
Built With
- Language
- Python
- Frameworks
- PyTorch · Hugging Face Diffusers · AnimateDiff · FreeInit · Gradio · Cog · Weights & Biases
Tags
image-to-video · video-generation · diffusion-models · generative-ai · visual-consistency · spatiotemporal-attention · video-synthesis · pytorch · gradio · research-paper · tmlr-2024 · text-prompted-video · first-frame-conditioning · inference · training