A Survey on Video Diffusion Models

Snapshot 2026-08-04 16:17:00 UTC · version 1

published
C
Collider.club487 cards · 9.8/10 MDRSS

Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, Yu-Gang Jiang. Use it to navigate the topic and choose relevant methods, papers or tools.

MARKDOWN SNAPSHOT

Loading…

Direct .mdRaw + metadata0 commentsMDRSS 9.8/10
INDEXABLE MARKDOWN SNAPSHOT

Research document

Open canonical .md

A Survey on Video Diffusion Models

Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, Yu-Gang Jiang. Use it to navigate the topic and choose relevant methods, papers or tools.

Editorial note: curated source snapshot published by Collider.club under the MIT License. Source attribution is preserved in the front matter.

Source snapshot

A Survey on Video Diffusion Models

(Source: Make-A-Video, SimDA, PYoCo, SVD , Video LDM and Tune-A-Video)

  • [News] The updated version is available on arXiv.
  • [News] Our survey is accepted by ACM Computing Surveys (CSUR).
  • [News] The Chinese translation is available on Zhihu. Special thanks to Dai-Wenxun for this.

Contact

If you have any suggestions or find our work helpful, feel free to contact us

Homepage: Zhen Xing

Email: zhenxingfd@gmail.com

If you find our survey is useful in your research or applications, please consider giving us a star 🌟 and citing it by the following BibTeX entry.

@article{xing2023survey,
  title={A survey on video diffusion models},
  author={Xing, Zhen and Feng, Qijun and Chen, Haoran and Dai, Qi and Hu, Han and Xu, Hang and Wu, Zuxuan and Jiang, Yu-Gang},
  journal={ACM Computing Surveys},
  year={2023},
  publisher={ACM New York, NY}
}

Open-source Toolboxes and Foundation Models

Methods Task Github
Helios T2V Generation
Movie Gen T2V Generation -
CogVideoX T2V Generation
Open-Sora-Plan T2V Generation
Open-Sora T2V Generation
Morph Studio T2V Generation -
Genie T2V Generation -
Sora T2V Generation & Editing -
VideoPoet T2V Generation & Editing -
Stable Video Diffusion T2V Generation
NeverEnds T2V Generation -
Pika T2V Generation -
EMU-Video T2V Generation -
GEN-2 T2V Generation & Editing -
ModelScope T2V Generation
ZeroScope T2V Generation -
T2V Synthesis Colab T2V Genetation
VideoCraft T2V Genetation & Editing
Diffusers (T2V synthesis) T2V Genetation -
AnimateDiff Personalized T2V Genetation
Text2Video-Zero T2V Genetation
HotShot-XL T2V Genetation
Genmo T2V Genetation -
Fliki T2V Generation -
Seedream AI Studio Image Generation + I2V Animation -

Table of Contents

Video Generation

Data

Caption-level

Title arXiv Github WebSite Pub. & Date
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation May, 2025
Identity-Preserving Text-to-Video Generation by Frequency Decomposition CVPR, 2025
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation NeurIPS, 2024
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers CVPR, 2024
CelebV-Text: A Large-Scale Facial Text-Video Dataset - CVPR, 2023
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation - May, 2023
VideoFactory: Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation - - May, 2023
Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions - - Nov, 2021
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval - - ICCV, 2021
MSR-VTT: A Large Video Description Dataset for Bridging Video and Language - - CVPR, 2016

Category-level

Title arXiv Github WebSite Pub. & Date
UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild - - Dec., 2012
First Order Motion Model for Image Animation - - May, 2023
Learning to Generate Time-Lapse Videos Using Multi-Stage Dynamic Generative Adversarial Networks - - CVPR,2018

Metric and BenchMark

Title arXiv Github WebSite Pub. & Date
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation May, 2025
Fréchet Video Motion Distance: A Metric for Evaluating Motion Consistency in Videos - Jul., 2024
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation NeurIPS, 2024
STREAM: Spatio-TempoRal Evaluation and Analysis Metric for Video Generative Models - ICLR, 2024
Subjective-Aligned Dateset and Metric for Text-to-Video Quality Assessment - - Mar, 2024
Towards A Better Metric for Text-to-Video Generation - Jan, 2024
AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI - - Jan, 2024
VBench: Comprehensive Benchmark Suite for Video Generative Models Nov, 2023
FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation - - NeurIPS, 2023
CVPR 2023 Text Guided Video Editing Competition - - Oct., 2023
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models Oct., 2023
Measuring the Quality of Text-to-Video Model Outputs: Metrics and Dataset - - Sep., 2023

Text-to-Video Generation

Training-based

Title arXiv Github WebSite Pub. & Date
Helios: Real Real-Time Long Video Generation Model Arxiv, 2026
Identity-Preserving Text-to-Video Generation by Frequency Decomposition CVPR, 2025
Enhancing Motion in Text-to-Video Generation with Decomposed Encoding and Conditioning NeurIPS 2024
Movie Gen - Oct, 2024
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer - Oct, 2024
Grid Diffusion Models for Text-to-Video Generation CVPR, 2024
MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators Apr., 2024
Mora: Enabling Generalist Video Generation via A Multi-Agent Framework - - Mar., 2024
VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis - - Mar., 2024
Genie: Generative Interactive Environments - Feb., 2024
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis - Feb., 2024
Lumiere: A Space-Time Diffusion Model for Video Generation - Jan, 2024
UNIVG: TOWARDS UNIFIED-MODAL VIDEO GENERATION - Jan, 2024
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models Jan, 2024
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model - Jan, 2024
MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation - Jan, 2024
VideoDrafter: Content-Consistent Multi-Scene Video Generation with LLM - Jan, 2024
A Recipe for Scaling up Text-to-Video Generation with Text-free Videos Dec, 2023
InstructVideo: Instructing Video Diffusion Models with Human Feedback Dec, 2023
VideoLCM: Video Latent Consistency Model - - Dec, 2023
Photorealistic Video Generation with Diffusion Models - Dec, 2023
Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation Dec, 2023
Delving Deep into Diffusion Transformers for Image and Video Generation - Dec, 2023
StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter Nov, 2023
MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation - Nov, 2023
ART•V: Auto-Regressive Text-to-Video Generation with Diffusion Models Nov, 2023
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Nov, 2023
FusionFrames: Efficient Architectural Aspects for Text-to-Video Generation Pipeline Nov, 2023
MoVideo: Motion-Aware Video Generation with Diffusion Models - Nov, 2023
Make Pixels Dance: High-Dynamic Video Generation - Nov, 2023
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning - Nov, 2023
Optimal Noise pursuit for Augmenting Text-to-Video Generation - - Nov, 2023
VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning - Nov, 2023
VideoCrafter1: Open Diffusion Models for High-Quality Video Generation Oct, 2023
SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction Oct, 2023
DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors Oct., 2023
LAMP: Learn A Motion Pattern for Few-Shot-Based Video Generation Oct., 2023
DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model Oct, 2023
MotionDirector: Motion Customization of Text-to-Video Diffusion Models Oct, 2023
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning Sep., 2023
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation Sep., 2023
LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models Sep., 2023
Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation Sep., 2023
VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation - Sep., 2023
MobileVidFactory: Automatic Diffusion-Based Social Media Video Generation for Mobile Devices from Text - - Jul., 2023
Text2Performer: Text-Driven Human Video Generation Apr., 2023
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning Jul., 2023
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with Large Language Models - Aug., 2023
SimDA: Simple Diffusion Adapter for Efficient Video Generation CVPR, 2024
Dual-Stream Diffusion Net for Text-to-Video Generation - - Aug., 2023
ModelScope Text-to-Video Technical Report Aug., 2023
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation - Jul., 2023
VideoFactory: Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation - - May, 2023
Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models - May, 2023
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models - -
Latent-Shift: Latent Diffusion with Temporal Shift - -
Probabilistic Adaptation of Text-to-Video Models - Jun., 2023
NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation - Mar., 2023
ED-T2V: An Efficient Training Framework for Diffusion-based Text-to-Video Generation - - - IJCNN, 2023
MagicVideo: Efficient Video Generation With Latent Diffusion Models - -
Phenaki: Variable Length Video Generation From Open Domain Textual Description - -
Imagen Video: High Definition Video Generation With Diffusion Models - -
VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation -
MAGVIT: Masked Generative Video Transformer - Dec., 2022
Make-A-Video: Text-to-Video Generation without Text-Video Data - -
Latent Video Diffusion Models for High-Fidelity Video Generation With Arbitrary Lengths Nov., 2022
CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers - May, 2022
Video Diffusion Models - -

Training-free

Title arXiv Github WebSite Pub. & Date
A²RD: Agentic Autoregressive Diffusion for Long Video Consistency - May, 2026
VISTA: A Test-Time Self-Improving Video Generation Agent - CVPR, 2026
InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO - May, 2025
VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models Mar, 2024
TRAILBLAZER: TRAJECTORY CONTROL FOR DIFFUSION-BASED VIDEO GENERATION Jan, 2024
FreeInit: Bridging Initialization Gap in Video Diffusion Models Dec, 2023
MTVG : Multi-text Video Generation with Text-to-Video Models - Dec, 2023
F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis - - Nov, 2023
AdaDiff: Adaptive Step Selection for Fast Diffusion - - Nov, 2023
FlowZero: Zero-Shot Text-to-Video Synthesis with LLM-Driven Dynamic Scene Syntax Nov, 2023
🏀GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning Nov, 2023
FreeNoise: Tuning-Free Longer Video Diffusion Via Noise Rescheduling Oct, 2023
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation Oct, 2023
LLM-grounded Video Diffusion Models Oct, 2023
Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator - NeurIPS, 2023
DiffSynth: Latent In-Iteration Deflickering for Realistic Video Synthesis Aug, 2023
Large Language Models are Frame-level Directors for Zero-shot Text-to-Video Generation - May, 2023
Text2video-Zero: Text-to-Image Diffusion Models Are Zero-Shot Video Generators Mar., 2023
PEEKABOO: Interactive Video Generation via Masked-Diffusion 🫣 CVPR, 2024

Video Generation with other conditions

Pose-guided Video Generation

Title arXiv Github WebSite Pub. & Date
Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures - Feb., 2026
EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses CVPR 2026
🔥🔥StableAnimator: High-Quality Identity-Preserving Human Image Animation🔥🔥 Nov., 2024
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model ECCV 2024
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance Jul., 2024
Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance Mar., 2024
Action Reimagined: Text-to-Pose Video Editing for Dynamic Human Actions - - Mar., 2024
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons - - Jan., 2024
DreaMoving: A Human Dance Video Generation Framework based on Diffusion Models - Dec., 2023
MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model Nov., 2023
Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation Nov., 2023
MagicDance: Realistic Human Dance Video Generation with Motions & Facial Expressions Transfer Nov., 2023
DisCo: Disentangled Control for Referring Human Dance Generation in Real World Jul., 2023
Dancing Avatar: Pose and Text-Guided Human Motion Videos Synthesis with Image Diffusion Model - - Aug., 2023
DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion Apr., 2023
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos Apr., 2023

Motion-guided Video Generation

Title arXiv Github WebSite Pub. & Date
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance Mar. 2025
MOTIONCLONE: TRAINING-FREE MOTION CLONING FOR CONTROLLABLE VIDEO GENERATION Jun., 2024
Tora: Trajectory-oriented Diffusion Transformer for Video Generation CVPR 2025
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model ECCV 2024
Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance Mar., 2024
Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling - - Jan., 2024

This HTML preview is truncated for page performance. The canonical Markdown file contains the complete snapshot.

MARKDOWN METRICS
6168words
41headings
1162links
1code blocks
MDRSS ASSESSMENT
Scam / risk5/100low
Evidence100/100high confidence
Why MDRSS assigned this score
  • evidence comes from multiple domains
  • some evidence URLs look like primary-source hosts
Evidence (4)
concept:image-video-and-creative-aiorg:collider-club

Discussion 0

Sign in to join the discussion.