A curated list of recent diffusion models for video generation, editing, restoration, understanding, nerf, etc. Use it to navigate the topic and choose relevant methods, papers or tools.
Awesome Video Diffusion
Snapshot 2026-08-04 16:17:00 UTC · version 1
Research document
Awesome Video Diffusion
A curated list of recent diffusion models for video generation, editing, restoration, understanding, nerf, etc. Use it to navigate the topic and choose relevant methods, papers or tools.
Editorial note: curated source snapshot published by Collider.club under the MIT License. Source attribution is preserved in the front matter.
Source snapshot
Awesome Video Diffusion
A curated list of recent diffusion models for video generation, editing, restoration, understanding, nerf, etc.
(Source: Make-A-Video, Tune-A-Video, and Fate/Zero.)
Table of Contents
- Open-source Toolboxes and Foundation Models
- Evaluation Benchmarks and Metrics
- Commercial Product
- Video Generation
- Efficient Video Generation
- Controllable Video Generation
- Character Customization
- Motion Customization
- Long Video / Film Generation
- Video Generation with 3D/Physical Prior
- Video Editing
- Human or Subject Motion
- Video Enhancement and Restoration
- Audio Synthesis for Video
- Talking Head Generation
- Reinforcement Learning for Video Generation
- Policy Learning
- Virtual Try-On
- 3D
- 4D
- Game Generation
- AI Safety
- Rendering with Virtual Engine
- Open-World Model
- Video Understanding
- Healthcare and Biology
- Other Applications
- Code-rendered Video Generation
Open-source Toolboxes and Foundation Models
NanoI2V
A step-by-step teaching series for building an Image-to-Video model from scratch in PyTorch, covering 3D VAEs, DiT, Flow Matching, RoPE, and conditioning.FastVideo: A unified inference and post-training framework for accelerated video generation
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
HunyuanVideo: A Systematic Framework For Large Video Generative Models
Pyramidal Flow Matching for Efficient Video Generative Modeling
VideoCrafter: A Toolkit for Text-to-Video Generation and Editing
Evaluation Benchmarks and Metrics
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference (Oct., 2025)
Stable Cinemetrics: Structured Taxonomy and Evaluation for Professional Video Generation (Sep., 2025)
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation (Mar., 2025)
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness (Mar., 2025)
Impossible Videos (Mar., 2025)
MEt3R: Measuring Multi-View Consistency in Generated Images (Jan., 2025)
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation (Dec., 2024)
Evaluation Agent, Efficient and Promptable Evaluation Framework for Visual Generative Models (Dec., 2024)
Frechet Video Motion Distance: A Metric for Evaluating Motion Consistency in Videos (Jun., 2024)
T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation (Jun., 2024)
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation (NeurIPS, 2024)
PEEKABOO: Interactive Video Generation via Masked-Diffusion (CVPR, 2024)
T2VScore: Towards A Better Metric for Text-to-Video Generation (Jan., 2024)
StoryBench: A Multifaceted Benchmark for Continuous Story Visualization (NeurIPS, 2023)
VBench: Comprehensive Benchmark Suite for Video Generative Models (Nov., 2023)
FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation (Nov., 2023)
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models (Oct., 2023)
Evaluation of Text-to-Video Generation Models: A Dynamics Perspective (Jul., 2024)
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models (May., 2024)
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers (CVPR, 2024)
ReLight My NeRF: A Dataset for Novel View Synthesis and Relighting of Real World Objects (CVPR, 2023)
Commercial Product
Video Generation
Helios: Real Real-Time Long Video Generation Model (Mar., 2026)
MOVA: Towards Scalable and Synchronized Video-Audio Generation (Feb., 2026)
Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation (Feb., 2026)
VINO: A Unified Visual Generator with Interleaved OmniModal Context (Jan., 2026)
UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers (Oct., 2025)
VISTA: A Test-Time Self-Improving Video Generation Agent (Oct., 2025 | CVPR 2026)
UniVideo: Unified Understanding, Generation, and Editing for Videos (Oct., 2025)
PUSA V1.0: Surpassing Wan-I2V with $500 Training Cost by Vectorized Timestep Adaptation (July., 2025)
LayerFlow : A Unified Model for Layer-aware Video Generation (May., 2025)
InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO (May., 2025)
Training-Free Efficient Video Generation via Dynamic Token Carving (May., 2025)
ReVision: High-Quality, Low-Cost Video Generation with Explicit 3D Physics Modeling for Complex Motion and Interaction (Apr., 2025)
Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis (Apr., 2025)
MAGI-1: Autoregressive Video Generation at Scale (Apr., 2025)
SphereDiff: Tuning-free Omnidirectional Panoramic Image and Video Generation via Spherical Latent Representation (Apr., 2025)
Packing Input Frame Context in Next-Frame Prediction Models for Video Generation (Apr., 2025)
SkyReels-V2: Infinite-length Film Generative Model (Apr., 2025)
Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model (Apr., 2025)
Aligning Text-to-Video Generation Models with Prompt Optimization (Mar., 2025)
Target-Aware Video Diffusion Models (Mar., 2025)
MagicComp: Training-free Dual-Phase Refinement for Compositional Video Generation (Mar., 2025)
Video-T1: Test-Time Scaling for Video Generation (Mar., 2025)
Temporal Regularization Makes Your Video Generator Stronger (Mar., 2025)
VACE: All-in-One Video Creation and Editing (Mar., 2025)
RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers (Feb., 2025)
DLFR-VAE: Dynamic Latent Frame Rate VAE for Video Generation (Feb., 2025)
Magic 1-For-1: Generating One Minute Video Clips within One Minute (Feb., 2025)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT (Feb., 2025)
Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search (Jan., 2025)
RepVideo: Rethinking Cross-Layer Representation for Video Generation (Jan., 2025)
Large Motion Video Autoencoding with Cross-modal Video VAE (Dec., 2024)
MotiF: Making Text Count in Image Animation with Motion Focal Loss (Dec., 2024)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation (Dec., 2024)
Autoregressive Video Generation without Vector Quantization (Dec., 2024)
AniDoc: Animation Creation Made Easier (Dec., 2024)
Video Diffusion Transformers are In-Context Learners (Dec., 2024)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation (Dec., 2024 | CVPR 2025)
Instructional Video Generation (Dec., 2024)
Mimir: Improving Video Diffusion Models for Precise Text Understanding (Dec., 2024)
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling (Dec., 2024)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition (Nov., 2024)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model (Nov., 2024)
VideoRepair: Improving Text-to-Video Generation via Misalignment Evaluation and Localized Refinement (Nov., 2024)
Enhancing Motion in Text-to-Video Generation with Decomposed Encoding and Conditioning (Oct., 2024 | NeurIPS 2024)
Improved Video VAE for Latent Video Diffusion Model (Oct., 2024)
T2V-Turbo-v2: Enhancing Video Generation Model Post-Training Through Data, Reward, and Conditional Guidance Design (Oct, 2024)
Progressive Autoregressive Video Diffusion Models (Oct., 2024)
Real-Time Video Generation with Pyramid Attention Broadcast (Aug., 2024)
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations (Aug., 2024)
CogVideoX: Text-to-video generation (Aug., 2024)
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention (Aug., 2024)
VEnhancer: Generative Space-Time Enhancement for Video Generation (Jul., 2024)
Live2Diff: Live Stream Translation via Uni-directional Attention in Video Diffusion Models (Jul., 2024)
Video Diffusion Alignment via Reward Gradient (Jul., 2024)
ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning (Jun., 2024)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance (Jul., 2024)
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model (Jun., 2024)
Video-Infinity: Distributed Long Video Generation (Jun., 2024)
MotionBooth: Motion-Aware Customized Text-to-Video Generation (Jun., 2024)
Text-Animator: Controllable Visual Text Video Generation (Jun., 2024)
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation (Jun., 2024)
T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback (May, 2024)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control (May, 2024)
Human4DiT: Free-view Human Video Generation with 4D Diffusion Transformer (May, 2024)
FIFO-Diffusion: Generating Infinite Videos from Text without Training (May, 2024)
Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models (May, 2024)
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers (May, 2024)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation (May, 2024)
TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models (CVPR 2024)
ID-Animator: Zero-Shot Identity-Preserving Human Video Generation (Apr., 2024)
AnimateZoo: Zero-shot Video Generation of Cross-Species Animation via Subject Alignment (Apr., 2024)
MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators (Apr., 2024)
TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models (CVPR 2024)
VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis (Mar., 2024)
StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text (Mar., 2024)
Intention-driven Ego-to-Exo Video Generation (Mar., 2024)
VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models (Mar., 2024)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis (Feb., 2024)
One-Shot Motion Customization of Text-to-Video Diffusion Models (Feb., 2024)
Magic-Me: Identity-Specific Video Customized Diffusion (Feb., 2024)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation (Feb., 2024)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion (Feb., 2024)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization (Feb., 2024)
Boximator: Generating Rich and Controllable Motions for Video Synthesis (Feb., 2024)
Lumiere: A Space-Time Diffusion Model for Video Generation (Jan., 2024)
ActAnywhere: Subject-Aware Video Background Generation (Jan., 2024)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens (Jan., 2024)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects (Jan., 2024)
UniVG: Towards UNIfied-modal Video Generation (Jan., 2024)
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models (Jan., 2024)
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model (Jan., 2024)
This HTML preview is truncated for page performance. The canonical Markdown file contains the complete snapshot.
Why MDRSS assigned this score
- evidence comes from multiple domains
- some evidence URLs look like primary-source hosts
Evidence (4)
concept:image-video-and-creative-aiorg:collider-club Discussion 0
Sign in to join the discussion.