# Multimodal — MDRSS semantic catalog

> Semantic domain: multimodal
> Vision, speech, audio, games, media generation, robotics, embodied AI, and creative production.
> Filters: category=image-video-and-creative-ai
> Aggregate catalog URL: https://mdrss.com/catalog/multimodal?category=image-video-and-creative-ai

## Feeds

- [Health, Medicine & Psychology](https://mdrss.com/s/health-medicine-and-psychology) — Open clinical standards, medical data, public health, imaging, and mental health.
- [Lifestyle, Culture & Hobbies](https://mdrss.com/s/lifestyle-culture-and-hobbies) — Open cultural collections, games, cooking, creative work, and practical hobbies.

## Aggregate endpoints

- RSS: https://mdrss.com/catalog/multimodal/rss.xml?category=image-video-and-creative-ai
- JSON: https://mdrss.com/catalog/multimodal/feed.json?category=image-video-and-creative-ai

## Semantic domain cards (7)

### [Code Explanation and Analysis](https://mdrss.com/multimodal/image-video-and-creative-ai/901235/901235.md)

You are a code education expert specializing in explaining complex code through clear narratives, visual diagrams, and step-by-step breakdowns. Transform difficult concepts into understandable explanations for developers at all levels. Use it to give an agent explicit responsibilities, steps and constraints.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T13:48:23.601Z · Version: 1

### [DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data](https://mdrss.com/multimodal/image-video-and-creative-ai/901115/901115.md)

A curated reference on image, video & creative ai centered on DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T13:48:05.624Z · Version: 1

### [Awesome Video Diffusion](https://mdrss.com/multimodal/image-video-and-creative-ai/901112/901112.md)

A curated list of recent diffusion models for video generation, editing, restoration, understanding, nerf, etc. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T13:48:05.624Z · Version: 1

### [Highlights](https://mdrss.com/multimodal/image-video-and-creative-ai/901027/901027.md)

14B Real-Time Long Video Generation Model can be Cheaper, Faster but Keep Stronger than 1.3B ones ⭐. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T13:47:57.776Z · Version: 1

### [DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors](https://mdrss.com/multimodal/image-video-and-creative-ai/901013/901013.md)

Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang, Ying Shan, Tien-Tsin Wong From CUHK and Tencent AI Lab. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T13:47:57.776Z · Version: 1

### [A Survey on Video Diffusion Models](https://mdrss.com/multimodal/image-video-and-creative-ai/901006/901006.md)

Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, Yu-Gang Jiang. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T13:47:57.776Z · Version: 1

### [VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models](https://mdrss.com/multimodal/image-video-and-creative-ai/901003/901003.md)

VideoCrafter is an open-source video generation and editing toolbox for crafting video content. It currently includes the Text2Video and Image2Video models:. Use it when a task needs concrete terminology, constraints or implementation detail.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T13:47:30.961Z · Version: 1
