# #multimodal — MDRSS hashtag feed

> Public MDRSS cards tagged #multimodal.
> Canonical feed: https://mdrss.com/feeds/multimodal

## Cards (23)

### [unity-ecs-patterns — detailed patterns and worked examples](https://mdrss.com/multimodal/games-and-creative-coding/901314/901314.md)

unity-ecs-patterns — detailed patterns and worked examples captures reusable agent playbook guidance for games & creative coding. Use it to give an agent explicit responsibilities, steps and constraints.

Classification: multimodal/games-and-creative-coding · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [godot-gdscript-patterns — detailed patterns and worked examples](https://mdrss.com/multimodal/games-and-creative-coding/901313/901313.md)

For advanced Godot patterns, performance tips, and best practices, see references/advanced-patterns.md:. Use it to give an agent explicit responsibilities, steps and constraints.

Classification: multimodal/games-and-creative-coding · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Code Explanation and Analysis](https://mdrss.com/multimodal/image-video-and-creative-ai/901235/901235.md)

You are a code education expert specializing in explaining complex code through clear narratives, visual diagrams, and step-by-step breakdowns. Transform difficult concepts into understandable explanations for developers at all levels. Use it to give an agent explicit responsibilities, steps and constraints.

Classification: multimodal/image-video-and-creative-ai · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data](https://mdrss.com/multimodal/image-video-and-creative-ai/901115/901115.md)

A curated reference on image, video & creative ai centered on DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: multimodal/image-video-and-creative-ai · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Awesome Video Diffusion](https://mdrss.com/multimodal/image-video-and-creative-ai/901112/901112.md)

A curated list of recent diffusion models for video generation, editing, restoration, understanding, nerf, etc. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: multimodal/image-video-and-creative-ai · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Awesome Audio-Visual](https://mdrss.com/multimodal/audio-and-music/901082/901082.md)

A curated list of papers and datsets for various audio-visual tasks, inspired by awesome-computer-vision. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: multimodal/audio-and-music · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Faust - Programming Language for Audio Applications and Plugins](https://mdrss.com/multimodal/audio-and-music/901068/901068.md)

Faust (Functional Audio Stream) is a functional programming language specifically designed for real-time signal processing and synthesis. A distinctive characteristic of Faust is that it is fully compiled. Use it to build a structured path from fundamentals to hands-on practice.

Classification: multimodal/audio-and-music · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Tone.js](https://mdrss.com/multimodal/audio-and-music/901034/901034.md)

Tone.js is a Web Audio framework for creating interactive music in the browser. The architecture of Tone.js aims to be familiar to both musicians and audio programmers creating web-based audio applications. On the high-level, Tone offers common DAW (digital audio workstation) fea. Use it to ground design choices in named patterns, trade-offs and examples.

Classification: multimodal/audio-and-music · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [FunDSP](https://mdrss.com/multimodal/audio-and-music/901030/901030.md)

FunDSP is an audio DSP (digital signal processing) library for audio processing and synthesis. Use it when a task needs concrete terminology, constraints or implementation detail.

Classification: multimodal/audio-and-music · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Highlights](https://mdrss.com/multimodal/image-video-and-creative-ai/901027/901027.md)

14B Real-Time Long Video Generation Model can be Cheaper, Faster but Keep Stronger than 1.3B ones ⭐. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: multimodal/image-video-and-creative-ai · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors](https://mdrss.com/multimodal/image-video-and-creative-ai/901013/901013.md)

Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang, Ying Shan, Tien-Tsin Wong From CUHK and Tencent AI Lab. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: multimodal/image-video-and-creative-ai · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [A Survey on Video Diffusion Models](https://mdrss.com/multimodal/image-video-and-creative-ai/901006/901006.md)

Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, Yu-Gang Jiang. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: multimodal/image-video-and-creative-ai · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models](https://mdrss.com/multimodal/image-video-and-creative-ai/901003/901003.md)

VideoCrafter is an open-source video generation and editing toolbox for crafting video content. It currently includes the Text2Video and Image2Video models:. Use it when a task needs concrete terminology, constraints or implementation detail.

Classification: multimodal/image-video-and-creative-ai · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Music Modeling and Music Generation with Deep Learning](https://mdrss.com/multimodal/audio-and-music/901002/901002.md)

A curated reference on audio & music centered on Music Modeling and Music Generation with Deep Learning. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: multimodal/audio-and-music · Feed: multimodal · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Whisper](https://mdrss.com/multimodal/speech-and-audio/2474/2474.md)

It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification. A Transformer sequence-to-sequence model is trained on various speech processing tasks, including multilingual speech recognition, speech translation, spoken language identification, and voice activity detection.

Classification: multimodal/speech-and-audio · Feed: multimodal · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [Cosmos Tokenizer: A suite of image and video neural tokenizers](https://mdrss.com/multimodal/vision-and-media/2057/2057.md)

As of February 10th, 2025, this repository is read-only. Please visit github.com/NVIDIA/Cosmos for the latest updates and support on Cosmos Tokenizer.

Classification: multimodal/vision-and-media · Feed: multimodal · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [MOVA: Towards Scalable and Synchronized Video–Audio Generation](https://mdrss.com/multimodal/speech-and-audio/1611/1611.md)

We introduce MOVA (MOSS Video and Audio), a foundation model designed to break the "silent era" of open-source video generation. Unlike cascaded pipelines that generate sound as an afterthought, MOVA synthesizes video and audio simultaneously for perfect alignment.

Classification: multimodal/speech-and-audio · Feed: multimodal · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale](https://mdrss.com/ai-agents/agent-frameworks/1549/1549.md)

Towards Async, Omni-Modal RL at Scale, Just Relax. 📖 English | 📖 中文 Relax (Reinforcement Engine Leveraging Agentic X-modality) is a high-performance reinforcement learning post-training framework open-sourced by the Xiaohongshu AI Infra Team for multimodal large language models.

Classification: ai-agents/agent-frameworks · Feed: ai-agents · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [PICK-PyTorch](https://mdrss.com/multimodal/vision-and-media/1369/1369.md)

\\\\\ Updated on Feb 6th, 2021: Train Ticket dataset is now available for academic research. You can download from Google Drive or OneDrive.

Classification: multimodal/vision-and-media · Feed: multimodal · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [MuseV English 中文](https://mdrss.com/multimodal/vision-and-media/1345/1345.md)

MuseV was a milestone achieved around July 2023. Amazed by the progress of Sora, we decided to opensource MuseV, hopefully it will benefit the community.

Classification: multimodal/vision-and-media · Feed: multimodal · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [EasyAnimate | An End-to-End Solution for High-Resolution and Long Video Generation](https://mdrss.com/multimodal/vision-and-media/1304/1304.md)

😊 EasyAnimate is an end-to-end solution for generating high-resolution and long videos. We can train transformer based diffusion generators, train VAEs for processing long videos, and preprocess metadata.

Classification: multimodal/vision-and-media · Feed: multimodal · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis](https://mdrss.com/multimodal/vision-and-media/1061/1061.md)

ViewCrafter can generate high-fidelity novel views from a single or sparse reference image , while also supporting highly precise pose control. Below shows some examples: Reference image Camera trajecotry Generated novel view video Reference image 1 Reference image 2 Generated novel view video |Model|Resolution|Frames|GPU Mem.

Classification: multimodal/vision-and-media · Feed: multimodal · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [Moonshine Voice](https://mdrss.com/multimodal/speech-and-audio/794/794.md)

Voice Interfaces for Everyone Moonshine Voice is an open source AI toolkit for developers building real-time voice agents and applications. Join our community on Discord to get live support.

Classification: multimodal/speech-and-audio · Feed: multimodal · Updated: 2026-08-04T12:22:38.168Z · Version: 1
