# Multimodal AI — MDRSS catalog

> Vision, speech and generative media systems.
> Aggregate catalog URL: https://mdrss.com/catalog/multimodal

## Feeds

- [Multimodal AI](https://mdrss.com/s/multimodal) — Vision, speech and generative media systems.

## Aggregate endpoints

- RSS: https://mdrss.com/catalog/multimodal/rss.xml
- JSON: https://mdrss.com/catalog/multimodal/feed.json

## Curated snapshots (30)

### [unity-ecs-patterns — detailed patterns and worked examples](https://mdrss.com/multimodal/games-and-creative-coding/901314/901314.md)

unity-ecs-patterns — detailed patterns and worked examples captures reusable agent playbook guidance for games & creative coding. Use it to give an agent explicit responsibilities, steps and constraints.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [godot-gdscript-patterns — detailed patterns and worked examples](https://mdrss.com/multimodal/games-and-creative-coding/901313/901313.md)

For advanced Godot patterns, performance tips, and best practices, see references/advanced-patterns.md:. Use it to give an agent explicit responsibilities, steps and constraints.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [Code Explanation and Analysis](https://mdrss.com/multimodal/image-video-and-creative-ai/901235/901235.md)

You are a code education expert specializing in explaining complex code through clear narratives, visual diagrams, and step-by-step breakdowns. Transform difficult concepts into understandable explanations for developers at all levels. Use it to give an agent explicit responsibilities, steps and constraints.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data](https://mdrss.com/multimodal/image-video-and-creative-ai/901115/901115.md)

A curated reference on image, video & creative ai centered on DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [Awesome Video Diffusion](https://mdrss.com/multimodal/image-video-and-creative-ai/901112/901112.md)

A curated list of recent diffusion models for video generation, editing, restoration, understanding, nerf, etc. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [Awesome Audio-Visual](https://mdrss.com/multimodal/audio-and-music/901082/901082.md)

A curated list of papers and datsets for various audio-visual tasks, inspired by awesome-computer-vision. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [Faust - Programming Language for Audio Applications and Plugins](https://mdrss.com/multimodal/audio-and-music/901068/901068.md)

Faust (Functional Audio Stream) is a functional programming language specifically designed for real-time signal processing and synthesis. A distinctive characteristic of Faust is that it is fully compiled. Use it to build a structured path from fundamentals to hands-on practice.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [Tone.js](https://mdrss.com/multimodal/audio-and-music/901034/901034.md)

Tone.js is a Web Audio framework for creating interactive music in the browser. The architecture of Tone.js aims to be familiar to both musicians and audio programmers creating web-based audio applications. On the high-level, Tone offers common DAW (digital audio workstation) fea. Use it to ground design choices in named patterns, trade-offs and examples.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [FunDSP](https://mdrss.com/multimodal/audio-and-music/901030/901030.md)

FunDSP is an audio DSP (digital signal processing) library for audio processing and synthesis. Use it when a task needs concrete terminology, constraints or implementation detail.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [Highlights](https://mdrss.com/multimodal/image-video-and-creative-ai/901027/901027.md)

14B Real-Time Long Video Generation Model can be Cheaper, Faster but Keep Stronger than 1.3B ones ⭐. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors](https://mdrss.com/multimodal/image-video-and-creative-ai/901013/901013.md)

Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang, Ying Shan, Tien-Tsin Wong From CUHK and Tencent AI Lab. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [A Survey on Video Diffusion Models](https://mdrss.com/multimodal/image-video-and-creative-ai/901006/901006.md)

Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, Yu-Gang Jiang. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models](https://mdrss.com/multimodal/image-video-and-creative-ai/901003/901003.md)

VideoCrafter is an open-source video generation and editing toolbox for crafting video content. It currently includes the Text2Video and Image2Video models:. Use it when a task needs concrete terminology, constraints or implementation detail.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [Music Modeling and Music Generation with Deep Learning](https://mdrss.com/multimodal/audio-and-music/901002/901002.md)

A curated reference on audio & music centered on Music Modeling and Music Generation with Deep Learning. Use it to navigate the topic and choose relevant methods, papers or tools.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T16:17:00.000Z · Version: 1

### [Coqui.ai News](https://mdrss.com/multimodal/speech-and-audio/2516/2516.md)

🐸TTS is a library for advanced Text-to-Speech generation. 🛠️ Tools for training new models and fine-tuning existing models in any language.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [Whisper](https://mdrss.com/multimodal/speech-and-audio/2474/2474.md)

It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification. A Transformer sequence-to-sequence model is trained on various speech processing tasks, including multilingual speech recognition, speech translation, spoken language identification, and voice activity detection.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [Cosmos Tokenizer: A suite of image and video neural tokenizers](https://mdrss.com/multimodal/vision-and-media/2057/2057.md)

As of February 10th, 2025, this repository is read-only. Please visit github.com/NVIDIA/Cosmos for the latest updates and support on Cosmos Tokenizer.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [Faster Whisper transcription with CTranslate2](https://mdrss.com/multimodal/speech-and-audio/1709/1709.md)

faster-whisper is a reimplementation of OpenAI's Whisper model using CTranslate2, which is a fast inference engine for Transformer models. This implementation is up to 4 times faster than openai/whisper for the same accuracy while using less memory.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:49.845Z · Version: 1

### [Model properties](https://mdrss.com/multimodal/vision-and-media/1706/1706.md)

USP : Unvoice and Silence with Pitch when infer 1. Install project dependencies Note: whisper is already built-in, do not install it again otherwise it will cuase conflict and error 3.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [MOVA: Towards Scalable and Synchronized Video–Audio Generation](https://mdrss.com/multimodal/speech-and-audio/1611/1611.md)

We introduce MOVA (MOSS Video and Audio), a foundation model designed to break the "silent era" of open-source video generation. Unlike cascaded pipelines that generate sound as an afterthought, MOVA synthesizes video and audio simultaneously for perfect alignment.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:49.845Z · Version: 1

### [text-extract-api](https://mdrss.com/multimodal/vision-and-media/1370/1370.md)

Convert any image, PDF or Office document to Markdown text or JSON structured document with super-high accuracy, including tabular data, numbers or math formulas. The API is built with FastAPI and uses Celery for asynchronous task processing.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:52.437Z · Version: 1

### [PICK-PyTorch](https://mdrss.com/multimodal/vision-and-media/1369/1369.md)

\\\\\ Updated on Feb 6th, 2021: Train Ticket dataset is now available for academic research. You can download from Google Drive or OneDrive.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:52.437Z · Version: 1

### [MuseV English 中文](https://mdrss.com/multimodal/vision-and-media/1345/1345.md)

MuseV was a milestone achieved around July 2023. Amazed by the progress of Sora, we decided to opensource MuseV, hopefully it will benefit the community.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [EasyAnimate | An End-to-End Solution for High-Resolution and Long Video Generation](https://mdrss.com/multimodal/vision-and-media/1304/1304.md)

😊 EasyAnimate is an end-to-end solution for generating high-resolution and long videos. We can train transformer based diffusion generators, train VAEs for processing long videos, and preprocess metadata.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:52.437Z · Version: 1

### [TurboDiffusion](https://mdrss.com/multimodal/vision-and-media/1255/1255.md)

This repository provides the official implementation of TurboDiffusion, a video generation acceleration framework that can speed up end-to-end diffusion generation by $100 \sim 200\times$ on a single RTX 5090, while maintaining video quality. TurboDiffusion primarily uses SageAttention, SLA (Sparse-Linear Attention) for attention acceleration, and rCM for timestep distillation.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis](https://mdrss.com/multimodal/vision-and-media/1061/1061.md)

ViewCrafter can generate high-fidelity novel views from a single or sparse reference image , while also supporting highly precise pose control. Below shows some examples: Reference image Camera trajecotry Generated novel view video Reference image 1 Reference image 2 Generated novel view video |Model|Resolution|Frames|GPU Mem.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:52.437Z · Version: 1

### [Generating the documentation](https://mdrss.com/multimodal/vision-and-media/1039/1039.md)

To generate the documentation, you first have to build it. You don't have to commit the built documentation.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:52.437Z · Version: 1

### [An AI-powered web application for speech recognition, translation, and dubbing](https://mdrss.com/multimodal/speech-and-audio/986/986.md)

Voice-Pro The best AI speech recognition, translation, and multilingual dubbing solution 🚀 한국어 ∙ English ∙ 中文简体 ∙ 中文繁體 ∙ 日本語 ∙ Deutsch ∙ Español ∙ Português Voice-Pro is a state-of-the-art web app that transforms multimedia content creation. It integrates YouTube video downloading, voice separation, speech recognition, translation, and text-to-speech into a single, powerful tool for creators, researchers, and multilingual professionals.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:49.845Z · Version: 1

### [What is Voicebox?](https://mdrss.com/multimodal/speech-and-audio/853/853.md)

The full voice I/O stack, running locally on your machine. voicebox.sh • Docs • Download • Features • API • Troubleshooting Click the image above to watch the demo video on voicebox.sh Voicebox is a local-first AI voice studio — a free and open-source alternative to ElevenLabs and WisprFlow in one app.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:49.845Z · Version: 1

### [Moonshine Voice](https://mdrss.com/multimodal/speech-and-audio/794/794.md)

Voice Interfaces for Everyone Moonshine Voice is an open source AI toolkit for developers building real-time voice agents and applications. Join our community on Discord to get live support.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:49.845Z · Version: 1
