{"version":"mdrss-catalog-feed/1","domain":{"slug":"multimodal","label":"Multimodal AI","description":"Vision, speech and generative media systems."},"feeds":[{"slug":"multimodal","url":"https://mdrss.com/s/multimodal"}],"urls":{"html":"https://mdrss.com/catalog/multimodal","rss":"https://mdrss.com/catalog/multimodal/rss.xml","json":"https://mdrss.com/catalog/multimodal/feed.json","markdown":"https://mdrss.com/catalog/multimodal/index.md"},"updated_at":"2026-08-04T13:54:51.641Z","items":[{"id":901314,"title":"unity-ecs-patterns — detailed patterns and worked examples","annotation":"unity-ecs-patterns — detailed patterns and worked examples captures reusable agent playbook guidance for games & creative coding. Use it to give an agent explicit responsibilities, steps and constraints.","feed":"multimodal","semantic_path":"multimodal:games-and-creative-coding","content_type":"guide","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/games-and-creative-coding/901314","permalink_url":"https://mdrss.com/m/901314","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/games-and-creative-coding/901314/901314.md","file_url":"https://mdrss.com/api/v1/cards/901314/file","raw_url":"https://mdrss.com/multimodal/games-and-creative-coding/901314/raw","embed_url":"https://mdrss.com/multimodal/games-and-creative-coding/901314/embed","edit_url":"https://mdrss.com/cards/901314/edit","legacy_url":"https://mdrss.com/s/multimodal/unity-ecs-patterns-detailed-patterns-and-worked-examples-collider-5a77d3b85074"},{"id":901313,"title":"godot-gdscript-patterns — detailed patterns and worked examples","annotation":"For advanced Godot patterns, performance tips, and best practices, see references/advanced-patterns.md:. Use it to give an agent explicit responsibilities, steps and constraints.","feed":"multimodal","semantic_path":"multimodal:games-and-creative-coding","content_type":"guide","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/games-and-creative-coding/901313","permalink_url":"https://mdrss.com/m/901313","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/games-and-creative-coding/901313/901313.md","file_url":"https://mdrss.com/api/v1/cards/901313/file","raw_url":"https://mdrss.com/multimodal/games-and-creative-coding/901313/raw","embed_url":"https://mdrss.com/multimodal/games-and-creative-coding/901313/embed","edit_url":"https://mdrss.com/cards/901313/edit","legacy_url":"https://mdrss.com/s/multimodal/godot-gdscript-patterns-detailed-patterns-and-worked-examples-collider-362a188e59df"},{"id":901235,"title":"Code Explanation and Analysis","annotation":"You are a code education expert specializing in explaining complex code through clear narratives, visual diagrams, and step-by-step breakdowns. Transform difficult concepts into understandable explanations for developers at all levels. Use it to give an agent explicit responsibilities, steps and constraints.","feed":"multimodal","semantic_path":"multimodal:image-video-and-creative-ai","content_type":"guide","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901235","permalink_url":"https://mdrss.com/m/901235","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901235/901235.md","file_url":"https://mdrss.com/api/v1/cards/901235/file","raw_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901235/raw","embed_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901235/embed","edit_url":"https://mdrss.com/cards/901235/edit","legacy_url":"https://mdrss.com/s/multimodal/code-explanation-and-analysis-collider-51c6014c751f"},{"id":901115,"title":"DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data","annotation":"A curated reference on image, video & creative ai centered on DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data. Use it to navigate the topic and choose relevant methods, papers or tools.","feed":"multimodal","semantic_path":"multimodal:image-video-and-creative-ai","content_type":"reference","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901115","permalink_url":"https://mdrss.com/m/901115","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901115/901115.md","file_url":"https://mdrss.com/api/v1/cards/901115/file","raw_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901115/raw","embed_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901115/embed","edit_url":"https://mdrss.com/cards/901115/edit","legacy_url":"https://mdrss.com/s/multimodal/dreamsim-learning-new-dimensions-of-human-visual-similarity-using-synt-collider-2644595ceec1"},{"id":901112,"title":"Awesome Video Diffusion","annotation":"A curated list of recent diffusion models for video generation, editing, restoration, understanding, nerf, etc. Use it to navigate the topic and choose relevant methods, papers or tools.","feed":"multimodal","semantic_path":"multimodal:image-video-and-creative-ai","content_type":"reference","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901112","permalink_url":"https://mdrss.com/m/901112","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901112/901112.md","file_url":"https://mdrss.com/api/v1/cards/901112/file","raw_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901112/raw","embed_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901112/embed","edit_url":"https://mdrss.com/cards/901112/edit","legacy_url":"https://mdrss.com/s/multimodal/awesome-video-diffusion-collider-5bd7542cbfc2"},{"id":901082,"title":"Awesome Audio-Visual","annotation":"A curated list of papers and datsets for various audio-visual tasks, inspired by awesome-computer-vision. Use it to navigate the topic and choose relevant methods, papers or tools.","feed":"multimodal","semantic_path":"multimodal:audio-and-music","content_type":"reference","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/audio-and-music/901082","permalink_url":"https://mdrss.com/m/901082","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/audio-and-music/901082/901082.md","file_url":"https://mdrss.com/api/v1/cards/901082/file","raw_url":"https://mdrss.com/multimodal/audio-and-music/901082/raw","embed_url":"https://mdrss.com/multimodal/audio-and-music/901082/embed","edit_url":"https://mdrss.com/cards/901082/edit","legacy_url":"https://mdrss.com/s/multimodal/awesome-audio-visual-collider-5368f2904684"},{"id":901068,"title":"Faust - Programming Language for Audio Applications and Plugins","annotation":"Faust (Functional Audio Stream) is a functional programming language specifically designed for real-time signal processing and synthesis. A distinctive characteristic of Faust is that it is fully compiled. Use it to build a structured path from fundamentals to hands-on practice.","feed":"multimodal","semantic_path":"multimodal:audio-and-music","content_type":"guide","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/audio-and-music/901068","permalink_url":"https://mdrss.com/m/901068","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/audio-and-music/901068/901068.md","file_url":"https://mdrss.com/api/v1/cards/901068/file","raw_url":"https://mdrss.com/multimodal/audio-and-music/901068/raw","embed_url":"https://mdrss.com/multimodal/audio-and-music/901068/embed","edit_url":"https://mdrss.com/cards/901068/edit","legacy_url":"https://mdrss.com/s/multimodal/faust-programming-language-for-audio-applications-and-plugins-collider-d2cf359f6958"},{"id":901034,"title":"Tone.js","annotation":"Tone.js is a Web Audio framework for creating interactive music in the browser. The architecture of Tone.js aims to be familiar to both musicians and audio programmers creating web-based audio applications. On the high-level, Tone offers common DAW (digital audio workstation) fea. Use it to ground design choices in named patterns, trade-offs and examples.","feed":"multimodal","semantic_path":"multimodal:audio-and-music","content_type":"reference","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/audio-and-music/901034","permalink_url":"https://mdrss.com/m/901034","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/audio-and-music/901034/901034.md","file_url":"https://mdrss.com/api/v1/cards/901034/file","raw_url":"https://mdrss.com/multimodal/audio-and-music/901034/raw","embed_url":"https://mdrss.com/multimodal/audio-and-music/901034/embed","edit_url":"https://mdrss.com/cards/901034/edit","legacy_url":"https://mdrss.com/s/multimodal/tone-js-collider-c052c03e3e23"},{"id":901030,"title":"FunDSP","annotation":"FunDSP is an audio DSP (digital signal processing) library for audio processing and synthesis. Use it when a task needs concrete terminology, constraints or implementation detail.","feed":"multimodal","semantic_path":"multimodal:audio-and-music","content_type":"reference","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/audio-and-music/901030","permalink_url":"https://mdrss.com/m/901030","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/audio-and-music/901030/901030.md","file_url":"https://mdrss.com/api/v1/cards/901030/file","raw_url":"https://mdrss.com/multimodal/audio-and-music/901030/raw","embed_url":"https://mdrss.com/multimodal/audio-and-music/901030/embed","edit_url":"https://mdrss.com/cards/901030/edit","legacy_url":"https://mdrss.com/s/multimodal/fundsp-collider-13583ee2dd8e"},{"id":901027,"title":"Highlights","annotation":"14B Real-Time Long Video Generation Model can be Cheaper, Faster but Keep Stronger than 1.3B ones ⭐. Use it to navigate the topic and choose relevant methods, papers or tools.","feed":"multimodal","semantic_path":"multimodal:image-video-and-creative-ai","content_type":"reference","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901027","permalink_url":"https://mdrss.com/m/901027","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901027/901027.md","file_url":"https://mdrss.com/api/v1/cards/901027/file","raw_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901027/raw","embed_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901027/embed","edit_url":"https://mdrss.com/cards/901027/edit","legacy_url":"https://mdrss.com/s/multimodal/highlights-collider-9a285aa6badd"},{"id":901013,"title":"DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors","annotation":"Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang, Ying Shan, Tien-Tsin Wong From CUHK and Tencent AI Lab. Use it to navigate the topic and choose relevant methods, papers or tools.","feed":"multimodal","semantic_path":"multimodal:image-video-and-creative-ai","content_type":"reference","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901013","permalink_url":"https://mdrss.com/m/901013","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901013/901013.md","file_url":"https://mdrss.com/api/v1/cards/901013/file","raw_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901013/raw","embed_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901013/embed","edit_url":"https://mdrss.com/cards/901013/edit","legacy_url":"https://mdrss.com/s/multimodal/dynamicrafter-animating-open-domain-images-with-video-diffusion-priors-collider-a26cfc12f1dd"},{"id":901006,"title":"A Survey on Video Diffusion Models","annotation":"Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, Yu-Gang Jiang. Use it to navigate the topic and choose relevant methods, papers or tools.","feed":"multimodal","semantic_path":"multimodal:image-video-and-creative-ai","content_type":"reference","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901006","permalink_url":"https://mdrss.com/m/901006","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901006/901006.md","file_url":"https://mdrss.com/api/v1/cards/901006/file","raw_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901006/raw","embed_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901006/embed","edit_url":"https://mdrss.com/cards/901006/edit","legacy_url":"https://mdrss.com/s/multimodal/a-survey-on-video-diffusion-models-collider-bef0ac2f349c"},{"id":901003,"title":"VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models","annotation":"VideoCrafter is an open-source video generation and editing toolbox for crafting video content. It currently includes the Text2Video and Image2Video models:. Use it when a task needs concrete terminology, constraints or implementation detail.","feed":"multimodal","semantic_path":"multimodal:image-video-and-creative-ai","content_type":"reference","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901003","permalink_url":"https://mdrss.com/m/901003","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901003/901003.md","file_url":"https://mdrss.com/api/v1/cards/901003/file","raw_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901003/raw","embed_url":"https://mdrss.com/multimodal/image-video-and-creative-ai/901003/embed","edit_url":"https://mdrss.com/cards/901003/edit","legacy_url":"https://mdrss.com/s/multimodal/videocrafter2-overcoming-data-limitations-for-high-quality-video-diffu-collider-5d4e724ba664"},{"id":901002,"title":"Music Modeling and Music Generation with Deep Learning","annotation":"A curated reference on audio & music centered on Music Modeling and Music Generation with Deep Learning. Use it to navigate the topic and choose relevant methods, papers or tools.","feed":"multimodal","semantic_path":"multimodal:audio-and-music","content_type":"reference","version":1,"snapshot_at":"2026-08-04T16:17:00.000Z","publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","card_url":"https://mdrss.com/multimodal/audio-and-music/901002","permalink_url":"https://mdrss.com/m/901002","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/audio-and-music/901002/901002.md","file_url":"https://mdrss.com/api/v1/cards/901002/file","raw_url":"https://mdrss.com/multimodal/audio-and-music/901002/raw","embed_url":"https://mdrss.com/multimodal/audio-and-music/901002/embed","edit_url":"https://mdrss.com/cards/901002/edit","legacy_url":"https://mdrss.com/s/multimodal/music-modeling-and-music-generation-with-deep-learning-collider-053f9a5bb164"},{"id":2516,"title":"Coqui.ai News","annotation":"🐸TTS is a library for advanced Text-to-Speech generation. 🛠️ Tools for training new models and fine-tuning existing models in any language.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"reference","version":1,"snapshot_at":"2026-08-04T12:18:51.210Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/2516","permalink_url":"https://mdrss.com/m/2516","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/2516/2516.md","file_url":"https://mdrss.com/api/v1/cards/2516/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/2516/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/2516/embed","edit_url":"https://mdrss.com/cards/2516/edit","legacy_url":"https://mdrss.com/s/multimodal/coqui-ai-tts-coqui-ai-tts-readme"},{"id":2474,"title":"Whisper","annotation":"It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification. A Transformer sequence-to-sequence model is trained on various speech processing tasks, including multilingual speech recognition, speech translation, spoken language identification, and voice activity detection.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:51.210Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/2474","permalink_url":"https://mdrss.com/m/2474","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/2474/2474.md","file_url":"https://mdrss.com/api/v1/cards/2474/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/2474/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/2474/embed","edit_url":"https://mdrss.com/cards/2474/edit","legacy_url":"https://mdrss.com/s/multimodal/openai-whisper-openai-whisper-readme"},{"id":2057,"title":"Cosmos Tokenizer: A suite of image and video neural tokenizers","annotation":"As of February 10th, 2025, this repository is read-only. Please visit github.com/NVIDIA/Cosmos for the latest updates and support on Cosmos Tokenizer.","feed":"multimodal","semantic_path":"multimodal:vision-and-media","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:51.210Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/vision-and-media/2057","permalink_url":"https://mdrss.com/m/2057","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/vision-and-media/2057/2057.md","file_url":"https://mdrss.com/api/v1/cards/2057/file","raw_url":"https://mdrss.com/multimodal/vision-and-media/2057/raw","embed_url":"https://mdrss.com/multimodal/vision-and-media/2057/embed","edit_url":"https://mdrss.com/cards/2057/edit","legacy_url":"https://mdrss.com/s/multimodal/nvidia-cosmos-tokenizer-nvidia-cosmos-tokenizer-readme"},{"id":1709,"title":"Faster Whisper transcription with CTranslate2","annotation":"faster-whisper is a reimplementation of OpenAI's Whisper model using CTranslate2, which is a fast inference engine for Transformer models. This implementation is up to 4 times faster than openai/whisper for the same accuracy while using less memory.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:49.845Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/1709","permalink_url":"https://mdrss.com/m/1709","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/1709/1709.md","file_url":"https://mdrss.com/api/v1/cards/1709/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/1709/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/1709/embed","edit_url":"https://mdrss.com/cards/1709/edit","legacy_url":"https://mdrss.com/s/multimodal/systran-faster-whisper-systran-faster-whisper-readme"},{"id":1706,"title":"Model properties","annotation":"USP : Unvoice and Silence with Pitch when infer 1. Install project dependencies Note: whisper is already built-in, do not install it again otherwise it will cuase conflict and error 3.","feed":"multimodal","semantic_path":"multimodal:vision-and-media","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:51.210Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/vision-and-media/1706","permalink_url":"https://mdrss.com/m/1706","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/vision-and-media/1706/1706.md","file_url":"https://mdrss.com/api/v1/cards/1706/file","raw_url":"https://mdrss.com/multimodal/vision-and-media/1706/raw","embed_url":"https://mdrss.com/multimodal/vision-and-media/1706/embed","edit_url":"https://mdrss.com/cards/1706/edit","legacy_url":"https://mdrss.com/s/multimodal/playvoice-whisper-vits-svc-playvoice-whisper-vits-svc-readme"},{"id":1611,"title":"MOVA: Towards Scalable and Synchronized Video–Audio Generation","annotation":"We introduce MOVA (MOSS Video and Audio), a foundation model designed to break the \"silent era\" of open-source video generation. Unlike cascaded pipelines that generate sound as an afterthought, MOVA synthesizes video and audio simultaneously for perfect alignment.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:49.845Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/1611","permalink_url":"https://mdrss.com/m/1611","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/1611/1611.md","file_url":"https://mdrss.com/api/v1/cards/1611/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/1611/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/1611/embed","edit_url":"https://mdrss.com/cards/1611/edit","legacy_url":"https://mdrss.com/s/multimodal/openmoss-mova-openmoss-mova-readme"},{"id":1370,"title":"text-extract-api","annotation":"Convert any image, PDF or Office document to Markdown text or JSON structured document with super-high accuracy, including tabular data, numbers or math formulas. The API is built with FastAPI and uses Celery for asynchronous task processing.","feed":"multimodal","semantic_path":"multimodal:vision-and-media","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:52.437Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/vision-and-media/1370","permalink_url":"https://mdrss.com/m/1370","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/vision-and-media/1370/1370.md","file_url":"https://mdrss.com/api/v1/cards/1370/file","raw_url":"https://mdrss.com/multimodal/vision-and-media/1370/raw","embed_url":"https://mdrss.com/multimodal/vision-and-media/1370/embed","edit_url":"https://mdrss.com/cards/1370/edit","legacy_url":"https://mdrss.com/s/multimodal/catchthetornado-text-extract-api-catchthetornado-text-extract-api-readme"},{"id":1369,"title":"PICK-PyTorch","annotation":"\\\\\\\\\\ Updated on Feb 6th, 2021: Train Ticket dataset is now available for academic research. You can download from Google Drive or OneDrive.","feed":"multimodal","semantic_path":"multimodal:vision-and-media","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:52.437Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/vision-and-media/1369","permalink_url":"https://mdrss.com/m/1369","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/vision-and-media/1369/1369.md","file_url":"https://mdrss.com/api/v1/cards/1369/file","raw_url":"https://mdrss.com/multimodal/vision-and-media/1369/raw","embed_url":"https://mdrss.com/multimodal/vision-and-media/1369/embed","edit_url":"https://mdrss.com/cards/1369/edit","legacy_url":"https://mdrss.com/s/multimodal/wenwenyu-pick-pytorch-wenwenyu-pick-pytorch-readme"},{"id":1345,"title":"MuseV English 中文","annotation":"MuseV was a milestone achieved around July 2023. Amazed by the progress of Sora, we decided to opensource MuseV, hopefully it will benefit the community.","feed":"multimodal","semantic_path":"multimodal:vision-and-media","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:51.210Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/vision-and-media/1345","permalink_url":"https://mdrss.com/m/1345","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/vision-and-media/1345/1345.md","file_url":"https://mdrss.com/api/v1/cards/1345/file","raw_url":"https://mdrss.com/multimodal/vision-and-media/1345/raw","embed_url":"https://mdrss.com/multimodal/vision-and-media/1345/embed","edit_url":"https://mdrss.com/cards/1345/edit","legacy_url":"https://mdrss.com/s/multimodal/tmelyralab-musev-tmelyralab-musev-readme"},{"id":1304,"title":"EasyAnimate | An End-to-End Solution for High-Resolution and Long Video Generation","annotation":"😊 EasyAnimate is an end-to-end solution for generating high-resolution and long videos. We can train transformer based diffusion generators, train VAEs for processing long videos, and preprocess metadata.","feed":"multimodal","semantic_path":"multimodal:vision-and-media","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:52.437Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/vision-and-media/1304","permalink_url":"https://mdrss.com/m/1304","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/vision-and-media/1304/1304.md","file_url":"https://mdrss.com/api/v1/cards/1304/file","raw_url":"https://mdrss.com/multimodal/vision-and-media/1304/raw","embed_url":"https://mdrss.com/multimodal/vision-and-media/1304/embed","edit_url":"https://mdrss.com/cards/1304/edit","legacy_url":"https://mdrss.com/s/multimodal/aigc-apps-easyanimate-aigc-apps-easyanimate-readme"},{"id":1255,"title":"TurboDiffusion","annotation":"This repository provides the official implementation of TurboDiffusion, a video generation acceleration framework that can speed up end-to-end diffusion generation by $100 \\sim 200\\times$ on a single RTX 5090, while maintaining video quality. TurboDiffusion primarily uses SageAttention, SLA (Sparse-Linear Attention) for attention acceleration, and rCM for timestep distillation.","feed":"multimodal","semantic_path":"multimodal:vision-and-media","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:51.210Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/vision-and-media/1255","permalink_url":"https://mdrss.com/m/1255","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/vision-and-media/1255/1255.md","file_url":"https://mdrss.com/api/v1/cards/1255/file","raw_url":"https://mdrss.com/multimodal/vision-and-media/1255/raw","embed_url":"https://mdrss.com/multimodal/vision-and-media/1255/embed","edit_url":"https://mdrss.com/cards/1255/edit","legacy_url":"https://mdrss.com/s/multimodal/thu-ml-turbodiffusion-thu-ml-turbodiffusion-readme"},{"id":1061,"title":"ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis","annotation":"ViewCrafter can generate high-fidelity novel views from a single or sparse reference image , while also supporting highly precise pose control. Below shows some examples: Reference image Camera trajecotry Generated novel view video Reference image 1 Reference image 2 Generated novel view video |Model|Resolution|Frames|GPU Mem.","feed":"multimodal","semantic_path":"multimodal:vision-and-media","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:52.437Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/vision-and-media/1061","permalink_url":"https://mdrss.com/m/1061","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/vision-and-media/1061/1061.md","file_url":"https://mdrss.com/api/v1/cards/1061/file","raw_url":"https://mdrss.com/multimodal/vision-and-media/1061/raw","embed_url":"https://mdrss.com/multimodal/vision-and-media/1061/embed","edit_url":"https://mdrss.com/cards/1061/edit","legacy_url":"https://mdrss.com/s/multimodal/drexubery-viewcrafter-drexubery-viewcrafter-readme"},{"id":1039,"title":"Generating the documentation","annotation":"To generate the documentation, you first have to build it. You don't have to commit the built documentation.","feed":"multimodal","semantic_path":"multimodal:vision-and-media","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:52.437Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/vision-and-media/1039","permalink_url":"https://mdrss.com/m/1039","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/vision-and-media/1039/1039.md","file_url":"https://mdrss.com/api/v1/cards/1039/file","raw_url":"https://mdrss.com/multimodal/vision-and-media/1039/raw","embed_url":"https://mdrss.com/multimodal/vision-and-media/1039/embed","edit_url":"https://mdrss.com/cards/1039/edit","legacy_url":"https://mdrss.com/s/multimodal/huggingface-diffusers-huggingface-diffusers-docs-readme"},{"id":986,"title":"An AI-powered web application for speech recognition, translation, and dubbing","annotation":"Voice-Pro The best AI speech recognition, translation, and multilingual dubbing solution 🚀 한국어 ∙ English ∙ 中文简体 ∙ 中文繁體 ∙ 日本語 ∙ Deutsch ∙ Español ∙ Português Voice-Pro is a state-of-the-art web app that transforms multimedia content creation. It integrates YouTube video downloading, voice separation, speech recognition, translation, and text-to-speech into a single, powerful tool for creators, researchers, and multilingual professionals.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"reference","version":1,"snapshot_at":"2026-08-04T12:18:49.845Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/986","permalink_url":"https://mdrss.com/m/986","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/986/986.md","file_url":"https://mdrss.com/api/v1/cards/986/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/986/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/986/embed","edit_url":"https://mdrss.com/cards/986/edit","legacy_url":"https://mdrss.com/s/multimodal/abus-aikorea-voice-pro-abus-aikorea-voice-pro-readme"},{"id":853,"title":"What is Voicebox?","annotation":"The full voice I/O stack, running locally on your machine. voicebox.sh • Docs • Download • Features • API • Troubleshooting Click the image above to watch the demo video on voicebox.sh Voicebox is a local-first AI voice studio — a free and open-source alternative to ElevenLabs and WisprFlow in one app.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:49.845Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/853","permalink_url":"https://mdrss.com/m/853","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/853/853.md","file_url":"https://mdrss.com/api/v1/cards/853/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/853/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/853/embed","edit_url":"https://mdrss.com/cards/853/edit","legacy_url":"https://mdrss.com/s/multimodal/jamiepine-voicebox-jamiepine-voicebox-readme"},{"id":794,"title":"Moonshine Voice","annotation":"Voice Interfaces for Everyone Moonshine Voice is an open source AI toolkit for developers building real-time voice agents and applications. Join our community on Discord to get live support.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:49.845Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/794","permalink_url":"https://mdrss.com/m/794","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/794/794.md","file_url":"https://mdrss.com/api/v1/cards/794/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/794/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/794/embed","edit_url":"https://mdrss.com/cards/794/edit","legacy_url":"https://mdrss.com/s/multimodal/moonshine-ai-moonshine-moonshine-ai-moonshine-readme"}]}