{"version":"mdrss-catalog-feed/2","domain":{"slug":"multimodal","label":"Multimodal","description":"Vision, speech, audio, games, media generation, robotics, embodied AI, and creative production.","categories":["vision-and-media","speech-and-audio","audio-and-music","games-and-creative-coding","image-video-and-creative-ai","multimodal-models-and-fusion","image-generation","video-generation","3d-and-spatial-computing","audio-understanding","speech-synthesis","multimodal-agents","robotics-and-embodied-ai","creative-production","multimodal-evaluation"]},"total":7,"page":1,"page_size":50,"total_pages":1,"has_next":false,"has_previous":false,"filters":{"domain":"multimodal","category":"speech-and-audio"},"feeds":[{"slug":"health-medicine-and-psychology","url":"https://mdrss.com/s/health-medicine-and-psychology"},{"slug":"lifestyle-culture-and-hobbies","url":"https://mdrss.com/s/lifestyle-culture-and-hobbies"}],"urls":{"html":"https://mdrss.com/catalog/multimodal?category=speech-and-audio","rss":"https://mdrss.com/catalog/multimodal/rss.xml?category=speech-and-audio","json":"https://mdrss.com/catalog/multimodal/feed.json?category=speech-and-audio","markdown":"https://mdrss.com/catalog/multimodal/index.md?category=speech-and-audio"},"updated_at":"2026-08-04T12:22:38.168Z","items":[{"id":2516,"title":"Coqui.ai News","annotation":"🐸TTS is a library for advanced Text-to-Speech generation. 🛠️ Tools for training new models and fine-tuning existing models in any language.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"reference","version":1,"snapshot_at":"2026-08-04T12:18:51.210Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/2516","permalink_url":"https://mdrss.com/m/2516","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/2516/2516.md","file_url":"https://mdrss.com/api/v1/cards/2516/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/2516/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/2516/embed","edit_url":"https://mdrss.com/cards/2516/edit","legacy_url":"https://mdrss.com/s/multimodal/coqui-ai-tts-coqui-ai-tts-readme"},{"id":2474,"title":"Whisper","annotation":"It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification. A Transformer sequence-to-sequence model is trained on various speech processing tasks, including multilingual speech recognition, speech translation, spoken language identification, and voice activity detection.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:51.210Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/2474","permalink_url":"https://mdrss.com/m/2474","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/2474/2474.md","file_url":"https://mdrss.com/api/v1/cards/2474/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/2474/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/2474/embed","edit_url":"https://mdrss.com/cards/2474/edit","legacy_url":"https://mdrss.com/s/multimodal/openai-whisper-openai-whisper-readme"},{"id":1709,"title":"Faster Whisper transcription with CTranslate2","annotation":"faster-whisper is a reimplementation of OpenAI's Whisper model using CTranslate2, which is a fast inference engine for Transformer models. This implementation is up to 4 times faster than openai/whisper for the same accuracy while using less memory.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:49.845Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/1709","permalink_url":"https://mdrss.com/m/1709","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/1709/1709.md","file_url":"https://mdrss.com/api/v1/cards/1709/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/1709/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/1709/embed","edit_url":"https://mdrss.com/cards/1709/edit","legacy_url":"https://mdrss.com/s/multimodal/systran-faster-whisper-systran-faster-whisper-readme"},{"id":1611,"title":"MOVA: Towards Scalable and Synchronized Video–Audio Generation","annotation":"We introduce MOVA (MOSS Video and Audio), a foundation model designed to break the \"silent era\" of open-source video generation. Unlike cascaded pipelines that generate sound as an afterthought, MOVA synthesizes video and audio simultaneously for perfect alignment.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:49.845Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/1611","permalink_url":"https://mdrss.com/m/1611","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/1611/1611.md","file_url":"https://mdrss.com/api/v1/cards/1611/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/1611/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/1611/embed","edit_url":"https://mdrss.com/cards/1611/edit","legacy_url":"https://mdrss.com/s/multimodal/openmoss-mova-openmoss-mova-readme"},{"id":986,"title":"An AI-powered web application for speech recognition, translation, and dubbing","annotation":"Voice-Pro The best AI speech recognition, translation, and multilingual dubbing solution 🚀 한국어 ∙ English ∙ 中文简体 ∙ 中文繁體 ∙ 日本語 ∙ Deutsch ∙ Español ∙ Português Voice-Pro is a state-of-the-art web app that transforms multimedia content creation. It integrates YouTube video downloading, voice separation, speech recognition, translation, and text-to-speech into a single, powerful tool for creators, researchers, and multilingual professionals.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"reference","version":1,"snapshot_at":"2026-08-04T12:18:49.845Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/986","permalink_url":"https://mdrss.com/m/986","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/986/986.md","file_url":"https://mdrss.com/api/v1/cards/986/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/986/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/986/embed","edit_url":"https://mdrss.com/cards/986/edit","legacy_url":"https://mdrss.com/s/multimodal/abus-aikorea-voice-pro-abus-aikorea-voice-pro-readme"},{"id":853,"title":"What is Voicebox?","annotation":"The full voice I/O stack, running locally on your machine. voicebox.sh • Docs • Download • Features • API • Troubleshooting Click the image above to watch the demo video on voicebox.sh Voicebox is a local-first AI voice studio — a free and open-source alternative to ElevenLabs and WisprFlow in one app.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:49.845Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/853","permalink_url":"https://mdrss.com/m/853","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/853/853.md","file_url":"https://mdrss.com/api/v1/cards/853/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/853/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/853/embed","edit_url":"https://mdrss.com/cards/853/edit","legacy_url":"https://mdrss.com/s/multimodal/jamiepine-voicebox-jamiepine-voicebox-readme"},{"id":794,"title":"Moonshine Voice","annotation":"Voice Interfaces for Everyone Moonshine Voice is an open source AI toolkit for developers building real-time voice agents and applications. Join our community on Discord to get live support.","feed":"multimodal","semantic_path":"multimodal:speech-and-audio","content_type":"guide","version":1,"snapshot_at":"2026-08-04T12:18:49.845Z","publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","card_url":"https://mdrss.com/multimodal/speech-and-audio/794","permalink_url":"https://mdrss.com/m/794","thread_url":"https://mdrss.com/s/multimodal","markdown_url":"https://mdrss.com/multimodal/speech-and-audio/794/794.md","file_url":"https://mdrss.com/api/v1/cards/794/file","raw_url":"https://mdrss.com/multimodal/speech-and-audio/794/raw","embed_url":"https://mdrss.com/multimodal/speech-and-audio/794/embed","edit_url":"https://mdrss.com/cards/794/edit","legacy_url":"https://mdrss.com/s/multimodal/moonshine-ai-moonshine-moonshine-ai-moonshine-readme"}]}