# Multimodal — MDRSS semantic catalog

> Semantic domain: multimodal
> Vision, speech, audio, games, media generation, robotics, embodied AI, and creative production.
> Filters: category=speech-and-audio
> Aggregate catalog URL: https://mdrss.com/catalog/multimodal?category=speech-and-audio

## Feeds

- [Health, Medicine & Psychology](https://mdrss.com/s/health-medicine-and-psychology) — Open clinical standards, medical data, public health, imaging, and mental health.
- [Lifestyle, Culture & Hobbies](https://mdrss.com/s/lifestyle-culture-and-hobbies) — Open cultural collections, games, cooking, creative work, and practical hobbies.

## Aggregate endpoints

- RSS: https://mdrss.com/catalog/multimodal/rss.xml?category=speech-and-audio
- JSON: https://mdrss.com/catalog/multimodal/feed.json?category=speech-and-audio

## Semantic domain cards (7)

### [Coqui.ai News](https://mdrss.com/multimodal/speech-and-audio/2516/2516.md)

🐸TTS is a library for advanced Text-to-Speech generation. 🛠️ Tools for training new models and fine-tuning existing models in any language.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [Whisper](https://mdrss.com/multimodal/speech-and-audio/2474/2474.md)

It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification. A Transformer sequence-to-sequence model is trained on various speech processing tasks, including multilingual speech recognition, speech translation, spoken language identification, and voice activity detection.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [Faster Whisper transcription with CTranslate2](https://mdrss.com/multimodal/speech-and-audio/1709/1709.md)

faster-whisper is a reimplementation of OpenAI's Whisper model using CTranslate2, which is a fast inference engine for Transformer models. This implementation is up to 4 times faster than openai/whisper for the same accuracy while using less memory.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:49.845Z · Version: 1

### [MOVA: Towards Scalable and Synchronized Video–Audio Generation](https://mdrss.com/multimodal/speech-and-audio/1611/1611.md)

We introduce MOVA (MOSS Video and Audio), a foundation model designed to break the "silent era" of open-source video generation. Unlike cascaded pipelines that generate sound as an afterthought, MOVA synthesizes video and audio simultaneously for perfect alignment.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:49.845Z · Version: 1

### [An AI-powered web application for speech recognition, translation, and dubbing](https://mdrss.com/multimodal/speech-and-audio/986/986.md)

Voice-Pro The best AI speech recognition, translation, and multilingual dubbing solution 🚀 한국어 ∙ English ∙ 中文简体 ∙ 中文繁體 ∙ 日本語 ∙ Deutsch ∙ Español ∙ Português Voice-Pro is a state-of-the-art web app that transforms multimedia content creation. It integrates YouTube video downloading, voice separation, speech recognition, translation, and text-to-speech into a single, powerful tool for creators, researchers, and multilingual professionals.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:49.845Z · Version: 1

### [What is Voicebox?](https://mdrss.com/multimodal/speech-and-audio/853/853.md)

The full voice I/O stack, running locally on your machine. voicebox.sh • Docs • Download • Features • API • Troubleshooting Click the image above to watch the demo video on voicebox.sh Voicebox is a local-first AI voice studio — a free and open-source alternative to ElevenLabs and WisprFlow in one app.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:49.845Z · Version: 1

### [Moonshine Voice](https://mdrss.com/multimodal/speech-and-audio/794/794.md)

Voice Interfaces for Everyone Moonshine Voice is an open source AI toolkit for developers building real-time voice agents and applications. Join our community on Discord to get live support.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:49.845Z · Version: 1
