# Multimodal — MDRSS semantic catalog

> Semantic domain: multimodal
> Vision, speech, audio, games, media generation, robotics, embodied AI, and creative production.
> Filters: category=vision-and-media
> Aggregate catalog URL: https://mdrss.com/catalog/multimodal?category=vision-and-media

## Feeds

- [Health, Medicine & Psychology](https://mdrss.com/s/health-medicine-and-psychology) — Open clinical standards, medical data, public health, imaging, and mental health.
- [Lifestyle, Culture & Hobbies](https://mdrss.com/s/lifestyle-culture-and-hobbies) — Open cultural collections, games, cooking, creative work, and practical hobbies.

## Aggregate endpoints

- RSS: https://mdrss.com/catalog/multimodal/rss.xml?category=vision-and-media
- JSON: https://mdrss.com/catalog/multimodal/feed.json?category=vision-and-media

## Semantic domain cards (9)

### [Cosmos Tokenizer: A suite of image and video neural tokenizers](https://mdrss.com/multimodal/vision-and-media/2057/2057.md)

As of February 10th, 2025, this repository is read-only. Please visit github.com/NVIDIA/Cosmos for the latest updates and support on Cosmos Tokenizer.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [Model properties](https://mdrss.com/multimodal/vision-and-media/1706/1706.md)

USP : Unvoice and Silence with Pitch when infer 1. Install project dependencies Note: whisper is already built-in, do not install it again otherwise it will cuase conflict and error 3.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [text-extract-api](https://mdrss.com/multimodal/vision-and-media/1370/1370.md)

Convert any image, PDF or Office document to Markdown text or JSON structured document with super-high accuracy, including tabular data, numbers or math formulas. The API is built with FastAPI and uses Celery for asynchronous task processing.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:52.437Z · Version: 1

### [PICK-PyTorch](https://mdrss.com/multimodal/vision-and-media/1369/1369.md)

\\\\\ Updated on Feb 6th, 2021: Train Ticket dataset is now available for academic research. You can download from Google Drive or OneDrive.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:52.437Z · Version: 1

### [MuseV English 中文](https://mdrss.com/multimodal/vision-and-media/1345/1345.md)

MuseV was a milestone achieved around July 2023. Amazed by the progress of Sora, we decided to opensource MuseV, hopefully it will benefit the community.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [EasyAnimate | An End-to-End Solution for High-Resolution and Long Video Generation](https://mdrss.com/multimodal/vision-and-media/1304/1304.md)

😊 EasyAnimate is an end-to-end solution for generating high-resolution and long videos. We can train transformer based diffusion generators, train VAEs for processing long videos, and preprocess metadata.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:52.437Z · Version: 1

### [TurboDiffusion](https://mdrss.com/multimodal/vision-and-media/1255/1255.md)

This repository provides the official implementation of TurboDiffusion, a video generation acceleration framework that can speed up end-to-end diffusion generation by $100 \sim 200\times$ on a single RTX 5090, while maintaining video quality. TurboDiffusion primarily uses SageAttention, SLA (Sparse-Linear Attention) for attention acceleration, and rCM for timestep distillation.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:51.210Z · Version: 1

### [ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis](https://mdrss.com/multimodal/vision-and-media/1061/1061.md)

ViewCrafter can generate high-fidelity novel views from a single or sparse reference image , while also supporting highly precise pose control. Below shows some examples: Reference image Camera trajecotry Generated novel view video Reference image 1 Reference image 2 Generated novel view video |Model|Resolution|Frames|GPU Mem.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:52.437Z · Version: 1

### [Generating the documentation](https://mdrss.com/multimodal/vision-and-media/1039/1039.md)

To generate the documentation, you first have to build it. You don't have to commit the built documentation.

Feed: [multimodal](https://mdrss.com/s/multimodal) · Snapshot: 2026-08-04T12:18:52.437Z · Version: 1
