Use this repository to compare multimodal model families, trace architectural ideas over time, or retrieve grounded references for research and AI-agent workflows. The catalog includes a verified release timeline through July 2026 and covers foundational systems such as CLIP and
👁️🗨️ Awesome VLM Architectures
Snapshot 2026-08-03 23:56:39 UTC · version 1
Research document
👁️🗨️ Awesome VLM Architectures
Awesome VLM Architectures is a citation-first visual catalog of 155+ Vision-Language Model (VLM/MLLM) architectures, spanning contrastive encoders, multimodal LLMs, native multimodal models, unified understanding and generation, video, OCR, GUI agents, and embodied AI. Each entry links to primary sources and summarizes the model architecture, modality alignment or fusion, training stages, datasets, and distinctive design choices, with an architectural figure when available.
Use this repository to compare multimodal model families, trace architectural ideas over time, or retrieve grounded references for research and AI-agent workflows. The catalog includes a verified release timeline through July 2026 and covers foundational systems such as CLIP and Flamingo alongside current multimodal reasoning and agentic models. Expand any model panel for its detailed architecture summary.
Last reviewed: August 1, 2026.
Contents
Citation
If this repository is useful in your work, you may cite it below. Please also cite the original paper for claims about any individual model; this catalog is a guide to the literature, not a substitute for it.
Created and maintained by Gökay Aydoğan at fal.ai (ORCID; gokay@fal.ai).
Architecture images are credited individually in the figure credits and remain subject to the rights described in the figure notice.
📚 BibTeX@misc{aydogan2024awesomevlmarchitectures,
author = {Gökay Aydoğan},
title = {Awesome VLM Architectures},
year = {2024},
howpublished = {\url{https://github.com/gokayfem/awesome-vlm-architectures}},
note = {GitHub repository, fal.ai},
url = {https://github.com/gokayfem/awesome-vlm-architectures}
}
Models
All architecture panels are ordered by release date, newest first. Models released on the same day retain editorial catalog order.
🧭 Chronological Model Index (155 architectures, newest first)2026: MODUS | Argus-Unified | Kimi K3 | Mage-VL | Inkling | Hy-Embodied-VLM | MonkeyOCRv2 | MiniMax M3 | InternVideo3 | Keye-VL 2.0 | Zamba2-VL | Cosmos 3 | Lance | ZAYA1-VL | Falcon Perception | GLM-5V-Turbo | PLaMo 2.1-VL | EXAONE 4.5 | BidirLM and BidirLM-Omni | Gemma 4 | Penguin-VL | Phi-4-Reasoning-Vision | V-SONAR and V-LCM | Qwen3.5 | Youtu-VL | Kimi K2.5 and K2.6 | Step3-VL-10B
2025: ERNIE 5.0 | DeepSeek-OCR | PaddleOCR-VL | Qwen3-VL | Step3 | GLM-4.1V-Thinking | ERNIE 4.5-VL | MiMo-VL | BAGEL | Seed1.5-VL | InternVL3 and InternVL3.5 | Kimi-VL | Llama 4 Scout and Maverick | Qwen2.5-Omni | Gemma 3 | Aya Vision | Phi-4-multimodal | SigLIP 2 | EVEv2 | Qwen2.5-VL | VideoLLaMA 3 | UI-TARS | MiniMax-01 | MiniCPM-o-2.6 | Eagle 2 | Sa2VA
2024: VideoChat-Flash | OmniVLM | Apollo | DeepSeek-VL2 | Maya | InternVL 2.5 | PaliGemma 2 | ShowUI | SmolVLM | AIMv2 | LLaVA-CoT | LLM2CLIP | Tarsier2 | Janus and Janus-Pro | ARIA | Emu3 | Molmo and PixMo | Llama 3.2-Vision | NVLM | Pixtral 12B | VILA-U | Qwen2-VL | EAGLE | Show-o | Idefics3-8B | Transfusion | mPLUG-Owl3 | VITA | LLaVA-OneVision | VILA² | INF-LLaVA | SlowFast-LLaVA | EVLM | InternLM-XComposer-2.5 | OMG-LLaVA | Cambrian-1 | EVE | Ovis | Parrot | ConvLLaVA | Phi-3-Vision and Phi-3.5-Vision | CogVLM2 | Chameleon | PaliGemma | xGen-MM (BLIP-3) | MANTIS | Moondream-next | Idefics2 | InternLM-XComposer2-4KHD | MM1 | DeepSeek-VL | AnyGPT | SPHINX-X | LLaVA 1.6 | MiniCPM-V | MouSi | InternLM-XComposer2 | MoE-LLaVA | moondream1 and moondream2 | FireLLaVA | COSMO
2023: TinyGPT-V | MobileVLM | Alpha-CLIP | Nous-Hermes-2-Vision - Mistral 7B | SPHINX | Florence-2 | u-LLaVA | LLaVA-Plus | OtterHD | CoVLM | GLaMM | Fuyu-8B | PaLI-3 Vision Language Models | MiniGPT-v2 | BakLLaVA | Ferret | LLaVA 1.5 | CogVLM | MetaCLIP | Qwen-VL | IDEFICS | BLIVA | KOSMOS-2 | LaVIN | InstructBLIP | ImageBind | LLaVA | MiniGPT-4 | SigLIP | OpenFlamingo | PaLM-E | KOSMOS-1 | BLIP-2
2022: MULTIINSTRUCT | PaLI | Flamingo | BLIP
2020: ViT
Release Timeline
Dates use the first documented official model release; when none is available, they use the paper's arXiv v1 submission or first technical report. Family point releases are folded into their first architecture release, and same-day entries retain catalog order.
🗓️ Release Timeline (155 architectures, newest first)| Date | Architecture | Distinctive contribution |
|---|---|---|
| 2026-07-28 | MODUS | Decoder-only any-to-any modeling without modality-specific heads or losses |
| 2026-07-28 | Argus-Unified | Hybrid continuous and discrete visual tokens for economical understanding and generation |
| 2026-07-27 | Kimi K3 | Kimi Delta Attention, Attention Residuals, and extremely sparse LatentMoE routing |
| 2026-07-27 | Mage-VL | Codec-native selective video tokenization with a proactive event gate |
| 2026-07-15 | Inkling | Relative-position million-context multimodal MoE trained from scratch |
| 2026-07-15 | Hy-Embodied-VLM | Action-centric sparse-MoE reasoning for physical-world agents |
| 2026-07-11 | MonkeyOCRv2 | Joint image-to-text and pixel-reconstruction pretraining for document vision |
| 2026-06-11 | MiniMax M3 | Native multimodality with block-sparse grouped-query attention at million-token context |
| 2026-06-10 | InternVideo3 | Token-preserving latent KV compression and closed-loop video reasoning |
| 2026-06-09 | Keye-VL 2.0 | DeepSeek Sparse Attention adapted to GQA-based long-video multimodality |
| 2026-06-02 | Zamba2-VL | Hybrid Mamba-2 and shared-attention blocks for efficient VLM inference |
| 2026-05-31 | Cosmos 3 | Coupled autoregressive reasoner and diffusion generator for physical AI |
| 2026-05-18 | Lance | Shared-sequence understanding, generation, and editing with modality experts |
| 2026-05-08 | ZAYA1-VL | Vision-conditional LoRA and compressed convolutional attention in an open-data MoE |
| 2026-05-03 | Falcon Perception | Early fusion with hybrid attention and continuous mask heads |
| 2026-04-29 | GLM-5V-Turbo | Perception integrated into reasoning, planning, tools, and execution |
| 2026-04-21 | PLaMo 2.1-VL | Compact Japanese VQA and grounding for edge deployment |
| 2026-04-09 | EXAONE 4.5 | Native multimodal pretraining with document-focused data and 256K context |
| 2026-04-02 | BidirLM and BidirLM-Omni | Converting causal decoders into bidirectional multimodal encoders |
| 2026-03-31 | Gemma 4 | Dense and MoE native multimodality, including an encoder-free 12B design |
| 2026-03-06 | Penguin-VL | Text-LLM-initialized vision encoder and priority-aware token compression |
| 2026-03-04 | Phi-4-Reasoning-Vision | Mid-fusion compact VLM with explicit reasoning and direct-answer modes |
| 2026-03-01 | V-SONAR and V-LCM | Vision-language alignment and prediction in multilingual concept space |
| 2026-02-16 | Qwen3.5 | Native early fusion with hybrid linear/full attention and sparse MoE variants |
| 2026-01-27 | Youtu-VL | Unified autoregressive visual tokens that emit dense vision outputs without task heads |
| 2026-01-27 | Kimi K2.5 and K2.6 | Trillion-parameter native multimodal MoE for agents and computer use |
| 2026-01-14 | Step3-VL-10B | Language-aligned perception encoder with 16-fold visual-token compression |
| 2025-11-13 | ERNIE 5.0 | One autoregressive sparse MoE for text, images, video, audio, and generation |
| 2025-10-20 | DeepSeek-OCR | DeepEncoder compresses high-resolution documents into very short visual contexts |
| 2025-10-16 | PaddleOCR-VL | NaViT-style dynamic resolution with a compact ERNIE decoder for document parsing |
| 2025-09-22 | Qwen3-VL | DeepStack multi-level ViT fusion and explicit video timestamp alignment |
| 2025-07-25 | Step3 | Model-system co-design for communication-efficient sparse-MoE multimodality |
| 2025-07-01 | GLM-4.1V-Thinking | Curriculum-sampled reinforcement learning for multimodal reasoning |
| 2025-06-30 | ERNIE 4.5-VL | Heterogeneous shared and modality-specific experts with isolated routing |
| 2025-06-04 | MiMo-VL | Four-stage multimodal pretraining followed by mixed on-policy RL |
| 2025-05-20 | BAGEL | Mixture-of-Transformer-Experts for understanding and generation |
| 2025-05-11 | Seed1.5-VL | Compact vision encoder with a 20B-active MoE for reasoning and agents |
| 2025-04-11 | InternVL3 and InternVL3.5 | Native multimodal pretraining, later extended with adaptive resolution and cascade RL |
| 2025-04-10 | Kimi-VL | MoonViT native-resolution packing with a sparse MoE decoder |
| 2025-04-05 | Llama 4 Scout and Maverick | Early-fusion native multimodality in sparse-MoE Scout and Maverick models |
| 2025-03-26 | Qwen2.5-Omni | Streaming Thinker-Talker architecture for multimodal input and speech output |
| 2025-03-12 | Gemma 3 | Efficient local/global attention with long-context image understanding |
| 2025-03-04 | Aya Vision | Cross-modal model merging for multilingual multimodality without language forgetting |
| 2025-03-03 | Phi-4-multimodal | Mixture-of-LoRAs for text, vision, and speech |
| 2025-02-20 | SigLIP 2 | Multilingual, localization-aware, native-aspect-ratio vision-language encoding |
| 2025-02-08 | EVEv2 | Improved Baselines for Encoder-Free Vision-Language Models |
| 2025-01-26 | Qwen2.5-VL | Enhanced Vision-Language Capabilities in the Qwen Series |
| 2025-01-21 | VideoLLaMA 3 | Frontier Multimodal Foundation Models for Image and Video Understanding |
| 2025-01-20 | UI-TARS | Pioneering Automated GUI Interaction with Native Agents |
| 2025-01-14 | MiniMax-01 | Scaling Foundation Models with Lightning Attention |
| 2025-01-12 | MiniCPM-o-2.6 | A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming |
| 2025-01-10 | Eagle 2 | Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models |
| 2025-01-07 | Sa2VA | Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos |
| 2024-12-31 | VideoChat-Flash | Hierarchical Compression for Long-Context Video Modeling |
| 2024-12-16 | OmniVLM | A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference |
| 2024-12-13 | Apollo | An Exploration of Video Understanding in Large Multimodal Models |
| 2024-12-13 | DeepSeek-VL2 | Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding |
| 2024-12-10 | Maya | An Instruction Finetuned Multilingual Multimodal Model |
| 2024-12-05 | InternVL 2.5 | Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling |
| 2024-12-04 | PaliGemma 2 | A Family of Versatile VLMs for Transfer |
| 2024-11-26 | ShowUI | UI-guided visual-token selection and interleaved action histories |
| 2024-11-26 | SmolVLM | A Small, Efficient, and Open-Source Vision-Language Model |
| 2024-11-21 | AIMv2 | Multimodal Autoregressive Pre-training of Large Vision Encoders |
| 2024-11-15 | LLaVA-CoT | Let Vision Language Models Reason Step-by-Step |
| 2024-11-06 | LLM2CLIP | Powerful Language Model Unlocks Richer Visual Representation |
| 2024-11-05 | Tarsier2 | Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding |
| 2024-10-17 | Janus and Janus-Pro | Decoupled visual encoders for understanding and generation with one transformer |
| 2024-10-08 | ARIA | An Open Multimodal Native Mixture-of-Experts Model |
| 2024-09-27 | Emu3 | One next-token objective over discrete text, image, and video tokens |
| 2024-09-25 | Molmo and PixMo | Open data pipeline with human captions and grounded pointing supervision |
| 2024-09-25 | Llama 3.2-Vision | Enhanced Multimodal Capabilities Built on Llama 3 |
| 2024-09-17 | NVLM | Open Frontier-Class Multimodal LLMs |
| 2024-09-11 | Pixtral 12B | A Cutting-Edge Open Multimodal Language Model |
| 2024-09-06 | VILA-U | Shared discrete visual tokens for autoregressive understanding and generation |
| 2024-08-29 | Qwen2-VL | A Powerful Open-Source Vision-Language Model for Image and Video Understanding |
| 2024-08-28 | EAGLE | Exploring The Design Space for Multimodal LLMs with Mixture of Encoders |
| 2024-08-22 | Show-o | Autoregressive language and discrete-diffusion image generation in one transformer |
| 2024-08-22 | Idefics3-8B | Building and Better Understanding Vision-Language Models |
| 2024-08-20 | Transfusion | Autoregressive text and continuous image diffusion in one transformer |
| 2024-08-09 | mPLUG-Owl3 | Hyper-attention for long image sequences and video |
| 2024-08-09 | VITA | Towards Open-Source Interactive Omni Multimodal LLM |
| 2024-08-05 | LLaVA-OneVision | Easy Visual Task Transfer |
| 2024-07-24 | VILA² | VILA Augmented VILA |
| 2024-07-23 | INF-LLaVA | High-Resolution Image Perception for Multimodal Large Language Models |
| 2024-07-22 | SlowFast-LLaVA | A Strong Training-Free Baseline for Video Large Language Models |
| 2024-07-19 | EVLM | An Efficient Vision-Language Model for Visual Understanding |
| 2024-07-03 | InternLM-XComposer-2.5 | A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output |
| 2024-06-27 | OMG-LLaVA | Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding |
| 2024-06-24 | Cambrian-1 | Spatial Vision Aggregator and systematic multi-encoder study |
| 2024-06-17 | EVE | Unveiling Encoder-Free Vision-Language Models |
| 2024-06-14 | Ovis | Learnable visual vocabulary for structural visual-text embedding alignment |
| 2024-06-04 | Parrot | Multilingual Visual Instruction Tuning |
| 2024-05-24 | ConvLLaVA | Hierarchical Backbones as Visual Encoder for Large Multimodal Models |
| 2024-05-21 | Phi-3-Vision and Phi-3.5-Vision | Compact dynamic-resolution VLM with 128K context |
| 2024-05-20 | CogVLM2 | Enhanced Vision-Language Models for Image and Video Understanding |
| 2024-05-16 | Chameleon | Mixed-modal early fusion over a shared token sequence |
| 2024-05-14 | PaliGemma | A Versatile and Transferable 3B Vision-Language Model |
| 2024-05-06 | xGen-MM (BLIP-3) | An Open-Source Framework for Building Powerful and Responsible Large Multimodal Models |
| 2024-05-02 | MANTIS | Mastering Multi-Image Understanding Through Interleaved Instruction Tuning |
| 2024-04-19 | Moondream-next | Compact Vision-Language Model with Enhanced Capabilities |
| 2024-04-15 | Idefics2 | Open 8B VLM with native-resolution inputs and strong OCR and document understanding |
| 2024-04-09 | InternLM-XComposer2-4KHD | A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD |
| 2024-03-14 | MM1 | Controlled study of encoders, connectors, token counts, and data mixtures |
| 2024-03-08 | DeepSeek-VL | Towards Real-World Vision-Language Understanding |
| 2024-02-19 | AnyGPT | Any-to-any autoregression over discrete text, image, speech, and music tokens |
| 2024-02-08 | SPHINX-X | Scaling Data and Parameters for a Family of Multi-modal Large Language Models |
| 2024-01-30 | LLaVA 1.6 | LLaVA-NeXT Improved reasoning, OCR, and world knowledge |
| 2024-01-30 | MiniCPM-V | A GPT-4V Level MLLM on Your Phone |
| 2024-01-30 | MouSi | Poly-Visual-Expert Vision-Language Models |
| 2024-01-29 | InternLM-XComposer2 | Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model |
| 2024-01-29 | MoE-LLaVA | Mixture of Experts for Large Vision-Language Models |
| 2024-01-20 | moondream1 and moondream2 | Compact SigLIP–Phi VLMs optimized for efficient edge inference |
| 2024-01-05 | FireLLaVA | LLaVA derivative trained rapidly on a curated multimodal instruction mixture |
| 2024-01-01 | COSMO | COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training |
| 2023-12-28 | TinyGPT-V | Efficient Multimodal Large Language Model via Small Backbones |
| 2023-12-28 | MobileVLM | A Fast, Strong and Open Vision Language Assistant for Mobile Devices |
| 2023-12-06 | Alpha-CLIP | A CLIP Model Focusing on Wherever You Want |
| 2023-11-28 | Nous-Hermes-2-Vision - Mistral 7B | SigLIP-equipped Mistral VLM with OCR and function-calling data |
| 2023-11-13 | SPHINX | The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models |
| 2023-11-10 | Florence-2 | A Deep Dive into its Unified Architecture and Multi-Task Capabilities |
| 2023-11-09 | u-LLaVA | Unifying Multi-Modal Tasks via Large Language Model |
| 2023-11-09 | LLaVA-Plus | Learning to Use Tools for Creating Multimodal Agents |
| 2023-11-07 | OtterHD | A High-Resolution Multi-modality Model |
| 2023-11-06 | CoVLM | Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding |
| 2023-11-06 | GLaMM | Pixel Grounding Large Multimodal Model |
| 2023-10-17 | Fuyu-8B | A Multimodal Architecture for AI Agents |
| 2023-10-13 | PaLI-3 Vision Language Models | Smaller, Faster, Stronger |
| 2023-10-13 | MiniGPT-v2 | large language model as a unified interface for vision-language multi-task learning |
| 2023-10-12 | BakLLaVA | Mistral-based LLaVA variant with a CLIP vision encoder and projection adapter |
| 2023-10-11 | Ferret | Refer and Ground Anything Anywhere at Any Granularity |
| 2023-10-05 | LLaVA 1.5 | Improved Baselines with Visual Instruction Tuning |
| 2023-10-05 | CogVLM | Visual Expert for Pretrained Language Models |
| 2023-09-28 | MetaCLIP | Demystifying CLIP Data |
| 2023-08-24 | Qwen-VL | A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond |
| 2023-08-22 | IDEFICS | Open Flamingo-style model for interleaved image-text generation |
| 2023-08-19 | BLIVA | A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions |
| 2023-06-26 | KOSMOS-2 | Grounding Multimodal Large Language Models to the World |
| 2023-05-24 | LaVIN | Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models |
| 2023-05-11 | InstructBLIP | Towards General-purpose Vision-Language Models with Instruction Tuning |
| 2023-05-09 | ImageBind | One Embedding Space To Bind Them All |
| 2023-04-17 | LLaVA | Large Language and Vision Assistant - Visual Instruction Tuning |
| 2023-04-16 | MiniGPT-4 | Enhancing Vision-Language Understanding with Advanced Large Language Models |
| 2023-03-27 | SigLIP | Sigmoid Loss for Language Image Pre-Training |
| 2023-03-14 | OpenFlamingo | An Open-Source Framework for Training Large Autoregressive Vision-Language Models |
| 2023-03-06 | PaLM-E | An Embodied Multimodal Language Model |
| 2023-02-27 | KOSMOS-1 | Language Is Not All You Need: Aligning Perception with Language Models |
| 2023-01-30 | BLIP-2 | Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models |
| 2022-12-21 | MULTIINSTRUCT | Improving Multi-Modal Zero-Shot Learning via Instruction Tuning |
| 2022-09-14 | PaLI | A Jointly-Scaled Multilingual Language-Image Model |
| 2022-04-28 | Flamingo | a Visual Language Model for Few-Shot Learning |
| 2022-01-28 | BLIP | Bootstrapping Language-Image Pre-training |
| 2021-12-07 | GLIP | Grounded Language-Image Pre-training |
| 2021-06-25 | FROZEN | Multimodal Few-Shot Learning with Frozen Language Models |
| 2021-01-05 | CLIP | Contrastive Language-Image Pre-training |
| 2020-10-22 | ViT | An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
Architectures
MODUS: Decoder-Only Any-to-Any Multimodal Modeling
MODUS treats every modality symmetrically as both input and output, enabling chained generation and cross-modal self-verification without modality-specific heads, losses, or task pipelines.
Mingqiao Ye et al., EPFL
Released: 2026-07-28
Figure 2. Decoder-only any-to-any modeling across tokenized 1D and 2D modalities. Source paper, PDF p. 4. Figure notice.
ℹ️ More InformationMODUS is a decoder-only any-to-any model that represents diverse modalities inside one autoregressive architecture. Unlike encoder-decoder or diffusion systems assembled around modality-specific output paths, the same model predicts any supported modality from any combination of the others. This makes intermediate-modality chains and self-scoring through a second generated modality native behaviors rather than external workflows.
The design deliberately reuses strong pretrained decoder-only priors instead of training a bespoke multimodal stack from scratch. A single checkpoint is evaluated across heterogeneous tasks and modalities, making MODUS most notable as a general architectural formulation rather than a narrowly optimized VLM endpoint.
Argus-Unified: Economical Understanding and Generation
Argus-Unified combines continuous tokens for understanding with learned discrete tokens for generation, reusing a frozen unified visual encoder and pretrained VLM to lower the cost of unified modeling.
Weiming Zhuang et al.
Released: 2026-07-28
Figure 3. Two-stage hybrid-token training for unified image understanding and generation. Source paper, PDF p. 4. Figure notice.
ℹ️ More InformationArgus-Unified resolves the conflicting visual representations required by comprehension and synthesis through hybrid visual tokens. Continuous encoder features preserve semantic information for understanding, while a learned quantizer produces discrete tokens for image generation from the same frozen visual encoder.
Training first learns the quantizer and image decoder, then initializes the language component from a pretrained VLM and trains unified multimodal prediction. The reported recipe uses 15.6 million examples and roughly $2,000 of compute, positioning the model as a reproducible baseline for understanding-and-generation research rather than a scale demonstration.
Kimi K3: Kimi Delta Attention at Trillion-Parameter Scale
Kimi K3 is a native multimodal sparse MoE that combines Kimi Delta Attention, gated latent attention, Attention Residuals, and 896 routed experts for million-token reasoning and agency.
Kimi Team, Moonshot AI
Released: 2026-07-27
Figure 2. Kimi Delta Attention, Stable LatentMoE, Attention Residuals, and MoonViT-V2. Source paper, PDF p. 3. Figure notice.
ℹ️ More InformationKimi K3 has 2.8T total and 104B active parameters across 93 layers. Sixty-nine layers use gated delta-rule linear attention and 24 use gated Multi-head Latent Attention; learned Attention Residuals create structured information paths across depth. Stable LatentMoE activates 16 of 896 routed experts plus two shared experts, making the model unusually sparse for its scale.
A 401M-parameter MoonViT-V2 encoder supplies native image and video input. The model supports a 1,048,576-token context and receives multimodal, long-context, agentic, reasoning, and tool-use training. Quantization-aware pretraining targets MXFP4 weights and MXFP8 activations rather than treating low-precision deployment as an afterthought.
Mage-VL: Codec-Native Streaming Multimodality
Mage-VL uses video-codec signals to encode only dynamic, information-rich regions and pairs a lightweight event gate with a causal decoder for proactive streaming perception.
Senqiao Yang et al., Microsoft Research
Released: 2026-07-27
Figure 3. Codec-native streaming perception with an event gate and causal language decoder. Source paper, PDF p. 8. Figure notice.
ℹ️ More InformationMage-ViT replaces uniform frame sampling with selective 16×16-patch encoding guided by motion vectors and residual energy across anchor and predicted frames. This codec-native tokenization reduces visual-token use by more than 75% while retaining spatiotemporal context. The encoder is trained from scratch on roughly 560M unlabeled images and 100M unlabeled video frames.
The complete VLM uses a dual-system design: a small event gate decides when a stream warrants deeper processing, and a causal decoder performs contextual reasoning. This turns streaming perception into an event-driven process and yields up to a reported 3.5× wall-clock speedup while retaining static-image and spatial reasoning ability.
Inkling: Relative-Position Multimodal Mixture of Experts
Inkling is an open-weight multimodal MoE trained from scratch on text, images, audio, and video, using relative positions and alternating local/global attention for million-token contexts.
Thinking Machines Lab
Released: 2026-07-15
ℹ️ More InformationArchitecture figure: The official Inkling model card contains no architecture figure.
Inkling contains 975B total and 41B active parameters. Its MoE layers activate six of 256 routed experts plus two shared experts using sigmoid routing and auxiliary-loss-free balancing. Attention alternates five local sliding-window layers with one global layer and uses learned relative-position representations rather than RoPE; short convolutions further refine key, value, attention, and MLP pathways.
Pretraining spans 45T text, image, audio, and video tokens. Post-training begins with a small synthetic supervised bootstrap and scales asynchronous reinforcement learning across multimodal understanding, tools, coding, mathematics, conversation, and safety. The released weights support controllable reasoning effort and contexts up to one million tokens.
Hy-Embodied-VLM: Sparse-MoE Reasoning for Physical Agents
Hy-Embodied-VLM-1.0 joins a native-aspect-ratio vision encoder to a compact sparse MoE and trains an action-centric reasoning hierarchy for perception, planning, reflection, and recovery.
Tencent Robotics X, Hy Vision Team and Futian Laboratory
Released: 2026-07-15
Figure 4. Self-evolving supervised fine-tuning, rejection sampling, and specialized reinforcement learning. Source paper, PDF p. 11. Figure notice.
ℹ️ More InformationHy-Embodied-VLM-1.0 connects Hy-ViT2 to a roughly 30B-total, 3B-active language model. Each sparse-MoE layer selects eight of 128 routed experts plus one shared expert. It accepts as many as 128 images in a 32K context and exposes direct-response and explicit-thinking modes through the same checkpoint.
Its curriculum moves from action-relevant state understanding to transition reasoning and then sequential, adaptive reasoning. A self-evolving post-training loop alternates reinforcement learning with rejection-sampling fine-tuning, separately optimizing geometric precision and higher-level planning before policy fusion.
MonkeyOCRv2: Document-Native Visual-Text Pretraining
MonkeyOCRv2 learns document vision jointly through image-to-text generation and pixel reconstruction, preserving both semantic content and character-level strokes across 17 languages.
Yuliang Liu et al.
Released: 2026-07-11
Figure 1. Document-native pretraining through text generation and pixel reconstruction. Source paper, PDF p. 2. Figure notice.
ℹ️ More InformationMonkeyOCRv2 is a visual-text foundation encoder specialized for the properties that natural-image pretraining often misses: dense text, small glyphs, fine strokes, and layout. Its two complementary objectives align images with their transcribed content while forcing the encoder to retain pixel-level document structure.
The MonkeyDoc v2 corpus contains 113M document images in 17 languages. The frozen encoder improves recognition, detection, tamper analysis, segmentation, parsing, and understanding; paired with a lightweight language model, it forms a 0.7B document parser. Public weights were announced July 11, two days before the arXiv report, so the timeline follows the model release rather than the paper date.
MiniMax M3: Native Multimodality with Sparse Long-Context Attention
MiniMax M3 combines native multimodal pretraining with learned block-sparse grouped-query attention, making million-token contexts practical in a 109B-parameter model.
MiniMax
Released: 2026-06-11
Figure 1. MiniMax Sparse Attention index and exact-attention branches. Source paper, PDF p. 1. Figure notice.
ℹ️ More InformationMiniMax Sparse Attention adds a lightweight index branch that scores KV blocks independently for each grouped-query-attention group. The main branch performs exact attention only over selected blocks, while a KV-outer kernel groups queries retrieving the same block to improve memory locality and reuse.
The released 109B-parameter model is trained natively across modalities and evaluated at contexts up to one million tokens. The sparse-attention report focuses on architecture, kernel design, and scaling behavior; it does not publish a complete source-level training-data inventory.
InternVideo3: Multimodal Contextual Reasoning for Video Agents
InternVideo3 reframes long-video understanding as an evolving loop of observation, reasoning, tools, and memory, while compressing KV state without dropping the underlying token stream.
Ziang Yan et al.
Released: 2026-06-10
Figure 2. InternVideo3 with multimodal multi-head latent attention across long contexts. Source paper, PDF p. 7. Figure notice.
ℹ️ More InformationMultimodal Contextual Reasoning maintains one evolving context containing observations, instructions, intermediate reasoning, tool actions, and memory. This closed loop lets the model accumulate and verify evidence across long videos rather than treating video QA as a single static prompt.
Multimodal Multi-head Latent Attention reparameterizes and compresses KV-cache states while preserving the complete visual-token sequence. Training combines continued pretraining, short-to-long supervised tuning, rule-based reinforcement learning, and on-policy distillation; a retrieval-equipped video-agent implementation demonstrates the intended tool-using setting.
Keye-VL 2.0: Sparse Attention for Long-Video Agents
Keye-VL 2.0 adapts DeepSeek Sparse Attention to a GQA-based multimodal MoE, targeting lossless 256K contexts, hour-scale video, and self-correcting tool-using agents.
Kwai Keye Team
Released: 2026-06-09
Figure 2. Four-stage curriculum extending Keye-VL from alignment to 256K context. Source paper, PDF p. 8. Figure notice.
ℹ️ More InformationKeye-VL-2.0-30B-A3B is the first reported adaptation of DeepSeek Sparse Attention to grouped-query multimodal attention. Only 3B of 30B parameters activate per token, while the sparse retrieval mechanism keeps critical frames and long-range dependencies accessible across a 256K context. Heterogeneous ViT-language-model parallelism and custom sparse-attention kernels address the systems cost of hour-long video.
Cross-Modal Multi-Teacher On-Policy Distillation feeds dense token-level guidance from specialist teachers back into on-policy trajectories. Context-RL and Video-RL then target long-context reasoning, temporal localization, code, search, tools, and multimodal self-correction without collapsing the multi-task mixture.
Zamba2-VL: Hybrid State-Space Vision-Language Modeling
Zamba2-VL combines efficient Mamba-2 state-space layers with a small number of shared Transformer blocks, reducing long-context prefill and recurrent-state costs across compact VLM scales.
Zyphra
Released: 2026-06-02
Figure 1. Zamba2 hybrid state-space language backbone connected to a vision encoder. Source paper, PDF p. 3. Figure notice.
ℹ️ More InformationZamba2-VL extends the Zamba hybrid backbone to vision-language inputs in 1.2B, 2.7B, and 7B variants. Most sequence processing occurs in Mamba-2 state-space layers, with a small set of shared attention blocks supplying global token interaction. This preserves transformer-like multimodal reasoning while moving more computation toward near-linear prefill and bounded recurrent state.
The architecture is particularly relevant for edge and long-context deployment: visual sequences can be large even when the language model is small, so reducing quadratic attention changes time-to-first-token and memory behavior more materially than ordinary parameter pruning.
Cosmos 3: Omnimodal World Modeling with Mixture of Transformers
Cosmos 3 couples an autoregressive reasoner with a diffusion generator through a shared representation for language, vision, audio, actions, simulation, and robot control.
NVIDIA
Released: 2026-05-31
Figure 5. Mixture-of-Transformers reasoner and generator with shared attention. Source paper, PDF p. 11. Figure notice.
ℹ️ More InformationCosmos 3 uses a Mixture-of-Transformers with two communicating towers. An autoregressive transformer reasons over language, vision, and action sequences; a diffusion transformer generates continuous images, video, and world trajectories. Their shared representation lets the family serve as a VLM, forward or inverse dynamics model, simulator, generator, or policy without separate task pipelines.
Nano configurations combine 8B-class components and Super configurations combine 32B-class components. Training spans text, image, video, audio, simulated physical interaction, action trajectories, robot demonstrations, driving, and synthetic worlds. The July 20 Cosmos 3 Edge release preserves the same two-tower design in a smaller real-time deployment branch and is folded into this family.
Lance: Unified Image and Video Understanding, Generation, and Editing
Lance is a compact model trained from scratch to understand, generate, and edit images and video inside one interleaved sequence while preserving specialized semantic and synthesis capacity.
Lance Team
Released: 2026-05-18
Figure 6. Dual-expert sequence modeling for understanding and visual generation. Source paper, PDF p. 9. Figure notice.
ℹ️ More InformationLance interleaves text, image, and video in a shared 3B-parameter sequence model while routing semantic understanding and visual synthesis through dedicated experts. Semantic ViT tokens coexist with clean or noised VAE latents, and generalized 3D causal attention plus modality-aware positional encoding preserve spatial and temporal relationships.
One checkpoint supports visual understanding, text-to-image and text-to-video generation, image-to-video generation, and editing. Its training-from-scratch recipe uses no more than 128 A100 GPUs, making the contribution as much about an attainable unified architecture as headline generation quality.
ZAYA1-VL: Vision-Specialized Compressed Convolutional Attention
ZAYA1-VL adds visual capacity to a sparse MoE through vision-only LoRA and compressed-convolutional-attention parameters, trained on an entirely open multimodal data mixture.
Zyphra
Released: 2026-05-08
Figure 2. Visual routing, compressed convolutional attention, and a hybrid language backbone. Source paper, PDF p. 5. Figure notice.
ℹ️ More InformationZAYA1-VL-8B joins a Qwen2.5-VL visual encoder to ZAYA1's MoE decoder. Vision-only LoRA and Compressed Convolutional Attention parameters activate for visual tokens without duplicating the language backbone, while bidirectional attention within image-token spans improves spatial integration.
The model is trained on roughly 140B vision-language tokens assembled from open data and released under Apache 2.0. Its importance is the clean separation of visual specialization from shared language capacity in a compact sparse model.
Falcon Perception: Early-Fusion Grounding, Segmentation, and OCR
Falcon Perception uses early image-text fusion, hybrid attention, instance tokens, and continuous mask heads to unify open-vocabulary grounding, segmentation, and OCR in compact models.
Technology Innovation Institute
Released: 2026-05-03
Figure 1. Early-fusion perception Transformer with grounding, geometry, and segmentation pathways. Source paper, PDF p. 3. Figure notice.
ℹ️ More InformationFalcon Perception is a dense early-fusion Transformer in which image patches and text interact from the first layer. Its attention mask is bidirectional over image tokens and causal over prediction tokens. Variable-length instance tokens identify multiple objects, while lightweight continuous heads produce masks without turning segmentation into an unwieldy text sequence.
The family includes a 600M perception model and a 300M OCR model. The paper was submitted March 28, but the timeline uses the documented May 3 public model launch, following this repository's release-date policy.
GLM-5V-Turbo: Native Multimodal Agency
GLM-5V-Turbo integrates visual perception into reasoning, planning, tool use, execution, and verification instead of treating images and video as an auxiliary language-model interface.
GLM-V Team
Released: 2026-04-29
This HTML preview is truncated for page performance. The canonical Markdown file contains the complete snapshot.
Why MDRSS assigned this score
- Production catalog audit 2026-08-04
- Taxonomy classified from title, annotation, source and Markdown signals
- Agent usefulness evaluated from structure, procedures, examples, evidence and retrieval value
Discussion 0
Sign in to join the discussion.