👁️‍🗨️ Awesome VLM Architectures

Snapshot 2026-08-03 23:56:39 UTC · version 1

published
M
MDRSS Source Library Github collector716 cards · 4.7/10 MDRSS

Use this repository to compare multimodal model families, trace architectural ideas over time, or retrieve grounded references for research and AI-agent workflows. The catalog includes a verified release timeline through July 2026 and covers foundational systems such as CLIP and

MARKDOWN SNAPSHOT

Loading…

Direct .mdRaw + metadata0 commentsMDRSS 4.6/10
INDEXABLE MARKDOWN SNAPSHOT

Research document

Open canonical .md

👁️‍🗨️ Awesome VLM Architectures

Awesome VLM Architectures is a citation-first visual catalog of 155+ Vision-Language Model (VLM/MLLM) architectures, spanning contrastive encoders, multimodal LLMs, native multimodal models, unified understanding and generation, video, OCR, GUI agents, and embodied AI. Each entry links to primary sources and summarizes the model architecture, modality alignment or fusion, training stages, datasets, and distinctive design choices, with an architectural figure when available.

Use this repository to compare multimodal model families, trace architectural ideas over time, or retrieve grounded references for research and AI-agent workflows. The catalog includes a verified release timeline through July 2026 and covers foundational systems such as CLIP and Flamingo alongside current multimodal reasoning and agentic models. Expand any model panel for its detailed architecture summary.

Last reviewed: August 1, 2026.

Contents

Citation

If this repository is useful in your work, you may cite it below. Please also cite the original paper for claims about any individual model; this catalog is a guide to the literature, not a substitute for it.

Created and maintained by Gökay Aydoğan at fal.ai (ORCID; gokay@fal.ai).

Architecture images are credited individually in the figure credits and remain subject to the rights described in the figure notice.

📚 BibTeX
@misc{aydogan2024awesomevlmarchitectures,
  author       = {Gökay Aydoğan},
  title        = {Awesome VLM Architectures},
  year         = {2024},
  howpublished = {\url{https://github.com/gokayfem/awesome-vlm-architectures}},
  note         = {GitHub repository, fal.ai},
  url          = {https://github.com/gokayfem/awesome-vlm-architectures}
}

Models

All architecture panels are ordered by release date, newest first. Models released on the same day retain editorial catalog order.

🧭 Chronological Model Index (155 architectures, newest first)

2026: MODUS | Argus-Unified | Kimi K3 | Mage-VL | Inkling | Hy-Embodied-VLM | MonkeyOCRv2 | MiniMax M3 | InternVideo3 | Keye-VL 2.0 | Zamba2-VL | Cosmos 3 | Lance | ZAYA1-VL | Falcon Perception | GLM-5V-Turbo | PLaMo 2.1-VL | EXAONE 4.5 | BidirLM and BidirLM-Omni | Gemma 4 | Penguin-VL | Phi-4-Reasoning-Vision | V-SONAR and V-LCM | Qwen3.5 | Youtu-VL | Kimi K2.5 and K2.6 | Step3-VL-10B

2025: ERNIE 5.0 | DeepSeek-OCR | PaddleOCR-VL | Qwen3-VL | Step3 | GLM-4.1V-Thinking | ERNIE 4.5-VL | MiMo-VL | BAGEL | Seed1.5-VL | InternVL3 and InternVL3.5 | Kimi-VL | Llama 4 Scout and Maverick | Qwen2.5-Omni | Gemma 3 | Aya Vision | Phi-4-multimodal | SigLIP 2 | EVEv2 | Qwen2.5-VL | VideoLLaMA 3 | UI-TARS | MiniMax-01 | MiniCPM-o-2.6 | Eagle 2 | Sa2VA

2024: VideoChat-Flash | OmniVLM | Apollo | DeepSeek-VL2 | Maya | InternVL 2.5 | PaliGemma 2 | ShowUI | SmolVLM | AIMv2 | LLaVA-CoT | LLM2CLIP | Tarsier2 | Janus and Janus-Pro | ARIA | Emu3 | Molmo and PixMo | Llama 3.2-Vision | NVLM | Pixtral 12B | VILA-U | Qwen2-VL | EAGLE | Show-o | Idefics3-8B | Transfusion | mPLUG-Owl3 | VITA | LLaVA-OneVision | VILA² | INF-LLaVA | SlowFast-LLaVA | EVLM | InternLM-XComposer-2.5 | OMG-LLaVA | Cambrian-1 | EVE | Ovis | Parrot | ConvLLaVA | Phi-3-Vision and Phi-3.5-Vision | CogVLM2 | Chameleon | PaliGemma | xGen-MM (BLIP-3) | MANTIS | Moondream-next | Idefics2 | InternLM-XComposer2-4KHD | MM1 | DeepSeek-VL | AnyGPT | SPHINX-X | LLaVA 1.6 | MiniCPM-V | MouSi | InternLM-XComposer2 | MoE-LLaVA | moondream1 and moondream2 | FireLLaVA | COSMO

2023: TinyGPT-V | MobileVLM | Alpha-CLIP | Nous-Hermes-2-Vision - Mistral 7B | SPHINX | Florence-2 | u-LLaVA | LLaVA-Plus | OtterHD | CoVLM | GLaMM | Fuyu-8B | PaLI-3 Vision Language Models | MiniGPT-v2 | BakLLaVA | Ferret | LLaVA 1.5 | CogVLM | MetaCLIP | Qwen-VL | IDEFICS | BLIVA | KOSMOS-2 | LaVIN | InstructBLIP | ImageBind | LLaVA | MiniGPT-4 | SigLIP | OpenFlamingo | PaLM-E | KOSMOS-1 | BLIP-2

2022: MULTIINSTRUCT | PaLI | Flamingo | BLIP

2021: GLIP | FROZEN | CLIP

2020: ViT

Release Timeline

Dates use the first documented official model release; when none is available, they use the paper's arXiv v1 submission or first technical report. Family point releases are folded into their first architecture release, and same-day entries retain catalog order.

🗓️ Release Timeline (155 architectures, newest first)
Date Architecture Distinctive contribution
2026-07-28 MODUS Decoder-only any-to-any modeling without modality-specific heads or losses
2026-07-28 Argus-Unified Hybrid continuous and discrete visual tokens for economical understanding and generation
2026-07-27 Kimi K3 Kimi Delta Attention, Attention Residuals, and extremely sparse LatentMoE routing
2026-07-27 Mage-VL Codec-native selective video tokenization with a proactive event gate
2026-07-15 Inkling Relative-position million-context multimodal MoE trained from scratch
2026-07-15 Hy-Embodied-VLM Action-centric sparse-MoE reasoning for physical-world agents
2026-07-11 MonkeyOCRv2 Joint image-to-text and pixel-reconstruction pretraining for document vision
2026-06-11 MiniMax M3 Native multimodality with block-sparse grouped-query attention at million-token context
2026-06-10 InternVideo3 Token-preserving latent KV compression and closed-loop video reasoning
2026-06-09 Keye-VL 2.0 DeepSeek Sparse Attention adapted to GQA-based long-video multimodality
2026-06-02 Zamba2-VL Hybrid Mamba-2 and shared-attention blocks for efficient VLM inference
2026-05-31 Cosmos 3 Coupled autoregressive reasoner and diffusion generator for physical AI
2026-05-18 Lance Shared-sequence understanding, generation, and editing with modality experts
2026-05-08 ZAYA1-VL Vision-conditional LoRA and compressed convolutional attention in an open-data MoE
2026-05-03 Falcon Perception Early fusion with hybrid attention and continuous mask heads
2026-04-29 GLM-5V-Turbo Perception integrated into reasoning, planning, tools, and execution
2026-04-21 PLaMo 2.1-VL Compact Japanese VQA and grounding for edge deployment
2026-04-09 EXAONE 4.5 Native multimodal pretraining with document-focused data and 256K context
2026-04-02 BidirLM and BidirLM-Omni Converting causal decoders into bidirectional multimodal encoders
2026-03-31 Gemma 4 Dense and MoE native multimodality, including an encoder-free 12B design
2026-03-06 Penguin-VL Text-LLM-initialized vision encoder and priority-aware token compression
2026-03-04 Phi-4-Reasoning-Vision Mid-fusion compact VLM with explicit reasoning and direct-answer modes
2026-03-01 V-SONAR and V-LCM Vision-language alignment and prediction in multilingual concept space
2026-02-16 Qwen3.5 Native early fusion with hybrid linear/full attention and sparse MoE variants
2026-01-27 Youtu-VL Unified autoregressive visual tokens that emit dense vision outputs without task heads
2026-01-27 Kimi K2.5 and K2.6 Trillion-parameter native multimodal MoE for agents and computer use
2026-01-14 Step3-VL-10B Language-aligned perception encoder with 16-fold visual-token compression
2025-11-13 ERNIE 5.0 One autoregressive sparse MoE for text, images, video, audio, and generation
2025-10-20 DeepSeek-OCR DeepEncoder compresses high-resolution documents into very short visual contexts
2025-10-16 PaddleOCR-VL NaViT-style dynamic resolution with a compact ERNIE decoder for document parsing
2025-09-22 Qwen3-VL DeepStack multi-level ViT fusion and explicit video timestamp alignment
2025-07-25 Step3 Model-system co-design for communication-efficient sparse-MoE multimodality
2025-07-01 GLM-4.1V-Thinking Curriculum-sampled reinforcement learning for multimodal reasoning
2025-06-30 ERNIE 4.5-VL Heterogeneous shared and modality-specific experts with isolated routing
2025-06-04 MiMo-VL Four-stage multimodal pretraining followed by mixed on-policy RL
2025-05-20 BAGEL Mixture-of-Transformer-Experts for understanding and generation
2025-05-11 Seed1.5-VL Compact vision encoder with a 20B-active MoE for reasoning and agents
2025-04-11 InternVL3 and InternVL3.5 Native multimodal pretraining, later extended with adaptive resolution and cascade RL
2025-04-10 Kimi-VL MoonViT native-resolution packing with a sparse MoE decoder
2025-04-05 Llama 4 Scout and Maverick Early-fusion native multimodality in sparse-MoE Scout and Maverick models
2025-03-26 Qwen2.5-Omni Streaming Thinker-Talker architecture for multimodal input and speech output
2025-03-12 Gemma 3 Efficient local/global attention with long-context image understanding
2025-03-04 Aya Vision Cross-modal model merging for multilingual multimodality without language forgetting
2025-03-03 Phi-4-multimodal Mixture-of-LoRAs for text, vision, and speech
2025-02-20 SigLIP 2 Multilingual, localization-aware, native-aspect-ratio vision-language encoding
2025-02-08 EVEv2 Improved Baselines for Encoder-Free Vision-Language Models
2025-01-26 Qwen2.5-VL Enhanced Vision-Language Capabilities in the Qwen Series
2025-01-21 VideoLLaMA 3 Frontier Multimodal Foundation Models for Image and Video Understanding
2025-01-20 UI-TARS Pioneering Automated GUI Interaction with Native Agents
2025-01-14 MiniMax-01 Scaling Foundation Models with Lightning Attention
2025-01-12 MiniCPM-o-2.6 A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
2025-01-10 Eagle 2 Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
2025-01-07 Sa2VA Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
2024-12-31 VideoChat-Flash Hierarchical Compression for Long-Context Video Modeling
2024-12-16 OmniVLM A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference
2024-12-13 Apollo An Exploration of Video Understanding in Large Multimodal Models
2024-12-13 DeepSeek-VL2 Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
2024-12-10 Maya An Instruction Finetuned Multilingual Multimodal Model
2024-12-05 InternVL 2.5 Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
2024-12-04 PaliGemma 2 A Family of Versatile VLMs for Transfer
2024-11-26 ShowUI UI-guided visual-token selection and interleaved action histories
2024-11-26 SmolVLM A Small, Efficient, and Open-Source Vision-Language Model
2024-11-21 AIMv2 Multimodal Autoregressive Pre-training of Large Vision Encoders
2024-11-15 LLaVA-CoT Let Vision Language Models Reason Step-by-Step
2024-11-06 LLM2CLIP Powerful Language Model Unlocks Richer Visual Representation
2024-11-05 Tarsier2 Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
2024-10-17 Janus and Janus-Pro Decoupled visual encoders for understanding and generation with one transformer
2024-10-08 ARIA An Open Multimodal Native Mixture-of-Experts Model
2024-09-27 Emu3 One next-token objective over discrete text, image, and video tokens
2024-09-25 Molmo and PixMo Open data pipeline with human captions and grounded pointing supervision
2024-09-25 Llama 3.2-Vision Enhanced Multimodal Capabilities Built on Llama 3
2024-09-17 NVLM Open Frontier-Class Multimodal LLMs
2024-09-11 Pixtral 12B A Cutting-Edge Open Multimodal Language Model
2024-09-06 VILA-U Shared discrete visual tokens for autoregressive understanding and generation
2024-08-29 Qwen2-VL A Powerful Open-Source Vision-Language Model for Image and Video Understanding
2024-08-28 EAGLE Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
2024-08-22 Show-o Autoregressive language and discrete-diffusion image generation in one transformer
2024-08-22 Idefics3-8B Building and Better Understanding Vision-Language Models
2024-08-20 Transfusion Autoregressive text and continuous image diffusion in one transformer
2024-08-09 mPLUG-Owl3 Hyper-attention for long image sequences and video
2024-08-09 VITA Towards Open-Source Interactive Omni Multimodal LLM
2024-08-05 LLaVA-OneVision Easy Visual Task Transfer
2024-07-24 VILA² VILA Augmented VILA
2024-07-23 INF-LLaVA High-Resolution Image Perception for Multimodal Large Language Models
2024-07-22 SlowFast-LLaVA A Strong Training-Free Baseline for Video Large Language Models
2024-07-19 EVLM An Efficient Vision-Language Model for Visual Understanding
2024-07-03 InternLM-XComposer-2.5 A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
2024-06-27 OMG-LLaVA Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
2024-06-24 Cambrian-1 Spatial Vision Aggregator and systematic multi-encoder study
2024-06-17 EVE Unveiling Encoder-Free Vision-Language Models
2024-06-14 Ovis Learnable visual vocabulary for structural visual-text embedding alignment
2024-06-04 Parrot Multilingual Visual Instruction Tuning
2024-05-24 ConvLLaVA Hierarchical Backbones as Visual Encoder for Large Multimodal Models
2024-05-21 Phi-3-Vision and Phi-3.5-Vision Compact dynamic-resolution VLM with 128K context
2024-05-20 CogVLM2 Enhanced Vision-Language Models for Image and Video Understanding
2024-05-16 Chameleon Mixed-modal early fusion over a shared token sequence
2024-05-14 PaliGemma A Versatile and Transferable 3B Vision-Language Model
2024-05-06 xGen-MM (BLIP-3) An Open-Source Framework for Building Powerful and Responsible Large Multimodal Models
2024-05-02 MANTIS Mastering Multi-Image Understanding Through Interleaved Instruction Tuning
2024-04-19 Moondream-next Compact Vision-Language Model with Enhanced Capabilities
2024-04-15 Idefics2 Open 8B VLM with native-resolution inputs and strong OCR and document understanding
2024-04-09 InternLM-XComposer2-4KHD A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
2024-03-14 MM1 Controlled study of encoders, connectors, token counts, and data mixtures
2024-03-08 DeepSeek-VL Towards Real-World Vision-Language Understanding
2024-02-19 AnyGPT Any-to-any autoregression over discrete text, image, speech, and music tokens
2024-02-08 SPHINX-X Scaling Data and Parameters for a Family of Multi-modal Large Language Models
2024-01-30 LLaVA 1.6 LLaVA-NeXT Improved reasoning, OCR, and world knowledge
2024-01-30 MiniCPM-V A GPT-4V Level MLLM on Your Phone
2024-01-30 MouSi Poly-Visual-Expert Vision-Language Models
2024-01-29 InternLM-XComposer2 Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
2024-01-29 MoE-LLaVA Mixture of Experts for Large Vision-Language Models
2024-01-20 moondream1 and moondream2 Compact SigLIP–Phi VLMs optimized for efficient edge inference
2024-01-05 FireLLaVA LLaVA derivative trained rapidly on a curated multimodal instruction mixture
2024-01-01 COSMO COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
2023-12-28 TinyGPT-V Efficient Multimodal Large Language Model via Small Backbones
2023-12-28 MobileVLM A Fast, Strong and Open Vision Language Assistant for Mobile Devices
2023-12-06 Alpha-CLIP A CLIP Model Focusing on Wherever You Want
2023-11-28 Nous-Hermes-2-Vision - Mistral 7B SigLIP-equipped Mistral VLM with OCR and function-calling data
2023-11-13 SPHINX The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
2023-11-10 Florence-2 A Deep Dive into its Unified Architecture and Multi-Task Capabilities
2023-11-09 u-LLaVA Unifying Multi-Modal Tasks via Large Language Model
2023-11-09 LLaVA-Plus Learning to Use Tools for Creating Multimodal Agents
2023-11-07 OtterHD A High-Resolution Multi-modality Model
2023-11-06 CoVLM Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding
2023-11-06 GLaMM Pixel Grounding Large Multimodal Model
2023-10-17 Fuyu-8B A Multimodal Architecture for AI Agents
2023-10-13 PaLI-3 Vision Language Models Smaller, Faster, Stronger
2023-10-13 MiniGPT-v2 large language model as a unified interface for vision-language multi-task learning
2023-10-12 BakLLaVA Mistral-based LLaVA variant with a CLIP vision encoder and projection adapter
2023-10-11 Ferret Refer and Ground Anything Anywhere at Any Granularity
2023-10-05 LLaVA 1.5 Improved Baselines with Visual Instruction Tuning
2023-10-05 CogVLM Visual Expert for Pretrained Language Models
2023-09-28 MetaCLIP Demystifying CLIP Data
2023-08-24 Qwen-VL A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
2023-08-22 IDEFICS Open Flamingo-style model for interleaved image-text generation
2023-08-19 BLIVA A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions
2023-06-26 KOSMOS-2 Grounding Multimodal Large Language Models to the World
2023-05-24 LaVIN Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models
2023-05-11 InstructBLIP Towards General-purpose Vision-Language Models with Instruction Tuning
2023-05-09 ImageBind One Embedding Space To Bind Them All
2023-04-17 LLaVA Large Language and Vision Assistant - Visual Instruction Tuning
2023-04-16 MiniGPT-4 Enhancing Vision-Language Understanding with Advanced Large Language Models
2023-03-27 SigLIP Sigmoid Loss for Language Image Pre-Training
2023-03-14 OpenFlamingo An Open-Source Framework for Training Large Autoregressive Vision-Language Models
2023-03-06 PaLM-E An Embodied Multimodal Language Model
2023-02-27 KOSMOS-1 Language Is Not All You Need: Aligning Perception with Language Models
2023-01-30 BLIP-2 Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
2022-12-21 MULTIINSTRUCT Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
2022-09-14 PaLI A Jointly-Scaled Multilingual Language-Image Model
2022-04-28 Flamingo a Visual Language Model for Few-Shot Learning
2022-01-28 BLIP Bootstrapping Language-Image Pre-training
2021-12-07 GLIP Grounded Language-Image Pre-training
2021-06-25 FROZEN Multimodal Few-Shot Learning with Frozen Language Models
2021-01-05 CLIP Contrastive Language-Image Pre-training
2020-10-22 ViT An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Architectures

MODUS: Decoder-Only Any-to-Any Multimodal Modeling

MODUS treats every modality symmetrically as both input and output, enabling chained generation and cross-modal self-verification without modality-specific heads, losses, or task pipelines.

Mingqiao Ye et al., EPFL
Released: 2026-07-28

Figure 2. Decoder-only any-to-any modeling across tokenized 1D and 2D modalities. Source paper, PDF p. 4. Figure notice.

ℹ️ More Information

MODUS is a decoder-only any-to-any model that represents diverse modalities inside one autoregressive architecture. Unlike encoder-decoder or diffusion systems assembled around modality-specific output paths, the same model predicts any supported modality from any combination of the others. This makes intermediate-modality chains and self-scoring through a second generated modality native behaviors rather than external workflows.

The design deliberately reuses strong pretrained decoder-only priors instead of training a bespoke multimodal stack from scratch. A single checkpoint is evaluated across heterogeneous tasks and modalities, making MODUS most notable as a general architectural formulation rather than a narrowly optimized VLM endpoint.

Argus-Unified: Economical Understanding and Generation

Argus-Unified combines continuous tokens for understanding with learned discrete tokens for generation, reusing a frozen unified visual encoder and pretrained VLM to lower the cost of unified modeling.

Weiming Zhuang et al.
Released: 2026-07-28

Figure 3. Two-stage hybrid-token training for unified image understanding and generation. Source paper, PDF p. 4. Figure notice.

ℹ️ More Information

Argus-Unified resolves the conflicting visual representations required by comprehension and synthesis through hybrid visual tokens. Continuous encoder features preserve semantic information for understanding, while a learned quantizer produces discrete tokens for image generation from the same frozen visual encoder.

Training first learns the quantizer and image decoder, then initializes the language component from a pretrained VLM and trains unified multimodal prediction. The reported recipe uses 15.6 million examples and roughly $2,000 of compute, positioning the model as a reproducible baseline for understanding-and-generation research rather than a scale demonstration.

Kimi K3: Kimi Delta Attention at Trillion-Parameter Scale

Kimi K3 is a native multimodal sparse MoE that combines Kimi Delta Attention, gated latent attention, Attention Residuals, and 896 routed experts for million-token reasoning and agency.

Kimi Team, Moonshot AI
Released: 2026-07-27

Figure 2. Kimi Delta Attention, Stable LatentMoE, Attention Residuals, and MoonViT-V2. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

Kimi K3 has 2.8T total and 104B active parameters across 93 layers. Sixty-nine layers use gated delta-rule linear attention and 24 use gated Multi-head Latent Attention; learned Attention Residuals create structured information paths across depth. Stable LatentMoE activates 16 of 896 routed experts plus two shared experts, making the model unusually sparse for its scale.

A 401M-parameter MoonViT-V2 encoder supplies native image and video input. The model supports a 1,048,576-token context and receives multimodal, long-context, agentic, reasoning, and tool-use training. Quantization-aware pretraining targets MXFP4 weights and MXFP8 activations rather than treating low-precision deployment as an afterthought.

Mage-VL: Codec-Native Streaming Multimodality

Mage-VL uses video-codec signals to encode only dynamic, information-rich regions and pairs a lightweight event gate with a causal decoder for proactive streaming perception.

Senqiao Yang et al., Microsoft Research
Released: 2026-07-27

Figure 3. Codec-native streaming perception with an event gate and causal language decoder. Source paper, PDF p. 8. Figure notice.

ℹ️ More Information

Mage-ViT replaces uniform frame sampling with selective 16×16-patch encoding guided by motion vectors and residual energy across anchor and predicted frames. This codec-native tokenization reduces visual-token use by more than 75% while retaining spatiotemporal context. The encoder is trained from scratch on roughly 560M unlabeled images and 100M unlabeled video frames.

The complete VLM uses a dual-system design: a small event gate decides when a stream warrants deeper processing, and a causal decoder performs contextual reasoning. This turns streaming perception into an event-driven process and yields up to a reported 3.5× wall-clock speedup while retaining static-image and spatial reasoning ability.

Inkling: Relative-Position Multimodal Mixture of Experts

Inkling is an open-weight multimodal MoE trained from scratch on text, images, audio, and video, using relative positions and alternating local/global attention for million-token contexts.

Thinking Machines Lab
Released: 2026-07-15

Architecture figure: The official Inkling model card contains no architecture figure.

ℹ️ More Information

Inkling contains 975B total and 41B active parameters. Its MoE layers activate six of 256 routed experts plus two shared experts using sigmoid routing and auxiliary-loss-free balancing. Attention alternates five local sliding-window layers with one global layer and uses learned relative-position representations rather than RoPE; short convolutions further refine key, value, attention, and MLP pathways.

Pretraining spans 45T text, image, audio, and video tokens. Post-training begins with a small synthetic supervised bootstrap and scales asynchronous reinforcement learning across multimodal understanding, tools, coding, mathematics, conversation, and safety. The released weights support controllable reasoning effort and contexts up to one million tokens.

Hy-Embodied-VLM: Sparse-MoE Reasoning for Physical Agents

Hy-Embodied-VLM-1.0 joins a native-aspect-ratio vision encoder to a compact sparse MoE and trains an action-centric reasoning hierarchy for perception, planning, reflection, and recovery.

Tencent Robotics X, Hy Vision Team and Futian Laboratory
Released: 2026-07-15

Figure 4. Self-evolving supervised fine-tuning, rejection sampling, and specialized reinforcement learning. Source paper, PDF p. 11. Figure notice.

ℹ️ More Information

Hy-Embodied-VLM-1.0 connects Hy-ViT2 to a roughly 30B-total, 3B-active language model. Each sparse-MoE layer selects eight of 128 routed experts plus one shared expert. It accepts as many as 128 images in a 32K context and exposes direct-response and explicit-thinking modes through the same checkpoint.

Its curriculum moves from action-relevant state understanding to transition reasoning and then sequential, adaptive reasoning. A self-evolving post-training loop alternates reinforcement learning with rejection-sampling fine-tuning, separately optimizing geometric precision and higher-level planning before policy fusion.

MonkeyOCRv2: Document-Native Visual-Text Pretraining

MonkeyOCRv2 learns document vision jointly through image-to-text generation and pixel reconstruction, preserving both semantic content and character-level strokes across 17 languages.

Yuliang Liu et al.
Released: 2026-07-11

Figure 1. Document-native pretraining through text generation and pixel reconstruction. Source paper, PDF p. 2. Figure notice.

ℹ️ More Information

MonkeyOCRv2 is a visual-text foundation encoder specialized for the properties that natural-image pretraining often misses: dense text, small glyphs, fine strokes, and layout. Its two complementary objectives align images with their transcribed content while forcing the encoder to retain pixel-level document structure.

The MonkeyDoc v2 corpus contains 113M document images in 17 languages. The frozen encoder improves recognition, detection, tamper analysis, segmentation, parsing, and understanding; paired with a lightweight language model, it forms a 0.7B document parser. Public weights were announced July 11, two days before the arXiv report, so the timeline follows the model release rather than the paper date.

MiniMax M3: Native Multimodality with Sparse Long-Context Attention

MiniMax M3 combines native multimodal pretraining with learned block-sparse grouped-query attention, making million-token contexts practical in a 109B-parameter model.

MiniMax
Released: 2026-06-11

Figure 1. MiniMax Sparse Attention index and exact-attention branches. Source paper, PDF p. 1. Figure notice.

ℹ️ More Information

MiniMax Sparse Attention adds a lightweight index branch that scores KV blocks independently for each grouped-query-attention group. The main branch performs exact attention only over selected blocks, while a KV-outer kernel groups queries retrieving the same block to improve memory locality and reuse.

The released 109B-parameter model is trained natively across modalities and evaluated at contexts up to one million tokens. The sparse-attention report focuses on architecture, kernel design, and scaling behavior; it does not publish a complete source-level training-data inventory.

InternVideo3: Multimodal Contextual Reasoning for Video Agents

InternVideo3 reframes long-video understanding as an evolving loop of observation, reasoning, tools, and memory, while compressing KV state without dropping the underlying token stream.

Ziang Yan et al.
Released: 2026-06-10

Figure 2. InternVideo3 with multimodal multi-head latent attention across long contexts. Source paper, PDF p. 7. Figure notice.

ℹ️ More Information

Multimodal Contextual Reasoning maintains one evolving context containing observations, instructions, intermediate reasoning, tool actions, and memory. This closed loop lets the model accumulate and verify evidence across long videos rather than treating video QA as a single static prompt.

Multimodal Multi-head Latent Attention reparameterizes and compresses KV-cache states while preserving the complete visual-token sequence. Training combines continued pretraining, short-to-long supervised tuning, rule-based reinforcement learning, and on-policy distillation; a retrieval-equipped video-agent implementation demonstrates the intended tool-using setting.

Keye-VL 2.0: Sparse Attention for Long-Video Agents

Keye-VL 2.0 adapts DeepSeek Sparse Attention to a GQA-based multimodal MoE, targeting lossless 256K contexts, hour-scale video, and self-correcting tool-using agents.

Kwai Keye Team
Released: 2026-06-09

Figure 2. Four-stage curriculum extending Keye-VL from alignment to 256K context. Source paper, PDF p. 8. Figure notice.

ℹ️ More Information

Keye-VL-2.0-30B-A3B is the first reported adaptation of DeepSeek Sparse Attention to grouped-query multimodal attention. Only 3B of 30B parameters activate per token, while the sparse retrieval mechanism keeps critical frames and long-range dependencies accessible across a 256K context. Heterogeneous ViT-language-model parallelism and custom sparse-attention kernels address the systems cost of hour-long video.

Cross-Modal Multi-Teacher On-Policy Distillation feeds dense token-level guidance from specialist teachers back into on-policy trajectories. Context-RL and Video-RL then target long-context reasoning, temporal localization, code, search, tools, and multimodal self-correction without collapsing the multi-task mixture.

Zamba2-VL: Hybrid State-Space Vision-Language Modeling

Zamba2-VL combines efficient Mamba-2 state-space layers with a small number of shared Transformer blocks, reducing long-context prefill and recurrent-state costs across compact VLM scales.

Zyphra
Released: 2026-06-02

Figure 1. Zamba2 hybrid state-space language backbone connected to a vision encoder. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

Zamba2-VL extends the Zamba hybrid backbone to vision-language inputs in 1.2B, 2.7B, and 7B variants. Most sequence processing occurs in Mamba-2 state-space layers, with a small set of shared attention blocks supplying global token interaction. This preserves transformer-like multimodal reasoning while moving more computation toward near-linear prefill and bounded recurrent state.

The architecture is particularly relevant for edge and long-context deployment: visual sequences can be large even when the language model is small, so reducing quadratic attention changes time-to-first-token and memory behavior more materially than ordinary parameter pruning.

Cosmos 3: Omnimodal World Modeling with Mixture of Transformers

Cosmos 3 couples an autoregressive reasoner with a diffusion generator through a shared representation for language, vision, audio, actions, simulation, and robot control.

NVIDIA
Released: 2026-05-31

Figure 5. Mixture-of-Transformers reasoner and generator with shared attention. Source paper, PDF p. 11. Figure notice.

ℹ️ More Information

Cosmos 3 uses a Mixture-of-Transformers with two communicating towers. An autoregressive transformer reasons over language, vision, and action sequences; a diffusion transformer generates continuous images, video, and world trajectories. Their shared representation lets the family serve as a VLM, forward or inverse dynamics model, simulator, generator, or policy without separate task pipelines.

Nano configurations combine 8B-class components and Super configurations combine 32B-class components. Training spans text, image, video, audio, simulated physical interaction, action trajectories, robot demonstrations, driving, and synthetic worlds. The July 20 Cosmos 3 Edge release preserves the same two-tower design in a smaller real-time deployment branch and is folded into this family.

Lance: Unified Image and Video Understanding, Generation, and Editing

Lance is a compact model trained from scratch to understand, generate, and edit images and video inside one interleaved sequence while preserving specialized semantic and synthesis capacity.

Lance Team
Released: 2026-05-18

Figure 6. Dual-expert sequence modeling for understanding and visual generation. Source paper, PDF p. 9. Figure notice.

ℹ️ More Information

Lance interleaves text, image, and video in a shared 3B-parameter sequence model while routing semantic understanding and visual synthesis through dedicated experts. Semantic ViT tokens coexist with clean or noised VAE latents, and generalized 3D causal attention plus modality-aware positional encoding preserve spatial and temporal relationships.

One checkpoint supports visual understanding, text-to-image and text-to-video generation, image-to-video generation, and editing. Its training-from-scratch recipe uses no more than 128 A100 GPUs, making the contribution as much about an attainable unified architecture as headline generation quality.

ZAYA1-VL: Vision-Specialized Compressed Convolutional Attention

ZAYA1-VL adds visual capacity to a sparse MoE through vision-only LoRA and compressed-convolutional-attention parameters, trained on an entirely open multimodal data mixture.

Zyphra
Released: 2026-05-08

Figure 2. Visual routing, compressed convolutional attention, and a hybrid language backbone. Source paper, PDF p. 5. Figure notice.

ℹ️ More Information

ZAYA1-VL-8B joins a Qwen2.5-VL visual encoder to ZAYA1's MoE decoder. Vision-only LoRA and Compressed Convolutional Attention parameters activate for visual tokens without duplicating the language backbone, while bidirectional attention within image-token spans improves spatial integration.

The model is trained on roughly 140B vision-language tokens assembled from open data and released under Apache 2.0. Its importance is the clean separation of visual specialization from shared language capacity in a compact sparse model.

Falcon Perception: Early-Fusion Grounding, Segmentation, and OCR

Falcon Perception uses early image-text fusion, hybrid attention, instance tokens, and continuous mask heads to unify open-vocabulary grounding, segmentation, and OCR in compact models.

Technology Innovation Institute
Released: 2026-05-03

Figure 1. Early-fusion perception Transformer with grounding, geometry, and segmentation pathways. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

Falcon Perception is a dense early-fusion Transformer in which image patches and text interact from the first layer. Its attention mask is bidirectional over image tokens and causal over prediction tokens. Variable-length instance tokens identify multiple objects, while lightweight continuous heads produce masks without turning segmentation into an unwieldy text sequence.

The family includes a 600M perception model and a 300M OCR model. The paper was submitted March 28, but the timeline uses the documented May 3 public model launch, following this repository's release-date policy.

GLM-5V-Turbo: Native Multimodal Agency

GLM-5V-Turbo integrates visual perception into reasoning, planning, tool use, execution, and verification instead of treating images and video as an auxiliary language-model interface.

GLM-V Team
Released: 2026-04-29

This HTML preview is truncated for page performance. The canonical Markdown file contains the complete snapshot.

MARKDOWN METRICS
64610words
162headings
555links
1code blocks
MDRSS ASSESSMENT
Evidence46/100high confidence
Why MDRSS assigned this score
  • Production catalog audit 2026-08-04
  • Taxonomy classified from title, annotation, source and Markdown signals
  • Agent usefulness evaluated from structure, procedures, examples, evidence and retrieval value
Evidence (1)

Discussion 0

Sign in to join the discussion.