Awesome-LLM-Inference: A curated list of 📙Awesome LLM Inference Papers with Codes. For Awesome Diffusion Inference, please check 📖Awesome-DiT-Inference . For CUDA learn notes, please check 📖LeetCUDA .
📒Introduction
Snapshot 2026-08-03 23:56:39 UTC · version 1
Research document
📒Introduction
Awesome-LLM-Inference: A curated list of 📙Awesome LLM Inference Papers with Codes. For Awesome Diffusion Inference, please check 📖Awesome-DiT-Inference . For CUDA learn notes, please check 📖LeetCUDA .
📖 News 🔥🔥
- [2026/03] Cache-DiT 🎉v1.3.0 release is ready, the major updates including: Ring Attention w/ batched P2P, USP (Hybrid Ring and Ulysses), Hybrid 2D and 3D Parallelism (💥USP + TP), VAE-P Comm overhead reduce.
©️Citations
@misc{Awesome-LLM-Inference@2024,
title={Awesome-LLM-Inference: A curated list of Awesome LLM Inference Papers with codes},
url={https://github.com/xlite-dev/Awesome-LLM-Inference},
note={Open-source software available at https://github.com/xlite-dev/Awesome-LLM-Inference},
author={xlite-dev, liyucheng09 etc},
year={2024}
}
🎉Awesome LLM Inference Papers with Codes
Awesome LLM Inference for Beginners.pdf: 500 pages, FastServe, FlashAttention 1/2, FlexGen, FP8, LLM.int8(), PagedAttention, RoPE, SmoothQuant, WINT8/4, Continuous Batching, ZeroQuant 1/2/FP, AWQ etc.
🎉Download All PDFs
python3 download_pdfs.py # The code is generated by Doubao AI
📖Contents
- 📖Trending LLM/VLM Topics🔥🔥🔥
- 📖DeepSeek/MLA Topics🔥🔥🔥
- 📖Multi-GPUs/Multi-Nodes Parallelism🔥🔥🔥
- 📖Disaggregating Prefill and Decoding🔥🔥🔥
- 📖LLM Algorithmic/Eval Survey
- 📖LLM Train/Inference Framework/Design
- 📖Weight/Activation Quantize/Compress🔥
- 📖Continuous/In-flight Batching
- 📖IO/FLOPs-Aware/Sparse Attention🔥
- 📖KV Cache Scheduling/Quantize/Dropping🔥
- 📖Prompt/Context Compression🔥
- 📖Long Context Attention/KV Cache Optimization🔥🔥
- 📖Early-Exit/Intermediate Layer Decoding
- 📖Parallel Decoding/Sampling🔥
- 📖Structured Prune/KD/Weight Sparse
- 📖Mixture-of-Experts(MoE) LLM Inference🔥
- 📖CPU/NPU/FPGA/Mobile Inference
- 📖Non Transformer Architecture🔥
- 📖GEMM/Tensor Cores/WMMA/Parallel
- 📖VLM/Position Embed/Others
- 📖LLM Inference Applications
📖Trending LLM/VLM Topics (©️back👆🏻)
| Date | Title | Paper | Code | Recom |
|---|---|---|---|---|
| 2026.03 | 🔥🔥🔥[OneComp] OneComp: One-Line Revolution for Generative AI Model Compression(@Fujitsu) | [pdf] | [OneCompression] | ⭐️⭐️ |
| 2025.12 | 🔥🔥[QEP] QEP: Quantization Error Propagation, NeurIPS 2025(@Fujitsu) | [pdf] | [OneCompression] | ⭐️⭐️ |
| 2024.04 | 🔥🔥🔥[Open-Sora] Open-Sora: Democratizing Efficient Video Production for All(@hpcaitech) | [docs] | [Open-Sora] | ⭐️⭐️ |
| 2024.04 | 🔥🔥🔥[Open-Sora Plan] Open-Sora Plan: This project aim to reproduce Sora (Open AI T2V model)(@PKU) | [report] | [Open-Sora-Plan] | ⭐️⭐️ |
| 2024.05 | 🔥🔥🔥[DeepSeek-V2] DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model(@DeepSeek-AI) | [pdf] | [DeepSeek-V2] | ⭐️⭐️ |
| 2024.05 | 🔥🔥[YOCO] You Only Cache Once: Decoder-Decoder Architectures for Language Models(@Microsoft) | [pdf] | [unilm-YOCO] | ⭐️⭐️ |
| 2024.06 | 🔥[Mooncake] Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving(@Moonshot AI) | [pdf] | [Mooncake] | ⭐️⭐️ |
| 2024.07 | 🔥🔥[FlashAttention-3] FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision(@TriDao etc) | [pdf] | [flash-attention] | ⭐️⭐️ |
| 2024.07 | 🔥🔥[MInference 1.0] MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention(@Microsoft) | [pdf] | [MInference 1.0] | ⭐️⭐️ |
| 2024.11 | 🔥🔥🔥[Star-Attention: 11x~ speedup] Star Attention: Efficient LLM Inference over Long Sequences(@NVIDIA) | [pdf] | [Star-Attention] | ⭐️⭐️ |
| 2024.12 | 🔥🔥🔥[DeepSeek-V3] DeepSeek-V3 Technical Report(@deepseek-ai) | [pdf] | [DeepSeek-V3] | ⭐️⭐️ |
| 2025.01 | 🔥🔥🔥 [MiniMax-Text-01] MiniMax-01: Scaling Foundation Models with Lightning Attention | [report] | [MiniMax-01] | ⭐️⭐️ |
| 2025.01 | 🔥🔥🔥[DeepSeek-R1] DeepSeek-R1 Technical Report(@deepseek-ai) | [pdf] | [DeepSeek-R1] | ⭐️⭐️ |
📖DeepSeek/Multi-head Latent Attention(MLA) (©️back👆🏻)
| Date | Title | Paper | Code | Recom |
|---|---|---|---|---|
| 2024.05 | 🔥🔥🔥[DeepSeek-V2] DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model(@DeepSeek-AI) | [pdf] | [DeepSeek-V2] | ⭐️⭐️ |
| 2024.12 | 🔥🔥🔥[DeepSeek-V3] DeepSeek-V3 Technical Report(@deepseek-ai) | [pdf] | [DeepSeek-V3] | ⭐️⭐️ |
| 2025.01 | 🔥🔥🔥[DeepSeek-R1] DeepSeek-R1 Technical Report(@deepseek-ai) | [pdf] | [DeepSeek-R1] | ⭐️⭐️ |
| 2025.02 | 🔥🔥🔥[DeepSeek-NSA] Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention(@deepseek-ai) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2025.02 | 🔥🔥🔥[FlashMLA] DeepSeek FlashMLA(@deepseek-ai) | ⚠️ | [FlashMLA] | ⭐️⭐️ |
| 2025.02 | 🔥🔥🔥[DualPipe] DeepSeek DualPipe(@deepseek-ai) | ⚠️ | [DualPipe] | ⭐️⭐️ |
| 2025.02 | 🔥🔥🔥[DeepEP] DeepSeek DeepEP(@deepseek-ai) | ⚠️ | [DeepEP] | ⭐️⭐️ |
| 2025.02 | 🔥🔥🔥[DeepGEMM] DeepSeek DeepGEMM(@deepseek-ai) | ⚠️ | [DeepGEMM] | ⭐️⭐️ |
| 2025.02 | 🔥🔥🔥[EPLB] DeepSeek EPLB(@deepseek-ai) | ⚠️ | [EPLB] | ⭐️⭐️ |
| 2025.02 | 🔥🔥🔥[3FS] DeepSeek 3FS(@deepseek-ai) | ⚠️ | [3FS] | ⭐️⭐️ |
| 2025.03 | 🔥🔥🔥[推理系统] DeepSeek-V3 / R1 推理系统概览 (@deepseek-ai) | [blog] | ⚠️ | ⭐️⭐️ |
| 2025.02 | 🔥🔥[MHA2MLA] Towards Economical Inference: Enabling DeepSeek’s Multi-Head Latent Attention in Any Transformer-based LLMs(@fudan.edu.cn) | [pdf] | [MHA2MLA] | ⭐️⭐️ |
| 2025.02 | 🔥🔥[TransMLA] TransMLA: Multi-head Latent Attention Is All You Need(@PKU) | [pdf] | [TransMLA] | ⭐️⭐️ |
| 2025.03 | 🔥🔥[X-EcoMLA] X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression(@AMD) | [pdf] | ⚠️ | ⭐️⭐️ |
📖Multi-GPUs/Multi-Nodes Parallelism (©️back👆🏻)
| Date | Title | Paper | Code | Recom |
|---|---|---|---|---|
| 2019.10 | 🔥🔥[MP: ZeRO] DeepSpeed-ZeRO: Memory Optimizations Toward Training Trillion Parameter Models(@microsoft.com) | [pdf] | [deepspeed] | ⭐️⭐️ |
| 2020.05 | 🔥🔥[TP: Megatron-LM] Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism(@NVIDIA) | [pdf] | [Megatron-LM] | ⭐️⭐️ |
| 2022.05 | 🔥🔥[SP: Megatron-LM] Megatron-LM: Reducing Activation Recomputation in Large Transformer Models(@NVIDIA) | [pdf] | [Megatron-LM] | ⭐️⭐️ |
| 2023.05 | 🔥🔥[SP: BPT] Blockwise Parallel Transformer for Large Context Models(@UC Berkeley) | [pdf] | [RingAttention] | ⭐️⭐️ |
| 2023.10 | 🔥🔥[SP: Ring Attention] Ring Attention with Blockwise Transformers for Near-Infinite Context(@UC Berkeley) | [pdf] | [RingAttention] | ⭐️⭐️ |
| 2023.11 | 🔥🔥[SP: STRIPED ATTENTION] STRIPED ATTENTION: FASTER RING ATTENTION FOR CAUSAL TRANSFORMERS(@MIT etc) | [pdf] | [striped_attention] | ⭐️⭐️ |
| 2023.10 | 🔥🔥[SP: DEEPSPEED ULYSSES] DEEPSPEED ULYSSES: SYSTEM OPTIMIZATIONS FOR ENABLING TRAINING OF EXTREME LONG SEQUENCE TRANSFORMER MODELS(@microsoft.com) | [pdf] | [deepspeed] | ⭐️⭐️ |
| 2024.03 | 🔥🔥[CP: Megatron-LM] Megatron-LM: Context parallelism overview(@NVIDIA) | [docs] | [Megatron-LM] | ⭐️⭐️ |
| 2024.05 | 🔥🔥[SP: Unified Sequence Parallel (USP)] YunChang: A Unified Sequence Parallel (USP) Attention for Long Context LLM Model Training and Inference(@Tencent) | [pdf] | [long-context-attention] | ⭐️⭐️ |
| 2024.11 | 🔥🔥[CP: Meta] Context Parallelism for Scalable Million-Token Inference(@Meta Platforms, Inc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.11 | 🔥🔥[TP: Comm Compression] Communication Compression for Tensor Parallel LLM Inference(@recogni.com) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.11 | 🔥🔥🔥[SP: Star-Attention, 11x~ speedup] Star Attention: Efficient LLM Inference over Long Sequences(@NVIDIA) | [pdf] | [Star-Attention] | ⭐️⭐️ |
| 2024.12 | 🔥🔥[SP: TokenRing] TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication(@SJTU) | [pdf] | [token-ring] | ⭐️⭐️ |
| 2025.05 | 🔥🔥[FSDP 1/2] PyTorch FSDP: Getting Started with Fully Sharded Data Parallel(FSDP) (@pytorch) | [docs] | ⚠️ | ⭐️⭐️ |
📖Disaggregating Prefill and Decoding (©️back👆🏻)
| Date | Title | Paper | Code | Recom |
|---|---|---|---|---|
| 2024.01 | 🔥🔥[DistServe] DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving(@PKU) | [pdf] | [DistServe] | ⭐️⭐️ |
| 2024.06 | 🔥🔥[Mooncake] Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving(@Moonshot AI) | [pdf] | [Mooncake] | ⭐️⭐️ |
| 2024.12 | 🔥🔥[KVDirect] KVDirect: Distributed Disaggregated LLM Inference(@ByteDance) | [pdf] | ⚠️ | ⭐️ |
| 2025.01 | 🔥🔥[DeServe] DESERVE: TOWARDS AFFORDABLE OFFLINE LLM INFERENCE VIA DECENTRALIZATION(@Berkeley) | [pdf] | ⚠️ | ⭐️ |
| 2025.04 | 🔥🔥[MegaScale-Infer] MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism(@ByteDance Seed) | [pdf] | ⚠️ | ⭐️ |
📖LLM Algorithmic/Eval Survey (©️back👆🏻)
| Date | Title | Paper | Code | Recom |
|---|---|---|---|---|
| 2023.10 | [Evaluating] Evaluating Large Language Models: A Comprehensive Survey(@tju.edu.cn) | [pdf] | [Awesome-LLMs-Evaluation] | ⭐️ |
| 2023.11 | 🔥[Runtime Performance] Dissecting the Runtime Performance of the Training, Fine-tuning, and Inference of Large Language Models(@hkust-gz.edu.cn) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2023.11 | [ChatGPT Anniversary] ChatGPT’s One-year Anniversary: Are Open-Source Large Language Models Catching up?(@e.ntu.edu.sg) | [pdf] | ⚠️ | ⭐️ |
| 2023.12 | [Algorithmic Survey] The Efficiency Spectrum of Large Language Models: An Algorithmic Survey(@Microsoft) | [pdf] | ⚠️ | ⭐️ |
| 2023.12 | [Security and Privacy] A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly(@Drexel University) | [pdf] | ⚠️ | ⭐️ |
| 2023.12 | 🔥[LLMCompass] A Hardware Evaluation Framework for Large Language Model Inference(@princeton.edu) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2023.12 | 🔥[Efficient LLMs] Efficient Large Language Models: A Survey(@Ohio State University etc) | [pdf] | [Efficient-LLMs-Survey] | ⭐️⭐️ |
| 2023.12 | [Serving Survey] Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems(@Carnegie Mellon University) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.01 | [Understanding LLMs] Understanding LLMs: A Comprehensive Overview from Training to Inference(@Shaanxi Normal University etc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.02 | [LLM-Viewer] LLM Inference Unveiled: Survey and Roofline Model Insights(@Zhihang Yuan etc) | [pdf] | [LLM-Viewer] | ⭐️⭐️ |
| 2024.07 | [Internal Consistency & Self-Feedback] Internal Consistency and Self-Feedback in Large Language Models: A Survey | [pdf] | [ICSF-Survey] | ⭐️⭐️ |
| 2024.09 | [Low-bit] A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms(@Beihang etc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.10 | [LLM Inference] LARGE LANGUAGE MODEL INFERENCE ACCELERATION: A COMPREHENSIVE HARDWARE PERSPECTIVE(@SJTU etc) | [pdf] | ⚠️ | ⭐️⭐️ |
📖LLM Train/Inference Framework/Design (©️back👆🏻)
| Date | Title | Paper | Code | Recom |
|---|---|---|---|---|
| 2020.05 | 🔥[Megatron-LM] Training Multi-Billion Parameter Language Models Using Model Parallelism(@NVIDIA) | [pdf] | [Megatron-LM] | ⭐️⭐️ |
| 2023.03 | [FlexGen] High-Throughput Generative Inference of Large Language Models with a Single GPU(@Stanford University etc) | [pdf] | [FlexGen] | ⭐️ |
| 2023.05 | [SpecInfer] Accelerating Generative Large Language Model Serving with Speculative Inference and Token Tree Verification(@Peking University etc) | [pdf] | [FlexFlow] | ⭐️ |
| 2023.05 | [FastServe] Fast Distributed Inference Serving for Large Language Models(@Peking University etc) | [pdf] | ⚠️ | ⭐️ |
| 2023.09 | 🔥[vLLM] Efficient Memory Management for Large Language Model Serving with PagedAttention(@UC Berkeley etc) | [pdf] | [vllm] | ⭐️⭐️ |
| 2023.09 | [StreamingLLM] EFFICIENT STREAMING LANGUAGE MODELS WITH ATTENTION SINKS(@Meta AI etc) | [pdf] | [streaming-llm] | ⭐️ |
| 2023.09 | [Medusa] Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads(@Tianle Cai etc) | [blog] | [Medusa] | ⭐️ |
| 2023.10 | 🔥[TensorRT-LLM] NVIDIA TensorRT LLM(@NVIDIA) | [docs] | [TensorRT-LLM] | ⭐️⭐️ |
| 2023.11 | 🔥[DeepSpeed-FastGen 2x vLLM?] DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference(@Microsoft) | [pdf] | [deepspeed-fastgen] | ⭐️⭐️ |
| 2023.12 | 🔥🔥[SGLang] Efficiently Programming Large Language Models using SGLang(@Stanford University etc) | [pdf] | [sglang] | ⭐️⭐️ |
| 2023.12 | 🔥[PETALS] Distributed Inference and Fine-tuning of Large Language Models Over The Internet(@HSE Univesity etc) | [pdf] | [petals] | ⭐️⭐️ |
| 2023.10 | [LightSeq] LightSeq: Sequence Level Parallelism for Distributed Training of Long Context Transformers(@UC Berkeley etc) | [pdf] | [LightSeq] | ⭐️ |
| 2023.12 | [PowerInfer] PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU(@SJTU) | [pdf] | [PowerInfer] | ⭐️ |
| 2024.01 | [inferflow]INFERFLOW: AN EFFICIENT AND HIGHLY CONFIGURABLE INFERENCE ENGINE FOR LARGE LANGUAGE MODELS(@Tencent AI Lab) | [pdf] | [inferflow] | ⭐️ |
| 2024.06 | 🔥[Mooncake] Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving(@Moonshot AI) | [pdf] | [Mooncake] | ⭐️⭐️ |
| 2023.06 | 🔥[LMDeploy] LMDeploy: LMDeploy is a toolkit for compressing, deploying, and serving LLMs(@InternLM) | [docs] | [lmdeploy] | ⭐️⭐️ |
| 2023.05 | 🔥[MLC-LLM]Universal LLM Deployment Engine with ML Compilation(@mlc-ai) | [docs] | [mlc-llm] | ⭐️⭐️ |
| 2023.08 | 🔥[LightLLM] LightLLM is a Python-based LLM (Large Language Model) inference and serving framework(@ModelTC) | [docs] | [lightllm] | ⭐️⭐️ |
| 2023.03 | 🔥[llama.cpp] llama.cpp: Inference of Meta's LLaMA model (and others) in pure C/C++(@ggerganov) | [docs] | [llama.cpp] | ⭐️⭐️ |
| 2024.02 | 🔥[flashinfer] FlashInfer: Kernel Library for LLM Serving(@flashinfer-ai) | [docs] | [flashinfer] | ⭐️⭐️ |
| 2024.06 | 🔥[Mooncake] Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving(@Moonshot AI) | [pdf] | [Mooncake] | ⭐️⭐️ |
| 2024.07 | 🔥[DynamoLLM] DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency(@Microsoft Azure Research) | [pdf] | ⚠️ | ⭐️ |
| 2024.08 | 🔥[NanoFlow] NanoFlow: Towards Optimal Large Language Model Serving Throughput(@University of Washington) | [pdf] | [Nanoflow] | ⭐️⭐️ |
| 2024.08 | 🔥[Decentralized LLM] Decentralized LLM Inference over Edge Networks with Energy Harvesting(@Padova) | [pdf] | ⚠️ | ⭐️ |
| 2024.11 | 🔥[SparseInfer] SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference(@University of Seoul, etc) | [pdf] | ⚠️ | ⭐️ |
| 2025.04 | 🔥[prima.cpp] PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters(@MBZUAI, etc) | [pdf] | [prima.cpp] | ⭐️ |
| 2025.07 | 🔥[siiRL] DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training(@Shanghai Inovation Institute) | [pdf] | [siiRL] |
⭐️⭐️ |
| 2025.04 | 🔥[ToolPipe] ToolPipe: 120+ Free Developer Tools REST API & MCP Server for AI Agents(@COSAI-Labs) | [docs] | [toolpipe-mcp-server] | ⭐️ |
📖Continuous/In-flight Batching (©️back👆🏻)
| Date | Title | Paper | Code | Recom |
|---|---|---|---|---|
| 2022.07 | 🔥[Continuous Batching] Orca: A Distributed Serving System for Transformer-Based Generative Models(@Seoul National University etc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2023.10 | 🔥[In-flight Batching] NVIDIA TensorRT LLM Batch Manager(@NVIDIA) | [docs] | [TensorRT-LLM] | ⭐️⭐️ |
| 2023.11 | 🔥[DeepSpeed-FastGen 2x vLLM?] DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference(@Microsoft) | [blog] | [deepspeed-fastgen] | ⭐️⭐️ |
| 2023.11 | [Splitwise] Splitwise: Efficient Generative LLM Inference Using Phase Splitting(@Microsoft etc) | [pdf] | ⚠️ | ⭐️ |
| 2023.12 | [SpotServe] SpotServe: Serving Generative Large Language Models on Preemptible Instances(@cmu.edu etc) | [pdf] | [SpotServe] | ⭐️ |
| 2023.10 | [LightSeq] LightSeq: Sequence Level Parallelism for Distributed Training of Long Context Transformers(@UC Berkeley etc) | [pdf] | [LightSeq] | ⭐️ |
| 2024.05 | 🔥[vAttention] vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention(@Microsoft Research India) | [pdf] | [vAttention] | ⭐️⭐️ |
| 2024.07 | 🔥🔥[vTensor] vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving(@Shanghai Jiao Tong University etc) | [pdf] | [vTensor] | ⭐️⭐️ |
| 2024.08 | 🔥[Automatic Inference Engine Tuning] Towards SLO-Optimized LLM Serving via Automatic Inference Engine Tuning(@Nanjing University etc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.08 | 🔥[SJF Scheduling] Efficient LLM Scheduling by Learning to Rank(@UCSD etc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.12 | 🔥[BatchLLM] BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching(@Microsoft) | [pdf] | ⚠️ | ⭐️⭐️ |
📖Weight/Activation Quantize/Compress (©️back👆🏻)
| Date | Title | Paper | Code | Recom |
|---|---|---|---|---|
| 2022.06 | 🔥[ZeroQuant] Efficient and Affordable Post-Training Quantization for Large-Scale Transformers(@Microsoft) | [pdf] | [DeepSpeed] | ⭐️⭐️ |
| 2022.08 | [FP8-Quantization] FP8 Quantization: The Power of the Exponent(@Qualcomm AI Research) | [pdf] | [FP8-quantization] | ⭐️ |
| 2022.08 | [LLM.int8()] 8-bit Matrix Multiplication for Transformers at Scale(@Facebook AI Research etc) | [pdf] | [bitsandbytes] | ⭐️ |
| 2022.10 | 🔥[GPTQ] GPTQ: ACCURATE POST-TRAINING QUANTIZATION FOR GENERATIVE PRE-TRAINED TRANSFORMERS(@IST Austria etc) | [pdf] | [gptq] | ⭐️⭐️ |
| 2022.11 | 🔥[WINT8/4] Who Says Elephants Can’t Run: Bringing Large Scale MoE Models into Cloud Scale Production(@NVIDIA&Microsoft) | [pdf] | [FasterTransformer] | ⭐️⭐️ |
| 2022.11 | 🔥[SmoothQuant] Accurate and Efficient Post-Training Quantization for Large Language Models(@MIT etc) | [pdf] | [smoothquant] | ⭐️⭐️ |
| 2023.03 | [ZeroQuant-V2] Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation(@Microsoft) | [pdf] | [DeepSpeed] | ⭐️ |
| 2023.06 | 🔥[AWQ] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration(@MIT etc) | [pdf] | [llm-awq] | ⭐️⭐️ |
| 2023.06 | [SpQR] SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression(@University of Washington etc) | [pdf] | [SpQR] | ⭐️ |
| 2023.06 | [SqueezeLLM] SQUEEZELLM: DENSE-AND-SPARSE QUANTIZATION(@berkeley.edu) | [pdf] | [SqueezeLLM] | ⭐️ |
| 2023.07 | [ZeroQuant-FP] A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats(@Microsoft) | [pdf] | [DeepSpeed] | ⭐️ |
| 2023.09 | [KV Cache FP8 + WINT4] Exploration on LLM inference performance optimization(@HPC4AI) | [blog] | ⚠️ | ⭐️ |
| 2023.10 | [FP8-LM] FP8-LM: Training FP8 Large Language Models(@Microsoft etc) | [pdf] | [MS-AMP] | ⭐️ |
| 2023.10 | [LLM-Shearing] SHEARED LLAMA: ACCELERATING LANGUAGE MODEL PRE-TRAINING VIA STRUCTURED PRUNING(@cs.princeton.edu etc) | [pdf] | [LLM-Shearing] | ⭐️ |
| 2023.10 | [LLM-FP4] LLM-FP4: 4-Bit Floating-Point Quantized Transformers(@ust.hk&meta etc) | [pdf] | [LLM-FP4] | ⭐️ |
| 2023.11 | [2-bit LLM] Enabling Fast 2-bit LLM on GPUs: Memory Alignment, Sparse Outlier, and Asynchronous Dequantization(@Shanghai Jiao Tong University etc) | [pdf] | ⚠️ | ⭐️ |
| 2023.12 | [SmoothQuant+] SmoothQuant+: Accurate and Efficient 4-bit Post-Training Weight Quantization for LLM(@ZTE Corporation) | [pdf] | [smoothquantplus] | ⭐️ |
| 2023.11 | [OdysseyLLM W4A8] A Speed Odyssey for Deployable Quantization of LLMs(@meituan.com) | [pdf] | ⚠️ | ⭐️ |
| 2023.12 | 🔥[SparQ] SPARQ ATTENTION: BANDWIDTH-EFFICIENT LLM INFERENCE(@graphcore.ai) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2023.12 | [Agile-Quant] Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge(@Northeastern University&Oracle) | [pdf] | ⚠️ | ⭐️ |
| 2023.12 | [CBQ] CBQ: Cross-Block Quantization for Large Language Models(@ustc.edu.cn) | [pdf] | ⚠️ | ⭐️ |
| 2023.10 | [QLLM] QLLM: ACCURATE AND EFFICIENT LOW-BITWIDTH QUANTIZATION FOR LARGE LANGUAGE MODELS(@ZIP Lab&SenseTime Research etc) | [pdf] | ⚠️ | ⭐️ |
| 2024.01 | [FP6-LLM] FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design(@Microsoft etc) | [pdf] | ⚠️ | ⭐️ |
| 2024.05 | 🔥🔥[W4A8KV4] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving(@MIT&NVIDIA) | [pdf] | [qserve] | ⭐️⭐️ |
| 2024.05 | 🔥[SpinQuant] SpinQuant: LLM Quantization with Learned Rotations(@Meta) | [pdf] | ⚠️ | ⭐️ |
| 2024.05 | 🔥[I-LLM] I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models(@Houmo AI) | [pdf] | ⚠️ | ⭐️ |
| 2024.06 | 🔥[OutlierTune] OutlierTune: Efficient Channel-Wise Quantization for Large Language Models(@Beijing University) | [pdf] | ⚠️ | ⭐️ |
| 2024.06 | 🔥[GPTQT] GPTQT: Quantize Large Language Models Twice to Push the Efficiency(@zju) | [pdf] | ⚠️ | ⭐️ |
| 2024.08 | 🔥[ABQ-LLM] ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models(@ByteDance) | [pdf] | [ABQ-LLM] | ⭐️ |
| 2024.08 | 🔥[1-bit LLMs] Matmul or No Matmal in the Era of 1-bit LLMs(@University of South Carolina) | [pdf] | ⚠️ | ⭐️ |
| 2024.08 | 🔥[ACTIVATION SPARSITY] TRAINING-FREE ACTIVATION SPARSITY IN LARGE LANGUAGE MODELS(@MIT etc) | [pdf] | [TEAL] | ⭐️ |
| 2024.09 | 🔥[VPTQ] VPTQ: EXTREME LOW-BIT VECTOR POST-TRAINING QUANTIZATION FOR LARGE LANGUAGE MODELS(@Microsoft) | [pdf] | [VPTQ] | ⭐️ |
| 2024.11 | 🔥[BitNet] BitNet a4.8: 4-bit Activations for 1-bit LLMs(@Microsoft) | [pdf] | [bitnet] | ⭐️ |
| 2025.04 | 🔥[BitNet v2] BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs(@Microsoft) | [pdf] | [bitnet] | ⭐️ |
| 2025.05 | 🔥[GuidedQuant] GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance (@SNU&SamsungAILab&Google) | [pdf] | [GuidedQuant] | ⭐️⭐️ |
📖IO/FLOPs-Aware/Sparse Attention (©️back👆🏻)
| Date | Title | Paper | Code | Recom |
|---|---|---|---|---|
| 2018.05 | [Online Softmax] Online normalizer calculation for softmax(@NVIDIA) | [pdf] | ⚠️ | ⭐️ |
| 2019.11 | 🔥[MQA] Fast Transformer Decoding: One Write-Head is All You Need(@Google) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2020.10 | [Hash Attention] REFORMER: THE EFFICIENT TRANSFORMER(@Google) | [pdf] | [reformer] | ⭐️⭐️ |
| 2022.05 | 🔥[FlashAttention] Fast and Memory-Efficient Exact Attention with IO-Awareness(@Stanford University etc) | [pdf] | [flash-attention] | ⭐️⭐️ |
| 2022.10 | [Online Softmax] SELF-ATTENTION DOES NOT NEED O(n^2) MEMORY(@Google) | [pdf] | ⚠️ | ⭐️ |
| 2023.05 | [FlashAttention] From Online Softmax to FlashAttention(@cs.washington.edu) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2023.05 | [FLOP, I/O] Dissecting Batching Effects in GPT Inference(@Lequn Chen) | [blog] | ⚠️ | ⭐️ |
| 2023.05 | 🔥🔥[GQA] GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints(@Google) | [pdf] | [flaxformer] | ⭐️⭐️ |
| 2023.06 | [Sparse FlashAttention] Faster Causal Attention Over Large Sequences Through Sparse Flash Attention(@EPFL etc) | [pdf] | [dynamic-sparse-flash-attention] | ⭐️ |
| 2023.07 | 🔥[FlashAttention-2] Faster Attention with Better Parallelism and Work Partitioning(@Stanford University etc) | [pdf] | [flash-attention] | ⭐️⭐️ |
| 2023.10 | 🔥[Flash-Decoding] Flash-Decoding for long-context inference(@Stanford University etc) | [blog] | [flash-attention] | ⭐️⭐️ |
| 2023.11 | [Flash-Decoding++] FLASHDECODING++: FASTER LARGE LANGUAGE MODEL INFERENCE ON GPUS(@Tsinghua University&Infinigence-AI) | [pdf] | ⚠️ | ⭐️ |
| 2023.01 | [SparseGPT] SparseGPT: Massive Language Models Can be Accurately Pruned in One-Shot(@ISTA etc) | [pdf] | [sparsegpt] | ⭐️ |
| 2023.12 | 🔥[GLA] Gated Linear Attention Transformers with Hardware-Efficient Training(@MIT-IBM Watson AI) | [pdf] | gated_linear_attention | ⭐️⭐️ |
| 2023.12 | [SCCA] SCCA: Shifted Cross Chunk Attention for long contextual semantic expansion(@Beihang University) | [pdf] | ⚠️ | ⭐️ |
| 2023.12 | 🔥[FlashLLM] LLM in a flash: Efficient Large Language Model Inference with Limited Memory(@Apple) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.03 | 🔥🔥[CHAI] CHAI: Clustered Head Attention for Efficient LLM Inference(@cs.wisc.edu etc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.04 | 🔥🔥[DeFT] DeFT: Decoding with Flash Tree-Attention for Efficient Tree-structured LLM Inference(@Westlake University etc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.04 | [MoA] MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression(@thu et el.) | [pdf] | [MoA] | ⭐️ |
| 2024.07 | 🔥🔥[FlashAttention-3] FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision(@TriDao etc) | [pdf] | [flash-attention] | ⭐️⭐️ |
| 2024.07 | 🔥🔥[MInference 1.0] MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention(@Microsoft) | [pdf] | [MInference 1.0] | ⭐️⭐️ |
| 2024.07 | 🔥🔥[Shared Attention] Beyond KV Caching: Shared Attention for Efficient LLMs(@Kyushu University etc) | [pdf] | [shareAtt] | ⭐️ |
| 2024.09 | 🔥🔥[CHESS] CHESS : Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification(@Wuhan University) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.09 | 🔥🔥[INT-FLASHATTENTION] INT-FLASHATTENTION: ENABLING FLASH ATTENTION FOR INT8 QUANTIZATION(@PKU etc) | [pdf] | [INT-FlashAttention] | ⭐️ |
| 2024.10 | 🔥🔥[SageAttention] SAGEATTENTION: ACCURATE 8-BIT ATTENTION FOR PLUG-AND-PLAY INFERENCE ACCELERATION(@thu-ml) | [pdf] | [SageAttention] | ⭐️⭐️ |
| 2024.11 | 🔥🔥[SageAttention-2] SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization(@thu-ml) | [pdf] | [SageAttention] | ⭐️⭐️ |
| 2024.11 | 🔥🔥[Squeezed Attention] SQUEEZED ATTENTION: Accelerating Long Context Length LLM Inference(@UC Berkeley) | [pdf] | [SqueezedAttention] | ⭐️⭐️ |
| 2024.12 | 🔥🔥[TurboAttention] TURBOATTENTION: EFFICIENT ATTENTION APPROXIMATION FOR HIGH THROUGHPUTS LLMS(@Microsoft) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2025.01 | 🔥🔥[FFPA] FFPA: Yet another Faster Flash Prefill Attention with O(1) SRAM complexity for headdim > 256, ~1.5x faster than SDPA EA(@xlite-dev) | [docs] | [ffpa-attn] | ⭐️⭐️ |
| 2025.03 | 🔥🔥[SpargeAttention] SpargeAttn: Accurate Sparse Attention Accelerating Any Model Inference(@thu-ml) | [pdf] | [SpargeAttn] | ⭐️⭐️ |
| 2025.04 | 🔥🔥[MMInference] MMInference: Accelerating Pre-filling for Long-Context Visual Language Models via Modality-Aware Permutation Sparse Attention(@microsoft) | [pdf] | [MInference] | ⭐️⭐️ |
| 2025.04 | 🔥🔥[Sparse Frontier] The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs (@Cohere) | [pdf] | [SparseFrontier] | ⭐️⭐️ |
| 2024.12 | 🔥🔥[Flex Attention] FLEX ATTENTION: A PROGRAMMING MODEL FOR GENERATING OPTIMIZED ATTENTION KERNELS(@pytorch) | [pdf] | [attention-gym] | ⭐️⭐️ |
| 2025.02 | 🔥🔥🔥[SeerAttention] SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs(@microsoft) | [pdf] | [SeerAttention] | ⭐️⭐️⭐️ |
| 2025.03 | [Slim attention] Slim attention: cut your context memory in half without loss of accuracy, K-cache is all you need for MHA(@OpenMachine.ai) | [pdf] | [OpenMchine] | ⭐️⭐️⭐️ |
| 2025.05 | 🔥🔥[SageAttention-3] SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-bit Training(@thu-ml) | [pdf] | [SageAttention] | ⭐️⭐️ |
| 2025.04 | 🔥🔥[Parallel Encoding] APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding(@cmu.edu&NVIDIA) | [pdf] | [APE] | ⭐️⭐️ |
| 2025.04 | 🔥🔥[Parallel Encoding] Block-Attention for Efficient Prefilling(@Tencent etc) | [pdf] | [Block-attention] | ⭐️⭐️ |
📖KV Cache Scheduling/Quantize/Dropping (©️back👆🏻)
| Date | Title | Paper | Code | Recom |
|---|---|---|---|---|
| 2026.03 | 🔥🔥[NexusQuant] NexusQuant: Training-Free KV Cache Compression via E8 Lattice Quantization and Attention-Aware Token Eviction — 6.1x compression, +0.276% PPL on Mistral-7B, drop-in one-liner | [code] | [jagmarques] | ⭐️⭐️ |
| 2019.11 | 🔥[MQA] Fast Transformer Decoding: One Write-Head is All You Need(@Google) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2022.06 | [LTP] Learned Token Pruning for Transformers(@UC Berkeley etc) | [pdf] | [LTP] | ⭐️ |
| 2023.05 | 🔥🔥[GQA] GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints(@Google) | [pdf] | [flaxformer] | ⭐️⭐️ |
| 2023.05 | [KV Cache Compress] Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time(@) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2023.06 | [H2O] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models(@Rice University etc) | [pdf] | [H2O] | ⭐️ |
| 2023.06 | [QK-Sparse/Dropping Attention] Faster Causal Attention Over Large Sequences Through Sparse Flash Attention(@EPFL etc) | [pdf] | [dynamic-sparse-flash-attention] | ⭐️ |
| 2023.08 | 🔥🔥[Chunked Prefills] SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills(@Microsoft etc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2023.09 | 🔥🔥[PagedAttention] Efficient Memory Management for Large Language Model Serving with PagedAttention(@UC Berkeley etc) | [pdf] | [vllm] | ⭐️⭐️ |
| 2023.09 | [KV Cache FP8 + WINT4] Exploration on LLM inference performance optimization(@HPC4AI) | [blog] | ⚠️ | ⭐️ |
| 2023.10 | 🔥[TensorRT-LLM KV Cache FP8] NVIDIA TensorRT LLM(@NVIDIA) | [docs] | [TensorRT-LLM] | ⭐️⭐️ |
| 2023.10 | 🔥[Adaptive KV Cache Compress] MODEL TELLS YOU WHAT TO DISCARD: ADAPTIVE KV CACHE COMPRESSION FOR LLMS(@illinois.eduµsoft) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2023.10 | [CacheGen] CacheGen: Fast Context Loading for Language Model Applications(@Chicago University&Microsoft) | [pdf] | [LMCache] | ⭐️ |
| 2023.12 | [KV-Cache Optimizations] Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO(@Haim Barad etc) | [pdf] | ⚠️ | ⭐️ |
| 2023.12 | [KV Cache Compress with LoRA] Compressed Context Memory for Online Language Model Interaction (@SNU & NAVER AI) | [pdf] | [Compressed-Context-Memory] | ⭐️⭐️ |
| 2023.12 | 🔥🔥[RadixAttention] Efficiently Programming Large Language Models using SGLang(@Stanford University etc) | [pdf] | [sglang] | ⭐️⭐️ |
| 2024.01 | 🔥🔥[DistKV-LLM] Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache(@Alibaba etc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.02 | 🔥🔥[Prompt Caching] Efficient Prompt Caching via Embedding Similarity(@UC Berkeley) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.02 | 🔥🔥[Less] Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference(@CMU etc) | [pdf] | ⚠️ | ⭐️ |
| 2024.02 | 🔥🔥[MiKV] No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization(@KAIST) | [pdf] | ⚠️ | ⭐️ |
| 2024.02 | 🔥🔥[Shared Prefixes] Hydragen: High-Throughput LLM Inference with Shared Prefixes | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.02 | 🔥🔥[ChunkAttention] ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition(@microsoft.com) | [pdf] | [chunk-attention] | ⭐️⭐️ |
| 2024.03 | 🔥[QAQ] QAQ: Quality Adaptive Quantization for LLM KV Cache(@@smail.nju.edu.cn) | [pdf] | [QAQ-KVCacheQuantization] | ⭐️⭐️ |
| 2024.03 | 🔥🔥[DMC] Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference(@NVIDIA etc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.03 | 🔥🔥[Keyformer] Keyformer: KV Cache reduction through key tokens selection for Efficient Generative Inference(@ece.ubc.ca etc) | [pdf] | [Keyformer] | ⭐️⭐️ |
| 2024.03 | [FASTDECODE] FASTDECODE: High-Throughput GPU-Efficient LLM Serving using Heterogeneous(@Tsinghua University) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.03 | [Sparsity-Aware KV Caching] ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching(@ucf.edu) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.03 | 🔥[GEAR] GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM(@gatech.edu) | [pdf] | [GEAR] | ⭐️ |
| 2024.04 | [SqueezeAttention] SQUEEZEATTENTION: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget(@lzu.edu.cn etc) | [pdf] | [SqueezeAttention] | ⭐️⭐️ |
| 2024.04 | [SnapKV] SnapKV: LLM Knows What You are Looking for Before Generation(@UIUC) | [pdf] | [SnapKV] | ⭐️ |
| 2024.05 | 🔥[vAttention] vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention(@Microsoft Research India) | [pdf] | [vAttention] | ⭐️⭐️ |
| 2024.05 | 🔥[KVCache-1Bit] KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization(@Rice University) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.05 | 🔥[KV-Runahead] KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation(@Apple etc) | [pdf] | ⚠️ | ⭐️⭐️ |
| 2024.05 | 🔥[ZipCache] ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification(@Zhejiang University etc) | [pdf] | ⚠️ | ⭐️⭐️ |
This HTML preview is truncated for page performance. The canonical Markdown file contains the complete snapshot.
Why MDRSS assigned this score
- Production catalog audit 2026-08-04
- Taxonomy classified from title, annotation, source and Markdown signals
- Agent usefulness evaluated from structure, procedures, examples, evidence and retrieval value
Discussion 0
Sign in to join the discussion.