# #inference-and-quantization — MDRSS hashtag feed

> Public MDRSS cards tagged #inference-and-quantization.
> Canonical feed: https://mdrss.com/feeds/inference-and-quantization

## Cards (9)

### [G1: CUDA 12/13 ABI Mismatch](https://mdrss.com/llm-engineering/inference-and-quantization/901285/901285.md)

Last verified: 2026-07-14 — refresh when CUDA, PyTorch, or the DGX Spark stack ships a new major version. Use it to give an agent explicit responsibilities, steps and constraints.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Spark Training Gotchas](https://mdrss.com/llm-engineering/inference-and-quantization/901284/901284.md)

DGX Spark's GB10 chip (Grace Blackwell, SM121, 128GB unified memory, aarch64) has ten recurring failure modes across launch, memory, thermals, bandwidth, and precision. Each is named G1–G10 so it can be checked by number — the numbering is load-bearing for tooling that runs these. Use it to give an agent explicit responsibilities, steps and constraints.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Spark Environment Setup](https://mdrss.com/llm-engineering/inference-and-quantization/901281/901281.md)

DGX Spark ships a GB10 Grace Blackwell chip: aarch64 CPU, SM121 GPU, 128GB unified memory, CUDA 13. This is a narrower and younger platform than a standard x86 CUDA 12 box, so package selection and ABI matching matter more than usual — the wheel ecosystem for aarch64 + CUDA 13 is. Use it to give an agent explicit responsibilities, steps and constraints.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Transfer Engine (TE)](https://mdrss.com/llm-engineering/inference-and-quantization/901084/901084.md)

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. Under real workloads, Mooncake’s innovative architecture enables Kimi to handle 75% more requests while adhering to SLOs. Use it to ground design choices in named patterns, trade-offs and examples.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [LLMTools: Run & Finetune LLMs on Consumer GPUs](https://mdrss.com/llm-engineering/inference-and-quantization/901083/901083.md)

LLMTools is a user-friendly library for running and finetuning LLMs in low-resource settings. Features include: 🔨 LLM finetuning in 2-bit, 3-bit, 4-bit precision using the ModuLoRA algorithm 🐍 Easy-to-use Python API for quantization, inference, and finetuning 🤖 Modular supp. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [TRL - Transformers Reinforcement Learning](https://mdrss.com/llm-engineering/inference-and-quantization/901071/901071.md)

TRL is a cutting-edge library designed for post-training foundation models using advanced techniques like Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO), and Direct Preference Optimization (DPO). Built on top of the 🤗 Transformers ecosystem, TRL supports. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Python Bindings for llama.cpp](https://mdrss.com/llm-engineering/inference-and-quantization/901042/901042.md)

Simple Python bindings for @ggerganov's llama.cpp library. This package provides:. Use it to ground design choices in named patterns, trade-offs and examples.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [AQLM](https://mdrss.com/llm-engineering/inference-and-quantization/901037/901037.md)

Official PyTorch implementation for Extreme Compression of Large Language Models via Additive Quantization. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Awesome Model Quantization](https://mdrss.com/llm-engineering/inference-and-quantization/901001/901001.md)

This repo collects papers, documents, and codes about model quantization for anyone who wants to research it. We are continuously improving the project. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1
