# #quantization — MDRSS hashtag feed

> Public MDRSS cards tagged #quantization.
> Canonical feed: https://mdrss.com/feeds/quantization

## Cards (9)

### [G1: CUDA 12/13 ABI Mismatch](https://mdrss.com/llm-engineering/inference-and-quantization/901285/901285.md)

Last verified: 2026-07-14 — refresh when CUDA, PyTorch, or the DGX Spark stack ships a new major version. Use it to give an agent explicit responsibilities, steps and constraints.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [AQLM](https://mdrss.com/llm-engineering/inference-and-quantization/901037/901037.md)

Official PyTorch implementation for Extreme Compression of Large Language Models via Additive Quantization. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Awesome Model Quantization](https://mdrss.com/llm-engineering/inference-and-quantization/901001/901001.md)

This repo collects papers, documents, and codes about model quantization for anyone who wants to research it. We are continuously improving the project. Use it to navigate the topic and choose relevant methods, papers or tools.

Classification: llm-engineering/inference-and-quantization · Feed: llm-engineering · Updated: 2026-08-04T13:54:51.641Z · Version: 1

### [Lit-LLaMA](https://mdrss.com/llm-engineering/serving-and-retrieval/2164/2164.md)

⚠️ Warning: Not Actively Maintained This repository is no longer actively maintained. For a more up-to-date alternative, please visit the LitGPT project: https://github.com/Lightning-AI/litgpt , which serves as the successor to this repository.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [rwkv.cpp](https://mdrss.com/llm-engineering/serving-and-retrieval/2075/2075.md)

This is a port of BlinkDL/RWKV-LM to ggerganov/ggml. Besides the usual FP32, it supports FP16, quantized INT4, INT5 and INT8 inference.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [Faster Whisper transcription with CTranslate2](https://mdrss.com/multimodal/speech-and-audio/1709/1709.md)

faster-whisper is a reimplementation of OpenAI's Whisper model using CTranslate2, which is a fast inference engine for Transformer models. This implementation is up to 4 times faster than openai/whisper for the same accuracy while using less memory.

Classification: multimodal/speech-and-audio · Feed: multimodal · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [QLoRA: Efficient Finetuning of Quantized LLMs](https://mdrss.com/llm-engineering/serving-and-retrieval/1671/1671.md)

This repo supports the paper "QLoRA: Efficient Finetuning of Quantized LLMs", an effort to democratize access to LLM research. QLoRA uses bitsandbytes for quantization and is integrated with Hugging Face's PEFT and transformers libraries.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [LLM Course](https://mdrss.com/llm-engineering/serving-and-retrieval/1089/1089.md)

𝕏 Follow me on X • 🤗 Hugging Face • 💻 Blog • 📙 LLM Engineer's Handbook The LLM course is divided into three parts: 1. 🧩 LLM Fundamentals is optional and covers fundamental knowledge about mathematics, Python, and neural networks.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [Core ML Stable Diffusion](https://mdrss.com/llm-engineering/serving-and-retrieval/938/938.md)

Run Stable Diffusion on Apple Silicon with Core ML  \ Blog Post\ (https://machinelearning.apple.com/research/stable-diffusion-coreml-apple-silicon)  \ BibTeX\ (#bibtex) This repository comprises: If you run into issues during installation or runtime, please refer to the FAQ section. Please refer to the System Requirements section before getting started.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1
