# #llama — MDRSS hashtag feed

> Public MDRSS cards tagged #llama.
> Canonical feed: https://mdrss.com/feeds/llama

## Cards (8)

### [medAlpaca: Finetuned Large Language Models for Medical Question Answering](https://mdrss.com/llm-engineering/evaluation/2676/2676.md)

MedAlpaca expands upon both Stanford Alpaca and AlpacaLoRA to offer an advanced suite of large language models specifically fine-tuned for medical question-answering and dialogue applications. Our primary objective is to deliver an array of open-source language models, paving the way for seamless development of medical chatbot solutions.

Classification: llm-engineering/evaluation · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [wllama - Wasm binding for llama.cpp](https://mdrss.com/llm-engineering/models-and-training/2180/2180.md)

For embeddings, please see examples/embeddings/index.html WebGPU support is introduced via PR #215. Upon updating to V3.1, WebGPU will be enabled automatically.

Classification: llm-engineering/models-and-training · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [Lit-LLaMA](https://mdrss.com/llm-engineering/serving-and-retrieval/2164/2164.md)

⚠️ Warning: Not Actively Maintained This repository is no longer actively maintained. For a more up-to-date alternative, please visit the LitGPT project: https://github.com/Lightning-AI/litgpt , which serves as the successor to this repository.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [koboldcpp](https://mdrss.com/llm-engineering/serving-and-retrieval/1899/1899.md)

KoboldCpp is an easy-to-use AI text-generation software for GGML and GGUF models, inspired by the original KoboldAI. It's a single self-contained distributable that builds off llama.cpp and adds many additional powerful features.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [REST API examples](https://mdrss.com/llm-engineering/serving-and-retrieval/1726/1726.md)

I've started to work on reimplementation of the library here: FastTensors Please star it if you'd like to see GGML-compatible implementation in pure Go. Please check out my related project Booster We dream of a world where fellow ML hackers are grokking REALLY BIG GPT models in their homelabs without having GPU clusters consuming a shit tons of $$$.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [llama.rn](https://mdrss.com/web-mobile/mobile-development/1537/1537.md)

React Native binding of llama.cpp - LLM inference in C/C++ Key Features: llama.rn downloads the pre-built ios/rnllama.xcframework and android/src/main/jniLibs from the matching GitHub release during postinstall. Existing downloads are reused, and each archive is verified with SHA-256 before extraction.

Classification: web-mobile/mobile-development · Feed: web-mobile · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU](https://mdrss.com/llm-engineering/serving-and-retrieval/1244/1244.md)

PowerInfer is a CPU/GPU LLM inference engine leveraging activation locality for your device. Project Kanban https://github.com/SJTU-IPADS/PowerInfer/assets/34213478/fe441a42-5fce-448b-a3e5-ea4abb43ba23 PowerInfer v.s.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [Get started](https://mdrss.com/llm-engineering/serving-and-retrieval/1049/1049.md)

Unsloth Studio lets you run and train models locally. Features • News • Quickstart • Notebooks • Documentation Unsloth Studio (Beta) lets you run and train text, audio, embedding, vision models on Windows, Linux and macOS.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1
