{"version":"mdrss-hashtag-feed/1","tag":"llama","urls":{"html":"https://mdrss.com/feeds/llama","rss":"https://mdrss.com/feeds/llama/rss.xml","json":"https://mdrss.com/feeds/llama/feed.json","markdown":"https://mdrss.com/feeds/llama/index.md"},"updated_at":"2026-08-04T12:22:38.168Z","items":[{"schema":"mdrss.card-summary/v1","id":2676,"version":1,"title":"medAlpaca: Finetuned Large Language Models for Medical Question Answering","annotation":"MedAlpaca expands upon both Stanford Alpaca and AlpacaLoRA to offer an advanced suite of large language models specifically fine-tuned for medical question-answering and dialogue applications. Our primary objective is to deliver an array of open-source language models, paving the way for seamless development of medical chatbot solutions.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"evaluation","content_type":"reference","tags":["python","models","llama","benchmarks","llm"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:18:22.299Z","urls":{"card_url":"https://mdrss.com/llm-engineering/evaluation/2676","permalink_url":"https://mdrss.com/m/2676","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/evaluation/2676/2676.md","file_url":"https://mdrss.com/api/v1/cards/2676/file","raw_url":"https://mdrss.com/llm-engineering/evaluation/2676/raw","embed_url":"https://mdrss.com/llm-engineering/evaluation/2676/embed","edit_url":"https://mdrss.com/cards/2676/edit","legacy_url":"https://mdrss.com/s/llm-engineering/kbressem-medalpaca-kbressem-medalpaca-readme"}},{"schema":"mdrss.card-summary/v1","id":2180,"version":1,"title":"wllama - Wasm binding for llama.cpp","annotation":"For embeddings, please see examples/embeddings/index.html WebGPU support is introduced via PR #215. Upon updating to V3.1, WebGPU will be enabled automatically.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"models-and-training","content_type":"guide","tags":["llama","llamacpp","llm","wasm","webassembly","typescript","models","documentation"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:18:31.127Z","urls":{"card_url":"https://mdrss.com/llm-engineering/models-and-training/2180","permalink_url":"https://mdrss.com/m/2180","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/models-and-training/2180/2180.md","file_url":"https://mdrss.com/api/v1/cards/2180/file","raw_url":"https://mdrss.com/llm-engineering/models-and-training/2180/raw","embed_url":"https://mdrss.com/llm-engineering/models-and-training/2180/embed","edit_url":"https://mdrss.com/cards/2180/edit","legacy_url":"https://mdrss.com/s/llm-engineering/ngxson-wllama-ngxson-wllama-readme"}},{"schema":"mdrss.card-summary/v1","id":2164,"version":1,"title":"Lit-LLaMA","annotation":"⚠️ Warning: Not Actively Maintained This repository is no longer actively maintained. For a more up-to-date alternative, please visit the LitGPT project: https://github.com/Lightning-AI/litgpt , which serves as the successor to this repository.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"serving-and-retrieval","content_type":"guide","tags":["python","models","llama","quantization","lora","fine-tuning","apache"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:18:49.845Z","urls":{"card_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/2164","permalink_url":"https://mdrss.com/m/2164","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/2164/2164.md","file_url":"https://mdrss.com/api/v1/cards/2164/file","raw_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/2164/raw","embed_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/2164/embed","edit_url":"https://mdrss.com/cards/2164/edit","legacy_url":"https://mdrss.com/s/llm-engineering/lightning-ai-lit-llama-lightning-ai-lit-llama-readme"}},{"schema":"mdrss.card-summary/v1","id":1899,"version":1,"title":"koboldcpp","annotation":"KoboldCpp is an easy-to-use AI text-generation software for GGML and GGUF models, inspired by the original KoboldAI. It's a single self-contained distributable that builds off llama.cpp and adds many additional powerful features.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"serving-and-retrieval","content_type":"reference","tags":["gemma","ggml","gguf","koboldai","koboldcpp","language-model","llama","llamacpp"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:18:48.518Z","urls":{"card_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1899","permalink_url":"https://mdrss.com/m/1899","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1899/1899.md","file_url":"https://mdrss.com/api/v1/cards/1899/file","raw_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1899/raw","embed_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1899/embed","edit_url":"https://mdrss.com/cards/1899/edit","legacy_url":"https://mdrss.com/s/llm-engineering/lostruins-koboldcpp-lostruins-koboldcpp-readme"}},{"schema":"mdrss.card-summary/v1","id":1726,"version":1,"title":"REST API examples","annotation":"I've started to work on reimplementation of the library here: FastTensors Please star it if you'd like to see GGML-compatible implementation in pure Go. Please check out my related project Booster We dream of a world where fellow ML hackers are grokking REALLY BIG GPT models in their homelabs without having GPU clusters consuming a shit tons of $$$.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"serving-and-retrieval","content_type":"guide","tags":["alpaca","chatgpt","dalai","gpt","gpt3","gpt4","gpt4all","llama"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:18:33.634Z","urls":{"card_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1726","permalink_url":"https://mdrss.com/m/1726","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1726/1726.md","file_url":"https://mdrss.com/api/v1/cards/1726/file","raw_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1726/raw","embed_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1726/embed","edit_url":"https://mdrss.com/cards/1726/edit","legacy_url":"https://mdrss.com/s/llm-engineering/gotzmann-llama-go-gotzmann-llama-go-readme"}},{"schema":"mdrss.card-summary/v1","id":1537,"version":1,"title":"llama.rn","annotation":"React Native binding of llama.cpp - LLM inference in C/C++ Key Features: llama.rn downloads the pre-built ios/rnllama.xcframework and android/src/main/jniLibs from the matching GitHub release during postinstall. Existing downloads are reused, and each archive is verified with SHA-256 before extraction.","catalog_feed":{"slug":"web-mobile","url":"https://mdrss.com/s/web-mobile"},"classification":{"domain":"web-mobile","category":"mobile-development","content_type":"guide","tags":["android","ios","llama","llama-cpp","llm","react-native","c++","mobile"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:19:52.567Z","urls":{"card_url":"https://mdrss.com/web-mobile/mobile-development/1537","permalink_url":"https://mdrss.com/m/1537","thread_url":"https://mdrss.com/s/web-mobile","markdown_url":"https://mdrss.com/web-mobile/mobile-development/1537/1537.md","file_url":"https://mdrss.com/api/v1/cards/1537/file","raw_url":"https://mdrss.com/web-mobile/mobile-development/1537/raw","embed_url":"https://mdrss.com/web-mobile/mobile-development/1537/embed","edit_url":"https://mdrss.com/cards/1537/edit","legacy_url":"https://mdrss.com/s/web-mobile/mybigday-llama-rn-mybigday-llama-rn-readme"}},{"schema":"mdrss.card-summary/v1","id":1244,"version":1,"title":"PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU","annotation":"PowerInfer is a CPU/GPU LLM inference engine leveraging activation locality for your device. Project Kanban https://github.com/SJTU-IPADS/PowerInfer/assets/34213478/fe441a42-5fce-448b-a3e5-ea4abb43ba23 PowerInfer v.s.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"serving-and-retrieval","content_type":"guide","tags":["large-language-models","llama","llm","llm-inference","local-inference","c++","models","architecture"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:18:32.550Z","urls":{"card_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1244","permalink_url":"https://mdrss.com/m/1244","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1244/1244.md","file_url":"https://mdrss.com/api/v1/cards/1244/file","raw_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1244/raw","embed_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1244/embed","edit_url":"https://mdrss.com/cards/1244/edit","legacy_url":"https://mdrss.com/s/llm-engineering/tiiny-ai-powerinfer-tiiny-ai-powerinfer-readme"}},{"schema":"mdrss.card-summary/v1","id":1049,"version":1,"title":"Get started","annotation":"Unsloth Studio lets you run and train models locally. Features • News • Quickstart • Notebooks • Documentation Unsloth Studio (Beta) lets you run and train text, audio, embedding, vision models on Windows, Linux and macOS.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"serving-and-retrieval","content_type":"guide","tags":["agent","deepseek","fine-tuning","gemma","gemma3","gpt-oss","llama","llama3"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:18:32.550Z","urls":{"card_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1049","permalink_url":"https://mdrss.com/m/1049","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1049/1049.md","file_url":"https://mdrss.com/api/v1/cards/1049/file","raw_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1049/raw","embed_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1049/embed","edit_url":"https://mdrss.com/cards/1049/edit","legacy_url":"https://mdrss.com/s/llm-engineering/unslothai-unsloth-unslothai-unsloth-readme"}}]}