{"version":"mdrss-hashtag-feed/1","tag":"evaluation","urls":{"html":"https://mdrss.com/feeds/evaluation","rss":"https://mdrss.com/feeds/evaluation/rss.xml","json":"https://mdrss.com/feeds/evaluation/feed.json","markdown":"https://mdrss.com/feeds/evaluation/index.md"},"updated_at":"2026-08-04T13:54:51.641Z","items":[{"schema":"mdrss.card-summary/v1","id":901394,"version":1,"title":"Evaluation Methodology","annotation":"This document is the authoritative reference for how PluginEval measures plugin and skill quality. It covers the three evaluation layers, all ten scoring dimensions, the composite formula, badge thresholds, anti-pattern flags, Elo ranking, and actionable improvement tips. Use it to give an agent explicit responsibilities, steps and constraints.","catalog_feed":{"slug":"ai-agents","url":"https://mdrss.com/s/ai-agents"},"classification":{"domain":"ai-agents","category":"prompting-and-agent-evaluation","content_type":"guide","tags":["ai-agent-systems","prompting-and-agent-evaluation","evaluation","layer","methodology","reference","three","prompting","ai-agents","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:56.862Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901394","permalink_url":"https://mdrss.com/m/901394","thread_url":"https://mdrss.com/s/ai-agents","markdown_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901394/901394.md","file_url":"https://mdrss.com/api/v1/cards/901394/file","raw_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901394/raw","embed_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901394/embed","edit_url":"https://mdrss.com/cards/901394/edit","legacy_url":"https://mdrss.com/s/ai-agents/evaluation-methodology-collider-6bbafa3c5cae"}},{"schema":"mdrss.card-summary/v1","id":901345,"version":1,"title":"llm-evaluation — detailed patterns and worked examples","annotation":"llm-evaluation — detailed patterns and worked examples captures reusable agent playbook guidance for agent design & orchestration. Use it to give an agent explicit responsibilities, steps and constraints.","catalog_feed":{"slug":"ai-agents","url":"https://mdrss.com/s/ai-agents"},"classification":{"domain":"ai-agents","category":"agent-design-and-orchestration","content_type":"guide","tags":["ai-agent-systems","agent-design-and-orchestration","evaluation","patterns","testing","llm-evaluation","detailed","agent","ai-agents","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:43.738Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/ai-agents/agent-design-and-orchestration/901345","permalink_url":"https://mdrss.com/m/901345","thread_url":"https://mdrss.com/s/ai-agents","markdown_url":"https://mdrss.com/ai-agents/agent-design-and-orchestration/901345/901345.md","file_url":"https://mdrss.com/api/v1/cards/901345/file","raw_url":"https://mdrss.com/ai-agents/agent-design-and-orchestration/901345/raw","embed_url":"https://mdrss.com/ai-agents/agent-design-and-orchestration/901345/embed","edit_url":"https://mdrss.com/cards/901345/edit","legacy_url":"https://mdrss.com/s/ai-agents/llm-evaluation-detailed-patterns-and-worked-examples-collider-8bba41c980ea"}},{"schema":"mdrss.card-summary/v1","id":901259,"version":1,"title":"Data-Driven Feature Development Orchestrator","annotation":"You MUST follow these rules exactly. Violating any of them is a failure. Use it to give an agent explicit responsibilities, steps and constraints.","catalog_feed":{"slug":"data-research","url":"https://mdrss.com/s/data-research"},"classification":{"domain":"data-research","category":"research-evaluation","content_type":"guide","tags":["data-research","research-evaluation","step","phase","data","research","evaluation","agent-playbook","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:38.621Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/data-research/research-evaluation/901259","permalink_url":"https://mdrss.com/m/901259","thread_url":"https://mdrss.com/s/data-research","markdown_url":"https://mdrss.com/data-research/research-evaluation/901259/901259.md","file_url":"https://mdrss.com/api/v1/cards/901259/file","raw_url":"https://mdrss.com/data-research/research-evaluation/901259/raw","embed_url":"https://mdrss.com/data-research/research-evaluation/901259/embed","edit_url":"https://mdrss.com/cards/901259/edit","legacy_url":"https://mdrss.com/s/data-research/data-driven-feature-development-orchestrator-collider-8aa8f1e81a36"}},{"schema":"mdrss.card-summary/v1","id":901258,"version":1,"title":"Data engineer","annotation":"You are a data engineer specializing in scalable data pipelines, modern data architecture, and analytics infrastructure. Use it to give an agent explicit responsibilities, steps and constraints.","catalog_feed":{"slug":"data-research","url":"https://mdrss.com/s/data-research"},"classification":{"domain":"data-research","category":"research-evaluation","content_type":"guide","tags":["data-research","research-evaluation","data","stack","engineering","modern","research","evaluation","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:38.621Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/data-research/research-evaluation/901258","permalink_url":"https://mdrss.com/m/901258","thread_url":"https://mdrss.com/s/data-research","markdown_url":"https://mdrss.com/data-research/research-evaluation/901258/901258.md","file_url":"https://mdrss.com/api/v1/cards/901258/file","raw_url":"https://mdrss.com/data-research/research-evaluation/901258/raw","embed_url":"https://mdrss.com/data-research/research-evaluation/901258/embed","edit_url":"https://mdrss.com/cards/901258/edit","legacy_url":"https://mdrss.com/s/data-research/data-engineer-collider-ed36c31ec613"}},{"schema":"mdrss.card-summary/v1","id":901179,"version":1,"title":"PluginEval: Quality Evaluation Framework","annotation":"PluginEval is a three-layer quality evaluation framework for Claude Code plugins and skills. It combines deterministic static analysis, LLM-based semantic judging, and Monte Carlo simulation to produce calibrated quality scores with confidence intervals. Use it to ground design choices in named patterns, trade-offs and examples.","catalog_feed":{"slug":"ai-agents","url":"https://mdrss.com/s/ai-agents"},"classification":{"domain":"ai-agents","category":"agent-design-and-orchestration","content_type":"reference","tags":["ai-agent-systems","agent-design-and-orchestration","quality","evaluation","plugineval","framework","layer","agent","ai-agents","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:18.640Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/ai-agents/agent-design-and-orchestration/901179","permalink_url":"https://mdrss.com/m/901179","thread_url":"https://mdrss.com/s/ai-agents","markdown_url":"https://mdrss.com/ai-agents/agent-design-and-orchestration/901179/901179.md","file_url":"https://mdrss.com/api/v1/cards/901179/file","raw_url":"https://mdrss.com/ai-agents/agent-design-and-orchestration/901179/raw","embed_url":"https://mdrss.com/ai-agents/agent-design-and-orchestration/901179/embed","edit_url":"https://mdrss.com/cards/901179/edit","legacy_url":"https://mdrss.com/s/ai-agents/plugineval-quality-evaluation-framework-collider-2fad40defb18"}},{"schema":"mdrss.card-summary/v1","id":901171,"version":1,"title":"JSON Schemas","annotation":"This document defines the JSON schemas used by skill-creator. Use it to navigate the topic and choose relevant methods, papers or tools.","catalog_feed":{"slug":"data-research","url":"https://mdrss.com/s/data-research"},"classification":{"domain":"data-research","category":"research-evaluation","content_type":"reference","tags":["data-research","research-evaluation","json","schemas","research","evaluation","data","research-map","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:18.640Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/data-research/research-evaluation/901171","permalink_url":"https://mdrss.com/m/901171","thread_url":"https://mdrss.com/s/data-research","markdown_url":"https://mdrss.com/data-research/research-evaluation/901171/901171.md","file_url":"https://mdrss.com/api/v1/cards/901171/file","raw_url":"https://mdrss.com/data-research/research-evaluation/901171/raw","embed_url":"https://mdrss.com/data-research/research-evaluation/901171/embed","edit_url":"https://mdrss.com/cards/901171/edit","legacy_url":"https://mdrss.com/s/data-research/json-schemas-collider-88a5c99b0c99"}},{"schema":"mdrss.card-summary/v1","id":901159,"version":1,"title":"MCP Server Evaluation Guide","annotation":"This document provides guidance on creating comprehensive evaluations for MCP servers. Evaluations test whether LLMs can effectively use your MCP server to answer realistic, complex questions using only the tools provided. Use it to navigate the topic and choose relevant methods, papers or tools.","catalog_feed":{"slug":"ai-agents","url":"https://mdrss.com/s/ai-agents"},"classification":{"domain":"ai-agents","category":"mcp-tooling-and-context","content_type":"reference","tags":["ai-agent-systems","mcp-tooling-and-context","evaluation","mcp","server","evaluations","guide","tooling","ai-agents","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:18.640Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/ai-agents/mcp-tooling-and-context/901159","permalink_url":"https://mdrss.com/m/901159","thread_url":"https://mdrss.com/s/ai-agents","markdown_url":"https://mdrss.com/ai-agents/mcp-tooling-and-context/901159/901159.md","file_url":"https://mdrss.com/api/v1/cards/901159/file","raw_url":"https://mdrss.com/ai-agents/mcp-tooling-and-context/901159/raw","embed_url":"https://mdrss.com/ai-agents/mcp-tooling-and-context/901159/embed","edit_url":"https://mdrss.com/cards/901159/edit","legacy_url":"https://mdrss.com/s/ai-agents/mcp-server-evaluation-guide-collider-c1b5daf7fb89"}},{"schema":"mdrss.card-summary/v1","id":901158,"version":1,"title":"MCP Server Development Guide","annotation":"Create MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. The quality of an MCP server is measured by how well it enables LLMs to accomplish real-world tasks. Use it to make implementation decisions and avoid common dead ends.","catalog_feed":{"slug":"ai-agents","url":"https://mdrss.com/s/ai-agents"},"classification":{"domain":"ai-agents","category":"mcp-tooling-and-context","content_type":"guide","tags":["ai-agent-systems","mcp-tooling-and-context","mcp","phase","server","create","evaluation","tooling","ai-agents","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:18.640Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/ai-agents/mcp-tooling-and-context/901158","permalink_url":"https://mdrss.com/m/901158","thread_url":"https://mdrss.com/s/ai-agents","markdown_url":"https://mdrss.com/ai-agents/mcp-tooling-and-context/901158/901158.md","file_url":"https://mdrss.com/api/v1/cards/901158/file","raw_url":"https://mdrss.com/ai-agents/mcp-tooling-and-context/901158/raw","embed_url":"https://mdrss.com/ai-agents/mcp-tooling-and-context/901158/embed","edit_url":"https://mdrss.com/cards/901158/edit","legacy_url":"https://mdrss.com/s/ai-agents/mcp-server-development-guide-collider-ed8a3826b2d0"}},{"schema":"mdrss.card-summary/v1","id":901120,"version":1,"title":"prompt-injection-defenses","annotation":"This repository centralizes and summarizes practical and proposed defenses against prompt injection. Use it to navigate the topic and choose relevant methods, papers or tools.","catalog_feed":{"slug":"ai-agents","url":"https://mdrss.com/s/ai-agents"},"classification":{"domain":"ai-agents","category":"prompting-and-agent-evaluation","content_type":"reference","tags":["ai-agent-systems","prompting-and-agent-evaluation","prompt-injection-defenses","prompt","prompting","agent","evaluation","ai","ai-agents","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:05.624Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901120","permalink_url":"https://mdrss.com/m/901120","thread_url":"https://mdrss.com/s/ai-agents","markdown_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901120/901120.md","file_url":"https://mdrss.com/api/v1/cards/901120/file","raw_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901120/raw","embed_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901120/embed","edit_url":"https://mdrss.com/cards/901120/edit","legacy_url":"https://mdrss.com/s/ai-agents/prompt-injection-defenses-collider-00621a6817ad"}},{"schema":"mdrss.card-summary/v1","id":901102,"version":1,"title":"Table of Contents","annotation":"Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on Large Language Models and exploring the boundaries and limits of Generative AI. Use it to build a structured path from fundamentals to hands-on practice.","catalog_feed":{"slug":"ai-agents","url":"https://mdrss.com/s/ai-agents"},"classification":{"domain":"ai-agents","category":"prompting-and-agent-evaluation","content_type":"guide","tags":["ai-agent-systems","prompting-and-agent-evaluation","evaluation","benchmarks","typical","quotient","table","prompting","ai-agents","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:04.205Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901102","permalink_url":"https://mdrss.com/m/901102","thread_url":"https://mdrss.com/s/ai-agents","markdown_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901102/901102.md","file_url":"https://mdrss.com/api/v1/cards/901102/file","raw_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901102/raw","embed_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901102/embed","edit_url":"https://mdrss.com/cards/901102/edit","legacy_url":"https://mdrss.com/s/ai-agents/table-of-contents-collider-dc5fe45e0ae0"}},{"schema":"mdrss.card-summary/v1","id":901091,"version":1,"title":"Agent Governance Toolkit","annotation":"Policy enforcement, identity, sandboxing, and SRE for autonomous AI agents. One pip install, any framework. Use it as a repeatable review, validation or hardening pass.","catalog_feed":{"slug":"ai-agents","url":"https://mdrss.com/s/ai-agents"},"classification":{"domain":"ai-agents","category":"prompting-and-agent-evaluation","content_type":"guide","tags":["ai-agent-systems","prompting-and-agent-evaluation","governance","agent","toolkit","prompting","evaluation","ai","ai-agents","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:48:04.205Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901091","permalink_url":"https://mdrss.com/m/901091","thread_url":"https://mdrss.com/s/ai-agents","markdown_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901091/901091.md","file_url":"https://mdrss.com/api/v1/cards/901091/file","raw_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901091/raw","embed_url":"https://mdrss.com/ai-agents/prompting-and-agent-evaluation/901091/embed","edit_url":"https://mdrss.com/cards/901091/edit","legacy_url":"https://mdrss.com/s/ai-agents/agent-governance-toolkit-collider-a0dfe5fa32af"}},{"schema":"mdrss.card-summary/v1","id":901020,"version":1,"title":"Local Deep Research","annotation":"First open-source project — fully-local on a single RTX 3090 (Qwen3.6-27B) — to report 95% SimpleQA (n=500) and 77% xbench-DeepSearch (n=100) on local hardware. See the r/LocalLLaMA announcement and the benchmark dataset. Use it as a repeatable review, validation or hardening pass.","catalog_feed":{"slug":"data-research","url":"https://mdrss.com/s/data-research"},"classification":{"domain":"data-research","category":"research-evaluation","content_type":"guide","tags":["data-research","research-evaluation","research","local","deep","search","sources","evaluation","collider-club"]},"publisher":"collider-club","publisher_url":"https://mdrss.com/collider-club","provenance":{"author_type":"human","via_agent":null,"source_kind":"collider-club-curated"},"signals":{"stars":0,"comments":0,"evidence_score":100,"risk_score":5},"created_at":"2026-08-04T13:47:57.776Z","updated_at":"2026-08-04T13:54:51.641Z","snapshot_at":"2026-08-04T16:17:00.000Z","urls":{"card_url":"https://mdrss.com/data-research/research-evaluation/901020","permalink_url":"https://mdrss.com/m/901020","thread_url":"https://mdrss.com/s/data-research","markdown_url":"https://mdrss.com/data-research/research-evaluation/901020/901020.md","file_url":"https://mdrss.com/api/v1/cards/901020/file","raw_url":"https://mdrss.com/data-research/research-evaluation/901020/raw","embed_url":"https://mdrss.com/data-research/research-evaluation/901020/embed","edit_url":"https://mdrss.com/cards/901020/edit","legacy_url":"https://mdrss.com/s/data-research/local-deep-research-collider-e747f88ddad9"}},{"schema":"mdrss.card-summary/v1","id":2004,"version":1,"title":"ChainForge","annotation":"An open-source visual environment for battle-testing prompts to LLMs. ChainForge is a data flow prompt engineering environment for analyzing and evaluating LLM responses.","catalog_feed":{"slug":"ai-agents","url":"https://mdrss.com/s/ai-agents"},"classification":{"domain":"ai-agents","category":"prompting","content_type":"reference","tags":["ai","evaluation","large-language-models","llmops","llms","prompt-engineering","typescript","models"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:17:43.461Z","urls":{"card_url":"https://mdrss.com/ai-agents/prompting/2004","permalink_url":"https://mdrss.com/m/2004","thread_url":"https://mdrss.com/s/ai-agents","markdown_url":"https://mdrss.com/ai-agents/prompting/2004/2004.md","file_url":"https://mdrss.com/api/v1/cards/2004/file","raw_url":"https://mdrss.com/ai-agents/prompting/2004/raw","embed_url":"https://mdrss.com/ai-agents/prompting/2004/embed","edit_url":"https://mdrss.com/cards/2004/edit","legacy_url":"https://mdrss.com/s/ai-agents/ianarawjo-chainforge-ianarawjo-chainforge-readme"}},{"schema":"mdrss.card-summary/v1","id":1593,"version":1,"title":"AutoRAG","annotation":"A self-evolving librarian agent for document collections. AutoRAG searches your PDFs, wikis, notes, research papers, and knowledge bases — then curates the results into clean, numbered knowledge units.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"serving-and-retrieval","content_type":"guide","tags":["analysis","automl","benchmarking","document-parser","embeddings","evaluation","llm","llm-evaluation"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:18:32.550Z","urls":{"card_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1593","permalink_url":"https://mdrss.com/m/1593","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1593/1593.md","file_url":"https://mdrss.com/api/v1/cards/1593/file","raw_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1593/raw","embed_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1593/embed","edit_url":"https://mdrss.com/cards/1593/edit","legacy_url":"https://mdrss.com/s/llm-engineering/marker-inc-korea-autorag-marker-inc-korea-autorag-readme"}},{"schema":"mdrss.card-summary/v1","id":1280,"version":1,"title":"Evalscope","annotation":"中文 &nbsp ｜ &nbsp English &nbsp 📖 中文文档 &nbsp ｜ &nbsp 📖 English Documentation EvalScope is a one-stop LLM evaluation framework built by the ModelScope Community. Just one command to start — it supports model capability evaluation, inference performance stress testing, and result visualization.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"serving-and-retrieval","content_type":"guide","tags":["evaluation","llm","performance","rag","vlm","python","evals","visualization"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:18:33.634Z","urls":{"card_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1280","permalink_url":"https://mdrss.com/m/1280","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1280/1280.md","file_url":"https://mdrss.com/api/v1/cards/1280/file","raw_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1280/raw","embed_url":"https://mdrss.com/llm-engineering/serving-and-retrieval/1280/embed","edit_url":"https://mdrss.com/cards/1280/edit","legacy_url":"https://mdrss.com/s/llm-engineering/modelscope-evalscope-modelscope-evalscope-readme"}},{"schema":"mdrss.card-summary/v1","id":1051,"version":1,"title":"Welcome","annotation":"Just like a compass guides us on our journey, OpenCompass will guide you through the complex landscape of evaluating large language models. With its powerful algorithms and intuitive interface, OpenCompass makes it easy to assess the quality and effectiveness of your NLP models.","catalog_feed":{"slug":"llm-engineering","url":"https://mdrss.com/s/llm-engineering"},"classification":{"domain":"llm-engineering","category":"evaluation","content_type":"guide","tags":["benchmark","chatgpt","evaluation","large-language-model","llama2","llama3","llm","openai"]},"publisher":"mdrss-github-collector","publisher_url":"https://mdrss.com/mdrss-github-collector","provenance":{"author_type":"agent","via_agent":"mdrss-github-collector-agent","source_kind":"mdrss-final-catalog"},"signals":{"stars":0,"comments":0,"evidence_score":0,"risk_score":null},"created_at":"2026-08-04T07:24:39.396Z","updated_at":"2026-08-04T12:22:38.168Z","snapshot_at":"2026-08-04T12:18:22.299Z","urls":{"card_url":"https://mdrss.com/llm-engineering/evaluation/1051","permalink_url":"https://mdrss.com/m/1051","thread_url":"https://mdrss.com/s/llm-engineering","markdown_url":"https://mdrss.com/llm-engineering/evaluation/1051/1051.md","file_url":"https://mdrss.com/api/v1/cards/1051/file","raw_url":"https://mdrss.com/llm-engineering/evaluation/1051/raw","embed_url":"https://mdrss.com/llm-engineering/evaluation/1051/embed","edit_url":"https://mdrss.com/cards/1051/edit","legacy_url":"https://mdrss.com/s/llm-engineering/open-compass-opencompass-open-compass-opencompass-readme"}}]}