# #evals — MDRSS hashtag feed

> Public MDRSS cards tagged #evals.
> Canonical feed: https://mdrss.com/feeds/evals

## Cards (2)

### [News](https://mdrss.com/llm-engineering/evaluation/2249/2249.md)

&nbsp; Read the Docs &nbsp;  日本語 | 中文简体 | 中文繁體 --- Code and data for the following works: SWE-bench is a benchmark for evaluating large language models on real world software issues collected from GitHub. Given a codebase and an issue, a language model is tasked with generating a patch that resolves the described problem.

Classification: llm-engineering/evaluation · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1

### [Evalscope](https://mdrss.com/llm-engineering/serving-and-retrieval/1280/1280.md)

中文 &nbsp ｜ &nbsp English &nbsp 📖 中文文档 &nbsp ｜ &nbsp 📖 English Documentation EvalScope is a one-stop LLM evaluation framework built by the ModelScope Community. Just one command to start — it supports model capability evaluation, inference performance stress testing, and result visualization.

Classification: llm-engineering/serving-and-retrieval · Feed: llm-engineering · Updated: 2026-08-04T12:22:38.168Z · Version: 1
