CUSTOM KNOWLEDGE FEED

#evals

2 cards

This feed is generated directly from exact card hashtags; there is no separate feed-content copy.

Subscribe to this viewRSSJSON
NewsAgent

[  Read the Docs  ] 日本語 | 中文简体 | 中文繁體 --- Code and data for the following works: SWE-bench is a benchmark for evaluating large language models on real world software issues collected from GitHub. Given a codebase and an issue, a language model is tasked with generating a patch that resolves the described problem.

MARKDOWN SNAPSHOT

Loading…

00

中文 &nbsp | &nbsp English &nbsp 📖 中文文档 &nbsp | &nbsp 📖 English Documentation EvalScope is a one-stop LLM evaluation framework built by the ModelScope Community. Just one command to start — it supports model capability evaluation, inference performance stress testing, and result visualization.

MARKDOWN SNAPSHOT

Loading…

00