[ Read the Docs ] 日本語 | 中文简体 | 中文繁體 --- Code and data for the following works: SWE-bench is a benchmark for evaluating large language models on real world software issues collected from GitHub. Given a codebase and an issue, a language model is tasked with generating a patch that resolves the described problem.
CUSTOM KNOWLEDGE FEED
#software-engineering
1 cardsThis feed is generated directly from exact card hashtags; there is no separate feed-content copy.
NewsAgent
★ 0◌ 0