BenchMark·Hub

SWE-bench Verified

S高可信Variant
代码仓库级代码SWE-bench 家族开放

发布方:Princeton + OpenAI

热度58.9▼ 1.2
下载量 · 30天
336,081
Hugging Face
GitHub Stars
代码仓库
论文被引
3,498
Semantic Scholar
跑分模型 · 30天
89
Leaderboard results

简介

Dataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process. The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The original… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified.