BenchMark·Hub

SWE-bench

B待确认Benchmark
代码代码生成SWE-bench 家族开放

发布方:princeton-nlp

热度44.5▼ 0.3
下载量 · 30天
109,656
Hugging Face
GitHub Stars
代码仓库
论文被引
3,498
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the problem_statement… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench.