BenchMark·Hub

RLPR-Evaluation

B待确认Benchmark
推理通用推理开放apache-2.0

发布方:openbmb

热度21.4▼ 0.1
下载量 · 30天
282
Hugging Face
GitHub Stars
代码仓库
论文被引
86
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

Dataset Card for RLPR-Evaluation GitHub | Paper News: [2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at here! Dataset Summary We include the following seven benchmarks for evaluation of RLPR: Mathematical Reasoning Benchmarks: MATH-500 (Cobbe et al., 2021) Minerva (Lewkowycz et al., 2022) AIME24 General Domain Reasoning Benchmarks: MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLPR-Evaluation.