BenchMark·Hub

C-Eval

S高可信Benchmark
中文中文通用开放cc-by-nc-sa-4.0

发布方:上交/清华

热度42.5▼ 0.3
下载量 · 30天
93,348
Hugging Face
GitHub Stars
代码仓库
论文被引
916
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

C-Eval is a comprehensive Chinese evaluation suite for foundation models. It consists of 13948 multi-choice questions spanning 52 diverse disciplines and four difficulty levels. Please visit our website and GitHub or check our paper for more details. Each subject consists of three splits: dev, val, and test. The dev set per subject consists of five exemplars with explanations for few-shot evaluation. The val set is intended to be used for hyperparameter tuning. And the test set is for model… See the full description on the dataset page: https://huggingface.co/datasets/ceval/ceval-exam.