BenchMark·Hub
Benchmark测评分类可信级别开放热度
MMLU-Pro
MMLU· TIGER-Lab
知识问答通用问答
S高可信66.0
MMLU
MMLU· UC Berkeley
知识问答通用问答
S高可信51.5
MMMLU
· openai
知识问答通用问答
B待确认44.3
boolq
BoolQ· google
知识问答通用问答
B待确认42.5
qasper
· allenai
知识问答通用问答
B待确认39.6
xquad
· google
知识问答通用问答
B待确认36.4
quac
· allenai
知识问答通用问答
B待确认33.1
WildBench
· allenai
知识问答通用问答
B待确认32.5
LiveBench
· Abacus.AI
知识问答通用问答
S高可信28.7
narrativeqa_manual
· deepmind
知识问答通用问答
B待确认25.8
IndicGenBench_xquad_in
· google
知识问答通用问答
B待确认23.5
IndicGenBench_xorqa_in
· google
知识问答通用问答
B待确认22.8
gooaq
· allenai
知识问答通用问答
B待确认22.1
hle-rolling
HLE· cais
知识问答通用问答
B待确认15.9
Quad_benchmark_cqa_latest
· LLM-Beetle
知识问答通用问答
B待确认4.1
yet-another-applied-llm-benchmark
· carlini
知识问答通用问答
B待确认0.0