BenchMark·Hub
Benchmark测评分类可信级别开放热度
LongBench v2
LongBench· THUDM
长上下文长文档
S高可信37.7
narrativeqa
· deepmind
长上下文长文档
B待确认36.4
mrcr
· openai
长上下文针检索
B待确认35.9
frames-benchmark
· google
长上下文检索上下文
B待确认35.5
pg19
· deepmind
长上下文长文档
B待确认35.1
HELMET
· princeton-nlp
长上下文长文档
B待确认27.6
LongRAG
· TIGER-Lab
长上下文检索上下文
B待确认25.5
LongICLBench
· TIGER-Lab
长上下文检索上下文
B待确认23.9
MMDocIR_Evaluation_Dataset
· MMDocIR
长上下文长文档
B待确认22.9
German-RAG-LLM-HARD-BENCHMARK
· avemio
长上下文检索上下文
B待确认21.4
pdfQA-Benchmark
· pdfqa
长上下文长文档
B待确认21.0
German-RAG-LLM-EASY-BENCHMARK
· avemio
长上下文检索上下文
B待确认20.6
BrowseCompLongContext
· openai
长上下文检索上下文
B待确认20.5
RULER
· NVIDIA
长上下文针检索
S高可信13.9
wmt-da-human-evaluation-long-context
· ymoslem
长上下文长文档
B待确认13.4
ragu_benchmarks
· RaguTeam
长上下文检索上下文
B待确认10.9
BRIGHT
· xlang-ai
长上下文检索上下文
B待确认8.9
Needle in a Haystack
· Greg Kamradt
长上下文针检索
S高可信0.0
Needle in a Haystack
· Greg Kamradt
长上下文针检索
S高可信0.0