| Benchmark | 测评分类 | 可信级别 | 开放 | 热度 |
|---|---|---|---|---|
LongBench v2 LongBench· THUDM | 长上下文长文档 | S高可信 | 37.7 | |
narrativeqa · deepmind | 长上下文长文档 | B待确认 | 36.4 | |
mrcr · openai | 长上下文针检索 | B待确认 | 35.9 | |
frames-benchmark · google | 长上下文检索上下文 | B待确认 | 35.5 | |
pg19 · deepmind | 长上下文长文档 | B待确认 | 35.1 | |
HELMET · princeton-nlp | 长上下文长文档 | B待确认 | 27.6 | |
LongRAG · TIGER-Lab | 长上下文检索上下文 | B待确认 | 25.5 | |
LongICLBench · TIGER-Lab | 长上下文检索上下文 | B待确认 | 23.9 | |
MMDocIR_Evaluation_Dataset · MMDocIR | 长上下文长文档 | B待确认 | 22.9 | |
German-RAG-LLM-HARD-BENCHMARK · avemio | 长上下文检索上下文 | B待确认 | 21.4 | |
pdfQA-Benchmark · pdfqa | 长上下文长文档 | B待确认 | 21.0 | |
German-RAG-LLM-EASY-BENCHMARK · avemio | 长上下文检索上下文 | B待确认 | 20.6 | |
BrowseCompLongContext · openai | 长上下文检索上下文 | B待确认 | 20.5 | |
RULER · NVIDIA | 长上下文针检索 | S高可信 | 13.9 | |
wmt-da-human-evaluation-long-context · ymoslem | 长上下文长文档 | B待确认 | 13.4 | |
ragu_benchmarks · RaguTeam | 长上下文检索上下文 | B待确认 | 10.9 | |
BRIGHT · xlang-ai | 长上下文检索上下文 | B待确认 | 8.9 | |
Needle in a Haystack · Greg Kamradt | 长上下文针检索 | S高可信 | 0.0 | |
Needle in a Haystack · Greg Kamradt | 长上下文针检索 | S高可信 | 0.0 |