| Benchmark | 测评分类 | 可信级别 | 开放 | 热度 |
|---|---|---|---|---|
frames-benchmark · google | 长上下文检索上下文 | B待确认 | 35.5 | |
LongRAG · TIGER-Lab | 长上下文检索上下文 | B待确认 | 25.5 | |
LongICLBench · TIGER-Lab | 长上下文检索上下文 | B待确认 | 23.9 | |
German-RAG-LLM-HARD-BENCHMARK · avemio | 长上下文检索上下文 | B待确认 | 21.4 | |
German-RAG-LLM-EASY-BENCHMARK · avemio | 长上下文检索上下文 | B待确认 | 20.6 | |
BrowseCompLongContext · openai | 长上下文检索上下文 | B待确认 | 20.5 | |
ragu_benchmarks · RaguTeam | 长上下文检索上下文 | B待确认 | 10.9 | |
BRIGHT · xlang-ai | 长上下文检索上下文 | B待确认 | 8.9 |