BenchMark·Hub
Benchmark测评分类可信级别开放热度
frames-benchmark
· google
长上下文检索上下文
B待确认35.5
LongRAG
· TIGER-Lab
长上下文检索上下文
B待确认25.5
LongICLBench
· TIGER-Lab
长上下文检索上下文
B待确认23.9
German-RAG-LLM-HARD-BENCHMARK
· avemio
长上下文检索上下文
B待确认21.4
German-RAG-LLM-EASY-BENCHMARK
· avemio
长上下文检索上下文
B待确认20.6
BrowseCompLongContext
· openai
长上下文检索上下文
B待确认20.5
ragu_benchmarks
· RaguTeam
长上下文检索上下文
B待确认10.9
BRIGHT
· xlang-ai
长上下文检索上下文
B待确认8.9