BenchMark·Hub

miriad-benchmark-200k

B待确认Benchmark
垂直领域医疗开放

发布方:tomaarsen

热度15.4▼ 1.1
下载量 · 30天
14
Hugging Face
GitHub Stars
代码仓库
论文被引
10
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

MIRIAD Benchmark 200k A medical retrieval benchmark: 1,000 questions searching 200,000 passages, in the BEIR layout (corpus, queries, qrels) so it works with standard IR evaluators out of the box. It is deliberately hard. Scoring questions against only their own ~10k source passages saturates above 0.97 NDCG@10 for almost every model, which tells you nothing. Adding 190,015 deduplicated distractor passages spreads the field out across roughly 0.43 to 0.91. [!TIP] This is the… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/miriad-benchmark-200k.