miriad-benchmark-200k
B待确认Benchmark垂直领域医疗开放
发布方:tomaarsen
热度15.4▼ 1.1
下载量 · 30天
14
Hugging Face
GitHub Stars
—
代码仓库
论文被引
10
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
MIRIAD Benchmark 200k A medical retrieval benchmark: 1,000 questions searching 200,000 passages, in the BEIR layout (corpus, queries, qrels) so it works with standard IR evaluators out of the box. It is deliberately hard. Scoring questions against only their own ~10k source passages saturates above 0.97 NDCG@10 for almost every model, which tells you nothing. Adding 190,015 deduplicated distractor passages spreads the field out across roughly 0.43 to 0.91. [!TIP] This is the… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/miriad-benchmark-200k.