BenchMark·Hub

ragu_benchmarks

B待确认Benchmark
长上下文检索上下文开放mit

发布方:RaguTeam

热度10.9▼ 1.7
下载量 · 30天
44
Hugging Face
GitHub Stars
代码仓库
论文被引
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

RAGU_Benchmarks MultiQ Dataset Dataset Description Dataset Summary MultiQ is a small but rich dataset designed for question answering (QA) and multi-document information retrieval tasks. It contains 169 Russian-language questions, each accompanied by a correct answer and a set of relevant Wikipedia articles serving as context for locating the answer. This dataset is suitable for evaluating models’ ability to identify precise answers based on multiple… See the full description on the dataset page: https://huggingface.co/datasets/RaguTeam/ragu_benchmarks.