ragu_benchmarks
B待确认Benchmark长上下文检索上下文开放mit
发布方:RaguTeam
热度10.9▼ 1.7
下载量 · 30天
44
Hugging Face
GitHub Stars
—
代码仓库
论文被引
—
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
RAGU_Benchmarks MultiQ Dataset Dataset Description Dataset Summary MultiQ is a small but rich dataset designed for question answering (QA) and multi-document information retrieval tasks. It contains 169 Russian-language questions, each accompanied by a correct answer and a set of relevant Wikipedia articles serving as context for locating the answer. This dataset is suitable for evaluating models’ ability to identify precise answers based on multiple… See the full description on the dataset page: https://huggingface.co/datasets/RaguTeam/ragu_benchmarks.