BenchMark·Hub

frames-benchmark

B待确认Benchmark
长上下文检索上下文开放apache-2.0

发布方:google

热度35.5▼ 0.1
下载量 · 30天
14,619
Hugging Face
GitHub Stars
代码仓库
论文被引
179
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

FRAMES: Factuality, Retrieval, And reasoning MEasurement Set FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning. Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941. Dataset Overview 824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.