BenchMark·Hub

mrcr

B待确认Benchmark
长上下文针检索开放mit

发布方:openai

热度35.9▲ 0.3
下载量 · 30天
5,400
Hugging Face
GitHub Stars
代码仓库
论文被引
79
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

OpenAI MRCR: Long context multiple needle in a haystack benchmark OpenAI MRCR (Multi-round co-reference resolution) is a long context dataset for benchmarking an LLM's ability to distinguish between multiple needles hidden in context. This eval is inspired by the MRCR eval first introduced by Gemini (https://arxiv.org/pdf/2409.12640v2). OpenAI MRCR expands the tasks's difficulty and provides opensource data for reproducing results. The task is as follows: The model is given a long… See the full description on the dataset page: https://huggingface.co/datasets/openai/mrcr.