BenchMark·Hub

HELMET

B待确认Benchmark
长上下文长文档开放

发布方:princeton-nlp

热度27.6▼ 0.2
下载量 · 30天
2,617
Hugging Face
GitHub Stars
代码仓库
论文被引
119
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

HELMET: How to Evaluate Long-context Language Models Effectively and Thoroughly [Paper][Code] HELMET is a comprehensive benchmark for long-context language models covering seven diverse categories of tasks. The datasets are application-centric and are designed to evaluate models at different lengths and levels of complexity. Please check out the paper for more details, and the code repo for how to process the data and run the evaluations