HELMET
B待确认Benchmark长上下文长文档开放
发布方:princeton-nlp
热度27.6▼ 0.2
下载量 · 30天
2,617
Hugging Face
GitHub Stars
—
代码仓库
论文被引
119
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
HELMET: How to Evaluate Long-context Language Models Effectively and Thoroughly [Paper][Code] HELMET is a comprehensive benchmark for long-context language models covering seven diverse categories of tasks. The datasets are application-centric and are designed to evaluate models at different lengths and levels of complexity. Please check out the paper for more details, and the code repo for how to process the data and run the evaluations