LLMBar
B待确认Benchmark推理指令遵循开放mit
发布方:princeton-nlp
热度11.7▼ 0.0
下载量 · 30天
146
Hugging Face
GitHub Stars
—
代码仓库
论文被引
—
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
LLMBar is a challenging meta-evaluation benchmark designed to test the ability of an LLM evaluator in discerning instruction-following outputs. LLMBar consists of 419 instances, where each entry contains an instruction paired with two outputs: one faithfully and correctly follows the instruction and the other deviates from it. There is also a gold preference label indicating which output is objectively better for each instance.