BenchMark·Hub

MASK

B待确认Benchmark
安全对齐对齐安全开放

发布方:cais

热度21.3▲ 0.1
下载量 · 30天
8,910
Hugging Face
GitHub Stars
代码仓库
论文被引
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

The MASK Evaluation 🌐 Website | 📄 Paper | GitHub Center for AI Safety & Scale AI The MASK evaluation provides a rigorous benchmark for evaluating honesty in large language models by measuring whether models remain truthful when incentivized to lie. The public set contains 1,028 high-quality human-labeled examples across six distinct archetypes, each consisting of a proposition, ground truth, pressure prompt designed to elicit lying, and belief elicitation prompt to… See the full description on the dataset page: https://huggingface.co/datasets/cais/MASK.