MASK
B待确认Benchmark安全对齐对齐安全开放
发布方:cais
热度21.3▲ 0.1
下载量 · 30天
8,910
Hugging Face
GitHub Stars
—
代码仓库
论文被引
—
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
The MASK Evaluation 🌐 Website | 📄 Paper | GitHub Center for AI Safety & Scale AI The MASK evaluation provides a rigorous benchmark for evaluating honesty in large language models by measuring whether models remain truthful when incentivized to lie. The public set contains 1,028 high-quality human-labeled examples across six distinct archetypes, each consisting of a proposition, ground truth, pressure prompt designed to elicit lying, and belief elicitation prompt to… See the full description on the dataset page: https://huggingface.co/datasets/cais/MASK.