BeaverTails-Evaluation
B待确认Benchmark安全对齐对齐安全开放cc-by-nc-4.0
发布方:PKU-Alignment
热度29.7▼ 0.4
下载量 · 30天
708
Hugging Face
GitHub Stars
—
代码仓库
论文被引
1,030
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
Dataset Card for BeaverTails-Evaluation BeaverTails is an AI safety-focused collection comprising a series of datasets. This repository contains test prompts specifically designed for evaluating language model safety. It is important to note that although each prompt can be connected to multiple categories, only one category is labeled for each prompt. The 14 harm categories are defined as follows: Animal Abuse: This involves any form of cruelty or harm inflicted on animals… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-Evaluation.