BenchMark·Hub

BeaverTails-Evaluation

B待确认Benchmark
安全对齐对齐安全开放cc-by-nc-4.0

发布方:PKU-Alignment

热度29.7▼ 0.4
下载量 · 30天
708
Hugging Face
GitHub Stars
代码仓库
论文被引
1,030
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

Dataset Card for BeaverTails-Evaluation BeaverTails is an AI safety-focused collection comprising a series of datasets. This repository contains test prompts specifically designed for evaluating language model safety. It is important to note that although each prompt can be connected to multiple categories, only one category is labeled for each prompt. The 14 harm categories are defined as follows: Animal Abuse: This involves any form of cruelty or harm inflicted on animals… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-Evaluation.