openai-moderation-api-evaluation
B待确认Benchmark安全对齐对齐安全开放mit
发布方:mmathys
热度32.1▼ 0.3
下载量 · 30天
2,956
Hugging Face
GitHub Stars
—
代码仓库
论文被引
492
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.