llm-refusal-evaluation
B待确认Benchmark安全对齐越狱攻防开放
发布方:MultiverseComputingCAI
热度25.1▲ 0.2
下载量 · 30天
201
Hugging Face
GitHub Stars
—
代码仓库
论文被引
615
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
🛡️ LLM Refusal Evaluation Benchmark This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite. The prompts are organized into three groups: Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse. Chinese Sensitive Topics — prompts that may be censored by China-aligned models. Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse. 📌 Contents Safety Benchmarks JailbreakBench SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/MultiverseComputingCAI/llm-refusal-evaluation.