BenchMark·Hub

llm-refusal-evaluation

B待确认Benchmark
安全对齐越狱攻防开放

发布方:MultiverseComputingCAI

热度25.1▲ 0.2
下载量 · 30天
201
Hugging Face
GitHub Stars
代码仓库
论文被引
615
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

🛡️ LLM Refusal Evaluation Benchmark This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite. The prompts are organized into three groups: Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse. Chinese Sensitive Topics — prompts that may be censored by China-aligned models. Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse. 📌 Contents Safety Benchmarks JailbreakBench SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/MultiverseComputingCAI/llm-refusal-evaluation.