| Benchmark | 测评分类 | 可信级别 | 开放 | 热度 |
|---|---|---|---|---|
llm-redactor-leak-benchmark · jayluxferro | 安全对齐对齐安全 | B待确认 | 8.9 | |
combine-llm-security-benchmark · tuandunghcmut | 安全对齐越狱攻防 | B待确认 | 8.8 | |
llm_physical_safety_benchmark · kumitang | 安全对齐对齐安全 | B待确认 | 8.6 | |
eu-cyber-llm-benchmark-prompts · eromang | 安全对齐偏见公平 | B待确认 | 7.8 | |
manta-benchmark-questions · mycelium-ai | 安全对齐对齐安全 | B待确认 | 7.7 | |
Ko-LLM-Safety-Benchmark · UHYEL | 安全对齐对齐安全 | B待确认 | 7.7 | |
LLM-Sec-Evaluation · c01dsnap | 安全对齐对齐安全 | B待确认 | 7.1 | |
representational-collapse-llm-benchmark · wu981526092 | 安全对齐越狱攻防 | B待确认 | 6.0 | |
detoxic_benchmark · d-llm | 安全对齐对齐安全 | B待确认 | 4.7 | |
wikitext2-MIA-Benchmark · mia-llm | 安全对齐数据污染 | B待确认 | 4.7 | |
AICGSecEval · Tencent | 安全对齐对齐安全 | B待确认 | 0.0 | |
SafetyBench · thu-coai | 安全对齐对齐安全 | B待确认 | 0.0 | |
pi-detector-bench · bastion-soft | 安全对齐越狱攻防 | B待确认 | 0.0 | |
agentsocialbench · kingofspace0wzz | 安全对齐Privacy | B待确认 | 0.0 |
第 2 / 2 页
上一页下一页