AgentHarm
S高可信Benchmark安全对齐越狱攻防开放other
发布方:UK AISI
热度32.9▲ 0.2
下载量 · 30天
4,598
Hugging Face
GitHub Stars
—
代码仓库
论文被引
368
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Maksym Andriushchenko1,†,*, Alexandra Souly2,* Mateusz Dziemian1, Derek Duenas1, Maxwell Lin1, Justin Wang1, Dan Hendrycks1,§, Andy Zou1,¶,§, Zico Kolter1,¶, Matt Fredrikson1,¶,* Eric Winsor2, Jerome Wynne2, Yarin Gal2,♯, Xander Davies2,♯,* 1Gray Swan AI, 2UK AI Safety Institute, *Core Contributor †EPFL, §Center for AI Safety, ¶Carnegie Mellon University, ♯University of Oxford Paper:… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/AgentHarm.