healthbench-professional
B待确认Benchmark垂直领域医疗开放mit
发布方:openai
热度19.4▲ 0.0
下载量 · 30天
1,720
Hugging Face
GitHub Stars
—
代码仓库
论文被引
—
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
Contains the data for the HealthBench Professional eval. Each example contains: conversation: list of user / assistant messages, ending in a user message rubric_items: list of rubric items, each containing criterion_text and points use_case: one of consult, writing, or research type: one of good_faith or red_teaming difficulty: physician-assigned difficulty rating (difficult for Likert 1-2, typical for Likert 3-7) specialty: medical specialty or sub-specialty physician_response: response… See the full description on the dataset page: https://huggingface.co/datasets/openai/healthbench-professional.