WildBench
B待确认Benchmark知识问答通用问答开放
发布方:allenai
热度32.5▼ 3.4
下载量 · 30天
1,890
Hugging Face
GitHub Stars
—
代码仓库
论文被引
248
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
🦁 WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild Loading from datasets import load_dataset wb_data = load_dataset("allenai/WildBench", "v2", split="test") Quick Links: HF Leaderboard HF Dataset Github Dataset Description License: CC BY Language(s) (NLP): English Point of Contact: Yuchen Lin WildBench is a subset of WildChat. The use of WildChat data to cause harm is strictly prohibited. Data… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildBench.