BenchMark·Hub

WildBench

B待确认Benchmark
知识问答通用问答开放

发布方:allenai

热度32.5▼ 3.4
下载量 · 30天
1,890
Hugging Face
GitHub Stars
代码仓库
论文被引
248
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

🦁 WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild Loading from datasets import load_dataset wb_data = load_dataset("allenai/WildBench", "v2", split="test") Quick Links: HF Leaderboard HF Dataset Github Dataset Description License: CC BY Language(s) (NLP): English Point of Contact: Yuchen Lin WildBench is a subset of WildChat. The use of WildChat data to cause harm is strictly prohibited. Data… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildBench.