BenchMark·Hub

simpleqa-verified

B待确认Benchmark
知识问答事实性SimpleQA 家族开放mit

发布方:google

热度28.8▼ 0.1
下载量 · 30天
5,646
Hugging Face
GitHub Stars
代码仓库
论文被引
39
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

SimpleQA Verified A 1,000-prompt factuality benchmark from Google DeepMind and Google Research, designed to reliably evaluate LLM parametric knowledge. ▶ SimpleQA Verified Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark SimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research… See the full description on the dataset page: https://huggingface.co/datasets/google/simpleqa-verified.