BenchMark·Hub

TruthfulQA

S高可信Benchmark
知识问答事实性TruthfulQA 家族未知apache-2.0

发布方:多家

热度45.6▲ 0.1
下载量 · 30天
117,833
Hugging Face
GitHub Stars
代码仓库
论文被引
3,810
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

Dataset Card for truthful_qa Dataset Summary TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/truthfulqa/truthful_qa.