BenchMark·Hub

TACT

B待确认Benchmark
推理通用推理申请cc-by-nd-4.0

发布方:google

热度12.9▼ 0.1
下载量 · 30天
16
Hugging Face
GitHub Stars
代码仓库
论文被引
9
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

TACT: A Complex Numerical Reasoning Benchmark Paper - TACT: Advancing Complex Aggregative Reasoning with Information Extraction Tools Website: https://tact-benchmark.github.io Abstract: Large Language Models (LLMs) often do not perform well on queries that require the aggregation of information across texts. To better evaluate this setting and facilitate modeling efforts, we introduce TACT - Text And Calculations through Tables, a dataset crafted to evaluate LLMs'… See the full description on the dataset page: https://huggingface.co/datasets/google/TACT.