pdfQA-Benchmark
B待确认Benchmark长上下文长文档开放mit
发布方:pdfqa
热度21.0▼ 0.1
下载量 · 30天
8,697
Hugging Face
GitHub Stars
—
代码仓库
论文被引
1
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research. The dataset is organized to support: Raw document processing research Structured extraction pipelines Retrieval-augmented QA End-to-end document reasoning systems It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Benchmark.