olmOCR-bench
B待确认Benchmark多模态文档/OCR开放odc-by
发布方:allenai
热度41.0▼ 0.6
下载量 · 30天
40,965
Hugging Face
GitHub Stars
—
代码仓库
论文被引
124
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
olmOCR-bench olmOCR-bench is a dataset of 1,403 PDF files, plus 7,010 unit test cases that capture properties of the output that a good OCR system should have. This benchmark evaluates the ability of OCR systems to accurately convert PDF documents to markdown format while preserving critical textual and structural information. Quick links: 📃 Paper 🛠️ Code 🎮 Demo Table 1. Distribution of Test Classes by Document Source Document Source Text Present Text… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-bench.