BenchMark·Hub

olmOCR-bench

B待确认Benchmark
多模态文档/OCR开放odc-by

发布方:allenai

热度41.0▼ 0.6
下载量 · 30天
40,965
Hugging Face
GitHub Stars
代码仓库
论文被引
124
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

olmOCR-bench olmOCR-bench is a dataset of 1,403 PDF files, plus 7,010 unit test cases that capture properties of the output that a good OCR system should have. This benchmark evaluates the ability of OCR systems to accurately convert PDF documents to markdown format while preserving critical textual and structural information. Quick links: 📃 Paper 🛠️ Code 🎮 Demo Table 1. Distribution of Test Classes by Document Source Document Source Text Present Text… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-bench.