BenchMark·Hub

MMDocIR_Evaluation_Dataset

B待确认Benchmark
长上下文长文档开放apache-2.0

发布方:MMDocIR

热度22.9▼ 0.2
下载量 · 30天
476
Hugging Face
GitHub Stars
代码仓库
论文被引
52
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

Evaluation Datasets Evaluation Set Overview MMDocIR evaluation set includes 313 long documents averaging 65.1 pages, categorized into ten main domains: research reports, administration&industry, tutorials&workshops, academic papers, brochures, financial reports, guidebooks, government documents, laws, and news articles. Different domains feature distinct distributions of multi-modal information. Overall, the modality distribution is: Text (60.4%), Image (18.8%), Table… See the full description on the dataset page: https://huggingface.co/datasets/MMDocIR/MMDocIR_Evaluation_Dataset.