BlindLoop-Evaluation
B待确认Benchmark多模态视觉语言申请
发布方:taesiri
热度8.4▼ 0.3
下载量 · 30天
12
Hugging Face
GitHub Stars
—
代码仓库
论文被引
—
Semantic Scholar
跑分模型 · 30天
—
Leaderboard results
简介
BlindLoop Evaluation The frozen evaluation cohorts for the BlindLoop paper. Each eligible generated task contributes exactly five deterministic, pixel-distinct image instances. The five rows share a task's selected question/prompt family while varying the rendered scene and gold answer as determined by the task's pixel oracle. Config Tasks Rows Documented exclusions section1_eval5 1,298 6,490 3 section2_eval5 874 4,370 1 combined_eval5 2,172 10,860 4 Each row… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/BlindLoop-Evaluation.