BenchMark·Hub

BlindLoop-Evaluation

B待确认Benchmark
多模态视觉语言申请

发布方:taesiri

热度8.4▼ 0.3
下载量 · 30天
12
Hugging Face
GitHub Stars
代码仓库
论文被引
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

BlindLoop Evaluation The frozen evaluation cohorts for the BlindLoop paper. Each eligible generated task contributes exactly five deterministic, pixel-distinct image instances. The five rows share a task's selected question/prompt family while varying the rendered scene and gold answer as determined by the task's pixel oracle. Config Tasks Rows Documented exclusions section1_eval5 1,298 6,490 3 section2_eval5 874 4,370 1 combined_eval5 2,172 10,860 4 Each row… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/BlindLoop-Evaluation.