BenchMark·Hub

vlm_evaluation_v1.0

B待确认Benchmark
多模态视觉语言开放

发布方:VLABench

热度25.2▼ 0.4
下载量 · 30天
2,282
Hugging Face
GitHub Stars
代码仓库
论文被引
177
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

Datacard This dataset is the evaluation VLM dataset used in VLABench. It is designed to evaluate the planning capabilities of Vision-Language Models (VLMs) in embodied scenarios. Source Project Page: https://vlabench.github.io/ Arxiv Paper: https://arxiv.org/abs/2412.18194 Code: https://github.com/OpenMOSS/VLABench Uses The dataset structure is as follows: vlm_evaluation_v1.0/ ├── CommenSence/ ├── add_condiment_common_sense/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlm_evaluation_v1.0.