BenchMark·Hub

LLM-Game-Benchmark

B待确认Benchmark
推理游戏推理开放

发布方:research-outcome

热度0.0±0
下载量 · 30天
Hugging Face
GitHub Stars
代码仓库
论文被引
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard