BenchMark·Hub

LLM_Benchmark

B待确认Benchmark
语言/语义多语低资源开放cc-by-4.0

发布方:Bokhbat

热度21.4▼ 0.1
下载量 · 30天
69
Hugging Face
GitHub Stars
1,570
代码仓库
论文被引
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

Mongolian LLM Benchmark A multi-task evaluation benchmark for large language models on the Mongolian language (Cyrillic script). Six task configurations cover open-ended QA, multiple-choice, code generation, instruction following, math, and culturally grounded knowledge. Configurations Config Rows Format Key fields culture 150 Multiple choice (A–D) prompt, options, answer, source_url math 150 Numeric / short answer prompt, answer, accepted_formats, rationale… See the full description on the dataset page: https://huggingface.co/datasets/Bokhbat/LLM_Benchmark.