BenchMark·Hub

TwinRouterBench

B待确认Benchmark
智能体工具调用开放

发布方:CommonstackAI

热度0.0±0
下载量 · 30天
Hugging Face
GitHub Stars
代码仓库
论文被引
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

Per-step LLM routing benchmark with 970 static labels, live SWE-bench evaluation, an open data pipeline, and a public leaderboard.