BenchMark·Hub

ClawBench

B待确认Benchmark
智能体网页导航开放apache-2.0

发布方:TIGER-Lab

热度31.1▲ 0.4
下载量 · 30天
1,112
Hugging Face
GitHub Stars
597
代码仓库
论文被引
20
Semantic Scholar
跑分模型 · 30天
Leaderboard results

简介

ClawBench Dataset ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. |💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website | 🚀 What's New [2026.05.12] Added the V2 corpus (130 newer tasks across 63 platforms) and 7 new models judged with… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ClawBench.