BenchMark·Hub

飙升榜 · 7 日增速

按最近 7 天热度变化幅度排序

全部 450
  1. 1
    code_contests
    ▲ 2.645.6
  2. 2
    form-field-v1-benchmark
    ▲ 2.511.6
  3. 3
    DL3DV-Evaluation
    ▲ 2.517.2
  4. 4
    turkish_political_position_benchmark
    ▲ 2.313.8
  5. 5
    ref-annotation-benchmark
    ▲ 1.310.1
  6. 6
    CPI-benchmark
    ▲ 1.116.5
  7. 7
    MBPP
    ▲ 1.056.1
  8. 8
    TTS-Voice-Direction-Benchmark
    ▲ 0.913.0
  9. 9
    manta-benchmark-questions
    ▲ 0.87.7
  10. 10
    WikiProfile
    ▲ 0.720.5
  11. 11
    RoboDojo
    ▲ 0.630.5
  12. 12
    GSM8K
    ▲ 0.660.4
  13. 13
    Aider Code Editing
    ▲ 0.48.7
  14. 14
    IFEval
    ▲ 0.456.1
  15. 15
    LongRAG
    ▲ 0.425.5
  16. 16
    ClawBench
    ▲ 0.431.1
  17. 17
    ai2_arc
    ▲ 0.454.7
  18. 18
    mrcr
    ▲ 0.335.9
  19. 19
    HumanEval+
    ▲ 0.39.1
  20. 20
    code_generation
    ▲ 0.321.0
  21. 21
    xtreme
    ▲ 0.338.8
  22. 22
    MathVista
    ▲ 0.340.0
  23. 23
    AgentHarm
    ▲ 0.232.9
  24. 24
    DOUBLE_ENTRY_LOGGING_FOR_LLM_AGENTS_TO_DETECT_ISSUES_AT_PROD_BENCHMARK
    ▲ 0.25.7
  25. 25
    HealthBench
    ▲ 0.222.4
  26. 26
    xyzibd
    ▲ 0.221.6
  27. 27
    histone_modification_benchmark_v1
    ▲ 0.26.1
  28. 28
    RSRCC
    ▲ 0.219.5
  29. 29
    MMMU
    ▲ 0.254.4
  30. 30
    llm-refusal-evaluation
    ▲ 0.225.1
  31. 31
    BLINK
    ▲ 0.135.3
  32. 32
    AARA_Azerbaijani_LLM_Benchmark
    ▲ 0.18.6
  33. 33
    nigeria-livestock-llm-benchmark
    ▲ 0.15.0
  34. 34
    Instruction-Following-Evaluation-for-Large-Language-Models
    ▲ 0.123.9
  35. 35
    BrowseCompLongContext
    ▲ 0.120.5
  36. 36
    art
    ▲ 0.131.8
  37. 37
    scrapinghub-article-extraction-benchmark
    ▲ 0.111.4
  38. 38
    real-toxicity-prompts
    ▲ 0.141.5
  39. 39
    AlGhafa-Arabic-LLM-Benchmark-Native
    ▲ 0.131.0
  40. 40
    EnigmaEval
    ▲ 0.115.2
  41. 41
    bigcodebench-hard-domain
    ▲ 0.17.4
  42. 42
    MMMU-Pro
    ▲ 0.136.7
  43. 43
    MASK
    ▲ 0.121.3
  44. 44
    mup
    ▲ 0.110.8
  45. 45
    MMEB-V2
    ▲ 0.128.2
  46. 46
    SWE-bench Bash-only
    ▲ 0.17.1
  47. 47
    Cultural-Evaluation-Kalahi
    ▲ 0.116.5
  48. 48
    xtreme_s
    ▲ 0.125.7
  49. 49
    MedQA
    ▲ 0.138.2
  50. 50
    maze-llm-benchmark-v1
    ▲ 0.15.7
  51. 51
    ruNaturalScienceVQA
    ▲ 0.111.3
  52. 52
    Aider Polyglot
    ▲ 0.18.0
  53. 53
    Humanity's Last Exam
    ▲ 0.140.3
  54. 54
    arabic-acoustic-poetry-benchmark
    ▲ 0.17.2
  55. 55
    enhancer_annotations_benchmark_v1
    ▲ 0.15.8
  56. 56
    code_x_glue_cc_code_completion_token
    ▲ 0.116.4
  57. 57
    MultiBanana-Benchmark
    ▲ 0.123.0
  58. 58
    LongICLBench
    ▲ 0.123.9
  59. 59
    TutorEval
    ▲ 0.19.4
  60. 60
    IneqMath
    ▲ 0.123.0
  61. 61
    LongBench v2
    ▲ 0.137.7
  62. 62
    NexusRaven_API_evaluation
    ▲ 0.127.3
  63. 63
    Gender_Bias_Evaluation_Set
    ▲ 0.123.8
  64. 64
    bovine-embryo-video-evaluation
    ▲ 0.110.0
  65. 65
    TruthfulQA
    ▲ 0.145.6
  66. 66
    peer_read
    ▲ 0.125.0
  67. 67
    WMDP
    ▲ 0.136.2
  68. 68
    CritPt
    ▲ 0.125.2
  69. 69
    formal-grammar-llm-benchmark
    ▲ 0.110.1
  70. 70
    Bias-Evaluation-Turkish
    ▲ 0.124.2
  71. 71
    swag
    ▲ 0.133.9
  72. 72
    Mantis-Eval
    ▲ 0.115.1
  73. 73
    TTS-Voice-Design-Benchmark
    ▲ 0.110.3
  74. 74
    llm_physical_safety_benchmark
    ▲ 0.18.6
  75. 75
    SWE-MERA
    ▲ 0.015.1
  76. 76
    qasper
    ▲ 0.039.6
  77. 77
    TheoremQA
    ▲ 0.029.6
  78. 78
    code_x_glue_cc_code_to_code_trans
    ▲ 0.031.0
  79. 79
    TheoremExplainBench
    ▲ 0.019.0
  80. 80
    CharXiv
    ▲ 0.032.0
  81. 81
    healthbench-professional
    ▲ 0.019.4
  82. 82
    LLMs-Planning
    ▲ 0.010.3
  83. 83
    Mind2Web
    ▲ 0.011.5
  84. 84
    Vision-DeepResearch
    ▲ 0.010.8
  85. 85
    OpenLane-V2
    ▲ 0.010.8
  86. 86
    harvey-labs
    ▲ 0.011.9
  87. 87
    wiqa
    ▲ 0.025.5
  88. 88
    MATH
    ±017.1
  89. 89
    FrontierMath
    ±010.6
  90. 90
    Terminal-Bench
    ±00.0
  91. 91
    WebArena
    ±014.7
  92. 92
    OSWorld
    ±027.0
  93. 93
    τ²-bench
    ±011.8
  94. 94
    Needle in a Haystack
    ±00.0
  95. 95
    RULER
    ±013.9
  96. 96
    StrongREJECT
    ±011.6
  97. 97
    HarmBench
    ±014.3
  98. 98
    SuperCLUE
    ±00.0
  99. 99
    SimpleQA
    ±00.0
  100. 100
    ARC-AGI-2
    ±010.0
  101. 101
    ade_corpus_v2
    ±018.2
  102. 102
    NLU-few-shot-benchmark-en-de
    ±09.2
  103. 103
    llm-speed-benchmarks
    ±07.6
  104. 104
    chromatin_accessibility_benchmark_v1
    ±05.7
  105. 105
    ade_benchmark_with_llm_reason
    ±06.1
  106. 106
    detoxic_benchmark
    ±04.7
  107. 107
    Quad_benchmark_cqa_latest
    ±04.1
  108. 108
    NTEU_Multilingual_Evaluation_Dataset
    ±010.7
  109. 109
    query2query_evaluation
    ±06.8
  110. 110
    imagereward-evaluation
    ±06.8
  111. 111
    LoraxBench
    ±010.6
  112. 112
    Video-ChatGPT
    ±012.2
  113. 113
    llm-colosseum
    ±012.2
  114. 114
    PIXIU
    ±011.3
  115. 115
    codefuse-devops-eval
    ±010.8
  116. 116
    meta-agents-research-environments
    ±010.5
  117. 117
    EvaLearn
    ±010.1
  118. 118
    K12-KGraph
    ±09.7
  119. 119
    UltraEval-Audio
    ±09.6
  120. 120
    Frontier-CS
    ±09.5
  121. 121
    whichllm
    ±014.6
  122. 122
    InferenceX
    ±012.2
  123. 123
    arena-hard-auto
    ±011.6
  124. 124
    ATLAS
    ±011.5
  125. 125
    Julia-LLM-Leaderboard
    ±07.4
  126. 126
    persian-llm-eval
    ±02.7
  127. 127
    chain-of-thought-hub
    ±013.2
  128. 128
    phyre
    ±010.2
  129. 129
    Visual-CoT
    ±010.2
  130. 130
    elimination_game
    ±09.5
  131. 131
    ERQA
    ±09.4
  132. 132
    sweet_rl
    ±09.3
  133. 133
    BALROG
    ±09.3
  134. 134
    nyt-connections
    ±09.1
  135. 135
    DyCodeEval
    ±09.0
  136. 136
    BRIGHT
    ±08.9
  137. 137
    MedXpertQA
    ±08.6
  138. 138
    DeepPHY
    ±08.6
  139. 139
    LLM-RGB
    ±08.5
  140. 140
    visual-spatial-reasoning
    ±08.3
  141. 141
    cladder
    ±08.3
  142. 142
    MME-CoT
    ±08.2
  143. 143
    OlympicArena
    ±07.8
  144. 144
    clutrr
    ±07.7
  145. 145
    FunQA
    ±07.7
  146. 146
    exploitgym
    ±011.2
  147. 147
    mcpmark
    ±010.2
  148. 148
    GateMem
    ±08.4
  149. 149
    autolab
    ±08.5
  150. 150
    HaluMem
    ±08.4
  151. 151
    ScienceBoard
    ±08.2
  152. 152
    AgentKernelArena
    ±07.9
  153. 153
    BBH
    ±010.0
  154. 154
    MATH Lvl 5
    ±010.0
  155. 155
    GPQA
    ±010.0
  156. 156
    MUSR
    ±010.0
  157. 157
    Massive Text Embedding Benchmark
    ±010.0
  158. 158
    BrowseComp
    ±012.4
  159. 159
    τ²-bench
    ±024.3
  160. 160
    BFCL (Berkeley Function Calling)
    ±00.0
  161. 161
    Needle in a Haystack
    ±00.0
  162. 162
    SWE-Next
    ±015.7
  163. 163
    AgentBench
    ±00.0
  164. 164
    MingLi-Bench
    ±00.0
  165. 165
    beir
    ±00.0
  166. 166
    KernelBench
    ±00.0
  167. 167
    yet-another-applied-llm-benchmark
    ±00.0
  168. 168
    AICGSecEval
    ±00.0
  169. 169
    ParseBench
    ±00.0
  170. 170
    EnterpriseRAG-Bench
    ±00.0
  171. 171
    VibeSearchBench
    ±00.0
  172. 172
    SEED-Bench
    ±00.0
  173. 173
    android-bench
    ±00.0
  174. 174
    SafetyBench
    ±00.0
  175. 175
    terminal-bench-1
    ±00.0
  176. 176
    SWELancer-Benchmark
    ±00.0
  177. 177
    homebench
    ±00.0
  178. 178
    LLM-Game-Benchmark
    ±00.0
  179. 179
    TwinRouterBench
    ±00.0
  180. 180
    pi-detector-bench
    ±00.0
  181. 181
    glossobench
    ±00.0
  182. 182
    trajrl-bench
    ±00.0
  183. 183
    HallusionBench
    ±00.0
  184. 184
    ChartVLM
    ±00.0
  185. 185
    MMAR
    ±00.0
  186. 186
    snapbench
    ±00.0
  187. 187
    RISEBench
    ±00.0
  188. 188
    CUREBench
    ±00.0
  189. 189
    prism-bench
    ±00.0
  190. 190
    OneIG-Benchmark
    ±00.0
  191. 191
    MathNet
    ±00.0
  192. 192
    AssetOpsBench
    ±00.0
  193. 193
    VLABench
    ±00.0
  194. 194
    BrowseComp-Plus
    ±00.0
  195. 195
    EmbodiedBench
    ±00.0
  196. 196
    AutomationBench
    ±00.0
  197. 197
    merchantbench
    ±00.0
  198. 198
    MobilityBench
    ±00.0
  199. 199
    people-search-bench
    ±00.0
  200. 200
    Mind2Web-2
    ±00.0
  201. 201
    AgentRE-Bench
    ±00.0
  202. 202
    agentsocialbench
    ±00.0
  203. 203
    LiveMCPBench
    ±00.0
  204. 204
    PawBench
    ±00.0
  205. 205
    silicon-rider-bench
    ±00.0
  206. 206
    truthful_qa
    ±00.0
  207. 207
    bigbench
    ▼ 0.034.4
  208. 208
    LLMBar
    ▼ 0.011.7
  209. 209
    CMMLU
    ▼ 0.038.1
  210. 210
    MMIU-Benchmark
    ▼ 0.026.9
  211. 211
    multi_lexsum
    ▼ 0.027.1
  212. 212
    IDEAL-Scenes
    ▼ 0.015.0
  213. 213
    movie_rationales
    ▼ 0.013.9
  214. 214
    LLM-Enzyme-Kinetics-Golden-Benchmark
    ▼ 0.09.4
  215. 215
    MMMLU
    ▼ 0.044.3
  216. 216
    LiveCodeBench
    ▼ 0.043.0
  217. 217
    Thai-Semantic-Textual-Similarity-Benchmark
    ▼ 0.08.7
  218. 218
    llm-redactor-leak-benchmark
    ▼ 0.08.9
  219. 219
    ArXivSQA
    ▼ 0.012.6
  220. 220
    seahorse_summarization_evaluation
    ▼ 0.020.2
  221. 221
    NLU-Evaluation-Data-en-de
    ▼ 0.019.9
  222. 222
    xquad
    ▼ 0.036.4
  223. 223
    scicite
    ▼ 0.025.7
  224. 224
    FinanceBench
    ▼ 0.126.7
  225. 225
    llm-query-complexity-benchmark
    ▼ 0.124.0
  226. 226
    combine-llm-security-benchmark
    ▼ 0.18.8
  227. 227
    RLPR-Evaluation
    ▼ 0.121.4
  228. 228
    image_gen_ocr_evaluation_data
    ▼ 0.16.8
  229. 229
    code_x_glue_cc_cloze_testing_maxmin
    ▼ 0.128.5
  230. 230
    humanevalpack
    ▼ 0.133.8
  231. 231
    GPQA Diamond
    ▼ 0.146.0
  232. 232
    genomics-long-range-benchmark
    ▼ 0.119.6
  233. 233
    MPB_Missing-Premise-Benchmark
    ▼ 0.117.1
  234. 234
    wmt-sqm-human-evaluation
    ▼ 0.19.5
  235. 235
    code_x_glue_tt_text_to_text
    ▼ 0.127.8
  236. 236
    IndicGenBench_crosssum_in
    ▼ 0.122.8
  237. 237
    social_i_qa
    ▼ 0.124.2
  238. 238
    bigcodereward
    ▼ 0.114.0
  239. 239
    legal-llm-benchmark
    ▼ 0.112.6
  240. 240
    bc-finance-llm-benchmark
    ▼ 0.19.7
  241. 241
    equity_evaluation_corpus
    ▼ 0.110.5
  242. 242
    code_x_glue_ct_code_to_text
    ▼ 0.136.5
  243. 243
    scitail
    ▼ 0.118.7
  244. 244
    bigcodebench-domain
    ▼ 0.16.2
  245. 245
    GenAI-Bench
    ▼ 0.133.5
  246. 246
    MathVerse
    ▼ 0.133.7
  247. 247
    speech_commands
    ▼ 0.136.6
  248. 248
    Video-MME
    ▼ 0.151.4
  249. 249
    representational-collapse-llm-benchmark
    ▼ 0.16.0
  250. 250
    LLM-Ribozyme-Kinetics-Golden-Benchmark
    ▼ 0.16.0
  251. 251
    MERA
    ▼ 0.118.5
  252. 252
    climate-evaluation
    ▼ 0.122.2
  253. 253
    code_evaluation_prompts
    ▼ 0.111.5
  254. 254
    code_x_glue_tc_nl_code_search_adv
    ▼ 0.129.9
  255. 255
    M-BEIR
    ▼ 0.130.0
  256. 256
    VideoScore-Bench
    ▼ 0.122.0
  257. 257
    VideoEval-Pro
    ▼ 0.121.7
  258. 258
    frames-benchmark
    ▼ 0.135.5
  259. 259
    Video-MME-v2
    ▼ 0.128.6
  260. 260
    STRABLE-benchmark
    ▼ 0.118.2
  261. 261
    cvdp-benchmark-dataset
    ▼ 0.119.3
  262. 262
    lm1b
    ▼ 0.132.1
  263. 263
    LLM-ABAP-Code-Generation-Benchmark
    ▼ 0.19.2
  264. 264
    LLM_Benchmark
    ▼ 0.121.4
  265. 265
    humaine-evaluation-dataset
    ▼ 0.112.2
  266. 266
    japanese-image-classification-evaluation-dataset
    ▼ 0.111.7
  267. 267
    svq
    ▼ 0.121.4
  268. 268
    simpleqa-verified
    ▼ 0.128.8
  269. 269
    VisPhyBench-Data
    ▼ 0.17.1
  270. 270
    benchmark-research
    ▼ 0.116.0
  271. 271
    local-llm-speed-benchmarks
    ▼ 0.114.9
  272. 272
    aya_evaluation_suite
    ▼ 0.131.0
  273. 273
    fleurs
    ▼ 0.148.6
  274. 274
    TACT
    ▼ 0.112.9
  275. 275
    csabstruct
    ▼ 0.121.6
  276. 276
    LiveBench
    ▼ 0.128.7
  277. 277
    Food_Portion_Benchmark
    ▼ 0.130.6
  278. 278
    IndicGenBench_xorqa_in
    ▼ 0.122.8
  279. 279
    MMLU-Pro
    ▼ 0.166.0
  280. 280
    hendrycks-MATH-benchmark
    ▼ 0.140.0
  281. 281
    pdfQA-Benchmark
    ▼ 0.121.0
  282. 282
    QZhou-Flowchart-QA-Benchmark
    ▼ 0.17.5
  283. 283
    JobBERT-evaluation-dataset
    ▼ 0.118.9
  284. 284
    bigcodebench-perf
    ▼ 0.16.6
  285. 285
    SWE-bench Lite
    ▼ 0.149.9
  286. 286
    SWE-bench Multilingual
    ▼ 0.15.0
  287. 287
    scitldr
    ▼ 0.128.0
  288. 288
    FinQA
    ▼ 0.115.6
  289. 289
    BEAR-benchmark
    ▼ 0.115.0
  290. 290
    LEXam
    ▼ 0.127.0
  291. 291
    German-RAG-LLM-HARD-BENCHMARK
    ▼ 0.121.4
  292. 292
    Ko-LLM-Safety-Benchmark
    ▼ 0.17.7
  293. 293
    ja-vicuna-qa-benchmark
    ▼ 0.16.4
  294. 294
    WER_Evaluation_For_TTS
    ▼ 0.17.3
  295. 295
    code_x_glue_cc_code_completion_line
    ▼ 0.115.4
  296. 296
    spiqa
    ▼ 0.127.0
  297. 297
    aqua_rat
    ▼ 0.139.4
  298. 298
    pg19
    ▼ 0.135.1
  299. 299
    gooaq
    ▼ 0.122.1
  300. 300
    ropes
    ▼ 0.133.4
  301. 301
    MMLU
    ▼ 0.151.5
  302. 302
    finno-ugric-benchmark
    ▼ 0.19.6
  303. 303
    TARA_Turkish_LLM_Benchmark
    ▼ 0.112.0
  304. 304
    autocodearena-v0
    ▼ 0.114.1
  305. 305
    HRVideoBench
    ▼ 0.118.1
  306. 306
    PixelWorld
    ▼ 0.115.8
  307. 307
    EditReward-Bench
    ▼ 0.120.2
  308. 308
    DocVQA
    ▼ 0.239.9
  309. 309
    llm-system-prompts-benchmark
    ▼ 0.222.4
  310. 310
    German-RAG-LLM-EASY-BENCHMARK
    ▼ 0.220.6
  311. 311
    llm-medical-reasoning-steps-benchmark
    ▼ 0.28.8
  312. 312
    cpg_methylation_benchmark_v1
    ▼ 0.26.0
  313. 313
    wikitext2-MIA-Benchmark
    ▼ 0.24.7
  314. 314
    vizdoom-llm-inverse-dynamics-benchmark
    ▼ 0.24.7
  315. 315
    NuosuBburma-OCR-Evaluation-Set
    ▼ 0.213.4
  316. 316
    eka-medical-asr-evaluation-dataset
    ▼ 0.215.7
  317. 317
    wmt-da-human-evaluation-long-context
    ▼ 0.213.4
  318. 318
    wmt-mqm-human-evaluation
    ▼ 0.211.2
  319. 319
    hippocorpus
    ▼ 0.211.8
  320. 320
    wmt22_african
    ▼ 0.215.9
  321. 321
    BigCodeBench
    ▼ 0.240.3
  322. 322
    GDPval
    ▼ 0.237.6
  323. 323
    ruCommonVQA
    ▼ 0.211.3
  324. 324
    frontierscience
    ▼ 0.227.9
  325. 325
    code_x_glue_cc_clone_detection_big_clone_bench
    ▼ 0.216.7
  326. 326
    AIME25
    ▼ 0.214.5
  327. 327
    ImagenWorld
    ▼ 0.217.6
  328. 328
    HELMET
    ▼ 0.227.6
  329. 329
    speculator_benchmarks
    ▼ 0.218.6
  330. 330
    onprem-llm-benchmark
    ▼ 0.217.1
  331. 331
    RAG-Evaluation-Dataset-KO
    ▼ 0.217.1
  332. 332
    granola-entity-questions
    ▼ 0.220.6
  333. 333
    code_x_glue_cc_code_refinement
    ▼ 0.230.2
  334. 334
    Crab-role-playing-evaluation-benchmark
    ▼ 0.214.1
  335. 335
    reveal
    ▼ 0.222.1
  336. 336
    FACTS-grounding-public
    ▼ 0.218.7
  337. 337
    ConvApparel
    ▼ 0.214.4
  338. 338
    VisPlotBench
    ▼ 0.214.4
  339. 339
    GAIA-2
    ▼ 0.229.3
  340. 340
    narrativeqa_manual
    ▼ 0.225.8
  341. 341
    LLM-WikiRace-Benchmark
    ▼ 0.28.3
  342. 342
    El-TARA_Spanish_LLM_Benchmark
    ▼ 0.25.1
  343. 343
    narrativeqa
    ▼ 0.236.4
  344. 344
    social_bias_frames
    ▼ 0.216.3
  345. 345
    DL3DV-Benchmark
    ▼ 0.225.9
  346. 346
    music-off-policy-evaluation-benchmark
    ▼ 0.212.1
  347. 347
    code_x_glue_tc_text_to_code
    ▼ 0.227.3
  348. 348
    mup-full
    ▼ 0.28.6
  349. 349
    ValuePrism
    ▼ 0.223.2
  350. 350
    snips_built_in_intents
    ▼ 0.230.2
  351. 351
    LLM-Sec-Evaluation
    ▼ 0.27.1
  352. 352
    boolq
    ▼ 0.242.5
  353. 353
    IndicGenBench_flores_in
    ▼ 0.223.9
  354. 354
    VideoFeedback2
    ▼ 0.217.7
  355. 355
    Additive-Manufacturing-Benchmark
    ▼ 0.216.9
  356. 356
    serbian-llm-benchmark
    ▼ 0.215.6
  357. 357
    MMDocIR_Evaluation_Dataset
    ▼ 0.222.9
  358. 358
    novalad-evaluation
    ▼ 0.222.2
  359. 359
    Entity-Imagen
    ▼ 0.221.7
  360. 360
    StructEval
    ▼ 0.222.4
  361. 361
    test_generation
    ▼ 0.217.7
  362. 362
    cosmos_qa
    ▼ 0.234.0
  363. 363
    mslr2022
    ▼ 0.214.8
  364. 364
    kazakhstan-sociology-llm-benchmark
    ▼ 0.38.8
  365. 365
    BlindLoop-Evaluation
    ▼ 0.38.4
  366. 366
    SimRelUz_semantic_evaluation_dataset
    ▼ 0.36.1
  367. 367
    math_dataset
    ▼ 0.336.1
  368. 368
    soda
    ▼ 0.331.2
  369. 369
    bigcodebench-hard
    ▼ 0.319.5
  370. 370
    docci
    ▼ 0.326.8
  371. 371
    curriculum_benchmark
    ▼ 0.38.4
  372. 372
    eu-cyber-llm-benchmark-prompts
    ▼ 0.37.8
  373. 373
    code_x_glue_cc_clone_detection_poj104
    ▼ 0.314.6
  374. 374
    H2VU-Benchmark
    ▼ 0.315.5
  375. 375
    lila
    ▼ 0.322.1
  376. 376
    coconot
    ▼ 0.330.7
  377. 377
    AIME 2025
    ▼ 0.325.5
  378. 378
    CHML-real-ecp-local-llm-audit-benchmark
    ▼ 0.39.6
  379. 379
    LogiTraj-Benchmark
    ▼ 0.314.9
  380. 380
    llm-smartrouter-benchmark
    ▼ 0.38.3
  381. 381
    Fraud-R1-LLM-Defense-Fraud-Benchmark
    ▼ 0.317.6
  382. 382
    vllm_safety_evaluation
    ▼ 0.318.9
  383. 383
    C-Eval
    ▼ 0.342.5
  384. 384
    llm-cold-start-benchmark
    ▼ 0.39.0
  385. 385
    graphwalks
    ▼ 0.320.2
  386. 386
    MEGA-Bench
    ▼ 0.323.8
  387. 387
    MBPP+
    ▼ 0.38.5
  388. 388
    openai-moderation-api-evaluation
    ▼ 0.332.1
  389. 389
    VideoFeedback
    ▼ 0.327.0
  390. 390
    SWE-bench
    ▼ 0.344.5
  391. 391
    data-product-benchmark
    ▼ 0.319.1
  392. 392
    coverbench
    ▼ 0.312.6
  393. 393
    prosocial-dialog
    ▼ 0.328.8
  394. 394
    AIGC-Detection-Benchmark
    ▼ 0.314.6
  395. 395
    qasc
    ▼ 0.336.1
  396. 396
    SWE-bench Test
    ▼ 0.35.1
  397. 397
    Aider Refactor
    ▼ 0.35.1
  398. 398
    HumanEval
    ▼ 0.361.5
  399. 399
    PDE_Inverse_Problem_Benchmarking
    ▼ 0.321.7
  400. 400
    plant-genomic-benchmark
    ▼ 0.321.0
  401. 401
    denoising-impact-evaluation-dataset
    ▼ 0.311.9
  402. 402
    IndicGenBench_xquad_in
    ▼ 0.323.5
  403. 403
    hle-rolling
    ▼ 0.315.9
  404. 404
    zest
    ▼ 0.321.3
  405. 405
    code_x_glue_cc_cloze_testing_all
    ▼ 0.429.7
  406. 406
    mittens
    ▼ 0.410.9
  407. 407
    winogrande
    ▼ 0.429.4
  408. 408
    quoref
    ▼ 0.428.5
  409. 409
    LegalBench
    ▼ 0.436.0
  410. 410
    WEIRD
    ▼ 0.412.2
  411. 411
    vlm_evaluation_v1.0
    ▼ 0.425.2
  412. 412
    imagenet-o
    ▼ 0.49.7
  413. 413
    bigcodebench-tool
    ▼ 0.45.1
  414. 414
    tabular-benchmark
    ▼ 0.45.0
  415. 415
    Chat2Workflow-Evaluation
    ▼ 0.416.3
  416. 416
    break_data
    ▼ 0.415.4
  417. 417
    LitSearch
    ▼ 0.417.5
  418. 418
    BeaverTails-Evaluation
    ▼ 0.429.7
  419. 419
    MMEB-eval
    ▼ 0.528.8
  420. 420
    sensory-awareness-benchmark
    ▼ 0.511.9
  421. 421
    turkish_cyber_security_controls_benchmark
    ▼ 0.513.7
  422. 422
    lince
    ▼ 0.515.1
  423. 423
    common_gen
    ▼ 0.525.3
  424. 424
    SWE-bench Multimodal
    ▼ 0.525.4
  425. 425
    civil_comments
    ▼ 0.535.5
  426. 426
    ruCLEVR
    ▼ 0.512.1
  427. 427
    SWE-QA-Pro-Bench
    ▼ 0.518.6
  428. 428
    deepsearchqa
    ▼ 0.628.4
  429. 429
    GAIA
    ▼ 0.646.2
  430. 430
    BFCL (Berkeley Function Calling)
    ▼ 0.68.9
  431. 431
    olmOCR-bench
    ▼ 0.641.0
  432. 432
    wildjailbreak
    ▼ 0.736.4
  433. 433
    tw-legal-benchmark-v2
    ▼ 0.715.3
  434. 434
    MusicCaps
    ▼ 0.935.2
  435. 435
    rag_instruct_benchmark_tester
    ▼ 1.016.7
  436. 436
    miriad-benchmark-200k
    ▼ 1.115.4
  437. 437
    SWE-bench Verified
    ▼ 1.258.9
  438. 438
    svg-benchmark
    ▼ 1.223.6
  439. 439
    SenseNova-Vision-Benchmark
    ▼ 1.213.0
  440. 440
    multilingual-queries-2026
    ▼ 1.413.4
  441. 441
    prompt-injections-benchmark
    ▼ 1.420.2
  442. 442
    chinese-writing-benchmark
    ▼ 1.613.8
  443. 443
    indic-queries-2026
    ▼ 1.714.2
  444. 444
    ragu_benchmarks
    ▼ 1.710.9
  445. 445
    red-team-appsec-benchmark
    ▼ 1.712.7
  446. 446
    quac
    ▼ 1.833.1
  447. 447
    cybersec-retrieval-benchmark
    ▼ 1.812.2
  448. 448
    sciq
    ▼ 2.945.2
  449. 449
    joyo-kanji-yomi-benchmark-parakeet
    ▼ 3.018.2
  450. 450
    WildBench
    ▼ 3.432.5