AI Arena Leaderboard

code Rankings

Arena comparison of top-performing AI models for the code benchmark.

Category
code
Date
2026-07-30
Models
104
Rank Model Score
1 claude-opus-5-max 1712
2 kimi-k3-max 1682
3 claude-opus-5-high 1669
4 claude-fable-5 1628
5 gpt-5.6-sol-xhigh (codex-harness) 1623
6 glm-5.2-max 1588
7 claude-opus-4-8-thinking 1568
8 claude-opus-4-7 1560
9 claude-opus-4-7-thinking 1556
10 grok-4.5 1550
11 claude-opus-4-6-thinking 1546
12 claude-sonnet-5-high 1544
13 claude-opus-4-8 1539
14 claude-opus-4-6 1538
15 muse-spark-1.1 1536
16 gemini-3.6-flash 1528
17 seed-2.1-pro-preview 1527
18 claude-sonnet-4-6 1524
19 glm-5.1 1518
20 hy3 1517
21 qwen3.7-max-20260517 1517
22 kimi-k2.6 1510
23 gpt-5.5-xhigh (codex-harness) 1507
24 claude-opus-4-5-20251101-thinking-32k 1494
25 minimax-m3 1494
26 gemini-3.5-flash 1492
27 gemini-3.5-flash-medium 1486
28 gpt-5.5-high (codex-harness) 1485
29 qwen3.6-max-preview 1478
30 mimo-v2.5-pro 1474
31 kimi-k2.7-code 1473
32 claude-opus-4-5-20251101 1467
33 deepseek-v4-pro-thinking 1464
34 gpt-5.4-high (codex-harness) 1462
35 qwen3.6-plus 1458
36 gpt-5.5 (codex-harness) 1455
37 gemini-3.5-flash-lite 1454
38 deepseek-v4-pro 1447
39 gemini-3.1-pro-preview 1446
40 gpt-5.4-medium (codex-harness) 1444
41 gemini-3-flash 1438
42 gemini-3-pro 1438
43 kimi-k2.5-thinking 1436
44 mimo-v2.5 1436
45 glm-5 1435
46 glm-4.7 1433
47 mimo-v2-pro 1433
48 gpt-5.2 1419
49 gpt-5-medium 1419
50 inkling 1417
51 gpt-5.3-codex (codex-harness) 1409
52 kimi-k2.5-instant 1405
53 qwen3.5-397b-a17b 1401
54 glm-5v-turbo 1400
55 gpt-5.4-mini-high 1399
56 minimax-m2.7 1398
57 claude-sonnet-4-5-20250929-thinking-32k 1392
58 gpt-5.1-medium 1390
59 gpt-5.4 1390
60 claude-opus-4-1-20250805 1389
61 minimax-m2.1-preview 1388
62 minimax-m2.5 1386
63 claude-sonnet-4-5-20250929 1385
64 gemini-3-flash (thinking-minimal) 1384
65 grok-4.20-beta-0309-reasoning 1374
66 gpt-5.3-codex (codex-harness) 1370
67 gemma-4-26b-a4b 1366
68 gemma-4-31b 1364
69 deepseek-v3.2-thinking 1361
70 qwen3.5-122b-a10b 1360
71 grok-4.3 1358
72 qwen3.5-27b 1357
73 hunyuan-hy3-preview 1357
74 laguna-m.1 1349
75 gpt-5.1 1341
76 glm-4.6 1339
77 gpt-5.2-codex 1338
78 gpt-5.1-codex 1336
79 mimo-v2-flash (non-thinking) 1331
80 claude-haiku-4-5-20251001 1325
81 deepseek-v3.2 1323
82 kimi-k2-thinking-turbo 1322
83 laguna-xs.2 1304
84 minimax-m2 1297
85 mimo-v2-flash (thinking) 1291
86 qwen3-coder-480b-a35b-instruct 1272
87 deepseek-v3.2-exp 1272
88 mistral-medium-3.5 1267
89 gemini-3.1-flash-lite-preview 1256
90 KAT-Coder-Pro-V1 1255
91 qwen3.5-35b-a3b 1251
92 gpt-5.1-codex-mini 1244
93 grok-4-1-fast-reasoning 1240
94 trinity-large-thinking 1239
95 qwen3.5-flash 1238
96 mistral-large-3 1230
97 gemini-2.5-pro 1224
98 grok-4.1-thinking 1210
99 granite-4.1-8b 1195
100 devstral-2 1194
101 mercury-2 1166
102 grok-code-fast-1 1163
103 grok-4-fast-reasoning 1160
104 devstral-medium-2507 1079