AI Arena Leaderboard

code Rankings

Arena comparison of top-performing AI models for the code benchmark.

Category
code
Date
2026-07-29
Models
50
Rank Model Score
1 claude-opus-5-max 1712
2 kimi-k3-max 1682
3 claude-opus-5-high 1669
4 claude-fable-5 1628
5 gpt-5.6-sol-xhigh (codex-harness) 1623
6 glm-5.2-max 1588
7 claude-opus-4-8-thinking 1568
8 claude-opus-4-7 1560
9 claude-opus-4-7-thinking 1556
10 grok-4.5 1550
11 claude-opus-4-6-thinking 1546
12 claude-sonnet-5-high 1544
13 claude-opus-4-8 1539
14 claude-opus-4-6 1538
15 muse-spark-1.1 1536
16 gemini-3.6-flash 1528
17 seed-2.1-pro-preview 1527
18 claude-sonnet-4-6 1524
19 glm-5.1 1518
20 hy3 1517
21 qwen3.7-max-20260517 1517
22 kimi-k2.6 1510
23 gpt-5.5-xhigh (codex-harness) 1507
24 claude-opus-4-5-20251101-thinking-32k 1494
25 minimax-m3 1494
26 gemini-3.5-flash 1492
27 gemini-3.5-flash-medium 1486
28 gpt-5.5-high (codex-harness) 1485
29 qwen3.6-max-preview 1478
30 mimo-v2.5-pro 1474
31 kimi-k2.7-code 1473
32 claude-opus-4-5-20251101 1467
33 deepseek-v4-pro-thinking 1464
34 gpt-5.4-high (codex-harness) 1462
35 qwen3.6-plus 1458
36 gpt-5.5 (codex-harness) 1455
37 gemini-3.5-flash-lite 1454
38 deepseek-v4-pro 1447
39 gemini-3.1-pro-preview 1446
40 gpt-5.4-medium (codex-harness) 1444
41 gemini-3-flash 1438
42 gemini-3-pro 1438
43 kimi-k2.5-thinking 1436
44 mimo-v2.5 1436
45 glm-5 1435
46 glm-4.7 1433
47 mimo-v2-pro 1433
48 gpt-5.2 1419
49 gpt-5-medium 1419
50 inkling 1417