AI Arena Leaderboard

code Rankings

Arena comparison of top-performing AI models for the code benchmark.

Category
code
Date
2026-08-06
Models
50
Rank Model Score
1 claude-opus-5-max 1705
2 kimi-k3-max 1676
3 claude-opus-5-high 1669
4 qwen3.8-max 1668
5 claude-fable-5 1630
6 gpt-5.6-sol-xhigh (codex-harness) 1620
7 glm-5.2-max 1586
8 deepseek-v4-flash-high 1577
9 claude-opus-4-8-thinking 1566
10 claude-opus-4-7 1561
11 claude-opus-4-7-thinking 1557
12 grok-4.5 1549
13 claude-opus-4-6-thinking 1545
14 claude-sonnet-5-high 1542
15 claude-opus-4-8 1539
16 claude-opus-4-6 1538
17 muse-spark-1.1 1536
18 gemini-3.6-flash 1533
19 seed-2.1-pro-preview 1527
20 gpt-5.6-luna-xhigh (codex-harness) 1523
21 claude-sonnet-4-6 1523
22 gpt-5.6-terra-xhigh (codex-harness) 1522
23 qwen3.7-max-20260517 1517
24 hy3 1517
25 glm-5.1 1516
26 kimi-k2.6 1510
27 gpt-5.5-xhigh (codex-harness) 1509
28 gemini-3.5-flash-high 1507
29 claude-opus-4-5-20251101-thinking-32k 1494
30 gemini-3.5-flash 1492
31 minimax-m3 1491
32 gemini-3.5-flash-medium 1486
33 gpt-5.5-high (codex-harness) 1485
34 qwen3.6-max-preview 1478
35 mimo-v2.5-pro 1474
36 kimi-k2.7-code 1473
37 claude-opus-4-5-20251101 1467
38 deepseek-v4-pro-high-preview 1464
39 gpt-5.4-high (codex-harness) 1462
40 qwen3.6-plus 1458
41 gpt-5.5 (codex-harness) 1458
42 gemini-3.5-flash-lite 1452
43 gemini-3.1-pro-preview 1447
44 deepseek-v4-pro 1446
45 gpt-5.4-medium (codex-harness) 1444
46 gemini-3-pro 1438
47 gemini-3-flash 1438
48 kimi-k2.5-thinking 1436
49 mimo-v2.5 1435
50 glm-5 1435