AI Arena Leaderboard

code Rankings

Arena comparison of top-performing AI models for the code benchmark.

Category
code
Date
2026-08-11
Models
50
Rank Model Score
1 claude-opus-5-max 1692
2 kimi-k3-max 1674
3 qwen3.8-max 1671
4 claude-opus-5-high 1667
5 claude-fable-5 1628
6 gpt-5.6-sol-xhigh (codex-harness) 1623
7 glm-5.2-max 1588
8 deepseek-v4-flash-high 1585
9 claude-opus-4-8-thinking 1564
10 claude-opus-4-7-thinking 1558
11 claude-opus-4-7 1558
12 grok-4.5 1554
13 claude-opus-4-6-thinking 1545
14 claude-sonnet-5-high 1542
15 claude-opus-4-8 1540
16 claude-opus-4-6 1538
17 muse-spark-1.1 1537
18 gemini-3.6-flash 1536
19 muse-spark-1.2 (xHigh) 1535
20 hy3 1524
21 claude-sonnet-4-6 1524
22 gpt-5.6-terra-xhigh (codex-harness) 1524
23 seed-2.1-pro-preview 1522
24 gpt-5.6-luna-xhigh (codex-harness) 1519
25 qwen3.7-max-20260517 1517
26 glm-5.1 1511
27 kimi-k2.6 1509
28 gpt-5.5-xhigh (codex-harness) 1508
29 gemini-3.5-flash-high 1506
30 claude-opus-4-5-20251101-thinking-32k 1495
31 gemini-3.5-flash 1492
32 minimax-m3 1491
33 gemini-3.5-flash-medium 1487
34 gpt-5.5-high (codex-harness) 1486
35 qwen3.6-max-preview 1479
36 mimo-v2.5-pro 1474
37 kimi-k2.7-code 1474
38 claude-opus-4-5-20251101 1468
39 deepseek-v4-pro-high-preview 1464
40 gpt-5.4-high (codex-harness) 1463
41 qwen3.6-plus 1459
42 gpt-5.5 (codex-harness) 1458
43 gemini-3.5-flash-lite 1450
44 gemini-3.1-pro-preview 1447
45 deepseek-v4-pro 1445
46 gpt-5.4-medium (codex-harness) 1442
47 gemini-3-flash 1438
48 mimo-v2.5 1438
49 gemini-3-pro 1438
50 kimi-k2.5-thinking 1436