LLM leaderboard

Every tracked language model by intelligence, with what it costs and how fast it answers.

ranked by intelligence · 611 models · updated Aug 30 · 51–100

#ModelIntelligence
51Motif 3
47.40
52GPT-5.6 Luna (high)
47.00
53GPT-5.6 Terra (medium)
46.80
54Gemini 3.5 Flash (medium)
46.70
55Qwen3.7 Max
46.70
56GPT-5.3 Codex (xhigh)
45.50
57MiniMax-M3
45.40
58Motif 3 (Beta)
45.30
59DeepSeek V4 Pro (Reasoning, Max Effort)
45.30
60Kimi K2.6
45.10
61Claude Opus 4.6 (Adaptive Reasoning, Max Effort)
44.90
62Qwen3.8 27B (medium)
44.50
63GPT-5.5 (low)
44.50
64Muse Spark
44.30
65Claude Opus 4.7 (Non-reasoning, High Effort)
43.90
66DeepSeek V4 Pro (Reasoning, High Effort)
43.70
67GPT-5.2 (xhigh)
43.30
68Kimi K2.7 Code
43.00
69MiMo-V2.5-Pro
42.90
70Qwen3.8 27B (low)
42.90
71Claude Sonnet 5 (Non-reasoning, High Effort)
42.60
72Inkling (xhigh)
42.30
73Hy3
42.20
74DeepSeek V4 Flash (Reasoning, Max Effort)
42.10
75Claude Opus 4.5 (Reasoning)
41.90
76GPT-5.6 Sol (Non-reasoning)
41.90
77Nex-N2-Pro
41.70
78Solar Pro 4
41.60
79MiMo-V2-Pro
41.40
80GPT-5.6 Terra (low)
41.30
81GPT-5.2 Codex (xhigh)
41.20
82Inkling Small
41.20
83Qwen3.6 Max Preview
41.10
84GLM-5.1 (Reasoning)
41.00
85GPT-5.4 mini (xhigh)
40.90
86Grok Build 0.1 0616
40.70
87GLM-5 (Reasoning)
40.60
88Gemini 3 Pro Preview (high)
40.60
89Qwen3.6 Plus
40.50
90GPT-5.4 (low)
40.20
91JT-4.1 Flash 236B A21B
39.90
92Agnes 2.5 Pro Alpha
39.70
93GPT-5.4 nano (xhigh)
39.70
94Qwen3.7 Plus
39.40
95GLM-5-Turbo
39.10
96DeepSeek V4 Flash (Reasoning, High Effort)
39.00
97MiniMax-M2.7
38.90
98GPT-5.6 Luna (medium)
38.90
99GPT-5.2 (medium)
38.90
100Claude Opus 4.6 (Non-reasoning, High Effort)
38.80

what intelligence means

A composite of nine independent evaluations covering reasoning, coding, agentic work and knowledge. Higher is better. It is versioned, so scores are comparable within a version rather than across all time.