LLM leaderboard

Every tracked language model by intelligence, with what it costs and how fast it answers.

ranked by intelligence · 611 models · updated Aug 30 · 201–250

#ModelIntelligence
201GPT-5.2 (Non-reasoning)
26.50
202Gemma 4 26B A4B (Reasoning)
26.10
203o4-mini (high)
26.10
204Claude 4 Opus (Non-reasoning)
26.00
205Step 3.5 Flash
26.00
206Claude 4 Sonnet (Non-reasoning)
26.00
207Gemini 2.5 Pro
25.90
208DeepSeek V3.2 Exp (Reasoning)
25.90
209GPT-5 mini (high)
25.80
210Nemotron 3 Super 120B A12B (Reasoning)
25.70
211Gemini 3.1 Flash-Lite
25.60
212Qwen3 Max Thinking (Preview)
25.50
213DeepSeek V3.2 (Non-reasoning)
25.10
214MiMo-V2-Flash (Non-reasoning)
25.10
215Grok 4.3 (Non-reasoning)
25.00
216Qwen3.6 35B A3B (Non-reasoning)
24.60
217Qwen3 Max
24.50
218Ling 3.0 Tiny
24.50
219Qwen3.5 35B A3B (Non-reasoning)
24.30
220Gemini 2.5 Flash Preview (Sep '25) (Reasoning)
24.20
221Claude 4.5 Haiku (Non-reasoning)
24.10
222gpt-oss-120b (high)
24.10
223Kimi K2 0905
24.00
224o1
23.90
225Claude 3.7 Sonnet (Non-reasoning)
23.90
226Granite 4.2 30B
23.70
227Nemotron 3.5 Lightning
23.60
228Gemini 2.5 Pro Preview (Mar' 25)
23.40
229GLM-4.6 (Non-reasoning)
23.40
230GLM-4.7-Flash (Reasoning)
23.30
231Grok 4.20 0309 (Non-reasoning)
22.90
232Grok 3 mini Reasoning (high)
22.90
233Command A+
22.80
234Gemini 2.5 Pro Preview (May' 25)
22.70
235DeepSeek V3.2 Speciale
22.60
236K-EXAONE (Reasoning)
22.50
237ERNIE 5.0 Thinking Preview
22.30
238Gemma 4 31B (Non-reasoning)
22.30
239Grok 4.20 0309 v2 (Non-reasoning)
22.20
240Gemma 4 12B (Reasoning)
22.20
241Nova 2.0 Pro Preview (medium)
22.10
242Grok Code Fast 1
22.00
243Mercury 2
21.90
244Qwen3.5 9B (Reasoning)
21.80
245DeepSeek V3.2 Exp (Non-reasoning)
21.70
246DeepSeek V3.1 Terminus (Non-reasoning)
21.70
247Apriel-v1.5-15B-Thinker
21.60
248DeepSeek V3.1 (Non-reasoning)
21.40
249Nova 2.0 Omni (medium)
21.30
250Qwen3 Coder Next
21.30

what intelligence means

A composite of nine independent evaluations covering reasoning, coding, agentic work and knowledge. Higher is better. It is versioned, so scores are comparable within a version rather than across all time.