LLM leaderboard

Every tracked language model by intelligence, with what it costs and how fast it answers.

ranked by intelligence · 611 models · updated Aug 30 · 301–350

#ModelIntelligence
301Grok 4.1 Fast (Non-reasoning)
17.00
302K-EXAONE (Non-reasoning)
16.90
303Qwen3 Next 80B A3B (Reasoning)
16.90
304GLM-4.6V (Reasoning)
16.90
305GPT-5.4 mini (Non-Reasoning)
16.80
306GLM-4.5-Air
16.70
307Nova 2.0 Omni (low)
16.70
308Grok 4 Fast (Non-reasoning)
16.60
309Mi:dm K 2.5 Pro
16.60
310Ring-1T
16.30
311G9v3-3B
16.20
312Qwen3.5 4B (Non-reasoning)
16.10
313Mistral Large 3
15.90
314o3-mini (high)
15.70
315INTELLECT-3
15.70
316GLM-4.7-Flash (Non-reasoning)
15.60
317GPT-5 (ChatGPT)
15.40
318gpt-oss-20b (high)
15.20
319Solar Open 100B (Reasoning)
15.20
320DeepSeek V3 0324
15.20
321Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning)
15.20
322Grok 3 Reasoning Beta
15.20
323Nemotron 3 Nano Omni 30B A3B Reasoning
15.00
324Mistral Small 3.1
14.90
325gpt-oss-120b (low)
14.90
326GPT-4.1 mini
14.80
327Mistral Medium 3.1
14.70
328Qwen3 30B A3B 2507 (Reasoning)
14.60
329MiniMax M1 40k
14.50
330NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)
14.50
331Llama 4 Maverick
14.50
332Solar Pro 3
14.50
333gpt-oss-20b (low)
14.40
334Nova 2.0 Pro Preview (Non-reasoning)
14.40
335Qwen3 VL 235B A22B Instruct
14.40
336Granite 4.2 3B
14.30
337GPT-5 mini (minimal)
14.30
338DeepSeek V3 (Dec '24)
14.20
339Gemini 2.5 Flash (Non-reasoning)
14.20
340Ling 2.6 Flash
14.20
341K2-V2 (high)
14.20
342o1-mini
14.00
343Qwen3 Next 80B A3B Instruct
13.80
344GPT-4.5 (Preview)
13.60
345Tri-21B-think Preview
13.60
346Qwen3 Coder 30B A3B Instruct
13.60
347Qwen3 235B A22B (Reasoning)
13.50
348DiffusionGemma 26B A4B
13.50
349Qwen3 VL 30B A3B (Reasoning)
13.40
350QwQ 32B
13.40

what intelligence means

A composite of nine independent evaluations covering reasoning, coding, agentic work and knowledge. Higher is better. It is versioned, so scores are comparable within a version rather than across all time.