LLM leaderboard
Every tracked language model by intelligence, with what it costs and how fast it answers.
ranked by intelligence · 611 models · updated Aug 30 · 451–500
| # | Model | Creator | Intelligence | $/1M in | $/1M out | tok/s |
|---|---|---|---|---|---|---|
| 451 | GPT-4o (ChatGPT) | OpenAI | 8.10 | — | — | — |
| 452 | LFM2.5-8B-A1B | Liquid AI | 8.10 | $0 | $0 | 339 |
| 453 | Llama 3.1 Tulu3 405B | Allen Institute for AI | 8.10 | — | — | — |
| 454 | Ring-flash-2.0 | InclusionAI | 8.00 | $0.14 | $0.57 | — |
| 455 | Pixtral Large | Mistral | 8.00 | — | — | — |
| 456 | Olmo 3.1 32B Think | Allen Institute for AI | 7.90 | $0 | $0 | — |
| 457 | GPT-5 nano (minimal) | OpenAI | 7.80 | $0.05 | $0.4 | 166 |
| 458 | Grok 2 (Dec '24) | SpaceXAI | 7.80 | — | — | — |
| 459 | Gemini 1.5 Flash (Sep '24) | 7.80 | — | — | — | |
| 460 | GPT-4 Turbo | OpenAI | 7.70 | $10 | $30 | 32 |
| 461 | Qwen3 VL 4B (Reasoning) | Alibaba | 7.70 | — | — | — |
| 462 | Solar Pro 2 (Non-reasoning) | Upstage | 7.60 | — | — | — |
| 463 | Command A | Cohere | 7.50 | $2.5 | $10 | 68 |
| 464 | Nova Pro | Amazon | 7.50 | $0.8 | $3.2 | — |
| 465 | Gemma 3 27B Instruct | 7.40 | $0 | $0 | — | |
| 466 | Qwen3.5 2B (Reasoning) | Alibaba | 7.40 | — | — | — |
| 467 | Llama 3.1 Nemotron Instruct 70B | NVIDIA | 7.40 | $1.2 | $1.2 | 75 |
| 468 | Llama 3.1 Instruct 8B | Meta | 7.40 | $0.02 | $0.05 | 137 |
| 469 | Grok Beta | SpaceXAI | 7.30 | — | — | — |
| 470 | Qwen2.5 Instruct 32B | Alibaba | 7.20 | — | — | — |
| 471 | NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) | NVIDIA | 7.20 | $0.05 | $0.2 | 268 |
| 472 | NVIDIA Nemotron Nano 9B V2 (Non-reasoning) | NVIDIA | 7.20 | $0.05 | $0.2 | 171 |
| 473 | Ministral 3 3B | Mistral | 7.10 | $0.1 | $0.1 | 174 |
| 474 | Mistral Large 2 (Jul '24) | Mistral | 7.00 | $2 | $6 | — |
| 475 | Qwen2.5 Coder Instruct 32B | Alibaba | 6.90 | — | — | — |
| 476 | Qwen3 4B 2507 Instruct | Alibaba | 6.90 | — | — | — |
| 477 | GLM-4.5V (Non-reasoning) | Z AI | 6.80 | $0.6 | $1.8 | 79 |
| 478 | GPT-4 | OpenAI | 6.80 | $30 | $60 | — |
| 479 | Qwen3 14B (Non-reasoning) | Alibaba | 6.80 | $0.35 | $1.4 | 60 |
| 480 | Gemini 2.5 Flash-Lite (Non-reasoning) | 6.70 | $0.1 | $0.4 | 300 | |
| 481 | Hermes 4 - Llama-3.1 70B (Non-reasoning) | Nous Research | 6.70 | $0.13 | $0.4 | 86 |
| 482 | Nova Lite | Amazon | 6.70 | $0.06 | $0.24 | 170 |
| 483 | GPT-4o mini | OpenAI | 6.70 | $0.15 | $0.6 | 86 |
| 484 | Mistral Small 3 | Mistral | 6.70 | $0.1 | $0.3 | 145 |
| 485 | Qwen3 30B A3B (Non-reasoning) | Alibaba | 6.60 | $0.2 | $0.8 | 106 |
| 486 | Llama 3.1 Instruct 70B | Meta | 6.50 | $0.56 | $0.56 | 55 |
| 487 | DeepSeek-V2.5 (Dec '24) | DeepSeek | 6.50 | — | — | — |
| 488 | Qwen3 4B (Non-reasoning) | Alibaba | 6.50 | — | — | — |
| 489 | Sarvam 30B (high) | Sarvam | 6.40 | $0.03 | $0.11 | — |
| 490 | Granite 4.1 8B | IBM | 6.40 | $0.05 | $0.1 | 119 |
| 491 | DeepSeek-V2.5 | DeepSeek | 6.40 | — | — | — |
| 492 | Gemini 2.0 Flash Thinking Experimental (Dec '24) | 6.40 | — | — | — | |
| 493 | DeepSeek R1 Distill Llama 8B | DeepSeek | 6.20 | — | — | — |
| 494 | Gemma 4 E2B (Non-reasoning) | 6.20 | — | — | — | |
| 495 | Mistral Saba | Mistral | 6.20 | — | — | — |
| 496 | Olmo 3.1 32B Instruct | Allen Institute for AI | 6.20 | — | — | — |
| 497 | Gemini 1.5 Pro (May '24) | 6.10 | — | — | — | |
| 498 | Olmo 3 32B Think | Allen Institute for AI | 6.10 | — | — | — |
| 499 | Qwen2.5 Turbo | Alibaba | 6.00 | $0.05 | $0.2 | 102 |
| 500 | R1 1776 | Perplexity | 6.00 | — | — | — |
what intelligence means
A composite of nine independent evaluations covering reasoning, coding, agentic work and knowledge. Higher is better. It is versioned, so scores are comparable within a version rather than across all time.
