LLM leaderboard
Every tracked language model by intelligence, with what it costs and how fast it answers.
Claude Opus 5 (Adaptive Reasoning, Max Effort) leads from Anthropic, ahead of 610 others.
- 1Anthropic
Claude Opus 5 (Adaptive Reasoning, Max Effort)
63.10 intelligence
- 2Anthropic
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
62.50 intelligence
- 3Anthropic
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
62.10 intelligence
ranked by intelligence · 611 models · updated Aug 30 · 1–50
| # | Model | Creator | Intelligence | $/1M in | $/1M out | tok/s |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (Adaptive Reasoning, Max Effort) | Anthropic | 63.10 | $5 | $25 | 55 |
| 2 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | 62.50 | $5 | $25 | 52 |
| 3 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 62.10 | $10 | $50 | 65 |
| 4 | Claude Opus 5 (Adaptive Reasoning, High Effort) | Anthropic | 61.50 | $5 | $25 | 54 |
| 5 | Grok 4.6 (high) | SpaceXAI | 60.90 | $2 | $6 | 58 |
| 6 | GPT-5.6 Sol (max) | OpenAI | 60.90 | $4 | $20 | 70 |
| 7 | Grok 4.6 (xhigh) | SpaceXAI | 60.00 | $2 | $6 | 60 |
| 8 | Kimi K3 (max) | Kimi | 59.70 | $3 | $15 | 38 |
| 9 | GLM-5.3 (max) | Z AI | 59.50 | $1.4 | $4.4 | 77 |
| 10 | Grok 4.6 (medium) | SpaceXAI | 59.00 | $2 | $6 | 57 |
| 11 | GPT-5.6 Sol (xhigh) | OpenAI | 59.00 | $4 | $20 | 73 |
| 12 | Claude Opus 5 (Adaptive Reasoning, Medium Effort) | Anthropic | 58.60 | $5 | $25 | 50 |
| 13 | Qwen3.8 Max | Alibaba | 58.10 | $2 | $6 | 21 |
| 14 | Qwen3.8 2.4T A95B | Alibaba | 57.70 | $2 | $6 | 24 |
| 15 | GLM-5.3-Flash | Z AI | 57.50 | $0.15 | $0.5 | 49 |
| 16 | GPT-5.6 Sol (high) | OpenAI | 57.30 | $4 | $20 | 75 |
| 17 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) | Anthropic | 57.30 | $5 | $25 | 58 |
| 18 | Muse Spark 1.2 (xhigh) | Meta | 56.80 | $1.25 | $4.25 | — |
| 19 | GPT-5.6 Terra (max) | OpenAI | 56.60 | $2 | $12 | 108 |
| 20 | GPT-5.5 (xhigh) | OpenAI | 56.30 | $5 | $30 | 83 |
| 21 | Gemini 3.7 Flash (high) | 56.00 | $0.75 | $3.75 | 322 | |
| 22 | Grok 4.5 (high) | SpaceXAI | 55.80 | $2 | $6 | 52 |
| 23 | Qwen3.8-Flash-Next | Alibaba | 55.80 | $0.15 | $0.47 | 73 |
| 24 | GPT-5.6 Sol (medium) | OpenAI | 55.60 | $4 | $20 | 73 |
| 25 | Claude Sonnet 5 (Adaptive Reasoning, Max Effort) | Anthropic | 55.30 | $2 | $10 | 84 |
| 26 | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) | Anthropic | 55.00 | $5 | $25 | 51 |
| 27 | GPT-5.5 (high) | OpenAI | 54.70 | $5 | $30 | 76 |
| 28 | Gemini 3.7 Flash (medium) | 53.40 | $0.75 | $3.75 | 318 | |
| 29 | DeepSeek V4 Pro 0813 (Reasoning, Max Effort) | DeepSeek | 53.20 | $1.32 | $3.96 | 66 |
| 30 | Muse Spark 1.1 (xhigh) | Meta | 53.20 | $1.25 | $4.25 | 205 |
| 31 | GPT-5.4 (xhigh) | OpenAI | 53.10 | $2.5 | $15 | 130 |
| 32 | GPT-5.6 Terra (xhigh) | OpenAI | 52.80 | $2 | $12 | 105 |
| 33 | GLM-5.2 (max) | Z AI | 52.60 | $1.4 | $4.4 | 71 |
| 34 | Claude Opus 5 (Adaptive Reasoning, Low Effort) | Anthropic | 52.50 | $5 | $25 | 51 |
| 35 | GPT-5.6 Luna (max) | OpenAI | 52.30 | $0.2 | $1.2 | 127 |
| 36 | Gemini 3.5 Flash (high) | 52.00 | $1.5 | $9 | 206 | |
| 37 | Qwen3.8 27B (xhigh) | Alibaba | 52.00 | $0.5 | $3 | 47 |
| 38 | DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | DeepSeek | 51.80 | $0.44 | $1.32 | 132 |
| 39 | Grok 4.6 (low) | SpaceXAI | 51.70 | $2 | $6 | 55 |
| 40 | Gemini 3.6 Flash (high) | 51.60 | $0.75 | $3.75 | 173 | |
| 41 | DeepSeek V4 Flash Vision (Reasoning, Max Effort) | DeepSeek | 51.50 | $0.44 | $1.32 | 119 |
| 42 | GPT-5.5 (medium) | OpenAI | 51.40 | $5 | $30 | 83 |
| 43 | Gemini 3.7 Flash (low) | 50.90 | $0.75 | $3.75 | 320 | |
| 44 | GPT-5.6 Sol (low) | OpenAI | 50.70 | $4 | $20 | 73 |
| 45 | GPT-5.6 Terra (high) | OpenAI | 50.10 | $2 | $12 | 104 |
| 46 | GPT-5.6 Luna (xhigh) | OpenAI | 50.10 | $0.2 | $1.2 | 115 |
| 47 | Agnes 2.5 Pro Beta | Sapiens AI | 49.10 | $0.1 | $0.3 | 159 |
| 48 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) | Anthropic | 48.40 | $3 | $15 | 57 |
| 49 | Kimi K3 (low) | Kimi | 48.30 | $3 | $15 | 36 |
| 50 | Gemini 3.1 Pro Preview | 47.70 | $2 | $12 | 115 |
what intelligence means
A composite of nine independent evaluations covering reasoning, coding, agentic work and knowledge. Higher is better. It is versioned, so scores are comparable within a version rather than across all time.
