LLM leaderboard
Every tracked language model by intelligence, with what it costs and how fast it answers.
ranked by intelligence · 611 models · updated Aug 30 · 551–600
| # | Model | Creator | Intelligence | $/1M in | $/1M out | tok/s |
|---|---|---|---|---|---|---|
| 551 | DeepSeek R1 Distill Qwen 1.5B | DeepSeek | 3.30 | — | — | — |
| 552 | Claude 2.0 | Anthropic | 3.30 | — | — | — |
| 553 | GPT-3.5 Turbo | OpenAI | 3.20 | $0.5 | $1.5 | — |
| 554 | Mistral Medium | Mistral | 3.20 | $1.5 | $7.5 | 140 |
| 555 | Mistral Small (Feb '24) | Mistral | 3.20 | $0.15 | $0.6 | 134 |
| 556 | Llama 3 Instruct 70B | Meta | 3.10 | $0.65 | $2.75 | — |
| 557 | LFM 40B | Liquid AI | 3.00 | — | — | — |
| 558 | Qwen Chat 72B | Alibaba | 3.00 | — | — | — |
| 559 | Llama 3.2 Instruct 11B (Vision) | Meta | 3.00 | $0.34 | $0.34 | 46 |
| 560 | Arctic Instruct | Snowflake | 3.00 | — | — | — |
| 561 | Qwen3.5 0.8B (Non-reasoning) | Alibaba | 2.90 | — | — | — |
| 562 | PALM-2 | 2.80 | — | — | — | |
| 563 | Gemini 1.0 Pro | 2.70 | — | — | — | |
| 564 | DeepSeek Coder V2 Lite Instruct | DeepSeek | 2.70 | — | — | — |
| 565 | DBRX Instruct | Databricks | 2.60 | — | — | — |
| 566 | Sarvam M (Reasoning) | Sarvam | 2.60 | $0 | $0 | — |
| 567 | Llama 2 Chat 70B | Meta | 2.60 | — | — | — |
| 568 | Llama 2 Chat 13B | Meta | 2.60 | — | — | — |
| 569 | Command-R+ (Apr '24) | Cohere | 2.60 | $3 | $15 | — |
| 570 | DeepSeek LLM 67B Chat (V1) | DeepSeek | 2.60 | — | — | — |
| 571 | OpenChat 3.5 (1210) | OpenChat | 2.60 | — | — | — |
| 572 | Exaone 4.0 1.2B (Reasoning) | LG AI Research | 2.50 | — | — | — |
| 573 | Exaone 4.0 1.2B (Non-reasoning) | LG AI Research | 2.40 | — | — | — |
| 574 | Olmo 3 7B Instruct | Allen Institute for AI | 2.40 | $0.1 | $0.2 | — |
| 575 | LFM2.5-1.2B-Instruct | Liquid AI | 2.30 | — | — | — |
| 576 | Jamba 1.5 Mini | AI21 Labs | 2.30 | $0.2 | $0.4 | — |
| 577 | LFM2.5-1.2B-Thinking | Liquid AI | 2.30 | — | — | — |
| 578 | LFM2 2.6B | Liquid AI | 2.30 | — | — | — |
| 579 | Jamba 1.7 Mini | AI21 Labs | 2.30 | — | — | — |
| 580 | Qwen3 1.7B (Reasoning) | Alibaba | 2.20 | — | — | — |
| 581 | Granite 4.0 H 1B | IBM | 2.20 | — | — | — |
| 582 | Jamba 1.6 Mini | AI21 Labs | 2.10 | — | — | — |
| 583 | Gemma 3 270M | 2.00 | — | — | — | |
| 584 | Granite 4.0 Micro | IBM | 2.00 | — | — | — |
| 585 | Mixtral 8x7B Instruct | Mistral | 2.00 | $0.45 | $0.7 | — |
| 586 | Apertus 70B Instruct | Swiss AI Initiative | 2.00 | $0.82 | $2.92 | — |
| 587 | DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) | Nous Research | 1.90 | — | — | — |
| 588 | Mistral 7B Instruct | Mistral | 1.70 | $0.25 | $0.25 | 125 |
| 589 | Claude Instant | Anthropic | 1.70 | — | — | — |
| 590 | Command-R (Mar '24) | Cohere | 1.70 | $0.5 | $1.5 | — |
| 591 | Qwen Chat 14B | Alibaba | 1.70 | — | — | — |
| 592 | Llama 65B | Meta | 1.70 | — | — | — |
| 593 | Molmo2-8B | Allen Institute for AI | 1.60 | — | — | — |
| 594 | Granite 4.0 1B | IBM | 1.60 | — | — | — |
| 595 | Granite 3.3 8B (Non-reasoning) | IBM | 1.30 | $0.03 | $0.25 | 15 |
| 596 | LFM2 8B A1B | Liquid AI | 1.30 | — | — | — |
| 597 | Qwen3 1.7B (Non-reasoning) | Alibaba | 1.10 | — | — | — |
| 598 | Qwen3 0.6B (Non-reasoning) | Alibaba | 1.00 | — | — | — |
| 599 | LFM2 1.2B | Liquid AI | 1.00 | — | — | — |
| 600 | Gemma 3 4B Instruct | 1.00 | $0 | $0 | — |
what intelligence means
A composite of nine independent evaluations covering reasoning, coding, agentic work and knowledge. Higher is better. It is versioned, so scores are comparable within a version rather than across all time.
