LLM leaderboard

Every tracked language model by intelligence, with what it costs and how fast it answers.

ranked by intelligence · 611 models · updated Aug 30 · 551–600

#ModelIntelligence
551DeepSeek R1 Distill Qwen 1.5B
3.30
552Claude 2.0
3.30
553GPT-3.5 Turbo
3.20
554Mistral Medium
3.20
555Mistral Small (Feb '24)
3.20
556Llama 3 Instruct 70B
3.10
557LFM 40B
3.00
558Qwen Chat 72B
3.00
559Llama 3.2 Instruct 11B (Vision)
3.00
560Arctic Instruct
3.00
561Qwen3.5 0.8B (Non-reasoning)
2.90
562PALM-2
2.80
563Gemini 1.0 Pro
2.70
564DeepSeek Coder V2 Lite Instruct
2.70
565DBRX Instruct
2.60
566Sarvam M (Reasoning)
2.60
567Llama 2 Chat 70B
2.60
568Llama 2 Chat 13B
2.60
569Command-R+ (Apr '24)
2.60
570DeepSeek LLM 67B Chat (V1)
2.60
571OpenChat 3.5 (1210)
2.60
572Exaone 4.0 1.2B (Reasoning)
2.50
573Exaone 4.0 1.2B (Non-reasoning)
2.40
574Olmo 3 7B Instruct
2.40
575LFM2.5-1.2B-Instruct
2.30
576Jamba 1.5 Mini
2.30
577LFM2.5-1.2B-Thinking
2.30
578LFM2 2.6B
2.30
579Jamba 1.7 Mini
2.30
580Qwen3 1.7B (Reasoning)
2.20
581Granite 4.0 H 1B
2.20
582Jamba 1.6 Mini
2.10
583Gemma 3 270M
2.00
584Granite 4.0 Micro
2.00
585Mixtral 8x7B Instruct
2.00
586Apertus 70B Instruct
2.00
587DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)
1.90
588Mistral 7B Instruct
1.70
589Claude Instant
1.70
590Command-R (Mar '24)
1.70
591Qwen Chat 14B
1.70
592Llama 65B
1.70
593Molmo2-8B
1.60
594Granite 4.0 1B
1.60
595Granite 3.3 8B (Non-reasoning)
1.30
596LFM2 8B A1B
1.30
597Qwen3 1.7B (Non-reasoning)
1.10
598Qwen3 0.6B (Non-reasoning)
1.00
599LFM2 1.2B
1.00
600Gemma 3 4B Instruct
1.00

what intelligence means

A composite of nine independent evaluations covering reasoning, coding, agentic work and knowledge. Higher is better. It is versioned, so scores are comparable within a version rather than across all time.