Flagship Models
The most capable general-purpose models from each provider.
9 models · Updated Jul 12, 2026
| # | ||
|---|---|---|
| 2 | Claude Opus 4.8 | 56.0 |
| 3 | GPT-5.5 | 52.5 |
| 4 | DeepSeek V4 Pro | 51.8 |
| 5 | Gemini 3.1 Pro Preview | 50.8 |
| 8 | Kimi K2.6 | 32.9 |
| 10 | GPT-5.5 Pro | 31.9 |
| 11 | Qwen3.7 Max | 31.7 |
| 12 | Grok 4.20 (Reasoning) | 31.6 |
| 15 | Mistral Large 3 | 21.7 |
Popular comparisons
Related model categories
Frequently asked questions
- How are flagship models ranked?
- Each model is scored on weighted, normalized metrics (pricing, capability, and benchmark data), then adjusted for data confidence and freshness. Sources are listed on every model page.
- How fresh is this data?
- Data is verified against named sources and flagged when older than 45 days. The last verification date is shown on each page.