Flagship Models
The most capable general-purpose models from each provider.
9 models · Updated Jul 12, 2026
| # | ||
|---|---|---|
| 2 | Claude Opus 4.8 | 49.7 |
| 3 | GPT-5.5 | 46.5 |
| 4 | DeepSeek V4 Pro | 45.9 |
| 5 | Gemini 3.1 Pro Preview | 45.1 |
| 8 | Kimi K2.6 | 29.1 |
| 9▲1 | GPT-5.5 Pro | 28.3 |
| 10▲1 | Qwen3.7 Max | 28.1 |
| 11▲1 | Grok 4.20 (Reasoning) | 28.0 |
| 14▲1 | Mistral Large 3 | 19.3 |
Popular comparisons
Related model categories
Frequently asked questions
- How are flagship models ranked?
- Each model is scored on weighted, normalized metrics (pricing, capability, and benchmark data), then adjusted for data confidence and freshness. Sources are listed on every model page.
- How fresh is this data?
- Data is verified against named sources and flagged when older than 45 days. The last verification date is shown on each page.