AI Model Leaderboard &
Empirical Comparisons
Independent, verifiable rankings of GPT, Claude, Gemini, Llama, DeepSeek, and 45+ multi-modality AI models. Evaluated continuously by Composite TrueSkill, coding benchmarks, output speed, and live API pricing.

Best AI for Coding
Code generation, debugging & refactoring
Best AI for Writing
Prose, tone matching & creative content
Best AI for Research
Analysis, synthesis & deep logical tasks
Best Image Generation
Text-to-image diffusion & editing
Price vs Performance (Pareto Frontier)
Blended token cost (8:1 input/output ratio) plotted against TrueSkill score. Models on the green line are Pareto-efficient.
Model Performance vs Blended Price
AI Model Comparison Table
| Rank | Model & Provider | Score | Reasoning | Coding | Arena ELO | Architecture | Context | Speed | Pricing $/1M | Action |
|---|---|---|---|---|---|---|---|---|---|---|
| 1NEW | ChatGPT Astra (GPT-6 Astra) OpenAI | 61.2 | 63.5 | 57.8 | 2,695 | Recurrent Depth (~1.8T Equiv) | 1.2M | 95 tok/s | $20.00 | Compare |
| 1NEW | Gemini 3.8 Flash Google | 59 | 56.5 | 54.2 | 2,190 | Sparse MoE (~140B Active) | 1.0M | 348 tok/s | $1.12 | Compare |
| 2NEW | Claude Opus 5.1 Anthropic | 58.8 | 57.2 | 51.5 | 2,710 | Dense (~250B) | 1.0M | 64 tok/s | $15.00 | Compare |
| 3NEW | Claude Fable 5.1 Anthropic | 58.2 | 56 | 53.5 | 2,160 | MoE (~380B) | 1.0M | 86 tok/s | $12.00 | Compare |
| 1 | GPT-5.6 Sol OpenAI | 57.4 | 56.8 | 50.6 | 2,134 | MoE (~1.8T) | 1.1M | 102 tok/s | $7.78 | Compare |
| 2 | Claude Opus 5 Anthropic | 56.2 | 55.3 | 42.7 | 2,668 | Dense (~220B) | 1.0M | 58 tok/s | $7.22 | Compare |
| 3 | Claude Fable 5 Anthropic | 56.1 | 53.5 | 48.8 | 2,003 | MoE (~350B) | 1.0M | 78 tok/s | $14.40 | Compare |
| 4NEW | Claude Mythos Preview Anthropic | 56 | 56.8 | 46.6 | 2,190 | Research Tier | 1.0M | 45 tok/s | Free / Open | Compare |
| 5 | Kimi K3 Moonshot AI | 54.9 | 53.9 | 45.9 | 1,816 | 2.8T (MoE) | 1.0M | 135 tok/s | $4.33 | Compare |
| 8NEW | Muse Spark 1.3 Meta | 54.8 | 54.5 | 47.2 | 1,840 | MoE (~110B Active) | 1.0M | 245 tok/s | $1.58 | Compare |
| 6NEW | GLM-5.3 Zhipu AI | 54.7 | 54.9 | 45.4 | 1,780 | 753B (MoE) | 1.0M | 142 tok/s | $1.73 | Compare |
| 7NEW | DeepSeek-V4-Pro-0813 DeepSeek | 54.5 | 52 | 44.3 | 1,980 | 1.6T (MoE) | 1.0M | 199 tok/s | $0.48 | Compare |
| 8NEW | Qwen3.8 Max Alibaba Cloud / Qwen Team | 53.2 | 52 | 42.1 | 1,850 | 2.4T (MoE) | 1.0M | 160 tok/s | $0.85 | Compare |
| 9 | GPT-5.6 Terra OpenAI | 52.8 | 51.1 | 46.4 | 1,185 | MoE (~500B) | 1.1M | 119 tok/s | $3.11 | Compare |
| 10 | Claude Opus 4.8 Anthropic | 51.9 | 51.4 | 43.8 | 1,689 | Dense | 1.0M | 125 tok/s | $7.22 | Compare |
| 11 | Muse Spark 1.1 Meta | 51.7 | 52.3 | 37.8 | 1,070 | Dense (~70B) | 1.0M | 229 tok/s | $1.58 | Compare |
| 12NEW | Gemini 3.7 Flash Google | 51.1 | 49.9 | 38.6 | 1,720 | Sparse MoE | 1.0M | 621 tok/s | $1.08 | Compare |
| 13 | Claude Sonnet 5 Anthropic | 49.6 | 49 | 40 | 1,588 | Dense (~100B) | 1.0M | 59 tok/s | $2.89 | Compare |
| 14 | GPT-5.5 OpenAI | 49.3 | 48.6 | 41.1 | 2,124 | MoE | 1.1M | 49 tok/s | $7.78 | Compare |
| 15NEW | DeepSeek-V4-Flash-Vision-Exp DeepSeek | 48.7 | 45.3 | 38.5 | 1,650 | Vision MoE | 1.0M | 310 tok/s | $0.27 | Compare |
| 16NEW | Grok 4.5 xAI | 48.2 | 47.8 | 37.2 | 1,540 | MoE (~300B) | 2.0M | 110 tok/s | $2.00 | Compare |
Popular 1v1 Model Comparisons
Direct head-to-head empirical evaluations across speed, pricing, and verified benchmark scores.
ChatGPT Astra vs Claude Fable 5.1
ChatGPT Astra vs Claude Opus 5.1
Gemini 3.8 Flash vs GPT-5.6 Sol
Muse Spark 1.3 vs GPT-5.6 Sol
Gemini 3.8 Flash vs Claude Fable 5.1
Muse Spark 1.3 vs DeepSeek-V4 Pro
Browse AI Models by Category
By Capability
By Modality
Latest AI Model Research & Releases
OpenAI Releases GPT-6 Astra: 99.9% ARC-AGI-3 & Critical Cybersecurity Rating
OpenAI deploys GPT-6 Astra featuring recurrent depth architecture, Symphony cross-modal encoding, 72.6% OSWorld 2.0 autonomous computer control, and 98.0% FrontierMath Tier 4.
Gemini 3.8 Flash Released: 90.8% Terminal-Bench & 71.0% DeepSWE at $0.75/M
Google officially deploys Gemini 3.8 Flash featuring 90.8% Terminal-Bench 2.1 shell coding, 71.0% DeepSWE v1.1, and 348-620 tokens/sec generation on TPU v6e.
Claude Fable 5.1 & Mythos 5.1: Non-Shortcut Long-Horizon Agent Engineering
Anthropic launches Fable 5.1 with deep agentic multi-file repository refactoring, reaching 62.8% on SWE-Bench Pro with strict anti-shortcut evaluation.
Meta Launches Muse Spark 1.3: 245 tok/s Throughput & Dual Engine Architecture
Meta releases Muse Spark 1.3 open-weights flagship with dual Max and XHigh engines, 1.0M context, and deep Muse Code integration at $1.58/M blended.
DeepSeek-V4-Flash-Vision-Exp: Vision at Flash Price
DeepSeek's first official V4-Flash vision API. Same $0.22/$0.66 off-peak list as V4-Flash, 1M context. Self-reported Chartography 64.3 and TerminalBench 83.9.