Home

Frontier model roster

The top 20, by AA-Omniscience

Ranked by AA-Omniscience Index, which nets a model’s knowledge against how often it invents an answer. Everything to the right of it is the trade you make to get that: how well it holds a long context, drives a tool, reads an image, and what a single task actually costs once you count the tokens it burns getting there.

Top 20 scored models by Omniscience · captured 14 August 2026Data: Artificial Analysis
Claude Fable 5 (with fallback)Anthropic1M43577762100.42s63$3.1419.70.003
Claude Opus 5 (max)Anthropic1M375979856353.48s53$2.3426.90.007
Gemini 3.1 Pro PreviewGoogle1M322379824827.13s113$0.331450.096
Grok 4.6 (high)SpaceXAI500K3059756132.30s66$0.8472.60.030
Claude Opus 4.8 (max)Anthropic1M2949735729.18s59$2.0328.10.013
Muse Spark 1.1 (xhigh)Meta1.1M284081531.69s230$0.291830.275
Muse Spark 1.2 (xhigh)Meta1.1M27498357$0.40143
Claude Opus 4.7 (max)Anthropic1M274675795515.36s46$2.2324.70.017
Gemini 3.7 Flash (high)Google1M26458185569.83s340$0.401400.221
Grok 4.5 (high)SpaceXAI500K25497480568.68s60$0.361560.164
GPT-5.6 Sol (max)OpenAI1M2258788361176.57s62$1.2349.60.004
Gemini 3.6 FlashGoogle1M224179835218.68s225$0.5692.90.085
GPT-5.5 (xhigh)OpenAI922K2147798156100.91s67$1.1747.90.008
Gemini 3.5 FlashGoogle1M214081845226.96s172$0.6975.40.049
Kimi K3 (max)Kimi1.1M20548381603.61s40$0.8471.40.018
Grok 4.3 (high)SpaceXAI1M182468783823.91s133$0.152530.241
Claude Sonnet 5 (max)Anthropic1M1650777755167.44s67$1.7232.00.003
Gemini 3 Pro Preview (high)Google1M15738041
Grok 4.20 0309 v2SpaceXAI2M1562753827.81s95
Claude Opus 4.6 (max)Anthropic1M1474754514.99s40

How to read it

The three cost columns are the ones that change decisions. Cost per task is the measured bill for one unit of work. Intelligence per cost says what that money buys in capability; E2E per cost says what it buys in responsiveness. A model can top the intelligence column and still lose both, which is usually the moment a cheaper model turns out to be the right one for the job in front of you.

Sort by any column. The two per-cost columns are shaded by rank, so the brighter the cell, the more capability or speed that model returns per dollar.

What these numbers are and are not

Cost is measured per completed task, not quoted per million tokens, so a verbose model is charged for its verbosity. Two models on the same headline token price can differ several-fold here.

An em dash means Artificial Analysis has not measured that model on that axis. It is absent from the comparison, not scored zero, and it sorts to the bottom of that column in both directions.

Figures are a snapshot taken on 14 August 2026, not a live feed. Every number is Artificial Analysis’ own measurement, not this site’s.