5shared-table models
6same-eval models
1unranked for now

Independent measurements lead. Vendor results stay in separate, source-linked lanes.

Independent comparison

One harness across 9 models.

Artificial Analysis Index v4.1 combines nine evaluations across coding, agentic work, science, knowledge, physics, and long-context reasoning. Higher is better.

Top index61Grok 4.6
Fastest output178GPT-5.6 Luna, tokens per second
Lowest blended price$0.87GPT-5.6 Luna, per 1M tokens
  1. 61AA source ↗
  2. Claude Fable 5Max, Opus 4.8 fallback
    60AA source ↗
  3. Claude Opus 5Adaptive reasoning, xhigh
    60AA source ↗
  4. 59AA source ↗
  5. Kimi K3Reasoning
    57AA source ↗
  6. 55AA source ↗
  7. GPT-5.5xhigh
    55AA source ↗
  8. 54AA source ↗
  9. 46AA source ↗

Configuration matters. Fable 5 includes its published Opus 4.8 fallback. GPT and Claude scores use the effort level shown beside each model.

Cost explorer

Price your actual workload.

Set the input and output tokens for one task. The explorer applies current API list prices and recalculates every model in the same scenario.

Example workloads
ModelAA IndexInput / 1MOutput / 1MSpeedYour task
Claude Fable 5Anthropic · Max, Opus 4.8 fallback60Index$10.00$50.0069.7tokens/sec$6.00AA data ↗Price ↗
Claude Opus 5Anthropic · Adaptive reasoning, xhigh60Index$5.00$25.0053.1tokens/sec$3.00AA data ↗Price ↗
GPT-5.6 SolOpenAI · Max59Index$5.00$30.0077.1tokens/sec$3.40AA data ↗Price ↗
Kimi K3Moonshot AI · Reasoning57Index$3.00$15.0032.0tokens/sec$1.80AA data ↗Price ↗
GPT-5.6 TerraOpenAI · Max55Index$2.50$15.00134.5tokens/sec$1.70AA data ↗Price ↗
GPT-5.5OpenAI · xhigh55Index$5.00$30.0072.5tokens/sec$3.40AA data ↗Price ↗
Grok 4.5Xai · High54Index$2.00$6.0067.1tokens/sec$0.88AA data ↗Price ↗
Grok 4.6Xai · High61Index$2.00$6.0065.1tokens/sec$0.88AA data ↗Price ↗
GPT-5.6 LunaOpenAI · High46Index$1.00$6.00178.0tokens/sec$0.68AA data ↗Price ↗

What is included: standard input and generated output at API prices verified 2026-08-14.

What is not included: cache writes, tool calls, storage, batch discounts, provider markups, or the extra tokens a model may consume to finish the same job.

Blended reference: Artificial Analysis also publishes a 7:2:1 cache-hit/input/output price. The lowest in this set is GPT-5.6 Luna at $0.87 per 1M blended tokens.

Comparable score lanes

One task. One scale. Full context.

Each plot is a single benchmark. Orange marks one source table. Outlined marks use the same named evaluation from another publisher.

Agentic coding

SWE-bench Pro

GPT-5.6 comparison table ↗

Public tasks. Scores below are published in cross-model tables; effort and harness are retained in each source.

  1. Claude Fable 5
    80.0%Shared table
  2. Grok 4.5
    64.7%Same evaluation ↗
  3. GPT-5.6 Sol
    64.6%Shared table
  4. GPT-5.6 Terra
    63.4%Shared table
  5. GPT-5.6 Luna
    62.7%Shared table
  6. GPT-5.5
    59.4%Shared table
  7. Inkling
    54.3%Same evaluation ↗

Published July 2026. Open the source before using these figures for a purchasing decision.

Terminal agents

Terminal-Bench 2.1

GPT-5.6 comparison table ↗

Command-line task completion. The primary figures use OpenAI’s shared table; the other reported scores are labelled separately rather than normalized.

  1. GPT-5.6 Sol
    88.8%Shared table
  2. GPT-5.6 Terra
    87.4%Shared table
  3. GPT-5.5
    85.6%Shared table
  4. GPT-5.6 Luna
    84.7%Shared table
  5. Grok 4.5
    83.3%Same evaluation ↗
  6. Claude Fable 5
    83.1%Shared table
  7. Inkling
    63.8%Same evaluation ↗

Published July 2026. Open the source before using these figures for a purchasing decision.

Science reasoning

GPQA Diamond

Inkling model card cross-model table ↗

No-tools science questions. This lane keeps the reported configuration in view because reasoning budgets differ by provider.

  1. GPT-5.6 Sol
    94.1%Shared table
  2. Claude Fable 5
    92.6%Shared table
  3. Inkling
    87.2%Shared table

Published July 2026. Open the source before using these figures for a purchasing decision.

Methodology

What this page will not do.

Artificial Analysis Index results come from one independent harness. The exact model configuration remains attached to every score.

Token cost scenarios multiply published input and output prices by the workload you enter. They exclude tools, storage, cache writes, and provider discounts.

A shared-table mark means the source published the compared models in one table. A same-eval mark means the benchmark name matches, but the source or configuration differs.

Missing data is not a low score. Models without a compatible public score remain unranked until a reproducible result is published.

Our Superbash visual runs are a separate, same-prompt evidence layer. Use the live run comparison to inspect build quality directly.

Open the same-prompt visual benchmarks →