AI model benchmarks

Choose a model by practical tier, then inspect every published same-prompt run behind the recommendation.

S

A

B

C

D

Additional benchmark runs

Published test sets that are not part of the current editorial tier ranking.

25 models across 5 tiers
223 runs across 19 tested models

Updated

Model hub next step

Compare the models before you commit.

Start with the editorial tier, then open the source-linked cost and intelligence tables for the workload you actually have. Tier placement is routing guidance, not a substitute for a current evaluation.

Last reviewed

Link map: Compare models and costsInspect test briefsLearn how tiers are set

Evidence and tools

Compare the evidence.

Compare independent scores, API costs, and the same prompts built by different models.

Benchmark hub next step

Turn a model choice into a testable brief.

Use a published prompt to compare the same job across models, then keep the output, method, and caveats together when you make a call.

Last reviewed

Link map: Choose a benchmark briefCompare model evidenceReview the testing skill