AI model benchmarks
Choose a model by practical tier, then inspect every published same-prompt run behind the recommendation.
S
A
B
C
D
Additional benchmark runs
Published test sets that are not part of the current editorial tier ranking.
Updated
Evidence and tools
Compare the evidence.
Compare independent scores, API costs, and the same prompts built by different models.