Retired benchmark. These historical results are preserved for inspection and excluded from active-suite run counts. Browse the archive →

Model compare

Petri Dish

Compare the same prompt across GPT-6 Astra, Claude Fable 5.1, Kimi K3, Grok 4.7.