Specialist test. Evaluate this separately from the eight core tests. Explore the active suite →

Focused model compare

City Scroll Journey

Grok 4.6 is the starting run, compared with Claude Opus 5.5, GPT-6 Astra, GPT-6 Sol.