Model compare

Mechanical Watch Simulator

Compare the same prompt across DeepSeek V4 Flash, Kimi K3, MiniMax M3.

Missing runs for gpt-sol-5-6-high on this benchmark.