Grok 4.6 just landed, so we did what we do every launch: skipped the leaderboard charts and put it to work. But this time we didn't stop at the newcomer. We lined up nine models — Grok 4.6, Fable 5, Opus 5, Opus 4.8, Sonnet 5, GPT-5.4, Gemini 3.1 Pro, Kimi K3, and GLM 5.2 — and ran every one of them through the same 16 jobs a PM, founder, engineer, or everyday operator does every week.
The bar moved. Here's exactly where.
Sixteen tasks. Four personas — engineer, PM, founder, generalist — four tasks each. Every model gets the identical prompt. All nine outputs go anonymized, in shuffled order, to a judge panel from neutral model families (Qwen, DeepSeek, MiniMax — none of the nine contestant families sit on the panel). Judges score 7 weighted dimensions, name a winner, and flag any claim not supported by the facts in the prompt. Real cost and latency metered where we ran fresh.
We're not walking through each test's prompt and trap here — the full audit pack (every prompt, every verbatim output, every judge's scores and rationale) is published in the repo for anyone who wants to check our work. This edition is the verdict sheet.
This is the fifth edition of the series, on the same suite as Opus 4.8 vs Fable 5, GLM 5.2 vs Opus 4.8, and the July Sonnet 5 four-way.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.