Marcin Dudek — Blog · Jul 17, 2026
A one-shot code-review benchmark: scoring restraint over recall across four Claude models
0Sign in to vote or save
This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.
I built a small benchmark to score the half of code review most evals ignore - restraint, not flagging correct code that only looks suspicious. Four Claude models (Opus 4.6, Opus 4.8, Sonnet 5, Haiku 4.5), four task families, one-shot and mechanically scored. The two Opus models tied and cost the same; Sonnet 5 matched them on almost everything while running fastest and costing a third as much;…
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.