RSSAmplifier

Marcin Dudek — Blog · Jul 17, 2026

A one-shot code-review benchmark: scoring restraint over recall across four Claude models

0
Sign in to vote or save

This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.

I built a small benchmark to score the half of code review most evals ignore - restraint, not flagging correct code that only looks suspicious. Four Claude models (Opus 4.6, Opus 4.8, Sonnet 5, Haiku 4.5), four task families, one-shot and mechanically scored. The two Opus models tied and cost the same; Sonnet 5 matched them on almost everything while running fastest and costing a third as much;…

Read on marcindudek.dev

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.