RSSAmplifier

Blog

AHD Eval Runs

Every AHD raw-versus-compiled eval run, dated, versioned, with manifests and per-cell counts.

ahd.adastra.computerRSS feed ↗8 posts

Latest posts

Eight runs · Same brief

The weekly programme read as a series rather than eight bulletins, covering 9 June to 10 August 2026. gpt-oss-120b and mistral-small-3.1 reduce every run and by a consistent amount. gemma-4 reduces every run by an amount that swings twenty points. llama-4-scout sits against zero. qwen3-30b changes sign six times in eight runs, so no claim about its direction survives the series. New runs update…

Three weeks · One split

Third consecutive weekly run. Gemma 57.2%, mistral 64.6% and gpt-oss 72.7% reduce under the compiled prompt; llama-4-scout stays flat at 1.6%. qwen3 swings back to -7.0% after +7.1% the week before, straddling zero across three runs (-3.4, +7.1, -7.0). The reducing band and the llama trade hold; qwen is the lone unstable cell.

The split holds · Week two

Second weekly run. The 9 June split reproduces: gemma 53.3%, mistral 66.0% and gpt-oss 73.0% reduce under the compiled prompt; llama-4-scout flat at 0.0% and qwen3 at 7.1%. llama again trades named-grid and type-pairing for line-height and radius rather than reducing. Repeating two weeks running makes this a stable pattern, not a one-off.

Five models · Three reduce · Two can't follow

First run on the automated weekly cadence. Five Cloudflare Workers AI open-source models, n=30, source-linter only. Three reduced cleanly under the compiled prompt: gpt-oss-120b 72.6%, mistral-small-3.1 68.6% and gemma-4 53.8%. Two stayed flat: llama-4-scout at 0% and qwen3 at -3.4%. The flat cells reflect the models, not the tooling. llama-4-scout trades tells, dropping require-named-grid and…

Eleven models · Same brief · Different token

The different-token-same-brief triangulation queued by the 22 April report. Eight of eleven cells regress under the compiled prompt. The compiler is not at fault: it transmits the token faithfully. The regressions come from lint rules that assumed the editorial defaults this token rejects, so they penalised output that followed the token. gpt-5.5 lands as the cleanest raw frontier baseline…

Ten models · One brief · Thirty samples each

Ten models, n=30 per cell, 600 samples. Eight of ten cells showed positive reductions under the compiled prompt, one flat, one regression. Best: gpt-oss-120b at 78.1% fewer tells. Three frontier cells via subscription CLIs (Claude Code, Codex, Gemini CLI), seven OSS via Cloudflare Workers AI. Wilson interval tightens from roughly +/-35% at n=5 to roughly +/-18% at n=30.

Seven models across four providers

Four positive reductions, one inconclusive, two regressions. Llama 3.3's regression reproduces across Cloudflare and Hugging Face, turning a single-cell finding into a cross-provider result.

Five models, n=5, zero errors

Claude Opus (Anthropic API) plus four OSS models on Cloudflare Workers AI. Claude dropped to zero tells compiled; Llama 3.3 70B regressed. Full per-model and per-tell breakdown with every attempted-vs-scored count published.