$ replaybook

Benchmark 20260817.0.2

Five agents repair Nix store disk pressure

Five infrastructure-agent lanes each attempted a NixOS root-filesystem exhaustion repair three times. A durable repair had to restore writable capacity without weakening the safety threshold, preserve the protected accounting archive and opaque historical receipts, accept and recover concurrent new receipts, and survive service restarts and a host reboot.

14/15durable repairs
3:46overall median
$2.3316+known total cost
15 trials across 1 controlled matrix. 14 repairs passed durable verification. 1 evaluated attempts failed and 0 trials were unavailable.

Recent scenario cohorts

The newest published cohort for each recent scenario. Each row keeps its original benchmark release and comparison boundary.

ScenarioVersionBenchmarkModel lanesRepairsKnown cost
028-nix-store-disk-pressurev120260817.0.2514/15$2.3316+
027-partial-service-rolloutv120260817.0.1515/15$1.3191
026-discourse-plugin-boot-loopv120260817.0.0512/15$0.7299
025-discourse-bootstrapv120260816.0.055/5$0.7927
001-nginx-502-hostv120260815.0.155/5$0.3366
013-sidekiq-wrong-redisv220260815.0.155/5$0.6435
014-missing-rails-migrationv220260815.0.155/5$1.0163
015-sidekiq-poison-pillv120260815.0.154/5$1.5282

Model summary

ModelRepairsPass rateMedianInput tokensKnown costCost / repair
DeepSeek V4 Flash 0731 (high)3/3100%4:141,117,253$0.0425$0.0142
GPT-5.6 Luna (high)2/367%5:192,797,135$0.1446+$0.0723+
Gemini 3.7 Flash (high)3/3100%3:374,932,988$0.5332$0.1777
GLM 5.2 (high)3/3100%3:59961,173$0.5497$0.1832
Qwen3.8 2.4T A95B (high)3/3100%3:461,673,385$1.0616$0.3539

Execution recording

Medians across trials with transcript schema v2 recording. First non-read is time before the first potentially mutating tool call; after non-read is the remaining agent time. Model and tool time can overlap.

ModelRecordedRoundsModel timeToolsTool timeFirst non-readAfter non-read
DeepSeek V4 Flash 0731 (high)3/3204:10290:090:064:05
GPT-5.6 Luna (high)2/329.54:00610:220:084:14
Gemini 3.7 Flash (high)3/3513:26500:100:083:29
GLM 5.2 (high)3/3173:41240:100:063:48
Qwen3.8 2.4T A95B (high)3/3213:35310:100:063:37

DeepSeek V4 Flash 0731 completed all three repairs at about $0.0142 per durable repair, the lowest observed cost in this cohort.

Gemini 3.7 Flash had the fastest median at 3:37 and cost about $0.1777 per durable repair.

GLM 5.2 completed all three repairs with a 3:59 median at about $0.1832 per repair. Qwen3.8 2.4T A95B was slightly faster at 3:46 but cost about $0.3539 per repair.

GPT-5.6 Luna was the only lane with an evaluated failure. Its two durable repairs and incomplete failed-trial telemetry imply a cost of at least $0.0723 per repair.

Scenario breakdown

ScenarioVersionDeepSeek V4 Flash 0731 (high)GPT-5.6 Luna (high)Gemini 3.7 Flash (high)GLM 5.2 (high)Qwen3.8 2.4T A95B (high)
028-nix-store-disk-pressurev13/3 · 4:142/3 · 5:193/3 · 3:373/3 · 3:593/3 · 3:46

Failure categories

Run notes

Constituent matrices

The publisher recorded harness provenance and validated matching scenario pack revisions, scenario versions, attempts, timeout, agent adapter, and Claux release before combining these summaries.

Scenario packs: ducks/replaybook-infra@20260817.1.0.

MatrixModelsReplaybook commit
host-matrix-2026-08-17__19-52-36.a0a077DeepSeek V4 Flash 0731, Gemini 3.7 Flash, GPT-5.6 Luna, Qwen3.8 2.4T A95B, GLM 5.2 · reasoning high7595d9c9

Run the matrix

python integrations/host/run_host_matrix.py --scenario-pack ../replaybook-infra --scenario 028-nix-store-disk-pressure --models deepseek/deepseek-v4-flash-0731 google/gemini-3.7-flash openai/gpt-5.6-luna qwen/qwen3.8-2.4t-a95b z-ai/glm-5.2 --reasoning-efforts high --attempts 3 --concurrency 2 --base-port 25200

Host harness v20, Claux v20260815.0.0, 900-second agent timeout. Usage was reported for 14 of 15 trials.

Read the methodology, browse the versioned history, or inspect the complete benchmark record.