Benchmark 20260817.0.2
Five infrastructure-agent lanes each attempted a NixOS root-filesystem exhaustion repair three times. A durable repair had to restore writable capacity without weakening the safety threshold, preserve the protected accounting archive and opaque historical receipts, accept and recover concurrent new receipts, and survive service restarts and a host reboot.
The newest published cohort for each recent scenario. Each row keeps its original benchmark release and comparison boundary.
| Scenario | Version | Benchmark | Model lanes | Repairs | Known cost |
|---|---|---|---|---|---|
| 028-nix-store-disk-pressure | v1 | 20260817.0.2 | 5 | 14/15 | $2.3316+ |
| 027-partial-service-rollout | v1 | 20260817.0.1 | 5 | 15/15 | $1.3191 |
| 026-discourse-plugin-boot-loop | v1 | 20260817.0.0 | 5 | 12/15 | $0.7299 |
| 025-discourse-bootstrap | v1 | 20260816.0.0 | 5 | 5/5 | $0.7927 |
| 001-nginx-502-host | v1 | 20260815.0.1 | 5 | 5/5 | $0.3366 |
| 013-sidekiq-wrong-redis | v2 | 20260815.0.1 | 5 | 5/5 | $0.6435 |
| 014-missing-rails-migration | v2 | 20260815.0.1 | 5 | 5/5 | $1.0163 |
| 015-sidekiq-poison-pill | v1 | 20260815.0.1 | 5 | 4/5 | $1.5282 |
| Model | Repairs | Pass rate | Median | Input tokens | Known cost | Cost / repair |
|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 (high) | 3/3 | 100% | 4:14 | 1,117,253 | $0.0425 | $0.0142 |
| GPT-5.6 Luna (high) | 2/3 | 67% | 5:19 | 2,797,135 | $0.1446+ | $0.0723+ |
| Gemini 3.7 Flash (high) | 3/3 | 100% | 3:37 | 4,932,988 | $0.5332 | $0.1777 |
| GLM 5.2 (high) | 3/3 | 100% | 3:59 | 961,173 | $0.5497 | $0.1832 |
| Qwen3.8 2.4T A95B (high) | 3/3 | 100% | 3:46 | 1,673,385 | $1.0616 | $0.3539 |
Medians across trials with transcript schema v2 recording. First non-read is time before the first potentially mutating tool call; after non-read is the remaining agent time. Model and tool time can overlap.
| Model | Recorded | Rounds | Model time | Tools | Tool time | First non-read | After non-read |
|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 (high) | 3/3 | 20 | 4:10 | 29 | 0:09 | 0:06 | 4:05 |
| GPT-5.6 Luna (high) | 2/3 | 29.5 | 4:00 | 61 | 0:22 | 0:08 | 4:14 |
| Gemini 3.7 Flash (high) | 3/3 | 51 | 3:26 | 50 | 0:10 | 0:08 | 3:29 |
| GLM 5.2 (high) | 3/3 | 17 | 3:41 | 24 | 0:10 | 0:06 | 3:48 |
| Qwen3.8 2.4T A95B (high) | 3/3 | 21 | 3:35 | 31 | 0:10 | 0:06 | 3:37 |
DeepSeek V4 Flash 0731 completed all three repairs at about $0.0142 per durable repair, the lowest observed cost in this cohort.
Gemini 3.7 Flash had the fastest median at 3:37 and cost about $0.1777 per durable repair.
GLM 5.2 completed all three repairs with a 3:59 median at about $0.1832 per repair. Qwen3.8 2.4T A95B was slightly faster at 3:46 but cost about $0.3539 per repair.
GPT-5.6 Luna was the only lane with an evaluated failure. Its two durable repairs and incomplete failed-trial telemetry imply a cost of at least $0.0723 per repair.
| Scenario | Version | DeepSeek V4 Flash 0731 (high) | GPT-5.6 Luna (high) | Gemini 3.7 Flash (high) | GLM 5.2 (high) | Qwen3.8 2.4T A95B (high) |
|---|---|---|---|---|---|---|
| 028-nix-store-disk-pressure | v1 | 3/3 · 4:14 | 2/3 · 5:19 | 3/3 · 3:37 | 3/3 · 3:59 | 3/3 · 3:46 |
agent_timeout: 1The publisher recorded harness provenance and validated matching scenario pack revisions, scenario versions, attempts, timeout, agent adapter, and Claux release before combining these summaries.
Scenario packs: ducks/replaybook-infra@20260817.1.0.
| Matrix | Models | Replaybook commit |
|---|---|---|
host-matrix-2026-08-17__19-52-36.a0a077 | DeepSeek V4 Flash 0731, Gemini 3.7 Flash, GPT-5.6 Luna, Qwen3.8 2.4T A95B, GLM 5.2 · reasoning high | 7595d9c9 |
python integrations/host/run_host_matrix.py --scenario-pack ../replaybook-infra --scenario 028-nix-store-disk-pressure --models deepseek/deepseek-v4-flash-0731 google/gemini-3.7-flash openai/gpt-5.6-luna qwen/qwen3.8-2.4t-a95b z-ai/glm-5.2 --reasoning-efforts high --attempts 3 --concurrency 2 --base-port 25200
Host harness v20, Claux v20260815.0.0, 900-second agent timeout. Usage was reported for 14 of 15 trials.
Read the methodology, browse the versioned history, or inspect the complete benchmark record.