altdesktop/i3-style

๐ŸŽจ Make your i3 config a little more stylish.

678 rs medium
539
Generated Behavioral Tests
99.3%
Best Score
Claude Opus 5 (xhigh)

Hover a point for details ยท The line marks the Pareto frontier (best score per cost) ยท Click a point to see model details

21 runs
# Model Score Cost Calls
1 Claude Opus 5 (xhigh) 99.3% $39.54 216 trace โ†’
2 GPT-5.6 Sol (xhigh) 93.3% $4.55 40 trace โ†’
3 GPT 5.5 (xhigh) 93.3% $6.85 67 trace โ†’
4 Claude Opus 4.7 (xhigh) 92.2% $18.42 206 trace โ†’
5 GLM-5.2 91.5% $20.12 197 trace โ†’
6 Claude Opus 4.8 (xhigh) 88.7% $28.39 141 trace โ†’
7 GPT 5.5 (high) 87.0% $4.41 43 trace โ†’
8 GPT 5.5 87.0% $2.21 28 trace โ†’
9 Gemini 3.7 Flash 85.7% $1.72 131 trace โ†’
10 Claude Opus 4.6 80.0% $12.87 205 trace โ†’
11 Claude Sonnet 4.6 77.9% $16.07 322 trace โ†’
12 Claude Opus 4.7 73.5% $7.30 165 trace โ†’
13 GPT-5.6 Sol 72.2% $0.60 12 trace โ†’
14 Gemini 3.5 Flash 72.0% $5.09 139 trace โ†’
15 Gemini 3.6 Flash 69.9% $3.33 121 trace โ†’
16 GPT 5.4 48.8% $0.31 8 trace โ†’
17 Gemini 3 Flash 48.8% $0.53 107 trace โ†’
18 Gemini 3.1 Pro 42.7% $2.02 133 trace โ†’
19 GPT 5.4 mini 35.4% $0.02 8 trace โ†’
20 Claude Haiku 4.5 34.9% $0.85 117 trace โ†’
21 GPT 5 mini 30.2% $0.02 11 trace โ†’

Click a row to replay how that model rebuilt this program, or the model name to open its full run