pocoo · Jul 24, 2026
The eval said fail. The baseline said the eval was wrong.
0Sign in to vote or save
This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.
A 25-hour, $100+ SFT run on Qwen2.5-Coder-32B came back with a FAIL verdict. Before retraining, we measured the un-tuned baseline against the same harness — and found the pass threshold had never been reachable by anything. Recalibrated the gate with real numbers, fixed a real publish bug, recovered the already-trained model with zero retraining, and made everything — weights and data — public by…
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.