This is where the final score lands, and what it says about four months of work. I finished four tenths of a foot from a medal, and everything I built in the second half was worth less than that gap.
A consultant sold us a six page strategy report that was generated, not written. The interesting question was never why he failed. It was why the company was so confident about what it was buying and so unclear about what it already had.
Six weeks before the deadline I worked out exactly how my models would fail, and built a hedge against it. The hedge was drawn from the same family as the thing it was hedging. Here is the question I should have asked instead.
For two months I optimized a score that had stopped ranking my models. It was still a valid measure of quality. It had just gone blind at the resolution I was working at, and three lines of code would have told me.
The competition closed and I pulled my final rank. The number was internally consistent, completely plausible, and wrong by nine hundred places. Here is the two-word API difference that caused it and how a competitor's blog post caught it.
Building on the Moon means making what you need from the dust already underfoot. A permanent settlement cannot import every beam and panel from Earth, and the materials to do it may be designed by AI before anyone builds them.
I pointed a 22-agent review at my own competition endgame expecting notes about the code. It came back with three errors in my reasoning instead, and all three were experiments that could only ever confirm what I already believed.
These systems learn from the open internet, and the open internet is getting worse on purpose. The quality of what goes in is falling while public trust in what comes out keeps climbing. Nobody is watching that gap.
Someone used our business information to make a fraudulent transaction look legitimate. No malware, no stolen passwords, no alert. They studied how the process worked and walked through the front door of trust.
The public leaderboard scores half the test set and says so on the page. I measured which half. It was not a random sample, it was the first 1,974,995 rows in time order, and that one fact invalidated six weeks of tuning.