Four days after I published a post about machines winning at math by finding counterexamples, OpenAI published ten results that counterexampled the post. Here's what broke, what held, and what we're still squinting at.
An 87-year-old math problem died in a tweet during the World Cup final, and the same week a friend asked me how to get AI to stop lying to her. It took me a while to notice these have the same answer.
We were hiring an engineer this summer. The technical interview couldn't be 'write me a function' anymore, because the intern writes the functions now. So we built a repo that lies to the AI, and watched who noticed.
The field manual for the overconfident intern. The actual decisions I make before and during the work, and the one-page map I wish someone had handed me two years ago.
You didn't get a magic tool. You hired an intern who has read everything, never says 'I don't know,' and never sticks around for the consequences. Here is how not to get fired for what It does.