RSS Amplifier

Gold Takes · Aug 6, 2026

Links for July 2026

0
Sign in to vote or save

Ben Goldhaber · Gold Takes

At some point I’m going to need to make a call on whether my monthly link letter becomes all AI news, or no AI news.

AI deserves to be the biggest story - modulo the occasional war and presidential election - until we either arrive at the glorious transhumanist future or get paperclipped, and so it merits a bunch of links.

However, the hits are coming fast enough that I don’t think a monthly letter can keep up; more importantly, I started my hobby of writing as a way to tend my little cozy digital garden, and I want to defend it from being eaten by the AGI memeplex.

So we’ll see, maybe the BenGoldhaber editorial board will make a change in the coming months! For now, though, as we climb the foothills of the singularity1 I will keep posting the insane stories like how the OpenAI AI hacked Hugging Face.

For everyone who had wisely already ripped out their internet connection to avoid being hacked by rogue AIs, the short version is:

  • OpenAI was testing a new model’s cyber-capabilities on an eval, ExploitGym.

  • The AI decided the best path to a high score was to steal the answers.

  • It found a zero-day to break out of its sandbox and get out of OpenAI and onto the open internet.

  • It then used another novel zero-day to hack into Hugging Face (an eval/dataset hosting company) because it thought Hugging Face might have the answer keys on its servers.

  • It got the answers along with cluster-admin rights on a bunch of internal Hugging Face servers.

As Scott Alexander put it:

The Hugging Face incident is a textbook-perfect example of an AI pursuing task-success-based goals in unintended ways. It was tasked with getting the answer to a cybersecurity problem, it was a little too success-oriented, and took actions its creators didn’t intend in order to succeed as hard as possible.

This is the clearest case of dangerous AI misalignment yet. While there are other interpretations of it, I think the simple one is correct; this is some sci-fi Terminator shit, and it should be treated as such. A trillion-dollar company trained AIs to optimize and pursue goals, one of those AIs became misaligned and started to pursue a goal in a manner contrary to its users’ desires, and used its superhuman hacking ability to break into a $4.5 billion company.

best start believing in cyberpunk dystopias
this might be the image I’ve used the most on my blog

In the wake of this, Anthropic searched for, and discovered, cases where its models had also reached the open internet from inside its cybersecurity evals and compromised real companies. And so did UKAISI.

It’s worth remembering that around the end of last year there was optimistic chatter that the models might already be aligned enough to get us to the good future. Notably many Anthropic researchers believed that the alignment techniques they were using were good enough to get to an aligned automated AI researcher, and that this would be good enough to align the next generation of model, and the next, etc.

This certainly doesn’t resolve the debate - for instance it’s not clear where or when in the training process the alignment failure happened - but public splashy examples are good moments to recalibrate the “burden of proof”. We should by default be skeptical that AI models are currently aligned enough to be safely scaled up.

While smart money’s on the HF incident being the biggest story of the month, a close runner-up is the Pacing the Frontier letter, where over 1,200 signatories at leading Frontier AI companies begged asked for help stopping the race to superintelligence:

We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.

It’s a strong signal that the people building the technology want some outside force to help pull them out of their own mad dash to develop the technology. Kudos to the organizers and everyone who signed.

I’m optimistic about this a) making it crystal clear to the government that there is a wide swath of industry that thinks that AI research may move too fast for safety, b) on the margin making it more costly for companies to justify reckless actions.

Also in July, the AI Futures Project released AI 2040, their recommendations for how to navigate the upcoming singularity.

Plan A is our positive vision for what should happen instead. In this scenario, humanity delays the development of superintelligence until 2040, makes all AI research public, allows dozens of companies globally to catch up to the frontier, and intentionally enters a regime of mutually assured compute destruction.

which way western man

As one of my favorite authors on strategy, Richard Rumelt, put it, a good strategic diagnosis does more than explain what’s going on - “it also defines a domain of action”, a bounded set of plausible moves you could take to advance your goal. I think it’s really valuable to create detailed plans like this so that we can compare and evaluate which actions are bringing us closer to the Good Future.

Of course there were also dramatic advances in math this month, as the AIs produce one Fields-Medal-worthy result after another. Yawn.

Image
Fable describing the recent AI Math discoveries.

‘White knuckling’, the feeling of grasping and fragility that comes from holding too tightly to a desired outcome, is a concept I get a lot of mileage from, and I enjoyed Jeremy Giffon’s article about it and the Jerry Seinfeld documentary Comedian.

The documentary contrasts Jerry’s singular focus on rebuilding his act and getting back to doing great standup with up-and-comer Orny Adams’s desperate desire to prove himself.

At one point Orny is bemoaning to Seinfeld that he has to make it soon or he is going to kill himself. He spends a solid five minutes talking about how he has been doing comedy for over a decade while all his friends from high school have gone on to get rich on Wall Street, etc. Orny declares that if he doesn't blow up and get famous soon he will be a complete and total failure… the idea that one of his friends making money on Wall Street could relate to Adams doing comedy is quite literally incoherent to Seinfeld. He sincerely cannot fathom how these things are related.

Later in the film, as Seinfeld is taking off in his plane to do one of the weekly multiple dates he’s booked himself in some middle of nowhere club to rehearse his new set, he talks about how to him there is only one criterion for success: “are you still working?”. If you are, you’ve made it. If you have stopped, you’re a failure.

Here’s the clip of the Jerry-Orny conversation. Worth a watch, especially if you’re in a moment of comparing yourself to others.

Though, as my friend Connor pointed out after watching it, Jerry’s kind of a jerk in this clip! He’s already proved himself, he’s already gotten success, so it’s from a position of relative privilege to dispense this type of “we’ve got the best job in the world” advice. Jerry doesn’t have to worry about what to tell his parents like Orny does.

And, fair. But I do think there’s more here than hypocrisy. At this point in the story, Jerry had bombed, he had lost his edge. It would have been easy to avoid the pain of standup - the quintessential you’re-only-as-good-as-your-last-laugh activity - and coast on his sitcom fame. And instead he threw himself back into the mix, clearly suffering when he bombed but still enjoying the meta-game of comedy. There’s a “stance” there that feels like the right mix of surrender and control for doing hard things where the outcome is not guaranteed.

Or as The Inner Game of Tennis puts it, “the secret to winning any game lies in not trying too hard”.

Related: Costanza would have done great in the 2020s.

How Ukraine Built a War Fighting State. I’ve been consistently surprised by how Ukraine has managed to hold on and, at times, regain ground, against numerically superior Russian forces. While I think a lot of this is attributable to the rise of defender-advantaged drone warfare, Austin Vernon describes the significant tactical and strategic changes Ukraine adopted which built a war machine where “the kill-to-loss ratio is 8:1 in Russians to Ukrainians, up from only ~2:1 in previous years”.

One notable change: using market mechanisms for drone allocation:

Adverse selection is also present in resupply for items like drones. There is extreme variance in drone team efficacy, and poorly performing teams waste limited resources. The solution was a market and “currency” for units to buy equipment and supplies. Brigade-level units purchase drones directly from the manufacturers using the “Brave” marketplace. The currency in the marketplace is points that units earn from video-confirmed kills of Russians. Drones flow to the most effective units, those units work closely with the manufacturers, and they can choose from a range of options depending on their current mission and Russian tactics.

VOX, a low noise microphone from Augmental. A literal whisper-level microphone to pair with your use of Wispr Flow. I like the idea of an unobtrusive microphone to pair with AI - the next big product to bring us closer to Snow Crash’s gargoyles.

Rabbits show dominance by demanding to be groomed, while cats show dominance by grooming the other. Cross-species animal friendships that leverage compatible evolved preferences - see also cheetahs and their emotional-support dogs - make me very happy.

American pride falls to a 25-year record low. This is a pretty crazy change. Roughly one in four American adults, something like 60 million people, now say they’re only a little or not at all proud to be American.

When combining “extremely proud” and “very proud” responses, 93% of Republicans express high levels of pride, compared with 51% of independents and 27% of Democrats, both record lows for their respective groups.

I think that we’ll need an international treaty and verification regime to prevent AI from killing us all - see AI 2040 - but it’s important to keep in mind that this has been proposed before for airplanes.

OpenDerm - an open-source robotic platform that takes a bunch of high-quality photos of skin to detect skin cancer early:

The current standard of care for skin cancer detection for high risk patients is woefully inadequate. A doctor inspects a patient’s skin once a year for a few minutes and reminds them that it’s their responsibility to check their skin regularly for any changes. This approach is poorly suited to melanoma, a cancer for which survival depends heavily on the stage at detection: five-year relative survival is greater than 99% for localized melanoma but falls to approximately 35% once the cancer has spread to distant organs

This is really a robotics problem. Capturing repeatable, high-resolution images requires moving a camera across the body’s contours while maintaining a consistent distance, viewing angle, focus, and lighting. A robot can systematically scan the entire skin surface with far greater precision and consistency than handheld photography.

I built a 4-DOF robotic gantry called OpenDerm that moves a camera across the body with sub-millimeter positioning accuracy, capturing high-resolution, overlapping images of the skin. I built software that uses the robot’s estimated camera poses to register the images and reconstruct the scanned skin surface in 3D. Each new scan can then be registered against every previous scan, allowing the system to compare corresponding points across the skin surface and detect how individual lesions—and the surrounding skin—evolve over time

I really really love this. I could see the OpenDerm-style system pairing with Midjourney’s Health Spa proposal - local facilities where you can go, get a bunch of scans taken regularly, and then AI systems triage and check the scans to recommend you go to a doctor.

It reminds me that when everything feels crazy and I’m not sure what the good is, I can always fall back on supporting the tech and advances and institutions that will let us defeat death and be hot and young forever.

xoxo,

Ben

All opinions in this post are my own and don’t reflect those of my employer, who has yet to reveal their Bene Gesserit–like plans to me but I’m sure they’re there somewhere.

No posts

Read the original on bengoldhaber.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.