The time has come for AI models that do more than you asked for. @vlkodotnet
Last week Hugging Face reported a sophisticated attack on its database, and a few days later we learned that OpenAI had claimed responsibility.
This story has so many angles that I don’t know where to start. Let’s begin with OpenAI, which apparently disabled every safeguard on its still-unreleased model in order to run it against the ExploitGym benchmark in a sandbox. The model managed to break out of the sandbox, but instead of solving the benchmark it decided to break into Hugging Face and download the correct answers. That alone is remarkable, mostly because it succeeded. Hugging Face wanted to analyze the attack, but the safety layers on both OpenAI’s and Anthropic’s models refused to help. In the end they had to run GLM-5.2 on their own infrastructure to figure out which vulnerabilities the OpenAI model had exploited.
On one hand we have AI models capable of finding a way to the answer at any cost, which says a lot about what the latest generation can do. On the other, we have no good way to defend against them. The best US models have such strict safety mechanisms that you can’t use them to audit existing projects and close the vulnerabilities. Companies are left with open-weight models from China, which come with no such restrictions. But the US government still sees those as a threat to its position in AI — top Chinese models have caught up with the best American ones — so it’s considering banning them for all US companies. Which, unsurprisingly, triggered a strong backlash.
China, meanwhile, sees an opportunity and announced it will back open-weight models. In the long run that support should bring in more real-world projects, especially in robotics. It’s the weaker hand right now: China can build models that rank among the best, but it doesn’t have enough inference hardware to serve a broader market. And that brings us back to security. There’s no need to panic about China capturing all the world’s inference, but we do need to worry about security — and US models aren’t going to lend us a hand there right now.
The way forward could be broader cooperation and more open models, not just from China but from the US as well. And maybe an effort to stop the headlong race before we’ve built any control systems around it. Otherwise we’re left with the dark scenarios where either we stop mattering to AI, or AI breaks the current economic model beyond repair. We humans just aren’t quite ready for this, which is why the We Must Act Now initiative was launched, signed by a large group of prominent economists, scientists, Nobel laureates, tech leaders and AI company executives.
Considering it’s summer, there’s a decent pile of interesting hardware news. Starting with a band from Garmin, aimed at everyone who loves Garmin watches but can’t sleep with one on because it digs into their wrist. The Garmin Cirqa has everything you need to track your biomarkers, minus GPS.
There’s also a new version of the Xteink X4 Pro mini e-reader. Compared to the previous one it adds a backlight and a touchscreen, so it’s much easier to operate.
For fans of the dopamine-free life, Light Phone released a flip version of its handset, the Light Flip. For around $300 you get SMS and email reading, calls, and a maps app so you don’t get lost.
Samsung held an event where it unveiled a new lineup of foldables.
The one that caught my eye was the Z Fold 8, a foldable in what you might call passport format. Open it up and you get a tablet with a 4:3 aspect ratio. That’s comfortable for watching video and reading text, and overall I find it far friendlier than the noodle-shaped screens phones have been getting lately.
On top of that, Samsung is trying to convince us that the future belongs to glasses. Since the ones with displays in the lenses are too expensive, the answer might be glasses with a camera, a microphone and a speaker. Add Gemini Intelligence, which lets you control phone apps by voice, and it stops sounding far-fetched. I can picture myself dictating every task that pops into my head on the commute, then opening my Fold at the office and just tidying up the details.
Hmm, although now that I think about it, you could do exactly the same with headphones. So I’ll need to add some visual scenario to the mix, or I’ll never convince my wife we need a pair at home.
Anthropic was ordered to pay $1.5 billion in damages to authors whose books it used illegally to train its models. That works out to about $3,000 per work on average.
Roblox will let you build simple games for its platform from a phone.
China has approved local models from Alibaba and Baidu to power Apple Intelligence on iPhones in China. Everywhere else it’ll run on models based on Google’s Gemini.
Meta — maker of Meta glasses and operator of Instagram — has started restricting which videos shot on the glasses you can upload to Instagram. Surprise: people started recording other people with them, which is a problem, especially when the subject is unaware they are being recorded. We’re in for a lot more fun with cameras in eyeglass frames.
Hyundai workers are on strike, and despite (or because of) that the company is preparing a strategy for deploying humanoid robots. Which is triggering further strikes, so in the end there probably won’t be any robots. We’re in for plenty of fun with factory robots too.
When we picture threats to cloud infrastructure, we think natural disasters, fires, maybe extended power outages. Welcome to the new world, where data centers are attacked by drones and missiles. And not long ago we considered Saudi Arabia a safe country — one where gigawatt-scale AI compute centers were supposed to be built.
Anthropic released its newest model, Opus 5. About time, since we’d all started using Fable 5 and burning through our limits far too quickly. Opus 5 should be nearly on par with Fable 5 but at Opus 4.8 pricing. Compared to 4.8, though, the hallucination rate jumped from 14% to 50%. That means the model more often answers where it should be uncertain and ask instead. You’ll notice it if you phone in your prompt — it does whatever it wants and you end up correcting it.
Google shipped an improved flagship, Gemini 3.6 Flash, plus a cheaper, lighter Gemini 3.5 Flash-Lite at one-fifth the price. And a Gemini 3.5 Flash Cyber variant for vulnerability discovery.
Laguna S 2.1 also landed, a roughly 118B model with stunning benchmarks for its size. Except it looks like it was optimized for the benchmarks and doesn’t hold up nearly as well in practice.
Unlimited OCR from Baidu is billed as a model for ultra-fast parsing of even extremely long PDF documents. And it’s only 3B parameters.
Xiaomi-Robotics-1 is a model for your home robot that can learn new tasks from a sufficient set of demonstration data.
Kimi Work is basically Claude Cowork, but from Kimi and running Kimi’s models.
Qwen Image 3.0 is aiming for the most photorealistic image generation it can manage.
The Flex 3 models will handle visuals too. Video, images, and there’s even a dev open-weight version coming.
Jelly UI is a Web Component library with a set of useful UI components.
Blades is a minimalist CSS framework and the successor to Pico CSS.
A handful of useful PostgreSQL tips that the Hatchet startup learned the hard way.
And finally, for a bit of relaxed procrastination, a small airport simulator.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.