I often get asked whether I think AI is a bubble. Usually because I work in an AI-native company. I took Andrew Ng’s machine learning course back when he still used Octave, and started experimenting with agentic frameworks 2 years ago. My primary thought: this is a revolution and we don’t yet know what the effects will be. No, “AI” is not a bubble. But… Yes, there is an AI bubble. The bubble isn’t…
I wish more product teams were asking themselves: should we be offering the user more “AI”, or should we be ensuring our product works well when they are using agents alongside it or even as an intermediary?
The recent Dwarkesh Podcast episode about AI “learning on the job” has no news or discoveries but provokes some thoughts: The true believers of AGI probably are also adherents to the Great Men of History view. Of course this view tends to correlate with seeing the PayPal Mafia boys as once-in-a-generation geniuses, meat-based AGI. I find it more optimistic to see them as competent-enough nerds who…
When I first started using coding agents: “Recent changes have caused an error rendering the template file /path/to/file.html . Please review the following error message and suggest necessary corrections. [relevant portion of stack trace]” Today: “Oops! [lazy copy-paste of stack trace]”
SpaceX is doing amazing work, and is full of engineers I really want to celebrate. I would find the usual regulatory capture antics gross, but that’s not why I cannot enjoy their achievement. It’s that every success of these brilliant people is directly empowering their actively-fascist oligarch. And I despise him all the more for this.
I have been trying models from 3 providers on 2 very different codebases and using 4 different agent/IDE toolchains. For any given task, all of these variables figure into which setup works best. Anyone who asserts “this is the best model for coding” is oversimplifying, grifting, or has found an optimal configuration that eludes me.
There is a growing body of research into LLM self-awareness but I am particularly fascinated by the “evaluation awareness” research by Anthropic into the Hawthorne effect in LLMs. (h/t to @futureshape.bsky.social for the connection)
Today in “things my coding model said”: This is a great example of premature optimization - the debouncing added complexity to solve a performance problem that doesn’t actually exist at the scale we’re operating. (It wrote the debouncing, until I asked it to simplify.)
As I enjoy watching models explain their reasoning through work I’ve given them, I keep thinking about the recent paper _ Chain-of-Thought Is Not Explainability and how much we should remind people that rationalization is not always insight: verbalised chains are frequently unfaithful, diverging from the true hidden computations that drive a model’s predictions, and giving an incorrect picture of…
My AI coding agent has been well coached. Describing a few discarded options for a caching architecture: Redis External Cache : Adds infrastructure complexity (against smol tech philosophy)
Over four months, LLM users consistently underperformed at neural, linguistic, and behavioral levels. – Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task
A lot of enterprise software design has always been serving users that work on behalf of others (customers, managers, etc.). It seems we should be able to transfer some of that know-how to designing for AI agents using software on behalf of users. (No, I don’t know how, exactly.)
I recently let Cline (I don’t remember which model) do a refactor that seemed quite harmless. Especially in Rust. But days later, I’m trying to figure out a bug and find that a large bit of code had been replaced with // ...rest of the function implementation... 🤦
It’s no surprise that training an LLM on material that demonstrates sound reasoning results in better performance, but I’m fascinated by this identification of “cognitive” behaviors (verification, backtracking, subgoal setting, backward chaining). They sound obvious but I also don’t recognize them from any study of human reasoning: https://nitter.net/gandhikanishk/status/1896988028893323675
“En ik zie alleen maar amateurisme.” “Minder op Twitter uiten en gewoon op een constructieve manier kijken hoe je het beste doet voor Nederland.” I love to see this multiparty democracy enforcing a meaningful balance of power . These days in particular.
In the way the dismantling of USAID is a practice run for DOGE dismantling the US government, I’m growing concerned the birthright citizenship attack is just a first step towards weakening the 14th amendment.
I find it worrying how often Github Copilot offers to complete my Rust database access code using .execute() , which is a method on the very popular sqlx crate, which I am not using. It feels like I’m constantly being judged by a dumb popular kid for making a less popular choice (in my current case, cornucopia). Our LLM friends could really regress us to the mean more subtly than media…
My biggest fear with AI is that we’ve reached a plateau. If we’re stuck with “smart enough to bullshit” but spend a decade struggling to actually automate meaningful solutions, I fear for human culture.
Germany, of all nations, should be keenly aware of the dangers of forgetting state attrocities. So it’s especially upsetting to see Berlin supporting Japan’s effort to censor history and remove memory of the ianfu. https://www.theleftberlin.com/berlin-mayor-insists-on-removing-statue-for-victims-of-sexual-slavery/
I’m old enough to remember when voting for local offices required researching the candidates because you couldn’t count on whether the Democrat or the Republican was the batshit crazy one.
Unfortunately, it seems that today’s big LLMs are more like physicists than philosophers. Rather than more knowledge improving their awareness of their own ignorance, it only makes them more cocksure about knowing everything .
I meet so many aspiring designers who proudly assert that they’re really good at “ideas”. The world does not need more idea people, at least until it catches up on the inspiring backlog of Matt Webb’s ideas. We need to make home-defragging robots before dreaming up new things to buy.
Klarna is going DIY on software thanks to AI . I bet if you’re on a dev team there, you are either: frustrated that your productive colleagues and “ignore the bullshit” team are having their good work credited solely to their use of copilots, or scared because you know it’s time to find another job and all you know how to do is consume user stories.
I just re-read this classic post about leadership in the metaverse and it strikes me how similar the experience seems, from my outsider perspective, to running a large open source project. Is OSS a metaverse?
Open Interpreter’s Local Model Computer protocol feels very important. Apple and Microsoft have struggled to bootstrap ubiquitous scripting APIs despite there being obvious user benefits. And not just to nerds; the Kin and Newton (and, farther back, BeOS) showed the tremendous UX potential of a UX-strong OS. They also demonstrated how difficult it is to sustain an ecosystem of developers willing…
Just because I think the tech is revolutionary doesn’t mean I disagree that most of today’s products are bullshit: https://www.wheresyoured.at/expectations-versus-reality/
I hope something comes of this. Marking up LLM output to explicitly quite source material, so even if the model hallucinates the reader can discern: https://mattyyeung.github.io/deterministic-quoting
A parable of our decade in two pieces of nearly-identical hardware: In 2012 we had the Descriptive Camera using Mechanical Turk, as a thought provoking student project and a blog post. In 2024 we get the Poetry Camera using AI, with a slick website and signups for buying one.
I see a lot of experts-of-the-day gnashing teeth over @rabbit@threads.net being “exposed” for using puppeted Android apps in the cloud. What did you all think they were using? Hand-wavy AI magic you’ve been pretending to understand? Some secret Android API that app developers have been quietly implementing just for them? This is the obvious way to pull off what they’re doing with today’s…
Back in 2011 when Microsoft demoed them, I didn’t see that communication avatars (sorry, “Personas”) would be useful for AR; strapping an iPad to my face was a distant dream. I think I still believe the real value will be in facing emotional states, though. https://twitter.com/gerwitz/status/84311406148718593