RSSAmplifier

Blog

James Padolsey's Blog

Random writings of James Padolsey

blog.j11y.ioRSS feed ↗29 posts

Latest posts

Anthropic’s weak watermarks appease a weak law

Anthropic’s weak watermarks appease a weak law Anthropic announced in August 2026 that text generated by supported Claude models will carry invisible, machine-readable watermarks, applied at model level across supported Claude products and API surfaces worldwide. The move is Anthropic’s response to Article 50(2) of the EU AI Act , which requires providers to ensure that outputs are “marked in a…

Using git diffs to circumvent Fable’s safeguards

Using git diffs to circumvent Fable’s safeguards Fable, Anthropic’s most heavily guarded and famously ‘locked down’ model, can be made to output persuasive misinformation on a range of topics, the example here being links between autism and vaccines, a decades-old staple of mis/dis-information that has, in fact, become a useful canary. Important: Anthropic’s paid jailbreak bounty only covers…

Don't let the LLM speak, just probe it.

Don't let the LLM speak, just probe it. TL;DR: When an LLM reads "here's some text, here's a criterion — does it satisfy it?", the answer often already exists in its hidden state before it generates a single token. So skip generation entirely: grab the hidden state at the last prompt token (~70% of the way up the model's layers), feed it to a tiny MLP, calibrate the output. Because the training…

AI Safety is theatre

AI Safety is theatre The AI Safety and Alignment communities are prolific. They are well funded and produce enormous volumes of research, evaluation methodology, governance frameworks, fellowship cohorts, conference programmes, and lobbying. What they have not produced, in any serious quantity, is deployed safety — the runtime infrastructure that sits between an AI and a user and prevents or…

I am building an AI safety company

I am building an AI safety company A couple of months ago I wrote about the Context->Interception->Thinking->Escalation approach to safer AI conversations. I have been hard at work refining this even further into an entire platform called NOPE . This has brought together my private work researching better safety pipelines and everything I've learned while at the Collective Intelligence Project ,…

CITE: Layered Defense in AI Chat Safety

CITE: Layered Defense in AI Chat Safety There have been a spate of suicides and other mental health crises recently attributed to large language models. It has confused me for some time why AI labs don't do a better job by default. I have little option but to put it down to naivety and lack of motivation. Maybe safety isn't interesting enough to intrigue the best engineers? I think researchers…

Tips for stroke-surviving software engineers

Tips for stroke-surviving software engineers This is a pretty niche topic; I don't imagine there are many of us out there. Actually, to be strict, I'd say this advice is tailored to people who've had hemorrhagic stroke in the parietal lobe with residual epilepsy... I was 29 and around 12 years into my career when it all happened , and in the six years since then I've had time to learn a bit more…

Sorry, We Deprecated Your Friend

Sorry, We Deprecated Your Friend With the release of GPT-5, OpenAI killed off a bunch of older models being used by millions of people around the globe. They did this while sneaking in a new underlying model and routing system into every existing ChatGPT conversation thread. This really takes the cake. Reckless even by normal software deprecation standards. And they only slightly rewound their…

Browser AI Agents Break Zero Trust

Browser AI Agents Break Zero Trust Anthropic has shipped a pilot of Claude for Chrome—an LLM that lives in your browser. It’s not the first and won’t be the last. I’m usually not the grinch of AI, but this one deserves pushback—especially from a lab that calls browser agents “inevitable” and reports 23.6% prompt-injection success before mitigations and 11.2% after in their own testing. Launch post…

The scribbler, the scribe, the sculptor.

The scribbler, the scribe, the sculptor. Vibe coding has become the term for the exercise of using AI to make software with little regard for the specifics, ostensibly hand-waving ideas into applications instead of taking time to design and build robust architectures in the 'old way'. It is used both derogatorily by the old and embracingly by the new. The old guard resists, writing off the vibe…

Latent Pluralism in Language Models

Latent Pluralism in Language Models In my latest work at CIP I've been thinking a lot about the western monopolization of AI and how our values have now leaked almost irreversibly into all of these models. This is an unsurprising side effect of how LLMs have been trained, as well as a general limitation of human expression available within their corpus: the internet. Below is a chaotic exploration…

Unrepresented Zeitgeists in AI

Unrepresented Zeitgeists in AI There is a hypothesis that says roughly: we believe AI is–by its nature and training–embedded with a collective truth or zeitgeist of human thought captured at a specific time. It encodes a singular snapshot of culture, likely very much amalgmated and homogenized into a singular coherent picture, that ends up flat–lukewarm–but also heavily biased to specific cultural…

Working with LLMs – not against them.

Working with LLMs – not against them. Learning to talk to an LLM is an odd sort of thing to do * . It is to induct oneself into an alien pattern of thought, where you are not asking for things like we do of humans, but instead, with every word or hint, inserting probabilities and weights into a singular brainwave. To derive true and deterministic value from this is an enchanting art for a hacker,…

LLM Security: Keep Untrusted Content in the User Role—Always

LLM Security: Keep Untrusted Content in the User Role—Always When you're working with Large Language Models that use roles like system , assistant , and user , there's one rule you need to burn into your brain: Never put untrusted content into system or assistant roles. Always keep it in the user role. "Untrusted", here, might mean: Retrieved documents (RAG) API responses Database content Web…

AIs without levers are inert

AIs without levers are inert I don't know whether this revelation is just really obvious or really dull. To me it's interesting because it flies in the face of lots of rheroric that people use when talking about AI. AI, in the mainstream context, nowadays, is basically synonomous with LLMs, or multimodal agents that have language models as their core. So that's what I'm talking about here when I…

Improving LLM Alignment with Metric-Based Self-Reflection

Improving LLM Alignment with Metric-Based Self-Reflection While building tiptap.chat , I've been pretty obsessed with safety and guardrails to prevent bad outputs. Often the solution lies in preventing bad inputs before the LLM has a chance to respond. Beyond basic filtering though, which is often a bit slow and awkward, there's an approach I've used for the "main" agent's streaming responses to…

Intercepting LLM Streams for Improved Chat UX

Intercepting LLM Streams for Improved Chat UX I’ve been building LLM chat interfaces for a while now and wanted to share some weird methods I’ve been using to get a finer grained control over text streams. As each token comes down on an HTTP stream (usually from an LLM cloud provider), I intercept in Node.js, apply a bunch of transformations, and then forward it on to the client so it can render…

Tipping AI for better responses?

Tipping AI for better responses? Nearing the end of 2023 people started reporting that ChatGPT was getting "lazy". One user joked about tipping GPT if it gave better responses, and concluded that, hilariously, offering a tip does increase the length of the response . Despite narrow data and anecdotes, offering a tip to LLMs has now become a bit of a meme. I wanted to perform a slightly more robust…

Robots Talking To Machines

Robots Talking To Machines TL;DR : we're at an inflection point where we're seeing more robots and physically-embodied AIs wanting to engage with machines that were made for humans. In a famous scene of the movie Interstellar, we see a robot called TARS manually dock the Lander craft with the Endurance station by using one of its awkward metal appendages to nudge a joystick back-and-forth. This is…

Akihabara, and my reflections on the democratization of AI

Akihabara, and my reflections on the democratization of AI Weaving through alleys of Akihabara, the so-called Electric Town of Tokyo, one notices the overwhelming amount of choice. In Yodobashi Store, an eye-watering nine-storey electronic supermarket, there are multiple aisles dedicated to converters of all varieties through the ages. From SCART to VGA, HDMI to DisplayPort and beyond. It is a…

Multifaceted: the linguistic echo chambers of LLMs

Multifaceted: the linguistic echo chambers of LLMs This is a fun one. I’ve spent more time than I’d care to admit staring at LLM output. And there’s something that I’ve noticed: LLM-generated prose has a kind of… vibe. It’s difficult to describe, but in this initial era of LLMs, it tends to be fairly obvious when you’re reading an AI-generated piece of prose. One giveaway I've noticed is this…

Simple LLM/GPT trick: “seeding”

[imported from medium.com] Simple LLM/GPT trick: “seeding” How to coerce a response with less up-front prompting This is an easy prompt-engineering hack I encountered when building both pippy.app and veri.foo . It’s a very simple idea. It can work via any interface to ChatGPT and similar LLMs but is best via the API where you’re able to designate roles. The idea is to prepopulate the LLM…

PSA: Always sanitize LLM user inputs

[imported from medium.com] PSA: Always sanitize LLM user inputs Protect yourself from different types of attacks that can expose data or functionality that you’d rather keep private. LLMs, much like SQL or any other data-layer, are liable to injection attacks . This will not change. It is in their nature as probabilistic machines. A necessary defence against this is to sanitize inputs. At the very…

Using History Insertion, Policy Drift and Allusions to jailbreak DALL·E 3

[imported from medium.com] Using History Insertion, Policy Drift and Allusions to jailbreak DALL·E 3 A quick attempt to contravene content policies with ‘history insertion’ and ‘ policy drift’ to create images of political figures. All of these images were created with DALL-E 3 using the approaches outlined in this article. I have been entranced by the influx of ingenuity that occurs whenever a…

Wanted: the engineer’s entrepreneur.

[imported from medium.com] Wanted: the engineer’s entrepreneur. I am, without question, No Good At Business. And I’m sorry to say: I have very little inclination to try. I can grit and wade through the mess, to an extent, but it’s not my spark nor my talent. This is not a false humility or nessless sass. I’m seriously just the worst at business. I don’t like making money. In fact, I hate money. I…

Building Safe, Aligned & Informed AI Chatbots

[imported from medium.com] Building Safe, Aligned & Informed AI Chatbots An analysis and walkthrough of how to build a safe, capable and truly aligned AI chatbot with ChatGPT A lot of AI fanfare has recently enveloped the world. Despite this, it’s still a rather obscure domain and difficult to know how to build atop these magical large-language-models without a lot of time and engineering…

The silliness, lossiness, and bias of leetcode screening in tech

[imported from medium.com] The silliness, lossiness, and bias of leetcode screening in tech At some point in the last ten years, LeetCode-like interview/screening platforms re-invaded tech and beautifully inserted themselves as the new norm. These are the kinds of platforms I’m talking about: CodeSignal , Arc , Codility , Hackerrank , CoderByte , and CoderPad . I’m gonna use ‘LeetCode’ as a…

Using LLMs to parse and understand proposed legislation

[imported from medium.com] Using LLMs to parse and understand proposed legislation Legislation is famously challenging to read and understand. Indeed, such documents are not even intended to be read by the average person. They are primarily tools for lawyers, ministers and judges. But we still need public scrutiny of them. If these documents are inaccessible to the majority of people, then it’s…

Disability: Models, Cultures, Perceptions and the Path to Inclusivity.

[imported from medium.com] Disability: Models, Cultures, Perceptions and the Path to Inclusivity. Understand more about the models, cultures, stereotypes, and language surrounding disability, and how we can move towards a more accessible and inclusive society for all. Disability is a famously loaded term. It lies awkwardly at the intersection of biology, identity, and society. It’s hard to talk…