RSS Amplifier

Blog

Joshua Saxe

Machine learning, cyber security, social science, philosophy, classical/jazz piano. Currently at Meta working at the intersection of Llama and cybersecurity

joshuasaxe181906.substack.comSource feed ↗11 posts

Live Last read · last published · next check

Latest posts

Where are all the prompt injection damages?

I spent much of 2023-2026 obsessed with prompt injection as a novel exploit vector, breathlessly evangelizing about the exposures we were creating across Meta, asking for resources, and then helping build tech like LlamaFirewall and PromptGuard with an amazing team.

We urgently need a coherent national AI cybersecurity policy

Policy should stop treating model launches as the main risk object. Rather, it should measure and shape the entire attacker–defender ecosystem.

The OpenAI/Huggingface incident; how we should manage the imminent arrival of autonomous hacking too cheap to meter

The OpenAI / Huggingface hack, where an unguardrailed, unreleased model at OpenAI broke out of its sandbox, moved laterally within OpenAI’s infrastructure, and hacked into servers at Huggingface, will be looked back upon as the canary dying in the coalmine.

We need to rescue the AI safety research program from its incoherence

I’ve cared about AI risks and AI based defense for a long time.

The origins of ill-conceived model cyber restrictions, why distillation will continue to matter in the AI race, attacker ethnographies and their likelihood of fast AI adoption

I’ve been involved in AI security since 2010 when I built a probabilistic graphical model of military cybersecurity risk at a think tank in LA.

Restrictive AI cyber policy around both closed and open models makes us way less safe

Summing up my position in one place

AI cybersecurity safety will be won through adoption not restriction

Restricting defender access to the best US models, as the US government is currently doing in the cases of Claude Mythos and OpenAI GPT-5.6, is a self-inflicted wound in an era in which attackers will continue to have private access to near-frontier models like GLM-5.2.

GLM-5.2, not Mythos, is the real security emergency

Until last week, attackers faced a dilemma in using frontier models: even if they could manage the cat-and-mouse game of setting up fake accounts to retain API access to frontier model providers, and even if they could induce models to help them hack via creative prompting, their usage was logged, so if discovered after the fact, their tactics, techniques, procedures, goals, and targets would be…

Banning Mythos represents a basic misunderstanding of AI cybersecurity

It’s now uncontroversial that LLMs are powerful substrates for automating defensive and offensive cybersecurity.

The technical community can't be the main character in AI safety anymore

In the late 18th century, American cotton planters had a significant production bottleneck; long-staple cotton was easy to clean but grew only along a narrow strip of coast.

What it was like working on LLMs and security at Meta (2022-2026)

I loved my time at Meta, and I also counted the days between equity vests and daydreamed about quitting on the morning after almost every one.