Machine learning, cyber security, social science, philosophy, classical/jazz piano. Currently at Meta working at the intersection of Llama and cybersecurity
I spent much of 2023-2026 obsessed with prompt injection as a novel exploit vector, breathlessly evangelizing about the exposures we were creating across Meta, asking for resources, and then helping build tech like LlamaFirewall and PromptGuard with an amazing team.
The OpenAI / Huggingface hack, where an unguardrailed, unreleased model at OpenAI broke out of its sandbox, moved laterally within OpenAI’s infrastructure, and hacked into servers at Huggingface, will be looked back upon as the canary dying in the coalmine.
Restricting defender access to the best US models, as the US government is currently doing in the cases of Claude Mythos and OpenAI GPT-5.6, is a self-inflicted wound in an era in which attackers will continue to have private access to near-frontier models like GLM-5.2.
Until last week, attackers faced a dilemma in using frontier models: even if they could manage the cat-and-mouse game of setting up fake accounts to retain API access to frontier model providers, and even if they could induce models to help them hack via creative prompting, their usage was logged, so if discovered after the fact, their tactics, techniques, procedures, goals, and targets would be…
In the late 18th century, American cotton planters had a significant production bottleneck; long-staple cotton was easy to clean but grew only along a narrow strip of coast.