McKinsey's Lilli Looks More Like an API Security Failure Than a Model Jailbreak
Why the reported Lilli incident looks like an application-security chain reaching an AI system, not a model jailbreak.
Articles on AI security, LLM red teaming, and trust & safety by Michael D'Angelo.
Why the reported Lilli incident looks like an application-security chain reaching an AI system, not a model jailbreak.
Announcing that Promptfoo has agreed to be acquired by OpenAI.
A changelog formatting change took down Claude Code. Lessons about parsing human docs as machine data.
Looking at what makes something a vulnerability versus a hardening opportunity in LLM applications.
Notes from using Claude Code in parallel git worktrees: Plan Mode, ultrathink, verification loops, and Chrome automation.
Why "AI compliance questions" appeared in security questionnaires and RFPs, and how policy becomes contract requirements.
ASR isn't portable across papers because measurement choices dominate the headline number. Includes math and a checklist for reading papers.
Day-zero red team of GPT-5.2 focusing on jailbreak resilience and harmful content.
Introduces search-rubric, an assertion where a search-enabled judge verifies time-sensitive claims during evals and CI.
Connects malware querying LLMs at runtime with "vibe hacking" case studies. Defense needs continuous testing.
RLVR gains are often "search compression" rather than new reasoning ability.
Jailbreaking targets model safety training; prompt injection targets application trust boundaries.
Safety protects people from harmful outputs; security protects systems from adversarial manipulation.
Announcing our Series A led by Insight Partners with participation from a16z.
Open methodology and dataset (2,500 political statements) to measure political leaning in models.
Promptfoo's journey from prompt evaluation to AI red teaming, marking 100,000 users.