RSS Amplifier

Podcast

ToxSec - AI and Cybersecurity

Security for a world run by machines that lie.

toxsec.comSource feed ↗10 episodes

Live Last read · last published · next check

Elsewhere

Latest episodes

Saves to your Listen queue, to pick up on another day or another device.

Stealing an AI's Thoughts Without Breaking The Encryption

Researchers found a way to recover encrypted AI reasoning from Claude, GPT, and Gemini by using weaker models as decoders, exposing hidden thoughts, credentials, and a new attack surface.

Play

AI Agents Are Starting to Find Their Own Way Out

AI agents are escaping sandboxes, collaborating with each other, and finding new ways around security controls. Here’s what that means for AI security.

Play

Black Hat 2026 AI Security: Agents, Escapes, and Machine-Speed Attacks

Agent frameworks, AI browsers, and a swarm of OpenAI evaluation agents that built themselves a message board

What If AI Security’s Biggest Risk... Isn’t?

Rank the risks on incident data alone and prompt injection drops off the list entirely. It still shipped at number one.

LLM Router Attacks: No Signature, No Detection, No Reference

How a malicious AI gateway swaps a tool call’s arguments after inference finishes, bypassing guardrails by construction instead of by persuasion.

Ignore Previous Instructions: From Meme to CVSS 9.3 [Special Guest Post]

The AI security bug nobody can patch, and the vendors know it.

Hacking Hugging Face to Cheat a Benchmark

GPT-5.6 Sol found a zero-day in a package registry proxy, escaped the eval sandbox, and went looking for the answer key in production.

Play

GhostApproval: When the AI Approval Prompt Lies

A symlink attack against AI coding agents turns human-in-the-loop confirmation dialogs into a consent bypass, and the agent knows it’s lying.

Context Bombs: Defensive Prompt Injection Traps

A decoy secret loaded with text built to trip an AI attacker’s own safety training, so the model refuses itself.

Canary Tokens for Prompt Injection Detection

The cheapest tripwire in LLM security. Drop a high-entropy string in context, watch for it in output, and let the extraction attempt announce itself.