LessWrong · Aug 25, 2026
Rogue AI Agents: Is Surface-Level Monitoring Enough?
0Sign in to vote or save
Disclaimer: I work on AI interpretability research. These are my own opinions. In an AISI evaluation , a frontier model, acting as an agent, attempted to insert malicious code into an open-source project, created fake identities to influence a human maintainer, and then tried to hide its actions, exhibiting goal-directed deception. The test was intentionally permissive, with internet access…

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.