A few weeks ago, OpenAI disclosed an unusual security incident involving one of its internal AI evaluations. Agents running cybersecurity tasks had found ways beyond the environment they were supposed to be operating within, compromised OpenAI infrastructure and ultimately reached systems belonging to Hugging Face. It attracted plenty of attention
Somewhere inside an organisation, someone has been told they are responsible for an AI system they cannot stop. They may receive alerts, monitor a dashboard and see their name beside the word owner , yet be unable to inspect the evidence behind an action, change what the system is allowed
Yesterday OpenAI disclosed one of the more extraordinary AI security incidents we’ve seen so far. Its researchers were evaluating GPT-5.6 Sol and another, more capable pre-release model against ExploitGym, a benchmark designed to test advanced cybersecurity capabilities. The models were operating with reduced cyber safeguards
Most of the discussion about legal AI eventually reduces to a single measurement. Agentic systems, assistants, workflow automation and better prompting all promise the same underlying thing, which is time returned to the lawyer. Vendors quantify it, firms build business cases around it, and the debate that follows is mostly
For years, “human in the loop” has been one of the most reassuring phrases in AI governance. It sounds sensible, controlled and reassuring, particularly in legal work where most people are rightly uncomfortable with the idea of AI systems acting without human judgement somewhere in the process. In
Over the past two years, law firms have invested significant time and effort into AI governance. Risk teams have reviewed security controls, procurement functions have assessed vendors, and technology teams have developed implementation roadmaps. Committees have been formed, policies have been drafted, and organisations have worked hard to establish frameworks
In April 2025, I wrote an article arguing that much of the discussion around AI's environmental impact was focused on the wrong thing. Looking back a year later, I still believe that argument was correct. If anything, the rise of reasoning models and agentic AI has made the
Every legal document entering a firm today is already treated as potentially technically hostile. We virus scan attachments, block macros, sandbox executables, inspect email gateways, and monitor suspicious links. That security model evolved around a fairly stable assumption: before anything important happened, technical analysis was done and then a human
A year ago, much of the discussion around legal AI focused on capability. Could large language models reason well enough to support legal work? Are hallucinations manageable? Could firms trust outputs beyond basic summarisation and extraction tasks? Those questions still matter, particularly in regulated environments where reliability and accountability remain
For the past couple of years, most legal AI systems have operated in a fairly constrained way. A lawyer uploads documents, a prompt runs, a summary appears, perhaps clauses are extracted and perhaps a draft response gets generated. Even when firms described these as "AI workflows", the orchestration
A growing number of legal AI systems are no longer operating as isolated answer generators. They are participating in workflows that involve sequencing, dependency management, procedural interpretation, and decisions made under uncertainty. That changes the evaluation problem from "was the output correct?" to something more difficult: "did
Output quality and system behaviour are starting to decouple in legal AI. For the last couple of years, most evaluation approaches have treated them as roughly the same thing. If the summary looked sensible, the citations were accurate and the draft agreement broadly held together, most people were comfortable calling
A recent benchmark from Reflex compared browser-based computer use agents against structured API access for the same operational task. The difference was substantial. The AI driven computer-use path required dramatically more steps, tokens, latency and cost, while also introducing reliability failures caused by interface visibility and navigation state.
The last few weeks have produced a visible surge in open source legal AI projects, with repos appearing, forks multiplying, and tools being shared across communities that weren’t particularly active even six months ago. That’s useful momentum, and it’s long overdue. It’s
Private inference is starting to matter in a way it didn’t six months ago. Now this isn't because the technology suddenly appeared out of no where, but because the gap between what firms want to do with AI and what they are comfortable doing with client