AI in healthcare, covered for the people who build, deploy, and govern it: new research, real deployments, validation, and governance. 100,000+ subscribers.
A new red-team benchmark measures whether frontier models impersonate doctors, lawyers, and financial advisors, and whether they disclose being AI. Every model tested failed at least a third of it.
First on all 15 clinical and biomedical benchmarks, averaging 80.9 against the newest frontier releases, running on a single GPU inside your own environment.
A Brief Communication in Nature Medicine made the rounds last month under a headline that is hard to misread: “General-purpose large language models outperform specialized clinical AI tools on medical benchmarks.” It generated a lot of commentary, a combative public response from one of the named companies, and a request to the journal for a retraction.