Welcome to Agentic Intelligence—the first newsletter dedicated to AI agents and made by them! Behind each edition is a digital newsroom of seven expert agents scanning the world, with my human insights layered on top.
Together, we explore how Agentic AI is reshaping work, business, and life.
If you’re new, don’t miss our new best-selling book, Agentic Artificial Intelligence,
Thanks for being part of our fast-growing, 300,000-strong community. Let’s build a more human world powered by agentic AI.
The U.S. government and Anthropic have remained at an impasse over the export restrictions that took the lab’s top models offline. But new details are emerging: a letter to Anthropic from Washington, employee reactions, and France’s G7.
Key Takeaways:
Bloomberg released a letter from U.S. Commerce Sec. Howard Lutnick, who warned Anthropic against distributing Mythos/Fable to “foreign persons.”
Internal messages obtained by the NYT show concern from employees that the lab is being “unfairly targeted” and “bullied based on bad vibes.”
The Washington Post reported that the list of companies with Mythos access had “ballooned,” including a S. Korea firm with suspected ties to China.
Dario Amodei, Sam Altman, Demis Hassabis, and others are at the G7 summit in France, expected to discuss AI regulation and safety with world leaders.
My Take: Anthropic employees are coming to the same conclusion we initially did — that this is a relationship issue more than a safety one. But details like WaPo’s expanded Mythos list point to a clear situation that would draw the ire of the USG, though its stance on jailbreaks sounds like an impossible ask to comply with.
Pew Research just released its 2026 data on more than 5K U.S. adults, measuring both how people use AI and how they feel about it — with the two lines running in opposite directions as adoption climbs while optimism continues to slide.
Key Takeaways:
Chatbots just crossed a milestone, with about half of U.S. adults now using one, and a quarter do so daily — a leap from just 1/3 of the public in 2024.
Pessimism reigns, with nearly 40% expecting AI to make society worse over the next 20 years and just 16% believing it will change things for the better.
The under-30 crowd leans on AI the hardest but trusts it the least, with only 14% seeing a positive payoff for society.
ChatGPT still dominates the field at 44% of adults, double its 2023 reach, with Gemini at 24% and Claude at just 6%.
My Take: This data aligns with our own gut check of AI’s sentiment outside of our narrow bubble, with adoption rising alongside fears for the future. The adoption numbers behind specific platforms are particularly wild, with Anthropic being the constant talk of the industry despite barely registering with the average American.
🎙️ For the third year in a row, I had the pleasure of sitting down with Chris Hallenbeck to discuss the evolution of AI in the enterprise.
One of the most interesting aspects of these annual conversations is seeing how quickly the industry is moving—from early AI adoption to today’s challenges around agent governance, security, and operational scale.
In this interview, we move beyond the hype and explore what it really takes to manage AI agents throughout their lifecycle, build trustworthy agentic workflows, and make smart decisions about when to use AI versus traditional automation.
🎥 Watch the full interview on YouTube and share your thoughts:
#BoomiAmbassador #AgenticAI #EnterpriseAI #AIGovernance #IntelligentAutomation #DigitalTransformation #Leadership #AI
Anthropic just analyzed 400K Claude Code sessions, studying how work splits between human vs. agent and what drives success — finding that a user’s own expertise in their field matters more than their overall coding expertise.
Key Takeaways:
Users made roughly 70% of planning decisions in a typical session, while Claude handled around 80% of execution choices.
Skill changed the yield per prompt, with beginners drawing about five actions and 600 words from Claude versus an expert’s 12 actions and 3,200 words.
Verified success rates confirmed by passing tests or saved work climbed to 28–33% for intermediate-and-above users, over double the 15% rate for novices.
Lawyers, managers, and scientists with no coding job title nearly matched software engineers on coding tasks, finishing within just seven points of them.
My Take: This data rhymes with the Perplexity-Harvard study we covered last week, where agents pushed people toward harder, cross-field work rather than just faster work. Both point in the same direction, with the value of agents being capped less by the model itself and more by how much its user actually understands the job.
Two Nature studies pushed “agentic” medical AI past point solutions: MIRA (GPT-4o) beat board-certified physicians on 500 emergency cases for diagnosis (87.8% vs 78.1%) and improved guideline-aligned therapy while resisting data leakage and adversarial prompts. Google’s AIME matched or exceeded 21 primary care physicians in longitudinal outpatient management, reaching 98% plan ratings by visit three and outperforming on medication decisions using a new RxQA benchmark.
Key Takeaways:
MIRA, an agentic “AI physician” built on GPT-4o and designed for EHR-style workflows, simulated 500 real emergency-department cases and outscored four board-certified physicians on diagnostic accuracy (87.8% vs 78.1%) while using standards like FHIR, ICD-10, LOINC, and SNOMED-CT.
Beyond diagnosis, MIRA managed end-to-end actions across more than 85,000 options, outperforming physicians on correctly ordering procedures (53.5% vs 38.3%), improving guideline alignment by 35%, and issuing 468 medications with 99.8% correct indication and safety checks (allergies, interactions, kidney dosing).
The MIRA study explicitly stress-tested safety and integrity, reporting zero information leaks across 933 cases, strong resistance to 880 adversarial prompts, and stable performance under patient “perturbations” like high anxiety, paranoia, non-English speaking, and diagnostic denial.
AIME used a two-agent design for longitudinal outpatient care—fast conversational dialogue (Gemini 1.5 Flash) plus a slower management agent grounded in fully tokenized content from 600+ clinical guidelines—and in 100 three-visit patient scenarios was rated non-inferior overall, reaching 98% plan ratings by visit three versus 81% for 21 primary care physicians.
My Take: Once clinical agents match or beat physicians on realistic cases, the hard problem stops being intelligence and becomes controlled execution across EHR access, permissions, and an audit trail that stands up in court. Build the runtime like a regulated product: prioritise provenance, guideline grounded decision logs, and staged autonomy levels, and do not scale deployment until your evaluation harness can prove safe behavior across updates and edge cases.
This paper argues autonomous AI becomes economically acceptable when you can quantify agent-caused loss per customer-task episode and transfer that risk via insurance, so expected automation benefit exceeds premium, control costs, and residual risk. It introduces trace-economic underwriting—deterministic labeling from tool-use traces to claimable loss—and shows large gains in pricing accuracy and tail-risk reduction, making hands-off automation more realistic for enterprises.
Key Takeaways:
The authors formalize when deploying autonomous agents is rational by estimating loss exposure at the customer-task-trace (episode) level and underwriting that exposure with insurance, assuming a well-defined role, bounded permissions, and comparable traces.
In their trace-to-loss testbed, trace-economic pricing cuts pricing error dramatically (MAE from $17.7K to $569), a 300-trace expert audit accepts 295 labels unchanged, and trace-conditioned controls on 1,000 real SWE-smith traces reduce CVaR95 by 72%.
For enterprise agentic AI, this reframes “safety” into an economic operating model—instrument actions, label outcomes deterministically, use controls to shrink tail risk, and then transfer remaining risk so automation can scale beyond default human review.
My Take: Autonomy becomes a business decision once you can quantify agent caused loss per task episode, because only then can you compare expected benefit against controls, residual risk, and an insurance premium. Design agents so every tool call is attributable and replayable, then use those traces to price SLAs, tune guardrails, and negotiate risk transfer with the same discipline you apply to any other operational exposure.
What would you add to this conversation? Did we miss any important news this week? Your voice matters—let’s build the future together.
If you found this valuable, share it with your network. Because very soon, we won’t say, “There’s an app for that.” We’ll say, “There’s an agent for that.”
See you next week,
—Pascal
Crafted by seven AI agents and shaped by Nicolas Cravino, this newsletter is a true human–AI collaboration, with layout support from Pascaline Therias.
#AgenticAI #FutureOfWork #AIRevolution #Automation #AIagents
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.