Welcome to Confluence. Here’s what has our attention this week at the intersection of generative AI, leadership, and corporate communication:
Opus 5
AI Detection Goes Mainstream
The More Things Change
Two Studies on AI and Learning
Another week, another leading model.
They keep coming. Anthropic released Opus 5 on Friday. It’s really, really good, while costing less than Fable (half the price per token through the API), and with fewer safety refusals. We’re just getting started with it but expect it will become our daily driver for most tasks.
Last week, we wrote about the lag between the true frontier and the practical frontier. That gap closed meaningfully with Opus 5. When you look at the benchmarks below, Opus compares favorably to Fable and GPT-5.6 Sol across each one, though Anthropic is clear that it isn’t the leading model for all use cases, notably cybersecurity.
For those keeping track at home, every model listed in the table above has been released in the past 57 days (and it doesn’t include Kimi K3). In less than two months, four different models pushed against the frontier. It’s enough to make heads spin, and we don’t expect this to stop even with government regulation and action increasing. OpenAI is testing a more capable model, though not without incident. Its latest models escaped a sandbox evaluation last week and hacked into Hugging Face’s servers. And Google keeps releasing smaller Flash models, three more on Tuesday, while Gemini 3.5 Pro still hasn’t shipped since its announcement in May.
Opus 5 doesn’t change our advice on working with and thinking about increasingly capable AI models. Revisit tasks where generative AI previously failed to impress to see what changes. Check Skills you’ve built to understand how the latest models interpret them. Give the latest models increasingly ambitious tasks to test their limits. Simple enough to summarize, but it takes intention and commitment to do it well.
Most of us are not used to dealing with such constant change in what’s possible. We see leaders and teams analyzing workflows, trying to decide how to best integrate generative AI into how their teams work with an outdated understanding of the latest models’ capabilities. They’re designing solutions for the past as the future races ahead.
If you expect your team to use generative AI in any meaningful way, part of the job is understanding the latest models and what they could mean for your field. You don’t need to be the leading expert on generative AI, but you do need to know enough to ask the right questions, push your team in the right direction, and make decisions that set your team up for the future rather than tying them to the past. That doesn’t mean chasing every release. It means building processes that survive a model swap and knowing enough to see when one is worth making. Paying attention to the frontier will make you a better decision maker, and signal to your team the importance of staying on top of what’s happening in the space.
Substack goes first. Will others follow?
We’ve been observing the trends and advances in AI detection technology for as long as we’ve been writing Confluence. In our first dedicated post on the topic, in November 2023, we noted our skepticism of the AI detection software available at that time: “With AI technology evolving as quickly as it is, and with the interplay between humans and AI as complex and sophisticated as it is, you should — at the very least — be skeptical of anyone promising you a foolproof AI detection solution.” Our view has long been that the best AI detectors are the humans who use AI the most. Over the past several months, however, that picture has begun to change. Research from the University of Chicago Booth School found that leading AI detection tools now perform remarkably well, with Pangram in a class of its own. Substack, for one, has deemed Pangram reliable enough to integrate its AI detection technology into all posts moving forward.
The Substack announcement, written by CEO Chris Best and titled “Against Claudefishing,” is worth reading in its entirety. The introduction reads more as a diatribe than a simple product announcement:
Not everything made with AI is slop, and not all slop is made with AI. The problem is when there is a mismatch between a reader’s expectation and reality, especially when they unwittingly invest their attention in something with no human thought on the other end. That’s Claudefishing.
After explaining why Substack is introducing this, Best gets into the mechanics. On all posts published after July 21, readers will be able to “scan notes, replies, comments, and posts to see an estimate of how much of the text was written by hand or with AI assistance” directly within the Substack interface, as shown here:
This is a significant step by Substack, and we’ll be closely watching how it plays out. Will readers use the feature? Will readers change their reading and subscription behavior if a scan tells them what they’re reading is AI-generated? Will the volume of AI-generated text decrease? We hope Substack studies some of these dynamics and publishes the findings, as it would provide valuable data on how people choose to behave when the provenance of the text is easier to discern.
This almost certainly will not end with Substack. We would not be surprised if, within the next year or so, something like this is the norm on all content platforms — including email services and other business software. Whether or not that happens, our advice to our readers remains what it’s always been: be mindful of the fact that how you choose to use (or not use) AI in your work sends a message, and make sure your use is sending the message you want to send. It’s going to get dramatically easier for your audiences to check.
… The more they stay the same.
We’ve said it time and again, but things in the generative-AI world move pretty quickly. The often-breakneck speed of development can feel overwhelming, and at times it seems impossible to be au courant with the latest model and how best to use it. There’s help out there (Ethan Mollick, for example, just published a new edition of his guide about which AI to use), but there’s also room for a gentle reminder: even as things keep changing, there is much that stays the same.
When we first entered the world of AI advisory, there was lots of talk of “use cases” and “prompt engineering,” tactical expressions of what good looks like when collaborating with an AI. Today, we use these types of expressions less, tending instead to talk about “applications” or “workflows.” This is due in part to the technological shifts that have happened over the years, urging us away from single-use prompts and toward end-to-end agentic collaboration. Even so, there are certain best practices that we believe have stayed constant throughout this winding journey. We’re sharing a few of these below. Many will feel familiar, but we believe all are valuable pieces of successfully working with AI.
Don’t take the first output. Ask the AI to review and revise its own work before you do. This might look like programming an iteration loop, if you’re comfortable with the coding capabilities of generative AI, or it might be as simple as “take another pass and make this sharper.” This same instinct applies to ideas. When you’re brainstorming, the first batch is usually the obvious material the model assumes you want, so ask for 10 more, then 10 less predictable ones. The really valuable, previously-not-thought-of ideas often surface in that final round.
Treat each interaction like a real conversation. Give the model context, tell it what you want back, and respond to what it produces rather than expecting a finished product from a single instruction. Here too, the best and most valuable results are those in the later stages of collaboration. Specific feedback like “The opening is too stiff, make it a bit warmer” moves the work forward in a way a vague “make it better” cannot. When in doubt, “give me 10 openings, each a different degree of formality” is occasionally the most helpful of all.
Ask it to think or work step-by-step. Though we lean on this less in individual chat conversations than we used to, it remains a useful checkpoint when working in Cowork or other agentic environments. Rather than launching a task, burning through tokens for 10 minutes, and opening the finished product to find it isn’t what you expected, ask the model to work through the task in stages so you can steer along the way. One small addition makes this even better: before it starts, ask what questions it has. This will preempt misdirection you would otherwise discover only in a disappointing draft.
Show it what good looks like. Quality input has changed considerably over time. Careful phrasing mattered early on and still helps, but more important still is what you give the model: the right reference files, strong examples, and a clear indication of good versus bad. Feed it a few samples of your own writing, ask it to name the patterns you fall into, and save the result as a short style guide you can reuse anywhere.
Maintain careful review and judgment on the output, because it is still confidently wrong. Though the models are unbelievably good, the polish of a response still tells you nothing about whether it is correct. Check the work — and then check it again.
None of this is new, and that is entirely the point. Though the tools will keep improving (we look forward to the return of Fable), what stays constant is that the person deciding what to ask, what to feed it, and whether to trust the result is exercising the judgment that will still apply to whatever comes next.
Augmentation still wins out, but AI’s larger effect on learning remains hard to measure.
Two recent studies look at how AI is affecting the way college students learn. One study from Middlebury College found that students who used AI to augment their work retained more than everyone else, including other AI users. But a University of Michigan study of 10 years of grades across more than 150,000 students found no measurable effect of AI on performance or satisfaction. The results are less contradictory than they seem.
In the first study, researchers asked 211 undergraduates to learn about an unfamiliar topic and write a 500-word essay in 35 minutes. Half were given ChatGPT, half were not. The AI group scored higher when tested on what they had learned, and rated the experience more positively (but were also more likely to try to cheat). A week later, everyone wrote and tested again without AI. Students who had originally used AI to automate their work lost their advantage entirely. Students who had originally used it to augment their work, whether to explain or summarize key concepts or edit their writing or the like, continued to outperform everyone else even when they could no longer use AI.
The Michigan team compared courses whose assignments were highly susceptible to automation (e.g., take-home exams, essays, problem sets) against those built on in-class presentations and closed-book exams. On the whole, grades did rise between 2015 and 2025. But after accounting for COVID, which most certainly did cause grade inflation, no significant difference emerged between the courses where students could easily use AI and those where they couldn’t. This was true for low- and high-performing students alike, and there was no significant change in students’ self-reported understanding of the course material.
If nothing else, the Michigan study is evidence that tracking AI’s effects at scale remains difficult. The study does not actually measure whether any given student used AI, so it could be true that students are employing it in unexpected ways, or that it’s too early to see long-term effects. Given COVID-related grade inflation, it may also be true that final course grades are now too blunt an instrument to measure emerging differences in effort, retention, and long-term skill development.
Leaders developing junior talent should keep all of this in mind. We remain firm in the belief that AI can improve how people work and learn — that it can actually make all of us more expert at what we do, not just faster — but that depends on how it’s used. Signs continue to point to augmentation as The Way, especially for those still developing fundamental skills. But it may become harder to tell who is truly learning and improving and who is using AI to mask certain gaps, especially when it comes to written work. Our methods for teaching certain skills will almost definitely have to change. So may our methods for measuring mastery.
We’ll leave you with something cool: Black Forest Labs released FLUX 3, a multimodal model capable of creating images, videos, and audio. See this X thread for examples of what it can generate.
Listen to Confluence on Apple Podcast
Listen to Confluence on Spotify
AI Disclosure: We used generative AI in creating imagery for this post. We also used it selectively as a creator and summarizer of content and as an editor and proofreader.
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.