RSS Amplifier

AI Customer Research · Jul 24, 2026

7 things worth seeing before your next round of research

0
Sign in to vote or save

Caitlin Sullivan · AI Customer Research

✌️ Hey, I’m Caitlin. I help product, design, and insights folks do better customer research with AI—without the hype.

Dive deeper: Claude Code for Customer Insights (August enrolling) | more coming soon

This will be an efficient edition - five things I think you should know about, a few that will be thought-provoking and two that are hands-on, use them today.

The commonality in this edition: I think all of them individually and together will help you think better as an insights person next week — better challenge ideas internally, set up better systems and do better work in your next project.

Skim the headlines, stop wherever it’s for you.

  • 📋 Field notes — the AI “respondent” that beat your screeners, AI interviews vs. surveys, and when synthetic users actually work

  • 🗺️ The route — the insight-to-bets workflow most teams can adopt immediately, my synthetic users readiness test, and a growing Skills Library for free

  • 📡 Weather report — NotebookLM just went agentic

Let’s dig in —

📋 FIELD NOTES

Three recent findings that change how much you can trust what you’re about to run.

A Dartmouth researcher built an AI “respondent” that passed 99.8% of attention checks across 43,000 tries, at about 5 cents each — and every AI-detector missed it (PNAS). Don’t torch your panel though. In real data across 10 platforms, under 1% of responses seemed to be AI-written (Mechanical Turk’s the outlier at up to 16%) (Prolific/Gordon). The capability is real and cheap and worth watching out for, but it’s probably not a problem for most of us yet.

The part to act on: your old screeners are not as effective. In head-to-head testing, classic attention checks flagged only about one in three AI agents (Prolific bench). Before your next field: stop treating “type ‘purple’ to continue” as proof of humanity. The real way to check if your participant is AI is to add signals AI is bad at: mouse movement, response timing, and dedicated bot-checks caught agents 94–100% where attention checks folded. And ask your panel recruiting provider what they actually verify.

〰️

The freshest evidence here is barely a month old — and it’s not a vendor’s. Researchers ran 571 people through either a standard survey or an AI-led conversational interview on the same topic, then compared the two (arXiv, June 18). The AI interviews surfaced reasoning the survey flattened: people with identical attitude scores turned out to hold completely different mental models — which only came out once the AI could ask “why?” (Unsurprising, since that’s what we do as humans).

The point isn’t “ditch your survey tool” — it’s that a well-run AI interview is becoming a real, established method, and its edge of depth is becoming more established too.

My take: Since I first started testing AI moderators for discovery, I figured they’d become one of the first tools we reach for, in place of surveys, because they scale and pull more depth out of people. What I didn’t anticipate is the rest of this study: that they’d also be relatively sharp at telling mental models apart (depending on the tool and host of other variables, of course). AI moderators vs. traditional surveys is still the comparison I’d watch — not a comparison to human-led interviews just yet.

〰️

New study from June: Researchers built digital twins of individual people from real, deep panel data and tested them across 2.1 million responses — at best, the twins matched real answers about 79% of the time (arXiv, June 3). So synthetic users can stand in for real people to an extent.

The catch is what it takes: that 79% was their ceiling, and it rests on years of genuine data about each person, not on prompting well or handing it just one little transcript. Give it thin data and accuracy falls off fast — worst on the small subgroups you most want to understand.

So the real question isn’t “can’t we just simulate this?” — it’s “do we have deep enough data to do what we want to do?

🗺️ THE ROUTE

A readout nobody acts on is the most expensive kind. The loop that fixes it: synthesize data into findings → log findings in one backlog (not five tools) → turn the critical ones into hypotheses you can test in two weekskill fast when the signal’s negative. It’s like atomic research nuggets 4.0.

This is the workflow the top 1% teams I’ve talked and worked with are building and iterating. It’s not about insights sitting on someone’s desk or waiting for the meeting where they’ll get officially presented, it’s about getting them into product work immediately, with the help of AI. The teams farthest ahead have nailed this and are speeding up experiments and builds because this is partly agentic (but still has human oversight).

Try it right now: take your last research readout and turn it into three hypotheses, each in this shape — “if we [change X], then [metric] will move, because [what the research told us].”

Level up: Turn the creation of the hypotheses (from any and all research results) and documentation of hypotheses -> test plans into a workflow with a file-based AI platform (ex: Claude Code, Cursor). Iterate until it happens well without you.

〰️

Half the questions I get about synthetic users are really the same one:

“Is my use case one where they’ll help, or one where they’ll dangerously mislead me?”

It’s honestly a tough call — and an easy one to get wrong. So I turned findings from synthetic users studies into a two-minute quiz to help you:

  • identify whether your target use case is a good one for synthetic users (or dangerous)

  • pinpoint whether the data you want to use is right for this

  • take the right first steps to prepping better for synthetic users

Take the quiz

The Synthetic Users Readiness Test asks seven questions about what you’re trying to learn, then hands you a straight verdict — a fit, proceed-with-guardrails, or a trap — plus next-step tips for your specific case. It’s built from 40+ studies’ findings, and it’s free.

Take it before your next “can’t we just simulate this?” debate. I’d love to hear what verdict you get and what you think of it — hit reply and tell me.

〰️

My free, subscriber-only Skills Library got a fresh batch of ready-to-run skills for customer-insights work — the same ones I use all the time, for things like getting from raw transcripts to something I’d actually put in front of a stakeholder.

Here’s the part I’m excited about: some of the best additions lately (and coming soon) haven’t been mine — they’ve come from course students automating their own real problems. I’m guessing their problems are yours, so I hope this growing collection keeps helping you move faster (and better) without starting from scratch.

Get the Skills!

Built a skill that saves you time on insights work? Send it my way. If it’s good, it goes in the library with your name on it and any link to your work you’d want me to use. Just reply to this email and show me what you’ve got.

📡 WEATHER REPORT

If you’ve filed NotebookLM under “chat with a PDF”, it started changing jobs in June. Every notebook now has a sandboxed computer that runs code on your sources — so “how many participants raised each theme?” returns a counted number, not an LLM guess. It’ll also go source-hunting on the open web from a blank notebook, and visualize then export results straight to slides, Excel, or CSV.

The catch: at launch these agentic features are gated to Google AI Ultra or a Workspace plan with the AI Ultra add-on — Google says cheaper tiers are coming but hasn’t said when. If that’s steep (it is), either put it on your watch list, or expense a single month of Ultra for a big analysis push and cancel. One caveat — it now pulls its own web sources, so quality checking might be back on you; open its show-your-work and trace a claim before it reaches a stakeholder. (Google · TechCrunch)

Keep moving.

— Caitlin Sullivan

No posts

Read the original on aicustomerresearch.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.