RSS Amplifier

AI Customer Research · Jul 10, 2026

How good is Fable really? And the intentional "poisoning" of your deep research

0
Sign in to vote or save

Caitlin Sullivan · AI Customer Research

✌️ Hey, I’m Caitlin. I help product, design, and insights folks do better customer research with AI—without the hype.

Dive deeper: Claude Code for Customer Insights (Aug-Sept cohort) | more coming soon

Fable 5 is back on non-US accounts like mine, and the internet has opinions — comments about 12-hour autonomous runs, a “relentlessly proactive” model, a real leap ahead. But almost all of those rave reviews are about coding. Nobody was answering the question I ask about every new model and it’s ability to reason:

Does it make you better at reading customer interviews? Can it identify the patterns, themes and hidden nuance like expert humans cans?

So I wanted to check. And it turned into the same lesson as the second thing that landed on my desk this month: AI hands you something that looks trustworthy — a benchmark score, a cited source — and the looking-trustworthy is doing a lot of work. Let’s dig in.

  • 📍 New from me — Future of UX podcast episode is out, and my course is filling

  • 📋 How good is Fable, really? — I went looking for the winner and found something weirder

  • 🗺️ Can you trust that model comparison? — a checklist, plus questions for the next “new model is better” debate

  • 📡 One Reddit comment poisoned an AI research agent — what to check before you trust AI market research

  • 🌰 Trail mix — Model updates, tools for in-person discovery, and a Claude Code spend tracker

Let’s dig in —

  • Patricia Reiners and I talked about the rather unsexy topic of files - but the files you put in to work with Claude Code and other systems, plus the files you get out for documentation are just about the most important part of working with AI and agents right now.

    Listen over here: Apple Podcasts, Spotify

  • The testimonials and reviews are getting better every cohort because people from teams like Mastercard, Hinge, and Spotify are turning spontaneous chats into repeatable systems within two weeks.

    Sign up here

📋 FIELD NOTES

The basic setup: I used three sets of ten real customer interviews, my own original analysis as the answer keys for each. Three models tested — Fable 5, Opus 4.8, Opus 4.6 — each completing 15 runs of the analysis. ChatGPT models graded what was delivered against my golden answer sets.

Worth knowing: This is a simple example of benchmarking - I run the same tests every time a new model comes out. But how I do it is important: I don’t use best-in-class prompting tactics or highly engineered context to help the models do a better job. I run a simple set of instructions that allows the models to reason and choose approaches on their own. That’s how we can see what each model is capable of with its defaults - and whether it’s an improvement over previous models.

Did one model find significantly more findings than the others, or do a closer analysis to mine? Not really. Roughly 71% findings match for Opus 4.6, 70% for Fable, 67% for Opus 4.8 when compared to my golden results set (e.g. my own thorough analysis).

That’s honestly too close to see a clear upgrade or winner among the three. Which also means the frontier Fable model didn’t clearly deliver better interview findings or themes. “Just use the newest model” isn’t the advice to follow every time for qualitative insights work. TBD if this looks better with quant work.

Fable walked right past the most human thing in the data.

Buried in one interview set: people were a little embarrassed to admit they meditate because of the cultural response they received to that kind of mental health approach in their work environments. So they downplayed how much they used the app. Reading between the lines and catching the “say versus do” gap is often the whole point of qual work. The older Opus 4.6 caught that gap every single run. Fable never did.

Hold these findings loosely — I ran three sets of tests, with three datasets, not hundreds of either. This should be a flag, but not necessarily a final verdict, depending on your own work and data.

Run the same model on the same interviews, with the same settings, and the scores don’t always match. Opus 4.6 landed at 61% one run and 76% the next — a 15-point swing on identical inputs. Not every model bounced that hard, but enough that a single try tells you a mood, not a measurement.

🗺️ THE ROUTE

Next time someone claims “the new model is way better”, here’s generally how to test if it’s real. It’s a blind taste test — same task, hidden answer key, a neutral grader, run enough times that one lucky run doesn’t fool you:

  • Run it more than a handful of times, and report the average and the range.

  • Use 2+ datasets (3+ is ideal) to catch where models might be particularly good or bad with certain kinds of data (ex: two very different sets of interviews)

  • Check whether the gap clears each model’s own bounce. Scores jump around run to run — Opus 4.6 swung 15 points above. If one model beats another by less than that swing, you’re looking at a lucky run, not a better model.

  • Don’t let a model grade its own family. Use a grader from a different model family, and hide which model wrote what. Ex: ChatGPT grades Claude.

  • Lock the rubric and the settings before you compare, then spot-check a few of the grader’s calls yourself.

  • Don’t treat your answer key as gospel. “Score” means agreement with one expert’s labels, not truth.

📡 WEATHER REPORT

There are many reasons to double-down on verification steps in your AI-assisted research workflows, but this adds one more.

Most researchers and PMs spot-check for hallucinations: the AI inventing something with no source. You typically catch those by asking AI and yourself “where did this come from?” and enforcing citation rules in outputs. But the “poisoning” threat slips past those spot-checks.

Poisoning: someone plants real content on a real platform — a Reddit comment, a forum post, a product review — deliberately worded to steer what the AI tells you. The source is real and looks legitimate, so “where did this come from?” comes back clean. It’s not the AI making something up; it’s the AI trusting a source that was set up to fool it.

Cornell Tech’s WARP study showed this spring how the attack lands on deep-research agents — the tools that scour the open web and hand you a synthesized report. Plant a single 13-word comment on a high-traffic page, and when the agent pulls that page into its research, the attacker’s fake product shows up in the report 38–51% of the time (on the open-source systems they tested).

The study didn’t run live attacks on consumer tools like Gemini or ChatGPT, but it did measure their potential exposure to the risk: Gemini’s Deep Research draws about 12% of its citations from user-generated content, most of it Reddit — the same surface the attack targets. There’s a name for doing this on purpose: Generative Engine Optimization (GEO).

You don’t have to stop using these tools. The rule: treat user-generated content as a lead, not a fact. Reddit, forums, and reviews are gold for finding language, complaints, and emerging problems — but two checks turn a lead into something you can act on:

  • Google shipped Gemini 3.5 Live Translate (June 9) — real-time, voice-to-voice translation in 70+ languages, now in Google Meet and Translate. If it holds up, you could run discovery with more customers more smoothly without avoiding markets where you can’t speak the languages.

  • Anthropic shipped Claude Sonnet 5 (June 30), the new default. They claim it hallucinates less and is less of a yes-man than the last Sonnet. If true, that makes it a stronger cheap starting point for a first analysis pass. But “near-Opus quality” isn’t “reads nuance like Opus” — test it on your own transcripts before you lean on it.

  • Ellis — Free AI notetaker built for face-to-face conversations: it records, tags who said what by voice - using a sample of your voice to identify when you’re speaking. It’s brand new, but a lot of my course students have wanted a better solution for in-person recording — this might be it?

  • WTClaude — free, open-source tool that shows your real Claude Code spend in the terminal, reading the actual numbers Anthropic bills you, not estimates. Sixty-second setup: npx wtclaude setup.

Keep moving.

— Caitlin Sullivan

No posts

Read the original on aicustomerresearch.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.