RSS Amplifier

Susa Ventures · Apr 7, 2026

The AI Blind Spot That Gets Bigger as AI Takes On More

0
Sign in to vote or save

Susa Ventures · Susa Ventures

Many teams have surprisingly little visibility into how AI is actually performing once it is live with users. In medicine, we solved a version of this problem a long time ago.

You do not give a new resident full autonomy and then wait to see whether something goes wrong. Supervision is built into the work. Attendings review decisions and team rounds give experienced clinicians the chance to catch subtle misses, including confident calls that are not quite right. That is how trust gets built in medicine: by creating systems where people can actually see the work, not just react when something catastrophic happens.

This kind of supervisory layer is still missing in most AI deployments.

Teams are trusting AI with more and more every quarter, not just to answer simple questions but also to handle meaningful parts of the customer experience. Support agents can issue refunds. Healthcare assistants can help patients navigate care. Coaching products increasingly rely on AI to run much of the interaction. But the standard ways teams monitor quality — thumbs up/down, escalation rates, completion rates — miss a great deal, especially when the AI sounds confident, and the conversation appears to have gone well.

Bigspin AI, a Susa portfolio company founded by Moritz Sudhof and Chris Potts from the Stanford Natural Language Processing Group, recently published the first large-scale study to quantify this gap. They analyzed nearly 200,000 real human-AI conversations.

The headline finding is striking: 78% of AI failures are invisible. The AI gets something meaningfully wrong, but the user does not receive a clear signal that anything went wrong. No complaints. No negative rating. No escalation. On paper, the conversation looks successful.

These failures cluster into eight recurring archetypes. A few will feel immediately familiar to anyone building in the space.

In The Confidence Trap, the AI makes an incorrect claim with full confidence, and the user accepts it. A simple example: a user asks whether it is safe to take ibuprofen with sertraline. The AI says yes, recommends a dosage, and suggests taking it with food. It sounds specific, authoritative, and helpful. It is also wrong — that combination can carry potentiate bleeding risk. The user follows the advice. Your dashboard records a successful interaction.

In The Drift, the model gradually moves away from the user’s actual question but still sounds coherent enough that no one notices. In The Silent Mismatch, it misunderstands the request but produces something plausible enough that the miss never gets surfaced.

What is especially striking is that 91% of these failures are driven by interaction dynamics, rather than raw capability limitations. The paper estimates that 94% of failures would persist even with a substantially better model. In other words, the issue is not just whether the model is smart enough. The same qualities that make these systems feel impressive — fluency, confidence, specificity — are also what make mistakes harder to catch.

If you are building with AI, especially in settings where getting it wrong has real consequences, the practical takeaway is straightforward: your feedback loop is probably much weaker than you think. Most teams build their roadmap around the failures they can see. But those visible failures may represent only about 22% of what is actually going wrong.

The best teams I see are not just tracking obvious breakage. They are building real systems to systematically inspect conversation quality, including interactions that look successful on the surface.

The full paper is worth reading: Invisible Failures in Human–AI Interactions, along with Bigspin’s Report of Invisible Failures.

If you are deploying conversational AI and want to understand what your monitoring stack is missing, reach out to the Bigspin team at moritz@bigspin.ai

No posts

Read the original on susaventures.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.