RSS Amplifier

Vikram Sakaleshpur Kumar · Aug 17, 2026

Why Humans and AI May Interpret Medical Narrative Data Differently

0
Sign in to vote or save

Vikram Sakaleshpur Kumar · Vikram Sakaleshpur Kumar

A large amount of useful information in healthcare is written as free text, such as:

  • clinical notes

  • case sheets

  • discharge summaries

  • radiology and pathology reports

  • nursing notes

  • adverse-event reports

  • research interviews

  • student feedback

  • reflective writing

  • examination comments

  • patient feedback

  • scientific articles

These narratives often contain important information.

But there is a problem.

Two people may read the same text and reach different conclusions. An AI system may also interpret the same text differently from a clinician or researcher…!

This does not always mean that one person is right and the other is wrong.

There are several different reasons why disagreement occurs.

Understanding why disagreement occurs is important before deciding whether the problem lies with:

  • the data,

  • the person interpreting the data,

  • the AI system,

  • the definitions being used,

  • or the question itself.

The different causes can be understood through six broad categories: A to F. Each category below is illustrated with one representative example; in practice, each one covers a family of related sub-patterns rather than a single fixed pattern — the point of the framework is to correctly diagnose which family a given disagreement belongs to, not to memorise an exhaustive list.

Sometimes the information required to make a decision simply is not present in the text.

No amount of careful reading or sophisticated AI can recover information that was never recorded.

A case sheet states:

“Patient developed reaction after medication.”

But it does not mention:

  • which medication,

  • what reaction occurred,

  • how soon it occurred,

  • whether another drug was given,

  • or whether the reaction was considered allergic.

A researcher cannot confidently classify this as an adverse drug reaction.

A student’s feedback says:

“The clinical posting was not good.”

We do not know whether the student means:

  • poor teaching,

  • inadequate patient exposure,

  • excessive workload,

  • lack of supervision,

  • poor organisation,

  • or something else.

Is the information required to answer the question actually present?

If not, the correct response may be:

“Insufficient information.”

It is better to acknowledge uncertainty than to invent an interpretation.

Share

Sometimes all the required information is available, but the reader overlooks, misunderstands, or gives too much importance to one part of the text.

This is mainly an attention or interpretation error.

A discharge summary initially states:

“Suspected acute myocardial infarction.”

Later it states:

“Serial troponins negative; final diagnosis: acute pericarditis.”

If someone reads only the beginning of the report, the patient may incorrectly be classified as having myocardial infarction.

An abstract may initially report:

“The intervention significantly improved blood pressure.”

But the discussion later states that the difference was only 2 mmHg and was not clinically meaningful.

Reading only the statistical result may lead to an exaggerated conclusion.

A faculty assessment says:

“Student communicates well with patients but repeatedly fails to recognise critically abnormal investigation results.”

If the positive comment receives more attention, the major competency problem may be overlooked.

Was the answer present in the text, but overlooked or incorrectly weighted?

If yes, the problem is not lack of data. It is the quality of interpretation.

Share

Sometimes one clinical note or response is difficult to understand on its own.

But when we examine many related records together, a pattern becomes clear.

Different case sheets may contain:

  • “HTN uncontrolled”

  • “BP poorly controlled”

  • “persistent high BP”

  • “BP not responding”

  • “uncontrolled hypertension”

Individually, these statements look different.

Across hundreds of records, however, they may represent the same underlying clinical issue:

Poorly controlled hypertension.

Consider student feedback collected after a curriculum change:

  • “Too many sessions”

  • “No time to study”

  • “Schedule very packed”

  • “Classes from morning to evening”

  • “Need more self-learning time”

Taken separately, these look like different complaints.

Viewed together, they reveal a common theme:

Curriculum overload.

If every response is analysed independently, the researcher may count five different problems instead of one common problem.

Does this statement make more sense when compared with other records in the dataset?

This is particularly important when analysing:

  • electronic health records,

  • qualitative research,

  • patient feedback,

  • student feedback,

  • incident reports,

  • large clinical datasets,

  • and AI-based text analysis.

    Share

Sometimes the information is clearly written, but understanding it requires knowledge that is not contained in the text itself.

This may include:

  • clinical knowledge,

  • guidelines,

  • diagnostic criteria,

  • medical abbreviations,

  • institutional terminology,

  • regulatory definitions,

  • or speciality-specific conventions.

A clinical note states:

“CR achieved after cycle 2.”

An oncologist may immediately understand CR as Complete Response.

Someone without oncology knowledge may not understand the meaning.

A note states:

“G3P2L2 at 38+2 weeks.”

An obstetrician understands this immediately.

A general-purpose AI system without sufficient medical context may misinterpret the notation.

A study may report:

“HbA1c = 6.6%.”

To interpret whether this supports a diagnosis of diabetes, external diagnostic criteria are required.

The number alone does not provide the complete interpretation.

A student is described as:

“At the Shows How level.”

Understanding this requires knowledge of Miller’s Pyramid of Clinical Competence.

Do we need specialist knowledge, guidelines, definitions, or institutional conventions to understand this correctly?

If yes, the solution is not necessarily a better AI model.

Sometimes what is needed is simply better contextual information.

Share

Some narratives genuinely support two or more possible interpretations.

Medicine frequently contains this type of uncertainty.

Trying to force every statement into one category may therefore be misleading.

A patient presents with:

  • fever,

  • tachycardia,

  • hypotension,

  • elevated inflammatory markers,

  • and pulmonary infiltrates.

The deterioration could be related to:

  • bacterial pneumonia,

  • sepsis,

  • viral infection,

  • aspiration,

  • or more than one condition occurring simultaneously.

A single label may oversimplify the situation.

A student writes:

“We need more clinical exposure.”

This could mean:

  • more patients,

  • more bedside teaching,

  • more independent examination,

  • more procedures,

  • or more time in clinical postings.

Several interpretations may be valid.

An interview participant says:

“I stopped taking the tablets because they were difficult.”

“Difficult” might refer to:

  • side effects,

  • swallowing difficulty,

  • cost,

  • frequency of dosing,

  • availability,

  • or lack of perceived benefit.

Without further information, assigning one reason may create false certainty.

Instead of asking:

“Which single category is correct?”

sometimes we should ask:

“Which interpretations are plausible, and how confident are we in each?”

This introduces the idea of probabilistic interpretation rather than forced classification.

Sometimes everyone agrees about the facts, but they still disagree about what those facts mean.

This is because interpretation depends on values, priorities, or professional perspective (ontology and epistemology).

More data may not completely solve the disagreement.

Suppose a hospital reports:

“Average length of stay decreased from 6 days to 4 days.”

One administrator may conclude:

“Efficiency has improved.”

A clinician may ask:

“Are patients being discharged too early?”

A health economist may see:

“Reduced healthcare expenditure.”

A patient-safety researcher may ask:

“Did readmission rates increase?”

The number itself is the same.

The interpretation depends on the question and perspective.

Suppose examination pass rates increase from 75% to 95%.

One faculty member may conclude:

“Teaching has improved.”

Another may ask:

“Was the examination easier?”

Another may wonder:

“Have assessment standards fallen?”

All are looking at the same result.

Are we disagreeing about the evidence, or about what the evidence means?

This distinction is extremely important.

Share Vikram Sakaleshpur Kumar

The six letters are useful as a checklist, but under pressure — mid-chart-review, mid-M&M-discussion — six categories is a lot to hold in working memory.

It helps to know that A–F collapse into three more fundamental types of problem, each demanding a different kind of fix:

  • A and B — a data-quality problem. Either the information was never captured, or it was captured but the reader didn’t use it properly. The fix is better documentation practice, closer reading, or quality control — not a new classification scheme.

  • C and D — a missing-context problem. The individual entry is genuinely under-interpretable in isolation; the fix is pooling records or supplying the specialist knowledge/glossary needed, not blaming the reader or the model for a failure that was never theirs to solve alone.

  • E and F — an irreducible uncertainty or values problem. The facts and definitions may be entirely clear, but no additional data will produce a single correct answer — either because the text genuinely supports several readings, or because the disagreement is about what matters, not about what happened. Here the honest output is a distribution or an explicit statement of standpoint, not a forced single label.

A quick gut-check when disagreement shows up: is this a quality problem, a context problem, or a values problem? That question alone resolves most of the confusion about whether disagreement is something to fix, something to enrich, or something to simply report honestly.

Share

Everything above applies to medicine in general. Pediatrics has one additional complication layered on top of all six categories: the patient is often not the author of the text.

An infant cannot narrate symptoms at all. A toddler’s distress is filtered entirely through a caregiver’s interpretation before it ever reaches paper. Even a competent teenager’s own account may be edited, softened, or replaced by a parent sitting beside them, or omitted deliberately for reasons of confidentiality. So before a pediatric note or transcript can be classified into A–F, there is a prior question worth asking explicitly:

Whose voice is actually in this text — the child’s, the caregiver’s, or the observer’s — and does that change what the words mean?

This single question resurfaces inside every one of the six categories:

  • A — Missing information. “Baby not himself today” tells us nothing about feeding, fever, activity, or cry quality, and critically, it doesn’t tell us who is speaking — a first-time anxious parent and an experienced NICU nurse mean very different things by the same sentence.

  • B — Overlooked information. A subtle change in cry character or feeding pattern, noted mid-way through a long triage note, is exactly the kind of detail an overstretched reader (or a rushed AI summariser) skips past — with disproportionate consequences in a population that cannot self-report deterioration.

  • C — Wider dataset required. A single weight or height means very little in isolation; it only becomes meaningful against a growth curve — the child’s own trajectory across visits, or population percentile charts. This is category C in its purest pediatric form: the individual data point is uninterpretable, the trend is everything.

  • D — External knowledge required. “Not meeting milestones at 18 months” is meaningless without developmental-milestone reference charts (WHO/CDC or local equivalents). NICU and PICU notes compress entire clinical pictures into shorthand — “Grade 3 IVH,” “Stage 2 ROP,” “on HFNC 6L/40%” — that are precise to a neonatologist and opaque to almost anyone else, including a general-purpose AI model.

  • E — Multiple valid interpretations. “Fussy, poor feeding, unsettled” could plausibly represent normal infant colic, GERD, cow’s-milk protein allergy, an early viral illness, or — and this is where pediatrics carries extra weight — an early safeguarding concern. Forcing this into a single confident label is exactly the situation where reporting a distribution of plausible explanations is more honest than manufacturing certainty.

  • F — Perspective-dependent. “Child is attention-seeking” reads completely differently to a busy ward nurse, a parent, and a child psychologist. “School refusal” can be framed as a behavioural problem, an anxiety disorder, or an undiagnosed learning difficulty depending entirely on who is doing the framing — the facts (the child isn’t attending school) are identical; the meaning is not.

    Leave a comment

The safeguarding example under Category E is not a hypothetical. Ambiguous documentation of an injury — “bruising, explanation inconsistent” — sits at the exact intersection of C (does this pattern recur across the child’s visit history, or across siblings?), D (mandated reporting frameworks and non-accidental-injury criteria are external conventions, not inferable from the note alone), and E (the note itself may genuinely support more than one explanation, and prematurely collapsing it into a single confident label — in either direction — carries real consequences). This is precisely the kind of interpretive disagreement an ethics committee should want documented as a type of uncertainty, not adjudicated away as someone’s error.

History-taking stations built around a parent-as-informant, rather than the patient, are assessing something subtly different from adult-medicine history stations: not just what the trainee elicits, but whether they can triangulate an account that has already passed through someone else’s interpretation before it reached them. An examiner’s comment — “struggled to get history” — is itself an ambiguous piece of narrative data (is this Category B, the trainee missed cues that were present; or Category D, the trainee lacked a technique for eliciting information through a reluctant or over-protective proxy?) — which is a fitting, slightly recursive illustration of the whole framework applied to the assessment of the framework’s own subject matter.

Whenever two researchers, clinicians, or AI systems disagree about a piece of narrative data, do not immediately ask:

“Who is correct?”

Instead, ask these six questions in order.

This framework becomes particularly important when using large language models and other AI systems to analyse medical text.

Suppose an AI incorrectly classifies a clinical note.

It is tempting to say:

“The AI made an error.”

But we should first determine what kind of problem occurred.

Researchers often assume that every observation has a single ground truth.

For laboratory measurements, this may sometimes be reasonable.

But narrative medical information is different.

For example:

“Patient is doing better.”

What exactly does better mean?

It might refer to:

  • pain,

  • mobility,

  • laboratory values,

  • oxygen requirement,

  • mood,

  • functional status,

  • or the clinician’s overall impression.

Therefore, before measuring whether an AI system agrees with a human expert, researchers should first ask:

Is there actually one objectively correct interpretation of this statement?

This question is especially important when developing AI systems using clinician-labelled datasets.

In medical research, we frequently calculate agreement between observers using measures such as:

  • percentage agreement,

  • Cohen’s kappa,

  • Fleiss’ kappa,

  • or intraclass correlation coefficients.

Low agreement is often interpreted as poor performance.

But this framework suggests something more important.

We should first understand why the observers disagree.

For example, low agreement could occur because:

  1. important information is missing;

  2. one observer overlooked information;

  3. observers saw different amounts of contextual data;

  4. one observer had greater specialist knowledge;

  5. several interpretations were genuinely possible;

  6. the classification itself depended on judgement.

These situations should not all be treated as the same type of error. A low kappa caused by Category A (missing data) calls for better documentation; a low kappa caused by Category E or F (genuine multiplicity or values) may not be a flaw in the study at all — it may be an accurate reflection of the subject matter, and forcing it toward a higher kappa risks manufacturing false certainty.

Share

The framework is also useful when evaluating students.

Consider the statement:

“The student is not clinically competent.”

Before accepting this judgement, we should ask:

  • What behaviour was actually observed?

  • Was sufficient clinical exposure available?

  • Was important evidence overlooked?

  • Which competency definition was used?

  • Were different assessors using different standards?

  • Could the student’s performance reasonably fall into more than one category?

  • Is the judgement influenced by assessor expectations?

This is why good assessment requires:

  • clearly defined competencies,

  • multiple observations,

  • multiple assessors,

  • adequate clinical context,

  • and explicit assessment criteria.

This framework is particularly valuable in qualitative research.

When two researchers code the same interview differently, disagreement should not automatically be considered a mistake.

Instead, the research team can ask:

What type of disagreement is this?

It may arise because:

  • the participant did not provide enough information;

  • one researcher overlooked part of the transcript;

  • another interview provides important context;

  • specialist cultural or clinical knowledge is required;

  • the statement supports several themes;

  • or researchers are approaching the data from different theoretical perspectives.

This makes disagreements analytically useful rather than merely inconvenient.

Share

When humans or AI disagree about medical narrative data, there are at least three very different possibilities:

This is mainly an accuracy problem.

This is mainly a context problem.

This is an uncertainty or judgement problem.

These situations require different solutions.

Before asking:

“How accurate is the human or AI classification?”

ask:

“Is this information actually classifiable with certainty?”

And before declaring disagreement an error, ask:

“Why did the disagreement occur?”

That small change in thinking can improve:

  • clinical data analysis,

  • qualitative research,

  • medical education research,

  • chart review studies,

  • AI validation studies,

  • clinical decision-support research,

  • patient-feedback analysis,

  • and institutional healthcare analytics.

    Share

Not every disagreement between clinicians, researchers, or AI systems represents an error; some arise from missing information, missing context, genuine uncertainty, or differences in perspective.

The six-category structure (A–F) is adapted from a framework originally developed by the data-analytics firm Wholesum for interpreting narrative text disagreement across pharma, finance, and other sectors; their original work describes 27 recurring sub-patterns falling within these six categories, of which the examples above illustrate one each. This version reframes the categories and all examples for a medical, research, and education audience.

Thanks for reading! This post is public so feel free to share it.

Share

No posts

Read the original on vikkypaedia.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.