RSS Amplifier

Dr. Tattvam A. Nair · Nov 16, 2025

Making Clinical Judgment Computable

0
Sign in to vote or save

Dr. Tattvam A. Nair · Dr. Tattvam A. Nair

“For all of the advantages of EMRs, they can create distance between the physician and patient if care is not taken to preserve face-to-face contact. EMRs also require training and time for data entry. Many providers spend significant time entering information to generate structured data and to meet billing requirements. They may feel pressured to take short cuts, such as ‘cutting and pasting’ parts of earlier notes into the daily record, thereby increasing the risk of errors. EMRs also structure information in a manner that disrupts the traditional narrative flow across time and among providers.”

— Harrison’s Principles of Internal Medicine, 21st Edition

  1. How chart-prep becomes the memory layer of an AI-first EHR
    1.1 What “better outcomes” actually means
    1.2 What has to be true for this to work

  2. Why Precise Meaning Is Required for Safe Automation

  3. Clinical Reasoning as a Graph of Linked Facts

  4. A clinically real walkthrough (a few examples)
    4.1 Drug-induced gout
    4.2 Heart failure
    4.3 Osteomyelitis
    4.4 Surveillance
    4.5 Anticoagulation

  5. What improves immediately, what improves by month three, and what compounds over a year

  6. The Memory Layer I Am Building and SNOMED-CT

  7. How will I evaluate this?
    8.1 Coding and graph quality
    8.2 Learning the clinical reasoning behind edits
    8.3 Note and workflow outcomes
    8.4 Safety checks and decision rule

  8. Reasonable concerns

  9. Conclusion

I believe AI-first EHRs can help improve patient outcomes, but only if they capture clinical intelligence instead of just storing text. That intelligence is most visible during chart-prep, when a clinician reviews scattered information, decides what is clinically relevant right now, and sets the next steps. Most of the reasoning behind those decisions is never written out. An AI-first EHR should treat that work as its main source of learning and turn it into memory: a structured, computable record of what is true about the patient, how we know it, and how the clinic chooses to respond.

To do that, the system needs to infer the reasoning from the pattern of edits - what was added, removed, or changed, and when - and how those changes line up with clinical reasoning, management protocols and outside evidence. It should then store each decision as a standardized, source-linked, time-stamped statement that software can reliably search, combine, and act on. Edits are not just cosmetic; they show how the clinician actually thinks and decides. As that computable memory is reused and refined, variation goes down, safety improves, and it becomes easier to explain why a given action was taken. Chart-prep is not just documentation; it is the moment when judgment turns into a plan, and the EHR should learn from it.

Most of what doctors train to do is to learn and apply knowledge to a specific patient’s context. Today, most of that context lives only in the clinician’s head: a mental model built over years of practice - understanding who this patient really is, what has changed, and what matters next, combined with clinical patterns learned from treating thousands of similar cases. With complete, computable context, AI can take on a large share of that work.

When I say “better clinical outcomes,” I do not mean a vague sense of efficiency. I mean fewer missed cancers, fewer avoidable myocardial infarctions, fewer preventable readmissions, fewer bleeding events on anticoagulants, better blood pressure, A1c, and LDL control across a panel, and lower complication rates. Today, those outcomes are constrained by the difficulty of delivering truly personalized care at scale. Guidelines exist, but they are written for generic cohorts, not for this specific patient with this specific history, these comorbidities, these medications, and these risk factors. Tailoring guideline-based care to each individual - consistently, for every patient, at every visit - is what becomes hard when context lives only in the clinician’s head.

To move those numbers, we need three things at the point of care:

  • Better longitudinal care. The entire patient context is computable, so it can be converted into specific clinical recommendations.

  • More personalized care. Guidelines applied to this specific patient’s situation, not to a generic cohort - tailored to the individual.

  • More proactive care. Risks identified early for the individual by learning patterns across patients at scale.

As that happens, the physician’s role shifts toward what must remain human. The doctor supervises the system, steps in on unclear or unusual cases, performs procedures, talks patients through trade-offs, and takes responsibility for the decision - essentially, the parts regulators and common sense agree should not be handed over to automation. If this works, one clinician will be able to manage a much larger panel safely, and the effective doctor–patient ratio that a health system can support will shift.

For this belief to hold, a few things have to line up.

On the AI track:

  1. The models themselves have to be excellent: accurate, robust, and able to represent and communicate uncertainty instead of hiding it.

  2. The applications built on top of those models - whether for diagnosis, treatment planning, or clinical decision support - have to be designed carefully and pass real clinical validation.

  3. Those applications must receive complete, and more importantly computable, clinical context at the exact moment of care. A model that cannot see the complete context of the patient cannot make safe, personalized suggestions, no matter how good it is in isolation.

On the day-to-day practice track:

4. The low-value administrative work that currently consumes clinician time has to be taken off their plate.

The memory layer I am describing is aimed squarely at point (3). It is the infrastructure that turns chart-prep into computable context so that, when the models and applications do mature, they are acting on a faithful picture of the patient rather than on guesses from text. The chart-prep tool we’re building also addresses point (4).

Memory helps only if the system knows exactly what it is remembering. Consider this, a diagnosis alone is too vague to generate safe clinical decision support. The system needs precise coordinates: what the problem is, where it is, which side, how severe, what caused it, when it started or resolved, and whether it is current or historical etc. It also needs the full context of the patient: age, comorbidities, allergies, medications, devices, any other relevant history etc.

With those coordinates and that context, the system can generate clinical decision support that is individualized to this specific patient and every recommendation can be traced back to the exact facts, sources, and reasoning that produced it.

Language models can read text and extract information from it. But even if the text explicitly says “open fracture of the distal radius, right side, on tramadol for pain, improving” the model is still operating on text, not on structured facts. It might hallucinate “right” when the text said “left.” It might confuse a historical fracture with a current one. It might miss a contraindication documented elsewhere in the chart. And even if it tries to link what it extracted back to a source, that linkage is not guaranteed to be correct.

Structured facts are different. The same clinical statement, stored in the backend, looks like this:

  • Base:

12676007 | Fracture of radius (disorder)

  • Attributes:

363698007 | Finding site = 75129005 | Bone structure of distal radius
272741003 | Laterality = 264180000 | Right sided
116676008 | Associated morphology = 59091005 | Open wound
22253000 | Pain (finding) = 182991002 | Treatment given = 386858008 | Tramadol = 385633008 | Improving

Each element - the base concept, each attribute, each value - has a unique SNOMED CT ID (we discuss SNOMED more later). The structure is explicit, verified, and traceable to a specific source and date. When recommendations are generated from this, you can audit exactly which facts drove the decision. When they are generated from text inference, you cannot trust that the inference is correct or that the source attribution is accurate. That is the difference between safe and unsafe.

Clinical care is not a list of isolated, well-labeled items. It is a graph of related facts that evolve over time. A diagnosis (finding) has attributes. A procedure has its own attributes. A therapy has dose and class and can be the cause of an adverse event. A lab or measurement carries a value, a unit, a reference range, and a recency window. A device persists after a procedure and constrains imaging safety and infection risk. Environmental and social contexts influence exposure and adherence. Time binds these pieces: what is current versus historical, what preceded what, and which intervals the clinic accepts as sufficient evidence.

Clinical meaning emerges from how these bases connect, not from any single node. These links are not ad-hoc English phrases; they are relationships the computer can rely on and clinicians can audit.
A few examples:

  • indication-of / indicated-by: a finding justifies a procedure.

  • performed-on / procedure site: a procedure targets a specific structure with laterality.

  • uses-device / device-in-situ: a procedure implants a device that then persists as a constraint.

  • supported-by: a diagnosis is grounded in labs or imaging with dates.

  • contraindicated-by / qualified-by: a therapy option is constrained by allergy, genotype, renal function, pregnancy, device presence, or specific thresholds.

  • due-after: surveillance is anchored to a prior event and interval.

  • resolved-on / worsened-since / progressed-to: state transitions tied to time and evidence.

  • same-anatomic-context-as: orders and imaging inherit the correct body site and side.

  • occurs-with-environment: recurrence or risk is linked to an exposure (e.g., persistent wet floors at work).

When the patient record is structured this way, clinical logic becomes queryable. A contraindication is a relationship in the graph between a therapy and the facts that make it unsafe for this patient. An eligibility criterion is a pattern of findings, labs, and history that either exists in the graph or does not. The system does not search text for keywords. It checks whether the required relationships exist.

Because every node has a source and a date, and every relationship is explicit, the system can explain itself. When it proposes or suppresses an action, it points to the specific facts and relationships that drove that decision.

A few examples.

A clinician documents gout in the right knee after ultrasound confirms crystals. The finding carries site = knee, side = right, morphology = urate crystal deposition, and course = acute. The medication list includes a thiazide diuretic; the gout is marked causative agent = thiazide. Labs show eGFR 38, which records chronic kidney disease and sets renal limits. History shows a prior allopurinol intolerance (rash) with a date. Genotype returns HLA-B*58:01 positive. Coronary disease is present on the problem list.

Those nodes and edges together tell a precise story: drug-induced gout in a patient with CKD, genotype risk for severe allopurinol reactions, and cardiovascular disease that complicates febuxostat. Without anyone writing an essay, the order composer infers that acute treatment should be renal-safe, that hypertension therapy should move off thiazide where possible, and that long-term urate-lowering should avoid allopurinol and treat febuxostat cautiously. The composer can also explain itself by pointing to the exact nodes and edges: the intolerance entry with date, the genotype report, the eGFR value and date, and the medication class.

The record documents heart failure with reduced ejection fraction, supported by an EF of 35% with a date. Serum potassium is 4.8 mmol/L with a date, and eGFR is 42 mL/min/1.73 m² with a date. An ACE inhibitor is on the active medication list. The clinic’s policy allows initiation of an ARNI for HFrEF when labs are within thresholds, the results are recent enough to be trusted, and a 36-hour ACE-to-ARNI washout is scheduled. That policy is stored as a computable rule attached to the heart-failure concept.

The system treats the EF of 35% as evidence for HFrEF and therefore as a valid indication to consider an ARNI. It then applies the clinic’s qualification checks by verifying that potassium and eGFR meet the clinic’s limits and that both results fall within their recency windows. It also enforces the temporal constraint that an ARNI start must occur at least 36 hours after the last ACE-inhibitor dose. The order composer shows the ARNI option only when all of those conditions are satisfied; if any requirement is missing, out of range, or stale, it withholds the suggestion and asks for the minimal next step, such as a fresh potassium or a scheduled washout time. When the conditions are met, it proposes the ARNI and drafts the plan in the clinic’s voice - dose, hold criteria, lab recheck timing, and a 72-hour nurse call - because prior edits trained that style on this decision. The system can show its work by pointing to the EF value and date, the potassium value and date, the eGFR value and date, and the ACE-inhibitor entry with its stop time.

A bone biopsy grows MRSA resistant to methicillin and TMP-SMX, and the report is stored with its date. During vancomycin therapy, creatinine rises; the record adds vancomycin-associated kidney injury with severity and course on the treatment timeline. A vascular stent was placed in the same leg last month and is recorded as a device-in-situ with side, implant date, and an MRI restriction window.

From these facts, the system treats the biopsy as evidence for MRSA osteomyelitis and constrains antibiotic choice by the resistance profile. It avoids vancomycin because of the documented kidney injury and proposes daptomycin with dosing adjusted for renal function. MRI safety checks trigger only when imaging is ordered for that same limb and only while the MRI restriction window is active, using the device-in-situ and shared-anatomic-context links. If cellulitis recurs in the same limb within weeks of the stent placement, the system classifies the episode as device-adjacent, explains the elevated risk, and adds device-care prevention steps to the plan per clinic policy. It can show its work by pointing to the biopsy organism and resistances, the adverse-event entry with dates and creatinine trend, and the stent implant note with side and date.

The record contains a procedure report that documents removal of a tubular adenoma with a procedure date. A separate report states “repeat colonoscopy in three years,” and it is stored with its own date and source. The memory links the procedure event to the recommended surveillance interval.

From these facts, the system computes a due date by applying the interval to the procedure date and records it with provenance to both documents. The chart then shows “Colonoscopy surveillance due March 2029” and can reveal the exact reports and dates that produced that due date. When the front desk asks, “Who is due next quarter?”, a cohort query returns the names whose due dates fall in that window, each entry carrying links to the two source documents. No spreadsheets are created, and no one re-reads PDFs. If a later report changes the histology or the interval, the system recomputes the due date and logs the reason for the change.

The clinic requires a recent INR before continuing warfarin. The memory stores the last INR value with its date and the clinic’s accepted recency window as attributes of the anticoagulation plan. The order composer checks those attributes every time continuation is considered.

If the INR is outside the recency window, the composer withholds the continuation order and asks for a repeat test, showing the last value and its date. If the INR is within the window, it stays silent and allows the plan to proceed. The system can show its work in either case by pointing to the stored value, the date, and the policy window that governed the decision. This is targeted safety rather than alert fatigue.

Across these patterns, two principles hold. The product acts on structured meaning with time and sources, not on guesses from text. It also reasons over links among findings, procedures, therapies, labs, devices, environments, and time because clinical reasoning lives in those links.

  • Day one. The system builds a structured graph of the patient. Every clinical fact - problems, labs, procedures, devices, medications etc - is stored with its attributes, its sources, and its timestamps. The graph is precise and traceable. When the system makes a suggestion, it can point to exactly which nodes and edges in that graph justify the action.

  • By month three. The system starts learning from your corrections. Every edit you make teaches it something: how you connect findings to decisions, which thresholds you trust, which patterns you treat as clinically significant. The system begins to anticipate your reasoning. Suggestions get closer to what you would have chosen. Routine work requires less cleanup.

  • After a year. The system has internalized your clinic’s clinical logic. It knows how you think about eligibility, contraindications, and risk. It can answer questions that depend on clinical meaning, not just codes. New clinicians inherit this accumulated reasoning immediately. You can run cohort queries based on clinical meaning. Large scale analysis becomes possible and research questions that once required manual chart review can be answered in a click. lThe knowledge that used to exist only in experienced clinicians’ heads now lives in the system, queryable and explainable.

The short answer: a memory layer that sits beside the EHR, turns unstructured input into computable facts, learns the clinical reasoning behind edits, and writes safe, explainable structure back into the chart.

Let me give you a primer on the foundation.

The system is built on SNOMED CT, a global clinical terminology with ~360,000 concepts. Each concept has a unique code and a precise definition. The terminology is polyhierarchical: a single concept can belong to multiple parent categories, so you can traverse the graph from different angles. The codes are granular: you can attach attributes to them and build relationships between them. Because concepts, attributes, and relationships are all structured and explicit, you can reason over the entire graph. You can query by meaning rather than keywords. You can detect when different expressions represent the same clinical idea. You can trace how a change in one part of the graph propagates to decisions elsewhere. That is what makes clinical reasoning computable.

The terminology also enforces rules about what relationships are clinically and legally valid. It also maps to ICD-10, LOINC, RxNorm, and CPT, so the structured facts you build internally translate to whatever codes external systems require.

Here is what each part of the memory layer does:

  • Ingestion and normalization. The system pulls information from the EHR and from outside documents such as faxes, PDFs, and referral letters. It converts unstructured text into structured facts (concepts with attributes, sources, and timestamps).

  • Concept store with canonical forms. Facts are stored in canonical form. Different phrasings of the same clinical idea resolve to the same underlying concept.

  • Legality checks. Every fact is validated clinically and legally, before it enters the chart. Invalid combinations are repaired or rejected.

  • Learning reasoning from edits. When a clinician edits a draft, those edits happen at the level of structure - concepts, attributes, and relationships. The system records what was added, removed, or changed. It records what facts were present when the decision was made. It identifies patterns ie if these facts appear together, this action follows. Over time, it builds a model of your clinic’s reasoning.

  • Query by meaning. Patients are retrieved based on clinical definitions, not billing codes. You define the criteria that matter clinically, no matter how complex. The system returns patients who match that definition, regardless of how the original text was worded.

  • Clinic style and policy as data. Your clinic’s phrasing, thresholds, and protocols are stored as structured data tied to concepts. When the system generates a draft, it uses the patterns your clinic has established, not generic templates.

  • Write-back adapters. The system writes structured facts back to the EHR into the correct fields via FHIR.

  • Evidence packet. An evidence packet is assembled for each visit, all linked to source documents. The packet serves as the basis for audits and quality review. No one has to compile it manually.

  • Interoperability. The system maps to standard terminologies ( ICD-10, LOINC, RxNorm, CPT, etc ) for billing, labs, medications, procedures, etc.

The goal is to know whether the memory layer improves clinical work beyond what a strong language model already does.

To test that, three elements are held constant: the language model, the de-identified cases, and the clinicians who review and edit. The only difference between study arms is how the system represents information and how it learns from edits.

Each encounter is assigned at random to one of two modes. A clinician does not see the same case in both modes, to avoid carryover effects.

There are two modes:

  1. Memory + SNOMED mode. The system converts text into SNOMED CT concepts and attributes, validates each combination against the MRCM, classifies and canonicalizes expressions, and generates the note from this structured graph. When a clinician edits the draft, those edits update the underlying concepts, attributes, links, and rules. In this arm, the memory layer changes over time.

  2. Text-only mode. The same model generates a readable draft, but the system does not build SNOMED structure. It stores no concepts, attributes, or links, and edits change only the surface wording. There is no durable memory in this arm.

The evaluation covers three layers:

  1. Coding and graph quality

  2. Learning of reasoning from edits

  3. Note and workflow outcomes

Explicit safety checks are applied throughout.

First, the system must be building the right structure.

For a subset of cases, clinicians and a terminologist construct a gold-standard graph. This graph contains the agreed SNOMED concepts, attributes, and key relationships for those encounters. Disagreements are adjudicated, and inter-rater agreement is recorded, so the reference standard is explicit.

Against this graph, the following is measured:

  • Concept coding. Whether everything is mapped to the correct SNOMED concepts. Precision, recall, and F1 for concept assignment are reported, along with exact-match rates for primary diagnoses and primary procedures.

  • Attributes and expressions. Whether attributes are assigned correctly. Whether post-coordinated expressions are logically equivalent to the gold expressions after canonicalization. Attribute-level precision and recall are reported, along with expression-level exact match after normalization.

  • Relationships (graph structure). Whether links between nodes are correct. Examples: finding → indicated procedure; diagnosis → supported-by labs or imaging; therapy → contraindicated-by allergy, genotype, or renal function; procedure → due-after interval → surveillance; and procedure → device-in-situ → imaging constraint. Precision, recall, and F1 for these links are reported.

For each proposed concept, attribute, and link, the model’s confidence score is recorded and compared to correctness to assess calibration. The goal is a system that abstains or asks for clarification when uncertain instead of acting with unjustified confidence.

In parallel, an error taxonomy for structure is maintained. Errors are classified as wrong concept, missing concept, wrong attribute, missing attribute, wrong link, or missing link. This makes clear where the representation fails and where the design needs to change.

For every structured statement, how often the model proposes illegal combinations and how often the MRCM checks catch and correct them is recorded.

Second, the system must be learning reasoning patterns from edits rather than just copying clinician wording.

For a set of common, high-stakes decisions (for example: ARNI initiation in HFrEF, warfarin continuation, imaging in patients with implanted devices, surveillance after polyp removal), clinicians define gold-standard decision rules. Each rule is an explicit “if–then” over the graph:

  • Conditions: which concepts, attributes, and links must be present and how recent they must be.

  • Actions: what the system should propose, what it should suppress, and what it should request.

These rules are drafted by multiple clinicians, with disagreements resolved by adjudication.

The memory mode is then evaluated on held-out cases that were not used to train its rules, along four axes:

  • Decision agreement. For each case, whether the system’s recommendation matches the action defined by the gold rule, given the same inputs. Each decision (for example, “propose ARNI now”, “withhold warfarin until INR repeated”, “suppress MRI suggestion”) is treated as a binary classification, and precision, recall, and F1 are reported.

  • Edit-driven learning and learning curves. The memory mode is exposed to clinician edits on a training set over time and evaluated at predefined checkpoints (for example, after 10, 50, 100 edited encounters per decision pattern) on a separate test set. At each checkpoint, measurements include how often the first suggestion matches the gold rule, how many edits are still needed, and how often the system correctly chooses to abstain. This produces a learning curve that shows whether exposure to edits is improving decision quality.

  • Rule reconstruction. The conditions under which the memory mode proposes or suppresses a given action are extracted and treated as its learned rule. This learned rule is compared to the gold rule by checking which conditions are present or missing and by measuring agreement on which cases are classified as eligible or ineligible. Whether adding or removing a single condition changes the recommendation in the same direction as it would for the clinician is also tested.

  • Explanation alignment. For each decision, whether the explanation the system provides — the facts, values, and dates it cites — matches the key facts clinicians themselves cite when asked to justify the same action. This tests whether the reasoning path is faithful, not just whether the final recommendation is correct.

A separate error taxonomy for reasoning failures is maintained, distinguishing between incorrect recommendations, missed recommendations, and unnecessary recommendations, so it is clear how the reasoning layer fails.

Third, what changes in day-to-day work is measured.

For both modes, on held-out encounters:

  • Edits and time. For each note, the number of edits needed to reach a signable version and the time from first draft to signature.

  • Required details. How often side, stage, severity, status, and dates are present in the draft without additional prompting.

  • Chart cleanliness. The rate of duplicate problems, wrong-side entries, and other structural errors that a correct graph should prevent. Using the structural error taxonomy, errors due to missing structure, incorrect structure, and failure to apply existing structure are distinguished.

  • Guardrail performance. Using pre-defined phenotypes for eligibility and contraindications, the precision and recall of safety checks and prompts against that ground truth.

  • Cohort fidelity. For cohorts defined by meaning (concepts, attributes, and links), how many included patients truly meet the intended definition. For each included patient, whether the system can show specific sources, values, and dates that justify inclusion.

  • Learning slope in workflow. At multiple time points, whether repeated edits for the same concepts and decisions decrease and whether clinic-specific phrasing and thresholds become stable. This shows whether the system is reducing friction in real use as it learns.

In the memory mode, all stored expressions must be MRCM-compliant, and all coded statements must be faithful to their sources so that a reviewer can trace each concept, attribute, and link back to the underlying text or document. Structured outputs must also be reasonably calibrated, with uncertainty expressed as abstention or clarification requests rather than confident but incorrect action.

If the memory + SNOMED mode does not clearly outperform the text-only mode on all three layers — coding and graph quality, learned reasoning, and workflow outcomes — and does not meet the legality, faithfulness, and calibration checks, the design is treated as unproven and changed.

Will this bury clinicians in prompts?
It should do the opposite. The system asks only for details that are essential for safety or correctness and not already known; once you record something like laterality, it is reused across orders and notes instead of being re-asked.

Why isn’t a strong language model alone enough?
A strong LLM can write fluent notes, but it is still guessing about the exact problem, site, side, severity, cause, and timing unless those are structured. Safe automation needs those coordinates to be explicit so the same situation produces the same actions and can be audited.

What happens when the system is wrong?
Clinicians remain the final check: every suggestion is editable, and disagreements are expected. When a clinician corrects the system, that edit is stored as a structured training signal on the concepts and rules involved, so the same mistake is less likely to repeat.

Could discrete write-back corrupt the EHR?
Write-back happens only into mapped fields that have been tested in a sandbox, and only with MRCM-valid statements carrying dates and sources. Every change is logged with old value, new value, and reason, so errors are traceable and reversible.

Is this overkill for a small practice?
Even in a small clinic, de-duplicated problem lists, correct side and status carried into orders, reliable surveillance due dates, and guardrails that can show their evidence remove a lot of rework. The machinery under the hood is complex; the visible effect is fewer avoidable mistakes and fewer repeated edits.

Are we locked into your system if we adopt this?
“Locked in” would mean the useful parts cannot leave the product. Here, those are built on SNOMED CT and standard maps, and clinic-specific phrasing and policies are stored as data tied to concepts, so they can be exported and moved to a different stack if needed.

AI-first EHRs can improve patient outcomes; but only if they capture clinical reasoning, not just text. That reasoning is most visible during chart-prep, when a clinician pieces together information across scattered sources to figure out what is actually true about the patient right now and what should happen next. Most of that reasoning never gets written down.

The memory layer captures that edits as structured, computable facts and learns the clinical reasoning behind it. When the system suggests an action, it shows exactly why. When a clinician corrects it, that correction feeds back so the system gets smarter over time.

The initial focus is internists and primary care physicians. They carry high patient loads, their patients return repeatedly over years, they reconcile information from multiple specialists, and they are responsible for long-term care.

This has not yet been proven to work better than a strong language model on its own. But the bet is sound: safe clinical AI requires computable meaning, not just fluent text, and chart-prep is where that meaning is created.

No posts

Read the original on drtattvamnair.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.