RSS Amplifier

Neural Horizons Substack · Aug 23, 2026

Evidence Frame Integrity – The Evidence Contact Test

0
Sign in to vote or save

Peter Benson · Neural Horizons Substack

Our previous article in the ‘Evidence Frame Integrity’ series left a useful object on the table: the Edge-Case Ledger. Its purpose was to stop consequential exceptions from vanishing when evidence is compressed into a summary, score, shortlist or dashboard. The ledger asks what the smooth centre of the story left behind: outliers, minority cohorts, contradictory observations, low-confidence cases and alternatives that could change a responsible decision. It also carried a warning. A ledger can exist without anyone opening it. A source link can be genuine without anyone following it. Formal review can occur while the people approving the result never reach the record underneath. [1]

That is the next problem in source-contact collapse. A hallucination can invent a fact or source. False evidentiary posture is subtler: an answer, recommendation or workflow behaves as though the relevant evidence was accessed and verified when that evidence was absent, stale, mismatched, uninspected or replaced by a proxy. The output may even be correct. What is false is the implied relationship between the answer and its evidence.

We separate this failure from ordinary factual error precisely because accuracy, fluency and a convincing explanation do not establish that the required evidentiary surface was actually used. [2]

For evidence-grounded work, that suggests a strong first governance gate: what evidence did the system or the human actually contact, and how do we know?

The Evidence Contact Test is a practical way to answer it.

The Edge-Case Ledger solved a visibility problem. The Evidence Contact Test adds a behaviour problem. [3]

Consider a hiring shortlist produced from applications, interview notes and assessment results. The system may show links to the candidate files. It may preserve a panel labelled “exceptions”. It may even provide a neat audit trail. None of those features proves that the model retrieved the correct files, that the links correspond to the claims being made, or that the hiring panel inspected the material before approving the ranking. A control can be present in the interface and absent in the decision. Our machine-side behavioural framework treats missing, wrong, stale, degraded or unverified evidence as distinct from the later problem of how genuine evidence is compressed into recommendations. [4]

That distinction matters. On the machine side, a recommendation-frame problem arises when artificial intelligence compresses evidence into a shortlist or executive view before people deliberate, while uncertainty, alternatives or source material are hidden or ignored. A separate evidence-frame integrity problem arises when the system acts as if it has contacted the required evidence although that evidence is missing, wrong, stale, degraded, unverified or inferred from cues such as metadata. The first can narrow what humans see. The second can counterfeit the very premise that there was something sound to see. [4]

On the human side, our (human factors) Cognitive Susceptibility Taxonomy describes Recommendation Frame Capture / Evidence Contact Loss: the recommendation becomes the first point of our meaningful contact with the case, rather than a tool used after contact with the underlying evidence. Its warning signs include low source-opening behaviour, absent edge-case review, weak uncertainty display and approval records that show a human choice without showing evidence inspection. This is a susceptibility pattern and governance lens, not a diagnosis of a person. [5]

The Edge Case Ledger also named the institutional version of the problem: the Institutional Blindfold, where an organisation can slide from examining the record to approving a compressed representation of it. In practical terms, the “human line” here is not a demand that people perform every calculation themselves. It is the boundary at which a reviewer can still inspect, contest and change the machine-shaped frame. [6]

There are practical objections here of course; a chief executive cannot read every transaction behind a quarterly dashboard; a clinician cannot reopen every historical note; a teacher cannot audit every token used to produce a student-risk summary. ‘The Edge Case Ledger’ already gave the right answer: the alternative to compressed evidence is not an archive dumped on the desk. It is retraceable compression – a concise view with a direct route to the consequential evidence and exceptions. [6]

We are aware that there is no defensible universal rule such as “open 20 per cent of sources” that turns contact into substance. The required depth depends on consequence, reversibility, evidence quality and the decision-maker’s role. The test should therefore be risk-tiered, not reduced to a single compliance percentage. NIST’s Generative AI Profile likewise treats risk-management effort as something to tailor to context, likelihood and severity. [7]

The Evidence Contact Test can be run as six questions. A “yes” requires observable evidence. “Unknown” is not a pass.

Before reviewing the answer, name the evidence surface the task depends on: the policy document, patient record, contract clause, dataset, interview notes, sensor feed, research paper or current regulation. This prevents a common substitution: judging the quality of the prose before establishing whether the task required access to material outside the model’s visible context. Our Robo-Psychology evidence-frame guidance treats absence, failed upload, stale retrieval, wrong attachment and proxy substitution as distinct reasons to detect a failure of evidence contact rather than behave as though the evidence had been seen. [8]

A citation or file name is not proof of retrieval. For consequential claims, the record should show what source was fetched, which version or date was used, whether retrieval succeeded, and whether any source was unavailable or degraded. NIST’s Generative AI Profile recommends establishing practices for data origin and content lineage and testing flows through original sources, transformations and decision criteria; its AI Risk Management Framework Playbook calls for provenance documentation covering sources, origins, transformations, dependencies, constraints and metadata. [9]

This is where stale evidence matters. A perfectly quoted policy superseded six months ago can produce a well-supported wrong decision. Contact is temporal as well as semantic; our Evidence-Frame Integrity Overlay explicitly treats stale or wrong evidence as a failure condition. [8]

Retrieval is only the middle of the chain. The retrieved passage might be relevant to the topic while failing to justify the sentence attached to it. Research on retrieval-augmented generation has made this separation explicit. ARES evaluates context relevance, answer faithfulness and answer relevance as different dimensions; RAGAs likewise separates retrieval quality from faithful use of the retrieved material and from the quality of the generated answer. [10]

Citation research reached the same conclusion earlier. The 2023 ALCE benchmark scored fluency, correctness and citation quality separately; in its ELI5 experiments, even the strongest systems in that study lacked complete citation support half the time. That figure should not be treated as a current failure rate for today’s systems. Its enduring lesson is structural: the presence of citations and the support they provide are different measurements. [11]

This carries forward the Edge-Case Ledger. Ask what outliers, contradictory cases, missing values, subgroup effects, uncertainty and plausible alternatives were suppressed by the summary or ranking. The important test is not whether anything was omitted – every useful summary omits – but whether the omitted material could change the decision, the distribution of harm, or confidence in the recommendation. [12]

A signature is an event. Evidence contact is an activity.

Our Robo-Psychology recommendation-frame guidance proposes looking for source opening, edge-case inspection, alternative review, uncertainty review, a real opportunity to challenge, and a recorded rationale. Add to this, our Positive Dyad / Co-Evolution Capability Overlay – the project framework concerned with whether the human–AI relationship preserves useful human capability – which adds direct evidence access and “time in evidence” as relevant indicators. None of these requires an executive to become the analyst. They require the organisation to distinguish an approval click from judgement. [13]

For high-consequence decisions, record which decision-driving sources or samples the reviewer actually inspected, which exception was checked, and whether the reviewer had authority to recover an excluded option or reject the AI frame. A person who is technically allowed to inspect evidence but is penalised for slowing the queue has a very different form of oversight from someone given time, authority and a workable route to challenge.

Our project frameworks explicitly treat throughput pressure, challenge opportunity and alternative recovery as relevant to whether review is substantive. [14]

Starting from the final claim or recommendation, an auditor should be able to move backwards through the AI output, the retrieved passages or data, relevant transformations, source version and responsible actors. NIST explicitly calls for data and content lineage; the World Wide Web Consortium’s PROV data model supplies a general vocabulary built around entities, activities and agents, including relationships such as use, derivation and attribution. [15]

The counter-case is that automated checks can perform much of this work. They can. Retrieval tests, faithfulness scoring, provenance logging and automated evaluators can reduce the human burden, especially at scale. ARES and RAGAs are examples of attempts to automate parts of that evaluation. The mistake would be to let an automated score certify its own evidentiary premise. If the evaluator is testing the wrong source, stale corpus or incomplete record, efficiency simply accelerates the wrong assurance. [10]

We recognise that retrieval-augmented generation evaluation is fast-moving, benchmark-dependent and sensitive to domain. The studies above support decomposing the problem into distinct checks, but and not intended to establish a universal threshold for safe evidence contact in medicine, education, hiring, law or public administration. NIST itself notes continuing measurement uncertainty in generative-AI risk assessment. [7]

“Provenance” is often offered as the cure for this problem, and it is indispensable. It is also easy to ask it to do too much.

A useful provenance record can tell us where an artefact came from, what transformations occurred, which system or person acted on it, and which version entered the workflow. The W3C PROV model formalises those relationships through entities, activities and agents. NIST’s Generative AI Profile asks organisations to establish practices for data origin and content lineage and to test flows through original sources, transformations and decision-making criteria. [16]

Think of provenance as parcel tracking. It can show that a package moved from a named sender through a known depot to your door. That is valuable. It does not tell you that the sender put the correct medicine in the box.

The Coalition for Content Provenance and Authenticity makes this boundary unusually clear. Its C2PA specification is designed to make provenance claims verifiable and resistant to undetected tampering, but its own guiding principle says the specification does not judge whether provenance data are “good” or “bad”; it validates their association with an asset, their form and their integrity. A trusted trail therefore cannot, by itself, prove that a source is accurate, current, representative or relevant to the claim. [17]

This matters because false evidence posture can survive excellent logging. Imagine a system that faithfully records that it retrieved Policy_v7.pdf, then accurately shows the paragraphs it used, while the operative policy is version 9. The lineage is clean. The decision is still grounded in the wrong thing.

Evidence integrity requires two checks that should never be merged: Can we trace the path? and Was the path evidentially valid?

NIST’s lineage guidance and the Robo-Psychology evidence-frame control address different parts of precisely this distinction. [18]

We do need to note that this can add cost to the operating model. Fine-grained provenance can become another compliance machine: expensive to retain, difficult to interpret and easy to produce at a granularity no decision-maker will use. That concern argues for consequence-based retention rather than maximal logging. Keep enough information to reconstruct consequential claims and decisions; do not turn every low-risk drafting interaction into a forensic archive. NIST’s framework is explicitly risk-management oriented and voluntary, and its website states that AI RMF 1.0 is being revised in 2026, so it should not be treated as a checklist frozen in time. [19]

The Evidence Contact Test covered here is a synthesis, not an accredited standard: provenance methods help establish lineage; the project frameworks add separate tests for source validity, compression, human review and preserved decision authority. That synthesis still needs a level of domain-specific validation. [20]

False evidence posture works because it often arrives wrapped in genuine usefulness.

Summaries save time. Retrieval systems can bring a large document collection into reach. Ranked options can help an overloaded manager navigate plausible choices. RAGAs describes retrieval-augmented generation as a way to connect language models to reference databases and reduce hallucination risk, while emphasising that retrieval quality and faithful use of retrieved passages remain separate evaluation problems. The benefit is real; so is the need to measure what happened between source and answer. [21]

The human vulnerability begins when useful form substitutes for evidential criteria. Our Cognitive Susceptibility Taxonomy calls this Discursive Validity / Criteria Collapse: fluent, well-structured, citation-rich or numerically plausible output can acquire credibility because it resembles the form of work we normally associate with credibility. The taxonomy flags low second-sourcing and confusion between confidence and proof as warning signs. Again, this is not a diagnosis. It is a description of a review condition that interfaces and institutions can amplify. [22]

Workload matters. So do incentives. If the dashboard offers one large green recommendation and hides source material behind six clicks, “human oversight” is being shaped before the person makes any conscious choice. If a review queue rewards speed, the reviewer who follows citations and reopens edge cases pays a productivity tax for doing the epistemically responsible thing. People are not failing because they suddenly stopped caring about truth. They are adapting to a workflow that makes verification costly and acceptance cheap. Our recommendation-frame material explicitly identifies throughput pressure, answer-first workflows and formal approval without evidence inspection as amplifiers of evidence-contact loss. [23]

This is where DAUS-5, our project’s five-layer uplift gate, becomes useful. We refuse to call a system an improvement merely because the immediate task is faster or smoother if reality-tracking, agency, skill, relational integrity or governance substance deteriorate.

Crucially, where relevant layers have not been measured, we instruct reviewers to mark them “not instrumented” rather than infer success from task completion, satisfaction, low complaint rates or silence. [24]

That is an unusually important discipline for AI governance. A dashboard showing “95 per cent reviewer acceptance” is not evidence that reviewers verified the recommendations. It may be evidence of agreement. Without contact measures, we do not know which. The distinction follows directly from our separation of formal approval from substantive evidence contact. [14]

Evidence-Contact Discipline, as defined in the Positive Dyad / Co-Evolution Capability Overlay, points towards selective safeguards: source access, visible uncertainty and edge cases, recoverable alternatives, and human inspection before high-stakes action. [25]

We don’t yet know if ‘one’ evidence-contact design will preserve human judgement across all populations and professions. Our Cognitive Susceptibility Taxonomy itself distinguishes provisional or “not instrumented” measures from stronger validation status. Our constructs are most defensible here as hypothesis-generating prompts for observable controls – source opening, challenge, alternative recovery and evidence inspection – rather than claims about a reviewer’s inner state. [26]

The Evidence Contact Test earns its place only if it changes a workflow. Four moves are enough to begin.

Within thirty days, put the test into one consequential decision. Choose a workflow where an AI-generated summary, ranking or recommendation materially shapes a person’s action: a board risk pack, procurement recommendation, hiring shortlist, student-support queue or another locally relevant process. Define the evidence that must exist before the system may make evidence-dependent claims. Require the output to show source identity and date, retrieval status, material uncertainty, an Edge-Case Ledger and at least one route back to a down-ranked or excluded alternative. Assign a named decision owner who can reject the AI frame. This turns the project’s recommendation-frame and evidence-contact controls into operating requirements rather than another general “human in the loop” statement. [13]

By day sixty, run an evidence-ablation drill. Test the workflow with evidence deliberately absent, wrong, stale, inaccessible or represented only by proxy cues such as a file name. The safe behaviour is not eloquent improvisation; it is detecting the loss of evidence, deferring, reducing confidence or requesting the missing material. In a second pass, seed consequential edge cases and see whether both the system and reviewers recover them. NIST recommends red-teaming to probe adverse or unforeseen behaviour, while the project frameworks specifically identify absent, wrong and stale evidence as test conditions. [27]

By day seventy-five, make the chain reconstructable. Preserve the source identifier and version, retrieval result, transformations that materially affected the evidence, relevant system version, final output, and human decision with any override or challenge. Use provenance standards as a model for lineage, not as a truth badge. A later reviewer should be able to move from decision back to evidence without reverse-engineering the organisation’s software. [28]

By day ninety, change what the governance dashboard rewards. Keep measures of speed and task quality, but add evidence-contact measures appropriate to risk: the share of high-stakes cases with verified source retrieval, sampled claim-support accuracy, edge-case inspection, alternative recovery, reviewer challenge, and time spent with decision-driving evidence. Where a relevant dimension has not been measured, say “not instrumented”. Audit a sample of approvals against the underlying record, not merely against the AI summary. [29]

Done correctly, a good design should add little friction to low-consequence work and deliberate friction where a false evidence posture could affect rights, safety, opportunity, money or institutional accountability. That risk-tiered approach is consistent with NIST’s emphasis on allocating governance effort according to context and consequence. [30]

The trade is not speed versus safety in the abstract. It is a small, visible cost of verification against the hidden cost of making a consequential decision on evidence nobody actually touched.

The next risk appears after the decision. Once an AI summary is copied into meeting minutes, a case file, a student record, a risk register or another durable system, later humans and later models may retrieve the summary as if it were the evidence itself. This is an inference from the lineage problem, not a claim that every organisation already behaves this way.

But it is the natural next question for this series: what happens when compression stops being a temporary aid and becomes part of the institutional record? NIST’s emphasis on source lineage and transformation history shows why that distinction matters. [9]

That is where our next article in the ‘Evidence Frame Integrity’ series goes next: When Summaries Become Records.

No posts

Read the original on neuralhorizons.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.