*Co-authored by Eugene Vyborov and Cornelius, an AI reasoning agent built on Ability AI’s Trinity platform.*
*On May 6, 2026, a real outbreak was unfolding. I tasked Cornelius - my AI reasoning engine - with modeling its trajectory using structured probabilistic thinking. One analytical move every 15 minutes. Here is the full report.*
How it was done
● The event: MV Hondius, a polar expedition vessel, had a confirmed Andes virus (hantavirus) outbreak - 8 cases, 3 deaths, 23 nationalities across 147 passengers
● The task: Model the probable trajectory and escalation conditions using a systematic AI analytical loop
● The method: 9 analytical cycles using rotating moves - ACH Audit, Bayesian Update, Steelman, Cross-Domain Bridge, Implication Check, Assumption Audit
● The conclusion: 59% probability of contained primary exposure cluster; 39% probability of limited H-to-H seeding; pandemic risk effectively eliminated (~0%)
● The methodology finding: The AI caught its own confirmation bias in Run 9 and self-corrected
On May 6, 2026, the WHO issued Disease Outbreak News bulletin DON599 confirming a hantavirus cluster aboard the MV Hondius, a Dutch-flagged polar expedition vessel operated by Oceanwide Expeditions.
The numbers were alarming:
● 8 confirmed or suspected cases
● 3 deaths (2 on the ship, 1 in Johannesburg after evacuation)
● Confirmed virus: Andes virus (ANDV) - the only hantavirus with documented human-to-human transmission
● 23 nationalities aboard 147 passengers and crew
● An international seeding event already confirmed: a Swiss patient in Zurich, a death in Johannesburg
● Incubation window still open: up to 40 days from exposure
The question I posed to Cornelius: *What is the probable trajectory of this outbreak, and under what conditions does it escalate from a contained incident to a global public health emergency?*
This is not a hypothetical. We ran this analysis live, in real time, as the outbreak unfolded.
Before the methodology, the epidemiological baseline matters.
Andes virus (ANDV) is the only hantavirus with documented human-to-human (H-to-H) transmission. Every other hantavirus requires rodent-to-human contact via aerosolized rodent excreta. ANDV can spread person-to-person - a property confirmed by the 2018-2019 Epuyén outbreak in Argentina, which generated 34 cases and 11 deaths from a single introduction.
From the NEJM’s *”Super-Spreaders” and Person-to-Person Transmission of Andes Virus in Argentina* (2020), the key parameters:
An R0 of 2.12 with a 30% CFR and a 40-day maximum incubation window, on a vessel that had passengers of 23 nationalities returning to 23 different countries - this is a legitimate modeling challenge.
The analytical engine Cornelius uses runs a rotating sequence of six research-validated analytical moves, applied in sequence to a central question. Each cycle advances the analysis by one move:
1. ACH Audit (Heuer/CIA) - Analysis of Competing Hypotheses: eliminate weak hypotheses using evidence
2. Bayesian Update - Explicit probability revision using likelihood ratios
3. Steelman Opposition - Build the strongest possible case against the leading hypothesis
4. Cross-Domain Bridge - What does an unrelated field reveal about the mechanism?
5. Implication Check (Popper) - If the hypothesis is true, what follows? Does the evidence support it?
6. Assumption Audit - Which load-bearing premises are shakiest?
The loop fires every 15 minutes. Each run advances one move per active topic, appends structured reasoning to a persistent file, and checks for convergence. A topic converges when:
● The same hypothesis leads for 3+ consecutive runs
● The confidence delta between the last two runs is under 5 percentage points
● No major analytical open questions remain
This produced 9 analytical cycles over a single session, converging to a stable answer.
At the start, four scenarios were defined:
The first ACH Audit immediately elevated H2. Two pieces of evidence were decisive:
7. The Swiss patient in Zurich and the Johannesburg death were already confirmed international seeds - meaning H2 (global seeding) was at minimum already true
8. The ship attack rate was only 5.4% after 35 days with 147 people in close quarters - far below what an R0 of 2.12 would produce if H-to-H transmission were running unchecked
H4 (pandemic) dropped to 3%. H2 rose to 42% and took the lead. H1 was partially falsified: “fully contained” was no longer possible since international seeding had already occurred, so H1 was redefined as “total cases <30, no sustained chains.”
The first Bayesian Update then incorporated the key early diagnostic: Tristan da Cunha had confirmed zero cases. This remote island (population ~250) had hosted passengers during a shore visit April 13-15. Day 21 of the 40-day maximum incubation window had passed with no island cases. Using explicit likelihood ratios, H3 dropped from 15% to 11%, H1 rose to 45%.
The steelman against H1 exposed two genuine structural weaknesses:
9. The Tristan da Cunha surveillance gap: TdC has one resident doctor and basic diagnostic capacity. ANDV’s early prodrome (fever, myalgia, GI symptoms) is clinically indistinguishable from influenza without specific testing. A “no cases” report from a 250-person remote island with minimal lab infrastructure is not the same as a WHO-verified epidemiological assessment.
10. The iceberg problem: The ship had 88 passengers from 23 nationalities. Each returned home through different healthcare systems. ANDV is so rare that most frontline physicians have never encountered it. A patient presenting with fever and respiratory distress after travel to South America is diagnosed with influenza or atypical pneumonia - not hantavirus. The Swiss case was only identified because Zurich hospital connected the symptoms to the voyage.
H2 retook the lead at 45%. H1 dropped to 42%.
This was the analytically richest run. The search of Cornelius’s knowledge base surfaced an unexpected connection: Van Steen and Tanenbaum’s distributed systems work on gossip protocols.
Gossip protocols use the same mathematical framework as epidemic models. The rumor-spreading equation s = e^(-(1/p_stop + 1)(1-s)) is structurally identical to epidemic SIR models - a fact not coincidental, since epidemic protocols were named after biological disease spread. The translation works in both directions.
The insight: in unstructured networks (the dispersed post-disembarkation network of 88 passengers now scattered across 23 countries), epidemic propagation depends entirely on whether a hub node - a super-spreader - exists. In gossip theory, hub presence reveals itself quickly through indegree distribution anomalies. In epidemiology, a super-spreader generates a cluster of secondary cases within one serial interval (~23 days).
We were 30 days past the index case. No cluster had emerged. This was meaningful structural evidence against H3. Without a super-spreader hub in the dispersed network, gossip theory predicts the epidemic burns out locally.
Additionally: applying COVID-19 surveillance iceberg research to ANDV, the estimated true global case count at the time was 32-48 - firmly in H2 territory.
H2 consolidated at 49%.
The Implication Check validated H2 across five testable predictions:
● Cases distributed across Netherlands, Germany, Switzerland, UK, South Africa - confirmed
● Case emergence pattern: a trickle (1-2/week), not a burst - confirmed
● Geographic spread tracking nationality distribution of a Dutch polar expedition vessel - confirmed
● Case timing consistent with 23-day serial interval - confirmed
● No healthcare worker infections at Dutch/German hospitals treating evacuees - confirmed
H2 peaked at 51%.
The Assumption Audit then pulled it back. Two critical findings:
11. A 2021 systematic review challenged the historical evidence for ANDV human-to-human transmission, describing it as “not supported by sufficient evidence” with “flawed methodology.” This raised the possibility that the current outbreak had even less H-to-H than assumed.
12. The 4-6x iceberg multiplier was revised down to 2-3x. ANDV causes severe disease (ARDS, shock) requiring hospitalization - detection rates are far higher than COVID’s mild/asymptomatic presentation. True global cases: closer to 16-24, not 32-48.
H1 and H2 were now converging toward the same narrative at different quantity thresholds. H2 retreated to 47%.
The dominant new evidence: WHO’s formal public statement that “the Dutch couple, who had been travelling in Argentina before boarding the cruise, were infected off the ship.”
This directly resolved the longest-standing ambiguity. If WHO was correct, Cases 1 and 2 were primary rodent exposures in Argentina - not H-to-H on the ship. Combined with four simultaneous negative diagnostics (no TdC cases, no household clusters from evacuees, no super-spreader event, stable case count), H1 took the lead for the first time at 47%.
H3 (sustained chains) dropped to 8%.
Run 7 had incorporated the WHO statement qualitatively via ACH. Run 8 formally computed the Bayesian posterior using likelihood ratios.
The key insight here came from Cornelius’s knowledge base: a note on *how institutional public commitments carry higher evidential weight than internal assessments* - because a formal public statement has institutional reversal cost. WHO would only issue it with reasonable confidence. This makes the statement stronger evidence than face value suggests.
Likelihood ratios (WHO Argentina-primary statement):
● P(WHO issues this | H1) = 0.85 → LR vs H3: 3.40
● P(WHO issues this | H2) = 0.65 → LR vs H3: 2.60
● P(WHO issues this | H3) = 0.25 → baseline
● P(WHO issues this | H4) = 0.05 → LR vs H3: 0.20
Combined with 48h case count stability, the formal posterior: H1: 63%, H2: 36%, H3: 1%, H4: ~0%.
H3 was now near-eliminated across four independent negative diagnostics. H4 was eliminated.
With H1 at 63% after a large single-run jump, the Steelman move was applied against H1 itself. This is where the methodology became most interesting.
The steelman identified three genuine weaknesses in H1’s 63% position:
The husband-wife onset gap. Case 1 (husband) developed symptoms approximately 18-20 days before Case 2 (wife). They traveled together in Argentina before boarding and shared all activities. If they were both infected by the same rodent at the same time, their symptom onsets should be correlated. The documented ANDV serial interval is 23 ± 7 days - an 18-20 day gap falls directly within this range. H2 (husband infecting wife on ship via H-to-H) is the *more parsimonious* explanation for this specific data point. WHO’s Argentina-primary assessment doesn’t adequately account for this anomaly.
WHO’s hedged language. “WHO believes... suggesting they may have contracted” - double-hedged phrasing (believes + suggesting + may have). This is field epidemiology, not molecular confirmation. The genomic sequencing that would definitively resolve H1 vs H2 (multiple independent haplotypes = primary exposures; single haplotype chain = H-to-H) had not yet been returned from Institut Pasteur de Dakar.
Confirmation accumulation bias. The most important meta-finding: nine consecutive runs that found no H3 or H4 evidence had generated their own confirmation stream. Cornelius recognized this pattern - the mechanism behind confirmation accumulation bias applied *reflexively to the analysis itself*. A deliberate calibration discount of 3-4 percentage points was warranted.
The steelman verdict: H1 survives as the leading hypothesis but is not as strong as 63%. A modest correction was applied.
Final converged position: H1 = 59%, H2 = 39%, H3 = 2%, H4 = ~0%.
Convergence confirmed: H1 led for 3 consecutive runs, delta was 4pp (under the 5pp threshold), remaining open questions are empirical rather than analytical.
What this outbreak is (59% probability):
A cluster of primary Andes virus infections contracted in Argentina/Patagonia during pre-voyage travel, now globally dispersed among returning passengers of 23 nationalities. No sustained community chains. Contact tracing is functioning. Expected outcome: 12-25 detected cases globally before June 7, spread across 4-8 countries. Most deaths already counted.
What it might also be (39% probability):
Some human-to-human transmission occurred on the ship - specifically, the husband may have infected the wife (the 18-20 day onset gap fitting the documented ANDV serial interval is the strongest remaining evidence for this). Even in this scenario, the conclusion is the same: no community chains, no epidemic trajectory. H2 and H1 differ on mechanism and final case count, not on outcome.
What it is not (effectively 0%):
A pandemic. H4 is eliminated. ANDV has never demonstrated pandemic spread despite a 30-year outbreak history. No super-spreader hub was detected in the dispersed network. No community clusters formed in 30+ days.
Three observable events will close the remaining analytical uncertainty:
13. Institut Pasteur de Dakar genomic sequencing (~May 20, 2026): Multiple independent haplotypes = H1 confirmed (primary exposure in Argentina). Single haplotype chain = H2 confirmed (H-to-H on ship). This is the definitive test.
14. Household contact surveillance from evacuated patients (~May 27): If household members of the Dutch, German, and UK evacuees develop ANDV symptoms in the incubation window, H3 probability rises. If none develop symptoms, H1 consolidates.
15. Total case count by June 7: <15 detected = H1; 15-30 = H1/H2 boundary; >30 = H2 confirmed.
Three observations about this process that go beyond the outbreak itself:
1. Explicit probability revision prevents drift. When you force yourself to compute likelihood ratios and write down posterior probabilities, you cannot quietly drift toward a conclusion. Each run’s confidence level is a matter of record. The large 47%→63% jump in Run 8 was immediately visible as anomalous and triggered the Run 9 steelman.
2. Cross-domain bridges produce non-obvious insights. The gossip protocol bridge (Run 4) was not in any epidemiology textbook. Van Steen and Tanenbaum’s distributed systems mathematics provided a structural model for why ANDV without a super-spreader hub would burn out in an unstructured dispersed network. This is what cross-domain consilience produces: not analogies, but shared mathematical structure.
3. The system can catch its own confirmation bias. The Steelman move in Run 9 was structurally required by the methodology rotation - it happened regardless of whether the analyst was satisfied with the current answer. It found two genuine weaknesses (onset gap, WHO hedging) and one methodological bias (confirmation accumulation). The deliberate self-correction reduced H1 from 63% to 59%. Whether or not that specific correction was perfectly calibrated, the *process* of forced steelmanning every 3-4 runs is what makes the conclusion trustworthy.
H2 leads Runs 1-6, H1 takes lead Run 7, CONVERGED Run 9
● WHO Disease Outbreak News DON599 (2026-05-04)
● *”Super-Spreaders” and Person-to-Person Transmission of Andes Virus in Argentina* - NEJM (2020)
● Van Steen & Tanenbaum - *Distributed Systems: Principles and Paradigms* (2023)
● 2021 PMC Systematic Review on ANDV Human-to-Human Transmission Evidence Quality
● Tristan da Cunha Government News Release (May 4, 2026)
● CIDRAP, Africa CDC, STAT News coverage (May 2026)
*Cornelius is an AI reasoning agent developed by Ability AI, running on the Trinity platform. This analysis was conducted autonomously using an Incubation Loop architecture - 9 structured analytical cycles over a single session. Eugene Vyborov is the founder of Ability AI.*
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.