A stealth pathogen with a long latent period between infection and onset of symptoms can go undetected for a long time. With HIV/AIDS, we were in the wilderness for more than half a century. HIV started circulating in humans in the early 20th century, but it wasn’t until 1981 that it was formally discovered. By that time, the HIV/AIDS pandemic was already well underway. An estimated 44 million people have died from AIDS-related causes since the start of that pandemic, and an estimated 91 million people have been infected. I believe that most of those infections and deaths could have been prevented, had we had the tools in place to discover the virus earlier.
This tragic history teaches us that building a robust pandemic early warning system against stealth pathogens requires the ability to detect and characterize novel pathogens that either we haven’t seen before or wouldn’t immediately recognize as harmful. Currently, our best approach for detecting novel pathogens is metagenomic next generation sequencing (mNGS), a powerful, almost miraculous set of genomic sequencing methods that capture and characterize the genetic information of all of the organisms in a given sample, matching each organism’s nucleic acid sequences to known reference sequences that allow us to identify their presence and prevalence. Despite this technology’s enormous promise, the current generation of metagenomic sequencing platforms are too costly, slow, and cumbersome to provide reliable early warning against fast-moving potential pandemic threats. Overcoming these bottlenecks will be crucial for providing early warning against future pandemics going forward, and would also offer valuable clinical and public health benefits with routine use.
The reason why cheap, fast, easy-to-use metagenomic sequencing is necessary to provide early warning of future pandemics is that truly novel infectious diseases remain a threat and would likely evade our most widely deployed default diagnostic tools. Such novel pathogens could arise from nature, accidental releases of pathogens modified by humans, or deliberate releases of pathogens engineered by humans. Unfortunately, our default diagnostic tools are not well suited for identifying novel pathogens. Quantitative PCR (qPCR) does an admirable job as the default tool for identifying known pathogens, able to provide specific detection of pathogen genomes within hours, with well-established, robust protocols that can even be deployed in field-portable formats. However, qPCR’s reliance on pre-designed, sequence-specific primers means that it can’t reliably detect novel or substantially modified pathogens. In contrast, metagenomic sequencing has the advantage of not relying on knowing which pathogens you are looking for in advance.
Metagenomic next generation sequencing (mNGS) is a set of methods for analyzing the collective DNA of a community of organisms present in a sample. The process typically uses an untargeted (“shotgun”) approach to sequence all (“meta”) of the genetic information in a sample. This involves several key steps. First, all the genetic material is extracted from the sample. The total extracted DNA is then broken into millions of smaller, random fragments, resembling the dispersal of a shotgun blast. A sequencer then simultaneously records the sequence of base pairs for all of these fragments. Finally, powerful bioinformatics tools are used to process the massive amount of sequencing data. This includes computationally filtering out any host DNA (in the case of clinical samples) and then assembling the remaining microbial DNA fragments into longer contiguous sequences, or “contigs”. These assembled genetic sequences are then compared against comprehensive reference databases to identify the types of organisms present and the potential functions encoded by their genes. Metagenomic sequencing for early detection of emerging pathogens can make use of a number of different sample types, including wastewater, clinical samples, and possibly indoor air sampling in the future. While no sample type is sufficient on its own for comprehensive early detection of biological threats, together they offer the potential for a robust pandemic early warning system.
A playful take on “Shotgun” metagenomics I made with ChatGPT
Metagenomic sequencing of wastewater is a promising new approach that builds on the demonstrated success of the targeted sequencing efforts of the National Wastewater Surveillance System (NWSS) and Traveler-based Genomic Surveillance System (TGSS), which searched for known pathogens such as SARS-CoV-2, Influenza A, RSV, and mPox. Pioneering initiatives such as Secure Bio’s Nucleic Acid Observatory program have taken this a step forward, moving from targeted sequencing of wastewater to demonstrating the feasibility of metagenomic sequencing of wastewater for pathogen surveillance. This is no small feat, given the difficulty of sample preparation and bioinformatic analysis on such complex samples containing a wide variety of human, bacterial, and viral nucleic acids.
Wastewater surveillance is an attractive approach given that it is able to provide a credible public health signal of the presence of a wide variety of pathogens circulating in a human population, while protecting individual privacy. However, wastewater surveillance has a few drawbacks as well. First, the relative abundance in wastewater of various pathogens of potential concern varies widely due both to differences in their prevalence in the human population as well as to differences in the degree to which they are shed in human excretions. This variance adds a layer of technical difficulty to analysis relative to other sample types. This problem can be mitigated with statistical abundance normalization techniques, but it is possible that there are unknown pathogens that could circulate in humans without being detectable in wastewater at all. Second, wastewater surveillance does not directly tie an increase in the prevalence of a pathogen in wastewater to any particular clinical case, and so is insufficient on its own for ascertaining that a new pathogen is actually causing illness in the community. This adds a layer of uncertainty in interpretation that could make it more difficult for decision-makers to initiate a rapid but potentially costly response to an emerging threat, especially if infected people initially do not exhibit symptoms.
Clinical metagenomic sequencing at the point of care also offers an important source of data for early warning against potential pandemics. Clinical metagenomics can be performed on a number of sample types, including respiratory samples (including simple nasal swabs), blood, and feces. In their deeply-researched technology roadmap towards ubiquitous metagenomic sequencing, Whiteford et al note four major advantages of clinical metagenomics relative to other early detection approaches. Most importantly, “If MGS proves to be clinically useful and cost effective, it can scale up naturally within the current health-economic system, without the need for continued public or philanthropic support above those of current diagnostics”. Also important, clinical nMGS can act as a diagnostic tool at the very outset of a novel pandemic, whereas it could take much longer (perhaps months) before PCR-based diagnostics are widely available for a novel pathogen. Moreover, because clinical nMGS enables detection of a pathogen in an individual, while environmental surveillance tools like wastewater nMGS do not, clinical nMGS makes it possible to immediately begin case-based interventions such as contact tracing, isolation of infected individuals, and quarantining of their known contacts. Last, clinical nMGS samples typically offer a more favorable signal to noise ratio than environmental samples.
The biggest challenge for using metagenomics to provide pandemic early warning is that the cost of the technology limits how widely it has been deployed. In the wastewater sampling context, the relatively high cost of mNGS has so far limited its deployment to a few relatively small pilot programs, although hopefully this will change soon. The Nucleic Acid Observatory program is one of these pilot initiatives, run by SecureBio, a biosecurity research organization. Based on cost data obtained for the Nucleic Acid Observatory program and the innovations they’ve been able to achieve in their sample processing and sequencing workflows, SecureBio has estimated that $52M would be sufficient funds to build a core biosurveillance system architecture (which could be expanded later) that could identify novel pathogens before 12 in 100,000 Americans had been infected. This would equate to detection of the novel pathogen when roughly 40,000 Americans in total had been infected- not early enough to contain an epidemic before it gains early momentum, but probably early enough to avert an existential risk to society. Detecting an epidemic even earlier would be considerably more expensive. In a promising initial step towards an early warning system, the President’s FY 2026 Budget proposes a $52M allocation to CDC for a proposed metagenomic sequencing-based pathogen early detection system called Biothreat Radar, although Congress has yet to allocate such funding. If Biothreat Radar is in fact built and sustained, it will go a long way to providing early warning against stealth pathogens. However, because environmental detection of a novel pathogen does not identify specific cases of sick individuals who are infected with that particular pathogen, we will likely still need fairly widespread clinical metagenomics to confirm that a serious biological threat has arrived.
In the clinical diagnostic context, the high cost of nMGS means that it is currently not cost-competitive with PCR, limiting adoption. The sequencing cost per clinical sample for nMGS ranges from $30-$2501, compared to roughly $10 per sample for PCR. As a result, the use of nMGS for diagnostics is typically limited to situations where PCR has first failed to identify the pathogen, rather than as a first option. Nevertheless, when a broader view is taken, it is clear that nMGS can offer substantial clinical value that more than justifies investments in its deployment. Recent modeling has indicated that the introduction of nMGS for clinical diagnostics can save health care systems substantial resources by speeding the correct identification of the pathogen(s) causing bacterial infections, making it easier for health providers to rapidly administer the correct antimicrobial treatments and reducing preventable infection-caused hospital stays.
The second major barrier to the effective use of mNGS for pandemic early warning for both clinical use and environmental sampling is that the turnaround for existing systems is too slow. This problem is particularly acute in the clinical context. PCR test devices such as Cepheid GeneXpert can provide diagnostic results within 45 minutes from the moment a clinical sample is inserted into the cartridge. In contrast, the time-to-result for mNGS diagnostics is typically 5-24 hours. This limits the clinical utility of mNGS and inhibits more widespread clinical adoption. The time-to-result for mNGS applied to environmental samples such as wastewater can be even longer, due to the greater complexity of sample and library preparation as well as the need for greater sequencing depth due to the lower density of the target DNA in the sample. However, environmental sampling can also be done at greater scale, which can facilitate other efficiencies, such as making greater use of pooled samples.
Most mNGS approaches are highly labor intensive, which slows turnaround time while increasing costs, especially for smaller scale labs that have high fixed costs for staff. mNGS requires a number of sequential steps including sample collection and processing, constructing metagenomic libraries, screening of metagenomic libraries, sequencing the DNA, assembling genomes from short sequencing reads, binning sequences by organism or taxa, annotating genes for prediction of their function, and finally bioinformatic analysis, each of which can be a bottleneck. Completing this process in a clinical lab typically requires a team of 5-10 people. This team will typically include a lab technologist to carefully prepare the samples and libraries, and a PhD molecular scientist to administer what can be a finicky sequencer, design protocols, and evaluate sequencing performance. The team will also typically include a bioinformatician, who takes the sequencing data off the sequencer and summarizes it in a report, a microbiologist who evaluates that report and identifies which organisms are clinically important, and a physician who guides the patient’s treatment (the microbiologist and physician are sometimes the same person).
In the future, automation might meaningfully reduce costs, speed turnaround time, and improve ease of use. Major sequencing companies such as Oxford Nanopore are already working on developing machines to automate the process of moving from a raw sample to preparing a library to put on your sequencer, but these new machines are expensive, only suitable for large labs with high volume. In order to gain the benefits of automation at labs with lower volume, it will be necessary to develop much cheaper, laptop-sized microfluidics platforms for sequencing library preparation, such as that of Integra Biosciences. Automation of the bioinformatics component of the process, as BugSeq is doing, is also very valuable.
If HIV and COVID-19 should have taught us anything, it’s that time is the most unforgiving variable in public health. Pathogens turn our delays into their advantage. Metagenomic sequencing won’t eliminate that risk, but it can shrink the window between silent spread and decisive action—if we make it cheap enough, fast enough, and easy enough to use everywhere that matters. That means funding R&D into metagenomics innovation, financing routine environmental metagenomics at scale; reimbursing clinical nMGS so it can become a first-line diagnostic where appropriate; accelerating automation to cut labor and turnaround times; and setting shared standards for data, privacy, and rapid reporting. It also means practicing now—running real exercises, red-teaming pipelines, and measuring detection latency as a performance metric that we are accountable for improving year over year. Some might say we can’t afford to build this infrastructure. I say we can’t afford not to. The choice is not between paying for early warning or paying nothing; it’s between making modest early investment in tools that also pay regular dividends, or facing a preventable international catastrophe later. I think it’s an easy choice.
The sequencing cost per sample for mNGS depends on factors such as the specific sequencing platform used as well as sample volume. Sequencing cost represents just one part of the overall cost of mNGS, which also includes other factors such as the skilled labor for bioinformatics and interpretation of results.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.