High-throughput nucleic acid sequencing has undoubtedly changed the world. Hereditary disease, oncology, GWAS, the entire genomics industry: all of it sits downstream of being able to read DNA cheaply and at scale. But the genome and transcriptome are just a blueprint, the parts list and the instructions; they are not the working machine. The machine is made of proteins.
Proteins are the actuators of life. Enzymes catalyze, receptors sense, antibodies defend, motor proteins move, structural proteins build, signaling proteins decide. Almost everything we’d recognize as biological work is a protein doing it. Almost all therapeutics that exist are aimed at modulating proteins in some way — whether directly activating or inhibiting them, disrupting protein-protein interactions, tuning their expression, or correcting missense/nonsense mutations; the list goes on. Your genome tells you what a cell could theoretically do. Your proteome is the one ‘doing’ right now — and that is where health, disease, and function actually live.
Proteomics has come a long way, in many ways thanks to nucleic acid sequencing, as the latter made it possible to infer what proteins can be out there. In today’s commercial proteomics, this has produced a powerful but still indirect toolkit. Mass spectrometry is the discovery workhorse: it can identify and quantify thousands of proteins per sample, compare protein abundance across conditions, and, with the right workflows, detect peptides, isoforms, and post-translational modifications. Affinity assays took a different route: platforms like Olink and SomaScan trade open-ended discovery for scale, measuring thousands to more than ten thousand predefined protein targets across very large cohorts. And adjacent sample-prep technologies, like nanoparticle-based enrichment, are trying to help mass spec see deeper into high-dynamic-range samples like plasma.
However, the current methods just scratch the surface of the proteome. For one, all of these methods mainly recognize what we already know — they need some form of a reference to compare against. Thus, they lose whatever we don’t have a reference for, and this ‘whatever’ is quite expansive — new proteins and peptides are continuously discovered by re-analyzing indirect data in new ways — imagine how much is still hidden. This is especially relevant in disease. Cancer is notorious for this, as mutations have the capacity to create entirely novel proteins that can have deleterious functions, but it’s not limited to cancer and we just don’t have good ways to discover this directly from the proteome.
The current methods also barely address the heterogeneity of proteins, like cell/tissue/disease-specific isoforms, post-translational modifications (PTMs) on proteins — and these are arguably the most important properties of the proteome. We likely know very little about the breadth of the impact these have — precisely because we don’t have the methods to scalably read this.
By this point, a reader well-versed in the art of proteomics is thinking that “this isn’t an original insight” — and they are right. The promise of unbiased, single-molecule protein sequencing has been discussed probably since NGS first appeared — after all, many of the early NGS pioneers like Jonathan Rothberg and Mark Chee have subsequently launched the first generation of protein sequencing companies over a decade ago.
Still, despite hundreds of millions in funding and many of the brightest minds working on this challenge, unbiased single-molecule protein sequencing remains elusive.
The first-principles explanation is quite literally the Central Dogma of biology — once sequential information has passed into a protein, it cannot be transferred backward (to nucleic acids). So unlike nucleic acids, there is no complementarity in a polypeptide, so nothing can template a copy of it — i.e., there is no PCR for proteins, and PCR is doing pretty much all of the heavy lifting in nucleic acid sequencing, addressing two main fundamental factors required for analyzing single molecules: depth of signal and signal acquisition.
For the former, it gets progressively harder to sequence something as it becomes less abundant, but in the case of nucleic acids, they can be reliably and predictably amplified using PCR, which allows for greatly expanding the signal in a reliable, predictable way. In proteins, the depth of signal is not as much of a problem when you have a lot of copies of a particular protein sequence that you want to sequence (mass spec can do this incredibly well), but it’s obviously not relevant for ‘real’ samples with rare proteins, the ones that actually matter. So with proteins with low abundance, you are essentially stuck with what you have.
The signal acquisition part in most methods of nucleic acid sequencing also depends on the P in PCR — the polymerase typically incorporates a fluorescent nucleotide step-by-step into the complementary sequence, allowing precise imaging of what the nucleotide sequence is. Proteins don’t have this luxury either.
In short, nucleic acids have been built by evolution to relay information, so collecting information from them is relatively trivial. Proteins have been built for function, so it’s not trivial at all.
Diving deeper, proteins are much more fickle chemically and biologically compared to nucleic acids:
The amino acid code contains 20 (or 22 if you’re in the know) canonical amino acids (vs 4 canonical nucleotides), which, at the very least, requires 5x the fluorophores needed for nucleic acids. Once post-translational modifications are added (remember, they are absolutely critical for functionality), the protein syllabary gets expanded considerably — over 650 different PTMs are known
Amino acids are smaller than nucleotides by about 2-3 times, so fundamentally telling one amino acid from the other is harder, which is compounded by the fact that some amino acids are incredibly similar (e.g. leucine and isoleucine, which have the same composition arranged differently)
A nucleic acid chain is uniformly negatively charged, while amino acids come in all flavors — negatively/positively charged, polar and non-polar. So the chemistry that can be used to analyze all amino acids equally is severely restricted
Proteins have significantly more complex secondary and tertiary structures, so it’s much harder to linearize them (a common requirement for sequencing) without breaking something
Dynamic range is something seen in both nucleic acids and proteins, but as we already know, there is no natural PCR to amplify signal in proteins. The dynamic range in proteins is also broader — in blood plasma, protein concentrations span ten or more orders of magnitude, and albumin alone accounts for roughly half of everything present. The proteins you actually care about tend to be the rarest — and with no amplification, there’s no way to dig down to them. Depletion and enrichment help, but they’re lossy and they introduce bias
Additionally, as proteins don’t have PCR, you can’t sequence them with sequential ‘addition’ — you need to either sequentially destroy them, which also often requires harsh chemistry or combinations of enzymes, or use non-destructive methods, where the whole protein is analyzed as is. In the former, read length can become an issue; in the latter, accurate discrimination of the amino acids is a predominant problem.
So what would an ideal method for sequencing single proteins be?
Most fundamentally, it would be a methodology to get signal from a single amino acid sequentially, and be able to discriminate between different amino acids with exceptional precision, ideally label-free, as designing discrete labels for 670+ individual ‘letters’ is not practical. Any method that is unable to do this gets relegated to fingerprinting, which, by definition, depends on some sort of reference, limiting the unbiased/de novo aspect
Be ultra-high throughput, being able to readily sequence hundreds of millions of individual molecules so as to address the dynamic range head-on
Have super-optimized chemistry that addresses all of the quirks seen in proteins
As one can imagine, this is quite an engineering challenge. Looking at the first pioneers of the single-molecule protein sequencing arena, Quantum-Si and Encodia, this becomes quite apparent.
Quantum-Si has developed a way to address the fundamental issue of signal acquisition. Instead of designing a binder against every single N-terminal amino acid, each of the binders recognizes several, and they are bound sequentially in each sequencing cycle, after which a peptidase cleaves off the amino acid. The individual amino acid sequence is then differentiated from the combination of fluorescent binder signals, as well as by the fluorescence kinetics of each single binder binding.
More importantly, it works — Quantum-Si’s Platinum can actually sequence single peptide molecules, detecting up to 17 of the 20 amino acids in certain sequence contexts — and you can buy the device right now.
However, as one can notice, Quantum-Si’s stock has collapsed 90% since going public, and while some of it can be explained by the general biotech downturn, another portion of it is explained by the limitations of the platform. Chiefly, the Platinum has just 2 million wells that can be occupied by single peptides, with each run yielding less than 150K high-quality reads — not nearly enough to address the dynamic range problem.
Moreover, since the binders are not specific to each amino acid, the approach can be considered fingerprinting, as many proteins, especially those with many PTMs, are directly unresolvable. Quantum-Si is working to address many of these problems, chiefly the throughput — the recently announced Proteus system promises significantly more wells, with 80 million in the first version, and scaling up to 10 billion in subsequent versions. The amino acid deconvolution method is allegedly also updated. Will this allow Quantum-Si scalable de novo protein sequencing? We will see.
Encodia, on the other hand, attempted to directly hack the Central Dogma through ‘reverse translation’. Unique binders bound to self-identifying DNA barcodes would recognize N-terminal amino acids, and upon recognition, would ligate the DNA barcode to a growing ‘foundation’ strand, thus creating a conventionally sequenceable DNA strand.
Conceptually, such an approach generally addresses both of the first-principles factors outlined above: repurposing existing and well-validated DNA sequencing machinery to solve signal acquisition, and by adding an additional PCR step, the depth-of-signal problem can also be tackled. It could also be massively scaled to cover the proteome’s dynamic range.
However, despite these advantages, Encodia quietly wound down operations at the end of 2025, after more than $170M in funding. While we are left to speculate, the most likely reason for this is the size of amino acids and the difficulty of telling them apart. Each binder had to be able to do this perfectly, which is a significant ‘physical’ challenge, compounded by the presence of amino acids directly adjacent to the N-terminal one that introduce noise stemming from steric interactions into that process. This is unlike the case with Quantum-Si, where the kinetic information meant that not each binder needed to be perfectly discriminatory. The DNA barcode building process was also likely not straightforward, with significant ‘trans’ reactions (DNA barcode from one peptide being added to another one). The fact that a unique binder had to be designed against each amino acid (and thus, PTM) also likely limited the future potential of the platform.
Despite these setbacks and failures, these pioneering technologies have shown that (a) protein sequencing is actually possible and (b) that other avenues of development can be pursued to address the issues that couldn’t be addressed with these technologies. Several technologies have since been adapted for single-molecule protein sequencing.
From the get-go, nanopore sequencing has been the obvious choice for protein sequencing, chiefly because, unlike with binder-based approaches, the unit of information (the amino acid) is measured ‘directly ’ – the sensor measures the fluctuations in electrical current as the amino acids from a protein pass through the nanopore. As each of the molecule’s ‘features’ is unique, this leads to unique fluctuations that can then be deconvoluted with a high degree of precision — hence, nanopore sequencing is by definition label-free. A nanopore can also thread very long sequences through, thus addressing long proteins with potentially less processing. This whole process can also readily be scaled to hundreds of millions of reads, covering the dynamic range.
Despite these positives and the amount of progress in nucleic acid nanopore sequencing, the aforementioned quirks of proteins create significant bottlenecks to reaching single-molecule protein sequencing on a nanopore:
The electrical signal is typically not deconvoluted feature-by-feature, but by k-mers. As amino acids are both smaller than nucleotides and have a larger alphabet (20 vs 4), there are exponentially more possible combinations of amino acids present in a pore at any given moment, making the deconvolution significantly more complex. For the 20 canonical amino acids, the k-mer combinations exceed 1 quadrillion (versus <1 million for DNA), even using conservative estimates
As mentioned before, proteins are heterogeneously charged, making it harder to thread them through the pore and further increasing electric signal deconvolution complexity. Proteins also generally lack natural enzymes (like helicase in DNA/RNA sequencing) that are able to thread them through the nanopore in very predictable, streamlined steps
The proteins also have to be linearized to go through the pore, and if they form secondary/tertiary structures while in the pore, this can affect the analysis or even bring it to a halt
Thus, the progress up to this point in protein nanopore sequencing has focused on addressing these issues through a variety of means:
Engineering new pores better suited for sequencing proteins, mainly by decreasing their size/length so that fewer amino acids fit in the pore at the same time. Examples include the CytK, FraC, novel variants of the alpha-hemolysin nanopore and more
Looking for and engineering enzymes that allow the protein to thread through the nanopore in a stepwise, predictable way, similar to helicase in nucleic acid nanopores, such as the ClpX unfoldase (Nivala lab) and the 20S proteasome from the archaeon Thermoplasma acidophilum (Maglia lab)
Other methods try to thread a protein through the nanopore directly, without any enzymes, using electro-osmotic flow. Examples include using guanidinium chloride (Aksimentiev group), and/or using engineered pores, some of which are outlined above
Some methodologies connect peptides to a nucleic acid leader that can then be wound into the nanopore via a traditional helicase, such as in this work by the Aksimentiev and Dekker groups (although this approach is not compatible with full-length proteins).
One method that significantly stands out among the others in protein nanopore space is the one developed by our portfolio company, Glyphic Biotechnologies — protein sequencing by expansion (ProSE).
In this methodology, somewhat similar to the recently released Roche Axelios SBX chemistry for DNA sequencing, proteins are first functionalized by attaching an initiating linker to one terminus of each peptide chain, and each amino acid is then sequentially added and uniformly spaced along a uniform linker made from DNA in its original order, constructing the ‘ProSE molecule’. The molecule is then threaded through a nanopore and read as if it were DNA, but every couple of nucleotides, there is an amino acid attached, producing a clean signal requiring no k-mer deconvolution.
Thus, this approach directly addresses all of the aforementioned bottlenecks in protein nanopore sequencing — the number of k-mers in the pore, the lack of uniform charge, and the ability to thread the protein predictably through the pore. More than that, this approach requires no optimization of the nanopore machinery itself, which makes it readily usable in already available, highly optimized nucleic acid nanopore sequencers. Similar to how Roche’s expansion technology for DNA sequencing has achieved records in throughput, speed, and accuracy, expansion technologies have the potential to achieve breakthroughs for protein sequencing as well. Wish we could say more - stay tuned!
While protein nanopore sequencing has been the main direction the industry has been taking in the last couple of years, another approach has been becoming increasingly popular — Raman spectroscopy. What separates it from nanopore sequencing is that in the case of Raman, the signal is deconvoluted directly from the analyte.
When monochromatic light (usually a laser) hits a molecule, most of the light scatters elastically (without changing color or energy); however, about 1 in 10 million photons interact with the molecule’s chemical bonds and scatter inelastically, transferring or absorbing a tiny bit of energy. This energy shift corresponds to the very specific vibrational frequencies of the molecule’s chemical bonds, which can be used to directly tell a molecule apart.
Unlike nanopores, where the signal depends on the size and charge of the molecule going through the pore, and thus in some instances can be confusing to deconvolute, the Raman spectra are much more specific — for instance, it can easily tell apart structural isomers. Raman spectroscopy also ‘cares’ less about the chemical characteristics of the protein, like amino acid size, charge and (to a certain degree) the propensity to form secondary/tertiary structures.
Theoretically, scaling Raman spectroscopy to very high throughput to address the proteome’s dynamic range shouldn’t be a problem — you just scale the amount of proteins immobilized and the cross-section of the laser shining on them. As such, Raman spectroscopy has seen significant interest in academia and industry in the last few years.
However, while Raman spectroscopy has many characteristics needed to make de novo single-molecule protein sequencing work, there are several Raman-specific bottlenecks that need to be addressed:
As mentioned earlier, Raman scattering happens only in 1 in 10 million interactions — so in order to enable single-molecule sequencing, the signal has to be significantly enhanced at least by the order of 108-1010. While this has been achieved in the laboratory setting on single molecules, scaling this to high throughput in a reproducible manner may be a significant challenge and hasn’t been reported in the literature to date.
The other problem is that when Raman scattering happens off a molecule, it happens off the whole molecule — in the case of an immobilized protein, well, the whole protein. While this can be directly used for protein fingerprinting, in order to achieve de novo protein sequencing, a working method would need to sequentially cleave off amino acids, and either measure the single amino acids, which may be challenging to reliably set up, or measure the remaining N-1 protein, which may be challenging to deconvolute the larger the protein is (similar to how the larger the k-mer in the nanopore, the harder it is to deconvolute)
It is worth noting that protein sequencing through nanopore and Raman scattering are not necessarily contradictory — several groups in Europe under the RamanProSeq project umbrella are working on marrying the two. They use a solid-state nanopore to thread the protein through, but also to concentrate the Raman signal at the end of the pore so as to read the protein signal through Raman scattering, and not the electrical current change as in standard nanopores.
Only time will tell. As history of technology shows, things tend to move slowly and arduously first, and then all of a sudden fast and parabolic — just look at AI. As you can probably tell by this article, we are strongly bullish on both the potential of de novo single-molecule protein sequencing and the next wave of technologies being developed now, like nanopore and Raman, being able to crack it.
We are not alone in this sentiment. Many of the companies building in the space have enjoyed considerable support as well, despite the difficulties experienced by the earlier generation.
Since we invested in Glyphic in 2024, another protein nanopore sequencing company, Portal Biotech, spun out from the Maglia Lab work, has raised $35M to scale the technologies they developed. Just a couple of weeks ago, Pumpkinseed, the leading Raman spectroscopy protein sequencing company, announced a $20M Series A. In both cases, there is strong support from government-related entities — Portal got funding from the NATO Innovation Fund, while Pumpkinseed was awarded a DARPA contract focused on protein sequencing, highlighting the importance of de novo single-molecule protein sequencing for society.
De novo single-molecule protein sequencing will change the world in a much more dramatic way than nucleic acid sequencing. If you are building something new in this space, don’t hesitate to reach out!
Special thanks to Josh Yang for reviewing this article and suggesting improvements.
Disclaimer: This article is intended for informational and educational purposes only. Nothing contained herein should be construed as investment advice, financial advice, or a recommendation to buy, sell, or hold any securities or other financial instruments. The biotechnology and pharmaceutical sectors involve significant risks, including regulatory uncertainty, clinical trial failures, and market volatility. Readers are strongly encouraged to conduct their own research and consult with qualified financial professionals before making any investment decisions. Past performance of any company or technology mentioned is not indicative of future results.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.