August 14, 2026
“It was very special when we realized: this really is something new,” Sebastiaan van Heesch said in the Princess Máxima Center and Oncode account of the work. The excitement was earned. The noun was not.
On May 6, 2026, a consortium of gene annotators, proteomics researchers, cancer biologists, and computational scientists published a result that sounded like a correction to the Human Genome Project. They had searched 7,264 stretches of the genome outside the conventional protein catalogue and found peptide fragments displayed by human leukocyte antigen (HLA) molecules mapping to 1,785 of them. The institutional announcement compressed the result into “1,700 new proteins.” The paper made a harder claim: translation products had been detected, but most had not earned conventional protein status.
A molecule can be real before biology knows what it means.
That is why the most important result in the TransCODE consortium’s Nature paper is not 1,785. It is a word: peptidein.
A peptidein is the paper’s label for an ORF with endogenous peptide evidence supporting a translation product, but not enough evidence for conventional protein-coding-gene status. It may matter only in cancer. It may regulate another gene. Its product may be transient material the cell quickly destroys. It may remain a detected signal with no established function.
That sounds like taxonomy. It is actually an argument about what science permits itself to see.
1,785 is not a count of newly established human proteins. It is the number of candidate noncanonical ORFs to which the HLA PeptideAtlas analysis mapped peptide evidence.
That distinction separates a remarkable detection result from a finished annotation, a drug target, or a treatment. The rest of the story turns on three questions: what was detected, what persists, and what changes biology.
One candidate tests every boundary. c10riboseqorf92 survives broad cancer-cell screens and an ORF rescue experiment, yet remains a peptidein because its evidence stops at transformed cells. Its journey through the paper will show why the name matters.
The Human Genome Project did not produce one final number of human genes and engrave it into stone. The number fell as the sequence improved and the definition tightened.
The 2001 draft human genome sequence estimated 30,000 to 40,000 protein-coding genes. The finished euchromatic sequence published in 2004 revised the estimate to roughly 20,000 to 25,000. Two decades of annotation, comparative genomics, RNA evidence, and protein evidence brought the reference set down again. The 2026 TransCODE paper begins from approximately 19,500 canonical protein-coding genes.
This history is sometimes retold as if the genome project failed to find a hidden continent of DNA. It did not. The sequence and the annotation are different maps.
A sequence map asks: what letters are here, and in what order?
An annotation map asks: which stretches are genes, which transcripts are made, which regions are translated, which molecules persist, and which products do something biologically meaningful?
The first map can approach completion while the second remains provisional. The genome can be sequenced without every sentence being understood.
For years, the shrinking protein count looked like maturity. The main protein-coding catalogue appeared stable. GENCODE and UniProt could refine boundaries, isoforms, start sites, and evidence levels, but the broad inventory no longer seemed likely to double. That stability mattered because nearly every layer of biomedical work depends on it. A gene name becomes a database entry. A database entry becomes an assay. An assay becomes a disease study, a knockout experiment, a diagnostic interpretation, or a drug-discovery program.
Reference annotation is therefore not clerical work. It determines what experiments are easy to imagine.
If a sequence is absent from the reference protein catalogue, a standard proteomics search may never look for its peptide. A disease-association pipeline may treat a variant inside it as noncoding background. A functional screen may not include it. A structural model may never be generated. The thing can exist in a cell and remain nearly invisible to the machinery built to study cells.
The catalogue did not merely describe biology. It decided what counted as searchable biology.
The old map was not foolish. It was conservative for good reasons. Cells produce enormous amounts of RNA. Ribosomes sample sequences that may never yield stable molecules. Short sequences are difficult to distinguish from chance. Mass spectrometry can misassign spectra. Evolutionary conservation becomes statistically slippery when an open reading frame is only a few dozen amino acids long. A database that promotes every signal into a protein would replace one blind spot with a swamp of false positives.
The unresolved problem was not whether to lower the standard. It was how to describe evidence that fell between the available categories.
“Protein” implied too much. “Noncoding” implied too little.
An open reading frame is a stretch of sequence that can be translated from a start signal until a stop signal. The familiar version sits inside a protein-coding transcript and produces the protein named by the gene. But transcripts can contain other reading frames before, after, across, or partly overlapping that main coding sequence. Transcripts labelled long noncoding RNAs can contain short reading frames too.
The terminology grew faster than the consensus. Researchers wrote about short ORFs, small ORFs, upstream ORFs, noncanonical ORFs, novel unannotated ORFs, small ORF-encoded peptides, micropeptides, and microproteins. These categories overlap. They are not synonyms.
An upstream ORF may regulate how efficiently a ribosome reaches the main coding sequence. Its act of translation can matter even if its peptide product does not. A sequence inside a nominally noncoding RNA may produce a stable microprotein with a cellular job. Another may produce a short-lived fragment that is degraded almost as soon as it appears. A tumor may translate sequences that normal tissue rarely uses. The immune system may display pieces of those products without the intact molecule ever accumulating to conventional protein abundance.
The first major change came from learning to watch ribosomes rather than merely measuring RNA.
In 2009, Nicholas Ingolia and colleagues introduced ribosome profiling at a resolution that could reveal where translating ribosomes sat on messenger RNAs. Their early yeast work found structured ribosome occupancy beyond known coding regions, including upstream reading frames. Mammalian studies then showed translation beginning at thousands of previously unannotated sites, including very short ORFs and non-AUG start sites.
Ribo-seq changed the question from “Is this RNA present?” to “Is a ribosome reading it?”
That was a breakthrough and a trap.
Ribosome occupancy is evidence of translation. It is not by itself proof of a stable protein. The distinction became one of the central rules in the 2022 TransCODE annotation framework. A ribosome may translate a regulatory upstream ORF. It may initiate imperfectly. The product may be destroyed. The reading event may matter because it changes translation downstream, not because the tiny peptide becomes an independent actor.
The field nevertheless had clear examples showing that some products were more than translational debris.
The 7-kilodalton microprotein NoBody, translated from the transcript now known as LINC01420, interacts with mRNA-decapping machinery and changes the number of processing bodies in cells. The 2017 Nature Chemical Biology paper established a named molecular interaction and a cellular phenotype.
The 54-amino-acid microprotein PIGBOS localizes to the mitochondrial outer membrane, interacts with the endoplasmic-reticulum protein CLCC1, and affects the unfolded-protein response. That conclusion came from fractionation, imaging, proximity labelling, co-immunoprecipitation, and knockout experiments, not from ribosome footprints alone. The evidence converged on a specific stress-response function at an organelle contact site.
The 56-amino-acid mitoregulin is associated with mitochondrial membrane biology, respiratory complexes, calcium retention, and fatty-acid oxidation. Its functional evidence spans human cells and knockout mice.
These examples do not prove that thousands of short translated ORFs are functional proteins. They prove that length and old annotation are not valid reasons to dismiss all of them.
The phrase “noncoding region” bundles together at least four possibilities:
sequence that is not detectably translated;
sequence whose translation regulates another coding event;
sequence that produces an unstable or context-specific product;
sequence that encodes a stable molecule with an independent function.
The cell does not owe the database a clean boundary between them.
A conventional protein leaves several kinds of evidence. Its RNA is transcribed. Ribosomes translate it. The intact molecule or its fragments can be detected. Its sequence may be conserved. Removing it may change a cell or organism. Restoring it may reverse that change. The protein may bind a partner, occupy a structure, catalyse a reaction, or alter a pathway.
For a canonical protein, these lines of evidence often reinforce one another. For a noncanonical ORF, they can point in different directions.
The TransCODE study treated that disagreement as the subject rather than the inconvenience. The consortium combined four major evidence systems.
Ribo-seq captures fragments of RNA protected by ribosomes. The three-nucleotide rhythm of those fragments, their location, and their reproducibility can support active translation. In the 2026 study, the starting catalogue consisted of 7,264 ORFs already supported through a GENCODE and TransCODE consensus process.
But ribosome profiling can be ambiguous around very short regions, overlapping frames, and alternative start sites. Translation can be regulatory. A strong Ribo-seq signal moves a candidate out of “nothing happened here.” It does not move it all the way to “functional protein.”
Mass spectrometry usually does not observe a protein as an intact object. Researchers digest proteins, most often with trypsin, into peptides, separate them, fragment them again, and match the resulting spectra against candidate sequences.
The TransCODE group assembled a non-HLA PeptideAtlas build containing about 3.5 billion MS/MS spectra. They searched those spectra against the 7,264 candidate ORFs and found 484 passing peptides mapping to 183 ncORFs, about 2.5% of the catalogue.
The low percentage is not automatically evidence that the other 97.5% are unreal. Small proteins create a geometry problem. Trypsin cuts after lysine or arginine. A tiny protein may generate no peptide long and unique enough to satisfy standard identification rules. The paper tested a manually curated set of 36 known proteins shorter than 50 amino acids. Only two, or 5.6%, met the conventional HUPO-HPP verification benchmark.
A standard designed to prevent false protein calls can therefore be excellent at preventing false calls and systematically poor at seeing the smallest true proteins.
Most of this resource was HLA class I. HLA-I molecules display processed peptides derived largely from inside cells at the cell surface; T cells recognize peptide-HLA complexes, not intact source proteins. HLA-II samples a different processing route that can include endosomal and extracellular material.
This is a different and selective sampling system from tryptic proteomics. Antigen processing and the HLA alleles represented in the datasets determine what becomes visible. A short or unstable translation product can therefore appear in HLA data even when conventional proteomics never detects it.
The TransCODE HLA build contained about 240 million MS/MS spectra. It yielded 3,116 peptides mapping to 1,785 of the 7,264 ncORFs, or 24.6%. Of those peptides, 2,937 were presented by HLA class I alone.
That tenfold difference between conventional and HLA-associated detection is the first major reveal. It is also a detection rate within this resource, not an unbiased census of translation.
Under the study’s assignment framework, an HLA peptide supports a candidate translation product from its mapped ORF. It does not establish that the intact product is abundant, stable, beneficial, harmful, or functional in normal physiology. Some displayed material may be defective, transient, stress-associated, or cancer-specific.
Conventional protein-coding analyses often ask whether amino-acid sequence is conserved across species. Very short and evolutionarily young ORFs are difficult to evaluate this way. The TransCODE group developed ORBL, ORF Relative Branch Length, to ask whether a start codon, open reading frame, and stop codon remain intact across a phylogenetic tree.
The refined ORBLq score compares each candidate with untranslated ORFs of similar length and type. In the study, 2,211 of 7,264 candidates, or 30.4%, had ORBLq above 0.9, compared with 10% expected from the matched null. Upstream ORFs were particularly enriched. Yet only 143 candidates, or 2.0%, had a positive amino-acid-level PhyloCSF score above 10.
Those measures ask different questions. ORBL tests preservation of the reading-frame architecture. PhyloCSF tests whether substitutions resemble protein-coding evolution. Neither reveals whether the product folds, functions, or matters in normal physiology.
CRISPR screens can target a candidate sequence and ask whether cells lose fitness. Rescue experiments can then restore the coding sequence and test whether the phenotype returns toward normal. These experiments bring the field closer to function, but context still controls the claim.
A knockout effect in a cancer cell is a cancer-cell result. It may expose a tumor dependency. It may reflect the host transcript, the DNA locus, or a neighbouring gene unless the design separates them. It does not automatically establish a normal physiological role.
c10riboseqorf92 will eventually clear part of the cancer-cell fitness rail while leaving normal physiology unresolved. That mismatch is exactly what the new category was built to preserve.
Our peptide evidence tier list makes the broader rule explicit: the strongest sentence an article can write is limited by the weakest evidentiary link it needs.
Ribo-seq says the cell read a sequence. Mass spectrometry says a fragment existed. Neither result says what the molecule does.
The study’s scale is easy to describe and easy to misuse.
The researchers started with 7,264 GENCODE-supported ncORFs. The non-HLA build represented hundreds of ProteomeXchange datasets and roughly 3.5 billion spectra. The HLA build added roughly 240 million spectra. The paper’s abstract compresses the combined resource into 95,520 proteomics experiments; the underlying PeptideAtlas pages distinguish datasets, experiments, runs, spectra, peptide-spectrum matches, and peptides. Those units should not be swapped merely because one produces a larger number.
Across the PeptideAtlas search workflow, the researchers used a decoy-estimated protein-level false-discovery rate below 0.1%. The conventional HUPO-HPP benchmark additionally required two distinct peptides, each at least nine amino acids, mapping uniquely and together spanning at least 18 amino acids of the ORF.
Then the consortium manually inspected the candidates.
Among 42 ncORFs initially supported by two unique conventional peptides, 30 survived manual validation. Among 141 supported by one peptide, 36 survived. Investigators also matched 29 of 30 nominated peptide-spectrum matches to synthetic reference spectra. They used targeted proteomics for selected candidates.
This is not the story of a permissive database accepting every match. It is the story of an enormous search producing a much smaller set when human scrutiny enters.
The HLA branch required its own checks. Investigators manually inspected 859 HLA-I spectra and 691 corresponding Ribo-seq profiles among candidates with stronger peptide support. They validated the Ribo-seq signal in 613 of 691 cases, or 88.7%. Candidates appearing in multiple Ribo-seq studies validated more often than single-study candidates: 96.1% versus 76.1%.
The consortium also compared modelled HLA binding with observed peptide presentation. Among runs with known HLA typing, 88.5% had more than 70% of their peptides classified as binders, and the reported model-to-detection concordance reached 94.8%.
The checks support different links in the inference chain. Manual MS review tests the spectrum and peptide assignment. Ribo-seq review tests translation of a plausible source ORF. HLA-binding models test compatibility with the reported HLA type. Together they strengthen peptide-to-ORF inference, not stability or function.
The final annotation framework divided candidates into evidence tiers. Tier 1A required enough conventional proteomics and Ribo-seq support to satisfy HUPO-HPP verification. Manual inspection initially produced 20 tier-1A candidates. Further scrutiny removed five because of pseudogenic sequences, a reference-genome assembly error, or insufficient Ribo-seq support. Fifteen remained.
By the paper’s publication, GENCODE had annotated three of the tier-1A ncORFs as protein-coding genes: candidates associated with CYP27B1, ERVH48-1, and PIDD1. The paper also describes a separate upstream ORF in GMCL1 whose case combined a loss-of-fitness phenotype, high ORBLq, positive PhyloCSF, translation in other mammals, and HLA-I and HLA-II peptides. GENCODE now annotates it as a protein-coding gene even though conventional tryptic peptides were absent.
That example matters because it shows the framework doing more than counting spectra. Different lines of evidence can compensate for a technology’s blind spot. But compensation is not relaxation. It requires a coherent case.
The study therefore did two things at once:
it expanded the searchable proteome;
it made the filter explicit.
The second achievement is less dramatic in a press release. It is more important for science.
The institutional announcement said researchers had discovered “1,700 new proteins.” That shorthand was understandable. It was also too broad.
The exact result was peptide evidence mapping to 1,785 ncORFs in the HLA build. The paper did not establish 1,785 new stable, functional human proteins. It did not add 1,785 conventional protein-coding genes to the reference catalogue. It did not identify 1,785 drug targets. It did not show that 1,785 molecules operate in normal physiology.
Those claims sit on different evidence rails:
Catalogue: Is there a candidate ORF worth testing?
Translation: Do ribosome data support that the sequence is read?
Detection: Does a peptide map to the candidate?
Persistence: Does an intact product accumulate?
Causality: Does the peptide product, or the act of translating its ORF, change biology?
Mechanism: Is the causal route known?
Physiology: Does the result hold outside an artificial or disease context?
Medical actionability: Can the biology be modulated safely and usefully in people?
The TransCODE candidates occupy different positions across those rails. Calling all 1,785 “new proteins” compresses every distinction into one familiar noun.
That compression matters. A protein sounds like an object with a structure, a job, and a place in the cell. “Translation product detected through an HLA-bound peptide” sounds like a method section. The second phrase is less memorable and more accurate.
If 1,785 new proteins have been established, then structure modelling, disease mapping, drug screening, and development programs appear ready to begin at scale. If 1,785 ORFs have peptide evidence, then the next job is triage: which products persist, which are regulatory, which are cancer-specific, which are degraded, which act through their peptide sequence, and which survive independent replication?
The discovery is not a finished protein catalogue. It is a disciplined queue of unfinished questions.
This is familiar territory for peptide science. The word peptide describes chemistry, not an evidence standard. A medicine such as semaglutide has a manufactured product, pharmacology, human dose-ranging data, randomized outcomes, safety evidence, and regulatory status. A newly detected peptidein acquires none of those merely because both are built from amino acids.
Scientific categories usually sound more settled than the evidence that created them. Peptideinwas designed to do the opposite.
The paper applies peptidein to an ORF with endogenous peptide evidence supporting a translation product, but not enough evidence for conventional protein-coding-gene status. The designation considers conventional peptide support, cancer-only detection, evolutionary constraint, and mechanistic evidence of function.
The consortium produced 121 initial peptidein annotations, not 1,785. HLA detection alone did not put every candidate into that curated set; the paper selected evidence-rich tiers for manual validation. Twenty-one of 39 tier-2A candidates, for example, combined one validated tryptic peptide with Ribo-seq and additional HLA peptides.
That middle category protects both discovery and scepticism. Without it, reproducible molecules can remain stranded under “noncoding.” With an undisciplined protein label, weak candidates can acquire functions, disease associations, and therapeutic narratives they have not earned.
The category can move when better proteomics, normal-tissue evidence, structural work, genetics, or carefully controlled perturbations strengthen the case. It can also remain still. Translation is not obliged to produce purpose.
Peptidein is therefore a record of what the evidence has earned: what was detected, in which context, what remains missing, and what result would promote or reject the candidate.
The study’s most revealing case sits inside a transcript whose name announces the old assumption: OLMALINC, also called LINC00263, is annotated as a long noncoding RNA.
That transcript contains six ncORFs recognized in the GENCODE catalogue. One of them, c10riboseqorf92, encodes a 123-amino-acid product. In the consortium’s screens, it behaved unlike the other five.
But a phenotype at a complicated locus does not identify its cause. A CRISPR cut can damage DNA regulation. Destroying the transcript can remove an RNA function. Six translated frames create six possible peptide products. The experiment had to separate the coding sequence from the host transcript.
Researchers began with CRISPR-Cas9 loss-of-function screens targeting more than 2,000 ncORFs across eight human cell lines. Across the catalogue, 51 ncORFs showed a pan-essential knockout signature. Six had enough HLA or other evidence to qualify as candidate peptideins or protein-coding genes. c10riboseqorf92 was one of them.
At OLMALINC, three observations narrowed the cause. Only c10riboseqorf92 among the transcript’s six ORFs produced the pan-essential signal. CRISPR-Cas13 degradation of the RNA supported a dependency. Re-expressing the c10riboseqorf92 coding sequence also rescued the viability effect caused by OLMALINC knockdown. That rescue supports an ORF-dependent effect consistent with the peptide product, but the reported design does not by itself exclude every effect of the rescue RNA.
The group then analysed pooled DepMap CRISPR screens covering 485 cancer cell lines. c10riboseqorf92 knockout was associated with loss of viability in 415 models, or 85.6%, consistent with pan-essentiality in that resource. Its dependency pattern correlated with canonical genes involved in mitosis and DNA-damage regulation. Transcriptome and single-cell analyses implicated chromosome-related, metabolic, hypoxic, and stress-response programs.
Here the inference should stop. The data support an ORF-dependent contribution to cancer-cell fitness in the tested models. They do not isolate the peptide’s precise biochemical mechanism, establish its necessity in an organism, show a role in healthy tissue, or reveal a therapeutic window.
The authors therefore kept c10riboseqorf92 classified as a peptidein. A less careful account would call it a newly discovered essential human protein and a pan-cancer target. The consortium showed coding-sequence rescue, quantified a striking dependency, mapped associated programs, and stopped one category short.
A cancer cell can depend on a molecule that normal biology has not yet explained.
That distinction protects the next experiment. If healthy cells use the product for chromosome maintenance, a drug against it could be broadly toxic. If tumors become unusually dependent on it, a therapeutic window may exist. If the phenotype is peculiar to transformed cell lines, the target may disappear in an organism. Each possibility demands a different study.
HLA presentation can make a noncanonical product medically interesting before conventional protein annotation catches up.
Cancer cells alter transcription, translation, splicing, RNA surveillance, and protein degradation. HLA molecules can display fragments from that altered intracellular world. A T cell may therefore encounter a noncanonical translation product even when conventional proteomics cannot find the intact molecule.
In 2022, Ouspenskaia and colleagues reported 3,555 translated noncanonical ORFs represented in the MHC-I immunopeptidome, including tumor-specific and somatically altered candidates across melanoma, chronic lymphocytic leukaemia, and glioblastoma datasets. The result expanded the searchable antigen space. It did not show that thousands of candidates could be treated successfully.
In 2024, researchers studying childhood medulloblastoma reported that translation from noncanonical ORFs could support cancer-cell survival. That work helps explain why the 2026 consortium integrated earlier functional screens instead of treating peptide presentation as the final word.
In 2025, a Science study found pancreatic-cancer-restricted cryptic antigens recognized by T cells. Recognition is necessary for some immunotherapies. It is not a safe, durable response in a person.
The asymmetry is real: annotation projects need a definition of protein; T cells do not. A displayed peptide can matter whether its source is a stable enzyme, a short-lived product, defective translation, or a cancer-specific event.
Turning that signal into a therapy is a different funnel. Each candidate needs source-ORF validation, a tumor-versus-normal presentation atlas, HLA restriction and population coverage, immune recognition, a workable modality and delivery route, and evidence that useful antitumor activity can be separated from toxicity.
A fragment on a tumor cell is the beginning of that program, not the end.
Commercial attention, not validation. ProFound Therapeutics announced an expanded-proteome discovery effort in 2022, a research agreement involving Pfizer in 2024, and a Novartis cardiovascular collaboration in 2025. Those company announcements show money entering the search space. They do not validate the TransCODE candidates or establish benefit in people.
An expanded proteome adds translated frames to variant interpretation. It does not turn those frames into personal forecasts.
An upstream ORF can influence how much protein is translated from the main coding sequence. A variant can create a new start site, remove a stop, change the distance between an upstream ORF and the main gene, or alter the peptide sequence itself. In 2020, a study of 15,708 human genomes found strong negative selection against classes of variants expected to create or disrupt upstream ORFs, especially when a new frame overlapped the main coding sequence. The authors described candidate disease mechanisms involving upstream-ORF disruption.
The important phrase is candidate mechanism.
A variant can change the peptide product, change ribosome traffic and the main protein’s abundance, or alter RNA structure and stability. The same coordinate can support several plausible routes. Peptide evidence, evolutionary constraint, and a perturbation phenotype can justify a sharper experiment; they do not identify which route operates in one person.
Our guide to peptide genetics draws the same boundary: pathway relevance is not an individual treatment forecast. The dark proteome adds possible biology to the map. It does not remove the need for phenotype, replication, and evidence in people.
The paper’s strongest objection is built into its method. Each instrument sees a biased slice of biology, and every relaxed threshold trades false negatives for false positives.
The paper reports that 2.36 billion of 3.53 billion non-HLA MS2 spectra, or 66.9%, came from cancer tissue or cancer cell lines. That matters because cancer and immortalization alter translation. Candidates supported only in those contexts may be real disease products without being components of normal physiology.
The consortium kept ncORFs in STK11, ZNF219, and CIRBP under the peptidein label despite multiple tryptic peptides because all supporting material came from cancer samples or immortalized lines. That is a strength of the framework. It is also evidence that sample bias reaches deep into the catalogue.
HUPO-HPP’s two-peptide rule protects the reference proteome from weak identifications. It also becomes physically difficult for very short sequences. The paper notes that 2,059 of the 7,264 candidates, or 28.3%, are shorter than 25 amino acids. A protein shorter than 18 amino acids cannot satisfy a rule requiring two peptides spanning at least 18 residues. A protein only slightly longer may still produce no useful tryptic pair.
Relaxing the standard creates the opposite risk. The smaller the candidate class and the larger the search space, the more damaging misassigned spectra become. Noncanonical databases contain many similar, overlapping, and short sequences. A global false-discovery rate can look acceptable while the error rate inside a rare subclass is much worse.
An independent 2026 analysis led by Aaron Wacholder examined the reproducibility of noncanonical proteogenomic claims. The Nature Communications paper reported that 96% of roughly 10,000 collected small-ORF peptides appeared in only one study. High-quality peptide-spectrum matches were far more common in HLA data than in non-HLA data. The authors argued for stronger class-specific error controls, synthetic-peptide comparisons, and orthogonal validation.
That critique does not erase TransCODE. It explains why TransCODE’s manual inspection, synthetic standards, Ribo-seq review, and conservative nomenclature are necessary.
The HLA system samples intracellular degradation. Some presented peptides may come from stable proteins. Others may come from newly synthesized defective products, rapidly degraded polypeptides, or stress-associated translation.
A 2023 Nature study on noncoding translation mitigation described cellular mechanisms that limit products from pervasive noncoding translation, including degradation associated with BAG6. A molecule can be translated, fragmented, and presented precisely because the cell does not want it to persist.
For immunology, that fragment can still matter. For conventional protein annotation, its instability matters too.
The immune system can prove a sequence was translated without proving that the product deserved a job description.
The consortium manually inspected spectra and ribosome profiles because automated pipelines could not resolve every ambiguity. The paper identifies that labour as a limitation. Thousands of future candidates across tissues, conditions, and disease states cannot all depend on a small group of experts examining plots by hand.
Deep learning may improve spectrum scoring, retention-time estimation, ion-mobility estimation, translation-site detection, and structural triage. But a model score can magnify whatever the training data reward. The paper therefore lists deep learning as a research question, not an annotation oracle.
This is an important limit for AI-heavy drug discovery. Our profile of AI-assisted biotechnologydescribes how computational tools can widen the space a small team explores. The dark proteome is an ideal use case for prioritization. It is also an ideal place to confuse ranking with proof.
The non-HLA effort focused on data-dependent acquisition. The authors note that data-independent acquisition paired with targeted parallel-reaction monitoring may increase sensitivity in specific contexts. Different proteases, enriched size fractions, post-translational-modification searches, and synthetic standards can reveal candidates missed by routine tryptic workflows.
A negative result therefore has at least two meanings: the molecule may not be present, or the method may be poorly suited to finding it.
A positive result has at least two meanings too: a matching peptide may have existed, or the assignment may be wrong.
The work is to narrow both ambiguities without pretending either can be wished away.
ORBL asks whether the reading-frame structure is preserved. That can capture regulatory upstream ORFs whose translation matters even if their amino-acid sequences do not. It can also be affected by other conserved sequence features, alignment errors, limited power in recent primate-specific ORFs, uncertain start sites, and the possibility that supposedly untranslated control ORFs are translated.
The authors put detailed ORBL limitations in the supplementary results. That placement should not make them optional reading. An ORBLq score can prioritize a candidate. It cannot tell researchers whether the product folds, binds a partner, acts in normal physiology, or becomes a useful drug target.
The evidence does not support “nothing was found.” It supports a narrower verdict:
the HLA result is stronger evidence of widespread noncanonical translation than of 1,785 conventional proteins;
the curated peptidein set is more defensible than the press-release number;
functional cases matter more than the raw search total;
replication across normal tissues will decide how much of the map becomes permanent.
The consortium’s seven closing questions reduce to four promotion tests:
Detection: Can standards built for larger proteins be adapted without loosening false-positive control? Synthetic standards, alternative proteases, targeted MS, intact-protein methods, and orthogonal translation evidence may need to work together.
Context: Should HLA-presented or cancer-only products count as proteins when they may be unstable, disease-specific, or absent from normal physiology?
Causality: Can experiments separate effects of the DNA locus, host RNA, act of translation, and peptide product, then show a mechanism in an organism?
Governance: What evidence promotes, retains, or rejects a peptidein, and how can machine learning scale review without hiding the evidence chain behind a score?
The answer must preserve the chain: sequence, translation evidence, peptide evidence, sample context, validation status, functional experiment, and unresolved contradiction.
The consortium has made the queue public through the Human HLA and non-HLA PeptideAtlas builds. A public queue can be challenged. A private score cannot.
Van Heesch’s “something new” was not 1,700 finished proteins. It was a way to name translation products that the instruments could detect before biology could explain them.
The consortium did not reopen the Human Genome Project because the original sequence was missing thousands of secret genes. It reopened the annotation contract. The genome had been mapped as letters. The proteome had been curated as a set of approved readings. New instruments showed that cells read more of the text than the official catalogue captured.
Some of those readings may earn protein-coding status.
Some may become cancer antigens.
Some may explain how upstream variants regulate familiar genes.
Some may remain transient products.
Some may disappear when better controls are applied.
Peptidein makes room for every outcome without pretending the verdict arrived with the discovery.
The human proteome was never finished. The catalogue had run out of names for what its instruments could see.
Peptidein is an evidence category, not a product category. For peptide medicines already studied in people, start with our evidence tier list. A newly detected translation product does not become a treatment by entering a catalogue.
This article is for educational and informational purposes only. It does not provide medical advice, diagnosis, treatment guidance, genetic interpretation, or a recommendation to use any peptide, product, test, or investigational therapy. Peptideins and other noncanonical ORF products discussed here are research findings; the cited studies did not evaluate them as medicines. Cancer-cell dependencies, HLA-presented peptides, genetic associations, and preclinical mechanisms do not establish safety or efficacy in humans. Consult a qualified healthcare professional for personal medical decisions.
Not in the conventional sense. The HLA PeptideAtlas analysis detected 3,116 peptides mapping to 1,785 of 7,264 candidate ncORFs. That supports widespread synthesis of noncanonical translation products. It does not establish 1,785 stable, functional proteins or 1,785 new protein-coding genes. The study produced 15 final tier-1A candidates after manual scrutiny and 121 initial peptidein annotations. Those numbers answer different questions and should not be combined.
The human proteome is the collection of proteins produced from the human genome, including different forms that can arise from alternative transcripts, processing, and modification. Reference projects do not merely list sequences; they evaluate whether a gene should be considered protein-coding and what evidence supports each protein. The 2026 TransCODE study argues that thousands of noncanonical reading frames produce detectable translation products outside the traditional catalogue, while only a smaller subset currently meets conventional annotation standards.
A noncanonical open reading frame, or ncORF, is a potentially translated sequence outside the main annotated coding sequence used to define a conventional protein. It may sit upstream or downstream of the main coding region, overlap it in another frame, occur inside an intron, or appear within a transcript labelled noncoding. “Noncanonical” describes its relationship to the reference annotation. It does not establish whether the sequence is functional, stable, or medically important.
Microprotein is an umbrella term for a very small protein or protein-like translation product, often encoded by a short ORF. The literature also uses micropeptide, small ORF-encoded peptide, and related terms. There is no single length threshold applied consistently across every field. Some microproteins have well-supported structures and functions. Others are detected products whose stability or role remains uncertain. The TransCODE framework uses evidence, not size alone, to distinguish stronger protein candidates from peptideins.
The TransCODE consortium applies peptidein to an ORF with endogenous peptide evidence supporting a translation product, but not enough evidence for conventional protein-coding-gene status. A peptidein may lack evidence of normal physiological function, may be supported mainly by cancer samples, or may have too little conventional proteomics evidence. The designation is provisional and can change when new data strengthen or weaken the case.
HLA molecules display peptide fragments generated inside cells, including fragments from short-lived or low-abundance products. Conventional proteomics usually depends on enzymatic digestion, often with trypsin, and requires peptides long and unique enough for confident matching. Very small proteins may produce no suitable tryptic peptides. HLA immunopeptidomics can therefore reveal translation products that conventional workflows miss, although presentation alone does not prove that the intact product is stable or functional.
Some noncanonical translation products may become cancer antigens or dependencies, and primary studies have shown T-cell recognition or cancer-cell fitness effects for selected candidates. That is discovery-stage evidence. A viable treatment additionally requires tumor specificity, reproducible presentation, population-relevant HLA coverage, a workable therapeutic modality, safety, dose finding, and benefit in people. No peptidein becomes a treatment merely by appearing in the catalogue.
Not from the TransCODE catalogue alone. A variant may affect an ncORF’s peptide sequence, alter translation of a neighbouring canonical gene, change RNA behavior, or do nothing measurable. Establishing personal relevance requires validated variant interpretation, phenotype evidence, population data, and often functional experiments. The discovery of a translated region raises a research question; it does not produce an individual diagnosis or peptide recommendation.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.