RSS Amplifier

Bits in Bio · Sep 26, 2025

The Bits In Bio Letter - September 25th 2025

0
Sign in to vote or save

Vincent Alessi, Erle Holgersen, Charlene Her · Bits in Bio

  • First AI-generated genomes create viable bacteria-killing phages

  • Lilly opens $1B research vault to its portfolio of biotech startups through new AI platform - TuneLabs

  • LB Pharma breaks the IPO drought with an increased $285M Nasdaq debut


    Not yet a member of our super awesome slack community of ~10,000? Join HERE
    🤗

I'm excited to share the second episode of the Bits In Bio Podcast Series!

Join your hosts Robbie Matthews and Vincent Alessi this month for deep dive into Science x AI with the builders-turned-executives leading the charge.

Join us this month to hear from Apheris CEO Robin Röhm, who’s built the federated learning infrastructure that’s letting the landmark OpenFold Consortium tap into billions of data points from big pharma’s jealously-guarded structural information.

Find it on Apple Podcasts and on Spotify


Join us next episode to hear from Armand B. Cognetta III, PhD to hear about his exciting journey from being the 1st intern at the legendary Alnylam pharmaceuticals to solo-founding General Proximity as one of the first biotechs to come out of Y Combinator (and the hilarious story about how that came about!).

Generative AI crosses the Rubicon from proteins to complete functional viral genomes

Led by Brian He, Arc Institute and Stanford researchers have achieved the first generative design of viable bacteriophage genomes using AI, with 16 functional phages emerging from 302 AI-designed genomes tested (~5.3% success rate) that successfully replicated and killed E. coli bacteria. The team fine-tuned Evo 1 and Evo 2 genome foundation models (trained on 2.7 million prokaryotic and phage genomes covering 300 billion nucleotides) to design variants of the historic ΦX174 phage, with several viable designs sharing less than 95% similarity to known phages—crossing the new species threshold. Most remarkably, AI-generated phage cocktails overcame ΦX174-resistance in three E. coli strains within 1-5 passages through mosaic genomes combining elements from different AI designs, while Evo-Φ36 successfully incorporated a distant G4 phage gene despite known incompatibility issues. Lead researcher Hie notes that “most biological functions are not achieved by any single gene,” positioning this breakthrough as essential groundwork for engineering complex functions beyond single-gene modifications. While J. Craig Venter dismisses AI methods as “just a faster version of trial-and-error,” the ability to generate functional genomes computationally—with all models freely available on HuggingFace—represents a fundamental shift in synthetic biology from tweaking nature to designing it, even if bacterial genomes 1000x larger present exponentially greater complexity ahead.

Big Pharma’s (semi) open data era begins with Lilly’s billion-dollar gambit

In an unprecedented turn, Eli Lilly launched TuneLab, an AI/ML platform offering biotech companies access to drug discovery models trained on over $1 billion worth of proprietary research data—a characteristically bold move to position itself as partner-of-choice rather than just acquirer. The federated learning system, powered by Rhino Federated Computing’s NVIDIA infrastructure, provides 18 initial models (12 small molecule, 6 antibody) while keeping everyone’s data separate and protected, solving the age-old problem of small biotechs lacking high-quality training data at scale. Limited to participants of Lilly’s select broader Catalyze360 initiative (which includes Gateway Labs facilities and Lilly Ventures capital), TuneLab represents a calculated response to the industry’s AI arms race, with about a dozen startups already enrolled including Insitro, Circle Pharma, and Superluminal Medicines. The platform’s continuous improvement model—where each partner’s contributions enhance the ecosystem—creates a network effect that could compress decades of learning into instantly accessible intelligence. While critics might argue this is simply Lilly crowdsourcing R&D, the potential 30% reduction in preclinical costs through better compound prioritization could prove transformative for cash-strapped biotechs racing toward IND.

Biotech IPO window cracks open as LB Pharmaceuticals prices an upsized $285M deal

In a welcome reprieve for public market biotechs, LB Pharmaceuticals, a CNS-focused biotech developing LB‑102 for acute schizophrenia, upsized its Nasdaq IPO to 19 million shares at $15, raising $285 million in gross proceeds—the first biotech listing in months. The company had initially targeted 16.7 million shares and $228.5 million; an underwriter greenshoe could add another 2.85 million shares (about $42.7 million). Proceeds are earmarked primarily for a Phase 3 study in acute schizophrenia (roughly $133 million) and a Phase 2 in bipolar depression (about $25 million), with the rest for general corporate needs. LB‑102 is a modified version of amisulpride, a dopamine D2/D3 antagonist used ex‑US, and a January Phase 2 readout showed statistically significant improvement on PANSS at four weeks. Shares trade under the ticker LBRX. The upsize follows cash‑conserving moves—including a May restructuring and a leadership change to CEO Heather Turner—and hints that investors remain open to late‑stage neuro bets even in a cautious tape. Whether this reception broadens to others will depend on clinical catalysts and how this debut trades in the coming weeks.

Machine learning meets DNA damage in promising early-stage trial targeting Glioblastoma

Lantern Pharma’s LP-184, guided by its RADR® AI platform with 200 billion+ oncology data points, met all primary endpoints in a 63-patient Phase 1a trial with favorable safety and early antitumor activity in advanced solid tumors including recurrent glioblastoma. The synthetic lethal compound, targeting DNA damage repair (DDR) deficiencies through PTGR1 activation specifically in cancer cells, showed clinical benefit in 4 of 16 recurrent GBM patients previously treated with temozolomide/lomustine/radiation—particularly impressive given this population’s dismal prognosis. Marked tumor reductions occurred in patients with CHK2, ATM, BRCA1, and STK11/KEAP1 mutations across colon cancer, thymic carcinoma, GIST, and NSCLC, with one NSCLC patient achieving nearly 2 years of clinical benefit. With Fast Track Designations for GBM and TNBC plus multiple Orphan Drug Designations, LP-184 targets an $11-13 billion global market opportunity while validating Lantern’s approach of developing programs for $1.0-2.5 million in 2-3 years versus traditional timelines. The company’s subsidiary Starlight Therapeutics is advancing a GBM-specific Phase 1b/2a trial (STAR-001), suggesting LP-184 could become a meaningful option in the notoriously difficult CNS cancer space.

Domain-specific training trumps scale in biological reasoning benchmark

Paris-based OG biotech AI, Owkin, has developed OwkinZero models—specialized 8-32B parameter LLMs that substantially outperform larger commercial models on biological reasoning tasks through their novel Reinforcement Learning from Verifiable Rewards strategy, according to their arXiv preprint. Testing across eight benchmark datasets with over 300,000 verifiable Q&A pairs covering target druggability, modality suitability, and drug perturbation effects, these smaller models consistently beat state-of-the-art LLMs despite having 10x fewer parameters. The key insight: specialist models trained on single biological tasks show remarkable cross-task generalization, outperforming base models on completely unseen challenges—with comprehensive models trained on dataset mixtures achieving even broader improvements. This challenges the industry assumption that bigger is always better, demonstrating that targeted training with verifiable biological rewards rather than human feedback produces superior performance on critical drug discovery bottlenecks. While everyone’s racing to build the largest models, Owkin’s results suggest the path to practical AI-driven drug discovery might lie in smaller, specialized models that actually understand biology rather than just pattern-matching at scale.

Google’s quantum computing offshoot lends researchers raw fuel for next-gen models

SandboxAQ, the characteristically over-funded $950M-funded Alphabet spinout, launched SAIR (Structurally Augmented IC50 Repository)—the largest-ever dataset of protein-ligand pairs with 5.2 million synthetic 3D molecular structures across 1+ million systems. Uncharacteristically the data is freely available for both commercial and non-commercial use. The physics-grounded dataset, generated through the company’s Large Quantitative Models (LQMs) running on NVIDIA DGX Cloud with 90%+ GPU utilization over 20 days, delivers AI model predictions 1,000+ times faster than traditional physics-based methods. Six pharma companies signed up within 48 hours of launch, reflecting the industry’s hunger for experimentally-validated training data rather than purely AI-generated predictions. The timing couldn’t be more interesting, aligning perfectly with FDA’s push toward computational validation methods to replace animal testing. While the promise of eliminating wet lab validation phases entirely remains ambitious, SandboxAQ’s approach—combining quantum mechanics with machine learning—represents a fundamentally different path than pure AI competitors, potentially setting a new pace for data-driven drug discovery.

Full-Length mRNA Design with Generative AI
While COVID-19 RNA vaccines demonstrated the therapeutic potential of mRNA, designing stable and deliverable RNA drugs across diverse indications remains challenging. This is largely because the sequence and structural features governing RNA stability and translational efficiency are still not fully understood. Cambridge-based Raina recently introduced its generative AI platform, GEMORNA, which designs full-length RNA sequences—including coding regions and untranslated regions—with optimized stability and expression. GEMORNA, a transformer-based model, was shown to increase protein expression levels by 10–45 fold. Beyond linear RNAs, GEMORNA can also design circular RNAs, whose closed structures confer greater stability than their linear counterparts. Notably, GEMORNA-designed circular RNAs, when combined with CAR-T therapy, significantly boosted anti-tumor cytotoxicity. Find the full report at Science.

Terabyte-scale spatial atlas feeds data-hungry biological foundation models

Bay Area native bioinformatics consulting company LatchBio has released a potent 25 million cell human spatial transcriptomics atlas spanning 11 spatial technologies, 45 tissue types, and 63 diseases—the largest open-source human spatial atlas to date, immediately available through their public portal. The company’s agentic spatial curation toolkit improves per-dataset curation times by 40x through human-in-the-loop AI frameworks, solving the bottleneck where processing raw data into publication figures previously took months but now takes days. Supporting everything from 10X Genomics Visium HD to Vizgen MERSCOPE to STOmics Stereo-seq, the standardized H5AD format with structured ontologies enables AI labs to model phenomena beyond binding and structure—critical for next-generation biological foundation models. CTO Kenny Workman notes that “progress in engineering biology increasingly depends on data-hungry statistical models,” with partners like AtlasXOmics and Takara Bio already leveraging the white-labeled platform for custom portals. While spatial transcriptomics has promised to revolutionize our understanding of tissue architecture and disease, the lack of comprehensive, standardized datasets has limited AI applications—a gap this atlas definitively closes.

Language models learn biology, deliver dramatically better reprogramming of select proteins

OpenAI has deepened its standing partnership with Retro Biosciences (where Sam Altman personally invested $180 million - conflict of interest, anyone?) to develop GPT-4b micro, a biology-specialized variant that achieved over 50x improvements in stem cell reprogramming efficiency through AI-redesigned Yamanaka factors. The model, trained using Reinforcement Learning from Verifiable Rewards on protein sequences, biological text, and tokenized 3D structures, suggested modifications to up to one-third of amino acids in proteins—with over 30% of AI-generated RetroSOX variants and nearly 50% of RetroKLF variants outperforming their natural or manually-engineered counterparts. These AI-designed factors showed 85% success in activating endogenous pluripotency genes within 12 days plus enhanced DNA damage repair capabilities, addressing the longstanding inefficiency of cellular reprogramming (typically less than 1% success rate). Joe Betts-Lacroix, Retro’s CEO, noted they “threw this model into the lab immediately and got real-world results,” with potential applications spanning regenerative medicine and Retro’s ambitious goal of extending human lifespan by 10 years. While OpenAI researcher John Hallman admits the proteins “seem better than what scientists were able to produce by themselves,” the real test will be whether GPT-4b micro’s cross-task generalization translates to therapeutic development beyond proof-of-concept reprogramming experiments.

Robotic experimentation seeks to close the loop on scientific AI’s data crisis

The San Francisco-based lab automation startup known as Medra has launched its Continuous Science Platform, combining “Physical AI” robots that automate 70% of existing basic lab instruments with “Scientific AI” reasoning models in a self-improving closed-loop system that addresses the fundamental data scarcity limiting scientific frontier models. CEO Michelle Lee notes that scientific models need 1,000X more training data to match multimodal reasoning capabilities—a gap Medra fills by capturing unprecedented “Infra-data” including pipette tip angles, reagent mixing timing, and granular experimental metadata never before available at scale. Already working with world’s largest biotech and pharma companies plus partners like Addition Therapeutics (RNA transfection) and Lila Sciences (protein characterization), the platform enables 24/7 experimentation without human fatigue. The approach tackles the stark reality that AlphaFold2 trained on 50 years of structural biology data represents just 0.3% of what multimodal models consume, suggesting scientific AI’s promise remains bottlenecked by data generation rather than algorithmic innovation. While the vision of autonomous scientific discovery has been pitched before, Medra’s focus on the unglamorous but critical data generation problem—rather than just fancy AI models—might actually deliver the infrastructure needed for AI-driven therapeutic breakthroughs.

Natural products AI startup hits unicorn status with back-to-back raises while first patient dosed

Colorado-based Enveda Biosciences raised an oversubscribed $150 million Series D led by Premji Invest, achieving unicorn status (>$1B valuation) just months after its $150M Series C—timing that coincides perfectly with Dr. Vince Clinical Research dosing the first patient in ENV-294’s Phase 1b trial for moderate-to-severe atopic dermatitis. The Boulder company’s AI-powered “library of life” analyzes 10,000+ molecules simultaneously from nature’s 99% unexplored chemical diversity, producing 12+ development candidates including ENV-294, a first-in-class compound designed to deliver JAK-inhibitor-like efficacy with IL-4/IL-13-like safety—addressing unmet need in 26 million Americans who face either injectable biologics or oral JAK inhibitors with black box warnings. With former Pfizer CSO Mikael Dolsten joining the board (bringing experience from 150+ clinical programs), ~250 employees across Boulder and Hyderabad, and $517M raised since 2019, Enveda’s platform claims to deliver candidates 4x faster and 10x cheaper than industry standard—a promise now being tested as ENV-294 advances in real patients after showing favorable Phase 1a safety with no dose-limiting toxicities.

D-Wave Quantum veterans pivot to molecular design, lands $349M of Merck’s attention

An 18-person Vancouver startup founded by former D-Wave Quantum AI researchers called Variational AI has entered a collaboration with Merck worth up to $349 million to apply its Enki™ generative AI platform against two undisclosed “challenging therapeutic targets.” The deal structure—upfront payment plus milestones with no royalties—validates the generative foundation model approach to drug discovery, particularly impressive given Variational AI operates without wet labs and competes against established players like Insilico Medicine and Exscientia. Merck will fine-tune Enki using its proprietary chemistry data, aiming to generate novel, selective, synthesizable lead-like structures in weeks rather than the traditional years-long timelines. The partnership reflects a broader industry trend where Big Pharma increasingly outsources AI innovation rather than building internally, with Variational AI CEO Handol Kim noting this positions Canada competitively in AI-driven life sciences despite lacking major pharma presence. While no AI-generated molecules have yet achieved regulatory approval, the potential to “redefine unit economics of drug discovery”—hundreds of thousands versus millions in costs—keeps attracting pharma checkbooks to these computational approaches.

Flagship’s latest venture attempts to automate the scientific method with AI

Cambridge-based Lila Sciences raised a $235 million Series A co-led by Braidwell and Collective Global, achieving a $1.2 billion valuation for its “scientific superintelligence” platform that combines AI, novel software, and custom hardware to autonomously execute the entire scientific method. The company, incubated within Flagship Pioneering for three years before emerging from stealth with a $200 million seed in March, has already driven thousands of discoveries across life science, chemistry, and materials science—including genetic medicine constructs outperforming commercial therapeutics and non-platinum hydrogen catalysts at a fraction of commercial costs. Led by CEO Geoffrey von Maltzahn (former Flagship partner) with George Church as Chief Scientist and former OpenAI researcher Kenneth Stanley joining the team, Lila’s AI agents complete in minutes or hours workflows taking human teams days or weeks. The platform’s real-world verification system distinguishes it from prediction-only approaches, with hundreds of thousands of AI-driven experiments already completed generating novel antibodies, green hydrogen catalysts, and carbon capture sorbents with better performance than leading products. As Collective Global’s Daniel Adamson describes it as “an IP factory par excellence,” Lila’s vision of AI running every step of scientific discovery at unprecedented scale positions it to tackle humanity’s greatest challenges—though whether autonomous science delivers transformative breakthroughs or just faster iterations remains to be proven.

GPCR targeting gets AI boost in Lilly’s post-GLP-1 diversification play

Eli Lilly signed a $1.3 billion collaboration with Superluminal Medicines to develop small molecule therapeutics targeting G protein-coupled receptors for cardiometabolic diseases and obesity, marking a strategic pivot as the pharma giant diversifies beyond GLP-1 drugs following mixed Phase III results for orforglipron. Superluminal’s AI-driven platform, integrating deep structural biology with proprietary pharmacokinetic/toxicology prediction tools, focuses on historically challenging GPCRs where functional selectivity and structural complexity have stymied traditional approaches—with 5 GPCR-targeting molecules already in development including an MC4R agonist for rare genetic obesity entering IND-enabling studies. The deal structure includes upfront and equity investments plus development milestones and tiered royalties, building on Lilly’s participation in Superluminal’s $120 million Series A just months earlier. Founded in 2022 and headquartered in Lilly’s own Gateway Labs, Superluminal represents the new breed of AI-native biotechs tackling “intractable” targets, with CEO Cony D’Cruz calling this “a defining moment” for the company. As the obesity drug market races toward $150 billion by early 2030s, Lilly’s bet on next-generation GPCR modulation through AI-designed compounds signals that even Ozempic’s success hasn’t exhausted pharma’s appetite for metabolic innovation.

Surface protein conformations unlock new cancer target universe

Immuto Scientific raised an oversubscribed $8 million Seed 2 round led by DYDX while announcing a drug discovery collaboration with Daiichi Sankyo focused on cancer-specific cell-surface targets identified through its AI-enabled structural surfaceomics platform. The company’s technology interrogates conformational differences across thousands of proteins simultaneously, revealing disease-specific surface protein conformations (SPCs) that represent a “hidden dimension of the cell surfaceome” invisible to conventional omics approaches. Unlike the industry’s obsession with new modalities, Immuto addresses what it sees as the real bottleneck—the shortage of truly novel, disease-specific targets that avoid healthy tissue overlap. The Daiichi Sankyo partnership validates this approach, with the pharma giant gaining options to license assets developed against solid tumor targets identified through Immuto’s structural epitope-mapping engine. With its lead program advancing through in vivo studies toward IND-enabling work, Immuto is creating and leveraging data that “no other group in the world has,” potentially representing a step-change in predictive protein structure for therapeutic discovery.

See our handy dandy Lu.ma event calendar HERE, please RSVP so folks can plan accordingly!

  • None up coming as folks get back into the swing of the academic year!

We’re partnering with Ginkgo Bioworks’ CRO group, the Datapoints Team (in collaboration with Hugging Face) to support the Antibody Developability Prediction Competition, a first-of-its-kind open benchmarking challenge in ML for antibody engineering!

With up to $60K in prizes/credits, we’ve spun up a dedicated channel, #competition-ginkgo-antibody-2025. Check it out!

Feedback: How is the Newsletter doing? We’re trying different formats/content. In case the hyperlink above didn’t get your attention, maybe a bright orange button will!

Start Survey

Volunteer: Want to get involved with Bits in Bio, meet new members across the community, and learn about the ecosystem? We are looking for volunteers to help us create great content and manage the community.

Volunteer

We want to deliver what matters most in Bio AI and would love your feedback on how we can do better. Please weigh in as anon here or DM me directly!

BiB Editor in Chief

No posts

Read the original on bitsinbio.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.