Historically, AI-driven drug discovery has been constrained by data scarcity. Earlier approaches in cellular imaging, also referred to as phenomics, relied heavily on supervised learning, which required costly manual labeling or many experimental replicates to mitigate biological and experimental batch effects. This paradigm has since evolved, pivoting toward data-rich, self-supervised foundation models that now underpin a new era of multimodal and generative biology.
The journey toward modern phenomics began in 2019 with Recursion’s RxRx1 Kaggle competition (Earnshaw, 2019). This supervised classification task, involving ~1,100 siRNA conditions across 125,510 images, in which perturbations were screened repeatedly in over 50 experimental batches, served as the initial proving ground for deep learning in cellular imaging. By challenging the ML community to identify perturbations across held-out experimental batches, RxRx1 established early benchmarks for handling systematic noise. This era proved that convolutional neural networks (CNNs) could decode cellular morphology (Sypetkowski 2023), but it also highlighted a critical bottleneck: the heavy reliance on labeled datasets from repeated experiments limited the scale at which models could learn.
In 2023, the field underwent a paradigm shift with the release of RxRx3 (Fay, 2023) and JUMP-CP (Chandrasekaran, 2023) and the transition to self-supervised learning (SSL) via Masked Autoencoders (MAE) and self-distillation (Cell-DINO, Moutakanni 2025). By moving away from supervised classification and instead training models to reconstruct masked regions of raw cellular images, Recursion broke the reliance on manual annotations. This allowed for the utilization of genome-scale datasets generated within standardized laboratory environments to mitigate batch effects. The resulting foundation models, including Phenom-1 (Kraus, 2024) and Phenom-2 (Kenyon-Dean, 2025), demonstrated an unprecedented ability to extract robust biological representations across the entire human genome.
To ensure these advancements were not restricted to enterprise-scale labs, Recursion released RxRx3-core (Kraus, 2025). By compressing 100 TB of raw imaging data into a manageable 18 GB format, and releasing it directly on public repositories like Hugging Face and Polaris Hub, RxRx3-core democratized access to high-content screening data. This accessibility has fueled a surge in academic innovation, evidenced by over 14,800 downloads and 32 unique research papers citing the RxRx3 ecosystem as of 2026.
Recently, foundation models are no longer just extracting features but are actively integrating modalities and generating new data. Multi-modal contrastive models like MolPhenix (Fradkin, 2024) have advanced the field by aligning molecular graph representations directly with visual phenotypic alterations in cells. Even more recently, the emergence of generative architectures has allowed researchers to visually simulate the morphological effects of unseen perturbations (Jones, 2026; Zhang, 2026). Morphological responses, which can be measured at scale with greater consistency than transcriptomics, offer a vital complementary modality for constructing predictive virtual cell models. These models represent a transition from representing experimental data to predictive simulation, enabling zero-shot evaluation of novel drug candidates.
The progression from supervised transfer learning to the generative, multimodal systems of today has transformed phenomics into a core pillar of drug discovery. However, the true legacy of the RxRx3 ecosystem extends beyond serving as a historical “proving ground.” As the field shifts toward predictive simulation and the pursuit of a Virtual Cell, structured, annotated phenomic datasets have become even more vital.
Generative models are often framed as a departure from traditional benchmarking, but they are fundamentally data-hungry and require high quality ground-truth datasets. In this new era, datasets like RxRx3 will be the core underlying datasets for generative biological models. By providing standardized, high-density, and curated phenotypic labels, these datasets prevent generative agents from drifting into biologically implausible space.
Looking ahead, I anticipate the RxRx3 ecosystem will be repurposed as the foundational training and validation ground for predictive simulations. Instead of merely classifying images, the next generation of methods will use this data to constrain the latent space of generative models, forcing them to learn not just the appearance of a cell, but the underlying mechanisms of perturbation response. Ultimately, the future of virtual cell research depends on our ability to map the nuance between generated biology and observed biology. In this context, RxRx3 is no longer just a benchmark, it is a foundational dataset that will keep generative AI for biology tethered to realistic cell states.
RxRx3: huggingface.co/datasets/recursionpharma/rxrx3-core
RxRx3-core: huggingface.co/datasets/recursionpharma/rxrx3
RxRx3-core benchmarks: github.com/recursionpharma/EFAAR_benchmarking/tree/trunk/RxRx3-core_benchmarks
Celik, S., et al. (2024). “EFAAR: A framework for evaluating AI in High-Content Screening.” PLOS Computational Biology. https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012463
Fay, M., et al. (2023). “RxRx3: Phenomics Map of Biology.” bioRxiv. https://doi.org/10.1101/2023.02.07.527350
Fradkin, P., et al. (2024). “MolPhenix: Aligning molecular graphs with visual phenotypes.” NeurIPS. https://arxiv.org/abs/2409.08302
Kraus, O., et al. (2024). “Masked Autoencoders for Microscopy are Scalable Learners of Cellular Biology.” CVPR. https://arxiv.org/abs/2404.10242
Kraus, O., et al. (2025). “RxRx3-core: Benchmarking drug-target interactions in High-Content Microscopy.” ICLR Workshop. https://arxiv.org/abs/2503.20158
Earnshaw, B., et al. (2019). “CellSignal: Disentangling biological signal from experimental noise in cellular images.” NeurIPS 2019 Competition. https://www.kaggle.com/c/recursion-cellular-image-classification
Jones, C., et al. (2026). “Generative Phenomics with Hooke.” arXiv preprint. https://arxiv.org/abs/2603.26790
Kenyon-Dean, K. et al (2025). “ViTally Consistent: Scaling Biological Representation Learning for Cell Microscopy”. ICML https://arxiv.org/abs/2411.02572
Chandrasekaran, S. N. et al (2023) JUMP Cell Painting dataset: morphological impact of 136,000 chemical and genetic perturbations. bioRxiv https://www.biorxiv.org/content/10.1101/2023.03.23.534023v2
Moutakanni, T. et al. (2025) Cell-DINO: Self-supervised image-based embeddings for cell fluorescent microscopy. PLOS Comput. Biol. 21, e1013828 https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1013828#pcbi.1013828.ref035
Zhang, Y. et al. (2026) CellFluxV2 : An Image Generative Foundation Model for Virtual Cell Modeling. bioRxiv 2026.01.19.696785 (2026) https://www.biorxiv.org/content/10.64898/2026.01.19.696785v1.full.pdf
Sypetkowski, M. et al. (2023) RxRx1: A Dataset for Evaluating Experimental Batch Correction Methods. CVMI at CVPR https://arxiv.org/abs/2301.05768
This post is part of “Inside Valence”, a series where you’ll get a behind-the-scenes look at our research, exploring new ways to predict, explain, and ultimately decode biology. If this resonates, consider subscribing!

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.