Over the past few months, I have been diving into the field of biology. One thing that surprised me was how we have incredible depth in specific domains but lack unified frameworks that connect them. People spend years exploring a particular pathway while we still lack a coherent theory that integrates size control, compartmentalization, force sensing, and timing into a unified model of cellular organization. It just seems to me that to escape the Eroom's law in biological sciences (drug innovation slows, costs rise) new paradigms must be found. My goal here is not to undermine the breakthroughs like insulin, antibiotics, and vaccines that have saved countless lives and spread pessimism. My intent in this series is to think about biology from first principles. I will share some experiments that I have been doing, questions that interest me, and potentially find like-minded people who are interested in a fresh view of biology.
One of the biggest questions worth pondering is what is a biological organism trying to do. Numerous theories have been proposed, from the foundational Darwinian views to the Optimal Foraging Theory. Surely one could write a whole book on this topic.
The one that I found to be the most cohesive and convincing is by Michael Levin. In short, he states that an organism is trying to navigate a spatiotemporal manifold (essentially moving through space and time) to achieve a specific goal. I like this definition because it does not discard evolution, does not look at cells as merely clockwork, and provides clear goals for the field of bioengineering:
1. Understand the implicit goal of a biological system
2. Redirect its trajectory to either correct errors or create new functions/forms.
We are still searching for answers to how goals are stored and maintained across scales, how the different levels communicate and coordinate, and let alone how we model and control these processes. Let me explore a few of these challenges and poke current approaches to see where they might be failing in helping us find an answer.
To understand the goal of a system we should first understand what exactly we want to be looking at. A cell, a tissue, or maybe a protein? Each of these systems has its variables and is organized in an integrated multiscale architecture. Wound healing in mammals consists of clotting factors (proteins) assembling at the wound site, platelets and immune cells detecting damage on the cellular level, and nervous/endocrine systems modulating healing speeds at an organismal level. Each of these processes happens on distinct time scales and each solves its own subproblem (either a protein folding on a scale of milliseconds or a piece of tissue trying to heal a wound in a few days). But what scale are we supposed to look at? Should we be looking at the electron spins, DNA methylation, intracellular communication schemes, or social dynamics among species? I think this is a lot more complicated than it might initially seem. Of course, the scale to choose depends on what exactly the problem we are trying to solve. Let's say that we are interested in solving aging. Virtually all traditional hallmarks of aging are grounded in molecular or cellular-level phenomena. However, there are emerging views that aging could be upstream of traditional hallmarks. For example, Peter Lidsky proposes a "Pathogen Control" theory where aging has evolved as a mechanism to protect the broader population from the infectious disease that older individuals have accumulated. This suggests a form of distributed social computation where individual aging trajectories are coordinated at the population level to optimize collective survival. For someone focused on the molecular and cellular-level phenomena, this view seems very confusing, and outright wrong (just check comments under some of his Twitter posts). But zooming in closer on his theory, we see that it holds up surprisingly well against the classic damage accumulation argument. This all leads me to think that we are still in search of the correct abstraction level to solve aging. There are of course a lot more problems to think about with their own unique abstraction levels which have different data modality implications. Whether it's bioelectricity in planaria morphogenesis, or immune memory formation across decades in humans, they operate on different levels of abstraction. However, we are currently largely stuck in DNA → RNA → Protein framework to solve all the problems.
This challenge becomes particularly evident when we look at AI in drug discovery. Specifically, the fact that it has not shown much meaningful results. Sure, there has been tremendous progress with protein folding, but this is one of the lower levels of organization and it is a relatively well-defined problem with a specific beginning and an end. Insilico’s “first end-to-end AI-designed drug” failed to show statistical efficacy in Phase 2a trials, and Recursion’s AI-derived candidate also failed to show benefit in its initial clinic study. There are huge datasets with millions of samples across various species with countless perturbations, yet something fundamental is missing. The traditional view is that we need more standardized datasets, and simply throwing more data at a problem will do it. But maybe the answer is not in more data, but rather which kind of data we are looking at. For example, we know that data quality in deep learning models can completely reshape the scaling law exponents, and sometimes offer better returns than simply enlarging the dataset. From an information-theoretic view, models cannot extract signals that aren't there. If your dataset doesn't capture the relevant dynamics of the phenomenon, the model will fail regardless of its size. In most current methods, whether it is scGPT, or velocity inference, you are looking at the static cell snapshots in time. As someone who is new to biology, this always looked strange to me. It's like trying to understand how you plan to move your right arm by looking at a single neuron spike once a day, or trying to understand music by looking at individual notes in isolation. If we are conceptualizing cells as dynamic agents that are trying to achieve goals, it would make sense to go all in on the spatiotemporal data. But the field is still largely tilted towards single-cell sequencing with bits of spatial omics here and there. I'm not suggesting that spatiotemporal tools exist but are simply being ignored. Rather, these tools remain underdeveloped, yet the community doesn't treat this as the central challenge it should be.
This is not meant to be a somber outlook on biology, aging, or AI. What it seems to me is that there is a huge opportunity to re-evaluate the field from first principles and look at things from a new perspective. I am not proposing a revolutionary idea here. This has been recognized by various people in the field who are approaching this from different perspectives whether it be from a systems-level viewpoint, a developmental framework, or even a philosophical critique of scientific paradigms. A common recognition unites them that the complexity of life cannot be fully captured by a gene-centric and reductionist framework.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.