RSS Amplifier

Complexity Thoughts · Jul 3, 2026

Decoding the “Architecture” of Living Systems: Chapter 2

0
Sign in to vote or save

Complexity Thoughts · Complexity Thoughts

If you find value in #ComplexityThoughts, consider helping it grow by subscribing and sharing it with friends, colleagues or on social media. Your support makes a real difference.

→ Don’t miss the podcast version of this post: click on “Spotify/Apple Podcast” above!

Can “architecture” be a scientific concept for life and not just a metaphor?

If life could simply “add more connections” it would already be fully connected. It is not, and the reason is not biological creativity, but physical constraint.

Based on the more technical arguments developed in my recent paper, this is the second step in our attempt to understand whether “architecture” in living systems is a real, measurable concept or just a convenient metaphor. I have introduced the key idea in the previous post:

where I have outlined how function in living systems emerges from the coupling between logic (what the system does) and circuitry (how it is physically implemented).

An important preliminary step is to ask — and understand, if possible — how life has learned to store, replicate, and maintain the intervening network across generations. In this post we will essentially deal with this point, that is where a fundamental paradox (only apparently) appears, while we leave for the future posts the discussion about how life build complex networks.

A common intuition — especially when thinking about brains, genomes or ecosystems — is that complexity could always be increased by adding more components and more connections1. After all, more links should mean more functionality, more robustness, more adaptability.

But this intuition implicitly assumes that structure is free, and it is not. In living systems, every connection has to be:

  • constructed during development

  • specified (directly or indirectly) by genetic information

  • maintained over time

  • and corrected when errors occur.

Each of the above steps carries a cost in energy, resources, time, and information processing. Note that this is not a secondary detail: it is a primary constraint shaping what kinds of architectures are even feasible in principle.

In complex systems terms, structure is not just a static property: it is a process under constraints. This idea resonates with constraint-based theories of living systems, or C-theory as I like to name it with Ricard Solè, which treat living organization as shaped by multiple constraints acting across scales.

To see why this matters, it helps to briefly recall how biological information is transmitted. A simplified version of the central dogma of molecular biology states that sequence information flows from DNA to RNA to proteins: DNA stores instructions, RNA mediates their expression, and proteins implement most cellular functions. Crucially, across generations, what is reliably copied is DNA, not the full physical state of the organism: cells, tissues and entire organisms are reconstructed each time from this limited description.

A cartoonish summary of the transfer of sequence information in molecular biology, according to Crick, 1958 and Crick, 1970.

This has a direct consequence: whatever “blueprint” evolution relies on must be encoded in the genome and passed on through replication. However the genome is finite, noisy and costly to maintain: should it explicitly list all the microscopic details of a mature organism, as well as the full connectivity of its internal networks?
So the unavoidable question becomes: what kind of information is actually stored, if not the full wiring diagram?

At first sight, this sounds counterintuitive. After all, reproduction looks like copying. If your dog has puppies, each of them is recognizably a “dog”: same body plan, same organs, similar behavior. Yet none of them is an exact replica of the parent: what is preserved is a set of instructions that reliably reconstruct a dog under a range of conditions, not necessarily a detailed specification of every cell and connection.
In other words, inheritance is not the transmission of a finished structure: we can better see it as the transmission of a generative process that produces structures within constrained variability.

The apparent paradox of the information storage becomes clearer when we compare two well-known biological systems. At the cellular scale, the budding yeast Saccharomyces cerevisiae has on the order of 6,000 genes, with roughly 170,000 pairwise genetic interactions identified experimentally.

At the organismal scale, the human brain contains approximately 86 billion neurons, with synaptic densities implying on the order of 10^{14} (i.e., 100 trillion) connections.

Now let us consider a simple (and intentionally naive) thought experiment. Suppose evolution had to encode the full connectivity of such a system explicitly, similarly to an engineer storing a wiring diagram. Each possible connection is represented by one bit: 1 if present, 0 if absent. Of course, the total number of possible configurations then grows combinatorially with system size.

For the brain, this leads to a space of possible networks of 2^{10^{22}}: so large that the information required to specify one exact configuration becomes astronomically high. In the language of statistical physics, the number of microstates explodes and the corresponding information needed to specify one configuration becomes enormous: the capacity of an ideal “device” to losslessly store this information is the maximum Shannon entropy log2 N, that in this case reduces to 10^{22} bits (about 1 ZB), which is energetically expensive for computational purposes.

Now compare this with the actual storage capacity of DNA. The human genome contains about 3 billion base pairs, corresponding to roughly 6 x 10^9 bits of information. This is an impressive storage system, but it is about 12–13 orders of magnitude smaller than what would be required to explicitly encode a full neural adjacency matrix. Even if we switch to more efficient encodings — such as adjacency lists rather than full matrices — the gap would still remain enormous.

The genome simply does not have the capacity to enumerate all connections: this is not a small mismatch, it is a structural impossibility.

However, the paradox resolves once we abandon the implicit engineering analogy and follow the causal chain more carefully, that we can summarize in 5 points.

The genome should not be understood as the whole “program” in isolation. It is better seen as a compressed handle into a much larger natural computation executed by cellular chemistry, developmental dynamics, boundary conditions, and the physical substrate itself. The fact that a relatively small genome can generate a vast regulatory or morphological architecture is therefore not paradoxical. It is exactly what one expects if the architecture is produced by a short generative description coupled to lawful dynamics that amplify, unfold, and stabilize that description. — Hector Zenil

1. Enumerating connections scales brutally: the number of possible connections in a network grows faster than linearly with system size. For large systems, explicitly listing all connections becomes combinatorially prohibitive, and this is a scaling problem, not a biological detail.

2. Information has physical cost: information is not abstract in living systems. Storing, copying and processing it requires physical operations that consume energy and produce dissipation. This principle is formalized in information thermodynamics: erasing or manipulating information has a minimum energetic cost, as established by Landauer in 1961 and further developed by Bennett in 1982. Thus, increasing informational detail is not free: it has thermodynamic consequences.

3. Repair and error-correction add recurring costs: biological systems are not static. They are far from equilibrium and continuously subject to noise, damage and turnover. DNA must be repaired, proteins must be refolded or degraded, and synapses are constantly formed and eliminated. Each of these processes requires control mechanisms — e.g., proofreading, signaling, feedback loops — that consume energy and resources: and they are not one-time costs, they are ongoing.

4. Evolution transmits generative rules, not networks: given these constraints, the solution adopted by evolution is conceptually simple but profound. Instead of encoding a full wiring diagram, organisms encode rules for constructing the network.

These rules operate locally:

  • how a neuron grows and forms synapses

  • how a gene regulates others under certain conditions

  • how cells signal and organize spatially.

From these local interactions, the global network emerges during development, and this is vastly more efficient. A compact set of generative instructions can produce a large, complex structure without specifying every connection explicitly.

Note. This perspective closely mirrors a well-known idea in information theory: the distinction between describing an object and describing a program that generates it. In algorithmic terms, a structure can often be represented much more efficiently by a set of rules than by an explicit listing of all its components. A long sequence that looks complex may in fact have a short “program” if it is generated by regularities or constraints, whereas a truly random structure cannot be compressed.

Living systems appear to exploit a similar principle, and this idea is currently being explored also by my colleagues, see Zenil et al (2016), Hernandez-Orozco et al (2018), Levin (2023), Mitchell and Cheney (2025)

Across animals as different as insects and vertebrates, homologous Hox genes orchestrate embryonic development by defining how structures are built, not what every structure must be. Their conservation over evolutionary time suggests that genomes encode generative rules— programs that reconstruct form under constraints — rather than detailed blueprints of the final organism. Source: Wikipedia

5. The consequence: architecture is probabilistic and constrained. There is a trade-off: generative rules introduce variability; development is not perfectly deterministic; noise, stochastic gene expression and environmental influences all contribute to differences in the final structure. As a result, biological architectures cannot be fixed blueprints, but operate as distributions of possible configurations constrained by genetic rules and environmental context. The brain is again a clear example: its connectivity is shaped by genetic programs, but also by activity, experience and developmental fluctuations. The final structure is reproducible in function, but not identical in detail.

In this sense, the “blueprint” of life is better understood as a program under constraints, not a static schematic.

Although the brain provides the most striking illustration, the same logic applies across biological scales:

  • Regulatory networks in cells are not explicitly encoded edge by edge. They emerge from biochemical interactions, binding affinities and feedback loops shaped by evolution.

  • Metabolic networks arise from chemical constraints and catalytic possibilities, then are stabilized through regulatory layers that manage flux and error correction.

  • At the ecological level, the same limitation appears in a different form. It is not feasible to realize all possible species interactions. Coordination, maintenance, and stability constraints limit which connections persist.

Across these systems, the pattern is consistent: networks are not written down, they are generated. So the next question is: how to study architecture when the blueprint is generative?

I would distinguish between abstract algorithmic complexity and what one might call effective or constrained biological complexity. The latter is not the length of the shortest program in the unrestricted sense, but the length of the shortest physically admissible generator capable of producing and maintaining a biological architecture under finite energetic, temporal, and control budgets, and with some required level of reliability in the presence of noise. Once one includes these constraints, one is no longer speaking of K in the strict canonical sense, but of a resource-bounded constrained version of K relative to a class of admissible generators which is also covered by Algorithmic Information Theory — Hector Zenil

If connectivity is the outcome of generative rules rather than explicit specification, then our analytical approach must shift accordingly:

  1. First, we should focus on generative signatures: instead of treating networks as fixed objects, we analyze recurring patterns such as motifs, degree distributions, modular organization and growth mechanisms, encoding the underlying rules.

  2. Second, we should quantify control and maintenance costs: where systems invest energy in repair, error correction and regulation reveals the constraints shaping their architecture.

  3. Third, we should model networks as ensembles rather than single realizations: the relevant object is not one graph, but a distribution of possible graphs generated by rules under stochasticity.

You might recognize why this perspective aligns naturally with statistical physics and network theory, where macroscopic behavior emerges from ensembles constrained by underlying principles.

I have discussed with Ricard Solè, Hector Zenil and Kevin Mitchell about the key points of this post. Below, you can find my questions and their answers.

A short generative program (genome + rules) can produce a vast, complex network (e.g., regulatory). However, biological systems are not abstract Turing machines; they operate under noise, dissipation and continuous error correction.

From the perspective of algorithmic information theory, as developed in your work, how should we reinterpret Kolmogorov complexity when the generating process itself has a physical cost and limited reliability? In particular, is there a meaningful way to define an effective algorithmic complexity of biological architectures that incorporates thermodynamic and control constraints, such that some structures are not just incompressible, but unrealizable in practice, even if they are algorithmically simple?

Hector: Biological systems are not literal Turing machines, but that is not the relevant objection, because Kolmogorov complexity is not tied to the peculiarities of tape and head, only to universality. I would therefore not redefine K itself: K remains the right measure of abstract generative compressibility. What changes in biology is not K, but the admissible class of generators. Once noise, dissipation, control, and error correction are taken seriously, the relevant quantity becomes the shortest physically executable and reliable generator of a biological architecture, not the shortest generator in the unrestricted mathematical sense.

In your recent framing of the genome as a generative model (a compressed latent-variable system decoded through development) there is a strong emphasis on how evolution encodes rules that reliably reconstruct organismal form. My argument, however, introduces an additional layer: that architecture is explicitly shaped by physical costs (construction, maintenance, repair), not just by representational efficiency.

How would you formalize the integration of these two perspectives? Specifically, can the latent generative space be extended to include explicit cost functionals (energetic, informational, control), such that developmental dynamics produce architectures that are near-optimal under these constraints? Or does your framework imply that such constraints are only implicitly embedded in the generative rules, making them effectively non-identifiable and therefore limiting the possibility of a predictive theory?

Kevin: I think the questions of physical costs and representational efficiency are intimately tied together. If you give a system infinite resources, it doesn’t need to do any compression. We see this with many artificial systems, which have effectively unlimited compute and energy, and which consequently can maintain huge amounts of data, often over-fitting as a result. By contrast, living systems always face resource constraints and are forced to be as efficient as possible.

In nervous systems, this means minimising wiring length and doing as much computation locally as possible and sending as little information as you can get away with and storing only the information you really need. (All of these principles are compellingly illustrated by Sterling and Laughlin in their 2015 book Principles of Neural Design). A consequence of these principles is the tendency towards compression, which in turn leads to a greater ability to generalise. By forcing the system to throw away lots of details, it encourages processes that abstract higher-order regularities instead. This turns out to be immensely powerful from a predictive point of view.

The same kinds of costs apply to the genome. Genes are expensive - to replicate, to proofread to ensure faithful replication, to transcribe and translate - all of those things take energy. This is why bacteria rapidly jettison genes they’re not using in a given environment - because a leaner genome means a more metabolically efficient cell and a faster reproductive cycle. These constraints thus impose the same tendency towards compression. In theory an organism could keep growing its genome and over-fitting its own phenotype. But it doesn’t need to and the efficiency constraints force it not to. What you get as a result is representational efficiency (along with robustness and evolvability).

If architecture is not a stored blueprint but a distribution over possible realizations generated under constraints, then C-theory shifts the object of study from single networks to ensembles.

Can this be elevated to a predictive framework by deriving these ensembles directly from first principles, combining information-theoretic limits (Shannon/Landauer), developmental dynamics and maintenance costs? More provocatively, do you expect the existence of forbidden regions in the (morpho)space of biological networks — architectures that are fundamentally unreachable regardless of evolutionary history — and what minimal set of constraints would be sufficient to derive such exclusions?

Ricard: I do think this can be made predictive if we treat biological organization as the intersection of fundamental constraints rather than a historical outcome: information-theoretic limits (Shannon bounds and Landauer costs), developmental dynamics (requiring robustness and generative capacity), and maintenance costs together carve out a restricted feasible morphospace.

Within this view, I would indeed expect forbidden regions, network architectures that are fundamentally unreachable because they would require infinite precision, violate energy–information bounds, or fail to sustain stable dynamics under noise. A minimal set of constraints to derive such exclusions would combine limits on information processing, energetic costs, dynamical stability (robust attractors), and scaling of maintenance with complexity, allowing the geometry of the feasible region itself to become predictive of what biological organizations can—and cannot—exist.

I have to thank my colleagues for their insightful answers and comments: this is a lot of food for (complexity) thoughts.

Therefore, the apparent “paradox” resolves in a precise way:

Living systems cannot afford to store and replicate full wiring diagrams. Instead, they transmit generative programs, complemented by control mechanisms that keep variability within functional bounds.

This shifts the meaning of “architecture” from a fixed design to a constrained generative process that produces stable yet variable structures.

However, as you might expect, this resolution introduces a new constraint: if connectivity cannot grow arbitrarily, and if dense networks are too costly to build and maintain, then robustness cannot come from simply “adding more links”; it must come from something more selective.

In the next post, we will look at the simplest candidate: loops (cycles), the minimal structural units that create alternative pathways and stabilizing feedback, and the reason robustness in living systems is never free.

→ Please, remind that if you find value in #ComplexityThoughts, you might consider helping it grow by subscribing, or by sharing it with friends, colleagues or on social media. See also this post to learn more about this space.

1

This is one of the fundamental assumptions feeding the AI hype, as I have discussed a while ago in this post

Read the original on manlius.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.