RSS Amplifier

Fernando’s Substack · Aug 9, 2026

How Einstein used intuition to discover general relativity

0
Sign in to vote or save

Fernando Palafox · Fernando’s Substack

Tom Zahavy, a Google DeepMind researcher, recently published a position paper called “LLMs Can’t Jump.” In it, Zahavy argues that large language models (LLMs) do not possess the capacity for generating axioms from which fundamentally new hypotheses about the universe may arise.

This may sound surprising given recent news of success for LLM’s in fields like mathematics (I wrote about one such case here). However, under Zahavy’s position, such success can be characterized as a logical deduction from an existing set of axioms. Impressive, but not the same process that led to fundamental discoveries like Einstein’s formulation of General Relativity (GR).

Zahavy argues that the process leading to paradigm-shifting discoveries (like GR) requires proposing fundamentally new axioms even when the empirical data is

  1. consistent with prevailing paradigms (e.g., Newtonian physics in Einstein’s case),

  2. scarce (or non-existent).

For example, at the time of Einstein’s invention of GR in 1915, Newtonian physics had been verified to a very small margin of error. And the data to verify GR as correct was extremely limited: an anomaly in the advance of Mercury’s perihelion. In fact, most of the pro-GR data arrived much later (e.g., Eddington’s light-bending experiment in 1919 or relativistic GPS correction in the 1970s)

A photograph of the total solar eclipse of 29 May 1919. Sir Arthur Eddington used it to measure the gravitational deflection of light—verifying Einstein’s theory. Image taken from this Wikipedia article.

This means that Einstein had to come up with a new set of axioms even when the available data was largely consistent with the accepted theory at the time. That is, data with a small “error signal.” This is bad news for existing LLM architectures which are, for the most part, inductive and great at forming hypotheses—but only when the data is plentiful and rich in error signal.

Moreover, as Zahavy argues, LLMs lack the crucial ability for embodied simulation: an active interaction with mental models of the universe, which provides the missing data from which scientists can pose new axioms. Einstein had very little data to use as inspiration, as most of it was already consistent with Newtonian physics. So he “generated” new data with one of his famous thought experiments:

Consider a physicist inside an elevator with no windows, floating in deep space. If we place the elevator on a planet, like Earth, the physicist will feel the pull of gravity and see any floating objects around him fall to the ground. If instead we strap rockets onto the elevator and accelerate it, from the physicist’s point of view the effect will be identical. He and any floating objects will be pulled towards the elevator’s floor.

Einstein’s thought experiment. Both physicists see the ball reach the floor.

This thought experiment, derived from what Einstein later referred to as “the happiest thought of [his] life,” had an important implication: if the physicist cannot distinguish between acceleration or the pull of a gravitational field, then gravity and acceleration must be physically identical (at least locally). And using this framing, Einstein eventually concluded that gravity is the physical curvature of spacetime, instead of a pull between two masses as Newtonian physics predicted.

Einstein essentially generated the missing data he needed by interacting with his mental model of the universe. And, according to Zahavy, this is something existing LLM architectures cannot do. They lack the high-fidelity models humans manipulate in search of inspiration for new axioms.

This limitation is largely driven by the fact that LLMs, for the most part, understand the world through language, a lossy abstraction which often fails to capture the richness of reality. And, as Zahavy explains,

“scientific revolutions are often driven by strong, pre-symbolic intuitions — whether Kepler’s Neoplatonic belief in the centrality of the Sun or the “objective anger” that drove Marx’s modeling of capital.”

Unlike LLMs, humans can manipulate models of the universe that are rooted in the rich sensory streams of, e.g., sight, touch, and smell. It was exactly this high-dimensional intuition of reality that Einstein leveraged when he imagined a physicist “feeling” the tug of acceleration and gravity as equal.

This was an extremely interesting paper, and I highly recommend reading it if you're interested in the structure of scientific revolutions and the role of AI and humans in them. I’m left with a few questions I’ll explore in future posts like: How can we quantify the level of sensory richness that humans experience? And what (if anything) prevents current AI architectures from capturing it? If AI can capture the richness of reality, what will the role of humans in scientific discovery? And what if it can’t? What is our role then?

Post 39 out of 100

No posts

Read the original on fernandopalafox.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.