RSS Amplifier

The Infinitesimal · Dec 8, 2025

More missing heritability discourse

0
Sign in to vote or save

Sasha Gusev · The Infinitesimal

Untitled (from “On a Clear Day”), Agnes Martin, 1973

There has been a flurry of discussion on missing heritability over the past few weeks, so I want to highlight a few articles worth reading for different perspectives on the topic. I have quoted some sections I found most relevant alongside my own commentary (which conveniently lets me get the last word in), but I encourage you to read them in full. I also recently wrote about the state of missing heritability here:

Matthew Yglesias makes the fundamental point that heritability is a snapshot of contemporary variation rather than an estimate of malleability. Just because a trait has high heritability does not mean it is unchangeable (emphasis mine):

What heritability means is that a large fraction of the observed variation in obesity is statistically explained by genetic variation. But that’s not to say there are “fat genes” that are making people fat for no reason. The world has changed in ways that make it a lot easier to overeat and a lot harder to avoid overeating, and there is evidently a genetic difference in the propensity to respond to those changes by gaining weight. But the genes don’t explain the large rise in obesity over time; they explain the cross-sectional variation in who specifically tends to get fat.

We now, of course, have surgical and pharmaceutical interventions that can act on and, to some extent, override these genetic predispositions. Similarly, we have eyeglasses that make it quite easy to remedy nearsightedness.

And I think in policy terms, we are mostly interested in questions about how changeable outcomes are, and changeability is not simply the inverse of heritability.


Freddie DeBoer emphasizes the same idea but from the other direction — just because something has low heritability does not mean that it is changeable either:

I have, as you know, invested an immense amount of time in aggregating research that demonstrates that the relative distribution of students in the academic performance spectrum is largely static. Here’s 18,000 words with links to nearly 200 sources, if you’d like to consider the evidence yourself. I’m not equipped to scientifically assess the heritability of intelligence, but I can certainly tell you that decade after decade of education research has demonstrated that students gravitate to a level of academic performance very early in life and tend to stay there, regardless of environment, school type, pedagogy, policy, or intervention. This is what the anti-hereditarians have worked relentlessly to avoid, the reality that regardless of cause, we all have academic constraints that we operate under that schooling cannot alter.


Eric Turkheimer reiterates that variance component estimates do not provide an answer to the question we actually care about — causal mechanisms:

On the other hand, our inability to define anything even resembling deterministic causal mechanisms underlying heritability coefficients places strict limitations on their application in the real world, or on our theoretical accounts of how human inequality (in the nonpejorative evolutionary sense) comes to be. If we ask, “if I had different genes, would I have a different IQ,” the answer is probably, yes. If we ask, “if I had these specific genes, how would my IQ be different?” (never mind why) the answer is almost entirely, we don’t know. We don’t know because it depends on a thousand other human contextual variables that we have no hope of controlling.


Greg Gibson considers what the heritability discussion would look like if we focused on measuring the Es as much as the Gs, and the potential for the emerging field of “exposomics” to get us there:

What I am excited about with exposomics is the potential to illuminate still-hidden secrets of genetic susceptibility. Whether it is environmental modulation of polygenic risk in small communities, or bespoke interaction measures explaining why two people with equivalent genetic risk have dissimilar outcomes, exposure measurement at scale might well be a bridge between heritability and inheritance.

I think these ideas are all circling around Turkheimer’s point that variance components do not estimate causal mechanisms. Causal mechanisms are the resolution to the concerns that Yglesias and DeBoer raised. For instance, PKU is “100%” heritable but we understand the mechanism and can fully ameliorate it with existing dietary changes. Knowing the mechanism allows us to reason about interventions and policy implications. What variance components can tell us is how much of a trait could be predicted, in principle, from different types of variation. This is a useful quantity for clinical risk screening and not much else. Under some strong assumptions, variance components can also inform where we should focus our search: as Turkheimer points out, if heritability is high and you think you’ve found an environmental risk factor, you should double check in adoptees. But if gene-environment interactions are involved (as is almost certainly the case), then the variance components you think you are estimating are incorrect, and many unexpected mechanisms are possible. I share Gibson’s excitement that better measures of the exposome can help identify some of these environmental and interactive causes and I think the gap between molecular and twin estimates (e.g. for BMI) suggests there is a lot to find.

Setting variance aside, I want to push back against DeBoer’s fatalism. I’m not an education policy expert but I’ve read enough to draw the following conclusions:

  1. High quality studies have shown that intensive one-on-one interventions (e.g. tutoring) or large-scale environmental changes (e.g. adoption, moving neighborhoods at an early age) can have a large effect on cognitive outcomes.

  2. High quality studies show that the crude intervention of “more schooling” has a significant impact on cognitive outcomes (with a wide range of effect-sizes, see here and here).

  3. A small number of randomized trials show that early pre-school interventions can have effects that are both positive or negative, fade out or fade back in.

  4. Individual systematic school-level interventions (e.g. phonics, tracking, gifted & talented programs) generally have small (but measurable) positive effects.

  5. Large-scale school disruption (e.g. COVID) has substantial and lasting effects on performance, particularly for low performing students.

DeBoer tends to focus on point (4), often by meta-analyzing the results of many education RCTs, showing that they are close to null, and concluding that “education doesn’t work”. This is intuitively impressive but doesn’t directly answer the question that DeBoer is asking. One could run the same meta-analysis across cancer clinical trials or drug RCTs and similarly conclude that “medicine doesn’t work”; yet we have clearly had remarkable advances in cancer treatment and breakthrough drugs. What such a meta-analysis actually tells us is that most attempts fail, which is not sufficient to conclude that nothing works (and, as in cancer, we should probably expect that most attempts will fail for any complex system). As to the broader question about mutability, the other four points I listed above clearly show that academic outcomes are significantly mutable! It is just that the mechanisms for doing so are either highly intensive (tutoring, moving neighborhoods) or still largely unknown (COVID learning loss heterogeneity, negative pre-k interventions).

Switching gears, David Bessis does a deep dive into studies of Twins Raised Apart (TRAs), pointing out that the twins are almost never actually “raised apart” — nevermind that they always share a womb — and the whole study design is a mirage:

I am not a behavioral geneticist and, paradoxically, this is precisely why I felt the need to write this detailed account of my thought process. Despite my solid background in mathematics and data analysis, it cost me serious effort to build a watertight debunk, and I thought it was worth sharing it with the other 3,499,999.

As I emerged from the rabbit hole, this is what struck me—this wasn’t really about Cremieux’s slide, this was about eradicating the “twins separated at birth” trope that had infected me as a teen and distorted my expectations of what was scientifically plausible.

It is fundamentally hard to believe that such a simple and beautiful design could be so profoundly flawed. And yet, people have tried it over and over again, and failed over and over again.

I really commiserate with Bessis here, TRAs feels like the perfect method for breaking the Gordian Knot of “nature-nurture” modeling assumptions. It is only when one looks at the data and thinks about the underlying mechanisms that it becomes clear that one set of assumptions has merely been substituted for another.

I’ll make two addendums to Bessis’ article. First, a recent analysis pulled together all of the publicly available twins raised apart data (a mere 87 pairs) and re-evaluated how similar their rearing environments were, ranking from “similar” to “very dissimilar”, and re-estimating the IQ differences (to be fair, these environmental rankings are extremely subjective). When you overlay these IQ differences on contemporaneous estimates from other relatedness classes the results are striking (figure mine):

At one end, TRAs that grew up in similar environments have IQ differences equivalent to that of one individual tested twice! At the other end, TRAs with very dissimilar environments have IQ differences that are indistinguishable from strangers. In other words, this data can give you whatever answer you like!

Second, Bessis zooms in on the MISTRA study (Bouchard et al. Science), which is not publicly available, and which famously reported a heritability of ~70% for IQ based on MZA twins. Bessis recounts how strange it is that, after laboriously collecting DZA data (a perfect control), the authors elected not to publish them, allegedly due to “space constraints”. Interestingly, these DZA correlations were recently published as a one sentence aside in Segal et al. (2025), and they produce vastly different heritability estimates: 24% for IQ and 52% for the general factor. The high DZA correlations are indicative of selective placement (i.e. twins assigned to adoptive homes non-randomly), which also undermines the entire premise of random, separated rearing environments. No wonder they weren’t reported!

Taken together with these newer data, Bessis has a strong case that the whole “raised apart” study design was essentially a dead end.

Kathryn Paige Harden has a typically thoughtful piece on what is actually interesting about missing heritability (and missing “environmentality”):

Heritability is missing, but so is environmentality. Let’s say we halve every heritability estimate from a classical twin study, presuming that the estimate is inflated, and attribute that variance to the “shared environment.” Where are the causal effects of specific environmental influences that add up to anything remotely close to that shared environmental variance component? They don’t exist. Even when you change literally everything about a child’s life by adopting them into an entirely new family, or adopting them out of hellacious institutional care, you still don’t get effect sizes big enough to explain the incredible similarity of identical twins. The “missing heritability problem” is just another manifestation of a much more general problem—the granularity problem, the reductionism problem. Human lives are both undeniably structured by naturenurtureluck and very poorly predicted by individual variables, at least the ones we currently know how to measure.

Harden and I have commented on overlapping topics in the past and I’m always intrigued to see how much our perspectives differ. In a previous post on embryo selection, I laid out my critiques of the expected gains and the disease modeling assumptions. Around the same time, Harden argued that (I’m paraphrasing) IVF is really fucking painful and difficult, and all anyone seems to care about is genetic “yield” and optimization. Now on the topic of missing heritability, I focused on itemizing the various estimators and estimands. Harden wrote about what actually makes twin studies interesting: the fact that MZ twins really are remarkably similar and we don’t know why. In this way, I think Harden’s piece is a useful counterpoint to Bessis’; set aside the flawed and/or fraudulent “raised apart” studies and there is still something genuinely magical about the similarity between twins even in high-quality registers.

But, Lord forgive me, I also have to get back on my bullshit: twins are not just studied to understand the magic of twins, they are used to draw extensive generalizations about the world around us and, often, to make specific policy arguments. To date, the primary presumed reason for MZ similarity was additive genetic variance in the general population. Twins were just the tool to itemize this variance. If the truth is that (a) MZ twins reshape their mutual environments in a highly unusual, twin-specific way or (b) additive genetics and environment interact and amplify in MZ twins; then the utility of generalizing from twin studies is lost. This is true not just of heritability estimates, but of the broader user of twins as a genetic control in the social sciences. So, yes, let’s understand the mystery, but let’s not lose track of the fact that we are also talking about a study design intended (and routinely wielded) to answer questions about non-twins too.

Finally, Scott Alexander casts the debate as the struggle between “hereditarians” and “nurturists” and concludes that neither side has won but both have proclaimed victory:

Here are the two stories you could tell, updated for this new paper:

Hereditarian: Most traits are 50 - 80% heritable, as per twin studies, adoption studies, and classic pedigree studies. Molecular genetics studies underestimate this because much of the heritability is in rare variants, as this new study demonstrates. Sib-regression, RDR, and this new study’s “pedigree-style” analysis underestimate this because they’re untested methods applied to problematic samples and the estimates are noisy; also, shut up.

Nurturist: Most traits are ~30% heritable, as per Sib-regression, RDR, molecular genetics, and this new study’s “pedigree-style” analysis. Twin studies, adoption studies, and pedigree studies overestimate this because of assortative mating and population stratification. This affects biomedical traits like white blood cell count just as much as behavioral traits, because shut up. The one sib-regression study that found very high heritability for IQ was just a weird sample, or noise.

It’s worth reiterating, again, that heritability is not malleability and the underlying “hereditarian versus nurturist” debate is really about causes. If, for instance, the heritability of educational attainment turns out to be very high but fully explained by discrimination on skin color or appearance (obviously heritable traits), would “hereditarians” really claim this as a victory?

Scott also makes a few technical errors that somewhat muddle the positions he defines, particularly around assortative mating and stratification (as Vinay Tummarakota explained on twitter). But since Scott has volunteered me as spokesperson for the “nurturists”, I would rewrite our position as follows:

Nurturist: Most traits are ~30% heritable, as per Sib-regression, RDR, molecular genetics, and this new study’s “pedigree-style” analysis [And the new study’s pedigree analysis shows a range of heritabilities up to 40%, likely inflated by shared environments]. Twin studies, adoption studies, and pedigree studies overestimate this because of assortative mating and population stratification [flawed equal environment assumptions and unmodeled gene-environment interactions]. This affects biomedical traits like white blood cell count just as much as behavioral traits, because shut up [every trait interacts with the shared environment, which the pedigree analysis tells us is substantial]. The one sib-regression study that found very high heritability for IQ was just a weird sample, or noise [and paradoxically found very low estimates of heritability for the education and labor outcomes that IQ is supposed to strongly predict].

I would also challenge the “hereditarian” claim that “Molecular genetics studies underestimate this because much of the heritability is in rare variants, as this new study demonstrates”. The reader is forgiven for taking this at face value, since Scott’s post never actually mentions the amount of heritability coming from rare variants in this new study. As a matter of fact, it was just 5% on average across traits. To put that 5% into context, Scott’s earlier post on missing heritability hypothesized that rare variants might explain the missing 20-25% of the variance in EA (presumably even more for IQ, which has an even larger twin heritability gap). 5%, it should be noted, is much less than 20-25%.

It might seem like I’m fixating on a detail, but the hypothesis that the missing heritability is explained by a massive tranche of rare variation has been put forth for my entire academic life. It’s there in the Discussion section of thousands of underpowered association studies that came up short. It’s there in the “disattenuated” polygenic score analyses. It’s even there in the 2008 article that coined the missing heritability debate, where Francis Collins (shortly before his rise to NIH director) makes the bold prediction that the 1,000 Genomes Project will “go a long way towards finding hidden heritability”. Half a million genomes later we have an answer: at least for the typical biomedical trait, this hypothesis is now untenable.

Read the original on theinfinitesimal.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.