RSS Amplifier

The Educable Mind · Jun 1, 2026

The Shortest Explanation That Survives

0
Sign in to vote or save

Jon Webster · The Educable Mind

Here is a market story that sounds simple: equities fall when rates rise.

It has the appeal of compression. A large part of the market reduces to one variable. Discount rates rise, present values fall, long-duration assets suffer, growth stocks derate.

Then reality interferes. Rates rise and equities rally.

So the story acquires a patch. Perhaps earnings expectations improved, or the rise reflected stronger nominal growth, or positioning was too bearish, or the real yield mattered, not the nominal one. Any of these may be true, but notice what has happened. The headline rule was five words; the rule that survives contact with the data now runs to a paragraph of conditions, each added after the previous version failed. The story did not explain the market; it hid the market in the exceptions.

The problem is not that these variables are irrelevant; several belong in any decent model. What makes them special pleading is timing: they were not part of the model when it made its call, but were added afterwards, one for each failure, to explain it away.

Every investor is in the compression business. The world is too big to hold in mind, to trade whole, or to explain to a client or a committee, so you reduce it to a theme, a factor, a regime, a multiple, a line on a pitch deck. The danger is not that you compress, but that you forget you have done it, and mistake a short slogan for a short explanation.

What kind of problem is this?

When an explanation looks short but runs long once you count what it leaves unexplained, when a model survives not by predicting new data but by absorbing each contradiction as a special case, the question is not “is this simple?” but “how much special pleading does it need?” A corner of information theory and statistics has formalised exactly that trade-off, in a principle called Minimum Description Length.

So let’s borrow.

This is the multi-model move: recognise the shape of a problem, find the discipline that has thought rigorously about that shape, and import its frameworks deliberately rather than reinventing them from scratch.

Minimum Description Length, introduced by the statistician Jorma Rissanen in 1978, rests on a plain intuition. The best explanation is not the one with the fewest words, nor the one with the most moving parts. It is the one that gives the shortest total description of the observed data: the cost of specifying the model, plus the cost of encoding the data given that model.

A model is a compression scheme. To describe a set of observations you need two things: the rule, and the data it does not already capture. A short rule that leaves a large pile of exceptions yields no short description at all; a longer rule that shrinks the pile enough is worth paying for. Minimum Description Length counts both. Formally, it picks the model that minimises L(M) + L(D|M): the bits needed to describe the model, plus the bits needed to describe the data once you have it. The question is not how elegant the explanation is, but how much special pleading it takes to keep standing.

So three things that look alike should be kept apart. Simplicity is a short explanation. Compression is a short explanation that preserves the structure that will later matter. Deletion is a short explanation that throws away the part of the world that would have made the decision hard. Markets reward compression and punish deletion, and from a distance all three look identical.

A valuation multiple is compression. A price-to-earnings ratio takes an entire business, its margins, capital structure, and competitive position, and reduces it to a number that gives comparison, discipline, speed. But it also deletes. A low multiple may mean undervaluation; it may equally mean terminal decline, peak earnings, hidden leverage, or profits that will never become cash. The multiple is wrong not when it compresses, but when the variable it deletes is the one that decides the outcome.

Factors carry the same hazard. “Value works.” “Momentum persists.” Each compresses a historical regularity and carries an implicit claim about which details are safe to ignore, and the trouble begins when the label replaces the mechanism. Value is a family of mechanisms, from mean reversion to distress risk to institutional constraint, each implying different conditions under which the strategy should work, stop working, or reverse. The label compresses the mechanism into a single word. Use the word without the mechanism, and you cannot say which conditions the strategy depends on, or when it will stop working.

Backtests offer the appearance of compression in a subtler form. A strategy with enough rules can describe the past with great precision: buy this signal, exclude that sector, winsorise this variable, drop the crisis months, neutralise beta, reintroduce it when the regime changes. The result looks like a model, but it is closer to a lookup table for the past: it has not compressed the data, it has memorised it. A useful model says that, given this structure, these observations are no longer surprising; an overfit one says that, given these observations, it can build a structure that would have predicted them. The first compresses; the second only rearranges the archive.

A market narrative compresses attention: it tells investors which facts matter, and a thousand details collapse into a few decisive questions. But a narrative can explain everything by forbidding nothing. If stocks rise, it explains why; if they fall, it explains why. If margins expand the thesis is confirmed; if they compress it is a better entry point. Nothing is ever disconfirming; every observation is reclassified as support.

This feels like explanatory power and is usually the opposite. A good model has exclusions: it tells you what would count against it. A narrative that absorbs all evidence is not compressing the data; it is refusing to count its errors as a cost. The thesis may still be right, but it is no longer simple, and its description length is growing. A model that explains everything predicts nothing. The same instinct hides exceptions off balance sheet: a private asset marked at par, stable until the discount rate moves or the exit market closes or no sale forces the question. The real test is not whether the story can be repeated but whether it keeps reducing surprise as new data arrive.

A model that compresses well makes new data cheaper to interpret; you know where to put the next observation. A bad one does the reverse. Each new fact makes it more expensive, until it exists mainly to protect the original conclusion: still named and defended, simple from a distance, but no longer compressing anything.

For over a thousand years, astronomers did this. When observations failed to match the geocentric model, they added circles upon circles to preserve it; it could always be made to fit, but at a description that would not stop growing. Investors do the same. A company misses guidance, but the miss is temporary; then margins fall, but that is strategic investment; then management changes the metric, then the thesis, though the position does not. At each step the explanation is plausible, but plausibility is not the test; the test is whether the total explanation is getting shorter or longer. Sometimes the right move is not to sell but to admit the compression has failed: the position may still be attractive, but it now needs a different model.

There is a failure in the other direction: worshipping simplicity. Some investors treat complexity itself as error and demand one variable, one chart, one clean causal line. “Money printing causes inflation.” “Deficits push bond yields up.” “Cheap stocks beat expensive ones.” Each compresses something real, and none is sufficient, because the omitted variables do the deciding: velocity, credibility, global savings, refinancing conditions, the policy reaction function. Minimum Description Length does not say the simplest model wins; it says the shortest adequate description wins, and adequacy is the operative word. A map with one road is useless if the terrain has many. The over-padded model fails for the mirror reason: a thesis promising revenue growth, margin expansion, re-rating, benign competition, and skilful capital allocation may hold five assumptions, or one assumption stated five times, that everything goes right. Pay for complexity only when it earns its cost.

Consider two explanations for a bank that fails. The first: depositors panicked, which captures the final motion but not the stored instability. The second: it funded long-duration assets with flight-prone deposits, hedged little, and met rising rates carrying unrealised losses that turned fatal once confidence broke. The second is longer, yet it compresses more, because it separates the trigger from the vulnerability and tells you where to look for the same risk elsewhere. The better compression is not the shorter sentence but the one that discards the most irrelevant detail without discarding the cause. That is the standard: not elegance, not cleverness, not how easily a thesis can be sold, but the shortest explanation that survives.

The principle does not promise the model is true. It tells you what to watch. How many assumptions must hold, and are they independent or the same optimism in different cells? How many observations are filed as “temporary”, and how long has that file been growing? When a contradiction arrives, does the explanation get shorter, because the structure was revised, or longer, because a clause was added to protect it?

The test applies wherever an explanation has to survive new facts. In medicine, a diagnosis that files each contradicting result as an atypical presentation is lengthening, not compressing, while the patient stays unexplained. In engineering, a root-cause account that adds a special condition for every failure it did not predict has begun cataloguing the system rather than explaining it. In law, a doctrine that needs a fresh distinction for each adverse precedent grows the same way. The move is always the same: protect the conclusion by adding clauses, and call the result understanding.

Knowing this does not place you outside the problem. You compress too, and the urge to rescue a failing thesis with one more clause is not a temptation only other people face. Understanding why descriptions lengthen does not stop your own from lengthening; it only tells you where to look first.

The discipline is in asking: not whether the explanation is short, nor whether it is persuasive, but whether it is compressing reality, or merely hiding the part it cannot explain?

References & Further Reading

  1. The Minimum Description Length Principle: Peter Grünwald

  2. Modeling by Shortest Data Description: Jorma Rissanen

  3. Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance: David Bailey, Jonathan Borwein, Marcos López de Prado & Qiji Zhu

No posts

Read the original on educablemind.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.