Some posts write themselves.
I was out for a walk recently, listening to an episode of BiggerPockets Money featuring Ben Felix. I’m a big fan of his - not because I agree with everything he says, but because he consistently forces me to examine my assumptions. If he says something that doesn’t fit my priors, I usually go back and check my priors.
So I did a double-take when I heard this exchange about portfolio construction and the value of backtesting.
The context was a discussion about Risk Parity-style portfolios versus a simple all-equity portfolio. Co-host Scott Trench asked what I think is a very good question: if a diversified portfolio appears to produce higher withdrawal rates in historical testing, why shouldn’t investors care about that?
Felix’s response was essentially that he would put “very, very little faith” in those kinds of backtests because future correlations are unknowable, fees matter, and future returns are uncertain.
To listen for yourself, here is the YouTube video, and the exchange I’m talking about starts at 38’32”:
My immediate reaction was: Ummm… what?
Not because I completely disagree with the conclusion. In fact, I agree with part of it. Investors absolutely should be cautious about drawing sweeping conclusions from backtests.
But the reasoning felt surprisingly thin coming from someone whose entire public persona is built around evidence-based investing.
The first issue is that all investing assumptions rely on historical data to some degree. If backtests are largely untrustworthy, then what exactly is the basis for believing equities will outperform cash over the long run?
The rest of the episode was essentially an argument for holding a highly concentrated equity portfolio. But where does confidence in equities come from if not decades of observed historical performance? It seems odd to dismiss portfolio backtests while simultaneously relying on the historical record of stocks to justify a stock-heavy allocation.
Perhaps the argument is that we can trust the history of individual assets but not the history of portfolios composed of those assets. But then the question becomes: why? What magical property suddenly makes historical evidence unreliable once multiple assets are combined together?
Felix is absolutely correct that correlations are unstable.
The stock-bond relationship that many investors had come to expect broke down in spectacular fashion during 2022. Stocks and bonds both struggled at the same time, and diversification failed when investors wanted it most.
But returns are unstable too.
Stocks have delivered wildly different results across different decades. Value stocks outperform for long stretches, then lag for years, yet Felix is firmly in the value investing camp. International stocks go through periods of dominance and irrelevance, and Felix consistently argues for holding global equities.
Nobody—not even Felix—would argue that because a factor or asset class underperformed for a few years, we should permanently abandon it. We understand that expected returns are estimated from large datasets, and the same should apply to correlations.
The more I thought about it, the more it seemed that the conversation had skipped over the strongest arguments for diversification altogether.
The case for combining non-correlated assets isn’t built on a handful of attractive backtests. It’s built on theory and practice.
The intellectual foundation goes back to Harry Markowitz and Modern Portfolio Theory itself. Markowitz’s work on diversification, which I summarized here, was important enough to earn a Nobel Prize. The idea that combining imperfectly correlated assets can improve portfolio outcomes isn’t some fringe internet theory.
Then there is the work of people like Claude Shannon and Antti Ilmanen, who explored how rebalancing and diversification can create value over time.
If theory isn’t enough, there is plenty of real-world evidence putting these theoretical ideas into profitable practice.
Ray Dalio built one of the largest hedge funds in history around diversification principles. More broadly, the idea of combining multiple return streams with low correlations is a cornerstone of institutional portfolio construction worldwide.
That doesn’t prove that creating diverse portfolios of multiple asset classes is optimal. But it does suggest the conversation deserves more than a quick dismissal of backtests.
This is where I think the discussion becomes genuinely useful.
Backtests are neither crystal balls nor worthless exercises.
They’re tools. Used badly, they’re dangerous. Used correctly, they’re incredibly informative.
The mistake many investors make is treating a backtest as a prediction.
They run a portfolio from 1995 to today, see an 11% CAGR, and assume they’re about to earn 11% going forward. That’s obviously nonsense. The future will be different.
But that doesn’t mean the exercise was useless.
The real value of backtesting is comparison.
Suppose you keep everything constant and change only one variable:
Replace intermediate Treasuries with long Treasuries.
Add gold.
Remove international stocks.
Add managed futures.
Now you’re learning something. You’re not predicting future returns. You’re observing how a portfolio historically behaved when a particular ingredient was added or removed.
That’s a very different exercise.
The longer the dataset, the more useful the exercise becomes.
Personally, I feel much more comfortable when I can test a portfolio across roughly fifty years of history.
That allows the portfolio to experience:
The inflationary 1970s
The interest-rate spike of the early 1980s
The 1987 crash
The dot-com bubble
The Global Financial Crisis
The COVID crash
The strange environment of 2022
No, the future won’t look exactly like those periods.
But if a portfolio can survive all of them, that’s valuable information.
For example, one of the reasons I prefer extended-duration Treasuries over intermediate-duration Treasuries is because historical backtests consistently show the tradeoff involved.
Higher duration has generally meant:
Higher returns
Higher Perpetual Withdrawal Rates
But also
Higher volatility
Lower Safe Withdrawal Rates
The backtest doesn’t guarantee those relationships will persist, and it doesn’t mean the investor will receive the same returns going forward. It also doesn’t imply that longer duration is the best choice for everybody.
But it does help us understand what role duration, for example, has historically played inside a diversified portfolio, and it does help investors to figure out what portfolio might generally be better for them.
And that’s useful.
Backtests shouldn’t be the beginning of the portfolio design process.
They should be near the end.
First:
Learn the theory.
Understand the asset classes.
Read opposing viewpoints.
Think about your goals.
Then:
Use backtests to see how those ideas behaved historically.
That’s a very different mindset than hunting through thousands of possible combinations until you find the prettiest chart.
So in one sense, I agree with Felix. Investors shouldn’t put blind faith in backtests. But neither should they dismiss them.
Historical testing isn’t a substitute for thinking. It’s one of the tools that helps us think better.
And if we’re willing to trust decades of historical evidence when discussing the expected return of stocks, it seems strange to suddenly become skeptical of history when discussing the behavior of diversified portfolios.
The trick isn’t to worship backtests.
It’s to understand what questions they’re actually capable of answering.
Thanks for reading Risk Parity Chronicles! This post is public so feel free to share it.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.