RSS Amplifier

Cassandra Unchained · Aug 21, 2026

Trading Post & Short Thoughts to the Max, August 18, 19, 20, 2026

0
Sign in to vote or save

Michael Burry · Cassandra Unchained

Torsten Slok, Apollo’s chief economist, runs a helpful blog The Daily Spark. He posts one chart or set of charts each morning. This morning, he posted a stunningly demonstrative one as part of his blog post, Why Higher Rates Are Slowing the Economy Less Than in Past Cycles.

Torsten summed it up.

When the Fed began hiking in 2022, traditional rate-sensitive sectors, including office construction, rolled over quickly, see chart below.

But data center construction continued to surge as investors and hyperscalers judged that AI-driven returns would exceed the higher cost of capital.

High expected returns and strategic capacity needs have made AI spending largely indifferent to higher interest rates, blunting one of the main channels through which tightening normally slows activity.

Combined with the growth impulse from the One Big Beautiful Bill, the industrial renaissance and prospective tariff refunds, all largely rate-insensitive, we expect growth to remain firm.

The emphasis is Torsten’s. My emphasis is somewhat different.

In the leadup to the Great Financial Crisis (GFC), fraud, leverage, and an economic and labor over-dependence on one factor, national housing price increases, comprised the tinder for the devastating conflagration that followed.

Remarkably, private non-residential construction excluding data centers fell 7.9% year over year in June. GDP growth is really counting on that data center buildout. With the leverage building and fraud growing, plus the complex special purpose vehicle off-balance sheet financing, a relatively new twist involving captive offshore and onshore insurers to soak up risky or fraudulent paper, we are not likely sitting pretty for long.

Recall, the shenanigans are apparent today for those that care to look.

As described in the post above, the Bank for International Settlements, a rather sober global institution, called out the risks in its 2026 annual report, released in June.

Also, a story not often told, and never heard on CNBC, is the degree to which American businesses are moving toward small language inference models built on open source and/or open weight models, regardless of national origin. Cost is the driver.

Journalists worth their salt could talk to the people building these SLMs and get a very good glimpse into the future of compression, and the implications for the massive multi-trillion dollar effort to advance frontier language models. They just do not, however.

I have been hearing this in the Valley for six months. It turns out that on August 7th, researchers Jon Saad-Falcon, Avanika Narayan, and others published Intelligence per Watt: Measuring Intelligence Efficiency of Local AI, a fifth revision of a paper first released on arXiv in November 2025.

Our work makes three primary contributions. (1) We introduce intelligence per watt as a unified metric for evaluating local inference viability, and conduct the first large-scale empirical study measuring its evolution across 1M+ queries, 20+ models, and 8 hardware accelerators spanning 2023-2025. (2) (Q1, Q2) We demonstrate that 88.7% of single-turn chat and reasoning queries can be successfully handled by small local models (with coverage varying by domain), and that IPW has improved 5.3× over two years through compounding model (3.1×) and hardware (1.7×) advances.2 (3) (Q3) We show that hybrid local-cloud routing yields 6080% reductions in energy, compute, and cost compared to a batched cloud baseline; even an 80%-accurate router (a realistic target) captures ∼80% of oracle gains while maintaining answer quality. Together, these findings establish local inference as a practical complement to centralized infrastructure whose viability continues expanding.

“Oracle” as used here has nothing to do with the company. An oracle is a perfect decision maker. In this case a perfect judge routing intelligence tasks to either the local model or to the AI cloud’s LLMs. This study looked at a modeled router/judge, and how the accuracy of the router affects data center workloads. They found perfection is not required of the router.

Read the original on michaeljburry.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.