As intelligence becomes abundant, knowing what questions to ask will matter more than ever.
This week, I want to share two: one I have discussed recently, and another that prompted some interesting responses from others (my What I Read This Week is still below).
If you find this useful, let me know in the comments, and we will do more every week.
As always, don’t just take my answers or anyone else’s. Think and then answer these questions for yourself.
Share your answer in the comments; we will repost the sharpest take.
Question 1) With massive AI bills hitting enterprises, how should companies think through maximizing ROI for their token costs?
Most CEOs and CFOs have no idea how much tokenmaxxing is happening within their own companies, and it will show up in the numbers.
At 8090, my CTO told me our token costs are doubling roughly every 45 days, and the incremental productivity we get from that doubling is maybe 5-10% at most.
Now run that against a Fortune 500 company that said it is all in on AI. The spend can quietly compound inside OpEx, and suddenly a quarter gets missed by a few pennies of EPS. The CEO would ask the CFO where the money went, and they would likely trace it back to the $56-per-million tokens of intelligence, when a $0.50 version could have done the job.
So how do you not become that company?
First, stop paying for the flagship where you don’t need it. The cheap models are now 80% to 95% as good as the frontier models on most tasks. If Pepsi is a fraction of the price of Coke and tastes almost the same in the dark, you at least have to manage that as a risk. Pick the model per task, route it through a control plane that gives you optionality, and pay frontier prices only for the few jobs that genuinely need frontier intelligence.
Second, watch where your data goes. When you pipe your proprietary workflows and your competitive logic through a closed frontier model, you are renting intelligence and giving away your edge to providers that may also want to compete with you.
Measure the output you get, and treat intelligence like any other input cost.
Question 2) If every smart enterprise starts to pull its AI in-house to protect its edge, who’s left to pay for the $1.4T buildout?
Gavin Baker’s answer:
Gavin Baker@GavinSBaker
The mega bull case for AI infrastructure would be *if* market share shifted away from certain frontier labs with 90%+ inference margins toward cheaper models, whether open-source or closed. It would increase the ROI on AI spend for end customers by increasing intelligence per
Cassandra Unchained @michaeljburry
This is true as I have heard this from contacts in the Valley. Goes with my pinned post. The AI race is shifting from bigger models to cheaper, smarter systems https://t.co/lS1YKxkjl0
6:15 PM · Jul 12, 2026 · 3.46M Views
419 Replies · 1.2K Reposts · 6K Likes
Replit CEO Amjad Masad’s answer:
Amjad Masad@amasad
@GavinSBaker Paradoxically is as frontier LLMs getting better might lead to less usage because they will enable enterprises to do their own ML research and move away from expensive models and distill known use-cases into smaller models. Programming languages might be a useful analogy. You
10:02 PM · Jul 12, 2026 · 15.2K Views
8 Replies · 10 Reposts · 154 Likes
1) Open-Weight AI Reaches the Coding Frontier
On July 16, Chinese lab Moonshot AI released Kimi K3, a 2.8-trillion-parameter model that took first place on the Frontend Code Arena, a leaderboard scored by blind developer votes on the web apps models generate. K3 reached 1,679 points, ahead of the closed flagships Claude Fable 5 and GPT-5.6 Sol, and led six of the seven frontend categories.
On broader intelligence tests it still trails both, but at $3 and $15 per million input and output tokens it undercuts Fable 5’s $10 and $50, and Moonshot plans to release the full weights by July 27, which would make it the largest open-weight model yet. Until then, every K3 spec is Moonshot’s own claim.
The same push is coming from US startups selling the means to build and own AI rather than rent it.
On July 15, Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, released Inkling, its first model, with open weights under an Apache license that lets anyone download and modify it.
Inkling is a 975-billion-parameter mixture-of-experts model that activates only about 41 billion parameters per request, which makes it cheaper to run. Thinking Machines describes Inkling as a base for customization rather than the strongest model available, and its full weights are free to download. The company’s paid product is Tinker, a fine-tuning tool that adapts models to a customer’s data. Bridgewater’s AIA Labs published a study that used Tinker to fine-tune an open model for financial document triage. The custom version scored 84.7% accuracy, compared to 78.2% for the best frontier model tested, at roughly one-fourteenth of the cost per task.
A week earlier, on July 8, Prime Intellect raised a $130 million Series A at a $1 billion valuation. Prime Intellect rents decentralized computing and open-source software that let a company train its own agents, reporting a $100 million annualized revenue run rate. A case study with Ramp showed their post-trained model beating Opus at spreadsheet search, and running 27% faster and far cheaper than Haiku.
2) Ex-SpaceX Engineers Raise $115M to Revolutionize Construction
On July 14, TerraFirma raised about $115 million, including a $100 million Series A led by Kleiner Perkins, and said it will hire 300 people and build a Texas factory. It was founded in 2024 by former SpaceX engineers Noah Schochet and Noah McGuinness, who met at Princeton and worked on Starlink and Starship.
The company retrofits heavy construction machinery, such as excavators, dozers, and loaders, into robots that a skilled operator operates from a screen rather than climbing into the machine. Using an interface it calls “click to dig”, an operator sketches the required work as a 3D model in about 60 seconds, and the machine runs the task for more than 20 minutes on its own, with an Xbox controller kept as a manual backup. One operator can currently handle three to five machines, which TerraFirma says makes each operator up to 300% more effective and the work safer, since no one has to sit in the machine.
US construction labor productivity has fallen about 0.6% every year since 1965 even as productivity across the broader economy has grown at roughly 1.6% annually. It’s estimated that the industry needs 349,000 more workers in 2026 just to keep pace. TerraFirma runs its own projects rather than selling the technology, which gives it live job sites and field data. The field is heating up as Caterpillar and Komatsu already offer remote control, and Bedrock Robotics raised $270 million to field operatorless excavators this year.
The New Private Asset
Jul 9
Over the past year, while every category of crypto token lost value, privacy tokens climbed 127.3%, becoming the best-performing sector in crypto. To understand why privacy coins are surging, we need to explore...
Making Fable Cheaper Than Opus (Joon Lee)
You Just Hired a Million Bad Employees (George Sivulka)
A Framework for Frontier AI and the Dawning of a New Age (Demis Hassabis)
The Reverse Information Paradox (Satya Nadella)
Tarek Mansour@mansourtarek_
Today, we launched GPU compute forward curves derived from our prediction market prices. Forward curves are now available on Nvidia B200. H200, and A100 chips. Forward curves track implied future prices. They are how mature commodity markets form expectations, allocate capital,

10:50 PM · Jul 14, 2026 · 1.23M Views
197 Replies · 222 Reposts · 2.02K Likes
Watcher.Guru@WatcherGuru
JUST IN: Stripe and Advent offer to buy PayPal $PYPL for $53 billion, Reuters reports.

3:38 AM · Jul 15, 2026 · 719K Views
273 Replies · 451 Reposts · 4.5K Likes
C3@C_3C_3
Nick Shirley was correct. NPR was wrong. They were off by only 1,000%. $110,000 to $110,000,000 Lie loudly then correct quietly. The Legacy Media is the enemy of the People.

5:29 PM · Jul 14, 2026 · 397K Views
261 Replies · 4.27K Reposts · 27.4K Likes
Haider.@haider1
i knew Codex was growing fast, but this is unbelievable: 1m → 8m users in just over five months even crazier, it went from 6m → 8m in only three days

Tibo @thsottiaux
Tomorrow might be 8M active user celebration day. Just saying
9:34 AM · Jul 14, 2026 · 429K Views
211 Replies · 225 Reposts · 5K Likes
Jukan@jukan05
KOREAN LOCAL MEDIA: SAMSUNG FOUNDRY HAS SECURED ANTHROPIC AS A FOUNDRY CUSTOMER AN INDUSTRY SOURCE CLOSE TO SAMSUNG ELECTRONICS SAID, “SAMSUNG ELECTRONICS’ FOUNDRY DIVISION HAS AGREED TO MANUFACTURE ANTHROPIC’S CUSTOM AI CHIPS.”
7:12 AM · Jul 14, 2026 · 262K Views
84 Replies · 116 Reposts · 1.19K Likes
zerohedge@zerohedge
Hyperscaler cash flow: 12 month fwd

1:19 PM · Jul 13, 2026 · 284K Views
60 Replies · 189 Reposts · 1.19K Likes
Mary Talley Bowden MD@MaryBowdenMD
Over the past 34 years, healthcare has replaced manufacturing as the top employer in the vast majority of states.

3:05 PM · Jul 11, 2026 · 3.7M Views
785 Replies · 1.81K Reposts · 7.05K Likes
Palmer Luckey@PalmerLuckey
Everyone who thinks AI slop will ruin code efficiency/performance is going to be so surprised when everything is absurdly well-optimized John Carmack style machine code.
2:21 AM · Jul 15, 2026 · 912K Views
602 Replies · 664 Reposts · 14.6K Likes
The Kobeissi Letter@KobeissiLetter
BREAKING: US officials say shipments of Nvidia, $NVDA, H200 chips to China have begun. Nvidia shares extend gains to a new high of day on the news.
2:57 PM · Jul 14, 2026 · 446K Views
126 Replies · 250 Reposts · 3.03K Likes
tobi lutke@tobi
A huge amount of the Anti-AI code sentiment massively overestimates the quality of human code outside of a very small set of open source and high quality company codebases. Human Slop is everywhere and can trivially be improved on by any opus level model.
9:07 PM · Jul 16, 2026 · 820K Views
422 Replies · 738 Reposts · 9.56K Likes
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.