(Campbell Brown (CEO, PredictHQ) on stage at Expedia Group's Explore, Las Vegas. PredictHQ's demand intelligence saw 147,000 extra visitors and $23M in predicted spend coming months out — the hotels that spot it early capture the rate.)
In April, Uber blew through its entire 2026 AI coding budget in four months. Its CTO burned $1,200 in tokens in a single two-hour demo, and the company capped everyone at $1,500 per tool just to stop the bleeding. We handed every employee an unlimited credit card for intelligence, and the statement just landed.
There are two ways to read a bill like that. One is panic. The other is Mercor, a ten-billion-dollar AI company that now spends more on tokens for its internal agents than it does on payroll, and whose CEO says it with pride. To him, inference crossing the wage bill is what a next-generation company looks like. The number that matters now is the ratio: for every dollar you pay a person, how many go to tokens? So the real question isn’t whether your token bill is climbing. It’s whether it’s climbing on purpose.
Most teams are climbing it by accident. We told ourselves model costs would halve every six months until tokens were basically free. The opposite happened. Chat became agents, agents started spawning agents, runs got longer and more autonomous, and the frontier keeps getting more expensive, not less. The bill compounds in three directions at once.
Here’s the part I’d sit with. The capability gap between the best open models and the best closed ones has closed far faster than the price gap. The intelligence is nearly level; the price isn’t close. Run the maths on a billion input and a billion output tokens a month: roughly $105,000 on the top-end frontier model, about $30,000 on Claude Opus, and around $5,000 on DeepSeek’s V4, which scores within a fraction of a point of Opus on the same coding benchmarks. Same workload, a spread of more than 20 times. Most teams are quietly paying the top of that range for everything, with no eval, no routing, no governance. That spread isn’t a cost. It’s margin you’re handing to someone else.
Everyone reading this is already running on these tools. The only question left is what to do about it. Two moves for your own spend this quarter.
Break your workflows into frontier and not-frontier. Go through what your agents actually do and split it: the planning, the hard reasoning, the reliability-critical calls need the best model. The high-volume, repetitive execution does not. Most teams have never drawn that line, which is why they pay frontier prices for work, a model at a fraction of the cost would do just as well. Frontier for the thinking. Something cheaper for the doing.
Build the plumbing to act on it. A view without the tooling is just a spreadsheet. Put an inference provider or router in front of your stack (we use OpenRouter) so swapping a model is a config change, not a rebuild. Run evals so you know, rather than guess, which model clears the bar for each task. Then offload the non-frontier workloads to smaller or open models and watch the bill fall. That’s how we are trying to do this at Rampersand: use the frontier to work out how to automate a job, then hand the running of it to cheap inference.
Two things follow from all this. The value flows to inference providers (all had big fundraising rounds recently) and open source, the people serving near-frontier intelligence at a fraction of the price. And the opening for founders is the layer that sits on top and tames it.
That opening is already showing up in how companies buy. The scrutiny that capped Uber’s spend is now the first thing a buyer brings to the table: a CFO asking does it work, and what will it cost. That same caution stalls them on build-versus-buy, where they go and build, the costs blow out, and they come back. The wedge for an application-layer founder is to win the deal by letting the buyer skip that detour entirely. If you can show referenceable results and hand them a hard cap on spend, you have answered both questions before they ever start a build. That is what the routing, evals, and governance layer buys them: it decides which model does what and proves the spend was worth it. Capped, token-based pricing stops being a liability and becomes the whole pitch. A harness for spend. You are not selling software, you are selling control, and control is what every CEO who just opened their bill now wants.
If you’re building something in that smart routing layer or anything equally interesting, I want to talk. Reach out.
What happened: Anthropic released Claude Fable 5, a guardrailed, generally available version of its Mythos-class model. It runs on the same underlying model as the restricted Mythos 5. The benchmarks are emphatic: 80.3% on SWE-Bench Pro versus 69.2% for Opus 4.8, 58.6% for GPT-5.5, and 54.2% for Gemini 3.1 Pro, with the lead widening the longer and more complex the task. Stripe says Fable compressed five months of engineering into days, finishing a migration in a 50-million-line Ruby codebase in one day that would have taken a team over two months.
“Qualitatively, this is a major-version-bump-deserving step change forward.” — Andrej Karpathy
Why this matters: For builders, the real shift is task length. When a model reliably runs long, ambitious loops, the unit of work moves from “fix this bug” to “own this outcome.”
What happened: The headline at Apple’s WWDC26 was an AI model that runs entirely on the phone itself, with nothing sent to the cloud. A model this capable would normally need a data centre. Apple’s trick is to keep the whole thing, around 20 billion parameters, in storage and load only the small slice it needs for each request, because a phone doesn’t have the memory to run all of it at once. It now handles images and photos, not just text. The catch: it only runs on the newest hardware, the iPhone 17 Pro and recent Macs, and it is not available in the EU or mainland China yet. Apple built the family with Google, and its bigger cloud companion model is based on Google’s Gemini.
Takeaway: If your product pays per-call API fees for things a 1–4B model can handle (dictation, voice, basic image understanding), model an on-device version. The Foundation Models framework now accepts images, so the on-device tier is a real, free, private distribution channel.
What happened: A Stanford and Together AI paper proposes “intelligence per watt” (accuracy per unit of power) and tests 20+ local models against a million real queries. The share of queries where a local model matches a frontier API rose from 23.2% in 2023 to 71.3% in 2025, roughly a 3.1x jump in two years. On single-turn chat and reasoning specifically, local models now answer 88.7% accurately.
Takeaway: The gap closed 3x in two years, so the model you picked as your default last year is probably the wrong one now. The harder idea is the unit: as intelligence moves on-device and toward open weights, the binding constraint stops being dollars and becomes watts. That is why energy (see Apple above, and the data-centre item below) is quietly becoming the thing that decides who can serve AI at scale.
What happened: Antares Nuclear brought its Mark-0 microreactor to initial criticality at Idaho National Laboratory, the first novel reactor design to achieve criticality at the lab in more than 50 years and the first private company to do so under the US Department of Energy’s Reactor Pilot Program. Mark-0 is the zero-power forerunner to Antares’s flagship R1, a sodium heat-pipe-cooled design fuelled by HALEU TRISO fuel. Zero-power criticality means the chain reaction is self-sustaining but not yet producing meaningful electricity. Antares, founded in 2023, has committed to electricity production in 2027 and a deployed reactor by 2028. Reaching criticality just three years from founding is a pace almost unheard of in a sector defined by decade-long delays.
Why this matters: The bigger tell is speed: the US fast-tracked Antares through a federal DOE pathway that skips the usual decade-long regulatory queue. Moving fast on nuclear is now a deliberate national choice. China now has roughly as many reactors under construction as the rest of the world combined and builds them in about six years against a ten-year global average, and the US has just signalled it will move faster too. Australia has a federal ban on nuclear power and AI-driven energy demand that keeps rising. The honest question, and we don’t pretend to have the answer, is what fills that gap: more renewables and storage, more gas, imported power, or a rethink of the ban?
Spotify Co-CEO Gustav Söderström gave David Senra a rare look at how the company operates, and the operating model is the part founders should steal. Spotify runs what he calls a synchronous model: senior leaders meet together every week for hours, work through every part of the business in one room, and are banned from saying “let’s take it offline.” The goal is to give every leader a CEO’s view of the whole company rather than a siloed view of their function, so calls get made with full context and in real time.
Two ideas sit underneath it:
You ship your org chart. Structure leaks into product, so design the org around the outcome you want, not the other way round.
Marry new tech to a contrarian business model. Technology alone rarely changes anything; the shift comes when you pair it with a model others won’t copy. Spotify’s “Time Well Spent” bet, optimising for users not regretting their time rather than for raw engagement, is why subscriptions anchor the business. Anonymous surveys found Gen Z valued roughly 90% of their time on Spotify, while regret rates on competing platforms topped 60%.
Takeaway: Audit your leadership rhythm. If your exec team meets only in functional silos and defers hard calls to “offline” follow-ups, you’re training people to optimise their patch, not the company. Put them in one room and ban “take it offline.”
What happened: NewLimit, the epigenetic reprogramming company co-founded by Coinbase CEO Brian Armstrong, raised a $435 million Series C led by Founders Fund at a $3.1 billion valuation, and confirmed its first-in-human trial will dose patients in Australia rather than the US. CEO Jacob Kimmel’s reasoning is the part worth noting for the ecosystem: Australia’s decentralised model lets individual hospital ethics committees (HRECs) approve trials, getting a first-in-human study into patients faster than the American pathway. It’s the same advantage that has quietly made Australia a hub for early-phase trials, now validated by one of the most closely watched biotechs in the world.
Takeaway: Australia’s clinical-trial regime is a real, underrated moat for the ecosystem. For ANZ founders in biotech, medtech, or health AI, “run your first trial here” is becoming a genuine pitch, not a consolation prize.
What happened: Supabase, the open-source Postgres platform co-founded by New Zealander Paul Copplestone, raised a $500 million Series F at a $10.5 billion valuation, led by Singapore’s GIC. That roughly doubles its valuation in eight months, riding the vibe-coding wave: database launches grew more than 600% year over year, over 60% of new ones are now created by some kind of AI tool, and Copplestone says Claude Code is the single largest contributor in 2026. The platform now claims close to 10 million developers. Notably, he got here partly by turning down multimillion-dollar enterprise deals that would have dictated the roadmap, betting on developer love instead.
Takeaway: The durable money in the AI coding boom may sit in the infrastructure layer, not the models, and an ANZ founder is proving it at global scale. Refusing early enterprise capture is also what let a developer-first product compound into a decacorn.
What happened: As Sam Altman talks up Australia as a potential “data centre capital of the world,” the backlash has fixated on water. Michael Vardon’s numbers in The Conversation suggest that’s the wrong worry: a megalitre of water in a data centre generates about $2.3 million in value, versus roughly $4,600 in agriculture. The real constraint is energy and location, with demand threatening to outstrip renewables and pull new gas plants onto the grid.
Takeaway: For founders chasing the AI infrastructure wave, the bottleneck isn’t water, it’s power and placement. That’s where the opportunity, and the policy gap, sits.
Parachute, the AI operating system for SME law firms, has gone live with an integration that pipes InfoTrack’s verified property and corporate data straight into its AI workflows. Lawyers can now run InfoTrack searches inside a Parachute matter without switching tools, grounding AI-drafted advice in sources a firm can stand behind rather than guesswork. It tackles the trust problem head-on: AI is only useful in legal practice if its outputs sit on verified data, and the integration also captures search costs at the point of work. Co-founder Ryan Zahrai and InfoTrack’s Brendan Smart walked firms through what it means in a 0.5 CPD webinar this month.
Mass Dynamics brought its agent-ready omics tooling to ASMS 2026 in San Diego, where AI was the dominant theme across the proteomics community. The standout was a benchmark from CSO Andrew Webb tackling a question every scientist using LLMs should be asking: can AI generate scientific visualisations that are reproducible, statistically correct, and trustworthy? Built on real-world proteomics tasks, it scored leading models on statistical correctness, structural accuracy, visual fidelity, and reproducibility. The finding: modern agents are remarkably capable but still need rigorous evaluation and domain-specific guardrails before they can be trusted in scientific workflows. As the team framed it, the future isn’t AI replacing scientists, it’s AI helping them move faster while preserving the standards that make science trustworthy.
Restoke, the AI platform for hospitality operations, featured in realcommercial.com.au as Johnny di Francesco’s Gradi Group detailed how it swapped a tangle of spreadsheets for one connected system covering recipe costs, stock, and COGS. The platform updates dish costs automatically as ingredient prices move, giving operators a real-time read on what each plate actually costs to produce. In an industry running on thin margins, that visibility is the difference between pricing with confidence and spotting a problem only after it has eaten into profit. Restoke now serves more than 2,000 venues across Australia, New Zealand, Singapore, and the US.
Location & type: Surry Hills, NSW, Australia, full-time
What they do: Keeyu monitors e-commerce post-purchase journeys to catch failures before they turn into churn.
What you’ll do: Build the agentic systems that detect and act on post-purchase failure signals, turning monitoring into automated intervention.
Location & type: Port Melbourne, VIC, hybrid, full-time
What they do: Restoke is an AI platform that automates hospitality operations, from recipe costing and inventory to supplier orders and team management.
What you’ll do: Own customer relationships after the sale, driving adoption, retention, and expansion across the restaurant groups on the platform.
Location & type: Melbourne, VIC, hybrid (remote + in-office), full-time
What they do: Mass Dynamics builds a Melbourne-based proteomics platform that makes mass-spectrometry analysis accessible and agent-ready for life scientists.
What you’ll do: Bridge bioinformatics and product, shaping how proteomics workflows show up in the platform for researchers.
Do you have a job you’d like us to promote? Add it to Hatch and share the link with us.
See all jobs across the Rampersand Portfolio.
Paul Naphtali will be in Sydney the week of 17 June. Reach out for a coffee.
📍 BYRON BAY — Jun 11, 2026
A relaxed evening at Lord Byron Distillery on how people are really putting AI to work day to day, 6:00–9:00pm. Andrew Poesaste will be there.
Thanks for being part of the Rampersand Community! You can stay updated with even more news and info on our LinkedIn and X/Twitter pages.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.