Say you inherit a building. Great location, solid bones, zoning already approved for hospitality. You furnish the rooms, install the plumbing, hang the signage. By any reasonable measure, you do not own a hotel.
What you own is a building shaped like a hotel. A hotel is what happens when someone answers the phone at 11pm, cleans a room between guests, runs a reservation system that does not double-book the same suite, and sends a bill that matches what the guest actually used. None of that comes bundled with the real estate.
This is exactly the situation power owners are walking into right now, except the building is a data center and the guests are GPU workloads.
Picture a family energy company in West Texas. Three generations, mostly gas peakers, a wind stake from a decade back. They have 40 megawatts that never gets fully monetized because the grid does not always want what they can produce, when they can produce it.
Then the calls start. A broker wants to lease the pad. A reseller wants to talk GPU orders. Somebody’s cousin in private equity sends a model showing what an H200 rents for per hour, times 8,000 hours a year, and the number at the bottom is large enough that the room goes quiet.
So they run the math. Power, check. Land, check. Shell and cooling, solvable. Every physical line item has a vendor and a quote.
Then someone asks the only question that matters. What are we selling, exactly?
The building is not the business. Nobody pays to stand in the lobby. The customer never touches your substation, they touch an API, a console, a login, and an invoice, and everything between the power and that invoice is the part nobody quotes you for.
Demand for accelerated compute is running ahead of what can be brought online, and the good large-scale capacity is mostly spoken for. That gap is pulling capital toward whoever controls power, which is why power-site developers and well-capitalized individuals are suddenly being courted as future cloud operators.
They are right about the scarcity. Power is the binding constraint.
The wrong conclusion is the arithmetic that follows: power plus hardware equals cloud platform. It gets you a depreciating asset with a clock running on it and no mechanism for turning it into revenue. It gets you infrastructure, not a service.
If a human touches a machine every time a customer signs up, you have built colocation with worse margins. The eleventh customer, on a Saturday, will find out.
RDMA fabrics and shared storage mean two tenants sharing hardware is not like two VMs on a hypervisor. Getting it wrong is not a performance issue, it’s a disclosure event.
When a tenant releases a node, something has to guarantee GPU memory and firmware state are clean before the next tenant lands. This is the likeliest way one customer’s model weights leak into another’s environment.
An idle GPU depreciates on schedule regardless of use. Fractional allocation (MIG, time-slicing) is what lets you sell two GPUs to a customer instead of turning them away because your smallest unit is a full node.
Graphing GPU usage gets you a fifth of the way there. Between a metric and an invoice sit rating rules, SKUs, quotas, and a number that survives a finance team’s scrutiny.
You likely have three sites of 8-15 megawatts, not one campus of 300. Customers won’t tolerate three consoles and three support paths.
These aren’t exotic problems. They’re the ordinary cost of being an operator, and hyperscalers spent a decade and thousands of engineers learning them.
The physical asset at the bottom of this stack and the top is identical. Same power, same GPUs. What changes as you climb is how much software sits between your hardware and your customer, and with it, who owns the brand, the price, and the relationship.
Rung one is a real business. Leasing power is fast, low risk, and produces revenue while you’re still learning what a tenant even is. Some owners should stop there.
But understand the trade. At rung one you’re infrastructure supply for someone else’s platform. Your upside is capped at the lease terms, and the customer belongs to them.
Lease the power, host someone else’s gear
✅ Fast revenue, no operational risk.
❌ Fails if your upside stays capped while compute services get more valuable around you.
⚖️ You give up the customer relationship and any path upward without starting over.
Rent raw GPU capacity yourself
✅ Fastest way to be a real operator.
❌ Fails when a competitor with cheaper power does the same thing and price is the only lever left.
⚖️ No differentiation. GPU hours are a commodity.
Build the control plane in house
✅ Wins with a real platform team and a multi-year horizon.
❌ Fails on timing: 12-24 months of building is 12-24 months of GPUs depreciating without revenue.
⚖️ You give up time-to-first-dollar, the expensive thing in this market.
Buy the control plane, operate the platform yourself
✅ Wins when you want the customer and would rather spend engineering time on service design than rebuilding tenant isolation.
❌ Fails if you treat the platform as magic and skip building real operations around it.
⚖️ Some architectural freedom, one vendor dependency.
Rafay is your stack
Two clocks start the moment your hardware is racked. One is depreciation, running from day zero. The other is revenue, which doesn’t start until a customer can self-serve something. Every month spent building provisioning and metering from scratch is a month where only the first clock runs.
This is where we need a platform layer. Rafay is one of the products that packages the middle column above: bare metal provisioning, GPU lifecycle and reclamation, multi-tenancy, MIG and time-slicing, metering that connects to billing, and a single control point across sites that aren’t physically together. What matters is that this layer now exists as a purchasable product, which wasn’t true at all five years ago. That’s the difference between first revenue in a quarter versus a fiscal year.
Want to learn more about the solutions? Attend the Rafay AI Infrastructure Leadership Summit in Barcelona Sept 8-10 for free.
If you have power but no site or capital, stay at rung one and revisit in 18 months. If you have power and financing but zero technical staff, aim for rung two or three on a bought control plane. If you have a platform team with genuinely unusual requirements, build only the pieces that are actually unusual and buy the rest. If you run three or more small sites, treat federation as a hard requirement, not a nice-to-have, since nothing that treats each site as an island will work. If you want to reach inference or token-metered services eventually, choose your foundation for that now, because retrofitting metering and multi-tenancy later is a rebuild, not an upgrade. And if you want to stay a landlord, that’s a real answer too, as long as it’s a decision and not a default.
The opportunity isn’t to supply capacity to someone else’s cloud. It’s to decide, deliberately, how far up that stack you intend to own. Most people won’t decide. They’ll drift into rung one because it’s the option that arrives with a contract already drafted.
This blog post is sponsored by Rafay, thanks for partnering with me on this post
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.