RSS Amplifier

Gal Ratner · Aug 9, 2026

I Priced a Rack at Switch. The Cloud Won.

0
Sign in to vote or save

Gal Ratner · Gal Ratner

A client came to me earlier this year with the proposal already half-built. Mid-market operation, a little under two hundred people, a line-of-business .NET application that has been earning revenue since before anyone at the company said the word “cloud” out loud. SQL Server on the back end. A document store. A handful of internal services and a reporting stack. No AI. No machine learning. No GPUs. Just software that works, running on infrastructure that costs money every month.

Their CFO had looked at the Azure bill and had the reaction every CFO has. Someone in the building said the word “colocation.” Someone else pointed out that Switch is right here in Las Vegas, on Decatur, an hour from most of the leadership team, and that you can walk into that building and touch your own hardware. By the time the question reached me it was not really a question anymore. It was a plan looking for validation.

So I ran the exercise properly. Not the vendor version, where you compare a reserved instance list price against a server you already depreciated. The version where you count everything, including the things nobody wants to put in the spreadsheet because they have a person’s name attached to them.

The cloud won. Not narrowly, not with an asterisk, and not because I like the cloud. It won on raw cost, and then it won again on every axis that was not cost. This is a company with zero AI workload, entirely predictable traffic, and a stable footprint. This is supposed to be the exact profile where on-prem wins. It did not.

Here is how that happened, where it stops being true, and what changes when you put AI into the picture.

The first spreadsheet is always wrong, and it is always wrong in the same direction. Somebody prices out the servers, prices out a cabinet, multiplies the monthly cloud bill by thirty-six, and the CapEx number looks like a rounding error against the OpEx number. The meeting ends early. Everyone feels smart.

That spreadsheet is comparing a purchase to a rental and pretending the purchase ends at the purchase. It counts the servers and the rack. It does not count the second set of servers you need so the first set can fail. It does not count the switches, the firewalls, the out-of-band management, the KVM, the PDUs, the cross-connects, the transit commit, the spare drives sitting in a box, the backup target that has to live somewhere else entirely, or the fact that a rack full of gear with no disaster recovery site is not a data center strategy, it is a single point of failure with a security badge.

It also does not count power correctly, and power is where colocation quietly eats you alive. Space is not the constraint in this market anymore. As of 2026, power availability rather than floor space is the binding constraint on new colocation supply in every major U.S. market. Wholesale colocation in primary North American markets was running roughly one hundred ninety-six dollars per kilowatt per month for deployments in the two hundred fifty to five hundred kilowatt range, and small single-cabinet deployments price worse than that per unit, not better. Switch will happily sell you a cabinet that supports up to fifty-five kilowatts. You will pay for the envelope you reserve, not the watts you happen to draw on a slow Tuesday.

Then there is the refresh cycle. Hardware you buy in 2026 is hardware you are replacing in 2029, or you are running a production database on a box that is out of vendor support and one failed backplane away from a very bad week. Amortize honestly over a real replacement horizon and the CapEx number stops looking like a rounding error and starts looking like a recurring cost with worse terms.

The largest cost in the on-prem column was not hardware. It was people.

Somebody has to patch the hypervisor. Somebody has to own firmware and BIOS updates across the fleet, which is genuinely miserable work that nobody has ever been promoted for doing well. Somebody has to test restores, not schedule backups, actually test restores, which is the only part of a backup that matters. Somebody has to be reachable at two in the morning when a drive fails, and somebody has to drive to Decatur when remote hands cannot resolve it. Somebody has to renew the certificates, rotate the keys, manage the VPN concentrator, and answer the auditor’s questions about physical access controls.

At this company that somebody was going to be one and a half people who did not exist yet, or three existing people who would each lose a third of their week to infrastructure and stop shipping features. Fully loaded, in this market, that was the single biggest number on the page. It was larger than the hardware. It was larger than the rack.

And this is the part where the on-prem advocates say that 37signals repatriated without adding headcount, which is true, and which I will come back to, because it is the strongest argument against everything I just wrote and it deserves better than a dismissal.

Even if the labor math had been a wash, the timing is indefensible.

The memory market has come apart. AI datacenter demand pulled DRAM and NAND capacity toward high-margin server parts and left everything else fighting for scraps. DRAM contract prices rose roughly fifty percent through 2025 and kept climbing, with TrendForce showing Q1 2026 contract prices up something on the order of ninety to ninety-five percent quarter over quarter. NAND overtook DRAM as the fastest-rising component in Q2 2026, with contract prices up seventy to seventy-five percent quarter over quarter after a fifty-five to sixty percent rise the quarter before. Memory is now as much as a quarter of the bill of materials on a typical server, and considerably more on the memory-dense virtualization hosts you would actually buy for this workload.

The downstream effect is exactly what you would expect. A server build that ran eight thousand dollars in Q4 2025 runs north of nine thousand two hundred today. Dell and Lenovo have both pushed through double-digit enterprise price increases. Cisco raised prices on compute and on anything containing memory. Production capacity for 2026 was already sold out with hyperscalers moving into 2027 negotiations, which means you are not at the front of that line and neither is your reseller.

Procurement got worse in ways that do not show up in a price comparison at all. Quote validity at many distributors has collapsed from thirty days to four or seven. Intel server CPU lead times stretch to six months, AMD runs eight to ten weeks, and a standard PowerEdge configuration that used to land in two or three weeks now takes eight to twelve. You are being asked to commit capital, at a cyclical price peak, for equipment that arrives a quarter late, to serve a capacity forecast you made before any of that was true.

Meanwhile the cloud provider absorbed that entire supply shock on your behalf and did not change your per-hour rate.

There is a real case on the other side and I am not going to strawman it, because the numbers are public and they are good.

37signals left AWS and documented every dollar. Roughly three point two million a year in cloud spend. A seven hundred thousand dollar Dell purchase that dropped cloud bills by about two million a year. Then eighteen petabytes moved off S3 onto a one and a half million dollar Pure Storage buy that costs under two hundred thousand a year to operate, replacing a one and a half million dollar annual S3 bill. AWS waived a quarter million in egress fees on the way out. Total infrastructure spend went from three point two million to well under a million, and David Heinemeier Hansson has been explicit that they did it without adding staff. Broadcom found forty to fifty percent lower total cost of ownership on private cloud for steady-state workloads. GEICO cut compute cost per core by half. Something like eighty-six percent of CIOs now say they plan to move at least some workloads off public cloud, the highest rate anyone has recorded.

That is not noise. That is a real signal and anyone selling you cloud-first as an ideology is lying to you.

But look at what actually makes that case work. 37signals runs at a scale where a two million dollar annual delta funds an entire infrastructure practice. They have deep, unusual, in-house operational expertise, and their CTO is personally invested in the problem as a matter of professional conviction. They own the product, so infrastructure was a product decision they were qualified to make. They migrated app by app over years, not in a quarter. And critically, their workload is enormous, stable, and homogeneous, which is the ideal shape for owned hardware.

My client had none of that. A hundred and eighty people, a six-figure annual cloud bill rather than a seven-figure one, no infrastructure practice, and a development team whose comparative advantage is domain logic in C#, not BGP. At 37signals scale the savings buy you a team. At mid-market scale the savings do not even cover the team, which means the savings are not savings, they are a cost transfer onto people who were doing something more valuable.

And the aggregate market data cuts the same way. Synergy put global cloud infrastructure spending at roughly a hundred twenty-nine billion dollars in Q1 2026 alone, up thirty-five percent year over year, with growth accelerating for nine consecutive quarters. Repatriation is real and cloud growth is accelerating at the same time, because they are happening to different workloads at different companies. Repatriation is a workload decision, not a doctrine. For the specific workload in front of me, the decision went the other way.

Here is where I will be harder on the cloud than most cloud advocates are willing to be, because if you migrate without fixing this you will end up with the worst of both worlds.

Cloud waste is structural and it is getting worse, not better. Flexera put wasted cloud spend at twenty-nine percent in 2026, rising after five straight years of decline, driven by AI workload complexity and the proliferation of new service SKUs. Independent analyses across hundreds of enterprise environments land in the same twenty-eight to thirty-five percent band. Ninety-one percent of enterprises report wasted cloud spend and three quarters say it is getting worse. Only about a quarter consider themselves efficient at managing it.

The mechanism is exactly what you think it is. Roughly forty-four percent of cloud spend covers non-production resources that are genuinely needed about forty hours a week and billed for all one hundred sixty-eight. Datadog found sixty-five percent of EC2 instances averaging below twenty percent CPU utilization over a thirty-day window. Kubernetes clusters average around ten percent CPU and twenty percent memory utilization, which means the overwhelming majority of container spend is buying reserved capacity that sits empty.

Translated into things that actually happen on a Tuesday: a developer spins up a test environment for a demo and never turns it off. Someone deletes a VM and leaves the managed disk, which keeps billing forever because nothing in the portal makes an orphaned disk feel urgent. A staging environment gets built to match production because that felt like the responsible thing to do, and now you are paying production prices to run smoke tests. Somebody picks an instance size by looking at the biggest number they can justify. Public IPv4 addresses became a line item across every provider and nobody noticed. An architect designs a chatty service that ships terabytes across a region boundary and discovers egress pricing in the same meeting where they discover the invoice.

None of that is Microsoft’s fault or Amazon’s fault. It is a governance failure, and it is the single most common reason a migration that penciled out on paper does not pencil out in production. The fix is not complicated and it is not optional: tag everything at creation and reject untagged resources with policy, auto-shutdown non-production compute on a schedule, put budgets and anomaly alerts on every subscription with an owner’s name attached, sweep for orphaned disks and snapshots and unattached IPs on a recurring job, right-size against actual telemetry rather than the architect’s anxiety, and buy reservations or savings plans only for the baseline you can prove from a year of data. Do that and the twenty-nine percent mostly evaporates. Skip it and you will spend a year proving the on-prem crowd right.

I priced my client’s migration with that governance built in from day one, because pricing it any other way is malpractice.

The security argument gets made badly in both directions, so let me make it precisely.

The provider gives you a platform floor that a mid-market company cannot build and cannot staff. Physical security and personnel screening at the facility. Firmware and hypervisor patching. Hardware root of trust. Encryption at rest on by default. Managed key vaults with rotation. DDoS scrubbing at the edge. An identity system with conditional access, MFA, and privileged identity management that would be a multi-year project to approximate on your own. Automated CVE remediation on managed services, so your SQL Server platform gets patched whether or not anyone on your team read the bulletin. When a critical vulnerability drops on a Friday, that is somebody else’s weekend and it is somebody whose entire business depends on getting it right.

Now the honest part. IBM’s breach data actually shows on-premises breaches costing slightly less on average than public cloud breaches, and something like forty-five percent of breaches now involve cloud environments. If you stop reading there you conclude the cloud is less secure. That is the wrong conclusion, and the reason is in the next number: Gartner has held for years that essentially all cloud security failures through 2026 are the customer’s fault. Misconfiguration shows up in roughly a quarter to a third of cloud breaches. Somewhere between eighty and ninety percent involve human error. Leaked credentials were the initial access vector in about sixty-five percent of analyzed cloud breaches.

None of those are platform failures. Those are public storage containers, over-permissive IAM roles, disabled audit logging, and secrets in a repository. Every single one of those mistakes is equally available to you on-prem, where you make them on top of a platform whose firmware nobody has updated in eighteen months, behind a firewall whose rules nobody has audited since the person who wrote them left, with an identity system that is a domain controller in a closet and a shared administrator password in a password manager if you are lucky. The cloud does not stop you from misconfiguring things. It just means the ninety percent of the stack underneath your misconfiguration is being run by people who are extremely good at it.

One more number, and it is the one that should actually drive your decision. Multi-environment breaches, meaning incidents that span cloud and on-premises, are the most expensive category at roughly five million dollars and take the longest to contain at two hundred seventy-six days. That is the real risk in this decision. Not cloud, not on-prem, but the half-migrated hybrid limbo companies land in when they cannot commit. Pick a side and finish.

The last argument is the one that does not appear on any spreadsheet and matters more than most of the ones that do.

When you own the hardware, capacity is a procurement event. Your biggest customer signs, traffic triples, and the answer is a purchase order, a six-to-twelve-week lead time in the current market, a change window, and a prayer. Your finance team asks why the sales win created a capital request. In the cloud the answer is a scale set parameter and a conversation about cost, which is a much better conversation to have.

It cuts the other direction too, and this half gets ignored. When you own the hardware, capacity you no longer need is capacity you still own. You cannot give it back. You cannot resize it. A workload that shrinks does not reduce your cost, it just lowers your utilization and makes the original purchase look worse. Elasticity is not only about handling a spike. It is about being wrong in either direction cheaply.

For a company whose growth curve is genuinely uncertain, which describes nearly every mid-market business I have ever worked with, the option value of being able to be wrong cheaply is worth real money. It just does not have a line item.

For this client and for the overwhelming majority of mid-market companies running conventional business applications, the cloud is the correct answer and it is not close.

It was cheaper on total cost of ownership once labor, redundancy, refresh, disaster recovery, and facility costs were counted honestly. It has a dramatically higher security floor with a compliance posture you can inherit rather than build. It removes an entire category of capital risk in the worst hardware market in a decade. It converts a capacity problem into a cost problem, which is always the better problem. And it lets a development team spend its time on the thing that actually makes the company money, which is the software, not the servers.

On-prem wins when you are large, stable, homogeneous, operationally sophisticated, and staffed for it. If you are reading this trying to decide, you are probably not that. And if you are, you already know it and you did not need me.

Everything above was about a company with no AI workload. The moment AI enters the conversation, someone in the room will suggest buying GPUs. This is the most expensive mistake available to a mid-market company in 2026 and I want to be very direct about it.

Cast AI analyzed roughly twenty-three thousand production clusters across AWS, Azure, and GCP and found average GPU utilization sitting at five percent. Ninety-five percent of provisioned accelerator capacity is idle at any given moment. That is measured telemetry, not a survey. And it arrived at exactly the moment NVIDIA raised H200 reserved pricing by around fifteen percent, breaking a twenty-year pattern of falling compute costs.

Now consider what owning that looks like. Enterprise GPU failure rates run five to ten percent annually, with Meta’s sixteen-thousand-GPU H100 cluster reporting roughly nine percent annualized. On-premises deployment requires N+1 or N+2 power and cooling redundancy, which means fifteen to twenty-five percent of your facility cost buys capacity you never use under normal conditions. Budget half to a full FTE per cluster just for drivers, firmware, failed hardware, and cluster operations. And the depreciation risk is not merely that the hardware ages. It is that model architectures change underneath you. If the open-weight ecosystem shifts toward different memory configurations or interconnect requirements, and it will, your cluster cannot adapt. You will be liquidating it.

There is a real break-even, and it is worth stating fairly: for sustained, genuinely high-utilization inference running around the clock, owned hardware beats cloud on-demand GPU pricing by a wide margin, and vendor TCO models claim order-of-magnitude advantages per million tokens at high utilization. Below roughly forty percent utilization you are burning money on idle silicon. Almost nobody in the mid-market is above that line, and the ones who think they are have not measured.

You do not have a GPU problem. You have a token problem. Solve it with tokens.

For the great majority of production agentic and RAG workloads I am building today, the DeepSeek API is the correct primary target, and the reason is that the price difference is not a discount, it is a different order of magnitude that changes what architectures are viable.

DeepSeek V4 Flash runs about fourteen cents per million input tokens on a cache miss and twenty-eight cents per million output. V4 Pro sits around forty-three and a half cents input and eighty-seven cents output. Both carry a one-million-token context window at no premium behind an OpenAI-compatible endpoint, which means your existing client code and your existing tooling largely just work. Automatic context caching is on by default, and on Flash a cache hit drops input to roughly a third of a cent per million, which is effectively free for the system prompt and tool definitions you resend on every single call in an agent loop.

That last detail is the one people miss. Agentic systems are pathologically repetitive. The same system prompt, the same tool schemas, the same retrieved context, resent on every turn of every loop. On a frontier-priced model that repetition is your entire bill. With aggressive prefix caching at DeepSeek rates it approaches noise, which means you can afford architectures, longer context, more tool calls, more reflection passes, that you would have to design around at ten or thirty times the token cost. Cheap tokens are not just cheaper. They are an architectural degree of freedom.

Two practical notes from production. Pin explicit model identifiers; the legacy aliases retired in July 2026 and code referencing them fails. And DeepSeek has announced a peak-hour pricing multiplier without publishing an effective date, so build your cost model with headroom and keep an eye on the pricing page rather than assuming today’s rates are permanent.

This is where the mid-market gets the architecture wrong in the opposite direction, by routing everything through a hyperscaler AI service because it feels safer.

Microsoft Foundry, which is what Azure AI Foundry became, now lists DeepSeek V4 Pro, V4 Flash, V3.2, and V3.2 Speciale among models sold directly by Azure, alongside Claude, GPT, Llama, Mistral, and a catalog running into the thousands. You get Entra ID for identity, your existing Azure commitments applying to spend, Azure Monitor for observability, Azure Policy for guardrails, and Microsoft’s contractual assurance that prompts and completions are not shared with model providers or used to train base models. For a .NET shop that already lives in Entra and Azure billing, adding a model vendor stops being a procurement cycle and becomes a routing decision. Bedrock offers the equivalent story on the AWS side, with over a hundred models across a dozen-plus providers inside your VPC boundary under IAM and KMS.

You pay for that. Bedrock lists DeepSeek V3.2 at roughly sixty-two cents input and a dollar eighty-five output per million, which is several times the first-party rate. That premium is not a ripoff; it buys data residency, enterprise controls, a single throat to choke, and integration with infrastructure you already govern. It is worth paying for the workloads that need it. It is not worth paying for the workloads that do not.

So the pattern I ship is a routing layer, not a vendor choice. Bulk work, classification, extraction, summarization, embedding, the high-volume interior of an agent loop, goes to DeepSeek first-party at first-party prices. Anything touching regulated data, anything where the data residency conversation with your customer’s security team is going to happen, anything where an enterprise contract and an SLA are the actual product, routes to Foundry or Bedrock. Frontier reasoning on the small percentage of calls that genuinely need it goes to whichever premium model is winning that month, through the same abstraction.

Build the routing layer on day one even if it only has one destination. Model pricing has moved by double-digit percentages within single quarters, entire model families have been deprecated with a few months’ notice, and the cheapest capable model in any given category has changed more than once a year for three years running. If swapping providers requires a refactor, you have not built an AI system, you have built a dependency. And note the meta-lesson: that volatility is itself the argument against owning hardware. You cannot re-architect a purchased GPU cluster in an afternoon. You can change a routing rule.

The final reason to stay out of the AI hardware business is that the ground is moving fast enough to make any purchase look foolish in twenty-four months.

Right now, in mid-2026, a hundred-twenty-eight-gigabyte unified memory machine on your desk runs a hundred-twenty-billion-parameter open-weight model at production-usable speeds. OpenAI’s open-weight flagship at Q6 quantization occupies about ninety-three gigabytes and turns out eighteen to twenty-eight tokens per second on an M2 Ultra, fourteen to twenty on an M4 Max. That is a workstation, not a rack. NVIDIA’s DGX Spark and Jetson Thor both land at a hundred twenty-eight gigabytes of unified memory in a form factor that sits on a shelf, AMD’s Ryzen AI Max reaches the same capacity, and Spark-class Windows laptops and small desktops from ASUS, Dell, HP, Lenovo, Microsoft, and MSI are scheduled for this fall, promising local hundred-twenty-billion-parameter inference with the full CUDA ecosystem in a three-pound machine.

Mixture-of-experts architecture is what makes this work. A model activates a small fraction of its parameters per token, so a hundred-twenty-billion-parameter MoE fits by total size while decoding at speeds a dense model that size never could. That architectural trend is not reversing, and it means the capability that fits on a personal machine keeps climbing without a proportional increase in memory bandwidth requirements.

I want to be precise rather than breathless, because the gap is real today. DeepSeek V4 Pro is roughly one-point-six trillion parameters with forty-nine billion active and does not fit on any single consumer machine. GLM-class open models at seven hundred forty-four billion parameters need well north of two hundred gigabytes even at aggressive two-bit quantization. Today, frontier means cloud.

But look at the direction and the slope. The consumer memory ceiling has roughly doubled in two years. MoE keeps improving the capability-per-gigabyte ratio. Quantization keeps getting better at preserving quality. Open-weight models keep landing within months of the closed frontier rather than years behind it. The trajectory says that a meaningful share of what you are paying API rates for today runs locally on a machine you were going to buy anyway, well within the depreciation window of anything you purchase now. Buying specialized inference hardware in 2026, at peak component pricing, with six-month lead times, to run models that will fit on a laptop before that hardware is paid off, is a decision you will have to defend in a board meeting.

Rent tokens. Keep the flexibility. Let someone else own the depreciating asset and the supply chain risk and the driver updates.

If you are running conventional business applications, go to the cloud, but go with governance from day one. Tagging enforced by policy, scheduled shutdown on non-production, budgets and anomaly detection with named owners, recurring sweeps for orphaned resources, right-sizing against telemetry, and commitments purchased only against a proven baseline. Migration without governance is how you end up as a cautionary tale in somebody else’s repatriation article.

If you are adding AI, do not buy GPUs. Route to DeepSeek for volume, use Foundry or Bedrock for the workloads where enterprise controls and data residency are the actual requirement, build the routing abstraction before you need it, and instrument your token spend the way you instrument everything else. In eighteen months, quietly move the workloads that make sense onto local models running on hardware your team already owns.

And do the exercise honestly before you commit either way. Count the people. Count the refresh. Count the disaster recovery site. Count the weekend somebody spends on a firmware bug. The answer is not always the cloud. For my client it was, decisively, and I went in expecting the opposite.

I do this work for a living. Application modernization, cloud migration and cost optimization, and production agentic AI systems on the Microsoft stack, built by someone who has been shipping .NET into production for close to thirty years and has the scar tissue to prove it. If you are staring at a colocation quote, an Azure bill you cannot explain, a legacy application that needs to move, or an AI initiative that needs to be architected by someone who has actually put one in production rather than presented a slide about it, reach out. I am happy to run the real numbers with you, including the ones that argue against hiring me.

Gal Ratner is the founder and CTO of Inverted Software and WhiteStar Labs, and Chief Architect at Prana Entertainment, a Las Vegas-based enterprise software and AI consultancy. He has spent nearly thirty years shipping production software on the Microsoft and .NET stack for clients including Microsoft, Sony, Rockstar Games, 2K Games, Best Buy, and Allegiant Air, and was employee number six at Break.com.

His current work centers on production agentic AI: MCP servers, the Microsoft Agent Framework, RAG pipelines, SQL Server 2025 vector search, and the PLogger observability framework. He also operates ShopSnap.

Outside of work he trains Brazilian jiu-jitsu under Sergio Penha in Las Vegas, rides motorcycles, co-hosts Edge Grip Podcast, and wrote the novel The Archive of Lost Suns. He writes about the gap between executive AI narratives and production reality on Substack and LinkedIn as the .NET AI guy

No posts

Read the original on galratner.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.

    Reading · Gal Ratner · RSS Amplifier