RSS Amplifier

The Chaos Guru · Apr 30, 2026

Stop trying to fail over from Cloudflare

0
Sign in to vote or save

Luka Kladarić · The Chaos Guru

On November 18, 2025, a database permissions change caused Cloudflare’s Bot Management feature file to balloon from 60 entries to over 200, exceeding a hard-coded limit in the core proxy. The proxy crashed. One in five webpages went dark. A third of the world’s 10,000 most popular websites stopped responding. X, ChatGPT, Spotify, Zoom, Coinbase, and Canva, all down. Even Downdetector went down. The outage started at 11:20 UTC; systems weren’t fully normal until 17:06.

Less than three weeks later, on December 5, Cloudflare engineers were rolling out a buffer size increase to protect against a React framework vulnerability. Their internal WAF testing tool didn’t support the new size, so they pushed a second config change through their global deployment system. It propagated to the entire fleet in seconds. 28% of all HTTP traffic served by Cloudflare dropped for 25 minutes.

The reaction was predictable. “We need a Cloudflare failover plan.” “We can’t be dependent on a single provider.” “Just point DNS at the origin until it comes back.”

The instinct makes sense. Removing a single point of failure is one of the first things you learn in reliability engineering. But the instinct is misguided here. Most people don’t fully understand what Cloudflare is actually doing for them, and once you do, the failover math looks very different.

Cloudflare’s free tier gives you a global CDN, DDoS protection, DNS, and a basic WAF. The $20/month Pro plan adds more WAF rulesets and better analytics. This is not normal. This is insane. Worth unpacking what each of those actually is.

Cloudflare operates a global anycast network across 330+ cities. When someone in Tokyo requests your website, they hit a Cloudflare edge node in Tokyo. The response is cached there. The next request from Tokyo never touches your origin server.

This matters more than people realize. Your origin might be a single server in Virginia. Without a CDN, every user everywhere is making a round trip to Virginia. With Cloudflare’s cache, most of that traffic never leaves the user’s region. Latency drops. Origin load drops. And your bandwidth bill drops with it.

That last part is the one that bites hardest if you ever turn Cloudflare off. If you’re serving images, videos, static assets, or anything heavy through Cloudflare’s cache, your origin is only seeing a fraction of the actual request volume. Pull the CDN out from under that and suddenly your origin is serving every byte directly. If your origin is an S3 bucket, you’re now paying AWS egress rates on 100% of your traffic instead of the sliver that was cache misses. AWS charges $0.09 per GB for egress in the first 10TB tier. A site doing 10TB/month through Cloudflare’s cache might only be sending 500GB from S3. Remove Cloudflare and that becomes 10TB at $0.09/GB, an extra ~$855/month on top of what you were already paying. And that’s a modest example.

Cloudflare’s CDN caching is free. On the free plan.

Cloudflare’s WAF sits between the internet and your application and inspects every request for malicious patterns, like SQL injection, cross-site scripting, path traversal, and all the OWASP Top 10. Cloudflare catches it before it reaches your code.

The rules are written in Wirefilter syntax, Cloudflare’s own expression language, referencing threat scores and bot detection signals that only exist inside Cloudflare’s network. These aren’t generic pattern matchers. They’re computed from traffic patterns across Cloudflare’s entire customer base.

To put that in context: the equivalent from AWS WAF requires you to configure CloudFront, attach WAF rules using JSON rule statements, manage rule groups, and pay per million requests. On Akamai you’re talking to an enterprise sales team and signing a contract. Cloudflare gives you a working WAF on the free plan.

The reason Cloudflare’s DDoS protection works is simple geometry. Cloudflare is more distributed than the attacks.

If you control tens of thousands of compromised devices worldwide, it’s easy to overwhelm any single thing. Having one server, one load balancer, one connection is an indefensible position in 2026. But Cloudflare operates a global network across 330+ cities. Your application potentially has hundreds of IPs associated with it at any given time, often multiple IPs per edge location. If attackers target one IP, Cloudflare delists it and stops using it. If they follow whatever IP comes up next for your service, the attack is already scattered across the globe in manageable chunks. Each edge node only sees its local slice of the traffic.

This means Cloudflare can absorb attacks orders of magnitude larger than anything your hosting provider can handle, unless that provider also offers a distributed DDoS mitigation service that you continuously operate under. The closest equivalents are Akamai Prolexic (enterprise contract) and AWS Shield Advanced ($3,000 per month base, AWS resources only). Cloudflare gives you this at L3 through L7 for free. On the free plan.

This connects directly to the failover question. By stepping out from under the Cloudflare umbrella for a few hours while they’re having problems, you’re telling the internet exactly where your origin is and how to hurt you whenever anyone wishes to do so later. You lose protection during the outage window. You lose the protection model permanently.

Cloudflare proxies roughly 20% of all web traffic on the internet. Their bot management uses machine learning models trained on traffic patterns across that entire network. When the same bot hits thousands of Cloudflare customers, the system learns its fingerprint. That cross-customer signal can’t be replicated by anyone seeing less traffic, which is basically everyone.

The November 2025 outage was, ironically, caused by the Bot Management system itself. I’ll come back to that.

Cloudflare handles certificate issuance, renewal, and termination automatically. Authenticated Origin Pulls let your origin verify that requests are actually coming through Cloudflare. Cloudflare Tunnel eliminates the need for a publicly routable origin IP entirely, your server calls out to Cloudflare instead of listening on a public port.

You don’t manage certificates. You don’t worry about renewal. You don’t deal with Let’s Encrypt cron jobs failing silently at 3am. It just works.

Cloudflare runs one of the fastest authoritative DNS services on the planet. Sub-millisecond response times in most regions. Free.

So who’s yelling about failover during these outages?

I keep seeing this from operators running their application in a single AWS region, single AZ, maybe a load balancer if they’re feeling fancy. No multi-region. No geo-redundancy. Their own uptime is worse than Cloudflare’s. Cloudflare had roughly seven hours of total major-incident downtime across 2025, call it 99.92% availability on the proxy layer. These same operators had their own production incidents last month that they blamed on a bad deploy, a database migration, or a dependency update that went sideways.

But Cloudflare goes down for five hours and suddenly multi-CDN failover is the priority.

Let me walk through what that actually looks like, and why it doesn’t work.

Cloudflare WAF rules don’t port to anything. Wirefilter syntax is Cloudflare-only. The cf.* fields are computed inside Cloudflare’s network. You’d need to manually rewrite every rule in whatever format the target platform uses. JSON for AWS WAF, VCL for Fastly, Akamai’s proprietary rule engine. Every rule rewritten, retested under real traffic, tuned for the new platform’s quirks. That’s weeks of work if you’re fast.

And even if you did it, configuration drift is inevitable. Every rule change on your primary must be replicated on the standby. In practice, the policies diverge within months and nobody notices until failover day, when the standby blocks real users or lets attacks through.

Cloudflare’s DDoS mitigation is architectural. It’s how their network works, not a product bolted on top. Replacing it means either paying for Akamai Prolexic (enterprise contract) or AWS Shield Advanced ($3k/month, AWS-only), or just... not having DDoS protection. During the outage window. When the internet is paying attention.

This is the suggestion that sounds reasonable for about ten seconds. What could go wrong?

The moment your DNS resolves to your origin IP instead of Cloudflare’s proxy, your origin is naked on the internet. No WAF. No DDoS protection. No bot management. No rate limiting. Full protection to zero protection, at the exact moment every script kiddie with a Shodan query is watching.

But it gets worse. The moment your origin IP appears in a DNS response, it’s recorded. Permanently.

SecurityTrails archives every DNS record it sees. Censys and Shodan continuously scan the entire IPv4 address space. Historical DNS databases like ViewDNS.info store everything. Even with Cloudflare’s default TTL of 300 seconds, a brief window is enough for these systems to permanently record your origin IP.

Once it’s known, it’s known forever. Or until you migrate to a new IP.

There’s a whole ecosystem of tools built specifically for this. unwaf discovers origin IPs using nothing but passive techniques (SPF records, MX records, subdomain enumeration, Certificate Transparency logs). No API keys needed. CloudFlair searches Censys for hosts presenting matching SSL certificates. There are dozens more. The now-defunct CrimeFlare maintained a database of domain-to-origin-IP mappings built through weekly scans of millions of domains. It’s gone, but its successors carry on.

As Detectify documented, once the origin IP is found, an attacker connects directly with a spoofed Host header and every Cloudflare protection becomes useless. Zafran’s “BreakingWAF” research in 2024 found nearly 40% of Fortune 100 companies had this exact misconfiguration. Certitude showed you can even use your own Cloudflare account to tunnel attacks through Cloudflare’s infrastructure against another customer’s origin.

And that’s just the security side. Remember the CDN? Your origin has been sitting behind Cloudflare’s cache, serving maybe 5-10% of actual request volume. Point DNS at the origin and your server eats 100% of traffic with a cold cache. If your origin can handle that spike without falling over, congratulations, you’ve been massively over-provisioning. If it can’t (and most can’t) you’ve just traded a Cloudflare outage for a self-inflicted one. Plus whatever your cloud provider charges for the bandwidth.

“Just point DNS at the origin” means permanently burning your origin IP for every future attacker, spiking your bandwidth bill, and probably crashing your origin under the traffic it was never sized to handle. That’s not a failover plan. That’s an own goal.

The more sophisticated version is maintaining a parallel CDN/WAF/DDoS stack on a second provider, ready to go at a moment’s notice.

Think about what that means. WAF rule parity across incompatible platforms. Configuration drift that you won’t catch until the day you need it. A cold CDN cache that causes a thundering herd against your origin the moment you fail over. Different logging formats, different metric definitions, different dashboards. SSL certificates managed in two places with different renewal cycles. Paying for two full platform subscriptions (Akamai’s enterprise pricing could easily be 5-10x what Cloudflare costs). And hidden shared dependencies where your “redundant” provider shares upstream transit with your primary anyway.

The operational complexity of keeping two incompatible security platforms in sync, untested under real traffic, is itself a reliability risk. You’ve added fragility in the name of reducing fragility. And you’ve done it at enormous cost, to recover from an event that happens a few hours per year.

If you have half a dozen full-time operations staff who can maintain that parity, test it regularly, and actually execute the failover under pressure, sure. Go for it. That’s a real multi-CDN strategy and it works for companies the size of Netflix.

For everyone else, this is fantasy architecture. The plan looks great in a slide deck and falls apart the first time you try to use it.

Cloudflare is also ridiculously simple to implement. The defaults are solid. You point your DNS at Cloudflare, flip the orange cloud on, and you have a CDN, a WAF, DDoS protection, bot management, and SSL termination working out of the box. If you need to customize, the configuration is flexible without being overwhelming. This is not a product that demands a dedicated team to operate.

Now try to replicate that yourself. Ignore the ability to survive massive attacks. Ignore that Cloudflare is continuously updating its protection in response to emerging threats. Just try to match the feature set in the most naive, basic terms. You’re stitching together a CDN, a WAF, a DDoS mitigation layer, certificate management, DNS, and bot detection. The cost of ownership for a self-maintained version of this easily blows through even the $200/month Business plan, at a laughable fraction of the power, volume, and feature set.

This is what makes the idea of maintaining a hot standby, or continuously splitting traffic between Cloudflare and your own parallel solution, completely batshit insane.

Yes, there’s math to be done about the level at which building your own solution becomes the better choice. But my guess is that if you had a hot take during Cloudflare’s outages, your numbers won’t back you.

I think people miss something about Cloudflare. This is a company that understands it has become world-critical infrastructure and acts like it.

After the November and December incidents, Cloudflare declared “Code Orange: Fail Small,” a company-wide engineering priority. Controlled rollouts. Fail-open error handling, so if a config is bad, the system defaults to passing traffic instead of dropping it. Blast-radius reduction so a single config change can’t take out the entire fleet in seconds. They published detailed post-mortems for both incidents. They explained exactly what went wrong, why, and what they’re changing.

Compare that to your cloud provider’s last incident report. “A subset of customers in us-east-1 experienced elevated error rates for approximately 47 minutes. The issue has been resolved.” No root cause. No timeline. No accountability. Cloudflare tells you what the hard-coded limit was, what the database query returned, and which engineer pushed the config. That’s not a company hiding from its responsibilities.

The February 2025 R2 incident is the same pattern. A routine phishing takedown accidentally disabled the entire R2 gateway for 59 minutes. They published every detail, admitted the tooling didn’t prevent an operator from nuking a production service, and laid out the fixes.

As the Pragmatic Engineer noted, the November and December outages both came from global config changes propagating too fast. Same failure mode, three weeks apart. That’s bad. But Cloudflare’s response wasn’t to minimize it. It was to declare a company-wide engineering emergency and restructure how config changes propagate. That’s what a company acting like critical infrastructure looks like.

Three more incidents have hit Cloudflare since the December 5 outage. A 6-hour BYOIP route withdrawal on February 20. A 54-minute Sites-and-Services degradation on April 3. A regional spate of Cloudflare Access HTTP 500s across April 21–24.

The February 20 incident was the worst of the three. A buggy cleanup subtask in the Addressing API ran a query passing the pending_delete flag with an empty value, which the server interpreted as “match everything,” and systematically withdrew 1,100 BYOIP prefixes from production, about 25% of all BYOIP prefixes Cloudflare advertises. Recovery took six hours. Cloudflare published a clinical post-mortem within 48 hours: the buggy query, the broken assumption, the recovery timeline, the fix.

The April 3 incident was a 54-minute degradation of the Sites and Services component, with increased latency and intermittent 502/503/504s across regions. The April 21–24 Cloudflare Access incident was small-blast: a subset of authentication requests in specific geographies, contained and recoverable.

That last one is exactly what Code Orange was designed to produce. The first two are Cloudflare continuing to find new ways to break, in places Code Orange wasn’t yet covering.

The response pattern is consistent. Where there’s a post-mortem, it’s clinical. Where there isn’t yet, the incident was small enough that the status page tells the story. The thesis isn’t that Cloudflare doesn’t fail. It’s that Cloudflare fails like a grown-up.

Cloudflare had roughly seven hours of major-incident downtime in 2025. That’s 99.92% availability on the proxy layer. Not perfect. Seven hours of downtime on a system proxying 20% of the internet is genuinely bad. No qualifier.

But here’s the question nobody in the failover crowd is asking: what’s your uptime?

If you’re running a single-region deployment (and most of you are) your own application probably had more downtime than Cloudflare did. Bad deploys. Database migrations. That one time someone fat-fingered a Terraform apply. The dependency update that broke production on a Friday afternoon. The SSL certificate that expired because nobody set up the renewal alert.

You tolerate all of that. You don’t demand a hot standby for your own application. You fix the problem, write a post-mortem (maybe), and move on.

When Cloudflare has an incident, you get the same thing: a problem, a fix, a post-mortem (always), and improvements. The difference is Cloudflare’s post-mortems are better than yours.

Should we demand Cloudflare do better? Absolutely. A system proxying a fifth of the internet should not deploy config changes globally and instantaneously. That lesson should have been absorbed the first time. The fact that it happened twice in three weeks is a real problem and Cloudflare rightly treated it as one.

Staged rollouts. Better blast-radius controls. Fail-open defaults on every component. Cloudflare is doing all of these things.

What we shouldn’t do is pretend we can architect around Cloudflare on a budget. The hot takes I saw on LinkedIn during and after those incidents all shared one blind spot: they assumed that reducing your Cloudflare dependency can only improve your position. Nobody considered the possibility that you could make your position significantly worse. Every alternative they proposed (pointing DNS at the origin, maintaining a parallel stack, splitting traffic) introduces risk that doesn’t exist when you just stay behind Cloudflare. The downside isn’t neutral. It’s negative.

And yet: you can’t replicate what they do at their price, with their global footprint, or with their traffic-level bot detection. The $20/month Pro plan gives you security infrastructure that would cost six figures to approximate on any other platform, and you’d need a team to maintain it.

If you have Netflix’s budget and Netflix’s ops team, run a multi-CDN setup. For everyone else, use Cloudflare. Use Cloudflare Tunnel so your origin IP never leaks. For AWS shops, lock down your origin’s security group to only accept traffic from Cloudflare’s proxy IPs — I built awesome-cloudflare-sg for exactly this, a CloudFormation stack that auto-syncs an EC2 security group with Cloudflare’s published IP ranges via Lambda. If your origin IP leaks, the security group makes the leak useless to anyone whose source IP isn’t on Cloudflare’s list. Build reasonable origin-level defenses as a degraded fallback. And when Cloudflare has a bad day, wait it out. They’ll fix it. They’ll tell you exactly what happened. And they’ll make it less likely to happen again.

And that’s not blind faith, it’s just looking at the track record and making a rational call.

Read the original on chaosguru.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.