RSS Amplifier

cybersins/sec · Aug 12, 2026

Pay per crawl: putting HTTP 402 to work, and the requirement I missed

0
Sign in to vote or save

Rishi Narang · cybersins/sec

HTTP 402 Payment Required has been in the specification since HTTP/1.1 and has never had a defined meaning. It was reserved for a digital cash system that did not arrive. MDN still files it as nonstandard and notes that no browser does anything useful with it, so a 402 renders as a generic error.1

Cloudflare’s pay per crawl is the first widely deployed thing I have seen that gives it a job. Rather than inventing a payment protocol, they took the one the web already had and pointed it at a client with no browser, no human waiting, and no objection to being charged a tenth of a cent.

I wanted to understand the mechanics properly, so I built a simulator of the whole thing: a small API with priced endpoints, a web console for configuring prices and crawler rules, an Ed25519 signing client, and a ledger to watch the money move. Every screenshot in this post is from that simulator running on my laptop. None of them are Cloudflare’s product, and nothing here is a Cloudflare dashboard. The protocol is theirs; the implementation and the interface are mine, written from their published specification so I could see the flow end to end.

I also got one of the signature requirements wrong, which I will come to, because it turns out to be the most interesting thing in the protocol.

This is part five of a seven part series on Cloudflare.

The bargain that broke#

The old arrangement was so obvious that nobody wrote it down. You let search engines crawl your site, they indexed it, someone searched and clicked and arrived, and you made money from their attention through an ad, an email capture, a subscription or an affiliate link. The crawler got content, you got a visitor, and that trade paid for a great deal of the web.

AI breaks the loop at the last step. The system reads your page, extracts the answer, hands it to the person, and the person never arrives. Your content still created the value; the visit did not happen.

Cloudflare has published the numbers. These are crawl-to-refer ratios from July 2025, showing how many times each company crawled for every one visitor it sent back.2

CompanyCrawls per referred visitor
Anthropic38,065 : 1
OpenAI1,091 : 1
Perplexity194 : 1
Microsoft41 : 1
Google5.4 : 1

Google’s 5.4:1 is roughly what the old bargain looked like when it worked. Narrowed to News and Publications, Anthropic still sits at 2,500:1 and OpenAI at 152:1.3 From the reader’s side, Pew found that when an AI summary appeared people clicked a traditional result in 8% of visits, against 15% when no summary was present, with clicks on links inside the summary itself at 1%.4 And about 80% of AI bot activity on Cloudflare’s network is training, roughly 17% search, and around 3% real-time user actions,2 so the slice that most resembles the old deal is also the smallest.

What pay per crawl is#

Announced on 1 July 2025 by Will Allen and Simon Newton, pay per crawl adds a third option to a choice that previously had two.5 Per crawler, you pick Allow, which is free access as today; Charge, which means 402 until paid; or Block, which refuses with no payment path and is functionally a 403.

Underneath sit four headers and two flows. The traces below are my simulator exercising both of them against my own API, so they show the shape of the protocol rather than anything Cloudflare serves.

Discovery: ask, get quoted#

An anonymous crawler sends a plain GET with no signature and no payment intent.

Terminal output from the author's simulator client showing an unsigned GET to /humor returning HTTP 402 Payment Required with a crawler-price header of USD 0.010
My simulator’s client in discovery mode, against my own API. No Web Bot Auth headers, no payment intent. The server quotes a price and serves nothing.

Back comes HTTP/1.1 402 Payment Required with crawler-price: USD 0.010. That is the discovery mechanism in full. The 402 is not a rejection so much as a price tag, and the crawler now knows what the page costs without having been able to find that out any other way.

Two details in that response I would have missed without implementing it. The body is JSON rather than an error page, because there is no human to read an error page. And access-control-expose-headers has to list crawler-price and crawler-charged, or a browser-based client sees the status code and none of the pricing. The second one cost me twenty minutes.

Payment: retry with intent#

The crawler decides the joke is worth a penny and retries, this time signed, with crawler-exact-price set to the quoted amount.

Terminal output from the author's simulator client showing a signed GET with Signature-Agent, Signature-Input, Signature and crawler-exact-price headers, returning HTTP 200 OK with crawler-charged USD 0.010
The same request forty seconds later, signed and carrying payment intent. My server returns 200 OK and crawler-charged confirms the debit against my simulated ledger.

200 OK, and crawler-charged: USD 0.010. Content delivered, money moved, receipt returned as a response header. No checkout page, no session, no account creation, no redirect: the request itself carries the transaction.

The alternative flow skips the negotiation entirely. The crawler puts crawler-max-price on the first request, declaring a ceiling, and if the site’s price is at or under it the content comes straight back with crawler-charged attached.

The proactive flow is the one that matters at scale. Declaring a budget up front collapses two requests into one, and for a fleet doing millions of fetches a day that is the difference between a workable protocol and doubling your request volume for nothing.

Cloudflare sits in the middle as merchant of record, aggregating charging events, billing the crawler and paying the publisher. That is what makes the arrangement plausible rather than merely elegant, because no publisher wants a billing relationship with forty AI companies and no AI company wants forty thousand publisher invoices.

A few things from the documentation surface only if you read it carefully.6 Charging fires only on successful responses, so errors are free even when you sent a price header. Crawlers are charged again for re-crawling the same page, with no free revisit. A handful of paths are permanently free: /robots.txt, /sitemap.xml, /security.txt, /.well-known/security.txt and /crawlers.json. And you cannot set different prices for different crawlers, only one price per zone for everyone on Charge, with the action varying per crawler rather than the number.

That last constraint is the one I would push back on hardest, and Cloudflare has half-fixed it. A June 2026 update added URI pattern exclusions and dynamic pricing driven by response headers or a Worker, which gets you per-path pricing even though per-crawler pricing is still off the table.7 Thirteen months after launch it remains in closed beta.8

Web Bot Auth, and the part I got wrong#

None of this works if you cannot tell who is asking. A price header attached to a user agent string is worthless, because user agents are free text and always have been.

The mechanism is Web Bot Auth, which applies HTTP Message Signatures (RFC 9421, February 2024) to bots.9 The crawler publishes an Ed25519 public key as a JWKS at /.well-known/http-message-signatures-directory over HTTPS, with Content-Type: application/http-message-signatures-directory+json, and the directory is itself signed. Ed25519 only; Cloudflare supports no other algorithm. Each request then carries a Signature-Agent header pointing at that directory, plus Signature-Input and Signature.

Here is what my signing client sent:

http

1Signature-Agent: "http://localhost:8402"
2Signature-Input: sig1=("@authority" "signature-agent");created=1786468338;
3  expires=1786468638;keyid="hUghpIfR4Brmmuo5oTKUEzTkl6ptWRlCdNPc_u3g-Y8";
4  tag="web-bot-auth";nonce="da13cd6fd548efb7"
5Signature: sig1=:V/sQC2czfzNyf8v7v1NiobgCd5vqC9JYkfFcUhDE0tn3JrBOiWv9EY/...:

keyid is the JWK thumbprint, which the verifier uses to select the right key from the directory. tag="web-bot-auth" marks this as bot authentication rather than any other use of RFC 9421. created and expires bound the validity window.

The author's own Web Bot Auth configuration panel showing strict Ed25519 verification, two registered crawler keys named pink and brown, and a clock skew setting of 300 seconds
The key registry in my console. Two crawlers I generated keys for; the key ID beginning hUghpIf is ‘brown’, which signed the paid request above.

Now the mistake. The covered component list my client builds is ("@authority" "signature-agent"). That is what the Web Bot Auth documentation requires in general, and it is what Cloudflare’s own example shows.9 It is not sufficient for pay per crawl, and has not been since 10 December 2025, when Cloudflare added a requirement I built straight past: payment headers must themselves be inside the signed component list.10

Their stated reasoning is worth quoting, because it is a tidy summary of what the signature is actually for:

Payment headers (crawler-exact-price or crawler-max-price) must now be included in the Web Bot Auth signature-input header components. This security enhancement prevents payment header tampering, ensures authenticated payment intent, validates crawler identity with payment commitment, and protects against replay attacks with modified pricing.

Signing crawler-exact-price binds the crawler’s identity to the specific amount it agreed to pay, so an intercepted request cannot be re-spent at a different price and the value cannot be rewritten in flight. My simulator authenticates the crawler and leaves the price unauthenticated, which is exactly the gap that requirement closes. My own server happily accepted it, because I wrote the verifier too and built the same assumption into both halves. Cloudflare would have rejected it, which is the difference between a simulator and the thing it simulates.

The reason this matters more than it first appears is what sits next to it. Cloudflare’s documentation is direct about the limits of the nonce: “there is currently no nonce validation, nor does Cloudflare guard against replay attacks using a database of seen nonces.”9 Replay protection rests on the expires window instead, and their advice is to keep it short, with “a minute is often sufficient.” My created and expires are 300 seconds apart, which I chose because five minutes is what everyone chooses, and that is not a good reason. With no nonce checking behind it, that window is the whole of your replay tolerance, and signing the payment header is what stops a replayed request inside it from being spent at a price the crawler never agreed to.

While I am cataloguing my own errors: Signature-Agent in that trace reads http://localhost:8402, because the simulator is a local server talking to itself. Cloudflare rejects http:// URIs outright, and rejects the dictionary form of the header too. Harmless on a laptop, immediately fatal against the real service.

The same December release also added a crawler-error header with eleven specific error codes, so a crawler can distinguish an expired signature from an unsupported currency from a price mismatch without parsing prose, and a Discovery API so crawlers can enumerate priced domains rather than probing for 402s.10 Both are the kind of thing you only add once real crawler operators have tried to integrate.

On standards, the position is better than I expected. Web Bot Auth is now a chartered IETF working group in the Web and Internet Transport area, chaired by David Schinazi and Rifaat Shekh-Yusef, with milestones targeting IESG submission during 2026.11 The drafts moved into the webbotauth namespace, so any tutorial citing draft-meunier-web-bot-auth-architecture is stale. Google is participating through Gary Illyes’ drafts on crawler best practices and publishing automated-client IP ranges.12 OpenAI signs Operator requests and has ChatGPT Agent in the signed agents programme, with Block’s Goose, Browserbase and Anchor Browser alongside.13 I found no evidence Perplexity has adopted it, and no first-party confirmation that Google runs it in production rather than only participating in the group.

What building the simulator changed my mind about#

One price per site is the wrong shape#

The author's own console showing five API endpoints with individual monetisation toggles and prices in USD, one of them toggled off and marked free
Per-endpoint pricing in my console. This is a design I chose for the simulator, not something Cloudflare offers: their pricing is one flat rate per zone. A dice roll at $0.002 and a classification call at $0.025 are not the same product.

I did not set out to argue with Cloudflare’s pricing model. When I came to build the configuration screen I had to decide what a price attaches to, and putting it on the zone rather than the resource felt wrong the moment I tried it, so I built mine per endpoint instead. A dice roll and an expensive classification call are not worth the same money, and a publisher’s archive is not worth what its front page is worth. I priced five endpoints between $0.002 and $0.025 and left one free, and each of those numbers felt like a decision rather than an arbitrary choice.

The flat rate exists because it is the simplest thing that bills reliably, which is a reasonable thing to ship first. The June 2026 dynamic pricing update is the admission that it was too simple. If you deploy this, plan on writing the Worker.

The log is worth more than the money#

The author's own request log showing ten requests with verification status, 402 price quotes, an exact-price mismatch, and successful charged responses
The log view I built, showing every state in one place. The 17:11:38 line is the discovery request from earlier; 17:12:18 is the paid retry.

This is the screen that changed how I think about the feature, and I only built it so I could debug the rest. Reading down it, every distinct outcome is visible: unsigned requests getting quoted, signed requests getting quoted, a signed request whose crawler-exact-price no longer matched and was rejected, and the ones that paid and were served.

What I had accidentally built was a bot analytics feed with money attached rather than a billing log. Knowing which AI companies want which of your pages, how often, and what they will pay is information you cannot currently obtain at any price, and for most sites I suspect it is worth more than the revenue will be for the next couple of years. Cloudflare appears to agree, having added an Attribution Business Insights dashboard covering AI consumption and referrals per company in July 2026.14

The surface is small, the operations are not#

The author's full simulator console showing endpoint pricing, crawler rules, Web Bot Auth configuration, the live request log and a billing ledger
The whole simulator: pricing, rules, signature verification, log and ledger. The ledger is fake money against a local server. The protocol underneath it is small; the policy questions are not.

Four headers, one status code, and a signature scheme that already had an RFC. Every previous attempt at web micropayments needed new infrastructure, new wallets, and a human consenting at the moment of purchase, and that last requirement killed all of them. Nobody will agree to pay a tenth of a cent to read a recipe; they will close the tab out of irritation. A program with a budget and a task will pay it a thousand times a day without noticing.

What a simulator lets you skip, and a real deployment does not, is the operational cover. Strict verification means fetching and caching the crawler’s key directory, which puts an outbound HTTP call to a third party in your request path. You need a cache, a timeout, and a decision made in advance about what happens when the fetch fails: serve free, or turn away a paying customer. On my laptop the directory was always up because I was serving it. Cloudflare handles all of that at the edge, and it is a real part of what you are buying.

What each side gets#

Publishers gain a third option where they had two poor ones, a price signal where they had none, and one payments counterparty instead of forty. Crawler operators gain certainty, which they want more than is generally assumed: an AI company paying a published price for signed, logged access has a defensible position that one scraping through a residential proxy does not, and a signed crawler stops being challenged, rate-limited or quietly served rubbish, which is worth real engineering time before you count the licensing risk.

The asymmetries deserve naming too. The price is set unilaterally, with no negotiation in the protocol beyond accept or decline, which is fine for a $0.001 page and untenable for a large archive, so bilateral licensing deals are not going anywhere. Nobody has published what publishers actually earn, and after thirteen months of beta there is no aggregate payout figure from Cloudflare, TollBit, ScalePost or Microsoft, so anyone telling you what this is worth is guessing. And Cloudflare’s take rate is not disclosed anywhere: the Open Markets Institute’s April 2026 survey puts it around 30%, against ScalePost at roughly 15% and TollBit charging the AI company rather than the publisher,15 but that figure is the researchers’ rather than Cloudflare’s, and for a company this transparent about its outages it is a strange gap to leave.

It is already being replaced#

The odd thing about writing this in August 2026 is that Cloudflare has largely conceded the central assumption.

Their July 2026 announcement moves from Pay Per Crawl to Pay Per Use, paying publishers when their content creates value in an AI answer rather than when a page is fetched.16 The justification reads as a self-critique: “Cloudflare data suggests that over 50% of crawl traffic from AI crawlers is spent re-fetching unchanged pages.”16 Set that beside the FAQ line confirming crawlers are charged again for re-crawls and the problem is plain. Per-fetch pricing bills for waste, and billing for waste gives neither party a reason to reduce it. Ceramic.ai, You.com, beehiiv, Condé Nast and Patreon are the named launch partners.

The same announcement sets a date. From 15 September 2026, on ad-monetised pages, Search bots stay allowed while Training and Agent bots are blocked by default, and multi-purpose crawlers that cannot separate the three follow the most restrictive rule.17 Free-plan customers who change nothing inherit the new defaults, so if you run a site behind Cloudflare that belongs in a calendar rather than being discovered.

The ambition has widened past crawling as well. The Monetization Gateway, announced the same day, applies the idea to any asset behind Cloudflare: pages, datasets, APIs, MCP tools. It runs on x402, Coinbase’s payment protocol, which also uses HTTP 402 but with an X-PAYMENT header and stablecoin settlement.18 The x402 Foundation launched under the Linux Foundation on 14 July 2026 with forty founding members including Cloudflare, Google, AWS, Visa, Mastercard and Stripe,19 and Cloudflare added agent wallets in August.20 So a status code that went unused for thirty years now carries two incompatible conventions, both shipped by the same company, with crawler-price and X-PAYMENT doing the same job in different vocabularies. I would expect convergence, and would rather it happened at the IETF than in a changelog.

The criticism, which is better than the fandom admits#

I like this feature. The objections are also stronger than most coverage allows, and they are not hard to find.

Luke Hogg and Tim Hwang at Tech Policy Press wrote the sharpest version days after launch: given Cloudflare’s share of the reverse proxy and DDoS markets, “Cloudflare’s policies become internet policy for a sizable portion of the web.”21 Cloudflare says it fronts more than 20% of web domains.17 A default that one company flips on 15 September across that share of the web is a policy decision with no democratic input and no appeal.

Farzaneh Badii at Digital Medusa raises one I have not seen answered. There is no stable definition of publisher or creator online, so “one party could profit from content created by another, just because they host it.”22 Every platform sitting on user-generated content is now positioned to sell access to writing it did not produce, and the protocol has no opinion about who gets the money.

The Open Markets Institute argues publishers are squeezed by firms that both strip their traffic and control the licensing alternatives, “occupying both sides of the value chain simultaneously,” and warns that today’s take rates and governance norms will be hard to revise later. It also found local, regional and independent outlets effectively absent from this market.15 Notably, that report mapped Human Native as an independent intermediary, and Cloudflare had already bought it in January 2026,23 so the consolidation it warns about was underway as it was published.

The EFF argues that anonymous crawling is essential to journalism, research and security work, and that identification regimes let sites block “not just bad actors, but also security professionals, researchers, dissidents, or anyone who has not paid for a license.”24 In fairness that piece does not name Cloudflare or Web Bot Auth and attacks the category rather than the product, but the category is what is being built here, and archives are already the visible casualty: the EFF has separately documented major publishers blocking the Internet Archive, where the archived page is often the only record of how a story originally ran.25 And there is the blunt version from Vercel’s Guillermo Rauch: “The fastest path to irrelevance: blocking progress. The answer to AI is: more AI. Not to block and stagnate.”26 I disagree, but it is an argument a serious person makes.

Having worked through the protocol properly, my position is that a priced, signed, logged crawl beats an unpriced, unsigned, unlogged one for everybody including the crawler. The protocol is not the problem. The problem is that the enforcement layer is proprietary and the defaults belong to one company, which is why the open pieces are the ones worth supporting: Web Bot Auth at the IETF, RSL as a robots.txt-native licensing vocabulary backed by Reddit, Yahoo, O’Reilly, Medium and Fastly among others,27 and the IAB Tech Lab’s Content Monetization Protocols, finalised at V1 in April 2026.28 Those work regardless of who your CDN is. Pay per crawl only works if it is Cloudflare.

What I would do#

If you run a site behind Cloudflare, look at AI Crawl Control before 15 September whatever you decide, because inheriting the new defaults by accident is the worst available outcome. If you have anything worth crawling, join the beta for the analytics and treat revenue as a bonus. Add a Content Signals line to your robots.txt today, since it costs nothing and states your intent independently of any vendor.29

If you operate a crawler, implement Web Bot Auth before somebody makes you: Ed25519 key, signed directory at the well-known path, tag="web-bot-auth" on the signature input. It is an afternoon of work and it stops you being treated as hostile traffic. Then read the pay per crawl changelog rather than only the Web Bot Auth documentation, because the payment-header signing requirement lives in one and not the other, and building from the wrong page is precisely how I got it wrong. Add crawler-max-price with a real budget, handle the crawler-error codes properly, and instrument your re-fetch rate, because if half your crawl volume is unchanged pages you are burning money and goodwill in the same request.

And if you are looking for something to build on top of this, the opportunities are not in publishing. They are the small, high-value resources an agent needs repeatedly and cannot derive for itself: a clean dataset behind a metered endpoint, a specialist index, an API that returns a verified fact. Those have never been viable as consumer products because no person will pay for them one query at a time. That constraint has just been removed.

For the wider context on what else Cloudflare shipped this year, the rest of the series takes one launch at a time.


  1. 402 Payment Required, MDN Web Docs; the status code is defined in RFC 9110 §15.5.3↩︎

  2. The crawl before the click: Cloudflare data on AI bots, training and referrals, Cloudflare blog. ↩︎ ↩︎

  3. A deeper look at AI crawlers, by purpose and industry, Cloudflare blog. ↩︎

  4. Google users are less likely to click on links when an AI summary appears in the results, Pew Research Center, 22 July 2025. ↩︎

  5. Introducing pay per crawl: enabling content owners to charge AI crawlers for access, Will Allen and Simon Newton, Cloudflare blog, 1 July 2025. ↩︎

  6. Pay per crawl FAQ, Cloudflare developer documentation. ↩︎

  7. Pay per crawl advanced configuration, Cloudflare changelog, 16 June 2026. ↩︎

  8. What is pay per crawl, Cloudflare developer documentation, last updated 23 April 2026. ↩︎

  9. RFC 9421: HTTP Message Signatures, IETF, February 2024; Web Bot Auth, Cloudflare developer documentation, which specifies the ("@authority" "signature-agent") requirement and states the current position on nonce validation and expires↩︎ ↩︎ ↩︎

  10. Pay Per Crawl (private beta): Discovery API, custom pricing, and advanced configuration, Cloudflare AI Crawl Control changelog, 10 December 2025. ↩︎ ↩︎

  11. Web Bot Auth (webbotauth) working group, IETF. ↩︎

  12. webbotauth working group documents, IETF Datatracker. ↩︎

  13. Web Bot Auth and Signed Agents, Cloudflare blog. ↩︎

  14. Your content, your rules, Cloudflare press release, 1 July 2026. ↩︎

  15. Courtney Radsch and Karina Montoya, Same Gatekeepers, New Tollbooths: Mapping the AI Content Licensing Market, Center for Journalism and Liberty / Open Markets Institute, April 2026. ↩︎ ↩︎

  16. Your content, your rules, Cloudflare press release, 1 July 2026. ↩︎ ↩︎

  17. Content Independence Day: your content, your rules, Jin-Hee Lee and Bryan Becker, Cloudflare blog, 1 July 2026. ↩︎ ↩︎

  18. Monetization Gateway, Rohin Lohe, Justin Ridgely and Will Papper, Cloudflare blog, 1 July 2026; x402, Coinbase. ↩︎

  19. Linux Foundation announces operational launch of the x402 Foundation, 14 July 2026. ↩︎

  20. Cloudflare Wallets, Cloudflare blog, 4 August 2026. ↩︎

  21. Luke Hogg and Tim Hwang, Cloudflare’s Troubling Shift From Guardian to Gatekeeper, Tech Policy Press, 9 July 2025. ↩︎

  22. Farzaneh Badii, The Enclosure of the Open Web and the Open Internet Toll Booth, Digital Medusa, 9 July 2025. ↩︎

  23. Cloudflare strengthens content offering to AI companies with acquisition of Human Native, 15 January 2026. ↩︎

  24. Tori Noble, ‘Stealth Crawlers’ Are Not a Threat to the Open Web. Bills Targeting Them Would Be., Electronic Frontier Foundation, 20 July 2026. The piece critiques mandatory crawler identification generally and does not name Cloudflare. ↩︎

  25. Joe Mullin, Blocking the Internet Archive Won’t Stop AI. It Will Erase the Web’s Historical Record., Electronic Frontier Foundation, 16 March 2026. ↩︎

  26. Quoted in Adam Hua, Cloudflare’s AI Tollbooth: Innovation or Gatekeeping?, AdMonsters, 7 August 2025. ↩︎

  27. Really Simple Licensing and the RSL 1.0 specification↩︎

  28. Content Monetization Protocols (CoMP), IAB Tech Lab. ↩︎

  29. Giving users choice with Cloudflare’s new Content Signals Policy, Cloudflare blog, 24 September 2025; contentsignals.org↩︎

Read the original on cybersins.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.