RSS Amplifier

Privacy Pointers · Aug 6, 2026

"Great Scrape" in the age of AI

0
Sign in to vote or save

Swati Popuri · Privacy Pointers

Ask most AI companies about their training data and you get the same line: it’s all publicly available. Clearview AI says it. OpenAI has said it. hiQ Labs said it when it scraped half a billion LinkedIn profiles to sell “people analytics” to employers.

On the product side, this feels almost self‑evident: if people put something on the open web, why shouldn’t we be allowed to read it, copy it, and learn from it at scale?

Daniel Solove and Woodrow Hartzog’s paper The Great Scrape: The Clash Between Scraping and Privacy pushes hard against that intuition. Their core claim is simple but uncomfortable for anyone building AI: the “it’s public” defense is doing far more work than it can bear, and scraping as we know it is fundamentally misaligned with how modern privacy law and norms are supposed to work.

Most privacy law quietly assumes a relationship model: a person gives data to an organization, and that organization takes on duties - inform, limit purposes, secure data, offer rights.

Scraping breaks that model.

Solove and Hartzog walk through the Fair Information Practice Principles and show how scraping collides with almost all of them.

  • Notice and consent are largely fictional. There’s no realistic way to inform or obtain opt‑in from millions of people before a crawler hoovers up their posts, profiles, or photos.

  • Purpose limitation is vague. “AI training” is not a specific, bounded purpose; it’s a commitment to future, open‑ended use.

  • Data minimization runs in reverse: collect first, decide what’s useful later.

  • Rights and control become mostly symbolic once data is scraped into opaque systems with no clear path back to the people underneath.

It’s not a relationship; it’s a one‑sided taking.

The move Solove and Hartzog show how slippery “public” actually is. In scraping debates, it can mean:

  • Descriptively accessible – technically reachable if you know where to look.

  • Officially designated public – like a court record.

  • “Not private” by assumption – the vague idea that anything on the internet is up for grabs.

Product teams tend to blur these categories and treat the loosest sense “someone could have seen it” as if it wipes away all privacy claims. That shortcut ignores what actually changes when we move from individual human access to industrial‑scale collection, aggregation, and repurposing.

In the AI era, that’s the key point: scale and reuse create new risks. The harm isn’t just that one more person saw a post; it’s that millions of posts, photos, and profiles become raw material for systems that profile, predict, and intervene without any real relationship to the people underneath.

For privacy and governance teams, the shift is less “stop scraping” and more “own the theory of scraping you’re willing to defend.”

  • Stop letting “public” end the discussion. When teams or vendors say “we only use public data,” treat that as the start of your risk analysis. Ask what’s being scraped, how it’s aggregated and repurposed, and who actually benefits versus who bears the risk.

  • Make scraping explicit in design and governance. Any product or model that leans on scraped data should trigger deliberate assessment: what sites and communities are affected, what legal theory you’re relying on, whether you’re honoring robots.txt and TOS, and what alternatives you considered.

  • Center people, not just platforms. Most scraping disputes so far hiQ v. LinkedIn being emblematic have been framed as fights over property and contracts between companies, with data subjects effectively absent. Your DPIAs, LIAs, and AI governance discussions don’t have to repeat that pattern; they can explicitly ask whose lives are being folded into training data and what they reasonably expect.

Try to answer a simple question: if your AI strategy depends on large‑scale scraping, can you defend that strategy without hiding behind the word “public”? This paper is a reminder that scraping is not a law of nature. It’s a design choice about whose data gets taken, for whose benefit, under whose terms.

Read the original on privacypointers.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.