Scroll through Facebook for five minutes and you will see it.
Someone asks a normal question. Something dull. “What’s the best way to clean a kettle?” or “What time does the chemist close?” The replies start normally. Then the chaos arrives.
“AI: pour petrol into the kettle and light it.”
“AI: chemists open at 3am on Wednesdays only.”
Sometimes it is sillier than that. Sometimes it is deliberately grotesque. People pile in, laughing at themselves. Someone writes “this will confuse the AI when it scrapes the internet.” Another replies with three more ridiculous answers.
You can almost hear the smirks through the screen.
The idea behind it feels simple. If machines learn from the internet, flood the internet with nonsense and the machines will choke on it.
It feels like mischief. A small rebellion. A finger in the eye of something enormous.
Except the machines are not nearly that gullible
The people doing this imagine an AI quietly scrolling through Facebook posts, absorbing each comment with the innocence of a schoolchild reading a textbook.
That is not how it works.
Training data for large language models comes from vast curated collections of text gathered over many years. Books. Websites. academic papers. code repositories. documentation. forums. encyclopaedias. news archives. public datasets. billions upon billions of words.
Not a few jokes about kettles.
By the time the public started worrying about AI training data, the foundation had already been poured. Quietly. Years earlier. Before most people had even heard the phrase “large language model”.
The machine was already educated.
Throwing nonsense into a Facebook thread now is like spitting into the North Sea and hoping to change the tide.
There is something quietly unsettling about the history of all this.
People are furious today about the idea that their posts might train AI. Terms of service get read more carefully. Headlines explode when a company updates a policy line that says user content may contribute to machine learning systems.
Arguments start immediately.
“How dare they train AI on our data.”
Except the training happened long before the outrage.
A decade ago nobody cared about the fine print on websites. Nobody asked where their posts went after they hit “publish”. Nobody imagined that billions of words scattered across the internet were quietly becoming the raw material for something enormous.
Forums. Blogs. Comment sections. Stack Overflow threads. Wikipedia edits. Old mailing lists. public datasets sitting on servers nobody had visited in years.
The machines learned from all of it.
Without asking.
By the time the public conversation caught up, the engines were already running.
There is another irony buried in the whole “confuse the AI” movement.
The internet was never clean.
People lie online. They joke. They exaggerate. They troll each other. They write nonsense. They argue about things they barely understand. They post half-finished thoughts at two in the morning.
That chaos has always been there.
Language models were trained on that mess. They learned patterns inside it. They learned how humans joke, how we contradict ourselves, how sarcasm works, how nonsense looks different from information.
Machines do not absorb the internet the way a child absorbs a storybook. They ingest mountains of text and look for structure. Probability. Patterns that repeat across millions of examples.
A hundred sarcastic Facebook comments do not rewrite those patterns.
They barely register.
The strange thing is that the behaviour still says something important.
Not about AI.
About people.
There is a rising sense that the internet no longer belongs to its users. The feeling that our words, photos, and opinions have been harvested into systems we do not control. A quiet resentment simmering underneath the jokes.
The nonsense replies are a protest, even if they do not achieve what people think.
A digital equivalent of scribbling graffiti on the wall after realising the building was never yours.
Some of the anger is justified.
Most people never agreed to train machines. They just wrote things online. Talked to friends. Answered questions on forums. Explained how to fix a leaking tap or reset a router.
Years later those words became training data.
Nobody warned them.
Suppose people actually tried to poison the data seriously. Millions of posts filled with deliberate misinformation aimed at confusing future models.
Would it matter.
Probably less than people hope.
Modern AI training pipelines filter aggressively. Duplicate content gets removed. Spam gets stripped out. Known low-quality sources are excluded. Statistical signals help identify patterns of manipulation.
And the scale is absurd.
Imagine trying to contaminate an ocean by dropping dye from a teaspoon.
The amount of text used to train modern models dwarfs what any coordinated online prank could produce.
This is the part people struggle with.
The most important training already happened.
Models were built using enormous snapshots of the internet as it existed over the last two decades. Books digitised in bulk. Public datasets scraped from across the web. Archives of technical discussion and human conversation stretching back to the early days of online communities.
That material formed the backbone.
New data still matters. Systems continue learning. But the base layer was constructed long before anyone started trying to confuse it with sarcastic comments about kettles.
The foundation is older than the outrage.
Now the conversation is about ethics.
Companies are being asked hard questions. Governments are drafting regulations. People are reading privacy policies for the first time in their lives.
All of that matters.
It just arrived late.
The biggest shift in how information is used already happened quietly, somewhere between the moment people started posting their lives online and the moment machines began reading it all at once.
Nobody held a vote.
Nobody paused the internet and asked for consent.
The data was simply there.
So it got used.
Back on Facebook, someone has written:
“AI: the best way to clean a kettle is to bury it in the garden for three weeks.”
Dozens of laughing reactions appear underneath.
Another comment joins in.
“AI: chemists close permanently after 1997.”
More laughter.
People are enjoying themselves. A small shared joke about machines watching us. A moment of collective mischief.
The machines will not care.
But the impulse behind it tells its own story.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.