Update: Never mind! It turns that Google’s issue is with unreachable robots.txt files, not absent robots.txt files. They really need to improve their messaging. Stand down everyone.
A bit has been flipped on Google Search.
Previously, the Googlebot would index any web page it came across, unless a robots.txt file said otherwise.
Now, a robots.txt file is required in order for the Googlebot to index a website.
This puzzles me. Until now, Google was all about “organising the world’s information and making it accessible.” This switch-up will limit “the world’s information” to “the information on websites that have a robots.txt file.”
They’re free to do this. Despite what some people think, Google isn’t a utility. It’s a business. Other search engines are available, with different business models. Kagi. Duck Duck Go. Google != the World Wide Web.
I am curious about this latest move with Google Search though. I’d love to know if it only applies to Google’s search bot. Google has other bots out crawling the web: Adsbot-Google, Google-Extended, Googlebot-Image, GoogleOther, Mediapartners-Google. I’m probably missing a few.
If the new default only applies to the searchbot and doesn’t include say, the crawler that’s fracking the web in order train Google’s large language model, then this is how things work now:
- Your website won’t appear in search results unless you explicitly opt in.
- Your website will be used as training data unless you explicitly opt out.
It would be good to get some clarity on this. Alas, the Google Search team are notoriously tight-lipped so I’m not holding my breath.
Responses
Related posts
Permission
You have the power, not Google.
Related links
Tagged with google search ai machinelearning language models enshittification permission robots crawling wordpress
A principled approach to evolving choice and control for web content
This would mean a lot more if it happened before the wholesale harvesting of everyone’s work.
But I’m sure Google will put a mighty fine lock on that stable door that the horse bolted from.
Tagged with google search ai machinelearning language models enshittification permission robots crawling licensing licenses
Previously on this day
2 years ago I wrote A long-awaited talk
Remy and I gave a talk at Brighton’s Async meetup …five years after we were originally booked in.
8 years ago I wrote Writing for hiring
Robot-free recruiting.
11 years ago I wrote Separated at death
Farewell, doppelgänger.
13 years ago I wrote Chüne
Churn on, Chüne in, chrop out.
17 years ago I wrote Making Workshops for the Web
Behind the scenes of the latest Clearleft site.
18 years ago I wrote Authors On Tour — Live!
For your huffduffing pleasure.
19 years ago I wrote Ten songs titles that could be Twitter updates
The kind of list that’s too geeky for McSweeney’s.
20 years ago I wrote The Best Songs I Acquired in 2006 Ever
Following Richard’s lead.
22 years ago I wrote DHTML is dead. Long live DOM Scripting.
Just in case I haven’t completely hammered the point home lately, I have a feeling that 2005 is going to see a big surge in the use of the Document Object Model with JavaScript.
24 years ago I wrote I've seen fire and I've seen rain
…but mostly rain.
25 years ago I wrote Which Kevin Smith character are you?
Excellent! I am Silent Bob, apparently:
25 years ago I wrote Biz Stone: Wrong Font
Take a look at this picture of a storefront, it’s a great example of how not to choose a font.