Why I Disallow Web Crawlers
When migrating my blog to Cloudflare a few days ago, I noticed that the robots.txt file had gone missing and fixed it. The robots.txt file has disallowed web crawlers on this blog almost from day one. I was reminded why I disallow web crawlers in the first place.
About 30% of the posts I’ve ever written for this blog are unpublished. They were either:
- never published because they didn’t meet my own quality control
- published and later removed because I changed my mind about whether they met my quality control
- published and later removed because I decided they were pointless
- only half written
In the modern world, every email, text message, and photograph is stored. Some of it indefinitely. In his autobiography, Edward Snowden called this a ‘permanent record.’ It’s unlikely you can even go shopping at the supermarket without recordings of you being stored somewhere.
I see my blog as a safe haven from the permanent record. One of the few places where I can be in control. I should be able to create, edit, and delete content at will and choose what versions the world sees. I don’t want my most inane writings or my most obvious mistakes to be available forever if I can help it.
Of course, non-compliant web crawlers may still scrape my blog anyway, and people may choose to store copies of my content for themselves. These copies are more ephemeral than they would be on the Web Archive or Google’s cache however.