RSS Amplifier

Chris Green Search Marketing (SEO/AEO) · Aug 25, 2026

When Blocking Bots Becomes 3D Chess

0
Sign in to vote or save

Chris Green · Chris Green Search Marketing (SEO/AEO)

For SEO’s bot control has gone from a relatively simple decision that concerned a handful of people to a full-blown battle ground which has many new stakeholders and is now more than just a “techie issue”.

Skip the “blah” and go direct to my “Robots Path” testing tool.

I started the year with some more detailed thoughts on this and I think the shape of it holds up, but it misses some of the actionable sides of this.

Blocking bots primarily takes place:

  • on a CDN or Edge or

  • you signal to bots that their crawling should be stopped/limited via the robots.txt OR

  • possibly using robots directives at the page level.

Most SEOs I’d argue aren’t having conversations about blocking bots at a CDN/Edge level & that is another post entirely - but it something you MUST be thinking about. I’d argue this is the most important element, whilst being least accessible of all the above options. It’s the only place you can block a bot with any degree of confidence, assuming they present themselves correctly to you.

Out of our other bot management methods - I spent a lot of time looking into the details of robots.txt and meta robots directives on my work leading the Web Almanac SEO Chapter. It seems clear that the shape of the web (robots.txt in particular) is starting to reflect the changes we have to deal with.

Understanding the implications of blocking/restricting different bots is not always straightforward. User agents change, some websites vary in complexity and there may be elements of search an organisation will except and others they will not.

The easiest way to check this - in theory - is via what is being set in the robots.txt. Even though bots can and will disrespect robots.txt at times, we can assume that most Search/AI bots display good behaviour, and that a robots.txt is a clear declaration of what they do and don’t want.

But if you have a large/complex site, even reading/understanding the rules in the robots.txt can be a challenge. What gets blocked, what doesn’t? What are the implications of blocking some URLs/Paths and not others?

To answer those questions I build “Robots Path” https://chr156r33n.github.io/robots-path/

Paste in the robots.txt file you want to check (or test changes with an existing one), add in the site origins and then the paths you wanted to check.

This provides an at-a-glance list for that URL:

Followed by a more detailed breakdown if you want to see the offending rules and potentially implications based on that bot user-agent and what it is thought to impact.

This is an early version, I’ll change and update more as I experiment and encounter new issues - but I’d love your thoughts/feedback if you have any!

No posts

Read the original on chrisgreenseo.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.