Gerris home
Gerris

Chris Abraham: technical consulting, AI tools, and technical SEO

Home › Guides › robots.txt

How robots.txt works, and the mistakes that cost traffic

robots.txt is a plain text file at the root of a site that tells crawlers which paths they may fetch. It's small, it's public, and a single wrong line in it can take a site out of Google. Here's how it actually behaves.

Where it lives and what it covers

The file must sit at the root of the host, such as https://example.com/robots.txt. It applies only to that exact protocol and host: www.example.com, shop.example.com, and the HTTP version each need their own. Google reads up to 500 kibibytes of it and generally caches it for up to a day, so changes take effect within about 24 hours.

The syntax

User-agent: *
Disallow: /cart/
Disallow: /search
Allow: /search/help

User-agent: GPTBot
Disallow: /

Sitemap: https://example.com/sitemap.xml

Google ignores Crawl-delay; Bing honors it. Google also stopped supporting Noindex lines in robots.txt in 2019.

Blocking crawling is different from blocking indexing

This is the most expensive misunderstanding. Disallow stops a crawler from fetching a page. It doesn't stop the URL from being indexed. If other pages link to a blocked URL, Google can index the address alone, with no content, and Search Console reports it as "Indexed, though blocked by robots.txt."

To keep a page out of results, let it be crawled and give it a noindex robots meta tag or an X-Robots-Tag: noindex header. Google has to fetch the page to see that instruction, so a page that's both disallowed and noindexed stays in the index, because the noindex is never read. For content that must stay private, use a login. robots.txt is public, and listing secret paths in it advertises them.

Status codes matter

A server or firewall that errors only on robots.txt can quietly stop all crawling, which is why it's one of the first URLs I check on a site that lost traffic.

AI crawlers

Each AI company publishes the user agents it uses, and each honors robots.txt. They split into crawlers that gather training data, such as GPTBot, ClaudeBot, and CCBot, and crawlers that fetch pages for live answers, such as OAI-SearchBot, Claude-SearchBot, and PerplexityBot. Google-Extended and Applebot-Extended are tokens rather than separate crawlers: they control whether content fetched by Googlebot and Applebot may be used for AI models, and blocking them leaves ordinary search crawling alone. The AI assistants guide lists which crawler feeds which product.

This site lets every AI crawler read everything, and keeps the Markdown copies of its pages out of traditional search engines, so they never compete with the HTML pages. Because a crawler follows only its own group, naming the search engines in one group and the AI crawlers in another gives each its own rules:

User-agent: Googlebot
User-agent: Bingbot
Allow: /
Disallow: /*.md$

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: PerplexityBot
Allow: /

Mistakes I find most often

Testing changes

Search Console's robots.txt report shows the version Google last fetched, when it fetched it, and any parsing errors. Before changing a live file, test the proposed rules against a list of important URLs with a parser that follows Google's rules, such as Google's open-source robots.txt library or a crawler like Screaming Frog with a custom robots.txt. Then confirm the change in Search Console after Google picks it up.

References