Insight

AI crawlers and robots.txt: what each token controls, and what blocking does not achieve

AI companies crawl the web to train models and to answer questions with live information. Robots.txt lets website owners allow or refuse many of these crawlers, but the tokens control different things, compliance is voluntary, and blocking one use can remove a site from another it wants.

Published by Somnium Digital

A wireframe of the Insight page: headline, supporting sections and a single call to action. Insight AI crawlers and robots.txt Get in touch 01 Why this became a bus… 02 How robots.txt works 03 The main tokens and w…

Why this became a business decision

For most of the web’s history, the main crawlers a business cared about belonged to search engines. Allowing them meant visibility; blocking them meant disappearing from search. The decision was simple.

Generative AI changed that. Some crawlers collect content to train large language models, others fetch pages to show answers or citations in AI assistants and AI search products, and some fetch pages when a user asks an assistant to read a specific link. A publisher selling subscriptions, a brand wanting to appear in AI answers and a company protecting proprietary documentation may want very different policies.

How robots.txt works

Robots.txt is a text file at the root of a domain that tells crawlers which paths they may fetch. Each rule group names a user agent, such as a specific crawler, and lists allowed and disallowed paths. The Robots Exclusion Protocol was standardised as RFC 9309 in 2022.

Robots.txt is a request, not a security control. Reputable crawlers follow it, but it does not prevent access by crawlers that ignore it, and it does not protect content that is publicly available. Anything genuinely confidential should be protected by authentication, not by robots rules.

It also only affects future crawling. Content collected before a rule was added may already be in datasets or indexes, and blocking a crawler does not remove it.

The main tokens and what they control

AI companies publish the user agent names they use and what each does. The details change, so documentation from each provider should be checked when setting policy. Several important examples illustrate the distinctions.

GPTBot
OpenAI’s crawler for collecting content that may be used to train its models. OpenAI separately documents a crawler for surfacing sites in ChatGPT search results.
Google-Extended
A robots.txt token that controls whether content crawled by Google can be used for Gemini models and related grounding. It does not affect inclusion or ranking in Google Search.
Applebot-Extended
A token that lets sites opt out of their content being used to train Apple’s foundation models, while Applebot can still crawl for Apple’s search features.
ClaudeBot
Anthropic’s crawler for collecting web content, which follows robots.txt directives.
CCBot
Common Crawl’s crawler, whose openly available datasets have been widely used in AI research and model training.
PerplexityBot
Perplexity’s crawler for surfacing and linking websites in its answer engine.

Google Search and AI Overviews

A common misunderstanding is that blocking Google-Extended removes a site from AI features in Google Search. It does not. AI Overviews and similar features are part of Google Search and use content crawled by Googlebot. Blocking Googlebot would remove the site from Search entirely.

Google’s controls for how content appears in Search, such as the nosnippet directive or data-nosnippet attributes, also limit what can be shown as snippets and in AI features. They can reduce how content appears, but they also affect ordinary search snippets, so they should be applied carefully to specific sections rather than whole sites.

Trade-offs for business websites

For most service businesses, agencies, manufacturers and B2B companies, public web content is marketing. Being cited or summarised accurately in AI assistants can create awareness and enquiries, much as search visibility does. Blocking all AI crawlers may reduce that visibility without meaningfully protecting anything.

Publishers, research firms, course providers and companies whose content is the product face a different calculation. Their content may be used to answer questions without visits or payment, and they may choose to block training crawlers, negotiate licences or place premium content behind logins.

A middle path is common: allow crawlers that surface and link to the site in search-like products, block crawlers used only for model training, and protect genuinely valuable material with authentication. Whatever the policy, it should be a deliberate decision recorded by the business, not an accident of a template or plugin.

Implementation and monitoring

Keep robots.txt simple and tested. A misplaced rule can block search engines from the whole site, and a single typo in a user agent name means the rule applies to nothing. Test the file with the tools search engines provide and review it after website migrations.

Check server logs or analytics from the hosting or content delivery network to see which crawlers actually visit, how often and which paths they request. Some infrastructure providers offer settings to identify and block known AI crawlers at the network level, which can be useful where crawlers ignore robots rules or create heavy load.

Review the policy periodically. AI crawler names, purposes and industry practices continue to change, and a rule set written a year ago may no longer reflect what each company does with the content it collects.

Questions

Does blocking Google-Extended remove my site from Google Search?

No. Google-Extended controls use for Gemini models and grounding; it does not affect Search inclusion or ranking.

Does blocking Google-Extended stop AI Overviews using my content?

No. AI Overviews are part of Google Search and use Googlebot’s crawl. Snippet controls can limit how content appears.

Is robots.txt legally binding?

It is a technical standard that reputable crawlers follow voluntarily. It does not physically prevent access.

What is GPTBot?

OpenAI’s crawler for collecting content that may be used to train its models.

Will blocking a crawler remove content already collected?

No. Robots.txt affects future crawling, not data already gathered.

Should a B2B company block AI crawlers?

Often not all of them, because public marketing content can gain visibility in AI assistants. Proprietary material should be behind authentication.

How can we see which AI crawlers visit our site?

Through server logs or reports from hosting and content delivery network providers.

Where this sits in what we do

This article covers one decision inside a wider engagement. The solution page sets out how that engagement runs, what it includes and what it costs to find out.

Need a deliberate policy for AI crawlers?

We review your robots rules, logs and content types, then configure crawler access and protection that match what you want to be visible.

Get in touch

Tell us what you are trying to change

Describe the problem rather than the service — the two frequently differ, and working out which is which is the useful part of a first conversation. We reply within one working day, and if it is outside what we do well you will hear that in the reply rather than after a call.

We use what you send to reply to you. Nothing else, and no list.

WhatsApp