One file, three different decisions
If you have been asked whether to “block AI”, the question is already too coarse to answer. The crawlers reaching your site today do at least three unrelated jobs, and the major vendors have separated them deliberately. One user agent collects content that may contribute to training a foundation model. A second builds the search index that an assistant consults when it needs current information. A third fetches a specific page because a person asked the assistant a question and the assistant went to look. These are different requests with different consequences for your business, and they are controlled by different lines in the same file.
The reason this matters commercially is simple. A company may have a considered position that its writing should not be absorbed into a training corpus, while also wanting its service pages to appear when a prospective client asks an assistant who does this kind of work. Those two positions are compatible. They only become contradictory when a single blanket rule is applied to every user agent whose name contains a vendor's brand. Plenty of sites have quietly removed themselves from assistant answers while intending to protect training data, because one line in robots.txt did both.
This guide walks through what each vendor currently documents, what a directive does and does not control, and how to write a file you can defend in a meeting. It does not tell you which policy is right for your organization. Licensing, competitive exposure, editorial rights and the value you place on assistant visibility are business judgments, and they differ legitimately between a law firm, a publisher and a B2B services company.
References: OpenAI — Overview of OpenAI CrawlersGoogle Search Central — Introduction to robots.txt
Training, indexing and fetching are not the same request
A training crawler reads pages to build a corpus. The content may influence a model's future behavior, and nothing about that request produces a link back to you. Whatever the outcome, it is not visibility. This is the request most organizations mean when they say they want to opt out, and it is the one with the clearest connection to questions about rights and compensation.
A search crawler is doing something much closer to what Googlebot does. It builds an index so an assistant can retrieve and cite current pages when a user's question requires information the model does not carry. When an assistant names a source and links to it, that link usually exists because a search crawler was allowed to read the page. Blocking this crawler is the choice that removes you from answers.
A user-initiated fetch is different again. Someone is sitting in front of the assistant, has asked about a specific topic or pasted a specific address, and the assistant retrieves that page to answer. Vendors describe this as an action taken on a person's behalf rather than autonomous crawling, and several document that robots.txt may not govern it for that reason. Whether you consider that acceptable is a real question, but it is a question about a person using a tool to read your public page, not about bulk collection.
Once you hold these three apart, the policy conversation becomes tractable. You are no longer asking whether you are for or against AI. You are asking three narrower questions: may our content contribute to training, do we want to be findable inside assistants, and are we comfortable with a reader's tool retrieving a public page on request.
References: OpenAI — Overview of OpenAI CrawlersAnthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler?Perplexity — PerplexityBot and Perplexity-User
What OpenAI documents
OpenAI's crawler documentation lists separate user agents for separate purposes. GPTBot is described as being used to make the company's generative AI foundation models more useful and safe; disallowing it indicates that your content should not be used in training those models. OAI-SearchBot is described as being used to surface websites in search results in ChatGPT's search features, and disallowing it is how a site opts out of appearing there. OAI-AdsBot validates the safety of pages submitted as ads, and only visits pages that were submitted.
ChatGPT-User covers certain user actions in ChatGPT and custom GPTs, where the visit happens because a person asked for it. OpenAI's documentation notes that robots.txt rules may not apply to these user-initiated actions, and that this agent is not used for automatic crawling or search indexing. That distinction is worth reading carefully before you assume a robots.txt line removes you from every possible retrieval path.
The practical consequence is that a rule written only for GPTBot does not affect ChatGPT search visibility, and a rule written for every OpenAI agent does. Sites that wanted the first and wrote the second are common. If your organization's decision was about training data, write the line that is about training data.
One implementation detail matters for anyone maintaining documentation or internal runbooks: the older documentation address at platform.openai.com now answers with a redirect to the developers.openai.com location. If your internal reference links still point at the old address they will still resolve, but the canonical page has moved, and anything you have copied from a year-old blog post deserves rechecking against the live document.
References: OpenAI — Overview of OpenAI Crawlers
What Anthropic documents
Anthropic publishes three crawlers with the same three-way split. ClaudeBot, in Anthropic's words, “helps enhance the utility and safety of our generative AI models by collecting web content that could potentially contribute to their training”. Restricting it signals that a site's future material should be excluded from those training datasets.
Claude-SearchBot “navigates the web to improve search result quality for users” and analyzes content to improve the relevance and accuracy of search responses. Anthropic states plainly what blocking it costs: doing so prevents the system from indexing your content for search, “which may reduce your site's visibility and accuracy in user search results”. Claude-User supports people asking Claude questions, and disabling it “prevents our system from retrieving your content in response to a user query, which may reduce your site's visibility for user-directed web search”.
It is unusual and useful that the vendor states the consequence of each block in its own documentation. You do not have to speculate about the trade-off or rely on a third-party estimate of it. If you block the search and user agents, the company that operates the assistant says you may become less visible inside it. That is the sentence to put in front of whoever is making the decision.
The page listing these agents carries its own last-updated date, which is a reminder rather than a footnote: this is a moving target. User agents have been renamed before, and earlier Anthropic identifiers were retired. A robots.txt written from a blog post published two years ago is likely to be naming agents that no longer exist while missing the ones that do.
References: Anthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler?
What Perplexity documents
Perplexity documents two agents. PerplexityBot is “designed to surface and link websites in search results on Perplexity”, and the documentation states that it “is not used to crawl content for AI foundation models”. For a site whose objection is specifically to training, that sentence changes the calculation: blocking this agent costs visibility without serving the stated policy goal.
Perplexity-User covers visits made because a user asked a question, and the documentation says that since a user requested the fetch, this fetcher generally ignores robots.txt rules. Again, that is an argument for deciding your position on user-initiated retrieval as a separate matter rather than assuming a file governs it.
Perplexity also publishes IP ranges for both agents. This is the part most teams skip and the part that makes verification possible. A user-agent string is self-reported and trivially forged, so a log line claiming to be PerplexityBot is evidence of nothing on its own. Published ranges let you confirm that traffic attributed to a vendor actually originated with it, which matters both for diagnosing whether your directives are being honored and for investigating scraping that is wearing someone else's name.
References: Perplexity — PerplexityBot and Perplexity-User
Google-Extended and Applebot-Extended: an AI opt-out that is not a search opt-out
Google and Apple handle this differently again, and the difference trips people up. Their AI-related controls are not crawlers at all in the usual sense. They are tokens you place in robots.txt to signal how content already crawled for search may be used for certain AI purposes. Blocking them is not the same as blocking the search crawler, and it does not remove your pages from the search index.
This design exists because the same crawl serves multiple purposes at those companies. It has a consequence worth stating explicitly to anyone hoping for a clean opt-out: Google's own guidance for AI features tells site owners that appearing in AI Overviews and AI Mode follows the same foundations as appearing in Search, and that no separate machine-readable files or special structured data are required. The eligibility question and the crawl question are not the same question, and neither is answered by adding a new file.
If your team's goal is to appear in Google's AI features, the work is the ordinary work: crawlable pages, useful content, accurate metadata, structured data that matches what a visitor can see. If the goal is to limit how content is used beyond search, the extended tokens are the documented lever, and the trade-offs are specific to each vendor's stated behavior rather than uniform across the industry.
References: Google Search Central — AI features and your websiteGoogle Search Central — Introduction to robots.txt
robots.txt is a crawl instruction, not access control
Three limits apply to everything above, and skipping them is how well-intentioned configurations fail. First, robots.txt asks compliant crawlers not to fetch a URL. It does not prevent a URL from being indexed if other pages link to it, and Google's documentation is explicit that blocking crawling is the wrong tool for keeping a page out of results. If a page must not appear, you need a noindex directive on a page that is allowed to be crawled, or authentication.
Second, it is not a security boundary. A file that lists your admin paths under Disallow is a public document announcing where they are. Anything confidential belongs behind authentication, and the site's own rule of keeping private routes protected applies regardless of what any crawler promises.
Third, compliance is voluntary. Well-known vendors publish their agents and generally honor directives, which is exactly why writing careful directives is worth doing. Scrapers that do not identify themselves are unaffected by any of this, and no line in a text file will change that. If unattributed scraping is your actual concern, the response is rate limiting, bot management at the edge, or authentication for the material that matters, not a longer robots.txt.
There is one more failure mode worth naming, because we have seen it in production: a site can serve a robots.txt that reflects an environment variable rather than an intention. Preview deployments that disallow everything are correct in preview and catastrophic in production if the flag is wrong. Whatever you decide, the check is to request the live file from the public address and read what it actually says.
References: Google Search Central — Introduction to robots.txt
A decision table you can defend
The table below maps the three decisions onto the agents each vendor documents. Use it to write down a position per column rather than per vendor. Most organizations we work with land on the same shape: no strong objection to being found, a considered position on training, and an acceptance that a reader using an assistant to open a public page is a reader.
Whatever you choose, record the reasoning next to the file. Six months from now someone will ask why a line exists, and “a consultant added it” is not an answer that survives a visibility investigation.
| Vendor | Model training | Search indexing | Fetch requested by a person |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User (robots.txt may not apply) |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | Not used for foundation models | PerplexityBot | Perplexity-User (generally ignores robots.txt) |
| Google-Extended token | Googlebot (also serves AI features) | Not separately documented as a distinct agent | |
| Apple | Applebot-Extended token | Applebot | Not separately documented as a distinct agent |
References: OpenAI — Overview of OpenAI CrawlersAnthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler?Perplexity — PerplexityBot and Perplexity-UserGoogle Search Central — AI features and your website
What a defensible file looks like
A robots.txt you can defend has three properties: every group exists because somebody decided it should, the groups are named specifically enough that a future change does not silently invert them, and the reasoning is recorded somewhere a person will find it. The file itself cannot hold the reasoning, which is why a short note in your repository next to it is worth more than another ten lines of directives.
The most common shape for a B2B services company that wants to be found is short. Allow general crawling. Keep write surfaces and private routes out of it, and protect them properly elsewhere. Then either say nothing further, which means every documented agent inherits the general rule, or name the vendor agents explicitly so that a reader can see the decision rather than infer it. Naming them changes nothing functionally when the policy is uniform; it changes a great deal when someone later wants to know what was intended.
A publisher with a licensing position will want the training groups separated and disallowed while the search and user-initiated groups stay open. That configuration is coherent and it is exactly the one a blanket rule destroys. The point of writing it out per purpose is that it survives the next person to edit the file.
Whatever shape you choose, run through the checks below before you consider it done. They take a few minutes and they catch the failures that otherwise surface as an unexplained absence six months later.
- Request the live file from the public address, over HTTPS, and read the bytes that are actually served. Do not review the source in your repository and assume they match.
- Confirm the environment logic. A preview build that disallows everything is correct in preview; make sure the production build is not inheriting that flag.
- Check each vendor's current documentation for agents added or renamed since you last wrote the file, and remove groups naming agents that no longer exist.
- Verify that nothing confidential appears as a Disallow line. If a path must stay private, it belongs behind authentication, not behind a published list of what to avoid.
- Confirm your sitemap line resolves, and that the sitemap itself returns valid XML rather than an HTML error page.
- Record the date of the review and what changed, next to the file.
References: Google Search Central — Introduction to robots.txtOpenAI — Overview of OpenAI Crawlers
Verify with your logs, not with a generator
A robots.txt file is a statement of intent. Whether it is working is a question about your server, and the answer is in your access logs. Filter by user agent, look at what was requested and what status code you returned, and compare the result with what you meant to allow. A search agent that has not appeared in months on a site you believe is open is a finding worth investigating, not a statistic to accept.
Where a vendor publishes IP ranges, use them. Confirm that requests claiming a vendor's name originated inside that vendor's published space. Traffic that claims to be a well-known crawler and arrives from elsewhere is either a misconfiguration or someone borrowing a reputable name, and both are useful to know about. Reverse DNS verification serves a similar purpose for crawlers that support it.
Be careful about what you conclude from absence. Crawl frequency varies, small sites are visited less often, and a quiet week is not proof of a block. Look at a period long enough to be meaningful, note what you could not observe, and record the date of the check. This is the same discipline that applies to measuring citations: an honest denominator and a recorded observation window beat a confident number with no method behind it.
Finally, make this a recurring review rather than a one-time task. Vendors add agents, rename them and retire them. A quarterly check of the live file against the current documentation, with a note of what changed and why, costs very little and prevents the most common outcome we see, which is a site quietly invisible inside an assistant for a year because of a line nobody remembers writing.
References: Perplexity — PerplexityBot and Perplexity-UserAnthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler?
What this does not buy you
Allowing every search crawler does not mean an assistant will cite you. Access is a precondition, not a cause. What a system retrieves and chooses to name depends on the question asked, what else exists on the topic, how well your page actually answers it and factors no publisher observes directly. Anyone who tells you a robots.txt edit produces citations is selling something.
The honest framing is the inverse. A blocked search crawler is a reliable way to be absent. An allowed one puts you in the same position as everyone else who is allowed, at which point the work that distinguishes you is the ordinary work: a clear answer near the top of the page, claims that survive being quoted out of context, evidence a reader can check, and a site that serves its content in HTML rather than assembling it after the fact.
That is a less exciting conclusion than a configuration trick, and it is the one the vendors' own documentation supports. Get the access decisions right so they stop costing you silently, then spend your effort on the page itself.
References: Google Search Central — AI features and your website