“Block AI bots” sounds like one switch. It is not. The same company may operate one bot for search citations, another for model training, a third that opens a page because a user asked, and a fourth that validates an advertising landing page. A blanket rule can protect content you meant to reserve—or quietly remove a public business from the AI answers where customers now compare suppliers.
The useful default for a public ecommerce or lead-generation site is to allow verified search crawlers, decide on training separately, allow user-requested retrieval only for genuinely public pages, and keep private areas behind real authentication. Treat that as a starting policy, not copied code: bot names, provider behaviour, edge settings, and your commercial interests can change.
Why this needs a business decision now
AI-mediated discovery is already material upstream of a sale. In a December 2025 survey of consumers in France, Germany, and the UK, McKinsey found that 38% used AI tools to research products and services or make purchase decisions. That does not mean 38% of your revenue comes from assistants. It does mean crawl policy now touches customer acquisition as well as infrastructure and intellectual property.
The controls have also become more granular. OpenAI documents separate agents for ChatGPT search, potential training, user-triggered access, and ad-page validation. Anthropic separates model-development, search, and user-requested bots. Perplexity separates its search indexer from user-requested retrieval. Google uses Googlebot for Search and a separate Google-Extended product token for some Gemini training and grounding uses.
That separation creates a better choice than “all AI is welcome” or “block everything.” A shop may want citations and qualified referrals while declining future-model training. A documentation company may want users to ask an assistant about public docs but protect licensed examples. A publisher may reserve articles for subscribers while opening summaries and product pages.
First, separate four jobs hidden behind “AI crawler”
| Purpose | Business outcome | Typical control |
|---|---|---|
| Search and citation | Your public pages can be discovered, summarised, linked, and recommended | robots.txt, crawlable HTML, sitemap, WAF allow rule |
| Model development | Your content may contribute to future foundation-model training | Provider-specific training token plus contract or licensing policy |
| User-requested retrieval | An assistant opens a URL or page because a person asked it to | Authentication and WAF policy; provider rules vary on robots.txt |
| Advertising validation | A submitted ad landing page can be checked for safety and relevance | Allow only when using that ad product; verify official IP ranges |
OpenAI’s current crawler documentation makes the distinction explicit: OAI-SearchBot supports search visibility; GPTBot covers content that may be used for model training; ChatGPT-User supports user actions; and OAI-AdsBot validates pages submitted as ChatGPT ads. OpenAI says each setting is independent. Blocking GPTBot is therefore not the same as leaving ChatGPT search.
Anthropic uses the comparable roles ClaudeBot, Claude-SearchBot, and Claude-User. Perplexity says PerplexityBot is for search results rather than foundation-model training, while Perplexity-User handles user actions and generally ignores robots.txt. These differences are why a static “top AI bots” snippet ages badly. Link each rule to the provider’s current documentation and assign an owner to review it.
Choose policy by content zone, not only by domain
Start with a content inventory. Separate public marketing pages, product and category pages, editorial articles, help documentation, licensed or paywalled material, user-generated content, account areas, checkout, partner portals, staging environments, and internal files. Then give each zone an intended outcome.
| Zone | Likely posture | Reason |
|---|---|---|
| Public products and services | Allow verified search bots; decide training separately | Discovery and citations can create qualified demand |
| Original articles and public docs | Allow search; choose open, blocked, or licensed training | Retrieval value and reuse value are different decisions |
| Licensed or subscriber content | Expose only the intended preview; enforce access at origin or edge | robots.txt is not a paywall |
| Accounts, checkout, admin, internal APIs | Authenticate; minimise bot access; noindex where appropriate | Privacy and security outweigh discoverability |
| Staging and unpublished work | Require authentication and remove public links | A disallow rule can reveal the path without securing it |
For a typical public business site, the conservative position is: keep Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, and PerplexityBot eligible to crawl public pages; make an explicit company decision on GPTBot, ClaudeBot, and Google-Extended; and use authentication rather than crawler etiquette for anything private.
Google is an important exception to simplistic bot lists. Google-Extended is a control token, not a separate HTTP user-agent string. Google says opting out does not affect inclusion or ranking in Google Search, while the token does affect stated Gemini training and grounding uses. Googlebot remains the relevant control for Search, including its generative search features.
A copy-ready starting point—with boundaries
This sample leaves public pages available to ordinary and AI search crawlers, blocks three documented training controls, points crawlers to the sitemap, and lists placeholder private paths. Replace the paths and the training choice after review. Do not paste it over existing rules without testing group precedence.
User-agent: *
Disallow: /account/
Disallow: /admin/
Disallow: /internal/
# Decline documented model-training uses
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
Sitemap: https://www.example.com/sitemap.xmlWhy are the search bots not named? With no matching specific block, they inherit the public wildcard policy. Fewer explicit groups reduce the chance that a future path restriction is accidentally bypassed by a more specific allow group. If your WAF blocks automated traffic by default, however, you may also need a verified allow rule for the search crawlers you value.
Do not use this file to protect customer data or unpublished material. The Robots Exclusion Protocol standard says its rules are not access authorisation. Use sessions, permissions, HTTP authentication, network controls, or another real security mechanism. Also remember that blocking crawl is not the same as removing a URL from results: a provider may learn the address elsewhere and show a navigational link. If a public page should disappear from search, use the provider-supported removal or noindex route while the bot can still read the directive.
The rule in Git is not necessarily the rule bots receive
A correct repository file can be contradicted at the edge. CDN bot protection, managed robots.txt, a web application firewall, rate limiting, geo rules, a security plugin, or a reverse proxy can return a different file or a 403 response. Cloudflare’s current AI Crawl Control, for example, can allow or block individual crawler categories and create WAF rules; its dashboard also reports requests and observed robots.txt violations.
Verify from outside the application stack:
- Fetch the live
/robots.txtfrom each hostname and subdomain, not just the source file. - Request representative public pages with the relevant user-agent, then verify the source IP against the provider’s published ranges before trusting or allowing it.
- Check the HTTP status, redirects, canonical URL, indexability, language alternates, sitemap inclusion, and server-rendered text.
- Confirm CSS, JavaScript, images, and APIs needed to understand the public page are not blocked.
- Inspect edge and origin logs for 200, 301, 403, 429, and 5xx patterns; a bot listed as allowed can still fail downstream.
User-agent text alone is easy to spoof. OpenAI and Perplexity publish IP range files, Google documents verification methods, and major bot-management products maintain verified categories. Combine identity and IP or verified-bot signals in WAF rules. Do not create a broad “allow anyone claiming to be GPTBot” bypass around security controls.
Crawl permission is eligibility, not a ranking strategy
Allowing a search bot only opens the door. OpenAI explicitly avoids promising placement, and Google says indexing and serving are not guaranteed. A useful public page still needs a stable URL, crawlable links, visible text, a correct canonical, current facts, descriptive headings, appropriate structured data, and corroborating evidence.
Google’s May 2026 guidance is unusually direct about shortcuts: its Search systems do not use llms.txt for visibility or ranking, there is no special AI schema, content does not need artificial “chunking,” and ordinary people-first SEO remains foundational. Maintain an llms.txt file if a specific system or documentation workflow uses it; do not sell it internally as a ticket into AI answers.
For interactive pages, accessibility is also machine-readability. Semantic buttons, labels, states, headings, forms, and predictable error messages help people using assistive technology and browser agents operating through the accessibility tree. That is product quality with a second distribution benefit—not a trick for bots.
Measure value before opening more
Keep three measurements separate. Crawl activity shows that a bot requested pages. AI referrals show identifiable visits from an assistant or search surface. AI-influenced demand includes recommendations with no observable click and is partly inferred. A crawl is not a lead, and a mention is not revenue.
Review requests by verified crawler, allowed versus blocked status, content zone, response code, bytes, and infrastructure cost. For traffic, track source, landing page, qualified action, pipeline or order, margin, and return behaviour. OpenAI’s publisher FAQ says ChatGPT search referrals include utm_source=chatgpt.com; use it as evidence when present, not proof that all influence is visible.
A quarterly policy review is usually enough for a small site, with an immediate review after a CDN migration, bot-protection change, paywall launch, new ad channel, provider documentation change, or incident. Version the policy, record the business owner, and keep a rollback path.
A seven-day crawler-policy audit
The deliverable should fit on one page: each content zone, intended outcome, bot or product token, enforcement layer, verification method, owner, approval date, and next review. The configuration can be short; the reasoning must be recoverable when a platform or colleague changes it six months later.
What this policy cannot guarantee
Compliant crawlers may change names, IP ranges, products, or behaviour. Some user-triggered agents do not treat robots.txt like scheduled crawlers. Unknown or hostile scrapers can ignore it completely. Search systems can discover a URL through third parties, and an allowed crawl does not guarantee a citation. Conversely, blocking future crawls cannot remove copies already collected or governed by another agreement.
This is a technical and commercial control framework, not a conclusion about copyright, text-and-data-mining reservations, privacy law, or a specific licence. If content has material licensing value or legal risk, align the technical signals with counsel and contracts.
Frequently asked questions
Should a business block GPTBot?
Decide based on the value and rights attached to your content. Blocking GPTBot signals that future crawled content should not be used for OpenAI foundation-model training; it does not require blocking OAI-SearchBot, which is the separate control for ChatGPT search.
Will blocking AI training bots hurt SEO?
Not automatically. OpenAI, Anthropic, and Google document separate controls for training and search. In particular, Google says Google-Extended does not affect inclusion or ranking in Google Search. Test the actual configuration because a blanket WAF rule can still block search bots unintentionally.
Do we need llms.txt to appear in ChatGPT or Google AI results?
No provider guarantees ChatGPT inclusion through an llms.txt file, and Google explicitly says it ignores llms.txt for Search visibility and ranking. Use it only when a system you care about documents a concrete purpose for it.
Is robots.txt enough to protect private or paid content?
No. The standard is a voluntary crawl protocol, not access control. Put private, licensed, account, and staging content behind authentication or an equivalent enforcement layer.
How can we tell whether an AI bot is genuine?
Match the request against the provider’s current published IP ranges or a trusted verified-bot signal, not the user-agent string alone. Monitor logs after allowing it and avoid broad firewall bypasses.
Make the policy boring, explicit, and measurable
The durable decision is not “AI: allow” or “AI: block.” It is a small matrix: which public content can support discovery, which material may support model development, what a user-requested agent may retrieve, what must remain authenticated, and how the live edge configuration proves the policy is working.
Rendframe can audit the crawl and delivery layer—live robots.txt, CDN and WAF rules, server rendering, structured data, accessibility, localized canonicals, logs, and conversion instrumentation—then implement the approved policy through our product engineering and agentic-commerce work. Send us the domains and content zones you need reviewed.
Continue reading: assess the commerce layer with the AI shopping-agent readiness audit, then measure results with the AI commerce attribution framework.
Sources and review date
Reviewed 20 August 2026 against the official OpenAI crawler documentation and publisher FAQ; Anthropic crawler controls; Perplexity crawler documentation; Google’s 2026 generative-search guidance and Google-Extended reference; Cloudflare AI Crawl Control; RFC 9309; and McKinsey’s European consumer research. Provider policies and identifiers change; verify current documentation before deployment.