Skip to content
EN

AI Crawlers, Search Access and Training Controls

Jack Lee

Written by Jack Lee

Developer managing three connected screens

Can an AI product find your pages in search without using them to train a model? On some platforms, yes. The right control depends on what the product is doing.

Three Different Jobs Behind an AI Visit

Search discovery is an automated visit that helps a product find pages it might link in an answer. Model training is a separate use of collected material to improve future models. User-triggered retrieval happens when someone asks a product to open or answer from a specific page. A company may use different agents or settings for these jobs.

A visit does not guarantee a citation. Blocking one agent may leave the others unaffected. For the wider picture, see our AI search introduction.

Which Controls Belong to Which Platform?

PlatformSearch or answer accessTraining or other control
OpenAIOAI-SearchBot is used for ChatGPT search.GPTBot concerns content that may be used for foundation-model training. ChatGPT-User is a separate user-triggered agent.
GoogleGooglebot handles ordinary Search crawling. Search Console has a separate site control for AI Overviews, AI Mode and covered generative Search features.Google-Extended is a robots token for future Gemini training and some Gemini grounding; it does not change ordinary Search inclusion or ranking.
AnthropicClaude-SearchBot supports search-result quality; Claude-User retrieves pages for user questions.ClaudeBot is the training crawler. Anthropic says its bots honor robots.txt.
PerplexityPerplexityBot helps surface websites in search results; Perplexity-User fetches pages after user requests.Perplexity says its search bot is not used to crawl for foundation-model training. Its user fetcher generally ignores robots.txt.

These are the providers' published roles as of October 2026, not a universal “AI bot” standard. Perplexity does not present its search crawler as a training toggle. Recheck current documentation before changing a site-wide rule.

Can You Allow Search While Declining Training?

For OpenAI, yes: it documents independent settings for OAI-SearchBot and GPTBot. A site that chooses search access but not future model training might use this example in its root robots.txt:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

This does not promise a link in ChatGPT or erase earlier data use. OpenAI says opting out of OAI-SearchBot keeps a site out of ChatGPT search answers, although a navigational link may still appear. It may take about a day for a changed policy to be reflected.

Google works differently. Google-Extended is a robots token, not a separate HTTP request agent; it does not control ordinary Search inclusion. Search Console's Search generative AI setting separately includes or excludes links and content from covered generative Search features. See our AI Overviews and AI Mode guide for that property's scope.

What Can Robots.txt Do—and What Can It Not Do?

A robots.txt rule tells cooperating crawlers which URLs not to fetch. It is not a password, a blanket license decision or a way to reverse past training. Public URLs may still be known through other pages. Protect private information with authentication.

Place the file at each host's root, such as https://example.com/robots.txt; subdomains need their own policies. Test the intended user-agent group: a broad Disallow: / for the wrong agent can remove search access. Our robots.txt guide covers syntax and tests. An optional llms.txt file does not replace access controls.

User-triggered fetching also differs: OpenAI says ChatGPT-User may not follow robots rules, Perplexity says Perplexity-User generally ignores them, and Anthropic says its bots, including Claude-User, honor them.

A Practical Access Audit for One Site

  1. Write down the goal. Decide whether you want ordinary Google Search visibility, a particular AI search surface, training restrictions, or a combination. Do not start with a list of bot names.
  2. Inspect current controls. Open the live robots.txt for each domain and subdomain. Check Google Search Console's generative AI setting separately. Review any CDN or firewall rules that could reject a permitted crawler.
  3. Change one policy at a time. Keep a copy of the previous file and record which user agent and URL group the rule targets. Test that public pages remain available to the search agents you intend to allow.
  4. Verify the result. Fetch robots.txt, inspect server or CDN logs for permitted agents, and use the provider's official checks where available. A matching user-agent string alone does not authenticate the requester; use published IP ranges where relevant. Allow for propagation time; missing AI citations are not proof of a blocked crawler.

A kitchen-appliance site might want its installation guides discoverable while declining one provider's training crawl. It should confirm the guides are public, then configure that provider's training control without blocking its search agent. This is a policy example, not a client result.

For page-level eligibility after access is settled, see the crawling and indexing guide.

Frequently Asked Questions

If I Block a Training Bot, Will I Disappear From AI Search?

Not necessarily. OpenAI separates GPTBot from OAI-SearchBot, so the two choices are independent. Other products use different controls; check the one you mean.

Does Blocking Google-Extended Remove My Site From Google Search?

No. Google says Google-Extended does not affect ordinary Search inclusion or ranking. Its separate Search Console control covers the listed generative Search features.

Can One Robots.txt Rule Block Every AI Use of My Content?

No. Providers use different agents and some user-triggered requests may not follow robots rules. robots.txt is a crawler instruction, not access protection; use authentication for private content.

Local design preview

This form is a design preview. Subscriptions and external tools are not connected.

Browse the article library