A service business owner has heard that AI "steals content" and wants to block bots from their website, but worries about disappearing from ChatGPT, Gemini, and Perplexity responses. However, robots.txt doesn't hide pages from Google — it controls which bots can visit and for what purpose. Every company has two types of bots: some serve search (showing your business in AI assistant responses), others train models (using your site content to teach AI). Blocking all bots with one line means disappearing from assistant search results.
On this page, you'll learn which bots from which companies visit your site, what happens when you block them, and how to configure robots.txt to preserve search visibility while protecting against unwanted training.

What robots.txt does and its limitations
The robots.txt file is used primarily to manage crawler traffic to your site, not to hide content from Google search results. As Google Search Central documentation explains, robots.txt tells crawlers which URLs they can access and serves mainly to control server load. Even if you block all bots, your site can still appear in Google — if someone links to it from another site, Google may index that link despite the robots.txt prohibition.
robots.txt operates at the server level — it's a guideline for robots, not a guarantee. Most legitimate bots (including those described below) respect these rules, but some may ignore them. That's why robots.txt is a management tool, not full protection.
OpenAI: two independent bots
OpenAI uses two separate robots that site owners control independently in robots.txt. OAI-SearchBot serves to display pages in ChatGPT search results — a company can block it if it doesn't want to appear in assistant responses. GPTBot collects content to train OpenAI's generative models. As the OpenAI documentation states, both settings are independent — you can block GPTBot so content isn't used for training while simultaneously allowing OAI-SearchBot so the site appears in search.
Changes to robots.txt may take approximately 24 hours before OpenAI systems reflect them.
Why allowing OAI-SearchBot matters
OAI-SearchBot is one of the few bots that literally shows your business page to ChatGPT users. When someone asks the assistant about services like "best hair salon in Warsaw", companies allowing this bot can appear in the response with a link. Allowing the bot doesn't guarantee a place in the answer, but blocking it practically rules one out.
Google: a token that doesn't affect the search engine
Google-Extended is a special token in robots.txt that lets you restrict content use for Gemini and Vertex AI training. The key point is that Google-Extended doesn't affect regular Google search — it only applies to Google's AI products. As Google Search Central documents explain, this token is independent from standard Google crawlers and doesn't affect indexing or ranking in search results.
For a service business owner, this means they can block Google-Extended if they don't want their site content used for Gemini model training, while their site remains visible in Google Maps, Search, and other standard Google services.
What is Vertex AI
Vertex AI is Google's platform for building and deploying applications based on generative models. Companies building their own chatbots or assistants on Google infrastructure can use Vertex AI. Google-Extended allows site owners to exclude their content from this process if they wish.
Apple: bots for search and for training
Applebot visits sites for Apple search functions — Spotlight, Siri, and Safari. Applebot-Extended is a separate token that lets site owners exclude their content from Apple generative foundation model training. According to Apple documentation, disallow for Applebot-Extended excludes content from training, but Applebot can still visit the site for search purposes.
This important distinction means blocking Applebot-Extended doesn't affect whether a business appears in Apple search results. The owner can therefore protect content from AI training while maintaining visibility in the Apple ecosystem.
Why Spotlight and Siri matter for business
When an iPhone user asks Siri about "nearest hair salon in Warsaw", Siri can also draw on content from pages Applebot has visited before. If you block Applebot, your site's content won't reach Apple's search — Siri and Spotlight won't be able to use it.
Anthropic: three bots with different roles
Anthropic operates three separate bots, each with a different purpose. ClaudeBot collects content that may be used for model training — blocking this bot excludes future content from training datasets. Claude-User fetches pages in response to user questions. Claude-SearchBot analyses online content to improve search result quality.
As Anthropic Help Center explains, blocking Claude-SearchBot may reduce site visibility in user-directed search results. All three bots respect robots.txt rules, but each has a different function — site owners can therefore precisely control access.
What "collecting for training" means
When a bot collects content for training, it doesn't publish or show it to users. The content becomes part of a massive dataset on which the model learns to recognize patterns, formulate responses, and generalize knowledge. For a business owner, this is a matter of ethics and preference — some don't want their content training competitive models, others have no objection.
Perplexity: a results bot that doesn't train models
PerplexityBot serves to display and link sites in Perplexity search results and, according to Perplexity documentation, is not used to train AI foundation models. The company recommends allowing PerplexityBot so your site appears in results.
Perplexity-User is a bot that fetches pages on user request — unlike PerplexityBot, this bot generally ignores robots.txt rules because it operates on explicit user request. You can't block it via robots.txt if a user asks for a specific page.
Perplexity as a next-generation search engine
Perplexity works differently from traditional search engines — instead of a list of links, it shows concise answers with citations from sources. For a service business, being present in Perplexity means a potential customer can see the company name and service description without clicking on a link. This is a new channel worth considering when building visibility.

Decision table: which bot does what
The following table shows which bots serve which companies, their purpose, and what happens when blocked.
| Bot | Company | Purpose | Effect of blocking | Recommendation for service business |
|---|---|---|---|---|
| OAI-SearchBot | OpenAI | Display in ChatGPT | Site disappears from ChatGPT results | Allow — maintains visibility |
| GPTBot | OpenAI | Model training | Content not used for training | Owner's decision |
| Google-Extended | Gemini/Vertex AI training | Content not used for AI training | Owner's decision | |
| Applebot | Apple | Search (Spotlight, Siri, Safari) | Site disappears from Apple search | Allow — maintains visibility |
| Applebot-Extended | Apple | Foundation AI training | Content not used for training | Owner's decision |
| ClaudeBot | Anthropic | Model training | Content not used for training | Owner's decision |
| Claude-SearchBot | Anthropic | Search quality | May reduce search visibility | Allow — maintains visibility |
| PerplexityBot | Perplexity | Search results | Site disappears from Perplexity | Allow — maintains visibility |
How to check and modify robots.txt yourself
Open your browser and navigate to yourdomain.com/robots.txt (replace with your actual domain). Look for the bot names listed above. If the file doesn't exist or contains no rules, all bots can visit your site by default.
Example of correct configuration that allows search bots and blocks training bots:
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot
Allow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: *
Disallow:This configuration allows bots responsible for search visibility (OAI-SearchBot, Applebot, Claude-SearchBot, PerplexityBot) and blocks training bots (GPTBot, Google-Extended, Applebot-Extended, ClaudeBot). The order of groups in the file doesn't matter — a bot follows the group that most specifically matches its name (bot names are case-insensitive), while paths after Allow and Disallow are case-sensitive.
After making changes, wait approximately 24 hours before bot systems recognise them. You can also verify file correctness using testing tools available in Google Search Console.
Visibility in AI under system control
When visibility in AI assistants becomes part of a business strategy, it's worth having a process that makes bot management not a one-time action but continuous monitoring. A system can regularly check the robots.txt file for presence of search and training bots, regularly repeat the query protocol in major assistants, and transform results into tasks to complete.
This approach lets a service business owner maintain control over who and why visits their site, without manually tracking changes in policies of individual AI companies. The system doesn't promise placement in assistant responses — it provides a tool for conscious visibility management.
See how this works in practice — check the visibility in AI assistants service, which handles visibility in ChatGPT and other assistants, go to SEO analysis to see the full picture of business visibility online, order a custom website optimized for bot readability, or add AI reports for visibility monitoring.
Read more about what data about your business AI assistants read, why customers don't leave inquiries, how AI agents differ from chatbots, and how business data looks in Google — this will help you understand how AI visibility affects customer acquisition.
Frequently asked questions
Will blocking all AI bots remove my site from Google?
Not entirely. robots.txt doesn't hide a page from Google in the sense that if other sites link to yours, Google may index it despite the prohibition. Blocking in robots.txt concerns bot visits, not indexing based on incoming links. If you genuinely want to hide a page, you need other methods such as meta tags or passwords.
Do I need to allow all search bots?
You don't have to, but if you block search bots (OAI-SearchBot, PerplexityBot, Applebot, Claude-SearchBot), your business won't appear in responses from the respective assistants. For a service business, presence in ChatGPT, Perplexity, or Spotlight can be a source of new customer inquiries — consider allowing these bots.
Can I block training bots but keep search bots?
Yes. Most AI companies offer separate bots for search and training. OpenAI has OAI-SearchBot (search) and GPTBot (training), Anthropic has Claude-SearchBot and ClaudeBot. You can allow the first and block the second in the same robots.txt configuration. The table above has the exact mapping.
Does Perplexity-User respect robots.txt?
No. Perplexity-User operates on explicit user request and generally ignores robots.txt rules. You can't block it at the file level — if a user asks for a specific page, Perplexity will fetch it. This distinguishes it from PerplexityBot, which respects robots.txt.
How quickly do robots.txt changes take effect?
Most bot systems need around 24 hours to recognise changes in robots.txt. OpenAI and Perplexity state about 24 hours in their documentation. In practice, it may take a bit longer depending on how often bots visit your site.