AI & Websites 10 min read

Should you block AI crawlers? GPTBot, Google-Extended and robots.txt explained

Training bots, AI search bots and user fetchers: what blocking each one does, and how to decide.

On this page 9 sections

If your website is public, bots such as GPTBot, ClaudeBot and PerplexityBot are probably already reading it. You can block AI crawlers like these with two lines of robots.txt each. The hard part is choosing which ones, because “AI crawler” now covers three different jobs, and blocking each has a very different effect.

Block a training crawler and you ask that company not to train on your content, while you keep appearing in AI answers. Block a search crawler and you drop out of that assistant’s answers. Block a user-triggered fetcher and an assistant can’t open your page even when a person, perhaps a buyer, asks it to.

Last reviewed: 27 September 2026. Bot names and behaviour were checked against each vendor’s documentation; they change often, so the vendor’s page is the final word.

The short answer

  • Training crawlers (GPTBot, ClaudeBot, CCBot, and the Google-Extended and Applebot-Extended tokens): blocking them asks those companies not to train AI models on your content from now on. You stay in Google Search, AI Overviews and AI search answers, though Google-Extended also covers grounding answers in the Gemini app.
  • Search and retrieval crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot): blocking them removes your pages from that assistant’s search answers.
  • User-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User): blocking them stops an assistant opening your page when someone asks, though several vendors say these fetchers may not follow robots.txt.
  • If your website exists to win enquiries, allow search crawlers and user fetchers; training is a judgement call, and mostly not a visibility decision. If your content is the product, block training and decide on search crawlers from your referral data.

Three kinds of AI bot, and what blocking each one does

The names matter less than the job each bot does, and new names keep appearing as AI changes how websites are found and used.

Training crawlers

OpenAI says disallowing GPTBot signals that your content shouldn’t be used to train its foundation models, and Anthropic says the same of ClaudeBot. Meta says Meta-ExternalAgent crawls both to train foundation AI models and to index content for its products. CCBot builds Common Crawl’s open archive of the web, which many AI developers have used as training data, so blocking it reaches further than any single company.

Blocking only works from now on: it doesn’t remove anything already collected. And Google-Extended and Applebot-Extended aren’t crawlers. Google’s usual crawlers fetch the pages, and the token decides whether that content may train future Gemini models or ground answers in Gemini Apps and Vertex AI’s Grounding with Google Search. Google states it doesn’t affect inclusion or ranking in Search. Apple says pages that disallow Applebot-Extended can still appear in its search results. Neither name ever shows up in your logs.

Search and retrieval crawlers

These build the indexes AI assistants search when they answer. OpenAI states that sites opted out of OAI-SearchBot won’t be shown in ChatGPT search answers, though they can still appear as navigational links. Anthropic warns that disabling Claude-SearchBot may reduce your visibility, and Perplexity says PerplexityBot surfaces and links sites in its results and isn’t used to train foundation models.

Google works differently. AI Overviews and AI Mode are part of Google Search, built from what Googlebot crawls, so there is no separate robots.txt switch for them. The controls are the ones that govern ordinary snippets: nosnippet, data-nosnippet, max-snippet and noindex. Google AI Overviews and your website explains what each costs you. Microsoft’s Copilot draws on Bing search, so blocking Bingbot to avoid AI answers takes you out of Bing too.

User-triggered fetchers and browsing agents

When someone asks an assistant to compare two suppliers, it fetches pages on their behalf. Perplexity says its user fetcher generally ignores robots.txt, OpenAI says robots.txt rules may not apply to ChatGPT-User, and Meta says Meta-ExternalFetcher may bypass them, all on the reasoning that a person asked. Anthropic says Claude-User follows robots.txt.

Browsing agents go further: they drive a real browser, click through pages and fill in forms. Some look like an ordinary visitor’s browser, so robots.txt can’t single them out. Treat them as visitors and make the site easy to use; AI browsing agents and your website covers what helps them succeed.

Copy-ready robots.txt rules for AI bots

One rule matters most here: a crawler follows only the most specific group that names it, and ignores the rest. Add a group for GPTBot and it stops following everything under User-agent: *, so repeat any path rules you want it to keep. Several bots can share a group, one User-agent line each.

Allow everything

You don’t need to name AI bots to allow them. A standard WordPress file already does:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://www.example.com/sitemap.xml
# Opt out of AI model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
Disallow: /

# Everyone else, including search engines and AI search crawlers
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://www.example.com/sitemap.xml

OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot aren’t named, so they follow the * group and keep crawling. If answers in the Gemini app matter to you, consider leaving Google-Extended out, since it covers grounding there as well as training.

Keep one section out of training

If only part of the site has licensing value, such as paid reports, limit the rule to that path and repeat your normal rules:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
Disallow: /reports/
Disallow: /wp-admin/

Block every AI bot you can name

To leave AI answers as well as training, add these lines to the training group above, before its Disallow: /:

User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User

Expect to lose citations and referral visits from those assistants.

Mistakes to avoid

  • Blocking Googlebot or Bingbot to stop AI. You leave those search engines entirely.
  • Hiding private pages with robots.txt. The file is public and advertises every path you list. Use a login.
  • Expecting instant results. Google generally caches robots.txt for up to 24 hours; OpenAI, Meta and Perplexity mention similar delays.
  • Forgetting subdomains. Each host needs its own file.

On WordPress, most SEO plugins include a robots.txt editor, and a physical file in the site root replaces WordPress’s virtual one. The technical SEO checklist covers the wider crawl checks.

robots.txt is a request, not a lock

robots.txt is voluntary. The major AI companies document that their crawlers follow it, but nothing enforces it, and any script can claim to be GPTBot. Its legal weight is also unsettled: AI training and copyright are still being tested in court, with different rules by country, so if your content has licensing value, take legal advice. Real enforcement happens at the network edge: firewall rules, rate limits and CDN bot management, and logins for anything private.

Your CDN or firewall may already be deciding for you

In July 2025, Cloudflare made blocking known AI crawlers the default for new domains, asking each new customer whether to allow them, and it offers a managed robots.txt that adds AI-training rules to yours. Other CDNs, hosting firewalls and security plugins have their own bot settings, and a single “block AI bots” switch can catch search crawlers and user fetchers along with training bots. The result can be a robots.txt that says “welcome” and a firewall that turns AI search crawlers away.

Check it yourself. In each of those dashboards, look for settings with “AI”, “bots” or “crawlers” in the name. Note which kinds of bot each blocks, not just whether it is on, then make every layer agree.

How to check which AI bots actually visit

Your analytics won’t show them: most crawlers don’t run your analytics script, and Google Analytics excludes known bots anyway. The evidence is in your server access logs, which most hosting panels let you download, or in your CDN’s bot dashboard. Search a downloaded log for the names:

grep -Ei "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|CCBot|Meta-External|Applebot|bingbot|Googlebot" access.log

Then read what you find:

  • Status codes. A 200 means the bot received the page. A 403 or a challenge page means your CDN or firewall stopped it.
  • Which pages. Are search crawlers reaching your service and pricing pages, or only old blog posts?
  • Whether it’s genuine. User agents are easy to fake. Google, OpenAI, Anthropic and Perplexity publish the IP ranges their bots use, so check a sample.

Then check the other side: visits referred by assistants such as chatgpt.com or perplexity.ai, in Google Analytics’ traffic acquisition report. That is what blocking search crawlers would cost. Generative engine optimisation explains how to track and grow it.

When to block AI crawlers, and when not to

Your website Training crawlers AI search crawlers User fetchers
Wins enquiries or sales Your choice Allow Allow
Content is the product Block, unless licensed Decide from referral data Usually allow
A mix of both Block paid sections only Allow Allow
Private or members-only Use a login Use a login Use a login

If your website exists to win enquiries

Service businesses, suppliers, clinics and schools publish content to be read, and some buyers now ask an assistant before they search. Blocking search crawlers or user fetchers makes you harder to recommend, and harder to check just when a buyer is comparing options.

Training is a genuine choice. Allowing it lets models learn what you do, which can matter when an assistant answers from memory rather than a live search, though nobody outside the AI companies can measure how much. Blocking it protects little when the pages exist to be read. With no strong view, allow it; if you object, block the training crawlers alone and your visibility in ChatGPT, Claude and Perplexity search stays intact.

If your content is the product

When people pay for your articles, research, courses or data, training use can compete with you directly, so blocking training crawlers is the sensible default. Search crawlers are harder to call: they send visits, but an answer that summarises your work can also replace the visit. Measure what each assistant sends you, and revisit as the numbers change.

Questions to settle first

What about llms.txt?

llms.txt is a proposal, first published by Jeremy Howard in September 2024, for a Markdown file at /llms.txt that points AI tools to a site’s most useful content in compact form.

  • It controls nothing. It can’t block or allow any bot; that is still robots.txt’s job.
  • It isn’t an established standard. At the time of writing, no major AI search provider has said it uses llms.txt to decide what to cite, and Google says no AI text files or special markup are needed to appear in AI Overviews or AI Mode.
  • Its clearest use is developer documentation, read by coding assistants.

For a business website it is optional: harmless, but unlikely to change much. Crawlable pages with clear headings and plain-text answers do far more.

Frequently asked questions

Will blocking GPTBot remove my website from ChatGPT?

No. GPTBot collects training data; ChatGPT’s search answers rely on OAI-SearchBot. You can block the first and still be cited through the second.

Does blocking Google-Extended keep my site out of AI Overviews?

No. AI Overviews use what Googlebot crawls for Search. Google-Extended covers Gemini model training and grounding in certain Gemini products, and has no effect on Search inclusion or ranking.

Does blocking AI crawlers hurt my Google rankings?

Not in itself: other companies’ bots play no part in Google’s ranking, and Google-Extended isn’t a ranking signal. The real risk is collateral damage, such as a firewall rule that also catches Googlebot.

Can AI tools still use my content if I block their crawlers?

Yes. Content collected earlier may remain in datasets and models, people can paste your text into a chat, and some user fetchers may ignore robots.txt. Blocking reduces future collection; it doesn’t make public pages private.

Make the setting a decision, not an accident

Many websites have no AI crawler policy, just whatever robots.txt, the CDN and a security plugin happen to say. Answer the questions above, make every layer agree, and check your logs every few months, because vendors add and rename bots.

Settings like these drift as plugins update. Our website care plans look after a WordPress site month to month: updates, regular backups, security and uptime monitoring, and a stated response time when something needs changing. For a second opinion first, request a free website audit.

Written by the PORVIX team

The people who design, build and maintain websites for growing businesses. We write about the questions that come up on real projects, in plain language, and update articles when the advice changes.

Published Updated

How we work

Is your current site costing you enquiries?

A free, plain-language audit, by email within 2 business days.

Get a free audit
Get a free audit

Keep reading

More plain-language guides.

More on AI and websites first, then other guides worth reading next.

All insights

Start here

Let’s build a website that brings in business.

Tell us about your project. You’ll hear back from a real person within one business day, with honest advice either way.

  • Free consultation
  • Fixed written quote
  • Your details stay private