Reference

robots.txt Reference

robots.txt is a plain text file placed at the root of a domain (https://example.com/robots.txt) that tells compliant crawlers which URLs they may and may not fetch. It is part of the Robots Exclusion Protocol, a convention rather than an enforced standard: well-behaved crawlers follow it, malicious ones do not.

For the SEO implications of crawl control, see Crawlability and robots.txt.


File location and format

The file must be at the root of the domain. It applies only to the domain it is hosted on: a robots.txt at example.com does not apply to subdomain.example.com, and vice versa.

Requirements:

  • Plain text, UTF-8 encoded
  • One directive per line
  • Lines beginning with # are comments and are ignored by crawlers
  • Blank lines separate rule groups

Directives

User-agent

Specifies which crawler the following rules apply to. Must appear at the start of each rule group, before any Disallow or Allow directives.

User-agent: Googlebot
User-agent: *
  • * matches all crawlers not covered by a specific rule group
  • Multiple User-agent lines can precede a shared set of rules
  • A crawler obeys exactly one group: the one whose User-agent line most specifically matches its name. Every other group is ignored
  • A named group does not combine with the * group. If a crawler has its own group, the * group is never consulted for it. See Group selection below

Disallow

Blocks the specified path. The crawler will not fetch any URL beginning with this path.

Disallow: /admin/
Disallow: /search
Disallow: /
  • An empty Disallow: value means “allow everything” (equivalent to no restriction)
  • Disallow: / blocks the entire site
  • Paths are case-sensitive: Disallow: /Admin/ and Disallow: /admin/ are different rules
  • The rule matches the beginning of the URL path: Disallow: /blog matches /blog, /blog/, and /blog-post

Allow

Explicitly permits a path within a disallowed directory. Used to create exceptions.

Disallow: /private/
Allow: /private/public-page/
  • Only meaningful within a Disallow context; Allow: / with no Disallow does nothing
  • When both Allow and Disallow rules match a URL, the more specific (longer) rule wins
  • If rules are the same length, Allow wins

Sitemap

Declares the location of an XML sitemap. Not a crawl rule: it is a hint to crawlers about where to find URLs.

Sitemap: https://example.com/sitemap.xml
Sitemap: https://example.com/sitemap-news.xml
  • Use absolute URLs, not relative paths
  • Multiple Sitemap directives are allowed
  • Can appear anywhere in the file, but convention places them at the end
  • Googlebot also discovers sitemaps submitted via Search Console; this directive covers other crawlers

Crawl-delay

Specifies a delay (in seconds) between successive requests from the crawler.

Crawl-delay: 2
  • Googlebot does not honour Crawl-delay. Googlebot now manages its own crawl rate automatically, backing off when the server is slow or returns errors. The manual crawl rate setting in Search Console was removed in January 2024; use Google’s Googlebot report form to flag unusual crawl activity.
  • Bing, Yandex, and many other crawlers do honour it
  • Useful on low-resource servers to prevent crawlers from overwhelming the host

Wildcards

* (asterisk)

Matches any sequence of characters, including none.

Disallow: /*.pdf$
Disallow: /search?*
Disallow: /*/print/
  • Supported by Googlebot; behaviour varies across other crawlers
  • Disallow: /search?* blocks all URLs containing /search? followed by anything
  • Disallow: /*/print/ blocks any path with /print/ as a segment

$ (end of string)

Anchors the pattern to the end of the URL.

Disallow: /*.pdf$
  • Matches only URLs ending with the specified pattern
  • Disallow: /*.pdf$ blocks /document.pdf but not /pdf-guide/
  • Supported by Googlebot; behaviour varies across other crawlers
  • Without $, the pattern matches anywhere in the URL

Group selection

Group selection and rule precedence are two different things, and conflating them causes the single most dangerous robots.txt mistake.

Group selection happens first. A crawler scans the file, finds the group whose User-agent line most specifically matches its own name, and obeys only that group. Every other group is ignored, including User-agent: *. Google’s spec states it plainly: “User agent specific groups and global groups (*) are not combined.”1

The * group is not a base layer that named groups build on. It is a fallback for crawlers with no group of their own. So a named group does not “win conflicts” against the * group, it replaces it wholesale. Non-conflicting Disallow rules in the * group are never even read for a crawler that has its own group.

Example:

User-agent: *
Disallow: /private/
Disallow: /cart/

User-agent: GPTBot
Disallow: /no-ai/

GPTBot matches its own group and follows only Disallow: /no-ai/. It never consults the * group, so /private/ and /cart/ are open to it. This fails open: the URLs you meant to block are silently exposed, with no error to warn you.

To apply shared rules to a named crawler, repeat them inside that crawler’s group:

User-agent: *
Disallow: /private/
Disallow: /cart/

User-agent: GPTBot
Disallow: /private/
Disallow: /cart/
Disallow: /no-ai/

This scales badly across many named bots, so generate the file from a template rather than maintaining each group by hand.

User agent specific groups and global groups (*) are not combined

The wildcard group (*) is a fallback, not a global rule. The moment a crawler has its own named group, the * group stops applying to it completely. If you want a shared rule to bind a named bot, it has to be written inside that bot's group. Nothing is inherited.


Rule precedence

Rule precedence applies within the single group a crawler has already selected. When several lines in that group match a URL, crawlers resolve them in this order:

  1. The most specific (longest) matching rule wins
  2. If two rules are equal length, Allow wins over Disallow

Example:

User-agent: *
Disallow: /private/
Allow: /private/public/

A request to /private/public/page matches both rules. /private/public/ (length 16) is longer than /private/ (length 9), so Allow wins and the URL is accessible.


Common user agent strings

Google

User agentCrawls
GooglebotWeb search (applies to all Googlebot variants unless overridden)
Googlebot-ImageGoogle Images
Googlebot-VideoGoogle Video
Googlebot-NewsGoogle News
Google-ExtendedGoogle AI (Gemini training and products)
AdsBot-GoogleGoogle Ads landing page quality
Mediapartners-GoogleAdSense

Other search engines

User agentCrawls
BingbotBing web search
SlurpYahoo (uses Bing index)
DuckDuckBotDuckDuckGo
BaiduspiderBaidu
YandexBotYandex

AI crawlers

User agentPurpose
GPTBotOpenAI training crawler
OAI-SearchBotOpenAI ChatGPT Search indexing
ChatGPT-UserOpenAI live retrieval at query time
ClaudeBotAnthropic training crawler
anthropic-aiAnthropic (alternative user agent)
PerplexityBotPerplexity indexing crawler
Perplexity-UserPerplexity live retrieval
cohere-aiCohere training crawler
CCBotCommon Crawl (used by many model providers)
meta-externalagentMeta AI training crawler

Common configurations

Block all crawlers

User-agent: *
Disallow: /

Use during development. Remove before launch: this is the most common cause of sites not appearing in search after deployment.

Block all AI training crawlers

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: anthropic-ai
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: meta-externalagent
Disallow: /

Note: blocking AI training crawlers does not prevent AI search products (ChatGPT Search, Perplexity) from citing your content if they retrieve it via their search crawlers (OAI-SearchBot, Perplexity-User). Training and retrieval use different bots.

Your CDN may be answering this question for you

robots.txt is a request, and it is no longer the only layer deciding whether an AI crawler reaches the site. Bot management at the CDN can block a crawler that robots.txt allows, and the block happens before the request ever reaches your server, so nothing in your robots.txt file will show it.

Cloudflare now groups AI crawlers into three categories rather than by vendor: training (crawlers taking content to train or fine-tune a model), search (crawlers indexing content to answer questions about it later) and agent (automated activity acting in real time on a person’s behalf).2 From 15 September 2026, for domains newly onboarded to Cloudflare, Training and Agent are blocked by default on pages that display ads, while Search remains allowed.2

Two things worth being precise about, because they are easy to conflate. The default change is narrow: it applies to newly onboarded domains, and only on pages carrying ads. A new editorial or business site with no ad units is not affected by it, and Search crawlers stay allowed either way, so AI citation visibility is not touched by the default.

The wider risk is one you opt into. Multi-purpose crawlers are judged on all of their behaviours, so if you choose to block the Training category, that also blocks Googlebot, Applebot and Bingbot.2 A setting picked to keep content out of model training can therefore take the site out of Google. That is not the default; it is a decision, and it is the one to get right. If the site is behind Cloudflare, read the security settings rather than assuming robots.txt is the whole picture.

Block internal search results

User-agent: *
Disallow: /search
Disallow: /search/

Block admin and account areas

User-agent: *
Disallow: /admin/
Disallow: /account/
Disallow: /login
Disallow: /checkout/

Block parameterised URLs

User-agent: *
Disallow: /*?

Blocks all parameterised URLs. Use with caution: this can inadvertently block legitimate pages if any use query strings.

Block print and feed versions

User-agent: *
Disallow: /*/print/
Disallow: /*/feed/
Disallow: /*/amp/

Allow Googlebot, block all others

User-agent: *
Disallow: /

User-agent: Googlebot
Disallow:

An empty Disallow after a specific user agent allows full access for that crawler. Useful for staging environments where you want Google to crawl but no other bots.


What robots.txt does not do

It does not prevent indexing. A URL blocked in robots.txt can still appear in search results if it has inbound links. Google will index the URL without crawling it, showing the page title from anchor text and no description. To prevent indexing, use a noindex meta tag, but the page must be crawlable for Google to read it.

It does not secure content. robots.txt is public and explicitly advertises what you are hiding. Do not rely on it for sensitive content; use authentication.

It does not apply to subdomains. example.com/robots.txt has no authority over blog.example.com. Each subdomain needs its own file.

It does not apply to non-HTTP protocols. FTP, email, and other protocols are not covered.

It does not stop malicious bots. Only well-behaved crawlers follow robots.txt. Scrapers, spam bots, and vulnerability scanners ignore it.

Footnotes

  1. robots.txt specification — Google Search Central

  2. Your site, your rules: new AI traffic options for all customers — Cloudflare 2 3