robots.txt Guide: Complete Syntax, Examples & AI Rules (2026)

Master robots.txt with syntax examples and AI crawler rules. Control GPTBot, ClaudeBot, and Googlebot to protect crawl budget and optimize GEO search readiness.

BugViso

16 min read

A developer makes a single typographical error during a production deployment, adding a single forward slash to their server configuration: Disallow: /. Overnight, Googlebot, Bingbot, and every major web crawler are instructed to halt all crawling across the entire domain. Within days, organic search traffic plummets to zero.

In 2026, the stakes of managing your robots.txt file are higher than ever. Beyond directing search engine bots like Googlebot, your robots.txt file now dictates how artificial intelligence crawlers (such as OpenAI's GPTBot, Anthropic's ClaudeBot, and PerplexityBot) interact with your content. A single misconfigured directive can either leak proprietary documentation to AI training models or inadvertently block your brand from appearing in AI-synthesized search answers.

In this comprehensive robots.txt guide, you will master the IETF RFC 9309 standard: explore core syntax and pattern-matching rules, implement copy-paste production templates, govern AI search and LLM crawlers for Generative Engine Optimization (GEO), and automate robots directive validation.


What Is robots.txt? RFC 9309 and the Robots Exclusion Protocol

The robots.txt file is a plain text document placed in the root directory of a web server that communicates crawling restrictions to automated web robots. Formally standardized in 2022 under IETF RFC 9309, the Robots Exclusion Protocol establishes universal rules for crawler discovery, syntax parsing, and error handling.

Diagram
+-------------------------------------------------------------------------+

|                  THE ROBOTS.TXT CRAWL PROTOCOL PIPELINE                 |
|                                                                         |
|  1. Crawler arrives at domain: `https://example.com/products/item-1`    |
|                                                                         |
|  2. BEFORE requesting the page, crawler fetches:                        |
|     `https://example.com/robots.txt`                                    |
|                                                                         |
|  3. PARSER CHECKS RULES:                                                |
|     +---> Allowed: Proceed to download HTML and render assets.          |
|     +---> Disallowed: ABORT REQUEST! (Zero bytes transferred).          |

+-------------------------------------------------------------------------+

According to Google's introduction to robots.txt, robots.txt is primarily a mechanism for managing crawl budget and server bandwidth. It prevents crawlers from overloading your backend infrastructure with low-value, duplicate, or administrative URLs.

Technical File Constraints (RFC 9309):

  1. Strict File Location: The file must reside at the root domain path (https://example.com/robots.txt). A file located at https://example.com/sub/robots.txt is completely ignored.
  2. File Size Limit: Search engines enforce a maximum file size of 500 Kilobytes (KB). Any directives beyond 500KB are truncated.
  3. Character Encoding: The file must be authored in standard UTF-8 text format.
  4. Protocol Scope: Each subdomain and protocol scheme requires its own distinct file (e.g., https://example.com/robots.txt does not govern https://blog.example.com/robots.txt).

Core robots.txt Directives and Syntax Rules

A valid robots.txt file consists of directive blocks composed of specific field-value pairs.

Diagram
+-------------------------------------------------------------------------+

|                  CORE DIRECTIVES & THEIR PROTOCOL ROLES                 |

+---------------+---------------------------------------------------------+

| Directive     | Purpose & Syntax                                        |

+---------------+---------------------------------------------------------+

| `User-agent:` | Identifies which crawler the following block applies to |
| `Disallow:`   | Specifies URL paths the crawler is forbidden to request |
| `Allow:`      | Carves out exceptions within a disallowed directory     |
| `Sitemap:`    | Declares absolute URL of your XML sitemap index         |

+---------------+---------------------------------------------------------+

1. The User-agent: Directive

Defines the target bot. A wildcard asterisk (*) applies the block to all web crawlers unless a more specific bot block is declared:

text
# Global rule block applying to all web crawlers
User-agent: *
Disallow: /admin/

# Specific rule block applying exclusively to Googlebot
User-agent: Googlebot
Disallow: /internal-search/

2. Path Matching, Wildcards (*), and Anchors ($)

Modern crawlers support pattern-matching extensions defined in RFC 9309:

  • Prefix Matching: Disallow: /private blocks /private, /private/, /private-doc.html, and /privatization.
  • Wildcard (*): Matches zero or more arbitrary characters.
  • End-of-URL Anchor ($): Matches the exact end of a URL path.
text
# Block all URLs containing the query parameter 'sort='
Disallow: /*?*sort=

# Block all standalone PDF documents across the entire site
Disallow: /*.pdf$

3. The Path Precedence Rule: Longest Match Wins

When both Allow: and Disallow: directives match a requested URL, search engines follow the directive with the longest matching character pattern:

text
User-agent: *
Disallow: /products/
Allow: /products/featured$

# EVALUATION:
# /products/shoes         -> DISALLOWED (Matches 10 chars: `/products/`)
# /products/featured      -> ALLOWED (Matches 19 chars: `/products/featured$`)

4. The Obsolete crawl-delay Directive

The crawl-delay: directive was historically used to slow down crawler request rates. Googlebot completely ignores crawl-delay. Google automatically calculates crawl rates based on your server's response time (TTFB) and error rates. If you need to manage Googlebot throughput, optimize your backend infrastructure rather than relying on crawl-delay.


Copy-Paste robots.txt Production Templates

Here are three battle-tested configuration templates for different web application architectures:

Template 1: Modern SaaS / B2B Web Application

text
# Standard Production Robots.txt for Web Applications
User-agent: *
Allow: /

# Disallow internal administrative, API, and authentication endpoints
Disallow: /admin/
Disallow: /api/
Disallow: /auth/
Disallow: /dashboard/
Disallow: /settings/
Disallow: /staging/

# Disallow internal search query parameters
Disallow: /search?*
Disallow: /*?*ref=

# Declare primary XML Sitemap
Sitemap: https://example.com/sitemap_index.xml

Template 2: E-Commerce Store with Parameter Filtering

text
# E-Commerce Optimization: Protect Crawl Budget from Faceted Filters
User-agent: *
Allow: /

# Prevent crawl loops from sorting and filter combinations
Disallow: /*?*sort=
Disallow: /*?*filter=
Disallow: /*?*price=
Disallow: /*?*color=
Disallow: /*?*size=
Disallow: /*?*sessionid=

# Disallow checkout, cart, and account portals
Disallow: /checkout/
Disallow: /cart/
Disallow: /account/
Disallow: /wishlist/

# Declare Sitemaps
Sitemap: https://example.com/sitemaps/sitemap-products.xml
Sitemap: https://example.com/sitemaps/sitemap-categories.xml

Template 3: Staging / Pre-Production Environment (Block All)

text
# Staging Environment: Block All Search Bots & AI Crawlers
User-agent: *
Disallow: /

Note: In staging environments, always combine Disallow: / with HTTP Basic Authentication or IP whitelisting to prevent sensitive endpoints from leaking.


Modern AI Search & LLM Crawler Directives (The 2026 AI Landscape)

The rapid rise of AI search engines and Large Language Models has split automated web crawlers into two distinct functional categories:

Diagram
+-------------------------------------------------------------------------+

|                  AI TRAINING CRAWLERS VS AI SEARCH AGENTS               |

+--------------------------+----------------------------------------------+

| 1. LLM TRAINING BOTS     | 2. LIVE AI SEARCH / GEO CITATION AGENTS      |

+--------------------------+----------------------------------------------+

| Examples: `GPTBot`,      | Examples: `ChatGPT-User`, `PerplexityBot`,   |
| `ClaudeBot`,             | `Google-Extended`                            |
| `Bytespider`             |                                              |
| Purpose: Scrapes data to | Purpose: Fetches real-time web context to    |
| train future AI models   | answer user queries with linked citations    |
| Commercial Impact: Consumes| Commercial Impact: Drives direct referral  |
| bandwidth; no clicks     | traffic and brand visibility in AI answers   |

+--------------------------+----------------------------------------------+

The Major AI User-Agents Matrix

According to the official OpenAI GPTBot documentation and Anthropic ClaudeBot documentation, you can selectively permit or restrict AI agents based on your intellectual property and marketing strategy:

Diagram
+-------------------------------------------------------------------------+

| User-Agent String | Operator      | Primary Purpose                     |

+-------------------+---------------+-------------------------------------+

| `GPTBot`          | OpenAI        | Model Training (GPT-5, future LLMs) |
| `ChatGPT-User`    | OpenAI        | Live Web Browsing in ChatGPT Search |
| `ClaudeBot`       | Anthropic     | Training & Live Context Retrieval   |
| `PerplexityBot`   | Perplexity AI | Real-Time AI Search Indexing        |
| `Google-Extended` | Google        | Gemini & Vertex AI Model Training   |
| `Bytespider`      | ByteDance     | TikTok / Doubao LLM Training        |
| `CCBot`           | Common Crawl  | Open-Source Web Scraping Archive    |

+-------------------+---------------+-------------------------------------+

The Strategic Generative Engine Optimization (GEO) Configuration

If your objective is to prevent bulk scraping of your proprietary data while maximizing your brand's visibility in AI search summaries (Perplexity, ChatGPT Search), configure your robots.txt to distinguish between training crawlers and live search agents:

text
# 1. Allow search engine crawlers and live AI retrieval agents
User-agent: Googlebot
Allow: /

User-agent: Bingbot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: ChatGPT-User
Allow: /

# 2. Block pure AI model training scrapers (Protects IP & saves bandwidth)
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: CCBot
Disallow: /

# 3. Global Default Rules
User-agent: *
Allow: /
Disallow: /admin/

Sitemap: https://example.com/sitemap_index.xml

Testing, Validating, and Debugging robots.txt

Before deploying changes to production, you must validate that your syntax complies with RFC 9309 and does not inadvertently block critical rendering assets.

Diagram
+-------------------------------------------------------------------------+

|                  ROBOTS.TXT TESTING WORKFLOW                            |
|                                                                         |
|  Step 1: Check Live File Delivery via cURL                              |
|          `curl -I https://example.com/robots.txt`                       |
|          -> Verify HTTP 200 OK & `Content-Type: text/plain`             |
|                                                                         |
|  Step 2: Inspect via Google Search Console Robots.txt Report            |
|          -> Check for syntax warnings and historical fetch timestamps   |
|                                                                         |
|  Step 3: Test Specific URL Paths via URL Inspection Tool                |
|          -> Verify Googlebot can fetch critical CSS, JS, and HTML       |

+-------------------------------------------------------------------------+

To diagnose why pages may fail to appear in search results despite valid syntax, consult our comprehensive guide on why is my page not indexed? technical audit guide.


How BugViso Audits Robots.txt and AI Search Readiness Automatically

Manually auditing robots.txt syntax, monitoring AI crawler policies, and verifying that critical stylesheet assets are not blocked is a complex maintenance task.

Diagram
+-------------------------------------------------------------------------+

|               BUGVISO ROBOTS & AI READINESS AUDIT PIPELINE              |
|                                                                         |
|  [Target Domain Crawled via Headless Chromium]                          |
|            |                                                            |
|            v                                                            |
|  [Multi-Engine Robots & AI Validation Pipeline]                         |
|            |                                                            |
|            +---> 1. AI Search Readiness (GEO) Engine (`ai_readiness.py`)|
|            |        (Audits GPTBot, ClaudeBot, PerplexityBot policies)  |
|            |        (Validates `/llms.txt` presence & structure)        |
|            |        (Computes 0-100 GEO Citability & Extractability)    |
|            |                                                            |
|            +---> 2. Robots Syntax & Resource Inspector (`seo_intel.py`) |
|            |        (Detects accidental site-wide blocks: `Disallow: /`)|
|            |        (Flags blocked CSS/JS bundles required for render)  |
|            |        (Validates XML sitemap URL declarations)            |
|            |                                                            |
|            v                                                            |
|  [Prioritized Remediation Playbook + Branded PDF Executive Report]      |

+-------------------------------------------------------------------------+

When you run an automated website scan with BugViso, the platform evaluates your technical directives across traditional SEO and modern AI search channels:

  1. AI Crawler & GEO Readiness Audit: BugViso evaluates your robots.txt file against major AI user-agents (GPTBot, ClaudeBot, PerplexityBot, Google-Extended), checks for the presence of an /llms.txt file, and calculates an overall 0–100 GEO Citability Score measuring how easily AI engines can cite your brand.
  2. Syntax & Resource Block Detection: The engine parses your directives against RFC 9309 standards, alerting you if critical CSS stylesheets or JavaScript bundles required for client-side hydration are inadvertently blocked.
  3. Sitemap Declaration Verification: BugViso confirms that your primary sitemap.xml or sitemap index is correctly declared and accessible to search bots. To maintain sitemap integrity, explore our guide on XML sitemaps: how to create, submit & maintain them.
  4. Prioritized Developer Remediation Playbook: All detected directive anomalies are organized into prioritized action items with exact lines of code and corrected syntax delivered in both the interactive dashboard and downloadable PDF report.

More detail is on the GEO audit for AI search feature page.


Common robots.txt Mistakes That Destroy SEO

Avoid these frequent engineering traps when managing robots.txt:

Common MistakeConsequence
Blocking CSS and JS BundlesBreaks mobile rendering in Googlebot
Using File for SecurityPublicly reveals secret admin paths
Disallowing noindex PagesTraps bare URL in Google Search index
Missing Leading SlashesBreaks path matching across subpaths

1. Blocking CSS and JavaScript Assets

Years ago, blocking /assets/ or /js/ was common to conserve bandwidth. Today, Googlebot operates as a full rendering engine. If Googlebot cannot download your CSS and JavaScript files, it cannot compute your layout, detect mobile responsiveness, or execute client-side hydration, resulting in catastrophic ranking drops.

2. Using robots.txt for Security or Secret URLs

The robots.txt file is publicly accessible to anyone on the internet. Adding Disallow: /super-secret-admin-portal-v2/ does not protect the page; it publicly advertises the exact path to malicious attackers. Use server-level authentication (OAuth, HTTP Basic Auth) to protect private endpoints.

3. Believing robots.txt Disallow Guarantees De-Indexation

robots.txt prevents crawlers from downloading a page, but it does not prevent Google from indexing the URL. If other websites link to https://example.com/disallowed-page, Google may index the bare URL without a description. To guarantee de-indexation, allow the page to be crawled and include a <meta name="robots" content="noindex"> tag.


Frequently Asked Questions About robots.txt

Does robots.txt stop Google from indexing a URL?

No. robots.txt prevents Googlebot from crawling and downloading the page content. However, if the disallowed URL has external backlinks or internal links pointing to it, Google can still index the bare URL in search results without a snippet. To completely remove a page from the index, use a noindex tag and keep the page crawlable.

What is the difference between Disallow: / and Disallow:?

  • Disallow: / blocks crawlers from accessing every URL on the domain (the entire site).
  • Disallow: (with an empty path) means nothing is disallowed, permitting crawlers to access all pages on the domain.

Where must the robots.txt file be located?

The robots.txt file must be located in the root directory of your website domain (https://example.com/robots.txt). Search engine bots will not look for or obey robots files placed in subdirectories (e.g., https://example.com/blog/robots.txt).

Should I block AI crawlers like GPTBot and ClaudeBot?

It depends on your business goals. If you want to prevent AI companies from scraping your content to train foundation models, block GPTBot and ClaudeBot. However, if you want your content to be cited and linked as a real-time source in AI search engines (like ChatGPT Search and Perplexity), you should allow search agents like ChatGPT-User and PerplexityBot.

Can I have different robots.txt rules for mobile vs desktop Googlebot?

No. Googlebot uses a single unified crawler infrastructure for both mobile and desktop crawling and does not distinguish between user agents for robots parsing. Standardize all Google directives under User-agent: Googlebot.


Summary and Action Plan

Your robots.txt file is the front gate of your technical web architecture: declare clear user-agent blocks according to RFC 9309, avoid blocking critical CSS and JavaScript rendering assets, protect crawl budget by filtering combinatorial query parameters, declare your primary sitemap index, and strategically manage AI search crawlers to maximize brand visibility in Generative Engine Optimization (GEO).

To validate your robots.txt syntax, eliminate accidental crawl blocks, and benchmark your domain's AI search readiness score, running a multi-engine BugViso site audit validates your robots.txt syntax and scores your AI crawler readiness.

Found this useful? Share it.

See where your site stands

Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.