Pick a template to instantly load a proven starting configuration.
Wildcards: * matches any characters (e.g. /checkout/*) and a trailing $ matches the end of a URL (e.g. /feed$). Rules are matched by length — the longest matching rule wins, and on a tie Allow beats Disallow.
Evaluated against the User-agent: * group.
The longest matching rule Allow: / (pattern length 1) won, so the URL is allowed.
Rule matching follows the RFC 9309 standard: longest pattern wins, ties favor Allow over Disallow.
A standard robots.txt file operates via rule groups composed of User-agent lines followed by Allow and Disallow directives:
User-agent: * for all crawlers, User-agent: Googlebot, or User-agent: GPTBot).Disallow: /admin/ or Disallow: /checkout/).Allow: /admin/login.css within a disallowed /admin/ directory).Sitemap: https://example.com/sitemap.xml).*) and End-of-Path Anchors ($)RFC 9309 standardizes pattern-matching symbols across major search engines:
):** Matches zero or more arbitrary characters (e.g. Disallow: /*.pdf` blocks all PDF files across the entire domain).Disallow: /*.json$ blocks URLs ending in .json without blocking /data.json?preview=true).For large e-commerce platforms and content websites, blocking low-value filter permutations and administrative query strings conserves crawler bandwidth ('crawl budget'), ensuring search engines prioritize indexing your highest-converting landing pages.
*, Googlebot, Bingbot, or AI Crawlers like GPTBot).* and $ end-of-path anchors).robots.txt file download.Scenario: Blocking admin and cart pages while allowing public products and attaching a sitemap.
User-agent: * | Disallow: /admin/, /cart/, /checkout/ | Sitemap: https://shop.com/sitemap.xml
User-agent: * Disallow: /admin/ Disallow: /cart/ Disallow: /checkout/ Sitemap: https://shop.com/sitemap.xml
Protects private transaction paths while ensuring catalog discoverability.
Scenario: Restricting automated LLM crawlers without affecting Google or Bing rankings.
User-agent: GPTBot, CCBot | Disallow: / | User-agent: * | Allow: /
User-agent: GPTBot User-agent: CCBot Disallow: / User-agent: * Allow: /
Restricts AI training scrapers while maintaining search engine indexation.
Generate XML sitemaps to link directly inside your robots.txt file.
Configure page-level 'noindex' and 'nofollow' search robot directives.
Simulate how your allowed pages appear in search listings.