The Magic Codex

robots.txt Generator

Build a valid robots.txt file — per-crawler allow/disallow rules, crawl-delay, and sitemap lines — with a live preview and one-click copy or download. Free, private, no signup.

Crawlers read this line to find your sitemap. It is a full URL, not a path.


        

How to write a robots.txt file

robots.txt is the gatekeeper of your website: a plain-text file that tells search-engine crawlers which corners of your site they may visit. Writing it by hand is easy to get wrong — one stray slash can unblock your admin panel or nuke your whole index — so this tool builds it for you.

  1. Add a crawler group above. The default User-agent: * group covers every crawler.
  2. Add Allow and Disallow rows. Paths start with / — /admin/ blocks a folder, /secret.pdf blocks one file.
  3. Set a crawl-delay if you need one (seconds between requests) and add your sitemap URL.
  4. Copy or download the preview, upload it to your site's root as robots.txt, and test it in Search Console.

When rules overlap

Crawlers match each URL against all Allow and Disallow lines in their group and follow the longest match. That means Disallow: /docs/ plus Allow: /docs/public/ opens just the public folder. The wildcard * matches any sequence of characters, and $ at the end anchors a path — /*.pdf$ blocks every PDF.

What robots.txt cannot do

robots.txt does not make pages private. Blocked URLs can still be visited by anyone, linked to by anyone, and — if other sites link to them — occasionally indexed by Google anyway. Use real passwords or a noindex meta tag for anything genuinely secret.

What is a robots.txt file?
robots.txt is a small text file at your site's root (for example yoursite.com/robots.txt) that tells search-engine crawlers which pages they may visit and which they should skip. It is a request, not a lock — polite crawlers obey it, malicious ones ignore it.
What does User-agent: * mean?
User-agent: * is the wildcard group — its rules apply to every crawler that doesn't have its own named group. A group named User-agent: Googlebot applies only to Google's crawler. Crawlers use the most specific group that matches them.
How do Allow and Disallow work together?
Disallow blocks a path; Allow carves an exception inside it. When both match the same URL, the longer (more specific) path wins. So Disallow: / with Allow: /public/ lets crawlers see /public/ but nothing else.
What is Crawl-delay?
Crawl-delay asks a crawler to wait that many seconds between requests, which protects small servers from being hammered. It is a courtesy rule: Google's crawler ignores it and expects you to set crawl speed in Search Console instead.
How do I block my whole site from search engines?
Use two lines: User-agent: * followed by Disallow: /. The slash blocks every path on the site. This tool's "Block everything" preset writes exactly that. Remember it blocks indexing, but the site is still public — anyone with the URL can visit it.
Where do I put the finished robots.txt?
Upload it to your website's root so it loads at yoursite.com/robots.txt, exactly that path. Add a Sitemap: line pointing at your sitemap.xml while you're at it, and test the result with Search Console's robots.txt tester.

More from the codex

Questions about this tool?

Crawler misbehaving or have a suggestion? Email [email protected] — every message is read by the indie dev behind the codex.