Robots.txt Generator

Generate a robots.txt file for your website. Choose default crawl access, add allow or disallow rules, include a sitemap URL, and copy the ready-to-use code.

Configuration

Custom Directories

Add specific directories to allow or disallow.

Generated Output

Create a Robots.txt File Online

A robots.txt file tells search-engine crawlers which areas of your website they may or may not crawl. Use this free Robots.txt Generator to create basic crawl rules without writing the file manually.

Choose whether crawlers should be allowed by default, add directories that should be allowed or disallowed, set an optional crawl delay, and include your XML sitemap URL. The tool creates the robots.txt code instantly so you can copy it and upload it to your website.


How to Use the Robots.txt Generator

  1. Choose the default rule for all crawlers:

    • Allow All for normal websites

    • Disallow All only when you want to block crawling across the whole site

  2. Add directories you want to allow or disallow.

  3. Enter a crawl delay if it is relevant for the crawlers you want to manage.

  4. Add the full URL of your XML sitemap.

  5. Copy the generated code.

  6. Save it as robots.txt.

  7. Upload it to the root directory of your website.

For example, your file should normally be available at:

https://yourwebsite.com/robots.txt

What Is a Robots.txt File?

A robots.txt file is a plain text file placed in the root directory of a website. It contains instructions for web crawlers, such as search-engine bots, about which URL paths they may access.

A basic robots.txt file can look like this:

User-agent: *
Disallow: /private/
Disallow: /wp-admin/
Sitemap: https://yourwebsite.com/sitemap.xml

In this example, the rules apply to all crawlers, block crawling of the /private/ and /wp-admin/ directories, and provide the location of the website sitemap.


Robots.txt Directives Explained

User-agent

The User-agent line identifies the crawler that a group of rules applies to.

User-agent: *

An asterisk means the rule applies to all crawlers.

Disallow

Use Disallow to request that crawlers do not access a specific path.

Disallow: /private/

This rule applies to URLs inside the /private/ directory.

Allow

Use Allow when you need to permit crawling of a specific path within a restricted area.

Disallow: /assets/
Allow: /assets/public/

This can be useful when most of a directory should not be crawled, but one subdirectory should remain accessible.

Sitemap

The Sitemap line tells search engines where they can find your XML sitemap.

Sitemap: https://yourwebsite.com/sitemap.xml

Use the full, absolute sitemap URL, including https://.

Crawl-delay

Some crawlers may respect a Crawl-delay rule, which requests a pause between their requests.

Crawl-delay: 10

However, not every crawler supports this directive. Google does not use Crawl-delay as part of its robots.txt interpretation, so do not rely on it to control Googlebot’s crawl rate.


Common Robots.txt Examples

Allow All Crawlers

Most public websites should allow crawling by default.

User-agent: *
Disallow:

Sitemap: https://yourwebsite.com/sitemap.xml

Block a Private Directory

User-agent: *
Disallow: /private/
Disallow: /internal/

Basic WordPress Example

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://yourwebsite.com/sitemap.xml

Block All Crawlers Temporarily

User-agent: *
Disallow: /

Use this carefully. It asks compliant crawlers not to crawl any pages on the website. Do not leave this rule active on a live public website unless that is your intention.


Where Should You Place Robots.txt?

Your robots.txt file must be placed at the top level of the relevant domain or subdomain.

Correct:

https://yourwebsite.com/robots.txt

Incorrect:

https://yourwebsite.com/files/robots.txt
https://yourwebsite.com/blog/robots.txt

A robots.txt file applies only to the host, protocol, and port where it is published. For example, rules at https://www.example.com/robots.txt do not automatically apply to https://example.com/ or https://shop.example.com/.


Important Limitations of Robots.txt

Robots.txt is useful for managing crawler access, but it is not a security tool.

Do Not Use Robots.txt to Protect Private Information

Robots.txt rules are voluntary instructions. Well-behaved crawlers may follow them, but malicious bots can ignore them. Never use robots.txt to protect confidential files, passwords, customer information, backups, or private documents.

Use authentication, password protection, server permissions, or other proper access controls for sensitive content.

Disallowed URLs Can Still Appear in Search Results

Blocking a page in robots.txt does not always guarantee that its URL will stay out of search results. If other websites link to that URL, a search engine may still discover and display the URL without crawling its content.

If you need a page removed from search results, use an appropriate noindex directive while allowing the page to be crawled, password-protect the page, or remove it entirely.

Avoid Blocking Important Resources

Do not block CSS, JavaScript, image, or other files required for search engines to properly render and understand your pages. Blocking essential resources can affect how search engines interpret your content.


Robots.txt Best Practices

  • Keep the file simple and easy to review.

  • Use paths that begin with /.

  • Use a full URL for the sitemap location.

  • Test important pages after updating crawl rules.

  • Do not block pages simply because they are low priority.

  • Avoid blocking essential CSS or JavaScript files.

  • Review your robots.txt file after website migrations, redesigns, plugin changes, or staging-to-live deployments.

  • Make sure the file returns a successful HTTP status code when opened in a browser.

Frequently Asked Questions

Not every website needs one. A robots.txt file is helpful when you need to manage crawler access to specific paths or provide the location of your sitemap. If you do not need crawl restrictions, you can use a simple allow-all file or leave it absent.

Not reliably. Robots.txt controls crawling, not guaranteed indexing. A blocked URL can still appear in search results if Google discovers it through links. Use noindex or password protection when you need to prevent a page from appearing in search.

A common rule is to disallow /wp-admin/ while allowing admin-ajax.php. Do not block important public pages, uploaded media, CSS, or JavaScript resources without understanding the impact.

Yes, adding your sitemap URL is a useful way to help crawlers discover it. Use the full URL, such as https://yourwebsite.com/sitemap.xml.

No. Google does not support the Crawl-delay directive in robots.txt. Other crawlers may interpret it differently.

Save the generated code in a file named robots.txt and upload it to your website’s root directory so it is accessible at https://yourwebsite.com/robots.txt.