Robots.txt Generator
Generate a robots.txt file for your website. Choose default crawl access, add allow or disallow rules, include a sitemap URL, and copy the ready-to-use code.
Configuration
Custom Directories
Add specific directories to allow or disallow.
Generated Output
Create a Robots.txt File Online
A robots.txt file tells search-engine crawlers which areas of your website they may or may not crawl. Use this free Robots.txt Generator to create basic crawl rules without writing the file manually.
Choose whether crawlers should be allowed by default, add directories that should be allowed or disallowed, set an optional crawl delay, and include your XML sitemap URL. The tool creates the robots.txt code instantly so you can copy it and upload it to your website.
How to Use the Robots.txt Generator
Choose the default rule for all crawlers:
Allow All for normal websites
Disallow All only when you want to block crawling across the whole site
Add directories you want to allow or disallow.
Enter a crawl delay if it is relevant for the crawlers you want to manage.
Add the full URL of your XML sitemap.
Copy the generated code.
Save it as
robots.txt.Upload it to the root directory of your website.
For example, your file should normally be available at:
https://yourwebsite.com/robots.txt
What Is a Robots.txt File?
A robots.txt file is a plain text file placed in the root directory of a website. It contains instructions for web crawlers, such as search-engine bots, about which URL paths they may access.
A basic robots.txt file can look like this:
User-agent: *
Disallow: /private/
Disallow: /wp-admin/
Sitemap: https://yourwebsite.com/sitemap.xmlIn this example, the rules apply to all crawlers, block crawling of the /private/ and /wp-admin/ directories, and provide the location of the website sitemap.
Robots.txt Directives Explained
User-agent
The User-agent line identifies the crawler that a group of rules applies to.
User-agent: *An asterisk means the rule applies to all crawlers.
Disallow
Use Disallow to request that crawlers do not access a specific path.
Disallow: /private/This rule applies to URLs inside the /private/ directory.
Allow
Use Allow when you need to permit crawling of a specific path within a restricted area.
Disallow: /assets/
Allow: /assets/public/This can be useful when most of a directory should not be crawled, but one subdirectory should remain accessible.
Sitemap
The Sitemap line tells search engines where they can find your XML sitemap.
Sitemap: https://yourwebsite.com/sitemap.xmlUse the full, absolute sitemap URL, including https://.
Crawl-delay
Some crawlers may respect a Crawl-delay rule, which requests a pause between their requests.
Crawl-delay: 10However, not every crawler supports this directive. Google does not use Crawl-delay as part of its robots.txt interpretation, so do not rely on it to control Googlebot’s crawl rate.
Common Robots.txt Examples
Allow All Crawlers
Most public websites should allow crawling by default.
User-agent: *
Disallow:
Sitemap: https://yourwebsite.com/sitemap.xmlBlock a Private Directory
User-agent: *
Disallow: /private/
Disallow: /internal/Basic WordPress Example
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://yourwebsite.com/sitemap.xmlBlock All Crawlers Temporarily
User-agent: *
Disallow: /Use this carefully. It asks compliant crawlers not to crawl any pages on the website. Do not leave this rule active on a live public website unless that is your intention.
Where Should You Place Robots.txt?
Your robots.txt file must be placed at the top level of the relevant domain or subdomain.
Correct:
https://yourwebsite.com/robots.txtIncorrect:
https://yourwebsite.com/files/robots.txt
https://yourwebsite.com/blog/robots.txtA robots.txt file applies only to the host, protocol, and port where it is published. For example, rules at https://www.example.com/robots.txt do not automatically apply to https://example.com/ or https://shop.example.com/.
Important Limitations of Robots.txt
Robots.txt is useful for managing crawler access, but it is not a security tool.
Do Not Use Robots.txt to Protect Private Information
Robots.txt rules are voluntary instructions. Well-behaved crawlers may follow them, but malicious bots can ignore them. Never use robots.txt to protect confidential files, passwords, customer information, backups, or private documents.
Use authentication, password protection, server permissions, or other proper access controls for sensitive content.
Disallowed URLs Can Still Appear in Search Results
Blocking a page in robots.txt does not always guarantee that its URL will stay out of search results. If other websites link to that URL, a search engine may still discover and display the URL without crawling its content.
If you need a page removed from search results, use an appropriate noindex directive while allowing the page to be crawled, password-protect the page, or remove it entirely.
Avoid Blocking Important Resources
Do not block CSS, JavaScript, image, or other files required for search engines to properly render and understand your pages. Blocking essential resources can affect how search engines interpret your content.
Robots.txt Best Practices
Keep the file simple and easy to review.
Use paths that begin with
/.Use a full URL for the sitemap location.
Test important pages after updating crawl rules.
Do not block pages simply because they are low priority.
Avoid blocking essential CSS or JavaScript files.
Review your robots.txt file after website migrations, redesigns, plugin changes, or staging-to-live deployments.
Make sure the file returns a successful HTTP status code when opened in a browser.