What robots.txt is actually for
Google describes robots.txt as a way to tell search engine crawlers which URLs they can access. It is useful for managing crawler traffic and preventing crawling of unimportant or repetitive areas.
noindex directive instead.
Place robots.txt at the root of the host it controls
For https://www.example.com/, the file must be available at:
https://www.example.com/robots.txt
The file applies only to the protocol, host and port where it is published.
A robots.txt file on www.example.com does not automatically control a different subdomain such as shop.example.com.
robots.txt is a public text file. Never use it as a security mechanism for confidential URLs or files. Private material should be protected through authentication or another real access-control method.
User-agent, Disallow, Allow and Sitemap
A group normally starts with a crawler name and then contains crawl rules.
User-agent: *
Disallow: /private-area/
Allow: /private-area/public-guide.html
Sitemap: https://www.example.com/sitemap.xml
- User-agent: identifies the crawler or crawler group.
- Disallow: defines a path the crawler should not access.
- Allow: can permit a more specific path inside a broader blocked area.
- Sitemap: gives the fully qualified URL of a sitemap file.
Google notes that rules are case-sensitive, and crawlers are allowed to access URLs that are not blocked by a matching rule.
Common robots.txt patterns
Allow normal crawling
If you have no restrictions, you do not need to add an explicit allow-all rule. If you prefer to show it:
User-agent: *
Allow: /
Block all compatible crawlers from all URLs
User-agent: *
Disallow: /
This is sometimes used temporarily on a non-public environment, but do not rely on robots.txt to protect confidential staging content.
Block one directory
User-agent: *
Disallow: /internal-search/
Block one page from crawling
User-agent: *
Disallow: /search-results.php
Rule for Googlebot only
User-agent: Googlebot
Disallow: /example-directory/
User-agent: Googlebot followed by Disallow: / blocks Googlebot from crawling the site;
it does not “allow Googlebot to all pages.”
Why robots.txt and noindex should not be treated as interchangeable
A URL blocked by robots.txt can still be discovered through external or internal links. Because Google cannot crawl the page content, the URL may still appear in search results without a normal description.
If your goal is to keep a normal web page out of Google Search, the safer pattern is usually:
- Allow Google to crawl the page.
- Use a supported
noindexmeta robots tag or X-Robots-Tag header.
For a complete decision guide, see Index, Noindex, Canonical or Sitemap?
Reference the sitemap in robots.txt
Sitemap: https://www.example.com/sitemap.xml
The sitemap URL must be fully qualified. You can also submit the sitemap through Google Search Console.
google.com/ping?sitemap=....
The sitemap itself should normally contain the canonical indexable URLs you want to appear in search results. See our new website SEO checklist for launch-related checks.
robots.txt problems worth checking
- Using robots.txt to hide sensitive information.
- Blocking CSS or JavaScript required for Google to understand an important page.
- Accidentally blocking an entire production site with
Disallow: /. - Trying to use robots.txt as a replacement for
noindex. - Publishing the file in a subdirectory instead of the host root.
- Writing rules for one host and assuming they automatically apply to another subdomain.
- Using stale sitemap ping instructions.
- Assuming malicious or non-compliant bots must obey the file.
Use robots.txt for crawler access, not for secrecy
robots.txt is a simple and useful crawl-control file. Keep the rules minimal, test changes carefully and use the correct indexing or security mechanism when your goal is something other than crawl management.