Updated September 2026

robots.txt: Control Crawling Without Confusing It With Indexing

A robots.txt file tells compatible crawlers which URLs they may access. It is mainly a crawl-management tool—not a reliable method for keeping a web page out of Google Search.

Core concept

What robots.txt is actually for

Google describes robots.txt as a way to tell search engine crawlers which URLs they can access. It is useful for managing crawler traffic and preventing crawling of unimportant or repetitive areas.

Important: blocking a URL in robots.txt does not guarantee that the URL will disappear from Google Search. If you need a crawlable page excluded from search results, use an appropriate noindex directive instead.
Quick navigation: location · syntax · examples · robots vs noindex · sitemap · mistakes
Location & scope

Place robots.txt at the root of the host it controls

For https://www.example.com/, the file must be available at:

https://www.example.com/robots.txt

The file applies only to the protocol, host and port where it is published. A robots.txt file on www.example.com does not automatically control a different subdomain such as shop.example.com.

robots.txt is a public text file. Never use it as a security mechanism for confidential URLs or files. Private material should be protected through authentication or another real access-control method.

Basic syntax

User-agent, Disallow, Allow and Sitemap

A group normally starts with a crawler name and then contains crawl rules.

User-agent: *
Disallow: /private-area/
Allow: /private-area/public-guide.html

Sitemap: https://www.example.com/sitemap.xml
  • User-agent: identifies the crawler or crawler group.
  • Disallow: defines a path the crawler should not access.
  • Allow: can permit a more specific path inside a broader blocked area.
  • Sitemap: gives the fully qualified URL of a sitemap file.

Google notes that rules are case-sensitive, and crawlers are allowed to access URLs that are not blocked by a matching rule.

Examples

Common robots.txt patterns

Allow normal crawling

If you have no restrictions, you do not need to add an explicit allow-all rule. If you prefer to show it:

User-agent: *
Allow: /

Block all compatible crawlers from all URLs

User-agent: *
Disallow: /

This is sometimes used temporarily on a non-public environment, but do not rely on robots.txt to protect confidential staging content.

Block one directory

User-agent: *
Disallow: /internal-search/

Block one page from crawling

User-agent: *
Disallow: /search-results.php

Rule for Googlebot only

User-agent: Googlebot
Disallow: /example-directory/
Correction from the older tutorial: a rule such as User-agent: Googlebot followed by Disallow: / blocks Googlebot from crawling the site; it does not “allow Googlebot to all pages.”
Crawling vs indexing

Why robots.txt and noindex should not be treated as interchangeable

A URL blocked by robots.txt can still be discovered through external or internal links. Because Google cannot crawl the page content, the URL may still appear in search results without a normal description.

If your goal is to keep a normal web page out of Google Search, the safer pattern is usually:

  1. Allow Google to crawl the page.
  2. Use a supported noindex meta robots tag or X-Robots-Tag header.

For a complete decision guide, see Index, Noindex, Canonical or Sitemap?

Sitemap discovery

Reference the sitemap in robots.txt

Sitemap: https://www.example.com/sitemap.xml

The sitemap URL must be fully qualified. You can also submit the sitemap through Google Search Console.

Old advice removed: Google's unauthenticated sitemap “ping” endpoint was deprecated. Do not use old URLs such as google.com/ping?sitemap=....

The sitemap itself should normally contain the canonical indexable URLs you want to appear in search results. See our new website SEO checklist for launch-related checks.

Common mistakes

robots.txt problems worth checking

  • Using robots.txt to hide sensitive information.
  • Blocking CSS or JavaScript required for Google to understand an important page.
  • Accidentally blocking an entire production site with Disallow: /.
  • Trying to use robots.txt as a replacement for noindex.
  • Publishing the file in a subdirectory instead of the host root.
  • Writing rules for one host and assuming they automatically apply to another subdomain.
  • Using stale sitemap ping instructions.
  • Assuming malicious or non-compliant bots must obey the file.
Summary

Use robots.txt for crawler access, not for secrecy

robots.txt is a simple and useful crawl-control file. Keep the rules minimal, test changes carefully and use the correct indexing or security mechanism when your goal is something other than crawl management.

Official references

References and related Plus2Net guides