Four controls that are often confused
Indexing, crawling, canonicalization and sitemaps are not the same thing. A common SEO mistake is to use one control to solve a different problem—for example, blocking a page in robots.txt when the real goal is to remove it from search results.
| Control | Primary purpose | What it does not guarantee |
|---|---|---|
noindex | Keep a page or resource out of search results | It does not secure private content |
robots.txt | Control crawling of URL patterns | A blocked URL can still be known or shown without its content |
rel="canonical" | Signal the preferred URL among duplicate or very similar pages | It is not a command to remove a page from Search |
| XML sitemap | Tell search engines which canonical URLs you want discovered and considered for Search | Submission does not guarantee crawling, indexing or ranking |
noindex or remove/password-protect the content—not a robots.txt block alone.
What should you do with common page types?
| Page type | Index? | In sitemap? | Recommended treatment |
|---|---|---|---|
| Main tutorials, articles, product or service pages | Yes | Yes | Return 200, use a self-referencing canonical, keep internally linked and include the preferred URL in the sitemap. |
| Duplicate or near-duplicate URLs | Usually one preferred version | Preferred URL only | Use redirects where possible or rel="canonical" to consolidate signals. Do not list competing duplicate URLs in the sitemap. |
| Thank-you / confirmation pages | Usually no | No | Use noindex if the page should remain accessible but not appear in Search. |
| Login, account, dashboard or confidential pages | No | No | Protect private content with authentication. noindex can be supplemental, but it is not a security control. |
| Internal search-result pages | Usually no | No | Keep low-value search-result URLs out of the sitemap. Use noindex or crawl controls depending on scale and architecture. |
| Filter / faceted-navigation URLs | Only when they provide unique search value | Only selected indexable URLs | Prevent crawl traps and duplicate URL combinations. If a filtered page is intended to rank, make it useful, stable and internally linked. |
| Pagination pages | Often yes | Optional | Do not automatically noindex page 2, 3, etc. Each useful page can be crawlable and self-canonical if it contains distinct items. |
| Privacy policy, terms, disclaimer | Depends | Optional | Do not noindex them simply because they have low SEO value. Decide based on whether you want them discoverable in Search. |
| 404 / removed URLs | No | No | Return a real 404 or 410 status. A noindex tag is not a substitute for the correct HTTP status. |
| PDFs or other non-HTML files you do not want indexed | No | No | Use an HTTP X-Robots-Tag: noindex response header. |
Use noindex when the URL can exist but should not appear in Search
For an HTML page, the common implementation is:
<meta name="robots" content="noindex">
For a PDF, image or other non-HTML resource, use an HTTP header such as:
X-Robots-Tag: noindex
robots.txt, the crawler may never see the noindex rule.
If the page contains confidential information, use login/password protection or remove the content. A noindex tag only controls search indexing; anyone with the URL may still be able to open the page.
Use canonicalization when several URLs contain the same or very similar content
Canonicalization tells Google which URL you prefer as the representative version of a duplicate cluster. Redirects and rel="canonical" are stronger signals than sitemap inclusion.
<link rel="canonical" href="https://www.example.com/preferred-page">
- Use a self-referencing canonical on important indexable pages.
- Keep internal links pointing to the preferred URL.
- Put the preferred canonical URL—not all duplicates—in your sitemap.
- Use a permanent redirect when the duplicate URL no longer needs to exist for users.
- Do not use
noindexas your normal duplicate-content consolidation method.
Your sitemap should describe the URLs you want shown in Search
A sitemap is a discovery and canonicalization signal. Include fully qualified canonical URLs that you actually want search engines to find and consider for indexing.
Good sitemap candidates
- Main articles and tutorials
- Important product/service pages
- Useful category or hub pages
- Canonical media or landing pages intended for Search
Usually exclude
- Noindex URLs
- Redirecting URLs
- 404/410 URLs
- Duplicate parameter versions
- Private, login or temporary utility pages
Submitting a sitemap is a hint, not a guarantee. Google can discover URLs through links and other methods, and it may decide not to index a submitted URL.
Use robots.txt to manage crawling—not as a noindex system
robots.txt is most useful when a site generates large groups of URLs that do not need to be crawled, such as certain filters, calendar combinations or other crawl traps.
User-agent: *
Disallow: /internal-search/
Sitemap: https://www.example.com/sitemap.xml
A blocked page can still be known to Google from links. Because Google cannot crawl the blocked content, it cannot reliably read a noindex tag on that page.
A five-question indexation decision process
Should a searcher land on this URL?
If yes, keep evaluating it as a possible indexable page. If no, consider noindex, removal or authentication.
Is this the preferred version?
If a better duplicate exists, redirect or canonicalize toward the preferred version.
Does Google need to crawl it?
If the page needs a noindex tag read, allow crawling. For large unwanted URL spaces, crawl controls may be more appropriate.
Should it be in the sitemap?
Include the canonical URL when you want it discovered and eligible for Search. Exclude noindex, broken, redirected and duplicate URLs.
Does the HTTP status match reality?
Serve 200 for real pages, 3xx for redirects, and 404/410 for URLs that do not exist or have been removed.
Check the result in Google Search Console
- Use URL Inspection to see whether Google can crawl the URL and which canonical it selected.
- Check Page Indexing for groups such as excluded by noindex, duplicate, redirect, not found or crawled/discovered but not indexed.
- Use the Sitemaps report to confirm that the file is processed successfully.
- After a meaningful fix, request recrawling for a small number of important URLs. Repeated requests do not make crawling instant.
Indexing mistakes to avoid
- Blocking a URL in robots.txt and expecting Google to read its noindex tag.
- Listing noindex or redirected URLs in the XML sitemap.
- Using noindex instead of a canonical for ordinary duplicate pages.
- Canonicalizing every paginated page to page 1 even though each page contains different items.
- Using noindex as a security mechanism for private data.
- Noindexing privacy or legal pages automatically without considering whether users may search for them.
- Serving an error message with HTTP 200 instead of a real 404/410 status.