robots.txt and sitemap.xml explained, with examples
Two small files at the root of your site tell crawlers where they may go and which pages you want found. Here's how each works, with examples.
Two plain files at the root of your website help search engines work with it. robots.txt tells crawlers which parts of the site they may and may not crawl. sitemap.xml lists the pages you want them to find. They work well together, but neither one guarantees that a page is indexed or kept out of search results, and misunderstanding that is behind most of the problems below.
What robots.txt does
robots.txt is a plain-text file that must live at the root of a host, for example https://example.com/robots.txt. It applies only to that exact protocol and host, so shop.example.com needs its own file. Crawlers that follow the Robots Exclusion Protocol, including Googlebot and Bingbot, read it before crawling and follow its rules.
It's a set of requests, not a security feature. Badly behaved bots can ignore it, and the file itself is public, so listing a "secret" folder in it actually advertises that folder. Protect private content with authentication.
A simple robots.txt
User-agent: *Disallow: /admin/Disallow: /cartDisallow: /searchSitemap: https://example.com/sitemap.xml
Line by line: the rules apply to every crawler (*); don't crawl anything under /admin/, or any URL starting with /cart or /search; and the sitemap is at the address given. The Sitemap line must be a full URL, and it applies regardless of which user-agent group it sits in.
How the rules are read
- Rules match the start of the URL path, and they're case-sensitive:
Disallow: /Admindoesn't block/admin. - Google supports two wildcards:
*for any sequence of characters and$for the end of the URL.Disallow: /*.pdf$blocks URLs ending in.pdf. - When
AllowandDisallowrules conflict, Google follows the most specific (longest) matching rule, andAllowwins a tie. - A crawler obeys only the most specific group that names it. If you have a
User-agent: Googlebotgroup, Googlebot ignores theUser-agent: *group entirely, so repeat any shared rules. Disallow: /blocks the whole site. An emptyDisallow:blocks nothing.
Blocking crawling is not blocking indexing
This is the most important point about robots.txt. A disallowed URL can still appear in search results, usually without a description, if other pages link to it: Google knows the URL exists but hasn't read it. To keep a page out of search results, let it be crawled and add a noindex directive, either <meta name="robots" content="noindex"> in the HTML or an X-Robots-Tag: noindex response header. If robots.txt blocks the page, Google never sees the noindex.
Rules Google ignores
Google doesn't support noindex inside robots.txt (it stopped honoring the unofficial rule in 2019) and ignores Crawl-delay. Bing does honor Crawl-delay. If Googlebot is putting too much load on your server, fix the server capacity or, in an emergency, temporarily return 503 or 429 responses.
What sitemap.xml does
An XML sitemap is a list of the URLs you'd like search engines to crawl and index. It matters most for large sites, new sites with few links from elsewhere, and pages that are hard to reach through internal links. A minimal sitemap looks like this:
<?xml version="1.0" encoding="UTF-8"?><urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"><url><loc>https://example.com/</loc><lastmod>2026-09-20</lastmod></url><url><loc>https://example.com/pricing</loc><lastmod>2026-08-02</lastmod></url></urlset>
- Size limits: each sitemap file can hold up to 50,000 URLs and be up to 50 MB uncompressed. Larger sites split their URLs across several sitemaps and list them in a sitemap index file.
- lastmod: Google uses the last-modified date if it's consistently accurate. Update it when the page's content meaningfully changes, not on every deploy.
- priority and changefreq: Google ignores both, so there's no need to tune them.
- Only list URLs you want indexed: canonical, indexable pages that return
200. Leave out redirects, error pages,noindexpages and duplicate URLs with tracking parameters.
Creating and submitting a sitemap
You rarely need to write one by hand. WordPress has generated a basic sitemap at /wp-sitemap.xml since version 5.5, and SEO plugins replace it with their own; Shopify and most hosted builders create /sitemap.xml automatically; in Next.js, an app/sitemap.ts file generates one. Once it exists, tell search engines where it is in two ways: add a Sitemap: line to robots.txt, and submit it in the Sitemaps report in Google Search Console (and in Bing Webmaster Tools). Search Console then shows whether it could be read and how many of its URLs are indexed.
How to check your robots.txt and sitemap
Start by opening both files in your browser: /robots.txt and the sitemap URL. Then run the page you care about through Rudra's free SEO checker. It fetches your robots.txt and reports whether it blocks the tested page for the generic * user-agent and whether it declares a sitemap. It then looks for the sitemap (at the declared location, otherwise /sitemap.xml and /sitemap_index.xml), checks that it parses as a sitemap or sitemap index, and reports whether the entries carry lastmod dates and whether the tested URL is listed.
For Google's own view, use the robots.txt report and the URL Inspection tool in Search Console, which show whether a specific URL is blocked and whether it's indexed.
llms.txt file, an emerging convention for pointing AI tools at your most useful content. It's optional and not a web standard, so it's reported as information only.Common mistakes
- Launching with the staging robots.txt. A leftover
Disallow: /from development can block the entire live site. Check robots.txt on launch day. - Blocking CSS and JavaScript that pages need to render, so Google can't see the page as visitors do.
- Using robots.txt to remove pages from Google. Use
noindex(and let the page be crawled), or the Removals tool in Search Console for urgent cases. - Combining a robots.txt block with noindex on the same page, which stops the
noindexfrom being seen. - Listing the wrong URLs in the sitemap: redirects, 404s,
noindexpages,http://versions, or a different domain. - Setting every lastmod date to today on each build, which makes the dates useless to search engines.
- A robots.txt that returns a server error. Google treats a missing robots.txt (404) as "crawl everything", but a 5xx error can make it stop crawling the site until the file is available again.
Frequently asked questions
Do I need a robots.txt file?
Not necessarily. If there's nothing you want to keep crawlers out of, you can skip it; a robots.txt that returns 404 means everything may be crawled. Many sites still add one to block low-value areas like internal search results and to point to the sitemap.
Do I need an XML sitemap?
Small sites with good internal linking are usually found without one, but a sitemap costs little and helps search engines find new and updated pages sooner. Large sites, new sites and sites with deep or poorly linked pages benefit most.
Will adding a page to my sitemap get it indexed?
No. A sitemap helps search engines discover URLs, but whether a page is indexed depends on its quality, whether it duplicates other pages, and whether it can be crawled and isn't marked noindex.
How do I block AI crawlers with robots.txt?
Add a user-agent group for each crawler you want to block, such as GPTBot (OpenAI) or CCBot (Common Crawl), with Disallow: /. Google-Extended controls whether Google may use your content for its AI models, without affecting Google Search. These rules only work for crawlers that choose to respect robots.txt.
Where should robots.txt and sitemap.xml be located?
robots.txt must be at the root of each host, such as https://example.com/robots.txt. A sitemap can live anywhere on the host, but /sitemap.xml is conventional; declare its location in robots.txt and submit it in Google Search Console.