Online Marketing Guy
Back to glossary

robots.txt

The robots.txt file tells search engine crawlers which areas of a website they may or may not crawl. It controls crawling, not indexing.

What is robots.txt?

The robots.txt file is a plain text file in the root directory of a website (e.g. example.com/robots.txt) that tells search engine crawlers which areas of the site they may or may not crawl. It uses simple rules such as User-agent, Disallow and Allow, and can also point crawlers to the XML sitemap.

Typical uses are keeping crawlers away from internal search results, filter combinations, shopping carts or staging areas – so that the crawl budget is spent on pages that matter.

Two important limitations:

  • robots.txt controls crawling, not indexing. A blocked URL can still appear in search results if other pages link to it. To keep a page out of the index, use a noindex meta tag instead – and do not block that page in robots.txt, otherwise search engines cannot see the noindex instruction.
  • It is not a security measure. The file is public, and badly behaved bots can ignore it.

For international websites, remember that robots.txt applies per host: every country domain or subdomain needs its own file. Also check that different search engines’ crawlers, such as Baiduspider or Yeti from Naver, are not accidentally blocked if those markets matter to you.