Robots.txt is a plain text file at the root of a website that tells search engine crawlers which parts of the site they're allowed to crawl.
Robots.txt lives at a predictable location — yoursite.com/robots.txt — and uses simple rules to allow or disallow crawlers from specific paths on a site. It's a request, not an enforcement mechanism: well-behaved crawlers like Googlebot respect it, but it doesn't actually block access the way a password or firewall would, so it shouldn't be relied on to keep sensitive content private. A common use is blocking crawlers from low-value areas like admin pages, internal search results, or duplicate filtered views, so crawl activity is spent on pages that actually matter. Robots.txt can also point to a site's sitemap.xml, helping crawlers discover a site's full page list more efficiently.
A misconfigured robots.txt file can accidentally block search engines from crawling your entire site, which is one of the more damaging and surprisingly common technical SEO mistakes — usually left over from a staging environment's stricter rules that never got updated for production.
Illumae's site audit reads your site's robots.txt file and respects its rules when crawling, so the pages it checks match what search engines are actually allowed to see.
Not reliably. Disallowing a page in robots.txt stops crawling, but Google can still index a URL it discovers through other links without crawling its content. A noindex meta tag is the more reliable way to keep a specific page out of search results.
At the root of the domain — yoursite.com/robots.txt. A robots.txt file placed anywhere else, like a subfolder, is not recognised by crawlers.
Accidentally leaving a blanket "Disallow: /" rule in place after moving a site from staging to production, which tells every crawler to avoid the entire site.
No credit card required. See real output for your own business before you pay anything.