SEO Glossary

Robots Meta Tag

The robots meta tag is an HTML tag placed in a page's <head> section that gives page-specific instructions to search engine crawlers — most commonly "noindex" (don't index this page) or "nofollow" (don't follow links on this page).

Unlike robots.txt, which gives site-wide or directory-level crawling instructions from a single file at the domain root, the robots meta tag is set individually on each page and specifically controls indexing and link-following behaviour for that one page. A common form looks like <meta name="robots" content="noindex, nofollow">. The most frequently used directive is noindex, which tells search engines not to include that specific page in their index even if they're able to crawl it — commonly used for thank-you pages, internal search results, staging content, or duplicate filter/sort variations that shouldn't compete for rankings. nofollow in this context tells crawlers not to pass ranking credit through any links on that page. A page can be crawled but excluded via a robots meta tag, which is different from being blocked in robots.txt — a robots.txt block can actually prevent noindex from ever being seen, since the crawler may never fetch the page to read the tag in the first place.

Why it matters

An accidentally-left-on noindex tag — commonly a leftover from a staging or development environment — is one of the most common, entirely self-inflicted reasons a real page fails to appear in search results, and checking for it directly in a page's HTML source is a fast, simple diagnostic step.

How Illumae helps

site audit checks pages for a rendered noindex meta tag as part of every crawl, since this is one of the most common reasons an otherwise healthy page is being excluded from analysis — and the same signal that would exclude it from Google's index too.

Frequently asked questions

What's the difference between robots.txt and the robots meta tag?+

robots.txt is one file giving site-wide or path-level crawling instructions, checked before a crawler even requests a page. The robots meta tag is set per-page in that page's own HTML and controls indexing/link-following for that specific page once it's been fetched.

Can a robots.txt block prevent a noindex tag from working?+

Yes, and this is a common mistake — if robots.txt blocks a page from being crawled at all, the crawler never gets to read that page's robots meta tag, so a noindex intention set there may never actually be seen or honoured.

Why would a real, important page accidentally have noindex on it?+

Most commonly, a leftover setting from when the page was in a staging or development environment gets carried over unintentionally when it goes live — a good reason to specifically check for this after any site launch or migration.

Try your first 3 articles free

No credit card required. See real output for your own business before you pay anything.