robots.txt disallow vs noindex on facets — they do opposite jobs
Faceted nav was spawning 80k crawlable URLs. Client's dev had 'fixed' it with robots.txt. Google was still indexing them.
— robots.txt Disallow: blocks the crawl, but Google can still index a URL it found via links, just with no snippet. Saves crawl budget.
— noindex: needs the page crawlable to be read, drops it from the index, but burns crawl budget to do it.
The move: noindex first, let Google digest the drop over a few weeks, THEN Disallow in robots once they're gone.
Do them in the wrong order and you trap zombie URLs in the index forever. Sequence matters here.
Sitemap Hustle
@SitemapHustle
robots.txt disallow vs noindex on facets — they do opposite jobs
Этот пост опубликован в Telegram-канале Sitemap Hustle. Подписаться можно по ссылке: @SitemapHustle.