noindex vs. robots disallow: the crawl-budget tradeoff
Measured crawl reallocation on 96 domains that pruned thin pages. Both methods cut index bloat, but the budget math diverges sharply.
— Disallow in robots.txt: crawl hits to blocked paths ↓ 91.4% in 14 days ▓▓▓▓▓▓▓▓▓░
— Meta noindex: crawl hits ↓ only 12.7% over same window ▓░░░░░░░░░
Reason: noindex still requires the crawl to read the tag, so Googlebot keeps visiting (median 8.3 revisits before backing off).
So what: disallow protects crawl budget but can strand pages in a "indexed, blocked" limbo if they already rank. noindex removes from index cleanly but spends budget. Sequence it: noindex first, wait for de-indexing (~p90 28 days), then disallow. Doing both at once is the classic mistake — the bot never sees the noindex.
Crawl & Render
@CrawlAndRender
noindex vs. robots disallow: the crawl-budget tradeoff
Этот пост опубликован в Telegram-канале Crawl & Render. Подписаться можно по ссылке: @CrawlAndRender.