Auditing only top URLs hides where 80% of waste lives
Reviewed 94 audit reports: 71% sampled from the homepage and top-nav only, covering a median 4% of total URLs. The crawl-budget leaks sit in the deep tail they never opened.
— Audits with shallow sampling: 71.0% ▓▓▓▓▓▓▓░░░
— Median URL coverage sampled: 4.0%
— Waste found in unsampled tail: 80%+ of total
— Detection rate of soft-404/facet bloat: ↓ 6x
Failure mode: a crawler set to depth 3 with a 10k cap reports a clean site while millions of facet URLs rot below.
Fix: stratified sampling — pull URLs by depth band, by template, and from server logs (what Google actually fetches), not just the menu. Audits that added log-based sampling surfaced ↑ 5.4x more crawl waste.
So what: if your audit can't see depth 8, your audit can't see your problem.
Crawl & Render
@CrawlAndRender
Auditing only top URLs hides where 80% of waste lives
Этот пост опубликован в Telegram-канале Crawl & Render. Подписаться можно по ссылке: @CrawlAndRender.