Playbook: triage crawl budget from 30 days of logs
Median site wastes 38.2% of crawl hits on non-indexable URLs (sample: 214 audited domains).
Run this in order:
— Step 1: parse 30d of access logs, filter UA to verified Googlebot via reverse DNS. Typical noise removed: 11.4%
— Step 2: bucket hits by status. Healthy split is 200s ▓▓▓▓▓▓▓▓░░ 82%, 301s ▓░░ 9%, 404/410 ▓░░ 6%
— Step 3: join hit-URLs against your sitemap. Orphan-crawled (in logs, not in sitemap) median = 23.7%
— Step 4: sort wasted hits by path prefix, fix the top 3 prefixes first — they hold p90 of the waste
So what: don't chase individual URLs. The top 3 path prefixes recover a median 71% of wasted budget in one robots.txt or noindex pass.
Crawl & Render
@CrawlAndRender
Playbook: triage crawl budget from 30 days of logs
Этот пост опубликован в Telegram-канале Crawl & Render. Подписаться можно по ссылке: @CrawlAndRender.