One site was crawled as 4 sites: http/https × www/non-www. Crawl efficiency: 25%.
Text-book waste, rarely measured. A site never forced a single canonical host, so Googlebot was crawling four full copies: http+www, http+non-www, https+www, https+non-www. All resolved, none redirected.
Log analysis: only ~25% of crawl requests hit the canonical version. The other 75% re-crawled identical duplicates.
Fix was a handful of 301s forcing everything to https://non-www. No content, no link changes.
Crawl efficiency on the canonical URLs rose to 90%+ within a month. Effective recrawl frequency of every real page roughly quadrupled, because Google stopped wasting three of every four fetches.
Here's the thing — you weren't short on budget. You were spending it four times on the same pages.
Force one host, one protocol. Day-one hygiene that quietly multiplies your crawl.
Budget Myths
@CrawlBudgetMyths
One site was crawled as 4 sites: http/https × www/non-www. Crawl efficiency: 25%.
Этот пост опубликован в Telegram-канале Budget Myths. Подписаться можно по ссылке: @CrawlBudgetMyths.