Log files vs. crawler simulation: the 41% gap
Across 188 audited domains, the URLs a desktop crawler reports as "crawlable" and the URLs Googlebot actually hit in server logs diverged by a median of 41.2%.
— Crawler-found, log-confirmed: 58.8% ▓▓▓▓▓▓░░░░
— Crawler-found, never crawled: 31.0% ▓▓▓░░░░░░░
— Log-only (orphan in crawl): 10.2% ▓░░░░░░░░░
p90 of the never-crawled bucket sat at 54.7% on sites over 100k URLs.
A crawl simulation tells you what's reachable; logs tell you what's valued. Use the simulator to find broken structure, but never size your crawl-budget problem without logs — the delta between intent and reality is where indexation actually leaks. Run both; reconcile the orphan set first.
Crawl & Render
@CrawlAndRender
Log files vs. crawler simulation: the 41% gap
Этот пост опубликован в Telegram-канале Crawl & Render. Подписаться можно по ссылке: @CrawlAndRender.