Playbook: read raw access logs before you touch anything else
Crawl problems hide in logs, not in dashboards. Dashboards are sampled and smoothed. Logs are what actually happened.
The pass I run:
— Pull at least four weeks. Less and you're reading noise.
— Split bot from human by user agent, then verify the bot hits by reverse lookup. Fake crawlers are common and they skew everything.
— Group requests by path pattern, not by URL. /product/* tells you more than ten thousand rows ever will.
— Sort by hit count descending. Look at the top twenty patterns and ask one question: do I want the crawler spending its time here?
— Now sort by status code. Anything that isn't a 200 or a 304 inside those top patterns is waste.
— Cross the list against your sitemap. Pages crawled but absent from the sitemap, and pages in the sitemap never crawled — both are findings.
Most sites discover the crawler is spending the bulk of its budget on parameter URLs, ancient redirects and internal search results.
You cannot optimise crawl until you can see it.
The Access Log
@TheAccessLog
Playbook: read raw access logs before you touch anything else
Этот пост опубликован в Telegram-канале The Access Log. Подписаться можно по ссылке: @TheAccessLog.