Playbook: pull real Googlebot logs instead of guessing
Crawl Stats is sampled and aggregated. Logs are ground truth. Get them:
— Ask your host/CDN for raw access logs, ideally 30 days. Nginx/Apache combined format is fine.
— Filter to Googlebot. Match user-agent AND verify by reverse DNS (the host must resolve to googlebot.com / google.com). Fake Googlebots are everywhere.
— Group hits by URL path. Sort descending.
— Now look at the top 50. Are they your money pages, or junk?
Actually, the first time most people do this they find Google hammering faceted filters, old paginations, and tracking-param URLs they forgot existed.
No verification, no analysis — half your "Googlebot traffic" is scrapers wearing its name. Verify before you conclude anything.
Budget Myths
@CrawlBudgetMyths
Playbook: pull real Googlebot logs instead of guessing
Этот пост опубликован в Telegram-канале Budget Myths. Подписаться можно по ссылке: @CrawlBudgetMyths.