SOP: protect crawl budget on a large set
Google won't crawl everything — make it crawl the right things.
— Step 1 (Ops): Pull server logs and count bot hits by path. Gate: if >30% of crawls land on noindex/parameter URLs, you're bleeding budget.
— Step 2 (Dev): Block crawl traps (infinite calendars, session params, sort permutations) in robots.txt. Gate: confirm the block doesn't hit indexable URLs.
— Step 3 (Dev): Return fast. Gate: template TTFB above your budget under load → optimize before scaling, slow pages get crawled less.
— Step 4 (SEO): Flatten depth — keep money pages within 3 clicks of the homepage via hubs. Gate: deep orphans don't get crawled.
— Step 5 (Ops): Re-pull logs after launch and compare the crawl distribution. Gate: budget still landing on junk → tighten step 2.
Guardrail: every crawl spent on a noindex URL is a crawl stolen from a money page.
Ship gate: don't publish until all boxes are checked.
Scale Engine SOP
@ScaleEngineSOP
SOP: protect crawl budget on a large set
Этот пост опубликован в Telegram-канале Scale Engine SOP. Подписаться можно по ссылке: @ScaleEngineSOP.