Budget Myths
Budget Myths
@CrawlBudgetMyths

Blocking crawl waste with robots.txt — the careful version

Blocking crawl waste with robots.txt — the careful version

robots.txt is the only tool that actually stops a fetch before it happens. But people wreck their indexing with it. Do this:

— List the exact URL patterns wasting crawl (from your logs).
— Confirm those URLs have no indexable value AND aren't linked-to for ranking.
— Write Disallow rules for the patterns, not blanket folders.
— Test every rule in the robots.txt Tester / a parser. One bad wildcard can disallow your whole site.
— Check you're NOT blocking URLs that are already indexed — blocked pages can't be removed or recanonicalized, they just freeze.
— Never block CSS/JS Google needs to render.

No: robots.txt does not deindex. Google's own docs say blocked URLs can still appear in results. Block to save crawl, not to remove pages.

Disallow stops crawling. It does not stop indexing. Don't confuse the two.
Этот пост опубликован в Telegram-канале Budget Myths. Подписаться можно по ссылке: @CrawlBudgetMyths.
tech

Свежие посты в категории «Tech Infrastructure»

Все каналы категории →

start

Готовы запустить рекламу через сеть public.tg?

Новый оффер, продукт, GEO, кейс, событие или партнёрский запуск — соберём маршрут под задачу и отдадим медиаплан.

Telegram для медиаплана: @AFFtop_connect. Быстрый тест: $20 за канал, $1000 за пакет по сети.