Budget Myths
Budget Myths
@CrawlBudgetMyths

They noindexed, then blocked in robots.txt. 80K pages stayed indexed for a year.

They noindexed, then blocked in robots.txt. 80K pages stayed indexed for a year.

Great case study in how two "fixes" cancel out. A site wanted 80,000 thin pages out of the index, so they added noindex — then, impatient, also blocked the path in robots.txt to "save crawl."

Wrong order, fatal combo. Robots.txt blocks crawling, so Googlebot could no longer fetch the pages to see the noindex. The pages stayed indexed, frozen, for over a year — showing as "indexed though blocked" in Search Console.

The fix: remove the robots.txt block, let Google recrawl and process the noindex, and only after the pages dropped did blocking become safe.

No — you can't noindex what you won't let Google crawl.

Process the noindex first. Block second, or never. Sequence matters more than speed.
Этот пост опубликован в Telegram-канале Budget Myths. Подписаться можно по ссылке: @CrawlBudgetMyths.
tech

Свежие посты в категории «Tech Infrastructure»

Все каналы категории →

start

Готовы запустить рекламу через сеть public.tg?

Новый оффер, продукт, GEO, кейс, событие или партнёрский запуск — соберём маршрут под задачу и отдадим медиаплан.

Telegram для медиаплана: @AFFtop_connect. Быстрый тест: $20 за канал, $1000 за пакет по сети.