"Just canonical the duplicate pages so they don't compete."
Depends what you actually want — and people reach for canonical when they need a different tool.
Three instruments, three jobs. rel=canonical is a hint that says "these are the same page, credit this one" — Google can and does ignore it. noindex is a directive that says "keep this out of the index" — obeyed, but the page still gets crawled. robots.txt Disallow says "don't crawl this" — but a disallowed URL can still get indexed if it's linked.
The classic blunder: noindex a page AND block it in robots.txt. Now Google can't crawl it, so it never sees the noindex, so the URL lingers in the index forever. The two cancel out.
So: near-duplicates that should consolidate equity — canonical. Pages that must vanish from results — noindex, and let them be crawled. Crawl-budget waste you never want fetched — robots.txt, and don't expect it to deindex anything.
(Canonical to a noindexed page is the other trap. You're pointing equity at a door you nailed shut.)
Myth Off Page
@MythOffPage
"Just canonical the duplicate pages so they don't compete."
Этот пост опубликован в Telegram-канале Myth Off Page. Подписаться можно по ссылке: @MythOffPage.