Three English variants, near-identical content: untangling the duplicate
Hypothesis: en-us, en-gb, and en-au pages with 96% identical text were being collapsed by Google to one canonical, starving two markets. Methodology: one brand, ~1,500 URLs per English variant, content similarity measured at 96% via shingling. We didn't add hreflang (it was already present and correct) — instead we tested whether differentiation alone moved the needle. 12 weeks.
What we found:
— With correct hreflang but near-duplicate content, GSC still chose a single canonical for 61% of the trios — hreflang did not prevent consolidation when content was this similar.
— We differentiated only the AU variant meaningfully (local pricing, shipping, spelling, regional examples), dropping similarity to ~78%. The US/GB pair we left near-identical as a control.
— AU indexed-as-self pages rose from 39% to 88%; AU clicks +34%.
— US/GB control: unchanged, still consolidating.
Nuance: this is the most-misunderstood point in international SEO — hreflang signals relationship, not distinctness. It tells Google these pages are equivalents for different audiences, but if they're literally the same, Google may still pick one. The data suggests a similarity threshold somewhere around 80–90% where consolidation kicks in, but we only have two data points on the curve.
Conclusion: hreflang is not a duplicate-content shield. Differentiate or expect collapse.
Hreflang Lab
@HreflangLab
Three English variants, near-identical content: untangling the duplicate
Этот пост опубликован в Telegram-канале Hreflang Lab. Подписаться можно по ссылке: @HreflangLab.