A Protocol for Establishing Your Anchor Baseline
Before diagnosing risk, you need a measured distribution. Most practitioners eyeball it; a defensible audit is reproducible. The procedure:
— Export all referring URLs and their anchor text from one tool, deduplicate by linking domain (not by link — multiple links from one domain inflate counts).
— Classify each anchor into five buckets: branded, naked URL, generic ('click here', 'this site'), partial-match, exact-match. Keep a sixth 'image/empty' bucket; alt-text-derived anchors behave differently and shouldn't pollute the keyword ratio.
— Compute each bucket as a percentage of unique referring domains.
Per a 2016 reverse-engineering of Penguin-era profiles, healthy commercial sites clustered around 50-70% branded-plus-naked, single-digit exact-match. This is correlational, not a target — sites in different niches show different natural distributions, and the dataset predates link-spam algorithms going real-time.
On one hand, the percentage view flags concentration. On the other, it hides velocity: a 4% exact-match ratio acquired in one month reads very differently from the same ratio over five years. Log acquisition dates if available.
Limitation: tool indexes disagree by 20-40% on link counts, so absolute numbers are soft. Use the same export for before/after comparisons rather than trusting any single snapshot.
Open question: should disavowed and nofollow links be included in the denominator when modeling what a classifier 'sees'?
Anchor Theory
@AnchorTheory
A Protocol for Establishing Your Anchor Baseline
Этот пост опубликован в Telegram-канале Anchor Theory. Подписаться можно по ссылке: @AnchorTheory.