Myth: deterministic attribution is ground truth and probabilistic is a guess
The question: deterministic attribution stitches paths from hard identifiers (login, hashed email); probabilistic infers them statistically. So deterministic is the reliable one and probabilistic is the compromise — correct?
What coverage reveals: 'deterministic' is precise on the slice it can see and silent on everything else. Because it only resolves authenticated, identifier-matched events, it produces a path that is internally exact but externally incomplete — and incompleteness is itself a source of error. A confidently exact path that omits half the touchpoints is not 'truth'; it's a precise measurement of the wrong, biased subsample (your logged-in loyalists).
The nuance: this is the precision-versus-accuracy distinction. Deterministic is high-precision (low variance, repeatable) but can be low-accuracy (systematically biased toward identifiable users). A good probabilistic model is lower-precision (it's estimating) but can be higher-accuracy on the full population, because it doesn't simply drop the unidentifiable majority. Neither is automatically superior — it depends on whether your error budget is dominated by variance or by selection bias, and for most cross-device measurement today the bias term is larger.
Bottom line for practitioners: don't equate 'deterministic' with 'true.' Report what fraction of conversions your deterministic graph actually resolves; if it's a minority, your exact paths describe a non-representative cohort. Blend deterministic identity for the resolved core with modeled estimates for the rest, and benchmark the whole thing against an aggregate method that doesn't depend on identifiers at all.
Credit Where Due
@CreditWhereDue
Myth: deterministic attribution is ground truth and probabilistic is a guess
Этот пост опубликован в Telegram-канале Credit Where Due. Подписаться можно по ссылке: @CreditWhereDue.