An AI auto-mod rollout and the precision number nobody quotes
A 2025 case from a 60,000-member Discord deployed AI-based auto-moderation (semantic, not keyword) on its busiest channels and published both the catch rate and — unusually — the false-positive rate over a 10-week measured run.
What the data shows
— Auto-removed messages rose to ~120/day, catching an estimated 70% of policy-violating content before any human saw it.
— False-positive rate: ~6% of removals were later overturned on appeal — legitimate messages wrongly axed.
— Human-mod workload on those channels fell roughly 45%.
Why it happens
Semantic models catch obfuscated and context-dependent violations keyword filters miss, which is why the catch rate jumped. But the 6% false-positive rate is the number most case studies omit — and it matters, because wrongful removals erode trust faster than the occasional missed violation.
The caveat
The 70% catch rate is an internal estimate against an unknown true denominator — you can't measure what no one flagged. Appeal-based false-positive rates undercount, since most wrongly-moderated users don't appeal, they just feel silenced and leave. Threshold tuning trades the two error types, and this server's settings won't transfer.
Open question: in moderation, is a 6% false-positive rate the acceptable price of a 45% workload cut — or the hidden tax that quietly churns your most-engaged members?
Server Signal
@ServerSignal
An AI auto-mod rollout and the precision number nobody quotes
Этот пост опубликован в Telegram-канале Server Signal. Подписаться можно по ссылке: @ServerSignal.