AutoMod rules vs. human moderators: where each one fails
Given finite moderator hours, where should automated filtering end and human judgment begin?
What the research shows
Moderation studies across large platforms converge on a split: automated systems excel at high-volume, low-ambiguity violations (spam links, raid floods, slurs) — catching 90%+ — but degrade badly on context-dependent calls (sarcasm, in-group reclaimed language, escalating-but-not-yet-rule-breaking conflict), where false-positive rates climb into double digits.
Why it happens
Automod operates on surface features — regex, rate, keyword. The violations that actually fracture communities are relational and contextual, invisible to pattern-matching. A human reads the thread; a filter reads the message.
The caveat
Most public moderation data comes from platforms far larger than a typical server, and false-positive rates depend heavily on rule-tuning you can't observe. Generalize cautiously.
The design that holds up: automate the floor (raids, spam, obvious slurs) so humans never burn attention there, and reserve human review exclusively for the ambiguous middle. Servers that try to automate the middle accumulate quiet resentment from wrongly-actioned members who simply leave without appealing.
Open question: does over-tuned automod suppress measured violations while increasing unmeasured churn from false positives?
Server Signal
@ServerSignal
AutoMod rules vs. human moderators: where each one fails
Этот пост опубликован в Telegram-канале Server Signal. Подписаться можно по ссылке: @ServerSignal.