Skip to content

Trust, Safety & Anomaly

Three case studies where the label is scarce, delayed, or actively fought over, and where the cost of a false negative differs sharply from the cost of a false positive. Fraud detection deals with censored labels and extreme imbalance. Content moderation deals with contested ground truth and the trap of measuring flagged precision instead of true prevalence. Spam and bot detection deals with adversarially polluted labels, where the signal is coordination rather than any single account.

The common move across all three is to stop trusting raw accuracy and reason in terms of cost weighted errors against an adversary who adapts.