In The World
Moderating a platform is a judgement problem that automation can only partly take on
Deciding what may stay up is done at a scale that forbids reading everything and requires context that no classifier holds, and the arrangement that resulted places people at the worst point in the pipeline.
By Daniel Okonkwo4 min read

The volume forbids the obvious approach
A platform receiving an enormous quantity of material every day cannot have a person look at all of it, and no amount of hiring changes that arithmetic. Automated classification is therefore not a cost-saving measure layered onto a working manual process; it is the only way the process exists at all, and everything else is arranged around it.
What the classifiers do well is the unambiguous end. Material that has been identified before can be matched against a stored fingerprint and removed on upload, reliably and at any volume. Categories with consistent visual or textual signatures are also handled acceptably, which covers a substantial share of the total.
What remains is the part that was always going to be hard, and it does not shrink in proportion to the improvements. Volume growth means that even a small residual fraction is an enormous absolute number of decisions requiring judgement.
The hard cases are hard because meaning depends on context
A great many rules cannot be applied by looking at the item in isolation. The same image is documentation, advocacy or celebration depending on who posted it and why. The same phrase is an insult, a reclaimed term or a quotation depending on speaker and audience. Satire and sincerity are frequently indistinguishable from the text alone.
A classifier sees the item and, at best, a thin surrounding of metadata. It does not know the history between two accounts, the local political situation, or the fact that a phrase acquired a new meaning last week. Speed is exactly the property that prevents it from acquiring any of that.
This is why the residual category is not merely the difficult tail of the same distribution. It is a qualitatively different problem, and the belief that better classifiers will eventually absorb it assumes that context can be inferred from the artefact, which is often false.
The guidelines are where the real decisions are made
Public policies are short and general. The documents reviewers actually work from are long, specific and full of worked examples, because a general rule cannot be applied consistently by thousands of people without being decomposed into cases. Those internal documents are the operative law of a platform in a way the published policy is not.
Writing them is a genuinely difficult editorial and ethical task, and the trade-offs are visible in the result. Rules precise enough to be applied uniformly are rules that will be obviously wrong in some cases; rules loose enough to accommodate judgement produce inconsistency that looks like bias from outside. There is no version without a cost.
The same documents then become training material, since the labels used to train classifiers come from decisions made under those guidelines. Every ambiguity in the written rule is inherited by the automated system, which will apply it faster and to more people.
The human part of the job is arranged badly and known to be
The work has been organised as high-volume piecework with short handling times and accuracy targets measured against other reviewers rather than against any external standard. Reviewers see concentrated quantities of exactly the material the classifiers could not confidently remove, which is by construction the most disturbing and ambiguous portion.
The psychological cost of this has been documented and acknowledged, and support arrangements vary widely between the organisations doing the work, much of which is contracted out across several countries. Improvements have been announced repeatedly over the years, and independent verification of them is limited by the same contractual distance that produced the problem.
Better automation genuinely reduces exposure by filtering the clearest cases before a person sees them. It also concentrates what remains, so the average item a reviewer handles gets worse as the system gets better. Both things are true at once.
Appeals, errors and where the disagreement sits
Any system operating at this scale makes errors in both directions, and the two are not symmetrical in how they are experienced. A wrongly removed post is visible to its author and generates a complaint; wrongly retained material is visible to everybody else and generates a different sort of complaint. Tuning towards one is tuning away from the other.
Appeal processes exist and are themselves partly automated, which introduces the familiar problem of a review conducted by something resembling what made the original decision. Regulation in several jurisdictions has begun requiring meaningful human review of certain decisions, and how meaningful that requirement proves in practice is not yet clear.
Reasonable people disagree about the direction. One view holds that more automation is the only humane answer, because it removes people from the worst exposure. Another holds that it entrenches unaccountable decisions about speech behind a technical process nobody outside can inspect. Both concerns are well founded and they do not resolve each other.
Common questions
Why do platforms not simply employ more reviewers?
Cost is part of it and not the whole. The work is difficult to staff, difficult to retain people in, and difficult to make consistent across a large workforce. Even a very large team cannot review a meaningful fraction of the total, so the automated first pass remains structural rather than optional.
Are language models replacing classifiers for this?
They are being used for some of it, and they bring a better handle on context along with higher cost per item and their own unpredictability. The consensus is that they shift the boundary of what can be automated rather than removing the boundary.
Does automated moderation work equally well in all languages?
No, and the disparity is large. Training data, evaluation data and reviewer availability are all concentrated in a handful of widely spoken languages, and enforcement quality follows that concentration closely.
Contributing editor, AI Worth Knowing
Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.





