In The World
Somebody wrote down the answers first, and that turned out to be a global industry
Systems that learn from examples need examples marked up by people, the work was organised into piecework distributed across continents, and the conditions of that work are visible in the finished models.
By Manish Trivedi3 min read

Supervision means a person supplied the answer
A large part of machine learning is supervised, which means every training example came with the correct output attached. A photograph paired with the word cat. A transcript paired with a recording. A support ticket paired with the department that should handle it. None of those pairings occur naturally.
Some labels are collected as a by-product of something people were doing anyway: a click, a purchase, a correction typed into a search box. Those are cheap and plentiful, and they measure behaviour rather than truth, which is a limitation that follows the resulting model everywhere it goes.
The rest has to be produced deliberately, by somebody sitting down and deciding. That work is not incidental to the technology. It is a substantial input, comparable in importance to the architecture, and it is described far less often because it is neither novel nor flattering.
The work was industrialised into very small pieces
The dominant arrangement breaks labelling into tasks lasting seconds and distributes them through online platforms and outsourcing firms to workers who may be anywhere. Payment is usually per task rather than per hour, quality is policed by inserting items whose answer is already known, and the pace is set by whoever is buying.
The economics push the work towards places where the rate on offer is worth taking, which has concentrated it in particular regions and made the composition of the workforce quite different from the composition of the teams commissioning it. That difference matters, because labelling is judgement and judgement is situated.
Journalistic and academic accounts of these arrangements have described low pay, unpredictable availability of work and limited recourse when a submission is rejected. The picture is not uniform — some annotation is done by salaried specialists on long contracts — but the piecework end of it is large and it is where most volume comes from.
The guidelines are where the real decisions live
Annotators don’t label from intuition. They work from a document specifying what counts as what, and that document is where a contested question gets resolved into a rule. Is a drawing of a dog a dog. Is sarcasm negative sentiment. Is this comment harassment or robust disagreement. Somebody decided, in writing.
Those documents grow long precisely because the edge cases are endless, and they are revised mid-project, which means a dataset can contain items labelled under two different regimes. The model then learns the average of two incompatible standards and nobody downstream knows the seam is there.
Where annotators disagree with each other, the disagreement is usually resolved by majority or by a senior reviewer, and the fact of disagreement is discarded. That is a real loss. A case that three careful people saw differently is genuinely ambiguous, and flattening it into one label teaches the model a certainty the world doesn’t contain.
Preference and safety work changed the character of the job
Turning a raw predictive model into something usable involves people comparing outputs and marking which is better, and reviewing material to decide what a system should refuse. That second category means reading the worst text and viewing the worst images that the internet contains, repeatedly, as a job.
The parallel with content moderation is exact, and the harms documented in reporting and research on moderation apply here too. Some organisations provide counselling and rotation; provision is inconsistent across the industry and difficult to verify from outside, since the work is largely subcontracted.
It is worth stating plainly what this means. A system that declines to help with something harmful behaves that way because people were shown harmful material and marked it. The refusal is not a property of the mathematics. It is the residue of somebody’s afternoon.
Where the work goes next is genuinely unsettled
One expectation is that models will increasingly label their own training data, with humans supervising the supervision. This already happens at scale, and it saves a great deal of money. The obvious risk is that errors and biases circulate rather than being corrected, since the judge and the student share an ancestry.
The opposing expectation is that demand for human labelling rises rather than falls, but shifts upmarket: fewer people marking whether an image contains a bicycle, more people with domain qualifications adjudicating difficult cases in medicine, law or code. There is visible movement in that direction, and it doesn’t obviously replace the volume work.
Both trends may run at once, which would be an uncomfortable outcome — a shrinking floor and a growing ceiling in the same labour market. Anyone predicting which dominates is guessing. What is not a guess is that the label came from somewhere, and the somewhere has people in it.
Common questions
Do the biggest modern models still need labelled data?
The bulk of their training uses raw text with no labels, since the objective is simply to predict what comes next. But the stages that make a model usable — preference comparisons, safety judgements, evaluation sets — are all human-labelled, and those stages shape the behaviour people actually encounter.
How do errors in labelling show up later?
As confident mistakes concentrated in whatever the guidelines handled badly. If the instructions were ambiguous about a category, the model learns the ambiguity and behaves inconsistently at exactly that boundary. Because the fault is in the data rather than the code, it survives retraining unless someone goes back to the guidelines.
Can a model be trained without any human judgement at all?
Not meaningfully. Even self-supervised training rests on a corpus that people selected and filtered, and evaluation requires some standard of correctness that ultimately traces to human judgement. The judgement can be pushed further back in the pipeline. It cannot be removed from it.
Consumer editor, AI Worth Knowing
Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.





