The Labour a Model’s Accuracy Score Never Mentions
The labour a model’s accuracy score never mentions.
Somewhere behind most large AI models sits a worker who spent a shift deciding whether a sentence or an image was too violent to leave in a training set. Her name is not on the model card.
A model does not learn to recognise what to refuse on its own!
Someone has to look at the output first, in bulk, and say yes, no, not this one. That work, usually called data labelling or content moderation depending on which stage of the pipeline it sits in, is not a footnote to AI development. It is a structural, ongoing part of it, and a great deal of it is outsourced through digital labour platforms or business-process firms, much of it based in the Global South, paid by the task rather than the hour.
Reporting over the past few years has been fairly consistent on what that work actually involves: reading or watching disturbing material for hours at a stretch, low and often piece-rate pay, close monitoring of speed and accuracy, and little in the way of counselling or psychological support for people whose job is, in effect, to absorb the worst of the internet so a model does not have to.
Workers describe the task itself as genuinely worthwhile: they are the reason a system learns not to reproduce that material back to someone else. The conditions around the task are frequently the opposite of worthwhile.
An outside check that already exists
You do not have to take a single company’s word for how well it treats this workforce. Fairwork, a research project based at the Oxford Internet Institute, rates platforms against five plain conditions:
- fair pay, at least a local minimum or living wage;
- fair conditions, including basic safety and paid leave;
- fair contracts, meaning real, secure terms rather than task-by-task ambiguity;
- fair management, meaning a worker can actually contest a decision made about their pay or their account; and
- fair representation, meaning workers can organise without being quietly deprioritised for it.
Each platform gets scored, in public, against a minimum and a more demanding advanced threshold. It is the closest thing this corner of the AI supply chain has to an outside audit, and it already exists!
The point for anyone procuring or deploying an AI system is not to feel bad about a supply chain three steps removed. It is that “we did not know” stops being much of a defence once a public rating exists! If your organisation has a choice of vendor, this is the kind of question a genuine due-diligence process asks alongside the security review.
The other kind of hidden labour: expertise going in, not coming back out
A second, less discussed group does related work at the opposite end of the skill scale.
Domain experts, in law, medicine, finance, are increasingly paid to review a model’s draft answers and correct them, effectively teaching the system to sound more like a competent practitioner in that field. It is better paid than data labelling, draws on real domain expertise, and by the measures in our companion piece on meaningful work, likely feels more meaningful to do.
It carries a slower, quieter risk that data labelling does not: the expertise being taught to the model is, over time, expertise the model may need fewer humans to supply.
A lawyer who spends a year correcting a model’s contract analysis is not necessarily spending that year sharpening her own contract analysis; she may be doing the opposite, watching the system get better at a task she is doing less of herself.
Skills erode when they stop being exercised, machine-assisted or otherwise.
The honest fix is not to avoid this work. It is to make sure the people doing it still get real, regular opportunities to exercise the underlying judgement directly, not only to mark someone else’s draft of it, or the expertise quietly changes hands over a few years without anyone deciding that should happen…
What this means for the people who use the model, not just the people who built it
Nobody reading this is likely to be setting piece rates for a labelling platform. But most organisations doing AI work of any kind are, at some point, a customer of one, directly or three vendors removed. Asking what a supplier’s Fairwork score is, or whether one exists at all, costs nothing and takes five minutes!
Treating the answer as part of a genuine procurement decision, is the difference between due diligence and its appearance. Our org self-audit asks a version of this same question about your own organisation’s practices; it is worth asking the same one, plainly, of whoever supplies your models.
Written by us at Ethics Directive, drawing on Fairwork’s published ratings and principles (Oxford Internet Institute) and the broader body of public reporting on AI data-labelling and content-moderation labour conditions over the past several years. Summarised and applied in our own words, not reproduced from any single source. If anything here needs correcting, we will say so in the open, dated.
Does your organisation actually back this work?
Real mandate, escalation, and evidence, or governance on paper. Five minutes, private, and not a certification.