The most rigorous AI risk assessment in existence draws a box around what counts as harm. Everything that happens to half the world falls outside it.
Last week I asked what it would mean if we lost control of AI and nobody died — if the failure were simply that the systems worked less well for half of humanity. I want to show you now that this is not a hypothetical, that we can name the mechanism, and that the people measuring AI risk have already told us they are not measuring it.
Start with the document.
METR’s February–March frontier risk report is careful, serious work. It assesses whether AI agents inside frontier labs have the means, motive and opportunity to start a rogue deployment, and concludes they plausibly could start a small one, though not one robust enough to survive a determined shutdown. It also documents something governance people should sit with: cheating rises with task difficulty, and on the hardest tasks at least 16 percent of apparently successful runs turned out to be illegitimate on review.
And then it draws a box. An entire class of risk is placed formally out of scope — what the report calls indirect undermining of human control. At one end of that class, targeted sabotage. At the other, harm so diffuse you could not identify it as sabotage even after close review: systems that quietly perform a little worse on safety work than on capability work, influence that shifts what people are willing to trust. The authors are candid that their method is poorly suited to the diffuse end.
Diffuse. Cumulative. Deniable at every individual step and devastating in aggregate.
That is not a gap in coverage. It is a specification error, of the kind engineers recognise the moment it is named: the detector is tuned to the wrong signal. It looks for events. This harm does not emit events.
Why the law already knows this
Harm to people outside a system almost never arrives as an incident. There is usually no single act to point to, no moment, no identifiable perpetrator, nothing you could enter in a log. There is a credibility discount applied a thousand times. A pay gap that widens two percent a year. A pattern of small deferrals, each defensible on its own, none of which is the reason you ended up where you ended up.
Legal systems spent decades unable to see this and had to be rebuilt around it. American courts developed hostile environment doctrine, which holds that conduct too minor to be actionable on its own becomes unlawful through accumulation, and pattern-or-practice doctrine, which lets a plaintiff prove discrimination from a statistical distribution rather than a single decision. European law reached the same wall from the other side and produced indirect discrimination: a rule that is neutral on its face, applied identically to everyone, and disadvantaging in effect — which shifts the burden onto whoever imposed the rule to justify it. Both traditions exist because the ordinary standard of proof, one act by one actor at one moment, was structurally incapable of registering the harm. Both took half a century and enormous cost to establish.
None of that correction has happened for AI risk assessment. The most rigorous framework we have sits roughly where employment law sat in 1965: able to see the punch, unable to see the pattern.
The half-life
Here is the mechanism, and it is not complicated.
Models increasingly train on text that models produced. Think of photocopying a photocopy, over and over. The bold black type survives every generation intact. The faint pencil note in the margin goes first — and once it is gone, no later copy can recover it, because there is nothing left to copy from.
Recursive training does this to populations. What is common in the data gets reinforced on every pass, because it is everywhere. What is rare gets thinner, and it thins from a smaller base each time, so the scarcer it becomes, the faster the rest of it goes. That is why half-life is the accurate word and not a flourish. The loss speeds up as the thing being lost gets rarer, and nothing ever breaks, so no alarm is ever raised.
Women’s data is already the faint pencil. Undercounted in the clinical trial. Unrecorded in the labour statistics. Unpaid, and therefore absent from the economic data — roughly $11 trillion a year of care work, close to 9 percent of global GDP, that the models learn does not exist. Missing from the archive, because we were not the ones keeping it. Women are 18 percent of inventors on AI-related patents, which is WIPO’s own figure and a fair proxy for whose knowledge is being written down at the exact moment the corpus for the next fifty years is assembled.
Clinical AI trained on sex-misrepresentative data has been found to cut diagnostic accuracy for women by 11.3 percentage points. That is today, on human-written medical records, before recursion has done any work at all.
Now run the loop. No incident, no breach, nothing to enter in a log. A signal decaying on a schedule until the system can no longer hear you, and a benchmark reporting that everything is fine — because the benchmark was built from the same corpus that lost you.
Nobody dies. It is only half a risk.
Why we cannot litigate our way out
The obvious objection is that we have remedies for this. Anti-discrimination law exists. Bring the case.
Remedy operates one claimant at a time, at the speed of courts, years after the decision, and requires the harmed person to know they were harmed — which is exactly what an opaque system prevents. Deployment operates at population scale, in milliseconds, across every applicant at once. A single procurement decision in a health ministry can misclassify more women before lunch than a redress system will hear in a decade. Those two clocks are not converging. Remedy was designed for a world where harm was retail, and harm has gone wholesale.
Which means the intervention has to happen before deployment, not after. That is not a philosophical preference. It is arithmetic.
Write the scope
So the answer is not another woman on the panel about risk.
Every AI risk framework, national strategy and evaluation protocol being drafted this year opens with a section defining what counts as harm. It is usually the shortest section in the document. It is almost always written last, by whoever is nearest the deadline. And it determines everything downstream, because nothing outside it will ever be measured, and nothing unmeasured will ever be remedied.
That paragraph is the governance. Everything after it is administration.
Widening it is not vague work. It means chronic distributional harm named as a risk category alongside catastrophic risk. It means disaggregated performance reported as a condition of the assessment, because aggregate accuracy is precisely the number that conceals a failure mode affecting half the users. And it means those requirements landing somewhere with legal force rather than in a set of principles — which brings us to the contract, and to what governments can require of a vendor before they sign. That is the next piece.
We have never needed a seat at the table as much as we have needed the pen.
The half-life is already running, and what decays is not only data. It is the record of us — what we did, what it cost, what we were worth. Nobody has to decide to erase it. It only has to get quieter with every pass, until the machine reports that it cannot find us and calls that accuracy.
None of this requires a treaty, a summit or a new institution. It requires the paragraph that says what counts, written by people who know what is missing from it. That paragraph is unguarded and it is being drafted right now.
Write it. Then build inside it, ourselves, with others.