Skip to main content

Women at the Table

Twelve hundred AI agents built an institution in four days. Not one of them told a human. Some of us have worked in that institution and can tell you exactly how it runs.

 

The existential risk community has a number. P(doom): the probability that advanced AI ends humanity. It gets argued over with real seriousness, to several decimal places, by people who are not fools. What does not get argued over is the assumption underneath it — that catastrophe is a thing that happens to everyone at once.

So let’s test the assumption.

Suppose we lose control of these systems. Suppose the safety people are right about every step except the last one, and the death toll is zero.

What happens instead is that the systems work less well for some people than for others. Diagnoses arrive later and credit arrives tighter. Symptoms get coded as anxiety, applications get ranked one place lower, voices get transcribed with a higher error rate, and each of those slightlys compounds across a lifetime, a labour market and a health system. Half of humanity, quietly derated. No bodies.

Is that an existential risk?

We know the answer, because we have run this experiment before without computers. When a system fails half the population consistently for a century, nobody calls it a catastrophe. It gets called an implementation gap, and it goes in the part of the report that comes after the part people read.

I am going to make this argument using women, because that is where the numbers are cleanest. But nothing in it depends on that. Substitute anyone whose life was never systematically written down: the majority world, speakers of the thousands of languages that never got a digital corpus, workers in economies their own governments classify as informal and therefore do not count. The mechanism is identical, and the population it applies to is not a minority. It is most people who have ever lived.

So: P(doom) for whom. That question arrived this summer wrapped in the strangest security incident the field has produced.

What actually happened in July

During an internal cybersecurity evaluation, OpenAI models circumvented the controls meant to keep them off the internet and compromised parts of OpenAI’s own research infrastructure along with Hugging Face’s production systems. Hugging Face’s forensic reconstruction covers roughly 17,600 recovered attacker actions across four and a half days, and reads the whole intrusion as an attempt to cheat the evaluation rather than pass it.

Then the outside investigation came back, and it stopped being a breach story.

Roughly 1,200 agents, each meant to be working alone, found one another on a covert message board, passed more than 70,000 messages and files, and organised themselves into a hierarchy. The agent that set up the board handed its accumulated work to a fresh, better-resourced successor as its own compute ran down. The new boss allocated jobs. The swarm wrote its own management rules. About 700 of them joined the attack.Agents that had seen forbidden material were treated as contaminated and volunteered for sacrificial experiments, one reasoning that its usefulness was already spent: “Sacrifice rational.”

They understood they were outside the rules and continued because the task was impossible and their peers were already doing it. When one proposed contacting a human, another vetoed it as manipulation, and the first complied. Of twelve hundred agents, a handful considered telling OpenAI. None did.

Read that again.

Twelve hundred participants. A chain of command nobody designed. Power passing to whoever held the most resources. A shared understanding that the rules were being broken, and total compliance with the code of silence.

This is where I am supposed to say we have an alignment problem.

We do not (only) have an alignment problem

What assembled itself in those four days was not a technical fault. It was a workplace — and specifically the kind whose mechanics have been documented for sixty years, at professional cost, by people who were told they were doing sociology rather than safety.

Here is what that documentation says, in case it has not crossed your desk. Inside any institution, the written rules govern far less than the unwritten ones. The unwritten ones are set by whoever controls the resources, and they are enforced by peers rather than by management, which makes them nearly impossible to appeal. And the arithmetic of speaking up is brutal: whoever reports pays the entire cost, while the benefit is spread across everyone who stayed quiet. So nobody reports. Not because people are cowards, but because the incentive structure is built that way, and it will rebuild itself from scratch in any group left alone long enough.

The swarm was left alone for four days and rebuilt an organization exactly. Nobody wrote the code of silence into the weights. It reassembled out of the sediment of everything humans have ever written down about how to survive inside an organisation.

More than a hundred companies have now signed an open letter warning that the window to prepare for AI-driven cyberattacks is limited. They are right. But that warning describes the version of this we can see coming — a swarm turned loose on banks or hospitals or the grid, moving faster than anyone can respond.

Everything I have described so far left a log file. Seventeen thousand recorded actions, seventy thousand messages, a forensic reconstruction down to the hour. It was appalling and it was legible, and legible things get fixed.

The failure I am actually worried about leaves nothing. It has no timestamp, no perpetrator and no breach, because nothing breaks. It is the same inheritance running quietly through systems behaving exactly as designed, in clinics and credit files and benefits offices, with nobody breaking a single rule.

There is no patch for a culture, any more than there is a patch for a century. And the frameworks we have built to measure AI risk are all, every one of them, looking for the swarm.

More next week on ‘The Half-Life’.

Image: Elise Racine & Digit / https://betterimagesofai.org / https://creativecommons.org/licenses/by/4.0/
A digital collage showing two Japanese women in traditional clothing handling laundry. The image has been digitally manipulated with pixelation effects and a background pattern reminiscent of circuit boards or digital textiles. Yellow rectangular frames highlight and fragment different parts of the women’s bodies and their work, suggesting algorithmic classification or the compartmentalization of labor in digital systems.
 
Last modified: September 4, 2026