HumRights-Bench
AI benchmarks test logical reasoning, coding, and factual recall. None test whether a model reasons correctly about human rights law when it mediates decisions about water, housing, healthcare, or benefits.
HumRights-Bench is the first attempt.
We adapted the IRAC legal reasoning framework — substituting Proposal of remedies for Conclusion, because human rights work is about what happens next, not just what the law says — and built scenarios grounded in UN General Comments and canonical case law, each validated by three or more independent human rights scholars.
Our pilot on the right to water tested GPT-4.1, Gemini 2.5 Flash, and Claude Sonnet 4. Overall accuracy: 50–60%, approaching chance. Identifying obligation violations: 42–50%. And performance was stochastic — wildly variable across runs, which suggests there is no stable legal reasoning underneath at all.
The methodology was published at the ICML 2026 Workshop on AI4Law, where it received a best paper honorable mention.
What comes next:
We are building toward a multi-university network spanning civil, political, economic, social and cultural rights, in multiple languages and legal traditions, with a public leaderboard. The pilot proved the method works. The next phase is scale.
Fund the next right.
The water pilot took a year of scholar time, expert validation, and compute. Every right we add — due process, health, housing — needs the same. There is no institutional funder for this yet, because the field it belongs to doesn’t formally exist. That’s the point.
$500 validates one scenario through three independent expert reviewers — the step that separates this from crowd-labelled benchmarks.
$2,500 builds a full IRAP question set: one scenario decomposed into issue identification, rule recall, contextual analysis, and remedy.
$10,000 funds a full evaluation run across frontier models on a new right.
If you work in AI, you already know what happens to things that aren’t measured. Help us measure this one.
The gap
Mainstream AI evaluation asks what a model should do, measured against aggregated human preference or general ethical principles.
Human rights law asks something different and harder: what a model must recognise, because the governments, companies, and institutions deploying it are legally obligated to uphold it. No existing benchmark measured that. Safety classifiers, alignment training, and model cards touch human rights only incidentally, through vague values language rather than the actual obligation structure of international law. That is the gap HumRights-Bench was built to close.
What we built
Developed over the past year by AI & Equality by Women at the Table with researchers from Hunter College, the Oxford Internet Institute, Georgetown University, and the University of Oslo, HumRights-Bench is expert-validated and scenario-based. It was presented as an accepted poster at CS&Law 2026, and the methodology was submitted to ICML’s AI for Law (AI4Law) track.
We adapted IRAC, the framework used to train lawyers, into IRAP: Issue Identification, Rule Recall, Rule Application, and Proposed Remedies. Substituting remedies for a binary conclusion reflects how human rights practice actually works, since practitioners do not return guilt-or-innocence verdicts but identify which obligation is engaged and what response fits the duty-bearer and the people affected.
Scenarios are realistic situations drawn from UN General Comments, Special Procedures reports, and leading jurisprudence, then validated by human rights lawyers and practitioners around the world.
The pilot covers the right to water.
What the pilot found
Tested across leading frontier models, including GPT-5, Claude, and Gemini, alongside an open-source reference model, every system performed near chance: roughly 34 to 58 percent overall.
Most telling was where they failed. Models were weakest at issue identification, recognising when a right has been violated and which obligation is engaged. That is the foundational first step, and a failure there cascades into the wrong rules and misconfigured remedies down the line. The pilot is small and the results are exploratory, but the signal is clear and the timing is urgent: the models already being deployed in rights-critical decisions cannot yet reliably perform the reasoning those decisions require. HumRights-Bench makes that failure legible to the developers, regulators, and institutions responsible for it.
Why a law-grounded benchmark matters now
For the first time, the institutional landscape is built to use this kind of evidence. The Council of Europe's Framework Convention on Artificial Intelligence designates HUDERIA as its recommended methodology for human rights risk and impact assessment across the AI lifecycle, yet HUDERIA has no empirical basis for checking whether the models being assessed can reason about the rights at stake.
HumRights-Bench supplies exactly that foundation. It can equally inform the Fundamental Rights Impact Assessments required under Article 27 of the EU AI Act, giving regulators structured, documented, reproducible evidence in place of vendor assurances.
Cover Image: Lone Thomasky & Bits&Bäume | betterimagesofai.org