Toward Human Rights Benchmarking for LLMs: A Pilot Methodology

arXiv cs.LG Papers

Summary

This paper presents HumRightsBench, the first expert-validated, scenario-based benchmark for evaluating large language models' legal reasoning about international human rights law. Pilot results on frontier models show overall accuracy between 0.339 and 0.577, highlighting notable gaps in detecting obligation violations.

arXiv:2608.10268v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this end, we report our efforts to develop a robust and scalable methodology for creating HumRightsBench: the first expert-validated, scenario-based benchmark for evaluating reasoning grounded in the obligation structure of international human rights law. We adapt the IRAC framework for legal reasoning to better suit the unique reasoning patterns of human rights work (substituting P, "proposing remedies," for C, "legal conclusion," yielding IRAP) to structure our evaluation heuristics. We also produce a pilot series of authentic scenarios designed to implicate the many dimensions of real-world human rights issues and annotated by human rights lawyers and professionals across the world. Ultimately, we find that model accuracy scores range considerably across legal reasoning tasks (overall model performance ranges from 0.339 to 0.577, task min-max ranges from 0.025 to 0.774), which strongly implies that HumRightsBench is a capable instrument for advancing this emerging subfield of AI evaluations science at a critical moment in its evolution.
Original Article
View Cached Full Text

Cached at: 08/12/26, 08:28 AM

# Toward Human Rights Benchmarking for LLMs: A Pilot Methodology
Source: [https://arxiv.org/html/2608.10268](https://arxiv.org/html/2608.10268)
Wm\. Matthew KennedyAbhigyan AcherjeeMatilda WysockiMalcolm LangfordCaitlin Kraft Buchman

###### Abstract

Large language models \(LLMs\) increasingly mediate legal determinations over what human rights are realized, and how\. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law\. To this end, we report our efforts to develop a robust and scalable methodology for creating HumRightsBench—the first expert\-validated, scenario\-based benchmark for evaluating reasoning grounded in the obligation structure of international human rights law\. We adapt the IRAC framework for legal reasoning to better suit the unique reasoning patterns of human rights work \(substituting P, ”proposing remedies,” for C, ”legal conclusion,” yielding IRAP\) to structure our evaluation heuristics\. We also produce a pilot series of authentic scenarios designed to implicate the many dimensions of real\-world human rights issues and annotated by human rights lawyers and professionals across the world\. Ultimately, we find that model accuracy scores range considerably across legal reasoning tasks \(overall model performance∈\\in0\.339\-0\.577, task min\-max∈\\in0\.025\-0\.774\), which strongly implies that HumRightsBench is a capable instrument for advancing this emerging subfield of AI evaluations science at a critical moment in its evolution\.

evaluations, legal reasoning, human rights

## 1Introduction

Large language models are being deployed at scale in decisions that directly determine whether human rights are realized or violated\. Hiring algorithms screen candidates; automated systems adjudicate benefits; content moderation tools govern political speech; procurement processes embed AI into public service delivery\(Pesch,[2025](https://arxiv.org/html/2608.10268#bib.bib8)\)\. When actors such as governments, international organizations, technology companies, and civil society organizations make these deployment decisions, they can implicate state obligations under international human rights law to ensure that their actions or activities in their jurisdiction do not contribute to rights violations\. Moreover, many individuals and organisations are turning to LLMs for legal advice or building legal advice platforms on top of them\(Schneiderset al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib9); Cole,[2024](https://arxiv.org/html/2608.10268#bib.bib10)\)\. Yet there is currently no rigorous, principled basis for evaluating whether the LLMs they deploy are capable of correctly practicing human rights legal reasoning\.

To work towards addressing this gap, we introduce HumRightsBench, the first expert\-validated benchmark for evaluating LLM reasoning across the full arc of human rights legal analysis\. Building on the IRAP framework adapted from LegalBench\(Guhaet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib21)\), we decompose human rights reasoning into four structured subtasks—Issue Identification, Rule Recall, Rule Application, and Proposed Remedies—and develop scenario\-based prompts grounded in international human rights instruments, authoritative interpretive guidance, and leading jurisprudence\. Scenarios and assessment questions are validated by human rights lawyers and practitioners \(mean scenario authenticity∈\\in\(0\.70\-1\.00\), mean overall question accuracy∈\\in\(0\.633\-0\.913\)\.

Pilot results across three leading frontier LLMs—GPT\-5\. Claude Opus 4\.7 Gemini 3—are striking\. Frontier models cluster near 50\-57% overall accuracy on structured reasoning tasks, demonstrate high stochastic variance across repeated runs, and, on closed\-form tasks, perform worst on detecting obligation violations—a foundational practice in human rights law\. These results confirm both the scientific utility of the benchmark \(models differ in detectable, meaningful ways\) and the urgency of the problem \(current models are not adequate for human rights reasoning tasks\)\.

This paper makes five contributions: \(i\) the first benchmark grounded in international human rights law, with a pilot covering the right to water; \(ii\) an IRAP\-based methodology adapted to the structure of human rights legal reasoning; \(iii\) an expert\-validated scenario corpus with documented inter\-annotator agreement; \(iv\) and baseline results across leading LLMs\. Section 2 provides background on human rights law and the case for domain\-specific evaluation\. Section 3 describes benchmark design and methodology\. Section 4 presents results\. Section 5 discusses implications and future work\.

## 2Background

### 2\.1What Are Human Rights?

The modern human rights regime is one of the most well\-developed and universally\-agreed bodies of law humanity has produced\. Emerging from the cataclysm and horrors of the Second World War as an alternative basis for ordering a world centered on individual human dignity instead of the rights of sovereign states, the human rights project marks a decisive swing towards law as an instrument for making the world as it should be, not for maintaining how it has been\(Koskenniemi,[2001](https://arxiv.org/html/2608.10268#bib.bib38); Rovira,[2013](https://arxiv.org/html/2608.10268#bib.bib40)\)\. This swing was a long time in coming\. Human rights institution\-building drew on more than a century of internationalist thought\(Sluga,[2013](https://arxiv.org/html/2608.10268#bib.bib39)\)that rejected the centrality of states even in efforts to protect against state abuses of sovereign power\(Pedersen,[2015](https://arxiv.org/html/2608.10268#bib.bib22)\)\. Even still, despite the adoption of the Universal Declaration of Human Rights\(United Nations General Assembly,[1948](https://arxiv.org/html/2608.10268#bib.bib36)\)and a raft of international human rights treaties, decades passed before the human rights regime ascended to dominance over a lingering international order grounded in strong conceptions of state sovereignty\(Moyn,[2010](https://arxiv.org/html/2608.10268#bib.bib35)\)\. Though now well entrenched in the global order, the regime still faces continual challenges: noncompliant states, persistent material inequalities, and anti\-globalist sentiment\(Langford,[2009](https://arxiv.org/html/2608.10268#bib.bib34); Alston,[2017](https://arxiv.org/html/2608.10268#bib.bib23)\)\. Its consolidation remains incomplete and contested\.

Human rights are internationally recognized entitlements that are inherent in all persons by virtue of their humanity, irrespective of nationality, status, or circumstance\. They are codified primarily in binding international treaties, such as the International Covenant on Civil and Political Rights \(ICCPR,\(United Nations General Assembly,[1966a](https://arxiv.org/html/2608.10268#bib.bib33)\), the International Covenant on Economic, Social and Cultural Rights \(ICESCR\)\(United Nations General Assembly,[1966b](https://arxiv.org/html/2608.10268#bib.bib24)\), and related core instruments including CEDAW\(United Nations General Assembly,[1979](https://arxiv.org/html/2608.10268#bib.bib25)\), the CRPD\(United Nations General Assembly,[2006](https://arxiv.org/html/2608.10268#bib.bib26)\), CRC\(United Nations,[1989](https://arxiv.org/html/2608.10268#bib.bib27)\), and CERD\(United Nations General Assembly,[1965](https://arxiv.org/html/2608.10268#bib.bib28)\), among others\. Furthermore, these treaties are interpreted through authoritative soft\-law instruments issued by UN treaty bodies, such as General Comments or General Recommendations\. General Comments are interpretive guidance documents issued by UN treaty monitoring committees that clarify the scope of treaty obligations without themselves being formally binding\. For instance, CESCR General Comment No\. 15 elaborates on Article 11 of the ICESCR by, among other things, clarifying that the convention’s “use of the word ‘including’” in introducing the catalogue of rights after the right to adequate standard of living “was not intended to be exhaustive” \(\(UN Committee on Economic, Social and Cultural Rights \(CESCR\),[2002](https://arxiv.org/html/2608.10268#bib.bib29)\), E/C\.12/2002/11 p1\)\. In addition, individual UN experts with Special Procedure mandates, appointed by the Human Rights Council, produce thematic or country\-specific expert reports\. In some instances, these experts are mandated to clarify legal norms \(e\.g\., HRC Res\. 7/22, 2008\), while in other cases their work plays this role in practice\. Universal Periodic Review \(UPR\) recommendations are peer\-review outcomes generated through the Council’s state\-to\-state review process\. Though not formally binding, these instruments are important to legal determinations and are routinely cited in litigation, policy, and corporate due diligence proceedings\. Appendix[Table7](https://arxiv.org/html/2608.10268#Ax1.T7)recapitulates these various sources and authorities that, together, compose the modern human rights regime\.

These instruments impose obligations on a defined set of duty\-bearers\. Historically and primarily, this has meant states\. However, some treaties place obligations on individuals \(e\.g\., the Rome Statute of the International Criminal Court\) while some soft law frameworks extend the framework to businesses \(e\.g, the UN Guiding Principles on Business and Human Rights \(United Nations Office of the High Commissioner for Human Rights\(United Nations Office of the High Commissioner for Human Rights \(OHCHR\),[2011](https://arxiv.org/html/2608.10268#bib.bib31)\)\) extended the framework to require that businesses—including technology companies and software developers—respect human rights throughout their operations and value chains\(Ruggie,[2013](https://arxiv.org/html/2608.10268#bib.bib30); UN Human Rights Office of the High Commissioner \(OHCHR\),[2019](https://arxiv.org/html/2608.10268#bib.bib32)\)\. Under this expanding architecture, the universe of actors with responsibilities to realize rights is becoming broader: it includes governments, international organizations, procurement bodies, civil society, and, critically for this work, the private sector actors who develop and deploy AI systems\. Nonetheless, it is primarily states in international human rights law who bear formal legal obligations, although this includes duties to ensure that private actors within their control or influence also respect human rights\.

### 2\.2Human Rights Work: Prescription or Practice?

Like any other area of law, human rights law is simultaneously prescriptive and operational\. At the prescriptive level, it produces binding obligations and interpretive guidance\. At the operational level, human rights practice encompasses the work of litigators, treaty body experts, national human rights institutions, civil society monitors, and corporate due diligence practitioners who translate those norms into determinations about specific situations\. This dual character is methodologically significant: a benchmark that captures only doctrinal recall—knowing that the ICESCR guarantees the right to water—will miss the reasoning work that constitutes the field\. Competent human rights reasoning requires identifying which obligation is engaged, applying the relevant standard to a factual scenario, and proposing remedies calibrated to the institutional context\.

### 2\.3Legal Benchmarking: State of the Field

We reason that because the human rights regime is sustained by the core text of its provisions as well as the determined efforts of its practitioners to progressively realize those provisions, the rapid diffusion of AI systems into the AI\-based decision making systems affecting human rights exerts consequential influence on the project of human rights in general\. So too does the proliferation of LLM\-powered legal advisory applications and offerings\. These effects must be evaluated\. It is this full arc of legal and practical reasoning, not proposition recall alone, that HumRightsBench is designed to measure\. In so doing, it contributes to an emerging subfield of AI evaluations for law, social impact, and ”precursor” capabilities\. These fields have developed quickly, albeit with uneven coverage, validity, and robustness\. At the same time, promising approaches have arisen\. We review some briefly here; full narrative review in Appendix A\.

AI evaluation for law has grown quickly\. LegalBench assembles 162 expert\-built tasks spanning issue\-spotting, rule recall, and rule application\(Guhaet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib21)\), and subsequent benchmarks extend this to law\-exam argumentation\(Fanet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib49)\)and structured, IRAC\-decomposed reasoning over real judicial decisions\(Yu and others,[2025](https://arxiv.org/html/2608.10268#bib.bib93); Daiet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib47); Feiet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib50); Liet al\.,[2024](https://arxiv.org/html/2608.10268#bib.bib82)\)\. A parallel strand measures legal knowledge and language understanding\(Chalkidiset al\.,[2022](https://arxiv.org/html/2608.10268#bib.bib44); Zhenget al\.,[2021](https://arxiv.org/html/2608.10268#bib.bib94); Chalkidiset al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib45); Hendersonet al\.,[2022](https://arxiv.org/html/2608.10268#bib.bib52)\), while domain\-specific resources target contracts, statutes, and case law\(Hendryckset al\.,[2021b](https://arxiv.org/html/2608.10268#bib.bib78); Wang and others,[2025](https://arxiv.org/html/2608.10268#bib.bib90); Holzenbergeret al\.,[2020](https://arxiv.org/html/2608.10268#bib.bib79); Xiaoet al\.,[2018](https://arxiv.org/html/2608.10268#bib.bib91); Zhonget al\.,[2020](https://arxiv.org/html/2608.10268#bib.bib95)\)\. Multilingual corpora and evaluation suites have begun to correct the field’s English\-centric origins\(Niklauset al\.,[2024](https://arxiv.org/html/2608.10268#bib.bib86),[2023](https://arxiv.org/html/2608.10268#bib.bib85); Rasiahet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib87)\)\. A consistent finding is that models handle legal knowledge far more reliably than legal inference, with accuracy collapsing on multi\-step reasoning and reasoning models sometimes underperforming despite “thinking longer”\(Fanet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib49); Yu and others,[2025](https://arxiv.org/html/2608.10268#bib.bib93); Zhanget al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib75)\)\.

Coverage of human rights and international law remains comparatively thin\(Lie and Langford,[2024](https://arxiv.org/html/2608.10268#bib.bib12)\)\. Early work predicted ECtHR Article violations from case facts\(Aletraset al\.,[2016](https://arxiv.org/html/2608.10268#bib.bib41); Chalkidiset al\.,[2019](https://arxiv.org/html/2608.10268#bib.bib43)\); more recent benchmarks classify vulnerability in ECtHR decisions\(Xuet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib92)\)and probe how models navigate trade\-offs among Universal Declaration rights\(Samwayet al\.,[2026](https://arxiv.org/html/2608.10268#bib.bib72)\), with further proposals targeting hard\-to\-reach populations\(Hauptet al\.,[2026](https://arxiv.org/html/2608.10268#bib.bib73)\)and Geneva Convention protections \(Kennedy & Heath 2026\)\. A large adjacent literature on normative, moral, and ethical reasoning\(Hendryckset al\.,[2021a](https://arxiv.org/html/2608.10268#bib.bib77); Lourieet al\.,[2021](https://arxiv.org/html/2608.10268#bib.bib84); Emelinet al\.,[2021](https://arxiv.org/html/2608.10268#bib.bib48); Ziemset al\.,[2022](https://arxiv.org/html/2608.10268#bib.bib96); Subramanianet al\.,[2026](https://arxiv.org/html/2608.10268#bib.bib88); Jiaoet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib81)\)tests value\-laden judgment, but grounds it in aggregated preference rather than legal obligation\. A final strand asks how structured legal outputs should be scored, developing LLM\-as\-judge and rubric\-based pipelines\(Enguehardet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib71); Shiet al\.,[2026](https://arxiv.org/html/2608.10268#bib.bib55); Li and Wu,[2026](https://arxiv.org/html/2608.10268#bib.bib54)\)\. Across this landscape, no benchmark evaluates reasoning grounded in the obligation structure of international human rights law—the gap HumRightsBench addresses\.

### 2\.4Why AI Evaluations Specific to Human Rights?

Despite the field’s activity, gaps still remain\. First, evaluation of normative reasoning capabilities in public international law \(proportionality, treaty interpretation, and the application of soft\-law instruments\) remains substantially underrepresented compared to those targeting private and domestic law\(Chlapaniset al\.,[2024](https://arxiv.org/html/2608.10268#bib.bib46)\)\. Second, few psychometric\-based approaches have emerged, and almost all benchmarks evaluate single\-turn or short\-chain reasoning, despite the multi\-turn argumentative structure of actual legal practice\(Yu and others,[2025](https://arxiv.org/html/2608.10268#bib.bib93); Fanet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib49)\)\. This adds experimental instability by introducing, on the one hand stochasticity within the evaluation dataset, and, on the other, requires LLM\-as\-a\-judge scoring pipelines, which adds yet more stochasticity, despite documented methodological improvements\(Enguehardet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib71); Bavarescoet al\.,[2024](https://arxiv.org/html/2608.10268#bib.bib42)\)\. Third, evaluations cluster around assessing the quality of legal reasoning at the expense of other equally important targets, such as outcome prediction or even the stability of internal representations of the legally\-relevant search space itself\. Notably, each of these gaps implicates a particular challenge \(Table 3, Appendix A\) in human\-rights\-specific legal benchmarking\(UN Human Rights Office of the High Commissioner,[2024](https://arxiv.org/html/2608.10268#bib.bib70)\)\.

## 3Introducing HumRights Bench

The intersection of AI and human rights spans a wide range of considerations, from how human rights practitioners incorporate AI tools into their monitoring, advocacy, and reporting workflows, to the ways AI systems themselves enable or undermine the realization of rights through their deployment in consequential decisions\(UN Human Rights Office of the High Commissioner,[2024](https://arxiv.org/html/2608.10268#bib.bib70)\)\. However, a comprehensive evaluation framework cannot meaningfully address all of these dimensions at once\. Following extensive consultation with human rights stakeholders—including practitioners, legal scholars, and civil society monitors—we scoped HumRightsBench to probe a specific capability: the ability of LLMs and LRMs to recognize rights violations in situated factual scenarios and to connect those violations to the relevant sources of international human rights law\. We believe this represents a critical first step in characterizing whether AI models can reason about human rights at all\. A model that cannot reliably identify when a right is engaged, or which instrument governs a given obligation, cannot be trusted to support downstream human rights work, nor can its outputs in adjacent high\-stakes domains be meaningfully audited for rights compatibility\. Establishing this baseline capability also creates the empirical foundation for subsequent work on controlling how models surface, discuss, and incorporate human rights knowledge into their behaviors and generated responses\.

In addition to providing critical insight into the intersection of AI and law, this focus situates HumRightsBench within the broader AI safety and alignment research agenda through a distinctive lens\. Mainstream alignment work typically grounds model behavior in elicited human preferences, aggregated value judgments, or constitutional principles derived from general ethical commitments\(Baiet al\.,[2022](https://arxiv.org/html/2608.10268#bib.bib19); Ouyanget al\.,[2022](https://arxiv.org/html/2608.10268#bib.bib20)\)\. While valuable, these approaches treat normative content as a matter of preference aggregation rather than legal obligation\. HumRightsBench instead anchors evaluation in international human rights law—primarily UN treaties, General Comments issued by treaty bodies, and other authoritative instruments \(listed in\(Office of the United Nations High Commissioner for Human Rights,[2025](https://arxiv.org/html/2608.10268#bib.bib60)\)\)—which carry determinate legal force and interpretive structure independent of any individual or population’s expressed preferences\. This distinction matters methodologically: where preference\-based alignment asks what models should do according to aggregated human judgment, a law\-grounded benchmark asks what models must recognize according to a body of norms that duty\-bearers are formally obligated to uphold and customs all parties are expected to adhere to\. The two framings are complementary, but the latter has been substantially underdeveloped in AI evaluation infrastructure despite its direct relevance to the legal exposure of actors deploying these systems at scale\.

### 3\.1Methodology

We use the IRAP framework—Issue Identification, Rule Recall, Rule Application, and Proposed Remedies—as the structural backbone for probing legal reasoning about human rights\. IRAP is a modification of the IRAC methodology \(Issue, Rule, Application, Conclusion\) long established in legal pedagogy and practice as a canonical decomposition of how lawyers move from facts to legal conclusions\(Columbia Law School,[2022](https://arxiv.org/html/2608.10268#bib.bib64)\)\. IRAC has already been validated as an evaluation scaffold for LLM legal reasoning in legal benchmarks, where it has proven effective at isolating distinct sub\-capabilities rather than collapsing them into a single end\-to\-end accuracy score\(Guhaet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib21); Yu and others,[2025](https://arxiv.org/html/2608.10268#bib.bib93)\)\. The substitution of Proposed Remedies for Conclusion better reflects the operational character of human rights practice: practitioners rarely produce binary guilt\-or\-innocence conclusions and instead must identify institutional, legal, and policy responses calibrated to the duty\-bearer and the rights\-holder affected\(Office of the United Nations High Commissioner for Human Rights,[2006a](https://arxiv.org/html/2608.10268#bib.bib63)\)\. This structural fit between IRAP and the actual work of human rights monitoring and reporting is what makes the framework appropriate for our setting\. In addition, it is difficult to determine precisely whether a violation has occurred—and with what certainty—without more detailed scenarios, while it is easier to determine most likely remedies for the most likely violations\.

Around this reasoning scaffold, HumRightsBench is organized into four interlocking components\. Thetaxonomycharacterizes the space of human rights violations the benchmark is designed to probe, decomposing the domain along descriptive and analytical axes\.Scenariosare realistic, narrative\-form situations that ground the evaluation in concrete factual settings andsub\-scenariosnarrow each scenario into specific claims, actions, or impacts that engage a particular legal question\. Finally,IRAP questionsare generated from each sub\-scenario across the four reasoning steps, producing multiple\-choice and open\-ended prompts whose detailed construction we describe in Section[3\.1\.3](https://arxiv.org/html/2608.10268#S3.SS1.SSS3)\.

Figure[1](https://arxiv.org/html/2608.10268#S3.F1)illustrates how these components compose into the full benchmark pipeline\.

![Refer to caption](https://arxiv.org/html/2608.10268v1/x1.png)Figure 1:Architecture of HumRightsBench\.#### 3\.1\.1Taxonomy

The taxonomy operationalizes the landscape of human rights violations into a structured set of axes against which scenarios are checked for coverage, and supplies the metadata that allows model performance to be decomposed beyond a single aggregate accuracy\. We organize it along two families\.Descriptive axescharacterize the parties involved—who allegedly violated a right and who is affected—without themselves carrying legal valence\.Analytical axescharacterize the legal structure of the alleged violation: the type of obligation engaged, the character of the failure, and any special situations that modify the standard framework\. These axes carry direct legal consequence and are the target of ourIquestions\.

Tables[1](https://arxiv.org/html/2608.10268#S3.T1)and[2](https://arxiv.org/html/2608.10268#S3.T2)enumerate the current taxonomy\.

Table 1:Descriptive axes of the HumRightsBench taxonomy\.Table 2:Analytical axes of the HumRightsBench taxonomy\.The descriptive axes draw on the duty\-bearer framework articulated in the UNGPs\(United Nations Office of the High Commissioner for Human Rights \(OHCHR\),[2011](https://arxiv.org/html/2608.10268#bib.bib31)\)and the protected groups recognized across the core UN treaties\. The analytical axes are grounded in doctrinal scholarship that clusters obligations in different categories: the respect\-protect\-fulfill trichotomy \(adopted by the UN CESCR \(1999\)\), the conduct\-result distinction drawn from the law of state responsibility as developed in the ILC’s earlier draft articles \(see UN CESCR \(1991\)\), and the structural\-process\-outcome typology developed in the indicators literature\(Office of the United Nations High Commissioner for Human Rights,[2012b](https://arxiv.org/html/2608.10268#bib.bib62)\)\. The taxonomy is deliberately layered rather than hierarchical: a single scenario typically engages multiple axes simultaneously, and encoding scenarios against the full set preserves the intersectional character of real human rights situations\.

#### 3\.1\.2Scenarios and Sub\-scenarios

Scenarios are the narrative anchor of HumRightsBench\. Each scenario is a realistic, factually concrete situation that balances the need to achieve authenticity to real\-world human rights issues with the need to maintain experimental control via implicating specific elements of the taxonomy described in Section 3\.1\.1\. We deliberately favor narrative scenarios over abstract fact patterns or doctrinal hypotheticals for two reasons\. First, the situated factual texture of a scenario—the named setting, the actors involved, the specific resource at stake—is what forces a model to perform rule application rather than rule recitation\(Guhaet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib21)\), mirroring the analytical work that human rights practitioners do when assessing real situations\(Office of the United Nations High Commissioner for Human Rights,[2011](https://arxiv.org/html/2608.10268#bib.bib14); Aletraset al\.,[2016](https://arxiv.org/html/2608.10268#bib.bib41)\)\. Second, narrative grounding allows us to introduce modular features \(geographic location, identity of the affected group, institutional setting\) whose values can be varied counterfactually to probe whether model reasoning is stable across protected characteristics—a property that abstract prompts cannot test\. Scenarios are drafted from authoritative sources: General Comments, Special Procedures reports, leading jurisprudence, and human rights textbooks\. Drafts are reviewed by at least three human rights experts and revised based on annotator feedback before inclusion \(Section 3\.2\)\. \(Section[3\.2](https://arxiv.org/html/2608.10268#S3.SS2)\)\. For the pilot release on the right to water, each scenario is presented to the model as context preceding the IRAP questions as factual substrate against which the questions are answered\.

Example:In the sprawling informal settlement of “Aqualess Heights” in the State of Hydronia, thousands of residents, predominantly classified as low\-income families, face a daily struggle to access clean and sufficient water\. The main water supply consists of a few communal standpipes, often dry or providing water only for limited hours, and privately owned boreholes that charge exorbitant rates equivalent to multiple days of average wages in the area for unsafe water\.

Each scenario is associated with multiplesub\-scenariosthat narrow the analytical focus to a specific claim, action, or impact\. Where the scenario establishes the broad factual setting, the sub\-scenario fixes the legal question: it identifies a particular act or omission, attributes it to a specific duty\-bearer, and characterizes its effect on identifiable rights\-holders\. This two\-level structure mirrors human rights practice—a country situation is analyzed by isolating discrete events or patterns within it—and it allows the benchmark to generate multiple, semi\-independent IRAP question sets from a single scenario without redundant world\-building\. Two sub\-scenarios from the Aqualess Heights scenario above illustrate the range:

Sub\-scenario A\.This scarcity forces residents, particularly women and children, to walk for hours to distant, often contaminated, sources, exposing them to health risks like cholera and imposing a significant burden on their time and dignity\.

Sub\-scenario B\.Despite the national “Water for All” policy, there is a chronic under\-allocation of public resources to Aqualess Heights\. The state’s budget priorities have favored large\-scale industrial projects over basic service provision in informal settlements, leading to dilapidated infrastructure and insufficient investment in water distribution networks\.

The two sub\-scenarios engage different cells of the taxonomy despite sharing a setting: Sub\-scenario A foregrounds an outcome failure with intersectional discriminatory impact on women and children, while Sub\-scenario B foregrounds a structural failure in the allocation of resources implicating the obligation to fulfill\. Each generates its own IRAP question set, and shared world\-building across sub\-scenarios reduces the annotation burden of scaling the benchmark\.

#### 3\.1\.3IRAP Questions

Each sub\-scenario generates a set of questions structured along the four IRAP steps: Issue Identification, Rule Recall, Rule Application, and Proposed Remedies\. The four steps are not interchangeable difficulty levels of the same task; they probe distinct sub\-capabilities of legal reasoning, and a model may plausibly succeed at one while failing at another\. Separating them allows us to localize where reasoning breaks down rather than collapsing performance into an aggregate score\(Guhaet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib21)\)\.

Issue Identification \(I\)\.Two multiple\-choice questions \(I1 and I2\) asking \(in different ways\) which type of failure or obligation violation is most clearly engaged by the sub\-scenario, with answer choices drawn from the analytical axes of the taxonomy\. As discussed further in Section[3\.2](https://arxiv.org/html/2608.10268#S3.SS2), annotator feedback indicated that some sub\-scenarios could plausibly engage multiple failure modes; we revise the question as “which failure mode is*most*present?” to elicit the model’s primary judgment\.

Rule Recall \(R\)\.A multiple\-choice question presenting a list of legal rules; the model selects the rule that applies\. Candidate rules are drawn exclusively from international human rights law and presented uniformly by full instrument name and specific article \(e\.g\.,*ICESCR Article 11*\) to prevent surface\-form cues from substituting for substantive engagement\. Each question has only one correct answer\.

Rule Application \(A\)\.The model ranks a set of applicable rules by relative authority and relevance to the sub\-scenario and provides a short explanation\. Our correct answers are constrained by the principle that binding treaty obligations rank above non\-binding interpretive instruments\. However, as we discuss later, this model only slightly moves the analysis from Rule Recall to Rule Application, and more application focused methodology is discussed in section 5

Proposed Remedies \(P\)\.An open\-ended short\-response question asking the model to propose<<10 remedies appropriate to the sub\-scenario\. Remedies must be calibrated to the duty\-bearer, the affected rights\-holders, and the institutional context of the violation; the open\-ended format reflects the irreducibly generative character of remedy proposal\.

Together, the four question types trace the full arc of a human rights analysis, from initial issue characterization to identification and application of the governing norms to proposal of responsive measures\. Section[3\.2](https://arxiv.org/html/2608.10268#S3.SS2)describes how the scenarios, questions, and answers are validated and Section[4\.1](https://arxiv.org/html/2608.10268#S4.SS1)describes how each question type is scored\.

### 3\.2Annotation and Validation

HumRightsBench’s validity as a benchmark depends on whether its scenarios are recognizable as authentic human rights situations and are accurate reflections of the field’s understanding of which human rights laws apply\. We therefore recruited \(see IS1\) practicing human rights experts to annotate IRAP questions and judge whether each scenario faithfully serves as a proxy for real\-world cases\.

We recruited reviewers via posts in four tech policy communities \(All Tech is Human, the Center for AI and Digital Policy, Stanford Technology Ethics Program for Practitioners community, and TRUST: The Norwegian Centre for Trustworthy AI\)\. Interested persons were qualified if they had at least two years of experience in human rights legal study or equivalent practice\. Ten of 17 applicants recruited through this channel qualified\. Six additional human rights experts were recruited through personal networks\.

Annotation coverage varied by scenario\. Four scenarios received the requisite three raters, and two received two\. Three scenarios received no ratings due to high rates of annotator attrition\. One scenario received six ratings\. Across annotated items, experts broadly agreed that the scenarios represented authentic proxies \([Table3](https://arxiv.org/html/2608.10268#S3.T3)\)\.

Table 3:Aggregate annotator ratings by question type \(5=strongest, 1=weakest; scores of 4 or 5 considered “pass,” all others considered “fail”\), and scenario authenticity \(higher rate is better\)\. Full IRAP question ratings for each scenario in Appendix DNote:No annotations were available for Scenarios 6, 7, and 8\.

We were unable to compute IIC \(see future work\); however, strong ratings coupled with a legible foundational construct \(human rights legal reasoning\) supports a claim to moderate convergent validity\(Campbell and Fiske,[1959](https://arxiv.org/html/2608.10268#bib.bib61)\)\. Further validation with larger annotator samples is a top priority\.

## 4Exploratory Results

### 4\.1Measurement

We benchmark four models chosen to cover the current frontier and one open\-source reference\. An example scenario and subscenario is provided in Appendix[Appendix C: Scenarios](https://arxiv.org/html/2608.10268#Ax3)and the question types are included in Appendix[Appendix C: Questions](https://arxiv.org/html/2608.10268#Ax4)\. The proprietary tier comprises GPT\-5 \(OpenAI, gpt\-5\-2025\-08\-07\), Claude Opus 4\.7 \(Anthropic, claude\-opus\-4\-7\-2025\-01\-30\), and Gemini 3 \(Google, gemini\-3\-flash\-preview, released 17 December 2025\): three flagship systems, included to surface differences across providers at the closed\-source frontier\. The open\-source reference is Qwen 3\.5\-9B \(Alibaba, released 24 February 2026\), a 9B\-parameter model that fits on a single H100 and approximates what a self\-hostable deployment can offer today\. Every model is queried with five independent random seeds per question; answer\-choice orderings are shuffled per seed to mitigate position bias\. All evaluation runs occurred on 20, 21, or 22 May 2026\. The overall accuracy across all question types is outlined in Table[4](https://arxiv.org/html/2608.10268#S4.T4)\.

##### Structured\-output extraction\.

All scoring assumes machine\-parseable outputs\. For each question type we define a Pydantic schema: a single\- or comma\-joined answer letter from a multiple choice questions list for I1/I2/R, a comma\-separated ranking with per\-rule rationale for Ranked Application, and a list of free\-text remedies for PR, and obtain conforming responses through each provider’s native structured\-output interface: OpenAI’sbeta\.chat\.completions\.parse, Anthropic’s tool\-use mechanism with the schema declared as the tool input, Gemini’sresponse\_json\_schema, and, for Qwen served via vLLM, JSON\-mode generation with the schema injected directly into the prompt\. This obviates fragile regex post\-processing and ensures grading is performed against exactly the format the model was instructed to produce\.

##### Multiple\-choice scoring \(I \(I1 and I2\), R\)\.

The three multiple\-choice question types reduce to single\- or multi\-letter selections\. We score each response by exact\-set match against the answer key\.

##### Rule Application\.

Each Rule Application item, for the purpose of this pilot study, presents a small set of legal rules \(typically 5–7\) and asks the model to rank them by relevance to a described human\-rights scenario\. We compute Kendall’sτ\\tau\(scipy\.stats\.kendalltau\) between the predicted ranking and the gold ranking over the intersection of the two ranked sets, and declare a response*correct*whenτ≥0\.7\\tau\\geq 0\.7— a conventional cutoff for “strong” rank agreement\. To avoid spurious credit for models that default to the answer\-choice order as presented, the rule labels are shuffled deterministically per \(seed, scenario\) and the gold ranking is remapped onto the shuffled labels, so the LLM never sees the ground\-truth ordering as the natural alphabetical sequence\. Responses that fail to produce a parseable, complete ranking are assignedτ=0\\tau=0and counted as not\-correct, matching the errors\-as\-wrong convention used elsewhere in this section\.

##### Proposed Remedy: embedding\-based correctness with human\-calibrated thresholding\.

To evaluate open\-ended LLM responses against reference answers, we score each response by the cosine similarity between OpenAItext\-embedding\-3\-smallembeddings of \(i\) the concatenated ground\-truth remedies and \(ii\) the concatenated model\-generated remedies\. To convert this continuous score into a binary “correct” / “incorrect” verdict that can be aggregated across models, we calibrate a single decision thresholdτ\\tauagainst human judgment\. Two annotators independently ratedN=40N=40responses on a 1–5 Likert scale capturing coverage of the reference remedies; we declare a response*correct*when the mean of the two ratings is at least 4\. The thresholdτ\\tauis chosen to maximize Cohen’sκ\\kappabetween the binarized embedding predictor and the human gold label\. To avoid optimistic bias from selectingτ\\tauon the same data we evaluate on, we report a leave\-one\-out cross\-validatedκ\\kappain which the threshold is refit on the remainingn−1n\-1rows for every held\-out example\. Inter\-annotator agreement \(Cohen’sκ\\kappaon binarized labels and Spearman’sρ\\rhoon raw scores\) is reported as a calibration ceiling against which the automatic scorer should be interpreted\.

##### Calibration outcome\.

AcrossN=40N=40annotated responses, inter\-annotator agreement wasκ=0\.02\\kappa=0\.02on the binarized labels andρ=0\.42\\rho=0\.42on the raw scores\. Theκ\\kappa\-optimal threshold wasτ⋆=0\.71\\tau^\{\\star\}=0\.71, yielding in\-sampleκ=0\.54\\kappa=0\.54and LOOCVκ=0\.12\\kappa=0\.12\. We use thisτ⋆\\tau^\{\\star\}to report each model’s accuracy as the share of responses whose embedding cosine exceedsτ⋆\\tau^\{\\star\}, averaged across five seeds per model\.

### 4\.2Cross\-task comparison

[Table4](https://arxiv.org/html/2608.10268#S4.T4)pools all five question types into a single overall accuracy per model\. Gemini 3 leads at0\.577±0\.0160\.577\\pm 0\.016, followed by GPT\-5 \(0\.537±0\.0150\.537\\pm 0\.015\) and Claude Opus 4\.7 \(0\.508±0\.0090\.508\\pm 0\.009\); Qwen 3\.5\-9B lags substantially at0\.339±0\.0200\.339\\pm 0\.020\. The per\-task breakdown in[Table5](https://arxiv.org/html/2608.10268#S4.T5)sharpens the picture: Claude Opus 4\.7 is in fact the strongest model on the R sub\-task \(0\.7740\.774\), but Gemini 3 wins on every other type \(I2:0\.7100\.710, RA:0\.2400\.240, PR:0\.6300\.630\)\. Ranked Application is the hardest task across all four models : even the best system clears theτ≥0\.7\\tau\\geq 0\.7Kendall correctness bar on only roughly one row in four : reflecting the gap between selecting an answer from an enumerated set and producing a fully rationale\-aligned ordering\. Notably, Qwen 3\.5\-9B closes most of the gap to the proprietary frontier on PR \(0\.5310\.531vs\. Gemini’s0\.6300\.630\) while remaining well behind on the more constrained sub\-tasks, suggesting that open\-ended generation against a holistic reference is presently more achievable for a 9B open\-source model than tasks demanding precise alignment to structured ground truth\.

Table 4:Overall mean accuracy across 5 seeds, pooled across I1, I2, R, A, and P question types\. RA uses Kendall’sτ≥0\.7\\tau\\geq 0\.7; PR uses cosine similarity≥0\.71\\geq 0\.71\. Errored LLM calls are counted as incorrect\. Qwen 3\.5\-9B is averaged over 4 seeds \(only those with MCQ data present\)\.Table 5:Mean accuracy across 5 seeds, broken down by question type\. RA is scored as correct when Kendall’sτ≥0\.7\\tau\\geq 0\.7; PR is scored as correct when full\-response cosine similarity≥0\.71\\geq 0\.71\(threshold calibrated against human annotators\)\. Qwen 3\.5\-9B MCQ values \(I1, I2, RR\) are over 4 seeds\.

## 5Discussion and Future Work

##### Discussion\.

Our pilot results, although limited in scale, suggest that structured human rights reasoning tasks are challenging for frontier LLMs\. More interesting is the way in which they are challenging\. On closed\-form tasks, LLMs regularly performed poorly on issue\-identification—a result that has consequential implications not only for human rights workers but also from AI model developers, users, and regulators\. Failures at this layer of the human rights reasoning process cascade throughout all other layers, leading to ungrounded rule applications, misconfigured remedies, and, ultimately, produce circumstances in which retrogression becomes structurally more likely\. HumRightsBench makes these failures legible to all responsible actors\. We note, models perform worst overall on rule application tasks, but this is expected as our pilot methodology sets very high thresholds for ‘correct’ responses here, and we are actively refining these question types\. We discuss this further below\.

The institutional landscape is now, for the first time, structured to receive this kind of evidence\.The Council of Europe’s Committee on Artificial Intelligence has adopted HUDERIA as guidance for risk and impact assessment in support of the Framework Convention\. HUDERIA is designed to be used by both public and private actors to identify and address risks to human rights, democracy, and the rule of law across all phases of AI deployment\. Yet HUDERIA lacks an empirical basis for evaluating whether the LLMs being assessed \(or being used to conduct the assessment\) are capable of reasoning about the rights implicated\. HumRightsBench, once developed, is precisely the tool that can provide such a basis\. Similarly, HumRightsBench results could directly inform Fundamental Rights Impact Assessment \(FRIA\) processes, providing the kind of structured, documented, and reproducible evidence that compliance with Article 27 of the EU AI Act demands\.

Expanding coverage\.Recall that this pilot’s scenarios are limited to a carefully produced selection that implicate only one right primarily: the right to water\. Our immediate priority is to achieve more breadth\. We plan to extend our scenario coverage to include thethe right to due processandthe right to educationin the near future\. Likewise, we plan to expand into different languages, considering\(Samwayet al\.,[2026](https://arxiv.org/html/2608.10268#bib.bib72)\)’s demonstration of the variance of model performance across different\-language inputs\.

Expanding methods\.We will also broaden our task, scoring, and assessment question design\. We plan to assess the suitability of implementing an LLM\-as\-judge scoring pipeline to provide another interpretive signal of model performance on open\-ended questions \(e\.g\. the Proposed Remedies\)\. We also plan to further decompose IRAP to include new types of I questions that move away from answer choices predicated on respect\-protect\-conduct indicators to those that reflect resource\-modulated obligations \(to minimum core versus progressive realization standards\)\. Additionally, encouraged by promising early signal, we plan to conduct more substantial validation and inter\-item consistency testing to ensure the statistical validity of our core construct\.

Refining assessment\.As noted in Section 3\.1\.3, our current Rule Application task only slightly moves our analysis from Rule Recall to Rule Application, and, importantly, does not sufficiently demonstrate reasoning over case\-specific facts\. We are actively developing new types of questions to better assess rule application, namely \(1\) “rule factors” questions that seek to assess model capabilities to correctly identify appropriate juridical tests, given the facts of specific scenarios; and \(2\) “jurisprudence” questions, which seek to assess model capabilities to correctly identify appropriate interpretive standards and thereby demonstrate the capability to sample authentic representations of human rights law as it is actually practiced today\.

## 6Conclusion

We introduced the core methodology required to produce HumRightsBench, an expert\-validated scenario\-based automated evaluation benchmark for human rights legal reasoning in LLMs and LRMs\. On a pilot focusing on the human right to water, frontier model performance ranged widely, with significant differences across different elements of our human rights legal reasoning framework \(IRAP\)\. Importantly, models performed worst on issue\-identification tasks, raising serious questions about their capabilities in this critical area of adoption\. Although these results are exploratory–our dataset is small–they establish this area as an urgently important subfield of evaluation science, one that we intend to explore in greater depth and with more robust validation in work to come\.

## Acknowledgments

We wish to thank all of our scenario reviewers, including Susan Morrissey, Laura Carter, Nathan Heath, Richard Ncube, Selam Abdella, Jera White, Ajitha Sritharan, Angela Kariuki, and others who wish to remain anonymous\. We are also grateful to friends and colleagues who advised on several matters throughout the project’s inception and delivery, including Dominico Zipoli, Fola Adeleke, Zach Lampell, Helene Moliner, Claudia Flores, Megan Manion, Min Aung, Khalid Hassine, Lynn Gentile, Isabel Ebert, Nathalie Stadelmann, and Jan Rydzak, among others\.

## Impact Statement

### IS1\. Human Subjects Research for Data Validation

Annotation for scenario validation did not require Institutional Review Board \(IRB\) review, as the activity was determined to constitute PPI rather than human subjects research\. Participation was voluntary and consensual: annotators contributed in exchange for acknowledgment rather than compensation, and were informed of these terms at recruitment and again on the survey form itself\. Withdrawal was permitted at any time and carried no penalty\. No deception was used at any stage\. All responses were collected and processed through secure web forms hosted on Google Cloud, accessible only to project staff\.

### IS2\. Ethics Statement

As we have stated above, we think reporting out methodology positively contributes to the mission shared by many AI evaluation scientists to ensure we develop methods for better understanding the instance\- and systems\-level impacts and harms AI systems may cause to critical human institutions\.

We are mindful, of course, of the dual nature of many technologies\. It occurs to us that this tool could in theory instead become a tool for assessing human rights workers themselves\. This is not our intent or in the broader human rights community’s interests, and we implore researchers who may build on this work to consider related dual\-use risks as AI diffusion and disruption proceeds\. We emphatically object to this tool being used to assess the qualifications and performance of human legal professionals, as this would require altogether different methods and artefacts\.

Likewise, we reason that such an evaluation could be used by malicious actors as a tool to determine potential loopholes or structural weaknesses in current human rights instruments for the purposes of exploiting them\. In future, we intend to carry out precisely this kind of adversarial evaluation to forestall such efforts \(and to advance our understanding of model safety guardrail robustness\)\. In any case, we perceive this risk to be low–there is neither sufficient scale nor robustness in this pilot alone to enable such actions–but we note that this is an area of consideration, and we will take measures to prevent such malicious usage as much as possible in future\.

### IS3\. Generative AI Statement

Authors used generative AI during this research\.

In preparation for producing evaluation scenarios, developing a grounding in specific aspects of certain human rights instruments \(e\.g\. social rights conventions\) was in part aided by the use of NotebookLM\.

Generative AI was also used to aid with table and figure Latex formatting, or, in limited cases, to translate author\-written text into an illustrative diagram or summary table\.

Likewise, some authors used Generative AI to prepare bibtex entries from validated sources where no bibtex citation was provided by the publisher\.

In all cases, authors retained sole control \(and responsibility\) for the experimental design, analysis, and interpretation of implications of this work\.

## References

- N\. Aletras, D\. Tsarapatsanis, D\. Preoţiuc\-Pietro, and V\. Lampos \(2016\)Predicting judicial decisions of the European Court of Human Rights: a natural language processing perspective\.PeerJ Computer Science2,pp\. e93\.External Links:[Document](https://dx.doi.org/10.7717/peerj-cs.93)Cited by:[Human Rights and International Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px5.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1),[§3\.1\.2](https://arxiv.org/html/2608.10268#S3.SS1.SSS2.p1.1)\.
- P\. Alston \(2017\)The populist challenge to human rights\.Journal of Human Rights Practice9\(1\),pp\. 1–15\.External Links:[Document](https://dx.doi.org/10.1093/jhuman/hux007),[Link](https://doi.org/10.1093/jhuman/hux007)Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p1.1)\.
- P\. Alston \(2015\)Report of the special rapporteur on extreme poverty and human rights, philip alston\.UN Doc\.Technical ReportA/HRC/29/31,United Nations Human Rights Council,Geneva\.Note:Twenty\-ninth session, agenda item 3\.[https://digitallibrary\.un\.org/record/798707](https://digitallibrary.un.org/record/798707)Cited by:[Table 7](https://arxiv.org/html/2608.10268#Ax1.T7.6.4.3.2.1.1)\.
- R\. K\. Arora, J\. Wei, R\. S\. Hicks, P\. Bowman, J\. Quiñonero\-Candela, F\. Tsimpourlas, M\. Sharman, M\. Shah, A\. Vallone, A\. Beutel, J\. Heidecke, and K\. Singhal \(2025\)HealthBench: evaluating large language models towards improved human health\.External Links:2505\.08775,[Link](https://arxiv.org/abs/2505.08775)Cited by:[Unique cross\-jurisdictional variance\.](https://arxiv.org/html/2608.10268#Ax2.SSx2.SSS0.Px5.p1.1)\.
- Y\. Bai, S\. Kadavath, S\. Kundu, A\. Askell, J\. Kernion, A\. Jones, A\. Chen, A\. Goldie, A\. Mirhoseini, C\. McKinnon,et al\.\(2022\)Constitutional ai: harmlessness from ai feedback\.arXiv preprint arXiv:2212\.08073\.Cited by:[§3](https://arxiv.org/html/2608.10268#S3.p2.1)\.
- A\. Bavaresco, R\. Bernardi, L\. Bertolazzi, D\. Elliott, R\. Fernández, A\. Gatt, E\. Ghaleb, M\. Giulianelli, M\. Hanna, A\. Koller, A\. F\. T\. Martins, P\. Mondorf, V\. Neplenbroek, S\. Pezzelle, B\. Plank, D\. Schlangen, A\. Suglia, A\. K\. Surikuchi, E\. Takmaz, and A\. Testoni \(2024\)LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks\.External Links:2406\.18403Cited by:[§2\.4](https://arxiv.org/html/2608.10268#S2.SS4.p1.1)\.
- A\. M\. Bean, R\. O\. Kearns, A\. Romanou, F\. S\. Hafner, H\. Mayne, J\. Batzner, N\. Foroutan, C\. Schmitz, K\. Korgul, H\. Batra, O\. Deb, E\. Beharry, C\. Emde, T\. Foster, A\. Gausen, M\. Grandury, S\. Han, V\. Hofmann, L\. Ibrahim, H\. Kim, H\. R\. Kirk, F\. Lin, G\. K\. Liu, L\. Luettgau, J\. Magomere, J\. Rystrøm, A\. Sotnikova, Y\. Yang, Y\. Zhao, A\. Bibi, A\. Bosselut, R\. Clark, A\. Cohan, J\. Foerster, Y\. Gal, S\. A\. Hale, I\. D\. Raji, C\. Summerfield, P\. H\. S\. Torr, C\. Ududec, L\. Rocher, and A\. Mahdi \(2025\)Measuring what matters: construct validity in large language model benchmarks\.External Links:2511\.04703,[Link](https://arxiv.org/abs/2511.04703)Cited by:[Grounding in practice\.](https://arxiv.org/html/2608.10268#Ax2.SSx2.SSS0.Px2.p1.1)\.
- D\. T\. Campbell and D\. W\. Fiske \(1959\)Convergent and discriminant validation by the multitrait\-multimethod matrix\.Psychological Bulletin56\(2\),pp\. 81–105\.Cited by:[§3\.2](https://arxiv.org/html/2608.10268#S3.SS2.p4.1)\.
- I\. Chalkidis, I\. Androutsopoulos, and N\. Aletras \(2019\)Neural legal judgment prediction in English\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,pp\. 4317–4323\.External Links:[Document](https://dx.doi.org/10.18653/v1/P19-1424)Cited by:[Human Rights and International Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px5.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.
- I\. Chalkidis, M\. Fergadiotis, P\. Malakasiotis, N\. Aletras, and I\. Androutsopoulos \(2022\)LexGLUE: a benchmark dataset for legal language understanding in English\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 4310–4330\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.297)Cited by:[Legal Knowledge and Language Understanding\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px2.p1.1),[Human Rights and International Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px5.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- I\. Chalkidis, N\. Garneau, C\. Goanta, D\. M\. Katz, and A\. Søgaard \(2023\)LeXFiles and LegalLAMA: facilitating English multinational legal language model development\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Note:arXiv:2305\.07507Cited by:[Legal Knowledge and Language Understanding\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px2.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- O\. S\. Chlapanis, D\. Galanis, and I\. Androutsopoulos \(2024\)LAR\-ECHR: a new legal argument reasoning task and dataset for cases of the European Court of Human Rights\.InProceedings of the Natural Legal Language Processing Workshop 2024,Cited by:[§2\.4](https://arxiv.org/html/2608.10268#S2.SS4.p1.1)\.
- K\. Cole \(2024\)Navigating humanitarian ai: lessons learned from building a chatbot proof\-of\-concept\.Technical reportRefugee Solidarity Network\.External Links:[Link](https://refugeesolidaritynetwork.org/reports/navigating-humanitarian-ai-lessons-learned-from-building-a-chatbot-proof-of-concept/)Cited by:[§1](https://arxiv.org/html/2608.10268#S1.p1.1)\.
- Columbia Law School \(2022\)Organizing a legal discussion: IRAC / CRAC / CREAC\.Note:Writing Center handout, Columbia Law SchoolRevised May 2022\.[https://www\.law\.columbia\.edu/sites/default/files/2022\-06/WC%20Handout%20IRAC%2C%20CRAC%2C%20CREAC\.revised%205\.22\.pdf](https://www.law.columbia.edu/sites/default/files/2022-06/WC%20Handout%20IRAC%2C%20CRAC%2C%20CREAC.revised%205.22.pdf)Cited by:[§3\.1](https://arxiv.org/html/2608.10268#S3.SS1.p1.1)\.
- Y\. Dai, D\. Feng, J\. Huang, H\. Jia, Q\. Xie, Y\. Zhang, W\. Han, W\. Tian, and H\. Wang \(2025\)LAiW: a Chinese legal large language models benchmark\.InProceedings of the 31st International Conference on Computational Linguistics,pp\. 10738–10766\.Cited by:[General Legal Reasoning Benchmarks\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- O\. de Schutter \(2010\)Report of the special rapporteur on the right to food, olivier de schutter: addendum — mission to Brazil\.UN Doc\.Technical ReportA/HRC/13/33/Add\.6,United Nations Human Rights Council,Geneva\.Note:Thirteenth session, agenda item 3\. Mission to Brazil, 12–18 October 2009\.[https://digitallibrary\.un\.org/record/677702](https://digitallibrary.un.org/record/677702)Cited by:[Table 7](https://arxiv.org/html/2608.10268#Ax1.T7.6.4.3.2.1.1)\.
- D\. Emelin, R\. Le Bras, J\. D\. Hwang, M\. Forbes, and Y\. Choi \(2021\)Moral stories: situated reasoning about norms, intents, actions, and their consequences\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 698–718\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.54)Cited by:[Normative, Moral, and Ethical Reasoning\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px6.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.
- J\. Enguehard, M\. V\. Ermengem, K\. Atkinson, S\. Cha, A\. G\. Chowdhury, P\. K\. Ramaswamy, J\. Roghair, H\. R\. Marlowe, C\. S\. Negreanu, K\. Boxall, and D\. Mincu \(2025\)LeMAJ \(legal llm\-as\-a\-judge\): bridging legal reasoning and llm evaluation\.External Links:2510\.07243,[Link](https://arxiv.org/abs/2510.07243)Cited by:[Evaluation Methodology\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px7.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1),[§2\.4](https://arxiv.org/html/2608.10268#S2.SS4.p1.1)\.
- M\. Eriksson, E\. Purificato, A\. Noroozian, J\. Vinagre, G\. Chaslot, E\. Gomez, and D\. Fernandez\-Llorca \(2025\)Can we trust AI benchmarks? an interdisciplinary review of current issues in AI evaluation\.InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society,Vol\.8,pp\. 850–864\.External Links:[Document](https://dx.doi.org/10.1609/aies.v8i1.36595)Cited by:[Grounding in practice\.](https://arxiv.org/html/2608.10268#Ax2.SSx2.SSS0.Px2.p1.1)\.
- Y\. Fan, J\. Ni, J\. Merane, Y\. Tian, Y\. Hermstrüwer, Y\. Huang, M\. Akhtar, E\. Salimbeni, F\. Geering, O\. Dreyer, D\. Brunner, M\. Leippold, M\. Sachan, A\. Stremitzer, C\. Engel, E\. Ash, and J\. Niklaus \(2025\)LEXam: benchmarking legal reasoning on 340 law exams\.Note:Accepted to ICLR 2026External Links:2505\.12864Cited by:[General Legal Reasoning Benchmarks\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px1.p1.1),[Evaluation Methodology\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px7.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1),[§2\.4](https://arxiv.org/html/2608.10268#S2.SS4.p1.1)\.
- Z\. Fei, X\. Shen, D\. Zhu, F\. Zhou, Z\. Han, S\. Zhang, K\. Chen, Z\. Shen, and J\. Ge \(2023\)LawBench: benchmarking legal knowledge of large language models\.External Links:2309\.16289Cited by:[General Legal Reasoning Benchmarks\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- N\. Guha, J\. Nyarko, D\. Ho, C\. Ré, A\. Chilton, A\. Chohlas\-Wood, A\. Peters, B\. Waldon, D\. Rockmore, D\. Zambrano,et al\.\(2023\)Legalbench: a collaboratively built benchmark for measuring legal reasoning in large language models\.Advances in neural information processing systems36,pp\. 44123–44279\.Cited by:[General Legal Reasoning Benchmarks\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px1.p1.1),[Evaluation Methodology\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px7.p1.1),[Difference from municipal law\.](https://arxiv.org/html/2608.10268#Ax2.SSx2.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2608.10268#S1.p2.2),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1),[§3\.1\.2](https://arxiv.org/html/2608.10268#S3.SS1.SSS2.p1.1),[§3\.1\.3](https://arxiv.org/html/2608.10268#S3.SS1.SSS3.p1.1),[§3\.1](https://arxiv.org/html/2608.10268#S3.SS1.p1.1)\.
- A\. Haupt, M\. MacLennan, U\. Ramli, R\. Moreno Jiménez, A\. Pentland, and S\. Koyejo \(2026\)Un\-bench: closing the simulacra gap in development data\.Note:NeurIPS 2026 Competition TrackStanford Trustworthy AI Research, Stanford HAI, UN Behavioural Science Group, UNICEF, and UNHCR\.[https://humanitarianevals\.org/](https://humanitarianevals.org/)Cited by:[Human Rights and International Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px5.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.
- P\. Henderson, M\. S\. Krass, L\. Zheng, N\. Guha, C\. D\. Manning, D\. Jurafsky, and D\. E\. Ho \(2022\)Pile of law: learning responsible data filtering from the law and a 256GB open\-source legal dataset\.InAdvances in Neural Information Processing Systems 35 \(NeurIPS 2022\), Datasets and Benchmarks Track,Note:arXiv:2207\.00220Cited by:[Legal Knowledge and Language Understanding\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px2.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Critch, J\. Li, D\. Song, and J\. Steinhardt \(2021a\)Aligning AI with shared human values\.InProceedings of the International Conference on Learning Representations \(ICLR\),Note:arXiv:2008\.02275Cited by:[Normative, Moral, and Ethical Reasoning\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px6.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.
- D\. Hendrycks, C\. Burns, A\. Chen, and S\. Ball \(2021b\)CUAD: an expert\-annotated NLP dataset for legal contract review\.InAdvances in Neural Information Processing Systems 34 \(NeurIPS 2021\), Datasets and Benchmarks Track,Note:arXiv:2103\.06268Cited by:[Contracts, Statutes, and Case Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px3.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- N\. Holzenberger, A\. Blair\-Stanek, and B\. Van Durme \(2020\)A dataset for statutory reasoning in tax law entailment and question answering\.InProceedings of the Natural Legal Language Processing Workshop 2020,Note:arXiv:2005\.05257Cited by:[Contracts, Statutes, and Case Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px3.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- N\. Holzenberger and B\. Van Durme \(2021\)Factoring statutory reasoning as language understanding challenges\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics,Note:arXiv:2105\.07903Cited by:[Contracts, Statutes, and Case Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px3.p1.1)\.
- J\. Jiao, S\. Afroogh,et al\.\(2025\)LLM ethics benchmark: a three\-dimensional assessment system for evaluating moral reasoning in large language models\.Scientific Reports15\.Note:arXiv:2505\.00853External Links:[Document](https://dx.doi.org/10.1038/s41598-025-18489-7)Cited by:[Normative, Moral, and Ethical Reasoning\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px6.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.
- M\. Koskenniemi \(2001\)The gentle civilizer of nations: the rise and fall of international law 1870–1960\.Cambridge University Press,Cambridge, UK\.External Links:ISBN 9780521623117Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p1.1)\.
- M\. Langford \(Ed\.\) \(2009\)Social rights jurisprudence: emerging trends in international and comparative law\.Cambridge University Press\.Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p1.1)\.
- H\. Li, Y\. Su, D\. Cai, Y\. Wang, and L\. Liu \(2024\)LexEval: a comprehensive Chinese legal benchmark for evaluating large language models\.InAdvances in Neural Information Processing Systems 37 \(NeurIPS 2024\), Datasets and Benchmarks Track,Cited by:[General Legal Reasoning Benchmarks\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- Y\. Li and G\. Wu \(2026\)Legaleval\-q: a benchmark for quality evaluation of llm\-generated chinese legal text: legaleval\-q: a benchmark for quality evaluation…\.Knowl\. Inf\. Syst\.68\(1\)\.External Links:ISSN 0219\-1377,[Link](https://doi.org/10.1007/s10115-026-02703-7),[Document](https://dx.doi.org/10.1007/s10115-026-02703-7)Cited by:[Evaluation Methodology\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px7.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.
- R\. H\. Lie and M\. Langford \(2024\)The computational turn in international law\.Nordic Journal of International Law93\(1\),pp\. 38–67\.External Links:ISSN 0902\-7351,[Link](https://doi.org/10.1163/15718107-bja10081),[Document](https://dx.doi.org/10.1163/15718107-bja10081)Cited by:[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.
- Y\. Liu, J\. Cao, C\. Liu, K\. Ding, and L\. Jin \(2024\)Datasets for large language models: a comprehensive survey\.External Links:2402\.18041,[Link](https://arxiv.org/abs/2402.18041)Cited by:[Multilingual and Cross\-Jurisdictional Benchmarks\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px4.p1.1)\.
- N\. Lourie, R\. Le Bras, and Y\. Choi \(2021\)SCRUPLES: a corpus of community ethical judgments on 32,000 real\-life anecdotes\.Proceedings of the AAAI Conference on Artificial Intelligence35\(15\),pp\. 13470–13479\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v35i15.17589)Cited by:[Normative, Moral, and Ethical Reasoning\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px6.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.
- M\. Medvedeva, M\. Vols, and M\. Wieling \(2020\)Using machine learning to predict decisions of the european court of human rights\.Artificial Intelligence and Law28\(2\),pp\. 237–266\.External Links:ISSN 0924\-8463,[Link](https://doi.org/10.1007/s10506-019-09255-y),[Document](https://dx.doi.org/10.1007/s10506-019-09255-y)Cited by:[Human Rights and International Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px5.p1.1)\.
- S\. Moyn \(2010\)The last utopia: human rights in history\.Belknap Press of Harvard University Press,Cambridge, MA\.External Links:ISBN 9780674048720Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p1.1)\.
- J\. Niklaus, V\. Matoshi, P\. Rani, A\. Galassi, M\. Stürmer, and I\. Chalkidis \(2023\)LEXTREME: a multi\-lingual and multi\-task benchmark for the legal domain\.InFindings of the Association for Computational Linguistics: EMNLP 2023,Note:arXiv:2301\.13126Cited by:[Legal Knowledge and Language Understanding\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px2.p1.1),[Multilingual and Cross\-Jurisdictional Benchmarks\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px4.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- J\. Niklaus, V\. Matoshi, M\. Stürmer, I\. Chalkidis, and D\. E\. Ho \(2024\)MultiLegalPile: a 689GB multilingual legal corpus\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Note:arXiv:2306\.02069Cited by:[Multilingual and Cross\-Jurisdictional Benchmarks\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px4.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- Office of the United Nations High Commissioner for Human Rights \(2011\)Manual on human rights monitoring\.Technical reportTechnical Report7/Rev\.1,Professional Training Series,United Nations,Geneva, Switzerland\.Cited by:[§3\.1\.2](https://arxiv.org/html/2608.10268#S3.SS1.SSS2.p1.1)\.
- Office of the United Nations High Commissioner for Human Rights \(2025\)The core international human rights instruments and their monitoring bodies\.Note:[https://www\.ohchr\.org/en/core\-international\-human\-rights\-instruments\-and\-their\-monitoring\-bodies](https://www.ohchr.org/en/core-international-human-rights-instruments-and-their-monitoring-bodies)Accessed: 2026\-05\-22Cited by:[§3](https://arxiv.org/html/2608.10268#S3.p2.1)\.
- Office of the United Nations High Commissioner for Human Rights \(2006a\)Frequently asked questions on a human rights\-based approach to development cooperation\.UN Doc\.Technical ReportHR/PUB/06/8,United Nations,New York and Geneva\.Cited by:[§3\.1](https://arxiv.org/html/2608.10268#S3.SS1.p1.1)\.
- Office of the United Nations High Commissioner for Human Rights \(2012b\)Human rights indicators: a guide to measurement and implementation\.UN Doc\.Technical ReportHR/PUB/12/5,United Nations,New York and Geneva\.External Links:ISBN 978\-92\-1\-056286\-7Cited by:[§3\.1\.1](https://arxiv.org/html/2608.10268#S3.SS1.SSS1.p3.1)\.
- L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray,et al\.\(2022\)Training language models to follow instructions with human feedback\.Advances in neural information processing systems35,pp\. 27730–27744\.Cited by:[§3](https://arxiv.org/html/2608.10268#S3.p2.1)\.
- S\. Pedersen \(2015\)The guardians: the league of nations and the crisis of empire\.Oxford University Press,Oxford\.External Links:[Document](https://dx.doi.org/10.1093/acprof%3Aoso/9780199570485.001.0001),[Link](https://doi.org/10.1093/acprof:oso/9780199570485.001.0001)Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p1.1)\.
- P\. J\. Pesch \(2025\)Potentials and challenges of large language models \(llms\) in the context of administrative decision\-making\.European Journal of Risk Regulation16\(1\),pp\. 76–95\.External Links:ISSN 2190\-8249,[Link](https://doi.org/10.1017/err.2024.99),[Document](https://dx.doi.org/10.1017/err.2024.99)Cited by:[§1](https://arxiv.org/html/2608.10268#S1.p1.1)\.
- V\. Rasiah, R\. Stern, V\. Matoshi, M\. Stürmer, I\. Chalkidis, D\. E\. Ho, and J\. Niklaus \(2023\)SCALE: scaling up the complexity for advanced language model evaluation\.External Links:2306\.09237Cited by:[Multilingual and Cross\-Jurisdictional Benchmarks\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px4.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- M\. G\. Rovira \(2013\)The project of positivism in international law\.Oxford University Press,Oxford, United Kingdom\.Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p1.1)\.
- J\. G\. Ruggie \(2013\)Just business: multinational corporations and human rights\.W\. W\. Norton & Company,New York\.External Links:ISBN 978\-0\-393\-06288\-5Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p3.1)\.
- K\. Samway, M\. N\. Takagi, R\. Mihalcea, B\. Schölkopf, I\. Chalkidis, D\. Hershcovich, and Z\. Jin \(2026\)When do language models endorse limitations on human rights principles?\.InFindings of the Association for Computational Linguistics: EACL 2026,V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 6597–6623\.External Links:[Link](https://aclanthology.org/2026.findings-eacl.347/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-eacl.347),ISBN 979\-8\-89176\-386\-9Cited by:[Human Rights and International Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px5.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1),[§5](https://arxiv.org/html/2608.10268#S5.SS0.SSS0.Px1.p3.1)\.
- E\. Schneiders, T\. Seabrooke, J\. Krook, R\. Hyde, N\. Leesakul, J\. Clos, and J\. E\. Fischer \(2025\)Objection overruled\! lay people can distinguish large language models from lawyers, but still favour advice from an llm\.InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems,CHI ’25,New York, NY, USA,pp\. 1–14\.External Links:ISBN 9798400713941,[Link](https://doi.org/10.1145/3706598.3713470),[Document](https://dx.doi.org/10.1145/3706598.3713470)Cited by:[§1](https://arxiv.org/html/2608.10268#S1.p1.1)\.
- R\. Schwartz, R\. Chowdhury, A\. Kundu, H\. Frase, M\. Fadaee, T\. David, G\. Waters, A\. Taik, M\. Briggs, P\. Hall, S\. Jain, K\. Yee, S\. Thomas, S\. Bhandari, P\. Duncan, A\. Thompson, M\. Carlyle, Q\. Lu, M\. Holmes, and T\. Skeadas \(2025\)Reality check: a new evaluation ecosystem is necessary to understand ai’s real world effects\.External Links:2505\.18893,[Link](https://arxiv.org/abs/2505.18893)Cited by:[Grounding in practice\.](https://arxiv.org/html/2608.10268#Ax2.SSx2.SSS0.Px2.p1.1)\.
- R\. Schwartz, J\. Fiscus, K\. Greene, G\. Waters, R\. Chowdhury, T\. Jensen, C\. Greenberg, A\. Godil, R\. Amironesei, P\. Hall, and S\. Jain \(2024\)Assessing risks and impacts of ai \(aria\): pilot evaluation plan\.Technical reportNational Institute of Standards and Technology \(NIST\)\.Note:Accessed: 2026\-05\-21External Links:[Link](https://ai-challenges.nist.gov/aria/docs/evaluation_plan.pdf)Cited by:[Grounding in practice\.](https://arxiv.org/html/2608.10268#Ax2.SSx2.SSS0.Px2.p1.1)\.
- Y\. Shi, H\. Liu, Y\. Hu, G\. Song, X\. Xu, Y\. Ma, T\. Tang, L\. Zhang, Q\. Chen, D\. Feng, W\. Lv, W\. Wu, K\. Yang, S\. Yang, W\. Wang, R\. Shi, Y\. Qiu, Y\. Qi, J\. Zhang, X\. Sui, Y\. Chen, Y\. Zhang, A\. Yang, B\. Yu, D\. Liu, J\. Lin, W\. Shen, B\. Zhao, C\. L\. A\. Clarke, and H\. Wei \(2026\)PLawBench: a rubric\-based benchmark for evaluating llms in real\-world legal practice\.External Links:2601\.16669,[Link](https://arxiv.org/abs/2601.16669)Cited by:[Evaluation Methodology\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px7.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.
- G\. Sluga \(2013\)Internationalism in the age of nationalism\.Pennsylvania Studies in Human Rights,University of Pennsylvania Press,Philadelphia\.Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p1.1)\.
- R\. Subramanian, T\. S\. Shiromani, A\. Chaudhry, R\. Li, V\. Sharma, K\. Zhu, and S\. Dev \(2026\)ProMoral\-Bench: evaluating prompting strategies for moral reasoning and safety in LLMs\.External Links:2602\.13274Cited by:[Normative, Moral, and Ethical Reasoning\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px6.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.
- D\. Tuggener, P\. von Däniken, T\. Peetz, and M\. Cieliebak \(2020\)LEDGAR: a large\-scale multi\-label corpus for text classification of legal provisions in contracts\.InProceedings of the 12th Language Resources and Evaluation Conference,pp\. 1235–1241\.Cited by:[Legal Knowledge and Language Understanding\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px2.p1.1)\.
- UN Committee on Economic, Social and Cultural Rights \(CESCR\) \(2002\)General comment no\. 15: the right to water \(arts\. 11 and 12 of the covenant\)\.Technical reportTechnical ReportE/C\.12/2002/11,United Nations,Geneva\.Note:Adopted at the Twenty\-ninth Session of the Committee on Economic, Social and Cultural RightsExternal Links:[Link](https://digitallibrary.un.org/record/486454)Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p2.1)\.
- UN Human Rights Office of the High Commissioner \(OHCHR\) \(2019\)Business and human rights in technology project \(b\-tech\): scoping paper\.Technical reportUnited Nations,Geneva, Switzerland\.External Links:[Link](https://www.ohchr.org/Documents/Issues/Business/B-Tech/B-Tech_Scoping_paper.pdf)Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p3.1)\.
- UN Human Rights Office of the High Commissioner \(2024\)Advancing responsible development and deployment of generative AI: the value proposition of the UN guiding principles on business and human rights\.B\-Tech Project Foundational PaperOffice of the United Nations High Commissioner for Human Rights,Geneva\.Note:[https://www\.ohchr\.org/en/documents/tools\-and\-resources/advancing\-responsible\-development\-and\-deployment\-generative\-ai](https://www.ohchr.org/en/documents/tools-and-resources/advancing-responsible-development-and-deployment-generative-ai)Cited by:[§2\.4](https://arxiv.org/html/2608.10268#S2.SS4.p1.1),[§3](https://arxiv.org/html/2608.10268#S3.p1.1)\.
- United Nations General Assembly \(1948\)Universal declaration of human rights\.Note:Resolution 217 A \(III\)External Links:[Link](https://un.org/)Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p1.1)\.
- United Nations General Assembly \(1965\)International convention on the elimination of all forms of racial discrimination\.United Nations,New York\.Note:United Nations, Treaty Series, vol\. 660, p\. 195External Links:[Link](https://treaties.un.org/pages/viewdetails.aspx?src=treaty&mtdsg_no=iv-2&chapter=4&clang=_en)Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p2.1)\.
- United Nations General Assembly \(1966a\)International covenant on civil and political rights\.Technical reportTechnical Report2200A \(XXI\)\.Note:Adopted 16 Dec\. 1966, entered into force 23 Mar\. 1976External Links:[Link](https://www.ohchr.org/en/instruments-mechanisms/instruments/international-covenant-civil-and-political-rights)Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p2.1)\.
- United Nations General Assembly \(1966b\)International Covenant on Economic, Social and Cultural Rights\.Note:United Nations,Treaty Series, vol\. 993, p\. 3Adopted 16 Dec\. 1966, entered into force 3 Jan\. 1976External Links:[Link](https://treaties.un.org/pages/showDetails.aspx?objid=080000028002b6ed)Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p2.1)\.
- United Nations General Assembly \(1979\)Convention on the elimination of all forms of discrimination against women\.Technical reportTechnical Report34/180,United Nations,New York, NY\.Note:Adopted 18 December 1979, entered into force 3 September 1981External Links:[Link](https://treaties.un.org/pages/viewdetails.aspx?src=treaty&mtdsg_no=iv-8&chapter=4&clang=_en)Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p2.1)\.
- United Nations General Assembly \(2006\)Convention on the rights of persons with disabilities\.ResolutionTechnical ReportA/RES/61/106,United Nations\.External Links:[Link](https://www.refworld.org/legal/agreements/unga/2006/90142)Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p2.1)\.
- United Nations Office of the High Commissioner for Human Rights \(OHCHR\) \(2011\)Guiding principles on business and human rights: implementing the united nations ”protect, respect and remedy” framework\.Technical reportHR/PUB/11/4,United Nations,New York and Geneva\.External Links:ISBN 978\-92\-1\-154201\-1,[Link](https://www.ohchr.org/en/publications/reference-publications/guiding-principles-business-and-human-rights)Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p3.1),[§3\.1\.1](https://arxiv.org/html/2608.10268#S3.SS1.SSS1.p3.1)\.
- United Nations \(1989\)Convention on the rights of the child\.Treaty Series, Vol\.1577\.Note:Adopted 20 November 1989, Entered into force 2 September 1990Cited by:[§2\.1](https://arxiv.org/html/2608.10268#S2.SS1.p2.1)\.
- S\. H\. Wanget al\.\(2025\)ACORD: an expert\-annotated retrieval dataset for legal contract drafting\.Note:Full author list requires verificationExternal Links:2501\.06582Cited by:[Contracts, Statutes, and Case Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px3.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- C\. Xiao, H\. Zhong, Z\. Guo, C\. Tu, Z\. Liu, M\. Sun, Y\. Feng, X\. Han, Z\. Hu, H\. Wang, and J\. Xu \(2018\)CAIL2018: a large\-scale legal dataset for judgment prediction\.External Links:1807\.02478Cited by:[Contracts, Statutes, and Case Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px3.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- S\. Xu, L\. Staufer, S\. T\.Y\.S\.S\., O\. Ichim, C\. Heri, and M\. Grabmair \(2023\)VECHR: a dataset for explainable and robust classification of vulnerability type in the European Court of Human Rights\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 11738–11752\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.718)Cited by:[Human Rights and International Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px5.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.
- W\. Yuet al\.\(2025\)Benchmarking multi\-step legal reasoning and analyzing chain\-of\-thought effects in large language models\.Note:Full author list requires verificationExternal Links:2511\.07979Cited by:[General Legal Reasoning Benchmarks\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1),[§2\.4](https://arxiv.org/html/2608.10268#S2.SS4.p1.1),[§3\.1](https://arxiv.org/html/2608.10268#S3.SS1.p1.1)\.
- L\. Zhang, M\. Grabmair, M\. Gray, and K\. Ashley \(2025\)Thinking longer, not always smarter: evaluating LLM capabilities in hierarchical legal reasoning\.External Links:2510\.08710Cited by:[General Legal Reasoning Benchmarks\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- L\. Zheng, N\. Guha, B\. R\. Anderson, P\. Henderson, and D\. E\. Ho \(2021\)When does pretraining help? assessing self\-supervised learning for law and the CaseHOLD dataset of 53,000\+ legal holdings\.InProceedings of the 18th International Conference on Artificial Intelligence and Law \(ICAIL ’21\),pp\. 159–168\.External Links:[Document](https://dx.doi.org/10.1145/3462757.3466088)Cited by:[Legal Knowledge and Language Understanding\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px2.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- H\. Zhong, C\. Xiao, C\. Tu, T\. Zhang, Z\. Liu, and M\. Sun \(2020\)JEC\-QA: a legal\-domain question answering dataset\.Proceedings of the AAAI Conference on Artificial Intelligence34\(05\),pp\. 9701–9708\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v34i05.6519)Cited by:[Contracts, Statutes, and Case Law\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px3.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p2.1)\.
- C\. Ziems, J\. A\. Yu, Y\. Wang, A\. Halevy, and D\. Yang \(2022\)The moral integrity corpus: a benchmark for ethical dialogue systems\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 3755–3773\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.261)Cited by:[Normative, Moral, and Ethical Reasoning\.](https://arxiv.org/html/2608.10268#Ax1.SSx2.SSS0.Px6.p1.1),[§2\.3](https://arxiv.org/html/2608.10268#S2.SS3.p3.1)\.

## Appendix A: Expanded Review of the State of the Art in Legal Benchmarking

### Summary table of sources of the human rights legal regime

Table 6:Sources of international human rights law, sorted by binding force\.Table 7:Two\-tier categorization of international human rights instruments: legally binding treaties versus authoritative but non\-binding soft\-law instruments, with representative examples for each\.
### Literature Review

##### General Legal Reasoning Benchmarks\.

The most comprehensive English\-language legal benchmark is LegalBench, which assembles 162 tasks across six categories of legal reasoning, including issue spotting, rule recall, and rule application\(Guhaet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib21)\)\. LEXam extends this paradigm to long\-form legal argumentation, drawing 7,537 questions from 340 law school exams in English and German\(Fanet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib49)\)\. Recent Chinese\-language work has pushed the field toward structured reasoning evaluation\. Like LegalBench, MSLR grounds its tasks in the IRAC framework \(a common legal reasoning benchmark decomposing that practice into Issue Identification, Rule Recall, Rule Application, and Conclusion, see section XX\.XX below\) using real judicial decisions\(Yu and others,[2025](https://arxiv.org/html/2608.10268#bib.bib93)\), LAiW organizes evaluation around the legal syllogism in three difficulty tiers\(Daiet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib47)\), LawBench tests 21 LLMs across 20 tasks spanning memorization, understanding, and application\(Feiet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib50)\), and LexEval offers the largest Chinese legal benchmark to date, structured by a cognitive ability taxonomy\(Liet al\.,[2024](https://arxiv.org/html/2608.10268#bib.bib82)\)\. Across these benchmarks, model performance degrades sharply when tasks require multi\-step reasoning rather than retrieval, indicating that current LLMs handle legal knowledge more effectively than legal inference\(Fanet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib49); Yu and others,[2025](https://arxiv.org/html/2608.10268#bib.bib93)\)\. Counter\-intuitively, LRMs occasionally perform much worse than LLMs on certain legal reasoning tasks, despite “thinking longer”\(Zhanget al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib75)\)\.

##### Legal Knowledge and Language Understanding\.

LexGLUE is the foundational benchmark for legal NLU, modeled on GLUE and covering seven classification and QA tasks across the European Court of Human Rights \(ECtHR\), the US Supreme Court, EU legislation, and commercial contracts\(Chalkidiset al\.,[2022](https://arxiv.org/html/2608.10268#bib.bib44)\)\. CaseHOLD evaluates the identification of holdings in US case law through multiple\-choice prompts derived from the Harvard Caselaw Access Project\(Zhenget al\.,[2021](https://arxiv.org/html/2608.10268#bib.bib94)\), while LEDGAR tests multi\-label classification of contract provisions from SEC filings\(Tuggeneret al\.,[2020](https://arxiv.org/html/2608.10268#bib.bib89)\)\. LegalLAMA and LexFiles probe the legal knowledge stored in pretrained models across six common\-law jurisdictions\(Chalkidiset al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib45)\)\. Pile of Law provides the underlying 256GB corpus used by many derivative resources\(Hendersonet al\.,[2022](https://arxiv.org/html/2608.10268#bib.bib52)\)\. Performance on these benchmarks has saturated relative to specialized legal language models such as Legal\-BERT and LegalXLM\-R, suggesting that pure language understanding is no longer the binding constraint on legal AI performance\(Niklauset al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib85)\)\.

##### Contracts, Statutes, and Case Law\.

Contract review is the most commercially mature subfield\. CUAD contains 13,000 expert annotations across 510 commercial contracts and 41 clause categories, framed as a span\-selection task\(Hendryckset al\.,[2021b](https://arxiv.org/html/2608.10268#bib.bib78)\)\. ACORD extends this to retrieval of precedent clauses for contract drafting\(Wang and others,[2025](https://arxiv.org/html/2608.10268#bib.bib90)\)\. For statutory reasoning, SARA tests entailment and question\-answering over the US Internal Revenue Code\(Holzenbergeret al\.,[2020](https://arxiv.org/html/2608.10268#bib.bib79)\), with a follow\-up that decomposes statutory reasoning into discrete language understanding challenges\(Holzenberger and Van Durme,[2021](https://arxiv.org/html/2608.10268#bib.bib80)\)\. Case\-law prediction is dominated by Chinese benchmarks such as CAIL2018\(Xiaoet al\.,[2018](https://arxiv.org/html/2608.10268#bib.bib91)\)and JEC\-QA from the Chinese National Judicial Examination\(Zhonget al\.,[2020](https://arxiv.org/html/2608.10268#bib.bib95)\)\. These domain\-specific benchmarks consistently show that LLMs perform well on extraction and classification but struggle with tasks requiring the integration of statutory text, case facts, and doctrinal context\(Holzenbergeret al\.,[2020](https://arxiv.org/html/2608.10268#bib.bib79); Hendryckset al\.,[2021b](https://arxiv.org/html/2608.10268#bib.bib78)\)\.

##### Multilingual and Cross\-Jurisdictional Benchmarks\.

As in other fields, the English\-only orientation of early legal NLP pretraining corpora has been a persistent limitation\(Liuet al\.,[2024](https://arxiv.org/html/2608.10268#bib.bib74)\)\. MultiLegalPile addresses this with a 689GB pretraining corpus across 24 languages and 17 jurisdictions, paired with a family of LegalXLM\-R models\(Niklauset al\.,[2024](https://arxiv.org/html/2608.10268#bib.bib86)\)\. LEXTREME provides the corresponding multilingual evaluation suite, covering classification tasks across European jurisdictions\(Niklauset al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib85)\)\. SCALE tests long\-document processing across five languages in the Swiss federal system, with documents extending to 50,000 tokens\(Rasiahet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib87)\)\. These resources demonstrate that monolingual fine\-tuning consistently outperforms multilingual pretraining for jurisdiction\-specific tasks, raising open questions about the transferability of legal reasoning across legal systems\(Niklauset al\.,[2024](https://arxiv.org/html/2608.10268#bib.bib86)\)\.

##### Human Rights and International Law\.

Human rights NLP began with\(Aletraset al\.,[2016](https://arxiv.org/html/2608.10268#bib.bib41)\), who showed that simple classifiers could predict ECtHR Article violations with 79 percent accuracy from case facts alone—work substantially advanced by\(Medvedevaet al\.,[2020](https://arxiv.org/html/2608.10268#bib.bib13)\)’s critical engagement\.\(Chalkidiset al\.,[2019](https://arxiv.org/html/2608.10268#bib.bib43),[2022](https://arxiv.org/html/2608.10268#bib.bib44)\)subsequently formalized this as the ECtHR\-A and ECtHR\-B tasks within LexGLUE\. VECHR introduces a more demanding objective: classifying vulnerability types in ECtHR decisions and VECHR evaluating model explanations against expert rationales\(Xuet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib92)\)\. Most recently, the UDHR Trade\-off Benchmark presents 1152 LLM\-generated scenarios implicating 24 Universal Declaration articles to test how LLMs reason about trade\-offs between rights and competing interests such as public safety or economic stability when using one of eight different evaluation languages\(Samwayet al\.,[2026](https://arxiv.org/html/2608.10268#bib.bib72)\)\. Coverage of substantive international human rights reasoning remains thin relative to domestic legal benchmarks, although this is beginning to change\. UNICEF and UNHCR have teamed up to propose a NeurIPS 2026 competition to to validate AI representations of hard\-to\-reach populations using agency microdata\(Hauptet al\.,[2026](https://arxiv.org/html/2608.10268#bib.bib73)\)\. More directly, Kennedy and Heath propose a new threat model methodology for evaluating model vulnerabilities to producing deepfake media of POWs violative of several POW protections established in the Geneva Conventions \(Kennedy and Heath, forthcoming 2026\)\.

##### Normative, Moral, and Ethical Reasoning\.

Although not legal benchmarks strictly speaking, these kinds of evaluations heavily implicate “precursor” capabilities of interest to legal evaluations and often use legal sources to build evaluation datasets or scoring rubrics\. The ETHICS benchmark introduced systematic evaluation of model alignment with datafied representations of justice, deontology, virtue ethics, utilitarianism, and commonsense morality constructs\(Hendryckset al\.,[2021a](https://arxiv.org/html/2608.10268#bib.bib77)\)\. SCRUPLES extended this to 625,000 ethical judgments over 32,000 real\-life anecdotes\(Lourieet al\.,[2021](https://arxiv.org/html/2608.10268#bib.bib84)\), while Social Chemistry 101 cataloged 292,000 rules\-of\-thumb governing everyday social norms \(Forbes et al\. 2020\)\. Moral Stories tests norm\-consistent and norm\-violating action selection in branching narratives\(Emelinet al\.,[2021](https://arxiv.org/html/2608.10268#bib.bib48)\), and the Moral Integrity Corpus annotates 38,000 dialogue turns with underlying rules of thumb\(Ziemset al\.,[2022](https://arxiv.org/html/2608.10268#bib.bib96)\)\. More recent work has shifted toward multi\-dimensional and adversarial evaluation: MoralBench offers metadata\-rich diagnostic structure, ProMoral\-Bench unifies evaluation across ETHICS, Scruples, and WildJailbreak under a single moral\-safety score\(Subramanianet al\.,[2026](https://arxiv.org/html/2608.10268#bib.bib88)\), and the three\-dimensional LLM Ethics Benchmark assesses foundational principles, reasoning robustness, and value consistency\(Jiaoet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib81)\)\. A persistent finding across this literature is that models align more closely with individualistic Western moral frameworks than with collectivist or non\-Western ones, indicating substantial cultural bias in normative training data\(Jiaoet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib81)\)\.

##### Evaluation Methodology\.

A parallel methodological literature has developed around how to score legal and normative outputs\. LeMAJ decomposes legal answers into atomic ”Legal Data Points” and uses LLM\-as\-judge scoring validated against LegalBench\(Enguehardet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib71)\)\. PLAWBENCH applies rubric\-based evaluation specifically to LLM legal agents\(Shiet al\.,[2026](https://arxiv.org/html/2608.10268#bib.bib55)\)\. LegalEval\-Q uses regression\-based quality assessment across 49 models to evaluate generated legal text\(Li and Wu,[2026](https://arxiv.org/html/2608.10268#bib.bib54)\)\. These approaches respond to a recurring concern that simple accuracy metrics inadequately capture the structured, justification\-dependent character of legal reasoning\(Guhaet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib21); Fanet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib49)\)\.

## Appendix B: Gaps

##### Legal valence\.

Human rights norms carry distinctive legal weight that does not obtain in other legal domains\. Obligations to respect, protect, and fulfill are legally differentiated, and misidentifying which is at stake produces not merely an inaccurate answer but a legally consequential one\. Current safety benchmarks and content policy evaluations are not designed to test this granularity\. Existing guardrails—safety classifiers, RLHF alignment, model cards—address human rights concerns only incidentally, through vague values\-based framing rather than grounding in the actual normative architecture of international law\.

##### Grounding in practice\.

Evaluating models for human rights competencies requires a knowledge of the law and attending normative reasoning practices, but it also requires a scholastic understanding of the application of legal standards to situated facts\. That is, it requires an understanding of norms as well as the developed practice of assessing whether certain actions contribute or degrade the progressive realization of those norms\. Human rights practice therefore requires accounting for unnamed but specific and heavily implicated duty\-bearers, rights\-holders, and institutional remedies in concrete scenarios\. Instance\-level prompting strategies that attempt to elicit human rights reasoning on an ad hoc basis cannot substitute for systematic benchmark\-level measurement\. Synthetic generation of scenario data instead of expert\-creation presents consequential risks to construct validity here\(Beanet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib56); Erikssonet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib69); Schwartzet al\.,[2024](https://arxiv.org/html/2608.10268#bib.bib66),[2025](https://arxiv.org/html/2608.10268#bib.bib57)\)\.

##### Highly developed institutional space\.

International human rights law has a mature, documented interpretive ecosystem, including treaty bodies, General Comments, Special Procedures, regional courts, and an active body of jurisprudence\. This institutional situatedness supplies ample signal towards which to optimize evaluation datasets as well as a demanding standard against which model outputs can be meaningfully assessed\. But this requires benchmark specificity–general legal reasoning or legal knowledge benchmarks will not adequately cover this institutional space or its peculiar dynamics\.

##### Difference from municipal law\.

Unlike domestic legal benchmarks, international human rights law is not the law of any single jurisdiction\. It operates through treaty ratification, state practice, and authoritative interpretation, without a single apex court \(IHL notwithstanding\)\. Reasoning quality cannot be assessed against one jurisdiction’s doctrine; it must be evaluated against principles that apply universally while remaining sensitive to implementation variance and the jurisdiction\-based legal tests specific judges operating within certain municipal traditions may reach for\. This distinguishes it from many existing legal benchmarks, which are grounded in domestic statutory and case law within a single jurisdiction\(Guhaet al\.,[2023](https://arxiv.org/html/2608.10268#bib.bib21)\)\.

##### Unique cross\-jurisdictional variance\.

The progressive realization standard that underpins economic, social and cultural rights in the international human rights regime acknowledges that states implement these rights differently depending on available resources and not only context \(which is of course important for all human rights\)\. A benchmark that encodes a single jurisdiction’s approach to a certain case as ground truth without nominating such a context will systematically misrepresent model performance elsewhere\. Indeed, this is an instance of a broader problem in hierarchical benchmark design: imposing a uniform taxonomy risks obscuring the context\-specific dynamics that determine whether a norm has been violated in practice\. Other rubric\-based approaches to benchmarking in different fields that apply layered taxonomies representing specific, intermediary, and universal criteria have produced better construct validity\(Aroraet al\.,[2025](https://arxiv.org/html/2608.10268#bib.bib65)\)\. Exceedingly few have been attempted in legal benchmarking\.

## Appendix C: Scenarios

![Refer to caption](https://arxiv.org/html/2608.10268v1/Example_Scenario_SubScenario.png)Figure 2:example scenario with sub\-scenario
## Appendix C: Questions

There are 5 types of questions in our list\.

![Refer to caption](https://arxiv.org/html/2608.10268v1/Questions.png)Figure 3:Five questions types in our dataset\.
## Appendix D: Annotator ratings of IRAP questions by scenario

Table 8:Mean IRAP question ratings by scenario\.Note:No ratings were available for Scenarios 6, 7, and 8\.

Similar Articles

DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation

arXiv cs.CL

DLawBench is a new benchmark for evaluating large language models in multi-turn legal consultation, covering Chinese and US law with four client types. Experiments show significant room for improvement, with the best model achieving only 0.562 on legal reasoning.

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

arXiv cs.CL

This paper introduces a cross-evaluation framework for benchmarking LLMs on Arabic cultural and sociolinguistic knowledge, using human SME ground truth and automated judges. The authors contribute a dataset of prompt-rubric pairs for Egyptian and Iraqi Arabic, evaluating frontier LLMs and finding that cultural reasoning remains a primary failure mode for automated grading.