Legal Research Bench:长时程法律研究智能体的端到端可靠性评测

arXiv cs.AI 论文

摘要

论文提出 Legal Research Bench(LRB),一个由专家撰写的包含 413 道题目的开放式美国法律研究基准,并对 13 个前沿模型的智能体进行评测;表现最强的 Claude Opus 4.8 全通过率也仅达 42.9%,表明法律研究智能体仍远未达到可靠水平。

arXiv:2610.00609v1 Announce Type: new Abstract: Legal research is a core and time-consuming legal workflow. Lawyers must identify controlling authority, verify that it remains valid, reconcile statutes and cases, and synthesize a grounded answer. Language model agents are a natural fit for this retrieval-intensive workflow, and automating even part of it would be valuable. But that value depends on reliability: a single missing authority, stale citation, or wrong legal conclusion can make an otherwise plausible answer unusable. We introduce \textbf{Legal Research Bench} (LRB), a benchmark of 413 open-ended U.S. legal research questions written by experts, each paired with a gold answer, supporting authorities, and a binary grading rubric. We evaluate thirteen frontier models in a harness with web search, case-law search, page parsing, and retrieval tools. We score agent responses through all-pass grading with source verification, where a response is correct only if every required criterion is satisfied and its cited authorities verify. We also validate the LLM judge against expert attorneys ensuring that benchmark scores track attorney judgment. Agents remain far from reliable: among the models we tested, the strongest, Claude Opus 4.8, is fully correct on 42.9\% of questions. Performance also varies substantially by task setting: all-pass rates differ across areas of law and are lower on questions requiring reconciliation of conflicting authorities. Across models, more turns, tool calls, and inference cost do not predict higher accuracy.
查看原文
查看缓存全文

缓存时间: 2026/10/02 09:46

# Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents
Source: [https://arxiv.org/html/2610.00609](https://arxiv.org/html/2610.00609)
Oliver ChenLangston NasholdRayan KrishnanAffiliation:Vals AIAffiliation:San Francisco, USA

###### Abstract

Legal research is a core and time\-consuming legal workflow\. Lawyers must identify controlling authority, verify that it remains valid, reconcile statutes and cases, and synthesize a grounded answer\. Language model agents are a natural fit for this retrieval\-intensive workflow, and automating even part of it would be valuable\. But that value depends on reliability: a single missing authority, stale citation, or wrong legal conclusion can make an otherwise plausible answer unusable\. We introduceLegal Research Bench\(LRB\), a benchmark of 413 open\-ended U\.S\. legal research questions written by experts, each paired with a gold answer, supporting authorities, and a binary grading rubric\. We evaluate thirteen frontier models in a harness with web search, case\-law search, page parsing, and retrieval tools\. We score agent responses through all\-pass grading with source verification, where a response is correct only if every required criterion is satisfied and its cited authorities verify\. We also validate the LLM judge against expert attorneys ensuring that benchmark scores track attorney judgment\. Agents remain far from reliable: among the models we tested, the strongest, Claude Opus 4\.8, is fully correct on 42\.9% of questions\. Performance also varies substantially by task setting: all\-pass rates differ across areas of law and are lower on questions requiring reconciliation of conflicting authorities\. Across models, more turns, tool calls, and inference cost do not predict higher accuracy\.

## 1Introduction

Legal research is one of the core tasks through which lawyers turn facts into legal advice, litigation strategy, compliance decisions, and client\-facing work product\. Doing it well requires locating controlling authority, verifying that it remains good law, reconciling conflicting statutes and cases, and synthesizing a grounded answer that withstands scrutiny\. Errors are consequential: a missed exception, stale rule, or wrong holding can change the legal conclusion, so partial correctness does not translate to partial value\.

This workflow is retrieval\-intensive, iterative, and tool\-mediated, making it a natural target for language model agents equipped with web search, case\-law databases, and document parsing\. If reliable, such agents could substantially reduce research time and help attorneys navigate large bodies of law\. The central question is therefore not whether they can retrieve and reason, but whether they can do so reliably enough to be depended upon in practice\.

However, most existing legal benchmarks report average correctness or partial credit, while legal research requires a standard closer to what lawyers need in practice: a legally correct answer supported by real, valid citations\([Guha et al\., 2023](https://arxiv.org/html/2610.00609#bib.bib1);[Pipitone and Alami, 2024](https://arxiv.org/html/2610.00609#bib.bib14);[Fei et al\., 2024](https://arxiv.org/html/2610.00609#bib.bib13);[Li et al\., 2025](https://arxiv.org/html/2610.00609#bib.bib12)\)\. To address this gap, we introduceLegal Research Bench\(LRB\), which counts a response as correct only if it gives the required legal answer and its cited authorities verify\.

Our main contributions are the following:

- •Benchmark\.We introduce LRB, a 413\-question benchmark authored by experts for open\-ended U\.S\. legal research tasks, with gold answers, verified authorities, and binary rubrics\.
- •Harness\.We release a reproducible agentic evaluation harness in which models search the web and case law, parse and fetch pages, recall retrieved content, and submit final answers\. Figure[1](https://arxiv.org/html/2610.00609#S1.F1)gives an overview of the evaluation setup\.
- •Metric\.We evaluate reliability with all\-pass grading: a response is correct only if it satisfies every rubric criterion and every cited authority verifies\.
- •Judge validation\.We validate an LLM judge against expert attorneys before using it to score model outputs\.
- •Empirical results\.We evaluate thirteen frontier models using this harness and find high partial credit but limited all\-pass reliability: the best model attains 42\.9% all\-pass, failures vary systematically across areas of law, reconciliation questions are especially difficult, and additional agentic effort does not predict higher accuracy\.

Figure 1:Overview of the LRB evaluation setup\. Each question is handed to a tool\-using agent that can search the web \(Tavily\) and case law \(CourtListener\), parse and fetch pages, and recall content it has collected into a working database, before submitting a final answer\. Runs are capped at three hours, and the agent manages its own errors\.
## 2Related work

#### Legal benchmarks\.

Most legal NLP benchmarks target static, single\-turn tasks: LegalBench\([Guha et al\., 2023](https://arxiv.org/html/2610.00609#bib.bib1)\), LegalBench\-RAG\([Pipitone and Alami, 2024](https://arxiv.org/html/2610.00609#bib.bib14)\), LawBench\([Fei et al\., 2024](https://arxiv.org/html/2610.00609#bib.bib13)\), and PRBench\([Akyürek et al\., 2025](https://arxiv.org/html/2610.00609#bib.bib15)\)test legal reasoning, retrieval, and domain knowledge, but generally do not measure whether a tool\-using agent can produce a complete final research answer over primary authority\. Recent agent benchmarks move closer: LegalAgentBench\([Li et al\., 2025](https://arxiv.org/html/2610.00609#bib.bib12)\)studies tool use in the Chinese legal domain, and Harvey’s Legal Agent Benchmark\([Grupen et al\., 2026](https://arxiv.org/html/2610.00609#bib.bib16)\), concurrent with our work, evaluates long\-horizon U\.S\. client\-matter tasks under all\-pass scoring\. Its independent adoption of all\-pass grading supports our metric choice, but the task setting differs: Harvey uses bounded matter\-specific document sets, whereas LRB evaluates open\-ended legal research across U\.S\. legal authority that an agent must locate, validate, and synthesize\. Prior work on legal hallucination further motivates authority\-grounded scoring: commercial legal AI tools still hallucinate on 17 to 33 percent of queries\([Magesh et al\., 2025](https://arxiv.org/html/2610.00609#bib.bib2)\), and general\-purpose models fabricate authority far more often\([Dahl et al\., 2024](https://arxiv.org/html/2610.00609#bib.bib3)\)\.

#### Long\-horizon agents\.

Benchmarks such as GAIA\([Mialon et al\., 2024](https://arxiv.org/html/2610.00609#bib.bib8)\), SWE\-bench\([Jimenez et al\., 2024](https://arxiv.org/html/2610.00609#bib.bib9)\), WebArena\([Zhou et al\., 2024](https://arxiv.org/html/2610.00609#bib.bib10)\), andτ\\tau\-bench\([Yao et al\., 2024](https://arxiv.org/html/2610.00609#bib.bib11)\)study multi\-step, tool\-grounded behavior rather than single\-turn accuracy\. These settings make trajectory matching a poor target: different paths can reach the same answer, and similar trajectories can produce different outcomes\. LRB follows this outcome\-level view, scoring the submitted answer rather than the path taken to it\.

#### Judges and measurement validity\.

Grading open\-ended answers at scale increasingly relies on model\-mediated judging\. The LLM\-as\-judge paradigm\([Zheng et al\., 2023](https://arxiv.org/html/2610.00609#bib.bib4)\)shows that strong models can approximate human judgments, and related work has used model judges for criteria\-based open\-ended assessment and rubric\-based grading\([Liu et al\., 2023](https://arxiv.org/html/2610.00609#bib.bib19);[Hashemi et al\., 2024](https://arxiv.org/html/2610.00609#bib.bib20)\)\. But judge scores are not self\-validating: they must agree with human or expert labels for the construct being measured, and their errors must not distort the benchmark’s primary metric\. Measurement\-focused work likewise emphasizes explicit constructs and measurement procedures\([Liang et al\., 2022](https://arxiv.org/html/2610.00609#bib.bib5);[Bowman and Dahl, 2021](https://arxiv.org/html/2610.00609#bib.bib6);[Jacobs and Wallach, 2021](https://arxiv.org/html/2610.00609#bib.bib7)\)\. LRB treats legal research reliability as a construct to be measured directly using all\-pass grading rather than partial credit, and validates its judge against expert attorneys before scoring\.

## 3Legal Research Bench

### 3\.1Data collection

Each LRB item is authored by practicing U\.S\. attorneys and consists of an open\-ended legal research question, a gold answer, a curated source list, and a binary grading rubric \(Section[3\.3](https://arxiv.org/html/2610.00609#S3.SS3)\)\. Items are designed to require legal reasoning rather than factual lookup, are written around novel fact patterns to reduce memorization, and are included only when the correct answer is objective enough for attorney agreement\. When legal authority may change over time, questions are temporally anchored with an explicit answer date \(e\.g\. “Answer as of \[date\]”\), and the agent is prompted to answer as of that date\. Each item is independently reviewed by at least one additional attorney, and ground\-truth materials are written without AI assistance beyond light copyediting\. In total, 15 practicing attorneys, with an average of 15 years of legal experience, contributed to authoring and review, and building each item required approximately four attorney\-hours\.

### 3\.2Dataset

LRB comprises 413 questions spanning eight areas of law and ten question types\. The corpus is partitioned into a 5\-question public sample, a 200\-question validation set, and a 208\-question test set\. We publicly release the five\-question sample together with the full agent harness111[https://github\.com/vals\-ai/legal\-research\-bench](https://github.com/vals-ai/legal-research-bench)to support reproducibility\. The validation and test questions are held out to limit contamination\. Unless noted otherwise, results in this paper are reported over the full 413\-question corpus\.

Questions are multi\-label: 37\.8% carry more than one area of law \(1\.41 areas on average\)\. For compact reporting, we group the ten fine\-grained question types into three overlappingprimary reasoning categoriesand twodifficulty attributes\. The primary reasoning categories describe the main body of authority a question turns on: Statutory Interpretation, Doctrinal Rule Reasoning, or Regulatory Framework Interpretation\. The difficulty attributes mark questions that are harder for a reason independent of the primary category: Reconciliation, where a question requires synthesizing multiple or conflicting authorities, and Temporal Validity, where the answer depends on whether a rule is still good law\. Because these labels overlap, they are descriptive marginals rather than a partition; Appendix[C](https://arxiv.org/html/2610.00609#A3)lists the ten underlying question types and their mapping to this scheme\. Table[1](https://arxiv.org/html/2610.00609#S3.T1)gives the full composition\.

Each question carries a curated source list \(6\.2 sources on average, median 6, up to 23\) and a binary rubric \(9\.4 items on average, median 9, up to 31\)\. The gold\-standard answers average 536 words \(median 429\) in length\.

Table 1:Dataset composition over all 413 questions: distribution across the eight areas of law \(left\) and the question\-type scheme of three overlapping primary reasoning categories and two difficulty attributes \(right\)\. Questions are multi\-label, so the counts sum to more than 413\.
### 3\.3Rubric and metrics

Each question is graded against binary rubric criteria that the authoring attorneys judged individually necessary and jointly sufficient for a correct answer\. Satisfying every criterion is therefore equivalent to producing a correct answer by that expert standard\. Criteria are weighted \[\+1\], \[\+2\], or \[\+3\], with at least one \[\+3\] item capturing the central legal conclusion\. We reportall\-passas the primary metric: a response is correct only if*every*rubric criterion is satisfied\.Weighted pass, the share of available rubric weight a response earns, is reported only as a secondary, partial\-credit measure\. Grounding is part of correctness: the rubric specifies which authorities a correct answer must cite, and any invalid cited URL fails asource checkthat marks the entire response wrong\. Following recent recommendations to report statistical uncertainty in language model evaluations\([Miller, 2024](https://arxiv.org/html/2610.00609#bib.bib21)\), we report standard errors for leaderboard estimates and confidence intervals clustered by question for category\-level analyses; Appendix[C](https://arxiv.org/html/2610.00609#A3)details the procedure\.

## 4Experimental setup

### 4\.1Agents

Models are evaluated as tool\-using agents \(Figure[1](https://arxiv.org/html/2610.00609#S1.F1)\)\. Each agent has access to five tools\. Web search via Tavily\([Tavily Inc\.,](https://arxiv.org/html/2610.00609#bib.bib18)\)returns links with short excerpts\. Case\-law search via CourtListener\([Free Law Project,](https://arxiv.org/html/2610.00609#bib.bib17)\)finds court opinions by keyword, by natural language, or by citation lookup, and fetches their full text; page parsing does the same for any other web page\. Both store what they fetch under a key the agent chooses, and retrieval then answers a prompt against the stored documents, so an opinion too long to hold in context can still be used\. A final tool submits the answer\. Runs are capped by a three\-hour wall\-clock budget per question rather than by a fixed number of turns\. A time budget better reflects real research workflows and makes the evaluation comparable across alternative harnesses, tools, and agent implementations\. The cap is far from binding in practice: no fully correct \(all\-pass\) answer took longer than 1\.9 hours, and 95% of all runs complete within roughly 50 minutes\. We evaluateN=13N=13frontier models on the corpus, all released between February and June 2026\.

### 4\.2Judge calibration

Since every score in LRB is assigned by a single LLM judge, we validate the judge against expert attorneys before using it for model evaluation\. In a held\-out calibration study, three attorneys and three candidate judges \(Claude Sonnet 4\.6, Gemini 3\.1 Pro Preview, and GPT\-5\.4\) scored 341 rubric items drawn from 45 model responses to 15 questions\. The attorney majority serves as the reference label, and human inter\-rater agreement sets the baseline: attorneys agree with the majority label on 83\.4% of rubric items, with Cohen’sκ=0\.644\\kappa=0\.644\. All three candidate judges meet or exceed this baseline, withκ\\kapparanging from 0\.708 to 0\.726\.

We select GPT\-5\.4 as the benchmark judge for two reasons\. First, its pass rate is closest to the attorney majority, reducing the risk that benchmark scores are systematically shifted by a judge that is much stricter or more lenient than expert graders\. Second, its false positives and false negatives are more balanced than those of the other candidate judges, at 5\.6% and 7\.0%, respectively\. This balance matters under all\-pass grading, where one\-sided item\-level errors can compound across multi\-item rubrics\. Aggregation at the task level does not degrade this agreement: translating GPT\-5\.4’s per\-item verdicts into the all\-pass score, it agrees with the expert majority on 86\.7% of responses \(κ=0\.710\\kappa=0\.710\), close to the human inter\-rater baseline of 87\.9% \(κ=0\.735\\kappa=0\.735\)\. Table[2](https://arxiv.org/html/2610.00609#S4.T2)reports the calibration results, and Appendix[B](https://arxiv.org/html/2610.00609#A2)gives additional robustness checks\.

Table 2:Judge–expert agreement vs\. the human inter\-rater baseline, over 341 rubric\-item comparisons\. Agree andκ\\kappascore each judge against the attorney majority; FP and FN are the false positive and false negative rates, respectively; Pass is the share of items a rater marks pass\. The Human row is pairwise attorney\-to\-attorney, so FP and FN are undefined, and its Pass entry is the majority’s\. All three judges exceed the human agreement baseline; GPT\-5\.4 is selected for its balanced error profile and high agreement with the human expert majority\.

## 5Results

### 5\.1Reliability remains limited

Table[3](https://arxiv.org/html/2610.00609#S5.T3)reports overall performance across the full corpus\. Claude Opus 4\.8 achieves the highest all\-pass rate at 42\.9%, followed by GPT\-5\.5 at 40\.0% and Claude Sonnet 4\.6 at 38\.5%; the weakest model reaches 12\.6%\. Weighted pass is much higher, from 68% to 86%: agents recover most of an answer while still missing at least one element it requires\. The distance, between nearly correct and complete, can be crucial in a legal setting\.

Table 3:Leaderboard over the full 413\-question corpus, sorted by all\-pass\. Pass columns are mean±\\pmstandard error \(%\); Sources is the mean number of authorities the model cites ; Turns and Tools are mean agent turns and tool calls per question; $/test is mean cost per question in USD\. Appendix[D](https://arxiv.org/html/2610.00609#A4)compares adjacent entries on paired per\-question outcomes\.
### 5\.2Partial correctness hides substantive omissions

All\-pass failures are not confined to minor omissions\. Across the corpus, models miss at least one high\-value \[\+3\] criterion on 29\.5% of questions, and these misses account for roughly 41% of all failures \(Appendix[E](https://arxiv.org/html/2610.00609#A5), Figure[9](https://arxiv.org/html/2610.00609#A5.F9)\)\. The remaining failures are also often substantive: lower\-weight criteria require a specific statute, rule, controlling authority, or legal interpretation\. A further 4\.3% of runs satisfy every rubric criterion but fail the source check, because a cited URL is invalid or does not point to the authority claimed\. Thus, the gap between weighted pass and all\-pass reflects missing required elements of the legal research answer, not a long tail of minor details\.

### 5\.3Reliability varies systematically

Reliability varies by practice area\. Averaged across models, all\-pass ranges from 44\.2% in Health and 32\.2% in Administrative/Regulatory to 14\.9% in Family, with most other areas in the low\-to\-mid 20s \(Figure[2](https://arxiv.org/html/2610.00609#S5.F2)\)\. No model is strong everywhere: the best model in the hardest area reaches only 24\.4% all\-pass\. Difficulty also varies by question type \(Figure[3](https://arxiv.org/html/2610.00609#S5.F3)\)\. Reconciliation questions, which require synthesizing multiple or conflicting authorities, pass at 20\.4% \(95% CI \[16\.1, 24\.9\]\) versus 29\.9% \(\[26\.4, 33\.4\]\) for questions without the attribute, a gap of 9\.4 percentage points, a penalty that holds for every model, while temporal\-validity questions sit near the benchmark average at 28\.7%\. The reconciliation penalty is not an artifact of rubric length: reconciliation questions carry slightly fewer criteria than the rest, and the effect persists in a logistic model controlling for rubric length and source count\. Appendix[C](https://arxiv.org/html/2610.00609#A3)gives the per\-model breakdown and the regression with these controls\.

![Refer to caption](https://arxiv.org/html/2610.00609v1/area_heatmap.png)Figure 2:All\-pass rate by area of law\.Figure 3:Pooled all\-pass rate by primary reasoning category \(blue\) and difficulty attribute \(red\), with 95% confidence intervals\. The reasoning categories cluster near the overall rate, whereas reconciliation is markedly harder\.
### 5\.4More turns, tool calls, and cost do not predict reliability

Higher turn counts, inference cost, and tool\-call counts do not reliably predict correctness across models \(Figure[4](https://arxiv.org/html/2610.00609#S5.F4); tool calls in Appendix[E](https://arxiv.org/html/2610.00609#A5), Figure[8](https://arxiv.org/html/2610.00609#A5.F8)\)\. Claude Opus 4\.8 achieves the highest all\-pass rate while using the fewest turns of any model \(11\.7\), whereas Kimi K2\.6 averages nearly 100 turns and 112 tool calls but reaches only 15\.5% all\-pass\. Cost shows the same pattern: GPT\-5\.5 is the most expensive run \($7\.34 per question\) yet trails Claude Opus 4\.8, which scores higher at roughly 40% of the cost\. Per\-question trajectories \(Appendix[G](https://arxiv.org/html/2610.00609#A7)\) instead show stronger agents reading and verifying authority rather than repeatedly issuing searches\. Appendix[F](https://arxiv.org/html/2610.00609#A6)reports answer length, the one effort measure that does correlate with all\-pass\.

Figure 4:All\-pass against inference cost \(left\) and mean agent turns \(right\)\. The red line on the cost panel marks the Pareto frontier: the models for which no other model is both cheaper and more accurate\. Tool calls are plotted separately in Appendix[E](https://arxiv.org/html/2610.00609#A5), Figure[8](https://arxiv.org/html/2610.00609#A5.F8)\.
### 5\.5Tool access is necessary but not sufficient

In a no\-tool, single\-shot setting, performance collapses across all thirteen models \(Table[4](https://arxiv.org/html/2610.00609#S5.T4)\): all\-pass falls by 22\.5 points on average, and every tool\-using model outperforms every no\-tool model\. The strongest models lose the most \(Claude Opus 4\.8 falls 34\.1 points\), indicating that their advantage comes from locating and reading authority rather than from memorized law\. Combined with Section[5\.4](https://arxiv.org/html/2610.00609#S5.SS4), this indicates that legal research ability rests on tool\-grounded retrieval and verification, yet once tools are available, using them moredoes not predict higher reliability\.

Table 4:Ablation over the full 413\-question corpus: each model run with its full tool environment versus single\-shot with no tool access, sorted by all\-pass with tools\. All values are percentages\. Removing tools lowers all\-pass by 22\.5 points and weighted pass by 25\.6 points on average, and every tool\-using model outscores every single\-shot one\.

## 6Limitations and future work

LRB is limited to U\.S\., English\-language legal research and to a fixed snapshot of models, tools, and legal authority\.222A live leaderboard is maintained at[https://www\.vals\.ai/benchmarks/legal\_research](https://www.vals.ai/benchmarks/legal_research)\.Its questions are also designed to have objective answers, which is necessary for all\-pass grading but excludes parts of legal work that involve strategy, judgment under uncertainty, client counseling, negotiation, or drafting choices with no single correct answer\. Although we validate the LLM judge against expert attorneys before using it for scoring, the reported results still depend on a model\-mediated grading procedure rather than full expert adjudication of every response\. Finally, each model was evaluated in a single run per question, so the reported rates carry sampling uncertainty over questions but not run\-to\-run variance\. Future work should extend this evaluation design to additional jurisdictions, languages, and legal tasks; study how reliability changes under different tool environments and human\-in\-the\-loop workflows; and develop evaluation protocols that measure not only final\-answer correctness, but also when agents appropriately surface uncertainty, ask for clarification, or defer to human review\.

## 7Conclusion

LRB makes end\-to\-end legal research reliability measurable by pairing all\-pass grading of final answers with judge validation against expert attorneys\. Under this evaluation, the frontier tool\-using models we test often earn substantial partial credit, but the strongest model at the time of evaluation reaches only 42\.9% all\-pass\. The observed failures are systematic rather than incidental: all\-pass rates vary across practice areas, and reconciliation tasks are especially difficult\. Across models, higher turn counts and tool\-call counts do not reliably predict higher accuracy\. LRB therefore suggests that legal agents should not be judged by whether they can produce plausible research, but by whether they can consistently produce work that lawyers could safely rely on\.

## Acknowledgments

We thank the practicing attorneys who authored and reviewed the LRB questions, gold answers, and rubrics, and took part in the judge calibration study\.

## Ethics Statement

Legal research is high\-stakes and highly sensitive work: errors can affect liberty, livelihood, and legal rights, and a single fabricated authority or misstated holding can cause real harm\. Our central finding is that current agents are not yet reliable at this task, and we present LRB as a measurement of that gap rather than as an endorsement of any system for practice\. The benchmark is not a certification of fitness for deployment, and its results should not be read as clearing models for unsupervised legal use; agent outputs are not legal advice and should be reviewed by a licensed attorney before they are relied upon\. We caution in particular against automation bias, since the fluent, well\-structured answers these agents produce can appear authoritative even when the central legal conclusion is wrong\.

The benchmark itself raises no human\-subjects or privacy concerns\. All questions are original hypotheticals authored by the participating attorneys around novel fact patterns; they contain no real client information, no personally identifiable information, and no privileged or confidential material\. The attorneys who wrote and reviewed the items did so as compensated contributors and consented to the use of their work\. We release a small public sample and the evaluation harness to support scrutiny and reproducibility, while holding out the remaining questions to limit contamination\.

## References

- Akyüreket al\.\(2025\)A\. F\. Akyürek, A\. Gosai, C\. B\. C\. Zhang, V\. Gupta, J\. Jeong, A\. Gunjal, T\. Rabbani, M\. Mazzone, D\. Randolph, M\. M\. Meymand,et al\.PRBench: large\-scale expert rubrics for evaluating high\-stakes professional reasoning\.arXiv preprint arXiv:2511\.11562\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px1.p1.1)\.
- Bowman and Dahl \(2021\)S\. Bowman and G\. DahlWhat will it take to fix benchmarking in natural language understanding?\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 4843–4855\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px3.p1.1)\.
- Dahlet al\.\(2024\)M\. Dahl, V\. Magesh, M\. Suzgun, and D\. E\. HoLarge legal fictions: profiling legal hallucinations in large language models\.Journal of Legal Analysis16\(1\),pp\. 64–93\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px1.p1.1)\.
- Feiet al\.\(2024\)Z\. Fei, X\. Shen, D\. Zhu, F\. Zhou, Z\. Han, A\. Huang, S\. Zhang, K\. Chen, Z\. Yin, Z\. Shen, J\. Ge, and V\. NgLawBench: benchmarking legal knowledge of large language models\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Miami, Florida, USA,pp\. 7933–7962\.Cited by:[§1](https://arxiv.org/html/2610.00609#S1.p3.1),[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px1.p1.1)\.
- \[5\]Free Law ProjectCourtListener\.Note:[https://www\.courtlistener\.com/](https://www.courtlistener.com/)Accessed: 2026\-06\-24Cited by:[§4\.1](https://arxiv.org/html/2610.00609#S4.SS1.p1.1)\.
- Grupenet al\.\(2026\)N\. Grupen, G\. Pereyra, and J\. PereyraHarvey’s legal agent benchmark \(LAB\): a long\-horizon benchmark for legal agents\.Note:[https://www\.harvey\.ai/blog/introducing\-harveys\-legal\-agent\-benchmark](https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmark)Open\-source benchmark;[https://github\.com/harveyai/harvey\-labs](https://github.com/harveyai/harvey-labs)Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px1.p1.1)\.
- Guhaet al\.\(2023\)N\. Guha, J\. Nyarko, D\. E\. Ho, C\. Ré, A\. Chilton,et al\.LegalBench: a collaboratively built benchmark for measuring legal reasoning in large language models\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track,Note:arXiv:2308\.11462Cited by:[§1](https://arxiv.org/html/2610.00609#S1.p3.1),[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px1.p1.1)\.
- Hashemiet al\.\(2024\)H\. Hashemi, J\. Eisner, C\. Rosset, B\. Van Durme, and C\. KedzieLlm\-rubric: a multidimensional, calibrated approach to automated evaluation of natural language texts\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13806–13834\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px3.p1.1)\.
- Jacobs and Wallach \(2021\)A\. Z\. Jacobs and H\. WallachMeasurement and fairness\.InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency,pp\. 375–385\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px3.p1.1)\.
- Jimenezet al\.\(2024\)C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. NarasimhanSwe\-bench: can language models resolve real\-world github issues?\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 54107–54157\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px2.p1.1)\.
- Liet al\.\(2025\)H\. Li, J\. Chen, J\. Yang, Q\. Ai, W\. Jia, Y\. Liu, K\. Lin, Y\. Wu, G\. Yuan, Y\. Hu,et al\.LegalAgentBench: evaluating llm agents in legal domain\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 2322–2344\.Cited by:[§1](https://arxiv.org/html/2610.00609#S1.p3.1),[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px1.p1.1)\.
- Lianget al\.\(2022\)P\. Liang, R\. Bommasani, T\. Lee, D\. Tsipras, D\. Soylu, M\. Yasunaga, Y\. Zhang, D\. Narayanan, Y\. Wu, A\. Kumar,et al\.Holistic evaluation of language models\.arXiv preprint arXiv:2211\.09110\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px3.p1.1)\.
- Liuet al\.\(2023\)Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. ZhuG\-eval: nlg evaluation using gpt\-4 with better human alignment\.InProceedings of the 2023 conference on empirical methods in natural language processing,pp\. 2511–2522\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px3.p1.1)\.
- Mageshet al\.\(2025\)V\. Magesh, F\. Surani, M\. Dahl, M\. Suzgun, C\. D\. Manning, and D\. E\. HoHallucination\-free? assessing the reliability of leading AI legal research tools\.Journal of Empirical Legal Studies22,pp\. 216–242\.Note:arXiv:2405\.20362, 2024Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px1.p1.1)\.
- Mialonet al\.\(2024\)G\. Mialon, C\. Fourrier, T\. Wolf, Y\. LeCun, and T\. ScialomGAIA: a benchmark for general ai assistants\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 9025–9049\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px2.p1.1)\.
- Miller \(2024\)E\. MillerAdding error bars to evals: a statistical approach to language model evaluations\.arXiv preprint arXiv:2411\.00640\.Cited by:[§3\.3](https://arxiv.org/html/2610.00609#S3.SS3.p1.1)\.
- Pipitone and Alami \(2024\)N\. Pipitone and G\. H\. AlamiLegalbench\-rag: a benchmark for retrieval\-augmented generation in the legal domain\.arXiv preprint arXiv:2408\.10343\.Cited by:[§1](https://arxiv.org/html/2610.00609#S1.p3.1),[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px1.p1.1)\.
- \[18\]Tavily Inc\.Tavily: web search api for ai agents\.Note:[https://www\.tavily\.com/](https://www.tavily.com/)Accessed: 2026\-06\-24Cited by:[§4\.1](https://arxiv.org/html/2610.00609#S4.SS1.p1.1)\.
- Yaoet al\.\(2024\)S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhanτ\\tau\-Bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.arXiv preprint arXiv:2406\.12045\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px2.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in Neural Information Processing Systems36,pp\. 46595–46623\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px3.p1.1)\.
- Zhouet al\.\(2024\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried,et al\.Webarena: a realistic web environment for building autonomous agents\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 15585–15606\.Cited by:[§2](https://arxiv.org/html/2610.00609#S2.SS0.SSS0.Px2.p1.1)\.

## Appendix ASample question

Below is a sample question from the public dataset split, P\-005 \(Business & Commercial; Statutory Interpretation\), used as the running example in Figure[11](https://arxiv.org/html/2610.00609#A7.F11)\. This question is graded against a binary rubric of 12 weighted criteria\.

#### Question\.

Client is a Virginia corporation \(“Client”\) that hauls refrigerated freight out of Rockbridge County\. In early 2024, Client needed two new reefer trailers and went to a national manufacturer \(“Manufacturer”\) to pick out a model and negotiate the refrigeration specs\. Client then arranged for a Virginia equipment financing company \(“Lessor”\) to buy the trailers \(“Trailers”\) from Manufacturer for $85,000 each and lease them to Client under a written 60\-month lease \(“Lease”\) at $2,200 per month per trailer, and Lessor gave Client a complete copy of its purchase agreement with Manufacturer, which included Manufacturer’s standard two\-year warranty on all mechanical and refrigeration components, before Client signed\.

Four months after delivery, both Trailers had compressor failures that left the units unable to hold temperature, and Manufacturer agreed to fix them under the warranty but filed for Chapter 7 two months later, where the trustee confirmed it wouldn’t honor outstanding warranty claims\. That said, the Lease doesn’t include any warranties from Lessor and tells Client to take equipment claims to Manufacturer\. It also says that Client’s payments are “absolute and unconditional, irrespective of any defense or any right of setoff, counterclaim, or recoupment\.” Client has stopped paying Lessor\. Given these facts, can Client stop its payments to Lessor, and is Lessor responsible for the defective Trailers?

## Appendix BJudge calibration: additional analyses

#### GPT\-5\.4 versus the other judge candidates

Gemini 3\.1 Pro Preview attains the top agreement \(κ=0\.726\\kappa=0\.726\), marginally above GPT\-5\.4 \(0\.7150\.715\), but we do not select onκ\\kappaalone\. GPT\-5\.4 has the most balanced error profile, with a false\-positive to false\-negative ratio of about1:1\.31\{:\}1\.3, versus roughly1:3\.21\{:\}3\.2to1:3\.31\{:\}3\.3for the stricter Claude and Gemini judges \(Table[2](https://arxiv.org/html/2610.00609#S4.T2)\); a balanced profile matters under all\-pass grading, where one\-sided item\-level errors compound across a multi\-item rubric\. Its pass rate \(66\.3%\) is also closest to the expert majority \(67\.7%\)\. Because the selected judge remains mildly strict \(false negatives modestly exceed false positives\), reported all\-pass rates are a slightly conservative estimate of human\-graded performance\.

Beyond the headline agreement in Section[4\.2](https://arxiv.org/html/2610.00609#S4.SS2), the alignment study supports the selected judge in two additional ways\. First, the agreement is not an artifact of easy questions: excluding four near\-uniform questions whose constant labels makeκ\\kappaundefined leaves 308 comparisons on the harder items, and every judge’sκ\\kappafalls by less than one point \(GPT\-5\.4:0\.715→0\.7050\.715\\to 0\.705\)\. Second, there is no same\-family leniency: each judge’s*lowest*agreement is on models from its own provider\.

## Appendix CFine\-grained question types and adjusted category effects

Every question is answered by all thirteen models, so observations are clustered within questions\. Pooled category\-level and area\-level rates use 95% confidence intervals from a question\-level cluster bootstrap, resampling questions with replacement and carrying all thirteen of a resampled question’s model outcomes with it\. The intraclass correlation of all\-pass within question is 0\.40, a design effect of 5\.8 at thirteen models\. Per\-model standard errors in Table[3](https://arxiv.org/html/2610.00609#S5.T3)are computed directly, since each model contributes one observation per question\.

In our analysis, ten fine\-grained question types are grouped into three overlapping primary reasoning categories \(the primary source of authority a question turns on\) and two difficulty attributes \(orthogonal sources of difficulty\)\. Table[5](https://arxiv.org/html/2610.00609#A3.T5)lists the ten types, the category or attribute each maps to, and the pooled all\-pass rate over the full 413\-question corpus\. Questions are multi\-label, so counts sum to more than 413 and the groups are descriptive marginals rather than a partition\.

Question typeMaps toNAll\-pass95% CI*Primary reasoning categories*Statutory InterpretationStatutory28825\.2\[21\.9, 28\.5\]Application of Precedent to a Fact PatternDoctrinal24023\.3\[20\.1, 26\.6\]Doctrinal Test IdentificationDoctrinal11624\.3\[19\.5, 29\.3\]Regulatory Framework InterpretationRegulatory9131\.0\[24\.6, 37\.6\]*Difficulty attributes*Cross\-Jurisdiction ComparisonReconciliation4821\.5\[14\.6, 28\.8\]Consensus vs\. Disagreement Across CourtsReconciliation2717\.9\[8\.5, 28\.2\]Resolving Conflicting PrecedentReconciliation2525\.6\[14\.2, 37\.8\]Interaction of Multiple Legal RegimesReconciliation7121\.4\[16\.0, 27\.2\]Validity of PrecedentTemporal Validity4229\.0\[20\.0, 38\.7\]Timeline / Doctrinal EvolutionTemporal Validity4527\.2\[18\.3, 36\.5\]Table 5:The ten question types, their mapping to the category/attribute scheme used in Section[5\.3](https://arxiv.org/html/2610.00609#S5.SS3), and pooled all\-pass over the full corpus \(percentages, with question\-level cluster\-bootstrap 95% confidence intervals, which account for every question being answered by all thirteen models\)\. N is the number of questions carrying each type\.Figure[5](https://arxiv.org/html/2610.00609#A3.F5)breaks the pooled rates of Section[5\.3](https://arxiv.org/html/2610.00609#S5.SS3)down by model, across the three primary reasoning categories and the two difficulty attributes\. The categories differ little across models, all sitting near the 27\.2% overall rate, while the Reconciliation attribute is consistently the clearest difficulty signal\.

![Refer to caption](https://arxiv.org/html/2610.00609v1/category_heatmap.png)Figure 5:Per\-model all\-pass rate by primary reasoning category and difficulty attribute\.Rubric length varies across questions and is itself a predictor of all\-pass, so here we check whether the reconciliation penalty reported in Section[5\.3](https://arxiv.org/html/2610.00609#S5.SS3)is driven by it\. Pooled across all thirteen models, all\-pass falls from 41\.8% on questions carrying six or fewer criteria to 15\.2% on questions carrying thirteen or more, a rank correlation of−0\.287\-0\.287between a question’s criterion count and its all\-pass rate\. To test whether this accounts for the penalty, we regress all\-pass on a reconciliation indicator, rubric length, and source count, with area\-of\-law indicators as further controls\. All eight areas enter without an omitted reference, since questions can carry more than one area and the indicators therefore do not sum to a constant\. Table[6](https://arxiv.org/html/2610.00609#A3.T6)compares the unadjusted and adjusted odds ratios\.

Table 6:All\-pass odds ratios for reconciliation before and after adjusting for rubric length and source count\. Area\-of\-law indicators are included in the adjusted model as controls but are not reported\. Rubric length and source count enter only the adjusted model\.An odds ratio of one would mean no difference\. Adjusting for rubric length and source count moves the reconciliation effect from 0\.60 to 0\.68, so the penalty weakens but stays below one \(coefficient−0\.380\-0\.380, 95% CI\[−0\.73,−0\.04\]\[\-0\.73,\-0\.04\]\)\. Both controls predict all\-pass in their own right, at 0\.89 per additional rubric criterion and 0\.93 per additional required source, yet neither accounts for the penalty: reconciliation questions carry 9\.0 criteria on average against 9\.5 for the rest, so the confound runs against the observed penalty rather than with it\.

## Appendix DPaired model comparisons

Every model answers every question, so adjacent leaderboard entries can be compared on paired per\-question outcomes rather than through independent standard errors\. Table[7](https://arxiv.org/html/2610.00609#A4.T7)reports McNemar’s test for each of the twelve adjacent pairs in Table[3](https://arxiv.org/html/2610.00609#S5.T3)\. The discordant counts give the number of questions the higher\-ranked model passes and the lower\-ranked model fails, followed by the reverse\. Only two adjacent pairs separate at the 5% level, so the leaderboard supports tiers rather than a strict ordering\.

Table 7:Adjacent\-pair comparisons in leaderboard order, over the full 413\-question corpus\. Discordant counts are higher\-ranked\-only passes / lower\-ranked\-only passes\. Boldpp\-values mark the two pairs that separate at the 5% level\.
## Appendix EAdditional results

#### Best model per area of law\.

Figure[6](https://arxiv.org/html/2610.00609#A5.F6)reports, for each area of law, the best single\-model all\-pass rate, summarizing the heatmap of Figure[2](https://arxiv.org/html/2610.00609#S5.F2)and showing that no single model leads across all areas\.

Figure 6:Best per\-model all\-pass by area of law\.
#### Tool\-usage composition\.

Figure[7](https://arxiv.org/html/2610.00609#A5.F7)breaks down each model’s tool calls by type\. Web search is the most\-used tool overall \(about 27 calls per session against about 11 for case\-law search\), but stronger models allocate more selectively: Claude Opus 4\.8 balances case\-law and web search in roughly equal proportion \(about 8–9 calls each\)\. The heaviest searchers sit at the bottom of the leaderboard, Kimi K2\.6 at 86\.4 web searches per question and GPT\-5\.4\-mini at 54\.5, but call volume does not otherwise track accuracy: Gemini 3\.1 Pro Preview issues the fewest calls of any model and ranks tenth\.

Figure 7:Tool\-usage composition by model\. Totals at the right are mean tool calls per question\.
#### Tool calls and correctness\.

Figure[8](https://arxiv.org/html/2610.00609#A5.F8)plots all\-pass against mean tool calls per question, as mentioned in Section[5\.4](https://arxiv.org/html/2610.00609#S5.SS4)\. Volume does not track correctness: Claude Opus 4\.8 leads the benchmark on 28\.7 calls per question, while Kimi K2\.6 issues 112\.3 and GPT\-5\.4\-mini 94\.4\.

Figure 8:All\-pass against mean tool calls per question\.
#### Failure decomposition\.

Figure[9](https://arxiv.org/html/2610.00609#A5.F9)classifies each evaluated question as all\-pass, a lower\-tier\-only failure \(errs solely on \[\+1\]/\[\+2\] items\), or central\-wrong \(misses a high\-value \[\+3\] item\)\. Central errors make up a large share of every model’s failures, accounting for roughly 41% of all failures across the corpus\.

Figure 9:Failure decomposition by model\. Each evaluated question is correct \(all\-pass\), fails only on lower\-tier \[\+1\]/\[\+2\] items, or is central\-wrong \(misses a high\-value \[\+3\] item\)\. Central errors \(right, red\) make up a large share of every model’s failures; the percentage of all questions that are central\-wrong is annotated\.

## Appendix FAnswer length and verbosity

Most models produce answers far longer than the gold\-standard references, which average 536 words\. Claude Sonnet 4\.6 is the most verbose at roughly 2,320 words, more than four times the reference length, and only GPT\-5\.4\-mini \(about 350 words\) falls below it\. Answer length is positively associated with accuracy across models \(Spearmanρ≈0\.7\\rho\\approx 0\.7with all\-pass; Figure[10](https://arxiv.org/html/2610.00609#A6.F10)\), though this appears to track underlying capability rather than verbosity itself: GPT\-5\.5 attains the second\-highest all\-pass with one of the shortest answers \(about 990 words\)\. The sharper contrast is with the expert references, which reach a complete and correct answer in far fewer words than any strong agent\. A reading of the public\-set answers indicates that this extra length is largely structural and expository rather than substantive: agents organize their answers with section headers and an issue\-rule\-application\-conclusion structure, walk through each statutory element explicitly even when it is uncontested, and recite background frameworks and citation detail that the question does not require, whereas the expert references state the holding and the controlling authorities directly\.

Figure 10:Mean answer length versus all\-pass per model, with the 536\-word gold reference marked\. Answer length is positively associated with all\-pass, yet every agent at the top of the leaderboard is several times more verbose than the expert references\.
## Appendix GAgent trajectories

To illustrate how models allocate effort, we trace the ordered sequence of tool calls on one question, P\-005 \(the Virginia equipment\-leasing problem under UCC Article 2A presented in Appendix[A](https://arxiv.org/html/2610.00609#A1)\), bucketing each call by type \(Figure[11](https://arxiv.org/html/2610.00609#A7.F11)\)\. Trajectory length varies from 19 calls \(Gemini 3\.1 Pro\) to 174 \(Kimi K2\.6\), while the five models that get a full score on the question cluster between 37 and 45 calls\. Claude Opus 4\.8 answers in 38 calls: web and case\-law search to frame the issue, CourtListener to read the controlling provisions and decisions, page reading and retrieval to pull specific passages, then a single submission\. It cites eight authorities, all of which verify, but misses one rubric criterion\. Kimi K2\.6 issues 174 calls, 117 of them web searches, and satisfies every criterion, but cites no authority and so fails the source check\. Grok 4\.3 exposes a different behavior: it never queries CourtListener, working from web search alone, and cites three authorities for 76% of the rubric weight\. On this question the three fail for distinct reasons: one grounded but incomplete, the other complete but ungrounded, the third never reaching primary authority\.

Figure 11:Per\-model tool\-call trajectories on the sample question P\-005, each call bucketed by type and shown in the order it was issued; bar length is the number of tool calls\. Trajectory length varies more than tenfold and does not track all\-pass performance \(models are ordered top\-to\-bottom by corpus\-wide all\-pass\)\.

相似文章

ResearchClawBench:面向端到端自主科学研究的基准测试

Hugging Face Daily Papers

ResearchClawBench 是一个用于评估端到端自主科学研究的基准测试,涵盖来自10个领域的40个任务,结果显示当前AI智能体和LLM的重新发现准确率较低,其中Claude Code平均得分为21.5,Claude-Opus-4.7平均得分为20.7(在可能的总分中)。

DLawBench:通过多轮法律咨询评估大语言模型

arXiv cs.CL

DLawBench是一个新的基准测试,用于评估大语言模型在多轮法律咨询中的表现,涵盖中国和美国法律,包含四种客户类型。实验表明仍有很大改进空间,最佳模型在法律推理上仅达到0.562。