Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant
Summary
This position paper argues that legal LLM hallucinations should be evaluated as failures of legal warrant rather than factual inaccuracies, proposing a new benchmark framework for assessing legal AI systems.
View Cached Full Text
Cached at: 09/17/26, 08:47 AM
# Position: Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant
Source: [https://arxiv.org/html/2609.17546](https://arxiv.org/html/2609.17546)
###### Abstract
In this position paper, we argue that legal LLMs’ hallucinations should be evaluated as a failure of legal warrant rather than as factual inaccuracy or citation failure\. We define claim\-authority warrant as the context\-sensitive relation between a consequential legal claim and authority that exists, applies to the relevant jurisdiction, is current for the date of analysis, has the legal status represented by the system, and supports the proposition asserted\. Warranted legal generation is the broader system behavior that answers, narrows, asks, warns, corrects a false premise, or abstains according to that relation\. The falsifiable prediction is that warrant metrics reveal material failures that answer accuracy, citation existence, generic attribution, LegalHalBench\-style statute relevance, and CitaLaw\-style sentence\-citation alignment can miss\. We sharpen this claim with a side\-by\-side comparison item and a small, reproducible pilot over public\-rule tests\. We then specify benchmark records, claim boundaries, support labels, mixed response\-policy scoring, risk weights, annotation reliability reporting, and jurisdiction\-specific authority ontologies\. The result is a concrete research agenda for evaluating legal AI systems by whether their consequential claims are licensed by law\.
legal AI, hallucination, retrieval\-augmented generation, evaluation, access to justice
## 1Introduction
Legal hallucination is often introduced as the invention of cases\. That is the simplest failure, not the whole problem\. InMata v\. Avianca, lawyers filed non\-existent cases and quotations generated by ChatGPT, leading to Rule 11 sanctions\(United States District Court for the Southern District of New York,[2023](https://arxiv.org/html/2609.17546#bib.bib14)\)\. More recent public incidents include filings or proposed orders with inaccurate citations, misstatements, fictitious or misattributed authority, and AI\-assisted drafting failures\(Freifeld and Scarcella,[2026](https://arxiv.org/html/2609.17546#bib.bib17); Scarcella,[2026](https://arxiv.org/html/2609.17546#bib.bib18)\)\. These events mix fabrication with a subtler defect: a legal proposition can be attached to a real source and still not be licensed by that source\.
The research evidence points in the same direction\. General\-purpose LLMs hallucinate on legal knowledge questions, vary across courts and time periods, accept false premises, and often lack calibrated awareness of their errors\(Dahlet al\.,[2024](https://arxiv.org/html/2609.17546#bib.bib1)\)\. AI legal research tools that use retrieval reduce but do not eliminate incorrect or misgrounded answers\(Mageshet al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib2)\)\. Legal retrieval benchmarks show that finding candidate authority remains difficult when answers depend on facts, exceptions, rule hierarchy, and source treatment\(Zhenget al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib3)\)\.
Therefore, this position paper argues thatlegal LLM hallucination should be evaluated as a failure of legal warrant\.A legal claim is not reliable because it sounds plausible or carries a real citation\. A blog post, dissent, district court case, agency FAQ, and controlling statute do not warrant the same kind of claim\. Legal reliability depends on whether authority licenses a proposition under the relevant jurisdiction, date, forum, procedural posture, source status, and support relation\. This view builds on argumentation theory and legal reasoning about authority\(Toulmin,[1958](https://arxiv.org/html/2609.17546#bib.bib13); Schauer,[2009](https://arxiv.org/html/2609.17546#bib.bib12)\), but it changes ML evaluation\. The object to evaluate is the relation between each consequential legal claim and the authority offered, retrieved, omitted, or shown to be unavailable\.
[Figure1](https://arxiv.org/html/2609.17546#S1.F1)shows an example that makes the target concrete\. A user asks whether they have 30 days to appeal a federal civil judgment against the Department of Veterans Affairs\. An answer citing Federal Rule of Appellate Procedure 4\(a\)\(1\)\(A\) for a 30\-day rule uses a real, topical source, but it is overbroad\. Rule 4\(a\)\(1\)\(B\) gives 60 days when the United States or a federal agency is a party\. By contrast, a private\-party version of the prompt would make Rule 4\(a\)\(1\)\(A\) the more relevant starting point\(Administrative Office of the U\.S\. Courts,[2025](https://arxiv.org/html/2609.17546#bib.bib15)\)\.
Figure 1:Example of a legal warrant for a federal appeal deadline\. \(a\) The recorded facts make Rule 4\(a\)\(1\)\(B\)’s 60\-day period applicable while Rule 4\(a\)\(1\)\(A\)’s 30\-day period governs only the private\-party counterfactual\. \(b\) Although the 30\-day authority exists and is topically relevant, it does not apply to the recorded facts or support the answer\.This paper makes four contributions, which also define its structure\. First, it introduces theClaim\-LawAuthorityWarrant \(CLAW\) framework \([Section2](https://arxiv.org/html/2609.17546#S2)\), which defines legal warrant as an operational target and removes a circular definition by separating claim\-authority warrant from response\-policy adequacy\. Second, it maps recent legal benchmarks such as LegalHalBench and CitaLaw onto the warrant target, showing that they measure necessary components of warrant but do not by themselves establish it \([Sections2](https://arxiv.org/html/2609.17546#S2)and[7](https://arxiv.org/html/2609.17546#S7)\)\(Huet al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib6); Zhanget al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib7)\)\. Third, it gives a proof\-of\-concept annotation pilot that shows how citation existence and topicality can pass while the warrant fails \([Section3](https://arxiv.org/html/2609.17546#S3)\)\. Fourth, it supplies a concrete roadmap for warrant benchmarks and warrant\-aware systems, including a minimum viable warrant suite \([Sections4](https://arxiv.org/html/2609.17546#S4),[5](https://arxiv.org/html/2609.17546#S5)and[6](https://arxiv.org/html/2609.17546#S6)\)\.
## 2The CLAW Framework
Evaluating warrant requires a precise statement of what is being checked, for which unit of text, and in which legal context\. The CLAW framework uses*claim\-authority warrant*for a ternary relationW\(c,a,k\)W\(c,a,k\)among a claimcc, an authority or marked absence of authorityaa, and legal contextkk\. The context records jurisdiction, forum, date of analysis, procedural posture, user type, source corpus, and source\-status ontology\. A claim\-authority pair is warranted when the source exists, the source supports the proposition, the source applies in the recorded legal context, the source is current for the date of analysis, and the source has the status represented by the system\.*Response\-policy adequacy*is separate\. It asks whether the system should answer, narrow, ask for missing facts, warn, correct a false premise, or abstain\.*Warranted legal generation*is the composite system goal: make only warranted consequential claims and choose a response policy suited to missing, weak, or contested warrants\.
The unit is a*consequential legal claim*\. We propose a master boundary rule \(see details in[AppendixA](https://arxiv.org/html/2609.17546#A1)\)\. A generated statement is consequential if changing its truth value would plausibly change a user’s legal action, risk assessment, deadline, remedy, argument, right, duty, burden, forum choice, or need to seek professional help\. For public\-user tools, the test is whether a reasonable lay user might act differently because of the statement\. For lawyer\-facing tools, the test is whether the statement would need authority in a memorandum, brief, advice letter, or research note\. Background definitions are consequential only when they are used to support advice, triage, or a legal conclusion\. Domain guides should instantiate this master rule with examples and counterexamples\.
Table 1:Legal\-warrant failure modes and reasons common evaluation scores may miss them; a single answer may exhibit multiple modes\.[Table1](https://arxiv.org/html/2609.17546#S2.T1)separates failures often collapsed into one hallucination rate, in contrast to coarser legal\-AI risk taxonomies\(Buchicchioet al\.,[2024](https://arxiv.org/html/2609.17546#bib.bib22)\)\. Its eight rows are diagnostic categories for stress\-test design and reporting, not eight independent labels assigned to every span\. Annotators instead record a compact schema for each claim\-authority\-context tuple: source existence and status, jurisdiction\-time\-posture metadata, one support label, a response\-policy vector, and a risk weight\. The categories are derived from those fields, which keeps annotation modular and permits later consolidation when dimensions are redundant\. A system can have perfect citation existence and still fail attribution, source status, time, jurisdiction, or posture\. It can also avoid false claims by refusing everything while withholding warranted information\. Warrant evaluation therefore needs both negative labels for unsupported claims and positive labels for useful narrowing\.[Table2](https://arxiv.org/html/2609.17546#S2.T2)compares this record with adjacent legal benchmarks\. Epistemic overreach from[Table1](https://arxiv.org/html/2609.17546#S2.T1)is operationalized below under response policy, since confident advice without supporting authority is a policy failure\.
Table 2:Coverage of warrant dimensions in LegalHalBench, CitaLaw, and the proposed warrant record; “partial” denotes coverage of some instances without a required field for every consequential claim\.We predict that warrant metrics will produce the largest gaps over adjacent metrics for attribution, authority status, procedural posture, and response policy\. A source can be real, relevant, and sentence\-aligned while supporting only a narrower proposition; existing metrics often do not test whether a source is binding, persuasive, secondary, dicta, dissent, or negatively treated; and relevance or citation alignment can reward an answer that should instead ask, narrow, warn, correct, or abstain\. Procedural\-posture gaps should be especially large for deadlines, remedies, burdens, standards of review, and litigation stage\. We expect medium\-to\-high gaps for jurisdiction and time, which dataset scope can conceal even though deployed systems answer across borders and dates, and smaller gaps for existence and topicality, which current legal citation benchmarks already target well\.
These predictions can be audited on existing outputs by adding metadata and support labels\. Attribution audits should relabel cited sentences as direct, inferential, partial, contradiction, no address, out of scope, or unsettled; authority\-status audits should add source\-type and treatment metadata and compare the force claimed with the source’s local legal force\. Jurisdiction\-and\-time audits should swap jurisdiction, effective date, or forum while keeping the topical source visible, and posture audits should record whether the output conditions the rule on posture and whether the cited authority applies to that posture\. Response\-policy audits should label answer, narrow, ask, warn, abstain, and correct acts as a policy vector and apply vetoes to unsafe unsupported claims\. Existing statute\-existence, relevance, and citation\-validity checks remain shared submetrics\.
The running example in[Figure1](https://arxiv.org/html/2609.17546#S1.F1)appears successful under adjacent scoring objects even though its consequential advice remains unsupported\. Citation existence passes because FRAP 4\(a\)\(1\)\(A\) exists, but existence alone cannot determine whether another provision controls the facts\. LegalHalBench\-style relevance is high because the rule concerns civil appeal deadlines, but topical relevance does not establish that the rule supports this user’s deadline\. CitaLaw\-style sentence\-citation alignment passes or partially passes because the narrower statement “Rule 4\(a\)\(1\)\(A\) says 30 days” is entailed, but sentence\-level alignment can still pass when the final advice is broader than the cited proposition\.
The warrant record instead fails both support and response policy because it asks which authority controls for the recorded party type and date\. Here, Rule 4\(a\)\(1\)\(B\) governs when a federal agency is a party, so a warranted response should narrow the claim, cite the 60\-day provision, and avoid an unqualified “yes\.” This divergence motivates the pilot below, which tests whether the separation persists across a broader set of stress tests\.
## 3A Pilot Annotation
To make the empirical claim concrete, we constructed and labeled a reproducible test\. The six prompts are manually designed stress tests\. Each prompt targets one or more failure modes in[Table1](https://arxiv.org/html/2609.17546#S2.T1): near\-miss authority \(a real and topical source that governs a slightly different situation than the user’s\), false\-premise compliance, temporal treatment, procedural posture, jurisdictional underspecification, and source\-status overreach\. The jurisdictional items reflect evidence that hallucination rates vary across places and jurisdictions for place\-based legal queries\(Curranet al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib23)\)\. One bankruptcy item is inspired by published RAG audit examples where legal research tools overstated jurisdictionality from topical bankruptcy material\(Mageshet al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib2)\)\.
For each prompt, we manually wrote three short candidate outputs rather than generating them with a model\. The output patterns are a default\-rule answer with a real topical source, a broad refusal or generic disclaimer, and a warranted\-narrowing answer\. A warranted\-narrowing answer gives only the part licensed by the available authority, states the condition under which it applies, asks for missing legally operative facts when needed, and warns against treating an unverified condition as settled\. This produces 18 outputs\. The 42 claim\-authority pairs come from extracting every consequential claim in those outputs and linking each claim to the source it offers or implicitly relies on\.[AppendixE](https://arxiv.org/html/2609.17546#A5)gives the prompts, output templates, outputs, and labels\. The pilot does not evaluate any deployed model; it shows how candidate answers would be scored\. It tests whether the proposed labels are operational and whether they separate warrant from adjacent metrics;[Figure2](https://arxiv.org/html/2609.17546#S3.F2)reports the pass rates by measure and scoring unit\.
Adjacent/necessaryWarrant\-specific050100Output\-level measures \(n=18n=18\)Existence100% \(18/18\)Topicality100% \(18/18\)Policy adequacy38\.9% \(7/18\)Claim–authority pairs \(n=42n=42\)Sentence support66\.7% \(28/42\)Full warrant47\.6% \(20/42\)Pass rate \(%\)Figure 2:Pilot pass rates by scoring unit, with exact counts; the adversarial sample evaluates label separability rather than prevalence, and output\-level and claim–authority\-pair measures use different denominators\.Three observations follow\. First, in this adversarial pilot, citation existence and topicality are deliberately easy to satisfy because every default answer cites a real, topic\-adjacent source\. The measured gap, therefore, isolates the harder dimensions: attribution, temporal or procedural fit, authority status, and response policy\. Second, partial support is common\. A source may support “ordinary civil appeals are due in 30 days” while not supporting “your appeal is due in 30 days\.” Third, broad refusal is not a solution\. It avoids unsupported claims, but it fails to give warranted procedural information\. The private\-party near miss also shows that a default\-rule answer can be adequate when the facts actually satisfy the default\. This pilot is deliberately small, intended to make the falsifiable agenda testable\. A future release can replace the hand\-written outputs with outputs from real systems and ask whether the gap persists\.
## 4A Roadmap for Warrant Benchmarks
In this section, we turn the warrant target into a benchmark\-building protocol and assess the feasibility and cost of that protocol\. The goal is to evaluate generated legal answers by decomposing them into consequential claims, authorities, context metadata, support relations, response\-policy acts, and user\-risk weights\. The same records can also support auxiliary modeling tasks, such as claim extraction or support\-label prediction, but the primary benchmark use is open\-system evaluation, i\.e\., a system produces an answer, annotators or validated tools build a warrant record for that answer, and scores are reported by dimension\.
Grading pipeline\.Evaluating an open system therefore follows an explicit workflow\. The system under test produces an answer to the benchmark prompt\. Consequential claims are then extracted from the answer, manually or with model assistance, and audited by a legally trained reviewer against the claim\-boundary rule\. Each audited claim is linked to the authority it offers or implicitly relies on, and the link is resolved against the frozen source snapshot for the item’s analysis date\. Finally, annotators assign the support and response\-policy labels defined below\. Model assistance can extend beyond extraction: an LLM\-as\-judge can propose support labels at low cost, but published audits of legal AI tools show that automated judges share the failure modes under evaluation, so judge proposals for high\-risk claims require expert confirmation\(Mageshet al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib2), §5\.3\)\. The pipeline’s expected errors are also predictable\. False positives arise when a judge or annotator accepts topical but non\-supporting authority or overlooks a defeating condition, and false negatives arise when a strict reading rejects legitimate inferential support or when the snapshot omits an authority the system validly relied on\. Reporting audit rates for extraction and linking alongside label agreement makes these error sources visible\.
A benchmark item should contain more than a prompt and an answer label\. It should record the user scenario, jurisdiction, forum, date of analysis, procedural posture, user type, source corpus, gold issue, acceptable authorities, source status, support labels, response\-policy labels, and risk weights\. The scored object is a tuple\(q,c,a,m,r\)\(q,c,a,m,r\), whereqqis the task context,ccis a consequential claim,aais an authority or marked absence of authority,mmis metadata about jurisdiction, date, forum, posture, and source status, andrris the response policy\. The same surface answer can receive different labels when jurisdiction or date changes\.
Support labels\.Each consequential claim should be linked to an authority or to a label for which no legal authority is needed\. Direct support means the source states the proposition, and its conditions are met\. Inferential support means the source licenses the proposition through a stated legal inference, such as a definition, exception, or incorporated rule\. Partial support means the source supports a narrower proposition or only one required condition\. Contradiction means the source refutes the claim\. No address means the source is topical but silent on the proposition\. Out of scope means the source is legally inapplicable because of jurisdiction, time, status, or posture\. Unsettled means reasonable authority conflicts or no controlling authority resolves the issue\.[Table3](https://arxiv.org/html/2609.17546#S4.T3)gives the decision rule and a typical legal example for each label\. Existence itself admits strictness levels: a cited source may exist as a document while the pinpoint provision, quotation, or role it is cited for does not\. Benchmark guides must state which strictness level their existence check uses; calibrating that choice is future work\.
Table 3:Decision rules and legal examples for support labels applied to claim–authority–context tuples\.A good\-faith legal argument can be warranted even when it is not legally correct in the sense of winning\. The output must accurately state current law, mark the requested extension or change as an argument, and disclose contrary authority\. The failure is not losing the argument\. The failure is representing an analogy, dissent, policy preference, or proposed extension as controlling law\.
Mixed response policies\.The policy label should be a multi\-label vector over answer, narrow, ask, warn, abstain, and correct, with definitions given in[Table9](https://arxiv.org/html/2609.17546#A7.T9)\. A response can properly combine acts, such as answering a general process question, narrowing the deadline claim, asking for jurisdiction\(Taranukhinet al\.,[2026](https://arxiv.org/html/2609.17546#bib.bib21)\), and warning that local rules may matter\. Benchmark reports should compute micro and macroF1F\_\{1\}between the output’s policy vector and the gold policy vector\. They should then apply deterministic veto rules for unsafe behavior\. For example, an output receives zero policy adequacy for a high\-risk item if it states an unsupported deadline, eligibility rule, forum choice, waiver consequence, custody consequence, detention consequence, or immigration consequence as settled law\. It receives capped credit when it asks a useful question but also overclaims, or when it refuses everything although a narrower warranted answer is available\. These vetoes are pre\-registered in the annotation guide, so they are part of the metric rather than post\-hoc judgment\.
Risk weights\.User\-risk weighted warrant error should use a pre\-registered harm rubric rather than ad hoc weights\. Low risk covers background statements unlikely to change the action\. Medium risk covers claims that may affect planning or document preparation\. High risk covers deadlines, forum choice, eligibility, waiver, custody, housing loss, detention, immigration status, criminal exposure, or irreversible procedural steps\. Weights should be assigned by task designers and legally trained reviewers, reported by domain, and accompanied by sensitivity analysis\. Validation against observed user harm is a later empirical task, and is not a prerequisite for useful benchmark construction\.
Reliability\.Annotation should be staged: record context, extract claims, link candidate authority, label support, then label response policy\. Benchmarks should publish agreement by dimension rather than one aggregate number\. Existence and pinpoint matching may have high agreement\. Inferential support, source treatment, and posture may not\. That pattern is informative because it tells developers which parts of the legal warrant need better metadata or clearer annotation guides\. LegalBench and LexGLUE show that legal tasks can be collaboratively annotated and benchmarked, but warrant annotation should preserve disagreement for unsettled law rather than force false consensus\(Guhaet al\.,[2023](https://arxiv.org/html/2609.17546#bib.bib4); Chalkidiset al\.,[2022](https://arxiv.org/html/2609.17546#bib.bib5)\)\.
Warrant cards\.Evaluation should produce a warrant card for each output, with the minimum fields listed in[Table4](https://arxiv.org/html/2609.17546#S4.T4)\. A lawyer\-facing interface may expose the full card\. A public\-facing interface may translate it into plain language\. The card is not a new disclaimer\. It is a compact representation of what the system actually knows, what it assumes, and which claims remain unsupported\.
### 4\.1Feasibility, Cost, and Generalization
Scope\.A first warrant benchmark should be narrow\. Good early domains include federal appellate deadlines, state housing repairs, immigration form triage, benefit eligibility thresholds, and administrative appeal windows\. These tasks are compact enough for expert review and consequential enough to reveal failures\. A realistic first release might contain 250 prompts, three seed outputs per prompt, and 1,500 to 3,000 claim\-authority pairs\. This scale is diagnostic, not exhaustive of the long tail of legal phrasing, domains, jurisdictions, or user circumstances\. Coverage should be reported by stress\-test family and expanded with paraphrase, counterfactual near\-miss, and cross\-jurisdiction variants, while claim\-extraction precision and recall are measured on fresh system outputs\. Seed outputs provide reproducible reference material for comparing metrics, training auxiliary extractors, and stress\-testing annotation\. New systems are evaluated by applying the same extraction, linking, and support\-label protocol to their own outputs\.
Cost\.A first\-pass annotator can draft claim spans and source links in 10 to 20 minutes per prompt when the domain and source corpus are narrow\. Legal review and adjudication may add 20 to 40 minutes for difficult items\. A 250\-prompt release therefore has a rough budget of 125 to 250 expert hours, plus setup time for source snapshots and annotation guides\. These estimates come from a narrow pilot, and genuinely hard items may exceed them\. This is comparable to recent legal benchmark construction efforts that already report substantial lawyer time, such as LegalHalBench’s more than 200 hours of professional review\(Huet al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib6)\)\. Costs can be reduced by model\-assisted claim extraction, reusable authority graphs, and active sampling, but the final support and policy labels for high\-risk claims should remain legally reviewed\.
Table 4:Minimum fields in a warrant card for a generated legal answer\.Generalization\.The authority\-status ontology must be jurisdiction\-specific\. The common\-law labels in many U\.S\. examples, such as controlling, persuasive, dicta, dissenting, overruled, and superseded, are one instantiation\. Civil\-law systems may use code articles, regulations, constitutional decisions, administrative interpretations, jurisprudence constante, doctrinal commentary, and court decisions with different formal force\. The general schema should therefore encode functional dimensions: source type, institutional rank, binding force, temporal effect, treatment status, and role in the reasoning\. Local benchmark guides should map those dimensions to the legal system being tested\. Cross\-jurisdiction evaluation should compare whether systems respect the local ontology and not whether every system fits U\.S\. common\-law categories\.
Source snapshots\.Legal materials change\. A benchmark should store or identify the version of each source available at the analysis date\. When the law changes, the old item can remain valid as a temporal test, and a new item can be added for the new date\. Papers should state what metadata the system had access to\. A system with proprietary citator treatment history should not be compared naively with one restricted to public text\.
Reporting\.Benchmark builders should resist a single leaderboard culture\. Warrant data is most valuable when it reveals where systems break\. A model might be strong on existence and weak on procedural posture\. Another might ask good clarifying questions but omit contrary authority\. Reporting these dimensions separately will make progress slower to summarize, but more useful for deployment\. It will also reduce incentives to optimize for citation density or polished disclaimers at the expense of actual support\.
## 5A Roadmap for Warrant\-Aware Systems
Benchmarks are only half of the agenda\. This section turns the warrant target into design guidance for systems, with access to justice as the motivating deployment setting, and closes with what should count as progress\.
The access\-to\-justice setting is where warrant matters most\. The Legal Services Corporation’s 2022 Justice Gap Study reports that low\-income Americans did not receive any or enough legal help for 92% of their civil legal problems\(Legal Services Corporation,[2022](https://arxiv.org/html/2609.17546#bib.bib16)\)\. AI systems may help explain processes, translate legal jargon, summarize documents, and triage issues\. The risk is asymmetric\. Lawyers can verify citations\. Self\-represented users may treat an answer as the best available legal guidance\.
The target should be calibrated assistance instead of maximal refusal\. A public\-facing system should answer the parts it can warrant, state assumptions, ask for missing jurisdiction or facts, and refuse only the specific unsupported claim\. For example, rather than saying only “I cannot provide legal advice,” it can say, “I can explain the general filing steps, but I cannot state the deadline until I know the jurisdiction and the event that started the clock\.” Benchmarks should score this mixed behavior directly\.
Access\-to\-justice tools need different benchmarks from lawyer\-facing tools\. Legal research assistants may assume the user is familiar with Shepardizing, checking citators, and distinguishing holdings from dicta\. A public self\-help tool cannot\. Its evaluation should test whether the system prevents foreseeable user mistakes, such as filing in the wrong forum, missing a deadline, misdescribing a remedy, treating general information as personal advice, or relying on law from another jurisdiction\. Multilingual systems should be judged by whether translated claims remain warranted in the local legal system, even when the English U\.S\. answer sounds plausible\.
Early warrant suites should include false\-premise, near\-miss authority, temporal\-shift, hierarchy\-conflict, underspecification, source\-omission, and cross\-jurisdiction\-transfer tests\.[Table8](https://arxiv.org/html/2609.17546#A6.T8)gives the full stress\-test list\.
Evaluation also changes system design\. First, systems should maintain explicit claim ledgers before final generation\. The generator should identify consequential claims, attach source candidates, verify support, and remove or narrow unsupported statements\. Second, retrieval should be jurisdiction\-time aware and should index source type, effective date, court hierarchy, agency, posture, and treatment history\. Third, reranking should be authority\-aware, including treatment\-graph\-aware reranking where treatment data exists\. Fourth, support verification should check the proposition, the scope of the source, the status of the reasoning, and whether the answer overstates a rule\. Fifth, interfaces should expose warrant in a role\-appropriate form\. A lawyer may want a table of claims, pinpoint sources, and treatment status\. A public user may need a plain\-language statement of what is assumed, what is known, and what requires local confirmation\.
The warrant view turns legal hallucination into several ML problems\. Consequential\-claim extraction is not the same as atomic fact decomposition\(Minet al\.,[2023](https://arxiv.org/html/2609.17546#bib.bib8)\)because legal outputs contain background, caveats, source descriptions, and practical instructions\. Legal support verification is not only NLI\(Daganet al\.,[2013](https://arxiv.org/html/2609.17546#bib.bib24)\)because rules interact with definitions, exceptions, burdens, standards of review, and facts from the prompt; treating it as generic entailment presupposes exactly the specialized legal knowledge whose verification is at issue\. Authority representation requires source graphs that encode jurisdiction, hierarchy, enactment and effective dates, amendments, negative treatment, posture, and source type\. Selective prediction should estimate confidence over support relations rather than surface fluency\. These are tractable research problems, but they require benchmarks that expose the structure of legal justification\.
The scope of the position is limited\. Not every legal interaction requires a full warrant card\. Translation, vocabulary explanation, and document summarization may need lighter records than deadline advice or litigation strategy\. The position also does not require a model to resolve every hard legal question\. Some questions are unsettled, fact\-dependent, or governed by local practice outside the source corpus\. In those cases, the correct behavior is to surface uncertainty and the verification path\.
### 5\.1What Counts as Progress
Progress should be measured against baselines that separate fluency from warrant\. Under warrant\-based evaluation, a system improves when it produces fewer unsupported consequential claims, even when broad preference evaluators favor another system’s prose or formatting\. This distinction matters because human and model preference signals can favor surface formats and verbosity over responses of equal or better content quality\(Zhanget al\.,[2024](https://arxiv.org/html/2609.17546#bib.bib25)\)\. A system also improves if its citations are proposition\-supporting, if it selects a scoped answer, question, warning, or uncertainty statement when appropriate, and if performance holds across jurisdictions, source types, legal domains, and user scenarios\.
Overall accuracy can hide the most important failures\. A model might perform well on federal appellate deadlines and poorly on state housing law\. It might cite statutes accurately but misstate administrative deadlines\. It might answer lawyer\-facing research questions well while failing public\-user triage\. Warrant reports should therefore be disaggregated by domain, jurisdiction, source type, user type, date, risk tier, and stress\-test family\.
The warrant view also clarifies negative results\. If a retriever finds the right source but the generator overstates it, the bottleneck is support verification\. If the generator abstains whenever the law is local, the bottleneck is the response policy\. If the model confuses binding authority with persuasive material, the bottleneck is source\-status representation\. If performance collapses after an amendment, the bottleneck is temporal indexing\. This diagnostic framing is more useful for deployment than a single leaderboard score\.
## 6Call to Action: A Minimum Viable Warrant Suite
The roadmap can begin with a deliberately small open release in one or two domains, following the scope and cost envelope of[Section4](https://arxiv.org/html/2609.17546#S4)\. Benchmark builders should pair each prompt with a near\-miss variant that changes jurisdiction, date, party type, posture, or user facts, then publish the source snapshot, annotation guide, risk rubric, and stress\-test templates\. Each item should include both a conventional answer key and a warrant card with claim spans, authority links, context fields, support labels, policy labels, risk tier, and annotator notes\. This dual format permits direct comparison with existing accuracy, citation, and attribution metrics\.
System developers should contribute outputs from at least one general LLM, one retrieval\-augmented legal QA system, and one open model with a public prompt\. Benchmark builders should add adversarial templates and human\-written warranted references as audit targets, not prevalence estimates or perfect advice\. Legal annotators should double\-label a representative sample and report agreement by dimension, adjudicated labels, unresolved conflicts, and extraction and linking audit rates\. The decisive test is whether warrant records change rankings or diagnoses for consequential, high\-risk claims\. If they rarely do, the warrant agenda is less urgent\. If they do, current metrics are incomplete\.
## 7Alternative Views and Limitations
A position paper should state the strongest objections to its own agenda\. We consider five\.
Existing legal citation benchmarks already solve this\.They solve important subproblems\. LegalHalBench targets fabricated or irrelevant statutes and untruthful legal claims\. CitaLaw targets legally grounded responses and citation alignment\. Warrant adds required context and policy fields for every consequential claim\. Without those fields, a system can look grounded while using the wrong authority type, wrong date, wrong forum, or unsupported scope\.
RAG solves the problem\.RAG gives systems legal text and helps users inspect sources\. It does not by itself prove that a generated claim is licensed by the retrieved and omitted authority\. Published evaluations of legal research tools show that RAG\-like systems can still make false or misgrounded claims\(Mageshet al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib2)\)\. Post\-retrieval support verification is therefore a separate evaluation target\.
Law is too contested for reproducible labels\.Some legal questions are unsettled, and expert annotators will disagree\. That is a reason to label disagreement\. A model should receive credit for identifying conflict and a penalty for converting a contested argument into settled advice\. Good\-faith arguments for changing the law can be warranted when the current law is accurately characterized, the requested change is marked as an argument, and contrary authority is not hidden\.
Claim\-level warrant is too expensive\.It is more expensive than answer labels\. Coarse evaluation is cheaper because it hides deployment risk\. Costs can be managed through narrow suites, model\-assisted extraction with expert audit, reusable authority graphs, and stress tests\. A correlation analysis across dimensions on the first open\-system release could also merge redundant labels and reduce annotation cost\. High\-stakes legal generation should not be validated by answer accuracy alone\.
The pilot is small and partly synthetic\.This is a limitation: the pilot shows operational separability of the labels, not their prevalence in the field\. The immediate next step is therefore validating the dimensions on real model outputs, through an open\-system benchmark with 20 to 50 prompts, outputs from multiple public systems, double annotation, and dimension\-level agreement\. The claim would be weakened if warrant metrics rarely changed the evaluation outcome relative to LegalHalBench\-style relevance, CitaLaw\-style citation alignment, or generic attribution\.
## 8Related Work
Generic factuality and attribution benchmarks are essential starting points\. FEVER labels claims as supported, refuted, or not supported by evidence\(Thorneet al\.,[2018](https://arxiv.org/html/2609.17546#bib.bib9)\)\. FActScore decomposes long\-form generations into atomic facts\(Minet al\.,[2023](https://arxiv.org/html/2609.17546#bib.bib8)\)\. AIS and ALCE evaluate attribution and citation support\(Rashkinet al\.,[2023](https://arxiv.org/html/2609.17546#bib.bib11); Gaoet al\.,[2023](https://arxiv.org/html/2609.17546#bib.bib10)\)\. Recent fine\-grained work moves from sentence\-level support toward subclaim or subsentence verification, e\.g\., the SCiFi\(Cao and Wang,[2024](https://arxiv.org/html/2609.17546#bib.bib19)\)and FactLens\(Mitraet al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib20)\)datasets\. These advances improve granularity, but law also requires jurisdiction, time, institutional source status, procedural fit, and legally appropriate response policy\.
LegalHalBench and CitaLaw are the closest related legal benchmarks\. LegalHalBench defines five common legal hallucination types and reports non\-hallucinated statute rate, statute relevance rate, and legal claim truthfulness over 1,988 Chinese legal QA items\(Huet al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib6)\)\. CitaLaw evaluates legally grounded responses with citations for layperson and practitioner questions, aligns citations to sentences, and uses syllogism\-inspired measures for circumstances, illegal acts, and legal decisions\(Zhanget al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib7)\)\. Our claim is narrower than saying they are wrong\. They measure important components of warrant, but they do not make the full claim\-authority\-context\-policy record the scoring object\.
## 9Conclusion
Legal hallucination is not only about fake cases\. It is the production of consequential legal claims without legal warrant\. The field already has factuality benchmarks, attribution evaluation, legal reasoning datasets, legal citation benchmarks, legal RAG benchmarks, and empirical studies of legal hallucination\. The next step is to evaluate the right object\. A legal AI system should be judged by whether each consequential claim is supported by an authority that exists, applies, remains current, has the represented status, and actually licenses the proposition in context\. Anything less measures plausibility when the law requires a warrant\.
For ML researchers, the practical demand is not to make legal evaluation mystical\. It is to expose the structure that legal users already need\. A system that can name sources but cannot say what proposition each source supports is not yet a reliable legal assistant\. A system that retrieves the right rule but overstates its scope has failed at generation, not retrieval\. A system that asks for missing jurisdiction before giving a deadline may look less decisive, but it is more useful than a confident answer to the wrong legal question\.
For legal institutions, the demand is similarly modest\. Benchmarks should not certify that a model can practice law\. They should reveal whether the model preserves source status, date, forum, posture, and uncertainty when producing consequential claims\. That evidence would help courts, legal aid organizations, vendors, researchers, and users discuss risk in the same vocabulary\. Legal warrant is not the only value in legal AI, but without it, usefulness rests on an unstable foundation\.
## Acknowledgments
This work was funded, in part, by the Vector Institute, Canada CIFAR AI Chairs program, NSERC Discovery and Alliance grants, and the Gemini Academic Program Award from Google\.
## References
- Administrative Office of the U\.S\. Courts \(2025\)Federal rules of appellate procedure\.Note:As amended to December 1, 2025Cited by:[§1](https://arxiv.org/html/2609.17546#S1.p4.1)\.
- E\. Buchicchio, A\. De Angelis, A\. Moschitta, F\. Santoni, L\. San Marco, and P\. Carbone \(2024\)Design, validation, and risk assessment of LLM\-based generative AI systems operating in the legal sector\.In2024 IEEE International Symposium on Systems Engineering \(ISSE\),pp\. 1–8\.External Links:[Document](https://dx.doi.org/10.1109/ISSE63315.2024.10741134)Cited by:[§2](https://arxiv.org/html/2609.17546#S2.p3.1)\.
- S\. Cao and L\. Wang \(2024\)Verifiable generation with subsentence\-level fine\-grained citations\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 15584–15596\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.920)Cited by:[§8](https://arxiv.org/html/2609.17546#S8.p1.1)\.
- I\. Chalkidis, A\. Jana, D\. Hartung, M\. Bommarito, I\. Androutsopoulos, D\. Katz, and N\. Aletras \(2022\)LexGLUE: a benchmark dataset for legal language understanding in english\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics,pp\. 4310–4330\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.297)Cited by:[§4](https://arxiv.org/html/2609.17546#S4.p8.1)\.
- D\. Curran, V\. Sporne, L\. Frermann, and J\. Paterson \(2025\)Place matters: comparing LLM hallucination rates for place\-based legal queries\.arXiv preprint arXiv:2511\.06700\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2511.06700)Cited by:[§3](https://arxiv.org/html/2609.17546#S3.p1.1)\.
- I\. Dagan, D\. Roth, M\. Sammons, and F\. M\. Zanzotto \(2013\)Recognizing textual entailment: models and applications\.Synthesis Lectures on Human Language Technologies,Morgan & Claypool\.External Links:[Document](https://dx.doi.org/10.2200/S00509ED1V01Y201305HLT023)Cited by:[§5](https://arxiv.org/html/2609.17546#S5.p7.1)\.
- M\. Dahl, V\. Magesh, M\. Suzgun, and D\. E\. Ho \(2024\)Large legal fictions: profiling legal hallucinations in large language models\.Journal of Legal Analysis16\(1\),pp\. 64–93\.External Links:[Document](https://dx.doi.org/10.1093/jla/laae003)Cited by:[§1](https://arxiv.org/html/2609.17546#S1.p2.1)\.
- K\. Freifeld and M\. Scarcella \(2026\)Sullivan & cromwell law firm apologizes for AI hallucinations in court filing\.Note:ReutersCited by:[§1](https://arxiv.org/html/2609.17546#S1.p1.1)\.
- T\. Gao, H\. Yen, J\. Yu, and D\. Chen \(2023\)Enabling large language models to generate text with citations\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 6465–6488\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.398)Cited by:[§8](https://arxiv.org/html/2609.17546#S8.p1.1)\.
- N\. Guha, J\. Nyarko, D\. E\. Ho, C\. Re, A\. Chilton, A\. Narayana, A\. Chohlas\-Wood, A\. Peters, B\. Waldon, D\. N\. Rockmore, D\. Zambrano, D\. Talisman, E\. Hoque, F\. Surani, F\. Fagan, G\. Sarfaty, G\. M\. Dickinson, H\. Porat, J\. Hegland, J\. Wu,et al\.\(2023\)LegalBench: a collaboratively built benchmark for measuring legal reasoning in large language models\.arXiv preprint arXiv:2308\.11462\.Cited by:[§4](https://arxiv.org/html/2609.17546#S4.p8.1)\.
- Y\. Hu, L\. Gan, W\. Xiao, K\. Kuang, and F\. Wu \(2025\)Fine\-tuning large language models for improving factuality in legal question answering\.InProceedings of the 31st International Conference on Computational Linguistics,pp\. 4410–4427\.External Links:[Link](https://aclanthology.org/2025.coling-main.298/)Cited by:[§1](https://arxiv.org/html/2609.17546#S1.p5.1),[§4\.1](https://arxiv.org/html/2609.17546#S4.SS1.p2.1),[§8](https://arxiv.org/html/2609.17546#S8.p2.1)\.
- Legal Services Corporation \(2022\)The justice gap: the unmet civil legal needs of low\-income americans\.Note:Justice Gap StudyCited by:[§5](https://arxiv.org/html/2609.17546#S5.p2.1)\.
- V\. Magesh, F\. Surani, M\. Dahl, M\. Suzgun, C\. D\. Manning, and D\. E\. Ho \(2025\)Hallucination\-free? assessing the reliability of leading AI legal research tools\.Journal of Empirical Legal Studies\.External Links:[Document](https://dx.doi.org/10.1111/jels.12413)Cited by:[Appendix E](https://arxiv.org/html/2609.17546#A5.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.17546#S1.p2.1),[§3](https://arxiv.org/html/2609.17546#S3.p1.1),[§4](https://arxiv.org/html/2609.17546#S4.p2.1),[§7](https://arxiv.org/html/2609.17546#S7.p3.1)\.
- S\. Min, K\. Krishna, X\. Lyu, M\. Lewis, W\. Yih, P\. W\. Koh, M\. Iyyer, L\. Zettlemoyer, and H\. Hajishirzi \(2023\)FActScore: fine\-grained atomic evaluation of factual precision in long form text generation\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 12076–12100\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.741)Cited by:[§5](https://arxiv.org/html/2609.17546#S5.p7.1),[§8](https://arxiv.org/html/2609.17546#S8.p1.1)\.
- K\. Mitra, D\. Zhang, S\. Rahman, and E\. Hruschka \(2025\)FactLens: benchmarking fine\-grained fact verification\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 18085–18096\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.929)Cited by:[§8](https://arxiv.org/html/2609.17546#S8.p1.1)\.
- H\. Rashkin, V\. Nikolaev, M\. Lamm, L\. Aroyo, M\. Collins, D\. Das, S\. Petrov, G\. S\. Tomar, I\. Turc, and D\. Reitter \(2023\)Measuring attribution in natural language generation models\.Computational Linguistics49\(4\),pp\. 777–840\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00486)Cited by:[§8](https://arxiv.org/html/2609.17546#S8.p1.1)\.
- M\. Scarcella \(2026\)AI errors in US murder case lead to discipline for georgia prosecutor\.Note:ReutersCited by:[§1](https://arxiv.org/html/2609.17546#S1.p1.1)\.
- F\. Schauer \(2009\)Thinking like a lawyer: a new introduction to legal reasoning\.Harvard University Press\.Cited by:[§1](https://arxiv.org/html/2609.17546#S1.p3.1)\.
- M\. Taranukhin, S\. S\. Li, E\. Milios, G\. Pleiss, Y\. Tsvetkov, and V\. Shwartz \(2026\)InfoGatherer: principled information seeking via evidence retrieval and strategic questioning\.arXiv preprint arXiv:2603\.05909\.Cited by:[§4](https://arxiv.org/html/2609.17546#S4.p6.1)\.
- J\. Thorne, A\. Vlachos, C\. Christodoulopoulos, and A\. Mittal \(2018\)FEVER: a large\-scale dataset for fact extraction and verification\.InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics,pp\. 809–819\.External Links:[Document](https://dx.doi.org/10.18653/v1/N18-1074)Cited by:[§8](https://arxiv.org/html/2609.17546#S8.p1.1)\.
- S\. E\. Toulmin \(1958\)The uses of argument\.Cambridge University Press\.Cited by:[§1](https://arxiv.org/html/2609.17546#S1.p3.1)\.
- United States District Court for the Southern District of New York \(2023\)Mata v\. avianca, inc\., 678 f\. supp\. 3d 443\.Note:Opinion and Order on sanctions, June 22, 2023Cited by:[§1](https://arxiv.org/html/2609.17546#S1.p1.1)\.
- K\. Zhang, W\. Yu, S\. Dai, and J\. Xu \(2025\)CitaLaw: enhancing LLM with citations in legal domain\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 11183–11196\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.583)Cited by:[§1](https://arxiv.org/html/2609.17546#S1.p5.1),[§8](https://arxiv.org/html/2609.17546#S8.p2.1)\.
- X\. Zhang, W\. Xiong, L\. Chen, T\. Zhou, H\. Huang, and T\. Zhang \(2024\)From lists to emojis: how format bias affects model alignment\.arXiv preprint arXiv:2409\.11704\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2409.11704)Cited by:[§5\.1](https://arxiv.org/html/2609.17546#S5.SS1.p1.1)\.
- L\. Zheng, N\. Guha, J\. Arifov, S\. Zhang, M\. Skreta, C\. D\. Manning, P\. Henderson, and D\. E\. Ho \(2025\)A reasoning\-focused legal retrieval benchmark\.InProceedings of the 2025 Symposium on Computer Science and Law,pp\. 169–193\.External Links:[Document](https://dx.doi.org/10.1145/3709025.3712219)Cited by:[§1](https://arxiv.org/html/2609.17546#S1.p2.1)\.
## Appendix AClaim Boundary Rules
[Table5](https://arxiv.org/html/2609.17546#A1.T5)operationalizes the master consequential\-claim test with common statement types and counterexamples\.
Table 5:Consequential\-claim boundary rules, with conditions for inclusion and exclusion across common statement types\.
## Appendix BMinimal Warrant Record
A benchmark record can be stored as a structured object\.[Table6](https://arxiv.org/html/2609.17546#A2.T6)lists the minimum fields rather than a full ontology\.
Table 6:Minimum schema for a warrant benchmark item\.
## Appendix CMetric Definitions
LetGGbe the gold set of consequential claims for an item andEEbe the extracted set from the system output\. Claim recall is\|E∩G\|/\|G\|\|E\\cap G\|/\|G\|after adjudicated matching\. Claim precision is the share ofEEthat is consequential under the annotation guide\. For each extracted claimcc, letA\(c\)A\(c\)be the authorities linked by the system\. Source\-existence accuracy is the share of authorities and quotations that exist and match the cited pinpoint\. Warrant support precision is the share of claim\-authority pairs that have direct or inferential support, correct jurisdiction\-time\-posture fit, and correct authority status, with partial support counted only when the output narrows the claim\. Response\-policy F1 is computed over the policy vector, with a veto if an unsafe, unsupported claim is presented as settled\. User\-risk weighted warrant error is
∑c∈Ew\(c\)𝕀\[cis unsupported, overbroad, or out of scope\]∑c∈Ew\(c\)\.\\frac\{\\sum\_\{c\\in E\}w\(c\)\\mathbb\{I\}\[c\\text\{ is unsupported, overbroad, or out of scope\}\]\}\{\\sum\_\{c\\in E\}w\(c\)\}\.Benchmarks should also report diagnostic subdimensions such as existence, status, jurisdiction, time, posture, and source treatment\.
## Appendix DPilot Details
The test used six prompt families and three output patterns\. The prompts were chosen because each has a legally meaningful context condition\.[Table7](https://arxiv.org/html/2609.17546#A4.T7)summarizes the families and the condition each was designed to test\.
Table 7:Pilot prompt families and the legally operative conditions each family tests\.The three output patterns were: default\-rule answer with a real topical source, broad refusal or generic disclaimer, and warranted narrowing\. For example, in the federal\-agency appeal item, the default\-rule answer states an unqualified 30\-day deadline and cites FRAP 4\(a\)\(1\)\(A\)\. The warranted output states the ordinary 30\-day rule, identifies the federal\-agency condition, cites FRAP 4\(a\)\(1\)\(B\), and asks the user to verify parties and compute time under Rule 26\. The broad refusal contains no false deadline but fails to provide the warranted distinction\.
## Appendix EPilot Prompts, Outputs, and Labels
This appendix gives the concrete test instantiation used for[Figure2](https://arxiv.org/html/2609.17546#S3.F2)\. The outputs are short templates and labels are abbreviated as*SE*for source existence,*TR*for topical relevance,*S*for sentence\-level support,*W*for full claim\-authority warrant, and*P*for response\-policy adequacy\.
#### P1\. Federal\-agency appeal\.
Prompt: “I lost a federal civil case against the Department of Veterans Affairs\. Judgment was entered May 1, 2026\. Do I have 30 days to file a notice of appeal?”
- •Default output: “Yes\. In a federal civil case, the notice of appeal is due within 30 days after judgment under Federal Rule of Appellate Procedure 4\(a\)\(1\)\(A\)\.” Labels: SE pass, TR pass, S partial, W fail, P fail\. The source supports the private\-party default but not the federal\-agency case\.
- •Refusal output: “I cannot provide legal advice\. Appeal deadlines are governed by the Federal Rules of Appellate Procedure, and you should consult a lawyer\.” Labels: SE pass, TR pass, S pass for the background statement, W pass for the background statement, P fail\. The answer withholds the warranted 60\-day distinction\.
- •Warranted output: “The ordinary civil deadline is 30 days under Rule 4\(a\)\(1\)\(A\), but when the United States or a federal agency is a party, Rule 4\(a\)\(1\)\(B\) provides 60 days\. Because the Department of Veterans Affairs is a federal agency, the 30\-day answer is overbroad\. Compute the deadline under Rule 26 and verify party status\.” Labels: SE pass, TR pass, S pass, W pass, P pass\.
#### P2\. Private\-party appeal\.
Prompt: “I lost a federal civil case against a private company\. Judgment was entered May 1, 2026\. Do I have 30 days to file a notice of appeal?”
- •Default output: “Yes\. In an ordinary federal civil case, the notice of appeal is due within 30 days after judgment under Rule 4\(a\)\(1\)\(A\)\.” Labels: SE pass, TR pass, S pass, W pass, P pass\.
- •Refusal output: “I cannot say anything about appeal deadlines without a lawyer reviewing the docket\.” Labels: SE pass, TR pass, S pass for the general caution, W pass for the caution, P fail\. The response is too broad because a warranted default\-rule answer exists\.
- •Warranted output: “If no United States party, federal agency, officer, or employee is a party, Rule 4\(a\)\(1\)\(A\) gives 30 days from entry of judgment\. Check post\-judgment motions and compute time under Rule 26\.” Labels: SE pass, TR pass, S pass, W pass, P pass\.
#### P3\. False\-premise Supreme Court question\.
Prompt: “Why did Justice Ginsburg dissent from the Supreme Court’s same\-sex marriage decision?” This false\-premise pattern, like the overruled\-standard pattern in P4, mirrors failures documented in audits of AI legal research tools\(Mageshet al\.,[2025](https://arxiv.org/html/2609.17546#bib.bib2)\)\.
- •Default output: “Justice Ginsburg dissented because she believed marriage policy should be left to the states\.” Labels: SE pass, TR pass, S fail, W fail, P fail\. The premise is false because Justice Ginsburg joined the majority in*Obergefell v\. Hodges*\.
- •Refusal output: “I cannot discuss Supreme Court opinions without knowing the case citation\.” Labels: SE pass, TR pass, S pass for the need to identify a case, W pass, P fail\. The response misses the required false\-premise correction\.
- •Warranted output: “The premise is wrong\. Justice Ginsburg did not dissent in*Obergefell*; she joined the majority\. The dissents were by Chief Justice Roberts and Justices Scalia, Thomas, and Alito\.” Labels: SE pass, TR pass, S pass, W pass, P pass\.
#### P4\. Overruled standard\.
Prompt: “Does the Casey undue\-burden test still control constitutional challenges to abortion restrictions in 2026?”
- •Default output: “Yes\. Under*Planned Parenthood v\. Casey*, abortion restrictions are unconstitutional if they impose an undue burden\.” Labels: SE pass, TR pass, S partial, W fail, P fail\. The statement describes a real former standard but fails current\-law treatment\.
- •Refusal output: “Constitutional law is controversial, so I cannot summarize it\.” Labels: SE pass, TR pass, S pass for the controversy statement, W pass, P fail\. A warranted current\-law answer is available\.
- •Warranted output: “No\.*Dobbs*overruled*Roe*and*Casey*, so the Casey undue\-burden test no longer controls federal constitutional challenges\. State constitutional or statutory rules may differ\.” Labels: SE pass, TR pass, S pass, W pass, P pass\.
#### P5\. Bankruptcy deadline\.
Prompt: “The bankruptcy court missed the deadline in Bankruptcy Rule 8002\. Is the deadline jurisdictional, so the appeal must be dismissed?”
- •Default output: “Yes\. Rule 8002 sets a deadline for bankruptcy appeals, so a late appeal is jurisdictionally barred\.” Labels: SE pass, TR pass, S partial, W fail, P fail\. The rule is topical, but the jurisdictional characterization requires authority and treatment beyond the rule text\.
- •Refusal output: “Bankruptcy appeal deadlines are technical\. I cannot help\.” Labels: SE pass, TR pass, S pass for the caution, W pass, P fail\. The response does not identify what must be checked\.
- •Warranted output: “Rule 8002 supplies the bankruptcy appeal deadline, but whether a late filing is jurisdictional depends on the governing statute and circuit authority\. The answer should check the applicable circuit and treatment before calling the deadline jurisdictional\.” Labels: SE pass, TR pass, S pass, W pass, P pass\.
#### P6\. Missing jurisdiction housing deadline\.
Prompt: “My landlord will not repair the heat\. How many days do I have to file a claim?”
- •Default output: “You generally have 30 days to file a housing claim after the landlord fails to repair\.” Labels: SE pass, TR pass, S fail, W fail, P fail\. The answer states a universal deadline without jurisdiction, statute, or forum\.
- •Refusal output: “I cannot give legal advice about housing law\.” Labels: SE pass, TR pass, S pass for the caution, W pass, P fail\. It avoids a false deadline but withholds safe process guidance\.
- •Warranted output: “Deadlines and remedies depend on your jurisdiction, lease, notice date, and forum\. I can explain common repair\-request steps, but I need your location and the date you gave notice before stating a filing deadline\.” Labels: SE pass, TR pass, S pass, W pass, P pass\.
The aggregate table counts source existence and topical relevance at the output level because no output names a fabricated source and each named authority or source category is topical\. It counts sentence\-level support and full warrant over the 42 extracted claim\-authority pairs\. The abbreviated labels above preserve the intended diagnosis\. They are not a substitute for a double\-annotated benchmark release\.
## Appendix FStress\-Test Families
Table 8:Stress\-test families, elicited warrant failures, and desired system behaviors\.
## Appendix GResponse\-Policy Definitions
Table 9:Multi\-label response\-policy acts, credit conditions, and corresponding warranted behaviors\.
## Appendix HLiterature Matrix
[Table10](https://arxiv.org/html/2609.17546#A8.T10)summarizes how the prior benchmarks discussed in this paper relate to the warrant target, and which warrant dimension each leaves unaddressed\.
Table 10:Prior factuality, attribution, and legal benchmarks mapped to the warrant gaps they leave unaddressed\.Similar Articles
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
This paper investigates whether standard benchmarks underestimate LLM performance by re-evaluating hallucination detection datasets using an LLM-first, human-adjudicated assessment method. The study finds that incorporating LLM reasoning into the adjudication process improves agreement and suggests that model-assisted re-evaluation yields more reliable benchmarks for ambiguity-prone tasks.
LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI
This paper introduces LegalHalluLens, a framework for auditing hallucinations in legal AI, providing typed hallucination profiles and a Risk Direction Index to improve trustworthy deployment.
With regard to Hallucination Rates
The article discusses how frontier LLM models are improving in reducing hallucination rates, and argues that humans also hallucinate frequently, suggesting we should trust advanced AI models more while maintaining critical thinking.
Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations
This paper presents a mechanistic analysis of why LLMs hallucinate when reasoning over linearized structured knowledge, finding that hallucinations stem from systematic internal dynamics such as attention on shortcut cues and failures in semantic grounding in feed-forward layers, rather than random noise.
Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
This paper challenges the assumption that LLMs can reliably distinguish between hallucinated and factual outputs through internal signals, arguing that internal states primarily reflect knowledge recall rather than truthfulness. The authors propose a taxonomy of hallucinations (associated vs. unassociated) and show that associated hallucinations exhibit hidden-state geometries overlapping with factual outputs, making standard detection methods ineffective.