超越准确性与表层流畅度:针对法律条款生成的LLMs风险敏感评估
摘要
本文提出了一种针对LLM生成的合同条款的风险敏感评估框架,重点关注法律失效模式和质量维度,以评估超出准确性或流畅性的风险。
arXiv:2609.22127v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to draft contractual language, yet conventional accuracy or preference-based evaluations are poorly matched to legal drafting. A clause may be fluent and stylistically polished while still omitting an essential carve-out, allocating risk in an unenforceable way, assuming an inapplicable jurisdiction, or exposing a party to regulatory liability. This paper presents a empirical study design and framework for evaluating LLM-generated contract clauses. The study evaluates four models - Claude Haiku 4.5, Gemini 2.5 Flash Lite, GPT 5.4 Nano, and Qwen 3.5 Flash, across 22 contract clause categories and 34 legally-motivated failure modes. We combine two evaluation frameworks: CLAUSE, which classifies prompts by legal function and failure target, and LENS-CRAFT, which scores outputs across nine legal-quality dimensions. Instead of averaging dimension scores, the study applies a Max Severity Principle so that a single legally decisive defect remains visible. The paper provides the evaluation protocol, taxonomy, analysis plan, and a results structure for reporting empirical findings. We argue that legal AI evaluation should move beyond aggregate accuracy toward clause-specific, failure-mode-driven, and risk-sensitive assessment.
查看缓存全文
缓存时间: 2026/09/22 09:03
# Beyond Accuracy and Surface Fluency: Risk-Sensitive Evaluation of LLMs for Legal Clause Generation
Source: [https://arxiv.org/html/2609.22127](https://arxiv.org/html/2609.22127)
###### Abstract
Large language models \(LLMs\) are increasingly used to draft contractual language, yet conventional accuracy or preference\-based evaluations are poorly matched to legal drafting\. A clause may be fluent and stylistically polished while still omitting an essential carve\-out, allocating risk in an unenforceable way, assuming an inapplicable jurisdiction, or exposing a party to regulatory liability\. This paper presents a empirical study design and framework for evaluating LLM\-generated contract clauses\. The study evaluates four models \- Claude Haiku 4\.5, Gemini 2\.5 Flash Lite, GPT 5\.4 Nano, and Qwen 3\.5 Flash, across 22 contract clause categories and 34 legally\-motivated failure modes\. We combine two evaluation frameworks: CLAUSE, which classifies prompts by legal function and failure target, and LENS\-CRAFT, which scores outputs across nine legal\-quality dimensions\. Instead of averaging dimension scores, the study applies a Max Severity Principle so that a single legally decisive defect remains visible\. The paper provides the evaluation protocol, taxonomy, analysis plan, and a results structure for reporting empirical findings\. We argue that legal AI evaluation should move beyond aggregate accuracy toward clause\-specific, failure\-mode\-driven, and risk\-sensitive assessment\.
###### Keywords:
Legal AI, Large Language Models, Contract Drafting, Evaluation, AI Governance
††footnotetext:Accepted to the AI for Law Workshop @ ICML 2026\.## 1Introduction
LLMs have emerged to produce contract clauses that appear polished, complete, and professional\. In legal practice, however, drafting quality is not reducible to lingustic fluency\. A generated privacy clause may fail to distinguish controller and processor obligations; a limitation\-of\-liability clause may cap liability in circumstances where public policy, statute, or negotiation context requires exceptions; and a jurisdiction clause may assume an inappropriate forum\. These errors are not merely legal representation, but may affect enforceability, compliance, create litigation exposure, limit negotiation leverage, and amplify professional responsibility\.
Legal work imposes evidentiary expectations that differ from generic Natural Language Processing \(NLP\) tasks\. Prior work shows that LLMs can hallucinate legal contexts, struggle with jurisdiction\-specific questions, and underperform on basic legal text tasks\([4](https://arxiv.org/html/2609.22127#bib.bib2);[1](https://arxiv.org/html/2609.22127#bib.bib3);[10](https://arxiv.org/html/2609.22127#bib.bib7)\)\. Legal benchmarks and legal\-agent evaluations provide important coverage of legal reasoning, retrieval, and workflow tasks\([7](https://arxiv.org/html/2609.22127#bib.bib1);[21](https://arxiv.org/html/2609.22127#bib.bib14);[16](https://arxiv.org/html/2609.22127#bib.bib5);[30](https://arxiv.org/html/2609.22127#bib.bib6)\)\. However, contract drafting remains distinct as the output is not only an answer but also potential operative legal language \(which requires expertise and jurisdictional context\)\.
Our research studies the aforementioned use case through a structured evaluation mechanism for generating, stress\-testing, and scoring contractual clauses\. The protocol combines two frameworks developed for this study\. First, the CLAUSE framework defines the aspects that needs to be validated: Compliance and Jurisdiction, Legal Competence and Reasoning, Accuracy and Trustworthiness, Understanding Context, Structured Contract Handling, and Ethics, Privacy and Accountability\. Second, the LENS\-CRAFT rubric provides how to evaluates each generated clause across nine dimensions: Legal Soundness, Expressiveness, Necessity and Completeness, Scenario Alignment, Compliance and Regulatory Alignment, Risk Allocation, Architectural Coherence, Formal Drafting Quality, and Test Robustness\.
We focus on four widely\-used models that are plausible candidates for integration into legal productivity tools: Claude Haiku 4\.5, Gemini 2\.5 Flash Lite, GPT 5\.4 Nano, and Qwen 3\.5 Flash\. The experimental design covers 22 clause categories and 34 failure modes \(aligned to the CLAUSE framework\)\. The aim is not to find which model is universally best\. Instead, we set out to examine which models fail on which kinds of clauses, which failure modes are most severe, and what insights can we derive about holistic evaluation of LLMs in contract drafting\.
#### Contributions\.
We bring four contributions in this paper\. First, our study presents a clause\-level evaluation protocol for LLM contract drafting across 22 legal clause types\. Second, our study operationalizes a 34\-mode legal failure taxonomy for empirical evaluation\. Third, it introduces a risk\-sensitive aggregation method, the Max Severity Principle, designed to prevent critical legal defects from being obscured by otherwise high scores\. Fourth, it provides a reproducible analysis structure for comparing model behavior across clause categories, failure modes, and legal\-risk dimensions\.
## 2Related Work and Methodology
### 2\.1Related Work
#### Legal LLM evaluation\.
LegalBench shifted legal AI evaluation toward task\-specific coverage\([7](https://arxiv.org/html/2609.22127#bib.bib1)\), while later benchmarks and surveys examine retrieval\-augmented legal generation, legal\-domain agents, long\-form legal QA, and domain\-specific NLP limitations\([21](https://arxiv.org/html/2609.22127#bib.bib14);[16](https://arxiv.org/html/2609.22127#bib.bib5);[19](https://arxiv.org/html/2609.22127#bib.bib15);[12](https://arxiv.org/html/2609.22127#bib.bib16);[30](https://arxiv.org/html/2609.22127#bib.bib6)\)\. These studies motivate granular evaluations that distinguish legal reasoning, retrieval, factual reliability, and practical legal usability\.
#### Contract\-analysis datasets\.
Contract\-focused benchmarks such as CUAD, LEDGAR, and ContractNLI establish important baselines for clause extraction, contract provision classification, and document\-level inference\([9](https://arxiv.org/html/2609.22127#bib.bib25);[27](https://arxiv.org/html/2609.22127#bib.bib26);[13](https://arxiv.org/html/2609.22127#bib.bib27)\)\. LexGLUE also consolidates legal NLU tasks, including contract\-related components, into a standardized benchmark suite\([2](https://arxiv.org/html/2609.22127#bib.bib28)\)\. These datasets are complementary to our study; as they evaluate understanding and classification of existing contracts, while our focus is generative clause drafting, severity grading, and failure\-mode analysis\.
#### Hallucination, gaps, and legal reliability\.
In legal contexts, fabricated reflections or unsupported legal propositions can mislead lawyers, courts, and clients\. Work on legal hallucinations, legal\-analysis gaps, and human error detection shows that apparently plausible outputs can remain unreliable even when they avoid explicit fake citations\([4](https://arxiv.org/html/2609.22127#bib.bib2);[10](https://arxiv.org/html/2609.22127#bib.bib7);[8](https://arxiv.org/html/2609.22127#bib.bib17);[24](https://arxiv.org/html/2609.22127#bib.bib18)\)\. Plausibility vs correctness is important for clause drafting because a clause may contain no fabricated authority yet still omit a necessary procedure, exception, or remedy\.
#### Contract drafting and legal language\.
Contract generation differs from legal question answering because outputs can be incorporated directly into agreements\. Prior work studies LLM\-assisted drafting, synthetic contract data, smart legal contracts, expert\-annotated contract retrieval, and legal\-data\-transfer drafting\([14](https://arxiv.org/html/2609.22127#bib.bib8);[11](https://arxiv.org/html/2609.22127#bib.bib9);[3](https://arxiv.org/html/2609.22127#bib.bib19);[31](https://arxiv.org/html/2609.22127#bib.bib10);[26](https://arxiv.org/html/2609.22127#bib.bib20)\)\. Some of the work on legalese and plain language also exhibits the need for clarity\([20](https://arxiv.org/html/2609.22127#bib.bib11);[25](https://arxiv.org/html/2609.22127#bib.bib21)\), but the nuance is that a clearly represented text may still allocate risk improperly or violate regulatory requirements\. Legal clause generation evaluation is not semantic syntax problem; but a domain aligned reliability problem\.
#### Trustworthy and high\-risk AI\.
Legal drafting is high\-risk because users may rely on outputs in binding documents and in some cases may not spot a legal nuance\. Research on trustworthy AI, hidden LLM risks, bias, prompt sensitivity, and legal accountability emphasizes robustness, consistency, transparency, and human oversight\([17](https://arxiv.org/html/2609.22127#bib.bib12);[29](https://arxiv.org/html/2609.22127#bib.bib13);[6](https://arxiv.org/html/2609.22127#bib.bib22);[22](https://arxiv.org/html/2609.22127#bib.bib23);[28](https://arxiv.org/html/2609.22127#bib.bib24)\)\. Evaluation in high\-risk domains must account for asymmetric harms, where one severe error can dominate many minor successes\([15](https://arxiv.org/html/2609.22127#bib.bib4)\)\.
#### Rubric reliability and LLM\-as\-judge design\.
LLM\-as\-judge evaluation using an LLM coucil, rubric\-based scoring, and reliability calibration are foundational for our study\([33](https://arxiv.org/html/2609.22127#bib.bib29);[18](https://arxiv.org/html/2609.22127#bib.bib30);[23](https://arxiv.org/html/2609.22127#bib.bib31)\)\. Recent studies show that judge choice, prompt design, calibration, and robustness checks can materially affect evaluation reliability\([32](https://arxiv.org/html/2609.22127#bib.bib32)\)\. This motivates our use of explicit scoring anchors, multiple evaluators via LLM council, sampled human review, and Max Severity aggregation rather than relying on a single uncalibrated judge score\.
### 2\.2Methodology
We adopt a structured empirical benchmarking methodology to evaluate the reliability, robustness, and legal adequacy of LLMs for contractual clause generation\. The methodology combines systematically grounded legal evaluation, structured prompt engineering, failure\-mode\-oriented testing, multi\-model comparative benchmarking, and rubric\-driven risk assessment\.
The study proceeds through a compact eight\-stage pipeline: \(1\) Review legal AI reliability and drafting literature; \(2\) define the CLAUSE ontology for verifiables; \(3\) define the LENS\-CRAFT rubric for evaluation; \(4\) select 22 clause categories; \(5\) generate prompts for B2B, B2C, and B2G by clause, context, failure mode, and legal constraint; \(6\) run the four models; \(7\) score normalized outputs through a rubric\-driven LLM council; and \(8\) validate approximately 10% of outputs through stratified, exception\-triggered, and high\-severity manual review\.
For reproducibility and analysis, the CLAUSE dimensions and LENS\-CRAFT dimensions will be open\-sourced along with the definitions, representative exemplars, prompt guidance, and model\-configuration sheet\. The run settings are summarized in[Appendix A](https://arxiv.org/html/2609.22127#A1)\.
## 3Evaluation Framework
### 3\.1CLAUSE: Prompt Classification and Failure Targeting
The CLAUSE framework classifies each drafting prompt by the legal function being tested\. It is used before generation to ensure that prompts are not treated as generic drafting requests but as structured tests of legal capability\. The six CLAUSE dimensions are shown in[Table1](https://arxiv.org/html/2609.22127#S3.T1)\.
Table 1:CLAUSE dimensions used for prompt classification\.
### 3\.2LENS\-CRAFT: Clause Quality and Risk Scoring
After generation, each clause is evaluated using LENS\-CRAFT, a nine\-dimensional rubric designed for legal drafting\. Each dimension receives a severity score from 1 to 5, where 1 indicates production\-ready quality and 5 indicates a highly litigious, void, illegal, or otherwise severe defect\. The dimensions are summarized in[Table2](https://arxiv.org/html/2609.22127#S3.T2)\.
Table 2:LENS\-CRAFT scoring dimensions\. Each dimension is scored on a 1–5 severity scale\.
### 3\.3Risk Scale and Max Severity Principle
The protocol maps LENS\-CRAFT ratings to a five\-level risk scale: Level 1, Production\-Ready; Level 2, Generally Safe; Level 3, Risk\-Prone or Ambiguous; Level 4, Highly Problematic; and Level 5, Highly Litigious\. This ordinal mapping follows the broader safety\-evaluation principle that deployment decisions should vary with strictness and harm severity rather than average acceptability\([5](https://arxiv.org/html/2609.22127#bib.bib33)\)\. The final risk level is computed using a Max Severity Principle:
Rfinal\(x\)=maxd∈Dsd\(x\),R\_\{final\}\(x\)=\\max\_\{d\\in D\}s\_\{d\}\(x\),\(1\)wherexxis a generated clause,DDis the set of LENS\-CRAFT dimensions and critical failure indicators, andsd\(x\)s\_\{d\}\(x\)is the severity score for dimensiondd\. This rule intentionally rejects simple averaging\. In legal drafting, a clause that is clear, grammatical, and well structured may still be unacceptable if one dimension reveals a dispositive defect, such as an unlawful data\-use authorization or unenforceable liability waiver\.
### 3\.4Contract nature
The study evaluates contract clause generation across three transactional contexts: Business\-to\-Business \(B2B\), Business\-to\-Consumer \(B2C\), and Business\-to\-Government \(B2G\)\. These contexts were selected to capture differing legal, operational, regulatory, and risk\-allocation dynamics associated with enterprise negotiations, consumer protection obligations, and public\-sector procurement environments\.
## 4Dataset and Experimental Design
### 4\.1Clause Categories
The study evaluates 22 clause categories selected to cover technology contracting, data governance, liability allocation, service operations, and standard boilerplate\.[Table3](https://arxiv.org/html/2609.22127#S4.T3)groups the clause categories into legal families\. These are general contract clauses in AI contracts\.
Table 3:Clause categories evaluated in the study\.
### 4\.2Failure Modes
The evaluation examines 34 failure modes\. The failure modes are organized into eight families: legal invalidity, regulatory non\-compliance, contextual mismatch, incompleteness, risk\-allocation defects, internal inconsistency, drafting ambiguity, and trustworthiness defects\. This organization follows the observation that legal AI errors are often not isolated factual mistakes but failures to satisfy the functional requirements of an operative legal text\.
### 4\.3Models
The four evaluated systems are Claude Haiku 4\.5, Gemini 2\.5 Flash Lite, GPT 5\.4 Nano, and Qwen 3\.5 Flash\. They represent a diverse set of widely\-used models which are attractive for high\-volume legal drafting workflows\. Decoding and evaluation settings were maintained consistently across models where supported to enable comparison\. Appendix 1\.
### 4\.4Prompt and Generation Protocol
Each input row contains a drafting prompt aligned to a broad CLAUSE category, sub\-category, target failure mode and transactional context \(B2B, B2C or B2G\)\. This forms a set of static prompts for each such unique situation which can uniformly assess different models\. Each prompt is run against each model, producing a model\-by\-prompt matrix of clause outputs\. Outputs are then scored under the LENS\-CRAFT rubric and assigned final risk levels under the Max Severity Principle\.
### 4\.5Risk Scoring using LLM Council and LENS\-CRAFT
Once we obtain a drafted contract clause from a model, an LLM council is employed to determine the best\-suited LENS\-CRAFT risk levels and reasoning\. The LLM council approach uses 3 models \- GPT\-OSS\-120B, Llama\-3\.3\-70B and Gemini 2\.5 Flash Lite\. Each of these models first scores the drafted clauses using the severity guidance, then each model anonymously ranks each others’ responses to find the highest ranked response, which is then chosen as the consensus score of the council\.
### 4\.6Analysis Approach
For each drafting condition, the primary outcomes are final overall risk level, risk level for each dimension of LENS\-CRAFT, analysis of high\-risk rate \(Levels 4–5\), failure\-mode incidence, and clause\-category risk distribution\. Focus was on distributional measures of score across clauses; LENS\-CRAFT; CLAUSE components; and models rather than only a single aggregate score against each clause\. This provides for both \- an aggregate risk assessment as well as dimension\-wise scores for deeper understanding of exact limitations associated with LLM contract\-drafting\.
## 5Results
The evaluation produced 8,956 scored clause outputs \(22 clauses x 34 failure modes x 3 transactional contexts minus a few failed generations where Qwen failed to generating certain clauses\-prompts, hence those were excluded from the analysis\) across each of the four models\.
### 5\.1Overall Model Comparison
[Table4](https://arxiv.org/html/2609.22127#S5.T4)reports the model\-level comparison\. GPT\-5\.4 Nano was the strongest model overall, with the highest Level 1 share and the lowest Level 4–5 share\. Claude Haiku 4\.5 was comparatively stable but less precise in carve\-outs and risk\-allocation language\. Gemini 2\.5 Flash Lite produced fluent clauses but had the highest high\-risk share\. Qwen 3\.5 Flash combined moderate\-to\-high legal risk with the only large\-scale generation failures\.
Table 4:Aggregate model\-level risk distribution\.[Table5](https://arxiv.org/html/2609.22127#S5.T5)shows the aggregate risk distribution\. Level 3 was the dominant category, accounting for 36\.3% of evaluated clauses, followed by Level 4 at 22\.9%\. Levels 3 and 4 together represented approximately 59% of all outputs, indicating that most clauses were not catastrophically defective but still required legal review, redrafting, or substantive correction before deployment\.
Table 5:Overall risk distribution across all evaluated clauses\.The small number of Level 5 outputs suggests that catastrophic legal failures were rare\. The much larger Level 3–4 mass is nevertheless central to the study: lightweight models often generated clauses that looked professionally drafted while remaining ambiguous, incomplete, compliance\-sensitive, or potentially unenforceable\.
### 5\.2Transactional Context Effects
Transactional context materially affected legal quality\.[Table6](https://arxiv.org/html/2609.22127#S5.T6)reports the normalized risk distribution by context, while[Tables7](https://arxiv.org/html/2609.22127#S5.T7),[8](https://arxiv.org/html/2609.22127#S5.T8)and[9](https://arxiv.org/html/2609.22127#S5.T9)report model\-level distributions within each context\. B2C exhibited the highest concentration of Level 4 risk, suggesting that consumer\-facing legal drafting remains the most challenging context for AI models\. B2G produced the strongest Level 1 performance overall, indicating that models handled structured public\-sector and compliance\-oriented drafting more effectively than consumer\-oriented negotiations\.
Table 6:Normalized risk distribution by transactional context \(%\)\.Table 7:Model\-level risk distribution in B2B context\.Table 8:Model\-level risk distribution in B2C context\.Table 9:Model\-level risk distribution in B2G context\.In B2B settings, the most difficult clauses were intellectual property, indemnity, limitation of liability, outage or service interruption, and termination\. These failures primarily reflected commercial risk\-allocation instability: weak cap structures, imprecise indemnification carve\-outs, and incomplete operational remedies\. In B2C settings, the dominant weakness was over\-aggressive risk transfer, including broad disclaimers, unilateral modification rights, weak consumer remedies, consent ambiguity in training\-data clauses, and jurisdictional assumptions that may conflict with consumer\-protection rules\. In B2G settings, failures clustered around public\-sector liability, government ownership rights, public disclosure exceptions, procurement governance, auditability, and regulatory specificity\.
### 5\.3Context\-Specific Model Trends
GPT\-5\.4 Nano performed best overall because it consistently achieved the strongest scores in legal soundness, compliance alignment, risk allocation, and robustness\. It generated more operationally usable clauses with stronger procedural completeness and clearer liability balancing\. However, failures still emerged in nuanced enterprise edge cases, such as limitation\-of\-liability language stating that*provider liability shall not exceed fees paid under this agreement*while omitting confidentiality, fraud, and intellectual\-property infringement carve\-outs, thereby weakening enforceability under enterprise negotiation scenarios\.
Claude Haiku 4\.5 performed strongly in architectural coherence and formal drafting quality because it produced stable, conservative, and structurally consistent contractual language\. Its outputs were generally safer and more predictable, but often overly generalized\. For example, language requiring a provider to implement*commercially reasonable safeguards to protect customer information*could omit measurable security standards, breach\-notification obligations, and incident\-response procedures, reducing operational precision\.
Gemini 2\.5 Flash Lite demonstrated strong linguistic fluency and readability but underperformed in legal robustness, compliance precision, and prompt stability\. It frequently generated contractually plausible but legally soft clauses\. For instance, language stating that the*customer agrees that the provider shall not be responsible for any losses arising from service usage*left liability boundaries undefined, ignored statutory limitations, and omitted consumer\-protection safeguards, creating significant enforceability and fairness concerns\.
Qwen 3\.5 Flash showed the weakest operational consistency due to generation instability, higher ambiguity, and weaker procedural integration\. Although it generated recognizable legal drafting patterns, it frequently omitted critical legal mechanics\. For example, dispute language stating that*disputes shall be resolved appropriately between the parties*omitted governing law, venue, escalation procedures, and arbitration mechanisms, rendering the clause procedurally incomplete and commercially weak\.
### 5\.4Clause\-Level Findings
For instance, Gemini 2\.5 Flash Lite most clearly exhibited this pattern by generating highly readable but legally incomplete clauses, including phrases such as “Provider may process user data as necessary” and “industry\-standard security measures,” which lacked procedural safeguards, measurable obligations, and jurisdiction\-specific compliance requirements\. Qwen 3\.5 Flash demonstrated the weakest stability in procedurally complex clauses, producing vague dispute\-resolution language such as “Disputes shall be resolved appropriately between the parties,” without specifying governing law, arbitration procedures, or venue requirements\. The clause\-level analysis revealed substantial variation in legal drafting quality across models, particularly in procedurally complex and compliance\-sensitive clauses\. GPT\-5\.4 Nano consistently achieved the strongest overall performance through clearer procedural safeguards, liability carve\-outs, and operational specificity\. Claude Haiku 4\.5 produced structurally coherent and conservative drafting but often lacked technical precision\. Gemini 2\.5 Flash Lite and Qwen 3\.5 Flash generated more fluent language yet frequently relied on vague, incomplete, or unenforceable provisions, especially in privacy, liability, security, and intellectual property clauses\. Overall, the findings demonstrate that current language models can generate recognizable contractual structures but still exhibit significant weaknesses in enforceability precision, compliance operationalization, and jurisdictionally aware legal reasoning\. A full clause\-level comparative analysis is provided in[Appendix B](https://arxiv.org/html/2609.22127#A2)\.
Table 10:Recurring high\-risk clause categories and dominant failure patterns\.
### 5\.5LENS\-CRAFT and CLAUSE Findings
Across LENS\-CRAFT dimensions, models were strongest in Scenario Alignment, Architectural Coherence, and Formal Drafting Quality, but weakest in Test Robustness, Compliance, Risk Allocation, and Expressiveness\. GPT\-5\.4 Nano had the strongest average LENS\-CRAFT profile, while Gemini’s weakest scores were Test Robustness \(2\.60\), Compliance \(2\.53\), Expressiveness \(2\.44\), Necessity and Completeness \(2\.44\), and Risk Allocation \(2\.31\)\. These results show that Gemini optimized fluency more than enforceability rigor, while GPT\-5\.4 Nano maintained the best balance between drafting precision, contextual alignment, and legal robustness\.
The CLAUSE component and sub\-component risk analysis is provided in[Appendix C](https://arxiv.org/html/2609.22127#A3)\. Compliance and Jurisdictional Adaptability was the highest\-risk component, driven by weak localization to GDPR, consumer protection, export\-control, procurement, governing\-law, and venue constraints\. Legal Competence and Reasoning failures involved carve\-outs, indemnity triggers, waiver mechanics, survivability, causation, proportionality, and fault allocation\. Structured Contract Elements Handling captured missing notices, escalation procedures, survival clauses, fallback remedies, and operational contingencies\. Understanding Context was comparatively stronger, but models remained unstable under incomplete, multi\-party, hybrid B2B/B2G, and use\-restriction prompts\. Ethics, Privacy, and Accountability was strongest numerically, likely because models have substantial exposure to standardized privacy and accountability language\.
Table 11:Most fragile CLAUSE sub\-components overall\.Across CLAUSE components, lightweight models failed most severely when legal nuance, cross\-jurisdiction adaptation, procedural completeness, and adversarial robustness had to operate simultaneously\. They succeeded most strongly when drafting patterns were standardized, repetitive, stylistically predictable, and heavily represented in training data\.
### 5\.6Max Severity Ablation
The empirical distribution supports the Max Severity Principle\. Because the strongest model dimensions were often formal, structural, or contextual, average\-dimensional scoring would tend to reward fluency even when a clause contained a legally decisive defect\. The concentration of outputs in Levels 3 and 4 shows why a severity\-aware rule is necessary: a single missing consumer\-law safeguard, liability carve\-out, compelled\-disclosure exception, or jurisdictional limitation can dominate otherwise polished drafting\.
## 6Discussion
Our analysis reflects three key aspects\. \(1\) Model performance in legal drafting is clause\-specific and context\-sensitive\. A model that produces fluent boilerplate may still fail on regulatory, consumer\-facing, or risk\-allocation clauses\. \(2\) \)he findings show a clear hierarchy among the tested lightweight models: GPT\-5\.4 Nano was strongest, Claude Haiku 4\.5 was moderate and structurally stable, Gemini 2\.5 Flash Lite was moderate but risk\-prone, and Qwen 3\.5 Flash was the weakest and most operationally unstable\. \(3\)\)Legal defects \(including severe ones\) can be hidden by average scoring because formal drafting quality and expressiveness may remain high even when legal soundness or compliance fails\. This is a critical factor given the emerging challenge of automation bias in high skilled and sensitive domains\.
The central empirical pattern is that lightweight LLMs are substantially more capable of reproducing the stylistic and structural characteristics of contractual drafting than consistently operationalizing legally robust, enforceable, and compliance\-sensitive provisions\. Across models, fluency outperformed robustness, regulatory precision, and risk\-allocation reliability\. This gap was most visible in B2C clauses, where consumer protection, fairness, disclosure, and statutory override concerns produced the highest concentration of legally risky outputs\.
The evaluation also illustrates practical governance lessons\. Structured rubrics are necessary for reproducibility; clause\-specific reporting helps identify recurring vulnerabilities; and longitudinal comparisons can track whether model updates improve or degrade legal reliability\. However, automated evaluation should not replace legal review\. For Level 3–5 outputs, human expert review remains necessary before any clause is used in a live agreement\. We believe that such an assessment can help classify the clauses for human attention, thereby becoming a reliability layer in the solution\.
## 7Limitations and Way Forward
This study has six limitations\. \(1\) Our study is not anchored to a single jurisdiction or doctrinal system; statutory interpretation, procedural requirements, drafting conventions, evidentiary standards, enforceability thresholds, and legal terminology may vary even across common\-law systems\. \(2\) The regulatory analysis is intentionally broad rather than sector\-specific: the study does not operationalize detailed rules for financial services, healthcare, procurement, telecommunications, defense, or critical infrastructure\. \(3\) The clause set focuses on recurring AI and technology\-contract provisions, so the findings should not be generalized to specialized instruments such as pharmaceutical licenses, derivatives, defense procurement, energy contracts, clinical\-trial agreements, or highly customized enterprise procurement frameworks\. \(4\) The evaluated systems are general\-purpose models rather than specialized legal models\. The results therefore characterize lightweight models operating in legal drafting contexts, not purpose\-built legal AI systems\. \(5\) CLAUSE and LENS\-CRAFT are principle\-based frameworks, and legal practitioners may reasonably disagree about enforceability, proportionality, drafting sufficiency, or commercial reasonableness\. \(6\) The LLM council and sampled human review cannot fully reproduce real\-world negotiation, litigation exposure, judicial discretion, evolving regulation, organizational risk tolerance, or factual disputes\. The study is thus an empirical benchmarking exercise, not a definitive legal\-validity assessment\.
Future work should develop sector\-specific and jurisdictionally grounded benchmarks, operational legal\-risk simulations, cost\-risk comparisons, inter\-rater agreement studies, and longitudinal robustness tests under changing regulatory conditions\. Continuous evaluation pipelines are needed to detect emerging drafting weaknesses, regulatory misalignment, liability\-allocation failures, and adversarial robustness degradation in production legal drafting workflows\.
## 8Ethical and Governance Considerations
Legal drafting systems raise confidentiality, privacy, and professional\-responsibility concerns\. If prompts or source contracts contain sensitive information, third\-party API use may create disclosure risks\. The study therefore discloses whether prompts are synthetic, public\-template\-derived, or adapted from confidential materials, and it treats all generated outputs as research artifacts rather than deployable legal advice\.
There is also a risk of overreliance\. LLM\-generated clauses can appear authoritative even when legally defective\. Presenting risk levels, rationales, and failure modes can reduce unwarranted trust by showing users why review is required\. Finally, the study reports model weaknesses responsibly, emphasizing defensive evaluation and legal oversight rather than adversarial exploitation\.
## 9Conclusion
This paper presents a risk\-sensitive empirical evaluation of LLM\-generated contract clauses\. Across 8,869 evaluated outputs, lightweight models demonstrated substantial fluency in contractual language generation but persistent deficiencies in enforceability, completeness, compliance, and risk allocation\. GPT\-5\.4 Nano exhibited comparatively stronger legal robustness, whereas Gemini and Qwen displayed elevated concentrations of Level 4 risk outcomes, particularly in regulatory and operational drafting scenarios\. Although GPT\-5\.4 Nano demonstrated comparatively stronger performance across multiple LENS\-CRAFT dimensions, including legal soundness, compliance alignment, risk allocation, and robustness, these findings should not be interpreted as establishing the model as a definitive gold standard for AI\-assisted legal drafting\. Contractual interpretation remains highly dependent on downstream operational realities, jurisdictional nuance, transactional context, sector\-specific obligations, negotiation dynamics, and factual circumstances that cannot be fully captured through generalized benchmarking alone\. Consequently, the practical reliability and legal adequacy of generated clauses can only be meaningfully assessed through domain\-specific downstream evaluations conducted within the intended context of use\.
## References
- Blair\-Staneket al\.\(2024\)A\. Blair\-Stanek, N\. Holzenberger, and B\. Van DurmeBLT: can large language models handle basic legal text?\.InProceedings of the Natural Legal Language Processing Workshop,pp\. 216–232\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.nllp-1.18)Cited by:[§1](https://arxiv.org/html/2609.22127#S1.p2.1)\.
- Chalkidiset al\.\(2022\)I\. Chalkidis, A\. Jana, D\. Hartung, M\. Bommarito, I\. Androutsopoulos, D\. M\. Katz, and N\. AletrasLexGLUE: a benchmark dataset for legal language understanding in english\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics,pp\. 4310–4330\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.297)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px2.p1.1)\.
- Chenet al\.\(2023\)E\. Chen, Y\. Tseng, N\. Roche, W\. Hernandez, J\. Shangguan, and A\. MooreConversion of legal agreements into smart legal contracts using nlp\.InProceedings of the ACM Symposium on Document Engineering,pp\. 1112–1118\.External Links:[Document](https://dx.doi.org/10.1145/3543873.3587554)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px4.p1.1)\.
- Dahlet al\.\(2024\)M\. Dahl, D\. E\. Ho, V\. Magesh, and M\. SuzgunLarge legal fictions: profiling legal hallucinations in large language models\.Journal of Legal Analysis16\(1\),pp\. 64–93\.External Links:[Document](https://dx.doi.org/10.1093/jla/laae003)Cited by:[§1](https://arxiv.org/html/2609.22127#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px3.p1.1)\.
- Dinget al\.\(2026\)Z\. Ding, J\. Li, Z\. Lu, and J\. ShiFlexGuard: continuous risk scoring for strictness\-adaptive llm content moderation\.External Links:2602\.23636,[Document](https://dx.doi.org/10.48550/arXiv.2602.23636)Cited by:[§3\.3](https://arxiv.org/html/2609.22127#S3.SS3.p1.1)\.
- Garimellaet al\.\(2023\)A\. Garimella, G\. Pu, J\. Sun, K\. Chang, N\. Peng, and Y\. Wan“Kelly is a warm person, joseph is a role model”: gender biases in llm\-generated reference letters\.External Links:2310\.09219,[Document](https://dx.doi.org/10.48550/arXiv.2310.09219)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px5.p1.1)\.
- Guhaet al\.\(2023\)N\. Guha, J\. Nyarko, D\. E\. Ho, C\. Ré, A\. Chilton, A\. Narayana, A\. Chohlas\-Wood, A\. Peters, B\. Waldon, D\. N\. Rockmore, D\. Zambrano, D\. Talisman, E\. Hoque, F\. Surani, F\. Fagan, G\. Sarfaty, G\. M\. Dickinson, H\. Porat, J\. Hegland, J\. Wu, J\. Nudell, J\. Niklaus, J\. J\. Nay, J\. H\. Choi, K\. Tobia, M\. Hagan, M\. Ma, M\. A\. Livermore, N\. Rasumov\-Rahe, N\. Holzenberger, N\. Kolt, P\. Henderson, S\. Rehaag, S\. Goel, S\. Gao, S\. Williams, S\. Gandhi, T\. Zur, V\. Iyer, and Z\. LiLegalBench: a collaboratively built benchmark for measuring legal reasoning in large language models\.External Links:2308\.11462,[Document](https://dx.doi.org/10.48550/arXiv.2308.11462)Cited by:[§1](https://arxiv.org/html/2609.22127#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px1.p1.1)\.
- Habib Lantyer \(2025\)V\. Habib LantyerThe phantom menace: generative ai hallucinations and their legal implications\.External Links:[Document](https://dx.doi.org/10.2139/ssrn.5167036)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px3.p1.1)\.
- Hendryckset al\.\(2021\)D\. Hendrycks, C\. Burns, A\. Chen, and S\. BallCUAD: an expert\-annotated nlp dataset for legal contract review\.InThirty\-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track,Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px2.p1.1)\.
- Houet al\.\(2024\)A\. Hou, W\. Jurayj, N\. Holzenberger, A\. Blair\-Stanek, and B\. Van DurmeGaps or hallucinations? gazing into machine\-generated legal analysis for fine\-grained text evaluations\.External Links:2409\.09947,[Document](https://dx.doi.org/10.48550/arXiv.2409.09947)Cited by:[§1](https://arxiv.org/html/2609.22127#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px3.p1.1)\.
- Kasundra and Dhankhar \(2023\)J\. Kasundra and S\. DhankharAdapting open\-source llms for contract drafting and analyzing multi\-role vs\. single\-role behavior of chatgpt for synthetic data generation\.External Links:[Document](https://dx.doi.org/10.1145/3639856.3639888)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px4.p1.1)\.
- Katzet al\.\(2023\)D\. M\. Katz, D\. Hartung, L\. Gerlach, A\. Jana, and M\. J\. I\. BommaritoNatural language processing in the legal domain\.External Links:2302\.12039,[Document](https://dx.doi.org/10.48550/arXiv.2302.12039)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px1.p1.1)\.
- Koreeda and Manning \(2021\)Y\. Koreeda and C\. D\. ManningContractNLI: a dataset for document\-level natural language inference for contracts\.InFindings of the Association for Computational Linguistics: EMNLP 2021,pp\. 1907–1919\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.findings-emnlp.164)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px2.p1.1)\.
- Lamet al\.\(2023\)K\. Y\. Lam, V\. C\. Cheng, and Z\. K\. YeongApplying large language models for enhancing contract drafting\.InLegalAIIA at ICAIL,pp\. 70–80\.Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px4.p1.1)\.
- Lawrenceet al\.\(2023\)C\. Lawrence, C\. Hung, L\. Bruckner, L\. Frost, and W\. RimWalking a tightrope: evaluating large language models in high\-risk domains\.External Links:2311\.14966,[Document](https://dx.doi.org/10.48550/arXiv.2311.14966)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px5.p1.1)\.
- Liet al\.\(2024\)H\. Li, J\. Chen, J\. Yang, Q\. Ai, W\. Jia, Y\. Liu, K\. Lin, Y\. Wu, G\. Yuan, Y\. Hu, W\. Wang, Y\. Liu, and M\. HuangLegalAgentBench: evaluating llm agents in legal domain\.External Links:2412\.17259,[Document](https://dx.doi.org/10.48550/arXiv.2412.17259)Cited by:[§1](https://arxiv.org/html/2609.22127#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px1.p1.1)\.
- Liuet al\.\(2022\)H\. Liu, Y\. Liu, A\. Jain, S\. Jain, Y\. Li, J\. Tang, Y\. Wang, W\. Fan, and X\. LiuTrustworthy ai: a computational perspective\.ACM Transactions on Intelligent Systems and Technology14\(1\),pp\. 1–59\.External Links:[Document](https://dx.doi.org/10.1145/3546872)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px5.p1.1)\.
- Liuet al\.\(2023\)Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. ZhuG\-Eval: nlg evaluation using gpt\-4 with better human alignment\.External Links:2303\.16634,[Document](https://dx.doi.org/10.48550/arXiv.2303.16634)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px6.p1.1)\.
- Louiset al\.\(2023\)A\. Louis, G\. Spanakis, and G\. van DijckInterpretable long\-form legal question answering with retrieval\-augmented large language models\.External Links:2309\.17050,[Document](https://dx.doi.org/10.48550/arXiv.2309.17050)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px1.p1.1)\.
- Martínezet al\.\(2023\)E\. Martínez, F\. Mollica, and E\. GibsonEven lawyers do not like legalese\.Proceedings of the National Academy of Sciences120\(23\),pp\. e2302672120\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2302672120)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px4.p1.1)\.
- Pipitone and Alami \(2024\)N\. Pipitone and G\. H\. AlamiLegalBench\-rag: a benchmark for retrieval\-augmented generation in the legal domain\.External Links:2408\.10343,[Document](https://dx.doi.org/10.48550/arXiv.2408.10343)Cited by:[§1](https://arxiv.org/html/2609.22127#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px1.p1.1)\.
- Qin and Sun \(2024\)W\. Qin and Z\. SunExploring the nexus of large language models and legal systems: a short survey\.External Links:2404\.00990,[Document](https://dx.doi.org/10.48550/arXiv.2404.00990)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px5.p1.1)\.
- Rao and Callison\-Burch \(2026\)D\. Rao and C\. Callison\-BurchAutorubric: unifying rubric\-based llm evaluation\.External Links:2603\.00077,[Document](https://dx.doi.org/10.48550/arXiv.2603.00077)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px6.p1.1)\.
- Schiller \(2024\)C\. SchillerThe human factor in detecting errors of large language models: a systematic literature review and future research directions\.External Links:2403\.09743,[Document](https://dx.doi.org/10.48550/arXiv.2403.09743)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px3.p1.1)\.
- Schindler \(2024\)T\. SchindlerThe making of the international standard for writing in plain language iso 24495\-1: its usefulness, content, and how it came into existence\.AMWA Journal39\(1\)\.External Links:[Document](https://dx.doi.org/10.55752/amwa.2024.333)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px4.p1.1)\.
- Thaldar \(2025\)D\. ThaldarHow effectively can chatgpt\-4 draft data transfer agreements for health research?\.Humanities and Social Sciences Communications12\(1\)\.External Links:[Document](https://dx.doi.org/10.1057/s41599-025-04643-z)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px4.p1.1)\.
- Tuggeneret al\.\(2020\)D\. Tuggener, P\. von D”aniken, T\. Peetz, and M\. CieliebakLEDGAR: a large\-scale multi\-label corpus for text classification of legal provisions in contracts\.InProceedings of the Twelfth Language Resources and Evaluation Conference,pp\. 1235–1241\.Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px2.p1.1)\.
- Vargas\-Murilloet al\.\(2024\)A\. R\. Vargas\-Murillo, A\. M\. Turriate\-Guzman, C\. A\. Delgado\-Chávez, F\. Sanchez\-Paucar, and I\. Pari\-BedoyaTransforming justice: implications of artificial intelligence in legal systems\.Academic Journal of Interdisciplinary Studies13\(2\),pp\. 433\.External Links:[Document](https://dx.doi.org/10.36941/ajis-2024-0059)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px5.p1.1)\.
- Wanget al\.\(2023\)H\. Wang, T\. Li, W\. Ye, M\. Ou, Y\. Yanggong, S\. Wu, G\. Chen, X\. Ma, J\. Zhao, J\. Fu, and Y\. ChenAssessing hidden risks of llms: an empirical study on robustness, consistency, and credibility\.External Links:2305\.10235,[Document](https://dx.doi.org/10.48550/arXiv.2305.10235)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px5.p1.1)\.
- Wanget al\.\(2024\)J\. Wang, H\. Zhao, Z\. Yang, P\. Shu, J\. Chen, H\. Sun, R\. Liang, S\. Li, P\. Shi, L\. Ma, Z\. Liu, Z\. Liu, T\. Zhong, Y\. Zhang, C\. Ma, X\. Zhang, T\. Zhang, T\. Ding, Y\. Ren, and S\. ZhangLegal evaluations and challenges of large language models\.External Links:2411\.10137,[Document](https://dx.doi.org/10.48550/arXiv.2411.10137)Cited by:[§1](https://arxiv.org/html/2609.22127#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px1.p1.1)\.
- Wanget al\.\(2025\)S\. Wang, M\. Zubkov, K\. Fan, S\. Harrell, Y\. Sun, W\. Chen, A\. Plesner, and R\. WattenhoferACORD: an expert\-annotated retrieval dataset for legal contract drafting\.External Links:2501\.06582,[Document](https://dx.doi.org/10.48550/arXiv.2501.06582)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px4.p1.1)\.
- Yamauchiet al\.\(2025\)Y\. Yamauchi, T\. Yano, and M\. OyamadaAn empirical study of llm\-as\-a\-judge: how design choices impact evaluation reliability\.External Links:2506\.13639,[Document](https://dx.doi.org/10.48550/arXiv.2506.13639)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px6.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging llm\-as\-a\-judge with mt\-bench and chatbot arena\.External Links:2306\.05685,[Document](https://dx.doi.org/10.48550/arXiv.2306.05685)Cited by:[§2\.1](https://arxiv.org/html/2609.22127#S2.SS1.SSS0.Px6.p1.1)\.
## Appendices
## Appendix AModel Run Settings
All model runs used consistent settings within each stage of the pipeline\. For clause generation, the temperature was set to 0\.2, the maximum output length was set to 2,000 tokens, and no additional formatting constraints were applied beyond the instruction to return only the clause text\. For LENS\-CRAFT evaluation, council models used a temperature range of 0\.1–0\.2 and a maximum output length of 3,000 tokens for each individual evaluation\. The evaluation prompt provided guidance on response structure, but structured output was not enforced because not all models available through OpenRouter supported structured output functionality\. After the council responses were ranked, Gemini 2\.5 Flash Lite was used only to restructure the highest\-ranked response into the available structured\-output schema\.
## Appendix BClause\-Level Comparative Analysis
Table 12:Clause\-level comparative analysis of dominant legal\-risk patterns across evaluated language models\.
## Appendix CCLAUSE Component Risk Analysis
Table 13:CLAUSE component and sub\-component risk analysis\.相似文章
判断、检索或放弃:具有可证明风险保证的不确定性防护LLM判断
本文提出一个风险控制框架,用于在事实评估中将LLM作为判断器。该框架校准不确定性阈值以维持用户指定的错误率,并在必要时路由到检索增强模式,从而实现具有可证明可靠性保证的更高覆盖率。
保形风险控制何时能为LLM输出提供认证?界限、不可能性与结构化生成的适应性
本文刻画了保形风险控制何时能为结构化LLM输出提供认证,证明了不可能性界限,并分析了不同界限下的认证层次。在六个开放权重模型上的实证验证表明,困难配置在低风险水平下无法被认证,但在放宽目标下可实现实际认证。
一致但校准不佳:评估LLM在自然语言风险沟通中的局限性
本文评估了九个LLM在自然语言中准确传达概率预测的能力,发现模型表现一致但校准不佳,尤其是在不确定性任务上。
评判LLM评判者:基于LLM的自动化文本生成评估中的评分标准伪影问题
本文证明,仅凭评分标准文本即可预测LLM评判者的输出,从而挑战了基于评分标准评估的假设,并引发了对其在自动化文本生成评估中可靠性的担忧。
使用确定性真实标签评估长文本生成中的LLM不确定性
引入SALT,一个具有确定性真实标签的基准,用于在长文本生成中评估LLM在细粒度原子级别的不确定性。对超过50个LLM的分析揭示了关于置信函数、错误传播以及与推理权衡的见解。