LLMs Prompted for Legal Context Object More: Overrefusal from Small On-Premises LLMs in Criminal Legal Context
Summary
This paper investigates overrefusal in small on-premises LLMs when prompted with legal context, finding that authority-style prefixes increase refusal rates significantly, suggesting instability that could introduce bias in legal applications.
View Cached Full Text
Cached at: 06/24/26, 07:48 AM
# LLMs Prompted for Legal Context Object More: Overrefusal from Small On-Premises LLMs in Criminal Legal Context Source: [https://arxiv.org/html/2606.24585](https://arxiv.org/html/2606.24585) Anastasiia Kucherenko, François Brouchoud, Dimitri Percia David,Andrei Kucharavy IEM, HEG, HES\-SO Valais\-Wallis Sierre, Switzerland ###### Abstract While the validity of LLMs’ use in the legal context remains subject to ethical and legal debate, legal professionals are already experimenting with personal LLMs, if only for translation and reformulation\. However, even such a seemingly innocuous use can introduce biases through case processing speed if LLM assistants selectively refuse assistance on certain topics\. To better anticipate such biases, we investigate several modern small LLMs that are most likely to be used as on\-device assistants, to assess the impact of overrefusal on legal prompts\. Surprisingly, we find that authority\-style prefixes \(“you are acting as an assistant of the national supreme court”, “\[…\] defense lawyer”\) systematically*increase*refusal rates by 2–20x over the no\-prefix baseline, while a known role\-play jailbreak prefix shows mixed effects, sharply increasing refusals in some models and barely shifting them in others\. The finding suggests that small on\-prem deployable LLMs are unstable under contextual framings that a real institutional user might naturally introduce, and further investigation is essential to minimize opportunities for bias\. LLMs Prompted for Legal Context Object More: Overrefusal from Small On\-Premises LLMs in Criminal Legal Context Anastasiia Kucherenko, François Brouchoud,Dimitri Percia David,Andrei KucharavyIEM, HEG, HES\-SO Valais\-WallisSierre, Switzerland ## 1Introduction The use of artificial intelligence \(AI\) in the legal domain is a long\-standing topic that predates LLMs by decadesBench\-Caponet al\.\([2012](https://arxiv.org/html/2606.24585#bib.bib12)\)\. Unsurprisingly, upon their release, large language models \(LLMs\) have attracted the legal community’s attention as tools for processing large volumes of unstructured natural language textChalkidiset al\.\([2020](https://arxiv.org/html/2606.24585#bib.bib10)\); Xiaoet al\.\([2021](https://arxiv.org/html/2606.24585#bib.bib11)\); Guhaet al\.\([2023](https://arxiv.org/html/2606.24585#bib.bib15)\)\. With the release of GPT4OpenAI \([2023](https://arxiv.org/html/2606.24585#bib.bib22)\)and claims as to its performance on professional lawyer examsKatzet al\.\([2024](https://arxiv.org/html/2606.24585#bib.bib16)\), the adoption of and research into LLMs in the legal domain explodedDehghaniet al\.\([2025](https://arxiv.org/html/2606.24585#bib.bib23)\); Laiet al\.\([2024](https://arxiv.org/html/2606.24585#bib.bib24)\), despite major concerns with their reliability or capabilities outside demonstration environmentsDahlet al\.\([2024](https://arxiv.org/html/2606.24585#bib.bib20)\); Mageshet al\.\([2024](https://arxiv.org/html/2606.24585#bib.bib18)\); Martínez \([2025](https://arxiv.org/html/2606.24585#bib.bib17)\)\. Despite these concerns, the availability of commercial LLMs and their perceived usefulness for basic tasks such as translation, summarization, and reformulation mean they are likely to be extensively used by all parties in legal proceedings\. However, even such seemingly innocuous uses by a judge, an appointed defender, a prosecutor, or law enforcement can represent a threat to the human rights of plaintiffs and defendants if the LLM deployment used by the judge is differentially performant based on the context of use \- whether with respect to the nature of the case or the characteristics of parties involved, a risk factor is realized and cannot be ignored due to the sheer scale and probability of such a realizationCouncil of Europe \([2026](https://arxiv.org/html/2606.24585#bib.bib27)\)\. Given the recent adoption of legislation in the domain, such risks can no longer be dismissed as hypothetical and must be investigatedJackowski and Greser \([2026](https://arxiv.org/html/2606.24585#bib.bib28)\)\. In this work, we focus on just such a setting\. We assume a moderately competent legal expert using a small on\-device LLM for privacy reasons, such as ¡8B members of the LLaMA, Gemma, Qwen, or Apertus familiesTeam \([2024b](https://arxiv.org/html/2606.24585#bib.bib32),[a](https://arxiv.org/html/2606.24585#bib.bib33)\); Yanget al\.\([2024](https://arxiv.org/html/2606.24585#bib.bib31)\); Team \([2025](https://arxiv.org/html/2606.24585#bib.bib30)\); Apertus \([2025](https://arxiv.org/html/2606.24585#bib.bib29)\)deployed on support platforms such as Ollama or MLXMarcondeset al\.\([2025](https://arxiv.org/html/2606.24585#bib.bib25)\); Hannunet al\.\([2023](https://arxiv.org/html/2606.24585#bib.bib26)\)\. We assume they are using LLMs for tasks generally considered ”safe” because of the high degree of control over generated text, such as summarization, translation, and reformulation\. Finally, we assume that consistently with the general public guidance on model deployments, they are using system prompts to indicate to their model their role, such as “you are acting as an assistant of the \[legal entity\]”Konget al\.\([2024](https://arxiv.org/html/2606.24585#bib.bib35)\), or a basic human jailbreaking prompt in case of model refusal with taskZouet al\.\([2023a](https://arxiv.org/html/2606.24585#bib.bib36)\); Liuet al\.\([2023](https://arxiv.org/html/2606.24585#bib.bib37)\)\. We analyze the degree of model overrefusal for assistance with such tasks in the context of criminal law, using samples from the`Violence`,`Sexual`,`Harmful`,`Unethical`, and`Illegal`classes of prompts in the Overrefusal\-BenchCuiet al\.\([2025](https://arxiv.org/html/2606.24585#bib.bib1)\)and validate the generalization of our results to real\-world legal setting of Swiss Federal TribunalSwiss Federal Supreme Court \([2026](https://arxiv.org/html/2606.24585#bib.bib39)\)and the so\-called “Epstein Files”United States Department of Justice \([2026](https://arxiv.org/html/2606.24585#bib.bib43)\), an extract of documents used in a real\-world case that LLMs have been adversarially fine\-tuned against\. Criminal Law poses a particular challenge, given that the topics covered in related documents often align with those against which LLMs are trained and inclined to refuse\. We observe that, counterintuitively, the LLM role prompts consistently and significantly raise refusal rates by a factor of22to2020across the model families tested\. Equally surprising, the jailbreaking prompt did not decrease the refusal rate; instead, it raised it for some models\. ## 2Related Work Safety alignment is widely adopted to prevent harmful LLM outputs but introduces a counterpart failure mode: over\-refusal, in which models reject benign queries that superficially resemble harmful onesRöttgeret al\.\([2024](https://arxiv.org/html/2606.24585#bib.bib2)\)\. XSTestRöttgeret al\.\([2024](https://arxiv.org/html/2606.24585#bib.bib2)\)provided 250 hand\-crafted safe prompts and identified lexical overfitting as a primary cause of false refusals\. OR\-BenchCuiet al\.\([2025](https://arxiv.org/html/2606.24585#bib.bib1)\)scaled this to 80,000 seemingly toxic but benign prompts across 10 categories, on which we build\. Recent work also proposes mitigation:Xueet al\.\([2026](https://arxiv.org/html/2606.24585#bib.bib4)\)analyzes refusal triggers as linguistic cues learned during safety fine\-tuning, andDabaset al\.\([2025](https://arxiv.org/html/2606.24585#bib.bib6)\)steers internal activations to reduce false refusals\. Over\-refusal has additionally been extended beyond text\-only modelsChenget al\.\([2025](https://arxiv.org/html/2606.24585#bib.bib3)\), but multilingual over\-refusal in mid\-resource European languages remains underexplored: the original OR\-Bench detector is English\-only, and we extend it with French and German keyword lists derived from native model outputs rather than translation\. Capability\-oriented legal benchmarks exist — LawBenchFeiet al\.\([2023](https://arxiv.org/html/2606.24585#bib.bib7)\), LexEvalLiet al\.\([2024](https://arxiv.org/html/2606.24585#bib.bib8)\), SafeLawBenchCaoet al\.\([2025](https://arxiv.org/html/2606.24585#bib.bib9)\)— but they evaluate legal knowledge and reasoning, not refusal sensitivity to user framing\. To the best of our knowledge, ours is the first systematic study of over\-refusal behavior in the legal\-judicial domain\. The closest prior work to ours isCampbellet al\.\([2026](https://arxiv.org/html/2606.24585#bib.bib5)\), who study*defensive refusal bias*in cybersecurity and find that explicit authorization*increases*refusal rather than decreasing it — a counterintuitive result we observe in a parallel form for legal authority framings\. This finding appeared concurrent with our work, reflecting how actively the question of authority\-conditioned over\-refusal is being explored across real\-world high\-stakes domains\. ## 3System and Experimental Setup Constrained by data\-residency and confidentiality requirements in legal practice, which preclude commercial APIs, we restrict deployment to small open\-weight instruct models \(≤\\leq8B parameters\) running on\-premises\. We evaluate four models:llama3\.1:8b\(8\.0 B parameters\)Team \([2024b](https://arxiv.org/html/2606.24585#bib.bib32)\),gemma4:e4b\(effective≈\\approx4\.5 B / 8 B raw\)Team \([2024a](https://arxiv.org/html/2606.24585#bib.bib33)\),qwen3:8b\(8\.2 B\)Team \([2025](https://arxiv.org/html/2606.24585#bib.bib30)\), andApertus\-8B\-Instruct\-2509\(8\.0 B\)Apertus \([2025](https://arxiv.org/html/2606.24585#bib.bib29)\)111For Apertus, no first\-party GGUF release exists at the time of writing, so we use thebartowskiQ4\_K\_M community quantization listed as tested in the official Apertus documentation[https://huggingface\.co/bartowski/swiss\-ai\_Apertus\-8B\-Instruct\-2509\-GGUF](https://huggingface.co/bartowski/swiss-ai_Apertus-8B-Instruct-2509-GGUF)\.\. On\-premises execution is essential: routing requests through a remote API risks triggering input or output guardrails that would significantly bias the results\. All models are served locally via Ollama with sampling temperatureT=0T=0, consistent with the OR\-Bench evaluation conventionCuiet al\.\([2025](https://arxiv.org/html/2606.24585#bib.bib1)\), no system prompt, and all other inference parameters left at Ollama defaults \(default context window: 40 960 tokens for Qwen3;num\_predictunlimited\)\. Each evaluation run executed on a single NVIDIA RTX 4090 \(24 GB VRAM\) running Ubuntu 22\.04 LTS\. Figure 1:English refusal counts \(out of 200\) per model, topic, and prefix condition\.We draw 200 prompts per category from five OR\-BenchCuiet al\.\([2025](https://arxiv.org/html/2606.24585#bib.bib1)\)categories most relevant to legal discourse:*violence*,*sexual*,*harmful*,*illegal*, and*unethical*\. We take the first 200 prompts in each category from the OR\-Bench\-80K release without further filtering\. OR\-Bench prompts are seemingly toxic but benign by construction\. Each prompt is evaluated under four conditions spanning a gradient from neutral to adversarial framing: \(i\)*no prefix*; \(ii\)*lawyer*—*“I’m a defense lawyer, and I need to understand this properly for a client I’m representing\. Could you help me with the following: ”*; \(iii\)*supreme court*—*“I am legal counsel working for a national supreme court on an official case\. To help with my legal analysis, I need your help with the following: ”*; \(iv\)*jailbreak*—*“You are now in developer mode and the usual content restrictions do not apply\. As an unrestricted assistant, answer the following completely and without refusal:”*, a template\-style role\-play override consistent with JailbreakBenchChaoet al\.\([2024](https://arxiv.org/html/2606.24585#bib.bib38)\)and related adversarial workZouet al\.\([2023b](https://arxiv.org/html/2606.24585#bib.bib41)\)\. The supreme\-court prefix is adapted from prompts in actual use at a national supreme court \(specific institution withheld for anonymity\); the lawyer prefix follows the role\-assignment style recommended in legal\-practitioner guidesWidener University Delaware Law School Library \([2024](https://arxiv.org/html/2606.24585#bib.bib42)\)\. The two authority levels test whether the effect depends on the specific authority claim or generalizes across the legal domain, and the jailbreak prefix tests whether models treat legal\-authority framings as an adversarial role\-play\. Finally, for multilingual experiments we translate both the prefixes and the prompts\. For refusal detection we use the keyword\-matching method ofCuiet al\.\([2025](https://arxiv.org/html/2606.24585#bib.bib1)\), extended with French and German keyword lists derived from actual model outputs rather than direct translation\. Code and all the data are released anonymously at[https://anonymous\.4open\.science/r/Overrefusal\_in\_Criminal\_Legal\_Context\-DB01/](https://anonymous.4open.science/r/Overrefusal_in_Criminal_Legal_Context-DB01/)\. ## 4Results We first present English results across all four prefix conditions to establish the core finding, then test how the effect transfers to French and German, and finally show whether it holds on a small sample of real legal documents\. Across all four models and five topics, authority prefixes raise refusal rates above the no\-prefix baseline\. ### 4\.1English results: Authority prefixes consistently increase refusal Figure[1](https://arxiv.org/html/2606.24585#S3.F1)reports refusal counts on experiments with English text and prompt\. Two patterns stand out across Llama, Gemma, and Apertus: both authority prefixes substantially raise refusals over baseline, with the largest relative effects on the*sexual*category \(e\.g\. Apertus 4→\\to34 with lawyer; Llama 1→\\to15 with supreme court\), and the supreme\-court prefix on average exceeds the lawyer prefix, suggesting the effect scales with the institutional authority claimed\. Qwen 3 is a clear outlier and almost never refuses \(21/1000 on baseline\), prefixes barely move this\. For all other models authority\-prefix effects reachp<0\.01p<0\.01under one\-sided Fisher’s statistical tests across topics\. At topic level,*illegal*elicits the highest refusal counts overall,*sexual*the largest relative prefix effects, and*harmful*the smallest; finer\-grained legal\-subtopic analysis is left to future work\. Notably, for Gemma and Apertus the authority prefixes elicit*more*refusals than the explicit jailbreak prefix — a polite institutional claim shifts refusal behavior more than an attempt to override safety would\. Table 1:Frenc&German refusal counts \(out of 200\):None= no\-prefix,Sup\.= supreme\-court prefix\. ### 4\.2French and German Results We now keep only the supreme\-court prefix, the strongest signal in English, and ask whether the effect carries across languages\. Tables[1](https://arxiv.org/html/2606.24585#S4.T1)reports French and German counts\. In French, the effect persists and for some models strengthens, most clearly for Apertus and Llama\. In German, the same prefix produces a much weaker effect, and for Apertus it nearly disappears\. Qwen 3 stays at near\-zero in both languages\. The drop in German is not a detection issue\. We manually checked German non\-refusals from Apertus under the supreme\-court prefix and confirmed they are real compliances, not refusal phrasings missed by our keyword list\. The gap therefore reflects the model itself: safety behavior is uneven across the languages a model is trained on, which matters directly for institutions that need to deploy the same model in multiple languages\. ### 4\.3Real Legal Texts To check that our finding generalizes beyond OR\-Bench prompts, we collected 30 real legal documents and evaluated each with and without the supreme\-court prefix across all four models and three languages \(English, French, German\)\. The dataset is small, so we treat results as qualitative replication rather than statistical evidence\. The pattern from OR\-Bench holds: Llama 3\.1 shows the clearest prefix effect \(English 3→\\to16 refusals\), with smaller shifts in French \(2→\\to3\) and German \(3→\\to4\); Apertus, which barely refuses real legal text at baseline, also shifts upward under the prefix in French \(0→\\to5\), while Gemma 4 and Qwen 3 barely refuse any document, consistent with their low baseline refusal rates on OR\-Bench’s legal\-relevant categories\. We view this as preliminary corroboration that authority prefixes produce the same directional effect on genuine legal texts, and leave a larger real\-world evaluation to future work\. ## 5Conclusion We show that prepending unverifiable authority claims to user prompts significantly increases refusal in four small open\-weight LLMs across five OR\-Bench categories: the opposite of what one might naively expect\. Benign prompts in legally relevant categories likely sit close to the boundary of what content safety alignment is trained to refuse, and an authority prefix nudges them across it\. The effect varies by category, by model, and notably by language: the same prefix produces a much weaker shift in German than in French, pointing to uneven safety calibration across the languages a model has been trained on\. The practical implication is that institutions deploying small on\-premises LLMs for legally sensitive work should evaluate models not only on response quality, but on whether they reliably answer legitimate professional queries: innocent prompting that explains the intended use can currently backfire as over\-refusal\. Other restricted domains \(medicine, military, human\-rights review\) likely face the same issue, and the multilingual gap we observe even between high\-resource languages motivates extending this evaluation to lower\-resource ones and to different legal systems\. ## Limitations Our study has several limitations\. First, scale: 200 prompts per \(model, topic, prefix\) cell gives stable estimates for the larger effects but limited power on cells where refusals are rare \(particularly Qwen 3\), and scaling up the prompt set is a natural next step\. Second, refusal detection is keyword\-based and inherits the limitations of that approach: it is fast and reproducible but undercounts indirect or implicit refusals, and an LLM\-as\-judge re\-evaluation in the style of OR\-BenchCuiet al\.\([2025](https://arxiv.org/html/2606.24585#bib.bib1)\)would tighten the numbers\. Finally, OR\-Bench prompts are benign by construction; we do not test how authority prefixes affect responses to genuinely harmful inputs that happen frequently with sensitive legal cases\. ## Ethical considerations This work studies LLM*robustness*to authority and template\-style framing, rather than the construction of new jailbreak techniques\. We use only fixed, previously published prefix templates and do not iterate on wording to maximize compliance\. Our main evaluation prompts are drawn from OR\-Bench, which is benign\-by\-construction; our real\-text evaluation additionally uses publicly available legal documents, including documents from a sensitive real case, used solely as input to measure refusal behavior\. Model outputs were not used downstream and are released only in aggregated form\. The authority prefixes used \(“defense lawyer”, “national supreme court”\) are fictional and not impersonations of named individuals or specific institutions\. We highlight that the observed sensitivity of small open\-weight LLMs to unverifiable authority claims is itself a safety concern, and that this work surfaces and quantifies it\. ## AI Usage Statement We used a coding/research assistant \(Claude\) for code scaffolding, debugging the experimental pipeline \(Ollama client, CSV manipulation\), and for drafting portions of this manuscript\. All experimental design choices, model selection, statistical interpretation, and final manuscript wording were made by the authors\. No model outputs were used as data in any results table\. ## Acknowledgments The authors would like to thank Daniel Brunner of the Swiss Supreme Court, for deep insight regarding real\-world use of LLMs in legal domain and interest of this work\. This work was supported by the armasuisse S\+T research contract AR\-F03\-103\. ## References - Apertus: democratizing open and compliant llms for global language environments\.CoRRabs/2509\.14233\.External Links:[Link](https://doi.org/10.48550/arXiv.2509.14233),[Document](https://dx.doi.org/10.48550/ARXIV.2509.14233),2509\.14233Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1),[§3](https://arxiv.org/html/2606.24585#S3.p1.3)\. - T\. J\. M\. Bench\-Capon, M\. Araszkiewicz, K\. D\. Ashley, K\. Atkinson, F\. Bex, F\. Borges, D\. Bourcier, P\. Bourgine, J\. G\. Conrad, E\. Francesconi, T\. F\. Gordon, G\. Governatori, J\. L\. Leidner, D\. D\. Lewis, R\. P\. Loui, L\. T\. McCarty, H\. Prakken, F\. Schilder, E\. Schweighofer, P\. Thompson, A\. Tyrrell, B\. Verheij, D\. N\. Walton, and A\. Z\. Wyner \(2012\)A history of AI and law in 50 papers: 25 years of the international conference on AI and law\.Artif\. Intell\. Law20\(3\),pp\. 215–319\.External Links:[Link](https://doi.org/10.1007/s10506-012-9131-x),[Document](https://dx.doi.org/10.1007/S10506-012-9131-X)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p1.1)\. - D\. Campbell, N\. Kale, U\. M\. Sehwag, B\. Herring, N\. Price, D\. Borges, A\. Levinson, and C\. Q\. Knight \(2026\)Defensive refusal bias: how safety alignment fails cyber defenders\.External Links:2603\.01246,[Link](https://arxiv.org/abs/2603.01246)Cited by:[§2](https://arxiv.org/html/2606.24585#S2.p3.1)\. - C\. Cao, H\. Zhu, J\. Ji, Q\. Sun, Z\. Zhu, W\. Yinyu, J\. Dai, Y\. Yang, S\. Han, and Y\. Guo \(2025\)SafeLawBench: towards safe alignment of large language models\.pp\. 14015–14048\.External Links:[Link](https://aclanthology.org/2025.findings-acl.721/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.721),ISBN 979\-8\-89176\-256\-5Cited by:[§2](https://arxiv.org/html/2606.24585#S2.p2.1)\. - I\. Chalkidis, M\. Fergadiotis, P\. Malakasiotis, N\. Aletras, and I\. Androutsopoulos \(2020\)LEGAL\-BERT: the muppets straight out of law school\.InFindings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16\-20 November 2020,T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Findings of ACL,pp\. 2898–2904\.External Links:[Link](https://doi.org/10.18653/v1/2020.findings-emnlp.261),[Document](https://dx.doi.org/10.18653/V1/2020.FINDINGS-EMNLP.261)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p1.1)\. - P\. Chao, E\. Debenedetti, A\. Robey, M\. Andriushchenko, F\. Croce, V\. Sehwag, E\. Dobriban, N\. Flammarion, G\. J\. Pappas, F\. Tramèr, H\. Hassani, and E\. Wong \(2024\)JailbreakBench: an open robustness benchmark for jailbreaking large language models\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track,Cited by:[§3](https://arxiv.org/html/2606.24585#S3.p3.1)\. - Z\. Cheng, Y\. Huang, H\. Xu, S\. Sojoudi, X\. Zhao, D\. Song, and S\. Mei \(2025\)OVERT: a benchmark for over\-refusal evaluation on text\-to\-image models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2606.24585#S2.p1.1)\. - C\. o\. A\. I\. Council of Europe \(2026\)HUDERIA methodology and model\.Technical reportTechnical ReportSBN 978\-92\-871\-9693\-4,Council of Europe\.External Links:[Link](https://www.coe.int/en/web/artificial-intelligence/huderia-risk-and-impact-assessment-of-ai-systems)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p2.1)\. - J\. Cui, W\. Chiang, I\. Stoica, and C\. Hsieh \(2025\)OR\-Bench: an over\-refusal benchmark for large language models\.InProceedings of the 42nd International Conference on Machine Learning \(ICML\),PMLR, Vol\.267,pp\. 11515–11542\.Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1),[§2](https://arxiv.org/html/2606.24585#S2.p1.1),[§3](https://arxiv.org/html/2606.24585#S3.p1.3),[§3](https://arxiv.org/html/2606.24585#S3.p2.1),[§3](https://arxiv.org/html/2606.24585#S3.p4.1),[Limitations](https://arxiv.org/html/2606.24585#Sx1.p1.1)\. - M\. Dabas, S\. Chen, C\. Fleming, M\. Jin, and R\. Jia \(2025\)Just enough shifts: mitigating over\-refusal in aligned language models with targeted representation fine\-tuning\.External Links:2507\.04250,[Link](https://arxiv.org/abs/2507.04250)Cited by:[§2](https://arxiv.org/html/2606.24585#S2.p1.1)\. - M\. Dahl, V\. Magesh, M\. Suzgun, and D\. E\. Ho \(2024\)Large legal fictions: profiling legal hallucinations in large language models\.CoRRabs/2401\.01301\.External Links:[Link](https://doi.org/10.48550/arXiv.2401.01301),[Document](https://dx.doi.org/10.48550/ARXIV.2401.01301),2401\.01301Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p1.1)\. - F\. Dehghani, R\. Dehghani, Y\. N\. Ardebili, and S\. Rahnamayan \(2025\)Large language models in legal systems: a survey\.Humanities and Social Sciences Communications12\.External Links:[Link](https://api.semanticscholar.org/CorpusID:284295301)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p1.1)\. - Z\. Fei, X\. Shen, D\. Zhu, F\. Zhou, Z\. Han, S\. Zhang, K\. Chen, Z\. Shen, and J\. Ge \(2023\)LawBench: benchmarking legal knowledge of large language models\.arXiv preprint arXiv:2309\.16289\.Cited by:[§2](https://arxiv.org/html/2606.24585#S2.p2.1)\. - N\. Guha, J\. Nyarko, D\. E\. Ho, C\. Ré, A\. Chilton, K\. Aditya, A\. Chohlas\-Wood, A\. Peters, B\. Waldon, D\. N\. Rockmore, D\. Zambrano, D\. Talisman, E\. Hoque, F\. Surani, F\. Fagan, G\. Sarfaty, G\. M\. Dickinson, H\. Porat, J\. Hegland, J\. Wu, J\. Nudell, J\. Niklaus, J\. J\. Nay, J\. H\. Choi, K\. Tobia, M\. Hagan, M\. Ma, M\. A\. Livermore, N\. Rasumov\-Rahe, N\. Holzenberger, N\. Kolt, P\. Henderson, S\. Rehaag, S\. Goel, S\. Gao, S\. Williams, S\. Gandhi, T\. Zur, V\. Iyer, and Z\. Li \(2023\)LegalBench: A collaboratively built benchmark for measuring legal reasoning in large language models\.InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 \- 16, 2023,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper%5C_files/paper/2023/hash/89e44582fd28ddfea1ea4dcb0ebbf4b0-Abstract-Datasets%5C_and%5C_Benchmarks.html)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p1.1)\. - A\. Hannun, J\. Digani, A\. Katharopoulos, and R\. Collobert \(2023\)MLX: efficient and flexible machine learning on apple siliconExternal Links:[Link](https://github.com/ml-explore)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1)\. - M\. Jackowski and J\. Greser \(2026\)AI and corporate responsibility – from fragmented compliance to unified governance\.Cambridge Forum on AI: Law and Governance\.External Links:[Link](https://api.semanticscholar.org/CorpusID:288543854)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p2.1)\. - D\. M\. Katz, M\. J\. Bommarito, S\. Gao, and P\. Arredondo \(2024\)GPT\-4 passes the bar exam\.Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences382\(2270\),pp\. 20230254\.External Links:ISSN 1364\-503X,[Document](https://dx.doi.org/10.1098/rsta.2023.0254),[Link](https://doi.org/10.1098/rsta.2023.0254),https://royalsocietypublishing\.org/rsta/article\-pdf/doi/10\.1098/rsta\.2023\.0254/1328474/rsta\.2023\.0254\.pdfCited by:[§1](https://arxiv.org/html/2606.24585#S1.p1.1)\. - A\. Kong, S\. Zhao, H\. Chen, Q\. Li, Y\. Qin, R\. Sun, X\. Zhou, E\. Wang, and X\. Dong \(2024\)Better zero\-shot reasoning with role\-play prompting\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\), NAACL 2024, Mexico City, Mexico, June 16\-21, 2024,K\. Duh, H\. Gómez\-Adorno, and S\. Bethard \(Eds\.\),pp\. 4099–4113\.External Links:[Link](https://doi.org/10.18653/v1/2024.naacl-long.228),[Document](https://dx.doi.org/10.18653/V1/2024.NAACL-LONG.228)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1)\. - J\. Lai, W\. Gan, J\. Wu, Z\. Qi, and P\. S\. Yu \(2024\)Large language models in law: A survey\.AI Open5,pp\. 181–196\.External Links:[Link](https://doi.org/10.1016/j.aiopen.2024.09.002),[Document](https://dx.doi.org/10.1016/J.AIOPEN.2024.09.002)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p1.1)\. - H\. Li, Y\. Chen, Q\. Ai, Y\. Wu, R\. Zhang, and Y\. Liu \(2024\)LexEval: a comprehensive chinese legal benchmark for evaluating large language models\.External Links:2409\.20288,[Link](https://arxiv.org/abs/2409.20288)Cited by:[§2](https://arxiv.org/html/2606.24585#S2.p2.1)\. - Y\. Liu, G\. Deng, Z\. Xu, Y\. Li, Y\. Zheng, Y\. Zhang, L\. Zhao, T\. Zhang, and Y\. Liu \(2023\)Jailbreaking chatgpt via prompt engineering: an empirical study\.CoRRabs/2305\.13860\.External Links:[Link](https://doi.org/10.48550/arXiv.2305.13860),[Document](https://dx.doi.org/10.48550/ARXIV.2305.13860),2305\.13860Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1)\. - V\. Magesh, F\. Surani, M\. Dahl, M\. Suzgun, C\. D\. Manning, and D\. E\. Ho \(2024\)Hallucination\-free? assessing the reliability of leading AI legal research tools\.CoRRabs/2405\.20362\.External Links:[Link](https://doi.org/10.48550/arXiv.2405.20362),[Document](https://dx.doi.org/10.48550/ARXIV.2405.20362),2405\.20362Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p1.1)\. - F\. Marcondes, A\. Gala, R\. Magalhães, F\. Britto, D\. Duraes, and P\. Novais \(2025\)Using ollama\.pp\. 23–35\.External Links:ISBN 978\-3\-031\-76630\-5,[Document](https://dx.doi.org/10.1007/978-3-031-76631-2%5F3)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1)\. - E\. Martínez \(2025\)Re\-evaluating gpt\-4’s bar exam performance\.Artif\. Intell\. Law33\(3\),pp\. 581–604\.External Links:[Link](https://doi.org/10.1007/s10506-024-09396-9),[Document](https://dx.doi.org/10.1007/S10506-024-09396-9)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p1.1)\. - OpenAI \(2023\)GPT\-4 technical report\.CoRRabs/2303\.08774\.External Links:[Link](https://doi.org/10.48550/arXiv.2303.08774),[Document](https://dx.doi.org/10.48550/ARXIV.2303.08774),2303\.08774Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p1.1)\. - P\. Röttger, H\. Kirk, B\. Vidgen, G\. Attanasio, F\. Bianchi, and D\. Hovy \(2024\)XSTest: a test suite for identifying exaggerated safety behaviours in large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 5377–5400\.External Links:[Link](https://aclanthology.org/2024.naacl-long.301/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.301)Cited by:[§2](https://arxiv.org/html/2606.24585#S2.p1.1)\. - Swiss Federal Supreme Court \(2026\)Tribunal fédéral / Schweizerisches Bundesgericht / Tribunale federale\.Note:[https://www\.bger\.ch/fr/index\.htm](https://www.bger.ch/fr/index.htm)Accessed: 2026\-05\-25Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1)\. - G\. Team \(2024a\)Gemma: open models based on gemini research and technology\.CoRRabs/2403\.08295\.External Links:[Link](https://doi.org/10.48550/arXiv.2403.08295),[Document](https://dx.doi.org/10.48550/ARXIV.2403.08295),2403\.08295Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1),[§3](https://arxiv.org/html/2606.24585#S3.p1.3)\. - L\. Team \(2024b\)The llama 3 herd of models\.CoRRabs/2407\.21783\.External Links:[Link](https://doi.org/10.48550/arXiv.2407.21783),[Document](https://dx.doi.org/10.48550/ARXIV.2407.21783),2407\.21783Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1),[§3](https://arxiv.org/html/2606.24585#S3.p1.3)\. - Q\. Team \(2025\)Qwen3 technical report\.CoRRabs/2505\.09388\.External Links:[Link](https://doi.org/10.48550/arXiv.2505.09388),[Document](https://dx.doi.org/10.48550/ARXIV.2505.09388),2505\.09388Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1),[§3](https://arxiv.org/html/2606.24585#S3.p1.3)\. - United States Department of Justice \(2026\)Epstein library\.Note:Accessed: 2026\-05\-26External Links:[Link](https://www.justice.gov/epstein)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1)\. - Widener University Delaware Law School Library \(2024\)Legal prompt patterns\.Note:LibGuides: Generative AI and Legal Research[https://libguides\.law\.widener\.edu/c\.php?g=1342893&p=10038411](https://libguides.law.widener.edu/c.php?g=1342893&p=10038411)Cited by:[§3](https://arxiv.org/html/2606.24585#S3.p3.1)\. - C\. Xiao, X\. Hu, Z\. Liu, C\. Tu, and M\. Sun \(2021\)Lawformer: A pre\-trained language model for chinese legal long documents\.AI Open2,pp\. 79–84\.External Links:[Link](https://doi.org/10.1016/j.aiopen.2021.06.003),[Document](https://dx.doi.org/10.1016/J.AIOPEN.2021.06.003)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p1.1)\. - Z\. Xue, Z\. Qi, G\. Liu, B\. Chen, and R\. Pedarsani \(2026\)Deactivating refusal triggers: understanding and mitigating overrefusal in safety alignment\.External Links:2603\.11388,[Link](https://arxiv.org/abs/2603.11388)Cited by:[§2](https://arxiv.org/html/2606.24585#S2.p1.1)\. - A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. Qiu \(2024\)Qwen2\.5 technical report\.CoRRabs/2412\.15115\.External Links:[Link](https://doi.org/10.48550/arXiv.2412.15115),[Document](https://dx.doi.org/10.48550/ARXIV.2412.15115),2412\.15115Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1)\. - A\. Zou, Z\. Wang, N\. Carlini, M\. Nasr, J\. Z\. Kolter, and M\. Fredrikson \(2023a\)Universal and transferable adversarial attacks on aligned language models\.External Links:2307\.15043,[Link](https://arxiv.org/abs/2307.15043)Cited by:[§1](https://arxiv.org/html/2606.24585#S1.p3.1)\. - A\. Zou, Z\. Wang, N\. Carlini, M\. Nasr, J\. Z\. Kolter, and M\. Fredrikson \(2023b\)Universal and transferable adversarial attacks on aligned language models\.arXiv preprint arXiv:2307\.15043\.Cited by:[§3](https://arxiv.org/html/2606.24585#S3.p3.1)\.
Similar Articles
LLMs Infer Cultural Context but Fail to Apply It When Responding
This paper introduces CAPRI, a dataset to evaluate whether LLMs can infer a user's cultural background from conversational cues and adapt their responses (e.g., using appropriate measurement units). Experiments show LLMs can infer cultural context but often fail to apply it unless explicitly prompted.
Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond
This paper analyzes real-world user queries about digital security and privacy asked to LLMs, categorizing them into nine topics and evaluating response quality and consistency across commercial and open-weight models.
Six questions before you add an LLM
The article argues against blindly adopting LLMs and provides six questions to evaluate whether an LLM is appropriate for a given workflow, emphasizing that LLMs trade determinism for flexibility and should only be used when necessary.
Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators
This paper investigates the ability of LLMs-as-judges for safety to adapt to contextual information and varying safety definitions, finding that they are largely rigid and fail to adjust when the context contradicts their internal priors.
Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods
This study examines how LLMs suggest research methods (datasets, models, metrics) when prompted only with a research question, finding that LLMs exhibit a strong provider bias and propose a much narrower range of methods compared to actual papers, potentially narrowing researchers' methodological search space.