Tag
LEGO is a dual-module framework that synergizes Expert GraphRAG and Expert Chain-of-Thought to enhance complex legal reasoning in large language models, achieving superior performance on benchmarks like LawExamQA_Civil.
This paper introduces ImmigrationReason, a large-scale structured dataset of U.S. immigration appeals for legal reasoning research, addressing the gap in administrative adjudication data for NLP studies.
This small-scale study evaluates LLM reasoning in legal case forecasting using European Court of Human Rights cases, finding that models produce structurally complete but substantively shallow analyses and that LLM-based evaluators align weakly with human annotators.
The paper constructs a benchmark to evaluate LLMs on temporal legal reasoning, revealing biases towards applying the most recently enacted laws and an inverse relationship between general reasoning ability and temporal performance.
This paper presents HumRightsBench, the first expert-validated, scenario-based benchmark for evaluating large language models' legal reasoning about international human rights law. Pilot results on frontier models show overall accuracy between 0.339 and 0.577, highlighting notable gaps in detecting obligation violations.
This paper proposes the Mecellem semantic protocol, an ontologically grounded framework for artificial legal intelligence, arguing that legal reasoning requires dynamic, context-dependent meaning construction rather than mere codification or statistical pattern recognition.
This paper presents adaptive pipelines for legal retrieval, entailment, and judgment prediction tasks in the COLIEE 2026 competition, using multi-stage retrieval, reranking, and LLM-based reasoning.
The paper introduces L-MAD, a framework for systematically evaluating multi-agent debate structures in legal textual entailment. It finds that increasing agent population improves accuracy but more debate rounds cause over-deliberation, with improvements of up to 8% over single-agent baselines.
This paper investigates multi-agent deliberation methods for legal reasoning tasks using LLMs, introducing two novel frameworks inspired by courtroom procedures. The experiments show that multi-agent systems achieve comparable overall performance to monolithic LLMs but produce distinct answers and can solve cases that baselines fail, highlighting the potential of multi-agent approaches for legal AI.
Meta FAIR's latest paper proposes the Autodata method, which uses an intelligent data scientist Agent to autonomously generate and optimize high-quality data, enabling a 4B small model to defeat a 397B large model on legal reasoning tasks. This indicates that data quality can bridge the gap in parameter count, providing new insights for data pipelines and scaling.
TW-LegalBench is a benchmark for evaluating large language models on Taiwanese legal understanding, including over 16,000 multiple-choice questions, 117 essay questions, and 14,000 legal judgment prediction instances. Results show top models exceed the passing threshold for lawyers but fall short for judges, highlighting challenges in reliable legal text generation.
This paper presents an overview of the QIAS 2026 shared task on Islamic inheritance reasoning, evaluating LLMs on multi-step legal and numerical reasoning using the MAWARITH benchmark.
This paper presents the participation of team PSL in the QIAS 2026 Shared Task on Arabic Islamic inheritance reasoning, comparing commercial and open-source large language models. Results show commercial models (e.g., Gemini 2.5 Flash) significantly outperform open-source models in structured legal reasoning with multi-step dependencies.
This paper introduces EP-HUBO, a quantum-inspired method that treats evidence selection in chain-of-thought reasoning as a combinatorial optimization problem, significantly improving performance on legal reasoning benchmarks like MMLU-Pro law and LEXam by allowing minority-but-correct hypotheses to override noisy majorities.
This paper introduces a relevance-sensitive evaluation suite for legal AI, demonstrating that LLMs are overly sensitive to legally irrelevant perturbations, and proposes LexGuard, an adversarial multi-agent framework using formal reasoning to improve legal reasoning reliability.
This paper empirically studies LLMs' legal reasoning in tax law, showing that data contamination inflates performance and that neuro-symbolic hybrid systems offer more reliable and robust generalization than monolithic LLMs.
This paper identifies a systematic gap between legal interpretation and formal logic in AI legal reasoning, proposes a neuro-symbolic approach to bridge it, and demonstrates substantial label shifts when re-annotating legal NLI data under strict formal entailment.
This paper presents Qatar University's multi-stage QLoRA fine-tuning approach on Qwen3-4B for Arabic Islamic inheritance reasoning, achieving 90% MIR-E score through domain adaptation on Islamic fatwa records followed by task-specific training on 12,000 structured inheritance cases, matching commercial systems like Gemini-2.5-flash with minimal computational resources.
VLegal-Bench is a cognitively grounded benchmark for evaluating large language models on Vietnamese legal reasoning tasks, containing 10,450 expert-annotated samples designed to address the gap in legal benchmarks for civil law systems. The benchmark assesses multiple levels of legal understanding through question answering, multi-step reasoning, and scenario-based problem solving, providing a replicable framework for evaluating LLMs in non-English, codified legal contexts.