LegalPincite: Multi-level Legal Information Retrieval Dataset
Summary
Introduces LegalPincite, a large-scale legal information retrieval dataset built from CJEU judgments, featuring masked queries, full corpora, and paragraph-level citation annotations to enable multi-level retrieval evaluation.
View Cached Full Text
Cached at: 08/05/26, 05:46 PM
Paper page - LegalPincite: Multi-level Legal Information Retrieval Dataset
Source: https://huggingface.co/papers/2608.03756 Published on Aug 4
·
Submitted byhttps://huggingface.co/theresiavr
TVRon Aug 5
Abstract
AcommontaskinlegalInformationRetrieval(IR)istofindrelevantlegalsourcesfromcase-lawcollections.Whilelegalpracticeoftenrequirespinpointcitations(pincites)tospecificcaseparagraphs,mostexistingpubliclegalIRdatasetslackparagraph-levelcitationannotations.Yet,publiclyavailabledatasetswithsuchinformationcontaindataleakageinthequerytextandexcludeparagraphsthatareneithercitingnorcitedfromthecorpora,creatinganunrealisticandoversimplifiedretrievalsetting,potentiallyleadingtoinflatedperformance.Toaddresstheselimitations,wecontributealarge-scalelegalIRdatasetconstructedfromCourtofJusticeoftheEuropeanUnion(CJEU)judgments.Thedatasetcontains:(i)maskedcase/paragraphqueries,withremovedcitationinformation;(ii)acorpusthatincludesallparagraphs;and(iii)case-andparagraph-levelground-truthcitations,withpartialhumanexpertvalidation.OurdatasetsupportsboththedevelopmentandrigorousevaluationoflegalIRmethods,atmultiplequery-documentlevels(case-to-case,paragraph-to-case,andparagraph-to-paragraphretrieval).Linktodataset:https://huggingface.co/datasets/theresiavr/legalpincite
View arXiv pageView PDFProject pageGitHubAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.03756 in a model README.md to link it from this page.
Datasets citing this paper1
#### theresiavr/legalpincite Viewer• Updated35 minutes ago • 1.47M • 320 • 1
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.03756 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
PIIBench: A Unified Multi-Source Benchmark Corpus for Personally Identifiable Information Detection
PIIBench presents a unified multi-source benchmark corpus for detecting personally identifiable information (PII) across diverse data sources. This resource addresses the need for standardized evaluation in PII detection tasks, which is critical for privacy-preserving NLP applications.
NOWJ@COLIEE 2026: Adaptive Pipelines for Legal Retrieval and Reasoning
This paper presents adaptive pipelines for legal retrieval, entailment, and judgment prediction tasks in the COLIEE 2026 competition, using multi-stage retrieval, reranking, and LLM-based reasoning.
Automatic Construction of a Legal Citation Graph from 100 Million Ukrainian Court Decisions: Large-Scale Extraction, Topological Analysis, and Ontology-Driven Clustering
This paper constructs the first large-scale citation graph from 100.7 million Ukrainian court decisions, extracting over 500 million citation links. It demonstrates that the citation structure can automatically recover legal domain boundaries and predict legislative importance with near-perfect accuracy, and releases the pipeline and data as open resources.
From Judgments to Issues: Structured Extraction of Legal Reasoning with Citation-Hallucination Control
This paper presents an automated pipeline that uses the DeepSeek V3 model to decompose Italian tax-court judgments into individual legal issues structured in XML following the IRAC framework, and includes a hallucination-detection filter using the Linkoln parser to validate citations, validated by expert annotators.
CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law
Introduces CanLegalRAGBench, a benchmark for evaluating retrieval-augmented generation on Canadian case law using realistic queries and expert-annotated answers. The evaluation reveals sensitivity to design choices, competitiveness of open-source embedding models, and persistent hallucinations in generated answers.