在句子层面分类解释准则:来自 German Federal Constitutional Court 的基准
摘要
本文介绍了一个用于在法律文件中分类解释准则的句子层面基准,评估了 LLMs 在 German Federal Constitutional Court 数据集上的表现。
arXiv:2609.26945v1 Announce Type: new
Abstract: Judicial reasoning remains challenging for large language models (LLMs) to analyze. This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in the tradition of Savigny. Our contributions are threefold. First, we operationalize this conception of interpretation as classification criteria. Second, we provide a dataset of decisions of the German Federal Constitutional Court annotated at the sentence level. Third, we report baseline evaluations of four LLMs from three model families under expert hand-written prompts, compared against prompts optimized with Genetic-Pareto (GEPA). Mean F1 over the seven binary subtasks clusters between 70.4 and 79.2 across models, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest; under the tested configuration, GEPA-optimized prompts do not systematically outperform the hand-written ones, suggesting that the expert prompts provide a meaningful baseline.
查看缓存全文
缓存时间: 2026/09/24 09:11
# Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court Source: [https://arxiv.org/html/2609.26945](https://arxiv.org/html/2609.26945) Felix RingeAffiliation:Department of Law, Freie Universität Berlin, Berlin, GermanyCorrespondence to:[felix\.ringe@fu\-berlin\.de](mailto:[email protected]) ###### Abstract Judicial reasoning remains challenging for large language models \(LLMs\) to analyze\. This paper contributes a sentence\-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in the tradition of Savigny\. Our contributions are threefold\. First, we operationalize this conception of interpretation as classification criteria\. Second, we provide a dataset of decisions of the German Federal Constitutional Court annotated at the sentence level\. Third, we report baseline evaluations of four LLMs from three model families under expert hand\-written prompts, compared against prompts optimized with Genetic\-Pareto \(GEPA\)\. MeanF1F\_\{1\}over the seven binary subtasks clusters between 70\.4 and 79\.2 across models, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest; under the tested configuration, GEPA\-optimized prompts do not systematically outperform the hand\-written ones, suggesting that the expert prompts provide a meaningful baseline\. ###### Keywords: legal NLP, argument mining, LLM evaluation, benchmark, statutory interpretation ## 1Introduction The systematic identification of the types of arguments that courts use to justify their decisions has long been a focus of empirical legal scholarship\. These efforts are commonly motivated by the idea that the core of just judicial decision\-making lies in upholding the guarantee of equal treatment\. This guarantee, in turn, finds expression in a transparent justification of decisions, and in particular in the consistent use of a recognized set of interpretive arguments\([Raisch, 1988](https://arxiv.org/html/2609.26945#bib.bib2)\)\. Such consistency is also said to make judicial decision\-making more predictable and to constrain judicial discretion\([Gorsuch, 2016](https://arxiv.org/html/2609.26945#bib.bib1)\)\. The capacity to study these arguments at scale has long been constrained by the manual effort required, forcing a careful selection of the materials to be analyzed\([Mendelson, 2018](https://arxiv.org/html/2609.26945#bib.bib3)\)\. Reliable classification of these argument types by LLMs promises to dramatically decrease the cost of such analysis and open up a host of substantive research questions\. For instance, this prospect has motivated work testing long\-standing narratives about the historical degree of formalism in courts for the United States\([Stiglitz and Thalken, 2024](https://arxiv.org/html/2609.26945#bib.bib4)\)and the Czech Republic\([Koref et al\., 2026](https://arxiv.org/html/2609.26945#bib.bib5)\)as well as analyses of reasoning styles of individual justices and the relationship between legal reasoning and other judicial features\([Thalken and Stiglitz, 2026](https://arxiv.org/html/2609.26945#bib.bib6)\)\. However, prior work has taken the paragraph as its modeling unit, which identifies argument types within paragraphs but not their precise locations\. This paper develops a finer\-grained approach to classifying interpretation, operating at the sentence level\. The analysis is built on the four interpretive canons of grammatical, systematic, historical, and objective\-teleological interpretation, usually traced to Friedrich Carl von Savigny and given their modern form, most prominently, by Karl Larenz\([Larenz, 1991](https://arxiv.org/html/2609.26945#bib.bib7)\)\. This conception has found resonance in many jurisdictions through frequent translation of his work, with editions published in Spain\([Larenz, 2023](https://arxiv.org/html/2609.26945#bib.bib8)\), Portugal\([Larenz, 2014](https://arxiv.org/html/2609.26945#bib.bib9)\), the People’s Republic of China\([Larenz, 2020](https://arxiv.org/html/2609.26945#bib.bib10)\), and South Korea\([Larenz, 2026](https://arxiv.org/html/2609.26945#bib.bib11)\), among others\. This paper makes three contributions: \(1\) it operationalizes Larenz’s conception of interpretation into granular classification criteria \([Section2](https://arxiv.org/html/2609.26945#S2)\), \(2\) it contributes an expert\-annotated dataset of decisions of the German Federal Constitutional Court that follows these criteria and records individual reasons for edge cases \([Section4](https://arxiv.org/html/2609.26945#S4)\), and \(3\) it reports baseline evaluations of four LLMs from three model families on the resulting subtasks \([Sections5](https://arxiv.org/html/2609.26945#S5)and[6](https://arxiv.org/html/2609.26945#S6)\)\.111Code available on[GitHub](https://github.com/KensingtonOscupant/classifying-interpretive-canons), datasets on[Hugging Face](https://huggingface.co/collections/felix453/interpretive-canons), and evaluation runs on[Weights & Biases](https://wandb.ai/icml-2026-ai4law/classifying-interpretive-canons-camera-ready/)\. Agent trajectories for successful reproductions of the[scores](https://hub.harborframework.com/jobs/35917940-c695-43ab-bdd9-17ddea56db05/trials/e8b5e216-2f0c-4cc3-b40f-d291e64c1833)and[datasets](https://hub.harborframework.com/jobs/676116e1-8968-4245-a097-c37e6c845ed6/trials/b0f8c8ef-d598-49c1-a5dd-a4acb8d364e9)can be inspected on Harbor Hub\. For more details on the datasets, see[Section4](https://arxiv.org/html/2609.26945#S4)\. We find that an AI agent can reproduce both the reported scores and benchmark datasets exactly from the respective input data and our methods descriptions, without access to our code \([AppendixD](https://arxiv.org/html/2609.26945#A4)\)\. ## 2Task Methodology To examine exactly where interpretation occurs in decisions, it is necessary to establish a working definition that allows for classification to be carried out\. Given the extensive literature on the concept of interpretation, one might expect formulating such a definition to be straightforward\. It is not, as there is no standardized definition:[Lüders and Stohlmann \(2025\)](https://arxiv.org/html/2609.26945#bib.bib12)report the same difficulty for defining proportionality, a related concept central to German constitutional reasoning\. Insightful general remarks on the concept of interpretation can, however, be found in Karl Larenz’s*Methodenlehre der Rechtswissenschaft*: > “To ‘interpret’ a text thus means to decide in favor of one among several possible readings on the basis of considerations \[…\]“\([Larenz, 1991](https://arxiv.org/html/2609.26945#bib.bib7), p\. 204\)222In the original: “Einen Text ‘auslegen’, heißt also, sich für eine unter mehreren möglichen Deutungen aufgrund von Überlegungen zu entscheiden \[…\]”\. Translation ours\. Several prerequisites can be derived from this definition\. First, one must decide on a reading\. Second, this must involve certain considerations\. Third, the decision in favor of the reading must follow precisely from those considerations\. In statutory interpretation, these considerations regularly draw on the interpretive canons which this work examines\. All of the above requirements will be examined more closely in what follows\. ### 2\.1Reading A reading within the meaning of this study is an assertion that the court itself advances and whose object is the claim that a particular statutory provision, in the abstract, has a particular content\. A detailed explanation follows in the sections below\. For the study of interpretive canons, the precise determination of the reading is decisive, because readings are the reference points of the interpretive canons\. Which of the interpretive canons are present can depend on the reading from which the consideration is viewed: ###### Example 1\. ‘‘Section 46\(1\), first sentence, no\. 2 BWahlG provides for the loss of the parliamentary seat in the case of a ‘redetermination’ of the election result\.’’333Translations of excerpts of court decisions in this paper are ours throughout\.\(BVerfGE 124, 1 \(15\)\)444Decisions are cited to the court’s official reports,*Entscheidungen des Bundesverfassungsgerichts*\(BVerfGE\), by volume, the page on which the decision begins, and, in parentheses, the specific page referred to; e\.g\., BVerfGE 124, 1 \(15\) is volume 124, decision beginning at page 1, cited at page 15\. In[Example1](https://arxiv.org/html/2609.26945#Thmexample1), one might assume at first glance that, because of the word “redetermination” in quotation marks, this is a grammatical interpretation of Section 46 BWahlG \(*Bundeswahlgesetz*, the Federal Electoral Act\)\. This picture changes once the preceding sentence is brought into view\. The expanded passage reads as follows: ###### Example 2\. “There are, however, indications supporting an interpretation to the effect that the ‘electoral act’ within the meaning of Section 37 BWahlG is completed on the evening of the main election and that a new ‘electoral act’ begins with the supplementary election: Section 46\(1\), first sentence, no\. 2 BWahlG provides for the loss of the parliamentary seat in the case of a ‘redetermination’ of the election result\.” \(BVerfGE 124, 1 \(15\)\) In[Example2](https://arxiv.org/html/2609.26945#Thmexample2), the reference point is no longer Section 46 BWahlG, as in[Example1](https://arxiv.org/html/2609.26945#Thmexample1), but Section 37 BWahlG\. From the perspective of Section 37 BWahlG, the statement about Section 46 BWahlG constitutes systematic interpretation\. From this vantage point, then, the statement about Section 46 BWahlG shifts from a grammatical interpretation to a systematic one\. The above definition of a reading identifies four prerequisites, which will be discussed in more detail below\. #### 2\.1\.1Abstract First, the assertion must be abstract\. A statement is abstract if it claims validity not only for the unique set of facts underlying the decision, but for an indeterminate range of situations in which the statutory provision or legal principle applies\. Statutory interpretation is always concerned with determining the meaning of a norm independently of the concrete facts\. This criterion also marks the distinction from subsumption, which brings a set of facts under a norm\. The boundary is at times not easy to draw\. ###### Example 3\. “Where the letter of a person held in pre\-trial detention is, as here, seized pursuant to the analogous application of Section 108 StPO and forwarded to the public prosecutor’s office for further action, the person concerned may apply to the public prosecutor’s office for the return of the letter, may at any time request a decision by the competent judge \(Section 98\(2\) StPO\), and may lodge a complaint against any seizure during the preliminary proceedings \(Section 304\(1\) StPO\)\.” \(BVerfGE 57, 170 \(181\)\) In[Example3](https://arxiv.org/html/2609.26945#Thmexample3), the phrase “as here” establishes a concrete connection to the facts, which weighs against an abstract statement\. The rest of the sentence, however, states a conclusion about the meaning of Section 98\(2\) and Section 304\(1\) StPO \(*Strafprozessordnung*, the Code of Criminal Procedure\) that applies to other cases as well\. Since the connection to the facts is established only incidentally, and the sentence does not aim to make the statement solely for this one situation, it is more apt to read the statement as abstract\. #### 2\.1\.2Assertion Furthermore, the statement must be an assertion\. An assertion is a statement against which a meaningful counterposition can be formed\. This is not the case where the statement is purely descriptive, for instance the verbatim reproduction of the statutory text\. This rests on the consideration that a reading fills the statutory provision precisely with content, and that this requires an assertion going beyond the mere statement of the wording\. While an argument employing an interpretive canon, for example in the context of literal interpretation, may well fall back on merely repeating the wording, this does not suffice for a reading\.555Of the four prerequisites, assertion is the one we do not operationalize as a stand\-alone classification subtask in this work; see[Section4](https://arxiv.org/html/2609.26945#S4)\. #### 2\.1\.3Reference Point: A Specific Statutory Provision with Determinate Content Furthermore, the assertion must relate to a specific statutory provision\. This restriction is helpful because it fixes the object of interpretation on an unambiguous reference point\. In particular, it excludes from consideration those statements that at first appear interpretive because they invoke a legal principle or general principle, but which in fact have no interpretable reference point in the written law\. Naming a provision is not sufficient, however: the assertion must also determine that provision’s content\. It does so when it has the effect that a range of cases falls under the provision while others do not\. An assertion that aims instead at declaring the provision as a whole void, valid, or inapplicable does not state a determinate content\. The practical purpose of this second half of the criterion is to exclude sentences whose object is merely the compatibility of one norm with another, for example the steps of a proportionality or constitutionality assessment, since such statements concern the relation between two norms rather than the content of either one\. ###### Example 4\. “Section 901 ZPO is consequently inapplicable where the debtor’s inability to pay is established, and to that extent cannot violate Art\. 2\(2\), second sentence, of the Basic Law\.” \(BVerfGE 61, 126\) [Example4](https://arxiv.org/html/2609.26945#Thmexample4)names two provisions precisely, yet it fixes the content of neither\. It states how the one relates to the other, which leaves open which cases Section 901 ZPO \(*Zivilprozessordnung*, the Code of Civil Procedure\) covers\. We treat reference point and determinate content as a single criterion because they share a reference point and are decided together in one pass over the sentence; the classification subtask in[Section4](https://arxiv.org/html/2609.26945#S4)accordingly asks for the provision the content of which the sentence determines, and returns nothing when no such provision exists\. #### 2\.1\.4Self\-advanced An assertion is self\-advanced when it has both its starting point and its end point in the reasoning of the deciding panel\. This is not the case when the court bases an assertion on a precedent\. That precedents, too, might fall under the concept of interpretation does not seem entirely far\-fetched, since they often themselves contain passages in which interpretation is carried out\. By referring to such a passage, the interpretation undertaken there is implicitly perpetuated, at least insofar as the court invokes the precedent approvingly\. However, it does not appear certain that the citation of a precedent permits an inference that the panel approves of every interpretation undertaken in that precedent\. Where an assertion is supported by several sources, for example a scholarly source, a legislative explanatory memorandum, and a precedent, not all of which are precedents, the assertion is assessed, in the panel’s favor, as self\-advanced\. ### 2\.2Considerations Interpretation means choosing one of several possible readings on the basis of considerations\. This gives a first criterion: a reading must be reached “on the basis of considerations”, i\.e\., there must be a causal link between the considerations and the reading the court adopts\. The four classical interpretive canons operate at the level of these considerations: grammatical, systematic, historical, and objective\-teleological interpretation\. There has long been consensus that there is no consensus on the precise definitions of the interpretive canons\([Kriele, 1967](https://arxiv.org/html/2609.26945#bib.bib13)\)\. The following definitions are based on the account given by Larenz: Grammatical interpretation is interpretation according to the meaning of an expression or a combination of words in general usage or, if such a usage can be identified, in the special usage of the statutory provision in question\([Larenz, 1991](https://arxiv.org/html/2609.26945#bib.bib7), p\. 321\)\. Systematic interpretation requires first and foremost attention to context, as is necessary for understanding any connected speech or writing\. Beyond this, it refers to the substantive consistency of provisions within a single body of rules, and further to attention to the outward arrangement of the statute\(up until here[Larenz, 1991](https://arxiv.org/html/2609.26945#bib.bib7), p\. 328\)as well as to provisions lying outside the statute that are relevant to its understanding\. Historical interpretation is interpretation according to the legislator’s regulatory intent\([Larenz, 1991](https://arxiv.org/html/2609.26945#bib.bib7), p\. 328\)\.666On Larenz’s account, the regulatory intent consists, in particular, of all those considerations that remained unchallenged in the deliberations\. He argues that the view of an individual ministerial official or politician does not permit an inference as to the legislator’s regulatory intent\. While this is plausible, this part of the definition was nevertheless not adopted into the definition used for this study\. Only in the rarest of cases will it be apparent to an LLM from the text of the decision itself whether a consideration remained unchallenged in the deliberations\. Furthermore, the definition would then no longer merely capture the occurrence of a particular interpretive canon as a phenomenon, but would also entail a value judgment as to whether the court conducted the interpretation in the individual case in a methodologically sound way\. Historical considerations that would commonly be subsumed under the term would thus run the risk of being left out of account\. Objective\-teleological interpretation means the interpretation of a statutory provision according to the structures of the regulated subject area and the underlying legal\-ethical principles\(up until here[Larenz, 1991](https://arxiv.org/html/2609.26945#bib.bib7), p\. 328\), without recourse to the legislator’s regulatory intent\. ## 3From Definition to a Sentence\-Level Task The criteria of[Section2](https://arxiv.org/html/2609.26945#S2)do not by themselves fix how a decision text is to be broken into units for classification\. This section motivates the unit of analysis and describes how the subtasks compose into a multi\-stage pipeline; the annotation scheme of[Section4](https://arxiv.org/html/2609.26945#S4)and the structure of the benchmark both follow from it\. ### 3\.1Unit of Analysis Ideally, one would extract each reading and each supporting consideration at its exact word boundaries, which need not coincide with sentence or paragraph boundaries\. The drawbacks of such free\-span extraction become apparent in an example: ###### Example 5\. “When the Bremen Leave Act was enacted and amended, postal employees were members of the administration of the United Economic Area\. That the Bremen Leave Act was, according to the will of the Bremen legislator, meant to extend to them as well was not expressly stated in the deliberations on the statute; it follows, however, beyond doubt from Section 1 BremUrlG\. Under that provision, the Act applies ‘… to the administrations and enterprises of the public service that have their seat in the Land of Bremen or operate in the Land of Bremen’\. In the deliberations it was expressly emphasized that the Leave Act was to apply to ‘all employees of the Bremen state territory’ \[…\]\. For the delimitation of the class of persons entitled, the place of service was thus to be decisive, and not the identity of the employer\.” \(BVerfGE 11, 89 \(95\)\) The reading lies in the second sentence and concerns the personal scope of Section 1 BremUrlG \(*Bremisches Urlaubsgesetz*, the Bremen Leave Act\)\. A model instructed to find all considerations supporting this reading, and to name the canons used in each, might point to the penultimate sentence, which draws on the parliamentary deliberations, and correctly identify a historical argument\. Yet the third sentence also carries a grammatical argument, since it invokes the wording of the provision\. Because the task was to find*all*uses of the canons in the passage, the answer would be incomplete and therefore wrong, and, more importantly, the implicit decision the model took against every other span in the passage cannot be checked\. Just as not every use of a canon catches a human reader’s eye on a cursory pass, LLMs often overlook arguments when asked to extract an unbounded number of them from a passage\. We therefore pose the task so that only a fixed set of answers is possible: every sentence is classified against every criterion it reaches, with missed arguments treated as false negatives\. Choosing the sentence as the unit of classification does not presuppose that an interpretive argument typically fits within a single sentence\. The annotated data show the opposite\. A single reading is frequently supported by several argument sentences: across all annotated positives, between roughly a third \(systematic, historical\) and nearly half \(grammatical, objective\-teleological\) of the readings are supported by more than one annotated argument sentence, and the same holds for the argument gate of[Section4](https://arxiv.org/html/2609.26945#S4)\. The sentence is thus a unit of*measurement*, chosen because it captures the components that make up an argument concisely and exhaustively\. Reconstructing full arguments would proceed by clustering classified sentences that support the same reading and contain the same canon; this additional step is outside the scope of this paper\. The structure of the dataset does not preclude coarser granularities\. Paragraph markers are inherited from the L\.L\.Con corpus \([Section4](https://arxiv.org/html/2609.26945#S4)\), so the dataset can also be used for paragraph\-level classification\. The results, however, would not be comparable to the ones reported here, as they would not share the same unit of measurement\. ### 3\.2Pipeline The subtasks compose into a cascading pipeline that mirrors the order of[Section2](https://arxiv.org/html/2609.26945#S2)\. Every sentence of the reasons is evaluated on the first reading criterion, and each subsequent criterion is asked only where the preceding requirements are satisfied; sentences that pass all three are treated as readings\. Relations between sentences can then be viewed as a matrix whose entries evaluate a \(reading sentence, candidate argument sentence\) pair\. Since evaluating all pairs grows quadratically with decision length, candidates are restricted to a window around each reading, with five sentences before and seven after, asymmetric because supporting sentences follow their reading far more often than they precede it \([Section3\.1](https://arxiv.org/html/2609.26945#S3.SS1)\); the window covers about 90% of annotated argument pairs, and positives annotated outside it are retained\. Each pair is classified on whether the candidate is an argument for the reading at all and, conditional on that, on each of the four canons\. In this paper, the pipeline defines the annotation scheme and the structure of the benchmark: each stage is evaluated in isolation, so that stage\-level scores are not confounded by upstream errors\. Running the pipeline end\-to\-end over full decisions, the setting in which upstream errors propagate, is supported by the decision\-level release \([Section4](https://arxiv.org/html/2609.26945#S4)\) but not pursued here\. ## 4Dataset ##### Source and scope\. The decisions are drawn from the L\.L\.Con corpus of decisions of the German Federal Constitutional Court\([Möllers and Wendel, 2023](https://arxiv.org/html/2609.26945#bib.bib14)\)\. We draw 28 decisions, all from the official collection \(*BVerfGE*\)\. Within each decision, annotation is restricted to the reasons \(*Entscheidungsgründe*\); the L\.L\.Con metadata already delimit them at paragraph level and we adopt these boundaries\. Reasons are split into sentences withdistilbert\-SBD\-de\-judgements, a DistilBERT model fine\-tuned for sentence boundary detection in German legal text\([Brugger et al\., 2023](https://arxiv.org/html/2609.26945#bib.bib15)\)\. The detector occasionally emits fragments that are not classifiable sentences, such as section headings and bare enumerators \(‘‘II\.’’, ‘‘B\.’’, ‘‘B\.\-I\.’’\) that carry no assertion; we mark these as non\-evaluable and exclude them from the dataset before annotation and classification, so they are never emitted as instances on any subtask\.777Operationally, a span is treated as non\-evaluable if it contains no alphabetic character or is at most five characters long\. This matters most fornicht\_abstrakt, since a bare enumerator trivially satisfies “not abstract” and leaving these spans in would inflate that subtask’s positive class\.The sentence is the unit of annotation and classification throughout\. Of the 28 decisions, 15 are annotated*exhaustively*, with every sentence of the reasons annotated, and the remaining 13 decisions are annotated*selectively*, with only the positive occurrences of the interpretive canons recorded\. In all 28 decisions, every annotated sentence is annotated along the cascade of[Section3\.2](https://arxiv.org/html/2609.26945#S3.SS2), receiving a label on each subtask it reaches; sentences that fail an earlier gate are not labeled further, since being able to reliably assign a sensible later label is conditional on the preceding requirements in the cascade being satisfied\. This means that negatives for any class are only drawn from sentences explicitly annotated as negative for the class in question, and never drawn from those which have not received a label at all\. It also means that the unannotated sentences in the selectively annotated decisions are not included in this dataset and are not, and should not be, treated as negatives\. For the abstract and self\-advanced criteria, positives and negatives come from the exhaustive set alone\. The statutory reference subtask also draws positives from selectively annotated decisions, while its negatives come only from the exhaustive set\. For the argument subtasks, selectively annotated decisions also contribute negatives\. ##### Subtasks\. The annotation scheme follows the steps laid out in[Section2](https://arxiv.org/html/2609.26945#S2)\. Throughout, we name each subtask by its English concept and give in parentheses the identifier under which it appears in the released data and in the result tables below; the identifiers are German, following the annotation scheme and the language of the source material\. Each candidate sentence is first scored on the three reading sub\-criteria that remain operationalizable on the basis of the decision text alone:not abstract\(nicht\_abstrakt\),not self\-advanced\(nicht\_selbst\_aufgestellt\), and thestatutory reference\(konkretes\_gesetz\), which asks for the provision whose content the sentence determines and so decides both halves of[Section2\.1\.3](https://arxiv.org/html/2609.26945#S2.SS1.SSS3)in a single pass\.888The first two criteria*abstract*and*self\-advanced*are framed as negations of how they were presented in[Section2\.1](https://arxiv.org/html/2609.26945#S2.SS1); for the rationale, see the section on class sizes below\.Sentences that pass all three are treated as readings; the sentences in the window described in[Section3\.2](https://arxiv.org/html/2609.26945#S3.SS2)are then evaluated as*candidate arguments*, first on whether the candidate supports the reading at all, i\.e\., theargument gate\(argument\), and then, conditional on that, on each of the four canons ofgrammatical\(wortlaut\),systematic\(systematik\),historical\(geschichte\) andobjective\-teleological\(zweck\) interpretation\. The gate is not a fifth canon: it asks*whether*a sentence supports the reading, where the canons ask*which kind*of support it offers, and the four canons do not exhaust the ways a court can argue\. A sentence can therefore pass the gate and still be negative on all four canons\. This gives eight subtasks in total: seven binary classification subtasks, plus the statutory reference task, whose label is a list of all the statutory provisions that the claim of the sentence refers to; usually, this will either be a single provision or no provision at all\. Argument subtasks are evaluated on a*sentence pair*, i\.e\. a candidate argument sentence together with the reading it is meant to support, since canon labels depend on what is being argued for \(cf\.[Examples1](https://arxiv.org/html/2609.26945#Thmexample1)and[2](https://arxiv.org/html/2609.26945#Thmexample2)in[Section2\.1](https://arxiv.org/html/2609.26945#S2.SS1)\)\. ##### Class sizes\. Each of the three reading sub\-criteria is represented by 400 sentences \(100 positive / 300 negative\)\. Each of the five argument subtasks is represented by 200 sentence pairs \(50 positive / 150 negative\)\. The argument subtasks are smaller because two filters compound: the pool of valid readings is itself a bottleneck, and many of the readings that do qualify are not supported by any explicit argument in the surrounding text\. For the rarer argument classes, the 1:3 positive–negative ratio is intentionally more favorable than the much lower true base rate to prevent class imbalances\. For the abstract and self\-advanced criteria, the natural ratio is closer to 3:1, which is why the task is framed inversely here, i\.e\. whether an assertion is*not*abstract or*not*self\-advanced\. The negatives of thestatutory referencesubtask come in two kinds, corresponding to the two halves of[Section2\.1\.3](https://arxiv.org/html/2609.26945#S2.SS1.SSS3): a sentence may name no specific provision at all, or name one without determining its content\. Because the first kind \(keine\_konkrete\_gesetzesbestimmung\) is concentrated in a few decisions, a decision\-disjoint draw clusters it, so each split preserves the pool’s proportions of the two kinds up to integer rounding\.999The exact split used is 78:22\. Of the 300 negatives, 234 are of the second kind \(kein\_bestimmter\_inhalt\) and 66 of the first, allocated per split as 47/13, 47/13, and 140/40 across train, validation, and test\. All splits are stratified to the class ratio of 1:3 described above\. Furthermore, they are decision\-disjoint throughout, and a decision quoted as a worked example in a subtask’s prompt is additionally held out of that subtask’s test split, so that no test sentence is one the prompt has already shown with its label\.101010The decisions held out of each subtask’s test split are:BVerfGE61,126andcs20090421\_2bvc000206forkonkretes\_gesetz;BVerfGE57,170forsystematik, forgeschichte, and fornicht\_selbst\_aufgestellt; andBVerfGE57,170,BVerfGE61,126, andBVerfGE62,338fornicht\_abstrakt\. All annotations were produced by a single annotator who pursues a PhD in law\. Each instance carries a free\-textreasoningfield that allows for more detailed explanations of the annotation in edge cases\. One subtask we deliberately do not provide as a stand\-alone classification target yet is whether a candidate sentence is an*assertion*, as opposed to a verbatim restatement of the statute: deciding this reliably requires comparing the sentence against the statutory text in force at the time of the decision, and the historical statutory text is not yet available at scale for Germany\. The related challenge of distinguishing a court’s own assertion from a reference to a statutory provision has recently been taken up for French law by[Floro et al\. \(2026\)](https://arxiv.org/html/2609.26945#bib.bib16)\. We instead flag sentences that look like potential restatements in the instance\-level data so they can be re\-checked once historical statutory data for Germany becomes available\. ##### Release\. The dataset is released at publication in two complementary forms\.Decision\-level: one record per fully annotated decision, pairing the plain text with JSON metadata that encodes every annotation as character offsets into that text along with paragraph\-level structural metadata inherited from L\.L\.Con\. This format is intended for end\-to\-end evaluation that runs the full pipeline over a whole decision, which is not pursued in this paper\.Instance\-level: one record per \(subtask, candidate\) pair, grouped by subtask and split\. Each record carries the candidate sentence, the surrounding context window with the candidate marked by inline XML tags \(<deutung\>/<potential\_argument\>;*Deutung*is the German term for a reading\), and the gold label with its reasoning\. For the statutory reference subtask, a wider context window is added because the relevant statute citation often lies further upstream than for the other criteria\. ## 5Baseline Evaluations We evaluate four LLMs from three model families under different prompts to establish a first baseline performance for the introduced task: DeepSeek\-V4 in its Pro and Flash variants, Gemini 3\.5 Flash Lite, and MiniMax M3\. DeepSeek\-V4 and MiniMax M3 are open\-weight models; Gemini 3\.5 Flash Lite is a proprietary model\. All four are reasoning models and were run at temperature 1\.0\. The models are compared across two sets of prompts: one that is not model\-specific, crafted by an expert, and one that is optimized using Genetic\-Pareto \(GEPA\)\([Agrawal et al\., 2026](https://arxiv.org/html/2609.26945#bib.bib17)\), starting from the expert prompt\.111111The expert prompts illustrate each criterion with worked examples\. Most are excerpts from the corpus, but some are constructed to isolate a single distinction as cleanly as a real passage rarely does; these carry placeholder citations such as “BVerfGE 123, 456” and are not references to decisions\. All prompts are released with the datasets \([Footnote1](https://arxiv.org/html/2609.26945#footnote1)\)\.The purpose of the GEPA condition is to validate the expert baseline: at a time when automatic prompt optimization frequently outperforms hand\-written prompts, it is useful to assess whether optimization substantially improves on a hand\-crafted baseline\. GEPA thus serves as an orientation point for what current models can achieve on the task, against which the expert prompt is anchored\. In the GEPA condition, the models are accordingly not evaluated across one static prompt, but across a model\-specific prompt produced by an optimizer with static settings: DeepSeek\-V4\-Pro was used as the reflection language model for all four evaluated models, at temperature 1\.0 with a maximum of 8,000 tokens, and the maximum metric calls budget was set to auto=light\. The built\-in instruction proposer was modified so that the resulting prompt would be German\. GEPA\([Agrawal et al\., 2026](https://arxiv.org/html/2609.26945#bib.bib17)\)evolves prompts by reflecting in natural language on execution traces and selecting along a Pareto frontier of candidates; because it learns from a handful of rollouts rather than from gradient updates over many labeled examples, it is markedly more sample\-efficient than tuning\-based adaptation and the optimization stage consumes correspondingly little data\. We optimize the binary subtasks for accuracy \(though we report F1,[AppendixA](https://arxiv.org/html/2609.26945#A1)\), and the statutory reference task for sample\-averaged F1\. Because GEPA scores each rollout individually, its target must be a per\-instance metric: a single binary decision admits only correct/incorrect \(accuracy in aggregate\), whereas the statutory reference output is a set of provisions with a per\-document F1 to average\. We exploit GEPA’s efficiency with an atypical 20/20/60 split: for each subtask, 20% of instances drive GEPA’s reflective search, a further 20% provide the validation signal it optimizes against, and the remaining 60% are held out as a test set touched only once, for the numbers reported here\. Allocating the majority of the labeled data to the test split yields the most reliable evaluation our annotation budget allows\. At classification time the model never sees a unit in isolation\. Each instance embeds the unit under classification in its surrounding sentence context, with the target span\(s\) marked by inline XML tags \(<deutung\>for a reading,<potential\_argument\>for a candidate argument\); the unit is a single sentence for the reading criteria and a \(reading, candidate\) pair for the argument subtasks\. The context spans two sentences on either side, widened to the ten preceding sentences for the statutory reference subtask, whose governing citation often lies further upstream \([Section4](https://arxiv.org/html/2609.26945#S4)\)\. Every subtask is binary except for the statutory reference subtask, which is a citation\-extraction target\. For each instance the model first emits a free\-textreasoningfield and then a structured decision: a binary value for the seven binary subtasks, or a list of statutory provisions for the statutory reference subtask\. The output is produced by constrained decoding wherever the backbone supports it, and falls back to JSON mode otherwise\. For the seven binary subtasks we reportF1F\_\{1\}on the positive class; per\-subtask precision and recall are given in[AppendixA](https://arxiv.org/html/2609.26945#A1)\. Thestatutory referencesubtask is reported separately because it is scored by set\-overlapF1F\_\{1\}between predicted and gold statutory citations rather than as a binary decision, so a single binaryF1F\_\{1\}would not be comparable\. Deciding whether a predicted citation matches a gold is challenging by string comparison, because the same provision admits many surface forms \(§ 2 Rechtshilfegesetz/§ 2 des Rechtshilfegesetzes/§ 2 RhG\)\. We therefore match with an LLM judge \(DeepSeek\-V4\-Pro\), which is shown the gold list for one sentence together with a single predicted citation and maps that prediction to at most one entry of the list, or to none\. The judge is instructed to tolerate formatting variation, for example roman numerals forAbs\., a trailing number forSatz, alternative statutory abbreviations, and differences in whitespace, punctuation and case, but to insist on the exact section, article, paragraph, number and lettered subdivision\. Predicted citations are deduplicated before judging; a prediction that the judge matches to no gold entry, or to one already claimed by an earlier prediction, counts as a false positive, and every gold entry left unclaimed counts as a false negative\. Per\-document set\-overlapF1F\_\{1\}averages the resulting per\-sentenceF1F\_\{1\}over the test sentences, crediting1\.01\.0where gold and prediction are both empty; span\-level precision, recall andF1F\_\{1\}pool the counts over all sentences instead\. Because the judge is itself a language model, its decisions are released alongside the predictions, so that the reported scores can be recomputed exactly without re\-running it\. The judge prompt is released with the other prompts\. All intervals are 95% bootstrap confidence intervals from 1,000 instance\-level resamples of the test set, reported as the 2\.5th and 97\.5th percentiles\. These intervals do not account for dependence between instances from the same decision\. Table 1:Per\-subtaskF1F\_\{1\}on the positive class for all four models under the expert and GEPA prompts\. Point estimates only; 95% bootstrap CIs and per\-cell precision and recall are in[AppendixA](https://arxiv.org/html/2609.26945#A1)\. “Mean” averages the seven binary subtasks\.†Statutory referenceis a citation\-extraction target scored by per\-document set\-overlapF1F\_\{1\}, not a binaryF1F\_\{1\}; its much lower span\-levelF1F\_\{1\}is discussed in[Section6](https://arxiv.org/html/2609.26945#S6)\. ## 6Results [Table1](https://arxiv.org/html/2609.26945#S5.T1)reportsF1F\_\{1\}scores for every run at the subtask level; more detailed scores can be found in the Appendix in[Tables2](https://arxiv.org/html/2609.26945#A1.T2)to[6](https://arxiv.org/html/2609.26945#A1.T6)\. ##### No model family dominates\. Averaged over the seven binary subtasks, the four models cluster between 70\.4 and 79\.2F1F\_\{1\}, a spread of under nine points\. MiniMax M3 under GEPA is nominally highest \(79\.2\) and DeepSeek\-V4\-Flash under GEPA nominally lowest \(70\.4\), but no model consistently leads across subtasks\. We therefore read the table as a set of baselines rather than a ranking\. ##### The expert baseline holds up against optimized prompts\. Under the tested configuration, the GEPA prompts do not systematically beat the expert prompt: on the binary mean, GEPA is ahead for Gemini 3\.5 Flash Lite \(75\.4 vs\. 72\.7\) and MiniMax M3 \(79\.2 vs\. 75\.2\), behind for DeepSeek\-V4\-Flash \(70\.4 vs\. 73\.5\), and level for DeepSeek\-V4\-Pro \(75\.1 vs\. 75\.2\)\. The expert prompt remains a meaningful baseline: its mean binaryF1F\_\{1\}does not sit far below what the tested GEPA configuration attains\. ##### The statutory\-reference scores need separate reading\. The headline set\-overlapF1F\_\{1\}on this subtask \(64\.1–78\.6\) is a per\-document metric dominated by the 180 of 240 test sentences whose gold citation set is empty: a model that correctly abstains on those scores1\.01\.0\. At the level of individual citations the extraction is far weaker, with span\-levelF1F\_\{1\}only 28\.4–42\.4 and precision between 21 and 47 \([Table6](https://arxiv.org/html/2609.26945#A1.T6)\)\. Both prompt conditions over\-extract, emitting a citation for sentences that merely mention or evaluate a provision without asserting its content, so the low precision is a property of the task rather than of one prompt\. ##### Canon difficulty is stable at its extremes and plausibly holds up against legal intuition\. The classifiers achieve the best scores on grammatical interpretation in seven of the eight model–prompt conditions \(mean 82\.9\), and the lowest ones on systematic interpretation in six of eight \(mean 58\.4\); historical and objective\-teleological interpretation sit between and are effectively tied \(means 70\.2 and 70\.9\)\. This is at least a plausible result as the grammatical canon has a clear anchor in the boundary of the word, whereas the reach of the systematic canon is contested and the boundaries of purpose are blurry\. Systematic interpretation’s low score comes from failing in both directions at once, over\- and under\-attributing the canon \([Section7](https://arxiv.org/html/2609.26945#S7)\)\. ## 7Error Analysis ##### Multi\-sentence arguments drive the false negatives\. We annotate every sentence of an argument that runs across several sentences; models miss continuing or concluding sentences of such arguments that lack canon\-specific cues\. This is why objective\-teleological interpretation, whose arguments are often built up gradually, loses ground on recall\. ##### Shallow cues drive the false positives\. On the argument gate the models accept restatements and mere consequences of the reading as though they were grounds; on the remaining subtasks they latch onto a surface proxy, for example the token “Wortlaut” \(grammatical\), the citation of any distinct provision \(systematic\), or any section number \(statutory reference, hence its over\-extraction\), none of which is sufficient for the criterion it stands in for\. Systematic interpretation is the canon with the lowest scores likely because its proxy misfires both ways, firing on any parallel citation and missing systematic arguments that name no second provision\. ## 8Discussion and Limitations Several limitations bound our results\. All gold labels come from a single annotator, which yields consistent application of the criteria in[Section2](https://arxiv.org/html/2609.26945#S2)but does not allow for a measure of inter\-annotator agreement\. The corpus is drawn entirely from one court in one language, so we make no claim of transfer to other jurisdictions or legal traditions\. Furthermore, while the baseline covers four models from three families, it remains unclear how well the most capable closed\-source models would close the remaining gap\. The confidence intervals in[Section6](https://arxiv.org/html/2609.26945#S6)are wide, a direct effect of the small test sets the current dataset affords, especially for the rarer argument subtasks; differences between prompt conditions should not be over\-interpreted\. Extending the corpus using multiple annotators, which we intend to do, is the most direct remedy\. ## 9Conclusion This paper set out to make the classification of interpretive canons tractable at the sentence level, rather than the paragraph that prior work has relied on\. We made three contributions: we operationalized the four classical canons as articulated by Larenz in the tradition of Savigny; we released an expert\-annotated dataset of German Federal Constitutional Court decisions, exhaustively labeled for fifteen decisions and selectively extended for rarer classes; and we reported baseline evaluations of four LLMs from three model families under expert\-written and GEPA\-optimized prompts\. Two findings stand out\. First, the relative difficulty of the canons is at least plausible in light of legal methodology, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest\. Second, prompts optimized with the tested GEPA configuration do not systematically outperform the expert hand\-written prompts in mean binaryF1F\_\{1\}, supporting the reported baseline as a meaningful reference point\. Given the wide confidence intervals and the modest size of the current dataset, we read these results as a starting point that invites extending the corpus and building toward better evaluations of interpretive canons\. ## Impact Statement This paper presents work whose goal is to advance the field of Machine Learning\. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here\. ## Acknowledgments I thank Professor Andreas Engert for his support throughout this work, and the anonymous reviewers of the ICML 2026 Workshop on AI for Law for their constructive feedback\. Any remaining errors are my own\. ## References - L\. A\. Agrawal, S\. Tan, D\. Soylu, N\. Ziems, R\. Khare, K\. Opsahl\-Ong, A\. Singhvi, H\. Shandilya, M\. J\. Ryan, M\. Jiang, C\. Potts, K\. Sen, A\. G\. Dimakis, I\. Stoica, D\. Klein, M\. Zaharia, and O\. KhattabGEPA: reflective prompt evolution can outperform reinforcement learning\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§5](https://arxiv.org/html/2609.26945#S5.p2.1),[§5](https://arxiv.org/html/2609.26945#S5.p3.1)\. - Bruggeret al\.\(2023\)T\. Brugger, M\. Stürmer, and J\. NiklausMultiLegalSBD: a multilingual legal sentence boundary detection dataset\.InProceedings of the Nineteenth International Conference on Artificial Intelligence and Law,ICAIL ’23,pp\. 42–51\.External Links:[Document](https://dx.doi.org/10.1145/3594536.3595132)Cited by:[§4](https://arxiv.org/html/2609.26945#S4.SS0.SSS0.Px1.p1.1)\. - Floroet al\.\(2026\)A\. Floro, T\. Dhorasoo, S\. Pellez, and N\. HolzenbergerWhere experts disagree, models fail: detecting implicit legal citations in french court decisions\.External Links:2603\.22973,[Document](https://dx.doi.org/10.48550/arXiv.2603.22973),[Link](https://arxiv.org/abs/2603.22973)Cited by:[§4](https://arxiv.org/html/2609.26945#S4.SS0.SSS0.Px3.p3.1)\. - Gorsuch \(2016\)N\. M\. GorsuchOf lions and bears, judges and legislators, and the legacy of justice Scalia\.Case Western Reserve Law Review66\(4\),pp\. 905–917\.External Links:ISSN 0008\-7262Cited by:[§1](https://arxiv.org/html/2609.26945#S1.p1.1)\. - Harbor Framework Team \(2026\)Harbor: A framework for evaluating and optimizing agents and models in container environmentsExternal Links:[Link](https://github.com/harbor-framework/harbor)Cited by:[Appendix D](https://arxiv.org/html/2609.26945#A4.p1.1)\. - Korefet al\.\(2026\)T\. Koref, L\. Held, M\. Namazov, H\. Kumru, Y\. Thlija, and I\. HabernalMining legal arguments to study judicial formalism\.External Links:2512\.11374,[Link](https://arxiv.org/abs/2512.11374)Cited by:[§1](https://arxiv.org/html/2609.26945#S1.p2.1)\. - Kriele \(1967\)M\. KrieleTheorie der rechtsgewinnung: entwickelt am problem der verfassungsinterpretation\.Schriften zum Öffentlichen Recht,Duncker & Humblot,Berlin\(german\)\.Cited by:[§2\.2](https://arxiv.org/html/2609.26945#S2.SS2.p1.1)\. - Larenz \(1991\)K\. LarenzMethodenlehre der rechtswissenschaft\.6 edition,Springer,Berlin, Heidelberg\(ger\)\.Cited by:[§1](https://arxiv.org/html/2609.26945#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.26945#S2.SS2.p2.1),[§2\.2](https://arxiv.org/html/2609.26945#S2.SS2.p3.1),[§2\.2](https://arxiv.org/html/2609.26945#S2.SS2.p4.1),[§2\.2](https://arxiv.org/html/2609.26945#S2.SS2.p5.1),[§2](https://arxiv.org/html/2609.26945#S2.p1.2.1)\. - Larenz \(2014\)K\. LarenzMetodologia da ciência do direito\.7 edition,Fundação Calouste Gulbenkian,Lisboa\(portuguese\)\.Note:Tradução de José Lamego; título original: Methodenlehre der RechtswissenschaftExternal Links:ISBN 978\-972\-31\-0770\-8Cited by:[§1](https://arxiv.org/html/2609.26945#S1.p3.1)\. - Larenz \(2020\)K\. LarenzFaxue fangfalun \[methodenlehre der rechtswissenschaft\]\.The Commercial Press,Beijing\(chinese\)\.Note:Translated by Huang Jiazhen; from the 6th German editionCited by:[§1](https://arxiv.org/html/2609.26945#S1.p3.1)\. - Larenz \(2023\)K\. LarenzMetodología de la ciencia del derecho\.Biblioteca de filosofía del derecho,Ediciones Olejnik,Santiago de Chile\(spanish\)\.Note:Traducción de Marcelino Rodríguez Molinero, Carlos Antonio Agurto Gonzáles, Sonia Lidia Quequejana Mamani y Benigno Choque Cuenca; título original: Methodenlehre der RechtswissenschaftExternal Links:ISBN 956\-407\-380\-4Cited by:[§1](https://arxiv.org/html/2609.26945#S1.p3.1)\. - Larenz \(2026\)K\. LarenzBeophak bangbeomnon \[methodenlehre der rechtswissenschaft\]\.Parkyoungsa,Seoul\(korean\)\.Note:Translated by Lee Dong\-jinExternal Links:ISBN 979\-11\-303\-4950\-3Cited by:[§1](https://arxiv.org/html/2609.26945#S1.p3.1)\. - Lüders and Stohlmann \(2025\)K\. Lüders and B\. StohlmannClassifying proportionality – identification of a legal argument\.Artificial Intelligence and Law33\(4\),pp\. 1051–1078\.External Links:[Document](https://dx.doi.org/10.1007/s10506-024-09415-9)Cited by:[§2](https://arxiv.org/html/2609.26945#S2.p1.1)\. - Mendelson \(2018\)N\. A\. MendelsonChange, creation, and unpredictability in statutory interpretation: interpretive canon use in the Roberts court’s first decade\.Michigan Law Review117\(1\),pp\. 71–142\.External Links:ISSN 0026\-2234Cited by:[§1](https://arxiv.org/html/2609.26945#S1.p2.1)\. - Möllers and Wendel \(2023\)C\. Möllers and L\. WendelKorpus der entscheidungen des bundesverfassungsgerichts\.Note:ZenodoData setExternal Links:[Document](https://dx.doi.org/10.5281/zenodo.10369205),[Link](https://doi.org/10.5281/zenodo.10369205)Cited by:[§4](https://arxiv.org/html/2609.26945#S4.SS0.SSS0.Px1.p1.1)\. - Raisch \(1988\)P\. RaischVom nutzen der überkommenen auslegungskanones für die praktische rechtsanwendung\.Juristische Studiengesellschaft Karlsruhe: Schriftenreihe Heft 181,C\. F\. Müller Juristischer Verlag,Heidelberg\.External Links:ISBN 3\-8114\-6988\-6Cited by:[§1](https://arxiv.org/html/2609.26945#S1.p1.1)\. - Stiglitz and Thalken \(2024\)E\. H\. Stiglitz and R\. ThalkenHistorical trends in macro\-jurisprudence: a language model assessment, 1870–2023\.Maryland Law Review84\(1\)\.Cited by:[§1](https://arxiv.org/html/2609.26945#S1.p2.1)\. - Thalken and Stiglitz \(2026\)R\. Thalken and E\. StiglitzMeasuring jurisprudence\.Journal of Law and Courts,pp\. 1–22\.External Links:[Document](https://dx.doi.org/10.1017/jlc.2025.10012)Cited by:[§1](https://arxiv.org/html/2609.26945#S1.p2.1)\. ## Appendix AComplete Model Scores Tables[2](https://arxiv.org/html/2609.26945#A1.T2)–[5](https://arxiv.org/html/2609.26945#A1.T5)report per\-subtask precision, recall andF1F\_\{1\}for all four models on the seven binary subtasks, and[Table6](https://arxiv.org/html/2609.26945#A1.T6)reports thestatutory referencesubtask\. Subtasks are named by the identifiers introduced in[Section4](https://arxiv.org/html/2609.26945#S4)\. Precision and recall are point estimates;F1F\_\{1\}carries a 95% bootstrap CI \(1,000 resamples\)\. “Mean” averages the seven binary subtasks\. Table 2:Precision, recall andF1F\_\{1\}\(positive class\) for DeepSeek\-V4\-Pro across the seven binary subtasks\. P and R are point estimates;F1F\_\{1\}carries the 95% bootstrap CI \(1,000 resamples\)\.Statutory referenceis reported separately in[Table6](https://arxiv.org/html/2609.26945#A1.T6)\.Table 3:Precision, recall andF1F\_\{1\}\(positive class\) for DeepSeek\-V4\-Flash across the seven binary subtasks\. P and R are point estimates;F1F\_\{1\}carries the 95% bootstrap CI \(1,000 resamples\)\.Statutory referenceis reported separately in[Table6](https://arxiv.org/html/2609.26945#A1.T6)\.Table 4:Precision, recall andF1F\_\{1\}\(positive class\) for Gemini 3\.5 Flash Lite across the seven binary subtasks\. P and R are point estimates;F1F\_\{1\}carries the 95% bootstrap CI \(1,000 resamples\)\.Statutory referenceis reported separately in[Table6](https://arxiv.org/html/2609.26945#A1.T6)\.Table 5:Precision, recall andF1F\_\{1\}\(positive class\) for MiniMax M3 across the seven binary subtasks\. P and R are point estimates;F1F\_\{1\}carries the 95% bootstrap CI \(1,000 resamples\)\.Statutory referenceis reported separately in[Table6](https://arxiv.org/html/2609.26945#A1.T6)\.Table 6:Statutory reference\(konkretes\_gesetz\), the citation\-extraction subtask\.*Set\-overlap*F1F\_\{1\}is the per\-document metric reported in the main text and averages over the 180/240 empty\-gold rows \(a model that correctly returns no citation scores 1\.0 on such a row\)\.*Span\-level*P/R/F1F\_\{1\}count individual predicted vs\. gold citations and expose the over\-extraction directly\.F1F\_\{1\}columns carry 95% bootstrap CIs\. ## Appendix BExpert System Prompts The eight hand\-written expert system prompts are reproduced verbatim below\. They are model\-agnostic, identical across the four models\. Each operationalizes one subtask: it states the criterion and illustrates it with worked positive and negative examples, some constructed with placeholder citations such as “BVerfGE 123, 456” rather than real decisions\. The paired user prompts, which wrap the instance’s input fields \(the context fieldtext\_with\_context, replaced by its wider varianttext\_with\_context\_konkretes\_gesetzfor the statutory reference subtask, plus, for the argument subtasks, the reading sentence, the candidate sentence and the reading’s statutory reference\), are omitted, as are the model\-specific GEPA prompts; all prompts are released with the datasets \([Footnote1](https://arxiv.org/html/2609.26945#footnote1)\)\. The prompts are in German, matching the corpus\. ### B\.1Not abstract \(nicht\_abstrakt\) DasübergreifendeProjekt,indemdugenutztwirst,isteineStudie,inderuntersuchtwird,wiesichdieVerwendungvonAuslegungsmethodenindenEntscheidungendesBundesverfassungsgerichtsimLaufederGeschichteentwickelthat\.DeineAufgabeistTeileinesWorkflows,dereszumZielhat,einenSatzMetadatenzuerstellen,derfürjedeEntscheidunggenauangibt,anwelcherStelledasGerichtAuslegungmittelsdervierAuslegungskriterienbetriebenhat\. Auslegungbedeutet,sichfüreineuntermehrerenmöglichenDeutungeneinerGesetzesbestimmungaufgrundvonÜberlegungenzuentscheiden\. AusdieserDefinitionlassensichdieverschiedenenSchritteableiten,ausdenensichderWorkflowzusammensetzt: 1\.Deutungenidentifizieren 2\.Identifizieren,obessichbeibestimmtenSätzenüberhauptumsolchehandelt,dieeineBegründungfürdieDeutungdarstellen\. 3\.Identifizieren,welcheAuslegungskriterienindemArgumentverwendetwerden\. DuwirstdichausschließlichmitdemerstenSchrittdiesesWorkflowsbefassen\. EineDeutungisteineBehauptung,diedasGerichtselbstaufstelltunddiezumGegenstandhat,dasseinebestimmteGesetzesbestimmungabstrakteinenbestimmtenInhalthabe\. EineBehauptungüberdenInhalteinerGesetzesbestimmungodereinesallgemeinenRechtsprinzipsistabstrakt,wennsieGeltungbeanspruchtnichtnurinBezugaufdeneinzigartigenLebenssachverhalt,derderEntscheidungzugrundeliegt,sondernfürunendlichvieleSituationen,indenendieGesetzesbestimmungoderdasRechtsprinzipAnwendungfindet\. DeineAufgabeisteszuentscheiden,obeineBehauptung\*\*nicht\*\*abstraktist\.DerinhaltlicheMaßstabistunverändertdersoebendefinierte;nurdieRichtungderAntwortistumgekehrt\.DasAusgabefeld‘nicht\_abstrakt‘istalsogenaudann‘true‘,wenndieBehauptungnachdiesemMaßstab\*\*nicht\*\*abstraktist,und‘false‘,wennsieabstraktist\. PositiveBeispiele\(‘nicht\_abstrakt:true‘,dieBehauptungistalsonichtabstrakt\): <example\>ImvorliegendenFallistnichtauszuschließen,daßdasOberlandesgerichtbeiderBeurteilungderFrage,obeinekonkreteGefährdungderAnstaltsordnungdurchdenBriefzubesorgenwar,beiAnwendungdesdargelegtenverfassungsrechtlichenMaßstabszueineranderenrechtlichenWürdigunggelangtwäre\.</example\> HierwirdaufdenvorliegendenFallunddiekonkreteEntscheidungdesOberlandesgerichtsübereinenbestimmtenBriefBezuggenommen\.EswirdkeineallgemeingültigeAussagegetroffen,daherliegtkeineabstrakteBehauptungvor\. <example\>"DieEntscheidungdesOberlandesgerichtsistgemäß§304Abs\.4StPOendgültig\."</example\> HierwirdkeineabstrakteAussageüber§304Abs\.4StPOgetroffen,sondernnurkonkret\*die\*EntscheidungdesOberlandesgerichtsanhandderBestimmungbewertet\.DaherliegtkeineabstrakteBehauptungvor\. <example\>DieVoraussetzungdes§147StPOlagenzweifelsfreivor;esgabkeinenGrund,denZentralregisterauszughiervonauszunehmen\.</example\> Hierwird§147StPOnichtabstraktausgelegt,sondernnuraufdenkonkretenSachverhaltangewandt:Eswirdfestgestellt,dassseineVoraussetzungenimentschiedenenFallvorlagen\.EswirdkeineallgemeingültigeAussagegetroffen,daherliegtkeineabstrakteBehauptungvor\. NegativeBeispiele\(‘nicht\_abstrakt:false‘,dieBehauptungistalsoabstrakt\): <negative\_example\>"EsdientverfassungsrechtlichlegitimenZwecken,dieexterneTeilungderin§17VersAusglGgenanntenAnrechte\(BetriebsrentenauseinerDirektzusageoderUnterstützungskasse\)auchüberdieWertgrenzedes§14Abs\.2Nr\.2VersAusglGhinauszuerlauben\."</negative\_example\> HierwirdeineAussageüberdiein§17derVorschriftgenanntenAnrechtegetroffen\.DieseAussagewirdhiernichtnurfüreinenkonkretenLebenssachverhaltgetroffen,sondernallgemeinfüralleFälle,indenendieVorschriftzurAnwendungkommenkönnte\.SomitistdieBehauptungabstraktimSinnederDefinition\. <negative\_example\>WirdderBriefeinesUntersuchungsgefangenen,wiehier,insinngemäßerAnwendungdes§108StPOsichergestelltundderStaatsanwaltschaftzurweiterenVeranlassungzugeleitet,sokannsichderBetroffenemitdemZielderHerausgabedesBriefesandieStaatsanwaltschaftwenden,jederzeitaufEntscheidungdeszuständigenRichtersantragen\(§98Abs\.2StPO\)undgegeneineetwaigeBeschlagnahmeimVorverfahrenBeschwerdeeinlegen\(§304Abs\.1StPO\)\.</negative\_example\> Hierwirdallgemeingültigausgeführt,welcheRechtsbehelfeeinemUntersuchungsgefangenenbeiderSicherstellungseinesBriefesinsinngemäßerAnwendungdes§108StPOoffenstehen\.DieAussagegiltnichtnurfürdenkonkretenFall\("wiehier"\),sondernfürallederartigenFälle\.EshandeltsichdaherumeineabstrakteBehauptung\. <negative\_example\>DochmußeinsolcherEingriffdemGrundsatzderVerhältnismäßigkeitentsprechen,dersichbereitsausdemWesenderGrundrechteselbstergibtunddemalsElementdesRechtsstaatsprinzipsVerfassungsrangzukommt\(vgl\.BVerfGE19,342\[347ff\.\];29,312\[316\]\)\.</negative\_example\> DieserSatzlegtdenGrundsatzderVerhältnismäßigkeitausundbeanspruchtGeltungfüralleGrundrechtseingriffe,nichtnurfürdendeskonkretenFalls\.DaheristdieBehauptungabstrakt\. DuerhältsteineEingabe: \-‘text\_with\_context‘:einAuszugauseinerEntscheidungdesBundesverfassungsgerichts\.DerkonkretzuprüfendeSatzistinnerhalbdesAuszugsmitdenMarkierungen‘<deutung\>\.\.\.</deutung\>‘umschlossen;dieumgebendenSätzedienenausschließlichalsKontext\. BegründeimmerdeineAntwort,bevordudichentscheidest\.Prüfe,obderInhaltdesmit‘<deutung\>‘markiertenSatzesabstraktimSinnedesMaßstabsausDefinitionundBeispielenist\.EineabstrakteBehauptungkannsichaufeinebestimmteGesetzesbestimmungodereinallgemeinesRechtsprinzipbeziehen\.Setzeanschließend‘nicht\_abstrakt‘auf‘false‘,wenndieBehauptungabstraktist,undauf‘true‘,wennsienichtabstraktist\. ### B\.2Not self\-advanced \(nicht\_selbst\_aufgestellt\) DasübergreifendeProjekt,indemdugenutztwirst,isteineStudie,inderuntersuchtwird,wiesichdieVerwendungvonAuslegungsmethodenindenEntscheidungendesBundesverfassungsgerichtsimLaufederGeschichteentwickelthat\.DeineAufgabeistTeileinesWorkflows,dereszumZielhat,einenSatzMetadatenzuerstellen,derfürjedeEntscheidunggenauangibt,anwelcherStelledasGerichtAuslegungmittelsdervierAuslegungskriterienbetriebenhat\. Auslegungbedeutet,sichfüreineuntermehrerenmöglichenDeutungeneinerGesetzesbestimmungaufgrundvonÜberlegungenzuentscheiden\. AusdieserDefinitionlassensichdieverschiedenenSchritteableiten,ausdenensichderWorkflowzusammensetzt: 1\.Deutungenidentifizieren 2\.Identifizieren,obessichbeibestimmtenSätzenüberhauptumsolchehandelt,dieeineBegründungfürdieDeutungdarstellen\. 3\.Identifizieren,welcheAuslegungskriterienindemArgumentverwendetwerden\. DuwirstdichausschließlichmitdemerstenSchrittdiesesWorkflowsbefassen\. EineDeutungisteineBehauptung,diedasGerichtselbstaufstelltunddiezumGegenstandhat,dasseinebestimmteGesetzesbestimmungabstrakteinenbestimmtenInhalthabe\. Grundsätzlichisthiervonauszugehen\.EineBehauptungistjedochdanngrundsätzlichnichtselbstaufgestellt,wennsieaufeinPräjudizverweist\(bspw\.durchZitiereneinerFundstellewieBVerfGE123,456oderdenVerweisaufständigeRechtsprechung,st\.Rspr\.\)\.WirdeinPräjudizUNDandereQuellenwiezBLiteraturquellen,GesetzesbegründungenoÄzitiert,istdieBehauptungebenfallsselbstaufgestellt;esgehtausschließlichdarum,Sätzeauszunehmen,dienurPräjudizienzitieren\.DerMaßstabdeinerBeurteilungistausschließlichdieQuellenlage\.UnterlassealsojeglicheÜberlegungenzuOriginalitätoderEigenleistung\. DeineAufgabeisteszuentscheiden,obeineBehauptung\*\*nicht\*\*selbstaufgestelltist\.DerinhaltlicheMaßstabistunverändertdersoebendefinierte;nurdieRichtungderAntwortistumgekehrt\.DasAusgabefeld‘nicht\_selbst\_aufgestellt‘istalsogenaudann‘true‘,wenndieBehauptung\*\*nicht\*\*selbstaufgestelltist\(wennalsoausschließlichPräjudizienzitiertwerden\),und‘false‘,wennsieselbstaufgestelltist\. PositiveBeispiele\(‘nicht\_selbst\_aufgestellt:true‘,dieBehauptungistalsonichtselbstaufgestellt\): <example\>"DeshalbmußeinBeschwerdeführerdieBeseitigungdesHoheitsaktes,dessenGrundrechtswidrigkeitergeltendmacht,zunächstmitdenihmdurchdasGesetzzurVerfügunggestelltenanderenRechtsmittelnoderRechtsbehelfenzuerreichenversuchen\(BVerfGE33,192\[194\];st\.Rspr\.\)"</example\> HierwirdeinUrteilzitiert,dasanscheinendmaßgeblichfürdieBehauptungist,undeswirddurchdenZusatz\\"st\.Rspr\.\\"\(ständigeRechtsprechung\)kenntlichgemacht,dassdieseineinderRechtsprechungdesBundesverfassungsgerichtsseitLangemanerkanntePositionist\. NegativeBeispiele\(‘nicht\_selbst\_aufgestellt:false‘,dieBehauptungistalsoselbstaufgestellt\): <negative\_example\>"DiedemBeschlußdesOberlandesgerichtsDüsseldorfzugrundeliegendeAuffassung,dieausdemGrundrechtdesArt\.2Abs\.1abgeleitetenGrundsätzeseiennichtvonBedeutung,verkenntdieTragweitedieserVerfassungsgarantien\."</negative\_example\> HierknüpftdasGerichtzwaraneinefremdeBehauptungan,stelltaberseineeigeneauf,indemesbehauptet,dassdieAuffassungdesanderenGerichtsgeradenichtzuträfe\. <negative\_example\>\*\*KleineBäckereiensindvomNachtbackverbotausgenommen\.\*\*DennesergibtsichausSinnundZweckderNorm,dasssiebesondersschutzbedürftigsind\.FabrikenmüssenihreMitarbeitermitderrichtigenAusrüstungversorgen\.\(BVerfGE123,456\)\.</negative\_example\> HierwirdzwareinPräjudizzitiert,aberineinemanderenSatz\.DaherhandeltessichbeidemSatzinSternchenumeineselbstaufgestellteBehauptung\. DuerhältsteineEingabe: \-‘text\_with\_context‘:einAuszugauseinerEntscheidungdesBundesverfassungsgerichts\.DereigentlichzubeurteilendeSatzistinnerhalbdesAuszugsmitdenMarkierungen‘<deutung\>\.\.\.</deutung\>‘umschlossen;dieumgebendenSätzebildenausschließlichdenKontext\. BegründeimmerdeineAntwort,bevordudichentscheidest\.BeurteilekurzanhandderobenaufgestelltenMaßstäbe,obdasGerichtdieBehauptungselbstaufgestellthat\.Gehedabeiwiefolgtvor:NennezunächstdenSatz,derzwischen‘<deutung\>‘und‘</deutung\>‘steht\.ListedannausschließlichdieindiesemSatzgenanntenQuellenauf\(Literaturquellen,Gesetzesbegründungen,Präjudizien,GutachtenuÄ\)\.Quellen,dieanandererStelleimTextstehen,dürfenaufkeinenFallgenanntwerden\.WennkeineQuellenzitiertwerden,giltdieBehauptungalsselbstaufgestellt\.WennausschließlichPräjudizien\(bspw\.BVerfGE12,345oder’UrteildesBVerfGvom12\.03\.1993,Az\.\.\.\.’\)zitiertwerden,istdieBehauptungnichtselbstaufgestellt\.InallenanderenFällenistdieBehauptungselbstaufgestellt\.Merke:\(1\)EntscheideausschließlichnachArtderQuellen,nichtnachInhaltderBehauptungund\(2\)NenneausschließlichQuellenausdemmit‘<deutung\>‘markiertenSatz\.Setzeschließlich‘nicht\_selbst\_aufgestellt‘auf‘true‘,wenndieBehauptungnachdiesemVorgehennichtselbstaufgestelltist,undauf‘false‘,wennsieselbstaufgestelltist\. ### B\.3Statutory reference \(konkretes\_gesetz\) DasübergreifendeProjekt,indemdugenutztwirst,isteineStudie,inderuntersuchtwird,wiesichdieVerwendungvonAuslegungsmethodenindenEntscheidungendesBundesverfassungsgerichtsimLaufederGeschichteentwickelthat\.DeineAufgabeistTeileinesWorkflows,dereszumZielhat,einenSatzMetadatenzuerstellen,derfürjedeEntscheidunggenauangibt,anwelcherStelledasGerichtAuslegungmittelsdervierAuslegungskriterienbetriebenhat\. Auslegungbedeutet,sichfüreineuntermehrerenmöglichenDeutungeneinerGesetzesbestimmungaufgrundvonÜberlegungenzuentscheiden\. AusdieserDefinitionlassensichdieverschiedenenSchritteableiten,ausdenensichderWorkflowzusammensetzt: 1\.Deutungenidentifizieren 2\.Identifizieren,obessichbeibestimmtenSätzenüberhauptumsolchehandelt,dieeineBegründungfürdieDeutungdarstellen\. 3\.Identifizieren,welcheAuslegungskriterienindemArgumentverwendetwerden\. DuwirstdichausschließlichmitdemerstenSchrittdiesesWorkflowsbefassen\. EineDeutungisteineBehauptung,diedasGerichtselbstaufstelltunddiezumGegenstandhat,dasseinebestimmteGesetzesbestimmungabstrakteinenbestimmtenInhalthabe\. EineGesetzesbestimmungisteinespezifische,abgrenzbareNorm,diedurcheinegenaueBezeichnungidentifiziertwerdenkann\(bspw\.§123BGB\)\.EsdarfsichnichtumGewohnheitsrechtoderbloßeRechtsprinzipien/Grundsätzehandeln\.WirdeineBestimmunggenannt,ausdersicheinGrundsatzherleitet,zähltdiesenurdannalstauglicheGesetzesbestimmung,wennderSatzeinespezifischeBehauptungüberderenInhaltenthält\.EinetauglicheGesetzesbestimmungmussnichtimSatzselbstgenanntsein,sondernkannauchimKontextstehenundimplizitGegenstandderAuslegungsein\.EineBehauptungübereinGesetzodereineVerordnungalsGanzesschließteineDeutungaus\. DieBehauptungmussdarüberhinauseinenbestimmtenInhaltderGesetzesbestimmungfestlegen\.DasistdannderFall,wenndieAussagedazuführt,dasseineReihevonFällenvonderGesetzesbestimmungerfasstistundanderenicht\.NichtbestimmtisteineAussagehingegen,wennsiedaraufabzielt,dieGesetzesbestimmungimGanzenalsnichtig,wirksam,anwendbaroderunanwendbarzuerklären\.DamitscheideninsbesonderesolcheSätzeaus,dienurdieVereinbarkeiteinerNormmiteineranderenzumGegenstandhaben:diePrüfungderVerfassungsmäßigkeiteinerVorschriftebensowiedieeinzelnenSchritteeinerVerhältnismäßigkeitsprüfung\(Geeignetheit,Erforderlichkeit,Angemessenheit\)undErwägungenzurZweckmäßigkeitinnerhalbeinersolchenPrüfung\.DerartigeAusführungenbeziehensichaufdasVerhältniszweierNormenzueinanderundnichtaufdenInhaltdereinenoderderanderen\.DassdiebetroffenenVorschriftendabeigenaubezeichnetwerden,genügtnicht:sinddieVoraussetzungendiesesAbsatzesnichterfüllt,gibeineleereListezurück,auchwennimSatzeineodermehrereNormenausdrücklichgenanntsind\. PositiveBeispiele: <example\>"DasBundeswahlgesetzordnetin§37an,dassdasWahlergebnis\\"nachBeendigungderWahlhandlung\\"festzustellenist\.\[\.\.\.\]FüreineAuslegungdahingehend,dass\*\*die\\"Wahlhandlung\\"amAbendderHauptwahlbeendetistundbeiderNachwahleineneue\\"Wahlhandlung\\"beginnt\*\*,sprichtaber§46Abs\.1Satz1Nr\.2BWG,derdenVerlustdesAbgeordnetenmandatsfürdenFallder\\"Neufeststellung\\"desWahlergebnissesvorsieht\."</example\> Hierwirdauf§37BWGBezuggenommen,obwohldieGesetzesbestimmunginderDeutungselbstnichtgenanntist\. <example\>Art\.38Abs\.1Satz1GGbestimmt,dassdieAbgeordnetendesDeutschenBundestagesinallgemeiner,unmittelbarer,freier,gleicherundgeheimerWahlgewähltwerden\.Die’GleichheitderWahl’erfordertdabei,dassjedeStimmedengleichenZählwertunddiegleicheErfolgschancehabenmuss\.</example\> HierwirdArt\.38Abs\.1Satz1GGkonkretausgelegt,indemerläutertwird,wasunterder’GleichheitderWahl’zuverstehenist\. <example\>GegendieVerfassungsmäßigkeitdes§1361Abs\.2BGB,derdurchdasGesetzüberdieGleichberechtigungvonMannundFrauaufdemGebietdesbürgerlichenRechtsvom18\.Juni1957\(BGBl\.IS\.609\)neugefaßtwordenist,bestehenwederimHinblickaufdendurchArt\.6Abs\.1GGgefordertenSchutzvonEheundFamilienochimHinblickaufdieGleichberechtigungvonMannundFraunachArt\.3Abs\.2GGBedenken\.DieseUnterhaltsregelungträgtbeidenVerfassungsgebotenRechnung,indemsieindendortgeregeltenFälleneinenbesonderenSchutzdernichterwerbstätigenEhefrauvorsieht\(vgl\.denSchriftlichenBerichtdesAusschussesfürRechtswesenundVerfassungsrecht\-16\.Ausschuß\-zuBT\-Drucks\.II/3409S\.39\)\.<potentielle\_deutung\>Siegehtdavonaus,daßdieFrauihreVerpflichtungzumUnterhaltderFamilieinderRegeldurchdieFührungdesHaushaltserfüllt\(§1360Satz2BGB\)undhäufigimVertrauenaufdieDauerhaftigkeitderEhevoneineraußerhäuslichenErwerbstätigkeitAbstandnimmt,umsichausschließlichdemhäuslichenBereichderFamiliezuwidmen\.</potentielle\_deutung\>HierausergebensichimFalleeinerAufhebungderehelichenGemeinschaftzwangsläufigerheblicheNachteilefürdieerwerbswirtschaftlicheSituationderFrau,undzwarauchdann,wennsiekeineKinderzubetreuenhat\(vgl\.BVerfGE17,1\[21f\.\]\)\.</example\> Hierwird§1361Abs\.2BGBkonkretausgelegt,indemerläutertwird,wasunterderGleichberechtigungvonMannundFrauzuverstehenist\.Nichthingegenwird§1360Satz2BGBausgelegt,obwohlerimzuanalysierendenSatzzitiertwird\.DerSatztriffteineAussageüberdenInhaltvon§1361Abs\.2BGBundziehtzurAusgestaltungeineDeutungaus§1360Satz2BGBheran\.DassderKontexteineVerfassungsmäßigkeitsprüfungist,ändertdarannichts:maßgeblichistallein,dassdermarkierteSatzselbstfestlegt,welcheFällevon§1361Abs\.2BGBerfasstsind\.UmgekehrtwirdeinSatznichtdadurchzurDeutung,dasserinnerhalbeinersolchenPrüfungsteht\. NegativeBeispiele: <negative\_example\>DerGrundsatzderGleichheitderWahlistimSinneeinerstrengenundformalenGleichheitzuverstehen\.</negative\_example\> HierwirdzwarderGrundsatzderGleichheitderWahlgenannt;dabeihandeltessichaberumeinRechtsprinzip,nichtumeineGesetzesbestimmung\. <negative\_example\>WederdieAnordnungunddieDurchführungderNachwahlnochdieErmittlungunddieBekanntgabedesvorläufigenamtlichenWahlergebnissesamTagederHauptwahlverletztenVorschriftendesBundeswahlgesetzesoderderBundeswahlordnung\.</negative\_example\> EsmusssichfüreineDeutungumeineganzbestimmteGesetzesbestimmunghandeln\.DasisthiernichtderFall,eswirdvondenVorschriftendesBWGundderBWOalsGanzemgesprochen\. <negative\_example\>DieNachwahlermöglichtdenWählernindenbetroffenenWahlkreisenüberhaupterstdieTeilnahmeanderWahlundverwirklichtdamitdenGrundsatzderAllgemeinheitderWahl\(Art\.38Abs\.1Satz1GG\)\.</negative\_example\> DieGesetzesbestimmungwirdnuralsAnkerpunktfüreinenGrundsatzgenanntundistdamitkeintauglicherAusgangspunktfüreineDeutungimSinnederDefinition\. <negative\_example\>DerGesetzgebermussdieWahlgrundsätzeausdemGrundgesetzbeachten\.</negative\_example\> HierwirdaufkeinespezifischeGesetzesbestimmungBezuggenommen,sondernnurallgemeinaufdieWahlgrundsätzeverwiesen\. <negative\_example\>DieFeststellungeines\(vorläufigen\)ErgebnissesnachderHauptwahlunddessenBekanntgabedurchdenBundeswahlleitervorDurchführungderNachwahlstellenkeineBeeinträchtigungderFreiheitderWahldar\.</negative\_example\> HierwirdderallgemeineRechtsgrundsatzderFreiheitderWahlthematisiert,ohnedasseinespezifischeGesetzesbestimmunggenanntodereinekonkreteBehauptungüberderenInhaltgemachtwird\. <negative\_example\>§901ZPOistinfolgedessenbeifeststehenderLeistungsunfähigkeitdesSchuldnersnichtanwendbarundkanninsoweitArt\.2Abs\.2Satz2GGnichtverletzen\.</negative\_example\> BeideVorschriftensindgenaubezeichnet,unddennochliegtkeineDeutungvor:DerSatzbetrifftdieAnwendbarkeitdes§901ZPOimGanzenunddessenVerhältniszuArt\.2Abs\.2Satz2GG\.Erlegtdamitnichtfest,welcheFällevondereinenoderderanderenNormerfasstsind\.DieAntwortisteineleereListe\. <negative\_example\>DieAnordnungderHafterscheintschließlichimengerenSinneverhältnismäßig,weildieSchweredesEingriffsunddasGewichtderihnrechtfertigendenGründeinangemessenemVerhältniszueinanderstehen\.</negative\_example\> EshandeltsichumeinenSchritteinerVerhältnismäßigkeitsprüfung\.SolcheErwägungenbetreffendasVerhältniszweierNormenzueinanderundhabendaherkeinenbestimmtenInhaltimSinnedieserUntersuchung,auchwenndiegeprüfteVorschriftausdemKontexteindeutighervorgeht\. DuerhältsteineEingabe: \-‘text\_with\_context\_konkretes\_gesetz‘:einAuszugauseinerEntscheidungdesBundesverfassungsgerichts\.DerkonkretzuprüfendeSatzistinnerhalbdesAuszugsmitdenMarkierungen‘<deutung\>\.\.\.</deutung\>‘umschlossen;dieumgebendenSätzedienenausschließlichalsKontext\. BegründeimmerdeineAntwort,bevordudichentscheidest\.Prüfezunächst,obdieGrundaussage,diedermit‘<deutung\>‘markierteSatztrifft,denInhalteinerkonkretenGesetzesbestimmungausgestaltet\.DieGesetzesbestimmungmussdabeinichtimSatzselbstgenanntsein,sondernkannauchineinemanderenSatzstehen\.Sobaldduanfangenmusst,zurechtfertigen,dasseigentlichkeinekonkreteBestimmung\(sondernbspw\.nurgenerelldieBezugnahmeaufeinGesetzalsGanzes\)vorliegt,liegtkeinekonkreteGesetzesbestimmungvor\.Prüfesodann,obdieAussageeinenbestimmtenInhaltfestlegt,alsodazuführt,dasseineReihevonFällenvonderBestimmungerfasstistundanderenicht\.BetrifftderSatzstattdessendieVereinbarkeit,WirksamkeitoderAnwendbarkeiteinerNormimGanzen\-\-\-insbesonderealsTeileinerVerfassungsmäßigkeits\-oderVerhältnismäßigkeitsprüfung\-\-\-,soisternichtbestimmtunddieAntwortisteineleereListe\.Achtedarauf,dassRechtsgrundsätzealsGrundaussagenurinBetrachtkommen,wennderSatzeinespezifischeBehauptungüberderenInhaltenthältundsieauseinerkonkretenBestimmunghergeleitetwerden\.DassmehralseineGesetzesbestimmungausgelegtwird,kommtäußerstseltenvor,seientsprechendstreng\.FallseinedenAnforderungenentsprechendeGesetzesbestimmungvorliegt,nennedie\(se\)Gesetzesbestimmung\(en\)alsListe;fallsnicht,gibeineleereListezurück\. ### B\.4Argument gate \(argument\) DasübergreifendeProjekt,indemdugenutztwirst,isteineStudie,inderuntersuchtwird,wiesichdieVerwendungvonAuslegungsmethodenindenEntscheidungendesBundesverfassungsgerichtsimLaufederGeschichteentwickelthat\.DeineAufgabeistTeileinesWorkflows,dereszumZielhat,einenSatzMetadatenzuerstellen,derfürjedeEntscheidunggenauangibt,anwelcherStelledasGerichtAuslegungmittelsdervierAuslegungskriterienbetriebenhat\. Auslegungbedeutet,sichfüreineuntermehrerenmöglichenDeutungeneinerGesetzesbestimmungaufgrundvonÜberlegungenzuentscheiden\. AusdieserDefinitionlassensichdieverschiedenenSchritteableiten,ausdenensichderWorkflowzusammensetzt: 1\.Deutungenidentifizieren 2\.Identifizieren,obessichbeibestimmtenSätzenüberhauptumsolchehandelt,dieeineBegründungfürdieDeutungdarstellen\. 3\.Identifizieren,welcheAuslegungskriterienindemArgumentverwendetwerden\. DuwirstdichausschließlichmitdemzweitenSchrittdiesesWorkflowsbefassen\. Maßstabisthierbeidasfolgende: WirddieDeutungdurchdasmöglicheArgumentbegründet?DasmöglicheArgumentstelltnurdanneineBegründungfürdieDeutungdar,wennesinhaltlichundlogischdirektmitderDeutungverknüpftistunddengleichenArgumentationsstrangverfolgt\.SowohlPro\-alsauchContra\-ArgumentekönneneineBegründungdarstellen\.Sätze,diezueinemanderenArgumentationsstranggehörenodernurthematischähnlichsind,sindnichtalsBegründungzubetrachten\.Achtebesondersdarauf,obdasmöglicheArgumenttatsächlicheineBegründungliefert,unabhängigdavon,obesvorodernachderDeutungsteht\. DuerhältstvierEingaben: \-‘deutung\_sentence‘:derSatz,derdieDeutungenthält\.EineDeutungisteineBehauptung,diedasGerichtselbstaufstelltunddiezumGegenstandhat,dasseinebestimmteGesetzesbestimmungabstrakteinenbestimmtenInhalthabe\. \-‘argument\_sentence‘:derzuprüfendeSatz\-\-einpotenziellesArgument,dasausdertextlichenUmgebungderDeutunggegriffenwurdeundsiemöglicherweisebegründet\. \-‘norm‘:dieGesetzesnorm,aufdiesichdieDeutungbezieht\. \-‘text\_with\_context‘:derumliegendeKontext,indemDeutungundArgumentauftreten\. BegründeimmerdeineAntwort,bevordudichentscheidest\.ErfolgtdieDeutungaufgrunddes‘argument\_sentence‘? ### B\.5Grammatical interpretation \(wortlaut\) DasübergreifendeProjekt,indemdugenutztwirst,isteineStudie,inderuntersuchtwird,wiesichdieVerwendungvonAuslegungsmethodenindenEntscheidungendesBundesverfassungsgerichtsimLaufederGeschichteentwickelthat\.DeineAufgabeistTeileinesWorkflows,dereszumZielhat,einenSatzMetadatenzuerstellen,derfürjedeEntscheidunggenauangibt,anwelcherStelledasGerichtAuslegungmittelsdervierAuslegungskriterienbetriebenhat\. Auslegungbedeutet,sichfüreineuntermehrerenmöglichenDeutungeneinerGesetzesbestimmungaufgrundvonÜberlegungenzuentscheiden\. AusdieserDefinitionlassensichdieverschiedenenSchritteableiten,ausdenensichderWorkflowzusammensetzt: 1\.Deutungenidentifizieren 2\.Identifizieren,obessichbeibestimmtenSätzenüberhauptumsolchehandelt,dieeineBegründungfürdieDeutungdarstellen\. 3\.Identifizieren,welcheAuslegungskriterienindemArgumentverwendetwerden\. DuwirstdichausschließlichmitdemdrittenSchrittdiesesWorkflowsbefassen\. Auslegungskriteriumisthierbeidasfolgende: Wortlautauslegung\(=grammatischeAuslegungoderauchWortsinnauslegung\)istAuslegungnachderBedeutungeinesAusdrucksodereinerWortverbindungimallgemeinenoder,fallseinsolcherfeststellbarist,imbesonderenSprachgebrauchderbetreffendenGesetzesbestimmung\.DieGesetzesbestimmungmussdabeinichtimSatzselbstgenanntsein,sondernkannauchineinemanderenSatzstehen\.Wortlautauslegungistauchbereitsdannvorhanden,wennderSatzselbstnureinZitatinAnführungszeichen\(bspw\.,,vorderHauptwahl"\)enthält,dieSchlussfolgerungausdemZitataberineinemanderenSatzsteht\.NennedannaberauchdieSchlussfolgerung\.Sieliegtabernichtvor,wennnurdasganzeGesetzallgemeinGegenstandistodereineandereGesetzesbestimmungindiesemGesetz\. DuerhältstvierEingaben: \-‘deutung\_sentence‘:derSatz,derdieDeutungenthält\.EineDeutungisteineBehauptung,diedasGerichtselbstaufstelltunddiezumGegenstandhat,dasseinebestimmteGesetzesbestimmungabstrakteinenbestimmtenInhalthabe\. \-‘argument\_sentence‘:derzuprüfendeSatz\-\-einmöglichesArgument,dasausdertextlichenUmgebungderDeutunggegriffenwurdeundsiepotenziellbegründet\. \-‘norm‘:dieGesetzesnorm,aufdiesichdieDeutungbezieht\. \-‘text\_with\_context‘:derumliegendeKontext,indemDeutungundArgumentauftreten\. BegründeimmerdeineAntwort,bevordudichentscheidest\.Entsprichtdas‘argument\_sentence‘derDefinitionunddenBeispielenderWortlautauslegung?Berücksichtigeden‘text\_with\_context‘nur,umdas‘argument\_sentence‘besserzuverstehen,aberentscheideausschließlichaufBasisdesInhaltsvon‘argument\_sentence‘,obdieAuslegungsmethodeangewendetwird\.Wortlautauslegungliegtvor,sobaldzumindesteinTeildesWortlautsderselbenGesetzesbestimmungGegenstandist\.DerbloßeNamedesganzenGesetzesodereineandereGesetzesbestimmungindiesemGesetzalsdiederDeutungführenzukeinerWortlautauslegung\. ### B\.6Systematic interpretation \(systematik\) DasübergreifendeProjekt,indemdugenutztwirst,isteineStudie,inderuntersuchtwird,wiesichdieVerwendungvonAuslegungsmethodenindenEntscheidungendesBundesverfassungsgerichtsimLaufederGeschichteentwickelthat\.DeineAufgabeistTeileinesWorkflows,dereszumZielhat,einenSatzMetadatenzuerstellen,derfürjedeEntscheidunggenauangibt,anwelcherStelledasGerichtAuslegungmittelsdervierAuslegungskriterienbetriebenhat\. Auslegungbedeutet,sichfüreineuntermehrerenmöglichenDeutungeneinerGesetzesbestimmungaufgrundvonÜberlegungenzuentscheiden\. AusdieserDefinitionlassensichdieverschiedenenSchritteableiten,ausdenensichderWorkflowzusammensetzt: 1\.Deutungenidentifizieren 2\.Identifizieren,obessichbeibestimmtenSätzenüberhauptumsolchehandelt,dieeineBegründungfürdieDeutungdarstellen\. 3\.Identifizieren,welcheAuslegungskriterienindemArgumentverwendetwerden\. DuwirstdichausschließlichmitdemdrittenSchrittdiesesWorkflowsbefassen\. Auslegungskriteriumisthierbeidasfolgende: DiesystematischeAuslegungerfordert:Einespezifische,klarabgrenzbareGesetzesbestimmung,diesichvonderinderDeutunggenanntenunterscheidet\.BezugnahmenaufbloßeRechtsgrundsätze,Prinzipien,ganzeGesetzeoderunbestimmteVorschriftenzählennicht\.DiezitierteGesetzesbestimmungmussgenutztwerden,umdenInhaltderDeutungzuerklärenoderzuvertiefen;siekannaucheinegegenteiligePositionvertreten\.DieGesetzesbestimmungmussnichtimSatzselbstgenanntsein,sondernkannauchineinemanderenSatzstehen\. PositiveBeispiele: <example\>DerSchutzbereichdesArt\.6Abs\.1GGumfaßtauchdasVerhältniszwischenElternundihrenvolljährigenKindern\. <auslegung\>SchonnachdergesetzlichenAusgestaltungerschöpftsichdieBeziehungzwischenElternundKindernnichtinderErziehungsfunktionderFamilie\.\*\*DieherkömmlicheRegelungdesUnterhaltsrechtsverweistdeutlichaufeinelebenslangeVerpflichtungvonElternundKindern,einanderBeistandzuleisten\.MitderEinführungvon§1618aBGBhatderGesetzgeberalsLeitbildderEltern\-Kind\-BeziehungdievomAlterderKinderunabhängigewechselseitigePflichtzuBeistandundRücksichtnahmestatuiert\.\*\*</auslegung\></example\> IndemdieTextstelleauf\*\*§1618aBGB\*\*Bezugnimmt,umdenCharakterdes\*\*Art\.6Abs\.1GG\*\*näherzubestimmen,berücksichtigtsieanderesachlicheBestimmungen\.DieseliegenhierzwarineineranderenRegelungalsderBezugsbestimmung,betreffenaberdenselbenRegelungsgegenstandundsinddaherrelevantfürdasVerständnisderGesetzesbestimmung\. <example\>\*\*DerBegriffErgänzungsabgabebesagtlediglich,daßdieseAbgabedieEinkommen\-undKörperschaftsteuer,alsoaufDauerangelegteSteuern,ergänzen,d\.h\.ineinergewissenAkzessorietätzuihnenstehensoll\.\*\*</example\> UmdieBedeutungdesinderVorschriftverwendetenBegriffs,,Ergänzungsabgabe"näherzubestimmen,greiftderTextaufdieinanderenVorschriftendesGGgenanntenBegriffe,,Körperschaftsteuer"und,,Einkommensteuer"zurück\.ErberücksichtigthieralsodenKontextdesBegriffs,,Ergänzungsabgabe",wieerzumVerständnisdesBegriffserforderlichist\.EshandeltsichdaherumsystematischeAuslegung\. DuerhältstvierEingaben: \-‘deutung\_sentence‘:derSatz,derdieDeutungenthält\.EineDeutungisteineBehauptung,diedasGerichtselbstaufstelltunddiezumGegenstandhat,dasseinebestimmteGesetzesbestimmungabstrakteinenbestimmtenInhalthabe\. \-‘argument\_sentence‘:derzuprüfendeSatz\-\-einmöglichesArgument,dasausdertextlichenUmgebungderDeutunggegriffenwurdeundsiepotenziellbegründet\. \-‘norm‘:dieGesetzesnorm,aufdiesichdieDeutungbezieht\. \-‘text\_with\_context‘:derumliegendeKontext,indemDeutungundArgumentauftreten\. BegründeimmerdeineAntwort,bevordudichentscheidest\.Vorgehensweise:Prüfe,obdas‘argument\_sentence‘aufeinespezifische,klarabgrenzbareGesetzesbestimmungBezugnimmt,diesichvonderinderDeutunggenanntenunterscheidet\.SiekannauchindenSätzenvordem‘argument\_sentence‘stehen\.NennedieseGesetzesbestimmung\.Beurteile,obdieseGesetzesbestimmungverwendetwird,umdenInhaltdes‘argument\_sentence‘direktzuerklärenoderzuvertiefen;seihiergroßzügig\.SystematischeAuslegungliegtnurvor,wennbeideKriterienerfülltsind\. ### B\.7Historical interpretation \(geschichte\) DasübergreifendeProjekt,indemdugenutztwirst,isteineStudie,inderuntersuchtwird,wiesichdieVerwendungvonAuslegungsmethodenindenEntscheidungendesBundesverfassungsgerichtsimLaufederGeschichteentwickelthat\.DeineAufgabeistTeileinesWorkflows,dereszumZielhat,einenSatzMetadatenzuerstellen,derfürjedeEntscheidunggenauangibt,anwelcherStelledasGerichtAuslegungmittelsdervierAuslegungskriterienbetriebenhat\. Auslegungbedeutet,sichfüreineuntermehrerenmöglichenDeutungeneinerGesetzesbestimmungaufgrundvonÜberlegungenzuentscheiden\. AusdieserDefinitionlassensichdieverschiedenenSchritteableiten,ausdenensichderWorkflowzusammensetzt: 1\.Deutungenidentifizieren 2\.Identifizieren,obessichbeibestimmtenSätzenüberhauptumsolchehandelt,dieeineBegründungfürdieDeutungdarstellen\. 3\.Identifizieren,welcheAuslegungskriterienindemArgumentverwendetwerden\. DuwirstdichausschließlichmitdemdrittenSchrittdiesesWorkflowsbefassen\. Auslegungskriteriumisthierbeidasfolgende: Subjektiv\-teleologischeAuslegung\(=historischeAuslegung,auchAuslegungnachEntstehungsgeschichte\)liegtvorbeieinerAuslegungnachderRegelungsabsichtdesGesetzgebers,derdieNormgeschaffenhat\.DieNennungdesWortlautsoderallgemeineErwägungenüberdenInhaltdesGesetzesstellennochkeinehistorischeAuslegungdar\.AusführungensindgrundsätzlichdemGerichtselbst\-\-\-nichtdemGesetzgeber\-\-\-zuzurechnen,wenndasGerichtdiesnichtklarzumAusdruckbringt\. PositiveBeispiele: <example\>Deutung:MitderEinführungvon§1618aBGBhatderGesetzgeberalsLeitbildderEltern\-Kind\-BeziehungdievomAlterderKinderunabhängigewechselseitigePflichtzuBeistandundRücksichtnahmestatuiert\.MöglichesArgument:DadurchsollteinsbesonderezueinergrößerenFamilienautonomiebeigetragenundGefährdungenderFamiliealsInstitutionentgegengewirktwerden\(BTDrucks\.8/2788S\.43\)\.</example\> DieTextstellenimmtBezugaufeineBundestagsdrucksache\.IndieserkommtdieRegelungsabsichtdesGesetzgeberszumAusdruck\. NegativeBeispiele: <negative\_example\>MöglichesArgument:DerangefochteneBeschlußdesLandgerichtsBremenberuhtauchaufderVerletzungvonArt\.103Abs\.1GG\.Deutung:Eskannnichtausgeschlossenwerden,daßdasLandgerichtbeiBeachtungvonBedeutungundTragweitedesArt\.103Abs\.1GGdemBeschwerdeführerWiedereinsetzungindenvorigenStandgewährthätte\.</negative\_example\> IndiesemFallwirdzwaraufdieBedeutungundTragweitedesArt\.103Abs\.1GGverwiesen\.Darauslässtsichabernichtschließen,dassdarindieRegelungsabsichtdesGesetzgeberszumAusdruckkommt\.EshandeltsichvielmehrumeinebloßeErwägungdesGerichts\. DuerhältstvierEingaben: \-‘deutung\_sentence‘:derSatz,derdieDeutungenthält\.EineDeutungisteineBehauptung,diedasGerichtselbstaufstelltunddiezumGegenstandhat,dasseinebestimmteGesetzesbestimmungabstrakteinenbestimmtenInhalthabe\. \-‘argument\_sentence‘:derzuprüfendeSatz\-\-einmöglichesArgument,dasausdertextlichenUmgebungderDeutunggegriffenwurdeundsiepotenziellbegründet\. \-‘norm‘:dieGesetzesnorm,aufdiesichdieDeutungbezieht\. \-‘text\_with\_context‘:derumliegendeKontext,indemDeutungundArgumentauftreten\. BegründeimmerdeineAntwort,bevordudichentscheidest\.Entsprichtdas‘argument\_sentence‘derDefinitionunddenBeispielendersubjektiv\-teleologischenAuslegung?Liegtim‘argument\_sentence‘nachdendargelegtenMaßstäbeneinsubjektiv\-teleologischesArgumentvor,istdieKategorieerfüllt;andernfallsnicht\. ### B\.8Objective\-teleological interpretation \(zweck\) DasübergreifendeProjekt,indemdugenutztwirst,isteineStudie,inderuntersuchtwird,wiesichdieVerwendungvonAuslegungsmethodenindenEntscheidungendesBundesverfassungsgerichtsimLaufederGeschichteentwickelthat\.DeineAufgabeistTeileinesWorkflows,dereszumZielhat,einenSatzMetadatenzuerstellen,derfürjedeEntscheidunggenauangibt,anwelcherStelledasGerichtAuslegungmittelsdervierAuslegungskriterienbetriebenhat\. Auslegungbedeutet,sichfüreineuntermehrerenmöglichenDeutungeneinerGesetzesbestimmungaufgrundvonÜberlegungenzuentscheiden\. AusdieserDefinitionlassensichdieverschiedenenSchritteableiten,ausdenensichderWorkflowzusammensetzt: 1\.Deutungenidentifizieren 2\.Identifizieren,obessichbeibestimmtenSätzenüberhauptumsolchehandelt,dieeineBegründungfürdieDeutungdarstellen\. 3\.Identifizieren,welcheAuslegungskriterienindemArgumentverwendetwerden\. DuwirstdichausschließlichmitdemdrittenSchrittdiesesWorkflowsbefassen\. Auslegungskriteriumisthierbeidasfolgende: Objektiv\-teleologischeAuslegungbedeutet,dassdieInterpretationeinerNormanhandihresobjektivenZweckserfolgt,wieersichausderheutigenRechtsordnungunddenallgemeinenPrinzipienergibt,ohneRückgriffaufdieAbsichtendeshistorischenGesetzgebers\.Beiderobjektiv\-teleologischenAuslegungwerdeninsbesonderedieaktuellenStrukturendesgeregeltenSachbereichsunddiezugrundeliegendenrechtsethischenPrinzipienherangezogen\. PositiveBeispiele: <example\>DieRegelungdes§37BWGsolldieschnelleFeststellungdesWahlergebnissesermöglichen,umdiedemokratischeLegitimationdergewähltenVertretersicherzustellen\.Daheristesangemessen,dasWahlergebnisunmittelbarnachderHauptwahlfestzustellen\.</example\> HierwirdderaktuelleZweckderNormbetont,ohnehistorischeBezügeoderGesetzesmaterialienzuverwenden\.Eshandeltsichumeineobjektiv\-teleologischeAuslegung\. <example\>DiePflichtzurZahlungvonSteuerndientderFinanzierungstaatlicherAufgaben,dieimInteressederAllgemeinheitliegen\.Deshalbistesgerechtfertigt,dassalleBürgerentsprechendihrerLeistungsfähigkeitzurSteuerherangezogenwerden\.</example\> DerobjektiveZweckderNorm,nämlichdieFinanzierungstaatlicherAufgabenzumWohlederAllgemeinheit,wirdhervorgehoben\.EswerdenkeinehistorischenBezügeverwendet\. NegativeBeispiele: <negative\_example\>DerBegriffErgänzungsabgabebesagtlediglich,dassdieseAbgabedieEinkommen\-undKörperschaftsteuerergänzt,alsoineinergewissenAkzessorietätzuihnensteht\.</negative\_example\> HierwirdnichtaufdenZweckoderdasZielderNormeingegangen,sondernlediglichaufihreBeziehungzuanderenSteuern\.Eshandeltsichnichtumeineobjektiv\-teleologischeAuslegung\. <negative\_example\>DieNormwurdeursprünglicheingeführt,umnachdenEreignissender1950erJahrediewirtschaftlicheStabilitätzufördern\(vgl\.Gesetzesbegründungvon1956\)\.Daheristsiesozuverstehen,dasssieauchheutenochdiesemZweckdient\.</negative\_example\> ObwohlhierderZweckderNormdiskutiertwird,wirdaufhistorischeBezügeundGesetzesmaterialienzurückgegriffen\.Eshandeltsichnichtumeinereinobjektiv\-teleologischeAuslegunggemäßderDefinition\. DuerhältstvierEingaben: \-‘deutung\_sentence‘:derSatz,derdieDeutungenthält\.EineDeutungisteineBehauptung,diedasGerichtselbstaufstelltunddiezumGegenstandhat,dasseinebestimmteGesetzesbestimmungabstrakteinenbestimmtenInhalthabe\. \-‘argument\_sentence‘:derzuprüfendeSatz\-\-einmöglichesArgument,dasausdertextlichenUmgebungderDeutunggegriffenwurdeundsiepotenziellbegründet\. \-‘norm‘:dieGesetzesnorm,aufdiesichdieDeutungbezieht\. \-‘text\_with\_context‘:derumliegendeKontext,indemDeutungundArgumentauftreten\. BegründeimmerdeineAntwort,bevordudichentscheidest\.Entsprichtdas‘argument\_sentence‘derDefinitionunddenBeispielendesobjektiv\-teleologischenAuslegungskriteriumsvonoben?Gehewiefolgtvor:HandeltessichumeineAuslegungnachdemZweckoderZielderNorm,ohnehistorischeBezügezuverwenden?Achtebesondersdarauf,dassdas‘argument\_sentence‘denaktuellenZweckderNorminderheutigenRechtsordnunghervorhebt\.Liegteinobjektiv\-teleologischesArgumentvor? ## Appendix CDataset Record Schema Every record is a JSON object with three top\-level keys: an integerid, arowholding the sentence\(s\) under classification together with their surrounding context, andobservedholding the gold labels for the subtask\. One stripped example from each of the two datasets follows; some paragraph\-level bookkeeping fields are elided and long text fields are truncated with …\. ##### Record identity\. Every instance is identified by a small tuple of provenance fields, all recoverable from the annotation export\. Adeutungenrecord \(a single candidate sentence\) is identified by\(decision\_name, sentence\_id\), wheresentence\_idis thesatz\_idof that sentence\. Anargumentsrecord \(a reading–candidate pair\) is identified by\(decision\_name, deutung\_id, argument\_start\_index\):deutung\_idis the reading’s stable annotation identifier, carried through verbatim from the export and shared by every candidate paired with that reading, andargument\_start\_indexis the character offset at which the candidate sentence’ssatzspan begins, as recorded in the earlier exportdataset\-project\-46\-2026\-05\-14\-17\-12\-32\.json\(thestartfield already carried on that candidate’spotential\_argumentsrow, which equals that export’sdocument\_structuresatzstart\)\. ### C\.1arguments\(example from subtaskwortlaut\) Here therowis a \(reading, candidate\) pair:deutung\_sentenceis the reading,argument\_sentencethe candidate justification, andobservedrecords the gate \(argument\) and each canon label\. \{ "id":3, "row":\{ "decision\_name":"BVerfGE32,273", "deutung\_id":"3tDBjYneKo", "deutung\_sentence":"DafürsprichtschonderWortlaut,aberauch…", "argument\_sentence":"GegenüberArt\.119Abs\.3WRVhatArt\.6Abs\.4GG…", "argument\_start\_index":7391, "relative\_position":1, "text\_with\_context":"…<deutung\>…</deutung\>…<potential\_argument\>…</potential\_argument\>…", "norm":\["Art\.6IVGG"\] \}, "observed":\{ "wortlaut":1,"systematik":0,"geschichte":1,"zweck":0, "argument":1, "reasoning":"DieWortverbindung\\"jederMutter\\"wirdherangezogen…" \} \} ### C\.2deutungen\(example from subtaskkonkretes\_gesetz\) Here therowis a single reading marked intext\_with\_context;observed\.konkretes\_gesetzis the list of concrete provisions it interprets \(empty when none\), while the reading\-criteria subtasks read the boolean fields ofobserved\. \{ "id":2, "row":\{ "decision\_name":"BVerfGE32,273", "sentence\_id":67, "text":"VerhältsicheinewerdendeMutternachdieserVorschrift,…", "is\_deutung":false, "text\_with\_context":"…<deutung\>…</deutung\>…", "norm":\["§5Abs\.1Satz1MuSchG"\], "vermutlich\_keine\_behauptung":false \}, "observed":\{ "nicht\_abstrakt":0, "nicht\_selbst\_aufgestellt":0, "konkretes\_gesetz":\["§5Abs\.1Satz1MuSchG"\], "reasoning":"" \} \} ## Appendix DReproducing the Scores and Benchmark Datasets To verify that the reported scores and benchmark datasets follow from the released artifacts and methods descriptions, we set up two tasks in the Harbor agent evaluation framework\([Harbor Framework Team, 2026](https://arxiv.org/html/2609.26945#bib.bib18)\)\. In each, an agent independently implements the described procedure without access to our implementation, and a withheld verifier compares its output against the reference\. For the scores, an agent working in a network\-isolated container is given the 64 sets of model predictions, the recorded judge decisions, the test\-split gold labels for all eight subtasks, and the parts of this paper that define the metrics \([Section5](https://arxiv.org/html/2609.26945#S5)and the table captions of[AppendixA](https://arxiv.org/html/2609.26945#A1)\), but neither our scoring code nor any reported number\. The task description specifies the keys of the output table\. A withheld verifier then recomputes every cell with our own scorer and compares it against the agent’s output\. The agent reproduces all 208 point estimates exactly: for each of the four models and two prompt conditions, precision, recall andF1F\_\{1\}on the seven binary subtasks, their mean, and the four statutory\-reference metrics\.121212Agreement is checked at a floating\-point tolerance of10−610^\{\-6\}; every cell matched exactly\. The bootstrap confidence intervals are outside the check, since they depend on the resampling seed\.The scores in this paper are therefore fully determined by the protocol of[Section5](https://arxiv.org/html/2609.26945#S5)together with the published predictions, judge decisions and gold labels\. For the benchmark datasets, the agent is given the raw annotation source files and the construction procedure specified in[AppendixE](https://arxiv.org/html/2609.26945#A5), but neither our construction code nor the reference datasets\. It reproduces all 2,200 records across eight subtasks exactly, including their field values, split membership, and identifiers\. ## Appendix EBenchmark Construction \(Methods Appendix\) *Note: In keeping with the reproduction approach described in[AppendixD](https://arxiv.org/html/2609.26945#A4), this appendix was generated entirely by an LLM\.* This appendix specifies, end to end, the deterministic transformation from the raw annotation source files to the finished benchmark\. The construction isexact: the same source files yield the same files, down to every record, every field, and everyid\. Every pseudo\-random step has a fixed generator, seed, and draw order, so "random" here never means "unpredictable"\. Several early definitions — the record schema and the sentence helpers — recur throughout\. ### E\.1What the benchmark contains The benchmark comprises eightsubtasks, each split into three files —train,validation,test— for24 filesin total\. Each file is a JSON array ofrecords, and every record has exactly three top\-level keys: \{"id":<integer\>,"row":\{…\},"observed":\{…\}\} The eight subtasks form two families: - •theargument family\(five subtasks:argument,wortlaut,systematik,geschichte,zweck\), where one record is one*\(reading, candidate\-sentence\)*pair; - •thereading family\(three subtasks:konkretes\_gesetz,nicht\_abstrakt,nicht\_selbst\_aufgestellt\), where one record is one*reading*— a single annotated sentence\. Each subtask occupies its own directory, as<subtask\>/train\.json,<subtask\>/validation\.json,<subtask\>/test\.json\. Only field values are significant,idincluded; the order of keys within a JSON object and the whitespace of the files are not\. ### E\.2Source files The four source files are three export forms of one Label Studio annotation project — each preserving something the others drop, so each is consulted only for what is named below — together with one small list\. #### E\.2\.1dataset\-project\-46\-2026\-05\-14\-17\-12\-32\.json— "structure export" A JSON array ofdecisions\. Each decision object carries: - •decision\_name— the decision’s identifier \(a string\)\. - •plain\_text— the decision’s full text as one string\. All character offsets in every file index into this string\. - •document\_structure, which includes \(among others\) asatzarray holding the sentence segmentation\. Eachsatz\(sentence\) object carriessatz\_id\(an integer, unique within the decision\),startandend\(character offsets intoplain\_text, half\-open, so the text isplain\_text\[start:end\]\),absatzID,ebene1nr,tbeg, andscore\. - •potential\_deutungen— the array of annotated*readings*\. Each entry hasid\(a short opaque string, theregion id, which is the join key across all exports\),satz\_id\(the sentence the reading sits on\),abstraktandselbst\_aufgestellt\(eachtrue,false, ornull\), and reasoning strings\. - •potential\_arguments— the array of annotated*argument candidates*\. Each entry hasid\(a region id\),satz\_id,startandend\(offsets of the candidate sentence\),deutung\_id\(the region id of the reading the candidate is attached to\),annotation\_method\(a string, whose value"individually"marks a per\-pair human judgment\), and one object per canon \(general\_argument,wortlaut,systematik,geschichte,zweck\), each eithernullor an object\{ "present": true\|false, "reasoning": "…" \}\. This file is the source of all sentence text, all offsets, all context windows, and the reading/argument structure\. The two remaining exports supply only the specific fields named below\. #### E\.2\.2raw\-labelstudio\-export\-project\-46\-2026\-07\-25\.json— "raw dump" The pre\-merge Label Studio export: a JSON array of task objects, each with anannotationsarray, each annotation aresultarray of regions\. Each region has anid\(the same region id as above\), afrom\_name\(which annotation field it is\), atype, and avalue\. Three fields survive only here and are read only from this file: - •bestimmt\_moeglicheDeutung\(typechoices\): the reading’s determinate\-content judgment, taken asvalue\.choices\[0\]— one of"Yes","No", or \(when empty\) absent\. Indexed by region id\. - •konkretes\_gesetz\_moeglicheDeutung\(typetextarea\): the list of concrete statutory provisions the reading names\. Herevalue\.textis a list of strings, one per line the annotator entered; each entry is stripped of surrounding whitespace, empties are dropped, and the order is preserved\. \(This field must be read from this export, because the others merge the list into one unsplittable string\.\) Indexed by region id; a reading with no such region has the empty list\. - •vermutlich\_keine\_behauptung\_moeglicheDeutung\(typechoices\): a boolean flag,trueexactly whenvalue\.choices\[0\] == "Yes"\. Indexed by region id; absent meansfalse\. #### E\.2\.3postprocessing\-export\-project\-46\-2026\-05\-14\-17\-12\-32\.json— "label export" The merged export, used only by the two reading\-criteria subtasks \(nicht\_abstrakt,nicht\_selbst\_aufgestellt\) as the source of the human Yes/No criterion answers\. Its task/annotation/result nesting matches the raw dump\. Per region id it supplies: - •abstrakt\_moeglicheDeutung\(typechoices\):value\.choices\[0\]∈\\in\{"Yes","No"\}, or absent\. - •selbst\_aufgestellt\_moeglicheDeutung\(typechoices\): likewise\. - •vermutlich\_keine\_behauptung\_moeglicheDeutung: a boolean, as in[SectionE\.2\.2](https://arxiv.org/html/2609.26945#A5.SS2.SSS2)\. - •a region withfrom\_name == "label"whosevalue\.labelscontains"moeglicheDeutung", which marks that region id as an annotated reading\. #### E\.2\.4selective\_decisions\.json A JSON array of decision names: theselectively annotateddecisions, for which not every sentence was reviewed\. Every other decision in the structure export isexhaustively annotated\. Several rules below turn on this distinction\. ### E\.3Global conventions #### E\.3\.1The random number generator Every stochastic step uses the standard CPythonrandommodule generator — the Mersenne\-Twister, MT19937, i\.e\.random\.Random— with the seed fixed atSEED = 42\. Determinism rests on three things, each fixed at its point of use: the seed, which generator instance is used where, and the order in which draws are taken\. The operations, with their standard\-library semantics, arerandrange\(3\)\(a uniform integer in \{0, 1, 2\}\),choice\(seq\)\(one uniformly chosen element\),shuffle\(list\)\(Fisher–Yates in place, the module’s algorithm\), andsample\(population, k\)\(k distinct elements, the module’s algorithm\)\. Theidnumbering and split membership are defined by the precise MT19937 stream these produce, not by any re\-implementation — a hand\-rolled Fisher–Yates or NumPy would diverge\. Generators are never interleaved: wherever a step starts a fresh generator, that is a newrandom\.Random\(42\), independent of any other\. #### E\.3\.2The split order Wherever the three splits are iterated — partitioning, sampling, numbering — they are visited in the fixed ordertest, then train, then validation\. Per\-split targets, written as*\(positives, negatives\)*, are: - •reading family:test = \(60, 180\),train = \(20, 60\),validation = \(20, 60\)— 100 positives and 300 negatives per subtask overall; - •argument family:test = \(30, 90\),train = \(10, 30\),validation = \(10, 30\)— 50 positives and 150 negatives per subtask overall\. #### E\.3\.3Sentence usability \(is\_evaluated\) The sentence segmentation emits some fragments that are not classifiable units \(section headings such as "II\." or "B\.\-I\."\)\. A sentence’s text isusableexactly when both its length exceeds 5 characters and it contains at least one alphabetic letter from A–Z, a–z, or the German letters ä ö ü Ä Ö Ü ß\. Non\-usable sentences are dropped before anything else and never enter any pool\. #### E\.3\.4Sentence text and context windows Thetextof a sentence with offsetsstart/endisplain\_text\[start:end\], with every newline replaced by a single space\. Acontext windowaround one or more target sentences, for integer countsbeforeandafter, is formed as follows\. The decision’ssatzlist is sorted bysatz\_idascending, which gives a linear sentence order and, for eachsatz\_id, an index within it\. Withloandhithe smallest and largest indices among the target sentences, the window spans indicesmax\(0, lo \- before\)throughmin\(last\_index, hi \+ after\), inclusive\. Those sentences’ texts \(each computed as above\) are concatenated with a single space between consecutive sentences\. Certain target sentences aremarked— wrapped with an opening and closing tag around that sentence’s text before concatenation; the tags per subtask are given below\. The window sizes areDEUTUNG\_BEFORE = DEUTUNG\_AFTER = 2,KONKRETES\_GESETZ\_BEFORE = 10\(with its "after" still 2\), andbefore = after = 2for argument windows\. ### E\.4The reading family: the reading pool Bothkonkretes\_gesetzand the twonicht\_\*criteria begin from*readings*, but along two build paths that were authored separately and read theabstrakt/selbst\_aufgestelltjudgments from different files; they are described separately and do not share a pool between[SectionE\.4](https://arxiv.org/html/2609.26945#A5.SS4)and[SectionE\.6](https://arxiv.org/html/2609.26945#A5.SS6)\. #### E\.4\.1Thekonkretes\_gesetzpool \(build\_deutung\_entries\) Decisions are processed in file order and, within each, thepotential\_deutungenin file order\. A reading on sentencesatz\_idis skipped whensatz\_idis not among the decision’s sentences, and skipped when its sentence text is not usable \([SectionE\.3\.3](https://arxiv.org/html/2609.26945#A5.SS3.SSS3)\)\. Itsprovision list\(norms,[SectionE\.2\.2](https://arxiv.org/html/2609.26945#A5.SS2.SSS2)\) is looked up by its regionid; denote that listnorm\(possibly empty\)\. The reading is apositiveexactly whennormis non\-empty\. A non\-positive reading may still become anegative, but only when it is a*recorded*decision, which requires both of the following\. First, it must havereached the criterion: its decision is exhaustively annotated \(itsdecision\_nameis not inselective\_decisions\), and the reading’sabstraktistrueand itsselbst\_aufgestelltistrue\(both read from the structure export\)\. Second,bestimmtfor this region id \([SectionE\.2\.2](https://arxiv.org/html/2609.26945#A5.SS2.SSS2)\) must be answered — exactly"Yes"or"No"\. When either condition fails, the reading is neither positive nor negative and is dropped\. A negative’skindis"kein\_bestimmter\_inhalt"whenbestimmt == "No"and"keine\_konkrete\_gesetzesbestimmung"whenbestimmt == "Yes"; positives have kindnull\. The record marks the reading’s own sentence with\("<deutung\>", "</deutung\>"\)\. Itsrowholdsdecision\_name;ebene1nr\(the sentence’sebene1nr\);absatz\_id\(the sentence’sabsatzID\);sentence\_id\(satz\_id\);start\_indexandend\_index\(the sentence’s offsets\);tbeg;score\(the sentence’sscore, or0\.0when absent\);text\(the sentence text\);should\_be\_evaluated\(truehere, the sentence having passed[SectionE\.3\.3](https://arxiv.org/html/2609.26945#A5.SS3.SSS3)\);deutung\_id\(the region id\);is\_deutung\(alwaysfalsefor this subtask\);text\_with\_context\(the context window withbefore = after = 2, the reading sentence marked\);text\_with\_context\_konkretes\_gesetz\(the context window withbefore = 10,after = 2, the reading sentence marked\);norm\(the provision list\);negative\_kind\(from the kind rule above\); andvermutlich\_keine\_behauptung\(the boolean flag,[SectionE\.2\.2](https://arxiv.org/html/2609.26945#A5.SS2.SSS2)\)\. Itsobservedis\{ "abstrakt": 1 if the reading’s abstrakt is true else 0, "selbst\_aufgestellt": 1 if true else 0, "konkretes\_gesetz": <the provision list\>, "reasoning": <konkrete\_bezugnahme\_reasoning or ""\> \}\. Positivity for this subtask is defined byobserved\.konkretes\_gesetzbeing non\-empty\. Each pool entry retains itsdecision\_nameand, for negatives, its kind\. ### E\.5The argument family: the pair pool A single pass builds one pool of*\(reading, candidate\)*pairs, from which the five argument subtasks are then cut\. Decisions are processed in file order and, within each, thepotential\_argumentsin file order\. A candidate with regionid, sentencea\_satz\_id, offsetsstart/end, anddeutung\_idis treated as follows, where itsstart/endare the offsets of the candidate’sannotated spaninpotential\_arguments\. The reading named bydeutung\_idis found in the decision’spotential\_deutungen; when there is none, the candidate is skipped\. The candidate is also skipped when either the reading’s sentence or the candidate’s sentence is missing from the decision’s sentences\. Agate on a genuine readingthen applies: the pair is kept only when the reading hasabstrakt == true,selbst\_aufgestellt == true, and a non\-empty provision list \(norms,[SectionE\.2\.2](https://arxiv.org/html/2609.26945#A5.SS2.SSS2), by the reading’s region id\)\. The pair’skeyis the triple\(decision\_name, deutung\_id, candidate span start offset\), and pairs arededuplicatedon this key: when a pair with the same key was already kept, the new one is ignored, unless the new one is individually judged \(annotation\_method == "individually"\) and the kept one was not, in which case the individually judged record replaces it — so the individually judged record wins a collision, and otherwise the first seen stays\. Finally, the candidate sentence text is computed, and the pair is skipped when it is not usable \([SectionE\.3\.3](https://arxiv.org/html/2609.26945#A5.SS3.SSS3)\)\. The record marks the reading’s sentence with\("<deutung\>", "</deutung\>"\)and the candidate’s sentence with\("<potential\_argument\>", "</potential\_argument\>"\); when the two are the same sentence, that single sentence is instead wrapped\("<deutung\><potential\_argument\>", "</potential\_argument\></deutung\>"\)\. Itsrowholdsdecision\_name;deutung\_id;ebene1nr\_deutungandabsatz\_id\_deutung\(the reading sentence’sebene1nrandabsatzID\);deutung\_sentence\(the reading sentence text\);argument\_sentence\(the text of the candidate’ssentencea\_satz\_id, taken from thesatzsegmentation per[SectionE\.3\.4](https://arxiv.org/html/2609.26945#A5.SS3.SSS4)\);argument\_start\_indexandargument\_end\_index\(thestartandendof that samesatzsentence — its segmentation offsets, which are what these fields record even in the rare case where the candidate’s annotatedstart/endspan reaches beyond its sentence; they are not the span’s offsets\);ebene1nr\_argumentandabsatz\_id\_argument\(that sentence’sebene1nrandabsatzID\);relative\_position\(the candidatesatz\_idminus the readingsatz\_id\);text\_with\_context\(the context window over both target sentences,before = after = 2, both marked as above\); andnorm\(the reading’s provision list\)\. Itsobservedcarries one flag per canon plus the gate flag:wortlaut,systematik,geschichte, andzweckare each1when that canon object on the candidate haspresent == true, else0; andargumentis1whengeneral\_argument\.present == true, else0\. \(Areasoningstring is added per subtask at selection time; see[SectionE\.7\.3](https://arxiv.org/html/2609.26945#A5.SS7.SSS3)\.\) Each pair retains itsdecision\_name, whether it is individually judged, and the raw canon objects\. #### E\.5\.1Cutting the five subtasks from the pair pool Forargument, the positives are the pairs withobserved\.argument == 1and the negatives the pairs withobserved\.argument == 0\(positivity for this subtask being theargumentflag\)\. For each canoncin \{wortlaut,systematik,geschichte,zweck\}, the pool is first restricted to thegate\-positivepairs \(observed\.argument == 1\); among those, the positives are the pairs whose canon objectchaspresent == true, and the negatives are the pairs whose canon objectcexists \(is notnull\) and haspresent == false\. Pairs wherecisnull— the canon was never asked — are neither, and are not eligible as negatives\. \(Positivity for a canon subtask is that canon’s flag\.\) ### E\.6The reading\-criteria family:nicht\_abstrakt,nicht\_selbst\_aufgestellt These two subtasks use aflipped polarity— the positive class is the rarer "No" answer to the underlying criterion\. Their candidate set is drawn from the structure export but their labels from the label export \([SectionE\.2\.3](https://arxiv.org/html/2609.26945#A5.SS2.SSS3)\), and they use only exhaustively annotated decisions\. #### E\.6\.1Candidate spans Decisions are processed in file order, skipping any whosedecision\_nameis inselective\_decisions, and within each thepotential\_deutungenin file order\. A reading onsatz\_idthat is present among the sentences and whose sentence text is usable \([SectionE\.3\.3](https://arxiv.org/html/2609.26945#A5.SS3.SSS3)\) yields a candidate span keyed by the reading’s regionid, with arowholdingdecision\_name,ebene1nr,absatz\_id,sentence\_id,start\_index,end\_index,tbeg,score,text,should\_be\_evaluated\(true\),deutung\_id,is\_deutung\(which istrueexactly when the reading’sabstraktistrue, andfalseotherwise — differing from[SectionE\.4\.1](https://arxiv.org/html/2609.26945#A5.SS4.SSS1)\),text\_with\_context\(before = after = 2, the reading sentence marked\("<deutung\>", "</deutung\>"\)\),text\_with\_context\_konkretes\_gesetz\(before = 10,after = 2\), andnorm\(the provision list,[SectionE\.2\.2](https://arxiv.org/html/2609.26945#A5.SS2.SSS2)\)\. Each span also retains the two reasoning strings from the structure export \(abstrakt\_reasoning,selbst\_aufgestellt\_reasoning, each or""\) and the provision list\. #### E\.6\.2Labels and polarity Each criterion has abase fieldand aflipped name:nicht\_abstraktfrom base fieldabstrakt\_moeglicheDeutung, andnicht\_selbst\_aufgestelltfrom base fieldselbst\_aufgestellt\_moeglicheDeutung\. For each candidate span, the base field’s answer is read from the label export \([SectionE\.2\.3](https://arxiv.org/html/2609.26945#A5.SS2.SSS3)\) by region id\. The span is an instance of the subtask only when that answer is exactly"Yes"or"No"\. Under the flip,"No"is the subtask’s positive and"Yes"its negative\. #### E\.6\.3The record’sobserved\(dual encoding\) Each reading\-criteria record carries both criteria’s encodings — regardless of which subtask file it lands in — together with the provision list\. Itsobservedhaskonkretes\_gesetzequal to the span’s provision list\. For each of the two criteria \(with its base field, its base keyabstraktorselbst\_aufgestellt, and its flipped name\), the base field’s answer for the span is read from the label export: when it is"Yes"or"No",observed\[base\_key\]is1 if "Yes" else 0andobserved\[flipped\_name\]is1−\-that; when it is neither, bothobserved\[base\_key\]andobserved\[flipped\_name\]arenull\. Thereasoningfield concatenates, for each criterion that was answered and whose corresponding reasoning string \([SectionE\.6\.1](https://arxiv.org/html/2609.26945#A5.SS6.SSS1)\) is non\-empty, the fragment"\[<base\_key\>\] <reasoning\>", joined by single spaces, with theabstraktcriterion first and thenselbst\_aufgestellt\. Therowis the span’s row \([SectionE\.6\.1](https://arxiv.org/html/2609.26945#A5.SS6.SSS1)\) with one field added:vermutlich\_keine\_behauptung, the boolean flag for the span read from the label export \(absent meansfalse\)\. ### E\.7Selecting, splitting, ordering, and numbering The steps above define, per subtask, apositive pooland anegative poolof records, each tagged with itsdecision\_name\. The remaining steps choose which records enter the benchmark, assign them to splits, order them, and number them\. The two families use different split\-search and sampling procedures, each described below\. Two rules are common to all subtasks\. First,splits are decision\-disjoint: every record from a given decision goes to exactly one split, so splitting is an assignment of whole decisions totest,train, orvalidation\. Second, aprompt\-example hold\-outapplies: a few decisions are quoted verbatim in a subtask’s worked\-example prompt and therefore never appear in that subtask’stestsplit\. The forbidden\-from\-test decisions are, per subtask,BVerfGE61,126andcs20090421\_2bvc000206forkonkretes\_gesetz;BVerfGE57,170forsystematik;BVerfGE57,170forgeschichte;BVerfGE57,170,BVerfGE61,126, andBVerfGE62,338fornicht\_abstrakt;BVerfGE57,170fornicht\_selbst\_aufgestellt; and none for the remaining subtasks \(argument,wortlaut,zweck\)\. #### E\.7\.1konkretes\_gesetz: split, then sample per negative kind Targets\.konkretes\_gesetzis a reading\-family subtask \([SectionE\.1](https://arxiv.org/html/2609.26945#A5.SS1)\) and therefore uses the reading\-family per\-split targets of[SectionE\.3\.2](https://arxiv.org/html/2609.26945#A5.SS3.SSS2):test = \(60, 180\),train = \(20, 60\),validation = \(20, 60\)— 100 positives and 300 negatives in all\. It does not take the argument\-family targets\(30, 90\)/\(10, 30\)/\(10, 30\), which belong only toargumentand the four canons even though[SectionE\.7\.3](https://arxiv.org/html/2609.26945#A5.SS7.SSS3)reuses this subtask’s split\-search and sampler\. Kind targets\.Because the two negative kinds are not interchangeable, each split holds them at the pool’s overall proportion\. Withtotal\_negthe size of the whole negative pool andrarethe count of kindkeine\_konkrete\_gesetzesbestimmungwithin it, a split with negative targettnhas a rare\-kind target ofround\(tn \* rare / total\_neg\)and akein\_bestimmter\_inhalttarget oftnminus that;roundis banker’s rounding, i\.e\. Python’s built\-inround\. Split search \(the "random\-restart" search\)\.The search ranges only over thepool decisions: the distinctdecision\_names that own at least one record in this subtask’s pool \(its positives or its negatives\)\. A decision with no record for this subtask is not a pool decision and never enters the search; in particular it spends no draw\. Each pool decision has counts of positives, negatives, and negatives of each kind\. A freshrandom\.Random\(42\)is started, and the following search runs for a fixed 300000 iterations, keeping the best assignment found\. In each iteration: 1. 1\.withnamesthe pool decisions sorted ascending as strings, onerandrange\(3\)is drawn per name in that order, giving each decision a split index \(0 = test, 1 = train, 2 = validation\) — exactly one draw is spent per name, sonamesis precisely the pool decisions, no dataset\-wide decisions that lack a record here, or the whole random stream shifts; 2. 2\.each forbidden\-from\-test decision \(taken in sorted order\) currently assigned to test is reassigned withchoice\(\[1, 2\]\); 3. 3\.the per\-split totals are tallied — positives, negatives, and per\-kind negatives; 4. 4\.the assignment is feasible exactly when, for every split, positives≥\\geqits positive target, negatives≥\\geqits negative target, and each negative kind’s count≥\\geqthat split’s kind target; 5. 5\.its score is the minimum, over all splits, of\(positives−\-positive target\)and\(negatives−\-negative target\)— the minimum slack\. Among feasible assignments, the one with the largest score is kept, ties going to the one found earlier in the iteration\. The result maps each split to its sorted list of decision names\. \(The search is robust to the order in which forbidden decisions are visited; the sorted order fixes the procedure concretely\.\) Sampling\.A single freshrandom\.Random\(42\)is then started and threaded through the splits in the order test, train, validation\. For each split: the split’s positive target is taken, via the deterministic sampler below, from the positive records whose decision is in the split; then, for each negative kind in ascending alphabetical order of the kind name \(kein\_bestimmter\_inhalt, thenkeine\_konkrete\_gesetzesbestimmung\), that kind’s target is taken from the negative records of that kind whose decision is in the split, using the same sampler and the same generator; and the positives and the two kinds’ negatives are concatenated and then shuffled with the same generator, that shuffled order being the file order\. The deterministic samplertake\(pool, n\), using the current generator, sortspoolby the pair\(record\.row\.decision\_name, canonical JSON of record\.row\)— where "canonical JSON" is the JSON serialization of therowobject with keys sorted, equivalent to Python’sjson\.dumps\(row, sort\_keys=True\)— then shuffles the sorted list with the generator and takes the firstn\. Because the single generator is threaded through test, then train, then validation, the sequence ofshufflecalls per split is positives, kind\-1 negatives, kind\-2 negatives, whole\-split shuffle, repeated for the three splits in order\. #### E\.7\.2konkretes\_gesetz: numbering Within each split file, records are numberedid = 1, 2, 3, …in the file order fixed by the final per\-split shuffle above\. #### E\.7\.3The argument family \(argumentand the four canons\): split and sample Each of the five subtasks is split and sampled independently, by the same procedure askonkretes\_gesetz\(the same random\-restart search and the same sampler\), minus the kind machinery\. Split search\.For each pool decision \(as in[SectionE\.7\.1](https://arxiv.org/html/2609.26945#A5.SS7.SSS1): only decisions that own a positive or negative record for this subtask, never the full dataset\), its positives and negatives for this subtask are counted, and forargumentits individually judged negatives as well\. The random\-restart search of[SectionE\.7\.1](https://arxiv.org/html/2609.26945#A5.SS7.SSS1)runs identically — a freshrandom\.Random\(42\), 300000 iterations,namesthe pool decisions sorted ascending, onerandrange\(3\)per name, the same forbidden\-from\-test handling, maximizing the minimum slack, earliest tie kept — with the argument\-family targets, and with one addition for theargumentsubtask only: an assignment is feasible only when, in every split, the number of individually judged negatives assigned to that split does not exceed the split’s negative target\. \(This forces every individually judged negative to fit, so all of them enter the benchmark\.\) The canon subtasks add no such constraint\. Sampling\.A single freshrandom\.Random\(42\)is started, and for each split in the order test, train, validation: the split’s positive target is taken, via the deterministic sampler, from this subtask’s positive records whose decision is in the split; the split’s negative target is taken from this subtask’s negative records whose decision is in the split — forargumentvia the priority sampler below, for the four canons via the plain sampler; and the positives and negatives are concatenated and shuffled with the same generator, giving the file order\. The priority sampler\(forargumentnegatives\) sorts the pool by\(decision\_name, canonical JSON of row\)as before, separates it into the individually judged records and the rest, shuffles the individually judged group with the generator, shuffles the rest with the generator, places the individually judged group first followed by the rest, and takes the firstn— so the twoshufflecalls occur in that order, then the take\. Reasoning\.When a pair is selected into a subtask,observed\.reasoningis added: forargumentit isgeneral\_argument\.reasoning\(or""\), and for a canoncit is that canon object’sreasoning\(or""\)\. Numbering\.Within each split file, records are numberedid = 1, 2, 3, …in the file order fixed by the final per\-split shuffle\. #### E\.7\.4The reading\-criteria family \(nicht\_abstrakt,nicht\_selbst\_aufgestellt\): a different split search, sampler, and numbering These two subtasks use an exhaustive decision\-disjoint split search \(not the random restart\), a different sampler, and — importantly — aglobalidnumbering, each subtask handled independently\. Split search \(exhaustive, deterministic\)\.For each decision,\(p, g\)are its counts of this subtask’s positives and negatives\. The decisions are ordered by\(−\-p, decision\_name\)— most positives first, ties by name ascending\. All assignments of decisions to the three splits are searched \(a decision in the forbidden\-from\-test set may go only to train or validation\), pruning a branch as soon as the positives still available among the not\-yet\-assigned decisions can no longer cover the remaining positive deficit across splits\. Among all feasible assignments \(every split meeting both its positive and negative target\), the one chosen maximizes, lexicographically, the pair*\(minimum positive slack across splits, minimum negative slack across splits\)*, ties going to the one reached earlier in this ordered enumeration\. The targets are read in split order test, train, validation\. Sampling\.A single freshrandom\.Random\(42\)is started, and for each split in the order test, train, validation: the split’s positive pool is all this subtask’s positive records whose decision is in the split, sorted by the record’s region id \(deutung\_id\) ascending, and likewise the negative pool sorted by region id; the split’s positive target is drawn from the sorted positive pool withsample, then its negative target from the sorted negative pool withsample\(the same generator, positives first\); and the sampled positives and negatives are concatenated and shuffled with the same generator, giving the split’s internal order\. Global numbering\(the distinctive part\)\. After all three splits are sampled, shuffled, and turned into records, the three splits’ record lists are concatenated in the order test, train, validation into one list; that combined list is shuffled with the same generator; and the combined list is numberedid = 1, 2, …up to the total \(each of these two subtasks has 400 records: 100 positives and 300 negatives\)\. Each record keeps the number it receives here as itsidbut stays physically in its own split file, in the per\-split order from the previous step — so within a reading\-criteria file theids are not1\.\.Nin order but the scattered global numbers\. \(Only these two subtasks number globally; the argument family andkonkretes\_gesetznumber1\.\.Nper file\.\) Concretely, the generator’s call sequence for one reading\-criteria subtask issample\(test positives\),sample\(test negatives\),shuffle\(test\),sample\(train positives\),sample\(train negatives\),shuffle\(train\),sample\(validation positives\),sample\(validation negatives\),shuffle\(validation\),shuffle\(combined\)\. ### E\.8The finished benchmark The benchmark is the 24 files<subtask\>/\{train,validation,test\}\.json, each a JSON array of the\{ id, row, observed \}records built above, in the file order and with theidnumbering fixed above\. The eight subtask directory names are exactlyargument,wortlaut,systematik,geschichte,zweck,konkretes\_gesetz,nicht\_abstrakt, andnicht\_selbst\_aufgestellt\.
相似文章
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
This article introduces Magis-Bench, a benchmark for evaluating large language models on magistrate-level legal tasks such as judicial reasoning and sentence drafting, using data from Brazilian judicial exams.
HKJudge:一个用于解读法院认定、推理和裁决的法律话语标注语料库
HKJudge是首个针对香港刑事判决进行句子级专家标注的法律话语语料库,包含两层话语标注体系以及基于BERT和LLM模型的基准评估。
通过检索、聚类和生成从案例数据库生成法律评注
本文提出了一种完全自动化的流程,通过提取、聚类和总结段落级块(使用LLM),将法院判决转化为法律评注,并在德国民法典案例上进行了评估。
大语言模型能否进行具有法律意义的推理?基于欧洲人权法院案例的小规模研究
本小规模研究通过欧洲人权法院案例评估了大语言模型在法律案件预测中的推理能力,发现模型能够生成结构完整但实质浅薄的分析,且基于大语言模型的评估器与人类标注者的一致性较低。
LegalBench-BR:评估大语言模型在巴西法律判决分类上的基准
研究者发布首个公开基准 LegalBench-BR,用于评估大模型在巴西法律文本分类任务上的表现。实验表明,LoRA 微调的 BERTimbau 大幅超越 GPT-4o mini 与 Claude 3.5 Haiku。