Schematize: An Agentic System for Generating and Refining Information-Extraction Schemas for Legal Research
Summary
Schematize is an open-source multi-agent system that interactively generates and refines information-extraction schemas for legal research, achieving top performance in human evaluations.
View Cached Full Text
Cached at: 09/22/26, 09:08 AM
# An Agentic System for Generating and RefiningInformation-Extraction Schemas for Legal Research
Source: [https://arxiv.org/html/2609.22209](https://arxiv.org/html/2609.22209)
## Schematize: An Agentic System for Generating and Refining Information\-Extraction Schemas for Legal Research
Albert Sawczyn†Jakub BinkowskiAffiliation:Wrocław University of Science and TechnologyKamil TagowskiAffiliation:Wrocław University of Science and TechnologyŁukasz AugustyniakAffiliation:Wrocław University of Science and TechnologyBerenika Kaczmarek\-TemplinAffiliation:Wrocław University of Science and TechnologyTomasz KajdanowiczAffiliation:Wrocław University of Science and Technology
###### Abstract
Empirical legal research often relies on turning research questions into structured data extracted from large collections of rulings and judgments\. Designing the extraction schema and then extracting the data remain a manual, expertise\-heavy bottleneck\. We presentschematize, an open\-source multi\-agent system that interactively turns a researcher’s problem statement into a validated extraction schema that can later be used for autonomous extraction\.Schematizecouples \(i\) a clarification dialogue that elicits implicit expert intent, \(ii\) iterative schema generation, \(iii\) data\-grounded refinement that tests the schema against documents, and \(iv\) chat\-based post\-editing\. We evaluated the system with human legal professionals, introducing our novel methodology, andschematizeachieves top performance in most of tested configurations\. While the system is designed to be domain\-agnostic and applicable to any document collection, we tailor and evaluate it on legal research problems\. We releaseschematizeas a pip\-installable Python package with full documentation\.
## 1Introduction
Empirical legal research often depends on structured representations of case law, rulings, and judgments: to study a question quantitatively, researchers must first extract fields that can be systematically analyzed across hundreds or thousands of documents\([Hwang et al\., 2022](https://arxiv.org/html/2609.22209#bib.bib11);[Mali et al\., 2024](https://arxiv.org/html/2609.22209#bib.bib9)\)\. Production legal\-analytics platforms likewise combine large\-scale document retrieval, expert\-defined criteria, and structured extraction to support quantitative studies\([Augustyniak et al\., 2026](https://arxiv.org/html/2609.22209#bib.bib22)\)\. The quality of the outcomes depends strongly on how researchers formulate the research problem and then design the extraction schema that specifies*what*to pull from each document\. This upstream formulation is itself difficult: an initial request often leaves scope, entities, exclusions, comparisons, and downstream analytical goals implicit, even when the researcher has a clear study in mind\. The process demands both legal expertise and data\-modeling skills, from defining atomic fields to judging which values an LLM can reliably extract, while being slow, iterative, and prone to bias\. For instance, a schema that omits a relevant field silently caps what the downstream study can ever discover\.
Modern information\-extraction systems share a common interface: they take a user\-supplied*schema*that declares what to pull from each document and ask a model to fill it in\([Zaratiana et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib6);[Shrimal et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib7);[Goel, 2025](https://arxiv.org/html/2609.22209#bib.bib8)\), and the same assumption underlies legal IE pipelines\([C R et al\., 2024](https://arxiv.org/html/2609.22209#bib.bib10);[Mali et al\., 2024](https://arxiv.org/html/2609.22209#bib.bib9)\)\. A growing line of work instead tries to*derive*schemas from a research question and a corpus\([Levy et al\., 2026](https://arxiv.org/html/2609.22209#bib.bib1);[Sadruddin et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib2);[Padmakumar et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib3)\), but it typically validates against a reference schema and offers little support for grounding the schema in the documents it will be applied to\. The upstream step – turning a research question into a good schema and checking it against real data – is thus left to the analyst, with no agreed way to tell a good schema from a bad one \(Section[2](https://arxiv.org/html/2609.22209#S2)\)\.
Schematizecloses this gap with a human\-in\-the\-loop, multi\-agent pipeline \(Figure[1](https://arxiv.org/html/2609.22209#S1.F1)\)\. A researcher states a problem in natural language; a*problem\-definition\-helper*agent asks clarifying questions, and once the expert answers, a*definer*agent turns the request and responses into a formal problem definition\. The system then generates search queries, drafts an initial schema, and improves it through two refinement loops: a criteria\-based refinement loop in which a*critic*agent scores the schema against quality criteria, and a data\-grounded refinement loop that retrieves real documents with the generated queries and tests whether the schema can actually be populated from them\. Finally, a summary explains how the schema was derived, and an interactive chat lets the expert request further edits\. The result is an extraction schema that is grounded in both expert intent and real documents before any large\-scale extraction begins\. Notably, we provide a novel methodology for the evaluation of schema quality by asking human experts to write exhaustive question sets for each research problem, which is specifically designed to limit the impact of biases in the evaluation process \([Section4](https://arxiv.org/html/2609.22209#S4)\), and then scoring the schema for coverage by an LLM\-as\-judge\([Zheng et al\., 2023](https://arxiv.org/html/2609.22209#bib.bib19)\)\.
- •System & library:an open\-source, pip\-installable library for research\-problem\-to\-schema generation that is model\-agnostic – any agent runs on either open\- or closed\-weight models – with pluggable data connectors and a schema\-based extractor\.
- •Evaluation methodology & data:an expert\-question coverage protocol – experts write exhaustive question sets for each research problem \(∼\\sim15–30\), and schemas are scored for coverage by an LLM\-as\-judge – together with expert\-curated legal research problems and their evaluation\-questions sets\.
- •Empirical study:a comprehensive evaluation on legal research problems, comparing small and large LLMs, single\- and multi\-agent configurations, system ablations, and cost–quality trade\-offs\.
#### Availability\.
UserRequestProblemDefinitionQueryGenerationSchemaGenerationSchemaRefinementData\-GroundedRefinementSummaryFinalSchemaDocument DatabaseHuggingFace/Weaviatequeriesdocs𝑚𝑖𝑛ref≤𝑚𝑎𝑥ref\\mathit\{min\}\_\{ref\}\\leq\\mathit\{max\}\_\{ref\}𝑚𝑖𝑛dref≤𝑚𝑎𝑥dref\\mathit\{min\}\_\{dref\}\\leq\\mathit\{max\}\_\{dref\}
Figure 1:Overview ofschematize’s agentic pipeline\. A user request is clarified, formalized into retrieval queries and an initial schema, refined via criteria\-based \(CBR, up to𝑚𝑎𝑥ref\\mathit\{max\}\_\{ref\}iterations\) and data\-grounded \(DGR, up to𝑚𝑎𝑥dref\\mathit\{max\}\_\{dref\}iterations\) loops, and finalized through a summary and interactive chat\.Blue\( \) marks automated agent steps,orange\( \) marks human\-in\-the\-loop steps,violetmarks the document store, andtealmarks refinement loops\.
## 2Related Work
Information extraction \(IE\) turns unstructured documents into structured records\([Jurafsky and Martin, 2026](https://arxiv.org/html/2609.22209#bib.bib13)\)\. As large language models have made extraction increasingly reliable\([Bai et al\., 2024](https://arxiv.org/html/2609.22209#bib.bib14)\), the field has settled on a common interface: rather than hand\-building an extractor for each task, modern systems take a user\-supplied*schema*that declares what to pull from each document and ask the model to fill it in\([Zaratiana et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib6)\)\.
#### Schema\-driven information extraction
treats this schema as the contract between user and model\. For instance, GLiNER2 unifies entity recognition, classification, and structured extraction behind a schema\-based API\([Zaratiana et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib6)\); PARSE optimizes existing JSON schemas for reliable LLM extraction\([Shrimal et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib7)\); LangExtract grounds each extracted field in its source span\([Goel, 2025](https://arxiv.org/html/2609.22209#bib.bib8)\)\. The same assumption carries into legal IE, where systems extract entities and relations against task\-specific schemas to support quantitative case\-law analysis\([C R et al\., 2024](https://arxiv.org/html/2609.22209#bib.bib10);[Hwang et al\., 2022](https://arxiv.org/html/2609.22209#bib.bib11);[Mali et al\., 2024](https://arxiv.org/html/2609.22209#bib.bib9)\)\.
#### Deriving schemas, rather than assuming them,
is the concern of the work closest to ours\.ScheMatiQ\([Levy et al\., 2026](https://arxiv.org/html/2609.22209#bib.bib1)\)turns a research question and a corpus into a query\-driven schema and creates a grounded database through interactive discovery; SchemaMiner builds schemas from scientific documents with human\-in\-the\-loop refinement\([Sadruddin et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib2)\); intent\-conditioned methods generate and refine literature\-review table schemas\([Padmakumar et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib3)\); SciDaSynth assembles structured tables from multi\-modal scientific sources\([Wang et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib4)\); and LOGOS induces hierarchical schemas for qualitative grounded\-theory analysis\([Pi et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib5)\)\.
Our work differs from prior work along three axes\. First, rather than assuming a corpus,schematizegenerates retrieval queries from the problem statement and pulls documents through pluggable connectors, so schema design is decoupled from having a pre\-assembled collection\. Second, we separate a document\-blind, criteria\-based critique – guarding against overfitting to any particular sample – from a data\-grounded refinement loop, whereas the systems above either skip document grounding or interleave it directly with generation\. Third, and most importantly, we evaluate schemas against expert\-authored*question sets*rather than treating reference schemas as gold: closely related work often evaluates by reconstructing manual or table\-derived schemas, even though such references can be ambiguous, shaped by feasibility constraints, or contain artifacts\([Levy et al\., 2026](https://arxiv.org/html/2609.22209#bib.bib1);[Padmakumar et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib3)\)\. A question\-coverage protocol is repeatable across system versions without re\-annotation\. Moreover, a good, robust schema for a given research problem can be constructed in multiple valid ways – field names, granularity, and typing choices vary even among experts solving the same problem – so schema design has no single ground\-truth target; this is precisely why we evaluate against expert\-authored questions rather than a single reference schema\.
#### Empirical legal research
is our main focus, where automation must augment rather than replace expert judgment\([Ramesh et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib12)\)\. Prior legal IE work delivers strong extractors but, as above, leaves schema design to the analyst\([Mali et al\., 2024](https://arxiv.org/html/2609.22209#bib.bib9);[Hwang et al\., 2022](https://arxiv.org/html/2609.22209#bib.bib11)\)\.Schematizefills this gap with a human\-in\-the\-loop pipeline that helps legal researchers express, ground obtained schema before any extraction begins\.
## 3TheschematizeSystem
Schematizeis organized as a pipeline of LLM agents \(Figure[1](https://arxiv.org/html/2609.22209#S1.F1)\)\. Each agent has a single responsibility and communicates through structured messages\. Users can configure the pipeline by choosing LLMs \(either closed\-weight API models or open\-weight models\), data connectors, and other parameters described in the library documentation\. We describe each pipeline stage below\. Notably, we tailored the prompts with the help of legal experts to be more specific to the legal domain\.
#### Problem Definition
First, the legal researcher sketches a problem of interest \(e\.g\., “study personal\-rights violations and assess their severity”\)\. A*problem\-definition\-helper*\(PDH\) agent reads this request and asks a short batch of clarifying questions – about scope, jurisdiction, the granularity of the target variables, and intended downstream analysis – surfacing implicit assumptions an expert could leave unstated\. The expert answers, and a*problem\-definer*agent compiles the request and prepares a formal problem definition, including a problem statement, the legal domain, the scope of judgments of interest, the legal concepts involved, and a set of research questions\.
#### Query Generation
From the formal problem definition, a*query\-generator*agent produces search query to find the documents the schema will eventually be applied to\. These queries feed the document retriever in the data\-grounded loop \(Section[3](https://arxiv.org/html/2609.22209#S3.SS0.SSS0.Px4)\)\.
#### Schema Generation and Criteria\-Based Refinement
A*generator*agent first drafts an extraction schema from the problem definition; Listingshows an excerpt\. The schema then enters a criteria\-based refinement \(CBR\) loop: an*assessment*agent scores it for coverage of the problem definition, field clarity and non\-redundancy, appropriate typing and granularity, and extractability, and a*refiner*agent applies the recommended changes\. To avoid overfitting the schema to the retrieved documents, the*generator*and*refiner*agents operate without document access and rely only on the problem\-definition instructions\. The loop repeats until it runs at least𝑚𝑖𝑛ref\\mathit\{min\}\_\{ref\}rounds and stops once the assessment reports no further changes are needed or a cap of𝑚𝑎𝑥ref\\mathit\{max\}\_\{ref\}rounds is reached\.
#### Data\-Grounded Assessment and Refinement
The second loop grounds the schema in the context of real documents\. Using the generated queries, a document retriever fetches a set of matching legal documents through a configured backend, such as Hugging Face datasets or a Weaviate vector store\. This data\-grounded refinement \(DGR\) loop starts when a*data\-assessment*agent examines each retrieved document to assess the schema from a practical perspective, and a*merger*agent consolidates these per\-document assessments into a single set of revisions\. Finally, a*data\-refiner*agent updates the schema accordingly\. Like the first loop, it runs between𝑚𝑖𝑛dref\\mathit\{min\}\_\{dref\}and𝑚𝑎𝑥dref\\mathit\{max\}\_\{dref\}rounds, with the*data\-assessment*agent examining a configurable number of retrieved documents per round\.
#### Summarization and Interactive Chat
After the refinement loops, a*summarizer*agent produces a short report explaining how the schema was derived – which fields were added or changed in each loop and why – so the expert can audit the process rather than trust an opaque output\. The session then enters an open\-ended chat: the expert can ask for further changes in natural language, and a chat agent edits the schema in place, keeping the human in control of the final artifact\.
### 3\.1Library Design
Schematizeships as an installable Python package so the pipeline can be run programmatically or embedded in other tools\. LLM access goes through LiteLLM\([Dholakia and Jaffer, 2023](https://arxiv.org/html/2609.22209#bib.bib17)\), and orchestrated through the langgraph framework\([LangChain Inc\., 2024](https://arxiv.org/html/2609.22209#bib.bib18)\), making the library modular and model\-agnostic: agents can be backed by either open\-weight models \(e\.g\., self\-hosted Llama or Qwen run locally or behind a vLLM[Kwon et al\. \(2023\)](https://arxiv.org/html/2609.22209#bib.bib15)/Ollama[Ollama \(2023\)](https://arxiv.org/html/2609.22209#bib.bib16)endpoint\) or closed\-weight API models \(e\.g\., GPT, Claude, or Gemini\)\. This lets users trade costs, latency, and data\-privacy requirements against capability\. Data access is abstracted through an abstract retriever base class, which ensures that the DGR loop receives documents in a uniform format regardless of their source\. The library includes connectors for Hugging Face datasets and Weaviate, and supports other document sources through extensions provided by users\.
## 4Empirical Evaluation
We introduce a custom methodology for assessing the quality of generated schemas while limiting potential annotation biases\. In particular, we evaluate the system on three Polish legal research problems in collaboration with seven legal experts\.
### 4\.1Methodology
In prior studies, experts were asked to assess schemas generated by the evaluated systems\. However, such holistic and intuitive evaluation can introduce confirmation and hindsight biases among annotators\([Mahdavi and Rahimian, 2017](https://arxiv.org/html/2609.22209#bib.bib23)\)\. Instead, we propose a protocol wherein domain experts cannot see the generated schema but independently write the atomic questions that the extracted data should be able to answer \(see[Figure5](https://arxiv.org/html/2609.22209#A2.F5)\)\. The experts wrote questions until they judged the topic to be exhausted; we suggested approximately 15–30 questions per problem\. By asking experts to formulate questions upfront, we shift their cognitive processing from post\-hoc schema evaluation to direct reasoning about the problem, mitigating potential biases\. For each research problem, experts wrote their questions from a single fixed PDH exchange: a user request, PDH clarifying questions from gpt\-5\.2, and the user’s answers \([Figure5](https://arxiv.org/html/2609.22209#A2.F5)\)\. In our study, we recruited 7 experts from the District Chamber of Legal Advisers in Wrocław and asked them to annotate three problems\. We randomly assigned 3 experts to each problem\. To ensure that experts understood the annotation task and were aligned on the protocol, we prepared detailed instructions \(see[AppendixB](https://arxiv.org/html/2609.22209#A2)\) and conducted a workshop session to explain the details and answer questions\.
### 4\.2Model configuration and baselines
We compareschematizewith two baselines: a*vanilla LLM*, which generates a schema in a single call, and the multi\-agent*ScheMatiQ*system\([Levy et al\., 2026](https://arxiv.org/html/2609.22209#bib.bib1)\)\. Becauseschematizeincludes the PDH stage, it receives information elicited from the user beyond the initial problem sketch\. To isolate the contribution of the pipeline rather than this additional input, we evaluate each baseline both without and with the same fixed PDH exchange \([Figure5](https://arxiv.org/html/2609.22209#A2.F5)\)\. This yields four baseline configurations:*vanilla LLM*,*vanilla LLM \+ PDH*,*ScheMatiQ*, and*ScheMatiQ \+ PDH*; the PDH\-augmented configurations generate schemas from the clarified problem\. Our study uses both closed\-weight models \(gpt\-5\.4, gpt\-5\.4\-mini, gpt\-5\.4\-nano, claude\-sonnet\-4\.6\) and open\-weight models \(qwen3\.6\-35b\-a3b, gemma\-4\-e4b\-it, llama\-4\-scout\-17b\) of varying sizes555Licenses \(all permit research use\): gpt\-5\.4/\-mini/\-nano: OpenAI Services Agreement; claude\-sonnet\-4\.6: Anthropic Commercial Terms; qwen3\.6\-35b\-a3b: Apache 2\.0; gemma\-4\-e4b\-it: Apache 2\.0; llama\-4\-scout\-17b: Llama 4 Community License\.\. In[Section4\.4](https://arxiv.org/html/2609.22209#S4.SS4), we report coverage results for the best hyperparameter configuration found, and provide a broader hyperparameter sensitivity study in[Section5\.1](https://arxiv.org/html/2609.22209#S5.SS1)along with component ablations in[Section5\.2](https://arxiv.org/html/2609.22209#S5.SS2)\.
Backbonevanilla LLMvanilla LLM \+ PDHScheMatiQScheMatiQ \+ PDHschematizegpt\-5\.436\.3%±\\pm11\.7pp64\.3%±\\pm19\.7pp43\.8%±\\pm15\.6pp64\.3%±\\pm17\.6pp79\.4%±\\pm12\.6ppgpt\-5\.4\-mini31\.7%±\\pm7\.2pp55\.5%±\\pm14\.9pp23\.9%±\\pm7\.9pp57\.6%±\\pm7\.6pp74\.8%±\\pm7\.2ppgpt\-5\.4\-nano22\.2%±\\pm18\.0pp59\.8%±\\pm10\.3pp24\.0%±\\pm13\.8pp42\.5%±\\pm20\.4pp34\.7%±\\pm35\.0ppclaude\-sonnet\-4\.637\.1%±\\pm10\.8pp60\.5%±\\pm12\.9pp49\.6%±\\pm29\.2pp67\.1%±\\pm14\.8pp70\.0%±\\pm18\.5ppqwen3\.6\-35b\-a3b28\.5%±\\pm8\.7pp51\.6%±\\pm2\.5pp40\.4%±\\pm15\.6pp67\.2%±\\pm15\.3pp66\.5%±\\pm16\.3ppgemma\-4\-e4b\-it3\.8%±\\pm4\.0pp38\.2%±\\pm14\.8pp34\.5%±\\pm18\.2pp46\.5%±\\pm4\.8pp53\.4%±\\pm3\.7ppllama\-4\-scout\-17b27\.7%±\\pm20\.0pp42\.0%±\\pm8\.8pp32\.7%±\\pm19\.3pp45\.3%±\\pm10\.1pp45\.1%±\\pm7\.9pp
Table 1:Expert\-question coverage \(mean±\\pmstd, in pp\) across backbones and configurations\. “\+ PDH” prepends the Problem\-Definition\-Helper stage to a baseline\. Bold marks the best, underline the second\-best, coverage per backbone\.
### 4\.3Metrics
We reportcoverage: the fraction of expert questions that the generated schema can answer, judged per question by an LLM\-as\-judge \(gpt\-5\.4\-mini\) that checks whether the schema contains the fields needed to resolve each question\. LLM\-as\-a\-judge evaluation is widely used and has demonstrated substantial alignment with human judgments in prior evaluation settings\([Zheng et al\., 2023](https://arxiv.org/html/2609.22209#bib.bib19);[Thakur et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib20);[Janiak et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib21)\)\. Thus, coverage provides a useful automatic proxy for schema quality\.
### 4\.4Results
Table[1](https://arxiv.org/html/2609.22209#S4.T1)reports coverage across model sizes, with a complementary visualization in[Figure2](https://arxiv.org/html/2609.22209#S4.F2), comparingschematizewith vanilla LLM and ScheMatiQ baselines\.Schematizeconsistently outperforms both the vanilla LLM and ScheMatiQ baselines, as well as their PDH\-augmented variants, across all tested backbones: it obtains the best or second\-best result in 6 of the 7 tested configurations, including the best result in 4 configurations\. This pattern holds for both open\- and closed\-weight models, suggesting that the proposed method remains effective in settings where privacy or deployment constraints favor open\-weight models\. Although larger models generally improve coverage,schematizesurpasses the average inter\-expert coverage \([Table2](https://arxiv.org/html/2609.22209#A2.T2)\) on 4 of the 7 tested backbones\. In other words,schematizeanticipates more of an expert’s questions than the other experts do themselves, on average\. On qwen3\.6\-35b\-a3b and llama\-4\-scout\-17b, ScheMatiQ \+ PDH edges outschematizeby only∼\\sim1pp, and that margin comes from PDH – our own contribution – not from ScheMatiQ itself\.
Figure 2:Comparison of theschematizeto all tested baselines\. The dashed line marks the average inter\-expert coverage \([Table2](https://arxiv.org/html/2609.22209#A2.T2)\)\.In addition,[Figure3](https://arxiv.org/html/2609.22209#S4.F3)shows that subsequent schema refinement iterations improve the coverage of the generated schema\. However, these gains plateau after several iterations, suggesting that comparable coverage can be achieved without many expensive refinement steps\. As[Figure3](https://arxiv.org/html/2609.22209#S4.F3)also shows, small models like gpt\-5\.4\-nano struggle to grasp that their step is part of a larger pipeline: coverage drops sharply once the schema enters the data\-grounded refinement \(DGR\) loop\. On our inspection, the model appears to overfit to whichever document it is currently looking at, dropping previously established fields that are not reflected in that document instead of merging the new evidence with the existing schema\. Further, we present detailed results in[AppendixC](https://arxiv.org/html/2609.22209#A3), inter\-expert coverage in[Table2](https://arxiv.org/html/2609.22209#A2.T2), and inference costs in[Table5](https://arxiv.org/html/2609.22209#A4.T5)\.
Figure 3:Coverage change over subsequent iterations of schema refinements for three considered problems\.
## 5Analysis
### 5\.1Hyperparameter Sensitivity
Our system is parametrized by the minimum and maximum number of CBR rounds,𝑚𝑖𝑛ref\\mathit\{min\}\_\{ref\}and𝑚𝑎𝑥ref\\mathit\{max\}\_\{ref\}, the minimum and maximum number of DGR rounds,𝑚𝑖𝑛dref\\mathit\{min\}\_\{dref\}and𝑚𝑎𝑥dref\\mathit\{max\}\_\{dref\}, and the number of documents the*data\-assessment*agent examines per DGR round\. To select the best hyperparameters, we run a grid search with gpt\-5\.4\-nano and choose the configuration with the highest coverage\.
Figure 4:Results of the hyperparameter sensitivity study\. We measure coverage of the tested configurations for three legal research problems\.We report coverage for the tested configurations in[Figure4](https://arxiv.org/html/2609.22209#S5.F4)\. Coverage remains stable across several hyperparameter settings, suggesting thatschematizeis robust to these choices\. This stability also indicates that users can often reduce costs by skipping hyperparameter search, especially when evaluation data is unavailable and the goal is schema generation\. The best configuration, used for the main results in[Table1](https://arxiv.org/html/2609.22209#S4.T1), sets𝑚𝑖𝑛ref=2\\mathit\{min\}\_\{ref\}=2,𝑚𝑎𝑥ref=3\\mathit\{max\}\_\{ref\}=3,𝑚𝑖𝑛dref=2\\mathit\{min\}\_\{dref\}=2, and𝑚𝑎𝑥dref=4\\mathit\{max\}\_\{dref\}=4, examining a single document per DGR round\.
### 5\.2Ablation Study
We assess the contribution of each pipeline stage by selectively ablating the problem\-definition\-helper \(PDH\), the criteria\-based refinement \(CBR\) loop, and the data\-grounded refinement \(DGR\) loop, individually and in combination; removing all three collapsesschematizeinto the single\-LLM baseline \(the vanilla LLM setting in[Table1](https://arxiv.org/html/2609.22209#S4.T1)\)\.[Table4](https://arxiv.org/html/2609.22209#A3.T4)shows that removing CBR or DGR components has only a modest effect, well within noise, whereas removing PDH substantially decreases the performance\. Likewise, removing two components together consistently leads to degradation, which indicates that they are complementary andschematizepeaks when PDH, CBR, and DGR operate together\. The full system’s larger variance reflects differences in problem difficulty rather than component instability\. Ablation runs are separate from the main results, so the “Full \(schematize\)” row may differ slightly due to LLM non\-determinism\.
## 6Conclusion
We contributedschematize, a human\-in\-the\-loop multi\-agent system that turns a research problem into a document\-grounded extraction schema through clarification, criteria\-based critique, and data\-grounded refinement against retrieved documents\. We rigorously tested the system by introducing a novel methodology to evaluate the schema quality with help of human legal professionals\. Empirical legal scholars and legal\-tech builders can useschematizeto generate extraction schemas from research problems without hand\-crafting them first\. The code and pip\-installable package are publicly available\.
## 7Limitations
Our evaluation is based on a single jurisdiction and language, and performance may vary for others\. We focused on Polish case law because we were able to recruit legal experts with relevant expertise in this jurisdiction\.
Coverage is scored by an LLM\-as\-judge, so the final estimate depends on the judge model and prompt\. To limit biases that can arise when assesing a schema post\-generation, experts authored question sets defining the target information needs prior to generation\. However, we did not compare the LLM judge’s individual coverage decisions with human annotations\.
Additionally, coverage measures schema only recall quality: whether the generated fields cover expert information needs\. While extra fields may hinder extraction of the fields of interest, they may also provide broader perspective and have only negligible effects on extraction quality\. Also, our evaluation measures schema coverage rather than end\-to\-end extraction accuracy\.
## 8Ethics Statement
Schematizetargets the legal domain, a high\-stakes setting where automation must augment rather than replace expert judgment\([Ramesh et al\., 2025](https://arxiv.org/html/2609.22209#bib.bib12)\)\. It is a research tool for designing extraction schemas and producing structured data for analysis; it does not provide legal advice, and its outputs should not be relied on for legal decisions\. Because automatically generated schemas can look authoritative, we keep the human in the loop at both problem definition and final editing, and the summarizer makes the derivation auditable to discourage over\-reliance\.
The study with human experts was carried out in collaboration with the District Chamber of Legal Advisers in Wrocław, Poland, under a formal agreement between the Chamber and Wrocław University of Science and Technology\. The experts participated voluntarily and without compensation, with the exception of the study coordinator, who is an employee funded by the National Science Centre grant\.
## Acknowledgments
#### Annotators
We gratefully thank the legal experts from the District Chamber of Legal Advisers in Wrocław, Poland, who contributed the annotations: Monika Bocheńska, Dominika Fikus, Dorota Jarzębowska, Berenika Kaczmarek\-Templin, Małgorzata Kozłowska, Maciej Kusyk, and Mateusz Lato\. Their domain expertise was essential for constructing the evaluation protocol and expert question sets\.
#### Funding
This work was co\-funded by the National Science Centre, Poland, under CHIST\-ERA Open & Re\-usable Research Data & Software \(grant no\. 2022/04/Y/ST6/00183\)\. We gratefully acknowledge the Wroclaw Center for Networking and Supercomputing for providing computing facilities and support\.
## References
- Augustyniaket al\.\(2026\)Ł\. Augustyniak, K\. Tagowski, A\. Szymczak, J\. Binkowski, A\. Sawczyn, M\. Skibiński, D\. Janiak, M\. Bystroński, G\. Piotrowski, M\. Bernaczyk, K\. Kamiński, and T\. J\. KajdanowiczBridging AI and law: a scalable multi\-agent platform for quantitative legal analytics across millions of documents\.InBridge between Artificial Intelligence and Law \(AILaw\),pp\. 207–214\.External Links:[Link](https://openreview.net/forum?id=hWjsyTSWrY)Cited by:[§1](https://arxiv.org/html/2609.22209#S1.p1.1)\.
- Baiet al\.\(2024\)F\. Bai, J\. Kang, G\. Stanovsky, D\. Freitag, M\. Dredze, and A\. RitterSchema\-driven information extraction from heterogeneous tables\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 10252–10273\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.600)Cited by:[§2](https://arxiv.org/html/2609.22209#S2.p1.1)\.
- C Ret al\.\(2024\)C\. C R, S\. Kulkarni, S\. R\. A\. V\. Sagi, S\. Pandey, R\. Yalavarthy, D\. Chakraborty, and P\. UpadhyayLeGen: complex information extraction from legal sentences using generative models\.InProceedings of the Natural Legal Language Processing Workshop 2024,Miami, FL, USA,pp\. 1–17\.External Links:[Link](https://aclanthology.org/2024.nllp-1.1/),[Document](https://dx.doi.org/10.18653/v1/2024.nllp-1.1)Cited by:[§1](https://arxiv.org/html/2609.22209#S1.p2.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px1.p1.1)\.
- Dholakia and Jaffer \(2023\)K\. Dholakia and I\. JafferLiteLLM: a unified interface to call 100\+ llm apis\.GitHub\.Note:MIT LicenseExternal Links:[Link](https://github.com/BerriAI/litellm)Cited by:[§3\.1](https://arxiv.org/html/2609.22209#S3.SS1.p1.1)\.
- Goel \(2025\)A\. GoelLangExtract: LLM\-powered structured information extraction from text with source grounding\.Google LLC\.Note:[https://github\.com/google/langextract](https://github.com/google/langextract)Open\-source software, Apache\-2\.0 license\. DOI:[10\.5281/zenodo\.17015089](https://doi.org/10.5281/zenodo.17015089)External Links:[Document](https://dx.doi.org/10.5281/zenodo.17015089),[Link](https://github.com/google/langextract)Cited by:[§1](https://arxiv.org/html/2609.22209#S1.p2.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px1.p1.1)\.
- Hwanget al\.\(2022\)W\. Hwang, S\. Eom, H\. Lee, H\. J\. Park, and M\. SeoData\-efficient end\-to\-end information extraction for statistical legal analysis\.InProceedings of the Natural Legal Language Processing Workshop 2022,Abu Dhabi, United Arab Emirates \(Hybrid\),pp\. 143–152\.External Links:[Link](https://aclanthology.org/2022.nllp-1.12/),[Document](https://dx.doi.org/10.18653/v1/2022.nllp-1.12)Cited by:[§1](https://arxiv.org/html/2609.22209#S1.p1.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px3.p1.1)\.
- Janiaket al\.\(2025\)D\. Janiak, J\. Binkowski, A\. Sawczyn, B\. Gabrys, R\. Shwartz\-Ziv, and T\. J\. KajdanowiczThe illusion of progress: re\-evaluating hallucination detection in LLMs\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 34728–34745\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.1761/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1761),ISBN 979\-8\-89176\-332\-6Cited by:[§4\.3](https://arxiv.org/html/2609.22209#S4.SS3.p1.1)\.
- Jurafsky and Martin \(2026\)D\. Jurafsky and J\. H\. MartinSpeech and language processing: an introduction to natural language processing, computational linguistics, and speech recognition, with language models\.3rd edition\.Note:Online manuscript released January 6, 2026External Links:[Link](https://web.stanford.edu/~jurafsky/slp3/)Cited by:[§2](https://arxiv.org/html/2609.22209#S2.p1.1)\.
- Kwonet al\.\(2023\)W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. E\. Gonzalez, H\. Zhang, and I\. StoicaEfficient memory management for large language model serving with PagedAttention\.InProceedings of the 29th Symposium on Operating Systems Principles \(SOSP ’23\),pp\. 611–626\.External Links:[Document](https://dx.doi.org/10.1145/3600006.3613165)Cited by:[§3\.1](https://arxiv.org/html/2609.22209#S3.SS1.p1.1)\.
- LangChain Inc\. \(2024\)LangChain Inc\.LangGraph: low\-level orchestration framework for building stateful, multi\-actor applications with llms\.GitHub\.Note:[https://github\.com/langchain\-ai/langgraph](https://github.com/langchain-ai/langgraph)Accessed: 2026\-07\-07Cited by:[§3\.1](https://arxiv.org/html/2609.22209#S3.SS1.p1.1)\.
- Levyet al\.\(2026\)S\. Levy, E\. Habba, R\. Mintz, B\. Raveh, R\. Keydar, and G\. StanovskyScheMatiQ: from research question to structured data through interactive schema discovery\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),San Diego, California, United States,pp\. 220–230\.External Links:[Link](https://aclanthology.org/2026.acl-demo.22/)Cited by:[§1](https://arxiv.org/html/2609.22209#S1.p2.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px2.p2.1),[§4\.2](https://arxiv.org/html/2609.22209#S4.SS2.p1.1)\.
- Mahdavi and Rahimian \(2017\)S\. Mahdavi and M\. A\. RahimianHindsight bias impedes learning\.InProceedings of the NIPS 2016 Workshop on Imperfect Decision Makers,T\. V\. Guy, M\. Kárný, D\. Rios\-Insua, and D\. H\. Wolpert \(Eds\.\),Proceedings of Machine Learning Research, Vol\.58,pp\. 111–127\.External Links:[Link](https://proceedings.mlr.press/v58/mahdavi17a.html)Cited by:[§4\.1](https://arxiv.org/html/2609.22209#S4.SS1.p1.1)\.
- Maliet al\.\(2024\)D\. Mali, R\. Mali, and C\. BaraleInformation extraction for planning court cases\.InProceedings of the Natural Legal Language Processing Workshop 2024,Miami, FL, USA,pp\. 97–114\.External Links:[Link](https://aclanthology.org/2024.nllp-1.8/),[Document](https://dx.doi.org/10.18653/v1/2024.nllp-1.8)Cited by:[§1](https://arxiv.org/html/2609.22209#S1.p1.1),[§1](https://arxiv.org/html/2609.22209#S1.p2.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px3.p1.1)\.
- Ollama \(2023\)OllamaOllama: get up and running with large language models locally\.Note:[https://github\.com/ollama/ollama](https://github.com/ollama/ollama)Cited by:[§3\.1](https://arxiv.org/html/2609.22209#S3.SS1.p1.1)\.
- Padmakumaret al\.\(2025\)V\. Padmakumar, J\. C\. Chang, K\. Lo, D\. Downey, and A\. NaikIntent\-aware schema generation and refinement for literature review tables\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 23450–23472\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1274/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1274)Cited by:[§1](https://arxiv.org/html/2609.22209#S1.p2.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px2.p2.1)\.
- Piet al\.\(2025\)X\. Pi, Q\. Yang, and C\. NguyenLOGOS: LLM\-driven end\-to\-end grounded theory development and schema induction for qualitative research\.External Links:2509\.24294,[Link](https://arxiv.org/abs/2509.24294)Cited by:[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px2.p1.1)\.
- Rameshet al\.\(2025\)K\. Ramesh, D\. Smolyak, Z\. Zhao, N\. Gandhi, R\. Agarwal, M\. V\. Bjarnadóttir, and A\. FieldSynthTextEval: synthetic text data generation and evaluation for high\-stakes domains\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,Suzhou, China,pp\. 487–499\.External Links:[Link](https://aclanthology.org/2025.emnlp-demos.35/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-demos.35)Cited by:[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px3.p1.1),[§8](https://arxiv.org/html/2609.22209#S8.p1.1)\.
- Sadruddinet al\.\(2025\)S\. Sadruddin, J\. D’Souza, E\. Poupaki, A\. Watkins, H\. Babaei Giglou, A\. Rula, B\. Karasulu, S\. Auer, A\. Mackus, and E\. KesselsLLMs4SchemaDiscovery: a human\-in\-the\-loop workflow for scientific schema mining with large language models\.InThe Semantic Web – ESWC 2025,Lecture Notes in Computer Science,pp\. 244–261\.External Links:[Document](https://dx.doi.org/10.1007/978-3-031-94578-6%5F14)Cited by:[§1](https://arxiv.org/html/2609.22209#S1.p2.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px2.p1.1)\.
- Shrimalet al\.\(2025\)A\. Shrimal, A\. Jain, S\. Chowdhury, and P\. YenigallaPARSE: LLM driven schema optimization for reliable entity extraction\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track,Suzhou, China,pp\. 2749–2763\.External Links:[Link](https://aclanthology.org/2025.emnlp-industry.184/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-industry.184)Cited by:[§1](https://arxiv.org/html/2609.22209#S1.p2.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px1.p1.1)\.
- Thakuret al\.\(2025\)A\. S\. Thakur, K\. Choudhary, V\. S\. Ramayapally, S\. Vaidyanathan, and D\. HupkesJudging the judges: evaluating alignment and vulnerabilities in LLMs\-as\-judges\.InProceedings of the Fourth Workshop on Generation, Evaluation and Metrics \(GEM²\),Vienna, Austria and virtual meeting,pp\. 404–430\.External Links:[Link](https://aclanthology.org/2025.gem-1.33/),ISBN 979\-8\-89176\-261\-9Cited by:[§4\.3](https://arxiv.org/html/2609.22209#S4.SS3.p1.1)\.
- Wanget al\.\(2025\)X\. Wang, S\. L\. Huey, R\. Sheng, S\. Mehta, and F\. WangSciDaSynth: interactive structured data extraction from scientific literature with large language model\.Campbell Systematic Reviews21\(4\),pp\. e70073\.External Links:[Document](https://dx.doi.org/10.1002/cl2.70073)Cited by:[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px2.p1.1)\.
- Zaratianaet al\.\(2025\)U\. Zaratiana, G\. Pasternak, O\. Boyd, G\. Hurn\-Maloney, and A\. LewisGLiNER2: schema\-driven multi\-task learning for structured information extraction\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,Suzhou, China,pp\. 130–140\.External Links:[Link](https://aclanthology.org/2025.emnlp-demos.10/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-demos.10)Cited by:[§1](https://arxiv.org/html/2609.22209#S1.p2.1),[§2](https://arxiv.org/html/2609.22209#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.22209#S2.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging llm\-as\-a\-judge with mt\-bench and chatbot arena\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§1](https://arxiv.org/html/2609.22209#S1.p3.1),[§4\.3](https://arxiv.org/html/2609.22209#S4.SS3.p1.1)\.
## Appendix AExample Schemas
Listingshows an excerpt from one generated schema\. Due to space constraints, we present only an illustrative portion of the full schema\.
1\{
2"fields":\[
3\{
4"name":"case\_id",
5"type\_":"string",
6"description":"Uniquejudgmentorrecordidentifier\."
7\},
8\{
9"name":"judgment\_year",
10"type\_":"integer",
11"description":"Yearwhenthejudgmentwasissued\."
12\},
13\{
14"name":"outcome",
15"type\_":"enum",
16"enum\_values":\[
17"skazanie\_bez\_zawieszenia",
18"skazanie\_z\_zawieszeniem",
19"warunkowe\_umorzenie",
20"inne\_nieustalone"
21\],
22"description":"Typeofjudgmentoutcome\."
23\},
24\{
25"name":"age\_years",
26"type\_":"integer",
27"description":"Ageinyearsfortheprimarydefendant\."
28\},
29\.\.\.
30\]
31\}
Listing 1:Excerpt of a generated extraction schema\.
## Appendix BExpert\-Based Evaluation Protocol
“Study personal\-rights violations and assess their severity\.”Problem definition Rights: reputation, privacy, image Domain: civil law Context: location, witnesses, actions Goal: common violations & consequencesExpert answers All rights⋅\\cdotcivil⋅\\cdotseverity 0–5 note online platform if relevantResearch ProblemExpert 1Expert 2Expert 3Q1\. What right was violated? Q2\. Was compensation granted? …Q1\. Which right was infringed? Q2\. How serious was it? …Q1\. What right per the court? Q2\. Consequences for the victim? …???Sets of questionsAgenticSystem\{ "violation\_type": \{ "type": "enum", "enum": \["privacy", "image", \.\.\.\] \}, "violation\_degree": \{ "type": "int 0\-\-5" \}, … \}\{ \}Extraction SchemaLLM\-as\-JudgeevaluationIs eachQiQ\_\{i\}coveredby the schema?CoverageExpert 1: 8/10Expert 2: 7/10Expert 3: 9/10Avg: 8/10Evaluation
Figure 5:Expert\-based evaluation pipeline\. Given a shared research problem, each domain expert writes an exhaustive question set the extracted data should answer, while theagentic systemindependently produces an extractionschema\. AnLLM\-as\-judgechecks whether each expert question is covered by the schema, yielding per\-expert and average coverage scores\.### B\.1Expert Instructions
Listingshows a shortened version of the annotation instructions given to the experts \(the detailed instructions referenced in Section[4](https://arxiv.org/html/2609.22209#S4)\)\.
### B\.2Inter\-expert Coverage
ProblemQuestions/expertInter\-expert coverageDrug offences28\.3±\\pm9\.656\.4%±\\pm7\.9ppMedical errors33\.0±\\pm13\.052\.7%±\\pm19\.0ppPersonal rights26\.0±\\pm3\.052\.3%±\\pm17\.1ppOverall29\.1±\\pm8\.853\.8%±\\pm14\.6ppTable 2:Average number of questions written per expert and inter\-expert coverage \([SectionB\.2](https://arxiv.org/html/2609.22209#A2.SS2)\) per research problem\.Since each research problem was annotated by three experts independently, we can measure how much their question sets overlap\. We compute*inter\-expert coverage*with the same LLM\-as\-judge protocol used for schema coverage \([Section4](https://arxiv.org/html/2609.22209#S4)\), but applied between experts: for each ordered pair of experts\(a,b\)\(a,b\)on the same problem, we check how many of expertaa’s questions are covered by expertbb’s question set alone\. Since this is not symmetric, we compute it for both\(a,b\)\(a,b\)and\(b,a\)\(b,a\)for every pair, then average over all ordered pairs and problems\.[Table2](https://arxiv.org/html/2609.22209#A2.T2)reports a mean ordered\-pair inter\-expert coverage of 53\.8%, underscoring that formulating a complete extraction schema is inherently difficult and admits no single ground truth\.Schematizeexceeds this average benchmark on several backbones \([Section4\.4](https://arxiv.org/html/2609.22209#S4.SS4)\), meaning that its schemas cover more benchmark questions than are covered, on average, by another expert’s question set; this does not imply superiority to every individual expert\.
AnnotatorInstructionsSummary
RoleandTask:
TheannotatorcreatesasetofquestionsthatadataextractionschemamustanswerbasedonPolishcourtjudgments\.Thesequestionsactasabenchmarktoverifyiftheautomatedschemacapturesallkeyinformationdefinedbytheexpert\.
ResearchProblemDefinition:
Theproblemisdefinedviaathree\-partdialogue:
1\.user\_input:Theresearcherdescribestheresearchgoal\.
2\.problem\_help:Thebotasksclarifyingquestions\.
3\.user\_feedback:Theresearcherconfirms,rejects,oraddsdetails\.
Questionsmuststemfromtheentiredialogue,coveringonlytheconfirmedscopeandignoringrejectedtopics\.
CriteriaforValidQuestions:
Questionsmustdeterminewhatthesystemcanextractforquantitativeanalysis\.
1\.Extractability:Questionsmustbeanswerabledirectlyfromasinglejudgment’stext,withoutexternalknowledgeorlegaldoctrine\.
2\.Unambiguity:Questionsmustbeclearinintent\.Subjectiveevaluationsarepermittedifgroundedinthetext\.
3\.Simplicity:Onequestionequalsonefact\.Multi\-threadedquestionsmustbesplit\.
4\.Coverage:Thequestionsmustcomprehensivelycoverallaspectsconfirmedinthedialogue\.
AnswerFormats:
AcceptableformatsincludeYes/No,Number,Scale,Text,andCategory\.Ifcategoricaloptionsarenotexhaustive,useaseriesofYes/Noquestionsinstead\.
Listing 2:Shortened version of the expert annotation instructions\.
## Appendix CDetailed results
Drug offencesMedical errorsPersonal rightsBackbonevanillaLLMvanillaLLM\+ PDHScheMatiQScheMatiQ\+ PDHschematizevanillaLLMvanillaLLM\+ PDHScheMatiQScheMatiQ\+ PDHschematizevanillaLLMvanillaLLM\+ PDHScheMatiQScheMatiQ\+ PDHschematizegpt\-5\.449\.0%87\.0%59\.3%84\.3%94\.0%25\.9%53\.4%28\.1%57\.2%71\.5%33\.9%52\.5%43\.9%51\.2%72\.8%gpt\-5\.4\-mini38\.0%72\.7%15\.9%66\.3%83\.2%23\.9%47\.9%31\.7%54\.1%70\.6%33\.3%46\.0%24\.1%52\.4%70\.6%gpt\-5\.4\-nano41\.2%71\.5%9\.2%63\.6%1\.9%20\.0%52\.3%26\.1%22\.8%30\.7%5\.5%55\.6%36\.6%41\.0%71\.4%claude\-sonnet\-4\.649\.5%74\.9%82\.6%84\.1%89\.9%29\.7%50\.1%27\.3%57\.2%53\.5%32\.1%56\.4%38\.8%59\.9%66\.7%qwen3\.6\-35b\-a3b27\.9%54\.5%47\.3%84\.0%83\.9%20\.2%50\.1%22\.6%54\.0%51\.5%37\.5%50\.2%51\.4%63\.5%64\.0%gemma\-4\-e4b\-it0\.0%55\.0%53\.0%52\.0%50\.1%3\.4%26\.9%16\.6%43\.2%57\.3%7\.9%32\.8%33\.9%44\.4%52\.9%llama\-4\-scout\-17b25\.0%50\.1%54\.2%55\.8%53\.7%9\.2%32\.7%16\.8%35\.6%38\.3%48\.8%43\.2%27\.3%44\.4%43\.2%
Table 3:Per\-problem expert\-question coverage \(%\) across backbones and configurations; columns match[Table1](https://arxiv.org/html/2609.22209#S4.T1)\. Bold marks the best, underline the second\-best, configuration per backbone and problem\.While[Table1](https://arxiv.org/html/2609.22209#S4.T1)averages coverage across the three research problems,[Table3](https://arxiv.org/html/2609.22209#A3.T3)reveals substantial variation across both problems and backbones\. Averaged across backbones, ScheMatiQ \+ PDH performs best on drug offences \(70\.0%\), whereasschematizeperforms best on medical errors \(53\.3%\) and personal rights \(63\.1%\)\. Across the 21 backbone\-problem combinations,schematizeranks first in 13 and first or second in 17\. Its performance is particularly consistent on personal rights, where it leads for six of the seven backbones, and on medical errors, where it ranks among the top two for every backbone\. Results on drug offences are more mixed: althoughschematizewins for three backbones, ScheMatiQ \+ PDH achieves the highest cross\-backbone average\.
### C\.1Ablation
AblationDrugoffencesMedicalerrorsPersonalrightsOverallFull \(schematize\)83\.9%51\.5%64\.0%66\.5%±\\pm16\.3ppw/o PDH\-31\.5pp\+5\.5pp\-16\.9pp\-14\.3pp±\\pm18\.7ppw/o CBR\+0\.1pp\+0\.2pp\+4\.4pp\+1\.5pp±\\pm2\.5ppw/o DGR\+0\.7pp\+3\.9pp\-7\.4pp\-1\.0pp±\\pm5\.8ppw/o PDH\+CBR\-36\.2pp\-7\.4pp\-4\.3pp\-16\.0pp±\\pm17\.5ppw/o PDH\+DGR\-29\.4pp\-13\.7pp\-18\.8pp\-20\.6pp±\\pm8\.0ppw/o CBR\+DGR\-40\.0pp\-6\.2pp\-9\.6pp\-18\.6pp±\\pm18\.6ppw/o PDH\+CBR\+DGR\-83\.9pp\-33\.1pp\-43\.3pp\-53\.4pp±\\pm26\.9pp
Table 4:Expert\-question coverage after selectively ablating components of theschematizepipeline\. Entries show the change in coverage \(percentage points, pp\) relative to the full system, averaged over the research problems\.
## Appendix DCosts
Backbonevanilla LLMScheMatiQschematizegpt\-5\.4$0\.024±\\pm0\.008$0\.176±\\pm0\.076$3\.785±\\pm0\.312gpt\-5\.4\-mini$0\.005±\\pm0\.000$0\.044±\\pm0\.020$0\.828±\\pm0\.059gpt\-5\.4\-nano$0\.002±\\pm0\.001$0\.014±\\pm0\.007$0\.234±\\pm0\.015claude\-sonnet\-4\.6$0\.027±\\pm0\.006$0\.243±\\pm0\.086$3\.057±\\pm0\.510qwen3\.6\-35b\-a3b$0\.005±\\pm0\.001$0\.033±\\pm0\.007$0\.183±\\pm0\.010gemma\-4\-e4b\-it$0\.000±\\pm0\.000$0\.016±\\pm0\.001$0\.041±\\pm0\.001llama\-4\-scout\-17b$0\.000±\\pm0\.000$0\.007±\\pm0\.001$0\.023±\\pm0\.002
Table 5:The LLM inference cost in USD per research problem \(mean±stdmean\\pm std\), per backbone, for the*vanilla LLM*,*ScheMatiQ*, andschematizeconfigurations\.[Table5](https://arxiv.org/html/2609.22209#A4.T5)reports the average inference cost per research problem for each configuration\.Schematizeis consistently the most expensive configuration, reflecting its multi\-agent, multi\-round design, while ScheMatiQ costs roughly an order of magnitude less thanschematizeand vanilla LLM is cheapest by a further order of magnitude\. This gap is most pronounced for large closed\-weight backbones \(e\.g\., gpt\-5\.4, claude\-sonnet\-4\.6\) and narrows substantially for smaller open\-weight models \(e\.g\., gemma\-4\-e4b\-it, llama\-4\-scout\-17b\), whereschematize’s absolute cost drops to a few cents per problem\. Combined with the coverage results in[Table1](https://arxiv.org/html/2609.22209#S4.T1), this suggests thatschematize’s largest coverage gains come at a real cost premium, but that premium can be kept small by pairingschematizewith a smaller backbone rather than the largest available model\.Similar Articles
Schema (2 minute read)
Schema is a harness that enables frontier AI models to achieve 99% on the ARC-AGI-3 benchmark by having them write executable programs to model game environments, test predictions, and plan.
SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction
SchemaRAG is a retrieval-augmented generation framework that dynamically reduces the output schema space for LLM-driven structured information extraction, achieving improved performance and efficiency on healthcare and e-commerce datasets.
Executable Schema Contracts: From Automatic Ingestion to Multi-Source Retrieval
This paper presents a system that automatically discovers an executable schema from raw multi-source data and uses it for knowledge graph construction and query-time retrieval, improving over baselines on QA benchmarks.
ASMR: Agentic Schema Generation for Ship Maintenance Report Writing
This paper proposes ASMR, an agentic framework with a Field Generation Agent and a Structural Optimizer Agent that uses reinforcement learning to automatically generate compact and informative schemas from historical ship maintenance reports, aiming to improve report completeness and consistency.
SAGE: Schema-Guided LLMs for Grant Review
SAGE is a schema-guided system that uses large language models to automate grant review by structuring rubrics and linking evidence, with human-in-the-loop validation to improve accuracy.