EviStreams: 医学系统综述中的人机协作 AI 数据提取

arXiv cs.CL 论文

摘要

EviStreams 是一个开源、无代码的网络平台,它使用人机协作 AI 进行系统综述中的数据提取,表明数据提取的质量更依赖于字段规范而非模型选择。

arXiv:2609.27418v1 Announce Type: new Abstract: Systematic reviews underpin clinical guidelines, yet their data-extraction step is a major expert-labor bottleneck bound by a protocolized workflow: two reviewers extract each study independently, an adjudicator resolves disagreements, and the team keeps an auditable record of how every value was produced. Large language models can assist with extraction, but that assistance must fit established review protocols and preserve reproducibility. We present EviStreams, a live, open-source, no-code web platform that puts review teams in control of AI-assisted extraction at three key stages: program design (a structured decomposition approved before any code runs), field specification (typed field definitions calibrated from a pilot), and extracted predictions (reviewer-blinded dual review with adjudication). Working through a form builder, a domain expert defines typed fields rather than prompts, runs extraction over uploaded PDFs, inspects every value alongside the supporting passage it came from, and resolves a reviewer-blinded dual review into an auditable consensus export. An evaluation across four clinical corpora and three frontier model families, released with the system, shows that extraction quality is shaped far more by the field specification than by the choice of model. EviStreams is live at https://evistreams.com/demo and released under Apache-2.0.
查看原文
查看缓存全文

缓存时间: 2026/09/24 09:24

# Human-in-the-Loop AI Data Extraction for Systematic Reviews in Medicine
Source: [https://arxiv.org/html/2609.27418](https://arxiv.org/html/2609.27418)
## evistreams: Human\-in\-the\-Loop AI Data Extraction for Systematic Reviews in MedicineThanks:Live demo:[https://evistreams\.com/demo](https://evistreams.com/demo)

Sai Karthik KosuriAffiliation:Center for Integrative Global Oral Health \(CIGOH\)Penn Dental Medicine, University of PennsylvaniaAnkita Shashikant BhosaleAffiliation:Center for Integrative Global Oral Health \(CIGOH\)Penn Dental Medicine, University of PennsylvaniaMichael GlickAffiliation:Center for Integrative Global Oral Health \(CIGOH\)Penn Dental Medicine, University of PennsylvaniaAlonso Carrasco\-LabraAffiliation:Center for Integrative Global Oral Health \(CIGOH\)Penn Dental Medicine, University of PennsylvaniaChris Callison\-BurchAffiliation:Department of Computer and Information Science, University of PennsylvaniaCorresponding author:karthik9@upenn\.edu

###### Abstract

Systematic reviews underpin clinical guidelines, yet their data\-extraction step is a major expert\-labor bottleneck bound by a protocolized workflow: two reviewers extract each study independently, an adjudicator resolves disagreements, and the team keeps an auditable record of how every value was produced\. Large language models can assist with extraction, but that assistance must fit established review protocols and preserve reproducibility\. We presentevistreams, a live, open\-source, no\-code web platform that puts review teams in control of AI\-assisted extraction at three key stages:*program design*\(a structured decomposition approved before any code runs\),*field specification*\(typed field definitions calibrated from a pilot\), and*extracted predictions*\(reviewer\-blinded dual review with adjudication\)\. Working through a form builder, a domain expert defines typed fields rather than prompts, runs extraction over uploaded PDFs, inspects every value alongside the supporting passage it came from, and resolves a reviewer\-blinded dual review into an auditable consensus export\. An evaluation across four clinical corpora and three frontier model families, released with the system, shows that extraction quality is shaped far more by the field specification than by the choice of model\.evistreamsis live at[https://evistreams\.com/demo](https://evistreams.com/demo)and released under Apache\-2\.0\.

## 1Introduction

Systematic reviews inform clinical and public health practice guidelines, health\-technology assessments, and payer and regulatory decisions\. Their authority rests on a demanding protocol: two reviewers extract each study independently, an adjudicator resolves disagreements, blinding guards against anchoring, and the team keeps an auditable record of how every value was produced\. Data extraction \(the transfer of numerical and categorical findings into structured tables\) is among the most labor\-intensive steps of a systematic review\([Higgins et al\., 2023](https://arxiv.org/html/2609.27418#bib.bib27)\)\. Large language models can fill in much of an extraction form, though how well varies sharply from field to field\([Simmons et al\., 2025](https://arxiv.org/html/2609.27418#bib.bib9)\), and an extraction error can be catastrophic: a single mis\-extracted sample size or effect estimate can propagate into a pooled estimate and a clinical recommendation\. Therefore we are motivated to design a human\-in\-the\-loop system for systematic reviews that lets humans oversee the review, trace each value, and verify the outputs\. We investigate three key stages that put humans in control of AI\-generated extraction while meeting the evidentiary standards of the discipline\.

Our main finding is that extraction quality is governed far less by*which*model reads a paper than by*how*the task is specified\. This echoes the growing observation in clinical information extraction that “success may increasingly hinge on the clear articulation of objectives, rather than on singular workflow methodologies”\([Hein et al\., 2025](https://arxiv.org/html/2609.27418#bib.bib8)\), and that customized, domain\-tuned prompts dominate generic ones\([Li et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib10)\), which we elevate into an architectural commitment\. We study extraction across four clinical corpora\. Choosing the model or changing the extraction pipeline changes performance by only a few F1 points, with mixed direction across corpora\. The field specification has the largest effect: a richer, more detailed specification produces better results\.

We presentevistreams\(short for evidence streams\), an open\-source, live platform running at[https://evistreams\.com/demo](https://evistreams.com/demo)that puts domain experts in control of AI\-assisted extraction through a no\-code interface\. Human review enters at three key stages: \(1\)*program*, where the system proposes a plan for extracting each field from the PDF and a reviewer approves, edits, or rejects it before any code runs; \(2\)*specification*, where each field’s definition is tested on a handful of papers and the reviewer refines it based on what comes back; and \(3\)*predictions*, where every extracted value is checked by two independent reviewers, and a third resolves any disagreement\. Figure[1](https://arxiv.org/html/2609.27418#S1.F1)shows how the domain expert controls extraction by adding hints, rules, and examples\.

![Refer to caption](https://arxiv.org/html/2609.27418v1/figures/field_editor_example.png)Figure 1:A structured field in the form builder \(surgery\_type\): description, extraction hints, rules, and examples, edited directly by the reviewer\.Every AI\-extracted value is paired with the verbatim quote it was extracted from, which reviewers can inspect\. Unlike prior work, which lets users correct only the final output,evistreamsgives many ways to control how the system behaves: from the extraction plan to the field specification to the final adjudication step\.

## 2TheevistreamsSystem

We designedevistreamsfor the ADA Living Guidelines Program, a collaboration between the American Dental Association and our lab, the Center for Integrative Global Oral Health \(CIGOH\), that keeps oral health guidelines updated as new evidence comes in\([Carrasco\-Labra et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib28)\)\. Our team usesevistreamsend to end: a methodologist designs a structured extraction form and approves the pipeline the system proposes for it, then reviewers run extraction over the uploaded PDFs and adjudicate the results, with human review entering at three key stages \(Figure[5](https://arxiv.org/html/2609.27418#A5.F5)\)\.

### 2\.1Extraction pipeline

A methodologist designs a*form*, giving only a name, a description, and optionally an example for each field\. The system groups related fields into stages and proposes this grouping as a plan \(Figure[2](https://arxiv.org/html/2609.27418#S2.F2)\) for the reviewer to approve: never executable code\. Once approved,evistreams111Backend:[https://github\.com/karthik\-strikes/evistream\_backend](https://github.com/karthik-strikes/evistream_backend)\. Frontend:[https://github\.com/karthik\-strikes/evistream\_frontend](https://github.com/karthik-strikes/evistream_frontend)\. Released under Apache\-2\.0\.adds hints and rules to each field and turns the plan into extraction code automatically: a fixed compiler maps the approved plan to predefined DSPy components; no LLM writes this code\. Separately, uploaded PDFs are stored in S3 and converted to Markdown using Datalab\([Datalab, 2024](https://arxiv.org/html/2609.27418#bib.bib13)\)\.

![Refer to caption](https://arxiv.org/html/2609.27418v1/figures/extraction_plan_example.png)Figure 2:An extraction plan: fields are grouped into tasks, and each stage runs its groups in parallel; later stages can use results from earlier ones\. Here Stage 1 extracts clinical setting, outcomes, and trial administration fields in parallel; Stage 2 synthesizes a study summary tag from three of Stage 1’s outputs\.Once fields are enriched, each group from the plan becomes a DSPy signature\([Khattab et al\., 2024](https://arxiv.org/html/2609.27418#bib.bib11)\): fields in the same group share one Chain\-of\-Thought signature\([Wei et al\., 2022](https://arxiv.org/html/2609.27418#bib.bib18)\), and independent groups run at the same time\. A DSPy module then runs each signature to extract the data\. We skip DSPy’s automatic prompt optimizers\([Opsahl\-Ong et al\., 2024](https://arxiv.org/html/2609.27418#bib.bib16);[Agrawal et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib17)\): at our scale \(5–20 example papers\), we prioritize direct expert editing over automatic prompt optimization\([Li et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib10)\)\. Every value comes back with its source text, in the same format across every model \(Figure[3](https://arxiv.org/html/2609.27418#S2.F3)\)\.

![Refer to caption](https://arxiv.org/html/2609.27418v1/figures/run_source_grounded_example.png)Figure 3:Run: every extracted value can be expanded to its verbatim source quote and page in the original PDF, so a reviewer can verify a value \(herecountry: “Brazil”\) against exactly the sentence it came from\.
### 2\.2Structured field specification

Each field’s specification has four elements \(Figure[1](https://arxiv.org/html/2609.27418#S1.F1), from the introduction\): adescriptionof the concept \(always present, user\-owned and never rewritten\);hintslocating the evidence within a paper;rulesimposing hard output constraints \(e\.g\. “return the analyzedNN, not the enrolledNN”\); andexamplesanchoring the expected value shape\. Hints and rules are filled in automatically: a LangGraph\([LangChain, 2024](https://arxiv.org/html/2609.27418#bib.bib12)\)chain runs a prompt over the field’s description and any user\-given examples to generate them, consistent with what the user wrote\.

Splitting*what*a field means from*how*to find it lets a methodologist localize an error to just one element\. These elements are also what pilot calibration edits later \(§[2\.4](https://arxiv.org/html/2609.27418#S2.SS4)\); the reviewer refines them after seeing pilot results\.

### 2\.3Table extraction

Medical\-study tables record one row perarm\(the different drugs or interventions a study tests for a condition, e\.g\., ibuprofen versus an alternative drug\), or one row per \(arm, follow\-up, outcome\) triplet\. Naive LLM extraction on these tables often truncates the table, drifts in meaning, or hallucinates columns that were never asked for\.

We avoid this by decomposing the extraction\([Khot et al\., 2023](https://arxiv.org/html/2609.27418#bib.bib14)\): instead of reading a table positionally, we first extract only a few*anchor columns*that name each row \(e\.g\., drug name and dose\), then fan out one call per row to fill in the rest\. For example, a 10\-column table repeated across 400 papers easily loses columns under naive extraction; anchoring each row to its drug and dose first, then filling in the remaining cells per row, is designed to reduce row and column drift, since each call already knows which row it belongs to\.

### 2\.4Human review at three key stages

Program review\(§[2\.1](https://arxiv.org/html/2609.27418#S2.SS1), Figure[2](https://arxiv.org/html/2609.27418#S2.F2)\) is implemented as a LangGraph\([LangChain, 2024](https://arxiv.org/html/2609.27418#bib.bib12)\)workflow: the proposed plan is a directed acyclic graph \(DAG\) of extraction sub\-tasks with field\-level dependencies, validated mechanically for cycles and for missing or duplicated field assignments \(regenerating with structured feedback on failure\), and paused at aninterrupt\_beforecheckpoint until the reviewer approves, edits, or rejects it\.

Specification calibrationoccurs after a form is active: the reviewer tests the form on a handful of papers \(typically 3–5\), sees the results, and re\-edits the fields \(Figure[1](https://arxiv.org/html/2609.27418#S1.F1)\) based on what comes back\. These corrections update the corresponding signatures at runtime: any change to a field’s hints, rules, or description is reflected directly in its prompt\.

![Refer to caption](https://arxiv.org/html/2609.27418v1/figures/adjudication_example.png)Figure 4:Adjudication: fields where the AI and both reviewers already agree are marked*Agreed*; fields with a disagreement \(heretrial\_designandnumber\_of\_centres\) show the AI’s and both reviewers’ entries side by side for the adjudicator to pick or correct\.Adjudication222Screencast \(≤\\leq2\.5 min\):[https://www\.youtube\.com/watch?v=j1Gxyu0KitU](https://www.youtube.com/watch?v=j1Gxyu0KitU)\.occurs after extraction: two reviewers independently complete the same form on the same paper under reviewer blinding \(each sees the AI pre\-fill but not the other’s entries\)\. Field\-level disagreements are computed on demand, and an adjudicator resolves each conflict \(Figure[4](https://arxiv.org/html/2609.27418#S2.F4)\)\. For every field, the system links the AI’s pre\-fill, both reviewers’ entries, and the adjudicator’s decision together, so the full history of how a value moved from AI guess to final answer can be looked up later\. When a reviewer changes an AI value they attach their own supporting quote, so every value still shows where in the paper it came from: the system keeps the AI’s value, the reviewer’s replacement, and a pointer into the PDF for each\. Reviewer blinding keeps one reviewer’s entry from anchoring the other, preserving the independent review that systematic\-review protocols require; it does not blind either reviewer to the AI pre\-fill itself\.

#### Implementation\.

evistreamsis a Next\.js/FastAPI stack with Celery workers, Redis\-backed live WebSocket progress, Supabase \(Postgres\) storage, Markdown ingestion \(Datalab\), and a multi\-model fallback chain over the three frontier families\.

## 3Evaluation

We evaluateevistreamson four clinical datasets, measure the differences among three frontier models, and analyze the errors that occur\. The choice of frontier model does not dramatically affect quality: all three perform similarly\. The field specification drives the largest variance in quality\.

#### Setup\.

We score results with a field\-typed harness, released publicly alongsideevistreamsand implemented independently from the extraction pipeline to reduce evaluation coupling\. We use four clinical corpora whose ground truth was drawn, as reported, from completed, peer\-reviewed systematic reviews in which two review authors extracted every study independently and resolved disagreements by consensus, the same protocolevistreamssupports:oral cancer\(three reviews of oral cancer diagnostic adjuncts\([Verdugo\-Paiva et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib32);[Bhosale et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib33);[Urquhart et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib34)\); 43 diagnostic\-accuracy studies, four forms, 47 fields\),antibiotic prophylaxis\(CD010266\([Brignardello\-Petersen et al\., 2015](https://arxiv.org/html/2609.27418#bib.bib30)\); 10 trials\),periodontitis\(CD004714\([Simpson et al\., 2022](https://arxiv.org/html/2609.27418#bib.bib29)\); 29 trials, outcomes per arm×\\timestime point\), andibuprofen\(CD015432\([Pessano et al\., 2024](https://arxiv.org/html/2609.27418#bib.bib31)\); 43 trials\); all three model families were run on every corpus \(Appendix[A](https://arxiv.org/html/2609.27418#A1)\)333Extraction models: Claude Sonnet 4\.6\([Anthropic, 2026](https://arxiv.org/html/2609.27418#bib.bib19)\), GPT\-5\.5\([OpenAI, 2026](https://arxiv.org/html/2609.27418#bib.bib20)\), and Gemini 3\.1 Pro\([Google DeepMind, 2026](https://arxiv.org/html/2609.27418#bib.bib21)\)\(preview\); the free\-text judge is Claude Sonnet 4\.6\. We score every study whose full\-text PDF was available: 10 of 11 antibiotic\-prophylaxis trials and 29 of 35 periodontitis studies \(the rest were excluded for PDF unavailability\); ibuprofen is complete at 43 of 43\.\.

A field’s initial definition comes from the review protocol, which sets out the PICO definitions, the eligibility criteria and the outcome definitions before anyone extracts anything, and which the reviewers who produced our reference data worked from as well\. Its hints, rules and examples were then refined on pilot papers drawn from the same corpus we later score against\.

Each field is scored by a type\-matched strategy, and record\-level forms are aligned by a globally optimal Hungarian assignment over row\-identity keys before scoring, so a missed arm counts against recall and a spurious one is an over\-extraction\. We report micro\- and macro\-F1, which agree within∼\\sim0\.02 on most forms \(up to∼\\sim0\.06 on the hardest oral\-cancer forms\)\.

We test the model, the pipeline decomposition, and the field specification one at a time, holding the other two fixed\. We score against the reference data, though this ground truth can itself contain reviewer conventions or occasional errors\. The free\-text judge \(Claude Sonnet 4\.6\) is itself one of the three extraction models, a potential circularity readers should weigh when interpreting free\-text scores\. That judge decides 34% of scored comparisons; the rest are settled by deterministic numeric or normalised\-string matching\.

### 3\.1Production quality and robustness

Table[1](https://arxiv.org/html/2609.27418#S3.T1)reports overall micro\-F1 per corpus for all three frontier families through the identical pipeline and specifications \(per\-form detail in Appendix[B](https://arxiv.org/html/2609.27418#A2)\)\. Quality varies little across the families tested \(Claude and Gemini within0\.030\.03micro\-F1 on every corpus, GPT trailing by up to∼\\sim0\.05 on antibiotic and periodontitis\), and is bounded far more by the form and its source text than by the model \(the between\-corpus range,0\.7410\.741–0\.8980\.898, dwarfs the between\-model one\)\. Oral cancer, our most complex corpus \(four forms, 47 fields\), is consistently hardest across all three models; antibiotic prophylaxis is consistently easiest, tracking form complexity rather than any one model’s weakness\. With specification and pipeline fixed, the provider can thus be chosen on cost, latency, or governance rather than accuracy\. Collapsing the multi\-signature pipeline into a single hand\-tuned prompt changes overall F1 by at most0\.0360\.036with mixed direction \(the one sizable swing, antibiotic tables, is GPT\-specific:Δ=−0\.096\\Delta=\-0\.096vs\.−0\.005\-0\.005for Claude and Gemini\), so decomposition provides governable structure, not consistent accuracy gains \(Appendix[D](https://arxiv.org/html/2609.27418#A4)\)\.

Table 1:Overall micro\-F1 per corpus, three frontier model families through the same pipeline with identical form specifications \(best per row in bold\)\. Macro\- and micro\-F1 differ by at most0\.030\.03at the corpus level; per\-form macro\-F1 detail in Appendix[B](https://arxiv.org/html/2609.27418#A2)\.
### 3\.2Field specification is the largest observed lever

Across the four calibrated corpora and three tested model families, richer field specifications improved macro\-F1 in 40 of 51 model–form comparisons \(78%78\\%\), with a meanΔ​F1\\Delta F\_\{1\}of0\.0490\.049and a median of0\.0140\.014\(full per\-form ablation in Appendix[C](https://arxiv.org/html/2609.27418#A3)\)\. The mean is strongly influenced by the oral\-cancerindex\_testform: removing hints, rules, and examples reduced its macro\-F1 from0\.7240\.724to0\.2470\.247for Claude, with similarly large declines for Gemini and GPT\. Excluding this form, the mean improvement was0\.0230\.023, and the full specification remained better in 37 of 48 comparisons \(77%77\\%\)\. Thus, richer specifications usually provide modest gains but can become decisive for fields requiring disambiguation, source localization, or non\-obvious extraction conventions\. Model and decomposition effects, reported above, were comparatively small and mixed in direction\.

### 3\.3Error profile

Table 2:Taxonomy of1,0031\{,\}003reviewed discrepancies from an error\-enriched sample of 52 studies\. Percentages are shares of reviewed discrepancies \(rounded; subtypes need not sum to the group total\), not extraction error rates\. Detailed definitions and causes appear in Appendix[E](https://arxiv.org/html/2609.27418#A5)\(Table[7](https://arxiv.org/html/2609.27418#A5.T7)\)\.We reviewed and classified every scored discrepancy in 52 studies \(all 10 antibiotic trials; 15/43 oral\-cancer, 15/43 ibuprofen, 12/29 periodontitis, each of the latter three by set\-cover over the lowest\-scoring fields\): an LLM\-assisted first\-pass classification, followed by manual verification of every label by an expert systematic\-review methodologist\. Table[2](https://arxiv.org/html/2609.27418#S3.T2)summarizes the discrepancy taxonomy and prevalence: of1,0031\{,\}003discrepancies, two\-thirds aremethodological\(ground\-truth, convention, ambiguity, or scoring\-side\), not model failures; under a third aregenuine model errors, concentrated in dense multi\-arm, multi\-table reporting; the rest are unadjudicable\. Among genuine model errors, omissions and context\-assignment mistakes were both common\. Missing rows and false\-NR omissions accounted for 138 discrepancies, while 160 involved incorrect arms, time points, cohorts, denominators, or regimens\. Pilot calibration is intended to reduce specification\-driven misses, while dual review and adjudication provide a final check on residual errors\. Detailed definitions and causes for each subtype are provided in Appendix[E](https://arxiv.org/html/2609.27418#A5)\.

## 4Demonstration

#### Demo walkthrough\.

In the reviewer demo, a methodologist first uploads a batch of PDFs, whichevistreamsconverts to Markdown via Datalab\([Datalab, 2024](https://arxiv.org/html/2609.27418#bib.bib13)\)in the background while the methodologist keeps working\. The methodologist then creates a form, naming each field with a description and an optional example;evistreamsproposes a decomposition that groups related fields into parallel extraction stages \(Figure[2](https://arxiv.org/html/2609.27418#S2.F2)\), and the methodologist reviews and approves the plan before any code runs\.

evistreamsthen runs a pilot study, extracting the form over 3–5 papers so the methodologist can check each AI value against the expected value and refine hints, rules, or examples on any field that came back wrong; edits save live and are reflected in the next pilot run\.

Once the pilot output matches expectations, the methodologist runs the approved form over the remaining documents, and every extracted value is shown alongside the supporting passage it was drawn from \(Figure[3](https://arxiv.org/html/2609.27418#S2.F3)\)\. Two reviewers then carry out manual extraction: each independently completes the same form on the same paper under reviewer blinding, accepting or correcting the AI pre\-fill\.evistreamssurfaces their field\-level disagreements, and an adjudicator resolves each conflict into the consensus export \(Figure[4](https://arxiv.org/html/2609.27418#S2.F4)\), which records the AI suggestion, both reviewer entries, the adjudicated value, and inter\-rater agreement\.

#### Target audience\.

Systematic\-review teams: clinical experts who extract data, and methodologists who need dual review, adjudication, and inter\-rater statistics for Cochrane\-style protocols\([Higgins et al\., 2023](https://arxiv.org/html/2609.27418#bib.bib27)\)\. A secondary audience is NLP researchers, for whom the per\-field decision chain \(AI pre\-fill, R1, R2, adjudicator\) is a reusable data asset\.

## 5Related Work

#### Systems for evidence synthesis\.

AutoForest\([Pronesti et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib22)\)targets a different endpoint \(Cochrane forest\-plot generation from PDFs\) with interactive correction of ICO elements, extracted values, and synthesis outputs\.ScheMatiQ\([Levy et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib23)\)discovers a schema from a research question plus a document collection and supports expert revision of the schema and the extracted table; it is our closest general\-purpose sibling\.

Established review tools\([Marshall et al\., 2016](https://arxiv.org/html/2609.27418#bib.bib4);[Sun et al\., 2025](https://arxiv.org/html/2609.27418#bib.bib5)\)and annotation platforms\([Pei et al\., 2022](https://arxiv.org/html/2609.27418#bib.bib6);[Klie et al\., 2018](https://arxiv.org/html/2609.27418#bib.bib7)\)place the human loop at the prediction level only, correcting outputs rather than the program that produces them\.

Table[3](https://arxiv.org/html/2609.27418#S5.T3)compares the extraction systems directly\. Several let a user define the extraction schema, and most let a reviewer correct extracted values, and the commercial platforms carry blinded dual review with adjudication\.evistreamsdiffers in one place outright, review at the program level before any extraction runs, and in combining that with specification calibration and blinded dual review in a single workflow\.

Table 3:Extraction\-assistance capabilities across systems:ScheMatiQ\([Levy et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib23)\),AutoForest\([Pronesti et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib22)\),ROBoto2\([Hevia et al\., 2025](https://arxiv.org/html/2609.27418#bib.bib1)\), and the commercial platforms Elicit, Covidence and DistillerSR, whose capabilities were verified against their official documentation\([Elicit, 2026](https://arxiv.org/html/2609.27418#bib.bib35);[Covidence, 2026a](https://arxiv.org/html/2609.27418#bib.bib36);[Covidence, 2026b](https://arxiv.org/html/2609.27418#bib.bib37);[DistillerSR, 2026](https://arxiv.org/html/2609.27418#bib.bib38);[DistillerSR, 2016](https://arxiv.org/html/2609.27418#bib.bib39)\);Trialstreamer\([Nye et al\., 2020](https://arxiv.org/html/2609.27418#bib.bib2)\)andClinicalTrialsHub\([Park et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib3)\)extract fixed elements rather than a review\-defined form and are not compared here\.UDS: user\-defined schema;AIX: AI\-generated pre\-fill;SRC: value\-level source grounding \(each value linked to its verbatim passage\);OUT: reviewer correction of extracted outputs;PRG: program\-level review \(extraction plan approved before it runs\);CAL: specification calibration loop \(pilot results→\\rightarrowspec edits with immediate effect\);DRA: reviewer\-blinded dual review with adjudication carried to the export\. ⚫ supported; ◗ partial \(AutoForest: control over ICO elements only, and free\-text extraction rationales without links to source locations; Covidence: field suggestions rather than form pre\-fill, and a supporting quote per suggestion without a link to its location; Elicit: iteration over free\-text column prompts rather than structured field elements, and dual review described for screening rather than extraction\); — not described in the cited work or documentation\. Systems target different endpoints and corpora, and no shared extraction benchmark exists; we therefore compare capabilities, and release our evaluation harness and four corpora to enable future head\-to\-head comparison\.
#### Specification, decomposition, and data extraction\.

Our thesis turns a growing observation \(that objective articulation can matter more than model choice\([Hein et al\., 2025](https://arxiv.org/html/2609.27418#bib.bib8);[Li et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib10)\)\) into a structured, editable control surface tested across models\. Concurrent clinical information\-extraction systems have also explored modular or agent\-based decomposition\([Aal Abdulsalam et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib25);[Hart and Bergamaschi, 2026](https://arxiv.org/html/2609.27418#bib.bib26)\)\.

On the data side,VaxScope\([İlgen et al\., 2026](https://arxiv.org/html/2609.27418#bib.bib24)\)benchmarks review\-level extraction with a fixed taxonomy, whereasevistreamsexposes user\-defined structured fields in a live loop\. We adopt a variant of the error taxonomy of[Simmons et al\. \(2025\)](https://arxiv.org/html/2609.27418#bib.bib9); the per\-field extractor is a Chain\-of\-Thought module\([Wei et al\., 2022](https://arxiv.org/html/2609.27418#bib.bib18)\)compiled through DSPy\([Khattab et al\., 2024](https://arxiv.org/html/2609.27418#bib.bib11)\), and our anchor\-column table extraction \(§[2\.3](https://arxiv.org/html/2609.27418#S2.SS3)\) is a domain instance of decomposed prompting\([Khot et al\., 2023](https://arxiv.org/html/2609.27418#bib.bib14);[Liu et al\., 2025](https://arxiv.org/html/2609.27418#bib.bib15)\)\.

We therefore do not claim novelty in pipeline decomposition or source\-grounding in themselves \(both appear in the systems above\), but in unifying human review over a structured, editable specification, and in showing empirically that, across our corpora, the field specification is the largest and most consistent control surface for quality\.

## 6Conclusion

evistreamsis a human\-in\-the\-loop platform for conducting systematic reviews via human–AI collaboration\. It is made up of several parts: the extraction*program*\(approved by users before it runs\), the*specification*\(iteratively refined by letting the user run test questions\), and the*extractions*\(validated through reviewer\-blinded dual review\)\. Our evaluation shows, across four clinical corpora and three frontier model families, that extraction quality is robust to model and pipeline decomposition, and that the structured field specification drives the biggest variance in quality\. Expert effort is therefore best spent on the specification, with adjudication the second\-most\-useful place for human intervention to catch residual errors\. Because a field’s specification can be edited and take effect immediately, without regenerating the pipeline, the lever our evaluation identifies as most valuable is also the cheapest one to pull\. An extraction workflow that is expert\-controlled and auditable is intended to reduce transcription work and supports systematic\-review workflows\.

## Limitations

#### Human agreement\.

Inter\-annotator agreement was not measured when the reference data was created \(§[3](https://arxiv.org/html/2609.27418#S3)\)\.

#### Specification calibration\.

We refined specifications using pilot papers from the same reviews used for evaluation\. Our results therefore measure the benefit of review\-specific, calibrated specifications\. Performance when specifications are developed for a new review without prior calibration remains untested\.

#### LLM\-judge overlap\.

The free\-text judge shares a provider family with one extraction model \(§[3](https://arxiv.org/html/2609.27418#S3.SS0.SSS0.Px1)\), creating potential evaluator dependence that we do not quantify\.

#### Corpus scope\.

We evaluated on four corpora totalling 125 studies\. All four come from completed, peer\-reviewed systematic reviews, and we chose them to differ in the kind of question they ask: diagnostic accuracy, surgical prophylaxis, treatment of a chronic disease, and relief of post\-operative pain\. We are extending the evaluation to reviews from further areas to see whether the field specification is still the biggest lever there\.

## Ethics Statement

evistreamsis a research tool for evidence synthesis, not a medical device, and its outputs are not clinical advice; this work adheres to the ACL Code of Ethics and the ACM Code of Ethics and Professional Conduct\. Its central design property is that no AI\-extracted value enters a final dataset without human review: the dual\-review and adjudication workflow \(§[2](https://arxiv.org/html/2609.27418#S2)\) enforces verification structurally rather than relying on user diligence\.

The evaluation corpora are published, peer\-reviewed articles; no patient\-level or identifiable data is processed\. Users upload PDFs they are licensed to access; uploaded text is sent to commercial LLM APIs \(Anthropic, OpenAI, Google\) subject to those providers’ data\-handling terms\. The hosted instance restricts project data to invited team members via role\-based access control\.

Automated extraction errors could propagate into clinical guidance if left unchecked; we report the error analysis openly \(§[3\.3](https://arxiv.org/html/2609.27418#S3.SS3)\), distinguishing genuine model errors from ground\-truth and annotation\-convention discrepancies, and the per\-field decision chain records every step from AI suggestion to human\-approved value\.

## Acknowledgments

This research was developed with funding from the Defense Advanced Research Projects Agency’s \(DARPA\) SciFy program \(Agreement No\. HR00112520300\) and is based upon work supported in part by the Office of the Director of National Intelligence \(ODNI\), Intelligence Advanced Research Projects Activity \(IARPA\), via 56000026C0019 \(the BENGAL program\)\. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of DARPA, ODNI, IARPA, the Department of Defense, or the U\.S\. Government\. The U\.S\. Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copyright annotation therein\.

## References

- A\. Aal Abdulsalam, A\. Al Zaabi, R\. Jeeballah, and H\. El KerabyA multi\-agent open\-source LLM for structured cancer registry information extraction from pathology and medical reports\.InProceedings of the 25th Workshop on Biomedical Language Processing \(BioNLP\),pp\. 531–551\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.bionlp-1.43)Cited by:[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px2.p1.1)\.
- Agrawalet al\.\(2026\)L\. A\. Agrawal, S\. Tan, D\. Soylu, N\. Ziems, R\. Khare, K\. Opsahl\-Ong, A\. Singhvi, H\. Shandilya, M\. J\. Ryan, M\. Jiang, C\. Potts, K\. Sen, A\. G\. Dimakis, I\. Stoica, D\. Klein, M\. Zaharia, and O\. KhattabGEPA: reflective prompt evolution can outperform reinforcement learning\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.1](https://arxiv.org/html/2609.27418#S2.SS1.p2.1)\.
- Anthropic \(2026\)AnthropicIntroducing Claude Sonnet 4\.6\.External Links:[Link](https://www.anthropic.com/news/claude-sonnet-4-6)Cited by:[footnote 3](https://arxiv.org/html/2609.27418#footnote3)\.
- Bhosaleet al\.\(2026\)A\. S\. Bhosale, F\. Verdugo\-Paiva, O\. Urquhart, C\. Martins\-Pfeifer, M\. W\. Weinstein, T\. Walsh, K\. K\. O’Brien, A\. M\. Rojas\-Gómez, M\. Glick, M\. W\. Lingen, and A\. Carrasco\-LabraVital staining adjuncts to determine the need for biopsy for early detection of oral squamous cell carcinoma and potentially malignant disorders: an evidence summary of a living systematic review, version 2026 1\.0\.The Journal of the American Dental Association157\(6\),pp\. 588–601\.External Links:[Document](https://dx.doi.org/10.1016/j.adaj.2026.01.008)Cited by:[§3](https://arxiv.org/html/2609.27418#S3.SS0.SSS0.Px1.p1.1)\.
- Brignardello\-Petersenet al\.\(2015\)R\. Brignardello\-Petersen, A\. Carrasco\-Labra, I\. Araya, N\. Yanine, L\. Cordova Jara, and J\. VillanuevaAntibiotic prophylaxis for preventing infectious complications in orthognathic surgery\.Cochrane Database of Systematic Reviews\(1\),pp\. CD010266\.External Links:[Document](https://dx.doi.org/10.1002/14651858.CD010266.pub2)Cited by:[§3](https://arxiv.org/html/2609.27418#S3.SS0.SSS0.Px1.p1.1)\.
- Carrasco\-Labraet al\.\(2026\)A\. Carrasco\-Labra, O\. Urquhart, A\. S\. Bhosale, F\. Verdugo\-Paiva, C\. C\. Martins\-Pfeifer, K\. Aravamudhan, M\. Vujicic, T\. Wright, and M\. GlickAmerican Dental Association living evidence\-informed guidelines: shaping the future of clinical and public oral health decision making\.The Journal of the American Dental Association157\(3\),pp\. 213–217\.External Links:[Document](https://dx.doi.org/10.1016/j.adaj.2025.07.025)Cited by:[§2](https://arxiv.org/html/2609.27418#S2.p1.1)\.
- Covidence \(2026a\)CovidenceAI feature: extraction suggestions\.Note:[https://support\.covidence\.org/help/ai\-feature\-extraction\-suggestions](https://support.covidence.org/help/ai-feature-extraction-suggestions)Accessed 30 August 2026Cited by:[Table 3](https://arxiv.org/html/2609.27418#S5.T3)\.
- Covidence \(2026b\)CovidenceHow to do comparison and consensus in extraction 1\.Note:[https://support\.covidence\.org/help/comparison\-and\-consensus](https://support.covidence.org/help/comparison-and-consensus)Accessed 30 August 2026Cited by:[Table 3](https://arxiv.org/html/2609.27418#S5.T3)\.
- Datalab \(2024\)DatalabDatalab: document intelligence API\.Note:[https://www\.datalab\.to](https://www.datalab.to/)Cited by:[§2\.1](https://arxiv.org/html/2609.27418#S2.SS1.p1.1),[§4](https://arxiv.org/html/2609.27418#S4.SS0.SSS0.Px1.p1.1)\.
- DistillerSR \(2016\)DistillerSRData extraction: weighing your options\.Note:[https://www\.distillersr\.com/resources/blog/data\-extraction\-weighing\-your\-options](https://www.distillersr.com/resources/blog/data-extraction-weighing-your-options)Published 14 July 2016; accessed 30 August 2026Cited by:[Table 3](https://arxiv.org/html/2609.27418#S5.T3)\.
- DistillerSR \(2026\)DistillerSRPurpose\-built GenAI for literature reviews\.Note:[https://www\.distillersr\.com/resources/guides\-white\-papers/purpose\-built\-genai\-for\-literature\-reviews](https://www.distillersr.com/resources/guides-white-papers/purpose-built-genai-for-literature-reviews)Accessed 30 August 2026Cited by:[Table 3](https://arxiv.org/html/2609.27418#S5.T3)\.
- Elicit \(2026\)ElicitSystematic literature reviews\.Note:[https://elicit\.com/solutions/literature\-review](https://elicit.com/solutions/literature-review)Accessed 30 August 2026Cited by:[Table 3](https://arxiv.org/html/2609.27418#S5.T3)\.
- Google DeepMind \(2026\)Google DeepMindGemini 3\.1 Pro model card\.External Links:[Link](https://deepmind.google/models/model-cards/gemini-3-1-pro/)Cited by:[footnote 3](https://arxiv.org/html/2609.27418#footnote3)\.
- Hart and Bergamaschi \(2026\)S\. N\. Hart and T\. S\. BergamaschiAgent\-based large language model system for extracting structured data from breast cancer synoptic reports: a dual\-validation study\.JAMIA Open9\(1\),pp\. ooag016\.External Links:[Document](https://dx.doi.org/10.1093/jamiaopen/ooag016)Cited by:[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px2.p1.1)\.
- Heinet al\.\(2025\)D\. Hein, A\. Christie, M\. Holcomb, B\. Xie, A\. Jain, J\. Vento, N\. Rakheja, A\. H\. Shakur, S\. Christley, L\. G\. Cowell, J\. Brugarolas, A\. R\. Jamieson, and P\. KapurIterative refinement and goal articulation to optimize large language models for clinical information extraction\.npj Digital Medicine8,pp\. 301\.External Links:[Document](https://dx.doi.org/10.1038/s41746-025-01686-z)Cited by:[§1](https://arxiv.org/html/2609.27418#S1.p2.1),[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px2.p1.1)\.
- Heviaet al\.\(2025\)A\. Hevia, S\. Chintalapati, V\. K\. W\. Lai, N\. T\. Tam, W\. Wong, T\. P\. Klassen, and L\. L\. WangROBOTO2: an interactive system and dataset for LLM\-assisted clinical trial risk of bias assessment\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,Suzhou, China,pp\. 12–25\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-demos.2)Cited by:[Table 3](https://arxiv.org/html/2609.27418#S5.T3)\.
- J\. P\. T\. Higgins, J\. Thomas, J\. Chandler, M\. Cumpston, T\. Li, M\. J\. Page, and V\. A\. Welch \(Eds\.\) \(2023\)J\. P\. T\. Higgins, J\. Thomas, J\. Chandler, M\. Cumpston, T\. Li, M\. J\. Page, and V\. A\. Welch \(Eds\.\)Cochrane handbook for systematic reviews of interventions\.version 6\.4 \(updated August 2023\) edition,Cochrane\.External Links:[Link](https://training.cochrane.org/handbook)Cited by:[§1](https://arxiv.org/html/2609.27418#S1.p1.1),[§4](https://arxiv.org/html/2609.27418#S4.SS0.SSS0.Px2.p1.1)\.
- İlgenet al\.\(2026\)B\. İlgen, E\. Awotoro, and G\. HattabVaxScope: document\-level structured evidence extraction from immunization systematic reviews\.InProceedings of the 25th Workshop on Biomedical Language Processing \(BioNLP\),pp\. 853–863\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.bionlp-1.69)Cited by:[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px2.p2.1)\.
- Khattabet al\.\(2024\)O\. Khattab, A\. Singhvi, P\. Maheshwari, Z\. Zhang, K\. Santhanam, S\. Vardhamanan, S\. Haq, A\. Sharma, T\. T\. Joshi, H\. Moazam, H\. Miller, M\. Zaharia, and C\. PottsDSPy: compiling declarative language model calls into self\-improving pipelines\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.1](https://arxiv.org/html/2609.27418#S2.SS1.p2.1),[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px2.p2.1)\.
- Khotet al\.\(2023\)T\. Khot, H\. Trivedi, M\. Finlayson, Y\. Fu, K\. Richardson, P\. Clark, and A\. SabharwalDecomposed prompting: a modular approach for solving complex tasks\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.3](https://arxiv.org/html/2609.27418#S2.SS3.p2.1),[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px2.p2.1)\.
- Klieet al\.\(2018\)J\. Klie, M\. Bugert, B\. Boullosa, R\. Eckart de Castilho, and I\. GurevychThe INCEpTION platform: machine\-assisted and knowledge\-oriented interactive annotation\.InProceedings of the 27th International Conference on Computational Linguistics: System Demonstrations,pp\. 5–9\.External Links:[Link](https://aclanthology.org/C18-2002/)Cited by:[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px1.p2.1)\.
- LangChain \(2024\)LangChainLangGraph: a framework for building stateful, multi\-actor applications with LLMs\.Note:[https://langchain\-ai\.github\.io/langgraph/](https://langchain-ai.github.io/langgraph/)Cited by:[§2\.2](https://arxiv.org/html/2609.27418#S2.SS2.p1.1),[§2\.4](https://arxiv.org/html/2609.27418#S2.SS4.p1.1)\.
- Levyet al\.\(2026\)S\. Levy, E\. Habba, R\. Mintz, B\. Raveh, R\. Keydar, and G\. StanovskyScheMatiQ: from research question to structured data through interactive schema discovery\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),pp\. 220–230\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-demo.22)Cited by:[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px1.p1.1),[Table 3](https://arxiv.org/html/2609.27418#S5.T3)\.
- Liet al\.\(2026\)L\. Li, A\. Mathrani, and T\. SusnjakWhat level of automation is “good enough”? A benchmark of large language models for meta\-analysis data extraction\.Research Synthesis Methods17\(4\),pp\. 671–692\.External Links:[Document](https://dx.doi.org/10.1017/rsm.2025.10066)Cited by:[§1](https://arxiv.org/html/2609.27418#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.27418#S2.SS1.p2.1),[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2025\)S\. Liu, Y\. Liu, Z\. Wang, Y\. Wang, H\. Wu, L\. Xiang, and Z\. HeSelect\-then\-decompose: from empirical analysis to adaptive selection strategy for task decomposition in large language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 5454–5477\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.278)Cited by:[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px2.p2.1)\.
- Marshallet al\.\(2016\)I\. J\. Marshall, J\. Kuiper, and B\. C\. WallaceRobotReviewer: evaluation of a system for automatically assessing bias in clinical trials\.Journal of the American Medical Informatics Association23\(1\),pp\. 193–201\.External Links:[Document](https://dx.doi.org/10.1093/jamia/ocv044)Cited by:[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px1.p2.1)\.
- Nyeet al\.\(2020\)B\. Nye, A\. Nenkova, I\. Marshall, and B\. C\. WallaceTrialstreamer: mapping and browsing medical evidence in real\-time\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations,pp\. 63–69\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.acl-demos.9)Cited by:[Table 3](https://arxiv.org/html/2609.27418#S5.T3)\.
- OpenAI \(2026\)OpenAIIntroducing GPT\-5\.5\.External Links:[Link](https://openai.com/index/introducing-gpt-5-5/)Cited by:[footnote 3](https://arxiv.org/html/2609.27418#footnote3)\.
- Opsahl\-Onget al\.\(2024\)K\. Opsahl\-Ong, M\. J\. Ryan, J\. Purtell, D\. Broman, C\. Potts, M\. Zaharia, and O\. KhattabOptimizing instructions and demonstrations for multi\-stage language model programs\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 9340–9366\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.525)Cited by:[§2\.1](https://arxiv.org/html/2609.27418#S2.SS1.p2.1)\.
- Parket al\.\(2026\)J\. Park, R\. Liu, A\. Jagdale, A\. Srisuwananukorn, J\. Zhao, L\. Li, P\. Zhang, and S\. KumarClinicalTrialsHub: bridging registries and literature for comprehensive clinical trial access\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),Rabat, Morocco\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.eacl-demo.26)Cited by:[Table 3](https://arxiv.org/html/2609.27418#S5.T3)\.
- Peiet al\.\(2022\)J\. Pei, A\. Ananthasubramaniam, X\. Wang, N\. Zhou, A\. Dedeloudis, J\. Sargent, and D\. JurgensPOTATO: the portable text annotation tool\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,pp\. 327–337\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-demos.33)Cited by:[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px1.p2.1)\.
- Pessanoet al\.\(2024\)S\. Pessano, N\. R\. Gloeck, L\. Tancredi, M\. Ringsten, A\. Hohlfeld, S\. Ebrahim, M\. Albertella, T\. Kredo, and M\. BruschettiniIbuprofen for acute postoperative pain in children\.Cochrane Database of Systematic Reviews\(1\),pp\. CD015432\.External Links:[Document](https://dx.doi.org/10.1002/14651858.CD015432.pub2)Cited by:[§3](https://arxiv.org/html/2609.27418#S3.SS0.SSS0.Px1.p1.1)\.
- Pronestiet al\.\(2026\)M\. Pronesti, A\. Miculescu, M\. Kapdi, P\. Flanagan, O\. Redmond, J\. Bettencourt\-Silva, G\. S\. Mannu, S\. Denaxas, R\. B\. Da Providencia E Costa, A\. Belz, and Y\. HouAutoForest: automatically generating forest plots from biomedical studies with end\-to\-end evidence extraction and synthesis\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),pp\. 128–137\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-demo.13)Cited by:[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px1.p1.1),[Table 3](https://arxiv.org/html/2609.27418#S5.T3)\.
- Simmonset al\.\(2025\)Z\. Simmons, B\. Evans, T\. Harris, H\. Woolnough, L\. Dunn, J\. Fuller, K\. Cella, and D\. DuvalAssessing the feasibility and acceptability of a bespoke large language model pipeline to extract data from different study designs for public health evidence reviews\.Cochrane Evidence Synthesis and Methods3\(6\),pp\. e70061\.External Links:[Document](https://dx.doi.org/10.1002/cesm.70061)Cited by:[§1](https://arxiv.org/html/2609.27418#S1.p1.1),[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px2.p2.1)\.
- Simpsonet al\.\(2022\)T\. C\. Simpson, J\. E\. Clarkson, H\. V\. Worthington, L\. MacDonald, J\. C\. Weldon, I\. Needleman, Z\. Iheozor\-Ejiofor, S\. H\. Wild, A\. Qureshi, A\. Walker, V\. A\. Patel, D\. Boyers, and J\. TwiggTreatment of periodontitis for glycaemic control in people with diabetes mellitus\.Cochrane Database of Systematic Reviews\(4\),pp\. CD004714\.External Links:[Document](https://dx.doi.org/10.1002/14651858.CD004714.pub4)Cited by:[§3](https://arxiv.org/html/2609.27418#S3.SS0.SSS0.Px1.p1.1)\.
- Sunet al\.\(2025\)Q\. Sun, S\. Li, T\. Bi, D\. Q\. Huynh, M\. Reynolds, Y\. Luo, and W\. LiuDocSpiral: a platform for integrated assistive document annotation through human\-in\-the\-spiral\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),pp\. 267–274\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-demo.26)Cited by:[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px1.p2.1)\.
- Urquhartet al\.\(2026\)O\. Urquhart, F\. Verdugo\-Paiva, C\. Martins\-Pfeifer, A\. S\. Bhosale, C\. Massignan, D\. Meisha, K\. K\. O’Brien, A\. M\. Rojas\-Gómez, M\. Glick, M\. W\. Lingen, and A\. Carrasco\-LabraLight\-based adjuncts to determine the need for biopsy for early detection of oral squamous cell carcinoma and potentially malignant disorders: an evidence summary of a living systematic review, version 2026 1\.0\.The Journal of the American Dental Association\.External Links:[Document](https://dx.doi.org/10.1016/j.adaj.2026.03.028)Cited by:[§3](https://arxiv.org/html/2609.27418#S3.SS0.SSS0.Px1.p1.1)\.
- Verdugo\-Paivaet al\.\(2026\)F\. Verdugo\-Paiva, C\. Martins\-Pfeifer, A\. S\. Bhosale, O\. Urquhart, T\. Walsh, M\. Balagopal, M\. Zakershahrak, K\. K\. O’Brien, C\. Ávila\-Oliver, M\. Glick, M\. W\. Lingen, and A\. Carrasco\-LabraCytology adjuncts to determine the need for biopsy for early detection of oral squamous cell carcinoma and potentially malignant disorders: an evidence summary of a living systematic review, version 2026 1\.0\.The Journal of the American Dental Association157\(3\),pp\. 235–246\.External Links:[Document](https://dx.doi.org/10.1016/j.adaj.2025.12.008)Cited by:[§3](https://arxiv.org/html/2609.27418#S3.SS0.SSS0.Px1.p1.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. H\. Chi, Q\. V\. Le, and D\. ZhouChain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2\.1](https://arxiv.org/html/2609.27418#S2.SS1.p2.1),[§5](https://arxiv.org/html/2609.27418#S5.SS0.SSS0.Px2.p2.1)\.

## Appendix AAblation and Model Coverage

All three model families \(Claude, GPT, Gemini\) were run through the production pipeline, the single\-prompt arm \(decomposition ablation\), and the description\-only arm \(specification ablation\) on all four corpora\.

## Appendix BPer\-Corpus Production Results

Table[4](https://arxiv.org/html/2609.27418#A2.T4)gives the per\-form macro\-F1 for all four corpora behind the main\-text corpus\-level micro\-F1 \(Table[1](https://arxiv.org/html/2609.27418#S3.T1)\)\. Interventions and outcomes are record\-level \(one row per arm, or per arm×\\timestime point\)\.

Table 4:Per\-form macro\-F1, three model families through the same pipeline with identical form specifications, across all four corpora \(best per row in bold; ties bolded on both\)\. Corpus\-level micro\-F1 is in Table[1](https://arxiv.org/html/2609.27418#S3.T1)\. Periodontitis is scored on the canonical 29\-field set \(OPTIONAL and structural row\-matching fields excluded\)\.
## Appendix CSpecification Ablation, Full Detail

Table[5](https://arxiv.org/html/2609.27418#A3.T5)gives the per\-form specification ablation \(Δ​F1\\Delta F\_\{1\}= full−\-description\-only\) for all three model families, the finding summarized in §[3\.2](https://arxiv.org/html/2609.27418#S3.SS2)\. Full\-pipeline per\-form macro\-F1 is in Table[4](https://arxiv.org/html/2609.27418#A2.T4)\(Appendix[B](https://arxiv.org/html/2609.27418#A2)\)\. The oral\-cancerindex\_testrecall collapse discussed in the main text reproduces across all three models \(Claude\+0\.476\+0\.476, Gemini\+0\.489\+0\.489, GPT\+0\.433\+0\.433\)\.

Table 5:Specification ablation,Δ​F1\\Delta F\_\{1\}= full−\-description\-only, macro\-F1, per form across all three model families\. Positive favors the full specification; bold marks the oral\-cancerindex\_testcollapse, reproduced by all three models\.
## Appendix DDecomposition Ablation

Table[6](https://arxiv.org/html/2609.27418#A4.T6)gives the full corpus\-by\-modality breakdown \(table vs\. scalar forms, averaged over all three model families\) behind the robustness finding in §[3\.1](https://arxiv.org/html/2609.27418#S3.SS1)\.

Table 6:Decomposition ablation: overallΔ​F1\\Delta F\_\{1\}\(full multi\-signature pipeline minus single\-prompt arm\), split by table and scalar forms and averaged over Claude, GPT, and Gemini\. Positive favors the full pipeline\.
## Appendix EDetailed Error Analysis

Table[7](https://arxiv.org/html/2609.27418#A5.T7)expands the taxonomy in Table[2](https://arxiv.org/html/2609.27418#S3.T2), giving for each discrepancy type what it is and why it occurs\.

![[Uncaptioned image]](https://arxiv.org/html/2609.27418v1/figures/architecture.png)

Figure 5:Theevistreamspipeline\. Human\-in\-the\-loop \(HITL\) review enters at three key stages: the extraction*program*\(HITL \#1\), its*specification*\(HITL \#2\), and its*predictions*\(HITL \#3\)\.Table 7:Detailed error definitions and causes for the discrepancy types in Table[2](https://arxiv.org/html/2609.27418#S3.T2)\(counts and shares are given there\)\.

相似文章

EviGraph:证据引导的自主研究代理

arXiv cs.AI

EviGraph 是一个自主研究框架,它将研究过程表示为类型化证据图,以在各阶段保持主张与证据的一致性,在 ARC-Bench-ML 和 NanoResearch-20 上提升了主张支持率和实验数据一致性。

DeepER-Med:通过智能体AI推进医学深度循证研究

arXiv cs.AI

DeepER-Med引入了一个用于循证医学研究的智能体AI框架,具有明确的证据评价标准和新的基准数据集(DeepER-MedQA),包含100个专家精选的医学问题,相比生产平台表现更优,并通过真实案例的临床验证。