EHR2Trace:面向患者世界模型与临床智能体的可审计电子病历数据基础设施

arXiv cs.LG 论文

摘要

EHR2Trace 是一个可配置的系统,能够将异构的电子病历转换为可追溯至数据源、可审计的患者事件,并导出至 OMOP 和 MEDS 格式,同时将事件发生时间与信息可用时间分离。在三个临床数据集上,该系统共转换了 8.464 亿条事件,并检测出全部 28 个注入的故障;研究还表明,基于信息可用性的过滤训练能提升留出集 AUROC(0.829 对比 0.642),并揭示了篡改回填病史所导致的虚高性能(AUROC 0.965)。

arXiv:2609.38193v1 Announce Type: new Abstract: Patient world models and clinical agents aim to predict changes in patients' health and support clinical work. Developing these systems requires reliable histories of patient conditions, treatments, and the information available at each decision. Electronic health records (EHRs) contain these histories, but differences in how events are recorded make them difficult to use consistently. We present EHR2Trace, a system that converts EHRs from different sources into traceable patient events for model training and evaluation. It links events to source records, separates event time from information availability, and distinguishes medication orders, dispensing, and administration. A shared event representation supports both OMOP and MEDS exports, with automated validation and reproducible builds. Across three clinical datasets, EHR2Trace converted 846.4 million events, with every applicable check passing except one unit-consistency check on MIMIC-IV, and detected all 28 injected faults. A controlled prediction experiment showed that assigning later diagnoses to admission time substantially inflated measured performance, and that a model trained on such data lost accuracy when deployed on histories filtered by availability. EHR2Trace provides a reusable data foundation for patient world models and clinical agents, helping researchers inspect patient histories, check conversion decisions, and evaluate models with explicit data rules.
查看原文
查看缓存全文

缓存时间: 2026/10/02 09:48

# EHR2Trace: Auditable EHR Data Infrastructure for Patient World Models and Clinical Agents
Source: [https://arxiv.org/html/2609.38193](https://arxiv.org/html/2609.38193)
Xinye YangAffiliation:Department of RadiologyAffiliation:University of Colorado Anschutz School of MedicineAffiliation:Aurora, CO, USAEmail:[yangxinye2001@gmail\.com](mailto:)Yuli WangAffiliation:Department of RadiologyAffiliation:Johns Hopkins University School of MedicineAffiliation:Baltimore, MD, USAEmail:[ywang687@jhmi\.edu](mailto:)Cheng Ting LinAffiliation:Department of RadiologyAffiliation:Johns Hopkins University School of MedicineAffiliation:Baltimore, MD, USAEmail:[clin97@jhmi\.edu](mailto:)Harrison Bai††thanks:Corresponding author\.Affiliation:Department of RadiologyAffiliation:University of Colorado Anschutz School of MedicineAffiliation:Aurora, CO, USAEmail:[harrison\.bai@cuanschutz\.edu](mailto:)

###### Abstract

Patient world models and clinical agents require histories that distinguish clinical events, treatment actions, and the information available at each decision\. We present EHR2Trace, a configurable system that converts heterogeneous electronic health records \(EHRs\) into source\-linked patient events and exports them to OMOP and MEDS\. The system separates event time from information availability, distinguishes medication orders, dispensing, and administration, and validates saved outputs against source and audit records\. Across three clinical datasets, EHR2Trace converted 846\.4 million events and detected all 28 faults in an injected catalogue\. Repeated builds produced 60 byte\-identical artifacts across 4 execution configurations on synthetic data\. In a retrospective mortality\-proxy experiment with availability\-filtered test histories, the model trained with availability filtering achieved six\-hour held\-out AUROC 0\.829, compared with 0\.642 for the model trained with backdated diagnoses\. Cross\-rule evaluation also exposed inflated AUROC of 0\.965 when backdated histories were used for both training and testing\. EHR2Trace delivers reusable, auditable data infrastructure for constructing patient histories and quantifying the effects of conversion choices on downstream models\.

## 1Introduction

Patient world models and clinical agents require longitudinal records that distinguish a patient’s condition, the care provided, and the information available at each decision\. World models use these records to predict changes in patient state following clinical actions\[[20](https://arxiv.org/html/2609.38193#bib.bib18),[16](https://arxiv.org/html/2609.38193#bib.bib19)\], while clinical agents retrieve information and perform tasks in EHR environments\[[12](https://arxiv.org/html/2609.38193#bib.bib21)\]\. Preserving these distinctions during conversion establishes a reliable basis for model training and evaluation\.

EHR conversion involves decisions about the meaning and timing of clinical records\. Medication orders, dispensing, and administration represent different stages of care\. Similarly, the time of an event may differ from the time at which it becomes available to a clinician or a model\. Conflating these distinctions introduces future information into predictions or represents a dispensed drug as an administered treatment, even when the exported data satisfy schema requirements\. Temporal leakage inflates measured performance\[[11](https://arxiv.org/html/2609.38193#bib.bib12)\]\. For offline reinforcement learning, treatment histories also require explicit definitions of actions, rewards, and episodes, together with methods for addressing confounding in observational data\[[3](https://arxiv.org/html/2609.38193#bib.bib20)\]\.

We developed EHR2Trace to preserve these distinctions during conversion and make the resulting histories auditable\. The work makes three contributions:

1. 1\.Traceable patient events\.A configurable pipeline represents source records as a shared set of patient events and exports them to OMOP and MEDS, preserving source links, timestamps, and medication action types\.
2. 2\.Automated validation\.Checks compare saved outputs with source and audit records to assess patient identity, temporal consistency, terminology mappings, and dataset splits\.
3. 3\.Empirical evaluation\.Conversions of three clinical datasets, injected faults, and repeated builds assess information preservation and reproducibility\. Comparisons with published conversions characterize the information retained, and a controlled prediction experiment measures the effects of timestamp errors\.

Figure[1](https://arxiv.org/html/2609.38193#S3.F1)locates EHR2Trace between source data and downstream trajectory construction\. The OMOP and MEDS exports integrate with existing cohort and modeling tools, with source links and temporal metadata retained for inspection\.

## 2Related work

ETHOS predicts patient trajectories from tokenized timelines, and EHRWorld studies patient states and actions for longitudinal simulation\[[20](https://arxiv.org/html/2609.38193#bib.bib18),[16](https://arxiv.org/html/2609.38193#bib.bib19)\]\. PhysicianBench evaluates clinical agents performing tasks in EHR environments\[[12](https://arxiv.org/html/2609.38193#bib.bib21)\]\. Wang et al\. integrate large language models with structured EHR data for predictive analytics\[[22](https://arxiv.org/html/2609.38193#bib.bib2)\]\. Controlled evaluations of clinical vision\-language models have further revealed sensitivity to workflow, prompt, and benchmark design\[[24](https://arxiv.org/html/2609.38193#bib.bib1)\]\. These applications rely on accurate records of available information and careful evaluation conditions\. Healthcare reinforcement learning adds requirements for action and reward definitions, confounding control, and policy evaluation\[[3](https://arxiv.org/html/2609.38193#bib.bib20)\]\.

The OMOP Common Data Model \(CDM\) 5\.4 organizes clinical data into tables with standard concepts\. The Medical Event Data Standard \(MEDS\) organizes data as patient event streams\[[4](https://arxiv.org/html/2609.38193#bib.bib3),[18](https://arxiv.org/html/2609.38193#bib.bib9),[1](https://arxiv.org/html/2609.38193#bib.bib4),[14](https://arxiv.org/html/2609.38193#bib.bib10)\]\. MEDS\-Transforms provides preprocessing tools, ACES defines cohorts and tasks, and MEDS\-DEV shares task definitions and model\-training workflows\[[13](https://arxiv.org/html/2609.38193#bib.bib11),[23](https://arxiv.org/html/2609.38193#bib.bib16),[15](https://arxiv.org/html/2609.38193#bib.bib22)\]\. EHR2Trace prepares and validates source data for use with these tools\.

Kahn et al\.’s framework organizes data\-quality assessment around conformance, completeness, and plausibility, and the OHDSI Data Quality Dashboard \(DQD\) implements checks for OMOP data\[[9](https://arxiv.org/html/2609.38193#bib.bib6),[2](https://arxiv.org/html/2609.38193#bib.bib5)\]\. EHR2Trace extends these approaches with comparisons across source records, canonical events, and exported data, including MEDS event ordering and dataset splits\. Terminology mapping relies on standard vocabularies and reviewed mappings\[[17](https://arxiv.org/html/2609.38193#bib.bib8)\]\.

## 3System design and validation

aFrom EHR data to patient world models and clinical agentsEHR sourcesAuditable datainfrastructureTrajectory layerWorld modelsand agentsEvaluationbConversion, audit records, and automated validationSource layerCanonical eventsoccurrence · availability · lineageOMOP CDM 5\.4MEDSAudit recordsextraction dates · cohort membership · quarantine · review decisionsValidation suite55 automated checks on saved outputs7Source rowsand links16Identityand time8OMOPchecks8MEDSchecks15Declarationsand references1Mappingreviewimplemented and evaluated heredownstream trajectory construction, learning, and evaluationFigure 1:\(a\)EHR2Trace forms the data layer for patient world models and clinical agents\. Dashed boxes show downstream stages\.\(b\)The system converts source records into shared patient events, exports OMOP and MEDS, and stores audit records alongside the data\. Automated checks validate the saved outputs\. Counts show the number of checks in each group\.### 3\.1Conversion with source links

EHR2Trace separates source preservation from the representation used for export\. The source layer retains raw values and row identifiers\. The canonical layer represents patient events with patient and encounter identifiers, event and availability times, codes, values, reported reference ranges, and quality flags\. Each event links to its source rows, whether several rows merge into one event or one row produces multiple events\. Extraction dates and cohort membership are stored separately from clinical events\.

Merge rules preserve distinctions that affect clinical interpretation\. Records are merged only when they describe the same fact\. For medication orders, dispensing, and administration, event identity includes dose, dose unit, route, status, and end time\. Conflicts in other fields are resolved by a rule specified in the dataset configuration\. A rule may retain the earliest value or set the field to null, flag the event, and quarantine the conflicting values\. Conflicts without a declared rule are flagged and fail validation\.

The OMOP exporter enforces the target model’s table requirements, concept domains, and patient eligibility rules\. Transfers, service changes, and intensive care unit \(ICU\) stays become visit details linked to a parent visit\. Details without a parent visit are excluded from OMOP and logged separately\. Measurement units are assigned concepts from declared unit tables and the vocabulary\. Source\-reported reference ranges take precedence over ranges configured for a code\. Infusion rates are preserved in the drug exposure’ssigtext\.

The MEDS exporter writes one code per event and retains the source code when a mapping is unavailable\. Additional fields preserve availability times, event identifiers, source links, quality flags, normalized values and units, infusion rates, order actions, and discharge destinations\. After identity resolution, a hash of each patient’s identifier assigns that patient to a training, tuning, or held\-out split\. Thetracecommand retrieves the source file and row associated with an exported record\.

Dataset\-specific conversion rules are specified in YAML files\. These files define input patterns, field aliases, patient identity and merge rules, time policies, unit declarations, and terminology normalization\. Adapters handle records with a single timestamp, narrative lines, measurements with multiple components, patient attributes, and visits\.

Preparation scripts select columns for MIMIC\-IV, CU\-CTPA, and two JHU\-CTPE tables, and join tables for MIMIC\-IV and CU\-CTPA\. Row counts and input/output hashes verify this step, and each script records which delivered files it does not read\. Two transformations add rows and record those additions: CU\-CTPA preparation separates combined blood\-pressure values into the two measurements required by OMOP, and MIMIC\-IV preparation converts wide emergency department vital\-sign and triage records to one row per value\. Before preparation, the formats of the two JHU\-CTPE tables are compared across the four cohort groups to check for export differences that could reveal cohort membership\.

### 3\.2Data for patient trajectories

A downstream trajectory layer builds a patient historyHtH\_\{t\}from these events using information available at decision timett\. A transition takes the form

\(ot,at,Δ​t,ot\+1,rt,dt,ct\),\(o\_\{t\},a\_\{t\},\\Delta t,o\_\{t\+1\},r\_\{t\},d\_\{t\},c\_\{t\}\),\(1\)whereoto\_\{t\}is the current observation,ata\_\{t\}the recorded action,Δ​t\\Delta tthe elapsed time, andot\+1o\_\{t\+1\}the next observation\. The remaining terms are a task\-defined rewardrtr\_\{t\}, a termination flagdtd\_\{t\}, and a truncation or censoring flagctc\_\{t\}\.

EHR2Trace supplies the events and timestamps from which observations, actions, and elapsed times are constructed\. Availability timestamps enable filtering at a decision time\. Medication action types distinguish orders, dispensing, and administration, while links between these stages verify whether an order was carried out\. Quality flags guide task\-specific record selection, including the handling of inconsistent time intervals\. The shared representation accommodates task\-specific rewards and episode boundaries\.

### 3\.3Time, missing data, and terminology

Time normalization follows a declared timezone\. Death records on the same local calendar date are merged using the more precise timestamp\. Records on different dates are retained and flagged, and neither is exported to the OMOP death table\. Under the strict birth\-year policy, patients without a derivable birth year and their dependent records are excluded from OMOP\. Approximate birth years require a reference date and an approval note, and are flagged in the output\. Cohort directory names remain source metadata, with labels and episodes defined downstream\.

Retired source codes are mapped through relationships retained in the vocabulary\. The output uses current standard concepts and flags the retired source code\. Dataset configurations also define filters for nonclinical rows, such as supply containers and specimen\-handling steps\. Source\-row accounting records these exclusions\. Delivered files and columns that the conversion does not read must be declared with a reason\. Records excluded from conversion are stored with reasons in a quarantine table for review\.

Unit normalization preserves the source value and unit alongside any normalized values and applies only exact conversions\. Corrections declared in the dataset configuration, such as interpreting Fahrenheit values under a Celsius label, require documented evidence and are flagged\. Values outside the plausible range for a code and unit retain their source value, receive no normalized value, and are flagged\. When a code appears in units that cannot be converted exactly, it is separated into one code per unit\.

EHR2Trace stores event time and availability time in separate columns\. For a timed eventee, the inclusion rule at prediction timeτ\\tauis

tevent​\(e\)≤τandtavailable​\(e\)≤τ\.t\_\{\\mathrm\{event\}\}\(e\)\\leq\\tau\\quad\\text\{and\}\\quad t\_\{\\mathrm\{available\}\}\(e\)\\leq\\tau\.\(2\)For MIMIC\-IV laboratory results, specimen collection defines event time andstoretimedefines availability time\. Table[1](https://arxiv.org/html/2609.38193#S3.T1)illustrates the filtering rule\. Missing availability times are replaced by event time and markedAVAILABILITY\_ASSUMED, so researchers can choose whether to include these events\. When duplicate records of a result have different availability times, the earliest is retained and markedAVAILABILITY\_MERGED\.

Records without a timestamp may use a declared fallback, such as emergency department arrival time for triage measurements, and receive the flagTIME\_FALLBACK\. Untimed attributes must appear on an approved list of stable baseline properties\. Other attributes require a timestamp or are withheld\. These policies define explicit, inspectable rules for selecting timed events and baseline attributes\.

event\_kindhours from admissionsource\_name, and the detail an action carriesatτ\\taueventavail\.Subject 1, one encounter of 75 h; decision timeτ\\tauat \+44\.08 hdrug\_admin4\.374\.38CeftriaXONE1 gm · Administered✓visit7\.98·EW EMER\.AVAILABILITY\_ASSUMED✓drug\_order12\.32·Vitamin D 1,000 Unit Tablet1000 · PO/NG · AVAILABILITY\_ASSUMED✓drug\_dispense12\.32·Docusate SodiumPO/NG · Discontinued via patient discharge · AVAILABILITY\_BEFORE\_EVENT✓drug\_order12\.32·Sodium Chloride 0\.9% Flush 10 mL Syringe3 · IV · AVAILABILITY\_ASSUMED✓measurement40\.1048\.07Thyroid Stimulating Hormonewithheldmeasurement40\.1048\.07Vitamin B12withheld261 further events in this encounter, all afterτ\\taudeath\+5882deathoutside the encounter window—Subject 2, one encounter of 2013 h; decision timeτ\\tauat \+1283\.29 hvisit0\.00·DIRECT EMER\.AVAILABILITY\_ASSUMED✓drug\_order0\.38·Influenza Vaccine Quadrivalent 0\.5 mL Syri0\.5 · IM · AVAILABILITY\_ASSUMED✓drug\_order0\.38·PNEUMOcoccal 23\-valent polysaccharide vacc0\.5 · IM · AVAILABILITY\_ASSUMED✓drug\_dispense0\.38·Sodium Chloride 0\.9% FlushIV · Discontinued · AVAILABILITY\_BEFORE\_EVENT✓drug\_dispense0\.38·Influenza Vaccine QuadrivalentIM · Discontinued · AVAILABILITY\_BEFORE\_EVENT✓drug\_admin1\.382\.52Sodium Chloride 0\.9% FlushNot Flushed · NOT\_ADMINISTERED✓measurement553\.872012\.72BRONCHOALVEOLAR LAVAGEwithheld45648 further events in this encounter, all afterτ\\taudeath\+772deathoutside the encounter window—Table 1:Two patient histories from the public MIMIC\-IV demonstration subset\[[6](https://arxiv.org/html/2609.38193#bib.bib14)\]\. Rows show canonical events in event\-time order\. A dot means event and availability times agree\. Events marked “withheld” occurred before decision timeτ\\taubut became available afterward\. Orders, dispensing, and administration remain separate records, with dose, route, status, and quality flags preserved\. TheNOT\_ADMINISTEREDflag records a treatment that was not given\. Omitted events appear as counts\.Terminology mapping uses reviewed mappings and vocabulary relationships\. Matching after punctuation normalization requires a unique vocabulary code\. When a code maps to several standard concepts, the event’s domain determines which concepts are eligible for the target column\. A fixed ordering resolves remaining ties, and the output records the ambiguity\. Drug matching uses ingredient, strength, and dose form and records the basis for each match\. Unresolved terms enter a review queue\. Optional local models suggest column roles or rank supplied candidates\. Human approval is required before their proposals enter the mapping registry\.

### 3\.4Validation of saved outputs

The validation suite assesses saved outputs using 55 checks\. Source checks account for input rows and event\-to\-source links\. Identity and time checks compare patient identifiers and timestamps with separately stored records\. Output checks examine concept domains, keys, source links, schemas, event ordering, dataset splits, and label fields\. Review checks compare model proposals with recorded decisions and accepted mappings\. Each check reports its applicability, status, and diagnostic details\.

Cross\-layer checks assess consistency between exported data and separately stored evidence\. They test whether merged source rows agree or follow a declared conflict rule, whether each source with parsed rows produces events, and whether each patient with a death event has a published death unless the dates conflict\. They also compare the concepts assigned to each term in MEDS and OMOP, which use the same vocabulary\.

Additional comparisons cover measurement and dose units, visit concepts, encounter links, repeated notes, excluded statuses, and the delivered files and columns\. Reference tables and dataset configurations define the expected values\. Configured thresholds and exceptions, such as an expected share of quarantined rows, pass validation when they satisfy their declared criteria\.

Validation operates on the saved artifacts, so it detects export errors and inconsistencies in outputs generated by earlier code versions\. Checks with empty or unavailable inputs are marked as skipped, and an OMOP database without patient rows is treated as absent\. The report records the coverage and outcome of each validation run\.

## 4Results

### 4\.1Conversion scale and concept coverage

We evaluated EHR2Trace on three datasets: a private Johns Hopkins CT pulmonary embolism extract \(JHU\-CTPE\), MIMIC\-IV v3\.1 hospital and emergency data with v2\.2 notes\[[7](https://arxiv.org/html/2609.38193#bib.bib7),[5](https://arxiv.org/html/2609.38193#bib.bib13),[8](https://arxiv.org/html/2609.38193#bib.bib15)\], and a private University of Colorado CT pulmonary angiography extract \(CU\-CTPA\)\. JHU\-CTPE comprised four cohort partitions from two extraction batches\. The MIMIC\-IV configuration covered 33 sources, and CU\-CTPA included flat comma\-separated files and complete clinical notes\.

The conversions produced 846\.4 million canonical events, with matching canonical and MEDS event counts \(Table[2](https://arxiv.org/html/2609.38193#S4.T2)\)\. Validation assessed the saved outputs of each dataset and reported pass, skip, and fail outcomes\. All applicable checks passed for JHU\-CTPE and CU\-CTPA\. The MIMIC\-IV unit\-consistency finding is detailed in Section[6](https://arxiv.org/html/2609.38193#S6)\.

Table 2:Conversion scale and validation results\. Canonical and MEDS event counts match\. Patient counts differ because some identities have no events or do not meet OMOP eligibility rules\. A source row can generate multiple quarantine entries\.Concept coverage measures the proportion of exported rows with a nonzero standard concept identifier \(Table[3](https://arxiv.org/html/2609.38193#S4.T3)\)\. The drug matcher mapped 129 of 139 reviewed free\-text medication terms to RxNorm using ingredient, strength, and dose form\. Among these resolved terms, 93\.8% matched the reviewer’s concept choice\.

Laboratory mappings used the community LOINC table for MIMIC\-IV\.111The MIT\-LCPmimic\-codefilemimic\-iv/mapping/d\_labitems\_to\_loinc\.csvuses SSSOM format; 1,397 of its entries were imported after review against LOINC component and system attributes\.Spirometry, electrocardiography, and echocardiography terms in the private extracts were reviewed against the vocabulary\. Configured filters removed nonclinical entries, including order\-entry placeholders, supplies, and physician names in result columns\.

MEDS preserves source codes and source links for events awaiting a standard concept mapping, keeping them available for source\-code analyses\. These records included 75,325 rows with ICD\-10\-CM place\-of\-occurrence codes and 450,357 JHU\-CTPE rows for 12\-lead electrocardiogram atrial rate\. Most other unresolved terms were medication names\.

Table 3:OMOP concept coverage by domain, weighted by row count\. “Mapped” indicates the presence of a standard concept identifier\.
### 4\.2Preserving time and medication actions

An audit of all canonical events confirmed that medication dispensing and administration remained distinct and that administration records retained source\-provided dose and route details \(Table[4](https://arxiv.org/html/2609.38193#S4.T4)\)\. The temporal layer also converted final vital status in JHU\-CTPE and MIMIC\-IV from an untimed attribute to a dated death event when a date was available; undated records were excluded\. CU\-CTPA supplied death dates directly\. Table[4](https://arxiv.org/html/2609.38193#S4.T4)also reports the share of assumed availability times for each dataset\.

Table 4:Time and medication\-action audit across all canonical events\. Dispensing records indicate drug supply, while administration records describe whether and how a drug was given\. Assumed availability is the share of timed events that default to event time because the source lacks an availability timestamp\. These events carryAVAILABILITY\_ASSUMED\.
### 4\.3Comparison with published conversions

We compared EHR2Trace with published OMOP and MEDS conversions of the open MIMIC\-IV demonstration subset\[[10](https://arxiv.org/html/2609.38193#bib.bib23),[21](https://arxiv.org/html/2609.38193#bib.bib24)\]\. Table[5](https://arxiv.org/html/2609.38193#S4.T5)reports measurements from the output files; Appendix[A](https://arxiv.org/html/2609.38193#A1)describes the procedure\. EHR2Trace produced both formats from one conversion and retained the source file and row for every exported event through 972,567 links\. The published OMOP conversion included no source\-row field, and the MEDS conversion retained source keys without their table names\.

EHR2Trace stored availability time separately from event time; availability time was later than event time for 71% of timed events\. Neither published conversion stored a separate availability field\. EHR2Trace also used an explicit field for medication orders, dispensing, and administration\. The OMOP baseline used one type concept for all drug rows\. The MEDS baseline used one medication code, with action types recoverable from retained source keys\. All three conversions read the hospital and ICU modules\. The published OMOP conversion maps ICU chart items through its own item tables\. EHR2Trace uses the community LOINC table for laboratory items and retains source codes for most ICU chart items\. Table[5](https://arxiv.org/html/2609.38193#S4.T5)reports the resulting concept coverage alongside each conversion’s retained information\.

Table 5:Information retained by three conversions of the 100\-patient MIMIC\-IV demonstration subset\. Results were measured from EHR2Trace outputs and the published baseline files\. Input modules are listed for each conversion\.
### 4\.4Fault detection

The validation suite detected all 28 faults in the injected catalogue \(Figure[2](https://arxiv.org/html/2609.38193#S4.F2)\)\. These faults target errors that row counts, schema checks, or small\-sample inspection would miss\. Detection relied on comparisons with source records, audit metadata, and configuration rules\. Removing the 21 cross\-layer checks reduced detection to 13 of 28 faults across the remaining 34 checks\.

DQD detected five of the nine injected faults that reached the OMOP database\. Eight faults affected data outside OMOP and were outside DQD’s scope\. The cross\-layer suite extends fault detection across ingest, canonical, OMOP, and MEDS artifacts\.

Detection on the 28\-fault catalogueDetectedMissedOutside scopeNot runIngest layer \(1\)Canonical layer \(14\)OMOP layer \(7\)MEDS layer \(6\)Raw column undeclaredAnchor as event timeCohort label a diagnosisDiagnosis read as a unitExcluded status publishedIdentity unresolvedBroken source linksNote text droppedNote under two encountersDifferent doses mergedPartition column leakedPost\-death rows deletedQuarantine reason erasedSource yielded nothingValue in another unitBirth year inventedConcept in wrong domainDeath date and time splitDose unit droppedInvented measurement datePrimary key collisionWithheld patient publishedAvailability before eventCode metadata incompleteCohort label in MEDSBuilt without vocabularyShard not time sortedSubject in two splitsOHDSI DQD2,374 checks5 / 9Reduced suite34 checks13 / 28Full suite55 checks28 / 28The reduced suite omits the 21 cross\-layer checks: 6 comparing saved content with the configuration and MEDS concepts with OMOP and 15 comparing merged rows, units, deaths, sources and delivered files with the dataset’s declarations and reference tables\. The 8 faults outside DQD’s reach never arrive in an OMOP table\.

Figure 2:Detection of injected faults, grouped by the affected data layer\. DQD was run on nine faults that reach OMOP\. The full suite covers all 28 faults\.
### 4\.5Effect of time errors on model evaluation

We retrospectively evaluated three temporal inclusion rules in a MIMIC\-IV cohort of 131,007 patients, holding the outcome labels, patient splits, model class, and fitting procedure fixed\. Rule A included timed events only when both event and availability times were at or before prediction\. Rule B filtered by event time alone\. Rule C assigned all diagnoses from the index admission to admission time, including diagnoses without individual timestamps\. Rule C represents backdating during conversion of admission\-level diagnosis tables\. We fitted a separate model under each rule\.

At six hours after admission, held\-out AUROC was 0\.829 under availability filtering and 0\.965 under diagnosis backdating \(Figure[3](https://arxiv.org/html/2609.38193#S4.F3)\)\. The difference was 0\.136 \(paired bootstrap interval, \[0\.116, 0\.157\]\), estimated from 2,000 resamples using the same patients for all three rules\. Diagnosis backdating increased measured performance at every evaluated horizon\. Filtering by event time alone had a smaller effect, indicating that the larger increase arose from assigning later admission information to earlier predictions\. Appendix[B](https://arxiv.org/html/2609.38193#A2)describes the cohort and model\.

Cross\-rule evaluation separated the effect of the test\-data rule from the effect of the training\-data rule \(Table[6](https://arxiv.org/html/2609.38193#S4.T6)\)\. At six hours, the model trained with backdated diagnoses achieved AUROC 0\.965 on backdated test histories \(C→\\rightarrowC\) and 0\.642 on availability\-filtered histories \(C→\\rightarrowA\)\. The model trained and evaluated with availability filtering achieved 0\.829 \(A→\\rightarrowA\)\.

For the same fitted model, changing the test rule reduced AUROC by 0\.323 \(paired interval, \[0\.295, 0\.352\]\)\. We refer to this contrast, C→\\rightarrowC minus C→\\rightarrowA, as evaluation inflation\. Holding the test rule fixed, the model trained with backdated diagnoses underperformed the availability\-filtered model by 0\.188 \(paired interval, \[0\.161, 0\.215\]\)\. This contrast, A→\\rightarrowA minus C→\\rightarrowA, measures the performance deficit associated with backdated training histories\. The within\-rule AUROC increase equals evaluation inflation minus the training deficit\.

AUPRC showed the same pattern at an outcome prevalence of 3\.2%\. It was 0\.071 for C→\\rightarrowA, compared with 0\.222 for A→\\rightarrowA and 0\.654 for C→\\rightarrowC\. Across the 4 horizons, the AUROC deficit under a common availability\-filtered test rule ranged from 0\.169 to 0\.188\. The model trained using event time alone performed within 0\.012 AUROC of the availability\-trained model when both were evaluated with availability filtering\.

Table 6:Retrospective evaluation under matched and mismatched training–test rules\. A denotes availability filtering and C diagnosis backdating\. Arrows indicate training→\\rightarrowtest rules\. Performance columns report held\-out AUROC, with paired percentile bootstrap intervals from 2,000 resamples for the two contrasts\. The final column reports AUPRC for A→\\rightarrowA and C→\\rightarrowA\.61224480\.800\.850\.900\.951\.00\(a\) AUROC61224480\.000\.200\.400\.600\.80Hours after admission\(b\) AUPRC61224480\.0010\.0020\.0050\.010\.050\.1\(c\) AUROC increase

Figure 3:Held\-out performance under three time rules for the mortality\-proxy task\. Backdating diagnoses exposes early predictions to later information\. Panel \(c\) shows AUROC increases relative to availability filtering on a logarithmic scale\. Prediction horizons use a base\-2 logarithmic axis\. In panel \(a\), the AUROC axis begins at 0\.80, and error bars show bootstrap intervals from 2,000 resamples of held\-out patients\.
### 4\.6Reproducibility and runtime

Reproducible builds rely on fixed merge ordering and output paths derived from input hashes, configuration, and code version\. We tested 4 builds of a synthetic dataset while varying worker count, input order, and cache reuse\. All builds produced 60 byte\-identical artifacts and passed validation\. A stage\-by\-stage rebuild of MIMIC\-IV from saved inputs reproduced the identity map, canonical events, MEDS outputs, and clinical OMOP tables\.

Runtime was measured one stage at a time on an otherwise idle machine with 48 CPU cores and 251 GB of memory\. Ingestion and canonical conversion used 12 worker processes\. The MIMIC\-IV build required 111 minutes for ingestion, 191 minutes for canonical conversion, 47 minutes for OMOP export, and 128 minutes to write 364,673 MEDS shards\. Peak memory use during canonical conversion was 167\.2 GB\.

## 5Discussion

EHR2Trace exposes every conversion decision along the path from source records to model inputs\. The shared event representation preserves clinical timing, medication actions, and source provenance across OMOP and MEDS exports\. Researchers can inspect an exported event, recover its source rows, and examine the rules that produced it\. This continuity enables reuse of the same clinical records across cohort construction, trajectory modeling, and agent evaluation\.

The temporal experiment demonstrates the value of explicit availability rules\. Under a common availability\-filtered test representation, the model trained with availability filtering achieved higher AUROC and AUPRC than the model trained with backdated diagnoses\. Cross\-rule evaluation also quantified the inflation introduced by backdated test histories\. Together, these comparisons demonstrate that temporal metadata directly improves history construction and reveals misleading performance estimates\.

For patient world models of the formpθ\(ot\+1,Δt∣Ht,at\)p\_\{\\theta\}\(o\_\{t\+1\},\\Delta t\\mid H\_\{t\},a\_\{t\}\), EHR2Trace provides explicit records for constructing the historyHtH\_\{t\}and actionata\_\{t\}\. Availability times enable filtering at each decision, while medication action types distinguish orders, dispensing, and administration\. Source links and quality flags let researchers inspect and vary these choices\. Separating conversion from task construction enables direct comparison of episode definitions, rewards, time rules, and action detail on a shared event base\.

Validation and reproducibility make these representations directly auditable and reusable\. Comparisons with source rows, reference tables, dataset declarations, and repeated builds verify the artifacts used in analysis\. The injected\-fault results demonstrate the coverage added by cross\-layer checks, and the repeated builds establish reproducibility under changes in execution settings\. Together with recorded conversion decisions, these capabilities enable inspection of data preparation at the level of individual events and complete datasets\.

#### Broader impact\.

Source\-linked histories and explicit temporal rules enable auditable training data and reproducible evaluation of clinical models\. Shared exports allow researchers to reuse existing cohort and modeling tools, while recorded source links and conversion decisions provide an inspectable account of how clinical records were processed\.

## 6Limitations and future work

#### Evaluation scope\.

The current evaluation spans three clinical datasets, a systematic fault catalogue, a reviewed medication reference set, and a retrospective mortality\-proxy task with logistic regression\. Independent error collections, additional datasets, and evaluations with patient world models, clinical agents, and prospective deployment are natural next steps\. Treatment\-effect estimation and policy learning also require methods that address confounding and censoring\[[3](https://arxiv.org/html/2609.38193#bib.bib20)\]\.

#### Source and mapping coverage\.

Validation establishes consistency with declared rules and reference information; source completeness and clinical interpretation remain the responsibility of data owners and domain experts\. Certain configurations, such as a CU\-CTPA temperature column interpreted as Fahrenheit based on value ranges despite its Celsius label, reflect documented interpretations of ambiguous source metadata\. The MIMIC\-IV full and demonstration conversions each failUNIT\_HOMOGENEOUS\_PER\_CODE\. In the full dataset, twelve laboratory codes mix unit families, which manual inspection attributed to notation differences for the same quantity\. Broader terminology coverage and unit harmonization are priorities for further development\. Reproducibility excludes the build\-date field in OMOP’scdm\_sourcemetadata table\.

#### Temporal coverage and use\.

Availability is assumed when the source provides no timestamp, including all timed CU\-CTPA events\. Individual\-field availability and revisions are not tracked, and untimed attributes or interval ends may extend beyond a prediction cutoff\. Field\-level timing, revision tracking, and task\-specific history construction will extend temporal filtering\. Deployment at scale requires compliance with applicable consent, data\-use, and de\-identification requirements\.

## 7Conclusion

EHR2Trace delivers auditable EHR infrastructure for patient world models and clinical agents\. Its shared event representation preserves source provenance, temporal metadata, and medication action types across OMOP and MEDS exports\. Conversion of 846\.4 million events across three clinical datasets, detection of all 28 injected faults, and reproducible builds demonstrate its scale and validation capabilities\. The temporal experiment demonstrates that explicit availability rules improve model training and expose inflated evaluation\. These capabilities make EHR preparation inspectable and reusable across downstream modeling tasks\.

#### Ethics approval\.

The JHU\-CTPE extract was collected at Johns Hopkins under Institutional Review Board protocol IRB00424745 \(“Radiology foundation model for chest and cardiac image interpretation”\), which approved this study\. The authors who collected the extract were then affiliated with Johns Hopkins\. The CU\-CTPA extract was collected at the University of Colorado under Institutional Review Board protocol 24\-1825\.

#### Data and code\.

EHR2Trace is available under Apache\-2\.0 at[https://github\.com/Yangxinyee/ehr2trace](https://github.com/Yangxinyee/ehr2trace), including dataset configurations, synthetic test data, validation checks, fault and reproducibility tests used in continuous integration, and the scripts and recorded results of every experiment in this paper underexperiments/\. JHU\-CTPE and CU\-CTPA are private\. MIMIC\-IV access follows PhysioNet credentialing and data\-use conditions\[[5](https://arxiv.org/html/2609.38193#bib.bib13),[8](https://arxiv.org/html/2609.38193#bib.bib15),[19](https://arxiv.org/html/2609.38193#bib.bib17)\]\. Patient\-level outputs and dataset\-specific run records are not redistributed\. Clinical vocabularies are obtained through Athena under the terms of their respective owners\[[17](https://arxiv.org/html/2609.38193#bib.bib8)\]\.

#### Writing assistance\.

OpenAI Codex and Anthropic Claude assisted with language editing, consistency checks, and figure\-generation code\.

## References

- \[1\]B\. Arnrich, E\. Choi, J\. Fries, M\. McDermott, J\. Oh, T\. Pollard, N\. Shah, E\. Steinberg, M\. Wornow, and R\. van de Water\(2024\)Medical event data standard \(MEDS\): facilitating machine learning for health\.InICLR 2024 Workshop on Learning from Time Series for Health,Cited by:[§2](https://arxiv.org/html/2609.38193#S2.p2.1)\.
- \[2\]\(2021\)Increasing trust in real\-world evidence through evaluation of observational data quality\.Journal of the American Medical Informatics Association28\(10\),pp\. 2251–2257\.External Links:[Document](https://dx.doi.org/10.1093/jamia/ocab132)Cited by:[§2](https://arxiv.org/html/2609.38193#S2.p3.1)\.
- \[3\]O\. Gottesman, F\. Johansson, M\. Komorowski, A\. Faisal, D\. Sontag, F\. Doshi\-Velez, and L\. A\. Celi\(2019\)Guidelines for reinforcement learning in healthcare\.Nature Medicine25,pp\. 16–18\.External Links:[Document](https://dx.doi.org/10.1038/s41591-018-0310-5)Cited by:[§1](https://arxiv.org/html/2609.38193#S1.p2.1),[§2](https://arxiv.org/html/2609.38193#S2.p1.1),[§6](https://arxiv.org/html/2609.38193#S6.SS0.SSS0.Px1.p1.1)\.
- \[4\]G\. Hripcsak, J\. D\. Duke, N\. H\. Shah, C\. G\. Reich, V\. Huser, M\. J\. Schuemie, M\. A\. Suchard, R\. W\. Park, I\. C\. K\. Wong, P\. R\. Rijnbeek, J\. van der Lei, N\. Pratt, G\. N\. Norén, Y\. Li, P\. E\. Stang, D\. Madigan, and P\. B\. Ryan\(2015\)Observational health data sciences and informatics \(OHDSI\): opportunities for observational researchers\.InMEDINFO 2015: eHealth\-enabled Health\. Studies in Health Technology and Informatics,Vol\.216,pp\. 574–578\.Cited by:[§2](https://arxiv.org/html/2609.38193#S2.p2.1)\.
- \[5\]A\. Johnson, L\. Bulgarelli, T\. Pollard, B\. Gow, B\. Moody, S\. Horng, L\. A\. Celi, and R\. Mark\(2024\)MIMIC\-IV\.PhysioNet\.Note:Version 3\.1External Links:[Document](https://dx.doi.org/10.13026/kpb9-mt58),[Link](https://physionet.org/content/mimiciv/3.1/)Cited by:[§4\.1](https://arxiv.org/html/2609.38193#S4.SS1.p1.1),[§7](https://arxiv.org/html/2609.38193#S7.SS0.SSS0.Px2.p1.1)\.
- \[6\]A\. Johnson, L\. Bulgarelli, T\. Pollard, S\. Horng, L\. A\. Celi, and R\. Mark\(2023\)MIMIC\-IV clinical database demo\.PhysioNet\.Note:Version 2\.2, Open Database LicenseExternal Links:[Document](https://dx.doi.org/10.13026/dp1f-ex47),[Link](https://physionet.org/content/mimic-iv-demo/2.2/)Cited by:[Table 1](https://arxiv.org/html/2609.38193#S3.T1)\.
- \[7\]A\. E\. W\. Johnson, L\. Bulgarelli, L\. Shen, A\. Gayles, A\. Shammout, S\. Horng, T\. J\. Pollard, S\. Hao, B\. Moody, B\. Gow, L\. H\. Lehman, L\. A\. Celi, and R\. G\. Mark\(2023\)MIMIC\-IV, a freely accessible electronic health record dataset\.Scientific Data10\(1\),pp\. 1\.External Links:[Document](https://dx.doi.org/10.1038/s41597-022-01899-x)Cited by:[§4\.1](https://arxiv.org/html/2609.38193#S4.SS1.p1.1)\.
- \[8\]A\. Johnson, T\. Pollard, S\. Horng, L\. A\. Celi, and R\. Mark\(2023\)MIMIC\-IV\-Note: deidentified free\-text clinical notes\.PhysioNet\.Note:Version 2\.2External Links:[Document](https://dx.doi.org/10.13026/1n74-ne17),[Link](https://physionet.org/content/mimic-iv-note/2.2/)Cited by:[§4\.1](https://arxiv.org/html/2609.38193#S4.SS1.p1.1),[§7](https://arxiv.org/html/2609.38193#S7.SS0.SSS0.Px2.p1.1)\.
- \[9\]M\. G\. Kahn, T\. J\. Callahan, J\. Barnard, A\. E\. Bauck, J\. Brown, B\. N\. Davidson, H\. Estiri, C\. Goerg, E\. Holve, S\. G\. Johnson, S\. Liaw, M\. Hamilton\-Lopez, D\. Meeker, T\. C\. Ong, P\. Ryan, N\. Shang, N\. G\. Weiskopf, C\. Weng, M\. N\. Zozus, and L\. Schilling\(2016\)A harmonized data quality assessment terminology and framework for the secondary use of electronic health record data\.eGEMs \(Generating Evidence & Methods to improve patient outcomes\)4\(1\),pp\. 18\.External Links:[Document](https://dx.doi.org/10.13063/2327-9214.1244)Cited by:[§2](https://arxiv.org/html/2609.38193#S2.p3.1)\.
- \[10\]M\. Kallfelz, A\. Tsvetkova, T\. Pollard, M\. Kwong, G\. Lipori, V\. Huser, J\. Osborn, S\. Hao, and A\. Williams\(2021\)MIMIC\-IV demo data in the OMOP common data model\.PhysioNet\.Note:Version 0\.9External Links:[Document](https://dx.doi.org/10.13026/p1f5-7x35),[Link](https://physionet.org/content/mimic-iv-demo-omop/0.9/)Cited by:[Appendix A](https://arxiv.org/html/2609.38193#A1.p1.1),[§4\.3](https://arxiv.org/html/2609.38193#S4.SS3.p1.1)\.
- \[11\]S\. Kapoor and A\. Narayanan\(2023\)Leakage and the reproducibility crisis in machine\-learning\-based science\.Patterns4,pp\. 100804\.External Links:[Document](https://dx.doi.org/10.1016/j.patter.2023.100804),[Link](https://arxiv.org/abs/2207.07048)Cited by:[§1](https://arxiv.org/html/2609.38193#S1.p2.1)\.
- \[12\]R\. Liu, I\. Q\. Mohiuddin, A\. J\. Schoeffler, K\. Renduchintala, A\. Nayak, P\. L\. Vemu, S\. C\. Vedak, K\. C\. Black, J\. L\. Havlik, I\. Ogunmola, S\. P\. Ma, R\. Dhatt, and J\. H\. Chen\(2026\)PhysicianBench: evaluating LLM agents in real\-world EHR environments\.Note:Version 1External Links:2605\.02240,[Link](https://arxiv.org/abs/2605.02240v1)Cited by:[§1](https://arxiv.org/html/2609.38193#S1.p1.1),[§2](https://arxiv.org/html/2609.38193#S2.p1.1)\.
- \[13\]M\. McDermott and contributorsMEDS\-Transforms: build and run complex pipelines over MEDS datasets via simple parts\.Note:[https://github\.com/mmcdermott/MEDS\_transforms](https://github.com/mmcdermott/MEDS_transforms)Accessed September 6, 2026Cited by:[§2](https://arxiv.org/html/2609.38193#S2.p2.1)\.
- \[14\]Medical Event Data Standard Working GroupMEDS: schema definitions and python types\.Note:[https://github\.com/Medical\-Event\-Data\-Standard/meds](https://github.com/Medical-Event-Data-Standard/meds)Accessed September 6, 2026Cited by:[§2](https://arxiv.org/html/2609.38193#S2.p2.1)\.
- \[15\]MEDS\-DEV contributorsMEDS\-DEV: establishing reproducibility and comparability in health AI\.Note:[https://github\.com/Medical\-Event\-Data\-Standard/MEDS\-DEV](https://github.com/Medical-Event-Data-Standard/MEDS-DEV)Accessed September 6, 2026Cited by:[§2](https://arxiv.org/html/2609.38193#S2.p2.1)\.
- \[16\]L\. Mu, Z\. Huang, Y\. Gu, S\. Qin, S\. Zhang, and X\. Zhang\(2026\)EHRWorld: a patient\-centric medical world model for long\-horizon clinical trajectories\.Note:Version 1External Links:2602\.03569,[Link](https://arxiv.org/abs/2602.03569v1)Cited by:[§1](https://arxiv.org/html/2609.38193#S1.p1.1),[§2](https://arxiv.org/html/2609.38193#S2.p1.1)\.
- \[17\]Observational Health Data Sciences and Informatics\(2026\)ATHENA: OHDSI standardized vocabularies\.Note:[https://athena\.ohdsi\.org](https://athena.ohdsi.org/)Conversion inventory: v5\.0, 29\-AUG\-26\. Historical experiments retain their recorded vocabulary snapshotsCited by:[§2](https://arxiv.org/html/2609.38193#S2.p3.1),[§7](https://arxiv.org/html/2609.38193#S7.SS0.SSS0.Px2.p1.1)\.
- \[18\]OHDSIOMOP Common Data Model v5\.4\.Note:[https://ohdsi\.github\.io/CommonDataModel/cdm54\.html](https://ohdsi.github.io/CommonDataModel/cdm54.html)Accessed September 6, 2026Cited by:[§2](https://arxiv.org/html/2609.38193#S2.p2.1)\.
- \[19\]T\. Pollard, B\. E\. Moody, L\. H\. Lehman, B\. J\. Gow, C\. Fernandes, C\. Xie, A\. Johnson, R\. G\. Mark, and T\. Heldt\(2026\)PhysioNet as a global platform for biomedical research\.Nature Health1\(8\),pp\. 792–795\.External Links:[Document](https://dx.doi.org/10.1038/s44360-026-00096-z)Cited by:[§7](https://arxiv.org/html/2609.38193#S7.SS0.SSS0.Px2.p1.1)\.
- \[20\]P\. Renc, Y\. Jia, A\. E\. Samir, J\. Was, Q\. Li, D\. W\. Bates, and A\. Sitek\(2024\)Zero shot health trajectory prediction using transformer\.npj Digital Medicine7,pp\. 256\.External Links:[Document](https://dx.doi.org/10.1038/s41746-024-01235-0)Cited by:[§1](https://arxiv.org/html/2609.38193#S1.p1.1),[§2](https://arxiv.org/html/2609.38193#S2.p1.1)\.
- \[21\]R\. P\. van de Water, E\. Steinberg, M\. Wornow, P\. Rockenschaub, and M\. McDermott\(2025\)MIMIC\-IV demo data in the Medical Event Data Standard \(MEDS\)\.PhysioNet\.Note:Version 0\.0\.1External Links:[Document](https://dx.doi.org/10.13026/t2y8-ea41),[Link](https://physionet.org/content/mimic-iv-demo-meds/0.0.1/)Cited by:[Appendix A](https://arxiv.org/html/2609.38193#A1.p1.1),[§4\.3](https://arxiv.org/html/2609.38193#S4.SS3.p1.1)\.
- \[22\]Y\. Wang, Y\. Dai, R\. Wang, T\. Mehta, P\. Trivedi, T\. Vu, C\. T\. Lin, L\. Yang, Z\. Jiao, I\. Kamel,et al\.\(2026\)Integrating large language models for enhanced predictive analytics in healthcare\.npj Digital Medicine\.Cited by:[§2](https://arxiv.org/html/2609.38193#S2.p1.1)\.
- \[23\]J\. Xu, J\. Gallifant, A\. E\. W\. Johnson, and M\. B\. A\. McDermott\(2025\)ACES: automatic cohort extraction system for event\-stream datasets\.Note:Version 3; ICLR 2025External Links:2406\.19653,[Link](https://arxiv.org/abs/2406.19653)Cited by:[§2](https://arxiv.org/html/2609.38193#S2.p2.1)\.
- \[24\]X\. Yang, Z\. Zhong, S\. Collins, G\. Baird, X\. Wang, and Z\. Jiao\(2026\)Reliability stress tests and decision\-time routing for chest X\-ray vision\-language models\.In2026 IEEE/ACM Conference on Connected Health: Applications, Systems and Engineering Technologies \(CHASE\),pp\. 428–433\.Cited by:[§2](https://arxiv.org/html/2609.38193#S2.p1.1)\.

## Appendix AConverter comparison details

We used the MIMIC\-IV demonstration release of 100 patients, available under the Open Data Commons Open Database License\. The OMOP baseline wasmimic\-iv\-demo\-omopv0\.9\[[10](https://arxiv.org/html/2609.38193#bib.bib23)\], with OMOP CDM 5\.3\.1, a September 2020 vocabulary release, and 423,057 clinical rows\. It included an OHDSI Data Quality Dashboard report\. The MEDS baseline wasmimic\-iv\-demo\-medsv0\.0\.1\[[21](https://arxiv.org/html/2609.38193#bib.bib24)\], with 916,166 events in MEDS 0\.3\.3 format, produced by MEDS\-Transforms\. We measured both baselines from their published files\. EHR2Trace used the full\-dataset configuration on the same subset and produced 971,545 events\. This conversion also runs in continuous integration on every code change and is reproducible from the repository\.

Table[5](https://arxiv.org/html/2609.38193#S4.T5)measures source recovery as the presence of both a source table name and a row identifier\. Availability requires a separate timestamp; the reported share is the proportion of timed events for which availability time is later than event time\. Medication action types are the values used to distinguish ordering, dispensing, and administration\. The OMOP baseline records one drug type concept\. The MEDS baseline uses one medication code and retains order and administration keys on 99\.8% of medication events\.

Concept coverage uses each model’s own denominator: clinical rows for OMOP and events for MEDS\. We report source modules alongside output counts because the modules determine which tables and events are included\. Counts and coverage should be compared within matching modules\. Concept coverage also depends on the vocabulary release\.

The scriptexperiments/run\_converter\_comparison\.pyin the software repository writes the measurements toexperiments/results/converter\_comparison\.json; the table is generated from that record\.

## Appendix BTemporal experiment details

The cohort included 131,007 MIMIC\-IV patients, each represented by the first qualifying admission lasting at least 72 hours\. A patient received a positive mortality\-proxy label if the earliest recorded death date was no later than one day after discharge\. This yielded 4,185 positive cases \(3\.2%\)\. Features were log\-transformed code counts, including prior history and untimed attributes\. We used logistic regression with anℓ2\\ell\_\{2\}penalty,liblinear,C=1C=1, a maximum of 2,000 iterations, and seed 20260826\. The training and held\-out sets contained 104,858 and 13,052 patients, respectively\. The tuning set was unused\.

Each arm built its code index and fitted its model using only its training split\. Codes found only in held\-out patients were excluded from the feature set\. We calculated percentile bootstrap intervals from 2,000 resamples of held\-out patients\. Each resample used the same patients for every arm, yielding paired estimates of the AUROC increase relative to availability filtering\.

For cross\-rule evaluation, each fitted model retained the code vocabulary of its training rule\. Test histories were represented in that vocabulary, with zeros for codes absent under the test rule\. All nine training–test combinations at a given horizon were evaluated using the same bootstrap resamples, yielding paired estimates for the reported contrasts\. Results for matched training and test rules agreed with the single\-rule experiments to within 0\.001\. A mismatch in the other direction also reduced performance: the model trained with availability filtering achieved AUROC 0\.721 on backdated test histories \(A→\\rightarrowC\) at six hours\. This combination corresponds to no realistic deployment scenario and is omitted from Table[6](https://arxiv.org/html/2609.38193#S4.T6)\. The training deficit compares the two models under the availability\-filtered rule, the condition under which either would be deployed\.

相似文章

MiGHT-EHR:面向异质时序电子健康记录的多任务图变换器

arXiv cs.LG

本文介绍了MiGHT-EHR,一种针对异质时序EHR数据的多任务图变换器,联合建模临床实体、时间轨迹和任务依赖。在MIMIC-III和MIMIC-IV上,它在药物推荐、住院时长、死亡率和再入院预测方面均优于现有最先进方法。