Patient-Reported Survey Data Improve Prediction of Opioid Use Disorder
Summary
This research evaluates the enhancement of opioid use disorder prediction by integrating patient-reported survey data with electronic health records, demonstrating improved performance across multiple machine learning models.
View Cached Full Text
Cached at: 09/14/26, 08:37 AM
# Patient-Reported Survey Data Improve Prediction of Opioid Use Disorder
Source: [https://arxiv.org/html/2609.12224](https://arxiv.org/html/2609.12224)
\\institutes\\amiasuper
1Department of Applied Mathematics & Statistics, Stony Brook University, Stony Brook, NY, USA;\\amiasuper2Department of Computer Science, Stony Brook University, Stony Brook, NY, USA;\\amiasuper3Department of Biomedical Informatics, Stony Brook University, Stony Brook, NY, USA;\\amiasuper4Department of Psychiatry, Stony Brook Medicine, Stony Brook, NY, USA
Zihan Ding\\amiasuper2†Grace Han\\amiasuper3Yinan Liu\\amiasuper2Richard N\. Rosenthal\\amiasuper4Fusheng WangPhD\\amiasuper2,3‡
†††Both authors contributed equally\.††‡Corresponding author: Fusheng Wang\(fusheng\.wang@stonybrook\.edu\)\.††Analysis code is available athttps://github\.com/StonyBrookDB/AllOfUsOUDPrediction\.††This work was supported by the Patient\-Centered Outcomes Research Institute \(PCORI\) under Contract No\. ME\-2023C3\-35532\.## Abstract
Electronic health records \(EHRs\) may incompletely capture patient\-reported factors associated with opioid use disorder \(OUD\)\. We evaluated whether survey data improve prediction of a first recorded OUD diagnosis among 267,747 All of Us participants with documented opioid exposure, including 15,287 OUD cases\. We compared EHR\-only and EHR\+survey models across 6\-, 12\-, and 24\-month look\-back windows using logistic regression, random forest, XGBoost, LightGBM, multilayer perceptron, LSTM, GRU, and Transformer\. Survey augmentation improved PR\-AUC across all 24 model\-window combinations by 0\.0087–0\.0505; the best 24\-month LightGBM model improved from 0\.6219 to 0\.6603\. Survey coverage increased with longer windows and differed by OUD status \(24 months: 21\.7% OUD\-positive vs\. 60\.7% OUD\-negative\)\. Permutation analysis ranked survey features as the second most important information domain at 24 months in both evaluated models\. Patient\-reported data provide complementary predictive signals beyond structured EHRs while highlighting the importance of survey availability\.
## Introduction
Opioid use disorder \(OUD\) is characterized by a problematic pattern of opioid use that causes clinically significant impairment or distress\[[1](https://arxiv.org/html/2609.12224#bib.bib1)\]\. It remains a substantial public health burden in the United States, affecting an estimated 4\.0 million people aged 12 years or older and resulting in 44,654 overdose deaths involving opioids in 2025\[[18](https://arxiv.org/html/2609.12224#bib.bib2),[2](https://arxiv.org/html/2609.12224#bib.bib3)\]\. Identifying individuals at elevated risk before OUD is first documented in the clinical record may create opportunities for earlier clinical assessment and facilitate timely intervention, thereby helping to reduce the risk of subsequent opioid\-related harms, including overdose and mortality\.
Electronic health records \(EHRs\) provide rich longitudinal information \(e\.g\., diagnoses, medications, laboratory results, procedures, physiological measurements, and healthcare utilization\) and have become an important data source for machine learning \(ML\)\-based prediction of OUD and other opioid\-related outcomes\[[15](https://arxiv.org/html/2609.12224#bib.bib4)\]\. Prior prediction studies have largely relied on routinely collected clinical and administrative data, highlighting patients’ demographic characteristics, psychiatric and substance use histories, pain\-related clinical history, medication exposure, and patterns of healthcare utilization as predictive features\[[17](https://arxiv.org/html/2609.12224#bib.bib5)\]\. Our group has similarly used EHR histories to predict Opioid\-related disease among patients prescribed opioids and, more recently, systematically evaluated how the selection of diagnosis\-based features influences OUD prediction\[[9](https://arxiv.org/html/2609.12224#bib.bib6),[8](https://arxiv.org/html/2609.12224#bib.bib7),[7](https://arxiv.org/html/2609.12224#bib.bib8),[10](https://arxiv.org/html/2609.12224#bib.bib9),[5](https://arxiv.org/html/2609.12224#bib.bib10),[6](https://arxiv.org/html/2609.12224#bib.bib11)\]\. However, models based solely on routinely documented clinical data may capture only part of the risk profile associated with OUD\.
OUD reflects the interaction of clinical, behavioral, and social factors\. Among patients prescribed opioids, prior substance use disorders and psychiatric comorbidities \(e\.g\., depression, anxiety disorders, and personality disorders\) have been associated with subsequent OUD or problematic opioid use\[[12](https://arxiv.org/html/2609.12224#bib.bib12),[19](https://arxiv.org/html/2609.12224#bib.bib13)\]\. Beyond these clinical and behavioral characteristics, social conditions may further shape vulnerability to opioid\-related harms\. A recent umbrella review found consistent associations between OUD and broader social determinants of health \(SDOH\), including unemployment and adverse childhood experiences, with additional socioeconomic factors linked to opioid overdose\[[13](https://arxiv.org/html/2609.12224#bib.bib14)\]\. Yet such factors are often incompletely or inconsistently represented in EHR data, which can introduce missingness and misclassification when these records are used for research\[[3](https://arxiv.org/html/2609.12224#bib.bib15)\]\. As a result, EHRs may not fully reflect the broader context associated with OUD risk\.
Patient\-reported information can complement EHRs by capturing additional factors that may be underrepresented in clinical data\. TheAll of UsResearch Program is well\-suited to evaluate their added value because it combines longitudinal EHR data with participant\-reported surveys in a large and diverse cohort\[[14](https://arxiv.org/html/2609.12224#bib.bib16),[4](https://arxiv.org/html/2609.12224#bib.bib17)\]\. Participants complete structured questionnaires through the program’s online portal, providing information on health behaviors, substance use, socioeconomic and social conditions, physical and mental health, daily functioning, disability, and healthcare access that may not be routinely documented in the EHR\. Prior work has shown that patient\-reported information derived from preoperative questionnaires \(e\.g\., measures of mental health, pain, substance use, and socioeconomic circumstances\) can improve EHR\-based prediction of persistent opioid use after surgery\[[11](https://arxiv.org/html/2609.12224#bib.bib18)\]\. Persistent opioid use, however, is distinct from OUD, and whether survey information provides incremental predictive value for OUD beyond longitudinal EHR data remains unknown\.
In this study, we evaluated whether incorporating participant\-reported survey information improves prediction of a first recorded OUD diagnosis beyond EHR data alone\. We compared EHR\-only and EHR\+survey models across multiple machine learning \(ML\) approaches using 6\-, 12\-, and 24\-month look\-back windows, selected to capture progressively longer pre\-diagnostic histories\. These windows were chosen because prior claims\-based research found that OUD diagnosis occurred, on average, approximately 10–12 months after initial prescription opioid exposure among patients who subsequently developed OUD\[[16](https://arxiv.org/html/2609.12224#bib.bib19)\]\. Our primary objective was to determine whether patient\-reported information consistently improves OUD prediction beyond information already captured in the EHR\.
Our study makes three contributions\. First, it quantifies the incremental predictive value of participant\-reported information beyond longitudinal clinical and demographic features\. Second, it tests whether this value is consistent across linear, tree\-based, neural\-network, and sequential models and across multiple observation windows\. Third, it examines the information domains and individual survey questions that contribute most strongly to OUD prediction\.
## Methods
### Data source and study cohort
Data were obtained from theAll of UsResearch Program Curated Data Repository, with data available through October 1, 2023\. The study cohort comprised 267,747 participants with at least one documented exposure to an opioid medication, identified using the Anatomical Therapeutic Chemical \(ATC\) level 3 code N02A\. Of those, 15,287 participants had a recorded OUD diagnosis \(based on ICD\-9\-CM codes 304\.00–304\.03 or ICD\-10\-CM F11 code family\) and were classified as OUD\-positive; the remaining 252,460 were classified as OUD\-negative, yielding a case\-to\-control ratio of approximately 1:16\.5\. Baseline characteristics of the cohort are summarized in Table[1](https://arxiv.org/html/2609.12224#Sx3.T1)\.
Table 1:Baseline characteristics of the cohort, overall and by opioid use disorder \(OUD\) status\.VariableOverallOUD\-positiveOUD\-negativeNumber of patients267,74715,287252,460GenderFemale165,095 \(61\.7%\)7,562 \(49\.5%\)157,533 \(62\.4%\)Male97,451 \(36\.4%\)7,321 \(47\.9%\)90,130 \(35\.7%\)Others/Unknown5,201 \(1\.9%\)404 \(2\.6%\)4,797 \(1\.9%\)RaceAsian5,329 \(2\.0%\)73 \(0\.5%\)5,256 \(2\.1%\)Black or African American46,156 \(17\.2%\)3,696 \(24\.2%\)42,460 \(16\.8%\)White150,777 \(56\.3%\)7,449 \(48\.7%\)143,328 \(56\.8%\)More than one race11,700 \(4\.4%\)936 \(6\.1%\)10,764 \(4\.3%\)Others/Unknown49,531 \(18\.5%\)2,618 \(17\.1%\)46,913 \(18\.6%\)Prior opioid overdose1,199 \(0\.4%\)1,044 \(6\.8%\)155 \(0\.1%\)Any other substance use disorder25,226 \(9\.4%\)6,809 \(44\.5%\)18,417 \(7\.3%\)Alcohol use disorder19,807 \(7\.4%\)4,513 \(29\.5%\)15,294 \(6\.1%\)Cocaine use disorder8,147 \(3\.0%\)3,517 \(23\.0%\)4,630 \(1\.8%\)Any mental health condition110,990 \(41\.5%\)11,588 \(75\.8%\)99,402 \(39\.4%\)Depression83,287 \(31\.1%\)9,578 \(62\.7%\)73,709 \(29\.2%\)Bipolar disorder14,888 \(5\.6%\)3,471 \(22\.7%\)11,417 \(4\.5%\)Anxiety disorder81,678 \(30\.5%\)9,149 \(59\.8%\)72,529 \(28\.7%\)PTSD4,532 \(1\.7%\)822 \(5\.4%\)3,710 \(1\.5%\)Any pain diagnosis136,318 \(50\.9%\)11,363 \(74\.3%\)124,955 \(49\.5%\)Back pain75,319 \(28\.1%\)7,106 \(46\.5%\)68,213 \(27\.0%\)Chronic pain \(general\)89,665 \(33\.5%\)9,623 \(62\.9%\)80,042 \(31\.7%\)Muscle/musculoskeletal/joint pain36,063 \(13\.5%\)3,784 \(24\.8%\)32,279 \(12\.8%\)Benzodiazepine use31,046 \(11\.6%\)3,645 \(23\.8%\)27,401 \(10\.9%\)Pain score, mean \(SD\)3\.67 \(3\.02\)5\.80 \(2\.85\)3\.54 \(2\.98\)Insured \(self\-reported\)248,889 \(95\.3%\)13,772 \(93\.8%\)235,117 \(95\.4%\)Income<25k66,900 \(31\.4%\)8,068 \(69\.4%\)58,832 \(29\.2%\)25k–50k40,530 \(19\.0%\)1,850 \(15\.9%\)38,680 \(19\.2%\)50k–100k50,089 \(23\.5%\)1,160 \(10\.0%\)48,929 \(24\.3%\)\>100k55,374 \(26\.0%\)542 \(4\.7%\)54,832 \(27\.2%\)Stable housing concern44,630 \(16\.9%\)6,337 \(42\.5%\)38,293 \(15\.4%\)
### Temporal design
An index date was defined for each participant to anchor longitudinal feature extraction\. For OUD\-positive cases, the index date was the first recorded OUD diagnosis; for OUD\-negative controls, it was the most recent medical encounter available in the record\. Features were extracted from the 6, 12, and 24 months preceding the index date \(Figure[1](https://arxiv.org/html/2609.12224#Sx3.F1)\), representing progressively longer pre\-diagnostic histories\. These windows were informed in part by unpublished analyses from our group using the Health Facts database, in which 34\.12% of patients who later received an OUD diagnosis were diagnosed within 1 year and 54\.07% within 2 years of their first documented opioid medication\. Data recorded on or after the index date were excluded to prevent temporal leakage\.
Figure 1:Temporal design for OUD\-positive participants and OUD\-negative controls across 6\-, 12\-, and 24\-month look\-back windows\.
### EHR\-Derived Features
The analysis incorporated EHR\-derived features \(i\.e\., demographics, diagnoses, medications, laboratory measurements, physical measurements, and clinical observations\)\. Candidate concepts within each domain were selecteda prioribased on their clinical relevance to OUD\. All features were constructed from information recorded within the specified look\-back window before the index date\. To reduce sparsity, clinical concepts observed in fewer than 0\.5% of participants within each look\-back window were excluded \(Table[2](https://arxiv.org/html/2609.12224#Sx3.T2)\)\. Feature\-inclusion thresholds were calculated using the full cohort before the train/validation/test split and were based only on feature frequency, without using OUD status\.
Table 2:Candidate and retained concepts/questions by data source across the 6\-, 12\-, and 24\-month look\-back windows\.Demographics\. Demographic features included age, gender, race, ethnicity, sex at birth, and self\-reported category \(a composite race/ethnicity variable\)\.
Conditions\. Condition features included diagnoses such as low back pain, sciatica, fibromyalgia, chronic pain syndrome, alcohol abuse, psychoactive substance abuse, major depressive disorder, generalized anxiety disorder, and posttraumatic stress disorder\. These features were represented by their occurrence counts within each look\-back window\. Participants with no recorded occurrence of a given concept were assigned a count of zero\.
Medications\. Medication features included opioid analgesics \(e\.g\., oxycodone, hydromorphone, morphine, fentanyl, and tramadol\), non\-opioid analgesics \(e\.g\., ibuprofen, ketorolac, and celecoxib\), muscle relaxants \(e\.g\., cyclobenzaprine, tizanidine\), benzodiazepines \(e\.g\., lorazepam, diazepam\), and antidepressants \(e\.g\., sertraline, duloxetine\)\. These features were represented using the same occurrence\-count approach described above\.
Laboratory and physical measurements\. Laboratory features included standard metabolic, liver, lipid, and hematology panels \(e\.g\., creatinine, alanine aminotransferase, total cholesterol, and complete blood count\), as well as a urine toxicology panel \(e\.g\., opiates, fentanyl, oxycodone, methadone, buprenorphine, cocaine, cannabinoids, benzodiazepines, and amphetamines\)\. Physical measurement features included body weight, height, body mass index, systolic and diastolic blood pressure, heart rate, and waist and hip circumference\. Any values recorded in different units were first converted to a common unit and then summarized within each look\-back window using three features: count, mean value, and most recent value\. When no measurement was available, the count was set to zero, while the mean and most recent values were retained as missing\.
Clinical observations\. Clinical observation features included both numeric and occurrence\-based concepts\. Numeric observations \(e\.g\., pain score, tobacco smoking, depression screening assessment\) were summarized using count, mean, and most recent value, following the same approach as laboratory and physical measurement features\. Occurrence\-based observations \(e\.g\., housing instability, documented abuse, treatment noncompliance, and self\-reported alcohol use\) were represented using occurrence counts within each look\-back window, with a count of zero assigned when no recorded occurrence was present\.
### Participant\-reported survey data
Survey Sources and Content\. Participant\-reported information was obtained from fourAll of UsResearch Program surveys:The Basics,Lifestyle,Overall Health, andSocial Determinants of Health16,17\.The Basicsis administered early in participation and captures demographic and socioeconomic characteristics, including employment, insurance, housing, and home\-life information\.LifestyleandOverall Healthbecome available after completion ofThe Basics\.Lifestyleassesses health behaviors including tobacco, alcohol, and recreational drug use, whereasOverall Healthassesses general health, daily functioning, pain, and physical and mental health\.Social Determinants of Healthis a follow\-up survey available after completion of the baseline surveys and captures broader social and environmental factors, including neighborhood characteristics, social relationships, stress, discrimination, loneliness, and social support\.
Survey Availability\. Only survey responses recorded before the index date and within the corresponding 6\-, 12\-, or 24\-month look\-back window were eligible for feature construction\. Because the source surveys differ in administration sequence, eligibility, and timing of completion, survey information was not uniformly available across participants or prediction windows\. We therefore characterized both cohort\-level survey coverage and respondent\-level completion behavior\. Survey coverage was defined as the proportion of participants with at least one eligible survey interaction within the corresponding look\-back window, including either an answered item or an explicitly recorded skip or decline response\. Among these respondents, we then summarized the number of questions answered, the number of questions explicitly skipped or declined, and the number of days between the most recent eligible survey response and the index date\. These measures separately characterize whether survey data were present, how much information was available, and how recently it was collected\.
Survey\-Derived Features\. Because the objective was to determine whether participant\-reported information provides predictive value beyond routinely documented EHR\-derived features, eligible survey questions were retained even when they reflected constructs also represented in the EHR, such as pain or substance use\. Among participants who contributed any survey data within a given look\-back window \(individuals with at least one valid survey response in that window\), questions with below 50% of response rates were excluded\. The remaining questions were encoded using binary indicators for individual response options\. Single\-response questions were represented using one\-hot encoding, whereas multiple\-response questions were represented using multi\-hot encoding\. A value of 1 indicated that a participant selected the corresponding response option, 0 indicated that the question was answered but the option was not selected, and missing indicated that no usable response was available for that item within the look\-back window\. The number of retained survey features consequently varied by observation window, with 30, 30, and 108 retained survey features at 6, 12, and 24 months, respectively \(Table[2](https://arxiv.org/html/2609.12224#Sx3.T2)\)\.
### Model development
We evaluated two complementary modeling approaches\. Tabular models used patient\-level features summarized over each look\-back window, whereas sequential models incorporated the temporal ordering of clinical events to assess whether longitudinal information improved OUD prediction\.
Cohort splitting and preprocessing\.Across all look\-back windows, participants were stratified by OUD status and split at the patient level into training \(68%\), validation \(12%\), and independent test \(20%\) sets\. The same patient splits were used for the corresponding EHR\-only and EHR\+survey models to enable paired comparisons\. The validation set was used for early stopping and hyperparameter selection, where applicable, as well as for decision\-threshold selection\. All preprocessing parameters, including those used for imputation and standardization, were estimated from the training set only and then applied unchanged to the validation and test sets\.
Class imbalance\.We evaluated four approaches to class imbalance: no adjustment, cost\-sensitive class weighting, synthetic minority oversampling using the synthetic minority over\-sampling technique \(SMOTE\), and random undersampling\. The comparison was performed using XGBoost with EHR and survey features at the 24\-month look\-back window\. Strategies were compared using validation area under the precision\-recall curve \(PR\-AUC\)\. Unadjusted training achieved the highest validation PR\-AUC \(0\.6533\), followed by class weighting \(0\.6475\), random undersampling \(0\.5906\), and SMOTE \(0\.4075\); therefore, final models were trained without class\-imbalance adjustment\.
Tabular models\.Five tabular models were evaluated: logistic regression, random forest, XGBoost, LightGBM, and a multilayer perceptron \(MLP\)\. Logistic regression served as the linear baseline\. Missing EHR numeric features were median\-imputed and standardized for logistic regression and the MLP, whereas random forest inputs were median\-imputed without standardization\. Missing survey feature values were also median\-imputed\. XGBoost, LightGBM, and the MLP used early stopping based on validation PR\-AUC\. Logistic regression was trained using the L\-BFGS solver with L2 penalty andC=1\.0C=1\.0\. The random forest consisted of 400 trees grown without a maximum depth\. XGBoost and LightGBM were each trained for up to 5,000 boosting rounds with early stopping after 200 rounds without improvement in validation PR\-AUC\. Both used a learning rate of 0\.03 and row and column subsampling rates of 0\.75 and 0\.70, respectively\. XGBoost additionally used a maximum depth of 5 and a minimum child weight of 16, whereas LightGBM used 31 leaves per tree\. The MLP consisted of two hidden layers with 256 and 128 units, and a dropout rate of 0\.3\. It was trained using the Adam optimizer and binary cross\-entropy loss with a batch size of 2,048 for up to 60 epochs, with early stopping based on validation PR\-AUC\.
Sequential models\.Three sequential architectures were evaluated: long short\-term memory \(LSTM\), gated recurrent unit \(GRU\), and Transformer models\. For each participant, the longitudinal input consisted of theKKmost recent dates with recorded clinical activity within the corresponding look\-back window\. Sequence lengths were set toK=20K=20, 30, and 50 for the 6\-, 12\-, and 24\-month windows, respectively\. These thresholds were chosen so that approximately 90% of participants had no more thanKKrecorded activity dates within each window\. This resulted in coverage of 93\.2%, 90\.9%, and 89\.7%, respectively\. Participants with fewer thanKKdates were zero\-padded, while those with more thanKKdates were truncated to theKKmost recent dates\. Padded positions were masked during model training\.
Each time step represented clinical events recorded on a single date\. The input for each date was a vector of clinical concept counts, with each feature corresponding to a diagnosis, medication, laboratory test, physical measurement, or clinical observation code and its value indicating the number of times that code was recorded that day\. Laboratory and physical measurement codes captured the occurrence of these events, while their numeric values were summarized across the full look\-back window using the mean and most recent observed value and included in the static feature vector along with demographic characteristics\. For EHR\+survey models, the same survey\-derived features used in the tabular models were added to the static feature vector\.
The LSTM and GRU processed the sequence chronologically and used the final hidden state as the patient\-level temporal representation\. The Transformer projected each time\-step vector to a 128\-dimensional embedding, added learnable positional embeddings, and processed the sequence using a two\-layer Transformer encoder\. Transformer outputs were aggregated into a single patient\-level temporal representation using masked mean pooling over non\-padded time steps\. For all three architectures, the temporal representation was concatenated with an encoded representation of the static features, obtained via a fully\-connected layer, before the final prediction layer\. EHR\-only and EHR\+survey models used the same longitudinal inputs and architecture, differing only in the addition of survey\-derived features to the static vector\.
The three sequential models shared a common overall architecture, differing primarily in their temporal encoders\. The LSTM and GRU used a hidden dimension of 128, while the Transformer used a 128\-dimensional embedding with four attention heads and two encoder layers\. In all three models, the temporal representation was concatenated with a 64\-dimensional static representation produced by a fully connected layer with ReLU activation and a dropout rate of 0\.3\. The combined representation was then passed through a two\-layer prediction head\. All models were trained using the Adam optimizer with a learning rate of1×10−31\\times 10^\{\-3\}and weight decay of1×10−51\\times 10^\{\-5\}, using binary cross\-entropy loss and a batch size of 512\. Training was performed for up to 30 epochs, with early stopping after five epochs without improvement in validation PR\-AUC\.
### Evaluation
Model performance was evaluated on the held\-out test set using PR\-AUC, area under the receiver operating characteristic curve \(ROC\-AUC\), positive\-class precision, recall, and F1\-score\. PR\-AUC was considered the primary performance metric due to class imbalance\. For each model, the classification threshold was selected on the validation set to maximize positive\-class F1 across 181 thresholds from 0\.05 to 0\.95\. The selected threshold was then applied unchanged to the test set\. F1\+ denotes the F1 score for the positive OUD class\.
Feature importance was evaluated at the 6\-, 12\-, and 24\-month look\-back windows at two levels: broad information domains and individual survey questions\. Both analyses used permutation importance on the held\-out test set and were performed for the LightGBM and GRU models across all look\-back windows\. Within each look\-back window, only the feature or feature group being evaluated was permuted across participants, while all other features remained unchanged\. Importance was defined as the mean decrease in PR\-AUC after permutation, averaged over five repetitions\. For the domain\-level analysis, features were grouped into demographics, conditions, medications, laboratory measurements, physical measurements, clinical observations, and survey\-derived features\. All features within a domain were jointly permuted to estimate the model’s reliance on that source of information\. For the survey question\-level analysis, each survey question was evaluated separately\. When a question was represented by multiple encoded response categories, all corresponding columns were permuted together as a single unit\. This analysis identified the individual survey questions that contributed most to model performance at each look\-back window\. The same permutation framework was applied across tabular and sequential models\. For sequential models, time\-varying EHR modalities were permuted at the patient level by reassigning entire temporal trajectories rather than individual time points, thereby preserving within\-patient temporal structure\.
## Results
### Survey Availability
Survey coverage increased with longer look\-back windows in both groups but was consistently lower among OUD\-positive participants \(Table[3](https://arxiv.org/html/2609.12224#Sx4.T3)\)\. The proportion of participants with at least one eligible survey response increased from 9\.5% at 6 months to 21\.7% at 24 months among OUD\-positive participants, compared with 21\.6% to 60\.7% among OUD\-negative participants\. Among respondents, the median number of questions answered was similar between groups at the 6\- and 12\-month windows, whereas at 24 months OUD\-negative participants contributed a larger volume of survey responses \(median 77 vs\. 28 questions\)\. Explicitly skipped or declined questions were uncommon across all windows, with median values of 0–1 per respondent\. The median interval between the most recent eligible survey response and the index date also increased with longer look\-back windows, from 51 and 63 days at 6 months to 217 and 270 days at 24 months for OUD\-positive and OUD\-negative participants, respectively\.
Table 3:Survey coverage and completion behavior by OUD status and look\-back window\.Note\.Survey coverage was calculated among all OUD\-positive \(n=15,287n=15\{,\}287\) and OUD\-negative \(n=252,460n=252\{,\}460\) participants at each look\-back window\.
### Model Evaluation
Across the evaluated models, ROC\-AUC ranged from 0\.8180 to 0\.9430, while PR\-AUC ranged from 0\.3239 to 0\.6603, reflecting the pronounced class imbalance in the study cohort\. Among the eight models, LightGBM and XGBoost achieved the strongest overall performance\. At the 24\-month look\-back window, LightGBM achieved the highest PR\-AUC of 0\.6603 with survey features, followed by XGBoost at 0\.6576\. Across all models, PR\-AUC increased with longer look\-back windows \(Table[4](https://arxiv.org/html/2609.12224#Sx4.T4)\)\.
Table 4:Predictive performance by model, feature set, and look\-back window \(months\)\.Adding survey\-derived features improved PR\-AUC in all 24 model–window combinations, with absolute gains ranging from 0\.0087 to 0\.0505 \(Table[5](https://arxiv.org/html/2609.12224#Sx4.T5)\)\. Similar improvements were observed in ROC\-AUC across linear, ensemble, gradient\-boosting, and neural\-network models, with gains ranging from 0\.0013 to 0\.0410\. The incremental contribution of survey information was greatest at the 24\-month look\-back window, where all eight models achieved their largest PR\-AUC improvement, with gains ranging from 0\.0181 to 0\.0505\. However, the magnitude of improvement did not increase monotonically across look\-back windows for all models\.
Table 5:Absolute increase in PR\-AUC and ROC\-AUC after adding survey\-derived features, by model and look\-back window \(months\)\. In each column,boldmarks the largest gain andunderlinethe second largest\.Model complexity did not consistently translate into better predictive performance\. Despite their ability to model nonlinear relationships or temporal structure, the LSTM, GRU, and Transformer models generally performed below the gradient\-boosting models, particularly at the longer look\-back windows\. Nevertheless, survey\-derived features improved performance even for the neural and sequential models, suggesting that their contribution was complementary to rather than dependent on a specific modeling architecture\.
Domain\-level permutation analysis showed that the contribution of survey\-derived features increased with longer look\-back windows\. In LightGBM, permutation of survey features resulted in PR\-AUC decreases of 0\.084, 0\.156, and 0\.172 at the 6\-, 12\-, and 24\-month windows, respectively\. The corresponding decreases in the GRU were 0\.046, 0\.073, and 0\.234, with the largest decrease observed at 24 months\. In both models, survey features ranked as the second most important information domain at 24 months, behind laboratory measurements\.
Figure 2:Permutation feature importance for the top 5 predictors of each model — LGB \(left\) and GRU \(right\) — at 6, 12, and 24 months\. Bars show mean decrease in PR\-AUC across five permutations; error bars indicate SD\.At the individual question level, the most influential survey features varied across models and look\-back windows, although several questions consistently ranked among the top contributors \(Figure[2](https://arxiv.org/html/2609.12224#Sx4.F2)\)\. Recreational drug use, difficulty concentrating, ability to complete errands independently, and employment status repeatedly appeared among the highest\-ranking questions\. Other recurring features included measures of alcohol and tobacco use, pain and general health, income, and education\. Overall, the highest\-ranking survey questions spanned behavioral, functional, socioeconomic, and general health domains\.
## Discussion
This study examined whether participant\-reported survey information adds predictive value beyond structured EHR data for identifying patients at risk of a first recorded OUD diagnosis\. Survey augmentation improved discrimination across all modeling approaches and observation windows, indicating that the added signal was not specific to one algorithmic family\. Although the gains were modest, their consistency supports patient\-reported information as a useful complement to EHR\-based OUD prediction\.
Feature\-importance analyses further clarified the source of this added value\. At the 24\-month window, survey\-derived features ranked among the most important information sources in both evaluated model architectures\. Influential questions spanned substance use, functional limitations, employment, income, pain, and general health, suggesting that the predictive signal was distributed across behavioral, functional, socioeconomic, and health\-related factors rather than driven by a single construct\. This pattern is consistent with the multifactorial nature of OUD risk and highlights information that may be incompletely represented in routine clinical records\.
The LSTM, GRU, and Transformer models generally did not outperform the gradient\-boosting approaches despite explicitly modeling event sequences\. Much of the useful longitudinal information may already have been captured by summary features such as event frequency, recent values, and average measurements\. In the sequential models, laboratory and physical measurement events were represented by their occurrence over time, while their numeric values were summarized across the look\-back window as static features\. This design may have limited the models’ ability to capture changes in these values over time\. More broadly, the results suggest that greater model complexity does not necessarily improve prediction when summary features already capture much of the relevant clinical history\.
The stronger contribution of survey information at longer observation windows should be interpreted in the context of data availability\. Longer look\-back periods increased both the opportunity to observe survey responses and the number of retained survey features, which rose from 30 at the 6\- and 12\-month windows to 108 at 24 months\. Thus, the larger gains at 24 months may reflect both accumulation of relevant historical information and greater availability of the survey modality itself\. The optimal observation horizon may therefore depend on both the timing of prior information and the likelihood that each data source is available\. The strong survey contribution at longer windows also suggests that patient\-reported information may provide useful signals for earlier risk assessment and future prevention\-oriented screening strategies\.
Survey coverage also differed markedly by OUD status, with OUD\-negative participants having greater coverage across all three look\-back windows\. This raises the possibility that missing survey data are informative rather than random\. Survey availability may reflect differences in program engagement, enrollment duration, eligibility, completion behavior, or timing relative to the index date\. Some of the observed gain may therefore arise from patterns of survey availability in addition to response content\. Future ablation analyses that separate availability\-related features from substantive survey responses will be important for clarifying these contributions\.
Several limitations should be considered\. First, OUD status was defined using recorded diagnostic codes and therefore depended on clinical recognition and documentation, which may have led to misclassification of participants with unrecorded OUD\. Second, regarding the temporal design, index dates were defined differently for cases and controls, potentially creating differences in observable history, follow\-up, and opportunities for survey completion\. Third, regarding survey availability, participation was incomplete and differed by outcome status, introducing possible selection effects and informative missingness\. Fourth, feature filtering was performed before data splitting, allowing the overall feature distribution to inform feature retention, although OUD status was not used\. Given the large cohort, this was unlikely to substantially alter which features were retained\. Finally, regarding evaluation, performance was assessed using an internal held\-out test set and focused primarily on discrimination\. External validation, calibration, subgroup performance, and prospective evaluation are needed before clinical use\.
Overall, these findings show both the value and the complexity of integrating patient\-reported information with longitudinal EHR data\. Survey data provided additional predictive signal, but their contribution depends not only on response content, but also on when and for whom those data are available\. Future work should separate survey content from availability effects, assess calibration and fairness, evaluate external generalizability, and determine whether these data improve clinically meaningful OUD screening and risk\-assessment workflows\.
## Conclusion
Patient\-reported survey information consistently improved OUD prediction beyond structured EHR data across eight modeling approaches and three look\-back windows\. These findings suggest that behavioral, social, functional, and self\-reported health information captures predictive signals that are not fully represented in routine clinical records\. More broadly, integrating patient\-reported data with longitudinal EHRs may strengthen multimodal risk prediction by providing a more complete view of patient context, while also highlighting the need to account for differences in survey availability and completion\. With further validation, such approaches may support earlier and more comprehensive OUD risk assessment and prevention\-oriented screening\.
## References
- \[1\]CDC\(2024\)Overdose prevention: opioid use disorder: diagnosis\.Internet\.External Links:[Link](https://www.cdc.gov/overdose-prevention/hcp/clinical-care/opioid-use-disorder-diagnosis.html)Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p1.1)\.
- \[2\]CDC\(2026\)NCHS pressroom: u\.s\. overdose deaths decrease for third consecutive year in 2025\.Internet\.External Links:[Link](https://www.cdc.gov/nchs/pressroom/releases/20260513.html)Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p1.1)\.
- \[3\]L\. A\. Cook, J\. Sachs, and N\. G\. Weiskopf\(2021\)The quality of social determinants data in the electronic health record: a systematic review\.J Am Med Inform Assoc29,pp\. 187–96\.Note:doi:10\.1093/jamia/ocab199Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p3.1)\.
- \[4\]R\. M\. Cronin, R\. N\. Jerome, B\. Mapes, R\. Andrade, R\. Johnston, J\. Ayala,et al\.\(2019\)Development of the initial surveys for the all of us research program\.Epidemiology30,pp\. 597–608\.Note:doi:10\.1097/EDE\.0000000000001028Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p4.1)\.
- \[5\]Z\. Ding, X\. Dong, Y\. Liu, T\. Ma, X\. Zhao, R\. Wong,et al\.\(2024\)HIBERT: a hybrid clustering BERT for interpretable opioid overdose risk prediction\.AMIA Annu Symp Proc2024,pp\. 303–12\.Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p2.1)\.
- \[6\]Z\. Ding, Y\. Liu, T\. Ma, R\. Wong, G\. Leibowitz, B\. Littenberg,et al\.\(2026\)A comparative study of feature selection paradigms for structured ehr diagnosis features in oud prediction\.Technical reportManuscript under review for the AMIA 2026 Annual Symposium\.Note:arXiv:2608\.04180\. doi:10\.48550/ARXIV\.2608\.04180External Links:[Link](https://arxiv.org/abs/2608.04180)Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p2.1)\.
- \[7\]X\. Dong, J\. Deng, W\. Hou, S\. Rashidian, R\. N\. Rosenthal, M\. Saltz,et al\.\(2021\)Predicting opioid overdose risk of patients with opioid prescriptions using electronic health records based on temporal deep learning\.J Biomed Inform116,pp\. 103725\.Note:doi:10\.1016/j\.jbi\.2021\.103725Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p2.1)\.
- \[8\]X\. Dong, J\. Deng, S\. Rashidian, K\. Abell\-Hart, W\. Hou, R\. N\. Rosenthal,et al\.\(2021\)Identifying risk of opioid use disorder for patients taking opioid medications with deep learning\.J Am Med Inform Assoc28,pp\. 1683–93\.Note:doi:10\.1093/jamia/ocab043Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p2.1)\.
- \[9\]X\. Dong, S\. Rashidian, Y\. Wang, J\. Hajagos, X\. Zhao, R\. N\. Rosenthal,et al\.\(2019\)Machine learning based opioid overdose prediction using electronic health records\.AMIA Annu Symp Proc2019,pp\. 389–98\.Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p2.1)\.
- \[10\]X\. Dong, R\. Wong, W\. Lyu, K\. Abell\-Hart, J\. Deng, Y\. Liu,et al\.\(2023\)An integrated LSTM\-HeteroRGNN model for interpretable opioid overdose risk prediction\.Artif Intell Med135,pp\. 102439\.Note:doi:10\.1016/j\.artmed\.2022\.102439Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p2.1)\.
- \[11\]A\. M\. Giladi, M\. M\. Shipp, K\. K\. Sanghavi, G\. Zhang, S\. Gupta, K\. E\. Miller,et al\.\(2023\)Patient\-reported data augment prediction models of persistent opioid use after elective upper extremity surgery\.Plast Reconstr Surg152,pp\. 358e–66e\.Note:doi:10\.1097/PRS\.0000000000010297Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p4.1)\.
- \[12\]J\. Klimas, L\. Gorfinkel, N\. Fairbairn, L\. Amato, K\. Ahamad, S\. Nolan,et al\.\(2019\)Strategies to identify patient risks of prescription opioid addiction when initiating opioids for pain: a systematic review\.JAMA Netw Open2,pp\. e193365\.Note:doi:10\.1001/jamanetworkopen\.2019\.3365Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p3.1)\.
- \[13\]L\. B\. Loeffel, H\. R\. Kwak, J\. Shin, M\. Lee, D\. J\. Jester, M\. Dawes,et al\.\(2026\)Associations among social determinants of health and opioid use disorder and overdose: an umbrella review\.Am J Addict35,pp\. 339–61\.Note:doi:10\.1111/ajad\.70133Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p3.1)\.
- \[14\]A\. H\. Ramirez, L\. Sulieman, D\. J\. Schlueter, A\. Halvorson, J\. Qian, F\. Ratsimbazafy,et al\.\(2022\)The all of us research program: data quality, utility, and diversity\.Patterns3,pp\. 100570\.Note:doi:10\.1016/j\.patter\.2022\.100570Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p4.1)\.
- \[15\]C\. R\. Ramírez Medina, J\. Benitez\-Aurioles, D\. A\. Jenkins, and M\. Jani\(2025\)A systematic review of machine learning applications in predicting opioid associated adverse events\.npj Digit Med8,pp\. 30\.Note:doi:10\.1038/s41746\-024\-01312\-4Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p2.1)\.
- \[16\]A\. J\. Schoenfeld, S\. Jeyakumar, J\. Morlando Geiger, N\. Princic, M\. Moynihan, H\. Varker,et al\.\(2026\)The incidence of opioid use disorder among people with acute and chronic pain managed with prescription opioids in the united states: economic and societal burden\.J Med Econ29,pp\. 1355–71\.Note:doi:10\.1080/13696998\.2026\.2655086Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p5.1)\.
- \[17\]S\. L\. Song, H\. G\. Dandapani, R\. S\. Estrada, N\. W\. Jones, E\. A\. Samuels, and M\. L\. Ranney\(2024\)Predictive models to assess risk of persistent opioid use, opioid use disorder, and overdose\.J Addict Med18,pp\. 218–39\.Note:doi:10\.1097/ADM\.0000000000001276Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p2.1)\.
- \[18\]Substance Abuse and Mental Health Services Administration\(2026\)Key substance use and mental health indicators in the united states: results from the 2025 national survey on drug use and health\.Technical reportSubstance Abuse and Mental Health Services Administration,Rockville, MD\.Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p1.1)\.
- \[19\]S\. M\. van Rijswijk, M\. H\. C\. T\. van Beek, G\. M\. Schoof, A\. H\. Schene, M\. Steegers, and A\. F\. Schellekens\(2019\)Iatrogenic opioid use disorder, chronic pain and psychiatric comorbidity: a systematic review\.Gen Hosp Psychiatry59,pp\. 37–50\.Note:doi:10\.1016/j\.genhosppsych\.2019\.04\.008Cited by:[Introduction](https://arxiv.org/html/2609.12224#Sx2.p3.1)\.Similar Articles
A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction
This paper compares five feature selection methods for EHR diagnosis codes in opioid use disorder prediction, finding that NTK sensitivity offers the best accuracy-stability balance while LLM-guided selection adds complementary clinical signals.
Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder
The study systematically assesses algorithmic fairness in machine learning models for predicting treatment retention in medication for opioid use disorder, finding performance gaps across patient subgroups and evaluating bias mitigation techniques with trade-offs.
Population Health-Based Machine Learning Reveals Associations Between Psychosocial Factors and Chronic Kidney Disease
The paper presents a two-part study using machine learning on large-scale telehealth data to classify self-reported chronic kidney disease status and identify key psychosocial and medical predictors, achieving balanced accuracy around 72-76% with SHAP analysis for interpretability.
Machine Learning-Based Pre-Test Risk Stratification for PCR-Confirmed Chlamydia Using Patient-Reported Data and Urine Biomarkers
This study evaluates machine learning models for pre-test risk stratification of Chlamydia trachomatis infection using non-invasive patient-reported data and urine biomarkers, demonstrating moderate predictive performance and the complementary value of both data types.
Leveraging Physiological Signals to Predict Exam Outcomes with Machine Learning
This study investigates machine learning models to predict exam outcomes using physiological data such as electrodermal activity, heart rate, and skin temperature, finding that both deep learning approaches and simpler models like random forests can be effective.