Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder

arXiv cs.LG Papers

Summary

The study systematically assesses algorithmic fairness in machine learning models for predicting treatment retention in medication for opioid use disorder, finding performance gaps across patient subgroups and evaluating bias mitigation techniques with trade-offs.

arXiv:2609.22113v1 Announce Type: new Abstract: Persistent low retention and completion rates in medications for opioid use disorder (MOUD) have driven the use of machine learning (ML) models to predict retention and identify patients at risk of premature discontinuation. However, the fairness of these models across patient populations remains largely unexplored, raising concerns about their application in treatment decision support. This study systematically assesses algorithmic fairness in ML models for predicting MOUD retention and premature discontinuation and investigates the effectiveness of bias mitigation techniques. Using the cross-sectional Treatment Episode Data Set-Discharges (TEDS-D), which includes treatment episodes for individuals in the U.S. discharged between 2015 and 2019, we trained four ML models to predict premature treatment discontinuation and retention beyond 180 days among individuals receiving outpatient MOUD. We evaluated overall performance and subgroup-level error rates across patient subgroups defined by race, ethnicity, age, and sex, complemented by model explanation analyses. We further assessed bias mitigation techniques and their effects on both fairness and predictive performance. Our findings demonstrate that ML models for MOUD outcome prediction can exhibit subgroup-level performance gaps even when overall predictive performance appears acceptable and that bias mitigation can reduce, but not fully eliminate, these gaps without trade-offs. By demonstrating the importance of fairness-aware evaluation and transparent reporting of subgroup performance, this study provides practical insights for the responsible and context-sensitive use of ML models for risk stratification and care prioritization in MOUD treatment settings.
Original Article
View Cached Full Text

Cached at: 09/22/26, 09:11 AM

# Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder
Source: [https://arxiv.org/html/2609.22113](https://arxiv.org/html/2609.22113)
Persistent low retention and completion rates in medications for opioid use disorder \(MOUD\) have driven the use of machine learning \(ML\) models to predict retention and identify patients at risk of premature discontinuation\. However, the fairness of these models across patient populations remains largely unexplored, raising concerns about their application in treatment decision support\. This study systematically assesses algorithmic fairness in ML models for predicting MOUD retention and premature discontinuation and investigates the effectiveness of bias mitigation techniques\. Using the cross\-sectional Treatment Episode Data Set–Discharges \(TEDS\-D\), which includes treatment episodes for individuals in the U\.S\. discharged between 2015 and 2019, we trained four ML models to predict premature treatment discontinuation and retention beyond 180 days among individuals receiving outpatient MOUD\. We evaluated overall performance and subgroup\-level error rates across patient subgroups defined by race, ethnicity, age, and sex, complemented by model explanation analyses\. We further assessed pre\-processing, in\-processing, and post\-processing bias mitigation techniques and their effects on both fairness and predictive performance\. The models exhibited substantial performance differences across patient subgroups, including overestimation of the likelihood of premature discontinuation for Black patients and of treatment retention beyond 180 days for older patients\. Model explanation analyses further identified race and age as influential predictors, but their impacts on model predictions varied substantially across patient subgroups\. Bias mitigation strategies reduced specific fairness gaps but often introduced trade\-offs, such as increased error rates for other subgroups or reductions in overall predictive performance\. These findings demonstrate that ML models for MOUD outcome prediction can exhibit subgroup\-level performance gaps even when overall predictive performance appears acceptable and that bias mitigation can reduce, but not fully eliminate, these gaps without trade\-offs\. By demonstrating the importance of fairness\-aware evaluation and transparent reporting of subgroup performance, this study provides practical insights for the responsible and context\-sensitive use of ML models for risk stratification and care prioritization in MOUD treatment settings\.

Tongnian WangEmail:[tongnian\-wang@utc\.edu](mailto:[email protected])Affiliation:Gary W\. Rollins College of Business, The University of Tennessee at Chattanooga, Chattanooga, Tennessee, USACarolina Vivas\-ValenciaEmail:[carolina\.vivasvalencia@utsa\.edu](mailto:[email protected])Affiliation:Department of Biomedical Engineering and Chemical Engineering, The University of Texas at San Antonio, San Antonio, Texas, USACici BauerEmail:[cici\.x\.bauer@uth\.tmc\.edu](mailto:[email protected])Affiliation:Department of Biostatistics and Data Science, School of Public Health, The University of Texas Health Science Center at Houston, Houston, Texas, USAAffiliation:Center for Spatial\-Temporal Modeling for Applications in Population Sciences, School of Public Health, The University of Texas Health Science Center at Houston, Houston, Texas, USAYanmin GongEmail:[yanmin\.gong@tamu\.edu](mailto:[email protected])Affiliation:School of Engineering Medicine, Texas A&M University, Houston, Texas, USAKim\-Kwang Raymond ChooEmail:[raymond\.choo@utsa\.edu](mailto:[email protected])Affiliation:Department of Information Systems and Cybersecurity, The University of Texas at San Antonio, San Antonio, Texas, USAYuanxiong GuoEmail:[yuanxiong\.guo@utsa\.edu](mailto:[email protected])Affiliation:Department of Information Systems and Cybersecurity, The University of Texas at San Antonio, San Antonio, Texas, USA

###### keywords

Medication for Opioid Use Disorder, Machine Learning, Group Fairness, Bias Mitigation, Interpretability

### 1Introduction

Opioid use disorder \(OUD\) continues to represent a critical public health crisis in the United States, with far\-reaching consequences for population health, healthcare systems, and the broader economy[76](https://arxiv.org/html/2609.22113#bib.bib7)\. OUD is a chronic, relapsing condition marked by persistent opioid use, including both prescribed and illicit opioids, despite adverse physical, psychological, and social consequences[28](https://arxiv.org/html/2609.22113#bib.bib5)\. Recent estimates indicate that approximately 6 million individuals in the United States were affected by OUD in 2022[17](https://arxiv.org/html/2609.22113#bib.bib4)\. In 2023 alone, opioid\-involved overdoses accounted for more than 81,000 deaths, representing roughly 75% of all drug overdose fatalities nationwide[6](https://arxiv.org/html/2609.22113#bib.bib1)\. Although data from 2024 suggest a modest decline in opioid\-related overdose deaths, indicators of ongoing harm persist[57](https://arxiv.org/html/2609.22113#bib.bib2)\. In particular, emergency department visits for non\-fatal opioid overdoses have continued to increase, rising by approximately 8% in February 2025[18](https://arxiv.org/html/2609.22113#bib.bib3)\.

Amid this public health emergency, medications for OUD \(MOUD\)—including methadone, buprenorphine, and naltrexone—are among the most effective interventions\. These medication treatments have been shown to significantly reduce both overdose\-related and all\-cause mortality while supporting long\-term stabilization and recovery[52](https://arxiv.org/html/2609.22113#bib.bib9)\. Yet, their effectiveness depends heavily on patient engagement and sustained treatment retention\. Despite their proven efficacy, many patients discontinue therapy prematurely and struggle to remain engaged in MOUD, making long\-term retention a persistent and complex challenge, particularly in outpatient treatment settings where sustained engagement is essential for achieving effective outcomes[55](https://arxiv.org/html/2609.22113#bib.bib24);[10](https://arxiv.org/html/2609.22113#bib.bib41)\. Short\-term regimens or early tapers are often associated with increased risks of relapse, overdose, and reduced long\-term stability[74](https://arxiv.org/html/2609.22113#bib.bib10)\. Prior research has shown that treatment discontinuation reflects a complex interplay of clinical, behavioral, and structural factors, including co\-occurring mental health conditions, unstable housing or transportation, employment constraints, stigma, limited access to supportive services, as well as demographic characteristics such as race and age[22](https://arxiv.org/html/2609.22113#bib.bib49);[66](https://arxiv.org/html/2609.22113#bib.bib59)\. These multifaceted factors make it difficult to identify which individuals are most likely to remain engaged in and successfully complete outpatient MOUD treatment[10](https://arxiv.org/html/2609.22113#bib.bib41)\. As a result, evidence\-based interventions known to improve retention, such as proactive appointment reminders and check\-ins, early engagement with counseling or behavioral therapy, transportation support or telehealth adjustments, more frequent dosing supervision for clinically unstable patients, and peer recovery coaching or contingency\-management interventions, often cannot be delivered uniformly across patients[37](https://arxiv.org/html/2609.22113#bib.bib76)\. This challenge is further compounded by the significant resource constraints under which many MOUD programs operate, limiting their capacity to provide intensive follow\-up and support to all patients[22](https://arxiv.org/html/2609.22113#bib.bib49);[37](https://arxiv.org/html/2609.22113#bib.bib76)\.

In this context, accurate prediction of patients at high risk of premature discontinuation can play a critical role in guiding clinical prioritization and resource allocation\. Machine learning \(ML\) models can leverage large\-scale datasets to identify patterns and risk factors associated with treatment disengagement[35](https://arxiv.org/html/2609.22113#bib.bib11);[50](https://arxiv.org/html/2609.22113#bib.bib12)\. By enabling early risk identification, ML models can support several practical uses in MOUD treatment settings, including risk stratification at intake to inform initial treatment planning, flagging early signs of disengagement within the first days or weeks of treatment, and triggering clinical decision support for targeted follow\-up interventions[69](https://arxiv.org/html/2609.22113#bib.bib8);[78](https://arxiv.org/html/2609.22113#bib.bib26)\.

However, differences in MOUD access and retention across patient sociodemographic subgroups are often reflected in the data used to train these models\. In addition to individual\-level complexity, prior research has identified consistent patterns in treatment retention and completion rates across different patient population subgroups\. Studies have shown that certain patient sociodemographic subgroups may experience lower rates of MOUD initiation and a higher likelihood of premature discontinuation[70](https://arxiv.org/html/2609.22113#bib.bib18);[45](https://arxiv.org/html/2609.22113#bib.bib17);[54](https://arxiv.org/html/2609.22113#bib.bib44)\. These patterns are often influenced by factors such as limited provider availability, lack of insurance coverage, and logistical obstacles related to transportation, scheduling, or clinic accessibility[38](https://arxiv.org/html/2609.22113#bib.bib46);[3](https://arxiv.org/html/2609.22113#bib.bib45);[63](https://arxiv.org/html/2609.22113#bib.bib47);[70](https://arxiv.org/html/2609.22113#bib.bib18)\. Additional external pressures, such as financial hardship or limited access to nearby treatment programs, can further reduce the likelihood of completing a full course of care[38](https://arxiv.org/html/2609.22113#bib.bib46);[2](https://arxiv.org/html/2609.22113#bib.bib48)\. As a result, there is growing concern that ML models may unintentionally produce inconsistent or less reliable predictions across different patient subgroups, leading to variation in model performance that could affect clinical decision\-making—an important consideration known as algorithmic fairness[59](https://arxiv.org/html/2609.22113#bib.bib13);[62](https://arxiv.org/html/2609.22113#bib.bib14);[20](https://arxiv.org/html/2609.22113#bib.bib15)\. These inconsistencies can arise at any stage of the ML model development cycle, ranging from data collection to model training and evaluation, and may result in treatment recommendations that do not generalize well across all patient subgroups[59](https://arxiv.org/html/2609.22113#bib.bib13);[73](https://arxiv.org/html/2609.22113#bib.bib25)\. When ML models consistently misestimate treatment outcomes for certain patient subgroups, patients may receive care that is not well\-aligned with their clinical profiles or treatment needs, ultimately limiting the overall reliability and clinical utility of these models\.

Although ML models have been extensively studied in MOUD research, prior work has paid limited attention to treatment outcome prediction specifically in outpatient treatment settings, as well as to systematic evaluation of how these models perform across different patient subgroups[69](https://arxiv.org/html/2609.22113#bib.bib8);[78](https://arxiv.org/html/2609.22113#bib.bib26);[23](https://arxiv.org/html/2609.22113#bib.bib27);[4](https://arxiv.org/html/2609.22113#bib.bib28);[49](https://arxiv.org/html/2609.22113#bib.bib29);[32](https://arxiv.org/html/2609.22113#bib.bib36)\. While the broader healthcare ML literature has increasingly explored issues related to algorithmic fairness and reliability across patient subgroups[31](https://arxiv.org/html/2609.22113#bib.bib30);[41](https://arxiv.org/html/2609.22113#bib.bib31);[77](https://arxiv.org/html/2609.22113#bib.bib35), there remains a significant gap in understanding how such concerns apply to ML models specifically designed for MOUD retention and completion\. Furthermore, although numerous strategies have been developed to improve model fairness across subgroups[36](https://arxiv.org/html/2609.22113#bib.bib32);[53](https://arxiv.org/html/2609.22113#bib.bib33), their effectiveness in the context of MOUD retention and premature discontinuation prediction remain largely unexplored\. Addressing these gaps is crucial to developing reliable and clinically useful ML decision support tools in addiction treatment\. Therefore, this study seeks to address the following research questions:

- •How accurately can ML models predict treatment retention and premature discontinuation in outpatient MOUD treatment settings using routinely collected administrative data?
- •To what extent do ML models show performance inconsistency across different patient subgroups when predicting MOUD retention and premature discontinuation?
- •How do model predictions and underlying feature attributions differ across patient subgroups, and what do these differences reveal about patterns in model behavior?
- •How effective are bias mitigation techniques—such as pre\-processing reweighting, in\-processing constraint optimization, and post\-processing threshold adjustment—in improving model fairness across patient subgroups while maintaining overall predictive accuracy?

To answer these questions, this study systematically assesses the fairness and reliability of ML models developed to predict MOUD retention and premature discontinuation\. We evaluate the predictive performance and fairness of ML models across patient subgroups defined by race, ethnicity, age, and sex, investigate subgroup\-level differences in model behavior using model explanation techniques, and evaluate representative bias mitigation strategies spanning the pre\-processing, in\-processing, and post\-processing stages of the ML pipeline\. Through this comprehensive analysis, our goal is to improve understanding of how ML models perform across different patient populations in the context of MOUD, and to offer practical guidance for developing more consistent and clinically reliable predictive tools to support treatment planning and decision\-making\. In summary, this paper makes the following key contributions:

1. 1\.We develop and evaluate ML models to predict MOUD treatment outcomes in outpatient treatment settings, including treatment retention and premature discontinuation\. This fills an important gap in the MOUD prediction literature, which has largely focused on mixed treatment settings\. Outpatient MOUD care is less structured, relies more heavily on sustained patient self\-engagement, and is more strongly influenced by dynamic social and contextual factors, making outcome prediction both more challenging and more operationally consequential than in other treatment settings\.
2. 2\.We conduct a comprehensive empirical evaluation of algorithmic fairness in ML models predicting MOUD treatment retention and premature discontinuation\. By integrating subgroup\-level performance assessment with model explanation analyses, we identify systematic differences in error rates, feature importance, and predictive behavior across patient sociodemographic subgroups\. This analysis contributes empirical evidence regarding the fairness and reliability of ML models for MOUD outcome prediction\.
3. 3\.We assess the effectiveness of widely used bias mitigation strategies across different stages of the ML development pipeline and compare their effects on predictive performance and fairness in MOUD outcome prediction\. This evaluation provides practical guidance for selecting bias mitigation approaches appropriate for different fairness objectives and predictive performance requirements in MOUD treatment settings\.

### 2Related Work

In this section, we review prior research across three key areas: \(1\) MOUD treatment outcomes and ML; \(2\) algorithmic fairness and bias mitigation in ML models\. We conclude by identifying gaps in the current literature and describing how our study addresses these challenges\.

#### 2\.1MOUD Treatment Outcomes and ML

The effectiveness of MOUD depends not only on treatment initiation but also on sustained engagement over time\. Prior literature indicates that fewer than one\-third of adults with prescription opioid use disorder ever receive treatment[14](https://arxiv.org/html/2609.22113#bib.bib60), highlighting substantial barriers at the initiation stage\. Among those who receive treatment, retention rates are discouragingly low and vary considerably across different follow\-up periods[74](https://arxiv.org/html/2609.22113#bib.bib10)\. Treatment retention, often measured by continued participation over clinically meaningful durations such as 90 or 180 days, has been consistently associated with lower relapse risk, reduced overdose, and improved stability[74](https://arxiv.org/html/2609.22113#bib.bib10)\. In outpatient MOUD settings, sustained engagement is especially important because treatment occurs in less supervised environments and depends on regular attendance, medication adherence, and continued follow\-up over time[10](https://arxiv.org/html/2609.22113#bib.bib41)\. However, premature treatment discontinuation remains common\. Prior studies have shown that a substantial proportion of patients discontinue therapy prematurely within the first weeks or months of treatment, particularly during early phases when patients may still be adjusting to medication, clinic requirements, and psychosocial supports[60](https://arxiv.org/html/2609.22113#bib.bib58);[66](https://arxiv.org/html/2609.22113#bib.bib59);[56](https://arxiv.org/html/2609.22113#bib.bib61)\. Retention outcomes vary widely across treatment settings, patient populations, and medication types, and are shaped by a complex interplay of clinical, behavioral, and structural factors[55](https://arxiv.org/html/2609.22113#bib.bib24);[65](https://arxiv.org/html/2609.22113#bib.bib62);[80](https://arxiv.org/html/2609.22113#bib.bib63)\. For example, prior research has documented substantial differences in treatment completion across sociodemographic subgroups, including by race, age, and sex, as well as by social and contextual conditions such as employment status, housing stability, and transportation access[22](https://arxiv.org/html/2609.22113#bib.bib49);[35](https://arxiv.org/html/2609.22113#bib.bib11);[55](https://arxiv.org/html/2609.22113#bib.bib24);[65](https://arxiv.org/html/2609.22113#bib.bib62)\. As a result, identifying patients at elevated risk of premature discontinuation in order to support targeted interventions remains challenging, particularly in outpatient MOUD settings, where engagement is less structured and programs must selectively prioritize limited resources such as care coordination time, counseling and peer\-support services, proactive outreach, and follow\-up visits\.

ML methods have increasingly been applied to support risk stratification, the process of identifying patients at relatively higher or lower risk for adverse outcomes in order to guide prioritization of care, using routinely collected clinical or administrative data in addiction treatment settings[69](https://arxiv.org/html/2609.22113#bib.bib8);[35](https://arxiv.org/html/2609.22113#bib.bib11);[7](https://arxiv.org/html/2609.22113#bib.bib64)\. This is particularly relevant in the context of MOUD treatment\. Previous studies have used conventional regression\-based methods to identify the significant predictors of OUD\-related outcomes[82](https://arxiv.org/html/2609.22113#bib.bib67);[64](https://arxiv.org/html/2609.22113#bib.bib68);[33](https://arxiv.org/html/2609.22113#bib.bib65)\. More recently, ML models can flexibly capture complex, non\-linear relationships among patient characteristics, treatment factors, and contextual variables, and have increasingly been used to develop predictive tools for OUD risk stratification[45](https://arxiv.org/html/2609.22113#bib.bib17);[10](https://arxiv.org/html/2609.22113#bib.bib41);[69](https://arxiv.org/html/2609.22113#bib.bib8);[81](https://arxiv.org/html/2609.22113#bib.bib66);[78](https://arxiv.org/html/2609.22113#bib.bib26);[23](https://arxiv.org/html/2609.22113#bib.bib27)\. In real\-world settings where data are limited and outcomes are influenced by various factors, ML\-based risk stratification offers a pragmatic approach to supporting population\-level decision\-making and quality improvement efforts\. Such models can inform decisions, such as prioritizing care coordination efforts, allocating behavioral health counseling resources, or triggering proactive follow\-up for patients at elevated risk of disengagement[35](https://arxiv.org/html/2609.22113#bib.bib11)\. At the same time, the use of ML\-based approaches raises important concerns regarding algorithmic fairness across patient subgroups\. If predictive models systematically over\- or under\-estimate risk for certain populations, risk\-informed decision\-making may inadvertently reinforce existing disparities in access to care or support services\. These concerns motivate the need for careful evaluation of model performance and the incorporation of algorithmic fairness considerations when developing ML models for MOUD treatment outcome prediction\.

#### 2\.2Algorithmic Fairness and Bias Mitigation

The growing use of ML has brought increasing attention to*algorithmic bias*, which occurs when AI models consistently produce less accurate or potentially discriminatory outcomes for certain subgroups of people\([61](https://arxiv.org/html/2609.22113#bib.bib51)\)\. Although developing separate prediction models for each patient subgroup may appear straightforward, social categories such as race and ethnicity are neither clear\-cut nor mutually exclusive[26](https://arxiv.org/html/2609.22113#bib.bib75)\. Rigidly assigning individuals to a single subgroup and conducting subgroup\-specific modeling can therefore misrepresent lived experiences and lead to biased model use\. This also risks reinforcing the false assumption that observed disparities reflect biological differences rather than the underlying structural or contextual factors[26](https://arxiv.org/html/2609.22113#bib.bib75)\. Therefore, evaluating differences in model performance across patient subgroups has become an emerging practice for assessing the algorithmic fairness of AI models\([34](https://arxiv.org/html/2609.22113#bib.bib22)\)\. Recent studies have established fairness evaluation criteria and model development procedures to quantify and mitigate differences in model performance across population subgroups defined by sensitive attributes such as race and sex\([34](https://arxiv.org/html/2609.22113#bib.bib22);[5](https://arxiv.org/html/2609.22113#bib.bib21)\)\. For example,*demographic parity*requires predictions to be statistically independent of sensitive attributes\([5](https://arxiv.org/html/2609.22113#bib.bib21)\)\.*Equal opportunity*\([34](https://arxiv.org/html/2609.22113#bib.bib22)\)requires equal true positive rates \(TPR\) across subgroups, while*equalized odds*\([34](https://arxiv.org/html/2609.22113#bib.bib22)\)adds an additional constraint that demands equal false positive rates \(FPR\)\. Both require that outcomes be conditionally independent of sensitive attributes given the true outcome\. All of these metrics aim to enforce performance parity across subgroups\. Since there is no consensus on a single metric or criterion for assessing algorithmic fairness, the choice of fairness definition often depends on the specific characteristics of the application\.

Various approaches have been developed to achieve fairness in ML, which are generally classified into three categories: pre\-processing, in\-processing, and post\-processing methods\. Pre\-processing methods focus on modifying input data to reduce biases before training, including techniques like resampling, adding new data, or adjusting labels\([24](https://arxiv.org/html/2609.22113#bib.bib52);[39](https://arxiv.org/html/2609.22113#bib.bib53);[1](https://arxiv.org/html/2609.22113#bib.bib54)\)\. Post\-processing methods adjust model predictions after training to meet fairness objectives, often by modifying decision thresholds or outcomes for specific subgroups\([44](https://arxiv.org/html/2609.22113#bib.bib55)\)\. In\-processing methods incorporate fairness constraints or objectives directly into the learning algorithm during training, penalizing the learning of discriminatory features and enabling a balance between fairness and predictive performance\. Examples include adversarial training, regularization, adaptive weighting, or fairness constraints on representations\([5](https://arxiv.org/html/2609.22113#bib.bib21);[27](https://arxiv.org/html/2609.22113#bib.bib56);[86](https://arxiv.org/html/2609.22113#bib.bib57)\)\.

In healthcare, algorithmic bias in ML models can arise from multiple sources, including structural inequities in access to care, differences in data completeness or quality across populations, and the use of proxy variables or shortcuts that reflect social and contextual factors rather than true clinical risk[59](https://arxiv.org/html/2609.22113#bib.bib13);[47](https://arxiv.org/html/2609.22113#bib.bib69);[20](https://arxiv.org/html/2609.22113#bib.bib15)\. As a result, models optimized for overall predictive performance may produce unfair outputs across patient subgroups, potentially reinforcing existing disparities\. Prior research has documented fairness concerns in a range of healthcare applications, including risk prediction, disease screening, and resource allocation, demonstrating that models with acceptable overall performance may still systematically disadvantage certain population subgroups[59](https://arxiv.org/html/2609.22113#bib.bib13);[20](https://arxiv.org/html/2609.22113#bib.bib15);[68](https://arxiv.org/html/2609.22113#bib.bib70);[48](https://arxiv.org/html/2609.22113#bib.bib71)\. In the context of MOUD, existing literature suggests that structural and contextual factors play an important role in patients’ sustained engagement in treatment\([22](https://arxiv.org/html/2609.22113#bib.bib49);[35](https://arxiv.org/html/2609.22113#bib.bib11)\)\. Community\-level context, including the availability of treatment providers, transportation infrastructure, and the broader socioeconomic environment, can further influence treatment adherence and outcomes\([43](https://arxiv.org/html/2609.22113#bib.bib50)\)\. When such variations and patterns are encoded in routinely collected clinical or administrative data, they may also be implicitly learned by ML models, causing unfair and biased outcomes across patient subgroups[20](https://arxiv.org/html/2609.22113#bib.bib15)\. However, fairness considerations have received limited attention in ML models predicting MOUD treatment outcomes\. As predictive models are increasingly used to support risk stratification and care prioritization, it is important to systematically examine whether these models perform fairly across patient subgroups and to assess whether bias mitigation strategies can reduce observed disparities\. The choice of fairness criteria should therefore be guided by the intended use of the ML model and the potential consequences of different types of prediction errors\. In this study, we focus on error\-rate–based fairness metrics, as the prediction model is intended to support risk stratification and care prioritization in MOUD treatment settings\. In this context, unfair false negative rates \(FNR\) may lead to missed opportunities for early intervention, while unfair FPR may result in unnecessary monitoring or resource allocation for certain patient subgroups\. Evaluating TPR and FPR parity therefore directly reflects whether patients from different subgroups have equitable access to follow\-up and support when risk\-based decisions are made\. Such fairness metrics have been widely used in prior healthcare literature where models inform disease screening, monitoring, or resource allocation decisions[68](https://arxiv.org/html/2609.22113#bib.bib70);[75](https://arxiv.org/html/2609.22113#bib.bib72)\.

### 3Methods

This retrospective cross\-sectional study used data from the U\.S\. Substance Abuse and Mental Health Services Administration’s \(SAMHSA\) Treatment Episode Data Set \- Discharges \(TEDS\-D\)[71](https://arxiv.org/html/2609.22113#bib.bib6), focusing on adult cases receiving MOUD in outpatient treatment facilities between 2015 and 2019\. In this study, we focused on two clinically important and complementary treatment outcomes in MOUD programs: premature treatment discontinuation and treatment retention exceeding 180 days\. Treatment retention is a key indicator of successful treatment engagement, whereas premature treatment discontinuation is associated with poor treatment outcomes and an increased risk of relapse, overdose, and mortality\. Our objective was not only to develop predictive models for these outcomes, but also to systematically examine how model performance varies across patient subgroups defined by sociodemographic characteristics such as race, ethnicity, and age, which are well\-documented attributes often linked to differences in treatment access and outcomes\. By treating these variables as sensitive attributes, we conducted a comprehensive assessment of whether prediction accuracy and behavior differed across these subgroups, and implemented bias mitigation strategies to improve the fairness of model performance across the patient sociodemographic subgroups\. The overall process of our pipeline is shown in Figure[1](https://arxiv.org/html/2609.22113#S3.F1)\.

![Refer to caption](https://arxiv.org/html/2609.22113v1/Fig1.png)Figure 1:Overall pipeline\.#### 3\.1Data

##### 3\.1\.1Data Sources and Study Population

This study utilized the 2015–2019 TEDS\-D dataset[71](https://arxiv.org/html/2609.22113#bib.bib6), a national data system of annual discharges from substance use treatment facilities licensed or certified by Single State Agencies and receiving federal funding\. TEDS\-D provides episode\-level information on discharges from substance use treatment services\. The dataset includes individuals aged 12 and older and covers a wide range of demographic, clinical and treatment\-related variables, as well as substance use characteristics and history for individuals admitted to treatment programs\. Importantly, each record in the dataset corresponds to a treatment episode rather than a unique individual\. TEDS\-D enables large\-scale, population\-level analyses of substance use treatment outcomes and is commonly used to examine patterns in service utilization, treatment completion, and other discharge\-related outcomes across a wide range of treatment settings in the United States\.

We derived cohort based on specific inclusion and exclusion criteria to ensure the robustness and consistency of the results\. First, the study cohort was restricted to adult patients \(aged 18 and older\) receiving MOUD, which was defined as the documented use of methadone, buprenorphine, or naltrexone in the patient’s treatment plan at the time of admission[70](https://arxiv.org/html/2609.22113#bib.bib18)\. Restricting the cohort to adults ensured demographic consistency and allowed us to focus on a population with well\-documented treatment trajectories and distinct clinical needs\. To focus on opioid\-related treatment outcomes, we further restricted the cohort to discharges where heroin or other opioids/synthetics were documented as the primary, secondary, or tertiary substance at admission, allowing us to specifically target patients with opioid use[69](https://arxiv.org/html/2609.22113#bib.bib8)\. We further restricted the cohort to individuals receiving treatment in outpatient settings to ensure consistency in treatment environments[45](https://arxiv.org/html/2609.22113#bib.bib17)\. Outpatient settings present a substantially more challenging prediction task than mixed service settings, as treatment engagement is less structured and outcomes are influenced by dynamic and often unobserved factors\. In such settings, ML may provide useful support for identifying patterns that are otherwise difficult to capture\. Finally, we excluded discharges classified as transfers to other treatment programs or facilities, preventing interruptions in treatment continuity from confounding the analysis and ensuring that the treatment outcomes evaluated correspond to a single, uninterrupted episode of care\. A total of 366,825 patients satisfied the respective inclusion/exclusion criteria for the model development\. Figure[9](https://arxiv.org/html/2609.22113#S7.F9)presents the complete cohort and variable selection process\. In addition, to assess the robustness of our analysis, we conducted additional experiments using the 2020–2023 TEDS\-D data following the same study design and analysis pipeline\. The results are presented in Appendix[7\.8](https://arxiv.org/html/2609.22113#S7.SS8)\.

##### 3\.1\.2Outcome Measures

This study examined two outcomes of interest: premature discontinuation of MOUD treatment and treatment retention beyond 180 days, defined as follows:

- \(1\)Premature Discontinuation: A binary outcome variable coded as positive for individuals who discontinue treatment prematurely due to potentially preventable attrition, that is, cases in which people “chose” not to complete treatment\. Specifically, this category included discharge reasons such as dropout for unknown reasons, loss to follow\-up, failure to return from leave, or discharge for administrative purposes following a prolonged absence from treatment\. These discharge reasons represent situations in which treatment discontinuation may be amenable to clinical or programmatic intervention \(e\.g\., counseling, outreach, or adjustments to the treatment plan\)\. The variable was coded as negative for all other discharge reasons, including treatment completion; termination by the facility due to non\-compliance or violations of program rules, laws, or policies; incarceration \(including jail, prison, house confinement, or release to or from the courts\); and death\. These discharge reasons were not considered premature discontinuation because they primarily reflect institutional decisions, external circumstances, or other reasons for treatment termination rather than patient disengagement from treatment, making them less directly amenable to interventions aimed at improving patient retention\. Transfers to another treatment program or facility were excluded[69](https://arxiv.org/html/2609.22113#bib.bib8)\. This outcome was derived from the “reason for discharge” variable in the original dataset and was based solely on the recorded discharge reason, consistent with prior work[69](https://arxiv.org/html/2609.22113#bib.bib8)\.
- \(2\)Length of Stay \(LOS\)\>\>180 Days: A binary variable measuring treatment retention or sustained engagement beyond 180 days\. The variable was coded as positive if the treatment episode lasted more than 180 days and negative otherwise, aligned with established treatment cascade metrics for MOUD continuity[45](https://arxiv.org/html/2609.22113#bib.bib17);[84](https://arxiv.org/html/2609.22113#bib.bib19);[8](https://arxiv.org/html/2609.22113#bib.bib16)\. The 180\-day threshold was selected based on prior studies demonstrating that patients who discontinue MOUD within the first six months are at substantially higher risk of relapse and adverse outcomes compared to those retained beyond this period[40](https://arxiv.org/html/2609.22113#bib.bib73);[83](https://arxiv.org/html/2609.22113#bib.bib74)\. This cutoff has been widely used in MOUD research and clinical trials as a clinically meaningful marker of sustained treatment engagement[40](https://arxiv.org/html/2609.22113#bib.bib73);[50](https://arxiv.org/html/2609.22113#bib.bib12)\. Accordingly, retention beyond 180 days identifies individuals who have achieved a level of continuity of treatment associated with improved outcomes, whereas shorter treatment episodes may indicate elevated risk and the need for additional support\. This outcome was derived from the “length of stay in treatment \(days\)” variable in the original dataset and was defined exclusively based on treatment duration over a clinically meaningful period\.

Because the TEDS\-D dataset does not contain a single variable that simultaneously captures both treatment duration and the reason for discharge, we operationalized treatment engagement using these two complementary outcomes\. Both outcomes were non\-missing for all treatment episodes, ensuring complete data availability for evaluation and analysis\.

##### 3\.1\.3Predictor Variables

To predict treatment outcomes, we included a range of variables collected at admission, spanning demographic characteristics \(e\.g\., race, ethnicity, sex, and age\)[46](https://arxiv.org/html/2609.22113#bib.bib82);[45](https://arxiv.org/html/2609.22113#bib.bib17), socioeconomic factors \(e\.g\., employment status, education, veteran status, living situation, geographic region, and arrest history\)[30](https://arxiv.org/html/2609.22113#bib.bib83);[12](https://arxiv.org/html/2609.22113#bib.bib85);[67](https://arxiv.org/html/2609.22113#bib.bib87);[72](https://arxiv.org/html/2609.22113#bib.bib84), substance use behaviors \(e\.g\., number of substance used\)[69](https://arxiv.org/html/2609.22113#bib.bib8), self\-reported substances \(e\.g\., heroin, alcohol, inhalants, and marijuana/hashish\)[69](https://arxiv.org/html/2609.22113#bib.bib8), other clinical factors \(e\.g\., presence of psychological problems\)[15](https://arxiv.org/html/2609.22113#bib.bib86);[67](https://arxiv.org/html/2609.22113#bib.bib87), and treatment\-related factors \(e\.g\., source of referral, service setting, and prior treatment history\)[16](https://arxiv.org/html/2609.22113#bib.bib81);[69](https://arxiv.org/html/2609.22113#bib.bib8)\. Specifically, we first excluded variables that were not relevant or applicable to the majority of the study population and therefore had no recorded values for more than 90% of individuals\. And we only considered the variables collected at admission\. We also derived several variables to better capture key patterns, including the highest frequency of non\-medical opioid use recorded among participants, a binary indicator for any current heroin use, and four distinct binary variables indicating whether participants had ever used substances within the broader categories of stimulants, hallucinogens, sedatives, or tranquilizers\. Finally, among the remaining candidate variables, those with more than 20% missing values were excluded\. In total, 27 variables were included, 2 were target variables \(i\.e\., retention and premature discontinuation\), leaving 25 variables for use as predictors\.

#### 3\.2Prediction Models

In this study, we assessed algorithmic fairness in the prediction of MOUD retention and premature discontinuation across four ML models: LR, RF, GBDT, and MLP\. These models were selected based on their widespread use in clinical prediction tasks and their complementary strengths in handling structured, tabular healthcare data[87](https://arxiv.org/html/2609.22113#bib.bib37);[25](https://arxiv.org/html/2609.22113#bib.bib38);[58](https://arxiv.org/html/2609.22113#bib.bib39)\. LR is a widely used baseline model for binary classification problems\. t assumes a linear relationship between predictors and the log\-odds of the outcome, making it particularly useful for establishing foundational performance benchmarks\. RF is an ensemble\-based method that constructs multiple decision trees and aggregates their outputs to improve accuracy and robustness\. Its ability to capture non\-linear interactions and rank variable importance has made it a standard model in medical applications involving complex clinical features\. GBDT is a boosting\-based ensemble method that iteratively reduces errors made by previous learners, which enhances its capacity to model subtle and complex relationships\. GBDT models are often considered state\-of\-the\-art in structured data scenarios due to their high predictive accuracy and versatility in handling heterogeneous variables\. MLP is a fully connected feedforward neural network that can learn complex, non\-linear patterns in data\. Although MLPs typically require more extensive tuning and are less interpretable than tree\-based models, they have shown competitive performance in structured clinical prediction when sufficient data are available\. By leveraging both traditional and neural network–based approaches, we aim to provide a comprehensive evaluation of predictive performance and subgroup\-level reliability across a range of ML models\.

#### 3\.3Empirical Analysis

Hyperparameters were optimized separately for each model using the procedures described below\. Data were randomly partitioned into 70% for training and 30% for testing using stratified sampling\. The training partition was further divided into 80% for training and 20% for validation\. The validation set was used for hyperparameter tuning for the MLP model, whereas LR, RF, and GBDT used five\-fold stratified cross\-validation on the training subset\. All variables were categorical variables and were converted to one\-hot encoding before being fed into ML models\. Independent models were trained and optimized for premature treatment discontinuation and treatment retention exceeding 180 days, as these outcomes are defined by different mechanisms\. Models were selected based on the area under the receiver operating characteristic \(AUC\), with results averaged over 10 independent runs using different random seeds to account for variability\. Best hyperparameter configurations for each model and task are summarized in Appendix[7\.2](https://arxiv.org/html/2609.22113#S7.SS2)Table[8](https://arxiv.org/html/2609.22113#S7.T8)\.

##### 3\.3\.1Subgroup\-Level Model Performance Evaluation

To evaluate how predictive performance varies across patient subgroups, we analyzed model predictions based on sociodemographic characteristics available in the dataset as sensitive attributes, with a focus on fields related to recorded race, ethnicity, and age\. The original TEDS\-D dataset includes nine racial categories, but due to small sample sizes in many of these subgroups, we limited our subgroup analysis to White and Black/African American patients\. Data samples from other racial categories were excluded from our evaluation\. Two ethnicity subgroups, Non\-Hispanic and Hispanic, were considered\. Patients were also stratified into three age subgroups: 18\-24 \(young adults\), 25\-54 \(middle\-aged adults\), and over 55 years \(older adults\)\.

To examine potential differences in predictive behavior, we used key classification metrics derived from the confusion matrix: true positive rate \(TPR or sensitivity\), false positive rate \(FPR\), false negative rate \(FNR\) and true negative rate \(TNR\), as listed in Table[1](https://arxiv.org/html/2609.22113#S3.T1)\. These metrics are particularly relevant to clinical decision\-making and resource allocation\. Since FNR and TNR can be directly derived from TPR and FPR, we primarily focused our analysis on TPR and FPR\. High FPR values indicate that patients who would have remained in treatment are incorrectly flagged as being at risk of early discontinuation, which can lead to unnecessary interventions, misallocated resources, and negative perceptions\. Low TPR values, on the other hand, suggest that many patients who are truly at risk are not identified in time, limiting the opportunity for timely intervention and support\. In the task of predicting whether patients will remain in treatment for more than 180 days, the same principles apply: accurate identification helps ensure that extended\-care resources are directed to those who truly need them, while minimizing both missed opportunities and inefficient resource use\.

Table 1:Confusion matrix and derived metricsPredicted \- PositivePredicted \- NegativeActual \- PositiveTrue Positive \(TP\)False Negative \(FN\)T​P​R=T​PT​P\+F​NTPR=\\frac\{TP\}\{TP\+FN\}F​N​R=1−T​P​RFNR=1\-TPRActual \- NegativeFalse Positive \(FP\)True Negative \(TN\)F​P​R=F​PF​P\+T​NFPR=\\frac\{FP\}\{FP\+TN\}T​N​R=1−F​P​RTNR=1\-FPRThe above metrics also align with two widely used subgroup fairness criteria:*Equal Opportunity*and*Predictive Equality*[34](https://arxiv.org/html/2609.22113#bib.bib22)\. Equal opportunity ensures that TPRs remain consistent across subgroups defined by a sensitive attribute\. It evaluates how well a model correctly identifies positive cases, ensuring that individuals with a positive ground truth outcome \(e\.g\., premature discontinuation or retention longer than 180 days\) have an equal likelihood of being predicted as such, regardless of their sensitive attribute\. The formal definition in the binary\-subgroup case is:

P⁡\(Y^=1\|Y=1,A=0\)=P⁡\(Y^=1\|Y=1,A=1\)P\(\\hat\{Y\}=1\|Y=1,A=0\)=P\(\\hat\{Y\}=1\|Y=1,A=1\)\(1\)whereY∈\{0,1\}Y\\in\\\{0,1\\\}denotes the actual outcome,Y^∈\{0,1\}\\hat\{Y\}\\in\\\{0,1\\\}is the model’s predicted outcome,A∈\{0,1\}A\\in\\\{0,1\\\}represents the sensitive attribute \(e\.g\., race, age\)\. In practice, however, sensitive attributes often include more than two subgroups—for instance, multiple racial/ethnic subgroups or age categories\. In such cases, equal opportunity can be generalized to require that the TPRs be similar across all subgroups𝒜=\{a1,a2,…,aN\}\\mathcal\{A\}=\\\{a\_\{1\},a\_\{2\},\\ldots,a\_\{N\}\\\}, whereNNdenotes the total number of subgroups\. To quantify the extent to which equal opportunity is satisfied across multiple subgroups, we measure the*Equal Opportunity Ratio*\(EOR\)\. This metric compares the TPRs between subgroups defined by the sensitive attributeAA\. Formally, it is defined as:

EOR=minai∈𝒜⁡P⁡\(Y^=1∣Y=1,A=ai\)maxai∈𝒜⁡P⁡\(Y^=1∣Y=1,A=ai\),\\text\{EOR\}=\\frac\{\\min\_\{a\_\{i\}\\in\\mathcal\{A\}\}P\(\\hat\{Y\}=1\\mid Y=1,A=a\_\{i\}\)\}\{\\max\_\{a\_\{i\}\\in\\mathcal\{A\}\}P\(\\hat\{Y\}=1\\mid Y=1,A=a\_\{i\}\)\},\(2\)where𝒜\\mathcal\{A\}is the set of all subgroups defined by sensitive attributeAA\. This ratio captures the worst\-case gap in TPRs across subgroups\. A value closer to 1 indicates that the model identifies positive cases at a similar rate for all subgroups, aligning with the principle of equal opportunity\. Lower values signal greater gaps, meaning the model favors certain subgroups over others in predicting positive outcomes\.

Predictive Equality focuses on ensuring that FPRs are consistent across subgroups defined by a sensitive attribute\. It ensures that individuals who do not actually exhibit the outcome of interest \(e\.g\., those who do not discontinue MOUD treatment prematurely, or those who do not stay in treatment longer than 180 days\) are not disproportionately predicted as high\-risk simply because they belong to a particular subgroup\. Formally, the definition in the binary\-subgroup setting is:

P⁡\(Y^=1\|Y=0,A=0\)=P⁡\(Y^=1\|Y=0,A=1\)P\(\\hat\{Y\}=1\|Y=0,A=0\)=P\(\\hat\{Y\}=1\|Y=0,A=1\)\(3\)Similarly, we use the*Predictive Equality Ratio*\(PER\) to quantify the extent to which predictive equality is satisfied:

PER=minai∈𝒜⁡P⁡\(Y^=1∣Y=0,A=ai\)maxai∈𝒜⁡P⁡\(Y^=1∣Y=0,A=ai\)\.\\text\{PER\}=\\frac\{\\min\_\{a\_\{i\}\\in\\mathcal\{A\}\}P\(\\hat\{Y\}=1\\mid Y=0,A=a\_\{i\}\)\}\{\\max\_\{a\_\{i\}\\in\\mathcal\{A\}\}P\(\\hat\{Y\}=1\\mid Y=0,A=a\_\{i\}\)\}\.\(4\)This ratio captures the worst\-case gap in FPRs across subgroups\. A PER close to 1 indicates that the model is making false positive errors at a similar rate across all subgroups, which aligns with the fairness goal of avoiding disproportionate over\-prediction for certain population subgroups\. Values far from 1 highlight discrepancies that could lead to unequal burdens, such as unnecessary interventions or stigma disproportionately affecting specific subgroups\.

In addition, we also evaluate Equalized Odds, which jointly measures gaps in TPRs and FPRs across patient subgroups\. To quantify deviations from this criterion, we use the Equalized Odds Difference \(EOD\), which is defined as the maximum of the TPR difference and the FPR difference across subgroups, as follows:

ΔTPR\\displaystyle\\Delta\_\{\\mathrm\{TPR\}\}=maxai∈𝒜⁡P⁡\(Y^=1∣Y=1,A=ai\)−minai∈𝒜⁡P⁡\(Y^=1∣Y=1,A=ai\),\\displaystyle=\\max\_\{a\_\{i\}\\in\\mathcal\{A\}\}P\(\\hat\{Y\}=1\\mid Y=1,A=a\_\{i\}\)\-\\min\_\{a\_\{i\}\\in\\mathcal\{A\}\}P\(\\hat\{Y\}=1\\mid Y=1,A=a\_\{i\}\),\(5\)ΔFPR\\displaystyle\\Delta\_\{\\mathrm\{FPR\}\}=maxai∈𝒜⁡P⁡\(Y^=1∣Y=0,A=ai\)−minai∈𝒜⁡P⁡\(Y^=1∣Y=0,A=ai\),\\displaystyle=\\max\_\{a\_\{i\}\\in\\mathcal\{A\}\}P\(\\hat\{Y\}=1\\mid Y=0,A=a\_\{i\}\)\-\\min\_\{a\_\{i\}\\in\\mathcal\{A\}\}P\(\\hat\{Y\}=1\\mid Y=0,A=a\_\{i\}\),\(6\)ΔEOD\\displaystyle\\Delta\_\{\\mathrm\{EOD\}\}=max⁡\(ΔTPR,ΔFPR\)\.\\displaystyle=\\max\\left\(\\Delta\_\{\\mathrm\{TPR\}\},\\Delta\_\{\\mathrm\{FPR\}\}\\right\)\.\(7\)
To assess the fairness of the ML models, we compared TPRs and FPRs across patient subgroups, as these metrics directly quantify gaps in prediction outcomes at the selected decision threshold\. We also evaluated subgroup AUROC as a complementary, threshold\-independent measure of overall discriminative performance, with the results presented in Appendix[7\.5](https://arxiv.org/html/2609.22113#S7.SS5)\.

##### 3\.3\.2Model Explanations

To better understand model behavior and enhance transparency, we computed SHAP values as a means of interpreting feature contributions\. SHAP \(SHapley Additive exPlanations\) provides a unified, theoretically grounded approach to feature attribution by quantifying the marginal contribution of each input feature to the model’s output[51](https://arxiv.org/html/2609.22113#bib.bib23)\. It is model\-agnostic and rooted in cooperative game theory, building on the concept of Shapley values, where each input feature is considered a “player” in a game contributing to the model’s prediction\. Formally, SHAP values are based on additive feature attribution methods, which assume the explanation model is a linear function of binary variables:

g⁡\(z′\)=ϕ0\+∑i=1Mϕi​zi′g\(z^\{\\prime\}\)=\\phi\_\{0\}\+\\sum\_\{i=1\}^\{M\}\\phi\_\{i\}z^\{\\prime\}\_\{i\}wherez′∈\{0,1\}Mz^\{\\prime\}\\in\\\{0,1\\\}^\{M\}represents the simplified input, a binary vector indicating the presence or absence of featureiiin the explanation,MMis the number of input features,ϕ0\\phi\_\{0\}is the expected model output when no features are present \(base value\), andϕi∈ℝ\\phi\_\{i\}\\in\\mathbb\{R\}represents the contribution \(SHAP value\) of featureii\. SHAP values capture how much each feature contributes to pushing the model’s prediction above or below the baseline \(i\.e\., the expected output across all instances\)\. A positive SHAP value indicates that a feature increases the model’s predicted probability of the target outcome \(e\.g\., premature discontinuation or retention longer than 180 days\), while a negative SHAP value suggests the opposite\. Importantly, SHAP values are computed for each instance, allowing for both local explanations \(individual\-level insights\) and global explanations \(aggregated feature importance across the dataset\)\.

#### 3\.4Bias Mitigation Strategies

We systematically assessed the effectiveness of several bias mitigation strategies to understand their ability to reduce performance inconsistencies and improve algorithmic fairness across patient subgroups in ML models\. To provide a comprehensive evaluation, we selected representative techniques that operate at three different stages of the ML development pipeline: pre\-processing, in\-processing, and post\-processing\. These methods are widely recognized in the ML fairness literature, and have been shown to be effective in improving fairness across a range of real\-world applications[19](https://arxiv.org/html/2609.22113#bib.bib34)\. Specifically, we evaluated three approaches\.

##### 3\.4\.1Reweighing

Reweighing[42](https://arxiv.org/html/2609.22113#bib.bib20)is a pre\-processing technique that assigns different weights to training instances based on their subgroup membership and outcome label\. The goal is to adjust the influence of each sample during training to create a more balanced dataset\. By increasing the weight of subgroups with limited data, this method encourages the model to learn from a more balanced representation of the population\. This helps mitigate learning patterns that disproportionately favor certain subgroups, without modifying the model architecture itself\. Specifically, let the training examples consist of triples\(X,A,Y\)\(X,A,Y\), whereX∈𝒳X\\in\\mathcal\{X\}is a feature vector,A∈𝒜A\\in\\mathcal\{A\}is the protected attribute \(e\.g\., race\) andY∈0,1Y\\in\{0,1\}is the outcome label\. Each training instance is assigned a weight based on its subgroup membershipAAand outcome labelYY:

w⁡\(a,y\)=P⁡\(A=a\)​P​\(Y=y\)P⁡\(A=a,Y=y\)\.w\(a,y\)=\\frac\{P\(A=a\)P\(Y=y\)\}\{P\(A=a,Y=y\)\}\.These weights reflect the ratio between the expected frequency of each\(A,Y\)\(A,Y\)combination under statistical independence and its observed frequency in the data\. Underrepresented subgroup–outcome combinations receive larger weights, while overrepresented combinations receive smaller weights\. The resulting weights are incorporated into the model’s loss function during training\.

##### 3\.4\.2Exponentiated Gradient Reduction \(EGR\)

Exponentiated Gradient Reduction \(EGR\)[5](https://arxiv.org/html/2609.22113#bib.bib21)is a model\-specific in\-processing technique that modifies the learning algorithm by incorporating fairness constraints during model training\. It works by adjusting how the model selects decision boundaries, seeking a balance between accuracy and fairness\. Specifically, EGR treats fairness as an optimization constraint—ensuring that prediction outcomes do not differ substantially across subgroups—while still aiming to minimize prediction error\. Letffdenote the prediction model,L⁡\(f\)L\(f\)denote the empirical prediction loss, andgk​\(f\)g\_\{k\}\(f\)denotekkfairness constraint functions that quantify gaps in error rates \(e\.g\., TPR or FPR\) across subgroups defined byAA, EGR formulates model training as a constrained optimization problem:

minf⁡L⁡\(f\)s\.t\.\|gk​\(f\)\|≤ϵ,∀k,\\min\_\{f\}L\(f\)\\quad\\text\{s\.t\.\}\\quad\|g\_\{k\}\(f\)\|\\leq\\epsilon,\\forall k,whereϵ\\epsilonspecifies the allowable tolerance for fairness violations\. This constrained problem is solved via a Lagrangian formulation, in which fairness constraints are incorporated through dual variables\. At each iteration, training samples are reweighted according to the current dual variables, and a base classifier is trained on the reweighted data\. The procedure alternates between minimizing prediction loss and adjusting weights to reduce fairness gaps\. EGR is a model\-agnostic approach that integrates directly into the training loop and can be applied to a wide range of predictive models\. The final predictor produced by EGR is a randomized mixture of classifiers obtained across iterations, which jointly balances predictive accuracy and fairness constraints\.

##### 3\.4\.3Threshold Optimizer

Threshold Optimizer[34](https://arxiv.org/html/2609.22113#bib.bib22)is a post\-processing method that adjusts the decision thresholds applied to predicted risk scores for different subgroups in order to satisfy a specified fairness criterion after the model has been trained\. By customizing thresholds per subgroup, the model can produce more consistent outputs across subgroups, helping to address fairness concerns without retraining or altering the underlying data\. Specifically, it learns subgroup\-specific decision thresholds for converting continuous risk scoress^=f⁡\(X\)∈\[0,1\]\\hat\{s\}=f\(X\)\\in\[0,1\]into binary predictionsY^\\hat\{Y\}\. Using the protected attributeAA, a thresholdτa\\tau\_\{a\}is chosen for each subgroupa∈Aa\\in Aso that:

Y^=𝕀\[s^≥τa\],\\hat\{Y\}=\\mathbb\{I\}\\big\[\\hat\{s\}\\geq\\tau\_\{a\}\\big\],where𝕀⁡\[⋅\]\\mathbb\{I\}\[\\cdot\]is the indicator function\. In practice, the Threshold Optimizer can also produce a randomized classifier \(a mixture of thresholds\) to achieve fairness constraints exactly, particularly when deterministic thresholds cannot satisfy the requirements\. It selects thresholds to satisfy the specified fairness constraints \(e\.g\., equal opportunity and predictive equality\) while minimizing a classification error metric in predictive performance\. Formally, this can be expressed as:

min\{τa\}⁡Error​\(Y^​\(\{τa\}\)\)s\.t\.\|μk​\(Y^\)\|≤ϵ,∀k,\\min\_\{\\\{\\tau\_\{a\}\\\}\}\\;\\text\{Error\}\(\\hat\{Y\}\(\\\{\\tau\_\{a\}\\\}\)\)\\quad\\text\{s\.t\.\}\\quad\|\\mu\_\{k\}\(\\hat\{Y\}\)\|\\leq\\epsilon,\\;\\forall k,whereμk​\(Y^\)\\mu\_\{k\}\(\\hat\{Y\}\)are fairness gap functions \(e\.g\., TPR or FPR differences across subgroups\) andϵ\\epsilonis an allowable tolerance\. From a practitioner perspective, this approach can be useful because it allows organizations to improve fairness in downstream decisions without retraining models or altering established workflows, which may be particularly valuable in resource\-constrained clinical environments\. In our study, the three approaches were optimized on validation data using Fairlearn[13](https://arxiv.org/html/2609.22113#bib.bib80)or AIF360[11](https://arxiv.org/html/2609.22113#bib.bib79)\.

Table 2:Characteristics of cohorts \(N=366,825N=366,825\) presented as the number of episodes, with percentages indicating their proportion within the entire cohort \(not within each outcome category\)\.Premature DiscontinuationTreatment RetentionDiscontinuationCompletion\>\>180 days≤\\leq180 daysSexMale115,613 \(31\.52%\)97,178 \(26\.49%\)81,857 \(22\.32%\)130,934 \(35\.69%\)Female83,008 \(22\.63%\)71,026 \(19\.36%\)59,910 \(16\.33%\)94,124 \(25\.66%\)RaceWhite163,842 \(44\.66%\)145,169 \(39\.48%\)116,872 \(31\.86%\)192,139 \(52\.28%\)Black34,779 \(9\.48%\)23,035 \(6\.28%\)24,895 \(6\.79%\)32,919 \(8\.97%\)EthnicityNon\-Hispanic182,285 \(49\.69%\)157,103 \(42\.83%\)130,078 \(35\.46%\)209,310 \(57\.06%\)Hispanic16,336 \(4\.45%\)11,101 \(3\.03%\)11,689 \(3\.19%\)15,748 \(4\.29%\)Age18\-2419,569 \(5\.33%\)16,526 \(4\.51%\)12,480 \(3\.40%\)23,615 \(6\.44%\)25\-54156,682 \(42\.71%\)135,436 \(36\.92%\)109,938 \(29\.97%\)182,179 \(49\.66%\)≥\\geq5522,370 \(6\.10%\)16,243 \(4\.43%\)19,349 \(5\.27%\)19,264 \(5\.25%\)EducationGrade 8 or less9,776 \(2\.67%\)6,714 \(1\.83%\)6,545 \(1\.78%\)9,945 \(2\.71%\)Grades 9\-1140,862 \(11\.14%\)33,146 \(9\.04%\)29,036 \(7\.92%\)44,972 \(12\.26%\)Grade 1295,113 \(25\.93%\)81,765 \(22\.29%\)66,775 \(18\.20%\)110,103 \(30\.02%\)College or more50,160 \(13\.67%\)41,467 \(11\.31%\)36,393 \(9\.92%\)55,234 \(15\.06%\)Unknown2,710 \(0\.74%\)5,112 \(1\.39%\)3,018 \(0\.82%\)4,804 \(1\.31%\)EmploymentFull\-time30,178 \(8\.23%\)28,396 \(7\.74%\)23,952 \(6\.53%\)34,622 \(9\.44%\)Part\-time16,123 \(4\.40%\)14,612 \(3\.98%\)12,431 \(3\.39%\)18,304 \(4\.99%\)Unemployed148,021 \(40\.35%\)119,799 \(32\.66%\)102,118 \(27\.84%\)165,702 \(45\.17%\)Unknown4,299 \(1\.17%\)5,397 \(1\.47%\)3,266 \(0\.89%\)6,430 \(1\.75%\)HousingIndependent153,761 \(41\.92%\)132,941 \(36\.24%\)114,551 \(31\.23%\)172,151 \(46\.93%\)Dependent25,535 \(6\.96%\)19,585 \(5\.34%\)16,105 \(4\.39%\)29,015 \(7\.91%\)Homeless17,032 \(4\.64%\)11,112 \(3\.03%\)8,720 \(2\.38%\)19,424 \(5\.30%\)Unknown2,293 \(0\.63%\)4,566 \(1\.24%\)2,391 \(0\.65%\)4,468 \(1\.22%\)RegionNortheast99,009 \(26\.99%\)104,350 \(28\.45%\)79,608 \(21\.70%\)123,751 \(33\.74%\)Midwest30,495 \(8\.31%\)21,196 \(5\.78%\)18,537 \(5\.05%\)33,154 \(9\.04%\)South17,614 \(4\.80%\)20,197 \(5\.51%\)9,393 \(2\.56%\)28,418 \(7\.75%\)West51,380 \(14\.01%\)22,430 \(6\.11%\)34,138 \(9\.31%\)39,672 \(10\.81%\)U\.S\. territories123 \(0\.03%\)31 \(0\.01%\)91 \(0\.02%\)63 \(0\.02%\)Arrest111Arrest: arrests in past 30 days prior to admission\.Arrest9,609 \(2\.62%\)10,190 \(2\.78%\)6,147 \(1\.68%\)13,652 \(3\.72%\)No arrest187,259 \(51\.05%\)155,115 \(42\.29%\)132,700 \(36\.18%\)209,674 \(57\.16%\)Unknown1,753 \(0\.48%\)2,899 \(0\.79%\)2,920 \(0\.80%\)1,732 \(0\.47%\)Prior Tx222Tx stands for treatment\.One or more145,219 \(39\.59%\)126,297 \(34\.43%\)105,161 \(28\.67%\)166,355 \(45\.35%\)No prior Tx50,941 \(13\.89%\)37,765 \(10\.30%\)33,418 \(9\.11%\)55,288 \(15\.07%\)Unknown2,461 \(0\.67%\)4,142 \(1\.13%\)3,188 \(0\.87%\)3,415 \(0\.93%\)Psych333Psych: has co\-occurring mental or behavioral health disorders\.Psych77,786 \(21\.21%\)74,807 \(20\.39%\)53,599 \(14\.61%\)98,994 \(26\.99%\)No psych107,025 \(29\.18%\)79,042 \(21\.55%\)76,215 \(20\.78%\)109,852 \(29\.95%\)Unknown13,810 \(3\.76%\)14,355 \(3\.91%\)11,953 \(3\.26%\)16,212 \(4\.42%\)Injection444Injection: have drugs used by injection\.Injection use108,649 \(29\.62%\)91,614 \(24\.97%\)75,440 \(20\.57%\)124,823 \(34\.03%\)OtherSubstancesHeroin163,042 \(44\.45%\)133,990 \(36\.53%\)115,034 \(31\.36%\)181,998 \(49\.61%\)Alcohol22,959 \(6\.26%\)22,502 \(6\.13%\)15,464 \(4\.22%\)29,997 \(8\.18%\)Inhalant56 \(0\.02%\)98 \(0\.03%\)49 \(0\.01%\)105 \(0\.03%\)Marijuana35,877 \(9\.78%\)33,938 \(9\.25%\)24,320 \(6\.63%\)45,495 \(12\.40%\)Stimulant67,514 \(18\.40%\)55,523 \(15\.14%\)40,951 \(11\.16%\)82,086 \(22\.38%\)Tranquilizer12,205 \(3\.33%\)13,707 \(3\.74%\)9,291 \(2\.53%\)16,621 \(4\.53%\)Sedative835 \(0\.23%\)842 \(0\.23%\)612 \(0\.17%\)1,065 \(0\.29%\)Hallucinogen4,298 \(1\.17%\)4,377 \(1\.19%\)3,149 \(0\.86%\)5,526 \(1\.51%\)Total198,621 \(54\.15%\)168,204 \(45\.85%\)141,767 \(38\.65%\)225,058 \(61\.35%\)Figure 2:FPRs and TPRs across different subgroups are evaluated for four ML models \(LR, RF, GBDT, MLP\) as the baseline models for predicting premature treatment discontinuation\. The subfigures presents results stratified by \(a\) race, \(b\) ethnicity, \(c\) age, and \(d\) sex\. Bars represent the mean values over 10 runs, while error bars indicate the standard deviation computed from these runs\. Higher values indicate better results\.

### 4Results

#### 4\.1Descriptive Results

A total of 366,825 patients satisfied the respective inclusion/exclusion criteria for the model development\. The cohort’s data characteristics are summarized in Table[2](https://arxiv.org/html/2609.22113#S3.T2)\. Of all treatment episodes, 54\.15% resulted in premature discontinuation, while 38\.65% retained longer than 180 days\. In terms of sociodemographic and clinical characteristics, the majority were aged 25–54 \(79\.63%\), White \(84\.14%\) and non\-Hispanic patients \(92\.52%\)\. The majority of patients \(72\.9%\) had an education level of grade 12 or below, and 73\.01% were unemployed at the time of admission\. Most patients \(78\.16%\) reported living in independent housing, and 55\.44% were from the Northeast region\. Only 5\.4% reported having been arrested in the 30 days prior to admission\. Additionally, 41\.6% had co\-occurring psychological disorders, 54\.6% reported injection drug use, and 80\.98% had used heroin\.

Table 3:Overall predictive performance, measured by area under the receiver operating characteristic curve \(AUROC\), for ML models predicting premature treatment discontinuation and treatment retention exceeding 180 days\.Premature DiscontinuationRetention\>\>180 daysLR0\.6562 \(0\.0011\)0\.6559 \(0\.0011\)RF0\.6863 \(0\.0013\)0\.6797 \(0\.0011\)GBDT0\.6892 \(0\.0013\)0\.6836 \(0\.0010\)MLP0\.6647 \(0\.0025\)0\.6657 \(0\.0014\)Figure 3:FPRs and TPRs across different subgroups are evaluated for four ML models \(LR, RF, GBDT, MLP\) as the baseline models for predicting treatment retention\>\>180 days\. The subfigures presents results stratified by \(a\) race, \(b\) ethnicity, \(c\) age, and \(d\) sex\. Bars represent the mean values over 10 runs, while error bars indicate the standard deviation computed from these runs\. Higher values indicate better results\.
#### 4\.2Model Performance Across Patient Subgroups

We began by systematically evaluating how each of the four ML models performed, with particular attention to how performance varied across patient subgroups for the two target outcomes\. Table[3](https://arxiv.org/html/2609.22113#S4.T3)reports the overall AUROC values for predicting premature treatment discontinuation and treatment retention exceeding 180 days\. Among the four models evaluated, the GBDT model achieved the highest AUROC\. Figure[2](https://arxiv.org/html/2609.22113#S3.F2)presents subgroup\-specific FPRs and TPRs for premature discontinuation predictions, stratified by different sensitive attributes for each of the four models\. Across all models, FPRs were consistently higher for Black patients compared to White patients, with a gap of 29\.67% for LR, 25\.26% for RF, 26\.81% for GBDT, and 28\.28% for MLP, indicating a tendency to incorrectly classify Black patients as likely to drop out of treatment prematurely\. Gaps were also observed in TPRs, with higher TPRs for Black patients across all models \(LR: 24\.91%, RF: 22\.13%, GBDT: 23\.04%, MLP: 24\.63%\)\. Our results show that the FPR and TPR gaps across ethnicity \(11%–19%\), age \(7%–18%\), and sex \(<<3%\) subgroups are less pronounced, compared to those observed for race\.

Figure[3](https://arxiv.org/html/2609.22113#S4.F3)shows FPRs and TPRs for predicting treatment retention\>\>180 days\. Our results indicate that both Black patients and patients over 55 years old consistently exhibited higher FPRs and TPRs compared to White patients and younger age subgroups, respectively\. The FPR gap between Black patients and White patients ranged from 13% to 15% across all models, while TPR gaps were even larger, ranging from 20% to 24%\. Age\-related gaps were evident, with older patients \(\>\>55\) having 33%–46% higher FPRs and 42%–57% higher TPRs than younger subgroups across all models\. The gaps were smaller among ethnic subgroups \(7%–15%\) and even lower among sex subgroups \(<<3%\)\. For comparison, we also trained models excluding sensitive attributes as input features \(see Tables[11](https://arxiv.org/html/2609.22113#S7.T11)\-[12](https://arxiv.org/html/2609.22113#S7.T12)\)\. To provide a threshold\-independent assessment of model performance across demographic groups, we additionally computed the AUROC separately for each subgroup defined by race, ethnicity, and age\. The detailed results are presented in Appendix[7\.5](https://arxiv.org/html/2609.22113#S7.SS5)\. In addition, to evaluate the robustness of our results under post\-pandemic changes in MOUD treatment, we additionally analyzed the 2020–2023 TEDS\-D data using the same experimental pipeline\. The results, presented in Appendix[7\.8](https://arxiv.org/html/2609.22113#S7.SS8), showed similar fairness patterns and slightly improved predictive performance, supporting the robustness of our findings across different time periods\.

![Refer to caption](https://arxiv.org/html/2609.22113v1/Fig4.png)Figure 4:Figures \(a\) and \(b\) present feature importance plots of the GBDT model for two prediction tasks: \(a\) premature treatment discontinuation and \(b\) treatment retention\>\>180 days\. The x\-axis represents feature attribution values, where positive values indicate that a feature increases the likelihood of the positive class, while negative values indicate a decrease\. Features are ranked by their average importance, with the most influential ones appearing at the top\.Figure 5:Figures \(a\) and \(b\) display histograms of feature attribution values for sensitive attributes \(race, ethnicity, and age\) for each task: \(a\) premature treatment discontinuation and \(b\) treatment retention\>\>180 days, to visualize how these attributes contribute to model predictions\.
#### 4\.3Interpreting Model Behavior Through Feature Attribution

To interpret how input features influenced the model’s predictions, we computed SHAP values as feature attributions, focusing on the GBDT model, which demonstrated the best overall performance among the four models evaluated\. In our analysis, all categorical predictors were one\-hot encoded prior to model training, and SHAP values were initially computed at the level of the one\-hot encoded features\. To facilitate interpretation, we aggregated SHAP values back to the original categorical variables by summing the SHAP values of all one\-hot encoded columns corresponding to each feature, following common practice\. So, the reported SHAP values represent the overall contribution of each original categorical feature, reflecting the combined effect of all its encoded categories\. Figure[4](https://arxiv.org/html/2609.22113#S4.F4)presents a global explanation of the GBDT model’s predictions for \(a\) premature treatment discontinuation and \(b\) treatment retention exceeding 180 days\. It displays the top 10 most influential features, ranked by their mean absolute SHAP values aggregated over the test set\. The feature attribution values indicate each feature’s contribution to increasing \(positive SHAP values\) or decreasing \(negative SHAP values\) the likelihood of positive class, while the color gradient represents the feature value\. For both outcome variables, the features that consistently had strong influence on model’s predictions included geographic region, patient’s referral source, type of service setting \(intensive or non\-intensive outpatient\), max frequency of substance use \(i\.e\., heroin, non\-prescription methadone, other opiates and synthetics\), the presence of co\-occurring mental or behavioral health disorders\.

As part of our analysis, we examined whether the observed performance differences reflected systematic variation in how the model treated different patient subgroups—that is, whether the model relied on sensitive attributes in a way that led to consistently different predictions across subgroups\. Figure[5](https://arxiv.org/html/2609.22113#S4.F5)shows the distributions of feature attributions for sensitive attributes over individual test points from different subgroups\. In predicting premature treatment discontinuation, race demonstrated relatively high feature attribution values, indicating that the model relies heavily on this attribute, whereas age and ethnicity contributed less\. Notably, the effect of race varied across subgroups: feature attribution values were predominantly positive for Black patients and negative for White patients, indicating differing influences on predicted risk\. Similarly, for treatment retention prediction, age emerged as a key predictive factor, as reflected in its higher SHAP values\. The impact of age also varied across subgroups, with patients aged over 55 receiving the most positive feature attribution values, suggesting the model associated older age with a higher likelihood of being retained in treatment for more than 180 days\. Table[4](https://arxiv.org/html/2609.22113#S4.T4)compares the feature profiles of false positives for White and Black patients\. Black patients who were falsely classified as at risk of premature discontinuation or as likely to remain in treatment beyond 180 days were generally older, had lower education levels, and exhibited higher rates of unemployment compared to the White patients\. Additionally, heroin use was more prevalent among Black patients, while injection drug use was more common among White patients\. These patterns suggest that the model may rely on different combinations of features—some of which may act as proxies for sensitive attributes—when making predictions for different subgroups\.

Table 4:Sample characteristics of false positive subgroup from GBDT model for premature treatment discontinuation and retention\>\>180 days, comparing Black and White patients\.Premature DiscontinuationRetention\>\>180 daysFalse PositiveWhite patients\(%\)False PositiveBlack patients\(%\)False PositiveWhite patients\(%\)False PositiveBlack patients\(%\)Age18\-2411\.32\.87\.10\.525\-5482\.667\.672\.654\.5\>\>556\.129\.620\.345\.0EducationLess than grade 1123\.935\.226\.233\.7Grade 1244\.645\.543\.745\.6College or more30\.018\.326\.719\.7Employment statusFull\-time/Part\-time25\.312\.532\.813\.3Unemployed73\.483\.764\.886\.0Used Heroin80\.791\.479\.394\.0Used by injection59\.228\.253\.327\.5Figure 6:FPRs across different subgroups are shown for three bias mitigation methods—Reweighing \(pre\-processing\), EGR \(in\-processing\), and Threshold \(post\-processing\)—in predicting premature treatment discontinuation, compared to a baseline without fairness considerations\. Results are stratified by \(a\) race, \(b\) ethnicity, and \(c\) age\. Bars represent the mean values over 10 independent runs, with error bars indicating the standard deviation across these runs\. Higher values indicate better results\.Figure 7:FPRs across different subgroups are shown for three bias mitigation methods—Reweighing \(pre\-processing\), EGR \(in\-processing\), and Threshold \(post\-processing\)—in predicting treatment retention\>\>180 days, compared to a baseline without fairness considerations\. Results are stratified by \(a\) race, \(b\) ethnicity, and \(c\) age\. Bars represent the mean values over 10 independent runs, with error bars indicating the standard deviation across these runs\. Higher values indicate better results\.
#### 4\.4Bias Mitigation Results

To address the observed performance inconsistencies across subgroups, we implemented and assessed three bias mitigation strategies, including Reweighing[42](https://arxiv.org/html/2609.22113#bib.bib20), EGR[5](https://arxiv.org/html/2609.22113#bib.bib21), and Threshold[34](https://arxiv.org/html/2609.22113#bib.bib22)\. As shown in Figure[6](https://arxiv.org/html/2609.22113#S4.F6), for predicting premature discontinuation, Reweighing significantly reduced gaps in FPR \(LR: 1\.36%, RF: 4\.52%, GBDT: 3\.19%, MLP: 4\.34%\) between Black and White patients\. Similarly, EGR reduced FPR gaps to 1\.08% for LR, 4\.89% for RF, and 4\.49% for GBDT\. Note that EGR is not compatible with MLPs, as they are non\-convex and typically do not support per\-sample reweighting in a stable manner as required\. Threshold reduced FPR gaps, achieving the lowest gaps in LR \(0\.56%\), RF \(3\.06%\), and GBDT \(1\.36%\) between White and Black patients, but was less effective for MLP \(19\.89%\)\. It is important to note that, although all three mitigation methods reduced subgroup gaps, these improvements often came at the expense of higher error rates for at least one subgroup\. For example, the Threshold method narrowed the gap by worsening the FPR for White patients while improving it for Black patients, and in some cases, both subgroups experienced worsened FPRs even though the overall gap was reduced\. For age subgroups, Reweighing reduced the FPR gap \(LR: 2\.63%, RF: 0\.43%, GBDT: 7\.54%, MLP: 7\.4%\) between younger \(18\-54\) and older \(\>\>55\) subgroups, compared to the baseline models \(LR: 10\.75%, RF: 8\.69%, GBDT: 5\.29%, MLP: 4\.39%\)\. EGR brings FPR gaps to 0\.7% \(LR\), 2\.47% \(RF\), and 1\.0% \(GBDT\), while Threshold produced the most minimal FPR gaps \(LR: 0\.71%, RF: 1\.24%, GBDT: 0\.15%, MLP: 2\.84%\)\. Similarly, for ethnicity, all three bias mitigation methods reduced the overall FPR gaps \(Reweighing—LR: 5\.22%, RF: 1\.16%, GBDT: 3\.33%, MLP: 4\.28%; EGR—LR: 1\.53%, RF: 7\.41%, GBDT: 4\.53%; Threshold—LR: 0\.99%, RF: 2\.03%, GBDT: 0\.88%, MLP: 9\.75%\)\. However, these reductions were often achieved by worsening the FPR for one subgroup or, in some cases, for both subgroups\. Such patterns were consistently observed across all three mitigation methods and across the subgroups examined\. A similar trend was observed in Figure[7](https://arxiv.org/html/2609.22113#S4.F7)for predicting treatment retention\>\>180 days, where all three mitigation strategies led to reductions in FPR gaps across subgroups with varying effectiveness\. Specifically, the FPR gap across race subgroups was substantially reduced by Reweighing \(LR: 1\.43%, RF: 5\.71%, GBDT: 4\.97%, MLP: 5\.00%\) and by EGR \(LR: 1\.21%, RF: 0\.37%, GBDT: 0\.78%\)\. The Threshold method also reduced race\-based FPR gaps for several models \(LR: 0\.47%, RF: 1\.34%, GBDT: 0\.49%\); however, for the MLP model, threshold adjustment instead exacerbated the gap, increasing the FPR gap to 14\.87%\. Similary, the FPR gap across age subgroups was significantly reduced with Reweighing \(LR: 2\.17%, RF: 3\.52%, GBDT: 1\.74%, MLP: 3\.55%\) and EGR \(LR: 2\.06%, RF: 2\.94%, GBDT: 2\.73%\)\. The Threshold method also substantially reduced the FPR gap for some models \(LR: 0\.96%, RF: 4\.01%, GBDT: 1\.70%\), but the gap persisted at 37\.16% for MLP\. For ethnicity subgroups, although the baseline gaps were relatively modest, all three mitigation methods generally reduced FPR gaps compared to the Baseline \(LR: 10\.30%, RF: 7\.59%, GBDT: 8\.05%, MLP: 9\.52%\)\. Reweighing reduced the gaps to 0\.50% for LR, 2\.60% for RF, 0\.58% for GBDT, and 1\.19% for MLP\. EGR reduced the gaps to 0\.71% for LR, 1\.08% for RF, and 1\.23% for GBDT\. The Threshold method also lowered the gaps for several models \(LR: 0\.90%, RF: 1\.82%, GBDT: 0\.97%\), but resulted in a large gap for MLP \(8\.65%\)\. All TPR results are presented in Figures[10](https://arxiv.org/html/2609.22113#S7.F10)and[11](https://arxiv.org/html/2609.22113#S7.F11), with more detailed results in Tables[13](https://arxiv.org/html/2609.22113#S7.T13)\-[16](https://arxiv.org/html/2609.22113#S7.T16)\.

Table 5:Group fairness metrics \(mean; standard deviation in parentheses\) for predicting premature treatment discontinuation\. Predictive Equality Ratio: FPR ratio of the subgroup with the lowest FPR to the subgroup with the highest FPR\. Equal Opportunity Ratio: TPR ratio of the subgroup with the lowest TPR to the subgroup with the highest TPR\. Values closer to 1 indicate more equitable model performance across subgroups\.Boldvalues indicate the best results\.GroupMethodPredictive Equality RatioEqual Opportunity RatioLRRFGBDTMLPLRRFGBDTMLPRaceBaseline0\.5444\(0\.0054\)0\.5296\(0\.0117\)0\.5291\(0\.0123\)0\.5010\(0\.0702\)0\.6966\(0\.0034\)0\.7075\(0\.0080\)0\.7070\(0\.0089\)0\.6715\(0\.0570\)Reweighing0\.9679\(0\.0150\)0\.8756\(0\.0099\)0\.9141\(0\.0175\)0\.8941\(0\.0539\)0\.9537\(0\.0111\)0\.9566\(0\.0125\)0\.9802\(0\.0129\)0\.9509\(0\.0322\)EGR0\.9787\(0\.0147\)0\.9104\(0\.0715\)0\.9198\(0\.0522\)\-\-555Results for MLP with EGR are not reported, as EGR requires models that support stable, per\-sample reweighting during training\. MLP are non\-convex and typically do not support per\-sample reweighting in a stable and standardized manner as required\.0\.9849\(0\.0091\)0\.9485\(0\.0505\)0\.9659\(0\.0155\)\-\-555Results for MLP with EGR are not reported, as EGR requires models that support stable, per\-sample reweighting during training\. MLP are non\-convex and typically do not support per\-sample reweighting in a stable and standardized manner as required\.Threshold0\.9888\(0\.0060\)0\.9256\(0\.0126\)0\.9645\(0\.0154\)0\.7024\(0\.0368\)0\.9959\(0\.0031\)0\.9805\(0\.0080\)0\.9896\(0\.0072\)0\.8344\(0\.0301\)EthnicityBaseline0\.6900\(0\.0120\)0\.7382\(0\.0205\)0\.6803\(0\.0314\)0\.6578\(0\.0431\)0\.7600\(0\.0081\)0\.8000\(0\.0109\)0\.7733\(0\.0194\)0\.7479\(0\.0394\)Reweighing0\.8710\(0\.0181\)0\.9638\(0\.0236\)0\.9023\(0\.0386\)0\.8753\(0\.0669\)0\.9739\(0\.0094\)0\.9583\(0\.0116\)0\.98230\.01530\.9677\(0\.0292\)EGR0\.9640\(0\.0208\)0\.8701\(0\.0119\)0\.9127\(0\.0592\)–0\.9871\(0\.0101\)0\.8991\(0\.0066\)0\.9510\(0\.0307\)–Threshold0\.9737\(0\.0220\)0\.9495\(0\.0193\)0\.9776\(0\.0230\)0\.8325\(0\.0359\)0\.9894\(0\.0088\)0\.9906\(0\.0065\)0\.9903\(0\.0063\)0\.8851\(0\.0259\)AgeBaseline0\.7734\(0\.0171\)0\.6857\(0\.0214\)0\.7644\(0\.0213\)0\.8061\(0\.0806\)0\.8056\(0\.0117\)0\.7527\(0\.0157\)0\.7932\(0\.0120\)0\.8219\(0\.0526\)Reweighing0\.9167\(0\.0249\)0\.8674\(0\.0250\)0\.8889\(0\.0258\)0\.8727\(0\.0655\)0\.9699\(0\.0106\)0\.8511\(0\.0138\)0\.9253\(0\.0141\)0\.9481\(0\.0310\)EGR0\.9590\(0\.0141\)0\.8259\(0\.0248\)0\.9225\(0\.0384\)–0\.9713\(0\.0119\)0\.8747\(0\.0147\)0\.9421\(0\.0430\)–Threshold0\.9706\(0\.0147\)0\.9628\(0\.0169\)0\.9713\(0\.0099\)0\.9167\(0\.0300\)0\.9806\(0\.0060\)0\.9790\(0\.0106\)0\.9820\(0\.0081\)0\.9259\(0\.0169\)
Table 6:Group fairness metrics \(mean; standard deviation in parentheses\) for predicting treatment retention\>\>180 days\. Predictive Equality Ratio: FPR ratio of the subgroup with the lowest FPR to the subgroup with the highest FPR\. Equal Opportunity Ratio: TPR ratio of the subgroup with the lowest TPR to the subgroup with the highest TPR\. Values closer to 1 indicate more equitable model performance across subgroups\.Boldvalues indicate the best results\.GroupMethodPredictive Equality RatioEqual Opportunity RatioLRRFGBDTMLPLRRFGBDTMLPRaceBaseline0\.4501\(0\.0108\)0\.4468\(0\.0113\)0\.4351\(0\.0133\)0\.4803\(0\.0823\)0\.5440\(0\.0091\)0\.5514\(0\.0069\)0\.5554\(0\.0096\)0\.5775\(0\.0769\)Reweighing0\.8984\(0\.0164\)0\.6828\(0\.0213\)0\.7241\(0\.0355\)0\.7659\(0\.1286\)0\.9490\(0\.0143\)0\.7503\(0\.0107\)0\.8007\(0\.0179\)0\.8109\(0\.0943\)EGR0\.9129\(0\.0161\)0\.9732\(0\.0377\)0\.9470\(0\.0351\)\-\-555Results for MLP with EGR are not reported, as EGR requires models that support stable, per\-sample reweighting during training\. MLP are non\-convex and typically do not support per\-sample reweighting in a stable and standardized manner as required\.0\.9625\(0\.0225\)0\.9747\(0\.0238\)0\.9771\(0\.0176\)\-\-555Results for MLP with EGR are not reported, as EGR requires models that support stable, per\-sample reweighting during training\. MLP are non\-convex and typically do not support per\-sample reweighting in a stable and standardized manner as required\.Threshold0\.9882\(0\.0089\)0\.9679\(0\.0121\)0\.9869\(0\.0105\)0\.4863\(0\.0491\)0\.9898\(0\.0093\)0\.9862\(0\.0083\)0\.9865\(0\.0068\)0\.5866\(0\.0431\)EthnicityBaseline0\.5457\(0\.0224\)0\.6262\(0\.0318\)0\.6234\(0\.0396\)0\.6019\(0\.0614\)0\.6552\(0\.0154\)0\.7309\(0\.0160\)0\.7485\(0\.0259\)0\.7200\(0\.0608\)Reweighing0\.9635\(0\.0246\)0\.8326\(0\.0326\)0\.9602\(0\.0290\)0\.9211\(0\.0582\)0\.9705\(0\.0114\)0\.8918\(0\.0186\)0\.9753\(0\.0153\)0\.9278\(0\.0534\)EGR0\.9595\(0\.0272\)0\.9270\(0\.0424\)0\.9194\(0\.0450\)–0\.9640\(0\.0220\)0\.9656\(0\.0229\)0\.9772\(0\.0178\)–Threshold0\.9897\(0\.0100\)0\.9577\(0\.0234\)0\.9745\(0\.0161\)0\.6498\(0\.0496\)0\.9862\(0\.0112\)0\.9794\(0\.0112\)0\.9825\(0\.0109\)0\.7658\(0\.0420\)AgeBaseline0\.0755\(0\.0049\)0\.1922\(0\.0107\)0\.1744\(0\.0113\)0\.1473\(0\.0496\)0\.1650\(0\.0088\)0\.3305\(0\.0154\)0\.3313\(0\.0161\)0\.2650\(0\.0641\)Reweighing0\.8255\(0\.0359\)0\.7479\(0\.0356\)0\.8708\(0\.0407\)0\.7917\(0\.0710\)0\.6824\(0\.0226\)0\.8373\(0\.0189\)0\.8608\(0\.0316\)0\.8211\(0\.1003\)EGR0\.8358\(0\.0514\)0\.8199\(0\.0336\)0\.8236\(0\.0324\)–0\.8925\(0\.0401\)0\.9195\(0\.0398\)0\.9579\(0\.0207\)–Threshold0\.9813\(0\.0150\)0\.9078\(0\.0149\)0\.9596\(0\.0158\)0\.1756\(0\.0399\)0\.9843\(0\.0070\)0\.9849\(0\.0087\)0\.9859\(0\.0065\)0\.3211\(0\.0474\)
Figure 8:FPRs for cross\-tabulated subgroups across all bias mitigation methods in predicting two tasks \(a\) premature discontinuation and \(b\) treatment retention\>\>180 days with GBDT model\. Lower FPR values indicate better results\.We also calculated group fairness metrics, including EOR and PER, for each experiment to evaluate the equity of model performance across subgroups, as presented in Table[5](https://arxiv.org/html/2609.22113#S4.T5)and Table[6](https://arxiv.org/html/2609.22113#S4.T6)\. All three bias mitigation methods enhanced both fairness metrics, bringing them closer to 1 compared to the baseline models\. Among them, Threshold demonstrated significant fairness improvements in most experiments\. However, it was less effective for the MLP model, likely because MLP models tend to produce more polarized probability outputs \(i\.e\., predictions concentrated closer to 0 or 1\), which leaves less flexibility for threshold adjustments to meaningfully alter subgroup\-specific error rates\. Moreover, the Threshold method often reduced subgroup gaps by worsening FPRs and TPRs for all the compared subgroups, as shown in the detailed subgroup results in Figures[6](https://arxiv.org/html/2609.22113#S4.F6)–[7](https://arxiv.org/html/2609.22113#S4.F7)and Figures[10](https://arxiv.org/html/2609.22113#S7.F10)–[11](https://arxiv.org/html/2609.22113#S7.F11)\. Reweighing and EGR generally performed more consistently across the four models in improving fairness metrics\. Notably, when examined alongside the subgroup\-specific results, a trade\-off between fairness and error rates remained evident across all three mitigation methods\. In addition to EOR and PER, we also evaluated EOD to provide a unified assessment of gaps in both TPRs and FPRs\. The detailed EOD results are provided in Table[7](https://arxiv.org/html/2609.22113#S7.T7)in Appendix[7\.3](https://arxiv.org/html/2609.22113#S7.SS3)\.

Finally, we conducted a comprehensive cross\-tabulated subgroup evaluation using the best\-performing GBDT model on both tasks to assess the effectiveness of bias mitigation strategies on more complex subgroups\. These subgroups, defined by combinations of sensitive attributes such as race\-age and race\-ethnicity, allow us to explore more complex patterns in model behavior that may not be evident when attributes are assessed independently\. Figure[8](https://arxiv.org/html/2609.22113#S4.F8)presents the FPR across all cross\-tabulated subgroup pairs for each mitigation strategy\. Overall, we observe that the evaluated methods reduce average FPR and narrow the performance gaps across intersecting subgroups\. This shows that these methods not only improve fairness along single dimensions, but also extend effectively to more complex, real\-world group intersections\.

### 5Discussion

This study developed ML models and conducted a systematic assessment of their algorithmic fairness to predict MOUD treatment outcomes, including treatment retention and premature discontinuation, specifically in outpatient treatment settings\. This setting represents a critical yet underexplored context in the MOUD literature, given the importance of sustained engagement for effective outcomes and the persistently high rates of premature discontinuation[55](https://arxiv.org/html/2609.22113#bib.bib24)\. In outpatient MOUD care, treatment retention is shaped by a complex interplay of clinical, behavioral, and structural factors, and prior research has documented substantial differences in retention and completion across patient subgroups[70](https://arxiv.org/html/2609.22113#bib.bib18);[8](https://arxiv.org/html/2609.22113#bib.bib16);[10](https://arxiv.org/html/2609.22113#bib.bib41)\. This study highlights the importance of evaluating not only overall predictive performance but also subgroup\-level fairness when developing ML models in such settings\. Across multiple models and outcome definitions, we observed consistent discrepancies in model performance across patient subgroups defined by sensitive attributes such as race, ethnicity, and age\. Our findings demonstrate that models with similar overall predictive performance can nonetheless produce substantially different error rates across patient sociodemographic subgroups\. In particular, we observed worse FPR for some patient subgroups, including Black patients and older patients, suggesting that the models may be relying on spurious patterns or proxies when making predictions\. This observation aligns with prior findings showing that ML models often learn spurious correlations or non\-generalizable patterns from training data, especially when model training data reflects long\-standing patterns or encodes existing structural variation across patient subgroups[85](https://arxiv.org/html/2609.22113#bib.bib43)\. In the context of outpatient MOUD treatment, this is particularly concerning because outcome predictions may directly inform the allocation of evidence\-based follow\-up strategies known to improve typically low treatment retention, such as proactive appointment reminders and check\-ins, early engagement with counseling or behavioral therapy, transportation support or telehealth adjustments, and peer recovery coaching[78](https://arxiv.org/html/2609.22113#bib.bib26)\. These tools may inadvertently amplify existing disparities in care and support if fairness is not explicitly examined\.

Furthermore, the prominent contributions of geographic region, patient referral source, service setting, substance use frequency, and co\-occurring mental or behavioral health disorders to model predictions are consistent with prior studies identifying these factors as important determinants of MOUD treatment outcomes, including treatment retention and premature discontinuation[69](https://arxiv.org/html/2609.22113#bib.bib8);[10](https://arxiv.org/html/2609.22113#bib.bib41)\. Sensitive attributes such as race and age emerged as influential features in model predictions, but their impact varied across subgroups\. Race contributed strongly to predictions of premature discontinuation, while age played a significant role in predicting long\-term treatment retention\. Notably, the impact of these features differed across subgroups, with being identified as Black individuals or being over 55 years old was associated with positive feature contributions in their respective tasks, while other subgroup characteristics showed negative or weaker associations\. This pattern suggests that the model may consistently assign higher risk scores for Black patients for premature discontinuation and older patients with a greater likelihood of long\-term retention\. This raises concerns about whether the model is over\-relying on subgroup\-specific patterns in ways that could affect prediction reliability or lead to differential treatment recommendations\. To further investigate this, we trained models without including sensitive attributes as input features \(Tables[11](https://arxiv.org/html/2609.22113#S7.T11)\-[12](https://arxiv.org/html/2609.22113#S7.T12)\)\. The overall prediction performance remained largely unchanged, suggesting that these attributes were not essential for predictive accuracy\. This finding indicates that the models may have relied on such attributes as shortcuts for prediction, reflecting structural patterns present in the training data rather than identifying clinically relevant relationships\. Such shortcut learning has been widely observed in ML systems, where models capture correlations that are statistically strong but clinically irrelevant or contextually misleading, as shown in previous studies[53](https://arxiv.org/html/2609.22113#bib.bib33);[79](https://arxiv.org/html/2609.22113#bib.bib40)\. Notably, removing the sensitive attributes from the input features did not eliminate the observed performance discrepancies across subgroups, indicating that other correlated features may act as proxies\. This finding underscores the difficulty of mitigating performance inconsistency simply by excluding subgroup identifiers alone, as structural signals may still be encoded through other variables\.

Our subgroup analysis of false positives revealed notable differences in the feature profiles of misclassified patients across subgroups\. Specifically, Black patients who were incorrectly classified as at risk of premature discontinuation or as likely to remain in treatment long\-term were generally older, less educated, more likely to be unemployed, and more likely to report heroin use\. In contrast, misclassified White patients were more often associated with injection drug use\. These findings echo long\-standing evidence in public health literature showing that social determinants—such as education, employment, and access to stable resources—are deeply connected to health outcomes and treatment engagement[59](https://arxiv.org/html/2609.22113#bib.bib13);[9](https://arxiv.org/html/2609.22113#bib.bib42)\. The prevalence of such social determinants among falsely flagged Black patients suggests that the model may be learning structural patterns from the training data, rather than focusing on clinical or behavioral risk factors\. Furthermore, these trends indicate that false positive predictions may disproportionately affect patients with certain social determinants, reflecting how model predictions are influenced by such non\-clinical factors\. Although such feature dependencies may reflect real\-world associations in the training data, they also raise important concerns about whether models are indirectly using sensitive characteristics through correlated variables\. These findings highlight the need for scrutinizing which features drive model predictions to understand how models make decisions and to identify when outputs may reflect non\-clinical factors embedded in data\.

To address these concerns, we implemented mitigation techniques aimed at reducing performance gaps across patient population subgroups\. These methods targeted different stages of the ML development pipeline, including pre\-processing, in\-processing, and post\-processing\. While these approaches could improve subgroup\-level fairness metrics, such as equal opportunity and predictive equality, they also introduced important trade\-offs that require careful consideration\. For example, reducing the FPR for one subgroup sometimes increased it for another, potentially leading to unintended consequences such as misallocated clinical resources or inappropriate interventions\. These findings highlight that fairness interventions must be assessed not only through the lens of statistical parity, but also with attention to their real\-world clinical and operational implications\. A technically fair model may still produce disproportionate harm if misclassification carries different consequences for different subgroups\. As such, healthcare providers and clinicians should scrutinize model outputs rather than treating them as neutral tools\. From a policy perspective, these findings suggest the need for formal guidelines and validation standards that ensure predictive fairness\. The demonstrated effectiveness of mitigation techniques further supports the integration of routine subgroup audits and fairness\-aware practices into health system development processes\.

On the other hand, no single bias mitigation strategy serves as a silver bullet, and the choice of method should be guided by real\-world implementation constraints and organizational capacity\. Pre\-processing methods such as Reweighing are often useful when models can be retrained but a low\-friction approach is preferred that does not modify the learning algorithm itself[42](https://arxiv.org/html/2609.22113#bib.bib20);[21](https://arxiv.org/html/2609.22113#bib.bib77)\. In practice, because Reweighing operates entirely before model fitting by adjusting sample weights, it can be incorporated into standard training workflows with minimal engineering overhead and is often a reasonable first\-line option for mitigating disparities arising from imbalanced representation in data\. In\-processing methods such as EGR are best suited for scenarios in which organizations can retrain models and seek explicit control over fairness constraints during training[5](https://arxiv.org/html/2609.22113#bib.bib21)\. EGR is particularly appropriate when institutions want to enforce specific parity criteria \(e\.g\., bounds on differences in TPR or FPR across groups\) as part of the optimization objective and have sufficient technical capacity to integrate constraint\-based training into their modeling pipelines\. This approach offers a principled way to balance predictive performance and fairness during model development, rather than relying on post hoc adjustments\. Post\-processing methods such as the Threshold Optimizer are often appropriate when the underlying prediction model is fixed, for example, when using a vendor\-supplied model or a model that has been operationally validated and cannot be easily retrained, and the primary degree of freedom lies in how predicted risks are translated into decisions\. This is practically meaningful because many clinical and operational workflows already rely on threshold\-based decision rules, such as determining which patients are flagged for proactive outreach, additional follow\-up, or care coordination[29](https://arxiv.org/html/2609.22113#bib.bib78)\. In such cases, threshold adjustment provides a feasible mechanism for reducing disparities in decision outcomes without altering the underlying risk model\. More broadly, ensuring model fairness and trustworthiness requires attention across the entire model development and deployment pipeline, from data curation and feature selection to model training, evaluation, and real\-world use, as biases can emerge at multiple stages\. From a practical perspective, reducing algorithmic bias may help promote more equitable decision support for patients receiving MOUD\. By reducing disparities in model predictions across demographic groups, fairness\-aware models may help ensure that patients with similar clinical risk profiles are identified more consistently regardless of race, ethnicity, age, or sex\. This could support more equitable treatment prioritization, referral decisions, and timely interventions for patients at increased risk of premature treatment discontinuation or poor treatment retention\. For example, clinicians may be better able to identify individuals who could benefit from additional counseling, case management, or closer follow\-up while reducing the likelihood that certain demographic groups are systematically under\- or over\-identified for these supportive services\. Such improvements may contribute to more equitable allocation of healthcare resources and promote more consistent access to evidence\-based interventions across diverse patient populations\. Nevertheless, algorithmic fairness alone cannot eliminate existing healthcare disparities\. Achieving equitable care also requires addressing broader structural, socioeconomic, and healthcare system factors that influence treatment access, engagement, and outcomes\. Therefore, fairness\-aware predictive models should be viewed as decision support tools that complement, rather than replace, clinical judgment and broader efforts to advance health equity\.

Our assessment of model fairness and performance inconsistencies has several important limitations\. First, from a predictive performance perspective, the overall performance achieved by the models was modest compared with that reported in some prior MOUD\-related prediction studies[69](https://arxiv.org/html/2609.22113#bib.bib8);[78](https://arxiv.org/html/2609.22113#bib.bib26)\. This is largely caused by challenges inherent to outpatient treatment settings and to the limitations of routinely collected administrative data\. Unlike inpatient, residential, or mixed treatment settings, outpatient MOUD care is less structured\. In this context, treatment retention is influenced by dynamic and often unobserved behavioral and social factors, and sustained engagement relies heavily on patient self\-management over extended periods[10](https://arxiv.org/html/2609.22113#bib.bib41)\. Although the TEDS\-D dataset is large and nationally representative, it lacks detailed and longitudinal clinical and contextual information, such as provider behavior, social support, and individual treatment adherence, all of which are important for understanding treatment retention and outcomes\. In the TEDS\-D dataset, MOUD is recorded as a single indicator of planned medication\-assisted opioid therapy and does not distinguish among methadone, buprenorphine, and naltrexone\. Our outcome definition may also be constrained by the limited information available in the dataset and may not fully distinguish the circumstances surrounding treatment termination\. In addition, restricting the cohort to outpatient settings increases population homogeneity and removes setting\-driven variation that may otherwise inflate predictive performance in mixed\-setting analyses\. As a result, achieving high overall performance for more complex, behaviorally driven outcomes such as treatment retention is inherently more challenging in outpatient MOUD settings\. Future MOUD research would benefit substantially from the availability of richer datasets that incorporate longitudinal clinical, behavioral, and social determinants of health, which may enable more accurate prediction while also supporting more nuanced fairness assessments\. Second, although our outcome definitions—such as retention over 180 days and premature discontinuation—are consistent with prior research, they may not fully reflect the nuanced and dynamic nature of treatment engagement and outcomes\. Treatment trajectories are often influenced by both personal and structural factors that are not easily captured by binary outcomes\. As such, some complexity of real\-world treatment behavior may be lost in this formulation, limiting the models’ ability to represent patient trajectories accurately\. Third, the retrospective nature of our analysis, which relies on administrative discharge data, may limit the generalizability of our findings to other clinical settings or populations\. Moreover, the dataset includes only facilities that report to TEDS\-D, which may not reflect the full diversity of MOUD treatment settings, such as private or community\-based clinics, limiting generalizability\. Finally, while we evaluated bias mitigation strategies across multiple stages of the ML pipeline, our analysis does not assess how these interventions would perform when deployed in real\-world clinical environments\. Factors such as clinical workflow integration, provider response to model outputs, and the downstream consequences of prediction errors remain unexamined\.

Looking ahead, future research can take several important directions to further advance the development of fair ML models in the context of MOUD treatment outcome prediction\. First, there is a continued need to develop strategies that can more effectively balance algorithmic fairness and overall predictive performance\. While many current approaches seek to improve one at the expense of the other, future work could explore adaptive or context\-aware algorithms that optimize both dimensions simultaneously, minimizing unintended trade\-offs\. More robust and domain\-specific mitigation strategies should be explored to ensure generalizability across diverse clinical settings and populations\. Second, evaluating the real\-world impact of these mitigation strategies is critical\. There is a need for prospective studies, implementation trials, and collaborations with health systems to assess how these strategies influence clinical workflows, decision\-making, and patient outcomes in practice over time\. Additionally, longitudinal studies could examine the durability of mitigation techniques, their acceptability among clinicians, and their integration into routine decision\-making workflows\. In addition, engaging stakeholders—including clinicians, patients, researchers, and implementation teams—through participatory design frameworks can help ensure that algorithmic tools are aligned with clinical priorities, ethically grounded, and contextually appropriate for the communities they are intended to serve\. Finally, identifying and understanding the root causes of observed model performance gaps across patient subgroups remains a crucial area for future work\. For instance, using causal inference techniques could help distinguish between variables that genuinely influence treatment outcomes and those that merely act as proxies for structural patterns in the data\. By uncovering these causal relationships, researchers can design models that focus on clinically meaningful signals while minimizing the influence of spurious or confounding patterns that may distort predictions\.

### 6Conclusion

This study systematically examined the use of ML models to predict MOUD treatment outcomes, including premature treatment discontinuation and treatment retention exceeding 180 days, with a specific focus on outpatient treatment settings\. Using a large, national administrative dataset, we evaluated both predictive performance and algorithmic fairness across patient sociodemographic subgroups, and assessed the effectiveness of commonly used bias mitigation strategies spanning different stages of the ML development pipeline\. Our analysis revealed notable discrepancies in model performance across patient subgroups defined by sensitive attributes such as race and age, as well as more complex cross\-tabulated subgroups\. We found that ML models optimized for overall performance can exhibit systematically different error rates and predictive patterns across sociodemographic groups, raising important concerns about fairness and equity when such models are used to inform care decisions\. Our interpretation of model outputs suggests that such performance variation across subgroups is likely caused by the model learning spurious patterns in the training data\. To address these concerns, we assessed a set of mitigation techniques aimed at improving model fairness across patient subgroups\. While these strategies were able to reduce performance gaps in many cases, they often introduced trade\-offs, such as shifts in prediction error between subgroups, highlighting the need for cautious implementation and continuous monitoring\. Our evaluation of bias mitigation strategies further demonstrates that no single approach is universally optimal\. Pre\-processing, in\-processing, and post\-processing methods each offer distinct advantages and trade\-offs depending on implementation constraints, technical capacity, and intended use\. These results provide practical guidance for researchers and practitioners seeking to integrate fairness\-aware ML into MOUD treatment planning\. Ultimately, our findings emphasize the importance of developing ML models that are not only accurate overall but also perform well across different patient populations\. Achieving this balance requires greater attention to the design, training, and evaluation of models, as well as clear guidelines for how subgroup\-level performance should be monitored and reported\. Future development efforts must integrate these principles to ensure that ML\-based tools support fair and clinically meaningful decision\-making for MOUD\.

### Statements and Declarations

#### Competing Interests

The authors have no competing interests to declare that are relevant to the content of this article and there are no financial interests\.

#### Funding Acknowledgements

This research was supported, in part, by the National Institutes of Health \(NIH\) under Agreement No\. 1OT2OD032581 and by the U\.S\. National Science Foundation under Grant No\. CMMI\-2222670\.

#### Ethics Approval

Not applicable\.

#### Consent to Participate

Not applicable\.

#### Consent for Publication

Not applicable\.

#### Availability of data and materials

### 7Supplementary Results

Figure 9:Cohort and variable selection\.#### 7\.1Data processing

Figure[9](https://arxiv.org/html/2609.22113#S7.F9)illustrates the process used for cohort construction and variable selection in this study\. Following this process, a total of 27 variables were included\. Of these, two served as outcome variables—treatment retention longer than 180 days and premature discontinuation—leaving 25 variables as predictors for model development\.

#### 7\.2Hyperparameter Tuning

For LR, RF, and GBDT, hyperparameters were optimized using grid search with 5\-fold stratified cross\-validation on the training set\. All candidate hyperparameter combinations in the predefined search grid were evaluated using the mean AUROC across the five validation folds\. The configuration achieving the highest mean cross\-validated AUROC was selected, and the corresponding model was refitted on the complete training set before evaluation on the held\-out test set\. Hyperparameter tuning for the MLP model was performed using grid search with Keras Tuner\. Candidate hyperparameter configurations were evaluated on the validation set using AUROC as the optimization criterion\. The hyperparameter configuration achieving the highest validation AUROC was selected for the final model\. Results are reported with averages over 10 independent runs using different random seeds to account for variability\. Best hyperparameter configurations for each model and task are summarized in Table[8](https://arxiv.org/html/2609.22113#S7.T8)\.

#### 7\.3Equalized Odds Results

Equalized odds is a fairness criterion that requires equal TPRs and FPRs across subgroups\. In this study, we evaluated it for all four ML models across subgroups defined by race, ethnicity, and age to provide a comprehensive assessment of fairness\. The results are presented in Table[7](https://arxiv.org/html/2609.22113#S7.T7)for both premature treatment discontinuation and treatment retention longer than 180 days\.

Table 7:Equalized Odds Difference \(mean; standard deviation in parentheses\) for predicting premature discontinuation and retention longer than 180 days\. The Equalized Odds Difference is defined as the maximum of the TPR difference and the FPR difference across subgroups\. Lower values indicate smaller gaps between subgroups, with 0 representing perfect equalized odds\.Boldvalues indicate the best results\.GroupMethodPremature DiscontinuationRetention\>180\>180daysLRRFGBDTMLPLRRFGBDTMLPRaceBaseline0\.2967\(0\.0060\)0\.2526\(0\.0104\)0\.2681\(0\.0095\)0\.2828\(0\.0410\)0\.2081\(0\.0073\)0\.2255\(0\.0048\)0\.2354\(0\.0075\)0\.2119\(0\.0394\)Reweighing0\.0289\(0\.0069\)0\.0452\(0\.0041\)0\.0319\(0\.0073\)0\.0434\(0\.0270\)0\.0147\(0\.0043\)0\.0987\(0\.0055\)0\.0792\(0\.0084\)0\.0831\(0\.0471\)EGR0\.0108\(0\.0073\)0\.0489\(0\.0434\)0\.0449\(0\.0304\)\-\-555Results for MLP with EGR are not reported, as EGR requires models that support stable, per\-sample reweighting during training\. MLP are non\-convex and typically do not support per\-sample reweighting in a stable and standardized manner as required\.0\.0121\(0\.0025\)0\.0082\(0\.0080\)0\.0078\(0\.0053\)\-\-555Results for MLP with EGR are not reported, as EGR requires models that support stable, per\-sample reweighting during training\. MLP are non\-convex and typically do not support per\-sample reweighting in a stable and standardized manner as required\.Threshold0\.0065\(0\.0036\)0\.0295\(0\.0060\)0\.0135\(0\.0053\)0\.2011\(0\.0222\)0\.0063\(0\.0058\)0\.0134\(0\.0051\)0\.0085\(0\.0043\)0\.2257\(0\.0267\)EthnicityBaseline0\.1895\(0\.0078\)0\.1406\(0\.0089\)0\.1709\(0\.0171\)0\.1786\(0\.0230\)0\.1437\(0\.0103\)0\.1133\(0\.0097\)0\.1101\(0\.0157\)0\.1226\(0\.0235\)Reweighing0\.0522\(0\.0073\)0\.0250\(0\.0071\)0\.0333\(0\.0132\)0\.0428\(0\.0190\)0\.0084\(0\.0032\)0\.0379\(0\.0074\)0\.00830\.00520\.0245\(0\.0170\)EGR0\.0153\(0\.0085\)0\.0830\(0\.0059\)0\.0453\(0\.0309\)–0\.0118\(0\.0072\)0\.0114\(0\.0077\)0\.0123\(0\.0076\)–Threshold0\.0099\(0\.0085\)0\.0203\(0\.0082\)0\.0088\(0\.0091\)0\.0975\(0\.0182\)0\.0090\(0\.0073\)0\.0182\(0\.0106\)0\.0112\(0\.0070\)0\.1091\(0\.0236\)AgeBaseline0\.1441\(0\.0100\)0\.1726\(0\.0127\)0\.1445\(0\.0096\)0\.1139\(0\.0315\)0\.5619\(0\.0095\)0\.4278\(0\.0151\)0\.4497\(0\.0152\)0\.4842\(0\.0428\)Reweighing0\.0345\(0\.0105\)0\.0930\(0\.0096\)0\.0457\(0\.0089\)0\.0482\(0\.0285\)0\.0217\(0\.0043\)0\.0497\(0\.0057\)0\.0447\(0\.0103\)0\.0607\(0\.0297\)EGR0\.0183\(0\.0078\)0\.1035\(0\.0126\)0\.0459\(0\.0351\)–0\.0212\(0\.0093\)0\.0294\(0\.0033\)0\.0273\(0\.0059\)–Threshold0\.0110\(0\.0037\)0\.0137\(0\.0064\)0\.0116\(0\.0053\)0\.0579\(0\.0123\)0\.0109\(0\.0050\)0\.0401\(0\.0061\)0\.0170\(0\.0062\)0\.4529\(0\.0344\)
Table 8:Best hyperparameter configurations\.LRRFGBDTMLPPremature Discontinuationsolversagan\_estimators500l2\_regularization0\.5hidden layer32 \(ReLU\)penaltyl1max\_depth20learning\_rate0\.01output layer1 \(Sigmoid\)C100min\_samples\_split20max\_iter500epoch \(batch size\)10 \(64\)max\_iter50max\_leaf\_nodes200optimizerAdamclass\_weightbalancedlearning rate0\.001Retention\>\>180 dayssolversagan\_estimators500l2\_regularization2\.0hidden layer32 \(ReLU\)penaltyl2max\_depth20learning\_rate0\.01output layer1 \(Sigmoid\)C10min\_samples\_split20max\_iter500epoch \(batch size\)10 \(64\)max\_iter50max\_leaf\_nodes200optimizerAdamclass\_weightNonelearning rate0\.001

#### 7\.4Accuracy and Fairness Trade\-off

Table[9](https://arxiv.org/html/2609.22113#S7.T9)reports overall model accuracy for both tasks\. When compared with the fairness metrics presented in Tables[5](https://arxiv.org/html/2609.22113#S4.T5)and[6](https://arxiv.org/html/2609.22113#S4.T6), the results reveal a clear trade\-off between accuracy and fairness\.

Table 9:Model accuracy \(mean; standard deviation in parentheses\) for predicting premature treatment discontinuation and treatment retention\.Boldvalues indicate the best results\.GroupMethodPremature DiscontinuationRetention\>\>180 daysLRRFGBDTMLPLRRFGBDTMLPRaceBaseline0\.6104\(0\.0011\)0\.6230\(0\.0014\)0\.6263\(0\.0014\)0\.6065\(0\.0061\)0\.6434\(0\.0010\)0\.6551\(0\.0007\)0\.6587\(0\.0009\)0\.6469\(0\.0014\)Reweighing0\.6058\(0\.0011\)0\.6218\(0\.0013\)0\.6251\(0\.0014\)0\.6193\(0\.0063\)0\.6416\(0\.0007\)0\.6544\(0\.0007\)0\.6576\(0\.0007\)0\.6521\(0\.0018\)EGR0\.6099\(0\.0022\)0\.6316\(0\.0021\)0\.6340\(0\.0010\)\-\-555Results for MLP with EGR are not reported, as EGR requires models that support stable, per\-sample reweighting during training\. MLP are non\-convex and typically do not support per\-sample reweighting in a stable and standardized manner as required\.0\.6416\(0\.0008\)0\.6538\(0\.0009\)0\.6568\(0\.0007\)\-\-555Results for MLP with EGR are not reported, as EGR requires models that support stable, per\-sample reweighting during training\. MLP are non\-convex and typically do not support per\-sample reweighting in a stable and standardized manner as required\.Threshold0\.6031\(0\.0016\)0\.6243\(0\.0014\)0\.6241\(0\.0026\)0\.6155\(0\.0036\)0\.6086\(0\.0055\)0\.6175\(0\.0029\)0\.6264\(0\.0027\)0\.6188\(0\.0055\)EthnicityBaseline0\.6104\(0\.0011\)0\.6230\(0\.0014\)0\.6263\(0\.0014\)0\.6065\(0\.0061\)0\.6434\(0\.0010\)0\.6551\(0\.0007\)0\.6587\(0\.0009\)0\.6469\(0\.0014\)Reweighing0\.6095\(0\.0013\)0\.6225\(0\.0013\)0\.6255\(0\.0012\)0\.6209\(0\.0050\)0\.6429\(0\.0010\)0\.6549\(0\.0010\)0\.6584\(0\.0009\)0\.6527\(0\.0012\)EGR0\.6104\(0\.0012\)0\.6332\(0\.0014\)0\.6347\(0\.0012\)–0\.6434\(0\.0010\)0\.6550\(0\.0008\)0\.6586\(0\.0008\)–Threshold0\.6062\(0\.0015\)0\.6259\(0\.0014\)0\.6276\(0\.0013\)0\.6230\(0\.0017\)0\.5957\(0\.0049\)0\.6185\(0\.0024\)0\.6283\(0\.0040\)0\.6148\(0\.0103\)AgeBaseline0\.6104\(0\.0011\)0\.6230\(0\.0014\)0\.6263\(0\.0014\)0\.6065\(0\.0061\)0\.6434\(0\.0010\)0\.6551\(0\.0007\)0\.6587\(0\.0009\)0\.6469\(0\.0014\)Reweighing0\.6097\(0\.0010\)0\.6233\(0\.0011\)0\.6264\(0\.0012\)0\.6222\(0\.0047\)0\.6378\(0\.0007\)0\.6515\(0\.0006\)0\.6539\(0\.0008\)0\.6481\(0\.0015\)EGR0\.6105\(0\.0011\)0\.6333\(0\.0014\)0\.6348\(0\.0007\)–0\.6241\(0\.0102\)0\.6520\(0\.0009\)0\.6542\(0\.0008\)–Threshold0\.5954\(0\.0029\)0\.6216\(0\.0025\)0\.6257\(0\.0027\)0\.6212\(0\.0040\)0\.5705\(0\.0064\)0\.6169\(0\.0039\)0\.6144\(0\.0074\)0\.6017\(0\.0078\)

#### 7\.5AUROC Results Across Subgroups

AUROC results across subgroups defined by race, ethnicity, and age are presented in Table[10](https://arxiv.org/html/2609.22113#S7.T10)\. Overall, AUROC values were generally comparable across demographic groups, although modest differences were observed\. However, disparities in TPRs and FPRs remained at the selected classification threshold\. Accordingly, the Threshold Optimizer was employed to mitigate these threshold\-dependent disparities\.

LRRFGBDTMLPPremature DiscontinuationRaceWhite0\.6566 \(0\.0011\)0\.6865 \(0\.0013\)0\.6895 \(0\.0013\)0\.6651 \(0\.0024\)Black0\.6291 \(0\.0033\)0\.6706 \(0\.0022\)0\.6729 \(0\.0024\)0\.6412 \(0\.0043\)EthnicityNon\-Hispanic0\.6535 \(0\.0011\)0\.6840 \(0\.0013\)0\.6872 \(0\.0013\)0\.6626 \(0\.0027\)Hispanic0\.6793 \(0\.0068\)0\.7072 \(0\.0049\)0\.7062 \(0\.0055\)0\.6828 \(0\.0056\)Age18\-240\.6455 \(0\.0038\)0\.6780 \(0\.0047\)0\.6805 \(0\.0049\)0\.6543 \(0\.0049\)25\-540\.6540 \(0\.0013\)0\.6843 \(0\.0011\)0\.6869 \(0\.0012\)0\.6629 \(0\.0025\)\>\>550\.6771 \(0\.0037\)0\.7062 \(0\.0042\)0\.7106 \(0\.0034\)0\.6829 \(0\.0041\)Retention\>\>180 daysRaceWhite0\.6531 \(0\.0010\)0\.6759 \(0\.0012\)0\.6796 \(0\.0011\)0\.6626 \(0\.0014\)Black0\.6600 \(0\.0039\)0\.6900 \(0\.0040\)0\.6943 \(0\.0044\)0\.6719 \(0\.0037\)EthnicityNon\-Hispanic0\.6562 \(0\.0013\)0\.6793 \(0\.0015\)0\.6833 \(0\.0013\)0\.6659 \(0\.0015\)Hispanic0\.6467 \(0\.0063\)0\.6788 \(0\.0054\)0\.6814 \(0\.0060\)0\.6576 \(0\.0065\)Age18\-240\.6465 \(0\.0018\)0\.6682 \(0\.0022\)0\.6743 \(0\.0022\)0\.6542 \(0\.0034\)25\-540\.6519 \(0\.0011\)0\.6768 \(0\.0013\)0\.6804 \(0\.0012\)0\.6621 \(0\.0012\)\>\>550\.6241 \(0\.0027\)0\.6578 \(0\.0033\)0\.6610 \(0\.0034\)0\.6416 \(0\.0041\)Table 10:AUROC results for predicting premature discontinuation and treatment retention\>\>180 days\.

#### 7\.6Ablation Results

All four ML models were also trained without including sensitive attributes as input features\. As shown in Table[11](https://arxiv.org/html/2609.22113#S7.T11)for predicting premature discontinuation and Table[12](https://arxiv.org/html/2609.22113#S7.T12)for treatment retention over 180 days, overall prediction performance remained largely unchanged\. However, performance gaps across patient subgroups persisted, indicating that excluding these attributes does not eliminate subgroup\-level differences and that models may have relied on correlated features as proxies for prediction\.

#### 7\.7Additional Results

Figure[10](https://arxiv.org/html/2609.22113#S7.F10)presents the TPR for each bias mitigation strategy compared to the baseline when predicting premature discontinuation\. All three methods substantially reduced TPR gaps across patient subgroups\. A similar pattern is observed in Figure[11](https://arxiv.org/html/2609.22113#S7.F11)for predictions of treatment retention beyond 180 days\. Tables[13](https://arxiv.org/html/2609.22113#S7.T13)and[14](https://arxiv.org/html/2609.22113#S7.T14)present detailed FPR and TPR results for each ML model across all patient subgroups for predicting premature discontinuation\. Similarly, Tables[15](https://arxiv.org/html/2609.22113#S7.T15)and[16](https://arxiv.org/html/2609.22113#S7.T16)report all FPR and TPR results for predicting treatment retention longer than 180 days\.

Table 11:Comparison of model performance with and without sensitive attribute as a feature for predicting premature treatment discontinuation\.LRRFGBDTMLPWith RaceOverall \(AUROC\)0\.6562 \(0\.0011\)0\.6863 \(0\.0013\)0\.6892 \(0\.0013\)0\.6647 \(0\.0025\)White \(FPR\)0\.3545 \(0\.0031\)0\.2842 \(0\.0047\)0\.3011 \(0\.0064\)0\.2878 \(0\.0666\)Black \(FPR\)0\.6512 \(0\.0067\)0\.5368 \(0\.0108\)0\.5692 \(0\.0085\)0\.5705 \(0\.0702\)Without RaceOverall \(AUROC\)0\.6548 \(0\.0011\)0\.6847 \(0\.0013\)0\.6881 \(0\.0013\)0\.6626 \(0\.0024\)White \(FPR\)0\.3882 \(0\.0035\)0\.3091 \(0\.0040\)0\.3191 \(0\.0052\)0\.3142 \(0\.0698\)Black \(FPR\)0\.5583 \(0\.0088\)0\.4522 \(0\.0046\)0\.4746 \(0\.0065\)0\.4363 \(0\.0857\)With EthnicityOverall \(AUROC\)0\.6562 \(0\.0011\)0\.6863 \(0\.0013\)0\.6892 \(0\.0013\)0\.6647 \(0\.0025\)Non\-Hispanic \(FPR\)0\.3839 \(0\.0033\)0\.3116 \(0\.0045\)0\.3277 \(0\.0067\)0\.3160 \(0\.0653\)Hispanic \(FPR\)0\.5565 \(0\.0085\)0\.4225 \(0\.0143\)0\.4824 \(0\.0168\)0\.4775 \(0\.0731\)Without EthnicityOverall \(AUROC\)0\.6560 \(0\.0012\)0\.6858 \(0\.0014\)0\.6890 \(0\.0014\)0\.6647 \(0\.0019\)Non\-Hispanic \(FPR\)0\.3947 \(0\.0030\)0\.3201 \(0\.0039\)0\.3367 \(0\.0058\)0\.3247 \(0\.0667\)Hispanic \(FPR\)0\.4714 \(0\.0077\)0\.3555 \(0\.0135\)0\.3824 \(0\.0122\)0\.3659 \(0\.0700\)With AgeOverall \(AUROC\)0\.6562 \(0\.0011\)0\.6863 \(0\.0013\)0\.6892 \(0\.0013\)0\.6647 \(0\.0025\)18\-24 \(FPR\)0\.4193 \(0\.0068\)0\.2850 \(0\.0086\)0\.3115 \(0\.0090\)0\.3394 \(0\.0602\)25\-54 \(FPR\)0\.3808 \(0\.0041\)0\.3116 \(0\.0049\)0\.3330 \(0\.0059\)0\.3175 \(0\.0684\)\>\>55 \(FPR\)0\.4926 \(0\.0082\)0\.4157 \(0\.0080\)0\.4076 \(0\.0090\)0\.3906 \(0\.0563\)Without AgeOverall \(AUROC\)0\.6559 \(0\.0012\)0\.6862 \(0\.0012\)0\.6891 \(0\.0014\)0\.6651 \(0\.0017\)18\-24 \(FPR\)0\.3355 \(0\.0055\)0\.2825 \(0\.0083\)0\.2958 \(0\.0073\)0\.2745 \(0\.0599\)25\-54 \(FPR\)0\.3820 \(0\.0036\)0\.3105 \(0\.0044\)0\.3249 \(0\.0054\)0\.3071 \(0\.0618\)\>\>55 \(FPR\)0\.5675 \(0\.0050\)0\.4643 \(0\.0059\)0\.4779 \(0\.0079\)0\.4866 \(0\.0664\)
Table 12:Comparison of model performance with and without sensitive attribute as a feature for predicting treatment retention\.LRRFGBDTMLPWith RaceOverall \(AUROC\)0\.6559 \(0\.0011\)0\.6797 \(0\.0011\)0\.6836 \(0\.0010\)0\.6657 \(0\.0014\)White \(FPR\)0\.1107 \(0\.0021\)0\.1114 \(0\.0016\)0\.1155 \(0\.0020\)0\.1288 \(0\.0397\)Black \(FPR\)0\.2459 \(0\.0050\)0\.2495 \(0\.0053\)0\.2656 \(0\.0068\)0\.2701 \(0\.0465\)Without RaceOverall \(AUROC\)0\.6556 \(0\.0011\)0\.6784 \(0\.0012\)0\.6823 \(0\.0009\)0\.6650 \(0\.0015\)White \(FPR\)0\.1179 \(0\.0021\)0\.1182 \(0\.0018\)0\.1216 \(0\.0017\)0\.1391 \(0\.0437\)Black \(FPR\)0\.1989 \(0\.0032\)0\.2236 \(0\.0051\)0\.2336 \(0\.0060\)0\.2259 \(0\.0522\)With EthnicityOverall \(AUROC\)0\.6559 \(0\.0011\)0\.6797 \(0\.0011\)0\.6836 \(0\.0010\)0\.6657 \(0\.0014\)Non\-Hispanic \(FPR\)0\.1233 \(0\.0021\)0\.1264 \(0\.0016\)0\.1319 \(0\.0018\)0\.1471 \(0\.0401\)Hispanic \(FPR\)0\.2264 \(0\.0092\)0\.2023 \(0\.0102\)0\.2124 \(0\.0139\)0\.2422 \(0\.0505\)Without EthnicityOverall \(AUROC\)0\.6559 \(0\.0011\)0\.6795 \(0\.0012\)0\.6835 \(0\.0010\)0\.6651 \(0\.0012\)Non\-Hispanic \(FPR\)0\.1269 \(0\.0021\)0\.1281 \(0\.0018\)0\.1332 \(0\.0019\)0\.1475 \(0\.0372\)Hispanic \(FPR\)0\.1765 \(0\.0041\)0\.1944 \(0\.0067\)0\.1947 \(0\.0048\)0\.2051 \(0\.0499\)With AgeOverall \(AUROC\)0\.6559 \(0\.0011\)0\.6797 \(0\.0011\)0\.6836 \(0\.0010\)0\.6657 \(0\.0014\)18\-24 \(FPR\)0\.1042 \(0\.0020\)0\.1084 \(0\.0022\)0\.1129 \(0\.0026\)0\.1320 \(0\.0403\)25\-54 \(FPR\)0\.1042 \(0\.0020\)0\.1084 \(0\.0022\)0\.1129 \(0\.0026\)0\.1320 \(0\.0403\)\>\>55 \(FPR\)0\.4901 \(0\.0096\)0\.4131 \(0\.0110\)0\.4414 \(0\.0073\)0\.4584 \(0\.0693\)Without AgeOverall \(AUROC\)0\.6481 \(0\.0010\)0\.6743 \(0\.0010\)0\.6773 \(0\.0009\)0\.6580 \(0\.0019\)18\-24 \(FPR\)0\.1040 \(0\.0032\)0\.1130 \(0\.0028\)0\.1079 \(0\.0044\)0\.1254 \(0\.0471\)25\-54 \(FPR\)0\.1120 \(0\.0021\)0\.1244 \(0\.0020\)0\.1193 \(0\.0020\)0\.1356 \(0\.0418\)\>\>¿55 \(FPR\)0\.2507 \(0\.0066\)0\.2747 \(0\.0042\)0\.2817 \(0\.0050\)0\.2769 \(0\.0518\)
Figure 10:TPRs across different subgroups are shown for three bias mitigation methods—Reweighing \(pre\-processing\), EGR \(in\-processing\), and Threshold \(post\-processing\)—in predicting premature treatment discontinuation, compared to a baseline without fairness considerations\. Results are stratified by \(a\) race, \(b\) ethnicity, and \(c\) age\. Bars represent the mean TPR values over 10 independent runs, with error bars indicating the standard deviation across these runs\.Figure 11:TPRs across different subgroups are shown for three bias mitigation methods—Reweighing \(pre\-processing\), EGR \(in\-processing\), and Threshold \(post\-processing\)—in predicting treatment retention\>\>180 days, compared to a baseline without fairness considerations\. Results are stratified by \(a\) race, \(b\) ethnicity, and \(c\) age\. Bars represent the mean TPR values over 10 independent runs, with error bars indicating the standard deviation across these runs\.LRRFGBDTMLPBaselineRaceWhite0\.3545 \(0\.0031\)0\.2842 \(0\.0047\)0\.3011 \(0\.0064\)0\.2878 \(0\.0666\)Black0\.6512 \(0\.0067\)0\.5368 \(0\.0108\)0\.5692 \(0\.0085\)0\.5705 \(0\.0702\)EthnicityNon\-Hispanic0\.3839 \(0\.0033\)0\.3116 \(0\.0045\)0\.3277 \(0\.0067\)0\.3160 \(0\.0653\)Hispanic0\.5565 \(0\.0085\)0\.4225 \(0\.0143\)0\.4824 \(0\.0168\)0\.4775 \(0\.0731\)Age18\-240\.4194 \(0\.0068\)0\.2850 \(0\.0086\)0\.3115 \(0\.0090\)0\.3394 \(0\.0602\)25\-540\.3808 \(0\.0041\)0\.3116 \(0\.0049\)0\.3330 \(0\.0059\)0\.3175 \(0\.0684\)\>\>550\.4926 \(0\.0082\)0\.4157 \(0\.0080\)0\.4076 \(0\.0090\)0\.3906 \(0\.0563\)ReweighingRaceWhite0\.4071 \(0\.0028\)0\.3176 \(0\.0040\)0\.3373 \(0\.0047\)0\.3559 \(0\.0592\)Black0\.4207 \(0\.0076\)0\.3628 \(0\.0058\)0\.3692 \(0\.0095\)0\.3896 \(0\.0710\)EthnicityNon\-Hispanic0\.4042 \(0\.0032\)0\.3219 \(0\.0040\)0\.3415 \(0\.0059\)0\.3599 \(0\.0565\)Hispanic0\.3521 \(0\.0078\)0\.3104 \(0\.0101\)0\.3081 \(0\.0146\)0\.3171 \(0\.0663\)Age18\-240\.4133 \(0\.0066\)0\.2919 \(0\.0083\)0\.3212 \(0\.0085\)0\.3644 \(0\.0581\)25\-540\.4048 \(0\.0034\)0\.3261 \(0\.0050\)0\.3492 \(0\.0066\)0\.3701 \(0\.0566\)\>\>550\.3790 \(0\.0081\)0\.3352 \(0\.0079\)0\.3115 \(0\.0082\)0\.3241 \(0\.0497\)EGRRaceWhite0\.4953 \(0\.0349\)0\.4940 \(0\.0139\)0\.4990 \(0\.0082\)–Black0\.5061 \(0\.0349\)0\.5160 \(0\.0684\)0\.5424 \(0\.0304\)–EthnicityNon\-Hispanic0\.4304 \(0\.0147\)0\.4956 \(0\.0014\)0\.4937 \(0\.0082\)–Hispanic0\.4151 \(0\.0211\)0\.5697 \(0\.0077\)0\.4950 \(0\.0616\)–Age18\-240\.4171 \(0\.0121\)0\.4846 \(0\.0133\)0\.4928 \(0\.0151\)–25\-540\.4114 \(0\.0092\)0\.4903 \(0\.0058\)0\.4993 \(0\.0051\)–\>\>550\.4042 \(0\.0119\)0\.5806 \(0\.0133\)0\.5078 \(0\.0347\)–ThresholdRaceWhite0\.5719 \(0\.0137\)0\.3652 \(0\.0146\)0\.3706 \(0\.0244\)0\.4732 \(0\.0524\)Black0\.5745 \(0\.0131\)0\.3958 \(0\.0194\)0\.3842 \(0\.0215\)0\.6722 \(0\.0441\)EthnicityNon\-Hispanic0\.3722 \(0\.0142\)0\.3775 \(0\.0098\)0\.3756 \(0\.0109\)0\.4917 \(0\.0491\)Hispanic0\.3674 \(0\.0183\)0\.3981 \(0\.0115\)0\.3807 \(0\.0156\)0\.5856 \(0\.0360\)Age18\-240\.3583 \(0\.0381\)0\.3518 \(0\.0206\)0\.3956 \(0\.0327\)0\.5166 \(0\.0503\)25\-540\.3585 \(0\.0369\)0\.3536 \(0\.0181\)0\.3936 \(0\.0354\)0\.4959 \(0\.0515\)\>\>550\.3633 \(0\.0388\)0\.3616 \(0\.0243\)0\.3962 \(0\.0329\)0\.5275 \(0\.0561\)Table 13:All FPR results for predicting premature treatment discontinuation\.LRRFGBDTMLPBaselineRaceWhite0\.5718 \(0\.0025\)0\.5352 \(0\.0047\)0\.5559 \(0\.0057\)0\.5069 \(0\.0700\)Black0\.8209 \(0\.0027\)0\.7564 \(0\.0078\)0\.7863 \(0\.0063\)0\.7532 \(0\.0588\)EthnicityNon\-Hispanic0\.5998 \(0\.0025\)0\.5623 \(0\.0048\)0\.5821 \(0\.0064\)0\.5352 \(0\.0673\)Hispanic0\.7893 \(0\.0064\)0\.7029 \(0\.0080\)0\.7530 \(0\.0118\)0\.7139 \(0\.0594\)Age18\-240\.6189 \(0\.0079\)0\.5249 \(0\.0069\)0\.5537 \(0\.0060\)0\.5483 \(0\.0605\)25\-540\.5970 \(0\.0033\)0\.5623 \(0\.0053\)0\.5869 \(0\.0058\)0\.5365 \(0\.0707\)\>\>550\.7411 \(0\.0080\)0\.6976 \(0\.0090\)0\.6982 \(0\.0079\)0\.6454 \(0\.0499\)ReweighingRaceWhite0\.6233 \(0\.0024\)0\.5712 \(0\.0037\)0\.5951 \(0\.0035\)0\.6013 \(0\.0601\)Black0\.5945 \(0\.0076\)0\.5972 \(0\.0085\)0\.6057 \(0\.0115\)0\.6068 \(0\.0744\)EthnicityNon\-Hispanic0\.6195 \(0\.0024\)0\.5727 \(0\.0044\)0\.5963 \(0\.0057\)0\.6034 \(0\.0566\)Hispanic0\.6034 \(0\.0068\)0\.5977 \(0\.0081\)0\.5887 \(0\.0123\)0\.5876 \(0\.0663\)Age18\-240\.6125 \(0\.0039\)0\.5312 \(0\.0064\)0\.5649 \(0\.0073\)0\.6012 \(0\.0597\)25\-540\.6202 \(0\.0023\)0\.5775 \(0\.0050\)0\.6036 \(0\.0059\)0\.6133 \(0\.0561\)\>\>550\.6312 \(0\.0081\)0\.6242 \(0\.0088\)0\.6088 \(0\.0078\)0\.6072 \(0\.0498\)EGRRaceWhite0\.6673 \(0\.0292\)0\.7413 \(0\.0133\)0\.7491 \(0\.0068\)–Black0\.6572 \(0\.0303\)0\.7333 \(0\.0620\)0\.7626 \(0\.0239\)–EthnicityNon\-Hispanic0\.6410 \(0\.0117\)0\.7396 \(0\.0022\)0\.7420 \(0\.0059\)–Hispanic0\.6477 \(0\.0092\)0\.8227 \(0\.0058\)0\.7613 \(0\.0468\)–Age18\-240\.6161 \(0\.0081\)0\.7241 \(0\.0118\)0\.7336 \(0\.0166\)–25\-540\.6261 \(0\.0088\)0\.7361 \(0\.0061\)0\.7465 \(0\.0039\)–\>\>550\.6322 \(0\.0118\)0\.8256 \(0\.0059\)0\.7753 \(0\.0271\)–ThresholdRaceWhite0\.7510 \(0\.0124\)0\.6209 \(0\.0137\)0\.6224 \(0\.0217\)0\.7141 \(0\.0465\)Black0\.7501 \(0\.0153\)0\.6095 \(0\.0173\)0\.6165 \(0\.0255\)0\.8573 \(0\.0296\)EthnicityNon\-Hispanic0\.5879 \(0\.0139\)0\.6299 \(0\.0100\)0\.6312 \(0\.0103\)0\.7291 \(0\.0420\)Hispanic0\.5847 \(0\.0122\)0\.6300 \(0\.0063\)0\.6287 \(0\.0087\)0\.8230 \(0\.0227\)Age18\-240\.5554 \(0\.0351\)0\.5960 \(0\.0178\)0\.6406 \(0\.0360\)0\.7462 \(0\.0415\)25\-540\.5569 \(0\.0372\)0\.6014 \(0\.0187\)0\.6430 \(0\.0341\)0\.7334 \(0\.0439\)\>\>550\.5544 \(0\.0397\)0\.6011 \(0\.0205\)0\.6422 \(0\.0329\)0\.7824 \(0\.0389\)Table 14:All TPR results for predicting premature treatment discontinuation\.LRRFGBDTMLPBaselineRaceWhite0\.1107 \(0\.0021\)0\.1114 \(0\.0016\)0\.1155 \(0\.0020\)0\.1288 \(0\.0397\)Black0\.2459 \(0\.0050\)0\.2495 \(0\.0053\)0\.2656 \(0\.0068\)0\.2701 \(0\.0465\)EthnicityNon\-Hispanic0\.1233 \(0\.0021\)0\.1264 \(0\.0016\)0\.1319 \(0\.0018\)0\.1471 \(0\.0401\)Hispanic0\.2264 \(0\.0092\)0\.2023 \(0\.0102\)0\.2124 \(0\.0139\)0\.2422 \(0\.0505\)Age18\-240\.1042 \(0\.0020\)0\.1084 \(0\.0022\)0\.1129 \(0\.0026\)0\.1320 \(0\.0403\)25\-540\.1042 \(0\.0020\)0\.1084 \(0\.0022\)0\.1129 \(0\.0026\)0\.1320 \(0\.0403\)\>\>550\.4901 \(0\.0096\)0\.4131 \(0\.0110\)0\.4414 \(0\.0073\)0\.4584 \(0\.0693\)ReweighingRaceWhite0\.1260 \(0\.0023\)0\.1227 \(0\.0013\)0\.1297 \(0\.0016\)0\.1543 \(0\.0322\)Black0\.1403 \(0\.0029\)0\.1798 \(0\.0058\)0\.1795 \(0\.0075\)0\.2042 \(0\.0401\)EthnicityNon\-Hispanic0\.1296 \(0\.0022\)0\.1286 \(0\.0017\)0\.1367 \(0\.0016\)0\.1612 \(0\.0310\)Hispanic0\.1342 \(0\.0035\)0\.1547 \(0\.0058\)0\.1420 \(0\.0054\)0\.1561 \(0\.0375\)Age18\-240\.1093 \(0\.0041\)0\.1041 \(0\.0043\)0\.1169 \(0\.0054\)0\.1430 \(0\.0294\)25\-540\.1242 \(0\.0023\)0\.1282 \(0\.0021\)0\.1325 \(0\.0020\)0\.1607 \(0\.0265\)\>\>550\.1032 \(0\.0056\)0\.1394 \(0\.0054\)0\.1299 \(0\.0061\)0\.1613 \(0\.0479\)EGRRaceWhite0\.1265 \(0\.0022\)0\.1329 \(0\.0040\)0\.1358 \(0\.0013\)–Black0\.1386 \(0\.0038\)0\.1359 \(0\.0050\)0\.1435 \(0\.0047\)–EthnicityNon\-Hispanic0\.1668 \(0\.0059\)0\.1343 \(0\.0053\)0\.1367 \(0\.0018\)–Hispanic0\.1731 \(0\.0087\)0\.1451 \(0\.0086\)0\.1487 \(0\.0091\)–Age18\-240\.1260 \(0\.0433\)0\.1374 \(0\.0080\)0\.1289 \(0\.0054\)–25\-540\.1083 \(0\.0376\)0\.1358 \(0\.0081\)0\.1291 \(0\.0022\)–\>\>550\.1262 \(0\.0420\)0\.1620 \(0\.0102\)0\.1539 \(0\.0058\)–ThresholdRaceWhite0\.3893 \(0\.0215\)0\.4040 \(0\.0112\)0\.3714 \(0\.0104\)0\.1377 \(0\.0261\)Black0\.3925 \(0\.0239\)0\.4169 \(0\.0110\)0\.3775 \(0\.0104\)0\.2799 \(0\.0416\)EthnicityNon\-Hispanic0\.4394 \(0\.0152\)0\.4071 \(0\.0095\)0\.3749 \(0\.0160\)0\.1517 \(0\.0275\)Hispanic0\.4349 \(0\.0118\)0\.4285 \(0\.0125\)0\.3743 \(0\.0170\)0\.2339 \(0\.0417\)Age18\-240\.5017 \(0\.0299\)0\.4006 \(0\.0201\)0\.4111 \(0\.0385\)0\.0804 \(0\.0196\)25\-540\.4982 \(0\.0269\)0\.3957 \(0\.0177\)0\.4096 \(0\.0330\)0\.1401 \(0\.0274\)\>\>550\.5019 \(0\.0324\)0\.4349 \(0\.0184\)0\.4251 \(0\.0318\)0\.4469 \(0\.0575\)Table 15:All FPR results for predicting treatment retention\>\>180 days\.LRRFGBDTMLPBaselineRaceWhite0\.2480 \(0\.0022\)0\.2771 \(0\.0028\)0\.2940 \(0\.0039\)0\.2856 \(0\.0646\)Black0\.4561 \(0\.0075\)0\.5026 \(0\.0041\)0\.5294 \(0\.0071\)0\.4991 \(0\.0623\)EthnicityNon\-Hispanic0\.2727 \(0\.0023\)0\.3073 \(0\.0023\)0\.3262 \(0\.0032\)0\.3204 \(0\.0626\)Hispanic0\.4165 \(0\.0108\)0\.4207 \(0\.0109\)0\.4363 \(0\.0168\)0\.4430 \(0\.0630\)Age18\-240\.1110 \(0\.0057\)0\.2111 \(0\.0084\)0\.2226 \(0\.0093\)0\.1779 \(0\.0561\)25\-540\.2356 \(0\.0294\)0\.2717 \(0\.0033\)0\.2885 \(0\.0047\)0\.2891 \(0\.0645\)\>\>550\.6729 \(0\.0072\)0\.6389 \(0\.0109\)0\.6723 \(0\.0092\)0\.6622 \(0\.0701\)ReweighingRaceWhite0\.2734 \(0\.0020\)0\.2964 \(0\.0022\)0\.3176 \(0\.0027\)0\.3417 \(0\.0472\)Black0\.2881 \(0\.0052\)0\.3951 \(0\.0058\)0\.3968 \(0\.0070\)0\.4248 \(0\.0621\)EthnicityNon\-Hispanic0\.2829 \(0\.0022\)0\.3111 \(0\.0025\)0\.3338 \(0\.0030\)0\.3585 \(0\.0478\)Hispanic0\.2745 \(0\.0031\)0\.3491 \(0\.0091\)0\.3307 \(0\.0102\)0\.3371 \(0\.0584\)Age18\-240\.2381 \(0\.0064\)0\.2558 \(0\.0063\)0\.2859 \(0\.0074\)0\.3169 \(0\.0433\)25\-540\.2691 \(0\.0021\)0\.3055 \(0\.0029\)0\.3208 \(0\.0029\)0\.3523 \(0\.0404\)\>\>550\.1837 \(0\.0071\)0\.2927 \(0\.0050\)0\.2774 \(0\.0099\)0\.3130 \(0\.0767\)EGRRaceWhite0\.2741 \(0\.0022\)0\.3129 \(0\.0058\)0\.3272 \(0\.0025\)–Black0\.2848 \(0\.0074\)0\.3189 \(0\.0081\)0\.3347 \(0\.0045\)–EthnicityNon\-Hispanic0\.3290 \(0\.0094\)0\.3205 \(0\.0081\)0\.3338 \(0\.0032\)–Hispanic0\.3185 \(0\.0140\)0\.3208 \(0\.0116\)0\.3385 \(0\.0101\)–Age18\-240\.2002 \(0\.0744\)0\.2982 \(0\.0128\)0\.3038 \(0\.0086\)–25\-540\.2087 \(0\.0744\)0\.3161 \(0\.0113\)0\.3150 \(0\.0030\)–\>\>550\.1891 \(0\.0724\)0\.3233 \(0\.0127\)0\.3122 \(0\.0065\)–ThresholdRaceWhite0\.6052 \(0\.0210\)0\.6530 \(0\.0111\)0\.6232 \(0\.0100\)0\.3152 \(0\.0420\)Black0\.6085 \(0\.0222\)0\.6616 \(0\.0098\)0\.6278 \(0\.0120\)0\.5340 \(0\.0552\)EthnicityNon\-Hispanic0\.6509 \(0\.0115\)0\.6622 \(0\.0097\)0\.6337 \(0\.0147\)0\.3442 \(0\.0424\)Hispanic0\.6513 \(0\.0192\)0\.6494 \(0\.0091\)0\.6312 \(0\.0216\)0\.4508 \(0\.0564\)Age18\-240\.6807 \(0\.0263\)0\.6395 \(0\.0220\)0\.6513 \(0\.0351\)0\.2179 \(0\.0356\)25\-540\.6814 \(0\.0269\)0\.6437 \(0\.0181\)0\.6567 \(0\.0323\)0\.3191 \(0\.0422\)\>\>550\.6820 \(0\.0303\)0\.6403 \(0\.0200\)0\.6524 \(0\.0360\)0\.6629 \(0\.0631\)Table 16:All TPR results for predicting treatment retention\>\>180 days\.
#### 7\.8Results for 2020–2023 TEDS\-D Data

To examine whether the findings remained consistent in more recent years and to evaluate the potential impact of changes in MOUD treatment following the COVID\-19 pandemic, we conducted additional experiments using the 2020–2023 TEDS\-D data\. The same cohort selection criteria, data preprocessing procedures, model development pipeline, and evaluation metrics were applied to the 2020–2023 cohort\. We evaluated the predictive performance and fairness of the four models for predicting premature treatment discontinuation and treatment retention exceeding 180 days\. The results are summarized in Table[17](https://arxiv.org/html/2609.22113#S7.T17), Figures[12](https://arxiv.org/html/2609.22113#S7.F12), and Figure[13](https://arxiv.org/html/2609.22113#S7.F13)\. Overall, the models achieved slightly better predictive performance than those obtained using the 2015–2019 cohort\. In addition, the fairness evaluation based on the Equalized Odds Difference \(EOD\) exhibited patterns similar to those observed in the primary analysis, with slightly smaller disparities across demographic subgroups\. Overall, the results demonstrate that the proposed modeling framework produces consistent predictive performance and fairness characteristics when applied to more recent TEDS\-D data\.

Table 17:Overall predictive performance, measured by area under the receiver operating characteristic curve \(AUROC\), as well as Equalized Odds Difference \(EOD\) across subgroups for ML models predicting premature treatment discontinuation and treatment retention exceeding 180 days using 2020\-2023 TEDS\-D data\.ModelPremature DiscontinuationRetention\>\>180 daysAUROCLR0\.7153 \(0\.0009\)0\.6427 \(0\.0015\)RF0\.7367 \(0\.0011\)0\.6770 \(0\.0017\)GBDT0\.7378 \(0\.0012\)0\.6781 \(0\.0018\)MLP0\.7173 \(0\.0022\)0\.6555 \(0\.0039\)EOD \- RaceLR0\.1633 \(0\.0110\)0\.1486 \(0\.0081\)RF0\.1425 \(0\.0094\)0\.1390 \(0\.0100\)GBDT0\.1538 \(0\.0090\)0\.1290 \(0\.0077\)MLP0\.1824 \(0\.0273\)0\.1549 \(0\.0382\)EOD \- EthnicityLR0\.0722 \(0\.0068\)0\.0525 \(0\.0090\)RF0\.0446 \(0\.0080\)0\.0154 \(0\.0081\)GBDT0\.0467 \(0\.0080\)0\.0218 \(0\.0087\)MLP0\.0702 \(0\.0356\)0\.0274 \(0\.0233\)EOD \- AgeLR0\.0637 \(0\.0112\)0\.6118 \(0\.0092\)RF0\.0829 \(0\.0134\)0\.4134 \(0\.0168\)GBDT0\.0471 \(0\.0101\)0\.4661 \(0\.0123\)MLP0\.0779 \(0\.0191\)0\.4601 \(0\.0367\)Figure 12:FPRs and TPRs across different subgroups are evaluated for four ML models \(LR, RF, GBDT, MLP\) for predicting premature treatment discontinuation with 2020\-2023 data\. The subfigures presents results stratified by \(a\) race, \(b\) ethnicity, and \(c\) age\. Bars represent the mean values over 10 runs, while error bars indicate the standard deviation computed from these runs\. Higher values indicate better results\.Figure 13:FPRs and TPRs across different subgroups are evaluated for four ML models \(LR, RF, GBDT, MLP\) for predicting treatment retention\>\>180 days with 2020\-2023 data\. The subfigures presents results stratified by \(a\) race, \(b\) ethnicity, and \(c\) age\. Bars represent the mean values over 10 runs, while error bars indicate the standard deviation computed from these runs\. Higher values indicate better results\.

## References

- J\. D\. Abernethy, P\. Awasthi, M\. Kleindessner, J\. Morgenstern, C\. Russell, and J\. ZhangActive sampling for min\-max fairness\.InInternational Conference on Machine Learning,pp\. 53–65\.Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p2.1)\.
- Abrahamet al\.\(2018\)A\. J\. Abraham, C\. M\. Andrews, M\. E\. Yingling, and J\. ShannonGeographic disparities in availability of opioid use disorder treatment for medicaid enrollees\.Health services research53\(1\),pp\. 389–404\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1111/1475-6773.12686)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p4.1)\.
- Acevedoet al\.\(2020\)A\. Acevedo, N\. Harvey, M\. Kamanu, S\. Tendulkar, and S\. FlearyBarriers, facilitators, and disparities in retention for adolescents in treatment for substance use disorders: a qualitative study with treatment providers\.Substance Abuse Treatment, Prevention, and Policy15,pp\. 1–13\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1186/s13011-020-00284-4)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p4.1)\.
- Acionet al\.\(2017\)L\. Acion, D\. Kelmansky, M\. van der Laan, E\. Sahker, D\. Jones, and S\. ArndtUse of a machine learning framework to predict substance use disorder treatment success\.PloS one12\(4\),pp\. e0175383\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1371/journal.pone.0175383)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p5.1)\.
- Agarwalet al\.\(2018\)A\. Agarwal, A\. Beygelzimer, M\. Dudík, J\. Langford, and H\. WallachA reductions approach to fair classification\.InInternational Conference on Machine Learning,pp\. 60–69\.Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p2.1),[§3\.4\.2](https://arxiv.org/html/2609.22113#S3.SS4.SSS2.p1.1),[§4\.4](https://arxiv.org/html/2609.22113#S4.SS4.p1.1),[§5](https://arxiv.org/html/2609.22113#S5.p5.1)\.
- Ahmadet al\.\(2025\)F\. Ahmad, J\. Cisewski, L\. Rossen, and P\. SuttonProvisional drug overdose death counts\.National Center for Health Statistics\.External Links:[Link](https://www.cdc.gov/nchs/nvss/vsrr/drug-overdose-data.htm)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p1.1)\.
- Al Faysalet al\.\(2024\)J\. Al Faysal, M\. Noor\-E\-Alam, G\. J\. Young, W\. Lo\-Ciganic, A\. J\. Goodin, J\. L\. Huang, D\. L\. Wilson, T\. W\. Park, and M\. M\. HasanAn explainable machine learning framework for predicting the risk of buprenorphine treatment discontinuation for opioid use disorder among commercially insured individuals\.Computers in Biology and Medicine177,pp\. 108493\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.compbiomed.2024.108493)Cited by:[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p2.1)\.
- Askariet al\.\(2020\)M\. S\. Askari, S\. S\. Martins, and P\. M\. MauroMedication for opioid use disorder treatment and specialty outpatient substance use treatment outcomes: differences in retention and completion among opioid\-related discharges in 2016\.Journal of Substance Abuse Treatment114,pp\. 108028\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jsat.2020.108028)Cited by:[item \(2\)](https://arxiv.org/html/2609.22113#S3.I1.ix2.p1.1),[§5](https://arxiv.org/html/2609.22113#S5.p1.1)\.
- Baileyet al\.\(2017\)Z\. D\. Bailey, N\. Krieger, M\. Agénor, J\. Graves, N\. Linos, and M\. T\. BassettStructural racism and health inequities in the usa: evidence and interventions\.The Lancet389\(10077\),pp\. 1453–1463\.Cited by:[§5](https://arxiv.org/html/2609.22113#S5.p3.1)\.
- Bairdet al\.\(2023\)A\. Baird, Y\. Cheng, and Y\. XiaDeterminants of outpatient substance use disorder treatment length\-of\-stay and completion: the case of a treatment program in the southeast us\.Scientific Reports13\(1\),pp\. 13961\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41598-023-41350-8)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p2.1),[§5](https://arxiv.org/html/2609.22113#S5.p1.1),[§5](https://arxiv.org/html/2609.22113#S5.p2.1),[§5](https://arxiv.org/html/2609.22113#S5.p6.1)\.
- Bellamyet al\.\(2018\)R\. K\. E\. Bellamy, K\. Dey, M\. Hind, S\. C\. Hoffman, S\. Houde, K\. Kannan, P\. Lohia, J\. Martino, S\. Mehta, A\. Mojsilovic, S\. Nagar, K\. N\. Ramamurthy, J\. Richards, D\. Saha, P\. Sattigeri, M\. Singh, K\. R\. Varshney, and Y\. ZhangAI Fairness 360: an extensible toolkit for detecting, understanding, and mitigating unwanted algorithmic bias\.External Links:[Link](https://arxiv.org/abs/1810.01943)Cited by:[§3\.4\.3](https://arxiv.org/html/2609.22113#S3.SS4.SSS3.p1.3)\.
- Binswangeret al\.\(2013\)I\. A\. Binswanger, P\. J\. Blatchford, S\. R\. Mueller, and M\. F\. SternMortality after prison release: opioid overdose and other causes of death, risk factors, and time trends from 1999 to 2009\.Annals of internal medicine159\(9\),pp\. 592–600\.Cited by:[§3\.1\.3](https://arxiv.org/html/2609.22113#S3.SS1.SSS3.p1.1)\.
- Birdet al\.\(2020\)S\. Bird, M\. Dudík, R\. Edgar, B\. Horn, R\. Lutz, V\. Milan, M\. Sameki, H\. Wallach, and K\. WalkerFairlearn: a toolkit for assessing and improving fairness in ai\.Cited by:[§3\.4\.3](https://arxiv.org/html/2609.22113#S3.SS4.SSS3.p1.3)\.
- Blancoet al\.\(2013\)C\. Blanco, M\. Iza, R\. P\. Schwartz, C\. Rafful, S\. Wang, and M\. OlfsonProbability and predictors of treatment\-seeking for prescription opioid use disorders: a national study\.Drug and Alcohol Dependence131\(1\-2\),pp\. 143–148\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.drugalcdep.2012.12.013)Cited by:[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p1.1)\.
- Blanco and Volkow \(2019\)C\. Blanco and N\. D\. VolkowManagement of opioid use disorder in the usa: present status and future directions\.The Lancet393\(10182\),pp\. 1760–1772\.Cited by:[§3\.1\.3](https://arxiv.org/html/2609.22113#S3.SS1.SSS3.p1.1)\.
- Brorsonet al\.\(2013\)H\. H\. Brorson, E\. A\. Arnevik, K\. Rand\-Hendriksen, and F\. DuckertDrop\-out from addiction treatment: a systematic review of risk factors\.Clinical psychology review33\(8\),pp\. 1010–1024\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.cpr.2013.07.007)Cited by:[§3\.1\.3](https://arxiv.org/html/2609.22113#S3.SS1.SSS3.p1.1)\.
- Centers for Disease Control and Prevention \(2024\)Centers for Disease Control and PreventionMedications for opioid use disorder \(MOUD\) study\.External Links:[Link](https://www.cdc.gov/overdoseprevention/data-research/facts-stats/moud-study.html)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p1.1)\.
- Centers for Disease Control and Prevention \(2025\)Centers for Disease Control and PreventionNonfatal Drug Overdose Surveillance and Epidemiology Syndromic Surveillance \(DOSE\-SYS\) System\.US Department of Health and Human Services, CDC\.External Links:[Link](https://www.cdc.gov/overdose-prevention/data-research/facts-stats/dose-dashboard-nonfatal-surveillance-data.html)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p1.1)\.
- Chenet al\.\(2024\)F\. Chen, L\. Wang, J\. Hong, J\. Jiang, and L\. ZhouUnmasking bias in artificial intelligence: a systematic review of bias detection and mitigation strategies in electronic health record\-based models\.Journal of the American Medical Informatics Association31\(5\),pp\. 1172–1183\.Cited by:[§3\.4](https://arxiv.org/html/2609.22113#S3.SS4.p1.1)\.
- Chenet al\.\(2023\)R\. J\. Chen, J\. J\. Wang, D\. F\. Williamson, T\. Y\. Chen, J\. Lipkova, M\. Y\. Lu, S\. Sahai, and F\. MahmoodAlgorithmic fairness in artificial intelligence for medicine and healthcare\.Nature Biomedical Engineering7\(6\),pp\. 719–742\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41551-023-01056-8)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p4.1),[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p3.1)\.
- Chintaet al\.\(2025\)S\. V\. Chinta, Z\. Wang, A\. Palikhe, X\. Zhang, A\. Kashif, M\. A\. Smith, J\. Liu, and W\. ZhangAI\-driven healthcare: Fairness in AI healthcare: A survey\.PLOS digital health4\(5\)\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1371/journal.pdig.0000864)Cited by:[§5](https://arxiv.org/html/2609.22113#S5.p5.1)\.
- Donget al\.\(2023\)H\. Dong, E\. J\. Stringfellow, W\. A\. Russell, and M\. S\. JalaliRacial and ethnic disparities in buprenorphine treatment duration in the US\.JAMA psychiatry80\(1\),pp\. 93–95\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1001/jamapsychiatry.2022.3673)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p3.1)\.
- Donget al\.\(2021\)X\. Dong, J\. Deng, S\. Rashidian, K\. Abell\-Hart, W\. Hou, R\. N\. Rosenthal, M\. Saltz, J\. H\. Saltz, and F\. WangIdentifying risk of opioid use disorder for patients taking opioid medications with deep learning\.Journal of the American Medical Informatics Association28\(8\),pp\. 1683–1693\.Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p2.1)\.
- Feldmanet al\.\(2015\)M\. Feldman, S\. A\. Friedler, J\. Moeller, C\. Scheidegger, and S\. VenkatasubramanianCertifying and removing disparate impact\.Inproceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining,pp\. 259–268\.Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p2.1)\.
- Fernandez\-Lozanoet al\.\(2021\)C\. Fernandez\-Lozano, P\. Hervella, V\. Mato\-Abad, M\. Rodríguez\-Yáñez, S\. Suárez\-Garaboa, I\. López\-Dequidt, A\. Estany\-Gestal, T\. Sobrino, F\. Campos, J\. Castillo,et al\.Random forest\-based prediction of stroke outcome\.Scientific Reports11\(1\),pp\. 10071\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41598-021-89434-7)Cited by:[§3\.2](https://arxiv.org/html/2609.22113#S3.SS2.p1.1)\.
- Ferryman and Pitcan \(2018\)K\. Ferryman and M\. PitcanFairness in precision medicine\.Data & Society1\(1\),pp\. 1–18\.Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p1.1)\.
- Fishet al\.\(2016\)B\. Fish, J\. Kun, and Á\. D\. LelkesA confidence\-based approach for balancing fairness and accuracy\.InProceedings of the 2016 SIAM International Conference on Data Mining,pp\. 144–152\.Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p2.1)\.
- Florenceet al\.\(2021\)C\. Florence, F\. Luo, and K\. RiceThe economic burden of opioid use disorder and fatal opioid overdose in the united states, 2017\.Drug and alcohol dependence218,pp\. 108350\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.drugalcdep.2020.108350)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p1.1)\.
- Foryciarzet al\.\(2022\)A\. Foryciarz, S\. R\. Pfohl, B\. Patel, and N\. ShahEvaluating algorithmic fairness in the presence of clinical guidelines: the case of atherosclerotic cardiovascular disease risk estimation\.BMJ Health & Care Informatics29\(1\),pp\. e100460\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1136/bmjhci-2021-100460)Cited by:[§5](https://arxiv.org/html/2609.22113#S5.p5.1)\.
- Gautam and Singh \(2020\)P\. Gautam and P\. SinghA machine learning approach to identify socio\-economic factors responsible for patients dropping out of substance abuse treatment\.Am J Public Health8\(5\),pp\. 140–6\.Cited by:[§3\.1\.3](https://arxiv.org/html/2609.22113#S3.SS1.SSS3.p1.1)\.
- Gianfrancescoet al\.\(2018\)M\. A\. Gianfrancesco, S\. Tamang, J\. Yazdany, and G\. SchmajukPotential biases in machine learning algorithms using electronic health record data\.JAMA Internal Medicine178\(11\),pp\. 1544–1547\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1001/jamainternmed.2018.3763)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p5.1)\.
- Guoet al\.\(2025\)X\. Guo, T\. Wang, Y\. Guo, C\. Vivas\-Valencia, C\. Bauer, and Y\. GongState\-specific explainable machine learning for predicting premature dropout in medication for opioid use disorder\.In2025 IEEE/ACM Conference on Connected Health: Applications, Systems and Engineering Technologies \(CHASE\),Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p5.1)\.
- Hanet al\.\(2020\)D\. Han, S\. Lee, and D\. SeoUsing machine learning to predict opioid misuse among us adolescents\.Preventive medicine130,pp\. 105886\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.ypmed.2019.105886)Cited by:[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p2.1)\.
- Hardtet al\.\(2016\)M\. Hardt, E\. Price, and N\. SrebroEquality of opportunity in supervised learning\.Advances in Neural Information Processing Systems29\.Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p1.1),[§3\.3\.1](https://arxiv.org/html/2609.22113#S3.SS3.SSS1.p3.1),[§3\.4\.3](https://arxiv.org/html/2609.22113#S3.SS4.SSS3.p1.1),[§4\.4](https://arxiv.org/html/2609.22113#S4.SS4.p1.1)\.
- Hasanet al\.\(2021\)M\. M\. Hasan, G\. J\. Young, J\. Shi, P\. Mohite, L\. D\. Young, S\. G\. Weiner, and M\. Noor\-E\-AlamA machine learning based two\-stage clinical decision support system for predicting patients’ discontinuation from opioid use disorder treatment: retrospective observational study\.BMC Medical Informatics and Decision Making21,pp\. 1–21\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1186/s12911-021-01692-7)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p2.1),[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p3.1)\.
- Huanget al\.\(2022\)J\. Huang, G\. Galal, M\. Etemadi, and M\. VaidyanathanEvaluation and mitigation of racial bias in clinical machine learning models: scoping review\.JMIR Medical Informatics10\(5\),pp\. e36388\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.2196/36388)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p5.1)\.
- Jacobsonet al\.\(2020\)N\. Jacobson, J\. Horst, L\. Wilcox\-Warren, A\. Toy, H\. K\. Knudsen, R\. Brown, E\. Haram, L\. Madden, and T\. MolfenterOrganizational facilitators and barriers to medication for opioid use disorder capacity expansion and use\.The journal of behavioral health services & research47\(4\),pp\. 439–448\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1007/s11414-020-09706-4)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p2.1)\.
- James and Jordan \(2018\)K\. James and A\. JordanThe opioid crisis in black communities\.The Journal of Law, Medicine & Ethics46\(2\),pp\. 404–421\.Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p4.1)\.
- Jiang and Nachum \(2020\)H\. Jiang and O\. NachumIdentifying and correcting label bias in machine learning\.InInternational conference on artificial intelligence and statistics,pp\. 702–712\.Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p2.1)\.
- Joneset al\.\(2022\)C\. M\. Jones, C\. Shoff, K\. Hodges, C\. Blanco, J\. L\. Losby, S\. M\. Ling, and W\. M\. ComptonReceipt of telehealth services, receipt and retention of medications for opioid use disorder, and medically treated overdose among medicare beneficiaries before and during the covid\-19 pandemic\.JAMA psychiatry79\(10\),pp\. 981–992\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1001/jamapsychiatry.2022.2284)Cited by:[item \(2\)](https://arxiv.org/html/2609.22113#S3.I1.ix2.p1.1)\.
- Juhnet al\.\(2022\)Y\. J\. Juhn, E\. Ryu, C\. Wi, K\. S\. King, M\. Malik, S\. Romero\-Brufau, C\. Weng, S\. Sohn, R\. R\. Sharp, and J\. D\. HalamkaAssessing socioeconomic bias in machine learning algorithms in health care: a case study of the houses index\.Journal of the American Medical Informatics Association29\(7\),pp\. 1142–1151\.Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p5.1)\.
- Kamiran and Calders \(2012\)F\. Kamiran and T\. CaldersData preprocessing techniques for classification without discrimination\.Knowledge and Information Systems33\(1\),pp\. 1–33\.Cited by:[§3\.4\.1](https://arxiv.org/html/2609.22113#S3.SS4.SSS1.p1.1),[§4\.4](https://arxiv.org/html/2609.22113#S4.SS4.p1.1),[§5](https://arxiv.org/html/2609.22113#S5.p5.1)\.
- Kennedyet al\.\(2022\)A\. J\. Kennedy, C\. B\. Wessel, R\. Levine, K\. Downer, M\. Raymond, D\. Osakue, I\. Hassan, J\. S\. Merlin, and J\. M\. LiebschutzFactors associated with long\-term retention in buprenorphine\-based addiction treatment programs: a systematic review\.Journal of general internal medicine37\(2\),pp\. 332–340\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1007/s11606-020-06448-z)Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p3.1)\.
- Kimet al\.\(2019\)M\. P\. Kim, A\. Ghorbani, and J\. ZouMultiaccuracy: black\-box post\-processing for fairness in classification\.InProceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society,pp\. 247–254\.Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p2.1)\.
- Krawczyket al\.\(2021\)N\. Krawczyk, A\. R\. Williams, B\. Saloner, and M\. CerdáWho stays in medication treatment for opioid use disorder? a national study of outpatient specialty treatment settings\.Journal of Substance Abuse Treatment126,pp\. 108329\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jsat.2021.108329)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p2.1),[item \(2\)](https://arxiv.org/html/2609.22113#S3.I1.ix2.p1.1),[§3\.1\.1](https://arxiv.org/html/2609.22113#S3.SS1.SSS1.p2.1),[§3\.1\.3](https://arxiv.org/html/2609.22113#S3.SS1.SSS3.p1.1)\.
- Lappanet al\.\(2020\)S\. N\. Lappan, A\. W\. Brown, and P\. S\. HendricksDropout rates of in\-person psychosocial substance use disorder treatments: a systematic review and meta\-analysis\.Addiction115\(2\),pp\. 201–217\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1111/add.14793)Cited by:[§3\.1\.3](https://arxiv.org/html/2609.22113#S3.SS1.SSS3.p1.1)\.
- Leslieet al\.\(2021\)D\. Leslie, A\. Mazumder, A\. Peppin, M\. K\. Wolters, and A\. HagertyDoes “AI” stand for augmenting inequality in the era of covid\-19 healthcare?\.bmj372\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1136/bmj.n304)Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p3.1)\.
- Liet al\.\(2023\)F\. Li, P\. Wu, H\. H\. Ong, J\. F\. Peterson, W\. Wei, and J\. ZhaoEvaluating and mitigating bias in machine learning models for cardiovascular disease prediction\.Journal of biomedical informatics138,pp\. 104294\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jbi.2023.104294)Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p3.1)\.
- Lo\-Ciganicet al\.\(2019\)W\. Lo\-Ciganic, J\. L\. Huang, H\. H\. Zhang, J\. C\. Weiss, Y\. Wu, C\. K\. Kwoh, J\. M\. Donohue, G\. Cochran, A\. J\. Gordon, D\. C\. Malone,et al\.Evaluation of machine\-learning algorithms for predicting opioid overdose risk among medicare beneficiaries with opioid prescriptions\.JAMA Network Open2\(3\),pp\. e190968–e190968\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1001/jamanetworkopen.2019.0968)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p5.1)\.
- Lopezet al\.\(2024\)I\. Lopez, S\. Fouladvand, S\. Kollins, C\. A\. Chen, J\. Bertz, T\. Hernandez\-Boussard, A\. Lembke, K\. Humphreys, A\. S\. Miner, and J\. H\. ChenPredicting premature discontinuation of medication for opioid use disorder from electronic medical records\.InAMIA Annual Symposium Proceedings,Vol\.2023,pp\. 1067\.Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p3.1),[item \(2\)](https://arxiv.org/html/2609.22113#S3.I1.ix2.p1.1)\.
- Lundberg and Lee \(2017\)S\. M\. Lundberg and S\. LeeA unified approach to interpreting model predictions\.Advances in Neural Information Processing Systems30\.Cited by:[§3\.3\.2](https://arxiv.org/html/2609.22113#S3.SS3.SSS2.p1.1)\.
- McCartyet al\.\(2014\)D\. McCarty, L\. Braude, D\. R\. Lyman, R\. H\. Dougherty, A\. S\. Daniels, S\. S\. Ghose, and M\. E\. Delphin\-RittmonSubstance abuse intensive outpatient programs: assessing the evidence\.Psychiatric Services65\(6\),pp\. 718–726\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1176/appi.ps.201300249)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p2.1)\.
- Mehrabiet al\.\(2021\)N\. Mehrabi, F\. Morstatter, N\. Saxena, K\. Lerman, and A\. GalstyanA survey on bias and fairness in machine learning\.ACM Computing Surveys \(CSUR\)54\(6\),pp\. 1–35\.Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p5.1),[§5](https://arxiv.org/html/2609.22113#S5.p2.1)\.
- Mintzet al\.\(2020\)C\. M\. Mintz, N\. J\. Presnall, J\. M\. Sahrmann, J\. T\. Borodovsky, P\. E\. Glaser, L\. J\. Bierut, and R\. A\. GruczaAge disparities in six\-month treatment retention for opioid use disorder\.Drug and alcohol dependence213,pp\. 108130\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.drugalcdep.2020.108130)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p4.1)\.
- Morganet al\.\(2018\)J\. R\. Morgan, B\. R\. Schackman, J\. A\. Leff, B\. P\. Linas, and A\. Y\. WalleyInjectable naltrexone, oral naltrexone, and buprenorphine utilization and discontinuation among individuals treated for opioid use disorder in a united states commercially insured population\.Journal of Substance Abuse Treatment85,pp\. 90–96\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jsat.2017.07.001)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p1.1),[§5](https://arxiv.org/html/2609.22113#S5.p1.1)\.
- Morganet al\.\(2021\)J\. R\. Morgan, A\. Y\. Walley, S\. M\. Murphy, A\. Chatterjee, S\. E\. Hadland, J\. Barocas, B\. P\. Linas, and S\. A\. AssoumouCharacterizing initiation, use, and discontinuation of extended\-release buprenorphine in a nationally representative united states commercially insured cohort\.Drug and alcohol dependence225,pp\. 108764\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.drugalcdep.2021.108764)Cited by:[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p1.1)\.
- National Center for Health Statistics \(2025\)National Center for Health StatisticsU\.s\. overdose deaths decrease almost 27% in 2024\.External Links:[Link](https://www.cdc.gov/nchs/%20pressroom/nchs_press_releases/2025/20250514.htm)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p1.1)\.
- Nusinoviciet al\.\(2020\)S\. Nusinovici, Y\. C\. Tham, M\. Y\. C\. Yan, D\. S\. W\. Ting, J\. Li, C\. Sabanayagam, T\. Y\. Wong, and C\. ChengLogistic regression was as good as machine learning for predicting major chronic diseases\.Journal of Clinical Epidemiology122,pp\. 56–69\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jclinepi.2020.03.002)Cited by:[§3\.2](https://arxiv.org/html/2609.22113#S3.SS2.p1.1)\.
- Obermeyeret al\.\(2019\)Z\. Obermeyer, B\. Powers, C\. Vogeli, and S\. MullainathanDissecting racial bias in an algorithm used to manage the health of populations\.Science366\(6464\),pp\. 447–453\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1126/science.aax2342)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p4.1),[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p3.1),[§5](https://arxiv.org/html/2609.22113#S5.p3.1)\.
- O’Brienet al\.\(2020\)P\. L\. O’Brien, K\. Schrader, A\. Waddell, and N\. Mulvaney\-DayModels for medication\-assisted treatment for opioid use disorder, retention, and continuity of care\.Report Prepared for Office of Disability, Aging and Long\-Term Care Policy, Office of the Assistant Secretary for Planning and Evaluation, US Department of Health and Human Services\.Cited by:[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p1.1)\.
- Parikhet al\.\(2019\)R\. B\. Parikh, S\. Teeple, and A\. S\. NavatheAddressing bias in artificial intelligence in health care\.JAMA322\(24\),pp\. 2377–2378\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1001/jama.2019.18058)Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p1.1)\.
- Paulus and Kent \(2020\)J\. K\. Paulus and D\. M\. KentPredictably unequal: understanding and addressing concerns that algorithmic clinical prediction may increase health disparities\.NPJ Digital Medicine3\(1\),pp\. 99\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41746-020-0304-9)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p4.1)\.
- Pinedo \(2019\)M\. PinedoA current re\-examination of racial/ethnic disparities in the use of substance abuse treatment: do disparities persist?\.Drug and alcohol dependence202,pp\. 162–167\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.drugalcdep.2019.05.017)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p4.1)\.
- Riceet al\.\(2012\)J\. B\. Rice, A\. G\. White, H\. G\. Birnbaum, M\. Schiller, D\. A\. Brown, and C\. L\. RolandA model to identify patients at risk for prescription opioid abuse, dependence, and misuse\.Pain Medicine13\(9\),pp\. 1162–1173\.Cited by:[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p2.1)\.
- Saloneret al\.\(2017\)B\. Saloner, M\. Daubresse, and G\. C\. AlexanderPatterns of buprenorphine\-naloxone treatment for opioid use disorder in a multistate population\.Medical care55\(7\),pp\. 669–676\.Cited by:[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p1.1)\.
- Sampleset al\.\(2018\)H\. Samples, A\. R\. Williams, M\. Olfson, and S\. CrystalRisk factors for discontinuation of buprenorphine treatment for opioid use disorders in a multi\-state sample of medicaid enrollees\.Journal of substance abuse treatment95,pp\. 9–17\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jsat.2018.09.001)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p1.1)\.
- Sayreet al\.\(2002\)S\. L\. Sayre, J\. M\. Schmitz, A\. L\. Stotts, P\. M\. Averill, H\. M\. Rhoades, and J\. J\. GrabowskiDetermining predictors of attrition in an outpatient substance abuse program\.The American Journal of Drug and Alcohol Abuse28\(1\),pp\. 55–72\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1081/ADA-120001281)Cited by:[§3\.1\.3](https://arxiv.org/html/2609.22113#S3.SS1.SSS3.p1.1)\.
- Seyyed\-Kalantariet al\.\(2021\)L\. Seyyed\-Kalantari, H\. Zhang, M\. B\. McDermott, I\. Y\. Chen, and M\. GhassemiUnderdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under\-served patient populations\.Nature medicine27\(12\),pp\. 2176–2182\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41591-021-01595-0)Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p3.1)\.
- Staffordet al\.\(2022\)C\. Stafford, W\. J\. Marrero, R\. B\. Naumann, K\. H\. Lich, S\. Wakeman, and M\. S\. JalaliIdentifying key risk factors for premature discontinuation of opioid use disorder treatment in the united states: a predictive modeling study\.Drug and Alcohol Dependence237,pp\. 109507\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.drugalcdep.2022.109507)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p3.1),[§1](https://arxiv.org/html/2609.22113#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p2.1),[item \(1\)](https://arxiv.org/html/2609.22113#S3.I1.ix1.p1.1),[§3\.1\.1](https://arxiv.org/html/2609.22113#S3.SS1.SSS1.p2.1),[§3\.1\.3](https://arxiv.org/html/2609.22113#S3.SS1.SSS3.p1.1),[§5](https://arxiv.org/html/2609.22113#S5.p2.1),[§5](https://arxiv.org/html/2609.22113#S5.p6.1)\.
- Stahleret al\.\(2021\)G\. J\. Stahler, J\. Mennis, and D\. A\. BaronRacial/ethnic disparities in the use of medications for opioid use disorder \(moud\) and their effects on residential drug treatment outcomes in the US\.Drug and Alcohol Dependence226,pp\. 108849\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.drugalcdep.2021.108849)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p4.1),[§3\.1\.1](https://arxiv.org/html/2609.22113#S3.SS1.SSS1.p2.1),[§5](https://arxiv.org/html/2609.22113#S5.p1.1)\.
- Substance Abuse and Mental Health Services Administration \(SAMHSA\) \(2021\)Substance Abuse and Mental Health Services Administration \(SAMHSA\)Treatment Episode Data Set \(TEDS\) Discharges\.Substance Abuse and Mental Health Services Administration\.External Links:[Link](https://www.samhsa.gov/data/data-we-collect/teds-treatment-episode-data-set)Cited by:[§3\.1\.1](https://arxiv.org/html/2609.22113#S3.SS1.SSS1.p1.1),[§3](https://arxiv.org/html/2609.22113#S3.p1.1)\.
- Sugarmanet al\.\(2020\)O\. K\. Sugarman, M\. A\. Bachhuber, A\. Wennerstrom, T\. Bruno, and B\. F\. SpringgateInterventions for incarcerated adults with opioid use disorder in the united states: a systematic review with a focus on social determinants of health\.PloS one15\(1\),pp\. e0227968\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1371/journal.pone.0227968)Cited by:[§3\.1\.3](https://arxiv.org/html/2609.22113#S3.SS1.SSS3.p1.1)\.
- Thompsonet al\.\(2021\)H\. M\. Thompson, B\. Sharma, S\. Bhalla, R\. Boley, C\. McCluskey, D\. Dligach, M\. M\. Churpek, N\. S\. Karnik, and M\. AfsharBias and fairness assessment of a natural language processing opioid misuse classifier: detection and mitigation of electronic health record data disadvantages across racial subgroups\.Journal of the American Medical Informatics Association28\(11\),pp\. 2393–2403\.Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p4.1)\.
- Timkoet al\.\(2016\)C\. Timko, N\. R\. Schultz, M\. A\. Cucciare, L\. Vittorio, and C\. Garrison\-DiehnRetention in medication\-assisted treatment for opiate dependence: a systematic review\.Journal of Addictive Diseases35\(1\),pp\. 22–35\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1080/10550887.2016.1100960)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p1.1)\.
- Vaidyaet al\.\(2024\)A\. Vaidya, R\. J\. Chen, D\. F\. Williamson, A\. H\. Song, G\. Jaume, Y\. Yang, T\. Hartvigsen, E\. C\. Dyer, M\. Y\. Lu, J\. Lipkova,et al\.Demographic bias in misdiagnosis by computational pathology models\.Nature Medicine30\(4\),pp\. 1174–1190\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41591-024-02885-z)Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p3.1)\.
- Volkow and Blanco \(2021\)N\. D\. Volkow and C\. BlancoThe changing opioid crisis: development, challenges and opportunities\.Molecular Psychiatry26\(1\),pp\. 218–233\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41380-020-0661-4)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p1.1)\.
- Wanget al\.\(2024\)T\. Wang, K\. Zhang, J\. Cai, Y\. Gong, K\. R\. Choo, and Y\. GuoAnalyzing the impact of personalization on fairness in federated learning for healthcare\.Journal of Healthcare Informatics Research8\(2\),pp\. 181–205\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1007/s41666-024-00164-7)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p5.1)\.
- Warrenet al\.\(2022\)D\. Warren, A\. Marashi, A\. Siddiqui, A\. A\. Eijaz, P\. Pradhan, D\. Lim, G\. Call, and M\. DrasUsing machine learning to study the effect of medication adherence in opioid use disorder\.PLoS One17\(12\),pp\. e0278988\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1371/journal.pone.0278988)Cited by:[§1](https://arxiv.org/html/2609.22113#S1.p3.1),[§1](https://arxiv.org/html/2609.22113#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p2.1),[§5](https://arxiv.org/html/2609.22113#S5.p1.1),[§5](https://arxiv.org/html/2609.22113#S5.p6.1)\.
- Weertset al\.\(2024\)H\. Weerts, A\. Kelly\-Lyth, R\. Binns, and J\. Adams\-PrasslUnlawful proxy discrimination: a framework for challenging inherently discriminatory algorithms\.InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency,pp\. 1850–1860\.Cited by:[§5](https://arxiv.org/html/2609.22113#S5.p2.1)\.
- Weinsteinet al\.\(2017\)Z\. M\. Weinstein, H\. W\. Kim, D\. M\. Cheng, E\. Quinn, D\. Hui, C\. T\. Labelle, M\. Drainoni, S\. S\. Bachman, and J\. H\. SametLong\-term retention in office based opioid treatment with buprenorphine\.Journal of substance abuse treatment74,pp\. 65–70\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jsat.2016.12.010)Cited by:[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p1.1)\.
- Welshet al\.\(2023\)J\. W\. Welsh, S\. I\. Sitar, B\. D\. Hunter, M\. D\. Godley, and M\. L\. DennisSubstance use severity as a predictor for receiving medication for opioid use disorder among adolescents: an analysis of the 2019 TEDS\.Drug and alcohol dependence246,pp\. 109850\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.drugalcdep.2023.109850)Cited by:[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p2.1)\.
- Whiteet al\.\(2009\)A\. G\. White, H\. G\. Birnbaum, M\. Schiller, J\. Tang, and N\. P\. KatzAnalytic models to identify patients at risk for prescription opioid abuse\.The American journal of managed care15\(12\),pp\. 897–906\.Cited by:[§2\.1](https://arxiv.org/html/2609.22113#S2.SS1.p2.1)\.
- Williamset al\.\(2019\)A\. R\. Williams, E\. V\. Nunes, A\. Bisaga, F\. R\. Levin, and M\. OlfsonDevelopment of a cascade of care for responding to the opioid epidemic\.The American journal of drug and alcohol abuse45\(1\),pp\. 1–10\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1080/00952990.2018.1546862)Cited by:[item \(2\)](https://arxiv.org/html/2609.22113#S3.I1.ix2.p1.1)\.
- Williamset al\.\(2018\)A\. R\. Williams, E\. V\. Nunes, A\. Bisaga, H\. A\. Pincus, K\. A\. Johnson, A\. N\. Campbell, R\. H\. Remien, S\. Crystal, P\. D\. Friedmann, F\. R\. Levin,et al\.Developing an opioid use disorder treatment cascade: a review of quality measures\.Journal of Substance Abuse Treatment91,pp\. 57–68\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jsat.2018.06.001)Cited by:[item \(2\)](https://arxiv.org/html/2609.22113#S3.I1.ix2.p1.1)\.
- Yanget al\.\(2024\)Y\. Yang, E\. Gan, G\. K\. Dziugaite, and B\. MirzasoleimanIdentifying spurious biases early in training through the lens of simplicity bias\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 2953–2961\.Cited by:[§5](https://arxiv.org/html/2609.22113#S5.p1.1)\.
- Zhanget al\.\(2018\)B\. H\. Zhang, B\. Lemoine, and M\. MitchellMitigating unwanted biases with adversarial learning\.InProceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society,pp\. 335–340\.Cited by:[§2\.2](https://arxiv.org/html/2609.22113#S2.SS2.p2.1)\.
- Zhanget al\.\(2019\)Z\. Zhang, Y\. Zhao, A\. Canes, D\. Steinberg, O\. Lyashevska,et al\.Predictive analytics with gradient boosting in clinical medicine\.Annals of Translational Medicine7\(7\),pp\. 152\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.21037/atm.2019.03.29)Cited by:[§3\.2](https://arxiv.org/html/2609.22113#S3.SS2.p1.1)\.

Similar Articles

Fair and Calibrated Toxicity Detection with Robust Training and Abstention

arXiv cs.LG

This paper studies fairness in toxicity classification across three axes: ranking, calibration, and abstention. It compares ERM, reweighted ERM, and Group DRO methods with post-hoc interventions, finding that calibration disparity is a hidden fairness violation and that abstention itself can be unfair.

Detecting and Mitigating Bias by Treating Fairness as a Symmetry Operation

arXiv cs.AI

The paper proposes treating fairness as a symmetry operation in machine learning classifiers, implementing loss-based regularization to enforce invariance under swapping of sensitive attributes while holding merit features fixed. The framework achieves over 90% bias reduction with minimal accuracy loss and requires no causal graph knowledge.