Improving Access to Essential Medicines via Decision-Aware Machine Learning

arXiv cs.LG Papers

Summary

A novel decision-aware machine learning framework was deployed nationwide in Sierra Leone to allocate essential medicines, achieving a 19% increase in consumption and covering 2 million women and children under five.

arXiv:2607.20542v1 Announce Type: new Abstract: A critical challenge in healthcare systems in low- and middle-income countries (LMICs) is the efficient and equitable allocation of scarce resources, particularly essential medicines. This problem is complicated by limited high-quality data, which restricts the applicability of traditional data-driven techniques. We propose a novel decision-aware machine learning framework for essential medicines allocation, which additionally leverages multi-task learning to ensure sample efficiency and catalytic priors to ensure equitable allocation. In collaboration with the Sierra Leone national government, we performed a staggered, nationwide deployment of our system as a decision support tool. Our econometric evaluation finds an estimated 19% increase in consumption of allocated products in treated districts, demonstrating its efficacy at improving access to essential medicines. Our tool was subsequently scaled nationwide, covering an estimated 2 million women and children under five. Our work demonstrates how machine learning methods can improve efficiency at very low cost in resource-constrained global health settings.
Original Article
View Cached Full Text

Cached at: 07/24/26, 05:11 AM

# Improving Access to Essential Medicines via Decision-Aware Machine Learning
Source: [https://arxiv.org/html/2607.20542](https://arxiv.org/html/2607.20542)
1Patrick Bayoh3Jatu Abdulai3Lawrence Sandi3Francis Smart4Hamsa Bastani1,∗Osbert Bastani2,∗ 1Department of Operations, Information, and Decisions, University of Pennsylvania, Jon M\. Huntsman Hall, 3730 Walnut St, Philadelphia, PA 19104, USA 2Department of Computer and Information Science, University of Pennsylvania, 3330 Walnut St, Philadelphia, PA 19104, USA 3National Medical Supplies Agency, Sierra Leone 31 Murray Town Road, Freetown, Sierra Leone 4Ministry of Health and Sanitation, Sierra Leone Youyi Building Freetown, Freetown, Sierra Leone ∗To whom correspondence should be addressed; E\-mail: hamsab@wharton\.upenn\.edu, obastani@seas\.upenn\.edu

> A critical challenge in healthcare systems in Low\- and Middle\-Income Countries \(LMICs\) is the efficient and equitable allocation of scarce resources, particularly essential medicines \(?\)\. This problem is complicated by limited high\-quality data, which restricts the applicability of traditional data\-driven techniques \(?,?,?,?\)\. We propose a novel decision\-aware machine learning framework for essential medicines allocation, which additionally leverages multi\-task learning to ensure sample efficiency and catalytic priors to ensure equitable allocation\. In collaboration with the Sierra Leone national government, we performed a staggered, nationwide deployment of our system as a decision support tool\. Our evaluation finds an estimated 19% increased consumption of allocated products in treated districts, demonstrating its efficacy at improving access to essential medicines\. Our tool was subsequently scaled nationwide, covering an estimated 2 million women and children under five\. Our work demonstrates how machine learning methods can improve efficiency at very low cost in resource\-constrained global health settings\.

## 1Introduction

Machine learning has demonstrated enormous potential for improving healthcare, with applications ranging from screening and diagnosis \(?,?,?\), targeted testing \(?\), and automating medical record generation \(?\)\. Success stories have largely occurred in developed nations, which have high\-quality data available to train predictive models\. Yet, machine learning has unique potential to help alleviate challenges caused by infrastructure limitations in developing nations, such as poor stock management systems \(?\) and insufficient staff training in logistics and inventory management \(?, ?\)\. These limitations can introduce operational challenges that make it difficult to match supply to demand, leading to shortages of common medicines and medical supplies taken for granted in developed nations\. An assessment by the World Health Organization \(WHO\) found that on average across a representative sample of health facilities in several African countries \(including Sierra Leone\), less than 40% of maternal essential medicines were readily available \(?,?\)\. This scarcity can force vulnerable patients to pay inflated private\-sector prices or forgo treatment \(?\)\. While there have been efforts to address these challenges through traditional methodologies for improving operations \(?,?,?,?\), recent work demonstrates that the impact of these efforts are limited by issues such as misalignment between digital systems and supply chain operations \(?\) or due to insufficient technical expertise and infrastructure to manage complex supply chains \(?\)\.

We focus on Sierra Leone, which launched the Free Health Care Initiative \(FHCI\) in 2010, one of the largest healthcare initiatives and a top priority in its post\-civil war recovery \(?\)\. This program provides free medical care and products to pregnant women and children under five\. However, these essential medicines are only distributed to healthcare facilities once a quarter, which can lead to significant shortages in one location even if there is adequate supply in other locations\. Efforts to design more adaptive supply chains have failed due to logistical challenges\. For example, inspired by The Coca\-Cola Company’s highly effective supply chain, Project Last Mile \(PLM\) operated since 2018 with funding from the United States Agency for International Development \(USAID\) to build software to digitize supply chain data, with the goal of supporting more frequent and targeted restocking\. However, despite operating for half a decade, their systems have been implemented in only about 8% of facilities \(15 government hospitals and 103 health facilities\) as of 2023 due to logistical challenges \(?,?,?\)\.

In this context, a low\-cost and promising avenue to scalably reduce shortages is to better match limited supply with patient demand using a combination of prediction and optimization \(?,?,?\)\. In particular, given accurate facility\-level demand forecasts for a product, existing supply chain optimization techniques \(?,?,?\) optimize facility\-level allocations to minimize unmet patient demand\.

However, accurately forecasting demand is difficult because of its high variability, both due to demand shocks caused by natural disasters or disease outbreaks \(?, ?, ?\) and periodic trends caused by seasonal diseases such as malaria\. Existing strategies employed in global health contexts include relying on historical consumption data, morbidity\-based forecasting, or proxy estimates \(?,?\); however, these approaches tend to have limited accuracy due to demand variability\. Furthermore, due to the lack of modern computing infrastructure, decision\-makers in practice rely primarily on ad\-hoc manual forecasting or on Excel spreadsheets with very high forecasting errors \(?,?\)\.

Modern machine learning tools provide a promising path to improve the accuracy of demand forecasting \(?\)\. By leveraging high\-dimensional covariates, they can predict complex trends in demand that are beyond the reach of simpler methodologies\. The key obstacle to realizing the promise of these approaches is the lack of training data\. What little data is available has significant rates of missing or unreliable entries due to the same infrastructure limitations that cause operational challenges\. For instance, in Sierra Leone, the only digitized data that can be used to train facility\-level demand forecasting models is monthly consumption data recorded in the District Health Information Software 2 \(DHIS2\) system \(used by 53 of the 54 countries in Africa\) \(?\)\. This data has only been collected systematically across all health facilities since 2020; furthermore, it has a high rate of missing or unreliable entries\. These limitations make it difficult to train accurate predictive models based on high\-dimensional covariates\.

In this paper, we introduce a novel machine learning framework for constrained resource allocation that satisfies two key criteria: \(1\) it makes effective use of limited and noisy historical data, and \(2\) it is scalable and can be deployed in limited\-compute environments\. At a high level, our system leverages machine learning to predict demand based on features constructed from historical data, and then applies stochastic optimization to compute the best allocation based on this model’s predictions\. To tackle data scarcity and quality challenges, our system uses a multi\-task learning strategy to share data across different healthcare facilities \(?\) along with a novel decision\-aware learning algorithm \(?,?,?,?,?,?\) that aligns the prediction loss with the downstream allocation objective\. Furthermore, it uses catalytic priors \(?\) with auxiliary data sources to mitigate data inequity \(i\.e\., where poorer facilities have lower\-quality data and therefore noisier forecasts\)\.

In close collaboration with the Sierra Leone national government, we deployed our system nationwide as a decision support tool to improve allocation decisions for Sierra Leone’s FHCI policy\. Our system was first piloted in 5 randomly selected districts, enabling us to econometrically evaluate our deployment\. A Synthetic Differences\-in\-Differences analysis \(?\) found that consumption increased by 19% in the 5 treated districts, suggesting that our system significantly improved patient access to essential medicines it helped allocate\. Following the pilot, our system has been scaled nationwide, covering an estimated 2 million women and children under five \(?,?\)\. Importantly, these gains were extremely cost\-effective, requiring only a $30 monthly server fee and no additional workforce \(see Supplement §[2\.7](https://arxiv.org/html/2607.20542#S2.SS7)for a cost\-effectiveness analysis\)\. These results highlight how novel data\-efficient machine learning can address critical challenges in healthcare delivery in resource\-constrained settings\.

## 2Essential Medicines Allocation System

Maternal mortality remains one of the most pressing global health challenges, particularly in developing countries where access to essential health products is often limited by inefficient allocation systems \(?\)\. In Sierra Leone, despite the FHCI providing free medical care to pregnant women and children, maternal mortality rate stands at 717 per 100,000 live births — one of the highest in the world \(?,?\)\.

Lack of access to essential medicines is one of the key contributors to preventable maternal deaths \(?\)\. In Sierra Leone, the National Medical Supplies Agency \(NMSA\) \(part of the Ministry of Health and Sanitation \(MoHS\)\) was created to manage the procurement and distribution of medicines and medical supplies to public health facilities across the nation, including over 70 essential medicines for women and children\. Supply of these medicines largely relies on donations from international organizations\. Prior to our collaboration, the NMSA distributed these medicines across the country following a centralized two\-level push system \(?\), where the NMSA first allocated supply to 16 districts, which districts then allocate to individual health facilities in their catchment \(see details in Supplement §[1\.3](https://arxiv.org/html/2607.20542#S1.SS3)\)\. The two\-level system is for administrative purposes only; all warehouses and health facilities are owned and operated by the NMSA\. Allocation decisions were largely made via an Excel tool; however, it was unable to capture key demand patterns such as seasonality and other time\-varying fluctuations\. As a result, district pharmacists often made significant manual adjustments because they perceived that the tool’s estimates did not accurately reflect the actual needs of health facilities\. According to NMSA officials, approximately 42% of facility requests were unfulfilled on average\.

In practice, this shortage occurred even when the total supply was adequate—some health facilities received a surplus of medicines while others suffered shortages \(surpluses often did not carry over fully to the next quarter due to significant reported waste and expiration\)\. For instance, DHIS2 data showed significant heterogeneity in both stockouts and wastage—e\.g\., for the Tonkolili district, in 2022 Q1, 10% of facilities that had previous stockouts continued to face supply shortages, while 22% of facilities with available stock had excess supply that could have been redistributed\. These findings suggest more efficient allocation strategies that better match supply and demand as a promising path to significantly reducing shortages\.

![Refer to caption](https://arxiv.org/html/2607.20542v1/x1.png)Figure 1:System Overview\.Every quarter, our system extracts and processes data including the total available central stock and historical facility\-level consumption records from mSupply and the DHIS2 government database\. It then trains a decision\-aware prediction model that informs a stochastic optimization procedure to make allocation decisions\. The system further provides a picking list for frontline workers to collect supplies from designated warehouses, enabling efficient distribution to local health facilities\. Finally, the resulting patient consumption data is recorded for informing predictions and decisions in subsequent quarters\.Fig\.[1](https://arxiv.org/html/2607.20542#S2.F1)illustrates the system we developed and deployed\. Unlike the earlier two\-level approach, our system directly generates facility\-level allocations, which are provided to district pharmacists to support their decision\-making\. Our system begins by pulling monthly consumption data from DHIS2, along with quarterly data on total central stock available for allocation and their expiry dates from the mSupply warehouse management system \(see Supplement §[1\.1](https://arxiv.org/html/2607.20542#S1.SS1)for details\)\. Then, it performs significant preprocessing to ensure data reliability and construct informative features for prediction \(see Section[1\.2](https://arxiv.org/html/2607.20542#S1.SS2)for details\)\.

Our prediction and optimization algorithm \(illustrated in the box in Fig\.[1](https://arxiv.org/html/2607.20542#S2.F1)\) can be divided into two components\. First, we predict the demand distribution for each facility\-product pair using a novel decision\-aware machine learning framework \(described in the next section\)\. Second, given the demand forecasts, our optimization algorithm outputs allocation decisions directly from the central stock to individual health facilities, which are designed to minimize the expected shortage of medicines\. In particular, for each product, it aims to minimize*unmet demand*—the total number of units of requested product that go unfulfilled—across facilities\. To account for the stochastic nature of demand, we use the expected unmet demand according to probabilistic demand forecasts\. This optimization problem can be solved efficiently via linear programming using a sample average approximation \(?\)\. Once the allocation is determined, our system assigns each batch of supplies to a warehouse based on proximity, central stock availability, and product expiration dates \(see Section[1\.4](https://arxiv.org/html/2607.20542#S1.SS4.SSS0.Px2)for details\)\. Finally, new patient consumption data is recorded and used to retrain our model each quarter just before making allocation decisions\.

## 3Machine Learning for Demand Forecasting

Next, we describe our machine learning framework for predicting the demand distribution used in our optimization algorithm\. Given a training dataset constructed from historical demand data, we could apply a traditional strategy for time series forecasting such as Autoregressive Integrated Moving Average \(ARIMA\) \(?\)\. However, these approaches work poorly due to data scarcity—e\.g\., when we trained our model for 2023 Q2 allocations, each facility\-product pair had at most 37 monthly observations \(i\.e\., from January 2020 to January 2023\)\.

Instead, we designed a machine learning framework that leverages three key techniques: \(1\) multi\-task learning to share data across facilities, \(2\) catalytic priors to regularize the model in data\-poor regions to mitigate data inequity, and \(3\) a novel decision\-aware learning algorithm to focus predictive power on facilities that are most relevant to the downstream optimization problem\. We summarize our techniques below, and provide details in Section[1\.5](https://arxiv.org/html/2607.20542#S1.SS5)\.

We developed a multi\-task learning strategy that trains one demand prediction model for each of our two product types \(medicines and medical supplies/equipment\) across all facility\-product pairs within that group\. Our multi\-task learning strategy enables knowledge transfer from locations with more available data to ones with less available data \(?,?,?\)\. Specifically, our system constructs features that facilitate generalization across facility\-product pairs \(e\.g\., average demand in the past year\), as well as features that capture trends specific to a given facility\-product pair \(e\.g\., the facility type and location, product fixed effect\)\. Then, it trains a random forest to predict demand from these features \(we use a random forest since it consistently outperforms other methods; see Supplement §[1\.5\.4](https://arxiv.org/html/2607.20542#S1.SS5.SSS4)\)\.

While multi\-task learning can improve performance in data\-poor locations, there are still systematic differences between data\-poor and data\-rich locations \(e\.g\., missing data often arises disproportionately in poorer regions due to staffing shortages\)\. Such*missing\-not\-at\-random*data \(?\) can lead to covariate shift, which reduces prediction accuracy for data\-poor locations\. To mitigate this data inequity, our system leverages catalytic priors \(?\) to regularize predictions for data\-poor locations towards a simpler, less biased population\-based model\. Specifically, it first uses census and satellite data to estimate catchment population, and then estimates demand proportionally to the catchment population\. This strategy ensures complete coverage of all facilities without biases due to low\-quality data\. Our system integrates this simple model as a catalytic prior for our machine learning model, regularizing predictions for data\-poor locations towards the population\-based ones\.

Finally, our system leverages decision\-aware learning \(?, ?,?, ?\) to focus predictive power on instances most relevant to optimizing allocation\. Intuitively, not all predictions matter equally—e\.g\., since our goal is to prevent unmet demand, we are primarily interested in how much we should stock facilities that arelikelyto be insufficiently stocked for a particular product\. Whereas a standard machine learning approach would treat all facility\-product pairs equally when training a prediction model, our decision\-aware learning algorithm focuses attention on facility\-product pairs deemed more important to the decision\-making objective; in this case, it upweights observations corresponding to facilities that are likely to be under\-stocked\. We found that existing decision\-aware learning algorithms were either computationally intractable at our scale or incompatible with the rest of our prediction and optimization pipeline; thus, we developed a novel decision\-aware learning algorithm that could be integrated more easily\. Prior to deploying our framework, we validated it on historical data, showing that it outperformed existing approaches on a held\-out test set; see Supplement §[1\.5\.4](https://arxiv.org/html/2607.20542#S1.SS5.SSS4)\.

## 4Deployment

In collaboration with the NMSA, in May 2023, we deployed our system in 5 of the 16 districts in Sierra Leone to perform facility\-level allocations for the second quarter \(June through August\)\. These districts—Tonkolili, Falaba, Karene, Kono, and Pujehun \(see map in Fig\.[2](https://arxiv.org/html/2607.20542#S4.F2)\)—were selected by the central government based on a randomized allocation schedule\. Prior to our deployment, the government had already established both the overall supply as well as the total amounts to be allocated to control districts\. The remaining supply was then to be allocated to the treatment districts, maintaining independence of supply quantities between the two groups\. Additional deployment details are in Supplement §[2\.1](https://arxiv.org/html/2607.20542#S2.SS1)\.

Beyond the performance of the allocation system, a critical priority was ensuring that it could be seamlessly integrated into NMSA’s existing workflows\. To this end, we conducted two targeted training sessions for policymakers and frontline workers, providing the necessary technical knowledge for operating our tool and understanding its implications\. Implementation emphasized stakeholder engagement at every level—outputs were presented in a familiar format that matched pre\-existing workflows and reviewed by NMSA management as well as district pharmaceutical managers prior to finalization\.

Furthermore, our system functions as a decision support tool, where central and district\-level planners retain ultimate authority and can override the system when necessary\. We found that decision\-makers closely followed the algorithmic allocations—the normalized overlap between the actual and algorithmic allocations ranged from 0\.89 to 1 \(see Supplement §[2\.4](https://arxiv.org/html/2607.20542#S2.SS4)for details\)\. Given the high compliance rates, stakeholder buy\-in, and early indications of improved efficiency \(detailed in the following section\), the national government adopted our system and expanded its use to all public\-sector health facilities in Sierra Leone beginning in the third quarter of 2023\.

![Refer to caption](https://arxiv.org/html/2607.20542v1/figures/treatmentMap.png)Figure 2:Map of treatment distribution in 2023 Q2\.Yellow dots denote treated facilities \(i\.e\., those in the Tonkolili, Falaba, Karene, Kono, and Pujehun districts\), and purple dots denote control facilities \(i\.e\., those in the Kailahun, Kenema, Bombali, Koinadugu, Kambia, Port Loko, Bo, Bonthe, Moyamba, Western Area Rural, and Western Area Urban districts\)\.
## 5Evaluation

We evaluated the effectiveness of our system at improving allocation efficiency\. While our optimization objective aims to minimize unmet demand, we do not directly observe this quantity \(we only observe when patients receive medicine, not when they are turned away due to stockouts\)\. Instead, examined the equivalent objective of maximizingpatient consumption, which is directly observed—since patient demand \(which is fixed but unobserved\) equals consumption plus unmet demand, maximizing total consumption is equivalent to minimizing total unmet demand\. \(This equivalence holds in a very general setting under the assumption that facilities do not wastefully give away medicines to patients that do not need them, which is unlikely in a such a resource\-constrained environment; we provide theoretical and empirical support for this equivalence in Supplement §[2\.3](https://arxiv.org/html/2607.20542#S2.SS3)\.\) We overview our analyses here, and provide details in Section[2\.2](https://arxiv.org/html/2607.20542#S2.SS2)and Supplement §[2](https://arxiv.org/html/2607.20542#S2a); our results are summarized in Fig\.[3](https://arxiv.org/html/2607.20542#S5.F3)and Table[1](https://arxiv.org/html/2607.20542#S5.T1)\.

Our main analysis focuses on the second quarter of 2023 \(2023 Q2\), where our system was deployed in a randomly selected subset of districts, providing natural variation enabling us to estimate causal effects\. Specifically, the NMSA performed allocation using standard procedures for the 11 control districts, after which our system was used to distribute remaining stock across the 5 treated districts\. We show descriptive time series trends of average normalized consumption for the treatment and control groups in Fig\.[3](https://arxiv.org/html/2607.20542#S5.F3)\(left\)\. We estimated how much our system changed patient consumption levels in treated districts \(i\.e\., districts where our system was deployed\) compared to what would have happened without our intervention, known as the Average Treatment Effect on the Treated \(ATT\)\. We used a balanced panel dataset of 312 facilities in treated districts and 746 facilities in control districts using time series data beginning in 2022 Q3 through 2023 Q3, after which our tool was used nationwide\. \(Prior to the implementation of the data\-driven Excel allocation tool by Crown Agents in 2022 Q3, allocation procedures were highly inconsistent, rendering the data unreliable\.\)

Given the limited number of treated districts, we used a Synthetic Difference\-in\-Differences \(SynthDiD\) analysis \(?\) to analyze the impact of our deployment\. SynthDiD combines the strengths of Difference\-in\-Differences \(?\) \(which assumes similar trends between treated and untreated groups in the pre\-treated period\) and Synthetic Controls \(?\) \(which constructs a synthetic control group that is similar to the treated units in terms of observed characteristics and outcome trends in the pre\-treated period\)\. SynthDiD ensures that differences between facilities in treated districts and synthetic control facilities remain stable prior to treatment\.

Fig\.[3](https://arxiv.org/html/2607.20542#S5.F3)\(right\) shows the corresponding time series trends for the treatment and synthetic control groups, and results are shown in the “SynthDiD” row of Table[1](https://arxiv.org/html/2607.20542#S5.T1)\. The 5 treated districts experienced a statistically significant increase of 19% \(p<0\.01p<0\.01\) in consumption; the improvement percentage is calculated from SynthDiD counterfactual estimates as:

Average treatment outcome−Average counterfactual outcomeAverage counterfactual outcome\.\\displaystyle\\frac\{\\text\{Average treatment outcome\}\-\\text\{Average counterfactual outcome\}\}\{\\text\{Average counterfactual outcome\}\}\.We validated our SynthDiD approach using a standard event\-study analysis \(?\), which showed that there are no statistically significant differences between treated and synthetic control facilities prior to our intervention, and that the change in consumption emerged only after our system was deployed\. These results suggest that our system substantially improved access to essential medicines in treated districts\. We provide additional details on our main analysis in Section[2\.2](https://arxiv.org/html/2607.20542#S2.SS2)\.

![Refer to caption](https://arxiv.org/html/2607.20542v1/x2.png)\(a\)Treatment vs Control Groups\.
![Refer to caption](https://arxiv.org/html/2607.20542v1/x3.png)\(b\)Treatment vs Synthetic Control Groups\.

Figure 3:Average normalized consumption time trends\.Panel \(a\) compares treatment \(green\) versus control \(orange\) groups using raw data\. Panel \(b\) compares treatment \(green\) versus synthetic control \(orange\) groups constructed via SynthDiD\. Thexx\-axis shows time in quarters and theyy\-axis denotes average normalized consumption\. The vertical dashed line at 2023 Q1 marks the quarter immediately before deployment\.Next, we studied how these improvements were distributed; we summarize our analyses here and provide details in Supplement §[2\.5](https://arxiv.org/html/2607.20542#S2.SS5)\. First, the NMSA seeks to prioritize larger facilities such as hospitals, since these facilities are most relied upon to provide quality care; indeed, we found that our system confers systematically greater benefits \(36%,p<0\.05p<0\.05\) for larger facilities, while preserving consumption for smaller ones\. In contrast, smaller health facilities showed a smaller, statistically insignificant improvement\. They may face more fundamental challenges \(e\.g\., inadequate staffing or equipment\) that limit their ability to provide services or utilize medical supplies regardless of availability—for instance, during a field visit, we found a small health post closed for two consecutive days, and local residents shared that they often seek essential care at larger, more reliable facilities\.

Next, we sought to ensure our consumption gains are not concentrated among wealthier patients, but also benefit poorer patients with the greatest need\. To this end, we examined previously under\-served facilities \(i\.e\., facilities that experienced at least one stockout in our data prior to our deployment\)\. We found that these facilities also saw a significant increase in consumption \(32%,p<0\.01p<0\.01\), suggesting that our system successfully addressed potential biases from uneven data quality and availability, leading to more equitable resource allocation\. Similarly, we found that rural facilities \(which tend to serve poorer populations\) benefited from our tool without compromising urban ones\.

Finally, we performed several robustness checks to validate our main results\. The main checks are described below and summarized in Table[1](https://arxiv.org/html/2607.20542#S5.T1); Supplement §[2\.6](https://arxiv.org/html/2607.20542#S2.SS6)provides details as well as additional checks\. First, to ensure consistency under a simpler methodology, we performed a standard DiD analysis, finding a consistent 21% increase in consumption \(p<0\.01p<0\.01\) \(“DiD” row in Table[1](https://arxiv.org/html/2607.20542#S5.T1)\)\. Second, to better account for geographic factors, we performed a DiD analysis on geographically matched facility pairs \(i\.e\., facilities within a short distance of a district border, which have a neighboring facility with the opposite treatment status\), again yielding a consistent 21% increase in consumption \(p<0\.05p<0\.05\) \(“Matching” row\)\.

Third, we considered an alternative analysis where we compared consumption trends across products instead of facilities; this robustness check uses a substantially different design with additional data to build confidence in our results\. In particular, we considered 25 other products that were concurrently allocated using a different, pre\-existing mechanism\. The consumption levels for these products can be used as a control group throughout our study period acrossalldistricts in the country using a staggered treatment—i\.e\., an advantage of this analysis is that it can be performed not just for the partial deployment in 2023 Q2, but also for the nationwide implementation starting in Q3\. We found a consistent, statistically significant increase in consumption of 18% \(p<0\.001p<0\.001\) \(“Alt\. Control” row\)\.

Next, a potential concern with our evaluation is due to missing data from facilities that failed to record consumption in a particular month, which may bias our results\. Our main analysis dropped observations with missing outcomes, but we also performed robustness checks using multiple standard imputation strategies—low\-rank matrix completion \(?\), population\-based methods, and historical average consumption\. The resulting SynthDiD ATT estimates are all statistically significant and consistent with our main analysis \(“Imputation” rows\)\. We also performed additional tests suggesting that our results are not driven by differential missingness; see Supplement §[2\.6\.4](https://arxiv.org/html/2607.20542#S2.SS6.SSS4)for details\.

Table 1:Average Treatment Effect on the Treated \(ATT\)\.ATT estimates across different estimation strategies\. SynthDiD, our main specification, shows a 19% increase in consumption, with similar magnitudes under DiD and several imputation approaches\. Estimation using alternative control product data or geographic facility\-level matching across district borders also produce consistent positive effects\. For stockouts as the outcome, the point estimate is negative but statistically insignificant; this is expected since it is not our objective\.ModelCoef\.Std\. ErrorObs\.Improv\.%Dependent Variable = ConsumptionSynthDiD0\.116∗\(0\.046\)5,29019%DiD0\.128∗∗\(0\.046\)5,29021%Matching \(15km\)0\.121∗∗\(0\.054\)2,41521%Alt\. Control0\.095∗∗∗\(0\.022\)10,52018%Imputation \(low rank\)0\.037∗∗\(0\.014\)5,45515%Imputation \(avg cons\)0\.067∗∗\(0\.031\)5,45521%Imputation \(pop\)0\.076∗∗\(0\.028\)5,45527%Dependent Variable = StockoutSynthDiD−\-0\.280\(0\.179\)5,290−\-4\.6% \(Insig\)
- •Notes:p∗<0\.05\{\}^\{\*\}p<0\.05,p∗∗<0\.01\{\}^\{\*\*\}p<0\.01,p∗⁣∗∗<0\.001\{\}^\{\*\*\*\}p<0\.001\. Standard errors in parentheses\. Improvement % is relative to counterfactual mean\.

Finally, we examined the impact of our system on stockouts\. While reducing stockouts might seem like a natural objective, optimizing for fewer stockouts can actually produce highly undesirable allocations—e\.g\., an optimal strategy is to allocate zero supply to a small number of high\-volume facilities, thereby ensuring that the remaining facilities are well\-stocked\. We found a directional reduction in stockouts but it is not statistically significant \(p≈0\.12p\\approx 0\.12\) \(“Stockouts” rows\)—i\.e\., our system does not inadvertently increase stockouts\.

## 6Discussion

Our findings provide strong field evidence of the effectiveness of our novel machine learning framework for resource allocation, significantly and equitably improving access to essential medicines in a highly constrained environment\. By replacing manual effort with data\-driven decision making, our system streamlined supply chain management, reduced administrative burdens, and adapted to real\-time changing consumption patterns, enabling facilities to better align limited supply with patient demand\. Efficient allocation becomes increasingly important as aid becomes more constrained to ensure the limited available supply meets demand\.

To ensure sustainable impact, we developed a web application with an intuitive user interface, which is now owned by the Sierra Leone government\. This system integrates with their government databases to automate the entire process, from data extraction and processing to generating final allocation results\. Notably, it is highly cost\-effective, operating without any additional workforce requirements and costing only $30 per month nationally in server fees\.

Beyond Sierra Leone, we believe our framework can be easily adapted to other contexts—by design, it relies on commonly available data: total available central stock, monthly health facility consumption, and stock expiry dates\. This data often already exists in a unified, digitized format—e\.g\., 53 \(out of 54 total\) African countries use DHIS2 \(?\) and 16 use mSupply \(?\)\. The key tasks required would be adapting our system with local data collection systems and discussing our optimization objective with policy makers to incorporate country\-specific constraints\.

## References

## End Notes

##### Acknowledgments

We are grateful for the close partnership offered by the Sierra Leone Ministry of Health and Sanitation and the National Medical Supplies Agency; this partnership stemmed from an earlier collaboration with Macro\-Eyes and its team members \(Suvrit Sra, Vahid Rostami, Ashley Schmidt, Musa Komeh, Lydia Bernard\-Jones, and Rene Ishiwe\)\. We acknowledge invaluable research assistance from our team of RAs \(Allan Zhang, Cheng\-Ying Wu, Norris Chen, and Hingis Chang\)\. This draft benefited from helpful feedback from Prashant Yadav, Gerard Cachon, Gad Allon, and participants at the Machine Learning for Health, Symposium on Artificial Intelligence in Learning Health Systems \(SAIL\), INFORMS Annual Conference, Marketplace Innovation Workshop, MSOM Healthcare SIG, Purdue Operations Conference, and Workshop on AI & Analytics for Social Good\.

##### Funding:

This research was supported by generous funding from the Wharton AI & Analytics Initiative, Wharton Global Initiatives, and the Wharton Mack Institute for Innovation Management\.

##### Author contributions:

HB, OB, and AC contributed to the design and methodology of the research and to the writing of the manuscript\. AC analyzed the results, supervised by HB and OB\. AC, PB, JA, LS, and FS contributed to the implementation of the research\.

##### Competing interests:

There are no competing interests to declare\.

##### Ethics & Inclusion Statement:

This research was conducted in close partnership with the Sierra Leone Ministry of Health and Sanitation \(MoHS\) and the National Medical Supplies Agency \(NMSA\), subject to a data\-sharing Non\-Disclosure Agreement \(NDA\) and a formal Memorandum of Understanding \(MOU\) with the Government of Sierra Leone\. The MoHS and NMSA oversaw the system’s deployment and integration into national workflows\. We collaborated closely with policymakers and frontline personnel, conducted training sessions, and aligned our approach with existing government systems and priorities\. Roles and responsibilities were clearly defined in advance, and system ownership was transferred to the government after nationwide deployment\.

##### Materials & Correspondence\.

Correspondence should be addressed to hamsab@wharton\.upenn\.edu and obatani@seas\.upenn\.edu

## Supplementary Information

§[1](https://arxiv.org/html/2607.20542#S1a)describes the data and methods supporting the design of our allocation system; §[2](https://arxiv.org/html/2607.20542#S2a)describes the deployment of our system and the empirical evaluation of its effectiveness\.

## 1Data and Methods

First, we outline our data sources \(§[1\.1](https://arxiv.org/html/2607.20542#S1.SS1)\) and detail how we processed the raw data for training our machine learning model and our evaluation \(§[1\.2](https://arxiv.org/html/2607.20542#S1.SS2)\)\. Second, we review the allocation mechanism used in Sierra Leone prior to adopting our approach \(§[1\.3](https://arxiv.org/html/2607.20542#S1.SS3)\)\. Then, we formalize resource allocation as an optimization problem \(§[1\.4](https://arxiv.org/html/2607.20542#S1.SS4.SSS0.Px1)\)\. Finally, we present our end\-to\-end machine learning pipeline for demand estimation and decision\-aware allocation \(§[1\.5](https://arxiv.org/html/2607.20542#S1.SS5)\), which spans:

- •Multi\-Task Learning \(§[1\.5\.1](https://arxiv.org/html/2607.20542#S1.SS5.SSS1)\): We exploited cross\-facility and cross\-product patterns to improve predictive accuracy with limited data\.
- •Catalytic Priors \(§[1\.5\.2](https://arxiv.org/html/2607.20542#S1.SS5.SSS2)\): We constructed catalytic priors from population estimates to address data quality issues such as censoring and missing values\.
- •Decision\-Aware Learning \(§[1\.5\.3](https://arxiv.org/html/2607.20542#S1.SS5.SSS3)\): We proposed a novel method for re\-weighting training data to align the prediction objective with the downstream optimization objective, thus shifting the focus from prediction accuracy to improving public health\.

### 1\.1Datasets

##### List of public health facilities\.

We obtained information on the ID, latitude, longitude, and type of public health facilities in Sierra Leone \(?\); this data was cross\-verified with frontline staff at the National Medical Supplies Agency \(NMSA\)\.

##### Consumption data\.

We constructed outcomes and features for demand prediction based on data extracted from the District Health Information Software 2 \(DHIS2\) used by Sierra Leone Ministry of Health and Sanitation \(MoHS\) to collect and manage health data\. The country transitioned from paper\-based to electronic reporting of health data in public health facilities in 2019 \(?\)\. We extracted monthly facility\-level data on consumption, opening balance, closing balance, and stockouts of 62 medicines and medical supplies across all the public health facilities from January 2020 to November 2023\. 36 of these products were allocated via our tool \(see Table[SI 10](https://arxiv.org/html/2607.20542#S2.T10)\) and the remaining 26 were allocated via existing mechanisms \(see Table[SI 11](https://arxiv.org/html/2607.20542#S2.T11)\); the latter was either because the allocation for these products was controlled by an external third party, or because there was too little historical data to use our system\. Data from products allocated via existing mechanisms was used only to support our multi\-task prediction model, and for the robustness check of our main analysis using product\-level controls \(“Alt\. Control”\)\.

##### Supply data\.

We collected expiry dates111In line with the NMSA’s existing practice, we prioritized allocating near\-expiry products to larger districts, where higher consumption minimizes waste by ensuring supplies are used up prior to expiration\.and available supply of all medicines and medical supplies per quarter from mSupply \(?\), a pharmaceutical logistics and warehouse management system used in Sierra Leone \(and over 40 other countries\)\. Prior to each allocation quarter, local staff are required to conduct central stock counts and record the information in mSupply\. In addition, we used invoice records from the mSupply system to determine the stock received by each public health facility to evaluate compliance for district\-level allocations\.

##### Catchment population\.

Estimating granular population is particularly challenging in developing countries\. To address this challenge, we leveraged multiple publicly available datasets to estimate each health facility’s catchment population \(i\.e\., the number of people each health facility is expected to serve\): the WorldPop Global Project Population Data \(?\), the global friction surface dataset \(`Oxford/MAP/friction\_surface\_2019`\) from Google Earth \(?\), and satellite imagery \(?\) also accessed through Google Earth\. We used this data to create population estimates for our catalytic priors \(see §[1\.5\.2](https://arxiv.org/html/2607.20542#S1.SS5.SSS2)\)\.

### 1\.2Data Processing

##### Training data for demand prediction model\.

The data for training our demand prediction model is derived from the historical consumption data extracted from DHIS2\. We used this time series data to construct both the demandξt,n∗\\xi\_\{t,n\}^\{\*\}in the current periodttfor facilitynnthat we are trying to predict as well as the featuresxt,nx\_\{t,n\}for prediction\. For example, we constructed facility\-specific features, including: consumption, product, facility ID, facility type, latitude and longitude of the facility’s geo\-location, district, average consumption of the product for the facility in the past\{1,2,3,4,5,6\}\\\{1,2,3,4,5,6\\\}months, standard deviation of the consumption in the past 3 and 6 months, total sample size for the facility\-product pair, year, month, average consumption of the product across facilities in the past\{1,2,3,4,5,6,10\}\\\{1,2,3,4,5,6,10\\\}months\. These features were determined based on domain knowledge and feature engineering\.

This data can sometimes be unreliable due to random or inconsistent data entry at individual facilities, requiring careful preprocessing\. First, we excluded any observations where the inflow and outflow are inconsistent—i\.e\.,

Closing Balance≠Opening Balance\+Quantity Received−Quantity Dispensed\+Adjustment/Loss\\displaystyle\\neq\\text\{Opening Balance\}\+\\text\{Quantity Received\}\-\\text\{Quantity Dispensed\}\+\\text\{Adjustment/Loss\}Second, we removed observations where all recorded quantities were zero, since this likely indicates the use of a default value\. Third, we excluded extreme outliers—specifically, the top and bottom 5% of values, to mitigate the impact of potential data entry errors\.

One major challenge is that we only observe consumption, which does not equal demand when there is a stockout; this issue is called*demand censoring*\(?\)\. To ensure we are predicting actual demand, we used a standard strategy where we dropped censored observations \(?\)\. In particular, when constructing\(xt,n,ξt,n∗\)\(x\_\{t,n\},\\xi\_\{t,n\}^\{\*\}\)pairs for training, we only included observations where no stockout occurred; then, consumption is equal to the demand, so we can take it to beξt,n∗\\xi\_\{t,n\}^\{\*\}\. One limitation of this strategy is that it introduces covariate shift, since there may be systematic differences between time periods/facilities where stockouts occur and those where they do not occur\. Covariate shift has the potential to degrade performance of the model compared to what is expected based on test set evaluation\. Thus, we used unbiased population\-based models that do not suffer from censoring as a catalytic prior when training our predictive models to mitigate this covariate shift \(see §[1\.5\.2](https://arxiv.org/html/2607.20542#S1.SS5.SSS2)\)\.

##### Panel data for evaluation\.

Our econometric evaluation uses the same historical DHIS2 data as above\. In this case, the main preprocessing we performed is to remove unreliable data—i\.e\., excluding any observations where the inflow and outflow are inconsistent and removing observations where all recorded quantities were zero, as described above\. Importantly, our evaluation is performed using consumption outcomes, which are not affected by demand censoring \(we also perform a number of checks to ensure robustness to data missingness; see §[2\.6](https://arxiv.org/html/2607.20542#S2.SS6)\)\. As we show in §[2\.3](https://arxiv.org/html/2607.20542#S2.SS3), increasing consumption is mathematically equivalent to reducing unmet demand\.

We used data from 2022 Q3 to 2023 Q3\. We started in 2022 Q3 since this was the quarter when the NMSA began using a standardized, data\-driven Excel allocation tool by Crown Agents; prior to this quarter, allocation procedures were inconsistent, which can prevent reliable counterfactual estimation\. We constructed a balanced panel dataset with 1,058 facilities for evaluation\.

The scale of quarterly consumption at each facility depends on the type and size of the facility’s catchment population\. To account for these differences, we normalized each product’s consumption at the facility level by subtracting the mean consumption of that product across all facilities and dividing by the standard deviation:

NormalizedConsumptionn,m\\displaystyle\\text\{NormalizedConsumption\}\_\{n,m\}=Consumptionn,m−MeanConsumptionmStdConsumptionm,\\displaystyle=\\frac\{\\text\{Consumption\}\_\{n,m\}\-\\text\{MeanConsumption\}\_\{m\}\}\{\\text\{StdConsumption\}\_\{m\}\},wherennrepresents the facility,mmrepresents the product,MeanConsumptionm\\text\{MeanConsumption\}\_\{m\}is the average consumption of productmmacross facilities, andStdConsumptionm\\text\{StdConsumption\}\_\{m\}is the standard deviation of productmmacross facilities\. For each facility in each quarter, we then calculated the average normalized consumption that is available at that facility as:

FacilityAverageNormalizedConsumptionn\\displaystyle\\text\{FacilityAverageNormalizedConsumption\}\_\{n\}=∑m∈AvailableProductsnNormalizedConsumptionn,mAvailableProductsn\.\\displaystyle=\\frac\{\\sum\_\{m\\in\\text\{AvailableProducts\}\_\{n\}\}\\text\{NormalizedConsumption\}\_\{n,m\}\}\{\\text\{AvailableProducts\}\_\{n\}\}\.Typically, this sum is only over the products that we allocated via our tool \(except in our staggered SynthDiD approach, where we compared average normalized consumption of products we allocated vs\. products we did not allocate\)\. This approach ensures that the facility\-level average consumption reflects the relative performance across products while accounting for the variability in the consumption level of different products\. We show the time series trends of average normalized consumption for treatment vs\. control groups in Fig\.[3\(a\)](https://arxiv.org/html/2607.20542#S5.F3.sf1)in the main paper; these raw data trends qualitatively support our main findings in the paper\.

### 1\.3Existing Allocation Approach

Each quarter, the National Medical Supplies Agency \(NMSA\) of Sierra Leone allocates approximately 70\-100 free healthcare products specifically for women and children under five years old, with supply primarily dependent on international donations\. Our system focuses exclusively on allocating these products\. Notably, the National Medical Supplies Agency \(NMSA\) Act of 2017 establishes the NMSA as the sole entity responsible for the procurement and distribution of these products to all public health institutions \(?\)\. Public facilities are the only source where patients can obtain these specific medications for free\. While private providers may exist, they are not part of this supply chain, and a majority of patients are unlikely to purchase expensive drugs that are available for free\. Thus, efficient allocation is critical for equitable access\.

The distribution of central stock to facilities occurs quarterly based on a centralized two\-stage push system \(?\), where supplies move from the central government to districts, and then to local health facilities\. We focused on a subset of 36 products that are regularly distributed—chosen in collaboration with NMSA officials prior to our deployment—ensuring sufficient historical data for model training\.

Until the deployment of our tool in 2023 Q2, the process of computing the allocation of central stock to each health facility relied primarily on a complex Excel tool\. An important aspect was organizing health facilities into three administrative categories: District Medical Stores \(DMS\), District Hospitals \(DH\), and Western Area Hospitals \(WAH\)\. The DMS includes four facility types: Community Health Centers \(CHC\), Community Health Posts \(CHP\), Maternal and Child Health Posts \(MCHP\), and Clinics\. Prior to each allocation cycle, all public health facilities submit requests based on their recent three\-month rolling average of consumption\. Workers can provide this information based on DHIS2, mSupply, or their professional judgment\.

Upon receiving all facility requests, the NMSA implements a structured allocation process\. First, the NMSA determines the distribution proportions among the three primary healthcare facility categories, typically 70% of total stock to DMS across 16 districts, 15% to DH, and 15% to WAH\. Following this initial distribution, each district receives a specific allocation based on multiple criteria, including district population, poverty levels, product types, and submitted requests\. For example, if a DMS in a particular district is allocated 10% of the DMS share, it receives a quantity calculated as

total available central stock×0\.7×0\.1\.\\displaystyle\\text\{total available central stock\}\\times 0\.7\\times 0\.1\.In cases where central stock remains after the initial distribution and requests remain partially fulfilled, the NMSA makes further allocations based on the unfulfilled requests\.

Once the allocations have been finalized, the distribution process follows a two\-tier delivery system\. The NMSA executes the “first mile” delivery to all districts, after which each district manages the “last mile” distribution to individual health facilities under the NMSA’s guidance and supervision\.

Our system was designed to address several limitations of this Excel\-based approach\. Most importantly, the existing approach predicted demand based on the average demand over the previous three months; while this strategy is standard, it fails to capture shifts in demand due to seasonal fluctuations or demand spikes due to natural disasters or disease outbreaks\. Our system achieves better demand forecasting through a combination of multi\-task learning, catalytic priors, and decision\-aware learning; in addition, it also significantly automates and improves the data processing pipeline\. The Excel tool relied on manually collecting requests at the district level; however, due to tight timelines for making allocation decisions, the government often used outdated records instead\. Our system also performs systematic data validation to identify potentially erroneous entries such as unexpectedly large numbers of zeros or inconsistencies in stock balance calculations\. Finally, our system incorporates stock expiry dates into allocation decisions, which prioritizes sending near\-expiring products to larger facilities with more consistent demand to minimize waste\.

### 1\.4Supply Chain Optimization Algorithm

We designed an optimization algorithm to allocate limited medical resources to health facilities to minimize total unmet demand, defined as the total number of units of requested product that goes unfulfilled\. The unmet demand objective is aligned with the goals of NMSA policymakers since it reflects the quality of care provision and directly affects patient outcomes, particularly in regions with limited access to alternative care options\.

Our system optimizes the allocation of each product separately\. There is a fixed total available central stockb∈ℝb\\in\\mathbb\{R\}\(which we refer to as the*budget*\) to be distributed acrossN∈ℕN\\in\\mathbb\{N\}facilities\. Each facilityn∈\[N\]n\\in\[N\]has an estimated demandΞn\\Xi\_\{n\}, whereΞ∈ℝN\\Xi\\in\\mathbb\{R\}^\{N\}is a real\-valued random vector with an estimated distributionℙΞ\\mathbb\{P\}\_\{\\Xi\}\. We denote the allocation decision and facility stock on hand asa∈ℝNa\\in\\mathbb\{R\}^\{N\}ands∈ℝNs\\in\\mathbb\{R\}^\{N\}, respectively, whereana\_\{n\}is the allocation intended for facilitynnandsns\_\{n\}is the stock on hand at each facilitynn\. DemandΞn\\Xi\_\{n\}, allocationana\_\{n\}, and facility stocksns\_\{n\}are all measured in units of product\. Then, the expected unmet demand measures the amount of unmet demand on average across facilities and overΞ\\Xi:

𝔼Ξ​\[∑n∈\[N\]max⁡\{Ξn−an−sn,0\}\]\.\\displaystyle\\mathbb\{E\}\_\{\\Xi\}\\left\[\\sum\_\{n\\in\[N\]\}\\max\\\{\\Xi\_\{n\}\-a\_\{n\}\-s\_\{n\},0\\\}\\right\]\.\(1\)If the budgetbbis very large, we can choose allana\_\{n\}sufficiently high to ensure that our objective is minimized at approximately0\. If the budgetbbis very low compared to the total excess demand∑n∈\[N\]\(Ξn−sn\)\+\\sum\_\{n\\in\[N\]\}\(\\Xi\_\{n\}\-s\_\{n\}\)^\{\+\}, then most facilities will suffer stockouts, and we cannot do much better than\(∑n∈\[N\]\(Ξn−sn\)\+\)−b\.\(\\sum\_\{n\\in\[N\]\}\(\\Xi\_\{n\}\-s\_\{n\}\)^\{\+\}\)\-b\.However, we found that the budget is often on the order of the total excess demand \(i\.e\.,b≈∑n∈\[N\]\(Ξn−sn\)\+b\\approx\\sum\_\{n\\in\[N\]\}\(\\Xi\_\{n\}\-s\_\{n\}\)^\{\+\}\), potentially because the budget is adjusted over time to meet demand\. In this case, many facilities are both over\- and under\-stocked; therefore, minimizing the objective requires setting the allocationsana\_\{n\}to be as close to the excess demand\(Ξn−sn\)\+\(\\Xi\_\{n\}\-s\_\{n\}\)^\{\+\}as possible\.

We focused on allocating over a single period instead of multi\-period allocation\. This is because we found that forecasting demand beyond one quarter is far too noisy to be of value\.

##### Optimization strategy\.

WhenΞ\\Xiis constant, the optimal policy can be straightforwardly expressed as a linear program\. To account for the uncertainty inΞ\\Xi, we use*sample average approximation*\(SAA\), which takesKKdemand samples from the estimated demand distributionξ\(k\)∼ℙΞ\\xi^\{\(k\)\}\\sim\\mathbb\{P\}\_\{\\Xi\}\(fork∈\[K\]k\\in\[K\]\), and then optimizes the objective on average across these samples\. The resulting optimization problem is

a∗=\\displaystyle a^\{\*\}=arg​mina∈ℝN⁡1K​∑k=1K∑n=1Ncn\(k\)\\displaystyle\\operatorname\*\{arg\\,min\}\_\{a\\in\\mathbb\{R\}^\{N\}\}\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\sum\_\{n=1\}^\{N\}c\_\{n\}^\{\(k\)\}\(2\)subj\. tocn\(k\)≥ξn\(k\)−an−sn,∀k∈\[K\],∀n∈\[N\],\\displaystyle\\text\{subj\. to\}\\quad c\_\{n\}^\{\(k\)\}\\geq\\xi\_\{n\}^\{\(k\)\}\-a\_\{n\}\-s\_\{n\},\\quad\\forall k\\in\[K\],\\ \\forall n\\in\[N\],cn\(k\)≥0,∀k∈\[K\],∀n∈\[N\],\\displaystyle\\phantom\{\\text\{subj\. to\}\\quad\}c\_\{n\}^\{\(k\)\}\\geq 0,\\quad\\forall k\\in\[K\],\\ \\forall n\\in\[N\],an≥0,∀n∈\[N\],\\displaystyle\\phantom\{\\text\{subj\. to\}\\quad\}a\_\{n\}\\geq 0,\\quad\\forall n\\in\[N\],∑n=1Nan≤b\.\\displaystyle\\phantom\{\\text\{subj\. to\}\\quad\}\\sum\_\{n=1\}^\{N\}a\_\{n\}\\leq b\.
where vector inequalities are element\-wise,cn\(k\)c\_\{n\}^\{\(k\)\}denotes the unmet demand for facilitynnin samplekk, andbbis the budget\. The first two constraints ensurecn\(k\)=max⁡\{ξn\(k\)−an−sn,0\}c\_\{n\}^\{\(k\)\}=\\max\\\{\\xi\_\{n\}^\{\(k\)\}\-a\_\{n\}\-s\_\{n\},0\\\}, with one of these constraints necessarily binding to minimize the objective\. The last constraint ensures the total allocation does not exceed the budget\.

##### Warehouse matching\.

On the basis of the final allocation generated by our system, we determined the specific inventory in warehouses that should be shipped to each facility\. In line with Sierra Leone Ministry of Health’s Free Healthcare Initiative policy, we preferentially allocated faster\-expiring stock to higher\-volume facilities in which it was more likely to be used prior to expiration\. In particular, for each product, we first ranked the stock based on time to expiration\. Then, we iterated through facilities on the basis of a ranking provided by the NMSA \(typically, facilities with larger catchment populations are ranked higher\)\. For each facility, we allocated stock with the earliest expiry date across all warehouses, continuing until the facility’s allocation was fully met\.

### 1\.5Machine Learning Framework

We present our machine learning framework for demand forecasting\. Our system trains a random forest using multi\-task learning and catalytic priors; to do so, it uses a novel decision\-aware learning approach to better align predictions with the downstream optimization loss\.

#### 1\.5\.1Multi\-Task Learning

A common approach for demand estimation in supply chain management and global health is to fit a distribution to historical consumption patterns, e\.g\., one could estimate the mean and variance of each facility\-product pair separately on the basis of its historical data\. Except where otherwise noted, we elide the product from our notation; our algorithm applies the procedure we describe to train one model per product category \(specifically, one for medicines and one for medical supplies\) and uses it to generate predictions for each product\. Mathematically, it can be viewed as solving the following maximum likelihood problem:

ℓ~​\(μ,σ\)=−∑t=1T∑n=1Nlog⁡𝒩​\(ξt,n∗;μn,σn2\),\\displaystyle\\tilde\{\\ell\}\(\\mu,\\sigma\)=\-\\sum\_\{t=1\}^\{T\}\\sum\_\{n=1\}^\{N\}\\log\\mathcal\{N\}\(\\xi\_\{t,n\}^\{\*\};\\mu\_\{n\},\\sigma\_\{n\}^\{2\}\),where𝒩​\(x;μn,σn2\)\\mathcal\{N\}\(x;\\mu\_\{n\},\\sigma\_\{n\}^\{2\}\)is the Gaussian probability density function inxxwith meanμn\\mu\_\{n\}and standard deviationσn\\sigma\_\{n\}andℓ~​\(μ,σ\)\\tilde\{\\ell\}\(\\mu,\\sigma\)is the negative log\-likelihood\. This objectiveℓ~\\tilde\{\\ell\}decomposes across facilitiesn∈\[N\]n\\in\[N\], and the solution for a givennnis the empirical mean and variance\. However, this strategy cannot learn dynamic patterns such as seasonal effects or demand that is elevated for a period of time \(e\.g\., due to an outbreak\)\. Time series models such as ARIMA are also infeasible, owning to the limited number of observations we have for each facility\-product pair—instead, we must leverage cross\-facility and cross\-product correlations\.

Multi\-task learning allows us to train a single model on multiple interrelated tasks \(i\.e\., the different facility\-product pairs\)\. By aggregating data and transferring knowledge across related tasks, multi\-task learning increases the effective sample size for each task \(?\)\. Intuitively, demand may exhibit similar patterns across different facilities \(e\.g\., seasonal trends, regional trends\), and training a single model enables trends observed in a large fraction of facilities to be extended to other facilities for which data may be more scarce\. Consider the following hypothetical example: suppose that the model observes a rise in demand for antimalarial medicines during the rainy season at several data\-rich rural clinics; then, it can apply this pattern to other rural clinics for which there is insufficient data to identify this pattern\. Thus, patterns in one facility directly inform predictions for other facilities, increasing the overall prediction accuracy with limited data\.

In more detail, we first categorized all products into two types \(medicines, or medical supplies and equipment, as shown in Tables[SI 10](https://arxiv.org/html/2607.20542#S2.T10)&[SI 11](https://arxiv.org/html/2607.20542#S2.T11)\); we trained separate predictive models for these two categories\. We associated each demand observationξt,n∗\\xi\_\{t,n\}^\{\*\}with a covariate vectorxt,n∈ℝdx\_\{t,n\}\\in\\mathbb\{R\}^\{d\}, which consists of features constructed from \(1\) the facilitynn, \(2\) the time steptt, and \(3\) the historical demand over thekkprevious stepsξt−k,n∗,ξt−k\+1,n∗,…,ξt−1,n∗\\xi\_\{t\-k,n\}^\{\*\},\\xi\_\{t\-k\+1,n\}^\{\*\},\.\.\.,\\xi\_\{t\-1,n\}^\{\*\}\(it also includes features of the product being allocated\), based on domain knowledge from consultations with local medical experts and policy makers along with extensive feature engineering\. Specifically, our prediction model uses the following covariates: lagged consumption, product, facility ID, facility type, latitude and longitude of the facility’s geo\-location, district, average consumption of the product for the facility in the past\{1,2,3,4,5,6\}\\\{1,2,3,4,5,6\\\}months, standard deviation of the consumption in the past 3 and 6 months, total sample size for the facility\-product pair, year, month, average consumption of the product across facilities in the past\{1,2,3,4,5,6,10\}\\\{1,2,3,4,5,6,10\\\}months\.

Then, for each category \(medicines or medical supply/equipment\), we trained a single model on the resulting dataset\{\(xt,n,ξt,n∗\)\}t∈\[T\],n∈\[N\]\\\{\(x\_\{t,n\},\\xi\_\{t,n\}^\{\*\}\)\\\}\_\{t\\in\[T\],n\\in\[N\]\}\. In particular, we aimed to train predictorsμθ​\(x\)\\mu\_\{\\theta\}\(x\)andσθ​\(x\)\\sigma\_\{\\theta\}\(x\)for the mean and standard deviation, respectively, whereθ∈Θ\\theta\\in\\Thetaare the parameters, using the following objective:

ℓ~​\(θ\)=−∑t=1T∑n=1Nlog⁡𝒩​\(ξt,n∗;μθ​\(xt,n\),σθ​\(xt,n\)2\)\.\\displaystyle\\tilde\{\\ell\}\(\\theta\)=\-\\sum\_\{t=1\}^\{T\}\\sum\_\{n=1\}^\{N\}\\log\\mathcal\{N\}\(\\xi\_\{t,n\}^\{\*\};\\mu\_\{\\theta\}\(x\_\{t,n\}\),\\sigma\_\{\\theta\}\(x\_\{t,n\}\)^\{2\}\)\.\(3\)In practice, we found that the following strategy works well\. First, we used a random forest to fit the meanμθ​\(xt,n\)\\mu\_\{\\theta\}\(x\_\{t,n\}\), assuming the variance is constant\. Then, we fitσθ​\(xt,n\)\\sigma\_\{\\theta\}\(x\_\{t,n\}\)only on the basis of historical data for the facilitynnand the current product in question\.

To evaluate our multi\-task learning strategy, we compared to two techniques that do not use multi\-task learning\. First, we considered the 3\-month rolling average, which is the prediction strategy used by the existing Excel tool and more broadly is a common strategy for demand forecasting in LMICs \(?, ?\)\. Second, we compared to a standard distribution modeling approach, where we fit a demand distribution for each facility\-product pair based only on historical data from that facility\-product pair\. We examined the fit of 46 well\-known candidate distributions \(e\.g\., normal, beta, gamma, exponential\) and chose the Nakagami distribution as the one that best minimized the sum of squared errors between the fitted probability density function \(PDF\) and a histogram of the historical data\.

We also compared to two other standard model families, LASSO regression and neural networks \(NN\), trained using the same multi\-task learning pipeline as our random forest\. In this comparison, we focused on evaluating the mean squared error \(MSE\) ofμθ\\mu\_\{\\theta\}, since we fitσθ\\sigma\_\{\\theta\}separately\. Results are shown in Fig\.[SI 1](https://arxiv.org/html/2607.20542#S1.F1); as can be seen, random forests consistently achieved lower MSE across representative products on a held\-out test set\.

![Refer to caption](https://arxiv.org/html/2607.20542v1/x4.png)Figure SI 1:Prediction error comparison between different model families\.Random forests \(RF\) yield the smallest prediction error for representative essential medicines chosen by the NMSA, compared to using a rolling 3\-month average \(3mth Avg\), distribution modeling \(Dist\.\) described in §[1\.5\.1](https://arxiv.org/html/2607.20542#S1.SS5.SSS1), LASSO regression \(LASSO\), and neural networks \(NN\)\.
#### 1\.5\.2Catalytic Priors

We implemented*catalytic priors*\(?\) as a safeguard to improve our model’s robustness to inequities in data quality \(arising from missing data or censoring\)\. Oftentimes, low quality data come from poorer districts—thus, a model trained only on the available data may be biased in poorer districts, which may create unintentional disparities in predictive performance and downstream resource allocation\.

A natural strategy for mitigating such bias is to incorporate auxiliary data sources that are less likely to suffer from bias\. For example, in public health, a standard approach is population\-based resource allocation \(PBRA\), which uses population estimates to guide proportional resource allocation needs \(?\)\. Although population\-based prediction is noisier \(i\.e\., higher variance\) because it cannot capture time\-dependent patterns \(e\.g\., seasonality of demand\), it is less biased since the observations do not suffer from nonrandom missingness or censoring\. Catalytic priors \(?\) allow us to suitably trade off bias and variance by regularizing our random forest with the simpler but less biased population\-based model, which predicts demand only on the basis of an estimate of the at\-risk population in the catchment of a health facility\.

For a given product, denote the prediction of the population\-based model byξt,n0=μC​\(pn\)=r⋅C⋅pn\\xi\_\{t,n\}^\{0\}=\\mu\_\{C\}\(p\_\{n\}\)=r\\cdot C\\cdot p\_\{n\}, wherepnp\_\{n\}is our estimated population in the catchment of facilitynn;r∈ℝr\\in\\mathbb\{R\}is an estimated multiplier derived from census data to account for the at\-risk population \(defined by the percentage of women and children in a given area\); andC∈ℝC\\in\\mathbb\{R\}is a singleproduct\-specificparameter estimated from our historical data \(i\.e\., the average quarterly demand per unit of at\-risk population\)\. The catchment populationpnp\_\{n\}served by each health facility is not readily available, so we estimated it as follows:

1. 1\.First, we collected data on geographic coordinates of health facilities in Sierra Leone from several sources, including Google Maps and Geo\-Referenced Infrastructure and Demographic Data for Development \(GRID3\) \(?\)\.
2. 2\.Next, we used Google Earth engine’s satellite imagery datasets to compute the normalized difference vegetation index \(NDVI\) on a 10km×\\times10km patch around each facility at monthly resolution between January 2022 and December 2022\. NDVI serves as a proxy for vegetation density, which can indicate human activity\.222NDVI is derived from satellite images using the formula:NDVI=\(NIR−red\)/\(NIR\+red\)\\text\{NDVI\}=\(\\text\{NIR\}\-\\text\{red\}\)/\(\\text\{NIR\}\+\\text\{red\}\)\. Since vegetation reflects light in the near\-infrared \(NIR\) spectrum and absorbs light in the red spectrum, areas with higher photosynthesis activity exhibit larger NDVI values\.
3. 3\.Then we used “friction surface” data from Google Earth \(?\) to obtain the travel time between every facility and pixel of the area with potential human activity\. This allows us to define the catchment area on the basis of minimal travel time\.
4. 4\.Finally, we estimatedpnp\_\{n\}for health facilitynnby using data from WorldPop \(?\), which provides population count estimates for each 100m×\\times100m grid cell\.

We estimatedrr, the proportion of women and children, using 2015 Sierra Leone Census data \(?\) at the chiefdom level\.333Chiefdoms are a more granular administrative unit than districts; there are 190 chiefdoms in Sierra Leone\. We further verified these estimates using data from the United Nations Office for the Coordination of Humanitarian Affairs \(OCHA\) \(?\)

Then, we followed \(?\), generating synthetic observations fromμC\\mu\_\{C\}to act as a Bayesian prior for our random forest\. In particular, we usedμC\\mu\_\{C\}to construct a single synthetic example\(xt,n,μC​\(pn\)\)\(x\_\{t,n\},\\mu\_\{C\}\(p\_\{n\}\)\)for each facilityn∈\[N\]n\\in\[N\]and time periodt∈\[T\]t\\in\[T\]\. We then trainedμθ^\\mu\_\{\\hat\{\\theta\}\}on a weighted combination of the original dataset and this synthetic dataset\. Intuitively, the synthetic dataset regularizesμθ^\\mu\_\{\\hat\{\\theta\}\}towards a stable estimateμC\\mu\_\{C\}in data\-poor regions of the covariate space\.

#### 1\.5\.3Decision\-Aware Learning

The next step is to incorporate our predictions into the optimization model to generate allocation decisions\. However, the model is usually trained to forecast demand using a standard objective such as mean\-squared error \(MSE\), which focuses on minimizing prediction error and ignores the decision error in the downstream optimization problem, making it*decision\-blind*\. This can result in poor performance since it may not focus the capacity of the machine learning model on predictions that are actually relevant to making decisions \(?, ?\)\.

Recent work has proposed algorithms that incorporate the downstream optimization objective as a*decision loss*in the training algorithm \(?, ?\)\. We found that existing decision\-aware learning algorithms were either computationally intractable at our scale or incompatible with the rest of our prediction and optimization pipeline\. Thus, we developed a novel and light\-weight decision\-aware learning approach; it relies only on re\-weighting observations in model training, which can be easily integrated with existing data pipelines\.

In our setting, the decision loss is

ℓ​\(μ^;μ∗\)=L​\(a∗​\(μ^\),μ∗\),\\displaystyle\\ell\(\\hat\{\\mu\};\\mu^\{\*\}\)=L\(a^\{\*\}\(\\hat\{\\mu\}\),\\mu^\{\*\}\),where

L​\(a;μ\)=∑n=1N𝔼Ξn∼𝒩​\(μn,σ2\)​\[max⁡\{Ξn−an−sn,0\}\],\\displaystyle L\(a;\\mu\)=\\sum\_\{n=1\}^\{N\}\\mathbb\{E\}\_\{\\Xi\_\{n\}\\sim\\mathcal\{N\}\(\\mu\_\{n\},\\sigma^\{2\}\)\}\\left\[\\max\\\{\\Xi\_\{n\}\-a\_\{n\}\-s\_\{n\},0\\\}\\right\],is the unmet demand for allocationa∈ℝNa\\in\\mathbb\{R\}^\{N\}assuming the true demand for facilityn∈\[N\]n\\in\[N\]is𝒩​\(μn,σ2\)\\mathcal\{N\}\(\\mu\_\{n\},\\sigma^\{2\}\)\(recall that for training the random forest, we have assumed that the standard deviation is a fixed valueσ\\sigma\), and where

a∗​\(μ\)\\displaystyle a^\{\*\}\(\\mu\)=arg⁡mina∈ℝN⁡L​\(a;μ\)\\displaystyle=\\arg\\min\_\{a\\in\\mathbb\{R\}^\{N\}\}L\(a;\\mu\)s\.t\.an≥0\(∀n∈\[N\]\),∑n=1Nan≤b\.\\displaystyle a\_\{n\}\\geq 0\\quad\(\\forall n\\in\[N\]\),\\quad\\sum\_\{n=1\}^\{N\}a\_\{n\}\\leq b\.
is the optimal allocation assuming the demand distributions forn∈\[N\]n\\in\[N\]is𝒩​\(μn,σ2\)\\mathcal\{N\}\(\\mu\_\{n\},\\sigma^\{2\}\)\. In other words, the decision loss is the expected unmet demand incurred when using the predictionsμ^\\hat\{\\mu\}in our optimization problem\. To address objective mismatch, we could trainμθ\\mu\_\{\\theta\}to directly minimize the decision loss:

θ^=arg⁡minθ∈Θ​∑t=1T∑n=1Nℓ​\(μθ​\(xt,n\);μt,n∗\)\.\\displaystyle\\hat\{\\theta\}=\\operatorname\*\{\\arg\\min\}\_\{\\theta\\in\\Theta\}\\sum\_\{t=1\}^\{T\}\\sum\_\{n=1\}^\{N\}\\ell\(\\mu\_\{\\theta\}\(x\_\{t,n\}\);\\mu\_\{t,n\}^\{\*\}\)\.Algorithms for doing so have been proposed in the setting of linear regression \(?\), and in the more general setting of differentiable model families by taking gradients through the optimization problem \(?, ?\)\. However, existing techniques are often limited to specific prediction setups or become computationally intractable for large\-scale problems \(?, ?, ?\)\.

Our strategy is to Taylor expand the optimal decision loss, which we will show can be interpreted as up\-weighting data more relevant to the downstream optimization problem\. This approach can also be easily integrated with existing pipelines and is flexible enough to handle a broad range of model families\. In particular, we approximated the decision loss by Taylor expanding it aroundμ^−μ0\\hat\{\\mu\}\-\\mu\_\{0\}\(whereμ0\\mu\_\{0\}is the current prediction andμ^=fθ​\(x\)\\hat\{\\mu\}=f\_\{\\theta\}\(x\)\), yielding

L​\(a∗​\(μ^\);μ∗\)≈L​\(a∗​\(μ0\);μ∗\)\+∇aL​\(a∗​\(μ0\);μ∗\)⊤​∇μa∗​\(μ0\)⊤​\(μ^−μ0\)\.\\displaystyle L\(a^\{\*\}\(\\hat\{\\mu\}\);\\mu^\{\*\}\)\\approx L\(a^\{\*\}\(\\mu\_\{0\}\);\\mu^\{\*\}\)\+\\nabla\_\{a\}L\(a^\{\*\}\(\\mu\_\{0\}\);\\mu^\{\*\}\)^\{\\top\}\\nabla\_\{\\mu\}a^\{\*\}\(\\mu\_\{0\}\)^\{\\top\}\(\\hat\{\\mu\}\-\\mu\_\{0\}\)\.\(4\)Since the first term is a constant, we can ignore it; in particular, we have

ℓ​\(μ^;μ∗\)\\displaystyle\\ell\(\\hat\{\\mu\};\\mu^\{\*\}\)≈∇aL​\(a∗​\(μ0\);μ∗\)⊤​∇μa∗​\(μ0\)⊤​\(μ^−μ0\)\+const\\displaystyle\\approx\\nabla\_\{a\}L\(a^\{\*\}\(\\mu\_\{0\}\);\\mu^\{\*\}\)^\{\\top\}\\nabla\_\{\\mu\}a^\{\*\}\(\\mu\_\{0\}\)^\{\\top\}\(\\hat\{\\mu\}\-\\mu\_\{0\}\)\+\\text\{const\}=∇aL​\(a∗​\(μ0\);μ∗\)⊤​∇μa∗​\(μ0\)⊤​\(μ^−μ∗\)\+const\.\\displaystyle=\\nabla\_\{a\}L\(a^\{\*\}\(\\mu\_\{0\}\);\\mu^\{\*\}\)^\{\\top\}\\nabla\_\{\\mu\}a^\{\*\}\(\\mu\_\{0\}\)^\{\\top\}\(\\hat\{\\mu\}\-\\mu^\{\*\}\)\+\\text\{const\}\.Here, we have replacedμ0\\mu\_\{0\}withμ∗\\mu^\{\*\}; since both of these are constants, it does not affect the optimal solution\. With this replacement, we can upper bound the term by its absolute value to avoid “overshooting”ξ∗\\xi^\{\*\}, resulting in a weighted absolute error loss:

ℓ​\(μ^;μ∗\)\\displaystyle\\ell\(\\hat\{\\mu\};\\mu^\{\*\}\)=∑t=1T∑n=1Nwt,n​\(μ^t,n−μt,n∗\)\+const\\displaystyle=\\sum\_\{t=1\}^\{T\}\\sum\_\{n=1\}^\{N\}w\_\{t,n\}\(\\hat\{\\mu\}\_\{t,n\}\-\\mu^\{\*\}\_\{t,n\}\)\+\\text\{const\}≤∑t=1T∑n=1N\|wt,n\|⋅\|μ^t,n−μt,n∗\|\+const,\\displaystyle\\leq\\sum\_\{t=1\}^\{T\}\\sum\_\{n=1\}^\{N\}\|w\_\{t,n\}\|\\cdot\|\\hat\{\\mu\}\_\{t,n\}\-\\mu^\{\*\}\_\{t,n\}\|\+\\text\{const\},where

wt,n=\(∇μa∗​\(μ0\)⊤​∇aL​\(a∗​\(μ0\);μ∗\)\)t,n\\displaystyle w\_\{t,n\}=\\left\(\\nabla\_\{\\mu\}a^\{\*\}\(\\mu\_\{0\}\)^\{\\top\}\\nabla\_\{a\}L\(a^\{\*\}\(\\mu\_\{0\}\);\\mu^\{\*\}\)\\right\)\_\{t,n\}Our algorithm uses this upper bound as the loss function, which works with any standard machine learning algorithm that can take weighted examples\. In particular, we trained our random forest to minimize this loss on the training data:

θ^=arg⁡minθ​∑t=1T∑n=1N\|wt,n\|⋅\|μθ​\(xt,n\)−ξt,n∗\|\.\\displaystyle\\hat\{\\theta\}=\\operatorname\*\{\\arg\\min\}\_\{\\theta\}\\sum\_\{t=1\}^\{T\}\\sum\_\{n=1\}^\{N\}\|w\_\{t,n\}\|\\cdot\|\\mu\_\{\\theta\}\(x\_\{t,n\}\)\-\\xi\_\{t,n\}^\{\*\}\|\.Note that we do not observe the true demandμt,n∗\\mu\_\{t,n\}^\{\*\}, so we useξt,n∗\\xi\_\{t,n\}^\{\*\}as an estimate\. Finally, as a heuristic, we replaced the absolute error with the squared error, which is more computationally efficient\.

An important insight here is that the weight can be interpreted as re\-weighting training examples\. Thewnw\_\{n\}can be decomposed as two gradients, and this can be computed numerically efficiently for a general class of convex programs \(?\)\. In our case, we derived the weights analytically by solving the optimization problem:

a∗​\(μ\)=\\displaystyle a^\{\*\}\(\\mu\)=arg⁡mina∈ℝN​∑n=1N𝔼Ξn∼𝒩​\(μn,σ2\)​\[max⁡\{Ξn−an−sn,0\}\]\\displaystyle\\operatorname\*\{\\arg\\min\}\_\{a\\in\\mathbb\{R\}^\{N\}\}\\sum\_\{n=1\}^\{N\}\\mathbb\{E\}\_\{\\Xi\_\{n\}\\sim\\mathcal\{N\}\(\\mu\_\{n\},\\sigma^\{2\}\)\}\\left\[\\max\\\{\\Xi\_\{n\}\-a\_\{n\}\-s\_\{n\},0\\\}\\right\]subj\. toan≥0​\(∀n∈\[N\]\),∑n=1Nan≤b\.\\displaystyle\\text\{subj\. to\}\\quad a\_\{n\}\\geq 0\\penalty 10000\\ \(\\forall n\\in\[N\]\),\\quad\\sum\_\{n=1\}^\{N\}a\_\{n\}\\leq b\.
LettingΞn=μn\+ηn\\Xi\_\{n\}=\\mu\_\{n\}\+\\eta\_\{n\}, whereηn∼𝒩​\(0,σ2\)\\eta\_\{n\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}\)i\.i\.d\., we can form the Lagrangian:

L​\(a,λ\)=∑n=1N𝔼ηn​\[max⁡\{μn\+ηn−an−sn,0\}\]\+λ0​\(b−∑n=1Nan\)−∑n=1Nλn​an\.\\displaystyle L\(a,\\lambda\)=\\sum\_\{n=1\}^\{N\}\\mathbb\{E\}\_\{\\eta\_\{n\}\}\\left\[\\max\\\{\\mu\_\{n\}\+\\eta\_\{n\}\-a\_\{n\}\-s\_\{n\},0\\\}\\right\]\+\\lambda\_\{0\}\\left\(b\-\\sum\_\{n=1\}^\{N\}a\_\{n\}\\right\)\-\\sum\_\{n=1\}^\{N\}\\lambda\_\{n\}a\_\{n\}\.The KKT conditions are

0\\displaystyle 0=∇anL​\(a∗​\(μ\),λ∗​\(μ\)\)=−ℙηn​\[an∗​\(μ\)≤μn\+ηn−sn\]\+λ0∗​\(μ\)−λn∗​\(μ\)\\displaystyle=\\nabla\_\{a\_\{n\}\}L\(a^\{\*\}\(\\mu\),\\lambda^\{\*\}\(\\mu\)\)=\-\\mathbb\{P\}\_\{\\eta\_\{n\}\}\\left\[a^\{\*\}\_\{n\}\(\\mu\)\\leq\\mu\_\{n\}\+\\eta\_\{n\}\-s\_\{n\}\\right\]\+\\lambda\_\{0\}^\{\*\}\(\\mu\)\-\\lambda\_\{n\}^\{\*\}\(\\mu\)0\\displaystyle 0=∇λ0L​\(a∗​\(μ\),λ∗​\(μ\)\)=∑n=1Nan∗​\(μ\)−b\\displaystyle=\\nabla\_\{\\lambda\_\{0\}\}L\(a^\{\*\}\(\\mu\),\\lambda^\{\*\}\(\\mu\)\)=\\sum\_\{n=1\}^\{N\}a^\{\*\}\_\{n\}\(\\mu\)\-b0\\displaystyle 0=λn∗​\(μ\)​an∗​\(μ\)\\displaystyle=\\lambda\_\{n\}^\{\*\}\(\\mu\)a^\{\*\}\_\{n\}\(\\mu\)0\\displaystyle 0≤λn∗​\(μ\)\\displaystyle\\leq\\lambda\_\{n\}^\{\*\}\(\\mu\)0\\displaystyle 0≤an∗​\(μ\)\.\\displaystyle\\leq a^\{\*\}\_\{n\}\(\\mu\)\.Let

ℐ​\(μ\)=\{n∈\[N\]∣λn∗​\(μ\)=0\}\.\\displaystyle\\mathcal\{I\}\(\\mu\)=\\\{n\\in\[N\]\\mid\\lambda\_\{n\}^\{\*\}\(\\mu\)=0\\\}\.Note that ifn∉ℐ​\(μ\)n\\not\\in\\mathcal\{I\}\(\\mu\), thenan∗​\(μ\)=0a\_\{n\}^\{\*\}\(\\mu\)=0, so the first condition becomes

an∗​\(μ\)\\displaystyle a\_\{n\}^\{\*\}\(\\mu\)=\{μn−sn\+Fηn−1​\(1−λ0∗​\(μ\)\)if​n∈ℐ​\(μ\),0otherwise\.\\displaystyle=\\begin\{cases\}\\mu\_\{n\}\-s\_\{n\}\+F\_\{\\eta\_\{n\}\}^\{\-1\}\(1\-\\lambda\_\{0\}^\{\*\}\(\\mu\)\)&\\text\{if \}n\\in\\mathcal\{I\}\(\\mu\),\\\\ 0&\\text\{otherwise\}\.\\end\{cases\}whereFηnF\_\{\\eta\_\{n\}\}is the CDF ofηn\\eta\_\{n\}\. Since theηn\\eta\_\{n\}are i\.i\.d\., we write this CDF as simplyFηF\_\{\\eta\}\. Summing overn∈\[N\]n\\in\[N\], we have

b=∑n=1Nan∗​\(μ\)\\displaystyle b=\\sum\_\{n=1\}^\{N\}a\_\{n\}^\{\*\}\(\\mu\)=∑n∈ℐ​\(μ\)\(μn−sn\+Fη−1​\(1−λ0∗​\(μ\)\)\)\\displaystyle=\\sum\_\{n\\in\\mathcal\{I\}\(\\mu\)\}\\left\(\\mu\_\{n\}\-s\_\{n\}\+F\_\{\\eta\}^\{\-1\}\(1\-\\lambda\_\{0\}^\{\*\}\(\\mu\)\)\\right\)=\(∑n∈ℐ​\(μ\)μn−sn\)\+\|ℐ​\(μ\)\|⋅Fη−1​\(1−λ0∗​\(μ\)\)\.\\displaystyle=\\left\(\\sum\_\{n\\in\\mathcal\{I\}\(\\mu\)\}\\mu\_\{n\}\-s\_\{n\}\\right\)\+\|\\mathcal\{I\}\(\\mu\)\|\\cdot F\_\{\\eta\}^\{\-1\}\(1\-\\lambda\_\{0\}^\{\*\}\(\\mu\)\)\.Thus, we have

Fη−1​\(1−λ0∗​\(μ\)\)=b−∑n∈ℐ​\(μ\)\(μn−sn\)\|ℐ​\(μ\)\|,\\displaystyle F\_\{\\eta\}^\{\-1\}\(1\-\\lambda\_\{0\}^\{\*\}\(\\mu\)\)=\\frac\{b\-\\sum\_\{n\\in\\mathcal\{I\}\(\\mu\)\}\(\\mu\_\{n\}\-s\_\{n\}\)\}\{\|\\mathcal\{I\}\(\\mu\)\|\},so

an∗​\(μ\)\\displaystyle a\_\{n\}^\{\*\}\(\\mu\)=\{μn−sn\+1\|ℐ​\(μ\)\|​\(b−∑m∈ℐ​\(μ\)\(μm−sm\)\)if​n∈ℐ​\(μ\),0otherwise\.\\displaystyle=\\begin\{cases\}\\mu\_\{n\}\-s\_\{n\}\+\\frac\{1\}\{\|\\mathcal\{I\}\(\\mu\)\|\}\\left\(b\-\\sum\_\{m\\in\\mathcal\{I\}\(\\mu\)\}\(\\mu\_\{m\}\-s\_\{m\}\)\\right\)&\\text\{if \}n\\in\\mathcal\{I\}\(\\mu\),\\\\ 0&\\text\{otherwise\}\.\\end\{cases\}Taking the derivative with respect toμm\\mu\_\{m\}\(m∈ℐ​\(μ\)m\\in\\mathcal\{I\}\(\\mu\)\), we have

∇μman∗​\(μ\)=δm,n−1\|ℐ​\(μ\)\|\.\\displaystyle\\nabla\_\{\\mu\_\{m\}\}a^\{\*\}\_\{n\}\(\\mu\)=\\delta\_\{m,n\}\-\\frac\{1\}\{\|\\mathcal\{I\}\(\\mu\)\|\}\.Thus, we have

∇μmL​\(a∗​\(μ\),λ∗​\(μ\)\)\\displaystyle\\nabla\_\{\\mu\_\{m\}\}L\(a^\{\*\}\(\\mu\),\\lambda^\{\*\}\(\\mu\)\)=∇μm​∑n=1N𝔼ηn​\[max⁡\{μn\+ηn−an∗​\(μ\)−sn,0\}\]\\displaystyle=\\nabla\_\{\\mu\_\{m\}\}\\sum\_\{n=1\}^\{N\}\\mathbb\{E\}\_\{\\eta\_\{n\}\}\\left\[\\max\\\{\\mu\_\{n\}\+\\eta\_\{n\}\-a^\{\*\}\_\{n\}\(\\mu\)\-s\_\{n\},0\\\}\\right\]=∑n=1Nℙηn​\[an∗​\(μ\)≤μn\+ηn−sn\]​∇μman∗​\(μ\)\\displaystyle=\\sum\_\{n=1\}^\{N\}\\mathbb\{P\}\_\{\\eta\_\{n\}\}\\left\[a^\{\*\}\_\{n\}\(\\mu\)\\leq\\mu\_\{n\}\+\\eta\_\{n\}\-s\_\{n\}\\right\]\\nabla\_\{\\mu\_\{m\}\}a^\{\*\}\_\{n\}\(\\mu\)=∑n∈ℐ​\(μ\)ℙηn​\[an∗​\(μ\)≤μn\+ηn−sn\]​\(δm,n−1\|ℐ​\(μ\)\|\)\\displaystyle=\\sum\_\{n\\in\\mathcal\{I\}\(\\mu\)\}\\mathbb\{P\}\_\{\\eta\_\{n\}\}\\left\[a^\{\*\}\_\{n\}\(\\mu\)\\leq\\mu\_\{n\}\+\\eta\_\{n\}\-s\_\{n\}\\right\]\\left\(\\delta\_\{m,n\}\-\\frac\{1\}\{\|\\mathcal\{I\}\(\\mu\)\|\}\\right\)=ℙηm​\[am∗​\(μ\)≤μm\+ηm−sm\]−1\|ℐ​\(μ\)\|​∑n∈ℐ​\(μ\)ℙηn​\[an∗​\(μ\)≤μn\+ηn−sn\]\\displaystyle=\\mathbb\{P\}\_\{\\eta\_\{m\}\}\\left\[a^\{\*\}\_\{m\}\(\\mu\)\\leq\\mu\_\{m\}\+\\eta\_\{m\}\-s\_\{m\}\\right\]\-\\frac\{1\}\{\|\\mathcal\{I\}\(\\mu\)\|\}\\sum\_\{n\\in\\mathcal\{I\}\(\\mu\)\}\\mathbb\{P\}\_\{\\eta\_\{n\}\}\[a\_\{n\}^\{\*\}\(\\mu\)\\leq\\mu\_\{n\}\+\\eta\_\{n\}\-s\_\{n\}\]=ℙηm​\[am∗​\(μ\)≤μm\+ηm−sm\]\+const\\displaystyle=\\mathbb\{P\}\_\{\\eta\_\{m\}\}\\left\[a\_\{m\}^\{\*\}\(\\mu\)\\leq\\mu\_\{m\}\+\\eta\_\{m\}\-s\_\{m\}\\right\]\+\\text\{const\}≈𝕀​\[am∗​\(μ\)≤μm−sm\]\+const,\\displaystyle\\approx\\mathbb\{I\}\[a\_\{m\}^\{\*\}\(\\mu\)\\leq\\mu\_\{m\}\-s\_\{m\}\]\+\\text\{const\},where the approximation on the last line holds whenσ\\sigmais small\. Using this gradient, our predictive model’s objective can be approximated as

θ^=arg⁡minθ​∑t=1T∑n∈ℐ​\(μ\)\(𝕀​\[at,n∗​\(μ\)≤μt,n−st,n\]\+c\)⋅\|μθ​\(xt,n\)−ξt,n∗\|\.\\displaystyle\\hat\{\\theta\}=\\operatorname\*\{\\arg\\min\}\_\{\\theta\}\\sum\_\{t=1\}^\{T\}\\sum\_\{n\\in\\mathcal\{I\}\(\\mu\)\}\(\\mathbb\{I\}\[a\_\{t,n\}^\{\*\}\(\\mu\)\\leq\\mu\_\{t,n\}\-s\_\{t,n\}\]\+c\)\\cdot\|\\mu\_\{\\theta\}\(x\_\{t,n\}\)\-\\xi\_\{t,n\}^\{\*\}\|\.for some constantcc\. This loss up\-weights the training examples that are more likely to experience unmet demand\. One remaining issue is that we do not knowat,n∗​\(μ\)a\_\{t,n\}^\{\*\}\(\\mu\), which is required for the computation of the weight\. We approximated it by first training a decision\-blind model, using it to optimize the allocations, and then using these allocations to determine the weights\.

#### 1\.5\.4Out\-of\-Sample Comparison

Prior to deploying our framework, we validated it on historical data by showing that it outperforms several baselines on a held\-out test set, including:

- •Excel tool:Data\-driven tool previously used in conjunction with DHIS2 by the NMSA to make allocations\.
- •Decision\-blind ablation:Follows our pipeline but uses the MSE loss to train the random forest instead of our decision\-aware learning algorithm\.
- •StochOptForest:Follows our pipeline but trains decision\-aware random forests using the end\-to\-end optimization strategy by \(?\)\.
- •Distribution modeling:Estimates demand via distribution modeling \(as described in §[1\.5\.1](https://arxiv.org/html/2607.20542#S1.SS5.SSS1)\), and then optimizes allocations based on these forecasts\.
- •Global Health:Estimates demand via a 3\-month rolling average \(described in §[1\.5\.1](https://arxiv.org/html/2607.20542#S1.SS5.SSS1), common in global health \(?, ?\)\), and then optimizes allocations based on these forecasts \(?\)\.
- •Population\-based:Allocates total available central stock to chiefdoms proportionally to their at\-risk population \(women and children\), as commonly done in global health \(?\)\. Within a chiefdom, all facilities are treated equally\.444Note that current data sources do not provide facility\-level population estimates in Sierra Leone; we construct such estimates from a variety of data sources in §[1\.5\.2](https://arxiv.org/html/2607.20542#S1.SS5.SSS2)\.

Here, we introduce the indexm∈\[M\]m\\in\[M\]to denote products\. We applied our framework and evaluate the unmet demand for each facility\-product pair:

UnmetDemandm,n=\{0if​μm,n−am,n−sm,n<0μm,n−am,n−sm,notherwise\\text\{UnmetDemand\}\_\{m,n\}=\\begin\{cases\}0&\\text\{if \}\\mu\_\{m,n\}\-a\_\{m,n\}\-s\_\{m,n\}<0\\\\ \\mu\_\{m,n\}\-a\_\{m,n\}\-s\_\{m,n\}&\\text\{otherwise\}\\end\{cases\}\(5\)for all productsm∈\[M\]m\\in\[M\]and facilitiesn∈\[N\]n\\in\[N\], wheream,na\_\{m,n\}is allocation decision andμm,n\\mu\_\{m,n\}is the true demand\. Then, we computed the total unmet demand across all facilities and divided it by total demand to obtain a normalized total unmet demand, and averaged across products:555Note that we approximate total demand by total consumption\.

NormalizedTotalUnmetDemand=1\#​products​∑m=1M∑n=1NUnmetDemandm,n∑n=1Nμm,n\.\\displaystyle\\text\{NormalizedTotalUnmetDemand\}=\\frac\{1\}\{\\\#\\text\{ products\}\}\\sum\_\{m=1\}^\{M\}\\frac\{\\sum\_\{n=1\}^\{N\}\\text\{UnmetDemand\}\_\{m,n\}\}\{\\sum\_\{n=1\}^\{N\}\\mu\_\{m,n\}\}\.\(6\)To test allocation performance in challenging environments, we focused our evaluation on lower budgets—specifically, we used the 25th percentile of quarterly budgets \(by product\) observed in our data\. We then measured how much our approach reduces unmet demand compared to the baseline:

NormalizedTotalUnmetDemandBaseline−NormalizedTotalUnmetDemandOursNormalizedTotalUnmetDemandBaseline\.\\displaystyle\\frac\{\\text\{NormalizedTotalUnmetDemand\}\_\{\\text\{Baseline\}\}\-\\text\{NormalizedTotalUnmetDemand\}\_\{\\text\{Ours\}\}\}\{\\text\{NormalizedTotalUnmetDemand\}\_\{\\text\{Baseline\}\}\}\.\(7\)Our results are shown in Table[SI 1](https://arxiv.org/html/2607.20542#S1.T1); they illustrate that our approach outperforms other baselines\. First, it performs significantly better than non\-machine learning approaches, demonstrating the predictive power of machine learning\. Next, StochOptForest most likely performs poorly since it is unable to integrate its decision\-aware learning strategy with multi\-task learning\. At a high level, their algorithm assumes that there is a single optimization problem for each example in the dataset \(i\.e\., one for each facility\-product pair\)\. However, in our problem, a single product’s optimization problem is associated with many examples \(i\.e\., observations across all facilities\)\. Thus, to apply their approach, we need to decouple the optimization problem into a separate optimization problem for each facility, which we do using the optimal dual variableλ\\lambda\. However, this eliminates cross\-learning between facilities, which may explain the poor results\. Finally, we also demonstrated a 5% improvement compared to a purely decision\-blind approach; while this improvement is comparatively smaller, a 5% reduction in unmet demand still has significant implications for social welfare\. Furthermore, it demonstrates the potential for decision\-aware learning to improve performance even compared to powerful state\-of\-the\-art models such as multi\-task random forests\.

Table SI 1:Average % Improvement in Unmet Demand\.We compare our framework vs\. various baselines on an out\-of\-sample historical dataset\.MethodImprovement %Our Framework0%Decision\-Blind Ablation5%Population Based Census27%Distribution Modeling82%Global Health \(3 Month Rolling Avg\)88%StochOptForest92%Existing Excel Tool98%

## 2Evaluation and Deployment

In this section, we describe how we deployed our system \(§[2\.1](https://arxiv.org/html/2607.20542#S2.SS1)\) and econometrically evaluated its impact using SynthDiD \(§[2\.2](https://arxiv.org/html/2607.20542#S2.SS2)\); we established an equivalence between consumption and unmet demand to justify this analysis \(§[2\.3](https://arxiv.org/html/2607.20542#S2.SS3)\)\. Then, we investigated real\-world compliance to our allocations \(§[2\.4](https://arxiv.org/html/2607.20542#S2.SS4)\), analyzed the distributional impacts of our deployment \(§[2\.5](https://arxiv.org/html/2607.20542#S2.SS5)\), and performed multiple robustness checks to validate our main findings \(§[2\.6](https://arxiv.org/html/2607.20542#S2.SS6)\)\. Finally, we performed a cost\-effectiveness analysis \(§[2\.7](https://arxiv.org/html/2607.20542#S2.SS7)\)\.

### 2\.1Deployment Details

The Sierra Leone national government deployed our system in 2023 Q2 for 5 randomly selected districts: Tonkolili, Falaba, Karene, Kono, and Pujehun; subsequently, it was rolled out nationwide in 2023 Q3\. We present a balance table for pre\-treatment covariates in Table[SI 12](https://arxiv.org/html/2607.20542#S2.T12);666This table is based on district\-level covariates from the Global Data Lab \(?\) and the Service Delivery Indicators Health Survey \(?\)\. We constructed the corresponding treatment and control values by averaging over districts in each respective group, weighted by their population\. One issue is that the districts in both data sources predated the 2017 creation of two new districts \(Koinadugu split in two into Falaba and Koinadugu; Bombali split in two into Karene and Bombali\); Falaba and Karene are treated districts whereas Koinadugu and Bombali are control districts\. We took the covariates for Falaba and Karene to be the same as the ones for Koinadugu and Bombali, respectively\.as can be seen, there are no statistically detectable differences between the two groups\. Furthermore, we present data missingness rates for treated and control districts in Table[SI 13](https://arxiv.org/html/2607.20542#S2.T13)\. Prior to the deployment, the government had established predetermined supply levels for this period and allocated resources to control districts\. The remaining supply was then assigned to the treatment districts, maintaining independence of supply quantities between the two groups\.

Before the implementation, we first conducted two training sessions for policymakers and frontline workers to provide them with a technical understanding of how our tool operated and what it did\. We ensured that our allocation tool was compatible with the same formatted inputs and outputs as the prior Excel allocation tool, allowing users to maintain their existing workflows with minimal changes\. Fig\.[SI 2](https://arxiv.org/html/2607.20542#S2.F2)shows a screenshot of our web interface\.

![Refer to caption](https://arxiv.org/html/2607.20542v1/figures/WebApp.png)Figure SI 2:Our System’s Web App Interface: Consistent with their prior workflow, users upload an Excel sheet with total available central stock information \(which is directly downloaded from the mSupply database\) and then run the tool to obtain downloadable allocation results\. This web app is now owned and operated by the Sierra Leone national government\.The deployment began in June 2023\. The implementation timeline proceeded as follows: in the Falaba and Tonkolili districts, last\-mile delivery to local health facilities was completed by mid\-July, while in the Karene, Kono, and Pujehun districts, implementation extended to the end of July due to logistical delays caused by the presidential election in Sierra Leone\. To evaluate the impact of the intervention, we primarily analyzed outcomes from 2023 Q2 to Q3\. The government did not conduct an allocation in Q4 due to insufficient supply, arising from a central\-level shortage of health aid from donors\.

Our system functioned as a decision support tool, where central and district\-level planners retain ultimate authority and can override the system when necessary\. During deployment, when planners identified discrepancies between our recommendations and their practices \(e\.g\., due to data lags, or a strategic need to maintain a larger safety stock for certain items that quarter\), they were able to easily adjust the allocations\. This human\-in\-the\-loop approach was crucial for building trust and enabling context\-specific adaptability\. We found that decision\-makers closely followed the algorithmic allocations—the normalized overlap between the actual and algorithmic allocations ranged from 0\.89 to 1 \(see §[2\.4](https://arxiv.org/html/2607.20542#S2.SS4)for details\)\. The system automatically re\-trains the machine learning model based on updated data for each new quarterly allocation, so that it continuously adapts to shifts in demand patterns\. The system is hosted on AWS with government\-controlled access, ensuring institutional ownership\. We provided knowledge transfer to local personnel to build internal capacity for system maintenance\.

### 2\.2Main Analysis Methodology

While the raw data in Fig\.[3\(a\)](https://arxiv.org/html/2607.20542#S5.F3.sf1)shows a descriptive increase in consumption, the small number of treated districts necessitates an evaluation strategy that appropriately matches pre\-treatment patterns across treated and control groups\. To this end, we used SynthDiD \(?\) to analyze the impact of our deployment on patient consumption\. SynthDiD identifies causal effects by ensuring that the difference between treated and synthetic control facilities remains stable before treatment\. To achieve this goal, SynthDiD assigns two sets of weights: one for control facilities and another for time periods\. The control facility weights are chosen so the weighted average of control facilities’ outcomes best match the unweighted average of treated facilities’ outcomes in the pre\-treatment period\. The time period weights ensure that the weighted average of the control facilities’ outcomes in the pre\-treatment period best match their unweighted average of the control facilities’ outcomes in the post\-treatment period\.

By combining these weights, SynthDiD constructs synthetic control facilities whose pre\-treatment trends align with those of the treated facilities, thereby providing a credible estimate of the treatment’s causal impact and addressing key limitations in other commonly used estimators\. Unlike standard difference\-in\-differences \(DiD\), which requires a parallel trends assumption, SynthDiD remains robust even when treatment and control groups show different trends before the intervention\. Furthermore, it can effectively control for variations in outcomes that arise from both time\-related and unit\-specific factors, unlike standard synthetic controls, which requires a near\-perfect match in pre\-treatment levels\.

Using unit weightsω^nsdid\\hat\{\\omega\}\_\{n\}^\{\\text\{sdid\}\}and time weightsλ^tsdid\\hat\{\\lambda\}\_\{t\}^\{\\text\{sdid\}\}derived from Eqs\. \(4\) & \(6\) in \(?\), the average effect of the treatment on the treated \(ATT\) is estimated as follows:

\(τ^sdid,μ^,α^,β^\)=argminτ,μ,α,β​∑n=1N∑t=1T\(Yn,t−μ−αn−βt−Wn,t​τ\)2​ω^nsdid​λ^tsdid\\left\(\\hat\{\\tau\}^\{\\text\{sdid\}\},\\hat\{\\mu\},\\hat\{\\alpha\},\\hat\{\\beta\}\\right\)=\\underset\{\\tau,\\mu,\\alpha,\\beta\}\{\\operatorname\{argmin\}\}\\;\\sum\_\{n=1\}^\{N\}\\sum\_\{t=1\}^\{T\}\\left\(Y\_\{n,t\}\-\\mu\-\\alpha\_\{n\}\-\\beta\_\{t\}\-W\_\{n,t\}\\tau\\right\)^\{2\}\\hat\{\\omega\}\_\{n\}^\{\\text\{sdid\}\}\\hat\{\\lambda\}\_\{t\}^\{\\text\{sdid\}\}\(8\)where the outcomeYn,tY\_\{n,t\}is the average consumption for facilitynnat timettacross all allocated products;μ\\muis the baseline average outcome,αn\\alpha\_\{n\}are facility fixed effects, andβt\\beta\_\{t\}are time period fixed effects\. The treatment assignmentWn,tW\_\{n,t\}indicates whether facilitynnreceived treatment at timett, andτ\\tauis the treatment effect\. Control facility weightsω^nsdid\\hat\{\\omega\}\_\{n\}^\{\\text\{sdid\}\}and time period weightsλ^tsdid\\hat\{\\lambda\}\_\{t\}^\{\\text\{sdid\}\}account for differences between treated and control groups over time\.

We independently performed our analysis using both thesynthdidpackage in R and thesdidpackage in STATA, finding consistent results\. We estimated standard errors using the jackknife \(Algorithm 3 in \(?\)\)\. To validate our use of SynthDiD, we also performed an event study \(?\), which showed that there are no statistically significant differences between treated facilities and control facilities prior to our intervention, and that the change in consumption emerged only after our system was deployed \(see Fig\.[SI 3](https://arxiv.org/html/2607.20542#S2.F3)\)\. We used Eq\. \(8\) in \(?\), which compares the treated\-minus\-synthetic\-control difference in each time periodttto a baseline pre\-treatment difference:

\(Y¯tT​r−Y¯tC​o\)−\(Y¯baselineT​r−Y¯baselineC​o\),\\displaystyle\(\\bar\{Y\}\_\{t\}^\{Tr\}\-\\bar\{Y\}\_\{t\}^\{Co\}\)\-\(\\bar\{Y\}\_\{\\text\{baseline\}\}^\{Tr\}\-\\bar\{Y\}\_\{\\text\{baseline\}\}^\{Co\}\),whereY¯baselineT​r\\bar\{Y\}\_\{\\text\{baseline\}\}^\{Tr\}andY¯baselineC​o\\bar\{Y\}\_\{\\text\{baseline\}\}^\{Co\}are the baseline means for the treated group and the synthetic control group, respectively\. Unlike conventional event studies that choose a single pre\-treatment period as the baseline, the SynthDiD framework selects the optimal pre\-treatment weightsλ^ts​d​i​d\\hat\{\\lambda\}\_\{t\}^\{sdid\}

Y¯baselineT​r=∑t=1Tpreλ^ts​d​i​d​Y¯tT​randY¯baselineC​o=∑t=1Tpreλ^ts​d​i​d​Y¯tC​o\.\\displaystyle\\bar\{Y\}\_\{\\text\{baseline\}\}^\{Tr\}=\\sum\_\{t=1\}^\{T\_\{\\text\{pre\}\}\}\\hat\{\\lambda\}\_\{t\}^\{sdid\}\\bar\{Y\}\_\{t\}^\{Tr\}\\quad\\text\{and\}\\quad\\bar\{Y\}\_\{\\text\{baseline\}\}^\{Co\}=\\sum\_\{t=1\}^\{T\_\{\\text\{pre\}\}\}\\hat\{\\lambda\}\_\{t\}^\{sdid\}\\bar\{Y\}\_\{t\}^\{Co\}\.Then, we constructed the event\-study estimates from SynthDiD using the following steps:

1. 1\.Initial Estimation: 1. \(a\)Fit SynthDiD on the full sample to obtain time period weightsλ^\\hat\{\\lambda\}and the pre\-treatment difference in outcomes\. 2. \(b\)Adjust post\-treatment differences by subtracting the pre\-treatment mean difference\.
2. 2\.Bootstrap Inference: 1. \(a\)Forb∈\{1,…,B\}b\\in\\\{1,\\ldots,B\\\}, resample the data and re\-estimate SynthDiD\. 2. \(b\)Compute the bootstrap\-adjusted difference series for each replication\. 3. \(c\)Use these bootstrap replicates to compute standard errors and form confidence intervals\.

![Refer to caption](https://arxiv.org/html/2607.20542v1/x5.png)Figure SI 3:SynthDiD Event Study\.This plot shows the estimated ATT across time\. Blue diamonds represent point estimates, and the shaded region represents 95% confidence intervals\. The dashed vertical line indicates the time period immediately prior to our 2023 Q2 deployment\. As expected, the pre\-treatment estimates are close to zero, while the post\-treatment estimates are positive and large, indicating that our deployment increased consumption\.
### 2\.3Equivalence of Consumption and Unmet Demand

One of the key challenges in evaluating our deployment is that unmet demand—the usual metric for evaluating the efficiency of inventory management policies—is censored since we do not observe when patient orders go unfulfilled\. Instead, we use consumption as our main metric, which measures the total amount of each product distributed to patients across all facilities\. First, we show that under a simple model of dispensing, maximizing total consumption and minimizing total unmet demand are mathematically equivalent in our optimization formulation\. Next, we show that this equivalence is robust to stock\-dependent dispensing \(such as rationing or over\-dispensing\) and product substitution\.

##### Basic Model\.

In each quartertt, each facilityn∈\[N\]n\\in\[N\]has current stocksm,n,t∈ℝ≥0s\_\{m,n,t\}\\in\\mathbb\{R\}\_\{\\geq 0\}of productm∈\[M\]m\\in\[M\]and needs to fulfill demandξm,n,t∈ℝ≥0\\xi\_\{m,n,t\}\\in\\mathbb\{R\}\_\{\\geq 0\}until it runs out of stock\. Our goal is to allocate medicines to each facility to minimize unmet demand\. The unmet demand and consumption for productmmand facilitynnin time periodttas a function of the allocationa∈ℝ≥0a\\in\\mathbb\{R\}\_\{\\geq 0\}of that product to that facility in that time period are

Um,n,t​\(a\)\\displaystyle U\_\{m,n,t\}\(a\)=max⁡\{ξm,n,t−\(sm,n,t\+a\),0\}\\displaystyle=\\max\\\{\\xi\_\{m,n,t\}\-\(s\_\{m,n,t\}\+a\),0\\\}Cm,n,t​\(a\)\\displaystyle C\_\{m,n,t\}\(a\)=min⁡\{ξm,n,t,sm,n,t\+a\},\\displaystyle=\\min\\\{\\xi\_\{m,n,t\},s\_\{m,n,t\}\+a\\\},respectively\. It is easy to check that

Um,n,t​\(a\)\+Cm,n,t​\(a\)=ξm,n,t\\displaystyle U\_\{m,n,t\}\(a\)\+C\_\{m,n,t\}\(a\)=\\xi\_\{m,n,t\}for any allocationaa\. Given allocation decisions\{an\}n\\\{a\_\{n\}\\\}\_\{n\}for each facilityn∈\[N\]n\\in\[N\], we have

Um,t​\(\{an\}n\)\+Cm,t​\(\{an\}n\)=ξm,t,\\displaystyle U\_\{m,t\}\(\\\{a\_\{n\}\\\}\_\{n\}\)\+C\_\{m,t\}\(\\\{a\_\{n\}\\\}\_\{n\}\)=\\xi\_\{m,t\},whereUm,t​\(\{an\}n\)=∑n=1NUm,n,t​\(an\)U\_\{m,t\}\(\\\{a\_\{n\}\\\}\_\{n\}\)=\\sum\_\{n=1\}^\{N\}U\_\{m,n,t\}\(a\_\{n\}\)is total unmet demand,Cm,t​\(\{an\}n\)=∑n=1NCm,n,t​\(an\)C\_\{m,t\}\(\\\{a\_\{n\}\\\}\_\{n\}\)=\\sum\_\{n=1\}^\{N\}C\_\{m,n,t\}\(a\_\{n\}\)is total consumption, andξm,t=∑n=1Nξm,n,t\\xi\_\{m,t\}=\\sum\_\{n=1\}^\{N\}\\xi\_\{m,n,t\}is the total demand for quartertt\. Thus, total unmet demand plus total consumption does not depend on our allocation decisions\{an\}n\\\{a\_\{n\}\\\}\_\{n\}; therefore, the allocation that minimizes unmet demand also maximizes consumption\.

Of course, this model makes a number of strong assumptions about how stock is distributed\. We describe how to handle alternative behaviors: \(1\) rationing \(i\.e\., when the facility is low on stock, it allocates less to patients than their need\), \(2\) over\-dispensing \(i\.e\., when the facility has excess stock, it allocates more to patients than they need\), and \(3\) substitution \(i\.e\., when stock of one medication runs out, the facility allocates a substitute\)\. We discuss rationing and over\-dispensing in §[2\.3\.1](https://arxiv.org/html/2607.20542#S2.SS3.SSS1), and substitution in §[2\.3\.2](https://arxiv.org/html/2607.20542#S2.SS3.SSS2)\.

#### 2\.3\.1Rationing and Over\-Dispensing\.

Providers may change their dispensing patterns dynamically as a function of the available stock remaining\. For instance, if product stock is running low and the next allocation isn’t anticipated for some time, a provider may ration stock by only allocating to high\-risk patients; alternatively, if product stock is relatively high and the next allocation is anticipated soon, a provider may dispense to even low\-risk patients who are not normally allocated\. We show using a Markov Decision Process analysis that any such stock\-dependent allocations still satisfy the equivalence of consumption and demand, as long as providers do notwastestock \(i\.e\., dispense to a patient with no demand, or dispense more than they could plausibly use\)\. Waste is unlikely to occur in our setting due to the limited supply\. We also performed an empirical analysis to try to detect rationing and over\-dispensing behavior in our data; we found no evidence for such behaviors\.

##### Theoretical analysis\.

For this section, we redefine the variablestt,sts\_\{t\},ata\_\{t\}, andξt\\xi\_\{t\}to be the index of the current patient, the current stock of a medicine at a given facility when patientttarrives, the amount of medicine distributed to patienttt, and the target demand of patienttt, respectively\. Specifically, we consider a model where \(1\) patients arrive at the facility sequentially overTTsteps \(indexed byt∈\[T\]t\\in\[T\]\), are distributedata\_\{t\}of the product and then leave; \(2\) each patient has a target demandξt\\xi\_\{t\}depending on their typext∈\[K\]x\_\{t\}\\in\[K\], and the facility distribution satisfiesat≤ξta\_\{t\}\\leq\\xi\_\{t\}; and \(3\) unmet demand for that patient isξt−at\\xi\_\{t\}\-a\_\{t\}\. Under this assumption, we prove that unmet demand and negative consumption are equivalent even in this general setting:

U¯​\(π\)\+C¯​\(π\)=const,\\displaystyle\\bar\{U\}\(\\pi\)\+\\bar\{C\}\(\\pi\)=\\text\{const\},whereU¯​\(π\)\\bar\{U\}\(\\pi\)andC¯​\(π\)\\bar\{C\}\(\\pi\)are the expected unmet demand and consumption when using policyπ\\pi, respectively\. This result establishes under general conditions that consumption is a reasonable proxy for unmet demand\.

We model the system as a finite\-horizon MDP\(S,A,R,P,T\)\(S,A,R,P,T\), where the state is a pair\(s,x\)∈S=\(ℕ∪\{0\}\)×\[K\]\(s,x\)\\in S=\(\\mathbb\{N\}\\cup\\\{0\\\}\)\\times\[K\]consisting of the current stockssand the typexxof the patient being served on steptt, the actiona∈A=ℕa\\in A=\\mathbb\{N\}indicates the amount of medication to dispense to the patient, the rewardR​\(s,x,a\)≥0R\(s,x,a\)\\geq 0indicates the payoff for dispensing medicine to the patient \(withR​\(s,x,0\)=0R\(s,x,0\)=0\), the transition measureP​\(s,x,s′,x′\)=𝟙​\(s′=s−a\)⋅P​\(x′\)P\(s,x,s^\{\\prime\},x^\{\\prime\}\)=\\mathbbm\{1\}\(s^\{\\prime\}=s\-a\)\\cdot P\(x^\{\\prime\}\)consists of a deterministic updates′=s−as^\{\\prime\}=s\-ato the stock and an i\.i\.d\. sample of the next patient typex′x^\{\\prime\}, and the time horizon isT∈ℕT\\in\\mathbb\{N\}\. We assume that patient typex=1x=1indicates no patient, andR​\(s,1,a\)=0R\(s,1,a\)=0for alla∈Aa\\in A\. We impose the constraint thata≤sa\\leq s\. We letξ​\(x\)∈ℕ\\xi\(x\)\\in\\mathbb\{N\}be the ideal amount of medication needed by patient of typexx\.

We assume the facility is using a state\-dependent and time\-dependent policyπt:S→A\\pi\_\{t\}:S\\to Ato distribute medicine to patients\. Since the state includes the current stock and the time encodes the time remaining until the end of the quarter, this policy can encode very complex rationing and over\-dispensing strategies\. As usual, a given policyπ\\piinduces a distributionD\(π\)D^\{\(\\pi\)\}over rolloutsζ=\(\(s1,x1,a1,r1\),…,\(sT,xT,aT,rT\)\)\\zeta=\(\(s\_\{1\},x\_\{1\},a\_\{1\},r\_\{1\}\),\.\.\.,\(s\_\{T\},x\_\{T\},a\_\{T\},r\_\{T\}\)\)\. We letU¯​\(π\)=𝔼ζ∼D\(π\)​\[U​\(ζ\)\]\\bar\{U\}\(\\pi\)=\\mathbb\{E\}\_\{\\zeta\\sim D^\{\(\\pi\)\}\}\[U\(\\zeta\)\]andC¯​\(π\)=𝔼ζ∼D\(π\)​\[C​\(ζ\)\]\\bar\{C\}\(\\pi\)=\\mathbb\{E\}\_\{\\zeta\\sim D^\{\(\\pi\)\}\}\[C\(\\zeta\)\]denote the expected unmet demand and consumption, respectively\. The following result shows that under this general model, unmet demand and negative consumption are equivalent metrics\.

###### Theorem 1\.

Assumeat≤ξta\_\{t\}\\leq\\xi\_\{t\}\(whereξt=ξ​\(xt\)\\xi\_\{t\}=\\xi\(x\_\{t\}\)\) for allt∈\[T\]t\\in\[T\]\. Then, we haveU¯​\(π\)=Ξ¯−C¯​\(π\)\\bar\{U\}\(\\pi\)=\\bar\{\\Xi\}\-\\bar\{C\}\(\\pi\), whereΞ¯\\bar\{\\Xi\}is a constant independent ofπ\\pi\.

###### Proof\.

Given an arbitrary rolloutζ=\(\(s1,x1,a1,r1\),…,\(sT,xT,aT,rT\)\)\\zeta=\(\(s\_\{1\},x\_\{1\},a\_\{1\},r\_\{1\}\),\.\.\.,\(s\_\{T\},x\_\{T\},a\_\{T\},r\_\{T\}\)\), the total unmet demand acrossζ\\zetais

U​\(ζ\)=∑t=1Tmax⁡\{ξt−at,0\}⋅𝟙​\(xt≠1\),\\displaystyle U\(\\zeta\)=\\sum\_\{t=1\}^\{T\}\\max\\\{\\xi\_\{t\}\-a\_\{t\},0\\\}\\cdot\\mathbbm\{1\}\(x\_\{t\}\\neq 1\),whereξt=ξ​\(xt\)\\xi\_\{t\}=\\xi\(x\_\{t\}\), and consumption is

C​\(ζ\)=∑t=1Tmin⁡\{at,ξt\}⋅𝟙​\(xt≠1\)\.\\displaystyle C\(\\zeta\)=\\sum\_\{t=1\}^\{T\}\\min\\\{a\_\{t\},\\xi\_\{t\}\\\}\\cdot\\mathbbm\{1\}\(x\_\{t\}\\neq 1\)\.Under our assumption thatat≤ξta\_\{t\}\\leq\\xi\_\{t\}, we have

U​\(ζ\)=∑t=1T\(ξt−at\)⋅𝟙​\(xt≠1\)=Ξ​\(ζ\)−∑t=1Tat⋅𝟙​\(xt≠1\),\\displaystyle U\(\\zeta\)=\\sum\_\{t=1\}^\{T\}\(\\xi\_\{t\}\-a\_\{t\}\)\\cdot\\mathbbm\{1\}\(x\_\{t\}\\neq 1\)=\\Xi\(\\zeta\)\-\\sum\_\{t=1\}^\{T\}a\_\{t\}\\cdot\\mathbbm\{1\}\(x\_\{t\}\\neq 1\),and

C​\(ζ\)=∑t=1Tat⋅𝟙​\(xt≠1\),\\displaystyle C\(\\zeta\)=\\sum\_\{t=1\}^\{T\}a\_\{t\}\\cdot\\mathbbm\{1\}\(x\_\{t\}\\neq 1\),whereΞ​\(ζ\)=∑t=1Tξt\\Xi\(\\zeta\)=\\sum\_\{t=1\}^\{T\}\\xi\_\{t\}; it follows that

U​\(ζ\)=Ξ​\(ζ\)−C​\(ζ\)\.\\displaystyle U\(\\zeta\)=\\Xi\(\\zeta\)\-C\(\\zeta\)\.Taking the expectation overζ∼D\(π\)\\zeta\\sim D^\{\(\\pi\)\}, whereD\(π\)D^\{\(\\pi\)\}is the distribution over rollouts induced by using policyπ\\pi, we have

U¯​\(π\)=Ξ¯−C¯​\(π\),\\displaystyle\\bar\{U\}\(\\pi\)=\\bar\{\\Xi\}\-\\bar\{C\}\(\\pi\),whereΞ¯=Ξ¯​\(π\)=𝔼ζ∼D\(π\)​\[Ξ​\(ζ\)\]\\bar\{\\Xi\}=\\bar\{\\Xi\}\(\\pi\)=\\mathbb\{E\}\_\{\\zeta\\sim D^\{\(\\pi\)\}\}\[\\Xi\(\\zeta\)\]is the expected total demand; note thatΞ​\(ζ\)\\Xi\(\\zeta\)only depends onζ\\zetathrough the sequence of patient covariates\(x1,…,xT\)\(x\_\{1\},\.\.\.,x\_\{T\}\), soΞ¯​\(π\)\\bar\{\\Xi\}\(\\pi\)is independent ofπ\\pi\. ∎

##### Empirical analysis\.

Next, we performed empirical analyses to try to detect rationing and over\-dispensing behaviors\. Ideally, we would have patient\-level data to understand if providers are allocating differently to different risk profiles as a function of available stock; however, such data does not exist in a digitized, unified form in many developing countries due to infrastructure limitations\. As a result, we are limited to analyzing aggregate monthly allocations\.

First, we considered rationing—facilities with low inventory may preemptively reduce consumption \(e\.g\., by allocating only to high\-risk patients\) even if they are sufficiently well stocked, in order to avoid a stockout in future months before the next allocation\. Consider a facility in montht−2t\-2\(with an upcoming allocation in monthtt\) that has sufficient stock for the current montht−2t\-2but insufficient stock for the coming montht−1t\-1\. If we see “less than expected” consumption int−2t\-2, that may indicate rationing behavior\.

To test this, we first identified facility\-month pairs\(n,τ\)\(n,\\tau\)in our panel data where the opening balance is sufficiently large to likely avoid a stockout in that month\.777If not, reduced consumption in that month would simply reflect a stockout rather than rationing behavior to protect stock for future months\.Specifically, we only considered\(n,τ\)\(n,\\tau\)pairs such that the initial stocksn,τs\_\{n,\\tau\}exceeded the average consumptionY¯n\\bar\{Y\}\_\{n\}by some multiplicative constant, i\.e\.,sn,τ\>C2​Y¯ns\_\{n,\\tau\}\>C\_\{2\}\\bar\{Y\}\_\{n\}for someC2≥1C\_\{2\}\\geq 1\. Then, we estimated the following regression model:

Yn,τ\\displaystyle Y\_\{n,\\tau\}=β1​LowStockn,τ\+β2​TimeIndicatorτ\+δ​\(LowStockn,τ×TimeIndicatorτ\)\\displaystyle=\\beta\_\{1\}\\text\{LowStock\}\_\{n,\\tau\}\+\\beta\_\{2\}\\text\{TimeIndicator\}\_\{\\tau\}\+\\delta\(\\text\{LowStock\}\_\{n,\\tau\}\\times\\text\{TimeIndicator\}\_\{\\tau\}\)\+αn\+λτ\+γ​Xn,τ\+ϵn,τ\\displaystyle\\qquad\+\\alpha\_\{n\}\+\\lambda\_\{\\tau\}\+\\gamma X\_\{n,\\tau\}\+\\epsilon\_\{n,\\tau\}\(9\)whereYn,τY\_\{n,\\tau\}is the consumption of facilitynnin monthτ\\tau;αn\\alpha\_\{n\}andλτ\\lambda\_\{\\tau\}are facility and time fixed effects, respectively;Xn,τX\_\{n,\\tau\}denotes controls \(including all the features in our demand prediction model\);LowStockn,τ=𝟙​\(sn,τ≤C1​Y¯n\)\\text\{LowStock\}\_\{n,\\tau\}=\\mathbbm\{1\}\(s\_\{n,\\tau\}\\leq C\_\{1\}\\bar\{Y\}\_\{n\}\)for someC1<2C\_\{1\}<2indicates whether facilitynnhas likely insufficient opening balance in monthτ\\tauto satisfy future demand in monthτ\+1\\tau\+1, as estimated by a multiplier of its average demandY¯n\\bar\{Y\}\_\{n\}; andTimeIndicatorτ=𝟙​\(∃t∈𝒯alloc​s\.t\.​τ=t−2\)\\text\{TimeIndicator\}\_\{\\tau\}=\\mathbbm\{1\}\(\\exists t\\in\\mathcal\{T\}\_\{\\text\{alloc\}\}\\text\{ s\.t\. \}\\tau=t\-2\)indicates whetherτ\\tauis two months prior to a planned allocationt∈𝒯alloct\\in\\mathcal\{T\}\_\{\\text\{alloc\}\}\. The coefficient of interest isδ\\delta; a negative value ofδ\\deltawould indicate that facilities reduce distribution in months where they have low stock in montht−2t\-2to save for montht−1t\-1, wherettis an allocation month\.

For robustness, we considered several choices of our constants—specifically,C1∈\{1\.5,1\.75\}C\_\{1\}\\in\\\{1\.5,1\.75\\\}andC2∈\{1\.0,1\.5\}C\_\{2\}\\in\\\{1\.0,1\.5\\\}; note that we requireC1\>C2C\_\{1\}\>C\_\{2\}\. Results from these regressions are shown in Table[SI 2](https://arxiv.org/html/2607.20542#S2.T2)\. As can be seen,δ\\deltais not statistically significantly different from zero in any of them, so we found no evidence of rationing behaviors\. We note that this empirical analysis is only suggestive, as we do not have the data to support a causal or more granular analysis\.

Table SI 2:Analysis of Rationing Behavior\.Condition on Opening BalanceC2C\_\{2\}Initial Supply LevelC1C\_\{1\}C2=1\.0C\_\{2\}=1\.0C2=1\.5C\_\{2\}=1\.5C1=1\.5C\_\{1\}=1\.50\.986–\(0\.592\)C1=1\.75C\_\{1\}=1\.750\.360−\-0\.822\(0\.496\)\(0\.810\)
- •Notes:The table reports regression coefficients of the interaction termδ\\delta\. Standard errors are in parentheses:p∗<0\.05\{\}^\{\*\}p<0\.05,p∗∗<0\.01\{\}^\{\*\*\}p<0\.01\.

Next, we considered over\-dispensing—facilities with high inventory may preemptively increase consumption \(e\.g\., by allocating to low\-risk patients\) in anticipation of an upcoming allocation\. Consider a facility in montht−1t\-1\(with an upcoming allocation in monthtt\) that has more than sufficient stock for the current month\. If we see “more than expected” consumption int−1t\-1, that may indicate a facility trying to over\-dispense a surplus\.

To test this, as before, we identified facility\-month pairs in our panel data where the opening balance is sufficiently large to avoid a stockout in that month, i\.e\.,\(n,τ\)\(n,\\tau\)pairs such thatsn,τ\>C2​Y¯ns\_\{n,\\tau\}\>C\_\{2\}\\bar\{Y\}\_\{n\}\.888If not, increased consumption in that month would simply reflect the lack of a stockout rather than over\-dispensing behavior to use up surplus stock\.Then, we estimated the following regression model:

Yn,τ\\displaystyle Y\_\{n,\\tau\}=β1​ExcessStockn,τ\+β2​TimeIndicatorτ\+δ​\(ExcessStockn,τ×TimeIndicatorτ\)\\displaystyle=\\beta\_\{1\}\\text\{ExcessStock\}\_\{n,\\tau\}\+\\beta\_\{2\}\\text\{TimeIndicator\}\_\{\\tau\}\+\\delta\(\\text\{ExcessStock\}\_\{n,\\tau\}\\times\\text\{TimeIndicator\}\_\{\\tau\}\)\+αn\+λτ\+γ​Xn,τ\+ϵn,τ\\displaystyle\\qquad\+\\alpha\_\{n\}\+\\lambda\_\{\\tau\}\+\\gamma X\_\{n,\\tau\}\+\\epsilon\_\{n,\\tau\}\(10\)whereExcessStockn,τ=𝟙​\(sn,τ≥C1​Y¯n\)\\text\{ExcessStock\}\_\{n,\\tau\}=\\mathbbm\{1\}\(s\_\{n,\\tau\}\\geq C\_\{1\}\\bar\{Y\}\_\{n\}\)for someC1\>2C\_\{1\}\>2indicates whether facilitynnwill likely have a significant surplus in monthτ\+1\\tau\+1\(given the upcoming allocation\), as estimated by a multiplier of its average demandY¯n\\bar\{Y\}\_\{n\}; we re\-definedTimeIndicatorτ=𝟙​\(∃t∈𝒯alloc​s\.t\.​τ=t−1\)\\text\{TimeIndicator\}\_\{\\tau\}=\\mathbbm\{1\}\(\\exists t\\in\\mathcal\{T\}\_\{\\text\{alloc\}\}\\text\{ s\.t\. \}\\tau=t\-1\)to indicate whetherτ\\tauis one month prior to a planned allocationt∈𝒯alloct\\in\\mathcal\{T\}\_\{\\text\{alloc\}\}\.

For robustness, we consideredC1∈\{3\.0,5\.0\}C\_\{1\}\\in\\\{3\.0,5\.0\\\}andC2∈\{1\.0,1\.5\}C\_\{2\}\\in\\\{1\.0,1\.5\\\}\. Results from the resulting regressions are shown in Table[SI 3](https://arxiv.org/html/2607.20542#S2.T3)\. As can be seen,δ\\deltais not statistically significantly different from zero in any of these regressions, so we found no evidence of over\-dispensing behaviors\. We note that this empirical analysis suffers from the same limitations noted earlier for rationing\.

Table SI 3:Analysis of Over\-Dispensing Behavior\.Condition on Opening BalanceC2C\_\{2\}Initial Supply LevelC1C\_\{1\}C2=1\.0C\_\{2\}=1\.0C2=1\.5C\_\{2\}=1\.5C1=3\.0C\_\{1\}=3\.00\.1740\.146\(0\.371\)\(0\.408\)C1=5\.0C\_\{1\}=5\.00\.4890\.503\(0\.368\)\(0\.380\)
- •Notes:The table reports regression coefficients of the interaction term\. Standard errors are in parentheses:p∗<0\.05\{\}^\{\*\}p<0\.05,p∗∗<0\.01\{\}^\{\*\*\}p<0\.01\.

One concern could be that our estimates are under\-powered\. Thus, as a robustness check, we took the point estimates for rationing and over\-dispensing at face value, and evaluated their impact on our main analysis\. Specifically, we took a conservative approach \(i\.e\., assume these behaviors only adversely affect consumption in treated facilities and do not affect control facilities\), which translated into an approximately 2 percentage point reduction in consumption in treated facilities\. Correspondingly, we adjusted our evaluation by reducing only the treated facilities’ consumption in the post\-treatment period by 2 percentage points\. We then re\-ran our main analysis and found that our deployment still improved consumption by a statistically significant 16%\. This improvement is quite similar to the 19% improvement found in our original analysis\.

#### 2\.3\.2Substitution

Another potential behavior we do not handle in our basic model is cross\-product substitution\. We show theoretically that consumption and unmet demand still remain equivalent under “perfect” product substitution, as long as we evaluate consumption at the level of product clusters \(that mutually substitute for each other\) with appropriate normalization, instead of evaluating consumption at the level of individual products\. We first formalize and prove this result\. Based on input from medical experts and field workers, we developed a list of all substitutable products among the ones we allocated; many of these can be considered perfect \(e\.g\., medications with different dosages\), for the remaining, perfect substitution was a reasonably good approximation\. Then, we performed a robustness check where we re\-ran our main analysis using the above strategy and obtained similar results as our main analysis\.

##### Theoretical analysis\.

We consider substitution under the following assumptions: \(1\) substitutions are not case\-dependent \(i\.e\., if two medicines can either always be substituted, or never be substituted\), \(2\) substitutions are transitive \(i\.e\., ifm1m\_\{1\}substitutes form2m\_\{2\}andm2m\_\{2\}substitutes form3m\_\{3\}, thenm1m\_\{1\}substitutes form3m\_\{3\}\), and \(3\) substitutions are always at the same rate \(i\.e\., ifm1m\_\{1\}andm2m\_\{2\}are substitutes, then if a patient requiresaaunits ofm1m\_\{1\}, then they requirecm1→m2⋅ac\_\{m\_\{1\}\\to m\_\{2\}\}\\cdot aunits ofm2m\_\{2\}, wherecm1→m2c\_\{m\_\{1\}\\to m\_\{2\}\}is constant across patients \(withcm→m=1c\_\{m\\to m\}=1\)\)\. While these assumptions are somewhat strong, they are accurate to first order for the products distributed by our system\.

With these assumptions, we can straightforwardly adapt our analysis to handle substitutions\. Specifically, rather than consider products separately, we first cluster products intoKKgroupsMk⊆\[M\]M\_\{k\}\\subseteq\[M\]that are mutually substitutable \(which is possible by Assumptions 1 & 2\)\. For each groupk∈\[K\]k\\in\[K\], we can choose one focal productm∈Mkm\\in M\_\{k\}in that group and standardize the other products relative tommto obtain the following aggregated stock, allocation, and demand:

s~k,n,t\\displaystyle\\tilde\{s\}\_\{k,n,t\}=∑m′∈Mkcm′→m⋅sm′,n,t\\displaystyle=\\sum\_\{m^\{\\prime\}\\in M\_\{k\}\}c\_\{m^\{\\prime\}\\to m\}\\cdot s\_\{m^\{\\prime\},n,t\}a~k,n,t\\displaystyle\\tilde\{a\}\_\{k,n,t\}=∑m′∈Mkcm′→m⋅am′,n,t\\displaystyle=\\sum\_\{m^\{\\prime\}\\in M\_\{k\}\}c\_\{m^\{\\prime\}\\to m\}\\cdot a\_\{m^\{\\prime\},n,t\}ξ~k,n,t\\displaystyle\\tilde\{\\xi\}\_\{k,n,t\}=∑m′∈Mkcm′→m⋅ξm′,n,t\.\\displaystyle=\\sum\_\{m^\{\\prime\}\\in M\_\{k\}\}c\_\{m^\{\\prime\}\\to m\}\\cdot\\xi\_\{m^\{\\prime\},n,t\}\.To account for substitution, we can simply analyze consumption using the normalized variables instead of the original ones\.

##### Empirical analysis\.

Using the strategy outlined above, we conducted a robustness check of our main analysis based on the normalized product consumption to account for potential substitution behavior; we discuss this analysis in detail in §2\.6\.5\. We found a consistent, statistically significant 18% increase in the consumption of these normalized product groups, demonstrating that substitution is unlikely to be the primary driver behind our treatment effect\.

### 2\.4Compliance Analysis

Our system was implemented as a decision support tool that could be overridden by district pharmaceutical managers, who may selectively override our recommendations to address shortcomings in our allocation decisions to improve performance \(though they may also make mistakes that lead to worse outcomes\)\. To measure compliance, we manually collected the local pharmacist’s allocation document and cross\-checked all the invoices pulled from mSupply\. To quantify compliance of a treated district, we normalized the allocation quantity for each facility\-product pair in that district, calculated the absolute difference between the actual allocation and our deployed suggestion, and then averaged this value across facility\-product pairs\. Results are shown in Table[SI 4](https://arxiv.org/html/2607.20542#S2.T4); we found that compliance was generally very high, with the Kono and Pujehun districts having slightly lower compliance than the other three districts\. This is likely due to logistical and communication issues that arose during the implementation of our system in 2023 Q2, which were resolved soon after\. In §[2\.6\.7](https://arxiv.org/html/2607.20542#S2.SS6.SSS7), we disentangled the effect of compliance with our recommendations in 2023 Q2 from the impact of our system through a standard instrumental variables analysis to compute the Local Average Treatment Effect \(LATE\)\. We found a statistically significant 37% improvement in consumption among compliers, suggesting that compliance with our system’s allocations was generally beneficial\.

Table SI 4:Compliance of Treated Districts in 2023 Q2DistrictNormalized Avg Absolute DiffTonkolili0\.000Falaba0\.028Karene0\.039Kono0\.073Pujehun0\.109
### 2\.5Distributional Impact Analysis

To understand the distributional impacts of our allocation system, we conducted additional analyses across several dimensions\.

The first dimension focuses on facility type; given limited resources, NMSA policymakers emphasized prioritizing larger facilities \(Hospitals and Community Health Centers \(CHCs\)\) to ensure that care remains available to the greatest number of people\. We checked whether these facilities are correctly prioritized by our system\. Specifically, we categorized facilities into two groups: \(1\) larger, better equipped facilities \(Hospitals and Community Health Centers \(CHCs\)\), and \(2\) smaller, community\-based facilities \(Community Health Posts \(CHPs\) and Maternal and Child Health Posts \(MCHPs\)\)\. Then, we conducted our main analysis separately on each group; results are shown in Table[SI 5](https://arxiv.org/html/2607.20542#S2.T5)\. As can be seen, our positive treatment effect remains strong and statistically significant for larger facilities, suggesting that our algorithm effectively improves consumption at facilities that are well\-equipped and consistently operational \(consistent with NMSA’s stated priorities\)\. In contrast, smaller health posts show a smaller, statistically insignificant improvement\. They may face more fundamental challenges \(e\.g\., inadequate staffing or equipment\) that limit their ability to provide services or utilize medical supplies, regardless of availability\. For instance, during a field visit, we found a small health post closed for two consecutive days, and local residents shared that they often seek essential care at larger, more reliable facilities\. Improving allocation efficiency for larger facilities is desirable since these facilities are relied on to provide consistent care; furthermore, we did so without impacting \(or even slightly improving\) efficiency for smaller facilities\.

Next, we performed two analyses to better understand the equity implications of our system\. First, we examined the impact on previously under\-served facilities; specifically, a facility is*under\-served*if it experienced at least one stockout in the data from 2020 until the quarter before our deployment\. Our results are shown in Table[SI 5](https://arxiv.org/html/2607.20542#S2.T5); as can be seen, our results are substantially larger \(compared to our overall effect of 19%\) and statistically significant for under\-served facilities\. Second, we examined distributional impacts on urban vs\. rural facilities, based on geographic location and accessibility; in Sierra Leone, rural facilities tend to be poorer than urban ones\. We categorized facilities as urban if they were within a walkable distance of a city or town \(?\) \(either within 0\.25 miles based on \(?\), or within 1 mile as a robustness check\) and rural otherwise\. As can be seen from Table[SI 5](https://arxiv.org/html/2607.20542#S2.T5), the treatment effect is positive and statistically significant for rural facilities \(it is also positive but statistically insignificant for urban facilities, but this may be due to the smaller sample size\)\. This suggests that our allocation system succeeded at reducing unmet demand at rural \(poor\) facilities, likely improving equity\.

Table SI 5:Distributional Analysis Results\.Columns 1\-2 compare large facilities \(Hospitals and CHCs\) vs\. smaller, community\-based facilities \(CHPs and MCHPs\)\. Columns 3\-6 analyze rural vs\. urban facilities based on distance to a city \(0\.25 or 1 mile\)\. Column 7 examines under\-served facilities that have experienced at least one stockout since 2020\. We found that our deployment increased consumption in large facilities, rural facilities, as well as under\-served facilities\.Dep\. var\.: Normalized ConsumptionBy Facility TypeBy LocationUnder\-ServedLargeSmall0\.25 miles1 mileFacilityFacilityRuralUrbanRuralUrban\(1\)\(2\)\(3\)\(4\)\(5\)\(6\)\(7\)Treatment Effect0\.368∗0\.0520\.107∗∗0\.2910\.090∗0\.3340\.182∗∗\(0\.183\)\(0\.037\)\(0\.043\)\(0\.395\)\(0\.044\)\(0\.315\)\(0\.069\)Observations1,0754,2154,8654254,4308604,055Improvement %36%10% \(Insig\)19%24% \(Insig\)16%32% \(Insig\)32%

- •Note:Standard errors in parentheses\.p∗<0\.05\{\}^\{\*\}p<0\.05,p∗∗<0\.01\{\}^\{\*\*\}p<0\.01,p∗⁣∗∗<0\.001\{\}^\{\*\*\*\}p<0\.001\.

### 2\.6Robustness Checks

We performed a number of robustness checks to support our main results, summarized in Table[SI 6](https://arxiv.org/html/2607.20542#S2.T6)\. Specifically, we performed difference\-in\-differences \(§2\.6\.1\), geographic matching \(§2\.6\.2\), imputation strategies \(§2\.6\.3\), differential missingness analysis \(§2\.6\.4\), and a substitution\-robust analysis \(§2\.6\.5\)\. We also re\-ran our main analysis using alternative controls \(§2\.6\.6\), examined the treatment effect on compliers \(§2\.6\.7\), and examined stockouts \(§2\.6\.8\)\.

Table SI 6:Robustness Analysis Results\.ModelCoefficientStd\. ErrorObservationsImprovement %Dependent Variable = Normalized ConsumptionSynthDiD0\.116∗\(0\.046\)5,29019%DiD0\.128∗∗\(0\.046\)5,29021%Matching \(5km\)0\.154∗\(0\.066\)1,12525%Matching \(10km\)0\.166∗∗\(0\.064\)1,81530%Matching \(15km\)0\.121∗\(0\.054\)2,41521%Imputation \(low rank\)0\.037∗∗\(0\.014\)5,45515%Imputation \(avg consump\)0\.067∗\(0\.031\)5,45521%Imputation \(pop\)0\.076∗∗\(0\.028\)5,45527%No Missingness Imbalance0\.132∗\(0\.058\)5,28518%Substitution \(Method 2\)0\.132∗∗\{\}^\{\*\}\*\(0\.048\)5,29020%Alt\. Control0\.095∗∗∗\(0\.022\)10,52018%Substitution \(Method 1\)0\.119∗\(0\.048\)5,29018%LATE \(Continuous\)0\.115∗∗\(0\.038\)5,29037%Dependent Variable = StockoutSynthDiD−\-0\.280\(0\.179\)5,290−\-4\.6% \(Insig\)Dependent Variable = MissingnessSynthDiD−\-0\.007\(0\.006\)5,290−\-1\.4% \(Insig\)
- •Notes:p∗<0\.05\{\}^\{\*\}p<0\.05,p∗∗<0\.01\{\}^\{\*\*\}p<0\.01,p∗⁣∗∗<0\.001\{\}^\{\*\*\*\}p<0\.001\. Standard errors in parentheses\. The balanced dataset includes 1,058 facilities across five quarters from 2022 Q3 to 2023 Q3\. The Improvement % column reports the relative change in consumption compared to the counterfactual mean for each specification\.

#### 2\.6\.1Difference\-in\-Differences \(DiD\)\.

First, we ensured our results are consistent under a simpler DiD analysis \(?\), which compares changes in outcomes over time between the treated and control groups\. Fig\.[SI 4](https://arxiv.org/html/2607.20542#S2.F4)shows the corresponding event study, which shows that there are no statistically significant differences between treated and control facilities prior to our intervention\. The DiD specification is:

Yn,t=μ\+αn\+βt\+τ​\(Treatn×Postt\)\+ϵn,t,Y\_\{n,t\}=\\mu\+\\alpha\_\{n\}\+\\beta\_\{t\}\+\\tau\(\\text\{Treat\}\_\{n\}\\times\\text\{Post\}\_\{t\}\)\+\\epsilon\_\{n,t\},\(11\)whereYn,tY\_\{n,t\}is the observed outcome \(i\.e\., normalized consumption\) for unitnnat timett;μ\\mudenotes the mean outcome;αn\\alpha\_\{n\}represents unit fixed effects;βt\\beta\_\{t\}represents time period fixed effects;ϵn,t\\epsilon\_\{n,t\}is the error term;Treatn\\text\{Treat\}\_\{n\}is a dummy indicator of treatment;Postt\\text\{Post\}\_\{t\}is a dummy indicator for post\-treatment periods; andτ\\tauis the treatment effect\. We find that the magnitude and significance of the increase in consumption are similar to those obtained using SynthDiD\.

![Refer to caption](https://arxiv.org/html/2607.20542v1/x6.png)Figure SI 4:DiD Event Study\.This plot shows estimated treatment effects across time\. Blue dots represent point estimates, and vertical bars denote 95% confidence intervals\. The black vertical dashed line shows the time period immediately preceding our 2023 Q2 deployment\. As expected, the pre\-treatment estimates are close to zero, while the post\-treatment estimates are larger, indicating a positive impact on consumption attributable to our deployment\.
#### 2\.6\.2Matching analysis\.

Next, we employed a geographic matching design, which is more robust to geographically\-varying unobserved confounders by restricting the DiD analysis to matching pairs of treatment and control facilities that are geographically close together\. Specifically, we first restricted our sample to facilities within some distance of a district border \(either 5km, 10km, or 15km\) and that have a neighboring facility with the opposite treatment status; Fig\.[SI 5](https://arxiv.org/html/2607.20542#S2.F5)shows the resulting set of facilities for the 15km cutoff\. Then, we ran our DiD analysis restricted to matched facilities\. Results for all three distances are shown in Table[SI 6](https://arxiv.org/html/2607.20542#S2.T6); they show a statistically significant increase in consumption for treated facilities that is consistent with our main analysis\.

![Refer to caption](https://arxiv.org/html/2607.20542v1/figures/MatchingMap.png)Figure SI 5:Map of Matched Border Facilities\.This map displays the geographic distribution of facilities included in the border matching analysis in the “Matching \(15km\)” row of Table[SI 6](https://arxiv.org/html/2607.20542#S2.T6)\. Purple dots represent control facilities and yellow dots represent treated facilities located within 15km of district borders\. Shorter distances \(5km, 10km\) yield similar results\.
#### 2\.6\.3Imputation strategies\.

Our main analysis simply dropped missing values; however, this may create concerns about bias from missing/unreliable data, and so we tested three different imputation strategies \(based on low\-rank completion, catchment population, and average demand\) and found consistent treatment effects across all of them \(see Table[SI 6](https://arxiv.org/html/2607.20542#S2.T6)\)\. Specifically:

- •Low\-rank imputation \(ImputedLowRank\):We used low\-rank matrix completion \(?\), a standard approach for handling missing data in large matrices\. We used thesoftImputeR package withrank=2\\text\{rank\}=2and regularization parameterλ=0\.01\\lambda=0\.01\(chosen using cross\-validation\) to impute missing consumption values\. Table[SI 7](https://arxiv.org/html/2607.20542#S2.T7)shows that our imputation results are robust to the choice of regularization parameterλ\\lambda\.
- •Population\-based imputation \(ImputedPop\):We first estimated demand in proportion to each facility’s catchment population, and then imputed consumption as the minimum of the estimated demand and the computed allocation \(we computed the allocation via the Excel tool in quarters prior to our tool’s deployment, and using our tool otherwise\)\.
- •Average imputation \(ImputedAvgConsump\)\.Third, using comprehensive data on unique quarter, facility, and product tuples, we imputed missing consumption values by assigning it to be the average quarterly consumption of the product across all facilities\.

We found that our estimates are qualitatively similar to our main analysis, suggesting that our results are robust to different imputation strategies\.

Table SI 7:Low\-Rank Imputation with Differentλ\\lambdaValues\.CoefficientStd\. ErrorObservationsImprovement %Dependent Variable = Normalized Consumption\(1\)λ=0\.01\\lambda=0\.010\.037∗∗\(0\.014\)5,45515%\(2\)λ=0\.1\\lambda=0\.10\.044∗∗∗\(0\.013\)5,45516%\(3\)λ=1\\lambda=10\.035∗∗\(0\.012\)5,45513%\(4\)λ=5\\lambda=50\.042∗∗∗\(0\.012\)5,45514%\(5\)λ=10\\lambda=100\.039∗∗\(0\.012\)5,45516%
- •Notes: Standard errors in parentheses\.\*p<0\.05p<0\.05,\*\*p<0\.01p<0\.01,\*\*\*p<0\.001p<0\.001\.

#### 2\.6\.4Missingness analysis\.

We checked that \(1\) our allocation tool did not drive differential missingness patterns, and \(2\) that products with differential missingness did not drive our results:

1. 1\.Missingness as the outcome variable: We used a missingness indicator as the dependent variable in our main specification; results are shown in the “Missingness” row of Table[SI 6](https://arxiv.org/html/2607.20542#S2.T6), and show no significant treatment effect\. These results suggest that our intervention did not cause differential missing data, which could bias our results\.
2. 2\.Analysis on a sub\-sample with no missingness imbalance: We also re\-ran our main analysis on a subset of products that showed no significant difference in missing data rates between the treatment and control groups after the treatment\. Results are shown in the “No Missingness Imbalance” row of Table[SI 6](https://arxiv.org/html/2607.20542#S2.T6); they are consistent with our main analysis, showing a statistically significant 18% increase in consumption\.

These results suggest that our findings are not driven by differential missing data patterns\.

#### 2\.6\.5Substitution\.

As discussed in §[2\.3\.2](https://arxiv.org/html/2607.20542#S2.SS3.SSS2), facilities may substitute related products when stock of one product is low\. As outlined in §[2\.3\.2](https://arxiv.org/html/2607.20542#S2.SS3.SSS2), we merged substitutable products into clusters, and performed our analysis on \(appropriately normalized\) consumption of product clusters\.

To this end, we developed groupings of substitutable products based on conversations with two medical experts, and validated them through discussions with field workers\. Table[SI 8](https://arxiv.org/html/2607.20542#S2.T8)shows substitutable groups of products among the ones we allocated; direct substitutes are perfect, functional substitutes are nearly perfect, and therapeutic substitutes are more complicated but can reasonably be approximated as being perfect\. Our discussions suggested that group 5 was context\-dependent, so we performed two analyses—one with all seven groups \(Method 1\) and one omitting group 5 \(Method 2\)\. We re\-ran our primary analysis using both groupings, which accounts for substitution effects according to our theoretical analysis\. Our results \(the “Substitution” rows in Table[SI 6](https://arxiv.org/html/2607.20542#S2.T6)\) show consistent and significant ATTs of 18% or 20%, suggesting that substitution behavior is unlikely to significantly affect our findings\.

Table SI 8:Medicine Substitutions\.This table lists clinically similar products grouped by substitution type: Direct Substitutions \(same ingredient but in different doses or forms\), Functional Substitutions \(different products with similar clinical purposes\), and Therapeutic Substitutions \(different medications for the same condition\)\. These expert\-provided substitution groups were used to test whether observed consumption increases could be driven by product substitution\. From our conversations, substitution group \(5\) \(marked with a†\\dagger\) was context\-dependent; thus, we performed checks both with and without this group\.Type of SubstitutionSubstitutesDirect Substitutions\(1\) Paracetamol \(Acetaminophen\) 500mg, Tab; Paracetamol \(Acetaminophen\) 250mg, Dispersible, TabFunctional Substitutions\(2\) Syringe, Luer, 2ml, Disposable, Pcs; Syringe, Luer, 10ml, Disposable, Pcs; Syringe, Luer, 20ml, Disposable, Pcs\(3\) Glucose \(Dextrose\) 5%, IV Inj, 500ml, Soft Bag; Glucose \(Dextrose\) Hypertonic 50%, IV Inj, 50ml, Bot\(4\) Cannula, IV, 18G, Short, Sterile, Disposable, Pcs; Cannula, IV, 24G, Short, Sterile, Disposable, PcsTherapeutic Substitutions\(5\)†Neomycin & Bacitracin 0\.5% & 500IU/g, Ointment, 15g, Tube; Povidone Iodine 10%, Solution, Bot; Chlorhexidine Gluconate 7\.1%, Gel\(6\) Oxytocin 10IU, Inj, Amp; Misoprostol 200mcg, Tab\(7\) Ampicillin 500mg, Pdr for IM/IV, Inj, Vial; Ceftriaxone 250mg, Pdr for InjNote:Every substitution requires a qualified healthcare professional to first assess the patient’s specific diagnosis and clinical context, then manage the substitution by adjusting dosage and following clinical or administration procedures \(e\.g\., dilution, etc\.\)\.
#### 2\.6\.6Alternative control group \(Alt\. Control\)\.

We also performed our analysis using an alternative control group of 26 products that were concurrently allocated using a different, pre\-existing mechanism \(see Table[SI 11](https://arxiv.org/html/2607.20542#S2.T11)\)\. The consumption levels for these products can be used as a control group throughout our study period across all districts in the country using a staggered treatment—i\.e\., an advantage of this analysis is that it can be performed not just for the partial deployment in 2023 Q2, but also for the nationwide implementation starting in Q3\. We used thesdidpackage in STATA \(?\), and found a consistent and statistically significant increase in consumption of 18% \(“Alt\. Control” row in Table[SI 6](https://arxiv.org/html/2607.20542#S2.T6)\)\.

#### 2\.6\.7Local average treatment effect\.

To disentangle the effect of compliance with our recommendations from the impact of our system, we performed a standard instrumental variables analysis to compute the Local Average Treatment Effect \(LATE\) \(?\)\. Following standard practice, we used the government’s random \(binary\) treatment assignment as the instrument, and compliance \(measured as a continuous variable, described in §[2\.4](https://arxiv.org/html/2607.20542#S2.SS4)\) as the treatment variable\.

In general, IV analyses require two conditions to be satisfied\. The first is relevance—our first\-stage results in Table[SI 9](https://arxiv.org/html/2607.20542#S2.T9)show a strong correlation between treatment assignment and compliance \(coefficient=0\.89\\text\{coefficient\}=0\.89,p<0\.001p<0\.001\), suggesting that the relevance assumption holds\. The second is the exclusion restriction, which requires that the assignment to treatment affects outcomes only through compliance\. This is highly plausible in our case\. For consumption to change, either patient/provider behavior or supply chain logistics must change\. Yet, patients and facility\-level providers had no knowledge of our deployment \(or more generally, NMSA’s operations\), and assignment alone did not alter logistics unless the district actually adopted our recommended allocation\. Thus, we believe assignment to treatment would not influence consumption except through compliance, satisfying the exclusion restriction\.

To estimate the LATE, we used Two\-Stage Least Squares \(2SLS\), with the first\-stage

Dn=π1​Zn\+π2​Xn\+μn,D\_\{n\}=\\pi\_\{1\}Z\_\{n\}\+\\pi\_\{2\}X\_\{n\}\+\\mu\_\{n\},\(12\)whereDnD\_\{n\}is compliance,ZnZ\_\{n\}is treatment assignment, andXnX\_\{n\}denotes covariates \(indicators for the district, quarter, and facility type\)\. Here,nnranges over all \(treated and control\) facilities\.

The second\-stage regression is given by

Yn=β1​Dn\+β2​Xn\+ϵn,Y\_\{n\}=\\beta\_\{1\}D\_\{n\}\+\\beta\_\{2\}X\_\{n\}\+\\epsilon\_\{n\},\(13\)whereYnY\_\{n\}is the outcome variable \(i\.e\., normalized consumption\)\. Table[SI 9](https://arxiv.org/html/2607.20542#S2.T9)shows results, yielding a significant increase of 0\.12 standard deviations\. To interpret this result, we make an additional assumption that the first\-stage relationship between assignment and compliance is monotonic and linear—then, this result can be interpreted as “a one\-unit increase in compliance induced by the instrument yields a 0\.12 standard deviation increase in outcome,” which can be translated to a 37% reduction in unmet demand among compliers\.999We can recover the IV coefficient in percentage terms: The implied percentage increase in consumption for each productppis computed as%ΔpLATE=100×β^norm×spY¯p\\%\\Delta^\{\\text\{LATE\}\}\_\{p\}=100\\times\\hat\{\\beta\}^\{\\text\{norm\}\}\\times\\frac\{s\_\{p\}\}\{\\bar\{Y\}\_\{p\}\}, whereβ^norm\\hat\{\\beta\}^\{\\text\{norm\}\}is the 2SLS coefficient on the standardized outcome,sps\_\{p\}is the product\-specific standard deviation, andY¯p\\bar\{Y\}\_\{p\}is its baseline mean consumption\. Then, we take the simple average of these product\-level effects to obtain an overall increase of approximately 37% for compliers\.

Table SI 9:LATE IV Result\.Column 1 shows the OLS estimate of treatment compliance \(continuous measure defined in Supplement §[2\.4](https://arxiv.org/html/2607.20542#S2.SS4)\) on normalized consumption\. Column 2 presents the first\-stage regression of treatment assignment \(instrument\) on compliance\. Column 3 reports the IV estimate, which identifies the causal effect for compliers\.\(1\)OLS\(2\)First\-Stage\(3\)IVDependent variable:NormalizedConsumptionTreated ComplierNormalizedConsumptionTreated Complier0\.075\*\*–0\.115\*\*\(0\.023\)\(0\.038\)IV \(treatment assignment\)–0\.887\*\*\*–\(0\.016\)Facility fixed effectYESYESYESQuarter fixed effectYESYESYESFacility Type fixed effectYESYESYESObservations5,2905,2905,290R2R^\{2\}0\.067\-0\.067First\-StageFF\-Stat–3,318–Notes: Standard errors in parentheses\.\*p<<0\.05,\*\*p<<0\.01,\*\*\*p<<0\.001\.
#### 2\.6\.8Stockouts\.

We also assessed the impact of our system on the number of facility\-product stockouts\. While reducing stockouts might seem like a natural objective, optimizing for fewer stockouts produces highly undesirable allocations—e\.g\., an optimal strategy is to allocate zero supply to a small number of high\-volume facilities, thereby ensuring that the remaining facilities are well\-stocked\. Using SynthDiD but replacing our unit\-level outcomes with binary stockout indicators, we found a directional reduction of 4\.6% in stockouts but it is not statistically significant \(p≈0\.12p\\approx 0\.12\) \(“Stockouts” rows in Table[SI 6](https://arxiv.org/html/2607.20542#S2.T6)\)—i\.e\., our system does not inadvertently increase stockouts\.

### 2\.7Cost\-Effectiveness Analysis

Our system was highly cost\-effective\. By design, it did not induce new costs for warehousing or logistics; instead, it integrated directly into the existing supply chain infrastructure, replacing the previous manual planning process without requiring additional physical resources or personnel\. The primary cost was a nationwide $30 USD monthly server fee for hosting the prediction and optimization pipeline and the interface by which NMSA workers interact with our system\.

We conducted an Incremental Cost\-Effectiveness Ratio \(ICER\) analysis using Disability\-Adjusted Life Years \(DALYs\) as our measure of effectiveness; one DALY represents one year of healthy life, with one year of life with some disease, disability, or other health condition represented by some fraction less than one\. To this end, we matched each medicine in our study to the condition it treats, and obtained the number of DALYs gained by treating that condition according to the World Health Organization \(?\)\. For products that can treat multiple conditions, we used a conservative approach by using the minimum DALY value that one unit of the medicine could avert; this strategy ensures that our effectiveness estimate is a lower bound\. The ICER is then calculated as the incremental cost of the interventionΔ​C\\Delta Cdivided by the incremental health effectΔ​E\\Delta E\. For every productm∈\[M\]m\\in\[M\], letq¯m\\bar\{q\}\_\{m\}denote the minimum number of DALYs that can be saved when one extra unit of productmmis consumed, and letF=360 USDF=\\text\{360 USD\}denote the yearly fixed cost of hosting our system for the entire nation\. Then, we have

ICER=Δ​CΔ​E=Fδ​∑m=1Mq¯m,\\displaystyle\\mathrm\{ICER\}=\\frac\{\\Delta C\}\{\\Delta E\}=\\frac\{F\}\{\\delta\\sum\_\{m=1\}^\{M\}\\bar\{q\}\_\{m\}\},whereδ\\deltais the treatment effect from our main impact evaluation\. We usedδ=0\.19\\delta=0\.19, representing the 19% average increase in medicine consumption attributed to our system, and we computed the yearly increase in DALYs from increased consumption of distributed products∑m=1Mq¯m=361\.3\\sum\_\{m=1\}^\{M\}\\bar\{q\}\_\{m\}=361\.3\. This analysis yields an average ICER of only $5\.24 USD per DALY\. This is much smaller than the ICER of typical interventions that are considered cost\-effective \($50,000–$100,000\) \(?,?,?\)\. While our ICER is extremely low, it is consistent with recent literature on the “unreasonable effectiveness of algorithms,” which finds that algorithms can significantly improve resource allocation at a very low marginal cost \(?\)\.

Table SI 10:Products Allocated via Our Tool\.Statistics for monthly facility\-level consumption between Jan 2020 and Nov 2023\. Missing values are excluded\.Product NameMeanMedianStd\. Dev\.MedicinesAlbendazole 400mg, Tab124\.8990205\.83Aluminium Hydroxide 500mg, Tab169\.7085246\.70Ampicillin 500mg, Pdr for IM/IV, Inj, Vial65\.5534203\.59Benzyl Benzoate 25%, Emulsion, 100ml, Bot12\.76165\.00Ceftriaxone 250mg, Pdr for Inj113\.3925226\.94Chlorhexidine Gluconate 7\.1%, Gel3\.58216\.11Ciprofloxacin 500mg, Tab171\.57100268\.18Erythromycin 125mg/5ml, Pdr for Susp, 100ml, Bot89\.2715224\.83Ferrous Sulphate 200mg, Tab512\.29300673\.05Folic Acid 5mg, Tab509\.04398745\.81Glucose \(Dextrose\) 5%, IV Inj, 500ml, Soft Bag40\.736155\.03Glucose \(Dextrose\) Hypertonic 50%, IV Inj, 50ml, Bot9\.60227\.87Methyldopa 250mg, Tab73\.7345168\.29Metoclopramide HCl 10mg, Tab147\.3130306\.85Misoprostol 200mcg, Tab31\.9110144\.94Neomycin & Bacitracin 0\.5% & 500IU/g, Ointment, 15g, Tube4\.70217\.82Oral Rehydration Salts \(ORS\), Sachet121\.27100187\.75Oxytocin 10IU, Inj, Amp19\.611086\.72Paracetamol \(Acetaminophen\) 250mg, Dispersible, Tab464\.35400475\.49Paracetamol \(Acetaminophen\) 500mg, Tab562\.41455635\.66Povidone Iodine 10%, Solution, Bot8\.64184\.45Prednisolone 5mg, Tab65\.2412134\.54Salbutamol 100mcg/dose, Aerosol, Inhaler1\.50014\.28Water for Injection 10ml, Inj, Amp33\.7520124\.35Zinc Sulphate 20mg, Dispersible, Tab221\.20130459\.66Medical Supplies & EquipmentApron, Plastic, Disposable, Pcs24\.6910124\.65Cannula, IV, 18G, Short, Sterile, Disposable, Pcs14\.901048\.70Cannula, IV, 24G, Short, Sterile, Disposable, Pcs16\.241061\.55Envelope, Dispensing, Plastic, 10cm x 7cm, Pcs150\.11100184\.44Glove, Exam, Latex, Medium, Nonsterile, Disposable, Pcs173\.18100338\.22IV Giving Set, Pcs19\.231149\.27Mask, Surgical, Pcs47\.7312215\.93Rapid test kit, Pregnancy, Pcs13\.731038\.64Syringe, Luer, 10ml, Disposable, Pcs37\.6517188\.69Syringe, Luer, 20ml, Disposable, Pcs25\.801075\.47Syringe, Luer, 2ml, Disposable, Pcs41\.232597\.43Table SI 11:Products Allocated via Other Mechanisms\.Statistics for monthly facility\-level consumption between Jan 2020 and Nov 2023\. Missing values are excluded\. Data from these products were only used to help train our prediction model and for one of our evaluations \(staggered SynthDiD using product\-level controls, called “Alt\. Control” in Table[1](https://arxiv.org/html/2607.20542#S5.T1)\)\.Product NameMeanMedianStd\. Dev\.MedicinesAmoxicillin 250mg, Dispersible, Tab592\.01456734\.22Calcium Gluconate 100mg/ml, Inj, 10ml, Amp5\.20052\.47Ceftriaxone 1g, Pdr for Inj, Vial164\.4520502\.68Chlorhexidine Gluconate 5%, Solution, 1000ml, Bot6\.381116\.60Cloxacillin 500mg, Tab/Cap300\.74100550\.75Epinephrine HCl \(Adrenaline\) 1mg/ml, Inj, 1ml, Amp9\.25078\.06Ibuprofen 400mg, Tab284\.31200346\.74Jadelle8\.07522\.51Levonorgestrel \(Emergency Contraceptive\) 1\.5mg, Tab4\.87015\.75Levoplant5\.33117\.90Nystatin 500,000IU, Tab8\.74424\.61Progesterone\-Only \(Microlut\) Levonorgestrel 30mcg, Tab, Cycle19\.633145\.70Sodium Chloride \(Normal Saline\) 0\.9%, IV Inj, 500ml, Bot12\.27850\.45Medical Supplies & EquipmentBandage, Elastic, 8cm x 4m, Roll10\.75143\.08Blade, Surgical, No\. 22, Sterile, Disposable, Pcs18\.23591\.90Cannula, IV, 22G, Short, Sterile, Disposable, Pcs14\.33651\.38Condom, Male114\.4060258\.42Copper\-Containing Device \(Copper T or Copper 7 or IUD\)7\.41028\.42Cotton Wool, Absorbent, 500g, Roll2\.73117\.93Glove, Surgical, Size 7\.5, Sterile, Disposable, Pair49\.3413112\.43Glove, Surgical, Size 8, Sterile, Disposable, Pair37\.9810108\.17Needle, Hypodermic, Luer, 21G, Sterile, Disposable69\.5650135\.77Needle, Hypodermic, Luer, 23G, Sterile, Disposable62\.095097\.86Sanitary Pads, Pcs9\.56220\.52Syringe, Luer, 5ml, Disposable57\.0645111\.68Tube, Asp/Feed, CH12, Sterile, Disposable, Pcs4\.16019\.59Table SI 12:Balance Table: District Characteristics by Treatment StatusControlTreatmentDifferenceVariableMean\(SD\)Mean\(SD\)p\-valueDays of service delivery0\.46\(0\.17\)0\.38\(0\.17\)0\.350\.71Hours of service delivery1\.40\(0\.56\)1\.26\(0\.55\)0\.140\.71Facilities where women give birth6\.44\(2\.54\)5\.23\(2\.47\)0\.190\.71Availability of basic emergency obstetric and neonatal care0\.43\(0\.38\)0\.19\(0\.18\)0\.350\.71Availability of priority drugs3\.62\(1\.39\)3\.12\(1\.58\)0\.780\.88Availability of all tracer drugs1\.84\(1\.41\)2\.40\(1\.78\)0\.240\.71Availability of vaccines6\.32\(2\.31\)5\.31\(2\.46\)0\.680\.88Vaccines storage: refrigerators22–8∘​C8\\,^\{\\circ\}\\mathrm\{C\}1\.72\(2\.72\)1\.06\(2\.14\)0\.870\.88Availability of communication equipment3\.65\(2\.56\)2\.56\(1\.50\)0\.410\.71Access to various forms of communication3\.65\(2\.56\)2\.56\(1\.50\)0\.410\.71Total proportion of facilities carrying out safe health care waste disposal5\.00\(2\.10\)3\.86\(1\.66\)0\.680\.88Availability of basic equipment1\.99\(0\.67\)2\.00\(0\.91\)0\.310\.71Availability of Standard Treatment Guidelines3\.63\(1\.97\)2\.53\(2\.56\)0\.440\.71Outpatient caseload \(median per facility\)0\.63\(0\.52\)0\.51\(0\.21\)0\.880\.88Facilities with community health workers5\.44\(2\.54\)4\.84\(2\.82\)0\.780\.88Average number of community health workers0\.69\(0\.24\)0\.64\(0\.33\)0\.570\.85Average health workers per facility0\.47\(0\.50\)0\.41\(0\.38\)0\.810\.88Population Share6\.95\(3\.13\)4\.72\(2\.45\)0\.150\.71Human Development Index \(HDI\)0\.45\(0\.05\)0\.43\(0\.02\)0\.160\.21Health Index \(health dimension of HDI based on life expectancy at birth\)0\.63\(0\.05\)0\.64\(0\.06\)0\.660\.66Income Index \(income dimension of HDI based on Gross National Income per capita\)0\.41\(0\.06\)0\.38\(0\.02\)0\.110\.21

Notes:Means and standard deviations are shown as Mean \(SD\)\. p\-values are from two\-sample Welch’s t\-tests\. Multiple testing adjustments control the false discovery rate \(FDR\) using the Benjamini–Hochberg \(BH\) procedure\.

Table SI 13:Percentage of Missing Data by Product\.We report missingness rates in consumption data for each product across control and treated facilities in pre\- and post\-treatment periods\. Thepp\-values are from two\-sample proportion tests of significant differences in missingness rates between control and treated groups, without adjusting for multiple hypothesis testing\.Pre\-TreatedPost\-TreatedProductControl\(%\)Treated\(%\)p\-valueControl\(%\)Treated\(%\)p\-valueAlbendazole 400mg, Tab0\.000\.001\.000\.100\.001\.00Aluminium Hydroxide 500mg, Tab97\.6099\.000\.1595\.2091\.701\.00Ampicillin 500mg, Pdr for IM/IV, Inj, Vial14\.6024\.200\.065\.5010\.100\.00\*\*\*Apron, Plastic, Disposable, Pcs73\.8042\.600\.00\*\*\*75\.9055\.200\.00\*\*\*Benzyl Benzoate 25%, Emulsion, 100ml, Bot98\.5098\.701\.0070\.2083\.800\.00\*\*\*Cannula, IV, 18G, Short, Sterile, Disposable, Pcs42\.5032\.300\.02\*66\.4053\.600\.00\*\*\*Cannula, IV, 24G, Short, Sterile, Disposable, Pcs39\.3026\.400\.00\*\*\*6\.503\.601\.00Ceftriaxone 250mg, Pdr for Inj96\.7099\.400\.6688\.9088\.200\.75Chlorhexidine Gluconate 7\.1%, Gel28\.7015\.900\.00\*\*\*3\.904\.800\.48Ciprofloxacin 500mg, Tab90\.6087\.900\.1967\.9022\.300\.00\*\*\*Envelope, Dispensing, Plastic, 10cm x 7cm, Pcs55\.4032\.700\.00\*\*\*70\.9052\.900\.00\*\*\*Erythromycin 125mg/5ml, Pdr for Susp, 100ml, Bot99\.5099\.000\.4382\.7085\.400\.32Ferrous Sulphate 200mg, Tab75\.7076\.900\.7183\.4054\.200\.00\*\*\*Folic Acid 5mg, Tab0\.100\.300\.151\.806\.100\.00\*\*\*Glove, Exam, Latex, Medium, Nonsterile, Disposable, Pcs4\.706\.900\.04\*6\.904\.300\.05Glucose \(Dextrose\) 5%, IV Inj, 500ml, Soft Bag91\.2096\.500\.1376\.1077\.600\.65Glucose \(Dextrose\) Hypertonic 50%, IV Inj, 50ml, Bot99\.6099\.400\.6487\.0092\.900\.38IV Giving Set, Pcs5\.004\.200\.464\.304\.800\.69Mask, Surgical, Pcs79\.1074\.001\.0083\.7077\.201\.00Methyldopa 250mg, Tab0\.200\.500\.300\.300\.301\.00Metoclopramide HCl 10mg, Tab90\.1088\.000\.3385\.1084\.000\.65Misoprostol 200mcg, Tab60\.6067\.901\.0070\.7075\.400\.10Neomycin & Bacitracin 0\.5% & 500IU/g, Ointment, 15g, Tube97\.6099\.000\.1569\.3068\.700\.83Oral Rehydration Salts \(ORS\), Sachet0\.000\.001\.000\.200\.800\.10Oxytocin 10IU, Inj, Amp1\.501\.300\.863\.303\.700\.68Paracetamol \(Acetaminophen\) 250mg, Dispersible, Tab0\.200\.200\.640\.100\.201\.00Paracetamol \(Acetaminophen\) 500mg, Tab38\.2028\.400\.01\*12\.1021\.000\.00\*\*Povidone Iodine 10%, Solution, Bot89\.7090\.000\.915\.5020\.100\.00\*\*\*Prednisolone 5mg, Tab100\.0099\.000\.5593\.6090\.700\.12Rapid test kit, Pregnancy, Pcs69\.8061\.600\.4962\.3059\.200\.35Salbutamol 100mcg/dose, Aerosol, Inhaler89\.1076\.800\.00\*\*\*40\.3044\.600\.20Syringe, Luer, 10ml, Disposable, Pcs80\.1054\.200\.00\*\*\*57\.9029\.000\.00\*\*\*Syringe, Luer, 20ml, Disposable, Pcs70\.4046\.000\.00\*\*\*61\.5031\.800\.00\*\*\*Syringe, Luer, 2ml, Disposable, Pcs75\.0047\.000\.00\*\*\*56\.8031\.200\.00\*\*\*Water for Injection 10ml, Inj, Amp25\.4025\.200\.9532\.3030\.800\.68Zinc Sulphate 20mg, Dispersible, Tab0\.000\.001\.000\.200\.300\.64
Notes:Significance levels:\*p<<0\.05,\*\*p<<0\.01,\*\*\*p<<0\.001\.

Similar Articles

Measuring the impact of learning with AI in Sierra Leone and beyond

Google DeepMind Blog

A pre-registered trial in Sierra Leone found that AI-powered Guided Learning significantly improved math scores, achieving 1.2 to 1.7 years of progress in eight weeks, while teachers reported enhanced professional growth and a shift toward facilitation roles.

Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Care

arXiv cs.LG

This paper introduces Manana, a non-parametric prompt-learning framework that teaches LLMs to recommend anti-seizure medications and defer uncertain cases in underrepresented epilepsy care settings, improving accuracy on Ugandan cohorts and enabling selective prediction with high precision.