Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments
Summary
This paper investigates how historical A/B test data can inform adaptive experiments using contextual bandits, providing a practical methodology for deciding when and how to deploy adaptive policies based on offline policy evaluation.
View Cached Full Text
Cached at: 09/29/26, 09:32 AM
# Offline Policy Evaluation as a decision‑support tool for designing Adaptive Experiments
Source: [https://arxiv.org/html/2609.30273](https://arxiv.org/html/2609.30273)
††footnotetext:1 joaovfalv@gmail\.com
2 eduardo\.laurentino@itau\-unibanco\.com\.br
3 gustavo\.kanno@itau\-unibanco\.com\.br
4 thiago\.rizuti\-rocha@itau\-unibanco\.com\.br
⋆\\starEqual contribution as first authors\.
†\\daggerCorresponding author\.Eduardo Rocha Laurentino2⋆\\star[https://orcid.org/0000-0001-5100-5029](https://orcid.org/0000-0001-5100-5029)Gustavo de Oliveira Kanno3[https://orcid.org/0009-0008-7329-3031](https://orcid.org/0009-0008-7329-3031)Thiago Costa Rizuti da Rocha4†\\dagger[https://orcid.org/0009-0005-9708-2253](https://orcid.org/0009-0005-9708-2253)
###### Abstract
We investigate how historical data from fixed randomized experiments \(A/B tests\) can be used to inform the deployment of adaptive experiments based on contextual bandits\. Given data collected under a static allocation, our goal is to assess which adaptive policies, if any, would have outperformed the original design and under what conditions\. To this end, we combine off\-policy evaluation \(OPE\) with a controlled warm\-start simulation\. From logged A/B test data exhibiting heterogeneous treatment effects, we estimate nuisance components and use doubly robust estimators to rank a portfolio of pre\-specified adaptive and non\-adaptive policies\. When ground truth is available, we then deploy the same offline\-trained policies in a simulator that reuses the exact data\-generating reward probabilities, providing a safe, ground\-truth\-anchored environment to study the offline\-to\-online transition under warm starting\. Using synthetic randomized controlled trials with known heterogeneity structures and an oracle policy, our results indicate that adaptive, context\-aware policies improve upon fixed allocations when meaningful heterogeneity is present, while providing little benefit in its absence\. We reinforce our findings on standard open benchmarks \(Hillstrom, Criteo Uplift, and LaLonde\), reinterpreted through a policy\-value and regret perspective\. Overall, our results provide a practical methodology for deciding when adaptive experimentation is worth deploying and how to select among competing adaptive policies using existing A/B test data\.
Keywords:Adaptive Experiments Contextual Multi Armed Bandits Offline Policy Evaluation\.
## 1Introduction
Randomized controlled trials \(RCTs\), commonly known as A/B tests, remain the gold standard for causal inference in technology\-based companies, clinical research, and social sciences\. In a typical A/B test, experimental units are assigned to one ofKKtreatment arms with fixed and uniform probabilities, and the experimenter collects outcome data until a pre\-specified sample size is reached\. This simplicity is both the method’s greatest strength \(since it yields unbiased estimates under minimal assumptions\) and its most consequential limitation: by treating all individuals identically, A/B tests forgo the opportunity to personalize treatment assignment based on observable characteristics that may moderate treatment effects\.
Adaptive experimental designs, Contextual Multi Armed Bandits \(CMAB\) in particular, address this limitation by sequentially adjusting the assignment policy as data accumulates\. At each roundtt, a contextual bandit algorithm observes a feature \(context\) vectorXtX\_\{t\}, selects an action \(treatment\)AtA\_\{t\}according to a policyπt\\pi\_\{t\}that depends on the history of past observations, and receives a stochastic reward \(outcome\)YtY\_\{t\}\. This adaptivity offers two potential advantages over static designs: it can improve outcomes for participants during the experiment itself, by routing individuals toward more promising treatments; and it can accelerate the identification of heterogeneous treatment effects \(HTE\), by concentrating exploration where uncertainty is high\. These properties have motivated increasing adoption of bandit\-based experimentation in domains ranging from digital advertising to clinical trial design and recommendation systems\[[7](https://arxiv.org/html/2609.30273#bib.bib28)\]\.
However, the transition from A/B testing to adaptive experimentation is rarely straightforward\. Deploying a contextual bandit requires confidence that it will outperform the existing static design, yet this confidence is difficult to obtain without running the adaptive experiment itself\. Moreover, even when the decision to adopt a contextual bandit is made, the initial deployment phase can still suffer from a cold\-start problem: when data are scarce, the algorithm must explore in order to learn which actions are effective for which contexts, and this exploration can reduce short\-term performance before sufficient evidence accumulates\.\[[12](https://arxiv.org/html/2609.30273#bib.bib21),[1](https://arxiv.org/html/2609.30273#bib.bib22)\]
In this paper, we argue that historical data from completed A/B tests can be systematically leveraged to address the pre\-deployment uncertainty problem of adaptive experimentation\. Our goal is to use data from fixed A/B tests to anticipate, before deployment, whether adaptive \(contextual bandit\) policies would improve outcomes relative to the original static design\. We propose a two\-level framework that transforms static experimental data into actionable intelligence for adaptive experimentation:Level 1 \- Off\-Policy Evaluation \(OPE\):using historical A/B test logs to counterfactually rank candidate adaptive policies before any live deployment;Level 2 \- Controlled Warm Simulator:deploying the OPE\-selected policies in a simulator whose reward probabilities are exactly known from the synthetic data\-generating process, providing a safe, ground\-truth\-anchored environment in which policies evaluations can be verified online\.
## 2Background
A principled response to this challenge is*off\-policy evaluation*\(OPE\): the problem of estimating the valueV\(π\)V\(\\pi\)of a target policyπ\\pi\(the expected reward underπ\\pi\) using data collected by a different, fixed logging policyπ0\\pi\_\{0\}, without deployingπ\\piin the real world\[[4](https://arxiv.org/html/2609.30273#bib.bib1)\]\. In practice, OPE combines nuisance estimation \(propensity scoresb^\(a∣x\)\\hat\{b\}\(a\\mid x\)and outcome modelsμ^a\(x\)\\hat\{\\mu\}\_\{a\}\(x\)\) with counterfactual value estimators such as Inverse Propensity Scoring \(IPS\), Self\-Normalized IPS \(SNIPS\), and especially Doubly Robust \(DR\) estimators, which are attractive in finite samples because they remain consistent when either the outcome model or the propensity model is correctly specified and can be implemented with sample splitting to reduce overfitting bias\[[4](https://arxiv.org/html/2609.30273#bib.bib1),[18](https://arxiv.org/html/2609.30273#bib.bib3),[20](https://arxiv.org/html/2609.30273#bib.bib7)\]\. In this paper, we restrict our attention to the contextual bandit setting rather than the full sequential reinforcement learning regime\.
Among standard OPE estimators, Doubly Robust \(DR\) is particularly attractive because it combines the strengths of direct modeling and importance weighting\. The direct method \(DM\) estimates the expected outcome under each action through nuisance modelsμ^a\(x\)≈𝔼^\[Y∣X=x,A=a\]\\hat\{\\mu\}\_\{a\}\(x\)\\approx\\hat\{\\mathbb\{E\}\}\[Y\\mid X=x,A=a\], which can yield low\-variance policy\-value estimates but may be biased if the outcome model is misspecified\. Inverse Propensity Scoring \(IPS\), in contrast, reweights observed rewards by the ratioπ\(a∣x\)/b^\(a∣x\)\\pi\(a\\mid x\)/\\hat\{b\}\(a\\mid x\)and is unbiased when the logging policy is correctly known, but it can have high variance when propensities are small or when the target policy differs substantially from the behavior policy\. DR combines both components: it uses the outcome model as a baseline prediction and adds a bias\-correction term based on inverse propensity weighting\. As a result, the estimator remains consistent if either the outcome model or the propensity model is correctly specified, which makes it a natural default in finite samples and a strong choice for practical policy screening\. In our setting, where the logging policy is fixed and known in the synthetic experiments, DR offers a favorable bias\-variance trade\-off while retaining a clear counterfactual interpretation\[[4](https://arxiv.org/html/2609.30273#bib.bib1),[18](https://arxiv.org/html/2609.30273#bib.bib3),[20](https://arxiv.org/html/2609.30273#bib.bib7),[17](https://arxiv.org/html/2609.30273#bib.bib31),[9](https://arxiv.org/html/2609.30273#bib.bib15)\]\.
For deployment, one may also consider Offline Policy Learning \(OPL\), which addresses a different problem\. Rather than merely estimating the base expected value of a policy or for pre\-specified set of candidate policies, OPL searches over a policy classΠ\\Pito findπ^∗=argmaxπ∈ΠV^\(π\)\\hat\{\\pi\}^\{\*\}=\\arg\\max\_\{\\pi\\in\\Pi\}\\hat\{V\}\(\\pi\), collapsing the evaluation into a single optimized output\[[18](https://arxiv.org/html/2609.30273#bib.bib3),[4](https://arxiv.org/html/2609.30273#bib.bib1)\]\. In this work, we deliberately do not perform OPL, mainly because our scientific goal is*comparative*: we wish to characterize the best available policy options among a diverse pre\-specified portfolio of contextual bandit exploration strategies\.
A parallel literature addresses uplift modeling, referring to the estimation of heterogeneous treatment effects \(HTE\), or conditional average treatment effects \(CATE\), expressed asτ\(x\)=𝔼\[Y\(1\)−Y\(0\)∣X=x\]\\tau\(x\)=\\mathbb\{E\}\[Y\(1\)\-Y\(0\)\\mid X=x\], whereY\(1\)Y\(1\)andY\(0\)Y\(0\)are potential outcomes\. Standard uplift evaluation relies on rank\-based metrics such as the Qini coefficient and the area under the uplift curve \(AUUC\), which measure how well a model orders individuals by their individual treatment effect\. While these metrics are natural for model selection, they do not directly capture the value of a deployed policy\.
Here, we reframe the uplift problem from the model\-centric AUUC perspective to a*policy\-value*perspective, measuring expected rewardV\(π\)V\(\\pi\)and regretΔ\(π\)=V\(π⋆\)−V\(π\)\\Delta\(\\pi\)=V\(\\pi^\{\\star\}\)\-V\(\\pi\)relative to the best achieved policyπ⋆\\pi^\{\\star\}\. For doing that, three open datasets serve as our empirical benchmarks\. The*Hillstrom Mine That Data*dataset\[[8](https://arxiv.org/html/2609.30273#bib.bib5)\]is a canonical three\-arm email marketing randomized experiment whose treatment effect heterogeneity, while real, is modest, particularly for conversion and spend outcomes, as documented by\[[13](https://arxiv.org/html/2609.30273#bib.bib6)\]\. The*Criteo Uplift*dataset\[[3](https://arxiv.org/html/2609.30273#bib.bib9)\]is a large\-scale benchmark characterized by severe treatment imbalance and rare conversion outcomes\. The*LaLonde/NSW*dataset\[[10](https://arxiv.org/html/2609.30273#bib.bib12)\]and its re\-analysis by\[[2](https://arxiv.org/html/2609.30273#bib.bib11)\]provide a benchmark for causal inference methods in which the experimental ground truth is known, here reinterpreted through the lens of policy regret rather than matching estimator bias\.
Our work is mainly related to three research streams\. First, there is a large literature on off\-policy evaluation \(OPE\) for contextual bandits, including inverse\-propensity, self\-normalized, doubly robust, and bias\-variance controlled estimators, as well as methods for valid inference under adaptive data collection\[[4](https://arxiv.org/html/2609.30273#bib.bib1),[18](https://arxiv.org/html/2609.30273#bib.bib3),[20](https://arxiv.org/html/2609.30273#bib.bib7),[17](https://arxiv.org/html/2609.30273#bib.bib31),[6](https://arxiv.org/html/2609.30273#bib.bib27)\]\. Second, there is growing interest in reproducible OPE benchmarking and controlled experimentation environments, including open\-source software packages such as the Open Bandit Pipeline \(OBP\) and its associated public bandit datasets, which expose the practical challenges of evaluating policies from logged bandit feedback and standardize empirical comparisons\[[15](https://arxiv.org/html/2609.30273#bib.bib18)\]\. Third, our use of Hillstrom, Criteo, and LaLonde connects to the uplift and heterogeneous treatment effect literature, which studies how treatment effects vary across individuals and typically evaluates models through uplift\-oriented metrics\[[3](https://arxiv.org/html/2609.30273#bib.bib9),[5](https://arxiv.org/html/2609.30273#bib.bib32),[14](https://arxiv.org/html/2609.30273#bib.bib33)\]\.
Beyond estimator design, a complementary line of work emphasizes the role of synthetic and semi\-synthetic environments for controlled benchmarking of contextual bandits and off\-policy methods\. Such environments make it possible to vary reward structure, delays, concept drift, and business constraints while retaining access to a known reward\-generating process\. The simulator module of the OBP ecosystem provides configurable contextual bandit environments and policy simulators, thereby bridging offline evaluation and online experimentation in a general\-purpose benchmarking framework\[[15](https://arxiv.org/html/2609.30273#bib.bib18),[16](https://arxiv.org/html/2609.30273#bib.bib30)\]\. Later extensions of OBP further incorporate industry\-relevant challenges such as delayed feedback, concept drift, reward design, and operational constraints\[[19](https://arxiv.org/html/2609.30273#bib.bib29)\]\. Broader empirical benchmarking efforts also share this motivation of understanding algorithm behavior under controlled but practically meaningful conditions\[[1](https://arxiv.org/html/2609.30273#bib.bib22)\]\.
What distinguishes the present paper from previous work is that we combine these strands under a deployment\-oriented perspective\. Rather than focusing only on estimator accuracy, synthetic benchmark design, or uplift ranking in isolation, we ask whether historical data from fixed randomized experiments can be used to decide if adaptive experimentation is worth deploying at all, and which adaptive policy classes are promising candidates before live interaction\. In this sense, our contribution is not merely another OPE benchmark, another synthetic simulator, or another uplift application, but an end\-to\-end decision\-support framework: fixed\-test logs are used to estimate counterfactual policy values, identify promising adaptive candidates, and then stress\-test those same offline\-selected policies in a controlled warm\-start simulator prior to roll\-out\.
## 3Methodology
We generate i\.i\.d\. samples\(Xi,Ai,Yi\)\(X\_\{i\},A\_\{i\},Y\_\{i\}\)with observed contextXiX\_\{i\}, binary actionA∈\{0,1\}A\\in\\\{0,1\\\}and binary outcomeY∈\{0,1\}Y\\in\\\{0,1\\\}under a balanced logging policyb\(A=1∣X\)≡0\.5b\(A=1\\mid X\)\\equiv 0\.5\. For each contextx∈Xx\\in X, the data\-generating process is defined by the arm\-specific response probabilitiesμ1\(x\)=Pr\(Y=1∣A=1,X=x\)\\mu\_\{1\}\(x\)=\\Pr\(Y=1\\mid A=1,X=x\),μ0\(x\)=Pr\(Y=1∣A=0,X=x\),\\mu\_\{0\}\(x\)=\\Pr\(Y=1\\mid A=0,X=x\),and the oracle policy isπ⋆\(x\)=𝕀\{μ1\(x\)≥μ0\(x\)\}\.\\pi^\{\\star\}\(x\)=\\mathbb\{I\}\\\{\\mu\_\{1\}\(x\)\\geq\\mu\_\{0\}\(x\)\\\}\.We consider four representative HTE regimes\.
##### No context\.
This is the limiting case in which no personalization is possible\. The response probabilities are constant,
μ1=0\.60andμ0=0\.55,\\mu\_\{1\}=0\.60\\text\{ and \}\\mu\_\{0\}=0\.55,\(1\)so the optimal rule is also constant\.
##### Categorical heterogeneity\.
For a discrete covariatex∈\{1,2,3\}x\\in\\\{1,2,3\\\}, we specify the piecewise response probabilitiesPr\(Y=1∣A=a,x=xi\)≡pa1\(xi\)\\Pr\(Y=1\\mid A=a,x=x\_\{i\}\)\\equiv p^\{1\}\_\{a\}\(x\_\{i\}\), with
\[p01\(1\),p11\(1\),p01\(2\),p11\(2\),p01\(3\),p11\(3\)\]=\[0\.1,0\.8,0\.5,0\.5,0\.8,0\.1\]\\left\[p^\{1\}\_\{0\}\(1\),p^\{1\}\_\{1\}\(1\),p^\{1\}\_\{0\}\(2\),p^\{1\}\_\{1\}\(2\),p^\{1\}\_\{0\}\(3\),p^\{1\}\_\{1\}\(3\)\\right\]=\\left\[0\.1,0\.8,0\.5,0\.5,0\.8,0\.1\\right\]\(2\)
so that treatment is beneficial in some categories, neutral in others, and harmful in the remainder\. This setting mimics segment\-level heterogeneity such as customer type, region, or income band\.
##### Linear heterogeneity\.
For a continuous covariatex∈\[0,1\]x\\in\[0,1\]with linear effect, we define
Pr\(Y=1∣A=a,x\)=αa\+βax,with\(α0,β0,α1,β1\)=\(0\.9,−0\.8,0\.1,0\.8\)\.\\Pr\(Y=1\\mid A=a,x\)=\\alpha\_\{a\}\+\\beta\_\{a\}x,\\text\{ with \}\(\\alpha\_\{0\},\\beta\_\{0\},\\alpha\_\{1\},\\beta\_\{1\}\)=\(0\.9,\-0\.8,0\.1,0\.8\)\.\(3\)This yields a monotone switching structure in which the preferred treatment depends on the covariate value\.
##### Oscillatory heterogeneity\.
For a continuous covariatex∈\[0,1\]x\\in\[0,1\]with oscillatory effect, we define
Pr\(Y=1∣A=a,x\)=12\[1\+cos\(2π\(αa\+βax\)\)\],\\Pr\(Y=1\\mid A=a,x\)=\\frac\{1\}\{2\}\\left\[1\+\\cos\\\!\\big\(2\\pi\(\\alpha\_\{a\}\+\\beta\_\{a\}x\)\\big\)\\right\],\(4\)with\(α0,β0,α1,β1\)=\(0,1,0\.5,1\)\.\\text\{ with \}\(\\alpha\_\{0\},\\beta\_\{0\},\\alpha\_\{1\},\\beta\_\{1\}\)=\(0,1,0\.5,1\)\.This yields alternating treatment advantage across the covariate range and captures periodic or seasonal heterogeneity\.
##### Two covariates\.
To study richer structure, we combine a linear and an oscillatory covariate\. In the main experiments, we use the average combination
Pr\(Y=1∣A=a,x=\(xlin,xosc\)\)=12\(μalin\(xlin\)\+μaosc\(xosc\)\),\\Pr\(Y=1\\mid A=a,x=\(x\_\{\\text\{lin\}\},x\_\{\\text\{osc\}\}\)\)=\\frac\{1\}\{2\}\\big\(\\mu^\{\\text\{lin\}\}\_\{a\}\(x\_\{\\text\{lin\}\}\)\+\\mu^\{\\text\{osc\}\}\_\{a\}\(x\_\{\\text\{osc\}\}\)\\big\),\(5\)
and the product combination
Pr\(Y=1∣A=a,x=\(xlin,xosc\)\)=μalin\(xlin\)×μaosc\(xosc\),\\Pr\(Y=1\\mid A=a,x=\(x\_\{\\text\{lin\}\},x\_\{\\text\{osc\}\}\)\)=\\mu^\{\\text\{lin\}\}\_\{a\}\(x\_\{\\text\{lin\}\}\)\\times\\mu^\{\\text\{osc\}\}\_\{a\}\(x\_\{\\text\{osc\}\}\),\(6\)whereμalin\(xlin\)\\mu^\{\\text\{lin\}\}\_\{a\}\(x\_\{\\text\{lin\}\}\)andμaosc\(xosc\)\\mu^\{\\text\{osc\}\}\_\{a\}\(x\_\{\\text\{osc\}\}\)are given by \([3](https://arxiv.org/html/2609.30273#S3.E3)\) and \([4](https://arxiv.org/html/2609.30273#S3.E4)\)\. These two combinations preserves the probability range and produces a smooth two\-dimensional decision surface\.
\(a\)Categorical heterogeneity\.
\(b\)Linear heterogeneity\.
\(c\)Oscillatory heterogeneity\.
\(d\)Two covariates with average combination\.
\(e\)Two covariates with product combination\.
Figure 1:Ground truth probability of effect given the Treatment and Covariate value\.For OPE, we split the data so that the observations used to fit the nuisance components are disjoint from those used to evaluate candidate policies\. This prevents optimistic bias and follows the standard sample\-splitting logic used for estimation in\[[4](https://arxiv.org/html/2609.30273#bib.bib1),[20](https://arxiv.org/html/2609.30273#bib.bib7)\]\. We evaluate a fixed portfolio of contextual bandit policies \(Bootstrapped Thompson Sampling, Bootstrapped UCB, Softmax/Boltzmann,ε\\varepsilon\-Greedy\), together with fixed baselines and an oracle policy whenμa\(x\)\\mu\_\{a\}\(x\)is known\. On the held\-out test split, we estimate policy value primarily with DR and use IPS and SNIPS as robustness checks\[[4](https://arxiv.org/html/2609.30273#bib.bib1),[18](https://arxiv.org/html/2609.30273#bib.bib3),[20](https://arxiv.org/html/2609.30273#bib.bib7)\]\. Uncertainty is quantified via bootstrap confidence intervals for the OPE estimates\. Inference for policy values in adaptive experiments is known to be delicate when assignment probabilities decay toward zero\[[6](https://arxiv.org/html/2609.30273#bib.bib27)\], but the synthetic OPE stage uses a fixed randomized logging policy with known propensities, so this extreme regime does not arise\.
Each policy is initialized with the parameters learned during the offline stage and then deployed sequentially in this matched environment, receiving rewards sampled from the sameμa\(x\)\\mu\_\{a\}\(x\)functions that generated the offline data\. Accordingly, the simulator should be interpreted as a controlled best\-case benchmark for offline\-to\-online transfer rather than as a realistic proxy for deployment under model misspecification, covariate shift, or temporal drift\.
The OPE\-predicted valueV^\(πk\)\\hat\{V\}\(\\pi\_\{k\}\)serves as a pre\-deployment forecast, but offline and online quantities are not identical objects in our setting\. Offline OPE evaluates a stationary target ruleπ\(a∣x\)\\pi\(a\\mid x\), whereas the online bandit policies considered here \(e\.g\., BootstrappedTS and BootstrappedUCB\) are history\-dependent and therefore induce learning\-phase behavior that changes with past observations\. Accordingly, discrepancies between offline OPE and realized online reward may arise not only from finite\-sample noise and stochastic action selection, but also from this mismatch between stationary\-policy evaluation and history\-dependent online learning dynamics\. For this reason, we interpret OPE primarily as a comparative diagnostic tool for screening policy classes and identifying promising candidates, rather than as a guarantee of exact online ranking among adaptive learners\.
A key practical motivation for our pipeline is to provide a controlled environment in which offline\-trained policies can be stress\-tested before any live rollout\. To this end, we construct a simulator that deliberately reuses the exact ground\-truth reward functionsμa\(x\)\\mu\_\{a\}\(x\)\. The online simulation proceeds as follows:
1. 1\.Drawxtx\_\{t\}i\.i\.d\. from the same generator and compute an appropriate featurizationzt=ϕ\(xt\)z\_\{t\}=\\phi\(x\_\{t\}\);
2. 2\.For each OPE pre\-trained policyπ\\pi, chooseat∼π\(⋅∣zt\)a\_\{t\}\\sim\\pi\(\\cdot\\mid z\_\{t\}\)\(deterministic policies useargmax\\arg\\max\);
3. 3\.Sampleyt∼\(μat\(xt\)\)y\_\{t\}\\sim\\\!\\left\(\\mu\_\{a\_\{t\}\}\(x\_\{t\}\)\\right\);
4. 4\.Record instantaneous reward, cumulative reward, cumulative regret relative to the oracle, and action accuracy𝕀\[at=π⋆\(xt\)\]\\mathbb\{I\}\[a\_\{t\}=\\pi^\{\\star\}\(x\_\{t\}\)\];
5. 5\.Aggregate performance across replications and visualize policy\-level trajectories and, when relevant, low\-dimensional policy/value maps\.
These simulations constitute the final stage of the pipeline\. By comparing realized cumulative reward and regret against both the oracle and the non\-adaptive baselines, we measure how well candidate policies exploit the offline\-learned structure when the online environment matches the offline data\-generating process\. Importantly, the online stage does not relearn a policy from logged feedback; it provides evidence about internal consistency and best\-case offline\-to\-online transfer\.
## 4Results
We report results in two complementary layers\. First, offline, we use Doubly Robust OPE to compare a portfolio of pre\-specified adaptive and non\-adaptive policies under a common policy\-value and regret criterion\. Second, when ground truth is available \(synthetic data\), we deploy the same offline\-trained policies in the controlled warm\-start simulator described above\.
This two\-stage design allows us to answer three practical questions:
1. 1\.whether adaptivity is beneficial relative to fixed allocation,
2. 2\.whether OPE\-based coarse policy screening is preserved under finite\-sample online deployment, and
3. 3\.whether the performance spread among strong adaptive candidates is large enough to justify additional offline policy optimization\.
Table[1](https://arxiv.org/html/2609.30273#S4.T1)presents the Summary of Experiments that produces the results we report in this section\.NN: Number of observations in the fixed experiment \(logging policy\);NCovN\_\{Cov\}: Number of covariate columns;NBootN\_\{Boot\}: Number of Bootstraps done for OPE Confidence Intervals;TOnT\_\{On\}: Number of Online Simulation steps;NOnN\_\{On\}: Number of Online Simulations repetitions\.
Table 1:Summary of Experiments ran for results generation\.NameGround TruthNNNCovN\_\{Cov\}Treatment/ControlNBootN\_\{Boot\}OnlineTOnT\_\{On\}NOnN\_\{On\}AllocationSimulation?No ContexEq\. \([1](https://arxiv.org/html/2609.30273#S3.E1)\)3000050/50300Yes500100CategoricalEq\. \([2](https://arxiv.org/html/2609.30273#S3.E2)\)3000150/50300Yes500100LinearEq\. \([3](https://arxiv.org/html/2609.30273#S3.E3)\)3000150/50300Yes500100OscillatoryEq\. \([4](https://arxiv.org/html/2609.30273#S3.E4)\)3000150/50300Yes500100Double AverageEq\. \([5](https://arxiv.org/html/2609.30273#S3.E5)\)3000250/50300Yes500100Double ProductEq\. \([6](https://arxiv.org/html/2609.30273#S3.E6)\)3000250/50300Yes500100HillstromNone64000833\.3/33\.3/33\.3300No\-\-CriteoNone2\.5m1184\.6/15\.4300No\-\-LalondeNone614830\.1/69\.9300No\-\-
Table 2:Structural differences between the two experiment families\.### 4\.1No Context
The no\-context setting functions as a sanity check for the pipeline\. Since no covariate information is available, the optimal rule is constant and there is no personalization problem to solve\. Accordingly, the relevant question is not whether adaptive policies outperform through contextual adaptation, but whether they correctly recover the arm with the larger expected reward\. Fig\.[2\(a\)](https://arxiv.org/html/2609.30273#S4.F2.sf1)shows that the strongest adaptive policies do identify the better arm and achieve strong OPE values, but this should not be interpreted as evidence that adaptivity is intrinsically useful in the absence of context\. Rather, this case illustrates the limiting regime in which a strong fixed policy is already sufficient and any further policy optimization is unlikely to add practical value\.
### 4\.2Categorical Heterogeneity
With discrete heterogeneity, adaptive policies consistently outperform non\-adaptive baselines, as shown in Fig\.[2\(b\)](https://arxiv.org/html/2609.30273#S4.F2.sf2)\. The important pattern is not that one specific adaptive policy dominates by a large margin, but that several adaptive candidates achieve similarly strong OPE values and low regret while the fixed baselines remain clearly inferior\. This indicates that once category\-level treatment differences are present, even relatively simple context\-aware rules are sufficient to capture most of the available gain from personalization\.
### 4\.3Continuous and Structured Heterogeneity
Figs\.[2\(c\)](https://arxiv.org/html/2609.30273#S4.F2.sf3),[2\(d\)](https://arxiv.org/html/2609.30273#S4.F2.sf4), and[2\(e\)](https://arxiv.org/html/2609.30273#S4.F2.sf5)summarize the offline OPE comparisons for the continuous and structured settings, while Figs\.[3\(b\)](https://arxiv.org/html/2609.30273#S4.F3.sf2)and[3\(c\)](https://arxiv.org/html/2609.30273#S4.F3.sf3)visualize the learned policy and value structures\. Online warm\-start performance is reported in Table[3](https://arxiv.org/html/2609.30273#S4.T3)\.
\(a\)No context\.
\(b\)Categorical heterogeneity\.
\(c\)Linear heterogeneity\.
\(d\)Oscillatory heterogeneity\.
\(e\)Two covariates \(average combination\)\.
\(f\)Two covariates \(product combination\)\.
Figure 2:OPE \(DR\) and regret\.Policies are ranked by DR value \(left\) and estimated mean regret relative to the best observed policy \(right\)\. Whenever contextual heterogeneity is present, adaptive policies substantially outperform fixed baselines; among strong adaptive candidates, differences are comparatively modest\.\(a\)Categorical heterogeneity\.
\(b\)Linear heterogeneity\.
\(c\)Oscillatory heterogeneity\.
Figure 3:Policy and value functions by bins\.In the continuous settings, the learned policies recover the main decision structure implied by the oracle: a monotone switching threshold in the linear case and alternating treatment advantage in the oscillatory case\.Across the linear \(Figs\.[2\(c\)](https://arxiv.org/html/2609.30273#S4.F2.sf3),[3\(b\)](https://arxiv.org/html/2609.30273#S4.F3.sf2)\), oscillatory \(Figs\.[2\(d\)](https://arxiv.org/html/2609.30273#S4.F2.sf4),[3\(c\)](https://arxiv.org/html/2609.30273#S4.F3.sf3)\), and two\-covariate settings \(Fig\.[2\(e\)](https://arxiv.org/html/2609.30273#S4.F2.sf5)\), the dominant result is stable: adaptive policies substantially outperform the fixed baselines, and the learned policy/value surfaces recover the structure implied by the oracle\. In the linear case, the learned policies recover a monotone switching threshold; in the oscillatory case, they recover alternating treatment advantage across the covariate range; and in the two\-covariate case, they assign probability mass to the regions favored by the oracle \(full policy\-surface plots are omitted for space\)\. In all three settings, the strongest adaptive candidates cluster near the oracle benchmark, whereas the fixed baselines remain far behind\. Thus, the practically important contrast is between adaptive and non\-adaptive designs, not between the strongest adaptive candidates themselves\.
The online warm\-start results reinforce this interpretation, as visible in Table[3](https://arxiv.org/html/2609.30273#S4.T3)and in Figure[4](https://arxiv.org/html/2609.30273#S4.F4): the top adaptive policies remain close to the oracle across all structured HTE settings, while the fixed baselines accumulate substantially larger regret\. When comparing the average and product combinations in the two\-covariate setup, we do not observe meaningful differences in policy ordering or qualitative fit; the overall mean reward changes, but the main conclusions remain the same\.
\(a\)No heterogeneity\.
\(b\)Categorical heterogeneity\.
\(c\)Linear heterogeneity\.
\(d\)Oscillatory heterogeneity\.
\(e\)Two covariates \(average combination\)\.
\(f\)Two covariates \(product combination\)\.
Figure 4:Warm\-start online cumulative metrics \(T=500,N=100T\{=\}500,N\{=\}100\)\.Left: cumulative reward; right: cumulative regret relative to the oracle\. Several adaptive policies remain close to the oracle and far outperform the fixed baselines\. The dominant contrast is between adaptive and non\-adaptive designs; among strong adaptive policies, realized differences are modest and partly attributable to finite\-sample variability\.\(a\)Always A policy\.
\(b\)Always B policy\.
\(c\)Uniform policy\.
\(d\)Oracle policy\.
\(e\)ε\\varepsilon\-Greedy policy\.
\(f\)Bootstrapped Thompson Sampling policy\.
Figure 5:Two\-covariate with average combination policy and value surfaces\.Panels showπ\(x1,x2\)\\pi\(x\_\{1\},x\_\{2\}\)\(left\) andVπ\(x1,x2\)V\_\{\\pi\}\(x\_\{1\},x\_\{2\}\)\(right\) for each policy\. Adaptive policies concentrate probability mass in regions favored by the oracle, while fixed baselines fail to exploit the two\-dimensional heterogeneity structure\. Softmax and UCB policies offers plots indistinguishable from the plots for Epsilon\-Greedy policy\.\(a\)Always A policy\.
\(b\)Always B policy\.
\(c\)Uniform policy\.
\(d\)Oracle policy\.
\(e\)ε\\varepsilon\-Greedy policy\.
\(f\)Bootstrapped Thompson Sampling policy\.
Figure 6:Two\-covariate with product combination policy and value surfaces\.Table 3:Warm\-start online simulations \(T=500T\{=\}500timesteps, average overN=100N=100simulations\)\.Accuracy is the fraction of oracle\-consistent actions; reward and regret are reported both per\-step and cumulatively\. Across heterogeneous settings, strong adaptive policies remain close to the oracle and far outperform the fixed baselines\.\(a\)No Covariates
\(b\)Categorical
\(c\)Linear
\(d\)Oscillatory
\(e\)Two covariates with average combination
\(f\)Two covariates with product combination
### 4\.4Validating on Open Datasets
To assess whether the same decision logic carries over beyond the synthetic experiments, we apply the OPE pipeline to three open benchmarks collected under fixed treatment assignment\. Since no ground\-truth reward surface is available in these datasets, we use only the offline stage of the framework\. Before doing so, we generalize the implementation to handle binary and numeric outcomes, multiple treatment arms, and automatic covariate handling\.
Taken together, the open\-dataset results show that OPE can serve as a practical deployment filter:
- •Hillstrom \(Fig\.[7\(a\)](https://arxiv.org/html/2609.30273#S4.F7.sf1)\):weak adaptivity signal\. The uncertainty bands of the top adaptive policies overlap substantially with those of the best fixed alternatives, so the available evidence does not justify the added complexity of adaptive deployment\.
- •Criteo \(Fig\.[7\(d\)](https://arxiv.org/html/2609.30273#S4.F7.sf4)\):strong adaptivity signal\. Adaptive policies achieve consistently higher estimated value and lower regret than non\-adaptive baselines, indicating a setting in which contextual adaptation is likely to be worth deploying\.
- •LaLonde \(Fig\.[7\(f\)](https://arxiv.org/html/2609.30273#S4.F7.sf6)\):overlap\-limited regime\. The best adaptive policies are not clearly separated from strong fixed alternatives, suggesting that the available support is insufficient for adaptivity to produce a decisive practical gain in this setting\.
These three cases reinforce that the value of adaptivity is not universal: it depends on the strength of contextual heterogeneity and on the quality of support provided by the data\.
\(a\)Hillstrom Datasetusing conversion as outcome\.
\(b\)Hillstrom Datasetusing visit as outcome\.
\(c\)Hillstrom Datasetusing spend as outcome\.
\(d\)Criteo Datasetusing conversion as outcome\.
\(e\)Criteo Datasetusing visit as outcome\.
\(f\)Lalonde Dataset\.
Figure 7:OPE \(DR\) and regret for open Datasets\.For Hillstrom and Lalonde, adaptive policies do not clearly separate from the strongest fixed alternatives\. For Criteo, adaptive policies consistently outperform non\-adaptive baselines\.Overall, the results support two interpretations\. First, OPE successfully distinguishes settings in which contextual adaptation is worth deploying from settings in which a strong fixed policy is already sufficient\. Second, when adaptive policies are beneficial, the warm\-start simulation shows that several reasonable candidates already operate close to the oracle benchmark from the start of online interaction\.
## 5Discussion
Our results support a decision\-oriented interpretation of adaptive experimentation\. Rather than asking only which policy achieves the highest estimated value, our pipeline addresses a more practical question: whether historical A/B test data provide sufficient evidence that contextual adaptation is worth deploying at all, and, if so, whether the performance spread among strong adaptive candidates is large enough to justify deployment\. Viewed through this lens, the main empirical contrast is not between the best two adaptive policies, but between adaptive and non\-adaptive designs\. Across synthetic settings with genuine heterogeneous treatment effects, adaptive policies consistently improve upon fixed baselines and remain close to the oracle benchmark; by contrast, in settings with weak or absent heterogeneity, the gains from adaptation are negligible\. This pattern clarifies the practical role of OPE in the regimes considered here\. This observation is directly relevant to the role of offline policy learning: since our objective is to compare the full policy portfolio rather than to optimize a single policy, selecting a strong candidate via OPE may already recover most of the available value, with limited marginal benefit from an additional policy\-learning step\.
The simulator\-based online stage should be interpreted as a matched best\-case benchmark for offline\-to\-online transfer rather than as a realistic proxy for deployment under arbitrary conditions\. By design, it isolates how much of the offline\-learned structure survives the transition to interaction when the data\-generating process is stable and shared across stages\. Our results show that OPE provides a reliable coarse deployment signal under this matched regime — clearly for broad contrasts \(adaptive versus non\-adaptive designs, or strong candidates versus obviously weak baselines\), though not necessarily for fine\-grained ranking among history\-dependent adaptive learners\. The present experiments therefore do not establish robustness to offline\-online mismatch; they evidence internal consistency and best\-case warm\-start transfer\.
The open benchmarks illustrate three distinct deployment regimes\.Hillstromrepresents a relatively clean randomized benchmark with modest heterogeneity: the top adaptive policies are not clearly separated from the best fixed alternatives, so the evidence does not strongly support the additional complexity of adaptive deployment\.Criteo, by contrast, is a large\-scale benchmark with meaningful covariate\-driven heterogeneity, and here adaptive policies consistently achieve higher estimated value and lower regret than non\-adaptive baselines; this is precisely the type of setting in which our framework would recommend adaptive experimentation\.LaLondeoccupies a third regime, in which selection bias and overlap limitations complicate counterfactual comparison; here the policy\-value perspective remains informative, but the results also highlight that poor support can limit the practical gains obtainable from adaptive decision rules\. Taken together, these benchmarks reinforce that the value of adaptivity is not universal: it depends on the strength of contextual heterogeneity and on the quality of support provided by the data\.
Our implementation also computes policy\-value estimates produced by the three OPE estimators \(DR, IPS, DM\)\. Across all considered scenarios, DR either achieves the highest estimated value or is statistically indistinguishable from the best\-performing estimator, as evidenced by overlapping confidence intervals\. If we compare the different policies estimators for each policy in the synthetic scenarios, we have a total of 38 combinations\. In the aggregate, across the 38 scenario\-policy combinations considered, DR attains the largest point estimate in 8 cases, its confidence interval overlaps that of the best estimator in all 38 cases, and we observe no case in which DR is clearly dominated by a competing estimator through non\-overlapping intervals\. This supports the use of DR as our primary estimator and suggests that the main qualitative conclusions are robust across standard OPE choices\.
While poor overlap can lead to unstable inverse\-propensity weights and inflated finite\-sample variance in weighting\-based and doubly robust estimators, this issue is most acute when propensity scores approach the boundaries of the unit interval\[[11](https://arxiv.org/html/2609.30273#bib.bib23),[21](https://arxiv.org/html/2609.30273#bib.bib24),[22](https://arxiv.org/html/2609.30273#bib.bib25)\]\. In our synthetic experiments, however, the logging propensities are known and uniformly bounded, so the heavy\-weight regime that motivates these concerns does not arise\. In the real\-data benchmarks, we likewise do not observe instability in the bootstrap estimates or qualitative disagreement between DR, IPS, and DM, suggesting that heavy\-tailed behavior does not dominate in the regimes considered here\.
Our study is related to, but distinct from, the classical tradeoff between reward maximization and best\-arm identification in adaptive experiments\[[7](https://arxiv.org/html/2609.30273#bib.bib28)\]\. The candidate contextual bandit policies we evaluate online are primarily reward\-oriented: they are assessed through cumulative reward and regret and are therefore closer in spirit to regret minimization than to pure best\-arm identification\. At the same time, the synthetic experiments provide an oracle policyπ⋆\(x\)\\pi^\{\\star\}\(x\), which allows us to inspect whether these reward\-oriented policies recover the correct context\-dependent decision structure\. Accordingly, the plots ofπ\(x\)\\pi\(x\)versusxxshould be interpreted as an ex post diagnostic of how closely the candidate policies approximate the oracle decision rule, not as evidence that the adaptive design itself targets a best\-arm\-identification objective\. Our contribution is therefore not to resolve the welfare\-versus\-identification tradeoff within an adaptive experiment, but to use offline policy evaluation as a pre\-deployment tool for deciding whether adaptive experimentation is worth pursuing and which reward\-oriented policies are promising candidates\.
## 6Conclusion and Future Work
In this work, we studied how historical data from fixed randomized experiments \(A/B tests\) can be used to inform the deployment of adaptive experiments based on contextual bandits\. Rather than optimizing a single policy from logged feedback, we focused on policy\-level comparison: determining whether contextual adaptation is worth deploying and which adaptive policies are reasonable candidates given the available evidence\. To this end, we combined off\-policy evaluation with a controlled warm\-start simulation\.
Across the settings considered here, our results show that adaptive policies improve upon fixed allocation when meaningful contextual heterogeneity is present, while offering little benefit in its absence\. More broadly, the paper shows that historical A/B test data can be used not only to evaluate candidate adaptive policies offline, but also to support deployment decisions and warm\-start initialization before live interaction\.
Our conclusions should be interpreted within the scope of the present study\. First, we do not perform offline policy learning; our aim is comparative evaluation of a pre\-specified policy portfolio rather than optimization over a large policy class\. Second, the warm\-start simulator is intentionally a matched best\-case benchmark; consequently, the present results do not establish robustness to outcome\-model misspecification, covariate shift, or temporal drift\. Third, our setting is contextual bandits rather than full sequential reinforcement learning, so the conclusions concern one\-step treatment assignment rather than long\-horizon control\. These scope choices are deliberate: they isolate the question the paper is meant to answer, namely whether existing A/B test data contain enough evidence to justify adaptive deployment and to select a strong adaptive policy candidate\.
Several directions naturally extend this study and emerges as potential future works\. First, while the present results suggest limited incremental value from offline policy learning in the regimes considered here, larger policy classes and richer action spaces may change this picture\. Second, extending the warm\-start framework to non\-stationary or partially observable environments would bring the analysis closer to realistic deployment conditions\. Third, incorporating operational constraints such as fairness, risk sensitivity, or budget limits would broaden the practical relevance of the framework\. Finally, additional large\-scale logged datasets or controlled live experiments would help evaluate the robustness of OPE\-based deployment decisions in more complex production settings\.
#### Acknowledgments
The authors would like to thank Daniel Vieira Batista, Marco Antonio Afonso Aragon, Adriana Laurindo Monteiro and Gabriel Mattos Langeloh for inspiring discussions that motivated the initial direction of this work\.
#### Disclosure of Interests
Any opinions, findings, conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of Itaú Unibanco and Instituto de Ciência e Tecnologia Itaú\. This document is not and does not constitute or intend to constitute investment advice or any investment service\. It is not and should not be deemed to be an offer to purchase or sell, or a solicitation of an offer to purchase or sell, or a recommendation to purchase or sell any securities or other financial instruments\. In addition, all data used in this study comply with the Brazilian General Data Protection Law\.
## References
- \[1\]A\. Bietti, A\. Agarwal, and J\. Langford\(2021\)A contextual bandit bake\-off\.Journal of Machine Learning Research22\(133\),pp\. 1–49\.Cited by:[§1](https://arxiv.org/html/2609.30273#S1.p3.1),[§2](https://arxiv.org/html/2609.30273#S2.p7.1)\.
- \[2\]R\. H\. Dehejia and S\. Wahba\(2002\-02\)Propensity score matching methods for non\-experimental causal studies\.Review of Economics and Statistics84\(1\),pp\. 151–161\.Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p5.3)\.
- \[3\]E\. Diemert, A\. Betlei, C\. Renaudin, and M\. Amini\(2018\)A large scale benchmark for uplift modeling\.InProceedings of the AdKDD and TargetAd Workshop, KDD,London, United Kingdom\.Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p5.3),[§2](https://arxiv.org/html/2609.30273#S2.p6.1)\.
- \[4\]M\. Dudík, D\. Erhan, J\. Langford, and L\. Li\(2014\)Doubly robust policy evaluation and optimization\.Statistical Science29\(4\),pp\. 485–511\.External Links:ISSN 08834237, 21688745Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p1.7),[§2](https://arxiv.org/html/2609.30273#S2.p2.2),[§2](https://arxiv.org/html/2609.30273#S2.p3.2),[§2](https://arxiv.org/html/2609.30273#S2.p6.1),[§3](https://arxiv.org/html/2609.30273#S3.SS0.SSS0.Px5.p4.2)\.
- \[5\]P\. Gutierrez and J\. Gérardy\(2017\)Causal inference and uplift modelling: a review of the literature\.InProceedings of The 3rd International Conference on Predictive Applications and APIs,Proceedings of Machine Learning Research, Vol\.67,pp\. 1–13\.External Links:[Link](https://proceedings.mlr.press/v67/gutierrez17a.html)Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p6.1)\.
- \[6\]V\. Hadad, D\. A\. Hirshberg, R\. Zhan, S\. Wager, and S\. Athey\(2021\)Confidence intervals for policy evaluation in adaptive experiments\.Proceedings of the National Academy of Sciences118\(15\),pp\. e2014602118\.Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p6.1),[§3](https://arxiv.org/html/2609.30273#S3.SS0.SSS0.Px5.p4.2)\.
- \[7\]V\. Hadad, L\. R\. Rosenzweig, S\. Athey, and D\. Karlan\(2021\-03\)Practitioner’s guide: designing adaptive experiments\.Technical reportGolub Capital Social Impact Lab\.Note:March 2021Cited by:[§1](https://arxiv.org/html/2609.30273#S1.p2.5),[§5](https://arxiv.org/html/2609.30273#S5.p6.3)\.
- \[8\]K\. Hillstrom\(2008\)The minethatdata e\-mail analytics and data mining challenge\.Note:Blog postAccessed: 2026\-03\-20Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p5.3)\.
- \[9\]N\. Kallus and M\. Uehara\(2020\)Double reinforcement learning for efficient and robust off\-policy evaluation\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 5078–5088\.Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p2.2)\.
- \[10\]R\. LaLonde\(1986\)Evaluating the econometric evaluations of training programs\.American Economic Review76\(4\),pp\. 604–620\.Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p5.3)\.
- \[11\]F\. Li, L\. E\. Thomas, and F\. Li\(2019\)Addressing extreme propensity scores via the overlap weights\.American Journal of Epidemiology188\(1\),pp\. 250–257\.Cited by:[§5](https://arxiv.org/html/2609.30273#S5.p5.1)\.
- \[12\]L\. Li, W\. Chu, J\. Langford, and R\. E\. Schapire\(2010\)A contextual\-bandit approach to personalized news article recommendation\.InProceedings of the 19th International Conference on World Wide Web \(WWW 2010\),pp\. 661–670\.Cited by:[§1](https://arxiv.org/html/2609.30273#S1.p3.1)\.
- \[13\]A\. Molak\(2023\)Heterogeneous treatment effects with experimental data – the uplift odyssey\.InCausal Inference and Discovery in Python: Unlock the Secrets of Modern Causal Machine Learning with DoWhy, EconML, PyTorch and More,Note:Accessed: 2026\-03\-20External Links:ISBN 9781804612989Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p5.3)\.
- \[14\]P\. Rzepakowski and S\. Jaroszewicz\(2012\)Decision trees for uplift modeling with single and multiple treatments\.Knowledge and Information Systems32\(2\),pp\. 303–327\.External Links:[Document](https://dx.doi.org/10.1007/s10115-011-0434-0)Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p6.1)\.
- \[15\]Y\. Saito, S\. Aihara, M\. Matsutani, and Y\. Narita\(2021\)Open bandit dataset and pipeline: towards realistic and reproducible off\-policy evaluation\.InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks \(NeurIPS Datasets and Benchmarks 2021\),Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p6.1),[§2](https://arxiv.org/html/2609.30273#S2.p7.1)\.
- \[16\]st\-tech/zr\-obp contributors\(2026\)Simulator module\.Note:[https://deepwiki\.com/st\-tech/zr\-obp/5\-simulator\-module](https://deepwiki.com/st-tech/zr-obp/5-simulator-module)Accessed: 2026\-06\-25Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p7.1)\.
- \[17\]Y\. Su, M\. Dimakopoulou, A\. Krishnamurthy, and M\. Dudík\(2020\)Doubly robust off\-policy evaluation with shrinkage\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 9167–9176\.External Links:[Link](https://proceedings.mlr.press/v119/su20a.html)Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p2.2),[§2](https://arxiv.org/html/2609.30273#S2.p6.1)\.
- \[18\]A\. Swaminathan and T\. Joachims\(2015\)The self\-normalized estimator for counterfactual learning\.InAdvances in Neural Information Processing Systems,C\. Cortes, N\. Lawrence, D\. Lee, M\. Sugiyama, and R\. Garnett \(Eds\.\),Vol\.28,pp\.\.Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p1.7),[§2](https://arxiv.org/html/2609.30273#S2.p2.2),[§2](https://arxiv.org/html/2609.30273#S2.p3.2),[§2](https://arxiv.org/html/2609.30273#S2.p6.1),[§3](https://arxiv.org/html/2609.30273#S3.SS0.SSS0.Px5.p4.2)\.
- \[19\]B\. van den Akker, N\. Weber, F\. Moraes, and D\. Goldenberg\(2022\)Extending open bandit pipeline to simulate industry challenges\.arXiv preprint arXiv:2209\.04147\.External Links:[Link](https://arxiv.org/abs/2209.04147)Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p7.1)\.
- \[20\]Y\. Wang, A\. Agarwal, and M\. Dudík\(2017\)Optimal and adaptive off\-policy evaluation in contextual bandits\.InProceedings of the 34th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.70,pp\. 3589–3597\.Cited by:[§2](https://arxiv.org/html/2609.30273#S2.p1.7),[§2](https://arxiv.org/html/2609.30273#S2.p2.2),[§2](https://arxiv.org/html/2609.30273#S2.p6.1),[§3](https://arxiv.org/html/2609.30273#S3.SS0.SSS0.Px5.p4.2)\.
- \[21\]C\. Yang, L\. E\. Thomas, and F\. Li\(2026\)Demystify doubly\-robust estimation: the role of overlap\.arXiv preprint arXiv:2602\.01648\.Cited by:[§5](https://arxiv.org/html/2609.30273#S5.p5.1)\.
- \[22\]M\. Zhang and B\. Zhang\(2022\)A stable and more efficient doubly robust estimator\.Statistica Sinica32,pp\. 1143–1163\.Cited by:[§5](https://arxiv.org/html/2609.30273#S5.p5.1)\.Similar Articles
When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits
This paper presents a diagnostic protocol for selecting rewards and policies in delayed-feedback contextual bandits, arguing that standard offline evaluation can mislead and validating the approach on benchmarks and a deployed push system.
Certifying when decision-time information justifies adaptive experimentation
This paper introduces Opal (Opportunity-aware Policy Authorization for Laboratories), a framework that certifies whether adaptive experimentation should be enabled by precommitting to non-trivial adaptation, controlled target risk, and positive executed value after cost. It establishes an impossibility boundary and demonstrates the method on a Cell Painting dataset, achieving risk control and positive value.
Off-Policy Evaluation with Strategic Agents via Local Disclosure
This paper studies off-policy evaluation (OPE) when decision subjects (agents) strategically modify their covariates in response to a policy. It proposes a method that uses local disclosure via post-hoc explanations to reveal agents' pre-strategic covariates and construct a doubly robust estimator for policy value.
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits
This paper introduces cross-domain off-policy evaluation and learning (OPE/L) for contextual bandits, allowing the use of logged data from multiple source domains to improve policy evaluation and learning in target domains with challenging conditions like few-shot data, deterministic logging policies, and new actions.
A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing
This paper proposes A/B Agent, a closed-loop agent framework that organizes historical A/B testing knowledge into a hierarchical experience tree, retrieves transferable strategies via multi-path Tree-RAG, and self-evolves through online experiment feedback, achieving a 4.829% GMV improvement in a short-video e-commerce recommendation system.