UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation

arXiv cs.LG 论文

摘要

UpliftBench is a benchmark paper showing that disagreements between uplift modeling evaluations often stem from metric choice rather than model quality, identifying specific mismatches between ranking metrics and deployment objectives across several dataset families.

arXiv:2608.00915v1 Announce Type: new Abstract: Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about metrics, not models. UpliftBench evaluates 12 uplift estimators under an outer-test-isolated, multi-objective protocol across seven dataset families; its two findings are identified where a reference objective exists -- F1 on the standard continuous benchmark (IHDP), F2 in a within-sample case study on Jobs. On that benchmark, Qini shows no detectable alignment with effect accuracy -- across all 100 IHDP realizations its mean rank correlation with effect accuracy is +0.07, 95% CI [-0.03, +0.16] -- while AUUC is consistently more aligned (paired prefix-mean-AUUC-over-Qini gap +0.49 [+0.40, +0.59]; the shipped cumulative-gain AUUC aligns better still, +0.73). On Jobs, ranking metrics are structurally insufficient for a sign-threshold policy because they discard the score level; empirically, within the released split-rotation analysis direct policy-risk selection yields lower benchmark regret than random model selection while Qini, AUUC, and uplift-at-$k$ do not (14-15% regret). Calibrating the decision threshold removes 81% of the Qini-selection regret. Both findings are bounded, not universal: F1 is not detected on either validation family (the ACIC and Revenue-Synthetic gaps are both indistinguishable from zero), and F2 vanishes under a budgeted-value objective where rank suffices. UpliftBench releases versioned loaders, fixed protocols, result artifacts, and a reproducible living leaderboard; the public repository accompanies the paper.
查看原文
查看缓存全文

缓存时间: 2026/08/04 07:44

# UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation
Source: [https://arxiv.org/html/2608.00915](https://arxiv.org/html/2608.00915)
###### Abstract\.

Uplift modeling \(conditional\-average\-treatment\-effect estimation\) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about*metrics*, not models\. UpliftBench evaluates 12 uplift estimators under an outer\-test\-isolated, multi\-objective protocol across seven dataset families; its two findings are identified where a reference objective exists — F1 on the standard continuous benchmark \(IHDP\), F2 in a within\-sample case study on Jobs\. On that benchmark, Qini shows no detectable alignment with effect accuracy — across all100100IHDP realizations its mean rank correlation with effect accuracy is\+0\.07\+0\.07, 95% CI\[−0\.03,\+0\.16\]\[\-0\.03,\+0\.16\]— while AUUC is consistently more aligned \(paired prefix\-mean\-AUUC\-over\-Qini gap\+0\.49\+0\.49\[\+0\.40,\+0\.59\+0\.40,\+0\.59\]; the shipped cumulative\-gain AUUC aligns better still,\+0\.73\+0\.73\)\. On Jobs, ranking metrics are structurally insufficient for a sign\-threshold policy because they discard the score level; empirically, within the released split\-rotation analysis direct policy\-risk selection yields lower benchmark regret than random model selection while Qini, AUUC, and uplift\-at\-kkdo not \(1414–15%15\\%regret\)\. Calibrating the decision threshold removes81%81\\%of the Qini\-selection regret\. Both findings are bounded, not universal: F1 is not detected on either validation family \(the ACIC and Revenue\-Synthetic gaps are both indistinguishable from zero\), and F2 vanishes under a budgeted\-value objective where rank suffices\. UpliftBench releases versioned loaders, fixed protocols, result artifacts, and a reproducible living leaderboard; the public repository accompanies the paper\.

uplift modeling, heterogeneous treatment effects, causal machine learning, evaluation metrics, benchmark design, model selection, policy evaluation, Qini coefficient

††ccs:Computing methodologies Machine learning††ccs:Computing methodologies Causal reasoning and diagnostics## 1\.Introduction

Every empirical benchmark ranks models by a metric, and the choice of metric quietly encodes a*scientific question*— the accuracy of the estimated effect, the quality of the induced ranking, or the value of the resulting decision\. When the metric is mismatched to the outcome type or to the deployment objective these questions come apart, and the “best” model changes with the measure rather than with the model\. This is a benchmarking and evaluation\-methodology problem, not a modeling one; it is central to how causal\-ML methods are compared and selected, and it is acute in uplift modeling \(conditional\-average\-treatment\-effect estimation for targeting\), where published benchmarks disagree on which estimator wins and no counterfactual is observed to adjudicate\.

Concretely, a retailer’s data team must choose an uplift model for a campaign\. They cannot observe, for any customer, the counterfactual outcome — the fundamental problem of causal inference — so they rank candidate models by a*proxy*computed from experimental data, usually the Qini coefficient or the area under the uplift curve \(AUUC\)\. The implicit assumption is that a model scoring well on the proxy also serves the deployment goal\.

We show this assumption fails in two structurally different ways that cannot be explained solely by a single unstable estimator or base learner\.They are not fold\-level noise; each has a distinct empirical signature and a different practical remedy, and conflating them \(as a single “rankings are unstable” observation would\) obscures both\. Throughout, we distinguish three levels of claim — exact identities, empirical signatures, and candidate mechanisms — and hold each finding to its level\. The evidence base is deliberately scoped: F1 is identified on the continuous benchmark with known effects — the semi\-synthetic IHDP benchmark\(Hill,[2011](https://arxiv.org/html/2608.00915#bib.bib14)\)\(10 realizations in the main panel, extended to all 100 simulations\) — and F2 on Jobs\(LaLonde,[1986](https://arxiv.org/html/2608.00915#bib.bib23); Shalit et al\.,[2017](https://arxiv.org/html/2608.00915#bib.bib32)\), the 10 re\-splits of the LaLonde job\-training study; every interval is read at that granularity\.What “regime” denotes here:an*observed dataset family in which the failure appears*, not a measurable property that predicts it\. No ex\-ante quantity we test predicts the metric\-specific*gap*within a family \(full search, including a post\-hoc cross\-family observation, in Appendix[E\.1](https://arxiv.org/html/2608.00915#A5.SS1)\), and*the failing family is not the heavy\-tailed one*\(Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)\)\. The operational diagnostic is cross\-metric stability on the user’s own data \(Section[6](https://arxiv.org/html/2608.00915#S6)\)\.

##### Finding 1 \(F1\) — outcome\-regime sensitivity of Qini\.

The unnormalized Qini coefficient sums outcomes, a construction designed for binary response\(Radcliffe and Surry,[2007](https://arxiv.org/html/2608.00915#bib.bib30)\); on continuous outcomes its reliability can become outcome\-distribution dependent\. We do not claim it is categorically invalid there; we show that on the evaluated continuous benchmark \(IHDP\) its ranking fails to track ground\-truth quality \(Section[5\.1](https://arxiv.org/html/2608.00915#S5.SS1)\), while the mean\-based area under the uplift curve \(AUUC\) stays informative\. This is practically relevant and potentially silent: continuous \(e\.g\. revenue\) outcomes are an active uplift setting\(He et al\.,[2024](https://arxiv.org/html/2608.00915#bib.bib13)\), and although scikit\-uplift requires a binary outcome\(Maksimov et al\.,[2020](https://arxiv.org/html/2608.00915#bib.bib27)\), the widely used causalml package\(Chen et al\.,[2020](https://arxiv.org/html/2608.00915#bib.bib4)\)computes Qini on a continuous outcome with no warning\.

##### Finding 2 \(F2\) — objective mismatch \(the metric measures the wrong goal\)\.

Even where ranking metrics are well constructed and agree with one another \(binary outcomes\), they can diverge from the*deployment objective*\. On Jobs, every ranking metric we test correlates negatively with RCT\-estimated policy value−Rpolicy\-R\_\{\\mathrm\{policy\}\}— equivalently, positively with policy risk\. The sharper statement is a selector contrast: Jobs carries policy\-selection signal that direct risk selection captures and the ranking metrics do not — metric selection is indistinguishable from random, at a1414–15%15\\%regret versus the risk\-selected candidate \(Section[5\.2](https://arxiv.org/html/2608.00915#S5.SS2)\)\. This gap echoes concurrent observations under injected structural bias\(Yang et al\.,[2026](https://arxiv.org/html/2608.00915#bib.bib38)\); we quantify it under an outer\-test\-isolated111No test\-fold information enters any reported number; the one disclosed deviation from strict*inner*nesting involves no outcomes and affects hyperparameter selection only \(Section[7](https://arxiv.org/html/2608.00915#S7)\)\.protocol with per\-regime uncertainty\.

##### Why this matters and what is new\.

This is not merely “the wrong gold standard”: our six\-metric design includes operational objectives \(policy value, policy risk, uplift\-at\-kk\), and the disagreements survive them\. Unlike prior work \(Section[2](https://arxiv.org/html/2608.00915#S2)\), UpliftBench connects production uplift proxies to reference effect and policy objectives*across distinct identification regimes*under a single outer\-test\-isolated protocol — the binary\-vs\-continuous Qini contrast and the regime boundary \(one of three continuous families fails, and not the heavy\-tailed one\) are observable only in such a multi\-regime design — and quantifies the operational model\-selection cost\. The distinction between effect estimation, ranking quality, and deployment value is a broader principle for benchmark design: choose the metric to match the question it must answer\.

##### Contributions\.

1. \(1\)Benchmark specification \(UpliftBench\)\.An outer\-test\-isolated benchmark of 12 uplift estimators over 25 instances from seven dataset families, scored by six primary objectives \(three ranking metrics, policy value,PEHE\\sqrt\{\\mathrm\{PEHE\}\}, and policy risk\) plus a calibration diagnostic, under repeated stratified cross\-validation with nested tuning and realization\-level uncertainty \(Sections[3](https://arxiv.org/html/2608.00915#S3)–[5](https://arxiv.org/html/2608.00915#S5)\)\.
2. \(2\)Reusable implementation and artifact\.Extensible dataset loaders, estimator and metric interfaces, a standardized result\-parquet schema stamped with git hash and config, tests, and documented reproduction targets; released with a living leaderboard, and a DOI\-bearing archival release will accompany the camera\-ready \(Section[9](https://arxiv.org/html/2608.00915#S9)\)\.
3. \(3\)Benchmark finding F1 — outcome\-regime sensitivity\.On IHDP, the Qini ranking shows no detectable alignment with effect accuracy \(across all100100IHDP realizations:\+0\.07\+0\.07, 95% CI\[−0\.03,\+0\.16\]\[\-0\.03,\+0\.16\]\) while AUUC and uplift\-at\-kkremain aligned; the divergence is affine\-invariant, has a treated\-count\-weighting structural pathway \(Lemma[5\.1](https://arxiv.org/html/2608.00915#S5.Thmtheorem1)\), and is not attributable to a single estimator\. It is not a tail effect, and it is not detected on either validation family — the ACIC andRevenue\-Syntheticgaps are both indistinguishable from zero \(Section[5\.1](https://arxiv.org/html/2608.00915#S5.SS1), Fig\.[1](https://arxiv.org/html/2608.00915#S5.F1)\)\.
4. \(4\)Case\-study finding F2 — objective mismatch on Jobs\.In a within\-sample Jobs case study, direct policy\-risk selection yields lower benchmark regret than random model selection while Qini, AUUC, and uplift\-at\-kkdo not — a cross\-repeat policy\-selection regret with selection and evaluation separated across repetitions \(Section[5\.2](https://arxiv.org/html/2608.00915#S5.SS2)\)\.
5. \(5\)Robustness and maintenance\.F1 replicates across all 100 IHDP realizations and survives estimator exclusions and implementation variants; both findings replicate under a second base learner \(XGBoost\) \(Table[6](https://arxiv.org/html/2608.00915#A7.T6), Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)\); the benchmark is versioned and maintained as a public leaderboard \(Section[9](https://arxiv.org/html/2608.00915#S9)\)\.

## 2\.Related Work

##### Uplift / CATE estimation\.

Estimator families include meta\-learners\(Künzel et al\.,[2019](https://arxiv.org/html/2608.00915#bib.bib22); Nie and Wager,[2021](https://arxiv.org/html/2608.00915#bib.bib29); Kennedy,[2020](https://arxiv.org/html/2608.00915#bib.bib20)\), class\-transformation methods\(Jaskowski and Jaroszewicz,[2012](https://arxiv.org/html/2608.00915#bib.bib16); Kane et al\.,[2014](https://arxiv.org/html/2608.00915#bib.bib17)\), tree\-based uplift models\(Rzepakowski and Jaroszewicz,[2012](https://arxiv.org/html/2608.00915#bib.bib31); Guelman et al\.,[2015](https://arxiv.org/html/2608.00915#bib.bib11)\), and causal forests\(Wager and Athey,[2018](https://arxiv.org/html/2608.00915#bib.bib35)\)\. Surveys\(Gutierrez and Gérardy,[2017](https://arxiv.org/html/2608.00915#bib.bib12); Devriendt et al\.,[2018](https://arxiv.org/html/2608.00915#bib.bib8)\)catalogue them without a unified evaluation\.

##### Benchmarks\.

Devriendt et al\.\([2018](https://arxiv.org/html/2608.00915#bib.bib8)\)compare uplift models on marketing data by Qini without nested CV or ground truth;Devriendt et al\.\([2021](https://arxiv.org/html/2608.00915#bib.bib7)\)argue for uplift over churn prediction\.Mahajan et al\.\([2024](https://arxiv.org/html/2608.00915#bib.bib26)\)carefully study CATE model\-*selection*criteria on semi\-synthetic data but do not contrast the proxy metrics used in marketing against reference objectives across regimes\. IHDP is conventionally scored byPEHE\\sqrt\{\\mathrm\{PEHE\}\}\(Hill,[2011](https://arxiv.org/html/2608.00915#bib.bib14); Shalit et al\.,[2017](https://arxiv.org/html/2608.00915#bib.bib32)\); marketing benchmarks use Qini/AUUC because no counterfactual exists\. We connect the two and show*which*proxy fails*how*\. Concerns about leakage in ML evaluation are broad\(Kapoor and Narayanan,[2023](https://arxiv.org/html/2608.00915#bib.bib18)\); our point is sharper — the choice of metric, not only the protocol, can reverse conclusions\.

##### Uplift evaluation metrics\.

A parallel line of work scrutinizes the metrics themselves\.Zhu et al\.\([2025](https://arxiv.org/html/2608.00915#bib.bib39)\)andVerbeken et al\.\([2025](https://arxiv.org/html/2608.00915#bib.bib34)\)identify limitations of Qini\-style curves on binary outcomes and propose replacements \(the Principled Uplift Curve and pROCini\);Bokelmann and Lessmann \([2024](https://arxiv.org/html/2608.00915#bib.bib3)\)reduce the variance of Qini estimates on RCT data\. Closest to our work,Yang et al\.\([2026](https://arxiv.org/html/2608.00915#bib.bib38)\)study metric*stability*and model*robustness*under injected structural biases \(selection, spillover, confounding\) on a single semi\-synthetic family, and likewise observe that targeting and effect\-estimation can be distinct objectives and that ATE\-aligned metrics rank more consistently\. We differ in question and scope: rather than proposing a replacement metric or perturbing bias, we measure how metric*choice*changes*benchmark conclusions*across*outcome regimes*and 25 instances from seven families, which isolates the binary\-vs\-continuous Qini contrast a continuous\-only design cannot reveal\. We therefore do not claim the targeting\-versus\-objective observation as wholly new; our contribution is its*benchmarked decomposition*\. Table[3](https://arxiv.org/html/2608.00915#A1.T3)itemizes the delta axis by axis\.

## 3\.Datasets

Table 1\.Benchmark datasets\.nn: sample size used \(Hillstromis used in full;Lenta/X5/MegaFonare capped at 10K by stratified subsampling, full size in parentheses\)\.π1\\pi\_\{1\}: treatment fraction\.y¯\\bar\{y\}: outcome mean \(a base rate where the outcome is binary; on the outcome scale for the continuous IHDP outcome; its 2–41 entry is the range ofy¯\\bar\{y\}across the 10 realizations; IHDPnnis the 672\-unit train portion of the standard 747\-unit release\)\.†Supplement, outside the 25\-instance accounting \(Appendix[I](https://arxiv.org/html/2608.00915#A9)\)\.Table[1](https://arxiv.org/html/2608.00915#S3.T1)lists the datasets, which we introduce here since several are used only by name below\.IHDPis a semi\-synthetic benchmark that pairs real covariates from the Infant Health and Development Program with simulated potential outcomes, so individual treatment effects are known\(Hill,[2011](https://arxiv.org/html/2608.00915#bib.bib14)\); we use 10 of the 100 standard simulation realizations released byShalit et al\.\([2017](https://arxiv.org/html/2608.00915#bib.bib32)\)\.Jobscombines the randomized LaLonde employment\-training experiment\(LaLonde,[1986](https://arxiv.org/html/2608.00915#bib.bib23)\)with observational PSID controls; policy value is evaluated on the randomized experimental subset, and the ten benchmark units are released re\-splits of that same sample\(Shalit et al\.,[2017](https://arxiv.org/html/2608.00915#bib.bib32)\)\.Syntheticis generated by our own documented data\-generating process \(DGP; datasheet indocs/DATASETS\.md\), andHillstrom,Lenta,X5, andMegaFonare public marketing randomized trials with binary outcomes and no counterfactual\(Hillstrom,[2008](https://arxiv.org/html/2608.00915#bib.bib15); Lenta,[2022](https://arxiv.org/html/2608.00915#bib.bib25); X5 Group,[2021](https://arxiv.org/html/2608.00915#bib.bib36); MegaFon,[2021](https://arxiv.org/html/2608.00915#bib.bib28)\)\. ACriteosupplement\(Diemert et al\.,[2018](https://arxiv.org/html/2608.00915#bib.bib9)\)ships outside this accounting \(†\\daggerin Table[1](https://arxiv.org/html/2608.00915#S3.T1); Appendix[I](https://arxiv.org/html/2608.00915#A9)\)\.

Crucially,SyntheticandIHDPprovide known simulated effect targets \(PEHE\\sqrt\{\\mathrm\{PEHE\}\}\) andJobsan experimentally identified policy\-value objective \(policy risk\), while the four marketing RCTs have no counterfactual\. The benchmark’s scope and the findings’ evidence base are deliberately distinct: UpliftBench spans 25 primary instances from seven families, but F1 is identified on the continuous benchmark with known effects — the 10 main IHDP realizations, replicated across all 100 — and F2 is evaluated on Jobs — the only benchmark whose deployment objective is experimentally identified\.

## 4\.Estimators, Metrics, and Protocol

##### Estimators \(12\)\.

Five meta\-learners \(S/T/X/R/DR\)\(Künzel et al\.,[2019](https://arxiv.org/html/2608.00915#bib.bib22); Nie and Wager,[2021](https://arxiv.org/html/2608.00915#bib.bib29); Kennedy,[2020](https://arxiv.org/html/2608.00915#bib.bib20)\); three class\-transformation models \(ClassTrans, TwoModel, SoloModel\)\(Jaskowski and Jaroszewicz,[2012](https://arxiv.org/html/2608.00915#bib.bib16); Kane et al\.,[2014](https://arxiv.org/html/2608.00915#bib.bib17)\); three uplift forests \(KL/ED/χ2\\chi^\{2\}\)\(Rzepakowski and Jaroszewicz,[2012](https://arxiv.org/html/2608.00915#bib.bib31)\); and the causal forest\(Wager and Athey,[2018](https://arxiv.org/html/2608.00915#bib.bib35); Battocchi et al\.,[2019](https://arxiv.org/html/2608.00915#bib.bib2)\)\. Meta\-learners and class\-transformation models share one LightGBM base learner\(Ke et al\.,[2017](https://arxiv.org/html/2608.00915#bib.bib19)\)and HP search space, so differences reflect the meta\-strategy\. We later swap in XGBoost to test robustness \(Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)\)\.

##### Six metrics\.

We compute three*ranking*metrics — Qini \(unnormalized\), AUUC, and uplift\-at\-k=0\.3k\{=\}0\.3; one*policy*metric — policy value atk=0\.3k\{=\}0\.3; and, where a reference objective exists,PEHE\\sqrt\{\\mathrm\{PEHE\}\}\(RMSE of the CATE\) and policy riskRpolicyR\_\{\\mathrm\{policy\}\}\(Shalit et al\.,[2017](https://arxiv.org/html/2608.00915#bib.bib32)\)\. Calibration is measured by expected calibration error \(ECE\)\(Kuleshov et al\.,[2018](https://arxiv.org/html/2608.00915#bib.bib21)\)\. Each dataset’s*primary*metric is what is knowable:PEHE\\sqrt\{\\mathrm\{PEHE\}\}\(Synthetic, IHDP\), policy risk \(Jobs\), Qini \(marketing\)\.Orientation\.PEHE\\sqrt\{\\mathrm\{PEHE\}\}andRpolicyR\_\{\\mathrm\{policy\}\}are error/risk quantities \(lower is better\); to compare them with the ranking metrics \(higher is better\) we always correlate against their negations−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}and−Rpolicy\-R\_\{\\mathrm\{policy\}\}, so a positive correlation means “tracks the reference objective” throughout \(PEHE\\sqrt\{\\mathrm\{PEHE\}\}is known exactly on IHDP/synthetic; Jobs policy risk is experimentally estimated\)\. Throughout, “Qini” means the canonical cumulative\-gain definition below; library implementations differ in baseline, normalization, and tie conventions, and our claims are about this canonical object\.

##### Why PolicyValue@kkis not in panel \(a\) of Fig\.[1](https://arxiv.org/html/2608.00915#S5.F1)\.

It evaluates one budgeted decision with treatment\-weighted noise, not global effect accuracy \(ρ≈0\.06\\rho\\approx 0\.06with−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}; no oracle reading on continuous data\), so it appears only where it is meaningful \(panels \(b\)–\(c\), Section[5\.2](https://arxiv.org/html/2608.00915#S5.SS2)\)\.

##### Policy risk and policy value \(exact definitions\)\.

These are the two deployment objectives, and they differ\.*Policy risk*followsShalit et al\.\([2017](https://arxiv.org/html/2608.00915#bib.bib32)\): a modelτ^\\hat\{\\tau\}induces the*sign\-thresholded*ruleπ​\(x\)=𝟙​\[τ^​\(x\)≥0\]\\pi\(x\)=\\mathbb\{1\}\[\\hat\{\\tau\}\(x\)\\geq 0\], and on the randomized Jobs experimental subset

Rpolicy​\(π\)\\displaystyle R\_\{\\mathrm\{policy\}\}\(\\pi\)=1−𝔼^​\[Y​\(π\)\],\\displaystyle=1\-\\widehat\{\\mathbb\{E\}\}\\big\[Y\(\\pi\)\\big\],𝔼^​\[Y​\(π\)\]\\displaystyle\\widehat\{\\mathbb\{E\}\}\\big\[Y\(\\pi\)\\big\]=1\|ℰ\|∑i∈ℰ\[yiπ^1𝟙\[πi=1,ti=1\]\\displaystyle=\\frac\{1\}\{\|\\mathcal\{E\}\|\}\\\!\\sum\_\{i\\in\\mathcal\{E\}\}\\Big\[\\tfrac\{y\_\{i\}\}\{\\hat\{\\pi\}\_\{1\}\}\\mathbb\{1\}\[\\pi\_\{i\}\{=\}1,t\_\{i\}\{=\}1\]\+yi1−π^1𝟙\[πi=0,ti=0\]\]\.\\displaystyle\\qquad\\qquad\+\\tfrac\{y\_\{i\}\}\{1\-\\hat\{\\pi\}\_\{1\}\}\\mathbb\{1\}\[\\pi\_\{i\}\{=\}0,t\_\{i\}\{=\}0\]\\Big\]\.withπ^1\\hat\{\\pi\}\_\{1\}the empirical treated fraction in the experimental subsetℰ\\mathcal\{E\}\(an inverse\-propensity\-weighted \(IPW\) estimate; lower risk is better; no further normalization\)\.*Policy value atkk*is a different, budgeted quantity: treat the top\-kkfraction by score,Vk=1n​\(∑i∈top\-​k,ti=1yi/e^i\+∑i∉top\-​k,ti=0yi/\(1−e^i\)\)V\_\{k\}=\\tfrac\{1\}\{n\}\\big\(\\sum\_\{i\\in\\text\{top\-\}k,t\_\{i\}=1\}y\_\{i\}/\\hat\{e\}\_\{i\}\+\\sum\_\{i\\notin\\text\{top\-\}k,t\_\{i\}=0\}y\_\{i\}/\(1\-\\hat\{e\}\_\{i\}\)\\big\)atk=0\.3k\{=\}0\.3\. Policy risk uses a sign threshold; policy value uses a budget — we report both\. The two use different propensity conventions by design: for Jobs policy risk we follow the established benchmark definition ofShalit et al\.\([2017](https://arxiv.org/html/2608.00915#bib.bib32)\)with the experimental treated fractionπ^1\\hat\{\\pi\}\_\{1\}, whereas PolicyValue@kkuses the general per\-dataset implementation across all benchmarks \(known propensities on the RCTs, train\-fold\-estimatede^i\\hat\{e\}\_\{i\}on observational data\)\.

##### Outer\-test\-isolated protocol\.

All preprocessing, propensity estimation and tuning are fit within the training fold only; we use stratified 3\-fold CV repeated over 3 seeds \(9 evaluations/cell\), tuning by random search \(B=10B\{=\}10, 2 inner folds\)\.The inner tuning objective is validation Qini on every dataset\(the released default\), so the candidate panel every metric is scored on was itself selected under one of the metrics we scrutinize; Section[7](https://arxiv.org/html/2608.00915#S7)states what this does and does not affect\.

##### Uncertainty: what is resampled\.

All F1/F2 CIs are cluster bootstraps whose resampling unit is the*benchmark realization*, so the effective sample size is the number of realizations \(10 continuous, 10 Jobs\) — never the fold rows \(Section[5](https://arxiv.org/html/2608.00915#S5)\)\. IHDP realizations are independent simulated draws; the 10 Jobs units are*re\-splits of one sample*, so their intervals can understate uncertainty, quantified by a design\-effect sensitivity in Section[5\.2](https://arxiv.org/html/2608.00915#S5.SS2.SSS0.Px4)\.

##### A note on Qini normalization\.

We report the*unnormalized*Qini throughout and make no claims based on the normalized variant: it divides by a perfect\-curve area that is frequently non\-positive on continuous outcomes, leaving the score undefined or orientation\-reversed there \(Appendix[A](https://arxiv.org/html/2608.00915#A1)\)\.

Before turning to results, Table[2](https://arxiv.org/html/2608.00915#S4.T2)summarizes the reusable benchmark artifact that operationalizes the analyses below\.

Table 2\.The UpliftBench artifact at a glance\.

## 5\.Results

The benchmark produced 2,112 fold\-level rows \(2,052 successful; 60 correctly skipped: binary\-only models on IHDP’s continuous outcome\)\. At the dataset×\\timesmodel level, all 300 cells are accounted for \(Criteo supplement outside this ledger; Appendix[I](https://arxiv.org/html/2608.00915#A9)\): 228 completed, 60 are inapplicable IHDP/binary\-model skips, and 12 timed out: the three uplift forests on each ofSynthetic,HillstromandMegaFon\(9\), DR onLenta, and R and DR onMegaFon\. TheSyntheticforest timeouts atn=2,000n\{=\}2\{,\}000are the fixed per\-job budget interacting with the∼189\{\\sim\}189tuning\-loop fits per fold evaluation, not a failure at that sample size\. Timeouts are marked, not imputed; skips are whole cells \(a completed cell is 9 fold rows\)\.

### 5\.1\.Finding 1: outcome\-regime sensitivity of Qini

![Refer to caption](https://arxiv.org/html/2608.00915v1/x1.png)

Three heatmap panels of mean cross\-metric rank agreement: on continuous outcomes Qini is isolated while AUUC and uplift\-at\-k track effect accuracy; on binary outcomes the ranking metrics agree with each other; and on Jobs alone every ranking metric diverges from the identified policy objective\.

Figure 1\.Two findings of metric disagreement\.Mean*across dataset instances*of the within\-instance Spearmanρ\\rhocomputed across models \(not across folds\)\. Each panel is computed on a single, stated instance set:\(a\)the 10 continuous instances \(IHDP×\\times10;Syntheticis generated with a binary outcome and is therefore grouped with the binary instances in \(b\)\) — Qini is isolated, agreeing neither with ground truth \(−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\},ρ=−0\.15\\rho=\-0\.15\) nor with sibling ranking metrics AUUC/uplift\-at\-kk\(ρ≈0\\rho\\approx 0\), which themselves correlate substantially with−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}\(ρ≈0\.6\\rho\\approx 0\.6\);\(b\)the 15 binary instances \(10 Jobs\+\+Synthetic\+\+4 marketing\) — ranking metrics agree tightly \(ρ≈0\.8\\rho\\approx 0\.8–0\.90\.9\);\(c\)the 10 Jobs splits, the only binary instances where the deployment objective is identified — every ranking metric diverges from it \(ρ≈−0\.25\\rho\\approx\-0\.25with policy value−Rpolicy\-R\_\{\\mathrm\{policy\}\}; equivalently, positively correlated with policy*risk*\)\. The reference objectives are oriented higher\-is\-better as−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}and−Rpolicy\-R\_\{\\mathrm\{policy\}\}, so a positiveρ\\rhomeans “tracks the objective\.”The failure appears on the evaluated continuous benchmark \(IHDP\) \(Fig\.[1](https://arxiv.org/html/2608.00915#S5.F1)a\)\. On the primary continuous panel \(the 10 IHDP splits and the six estimators applicable to continuous outcomes — binary\-only models are skipped by design\), the Qini ranking shows no detectable alignment with effect accuracy \(ρ=−0\.15\\rho=\-0\.15vs\.−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\};\+0\.02\+0\.02vs\. AUUC\) while AUUC and uplift\-at\-kkcorrelate substantially with−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}\(\+0\.68\+0\.68and\+0\.69\+0\.69\)\. Tuning off gives a paired gap of\+0\.55\+0\.55\[\+0\.33,\+0\.77\+0\.33,\+0\.77\]; inverting the inner objective to AUUC gives\+0\.56\+0\.56\[\+0\.34,\+0\.78\+0\.34,\+0\.78\] \(Appendix[E\.2](https://arxiv.org/html/2608.00915#A5.SS2)\) — F1 is an artifact of neither metric’s selection\. Nor is the six\-model rank correlation doing the work: as pairwise concordance over15001500model\-pair comparisons across all 100 realizations, Qini orders pairs at0\.520\.52\[0\.49,0\.560\.49,0\.56\] — a coin flip — against AUUC’s0\.730\.73\[0\.69,0\.760\.69,0\.76\]\. On the extended common panel — all100100IHDP realizations — Qini’s mean rank correlation with effect accuracy is\+0\.07\+0\.07, 95% CI\[−0\.03,\+0\.16\]\[\-0\.03,\+0\.16\], against a paired AUUC\-over\-Qini gap of\+0\.49\+0\.49\[\+0\.40,\+0\.59\+0\.40,\+0\.59\]\. Within the evaluated benchmark panel, the observed isolation is specific to Qini rather than a generic ranking\-versus\-estimation gap\.

##### F1 is not a tail effect\.

The regime that*indexes*F1 is not a mechanism for it: the panel where F1 appears is*mild*\-tailed, and on a far heavier\-tailed family Qini tracks effect accuracy \(Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)\)\. In short:*Qini is uninformative on IHDP, not because of heavy tails, and we cannot yet say why*\.

##### Definitions and identities\.

Sort units by descending predicted uplift \(deterministic tie\-breaking\), with cumulative treated/control countsTk,CkT\_\{k\},C\_\{k\}at depthkk\. Qini integrates the count\-corrected cumulative*sum*g​\(k\)=∑i≤kyi​ti−\(∑i≤kyi​\(1−ti\)\)​Tk/Ckg\(k\)=\\sum\_\{i\\leq k\}y\_\{i\}t\_\{i\}\-\\big\(\\sum\_\{i\\leq k\}y\_\{i\}\(1\-t\_\{i\}\)\\big\)T\_\{k\}/C\_\{k\}against a chord baseline; our AUUC — the*prefix\-mean*AUUC; “AUUC” unqualified always means this variant — integrates the difference of cumulative*means*u​\(k\)u\(k\)minus the ATE triangle \(a*shifted*convention: a random ranking scores≈ATE/2\{\\approx\}\\mathrm\{ATE\}/2, identical for every model on a fold, so it cancels from every reported statistic\); both use the population\-fraction axis and we report raw \(unnormalized\) areas \(edge cases in Appendix[D](https://arxiv.org/html/2608.00915#A4)\); the shippedcausalml/scikit\-upliftAUUC instead integrates the cumulative\-gain curvek​u​\(k\)k\\,u\(k\)— the*cumulative\-gain*AUUC of Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)\. The two are linked exactly:

###### Lemma 5\.1 \(Qini is treated\-count\-weighted AUUC\)\.

On any prefix where both arms are present,g​\(k\)=Tk​u​\(k\)g\(k\)=T\_\{k\}\\,u\(k\)\.

###### Lemma 5\.2 \(Score\-level relationship\)\.

With baseline\-subtracted residualsq​\(k\)=g​\(k\)−kn​g​\(n\)q\(k\)=g\(k\)\-\\tfrac\{k\}\{n\}g\(n\)anda​\(k\)=u​\(k\)−kn​u​\(n\)a\(k\)=u\(k\)\-\\tfrac\{k\}\{n\}u\(n\),

q​\(k\)=Tk​a​\(k\)\+kn​u​\(n\)​\(Tk−Tn\)\.q\(k\)=T\_\{k\}\\,a\(k\)\+\\tfrac\{k\}\{n\}\\,u\(n\)\\,\(T\_\{k\}\-T\_\{n\}\)\.

Thus Qini introduces treated\-count depth weighting and an additional treatment\-interleaving term relative to AUUC; these identities provide a structural pathway, not a complete explanation of the observed IHDP separation \(proofs in Appendix[D](https://arxiv.org/html/2608.00915#A4)\)\.

##### Scale ruled out; the failure in one construction\.

Positive affine transformations preserve every within\-split Qini, AUUC, and uplift\-at\-kkmodel ranking, ruling out scale and location as explanations \(Prop\.[D\.1](https://arxiv.org/html/2608.00915#A4.Thmtheorem1)and empirical confirmation, Appendix[D](https://arxiv.org/html/2608.00915#A4)\)\. The distribution’s*shape*, however, can bite: on a hand\-checkable 12\-unit construction, a single large control outcome \(y=40y\{=\}40\) makes Qini prefer a near\-random model over a near\-true one while AUUC and−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}prefer the better model, and deleting that one outcome flips Qini back \(Table[9](https://arxiv.org/html/2608.00915#A7.T9)\)\.

##### A controlled test, and an honest boundary\.

Holding models and latent CATE fixed and varying only the outcome distribution \(Fig\.[3](https://arxiv.org/html/2608.00915#A1.F3)\), affine transforms change nothing and rising kurtosis degrades*all*cumulative ranking metrics — Qini and AUUC*together*— so the controlled experiment does not reproduce Qini’s isolation, and candidate mechanisms remain open \(Appendix[D](https://arxiv.org/html/2608.00915#A4)\)\.

The clearest symptom is theDR\-Learneron IHDP: it tops Qini \(199\.6199\.6on s8\) yet is ranked worst byPEHE\\sqrt\{\\mathrm\{PEHE\}\}, which diverges into the tens of thousands \(max153,013153\{,\}013, one fold of the main run’s s3 cell; the per\-realization cells of Fig\.[7](https://arxiv.org/html/2608.00915#A7.F7)average nine folds and sit lower, s3 cell mean≈2\.9×104\\approx 2\.9\{\\times\}10^\{4\}\) where EconML’s augmented\-IPW \(AIPW\) nuisance estimation degenerates — a known numerical risk of inverse\-propensity\-weighted pseudo\-outcomes under extreme estimated propensities\(cf\. Kennedy,[2020](https://arxiv.org/html/2608.00915#bib.bib20)\)\. Because F1 is rank\-based this magnitude is irrelevant, and F1 survives excluding DR \(and DR\+\+R\) entirely \(Appendix[D](https://arxiv.org/html/2608.00915#A4)\)\.

### 5\.2\.Finding 2: ranking metrics are insufficient for level\-dependent deployment objectives

Figure[1](https://arxiv.org/html/2608.00915#S5.F1)\(b\)–\(c\) illustrates F2, which has two parts: a*structural*insufficiency of ranking metrics for level\-dependent objectives, and its operational magnitude on Jobs\. On binary outcomes Qini behaves like a normal ranking metric — it agrees with AUUC atρ=0\.90\\rho=0\.90\(Jobs\) and0\.720\.72\(marketing\) — yet it does not recover the*policy*objective on Jobs, measured by the RCT\-estimated policy risk \(an IPW value on the randomized experimental subset; Section[4](https://arxiv.org/html/2608.00915#S4)\)\.

###### Proposition 5\.3 \(Rank\-invariance boundary\)\.

Any evaluation metric that depends onτ^\\hat\{\\tau\}only through the induced ranking of units is invariant to every strictly increasing transformation ofτ^\\hat\{\\tau\}, including shiftsτ^↦τ^\+c\\hat\{\\tau\}\\mapsto\\hat\{\\tau\}\+c\. The sign\-threshold policyπ​\(x\)=𝟙​\[τ^​\(x\)≥0\]\\pi\(x\)=\\mathbb\{1\}\[\\hat\{\\tau\}\(x\)\\geq 0\]is not: a shift changes which units are treated\. Hence for two models whose scores share a ranking but differ by a shift, ranking metrics assign identical values while their sign policies — and the resulting policy values — can differ\. Ranking metrics therefore cannot, from rank information alone, identify the sign\-threshold\-optimal model; that requires the score*level*\(calibration or an explicit threshold\)\.

This is a*non\-identifiability*result, not a claim that ranking metrics carry no information in practice — real estimators do not differ by pure shifts — so both halves are load\-bearing: the proposition gives insufficiency in principle, and Jobs supplies the empirical claim that here the signal is not transmitted\. The headline is a selector contrast:*within the released split rotation, direct policy\-risk selection lowers Jobs benchmark regret versus random; Qini, AUUC, and uplift\-at\-kkdo not\.*

##### On Jobs, the benchmark contains detectable policy\-selection signal that the evaluated ranking metrics do not transmit\.

Under a cross\-repeat rotation \(select on two repeated\-CV seeds, evaluate on the held\-out seed; cluster bootstrap over the 10 Jobs splits\), selecting by*cross\-fitted policy risk itself*reduced benchmark risk relative to random selection by\+0\.0209\+0\.0209\(95% CI\[\+0\.0071,\+0\.0358\]\[\+0\.0071,\+0\.0358\], excluding zero\), whereas the Qini, AUUC, and uplift\-at\-kkselectors changed it by−0\.0070\-0\.0070,−0\.0070\-0\.0070, and−0\.0050\-0\.0050respectively — none distinguishable from zero \(ladder details in Appendix[E](https://arxiv.org/html/2608.00915#A5)\)\.

##### The operational cost: policy\-selection regret\.

With both the metric winner and the risk\-minimizing reference chosen on the selection seeds and evaluated on the held\-out seed \(policyπ​\(x\)=1\\pi\(x\)\{=\}1iffτ^​\(x\)≥0\\hat\{\\tau\}\(x\)\{\\geq\}0\), metric selection incurs mean regret0\.0260\.026–0\.0280\.028risk —1414–15%15\\%of the best selected\-candidate risk — and is comparable to the random\-selection baseline \(paired differences’ CIs span zero\)\. Because repeated\-CV seeds re\-partition the*same*observations, this is cross\-repeat benchmark\-selection regret, not evaluation on a new sample; secondary statistics \(within\-margin rates, paired differences per metric\) are in Appendix[E](https://arxiv.org/html/2608.00915#A5)\.

##### Calibrating the threshold closes most of the gap\.

As Proposition[5\.3](https://arxiv.org/html/2608.00915#S5.Thmtheorem3)predicts, replacing the fixed zero threshold with one calibrated on the selection folds cuts the Qini\-selection regret by81%81\\%\(to\+0\.0049\+0\.0049, CI no longer excluding zero\): F2 on the sign\-threshold objective is largely a score\-*level*effect, so pair a ranking metric with a calibrated threshold \(Appendix[A](https://arxiv.org/html/2608.00915#A1)\)\.

##### How strong is this evidence?

The Jobs instances re\-split one sample and the IPW risk is noisy \(SE≈0\.023\\approx 0\.023\); the regret interval excludes zero only for between\-split dependenceρ¯≤0\.12\\bar\{\\rho\}\\leq 0\.12, and the splits’ per\-model risk vectors already correlate at\+0\.41\+0\.41across the 10 splits \(a raw correlation reflecting both between\-model signal and split dependence, not itselfρ¯\\bar\{\\rho\}of the differenced statistic\)\. We therefore read the regret as*modest*and rest F2 on its convergent parts — the structural boundary \(Prop\.[5\.3](https://arxiv.org/html/2608.00915#S5.Thmtheorem3)\), the selector contrast, and threshold calibration \(Appendix[E](https://arxiv.org/html/2608.00915#A5)\)\.

Takeaway \(F2\)\.*Jobs contains usable policy\-selection signal, but the common ranking metrics fail to capture it:*direct policy\-risk selection yields lower benchmark regret than random within the released split rotation, whereas Qini, AUUC, and uplift\-at\-kkdo not; their selected policies incur a1414–15%15\\%regret relative to the risk\-selected candidate\. For sign\-threshold policies the mismatch is structural — rank\-only metrics discard the score level — andcalibrating the threshold removes81%81\\%of the Qini\-selection regret, which is the paper’s most directly actionable recommendation; the gap disappears when model selection and evaluation are aligned to budgeted policy value\.

##### Which model wins depends on regime and metric\.

No estimator dominates: simple meta\-learners and the causal forest lead the semi\-synthetic regimes, simple estimators lead the marketing RCTs \(per\-dataset winners and critical\-difference diagrams in Appendix[A](https://arxiv.org/html/2608.00915#A1)\)\. These standings are protocol\-scoped — samples in the hundreds, 3\-fold CV, and aB=10B\{=\}10tuning budget can under\-serve nuisance\-heavy learners \(DR/R\) — not universal estimator verdicts \(Section[6](https://arxiv.org/html/2608.00915#S6), Step 3\)\.

### 5\.3\.Robustness and boundaries

#### 5\.3\.1\.Robustness of F1

##### Paired, decomposed statistics\.

We do not rest on a pooled Qini\-vs\-reference correlation: its CI crosses zero, and pooling outcome regimes conflates the two findings \(Appendix[E](https://arxiv.org/html/2608.00915#A5)\)\. The load\-bearing statistic is the paired AUUC\-over\-Qini gap on the continuous panel,Δ=\+0\.63\\Delta=\+0\.63\[\+0\.35,\+0\.91\+0\.35,\+0\.91\], positive in all six exclusion/base\-learner cells and with its CI excluding zero in five \(Table[6](https://arxiv.org/html/2608.00915#A7.T6)\)\.

##### Uncertainty, and all 100 realizations\.

On the 10 main splits the paired gap is positive on99/10 \(Qini\+0\.02\+0\.02\[−0\.33,\+0\.37\-0\.33,\+0\.37\]\)\. Extending to*all*100100IHDP realizations \(split indices fixed before analysis; meta/forest learners on the added splits\) strengthens the evidence: Qini’s correlation with effect accuracy is\+0\.07\+0\.07\[−0\.03,\+0\.16\-0\.03,\+0\.16\] while AUUC and uplift\-at\-kktrack it at\+0\.56\+0\.56and\+0\.64\+0\.64; the paired gap is\+0\.49\+0\.49\[\+0\.40,\+0\.59\+0\.40,\+0\.59\], positive on8080/100100realizations \(run documentation in Table[10](https://arxiv.org/html/2608.00915#A7.T10)\)\. Extending from the 10\-split primary panel to all100100realizations lowers the point estimate \(\+0\.63\+0\.63to\+0\.49\+0\.49\) and tightens the interval, so the primary panel did not understate the gap\. With six common estimators in the extension, each per\-realization Spearman is discrete and coarse; inference therefore concerns the mean paired rank statistic across realizations, not fine\-grained estimation within any single one\.

##### Estimator exclusion\.

Excluding the DR\-Learner, and both DR\- and R\-Learners, keeps Qini’s correlation with ground truth indistinguishable from zero in every configuration and the pairedΔ\\Deltapositive in all six \(base learner×\\timesexclusion\) cells, its CI excluding zero in five \(Table[6](https://arxiv.org/html/2608.00915#A7.T6)\); AUUC’s absolute alignment falls as estimators are removed, so we state F1 relatively \(AUUC consistently more aligned than Qini\), not as a fixed level\.Δ\\Deltadeclines \(\+0\.63→\+0\.40→\+0\.42\+0\.63\\rightarrow\+0\.40\\rightarrow\+0\.42;\+0\.69→\+0\.44→\+0\.34\+0\.69\\rightarrow\+0\.44\\rightarrow\+0\.34\) and the decline runs entirely through AUUC’s alignment, so part of the advantage is AUUC penalizing the unstable learners’PEHE\\sqrt\{\\mathrm\{PEHE\}\}blow\-ups, to which a rank\-only statistic is indifferent:the advantage attenuates but does not vanish— on all100100realizations the excluded\-panel gaps are\+0\.28\+0\.28\[\+0\.18,\+0\.38\+0\.18,\+0\.38\] and\+0\.34\+0\.34\[\+0\.21,\+0\.46\+0\.21,\+0\.46\], CIs excluding zero\.

##### Base learner\.

Switching to XGBoost \(Fig\.[2](https://arxiv.org/html/2608.00915#A1.F2)\) leaves the F1 signature intact \(Qini⟂PEHE\\perp\\sqrt\{\\mathrm\{PEHE\}\}, AUUC∥PEHE\\parallel\\sqrt\{\\mathrm\{PEHE\}\}\) and preserves the ranking metrics’ mutual agreement on binary outcomes\. The conclusions are about*metrics*, not a LightGBM artifact\.

##### Implementation variants\.

Scoring identical predictions withcausalml’s shipped Qini yields nearly identical rankings \(rankρ\\rho\+0\.99\+0\.99; winner agreement100%100\\%\) and the same F1 result\. Correcting the R\-Learner/Causal Forest treatment nuisance to a classifier \(an EconML requirement; a regressor through v1\.4\.0, caught by an external audit\) and rerunning the full benchmark leaves F1 intact: corrected gap\+0\.63\+0\.63\[\+0\.35,\+0\.91\+0\.35,\+0\.91\] vs\.\+0\.83\+0\.83\[\+0\.58,\+1\.09\+0\.58,\+1\.09\] pre\-fix, Causal Forest cells unchanged \(ρ=\+1\.00\\rho=\+1\.00; Appendix[E](https://arxiv.org/html/2608.00915#A5)\)\.

#### 5\.3\.2\.Boundary of F1

##### Two boundary\-case validation families\.

ACIC 2016\(Dorie et al\.,[2019](https://arxiv.org/html/2608.00915#bib.bib10)\)supplies1818instances with real covariates and known effects\. There Qini behaves normally \(\+0\.48\+0\.48\[\+0\.25,\+0\.68\+0\.25,\+0\.68\] vs−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}\), AUUC gives\+0\.55\+0\.55\[\+0\.34,\+0\.73\+0\.34,\+0\.73\], and the paired AUUC\-over\-Qini gap — the statistic F1 is defined by — is\+0\.11\+0\.11\[−0\.03,\+0\.25\-0\.03,\+0\.25\], CI covering zero; its outcome kurtosis \(≈0\.4\{\\approx\}0\.4\) is similarly mild\.

##### A third continuous family ends the tail explanation\.

We add a third family with known effects and a response surface unlike IHDP’s \(Revenue\-Synthetic: lognormal spend,*multiplicative*effect;1010realizations, same protocol\)\.F1 does not replicate: Qini tracks effect accuracy \(\+0\.39\+0\.39\[\+0\.00,\+0\.73\+0\.00,\+0\.73\]\) and the paired gap vanishes,−0\.02\-0\.02\[−0\.19,\+0\.17\-0\.19,\+0\.17\]\. The family that fails F1 is the least heavy\-tailed of the three \(medians in Sec\.[5\.1](https://arxiv.org/html/2608.00915#S5.SS1)\), so heavy tails are neither necessary nor sufficient\. F1 appears on one of three evaluated continuous families: Qini is*sometimes*uninformative and*sometimes*the best of the ranking metrics, spanning\+0\.02\+0\.02to\+0\.48\+0\.48; no measured property yet predicts which case a practitioner is in \(Appendix[E\.1](https://arxiv.org/html/2608.00915#A5.SS1)\) — which is why Qini cannot be trusted unvalidated \(Appendix[A](https://arxiv.org/html/2608.00915#A1)\)\. This is consistent withCurth et al\.\([2021](https://arxiv.org/html/2608.00915#bib.bib5)\), who argue IHDP is idiosyncratic; we add that the idiosyncrasy is*metric\-facing*, and that no proposed property yet predicts the metric\-specific gap ex ante — the one surviving candidate is the lemma\-derived composite of Appendix[E\.1](https://arxiv.org/html/2608.00915#A5.SS1), pending its controlled sweep\.

##### Mechanism probes, summarized\.

RATE and oracle residualization improve Qini only modestly, and controlled sweeps of kurtosis, imbalance, and correlated score error do not isolate the separation\. A weighting decomposition localizes it: the field’s shipped cumulative\-gain AUUC \(checked againstcausalmlon all54005400folds,\+1\.00\+1\.00\) tracks−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}at\+0\.73\+0\.73\[\+0\.66,\+0\.79\+0\.66,\+0\.79\] on identical predictions over all100100realizations — better than our prefix\-mean\+0\.56\+0\.56— while Qini sits at≈0\{\\approx\}0: among these functionals the discrepancy localizes to*treated\-count*, not depth, weighting \(Lemma[5\.2](https://arxiv.org/html/2608.00915#S5.Thmtheorem2)’s interleaving term\)\. Mechanism otherwise open \(values in Appendix[A](https://arxiv.org/html/2608.00915#A1)\)\.

#### 5\.3\.3\.Robustness and scope of F2

##### Base\-learner replication\.

Under XGBoost, all three ranking metrics correlate negatively with policy value \(Qini−0\.35\-0\.35\[−0\.53,−0\.18\-0\.53,\-0\.18\], there*not*borderline\) and the selection regret stays positive \(0\.0170\.017–0\.0230\.023\); Table[11](https://arxiv.org/html/2608.00915#A7.T11)audits both runs side by side\.

##### Budget grid\.

The F2 selection regret is positive with the CI excluding zero at*every*budgetk∈\{0\.1,0\.2,0\.3,0\.5\}k\\in\\\{0\.1,0\.2,0\.3,0\.5\\\}and under the area\-under\-curve selector; within\-dataset rankings are moderately stable across budgets \(mean pairwise Spearman\+0\.75\+0\.75; per\-budget values in Appendix[E](https://arxiv.org/html/2608.00915#A5)\)\.

##### Objective specificity\.

The F2 gap is objective\-specific: direct risk selection captures signal for the sign\-threshold policy, while ranking metrics transmit signal when the reference objective is budgeted policy value\. With policy value at a fixed budget as both reference and objective, every selector beats random and metric\-vs\-reference regrets are indistinguishable from zero at every budget \(Appendix[E](https://arxiv.org/html/2608.00915#A5)\) — which is exactly why the metric must match the intended deployment rule\.

### 5\.4\.Calibration: an identification check, not a third finding

Mis\-calibration is*not*a third finding: the diff\-in\-means reliability estimator behind uplift\-ECE is unbiased only under randomized assignment, and once the analysis is stratified by identification regime the pooled Qini\-vs\-ECE disagreement is revealed as an identification artifact — where ECE is identified, calibration does*not*disagree with Qini\. The full stratified analysis, figure, and tables are in Appendix[C](https://arxiv.org/html/2608.00915#A3)\.

## 6\.Practitioner’s Guide: Choose a Metric, Then a Model

Match the metric to the*outcome type*\(on continuous outcomes avoid the unnormalized Qini as the sole criterion; preferPEHE\\sqrt\{\\mathrm\{PEHE\}\}where identifiable, and compare ranking metrics against each other otherwise\) and to the*deployment objective*\(evaluate policy value at the operating budget, and assess calibration with an identified estimator when allocation magnitude matters\), then validate rankings on more than one dataset\. For sign\-threshold deployment, pair the ranking metric with a threshold calibrated on held\-out selection data \(Section[5\.2](https://arxiv.org/html/2608.00915#S5.SS2)\)\. On continuous outcomes, bootstrap independent evaluation units \(or benchmark realizations where available\), compare Qini and AUUC rankings, and treat the winner as unstable where they diverge\. The full step\-by\-step guide is in Appendix[B](https://arxiv.org/html/2608.00915#A2)\.

## 7\.Limitations

##### No ground truth on marketing data\.

We establish F1 on semi\-synthetic benchmarks with known effects, and F2 using experimentally identified policy value on Jobs; we extrapolate the result only as a caution for marketing applications; we cannot establish whether Qini misranks on Lenta\. The failure appears on*one of the three*evaluated continuous families with reference effects \(IHDP\); across the three the paired gap forms a gradient \(\+0\.49\+0\.49,\+0\.11\+0\.11,−0\.02\-0\.02\) rather than a binary split; F1 should be read as a demonstrated failure case on a standard benchmark, not as a property of continuous outcomes in general\.

##### Protocol limitations\.

Three imperfections in the released protocol bound how F1/F2 should be read, and each is disclosed in full in Appendix[A](https://arxiv.org/html/2608.00915#A1)\. \(i\)*The candidate panel is Qini\-tuned*, conditioning the candidate set \(all metrics score identical out\-of\-fold predictions\); the finding persists with tuning off \(\+0\.55\+0\.55\) and with the inner objective switched to AUUC \(\+0\.56\+0\.56, Appendix[E\.2](https://arxiv.org/html/2608.00915#A5.SS2)\)\. \(ii\)*Propensity nesting is imperfect inside the tuning loop*: on IHDP/Jobs the propensity model is fit on the whole outer\-training fold and sliced for the inner folds — a covariate/treatment\-only deviation \(never outcomes\) affecting hyperparameter selection only\. \(iii\)*F2’s empirical half is within\-sample*: selection, calibration, and evaluation draw on the same Jobs observations — a descriptive case study whose intervals describe benchmark\-split variability, not population sampling error\. Fixes for \(ii\) and \(iii\) are planned follow\-up work\.

##### Scale and compute\.

Lenta/X5/MegaFon are capped at 10K by documented stratified subsampling;Criteo\(14M\) is outside the headline results; a v1\.4\.3 supplement adds its leaderboard — the 1M tier subsampled to the released 10K cap for comparability \(Appendix[I](https://arxiv.org/html/2608.00915#A9); binary\-regime agreement replicates, Qini–AUUC\+0\.74\+0\.74; enters neither F1 nor F2\) — and a 100K probe: resolution is cap\-limited — 10K gaps are small relative to fold noise and the 10K and 100K orderings did not agree \(≈0\{\\approx\}0rank correlation\); capped leaderboards are subsample\-scoped\. A handful of uplift\-forest and R/DR cells timed out \(marked, not imputed\); none affects F1/F2, whose load\-bearing cells have no timeouts \(correlations complete\-case\)\. A full\-scale sweep remains versioned future work\.

##### Scope of F1\.

We do not claim Qini is categorically invalid on continuous outcomes: the three families form a gradient in the paired gap \(\+0\.49\+0\.49,\+0\.11\+0\.11,−0\.02\-0\.02\), not a binary split, and Qini tracks effect accuracy on both validation families \(\+0\.48\+0\.48,\+0\.39\+0\.39\)\. We establish that on the standard continuous benchmark the Qini ranking is effectively unrelated to effect accuracy while AUUC and uplift\-at\-kkstay informative, and report it because practitioners and libraries apply Qini to continuous outcomes and the failure is silent — and regime\-dependent, which is what makes it dangerous\.

## 8\.Reproducibility

The public reproducibility package — UpliftBench \(code, configs, data loaders, and result parquets\) — is released at[https://github\.com/binshuangli/uplift\-bench](https://github.com/binshuangli/uplift-bench); a DOI\-bearing archival release will accompany any archival version\. Dedicated analysis targets regenerate all reported tables, figures, and numerical macros from the committed result artifacts\. The core benchmark runs are driven bymake repro\(smoke pathmake repro\-smoke\); the prediction\-level auxiliary analyses \(the causalml variant, threshold calibration, and the F2 robustness checks\) are reproduced through the documentedmake repro\-r4predtarget\. Every result parquet is stamped with git hash and config; seeds are globally controlled\. Settings: 3 folds×\\times3 seeds,B=10B\{=\}10tuning configs, 2 inner folds; base learners LightGBM and XGBoost\. Each analysis regenerates from the raw result rows via a dedicated script underscripts/\(per\-artifact mapping in the repository README; artifact inventory in Table[2](https://arxiv.org/html/2608.00915#S4.T2)\)\.

##### Generative\-AI assistance\.

An AI coding assistant \(Claude Code\) assisted with code, documentation, manuscript drafting and editing, and brainstorming candidate analyses\. The author independently selected the research questions, specifications, methods, and analyses; reviewed, executed, and validated all generated code and outputs against the committed artifacts; and takes responsibility for all results and claims\.

## 9\.Benchmark Documentation, Availability, and Ethics

##### Provenance and curation\.

Dataset provenance is specified in Section[3](https://arxiv.org/html/2608.00915#S3)and per\-family datasheets ship indocs/DATASETS\.md\(source, licence, preprocessing, caveats\); no new human\-subjects data is introduced\. The two*auxiliary*continuous families \(Revenue\-Synthetic;ACIC 2016\) are regenerated deterministically and serve as F1’s boundary cases \(Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)\)\. Every dataset loads through a versioned loader recording size, treatment fraction, outcome type, and preprocessing; no full\-dataset statistic precedes the train/test split\.

##### Availability, licensing, and maintenance\.

Code, configs, loaders, result parquets, and the scripts regenerating every figure and table are MIT\-licensed; each dataset keeps its original license and is fetched by its loader from the original host, never redistributed \(Criteounder Criteo Research’s research\-use terms\)\. The repository is public; an archival release with a Zenodo DOI will accompany the conference version\.

##### Bring your own data\.

A documented adapter \(UpliftDataset\.from\_frame\) andregister\_loaderrun every estimator and metric on user data under the released protocol\.

##### Leaderboard governance\.

Submissions are pull requests adding a model card plus estimator code or per\-fold predictions on the*fixed*released folds/seeds \(RESULTS\.md,CONTRIBUTING\.md\); the maintainer re\-runs each within the released budget \(single\-maintainer; review latency scales accordingly\); results are frozen per version and timed\-out runs are shown, not hidden\. We do*not*claim overfitting resistance: folds and outcomes are public — a reusable\-target risk\(Kapoor and Narayanan,[2023](https://arxiv.org/html/2608.00915#bib.bib18)\)\.

##### Representativeness of the synthetic data\.

The continuous synthetic DGPs are checked against the real semi\-synthetic IHDP data \(all100100realizations\) and a controlled kurtosis sweep spanning0\.20\.2–7070\(Fig\.[3](https://arxiv.org/html/2608.00915#A1.F3)\)\.

##### Ethical considerations\.

All data are public and de\-identified by their publishers; no new personal data is introduced\. Mis\-specified evaluation misdirects consequential targeting resources — the risk this paper reduces\. UpliftBench evaluates neither subgroup fairness nor treatment burden and does not validate high\-stakes deployment\.

## 10\.Conclusion

In uplift evaluation, the metric is the message\. On IHDP, AUUC tracks effect accuracy while Qini does not \(\+0\.49\+0\.49\[\+0\.40,\+0\.59\+0\.40,\+0\.59\], F1\); on Jobs only policy\-risk selection lowers regret \(F2\), largely removed by threshold calibration\. Both are bounded: F1 holds on one of three continuous families, F2 vanishes under a budgeted objective\. Match the metric to outcome type and goal: UpliftBench makes that the default\.

## References

- \(1\)
- Battocchi et al\.\(2019\)Keith Battocchi, Eleanor Dillon, Maggie Hei, Greg Lewis, Paul Oka, Miruna Oprescu, and Vasilis Syrgkanis\. 2019\.EconML: A Python Package for ML\-Based Heterogeneous Treatment Effects Estimation\.[https://github\.com/py\-why/EconML](https://github.com/py-why/EconML)\.
- Bokelmann and Lessmann \(2024\)Björn Bokelmann and Stefan Lessmann\. 2024\.Improving uplift model evaluation on randomized controlled trial data\.*European Journal of Operational Research*313, 2 \(2024\), 691–707\.
- Chen et al\.\(2020\)Huigang Chen, Totte Harinen, Jeong\-Yoon Lee, Mike Yung, and Zhenyu Zhao\. 2020\.CausalML: Python package for causal machine learning\.*arXiv preprint arXiv:2002\.11631*\(2020\)\.
- Curth et al\.\(2021\)Alicia Curth, David Svensson, James Weatherall, and Mihaela van der Schaar\. 2021\.Really Doing Great at Estimating CATE? A Critical Look at ML Benchmarking Practices in Treatment Effect Estimation\. In*Advances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track*\.
- Demšar \(2006\)Janez Demšar\. 2006\.Statistical comparisons of classifiers over multiple data sets\.*Journal of Machine Learning Research*7 \(2006\), 1–30\.
- Devriendt et al\.\(2021\)Floris Devriendt, Jeroen Berrevoets, and Wouter Verbeke\. 2021\.Why you should stop predicting customer churn and start using uplift models\.*Information Sciences*548 \(2021\), 497–515\.
- Devriendt et al\.\(2018\)Floris Devriendt, Darie Moldovan, and Wouter Verbeke\. 2018\.A literature survey and experimental evaluation of the state\-of\-the\-art in uplift modeling: A stepping stone toward the development of prescriptive analytics\.*Big Data*6, 1 \(2018\), 13–41\.
- Diemert et al\.\(2018\)Eustache Diemert, Artem Betlei, Christophe Broisin, and Massih\-Reza Amini\. 2018\.A large scale benchmark for uplift modeling\.*KDD Workshop on Causal Discovery, Prediction and Decision*\(2018\)\.
- Dorie et al\.\(2019\)Vincent Dorie, Jennifer Hill, Uri Shalit, Marc Scott, and Dan Cervone\. 2019\.Automated versus do\-it\-yourself methods for causal inference: Lessons learned from a data analysis competition\.*Statist\. Sci\.*34, 1 \(2019\), 43–68\.
- Guelman et al\.\(2015\)Leo Guelman, Montserrat Guillén, and Ana M Pérez\-Marín\. 2015\.Uplift random forests\.*Cybernetics & Systems*46, 3\-4 \(2015\), 230–248\.
- Gutierrez and Gérardy \(2017\)Pierre Gutierrez and Jean\-Yves Gérardy\. 2017\.Causal inference and uplift modelling: A review of the literature\.*ICML Workshop on Predictive Causality*67 \(2017\), 1–13\.
- He et al\.\(2024\)Bowei He, Yunpeng Weng, Xing Tang, Ziqiang Cui, Zexu Sun, Liang Chen, Xiuqiang He, and Chen Ma\. 2024\.Rankability\-enhanced revenue uplift modeling framework for online marketing\. In*Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining \(KDD\)*\.
- Hill \(2011\)Jennifer L Hill\. 2011\.Bayesian nonparametric modeling for causal inference\.*Journal of Computational and Graphical Statistics*20, 1 \(2011\), 217–240\.
- Hillstrom \(2008\)Kevin Hillstrom\. 2008\.Mine that data\! Kevin Hillstrom’s email marketing challenge\.[https://blog\.minethatdata\.com/2008/03/minethatdata\-e\-mail\-analytics\-and\-data\.html](https://blog.minethatdata.com/2008/03/minethatdata-e-mail-analytics-and-data.html)\.
- Jaskowski and Jaroszewicz \(2012\)Maciej Jaskowski and Szymon Jaroszewicz\. 2012\.Uplift modeling for clinical trial data\.*ICML Workshop on Clinical Data Analysis*\(2012\)\.
- Kane et al\.\(2014\)Kevin Kane, Victor SY Lo, and Jianying Zheng\. 2014\.Mining for the truly responsive customers and prospects using true\-lift modeling\.*Journal of Marketing Analytics*2, 4 \(2014\), 218–238\.
- Kapoor and Narayanan \(2023\)Sayash Kapoor and Arvind Narayanan\. 2023\.Leakage and the reproducibility crisis in machine\-learning\-based science\.*Patterns*4, 9 \(2023\)\.
- Ke et al\.\(2017\)Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie\-Yan Liu\. 2017\.LightGBM: A highly efficient gradient boosting decision tree\.*Advances in Neural Information Processing Systems*30 \(2017\)\.
- Kennedy \(2020\)Edward H Kennedy\. 2020\.Optimal doubly robust estimation of heterogeneous causal effects\.*arXiv preprint arXiv:2004\.14497*\(2020\)\.
- Kuleshov et al\.\(2018\)Volodymyr Kuleshov, Nathan Fenner, and Stefano Ermon\. 2018\.Accurate uncertainties for deep learning using calibrated regression\.*International Conference on Machine Learning*\(2018\), 2796–2804\.
- Künzel et al\.\(2019\)Sören R Künzel, Jasjeet S Sekhon, Peter J Bickel, and Bin Yu\. 2019\.Metalearners for estimating heterogeneous treatment effects using machine learning\.*Proceedings of the National Academy of Sciences*116, 10 \(2019\), 4156–4165\.
- LaLonde \(1986\)Robert J LaLonde\. 1986\.Evaluating the econometric evaluations of training programs with experimental data\.*The American Economic Review*76, 4 \(1986\), 604–620\.
- Leng and Dimmery \(2024\)Yan Leng and Drew Dimmery\. 2024\.Calibration of Heterogeneous Treatment Effects in Randomized Experiments\.*Information Systems Research*\(2024\)\.
- Lenta \(2022\)Lenta\. 2022\.Lenta Uplift Modelling Dataset\.[https://www\.uplift\-modeling\.com/en/latest/api/datasets/fetch\_lenta\.html](https://www.uplift-modeling.com/en/latest/api/datasets/fetch_lenta.html)\.
- Mahajan et al\.\(2024\)Divyat Mahajan, Ioannis Mitliagkas, Brady Neal, and Vasilis Syrgkanis\. 2024\.Empirical analysis of model selection for heterogeneous causal effect estimation\. In*International Conference on Learning Representations \(ICLR\)*\.
- Maksimov et al\.\(2020\)Nikita Maksimov, Anvar Kurmukov, and Elena Shevchenko\. 2020\.scikit\-uplift: uplift modeling in scikit\-learn style in Python\.[https://github\.com/maks\-sh/scikit\-uplift](https://github.com/maks-sh/scikit-uplift)\.
- MegaFon \(2021\)MegaFon\. 2021\.MegaFon Uplift Competition Dataset\.[https://www\.uplift\-modeling\.com/en/latest/api/datasets/fetch\_megafon\.html](https://www.uplift-modeling.com/en/latest/api/datasets/fetch_megafon.html)\.
- Nie and Wager \(2021\)Xinkun Nie and Stefan Wager\. 2021\.Quasi\-oracle estimation of heterogeneous treatment effects\.*Biometrika*108, 2 \(2021\), 299–319\.
- Radcliffe and Surry \(2007\)Nicholas J Radcliffe and Patrick D Surry\. 2007\.Using control groups to target on predicted lift: Building and assessing uplift models\.*Direct Marketing Analytics Journal*1 \(2007\), 14–21\.
- Rzepakowski and Jaroszewicz \(2012\)Piotr Rzepakowski and Szymon Jaroszewicz\. 2012\.Decision trees for uplift modeling with single and multiple treatments\.*Knowledge and Information Systems*32, 2 \(2012\), 303–327\.
- Shalit et al\.\(2017\)Uri Shalit, Fredrik D Johansson, and David Sontag\. 2017\.Estimating individual treatment effect: generalization bounds and algorithms\.*International Conference on Machine Learning*\(2017\), 3076–3085\.
- van der Laan et al\.\(2023\)Lars van der Laan, Ernesto Ulloa\-Rivero, Marco Carone, and Alex Luedtke\. 2023\.Causal Isotonic Calibration for Heterogeneous Treatment Effects\. In*Proceedings of the 40th International Conference on Machine Learning \(ICML\)**\(PMLR, Vol\. 202\)*\.
- Verbeken et al\.\(2025\)Brecht Verbeken, Marie\-Anne Guerry, Wouter Verbeke, and Sam Verboven\. 2025\.Uplift model evaluation with ordinal dominance graphs\.*Journal of Machine Learning Research*26 \(2025\)\.
- Wager and Athey \(2018\)Stefan Wager and Susan Athey\. 2018\.Estimation and inference of heterogeneous treatment effects using random forests\.*J\. Amer\. Statist\. Assoc\.*113, 523 \(2018\), 1228–1242\.
- X5 Group \(2021\)X5 Group\. 2021\.X5 Retail Group Uplift Dataset\.[https://www\.uplift\-modeling\.com/en/latest/api/datasets/fetch\_x5\.html](https://www.uplift-modeling.com/en/latest/api/datasets/fetch_x5.html)\.
- Yadlowsky et al\.\(2025\)Steve Yadlowsky, Scott Fleming, Nigam Shah, Emma Brunskill, and Stefan Wager\. 2025\.Evaluating Treatment Prioritization Rules via Rank\-Weighted Average Treatment Effects\.*J\. Amer\. Statist\. Assoc\.*120, 549 \(2025\), 38–51\.
- Yang et al\.\(2026\)Yuxuan Yang, Dugang Liu, and Yiyan Huang\. 2026\.Evaluating uplift modeling under structural biases: Insights into metric stability and model robustness\. In*Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining \(KDD\)*\.
- Zhu et al\.\(2025\)Minqin Zhu, Zexu Sun, Ruoxuan Xiong, Anpeng Wu, Baohong Li, Caizhi Tang, Jun Zhou, Fei Wu, and Kun Kuang\. 2025\.Rethinking causal ranking: A balanced perspective on uplift model evaluation\. In*Proceedings of the 42nd International Conference on Machine Learning \(ICML\)*\.

## Appendix ADeferred discussion and extended material

![Refer to caption](https://arxiv.org/html/2608.00915v1/x2.png)

Bar chart comparing metric\-versus\-reference\-objective correlations under LightGBM and XGBoost base learners, showing both mechanisms persist\.

Figure 2\.The findings are robust to the base learner\. Switching LightGBM→\\rightarrowXGBoost, Qini still diverges from effect accuracy \(−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}\) on IHDP \(ρ:−0\.15→−0\.33\\rho:\-0\.15\\rightarrow\-0\.33\) while AUUC still tracks it \(\+0\.68→\+0\.53\+0\.68\\rightarrow\+0\.53\) — F1 — and ranking metrics still agree with each other on binary \(\+0\.90→\+0\.91\+0\.90\\rightarrow\+0\.91\), the mutual\-agreement premise of F2\. The deployment\-objective half of F2 is retested under XGBoost in Section[5\.2](https://arxiv.org/html/2608.00915#S5.SS2): the policy\-value correlations and the cross\-repeat selection regret both replicate\.![Refer to caption](https://arxiv.org/html/2608.00915v1/x3.png)

Two panels: metric agreement with the true quality order is invariant to affine outcome transforms, and all cumulative ranking metrics degrade together as outcome kurtosis increases\.

Figure 3\.Controlled probes of the F1 failure signature\(controlled DGP; 8 fixed models of graded quality, latent CATE fixed, only the outcome distribution varies; 30 seeds\)\.\(a\)Multiplying or shifting the outcome \(y,10​y,y\+100,3​y\+50y,\\,10y,\\,y\{\+\}100,\\,3y\{\+\}50\) leaves every metric’s agreement with the true quality order unchanged — the failure is not a scale artifact \(Prop\.[D\.1](https://arxiv.org/html/2608.00915#A4.Thmtheorem1)\)\.\(b\)As outcome excess kurtosis grows, all cumulative ranking metrics lose fidelity to true quality, and Qini and AUUC degrade*together*— so the AUUC\-over\-Qini advantage on the*real*continuous data \(Fig\.[1](https://arxiv.org/html/2608.00915#S5.F1)a\) is a dataset\-specific empirical observation, not a guarantee\.##### Which model wins depends on regime and metric \(full commentary\)\.

No estimator dominates\. By the appropriate metric, simple meta\-learners and the causal forest lead the semi\-synthetic regimes, and simple estimators lead the marketing RCTs \(on the released 10K subsamples\): SoloModel on Hillstrom \(49\.1\), ClassTrans on Lenta \(20\.3\), the X\- and S\-Learner statistically tied on MegaFon \(≈\\approx40\.3\), and on the released 10K X5 subsample no evaluated model produces a materially positive Qini signal \(Qini≈\\approx0\)\. These standings are subsample\-scoped, and fold\-level standard errors say which are real at the evaluatednn:Lenta’s leader shows a large descriptive separation \(\>5\{\>\}5combined SEs\), while theHillstrom/X5/MegaFonleads are within fold noise \(Appendix[I](https://arxiv.org/html/2608.00915#A9);scripts/leaderboard\_resolution\.py\)\. Nemenyi critical\-difference diagrams \(reported*descriptively*, not as confirmatory tests\) show wide cliques: on IHDP byPEHE\\sqrt\{\\mathrm\{PEHE\}\},S\-Learnerhas the best mean rank \(2\.00, CD=2\.38=2\.38\) with DR worst; on binary data by Qini,ClassTransleads \(2\.07, CD=2\.41=2\.41\)\. These standings are protocol\-scoped: samples in the hundreds, 3\-fold CV, and aB=10B\{=\}10tuning budget can under\-serve nuisance\-heavy learners \(DR/R\), so “simple learners lead” is an observation about this regime, not a general estimator verdict \(Section[6](https://arxiv.org/html/2608.00915#S6), Step 3\)\.

##### Calibrating the threshold closes most of the gap\.

Proposition[5\.3](https://arxiv.org/html/2608.00915#S5.Thmtheorem3)predicts that the sign\-threshold gap should shrink once the score*level*is supplied\. It does: replacing the fixed zero threshold with one*calibrated on the selection folds*\(per model, the threshold minimizing the same out\-of\-sample IPW policy risk over a 41\-point score\-quantile grid on the selection seeds’ pooled out\-of\-fold predictions; the held\-out evaluation seed is untouched\) cuts the Qini\-selection regret from\+0\.0253\+0\.0253\[\+0\.0064,\+0\.0437\+0\.0064,\+0\.0437\]222This experiment recomputes its rotation baseline from the stored per\-unit predictions, so it differs in the third decimal from the selector\-ladder rotation’s0\.0280\.028\(Appendix[E](https://arxiv.org/html/2608.00915#A5)\); both are the same quantity under slightly different rotations\.to\+0\.0049\+0\.0049\[−0\.0082,\+0\.0198\-0\.0082,\+0\.0198\] — a81%81\\%reduction, with the calibrated interval no longer excluding zero\. So F2 on the sign\-threshold objective is largely a score\-*level*effect, exactly as the invariance boundary implies; the actionable message is to pair a ranking metric with a calibrated threshold rather than trustsgn​\(τ^\)\\mathrm\{sgn\}\(\\hat\{\\tau\}\)\(Section[6](https://arxiv.org/html/2608.00915#S6); full protocol in Appendix[E](https://arxiv.org/html/2608.00915#A5)\)\.

##### Mechanism probes, summarized\.

RATE and oracle residualization improve Qini only modestly \(\+0\.15\+0\.15and\+0\.22\+0\.22vs\. AUUC’s\+0\.56\+0\.56; the oracle adjustment is an optimistic oracle benchmark for outcome\-residualization\-based variance reduction; PUC and pROCini were developed for*binary*\-outcome uplift evaluation and are not directly targeted at the continuous\-outcome setting studied in F1 — evaluating compatible variants remains useful future work — whereas the variance\-reduced Qini ofBokelmann and Lessmann \([2024](https://arxiv.org/html/2608.00915#bib.bib3)\)is the highest\-priority planned extension because it directly targets Qini variance\), while controlled sweeps of kurtosis, treatment imbalance, and outcome\-correlated score error do not isolate the observed Qini/AUUC separation\. A weighting decomposition sharpens the localization considerably: by Lemma[5\.1](https://arxiv.org/html/2608.00915#S5.Thmtheorem1), Qini, the field’s shipped cumulative\-gain AUUC \(causalml/scikit\-uplift\), and our prefix\-mean AUUC are all integrals∫w​\(k\)​u​\(k\)​𝑑k\\int w\(k\)\\,u\(k\)\\,dkwithw​\(k\)=Tkw\(k\)=T\_\{k\},k​nk\\,n, and11respectively\. Scoring the identical stored predictions on all100100IHDP realizations \(our implementation checked againstcausalml\.auuc\_scoreon all54005400evaluated folds: rank agreement\+1\.00\+1\.00, within5\.3%5\.3\\%of the exactn×n\{\\times\}scaling identity\): the cumulative\-gain AUUC correlates with−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}at\+0\.73\+0\.73\[\+0\.66,\+0\.79\+0\.66,\+0\.79\] —*better*than our prefix\-mean AUUC \(\+0\.56\+0\.56\[\+0\.49,\+0\.63\+0\.49,\+0\.63\]\) — while Qini sits at≈0\{\\approx\}0\. Among the evaluated functionals the discrepancy therefore localizes not to depth weighting but to*treated\-count*weighting:TkT\_\{k\}differs fromk​n​π1k\\,n\\,\\pi\_\{1\}exactly by the treatment\-interleaving fluctuation that Lemma[5\.2](https://arxiv.org/html/2608.00915#S5.Thmtheorem2)isolates and that the compositeRRscales with\. This also validates the practical advice for the shipped implementations: cumulative\-gain AUUC tracks effect accuracy on IHDP at least as well as any variant we compute\. F1 therefore remains a bounded, benchmark\-discovered failure signature, now localized to treated\-count weighting, whose sufficient mechanism is open \(all values in Appendix[E](https://arxiv.org/html/2608.00915#A5);scripts/gain\_auuc\_check\.py\)\.

##### Estimator exclusion \(full detail\)\.

Recomputing with the DR\-Learner excluded, and with both DR\- and R\-Learners excluded \(Table[6](https://arxiv.org/html/2608.00915#A7.T6)\), Qini’s correlation with ground truth remains indistinguishable from zero in every configuration \(LightGBM:\+0\.02→\+0\.12→−0\.08\+0\.02\\rightarrow\+0\.12\\rightarrow\-0\.08; XGBoost:−0\.21→−0\.03→−0\.12\-0\.21\\rightarrow\-0\.03\\rightarrow\-0\.12\) and the pairedΔ\\Deltais positive in all six \(base learner×\\timesexclusion\) configurations, with its CI excluding zero in five of the six \(the exception is XGBoost with DR\- and R\-Learners excluded,Δ=\+0\.34\\Delta=\+0\.34\[−0\.16,\+0\.84\-0\.16,\+0\.84\]\)\. The complement bounds the claim: AUUC’s*absolute*alignment falls as unstable estimators are removed, so we state F1 relatively — AUUC consistently*more*aligned than Qini — not as a fixed absolute level\.

We name the pattern rather than leave it to be inferred\.Δ\\Deltadeclines with exclusion under both base learners \(\+0\.63→\+0\.40→\+0\.42\+0\.63\\rightarrow\+0\.40\\rightarrow\+0\.42;\+0\.69→\+0\.44→\+0\.34\+0\.69\\rightarrow\+0\.44\\rightarrow\+0\.34\), and the decline is entirely in AUUC’s alignment \(\+0\.65→\+0\.52→\+0\.34\+0\.65\\rightarrow\+0\.52\\rightarrow\+0\.34\) while Qini’s stays flat and near zero\. Part of AUUC’s advantage is therefore AUUC correctly penalizing the numerically unstable learners — whose fold\-levelPEHE\\sqrt\{\\mathrm\{PEHE\}\}reaches153,013153\{,\}013\(Appendix[D](https://arxiv.org/html/2608.00915#A4)\) — to which a rank\-only statistic is indifferent\. That is a real mechanism, not an artifact, and it does not exhaust the effect: with both unstable learners removed the LightGBM advantage is still\+0\.42\+0\.42\[\+0\.06,\+0\.80\+0\.06,\+0\.80\] on the 10\-split panel, and the wide interval there is a small\-nnartifact — on all100100realizations the excluded\-panel gaps are\+0\.28\+0\.28\[\+0\.18,\+0\.38\+0\.18,\+0\.38\] \(DR removed\) and\+0\.34\+0\.34\[\+0\.21,\+0\.46\+0\.21,\+0\.46\] \(DR and R removed\), both CIs comfortably excluding zero\.The advantage attenuates but does not vanish when the numerically unstable learners are removed\.

##### The candidate panel is Qini\-tuned\.

Every estimator’s hyperparameters are selected by inner\-fold validation Qini \(Section[4](https://arxiv.org/html/2608.00915#S4)\), including on the continuous benchmarks where we then report that Qini misranks\. This conditions the*candidate set*, not the comparison: all six metrics are computed on identical fitted models and identical out\-of\-fold predictions, so the paired AUUC\-over\-Qini gap is a statement about metrics scoring the same candidates\. It does mean the panel is not neutral — a Qini\-tuned search may favour configurations Qini rates highly\. We therefore rerun the primary continuous panel with tuning switched off entirely \(Appendix[E\.2](https://arxiv.org/html/2608.00915#A5.SS2)\): at fixed default hyperparameters, with no inner Qini search anywhere in the pipeline, the separation persists \(paired gap\+0\.55\+0\.55\[\+0\.33,\+0\.77\+0\.33,\+0\.77\]\), so F1 is not induced by Qini\-tuned selection\. We also run the intermediate case, keeping the released budget but switching the inner objective from Qini to AUUC: the gap is\+0\.56\+0\.56\[\+0\.34,\+0\.78\+0\.34,\+0\.78\], slightly smaller than as released, so letting the favoured metric choose the panel does not manufacture its advantage either\. Tuning byPEHE\\sqrt\{\\mathrm\{PEHE\}\}or policy risk remains untested and is not a like\-for\-like control \(PEHE\\sqrt\{\\mathrm\{PEHE\}\}requires the ground truth the metric substitutes for; policy risk targets a different objective\)\. The bounded budget \(B=10B\{=\}10\) still compresses candidate spread in both directions \(Appendix[E](https://arxiv.org/html/2608.00915#A5)\)\.

##### Propensity nesting is imperfect inside the tuning loop\.

For the observational and semi\-synthetic families \(IHDP, Jobs\), the propensity model is fit once on the whole outer\-training fold and its predictions are then sliced for the inner tuning folds, so inner\-validation rows influenced the propensity values used during inner training\. The deviation involves covariates and treatment assignments only — the propensity model never sees outcomes — and the inner selection criterion \(validation Qini\) does not use the propensity at all; it enters only through the fitted candidate\. It therefore affects hyperparameter*selection*only: every reported test\-fold metric uses a propensity model fit without that fold, and the RCT families use their known assignment probabilities\. Refitting propensity strictly within each inner\-training fold, with a regression test enforcing it, is planned for the next release; we do not expect it to move F1 \(rank\-based, and unchanged under estimator exclusions and base\-learner swaps\) but we have not demonstrated that\.

##### F2’s empirical half is within\-sample\.

Selection, threshold calibration, and evaluation all draw on the same underlying Jobs observations: the cross\-repeat rotation re\-partitions one sample rather than holding out fresh units, and the ten released “instances” are overlapping re\-splits of that sample, so treating them as ten exchangeable clusters overstates the independent information available \(their per\-model risk vectors correlate at\+0\.41\+0\.41; Section[5\.2](https://arxiv.org/html/2608.00915#S5.SS2)\)\. The empirical half of F2 is therefore a*descriptive within\-sample case study*of the selection cost, and its intervals describe benchmark\-split variability, not population sampling error; selection and evaluation sharing outcome noise can also flatter the direct risk selector\. The structural half \(Prop\.[5\.3](https://arxiv.org/html/2608.00915#S5.Thmtheorem3)\) does not depend on this design\. A genuinely disjoint protocol — fit on the training portion, select and calibrate on a validation subset, then estimate policy value on the untouched experimental test units the loader already exposes, accounting for repeated unit membership across the ten splits — is the committed next step\.

##### PEHE\\sqrt\{\\mathrm\{PEHE\}\}is not the sole standard — by design\.

F1 is shown againstPEHE\\sqrt\{\\mathrm\{PEHE\}\}*and*sibling ranking metrics, F2 against operational*policy risk*; the conclusions do not rest onPEHE\\sqrt\{\\mathrm\{PEHE\}\}alone\.

Table 3\.Scope delta against the closest prior work,Yang et al\.\([2026](https://arxiv.org/html/2608.00915#bib.bib38)\)\. Both papers conclude that targeting quality and effect\-estimation quality can come apart; the designs that support the conclusion differ, and the binary\-vs\-continuous contrast is observable only in a multi\-regime design\.
##### A third continuous family, and the end of the tail explanation\.

Two continuous families is a two\-point sample, so we add a third with known per\-unit effects and a response surface that differs from IHDP’s*in kind*:Revenue\-Synthetic, in which baseline spend is lognormal and the treatment effect is*multiplicative*, making the row\-level ITE heavy\-tailed by construction — the setting in which a cumulative\-sum statistic should be most exposed to single large outcomes\. We run1010realizations with the same66estimators, fold protocol, and tuning budget as the primary panel \(make repro\-f1\-revsynth; 60 cells, all completing\)\.

F1 does not replicate there, and the reason matters\.Qini tracks effect accuracy on this family \(\+0\.39\+0\.39\[\+0\.00,\+0\.73\+0\.00,\+0\.73\]\) on par with AUUC \(\+0\.37\+0\.37\[−0\.01,\+0\.70\-0\.01,\+0\.70\]\), so the paired AUUC\-over\-Qini gap vanishes:−0\.02\-0\.02\[−0\.19,\+0\.17\-0\.19,\+0\.17\], positive on only44of1010realizations\. This is not a DR\-Learner artifact — excluding it moves the gap to−0\.08\-0\.08, and excluding both DR and R to−0\.16\-0\.16— even though the DR\-Learner’sPEHE\\sqrt\{\\mathrm\{PEHE\}\}diverges here as it does on IHDP\.

Taken with the outcome distributions, this*falsifies*the heavy\-tail explanation rather than merely failing to support it\.Revenue\-Synthetic’s outcome has median excess kurtosis4444\(range1919–109109\) and Qini is fine on it; IHDP’s primary panel has median excess kurtosis−0\.4\-0\.4— it is*mild*\-tailed, comparable to ACIC — and Qini fails on it\. The family that breaks F1 is the least heavy\-tailed of the three\. \(Precisely: excess kurtosis of the*factual outcome*, median−0\.4\-0\.4over the 10 primary IHDP splits\. Across all100100realizations the median is0\.30\.3with range−1\-1–2727, so a minority of IHDP realizations are heavy\-tailed while the primary panel as a whole is not; both statements are used consistently below\.\) Heavy tails are therefore neither necessary nor sufficient for the failure, which is consistent with the controlled experiment \(Table[8](https://arxiv.org/html/2608.00915#A7.T8)\), with the covariate nulls \(Table[4](https://arxiv.org/html/2608.00915#A5.T4)\), and with nothing else we have tested\.

##### What F1 is, stated at its true scope\.

Across three continuous families with reference effects, Qini’s alignment spans\+0\.02\+0\.02\(IHDP\) to\+0\.48\+0\.48\(ACIC\), withRevenue\-Syntheticat\+0\.39\+0\.39: it is*sometimes*uninformative and*sometimes*the best of the ranking metrics, and no measured property yet predicts which case a practitioner is in — the lemma\-derived composite of Appendix[E\.1](https://arxiv.org/html/2608.00915#A5.SS1), whose per\-family medians order the three families, is the one candidate still standing, pending its controlled sweep\. That is the operationally important claim, and it is weaker than “Qini fails on continuous outcomes” while being harder to dismiss: a metric whose reliability varies unpredictably across datasets cannot be trusted unvalidated on a new one, which is exactly the diagnostic we recommend \(Section[6](https://arxiv.org/html/2608.00915#S6)\)\. F1 names the failure case and demonstrates that it occurs on a standard, widely used benchmark; it does not claim the failure is general, and Section[7](https://arxiv.org/html/2608.00915#S7)records that one of three evaluated continuous families exhibits it\.

##### A note on Qini normalization\.

The optional Qini normalization divides every model’s score within a split by the same perfect\-curve area\. When that area is strictly positive, dividing by a common positive constant*preserves*the within\-split model ranking; when it is zero the normalized score is undefined; and when it is negative the division*reverses*score orientation\. On continuous outcomes the standard perfect curve is built from outcomes treated as counts and its area is frequently non\-positive, so the normalized score is degenerate there\. We therefore report the*unnormalized*Qini throughout and make no claims based on normalized Qini\.

##### ACIC 2016: an independent boundary\-case validation family\.

ACIC provides an independent test of the finding’s scope: Qini tracks effect accuracy on its continuous settings, showing that F1 is not a universal continuous\-outcome law\. Concretely, we add1818instances from the ACIC 2016 competition\(Dorie et al\.,[2019](https://arxiv.org/html/2608.00915#bib.bib10)\)— real Collaborative Perinatal Project covariates \(n=4802n\{=\}4802\), nine heterogeneous\-effect settings stratified over response model×\\timesheterogeneity×\\timesoverlap, two replicates each, generated by the official package \(a validation family, not one of the seven primary benchmark families\)\. ACIC outcomes are continuous with excess kurtosis≈0\.4\{\\approx\}0\.4— similarly mild \(IHDP’s primary\-panel median is−0\.4\-0\.4\) — and there Qini behaves normally: rank correlation with−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}is\+0\.48\+0\.48\[\+0\.25,\+0\.68\+0\.25,\+0\.68\], close to uplift\-at\-kk\(\+0\.45\+0\.45\) and the rank\-weighted average treatment effect \(RATE\-AUTOC,\+0\.65\+0\.65\); AUUC’s is\+0\.55\+0\.55\[\+0\.34,\+0\.73\+0\.34,\+0\.73\], so the paired AUUC\-over\-Qini gap — the statistic F1 is defined by — is\+0\.11\+0\.11\[−0\.03,\+0\.25\-0\.03,\+0\.25\] on ACIC: mildly positive in point estimate, with a CI covering zero\. The breakdown is observed on the evaluated IHDP family and not on the evaluated ACIC family; the two differ in many ways, and Table[4](https://arxiv.org/html/2608.00915#A5.T4)shows tail weight does not predict the metric gap\.

##### A controlled test, and an honest boundary\.

Holding models and latent CATE fixed and varying only the outcome distribution \(Fig\.[3](https://arxiv.org/html/2608.00915#A1.F3)\), affine transforms change nothing, and rising excess kurtosis degrades*all*cumulative ranking metrics — Qini and AUUC*together*— so the controlled experiment does not by itself reproduce Qini’s isolation\. We therefore separate what is mathematically established \(the identities; affine invariance\) from the observed signature \(Qini’s isolation on the evaluated IHDP benchmark\) and from candidate mechanisms, which remain open after targeted tests \(Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3); full discussion in Appendix[D](https://arxiv.org/html/2608.00915#A4)\)\.

##### Scope of F1\.

We do not claim Qini is categorically invalid on continuous outcomes — in a controlled sweep all cumulative ranking metrics degrade together as tails thicken \(Fig\.[3](https://arxiv.org/html/2608.00915#A1.F3)\), and on both validation families Qini tracks effect accuracy \(\+0\.48\+0\.48on ACIC 2016,\+0\.39\+0\.39onRevenue\-Synthetic\); the three families form a gradient in the paired gap \(\+0\.49\+0\.49,\+0\.11\+0\.11,−0\.02\-0\.02\) rather than a binary split \(Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)\)\. We establish that on the standard continuous benchmark \(IHDP, all100100realizations\) the Qini ranking is effectively unrelated to effect accuracy while AUUC and uplift\-at\-kkstay informative, not explained by a single unstable estimator, base learner, Qini implementation variant, or the tested imbalance\-×\\times\-tails mechanism\. We report it because practitioners and libraries apply Qini to continuous outcomes and the failure is silent — and regime\-dependent, which is precisely what makes it dangerous\.

Nor is it an artifact of the six\-model rank correlation being coarse: restated as*pairwise concordance*— for each realization and each model pair, does the metric order the pair as−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}does? — Qini agrees with effect accuracy on0\.520\.52\[0\.49,0\.560\.49,0\.56\] of the15001500comparisons across all100100realizations \(15 pairs each\), indistinguishable from a coin flip, while AUUC agrees on0\.730\.73\[0\.69,0\.760\.69,0\.76\]\.

This reading is consistent withCurth et al\.\([2021](https://arxiv.org/html/2608.00915#bib.bib5)\), who argue that IHDP is idiosyncratic and that conclusions drawn on it need not generalize: our contribution is to show the idiosyncrasy is also*metric\-facing*— the field’s standard continuous benchmark silently breaks its most common evaluation metric — and that no proposed property, theirs or ours, has yet been shown to predict the metric\-specific gap ex ante; the one surviving candidate is the lemma\-derived composite whose per\-family medians track the cross\-family gradient \(Appendix[E\.1](https://arxiv.org/html/2608.00915#A5.SS1)\), pending its controlled sweep\.

##### Uncertainty: what is resampled\.

The metric\-agreement statistics \(F1, F2\) are per\-dataset Spearman correlations*across models*, and each cross\-metric claim is a mean of these over datasets\. Their CIs are*cluster bootstraps whose resampling unit is the benchmark realization*\(not the fold\), keeping all models and metrics within a resampled realization paired\. The two families differ in what a “realization” is: IHDP realizations are conditionally independent*simulated potential\-outcome draws*over fixed covariates, whereas the 10 Jobs units are the train/test*re\-splits*ofShalit et al\.\([2017](https://arxiv.org/html/2608.00915#bib.bib32)\)over the same underlying LaLonde\+\+PSID observations — exchangeable partitions of one sample, not independent draws\. The cluster bootstrap treats both as exchangeable clusters; for Jobs this can understate uncertainty to the extent of between\-split dependence, which we quantify with a design\-effect sensitivity in Section[5\.2](https://arxiv.org/html/2608.00915#S5.SS2.SSS0.Px4)\. These intervals quantify variation across realizations of the benchmark — not generalization across unrelated real\-world datasets — and the effective sample size is the number of realizations \(10 continuous; 10 Jobs\), never the fold rows \(2,052; Section[5](https://arxiv.org/html/2608.00915#S5)\)\. Within a realization, the model\-level metric means already average the 9 folds\. We also report mean pairwise Spearman correlations \(complete\-case per pair\) and Nemenyi critical\-difference diagrams\(Demšar,[2006](https://arxiv.org/html/2608.00915#bib.bib6)\)atα=0\.05\\alpha\{=\}0\.05on a comparable model/dataset group\.

baseline; our AUUC — the*prefix\-mean*AUUC; “AUUC” unqualified always means this variant — integrates the difference of cumulative*means*u​\(k\)u\(k\)minus the ATE triangle — a*shifted*convention: becauseu​\(k\)u\(k\)is a prefix mean, a random ranking scores≈ATE/2\{\\approx\}\\mathrm\{ATE\}/2under it, not0; the shift depends only on the fold’s data, so it is identical for every model and cancels from every reported statistic — and both use the population\-fraction axis; we report raw \(unnormalized\) areas \(edge\-case conventions in Appendix[D](https://arxiv.org/html/2608.00915#A4)\)\. The AUUC shipped bycausalml/scikit\-upliftinstead integrates the cumulative\-gain curvek​u​\(k\)k\\,u\(k\); we call that variant the*cumulative\-gain*AUUC and evaluate it in Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)\.

All released results use a classifier for the R\-Learner and Causal Forest treatment nuisance \(an EconML requirement fordiscrete\_treatment\); releases up to v1\.4\.0 passed a regressor, caught by an external audit, and the full benchmark was rerun under the corrected nuisance\. The fix does not change F1: the corrected panel’s gap is\+0\.63\+0\.63\[\+0\.35,\+0\.91\+0\.35,\+0\.91\] against\+0\.83\+0\.83\[\+0\.58,\+1\.09\+0\.58,\+1\.09\] for the archived pre\-fix panel, with the Causal Forest cells essentially unchanged across the fix \(cell\-levelρ=\+1\.00\\rho=\+1\.00\); full comparisons in Appendix[E](https://arxiv.org/html/2608.00915#A5)\.

TheSyntheticforest timeouts atn=2,000n\{=\}2\{,\}000reflect the per\-job wall\-clock budget interacting with theB=10×2B\{=\}10\\times 2inner tuning loop \(∼189\{\\sim\}189forest fits per fold evaluation\), not a failure at that sample size; the budget, like everything else, is fixed and released\.

## Appendix BFull practitioner’s guide

##### Continuous\-outcome diagnostic\.

Do not use Qini as the sole selector\. Bootstrap independent evaluation units \(or benchmark realizations where available; folds re\-partition the same sample and are not the inferential unit\), compare Qini and AUUC model rankings, and report their rank correlation and winner agreement\. When the rankings diverge materially, prefer an effect\-accuracy or deployment\-aligned objective where identifiable and treat the winner as unstable\.

1. Step 1\.Match the metric to the outcome type\.For*continuous*outcomes, do not use the unnormalized Qini as the*sole*model\-selection criterion without validating it against an effect\-accuracy, policy, or average\-uplift criterion\. In our IHDP benchmark it fails to track ground\-truth quality while AUUC and uplift\-at\-kkperform substantially better \(F1\) — but onRevenue\-Syntheticthe two are on par \(−0\.02\-0\.02\[−0\.19,\+0\.17\-0\.19,\+0\.17\]\), so*no standing preference between the ranking metrics is warranted*\. PreferPEHE\\sqrt\{\\mathrm\{PEHE\}\}where effects are known or simulable; otherwise compare the ranking metrics against each other on your own data and distrust the winner wherever they diverge\. For binary outcomes, Qini can be used as a ranking metric, subject to the usual variance, tie, and objective\-alignment checks\.
2. Step 2\.Match the metric to the objective\.For*ranked targeting*at a fixed cutoff, a ranking metric suffices\. For*sign\-threshold deployment*\(π​\(x\)=𝟙​\[τ^​\(x\)≥0\]\\pi\(x\)=\\mathbb\{1\}\[\\hat\{\\tau\}\(x\)\\geq 0\]\), a ranking metric alone is structurally insufficient \(Prop\.[5\.3](https://arxiv.org/html/2608.00915#S5.Thmtheorem3)\): pair it with a decision threshold calibrated on held\-out selection data, which closed most of the observed selection gap on Jobs \(Section[5\.2](https://arxiv.org/html/2608.00915#S5.SS2)\)\. For*budget\-constrained allocation*, evaluate policy value directly at the relevant budget \(F2\); if allocation magnitude depends on predicted effect size rather than rank alone, additionally assess calibration with an identified estimator appropriate to the treatment\-assignment regime \(Section[5\.4](https://arxiv.org/html/2608.00915#S5.SS4)\)\.
3. Step 3\.Respect the regime\.In our randomized, small\-fold benchmark settings, R/DR\-learners sometimes became numerically unstable while simpler learners \(ClassTrans on binary outcomes, S\-Learner\) were competitive\. Include simple baselines, and reach for orthogonal/doubly\-robust learners when nuisance estimates can be supported reliably — typically observational data with adequate sample size\.
4. Step 4\.Validate across datasets\.Single\-dataset rankings transfer only moderately \(within\-regime reference\-objectiveρ≈0\.5\\rho\\approx 0\.5\); confirm on more than one dataset\.

## Appendix CCalibration: the full identification\-stratified analysis

![Refer to caption](https://arxiv.org/html/2608.00915v1/x4.png)

Scatter of within\-dataset Qini z\-score against calibration ECE by regime, shown descriptively\.

Figure 4\.Qini \(z\-scored within dataset\) vs\. calibration ECE, by regime, shown*descriptively*\(no pooled inferential fit; observations are clustered within instances\)\. Only the randomized regimes identify ECE \(see text\)\.![Refer to caption](https://arxiv.org/html/2608.00915v1/x5.png)

Line chart of estimated policy value against targeting budget for the marketing datasets; the budget\-optimal model changes with the budget\.

Figure 5\.Policy value vs\. targeting budgetkk\(marketing RCTs\)\. The budget\-optimal model changes withkk— e\.g\. on Hillstrom the best model shifts from ClassTrans atk=0\.1k\{=\}0\.1to SoloModel atk=0\.9k\{=\}0\.9— so no single ranking fixes the deployment choice\.##### Diagnostic and its identification requirement\.

We measure uplift calibration by a bin\-based diagnostic analogous to expected calibration error \(ECE\): on the held\-out test fold, units are grouped into 10 equal\-count \(quantile\) bins of predicted uplift, and we report the bin\-size\-weighted mean absolute gap between the mean predicted uplift and the observed within\-bin uplift, excluding bins that lack a treated or a control unit\. The within\-bin treated\-minus\-control difference identifies the bin\-average uplift*only under randomized treatment*\. This holds on the marketing RCTs and the synthetic RCT; it does*not*hold on IHDP \(a confounded semi\-synthetic design\) or on Jobs \(which mixes an experimental arm with observational PSID controls\), where the contrast conflates uplift with selection\. Calibration of treatment\-effect predictions is a recognized concern with dedicated estimators\(Leng and Dimmery,[2024](https://arxiv.org/html/2608.00915#bib.bib24); van der Laan et al\.,[2023](https://arxiv.org/html/2608.00915#bib.bib33); Yadlowsky et al\.,[2025](https://arxiv.org/html/2608.00915#bib.bib37)\); ours is a coarse instance, valid only where treatment is randomized\.

##### Where it is identified, calibration does*not*disagree with Qini\.

On the55randomized instances the within\-instance Spearman correlation between Qini and−\-ECE is*positive*\(\+0\.43\+0\.43,95%95\\%cluster\-bootstrap CI\[\+0\.07,\+0\.80\]\[\+0\.07,\+0\.80\]\), and the paired ECE cost of selecting by Qini is small \(median0\.0070\.007, CI\[0\.004,0\.282\]\[0\.004,0\.282\]\)\. Mis\-calibration is therefore*not*a third instance of F2: where it can be measured without confounding, calibration is moderately*aligned*with Qini and the ECE cost of selecting by Qini is small in the median, though imprecisely bounded \(the interval’s upper end exceeds most binary\-outcome ECEs\)\. \(Pooling over all2525instances instead drives the correlation to\+0\.18\+0\.18— driven by applying an unidentified within\-bin contrast to the confounded IHDP/Jobs data; the pooled contrast and a bin\-count robustness check are in Appendix[F](https://arxiv.org/html/2608.00915#A6)\.\) An identified analysis on the semi\-synthetic datasets would need oracle\-CATE targets and stored per\-unit predictions, which we leave to the release\. The separate observation that the budget\-optimal model*changes with the budget*\(Fig\.[5](https://arxiv.org/html/2608.00915#A3.F5)\) still holds and reinforces F2: there is no single ranking a practitioner can read off in advance\.

## Appendix DF1: proofs, derivations, and detailed constructions

##### Edge\-case conventions \(as shipped\)\.

The curves are anchored at\(0,0\)\(0,0\); on any prefix withCk=0C\_\{k\}=0the gain contribution is set to0, and whereTk=0T\_\{k\}=0orCk=0C\_\{k\}=0the uplift valueu​\(k\)u\(k\)is treated as0\(undefined prefixes do not contribute\), so both integrals begin effectively once each arm appears\. Ties are broken by input order, deterministically\.

###### Proof of Lemma[5\.1](https://arxiv.org/html/2608.00915#S5.Thmtheorem1)\.

Tk​u​\(k\)\\displaystyle T\_\{k\}\\,u\(k\)=Tk​\(1Tk​∑i≤kyi​ti−1Ck​∑i≤kyi​\(1−ti\)\)\\displaystyle=T\_\{k\}\\big\(\\tfrac\{1\}\{T\_\{k\}\}\\\!\\sum\_\{i\\leq k\}y\_\{i\}t\_\{i\}\-\\tfrac\{1\}\{C\_\{k\}\}\\\!\\sum\_\{i\\leq k\}y\_\{i\}\(1\-t\_\{i\}\)\\big\)=∑i≤kyi​ti−TkCk​∑i≤kyi​\(1−ti\)=g​\(k\)\.∎\\displaystyle=\\sum\_\{i\\leq k\}y\_\{i\}t\_\{i\}\-\\tfrac\{T\_\{k\}\}\{C\_\{k\}\}\\sum\_\{i\\leq k\}y\_\{i\}\(1\-t\_\{i\}\)=g\(k\)\.\\qed

###### Proof of Lemma[5\.2](https://arxiv.org/html/2608.00915#S5.Thmtheorem2)\.

By Lemma[5\.1](https://arxiv.org/html/2608.00915#S5.Thmtheorem1),q​\(k\)=Tk​u​\(k\)−kn​Tn​u​\(n\)=Tk​\[u​\(k\)−kn​u​\(n\)\]\+kn​u​\(n\)​\(Tk−Tn\)=Tk​a​\(k\)\+kn​u​\(n\)​\(Tk−Tn\)q\(k\)=T\_\{k\}u\(k\)\-\\tfrac\{k\}\{n\}T\_\{n\}u\(n\)=T\_\{k\}\\big\[u\(k\)\-\\tfrac\{k\}\{n\}u\(n\)\\big\]\+\\tfrac\{k\}\{n\}u\(n\)\(T\_\{k\}\-T\_\{n\}\)=T\_\{k\}\\,a\(k\)\+\\tfrac\{k\}\{n\}u\(n\)\(T\_\{k\}\-T\_\{n\}\)\. ∎

Two effects therefore separate the scores: \(i\) Qini depth\-weights the uplift contrast byTkT\_\{k\}; and \(ii\) an additional term proportional to the deviation of the cumulative treated countTkT\_\{k\}from its depth\-proportional valuekn​Tn\\tfrac\{k\}\{n\}T\_\{n\}, which depends on how a model’s ranking interleaves treated and control units\. Continuous outcome magnitudes and the local treated/control composition thus enter the two integrated scores differently\.

###### Proposition D\.1 \(Affine invariance of within\-split rankings\)\.

Fix a split and replace each outcomeyiy\_\{i\}byα​yi\+β\\alpha y\_\{i\}\+\\betawithα\>0\\alpha\>0\. Then every model’sQQ,AA, and uplift\-at\-kkis multiplied byα\\alpha\(the additiveβ\\betacancels through the treated/control count correction\), so the induced ranking of models is unchanged\.

###### Proof\.

Shift:g​\(k\)→g​\(k\)\+β​\[Tk−Ck​\(Tk/Ck\)\]=g​\(k\)g\(k\)\\\!\\to\\\!g\(k\)\+\\beta\\big\[T\_\{k\}\-C\_\{k\}\\,\(T\_\{k\}/C\_\{k\}\)\\big\]=g\(k\), andu​\(k\)→u​\(k\)\+β​\(1−1\)=u​\(k\)u\(k\)\\\!\\to\\\!u\(k\)\+\\beta\(1\-1\)=u\(k\); both areas are shift\-invariant\. Scale multiplies every outcome, henceg,ug,uand their areas, byα\\alpha\. Within a split all models share the same\(α,β\)\(\\alpha,\\beta\), and a common positive factor preserves order\. ∎

Table[7](https://arxiv.org/html/2608.00915#A7.T7)confirms this empirically: multiplying or shifting the outcome leaves every metric’s agreement with the true quality order unchanged to two decimals\. The F1 failure is therefore not driven by outcome scale or location; distributional shape remains relevant but does not by itself explain the observed Qini/AUUC separation\.

##### The 12\-unit counterexample in full\.

Onn=12n\{=\}12units with balanced treatment and one large control outcome \(y=40y\{=\}40\), a near\-random model A and a near\-true model B satisfy: Qini prefers the*worse*A \(QA=−1\.18Q\_\{A\}\{=\}\-1\.18vs\.QB=−1\.63Q\_\{B\}\{=\}\-1\.63\), while both AUUC \(−3\.18\-3\.18vs\.−1\.71\-1\.71\) and−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}\(PEHE\\sqrt\{\\mathrm\{PEHE\}\}1\.241\.24vs\.0\.060\.06\) prefer B\. Deleting that one outcome flips Qini to prefer B \(−1\.54\-1\.54vs\.\+1\.13\+1\.13\), establishing the reversal is caused by outcome magnitude — not treatment imbalance or estimator instability\. Table[9](https://arxiv.org/html/2608.00915#A7.T9)lists all twelve units \(treatment, outcome, true CATE, both models’ scores and induced ranks\) so the example is verifiable by hand\.

##### The three\-level claim taxonomy in full\.

*Mathematically established:*rankings are affine\-invariant \(Prop\.[D\.1](https://arxiv.org/html/2608.00915#A4.Thmtheorem1)\) andQQintegrates a depth\-weighted version of AUUC’s contrast \(Lemma[5\.1](https://arxiv.org/html/2608.00915#S5.Thmtheorem1)\)\.*Observed signature:*on the continuous benchmark family we evaluate \(the IHDP realizations; the shippedSyntheticgenerator is binary\-outcome and belongs to the binary panel\), Qini’s ranking — unlike AUUC’s — fails to track effect accuracy, and a single large outcome can reverse Qini against it\.*Candidate explanation, not yet isolated:*exactly why AUUC remains informative there while Qini does not; the controlled experiment shows heavy tails degrade both together, so we do not claim heavy tails alone explain the separation, nor that AUUC is universally tail\-robust — its advantage on these data is a dataset\-specific empirical observation\. This is why F1 is named*outcome\-regime sensitivity*rather than a proven binary\-vs\-continuous theorem\.

##### DR\-Learner detail and the estimator\-panel caveat\.

The DR\-Learner’s nested nuisance estimation diverges on small folds \(median per\-realizationPEHE≈24\\sqrt\{\\mathrm\{PEHE\}\}\\approx 24on IHDP — with realization means reaching≈2\.9×104\\approx 2\.9\{\\times\}10^\{4\}— vs\.≈0\.9\\approx 0\.9for T\-/S\-Learner; on synthetic it degenerates toPEHE≈721\\sqrt\{\\mathrm\{PEHE\}\}\\approx 721while the CausalForest and S\-Learner lead at≈0\.09\\approx 0\.09\)\. Metric agreement is inherently evaluated over the candidate estimators, so a very different panel could yield different correlations; we claim robustness to exclusions and base\-learner swaps, not estimator\-panel independence\.

##### DR\-Learner instability\.

The DR\-Learner is EconML’sDRLearnerwith default nuisance settings \(defaultmin\_propensity; no additional clipping\)\. On IHDP’s≈450\{\\approx\}450\-unit folds the propensity model can predict near0or11, so the inverse\-weighted AIPW pseudo\-outcome — and hencePEHE\\sqrt\{\\mathrm\{PEHE\}\}— diverges on particular seeds/folds \(up to153,013153\{,\}013\) — a known numerical risk of inverse\-propensity\-weighted pseudo\-outcomes under extreme estimated propensities\(cf\. Kennedy,[2020](https://arxiv.org/html/2608.00915#bib.bib20)\)\. It is seed/fold\-dependent rather than a tuning\-budget artifact: raising the budget fromB=10B\{=\}10toB=50B\{=\}50at fixed seed does not remove it \(split\-0 mean1\.4→2\.11\.4\\rightarrow 2\.1over three folds — a small audit, so read as*not explained by*the budget rather than as invariance\)\. Because the F1 analysis is rank\-based, this magnitude does not affect any reported correlation, and the paired AUUC\-over\-Qini gap holds in all six estimator\-exclusion×\\timesbase\-learner configurations\.

## Appendix EExtended robustness and uncertainty analyses

![Refer to caption](https://arxiv.org/html/2608.00915v1/x6.png)Figure; see the caption for details\.

Figure 6\.Within IHDP, outcome kurtosis does not predict Qini’s misalignment\(100 realizations\)\. \(a\) Per\-realization rank correlation of Qini and AUUC with−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}vs\. outcome excess kurtosis; \(b\) the paired AUUC−\-Qini gap vs\. kurtosis\. Spearman\(kurtosis, Qini alignment\)=\+0\.14=\+0\.14\[−0\.07,\+0\.35\-0\.07,\+0\.35\]; Spearman\(kurtosis, gap\)=−0\.07=\-0\.07\[−0\.29,\+0\.15\-0\.29,\+0\.15\] — weakly positive and weakly negative respectively, neither useful as a separating diagnostic, so the F1 signature is not a within\-family tail effect and no scalar kurtosis threshold separates failing from non\-failing realizations\.### E\.1\.Searching for a covariate that separates the realizations

Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)reports that per\-realization outcome kurtosis does not predict Qini’s misalignment within IHDP\. Kurtosis is one summary among many, and it is a natural objection that it is simply the wrong one: the realizations differ by more than an order of magnitude in outcome*level*, and thePEHE\\sqrt\{\\mathrm\{PEHE\}\}blow\-ups concentrate on particular splits\. We therefore repeat the search with three further covariates, all computable ex ante from observed data with no ground\-truth effects: the mean outcome, its coefficient of variation, and the ratio of mean outcomes between arms\.

Two points of interpretation matter before the numbers\. First, by Prop\.[D\.1](https://arxiv.org/html/2608.00915#A4.Thmtheorem1)every metric’s within\-split ranking is invariant to an affine transform of the outcome, so outcome level or scale*per se*cannot drive F1 — rescaling a dataset changes nothing\. These covariates vary across realizations because the response surfaces differ, so any correlation identifies a property the covariate*proxies*, not an effect of scale\. The between\-arm ratio is the most interesting candidate precisely because it is arm\-asymmetric and therefore not affine\-invariant\. Second, F1 is a claim about one metric*relative to another*: a covariate that predicts how well both Qini and AUUC align merely marks a harder realization, whereas only a covariate predicting the AUUC\-over\-Qini*gap*would be a candidate mechanism\. Table[4](https://arxiv.org/html/2608.00915#A5.T4)therefore reports all three\.

The result is a clean negative on the mechanism question and a modest positive on difficulty\. Mean outcome and the arm ratio both predict alignment for both metrics — higher outcome levels go with better alignment, larger between\-arm ratios with worse — so how hard a realization is for cumulative metrics is partly foreseeable\. But no covariate predicts the gap: the arm ratio, the strongest single\-metric correlate at−0\.33\-0\.33\[−0\.49,−0\.14\-0\.49,\-0\.14\] against Qini’s alignment, moves the gap by\+0\.05\+0\.05\[−0\.15,\+0\.25\-0\.15,\+0\.25\]\. This mirrors the controlled tail experiment \(Table[8](https://arxiv.org/html/2608.00915#A7.T8)\), where thickening the tail degrades Qini and AUUC together\. With4×34\\times 3correlations tested, marginal intervals deserve scepticism; the load\-bearing reading is the null on the gap, which is the quantity F1 is about\.

##### A pre\-specified, lemma\-derived composite\.

Unlike the exploratory covariates above, one predictor was fixed — direction and specification — before being computed\. Lemma[5\.2](https://arxiv.org/html/2608.00915#S5.Thmtheorem2)’s interleaving term is scaled by the full\-sample ATE and its noise amplified by treatment imbalance, implying thatR=\|ATE\|/SD​\(τ\)×\(1−π1\)/π1R=\|\\mathrm\{ATE\}\|/\\mathrm\{SD\}\(\\tau\)\\times\\sqrt\{\(1\-\\pi\_\{1\}\)/\\pi\_\{1\}\}\(oracleSD​\(τ\)\\mathrm\{SD\}\(\\tau\), used to identify, not to recommend\) should predict worse Qini alignment and a larger gap\. Tested within family, the first half*confirms in both*:corr​\(R,ρQini\)\\mathrm\{corr\}\(R,\\rho\_\{\\mathrm\{Qini\}\}\)is−0\.22\-0\.22\[−0\.40,−0\.02\-0\.40,\-0\.02\] across the 100 IHDP realizations and−0\.55\-0\.55\[−0\.88,−0\.02\-0\.88,\-0\.02\] across the 18 ACIC instances — the first replicated ex\-ante predictor of anything in this paper\. But the gap half is null in both \(−0\.01\-0\.01\[−0\.22,\+0\.19\-0\.22,\+0\.19\];\+0\.06\+0\.06\[−0\.43,\+0\.58\-0\.43,\+0\.58\]; the ACIC gap CI is wide enough that 18 instances are simply underpowered for the gap test\), and on IHDPRRpredicts AUUC’s alignment about as strongly \(−0\.27\-0\.27\[−0\.46,−0\.06\-0\.46,\-0\.06\]; on ACIC the AUUC half is directionally weaker,−0\.50\-0\.50\[−0\.84,−0\.01\-0\.84,\-0\.01\], but its CI covers zero atn=18n\{=\}18\)\. So*within*a family, by this appendix’s own discipline,RRis a*difficulty*covariate: it says when cumulative metrics degrade, not why Qini alone decouples\.

*Across*families, however,RRdoes what no other quantity in this paper has: its per\-family medians order the three continuous families exactly as the observed gap gradient does — IHDP3\.523\.52, ACIC1\.071\.07,Revenue\-Synthetic0\.390\.39against gaps\+0\.49\+0\.49,\+0\.11\+0\.11,−0\.02\-0\.02\. Three ordered medians is, by itself, a one\-in\-six chance event, so we report this as a*candidate*family\-level separating property, not a mechanism; and unlike the direction and specification ofRR, which were fixed before computation, the family\-level comparison was specified*after*the gradient was observed, so it carries no pre\-registration credit\. But it is the first candidate to survive a test, and its*proxy*form —\|ATE^\|\|\\widehat\{\\mathrm\{ATE\}\}\|over the spread of predicted CATEs, pooled across candidates — is genuinely ex\-ante computable and tracks the oracle where estimates are stable: restricted to the four numerically stable learners, oracle\-vs\-proxy Spearman is\+0\.99\+0\.99on IHDP and\+1\.00\+1\.00on ACIC \(proxy medians4\.644\.64,1\.161\.16against oracles3\.523\.52,1\.071\.07\)\. Pooled over all six candidates it fails on IHDP \(−0\.00\-0\.00\): DR/R prediction blow\-ups dominate the predicted\-CATE spread — consistent with this paper’s DR account, and a caution that the proxy is usable only alongside an estimate\-stability check\. All of this sharply motivates the controlled sweep ofRR\(the imbalance×\\timeseffect\-to\-heterogeneity cell our five sweeps did not visit\), which is committed next\-version work\.python scripts/f1\_separating\_covariates\.pyandscripts/r19\_analyses\.pyregenerate this appendix\.

Table 4\.Searching for a realization\-level covariate that separates the IHDP realizations where Qini’s ranking tracks effect accuracy from those where it does not\. Spearman correlation across the100100realizations between each ex\-ante covariate \(computable from observed data alone\) and two alignment outcomes plus AUUC’s alignment, with 95% paired bootstrap CIs\. Only a covariate predicting the*gap*would be a candidate mechanism for F1; covariates predicting both alignments merely mark harder realizations\. By Prop\.[D\.1](https://arxiv.org/html/2608.00915#A4.Thmtheorem1)an affine transform of the outcome leaves every ranking unchanged, so outcome level or scale cannot itself drive F1; these covariates vary because the response surfaces differ, and the arm ratio is the one candidate that is not affine\-invariant\.

### E\.2\.F1 without tuning: is the separation an artifact of the Qini\-tuned panel?

Because the released protocol tunes every estimator by inner\-fold validation Qini \(Section[7](https://arxiv.org/html/2608.00915#S7)\), the candidate panel that all metrics score was itself selected under one of the metrics we scrutinize\. The sharpest test of whether this induces F1 is to remove tuning altogether\. We rerun the primary continuous panel — the same 10 IHDP splits, the same66estimators applicable to continuous outcomes, the same repeated stratified3×33\\times 3fold protocol — withtuning\.enabled=false, so every estimator runs at its library default hyperparameters and no inner Qini search occurs anywhere in the pipeline \(make repro\-f1\-notune; 60 cells, all completing\)\.

The separation persists\. Rank correlation with−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}across models, averaged over the 10 splits with a dataset\-level bootstrap CI: Qini\+0\.23\+0\.23\[−0\.07,\+0\.52\-0\.07,\+0\.52\] — again no detectable alignment — against AUUC\+0\.78\+0\.78\[\+0\.66,\+0\.89\+0\.66,\+0\.89\] and uplift\-at\-kk\+0\.77\+0\.77\[\+0\.66,\+0\.87\+0\.66,\+0\.87\]\. The load\-bearing paired AUUC\-over\-Qini gap is\+0\.55\+0\.55\[\+0\.33,\+0\.77\+0\.33,\+0\.77\], positive on99of 10 splits, with the CI excluding zero\. This is the same qualitative picture as the tuned panel and a gap of comparable magnitude, obtained with the Qini objective removed from model selection entirely\. F1 is therefore not an artifact of Qini\-tuned hyperparameter search: it is reproduced when nothing in the pipeline optimizes Qini\.

##### The other tuning arm: selecting the panel by AUUC

Removing tuning is not a neutral control either, because library defaults are not equally favourable to all six estimators\. The complementary test is to keep tuning at the released budget and*switch*the objective, selecting every candidate by inner\-fold validation AUUC — the metric F1 says is the reliable one on these data\. If the AUUC\-over\-Qini gap were an artifact of the selection criterion, this arm is where it should be largest; the mirror\-image objection, that AUUC only looks good because we never let it choose the panel, is settled here too \(make repro\-f1\-auuctune; 60 cells, all completing\)\.

The gap does not grow\. With the inner objective set to AUUC, Qini’s rank correlation with−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}is\+0\.03\+0\.03\[−0\.24,\+0\.33\-0\.24,\+0\.33\] — still no detectable alignment — against AUUC’s\+0\.59\+0\.59\[\+0\.43,\+0\.75\+0\.43,\+0\.75\], for a paired gap of\+0\.56\+0\.56\[\+0\.34,\+0\.78\+0\.34,\+0\.78\], positive on99of 10 splits\. That is slightly*smaller*than the released Qini\-tuned panel’s\+0\.63\+0\.63\[\+0\.35,\+0\.91\+0\.35,\+0\.91\] and close to the untuned arm’s\+0\.55\+0\.55\. Across all three arms — panel selected by Qini, by nothing, and by AUUC — Qini’s correlation with effect accuracy is indistinguishable from zero and AUUC’s is substantial, so the finding is invariant to which of the two metrics does the selecting\. What remains untested is tuning byPEHE\\sqrt\{\\mathrm\{PEHE\}\}or by policy risk; the first is unavailable in practice \(it needs the ground truth the metric is standing in for\) and the second targets a different objective, so neither is a like\-for\-like control\.

Table 5\.F1 under three inner\-loop selection criteria, on the same 10 IHDP splits, the same six estimators, and the same3×33\\times 3fold protocol\. Rank correlations are with−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}across models, averaged over splits with a dataset\-level bootstrap CI;Δ\\Deltais the paired AUUC\-over\-Qini gap\. The released panel is Qini\-tuned; the second arm removes tuning; the third selects candidates by AUUC at the released budget\. Qini shows no detectable alignment in every arm andΔ\\Deltastays positive with its CI excluding zero, so F1 does not depend on which metric selects the candidates\.
##### Pooled statistic \(why we do not lead with it\)\.

Collapsing Qini\-vs\-reference\-objective into a single mean rank\-correlation gives−0\.18\-0\.18over the 21 datasets with a reference objective, but its 95% bootstrap CI\[−0\.37,\+0\.01\]\[\-0\.37,\+0\.01\]crosses zero, and pooling continuous\- and binary\-outcome datasets conflates the two findings\. The decomposed, per\-finding statistics in the main text are the paper’s quantitative claims; the pooled number is reported only for completeness\.

##### F2 selector ladder, full detail\.

The oracle bound \(selecting on the evaluation seed itself\) gains\+0\.0450\+0\.0450; the reference \(risk\-minimizing\) winner is stable across rotations with modal agreement77%77\\%; Qini and AUUC pick identical winners on Jobs, consistent with their mutual agreement on binary outcomes\. Metric gains over random: Qini−0\.0070\-0\.0070\[−0\.0195,\+0\.0073\-0\.0195,\+0\.0073\], AUUC−0\.0070\-0\.0070\[−0\.0195,\+0\.0073\-0\.0195,\+0\.0073\], uplift\-at\-kk−0\.0050\-0\.0050\[−0\.0183,\+0\.0097\-0\.0183,\+0\.0097\]\.

##### F2 regret, secondary statistics\.

Per\-metric regrets: Qini0\.0280\.028\[0\.009,0\.0460\.009,0\.046\], AUUC0\.0280\.028\[0\.009,0\.0460\.009,0\.046\], uplift\-at\-kk0\.0260\.026\[0\.006,0\.0450\.006,0\.045\] \(cluster bootstrap over1010Jobs splits; random baseline0\.0210\.021\)\. The1414–15%15\\%relative figure is the regret as a fraction of the*mean*best\-candidate risk across rotations \(the per\-rotation best risk ranges0\.140\.14–0\.230\.23\)\. Paired differences vs\. random span zero for all three metrics \(\[−0\.007,\+0\.019\]\[\-0\.007,\+0\.019\],\[−0\.007,\+0\.019\]\[\-0\.007,\+0\.019\],\[−0\.010,\+0\.018\]\[\-0\.010,\+0\.018\]\); the metric\-selected model is within0\.0050\.005/0\.010\.01/0\.020\.02risk of the reference only30%30\\%/43%43\\%/43%43\\%of the time\.

##### F2 rank correlations, read carefully\.

The regret is mirrored by negative rank correlations with policy*value*: AUUC and uplift\-at\-kkcorrelate at−0\.24\-0\.24\[−0\.44,−0\.02\-0\.44,\-0\.02\] and−0\.30\-0\.30\[−0\.51,−0\.05\-0\.51,\-0\.05\] \(excluding zero\), while Qini is−0\.22\-0\.22but borderline \(\[−0\.43,\+0\.02\]\[\-0\.43,\+0\.02\]\)\. F2 therefore does not rest on the Qini correlation alone; the regret and selector\-ladder results establish the gap for all three metrics directly\. By the appropriate metric theDR\-LearnerandS\-Learnergive the lowest Jobs policy risk, in a narrow band of0\.210\.21–0\.250\.25\(leaderboards in Appendix[G](https://arxiv.org/html/2608.00915#A7)\)\.

##### causalml implementation comparison, full detail\.

Over the100100continuous\-outcome datasets \(all IHDP realizations\),causalml’s shippedqini\_score\(v0\.16\.0, default normalization\) and our canonical unnormalized definition produce near\-identical model rankings \(mean per\-dataset Spearman\+0\.99\+0\.99\[\+0\.98,\+0\.99\+0\.98,\+0\.99\]; Qini\-best model agreement100%100\\%\), and the F1 statistic is unchanged:\+0\.00\+0\.00\[−0\.10,\+0\.10\-0\.10,\+0\.10\] canonical vs\.−0\.01\-0\.01\[−0\.11,\+0\.09\-0\.11,\+0\.09\] undercausalml\. The silently computed object practitioners receive behaves exactly as the object we analyze\.

##### Budget grid, per\-budget values\.

Regret at each budget:\+0\.029\+0\.029\[\+0\.009,\+0\.051\+0\.009,\+0\.051\] \(k=0\.1k\{=\}0\.1\),\+0\.029\+0\.029\[\+0\.012,\+0\.045\+0\.012,\+0\.045\] \(k=0\.2k\{=\}0\.2\),\+0\.026\+0\.026\[\+0\.008,\+0\.044\+0\.008,\+0\.044\] \(k=0\.3k\{=\}0\.3\),\+0\.027\+0\.027\[\+0\.007,\+0\.048\+0\.007,\+0\.048\] \(k=0\.5k\{=\}0\.5\); PV\-AUC selector\+0\.024\+0\.024\[\+0\.009,\+0\.038\+0\.009,\+0\.038\]; rank stability across budgets\+0\.75\+0\.75\[\+0\.73,\+0\.77\+0\.73,\+0\.77\]\. Value\-reference rotation atk=0\.3k\{=\}0\.3: Qini regret−0\.002\-0\.002\[−0\.004,\+0\.002\-0\.004,\+0\.002\], gain over random\+0\.007\+0\.007\[\+0\.006,\+0\.009\+0\.006,\+0\.009\]\.

##### Estimation noise and split dependence \(full detail\)\.

Two structural caveats bound what the F2 intervals can claim, and we measure both\. First, the IPW policy risk is a noisy estimand on the small experimental subset: unit\-level bootstrapping the experimental units within each evaluation fold \(predictions held fixed\) gives a per\-split risk standard error with median0\.0230\.023\(10–90% range\[0\.021,0\.024\]\[0\.021,0\.024\]\) — the same order as the regret itself, so no single split is informative; the inference runs entirely through pairing and averaging \(3 rotations×\\times10 splits\)\. Second, because the Jobs splits re\-partition the same observations, the cluster bootstrap’s exchangeability assumption may understate uncertainty\. A design\-effect sensitivity makes the assumption’s role explicit: writingSEtrue=SEiid​1\+\(n−1\)​ρ¯\\mathrm\{SE\}\_\{\\mathrm\{true\}\}=\\mathrm\{SE\}\_\{\\mathrm\{iid\}\}\\sqrt\{1\+\(n\{\-\}1\)\\bar\{\\rho\}\}for average between\-split dependenceρ¯\\bar\{\\rho\}, the 95% regret interval continues to exclude zero only forρ¯≤0\.12\\bar\{\\rho\}\\leq 0\.12\(Qini and AUUC;0\.060\.06for uplift\-at\-kk\); the leave\-one\-split\-out jackknife matches the iid SE, and the regret is positive in88/10 splits\. The same caveat applies to the selector\-ladder gains above\. We therefore read F2 as a*consistently positive regret of modest statistical strength*, resting on the convergence of four diagnostics — convergent, not statistically independent, since all four read the same underlying observations — the rotation regret, the sign pattern, the negative rank correlations \(which replicate under XGBoost\), and the risk\-vs\-metric selector contrast — rather than on any single interval\.

##### Replacement metrics \(full detail\)\.

We add the rank\-weighted average treatment effect\(Yadlowsky et al\.,[2025](https://arxiv.org/html/2608.00915#bib.bib37)\)to the harness \(IPW effect scores from the fold propensities; AUTOC and Qini weightings\) and score the same stored predictions\. On the IHDP continuous panel RATE improves on Qini but only marginally: rank correlation with−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}is\+0\.15\+0\.15\[\+0\.05,\+0\.24\+0\.05,\+0\.24\] \(AUTOC\) and\+0\.17\+0\.17\[\+0\.08,\+0\.26\+0\.08,\+0\.26\] \(Qini\-weighted\) — statistically positive, unlike Qini’s\+0\.00\+0\.00, yet far below AUUC’s\+0\.56\+0\.56\[\+0\.49,\+0\.62\+0\.49,\+0\.62\] on the identical frame\. RATE alone does not restore effect\-accuracy tracking here\. Nor does variance reduction: recomputing Qini on*oracle*baseline\-adjusted outcomesy−μ0​\(x\)y\-\\mu\_\{0\}\(x\)\(an optimistic benchmark for outcome\-residualization\-based variance reduction, cf\. the corollary to Lemma[5\.2](https://arxiv.org/html/2608.00915#S5.Thmtheorem2)\) lifts the correlation only to\+0\.22\+0\.22\[\+0\.12,\+0\.31\+0\.12,\+0\.31\] on the100100IHDP realizations — the same modest level as RATE, and far from AUUC — so estimation variance of the Qini functional is not the driver of F1\. These two probes do not exhaust the proposed replacements: PUC, pROCini, and the exact variance\-reduced estimator ofBokelmann and Lessmann \([2024](https://arxiv.org/html/2608.00915#bib.bib3)\)remain unbenchmarked \(the oracle adjustment upper\-bounds the outcome\-residualization class considered here, which is the question relevant to F1\)\.

##### Mechanism\-candidate sweeps \(full detail\)\.

A concrete mechanism candidate for F1 lives in Lemma[5\.2](https://arxiv.org/html/2608.00915#S5.Thmtheorem2): the interleaving term proportional toTk−kn​TnT\_\{k\}\-\\tfrac\{k\}\{n\}T\_\{n\}multiplies the full\-sample ATE estimate, whose noise is heavy\-tailed in the controlled sweep, and treatment*imbalance*\(IHDP:π1=0\.19\\pi\_\{1\}\{=\}0\.19\) inflates those count fluctuations\. We tested it directly: a controlled sweep ofπ1∈\{0\.50,0\.35,0\.19,0\.10\}\\pi\_\{1\}\\in\\\{0\.50,0\.35,0\.19,0\.10\\\}crossed with the four outcome regimes \(identical latent CATE and model panel;200200seeds per cell\)\. The result is null in all 16 cells: the paired Qini−\-AUUC fidelity gap is never significantly negative — at the IHDP\-matched cell \(π1=0\.19\\pi\_\{1\}\{=\}0\.19, heavy log\-normal\) it is\+0\.015\+0\.015\[−0\.024,\+0\.056\-0\.024,\+0\.056\] — so imbalance×\\timestails does not reproduce the benchmark’s Qini isolation either\. This tightens the boundary of Section[5\.1](https://arxiv.org/html/2608.00915#S5.SS1.SSS0.Px4): scale artifacts, single estimators, base learners, implementation variants, kurtosis alone, and now imbalance×\\timeskurtosis are all eliminated as sole drivers\. We then tested the last candidate we considered plausible — score errors*correlated with the outcome model*\(heteroskedastic in the realized potential\-outcome magnitude\), mixed into the panel atλ∈\[0,1\]\\lambda\\in\[0,1\]and crossed with the kurtosis ladder \(200200seeds/cell\)\. It is also null, and informatively so: under correlated errors and heavy tails the paired gap moves in Qini’s*favor*\(\+0\.184\+0\.184\[\+0\.136,\+0\.230\+0\.136,\+0\.230\] atλ=0\.75\\lambda\{=\}0\.75, heavy log\-normal\), never reproducing the benchmark’s isolation in any cell\. With scale, kurtosis alone, imbalance×\\timeskurtosis, correlated errors×\\timeskurtosis, and estimation variance \(the oracle\-adjusted benchmark above\) all eliminated, we position F1 as a benchmark\-discovered failure signature whose isolating mechanism remains open despite targeted tests — the signature itself is what the multi\-regime benchmark contributes\.

##### Tuning budget \(full detail\)\.

A bounded tuning budget \(B=10B\{=\}10\) compresses the quality spread of the candidate panel, which can attenuate*all*model\-level correlations and makes small\-panel Spearman statistics noisier in both directions; we do not claim limited tuning cannot interact with individual metrics\. Two design features bound the concern: every metric scores the*identical*out\-of\-fold predictions, and the load\-bearing statistic is the paired AUUC−\-Qini gapΔ\\Deltawith its CI — robust across base learners and estimator exclusions — not any absolute correlation level\.

## Appendix FCalibration: pooled contrast and bin\-count robustness

This appendix supports Section[5\.4](https://arxiv.org/html/2608.00915#S5.SS4)\. Restricting to the55randomized instances \(where the treated\-minus\-control bin contrast identifies uplift\), Qini and−\-ECE correlate at\+0\.43\+0\.43\[\+0\.07,\+0\.80\+0\.07,\+0\.80\] and the best\-Qini and best\-ECE models differ in55/55instances with a small paired ECE cost \(median0\.0070\.007\)\. Pooling over all2525instances —*including*the confounded IHDP/Jobs datasets, where the within\-bin contrast conflates uplift with selection and the ECE is biased — instead drives the correlation to\+0\.18\+0\.18\[\+0\.02,\+0\.34\+0\.02,\+0\.34\] and inflates the winner disagreement to2121/2525\. The apparent pooled “disagreement” is thus driven by applying an unidentified within\-bin contrast to confounded datasets, not by a Qini–calibration gap\. The ECE\-induced ranking is itself bin\-count\-robust: on a controlled panel \(88models,2020seeds\) the rankings at55,1010and2020bins are identical, with every pairwise Spearman correlation equal to1\.0001\.000\.

## Appendix GLeaderboards and tables

![Refer to caption](https://arxiv.org/html/2608.00915v1/x7.png)

Per\-regime leaderboard heatmap; cell color encodes within\-dataset model rank by the regime’s appropriate metric\.

Figure 7\.Per\-regime leaderboard, each panel scored by its appropriate metric \(color = within\-dataset rank, green best\); “—” marks binary\-only models on continuous IHDP\. IHDP cells are the arithmetic mean of fold\-levelPEHE\\sqrt\{\\mathrm\{PEHE\}\}\(notmean PEHE\\sqrt\{\\text\{mean PEHE\}\}\), so single divergent folds inflate the affected DR\-Learner cells; rankings are unaffected\.![Refer to caption](https://arxiv.org/html/2608.00915v1/x8.png)Bar chart of mean within\-dataset rank stability by metric\.

\(a\)Within\-regime rank stability by metric\.
![Refer to caption](https://arxiv.org/html/2608.00915v1/x9.png)Nemenyi critical\-difference diagrams over the estimators\.

\(b\)Complete\-case critical\-difference diagrams\.

Figure 8\.Supporting analyses for Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)–[5\.2](https://arxiv.org/html/2608.00915#S5.SS2)\.Table[12](https://arxiv.org/html/2608.00915#A7.T12)lists, per dataset with a reference objective, the best model by that objective vs\. by Qini; Tables[13](https://arxiv.org/html/2608.00915#A7.T13)and[14](https://arxiv.org/html/2608.00915#A7.T14)report calibration ECE on the continuous\- and binary\-outcome datasets respectively\. ECE is tabulated per outcome family because the IHDP and Jobs realization labels both abbreviate to s0–s9; the DR\-Learner’s large IHDP ECE is consistent with the numerical instability documented above\. Reporting convention: leaderboard and ECE cells are means of the corresponding fold\-level quantity \(for IHDP, of fold\-levelPEHE\\sqrt\{\\mathrm\{PEHE\}\}\), so a single divergent fold can dominate a cell while leaving every rank\-based statistic unchanged\.

Table 6\.F1 estimator\-exclusion sensitivity on the continuous\-outcome datasets \(IHDP×\\times10\)\.ρ\\rho: mean within\-dataset Spearman correlation with−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}\.Δ\\Delta: paired per\-dataset differenceρ​\(AUUC\)−ρ​\(Qini\)\\rho\(\\mathrm\{AUUC\}\)\-\\rho\(\\mathrm\{Qini\}\)with 95% bootstrap CI\. The AUUC advantage is positive in all six configurations; its CI excludes zero in five of six \(the exception is XGBoost with DR\- and R\-Learner excluded\)\.Table 7\.Affine invariance\. Multiplying or shifting a Gaussian outcome leaves every metric’s within\-split model ranking \(and thus itsρ\\rhowith truth\) unchanged: the count/mean corrections cancel affine transforms\. The F1 failure is therefore not driven by outcome scale or location; distributional shape remains relevant but does not by itself explain the observed Qini/AUUC separation\.Table 8\.Controlled outcome\-distribution experiment\. Spearmanρ\\rho\(mean over 30 seeds,±\\pms\.e\.\) between each metric’s ranking of a fixed 8\-model panel and the*known*quality order \(−PEHE\-\\sqrt\{\\mathrm\{PEHE\}\}against the latent CATE\), with the models and latent CATE held fixed as only the outcome distribution’s tail varies\. All cumulative ranking metrics lose fidelity as the tail thickens, and Qini and AUUC degrade*together*: the AUUC\-over\-Qini advantage seen on IHDP \(Sec\.[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)\) is thus an empirical property of that dataset, not a property guaranteed by construction — and onRevenue\-Synthetic, whose outcome is far heavier\-tailed, the advantage reverses\.Table 9\.Full data for the F1 existence counterexample \(Section[5\.1](https://arxiv.org/html/2608.00915#S5.SS1)\)\. Twelve units;tttreatment,yyobserved outcome \(one large control outcome at unit 3\),τ\\tautrue CATE; models A and B are the scored uplift predictions, with their induced descending\-sort ranks\. Metrics computed by the shipped functions giveQA=−1\.18\>QB=−1\.63Q\_\{A\}\{=\}\-1\.18\{\>\}Q\_\{B\}\{=\}\-1\.63\(Qini prefers the worse A\) butAUUCA=−3\.18<AUUCB=−1\.71\\mathrm\{AUUC\}\_\{A\}\{=\}\-3\.18\{<\}\\mathrm\{AUUC\}\_\{B\}\{=\}\-1\.71andPEHEA=1\.24\>PEHEB=0\.06\\sqrt\{\\mathrm\{PEHE\}\}\_\{A\}\{=\}1\.24\{\>\}\\sqrt\{\\mathrm\{PEHE\}\}\_\{B\}\{=\}0\.06\(both prefer B\)\. Settingyyat unit 3 to0yieldsQA=−1\.54<QB=\+1\.13Q\_\{A\}\{=\}\-1\.54\{<\}Q\_\{B\}\{=\}\+1\.13\.Table 10\.Documentation of the all\-100 IHDP validation \(Section[5\.3](https://arxiv.org/html/2608.00915#S5.SS3)\)\. Splits 0–9 come from the main benchmark, 10–29 from the F1\-extension run, and 30–99 from a dedicated validation run; all use the same loaders, LightGBM base learner, outer\-test\-isolated protocol and3×33\\times 3outer CV\. The validation run uses a smaller tuning budget \(B=6B\{=\}6rather thanB=10B\{=\}10\) and restricts to the six meta/forest learners common to every run\. Correlations are complete\-case per split \(a split enters if≥4\\geq 4models have a defined score\)\.Table 11\.Audit of the F2 Jobs policy\-regret experiment under both base learners \(Section[5\.2](https://arxiv.org/html/2608.00915#S5.SS2)\)\. Regret is the cross\-repeat selection/evaluation rotation regret \(mean over the ten Jobs splits, 95% cluster bootstrap\);ρ​\(Qini,−Rpolicy\)\\rho\(\\text\{Qini\},\-R\_\{\\mathrm\{policy\}\}\)is the mean per\-split Spearman correlation with policy value\. Both halves of F2 — the ranking metrics’ negative association with policy value and the positive selection regret — replicate under XGBoost \(indeed Qini’s correlation is no longer borderline there\)\.Table 12\.Best model by reference objective vs by Qini, and their rank correlation\. The reference objective is ground\-truthPEHE\\sqrt\{\\mathrm\{PEHE\}\}on IHDP and*RCT\-estimated*policy risk on Jobs \(no per\-unit ground truth exists on Jobs\)\.Table 13\.Calibration ECE, continuous\-outcome datasets \(IHDP s0–s9\); mean across folds, lower=better\. Uplift ECE is∑bwb​\|s¯b−τ^b\|\\sum\_\{b\}w\_\{b\}\\lvert\\bar\{s\}\_\{b\}\-\\widehat\{\\tau\}\_\{b\}\\rverton the*score*scale, so it is unbounded above: the large DR\-Learner entries inherit the same AIPW pseudo\-outcome divergence documented forPEHE\\sqrt\{\\mathrm\{PEHE\}\}\(Appendix[D](https://arxiv.org/html/2608.00915#A4)\), not a separate defect\.Table 14\.Calibration ECE, binary\-outcome datasets \(Jobs s0–s9, Synthetic, marketing RCTs\); mean across folds, lower=better\. Values are not bounded by 1 even for binary outcomes, because the predicted uplift entering the score\-scale ECE is unbounded; the two DR\-Learner outliers \(Jobs s7, s9\) are that divergence\. Every ECE statistic quoted in the text is a*median*paired cost, so these cells do not drive any reported number\.
## Appendix HRun statistics

Table[15](https://arxiv.org/html/2608.00915#A8.T15)accounts for every fold\-level evaluation and every dataset×\\timesmodel cell in the released run \(completed, correctly skipped, timed out\)\. TheCriteosupplement \(Appendix[I](https://arxiv.org/html/2608.00915#A9)\) sits outside this ledger by design: it is released and versioned, but not part of the 300\-cell accounting\.

Table 15\.Benchmark run statistics \(current data\)\.
## Appendix IThe Criteo supplement

CriteoUplift v2\.1\(Diemert et al\.,[2018](https://arxiv.org/html/2608.00915#bib.bib9)\)is the field’s largest public uplift RCT: 13\.9M rows, 12 features, treatment fractionπ1=0\.85\\pi\_\{1\}\{=\}0\.85, binaryvisitoutcome with base rate0\.0470\.047\(1M tier\)\. The supplement \(repository v1\.4\.3,results\_criteo/andresults\_criteo100k/\) runs all 12 estimators under the released protocol on the 1M tier subsampled to the 10K evaluation cap for comparability with the other large RCTs, plus a 100K probe of nine of the twelve estimators \(the three uplift forests are omitted on compute grounds:∼17\{\\sim\}17minutes per cell already at 10K, scaling superlinearly\) — the largest\-scale released result\. It sits outside the 25\-instance/300\-cell ledger and enters neither F1 nor F2 \(no ground\-truth effects\)\. Table[16](https://arxiv.org/html/2608.00915#A9.T16)reports per\-model Qini with fold\-level standard errors at both scales\.

##### Subsampling design\.

The cap subsamples the*entire dataset*before outer folding \(each tier’snnis then split by the standard3×33\{\\times\}3repeated CV, so a 10K tier has≈6,666\{\\approx\}6\{,\}666training and≈3,334\{\\approx\}3\{,\}334test rows per fold\), stratified jointly by treatment×\\timesoutcome via a seededStratifiedShuffleSplit\. Both tiers draw from the same 1M tier with the same seed and identical fold protocol, but the 10K draw is*not nested*in the 100K draw \(stratified draws are not prefix\-stable across sizes\) — so the cross\-tier rank comparison spans one pair of separately generated, non\-nested stratified subsamples, and training and evaluation sizes change together across tiers\. Within a tier, all estimators share identical subsampled rows and identical folds, which is what licenses the paired per\-fold differences below\.

##### Replication \(descriptive\)\.

The binary\-regime cross\-metric rank agreement appears on Criteo as well \(Qini–AUUC\+0\.74\+0\.74, Qini–uplift@kk\+0\.80\+0\.80, AUUC–uplift@kk\+0\.92\+0\.92across the 12 estimators — inside the released binary\-panel range\); this is a descriptive external replication on one fixed stratified subsample, not a formal test\.ClassTransposts the highest point estimate at the cap \(\+11\.1±0\.9\+11\.1\\pm 0\.9\), a lead overCausalForestthat is within fold\-level noise \(paired per\-fold difference\+6\.0±4\.5\+6\.0\\pm 4\.5, positive on 6 of 9 folds\); its fold\-level variance is substantially smaller than most alternatives’ — a single classifier on the transformed label, with no nuisance stacking\. At 100K the same two lead in the same order \(ClassTrans\+96\.9±6\.0\+96\.9\\pm 6\.0vs\.CausalForest\+88\.1±8\.8\+88\.1\\pm 8\.8; paired\+8\.8±11\.0\+8\.8\\pm 11\.0, positive on 6 of 9 folds\): the capped point\-estimate winner survives the tenfold resolution increase, though the two leaders remain within noise of each other at both tiers\.

##### Resolution\.

At the 10K cap, pairwise gaps among the six meta/forest learners were small relative to fold\-to\-fold variation — descriptive non\-resolution, not an inferential equivalence result — and their rankings did not agree between the 10K and 100K probes \(Spearman−0\.03\-0\.03across those six;X\-Learnermoves last→\{\\to\}second,R\-Learnerthird→\{\\to\}last\); over all nine common estimators the correlation is\+0\.20\+0\.20, the agreement carried by the two leaders\. Because this is one subsample comparison and training size also changed, we read it as evidence consistent with limited resolution at the cap — not as the cap actively misranking\. At 100K the signal\-to\-noise ratio improves sharply \(the top point estimates’ fold\-levelttrises from≈1\.2\{\\approx\}1\.2to≈10\{\\approx\}10\), and paired per\-fold differences on identical folds \(descriptive — repeated\-CV folds are dependent, so these are not formal tests\) putCausalForestahead of every meta\-learner on 6–7 of 9 folds \(mean differences\+19\.6\+19\.6to\+44\.1\+44\.1\), while theClassTrans–CausalForestpair stays unresolved at both tiers\.

##### Confound and mechanism\.

Training and evaluation samples grew together, so isolating evaluation\-side resolution would need a train\-at\-100K/evaluate\-at\-10K contrast, which we have not run\. Criteo is especially cap\-limited by design:π1=0\.85\\pi\_\{1\}\{=\}0\.85with a control visit rate of0\.0380\.038leaves≈19\{\\approx\}19positive control events per test fold at the cap — supporting evidence for the resolution reading\.

##### Scope and licensing\.

This is why every capped leaderboard in this paper is stated as subsample\-scoped\.scripts/leaderboard\_resolution\.pyreproduces the analysis for all marketing RCTs:Lenta’s leader shows a large descriptive separation at the cap \(\>5\{\>\}5combined SEs\) while theHillstrom/X5/MegaFonleads are within fold\-to\-fold noise\. Criteo Uplift v2\.1 is distributed under Criteo Research’s own research\-use terms; the harness downloads it from Criteo’s server at build time and never redistributes it\.

Table 16\.Criteo supplement: per\-model Qini±\\pmfold\-level, descriptive SE at the 10K cap \(12 estimators\) and the 100K probe \(nine; uplift forests omitted for compute\)\. Unnormalized Qini scales withnn: compare orderings within a tier only\. Paired per\-fold differences are in the text\.

相似文章

AI模型构建者的不稳定指标与基准测试文化

arXiv cs.AI

本文介绍了Benchmarking-Cultures-25数据集,该数据集分析了AI模型构建者如何在新闻稿中选择性突出基准测试。研究发现评估格局碎片化,跨模型可比性有限,并指出基准测试更多被用作市场定位的叙事工具,而非标准化的科学测量手段。

当AI基准陷入平台期:基准饱和的系统性研究

Hacker News Top

一项系统性研究,定义并分析了60个语言模型基准中的基准饱和现象,发现近一半的基准已出现饱和,且专家精选的基准更具韧性,为构建持久评估提供了设计选择。