CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning
Summary
CRISP is a scalable coreset method for imbalanced tabular learning that efficiently reduces dataset size while maintaining high accuracy in tasks like fraud detection.
View Cached Full Text
Cached at: 09/24/26, 09:35 AM
# Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning Source: [https://arxiv.org/html/2609.26962](https://arxiv.org/html/2609.26962) ## CRISP: Scalable Importance\-Stratified Coresets for Imbalanced Tabular LearningNote:The views expressed in this article are those of the authors alone and do not necessarily reflect the views of Coinbase or its affiliates\. CCS:Computing methodologies Supervised learning by classificationCCS:Mathematics of computing Combinatorial optimizationCCS:Information systems Data miningHardhik MohantyAffiliation:University of Southern California,Los Angeles,California,USAemail:[hmohanty@usc\.edu](mailto:[email protected])Indrayana RustandiAffiliation:Coinbase,San Francisco,California,USAemail:[indrayana\.rustandi@coinbase\.com](mailto:[email protected])andMohamadreza SheibaniAffiliation:Coinbase,San Francisco,California,USAemail:[mohamadreza\.sheibani@coinbase\.com](mailto:[email protected]) ###### Abstract\. Large imbalanced tabular datasets make repeated gradient\-boosted tree training expensive\. Existing coreset methods often lose accuracy when most majority examples are removed\. We presentCRISP\(CoresetReduction viaImportance\-StratifiedPruning\), a linear\-time method that allocates a negative\-class budget across quantile strata of a proxy\-model score\. Sample weights account for unequal inclusion probabilities\. At95%95\\%negative\-class reduction on a production fraud dataset, CRISP trains on approximately1\.701\.70M of2525M rows and retains99\.7%99\.7\\%of full\-data Average Precision\. This is a93\.2%93\.2\\%reduction in total training rows\. On public CriteoPrivateAds, CRISP has the highest mean Average Precision at each tested rate from90%90\\%to99\.4%99\.4\\%majority reduction\. Sparkov results are mixed at lower rates, but CRISP has the highest mean at99\.2%99\.2\\%and99\.4%99\.4\\%\. Ablations identify budget allocation and inverse\-propensity weighting as the main sources of the production\-dataset gain\. ###### Keywords: coreset selection, class imbalance, fraud detection, gradient boosting, inverse propensity weighting ## 1\.Introduction Large tabular datasets create a recurring cost rather than a one\-time training expense\. A production model is retrained as labels mature, transaction patterns change, and new features become available\. Each candidate model may also be trained many times during validation, calibration, and hyperparameter search\. U\.S\. ACH volume rose from11\.611\.6billion transactions in 2014 to21\.621\.6billion in 2025\([Board of Governors of the Federal Reserve System, 2025](https://arxiv.org/html/2609.26962#bib.bib3)\), illustrating the growth faced by financial systems\. Similar scale appears in advertising and commerce\. Gradient\-boosted decision trees \(GBDTs\) such as XGBoost\([Chen and Guestrin, 2016](https://arxiv.org/html/2609.26962#bib.bib5)\)and LightGBM\([Ke et al\., 2017](https://arxiv.org/html/2609.26962#bib.bib12)\)remain strong choices for these data, but repeated training on tens of millions of rows can dominate the development cycle\. The cost is not distributed evenly across examples\. In fraud and conversion prediction, most rows belong to the negative class and are easy to classify\. Many occupy dense interior regions that add little new information once the model has seen enough similar examples\. The rare positive class and the smaller set of boundary negatives carry much more of the signal needed to distinguish classes\. Uniformly discarding rows ignores this difference\. Uniform sampling is attractive because it is simple and preserves the population distribution in expectation, yet it removes informative and redundant negatives at the same rate\. A coreset attempts to retain the information needed for training in a smaller weighted subset\([Feldman, 2020](https://arxiv.org/html/2609.26962#bib.bib8)\)\. We focus on one\-shot selection, where the subset is created before target\-model training and reused across later runs\. Reuse matters in practice\. A selector that costs as much as the full hyperparameter sweep offers little benefit, while a selector that can be amortized across models may save substantial computation\. This requirement favors methods with one lightweight proxy\-model fit and a small number of passes over the data\. Class imbalance places a strong constraint on subset selection for fraud and conversion data\. Removing positives is risky when they represent less than2%2\\%of the data\([Dal Pozzolo et al\., 2015](https://arxiv.org/html/2609.26962#bib.bib7);[Mienye and Jere, 2024](https://arxiv.org/html/2609.26962#bib.bib17)\)\. Our pipeline therefore retains all positive examples and applies the reduction budget to negatives only\. This distinction affects how reduction is reported\. A95%95\\%negative\-class reduction is not a95%95\\%reduction in all rows because the positive class remains intact\. We report both the selector’s negative\-class rate and the corresponding target\-training size\. Existing selection strategies make different tradeoffs between coverage, model awareness, and cost\. Geometry\-based methods preserve feature\-space coverage but can become expensive or unreliable in high dimensions\. Gradient\-matching methods use a more direct learning objective, although pairwise similarities and iterative updates limit their use at production scale\([Mirzasoleiman et al\., 2020](https://arxiv.org/html/2609.26962#bib.bib18);[Killamsetty et al\., 2021b](https://arxiv.org/html/2609.26962#bib.bib13)\)\. Training\-dynamics methods identify easy, ambiguous, and hard regions from predictions collected across checkpoints\([Hadar et al\., 2024](https://arxiv.org/html/2609.26962#bib.bib9);[Swayamdipta et al\., 2020](https://arxiv.org/html/2609.26962#bib.bib24)\)\. They are informative but require repeated inference\. Coverage\-Centric Coreset Selection \(CCS\) avoids concentrating the subset on only the hardest examples by sampling across score strata\([Zheng et al\., 2023](https://arxiv.org/html/2609.26962#bib.bib30)\)\. Its uniform allocation, however, gives the same budget to strata whose estimated learning value may differ sharply\. Direct importance sampling addresses allocation but can lose coverage\. At high reduction rates, most of the selected rows may come from a narrow upper tail of the score distribution\. Zheng et al\. report that this behavior can perform worse than random sampling\([Zheng et al\., 2023](https://arxiv.org/html/2609.26962#bib.bib30)\)\. Unequal sampling also changes the effective training distribution unless the target learner receives appropriate weights\. These two issues motivate the design of CRISP: preserve broad score coverage through quantile strata, then spend more of the budget on strata with larger proxy\-model scores and correct unequal inclusion with inverse\-propensity weights\. CRISP trains a lightweight proxy GBDT whose sole role is to assign importance scores to negative rows before target\-model training\. The proxy model is separate from the downstream model and is used only during coreset construction\. Equal\-count score quantiles prevent a skewed score distribution from collapsing into one dominant bin\. The mean score of each stratum determines its share of the negative budget\. Sampling then occurs within each stratum, and selected negatives receive clipped inverse inclusion weights\. Only the fixed\-size proxy training sample is collected centrally\. Scoring, quantiles, aggregation, and sampling remain distributed, giving linear cost in the number of training rows\. We evaluate this design on three datasets with different scales and domains\. The proprietary production dataset contains about2525million rows, of which1\.9%1\.9\\%are positive examples\. This proportion describes the constructed experimental training dataset and is not an overall or population\-level payment\-reversal rate\. CriteoPrivateAds provides a public advertising benchmark with roughly8585million training rows\([Sebbar et al\., 2025](https://arxiv.org/html/2609.26962#bib.bib21)\)\. Sparkov\([Shenoy, 2020](https://arxiv.org/html/2609.26962#bib.bib23)\)supplies a smaller synthetic fraud benchmark where additional baselines are feasible\. On the production dataset, CRISP retains99\.7%99\.7\\%of full\-data monthly AP after removing95%95\\%of negatives, while using approximately6\.8%6\.8\\%of the original rows\. On CriteoPrivateAds, CRISP has the highest mean AP at every tested rate\. Sparkov is less uniform and exposes the sensitivity of five\-run averages to outliers\. The contributions are: 1. \(1\)A linear\-time, importance\-allocated stratified sampling framework for highly imbalanced tabular data, with explicit accounting for the positive\-class retention floor\. 2. \(2\)An estimator analysis that separates the unbiased, unclipped inverse\-propensity estimator from the biased clipped estimator used for stable GBDT training\. 3. \(3\)Production\-scale evidence at up to95%95\\%negative\-class reduction, public cross\-domain evidence from CriteoPrivateAds, and component ablations covering stratification, score construction, clipping, and proxy\-model capacity\. 4. \(4\)Row\-count tradeoff plots and public benchmark configurations that distinguish target\-training size from end\-to\-end selection cost\. CRISP is designed for high reduction rates, where a small negative budget makes allocation and coverage especially important\. The experiments emphasize this regime while retaining lower\-rate results where available\. Section[2](https://arxiv.org/html/2609.26962#S2)reviews coreset selection, training\-dynamics methods, and unequal\-probability sampling\. Section[3](https://arxiv.org/html/2609.26962#S3)presents CRISP and its distributed implementation, followed by the estimator properties in Section[4](https://arxiv.org/html/2609.26962#S4)\. Section[5](https://arxiv.org/html/2609.26962#S5)describes the datasets, baselines, ablations, and cross\-dataset results\. Sections[6](https://arxiv.org/html/2609.26962#S6)and[7](https://arxiv.org/html/2609.26962#S7)discuss reproducibility, limitations, and ethical considerations, and Section[8](https://arxiv.org/html/2609.26962#S8)concludes the paper\. ## 2\.Related Work Imbalanced instance selection\.Early instance\-selection methods remove majority examples with local geometric rules\. Condensed Nearest Neighbors retains a small set that preserves nearest\-neighbor decisions, while Tomek Links removes overlapping pairs near a class boundary\([Hart, 1968](https://arxiv.org/html/2609.26962#bib.bib10);[Tomek, 1976](https://arxiv.org/html/2609.26962#bib.bib25)\)\. NearMiss instead keeps majority points close to minority examples\([Mani and Zhang, 2003](https://arxiv.org/html/2609.26962#bib.bib15)\)\. These methods are intuitive on small metric datasets, but distance computations become expensive and less informative as dimension and sample count grow\. LSH\-based selection reduces this cost by replacing exact neighborhoods with hash\-based approximations\([Melo\-Acosta et al\., 2022](https://arxiv.org/html/2609.26962#bib.bib16)\)\. Its selection criterion is still independent of the model eventually trained on the subset\. Proxy and geometry\-based coresets\.Active\-learning coresets use feature\-space coverage to select points that are far from the current labeled set\([Sener and Savarese, 2018](https://arxiv.org/html/2609.26962#bib.bib22)\)\. Selection via Proxy reduces the cost of repeated ranking by using a smaller proxy model to score examples\([Coleman et al\., 2020](https://arxiv.org/html/2609.26962#bib.bib6)\)\. FCTR combines gradient approximations with clustering and product quantization, reporting a three\- to ten\-fold selection speedup over gradient\-based competitors\([Chai et al\., 2023](https://arxiv.org/html/2609.26962#bib.bib4)\)\. More recent approaches reconstruct a learned decision boundary\([Yang et al\., 2024](https://arxiv.org/html/2609.26962#bib.bib29)\)or optimize the smallest subset that satisfies a model\-performance constraint\([Xia et al\., 2024](https://arxiv.org/html/2609.26962#bib.bib28)\)\. These methods strengthen geometric or optimization objectives, but most were evaluated outside large\-scale tabular settings and at moderate pruning rates\. Training dynamics and sample importance\.Example forgetting counts how often a model changes an example from correct to incorrect during training\([Toneva et al\., 2019](https://arxiv.org/html/2609.26962#bib.bib26)\)\. Dataset Cartography summarizes confidence and variability across checkpoints to identify easy, ambiguous, and hard regions\([Swayamdipta et al\., 2020](https://arxiv.org/html/2609.26962#bib.bib24)\)\. GraNd and EL2N estimate importance from early loss gradients or prediction errors\([Paul et al\., 2021](https://arxiv.org/html/2609.26962#bib.bib19)\)\. TracIn measures influence through gradients saved at several training checkpoints\([Pruthi et al\., 2020](https://arxiv.org/html/2609.26962#bib.bib20)\)\. SAMIS uses the discrepancy between sharpness\-aware and standard optimization as an estimate of memorization\([Agiollo et al\., 2024](https://arxiv.org/html/2609.26962#bib.bib2)\)\. These scores can identify useful or atypical examples, although collecting trajectories or training multiple models can cost as much as the downstream training that pruning is meant to avoid\. Gradient and validation matching\.CRAIG selects a subset whose aggregate gradient approximates the full\-data gradient\([Mirzasoleiman et al\., 2020](https://arxiv.org/html/2609.26962#bib.bib18)\)\. GradMatch directly minimizes gradient mismatch, and GLISTER optimizes a validation objective\([Killamsetty et al\., 2021b](https://arxiv.org/html/2609.26962#bib.bib13);[Killamsetty et al\., 2021a](https://arxiv.org/html/2609.26962#bib.bib14)\)\. These methods give a clearer optimization target than scalar difficulty scores\. Their pairwise similarities, repeated gradient evaluations, or iterative subset updates limit their use on tens of millions of rows\. A comparison on a small downsample can still be informative, but it does not answer whether the selector can run on the production\-scale dataset itself\. High reduction and tabular models\.CCS addresses a failure that appears when a score\-based selector keeps only the hardest examples\. It divides examples into difficulty strata and allocates budget uniformly so that easy and moderate examples remain represented\([Zheng et al\., 2023](https://arxiv.org/html/2609.26962#bib.bib30)\)\. Uniform allocation protects coverage, but it does not distinguish between strata when the final budget is very small\. CoreTab adapts training\-dynamics ideas to GBDTs by building datamaps from predictions at several boosting checkpoints\([Hadar et al\., 2024](https://arxiv.org/html/2609.26962#bib.bib9)\)\. This is one of the few coreset methods designed for modern tabular models, though its repeated prediction passes are costly on the largest datasets\. Connection to survey sampling\.Unequal\-probability sampling has long used inverse inclusion weights to estimate a finite\-population total\([Horvitz and Thompson, 1952](https://arxiv.org/html/2609.26962#bib.bib11)\)\. CRISP follows this principle after allocating a negative\-class budget across proxy\-score strata\. It differs from CCS in the allocation rule and from direct importance sampling in its use of quantile coverage\. It also differs from gradient\-matching coresets because it never forms pairwise gradient similarities\. The practical gap addressed here is narrower than general coreset construction: a reusable subset for imbalanced tabular GBDTs, selected with linear passes at negative\-class reduction rates above90%90\\%\. Weight clipping means the deployed estimator is not exactly unbiased, so Section[4](https://arxiv.org/html/2609.26962#S4)states the resulting bias rather than claiming a general variance guarantee\. ## 3\.Methodology: CRISP ### 3\.1\.Problem Formulation Let𝒱=\{\(xi,yi\)\}i=1n\\mathcal\{V\}=\\\{\(x\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{n\}withyi∈\{0,1\}y\_\{i\}\\in\\\{0,1\\\}\. The positive and negative index sets are𝒞\+\\mathcal\{C\}\_\{\+\}and𝒞−\\mathcal\{C\}\_\{\-\}, with sizesn\+n\_\{\+\}andn−n\_\{\-\}\. CRISP retains every positive row and reduces only the negative class\. For a negative\-class reduction rater−r\_\{\-\}, the target negative budget is \(1\)k−=⌊\(1−r−\)n−⌋\.k\_\{\-\}=\\left\\lfloor\(1\-r\_\{\-\}\)n\_\{\-\}\\right\\rfloor\.The expected training size is thereforen\+\+k−n\_\{\+\}\+k\_\{\-\}rather than\(1−r−\)n\(1\-r\_\{\-\}\)n\. This distinction matters at high reduction rates because positives become a large share of the retained rows\. ### 3\.2\.Budget Interpretation Negative\-class reduction and total\-row reduction are related but not identical\. If the negative budget is met exactly, the fraction of all rows removed is \(2\)rall=r−n−n\.r\_\{\\mathrm\{all\}\}=r\_\{\-\}\\frac\{n\_\{\-\}\}\{n\}\.The retained positive fraction places a floor on target\-training size\. On the production dataset,95%95\\%negative reduction removes about93\.2%93\.2\\%of all rows because positives account for roughly1\.9%1\.9\\%of the experimental training table\. The difference is larger on Sparkov and Criteo at rates above99%99\\%, where retained positives can outnumber retained negatives\. Table 1\.Training rows implied by the retain\-all\-positives rule\.We user−r\_\{\-\}throughout the paper because it is the parameter passed to every selector\. Figures that discuss training cost use the derived total rows instead\. This keeps method comparisons at the same negative budget while making the actual target\-model input size visible\. ### 3\.3\.Proxy Model and Importance Scores A lightweight proxy GBDT is trained on at mostNsubN\_\{\\mathrm\{sub\}\}negatives and all positives\. This model is used only to rank examples for selection and is not reused as the downstream classifier\. Class weighting is used during its fit\. Its outputpip\_\{i\}is a ranking score for fraud risk, not a calibrated estimate of the population fraud rate\. For a negative row, CRISP computes \(3\)si=max\{ϵ,αpi\+βpi\(1−pi\)\},i∈𝒞−,s\_\{i\}=\\max\\\{\\epsilon,\\alpha p\_\{i\}\+\\beta p\_\{i\}\(1\-p\_\{i\}\)\\\},\\qquad i\\in\\mathcal\{C\}\_\{\-\},whereϵ\>0\\epsilon\>0prevents zero importance scores\. A strictly positive inclusion probability also requires the rounded stratum allocationkqk\_\{q\}to be positive\. The experiments useα=β=1\\alpha=\\beta=1\. In that settings\(p\)=2p−p2s\(p\)=2p\-p^\{2\}is monotone forp∈\[0,1\]p\\in\[0,1\]\. The proxy\-model ranking is therefore driven bypip\_\{i\}, while the nonlinear transformation changes the relative budget assigned to score strata\. ### 3\.4\.Budget Allocation and Sampling The negative rows are partitioned intoQQequal\-count quantile strataB1,…,BQB\_\{1\},\\ldots,B\_\{Q\}\. Lets¯q\\bar\{s\}\_\{q\}be the mean score in stratumqq\. CRISP assigns \(4\)kq∝s¯q,∑q=1Qkq=k−,k\_\{q\}\\propto\\bar\{s\}\_\{q\},\\qquad\\sum\_\{q=1\}^\{Q\}k\_\{q\}=k\_\{\-\},with integer rounding and redistribution whenkq\>\|Bq\|k\_\{q\}\>\|B\_\{q\}\|\. Within a stratum, defineui=siγu\_\{i\}=s\_\{i\}^\{\\gamma\}\. The inclusion probability used by the evaluated Bernoulli sampler is \(5\)πi=min\(1,kqui∑j∈Bquj\)\.\\pi\_\{i\}=\\min\\\!\\left\(1,\\frac\{k\_\{q\}u\_\{i\}\}\{\\sum\_\{j\\in B\_\{q\}\}u\_\{j\}\}\\right\)\.The production implementation usesγ=1\\gamma=1, which favors larger scores inside each stratum\. Sparkov and CriteoPrivateAds useγ=0\\gamma=0, which samples uniformly inside each stratum\. All variants share the same quantile construction and importance\-based allocation\. Since sampling is Bernoulli,kqk\_\{q\}is an expected budget before probability saturation and the realized coreset size varies by seed\. Selected negatives receive a clipped inverse inclusion weight \(6\)w~i=min\{1/πi,wmax\}\.\\widetilde\{w\}\_\{i\}=\\min\\\{1/\\pi\_\{i\},w\_\{\\max\}\\\}\.All positives receive weight one\. The target GBDT consumes these values as sample weights\. Algorithm 1CRISP: Coreset SelectionInput:Dataset𝒱\\mathcal\{V\}, negative reductionr−r\_\{\-\}, strataQQ, score weightsα,β\\alpha,\\beta, within\-stratum exponentγ\\gamma Output:Coreset 𝒮\\mathcal\{S\}with weights \{wi\}\\\{w\_\{i\}\\\} 1 𝒞\+,𝒞−←Partition\(𝒱\)\\mathcal\{C\}\_\{\+\},\\mathcal\{C\}\_\{\-\}\\leftarrow\\textsc\{Partition\}\(\\mathcal\{V\}\); 2 k−←⌊\(1−r−\)⋅\|𝒞−\|⌋k\_\{\-\}\\leftarrow\\lfloor\(1\-r\_\{\-\}\)\\cdot\|\\mathcal\{C\}\_\{\-\}\|\\rfloor 3 fproxy←TrainProxy\(subsample of𝒱\)f\_\{\\mathrm\{proxy\}\}\\leftarrow\\textsc\{TrainProxy\}\(\\text\{subsample of \}\\mathcal\{V\}\) 4foreach*i∈𝒞−i\\in\\mathcal\{C\}\_\{\-\}*do 5 pi←fproxy\(xi\)p\_\{i\}\\leftarrow f\_\{\\mathrm\{proxy\}\}\(x\_\{i\}\); 6 si←max\{ϵ,αpi\+βpi\(1−pi\)\}s\_\{i\}\\leftarrow\\max\\\{\\epsilon,\\alpha p\_\{i\}\+\\beta p\_\{i\}\(1\-p\_\{i\}\)\\\} 7 \{Bq\}←QuantilePartition\(\{si\},Q\)\\\{B\_\{q\}\\\}\\leftarrow\\textsc\{QuantilePartition\}\(\\\{s\_\{i\}\\\},Q\) 8foreach*stratumBqB\_\{q\}*do 9 kq←k−⋅s¯q/∑q′s¯q′k\_\{q\}\\leftarrow k\_\{\-\}\\cdot\\bar\{s\}\_\{q\}/\\sum\_\{q^\{\\prime\}\}\\bar\{s\}\_\{q^\{\\prime\}\} 10foreach*i∈Bqi\\in B\_\{q\}*do 11 πi←min\(1,kqsiγ/∑j∈Bqsjγ\)\\pi\_\{i\}\\leftarrow\\min\(1,k\_\{q\}s\_\{i\}^\{\\gamma\}/\\sum\_\{j\\in B\_\{q\}\}s\_\{j\}^\{\\gamma\}\) 12Sample iiw\.p\. πi\\pi\_\{i\}; set wi←min\(1/πi,wmax\)w\_\{i\}\\leftarrow\\min\(1/\\pi\_\{i\},w\_\{\\max\}\) 13return*𝒞\+∪SelectedNegatives,\{wi\}\\mathcal\{C\}\_\{\+\}\\cup\\textsc\{SelectedNegatives\},\\\{w\_\{i\}\\\}* ### 3\.5\.Distributed Implementation Figure[1](https://arxiv.org/html/2609.26962#S3.F1)summarizes the selection pipeline\. The selector keeps the full negative table in Spark\. A fixed\-size negative sample and all positives are transferred to the driver for proxy\-model fitting\. The fitted model is then broadcast to the workers, which score negative rows in vectorized batches\. Approximate quantiles form the stratum boundaries\. A grouped aggregation returns each stratum’s row count, mean score, and score sum\. Since this summary contains onlyQQrows, the allocation table can be broadcast back to the workers before sampling\. This design avoids collecting the full score vector on one machine\. It also makes the distinction between target and realized size explicit\. The allocation routine produces the targetkqk\_\{q\}for each stratum, while Bernoulli sampling determines the count observed in a given run\. The implementation records both the target and realized negative counts, along with the range of sample weights\. These diagnostics are especially important when the positive class forms a substantial fraction of the final training set\. Complexity\.Proxy\-model training uses a sample capped independently ofnn\. Scoring, approximate quantiles, aggregation, and sampling require a small constant number of distributed passes\. The overall selection cost is therefore𝒪\(n\)\\mathcal\{O\}\(n\)in the number of training rows, with memory distributed across the Spark workers\. Figure 1\.CRISP selection pipeline\. The input rater−r\_\{\-\}reduces the negative class only\. Quantile strata preserve score coverage, while their mean scores determine budget\. The realized negative count is random under Bernoulli sampling\.CRISP keeps all positive rows, scores negative rows with a proxy model, forms score quantiles, allocates the negative budget by mean score, samples within each stratum, and applies clipped inverse\-propensity weights before target\-model training\. ## 4\.Design Rationale and Estimator Properties ### 4\.1\.Coverage Under a Small Budget Proxy\-model scores on imbalanced data are typically concentrated near zero with a much smaller upper tail\. Equal\-width bins can therefore place most negatives into one stratum and leave several bins nearly empty\. Equal\-count quantiles avoid that collapse\. Every stratum starts with a comparable number of candidates, and the allocation step decides how many to retain from each part of the ranking\. This construction also separates two decisions that direct importance sampling combines\. The strata determine which score regions remain represented\. The allocation determines how much resolution each region receives\. A high\-score stratum may receive most of the budget, but low and moderate strata retain nonzero inclusion probabilities\. TheQ=1Q=1ablation removes this separation and recovers the direct importance\-sampling behavior discussed by Zheng et al\.\([Zheng et al\., 2023](https://arxiv.org/html/2609.26962#bib.bib30)\)\. ### 4\.2\.Weighting and Estimator Scope The weighting argument follows unequal\-probability survey sampling\([Horvitz and Thompson, 1952](https://arxiv.org/html/2609.26962#bib.bib11)\)\. Fix the dataset𝒱\\mathcal\{V\}and model parameterθ\\theta\. Forℓi\(θ\)=ℓ\(fθ\(xi\),yi\)\\ell\_\{i\}\(\\theta\)=\\ell\(f\_\{\\theta\}\(x\_\{i\}\),y\_\{i\}\), define the full\-data empirical risk \(7\)L𝒱\(θ\)=1n∑i=1nℓi\(θ\)\.L\_\{\\mathcal\{V\}\}\(\\theta\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\ell\_\{i\}\(\\theta\)\.Let𝒟\\mathcal\{D\}denote the realized sampling design, including the fitted proxy model, scores, strata, rounded allocations, and resulting inclusion probabilities\. ###### Proposition 0 \(Conditional unbiasedness of unclipped IPW\)\. LetIiI\_\{i\}indicate whether negative rowiiis selected\. Assumeπi∈\(0,1\]\\pi\_\{i\}\\in\(0,1\]almost surely and𝔼\[Ii∣𝒟\]=πi\\mathbb\{E\}\[I\_\{i\}\\mid\\mathcal\{D\}\]=\\pi\_\{i\}for eachi∈𝒞−i\\in\\mathcal\{C\}\_\{\-\}\. If selected negatives usewi=1/πiw\_\{i\}=1/\\pi\_\{i\}, then \(8\)L^\(θ\)=1n\(∑i∈𝒞\+ℓi\(θ\)\+∑i∈𝒞−Iiℓi\(θ\)πi\)\\widehat\{L\}\(\\theta\)=\\frac\{1\}\{n\}\\left\(\\sum\_\{i\\in\\mathcal\{C\}\_\{\+\}\}\\ell\_\{i\}\(\\theta\)\+\\sum\_\{i\\in\\mathcal\{C\}\_\{\-\}\}I\_\{i\}\\frac\{\\ell\_\{i\}\(\\theta\)\}\{\\pi\_\{i\}\}\\right\)satisfies \(9\)𝔼\[L^\(θ\)∣𝒟\]=L𝒱\(θ\),\\mathbb\{E\}\[\\widehat\{L\}\(\\theta\)\\mid\\mathcal\{D\}\]=L\_\{\\mathcal\{V\}\}\(\\theta\),and therefore𝔼\[L^\(θ\)\]=L𝒱\(θ\)\\mathbb\{E\}\[\\widehat\{L\}\(\\theta\)\]=L\_\{\\mathcal\{V\}\}\(\\theta\)\. The result uses only first\-order inclusion probabilities and does not require independent selections\. The deployed estimator clips each negative weight atwmaxw\_\{\\max\}\. For a nonnegative loss, its conditional downward bias is \(10\)L𝒱\(θ\)−𝔼\[L^clip\(θ\)∣𝒟\]=1n∑i:πi<1/wmax\(1−πiwmax\)ℓi\(θ\)\.L\_\{\\mathcal\{V\}\}\(\\theta\)\-\\mathbb\{E\}\[\\widehat\{L\}\_\{\\mathrm\{clip\}\}\(\\theta\)\\mid\\mathcal\{D\}\]=\\frac\{1\}\{n\}\\sum\_\{i:\\pi\_\{i\}<1/w\_\{\\max\}\}\\left\(1\-\\pi\_\{i\}w\_\{\\max\}\\right\)\\ell\_\{i\}\(\\theta\)\.Clipping is therefore a stability choice rather than an unbiased correction\. The ablation in Section[5\.3](https://arxiv.org/html/2609.26962#S5.SS3)measures its empirical effect\. Proposition[1](https://arxiv.org/html/2609.26962#S4.Thmtheorem1)concerns empirical loss at a fixedθ\\theta\. It does not imply that training on the weighted sample produces the same model, nor does it guarantee AP retention after optimization\. CRISP is therefore a task\-aware subset selector rather than a strong coreset that approximates every parameter setting\. The empirical evaluation tests whether the weaker estimator property is useful for downstream GBDT training\. ### 4\.3\.Allocation as a Practical Heuristic Classical stratified sampling allocates more rows to strata with larger within\-stratum variation\. The optimal allocation depends on quantities that are unavailable before target training\. CRISP substitutes the mean proxy\-model score for the unknown variation\. This resembles Neyman allocation when score and loss variation are correlated, but it is not an estimate of the exact Neyman solution\. The design makes a testable prediction\. Importance\-aware allocation should matter little when the budget is large and should separate from uniform allocation ask−k\_\{\-\}shrinks\. The production dataset and Criteo follow that pattern\. Sparkov is less consistent, which is evidence against treating the allocation rule as universally optimal\. The uniform\-allocation baseline, theQ=1Q=1ablation, and the no\-weight ablation test different parts of this prediction\. ## 5\.Experiments Table[2](https://arxiv.org/html/2609.26962#S5.T2)lists the three evaluation datasets\. The production dataset contains financial transactions\. Sparkov is synthetic fraud data, and CriteoPrivateAds is a public advertising conversion dataset\. Table 2\.Training\-set statistics\. Criteo positives are conversions\. Production and Criteo counts are approximate; production counts describe the constructed experimental training dataset rather than a population\-level rate\.### 5\.1\.Datasets and Protocol Production dataset\.The production task predicts payment reversal fraud\. Training covers April through September 2025, followed by temporal calibration and validation windows through January 2026\. The reported metrics are Average Precision on monthly and weekly temporal holds, written APmonthand APweek\. Each selected dataset receives a FLAML zero\-shot LightGBM configuration\([Wang et al\., 2021](https://arxiv.org/html/2609.26962#bib.bib27)\)and follows the same calibration procedure\. Because the zero\-shot suggestion is computed from the selected data, target hyperparameters can differ across methods\. We state this source of variation rather than treating the comparison as fixed\-hyperparameter isolation\. Sparkov\.We use the dataset’s standard public training and test split\([Shenoy, 2020](https://arxiv.org/html/2609.26962#bib.bib23)\)\. Fourteen engineered numeric features cover amount, location, time, demographics, and merchant category\. The benchmark searches LightGBM and XGBoost with FLAML\. The five seeds change both coreset selection and model search, which explains part of the observed dispersion\. CriteoPrivateAds\.We process the full∼100\{\\sim\}100M\-row corpus and follow the temporal protocol intended by the dataset authors\([Sebbar et al\., 2025](https://arxiv.org/html/2609.26962#bib.bib21)\): days 1–24 form the training split and days 25–30 form the test split\. The conversion label isnb\_sales\>0\\texttt\{nb\\\_sales\}\>0, with a0\.77%0\.77\\%positive rate\. Identifier fields, all\-null fields, and high\-cardinality hash identifiers are excluded, leaving 51 features\. Every method uses the same zero\-shot FLAML LightGBM procedure and five random seeds\. The full\-data baseline is one reference run\. The CC\-BY\-SA 4\.0 dataset is used under the applicable license, and no dataset rows are included in the paper\. Methods\.Random retains all positives and samples negatives uniformly\. CCS uses ten quantile strata, uniform allocation, and a1%1\\%hard\-example cutoff\. CoreTab uses the shared datamap implementation\. The production CRISP runs use score\-proportional sampling within each stratum \(γ=1\\gamma=1\), while Sparkov and CriteoPrivateAds use uniform within\-stratum sampling \(γ=0\\gamma=0\)\. All CRISP runs useQ=10Q=10,α=β=1\\alpha=\\beta=1,wmax=20w\_\{\\max\}=20, and a proxy model trained on up to one million negatives plus all positives\. CriteoPrivateAds runs the same CCS, CoreTab, and importance\-sampling implementations used by the other benchmarks\. Metric and repetition\.Average Precision is the primary metric because each task is highly imbalanced\. The production dataset uses three seeds\. Sparkov and Criteo use five\. We report means and sample standard deviations across runs\. We do not interpret seed dispersion as a formal significance test\. Table 3\.Experiment configurations used by the three evaluations\. ### 5\.2\.Production Dataset Results Table[4](https://arxiv.org/html/2609.26962#S5.T4)focuses on rates where the methods separate\. Figure[2](https://arxiv.org/html/2609.26962#S5.F2)shows the corresponding results across all tested reduction rates\. At95%95\\%negative reduction, CRISP retains99\.7%99\.7\\%of full\-data APmonth\. Random retains94\.2%94\.2\\%, and CCS retains94\.1%94\.1\\%\. The expected CRISP training set contains about1\.701\.70M rows, or6\.8%6\.8\\%of the original data\. At80%80\\%negative reduction, Random and CCS use about5\.385\.38M rows\. CRISP therefore trains on roughly one\-third as many rows at the comparison point while retaining more AP\. Table[5](https://arxiv.org/html/2609.26962#S5.T5)summarizes retention at the two highest production reduction rates\. Table 4\.Production Dataset AP retained relative to the corresponding full\-data model\. Values are means over three runs\.Figure 2\.Production Dataset AP retained relative to full\-data AP\. Lines show means over three runs, and shaded bands show one sample standard deviation\.Two line charts compare monthly and weekly full\-data Average Precision retained on the Production Dataset across negative\-class reduction rates for Random, CCS, and CRISP\.Table 5\.Percentage of full\-data Production Dataset AP retained at high negative\-class reduction\. ### 5\.3\.Ablation Figure[3](https://arxiv.org/html/2609.26962#S5.F3)summarizes the ablation suite, which used a separate set of jobs from Table[4](https://arxiv.org/html/2609.26962#S5.T4), so its reference values should not be compared as if they were the same runs\. At95%95\\%negative reduction, theQ=10Q=10setting retains99\.0%99\.0\\%of full\-data monthly AP, whileQ=1Q=1retains95\.2%95\.2\\%\. Constant scores retain94\.6%94\.6\\%\. Removing IPW retains97\.4%97\.4\\%, compared with99\.2%99\.2\\%for the matched clipped\-IPW setting\. Gradient\-only scoring matches the combined score in this suite\. Hessian\-only scoring is worse becausep\(1−p\)p\(1\-p\)falls when a negative receives a fraud score near one\. Proxy\-model depth and tree count change retained AP by at most1\.11\.1percentage points in the collected runs\. Figure 3\.Production Dataset ablation suite at90%90\\%and95%95\\%negative\-class reduction\. Error bars span the observed minimum and maximum because some settings have only one or two completed seeds\.Four grouped bar charts show full\-data monthly Average Precision retained on the Production Dataset when varying the number of strata, score weights, inverse\-propensity clipping, and proxy\-model capacity\. ### 5\.4\.Sparkov Sparkov is less conclusive than the production dataset\. CRISP has the highest mean at99\.2%99\.2\\%and99\.4%99\.4\\%negative reduction, but it loses to Random and CCS at99%99\\%\. Its99%99\\%standard deviation is also the largest in the table because one run falls to0\.66470\.6647\. Random has the highest mean at95%95\\%because one run reaches0\.96290\.9629\. Appendix[D](https://arxiv.org/html/2609.26962#A4)reports medians and interquartile ranges so that these outliers remain visible without determining the entire interpretation\. Table 6\.Sparkov AP mean and sample standard deviation over five runs\. Full\-data mean AP is0\.86290\.8629\. Bold marks the best mean\. ### 5\.5\.CriteoPrivateAds The Criteo training split contains about8585M rows with a0\.77%0\.77\\%positive rate\. CRISP has the highest mean in each column of Table[7](https://arxiv.org/html/2609.26962#S5.T7)\. At90%90\\%negative reduction, it reaches0\.17340\.1734AP on about9\.099\.09M target\-training rows, or10\.7%10\.7\\%of the full split, and retains99\.8%99\.8\\%of the full\-data AP\. At99\.4%99\.4\\%, about1\.161\.16M rows remain, or1\.37%1\.37\\%of the full split\. CRISP retains94\.2%94\.2\\%of full\-data AP, compared with approximately90%90\\%for Random and CCS and76\.7%76\.7\\%for CoreTab\. The lead over the strongest non\-CRISP baseline grows from0\.00340\.0034AP at90%90\\%reduction to0\.00740\.0074at99\.4%99\.4\\%\. At every rate, this gap is at least eight times the larger reported run\-to\-run standard deviation, although this descriptive ratio is not a formal hypothesis test\. Table 7\.CriteoPrivateAds AP mean and sample standard deviation over five runs\. The single full\-data reference has AP0\.17370\.1737\.Figure 4\.Target\-training row tradeoffs\. Row counts are derived from reported class counts and the keep\-all\-positives rule\. They approximate target\-model training cost and exclude selection overhead\.Four panels plot relative or absolute Average Precision against approximate target\-model training rows on a log scale for the Production Dataset monthly and weekly holds, Sparkov, and CriteoPrivateAds\. ### 5\.6\.Target\-Model Transfer Figure[5](https://arxiv.org/html/2609.26962#A5.F5)reports the existing Production Dataset XGBoost runs, which provide a second target model without new training\. At95%95\\%negative reduction, CRISP retains98\.9%98\.9\\%of monthly full\-data AP and98\.2%98\.2\\%of weekly full\-data AP\. Random retains96\.5%96\.5\\%and94\.7%94\.7\\%, while CCS retains96\.2%96\.2\\%and95\.6%95\.6\\%\. This result supports transfer across two tree implementations, although it remains one dataset\. ### 5\.7\.Compute Interpretation Figure[4](https://arxiv.org/html/2609.26962#S5.F4)replots AP against derived target\-training rows\. The x\-axis is not wall\-clock time\. CRISP adds a proxy\-model fit and a distributed scoring pass before target training\. Random does not\. CCS also fits a proxy model, while CoreTab performs several prediction passes\. The figure shows how target\-model AP changes with training\-set size\. It does not establish end\-to\-end speedup\. ### 5\.8\.Cross\-Dataset Analysis The clearest pattern is that allocation matters more as the negative budget shrinks\. On the production dataset, Random and CCS remain close to the full\-data result through moderate reduction, then decline after60%60\\%\. CRISP retains at least99\.7%99\.7\\%of monthly full\-data AP through95%95\\%negative reduction\. The value above100%100\\%retention at90%90\\%should not be read as evidence that pruning improves the population optimum\. It is small relative to run\-to\-run variation and may reflect regularization, a changed class ratio, or ordinary training noise\. Criteo shows the same ordering with less seed variation\. The CRISP margin over Random grows from0\.00340\.0034AP at90%90\\%negative reduction to0\.00750\.0075at99\.4%99\.4\\%\. This pattern is consistent with the allocation argument: uniform negative sampling becomes less likely to retain rare high\-score regions as the budget contracts\. The result is useful because Criteo differs from the production dataset in domain, feature construction, and target label, while using the same coreset implementations and configuration family\. Sparkov is a useful counterexample to a simple success narrative\. Mean AP favors Random or CCS at95%95\\%and99%99\\%, while CRISP leads at99\.2%99\.2\\%and99\.4%99\.4\\%\. The medians in Appendix[D](https://arxiv.org/html/2609.26962#A4)favor CRISP more consistently, which shows how strongly a single AutoML run can affect a five\-seed mean\. At99\.4%99\.4\\%negative reduction, the expected target\-training set has only about15,24115\{,\}241rows\. Model search and coreset randomness are both visible at that scale\. The ablation narrows the source of the production gain\. Removing strata reduces coverage, while removing inverse weighting changes the effective training distribution\. Proxy\-model depth and tree count have much smaller effects in the tested range\. These observations support the design choices, but they do not prove that the same allocation is optimal on every dataset\. A fixed\-hyperparameter study and end\-to\-end timing measurements remain necessary for a cleaner causal and systems comparison\. ## 6\.Reproducibility Sparkov and CriteoPrivateAds provide public benchmark data for evaluating the method outside the proprietary setting\. CriteoPrivateAds is the primary public\-data evaluation at scale\. The paper documents the data splits, method and configuration family, reduction rates, random seeds, evaluation metrics, aggregate results, and public\-benchmark per\-run values in the appendix\. No external artifact release is pledged as part of this submission\. The proprietary dataset cannot be released due to confidentiality restrictions\. The production implementation, configurations, feature definitions, logs, run metadata, and underlying records are likewise not released\. Absolute production AP values are withheld, and all reported production results are normalized by the corresponding full\-data model\. The public datasets are cited and used under their applicable licenses; their rows are not redistributed\. Hardware details were not preserved consistently across historical runs, which is another reason the paper reports target\-training rows rather than wall\-clock time\. ## 7\.Limitations and Ethics Method limitations\.Retaining every positive works only when the positive class is small enough to leave a useful negative budget\. The production dataset uses score\-proportional within\-stratum sampling, while Sparkov and CriteoPrivateAds use uniform within\-stratum sampling\. Their shared evidence therefore concerns score strata, budget allocation, and weighting rather than one fixed within\-stratum rule\. Bernoulli sampling introduces small differences between target and realized coreset sizes, and weight clipping introduces the bias in Equation[10](https://arxiv.org/html/2609.26962#S4.E10)\. Evaluation limitations\.The row\-count tradeoff is not an end\-to\-end timing result\. It omits proxy\-model fitting, distributed scoring, data movement, and calibration\. Production target hyperparameters can vary because FLAML makes a new zero\-shot suggestion for each selected dataset\. This mirrors the existing training pipeline but weakens isolation of the selection method\. The evaluation reports AP over three or five seeds and does not include a paired test, subgroup metrics, or a fixed\-hyperparameter replication\. The strongest claims are therefore about observed AP retention, not universal dominance or deployment cost\. Ethics\.The production dataset is handled under institutional data\-governance controls\. The paper reports aggregate metrics and does not release transactions or user identifiers\. The payment reversal fraud label reflects an operational outcome and should not be interpreted as evidence of a customer’s intent\. Direct identifiers are removed from the public datasets during preprocessing\. Coreset selection can still alter subgroup representation and false\-positive rates even when aggregate AP is stable\. Before deployment, practitioners should compare subgroup recall, calibration, and manual\-review burden against the full\-data model\. They should also check whether high\-score strata overrepresent particular geographies, payment methods, or customer cohorts\. ## 8\.Conclusion CRISP addresses a specific large\-scale setting: binary tabular learning with a small positive class and a highly redundant negative class\. A proxy GBDT places negatives into score quantiles, the mean score determines each stratum’s budget, and inverse inclusion probabilities reduce the distribution shift introduced by unequal sampling\. The procedure needs only a fixed\-size proxy\-model fit and distributed passes over the training table\. It does not require pairwise gradients or predictions from many training checkpoints\. The strongest result comes from the production dataset\. At95%95\\%negative\-class reduction, CRISP retains99\.7%99\.7\\%of full\-data monthly AP while using about1\.701\.70M of2525M rows\. CriteoPrivateAds shows the same mean ordering across five reduction rates and provides evidence outside fraud detection\. Sparkov is less stable\. Its outlier runs and changing method ranking show why a high mean at one rate is not enough to claim general superiority\. The XGBoost experiment adds evidence that the selected production rows transfer across two tree implementations\. These experiments support importance\-aware allocation when the negative budget is scarce\. They do not establish a universal variance guarantee or an end\-to\-end wall\-clock speedup\. The next evaluation should use one within\-stratum rule across all datasets, fixed target hyperparameters, actual selection and training times, and enough repeated runs for paired uncertainty estimates\. Subgroup checks are also needed before the empirical package is complete\. ## References - Agiollo et al\.\(2024\)Andrea Agiollo, Young In Kim, and Rajiv Khanna\. 2024\.Approximating Memorization Using Loss Surface Geometry for Dataset Pruning and Summarization\. In*Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining*\. Association for Computing Machinery, Barcelona, Spain, 17–28\.[doi:10\.1145/3637528\.3671985](https://doi.org/10.1145/3637528.3671985) - Board of Governors of the Federal Reserve System \(2025\)Board of Governors of the Federal Reserve System\. 2025\.Commercial ACH Volume\.[https://www\.federalreserve\.gov/paymentsystems/fedach\_yearlycomm\.htm](https://www.federalreserve.gov/paymentsystems/fedach_yearlycomm.htm)\.Accessed April 2026\. - Chai et al\.\(2023\)Chengliang Chai, Jiayi Wang, Nan Tang, Ye Yuan, Jiabin Liu, Yuhao Deng, and Guoren Wang\. 2023\.Efficient Coreset Selection with Cluster\-Based Methods\. In*Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining*\. Association for Computing Machinery, Long Beach, CA, USA, 167–178\.[doi:10\.1145/3580305\.3599326](https://doi.org/10.1145/3580305.3599326) - Chen and Guestrin \(2016\)Tianqi Chen and Carlos Guestrin\. 2016\.XGBoost: A Scalable Tree Boosting System\. In*Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining \(KDD\)*\. Association for Computing Machinery, San Francisco, CA, USA, 785–794\.[doi:10\.1145/2939672\.2939785](https://doi.org/10.1145/2939672.2939785) - Coleman et al\.\(2020\)Cody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman, Peter Bailis, Percy Liang, Jure Leskovec, and Matei Zaharia\. 2020\.Selection via Proxy: Efficient Data Selection for Deep Learning\.International Conference on Learning Representations \(ICLR\)\.arXiv:1906\.11829 \[cs\.LG\] - Dal Pozzolo et al\.\(2015\)Andrea Dal Pozzolo, Olivier Caelen, Reid A\. Johnson, and Gianluca Bontempi\. 2015\.Calibrating Probability with Undersampling for Unbalanced Classification\. In*IEEE Symposium Series on Computational Intelligence \(SSCI\)*\. IEEE, Cape Town, South Africa, 159–166\.[doi:10\.1109/SSCI\.2015\.33](https://doi.org/10.1109/SSCI.2015.33) - Feldman \(2020\)Dan Feldman\. 2020\.Introduction to Core\-Sets: An Updated Survey\.*WIREs Data Mining and Knowledge Discovery*10, 1 \(2020\), e1335\.[doi:10\.1002/widm\.1335](https://doi.org/10.1002/widm.1335) - Hadar et al\.\(2024\)Aviv Hadar, Tova Milo, and Kathy Razmadze\. 2024\.Datamap\-Driven Tabular Coreset Selection for Classifier Training\.*Proceedings of the VLDB Endowment*18, 3 \(2024\), 876–888\.[doi:10\.14778/3712221\.3712249](https://doi.org/10.14778/3712221.3712249) - Hart \(1968\)Peter Hart\. 1968\.The Condensed Nearest Neighbor Rule\.*IEEE Transactions on Information Theory*14, 3 \(1968\), 515–516\.[doi:10\.1109/TIT\.1968\.1054155](https://doi.org/10.1109/TIT.1968.1054155) - Horvitz and Thompson \(1952\)Daniel G\. Horvitz and Donovan J\. Thompson\. 1952\.A Generalization of Sampling Without Replacement from a Finite Universe\.*J\. Amer\. Statist\. Assoc\.*47, 260 \(1952\), 663–685\.[doi:10\.1080/01621459\.1952\.10483446](https://doi.org/10.1080/01621459.1952.10483446) - Ke et al\.\(2017\)Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie\-Yan Liu\. 2017\.LightGBM: A Highly Efficient Gradient Boosting Decision Tree\. In*Advances in Neural Information Processing Systems \(NeurIPS\)*\. Curran Associates, Inc\., Long Beach, CA, USA, 3146–3154\. - Killamsetty et al\.\(2021b\)Krishnateja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De, and Rishabh Iyer\. 2021b\.GRAD\-MATCH: Gradient Matching Based Data Subset Selection for Efficient Deep Model Training\. In*Proceedings of the International Conference on Machine Learning \(ICML\)*\. PMLR, Online, 5464–5474\.arXiv:2103\.00123 \[cs\.LG\] - Killamsetty et al\.\(2021a\)Krishnateja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, and Rishabh Iyer\. 2021a\.GLISTER: Generalization based Data Subset Selection for Efficient and Robust Learning\.International Conference on Learning Representations \(ICLR\)\.arXiv:2012\.10630 \[cs\.LG\] - Mani and Zhang \(2003\)Inderjeet Mani and I\. Zhang\. 2003\.kNN Approach to Unbalanced Data Distributions\.ICML Workshop on Learning from Imbalanced Data Sets\. - Melo\-Acosta et al\.\(2022\)German E\. Melo\-Acosta, Freddy Duitama\-Munoz, and Julian D\. Arias\-Londono\. 2022\.An Instance Selection Algorithm for Big Data in High Imbalanced Datasets Based on LSH\.arXiv:2210\.04310\.arXiv:2210\.04310 \[cs\.LG\] - Mienye and Jere \(2024\)Ibomoiye Domor Mienye and Nobert Jere\. 2024\.Deep Learning for Credit Card Fraud Detection: A Review of Algorithms, Challenges, and Solutions\.*IEEE Access*12 \(2024\), 96893–96910\.[doi:10\.1109/ACCESS\.2024\.3426955](https://doi.org/10.1109/ACCESS.2024.3426955) - Mirzasoleiman et al\.\(2020\)Baharan Mirzasoleiman, Jeff Bilmes, and Jure Leskovec\. 2020\.Coresets for Data\-Efficient Training of Machine Learning Models\. In*Proceedings of the International Conference on Machine Learning \(ICML\)*\. PMLR, Online, 6950–6960\.arXiv:1906\.01827 \[cs\.LG\] - Paul et al\.\(2021\)Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite\. 2021\.Deep Learning on a Data Diet: Finding Important Examples Early in Training\.Advances in Neural Information Processing Systems \(NeurIPS\)\.arXiv:2107\.07075 \[cs\.LG\] - Pruthi et al\.\(2020\)Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan\. 2020\.Estimating Training Data Influence by Tracing Gradient Descent\. In*Advances in Neural Information Processing Systems \(NeurIPS\)*\. Curran Associates, Inc\., Online, 19920–19930\.arXiv:2002\.08484 \[cs\.LG\] - Sebbar et al\.\(2025\)Mehdi Sebbar, Corentin Odic, Mathieu Léchine, Aloïs Bissuel, Nicolas Chrysanthos, Anthony D’Amato, Alexandre Gilotte, Fabian Höring, Sarah Nogueira, and Maxime Vono\. 2025\.CriteoPrivateAds: A Real\-World Bidding Dataset to Design Private Advertising Systems\.arXiv:2502\.12103\.arXiv:2502\.12103 \[cs\.LG\] - Sener and Savarese \(2018\)Ozan Sener and Silvio Savarese\. 2018\.Active Learning for Convolutional Neural Networks: A Core\-Set Approach\.International Conference on Learning Representations \(ICLR\)\.arXiv:1708\.00489 \[stat\.ML\] - Shenoy \(2020\)Kartik Shenoy\. 2020\.Credit Card Transactions Fraud Detection Dataset\.Kaggle\.[https://www\.kaggle\.com/datasets/kartik2112/fraud\-detection](https://www.kaggle.com/datasets/kartik2112/fraud-detection)Accessed July 2026\. - Swayamdipta et al\.\(2020\)Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A\. Smith, and Yejin Choi\. 2020\.Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics\. In*Proceedings of the Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*\. Association for Computational Linguistics, Online, 9275–9293\.arXiv:2009\.10795 \[cs\.CL\][doi:10\.18653/v1/2020\.emnlp\-main\.746](https://doi.org/10.18653/v1/2020.emnlp-main.746) - Tomek \(1976\)Ivan Tomek\. 1976\.Two Modifications of CNN\.*IEEE Transactions on Systems, Man, and Cybernetics*6, 11 \(1976\), 769–772\.[doi:10\.1109/TSMC\.1976\.4309452](https://doi.org/10.1109/TSMC.1976.4309452) - Toneva et al\.\(2019\)Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J\. Gordon\. 2019\.An Empirical Study of Example Forgetting during Deep Neural Network Learning\.International Conference on Learning Representations \(ICLR\)\.arXiv:1812\.05159 \[cs\.LG\] - Wang et al\.\(2021\)Chi Wang, Qingyun Wu, Markus Weimer, and Erkang Zhu\. 2021\.FLAML: A Fast and Lightweight AutoML Library\. In*Proceedings of Machine Learning and Systems \(MLSys\)*, Vol\. 3\. MLSys, Online, 434–447\.arXiv:1911\.04706 \[cs\.LG\] - Xia et al\.\(2024\)Xiaobo Xia, Jiale Liu, Shaokun Zhang, Qingyun Wu, Hongxin Wei, and Tongliang Liu\. 2024\.Refined Coreset Selection: Towards Minimal Coreset Size under Model Performance Constraints\. In*Proceedings of the 41st International Conference on Machine Learning**\(Proceedings of Machine Learning Research, Vol\. 235\)*\. PMLR, Vienna, Austria, 54082–54103\.[https://proceedings\.mlr\.press/v235/xia24b\.html](https://proceedings.mlr.press/v235/xia24b.html) - Yang et al\.\(2024\)Shuo Yang, Zhe Cao, Sheng Guo, Ruiheng Zhang, Ping Luo, Shengping Zhang, and Liqiang Nie\. 2024\.Mind the Boundary: Coreset Selection via Reconstructing the Decision Boundary\. In*Proceedings of the 41st International Conference on Machine Learning**\(Proceedings of Machine Learning Research, Vol\. 235\)*\. PMLR, Vienna, Austria, 55948–55960\.[https://proceedings\.mlr\.press/v235/yang24b\.html](https://proceedings.mlr.press/v235/yang24b.html) - Zheng et al\.\(2023\)Haizhong Zheng, Rui Liu, Fan Lai, and Atul Prakash\. 2023\.Coverage\-Centric Coreset Selection for High Pruning Rates\.International Conference on Learning Representations \(ICLR\)\.arXiv:2210\.15809 \[cs\.LG\] ## Appendix AEstimator Derivation ###### Proof of Proposition[1](https://arxiv.org/html/2609.26962#S4.Thmtheorem1)\. Condition on the realized design𝒟\\mathcal\{D\}\. The positive rows are retained deterministically\. For each negative row,πi\>0\\pi\_\{i\}\>0and \(11\)𝔼\[Iiℓi\(θ\)πi\|𝒟\]=ℓi\(θ\)πi𝔼\[Ii∣𝒟\]=ℓi\(θ\)\.\\mathbb\{E\}\\left\[I\_\{i\}\\frac\{\\ell\_\{i\}\(\\theta\)\}\{\\pi\_\{i\}\}\\,\\middle\|\\,\\mathcal\{D\}\\right\]=\\frac\{\\ell\_\{i\}\(\\theta\)\}\{\\pi\_\{i\}\}\\mathbb\{E\}\[I\_\{i\}\\mid\\mathcal\{D\}\]=\\ell\_\{i\}\(\\theta\)\.Summing over the negative rows and adding the deterministic positive sum gives𝔼\[L^\(θ\)∣𝒟\]=L𝒱\(θ\)\\mathbb\{E\}\[\\widehat\{L\}\(\\theta\)\\mid\\mathcal\{D\}\]=L\_\{\\mathcal\{V\}\}\(\\theta\)\. This step uses linearity of expectation and does not require the indicators to be independent\. Unconditional unbiasedness follows from the law of total expectation\. ∎ Clipping\.Conditional on𝒟\\mathcal\{D\},𝔼\[Iiw~iℓi∣𝒟\]=min\(1,πiwmax\)ℓi\\mathbb\{E\}\[I\_\{i\}\\widetilde\{w\}\_\{i\}\\ell\_\{i\}\\mid\\mathcal\{D\}\]=\\min\(1,\\pi\_\{i\}w\_\{\\max\}\)\\ell\_\{i\}\. Ifπi≥1/wmax\\pi\_\{i\}\\geq 1/w\_\{\\max\}, the expected contribution remainsℓi\\ell\_\{i\}\. Otherwise it isπiwmaxℓi\\pi\_\{i\}w\_\{\\max\}\\ell\_\{i\}\. Subtracting these terms from the full loss gives Equation[10](https://arxiv.org/html/2609.26962#S4.E10)\. Evaluated variants\.The production dataset samples within each stratum with probability proportional to score\. Sparkov and CriteoPrivateAds use a constant Bernoulli fraction inside each stratum after the same importance\-based allocation\. ## Appendix BFull Per\-Run Experimental Results Table[8](https://arxiv.org/html/2609.26962#A2.T8)reports relative per\-run Production Dataset results at the two highest reduction rates\. Per\-run Sparkov and CriteoPrivateAds tables follow\. Table 8\.Per\-run Production Dataset AP retained as a percentage of the corresponding full\-data AP\. ## Appendix CFull Per\-Run Sparkov Results Table[9](https://arxiv.org/html/2609.26962#A3.T9)reports per\-run AP on the Sparkov dataset\. Full\-data baselines are0\.85100\.8510,0\.85940\.8594,0\.87360\.8736,0\.85630\.8563, and0\.87420\.8742for seeds 1 through 5\. Table 9\.Sparkov per\-run AP across negative\-class reduction rates\. ## Appendix DSparkov Robust Summaries Table[10](https://arxiv.org/html/2609.26962#A4.T10)reports medians and interquartile ranges across the same five runs as Table[6](https://arxiv.org/html/2609.26962#S5.T6)\. The mean and median rankings differ because Random, CoreTab, and CRISP each contain at least one large outlier\. Table 10\.Sparkov AP median with interquartile range in brackets\. ## Appendix EProduction Dataset XGBoost Transfer Results Figure 5\.Production Dataset XGBoost target results over three runs, reported relative to full\-data AP\.Two line charts compare monthly and weekly full\-data Average Precision retained on the Production Dataset for XGBoost targets across negative\-class reduction rates\. ## Appendix FFull Per\-Run CriteoPrivateAds Results Table[11](https://arxiv.org/html/2609.26962#A6.T11)reports per\-run AP on CriteoPrivateAds for seeds 42–46\. The full\-data baseline \(AP=0\.1737=0\.1737\) is a single reference run trained on all∼85\{\\sim\}85M training\-split examples\. CoreTab is reported only in aggregate in Table[7](https://arxiv.org/html/2609.26962#S5.T7)because its per\-seed values were not retained in the exported summary\. Table 11\.CriteoPrivateAds per\-run AP across negative\-class reduction rates\.
Similar Articles
CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
CRISP introduces a method for improving sparse prefilling in long-context LLM inference by using structural routing to address noise accumulation, achieving speedups and performance gains on retrieval tasks.
K-IPO: Kendall-constrained Importance Preserving Oversampling for Imbalanced Tabular Data
Introduces K-IPO, a generate-then-select oversampling framework that preserves the original data's feature importance ranking (measured by Kendall's tau) during augmentation for imbalanced tabular data, showing improved preservation, explanation consistency, and predictive performance across 20 datasets.
Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks
This paper proposes a submodular coreset selection method for LLM benchmarks that selects a subset of prompts without using model evaluation outcomes, achieving score preservation across 35 benchmarks and 18 LLMs.
When, Where, and How: Adaptive Binning for Tabular Self-Supervised Learning
This paper proposes Adaptive Binning, a learning-coupled feature-wise coarse-to-fine curriculum for tabular self-supervised learning that adaptively discretizes features, improving representations on medical datasets and establishing a unified benchmark.
Ensemble of Unsupervised Deep Learning for Clustering Imbalanced Tabular Data
This paper investigates deep clustering methods on imbalanced tabular data and proposes two novel ensemble approaches that aggregate clustering assignments across embedding dimensions or via majority voting, outperforming individual methods on 16 datasets.