Can Training Logs Make Model Comparisons More Precise?

arXiv cs.LG Papers

Summary

This paper studies whether training logs from stochastic runs can improve the precision of model comparisons via arm-specific covariate adjustment, finding that simple adjustments can reduce uncertainty, though careful covariate selection is needed to avoid noise.

arXiv:2608.02705v1 Announce Type: new Abstract: Comparing stochastically trained models requires estimating both a performance difference and its uncertainty from repeated runs. We study whether training logs from those same runs can make such comparisons more precise. Because training-log covariates are produced during training rather than measured before it, we use arm-specific covariate adjustment: each model is adjusted only with statistics from its own runs, and the raw mean difference remains the reported effect. In a vision study spanning three architectures and three datasets, simple adjustments based on early training logs often reduce uncertainty in model comparisons. The main limitation is covariate selection. Broadly searching the log pool for the most correlated statistic often adds more noise than it removes, even when useful statistics exist in hindsight. Training logs therefore appear useful for more precise model comparisons, but only when the adjustment avoids large selection noise.
Original Article
View Cached Full Text

Cached at: 08/05/26, 07:41 AM

# Can Training Logs Make Model Comparisons More Precise?
Source: [https://arxiv.org/html/2608.02705](https://arxiv.org/html/2608.02705)
###### Abstract

Comparing stochastically trained models requires estimating both a performance difference and its uncertainty from repeated runs\. We study whether training logs from those same runs can make such comparisons more precise\. Because training\-log covariates are produced during training rather than measured before it, we use arm\-specific covariate adjustment: each model is adjusted only with statistics from its own runs, and the raw mean difference remains the reported effect\. In a vision study spanning three architectures and three datasets, simple adjustments based on early training logs often reduce uncertainty in model comparisons\. The main limitation is covariate selection\. Broadly searching the log pool for the most correlated statistic often adds more noise than it removes, even when useful statistics exist in hindsight\. Training logs therefore appear useful for more precise model comparisons, but only when the adjustment avoids large selection noise\.

variance reduction, covariate adjustment, hypothesis testing, model comparison, stochastic training

## 1Introduction

Every comparison between stochastically trained models rests on repeated runs\. Because each run depends on initialization, data order, and nondeterministic implementation details, the performance difference between two models is not a single number but a statistical estimate with uncertainty\. Run\-to\-run variation directly determines the standard error of this estimate and can change reported confidence intervals and model rankings\(Bouthillier et al\.,[2021](https://arxiv.org/html/2608.02705#bib.bib1); Henderson et al\.,[2018](https://arxiv.org/html/2608.02705#bib.bib10); Dehghani et al\.,[2021](https://arxiv.org/html/2608.02705#bib.bib3)\)\. The standard remedy is to run more repetitions, but this scales linearly in compute\.

If early training statistics correlate with final test accuracy across repeated runs, subtracting their centered contribution can reduce run\-level variation\. Applied separately to each model, this adjustment can tighten the confidence interval for the performance difference without changing which models are being compared\. The idea is closely related to regression adjustment \(CUPED\) in online A/B testing\(Deng et al\.,[2013](https://arxiv.org/html/2608.02705#bib.bib4)\), where pre\-experiment user metrics reduce outcome variance and thereby tighten the confidence interval for the treatment effect\.

An important difference is that A/B testing covariates are measured before treatment assignment, whereas training\-log covariates such as early validation accuracy, gradient norms, and batch\-loss summaries are co\-produced with the outcome\. Because these covariates may differ across models for scientific reasons, a pooled adjustment could remove part of the model difference itself\. We therefore use*arm\-specific*adjustment: each model is adjusted using only its own training\-log covariates, and the adjusted dispersions are combined in the standard error for the performance difference\.

We address three empirical questions: whether training logs contain useful run\-level signal about final performance, whether this signal can be selected reliably at typical run budgets, and whether arm\-specific adjustment tightens confidence intervals for model\-performance differences when the statistic used for adjustment is fixed in advance\. The experimental design is a3×33\\times 3factorial study \(ResNet\-18, ViT\-Tiny, ConvNeXt\-Tiny×\\timesCIFAR\-10, CIFAR\-100, Tiny\-ImageNet; 50 runs per cell, 450 runs total\)\. Our contributions are:

1. 1\.An arm\-specific adjustment framework for model comparison\.Each model is adjusted using only its own training\-log covariates, with cross\-fitted coefficient estimation, coverage diagnostics, and null calibration checks \(§[3](https://arxiv.org/html/2608.02705#S3), §[5\.1](https://arxiv.org/html/2608.02705#S5.SS1), §[5\.2](https://arxiv.org/html/2608.02705#S5.SS2)\)\.
2. 2\.Model\-specific covariate signal and selection risk\.A fixed early validation\-loss adjustment and PCA summaries of training\-log families can reduce single\-arm variance, in some cells by sizable margins, but which training statistic helps depends on the model and training recipe, and selecting the best one from a large candidate pool often backfires \(§[5\.4](https://arxiv.org/html/2608.02705#S5.SS4), §[5\.5](https://arxiv.org/html/2608.02705#S5.SS5)\)\.
3. 3\.Pairwise comparison evidence\.Arm\-specific adjustment narrows confidence intervals for performance differences at sufficient run budgets but widens them when runs are too few for stable estimation \(§[5\.3](https://arxiv.org/html/2608.02705#S5.SS3), §[5\.6](https://arxiv.org/html/2608.02705#S5.SS6)\)\.
4. 4\.Practical guidance\.A training\-log statistic should be chosen before it is used as the primary adjustment for a model comparison; otherwise, the adjustment should be reported as exploratory\. Reliable selection under limited run budgets remains the main open problem \(§[6](https://arxiv.org/html/2608.02705#S6)\)\.

## 2Related Work

#### Run\-to\-run variation in deep learning\.

The sensitivity of deep learning results to random initialization has been documented in reinforcement learning\(Henderson et al\.,[2018](https://arxiv.org/html/2608.02705#bib.bib10)\), GANs\(Lucic et al\.,[2018](https://arxiv.org/html/2608.02705#bib.bib16)\), NLP fine\-tuning\(Dodge et al\.,[2020](https://arxiv.org/html/2608.02705#bib.bib5); Sellam et al\.,[2022](https://arxiv.org/html/2608.02705#bib.bib21)\), and supervised vision\(Bouthillier et al\.,[2021](https://arxiv.org/html/2608.02705#bib.bib1); Picard,[2021](https://arxiv.org/html/2608.02705#bib.bib18)\)\.Dehghani et al\. \([2021](https://arxiv.org/html/2608.02705#bib.bib3)\)argued that run\-to\-run variation and single\-split evaluation create a “benchmark lottery” in which perceived rankings are fragile\. These studies motivate running more repetitions or reporting variance more carefully\. Complementing this line of work, we ask whether training logs can reduce uncertainty from a fixed run budget\.

#### Regression adjustment and variance reduction\.

Classical variance\-reduction methods subtract a correlated signal with known mean to reduce Monte Carlo variance\(Owen,[2013](https://arxiv.org/html/2608.02705#bib.bib17)\)\. Regression adjustment\(Deng et al\.,[2013](https://arxiv.org/html/2608.02705#bib.bib4)\)applies this in A/B testing using pre\-experiment user metrics, with nonlinear extensions via boosted trees\(Poyarkov et al\.,[2016](https://arxiv.org/html/2608.02705#bib.bib19)\)\. In the randomized\-experiment literature,Freedman \([2008](https://arxiv.org/html/2608.02705#bib.bib7)\)showed that regression adjustment can improve or worsen precision, can introduce finite\-sample bias, and can make usual standard errors misleading under randomization\.Lin \([2013](https://arxiv.org/html/2608.02705#bib.bib14)\)showed that these concerns are minor or fixable in large samples when the adjustment interacts treatment with centered covariates and uses robust standard errors\. Our setting differs in two respects\. First, the estimand is a model’s expected performance over repeated training runs, and pairwise comparisons combine two such arm\-specific estimators\. Second, the covariates are co\-produced rather than pre\-treatment\. Out\-of\-fold estimation mitigates direct overfitting, but the co\-produced setting lacks the pre\-treatment independence guarantee, so we check interval calibration by resampling from the 50 runs we actually trained \(Section[5\.1](https://arxiv.org/html/2608.02705#S5.SS1)\)\. Ridge and James\-Stein shrinkage could further stabilize coefficient estimates in the multivariate case; we focus on ordinary least squares \(OLS\) and principal component analysis \(PCA\) for interpretability\. Bayesian partial pooling\(Gelman et al\.,[2013](https://arxiv.org/html/2608.02705#bib.bib8)\)offers an alternative; we adopt the frequentist approach because it requires fewer distributional assumptions and gives interval estimates whose coverage can be checked empirically in our diagnostics\.

#### Learning curve prediction\.

Early\-training signals predict converged performance\(Domhan et al\.,[2015](https://arxiv.org/html/2608.02705#bib.bib6)\), a structure routinely exploited for early stopping and hyperparameter search\. We repurpose the same correlations for a different goal: rather than predicting final accuracy or deciding when to stop, we use them to reduce run\-level uncertainty after training completes\. The fixed validation\-loss covariate in our experiments \(validation loss at the first\-third epoch\) is a direct application of this learning\-curve structure for covariate adjustment\. No modification to the training pipeline is needed\.

## 3Method

### 3\.1Arm\-Specific Covariate Adjustment

Let modelm∈\{A,B\}m\\in\\\{A,B\\\}produce final test accuracyYm,sY\_\{m,s\}from runs∈\{1,…,nm\}s\\in\\\{1,\\ldots,n\_\{m\}\\\}, and letXm,s∈ℝpmX\_\{m,s\}\\in\\mathbb\{R\}^\{p\_\{m\}\}be training\-log covariates from the same run\. The model\-comparison estimand is

Δ=μA−μB,μm=𝔼​\[Ym,s\],\\Delta=\\mu\_\{A\}\-\\mu\_\{B\},\\quad\\mu\_\{m\}=\\mathbb\{E\}\[Y\_\{m,s\}\],\(1\)where the expectation is over repeated stochastic training runs for the same model and recipe\. The null hypothesis isH0:Δ=0H\_\{0\}\\colon\\Delta=0\. The reported point estimate ofΔ\\Deltais always the raw mean differenceY¯A−Y¯B\\bar\{Y\}\_\{A\}\-\\bar\{Y\}\_\{B\}; covariate adjustment changes only the estimated standard error and confidence interval width, not the point estimate itself\.

For a fixed training\-log statistic, or fixed vector of statistics, in armmm, define the adjusted run value

Ym,sadj=Ym,s−θm⊤​\(Xm,s−μX,m\),Y^\{\\mathrm\{adj\}\}\_\{m,s\}=Y\_\{m,s\}\-\\theta\_\{m\}^\{\\top\}\(X\_\{m,s\}\-\\mu\_\{X,m\}\),\(2\)whereμX,m=𝔼​\[Xm,s\]\\mu\_\{X,m\}=\\mathbb\{E\}\[X\_\{m,s\}\]\. IfμX,m\\mu\_\{X,m\}is known, the centered covariate has mean zero and𝔼​\[Ym,sadj\]=μm\\mathbb\{E\}\[Y^\{\\mathrm\{adj\}\}\_\{m,s\}\]=\\mu\_\{m\}for any fixedθm\\theta\_\{m\}\. The coefficient affects variance, not the target\. For a single covariate with the population\-optimal coefficient, the variance reduction \(VR\) is

VRm=1−Var​\(Ymadj\)Var​\(Ym\)=ρXm​Ym2\.\\mathrm\{VR\}\_\{m\}=1\-\\frac\{\\mathrm\{Var\}\(Y^\{\\mathrm\{adj\}\}\_\{m\}\)\}\{\\mathrm\{Var\}\(Y\_\{m\}\)\}=\\rho\_\{X\_\{m\}Y\_\{m\}\}^\{2\}\.\(3\)In finite samples, the realized VR is smaller thanρ2\\rho^\{2\}becauseθm\\theta\_\{m\}must be estimated; the cost scales roughly as1/ntrain1/n\_\{\\mathrm\{train\}\}\(Section[3\.2](https://arxiv.org/html/2608.02705#S3.SS2)\)\. Thus, a training statistic that explains run\-level variation can reduce the standard error ofμm\\mu\_\{m\}and, when applied independently in both arms, reduce the standard error ofΔ\\Delta\.

The adjustment must be arm\-specific\. A pooled regression of outcomes on model identity and co\-produced training statistics can remove part of the model difference, because early validation loss or gradient norms may differ between models precisely because the models train differently\. We therefore fit the adjustment separately within each model and combine the two arm\-specific standard errors only after adjustment\.

### 3\.2Covariate Selection Risk

If the covariate is chosen after screening many candidates, the apparent correlation can be optimistic and the realized VR can be negative\. Intuitively, the covariate that looks most correlated on a small training sample may not remain correlated on held\-out runs, causing the adjustment to add noise rather than remove it\. The estimation penalty scales as∼p/ntrain\\sim p/n\_\{\\mathrm\{train\}\}, whereppis the effective number of covariates\. For a single selected covariate, the nominalp=1p=1, but the implicit search over hundreds of candidates inflates the effective degrees of freedom\. This risk motivates out\-of\-fold estimation \(Section[3\.4](https://arxiv.org/html/2608.02705#S3.SS4)\) and is analyzed empirically in Section[5\.5](https://arxiv.org/html/2608.02705#S5.SS5)\.

### 3\.3Why Co\-Produced Covariates Need Care

In standard regression adjustment for A/B tests, covariates are measured before the treatment, which separates covariate measurement from treatment assignment\. This separation helps prevent a covariate adjustment from absorbing part of the treatment effect\(Rosenbaum,[1984](https://arxiv.org/html/2608.02705#bib.bib20)\)\. Training\-log covariates are different: they are produced by the same stochastic run whose final accuracy is being evaluated\.

Two design choices make the adjustment scientifically interpretable in our setting\. First, we use covariates only within the arm that produced them\. This keeps the target for each arm as the expected final accuracy of that model, rather than a performance contrast after controlling for a shared post\-treatment variable\. Second, we evaluate each proposed training\-log adjustment by its out\-of\-fold variance reduction and by coverage diagnostics based on resampling from the 50 runs we actually trained, rather than assuming that a reduced residual variance automatically yields a valid hypothesis test\.

If the training\-log statistic and population covariate mean were fixed in advance, the centering in Equation[2](https://arxiv.org/html/2608.02705#S3.E2)would preserve the arm mean\. In practice, both the coefficient and the covariate mean are estimated from the same limited run budget\. Cross\-fitting reduces direct overfitting, but the results still report both variance reduction and interval coverage to verify the adjustment empirically\.

### 3\.4Cross\-Fitted Evaluation

The population expression above assumes thatθm\\theta\_\{m\}andμX,m\\mu\_\{X,m\}are known\. In practice, both must be estimated from the same runs used for evaluation\. This creates a risk: the adjustment might reduce apparent variance by memorizing run\-level noise rather than capturing stable structure\. We useKK\-fold cross\-fitting\(Chernozhukov et al\.,[2018](https://arxiv.org/html/2608.02705#bib.bib2)\)withK=5K=5throughout\. Runs are partitioned intoKKfolds, and for runssin foldf​\(s\)f\(s\), nuisance parameters are estimated from the complement:

Y~m,s=Ym,s−θ^m,−f​\(s\)⊤​\(Xm,s−μ^X,m,−f​\(s\)\)\.\\tilde\{Y\}\_\{m,s\}=Y\_\{m,s\}\-\\hat\{\\theta\}\_\{m,\-f\(s\)\}^\{\\top\}\\bigl\(X\_\{m,s\}\-\\hat\{\\mu\}\_\{X,m,\-f\(s\)\}\\bigr\)\.\(4\)For single\-arm diagnostics, we compare the sample variance of the cross\-fitted adjusted outcomes to the raw outcome variance\. For pairwise comparisons, we recenter adjusted outcomes within each arm to keep the reported performance difference equal to the raw difference, then compute a Welch\-style interval using the adjusted within\-arm variances\. This conservative reporting choice separates the model ranking from the variance estimate: covariates can change the reported uncertainty, but not the reported mean performance difference\.

Because adjusted outcomes within a fold share fitted coefficients, and because we recenter to preserve the raw point estimate, Section[5\.1](https://arxiv.org/html/2608.02705#S5.SS1)checks adjusted\-interval coverage by repeatedly resampling from the 50 runs we actually trained\.

Cross\-fitting prevents the coefficient from directly memorizing held\-out runs\. However, if the covariate is*selected*from a large pool using the same runs, the selection step must also be kept out of fold to avoid choosing a statistic that correlates with noise in the evaluation runs\. As we show in Section[5\.5](https://arxiv.org/html/2608.02705#S5.SS5), this selection problem is the primary practical bottleneck\.

### 3\.5Pairwise Intervals

For two models A and B, arm\-specific adjustment gives estimated within\-arm standard deviationssadj,As\_\{\\mathrm\{adj\},A\}andsadj,Bs\_\{\\mathrm\{adj\},B\}\. We use the raw mean difference as the point estimate and compute the adjusted standard error

SE^adj​\(Δ^\)=sadj,A2nA\+sadj,B2nB\.\\widehat\{\\mathrm\{SE\}\}\_\{\\mathrm\{adj\}\}\(\\hat\{\\Delta\}\)=\\sqrt\{\\frac\{s\_\{\\mathrm\{adj\},A\}^\{2\}\}\{n\_\{A\}\}\+\\frac\{s\_\{\\mathrm\{adj\},B\}^\{2\}\}\{n\_\{B\}\}\}\.\(5\)The corresponding confidence interval uses the Welch degrees of freedom\. When both adjusted variances are smaller than their raw counterparts, the interval for the performance difference tightens\. If one arm has a poor covariate, the interval can widen\. The minimum detectable effect size \(MDE\) at power1−β1\-\\betaand significanceα\\alphais

MDE≈\(z1−α/2\+z1−β\)⋅sadj,A2\+sadj,B2n,\\mathrm\{MDE\}\\approx\(z\_\{1\-\\alpha/2\}\+z\_\{1\-\\beta\}\)\\cdot\\sqrt\{\\frac\{s\_\{\\mathrm\{adj\},A\}^\{2\}\+s\_\{\\mathrm\{adj\},B\}^\{2\}\}\{n\}\},\(6\)for equal run counts \(the normal quantiles are asymptotic; our intervals use the Welchttapproximation\)\. Thus, interval tightening also corresponds to improved test sensitivity when uncertainty is the limiting factor\.

We now describe the experimental design used to test these claims empirically\.

## 4Experimental Setup

### 4\.1Design

We conduct a3×33\\times 3factorial experiment: three architectures×\\timesthree datasets, with 50 runs per cell \(450 runs total\)\. All experiments run on NVIDIA L4 GPUs with bfloat16 mixed precision\. Sources of randomness \(Python, NumPy, PyTorch, CUDA, DataLoader workers\) are seeded deterministically\. The primary outcome isfinal\-epoch test accuracy, fixed before the covariate analysis\. Using the best\-validation\-accuracy checkpoint would make validation metrics part of the outcome selection, preventing their use as covariates\. With a fixed checkpoint, validation logs can serve as covariates without leakage\. Across the nine cells, baseline run standard deviations range from 0\.13 pp \(ConvNeXt\-Tiny / CIFAR\-10\) to 0\.69 pp \(ConvNeXt\-Tiny / Tiny\-ImageNet\); full results are in Appendix[A](https://arxiv.org/html/2608.02705#A1)\.

### 4\.2Architectures and Recipes

ResNet\-18\(He et al\.,[2016](https://arxiv.org/html/2608.02705#bib.bib9)\)with a stem adapted for32×3232\\\!\\times\\\!32inputs \(stride\-13×33\\\!\\times\\\!3conv, no max\-pool\), trained with SGD \(lr 0\.1, cosine schedule, 100 to 200 epochs depending on dataset\)\.ViT\-Tiny\(Touvron et al\.,[2021](https://arxiv.org/html/2608.02705#bib.bib22)\)\(embed dim 192, 12 layers, 3 heads, patch size 4 on CIFAR, 8 on Tiny\-ImageNet\) andConvNeXt\-Tiny\(Liu et al\.,[2022](https://arxiv.org/html/2608.02705#bib.bib15)\)\(resolution\-adapted stems\), both trained with AdamW \(cosine schedule, RandAugment, Mixup, CutMix, 300 epochs\)\. ViT\-Tiny applies gradient clipping at 1\.0; ConvNeXt\-Tiny omits clipping and EMA to preserve run\-specific variance signals\.

### 4\.3Datasets

CIFAR\-10 and CIFAR\-100\(Krizhevsky,[2009](https://arxiv.org/html/2608.02705#bib.bib12)\)\(32×3232\\\!\\times\\\!32, 50k train / 10k test\) and Tiny\-ImageNet\(Le & Yang,[2015](https://arxiv.org/html/2608.02705#bib.bib13)\)\(64×6464\\\!\\times\\\!64, 100k train / 10k test, 200 classes\)\. From each training set, 5,000 images are held out as a fixed validation split \(split seed 2026\)\.

### 4\.4Covariates

Every epoch logs train/val/test loss and accuracy and gradient norm statistics\. The first 1,000 training steps log per\-batch loss, accuracy, gradient norm, and parameter norm\. At completion,∼\\sim200 summary statistics are computed per run \(prefix means, standard deviations, slopes, etc\.\)\. We restrict covariates to the first third of training and exclude test\-derived statistics\. Candidates are organized into five families:

- •Validation snapshots: val\_acc and val\_loss at early epochs \(roughly the first third of training\)\.
- •Training\-loss summaries: epoch losses, per\-batch loss summaries, and first\-step loss statistics\.
- •Training\-accuracy summaries: epoch training accuracy, per\-batch accuracy summaries, and first\-step accuracy statistics\.
- •Gradient\-norm summaries: per\-epoch and per\-step gradient norms and their derived statistics\.
- •Parameter\-norm summaries: per\-epoch and per\-step parameter norms, learning rate, and their derived statistics\.

The total candidate count ranges from 672 \(ResNet\-18 / CIFAR\-10\) to 1,719 \(ViT\-Tiny / CIFAR\-10\)\. Two training details affect covariate interpretation: for ViT\-Tiny, gradient norms are logged post\-clipping \(capped at≤1\.0\\leq 1\.0\), limiting their informativeness; and for ViT\-Tiny and ConvNeXt\-Tiny, training accuracy is computed against hard labels while training uses mixed soft targets from Mixup/CutMix\.

### 4\.5Choosing Training\-Log Covariates

The primary pairwise experiments adjust each arm using validation loss at the first\-third epoch\. This statistic is chosen before comparing the empirical covariance structure across the candidate pool and is motivated by prior work on learning\-curve prediction\(Domhan et al\.,[2015](https://arxiv.org/html/2608.02705#bib.bib6)\)\. It provides a prospective baseline for testing whether training logs can reduce uncertainty without searching over hundreds of candidates\.

We also evaluate two data\-driven alternatives to diagnose whether stronger adjustments can be obtained from the log pool\.Single\-best OLSasks whether directly searching the logs can find a useful statistic: within each training fold, it chooses the candidate with the largestρ2\\rho^\{2\}with final accuracy, then applies the fitted coefficient to held\-out runs\. This procedure is intuitive, but it uses the outcome during selection and can overfit when the candidate pool is large\.PCA \(k=1k=1\)takes a more restricted route: it replaces each family of training\-log statistics with its first principal component and fits OLS on those components\. PCA does not use final accuracy when constructing the components within a fold, while single\-best OLS does use final accuracy to choose among candidates\. Comparing the two helps separate useful signal in the logs from the noise introduced by searching over many candidate covariates\.

## 5Results

We report results for all nine configurations in the3×33\\times 3design\. Early\-training covariates often reduce variance, but choosing among many training\-log statistics can introduce more estimation error than it removes\. We first check interval calibration and null behavior, then evaluate pairwise model comparisons, and finally use single\-arm variance reduction, covariate\-selection diagnostics, and run\-budget analysis to explain when adjustment helps or hurts\.

### 5\.1Interval Calibration

Before interpreting variance reduction or interval width, we assess whether adjusted intervals become misleadingly narrow\. In standard A/B testing, covariates are measured before treatment, which separates covariate measurement from treatment assignment\. Training\-log covariates are co\-produced with the outcome, so this separation does not hold\. If adjustment distorts coverage, tighter intervals would reflect false confidence rather than improved precision\.

We test this with a subsampling diagnostic\. For each model\-dataset cell and run budgetn∈\{10,15,20,30,50\}n\\in\\\{10,15,20,30,50\\\}, we repeatedly draw a random subset ofnnruns from the full pool of 50 \(without replacement within each draw\), build a cross\-fitted 95% confidence interval from that subset, and check whether it contains the 50\-run mean\. We repeat this 5,000 times per setting\. Under\-coverage in this finite\-run diagnostic would indicate that adjustment is making intervals too narrow; coverage at or above 95% is a minimal calibration check\.

Table 1:Interval coverage \(%\) across run budgets for CIFAR\-100\.Raw = unadjusted; Pre = validation loss at first\-third epoch\.

Table[1](https://arxiv.org/html/2608.02705#S5.T1)shows results for CIFAR\-100; the same diagnostic on the other datasets also shows no under\-coverage\. All entries are at or above 95%, suggesting that adjustment does not make intervals anti\-conservative when we resample from the 50 runs we actually trained\. Coverage exceeds 95% at smallnnbecause thetn−1t\_\{n\-1\}quantile is conservative with few observations\. Then=50n=50row is a deterministic limit case: drawing all 50 runs yields one subset, which covers its own mean by construction\. We include it as an implementation check\. Because the 5,000 subsamples are drawn from the same 50 runs, the coverage indicators are positively correlated; the decimal precision should not be over\-interpreted\.

### 5\.2Null Calibration

The coverage diagnostic above checks whether adjusted intervals cover the mean of the 50 runs we actually trained within each model\-dataset cell, but it does not test behavior under a zero model difference\. To probe Type I behavior under a known\-zero contrast within the observed run pool, we construct a synthetic null: for each cell, we randomly split the 50 runs from the*same*model into two fake arms of 25\. The two arms are exchangeable under this construction, so the null hypothesis of equal expected performance holds even though any particular split can have a nonzero sample mean difference\. We then compute the adjusted Welch interval and check whether it covers zero\. Across 5,000 random splits per cell, adjusted coverage averages 94\.8% \(range 92\.9% to 95\.9%\), close to the nominal 95%\. The raw interval averages 95\.1%\. Adjusted rejection rates average 5\.2%, compared to 4\.9% raw\. One cell \(CIFAR\-10 / ViT\-Tiny\) shows elevated rejection at 7\.1%\. Because the random splits share runs and we examine nine cells, we treat this elevation as a diagnostic warning rather than a formal significance claim\. The pattern suggests that arm\-specific validation\-loss adjustment does not systematically inflate rejection rates, but that null calibration should be reported rather than assumed, especially for cells with strong covariate signal\.

### 5\.3Pairwise Model Comparisons

After the calibration checks, we evaluate the main inferential target: confidence intervals for model\-performance differences\. We use the fixed validation\-loss adjustment from Section[4](https://arxiv.org/html/2608.02705#S4)and fit it independently per arm with cross\-fitting \(2,000 subsamples per run budget\)\. Table[2](https://arxiv.org/html/2608.02705#S5.T2)reports the percent change in the 95% Welch interval half\-width; positive values mean the adjusted interval is narrower\.

Table 2:Pairwise CI diagnostics from arm\-specific validation\-loss adjustment\.Δ\\DeltaCI = percent change in half\-width \(positive = narrower\)\. HW = absolute half\-width in pp atn=50n=50\. Cov = minimum adjusted coverage overn∈\{10,15,20,30,50\}n\\in\\\{10,15,20,30,50\\\}when resampling from the trained runs \(%\)\.

The pairwise results show conditional gains\. Atn=50n=50, adjusted intervals narrow in eight of nine model pairs, with reductions from 1\.9% to 8\.8% among the successful pairs\. The only exception is CIFAR\-10 / ResNet\-18 vs\. ConvNeXt\-Tiny \(−\-1\.4%\), where the later single\-arm analysis shows that validation\-loss adjustment is not useful in either arm\. The absolute magnitudes are modest: the largest improvement is CIFAR\-100 / ViT\-Tiny vs\. ConvNeXt\-Tiny, where the half\-width shrinks from 0\.158 pp to 0\.144 pp\. These gains require a sufficient run budget; at very small budgets \(e\.g\.,n=10n=10\), the cost of estimating the covariate coefficient systematically exceeds the variance removed, and the adjusted interval widens across all pairs\. Adjustment pays off only when the run budget is large enough for stable coefficient estimation\. When resampling from the trained runs, adjusted pairwise coverage remains≥\\geq97\.2% across all pairs and run budgets, consistent with no clear anti\-conservatism in this diagnostic\.

Since MDE scales with the same standard error \(Equation[6](https://arxiv.org/html/2608.02705#S3.E6)\), these interval reductions correspond to higher sensitivity in comparisons where uncertainty is practically relevant\.

### 5\.4Single\-Arm Variance Reduction

To explain why the pairwise gains are conditional, we measure single\-arm variance reduction \(VR\) for the three adjustments defined in Section[4](https://arxiv.org/html/2608.02705#S4): validation loss at the first\-third epoch,single\-best OLS, andPCA \(kk=1\)\. We defineVR=1−Var​\(Y~\)/Var​\(Y\)\\mathrm\{VR\}=1\-\\mathrm\{Var\}\(\\tilde\{Y\}\)/\\mathrm\{Var\}\(Y\), where variances are computed over cross\-fitted adjusted and raw outcomes using repeated 5\-fold cross\-fitting \(1,000 repeats\)\. Negative VR means the adjustment*increases*variance: the estimation cost of fittingθ\\thetaexceeds the variance removed\.

Table 3:Variance reduction \(%\) by adjustment\.Parentheses in OLS and PCA columns: percentage of repeated cross\-fitting splits with VR<<0\. Bold = best VR per cell\. Pre = validation loss at first\-third epoch; OLS = single\-best out\-of\-fold selector; PCA = first PC per family\.

#### Fixed validation\-loss adjustment\.

Validation loss at the first\-third epoch gives positive VR for ViT\-Tiny on all three datasets \(7\.0 to 17\.8%\) and for ConvNeXt\-Tiny on two of three, but is uniformly negative for ResNet\-18\. Because no selection is involved, these results isolate covariate quality from selection noise: the fixed validation\-loss adjustment helps some architecture and training\-recipe combinations but not others\.

#### Data\-driven selection\.

The single\-best OLS selector produces negative VR in six of nine cells\. The failures are not small: ViT\-Tiny on CIFAR\-10 hasρ2=23%\\rho^\{2\}=23\\%for its best early validation covariate, yet the selector yields−38\.1%\-38\.1\\%VR\. This pattern identifies covariate selection as the main empirical failure mode in our study\.

#### PCA\.

PCA \(kk=1\) avoids outcome\-based selection and recovers positive VR for ViT\-Tiny on all datasets and for Tiny\-ImageNet / ConvNeXt\-Tiny, outperforming both alternatives in three of nine cells\. Increasing tok=2k=2ork=3k=3PCs per family degrades performance, consistent with overfitting when the number of fitted coefficients grows relative to the run budget\. PCA’s stability comes from the fact that the first principal component is a deterministic function of the covariate matrix within each fold, avoiding outcome\-based discrete selection entirely\.

#### Post hoc evidence of signal\.

To assess whether useful covariates exist even when selection fails, we run exploratory scans using all 50 runs for selection \(but cross\-fitting the coefficient\)\. Every cell has at least one first\-third covariate with positive VR \(11\.7% to 44\.3%\)\. The examples are heterogeneous: ViT\-Tiny’s best covariates are early validation snapshots \(ρ2\\rho^\{2\}23 to 30%\), whereas ResNet\-18’s are gradient and batch\-loss summaries \(ρ2≤18\.8%\\rho^\{2\}\\leq 18\.8\\%\)\. These discovery results suggest that signal exists but is difficult to extract reliably at current run budgets\. Figure[1](https://arxiv.org/html/2608.02705#S5.F1)illustrates the contrast using first\-third covariates: ViT\-Tiny shows stronger linear covariate\-outcome relationships, while ResNet\-18 shows weaker, noisier correlations\.

![Refer to caption](https://arxiv.org/html/2608.02705v1/scatter_covariates.png)

Figure 1:Covariate\-outcome scatter for four cells\. Each point is one run\.ViT\-Tiny \(bottom\) shows stronger first\-third relationships \(ρ2=23\.0%\\rho^\{2\}=23\.0\\%and30\.3%30\.3\\%\); ResNet\-18 \(top\) shows weaker relationships \(ρ2=13\.8%\\rho^\{2\}=13\.8\\%and14\.6%14\.6\\%\)\.

### 5\.5Why Data\-Driven Selection Fails

The discrepancy between post hoc signal and nested\-selection performance suggests that the main difficulty is not signal absence, but finite\-sample selection noise\. We examine this mechanism with a family\-level diagnostic\.

#### Diagnostic setup\.

We split candidates into five families \(validation, training\-loss, training\-accuracy, gradient\-norm, parameter\-norm\) and compare two adjustment procedures within each:

- •*Post hoc best*: use all 50 runs to identify the highest\-R2R^\{2\}covariate, then cross\-fit only the coefficient\. This is biased upward because the selection step uses evaluation runs\.
- •*Nested selector*: use only the∼\\sim40 training\-fold runs for both selection and coefficient estimation\.

#### Results\.

The post hoc best is positive in 44 of 45 family\-cell combinations; the nested selector is negative in 39 of 45 \(Figure[2](https://arxiv.org/html/2608.02705#S5.F2)\)\. Even within a single family, the candidate count is large relative to the training\-fold size, so the nested selector overfits to fold\-specific noise\.

![Refer to caption](https://arxiv.org/html/2608.02705v1/x1.png)

Figure 2:Family\-level selection diagnostic\.Post hoc best covariates \(using all 50 runs for selection\) are usually positive; nested within\-fold selection shifts VR below zero\. The gap is selection noise\.

#### Interpretation\.

The estimation penalty scales as∼p/ntrain\\sim p/n\_\{\\mathrm\{train\}\}, whereppis the number of covariates\. For a single selected covariate, the nominalp=1p=1, but searching over hundreds of candidates inflates the effective degrees of freedom\. The ConvNeXt\-Tiny cells provide independent corroboration: the nested selector gives positive VR in all three ConvNeXt\-Tiny configurations, and the selected covariates are interpretable early\-training summaries\. Two regimes emerge: when signal is distributed broadly across a family, PCA captures the dominant relationship without selection; when signal is concentrated in a few covariates \(as in ConvNeXt\-Tiny\), single\-best OLS identifies them directly\.

### 5\.6Variance Reduction vs\. Run Budget

At smaller run budgets, fixed validation\-loss adjustment degrades more gracefully than data\-driven selection, which is not reliably positive belown=50n=50in these experiments \(Figure[3](https://arxiv.org/html/2608.02705#S5.F3)\)\. For ViT\-Tiny, validation loss at the first\-third epoch gives positive median VR fromn=15n=15onward \(\+2\.8% atn=15n=15, \+14\.1% atn=50n=50on CIFAR\-10\)\. Atn<15n<15, the cost of estimatingθ\\thetafrom too few runs exceeds the variance removed\. This threshold is consistent with the heuristic that realized VR is approximatelyρ2−1/ntrain\\rho^\{2\}\-1/n\_\{\\mathrm\{train\}\}: a covariate needsρ2\>1/ntrain\\rho^\{2\}\>1/n\_\{\\mathrm\{train\}\}to pay for itself\.

![Refer to caption](https://arxiv.org/html/2608.02705v1/x2.png)

Figure 3:Median VR vs\. run budget\.Solid: fixed validation\-loss adjustment; dashed: single\-best selector\. The fixed adjustment degrades more smoothly; automatic selection remains brittle at all tested budgets\.

## 6Discussion

The main lesson is modest but useful: training logs can make model comparisons more precise when the adjustment is chosen before the comparison or avoids outcome\-based selection\. At the largest run budget we study, the fixed early validation\-loss adjustment narrows eight of nine pairwise confidence intervals\. The largest single\-arm gains come from ConvNeXt\-Tiny, where single\-best OLS reaches 34\.1% variance reduction and PCA reaches 34\.0%\. The 34\.1% reduction is equivalent to 75\.9 effective runs from a 50\-run budget\. This effective\-run calculation is a diagnostic scale, not a recommendation to train that many repeats in routine comparisons\.

The failure cases matter just as much\. The log pool contains useful statistics, as shown by the post hoc family scans, but choosing the most correlated statistic within each training fold usually adds more noise than it removes\. In the family diagnostic, the post hoc best statistic is positive in 44 of 45 family\-cell combinations, while nested within\-fold selection is negative in 39 of 45\. In these experiments, the bottleneck is estimating and selecting the adjustment from limited runs, not simply a lack of signal in the logs\.

The useful signal also varies across models\. ConvNeXt\-Tiny gives the largest gains, ViT\-Tiny benefits from the fixed validation\-loss adjustment, and ResNet\-18 shows little benefit under the adjustments tested here\. A training\-log statistic that helps one model should not be assumed to help another\. Pairwise comparisons should therefore adjust each arm separately and show the raw interval alongside the adjusted interval\.

Reports of adjusted comparisons should include the raw mean difference and confidence interval, the adjusted interval, the exact training\-log statistic or summary used in each arm, and the out\-of\-fold variance reduction, including how often it is negative across cross\-fitting splits\. If the statistic was chosen from a large candidate pool after looking at outcome correlations, the adjusted interval is better treated as exploratory\.

## 7Limitations

The adjusted results have an important inferential limit\. Training\-log covariates are recorded during the same stochastic runs that produce the final accuracies\. As a result, the adjusted intervals should be read as empirically checked precision estimates, not as distribution\-free finite\-sample guarantees\. A stronger guarantee would require the adjustment to be fixed in advance and the covariate mean to be known or estimated independently\. For this reason, we keep the raw interval as the baseline, report adjusted intervals separately, and check the adjusted intervals by resampling from the 50 runs we actually trained and by using the synthetic\-null diagnostic\. The synthetic null is broadly consistent with nominal rejection but has one elevated cell\.

The experiments use CIFAR\-10, CIFAR\-100, and Tiny\-ImageNet\. Larger datasets, larger models, and different run budgets may change both the available covariate signal and the cost of estimating the adjustment\. Because architecture and training recipe vary together, we treat the differences across ResNet\-18, ViT\-Tiny, and ConvNeXt\-Tiny as configuration\-level evidence rather than as an isolated architecture effect\. The adjustment model is also linear\. Nonlinear adjustment might recover additional signal, but it would increase the risk of overfitting at the run budgets studied here\.

Finally, the outcome is final\-epoch test accuracy\. Validation metrics can be used as training\-log covariates in this setting because validation accuracy does not choose the reported checkpoint\. If validation performance determines checkpoint selection or early stopping, validation\-derived covariates would introduce leakage and should not be used for the adjustment\. Sequential evaluation is another open direction: the fixed\-sample interval studied here could be embedded within an always\-valid confidence sequence\(Johari et al\.,[2022](https://arxiv.org/html/2608.02705#bib.bib11)\), where variance reduction from covariate adjustment would translate to faster stopping once the adjusted interval reaches a target width\.

## 8Conclusion

Training logs are usually treated as optimization diagnostics\. This paper asks whether they can also make statistical comparisons between stochastically trained models more precise\. Arm\-specific covariate adjustment leaves the raw mean difference unchanged, but can reduce the standard error of that difference when each arm has a stable training\-log statistic\. In our 450\-run study, a fixed early validation\-loss adjustment narrows most pairwise confidence intervals at the largest run budget we study\. PCA summaries of training\-log families also show that some log families contain additional single\-arm variance signal\.

The results should not be read as recommending dozens of repeated runs for routine model comparisons\. They show a different point: training logs can reduce uncertainty once enough repeated runs are available to estimate the adjustment stably\. Selecting the most correlated statistic from a large log pool usually worsens variance after cross\-fitting, even when useful statistics exist in hindsight\. Training logs can help make model comparisons more precise, but the practical bottleneck is reliable covariate selection under limited run budgets\.

## Impact Statement

This paper presents work whose goal is to advance the methodology of evaluating machine learning systems\. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here\.

## References

- Bouthillier et al\. \(2021\)Bouthillier, X\., Delaunay, P\., Bronzi, M\., Trofimov, A\., et al\.Accounting for variance in machine learning benchmarks\.In*Proceedings of Machine Learning and Systems*, volume 3, pp\. 747–769, 2021\.
- Chernozhukov et al\. \(2018\)Chernozhukov, V\., Chetverikov, D\., Demirer, M\., Duflo, E\., Hansen, C\., Newey, W\., and Robins, J\.Double/debiased machine learning for treatment and structural parameters\.*The Econometrics Journal*, 21\(1\):C1–C68, 2018\.
- Dehghani et al\. \(2021\)Dehghani, M\., Tay, Y\., Gritsenko, A\. A\., Zhao, Z\., Houlsby, N\., Diaz, F\., Metzler, D\., and Vinyals, O\.The Benchmark Lottery\.*arXiv preprint arXiv:2107\.07002*, 2021\.
- Deng et al\. \(2013\)Deng, A\., Xu, Y\., Kohavi, R\., and Walker, T\.Improving the sensitivity of online controlled experiments by utilizing pre\-experiment data\.In*Proceedings of the Sixth ACM International Conference on Web Search and Data Mining*, pp\. 123–132, 2013\.
- Dodge et al\. \(2020\)Dodge, J\., Ilharco, G\., Schwartz, R\., Farhadi, A\., Hajishirzi, H\., and Smith, N\.Fine\-tuning pretrained language models: Weight initializations, data orders, and early stopping\.*arXiv preprint arXiv:2002\.06305*, 2020\.
- Domhan et al\. \(2015\)Domhan, T\., Springenberg, J\. T\., and Hutter, F\.Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves\.In*Proceedings of the Twenty\-Fourth International Joint Conference on Artificial Intelligence*, pp\. 3460–3468, 2015\.
- Freedman \(2008\)Freedman, D\. A\.On regression adjustments to experimental data\.*Advances in Applied Mathematics*, 40\(2\):180–193, 2008\.
- Gelman et al\. \(2013\)Gelman, A\., Carlin, J\. B\., Stern, H\. S\., Dunson, D\. B\., Vehtari, A\., and Rubin, D\. B\.*Bayesian Data Analysis*\.Chapman and Hall/CRC, 3rd edition, 2013\.
- He et al\. \(2016\)He, K\., Zhang, X\., Ren, S\., and Sun, J\.Deep residual learning for image recognition\.In*Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition*, pp\. 770–778, 2016\.
- Henderson et al\. \(2018\)Henderson, P\., Islam, R\., Bachman, P\., Pineau, J\., Precup, D\., and Meger, D\.Deep reinforcement learning that matters\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 32, 2018\.
- Johari et al\. \(2022\)Johari, R\., Koomen, P\., Pekelis, L\., and Walsh, D\.Always valid inference: Continuous monitoring of a/b tests\.*Operations Research*, 70\(3\):1806–1821, 2022\.
- Krizhevsky \(2009\)Krizhevsky, A\.Learning multiple layers of features from tiny images\.*Technical report, University of Toronto*, 2009\.
- Le & Yang \(2015\)Le, Y\. and Yang, X\.Tiny imagenet visual recognition challenge\.*CS 231N*, 7\(7\):3, 2015\.
- Lin \(2013\)Lin, W\.Agnostic notes on regression adjustments to experimental data: Reexamining Freedman’s critique\.*The Annals of Applied Statistics*, 7\(1\):295–318, 2013\.
- Liu et al\. \(2022\)Liu, Z\., Mao, H\., Wu, C\.\-Y\., Feichtenhofer, C\., Darrell, T\., and Xie, S\.A ConvNet for the 2020s\.In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, pp\. 11976–11986, 2022\.
- Lucic et al\. \(2018\)Lucic, M\., Kurach, K\., Michalski, M\., Gelly, S\., and Bousquet, O\.Are GANs created equal? A large\-scale study\.In*Advances in Neural Information Processing Systems*, volume 31, 2018\.
- Owen \(2013\)Owen, A\. B\.*Monte Carlo theory, methods and examples*\.Self\-published, 2013\.
- Picard \(2021\)Picard, D\.Torch\.manual\_seed\(3407\) is all you need: On the influence of random seeds in deep learning architectures for computer vision\.*arXiv preprint arXiv:2109\.08203*, 2021\.
- Poyarkov et al\. \(2016\)Poyarkov, A\., Drutsa, A\., Khalyavin, A\., Gusev, G\., and Serdyukov, P\.Boosted decision tree regression adjustment for variance reduction in online controlled experiments\.In*Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining*, pp\. 235–244, 2016\.
- Rosenbaum \(1984\)Rosenbaum, P\. R\.The consequences of adjustment for a concomitant variable that has been affected by the treatment\.*Journal of the Royal Statistical Society: Series A*, 147\(5\):656–666, 1984\.
- Sellam et al\. \(2022\)Sellam, T\., Yadlowsky, S\., Tenney, I\., Wei, J\., Saphra, N\., D’Amour, A\., Linzen, T\., Bastings, J\., Turc, I\. R\., Eisenstein, J\., Das, D\., and Pavlick, E\.MultiBERTs: BERT reproductions for robustness analysis\.In*International Conference on Learning Representations*, 2022\.
- Touvron et al\. \(2021\)Touvron, H\., Cord, M\., Douze, M\., Massa, F\., Sablayrolles, A\., and Jégou, H\.Training data\-efficient image transformers & distillation through attention\.In*International Conference on Machine Learning*, pp\. 10347–10357\. PMLR, 2021\.

## Appendix ARaw Run Variability

Table 4:Mean test accuracy \(%\) and run\-to\-run standard deviation \(pp\) across 50 runs per cell\.

Similar Articles

Which Pairs to Compare for LLM Post-Training?

arXiv cs.AI

This paper studies the problem of selecting which completion pairs to label for human preference feedback in LLM post-training. It formulates comparison curation as a sampling-design problem, provides theoretical bounds on DPO's policy optimality gap, and proposes practical sampling designs that improve sample efficiency over common heuristics on synthetic and real benchmarks.

Diff Mining: Logit Differences Reveal Finetuning Objectives

arXiv cs.LG

The paper introduces Diff Mining, a framework for identifying finetuning objectives in language models by analyzing logit differences between finetuned and base models, enabling interpretable auditing of learned behaviors.

Don't let the model write the audit log

Reddit r/AI_Agents

The article warns against using model-generated narration as the authoritative audit log for AI agents, advocating for persisting raw tool call data instead, and suggests a simple diff check to catch discrepancies.