Sample Complexity of Multicalibration for Multilevel Properties
Summary
This paper studies the sample complexity of multicalibration for a sequence of properties that are sequentially identifiable, establishing matching upper and lower bounds up to logarithmic factors.
View Cached Full Text
Cached at: 08/06/26, 07:48 AM
# Sample Complexity of Multicalibration for Multilevel Properties
Source: [https://arxiv.org/html/2608.04288](https://arxiv.org/html/2608.04288)
Jiuyao LuThis work was done while interning at Amazon\.Department of Statistics and Data Science, The Wharton School, University of PennsylvaniaAleksandr PodkopaevAmazonShiva Prasad KasiviswanathanAmazon
###### Abstract
Calibration requires a predictor to be unbiased after conditioning on its own predictions\. Multicalibration asks for this guarantee simultaneously across a collection of groups\. Many prediction tasks ask for several related features of the same conditional outcome distribution: variance is defined relative to the mean, skewness relative to both mean and variance, and conditional value at risk relative to a quantile\. We study multicalibration for a sequence ofkkproperties in which each property is identifiable once the preceding properties are fixed\. This framework includes Bayes pairs but does not require the properties to arise from a single loss\.
For every fixedk≥2k\\geq 2, we establish matching upper and lower sample\-complexity bounds up to logarithmic factors under regularity conditions\. Even with only polylogarithmically many binary groups, achieving multicalibration errorε\\varepsilonrequiresΩ~\(ε−\(k\+2\)\)\\widetilde\{\\Omega\}\(\\varepsilon^\{\-\(k\+2\)\}\)samples\. Conversely, for any finite group family𝒢\\mathcal\{G\}, we give a randomized learner usingO\(ε−\(k\+2\)\+ε−2log\|𝒢\|\)O\(\\varepsilon^\{\-\(k\+2\)\}\+\\varepsilon^\{\-2\}\\log\|\\mathcal\{G\}\|\)samples\. Thus the sample complexity isΘ~\(ε−\(k\+2\)\)\\widetilde\{\\Theta\}\(\\varepsilon^\{\-\(k\+2\)\}\)for polynomial\-size group families\. We instantiate the theory for three canonical examples\.
###### Contents
1. [1Introduction](https://arxiv.org/html/2608.04288#S1)1. [1\.1Main Results](https://arxiv.org/html/2608.04288#S1.SS1) 2. [1\.2Proof Overview](https://arxiv.org/html/2608.04288#S1.SS2) 3. [1\.3Related Work](https://arxiv.org/html/2608.04288#S1.SS3)
2. [2Basic Setup](https://arxiv.org/html/2608.04288#S2)1. [2\.1Sequential Conditional Identifiability](https://arxiv.org/html/2608.04288#S2.SS1) 2. [2\.2Multicalibration Error](https://arxiv.org/html/2608.04288#S2.SS2) 3. [2\.3Minimax Sample Complexity](https://arxiv.org/html/2608.04288#S2.SS3)
3. [3The Lower Bound](https://arxiv.org/html/2608.04288#S3)1. [3\.1Regularity Assumptions](https://arxiv.org/html/2608.04288#S3.SS1) 2. [3\.2Main Result](https://arxiv.org/html/2608.04288#S3.SS2) 3. [3\.3The Candidate Distributions](https://arxiv.org/html/2608.04288#S3.SS3) 4. [3\.4Approximation of Threshold Signs with Walsh Functions](https://arxiv.org/html/2608.04288#S3.SS4) 5. [3\.5The Group Family](https://arxiv.org/html/2608.04288#S3.SS5) 6. [3\.6From Multicalibration to Prediction Accuracy](https://arxiv.org/html/2608.04288#S3.SS6) 7. [3\.7Recovering the Sign Assignment](https://arxiv.org/html/2608.04288#S3.SS7) 8. [3\.8Proof of the Main Theorem](https://arxiv.org/html/2608.04288#S3.SS8)
4. [4The Upper Bound](https://arxiv.org/html/2608.04288#S4)1. [4\.1The Online Formulation and Its Objectives](https://arxiv.org/html/2608.04288#S4.SS1) 2. [4\.2Regularity Assumptions](https://arxiv.org/html/2608.04288#S4.SS2) 3. [4\.3The Online Forecaster and Its Empirical Guarantee](https://arxiv.org/html/2608.04288#S4.SS3) 4. [4\.4Analysis of the Online Forecaster](https://arxiv.org/html/2608.04288#S4.SS4)1. [4\.4\.1Exponential Weights](https://arxiv.org/html/2608.04288#S4.SS4.SSS1) 2. [4\.4\.2The Minimax Value](https://arxiv.org/html/2608.04288#S4.SS4.SSS2) 5. [4\.5The Averaged Batch Predictor and Sample Complexity](https://arxiv.org/html/2608.04288#S4.SS5)1. [4\.5\.1Online\-to\-Batch Reduction](https://arxiv.org/html/2608.04288#S4.SS5.SSS1) 2. [4\.5\.2Sample Complexity](https://arxiv.org/html/2608.04288#S4.SS5.SSS2) 6. [4\.6Polynomial\-Time Implementation](https://arxiv.org/html/2608.04288#S4.SS6)1. [4\.6\.1Computing Exponential Weights Implicitly](https://arxiv.org/html/2608.04288#S4.SS6.SSS1) 2. [4\.6\.2The Optimization Oracle](https://arxiv.org/html/2608.04288#S4.SS6.SSS2)
5. [5Canonical Instantiations](https://arxiv.org/html/2608.04288#S5)1. [5\.1Mean, Mean Absolute Deviation](https://arxiv.org/html/2608.04288#S5.SS1) 2. [5\.2Mean, Variance, Skewness](https://arxiv.org/html/2608.04288#S5.SS2) 3. [5\.3Quantile, CVaR](https://arxiv.org/html/2608.04288#S5.SS3)
6. [References](https://arxiv.org/html/2608.04288#bib)
7. [AComparison with Multi\-Level Stochastic Optimization](https://arxiv.org/html/2608.04288#A1)
8. [BMeasurability of the Online Forecaster](https://arxiv.org/html/2608.04288#A2)
## 1Introduction
Calibration asks whether predictions have their intended statistical meaning\. For a predictor of the conditional mean, the basic requirement is that among the individuals who receive predictionpp, the average outcome is alsopp\[[11](https://arxiv.org/html/2608.04288#bib.bib13)\]\. This guarantee can still hide systematic errors: a predictor may underestimate outcomes on one structured subpopulation and compensate by overestimating them elsewhere\. Given a family𝒢\\mathcal\{G\}of groups, multicalibration requires the predictor to remain calibrated after restricting attention to any one of them\[[29](https://arxiv.org/html/2608.04288#bib.bib7)\]\. Multicalibration is useful when the same predictions inform many downstream decisions\. It supports predictors that perform well for many losses, as in omniprediction\[[23](https://arxiv.org/html/2608.04288#bib.bib19),[21](https://arxiv.org/html/2608.04288#bib.bib11),[39](https://arxiv.org/html/2608.04288#bib.bib29)\], and it has found applications in complexity theory and information aggregation\[[3](https://arxiv.org/html/2608.04288#bib.bib12),[14](https://arxiv.org/html/2608.04288#bib.bib35),[7](https://arxiv.org/html/2608.04288#bib.bib8),[6](https://arxiv.org/html/2608.04288#bib.bib10)\]\.
Prior work has developed multicalibration for conditional means, moments, and quantiles\[[32](https://arxiv.org/html/2608.04288#bib.bib22),[27](https://arxiv.org/html/2608.04288#bib.bib23),[33](https://arxiv.org/html/2608.04288#bib.bib25),[12](https://arxiv.org/html/2608.04288#bib.bib9),[38](https://arxiv.org/html/2608.04288#bib.bib31)\]\. In many applications, however, several related features of the conditional outcome law must be predicted together, and some are defined relative to others\. Quantiles describe coverage and tail events, dispersion measures describe uncertainty, and measures of tail risk summarize rare outcomes with potentially large costs\. For example, a three\-level example arises in hospital capacity planning, where a model may need to predict a patient’s expected length of stay, the variability around that expectation, and the degree of right\-tail asymmetry caused by occasional very long stays\. The model can output a prediction vectorp=\(m,σ2,γ\)p=\(m,\\sigma^\{2\},\\gamma\), representing the conditional mean, variance, and skewness of length of stay\. These quantities have a natural sequential structure\. At the first level, the mean predictionmmseparates patient groups whose stays are centered at different values\. At the second level, after the mean has been fixed, the variance predictionσ2\\sigma^\{2\}describes variability around that mean, so that differences in expected length of stay are not mistaken for uncertainty\. At the third level, after the mean and variance have been fixed, the skewness predictionγ\\gammadescribes asymmetry relative to that center and scale\. The third coordinate is important because two patient groups may have the same expected stay and the same overall variability but very different tail behavior: one may have roughly symmetric fluctuations around its mean, while the other may usually be discharged near the expected date but occasionally require a much longer hospitalization\. Calibrating skewness without also fixing the mean and variance could pool groups with different centers or scales, causing those lower\-order differences to appear spuriously as tail asymmetry\. Three\-level multicalibration therefore ensures that each prediction describes a distinct feature of the conditional outcome distribution: its location, its dispersion around that location, and the shape of its tail after location and dispersion have been accounted for\.
Formally, letℳ\\mathcal\{M\}be a class of probability laws on the outcome space𝒴\\mathcal\{Y\}, and letΓj:ℳ→ℝ\\Gamma\_\{j\}:\\mathcal\{M\}\\to\\mathbb\{R\}denote thejj\-th distributional property of interest\. For each contextxx, writeμx=ℒ\(Y∣X=x\)\\mu\_\{x\}=\\mathcal\{L\}\(Y\\mid X=x\)for the conditional law of the outcome\. Calibrating the propertiesΓ1,Γ2,Γ3\\Gamma\_\{1\},\\Gamma\_\{2\},\\Gamma\_\{3\}separately need not yield a jointly coherent predictor\. Indeed, consider the oracle scalar predictorfj\(x\)=Γj\(μx\)f\_\{j\}\(x\)=\\Gamma\_\{j\}\(\\mu\_\{x\}\)\. Conditioning on the eventfj\(X\)=af\_\{j\}\(X\)=apools the conditional lawsμx\\mu\_\{x\}into the mixture
μ¯a=ℒ\(Y∣fj\(X\)=a\)=∫μx𝑑ℙ\(x∣fj\(X\)=a\)\.\\bar\{\\mu\}\_\{a\}=\\mathcal\{L\}\(Y\\mid f\_\{j\}\(X\)=a\)=\\int\\mu\_\{x\}\\,d\\mathbb\{P\}\(x\\mid f\_\{j\}\(X\)=a\)\.Although every component of this mixture satisfiesΓj\(μx\)=a\\Gamma\_\{j\}\(\\mu\_\{x\}\)=a, calibration additionally requiresΓj\(μ¯a\)=a\\Gamma\_\{j\}\(\\bar\{\\mu\}\_\{a\}\)=a\. This implication holds only when the level setΓj−1\(a\)=\{μ∈ℳ:Γj\(μ\)=a\}\\Gamma\_\{j\}^\{\-1\}\(a\)=\\\{\\mu\\in\\mathcal\{M\}:\\Gamma\_\{j\}\(\\mu\)=a\\\}is closed under mixtures\. Thus, ifΓj\\Gamma\_\{j\}lacks convex level sets, even an oracle predictor can fail to be calibrated after conditioning on its own output\[[38](https://arxiv.org/html/2608.04288#bib.bib31)\]\. In a multilevel setting, the failure may arise because distributions that agree on a higher\-level propertyΓj\\Gamma\_\{j\}, but differ in lower\-level propertiesΓ1,…,Γj−1\\Gamma\_\{1\},\\ldots,\\Gamma\_\{j\-1\}, are pooled together, and these lower\-level differences change the value ofΓj\\Gamma\_\{j\}under mixing\.
Variance provides a concrete example of this failure\. Suppose
ℙ\[\(X,Y\)=\(1,1\)\]=ℙ\[\(X,Y\)=\(2,2\)\]=12,\\mathbb\{P\}\[\(X,Y\)=\(1,1\)\]=\\mathbb\{P\}\[\(X,Y\)=\(2,2\)\]=\\frac\{1\}\{2\},and consider the groups
g1\(x\)=𝟏\{x=1\},g2\(x\)=𝟏\{x=2\},gall\(x\)≡1\.g\_\{1\}\(x\)=\\mathbf\{1\}\\\{x=1\\\},\\qquad g\_\{2\}\(x\)=\\mathbf\{1\}\\\{x=2\\\},\\qquad g\_\{\\mathrm\{all\}\}\(x\)\\equiv 1\.The oracle conditional\-variance predictor outputs zero at both contexts because the outcome is deterministic conditional onXX\. It is calibrated withing1g\_\{1\}andg2g\_\{2\}, but not withingallg\_\{\\mathrm\{all\}\}\. Its single prediction cell pools the two outcomes, whose variance is1/41/4\. In fact,g1g\_\{1\}andg2g\_\{2\}force any variance\-only predictor to output zero at both contexts, whereas calibration withingallg\_\{\\mathrm\{all\}\}would require their common prediction to be1/41/4\. Thus no scalar variance predictor can be multicalibrated with respect to all three groups\. Predicting the mean together with the variance separates the two contexts through their mean coordinates and removes this conflict\.
This motivates predicting the ordered vectorf\(x\)=\(f1\(x\),f2\(x\),f3\(x\)\)f\(x\)=\\bigl\(f\_\{1\}\(x\),f\_\{2\}\(x\),f\_\{3\}\(x\)\\bigr\)and calibrating its coordinates jointly\. For simplicity, suppose that there are residual functionsR1,R2,R3R\_\{1\},R\_\{2\},R\_\{3\}such thatR1R\_\{1\}identifiesΓ1\\Gamma\_\{1\},R2R\_\{2\}identifiesΓ2\\Gamma\_\{2\}once the value ofΓ1\\Gamma\_\{1\}is fixed, andR3R\_\{3\}identifiesΓ3\\Gamma\_\{3\}once the values of bothΓ1\\Gamma\_\{1\}andΓ2\\Gamma\_\{2\}are fixed\. That is,
𝔼μ\[R1\(p1,Y\)\]=0⟺p1=Γ1\(μ\),\\mathbb\{E\}\_\{\\mu\}\\\!\\left\[R\_\{1\}\(p\_\{1\},Y\)\\right\]=0\\quad\\Longleftrightarrow\\quad p\_\{1\}=\\Gamma\_\{1\}\(\\mu\),and, conditional onp1=Γ1\(μ\)p\_\{1\}=\\Gamma\_\{1\}\(\\mu\),
𝔼μ\[R2\(p1,p2,Y\)\]=0⟺p2=Γ2\(μ\),\\mathbb\{E\}\_\{\\mu\}\\\!\\left\[R\_\{2\}\(p\_\{1\},p\_\{2\},Y\)\\right\]=0\\quad\\Longleftrightarrow\\quad p\_\{2\}=\\Gamma\_\{2\}\(\\mu\),while, conditional onp1=Γ1\(μ\)p\_\{1\}=\\Gamma\_\{1\}\(\\mu\)andp2=Γ2\(μ\)p\_\{2\}=\\Gamma\_\{2\}\(\\mu\),
𝔼μ\[R3\(p1,p2,p3,Y\)\]=0⟺p3=Γ3\(μ\)\.\\mathbb\{E\}\_\{\\mu\}\\\!\\left\[R\_\{3\}\(p\_\{1\},p\_\{2\},p\_\{3\},Y\)\\right\]=0\\quad\\Longleftrightarrow\\quad p\_\{3\}=\\Gamma\_\{3\}\(\\mu\)\.Three\-level multicalibration then requires these residuals to have mean zero on every relevant subgroup and within each prediction cell indexed byf=\(f1,f2,f3\)f=\(f\_\{1\},f\_\{2\},f\_\{3\}\)\. The earlier coordinates therefore prevent contexts with different predictions at earlier levels from being pooled when calibrating a later coordinate, allowing a property that is not calibratable in isolation to become calibratable relative to the preceding predictions\.
This motivates us to study the sample complexity of multilevel multicalibration under an ECE\-type metric defined from these residuals, where ECE denotes expected calibration error\[[36](https://arxiv.org/html/2608.04288#bib.bib46),[9](https://arxiv.org/html/2608.04288#bib.bib43)\]\. Let\(Γ1,…,Γk\)\(\\Gamma\_\{1\},\\ldots,\\Gamma\_\{k\}\)be an ordered collection of distributional properties such that, for each levelj∈\[k\]j\\in\[k\], the propertyΓj\\Gamma\_\{j\}is identifiable conditional on the preceding properties\(Γ1,…,Γj−1\)\(\\Gamma\_\{1\},\\ldots,\\Gamma\_\{j\-1\}\)\. Givennnindependent samples and a target accuracyε\\varepsilon, we ask when a learner can produce a predictorf=\(f1,…,fk\)f=\(f\_\{1\},\\ldots,f\_\{k\}\)with multicalibration error at mostε\\varepsilonover every relevant subgroup\. Our goal is to characterize how the required sample sizennscales withε\\varepsilonand the number of groups, for each fixedkk\. In particular, we ask whether exploiting the conditional structure permits sample complexity comparable to scalar or vector\-valued mean multicalibration, or whether the dependence among levels introduces an unavoidable additional statistical cost\.
### 1\.1Main Results
We prove upper and lower bounds on the sample complexity of multicalibration that match up to logarithmic factors for sequentially conditionally identifiable multilevel properties\. For finitely supported predictors, our multicalibration error is the maximum over groups of theℓ1\\ell\_\{1\}expected calibration error computed by bucketing according to the full prediction vector\. \(See Section[2](https://arxiv.org/html/2608.04288#S2)for the definitions of sequential conditional identifiability and multicalibration error for general predictors\.\) The lower and upper bounds rely on different regularity assumptions, all of which are satisfied by our canonical examples\.
##### Lower Bound\.
For every sufficiently smallε\\varepsilon, we construct a finite context space, polylogarithmically many binary groups, and a collection of data distributions such that any learner achieving multicalibration error at mostε\\varepsilonwith probability at least2/32/3must have sample size
n=Ω\(ε−\(k\+2\)logk\+2\(1/ε\)\)=Ω~\(ε−\(k\+2\)\)\.n=\\Omega\\\!\\left\(\\frac\{\\varepsilon^\{\-\(k\+2\)\}\}\{\\log^\{k\+2\}\(1/\\varepsilon\)\}\\right\)=\\widetilde\{\\Omega\}\\\!\\left\(\\varepsilon^\{\-\(k\+2\)\}\\right\)\.The result holds against randomized learners\. Since the constructed group family has polylogarithmic size, it lies within every group\-size budget of the formε−κ\\varepsilon^\{\-\\kappa\}with fixedκ\>0\\kappa\>0, onceε\\varepsilonis sufficiently small\.
##### Upper Bound\.
We give an online forecaster and convert its sequence of prediction rules into a randomized predictor\. For a finite family𝒢\\mathcal\{G\}of groups, the resulting learner uses
n=O\(ε−\(k\+2\)\+ε−2log\|𝒢\|\)\.n=O\\\!\\left\(\\varepsilon^\{\-\(k\+2\)\}\+\\varepsilon^\{\-2\}\\log\|\\mathcal\{G\}\|\\right\)\.Thus, whenever\|𝒢\|\|\\mathcal\{G\}\|is polynomial in1/ε1/\\varepsilon, the upper and lower bounds have the same exponentk\+2k\+2\. We also give conditions under which the learner can be implemented in polynomial time\.
##### Canonical Properties\.
We verify the conditions of both bounds in three cases: mean and mean absolute deviation; mean, variance, and skewness; and quantile and conditional value at risk\. The two\-level examples have sample complexityΘ~\(ε−4\)\\widetilde\{\\Theta\}\(\\varepsilon^\{\-4\}\), while the three\-level example has sample complexityΘ~\(ε−5\)\\widetilde\{\\Theta\}\(\\varepsilon^\{\-5\}\)\.
The results ofCollinaet al\.\[[9](https://arxiv.org/html/2608.04288#bib.bib43)\]are most closely related to ours\. They establish the exponent33for scalar mean multicalibration andd\+2d\+2fordd\-dimensional vector means, and also treat regular scalar elicitable properties\. WhenΓ\(μ\)=𝔼μ\[ψ\(Y\)\]\\Gamma\(\\mu\)=\\mathbb\{E\}\_\{\\mu\}\[\\psi\(Y\)\]for a fixed vector\-valued functionψ\\psi, vector mean multicalibration applies directly by treatingψ\(Y\)\\psi\(Y\)as the outcome\. More generally, a multilevel property need not admit such a representation, and a residual at a later level may depend on the earlier predictions\. Our bounds allow this structure\.
### 1\.2Proof Overview
##### Lower Bound\.
The lower bound proceeds by considering the task of determining which distribution generated the samples\. The proof has four steps\. First, we construct a large finite collection of candidate distributions\. The true distribution is an unknown member of this collection, and the samples are drawn from it\. Second, we show that a predictor with small multicalibration error must have small prediction error\. Third, we use such a predictor to determine which member generated the samples\. Finally, we show that determining the true distribution requires many samples\.
Fix a resolution parameterQQ\. We take the context space to be the discrete grid\{0,…,Q−1\}k\\\{0,\\ldots,Q\-1\\\}^\{k\}, which containsQkQ^\{k\}contexts\. We associate each context with an unperturbed vector in thekk\-dimensional property space; these vectors form a Cartesian grid\. We assign one hidden sign in\{−1,\+1\}\\\{\-1,\+1\\\}to each context\. According to this sign assignment, we shift the final coordinate of the corresponding property vector by an amount of order1/Q1/Q, while leaving the firstk−1k\-1coordinates unchanged\. We selectexp\(Ω\(Qk\)\)\\exp\(\\Omega\(Q^\{k\}\)\)sign assignments such that any two of them disagree on a constant fraction of the contexts\. Our regularity assumptions \(Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(i\)\) ensure that every resulting vector is the property value of some outcome distribution\. Taking the context to be uniform and using the corresponding outcome distribution conditionally on each context produces one candidate data distribution for every sign assignment\. The resulting statistical task is to identify which candidate distribution generated the data, or equivalently, its sign assignment\.
The second step converts multicalibration error into prediction error\. Consider a given level, and suppose first that the predictions at all preceding levels are correct\. Under our regularity assumptions \(Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(ii\)\), the expected residual at the current level has the same sign as the difference between the prediction at that level and its true value, and its magnitude is at least proportional to the absolute value of this difference\. Incorrect predictions at preceding levels can change the current residual, but the change is controlled by the corresponding prediction errors \(see Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(iv\)\)\. Thus, the signed residual controls the current prediction error up to the prediction errors at preceding levels\. The required sign depends on both the prediction and the context, whereas a group can depend only on the context\. For a fixed prediction, however, this sign is a threshold function of one context coordinate\. We construct a common collection of only polylogarithmically many binary groups whose linear combinations approximate all the threshold functions needed in the proof \(see Lemma[3\.3](https://arxiv.org/html/2608.04288#S3.Thmtheorem3)\)\. Multiplying the residual by such an approximation produces, up to the approximation error, a linear combination of group\-weighted residuals\. Each of these residual terms is controlled by the multicalibration error \(see Lemmas[3\.4](https://arxiv.org/html/2608.04288#S3.Thmtheorem4)and[3\.6](https://arxiv.org/html/2608.04288#S3.Thmtheorem6)\)\. Since the sum of the absolute values of the coefficients isO\(logQ\)O\(\\log Q\), the signed residual, and hence the current prediction error, is controlled byO\(logQ\)O\(\\log Q\)times the multicalibration error, together with the prediction errors at preceding levels\.
Because a residual at leveljjmay depend on all preceding predictions, we apply this argument successively\. After controlling the errors in the firstj−1j\-1coordinates, the Lipschitz dependence of the level\-jjresidual on preceding predictions allows us to control the error in coordinatejj\. This gives bounds for the firstk−1k\-1coordinates\. For the final coordinate, we additionally separate predictions according to whether they lie inside or outside disjoint neighborhoods of the unperturbed property vectors\. Threshold functions control the predictions outside these neighborhoods, while the group containing all contexts controls those inside\. Altogether, multicalibration error at mostε\\varepsilonforces the expectedℓ1\\ell\_\{1\}distance between the prediction and the true property vector to beO\(εlogQ\)O\(\\varepsilon\\log Q\), where the expectation is over a uniformly random context and the predictor’s randomization \(see Proposition[3\.7](https://arxiv.org/html/2608.04288#S3.Thmtheorem7)\)\.
The third step uses this prediction guarantee to recover the sign assignment\. Because any two candidate sign assignments disagree on a constant fraction of the contexts, their associated property vectors are separated by order1/Q1/Qon average\. Given a possibly randomized predictor, we take its mean output at every context and select the sign assignment whose associated property vectors are closest to these mean predictions\. If the average prediction error is smaller than a sufficiently small constant multiple of1/Q1/Q, the true sign assignment is the unique closest one \(see Lemmas[3\.8](https://arxiv.org/html/2608.04288#S3.Thmtheorem8)and[3\.9](https://arxiv.org/html/2608.04288#S3.Thmtheorem9)\)\. Consequently, any learner that produces a predictor with multicalibration error at mostε\\varepsilonalso provides a way to determine the true distribution whenever
εlogQ≲1Q\.\\varepsilon\\log Q\\lesssim\\frac\{1\}\{Q\}\.
Finally, our regularity assumptions \(Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(v\)\) ensure that the Kullback–Leibler divergence between nearby outcome distributions is at most a constant times the squared distance between their property vectors\. Since the candidates differ by onlyO\(1/Q\)O\(1/Q\)at each context, the pairwise Kullback–Leibler divergence between their one\-sample data distributions isO\(1/Q2\)O\(1/Q^\{2\}\), and fornnindependent samples it isO\(n/Q2\)O\(n/Q^\{2\}\)\(see Lemma[3\.10](https://arxiv.org/html/2608.04288#S3.Thmtheorem10)\)\. On the other hand, the logarithm of the number of candidates isΩ\(Qk\)\\Omega\(Q^\{k\}\)\. Using Fano’s inequality, we show in Proposition[3\.11](https://arxiv.org/html/2608.04288#S3.Thmtheorem11)that identifying the true candidate with constant success probability requires
nQ2≳Qk,and hencen=Ω\(Qk\+2\)\.\\frac\{n\}\{Q^\{2\}\}\\gtrsim Q^\{k\},\\qquad\\text\{and hence\}\\qquad n=\\Omega\(Q^\{k\+2\}\)\.ChoosingQQat the largest scale allowed byεlogQ≲1/Q\\varepsilon\\log Q\\lesssim 1/Q, namely
Q≍1εlog\(1/ε\),Q\\asymp\\frac\{1\}\{\\varepsilon\\log\(1/\\varepsilon\)\},gives
n=Ω\(ε−\(k\+2\)logk\+2\(1/ε\)\)\.n=\\Omega\\left\(\\frac\{\\varepsilon^\{\-\(k\+2\)\}\}\{\\log^\{k\+2\}\(1/\\varepsilon\)\}\\right\)\.The constructed family contains only polylogarithmically many groups, so it lies within every polynomial group\-size budget for sufficiently smallε\\varepsilon\.
##### Upper Bound\.
Our upper\-bound construction proceeds through an online forecasting problem\. Fix a finite grid𝒫Q\\mathcal\{P\}\_\{Q\}ofkk\-dimensional prediction vectors, whereQQcontrols the resolution and\|𝒫Q\|=O\(Qk\)\|\\mathcal\{P\}\_\{Q\}\|=O\(Q^\{k\}\)\. At roundtt, the forecaster uses the history from the preceding rounds to construct a randomized prediction rule
πt:𝒳→Δ\(𝒫Q\),\\pi\_\{t\}:\\mathcal\{X\}\\to\\Delta\(\\mathcal\{P\}\_\{Q\}\),whereΔ\(𝒫Q\)\\Delta\(\\mathcal\{P\}\_\{Q\}\)denotes the set of distributions over the grid\. The current contextXtX\_\{t\}is then revealed, and the forecaster predicts the distributionπt\(⋅∣Xt\)\\pi\_\{t\}\(\\cdot\\mid X\_\{t\}\)\. After observing the outcomeYtY\_\{t\}, it evaluates the residuals associated with this distribution\. We first control the empirical analogue of multicalibration error accumulated over these online rounds\.
Given a dataset consisting ofTTi\.i\.d\. context–outcome pairs, we run this online forecasting procedure forTTrounds, using thett\-th pair on roundtt\. The contextXtX\_\{t\}is revealed before the prediction and the outcomeYtY\_\{t\}afterward\. This run producesTTprediction rulesπ1,…,πT\\pi\_\{1\},\\ldots,\\pi\_\{T\}\. We return their average:
Π\(x\)=1T∑t=1Tπt\(⋅∣x\)\.\\Pi\(x\)=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\pi\_\{t\}\(\\cdot\\mid x\)\.An online\-to\-batch argument transfers the empirical guarantee for the online transcript to a population multicalibration guarantee forΠ\\Pi\.
The empirical objective contains many calibration requirements\. For every groupgg, pointp∈𝒫Qp\\in\\mathcal\{P\}\_\{Q\}, and leveljj, it records the cumulative level\-jjresidual associated withgg, with the residual on roundttweighted by the probabilityπt\(p∣Xt\)\\pi\_\{t\}\(p\\mid X\_\{t\}\)assigned topp\. For each group, the empirical error sums the absolute values of these cumulative residuals over all pairs\(p,j\)\(p,j\)\. We can express this sum as a maximum over sign arrayss=\(sp,j\)s=\(s\_\{p,j\}\), with one sign assigned to each pair\(p,j\)\(p,j\)\. After also taking the maximum over groups, the empirical error becomes the maximum of a collection of signed residual objectives indexed by pairs\(g,s\)\(g,s\)\. Controlling all of these objectives is therefore equivalent to controlling the empirical multicalibration error\.
The forecaster assigns a weight to each objective at the beginning of every round\. These weights depend on the signed residuals accumulated on previous rounds\. If a calibration violation has grown large for some group and some pattern of signs across prediction cells and levels, the corresponding objective receives greater weight, so the next prediction places more emphasis on reducing that violation\. Objectives whose cumulative violations remain small receive less relative emphasis\. Once the weights have been fixed, every candidate distribution on𝒫Q\\mathcal\{P\}\_\{Q\}has a weighted expected residual under each possible conditional outcome law\. The forecaster chooses among these distributions in a minimax manner by minimizing the largest such weighted expected residual over all admissible outcome laws\.
Formally, we use the framework of online learning with expert advice\. We treat each objective as an expert and its signed residual on each round as the expert’s gain\. The exponential weights algorithm assigns weights to the experts according to their cumulative gains, producing the objective weights described above\. Its regret guarantee bounds the largest cumulative gain of any expert by the sum of the expected gains under these weights, plus a regret term \(Lemma[4\.4](https://arxiv.org/html/2608.04288#S4.Thmtheorem4)\)\. Because the empirical multicalibration error is exactly the largest expert gain divided byTT, it remains to control these expected gains\. Conditional on the past and the current context, the expected gain on each round is bounded by the value of the minimax problem solved on that round\.
We bound this value by reversing the order of play using Sion’s minimax theorem \(Lemma[4\.5](https://arxiv.org/html/2608.04288#S4.Thmtheorem5)\)\. In the reversed game, an admissible outcome lawμ\\muis fixed before the prediction is selected\. For high\-dimensional mean prediction, it suffices to choose a grid point close to the revealed outcome vector\[[37](https://arxiv.org/html/2608.04288#bib.bib33),[9](https://arxiv.org/html/2608.04288#bib.bib43)\]\. In our setting, the residual need not be the difference between a prediction and an outcome, and a residual at a later level may depend on the entire preceding prediction\. Instead, we assume that, for everyμ\\mu, there is a single grid point at which all expected residuals are simultaneouslyO\(1/Q\)O\(1/Q\)\(Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(iii\)\)\. Choosing this point in the reversed game makes every weighted objectiveO\(1/Q\)O\(1/Q\)\. The minimax theorem then guarantees that, in the original order of play, there is a distribution on𝒫Q\\mathcal\{P\}\_\{Q\}whose worst\-case weighted expected residual is equally small, even though the forecaster does not knowμ\\mu\.
Combining this bound with the regret guarantee of exponential weights, Theorem[4\.3](https://arxiv.org/html/2608.04288#S4.Thmtheorem3)gives
O\(1Q\+Qk\+log\|𝒢\|T\)O\\left\(\\frac\{1\}\{Q\}\+\\sqrt\{\\frac\{Q^\{k\}\+\\log\|\\mathcal\{G\}\|\}\{T\}\}\\right\)expected empirical multicalibration error, apart from the chosen accuracy for solving the minimax problem\.
Finally, we convert the online guarantee into a batch guarantee\. A martingale concentration argument shows that the population multicalibration error of the averaged predictorΠ\\Piexceeds the empirical multicalibration error of the online transcript by at most a term of the same order \(Lemma[4\.7](https://arxiv.org/html/2608.04288#S4.Thmtheorem7)\)\. TakingQ=Θ\(1/ε\)Q=\\Theta\(1/\\varepsilon\)and
T=O\(ε−\(k\+2\)\+ε−2log\|𝒢\|\)T=O\\left\(\\varepsilon^\{\-\(k\+2\)\}\+\\varepsilon^\{\-2\}\\log\|\\mathcal\{G\}\|\\right\)gives the upper bound on sample complexity in Corollary[4\.9](https://arxiv.org/html/2608.04288#S4.Thmtheorem9)\.
Although the expert family is exponentially large, its weights can be computed without enumerating all experts\. The minimax step reduces to linear optimization overℳ\\mathcal\{M\}, as specified in Assumption[4\.10](https://arxiv.org/html/2608.04288#S4.Thmtheorem10)\. For each of our canonical examples, we implement this oracle in polynomial time \(Propositions[5\.3](https://arxiv.org/html/2608.04288#S5.Thmtheorem3),[5\.7](https://arxiv.org/html/2608.04288#S5.Thmtheorem7), and[5\.12](https://arxiv.org/html/2608.04288#S5.Thmtheorem12)\), yielding the polynomial\-time learner in Theorem[4\.11](https://arxiv.org/html/2608.04288#S4.Thmtheorem11)\.
### 1\.3Related Work
##### Multicalibration and Sample Complexity\.
Hebert\-Johnsonet al\.\[[29](https://arxiv.org/html/2608.04288#bib.bib7)\]introduced multicalibration\. Subsequent work developed online algorithms and efficient implementations, and connected multicalibration to multiobjective optimization and omniprediction\[[34](https://arxiv.org/html/2608.04288#bib.bib28),[28](https://arxiv.org/html/2608.04288#bib.bib16),[17](https://arxiv.org/html/2608.04288#bib.bib17),[18](https://arxiv.org/html/2608.04288#bib.bib15)\]\. Other formulations measure calibration differently\.Gopalanet al\.\[[24](https://arxiv.org/html/2608.04288#bib.bib20)\]consider a multicalibration condition that tests residuals after restricting attention to observations whose predictions fall in an interval\.Globus\-Harriset al\.\[[20](https://arxiv.org/html/2608.04288#bib.bib21)\]study multicalibration under a weightedL2L\_\{2\}metric and give a boosting algorithm for controlling it\. When\|𝒢\|\|\\mathcal\{G\}\|is polynomial in1/ε1/\\varepsilon, the guarantees ofGopalanet al\.\[[24](https://arxiv.org/html/2608.04288#bib.bib20)\]andGlobus\-Harriset al\.\[[20](https://arxiv.org/html/2608.04288#bib.bib21)\]translate into sample complexities ofO~\(ε−8\)\\widetilde\{O\}\(\\varepsilon^\{\-8\}\)andO~\(ε−10\)\\widetilde\{O\}\(\\varepsilon^\{\-10\}\)for mean ECE, respectively\[[9](https://arxiv.org/html/2608.04288#bib.bib43)\]\.
Gopalanet al\.\[[25](https://arxiv.org/html/2608.04288#bib.bib18)\]introduce swap multicalibration, a stronger condition in which the group function used to test residual bias may vary across prediction values\.Luoet al\.\[[35](https://arxiv.org/html/2608.04288#bib.bib32)\]improve the online rates and sample\-complexity upper bounds for swap multicalibration of conditional means when residual bias is tested using bounded linear functions of the context\.
The weaker notion of calibrated multiaccuracy asks a predictor to be globally calibrated and multiaccurate with respect to a class of groups, meaning that its residuals are uncorrelated with every group function in the class\[[4](https://arxiv.org/html/2608.04288#bib.bib27)\]\.Gibbs and Tibshirani \[[19](https://arxiv.org/html/2608.04288#bib.bib1)\]establish anΩ\(ε−5/2\)\\Omega\(\\varepsilon^\{\-5/2\}\)sample\-complexity lower bound for learning deterministic predictors of conditional means under this requirement\.
Collinaet al\.\[[9](https://arxiv.org/html/2608.04288#bib.bib43)\]study batch sample complexity for scalar and vector means, weighted error metrics, and regular scalar elicitable properties\.Collinaet al\.\[[8](https://arxiv.org/html/2608.04288#bib.bib26)\]study the online setting, establish sharp lower bounds, and develop the subsampled Walsh family used in our group construction\.
A related but distinct statistical question is whether empirical multicalibration error reliably estimates population multicalibration error throughout a chosen predictor class\.Shabatet al\.\[[44](https://arxiv.org/html/2608.04288#bib.bib14)\]establish bounds depending on the size or graph dimension of the class, whileRosenberget al\.\[[41](https://arxiv.org/html/2608.04288#bib.bib30)\]connect this question to standard ERM generalization bounds\. This guarantee concerns estimation within the class, not the existence of a low\-error predictor in it\.
##### Calibration of Distributional Properties\.
Junget al\.\[[32](https://arxiv.org/html/2608.04288#bib.bib22)\]introduce mean\-conditioned moment multicalibration, andGuptaet al\.\[[27](https://arxiv.org/html/2608.04288#bib.bib23)\]give an online algorithm for this notion\. Multivalid prediction intervals have also been studied in online and batch settings\[[27](https://arxiv.org/html/2608.04288#bib.bib23),[2](https://arxiv.org/html/2608.04288#bib.bib24),[33](https://arxiv.org/html/2608.04288#bib.bib25)\]\.Noarov and Roth \[[38](https://arxiv.org/html/2608.04288#bib.bib31)\]connect property multicalibration to elicitation and identification\. They characterize which continuous scalar properties are sensible for calibration under mild conditions and give an algorithm for jointly multicalibrating two\-level conditionally elicitable properties, including Bayes pairs\.Huet al\.\[[31](https://arxiv.org/html/2608.04288#bib.bib34)\]give an oracle\-efficient online algorithm for swap multicalibration of elicitable properties\. We study the sample complexity of joint calibration for an arbitrary fixed number of conditional levels, including sequences that are not generated by a single loss and its Bayes risk\.
##### High\-Dimensional Prediction\.
Recent work studies calibration and related unbiasedness guarantees for vector\-valued predictions, including computational questions and online rates that depend on dimension\[[22](https://arxiv.org/html/2608.04288#bib.bib42),[40](https://arxiv.org/html/2608.04288#bib.bib40),[16](https://arxiv.org/html/2608.04288#bib.bib41),[37](https://arxiv.org/html/2608.04288#bib.bib33)\]\.Noarovet al\.\[[37](https://arxiv.org/html/2608.04288#bib.bib33)\]give an online algorithm for predicting high\-dimensional states such that, on every event in a specified family, the accumulated prediction error is small in every coordinate\. These events may depend on the past transcript, the current context, and the prediction\. In our setting, each coordinate predicts one property in an ordered sequence rather than one component of a state vector, and the residual for a later property may depend on the preceding predictions\.
##### Organization\.
Section[2](https://arxiv.org/html/2608.04288#S2)defines sequential conditional identifiability and the multicalibration error used throughout\. Sections[3](https://arxiv.org/html/2608.04288#S3)and[4](https://arxiv.org/html/2608.04288#S4)prove the lower and upper bounds, respectively\. Section[5](https://arxiv.org/html/2608.04288#S5)verifies the assumptions in our three canonical examples\.
## 2Basic Setup
### 2\.1Sequential Conditional Identifiability
Let𝒴\\mathcal\{Y\}be a standard Borel outcome space and letℳ\\mathcal\{M\}be a class of probability measures on𝒴\\mathcal\{Y\}\. Fix an integerk≥2k\\geq 2and a compact rectangular prediction space
𝒫=𝒫1×⋯×𝒫k⊂ℝk\.\\mathcal\{P\}=\\mathcal\{P\}\_\{1\}\\times\\cdots\\times\\mathcal\{P\}\_\{k\}\\subset\\mathbb\{R\}^\{k\}\.Throughout,kkis treated as fixed\. Constants suppressed by asymptotic notation may depend onkk, as well as on the stated regularity parameters\. A prediction value is denotedp=\(p1,…,pk\)p=\(p\_\{1\},\\ldots,p\_\{k\}\)\. Forj∈\[k\]j\\in\[k\], write
p<j:=\(p1,…,pj−1\),𝒫<j:=𝒫1×⋯×𝒫j−1,p\_\{<j\}:=\(p\_\{1\},\\ldots,p\_\{j\-1\}\),\\qquad\\mathcal\{P\}\_\{<j\}:=\\mathcal\{P\}\_\{1\}\\times\\cdots\\times\\mathcal\{P\}\_\{j\-1\},with𝒫<1\\mathcal\{P\}\_\{<1\}interpreted as a one\-point set\.
###### Definition 2\.1\(Sequentially Conditionally Identifiable Property\)\.
Akk\-level property is a map
Γ=\(Γ1,…,Γk\):ℳ→𝒫\.\\Gamma=\(\\Gamma\_\{1\},\\ldots,\\Gamma\_\{k\}\):\\mathcal\{M\}\\to\\mathcal\{P\}\.Forμ∈ℳ\\mu\\in\\mathcal\{M\}, write
Γ<j\(μ\):=\(Γ1\(μ\),…,Γj−1\(μ\)\)\.\\Gamma\_\{<j\}\(\\mu\):=\(\\Gamma\_\{1\}\(\\mu\),\\ldots,\\Gamma\_\{j\-1\}\(\\mu\)\)\.The propertyΓ\\Gammais sequentially conditionally identifiable if, for every levelj∈\[k\]j\\in\[k\], there is a jointly measurable residual function
Rj:𝒫<j×𝒫j×𝒴→ℝR\_\{j\}:\\mathcal\{P\}\_\{<j\}\\times\\mathcal\{P\}\_\{j\}\\times\\mathcal\{Y\}\\to\\mathbb\{R\}with the following properties\. For everyμ∈ℳ\\mu\\in\\mathcal\{M\}and every\(a,q\)∈𝒫<j×𝒫j\(a,q\)\\in\\mathcal\{P\}\_\{<j\}\\times\\mathcal\{P\}\_\{j\}, the functionRj\(a,q,⋅\)R\_\{j\}\(a,q,\\cdot\)isμ\\mu\-integrable\. Moreover, for everyμ∈ℳ\\mu\\in\\mathcal\{M\}and everyq∈𝒫jq\\in\\mathcal\{P\}\_\{j\},
𝔼Y∼μRj\(Γ<j\(μ\),q,Y\)=0⟺q=Γj\(μ\)\.\\mathbb\{E\}\_\{Y\\sim\\mu\}R\_\{j\}\(\\Gamma\_\{<j\}\(\\mu\),q,Y\)=0\\quad\\Longleftrightarrow\\quad q=\\Gamma\_\{j\}\(\\mu\)\.
Thus leveljjis identified after the previous true levels have been fixed\. The residual functions are also evaluated away from the true prefix, because predictions need not have correct earlier coordinates\. The residual vector used throughout the paper is
R\(p,y\):=\(R1\(p<1,p1,y\),…,Rk\(p<k,pk,y\)\)\.R\(p,y\):=\\bigl\(R\_\{1\}\(p\_\{<1\},p\_\{1\},y\),\\ldots,R\_\{k\}\(p\_\{<k\},p\_\{k\},y\)\\bigr\)\.When the final argument is a distribution, expectation over that distribution is denoted by:
Rj\(a,q,μ\):=𝔼Y∼μRj\(a,q,Y\),R\(p,μ\):=𝔼Y∼μR\(p,Y\)\.R\_\{j\}\(a,q,\\mu\):=\\mathbb\{E\}\_\{Y\\sim\\mu\}R\_\{j\}\(a,q,Y\),\\qquad R\(p,\\mu\):=\\mathbb\{E\}\_\{Y\\sim\\mu\}R\(p,Y\)\.Sequential conditional identifiability implies
R\(Γ\(μ\),μ\)=0∀μ∈ℳ\.R\(\\Gamma\(\\mu\),\\mu\)=0\\qquad\\forall\\mu\\in\\mathcal\{M\}\.
The familiar Bayes pair setting is the special casek=2k=2\[[38](https://arxiv.org/html/2608.04288#bib.bib31)\]\. If a scalar propertyγ\\gammahas identification functionrγr\_\{\\gamma\}, and ifLLis a strictly consistent loss forγ\\gamma\(soγ\(μ\)\\gamma\(\\mu\)uniquely minimizes𝔼μL\(r,Y\)\\mathbb\{E\}\_\{\\mu\}L\(r,Y\)\), then its Bayes risk is
ℬ\(μ\):=𝔼Y∼μL\(γ\(μ\),Y\),\\mathcal\{B\}\(\\mu\):=\\mathbb\{E\}\_\{Y\\sim\\mu\}L\(\\gamma\(\\mu\),Y\),and the Bayes pair isΓ\(μ\)=\(γ\(μ\),ℬ\(μ\)\)\\Gamma\(\\mu\)=\(\\gamma\(\\mu\),\\mathcal\{B\}\(\\mu\)\)\. It is sequentially conditionally identifiable with residual functions
R1\(∅,p1,y\)=rγ\(p1,y\),R2\(p1,p2,y\)=p2−L\(p1,y\)\.R\_\{1\}\(\\varnothing,p\_\{1\},y\)=r\_\{\\gamma\}\(p\_\{1\},y\),\\qquad R\_\{2\}\(p\_\{1\},p\_\{2\},y\)=p\_\{2\}\-L\(p\_\{1\},y\)\.Forτ∈\(0,1\)\\tau\\in\(0,1\), the quantile and conditional\-value\-at\-risk example in Section[5\.3](https://arxiv.org/html/2608.04288#S5.SS3)uses the loss
Lτ\(q,y\):=q\+\(y−q\)\+1−τL\_\{\\tau\}\(q,y\):=q\+\\frac\{\(y\-q\)\_\{\+\}\}\{1\-\\tau\}which is strictly consistent forqτq\_\{\\tau\}on the distribution class considered there and has Bayes riskCVaRτ\\operatorname\{CVaR\}\_\{\\tau\}\.
### 2\.2Multicalibration Error
Let𝒳\\mathcal\{X\}be a standard Borel context space and let𝖣\\mathsf\{D\}be a distribution on𝒳×𝒴\\mathcal\{X\}\\times\\mathcal\{Y\}\. Let𝖣Y∣X=x\\mathsf\{D\}\_\{Y\\mid X=x\}denote a regular conditional law ofYYgivenX=xX=x\. We call𝖣\\mathsf\{D\}ℳ\\mathcal\{M\}\-compatible if
𝖣Y∣X=x∈ℳfor𝖣X\-almost everyx\.\\mathsf\{D\}\_\{Y\\mid X=x\}\\in\\mathcal\{M\}\\qquad\\text\{for $\\mathsf\{D\}\_\{X\}$\-almost every $x$\}\.A randomized predictor assigns to each contextx∈𝒳x\\in\\mathcal\{X\}a distributionΠx∈Δ\(𝒫\)\\Pi\_\{x\}\\in\\Delta\(\\mathcal\{P\}\)over prediction values\. This assignment is measurable in the sense that, for every Borel setA⊆𝒫A\\subseteq\\mathcal\{P\}, the mapx↦Πx\(A\)x\\mapsto\\Pi\_\{x\}\(A\)is measurable\. We write
Π=\(Πx\)x∈𝒳\.\\Pi=\(\\Pi\_\{x\}\)\_\{x\\in\\mathcal\{X\}\}\.
For a given𝖣\\mathsf\{D\}andΠ\\Pi, consider the absolute\-integrability condition
𝔼\(X,Y\)∼𝖣\[∫𝒫‖R\(p,Y\)‖1ΠX\(dp\)\]<∞\.\\mathbb\{E\}\_\{\(X,Y\)\\sim\\mathsf\{D\}\}\\left\[\\int\_\{\\mathcal\{P\}\}\\\|R\(p,Y\)\\\|\_\{1\}\\,\\Pi\_\{X\}\(dp\)\\right\]<\\infty\.\(1\)When \([1](https://arxiv.org/html/2608.04288#S2.E1)\) holds, a measurable weightw:𝒳→\[−1,1\]w:\\mathcal\{X\}\\to\[\-1,1\]defines a finite vector signed measure on𝒫\\mathcal\{P\}by
νw𝖣,Π\(A\)\\displaystyle\\nu^\{\\mathsf\{D\},\\Pi\}\_\{w\}\(A\):=𝔼\(X,Y\)∼𝖣\[w\(X\)∫AR\(p,Y\)ΠX\(dp\)\]\\displaystyle=\\mathbb\{E\}\_\{\(X,Y\)\\sim\\mathsf\{D\}\}\\left\[w\(X\)\\int\_\{A\}R\(p,Y\)\\,\\Pi\_\{X\}\(dp\)\\right\]=𝔼\(X,Y\)∼𝖣P∼ΠX\[w\(X\)𝟏\{P∈A\}R\(P,Y\)\],for measurableA⊆𝒫\.\\displaystyle=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}\(X,Y\)\\sim\\mathsf\{D\}\\\\ P\\sim\\Pi\_\{X\}\\end\{subarray\}\}\\left\[w\(X\)\\mathbf\{1\}\\\{P\\in A\\\}R\(P,Y\)\\right\],\\qquad\\text\{for measurable $A\\subseteq\\mathcal\{P\}$\}\.
The population expected calibration error forΓ\\Gamma\(Γ\\Gamma\-ECE\) ofΠ\\Piwith respect towwis the coordinatewise total variation
Err𝖣Γ\(Π;w\)\\displaystyle\\operatorname\{Err\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi;w\):=∑j=1k\|νw,j𝖣,Π\|\(𝒫\)\\displaystyle=\\sum\_\{j=1\}^\{k\}\|\\nu^\{\\mathsf\{D\},\\Pi\}\_\{w,j\}\|\(\\mathcal\{P\}\)=∑j=1k𝔼\(X,Y\)∼𝖣P∼ΠX\[\|𝔼\[w\(X\)Rj\(P<j,Pj,Y\)∣P\]\|\]\.\\displaystyle=\\sum\_\{j=1\}^\{k\}\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}\(X,Y\)\\sim\\mathsf\{D\}\\\\ P\\sim\\Pi\_\{X\}\\end\{subarray\}\}\\left\[\\left\|\\mathbb\{E\}\\left\[w\(X\)R\_\{j\}\(P\_\{<j\},P\_\{j\},Y\)\\mid P\\right\]\\right\|\\right\]\.IfRRis not absolutely integrable under\(𝖣,Π\)\(\\mathsf\{D\},\\Pi\), we setErr𝖣Γ\(Π;w\)=\+∞\\operatorname\{Err\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi;w\)=\+\\infty\. A group function is a measurable mapg:𝒳→\[0,1\]g:\\mathcal\{X\}\\to\[0,1\]\. It is binary if it takes values in\{0,1\}\\\{0,1\\\}\. For a finite family𝒢\\mathcal\{G\}of group functions, define
MCErr𝖣Γ\(Π;𝒢\):=maxg∈𝒢Err𝖣Γ\(Π;g\)\.\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi;\\mathcal\{G\}\):=\\max\_\{g\\in\\mathcal\{G\}\}\\operatorname\{Err\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi;g\)\.
### 2\.3Minimax Sample Complexity
We use the minimax formulation for batch multicalibration ofCollinaet al\.\[[9](https://arxiv.org/html/2608.04288#bib.bib43)\], with the distributions restricted to those that areℳ\\mathcal\{M\}\-compatible\. A learner𝖠=\(𝖠n\)n≥1\\mathsf\{A\}=\(\\mathsf\{A\}\_\{n\}\)\_\{n\\geq 1\}receives a finite group family𝒢\\mathcal\{G\}, a sampleS∈\(𝒳×𝒴\)nS\\in\(\\mathcal\{X\}\\times\\mathcal\{Y\}\)^\{n\}, and an independent random seedζ\\zeta, and returns a randomized predictor𝖠n\(𝒢,S,ζ\)\\mathsf\{A\}\_\{n\}\(\\mathcal\{G\},S,\\zeta\)with prediction space𝒫\\mathcal\{P\}\.
LetB:\(0,1\)→ℕB:\(0,1\)\\to\\mathbb\{N\}be a group\-size budget\. For a learner𝖠\\mathsf\{A\}, letn𝖠Γ\(ε;B,ℳ\)n\_\{\\mathsf\{A\}\}^\{\\Gamma\}\(\\varepsilon;B,\\mathcal\{M\}\)be the smallestn≥1n\\geq 1such that, for every standard Borel context space𝒳\\mathcal\{X\}, everyℳ\\mathcal\{M\}\-compatible distribution𝖣\\mathsf\{D\}on𝒳×𝒴\\mathcal\{X\}\\times\\mathcal\{Y\}, and every finite group family𝒢\\mathcal\{G\}with1≤\|𝒢\|≤B\(ε\)1\\leq\|\\mathcal\{G\}\|\\leq B\(\\varepsilon\),
ℙS∼𝖣n,ζ\[MCErr𝖣Γ\(𝖠n\(𝒢,S,ζ\);𝒢\)≤ε\]≥23\.\\mathbb\{P\}\_\{S\\sim\\mathsf\{D\}^\{n\},\\,\\zeta\}\\left\[\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\\bigl\(\\mathsf\{A\}\_\{n\}\(\\mathcal\{G\},S,\\zeta\);\\mathcal\{G\}\\bigr\)\\leq\\varepsilon\\right\]\\geq\\frac\{2\}\{3\}\.If no suchnnexists, setn𝖠Γ\(ε;B,ℳ\)=\+∞n\_\{\\mathsf\{A\}\}^\{\\Gamma\}\(\\varepsilon;B,\\mathcal\{M\}\)=\+\\infty\. The minimax sample complexity is
SCΓ,ℳ\(ε;B\):=inf𝖠n𝖠Γ\(ε;B,ℳ\)\.\\operatorname\{SC\}\_\{\\Gamma,\\mathcal\{M\}\}\(\\varepsilon;B\):=\\inf\_\{\\mathsf\{A\}\}n\_\{\\mathsf\{A\}\}^\{\\Gamma\}\(\\varepsilon;B,\\mathcal\{M\}\)\.For a fixedκ\>0\\kappa\>0, define
Bκ\(ε\):=⌈ε−κ⌉,SCΓ,ℳ\(κ\)\(ε\):=SCΓ,ℳ\(ε;Bκ\)\.B\_\{\\kappa\}\(\\varepsilon\):=\\left\\lceil\\varepsilon^\{\-\\kappa\}\\right\\rceil,\\qquad\\operatorname\{SC\}^\{\(\\kappa\)\}\_\{\\Gamma,\\mathcal\{M\}\}\(\\varepsilon\):=\\operatorname\{SC\}\_\{\\Gamma,\\mathcal\{M\}\}\(\\varepsilon;B\_\{\\kappa\}\)\.Constants suppressed by asymptotic notation may depend on the fixed parameterκ\\kappa\. Thus the context space, distribution, and group family may all vary withε\\varepsilon, subject to the compatibility and size requirements above\.
## 3The Lower Bound
This section proves the lower bound by reducing multicalibration to the task of identifying which distribution in a finite collection generated the observed data\. We construct this collection so that a predictor with small multicalibration error is accurate enough to reveal the true member\. We then use an information\-theoretic argument to show that this identification cannot be accomplished from too few samples\.
The construction must therefore satisfy two requirements\. The same candidate distributions must yield sufficiently different property vectors at the contexts for an accurate predictor to distinguish them, yet remain statistically close enough that distinguishing them from the observed data requires many samples\. We also construct a small family of groups such that small multicalibration error with respect to these groups guarantees the required prediction accuracy\. The regularity assumptions below make these ingredients possible\.
### 3\.1Regularity Assumptions
Sequential conditional identifiability determines the value at which the expected residual at each level vanishes, but it does not quantify how the residual changes when the prediction moves away from that value\. For the lower bound, we require this quantitative control for a family of outcome distributions indexed by property vectors in a fixed rectangleℛ0\\mathcal\{R\}\_\{0\}\. Specifically, for everyv∈ℛ0v\\in\\mathcal\{R\}\_\{0\}, the family contains a distributionμv\\mu\_\{v\}on𝒴\\mathcal\{Y\}satisfyingΓ\(μv\)=v\\Gamma\(\\mu\_\{v\}\)=v\(Condition \(i\) below\)\. We will later use these outcome distributions as conditional laws at different contexts to construct the finite collection of candidate data distributions\.
The remaining conditions govern how the residuals and distributions vary within this family\. When the predictions at preceding levels are correct, the expected residual at the current level must have the same sign as the current prediction error, and its magnitude must be bounded above and below by constant multiples of that error \(Conditions \(ii\) and \(iii\)\)\. Errors in preceding predictions may change a later residual, but only in proportion to those errors \(Condition \(iv\)\)\. Finally, distributions with nearby property vectors must have small Kullback–Leibler divergence, so that the candidate distributions we construct are difficult to distinguish \(Condition \(v\)\)\.
Formally, let
ℛ0=I10×⋯×Ik0⊂int\(𝒫\)\\mathcal\{R\}\_\{0\}=I\_\{1\}^\{0\}\\times\\cdots\\times I\_\{k\}^\{0\}\\subset\\operatorname\{int\}\(\\mathcal\{P\}\)be a nondegenerate compact rectangle\. A point inℛ0\\mathcal\{R\}\_\{0\}is denoted
v=\(v1,…,vk\)\.v=\(v\_\{1\},\\ldots,v\_\{k\}\)\.Write
Ij0=\[ℓj,ℓj\+Wj\],Wj\>0,j∈\[k\]\.I\_\{j\}^\{0\}=\[\\ell\_\{j\},\\ell\_\{j\}\+W\_\{j\}\],\\qquad W\_\{j\}\>0,\\qquad j\\in\[k\]\.Forv∈ℛ0v\\in\\mathcal\{R\}\_\{0\}, writev<j:=\(v1,…,vj−1\)v\_\{<j\}:=\(v\_\{1\},\\ldots,v\_\{j\-1\}\)\. We interpretv<1v\_\{<1\}as the empty prefix\. ThusRj\(v<j,q,μv\)R\_\{j\}\(v\_\{<j\},q,\\mu\_\{v\}\)is the expected level\-jjresidual evaluated with the true prefix\.
###### Assumption 3\.1\(Regularity of the Witness Family\)\.
There exists a family of outcome distributions
\{μv:v∈ℛ0\}⊂ℳ\\\{\\mu\_\{v\}:v\\in\\mathcal\{R\}\_\{0\}\\\}\\subset\\mathcal\{M\}and constantscanti,Cprev,Clip,CKL∈\(0,∞\)c\_\{\\mathrm\{anti\}\},C\_\{\\mathrm\{prev\}\},C\_\{\\mathrm\{lip\}\},C\_\{\\mathrm\{KL\}\}\\in\(0,\\infty\)such that the following hold\.
1. \(i\)For everyv∈ℛ0v\\in\\mathcal\{R\}\_\{0\}, Γ\(μv\)=v\.\\Gamma\(\\mu\_\{v\}\)=v\.
2. \(ii\)For everyv∈ℛ0v\\in\\mathcal\{R\}\_\{0\}, everyj∈\[k\]j\\in\[k\], and everyq∈𝒫jq\\in\\mathcal\{P\}\_\{j\}, sgn\(q−vj\)\(Rj\(v<j,q,μv\)−Rj\(v<j,vj,μv\)\)≥canti\|q−vj\|\.\\operatorname\{sgn\}\(q\-v\_\{j\}\)\\left\(R\_\{j\}\(v\_\{<j\},q,\\mu\_\{v\}\)\-R\_\{j\}\(v\_\{<j\},v\_\{j\},\\mu\_\{v\}\)\\right\)\\geq c\_\{\\mathrm\{anti\}\}\|q\-v\_\{j\}\|\.Here and throughout this section, we use the conventionsgn\(0\)=0\\operatorname\{sgn\}\(0\)=0\.
3. \(iii\)For everyv∈ℛ0v\\in\\mathcal\{R\}\_\{0\}, everyj∈\[k\]j\\in\[k\], and everyq∈𝒫jq\\in\\mathcal\{P\}\_\{j\}, \|Rj\(v<j,q,μv\)−Rj\(v<j,vj,μv\)\|≤Clip\|q−vj\|\.\\left\|R\_\{j\}\(v\_\{<j\},q,\\mu\_\{v\}\)\-R\_\{j\}\(v\_\{<j\},v\_\{j\},\\mu\_\{v\}\)\\right\|\\leq C\_\{\\mathrm\{lip\}\}\|q\-v\_\{j\}\|\.
4. \(iv\)For everyv∈ℛ0v\\in\\mathcal\{R\}\_\{0\}, everyp∈𝒫p\\in\\mathcal\{P\}, and everyj∈\[k\]j\\in\[k\], \|Rj\(p<j,pj,μv\)−Rj\(v<j,pj,μv\)\|≤Cprev∑i<j\|pi−vi\|\.\\left\|R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\_\{v\}\)\-R\_\{j\}\(v\_\{<j\},p\_\{j\},\\mu\_\{v\}\)\\right\|\\leq C\_\{\\mathrm\{prev\}\}\\sum\_\{i<j\}\|p\_\{i\}\-v\_\{i\}\|\.The sum overi<1i<1is interpreted as zero\.
5. \(v\)For everyv,v′∈ℛ0v,v^\{\\prime\}\\in\\mathcal\{R\}\_\{0\}, DKL\(μv∥μv′\)≤CKL‖v−v′‖22\.D\_\{\\mathrm\{KL\}\}\(\\mu\_\{v\}\\,\\\|\\,\\mu\_\{v^\{\\prime\}\}\)\\leq C\_\{\\mathrm\{KL\}\}\\\|v\-v^\{\\prime\}\\\|\_\{2\}^\{2\}\.
Collinaet al\.\[[9](https://arxiv.org/html/2608.04288#bib.bib43)\]use counterparts of Conditions \(i\), \(ii\), and \(v\) in their lower bounds for scalar properties\. Our proof additionally uses \(iii\) to bound the magnitude of a residual by the current prediction error and \(iv\) to bound the effect of errors at earlier levels\.
As a quick sanity check, for mean and mean absolute deviation, the expected level\-jjresidual at the true prefix is exactlyq−vjq\-v\_\{j\}\. Thus, \(ii\) and \(iii\) hold as equalities with constant11, and \(iv\) also holds with constant11\. For quantile and CVaR, \(ii\) and \(iii\) in the quantile coordinate follow directly from lower and upper bounds on the density, while in the CVaR coordinate they hold as equalities with constant11\. Section[5](https://arxiv.org/html/2608.04288#S5)gives the complete verifications\.
### 3\.2Main Result
The conditions above allow us to construct a finite collection of candidate data distributions for every sufficiently smallε\\varepsilon\. Any learner that achieves multicalibration error at mostε\\varepsilonon every distribution in the collection must useΩ~\(ε−\(k\+2\)\)\\widetilde\{\\Omega\}\(\\varepsilon^\{\-\(k\+2\)\}\)samples\. The corresponding group family has only polylogarithmically many members and therefore satisfies\|𝒢ε\|≤ε−κ\|\\mathcal\{G\}\_\{\\varepsilon\}\|\\leq\\varepsilon^\{\-\\kappa\}for every fixedκ\>0\\kappa\>0onceε\\varepsilonis sufficiently small\. The following theorem states this result precisely\.
###### Theorem 3\.2\(Lower Bound for Multicalibration ofkk\-Level Properties\)\.
Suppose Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)holds\. Fix anyκ\>0\\kappa\>0\. Then, for every sufficiently smallε\>0\\varepsilon\>0, there exist
1. \(i\)a finite context space𝒳ε\\mathcal\{X\}\_\{\\varepsilon\};
2. \(ii\)a binary group family𝒢ε\\mathcal\{G\}\_\{\\varepsilon\}on𝒳ε\\mathcal\{X\}\_\{\\varepsilon\}satisfying \|𝒢ε\|≤ε−κ;\|\\mathcal\{G\}\_\{\\varepsilon\}\|\\leq\\varepsilon^\{\-\\kappa\};
3. \(iii\)a finite collection𝔇ε\\mathfrak\{D\}\_\{\\varepsilon\}ofℳ\\mathcal\{M\}\-compatible distributions on𝒳ε×𝒴\\mathcal\{X\}\_\{\\varepsilon\}\\times\\mathcal\{Y\};
such that any possibly randomized learner which, givennni\.i\.d\. samples from any𝖣∈𝔇ε\\mathsf\{D\}\\in\\mathfrak\{D\}\_\{\\varepsilon\}, outputs a predictorΠ\\Pisatisfying
MCErr𝖣Γ\(Π;𝒢ε\)≤ε\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi;\\mathcal\{G\}\_\{\\varepsilon\}\)\\leq\\varepsilonwith probability at least2/32/3, must use
n=Ω\(ε−\(k\+2\)logk\+2\(1/ε\)\)=Ω~\(ε−\(k\+2\)\)\.n=\\Omega\\\!\\left\(\\frac\{\\varepsilon^\{\-\(k\+2\)\}\}\{\\log^\{k\+2\}\(1/\\varepsilon\)\}\\right\)=\\widetilde\{\\Omega\}\\\!\\left\(\\varepsilon^\{\-\(k\+2\)\}\\right\)\.The implicit constants, and the threshold for sufficiently smallε\\varepsilon, depend only onkk,κ\\kappa,ℛ0\\mathcal\{R\}\_\{0\}, and the constants in Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\.
Equivalently, in the notation of Section[2\.3](https://arxiv.org/html/2608.04288#S2.SS3),
SCΓ,ℳ\(κ\)\(ε\)=Ω\(ε−\(k\+2\)logk\+2\(1/ε\)\)=Ω~\(ε−\(k\+2\)\)\.\\operatorname\{SC\}^\{\(\\kappa\)\}\_\{\\Gamma,\\mathcal\{M\}\}\(\\varepsilon\)=\\Omega\\\!\\left\(\\frac\{\\varepsilon^\{\-\(k\+2\)\}\}\{\\log^\{k\+2\}\(1/\\varepsilon\)\}\\right\)=\\widetilde\{\\Omega\}\\\!\\left\(\\varepsilon^\{\-\(k\+2\)\}\\right\)\.
We prove Theorem[3\.2](https://arxiv.org/html/2608.04288#S3.Thmtheorem2)in the remainder of this section, starting with the construction of the candidate distributions\.
### 3\.3The Candidate Distributions
Fix a power of twoQ≥2Q\\geq 2\. We take the context space to be the discrete grid
𝒳Q=\{0,…,Q−1\}k\.\\mathcal\{X\}\_\{Q\}=\\\{0,\\dots,Q\-1\\\}^\{k\}\.We write a context as
a=\(a1,…,ak\)∈𝒳Q\.a=\(a\_\{1\},\\ldots,a\_\{k\}\)\\in\\mathcal\{X\}\_\{Q\}\.
Recall thatIj0=\[ℓj,ℓj\+Wj\]I\_\{j\}^\{0\}=\[\\ell\_\{j\},\\ell\_\{j\}\+W\_\{j\}\]is thejj\-th coordinate interval ofℛ0\\mathcal\{R\}\_\{0\}, with lower endpointℓj\\ell\_\{j\}and widthWj\>0W\_\{j\}\>0\. Forj∈\[k\]j\\in\[k\]andu∈\{0,…,Q−1\}u\\in\\\{0,\\dots,Q\-1\\\}, define
bj\(u\):=ℓj\+Wj3\+uWj3\(Q−1\)\.b\_\{j\}\(u\):=\\ell\_\{j\}\+\\frac\{W\_\{j\}\}\{3\}\+\\frac\{u\\,W\_\{j\}\}\{3\(Q\-1\)\}\.Thusbj\(0\),…,bj\(Q−1\)b\_\{j\}\(0\),\\ldots,b\_\{j\}\(Q\-1\)areQQevenly spaced values in thejj\-th property coordinate\. Fora∈𝒳Qa\\in\\mathcal\{X\}\_\{Q\}, define the unperturbed property vector associated with contextaa
b\(a\):=\(b1\(a1\),…,bk\(ak\)\)\.b\(a\):=\(b\_\{1\}\(a\_\{1\}\),\\ldots,b\_\{k\}\(a\_\{k\}\)\)\.The vectorsb\(a\)b\(a\), asaaranges over𝒳Q\\mathcal\{X\}\_\{Q\}, form a Cartesian grid in the property space\. This grid lies in the central third
∏j=1k\[ℓj\+Wj3,ℓj\+2Wj3\]⊂ℛ0\.\\prod\_\{j=1\}^\{k\}\\left\[\\ell\_\{j\}\+\\frac\{W\_\{j\}\}\{3\},\\,\\ell\_\{j\}\+\\frac\{2W\_\{j\}\}\{3\}\\right\]\\subset\\mathcal\{R\}\_\{0\}\.The distance between consecutive valuesbj\(u\)b\_\{j\}\(u\)isWj/\[3\(Q−1\)\]W\_\{j\}/\[3\(Q\-1\)\]\. LetdQd\_\{Q\}be the smallest of these distances:
dQ:=minj∈\[k\]Wj3\(Q−1\)\.d\_\{Q\}:=\\frac\{\\min\_\{j\\in\[k\]\}W\_\{j\}\}\{3\(Q\-1\)\}\.In particular,
dQ≍Q−1,d\_\{Q\}\\asymp Q^\{\-1\},with constants depending only onℛ0\\mathcal\{R\}\_\{0\}\.
As in the construction used byCollinaet al\.\[[9](https://arxiv.org/html/2608.04288#bib.bib43)\]for vector means, we assign one hidden sign to each context\. At each context, this sign determines the direction in which we perturb the final property coordinate\. Keeping the preceding coordinates at their unperturbed values allows their prediction errors to be controlled successively\. The direction in which the final coordinate is perturbed can be chosen separately at each of theQkQ^\{k\}contexts, yielding exponentially many candidate data distributions\.
We choose the hidden perturbation to be small relative todQd\_\{Q\}\. This keeps every perturbed property vector inside the witness rectangle and ensures that the property vectors assigned to different contexts remain well separated\. Choose a constantcη\>0c\_\{\\eta\}\>0satisfying
cη<120,c\_\{\\eta\}<\\frac\{1\}\{20\},and set
η:=cηdQ\.\\eta:=c\_\{\\eta\}d\_\{Q\}\.Then, for everya∈\{0,…,Q−1\}ka\\in\\\{0,\\dots,Q\-1\\\}^\{k\}and everyθa∈\{−1,\+1\}\\theta\_\{a\}\\in\\\{\-1,\+1\\\},
b\(a\)\+ηθaek∈ℛ0\.b\(a\)\+\\eta\\theta\_\{a\}e\_\{k\}\\in\\mathcal\{R\}\_\{0\}\.
We need to choose exponentially many sign assignments while ensuring that any two disagree on a constant fraction of the contexts\. Let
ΘQ⊆\{−1,\+1\}𝒳Q\\Theta\_\{Q\}\\subseteq\\\{\-1,\+1\\\}^\{\\mathcal\{X\}\_\{Q\}\}satisfy
log\|ΘQ\|≥cpackQk,\|\{a∈𝒳Q:θa≠θa′\}\|≥ρpackQk∀θ≠θ′,\\log\|\\Theta\_\{Q\}\|\\geq c\_\{\\mathrm\{pack\}\}Q^\{k\},\\qquad\\bigl\|\\\{a\\in\\mathcal\{X\}\_\{Q\}:\\theta\_\{a\}\\neq\\theta^\{\\prime\}\_\{a\}\\\}\\bigr\|\\geq\\rho\_\{\\mathrm\{pack\}\}Q^\{k\}\\quad\\forall\\theta\\neq\\theta^\{\\prime\},for universal constantscpack,ρpack\>0c\_\{\\mathrm\{pack\}\},\\rho\_\{\\mathrm\{pack\}\}\>0\. Such a collection exists by the Gilbert packing bound\[[42](https://arxiv.org/html/2608.04288#bib.bib38),[46](https://arxiv.org/html/2608.04288#bib.bib39)\]\.
Each sign assignment specifies the true property vector at every context, and hence one candidate distribution\. Forθ∈ΘQ\\theta\\in\\Theta\_\{Q\}, define the true property vector at contextaaby
vθ\(a\):=b\(a\)\+ηθaek\.v\_\{\\theta\}\(a\):=b\(a\)\+\\eta\\theta\_\{a\}e\_\{k\}\.Thus
vθ,j\(a\)=bj\(aj\)\(j<k\),vθ,k\(a\)=bk\(ak\)\+ηθa\.v\_\{\\theta,j\}\(a\)=b\_\{j\}\(a\_\{j\}\)\\quad\(j<k\),\\qquad v\_\{\\theta,k\}\(a\)=b\_\{k\}\(a\_\{k\}\)\+\\eta\\theta\_\{a\}\.The distribution𝖣θ\\mathsf\{D\}\_\{\\theta\}on𝒳Q×𝒴\\mathcal\{X\}\_\{Q\}\\times\\mathcal\{Y\}is
X∼Unif\(𝒳Q\),Y∣X=a∼μvθ\(a\)\.X\\sim\\operatorname\{Unif\}\(\\mathcal\{X\}\_\{Q\}\),\\qquad Y\\mid X=a\\sim\\mu\_\{v\_\{\\theta\}\(a\)\}\.
We next construct a group family for which small multicalibration error forces a predictor’s outputs to be close tovθ\(a\)v\_\{\\theta\}\(a\)\. This prediction guarantee will allow us to identifyθ\\theta\.
### 3\.4Approximation of Threshold Signs with Walsh Functions
The key relation, formalized later in Lemma[3\.5](https://arxiv.org/html/2608.04288#S3.Thmtheorem5), has the schematic form
\|pj−vθ,j\(a\)\|≲sgn\(pj−vθ,j\(a\)\)Rj\(p<j,pj,μvθ\(a\)\)\+∑i<j\|pi−vθ,i\(a\)\|\.\|p\_\{j\}\-v\_\{\\theta,j\}\(a\)\|\\lesssim\\operatorname\{sgn\}\\\!\\bigl\(p\_\{j\}\-v\_\{\\theta,j\}\(a\)\\bigr\)R\_\{j\}\\\!\\left\(p\_\{<j\},p\_\{j\},\\mu\_\{v\_\{\\theta\}\(a\)\}\\right\)\+\\sum\_\{i<j\}\|p\_\{i\}\-v\_\{\\theta,i\}\(a\)\|\.Thus, once the errors in the preceding coordinates have been controlled, the error in coordinatejjcan be bounded through a residual multiplied by the sign of the current error\.
Multicalibration, on the other hand, controls residuals multiplied by groupsg\(a\)g\(a\)\. We therefore seek to approximate the sign in the inequality above by a linear combination of a common family of functions of the context\. This will allow us to bound the prediction error by the multicalibration error multiplied by a factor\. The sum of the absolute values of the coefficients in the linear combination determines this factor, while the number of functions determines the size of the resulting group family\.
Forj<kj<k, we havevθ,j\(a\)=bj\(aj\)v\_\{\\theta,j\}\(a\)=b\_\{j\}\(a\_\{j\}\), so for a fixed predictionpjp\_\{j\}, the sign
a↦sgn\(pj−bj\(aj\)\)a\\mapsto\\operatorname\{sgn\}\\\!\\bigl\(p\_\{j\}\-b\_\{j\}\(a\_\{j\}\)\\bigr\)is a threshold function ofaja\_\{j\}\. For the final coordinate, the two possible true valuesbk\(ak\)−ηb\_\{k\}\(a\_\{k\}\)\-\\etaandbk\(ak\)\+ηb\_\{k\}\(a\_\{k\}\)\+\\etalie in a small interval aroundbk\(ak\)b\_\{k\}\(a\_\{k\}\)\. Outside a slightly larger interval, the sign of the error is the same for both values\. We use a two\-sided threshold that agrees with this common sign outside the interval and is zero inside it, where predictions will be handled separately\.
Because the group family must be fixed independently of the prediction and ofθ\\theta, we need a common set of functions that can represent all these one\-sided and two\-sided threshold signs, with uniformly controlled coefficients\. The full Walsh basis gives an exact representation with coefficient boundO\(logQ\)O\(\\log Q\), but it would requireO\(Q\)O\(Q\)groups\. By retaining a suitable common subset of the Walsh functions, we can uniformly approximate all the required threshold signs while preserving theO\(logQ\)O\(\\log Q\)coefficient bound\. Only polylogarithmically many Walsh functions are needed\.
BecauseQQis a power of two, identify eachu,ℓ∈\{0,…,Q−1\}u,\\ell\\in\\\{0,\\ldots,Q\-1\\\}with its length\-log2Q\\log\_\{2\}Qbinary expansion\. Define the Walsh function
ψℓ\(u\):=\(−1\)⟨ℓ,u⟩𝔽2\.\\psi\_\{\\ell\}\(u\):=\(\-1\)^\{\\langle\\ell,u\\rangle\_\{\\mathbb\{F\}\_\{2\}\}\}\.Here
⟨ℓ,u⟩𝔽2:=∑i=1log2Qℓiui\(mod2\)\.\\langle\\ell,u\\rangle\_\{\\mathbb\{F\}\_\{2\}\}:=\\sum\_\{i=1\}^\{\\log\_\{2\}Q\}\\ell\_\{i\}u\_\{i\}\\pmod\{2\}\.The Walsh functions form an orthonormal basis for real\-valued functions on\{0,…,Q−1\}\\\{0,\\ldots,Q\-1\\\}\.
Forσ∈\{−1,\+1\}\\sigma\\in\\\{\-1,\+1\\\}andt∈ℝt\\in\\mathbb\{R\}, define the one\-sided function
sσ,t\(u\):=σsgn\(t−u\)\.s\_\{\\sigma,t\}\(u\):=\\sigma\\operatorname\{sgn\}\(t\-u\)\.Forσ∈\{−1,\+1\}\\sigma\\in\\\{\-1,\+1\\\}andt−≤t\+t\_\{\-\}\\leq t\_\{\+\}, define the two\-sided function
sσ,t−,t\+\(u\):=σ2\(sgn\(t−−u\)\+sgn\(t\+−u\)\)\.s\_\{\\sigma,t\_\{\-\},t\_\{\+\}\}\(u\):=\\frac\{\\sigma\}\{2\}\\bigl\(\\operatorname\{sgn\}\(t\_\{\-\}\-u\)\+\\operatorname\{sgn\}\(t\_\{\+\}\-u\)\\bigr\)\.Let𝒯Q\\mathcal\{T\}\_\{Q\}be the class of all these functions\. A one\-sided function changes sign at one threshold\. A two\-sided function is zero between two thresholds and has opposite signs beyond them, with the boundary values determined bysgn\(0\)=0\\operatorname\{sgn\}\(0\)=0\. A two\-sided function tests whether the final prediction lies outside a prescribed interval\.
Collinaet al\.\[[8](https://arxiv.org/html/2608.04288#bib.bib26)\]show that one common polylogarithmic subset of the Walsh basis approximates every one\-sided discrete threshold sign\. Their coefficients satisfy the uniformO\(logQ\)O\(\\log Q\)bound in Lemma[3\.3](https://arxiv.org/html/2608.04288#S3.Thmtheorem3)\.
###### Lemma 3\.3\(Walsh Approximation of Threshold Signs\)\.
There is a universal constantCWC\_\{\\mathrm\{W\}\}such that the following holds\. LetQ≥2Q\\geq 2be a power of two\. For every fixed accuracyα∈\(0,1\)\\alpha\\in\(0,1\), there exists a set
𝒮α⊆\{1,…,Q−1\}\\mathcal\{S\}\_\{\\alpha\}\\subseteq\\\{1,\\dots,Q\-1\\\}of indices of Walsh functions satisfying
\|𝒮α\|≤CWα−2log3\(Q\+1\)\|\\mathcal\{S\}\_\{\\alpha\}\|\\leq C\_\{\\mathrm\{W\}\}\\alpha^\{\-2\}\\log^\{3\}\(Q\+1\)and coefficient functions
β0:𝒯Q→ℝ,βℓ:𝒯Q→ℝ\(ℓ∈𝒮α\),\\beta\_\{0\}:\\mathcal\{T\}\_\{Q\}\\to\\mathbb\{R\},\\qquad\\beta\_\{\\ell\}:\\mathcal\{T\}\_\{Q\}\\to\\mathbb\{R\}\\quad\(\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\),such that, for everys∈𝒯Qs\\in\\mathcal\{T\}\_\{Q\},
s^\(u\)=β0\(s\)\+∑ℓ∈𝒮αβℓ\(s\)ψℓ\(u\)\\widehat\{s\}\(u\)=\\beta\_\{0\}\(s\)\+\\sum\_\{\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\}\\beta\_\{\\ell\}\(s\)\\psi\_\{\\ell\}\(u\)is a uniform approximation:
‖s−s^‖∞≤α\\\|s\-\\widehat\{s\}\\\|\_\{\\infty\}\\leq\\alphaand the coefficients satisfy
sups∈𝒯Q\|β0\(s\)\|\+∑ℓ∈𝒮αsups∈𝒯Q\|βℓ\(s\)\|≤CWlog\(Q\+1\)\.\\sup\_\{s\\in\\mathcal\{T\}\_\{Q\}\}\|\\beta\_\{0\}\(s\)\|\+\\sum\_\{\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\}\\sup\_\{s\\in\\mathcal\{T\}\_\{Q\}\}\|\\beta\_\{\\ell\}\(s\)\|\\leq C\_\{\\mathrm\{W\}\}\\log\(Q\+1\)\.
###### Proof\.
For eachr∈\{0,…,Q\}r\\in\\\{0,\\ldots,Q\\\}, define
fr:\{0,…,Q−1\}→\{−1,\+1\}f\_\{r\}:\\\{0,\\ldots,Q\-1\\\}\\to\\\{\-1,\+1\\\}by
fr\(u\):=\{\+1,u<r,−1,u≥r\.f\_\{r\}\(u\):=\\begin\{cases\}\+1,&u<r,\\\\ \-1,&u\\geq r\.\\end\{cases\}Collinaet al\.\[[8](https://arxiv.org/html/2608.04288#bib.bib26)\]construct a common set
𝒮α⊆\{1,…,Q−1\},\|𝒮α\|≤Cα−2log3\(Q\+1\),\\mathcal\{S\}\_\{\\alpha\}\\subseteq\\\{1,\\ldots,Q\-1\\\},\\qquad\|\\mathcal\{S\}\_\{\\alpha\}\|\\leq C\\alpha^\{\-2\}\\log^\{3\}\(Q\+1\),together with coefficientsa0\(r\)a\_\{0\}\(r\)andcℓ\(r\)c\_\{\\ell\}\(r\)for allr∈\{0,…,Q\}r\\in\\\{0,\\ldots,Q\\\}\. The corresponding approximations
f~r\(u\):=a0\(r\)\+∑ℓ∈𝒮αcℓ\(r\)ψℓ\(u\)\\widetilde\{f\}\_\{r\}\(u\):=a\_\{0\}\(r\)\+\\sum\_\{\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\}c\_\{\\ell\}\(r\)\\psi\_\{\\ell\}\(u\)satisfy
maxr∈\{0,…,Q\}‖fr−f~r‖∞≤α\\max\_\{r\\in\\\{0,\\ldots,Q\\\}\}\\\|f\_\{r\}\-\\widetilde\{f\}\_\{r\}\\\|\_\{\\infty\}\\leq\\alphaand
maxr∈\{0,…,Q\}\|a0\(r\)\|\+∑ℓ∈𝒮αmaxr∈\{0,…,Q\}\|cℓ\(r\)\|≤Clog\(Q\+1\)\\max\_\{r\\in\\\{0,\\ldots,Q\\\}\}\|a\_\{0\}\(r\)\|\+\\sum\_\{\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\}\\max\_\{r\\in\\\{0,\\ldots,Q\\\}\}\|c\_\{\\ell\}\(r\)\|\\leq C\\log\(Q\+1\)for a universal constantCC\.
We use these coefficients to defineβ0\\beta\_\{0\}andβℓ\\beta\_\{\\ell\}on𝒯Q\\mathcal\{T\}\_\{Q\}\. For each functionssthat has a one\-sided representation, fix one such representations=sσ,ts=s\_\{\\sigma,t\}\. Ift∉\{0,…,Q−1\}t\\notin\\\{0,\\ldots,Q\-1\\\}, the restriction ofu↦sgn\(t−u\)u\\mapsto\\operatorname\{sgn\}\(t\-u\)to\{0,…,Q−1\}\\\{0,\\ldots,Q\-1\\\}equals a uniquefrf\_\{r\}\. Define
β0\(s\):=σa0\(r\),βℓ\(s\):=σcℓ\(r\)\(ℓ∈𝒮α\)\.\\beta\_\{0\}\(s\):=\\sigma a\_\{0\}\(r\),\\qquad\\beta\_\{\\ell\}\(s\):=\\sigma c\_\{\\ell\}\(r\)\\quad\(\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\)\.Ift∈\{0,…,Q−1\}t\\in\\\{0,\\ldots,Q\-1\\\}, then
sgn\(t−u\)=ft\(u\)\+ft\+1\(u\)2\.\\operatorname\{sgn\}\(t\-u\)=\\frac\{f\_\{t\}\(u\)\+f\_\{t\+1\}\(u\)\}\{2\}\.In this case, define
β0\(s\):=σ2\(a0\(t\)\+a0\(t\+1\)\)\\beta\_\{0\}\(s\):=\\frac\{\\sigma\}\{2\}\\bigl\(a\_\{0\}\(t\)\+a\_\{0\}\(t\+1\)\\bigr\)and
βℓ\(s\):=σ2\(cℓ\(t\)\+cℓ\(t\+1\)\)\(ℓ∈𝒮α\)\.\\beta\_\{\\ell\}\(s\):=\\frac\{\\sigma\}\{2\}\\bigl\(c\_\{\\ell\}\(t\)\+c\_\{\\ell\}\(t\+1\)\\bigr\)\\quad\(\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\)\.
For each functions∈𝒯Qs\\in\\mathcal\{T\}\_\{Q\}not already covered, fix a two\-sided representations=sσ,t−,t\+s=s\_\{\\sigma,t\_\{\-\},t\_\{\+\}\}\. By definition,
sσ,t−,t\+\(u\)=sσ,t−\(u\)\+sσ,t\+\(u\)2\.s\_\{\\sigma,t\_\{\-\},t\_\{\+\}\}\(u\)=\\frac\{s\_\{\\sigma,t\_\{\-\}\}\(u\)\+s\_\{\\sigma,t\_\{\+\}\}\(u\)\}\{2\}\.Using the coefficients already defined for these two one\-sided functions, set
β0\(s\):=β0\(sσ,t−\)\+β0\(sσ,t\+\)2\\beta\_\{0\}\(s\):=\\frac\{\\beta\_\{0\}\(s\_\{\\sigma,t\_\{\-\}\}\)\+\\beta\_\{0\}\(s\_\{\\sigma,t\_\{\+\}\}\)\}\{2\}and
βℓ\(s\):=βℓ\(sσ,t−\)\+βℓ\(sσ,t\+\)2\(ℓ∈𝒮α\)\.\\beta\_\{\\ell\}\(s\):=\\frac\{\\beta\_\{\\ell\}\(s\_\{\\sigma,t\_\{\-\}\}\)\+\\beta\_\{\\ell\}\(s\_\{\\sigma,t\_\{\+\}\}\)\}\{2\}\\quad\(\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\)\.
In each case,s^\\widehat\{s\}is obtained by applying the same multiplication and averaging operations to the corresponding approximationsf~r\\widetilde\{f\}\_\{r\}\. The triangle inequality therefore gives
‖s−s^‖∞≤α\.\\\|s\-\\widehat\{s\}\\\|\_\{\\infty\}\\leq\\alpha\.Every coefficientβ0\(s\)\\beta\_\{0\}\(s\)is a signed average of coefficientsa0\(r\)a\_\{0\}\(r\), and everyβℓ\(s\)\\beta\_\{\\ell\}\(s\)is the corresponding signed average of coefficientscℓ\(r\)c\_\{\\ell\}\(r\)\. Consequently,
sups∈𝒯Q\|β0\(s\)\|≤maxr∈\{0,…,Q\}\|a0\(r\)\|\\sup\_\{s\\in\\mathcal\{T\}\_\{Q\}\}\|\\beta\_\{0\}\(s\)\|\\leq\\max\_\{r\\in\\\{0,\\ldots,Q\\\}\}\|a\_\{0\}\(r\)\|and, for everyℓ∈𝒮α\\ell\\in\\mathcal\{S\}\_\{\\alpha\},
sups∈𝒯Q\|βℓ\(s\)\|≤maxr∈\{0,…,Q\}\|cℓ\(r\)\|\.\\sup\_\{s\\in\\mathcal\{T\}\_\{Q\}\}\|\\beta\_\{\\ell\}\(s\)\|\\leq\\max\_\{r\\in\\\{0,\\ldots,Q\\\}\}\|c\_\{\\ell\}\(r\)\|\.The required coefficient bound now follows from the corresponding bound ona0\(r\)a\_\{0\}\(r\)andcℓ\(r\)c\_\{\\ell\}\(r\), after enlarging the universal constant if necessary\. ∎
### 3\.5The Group Family
We now use the Walsh approximation above to construct a family of binary groups\.
Fix a sufficiently small constantα∈\(0,1\)\\alpha\\in\(0,1\), and let𝒮α\\mathcal\{S\}\_\{\\alpha\}be the set supplied by Lemma[3\.3](https://arxiv.org/html/2608.04288#S3.Thmtheorem3)\. For each coordinatej∈\[k\]j\\in\[k\]and eachℓ∈𝒮α\\ell\\in\\mathcal\{S\}\_\{\\alpha\}, define the signed Walsh weight
wj,ℓ\(a\):=ψℓ\(aj\)\.w\_\{j,\\ell\}\(a\):=\\psi\_\{\\ell\}\(a\_\{j\}\)\.Convert it into two binary groups
gj,ℓ\+\(a\):=1\+wj,ℓ\(a\)2,gj,ℓ−\(a\):=1−wj,ℓ\(a\)2\.g^\{\+\}\_\{j,\\ell\}\(a\):=\\frac\{1\+w\_\{j,\\ell\}\(a\)\}\{2\},\\qquad g^\{\-\}\_\{j,\\ell\}\(a\):=\\frac\{1\-w\_\{j,\\ell\}\(a\)\}\{2\}\.Let
gall≡1,g\_\{\\mathrm\{all\}\}\\equiv 1,and define
𝒢Q:=\{gall\}∪\{gj,ℓ\+,gj,ℓ−:j∈\[k\],ℓ∈𝒮α\}\.\\mathcal\{G\}\_\{Q\}:=\\\{g\_\{\\mathrm\{all\}\}\\\}\\cup\\\{g^\{\+\}\_\{j,\\ell\},g^\{\-\}\_\{j,\\ell\}:\\ j\\in\[k\],\\ \\ell\\in\\mathcal\{S\}\_\{\\alpha\}\\\}\.Then𝒢Q\\mathcal\{G\}\_\{Q\}is binary\-valued and satisfies
\|𝒢Q\|≤Cα,klog3\(Q\+1\)\.\|\\mathcal\{G\}\_\{Q\}\|\\leq C\_\{\\alpha,k\}\\log^\{3\}\(Q\+1\)\.
The analysis uses the signed Walsh weightswj,ℓw\_\{j,\\ell\}, although the group family itself is binary\-valued\. The following elementary lemma passes between these two forms\.
###### Lemma 3\.4\(Signed Weights from Binary Groups\)\.
For everyj∈\[k\]j\\in\[k\]andℓ∈𝒮α\\ell\\in\\mathcal\{S\}\_\{\\alpha\}, every distribution𝖣\\mathsf\{D\}on𝒳Q×𝒴\\mathcal\{X\}\_\{Q\}\\times\\mathcal\{Y\}, and every predictorΠ\\Pi,
Err𝖣Γ\(Π;wj,ℓ\)≤2MCErr𝖣Γ\(Π;𝒢Q\)\.\\operatorname\{Err\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi;w\_\{j,\\ell\}\)\\leq 2\\,\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\.
###### Proof\.
Sincewj,ℓ=gj,ℓ\+−gj,ℓ−w\_\{j,\\ell\}=g^\{\+\}\_\{j,\\ell\}\-g^\{\-\}\_\{j,\\ell\}, its vector signed measure is the difference of the vector signed measures forgj,ℓ\+g^\{\+\}\_\{j,\\ell\}andgj,ℓ−g^\{\-\}\_\{j,\\ell\}\. Both groups belong to𝒢Q\\mathcal\{G\}\_\{Q\}, and coordinatewise total variation is subadditive\. ∎
### 3\.6From Multicalibration to Prediction Accuracy
Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(ii\) gives a lower bound when the preceding predictions equal their true values, but a predictor need not use the true prefix\. The following lemma relates the residual at an arbitrary prediction to the prediction error at the current level and the errors at preceding levels\.
###### Lemma 3\.5\(Relating Residuals to Prediction Errors\)\.
Under Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1), for everyv∈ℛ0v\\in\\mathcal\{R\}\_\{0\}, everyp∈𝒫p\\in\\mathcal\{P\}, and everyj∈\[k\]j\\in\[k\],
sgn\(pj−vj\)\(Rj\(p<j,pj,μv\)−Rj\(v<j,vj,μv\)\)≥canti\|pj−vj\|−Cprev∑i<j\|pi−vi\|,\\operatorname\{sgn\}\(p\_\{j\}\-v\_\{j\}\)\\left\(R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\_\{v\}\)\-R\_\{j\}\(v\_\{<j\},v\_\{j\},\\mu\_\{v\}\)\\right\)\\geq c\_\{\\mathrm\{anti\}\}\|p\_\{j\}\-v\_\{j\}\|\-C\_\{\\mathrm\{prev\}\}\\sum\_\{i<j\}\|p\_\{i\}\-v\_\{i\}\|,and
\|Rj\(p<j,pj,μv\)−Rj\(v<j,vj,μv\)\|≤Clip\|pj−vj\|\+Cprev∑i<j\|pi−vi\|\.\\left\|R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\_\{v\}\)\-R\_\{j\}\(v\_\{<j\},v\_\{j\},\\mu\_\{v\}\)\\right\|\\leq C\_\{\\mathrm\{lip\}\}\|p\_\{j\}\-v\_\{j\}\|\+C\_\{\\mathrm\{prev\}\}\\sum\_\{i<j\}\|p\_\{i\}\-v\_\{i\}\|\.
###### Proof\.
BecauseΓ\(μv\)=v\\Gamma\(\\mu\_\{v\}\)=vby Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(i\),Rj\(v<j,vj,μv\)=0R\_\{j\}\(v\_\{<j\},v\_\{j\},\\mu\_\{v\}\)=0\. Thus it suffices to prove the same two inequalities withRj\(p<j,pj,μv\)R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\_\{v\}\)in place ofRj\(p<j,pj,μv\)−Rj\(v<j,vj,μv\)R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\_\{v\}\)\-R\_\{j\}\(v\_\{<j\},v\_\{j\},\\mu\_\{v\}\)\. For the lower bound,
sgn\(pj−vj\)Rj\(p<j,pj,μv\)≥sgn\(pj−vj\)Rj\(v<j,pj,μv\)−\|Rj\(p<j,pj,μv\)−Rj\(v<j,pj,μv\)\|\.\\operatorname\{sgn\}\(p\_\{j\}\-v\_\{j\}\)R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\_\{v\}\)\\geq\\operatorname\{sgn\}\(p\_\{j\}\-v\_\{j\}\)R\_\{j\}\(v\_\{<j\},p\_\{j\},\\mu\_\{v\}\)\-\\left\|R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\_\{v\}\)\-R\_\{j\}\(v\_\{<j\},p\_\{j\},\\mu\_\{v\}\)\\right\|\.Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(ii\) controls the first term, and condition \(iv\) controls the second\. For the upper bound, use the triangle inequality andRj\(v<j,vj,μv\)=0R\_\{j\}\(v\_\{<j\},v\_\{j\},\\mu\_\{v\}\)=0, followed by the bounds in conditions \(iii\) and \(iv\)\. ∎
The preceding lemma relates prediction error to a residual multiplied by the sign of the prediction error\. We approximate the threshold signs that arise in this argument by linear combinations of the fixed Walsh functions that define𝒢Q\\mathcal\{G\}\_\{Q\}\. Lemma[3\.4](https://arxiv.org/html/2608.04288#S3.Thmtheorem4)shows thatΓ\\Gamma\-ECE controls residuals weighted by each of these Walsh functions\. The next lemma shows thatΓ\\Gamma\-ECE also controls a residual weighted by such a linear combination\.
###### Lemma 3\.6\(Multicalibration Controls Weighted Residuals\)\.
Fixθ\\theta, a predictorΠ\\Pi, and a levelj∈\[k\]j\\in\[k\]\. Suppose a functionh:𝒫×𝒳Q→ℝh:\\mathcal\{P\}\\times\\mathcal\{X\}\_\{Q\}\\to\\mathbb\{R\}has the form
h\(p,a\)=β0\(p\)\+∑ℓ∈𝒮αβℓ\(p\)wi,ℓ\(a\),h\(p,a\)=\\beta\_\{0\}\(p\)\+\\sum\_\{\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\}\\beta\_\{\\ell\}\(p\)w\_\{i,\\ell\}\(a\),for some context coordinatei∈\[k\]i\\in\[k\]and measurable coefficient functions
β0:𝒫→ℝ,βℓ:𝒫→ℝ\(ℓ∈𝒮α\),\\beta\_\{0\}:\\mathcal\{P\}\\to\\mathbb\{R\},\\qquad\\beta\_\{\\ell\}:\\mathcal\{P\}\\to\\mathbb\{R\}\\quad\(\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\),satisfying
supp∈𝒫\|β0\(p\)\|\+∑ℓ∈𝒮αsupp∈𝒫\|βℓ\(p\)\|≤Ccoef\.\\sup\_\{p\\in\\mathcal\{P\}\}\|\\beta\_\{0\}\(p\)\|\+\\sum\_\{\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\}\\sup\_\{p\\in\\mathcal\{P\}\}\|\\beta\_\{\\ell\}\(p\)\|\\leq C\_\{\\mathrm\{coef\}\}\.Then
\|1Qk∑a∈𝒳Q∫𝒫h\(p,a\)Rj\(p<j,pj,μvθ\(a\)\)Πa\(dp\)\|≤2CcoefMCErr𝖣θΓ\(Π;𝒢Q\)\.\\left\|\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\\in\\mathcal\{X\}\_\{Q\}\}\\int\_\{\\mathcal\{P\}\}h\(p,a\)\\,R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\_\{v\_\{\\theta\}\(a\)\}\)\\,\\Pi\_\{a\}\(dp\)\\right\|\\leq 2C\_\{\\mathrm\{coef\}\}\\,\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\.
###### Proof\.
For a signed weightww, thejj\-th coordinate signed measure under\(𝖣θ,Π\)\(\\mathsf\{D\}\_\{\\theta\},\\Pi\)is
νw,j𝖣θ,Π\(A\):=1Qk∑a∈𝒳Qw\(a\)∫ARj\(p<j,pj,μvθ\(a\)\)Πa\(dp\),A⊆𝒫\.\\nu^\{\\mathsf\{D\}\_\{\\theta\},\\Pi\}\_\{w,j\}\(A\):=\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\\in\\mathcal\{X\}\_\{Q\}\}w\(a\)\\int\_\{A\}R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\_\{v\_\{\\theta\}\(a\)\}\)\\,\\Pi\_\{a\}\(dp\),\\qquad A\\subseteq\\mathcal\{P\}\.Therefore
1Qk∑a∫h\(p,a\)Rj\(p<j,pj,μvθ\(a\)\)Πa\(dp\)=∫β0\(p\)νgall,j𝖣θ,Π\(dp\)\+∑ℓ∈𝒮α∫βℓ\(p\)νwi,ℓ,j𝖣θ,Π\(dp\)\.\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\int h\(p,a\)R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\_\{v\_\{\\theta\}\(a\)\}\)\\Pi\_\{a\}\(dp\)=\\int\\beta\_\{0\}\(p\)\\,\\nu^\{\\mathsf\{D\}\_\{\\theta\},\\Pi\}\_\{g\_\{\\mathrm\{all\}\},j\}\(dp\)\+\\sum\_\{\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\}\\int\\beta\_\{\\ell\}\(p\)\\,\\nu^\{\\mathsf\{D\}\_\{\\theta\},\\Pi\}\_\{w\_\{i,\\ell\},j\}\(dp\)\.By total variation,
\|∫β0𝑑νgall,j𝖣θ,Π\|≤supp\|β0\(p\)\|\|νgall,j𝖣θ,Π\|\(𝒫\)≤supp\|β0\(p\)\|MCErr𝖣θΓ\(Π;𝒢Q\)\.\\left\|\\int\\beta\_\{0\}\\,d\\nu^\{\\mathsf\{D\}\_\{\\theta\},\\Pi\}\_\{g\_\{\\mathrm\{all\}\},j\}\\right\|\\leq\\sup\_\{p\}\|\\beta\_\{0\}\(p\)\|\\,\|\\nu^\{\\mathsf\{D\}\_\{\\theta\},\\Pi\}\_\{g\_\{\\mathrm\{all\}\},j\}\|\(\\mathcal\{P\}\)\\leq\\sup\_\{p\}\|\\beta\_\{0\}\(p\)\|\\,\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\.For each signed Walsh weightwi,ℓ=gi,ℓ\+−gi,ℓ−w\_\{i,\\ell\}=g^\{\+\}\_\{i,\\ell\}\-g^\{\-\}\_\{i,\\ell\}, Lemma[3\.4](https://arxiv.org/html/2608.04288#S3.Thmtheorem4)gives
\|νwi,ℓ,j𝖣θ,Π\|\(𝒫\)≤Err𝖣θΓ\(Π;wi,ℓ\)≤2MCErr𝖣θΓ\(Π;𝒢Q\)\.\|\\nu^\{\\mathsf\{D\}\_\{\\theta\},\\Pi\}\_\{w\_\{i,\\ell\},j\}\|\(\\mathcal\{P\}\)\\leq\\operatorname\{Err\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;w\_\{i,\\ell\}\)\\leq 2\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\.Therefore
\|1Qk∑a∫h\(p,a\)Rj\(p<j,pj,μvθ\(a\)\)Πa\(dp\)\|\\displaystyle\\left\|\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\int h\(p,a\)R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\_\{v\_\{\\theta\}\(a\)\}\)\\Pi\_\{a\}\(dp\)\\right\|≤\(supp\|β0\(p\)\|\+2∑ℓ∈𝒮αsupp\|βℓ\(p\)\|\)MCErr𝖣θΓ\(Π;𝒢Q\)\\displaystyle\\leq\\left\(\\sup\_\{p\}\|\\beta\_\{0\}\(p\)\|\+2\\sum\_\{\\ell\\in\\mathcal\{S\}\_\{\\alpha\}\}\\sup\_\{p\}\|\\beta\_\{\\ell\}\(p\)\|\\right\)\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)≤2CcoefMCErr𝖣θΓ\(Π;𝒢Q\)\.\\displaystyle\\leq 2C\_\{\\mathrm\{coef\}\}\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\.∎
We now combine the residual inequalities with the group family constructed above\. To establish their lower bound for multicalibration of vector\-valued means,Collinaet al\.\[[9](https://arxiv.org/html/2608.04288#bib.bib43)\]divide the prediction error into two parts\. Coordinatewise signs control the contribution from predictions that are far from the unperturbed mean for their context, while the all\-ones group controls the contribution from predictions that are nearby\. Splitting predictions into these two cases is not enough here, because an error in an earlier coordinate can alter every later residual\. We first use threshold signs to control coordinates1,…,k−11,\\ldots,k\-1successively\. With the prefix errors controlled, two\-sided threshold signs bound the error in the final coordinate outside the correct box, and the all\-ones group handles predictions inside it\. The resulting bound controls the average prediction error byΓ\\Gamma\-ECE\.
For a predictorΠ=\(Πa\)a∈𝒳Q\\Pi=\(\\Pi\_\{a\}\)\_\{a\\in\\mathcal\{X\}\_\{Q\}\}, define its average error in predicting the property underθ\\thetaby
PredErrθ\(Π\):=1Qk∑a∈𝒳Q∫𝒫‖p−vθ\(a\)‖1Πa\(dp\)\.\\operatorname\{PredErr\}\_\{\\theta\}\(\\Pi\):=\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\\in\\mathcal\{X\}\_\{Q\}\}\\int\_\{\\mathcal\{P\}\}\\\|p\-v\_\{\\theta\}\(a\)\\\|\_\{1\}\\,\\Pi\_\{a\}\(dp\)\.Also define the coordinate prediction errors
PredErrθ,j\(Π\):=1Qk∑a∈𝒳Q∫𝒫\|pj−vθ,j\(a\)\|Πa\(dp\),j∈\[k\]\.\\operatorname\{PredErr\}\_\{\\theta,j\}\(\\Pi\):=\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\\in\\mathcal\{X\}\_\{Q\}\}\\int\_\{\\mathcal\{P\}\}\|p\_\{j\}\-v\_\{\\theta,j\}\(a\)\|\\,\\Pi\_\{a\}\(dp\),\\qquad j\\in\[k\]\.Thus
PredErrθ\(Π\)=∑j=1kPredErrθ,j\(Π\)\.\\operatorname\{PredErr\}\_\{\\theta\}\(\\Pi\)=\\sum\_\{j=1\}^\{k\}\\operatorname\{PredErr\}\_\{\\theta,j\}\(\\Pi\)\.
###### Proposition 3\.7\(Γ\\Gamma\-ECE Controls Prediction Error\)\.
Under Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1), there exists a constantCpred<∞C\_\{\\mathrm\{pred\}\}<\\infty, depending only onkk,ℛ0\\mathcal\{R\}\_\{0\}, and the constants in Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1), such that for every sufficiently large power of twoQQ, everyθ∈ΘQ\\theta\\in\\Theta\_\{Q\}, and every randomized predictorΠ\\Pi,
PredErrθ\(Π\)≤Cpredlog\(Q\+1\)MCErr𝖣θΓ\(Π;𝒢Q\)\.\\operatorname\{PredErr\}\_\{\\theta\}\(\\Pi\)\\leq C\_\{\\mathrm\{pred\}\}\\log\(Q\+1\)\\,\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\.
###### Proof\.
All constants in this proof may depend onkk,ℛ0\\mathcal\{R\}\_\{0\}, and the constants in Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1), but not onQQ,θ\\theta, orΠ\\Pi\. The accuracyα\\alphain the Walsh approximation was fixed small enough that the termsCαPredErrθ,j\(Π\)C\\alpha\\operatorname\{PredErr\}\_\{\\theta,j\}\(\\Pi\)that arise can be absorbed into the left\-hand side\. For brevity, write
ei\(a,p\):=\|pi−vθ,i\(a\)\|,𝖱a,i\(p\):=Ri\(p<i,pi,μvθ\(a\)\)\.e\_\{i\}\(a,p\):=\|p\_\{i\}\-v\_\{\\theta,i\}\(a\)\|,\\qquad\\mathsf\{R\}\_\{a,i\}\(p\):=R\_\{i\}\(p\_\{<i\},p\_\{i\},\\mu\_\{v\_\{\\theta\}\(a\)\}\)\.BecauseΓ\(μvθ\(a\)\)=vθ\(a\)\\Gamma\(\\mu\_\{v\_\{\\theta\}\(a\)\}\)=v\_\{\\theta\}\(a\)by Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(i\),𝖱a,i\(vθ\(a\)\)=0\\mathsf\{R\}\_\{a,i\}\(v\_\{\\theta\}\(a\)\)=0\.
##### Controlling the Firstk−1k\-1Coordinates\.
Forj<kj<kandpj∈𝒫jp\_\{j\}\\in\\mathcal\{P\}\_\{j\}, define
spj\(j\)\(a\):=sgn\(pj−bj\(aj\)\)\.s^\{\(j\)\}\_\{p\_\{j\}\}\(a\):=\\operatorname\{sgn\}\(p\_\{j\}\-b\_\{j\}\(a\_\{j\}\)\)\.Sincevθ,j\(a\)=bj\(aj\)v\_\{\\theta,j\}\(a\)=b\_\{j\}\(a\_\{j\}\)forj<kj<k, this issgn\(pj−vθ,j\(a\)\)\\operatorname\{sgn\}\(p\_\{j\}\-v\_\{\\theta,j\}\(a\)\)\. As a function ofaja\_\{j\}, it is one of the one\-sided threshold signs in𝒯Q\\mathcal\{T\}\_\{Q\}\. Lets^pj\(j\)\\widehat\{s\}^\{\(j\)\}\_\{p\_\{j\}\}be its Walsh approximation from Lemma[3\.3](https://arxiv.org/html/2608.04288#S3.Thmtheorem3)\. Lemma[3\.5](https://arxiv.org/html/2608.04288#S3.Thmtheorem5)gives
spj\(j\)\(a\)𝖱a,j\(p\)≥cantiej\(a,p\)−Cprev∑i<jei\(a,p\),s^\{\(j\)\}\_\{p\_\{j\}\}\(a\)\\mathsf\{R\}\_\{a,j\}\(p\)\\geq c\_\{\\mathrm\{anti\}\}e\_\{j\}\(a,p\)\-C\_\{\\mathrm\{prev\}\}\\sum\_\{i<j\}e\_\{i\}\(a,p\),and also
\|𝖱a,j\(p\)\|≤Cej\(a,p\)\+C∑i<jei\(a,p\)\.\|\\mathsf\{R\}\_\{a,j\}\(p\)\|\\leq Ce\_\{j\}\(a,p\)\+C\\sum\_\{i<j\}e\_\{i\}\(a,p\)\.Therefore, using‖s^pj\(j\)−spj\(j\)‖∞≤α\\\|\\widehat\{s\}^\{\(j\)\}\_\{p\_\{j\}\}\-s^\{\(j\)\}\_\{p\_\{j\}\}\\\|\_\{\\infty\}\\leq\\alpha,
s^pj\(j\)\(a\)𝖱a,j\(p\)≥cej\(a,p\)−C∑i<jei\(a,p\)\\widehat\{s\}^\{\(j\)\}\_\{p\_\{j\}\}\(a\)\\mathsf\{R\}\_\{a,j\}\(p\)\\geq ce\_\{j\}\(a,p\)\-C\\sum\_\{i<j\}e\_\{i\}\(a,p\)after fixingα\\alphasufficiently small and decreasingc\>0c\>0\. Integrating over\(a,p\)\(a,p\)and rearranging gives
cPredErrθ,j\(Π\)≤\|1Qk∑a∫s^pj\(j\)\(a\)𝖱a,j\(p\)Πa\(dp\)\|\+C∑i<jPredErrθ,i\(Π\)\.c\\operatorname\{PredErr\}\_\{\\theta,j\}\(\\Pi\)\\leq\\left\|\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\int\\widehat\{s\}^\{\(j\)\}\_\{p\_\{j\}\}\(a\)\\mathsf\{R\}\_\{a,j\}\(p\)\\,\\Pi\_\{a\}\(dp\)\\right\|\+C\\sum\_\{i<j\}\\operatorname\{PredErr\}\_\{\\theta,i\}\(\\Pi\)\.The Walsh coefficients ofs^pj\(j\)\(a\)\\widehat\{s\}^\{\(j\)\}\_\{p\_\{j\}\}\(a\), viewed as a function ofaa, are step functions ofpjp\_\{j\}and hence measurable\. They also satisfy the uniform boundClog\(Q\+1\)C\\log\(Q\+1\)\. Lemma[3\.6](https://arxiv.org/html/2608.04288#S3.Thmtheorem6)therefore bounds the term inside the absolute value by
Clog\(Q\+1\)MCErr𝖣θΓ\(Π;𝒢Q\),C\\log\(Q\+1\)\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\),and hence
PredErrθ,j\(Π\)≤Clog\(Q\+1\)MCErr𝖣θΓ\(Π;𝒢Q\)\+C∑i<jPredErrθ,i\(Π\)\.\\operatorname\{PredErr\}\_\{\\theta,j\}\(\\Pi\)\\leq C\\log\(Q\+1\)\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\+C\\sum\_\{i<j\}\\operatorname\{PredErr\}\_\{\\theta,i\}\(\\Pi\)\.Induction overj=1,…,k−1j=1,\\ldots,k\-1yields
∑j<kPredErrθ,j\(Π\)≤Clog\(Q\+1\)MCErr𝖣θΓ\(Π;𝒢Q\)\.\\sum\_\{j<k\}\\operatorname\{PredErr\}\_\{\\theta,j\}\(\\Pi\)\\leq C\\log\(Q\+1\)\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\.\(2\)
##### Predictions Outside the Boxes\.
Set
Sincecη<1/20c\_\{\\eta\}<1/20, we haver<dQ/4r<d\_\{Q\}/4\. Define boxes around the unperturbed property vectors by
Ca:=\{p∈𝒫:\|pi−bi\(ai\)\|≤rfor everyi∈\[k\]\}\.C\_\{a\}:=\\\{p\\in\\mathcal\{P\}:\\ \|p\_\{i\}\-b\_\{i\}\(a\_\{i\}\)\|\\leq r\\ \\text\{for every \}i\\in\[k\]\\\}\.Each true property vectorvθ\(a\)v\_\{\\theta\}\(a\)lies inCaC\_\{a\}, and the boxes are pairwise disjoint because2r<dQ2r<d\_\{Q\}\. Let
Ekout:=1Qk∑a∫𝒫∖Ca\|pk−vθ,k\(a\)\|Πa\(dp\)\.E\_\{k\}^\{\\mathrm\{out\}\}:=\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\int\_\{\\mathcal\{P\}\\setminus C\_\{a\}\}\|p\_\{k\}\-v\_\{\\theta,k\}\(a\)\|\\,\\Pi\_\{a\}\(dp\)\.Decompose this error according to whether the final prediction is nearbk\(ak\)b\_\{k\}\(a\_\{k\}\)\. Define
Ekout,near\\displaystyle E\_\{k\}^\{\\mathrm\{out,near\}\}:=1Qk∑a∫𝒫∖Ca𝟏\{\|pk−bk\(ak\)\|≤r\}\|pk−vθ,k\(a\)\|Πa\(dp\),\\displaystyle=\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\int\_\{\\mathcal\{P\}\\setminus C\_\{a\}\}\\mathbf\{1\}\\\{\|p\_\{k\}\-b\_\{k\}\(a\_\{k\}\)\|\\leq r\\\}\|p\_\{k\}\-v\_\{\\theta,k\}\(a\)\|\\,\\Pi\_\{a\}\(dp\),Ekout,far\\displaystyle E\_\{k\}^\{\\mathrm\{out,far\}\}:=1Qk∑a∫𝒫∖Ca𝟏\{\|pk−bk\(ak\)\|\>r\}\|pk−vθ,k\(a\)\|Πa\(dp\)\.\\displaystyle=\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\int\_\{\\mathcal\{P\}\\setminus C\_\{a\}\}\\mathbf\{1\}\\\{\|p\_\{k\}\-b\_\{k\}\(a\_\{k\}\)\|\>r\\\}\|p\_\{k\}\-v\_\{\\theta,k\}\(a\)\|\\,\\Pi\_\{a\}\(dp\)\.Then
Ekout=Ekout,near\+Ekout,far\.E\_\{k\}^\{\\mathrm\{out\}\}=E\_\{k\}^\{\\mathrm\{out,near\}\}\+E\_\{k\}^\{\\mathrm\{out,far\}\}\.On the near part,p∉Cap\\notin C\_\{a\}implies that some prefix coordinatei<ki<khas\|pi−bi\(ai\)\|\>r\|p\_\{i\}\-b\_\{i\}\(a\_\{i\}\)\|\>r, while
\|pk−vθ,k\(a\)\|≤r\+η≤C∑i<k\|pi−vθ,i\(a\)\|\.\|p\_\{k\}\-v\_\{\\theta,k\}\(a\)\|\\leq r\+\\eta\\leq C\\sum\_\{i<k\}\|p\_\{i\}\-v\_\{\\theta,i\}\(a\)\|\.The second inequality usesvθ,i\(a\)=bi\(ai\)v\_\{\\theta,i\}\(a\)=b\_\{i\}\(a\_\{i\}\)fori<ki<k, and the fact that one prefix error is at leastrrwhilerrandη\\etaare fixed constant multiples of one another\. Therefore
Ekout,near≤C∑i<kPredErrθ,i\(Π\)\.E\_\{k\}^\{\\mathrm\{out,near\}\}\\leq C\\sum\_\{i<k\}\\operatorname\{PredErr\}\_\{\\theta,i\}\(\\Pi\)\.
For the far part, define the two\-sided threshold sign
spk\(k\)\(a\):=12\(sgn\(pk−r−bk\(ak\)\)\+sgn\(pk\+r−bk\(ak\)\)\)\.s^\{\(k\)\}\_\{p\_\{k\}\}\(a\):=\\frac\{1\}\{2\}\\left\(\\operatorname\{sgn\}\(p\_\{k\}\-r\-b\_\{k\}\(a\_\{k\}\)\)\+\\operatorname\{sgn\}\(p\_\{k\}\+r\-b\_\{k\}\(a\_\{k\}\)\)\\right\)\.As a function ofaka\_\{k\}, this belongs to𝒯Q\\mathcal\{T\}\_\{Q\}\. On the event\|pk−bk\(ak\)\|\>r\|p\_\{k\}\-b\_\{k\}\(a\_\{k\}\)\|\>r, it agrees withsgn\(pk−vθ,k\(a\)\)\\operatorname\{sgn\}\(p\_\{k\}\-v\_\{\\theta,k\}\(a\)\), because\|vθ,k\(a\)−bk\(ak\)\|=η<r\|v\_\{\\theta,k\}\(a\)\-b\_\{k\}\(a\_\{k\}\)\|=\\eta<r\. Lemma[3\.5](https://arxiv.org/html/2608.04288#S3.Thmtheorem5)gives
canti𝟏\{\|pk−bk\(ak\)\|\>r\}\|pk−vθ,k\(a\)\|≤spk\(k\)\(a\)𝖱a,k\(p\)\+C∑i<k\|pi−vθ,i\(a\)\|\.c\_\{\\mathrm\{anti\}\}\\mathbf\{1\}\\\{\|p\_\{k\}\-b\_\{k\}\(a\_\{k\}\)\|\>r\\\}\|p\_\{k\}\-v\_\{\\theta,k\}\(a\)\|\\leq s^\{\(k\)\}\_\{p\_\{k\}\}\(a\)\\mathsf\{R\}\_\{a,k\}\(p\)\+C\\sum\_\{i<k\}\|p\_\{i\}\-v\_\{\\theta,i\}\(a\)\|\.Lets^pk\(k\)\\widehat\{s\}^\{\(k\)\}\_\{p\_\{k\}\}be the Walsh approximation tospk\(k\)s^\{\(k\)\}\_\{p\_\{k\}\}\. Replacingspk\(k\)s^\{\(k\)\}\_\{p\_\{k\}\}bys^pk\(k\)\\widehat\{s\}^\{\(k\)\}\_\{p\_\{k\}\}, integrating, and controlling the approximation error gives
cantiEkout,far\\displaystyle c\_\{\\mathrm\{anti\}\}E\_\{k\}^\{\\mathrm\{out,far\}\}≤\|1Qk∑a∫s^pk\(k\)\(a\)𝖱a,k\(p\)Πa\(dp\)\|\\displaystyle\\leq\\left\|\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\int\\widehat\{s\}^\{\(k\)\}\_\{p\_\{k\}\}\(a\)\\mathsf\{R\}\_\{a,k\}\(p\)\\,\\Pi\_\{a\}\(dp\)\\right\|\+C∑i<kPredErrθ,i\(Π\)\+CαPredErrθ,k\(Π\)\.\\displaystyle\\quad\+C\\sum\_\{i<k\}\\operatorname\{PredErr\}\_\{\\theta,i\}\(\\Pi\)\+C\\alpha\\operatorname\{PredErr\}\_\{\\theta,k\}\(\\Pi\)\.Indeed, the extraCαPredErrθ,k\(Π\)C\\alpha\\operatorname\{PredErr\}\_\{\\theta,k\}\(\\Pi\)term comes from
\|s^pk\(k\)\(a\)−spk\(k\)\(a\)\|\|𝖱a,k\(p\)\|≤α\(Cek\(a,p\)\+C∑i<kei\(a,p\)\),\\left\|\\widehat\{s\}^\{\(k\)\}\_\{p\_\{k\}\}\(a\)\-s^\{\(k\)\}\_\{p\_\{k\}\}\(a\)\\right\|\\,\|\\mathsf\{R\}\_\{a,k\}\(p\)\|\\leq\\alpha\\left\(Ce\_\{k\}\(a,p\)\+C\\sum\_\{i<k\}e\_\{i\}\(a,p\)\\right\),where the residual bound is the upper inequality in Lemma[3\.5](https://arxiv.org/html/2608.04288#S3.Thmtheorem5)\. The Walsh coefficients again satisfy the conditions of Lemma[3\.6](https://arxiv.org/html/2608.04288#S3.Thmtheorem6), with coefficient boundClog\(Q\+1\)C\\log\(Q\+1\), so that lemma bounds the term inside the absolute value by
Clog\(Q\+1\)MCErr𝖣θΓ\(Π;𝒢Q\)\.C\\log\(Q\+1\)\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\.Therefore
Ekout,far≤Clog\(Q\+1\)MCErr𝖣θΓ\(Π;𝒢Q\)\+C∑i<kPredErrθ,i\(Π\)\+CαPredErrθ,k\(Π\)\.E\_\{k\}^\{\\mathrm\{out,far\}\}\\leq C\\log\(Q\+1\)\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\+C\\sum\_\{i<k\}\\operatorname\{PredErr\}\_\{\\theta,i\}\(\\Pi\)\+C\\alpha\\operatorname\{PredErr\}\_\{\\theta,k\}\(\\Pi\)\.Combining the near and far bounds with \([2](https://arxiv.org/html/2608.04288#S3.E2)\),
Ekout≤Clog\(Q\+1\)MCErr𝖣θΓ\(Π;𝒢Q\)\+CαPredErrθ,k\(Π\)\.E\_\{k\}^\{\\mathrm\{out\}\}\\leq C\\log\(Q\+1\)\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\+C\\alpha\\operatorname\{PredErr\}\_\{\\theta,k\}\(\\Pi\)\.\(3\)
##### Predictions Inside the Boxes\.
Inside the boxes, we use the all\-ones group rather than another threshold sign\. Because the boxesCaC\_\{a\}are pairwise disjoint, the residual contributions localized to different contexts occupy disjoint regions of the prediction space and therefore do not cancel in total variation\. The all\-ones residual also includes predictions outside the boxes, whose contribution is controlled by the preceding bounds\.
Let
Ekloc:=1Qk∑a∫Ca\|pk−vθ,k\(a\)\|Πa\(dp\),E\_\{k\}^\{\\mathrm\{loc\}\}:=\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\int\_\{C\_\{a\}\}\|p\_\{k\}\-v\_\{\\theta,k\}\(a\)\|\\,\\Pi\_\{a\}\(dp\),so that
PredErrθ,k\(Π\)=Ekloc\+Ekout\.\\operatorname\{PredErr\}\_\{\\theta,k\}\(\\Pi\)=E\_\{k\}^\{\\mathrm\{loc\}\}\+E\_\{k\}^\{\\mathrm\{out\}\}\.Letνgall,k𝖣θ,Π\\nu^\{\\mathsf\{D\}\_\{\\theta\},\\Pi\}\_\{g\_\{\\mathrm\{all\}\},k\}be the signed measure for the final coordinate and the all\-ones group, and define its localized part by
νkloc\(A\):=1Qk∑a∫A∩Ca𝖱a,k\(p\)Πa\(dp\)\.\\nu^\{\\mathrm\{loc\}\}\_\{k\}\(A\):=\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\int\_\{A\\cap C\_\{a\}\}\\mathsf\{R\}\_\{a,k\}\(p\)\\,\\Pi\_\{a\}\(dp\)\.Because the boxesCaC\_\{a\}are pairwise disjoint,
\|νkloc\|\(𝒫\)=1Qk∑a∫Ca\|𝖱a,k\(p\)\|Πa\(dp\)\.\|\\nu^\{\\mathrm\{loc\}\}\_\{k\}\|\(\\mathcal\{P\}\)=\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\int\_\{C\_\{a\}\}\|\\mathsf\{R\}\_\{a,k\}\(p\)\|\\,\\Pi\_\{a\}\(dp\)\.The lower inequality in Lemma[3\.5](https://arxiv.org/html/2608.04288#S3.Thmtheorem5)implies, pointwise,
cantiek\(a,p\)≤\|𝖱a,k\(p\)\|\+C∑i<kei\(a,p\)\.c\_\{\\mathrm\{anti\}\}e\_\{k\}\(a,p\)\\leq\|\\mathsf\{R\}\_\{a,k\}\(p\)\|\+C\\sum\_\{i<k\}e\_\{i\}\(a,p\)\.Integrating this bound overCaC\_\{a\}, summing overaa, and dividing byQkQ^\{k\}gives
Ekloc≤C\|νkloc\|\(𝒫\)\+C∑i<kPredErrθ,i\(Π\)\.E\_\{k\}^\{\\mathrm\{loc\}\}\\leq C\|\\nu^\{\\mathrm\{loc\}\}\_\{k\}\|\(\\mathcal\{P\}\)\+C\\sum\_\{i<k\}\\operatorname\{PredErr\}\_\{\\theta,i\}\(\\Pi\)\.Write
νgall,k𝖣θ,Π=νkloc\+νkout,\\nu^\{\\mathsf\{D\}\_\{\\theta\},\\Pi\}\_\{g\_\{\\mathrm\{all\}\},k\}=\\nu^\{\\mathrm\{loc\}\}\_\{k\}\+\\nu^\{\\mathrm\{out\}\}\_\{k\},whereνkout\\nu^\{\\mathrm\{out\}\}\_\{k\}is the contribution fromp∉Cap\\notin C\_\{a\}\. Lemma[3\.5](https://arxiv.org/html/2608.04288#S3.Thmtheorem5)gives
\|νkout\|\(𝒫\)≤CEkout\+C∑i<kPredErrθ,i\(Π\)\.\|\\nu^\{\\mathrm\{out\}\}\_\{k\}\|\(\\mathcal\{P\}\)\\leq CE\_\{k\}^\{\\mathrm\{out\}\}\+C\\sum\_\{i<k\}\\operatorname\{PredErr\}\_\{\\theta,i\}\(\\Pi\)\.Here the total variation of the outside contribution is bounded by the integral of\|𝖱a,k\(p\)\|\|\\mathsf\{R\}\_\{a,k\}\(p\)\|overp∉Cap\\notin C\_\{a\}; the boxesCaC\_\{a\}are disjoint, but their complements need not be\. The upper inequality in Lemma[3\.5](https://arxiv.org/html/2608.04288#S3.Thmtheorem5)then bounds this integral byEkoutE\_\{k\}^\{\\mathrm\{out\}\}and the prefix errors\. Since
\|νgall,k𝖣θ,Π\|\(𝒫\)≤MCErr𝖣θΓ\(Π;𝒢Q\),\|\\nu^\{\\mathsf\{D\}\_\{\\theta\},\\Pi\}\_\{g\_\{\\mathrm\{all\}\},k\}\|\(\\mathcal\{P\}\)\\leq\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\),the measure decomposition also gives
\|νkloc\|\(𝒫\)≤\|νgall,k𝖣θ,Π\|\(𝒫\)\+\|νkout\|\(𝒫\)\.\|\\nu^\{\\mathrm\{loc\}\}\_\{k\}\|\(\\mathcal\{P\}\)\\leq\|\\nu^\{\\mathsf\{D\}\_\{\\theta\},\\Pi\}\_\{g\_\{\\mathrm\{all\}\},k\}\|\(\\mathcal\{P\}\)\+\|\\nu^\{\\mathrm\{out\}\}\_\{k\}\|\(\\mathcal\{P\}\)\.Combining these bounds gives
Ekloc≤CMCErr𝖣θΓ\(Π;𝒢Q\)\+CEkout\+C∑i<kPredErrθ,i\(Π\)\.E\_\{k\}^\{\\mathrm\{loc\}\}\\leq C\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\+CE\_\{k\}^\{\\mathrm\{out\}\}\+C\\sum\_\{i<k\}\\operatorname\{PredErr\}\_\{\\theta,i\}\(\\Pi\)\.Using \([2](https://arxiv.org/html/2608.04288#S3.E2)\) and \([3](https://arxiv.org/html/2608.04288#S3.E3)\),
PredErrθ,k\(Π\)≤Clog\(Q\+1\)MCErr𝖣θΓ\(Π;𝒢Q\)\+CαPredErrθ,k\(Π\)\.\\operatorname\{PredErr\}\_\{\\theta,k\}\(\\Pi\)\\leq C\\log\(Q\+1\)\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\+C\\alpha\\operatorname\{PredErr\}\_\{\\theta,k\}\(\\Pi\)\.Absorbing the final term into the left\-hand side gives
PredErrθ,k\(Π\)≤Clog\(Q\+1\)MCErr𝖣θΓ\(Π;𝒢Q\)\.\\operatorname\{PredErr\}\_\{\\theta,k\}\(\\Pi\)\\leq C\\log\(Q\+1\)\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\.Together with \([2](https://arxiv.org/html/2608.04288#S3.E2)\), this proves the proposition\. ∎
### 3\.7Recovering the Sign Assignment
Proposition[3\.7](https://arxiv.org/html/2608.04288#S3.Thmtheorem7)converts smallΓ\\Gamma\-ECE into small average prediction error\. Distinct sign assignments produce different true property vectors on a constant fraction of the contexts, so sufficiently small prediction error identifies the true assignment\. The KL bound in Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(v\) limits how much informationnnsamples contain about that assignment, and Fano’s inequality then gives the required lower bound on sample complexity\.
###### Lemma 3\.8\(Different Sign Assignments Produce Separated Property Vectors\)\.
For every distinctθ,θ′∈ΘQ\\theta,\\theta^\{\\prime\}\\in\\Theta\_\{Q\},
1Qk∑a∈𝒳Q‖vθ\(a\)−vθ′\(a\)‖1≥2ρpackη\.\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\\in\\mathcal\{X\}\_\{Q\}\}\\\|v\_\{\\theta\}\(a\)\-v\_\{\\theta^\{\\prime\}\}\(a\)\\\|\_\{1\}\\geq 2\\rho\_\{\\mathrm\{pack\}\}\\eta\.
###### Proof\.
For every contextaa, the vectorsvθ\(a\)v\_\{\\theta\}\(a\)andvθ′\(a\)v\_\{\\theta^\{\\prime\}\}\(a\)agree in their firstk−1k\-1coordinates\. Wheneverθa≠θa′\\theta\_\{a\}\\neq\\theta^\{\\prime\}\_\{a\}, their final coordinates satisfy
\|vθ,k\(a\)−vθ′,k\(a\)\|=2η\.\|v\_\{\\theta,k\}\(a\)\-v\_\{\\theta^\{\\prime\},k\}\(a\)\|=2\\eta\.The packing condition gives at leastρpackQk\\rho\_\{\\mathrm\{pack\}\}Q^\{k\}such contexts\. ∎
###### Lemma 3\.9\(Prediction Error Implies Exact Decoding\)\.
There exists a decoderθ^\(Π\)\\widehat\{\\theta\}\(\\Pi\)such that, if
PredErrθ\(Π\)≤ρpackη2,\\operatorname\{PredErr\}\_\{\\theta\}\(\\Pi\)\\leq\\frac\{\\rho\_\{\\mathrm\{pack\}\}\\eta\}\{2\},then
θ^\(Π\)=θ\.\\widehat\{\\theta\}\(\\Pi\)=\\theta\.
###### Proof\.
For eachaa, define the mean prediction value
p¯Π\(a\):=∫𝒫pΠa\(dp\)\.\\bar\{p\}\_\{\\Pi\}\(a\):=\\int\_\{\\mathcal\{P\}\}p\\,\\Pi\_\{a\}\(dp\)\.By Jensen’s inequality,
1Qk∑a‖p¯Π\(a\)−vθ\(a\)‖1≤PredErrθ\(Π\)\.\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\\|\\bar\{p\}\_\{\\Pi\}\(a\)\-v\_\{\\theta\}\(a\)\\\|\_\{1\}\\leq\\operatorname\{PredErr\}\_\{\\theta\}\(\\Pi\)\.Decode by nearest neighbor:
θ^\(Π\)∈argminθ′∈ΘQ1Qk∑a‖p¯Π\(a\)−vθ′\(a\)‖1\.\\widehat\{\\theta\}\(\\Pi\)\\in\\arg\\min\_\{\\theta^\{\\prime\}\\in\\Theta\_\{Q\}\}\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\\|\\bar\{p\}\_\{\\Pi\}\(a\)\-v\_\{\\theta^\{\\prime\}\}\(a\)\\\|\_\{1\}\.Ifθ′≠θ\\theta^\{\\prime\}\\neq\\theta, then Lemma[3\.8](https://arxiv.org/html/2608.04288#S3.Thmtheorem8)and the triangle inequality imply
1Qk∑a‖p¯Π\(a\)−vθ′\(a\)‖1≥2ρpackη−ρpackη2\>ρpackη2≥1Qk∑a‖p¯Π\(a\)−vθ\(a\)‖1\.\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\\|\\bar\{p\}\_\{\\Pi\}\(a\)\-v\_\{\\theta^\{\\prime\}\}\(a\)\\\|\_\{1\}\\geq 2\\rho\_\{\\mathrm\{pack\}\}\\eta\-\\frac\{\\rho\_\{\\mathrm\{pack\}\}\\eta\}\{2\}\>\\frac\{\\rho\_\{\\mathrm\{pack\}\}\\eta\}\{2\}\\geq\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}\\\|\\bar\{p\}\_\{\\Pi\}\(a\)\-v\_\{\\theta\}\(a\)\\\|\_\{1\}\.Thus the unique nearest neighbor isθ\\theta\. ∎
###### Lemma 3\.10\(Pairwise KL Bound\)\.
For everyθ,θ′∈ΘQ\\theta,\\theta^\{\\prime\}\\in\\Theta\_\{Q\}and every sample sizenn,
DKL\(𝖣θn∥𝖣θ′n\)≤4CKLnη2\.D\_\{\\mathrm\{KL\}\}\(\\mathsf\{D\}\_\{\\theta\}^\{n\}\\,\\\|\\,\\mathsf\{D\}\_\{\\theta^\{\\prime\}\}^\{n\}\)\\leq 4C\_\{\\mathrm\{KL\}\}\\,n\\eta^\{2\}\.
###### Proof\.
SinceXXis uniform on𝒳Q\\mathcal\{X\}\_\{Q\}under both distributions,
DKL\(𝖣θ∥𝖣θ′\)=1Qk∑aDKL\(μvθ\(a\)∥μvθ′\(a\)\)\.D\_\{\\mathrm\{KL\}\}\(\\mathsf\{D\}\_\{\\theta\}\\,\\\|\\,\\mathsf\{D\}\_\{\\theta^\{\\prime\}\}\)=\\frac\{1\}\{Q^\{k\}\}\\sum\_\{a\}D\_\{\\mathrm\{KL\}\}\\left\(\\mu\_\{v\_\{\\theta\}\(a\)\}\\,\\middle\\\|\\,\\mu\_\{v\_\{\\theta^\{\\prime\}\}\(a\)\}\\right\)\.By Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(v\),
DKL\(μvθ\(a\)∥μvθ′\(a\)\)≤CKL∥vθ\(a\)−vθ′\(a\)∥22≤4CKLη2\.D\_\{\\mathrm\{KL\}\}\\left\(\\mu\_\{v\_\{\\theta\}\(a\)\}\\,\\middle\\\|\\,\\mu\_\{v\_\{\\theta^\{\\prime\}\}\(a\)\}\\right\)\\leq C\_\{\\mathrm\{KL\}\}\\\|v\_\{\\theta\}\(a\)\-v\_\{\\theta^\{\\prime\}\}\(a\)\\\|\_\{2\}^\{2\}\\leq 4C\_\{\\mathrm\{KL\}\}\\eta^\{2\}\.Tensorization overnni\.i\.d\. samples gives the claim\. ∎
Combining Proposition[3\.7](https://arxiv.org/html/2608.04288#S3.Thmtheorem7), Lemma[3\.9](https://arxiv.org/html/2608.04288#S3.Thmtheorem9), and Lemma[3\.10](https://arxiv.org/html/2608.04288#S3.Thmtheorem10)gives a lower bound on sample complexity at a fixed resolution\.
###### Proposition 3\.11\(Lower Bound at ResolutionQQ\)\.
There are constantsc∗,C∗\>0c\_\{\*\},C\_\{\*\}\>0, depending only onkk,ℛ0\\mathcal\{R\}\_\{0\}, and the constants in Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1), such that the following holds for all sufficiently large powers of twoQQ\. If a possibly randomized learner, givennni\.i\.d\. samples from𝖣θ\\mathsf\{D\}\_\{\\theta\}, outputs a predictorΠ\\Pisatisfying
MCErr𝖣θΓ\(Π;𝒢Q\)≤c∗ηlog\(Q\+1\)\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\_\{\\theta\}\}\(\\Pi;\\mathcal\{G\}\_\{Q\}\)\\leq c\_\{\*\}\\frac\{\\eta\}\{\\log\(Q\+1\)\}with probability at least2/32/3for everyθ∈ΘQ\\theta\\in\\Theta\_\{Q\}, then
n≥C∗Qk\+2\.n\\geq C\_\{\*\}Q^\{k\+2\}\.
###### Proof\.
Choosec∗\>0c\_\{\*\}\>0small enough that Proposition[3\.7](https://arxiv.org/html/2608.04288#S3.Thmtheorem7)implies
PredErrθ\(Π\)≤ρpackη2\\operatorname\{PredErr\}\_\{\\theta\}\(\\Pi\)\\leq\\frac\{\\rho\_\{\\mathrm\{pack\}\}\\eta\}\{2\}whenever the learner achieves the multicalibration error required in the proposition\. By Lemma[3\.9](https://arxiv.org/html/2608.04288#S3.Thmtheorem9), the learner then induces an exact decoder forθ\\thetawith success probability at least2/32/3\.
LetΘ\\Thetabe uniformly distributed overΘQ\\Theta\_\{Q\}, and let the sample be drawn from𝖣Θn\\mathsf\{D\}\_\{\\Theta\}^\{n\}\. If the learner is randomized, include its independent random seed in the decoder’s observation; this does not change the KL bound\. Fano’s inequality\[[10](https://arxiv.org/html/2608.04288#bib.bib36)\]gives
ℙ\[Θ^≠Θ\]≥1−I\(Θ;sample\)\+log2log\|ΘQ\|\.\\mathbb\{P\}\[\\widehat\{\\Theta\}\\neq\\Theta\]\\geq 1\-\\frac\{I\(\\Theta;\\text\{sample\}\)\+\\log 2\}\{\\log\|\\Theta\_\{Q\}\|\}\.The standard bound on mutual information in terms of pairwise KL divergence, together with Lemma[3\.10](https://arxiv.org/html/2608.04288#S3.Thmtheorem10), implies
I\(Θ;sample\)≤4CKLnη2\.I\(\\Theta;\\text\{sample\}\)\\leq 4C\_\{\\mathrm\{KL\}\}n\\eta^\{2\}\.A decoder with error probability at most1/31/3therefore requires
4CKLnη2\+log2≥23log\|ΘQ\|≥cQk4C\_\{\\mathrm\{KL\}\}n\\eta^\{2\}\+\\log 2\\geq\\frac\{2\}\{3\}\\log\|\\Theta\_\{Q\}\|\\geq cQ^\{k\}for a constantc\>0c\>0\. For all sufficiently largeQQ, thelog2\\log 2term is absorbed into the right\-hand side, so
4CKLnη2≥cQk4C\_\{\\mathrm\{KL\}\}n\\eta^\{2\}\\geq cQ^\{k\}after decreasingcc\. Sinceη≍Q−1\\eta\\asymp Q^\{\-1\},
n≥C∗Qk\+2\.n\\geq C\_\{\*\}Q^\{k\+2\}\.∎
### 3\.8Proof of the Main Theorem
ChoosingQQas a function of the target errorε\\varepsilonconverts Proposition[3\.11](https://arxiv.org/html/2608.04288#S3.Thmtheorem11)into Theorem[3\.2](https://arxiv.org/html/2608.04288#S3.Thmtheorem2)\.
###### Proof of Theorem[3\.2](https://arxiv.org/html/2608.04288#S3.Thmtheorem2)\.
For each large power of twoQQ, Proposition[3\.11](https://arxiv.org/html/2608.04288#S3.Thmtheorem11)gives a finite collection of candidate distributions for whichn≥C∗Qk\+2n\\geq C\_\{\*\}Q^\{k\+2\}samples are required at error level
εQ:=c∗ηlog\(Q\+1\)≍1QlogQ\.\\varepsilon\_\{Q\}:=c\_\{\*\}\\frac\{\\eta\}\{\\log\(Q\+1\)\}\\asymp\\frac\{1\}\{Q\\log Q\}\.Given sufficiently smallε\>0\\varepsilon\>0, chooseQQto be the largest power of two such that
ε≤εQ\.\\varepsilon\\leq\\varepsilon\_\{Q\}\.Then
Q≥c′1εlog\(C′/ε\)Q\\geq c^\{\\prime\}\\frac\{1\}\{\\varepsilon\\log\(C^\{\\prime\}/\\varepsilon\)\}for constantsc′,C′\>0c^\{\\prime\},C^\{\\prime\}\>0\. Therefore
n≥C∗Qk\+2≥cε−\(k\+2\)logk\+2\(C/ε\)n\\geq C\_\{\*\}Q^\{k\+2\}\\geq c\\frac\{\\varepsilon^\{\-\(k\+2\)\}\}\{\\log^\{k\+2\}\(C/\\varepsilon\)\}for constantsc,C\>0c,C\>0depending only onkk,κ\\kappa,ℛ0\\mathcal\{R\}\_\{0\}, and the constants in Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\. The group family satisfies
\|𝒢Q\|≤Cα,klog3\(Q\+1\)≤ε−κ\|\\mathcal\{G\}\_\{Q\}\|\\leq C\_\{\\alpha,k\}\\log^\{3\}\(Q\+1\)\\leq\\varepsilon^\{\-\\kappa\}for every fixedκ\>0\\kappa\>0and all sufficiently smallε\\varepsilon\. Taking
𝒳ε:=𝒳Q,𝒢ε:=𝒢Q,𝔇ε:=\{𝖣θ:θ∈ΘQ\}\\mathcal\{X\}\_\{\\varepsilon\}:=\\mathcal\{X\}\_\{Q\},\\qquad\\mathcal\{G\}\_\{\\varepsilon\}:=\\mathcal\{G\}\_\{Q\},\\qquad\\mathfrak\{D\}\_\{\\varepsilon\}:=\\\{\\mathsf\{D\}\_\{\\theta\}:\\theta\\in\\Theta\_\{Q\}\\\}proves the theorem\. ∎
## 4The Upper Bound
Our route to the upper bound is through an online forecasting problem\. On each round, the forecaster constructs a prediction rule from the preceding rounds, applies it to the current context, and then observes the outcome\. We first control the empirical multicalibration error accumulated over these rounds\.
Given a dataset consisting ofTTi\.i\.d\. context–outcome pairs, we run the online procedure forTTrounds, using one pair on each round\. The context component of thett\-th pair is revealed before the prediction and its outcome component afterward\. The procedure producesTTprediction rules\. We return their average as our predictor\. A martingale argument then transfers the empirical guarantee for the online transcript to a population multicalibration guarantee for this predictor\.
We begin by formulating the online problem and expressing its empirical multicalibration error as a collection of objectives that the forecaster must control simultaneously\. We then state the regularity conditions, construct and analyze the forecaster, and carry out the online\-to\-batch reduction\. Finally, we show how to implement the forecaster efficiently using linear optimization overℳ\\mathcal\{M\}\.
### 4\.1The Online Formulation and Its Objectives
We first describe the online formulation for an arbitrary finite set𝒫Q⊂𝒫\\mathcal\{P\}\_\{Q\}\\subset\\mathcal\{P\}of possible predictions\. Fix a horizonTTand a finite group family𝒢\\mathcal\{G\}\.
Letℱt−1\\mathcal\{F\}\_\{t\-1\}be the sigma\-field generated by the observations through roundt−1t\-1and any internal randomness used by the forecaster before roundtt\. At the beginning of roundtt, the forecaster uses this history to define anℱt−1\\mathcal\{F\}\_\{t\-1\}\-measurable ruleπt:𝒳→Δ\(𝒫Q\)\\pi\_\{t\}:\\mathcal\{X\}\\to\\Delta\(\\mathcal\{P\}\_\{Q\}\)\. After the contextXtX\_\{t\}is revealed, it outputsπt\(⋅∣Xt\)\\pi\_\{t\}\(\\cdot\\mid X\_\{t\}\)as the randomized prediction and then observes the outcomeYtY\_\{t\}\.
For any sequence of such rules, define the empirical multicalibration error over theTTonline rounds by
MCErr^TΓ:=1Tmaxg∈𝒢∑p∈𝒫Q∥∑t=1Tg\(Xt\)πt\(p∣Xt\)R\(p,Yt\)∥1\.\\widehat\{\\operatorname\{MCErr\}\}^\{\\Gamma\}\_\{T\}:=\\frac\{1\}\{T\}\\max\_\{g\\in\\mathcal\{G\}\}\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\left\\\|\\sum\_\{t=1\}^\{T\}g\(X\_\{t\}\)\\pi\_\{t\}\(p\\mid X\_\{t\}\)R\(p,Y\_\{t\}\)\\right\\\|\_\{1\}\.LetΣ:=\{±1\}𝒫Q×\[k\]\\Sigma:=\\\{\\pm 1\\\}^\{\\mathcal\{P\}\_\{Q\}\\times\[k\]\}, and writesp,js\_\{p,j\}for the sign assigned bys∈Σs\\in\\Sigmato the pair\(p,j\)\(p,j\)\. Representing each absolute value in theℓ1\\ell\_\{1\}\-norm as a maximum over its sign gives
MCErr^TΓ\\displaystyle\\widehat\{\\operatorname\{MCErr\}\}^\{\\Gamma\}\_\{T\}=1Tmaxg∈𝒢s∈Σ∑p∈𝒫Q∑j=1ksp,j∑t=1Tg\(Xt\)πt\(p∣Xt\)Rj\(p<j,pj,Yt\)\\displaystyle=\\frac\{1\}\{T\}\\max\_\{\\begin\{subarray\}\{c\}g\\in\\mathcal\{G\}\\\\ s\\in\\Sigma\\end\{subarray\}\}\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\sum\_\{j=1\}^\{k\}s\_\{p,j\}\\sum\_\{t=1\}^\{T\}g\(X\_\{t\}\)\\pi\_\{t\}\(p\\mid X\_\{t\}\)R\_\{j\}\(p\_\{<j\},p\_\{j\},Y\_\{t\}\)=1Tmaxg∈𝒢s∈Σ∑t=1Tg\(Xt\)𝔼p∼πt\(⋅∣Xt\)\[∑j=1ksp,jRj\(p<j,pj,Yt\)\]\.\\displaystyle=\\frac\{1\}\{T\}\\max\_\{\\begin\{subarray\}\{c\}g\\in\\mathcal\{G\}\\\\ s\\in\\Sigma\\end\{subarray\}\}\\sum\_\{t=1\}^\{T\}g\(X\_\{t\}\)\\mathbb\{E\}\_\{p\\sim\\pi\_\{t\}\(\\cdot\\mid X\_\{t\}\)\}\\left\[\\sum\_\{j=1\}^\{k\}s\_\{p,j\}R\_\{j\}\(p\_\{<j\},p\_\{j\},Y\_\{t\}\)\\right\]\.The above equalities express the empirical error as the largest cumulative signed residual among the objectives indexed by\(g,s\)∈𝒢×Σ\(g,s\)\\in\\mathcal\{G\}\\times\\Sigma\. Thus, the online task is a multiobjective learning problem: the forecaster must control all of these signed objectives simultaneously\[[34](https://arxiv.org/html/2608.04288#bib.bib28)\]\. We use the framework of online learning with expert advice\. Each pair\(g,s\)\(g,s\)is treated as an expert, whose gain on roundttis
ht\(g,s\):=g\(Xt\)𝔼p∼πt\(⋅∣Xt\)\[∑j=1ksp,jRj\(p<j,pj,Yt\)\]\.h\_\{t\}\(g,s\):=g\(X\_\{t\}\)\\mathbb\{E\}\_\{p\\sim\\pi\_\{t\}\(\\cdot\\mid X\_\{t\}\)\}\\left\[\\sum\_\{j=1\}^\{k\}s\_\{p,j\}R\_\{j\}\(p\_\{<j\},p\_\{j\},Y\_\{t\}\)\\right\]\.Since the empirical error is the largest cumulative gain of any expert divided byTT, our goal is to control the cumulative gain of every expert\. The exponential weights algorithm chooses a distributionwtw\_\{t\}over the experts on each round\. We call the expectation ofht\(g,s\)h\_\{t\}\(g,s\)under this distribution the expected gain underwtw\_\{t\}\. Its regret guarantee bounds the cumulative gain of every individual expert by the sum of the expected gains underwtw\_\{t\}, plus a sublinear regret term\. It therefore remains to keep these expected gains small\.
We use the same weightswtw\_\{t\}to choose the forecaster’s distribution on𝒫Q\\mathcal\{P\}\_\{Q\}\. For any candidate distributionπ∈Δ\(𝒫Q\)\\pi\\in\\Delta\(\\mathcal\{P\}\_\{Q\}\)and any admissible outcome law, we can evaluate the expected gain underwtw\_\{t\}from the corresponding signed residuals\. The forecaster choosesπ\\piin a minimax manner, minimizing the largest possible value of this expected gain over admissible outcome laws\. We allow an additive errorρ≥0\\rho\\geq 0in this minimization\.Noarovet al\.\[[37](https://arxiv.org/html/2608.04288#bib.bib33)\]use a related construction for high\-dimensional unbiased prediction, in which an online learning algorithm weights signed objectives and those weights guide the randomized forecast\.
### 4\.2Regularity Assumptions
The online formulation can be stated for any finite set𝒫Q\\mathcal\{P\}\_\{Q\}, but its analysis requires additional structure\. Convexity, compactness, and continuity \(Condition \(i\) below\) allow us to apply Sion’s minimax theorem\. Although the forecaster must choose a distribution on𝒫Q\\mathcal\{P\}\_\{Q\}without knowing the conditional outcome law, the theorem allows us to analyze the one\-round objective by first fixing that law and then choosing the distribution\. A uniform bound on the residual \(Condition \(ii\)\) controls both the expert gains and the martingale deviations\. Finally, Condition \(iii\) places two requirements on the grid\. It has onlyO\(Qk\)O\(Q^\{k\}\)points, and for every admissible outcome lawμ\\mu, somep∈𝒫Qp\\in\\mathcal\{P\}\_\{Q\}makes every coordinate ofR\(p,μ\)R\(p,\\mu\)small\. Afterμ\\muis fixed in the reversed analysis, choosing the point mass at such appmakes the expected gain underwtw\_\{t\}small, which bounds the one\-round minimax value\. The following assumption makes these requirements precise\.
###### Assumption 4\.1\(Regularity for the Upper Bound\)\.
There are constantsRmax,Cgrid<∞R\_\{\\max\},C\_\{\\mathrm\{grid\}\}<\\infty\. For each integerQ≥1Q\\geq 1, there is a finite grid𝒫Q⊂𝒫\\mathcal\{P\}\_\{Q\}\\subset\\mathcal\{P\}such that the following hold\.
1. \(i\)ℳ\\mathcal\{M\}is a convex compact subset of a topological vector space of finite signed measures on𝒴\\mathcal\{Y\}, and for everyp∈𝒫Qp\\in\\mathcal\{P\}\_\{Q\}and coordinatej∈\[k\]j\\in\[k\], the map μ↦Rj\(p<j,pj,μ\)\\mu\\mapsto R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\)is affine and continuous onℳ\\mathcal\{M\}\.
2. \(ii\)For everyp∈𝒫p\\in\\mathcal\{P\}andy∈𝒴y\\in\\mathcal\{Y\}, ‖R\(p,y\)‖∞≤Rmax\.\\\|R\(p,y\)\\\|\_\{\\infty\}\\leq R\_\{\\max\}\.
3. \(iii\)The grid size satisfies \|𝒫Q\|≤CgridQk,\|\\mathcal\{P\}\_\{Q\}\|\\leq C\_\{\\mathrm\{grid\}\}Q^\{k\},and the expected residual can be rounded at rate1/Q1/Q: ΔQ:=supμ∈ℳminp∈𝒫Q‖R\(p,μ\)‖∞≤CgridQ\.\\Delta\_\{Q\}:=\\sup\_\{\\mu\\in\\mathcal\{M\}\}\\min\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\\|R\(p,\\mu\)\\\|\_\{\\infty\}\\leq\\frac\{C\_\{\\mathrm\{grid\}\}\}\{Q\}\.
### 4\.3The Online Forecaster and Its Empirical Guarantee
Fix an integerQ≥1Q\\geq 1, and let𝒫Q\\mathcal\{P\}\_\{Q\}be the grid supplied by Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\. Algorithm[1](https://arxiv.org/html/2608.04288#alg1)specifies how the forecaster operates on each round\.
Algorithm 1Online Forecaster1:Horizon
TT, group family
𝒢\\mathcal\{G\}, grid
𝒫Q\\mathcal\{P\}\_\{Q\}, law class
ℳ\\mathcal\{M\}, level count
kk, additive optimization error
ρ\\rho
2:Set
ηEW←2log\|𝒢×Σ\|kRmaxT\.\\eta\_\{\\mathrm\{EW\}\}\\leftarrow\\frac\{\\sqrt\{2\\log\|\\mathcal\{G\}\\times\\Sigma\|\}\}\{kR\_\{\\max\}\\sqrt\{T\}\}\.
3:Initialize cumulative gains
C0\(g,s\)←0C\_\{0\}\(g,s\)\\leftarrow 0for all
\(g,s\)∈𝒢×Σ\(g,s\)\\in\\mathcal\{G\}\\times\\Sigma\.
4:for
t=1,…,Tt=1,\\ldots,Tdo
5:Compute
wt\(g,s\)←exp\{ηEWCt−1\(g,s\)\}∑\(g′,s′\)∈𝒢×Σexp\{ηEWCt−1\(g′,s′\)\}\.w\_\{t\}\(g,s\)\\leftarrow\\frac\{\\exp\\\{\\eta\_\{\\mathrm\{EW\}\}C\_\{t\-1\}\(g,s\)\\\}\}\{\\sum\_\{\(g^\{\\prime\},s^\{\\prime\}\)\\in\\mathcal\{G\}\\times\\Sigma\}\\exp\\\{\\eta\_\{\\mathrm\{EW\}\}C\_\{t\-1\}\(g^\{\\prime\},s^\{\\prime\}\)\\\}\}\.
6:Define
πt\(⋅∣x\)\\pi\_\{t\}\(\\cdot\\mid x\)pointwise for
x∈𝒳x\\in\\mathcal\{X\}to satisfy
maxμ∈ℳ𝔼p∼πt\(⋅∣x\)\[∑g∈𝒢∑s∈Σwt\(g,s\)g\(x\)∑j=1ksp,jRj\(p<j,pj,μ\)\]\\displaystyle\\max\_\{\\mu\\in\\mathcal\{M\}\}\\mathbb\{E\}\_\{p\\sim\\pi\_\{t\}\(\\cdot\\mid x\)\}\\left\[\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{s\\in\\Sigma\}w\_\{t\}\(g,s\)g\(x\)\\sum\_\{j=1\}^\{k\}s\_\{p,j\}R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\)\\right\]\(4\)≤minπ∈Δ\(𝒫Q\)maxμ∈ℳ𝔼p∼π\[∑g∈𝒢∑s∈Σwt\(g,s\)g\(x\)∑j=1ksp,jRj\(p<j,pj,μ\)\]\+ρ\.\\displaystyle\\qquad\\leq\\min\_\{\\pi\\in\\Delta\(\\mathcal\{P\}\_\{Q\}\)\}\\max\_\{\\mu\\in\\mathcal\{M\}\}\\mathbb\{E\}\_\{p\\sim\\pi\}\\left\[\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{s\\in\\Sigma\}w\_\{t\}\(g,s\)g\(x\)\\sum\_\{j=1\}^\{k\}s\_\{p,j\}R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\)\\right\]\+\\rho\.
7:Observe
XtX\_\{t\}and output
πt\(⋅∣Xt\)\\pi\_\{t\}\(\\cdot\\mid X\_\{t\}\)as the randomized prediction\.
8:Observe
YtY\_\{t\}\.
9:For every
\(g,s\)∈𝒢×Σ\(g,s\)\\in\\mathcal\{G\}\\times\\Sigma, set
ht\(g,s\)←g\(Xt\)𝔼p∼πt\(⋅∣Xt\)\[∑j=1ksp,jRj\(p<j,pj,Yt\)\]\.h\_\{t\}\(g,s\)\\leftarrow g\(X\_\{t\}\)\\mathbb\{E\}\_\{p\\sim\\pi\_\{t\}\(\\cdot\\mid X\_\{t\}\)\}\\left\[\\sum\_\{j=1\}^\{k\}s\_\{p,j\}R\_\{j\}\(p\_\{<j\},p\_\{j\},Y\_\{t\}\)\\right\]\.
10:Set
Ct\(g,s\)←Ct−1\(g,s\)\+ht\(g,s\)C\_\{t\}\(g,s\)\\leftarrow C\_\{t\-1\}\(g,s\)\+h\_\{t\}\(g,s\)for every
\(g,s\)∈𝒢×Σ\(g,s\)\\in\\mathcal\{G\}\\times\\Sigma\.
11:endfor
12:Prediction rules
π1,…,πT\\pi\_\{1\},\\ldots,\\pi\_\{T\}\.
The prediction rules in Algorithm[1](https://arxiv.org/html/2608.04288#alg1)can be chosen jointly measurable, as shown in Appendix[B](https://arxiv.org/html/2608.04288#A2)\.
The forecaster choosesπt\\pi\_\{t\}to keep the expected gain underwtw\_\{t\}small on each round, while the no\-regret guarantee of the exponential weights algorithm bounds the cumulative gain of every expert by the sum of these expected gains plus a sublinear regret term\. Since the empirical multicalibration error is the largest cumulative expert gain divided byTT, this gives the following guarantee\.
###### Theorem 4\.3\(Empirical Guarantee for the Online Forecaster\)\.
Suppose Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)holds\. Let\(Xt,Yt\)t=1T\(X\_\{t\},Y\_\{t\}\)\_\{t=1\}^\{T\}be any stochastic process such that, conditionally onℱt−1\\mathcal\{F\}\_\{t\-1\}andXtX\_\{t\}, the conditional law ofYtY\_\{t\}belongs toℳ\\mathcal\{M\}\. Then Algorithm[1](https://arxiv.org/html/2608.04288#alg1)satisfies
𝔼MCErr^TΓ≤ρ\+kΔQ\+CkRmaxlog\|𝒢\|\+k\|𝒫Q\|T,\\mathbb\{E\}\\widehat\{\\operatorname\{MCErr\}\}^\{\\Gamma\}\_\{T\}\\leq\\rho\+k\\Delta\_\{Q\}\+C\_\{k\}R\_\{\\max\}\\sqrt\{\\frac\{\\log\|\\mathcal\{G\}\|\+k\|\\mathcal\{P\}\_\{Q\}\|\}\{T\}\},whereCkC\_\{k\}depends only onkk\. Consequently, the two bounds in Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(iii\) give
𝔼MCErr^TΓ≤ρ\+O\(1Q\+RmaxQk\+log\|𝒢\|T\)\.\\mathbb\{E\}\\widehat\{\\operatorname\{MCErr\}\}^\{\\Gamma\}\_\{T\}\\leq\\rho\+O\\\!\\left\(\\frac\{1\}\{Q\}\+R\_\{\\max\}\\sqrt\{\\frac\{Q^\{k\}\+\\log\|\\mathcal\{G\}\|\}\{T\}\}\\right\)\.\(5\)The expectation is over the stochastic transcript and any internal randomization of the forecaster\.
We prove Theorem[4\.3](https://arxiv.org/html/2608.04288#S4.Thmtheorem3)in the next subsection\.
### 4\.4Analysis of the Online Forecaster
Throughout the analysis,πt\\pi\_\{t\}denotes theℱt−1\\mathcal\{F\}\_\{t\-1\}\-measurable rule selected at roundtt\. Since every group function takes values in\[0,1\]\[0,1\]and the residual bound in Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(ii\) gives‖R‖∞≤Rmax\\\|R\\\|\_\{\\infty\}\\leq R\_\{\\max\}, every signed gain satisfies
\|ht\(g,s\)\|≤kRmax\.\|h\_\{t\}\(g,s\)\|\\leq kR\_\{\\max\}\.By the signed\-objective representation of the empirical multicalibration error,
MCErr^TΓ=1Tmax\(g,s\)∈𝒢×Σ∑t=1Tht\(g,s\)\.\\widehat\{\\operatorname\{MCErr\}\}^\{\\Gamma\}\_\{T\}=\\frac\{1\}\{T\}\\max\_\{\(g,s\)\\in\\mathcal\{G\}\\times\\Sigma\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(g,s\)\.Thus controlling the gains of all experts\(g,s\)∈𝒢×Σ\(g,s\)\\in\\mathcal\{G\}\\times\\Sigmacontrols empiricalΓ\\Gamma\-ECE\.
#### 4\.4\.1Exponential Weights
The first step compares the largest cumulative signed residual with the sum of the expected gains under the distributionswtw\_\{t\}produced by exponential weights\.
###### Lemma 4\.4\(Exponential Weights Bound\)\.
For every realized transcript of Algorithm[1](https://arxiv.org/html/2608.04288#alg1),
max\(g,s\)∈𝒢×Σ∑t=1Tht\(g,s\)≤∑t=1T∑g∈𝒢∑s∈Σwt\(g,s\)ht\(g,s\)\+CkRmaxT\(log\|𝒢\|\+k\|𝒫Q\|\),\\max\_\{\(g,s\)\\in\\mathcal\{G\}\\times\\Sigma\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(g,s\)\\leq\\sum\_\{t=1\}^\{T\}\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{s\\in\\Sigma\}w\_\{t\}\(g,s\)h\_\{t\}\(g,s\)\+C\_\{k\}R\_\{\\max\}\\sqrt\{T\\bigl\(\\log\|\\mathcal\{G\}\|\+k\|\\mathcal\{P\}\_\{Q\}\|\\bigr\)\},whereCkC\_\{k\}depends only onkkandht\(g,s\)h\_\{t\}\(g,s\)is the signed gain from Algorithm[1](https://arxiv.org/html/2608.04288#alg1),
ht\(g,s\)=g\(Xt\)𝔼p∼πt\(⋅∣Xt\)\[∑j=1ksp,jRj\(p<j,pj,Yt\)\]\.h\_\{t\}\(g,s\)=g\(X\_\{t\}\)\\mathbb\{E\}\_\{p\\sim\\pi\_\{t\}\(\\cdot\\mid X\_\{t\}\)\}\\left\[\\sum\_\{j=1\}^\{k\}s\_\{p,j\}R\_\{j\}\(p\_\{<j\},p\_\{j\},Y\_\{t\}\)\\right\]\.
###### Proof\.
Using the learning rateηEW\\eta\_\{\\mathrm\{EW\}\}from Algorithm[1](https://arxiv.org/html/2608.04288#alg1), define the exponential weights potential
Wt:=∑g∈𝒢∑s∈Σexp\{ηEWCt−1\(g,s\)\}\.W\_\{t\}:=\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{s\\in\\Sigma\}\\exp\\\{\\eta\_\{\\mathrm\{EW\}\}C\_\{t\-1\}\(g,s\)\\\}\.SinceCt\(g,s\)=Ct−1\(g,s\)\+ht\(g,s\)C\_\{t\}\(g,s\)=C\_\{t\-1\}\(g,s\)\+h\_\{t\}\(g,s\), and since\|ht\(g,s\)\|≤kRmax\|h\_\{t\}\(g,s\)\|\\leq kR\_\{\\max\}, Hoeffding’s lemma\[[30](https://arxiv.org/html/2608.04288#bib.bib37)\]gives
logWt\+1Wt=log∑g∈𝒢∑s∈Σwt\(g,s\)exp\{ηEWht\(g,s\)\}≤ηEW∑g∈𝒢∑s∈Σwt\(g,s\)ht\(g,s\)\+ηEW2k2Rmax22\.\\log\\frac\{W\_\{t\+1\}\}\{W\_\{t\}\}=\\log\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{s\\in\\Sigma\}w\_\{t\}\(g,s\)\\exp\\\{\\eta\_\{\\mathrm\{EW\}\}h\_\{t\}\(g,s\)\\\}\\leq\\eta\_\{\\mathrm\{EW\}\}\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{s\\in\\Sigma\}w\_\{t\}\(g,s\)h\_\{t\}\(g,s\)\+\\frac\{\\eta\_\{\\mathrm\{EW\}\}^\{2\}k^\{2\}R\_\{\\max\}^\{2\}\}\{2\}\.Summing overttyields
logWT\+1≤log\|𝒢×Σ\|\+ηEW∑t=1T∑g∈𝒢∑s∈Σwt\(g,s\)ht\(g,s\)\+ηEW2k2Rmax2T2\.\\log W\_\{T\+1\}\\leq\\log\|\\mathcal\{G\}\\times\\Sigma\|\+\\eta\_\{\\mathrm\{EW\}\}\\sum\_\{t=1\}^\{T\}\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{s\\in\\Sigma\}w\_\{t\}\(g,s\)h\_\{t\}\(g,s\)\+\\frac\{\\eta\_\{\\mathrm\{EW\}\}^\{2\}k^\{2\}R\_\{\\max\}^\{2\}T\}\{2\}\.On the other hand, for every\(g,s\)∈𝒢×Σ\(g,s\)\\in\\mathcal\{G\}\\times\\Sigma,
logWT\+1≥ηEWCT\(g,s\)=ηEW∑t=1Tht\(g,s\)\.\\log W\_\{T\+1\}\\geq\\eta\_\{\\mathrm\{EW\}\}C\_\{T\}\(g,s\)=\\eta\_\{\\mathrm\{EW\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(g,s\)\.Combining the two inequalities and maximizing over\(g,s\)\(g,s\)gives
max\(g,s\)∈𝒢×Σ∑t=1Tht\(g,s\)≤∑t=1T∑g∈𝒢∑s∈Σwt\(g,s\)ht\(g,s\)\+log\|𝒢×Σ\|ηEW\+ηEWk2Rmax2T2\.\\max\_\{\(g,s\)\\in\\mathcal\{G\}\\times\\Sigma\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(g,s\)\\leq\\sum\_\{t=1\}^\{T\}\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{s\\in\\Sigma\}w\_\{t\}\(g,s\)h\_\{t\}\(g,s\)\+\\frac\{\\log\|\\mathcal\{G\}\\times\\Sigma\|\}\{\\eta\_\{\\mathrm\{EW\}\}\}\+\\frac\{\\eta\_\{\\mathrm\{EW\}\}k^\{2\}R\_\{\\max\}^\{2\}T\}\{2\}\.The expert class satisfies
log\|𝒢×Σ\|=log\|𝒢\|\+k\|𝒫Q\|log2\.\\log\|\\mathcal\{G\}\\times\\Sigma\|=\\log\|\\mathcal\{G\}\|\+k\|\\mathcal\{P\}\_\{Q\}\|\\log 2\.Substituting this expression and the value ofηEW\\eta\_\{\\mathrm\{EW\}\}from Algorithm[1](https://arxiv.org/html/2608.04288#alg1)into the regret terms completes the proof\. ∎
#### 4\.4\.2The Minimax Value
Conditional on the history and the current context, the expected gain underwtw\_\{t\}is obtained by evaluating the objective in \([4](https://arxiv.org/html/2608.04288#S4.E4)\) at the conditional law of the outcome\. By the choice ofπt\\pi\_\{t\}, it is bounded by the worst\-case value of the one\-round minimax problem, up to the additive accuracyρ\\rho\. We next bound this value\.
###### Lemma 4\.5\(Upper Bound on the Minimax Value\)\.
Under Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1), the minimax term on the right\-hand side of \([4](https://arxiv.org/html/2608.04288#S4.E4)\) is at mostkΔQk\\Delta\_\{Q\}for every history and every contextxx\.
###### Proof\.
Fix the history and the contextxx, and letF\(π,μ\)F\(\\pi,\\mu\)denote the objective in \([4](https://arxiv.org/html/2608.04288#S4.E4)\)\. We first bound the value when the outcome law is chosen before the prediction:
maxμ∈ℳminπ∈Δ\(𝒫Q\)F\(π,μ\)\.\\max\_\{\\mu\\in\\mathcal\{M\}\}\\min\_\{\\pi\\in\\Delta\(\\mathcal\{P\}\_\{Q\}\)\}F\(\\pi,\\mu\)\.Fixμ∈ℳ\\mu\\in\\mathcal\{M\}\. Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(iii\) providespμ∈𝒫Qp\_\{\\mu\}\\in\\mathcal\{P\}\_\{Q\}such that
‖R\(pμ,μ\)‖∞≤ΔQ\.\\\|R\(p\_\{\\mu\},\\mu\)\\\|\_\{\\infty\}\\leq\\Delta\_\{Q\}\.Choose the point massπ=δpμ\\pi=\\delta\_\{p\_\{\\mu\}\}atpμp\_\{\\mu\}\. For every\(g,s\)\(g,s\), the bound0≤g\(x\)≤10\\leq g\(x\)\\leq 1gives
g\(x\)∑j=1kspμ,jRj\(\(pμ\)<j,\(pμ\)j,μ\)≤∑j=1k\|Rj\(\(pμ\)<j,\(pμ\)j,μ\)\|≤kΔQ\.g\(x\)\\sum\_\{j=1\}^\{k\}s\_\{p\_\{\\mu\},j\}R\_\{j\}\(\(p\_\{\\mu\}\)\_\{<j\},\(p\_\{\\mu\}\)\_\{j\},\\mu\)\\leq\\sum\_\{j=1\}^\{k\}\|R\_\{j\}\(\(p\_\{\\mu\}\)\_\{<j\},\(p\_\{\\mu\}\)\_\{j\},\\mu\)\|\\leq k\\Delta\_\{Q\}\.Taking the expectation underwtw\_\{t\}preserves the bound\. Since the argument holds for everyμ\\mu,
maxμ∈ℳminπ∈Δ\(𝒫Q\)F\(π,μ\)≤kΔQ\.\\max\_\{\\mu\\in\\mathcal\{M\}\}\\min\_\{\\pi\\in\\Delta\(\\mathcal\{P\}\_\{Q\}\)\}F\(\\pi,\\mu\)\\leq k\\Delta\_\{Q\}\.SinceFFis affine inπ\\pi, Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(i\) allows us to apply Sion’s minimax theorem\[[45](https://arxiv.org/html/2608.04288#bib.bib44)\]\. It identifies this value with the value in which the prediction is chosen first:
minπmaxμF\(π,μ\)=maxμminπF\(π,μ\)≤kΔQ\.\\min\_\{\\pi\}\\max\_\{\\mu\}F\(\\pi,\\mu\)=\\max\_\{\\mu\}\\min\_\{\\pi\}F\(\\pi,\\mu\)\\leq k\\Delta\_\{Q\}\.∎
The exponential weights lemma bounds the largest total gain of any expert by the sum of the expected gains underwtw\_\{t\}, together with a regret term\. Lemma[4\.5](https://arxiv.org/html/2608.04288#S4.Thmtheorem5)bounds each round’s expected gain after conditioning on the history and the current context\. We now combine the two bounds\.
###### Proof of Theorem[4\.3](https://arxiv.org/html/2608.04288#S4.Thmtheorem3)\.
Lemma[4\.4](https://arxiv.org/html/2608.04288#S4.Thmtheorem4)gives
𝔼\[max\(g,s\)∈𝒢×Σ∑t=1Tht\(g,s\)\]≤𝔼\[∑t=1T∑g∈𝒢∑s∈Σwt\(g,s\)ht\(g,s\)\]\+CkRmaxT\(log\|𝒢\|\+k\|𝒫Q\|\)\.\\mathbb\{E\}\\left\[\\max\_\{\(g,s\)\\in\\mathcal\{G\}\\times\\Sigma\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(g,s\)\\right\]\\leq\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{s\\in\\Sigma\}w\_\{t\}\(g,s\)h\_\{t\}\(g,s\)\\right\]\+C\_\{k\}R\_\{\\max\}\\sqrt\{T\\bigl\(\\log\|\\mathcal\{G\}\|\+k\|\\mathcal\{P\}\_\{Q\}\|\\bigr\)\}\.Condition onℱt−1\\mathcal\{F\}\_\{t\-1\}andXtX\_\{t\}, and letμt\\mu\_\{t\}be the conditional law ofYtY\_\{t\}\. By the hypothesis of Theorem[4\.3](https://arxiv.org/html/2608.04288#S4.Thmtheorem3),μt∈ℳ\\mu\_\{t\}\\in\\mathcal\{M\}\. The conditional expected gain underwtw\_\{t\}is the objective in \([4](https://arxiv.org/html/2608.04288#S4.E4)\) evaluated atπt\(⋅∣Xt\)\\pi\_\{t\}\(\\cdot\\mid X\_\{t\}\)andμt\\mu\_\{t\}\. It is therefore at most the worst\-case value forπt\(⋅∣Xt\)\\pi\_\{t\}\(\\cdot\\mid X\_\{t\}\)\. The choice ofπt\\pi\_\{t\}bounds this worst\-case value by the minimax value plusρ\\rho, and Lemma[4\.5](https://arxiv.org/html/2608.04288#S4.Thmtheorem5)bounds the minimax value bykΔQk\\Delta\_\{Q\}\. Hence
𝔼\[∑g∈𝒢∑s∈Σwt\(g,s\)ht\(g,s\)\|ℱt−1,Xt\]≤kΔQ\+ρ\.\\mathbb\{E\}\\left\[\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{s\\in\\Sigma\}w\_\{t\}\(g,s\)h\_\{t\}\(g,s\)\\middle\|\\mathcal\{F\}\_\{t\-1\},X\_\{t\}\\right\]\\leq k\\Delta\_\{Q\}\+\\rho\.Summing overttand taking expectations bounds the cumulative expected gain byT\(kΔQ\+ρ\)T\(k\\Delta\_\{Q\}\+\\rho\)\. Dividing the exponential weights bound byTTand using the representation at the beginning of this subsection proves the first bound in Theorem[4\.3](https://arxiv.org/html/2608.04288#S4.Thmtheorem3)\. The second follows from the two bounds in Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(iii\)\. ∎
### 4\.5The Averaged Batch Predictor and Sample Complexity
Let𝖣\\mathsf\{D\}be anℳ\\mathcal\{M\}\-compatible distribution, and letS=\(\(X1,Y1\),…,\(XT,YT\)\)∼𝖣TS=\(\(X\_\{1\},Y\_\{1\}\),\\ldots,\(X\_\{T\},Y\_\{T\}\)\)\\sim\\mathsf\{D\}^\{T\}\. We run Algorithm[1](https://arxiv.org/html/2608.04288#alg1)on the sample sequence and return the averaged predictorΠS:𝒳→Δ\(𝒫Q\)\\Pi\_\{S\}:\\mathcal\{X\}\\to\\Delta\(\\mathcal\{P\}\_\{Q\}\)defined by
ΠS\(p∣x\):=1T∑t=1Tπt\(p∣x\)\.\\Pi\_\{S\}\(p\\mid x\):=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\pi\_\{t\}\(p\\mid x\)\.\(6\)
#### 4\.5\.1Online\-to\-Batch Reduction
The population multicalibration error of the averaged predictor need not equal the empirical multicalibration error of the online transcript\. After expressing both errors as maxima over groups and sign arrays, we compare the corresponding population and empirical objectives\. As inCollinaet al\.\[[9](https://arxiv.org/html/2608.04288#bib.bib43)\], for each fixed group and sign array, the difference between these objectives is a martingale sum\. The following maximal inequality controls all these sums simultaneously\.
###### Lemma 4\.6\(A Maximal Inequality for Finitely Many Martingales\)\.
Let𝒜\\mathcal\{A\}be a finite set of sizeNN\. For eache∈𝒜e\\in\\mathcal\{A\}, let\(Mt\(e\)\)t=1T\(M\_\{t\}\(e\)\)\_\{t=1\}^\{T\}be martingale differences with respect to the same filtration, and assume\|Mt\(e\)\|≤G\|M\_\{t\}\(e\)\|\\leq Galmost surely for allt,et,e\. Then
𝔼maxe∈𝒜∑t=1TMt\(e\)≤G2TlogN\.\\mathbb\{E\}\\max\_\{e\\in\\mathcal\{A\}\}\\sum\_\{t=1\}^\{T\}M\_\{t\}\(e\)\\leq G\\sqrt\{2T\\log N\}\.
###### Proof\.
For anyη\>0\\eta\>0, conditional Hoeffding gives
𝔼\[exp\(η∑t=1TMt\(e\)\)\]≤exp\(η2G2T2\)\.\\mathbb\{E\}\\left\[\\exp\\left\(\\eta\\sum\_\{t=1\}^\{T\}M\_\{t\}\(e\)\\right\)\\right\]\\leq\\exp\\left\(\\frac\{\\eta^\{2\}G^\{2\}T\}\{2\}\\right\)\.Therefore
𝔼maxe∑tMt\(e\)≤1ηlog𝔼∑eexp\(η∑tMt\(e\)\)≤logNη\+ηG2T2\.\\mathbb\{E\}\\max\_\{e\}\\sum\_\{t\}M\_\{t\}\(e\)\\leq\\frac\{1\}\{\\eta\}\\log\\mathbb\{E\}\\sum\_\{e\}\\exp\\left\(\\eta\\sum\_\{t\}M\_\{t\}\(e\)\\right\)\\leq\\frac\{\\log N\}\{\\eta\}\+\\frac\{\\eta G^\{2\}T\}\{2\}\.Optimizing overη\\etaproves the bound\. ∎
We now apply Lemma[4\.6](https://arxiv.org/html/2608.04288#S4.Thmtheorem6)to compare the population multicalibration error of the averaged predictor with the empirical multicalibration error of the online transcript\.
###### Lemma 4\.7\(Online\-to\-Batch Transfer\)\.
Under Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1), for the averaged predictorΠS\\Pi\_\{S\}in \([6](https://arxiv.org/html/2608.04288#S4.E6)\),
𝔼\[MCErr𝖣Γ\(ΠS;𝒢\)\]≤𝔼MCErr^TΓ\+CkRmaxQk\+log\|𝒢\|T,\\mathbb\{E\}\\left\[\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi\_\{S\};\\mathcal\{G\}\)\\right\]\\leq\\mathbb\{E\}\\widehat\{\\operatorname\{MCErr\}\}^\{\\Gamma\}\_\{T\}\+C\_\{k\}R\_\{\\max\}\\sqrt\{\\frac\{Q^\{k\}\+\\log\|\\mathcal\{G\}\|\}\{T\}\},whereCkC\_\{k\}depends only onkk\. Both expectations are over the sample and any internal randomization used by the forecaster\.
###### Proof\.
For\(g,s\)∈𝒢×Σ\(g,s\)\\in\\mathcal\{G\}\\times\\Sigma, set
ℒg,s\(ΠS\):=∑p∈𝒫Q∑j=1ksp,jνg,j𝖣,ΠS\(\{p\}\)\.\\mathcal\{L\}\_\{g,s\}\(\\Pi\_\{S\}\):=\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\sum\_\{j=1\}^\{k\}s\_\{p,j\}\\nu^\{\\mathsf\{D\},\\Pi\_\{S\}\}\_\{g,j\}\(\\\{p\\\}\)\.Maximizing over eachsp,j∈\{±1\}s\_\{p,j\}\\in\\\{\\pm 1\\\}gives
MCErr𝖣Γ\(ΠS;𝒢\)=maxg∈𝒢maxs∈Σℒg,s\(ΠS\)\.\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi\_\{S\};\\mathcal\{G\}\)=\\max\_\{g\\in\\mathcal\{G\}\}\\max\_\{s\\in\\Sigma\}\\mathcal\{L\}\_\{g,s\}\(\\Pi\_\{S\}\)\.Its empirical counterpart is
ℒ^g,s\(S\):=1T∑t=1Tg\(Xt\)∑p∈𝒫Qπt\(p∣Xt\)∑j=1ksp,jRj\(p<j,pj,Yt\)\.\\widehat\{\\mathcal\{L\}\}\_\{g,s\}\(S\):=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}g\(X\_\{t\}\)\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\pi\_\{t\}\(p\\mid X\_\{t\}\)\\sum\_\{j=1\}^\{k\}s\_\{p,j\}R\_\{j\}\(p\_\{<j\},p\_\{j\},Y\_\{t\}\)\.Then
MCErr^TΓ=maxg,sℒ^g,s\(S\)\.\\widehat\{\\operatorname\{MCErr\}\}^\{\\Gamma\}\_\{T\}=\\max\_\{g,s\}\\widehat\{\\mathcal\{L\}\}\_\{g,s\}\(S\)\.Therefore
MCErr𝖣Γ\(ΠS;𝒢\)≤MCErr^TΓ\+maxg,s\(ℒg,s\(ΠS\)−ℒ^g,s\(S\)\)\.\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi\_\{S\};\\mathcal\{G\}\)\\leq\\widehat\{\\operatorname\{MCErr\}\}^\{\\Gamma\}\_\{T\}\+\\max\_\{g,s\}\\left\(\\mathcal\{L\}\_\{g,s\}\(\\Pi\_\{S\}\)\-\\widehat\{\\mathcal\{L\}\}\_\{g,s\}\(S\)\\right\)\.Fix\(g,s\)\(g,s\)\. For eachtt, let
Ztg,s:=g\(Xt\)∑p∈𝒫Qπt\(p∣Xt\)∑j=1ksp,jRj\(p<j,pj,Yt\)\.Z\_\{t\}^\{g,s\}:=g\(X\_\{t\}\)\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\pi\_\{t\}\(p\\mid X\_\{t\}\)\\sum\_\{j=1\}^\{k\}s\_\{p,j\}R\_\{j\}\(p\_\{<j\},p\_\{j\},Y\_\{t\}\)\.Define
Z¯tg,s:=𝔼\(X,Y\)∼𝖣\[g\(X\)∑p∈𝒫Qπt\(p∣X\)∑j=1ksp,jRj\(p<j,pj,Y\)\|ℱt−1\]\.\\bar\{Z\}\_\{t\}^\{g,s\}:=\\mathbb\{E\}\_\{\(X,Y\)\\sim\\mathsf\{D\}\}\\left\[g\(X\)\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\pi\_\{t\}\(p\\mid X\)\\sum\_\{j=1\}^\{k\}s\_\{p,j\}R\_\{j\}\(p\_\{<j\},p\_\{j\},Y\)\\middle\|\\mathcal\{F\}\_\{t\-1\}\\right\]\.Conditionally onℱt−1\\mathcal\{F\}\_\{t\-1\}, the ruleπt\\pi\_\{t\}is fixed\. The linearity ofℒg,s\\mathcal\{L\}\_\{g,s\}in the predictor and the definition ofΠS\\Pi\_\{S\}therefore give
ℒg,s\(ΠS\)=1T∑t=1TZ¯tg,s\.\\mathcal\{L\}\_\{g,s\}\(\\Pi\_\{S\}\)=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\bar\{Z\}\_\{t\}^\{g,s\}\.The empirical quantity satisfies
ℒ^g,s\(S\)=1T∑t=1TZtg,s\.\\widehat\{\\mathcal\{L\}\}\_\{g,s\}\(S\)=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}Z\_\{t\}^\{g,s\}\.Because the sample is i\.i\.d\.,Z¯tg,s\\bar\{Z\}\_\{t\}^\{g,s\}is the conditional expectation ofZtg,sZ\_\{t\}^\{g,s\}givenℱt−1\\mathcal\{F\}\_\{t\-1\}\. Hence the differences
Mtg,s:=Z¯tg,s−Ztg,sM\_\{t\}^\{g,s\}:=\\bar\{Z\}\_\{t\}^\{g,s\}\-Z\_\{t\}^\{g,s\}are martingale differences with respect to the filtration generated by the transcript and the forecaster’s internal randomization\. Since group functions take values in\[0,1\]\[0,1\], we have\|Ztg,s\|≤kRmax\|Z\_\{t\}^\{g,s\}\|\\leq kR\_\{\\max\}, and hence\|Mtg,s\|≤2kRmax\|M\_\{t\}^\{g,s\}\|\\leq 2kR\_\{\\max\}\. Subtracting the two equalities gives
ℒg,s\(ΠS\)−ℒ^g,s\(S\)=1T∑t=1TMtg,s\.\\mathcal\{L\}\_\{g,s\}\(\\Pi\_\{S\}\)\-\\widehat\{\\mathcal\{L\}\}\_\{g,s\}\(S\)=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}M\_\{t\}^\{g,s\}\.Lemma[4\.6](https://arxiv.org/html/2608.04288#S4.Thmtheorem6), applied to the\|𝒢\|\|Σ\|=\|𝒢\|2k\|𝒫Q\|\|\\mathcal\{G\}\|\|\\Sigma\|=\|\\mathcal\{G\}\|2^\{k\|\\mathcal\{P\}\_\{Q\}\|\}martingale\-difference sequences, gives
𝔼maxg,s∑t=1TMtg,s≤CkRmaxT\(log\|𝒢\|\+k\|𝒫Q\|\)\.\\mathbb\{E\}\\max\_\{g,s\}\\sum\_\{t=1\}^\{T\}M\_\{t\}^\{g,s\}\\leq C\_\{k\}R\_\{\\max\}\\sqrt\{T\\bigl\(\\log\|\\mathcal\{G\}\|\+k\|\\mathcal\{P\}\_\{Q\}\|\\bigr\)\}\.Dividing byTTproves the lemma\. ∎
#### 4\.5\.2Sample Complexity
Combining the empirical guarantee for the online forecaster with the online\-to\-batch transfer gives the sample\-complexity bound\.
###### Theorem 4\.8\(Sample Complexity under Approximate Minimax Optimization\)\.
Suppose Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)holds\. Let𝖣\\mathsf\{D\}be anyℳ\\mathcal\{M\}\-compatible distribution on𝒳×𝒴\\mathcal\{X\}\\times\\mathcal\{Y\}, and let𝒢\\mathcal\{G\}be any finite group family\. Fix a sufficiently small target errorε\>0\\varepsilon\>0and a minimax accuracy
0≤ρ<ε3\.0\\leq\\rho<\\frac\{\\varepsilon\}\{3\}\.There is a randomized learner using minimax accuracyρ\\rhoand a grid with
Q=Θ\(\(ε−3ρ\)−1\)Q=\\Theta\\\!\\left\(\(\\varepsilon\-3\\rho\)^\{\-1\}\\right\)which, from
n≥C\(\(ε−3ρ\)−\(k\+2\)\+\(ε−3ρ\)−2log\|𝒢\|\)n\\geq C\\left\(\(\\varepsilon\-3\\rho\)^\{\-\(k\+2\)\}\+\(\\varepsilon\-3\\rho\)^\{\-2\}\\log\|\\mathcal\{G\}\|\\right\)i\.i\.d\. samples from𝖣\\mathsf\{D\}, outputs a finite\-support predictorΠS:𝒳→Δ\(𝒫Q\)\\Pi\_\{S\}:\\mathcal\{X\}\\to\\Delta\(\\mathcal\{P\}\_\{Q\}\)satisfying
ℙ\[MCErr𝖣Γ\(ΠS;𝒢\)≤ε\]≥23\.\\mathbb\{P\}\\left\[\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi\_\{S\};\\mathcal\{G\}\)\\leq\\varepsilon\\right\]\\geq\\frac\{2\}\{3\}\.The probability is over the sample and any internal randomization of the learner, andCCdepends only onkk,RmaxR\_\{\\max\}, andCgridC\_\{\\mathrm\{grid\}\}\.
###### Proof of Theorem[4\.8](https://arxiv.org/html/2608.04288#S4.Thmtheorem8)\.
Run the online forecaster forT=nT=nrounds\. Because the sample is i\.i\.d\., the conditional law ofYtY\_\{t\}givenℱt−1\\mathcal\{F\}\_\{t\-1\}andXtX\_\{t\}is𝖣Y∣X=Xt\\mathsf\{D\}\_\{Y\\mid X=X\_\{t\}\}\. Theℳ\\mathcal\{M\}\-compatibility of𝖣\\mathsf\{D\}therefore ensures that the hypothesis of Theorem[4\.3](https://arxiv.org/html/2608.04288#S4.Thmtheorem3)holds\. Theorem[4\.3](https://arxiv.org/html/2608.04288#S4.Thmtheorem3)and Lemma[4\.7](https://arxiv.org/html/2608.04288#S4.Thmtheorem7)now give
𝔼MCErr𝖣Γ\(ΠS;𝒢\)≤ρ\+C′\(1Q\+RmaxQk\+log\|𝒢\|n\)\.\\mathbb\{E\}\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi\_\{S\};\\mathcal\{G\}\)\\leq\\rho\+C^\{\\prime\}\\left\(\\frac\{1\}\{Q\}\+R\_\{\\max\}\\sqrt\{\\frac\{Q^\{k\}\+\\log\|\\mathcal\{G\}\|\}\{n\}\}\\right\)\.Becauseε−3ρ\>0\\varepsilon\-3\\rho\>0, take
Q=⌈6C′ε−3ρ⌉\.Q=\\left\\lceil\\frac\{6C^\{\\prime\}\}\{\\varepsilon\-3\\rho\}\\right\\rceil\.If
n≥C\(\(ε−3ρ\)−\(k\+2\)\+\(ε−3ρ\)−2log\|𝒢\|\)n\\geq C\\left\(\(\\varepsilon\-3\\rho\)^\{\-\(k\+2\)\}\+\(\\varepsilon\-3\\rho\)^\{\-2\}\\log\|\\mathcal\{G\}\|\\right\)for a sufficiently large constantCC, then the expectation is at most
ρ\+ε−3ρ3=ε3\.\\rho\+\\frac\{\\varepsilon\-3\\rho\}\{3\}=\\frac\{\\varepsilon\}\{3\}\.Markov’s inequality therefore gives
ℙ\[MCErr𝖣Γ\(ΠS;𝒢\)\>ε\]≤13\.\\mathbb\{P\}\\left\[\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi\_\{S\};\\mathcal\{G\}\)\>\\varepsilon\\right\]\\leq\\frac\{1\}\{3\}\.∎
Ifρ≤ε/6\\rho\\leq\\varepsilon/6, thenε−3ρ≥ε/2\\varepsilon\-3\\rho\\geq\\varepsilon/2, so the theorem immediately gives the following bound\.
###### Corollary 4\.9\(Upper Bound on Sample Complexity\)\.
Under the conditions of Theorem[4\.8](https://arxiv.org/html/2608.04288#S4.Thmtheorem8), suppose each inner minimax problem is solved to additive accuracy
ρ≤ε6\.\\rho\\leq\\frac\{\\varepsilon\}\{6\}\.Then there is a randomized learner usingQ=Θ\(1/ε\)Q=\\Theta\(1/\\varepsilon\)and
n≥C\(ε−\(k\+2\)\+ε−2log\|𝒢\|\)n\\geq C\\left\(\\varepsilon^\{\-\(k\+2\)\}\+\\varepsilon^\{\-2\}\\log\|\\mathcal\{G\}\|\\right\)i\.i\.d\. samples from𝖣\\mathsf\{D\}whose output satisfies
ℙ\[MCErr𝖣Γ\(ΠS;𝒢\)≤ε\]≥23\.\\mathbb\{P\}\\left\[\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi\_\{S\};\\mathcal\{G\}\)\\leq\\varepsilon\\right\]\\geq\\frac\{2\}\{3\}\.In particular, if\|𝒢\|≤⌈ε−κ⌉\|\\mathcal\{G\}\|\\leq\\lceil\\varepsilon^\{\-\\kappa\}\\rceilfor a fixedκ\>0\\kappa\>0, then
n=O\(ε−\(k\+2\)\)\.n=O\(\\varepsilon^\{\-\(k\+2\)\}\)\.
Consequently,
SCΓ,ℳ\(κ\)\(ε\)=O\(ε−\(k\+2\)\)\.\\operatorname\{SC\}^\{\(\\kappa\)\}\_\{\\Gamma,\\mathcal\{M\}\}\(\\varepsilon\)=O\(\\varepsilon^\{\-\(k\+2\)\}\)\.Binary groups are included among the group functions allowed in the minimax definition\. Thus, if Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)also holds, Theorem[3\.2](https://arxiv.org/html/2608.04288#S3.Thmtheorem2)and Corollary[4\.9](https://arxiv.org/html/2608.04288#S4.Thmtheorem9)give
SCΓ,ℳ\(κ\)\(ε\)=Θ~\(ε−\(k\+2\)\)\.\\operatorname\{SC\}^\{\(\\kappa\)\}\_\{\\Gamma,\\mathcal\{M\}\}\(\\varepsilon\)=\\widetilde\{\\Theta\}\\\!\\left\(\\varepsilon^\{\-\(k\+2\)\}\\right\)\.
### 4\.6Polynomial\-Time Implementation
The exponential weights procedure analyzed above is information\-theoretic and maintains weights over\|𝒢\|2k\|𝒫Q\|\|\\mathcal\{G\}\|2^\{k\|\\mathcal\{P\}\_\{Q\}\|\}experts indexed by sign patterns\. This section gives a polynomial\-time implementation under an explicit linear optimization oracle overℳ\\mathcal\{M\}\.
All running\-time statements in this paper use a unit\-cost real\-arithmetic model\. Arithmetic operations, comparisons, evaluations of the elementary functions used in the implementation, and evaluations of the supplied group and residual functions are exact and have unit cost\. The running time also includes the operations performed by the linear optimization oracle, as specified in Assumption[4\.10](https://arxiv.org/html/2608.04288#S4.Thmtheorem10)\.
There are two computational tasks\. We first compute the exponential weights distribution without enumerating all sign arrays\. We then solve the minimax problem using linear optimization overℳ\\mathcal\{M\}\.
#### 4\.6\.1Computing Exponential Weights Implicitly
Let
CumResg,p,jt−1:=∑τ<tg\(Xτ\)πτ\(p∣Xτ\)Rj\(p<j,pj,Yτ\)\.\\operatorname\{CumRes\}^\{t\-1\}\_\{g,p,j\}:=\\sum\_\{\\tau<t\}g\(X\_\{\\tau\}\)\\pi\_\{\\tau\}\(p\\mid X\_\{\\tau\}\)R\_\{j\}\(p\_\{<j\},p\_\{j\},Y\_\{\\tau\}\)\.The definition of the gain gives
Ct−1\(g,s\)=∑τ<thτ\(g,s\)=∑p∈𝒫Q∑j=1ksp,jCumResg,p,jt−1\.C\_\{t\-1\}\(g,s\)=\\sum\_\{\\tau<t\}h\_\{\\tau\}\(g,s\)=\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\sum\_\{j=1\}^\{k\}s\_\{p,j\}\\operatorname\{CumRes\}^\{t\-1\}\_\{g,p,j\}\.Hence the exponential weights distribution over\(g,s\)∈𝒢×Σ\(g,s\)\\in\\mathcal\{G\}\\times\\Sigma, with learning rateηEW\>0\\eta\_\{\\mathrm\{EW\}\}\>0, has weight proportional to
exp\(ηEW∑p∈𝒫Q∑j=1ksp,jCumResg,p,jt−1\)\.\\exp\\left\(\\eta\_\{\\mathrm\{EW\}\}\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\sum\_\{j=1\}^\{k\}s\_\{p,j\}\\operatorname\{CumRes\}^\{t\-1\}\_\{g,p,j\}\\right\)\.For fixedgg, the exponent is a sum of terms that each depend on only one sign coordinate, so the conditional distribution of the signs factorizes\. Using
∑σ∈\{±1\}eησa=2cosh\(ηa\),\\sum\_\{\\sigma\\in\\\{\\pm 1\\\}\}e^\{\\eta\\sigma a\}=2\\cosh\(\\eta a\),the group marginal is
ωt\(g\)=∏p,j2cosh\(ηEWCumResg,p,jt−1\)∑g′∈𝒢∏p,j2cosh\(ηEWCumResg′,p,jt−1\),\\omega\_\{t\}\(g\)=\\frac\{\\prod\_\{p,j\}2\\cosh\\bigl\(\\eta\_\{\\mathrm\{EW\}\}\\operatorname\{CumRes\}^\{t\-1\}\_\{g,p,j\}\\bigr\)\}\{\\sum\_\{g^\{\\prime\}\\in\\mathcal\{G\}\}\\prod\_\{p,j\}2\\cosh\\bigl\(\\eta\_\{\\mathrm\{EW\}\}\\operatorname\{CumRes\}^\{t\-1\}\_\{g^\{\\prime\},p,j\}\\bigr\)\},and, conditional ongg, each sign has mean
mt,g,p,j:=𝔼\[sp,j∣g\]=tanh\(ηEWCumResg,p,jt−1\)\.m\_\{t,g,p,j\}:=\\mathbb\{E\}\[s\_\{p,j\}\\mid g\]=\\tanh\\bigl\(\\eta\_\{\\mathrm\{EW\}\}\\operatorname\{CumRes\}^\{t\-1\}\_\{g,p,j\}\\bigr\)\.Conditional ongg, the objective in \([4](https://arxiv.org/html/2608.04288#S4.E4)\) is linear in the signs, so its average over the signs depends only on the meansmt,g,p,jm\_\{t,g,p,j\}\. Averaging also over the group marginalωt\\omega\_\{t\}, the objective becomes
∑p∈𝒫Qπp∑j=1kat,p,j\(x\)Rj\(p<j,pj,μ\),\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\pi\_\{p\}\\sum\_\{j=1\}^\{k\}a\_\{t,p,j\}\(x\)R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\),where
at,p,j\(x\):=∑g∈𝒢ωt\(g\)g\(x\)mt,g,p,j\.a\_\{t,p,j\}\(x\):=\\sum\_\{g\\in\\mathcal\{G\}\}\\omega\_\{t\}\(g\)g\(x\)m\_\{t,g,p,j\}\.Thus it suffices to maintain theO\(k\|𝒢\|\|𝒫Q\|\)O\(k\|\\mathcal\{G\}\|\|\\mathcal\{P\}\_\{Q\}\|\)valuesCumResg,p,jt−1\\operatorname\{CumRes\}^\{t\-1\}\_\{g,p,j\}, rather than enumerate2k\|𝒫Q\|2^\{k\|\\mathcal\{P\}\_\{Q\}\|\}sign patterns\.
#### 4\.6\.2The Optimization Oracle
Once the coefficientsat,p,j\(x\)a\_\{t,p,j\}\(x\)have been computed, it remains to solve the minimax problem\.Noarovet al\.\[[37](https://arxiv.org/html/2608.04288#bib.bib33)\]solve an analogous problem for high\-dimensional unbiased prediction by expressing it as a finite\-dimensional linear program with continuously many constraints and using an approximate separation oracle\. Here the required approximate separation step reduces to linear optimization overℳ\\mathcal\{M\}\.
###### Assumption 4\.10\(Linear Optimization Oracle overℳ\\mathcal\{M\}\)\.
There is an oracle𝖮𝖯𝖳ℳ\\mathsf\{OPT\}\_\{\\mathcal\{M\}\}which, for every requested accuracyξ∈\(0,1\)\\xi\\in\(0,1\)and given coefficientscp,j∈ℝc\_\{p,j\}\\in\\mathbb\{R\}, returns a lawμ^∈ℳ\\widehat\{\\mu\}\\in\\mathcal\{M\}satisfying
∑p∈𝒫Q∑j=1kcp,jRj\(p<j,pj,μ^\)≥supμ∈ℳ∑p∈𝒫Q∑j=1kcp,jRj\(p<j,pj,μ\)−ξ\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\sum\_\{j=1\}^\{k\}c\_\{p,j\}R\_\{j\}\(p\_\{<j\},p\_\{j\},\\widehat\{\\mu\}\)\\geq\\sup\_\{\\mu\\in\\mathcal\{M\}\}\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\sum\_\{j=1\}^\{k\}c\_\{p,j\}R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\)\\;\-\\xiusing a number of operations polynomial in\|𝒫Q\|\|\\mathcal\{P\}\_\{Q\}\|andlog\(1/ξ\)\\log\(1/\\xi\)\. Within the same bound, the objective value and the residual expectations atμ^\\widehat\{\\mu\}can be evaluated to additive accuracyξ\\xi\.
For the computational implementation, fix one deterministic, Borel\-measurable realization of this oracle, including deterministic tie\-breaking whenever more than one admissible output is available\.
Given the coefficientsat,p,j\(x\)a\_\{t,p,j\}\(x\), the inner minimax problem is the semi\-infinite linear program
minπ,ζ\\displaystyle\\min\_\{\\pi,\\zeta\}ζ\\displaystyle\\zetas\.t\.π∈Δ\(𝒫Q\),\\displaystyle\\pi\\in\\Delta\(\\mathcal\{P\}\_\{Q\}\),∑p,jπpat,p,j\(x\)Rj\(p<j,pj,μ\)≤ζ∀μ∈ℳ\.\\displaystyle\\sum\_\{p,j\}\\pi\_\{p\}a\_\{t,p,j\}\(x\)R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\)\\leq\\zeta\\qquad\\forall\\mu\\in\\mathcal\{M\}\.For a candidate\(π,ζ\)\(\\pi,\\zeta\), a violated constraint is found by calling𝖮𝖯𝖳ℳ\\mathsf\{OPT\}\_\{\\mathcal\{M\}\}with coefficients
cp,j=πpat,p,j\(x\)\.c\_\{p,j\}=\\pi\_\{p\}a\_\{t,p,j\}\(x\)\.To see how the two optimization accuracies enter, call the oracle with internal accuracyξ\\xi, letμ^\\widehat\{\\mu\}be the returned law, and evaluate its objective to the same accuracy\. If the resulting estimate exceedsζ\+ξ\\zeta\+\\xi, then the constraint associated withμ^\\widehat\{\\mu\}is violated\. Otherwise, the oracle and evaluation guarantees certify that every constraint holds with additive slack at most3ξ3\\xi\. Since\|at,p,j\(x\)\|≤1\|a\_\{t,p,j\}\(x\)\|\\leq 1and the residuals are bounded,ζ\\zetamay be restricted to a bounded interval\. Thus the linear program has a bounded finite\-dimensional feasible region, and the oracle supplies approximate separation for its continuously many constraints\. Given a target minimax accuracyρ∈\(0,1\)\\rho\\in\(0,1\), the weak ellipsoid method\[[26](https://arxiv.org/html/2608.04288#bib.bib45)\]chooses its internal tolerances, includingξ\\xi, withlog\(1/ξ\)\\log\(1/\\xi\)polynomial in\|𝒫Q\|\|\\mathcal\{P\}\_\{Q\}\|andlog\(1/ρ\)\\log\(1/\\rho\)\. It returns a distributionπt\(⋅∣x\)\\pi\_\{t\}\(\\cdot\\mid x\)satisfying \([4](https://arxiv.org/html/2608.04288#S4.E4)\) using a number of operations and oracle calls polynomial in the same quantities\.
Fix a deterministic implementation of the weak ellipsoid method, including deterministic tie\-breaking, and combine it with the fixed oracle from Assumption[4\.10](https://arxiv.org/html/2608.04288#S4.Thmtheorem10)\. Denote the resulting solver by
𝖲𝗈𝗅𝗏𝖾ρ\(H,x\)∈Δ\(𝒫Q\),\\mathsf\{Solve\}\_\{\\rho\}\(H,x\)\\in\\Delta\(\\mathcal\{P\}\_\{Q\}\),whereH=\(CumResg,p,j\)g,p,jH=\(\\operatorname\{CumRes\}\_\{g,p,j\}\)\_\{g,p,j\}is the array of cumulative residuals available before the round\. The internal tolerances are chosen so that𝖲𝗈𝗅𝗏𝖾ρ\(H,x\)\\mathsf\{Solve\}\_\{\\rho\}\(H,x\)satisfies \([4](https://arxiv.org/html/2608.04288#S4.E4)\) with additive error at mostρ\\rho\. In the computational implementation, the rule selected at roundttis defined for every context by
πt\(⋅∣x\):=𝖲𝗈𝗅𝗏𝖾ρ\(\(CumResg,p,jt−1\)g,p,j,x\)\.\\pi\_\{t\}\(\\cdot\\mid x\):=\\mathsf\{Solve\}\_\{\\rho\}\\\!\\left\(\\bigl\(\\operatorname\{CumRes\}^\{t\-1\}\_\{g,p,j\}\\bigr\)\_\{g,p,j\},x\\right\)\.Thus the rule is fixed by the pre\-round state and the context, and the same rule can be reproduced after training\.
We can therefore compute the exponential weights distribution and solve the minimax problem in polynomial time; the following theorem records the resulting runtime guarantee\.
###### Theorem 4\.11\(Polynomial\-Time Implementation\)\.
Assume Assumptions[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)and[4\.10](https://arxiv.org/html/2608.04288#S4.Thmtheorem10)\. Then, for every minimax accuracyρ∈\(0,1\)\\rho\\in\(0,1\), the learner can be implemented with training time polynomial in
n,Qk,\|𝒢\|,log\(1/ρ\)\.n,\\quad Q^\{k\},\\quad\|\\mathcal\{G\}\|,\\quad\\log\(1/\\rho\)\.Choosingρ=ε/6\\rho=\\varepsilon/6and the remaining parameters as in Corollary[4\.9](https://arxiv.org/html/2608.04288#S4.Thmtheorem9)implements its learner\. In particular, for every fixedκ\>0\\kappa\>0, if\|𝒢\|≤ε−κ\|\\mathcal\{G\}\|\\leq\\varepsilon^\{\-\\kappa\}, then the learner runs in time polynomial in1/ε1/\\varepsilon\.
The output predictor has a representation of polynomial size, and for any queried contextxx, either its full prediction distribution or an exact sample from that distribution can be computed in time polynomial in the same parameters\.
###### Proof\.
The formulas forωt\(g\)\\omega\_\{t\}\(g\)andmt,g,p,jm\_\{t,g,p,j\}represent the exact exponential weights distribution over the sign\-pattern experts without enumerating them, while the formula forat,p,j\(x\)a\_\{t,p,j\}\(x\)computes the resulting coefficients in the minimax objective\. These calculations require storing only theO\(k\|𝒢\|\|𝒫Q\|\)O\(k\|\\mathcal\{G\}\|\|\\mathcal\{P\}\_\{Q\}\|\)valuesCumResg,p,jt−1\\operatorname\{CumRes\}^\{t\-1\}\_\{g,p,j\}\. At each round, the coefficientsat,p,j\(x\)a\_\{t,p,j\}\(x\)can be computed in time polynomial ink\|𝒢\|\|𝒫Q\|k\|\\mathcal\{G\}\|\|\\mathcal\{P\}\_\{Q\}\|for any queried contextxx\. During training, they need only be computed at the observed contextXtX\_\{t\}\. The inner minimax problem has\|𝒫Q\|\+1\|\\mathcal\{P\}\_\{Q\}\|\+1variables, and the approximate\-separation argument in the preceding subsection computes aρ\\rho\-approximate solution in the stated running time\. Theρ\\rho\-approximate minimax solution contributes the additiveρ\\rhoterm in \([5](https://arxiv.org/html/2608.04288#S4.E5)\)\. Finally, chooseρ=ε/6\\rho=\\varepsilon/6,Q=Θ\(1/ε\)Q=\\Theta\(1/\\varepsilon\), andn=O\(ε−\(k\+2\)\+ε−2log\|𝒢\|\)n=O\(\\varepsilon^\{\-\(k\+2\)\}\+\\varepsilon^\{\-2\}\\log\|\\mathcal\{G\}\|\)as in Corollary[4\.9](https://arxiv.org/html/2608.04288#S4.Thmtheorem9)\. These choices give the claimed polynomial runtime in1/ε1/\\varepsilon\.
To represent the output, store the fixed solver specification together with theTTpre\-round arrays
Ht−1:=\(CumResg,p,jt−1\)g,p,j,t∈\[T\]\.H\_\{t\-1\}:=\\bigl\(\\operatorname\{CumRes\}^\{t\-1\}\_\{g,p,j\}\\bigr\)\_\{g,p,j\},\\qquad t\\in\[T\]\.This usesO\(Tk\|𝒢\|\|𝒫Q\|\)O\(Tk\|\\mathcal\{G\}\|\|\\mathcal\{P\}\_\{Q\}\|\)real numbers\. For a queried contextxx, rerunning the same deterministic map𝖲𝗈𝗅𝗏𝖾ρ\(Ht−1,x\)\\mathsf\{Solve\}\_\{\\rho\}\(H\_\{t\-1\},x\)reproduces the ruleπt\(⋅∣x\)\\pi\_\{t\}\(\\cdot\\mid x\)used by the learner\. The full averaged distribution
ΠS\(⋅∣x\)=1T∑t=1T𝖲𝗈𝗅𝗏𝖾ρ\(Ht−1,x\)\\Pi\_\{S\}\(\\cdot\\mid x\)=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\mathsf\{Solve\}\_\{\\rho\}\(H\_\{t\-1\},x\)is obtained byTTsolver calls\. Alternatively, an exact sample fromΠS\(⋅∣x\)\\Pi\_\{S\}\(\\cdot\\mid x\)is obtained by drawingt∼Unif\{1,…,T\}t\\sim\\operatorname\{Unif\}\\\{1,\\ldots,T\\\}, computing𝖲𝗈𝗅𝗏𝖾ρ\(Ht−1,x\)\\mathsf\{Solve\}\_\{\\rho\}\(H\_\{t\-1\},x\), and sampling from that distribution\. Both the representation size and these evaluation procedures are polynomial innn,QkQ^\{k\},\|𝒢\|\|\\mathcal\{G\}\|, andlog\(1/ρ\)\\log\(1/\\rho\)\. ∎
## 5Canonical Instantiations
We now show that the general lower and upper bounds apply to three multilevel properties: mean and mean absolute deviation; mean, variance, and skewness; and quantile and CVaR\. In each case, they give matching sample\-complexity bounds up to logarithmic factors\. The lower bound uses a local family of outcome distributions, while the upper bound may hold over a larger class and uses an explicit optimization oracle\.
### 5\.1Mean, Mean Absolute Deviation
Letℳ\\mathcal\{M\}be the class of all probability laws on\[0,1\]\[0,1\], and use the prediction space
𝒫=\[0,1\]2\.\\mathcal\{P\}=\[0,1\]^\{2\}\.Forμ∈ℳ\\mu\\in\\mathcal\{M\}, define
m\(μ\):=𝔼μY,d\(μ\):=𝔼μ\|Y−m\(μ\)\|\.m\(\\mu\):=\\mathbb\{E\}\_\{\\mu\}Y,\\qquad d\(\\mu\):=\\mathbb\{E\}\_\{\\mu\}\|Y\-m\(\\mu\)\|\.The two\-level property is
Γ\(μ\)=\(m\(μ\),d\(μ\)\)\.\\Gamma\(\\mu\)=\(m\(\\mu\),d\(\\mu\)\)\.It is sequentially conditionally identifiable with residual functions
R1\(∅,m,y\)=m−y,R2\(m,d,y\)=d−\|y−m\|\.R\_\{1\}\(\\varnothing,m,y\)=m\-y,\\qquad R\_\{2\}\(m,d,y\)=d\-\|y\-m\|\.The second coordinate is the expectedℓ1\\ell\_\{1\}loss evaluated at the first coordinate\. It is therefore analogous to variance, but uses absolute loss rather than squared loss\. It is not a bilevel Bayes pair, since absolute loss elicits medians rather than means\.
We first construct a local witness family on the three\-point support
𝒴MAD=\{0,1/2,1\}\.\\mathcal\{Y\}\_\{\\mathrm\{MAD\}\}=\\\{0,1/2,1\\\}\.Letv∘=\(7/20,7/20\)v^\{\\circ\}=\(7/20,7/20\), and letℛMAD\\mathcal\{R\}\_\{\\mathrm\{MAD\}\}be a sufficiently small closed rectangle aroundv∘v^\{\\circ\}, contained in\(0,1/2\)×\(0,1\)\(0,1/2\)\\times\(0,1\)\. Forv=\(v1,v2\)∈ℛMADv=\(v\_\{1\},v\_\{2\}\)\\in\\mathcal\{R\}\_\{\\mathrm\{MAD\}\}, defineμv\\mu\_\{v\}by
ℙY∼μv\(Y=0\)=v22v1,\\mathbb\{P\}\_\{Y\\sim\\mu\_\{v\}\}\(Y=0\)=\\frac\{v\_\{2\}\}\{2v\_\{1\}\},ℙY∼μv\(Y=1/2\)=2\(1−v1\)−v2v1,\\mathbb\{P\}\_\{Y\\sim\\mu\_\{v\}\}\(Y=1/2\)=2\(1\-v\_\{1\}\)\-\\frac\{v\_\{2\}\}\{v\_\{1\}\},ℙY∼μv\(Y=1\)=2v1−1\+v22v1\.\\mathbb\{P\}\_\{Y\\sim\\mu\_\{v\}\}\(Y=1\)=2v\_\{1\}\-1\+\\frac\{v\_\{2\}\}\{2v\_\{1\}\}\.Atv∘v^\{\\circ\}, these probabilities are1/2,3/10,1/51/2,3/10,1/5\. Since they depend continuously onvv,ℛMAD\\mathcal\{R\}\_\{\\mathrm\{MAD\}\}can be chosen small enough that each is uniformly bounded below by a positive constant throughout the rectangle\. Sincev1<1/2v\_\{1\}<1/2onℛMAD\\mathcal\{R\}\_\{\\mathrm\{MAD\}\}, direct calculation gives
𝔼μvY=v1,𝔼μv\|Y−v1\|=v2\.\\mathbb\{E\}\_\{\\mu\_\{v\}\}Y=v\_\{1\},\\qquad\\mathbb\{E\}\_\{\\mu\_\{v\}\}\|Y\-v\_\{1\}\|=v\_\{2\}\.
###### Proposition 5\.1\(Mean, Mean Absolute Deviation: Witness\)\.
The family\{μv:v∈ℛMAD\}\\\{\\mu\_\{v\}:v\\in\\mathcal\{R\}\_\{\\mathrm\{MAD\}\}\\\}satisfies Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\.
###### Proof\.
We verified above thatΓ\(μv\)=v\\Gamma\(\\mu\_\{v\}\)=v, as required by Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(i\)\. The residuals evaluated at the true prefix are
R1\(v<1,q,μv\)=q−v1,R2\(v1,q,μv\)=q−v2\.R\_\{1\}\(v\_\{<1\},q,\\mu\_\{v\}\)=q\-v\_\{1\},\\qquad R\_\{2\}\(v\_\{1\},q,\\mu\_\{v\}\)=q\-v\_\{2\}\.Therefore conditions \(ii\) and \(iii\) in Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)hold\.
It remains to verify Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(iv\)\. The first coordinate has no preceding predictions\. For the second coordinate,
\|R2\(p1,p2,μv\)−R2\(v1,p2,μv\)\|\\displaystyle\\left\|R\_\{2\}\(p\_\{1\},p\_\{2\},\\mu\_\{v\}\)\-R\_\{2\}\(v\_\{1\},p\_\{2\},\\mu\_\{v\}\)\\right\|=\|𝔼μv\[\|Y−v1\|−\|Y−p1\|\]\|≤\|p1−v1\|,\\displaystyle\\qquad=\\left\|\\mathbb\{E\}\_\{\\mu\_\{v\}\}\\bigl\[\|Y\-v\_\{1\}\|\-\|Y\-p\_\{1\}\|\\bigr\]\\right\|\\leq\|p\_\{1\}\-v\_\{1\}\|,where the last step uses the reverse triangle inequality\. This verifies Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(iv\)\.
Finally, the probability vector ofμv\\mu\_\{v\}is a smooth function ofv=\(v1,v2\)v=\(v\_\{1\},v\_\{2\}\)onℛMAD\\mathcal\{R\}\_\{\\mathrm\{MAD\}\}, and all three coordinates are uniformly bounded below\. Hence, for some constantC<∞C<\\infty,
DKL\(μv∥μv′\)≤C‖v−v′‖22\.D\_\{\\mathrm\{KL\}\}\(\\mu\_\{v\}\\,\\\|\\,\\mu\_\{v^\{\\prime\}\}\)\\leq C\\\|v\-v^\{\\prime\}\\\|\_\{2\}^\{2\}\.This verifies Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(v\)\. ∎
###### Corollary 5\.2\(Mean, Mean Absolute Deviation: Lower Bound\)\.
Fixκ\>0\\kappa\>0\. There exist constantsc,C,ε0\>0c,C,\\varepsilon\_\{0\}\>0, depending only onκ\\kappa, such that the following holds for every0<ε≤ε00<\\varepsilon\\leq\\varepsilon\_\{0\}\. For the property
Γ=\(𝔼Y,𝔼\|Y−𝔼Y\|\),\\Gamma=\(\\mathbb\{E\}Y,\\mathbb\{E\}\|Y\-\\mathbb\{E\}Y\|\),one can construct a finite context space, a binary group family of size at mostε−κ\\varepsilon^\{\-\\kappa\}, and a finite collection of data distributions whose conditional outcome laws are supported on three points in\[0,1\]\[0,1\], such that every learner achievingΓ\\Gamma\-ECE at mostε\\varepsilonwith probability at least2/32/3must use
n≥cε−4log4\(C/ε\)\.n\\geq c\\frac\{\\varepsilon^\{\-4\}\}\{\\log^\{4\}\(C/\\varepsilon\)\}\.Equivalently,
SCΓ,ℳ\(κ\)\(ε\)=Ω~\(ε−4\)\.\\operatorname\{SC\}^\{\(\\kappa\)\}\_\{\\Gamma,\\mathcal\{M\}\}\(\\varepsilon\)=\\widetilde\{\\Omega\}\(\\varepsilon^\{\-4\}\)\.
###### Proof\.
Apply Theorem[3\.2](https://arxiv.org/html/2608.04288#S3.Thmtheorem2)using Proposition[5\.1](https://arxiv.org/html/2608.04288#S5.Thmtheorem1)\. ∎
We next verify the upper\-bound conditions overℳ\\mathcal\{M\}and give an exact optimization oracle\.
###### Proposition 5\.3\(Mean, Mean Absolute Deviation: Conditions for the Upper Bound\)\.
The regularity conditions in Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)hold overℳ\\mathcal\{M\}\. Moreover, the linear optimization oracle in Assumption[4\.10](https://arxiv.org/html/2608.04288#S4.Thmtheorem10)has an exact polynomial\-time implementation\.
###### Proof\.
The law classℳ\\mathcal\{M\}is convex and compact in the weak topology\. For each fixed prediction\(m,d\)\(m,d\), the maps
μ↦𝔼μ\[m−Y\],μ↦𝔼μ\[d−\|Y−m\|\]\\mu\\mapsto\\mathbb\{E\}\_\{\\mu\}\[m\-Y\],\\qquad\\mu\\mapsto\\mathbb\{E\}\_\{\\mu\}\\\!\\left\[d\-\|Y\-m\|\\right\]are affine and continuous\. These observations verify Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(i\)\. Condition \(ii\) also holds because both coordinates of the prediction and the outcome lie in\[0,1\]\[0,1\]\.
For two predictions\(m,d\),\(m′,d′\)\(m,d\),\(m^\{\\prime\},d^\{\\prime\}\),
\|R1\(∅,m,μ\)−R1\(∅,m′,μ\)\|=\|m−m′\|,\\left\|R\_\{1\}\(\\varnothing,m,\\mu\)\-R\_\{1\}\(\\varnothing,m^\{\\prime\},\\mu\)\\right\|=\|m\-m^\{\\prime\}\|,and
\|R2\(m,d,μ\)−R2\(m′,d′,μ\)\|≤\|d−d′\|\+\|m−m′\|\.\\left\|R\_\{2\}\(m,d,\\mu\)\-R\_\{2\}\(m^\{\\prime\},d^\{\\prime\},\\mu\)\\right\|\\leq\|d\-d^\{\\prime\}\|\+\|m\-m^\{\\prime\}\|\.ThusR\(⋅,μ\)R\(\\cdot,\\mu\)is uniformly Lipschitz\. A rectangular grid with meshO\(1/Q\)O\(1/Q\)therefore satisfies the required approximation bound and hasO\(Q2\)O\(Q^\{2\}\)points, verifying Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(iii\)\.
For the oracle, given coefficientscp,jc\_\{p,j\}, write each grid point asp=\(mp,dp\)p=\(m\_\{p\},d\_\{p\}\), and define
ϕ\(y\):=∑p∈𝒫Q\[cp,1\(mp−y\)\+cp,2\(dp−\|y−mp\|\)\]\.\\phi\(y\):=\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\left\[c\_\{p,1\}\(m\_\{p\}\-y\)\+c\_\{p,2\}\\bigl\(d\_\{p\}\-\|y\-m\_\{p\}\|\\bigr\)\\right\]\.Maximizing𝔼μϕ\(Y\)\\mathbb\{E\}\_\{\\mu\}\\phi\(Y\)over all probability laws on\[0,1\]\[0,1\]is the same as maximizingϕ\(y\)\\phi\(y\)overy∈\[0,1\]y\\in\[0,1\], because an optimizer may be taken to be a point mass at a maximizer\. The functionϕ\\phiis piecewise affine with breakpoints among the grid valuesmpm\_\{p\}, so a maximizer lies in\{0,1\}∪\{mp:p∈𝒫Q\}\\\{0,1\\\}\\cup\\\{m\_\{p\}:p\\in\\mathcal\{P\}\_\{Q\}\\\}\. Checking these points gives an exact polynomial\-time oracle\. ∎
###### Corollary 5\.4\(Mean, Mean Absolute Deviation: Upper Bound\)\.
Fixκ\>0\\kappa\>0\. There exist constantsC,ε0\>0C,\\varepsilon\_\{0\}\>0, depending only onκ\\kappa, such that the following holds for every0<ε≤ε00<\\varepsilon\\leq\\varepsilon\_\{0\}\. Let
Γ=\(𝔼Y,𝔼\|Y−𝔼Y\|\)\\Gamma=\(\\mathbb\{E\}Y,\\mathbb\{E\}\|Y\-\\mathbb\{E\}Y\|\)be the property on\[0,1\]\[0,1\]\. For every data distribution𝖣\\mathsf\{D\}on𝒳×\[0,1\]\\mathcal\{X\}\\times\[0,1\]and every finite group family satisfying\|𝒢\|≤⌈ε−κ⌉\|\\mathcal\{G\}\|\\leq\\lceil\\varepsilon^\{\-\\kappa\}\\rceil, there is a randomized learner using at mostCε−4C\\varepsilon^\{\-4\}samples whose outputΠS\\Pi\_\{S\}satisfies
ℙ\[MCErr𝖣Γ\(ΠS;𝒢\)≤ε\]≥23\.\\mathbb\{P\}\\left\[\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi\_\{S\};\\mathcal\{G\}\)\\leq\\varepsilon\\right\]\\geq\\frac\{2\}\{3\}\.The exact oracle in Proposition[5\.3](https://arxiv.org/html/2608.04288#S5.Thmtheorem3)gives training time polynomial in1/ε1/\\varepsilon\.
###### Proof\.
Apply Corollary[4\.9](https://arxiv.org/html/2608.04288#S4.Thmtheorem9)and Theorem[4\.11](https://arxiv.org/html/2608.04288#S4.Thmtheorem11)using Proposition[5\.3](https://arxiv.org/html/2608.04288#S5.Thmtheorem3)\. ∎
### 5\.2Mean, Variance, Skewness
Let
𝒴MVS=\{0,1/3,2/3,1\}\.\\mathcal\{Y\}\_\{\\mathrm\{MVS\}\}=\\\{0,1/3,2/3,1\\\}\.Fixh∈\(0,1/4\)h\\in\(0,1/4\), and letℳh\\mathcal\{M\}\_\{h\}be the set of probability laws on𝒴MVS\\mathcal\{Y\}\_\{\\mathrm\{MVS\}\}that assign mass at leasthhto every atom\. This class is convex and compact\. Moreover, everyμ∈ℳh\\mu\\in\\mathcal\{M\}\_\{h\}has variance bounded below:
Varμ\(Y\)=12∑y,y′∈𝒴MVSμ\(y\)μ\(y′\)\(y−y′\)2≥μ\(0\)μ\(1\)≥h2\.\\operatorname\{Var\}\_\{\\mu\}\(Y\)=\\frac\{1\}\{2\}\\sum\_\{y,y^\{\\prime\}\\in\\mathcal\{Y\}\_\{\\mathrm\{MVS\}\}\}\\mu\(y\)\\mu\(y^\{\\prime\}\)\(y\-y^\{\\prime\}\)^\{2\}\\geq\\mu\(0\)\\mu\(1\)\\geq h^\{2\}\.Forμ∈ℳh\\mu\\in\\mathcal\{M\}\_\{h\}, define
m\(μ\):=𝔼μY,σ2\(μ\):=𝔼μ\(Y−m\(μ\)\)2,m\(\\mu\):=\\mathbb\{E\}\_\{\\mu\}Y,\\qquad\\sigma^\{2\}\(\\mu\):=\\mathbb\{E\}\_\{\\mu\}\(Y\-m\(\\mu\)\)^\{2\},and the standardized skewness
γ\(μ\):=𝔼μ\(Y−m\(μ\)\)3\(σ2\(μ\)\)3/2\.\\gamma\(\\mu\):=\\frac\{\\mathbb\{E\}\_\{\\mu\}\(Y\-m\(\\mu\)\)^\{3\}\}\{\\bigl\(\\sigma^\{2\}\(\\mu\)\\bigr\)^\{3/2\}\}\.The three\-level property is
Γ\(μ\)=\(m\(μ\),σ2\(μ\),γ\(μ\)\)\.\\Gamma\(\\mu\)=\(m\(\\mu\),\\sigma^\{2\}\(\\mu\),\\gamma\(\\mu\)\)\.Since\|Y−m\(μ\)\|≤1\|Y\-m\(\\mu\)\|\\leq 1andσ2\(μ\)≥h2\\sigma^\{2\}\(\\mu\)\\geq h^\{2\}, its range lies in
𝒫:=\[0,1\]×\[h2,1/4\]×\[−h−3,h−3\]\.\\mathcal\{P\}:=\[0,1\]\\times\[h^\{2\},1/4\]\\times\[\-h^\{\-3\},h^\{\-3\}\]\.It is sequentially conditionally identifiable with residual functions
R1\(∅,m,y\)=m−y,R\_\{1\}\(\\varnothing,m,y\)=m\-y,R2\(m,σ2,y\)=σ2−\(y−m\)2,R\_\{2\}\(m,\\sigma^\{2\},y\)=\\sigma^\{2\}\-\(y\-m\)^\{2\},R3\(\(m,σ2\),γ,y\)=γ\(σ2\)3/2−\(y−m\)3\.R\_\{3\}\(\(m,\\sigma^\{2\}\),\\gamma,y\)=\\gamma\\,\(\\sigma^\{2\}\)^\{3/2\}\-\(y\-m\)^\{3\}\.The third residual depends on both preceding coordinates and identifies skewness only after the mean and variance have been fixed\.
We now construct a local three\-dimensional witness family\. Letμ∘\\mu^\{\\circ\}be the uniform distribution on𝒴MVS\\mathcal\{Y\}\_\{\\mathrm\{MVS\}\}, and let
v∘:=Γ\(μ∘\)=\(12,536,0\)\.v^\{\\circ\}:=\\Gamma\(\\mu^\{\\circ\}\)=\\left\(\\frac\{1\}\{2\},\\frac\{5\}\{36\},0\\right\)\.For a target vectorv=\(v1,v2,v3\)v=\(v\_\{1\},v\_\{2\},v\_\{3\}\)withv2\>0v\_\{2\}\>0, define the corresponding raw moments
r1\(v\):=v1,r2\(v\):=v12\+v2,r3\(v\):=v13\+3v1v2\+v3v23/2\.r\_\{1\}\(v\):=v\_\{1\},\\qquad r\_\{2\}\(v\):=v\_\{1\}^\{2\}\+v\_\{2\},\\qquad r\_\{3\}\(v\):=v\_\{1\}^\{3\}\+3v\_\{1\}v\_\{2\}\+v\_\{3\}v\_\{2\}^\{3/2\}\.Define weights on𝒴MVS\\mathcal\{Y\}\_\{\\mathrm\{MVS\}\}by
μv\(0\)\\displaystyle\\mu\_\{v\}\(0\):=1−112r1\(v\)\+9r2\(v\)−92r3\(v\),\\displaystyle=1\-\\frac\{11\}\{2\}r\_\{1\}\(v\)\+9r\_\{2\}\(v\)\-\\frac\{9\}\{2\}r\_\{3\}\(v\),μv\(1/3\)\\displaystyle\\mu\_\{v\}\(1/3\):=9r1\(v\)−452r2\(v\)\+272r3\(v\),\\displaystyle=9r\_\{1\}\(v\)\-\\frac\{45\}\{2\}r\_\{2\}\(v\)\+\\frac\{27\}\{2\}r\_\{3\}\(v\),μv\(2/3\)\\displaystyle\\mu\_\{v\}\(2/3\):=−92r1\(v\)\+18r2\(v\)−272r3\(v\),\\displaystyle=\-\\frac\{9\}\{2\}r\_\{1\}\(v\)\+8r\_\{2\}\(v\)\-\\frac\{27\}\{2\}r\_\{3\}\(v\),μv\(1\)\\displaystyle\\mu\_\{v\}\(1\):=r1\(v\)−92r2\(v\)\+92r3\(v\)\.\\displaystyle=r\_\{1\}\(v\)\-\\frac\{9\}\{2\}r\_\{2\}\(v\)\+\\frac\{9\}\{2\}r\_\{3\}\(v\)\.A direct calculation gives
∑yμv\(y\)=1,∑yμv\(y\)yj=rj\(v\)forj∈\{1,2,3\}\.\\sum\_\{y\}\\mu\_\{v\}\(y\)=1,\\qquad\\sum\_\{y\}\\mu\_\{v\}\(y\)y^\{j\}=r\_\{j\}\(v\)\\quad\\text\{for \}j\\in\\\{1,2,3\\\}\.Atv=v∘v=v^\{\\circ\}, all four weights equal1/41/4\. Sinceh<1/4h<1/4, continuity gives a nondegenerate closed box
ℛMVS⊂int\(𝒫\)\\mathcal\{R\}\_\{\\mathrm\{MVS\}\}\\subset\\operatorname\{int\}\(\\mathcal\{P\}\)aroundv∘v^\{\\circ\}, small enough thatμv\(y\)≥h\\mu\_\{v\}\(y\)\\geq hfor everyv∈ℛMVSv\\in\\mathcal\{R\}\_\{\\mathrm\{MVS\}\}andy∈𝒴MVSy\\in\\mathcal\{Y\}\_\{\\mathrm\{MVS\}\}\. Thusμv∈ℳh\\mu\_\{v\}\\in\\mathcal\{M\}\_\{h\}throughout this box\. Expanding the centered moments gives
𝔼μv\(Y−v1\)2=v2,𝔼μv\(Y−v1\)3=v3v23/2,\\mathbb\{E\}\_\{\\mu\_\{v\}\}\(Y\-v\_\{1\}\)^\{2\}=v\_\{2\},\\qquad\\mathbb\{E\}\_\{\\mu\_\{v\}\}\(Y\-v\_\{1\}\)^\{3\}=v\_\{3\}v\_\{2\}^\{3/2\},and hence
Γ\(μv\)=v∀v∈ℛMVS\.\\Gamma\(\\mu\_\{v\}\)=v\\qquad\\forall v\\in\\mathcal\{R\}\_\{\\mathrm\{MVS\}\}\.
###### Proposition 5\.5\(Mean, Variance, Skewness: Witness\)\.
The family\{μv:v∈ℛMVS\}\\\{\\mu\_\{v\}:v\\in\\mathcal\{R\}\_\{\\mathrm\{MVS\}\}\\\}satisfies Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\.
###### Proof\.
The identityΓ\(μv\)=v\\Gamma\(\\mu\_\{v\}\)=vholds by construction, as required by Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(i\)\. The residuals evaluated at the true prefix are
R1\(v<1,q,μv\)=q−v1,R2\(v1,q,μv\)=q−v2,R\_\{1\}\(v\_\{<1\},q,\\mu\_\{v\}\)=q\-v\_\{1\},\\qquad R\_\{2\}\(v\_\{1\},q,\\mu\_\{v\}\)=q\-v\_\{2\},and
R3\(\(v1,v2\),q,μv\)=v23/2\(q−v3\)\.R\_\{3\}\(\(v\_\{1\},v\_\{2\}\),q,\\mu\_\{v\}\)=v\_\{2\}^\{3/2\}\(q\-v\_\{3\}\)\.SinceℛMVS\\mathcal\{R\}\_\{\\mathrm\{MVS\}\}is compact andv2≥h2v\_\{2\}\\geq h^\{2\}, the coefficients11andv23/2v\_\{2\}^\{3/2\}are uniformly bounded above and away from zero\. Therefore conditions \(ii\) and \(iii\) in Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)hold\.
It remains to verify Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(iv\)\. The first coordinate has no preceding predictions\. For the second coordinate,
R2\(p1,p2,μv\)−R2\(v1,p2,μv\)=𝔼μv\[\(v1−Y\)2−\(p1−Y\)2\]\.R\_\{2\}\(p\_\{1\},p\_\{2\},\\mu\_\{v\}\)\-R\_\{2\}\(v\_\{1\},p\_\{2\},\\mu\_\{v\}\)=\\mathbb\{E\}\_\{\\mu\_\{v\}\}\\\!\\left\[\(v\_\{1\}\-Y\)^\{2\}\-\(p\_\{1\}\-Y\)^\{2\}\\right\]\.Sincep1,v1,Y∈\[0,1\]p\_\{1\},v\_\{1\},Y\\in\[0,1\], the above equation gives
\|R2\(p1,p2,μv\)−R2\(v1,p2,μv\)\|≤2\|p1−v1\|\.\\left\|R\_\{2\}\(p\_\{1\},p\_\{2\},\\mu\_\{v\}\)\-R\_\{2\}\(v\_\{1\},p\_\{2\},\\mu\_\{v\}\)\\right\|\\leq 2\|p\_\{1\}\-v\_\{1\}\|\.For the third coordinate,
\|R3\(\(p1,p2\),p3,μv\)−R3\(\(v1,v2\),p3,μv\)\|\\displaystyle\\left\|R\_\{3\}\(\(p\_\{1\},p\_\{2\}\),p\_\{3\},\\mu\_\{v\}\)\-R\_\{3\}\(\(v\_\{1\},v\_\{2\}\),p\_\{3\},\\mu\_\{v\}\)\\right\|≤\|p3\|\|p23/2−v23/2\|\+𝔼μv\|\(Y−p1\)3−\(Y−v1\)3\|\.\\displaystyle\\qquad\\leq\|p\_\{3\}\|\\,\|p\_\{2\}^\{3/2\}\-v\_\{2\}^\{3/2\}\|\+\\mathbb\{E\}\_\{\\mu\_\{v\}\}\\left\|\(Y\-p\_\{1\}\)^\{3\}\-\(Y\-v\_\{1\}\)^\{3\}\\right\|\.On𝒫\\mathcal\{P\},\|p3\|≤h−3\|p\_\{3\}\|\\leq h^\{\-3\}, the mapu↦u3/2u\\mapsto u^\{3/2\}is Lipschitz on\[h2,1/4\]\[h^\{2\},1/4\], and\(y,a\)↦\(y−a\)3\(y,a\)\\mapsto\(y\-a\)^\{3\}is Lipschitz fory,a∈\[0,1\]y,a\\in\[0,1\]\. Therefore
\|R3\(\(p1,p2\),p3,μv\)−R3\(\(v1,v2\),p3,μv\)\|≤C\(\|p1−v1\|\+\|p2−v2\|\),\\left\|R\_\{3\}\(\(p\_\{1\},p\_\{2\}\),p\_\{3\},\\mu\_\{v\}\)\-R\_\{3\}\(\(v\_\{1\},v\_\{2\}\),p\_\{3\},\\mu\_\{v\}\)\\right\|\\leq C\\bigl\(\|p\_\{1\}\-v\_\{1\}\|\+\|p\_\{2\}\-v\_\{2\}\|\\bigr\),withCCdepending only onhh\. This verifies Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(iv\)\.
Finally, the explicit formulas above show thatv↦μvv\\mapsto\\mu\_\{v\}is Lipschitz onℛMVS\\mathcal\{R\}\_\{\\mathrm\{MVS\}\}\. Sinceμv′\(y\)≥h\\mu\_\{v^\{\\prime\}\}\(y\)\\geq hfor every atom, the inequalitylogu≤u−1\\log u\\leq u\-1gives
DKL\(μv∥μv′\)\\displaystyle D\_\{\\mathrm\{KL\}\}\(\\mu\_\{v\}\\,\\\|\\,\\mu\_\{v^\{\\prime\}\}\)≤∑y∈𝒴MVS\(μv\(y\)−μv′\(y\)\)2μv′\(y\)\\displaystyle\\leq\\sum\_\{y\\in\\mathcal\{Y\}\_\{\\mathrm\{MVS\}\}\}\\frac\{\\bigl\(\\mu\_\{v\}\(y\)\-\\mu\_\{v^\{\\prime\}\}\(y\)\\bigr\)^\{2\}\}\{\\mu\_\{v^\{\\prime\}\}\(y\)\}≤1h∑y∈𝒴MVS\(μv\(y\)−μv′\(y\)\)2\\displaystyle\\leq\\frac\{1\}\{h\}\\sum\_\{y\\in\\mathcal\{Y\}\_\{\\mathrm\{MVS\}\}\}\\bigl\(\\mu\_\{v\}\(y\)\-\\mu\_\{v^\{\\prime\}\}\(y\)\\bigr\)^\{2\}≤C‖v−v′‖22\.\\displaystyle\\leq C\\\|v\-v^\{\\prime\}\\\|\_\{2\}^\{2\}\.This verifies Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(v\)\. ∎
###### Corollary 5\.6\(Mean, Variance, Skewness: Lower Bound\)\.
Fixh∈\(0,1/4\)h\\in\(0,1/4\)andκ\>0\\kappa\>0\. There exist constantsc,C,ε0\>0c,C,\\varepsilon\_\{0\}\>0, depending only onhhandκ\\kappa, such that the following holds for every0<ε≤ε00<\\varepsilon\\leq\\varepsilon\_\{0\}\. For the property
Γ=\(mean,variance,skewness\),\\Gamma=\(\\text\{mean\},\\text\{variance\},\\text\{skewness\}\),one can construct a finite context space, a binary group family of size at mostε−κ\\varepsilon^\{\-\\kappa\}, and a finite collection of data distributions whose conditional outcome laws belong toℳh\\mathcal\{M\}\_\{h\}, such that every learner achievingΓ\\Gamma\-ECE at mostε\\varepsilonwith probability at least2/32/3must use
n≥cε−5log5\(C/ε\)\.n\\geq c\\frac\{\\varepsilon^\{\-5\}\}\{\\log^\{5\}\(C/\\varepsilon\)\}\.Equivalently,
SCΓ,ℳh\(κ\)\(ε\)=Ω~\(ε−5\)\.\\operatorname\{SC\}^\{\(\\kappa\)\}\_\{\\Gamma,\\mathcal\{M\}\_\{h\}\}\(\\varepsilon\)=\\widetilde\{\\Omega\}\(\\varepsilon^\{\-5\}\)\.
###### Proof\.
Apply Theorem[3\.2](https://arxiv.org/html/2608.04288#S3.Thmtheorem2)using Proposition[5\.5](https://arxiv.org/html/2608.04288#S5.Thmtheorem5)\. ∎
We next verify the upper\-bound conditions overℳh\\mathcal\{M\}\_\{h\}and give an exact optimization oracle\.
###### Proposition 5\.7\(Mean, Variance, Skewness: Conditions for the Upper Bound\)\.
The regularity conditions in Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)hold overℳh\\mathcal\{M\}\_\{h\}\. Moreover, the linear optimization oracle in Assumption[4\.10](https://arxiv.org/html/2608.04288#S4.Thmtheorem10)has an exact polynomial\-time implementation\.
###### Proof\.
The law classℳh\\mathcal\{M\}\_\{h\}is a convex, compact subset of the four\-dimensional probability simplex\. For each fixed predictionp=\(m,σ2,γ\)p=\(m,\\sigma^\{2\},\\gamma\), the maps
μ↦𝔼μ\[m−Y\],\\mu\\mapsto\\mathbb\{E\}\_\{\\mu\}\[m\-Y\],μ↦𝔼μ\[σ2−\(Y−m\)2\],\\mu\\mapsto\\mathbb\{E\}\_\{\\mu\}\\\!\\left\[\\sigma^\{2\}\-\(Y\-m\)^\{2\}\\right\],μ↦𝔼μ\[γ\(σ2\)3/2−\(Y−m\)3\]\\mu\\mapsto\\mathbb\{E\}\_\{\\mu\}\\\!\\left\[\\gamma\\,\(\\sigma^\{2\}\)^\{3/2\}\-\(Y\-m\)^\{3\}\\right\]are affine and continuous\. These observations verify Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(i\)\. Condition \(ii\) also holds because𝒫\\mathcal\{P\}and𝒴MVS\\mathcal\{Y\}\_\{\\mathrm\{MVS\}\}are compact\.
It remains to verify Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(iii\)\. For two predictionsp=\(m,σ2,γ\)p=\(m,\\sigma^\{2\},\\gamma\)andp′=\(m′,σ′2,γ′\)p^\{\\prime\}=\(m^\{\\prime\},\{\\sigma^\{\\prime\}\}^\{2\},\\gamma^\{\\prime\}\)in𝒫\\mathcal\{P\},
\|R1\(∅,m,μ\)−R1\(∅,m′,μ\)\|=\|m−m′\|,\\left\|R\_\{1\}\(\\varnothing,m,\\mu\)\-R\_\{1\}\(\\varnothing,m^\{\\prime\},\\mu\)\\right\|=\|m\-m^\{\\prime\}\|,\|R2\(m,σ2,μ\)−R2\(m′,σ′2,μ\)\|≤\|σ2−σ′2\|\+2\|m−m′\|,\\left\|R\_\{2\}\(m,\\sigma^\{2\},\\mu\)\-R\_\{2\}\(m^\{\\prime\},\{\\sigma^\{\\prime\}\}^\{2\},\\mu\)\\right\|\\leq\|\\sigma^\{2\}\-\{\\sigma^\{\\prime\}\}^\{2\}\|\+2\|m\-m^\{\\prime\}\|,and
\|R3\(\(m,σ2\),γ,μ\)−R3\(\(m′,σ′2\),γ′,μ\)\|≤Ch\(\|m−m′\|\+\|σ2−σ′2\|\+\|γ−γ′\|\),\\left\|R\_\{3\}\(\(m,\\sigma^\{2\}\),\\gamma,\\mu\)\-R\_\{3\}\(\(m^\{\\prime\},\{\\sigma^\{\\prime\}\}^\{2\}\),\\gamma^\{\\prime\},\\mu\)\\right\|\\leq C\_\{h\}\\bigl\(\|m\-m^\{\\prime\}\|\+\|\\sigma^\{2\}\-\{\\sigma^\{\\prime\}\}^\{2\}\|\+\|\\gamma\-\\gamma^\{\\prime\}\|\\bigr\),becauseγ\\gammais bounded,u↦u3/2u\\mapsto u^\{3/2\}is Lipschitz on\[h2,1/4\]\[h^\{2\},1/4\], and\(y,m\)↦\(y−m\)3\(y,m\)\\mapsto\(y\-m\)^\{3\}is Lipschitz on𝒴MVS×\[0,1\]\\mathcal\{Y\}\_\{\\mathrm\{MVS\}\}\\times\[0,1\]\. ThusR\(⋅,μ\)R\(\\cdot,\\mu\)is uniformly Lipschitz on𝒫\\mathcal\{P\}\. A rectangular grid with meshO\(1/Q\)O\(1/Q\)therefore satisfies the required approximation bound and hasO\(Q3\)O\(Q^\{3\}\)points\.
For the oracle, given coefficientscp,jc\_\{p,j\}, write each grid point asp=\(mp,σp2,γp\)p=\(m\_\{p\},\\sigma\_\{p\}^\{2\},\\gamma\_\{p\}\), and define
ϕ\(y\):=∑p∈𝒫Q\[cp,1\(mp−y\)\+cp,2\(σp2−\(y−mp\)2\)\+cp,3\(γp\(σp2\)3/2−\(y−mp\)3\)\]\.\\phi\(y\):=\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\left\[c\_\{p,1\}\(m\_\{p\}\-y\)\+c\_\{p,2\}\\bigl\(\\sigma\_\{p\}^\{2\}\-\(y\-m\_\{p\}\)^\{2\}\\bigr\)\+c\_\{p,3\}\\bigl\(\\gamma\_\{p\}\\,\(\\sigma\_\{p\}^\{2\}\)^\{3/2\}\-\(y\-m\_\{p\}\)^\{3\}\\bigr\)\\right\]\.The oracle problem is
supμ∈ℳh∑y∈𝒴MVSμ\(y\)ϕ\(y\)\.\\sup\_\{\\mu\\in\\mathcal\{M\}\_\{h\}\}\\sum\_\{y\\in\\mathcal\{Y\}\_\{\\mathrm\{MVS\}\}\}\\mu\(y\)\\phi\(y\)\.This is a linear optimization problem over the simplex with lower boundsμ\(y\)≥h\\mu\(y\)\\geq h\. An optimizer assigns masshhto each atom and the remaining mass1−4h1\-4hto an atom maximizingϕ\(y\)\\phi\(y\)\. Thus the oracle is exact and polynomial time\. ∎
###### Corollary 5\.8\(Mean, Variance, Skewness: Upper Bound\)\.
Fixh∈\(0,1/4\)h\\in\(0,1/4\)andκ\>0\\kappa\>0\. There exist constantsC,ε0\>0C,\\varepsilon\_\{0\}\>0, depending only onhhandκ\\kappa, such that the following holds for every0<ε≤ε00<\\varepsilon\\leq\\varepsilon\_\{0\}\. Let
Γ=\(mean,variance,skewness\)\\Gamma=\(\\text\{mean\},\\text\{variance\},\\text\{skewness\}\)be the property overℳh\\mathcal\{M\}\_\{h\}\. For everyℳh\\mathcal\{M\}\_\{h\}\-compatible data distribution𝖣\\mathsf\{D\}and every finite group family satisfying\|𝒢\|≤⌈ε−κ⌉\|\\mathcal\{G\}\|\\leq\\lceil\\varepsilon^\{\-\\kappa\}\\rceil, there is a randomized learner using at mostCε−5C\\varepsilon^\{\-5\}samples whose outputΠS\\Pi\_\{S\}satisfies
ℙ\[MCErr𝖣Γ\(ΠS;𝒢\)≤ε\]≥23\.\\mathbb\{P\}\\left\[\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi\_\{S\};\\mathcal\{G\}\)\\leq\\varepsilon\\right\]\\geq\\frac\{2\}\{3\}\.The exact oracle in Proposition[5\.7](https://arxiv.org/html/2608.04288#S5.Thmtheorem7)gives training time polynomial in1/ε1/\\varepsilon\.
###### Proof\.
Apply Corollary[4\.9](https://arxiv.org/html/2608.04288#S4.Thmtheorem9)and Theorem[4\.11](https://arxiv.org/html/2608.04288#S4.Thmtheorem11)using Proposition[5\.7](https://arxiv.org/html/2608.04288#S5.Thmtheorem7)\. ∎
### 5\.3Quantile, CVaR
Fixτ∈\(0,1\)\\tau\\in\(0,1\)and constants0<h<1<H<∞0<h<1<H<\\infty\. Letℳh,H\\mathcal\{M\}\_\{h,H\}be the class of probability laws on\[0,1\]\[0,1\]with densitiesffsatisfying
h≤f\(y\)≤Hfor almost everyy∈\[0,1\]\.h\\leq f\(y\)\\leq H\\qquad\\text\{for almost every \}y\\in\[0,1\]\.We use the prediction space
𝒫=\[0,1\]2\.\\mathcal\{P\}=\[0,1\]^\{2\}\.Forμ∈ℳh,H\\mu\\in\\mathcal\{M\}\_\{h,H\}, define theτ\\tau\-quantile
qτ\(μ\):=inf\{q:μ\(Y≤q\)≥τ\}\.q\_\{\\tau\}\(\\mu\):=\\inf\\\{q:\\mu\(Y\\leq q\)\\geq\\tau\\\}\.Use the upper\-tail convention
CVaRτ\(μ\):=qτ\(μ\)\+𝔼μ\(Y−qτ\(μ\)\)\+1−τ\.\\operatorname\{CVaR\}\_\{\\tau\}\(\\mu\):=q\_\{\\tau\}\(\\mu\)\+\\frac\{\\mathbb\{E\}\_\{\\mu\}\(Y\-q\_\{\\tau\}\(\\mu\)\)\_\{\+\}\}\{1\-\\tau\}\.Consider the scoring function
Lτ\(q,y\):=q\+\(y−q\)\+1−τ\.L\_\{\\tau\}\(q,y\):=q\+\\frac\{\(y\-q\)\_\{\+\}\}\{1\-\\tau\}\.Overℳh,H\\mathcal\{M\}\_\{h,H\},LτL\_\{\\tau\}is strictly consistent forqτq\_\{\\tau\}, and its Bayes risk isCVaRτ\\operatorname\{CVaR\}\_\{\\tau\}\[[38](https://arxiv.org/html/2608.04288#bib.bib31)\]\. Thus
\(qτ,CVaRτ\)\(q\_\{\\tau\},\\operatorname\{CVaR\}\_\{\\tau\}\)is a Bayes pair\. This is thek=2k=2case of the sequential framework with residual functions
R1\(∅,q,y\)=𝟏\{y≤q\}−τ,R2\(q,r,y\)=r−Lτ\(q,y\)\.R\_\{1\}\(\\varnothing,q,y\)=\\mathbf\{1\}\\\{y\\leq q\\\}\-\\tau,\\qquad R\_\{2\}\(q,r,y\)=r\-L\_\{\\tau\}\(q,y\)\.
We now construct a two\-dimensional local witness\. Forλ=\(λ1,λ2\)\\lambda=\(\\lambda\_\{1\},\\lambda\_\{2\}\)in a small neighborhood of0∈ℝ20\\in\\mathbb\{R\}^\{2\}, define the density
fλ\(y\)=exp\(λ1y\+λ2y2−A\(λ\)\),0≤y≤1,f\_\{\\lambda\}\(y\)=\\exp\(\\lambda\_\{1\}y\+\\lambda\_\{2\}y^\{2\}\-A\(\\lambda\)\),\\qquad 0\\leq y\\leq 1,whereA\(λ\)A\(\\lambda\)is the log normalizer\. Letμλ\\mu\_\{\\lambda\}be the corresponding distribution\. Sincef0≡1f\_\{0\}\\equiv 1, we may choose the neighborhood of0small enough thatμλ∈ℳh,H\\mu\_\{\\lambda\}\\in\\mathcal\{M\}\_\{h,H\}throughout it\.
Define
Ψ\(λ\):=\(qτ\(μλ\),CVaRτ\(μλ\)\)\.\\Psi\(\\lambda\):=\\left\(q\_\{\\tau\}\(\\mu\_\{\\lambda\}\),\\operatorname\{CVaR\}\_\{\\tau\}\(\\mu\_\{\\lambda\}\)\\right\)\.Atλ=0\\lambda=0,μ0=Unif\[0,1\]\\mu\_\{0\}=\\operatorname\{Unif\}\[0,1\], so
Ψ\(0\)=\(τ,1\+τ2\)\.\\Psi\(0\)=\\left\(\\tau,\\frac\{1\+\\tau\}\{2\}\\right\)\.
###### Lemma 5\.9\(Local Two\-Dimensional Parameterization\)\.
The JacobianDΨ\(0\)D\\Psi\(0\)is nonsingular\. Consequently, there exists a nondegenerate rectangle
ℛQC⊂range\(Ψ\)\\mathcal\{R\}\_\{\\mathrm\{QC\}\}\\subset\\operatorname\{range\}\(\\Psi\)around\(τ,\(1\+τ\)/2\)\\bigl\(\\tau,\(1\+\\tau\)/2\\bigr\)and a smooth inverse map
λ=λ\(v\),v=\(v1,v2\)∈ℛQC,\\lambda=\\lambda\(v\),\\qquad v=\(v\_\{1\},v\_\{2\}\)\\in\\mathcal\{R\}\_\{\\mathrm\{QC\}\},such that
qτ\(μλ\(v\)\)=v1,CVaRτ\(μλ\(v\)\)=v2\.q\_\{\\tau\}\(\\mu\_\{\\lambda\(v\)\}\)=v\_\{1\},\\qquad\\operatorname\{CVaR\}\_\{\\tau\}\(\\mu\_\{\\lambda\(v\)\}\)=v\_\{2\}\.
###### Proof\.
Sincefλ\(y\)f\_\{\\lambda\}\(y\)is smooth in\(λ,y\)\(\\lambda,y\)and strictly positive, the implicit equation∫0qτ\(μλ\)fλ\(y\)𝑑y=τ\\int\_\{0\}^\{q\_\{\\tau\}\(\\mu\_\{\\lambda\}\)\}f\_\{\\lambda\}\(y\)\\,dy=\\tauand the implicit function theorem show thatλ↦qτ\(μλ\)\\lambda\\mapsto q\_\{\\tau\}\(\\mu\_\{\\lambda\}\)is smooth near0\. The defining integral forCVaRτ\(μλ\)\\operatorname\{CVaR\}\_\{\\tau\}\(\\mu\_\{\\lambda\}\)is then smooth as well, soΨ\\Psiis smooth near0\.
Atλ=0\\lambda=0, direct differentiation gives
DΨ\(0\)=\(τ\(1−τ\)2τ\(1−τ2\)3\(1−τ\)\(2τ\+1\)12\(1−τ\)\(τ\+1\)212\)\.D\\Psi\(0\)=\\begin\{pmatrix\}\\dfrac\{\\tau\(1\-\\tau\)\}\{2\}&\\dfrac\{\\tau\(1\-\\tau^\{2\}\)\}\{3\}\\\\\[11\.00008pt\] \\dfrac\{\(1\-\\tau\)\(2\\tau\+1\)\}\{12\}&\\dfrac\{\(1\-\\tau\)\(\\tau\+1\)^\{2\}\}\{12\}\\end\{pmatrix\}\.Its determinant is
detDΨ\(0\)=τ\(1−τ\)3\(1\+τ\)72\>0\.\\det D\\Psi\(0\)=\\frac\{\\tau\(1\-\\tau\)^\{3\}\(1\+\\tau\)\}\{72\}\>0\.The inverse function theorem gives the claim\. ∎
Forv∈ℛQCv\\in\\mathcal\{R\}\_\{\\mathrm\{QC\}\}, define
μv:=μλ\(v\)\.\\mu\_\{v\}:=\\mu\_\{\\lambda\(v\)\}\.
It remains to check that the local parameterization from Lemma[5\.9](https://arxiv.org/html/2608.04288#S5.Thmtheorem9)has the residual and KL properties required by Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\.
###### Proposition 5\.10\(Quantile, CVaR: Witness\)\.
ForℛQC\\mathcal\{R\}\_\{\\mathrm\{QC\}\}sufficiently small, the family
\{μv:v∈ℛQC\}\\\{\\mu\_\{v\}:v\\in\\mathcal\{R\}\_\{\\mathrm\{QC\}\}\\\}is contained inℳh,H\\mathcal\{M\}\_\{h,H\}and satisfies Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)for the property
\(qτ,CVaRτ\)\.\(q\_\{\\tau\},\\operatorname\{CVaR\}\_\{\\tau\}\)\.
###### Proof\.
By Lemma[5\.9](https://arxiv.org/html/2608.04288#S5.Thmtheorem9),Γ\(μv\)=v\\Gamma\(\\mu\_\{v\}\)=v, as required by Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(i\)\.
By the choice of the parameter neighborhood,μv∈ℳh,H\\mu\_\{v\}\\in\\mathcal\{M\}\_\{h,H\}for everyv∈ℛQCv\\in\\mathcal\{R\}\_\{\\mathrm\{QC\}\}\. LetFvF\_\{v\}be the CDF ofμv\\mu\_\{v\}\. SinceFv\(v1\)=τF\_\{v\}\(v\_\{1\}\)=\\tau,
R1\(v<1,q,μv\)=Fv\(q\)−τ=Fv\(q\)−Fv\(v1\),R2\(v1,r,μv\)=r−v2\.R\_\{1\}\(v\_\{<1\},q,\\mu\_\{v\}\)=F\_\{v\}\(q\)\-\\tau=F\_\{v\}\(q\)\-F\_\{v\}\(v\_\{1\}\),\\qquad R\_\{2\}\(v\_\{1\},r,\\mu\_\{v\}\)=r\-v\_\{2\}\.Therefore,
sgn\(q−v1\)\(R1\(v<1,q,μv\)−R1\(v<1,v1,μv\)\)≥h\|q−v1\|\.\\operatorname\{sgn\}\(q\-v\_\{1\}\)\\left\(R\_\{1\}\(v\_\{<1\},q,\\mu\_\{v\}\)\-R\_\{1\}\(v\_\{<1\},v\_\{1\},\\mu\_\{v\}\)\\right\)\\geq h\|q\-v\_\{1\}\|\.The second residual evaluated at the true prefix satisfies
sgn\(r−v2\)\(R2\(v1,r,μv\)−R2\(v1,v2,μv\)\)=\|r−v2\|\.\\operatorname\{sgn\}\(r\-v\_\{2\}\)\\left\(R\_\{2\}\(v\_\{1\},r,\\mu\_\{v\}\)\-R\_\{2\}\(v\_\{1\},v\_\{2\},\\mu\_\{v\}\)\\right\)=\|r\-v\_\{2\}\|\.Thus Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(ii\) holds\. The bound in condition \(iii\) follows from
\|R1\(v<1,q,μv\)−R1\(v<1,v1,μv\)\|=\|Fv\(q\)−Fv\(v1\)\|≤H\|q−v1\|\\left\|R\_\{1\}\(v\_\{<1\},q,\\mu\_\{v\}\)\-R\_\{1\}\(v\_\{<1\},v\_\{1\},\\mu\_\{v\}\)\\right\|=\|F\_\{v\}\(q\)\-F\_\{v\}\(v\_\{1\}\)\|\\leq H\|q\-v\_\{1\}\|and
\|R2\(v1,r,μv\)−R2\(v1,v2,μv\)\|=\|r−v2\|\.\\left\|R\_\{2\}\(v\_\{1\},r,\\mu\_\{v\}\)\-R\_\{2\}\(v\_\{1\},v\_\{2\},\\mu\_\{v\}\)\\right\|=\|r\-v\_\{2\}\|\.
It remains to verify Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(iv\)\. The first coordinate has no preceding predictions\. For the second coordinate, fixy∈\[0,1\]y\\in\[0,1\]\. The map
q↦Lτ\(q,y\)=q\+\(y−q\)\+1−τq\\mapsto L\_\{\\tau\}\(q,y\)=q\+\\frac\{\(y\-q\)\_\{\+\}\}\{1\-\\tau\}is Lipschitz with constant
Lτ∗:=max\{1,τ1−τ\}\.L\_\{\\tau\}^\{\*\}:=\\max\\left\\\{1,\\frac\{\\tau\}\{1\-\\tau\}\\right\\\}\.Hence
\|𝔼μvLτ\(p1,Y\)−𝔼μvLτ\(v1,Y\)\|≤Lτ∗\|p1−v1\|\.\\left\|\\mathbb\{E\}\_\{\\mu\_\{v\}\}L\_\{\\tau\}\(p\_\{1\},Y\)\-\\mathbb\{E\}\_\{\\mu\_\{v\}\}L\_\{\\tau\}\(v\_\{1\},Y\)\\right\|\\leq L\_\{\\tau\}^\{\*\}\|p\_\{1\}\-v\_\{1\}\|\.Since
R2\(p1,p2,μv\)−R2\(v1,p2,μv\)=−\(𝔼μvLτ\(p1,Y\)−𝔼μvLτ\(v1,Y\)\),R\_\{2\}\(p\_\{1\},p\_\{2\},\\mu\_\{v\}\)\-R\_\{2\}\(v\_\{1\},p\_\{2\},\\mu\_\{v\}\)=\-\\left\(\\mathbb\{E\}\_\{\\mu\_\{v\}\}L\_\{\\tau\}\(p\_\{1\},Y\)\-\\mathbb\{E\}\_\{\\mu\_\{v\}\}L\_\{\\tau\}\(v\_\{1\},Y\)\\right\),this verifies Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(iv\)\.
Finally, the family\{fλ\}\\\{f\_\{\\lambda\}\\\}is a regular two\-parameter exponential family with bounded sufficient statisticsyyandy2y^\{2\}\. On a compact neighborhood of0,
DKL\(μλ∥μλ′\)≤C‖λ−λ′‖22\.D\_\{\\mathrm\{KL\}\}\(\\mu\_\{\\lambda\}\\,\\\|\\,\\mu\_\{\\lambda^\{\\prime\}\}\)\\leq C\\\|\\lambda\-\\lambda^\{\\prime\}\\\|\_\{2\}^\{2\}\.Becausev↦λ\(v\)v\\mapsto\\lambda\(v\)is smooth onℛQC\\mathcal\{R\}\_\{\\mathrm\{QC\}\},
DKL\(μv∥μv′\)≤C′‖v−v′‖22\.D\_\{\\mathrm\{KL\}\}\(\\mu\_\{v\}\\,\\\|\\,\\mu\_\{v^\{\\prime\}\}\)\\leq C^\{\\prime\}\\\|v\-v^\{\\prime\}\\\|\_\{2\}^\{2\}\.Thus Assumption[3\.1](https://arxiv.org/html/2608.04288#S3.Thmtheorem1)\(v\) holds\. ∎
###### Corollary 5\.11\(Quantile, CVaR: Lower Bound\)\.
Fixτ∈\(0,1\)\\tau\\in\(0,1\), constants0<h<1<H<∞0<h<1<H<\\infty, andκ\>0\\kappa\>0\. There exist constantsc,C,ε0\>0c,C,\\varepsilon\_\{0\}\>0, depending only onτ,h,H\\tau,h,H, andκ\\kappa, such that the following holds for every0<ε≤ε00<\\varepsilon\\leq\\varepsilon\_\{0\}\. Let
Γ=\(qτ,CVaRτ\)\\Gamma=\(q\_\{\\tau\},\\operatorname\{CVaR\}\_\{\\tau\}\)be the property overℳh,H\\mathcal\{M\}\_\{h,H\}\. One can construct a finite context space, a binary group family of size at mostε−κ\\varepsilon^\{\-\\kappa\}, and a finite collection ofℳh,H\\mathcal\{M\}\_\{h,H\}\-compatible data distributions, such that every learner achievingΓ\\Gamma\-ECE at mostε\\varepsilonwith probability at least2/32/3must use
n≥cε−4log4\(C/ε\)\.n\\geq c\\frac\{\\varepsilon^\{\-4\}\}\{\\log^\{4\}\(C/\\varepsilon\)\}\.Equivalently,
SCΓ,ℳh,H\(κ\)\(ε\)=Ω~\(ε−4\)\.\\operatorname\{SC\}^\{\(\\kappa\)\}\_\{\\Gamma,\\mathcal\{M\}\_\{h,H\}\}\(\\varepsilon\)=\\widetilde\{\\Omega\}\(\\varepsilon^\{\-4\}\)\.
###### Proof\.
Apply Theorem[3\.2](https://arxiv.org/html/2608.04288#S3.Thmtheorem2)using Proposition[5\.10](https://arxiv.org/html/2608.04288#S5.Thmtheorem10)\. ∎
For the upper bound, we verify the required conditions over the same classℳh,H\\mathcal\{M\}\_\{h,H\}\.
###### Proposition 5\.12\(Quantile, CVaR: Conditions for the Upper Bound\)\.
The regularity conditions in Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)hold overℳh,H\\mathcal\{M\}\_\{h,H\}\. Moreover, the linear optimization oracle in Assumption[4\.10](https://arxiv.org/html/2608.04288#S4.Thmtheorem10)has a polynomial\-time implementation to any additive accuracyξ\>0\\xi\>0\.
###### Proof\.
Recall that
R\(\(q,r\),y\)=\(𝟏\{y≤q\}−τ,r−Lτ\(q,y\)\)\.R\(\(q,r\),y\)=\\left\(\\mathbf\{1\}\\\{y\\leq q\\\}\-\\tau,\\ r\-L\_\{\\tau\}\(q,y\)\\right\)\.The classℳh,H\\mathcal\{M\}\_\{h,H\}is convex and compact in the weak topology\. For each fixed\(q,r\)\(q,r\), the map
μ↦Fμ\(q\)−τ\\mu\\mapsto F\_\{\\mu\}\(q\)\-\\tauis affine and continuous onℳh,H\\mathcal\{M\}\_\{h,H\}, because every law in this class has no atom atqq\. The map
μ↦𝔼μ\[r−Lτ\(q,Y\)\]\\mu\\mapsto\\mathbb\{E\}\_\{\\mu\}\[r\-L\_\{\\tau\}\(q,Y\)\]is affine and continuous becauseLτ\(q,⋅\)L\_\{\\tau\}\(q,\\cdot\)is bounded and continuous\. These observations verify Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(i\)\. Condition \(ii\) also holds becauseq,r,y∈\[0,1\]q,r,y\\in\[0,1\]andτ∈\(0,1\)\\tau\\in\(0,1\)\.
To verify Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(iii\), letFμF\_\{\\mu\}be the CDF ofμ\\mu\. Forq,q′∈\[0,1\]q,q^\{\\prime\}\\in\[0,1\],
\|R1\(∅,q,μ\)−R1\(∅,q′,μ\)\|=\|Fμ\(q\)−Fμ\(q′\)\|≤H\|q−q′\|\.\\left\|R\_\{1\}\(\\varnothing,q,\\mu\)\-R\_\{1\}\(\\varnothing,q^\{\\prime\},\\mu\)\\right\|=\|F\_\{\\mu\}\(q\)\-F\_\{\\mu\}\(q^\{\\prime\}\)\|\\leq H\|q\-q^\{\\prime\}\|\.Also, the mapq↦Lτ\(q,y\)q\\mapsto L\_\{\\tau\}\(q,y\)is Lipschitz uniformly inyy, with a constant depending only onτ\\tau\. ThereforeR\(⋅,μ\)R\(\\cdot,\\mu\)is uniformly Lipschitz in\(q,r\)\(q,r\)\. A rectangular grid with meshO\(1/Q\)O\(1/Q\)satisfies the required approximation bound and hasO\(Q2\)O\(Q^\{2\}\)points\.
It remains to implement the oracle\. Given coefficientscp,jc\_\{p,j\}, write each grid point asp=\(qp,rp\)p=\(q\_\{p\},r\_\{p\}\)\. The integrand
ϕ\(y\)=∑p∈𝒫Q\[cp,1\(𝟏\{y≤qp\}−τ\)\+cp,2\(rp−qp−\(y−qp\)\+/\(1−τ\)\)\]\\phi\(y\)=\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\left\[c\_\{p,1\}\(\\mathbf\{1\}\\\{y\\leq q\_\{p\}\\\}\-\\tau\)\+c\_\{p,2\}\\bigl\(r\_\{p\}\-q\_\{p\}\-\(y\-q\_\{p\}\)\_\{\+\}/\(1\-\\tau\)\\bigr\)\\right\]is affine on each open interval between consecutive distinct grid valuesqpq\_\{p\}, and it may have jump discontinuities at those values\. Since there are only finitely many such points, their values do not affect the oracle objective\. The oracle must maximize
∫01ϕ\(y\)f\(y\)𝑑ysubject toh≤f≤H,∫01f=1\.\\int\_\{0\}^\{1\}\\phi\(y\)f\(y\)\\,dy\\qquad\\text\{subject to\}\\qquad h\\leq f\\leq H,\\quad\\int\_\{0\}^\{1\}f=1\.Writef=h\+\(H−h\)uf=h\+\(H\-h\)u, where0≤u≤10\\leq u\\leq 1and
∫01u\(y\)dy=1−hH−h=:m0\.\\int\_\{0\}^\{1\}u\(y\)\\,dy=\\frac\{1\-h\}\{H\-h\}=:m\_\{0\}\.The objective differs by the constanth∫ϕh\\int\\phifrom maximizing∫ϕu\\int\\phi uover0≤u≤10\\leq u\\leq 1with massm0m\_\{0\}\. For a thresholdtt, let
A\(t\):=∫01𝟏\{ϕ\(y\)\>t\}𝑑y,B\(t\):=∫01𝟏\{ϕ\(y\)≥t\}𝑑y\.A\(t\):=\\int\_\{0\}^\{1\}\\mathbf\{1\}\\\{\\phi\(y\)\>t\\\}\\,dy,\\qquad B\(t\):=\\int\_\{0\}^\{1\}\\mathbf\{1\}\\\{\\phi\(y\)\\geq t\\\}\\,dy\.Choosettsuch thatA\(t\)≤m0≤B\(t\)A\(t\)\\leq m\_\{0\}\\leq B\(t\), and set
γ:=\{m0−A\(t\)B\(t\)−A\(t\),B\(t\)\>A\(t\),0,B\(t\)=A\(t\)\.\\gamma:=\\begin\{cases\}\\dfrac\{m\_\{0\}\-A\(t\)\}\{B\(t\)\-A\(t\)\},&B\(t\)\>A\(t\),\\\\\[8\.00003pt\] 0,&B\(t\)=A\(t\)\.\\end\{cases\}Then
ut\(y\):=𝟏\{ϕ\(y\)\>t\}\+γ𝟏\{ϕ\(y\)=t\}u\_\{t\}\(y\):=\\mathbf\{1\}\\\{\\phi\(y\)\>t\\\}\+\\gamma\\mathbf\{1\}\\\{\\phi\(y\)=t\\\}has massm0m\_\{0\}\. It is optimal by the bathtub principle, which places the available mass whereϕ\\phiis largest\. This includes the case in whichϕ\\phiis constant on an interval:γ\\gammasupplies exactly the required fraction of that level set\.
Becauseϕ\\phihasO\(\|𝒫Q\|\)O\(\|\\mathcal\{P\}\_\{Q\}\|\)affine pieces,A\(t\)A\(t\)is affine on each interval between consecutive critical values\. These critical values are the values ofϕ\\phiat the endpoints of its affine pieces, together with the values on any constant pieces\. Sorting them and solving one linear equation locates an admissiblettandγ\\gammain polynomial time\. The resulting densityft=h\+\(H−h\)utf\_\{t\}=h\+\(H\-h\)u\_\{t\}is piecewise constant onO\(\|𝒫Q\|\)O\(\|\\mathcal\{P\}\_\{Q\}\|\)intervals, which gives a polynomial\-size representation\. Integrating the residuals interval by interval computes the objective and all residual expectations to additive accuracyξ\\xiin time polynomial in\|𝒫Q\|\|\\mathcal\{P\}\_\{Q\}\|andlog\(1/ξ\)\\log\(1/\\xi\)\. This implements Assumption[4\.10](https://arxiv.org/html/2608.04288#S4.Thmtheorem10)\. ∎
###### Corollary 5\.13\(Quantile, CVaR: Upper Bound\)\.
Fixτ∈\(0,1\)\\tau\\in\(0,1\), constants0<h<1<H<∞0<h<1<H<\\infty, andκ\>0\\kappa\>0\. There exist constantsC,ε0\>0C,\\varepsilon\_\{0\}\>0, depending only onτ,h,H\\tau,h,H, andκ\\kappa, such that the following holds for every0<ε≤ε00<\\varepsilon\\leq\\varepsilon\_\{0\}\. Let
Γ=\(qτ,CVaRτ\)\\Gamma=\(q\_\{\\tau\},\\operatorname\{CVaR\}\_\{\\tau\}\)be the property overℳh,H\\mathcal\{M\}\_\{h,H\}\. For everyℳh,H\\mathcal\{M\}\_\{h,H\}\-compatible data distribution𝖣\\mathsf\{D\}and every finite group family satisfying\|𝒢\|≤⌈ε−κ⌉\|\\mathcal\{G\}\|\\leq\\lceil\\varepsilon^\{\-\\kappa\}\\rceil, there is a randomized learner using at mostCε−4C\\varepsilon^\{\-4\}samples whose outputΠS\\Pi\_\{S\}satisfies
ℙ\[MCErr𝖣Γ\(ΠS;𝒢\)≤ε\]≥23\.\\mathbb\{P\}\\left\[\\operatorname\{MCErr\}^\{\\Gamma\}\_\{\\mathsf\{D\}\}\(\\Pi\_\{S\};\\mathcal\{G\}\)\\leq\\varepsilon\\right\]\\geq\\frac\{2\}\{3\}\.The oracle in Proposition[5\.12](https://arxiv.org/html/2608.04288#S5.Thmtheorem12)gives training time polynomial in1/ε1/\\varepsilon\.
###### Proof\.
Apply Corollary[4\.9](https://arxiv.org/html/2608.04288#S4.Thmtheorem9)and Theorem[4\.11](https://arxiv.org/html/2608.04288#S4.Thmtheorem11)using Proposition[5\.12](https://arxiv.org/html/2608.04288#S5.Thmtheorem12)\. ∎
### Statement of AI Use
GPT\-5\.6 Sol Pro was used during manuscript preparation for brainstorming, drafting, editing, and formatting assistance, including preliminary drafts of proof arguments\. The LLM\-generated material was not used without author review\. All proof arguments were independently checked, corrected, and finalized by the authors, who take full responsibility for the correctness, originality, and presentation of the paper\.
## References
- \[1\]K\. Balasubramanian, S\. Ghadimi, and A\. Nguyen\(2022\)Stochastic multilevel composition optimization algorithms with level\-independent convergence rates\.SIAM Journal on Optimization32\(2\),pp\. 519–544\.Cited by:[Appendix A](https://arxiv.org/html/2608.04288#A1.p2.9)\.
- \[2\]O\. Bastani, V\. Gupta, C\. Jung, G\. Noarov, R\. Ramalingam, and A\. Roth\(2022\)Practical adversarial multivalid conformal prediction\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 29362–29373\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px2.p1.1)\.
- \[3\]S\. Casacuberta, C\. Dwork, and S\. Vadhan\(2024\)Complexity\-theoretic implications of multicalibration\.InProceedings of the 56th Annual ACM Symposium on Theory of Computing,pp\. 1071–1082\.Cited by:[§1](https://arxiv.org/html/2608.04288#S1.p1.3)\.
- \[4\]S\. Casacuberta, P\. Gopalan, V\. Kanade, and O\. Reingold\(2025\)How global calibration strengthens multiaccuracy\.In2025 IEEE 66th Annual Symposium on Foundations of Computer Science \(FOCS\),pp\. 1198–1227\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p3.1)\.
- \[5\]T\. Chen, Y\. Sun, and W\. Yin\(2021\)Solving stochastic compositional optimization is nearly as easy as solving stochastic optimization\.IEEE Transactions on Signal Processing69,pp\. 4937–4948\.Cited by:[Appendix A](https://arxiv.org/html/2608.04288#A1.p2.9)\.
- \[6\]N\. Collina, I\. Globus\-Harris, S\. Goel, V\. Gupta, A\. Roth, and M\. Shi\(2026\)Collaborative prediction: tractable information aggregation via agreement\.InProceedings of the 2026 Annual ACM\-SIAM Symposium on Discrete Algorithms \(SODA\),pp\. 4712–4798\.Cited by:[§1](https://arxiv.org/html/2608.04288#S1.p1.3)\.
- \[7\]N\. Collina, S\. Goel, V\. Gupta, and A\. Roth\(2025\)Tractable agreement protocols\.InProceedings of the 57th Annual ACM Symposium on Theory of Computing,pp\. 1532–1543\.Cited by:[§1](https://arxiv.org/html/2608.04288#S1.p1.3)\.
- \[8\]N\. Collina, J\. Lu, G\. Noarov, and A\. Roth\(2026\)Optimal lower bounds for online multicalibration\.In2026 IEEE 67th Annual Symposium on Foundations of Computer Science \(FOCS\),Note:To appearCited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p4.1),[§3\.4](https://arxiv.org/html/2608.04288#S3.SS4.1.p1.7),[§3\.4](https://arxiv.org/html/2608.04288#S3.SS4.p7.1)\.
- \[9\]N\. Collina, J\. Lu, G\. Noarov, and A\. Roth\(2026\)The sample complexity of multicalibration\.arXiv preprint arXiv:2604\.21923\.Cited by:[§1\.1](https://arxiv.org/html/2608.04288#S1.SS1.SSS0.Px3.p2.6),[§1\.2](https://arxiv.org/html/2608.04288#S1.SS2.SSS0.Px2.p6.6),[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p1.5),[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p4.1),[§1](https://arxiv.org/html/2608.04288#S1.p6.11),[§2\.3](https://arxiv.org/html/2608.04288#S2.SS3.p1.7),[Remark 2\.2](https://arxiv.org/html/2608.04288#S2.Thmtheorem2.p1.4.2),[§3\.1](https://arxiv.org/html/2608.04288#S3.SS1.p4.1),[§3\.3](https://arxiv.org/html/2608.04288#S3.SS3.p3.1),[§3\.6](https://arxiv.org/html/2608.04288#S3.SS6.p3.2),[§4\.5\.1](https://arxiv.org/html/2608.04288#S4.SS5.SSS1.p1.1)\.
- \[10\]T\. M\. Cover and J\. A\. Thomas\(2006\)Elements of information theory\.Second edition,Wiley\-Interscience\.Cited by:[§3\.7](https://arxiv.org/html/2608.04288#S3.SS7.5.p2.3)\.
- \[11\]A\. P\. Dawid\(1982\)The well\-calibrated Bayesian\.Journal of the American Statistical Association77\(379\),pp\. 605–610\.Cited by:[§1](https://arxiv.org/html/2608.04288#S1.p1.3)\.
- \[12\]Z\. Deng, C\. Dwork, and L\. Zhang\(2023\)HappyMap: a generalized multicalibration method\.In14th Innovations in Theoretical Computer Science Conference \(ITCS 2023\),Leibniz International Proceedings in Informatics \(LIPIcs\), Vol\.251,pp\. 41:1–41:23\.Cited by:[§1](https://arxiv.org/html/2608.04288#S1.p2.4)\.
- \[13\]D\. Dentcheva, S\. Penev, and A\. Ruszczyński\(2017\)Statistical estimation of composite risk functionals and risk optimization problems\.Annals of the Institute of Statistical Mathematics69\(4\),pp\. 737–760\.Cited by:[Appendix A](https://arxiv.org/html/2608.04288#A1.p2.9)\.
- \[14\]C\. Dwork and P\. Tankala\(2025\)Supersimulators\.arXiv preprint arXiv:2509\.17994\.Cited by:[§1](https://arxiv.org/html/2608.04288#S1.p1.3)\.
- \[15\]Y\. M\. Ermoliev and V\. I\. Norkin\(2013\)Sample average approximation method for compound stochastic optimization problems\.SIAM Journal on Optimization23\(4\),pp\. 2231–2263\.Cited by:[Appendix A](https://arxiv.org/html/2608.04288#A1.p2.9)\.
- \[16\]M\. Fishelson, N\. Golowich, M\. Mohri, and J\. Schneider\(2025\)High\-dimensional calibration from swap regret\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 50433–50465\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px3.p1.1)\.
- \[17\]S\. Garg, C\. Jung, O\. Reingold, and A\. Roth\(2024\)Oracle efficient online multicalibration and omniprediction\.InProceedings of the 2024 Annual ACM\-SIAM Symposium on Discrete Algorithms \(SODA\),pp\. 2725–2792\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p1.5)\.
- \[18\]R\. Ghuge, V\. Muthukumar, and S\. Singla\(2025\)Improved and oracle\-efficient onlineℓ1\\ell\_\{1\}\-multicalibration\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 19437–19457\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p1.5)\.
- \[19\]I\. Gibbs and R\. J\. Tibshirani\(2026\)Sample\-efficient omniprediction for proper losses\.InProceedings of Thirty Ninth Conference on Learning Theory,Proceedings of Machine Learning Research, Vol\.336,pp\. 2679–2719\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p3.1)\.
- \[20\]I\. Globus\-Harris, D\. Harrison, M\. Kearns, A\. Roth, and J\. Sorrell\(2023\)Multicalibration as boosting for regression\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 11459–11492\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p1.5)\.
- \[21\]P\. Gopalan, L\. Hu, M\. P\. Kim, O\. Reingold, and U\. Wieder\(2023\)Loss minimization through the lens of outcome indistinguishability\.In14th Innovations in Theoretical Computer Science Conference \(ITCS 2023\),Leibniz International Proceedings in Informatics \(LIPIcs\), Vol\.251,pp\. 60:1–60:20\.Cited by:[§1](https://arxiv.org/html/2608.04288#S1.p1.3)\.
- \[22\]P\. Gopalan, L\. Hu, and G\. N\. Rothblum\(2024\)On computationally efficient multi\-class calibration\.InProceedings of Thirty Seventh Conference on Learning Theory,Proceedings of Machine Learning Research, Vol\.247,pp\. 1983–2026\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px3.p1.1)\.
- \[23\]P\. Gopalan, A\. T\. Kalai, O\. Reingold, V\. Sharan, and U\. Wieder\(2022\)Omnipredictors\.In13th Innovations in Theoretical Computer Science Conference \(ITCS 2022\),Leibniz International Proceedings in Informatics \(LIPIcs\), Vol\.215,pp\. 79:1–79:21\.Cited by:[§1](https://arxiv.org/html/2608.04288#S1.p1.3)\.
- \[24\]P\. Gopalan, M\. P\. Kim, M\. A\. Singhal, and S\. Zhao\(2022\)Low\-degree multicalibration\.InProceedings of Thirty Fifth Conference on Learning Theory,Proceedings of Machine Learning Research, Vol\.178,pp\. 3193–3234\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p1.5)\.
- \[25\]P\. Gopalan, M\. Kim, and O\. Reingold\(2023\)Swap agnostic learning, or characterizing omniprediction via multicalibration\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 39936–39956\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p2.1)\.
- \[26\]M\. Grötschel, L\. Lovász, and A\. Schrijver\(1988\)Geometric algorithms and combinatorial optimization\.Algorithms and Combinatorics, Vol\.2,Springer\.Cited by:[§4\.6\.2](https://arxiv.org/html/2608.04288#S4.SS6.SSS2.p2.16)\.
- \[27\]V\. Gupta, C\. Jung, G\. Noarov, M\. M\. Pai, and A\. Roth\(2022\)Online multivalid learning: means, moments, and prediction intervals\.In13th Innovations in Theoretical Computer Science Conference \(ITCS 2022\),Leibniz International Proceedings in Informatics \(LIPIcs\), Vol\.215,pp\. 82:1–82:24\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2608.04288#S1.p2.4)\.
- \[28\]N\. Haghtalab, M\. I\. Jordan, and E\. Zhao\(2023\)A unifying perspective on multi\-calibration: game dynamics for multi\-objective learning\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 72464–72506\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p1.5)\.
- \[29\]U\. Hebert\-Johnson, M\. Kim, O\. Reingold, and G\. Rothblum\(2018\)Multicalibration: calibration for the \(Computationally\-identifiable\) masses\.InProceedings of the 35th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.80,pp\. 1939–1948\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p1.5),[§1](https://arxiv.org/html/2608.04288#S1.p1.3)\.
- \[30\]W\. Hoeffding\(1963\)Probability inequalities for sums of bounded random variables\.Journal of the American Statistical Association58\(301\),pp\. 13–30\.Cited by:[§4\.4\.1](https://arxiv.org/html/2608.04288#S4.SS4.SSS1.1.p1.3)\.
- \[31\]L\. Hu, H\. Luo, S\. Senapati, and V\. Sharan\(2026\)Efficient swap multicalibration of elicitable properties\.InProceedings of Thirty Ninth Conference on Learning Theory,Proceedings of Machine Learning Research, Vol\.336,pp\. 3314–3348\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px2.p1.1)\.
- \[32\]C\. Jung, C\. Lee, M\. Pai, A\. Roth, and R\. Vohra\(2021\)Moment multicalibration for uncertainty estimation\.InProceedings of Thirty Fourth Conference on Learning Theory,Proceedings of Machine Learning Research, Vol\.134,pp\. 2634–2678\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2608.04288#S1.p2.4)\.
- \[33\]C\. Jung, G\. Noarov, R\. Ramalingam, and A\. Roth\(2023\)Batch multivalid conformal prediction\.InThe Eleventh International Conference on Learning Representations,Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2608.04288#S1.p2.4)\.
- \[34\]D\. D\. Lee, G\. Noarov, M\. Pai, and A\. Roth\(2022\)Online minimax multiobjective optimization: multicalibeating and other applications\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 29051–29063\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p1.5),[§4\.1](https://arxiv.org/html/2608.04288#S4.SS1.p3.9)\.
- \[35\]H\. Luo, S\. Senapati, and V\. Sharan\(2025\)Improved bounds for swap multicalibration and swap omniprediction\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 20425–20467\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p2.1)\.
- \[36\]M\. P\. Naeini, G\. F\. Cooper, and M\. Hauskrecht\(2015\)Obtaining well calibrated probabilities using Bayesian binning\.InProceedings of the Twenty\-Ninth AAAI Conference on Artificial Intelligence,pp\. 2901–2907\.Cited by:[§1](https://arxiv.org/html/2608.04288#S1.p6.11)\.
- \[37\]G\. Noarov, R\. Ramalingam, A\. Roth, and S\. Xie\(2025\)High\-dimensional prediction for sequential decision making\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 46762–46783\.Cited by:[§1\.2](https://arxiv.org/html/2608.04288#S1.SS2.SSS0.Px2.p6.6),[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2608.04288#S4.SS1.p4.6),[§4\.6\.2](https://arxiv.org/html/2608.04288#S4.SS6.SSS2.p1.2)\.
- \[38\]G\. Noarov and A\. Roth\(2023\)The statistical scope of multicalibration\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 26283–26310\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2608.04288#S1.p2.4),[§1](https://arxiv.org/html/2608.04288#S1.p3.17),[§2\.1](https://arxiv.org/html/2608.04288#S2.SS1.p3.7),[Remark 2\.2](https://arxiv.org/html/2608.04288#S2.Thmtheorem2.p1.4.2),[§5\.3](https://arxiv.org/html/2608.04288#S5.SS3.p1.11)\.
- \[39\]P\. Okoroafor, R\. Kleinberg, and M\. P\. Kim\(2025\)Near\-optimal algorithms for omniprediction\.In2025 IEEE 66th Annual Symposium on Foundations of Computer Science \(FOCS\),pp\. 1595–1609\.Cited by:[§1](https://arxiv.org/html/2608.04288#S1.p1.3)\.
- \[40\]B\. Peng\(2025\)High dimensional online calibration in polynomial time\.In2025 IEEE 66th Annual Symposium on Foundations of Computer Science \(FOCS\),Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px3.p1.1)\.
- \[41\]H\. Rosenberg, R\. Bhattacharjee, K\. Fawaz, and S\. Jha\(2022\)An exploration of multicalibration uniform convergence bounds\.arXiv preprint arXiv:2202\.04530\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p5.1)\.
- \[42\]R\. M\. Roth\(2006\)Introduction to coding theory\.Cambridge University Press\.Cited by:[§3\.3](https://arxiv.org/html/2608.04288#S3.SS3.p5.1)\.
- \[43\]A\. Ruszczyński\(2021\)A stochastic subgradient method for nonsmooth nonconvex multilevel composition optimization\.SIAM Journal on Control and Optimization59\(3\),pp\. 2301–2320\.Cited by:[Appendix A](https://arxiv.org/html/2608.04288#A1.p2.9)\.
- \[44\]E\. Shabat, L\. Cohen, and Y\. Mansour\(2020\)Sample complexity of uniform convergence for multicalibration\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 13331–13340\.Cited by:[§1\.3](https://arxiv.org/html/2608.04288#S1.SS3.SSS0.Px1.p5.1)\.
- \[45\]M\. Sion\(1958\)On general minimax theorems\.Pacific Journal of Mathematics8\(1\),pp\. 171–176\.Cited by:[§4\.4\.2](https://arxiv.org/html/2608.04288#S4.SS4.SSS2.1.p1.12)\.
- \[46\]A\. Ta\-Shma\(2017\)Explicit, almost optimal, epsilon\-balanced codes\.InProceedings of the 49th Annual ACM SIGACT Symposium on Theory of Computing,pp\. 238–251\.Cited by:[§3\.3](https://arxiv.org/html/2608.04288#S3.SS3.p5.1)\.
## Appendix AComparison with Multi\-Level Stochastic Optimization
Here, we discuss the difference betweenkk\-level multicalibration problem and the well\-studiedkk\-level stochastic composite optimization problem\. The key distinction is that the levels are*sequential*in compositional optimization but*Cartesian*in multicalibration\.
Akk\-level composition optimization problem involves solving
F\(x\)=f1∘f2∘⋯∘fk\(x\),F\(x\)=f\_\{1\}\\circ f\_\{2\}\\circ\\cdots\\circ f\_\{k\}\(x\),given access to stochastic evaluations of the gradients and function values of functionsfjf\_\{j\}\. In this setup, typically, the quantity of interest is a fixed pointxxthat determines only one chain of intermediate values,
yk=x,yj−1=fj\(yj\),F\(x\)=f1\(y1\)\.y\_\{k\}=x,\\qquad y\_\{j\-1\}=f\_\{j\}\(y\_\{j\}\),\\qquad F\(x\)=f\_\{1\}\(y\_\{1\}\)\.The optimization algorithm must estimate and track the quantities along this chain, but its statistical goal is still to find a single pointx^N\\widehat\{x\}\_\{N\}\(whereNNis the sample or iteration complexity\) satisfying a first\-order condition, such as
𝔼\[‖∇F\(x^N\)‖\]≤ε\.\\mathbb\{E\}\\bigl\[\\\|\\nabla F\(\\widehat\{x\}\_\{N\}\)\\\|\\bigr\]\\leq\\varepsilon\.It is not required to certify stationarity separately for every possible vector of intermediate values\(y1,…,yk−1\)\(y\_\{1\},\\ldots,y\_\{k\-1\}\)\. Smoothness allows the estimation errors arising at the different levels to be propagated along the chain and controlled on the same stochastic time scale\. For the stationarity criterion considered in the multi\-level optimization literature, the resulting bound has the schematic form
𝔼\[‖∇F\(x^N\)‖2\]≲N−1/2,and hence𝔼\[‖∇F\(x^N\)‖\]≲N−1/4\.\\mathbb\{E\}\\bigl\[\\\|\\nabla F\(\\widehat\{x\}\_\{N\}\)\\\|^\{2\}\\bigr\]\\lesssim N^\{\-1/2\},\\qquad\\text\{and hence\}\\qquad\\mathbb\{E\}\\bigl\[\\\|\\nabla F\(\\widehat\{x\}\_\{N\}\)\\\|\\bigr\]\\lesssim N^\{\-1/4\}\.ThusN≳ε−4N\\gtrsim\\varepsilon^\{\-4\}suffices, independently of the number of compositional levels in the exponent ofε\\varepsilon\. Additional levels make the gradient estimator more involved, but they do not create additionalε\\varepsilon\-scale locations at which the desired guarantee must hold\. This sequential propagation of estimation error also underlies level\-independent accuracy rates in statistical and sample\-average analyses of compound objectives\[[13](https://arxiv.org/html/2608.04288#bib.bib4),[15](https://arxiv.org/html/2608.04288#bib.bib3)\], as well as in stochastic algorithms for nested composition problems\[[43](https://arxiv.org/html/2608.04288#bib.bib2),[5](https://arxiv.org/html/2608.04288#bib.bib5),[1](https://arxiv.org/html/2608.04288#bib.bib6)\]\.
Multicalibration has a different geometry\. The predictor outputs a prediction vector
p=\(p1,…,pk\)∈𝒫1×⋯×𝒫k,p=\(p\_\{1\},\\ldots,p\_\{k\}\)\\in\\mathcal\{P\}\_\{1\}\\times\\cdots\\times\\mathcal\{P\}\_\{k\},and calibration is imposed after conditioning on the*entire*prediction vector\. Although the residuals are sequentially conditionally identifiable, the collection of prediction values over which calibration must be controlled is a Cartesian product\. Discretizing each coordinate at resolutionQ−1Q^\{\-1\}therefore produces
\|𝒫Q\|≍Qk\|\\mathcal\{P\}\_\{Q\}\|\\asymp Q^\{k\}prediction values\. Correspondingly, the statistical error has the form
MCErrDΓ\(Π;𝒢\)≲1Q\+Qk\+log\|𝒢\|n\.\\operatorname\{MCErr\}^\{\\Gamma\}\_\{D\}\(\\Pi;\\mathcal\{G\}\)\\lesssim\\frac\{1\}\{Q\}\+\\sqrt\{\\frac\{Q^\{k\}\+\\log\|\\mathcal\{G\}\|\}\{n\}\}\.The difference from ordinary mean estimation can be seen directly from the second term\. Suppose, for intuition, that theM=QkM=Q^\{k\}prediction buckets have comparable probability\. A bucket then contains aboutn/Mn/Mobservations\. Its conditional residual is estimated with error of orderM/n\\sqrt\{M/n\}; after multiplying by the bucket probability1/M1/M, its contribution to the calibration error is of order1/Mn1/\\sqrt\{Mn\}\. Since the calibration error sums absolute residual contributions over allMMbuckets, their total stochastic contribution is
M⋅1Mn=Mn=Qkn\.M\\cdot\\frac\{1\}\{\\sqrt\{Mn\}\}=\\sqrt\{\\frac\{M\}\{n\}\}=\\sqrt\{\\frac\{Q^\{k\}\}\{n\}\}\.To make the discretization error at mostε\\varepsilon, one takesQ≍ε−1Q\\asymp\\varepsilon^\{\-1\}\. Requiring the stochastic term to be at mostε\\varepsilonthen gives
Qkn≲ε⟹n≳Qkε2≍ε−\(k\+2\)\.\\sqrt\{\\frac\{Q^\{k\}\}\{n\}\}\\lesssim\\varepsilon\\qquad\\Longrightarrow\\qquad n\\gtrsim\\frac\{Q^\{k\}\}\{\\varepsilon^\{2\}\}\\asymp\\varepsilon^\{\-\(k\+2\)\}\.Thus the factorε−2\\varepsilon^\{\-2\}is the usual cost of estimating expectations to accuracyε\\varepsilon, while the additional factorε−k\\varepsilon^\{\-k\}is the number of distinguishable prediction values at that resolution\. Each additional calibrated level adds another prediction coordinate, multiplies the size of the joint grid byε−1\\varepsilon^\{\-1\}, and therefore increases the exponent by one\.
The lower bound confirms that this difference is intrinsic rather than an artifact of discretization\. At resolutionQ−1Q^\{\-1\}, one can encode independent perturbations of sizeΘ\(Q−1\)\\Theta\(Q^\{\-1\}\)acrossΘ\(Qk\)\\Theta\(Q^\{k\}\)grid points\. The logarithm of the number of distinguishable alternatives is thenΩ\(Qk\)\\Omega\(Q^\{k\}\), whereas one sample contributes onlyO\(Q−2\)O\(Q^\{\-2\}\)Kullback–Leibler information\. Fano’s inequality consequently requires
nQ−2≳Qk,or equivalentlyn≳Qk\+2\.nQ^\{\-2\}\\gtrsim Q^\{k\},\\qquad\\text\{or equivalently\}\\qquad n\\gtrsim Q^\{k\+2\}\.The reduction from calibration error to prediction accuracy loses a factor oflogQ\\log Q, so a construction at resolutionQ−1Q^\{\-1\}applies whenε≍1/\(QlogQ\)\\varepsilon\\asymp 1/\(Q\\log Q\)\. Thus
Q≍1εlog\(1/ε\),n≳ε−\(k\+2\)logk\+2\(1/ε\)=Ω~\(ε−\(k\+2\)\)\.Q\\asymp\\frac\{1\}\{\\varepsilon\\log\(1/\\varepsilon\)\},\\qquad n\\gtrsim\\frac\{\\varepsilon^\{\-\(k\+2\)\}\}\{\\log^\{k\+2\}\(1/\\varepsilon\)\}=\\widetilde\{\\Omega\}\\\!\\left\(\\varepsilon^\{\-\(k\+2\)\}\\right\)\.
In summary, stochastic composition optimization follows one nested chain and asks for stationarity at one output point, whereas multicalibration must resolve and control akk\-dimensional Cartesian grid of prediction values\.
## Appendix BMeasurability of the Online Forecaster
We justify that the prediction rules in Algorithm[1](https://arxiv.org/html/2608.04288#alg1)can be chosen jointly measurably\. For the fixed grid𝒫Q\\mathcal\{P\}\_\{Q\}, define the finite\-dimensional residual image
𝒦Q:=\{\(Rj\(p<j,pj,μ\)\)\(p,j\)∈𝒫Q×\[k\]:μ∈ℳ\}⊆ℝk\|𝒫Q\|\.\\mathcal\{K\}\_\{Q\}:=\\left\\\{\\bigl\(R\_\{j\}\(p\_\{<j\},p\_\{j\},\\mu\)\\bigr\)\_\{\(p,j\)\\in\\mathcal\{P\}\_\{Q\}\\times\[k\]\}:\\mu\\in\\mathcal\{M\}\\right\\\}\\subseteq\\mathbb\{R\}^\{k\|\\mathcal\{P\}\_\{Q\}\|\}\.Assumption[4\.1](https://arxiv.org/html/2608.04288#S4.Thmtheorem1)\(i\) implies that𝒦Q\\mathcal\{K\}\_\{Q\}is compact, because it is the continuous image of the compact setℳ\\mathcal\{M\}\. Forw∈Δ\(𝒢×Σ\)w\\in\\Delta\(\\mathcal\{G\}\\times\\Sigma\),z=\(zg\)g∈𝒢∈\[0,1\]𝒢z=\(z\_\{g\}\)\_\{g\\in\\mathcal\{G\}\}\\in\[0,1\]^\{\\mathcal\{G\}\},π∈Δ\(𝒫Q\)\\pi\\in\\Delta\(\\mathcal\{P\}\_\{Q\}\), andr=\(rp,j\)p,j∈𝒦Qr=\(r\_\{p,j\}\)\_\{p,j\}\\in\\mathcal\{K\}\_\{Q\}, set
F\(π,r;w,z\):=∑p∈𝒫Qπp∑g∈𝒢∑s∈Σw\(g,s\)zg∑j=1ksp,jrp,j\.F\(\\pi,r;w,z\):=\\sum\_\{p\\in\\mathcal\{P\}\_\{Q\}\}\\pi\_\{p\}\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{s\\in\\Sigma\}w\(g,s\)z\_\{g\}\\sum\_\{j=1\}^\{k\}s\_\{p,j\}r\_\{p,j\}\.This function is jointly continuous on finite\-dimensional compact spaces\. Consequently,
Φ\(π;w,z\):=maxr∈𝒦QF\(π,r;w,z\),V\(w,z\):=minπ∈Δ\(𝒫Q\)Φ\(π;w,z\),\\Phi\(\\pi;w,z\):=\\max\_\{r\\in\\mathcal\{K\}\_\{Q\}\}F\(\\pi,r;w,z\),\\qquad V\(w,z\):=\\min\_\{\\pi\\in\\Delta\(\\mathcal\{P\}\_\{Q\}\)\}\\Phi\(\\pi;w,z\),are continuous by the maximum theorem\. For everyρ≥0\\rho\\geq 0, the correspondence
𝒜ρ\(w,z\):=\{π∈Δ\(𝒫Q\):Φ\(π;w,z\)≤V\(w,z\)\+ρ\}\\mathcal\{A\}\_\{\\rho\}\(w,z\):=\\left\\\{\\pi\\in\\Delta\(\\mathcal\{P\}\_\{Q\}\):\\Phi\(\\pi;w,z\)\\leq V\(w,z\)\+\\rho\\right\\\}has nonempty compact values and a measurable graph\. The measurable selection theorem therefore provides a Borel mapaρ\(w,z\)∈𝒜ρ\(w,z\)a\_\{\\rho\}\(w,z\)\\in\\mathcal\{A\}\_\{\\rho\}\(w,z\)\.
At roundtt, take
z\(x\):=\(g\(x\)\)g∈𝒢,πt\(⋅∣x\):=aρ\(wt,z\(x\)\)\.z\(x\):=\(g\(x\)\)\_\{g\\in\\mathcal\{G\}\},\\qquad\\pi\_\{t\}\(\\cdot\\mid x\):=a\_\{\\rho\}\(w\_\{t\},z\(x\)\)\.The mapx↦z\(x\)x\\mapsto z\(x\)is measurable, andwtw\_\{t\}is a measurable function of the cumulative gains through roundt−1t\-1\. Hence\(history,x\)↦πt\(⋅∣x\)\(\\text\{history\},x\)\\mapsto\\pi\_\{t\}\(\\cdot\\mid x\)is jointly measurable and satisfies \([4](https://arxiv.org/html/2608.04288#S4.E4)\)\.Similar Articles
Smoothed Elicitation Complexity for Approximate $\Gamma$-calibration of Discrete Classification Tasks
This paper characterizes approximate property calibration for discrete properties in multiclass classification, using Lipschitz continuous properties as an intermediary to reduce complexity from the number of classes to the elicitation complexity dimension.
Finite Sample Bounds for Learning with Score Matching
This paper provides the first non-asymptotic sample complexity bounds for learning exponential families of polynomials with score matching, showing polynomial dependence on model dimension.
Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs
This paper introduces a validity-diversity framework attributing diversity collapse in LLMs to order and shape miscalibration during decoding, validated across 14 language models.
From token probabilities to calibrated confidence: An empirical study of mathematical question answering
This paper empirically studies token-probability-based confidence estimation and calibration for LLMs in mathematical question answering, comparing single-pass and multi-pass estimators and evaluating post-hoc calibration methods.
Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?
This paper presents the first study of probability calibration as a mitigation for evaluator preference coupling in LLM agent feedback loops, showing that calibrated evaluator judgments reduce coupling coefficients by 20-49% and divergence by 45-67%.