Beyond Client Averaging: A Client-Independent Second-Order Stationary-Bias Component in Stochastic SCAFFOLD

arXiv cs.LG Papers

Summary

The paper proves that in stochastic SCAFFOLD for federated learning, client averaging suppresses the leading stationary bias but leaves a client-independent second-order bias component due to control fluctuations and nonquadratic curvature.

arXiv:2608.26765v1 Announce Type: new Abstract: Existing constant-step analysis of stochastic \Scaf{} identifies a leading $O(\gamma/N)$ stationary mean bias and shows that higher-order bias can persist as the client count increases, but does not identify the first client-independent contribution at coefficient level. For full-participation stochastic \Scaf{} with one-dimensional homogeneous clients, fixed local-step count $H$, and bounded additive gradient noise, we prove, uniformly over $N\ge2$, $$ \begin{aligned} \mathbb{E}_{\pi_{\gamma,N,H}}[x]-x^\star ={}& -\frac{f'''(x^\star)\sigma^2}{4f''(x^\star)^2}\frac{\gamma}{N}\\ &- \frac{f'''(x^\star)\sigma^2}{12f''(x^\star)} \frac{(H-1)(5H-1)}{H}\gamma^2 +O_H\!\left(\frac{\gamma^2}{N}+\gamma^3\right). \end{aligned} $$ Hence client averaging suppresses the leading $O(\gamma/N)$ bias but does not remove the client-independent $O(\gamma^2)$ component when its coefficient is nonzero. The mechanism is indirect: although the direct control contribution cancels pathwise in the linear global average, the controls still alter within-round local trajectories and their second moments. Fresh gradient noise and persistent control fluctuations therefore generate local second-moment corrections that nonquadratic curvature converts into stationary mean bias. The coefficient vanishes for quadratic objectives. Numerical experiments are consistent with the predicted coefficient, its persistence as client count increases, and the stated joint remainder. The result is restricted to the one-dimensional homogeneous fixed-$H$ setting.
Original Article
View Cached Full Text

Cached at: 08/28/26, 09:44 AM

# Beyond Client Averaging: A Client-Independent Second-Order Stationary-Bias Component in Stochastic SCAFFOLD
Source: [https://arxiv.org/html/2608.26765](https://arxiv.org/html/2608.26765)
###### Abstract

Existing constant\-step analysis of stochasticSCAFFOLDidentifies a leadingO⁡\(γ/N\)O\(\\gamma/N\)stationary mean bias and shows that higher\-order bias can persist as the client count increases, but does not identify the first client\-independent contribution at coefficient level\. For full\-participation stochasticSCAFFOLDwith one\-dimensional homogeneous clients, fixed local\-step countHH, and bounded additive gradient noise, we prove, uniformly overN≥2N\\geq 2,

𝔼πγ,N,H​\[x\]−x⋆=\\displaystyle\\mathbb\{E\}\_\{\\pi\_\{\\gamma,N,H\}\}\[x\]\-x^\{\\star\}=\{\}−f′′′​\(x⋆\)​σ24​f′′​\(x⋆\)2​γN\\displaystyle\-\\frac\{f^\{\\prime\\prime\\prime\}\(x^\{\\star\}\)\\sigma^\{2\}\}\{4f^\{\\prime\\prime\}\(x^\{\\star\}\)^\{2\}\}\\frac\{\\gamma\}\{N\}−f′′′​\(x⋆\)​σ212​f′′​\(x⋆\)​\(H−1\)​\(5​H−1\)H​γ2\+OH​\(γ2N\+γ3\)\.\\displaystyle\-\\frac\{f^\{\\prime\\prime\\prime\}\(x^\{\\star\}\)\\sigma^\{2\}\}\{12f^\{\\prime\\prime\}\(x^\{\\star\}\)\}\\frac\{\(H\-1\)\(5H\-1\)\}\{H\}\\gamma^\{2\}\+O\_\{H\}\\\!\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\)\.Hence client averaging suppresses the leadingO⁡\(γ/N\)O\(\\gamma/N\)bias but does not remove the client\-independentO⁡\(γ2\)O\(\\gamma^\{2\}\)component when its coefficient is nonzero\.

The mechanism is indirect: although the direct control contribution cancels pathwise in the linear global average, the controls still alter within\-round local trajectories and their second moments\. Fresh gradient noise and persistent control fluctuations therefore generate local second\-moment corrections that nonquadratic curvature converts into stationary mean bias\. The coefficient vanishes for quadratic objectives\. Numerical experiments are consistent with the predicted coefficient, its persistence as client count increases, and the stated joint remainder\. The result is restricted to the one\-dimensional homogeneous fixed\-HHsetting\.

## 1Introduction

Federated learning trades communication for local computation: each participating client performs several stochastic\-gradient steps before the server averages the resulting updates\. This design can substantially reduce synchronization, but the local trajectories may drift toward client\-specific directions\.SCAFFOLDaddresses this problem by attaching control variates to local updates, enabling clients to track a common optimization direction more closely\[[6](https://arxiv.org/html/2608.26765#bib.bib1)\]\. Most analyses study finite\-time convergence\. Under a constant step size, however, stochastic iterates continue to fluctuate around the optimum, and their long\-run behavior is described by an invariant distribution whose mean need not equal the minimizer\.

Recent constant\-step analysis of stochasticSCAFFOLDmakes the usual “more clients help” intuition precise and identifies its higher\-order limit\.[Mangold et al\. \[12\]](https://arxiv.org/html/2608.26765#bib.bib2)establish stationarity, a client\-number linear speed\-up up to higher\-order terms, and a higher\-order stationary\-bias component that is not eliminated by increasing the client count\. In the scalar homogeneous specialization, their leading stationary mean\-bias layer is of orderγ/N\\gamma/N, whereγ\\gammais the step size andNNis the number of participating clients\. Still, their analysis leaves the client\-independent higher\-order contribution unresolved at the coefficient level\. We ask whether client averaging suppresses that first client\-independent contribution and, if not, which term first carries it and what mechanism generates it\.

Our answer is negative in this setting\. For full\-participation stochasticSCAFFOLDwith one\-dimensional homogeneous clients, fixed local\-step countHH, sufficiently small constant step sizeγ\\gamma, and bounded additive gradient noise, we prove

𝔼πγ,N,H​\[x\]−x⋆=−τ​σ24​a2​γN−τ​σ212​a​\(H−1\)​\(5​H−1\)H​γ2\+OH​\(γ2N\+γ3\),\\mathbb\{E\}\_\{\\pi\_\{\\gamma,N,H\}\}\[x\]\-x^\{\\star\}=\-\\frac\{\\tau\\sigma^\{2\}\}\{4a^\{2\}\}\\frac\{\\gamma\}\{N\}\-\\frac\{\\tau\\sigma^\{2\}\}\{12a\}\\frac\{\(H\-1\)\(5H\-1\)\}\{H\}\\gamma^\{2\}\+O\_\{H\}\\\!\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\),\(1\)wherea=f′′​\(x⋆\)a=f^\{\\prime\\prime\}\(x^\{\\star\}\)andτ=f′′′​\(x⋆\)\\tau=f^\{\\prime\\prime\\prime\}\(x^\{\\star\}\)\. The first displayed layer decreases as the client count increases; the second does not\. The uniform remainder further shows that, after the knownγ/N\\gamma/Nlayer is removed, the normalized residual converges to the displayedγ2\\gamma^\{2\}coefficient along every joint sequenceN→∞N\\to\\inftyandγ→0\\gamma\\to 0\.

#### Why this matters\.

Client averaging reduces both stationary fluctuations and the known leading mean\-bias layer, which suggests that increasing participation should continually improve long\-run accuracy\. Equation \([1](https://arxiv.org/html/2608.26765#S1.E1)\) identifies a limit of that mechanism within the expansion: when its coefficient is nonzero, the client\-independentγ2\\gamma^\{2\}component is not reduced by further averaging after theγ/N\\gamma/Nlayer becomes smaller\. IncreasingNNalone cannot reduce this component; decreasing the step size does so, and coefficient\-aware bias correction is a natural direction for future work\.

LetYYdenote a local iterate obtained by sampling uniformly over clientsccand local\-step indicesh=0,…,H−1h=0,\\ldots,H\-1under stationarity\. The exact stationary mean\-balance identity gives𝔼​\[f′​\(Y\)\]=0\\mathbb\{E\}\[f^\{\\prime\}\(Y\)\]=0, and a local Taylor expansion then gives the heuristic

0=𝔼⁡\[f′​\(Y\)\]≈a​𝔼​\[Y\]\+τ2​𝔼​\[Y2\],𝔼⁡\[Y\]≈−τ2​a​𝔼​\[Y2\]\.0=\\mathbb\{E\}\[f^\{\\prime\}\(Y\)\]\\approx a\\,\\mathbb\{E\}\[Y\]\+\\frac\{\\tau\}\{2\}\\mathbb\{E\}\[Y^\{2\}\],\\qquad\\mathbb\{E\}\[Y\]\\approx\-\\frac\{\\tau\}\{2a\}\\mathbb\{E\}\[Y^\{2\}\]\.Thus, the second\-moment\-to\-mean intuition follows directly from exact stationary balance plus a local Taylor expansion: nonquadratic curvature converts a second\-moment correction into a mean shift\. InSCAFFOLD, the controls cancel out in the linear client average, but they still affect the intermediate local trajectories\. Fresh gradient noise and persistent control fluctuations therefore modify the second moments of local trajectories before curvature converts them into stationary bias;[Figure1](https://arxiv.org/html/2608.26765#S1.F1)summarizes this indirect channel\.

![Refer to caption](https://arxiv.org/html/2608.26765v1/figures/scaffold_bias_mechanism.png)Figure 1:Mechanism behind the client\-independent stationary\-bias layer\. Linear control cancellation at the server does not erase the controls’ effect on intermediate local trajectories\. Fresh noise and persistent control fluctuations alter local second moments, which nonquadratic curvature converts into a mean shift\.The contributions are as follows\.

- •A limit to client averaging at second order\.We identify an explicit client\-independentγ2\\gamma^\{2\}stationary\-bias coefficient, with a remainder uniform jointly in\(γ,1/N\)\(\\gamma,1/N\)for fixedHH\.
- •A trajectory\-level mechanism insideSCAFFOLD\.We trace the coefficient to two local second\-moment sources: fresh stochastic\-gradient noise and persistent control fluctuations\.
- •Targeted numerical checks\.We test persistence with increasing client count, convergence toward the predicted coefficient, and behavior when stochasticity or nonquadratic curvature is removed\.

The scalar analysis is informative beyond one dimension because it isolates three ingredients that also occur in higher\-dimensional dynamics: persistent control second moments, within\-round local trajectories, and nonlinear curvature\. Their interaction would involve covariance–third\-derivative contractions; establishing such an extension remains open\.

#### Claim boundaries\.

The theorem concerns full participation, one\-dimensional homogeneous clients, bounded additive noise, and fixedHH\. It is uniform inNNand smallγ\\gamma, but not asH→∞H\\to\\infty, and it does not identify the complete finite\-NNcoefficient ofγ2\\gamma^\{2\}\. The source decomposition is internal toSCAFFOLDand is not a theorem\-level comparison with FedAvg\. The crossover of the two displayed layers is an asymptotic scale comparison, not a monotonicity statement for the total finite\-step bias\. The numerical experiments are finite\-setting consistency checks\. Multidimensional, heterogeneous\-client, state\-dependent\-noise, and debiasing extensions remain open\.

[Section2](https://arxiv.org/html/2608.26765#S2)positions the result\.[Section3](https://arxiv.org/html/2608.26765#S3)connects the standardSCAFFOLDcontrols to the zero\-sum parametrization used in the proof\.[Sections4](https://arxiv.org/html/2608.26765#S4)and[5](https://arxiv.org/html/2608.26765#S5)state the theorem and derive its mechanism, while[Sections6](https://arxiv.org/html/2608.26765#S6)and[7](https://arxiv.org/html/2608.26765#S7)summarize the proof and numerical evidence\.

## 2Related work

#### Federated averaging and local SGD\.

Federated Averaging \(FedAvg\) introduced iterative local model averaging as a communication\-efficient primitive for federated learning\[[13](https://arxiv.org/html/2608.26765#bib.bib9)\]; the broader optimization and systems landscape is surveyed by[Kairouz et al\. \[5\]](https://arxiv.org/html/2608.26765#bib.bib16)\. Its optimization core is closely related to LocalSGD, in which workers take several stochastic gradient steps before synchronization\. Early theory established that LocalSGD can retain the statistical rate of minibatch SGD while communicating less\[[14](https://arxiv.org/html/2608.26765#bib.bib10)\], and subsequent analyses clarified the distinct roles of identical and heterogeneous client objectives\[[7](https://arxiv.org/html/2608.26765#bib.bib11)\]and the regimes in which local updates do or do not improve on minibatch SGD\[[16](https://arxiv.org/html/2608.26765#bib.bib12)\]\.

#### SCAFFOLD and control correction\.

SCAFFOLDaugments local updates with control variates to correct client drift\[[6](https://arxiv.org/html/2608.26765#bib.bib1)\]\. More recent work sharpens the finite\-time picture:[Luo et al\. \[8\]](https://arxiv.org/html/2608.26765#bib.bib3)obtain improved LocalSGD andSCAFFOLDrates under gradient/Hessian similarity and higher\-order smoothness conditions, while[Mangold and Moulines \[11\]](https://arxiv.org/html/2608.26765#bib.bib5)give a refined quadratic analysis for arbitrary local\-step counts\. These results explain when local computation and control correction improve finite\-time optimization, but they do not characterize the mean of the invariant distribution generated by constant\-step stochasticSCAFFOLD\.

#### Constant\-step stochastic optimization and stationary laws\.

Under suitable stability and regularity conditions, constant\-step stochastic\-gradient methods are naturally studied in terms of an invariant distribution\. This viewpoint appears in diffusion approximations of constant\-step SGD\[[9](https://arxiv.org/html/2608.26765#bib.bib14)\]and in rigorous small\-step characterizations of stationary stochastic\-gradient dynamics\[[3](https://arxiv.org/html/2608.26765#bib.bib13),[2](https://arxiv.org/html/2608.26765#bib.bib15)\]\. In particular,[Chen et al\. \[2\]](https://arxiv.org/html/2608.26765#bib.bib15)obtain Gaussian/Lyapunov characterizations of scaled stationary laws for smooth strongly convex SGD and related stochastic\-approximation models under their conditions, while[Dieuleveut et al\. \[3\]](https://arxiv.org/html/2608.26765#bib.bib13)derive stationary/asymptotic moment expansions and use Richardson–Romberg extrapolation to reduce step\-size bias\. The same perspective has recently been developed for federated and decentralized algorithms\.[Mangold et al\. \[10\]](https://arxiv.org/html/2608.26765#bib.bib4)analyze constant\-step FedAvg through its Markov structure, characterize stationary bias and variance, derive a first\-order bias expansion, and construct a federated Richardson–Romberg correction\.[Mangold et al\. \[12\]](https://arxiv.org/html/2608.26765#bib.bib2)extend this framework to stochasticSCAFFOLD, proving existence and geometric convergence of a stationary law, client\-number variance reduction, and a first\-order stationary mean\-bias expansion whose scalar homogeneous leading layer isO⁡\(γ/N\)O\(\\gamma/N\)\.[Versini et al\. \[15\]](https://arxiv.org/html/2608.26765#bib.bib8)obtain related first\-order stationary bias and variance expansions for decentralized SGD, separating stochastic, heterogeneity, and network effects\.

#### Nonlinearity and higher\-order stationary bias\.

The mechanism studied here belongs to a broader class of noise–nonlinearity interactions in constant\-step stochastic approximation\.[Allmeier and Gast \[1\]](https://arxiv.org/html/2608.26765#bib.bib7)allow Markovian noise to depend on the iterate, prove anO⁡\(α\)O\(\\alpha\)stationary\-bias bound, and refine the Polyak–Ruppert time\-averaged bias toα​V\+O⁡\(α2\)\\alpha V\+O\(\\alpha^\{2\}\)\.[Huo et al\. \[4\]](https://arxiv.org/html/2608.26765#bib.bib6)show that Markovian memory and nonlinearity can interact to create distinct components of the constant\-step stationary bias\. These results establish that persistent stochastic fluctuations can be converted into mean shifts by nonlinear dynamics\. They do not, however, determine the algorithm\-specific higher\-order coefficient generated bySCAFFOLD’s local control recursion\.

#### Position of the present result\.

The closest starting point is[Mangold et al\. \[12\]](https://arxiv.org/html/2608.26765#bib.bib2)\. We use their stationary\-law existence and geometric convergence, coarse iterate\-moment bounds, leading control second moment, and knownO⁡\(γ/N\)O\(\\gamma/N\)bias coefficient\. Mangold et al\. establish that stochasticSCAFFOLDretains a higher\-order stationary bias that is not eliminated by increasing the client count\. Still, their analysis does not identify the client\-independent contribution at the coefficient level\. In the one\-dimensional homogeneous fixed\-HHsetting, we resolve this structure explicitly: the first client\-independent term isN0​γ2N^\{0\}\\gamma^\{2\}, with a closed\-form coefficient that decomposes into fresh within\-round local\-gradient noise and persistentSCAFFOLDcontrol second moments\. The result is complementary to the quadratic theory of[Mangold and Moulines \[11\]](https://arxiv.org/html/2608.26765#bib.bib5): quadratic objectives support sharper arbitrary\-HHconvergence analysis, butf′′′≡0f^\{\\prime\\prime\\prime\}\\equiv 0removes the second\-moment\-to\-mean conversion isolated here\. The source decomposition is internal toSCAFFOLD; it is not a theorem\-level comparison with FedAvg, which would require a matched second\-order FedAvg expansion\.

## 3Setting and exact identities

We deliberately use a homogeneous scalar regime, removing client heterogeneity as a confounding source of bias and isolating the stochastic effect of theSCAFFOLDcontrol recursion itself\. AllNNclients participate in every communication round\. Even with identical objectives, the stochastic client controls remain nontrivial because they are continually refreshed from noisy local trajectories\. The unique minimizer is translated tox⋆=0x^\{\\star\}=0, allN≥2N\\geq 2clients share the same objectivefc=ff\_\{c\}=f, and the number of local stepsH≥2H\\geq 2is fixed\. The objective satisfiesf∈C5​\(ℝ\)f\\in C^\{5\}\(\\mathbb\{R\}\)and

0<μ≤f′′​\(y\)≤L<∞,0<\\mu\\leq f^\{\\prime\\prime\}\(y\)\\leq L<\\infty,with bounded third through fifth derivatives\. We write

a:=f′′​\(0\)\>0,τ:=f′′′​\(0\),κ:=f′′′′​\(0\)\.a:=f^\{\\prime\\prime\}\(0\)\>0,\\qquad\\tau:=f^\{\\prime\\prime\\prime\}\(0\),\\qquad\\kappa:=f^\{\\prime\\prime\\prime\\prime\}\(0\)\.At local stephhof clientcc, the stochastic gradient is

∇Fc,h​\(y\)=f′​\(y\)\+εc,h,\\nabla F\_\{c,h\}\(y\)=f^\{\\prime\}\(y\)\+\\varepsilon\_\{c,h\},where the fresh noises are independent across clients, local steps, and communication rounds, independent of the round\-start state, satisfy𝔼⁡\[εc,h\]=0\\mathbb\{E\}\[\\varepsilon\_\{c,h\}\]=0and𝔼⁡\[εc,h2\]=σ2\\mathbb\{E\}\[\\varepsilon\_\{c,h\}^\{2\}\]=\\sigma^\{2\}, and are almost surely bounded\. No symmetry assumption is imposed\.

At the start of a communication round, the server model isxx, and every client initializes atθc,0=x\\theta\_\{c,0\}=x\. Each client then takesHHcorrected stochastic\-gradient steps\. We first connect the parametrization used below to the standardSCAFFOLDcontrols\.

Equivalent full\-participation SCAFFOLD round\.Letcsrvc\_\{\\rm srv\}be the server control andccc\_\{c\}the control stored by clientcc, withcsrv=N−1​∑cccc\_\{\\rm srv\}=N^\{\-1\}\\sum\_\{c\}c\_\{c\}\. StandardSCAFFOLDuses the correctioncsrv−ccc\_\{\\rm srv\}\-c\_\{c\}in each local step\. Defineξc:=csrv−cc,1N​∑c=1Nξc=0\.\\xi\_\{c\}:=c\_\{\\rm srv\}\-c\_\{c\},\\qquad\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\xi\_\{c\}=0\.With full participation, global mixing parameter one, and the standard Option\-II control refresh\[[6](https://arxiv.org/html/2608.26765#bib.bib1)\],cc\+=cc−csrv\+x−θc,Hγ​H,csrv\+=csrv\+1N​∑j=1N\(cj\+−cj\)\.c\_\{c\}^\{\+\}=c\_\{c\}\-c\_\{\\rm srv\}\+\\frac\{x\-\\theta\_\{c,H\}\}\{\\gamma H\},\\qquad c\_\{\\rm srv\}^\{\+\}=c\_\{\\rm srv\}\+\\frac\{1\}\{N\}\\sum\_\{j=1\}^\{N\}\(c\_\{j\}^\{\+\}\-c\_\{j\}\)\.Usingx\+=N−1​∑jθj,Hx^\{\+\}=N^\{\-1\}\\sum\_\{j\}\\theta\_\{j,H\}givesξc\+=csrv\+−cc\+=ξc\+θc,H−x\+γ​H\.\\xi\_\{c\}^\{\+\}=c\_\{\\rm srv\}^\{\+\}\-c\_\{c\}^\{\+\}=\\xi\_\{c\}\+\\frac\{\\theta\_\{c,H\}\-x^\{\+\}\}\{\\gamma H\}\.Thus the zero\-sum variables\{ξc\}\\\{\\xi\_\{c\}\\\}are an exact reparametrization of the usual server/client controls in the regime analyzed here\.

Using this parametrization, the local, server, and control updates are

θc,h\+1\\displaystyle\\theta\_\{c,h\+1\}=θc,h−γ⁡\{f′​\(θc,h\)\+ξc\+εc,h\+1\},θc,0=x,\\displaystyle=\\theta\_\{c,h\}\-\\gamma\\\{f^\{\\prime\}\(\\theta\_\{c,h\}\)\+\\xi\_\{c\}\+\\varepsilon\_\{c,h\+1\}\\\},\\qquad\\theta\_\{c,0\}=x,\(2\)x\+\\displaystyle x^\{\+\}=1N​∑c=1Nθc,H,\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\theta\_\{c,H\},\(3\)ξc\+\\displaystyle\\xi\_\{c\}^\{\+\}=ξc\+θc,H−x\+γ​H\.\\displaystyle=\\xi\_\{c\}\+\\frac\{\\theta\_\{c,H\}\-x^\{\+\}\}\{\\gamma H\}\.\(4\)Summing Equation \([4](https://arxiv.org/html/2608.26765#S3.E4)\) over clients shows directly that the zero\-sum condition is preserved\.

Two exact identities expose the distinction between global cancellation and local influence\. First, averaging the unrolled local recursion gives

x\+−x=−γ∑h=0H−1f′​\(θh\)¯−γ∑h=0H−1ε¯h\+1,x^\{\+\}\-x=\-\\gamma\\sum\_\{h=0\}^\{H\-1\}\\overline\{f^\{\\prime\}\(\\theta\_\{h\}\)\}\-\\gamma\\sum\_\{h=0\}^\{H\-1\}\\bar\{\\varepsilon\}\_\{h\+1\},\(5\)where bars denote client averages\. The controls cancel pathwise from this linear average\. Second, stationarity gives

0=∑h=0H−11N​∑c=1N𝔼⁡\[f′​\(θc,h\)\]\.0=\\sum\_\{h=0\}^\{H\-1\}\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\[f^\{\\prime\}\(\\theta\_\{c,h\}\)\]\.\(6\)These identities contain the central tension of the paper\. Equation \([5](https://arxiv.org/html/2608.26765#S3.E5)\) rules out a direct linear control contribution to the global update, but it does not erase the effect ofξc\\xi\_\{c\}on the intermediate trajectoriesθc,h\\theta\_\{c,h\}\. Equation \([6](https://arxiv.org/html/2608.26765#S3.E6)\) is where those trajectory\-level changes re\-enter the stationary mean through the nonlinear mapf′f^\{\\prime\}\. Formal assumptions and the earlier stationary results used below are stated in Appendix[A](https://arxiv.org/html/2608.26765#A1)\.

## 4Main result: a client\-independent second\-order layer

Letπγ,N,H\\pi\_\{\\gamma,N,H\}denote the unique stationary law of the global/control process, whose existence for sufficiently smallγ\\gammafollows from[PropositionA\.3](https://arxiv.org/html/2608.26765#A1.Thmtheorem3), and define

bγ,N,H:=𝔼πγ,N,H​\[x\]−x⋆\.b\_\{\\gamma,N,H\}:=\\mathbb\{E\}\_\{\\pi\_\{\\gamma,N,H\}\}\[x\]\-x^\{\\star\}\.
###### Theorem 4\.1\(Uniform joint stationary\-bias expansion\)\.

Under the assumptions of[Section3](https://arxiv.org/html/2608.26765#S3), including full participation, for fixedH≥2H\\geq 2there existγ0,H\>0\\gamma\_\{0,H\}\>0andCH<∞C\_\{H\}<\\infty, independent ofNNandγ\\gamma, such that for everyN≥2N\\geq 2and0<γ≤γ0,H0<\\gamma\\leq\\gamma\_\{0,H\},

bγ,N,H=−τ​σ24​a2​γN−τ​σ212​a​\(H−1\)​\(5​H−1\)H​γ2\+Rγ,N,H\\boxed\{b\_\{\\gamma,N,H\}=\-\\frac\{\\tau\\sigma^\{2\}\}\{4a^\{2\}\}\\frac\{\\gamma\}\{N\}\-\\frac\{\\tau\\sigma^\{2\}\}\{12a\}\\frac\{\(H\-1\)\(5H\-1\)\}\{H\}\\gamma^\{2\}\+R\_\{\\gamma,N,H\}\}\(7\)with

\|Rγ,N,H\|≤CH​\(γ2N\+γ3\)\.\\boxed\{\|R\_\{\\gamma,N,H\}\|\\leq C\_\{H\}\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\)\.\}\(8\)

The client\-independentγ2\\gamma^\{2\}coefficient is

B20SCAF​\(H\)=−f′′′​\(x⋆\)​σ212​f′′​\(x⋆\)​\(H−1\)​\(5​H−1\)H\.\\boxed\{B\_\{20\}^\{\\mathrm\{SCAF\}\}\(H\)=\-\\frac\{f^\{\\prime\\prime\\prime\}\(x^\{\\star\}\)\\sigma^\{2\}\}\{12f^\{\\prime\\prime\}\(x^\{\\star\}\)\}\\frac\{\(H\-1\)\(5H\-1\)\}\{H\}\.\}\(9\)The subscript2020records the bivariate orderγ2​N0\\gamma^\{2\}N^\{0\}: second order in the step size and independent of the client count\.

### 4\.1Interpretation for client averaging and local computation

#### More clients suppress only the leading displayed layer\.

The knownO⁡\(γ/N\)O\(\\gamma/N\)contribution decreases at rate1/N1/N, whereas the displayedγ2\\gamma^\{2\}contribution is independent ofNN\. Thus, within the theorem’s stationary and small\-step regime, client averaging suppresses the leading bias layer but does not remove the client\-independent second\-order layer when its coefficient is nonzero\. Whenτ​σ2≠0\\tau\\sigma^\{2\}\\neq 0, equating the magnitudes of the two displayed terms gives

N×=3​Ha​\(H−1\)​\(5​H−1\)​γ−1\.N\_\{\\times\}=\\frac\{3H\}\{a\(H\-1\)\(5H\-1\)\}\\,\\gamma^\{\-1\}\.The two displayed coefficients then have the same sign\. This is a crossover scale for the two displayed asymptotic layers, not an exact finite\-step phase transition or a monotonicity statement for the full bias;[Figure2](https://arxiv.org/html/2608.26765#S4.F2)visualizes this scale comparison\.

number of participating clientsNNschematic magnitude of displayed bias layersN×N\_\{\\times\}Displayed\-layer crossoverN×=3​Ha​\(H−1\)​\(5​H−1\)​γ−1\\displaystyle N\_\{\\times\}=\\frac\{3H\}\{a\(H\-1\)\(5H\-1\)\}\\,\\gamma^\{\-1\}O⁡\(γ/N\)O\(\\gamma/N\)componentdecreases withNNclient\-independentO⁡\(γ2\)O\(\\gamma^\{2\}\)componentdoes not decay withNNFigure 2:Schematic magnitude of the two displayed terms in[Theorem4\.1](https://arxiv.org/html/2608.26765#S4.Thmtheorem1)\. TheO⁡\(γ/N\)O\(\\gamma/N\)component decreases withNN, whereas the client\-independentO⁡\(γ2\)O\(\\gamma^\{2\}\)component does not\. When both are nonzero, equating their displayed coefficients gives the marked crossover scale\. It is not an exact finite\-step threshold or a curve for the total bias\.
#### The coefficient records a local\-computation effect\.

For fixed problem parameters, the factor

\(H−1\)​\(5​H−1\)H\\frac\{\(H\-1\)\(5H\-1\)\}\{H\}is increasing over integersH≥2H\\geq 2, while the theorem treats eachHHas fixed\. The coefficient therefore links the client\-independent stationary effect to computation performed inside each communication round\.

#### Quadratic models hide the mechanism\.

When the objective is quadratic,f′′′​\(x⋆\)=0f^\{\\prime\\prime\\prime\}\(x^\{\\star\}\)=0and both displayed coefficients vanish\. A purely quadratic analysis can therefore characterize stability and convergence sharply while remaining blind to the second\-moment\-to\-mean conversion isolated here\.

### 4\.2Precise joint\-limit meaning

The uniform remainder gives the exact asymptotic interpretation\. After subtracting the knownγ/N\\gamma/Nlayer and normalizing byγ2\\gamma^\{2\},

\|bγ,N,H\+τ​σ24​a2​γNγ2−B20SCAF​\(H\)\|≤CH​\(1N\+γ\)\.\\left\|\\frac\{b\_\{\\gamma,N,H\}\+\\frac\{\\tau\\sigma^\{2\}\}\{4a^\{2\}\}\\frac\{\\gamma\}\{N\}\}\{\\gamma^\{2\}\}\-B\_\{20\}^\{\\mathrm\{SCAF\}\}\(H\)\\right\|\\leq C\_\{H\}\\left\(\\frac\{1\}\{N\}\+\\gamma\\right\)\.\(10\)Hence, the normalized residual converges to the value in Equation \([9](https://arxiv.org/html/2608.26765#S4.E9)\) along any joint sequenceN→∞N\\to\\infty,γ→0\\gamma\\to 0, without a relative\-rate condition\. Terms of orderγ2/N\\gamma^\{2\}/Nremain in the remainder, so this identifies the client\-independent coefficient but not the complete fixed\-NNsecond\-order expansion\.

## 5Mechanism: how controls survive linear cancellation

[Figure1](https://arxiv.org/html/2608.26765#S1.F1)gives the qualitative picture\. The coefficient follows from making the two sources of its fluctuations precise\.

#### Linear aggregation\.

Under the zero\-sum control state, Equation \([5](https://arxiv.org/html/2608.26765#S3.E5)\) shows that the client\-average control term cancels pathwise from the global update\. There is no direct linear control term shifting the server model\.

#### Within\-round local dynamics\.

Eachξc\\xi\_\{c\}nevertheless changes the intermediate trajectory

θc,0,θc,1,…,θc,H,\\theta\_\{c,0\},\\theta\_\{c,1\},\\ldots,\\theta\_\{c,H\},and hence the points at which subsequent stochastic gradients are evaluated\. The nonlinear local updates act along these control\-dependent paths before the server observes their average\.

#### Two second\-moment sources\.

Let

Qγ,N,H:=1N​∑c=1N𝔼⁡\[ξc2\]Q\_\{\\gamma,N,H\}:=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\[\\xi\_\{c\}^\{2\}\]be the stationary average control second moment\. In the homogeneous additive\-noise setting,

Qγ,N,H=σ2H​\(1−1N\)\+OH​\(γ\)\.Q\_\{\\gamma,N,H\}=\\frac\{\\sigma^\{2\}\}\{H\}\\left\(1\-\\frac\{1\}\{N\}\\right\)\+O\_\{H\}\(\\gamma\)\.\(11\)The controls therefore retain nonzero fluctuations as the number of clients grows\. For the change in the local second moment relative to the round\-start server model,

U¯h:=1N​∑c=1N\(𝔼⁡\[θc,h2\]−𝔼⁡\[x2\]\),\\bar\{U\}\_\{h\}:=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\bigl\(\\mathbb\{E\}\[\\theta\_\{c,h\}^\{2\}\]\-\\mathbb\{E\}\[x^\{2\}\]\\bigr\),we obtain

U¯h=γ2​\(h​σ2\+h2​Qγ,N,H\)\+OH​\(γ2N\+γ3\)\.\\bar\{U\}\_\{h\}=\\gamma^\{2\}\\bigl\(h\\sigma^\{2\}\+h^\{2\}Q\_\{\\gamma,N,H\}\\bigr\)\+O\_\{H\}\\\!\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\)\.\(12\)The first displayed term is generated by fresh within\-round gradient noise\. The persistent round\-start controls generate the second\. Because Equation \([11](https://arxiv.org/html/2608.26765#S5.E11)\) approachesσ2/H\\sigma^\{2\}/HasNNgrows, both sources contribute at the client\-independentγ2\\gamma^\{2\}scale\.

#### Nonlinear conversion\.

Expanding the gradient near the optimum,

f′​\(y\)=a​y\+τ2​y2\+O⁡\(y3\),f^\{\\prime\}\(y\)=ay\+\\frac\{\\tau\}\{2\}y^\{2\}\+O\(y^\{3\}\),and inserting the local moments into the exact stationary balance in Equation \([6](https://arxiv.org/html/2608.26765#S3.E6)\) converts the second\-moment correction into a mean shift\. The resulting coefficient separates into

B20local​\(H\)\\displaystyle B\_\{20\}^\{\\mathrm\{local\}\}\(H\)=−τ​σ24​a​\(H−1\),\\displaystyle=\-\\frac\{\\tau\\sigma^\{2\}\}\{4a\}\(H\-1\),\(13\)B20ctrl​\(H\)\\displaystyle B\_\{20\}^\{\\mathrm\{ctrl\}\}\(H\)=−τ​σ212​a​\(H−1\)​\(2​H−1\)H,\\displaystyle=\-\\frac\{\\tau\\sigma^\{2\}\}\{12a\}\\frac\{\(H\-1\)\(2H\-1\)\}\{H\},\(14\)B20SCAF​\(H\)\\displaystyle B\_\{20\}^\{\\mathrm\{SCAF\}\}\(H\)=B20local​\(H\)\+B20ctrl​\(H\)\.\\displaystyle=B\_\{20\}^\{\\mathrm\{local\}\}\(H\)\+B\_\{20\}^\{\\mathrm\{ctrl\}\}\(H\)\.\(15\)The first component records direct local\-gradient noise; the second records the additional second moment carried by theSCAFFOLDcontrols\. Their sum is the coefficient in Equation \([9](https://arxiv.org/html/2608.26765#S4.E9)\)\.

## 6Proof overview

The coefficient suggested by a formal Taylor expansion is not automatically identifiable\. The stationary balance contains the global second moment, within\-round local second\-moment corrections, and higher signed moments\. Any one of these quantities could, in principle, carry an additional client\-independentγ2\\gamma^\{2\}contribution\. The proof must therefore isolate the two intended local sources of fluctuation and rule out all competing channels at the same scale\.

The argument proceeds in the order shown in the dependency map in[Figure4](https://arxiv.org/html/2608.26765#A8.F4)of Appendix[H](https://arxiv.org/html/2608.26765#A8)\. Coarse moment bounds first place the global iterate, controls, and local displacements on scales that are uniform inNN\. The local recursion is then expanded sharply enough to expose the two terms in Equation \([12](https://arxiv.org/html/2608.26765#S5.E12)\): fresh noise and the round\-start control second moment\. Before these terms can be converted to the stationary mean, the proof controls the global fourth moment and the signed third moment, ensuring that cubic and quartic Taylor contributions remain within the target remainder\.

The key identifiability step is the sharp global second moment

𝔼⁡\[x2\]=σ22​a​γN\+OH​\(γ2N\+γ3\)\.\\mathbb\{E\}\[x^\{2\}\]=\\frac\{\\sigma^\{2\}\}\{2a\}\\frac\{\\gamma\}\{N\}\+O\_\{H\}\\\!\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\)\.\(16\)A coarserO⁡\(γ/N\+γ2\)O\(\\gamma/N\+\\gamma^\{2\}\)estimate would leave open an additional client\-independentγ2\\gamma^\{2\}term in𝔼⁡\[x2\]\\mathbb\{E\}\[x^\{2\}\], which would contaminate the coefficient attributed to local trajectories\. Establishing Equation \([16](https://arxiv.org/html/2608.26765#S6.E16)\) rules out that competing source\.

The delicate covariance in this step couples fresh round noise with the nonlinear increment\. A generic Cauchy–Schwarz bound is too loose because it loses the extra factor1/N1/Nrequired by the joint remainder\. The proof uses a coordinate\-replacement argument that sets one fresh\-noise coordinate to zero\. Only one client’s within\-round path changes, and the subsequent client average contributes an additional1/N1/Nsensitivity\. Summing over all noise coordinates then recovers the required scale\.

With the global second moment sharpened and the local cubic and quartic terms placed inside the remainder, the exact stationary balance reduces to

bγ,N,H=−τ2​a​𝔼​\[x2\]−τ2​a​H​∑h=0H−1U¯h\+OH​\(γ2N\+γ3\)\.b\_\{\\gamma,N,H\}=\-\\frac\{\\tau\}\{2a\}\\mathbb\{E\}\[x^\{2\}\]\-\\frac\{\\tau\}\{2aH\}\\sum\_\{h=0\}^\{H\-1\}\\bar\{U\}\_\{h\}\+O\_\{H\}\\\!\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\)\.Substituting Equations \([16](https://arxiv.org/html/2608.26765#S6.E16)\) and \([12](https://arxiv.org/html/2608.26765#S5.E12)\), then summinghhandh2h^\{2\}, yields[Theorem4\.1](https://arxiv.org/html/2608.26765#S4.Thmtheorem1)\. The full proof begins in Appendix[A](https://arxiv.org/html/2608.26765#A1)and is assembled in Appendix[H](https://arxiv.org/html/2608.26765#A8)\.

## 7Numerical checks of the predicted effect

The numerical study follows the paper’s main claim in the same order: first, whether the debiased statistic remains nonzero as the client count grows; second, whether it approaches the predicted coefficient as\(γ,1/N\)→\(0,0\)\(\\gamma,1/N\)\\to\(0,0\); and third, whether the observations are consistent with removing the effect when stochasticity or nonquadratic curvature is absent\.

We use the smooth, strongly convex objective

f′​\(x\)=x\+12​log⁡cosh⁡\(x\),f^\{\\prime\}\(x\)=x\+\\frac\{1\}\{2\}\\log\\cosh\(x\),\(17\)for which

12≤f′′​\(x\)=1\+12​tanh⁡\(x\)≤32,a=1,τ=12\.\\frac\{1\}\{2\}\\leq f^\{\\prime\\prime\}\(x\)=1\+\\frac\{1\}\{2\}\\tanh\(x\)\\leq\\frac\{3\}\{2\},\\qquad a=1,\\qquad\\tau=\\frac\{1\}\{2\}\.The additive noise is independent Rademacher noise with variance one\. The predicted client\-independent coefficient becomes

B20​\(H\)=−\(H−1\)​\(5​H−1\)24​H,B\_\{20\}\(H\)=\-\\frac\{\(H\-1\)\(5H\-1\)\}\{24H\},\(18\)soB20​\(2\)=−0\.1875B\_\{20\}\(2\)=\-0\.1875andB20​\(4\)=−0\.59375B\_\{20\}\(4\)=\-0\.59375\.

We estimate the stationary mean both directly and through the exact mean\-balance identity\. Unless stated otherwise,b^γ,N,H\\hat\{b\}\_\{\\gamma,N,H\}in the main text denotes the mean\-balance estimate; the direct sample mean is retained as a diagnostic\. The primary statistic removes the known leading layer and normalizes the residual:

Cγ,N,H:=b^γ,N,H\+18​γNγ2\.C\_\{\\gamma,N,H\}:=\\frac\{\\hat\{b\}\_\{\\gamma,N,H\}\+\\frac\{1\}\{8\}\\frac\{\\gamma\}\{N\}\}\{\\gamma^\{2\}\}\.\(19\)Thus, the theorem predicts that this normalized residual should approach a nonzero constant asNNgrows andγ\\gammadecreases\. By[Theorem4\.1](https://arxiv.org/html/2608.26765#S4.Thmtheorem1),

Cγ,N,H=B20​\(H\)\+OH​\(1N\+γ\)\.C\_\{\\gamma,N,H\}=B\_\{20\}\(H\)\+O\_\{H\}\\\!\\left\(\\frac\{1\}\{N\}\+\\gamma\\right\)\.\(20\)Each setting uses eight independent chains, and uncertainty is reported by 95% Student\-ttintervals across chain estimates\. The complete burn\-in, batching, estimator agreement, and stability protocol are given in Appendix[J](https://arxiv.org/html/2608.26765#A10)\.

### 7\.1Does the component survive increasing client count?

We fixH=2H=2andγ=1/24\\gamma=1/24and increaseN∈\{8,16,32,64,128\}N\\in\\\{8,16,32,64,128\\\}\. After removing the knownγ/N\\gamma/Ncontribution, the normalized residual approaches a nonzero plateau near the predicted value−0\.1875\-0\.1875; see[Figure3](https://arxiv.org/html/2608.26765#S7.F3)\. A regression in1/N1/Ngives large\-NNintercept−0\.189145\-0\.189145with 95% bootstrap interval\[−0\.190398,−0\.188035\]\[\-0\.190398,\-0\.188035\]\. This is the numerical check that most directly mirrors the title claim: averaging suppresses the leading layer but does not remove the client\-independent component\.

![Refer to caption](https://arxiv.org/html/2608.26765v1/figures/client_count_scaling.png)Figure 3:Client\-count persistence atH=2H=2andγ=1/24\\gamma=1/24\. The normalized residual approaches a nonzero plateau; largerNNlies to the right on the horizontal axis\.
### 7\.2Does the normalized residual approach the coefficient?

Along the joint pathN=1/γN=1/\\gamma, Equation \([20](https://arxiv.org/html/2608.26765#S7.E20)\) permits anOH​\(γ\)O\_\{H\}\(\\gamma\)error\. The coarse\-gridH=4H=4offset is therefore compatible with finite\-step effects allowed by the theorem; the smaller\-step experiment tests whether the discrepancy contracts at the predicted order\. ForH=2H=2, the initial gridγ∈\{1/12,1/16,1/24,1/32,1/48\}\\gamma\\in\\\{1/12,1/16,1/24,1/32,1/48\\\}gives extrapolated intercept−0\.187214\-0\.187214with 95% bootstrap interval\[−0\.189250,−0\.185175\]\[\-0\.189250,\-0\.185175\], consistent withB20​\(2\)B\_\{20\}\(2\)\.

The same coarse grid atH=4H=4showed a measurable finite\-step offset, so we ran a separately specified smaller\-step study with fresh chains at

\(γ,N\)∈\{\(1/64,64\),\(1/96,96\),\(1/128,128\),\(1/192,192\)\}\.\(\\gamma,N\)\\in\\\{\(1/64,64\),\(1/96,96\),\(1/128,128\),\(1/192,192\)\\\}\.The normalized statistic moves towardB20​\(4\)B\_\{20\}\(4\), while the scaled error\{Cγ,1/γ,4−B20​\(4\)\}/γ\\\{C\_\{\\gamma,1/\\gamma,4\}\-B\_\{20\}\(4\)\\\}/\\gammaremains between1\.381\.38and1\.521\.52; see[Figure6](https://arxiv.org/html/2608.26765#A10.F6)in Appendix[J](https://arxiv.org/html/2608.26765#A10)\. A linear fit over these four settings yields an intercept of−0\.592953\-0\.592953with a 95% bootstrap interval of\[−0\.595526,−0\.590236\]\[\-0\.595526,\-0\.590236\]\. The smaller\-step figure, initial discrepancy, follow\-up criteria, and numerical table are retained in Appendix[J](https://arxiv.org/html/2608.26765#A10)\.

### 7\.3Do the required ingredients matter?

The supplementary negative controls remove both ingredients from the mechanism\. With deterministic gradients, the stationary\-mean intervals contain zero\. With a quadratic objective, wheref′′′​\(x⋆\)=0f^\{\\prime\\prime\\prime\}\(x^\{\\star\}\)=0, direct\-mean intervals also contain zero\. The settings, intervals, and an additionalH=1H=1boundary check are reported in Appendix[J](https://arxiv.org/html/2608.26765#A10)\.

Together, the experiments are consistent with three distinct implications of the expansion: persistence as the client count increases, convergence toward the predicted coefficient, and dependence on stochastic, nonquadratic local dynamics\.

## 8Conclusion

In the homogeneous scalar fixed\-HHsetting, stochasticSCAFFOLDhas two stationary\-bias layers: the knownO⁡\(γ/N\)O\(\\gamma/N\)contribution, which client averaging suppresses, and a client\-independentO⁡\(γ2\)O\(\\gamma^\{2\}\)contribution that is unaffected by the client count when its coefficient is nonzero\. Fresh gradient noise and persistent control fluctuations generate the latter through local second moments, and nonquadratic curvature converts it into a stationary mean shift\.

Control cancellation is only linear: the controls vanish from the server average but still change local trajectories\. Numerically, the residual forms a client\-count plateau, approaches the coefficient along smaller\-step paths, and disappears without stochasticity or nonlinearity\. In higher dimensions, the scalar second\-moment\-to\-mean conversion would be replaced by contractions between stationary covariance structure and the third\-derivative tensor, while client heterogeneity may introduce additional weighted moment channels absent from the present homogeneous argument\. Multidimensional analysis, client heterogeneity, and coefficient\-based bias cancellation remain open\.

## References

- \[1\]S\. Allmeier and N\. Gast\(2024\)Computing the bias of constant\-step stochastic approximation with Markovian noise\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 137873–137902\.External Links:[Document](https://dx.doi.org/10.52202/079017-4379),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/f949c1f490beb42124a267b7476cd353-Abstract-Conference.html)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px4.p1.1)\.
- \[2\]Z\. Chen, S\. Mou, and S\. T\. Maguluri\(2022\)Stationary behavior of constant stepsize SGD\-type algorithms: an asymptotic characterization\.Proceedings of the ACM on Measurement and Analysis of Computing Systems6\(1\),pp\. 19:1–19:24\.External Links:[Document](https://dx.doi.org/10.1145/3508039),[Link](https://doi.org/10.1145/3508039)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px3.p1.1)\.
- \[3\]A\. Dieuleveut, A\. Durmus, and F\. Bach\(2020\)Bridging the gap between constant step size stochastic gradient descent and Markov chains\.The Annals of Statistics48\(3\),pp\. 1348–1382\.External Links:[Document](https://dx.doi.org/10.1214/19-AOS1850),[Link](https://doi.org/10.1214/19-AOS1850)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px3.p1.1)\.
- \[4\]D\. L\. Huo, Y\. Zhang, Y\. Chen, and Q\. Xie\(2024\)The collusion of memory and nonlinearity in stochastic approximation with constant stepsize\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 21699–21762\.External Links:[Document](https://dx.doi.org/10.52202/079017-0684),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/2676109d49d1eb26d6bc584a8f556305-Abstract-Conference.html)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px4.p1.1)\.
- \[5\]P\. Kairouz, H\. B\. McMahan, B\. Avent, A\. Bellet, M\. Bennis, A\. N\. Bhagoji, K\. A\. Bonawitz, Z\. Charles, G\. Cormode, R\. Cummings, R\. G\. L\. D’Oliveira, H\. Eichner, S\. El Rouayheb, D\. Evans, J\. Gardner, Z\. Garrett, A\. Gascón, B\. Ghazi, P\. B\. Gibbons, M\. Gruteser, Z\. Harchaoui, C\. He, L\. He, Z\. Huo, B\. Hutchinson, J\. Hsu, M\. Jaggi, T\. Javidi, G\. Joshi, M\. Khodak, J\. Konecný, A\. Korolova, F\. Koushanfar, S\. Koyejo, T\. Lepoint, Y\. Liu, P\. Mittal, M\. Mohri, R\. Nock, A\. Özgür, R\. Pagh, H\. Qi, D\. Ramage, R\. Raskar, M\. Raykova, D\. Song, W\. Song, S\. U\. Stich, Z\. Sun, A\. T\. Suresh, F\. Tramèr, P\. Vepakomma, J\. Wang, L\. Xiong, Z\. Xu, Q\. Yang, F\. X\. Yu, H\. Yu, and S\. Zhao\(2021\)Advances and open problems in federated learning\.Foundations and Trends in Machine Learning14\(1–2\),pp\. 1–210\.External Links:[Document](https://dx.doi.org/10.1561/2200000083),[Link](https://doi.org/10.1561/2200000083)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px1.p1.1)\.
- \[6\]S\. P\. Karimireddy, S\. Kale, M\. Mohri, S\. Reddi, S\. Stich, and A\. T\. Suresh\(2020\)SCAFFOLD: stochastic controlled averaging for federated learning\.InProceedings of the 37th International Conference on Machine Learning,H\. D\. III and A\. Singh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.119,pp\. 5132–5143\.External Links:[Link](https://proceedings.mlr.press/v119/karimireddy20a.html)Cited by:[§1](https://arxiv.org/html/2608.26765#S1.p1.1),[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2608.26765#S3.p3.1.1.2)\.
- \[7\]A\. Khaled, K\. Mishchenko, and P\. Richtarik\(2020\)Tighter theory for Local SGD on identical and heterogeneous data\.InProceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics,S\. Chiappa and R\. Calandra \(Eds\.\),Proceedings of Machine Learning Research, Vol\.108,pp\. 4519–4529\.External Links:[Link](https://proceedings.mlr.press/v108/bayoumi20a.html)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px1.p1.1)\.
- \[8\]R\. Luo, S\. U\. Stich, S\. Horváth, and M\. Takáč\(2025\)Revisiting LocalSGD and SCAFFOLD: improved rates and missing analysis\.InProceedings of The 28th International Conference on Artificial Intelligence and Statistics,Y\. Li, S\. Mandt, S\. Agrawal, and E\. Khan \(Eds\.\),Proceedings of Machine Learning Research, Vol\.258,pp\. 2539–2547\.External Links:[Link](https://proceedings.mlr.press/v258/luo25c.html)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px2.p1.1)\.
- \[9\]S\. Mandt, M\. D\. Hoffman, and D\. M\. Blei\(2017\)Stochastic gradient descent as approximate Bayesian inference\.Journal of Machine Learning Research18\(134\),pp\. 1–35\.External Links:[Link](https://jmlr.org/papers/v18/17-214.html)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px3.p1.1)\.
- \[10\]P\. Mangold, A\. O\. Durmus, A\. Dieuleveut, S\. Samsonov, and E\. Moulines\(2025\)Refined analysis of constant step size Federated Averaging and Federated Richardson–Romberg Extrapolation\.InProceedings of The 28th International Conference on Artificial Intelligence and Statistics,Y\. Li, S\. Mandt, S\. Agrawal, and E\. Khan \(Eds\.\),Proceedings of Machine Learning Research, Vol\.258,pp\. 5023–5031\.External Links:[Link](https://proceedings.mlr.press/v258/mangold25a.html)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px3.p1.1)\.
- \[11\]P\. Mangold and E\. Moulines\(2025\)A sharper analysis of SCAFFOLD on quadratics\.In2025 3rd International Conference on Federated Learning Technologies and Applications, FLTA 2025,F\. M\. Awaysheh and S\. Alawadi \(Eds\.\),pp\. 332–339\.External Links:[Document](https://dx.doi.org/10.1109/FLTA67013.2025.11336626),[Link](https://ieeexplore.ieee.org/document/11336626)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px5.p1.1)\.
- \[12\]P\. Mangold, A\. Oliviero Durmus, A\. Dieuleveut, and E\. Moulines\(2025\)SCAFFOLD with stochastic gradients: new analysis with linear speed\-up\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 42902–42946\.External Links:[Link](https://proceedings.mlr.press/v267/mangold25a.html)Cited by:[item \(L1\)](https://arxiv.org/html/2608.26765#A1.I1.i1.p1.1),[item \(L2\)](https://arxiv.org/html/2608.26765#A1.I1.i2.p1.1),[item \(L3\)](https://arxiv.org/html/2608.26765#A1.I1.i3.p1.1),[item \(L4\)](https://arxiv.org/html/2608.26765#A1.I1.i4.p1.1),[§A\.3](https://arxiv.org/html/2608.26765#A1.SS3.p1.1),[Remark A\.1](https://arxiv.org/html/2608.26765#A1.Thmtheorem1.p1.3),[Proposition A\.3](https://arxiv.org/html/2608.26765#A1.Thmtheorem3.p1.1.1),[Appendix B](https://arxiv.org/html/2608.26765#A2.p1.1),[§C\.1](https://arxiv.org/html/2608.26765#A3.SS1.p1.4.1),[§C\.2](https://arxiv.org/html/2608.26765#A3.SS2.p1.1.1),[Remark D\.3](https://arxiv.org/html/2608.26765#A4.Thmtheorem3.p1.1),[§I\.1](https://arxiv.org/html/2608.26765#A9.SS1.p1.1.1),[§1](https://arxiv.org/html/2608.26765#S1.p2.1),[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px5.p1.1)\.
- \[13\]B\. McMahan, E\. Moore, D\. Ramage, S\. Hampson, and B\. A\. y\. Arcas\(2017\)Communication\-efficient learning of deep networks from decentralized data\.InProceedings of the 20th International Conference on Artificial Intelligence and Statistics,A\. Singh and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.54,pp\. 1273–1282\.External Links:[Link](https://proceedings.mlr.press/v54/mcmahan17a.html)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px1.p1.1)\.
- \[14\]S\. U\. Stich\(2019\)Local SGD converges fast and communicates little\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=S1g2JnRcFX)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px1.p1.1)\.
- \[15\]L\. Versini, P\. Mangold, and A\. Dieuleveut\(2026\)Tight analysis of decentralized SGD: a Markov chain perspective\.InProceedings of the 29th International Conference on Artificial Intelligence and Statistics,Vol\.300\.External Links:[Link](https://openreview.net/forum?id=5ob5u8lZeL)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px3.p1.1)\.
- \[16\]B\. Woodworth, K\. K\. Patel, S\. Stich, Z\. Dai, B\. Bullins, B\. Mcmahan, O\. Shamir, and N\. Srebro\(2020\)Is Local SGD better than Minibatch SGD?\.InProceedings of the 37th International Conference on Machine Learning,H\. D\. III and A\. Singh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.119,pp\. 10334–10343\.External Links:[Link](https://proceedings.mlr.press/v119/woodworth20a.html)Cited by:[§2](https://arxiv.org/html/2608.26765#S2.SS0.SSS0.Px1.p1.1)\.

## Appendix AScope, notation, and imported results

### A\.1Theorem scope

We translate the optimum to the origin and work only in the following restricted setting\.

###### Assumption 1\(One\-dimensional homogeneous fixed\-HHsetting\)\.

Letd=1d=1, letN≥2N\\geq 2, and fix an integerH≥2H\\geq 2\. AllNNclients participate in every communication round and have the same objectivefc=ff\_\{c\}=f\. The unique minimizer is translated tox⋆=0x^\{\\star\}=0, sof′​\(0\)=0f^\{\\prime\}\(0\)=0\. The integerHHis fixed: no bound in this document is claimed to be uniform asH→∞H\\to\\infty\.

###### Assumption 2\(Strong convexity, smoothness, and higher regularity\)\.

There exist0<μ≤L<∞0<\\mu\\leq L<\\inftysuch that

μ≤f′′​\(y\)≤L,y∈ℝ\.\\mu\\leq f^\{\\prime\\prime\}\(y\)\\leq L,\\qquad y\\in\\mathbb\{R\}\.Moreover,f∈C5​\(ℝ\)f\\in C^\{5\}\(\\mathbb\{R\}\)and, forj=3,4,5j=3,4,5, there are finite constantsKjK\_\{j\}such that

supy∈ℝ\|f\(j\)​\(y\)\|≤Kj\.\\sup\_\{y\\in\\mathbb\{R\}\}\|f^\{\(j\)\}\(y\)\|\\leq K\_\{j\}\.Write

a:=f′′​\(0\)\>0,τ:=f′′′​\(0\),κ:=f′′′′​\(0\)\.a:=f^\{\\prime\\prime\}\(0\)\>0,\\qquad\\tau:=f^\{\\prime\\prime\\prime\}\(0\),\\qquad\\kappa:=f^\{\\prime\\prime\\prime\\prime\}\(0\)\.

###### Assumption 3\(Bounded additive fresh gradient noise\)\.

At local stephhof clientcc, the stochastic gradient is

∇Fc,h​\(y\)=f′​\(y\)\+εc,h\.\\nabla F\_\{c,h\}\(y\)=f^\{\\prime\}\(y\)\+\\varepsilon\_\{c,h\}\.The variables\{εc,h\}\\\{\\varepsilon\_\{c,h\}\\\}are independent across clients, local steps, and communication rounds, independent of the round\-start state, and satisfy

𝔼\[εc,h\]=0,𝔼\[εc,h2\]=σ2,\|εc,h\|≤Bεa\.s\.\\mathbb\{E\}\[\\varepsilon\_\{c,h\}\]=0,\\qquad\\mathbb\{E\}\[\\varepsilon\_\{c,h\}^\{2\}\]=\\sigma^\{2\},\\qquad\|\\varepsilon\_\{c,h\}\|\\leq B\_\{\\varepsilon\}\\quad\\text\{a\.s\.\}No symmetry assumption is imposed\. In particular,𝔼⁡\[ε3\]\\mathbb\{E\}\[\\varepsilon^\{3\}\]may be nonzero\.

###### Assumption 4\(Zero\-sum control state\)\.

The SCAFFOLD state is restricted to the invariant subspace

𝒳0:=\{\(x,ξ1,…,ξN\)∈ℝN\+1:∑c=1Nξc=0\}\.\\mathcal\{X\}\_\{0\}:=\\left\\\{\(x,\\xi\_\{1\},\\ldots,\\xi\_\{N\}\)\\in\\mathbb\{R\}^\{N\+1\}:\\sum\_\{c=1\}^\{N\}\\xi\_\{c\}=0\\right\\\}\.Equivalently, the initial controls satisfy∑cξc0=0\\sum\_\{c\}\\xi\_\{c\}^\{0\}=0\.

### A\.2Stationary convention and scale notation

Throughout the proof,\(x,Ξ\)=\(x,ξ1,…,ξN\)\(x,\\Xi\)=\(x,\\xi\_\{1\},\\ldots,\\xi\_\{N\}\)denotes a round\-start state distributed according to the stationary lawπγ,N,H\\pi\_\{\\gamma,N,H\}, and all expectations also include the fresh noises drawn during the next communication round\. The communication\-round index is suppressed\.

Table 1:Core notation used in the main text and appendices\.Define the three scales

sγ,N:=γN\+γ2,rγ,N:=γ2N\+γ3=γ​sγ,N,vγ,N:=γ2N2\+γ3\.s\_\{\\gamma,N\}:=\\frac\{\\gamma\}\{N\}\+\\gamma^\{2\},\\qquad r\_\{\\gamma,N\}:=\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}=\\gamma s\_\{\\gamma,N\},\\qquad v\_\{\\gamma,N\}:=\\frac\{\\gamma^\{2\}\}\{N^\{2\}\}\+\\gamma^\{3\}\.\(21\)The notationOH​\(⋅\)O\_\{H\}\(\\cdot\)means that the hidden constant may depend on fixedHHand the fixed problem/noise parameters, but not onNNorγ\\gamma\.

### A\.3Earlier stationary results used in the analysis

The following inputs are taken from[Mangold et al\. \[12\]](https://arxiv.org/html/2608.26765#bib.bib2); they are not new claims of this document\.

Table 2:Imported\-result interface\. The middle column records how the assumptions underlying the earlier results are satisfied in the present setting; the last column states how each input is used here\.###### Proposition A\.3\(Earlier stationary inputs\)\.

Under the matched assumptions summarized in[Table2](https://arxiv.org/html/2608.26765#A1.T2)and sufficiently small step size,[Mangold et al\. \[12\]](https://arxiv.org/html/2608.26765#bib.bib2)establish the following facts\.

1. \(L1\)Stationarity\.The joint global/control process is a time\-homogeneous Markov chain with a unique stationary distribution and geometric convergence inW2W\_\{2\}; see[Mangold et al\. \[12, Theorem 4\.2\]](https://arxiv.org/html/2608.26765#bib.bib2)\.
2. \(L2\)Coarse moments\.The stationary global/local iterates have coarse second moments of orderOH​\(γ\)O\_\{H\}\(\\gamma\), the control second moments areOH​\(1\)O\_\{H\}\(1\), and Appendix B of[Mangold et al\. \[12\]](https://arxiv.org/html/2608.26765#bib.bib2)provides coarse fourth/sixth\-moment estimates for the global/local iterates\. We use those iterate moment bounds, but we do*not*rely on an imported control sixth\-moment bound\.
3. \(L3\)Control second moment\.In the one\-dimensional homogeneous additive\-noise specialization of[Mangold et al\. \[12, Lemma 5\.1 and Appendix E\.2\]](https://arxiv.org/html/2608.26765#bib.bib2), Qγ,N,H:=1N​∑c=1N𝔼⁡\[ξc2\]=σ2H​\(1−1N\)\+OH​\(γ\)\.Q\_\{\\gamma,N,H\}:=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\[\\xi\_\{c\}^\{2\}\]=\\frac\{\\sigma^\{2\}\}\{H\}\\left\(1\-\\frac\{1\}\{N\}\\right\)\+O\_\{H\}\(\\gamma\)\.\(22\)In this homogeneous setting, the ideal controls satisfyξc⋆=−f′​\(0\)=0\\xi\_\{c\}^\{\\star\}=\-f^\{\\prime\}\(0\)=0, so the squared deviation from the ideal control is the raw second moment displayed above\.
4. \(L4\)Known first bias layer\.[Mangold et al\. \[12, Theorem 5\.3\]](https://arxiv.org/html/2608.26765#bib.bib2)prove a leading stationary bias of orderγ/N\\gamma/Nand leaves a higher\-order remainderO⁡\(γ2​H\+γ3/2\)O\(\\gamma^\{2\}H\+\\gamma^\{3/2\}\)\. In the present scalar homogeneous specialization, the leading coefficient is −τ​σ24​a2​γN\.\-\\frac\{\\tau\\sigma^\{2\}\}\{4a^\{2\}\}\\frac\{\\gamma\}\{N\}\.\(23\)

## Appendix BExact SCAFFOLD identities

All identities in this section are pathwise and precede any small\-step expansion\. The one\-round update below is the stochasticSCAFFOLDrecursion specialized to[Assumptions1](https://arxiv.org/html/2608.26765#Thmassumption1)and[3](https://arxiv.org/html/2608.26765#Thmassumption3); it serves as the complete algorithmic definition used in the analysis\[[12](https://arxiv.org/html/2608.26765#bib.bib2)\]\.

###### Lemma B\.1\(Exact local/global/control recursions\)\.

For one communication round,

θc,h\+1\\displaystyle\\theta\_\{c,h\+1\}=θc,h−γ⁡\{f′​\(θc,h\)\+ξc\+εc,h\+1\},θc,0=x,\\displaystyle=\\theta\_\{c,h\}\-\\gamma\\\{f^\{\\prime\}\(\\theta\_\{c,h\}\)\+\\xi\_\{c\}\+\\varepsilon\_\{c,h\+1\}\\\},\\qquad\\theta\_\{c,0\}=x,\(24\)x\+\\displaystyle x^\{\+\}=1N​∑c=1Nθc,H,\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\theta\_\{c,H\},\(25\)ξc\+\\displaystyle\\xi\_\{c\}^\{\+\}=ξc\+1γ​H​\(θc,H−x\+\)\.\\displaystyle=\\xi\_\{c\}\+\\frac\{1\}\{\\gamma H\}\(\\theta\_\{c,H\}\-x^\{\+\}\)\.\(26\)Moreover,

∑c=1Nξc\+=∑c=1Nξc\.\\sum\_\{c=1\}^\{N\}\\xi\_\{c\}^\{\+\}=\\sum\_\{c=1\}^\{N\}\\xi\_\{c\}\.\(27\)Hence[Assumption4](https://arxiv.org/html/2608.26765#Thmassumption4)is invariant under the dynamics\.

###### Proof\.

Equations \([24](https://arxiv.org/html/2608.26765#A2.E24)\)–\([26](https://arxiv.org/html/2608.26765#A2.E26)\) are the defining one\-round updates in the present notation\. Summing Equation \([26](https://arxiv.org/html/2608.26765#A2.E26)\) over clients and using Equation \([25](https://arxiv.org/html/2608.26765#A2.E25)\),

∑cξc\+=∑cξc\+1γ​H​\(∑cθc,H−N​x\+\)=∑cξc\.\\sum\_\{c\}\\xi\_\{c\}^\{\+\}=\\sum\_\{c\}\\xi\_\{c\}\+\\frac\{1\}\{\\gamma H\}\\left\(\\sum\_\{c\}\\theta\_\{c,H\}\-Nx^\{\+\}\\right\)=\\sum\_\{c\}\\xi\_\{c\}\.∎

###### Corollary B\.2\(Exact global update and linear control cancellation\)\.

Under[Assumption4](https://arxiv.org/html/2608.26765#Thmassumption4),

x\+−x=−γ∑h=0H−1f′​\(θh\)¯−γ∑h=0H−1ε¯h\+1,x^\{\+\}\-x=\-\\gamma\\sum\_\{h=0\}^\{H\-1\}\\overline\{f^\{\\prime\}\(\\theta\_\{h\}\)\}\-\\gamma\\sum\_\{h=0\}^\{H\-1\}\\bar\{\\varepsilon\}\_\{h\+1\},\(28\)where

f′​\(θh\)¯:=1N​∑c=1Nf′​\(θc,h\),ε¯h:=1N​∑c=1Nεc,h\.\\overline\{f^\{\\prime\}\(\\theta\_\{h\}\)\}:=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}f^\{\\prime\}\(\\theta\_\{c,h\}\),\\qquad\\bar\{\\varepsilon\}\_\{h\}:=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\varepsilon\_\{c,h\}\.The cancellation of the linear control term is pathwise and uses no independence assumption\.

###### Proof\.

Unrolling Equation \([24](https://arxiv.org/html/2608.26765#A2.E24)\) gives

θc,H−x=−γ∑h=0H−1f′\(θc,h\)−γHξc−γ∑h=0H−1εc,h\+1\.\\theta\_\{c,H\}\-x=\-\\gamma\\sum\_\{h=0\}^\{H\-1\}f^\{\\prime\}\(\\theta\_\{c,h\}\)\-\\gamma H\\xi\_\{c\}\-\\gamma\\sum\_\{h=0\}^\{H\-1\}\\varepsilon\_\{c,h\+1\}\.Average overccand useN−1​∑cξc=0N^\{\-1\}\\sum\_\{c\}\\xi\_\{c\}=0\. ∎

###### Corollary B\.3\(Exact stationary mean balance\)\.

At stationarity,

0=∑h=0H−11N​∑c=1N𝔼⁡\[f′​\(θc,h\)\]\.0=\\sum\_\{h=0\}^\{H\-1\}\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\[f^\{\\prime\}\(\\theta\_\{c,h\}\)\]\.\(29\)

###### Proof\.

Take the expectation in Equation \([28](https://arxiv.org/html/2608.26765#A2.E28)\)\. Stationarity gives𝔼⁡\[x\+\]=𝔼⁡\[x\]\\mathbb\{E\}\[x^\{\+\}\]=\\mathbb\{E\}\[x\], while fresh mean\-zero noise gives𝔼⁡\[ε¯h\]=0\\mathbb\{E\}\[\\bar\{\\varepsilon\}\_\{h\}\]=0\. ∎

###### Lemma B\.4\(Exact new\-control representation\)\.

The new control admits the exact representation

ξc\+=\\displaystyle\\xi\_\{c\}^\{\+\}=\{\}−1H∑h=0H−1\(f′\(θc,h\)−f′​\(θh\)¯\)\\displaystyle\-\\frac\{1\}\{H\}\\sum\_\{h=0\}^\{H\-1\}\\left\(f^\{\\prime\}\(\\theta\_\{c,h\}\)\-\\overline\{f^\{\\prime\}\(\\theta\_\{h\}\)\}\\right\)\(30\)−1H∑h=0H−1\(εc,h\+1−ε¯h\+1\)\.\\displaystyle\-\\frac\{1\}\{H\}\\sum\_\{h=0\}^\{H\-1\}\\left\(\\varepsilon\_\{c,h\+1\}\-\\bar\{\\varepsilon\}\_\{h\+1\}\\right\)\.In particular, the old controlξc\\xi\_\{c\}has no direct term on the right\-hand side\.

###### Proof\.

Subtract Equation \([28](https://arxiv.org/html/2608.26765#A2.E28)\) from the unrolled local endpoint, substitute the result into Equation \([26](https://arxiv.org/html/2608.26765#A2.E26)\), and cancel the termξc−ξc\\xi\_\{c\}\-\\xi\_\{c\}\. ∎

## Appendix CControl moments and coarse global second moment

### C\.1Uniform control moments

###### Lemma C\.1\(Uniform sixth control moment\)\.

There existγ0,H\(C​6\)\>0\\gamma\_\{0,H\}^\{\(C6\)\}\>0andCξ,6,H<∞C\_\{\\xi,6,H\}<\\infty, independent ofNNandγ\\gamma, such that for0<γ≤γ0,H\(C​6\)0<\\gamma\\leq\\gamma\_\{0,H\}^\{\(C6\)\},

1N​∑c=1N𝔼​\|ξc\|6≤Cξ,6,H\.\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\|\\xi\_\{c\}\|^\{6\}\\leq C\_\{\\xi,6,H\}\.\(31\)Consequently, the client\-averaged second and fourth control moments are uniformly bounded as well\.

###### Proof\.

Write Equation \([30](https://arxiv.org/html/2608.26765#A2.E30)\) asξc\+=−Ac−Bc\\xi\_\{c\}^\{\+\}=\-A\_\{c\}\-B\_\{c\}, where

Ac:=1H​∑h=0H−1\(f′​\(θc,h\)−f′​\(θh\)¯\),Bc:=1H​∑h=0H−1\(εc,h\+1−ε¯h\+1\)\.A\_\{c\}:=\\frac\{1\}\{H\}\\sum\_\{h=0\}^\{H\-1\}\\left\(f^\{\\prime\}\(\\theta\_\{c,h\}\)\-\\overline\{f^\{\\prime\}\(\\theta\_\{h\}\)\}\\right\),\\qquad B\_\{c\}:=\\frac\{1\}\{H\}\\sum\_\{h=0\}^\{H\-1\}\(\\varepsilon\_\{c,h\+1\}\-\\bar\{\\varepsilon\}\_\{h\+1\}\)\.By convexity ofz↦\|z\|6z\\mapsto\|z\|^\{6\}and\|u−v\|6≤32​\(\|u\|6\+\|v\|6\)\|u\-v\|^\{6\}\\leq 32\(\|u\|^\{6\}\+\|v\|^\{6\}\),

1N​∑c𝔼​\|Ac\|6≤64​1H​∑h=0H−11N​∑c𝔼​\|f′​\(θc,h\)\|6\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|A\_\{c\}\|^\{6\}\\leq 64\\frac\{1\}\{H\}\\sum\_\{h=0\}^\{H\-1\}\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|f^\{\\prime\}\(\\theta\_\{c,h\}\)\|^\{6\}\.Sincef′​\(0\)=0f^\{\\prime\}\(0\)=0andf′f^\{\\prime\}isLL\-Lipschitz,

\|f′​\(y\)\|6≤L6​\|y\|6\.\|f^\{\\prime\}\(y\)\|^\{6\}\\leq L^\{6\}\|y\|^\{6\}\.The coarse local sixth\-moment estimate of[Mangold et al\. \[12\]](https://arxiv.org/html/2608.26765#bib.bib2), with the sixth\-moment noise envelope supplied by boundedness under[Assumption3](https://arxiv.org/html/2608.26765#Thmassumption3), gives

1N​∑c𝔼​\|θc,h\|6≤CH​γ3\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|\\theta\_\{c,h\}\|^\{6\}\\leq C\_\{H\}\\gamma^\{3\}\.HenceN−1​∑c𝔼​\|Ac\|6≤CH​γ3N^\{\-1\}\\sum\_\{c\}\\mathbb\{E\}\|A\_\{c\}\|^\{6\}\\leq C\_\{H\}\\gamma^\{3\}\. On the other hand,

\|εc,h−ε¯h\|≤2​Bε\|\\varepsilon\_\{c,h\}\-\\bar\{\\varepsilon\}\_\{h\}\|\\leq 2B\_\{\\varepsilon\}pathwise, soN−1​∑c𝔼​\|Bc\|6≤\(2​Bε\)6N^\{\-1\}\\sum\_\{c\}\\mathbb\{E\}\|B\_\{c\}\|^\{6\}\\leq\(2B\_\{\\varepsilon\}\)^\{6\}\. Therefore

1N​∑c𝔼​\|ξc\+\|6≤CH\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|\\xi\_\{c\}^\{\+\}\|^\{6\}\\leq C\_\{H\}forγ≤1\\gamma\\leq 1\. Stationarity transfers this bound fromξc\+\\xi\_\{c\}^\{\+\}toξc\\xi\_\{c\}\. The second and fourth moment bounds follow from Hölder/Jensen on the joint probability–client average\. ∎

### C\.2Restricted uniformity of the control second\-moment expansion

###### Lemma C\.3\(Uniform restricted control second moment\)\.

In the present homogeneous additive\-noise setting,

Qγ,N,H=σ2H​\(1−1N\)\+RQ,H,\|RQ,H\|≤CQ,H​γ,Q\_\{\\gamma,N,H\}=\\frac\{\\sigma^\{2\}\}\{H\}\\left\(1\-\\frac\{1\}\{N\}\\right\)\+R\_\{Q,H\},\\qquad\|R\_\{Q,H\}\|\\leq C\_\{Q,H\}\\gamma,\(32\)whereCQ,HC\_\{Q,H\}is independent ofNNandγ\\gamma\.

###### Proof\.

This is consistent with the specialization of[Mangold et al\. \[12, Lemma 5\.1\]](https://arxiv.org/html/2608.26765#bib.bib2); we record a restricted argument to make the uniformity explicit\. From Equation \([30](https://arxiv.org/html/2608.26765#A2.E30)\), write

ξc\+=−Gc−Ec,\\xi\_\{c\}^\{\+\}=\-G\_\{c\}\-E\_\{c\},withGcG\_\{c\}the averaged gradient\-disagreement term andEcE\_\{c\}the averaged fresh\-noise disagreement\. For each local step,

Var⁡\(εc,h−ε¯h\)=σ2​\(1−1N\),\\operatorname\{Var\}\(\\varepsilon\_\{c,h\}\-\\bar\{\\varepsilon\}\_\{h\}\)=\\sigma^\{2\}\\left\(1\-\\frac\{1\}\{N\}\\right\),and different local\-step noises are independent\. Thus

𝔼⁡\[Ec2\]=σ2H​\(1−1N\)\.\\mathbb\{E\}\[E\_\{c\}^\{2\}\]=\\frac\{\\sigma^\{2\}\}\{H\}\\left\(1\-\\frac\{1\}\{N\}\\right\)\.\(33\)We first establish the finite\-step estimate needed here without appealing forward to[LemmaC\.4](https://arxiv.org/html/2608.26765#A3.Thmtheorem4)\. Letδc,h:=θc,h−x\\delta\_\{c,h\}:=\\theta\_\{c,h\}\-x\. From Equation \([24](https://arxiv.org/html/2608.26765#A2.E24)\),

δc,h\+1=δc,h−γ⁡\{f′​\(x\+δc,h\)\+ξc\+εc,h\+1\},δc,0=0\.\\delta\_\{c,h\+1\}=\\delta\_\{c,h\}\-\\gamma\\\{f^\{\\prime\}\(x\+\\delta\_\{c,h\}\)\+\\xi\_\{c\}\+\\varepsilon\_\{c,h\+1\}\\\},\\qquad\\delta\_\{c,0\}=0\.Sincef′​\(0\)=0f^\{\\prime\}\(0\)=0andf′f^\{\\prime\}isLL\-Lipschitz, a finite\-step discrete Grönwall bound gives, for everyh≤Hh\\leq H,

\|δc,h\|≤CH​γ​\(\|x\|\+\|ξc\|\+∑j=1h\|εc,j\|\)\.\|\\delta\_\{c,h\}\|\\leq C\_\{H\}\\gamma\\left\(\|x\|\+\|\\xi\_\{c\}\|\+\\sum\_\{j=1\}^\{h\}\|\\varepsilon\_\{c,j\}\|\\right\)\.Averaging the square over clients and expectations, and using the earlier coarse bound𝔼⁡\[x2\]=OH​\(γ\)\\mathbb\{E\}\[x^\{2\}\]=O\_\{H\}\(\\gamma\), the coarse client\-averaged control second moment, and bounded noise, yields

1N​∑c𝔼​\|θc,h−x\|2=1N​∑c𝔼​\|δc,h\|2≤CH​γ2\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|\\theta\_\{c,h\}\-x\|^\{2\}=\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|\\delta\_\{c,h\}\|^\{2\}\\leq C\_\{H\}\\gamma^\{2\}\.Thus, the estimate used in this lemma has already been proved at this point;[LemmaC\.4](https://arxiv.org/html/2608.26765#A3.Thmtheorem4)records it separately for later reuse\. Using smoothness and Jensen givesN−1​∑c𝔼⁡\[Gc2\]≤CH​γ2N^\{\-1\}\\sum\_\{c\}\\mathbb\{E\}\[G\_\{c\}^\{2\}\]\\leq C\_\{H\}\\gamma^\{2\}\. Hence

\|1N​∑c𝔼⁡\[Gc​Ec\]\|≤CH​γ,\\left\|\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[G\_\{c\}E\_\{c\}\]\\right\|\\leq C\_\{H\}\\gamma,whileN−1​∑c𝔼⁡\[Gc2\]=OH​\(γ2\)N^\{\-1\}\\sum\_\{c\}\\mathbb\{E\}\[G\_\{c\}^\{2\}\]=O\_\{H\}\(\\gamma^\{2\}\)\. Expanding𝔼⁡\[\(ξc\+\)2\]\\mathbb\{E\}\[\(\\xi\_\{c\}^\{\+\}\)^\{2\}\], averaging clients, and using stationarity yields Equation \([32](https://arxiv.org/html/2608.26765#A3.E32)\)\. All constants are uniform inNNbecause0≤1−1/N≤10\\leq 1\-1/N\\leq 1and the preceding average\-moment estimates are uniform\. ∎

### C\.3Coarse global second moment

###### Lemma C\.4\(Finite\-step local displacement\)\.

For everyh∈\{0,…,H\}h\\in\\\{0,\\ldots,H\\\},

1N​∑c=1N𝔼⁡\[\(θc,h−x\)2\]≤CH​γ2\.\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\[\(\\theta\_\{c,h\}\-x\)^\{2\}\]\\leq C\_\{H\}\\gamma^\{2\}\.\(34\)

###### Proof\.

Letδc,h:=θc,h−x\\delta\_\{c,h\}:=\\theta\_\{c,h\}\-x\. From Equation \([24](https://arxiv.org/html/2608.26765#A2.E24)\),

δc,h\+1=δc,h−γ⁡\{f′​\(x\+δc,h\)\+ξc\+εc,h\+1\},δc,0=0\.\\delta\_\{c,h\+1\}=\\delta\_\{c,h\}\-\\gamma\\\{f^\{\\prime\}\(x\+\\delta\_\{c,h\}\)\+\\xi\_\{c\}\+\\varepsilon\_\{c,h\+1\}\\\},\\qquad\\delta\_\{c,0\}=0\.Since\|f′​\(y\)\|≤L​\|y\|\|f^\{\\prime\}\(y\)\|\\leq L\|y\|, finite\-step discrete Grönwall gives, forh≤Hh\\leq H,

\|δc,h\|≤CH​γ​\(\|x\|\+\|ξc\|\+∑j=1h\|εc,j\|\)\.\|\\delta\_\{c,h\}\|\\leq C\_\{H\}\\gamma\\left\(\|x\|\+\|\\xi\_\{c\}\|\+\\sum\_\{j=1\}^\{h\}\|\\varepsilon\_\{c,j\}\|\\right\)\.Use the earlier coarse bound𝔼⁡\[x2\]=OH​\(γ\)\\mathbb\{E\}\[x^\{2\}\]=O\_\{H\}\(\\gamma\), the coarse control second moment, and bounded noise\. No sharp coefficient from[LemmaC\.3](https://arxiv.org/html/2608.26765#A3.Thmtheorem3)is needed\. ∎

###### Lemma C\.5\(Coarse global second moment\)\.

There are constants independent ofN,γN,\\gammasuch that

M2:=𝔼⁡\[x2\]≤CH​sγ,N=CH​\(γN\+γ2\)\.M\_\{2\}:=\\mathbb\{E\}\[x^\{2\}\]\\leq C\_\{H\}s\_\{\\gamma,N\}=C\_\{H\}\\left\(\\frac\{\\gamma\}\{N\}\+\\gamma^\{2\}\\right\)\.\(35\)

###### Proof\.

From[CorollaryB\.2](https://arxiv.org/html/2608.26765#A2.Thmtheorem2), define the exact decomposition

x\+=Φγ​\(x\)\+ζ\+𝒯,x^\{\+\}=\\Phi\_\{\\gamma\}\(x\)\+\\zeta\+\\mathcal\{T\},\(36\)where

Φγ\(x\):=x−γHf′\(x\),ζ:=−γ∑h=0H−1ε¯h\+1,\\Phi\_\{\\gamma\}\(x\):=x\-\\gamma Hf^\{\\prime\}\(x\),\\qquad\\zeta:=\-\\gamma\\sum\_\{h=0\}^\{H\-1\}\\bar\{\\varepsilon\}\_\{h\+1\},and

𝒯:=−γ∑h=0H−1\(f′​\(θh\)¯−f′\(x\)\)\.\\mathcal\{T\}:=\-\\gamma\\sum\_\{h=0\}^\{H\-1\}\\left\(\\overline\{f^\{\\prime\}\(\\theta\_\{h\}\)\}\-f^\{\\prime\}\(x\)\\right\)\.We call𝒯\\mathcal\{T\}the*local\-trajectory deviation term*; it is not purely a nonlinear remainder and need not vanish for a quadratic objective\.

Client and local\-step independence give

𝔼⁡\[ζ2\]=γ2​H​σ2N,𝔼⁡\[Φγ​\(x\)​ζ\]=0\.\\mathbb\{E\}\[\\zeta^\{2\}\]=\\frac\{\\gamma^\{2\}H\\sigma^\{2\}\}\{N\},\\qquad\\mathbb\{E\}\[\\Phi\_\{\\gamma\}\(x\)\\zeta\]=0\.\(37\)By smoothness, Jensen, and[LemmaC\.4](https://arxiv.org/html/2608.26765#A3.Thmtheorem4),

𝔼⁡\[𝒯2\]≤CH​γ4\.\\mathbb\{E\}\[\\mathcal\{T\}^\{2\}\]\\leq C\_\{H\}\\gamma^\{4\}\.\(38\)Strong convexity andγ​H​L≤1\\gamma HL\\leq 1imply

Φγ​\(x\)2≤\(1−μ​H​γ\)​x2\.\\Phi\_\{\\gamma\}\(x\)^\{2\}\\leq\(1\-\\mu H\\gamma\)x^\{2\}\.Stationarity in Equation \([36](https://arxiv.org/html/2608.26765#A3.E36)\), Young’s inequality for theΦγ​𝒯\\Phi\_\{\\gamma\}\\mathcal\{T\}term, and Cauchy–Schwarz forζ​𝒯\\zeta\\mathcal\{T\}give

μ​H​γ2​M2≤CH​\(γ2N\+γ3\+γ4\)\.\\frac\{\\mu H\\gamma\}\{2\}M\_\{2\}\\leq C\_\{H\}\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\+\\gamma^\{4\}\\right\)\.Here

γ3N≤12​\(γ2N\+γ4\)\\frac\{\\gamma^\{3\}\}\{\\sqrt\{N\}\}\\leq\\frac\{1\}\{2\}\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{4\}\\right\)controls the mixed noise/trajectory term without any relative\-rate condition\. Divide byγ\\gammaand takeγ≤1\\gamma\\leq 1\. ∎

## Appendix DLocal second\-moment expansion

This section refines an earlier coarse local second\-moment bound to the coefficient level\. The direct fresh\-noise term and the persistent control second moment enter through an exact square\.

###### Lemma D\.1\(Local second\-moment expansion\)\.

Forh∈\{0,…,H\}h\\in\\\{0,\\ldots,H\\\}define

U¯h:=1N​∑c=1N\(𝔼⁡\[θc,h2\]−M2\)\.\\bar\{U\}\_\{h\}:=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\left\(\\mathbb\{E\}\[\\theta\_\{c,h\}^\{2\}\]\-M\_\{2\}\\right\)\.Then

U¯h=γ2​\(h2​Qγ,N,H\+h​σ2\)\+OH​\(rγ,N\),\\bar\{U\}\_\{h\}=\\gamma^\{2\}\\left\(h^\{2\}Q\_\{\\gamma,N,H\}\+h\\sigma^\{2\}\\right\)\+O\_\{H\}\(r\_\{\\gamma,N\}\),\(39\)and, using[LemmaC\.3](https://arxiv.org/html/2608.26765#A3.Thmtheorem3),

U¯h=γ2​σ2​\(h\+h2H\)\+OH​\(γ2N\+γ3\)\.\\bar\{U\}\_\{h\}=\\gamma^\{2\}\\sigma^\{2\}\\left\(h\+\\frac\{h^\{2\}\}\{H\}\\right\)\+O\_\{H\}\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\)\.\(40\)The remainder is uniform inNNandγ\\gammafor fixedHH\.

###### Proof\.

Unroll Equation \([24](https://arxiv.org/html/2608.26765#A2.E24)\) forhhsteps and write the exact decomposition

θc,h=x\+Dc,h\+Rc,h,\\theta\_\{c,h\}=x\+D\_\{c,h\}\+R\_\{c,h\},\(41\)where

Dc,h:=−γ\(hξc\+∑j=1hεc,j\),Rc,h:=−γ∑ℓ=0h−1f′\(θc,ℓ\)\.D\_\{c,h\}:=\-\\gamma\\left\(h\\xi\_\{c\}\+\\sum\_\{j=1\}^\{h\}\\varepsilon\_\{c,j\}\\right\),\\qquad R\_\{c,h\}:=\-\\gamma\\sum\_\{\\ell=0\}^\{h\-1\}f^\{\\prime\}\(\\theta\_\{c,\\ell\}\)\.Expanding the square gives

U¯h=\\displaystyle\\bar\{U\}\_\{h\}=\{\}2​1N​∑c𝔼⁡\[x​Dc,h\]\+1N​∑c𝔼⁡\[Dc,h2\]\\displaystyle 2\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[xD\_\{c,h\}\]\+\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[D\_\{c,h\}^\{2\}\]\(42\)\+21N∑c𝔼\[xRc,h\]\+21N∑c𝔼\[Dc,hRc,h\]\+1N∑c𝔼\[Rc,h2\]\.\\displaystyle\+2\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[xR\_\{c,h\}\]\+2\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[D\_\{c,h\}R\_\{c,h\}\]\+\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[R\_\{c,h\}^\{2\}\]\.
The first term vanishes exactly\. The control part vanishes after the client average because∑cξc=0\\sum\_\{c\}\\xi\_\{c\}=0pathwise, and the fresh\-noise part vanishes because the round\-startxxis independent of the mean\-zero fresh noises\.

For the square ofDc,hD\_\{c,h\}, setSc,h:=∑j=1hεc,jS\_\{c,h\}:=\\sum\_\{j=1\}^\{h\}\\varepsilon\_\{c,j\}\. Conditional on the round\-start state,𝔼⁡\[Sc,h\]=0\\mathbb\{E\}\[S\_\{c,h\}\]=0and𝔼⁡\[Sc,h2\]=h​σ2\\mathbb\{E\}\[S\_\{c,h\}^\{2\}\]=h\\sigma^\{2\}\. Hence

𝔼⁡\[ξc​Sc,h\]=0\\mathbb\{E\}\[\\xi\_\{c\}S\_\{c,h\}\]=0and therefore the client average satisfies the exact identity

1N​∑c𝔼⁡\[Dc,h2\]=γ2​\(h2​Qγ,N,H\+h​σ2\)\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[D\_\{c,h\}^\{2\}\]=\\gamma^\{2\}\\left\(h^\{2\}Q\_\{\\gamma,N,H\}\+h\\sigma^\{2\}\\right\)\.\(43\)
It remains to bound the terms containingRc,hR\_\{c,h\}\. By\|f′​\(y\)\|≤L​\|y\|\|f^\{\\prime\}\(y\)\|\\leq L\|y\|,

Rc,h2≤γ2​h​L2​∑ℓ=0h−1θc,ℓ2\.R\_\{c,h\}^\{2\}\\leq\\gamma^\{2\}hL^\{2\}\\sum\_\{\\ell=0\}^\{h\-1\}\\theta\_\{c,\\ell\}^\{2\}\.Using[LemmasC\.5](https://arxiv.org/html/2608.26765#A3.Thmtheorem5)and[C\.4](https://arxiv.org/html/2608.26765#A3.Thmtheorem4),

1N​∑c𝔼⁡\[θc,ℓ2\]≤CH​sγ,N,\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\theta\_\{c,\\ell\}^\{2\}\]\\leq C\_\{H\}s\_\{\\gamma,N\},so

1N​∑c𝔼⁡\[Rc,h2\]≤CH​γ2​sγ,N=CH​\(γ3N\+γ4\)\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[R\_\{c,h\}^\{2\}\]\\leq C\_\{H\}\\gamma^\{2\}s\_\{\\gamma,N\}=C\_\{H\}\\left\(\\frac\{\\gamma^\{3\}\}\{N\}\+\\gamma^\{4\}\\right\)\.\(44\)Cauchy–Schwarz then gives

\|1N​∑c𝔼⁡\[x​Rc,h\]\|≤CH​γ​sγ,N=CH​rγ,N\.\\left\|\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[xR\_\{c,h\}\]\\right\|\\leq C\_\{H\}\\gamma s\_\{\\gamma,N\}=C\_\{H\}r\_\{\\gamma,N\}\.Likewise, Equation \([43](https://arxiv.org/html/2608.26765#A4.E43)\) and the coarse boundQγ,N,H=OH​\(1\)Q\_\{\\gamma,N,H\}=O\_\{H\}\(1\)give

\|1N​∑c𝔼⁡\[Dc,h​Rc,h\]\|≤CH​γ2​sγ,N\.\\left\|\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[D\_\{c,h\}R\_\{c,h\}\]\\right\|\\leq C\_\{H\}\\gamma^\{2\}\\sqrt\{s\_\{\\gamma,N\}\}\.Since

γ2​sγ,N≤γ5/2N\+γ3≤C⁡\(γ2N\+γ3\),\\gamma^\{2\}\\sqrt\{s\_\{\\gamma,N\}\}\\leq\\frac\{\\gamma^\{5/2\}\}\{\\sqrt\{N\}\}\+\\gamma^\{3\}\\leq C\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\),all remaining terms in Equation \([42](https://arxiv.org/html/2608.26765#A4.E42)\) areOH​\(rγ,N\)O\_\{H\}\(r\_\{\\gamma,N\}\)\. This proves Equation \([39](https://arxiv.org/html/2608.26765#A4.E39)\)\.

Finally, substitute Equation \([32](https://arxiv.org/html/2608.26765#A3.E32)\):

γ2​h2​Qγ,N,H=γ2​h2​σ2H\+OH​\(γ2N\+γ3\),\\gamma^\{2\}h^\{2\}Q\_\{\\gamma,N,H\}=\\gamma^\{2\}\\frac\{h^\{2\}\\sigma^\{2\}\}\{H\}\+O\_\{H\}\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\),which gives Equation \([40](https://arxiv.org/html/2608.26765#A4.E40)\)\. ∎

## Appendix EMean bootstrap and global higher moments

### E\.1Coarse mean bootstrap without signed\-moment circularity

Define

q⁡\(y\):=f′​\(y\)−a​y,a=f′′​\(0\),q\(y\):=f^\{\\prime\}\(y\)\-ay,\\qquad a=f^\{\\prime\\prime\}\(0\),and

mh:=1N​∑c=1N𝔼⁡\[θc,h\],νh:=1N​∑c=1N𝔼⁡\[q⁡\(θc,h\)\],bγ,N,H:=𝔼⁡\[x\]\.m\_\{h\}:=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\[\\theta\_\{c,h\}\],\\qquad\\nu\_\{h\}:=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\[q\(\\theta\_\{c,h\}\)\],\\qquad b\_\{\\gamma,N,H\}:=\\mathbb\{E\}\[x\]\.
###### Lemma E\.1\(Mean bootstrap\)\.

There existsCH<∞C\_\{H\}<\\inftysuch that

\|νh\|\\displaystyle\|\\nu\_\{h\}\|≤CHsγ,N,h=0,…,H,\\displaystyle\\leq C\_\{H\}s\_\{\\gamma,N\},\\qquad h=0,\\ldots,H,\(45\)\|bγ,N,H\|\\displaystyle\|b\_\{\\gamma,N,H\}\|≤CH​sγ,N,\\displaystyle\\leq C\_\{H\}s\_\{\\gamma,N\},\(46\)∑h=0H−1mh\\displaystyle\\sum\_\{h=0\}^\{H\-1\}m\_\{h\}=H​bγ,N,H\+OH​\(rγ,N\)\.\\displaystyle=Hb\_\{\\gamma,N,H\}\+O\_\{H\}\(r\_\{\\gamma,N\}\)\.\(47\)This lemma uses only the coarse second\-moment result[LemmaC\.5](https://arxiv.org/html/2608.26765#A3.Thmtheorem5)and the finite\-step displacement estimate; it does not use[LemmaD\.1](https://arxiv.org/html/2608.26765#A4.Thmtheorem1)or any signed third moment\.

###### Proof\.

Sinceq⁡\(0\)=q′​\(0\)=0q\(0\)=q^\{\\prime\}\(0\)=0andq′′=f′′′q^\{\\prime\\prime\}=f^\{\\prime\\prime\\prime\}, the integral Taylor formula and\|f′′′\|≤K3\|f^\{\\prime\\prime\\prime\}\|\\leq K\_\{3\}give the global bound

\|q⁡\(y\)\|≤K32​\|y\|2\.\|q\(y\)\|\\leq\\frac\{K\_\{3\}\}\{2\}\|y\|^\{2\}\.\(48\)By[LemmasC\.5](https://arxiv.org/html/2608.26765#A3.Thmtheorem5)and[C\.4](https://arxiv.org/html/2608.26765#A3.Thmtheorem4),

1N​∑c𝔼⁡\[θc,h2\]≤2​M2\+2​1N​∑c𝔼⁡\[\(θc,h−x\)2\]≤CH​sγ,N\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\theta\_\{c,h\}^\{2\}\]\\leq 2M\_\{2\}\+2\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\(\\theta\_\{c,h\}\-x\)^\{2\}\]\\leq C\_\{H\}s\_\{\\gamma,N\}\.Thus Equation \([45](https://arxiv.org/html/2608.26765#A5.E45)\) follows directly from Equation \([48](https://arxiv.org/html/2608.26765#A5.E48)\)\.

Average Equation \([24](https://arxiv.org/html/2608.26765#A2.E24)\) over clients and expectations\. The zero\-sum control condition and mean\-zero fresh noise yield the exact recursion

mh\+1=λγ​mh−γ​νh,λγ:=1−a​γ,m0=bγ,N,H\.m\_\{h\+1\}=\\lambda\_\{\\gamma\}m\_\{h\}\-\\gamma\\nu\_\{h\},\\qquad\\lambda\_\{\\gamma\}:=1\-a\\gamma,\\qquad m\_\{0\}=b\_\{\\gamma,N,H\}\.\(49\)Unrolling,

mh=λγh​bγ,N,H−γ​∑j=0h−1λγh−1−j​νj\.m\_\{h\}=\\lambda\_\{\\gamma\}^\{h\}b\_\{\\gamma,N,H\}\-\\gamma\\sum\_\{j=0\}^\{h\-1\}\\lambda\_\{\\gamma\}^\{h\-1\-j\}\\nu\_\{j\}\.\(50\)At stationarity,mH=𝔼⁡\[x\+\]=𝔼⁡\[x\]=bγ,N,Hm\_\{H\}=\\mathbb\{E\}\[x^\{\+\}\]=\\mathbb\{E\}\[x\]=b\_\{\\gamma,N,H\}\. Hence

bγ,N,H=−1a​AH​\(λγ\)∑j=0H−1λγH−1−jνj,AH\(λ\):=∑k=0H−1λk\.b\_\{\\gamma,N,H\}=\-\\frac\{1\}\{aA\_\{H\}\(\\lambda\_\{\\gamma\}\)\}\\sum\_\{j=0\}^\{H\-1\}\\lambda\_\{\\gamma\}^\{H\-1\-j\}\\nu\_\{j\},\\qquad A\_\{H\}\(\\lambda\):=\\sum\_\{k=0\}^\{H\-1\}\\lambda^\{k\}\.\(51\)Choose the common step\-size threshold so that0≤λγ≤10\\leq\\lambda\_\{\\gamma\}\\leq 1\. Then the weights in Equation \([51](https://arxiv.org/html/2608.26765#A5.E51)\) are nonnegative and sum toAH​\(λγ\)A\_\{H\}\(\\lambda\_\{\\gamma\}\), giving

\|bγ,N,H\|≤a−1​maxj​\|νj\|≤CH​sγ,N\.\|b\_\{\\gamma,N,H\}\|\\leq a^\{\-1\}\\max\_\{j\}\|\\nu\_\{j\}\|\\leq C\_\{H\}s\_\{\\gamma,N\}\.This proves Equation \([46](https://arxiv.org/html/2608.26765#A5.E46)\) without any signed\-third\-moment input\.

Finally, subtractbγ,N,Hb\_\{\\gamma,N,H\}from Equation \([50](https://arxiv.org/html/2608.26765#A5.E50)\)\. For fixedh≤Hh\\leq H,

\|λγh−1\|≤a​h​γ,\|\\lambda\_\{\\gamma\}^\{h\}\-1\|\\leq ah\\gamma,so Equations \([45](https://arxiv.org/html/2608.26765#A5.E45)\) and \([46](https://arxiv.org/html/2608.26765#A5.E46)\) give

\|mh−bγ,N,H\|≤CH​γ​sγ,N=CH​rγ,N\.\|m\_\{h\}\-b\_\{\\gamma,N,H\}\|\\leq C\_\{H\}\\gamma s\_\{\\gamma,N\}=C\_\{H\}r\_\{\\gamma,N\}\.Sum over fixedHHto obtain Equation \([47](https://arxiv.org/html/2608.26765#A5.E47)\)\. ∎

### E\.2Exact linear/nonlinear global split

For the remaining global\-moment arguments, define

ργ:=λγH,\\rho\_\{\\gamma\}:=\\lambda\_\{\\gamma\}^\{H\},and the exact decomposition obtained by unrolling the local recursion written asf′​\(y\)=a​y\+q⁡\(y\)f^\{\\prime\}\(y\)=ay\+q\(y\):

x\+=ργ​x\+η\+Δ,x^\{\+\}=\\rho\_\{\\gamma\}x\+\\eta\+\\Delta,\(52\)where

η\\displaystyle\\eta:=−γ∑j=1HλγH−jε¯j,\\displaystyle:=\-\\gamma\\sum\_\{j=1\}^\{H\}\\lambda\_\{\\gamma\}^\{H\-j\}\\bar\{\\varepsilon\}\_\{j\},\(53\)Δ\\displaystyle\\Delta:=−γ∑h=0H−1λγH−1−hq¯h,q¯h:=1N∑c=1Nq\(θc,h\)\.\\displaystyle:=\-\\gamma\\sum\_\{h=0\}^\{H\-1\}\\lambda\_\{\\gamma\}^\{H\-1\-h\}\\bar\{q\}\_\{h\},\\qquad\\bar\{q\}\_\{h\}:=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}q\(\\theta\_\{c,h\}\)\.\(54\)The linear control term cancels exactly after the client average\.

The fresh\-noise term satisfies

𝔼⁡\[η\]\\displaystyle\\mathbb\{E\}\[\\eta\]=0,\\displaystyle=0,\(55\)𝔼⁡\[η2\]\\displaystyle\\mathbb\{E\}\[\\eta^\{2\}\]=γ2​σ2N​∑k=0H−1λγ2​k,\\displaystyle=\\frac\{\\gamma^\{2\}\\sigma^\{2\}\}\{N\}\\sum\_\{k=0\}^\{H\-1\}\\lambda\_\{\\gamma\}^\{2k\},\(56\)\|𝔼⁡\[η3\]\|\\displaystyle\|\\mathbb\{E\}\[\\eta^\{3\}\]\|≤CH​γ3N2,\\displaystyle\\leq C\_\{H\}\\frac\{\\gamma^\{3\}\}\{N^\{2\}\},\(57\)𝔼⁡\[η4\]\\displaystyle\\mathbb\{E\}\[\\eta^\{4\}\]≤CH​γ4N2\.\\displaystyle\\leq C\_\{H\}\\frac\{\\gamma^\{4\}\}\{N^\{2\}\}\.\(58\)The third\-moment estimate holds without noise symmetry\.

###### Lemma E\.2\(Nonlinear\-increment moments\)\.

For theΔ\\Deltain Equation \([54](https://arxiv.org/html/2608.26765#A5.E54)\),

𝔼⁡\[Δ2\]\\displaystyle\\mathbb\{E\}\[\\Delta^\{2\}\]≤CH​γ2​sγ,N,\\displaystyle\\leq C\_\{H\}\\gamma^\{2\}s\_\{\\gamma,N\},\(59\)𝔼⁡\[Δ4\]\\displaystyle\\mathbb\{E\}\[\\Delta^\{4\}\]≤CH​γ7,\\displaystyle\\leq C\_\{H\}\\gamma^\{7\},\(60\)𝔼​\|Δ\|6\\displaystyle\\mathbb\{E\}\|\\Delta\|^\{6\}≤CH​γ9\.\\displaystyle\\leq C\_\{H\}\\gamma^\{9\}\.\(61\)A second, weaker but sometimes useful estimate is

𝔼⁡\[Δ4\]≤CH​γ4​\(M4\+γ4\),M4:=𝔼⁡\[x4\]\.\\mathbb\{E\}\[\\Delta^\{4\}\]\\leq C\_\{H\}\\gamma^\{4\}\(M\_\{4\}\+\\gamma^\{4\}\),\\qquad M\_\{4\}:=\\mathbb\{E\}\[x^\{4\}\]\.\(62\)

###### Proof\.

Besides the quadratic bound in Equation \([48](https://arxiv.org/html/2608.26765#A5.E48)\), smoothness gives the global linear bound

\|q⁡\(y\)\|≤\(L\+a\)​\|y\|\.\|q\(y\)\|\\leq\(L\+a\)\|y\|\.The earlier coarse local fourth\- and sixth\-moment bounds give, uniformly forh≤Hh\\leq H,

1N​∑c𝔼​\|θc,h\|4≤CH​γ2,1N​∑c𝔼​\|θc,h\|6≤CH​γ3\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|\\theta\_\{c,h\}\|^\{4\}\\leq C\_\{H\}\\gamma^\{2\},\\qquad\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|\\theta\_\{c,h\}\|^\{6\}\\leq C\_\{H\}\\gamma^\{3\}\.Using the quadratic bound, Jensen, and fixedHHyields𝔼⁡\[Δ2\]≤CH​γ4\\mathbb\{E\}\[\\Delta^\{2\}\]\\leq C\_\{H\}\\gamma^\{4\}, which is stronger than Equation \([59](https://arxiv.org/html/2608.26765#A5.E59)\) becausesγ,N≥γ2s\_\{\\gamma,N\}\\geq\\gamma^\{2\}\.

For the fourth moment, combine the linear and quadratic bounds to obtain

\|q⁡\(y\)\|4≤C​\|y\|6\.\|q\(y\)\|^\{4\}\\leq C\|y\|^\{6\}\.Jensen then yields

𝔼⁡\[Δ4\]≤CH​γ4​∑h=0H−11N​∑c𝔼​\|θc,h\|6≤CH​γ7,\\mathbb\{E\}\[\\Delta^\{4\}\]\\leq C\_\{H\}\\gamma^\{4\}\\sum\_\{h=0\}^\{H\-1\}\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|\\theta\_\{c,h\}\|^\{6\}\\leq C\_\{H\}\\gamma^\{7\},which is the sharpened estimate needed for the fourth\-moment absorption argument\.

For completeness, we now prove the weaker estimate in Equation \([62](https://arxiv.org/html/2608.26765#A5.E62)\) without appealing to a later lemma\. The same finite\-step argument used in[LemmaC\.4](https://arxiv.org/html/2608.26765#A3.Thmtheorem4)gives pathwise, forδc,h:=θc,h−x\\delta\_\{c,h\}:=\\theta\_\{c,h\}\-xand fixedh≤Hh\\leq H,

\|δc,h\|≤CH​γ​\(\|x\|\+\|ξc\|\+∑j=1h\|εc,j\|\)\.\|\\delta\_\{c,h\}\|\\leq C\_\{H\}\\gamma\\left\(\|x\|\+\|\\xi\_\{c\}\|\+\\sum\_\{j=1\}^\{h\}\|\\varepsilon\_\{c,j\}\|\\right\)\.Raise this inequality to the fourth power, average over clients and expectations, and use the earlier coarse global fourth\-moment bound, the client\-averaged fourth control moment from[LemmaC\.1](https://arxiv.org/html/2608.26765#A3.Thmtheorem1), and bounded noise\. This yields

1N​∑c𝔼​\|δc,h\|4≤CH​γ4\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|\\delta\_\{c,h\}\|^\{4\}\\leq C\_\{H\}\\gamma^\{4\}\.Consequently,

1N​∑c𝔼​\|θc,h\|4≤8​M4\+8​1N​∑c𝔼​\|δc,h\|4≤CH​\(M4\+γ4\)\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|\\theta\_\{c,h\}\|^\{4\}\\leq 8M\_\{4\}\+8\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|\\delta\_\{c,h\}\|^\{4\}\\leq C\_\{H\}\(M\_\{4\}\+\\gamma^\{4\}\)\.Using the global linear bound\|q⁡\(y\)\|≤\(L\+a\)​\|y\|\|q\(y\)\|\\leq\(L\+a\)\|y\|, Jensen over clients and the fixed sum over local steps now gives

𝔼⁡\[Δ4\]≤CH​γ4​∑h=0H−11N​∑c𝔼​\|θc,h\|4≤CH​γ4​\(M4\+γ4\),\\mathbb\{E\}\[\\Delta^\{4\}\]\\leq C\_\{H\}\\gamma^\{4\}\\sum\_\{h=0\}^\{H\-1\}\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\|\\theta\_\{c,h\}\|^\{4\}\\leq C\_\{H\}\\gamma^\{4\}\(M\_\{4\}\+\\gamma^\{4\}\),which is exactly Equation \([62](https://arxiv.org/html/2608.26765#A5.E62)\)\. Thus no forward reference to the later displacement\-fourth\-moment lemma is needed\.

Finally,\|q⁡\(y\)\|6≤C​\|y\|6\|q\(y\)\|^\{6\}\\leq C\|y\|^\{6\}by the linear bound, so the same argument gives Equation \([61](https://arxiv.org/html/2608.26765#A5.E61)\)\. ∎

### E\.3Global fourth moment

###### Lemma E\.3\(Global fourth moment\)\.

There existsCH<∞C\_\{H\}<\\inftysuch that

M4:=𝔼⁡\[x4\]≤CH​vγ,N=CH​\(γ2N2\+γ3\)\.M\_\{4\}:=\\mathbb\{E\}\[x^\{4\}\]\\leq C\_\{H\}v\_\{\\gamma,N\}=C\_\{H\}\\left\(\\frac\{\\gamma^\{2\}\}\{N^\{2\}\}\+\\gamma^\{3\}\\right\)\.\(63\)

###### Proof\.

LetY:=ργ​x\+ηY:=\\rho\_\{\\gamma\}x\+\\eta\. Stationarity and Equation \([52](https://arxiv.org/html/2608.26765#A5.E52)\) give

M4=\\displaystyle M\_\{4\}=\{\}𝔼⁡\[Y4\]\+4​𝔼​\[Y3​Δ\]\+6​𝔼​\[Y2​Δ2\]\+4​𝔼​\[Y​Δ3\]\+𝔼⁡\[Δ4\]\.\\displaystyle\\mathbb\{E\}\[Y^\{4\}\]\+4\\mathbb\{E\}\[Y^\{3\}\\Delta\]\+6\\mathbb\{E\}\[Y^\{2\}\\Delta^\{2\}\]\+4\\mathbb\{E\}\[Y\\Delta^\{3\}\]\+\\mathbb\{E\}\[\\Delta^\{4\}\]\.\(64\)Becauseη\\etais independent of the round\-startxxand has zero mean,

𝔼⁡\[Y4\]=ργ4​M4\+6​ργ2​M2​𝔼​\[η2\]\+4​ργ​bγ,N,H​𝔼​\[η3\]\+𝔼⁡\[η4\]\.\\mathbb\{E\}\[Y^\{4\}\]=\\rho\_\{\\gamma\}^\{4\}M\_\{4\}\+6\\rho\_\{\\gamma\}^\{2\}M\_\{2\}\\mathbb\{E\}\[\\eta^\{2\}\]\+4\\rho\_\{\\gamma\}b\_\{\\gamma,N,H\}\\mathbb\{E\}\[\\eta^\{3\}\]\+\\mathbb\{E\}\[\\eta^\{4\}\]\.\(65\)By[LemmaC\.5](https://arxiv.org/html/2608.26765#A3.Thmtheorem5), Equations \([56](https://arxiv.org/html/2608.26765#A5.E56)\)–\([58](https://arxiv.org/html/2608.26765#A5.E58)\), and\|bγ,N,H\|≤M2\|b\_\{\\gamma,N,H\}\|\\leq\\sqrt\{M\_\{2\}\},

M2​𝔼​\[η2\]\+\|bγ,N,H​𝔼​\[η3\]\|\+𝔼⁡\[η4\]≤CH​γ​vγ,N\.M\_\{2\}\\mathbb\{E\}\[\\eta^\{2\}\]\+\|b\_\{\\gamma,N,H\}\\mathbb\{E\}\[\\eta^\{3\}\]\|\+\\mathbb\{E\}\[\\eta^\{4\}\]\\leq C\_\{H\}\\gamma v\_\{\\gamma,N\}\.
For the nonlinear terms, fixϵ\>0\\epsilon\>0\. Young’s inequality and the sharpened Equation \([60](https://arxiv.org/html/2608.26765#A5.E60)\) give

\|𝔼⁡\[Y3​Δ\]\|\\displaystyle\|\\mathbb\{E\}\[Y^\{3\}\\Delta\]\|≤ϵ​γ​𝔼​\[Y4\]\+CH,ϵ​γ−3​𝔼​\[Δ4\]≤CH​ϵ​γ​M4\+CH,ϵ​γ​vγ,N,\\displaystyle\\leq\\epsilon\\gamma\\mathbb\{E\}\[Y^\{4\}\]\+C\_\{H,\\epsilon\}\\gamma^\{\-3\}\\mathbb\{E\}\[\\Delta^\{4\}\]\\leq C\_\{H\}\\epsilon\\gamma M\_\{4\}\+C\_\{H,\\epsilon\}\\gamma v\_\{\\gamma,N\},\|𝔼⁡\[Y2​Δ2\]\|\\displaystyle\|\\mathbb\{E\}\[Y^\{2\}\\Delta^\{2\}\]\|≤ϵ​γ​𝔼​\[Y4\]\+CH,ϵ​γ−1​𝔼​\[Δ4\]≤CH​ϵ​γ​M4\+CH,ϵ​γ​vγ,N,\\displaystyle\\leq\\epsilon\\gamma\\mathbb\{E\}\[Y^\{4\}\]\+C\_\{H,\\epsilon\}\\gamma^\{\-1\}\\mathbb\{E\}\[\\Delta^\{4\}\]\\leq C\_\{H\}\\epsilon\\gamma M\_\{4\}\+C\_\{H,\\epsilon\}\\gamma v\_\{\\gamma,N\},\|𝔼⁡\[Y​Δ3\]\|\\displaystyle\|\\mathbb\{E\}\[Y\\Delta^\{3\}\]\|≤ϵγ𝔼\[Y4\]\+CH,ϵγ−1/3𝔼\[Δ4\]≤CHϵγM4\+CH,ϵγvγ,N,\\displaystyle\\leq\\epsilon\\gamma\\mathbb\{E\}\[Y^\{4\}\]\+C\_\{H,\\epsilon\}\\gamma^\{\-1/3\}\\mathbb\{E\}\[\\Delta^\{4\}\]\\leq C\_\{H\}\\epsilon\\gamma M\_\{4\}\+C\_\{H,\\epsilon\}\\gamma v\_\{\\gamma,N\},and𝔼⁡\[Δ4\]≤CH​γ​vγ,N\\mathbb\{E\}\[\\Delta^\{4\}\]\\leq C\_\{H\}\\gamma v\_\{\\gamma,N\}forγ≤1\\gamma\\leq 1\.

Finally,

1−ργ4=a​γ​∑k=04​H−1\(1−a​γ\)k≥a​γ1\-\\rho\_\{\\gamma\}^\{4\}=a\\gamma\\sum\_\{k=0\}^\{4H\-1\}\(1\-a\\gamma\)^\{k\}\\geq a\\gammawhen0≤1−a​γ≤10\\leq 1\-a\\gamma\\leq 1\. Move theργ4​M4\\rho\_\{\\gamma\}^\{4\}M\_\{4\}term to the left of Equation \([64](https://arxiv.org/html/2608.26765#A5.E64)\), choose a fixedϵ\\epsilonsmall enough that theϵ​γ​M4\\epsilon\\gamma M\_\{4\}terms are absorbed, and divide byγ\\gammato obtain Equation \([63](https://arxiv.org/html/2608.26765#A5.E63)\)\. ∎

###### Proof note E\.4\(Why Equation \([60](https://arxiv.org/html/2608.26765#A5.E60)\) is needed\)\.

The weaker estimate in Equation \([62](https://arxiv.org/html/2608.26765#A5.E62)\) is true but does not by itself justify theY3​ΔY^\{3\}\\Deltaabsorption: after the factorγ−3\\gamma^\{\-3\}from Young’s inequality, it produces a coefficient of orderCϵ​γ​M4C\_\{\\epsilon\}\\gamma M\_\{4\}, whose constant worsens asϵ↓0\\epsilon\\downarrow 0\. The independent bound𝔼⁡\[Δ4\]≤CH​γ7\\mathbb\{E\}\[\\Delta^\{4\}\]\\leq C\_\{H\}\\gamma^\{7\}removes this defect\. The sharper bound is needed to complete the absorption argument\.

### E\.4Global signed third moment

###### Lemma E\.5\(Global signed third moment\)\.

There existsCH<∞C\_\{H\}<\\inftysuch that

\|M3\|:=\|𝔼⁡\[x3\]\|≤CH​rγ,N\.\|M\_\{3\}\|:=\|\\mathbb\{E\}\[x^\{3\}\]\|\\leq C\_\{H\}r\_\{\\gamma,N\}\.\(66\)In fact, the proof below yields the stronger internal estimate

\|M3\|≤CH​vγ,N\.\|M\_\{3\}\|\\leq C\_\{H\}v\_\{\\gamma,N\}\.\(67\)Only Equation \([66](https://arxiv.org/html/2608.26765#A5.E66)\) is needed for the main theorem\.

###### Proof\.

Again setY:=ργ​x\+ηY:=\\rho\_\{\\gamma\}x\+\\eta\. Stationarity gives the exact identity

\(1−ργ3\)​M3=\\displaystyle\(1\-\\rho\_\{\\gamma\}^\{3\}\)M\_\{3\}=\{\}3​ργ​bγ,N,H​𝔼​\[η2\]\+𝔼⁡\[η3\]\\displaystyle 3\\rho\_\{\\gamma\}b\_\{\\gamma,N,H\}\\mathbb\{E\}\[\\eta^\{2\}\]\+\\mathbb\{E\}\[\\eta^\{3\}\]\(68\)\+3​𝔼​\[Y2​Δ\]\+3​𝔼​\[Y​Δ2\]\+𝔼⁡\[Δ3\]\.\\displaystyle\+3\\mathbb\{E\}\[Y^\{2\}\\Delta\]\+3\\mathbb\{E\}\[Y\\Delta^\{2\}\]\+\\mathbb\{E\}\[\\Delta^\{3\}\]\.The first two terms areOH​\(γ​vγ,N\)O\_\{H\}\(\\gamma v\_\{\\gamma,N\}\)by[LemmaE\.1](https://arxiv.org/html/2608.26765#A5.Thmtheorem1)and Equations \([56](https://arxiv.org/html/2608.26765#A5.E56)\) and \([57](https://arxiv.org/html/2608.26765#A5.E57)\)\.

The delicate term is𝔼⁡\[Y2​Δ\]\\mathbb\{E\}\[Y^\{2\}\\Delta\]\. A generic Cauchy–Schwarz bound is too coarse, so use the structure in Equations \([54](https://arxiv.org/html/2608.26765#A5.E54)\) and \([48](https://arxiv.org/html/2608.26765#A5.E48)\):

\|𝔼⁡\[Y2​Δ\]\|≤CH​γ​∑h=0H−11N​∑c𝔼⁡\[Y2​θc,h2\]\.\|\\mathbb\{E\}\[Y^\{2\}\\Delta\]\|\\leq C\_\{H\}\\gamma\\sum\_\{h=0\}^\{H\-1\}\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[Y^\{2\}\\theta\_\{c,h\}^\{2\}\]\.Letδc,h:=θc,h−x\\delta\_\{c,h\}:=\\theta\_\{c,h\}\-x\. A finite\-step fourth\-moment argument using[LemmasC\.1](https://arxiv.org/html/2608.26765#A3.Thmtheorem1)and[E\.3](https://arxiv.org/html/2608.26765#A5.Thmtheorem3)gives

1N​∑c𝔼⁡\[δc,h4\]≤CH​γ4\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\delta\_\{c,h\}^\{4\}\]\\leq C\_\{H\}\\gamma^\{4\}\.\(69\)SinceY2≤CH​\(x2\+η2\)Y^\{2\}\\leq C\_\{H\}\(x^\{2\}\+\\eta^\{2\}\)andθc,h2≤2​x2\+2​δc,h2\\theta\_\{c,h\}^\{2\}\\leq 2x^\{2\}\+2\\delta\_\{c,h\}^\{2\}, the joint probability–client average is bounded by a constant times

M4\+1N​∑c𝔼⁡\[x2​δc,h2\]\+𝔼⁡\[η2​x2\]\+1N​∑c𝔼⁡\[η2​δc,h2\]\.M\_\{4\}\+\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[x^\{2\}\\delta\_\{c,h\}^\{2\}\]\+\\mathbb\{E\}\[\\eta^\{2\}x^\{2\}\]\+\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\eta^\{2\}\\delta\_\{c,h\}^\{2\}\]\.The first three terms areOH​\(vγ,N\)O\_\{H\}\(v\_\{\\gamma,N\}\),OH​\(vγ,N​γ4\)=OH​\(vγ,N\)O\_\{H\}\(\\sqrt\{v\_\{\\gamma,N\}\\,\\gamma^\{4\}\}\)=O\_\{H\}\(v\_\{\\gamma,N\}\), andOH​\(𝔼⁡\[η2\]​M2\)=OH​\(vγ,N\)O\_\{H\}\(\\mathbb\{E\}\[\\eta^\{2\}\]M\_\{2\}\)=O\_\{H\}\(v\_\{\\gamma,N\}\)\. The fourth is

OH​\(𝔼⁡\[η4\]​1N​∑c𝔼⁡\[δc,h4\]\)=OH​\(γ4N\)=OH​\(vγ,N\)\.O\_\{H\}\\\!\\left\(\\sqrt\{\\mathbb\{E\}\[\\eta^\{4\}\]\\,\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\delta\_\{c,h\}^\{4\}\]\}\\right\)=O\_\{H\}\\left\(\\frac\{\\gamma^\{4\}\}\{N\}\\right\)=O\_\{H\}\(v\_\{\\gamma,N\}\)\.Thus

\|𝔼⁡\[Y2​Δ\]\|≤CH​γ​vγ,N\.\|\\mathbb\{E\}\[Y^\{2\}\\Delta\]\|\\leq C\_\{H\}\\gamma v\_\{\\gamma,N\}\.\(70\)
Next,

\|𝔼⁡\[Y​Δ2\]\|≤𝔼⁡\[Y2\]​𝔼​\[Δ4\]≤CH​γ4≤CH​γ​vγ,N,\|\\mathbb\{E\}\[Y\\Delta^\{2\}\]\|\\leq\\sqrt\{\\mathbb\{E\}\[Y^\{2\}\]\\mathbb\{E\}\[\\Delta^\{4\}\]\}\\leq C\_\{H\}\\gamma^\{4\}\\leq C\_\{H\}\\gamma v\_\{\\gamma,N\},where[LemmasC\.5](https://arxiv.org/html/2608.26765#A3.Thmtheorem5)and[E\.2](https://arxiv.org/html/2608.26765#A5.Thmtheorem2)were used\. Finally,

\|𝔼⁡\[Δ3\]\|≤𝔼​\|Δ\|6≤CH​γ9/2≤CH​γ​vγ,N\.\|\\mathbb\{E\}\[\\Delta^\{3\}\]\|\\leq\\sqrt\{\\mathbb\{E\}\|\\Delta\|^\{6\}\}\\leq C\_\{H\}\\gamma^\{9/2\}\\leq C\_\{H\}\\gamma v\_\{\\gamma,N\}\.The restoring factor satisfies

1−ργ3=a​γ​∑k=03​H−1\(1−a​γ\)k≥a​γ\.1\-\\rho\_\{\\gamma\}^\{3\}=a\\gamma\\sum\_\{k=0\}^\{3H\-1\}\(1\-a\\gamma\)^\{k\}\\geq a\\gamma\.Divide Equation \([68](https://arxiv.org/html/2608.26765#A5.E68)\) by this factor to obtain Equation \([67](https://arxiv.org/html/2608.26765#A5.E67)\), and hence Equation \([66](https://arxiv.org/html/2608.26765#A5.E66)\)\. Notice that the skewness contribution𝔼⁡\[η3\]\\mathbb\{E\}\[\\eta^\{3\}\]is retained explicitly; symmetry is unnecessary\. ∎

## Appendix FSharp global second moment

This is the key identifiability step\. It rules out an unresolved client\-number\-independentN0​γ2N^\{0\}\\gamma^\{2\}contribution from the global raw second moment itself\.

###### Lemma F\.1\(Averaged local displacement\)\.

Let

δ¯h:=1N​∑c=1N\(θc,h−x\)\.\\bar\{\\delta\}\_\{h\}:=\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\(\\theta\_\{c,h\}\-x\)\.Then, for fixedh≤Hh\\leq H,

𝔼⁡\[δ¯h2\]≤CH​γ2​\(sγ,N\+1N\)\.\\mathbb\{E\}\[\\bar\{\\delta\}\_\{h\}^\{2\}\]\\leq C\_\{H\}\\gamma^\{2\}\\left\(s\_\{\\gamma,N\}\+\\frac\{1\}\{N\}\\right\)\.\(71\)

###### Proof\.

Average the local displacement recursion\. The control term cancels pathwise, so

δ¯h\+1=δ¯h−γ​f′​\(θh\)¯−γ​ε¯h\+1,δ¯0=0\.\\bar\{\\delta\}\_\{h\+1\}=\\bar\{\\delta\}\_\{h\}\-\\gamma\\overline\{f^\{\\prime\}\(\\theta\_\{h\}\)\}\-\\gamma\\bar\{\\varepsilon\}\_\{h\+1\},\\qquad\\bar\{\\delta\}\_\{0\}=0\.Unrolling,

δ¯h=−γ∑ℓ=0h−1f′​\(θℓ\)¯−γ∑j=1hε¯j\.\\bar\{\\delta\}\_\{h\}=\-\\gamma\\sum\_\{\\ell=0\}^\{h\-1\}\\overline\{f^\{\\prime\}\(\\theta\_\{\\ell\}\)\}\-\\gamma\\sum\_\{j=1\}^\{h\}\\bar\{\\varepsilon\}\_\{j\}\.For fixedHH, Jensen’s inequality, smoothness, and the coarse local second\-moment estimate give

𝔼​\|f′​\(θℓ\)¯\|2≤L2​1N​∑c𝔼⁡\[θc,ℓ2\]≤CH​sγ,N,\\mathbb\{E\}\|\\overline\{f^\{\\prime\}\(\\theta\_\{\\ell\}\)\}\|^\{2\}\\leq L^\{2\}\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\theta\_\{c,\\ell\}^\{2\}\]\\leq C\_\{H\}s\_\{\\gamma,N\},while𝔼⁡\[ε¯j2\]=σ2/N\\mathbb\{E\}\[\\bar\{\\varepsilon\}\_\{j\}^\{2\}\]=\\sigma^\{2\}/N\. This gives Equation \([71](https://arxiv.org/html/2608.26765#A6.E71)\)\. ∎

###### Lemma F\.2\(Sharp global second moment\)\.

There existsCH<∞C\_\{H\}<\\inftysuch that

M2=σ22​a​γN\+OH​\(γ2N\+γ3\)\.M\_\{2\}=\\frac\{\\sigma^\{2\}\}\{2a\}\\frac\{\\gamma\}\{N\}\+O\_\{H\}\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\)\.\(72\)Equivalently, the nonlinear correction to the stationary second moment has no unresolvedN0​γ2N^\{0\}\\gamma^\{2\}term\.

###### Proof\.

We use the exact decomposition in Equation \([52](https://arxiv.org/html/2608.26765#A5.E52)\)\.

#### Step 1: exact linearized stationary second moment\.

For the linearized chainxlin\+=ργ​xlin\+ηx\_\{\\mathrm\{lin\}\}^\{\+\}=\\rho\_\{\\gamma\}x\_\{\\mathrm\{lin\}\}\+\\eta, stationarity gives

M2,lin=𝔼⁡\[η2\]1−ργ2\.M\_\{2,\\mathrm\{lin\}\}=\\frac\{\\mathbb\{E\}\[\\eta^\{2\}\]\}\{1\-\\rho\_\{\\gamma\}^\{2\}\}\.By Equation \([56](https://arxiv.org/html/2608.26765#A5.E56)\),

𝔼⁡\[η2\]=γ2​σ2N​∑k=0H−1λγ2​k\.\\mathbb\{E\}\[\\eta^\{2\}\]=\\frac\{\\gamma^\{2\}\\sigma^\{2\}\}\{N\}\\sum\_\{k=0\}^\{H\-1\}\\lambda\_\{\\gamma\}^\{2k\}\.Since

1−ργ2=1−λγ2​H=\(1−λγ2\)​∑k=0H−1λγ2​k,1\-\\rho\_\{\\gamma\}^\{2\}=1\-\\lambda\_\{\\gamma\}^\{2H\}=\(1\-\\lambda\_\{\\gamma\}^\{2\}\)\\sum\_\{k=0\}^\{H\-1\}\\lambda\_\{\\gamma\}^\{2k\},theHH\-dependent geometric sum cancels exactly:

M2,lin=γ​σ2a​N​\(2−a​γ\)\.M\_\{2,\\mathrm\{lin\}\}=\\frac\{\\gamma\\sigma^\{2\}\}\{aN\(2\-a\\gamma\)\}\.\(73\)Therefore

M2,lin=σ22​a​γN\+O⁡\(γ2N\)\.M\_\{2,\\mathrm\{lin\}\}=\\frac\{\\sigma^\{2\}\}\{2a\}\\frac\{\\gamma\}\{N\}\+O\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\\right\)\.\(74\)The explicit\+σ2γ2/\(4N\)\+\\sigma^\{2\}\\gamma^\{2\}/\(4N\)term in the Taylor expansion of Equation \([73](https://arxiv.org/html/2608.26765#A6.E73)\) is*not*claimed to be the complete nonlinear finite\-NNcoefficient\.

#### Step 2: exact nonlinear stationary equation\.

Squaring Equation \([52](https://arxiv.org/html/2608.26765#A5.E52)\) and using𝔼⁡\[x​η\]=0\\mathbb\{E\}\[x\\eta\]=0gives

\(1−ργ2\)​M2=𝔼⁡\[η2\]\+2​ργ​𝔼​\[x​Δ\]\+2​𝔼​\[η​Δ\]\+𝔼⁡\[Δ2\]\.\(1\-\\rho\_\{\\gamma\}^\{2\}\)M\_\{2\}=\\mathbb\{E\}\[\\eta^\{2\}\]\+2\\rho\_\{\\gamma\}\\mathbb\{E\}\[x\\Delta\]\+2\\mathbb\{E\}\[\\eta\\Delta\]\+\\mathbb\{E\}\[\\Delta^\{2\}\]\.\(75\)By[LemmaE\.2](https://arxiv.org/html/2608.26765#A5.Thmtheorem2),

𝔼⁡\[Δ2\]≤CH​γ​rγ,N\.\\mathbb\{E\}\[\\Delta^\{2\}\]\\leq C\_\{H\}\\gamma r\_\{\\gamma,N\}\.\(76\)It remains to obtain equally sharp bounds for the two covariance terms\.

#### Step 3: the signed covariance𝔼⁡\[x​Δ\]\\mathbb\{E\}\[x\\Delta\]\.

Forδc,h:=θc,h−x\\delta\_\{c,h\}:=\\theta\_\{c,h\}\-x, Taylor\-expandqqaround the round\-startxx:

q⁡\(x\+δc,h\)=q⁡\(x\)\+q′​\(x\)​δc,h\+rc,h,\|rc,h\|≤K32​δc,h2\.q\(x\+\\delta\_\{c,h\}\)=q\(x\)\+q^\{\\prime\}\(x\)\\delta\_\{c,h\}\+r\_\{c,h\},\\qquad\|r\_\{c,h\}\|\\leq\\frac\{K\_\{3\}\}\{2\}\\delta\_\{c,h\}^\{2\}\.\(77\)After the client average, we write

q¯h=q⁡\(x\)\+q′​\(x\)​δ¯h\+r¯h\\bar\{q\}\_\{h\}=q\(x\)\+q^\{\\prime\}\(x\)\\bar\{\\delta\}\_\{h\}\+\\bar\{r\}\_\{h\}and hence

Δ=Δ0\+R1\+R2,\\Delta=\\Delta\_\{0\}\+R\_\{1\}\+R\_\{2\},\(78\)with

Δ0\\displaystyle\\Delta\_\{0\}:=−γ​AH​\(λγ\)​q​\(x\),\\displaystyle:=\-\\gamma A\_\{H\}\(\\lambda\_\{\\gamma\}\)q\(x\),R1\\displaystyle R\_\{1\}:=−γ∑h=0H−1λγH−1−hq′\(x\)δ¯h,\\displaystyle:=\-\\gamma\\sum\_\{h=0\}^\{H\-1\}\\lambda\_\{\\gamma\}^\{H\-1\-h\}q^\{\\prime\}\(x\)\\bar\{\\delta\}\_\{h\},R2\\displaystyle R\_\{2\}:=−γ∑h=0H−1λγH−1−hr¯h\.\\displaystyle:=\-\\gamma\\sum\_\{h=0\}^\{H\-1\}\\lambda\_\{\\gamma\}^\{H\-1\-h\}\\bar\{r\}\_\{h\}\.Sinceq⁡\(y\)=τ2​y2\+O⁡\(\|y\|3\)q\(y\)=\\frac\{\\tau\}\{2\}y^\{2\}\+O\(\|y\|^\{3\}\), we have

\|𝔼⁡\[x​q​\(x\)\]\|≤C⁡\(\|M3\|\+M4\)≤CH​rγ,N\|\\mathbb\{E\}\[xq\(x\)\]\|\\leq C\(\|M\_\{3\}\|\+M\_\{4\}\)\\leq C\_\{H\}r\_\{\\gamma,N\}by[LemmasE\.5](https://arxiv.org/html/2608.26765#A5.Thmtheorem5)and[E\.3](https://arxiv.org/html/2608.26765#A5.Thmtheorem3)\. Thus

\|𝔼⁡\[x​Δ0\]\|≤CH​γ​rγ,N\.\|\\mathbb\{E\}\[x\\Delta\_\{0\}\]\|\\leq C\_\{H\}\\gamma r\_\{\\gamma,N\}\.\(79\)Moreover\|q′​\(x\)\|≤K3​\|x\|\|q^\{\\prime\}\(x\)\|\\leq K\_\{3\}\|x\|, so[LemmasE\.3](https://arxiv.org/html/2608.26765#A5.Thmtheorem3)and[F\.1](https://arxiv.org/html/2608.26765#A6.Thmtheorem1)give

\|𝔼⁡\[x​R1\]\|≤CH​γ​M4​𝔼​\[δ¯h2\]≤CH​γ2​vγ,N​\(sγ,N\+1N\)\.\|\\mathbb\{E\}\[xR\_\{1\}\]\|\\leq C\_\{H\}\\gamma\\sqrt\{M\_\{4\}\\,\\mathbb\{E\}\[\\bar\{\\delta\}\_\{h\}^\{2\}\]\}\\leq C\_\{H\}\\gamma^\{2\}\\sqrt\{v\_\{\\gamma,N\}\\left\(s\_\{\\gamma,N\}\+\\frac\{1\}\{N\}\\right\)\}\.Becausevγ,N≤γ​sγ,Nv\_\{\\gamma,N\}\\leq\\gamma s\_\{\\gamma,N\}andγ⁡\(sγ,N\+1/N\)≤2​sγ,N\\gamma\(s\_\{\\gamma,N\}\+1/N\)\\leq 2s\_\{\\gamma,N\}forγ≤1\\gamma\\leq 1,

\|𝔼⁡\[x​R1\]\|≤CH​γ2​sγ,N=CH​γ​rγ,N\.\|\\mathbb\{E\}\[xR\_\{1\}\]\|\\leq C\_\{H\}\\gamma^\{2\}s\_\{\\gamma,N\}=C\_\{H\}\\gamma r\_\{\\gamma,N\}\.\(80\)ForR2R\_\{2\}, Equation \([69](https://arxiv.org/html/2608.26765#A5.E69)\) and Jensen imply𝔼⁡\[r¯h2\]≤CH​γ4\\mathbb\{E\}\[\\bar\{r\}\_\{h\}^\{2\}\]\\leq C\_\{H\}\\gamma^\{4\}, hence

\|𝔼⁡\[x​R2\]\|≤CH​γ​M2​γ4=CH​γ3​sγ,N≤CH​γ2​sγ,N\.\|\\mathbb\{E\}\[xR\_\{2\}\]\|\\leq C\_\{H\}\\gamma\\sqrt\{M\_\{2\}\\gamma^\{4\}\}=C\_\{H\}\\gamma^\{3\}\\sqrt\{s\_\{\\gamma,N\}\}\\leq C\_\{H\}\\gamma^\{2\}s\_\{\\gamma,N\}\.Combining the three pieces,

\|𝔼⁡\[x​Δ\]\|≤CH​γ​rγ,N\.\|\\mathbb\{E\}\[x\\Delta\]\|\\leq C\_\{H\}\\gamma r\_\{\\gamma,N\}\.\(81\)

#### Step 4: the fresh\-noise covariance𝔼⁡\[η​Δ\]\\mathbb\{E\}\[\\eta\\Delta\]by coordinate replacement\.

A naive Cauchy–Schwarz bound is too coarse\. First note thatΔ0\\Delta\_\{0\}in Equation \([78](https://arxiv.org/html/2608.26765#A6.E78)\) depends only on the round\-startxx, so

𝔼⁡\[η​Δ0\]=0\.\\mathbb\{E\}\[\\eta\\Delta\_\{0\}\]=0\.\(82\)LetR:=Δ−Δ0R:=\\Delta\-\\Delta\_\{0\}\. Index theN​HNHfresh noise coordinates byi=\(c,j\)i=\(c,j\)and write

η=∑iαi​εi,\|αi\|≤γN\.\\eta=\\sum\_\{i\}\\alpha\_\{i\}\\varepsilon\_\{i\},\\qquad\|\\alpha\_\{i\}\|\\leq\\frac\{\\gamma\}\{N\}\.For a fixed coordinateii, letZ\(i,0\)Z^\{\(i,0\)\}denote the noise array obtained by replacingεi\\varepsilon\_\{i\}by00while keeping all other coordinates unchanged\. SinceR⁡\(Z\(i,0\)\)R\(Z^\{\(i,0\)\}\)is independent ofεi\\varepsilon\_\{i\}and𝔼⁡\[εi\]=0\\mathbb\{E\}\[\\varepsilon\_\{i\}\]=0,

𝔼⁡\[εi​R​\(Z\)\]=𝔼⁡\[εi​\{R⁡\(Z\)−R⁡\(Z\(i,0\)\)\}\]\.\\mathbb\{E\}\[\\varepsilon\_\{i\}R\(Z\)\]=\\mathbb\{E\}\\left\[\\varepsilon\_\{i\}\\\{R\(Z\)\-R\(Z^\{\(i,0\)\}\)\\\}\\right\]\.\(83\)We now justify the required path sensitivity deterministically\. Writei=\(c0,j0\)i=\(c\_\{0\},j\_\{0\}\)and couple the two within\-round trajectories by using the same round\-start state and the same fresh noises except that the coordinateεc0,j0\\varepsilon\_\{c\_\{0\},j\_\{0\}\}is replaced by00inZ\(i,0\)Z^\{\(i,0\)\}\. All clientsc≠c0c\\neq c\_\{0\}then have identical trajectories\. For the affected client, we define

dh:=θc0,h​\(Z\)−θc0,h​\(Z\(i,0\)\)\.d\_\{h\}:=\\theta\_\{c\_\{0\},h\}\(Z\)\-\\theta\_\{c\_\{0\},h\}\(Z^\{\(i,0\)\}\)\.Before the perturbed noise is used,dh=0d\_\{h\}=0forh<j0h<j\_\{0\}, while the update containing that coordinate gives

dj0=−γ​εi\.d\_\{j\_\{0\}\}=\-\\gamma\\varepsilon\_\{i\}\.For every subsequent local steph≥j0h\\geq j\_\{0\}, the two recursions use the same control and the same remaining fresh noises, hence

dh\+1=dh−γ⁡\{f′​\(θc0,h​\(Z\)\)−f′​\(θc0,h​\(Z\(i,0\)\)\)\}\.d\_\{h\+1\}=d\_\{h\}\-\\gamma\\\{f^\{\\prime\}\(\\theta\_\{c\_\{0\},h\}\(Z\)\)\-f^\{\\prime\}\(\\theta\_\{c\_\{0\},h\}\(Z^\{\(i,0\)\}\)\)\\\}\.ByLL\-smoothness,

\|dh\+1\|≤\(1\+γ​L\)​\|dh\|,\|d\_\{h\+1\}\|\\leq\(1\+\\gamma L\)\|d\_\{h\}\|,and therefore, forj0≤h≤Hj\_\{0\}\\leq h\\leq H,

\|dh\|≤γ​\(1\+γ​L\)h−j0​\|εi\|≤CH​γ​\|εi\|,\|d\_\{h\}\|\\leq\\gamma\(1\+\\gamma L\)^\{h\-j\_\{0\}\}\|\\varepsilon\_\{i\}\|\\leq C\_\{H\}\\gamma\|\\varepsilon\_\{i\}\|,where the last constant is uniform inNNandγ\\gammaunder the common small\-step restriction\.

BecauseΔ0\\Delta\_\{0\}depends only on the round\-startxx, the differenceR=Δ−Δ0R=\\Delta\-\\Delta\_\{0\}is the same as the difference ofΔ\\Delta\. Only clientc0c\_\{0\}contributes to the change in each client averageq¯h\\bar\{q\}\_\{h\}\. Since

q′​\(y\)=f′′​\(y\)−a,\|q′​\(y\)\|≤L\+a,q^\{\\prime\}\(y\)=f^\{\\prime\\prime\}\(y\)\-a,\\qquad\|q^\{\\prime\}\(y\)\|\\leq L\+a,we obtain

\|q¯h​\(Z\)−q¯h​\(Z\(i,0\)\)\|≤L\+aN​\|dh\|≤CH​γN​\|εi\|\.\|\\bar\{q\}\_\{h\}\(Z\)\-\\bar\{q\}\_\{h\}\(Z^\{\(i,0\)\}\)\|\\leq\\frac\{L\+a\}\{N\}\|d\_\{h\}\|\\leq C\_\{H\}\\frac\{\\gamma\}\{N\}\|\\varepsilon\_\{i\}\|\.Moreover, under the same threshold0≤λγ≤10\\leq\\lambda\_\{\\gamma\}\\leq 1, the weights in Equation \([54](https://arxiv.org/html/2608.26765#A5.E54)\) have absolute value at most one\. Summing at mostHHaffected local steps and using the outer factorγ\\gammain Equation \([54](https://arxiv.org/html/2608.26765#A5.E54)\) yields the deterministic sensitivity bound

\|R⁡\(Z\)−R⁡\(Z\(i,0\)\)\|≤CH​γ2N​\|εi\|\.\|R\(Z\)\-R\(Z^\{\(i,0\)\}\)\|\\leq C\_\{H\}\\frac\{\\gamma^\{2\}\}\{N\}\|\\varepsilon\_\{i\}\|\.\(84\)Hence

\|𝔼⁡\[εi​R\]\|≤CH​γ2N​σ2\.\|\\mathbb\{E\}\[\\varepsilon\_\{i\}R\]\|\\leq C\_\{H\}\\frac\{\\gamma^\{2\}\}\{N\}\\sigma^\{2\}\.Summing theN​HNHcoordinates inη\\eta,

\|𝔼⁡\[η​Δ\]\|=\|𝔼⁡\[η​R\]\|≤N​H​\(CH​γN\)​\(CH​γ2N\)≤CH​γ3N\.\|\\mathbb\{E\}\[\\eta\\Delta\]\|=\|\\mathbb\{E\}\[\\eta R\]\|\\leq NH\\left\(C\_\{H\}\\frac\{\\gamma\}\{N\}\\right\)\\left\(C\_\{H\}\\frac\{\\gamma^\{2\}\}\{N\}\\right\)\\leq C\_\{H\}\\frac\{\\gamma^\{3\}\}\{N\}\.\(85\)This is the extra1/N1/Ngain that the generic Cauchy–Schwarz route misses\.

#### Step 5: conclude\.

Insert Equations \([76](https://arxiv.org/html/2608.26765#A6.E76)\), \([81](https://arxiv.org/html/2608.26765#A6.E81)\), and \([85](https://arxiv.org/html/2608.26765#A6.E85)\) into Equation \([75](https://arxiv.org/html/2608.26765#A6.E75)\):

\(1−ργ2\)​M2=𝔼⁡\[η2\]\+OH​\(γ3N\+γ4\)\.\(1\-\\rho\_\{\\gamma\}^\{2\}\)M\_\{2\}=\\mathbb\{E\}\[\\eta^\{2\}\]\+O\_\{H\}\\left\(\\frac\{\\gamma^\{3\}\}\{N\}\+\\gamma^\{4\}\\right\)\.Because

1−ργ2=a​γ​∑k=02​H−1\(1−a​γ\)k≥a​γ,1\-\\rho\_\{\\gamma\}^\{2\}=a\\gamma\\sum\_\{k=0\}^\{2H\-1\}\(1\-a\\gamma\)^\{k\}\\geq a\\gamma,we obtain

M2=M2,lin\+OH​\(γ2N\+γ3\)\.M\_\{2\}=M\_\{2,\\mathrm\{lin\}\}\+O\_\{H\}\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\)\.Use Equation \([74](https://arxiv.org/html/2608.26765#A6.E74)\) to conclude Equation \([72](https://arxiv.org/html/2608.26765#A6.E72)\)\. ∎

## Appendix GLocal higher moments and Taylor control

This section places the cubic and quartic pieces of the gradient expansion inside the uniform remainder, thereby excluding an additional client\-number\-independentN0​γ2N^\{0\}\\gamma^\{2\}contribution from those terms\.

###### Lemma G\.1\(Local displacement fourth moment\)\.

For every fixedh≤Hh\\leq H,

1N​∑c=1N𝔼​\|θc,h−x\|4≤CH​γ4\.\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\|\\theta\_\{c,h\}\-x\|^\{4\}\\leq C\_\{H\}\\gamma^\{4\}\.\(86\)

###### Proof\.

The finite\-step bound used in[LemmaC\.4](https://arxiv.org/html/2608.26765#A3.Thmtheorem4)also yields

\|θc,h−x\|≤CH​γ​\(\|x\|\+\|ξc\|\+∑j=1h\|εc,j\|\)\.\|\\theta\_\{c,h\}\-x\|\\leq C\_\{H\}\\gamma\\left\(\|x\|\+\|\\xi\_\{c\}\|\+\\sum\_\{j=1\}^\{h\}\|\\varepsilon\_\{c,j\}\|\\right\)\.Raise to the fourth power, average over clients, and use[LemmasE\.3](https://arxiv.org/html/2608.26765#A5.Thmtheorem3)and[C\.1](https://arxiv.org/html/2608.26765#A3.Thmtheorem1)together with bounded noise\. The resulting constant is independent ofNNandγ\\gammafor fixedHH\. ∎

###### Lemma G\.2\(Local signed cubic and fourth moments\)\.

For every fixedh≤Hh\\leq H,

\|1N​∑c=1N𝔼⁡\[θc,h3\]\|\\displaystyle\\left\|\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\[\\theta\_\{c,h\}^\{3\}\]\\right\|≤CH​rγ,N,\\displaystyle\\leq C\_\{H\}r\_\{\\gamma,N\},\(87\)1N​∑c=1N𝔼⁡\[θc,h4\]\\displaystyle\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\[\\theta\_\{c,h\}^\{4\}\]≤CH​vγ,N\.\\displaystyle\\leq C\_\{H\}v\_\{\\gamma,N\}\.\(88\)The proof uses the coarseM2M\_\{2\}bound, the globalM3/M4M\_\{3\}/M\_\{4\}bounds, and control/displacement moments; it does not require the sharp coefficient in[LemmaF\.2](https://arxiv.org/html/2608.26765#A6.Thmtheorem2)\.

###### Proof\.

Writeδc,h:=θc,h−x\\delta\_\{c,h\}:=\\theta\_\{c,h\}\-xandδ¯h=N−1​∑cδc,h\\bar\{\\delta\}\_\{h\}=N^\{\-1\}\\sum\_\{c\}\\delta\_\{c,h\}\. We already have

1N​∑c𝔼⁡\[δc,h2\]≤CH​γ2,1N​∑c𝔼⁡\[δc,h4\]≤CH​γ4,\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\delta\_\{c,h\}^\{2\}\]\\leq C\_\{H\}\\gamma^\{2\},\\qquad\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\delta\_\{c,h\}^\{4\}\]\\leq C\_\{H\}\\gamma^\{4\},from[LemmasC\.4](https://arxiv.org/html/2608.26765#A3.Thmtheorem4)and[G\.1](https://arxiv.org/html/2608.26765#A7.Thmtheorem1), and

𝔼⁡\[δ¯h2\]≤CH​γ2​\(sγ,N\+1N\)\\mathbb\{E\}\[\\bar\{\\delta\}\_\{h\}^\{2\}\]\\leq C\_\{H\}\\gamma^\{2\}\\left\(s\_\{\\gamma,N\}\+\\frac\{1\}\{N\}\\right\)from[LemmaF\.1](https://arxiv.org/html/2608.26765#A6.Thmtheorem1)\.

For the signed cubic, expand exactly:

1N​∑c𝔼⁡\[θc,h3\]=\\displaystyle\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\theta\_\{c,h\}^\{3\}\]=\{\}M3\+3​𝔼​\[x2​δ¯h\]\+3​𝔼​\[x​1N​∑cδc,h2\]\\displaystyle M\_\{3\}\+3\\mathbb\{E\}\[x^\{2\}\\bar\{\\delta\}\_\{h\}\]\+3\\mathbb\{E\}\\left\[x\\frac\{1\}\{N\}\\sum\_\{c\}\\delta\_\{c,h\}^\{2\}\\right\]\(89\)\+𝔼⁡\[1N​∑cδc,h3\]\.\\displaystyle\+\\mathbb\{E\}\\left\[\\frac\{1\}\{N\}\\sum\_\{c\}\\delta\_\{c,h\}^\{3\}\\right\]\.The first term isOH​\(rγ,N\)O\_\{H\}\(r\_\{\\gamma,N\}\)by[LemmaE\.5](https://arxiv.org/html/2608.26765#A5.Thmtheorem5)\. For the second term,

\|𝔼⁡\[x2​δ¯h\]\|≤M4​𝔼​\[δ¯h2\]≤CH​rγ,N;\|\\mathbb\{E\}\[x^\{2\}\\bar\{\\delta\}\_\{h\}\]\|\\leq\\sqrt\{M\_\{4\}\\mathbb\{E\}\[\\bar\{\\delta\}\_\{h\}^\{2\}\]\}\\leq C\_\{H\}r\_\{\\gamma,N\};indeedvγ,N=γ2​\(N−2\+γ\)v\_\{\\gamma,N\}=\\gamma^\{2\}\(N^\{\-2\}\+\\gamma\)andsγ,N\+N−1≤C⁡\(N−1\+γ\)s\_\{\\gamma,N\}\+N^\{\-1\}\\leq C\(N^\{\-1\}\+\\gamma\)forγ≤1\\gamma\\leq 1\. For the third term, Jensen and Cauchy–Schwarz give

\|𝔼⁡\[x​1N​∑cδc,h2\]\|≤M2​1N​∑c𝔼⁡\[δc,h4\]≤CH​γ2​sγ,N≤CH​rγ,N\.\\left\|\\mathbb\{E\}\\left\[x\\frac\{1\}\{N\}\\sum\_\{c\}\\delta\_\{c,h\}^\{2\}\\right\]\\right\|\\leq\\sqrt\{M\_\{2\}\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\delta\_\{c,h\}^\{4\}\]\}\\leq C\_\{H\}\\gamma^\{2\}\\sqrt\{s\_\{\\gamma,N\}\}\\leq C\_\{H\}r\_\{\\gamma,N\}\.Finally,

\|𝔼⁡\[1N​∑cδc,h3\]\|≤\(1N​∑c𝔼​δc,h2\)1/2​\(1N​∑c𝔼​δc,h4\)1/2≤CH​γ3≤CH​rγ,N\.\\left\|\\mathbb\{E\}\\left\[\\frac\{1\}\{N\}\\sum\_\{c\}\\delta\_\{c,h\}^\{3\}\\right\]\\right\|\\leq\\left\(\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\\delta\_\{c,h\}^\{2\}\\right\)^\{1/2\}\\left\(\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\\delta\_\{c,h\}^\{4\}\\right\)^\{1/2\}\\leq C\_\{H\}\\gamma^\{3\}\\leq C\_\{H\}r\_\{\\gamma,N\}\.This proves Equation \([87](https://arxiv.org/html/2608.26765#A7.E87)\)\. Notice that only the*signed client average*is controlled; no bound of the stronger formN−1​∑c\|𝔼⁡\[θc,h3\]\|N^\{\-1\}\\sum\_\{c\}\|\\mathbb\{E\}\[\\theta\_\{c,h\}^\{3\}\]\|is claimed\.

For the fourth moment,

\|x\+δ\|4≤8​\(\|x\|4\+\|δ\|4\),\|x\+\\delta\|^\{4\}\\leq 8\(\|x\|^\{4\}\+\|\\delta\|^\{4\}\),so

1N​∑c𝔼⁡\[θc,h4\]≤8​M4\+8​1N​∑c𝔼⁡\[δc,h4\]≤CH​\(vγ,N\+γ4\)≤CH​vγ,N\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\theta\_\{c,h\}^\{4\}\]\\leq 8M\_\{4\}\+8\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\delta\_\{c,h\}^\{4\}\]\\leq C\_\{H\}\(v\_\{\\gamma,N\}\+\\gamma^\{4\}\)\\leq C\_\{H\}v\_\{\\gamma,N\}\.∎

### G\.1Gradient Taylor remainder

###### Lemma G\.4\(GeneralC5C^\{5\}Taylor control\)\.

For

a=f′′​\(0\),τ=f′′′​\(0\),κ=f′′′′​\(0\),a=f^\{\\prime\\prime\}\(0\),\\qquad\\tau=f^\{\\prime\\prime\\prime\}\(0\),\\qquad\\kappa=f^\{\\prime\\prime\\prime\\prime\}\(0\),the gradient admits

f′​\(y\)=a​y\+τ2​y2\+κ6​y3\+ρ5​\(y\),\|ρ5​\(y\)\|≤K524​\|y\|4\.f^\{\\prime\}\(y\)=ay\+\\frac\{\\tau\}\{2\}y^\{2\}\+\\frac\{\\kappa\}\{6\}y^\{3\}\+\\rho\_\{5\}\(y\),\\qquad\|\\rho\_\{5\}\(y\)\|\\leq\\frac\{K\_\{5\}\}\{24\}\|y\|^\{4\}\.\(90\)Moreover,

\|∑h=0H−11N​∑c=1N𝔼⁡\[f′​\(θc,h\)−a​θc,h−τ2​θc,h2\]\|≤CH​rγ,N\.\\left\|\\sum\_\{h=0\}^\{H\-1\}\\frac\{1\}\{N\}\\sum\_\{c=1\}^\{N\}\\mathbb\{E\}\\left\[f^\{\\prime\}\(\\theta\_\{c,h\}\)\-a\\theta\_\{c,h\}\-\\frac\{\\tau\}\{2\}\\theta\_\{c,h\}^\{2\}\\right\]\\right\|\\leq C\_\{H\}r\_\{\\gamma,N\}\.\(91\)

###### Proof\.

Apply Taylor’s theorem tog=f′g=f^\{\\prime\}through cubic order\. Sinceg\(4\)=f\(5\)g^\{\(4\)\}=f^\{\(5\)\},

ρ5​\(y\)=y43\!​∫01\(1−s\)3​f\(5\)​\(s​y\)​𝑑s,\\rho\_\{5\}\(y\)=\\frac\{y^\{4\}\}\{3\!\}\\int\_\{0\}^\{1\}\(1\-s\)^\{3\}f^\{\(5\)\}\(sy\)\\,ds,which gives the coefficientK5/24K\_\{5\}/24in Equation \([90](https://arxiv.org/html/2608.26765#A7.E90)\)\.

For eachhh, subtract the linear and quadratic pieces:

1N​∑c𝔼⁡\[f′​\(θc,h\)−a​θc,h−τ2​θc,h2\]=κ6​1N​∑c𝔼⁡\[θc,h3\]\+1N​∑c𝔼⁡\[ρ5​\(θc,h\)\]\.\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\\left\[f^\{\\prime\}\(\\theta\_\{c,h\}\)\-a\\theta\_\{c,h\}\-\\frac\{\\tau\}\{2\}\\theta\_\{c,h\}^\{2\}\\right\]=\\frac\{\\kappa\}\{6\}\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\theta\_\{c,h\}^\{3\}\]\+\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\rho\_\{5\}\(\\theta\_\{c,h\}\)\]\.By[LemmaG\.2](https://arxiv.org/html/2608.26765#A7.Thmtheorem2), the first term isOH​\(rγ,N\)O\_\{H\}\(r\_\{\\gamma,N\}\)and the second isOH​\(vγ,N\)O\_\{H\}\(v\_\{\\gamma,N\}\)\. Sincevγ,N≤rγ,Nv\_\{\\gamma,N\}\\leq r\_\{\\gamma,N\}forN≥2N\\geq 2, each local\-step contribution isOH​\(rγ,N\)O\_\{H\}\(r\_\{\\gamma,N\}\), and summation over fixedHHpreserves this order\. ∎

## Appendix HAssembly proof of the main expansion

This appendix proves[Theorem4\.1](https://arxiv.org/html/2608.26765#S4.Thmtheorem1)from the preceding lemmas\.

Stationarity inputs \+ exactSCAFFOLDrecursion\+ zero\-sum control stateUniform control moments and coarse global/local second momentsCoefficient\-level local second\-moment expansionMean bootstrap \+ global fourth moment\+ signed global third momentSharp global second moment: no unresolved client\-independentγ2\\gamma^\{2\}term in𝔼⁡\[x2\]\\mathbb\{E\}\[x^\{2\}\]Local higher moments \+C5C^\{5\}Taylor remainder controlStationary mean assembly and uniform joint remainderFigure 4:Proof dependency map\. Each stage either identifies a source term or excludes a competing client\-independent contribution at orderγ2\\gamma^\{2\}\. The detailed lemmas and constants are given in the surrounding appendices\.###### Proof of[Theorem4\.1](https://arxiv.org/html/2608.26765#S4.Thmtheorem1)\.

The exact stationary mean balance in Equation \([29](https://arxiv.org/html/2608.26765#A2.E29)\) and[LemmaG\.4](https://arxiv.org/html/2608.26765#A7.Thmtheorem4)give

0=a​∑h=0H−1mh\+τ2​∑h=0H−1sh\+OH​\(rγ,N\),0=a\\sum\_\{h=0\}^\{H\-1\}m\_\{h\}\+\\frac\{\\tau\}\{2\}\\sum\_\{h=0\}^\{H\-1\}s\_\{h\}\+O\_\{H\}\(r\_\{\\gamma,N\}\),\(92\)where

sh:=1N​∑c𝔼⁡\[θc,h2\]\.s\_\{h\}:=\\frac\{1\}\{N\}\\sum\_\{c\}\\mathbb\{E\}\[\\theta\_\{c,h\}^\{2\}\]\.By[LemmaE\.1](https://arxiv.org/html/2608.26765#A5.Thmtheorem1),

∑h=0H−1mh=H​bγ,N,H\+OH​\(rγ,N\)\.\\sum\_\{h=0\}^\{H\-1\}m\_\{h\}=Hb\_\{\\gamma,N,H\}\+O\_\{H\}\(r\_\{\\gamma,N\}\)\.By definition of[LemmaD\.1](https://arxiv.org/html/2608.26765#A4.Thmtheorem1),

sh=M2\+U¯h\.s\_\{h\}=M\_\{2\}\+\\bar\{U\}\_\{h\}\.Substituting into Equation \([92](https://arxiv.org/html/2608.26765#A8.E92)\),

a​H​bγ,N,H\+τ2​H​M2\+τ2​∑h=0H−1U¯h\+OH​\(rγ,N\)=0\.aHb\_\{\\gamma,N,H\}\+\\frac\{\\tau\}\{2\}HM\_\{2\}\+\\frac\{\\tau\}\{2\}\\sum\_\{h=0\}^\{H\-1\}\\bar\{U\}\_\{h\}\+O\_\{H\}\(r\_\{\\gamma,N\}\)=0\.Hence

bγ,N,H=−τ2​a​M2−τ2​a​H​∑h=0H−1U¯h\+OH​\(rγ,N\)\.b\_\{\\gamma,N,H\}=\-\\frac\{\\tau\}\{2a\}M\_\{2\}\-\\frac\{\\tau\}\{2aH\}\\sum\_\{h=0\}^\{H\-1\}\\bar\{U\}\_\{h\}\+O\_\{H\}\(r\_\{\\gamma,N\}\)\.\(93\)The sharp global second moment[LemmaF\.2](https://arxiv.org/html/2608.26765#A6.Thmtheorem2)yields

−τ2​a​M2=−τ​σ24​a2​γN\+OH​\(rγ,N\)\.\-\\frac\{\\tau\}\{2a\}M\_\{2\}=\-\\frac\{\\tau\\sigma^\{2\}\}\{4a^\{2\}\}\\frac\{\\gamma\}\{N\}\+O\_\{H\}\(r\_\{\\gamma,N\}\)\.\(94\)For the local correction,[LemmaD\.1](https://arxiv.org/html/2608.26765#A4.Thmtheorem1)gives

∑h=0H−1U¯h=γ2​σ2​\(∑h=0H−1h\+1H​∑h=0H−1h2\)\+OH​\(rγ,N\)\.\\sum\_\{h=0\}^\{H\-1\}\\bar\{U\}\_\{h\}=\\gamma^\{2\}\\sigma^\{2\}\\left\(\\sum\_\{h=0\}^\{H\-1\}h\+\\frac\{1\}\{H\}\\sum\_\{h=0\}^\{H\-1\}h^\{2\}\\right\)\+O\_\{H\}\(r\_\{\\gamma,N\}\)\.Use

∑h=0H−1h=H⁡\(H−1\)2,∑h=0H−1h2=H​\(H−1\)​\(2​H−1\)6\.\\sum\_\{h=0\}^\{H\-1\}h=\\frac\{H\(H\-1\)\}\{2\},\\qquad\\sum\_\{h=0\}^\{H\-1\}h^\{2\}=\\frac\{H\(H\-1\)\(2H\-1\)\}\{6\}\.Then

∑h=0H−1h\+1H​∑h=0H−1h2=\(H−1\)​\(5​H−1\)6,\\sum\_\{h=0\}^\{H\-1\}h\+\\frac\{1\}\{H\}\\sum\_\{h=0\}^\{H\-1\}h^\{2\}=\\frac\{\(H\-1\)\(5H\-1\)\}\{6\},and therefore

−τ2​a​H∑h=0H−1U¯h=−τ​σ212​a\(H−1\)​\(5​H−1\)Hγ2\+OH\(rγ,N\)\.\-\\frac\{\\tau\}\{2aH\}\\sum\_\{h=0\}^\{H\-1\}\\bar\{U\}\_\{h\}=\-\\frac\{\\tau\\sigma^\{2\}\}\{12a\}\\frac\{\(H\-1\)\(5H\-1\)\}\{H\}\\gamma^\{2\}\+O\_\{H\}\(r\_\{\\gamma,N\}\)\.\(95\)Combining Equations \([93](https://arxiv.org/html/2608.26765#A8.E93)\)–\([95](https://arxiv.org/html/2608.26765#A8.E95)\) gives Equation \([7](https://arxiv.org/html/2608.26765#S4.E7)\)\. The common small\-step threshold andN,γN,\\gamma\-uniform remainder constant are justified in Appendix[I](https://arxiv.org/html/2608.26765#A9)\. ∎

### H\.1Canonical source\-wise decomposition

###### Corollary H\.1\(Direct local\-noise and SCAFFOLD control\-second\-moment sources\)\.

TheN0​γ2N^\{0\}\\gamma^\{2\}coefficient in Equation \([7](https://arxiv.org/html/2608.26765#S4.E7)\) admits the canonical source\-wise decomposition induced by the exact local second\-moment identity in Equation \([43](https://arxiv.org/html/2608.26765#A4.E43)\):

B20local​\(H\)\\displaystyle B\_\{20\}^\{\\mathrm\{local\}\}\(H\):=−τ​σ24​a​\(H−1\),\\displaystyle:=\-\\frac\{\\tau\\sigma^\{2\}\}\{4a\}\(H\-1\),\(96\)B20ctrl​\(H\)\\displaystyle B\_\{20\}^\{\\mathrm\{ctrl\}\}\(H\):=−τ​σ212​a​\(H−1\)​\(2​H−1\)H,\\displaystyle:=\-\\frac\{\\tau\\sigma^\{2\}\}\{12a\}\\frac\{\(H\-1\)\(2H\-1\)\}\{H\},\(97\)B20SCAF​\(H\)\\displaystyle B\_\{20\}^\{\\mathrm\{SCAF\}\}\(H\):=B20local​\(H\)\+B20ctrl​\(H\)=−τ​σ212​a​\(H−1\)​\(5​H−1\)H\.\\displaystyle:=B\_\{20\}^\{\\mathrm\{local\}\}\(H\)\+B\_\{20\}^\{\\mathrm\{ctrl\}\}\(H\)=\-\\frac\{\\tau\\sigma^\{2\}\}\{12a\}\\frac\{\(H\-1\)\(5H\-1\)\}\{H\}\.\(98\)

###### Proof\.

The direct fresh\-noise source in Equation \([43](https://arxiv.org/html/2608.26765#A4.E43)\) isγ2​h​σ2\\gamma^\{2\}h\\sigma^\{2\}\. Its contribution through the factor−τ/\(2aH\)\-\\tau/\(2aH\)in Equation \([93](https://arxiv.org/html/2608.26765#A8.E93)\) is

−τ2​a​Hσ2∑h=0H−1h=−τ​σ24​a\(H−1\)\.\-\\frac\{\\tau\}\{2aH\}\\sigma^\{2\}\\sum\_\{h=0\}^\{H\-1\}h=\-\\frac\{\\tau\\sigma^\{2\}\}\{4a\}\(H\-1\)\.TheN0N^\{0\}part of the control source is

γ2​h2​σ2H\.\\gamma^\{2\}h^\{2\}\\frac\{\\sigma^\{2\}\}\{H\}\.Therefore, its contribution is

−τ2​a​Hσ2H∑h=0H−1h2=−τ​σ212​a\(H−1\)​\(2​H−1\)H\.\-\\frac\{\\tau\}\{2aH\}\\frac\{\\sigma^\{2\}\}\{H\}\\sum\_\{h=0\}^\{H\-1\}h^\{2\}=\-\\frac\{\\tau\\sigma^\{2\}\}\{12a\}\\frac\{\(H\-1\)\(2H\-1\)\}\{H\}\.Adding the two gives Equation \([98](https://arxiv.org/html/2608.26765#A8.E98)\)\. ∎

## Appendix IUniform constants and joint\-limit semantics

### I\.1Uniformity inNNandγ\\gammafor fixedHH

###### Proposition I\.1\(Common small\-step threshold and uniform remainder constant\)\.

For fixedHH, all lemmas used in[Theorem4\.1](https://arxiv.org/html/2608.26765#S4.Thmtheorem1)admit a common thresholdγ0,H\>0\\gamma\_\{0,H\}\>0and constants that are independent ofNNandγ\\gamma\. Consequently, the remainder in Equation \([8](https://arxiv.org/html/2608.26765#S4.E8)\) is uniform over

N≥2,0<γ≤γ0,H\.N\\geq 2,\\qquad 0<\\gamma\\leq\\gamma\_\{0,H\}\.

###### Proof\.

Each preceding argument requires only finitely many small\-step restrictions of the formγ≤cH\\gamma\\leq c\_\{H\}, wherecH\>0c\_\{H\}\>0depends on fixedHHand fixed problem/noise parameters\. These include the stationarity and moment restrictions inherited from[Mangold et al\. \[12\]](https://arxiv.org/html/2608.26765#bib.bib2),γ≤1\\gamma\\leq 1, conditions such asa​γ≤1a\\gamma\\leq 1orγ​H​L≤1\\gamma HL\\leq 1, finite\-step stability restrictions, and the fixed\-ϵ\\epsilonabsorption conditions in[LemmaE\.3](https://arxiv.org/html/2608.26765#A5.Thmtheorem3)\. Define

γ0,H:=minj⁡γ0,H\(j\)\.\\gamma\_\{0,H\}:=\\min\_\{j\}\\gamma\_\{0,H\}^\{\(j\)\}\.Because the collection is finite and everyγ0,H\(j\)\>0\\gamma\_\{0,H\}^\{\(j\)\}\>0, the common threshold is positive and independent ofNN\.

The constant bookkeeping is uniform inNNfor the following reasons\.

- •The imported coarse iterate moments and the restricted control expansion[LemmaC\.3](https://arxiv.org/html/2608.26765#A3.Thmtheorem3)have constants independent ofNN\.
- •Client averaging introduces factorsN−1N^\{\-1\}orN−2N^\{\-2\}explicitly; Jensen and Cauchy–Schwarz do not introduce hidden powers ofNN\.
- •Mixed fractional scales are reduced using algebraic inequalities such as γ5/2N≤12​\(γ2N\+γ3\),\\frac\{\\gamma^\{5/2\}\}\{\\sqrt\{N\}\}\\leq\\frac\{1\}\{2\}\\left\(\\frac\{\\gamma^\{2\}\}\{N\}\+\\gamma^\{3\}\\right\),with constants independent of the relative rate ofNNandγ\\gamma\.
- •In the coordinate\-replacement argument in Equation \([85](https://arxiv.org/html/2608.26765#A6.E85)\), theN​HNHfresh coordinates are multiplied by one coefficient of orderγ/N\\gamma/Nand one path\-sensitivity factor of orderγ2/N\\gamma^\{2\}/N, leavingCH​γ3/NC\_\{H\}\\gamma^\{3\}/N, not anN0N^\{0\}term\.
- •Restoring factors satisfy lower bounds of ordera​γa\\gamma, anda≥μ\>0a\\geq\\mu\>0is fixed; the resulting division byγ\\gammais matched by an extra factor ofγ\\gammain the numerator of each moment estimate\.
- •The fourth\-moment absorption fixesϵ\>0\\epsilon\>0after theHH\-dependent constants are known;ϵ\\epsilonis never chosen as a function ofNNorγ\\gamma\.

Thus the final constantCHC\_\{H\}may depend onH,μ,L,K3,K4,K5,Bε,σ2H,\\mu,L,K\_\{3\},K\_\{4\},K\_\{5\},B\_\{\\varepsilon\},\\sigma^\{2\}and related fixed quantities, but not onNNorγ\\gamma\. ∎

### I\.2Uniform joint\(γ,1/N\)\(\\gamma,1/N\)interpretation

###### Corollary I\.4\(Joint\-limit characterization of theN0​γ2N^\{0\}\\gamma^\{2\}coefficient\)\.

For any sequencesγk→0\\gamma\_\{k\}\\to 0and integersNk→∞N\_\{k\}\\to\\infty, with fixedHHand no relative\-rate condition,

limk→∞bγk,Nk,H\+τ​σ24​a2​γkNkγk2=−τ​σ212​a​\(H−1\)​\(5​H−1\)H\.\\lim\_\{k\\to\\infty\}\\frac\{b\_\{\\gamma\_\{k\},N\_\{k\},H\}\+\\dfrac\{\\tau\\sigma^\{2\}\}\{4a^\{2\}\}\\dfrac\{\\gamma\_\{k\}\}\{N\_\{k\}\}\}\{\\gamma\_\{k\}^\{2\}\}=\-\\frac\{\\tau\\sigma^\{2\}\}\{12a\}\\frac\{\(H\-1\)\(5H\-1\)\}\{H\}\.\(99\)

###### Proof\.

Divide Equation \([8](https://arxiv.org/html/2608.26765#S4.E8)\) byγ2\\gamma^\{2\}:

\|Rγ,N,Hγ2\|≤CH​\(1N\+γ\)\.\\left\|\\frac\{R\_\{\\gamma,N,H\}\}\{\\gamma^\{2\}\}\\right\|\\leq C\_\{H\}\\left\(\\frac\{1\}\{N\}\+\\gamma\\right\)\.Both terms converge to zero along any such sequence, independently of their relative rate\. ∎

### I\.3Precise claim boundary

The strongest claim supported by the present proof is:

> For fixedHH, we derive a stationary\-bias expansion for one\-dimensional homogeneous stochastic SCAFFOLD that is uniform jointly in\(γ,1/N\)\(\\gamma,1/N\)\. After removing the knownO⁡\(γ/N\)O\(\\gamma/N\)layer, the client\-number\-independentN0​γ2N^\{0\}\\gamma^\{2\}coefficient is −f′′′​\(x⋆\)​σ212​f′′​\(x⋆\)​\(H−1\)​\(5​H−1\)H\.\-\\frac\{f^\{\\prime\\prime\\prime\}\(x^\{\\star\}\)\\sigma^\{2\}\}\{12f^\{\\prime\\prime\}\(x^\{\\star\}\)\}\\frac\{\(H\-1\)\(5H\-1\)\}\{H\}\.The coefficient admits a canonical source\-wise decomposition into a direct local\-gradient\-noise contribution and an additional contribution induced by the persistent SCAFFOLD control second moment\.

The following claims are*not*established here:

- •a complete second\-order expansion at fixed finiteNN;
- •discovery of the fact that stochastic SCAFFOLD is biased;
- •a claim that all higher\-order bias is caused by control variates;
- •uniformity asH→∞H\\to\\infty;
- •heterogeneous\-client, multidimensional, or state\-dependent\-noise extensions;
- •equality ofB20localB\_\{20\}^\{\\mathrm\{local\}\}with a separately proved general\-HHFedAvg coefficient\.

## Appendix JNumerical protocol and supplementary checks

This appendix records the simulation protocol, the initial coefficient grid, and the negative controls summarized in[Section7](https://arxiv.org/html/2608.26765#S7)\.

### J\.1Full simulation protocol

For each setting, we estimate the stationary mean in two ways: by the direct sample mean and by the exact mean\-balance identity used in the proof\. Unless an entry is explicitly labeled “direct,” the reportedb^γ,N,H\\hat\{b\}\_\{\\gamma,N,H\}and normalized statisticCγ,N,HC\_\{\\gamma,N,H\}use the mean\-balance estimator; the direct sample mean is retained as a diagnostic\. Let

mγ,H=⌈1μ​H​γ⌉\.m\_\{\\gamma,H\}=\\left\\lceil\\frac\{1\}\{\\mu H\\gamma\}\\right\\rceil\.We discard100​mγ,H100m\_\{\\gamma,H\}communication rounds as burn\-in and use batches of length20​mγ,H20m\_\{\\gamma,H\}, with at least 100 batches per chain\. Each setting uses eight independent chains\. Pointwise uncertainty is reported using 95% Student\-ttintervals across the eight chain estimates\.

For extrapolated intercepts, we use a nonparametric bootstrap at the independent\-chain level\. Within each parameter setting, the eight chain estimates are resampled with replacement, the setting mean is recomputed, and the same unweighted linear regression is refit\. The reported 95% bootstrap interval is the interval between the empirical 2\.5% and 97\.5% quantiles of the refitted intercepts\.

For every reported setting, we monitor finiteness, preservation of the zero\-sum control invariant, agreement between the direct and mean\-balance estimators, and stability between the first and second halves of the post\-burn\-in sample\. These checks diagnose simulation failures or insufficient equilibration; they are not tests of the theorem\.

### J\.2Initial joint path and the recordedH=4H=4discrepancy

The initial study followedN=1/γN=1/\\gammaat

γ∈\{1/12,1/16,1/24,1/32,1/48\},H∈\{2,4\}\.\\gamma\\in\\\{1/12,1/16,1/24,1/32,1/48\\\},\\qquad H\\in\\\{2,4\\\}\.ForH=2H=2, a linear extrapolation gives intercept−0\.187214\-0\.187214with 95% bootstrap interval\[−0\.189250,−0\.185175\]\[\-0\.189250,\-0\.185175\], consistent withB20​\(2\)=−0\.1875B\_\{20\}\(2\)=\-0\.1875\. ForH=4H=4, the same coarse\-grid extrapolation gives−0\.588633\-0\.588633with interval\[−0\.590776,−0\.586514\]\[\-0\.590776,\-0\.586514\], which does not containB20​\(4\)=−0\.59375B\_\{20\}\(4\)=\-0\.59375\. Under the criterion specified for the initial study, this outcome was recorded as a contradiction candidate\. The theorem permits anOH​\(γ\)O\_\{H\}\(\\gamma\)\-normalized remainder and does not imply exact linearity over this finite grid, so we specified a smaller\-step experiment with fresh random seeds before inspecting its outcomes\.

![Refer to caption](https://arxiv.org/html/2608.26765v1/figures/coefficient_convergence_H2.png)\(a\)H=2H=2\.
![Refer to caption](https://arxiv.org/html/2608.26765v1/figures/coefficient_convergence_H4.png)\(b\)H=4H=4\.

Figure 5:Initial coarse\-grid coefficient study\. TheH=2H=2extrapolation is consistent with the target coefficient, whereas theH=4H=4extrapolation is offset on this finite\-step grid\.The follow\-up used

\(γ,N\)∈\{\(1/64,64\),\(1/96,96\),\(1/128,128\),\(1/192,192\)\}\.\(\\gamma,N\)\\in\\\{\(1/64,64\),\(1/96,96\),\(1/128,128\),\(1/192,192\)\\\}\.WritingE⁡\(γ\):=Cγ,1/γ,4−B20​\(4\)E\(\\gamma\):=C\_\{\\gamma,1/\\gamma,4\}\-B\_\{20\}\(4\), the observed errors decrease from0\.021590\.02159to0\.007870\.00787, whileE⁡\(γ\)/γE\(\\gamma\)/\\gammaremains of constant order\. The follow\-up met its specified validity, directional convergence, smallest\-step proximity, andO⁡\(γ\)O\(\\gamma\)\-compatibility criteria\.

![Refer to caption](https://arxiv.org/html/2608.26765v1/figures/coefficient_vs_step_size.png)\(a\)Normalized residual\.
![Refer to caption](https://arxiv.org/html/2608.26765v1/figures/scaled_error_vs_step_size.png)\(b\)Error divided byγ\\gamma\.

Figure 6:Smaller\-stepH=4H=4study\. The normalized residual moves toward the predicted coefficient, while the scaled error remains of constant order, consistent with theOH​\(γ\)O\_\{H\}\(\\gamma\)remainder permitted by Equation \([20](https://arxiv.org/html/2608.26765#S7.E20)\) alongN=1/γN=1/\\gamma\.Table 3:Fresh smaller\-step results forH=4H=4\. Intervals are 95% Student\-ttintervals across eight independent chains\.
### J\.3Negative controls and theH=1H=1boundary

With deterministic gradients \(σ2=0\\sigma^\{2\}=0\), both displayed stochastic coefficients vanish\. For the quadratic objectivef′​\(x\)=xf^\{\\prime\}\(x\)=x, one hasf′′′​\(x⋆\)=0f^\{\\prime\\prime\\prime\}\(x^\{\\star\}\)=0; the mean\-balance estimator is then structurally zero, so we report the direct sample mean\. Finally,H=1H=1is a boundary check outside the theorem’s statedH≥2H\\geq 2domain\. The statistic in that case is the normalized residual after removing the leadingγ/N\\gamma/Nterm\.

Table 4:Negative controls and theH=1H=1boundary check\.The deterministic and quadratic intervals contain zero at both tested step sizes\. ForH=1H=1, the interval is slightly separated from zero atγ=1/24\\gamma=1/24but contains zero atγ=1/48\\gamma=1/48, with a smaller point estimate at the smaller step size\.

Similar Articles