Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning

arXiv cs.LG Papers

Summary

This paper presents a layer-wise information-theoretic framework for replay-based continual learning, decomposing the generalization gap into replay-induced representation drift and optimization-dependence terms, with refinements via Wasserstein relaxation and SGLD instantiation.

arXiv:2608.11690v1 Announce Type: new Abstract: Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies for catastrophic forgetting. Yet its generalization behavior is shaped by two coupled effects that existing analyses fold into a single hypothesis-level quantity: finite memory replaces each past distribution with an empirical proxy, and repeated reuse couples the buffer, the current data, and the final hypothesis through a shared optimization trajectory. We develop a layer-wise information-theoretic framework that separates these effects at every depth. Our main result decomposes the expected generalization gap into a replay-induced representation drift and an optimization-dependence term, the latter further resolved into stability, plasticity, interaction, and residual-coupling components. Two refinements make the framework operational. A Wasserstein relaxation of the drift term, valid under support mismatch, yields a depth-dependent drift--sensitivity trade-off whose minimizer identifies which interior layer to stabilize. An SGLD instantiation of the optimization term reduces it to a trajectory-level log-determinant budget, exposing a curvature-aware gradient-alignment statistic that serves as an online diagnostic of task-wise forgetting. Controlled and benchmark experiments confirm the predicted memory scaling, the interior funnel, and the alignment signal's link to forgetting.
Original Article
View Cached Full Text

Cached at: 08/13/26, 03:38 PM

# Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning
Source: [https://arxiv.org/html/2608.11690](https://arxiv.org/html/2608.11690)
Zhongbo ZhangWen WenYong\-Jin LiuThanks:This work was supported by National Natural Science Foundation of China under Grant 62576268, the Fundamental Research Funds for the Central Universities and the Key Research and Development Project in Shaanxi Province No\. 2024PT\-ZCK\-89\. \(Corresponding author: Tieliang Gong\.\)Thanks:Tieliang Gong, Zhongbo Zhang and Wen Wen are with the School of Computer Science and Technology, Xi’an Jiaotong University, China \(e\-mail: adidasgtl@gmail\.com; shuaishuai737@gmail\.com; wen190329@gmail\.com\)\.Thanks:Yong\-Jin Liu is with Department of Computer Science and Technology, Tsinghua University, China \(e\-mail: liuyongjin@tsinghua\.edu\.cn\)

###### Abstract

Continual learning must absorb new tasks without erasing old ones, and replay—mixing a small buffer of past examples into current training—is among the most effective remedies for catastrophic forgetting\. Yet its generalization behavior is shaped by two coupled effects that existing analyses fold into a single hypothesis\-level quantity: finite memory replaces each past distribution with an empirical proxy, and repeated reuse couples the buffer, the current data, and the final hypothesis through a shared optimization trajectory\. We develop a layer\-wise information\-theoretic framework that separates these effects at every depth\. Our main result decomposes the expected generalization gap into a replay\-induced representation drift and an optimization\-dependence term, the latter further resolved into stability, plasticity, interaction, and residual\-coupling components\. Two refinements make the framework operational\. A Wasserstein relaxation of the drift term, valid under support mismatch, yields a depth\-dependent drift–sensitivity trade\-off whose minimizer identifies which interior layer to stabilize\. An SGLD instantiation of the optimization term reduces it to a trajectory\-level log\-determinant budget, exposing a curvature\-aware gradient\-alignment statistic that serves as an online diagnostic of task\-wise forgetting\. Controlled and benchmark experiments confirm the predicted memory scaling, the interior funnel, and the alignment signal’s link to forgetting\.

###### Index Terms:

Continual learning, experience replay, information theory, Wasserstein distance, stochastic gradient Langevin dynamics\.

## IIntroduction

Continual learning \(CL\) studies how a model can incrementally learn a sequence of tasks while retaining knowledge acquired from previous ones\[[21](https://arxiv.org/html/2608.11690#bib.bib12),[31](https://arxiv.org/html/2608.11690#bib.bib1),[8](https://arxiv.org/html/2608.11690#bib.bib2)\]\. A central obstacle in CL is catastrophic forgetting: when the model adapts to a new task, its performance on earlier tasks may deteriorate significantly\[[29](https://arxiv.org/html/2608.11690#bib.bib22),[14](https://arxiv.org/html/2608.11690#bib.bib3)\]\. In recent years, extensive efforts have been made to mitigate catastrophic forgetting\[[20](https://arxiv.org/html/2608.11690#bib.bib5),[24](https://arxiv.org/html/2608.11690#bib.bib6)\]\. Among them, replay\-based methods\[[36](https://arxiv.org/html/2608.11690#bib.bib13),[33](https://arxiv.org/html/2608.11690#bib.bib14),[16](https://arxiv.org/html/2608.11690#bib.bib7)\]have emerged as one of the most successful practical mechanisms, where the learner stores a small buffer of past examples and mixes them with the current task during training\.

Despite this success, the generalization behavior of replay\-based CL remains insufficiently understood since replay changes the statistical object being optimized\. The learner never minimizes risk against the past\-task populations\. It minimizes risk against finite proxies of them, and does so repeatedly as the model continues to evolve\. This substitution acts in two channels\. First, replacing a population by a finite stored subset produces a proxy bias, say, the empirical objective on the buffer is a biased stand\-in for the population objective, and the bias does not disappear merely because the data are i\.i\.d\. Second, reusing the same stored samples across many updates produces a dependence coupling among the old samples, the current samples, and the final hypothesis, since all three are tied together through a shared optimization trajectory\. Note that neither channel presupposes that the model learns a representation, and both are already present when replay is performed directly on raw inputs\.

What a deep network adds is not a third channel but a change in how these two channels behave with depth, and it does so asymmetrically\. Because the buffer influences later tasks only through a learned feature map, and because that map is fitted to the particular stored exemplars, the proxy bias is no longer fixed: it is generated and reshaped at every layer, and a discrepancy that is negligible in input space can be amplified by the sub\-network above it\. The dependence coupling, by contrast, is set by the algorithm’s reuse of data and survives even when no representation is learned\. The learned representation therefore turns the proxy bias into a depth\-dependent object while leaving the coupling channel essentially intact\. As the boundary cases of our bound make precise, the proxy\-bias channel vanishes at the input layer and grows with depth, whereas the coupling channel survives even there\.

Existing theory does not separate these channels\. Analyses based on replay losses, trade\-off parameters, or domain\-adaptation arguments\[[35](https://arxiv.org/html/2608.11690#bib.bib8),[40](https://arxiv.org/html/2608.11690#bib.bib9)\]blend the intrinsic difficulty of learning from the full stream with the approximation cost of compressing it into a buffer, so the finite capacity loss and the reuse\-induced dependence never surface as distinct quantities\. Classic complexity bounds fare no better: their capacity terms grow far faster than the sample size, leaving them numerically vacuous in precisely the overparameterized regime of interest\. Stability and convergence analyses under smoothness or Lipschitz assumptions\[[23](https://arxiv.org/html/2608.11690#bib.bib10),[22](https://arxiv.org/html/2608.11690#bib.bib11)\]take the opposite tack—they track optimization dynamics closely but say little about how buffer size, sequence length, depth, and algorithmic dependence jointly shape the gap\. The information\-theoretic framework comes closest: it relates the learned parameters to the data through mutual information and its refinements\[[50](https://arxiv.org/html/2608.11690#bib.bib18),[44](https://arxiv.org/html/2608.11690#bib.bib19),[17](https://arxiv.org/html/2608.11690#bib.bib38)\]and recent work carries this analysis across the layers of a deep network\[[19](https://arxiv.org/html/2608.11690#bib.bib41)\]\. This is our starting point, but replay changes the objects the framework must measure, so it cannot be applied off the shelf\. The relevant quantities are no longer parameter–data dependencies but buffer\-induced representation distributions, and the coupling created by reuse and by sampling without replacement is intrinsic to the problem rather than a nuisance term: it must be built into the analysis from the outset\. What makes a layer\-wise treatment unavoidable is the asymmetry between the two channels\. Because the finite\-memory channel is reshaped at every layer while the coupling channel is not, a black\-box, hypothesis\-level analysis averages over exactly the depth structure the buffer induces, collapsing the very quantity we set out to resolve—the depth at which the buffer’s finite\-sample bias is least amplified by the network above it\. We therefore introduce two layer\-wise objects: per\-task replay centroids, which make the finite\-memory discrepancy well\-defined at every depth, and a decoupled reference, which isolates the dependence due to reuse\. Together these let finite\-memory bias and optimization coupling sit in a single decomposition—something hypothesis\-level information\-theoretic bounds cannot express\.

There are two difficulties in carrying this out rigorously\. The first is circularity: the feature map at a given depth is learned online and shifts with whichever examples enter the buffer, so the map we use to measure replay discrepancy is shaped by the very buffer we are measuring\. The second is accumulation: the dependence entering a new task is not generated fresh but inherited from earlier tasks and built up along the training trajectory\. Both must be handled jointly and at an arbitrary depth, not just at the input or output\. This is what the paper’s three constructions provide: a replay centroid that makes the finite\-memory discrepancy well\-defined at every layer and separates it from the dependence term; a Wasserstein relaxation of that discrepancy for the regime where its information\-theoretic form is ill\-posed under support mismatch; and a trajectory\-level decomposition that, for a concrete noisy optimizer, splits the accumulated dependence into per\-step contributions\.

At a high level, our analysis separates replay generalization into two coupled effects: a replay\-induced drift term, arising from the finite\-memory approximation of past task distributions, and an optimization\-dependence term, arising from the reuse of replayed samples through shared downstream parameters\. The former quantifies the mismatch between the replay buffer and the corresponding past\-task population in representation space, while the latter captures how strongly the learned suffix parameters encode information about both old and current task representations\. This decomposition cleanly distinguishes the statistical cost of memory compression from the dependence created by repeated optimization\. We then further refine the drift branch with a Wasserstein relaxation when the KL\-based form becomes uninformative under support mismatch, and instantiate the dependence branch for stochastic gradient Langevin dynamics \(SGLD\)\. The resulting quantities are intended as analytic diagnostics and control signals for replay\-based training, rather than as a replacement for existing continual learning algorithms\.

In summary, our contributions are as follows:

- •We develop a layer\-wise information\-theoretic bound for replay\-based continual learning that explicitly separates replay\-induced drift from optimization dependence\. The bound further clarifies how old\-task dependence, current\-task dependence, their interaction, and residual coupling jointly shape the expected generalization gap\.
- •When the KL\-based drift term becomes vacuous under empirical\-population support mismatch, we derive a Wasserstein relaxation that yields a depth\-dependent drift–sensitivity trade\-off and identifies an interior generalization\-funnel layer, indicating which layer to stabilize\.
- •We instantiate the optimization\-dependence branch with SGLD and obtain a trajectory\-level log\-determinant budget\. This leads to sensitivity\-aware gradient\-alignment diagnostics and a local replay–current mixing rule\.
- •We validate the theory through controlled and benchmark experiments, confirming the predicted finite\-memory scaling, funnel behavior, and alignment patterns in replay dynamics\.

## IIRelated Work

### II\-AReplay\-based continual learning

Catastrophic forgetting under sequential training has long been recognized as a central obstacle for neural networks\[[29](https://arxiv.org/html/2608.11690#bib.bib22),[13](https://arxiv.org/html/2608.11690#bib.bib23)\], and modern CL methods are commonly grouped into replay\-based, regularization\-based, and architectural/parameter\-isolation approaches\[[46](https://arxiv.org/html/2608.11690#bib.bib25),[31](https://arxiv.org/html/2608.11690#bib.bib1),[48](https://arxiv.org/html/2608.11690#bib.bib15)\]\. Regularization methods constrain updates through parameter\-importance penalties such as Fisher information and synapse\-importance surrogates\[[21](https://arxiv.org/html/2608.11690#bib.bib12),[51](https://arxiv.org/html/2608.11690#bib.bib26),[1](https://arxiv.org/html/2608.11690#bib.bib27)\], while architectural strategies expand or partition parameters across tasks\[[37](https://arxiv.org/html/2608.11690#bib.bib28),[28](https://arxiv.org/html/2608.11690#bib.bib29)\]\. Replay remains a strong and competitive option because it directly counteracts forgetting by revisiting stored examples, through exemplar rehearsal\[[33](https://arxiv.org/html/2608.11690#bib.bib14)\], generic experience replay\[[36](https://arxiv.org/html/2608.11690#bib.bib13)\], and generative replay\[[41](https://arxiv.org/html/2608.11690#bib.bib35)\]\. Modern replay systems differ mainly in how they allocate memory and select exemplars: a typical choice is reservoir sampling\[[7](https://arxiv.org/html/2608.11690#bib.bib30)\], while more informed strategies prioritize examples that maximize future utility or reduce interference\[[45](https://arxiv.org/html/2608.11690#bib.bib36)\], and further work corrects the representation drift and class imbalance that erode replay in online CL\[[4](https://arxiv.org/html/2608.11690#bib.bib37)\]\. Replay also connects to constrained\-optimization views that explicitly control gradient interference\[[2](https://arxiv.org/html/2608.11690#bib.bib31)\], including GEM\[[27](https://arxiv.org/html/2608.11690#bib.bib32)\], its averaged variant A\-GEM\[[5](https://arxiv.org/html/2608.11690#bib.bib33)\], and orthogonal projection\[[38](https://arxiv.org/html/2608.11690#bib.bib34)\]\. These methods align with the view taken here, that replay replaces past populations with finite proxy distributions while evolving representations cause the buffer to shift within feature space\. Our goal is different from any single update rule: we begin from a layer\-wise decomposition of replay generalization that separates distributional drift from optimization coupling, and that explains when replay acts as a stable rehearsal mechanism and when it becomes a biased or strongly coupled proxy for the past\.

### II\-BInformation\-theoretic analysis

Recent theoretical work studies forgetting, task order, and generalization in sequential learning through statistical or optimization lenses\. Closest to us are information\-theoretic and statistical analyses of forgetting\[[25](https://arxiv.org/html/2608.11690#bib.bib16),[9](https://arxiv.org/html/2608.11690#bib.bib17)\], which operate in linear or overparameterized regimes and characterize how task similarity and ordering affect generalization\. For the replay setting in particular, information\-theoretic bounds have been derived from the dependence between the hypothesis and the memory buffer\[[49](https://arxiv.org/html/2608.11690#bib.bib4)\], separating an ideal full\-data risk from a memory\-compression cost\. Because these bounds are read at the level of the whole hypothesis, their finite\-memory term stays a raw\-sample compression cost that does not resolve the representation\-level drift or the depth at which the gap arises\. Our contribution is different in kind: we retain learned hierarchical representations and expose layer\-dependent quantities \(𝒦\(l\)\\mathcal\{K\}^\{\(l\)\},𝒮\(l\),𝒫\(l\),ℛ\(l\)\\mathcal\{S\}^\{\(l\)\},\\mathcal\{P\}^\{\(l\)\},\\mathcal\{R\}^\{\(l\)\}\) that such hypothesis\- or input\-layer analysis cannot see\. We do not claim a tighter constant; we claim a finer\-grained decomposition\. More broadly, the information\-theoretic framework relates learned parameters to data via mutual information\[[50](https://arxiv.org/html/2608.11690#bib.bib18)\]and its conditional refinements\[[44](https://arxiv.org/html/2608.11690#bib.bib19),[17](https://arxiv.org/html/2608.11690#bib.bib38)\], and has been specialized to the trajectory of noisy iterative algorithms such as SGLD\[[30](https://arxiv.org/html/2608.11690#bib.bib39),[11](https://arxiv.org/html/2608.11690#bib.bib40)\]; the dependence of generalization on network depth has also been studied directly\[[19](https://arxiv.org/html/2608.11690#bib.bib41),[42](https://arxiv.org/html/2608.11690#bib.bib42)\]\. Our main theorem combines these perspectives for replay: it keeps the information\-theoretic dependence structure, but introduces replay centroids and layer\-wise old/new task representations so that finite replay bias and optimization coupling appear in the same decomposition\.

### II\-CGeometry and gradient interaction

Wasserstein and optimal\-transport arguments provide finite notions of mismatch when KL\-type divergences are unstable under support shift, both as a general geometric tool\[[32](https://arxiv.org/html/2608.11690#bib.bib46)\]and in domain\-adaptation and generalization bounds that control risk through a transport discrepancy\[[34](https://arxiv.org/html/2608.11690#bib.bib43),[26](https://arxiv.org/html/2608.11690#bib.bib45),[47](https://arxiv.org/html/2608.11690#bib.bib44)\]; these works, however, remain confined to single\-task learning\. We use this geometry only to relax the drift branch of our main decomposition\. Conversely, our SGLD analysis targets the optimization branch and yields gradient\-moment diagnostics\. This distinguishes our alignment quantities from Euclidean projection rules in GEM/OGD\-style methods: here alignment is derived as a bound\-driven, sensitivity\-aware diagnostic of optimization coupling rather than introduced as a standalone projection heuristic\.

## IIIProblem Setup and Layer\-wise Replay Distributions

Notations\.Throughout this paper, we denote random variables by uppercase letters \(e\.g\.,XX\), their realizations by lowercase letters \(e\.g\.,xx\), and their domains by calligraphic letters \(e\.g\.,𝒳\\mathcal\{X\}\)\. The distribution ofXXis denotedPXP\_\{X\}\. Expectation, variance, and covariance are written𝔼⁡\[X\]\\mathbb\{E\}\[X\],Var⁡\(X\)\\mathrm\{Var\}\(X\), andCov⁡\(X\)\\mathrm\{Cov\}\(X\), respectively\. For conditional distributions, we writePX\|Y=yP\_\{X\|Y=y\}\. We use the standard information\-theoretic quantities: entropyH⁡\(X\)H\(X\), mutual informationI⁡\(X,Y\)I\(X;Y\), Kullback\-Leibler \(KL\) divergenceDKL\(P∥Q\)D\_\{\\mathrm\{KL\}\}\(P\\\|Q\), and thepp\-Wasserstein distance𝒲p\\mathcal\{W\}\_\{p\}between probability measures\. For vectors,‖v‖\\\|v\\\|denotes the Euclidean \(ℓ2\\ell\_\{2\}\) norm, and for matrices,‖A‖op\\\|A\\\|\_\{\\mathrm\{op\}\}denotes the operator norm\. We write\[T\]≜\{1,…,T\}\[T\]\\triangleq\\\{1,\\dots,T\\\}\.

### III\-AReplay\-Based Continual Learning

We consider supervised classification with input space𝒳\\mathcal\{X\}, label space𝒴\\mathcal\{Y\}and joint space𝒵=𝒳×𝒴\\mathcal\{Z\}=\\mathcal\{X\}\\times\\mathcal\{Y\}\. A sequence ofTTtasksD1:T=\{D1,…,DT\}D^\{1:T\}=\\\{D^\{1\},\\ldots,D^\{T\}\\\}arrives incrementally\. For eacht∈\[T\]t\\in\[T\], taskttprovides a datasetDt=\{Zjt\}j=1n∈𝒵nD^\{t\}=\\\{Z\_\{j\}^\{t\}\\\}\_\{j=1\}^\{n\}\\in\\mathcal\{Z\}^\{n\}, consisting ofnni\.i\.d\. samples drawn from an unknown task distribution𝒟t\\mathcal\{D\}\_\{t\}on𝒵\\mathcal\{Z\}\. We do not impose restrictions on the relationships among\{𝒟t\}t=1T\\\{\\mathcal\{D\}\_\{t\}\\\}\_\{t=1\}^\{T\}, permitting arbitrary distributional heterogeneity across tasks\. Moreover, the datasets\{Dt\}t=1T\\\{D^\{t\}\\\}\_\{t=1\}^\{T\}are assumed independent across tasks\. At the beginning of tasktt, the learner receives a fresh datasetDtD^\{t\}and a replay bufferℳ1:t−1=\{Mi\}i=1t−1\\mathcal\{M\}^\{1:t\-1\}=\\\{M^\{i\}\\\}\_\{i=1\}^\{t\-1\}holding exemplars from previous tasks\. For each past taski∈\[t−1\]i\\in\[t\-1\], the buffer stores a subsetMi=\{Zji\}j∈ℐi⊂DiM^\{i\}=\\\{Z\_\{j\}^\{i\}\\\}\_\{j\\in\\mathcal\{I\}\_\{i\}\}\\subset D^\{i\}, where the index setℐi⊂\[n\]\\mathcal\{I\}\_\{i\}\\subset\[n\]has fixed size\|ℐi\|=m\|\\mathcal\{I\}\_\{i\}\|=mand is drawn uniformly without replacement from\[n\]\[n\]\. We focus on the memory\-limited regimem≪nm\\ll n\. Importantly, a past task never recurs as a full dataset: once taskiiis completed, it can be revisited only through themmexemplars retained inMiM^\{i\}\.

Let𝒲\\mathcal\{W\}denote the hypothesis parameter space, and letℓ:𝒲×𝒵→ℝ\+\\ell:\\mathcal\{W\}\\times\\mathcal\{Z\}\\rightarrow\\mathbb\{R\}\_\{\+\}be a fixed nonnegative loss function\. After receiving tasktt, training uses the union of the current dataset and all replayed exemplars\. We define the empirical risk forW∈𝒲W\\in\\mathcal\{W\}as

ℒ^\(t\)​\(W\)=∑i=1t−11m​∑j∈ℐiℓ⁡\(W,Zji\)\+1n​∑j=1nℓ⁡\(W,Zjt\),\\hat\{\\mathcal\{L\}\}^\{\(t\)\}\(W\)\\;=\\;\\sum\_\{i=1\}^\{t\-1\}\\frac\{1\}\{m\}\\sum\_\{j\\in\\mathcal\{I\}\_\{i\}\}\\ell\(W,Z\_\{j\}^\{i\}\)\\;\+\\;\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}\\ell\(W,Z\_\{j\}^\{t\}\),\(1\)with the population risk

ℒ\(t\)​\(W\)=∑i=1t𝔼Z∼𝒟i​\[ℓ⁡\(W,Z\)\]\.\\mathcal\{L\}^\{\(t\)\}\(W\)\\;=\\;\\sum\_\{i=1\}^\{t\}\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{i\}\}\\bigl\[\\ell\(W,Z\)\\bigr\]\.\(2\)Since performance is evaluated after training over all tasks, we writeℒ​\(W\)≡ℒ\(T\)​\(W\)\\mathcal\{L\}\(W\)\\equiv\\mathcal\{L\}^\{\(T\)\}\(W\)andℒ^​\(W\)≡ℒ^\(T\)​\(W\)\\hat\{\\mathcal\{L\}\}\(W\)\\equiv\\hat\{\\mathcal\{L\}\}^\{\(T\)\}\(W\), and define the expected generalization error of replay\-based continual learning as

genW≜𝔼W,D1:T,\{ℐi\}i<T\[ℒ\(W\)−ℒ^\(W\)\],\\mathrm\{gen\}\_\{W\}\\;\\triangleq\\;\\mathbb\{E\}\_\{W,D^\{1:T\},\\\{\\mathcal\{I\}\_\{i\}\\\}\_\{i<T\}\}\\bigl\[\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)\\bigr\],\(3\)where the expectation is over the task data, the algorithm randomness, and the replay\-sampling procedure\. HereWWis not a fixed parameter but the random output of the learning algorithm: its law is induced jointly by the streamD1:TD^\{1:T\}, the random buffer construction, and any stochasticity of the optimizer\.

### III\-BLayer\-wise Replay Distributions

Layer\-wise representations\.We analyze a depth\-LLnetworkfW:𝒳→ℝKf\_\{W\}:\\mathcal\{X\}\\to\\mathbb\{R\}^\{K\}with weightsW=\(W1,…,WL\)W=\(W\_\{1\},\\ldots,W\_\{L\}\), whereWl∈ℝdl×dl−1W\_\{l\}\\in\\mathbb\{R\}^\{d\_\{l\}\\times d\_\{l\-1\}\}\. Settinga0​\(x\)=xa\_\{0\}\(x\)=x, each layer defines the intermediate representation

al​\(x\)=ϕl​\(Wl​al−1​\(x\)\),l∈\[L\],a\_\{l\}\(x\)\\;=\\;\\phi\_\{l\}\\\!\\left\(W\_\{l\}\\,a\_\{l\-1\}\(x\)\\right\),\\qquad l\\in\[L\],\(4\)for an element\-wise activationϕl\\phi\_\{l\}, and the output isy^=fW​\(x\)=aL​\(x\)\\hat\{y\}=f\_\{W\}\(x\)=a\_\{L\}\(x\)\. We writeAl=al​\(X\)A\_\{l\}=a\_\{l\}\(X\)for the layer\-llrepresentation of a random inputXX, andAj,lt=al​\(Xjt\)A\_\{j,l\}^\{t\}=a\_\{l\}\(X\_\{j\}^\{t\}\)for the representation of thejj\-th example of tasktt\. We useW1:l=\(W1,…,Wl\)W\_\{1:l\}=\(W\_\{1\},\\dots,W\_\{l\}\)for the bottom \(feature\) sub\-network up to layerllandWl\+1:LW\_\{l\+1:L\}for the remaining suffix\.

Split\-layer distributions\.To analyze generalization at an arbitrary depth, we fix a split layerl∈\{0,…,L\}l\\in\\\{0,\\dots,L\\\}and condition on the learned bottom parametersW1:lW\_\{1:l\}, with the conventions thatW1:0W\_\{1:0\}andWL\+1:LW\_\{L\+1:L\}are empty\. We unify the data source of each task asSiS^\{i\}, whereSi=MiS^\{i\}=M^\{i\}for an old taski<Ti<TandSi=DTS^\{i\}=D^\{T\}for the current taski=Ti=T, and writeS≜⋃i=1TSiS\\triangleq\\bigcup\_\{i=1\}^\{T\}S^\{i\}for the full training set at timeTT\. Replay generalization is governed by how three laws of the representation–label pair\(Al,Y\)\(A\_\{l\},Y\)relate to one another, which we introduce in turn\.

*\(i\) Population law\.*PAl,Y\|i,W1:lP\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}is the population distribution of\(Al,Y\)\(A\_\{l\},Y\)obtained by pushing\(X,Y\)∼𝒟i\(X,Y\)\\sim\\mathcal\{D\}\_\{i\}through the learned features\. Here, “pushing throughW1:lW\_\{1:l\}” acts on the input only:X↦Al=al​\(X\)X\\mapsto A\_\{l\}=a\_\{l\}\(X\), while the labelYYis carried over unchanged\. This is the target the learner would like to control but never observes directly\.

*\(ii\) Empirical proxy\.*P^Al,Y\|Si,W1:l\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\}is the uniform distribution over the representation\-label pairs induced by the finite sampleSiS^\{i\}under the learned featuresW1:lW\_\{1:l\}\. This is the actual objective used in empirical risk minimization\.

*\(iii\) Replay centroid\.*To quantify the statistical bias introduced by the finite buffer, we average over the randomness of buffer construction and define, fori<Ti<T,

QAl,Y\|i,W1:l≜𝔼Si\[P^Al,Y\|Si,W1:l\|W1:l\]\.Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\;\\triangleq\\;\\mathbb\{E\}\_\{S^\{i\}\}\\left\[\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\}\\;\\middle\|\\;W\_\{1:l\}\\right\]\.\(5\)The subscriptW1:lW\_\{1:l\}and the outer expectation are not redundant: the inner objectP^Al,Y\|Si,W1:l\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\}is the proxy for a single realized bufferSiS^\{i\}, whereas the outer𝔼Si\[⋅∣W1:l\]\\mathbb\{E\}\_\{S^\{i\}\}\[\\,\\cdot\\mid W\_\{1:l\}\]averages over whichmmexemplars happened to be stored, holding the features at their learned value\. Equivalently,QAl,Y\|i,W1:lQ\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}is the conditional distribution of a random training example from taskiigiven the learned representation parameters\. For old tasks, in general, one hasQAl,Y\|i,W1:l≠PAl,Y\|i,W1:lQ\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\neq P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\. The cause is a representation\-induced selection effect: although the raw exemplars are drawn i\.i\.d\. from𝒟i\\mathcal\{D\}\_\{i\}—so that at the input layer a uniformly chosen stored example already carries the task marginal—the feature mapW1:lW\_\{1:l\}is itself fitted to the particular realized buffer\. Conditioning on this data\-dependentW1:lW\_\{1:l\}ties the stored representations to the very sample that shaped them, shifting their averaged law away from the population\.

*\(iv\) Stacked training representations\.*Collecting the layer\-llrepresentations of the entire training set gives

𝐔\(l\)≜\{\(al​\(x\),y\):\(x,y\)∈S\}≡\(𝐔old\(l\),𝐔new\(l\)\),\\mathbf\{U\}^\{\(l\)\}\\;\\triangleq\\;\\big\\\{\(a\_\{l\}\(x\),y\):\(x,y\)\\in S\\big\\\}\\;\\equiv\\;\\big\(\\mathbf\{U\}\_\{\\mathrm\{old\}\}^\{\(l\)\},\\;\\mathbf\{U\}\_\{\\mathrm\{new\}\}^\{\(l\)\}\\big\),\(6\)where𝐔old\(l\)\\mathbf\{U\}\_\{\\mathrm\{old\}\}^\{\(l\)\}gathers the representations of the buffered exemplars⋃i<TMi\\bigcup\_\{i<T\}M^\{i\}and𝐔new\(l\)\\mathbf\{U\}\_\{\\mathrm\{new\}\}^\{\(l\)\}those of the current taskDTD^\{T\}\. We emphasize that𝐔old\(l\)\\mathbf\{U\}\_\{\\mathrm\{old\}\}^\{\(l\)\}is a collection of representations, not of raw exemplars: the stored inputs are frozen once written to the buffer, but their imagesal​\(⋅\)a\_\{l\}\(\\cdot\)shift as the shared featuresW1:lW\_\{1:l\}are learned\. This is exactly why replayed data can carry information about the current optimization even though the buffer content is fixed\.

*\(v\) Decoupled reference\.*LetP𝐔\(l\)∣W1:lP\_\{\\mathbf\{U\}^\{\(l\)\}\\mid W\_\{1:l\}\}denote the joint law of this stacked sequence\. Because all coordinates share the data\-dependent mapW1:lW\_\{1:l\}and the buffer is drawn without replacement, the entries of𝐔\(l\)\\mathbf\{U\}^\{\(l\)\}are statistically dependent\. To isolate this coupling, we compare against an idealized Decoupled Reference in which every coordinate is drawn independently, with old\-task coordinates from their centroid and current\-task coordinates from the population:

Q~𝐔\(l\)\|W1:l≜\(⨂i=1T−1\(QAl,Y\|i,W1:l\)⊗m\)⊗\(PAl,Y\|T,W1:l\)⊗n\.\\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(l\)\}\|W\_\{1:l\}\}\\\!\\triangleq\\\!\\left\(\\bigotimes\_\{i=1\}^\{T\-1\}\(Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\)^\{\\otimes m\}\\right\)\\otimes\(P\_\{A\_\{l\},Y\|T,W\_\{1:l\}\}\)^\{\\otimes n\}\.\(7\)The tensor exponents record the sample budgets:\(⋅\)⊗m\(\\cdot\)^\{\\otimes m\}placesmmi\.i\.d\. copies for each of theT−1T\-1past tasks \(one per stored exemplar\), and\(⋅\)⊗n\(\\cdot\)^\{\\otimes n\}placesnni\.i\.d\. copies for the current task\. The reference is therefore an i\.i\.d\. process across the stored and current coordinates: it preserves every coordinate’s marginal but strips away the cross\-sample coupling created by non\-replacement sampling and by the shared feature map\.

## IVA Layer\-wise Decomposition of Replay Generalization

This section contains the backbone result of the paper\. Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)is the only place where the full replay generalization gap is decomposed\. The remainder of the paper revisits its two components separately: the next subsection presents the decomposition itself, Section[IV\-B](https://arxiv.org/html/2608.11690#S4.SS2)relaxes the drift branch geometrically when KL control becomes vacuous, and Section[V](https://arxiv.org/html/2608.11690#S5)upper bounds the optimization branch for a concrete noisy optimizer\.

### IV\-AMain Decomposition: The Synergy\-Drift Bound

The classical stability–plasticity framework is useful\[[48](https://arxiv.org/html/2608.11690#bib.bib15)\], but it does not by itself isolate the replay\-specific bias induced by finite memory nor the dependence created by repeatedly optimizing on replayed and current samples\. The next theorem addresses both issues simultaneously: it isolates replay\-induced drift and then decomposes the remaining optimization complexity into stability, plasticity, interaction, and residual coupling\.

###### Theorem IV\.1\(Hierarchical Synergy–Drift Bound\)\.

LetWWbe the output of a learner minimizing the empirical riskℒ^​\(W\)\\hat\{\\mathcal\{L\}\}\(W\)\. Assume the loss functionℓ⁡\(W,Z\)\\ell\(W,Z\)isσ\\sigma\-subgaussian for allZ∈𝒵Z\\in\\mathcal\{Z\}andW∈𝒲W\\in\\mathcal\{W\}\. For any split layerl∈\{0,…,L\}l\\in\\\{0,\\dots,L\\\}, we have

\|genW\|≤\(T−1\)​2​σ2​𝒦\(l\)⏟Replay\-centroid Drift\+2​σ2Neff​\(𝒮\(l\)\+𝒫\(l\)−ℛ\(l\)\+𝒞\(l\)\)⏟Optimization Variance,\\displaystyle\|\\mathrm\{gen\}\_\{W\}\|\\leq\\underbrace\{\(T\-1\)\\sqrt\{2\\sigma^\{2\}\\,\\mathcal\{K\}^\{\(l\)\}\}\}\_\{\\text\{\\bf Replay\-centroid Drift\}\}\+\\underbrace\{\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\\Big\(\\mathcal\{S\}^\{\(l\)\}\+\\mathcal\{P\}^\{\(l\)\}\-\\mathcal\{R\}^\{\(l\)\}\+\\mathcal\{C\}^\{\(l\)\}\\Big\)\}\}\_\{\\text\{\\bf Optimization Variance\}\},whereNeff≜1\(T−1\)/m\+1/nN\_\{\\mathrm\{eff\}\}\\triangleq\\frac\{1\}\{\(T\-1\)/m\+1/n\}is the effective sample size,𝒦\(l\)≜1T−1∑i=1T−1𝔼W1:l\[DKL\(QAl,Y\|i,W1:l∥PAl,Y\|i,W1:l\)\]\\,\\mathcal\{K\}^\{\(l\)\}\\,\\triangleq\\,\\frac\{1\}\{T\-1\}\\sum\_\{i=1\}^\{T\-1\}\\mathbb\{E\}\_\{W\_\{1:l\}\}\[D\_\{\\mathrm\{KL\}\}\(Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\\|P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\)\],𝒞\(l\)≜𝔼W1:l\[DKL\(P𝐔\(l\)∣W1:l∥Q~𝐔\(l\)\|W1:l\)\]\\mathcal\{C\}^\{\(l\)\}\\triangleq\\mathbb\{E\}\_\{W\_\{1:l\}\}\[D\_\{\\mathrm\{KL\}\}\(P\_\{\\mathbf\{U\}^\{\(l\)\}\\mid W\_\{1:l\}\}\\\|\\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(l\)\}\|W\_\{1:l\}\}\)\]\.

Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)splits replay generalization into two parts\. The replay\-centroid drift term captures the bias induced by finite memory: at layerll, the mismatch between the replay centroidQAl,Y\|i,W1:lQ\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}and the corresponding population distributionPAl,Y\|i,W1:lP\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\. A key feature of this term is that it generally does not vanish as training proceeds, even though the raw exemplars are drawn i\.i\.d\. from the task\. The i\.i\.d\. sampling guarantees only that the stored data match the population in input space, which is why the drift is zero at the input layer\. In feature space the situation differs, because the mapW1:lW\_\{1:l\}is itself fitted to the particular stored buffer\. Conditioning on this buffer\-dependent map couples the stored representations to the same samples that shaped the map, and this selection effect displaces their averaged law away from the population\. The second term measures the optimization dependence under the actual training representations\.

Its prefactor2​σ2/Neff\\sqrt\{2\\sigma^\{2\}/N\_\{\\mathrm\{eff\}\}\}is set by the effective sample sizeNeff=1\(T−1\)/m\+1/nN\_\{\\mathrm\{eff\}\}=\\frac\{1\}\{\(T\-1\)/m\+1/n\}for the replay\-current empirical objective, and the variance term scales as1/Neff1/\\sqrt\{N\_\{\\mathrm\{eff\}\}\}\. When replay is effectively unlimited \(m→∞m\\to\\infty\), the old\-task contribution\(T−1\)/m\(T\-1\)/mvanishes and the current\-task dependence recovers the usual𝒪⁡\(1/n\)\\mathcal\{O\}\(1/\\sqrt\{n\}\)rate\. In contrast, when the current\-task size grows under a fixed replay budget \(n→∞n\\to\\infty\),NeffN\_\{\\mathrm\{eff\}\}approaches the finite limitm/\(T−1\)m/\(T\-1\), so the variance term does not vanish but settles at a finite floor of order\(T−1\)/m\\sqrt\{\(T\-1\)/m\}, which we call the finite\-memory variance floor\. The variance term should not be considered a measure of forgetting\. It bounds the train\-population discrepancy, whereas forgetting is the degradation of old\-task population risk, carried by the drift term𝒦\(l\)\\mathcal\{K\}^\{\(l\)\}, which does not vanish during training, together with the cross\-task dependence accumulated along the optimization trajectory \(Corollary[V\.2](https://arxiv.org/html/2608.11690#S5.Thmtheorem2)\)\.

The optimization term further admits a stability–plasticity–synergy \(SPS\) decomposition\. Here𝒮\(l\)≜I\(𝐔old\(l\);Wl\+1:L∣W1:l\)\\mathcal\{S\}^\{\(l\)\}\\triangleq I\(\\mathbf\{U\}\_\{\\mathrm\{old\}\}^\{\(l\)\};W\_\{l\+1:L\}\\mid W\_\{1:l\}\)measures how much information the suffix weights retain about replayed old\-task representations; this is the stability term, the learner’s information\-theoretic memory of past tasks\. Correspondingly,𝒫\(l\)≜I\(𝐔new\(l\);Wl\+1:L∣W1:l\)\\mathcal\{P\}^\{\(l\)\}\\triangleq I\(\\mathbf\{U\}\_\{\\mathrm\{new\}\}^\{\(l\)\};W\_\{l\+1:L\}\\mid W\_\{1:l\}\)measures the dependence on the current task; this is the plasticity term, the extent to which the suffix encodes and adapts to the new task\. The interaction termℛ\(l\)≜I\(𝐔old\(l\);𝐔new\(l\);Wl\+1:L∣W1:l\)\\mathcal\{R\}^\{\(l\)\}\\triangleq I\(\\mathbf\{U\}\_\{\\mathrm\{old\}\}^\{\(l\)\};\\mathbf\{U\}\_\{\\mathrm\{new\}\}^\{\(l\)\};W\_\{l\+1:L\}\\mid W\_\{1:l\}\)is signed: positive values tighten the bound and correspond to beneficial sharing between old and current representations, whereas negative values enlarge the bound and correspond to interference\. The theorem does not assert thatℛ\(l\)\\mathcal\{R\}^\{\(l\)\}is always nonnegative\. Finally,𝒞\(l\)\\mathcal\{C\}^\{\(l\)\}measures the residual dependence between the actual replay sequence and an idealized independent reference\. This term accounts for statistical coupling introduced by sampling without replacement and by the shared learned representation mapW1:lW\_\{1:l\}\.

###### Corollary IV\.2\(Input\-Layer Bound\)\.

At the input layerl=0l=0, the hierarchical generalization bound simplifies to:

\|genW\|≤2​σ2Neff​I​\(S,W\),\\displaystyle\|\\mathrm\{gen\}\_\{W\}\|\\;\\leq\\;\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\\,I\\\!\\big\(S;W\\big\)\},whereSSis the original\(X,Y\)\(X,Y\)pair in the replay buffer and the current task dataset\.

Specifically, at the input layer,𝐔\(0\)\\mathbf\{U\}^\{\(0\)\}represents the original data pair\. From Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1), we haveI⁡\(S,W\)=I⁡\(𝐔old\(0\),W\)\+I⁡\(𝐔new\(0\),W\)−I⁡\(𝐔old\(0\),𝐔new\(0\),W\)I\(S;W\)=I\(\\mathbf\{U\}\_\{\\mathrm\{old\}\}^\{\(0\)\};W\)\+I\(\\mathbf\{U\}\_\{\\mathrm\{new\}\}^\{\(0\)\};W\)\-I\(\\mathbf\{U\}\_\{\\mathrm\{old\}\}^\{\(0\)\};\\mathbf\{U\}\_\{\\mathrm\{new\}\}^\{\(0\)\};W\)\. The replay centroidQX,Y\|i,WQ\_\{X,Y\|i,W\}coincides with the populationPX,Y\|i,WP\_\{X,Y\|i,W\}and the training sequence is block\-wise i\.i\.d\., therefore𝒦\(0\)=0\\mathcal\{K\}^\{\(0\)\}=0and𝒞\(0\)=0\\mathcal\{C\}^\{\(0\)\}=0\. In the special caseT=1T=1, replay disappears: the empirical and population risks reduce to their standard single\-task form,Neff=nN\_\{\\mathrm\{eff\}\}=n, and the bound recovers the result of\[[50](https://arxiv.org/html/2608.11690#bib.bib18)\]\. Whenm=nm=n, the buffer retains every past example andSSconsists ofT​nTni\.i\.d\. samples; the bound then matches the classical on\-average stability bound\[[44](https://arxiv.org/html/2608.11690#bib.bib19)\]forT​nTnsamples\. Corollary[IV\.2](https://arxiv.org/html/2608.11690#S4.Thmtheorem2)thus extends the standard mutual\-information framework to replay\-based continual learning\. The CL structure therefore does not arise at the input layer itself\. Instead, replay\-specific effects emerge through the finite\-memory factorNeffN\_\{\\mathrm\{eff\}\}, the representation\-level drift term𝒦\(l\)\\mathcal\{K\}^\{\(l\)\}, and the interaction terms appearing forl\>0l\>0, which become relevant only when multiple tasks are learned under memory constraints\.

Then, when consideringl=Ll=L, we obtain a bound dominated by pure drift:

###### Corollary IV\.3\(Output\-Layer Bound\)\.

At the output layerl=Ll=L, the hierarchical generalization bound simplifies to:

\|genW\|≤\(T−1\)​2​σ2​𝒦\(L\)\+2​σ2Neff​𝒞\(L\),\|\\mathrm\{gen\}\_\{W\}\|\\;\\leq\\;\(T\-1\)\\sqrt\{2\\sigma^\{2\}\\,\\mathcal\{K\}^\{\(L\)\}\}\\;\+\\;\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\\,\\mathcal\{C\}^\{\(L\)\}\},where𝒦\(L\)=1T−1∑i=1T−1𝔼W\[DKL\(QAL,Y\|i,W∥PAL,Y\|i,W\)\]\\mathcal\{K\}^\{\(L\)\}=\\frac\{1\}\{T\-1\}\\sum\_\{i=1\}^\{T\-1\}\\mathbb\{E\}\_\{W\}\[D\_\{\\mathrm\{KL\}\}\(Q\_\{A\_\{L\},Y\|i,W\}\\\|P\_\{A\_\{L\},Y\|i,W\}\)\],𝒞\(L\)=𝔼W\[DKL\(P𝐔\(L\)\|W∥Q~𝐔\(L\)\|W\)\]\\mathcal\{C\}^\{\(L\)\}=\\mathbb\{E\}\_\{W\}\[D\_\{\\mathrm\{KL\}\}\(P\_\{\\mathbf\{U\}^\{\(L\)\}\\mid W\}\\\|\\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(L\)\}\|W\}\)\]\.

In particular, Corollary[IV\.3](https://arxiv.org/html/2608.11690#S4.Thmtheorem3)can be viewed as a degenerate limit of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)\. At the output layer, there is no trainable mapping downstream of the representation, hence the stability–plasticity–synergy tradeoff overWl\+1:LW\_\{l\+1:L\}disappears and𝒮\(L\)=𝒫\(L\)=ℛ\(L\)=0\\mathcal\{S\}^\{\(L\)\}=\\mathcal\{P\}^\{\(L\)\}=\\mathcal\{R\}^\{\(L\)\}=0, leaving generalization solely governed by the replay\-centroid drift𝒦\(L\)\\mathcal\{K\}^\{\(L\)\}and dependence divergence𝒞\(L\)\\mathcal\{C\}^\{\(L\)\}\. Read together with Corollary[IV\.2](https://arxiv.org/html/2608.11690#S4.Thmtheorem2), this highlights a depth\-wise shift of the dominant error: from parameter\-information complexity at shallow splits to representational mismatch at deep ones\. This shift is directly actionable\. Because neither end of the network balances the two costs, the bound predicts that stabilization techniques such as feature distillation or partial freezing\[[12](https://arxiv.org/html/2608.11690#bib.bib20)\]should target the interior layer that minimizes the drift–sensitivity product, rather than being applied uniformly across layers\. Figure[3](https://arxiv.org/html/2608.11690#S6.F3)locates this layer empirically, and the feature\-distillation probe confirms that stabilizing the interior basin outperforms stabilizing either end\.

### IV\-BGeometric Relaxation of the Drift Term

Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)identifies replay\-induced drift as one of the two sources of replay generalization error\. Its KL form is natural from the information\-theoretic proof, but it is not always the most informative geometry for replay\. When an old task is represented by finitely many stored examples while the corresponding population representation distribution is continuous, KL\-based drift can become vacuous under support mismatch even if the two measures remain geometrically close\. The next result therefore revisits only the drift branch of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1): it keeps the layer\-wise representation viewpoint but replaces KL by Wasserstein geometry\.

###### Theorem IV\.4\(Hierarchical Wasserstein Bound\)\.

Assume the loss functionℓ⁡\(⋅,y\)\\ell\(\\cdot,y\)isρ0\\rho\_\{0\}\-Lipschitz and the activation functionsϕl\\phi\_\{l\}areρl\\rho\_\{l\}\-Lipschitz\. Then, at timeTT, we have:

\|genW\|≤minl∈\{0,…,L\}𝔼\[ρ¯l\(W\)⋅∑i=1T𝒲1\(P^Al,Y\|Si,W1:l,PAl,Y\|i,W1:l\)\],\\displaystyle\\big\|\\mathrm\{gen\}\_\{W\}\\big\|\\leq\\min\_\{l\\in\\\{0,\\dots,L\\\}\}\\mathbb\{E\}\\bigg\[\\bar\{\\rho\}\_\{l\}\(W\)\\cdot\\sum\_\{i=1\}^\{T\}\\mathcal\{W\}\_\{1\}\\\!\\Big\(\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\},\\,P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\Big\)\\bigg\],whereρ¯l​\(W\):=ρ0​\(1∨∏h=l\+1Lρh​‖Wh‖op\)\.\\bar\{\\rho\}\_\{l\}\(W\):=\\rho\_\{0\}\\left\(1\\vee\\prod\_\{h=l\+1\}^\{L\}\\rho\_\{h\}\\\|W\_\{h\}\\\|\_\{\\mathrm\{op\}\}\\right\)\.

Theorem[IV\.4](https://arxiv.org/html/2608.11690#S4.Thmtheorem4)relaxes the drift component of the generalization bound into a geometric form, complementing rather than replacing the full stability\-plasticity\-interaction \(SPS\) decomposition of the optimization term in Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)\. The Wasserstein distance𝒲1\\mathcal\{W\}\_\{1\}quantifies the distribution mismatch between replay\-induced representation–label pairs\(Al,Y\)\(A\_\{l\},Y\)and their task\-iipopulation counterparts, whileρ¯l​\(W\)\\bar\{\\rho\}\_\{l\}\(W\)measures how strongly the suffix network amplifies this mismatch\. This relaxation admits a finite, geometrically interpretable bound even under support mismatch, at the cost of losing the explicit structural decomposition of the optimization term\.

The layer minimization reveals a depth\-dependent trade\-off\. At shallow layers, representations tend to be shared across tasks, keeping𝒲1\\mathcal\{W\}\_\{1\}small, but any residual mismatch is amplified by a long suffix network, inflatingρ¯l​\(W\)\\bar\{\\rho\}\_\{l\}\(W\)\. At deeper layers, this amplification shrinks while representations become increasingly task\-specific, enlarging the drift\. The optimal layerl⋆l^\{\\star\}is where these two competing effects are balanced, yielding the tightest bound — a phenomenon we formalize as the “generalization funnel layer” in the following\.

###### Corollary IV\.5\(Generalization Funnel Layer\)\.

Define aGeneralization Funnel Layeras any minimizer of the geometric upper\-bound proxy:

l⋆∈arg⁡minl∈\{0,…,L\}𝔼\[ρ0\(∏j=l\+1Lρj∥Wj∥op\)⋅∑i=1T𝒲1\(P^Al,Y\|Si,W1:l,PAl,Y\|i,W1:l\)\]\.\\displaystyle l^\{\\star\}\\in\\mathop\{\\arg\\min\}\_\{l\\in\\\{0,\\dots,L\\\}\}\\mathbb\{E\}\\bigg\[\\rho\_\{0\}\\big\(\\\!\\prod\_\{j=l\+1\}^\{L\}\\rho\_\{j\}\\\|W\_\{j\}\\\|\_\{\\mathrm\{op\}\}\\big\)\\cdot\\sum\_\{i=1\}^\{T\}\\mathcal\{W\}\_\{1\}\\big\(\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\},P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\big\)\\bigg\]\.

Corollary[IV\.5](https://arxiv.org/html/2608.11690#S4.Thmtheorem5)does not claim that every architecture possesses a unique, intrinsic funnel layer\. Rather, it identifies a principled criterion: minimizing the product of suffix sensitivity and representation drift for selecting which layer to be stabilized\. When drift increases with depth while suffix sensitivity decreases, this product attains its minimum at an intermediate layer, providing a bound\-level rationale for intermediate\-layer stabilization strategies such as feature distillation or partial freezing\[[39](https://arxiv.org/html/2608.11690#bib.bib21),[43](https://arxiv.org/html/2608.11690#bib.bib50)\]\.

## VOperationalizing the Optimization Term: An SGLD Instantiation

Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)leaves the optimization branch𝒮\(l\)\+𝒫\(l\)−ℛ\(l\)\+𝒞\(l\)\\mathcal\{S\}^\{\(l\)\}\+\\mathcal\{P\}^\{\(l\)\}\-\\mathcal\{R\}^\{\(l\)\}\+\\mathcal\{C\}^\{\(l\)\}in an optimizer\-agnostic form\. This section makes it concrete for one explicit noisy optimizer, stochastic gradient Langevin dynamics \(SGLD\)\[[6](https://arxiv.org/html/2608.11690#bib.bib51)\]\. The result is specific to this branch and to SGLD: it bounds the optimization complexity, complementing the geometric treatment of the drift term in Section[IV\-B](https://arxiv.org/html/2608.11690#S4.SS2)\.

### V\-AAlgorithm\-Based Hierarchical Bounds

For analytical convenience, we condition on the prefix parametersW1:lW\_\{1:l\}\(the bottomlllayers\) and analyze the generalization performance of the suffix parametersΘ=Wl\+1:L\\Theta=W\_\{l\+1:L\}\. Conducting this analysis for an arbitrary split layer l produces a unified layer\-wise framework\. Let\{Wr\}r=0R\\\{W\_\{r\}\\\}\_\{r=0\}^\{R\}denote the training trajectory of the SGLD\-based CL algorithm for current taskTT, whereW0W\_\{0\}is initialized from the optimized parameters of taskT−1T\-1\. To derive a generalization bound applicable to an arbitrary layerll, we fixW1:lW\_\{1:l\}and update only the remaining parametersΘr\(T\):=Wl\+1:L,r∈ℝdl\+1:L\\Theta\_\{r\}^\{\(T\)\}:=W\_\{l\+1:L,r\}\\in\\mathbb\{R\}^\{d\_\{l\+1:L\}\}\. At each iterationr=1,⋯,Rr=1,\\cdots,R, we sample a replay mini\-batchBroldB\_\{r\}^\{\\mathrm\{old\}\}of sizeboldb\_\{\\mathrm\{old\}\}fromℳ1:T−1\\mathcal\{M\}^\{1:T\-1\}and a current\-task mini\-batchBrnewB\_\{r\}^\{\\mathrm\{new\}\}of sizebnewb\_\{\\mathrm\{new\}\}fromDTD^\{T\}\. Define the corresponding stochastic gradient descent directionsGrold:=−1bold∑Z∈Brold∇Θℓ\(\(W1:l,Θr−1\),Z\)G\_\{r\}^\{\\mathrm\{old\}\}:=\-\\frac\{1\}\{b\_\{\\mathrm\{old\}\}\}\\sum\_\{Z\\in B\_\{r\}^\{\\mathrm\{old\}\}\}\\nabla\_\{\\Theta\}\\ell\(\(W\_\{1:l\},\\Theta\_\{r\-1\}\),Z\)andGrnew:=−1bnew∑Z∈Brnew∇Θℓ\(\(W1:l,Θr−1\),Z\)\.G\_\{r\}^\{\\mathrm\{new\}\}:=\-\\frac\{1\}\{b\_\{\\mathrm\{new\}\}\}\\sum\_\{Z\\in B\_\{r\}^\{\\mathrm\{new\}\}\}\\nabla\_\{\\Theta\}\\ell\(\(W\_\{1:l\},\\Theta\_\{r\-1\}\),Z\)\.The update rule at iterationrrcan then be formalized by

Θr\(T\)=Θr−1\(T\)\+ηrGr\+Nr,Nr∼𝒩\(0,τr2Idl\+1:L\),\\Theta\_\{r\}^\{\(T\)\}=\\Theta\_\{r\-1\}^\{\(T\)\}\+\\eta\_\{r\}G\_\{r\}\+N\_\{r\},\\qquad N\_\{r\}\\sim\\mathcal\{N\}\(0,\\tau\_\{r\}^\{2\}I\_\{d\_\{l\+1:L\}\}\),where the gradient mixture isGr=λold​Grold\+λnew​GrnewG\_\{r\}=\\lambda\_\{\\mathrm\{old\}\}G\_\{r\}^\{\\mathrm\{old\}\}\+\\lambda\_\{\\mathrm\{new\}\}G\_\{r\}^\{\\mathrm\{new\}\}with weightsλold=T−1T\\lambda\_\{\\mathrm\{old\}\}=\\frac\{T\-1\}\{T\}andλnew=1T\\lambda\_\{\\mathrm\{new\}\}=\\frac\{1\}\{T\}\. These weights point along the gradient ofℒ^​\(W\)\\hat\{\\mathcal\{L\}\}\(W\), whose overall scale is absorbed into the learning rateηr\\eta\_\{r\}, so the update shares the minimizer ofℒ^​\(W\)\\hat\{\\mathcal\{L\}\}\(W\)\. The termNrN\_\{r\}is the isotropic Gaussian noise injected at each step\.

Notably, in traditional single\-task SGLD analysis, the initializationW0W\_\{0\}is typically assumed to be independent of the current dataset\. However, this independence generally fails in CL\. At taskTT, the initializationΘ0\(T\)\\Theta\_\{0\}^\{\(T\)\}is inherited from the previous taskΘ0\(T\)=ΘR\(T−1\)\\Theta\_\{0\}^\{\(T\)\}=\\Theta\_\{R\}^\{\(T\-1\)\}\. Consequently, it may already encode information about the current representation𝐔T\(l\)=\(𝐔old\(l\),𝐔new\(l\)\)\\mathbf\{U\}\_\{T\}^\{\(l\)\}=\(\\mathbf\{U\}\_\{\\mathrm\{old\}\}^\{\(l\)\},\\mathbf\{U\}\_\{\\mathrm\{new\}\}^\{\(l\)\}\)especially when𝐔old\(l\)\\mathbf\{U\}\_\{\\mathrm\{old\}\}^\{\(l\)\}is constructed by replaying past data\. To accurately characterize the optimization at timeTT, we must disentangle the inherited information carried byΘ0\(T\)\\Theta\_\{0\}^\{\(T\)\}from the new information created during task\-TToptimization, and then further decompose the latter along the SGLD trajectory\.

###### Proposition V\.1\(Heritage\)\.

Fix a layerlland condition onW1:lW\_\{1:l\}\. At taskTT, we have

I\(𝐔T\(l\);ΘR\(T\)\|W1:l\)\\displaystyle I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta\_\{R\}^\{\(T\)\}\|W\_\{1:l\}\)≤I\(𝐔T\(l\);Θ0\(T\)\|W1:l\)\+I\(𝐔T\(l\);ΘR\(T\)\|Θ0\(T\),W1:l\)\\displaystyle\\leq I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta\_\{0\}^\{\(T\)\}\|W\_\{1:l\}\)\+I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta\_\{R\}^\{\(T\)\}\|\\Theta\_\{0\}^\{\(T\)\},W\_\{1:l\}\)=:HT\(l\)\+ΔT\(l\)\.\\displaystyle=:H\_\{T\}^\{\(l\)\}\+\\Delta\_\{T\}^\{\(l\)\}\.

This proposition ties the SGLD trajectory back to the main decomposition\. By the interaction identity, the optimization complexity in Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)is a single conditional mutual information,𝒮\(l\)\+𝒫\(l\)−ℛ\(l\)=I\(𝐔T\(l\);Wl\+1:L∣W1:l\)=I\(𝐔T\(l\);ΘR\(T\)∣W1:l\)\\mathcal\{S\}^\{\(l\)\}\+\\mathcal\{P\}^\{\(l\)\}\-\\mathcal\{R\}^\{\(l\)\}=I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};W\_\{l\+1:L\}\\mid W\_\{1:l\}\)=I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta\_\{R\}^\{\(T\)\}\\mid W\_\{1:l\}\), whereΘR\(T\)\\Theta\_\{R\}^\{\(T\)\}is the task\-TTterminal suffixWl\+1:LW\_\{l\+1:L\}\. Proposition[V\.1](https://arxiv.org/html/2608.11690#S5.Thmtheorem1)splits it into a heritage termHT\(l\)H\_\{T\}^\{\(l\)\}, the dependence already present in the inherited initialization, and an incrementΔT\(l\)\\Delta\_\{T\}^\{\(l\)\}, the dependence injected by task\-TToptimization\. Crucially, becauseΘ0\(t\+1\)=ΘR\(t\)\\Theta\_\{0\}^\{\(t\+1\)\}=\\Theta\_\{R\}^\{\(t\)\}, the heritage term is not generated afresh at taskTTbut carries forward the dependent build up over all previous tasks\. Corollary[V\.2](https://arxiv.org/html/2608.11690#S5.Thmtheorem2)makes this accumulation precise as a bound by past increments\.

###### Corollary V\.2\(Cumulative information budget\)\.

Fix a layerlland condition onW1:lW\_\{1:l\}\. For anyT≥2T\\geq 2, the heritage term admits the following cumulative\-budget upper bound:

HT\(l\)=I\(𝐔T\(l\);Θ0\(T\)∣W1:l\)≤∑t=1T−1Δt\(l\),H\_\{T\}^\{\(l\)\}=I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta\_\{0\}^\{\(T\)\}\\mid W\_\{1:l\}\)\\leq\\sum\_\{t=1\}^\{T\-1\}\\Delta\_\{t\}^\{\(l\)\},whereΔt\(l\):=I\(𝐔t\(l\);ΘR\(t\)∣Θ0\(t\),W1:l\)\\Delta\_\{t\}^\{\(l\)\}:=I\(\\mathbf\{U\}\_\{t\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\\mid\\Theta\_\{0\}^\{\(t\)\},W\_\{1:l\}\)is the within\-task information increment in Proposition[V\.1](https://arxiv.org/html/2608.11690#S5.Thmtheorem1)applied to tasktt\.

Corollary[V\.2](https://arxiv.org/html/2608.11690#S5.Thmtheorem2)reveals thatHT\(l\)H\_\{T\}^\{\(l\)\}can be upper bounded by the cumulative sum of past increments\. Therefore, to control the cross\-task heritageHT\(l\)H\_\{T\}^\{\(l\)\}, it suffices to control the sequence of within\-task increments\{Δt\(l\)\}t<T\\\{\\Delta\_\{t\}^\{\(l\)\}\\\}\_\{t<T\}\. This insight reduces the problem of bounding the total complexity to controlling the within\-task incrementΔt\(l\)\\Delta\_\{t\}^\{\(l\)\}\. Thus, our analysis now focuses on deriving a sharp, algorithm\-dependent bound for general single\-task updateΔt\(l\)\\Delta\_\{t\}^\{\(l\)\}, which we instantiate for SGLD in the next theorem\.

###### Theorem V\.3\(Task\-Wise SGLD\-Based Bound\)\.

Fix a split layerl∈\{0,…,L−1\}l\\in\\\{0,\\ldots,L\-1\\\}and condition on the base parametersW1:lW\_\{1:l\}, so that the suffix network remains nonempty\. Assume the loss isσ\\sigma\-subgaussian as in Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)\. For tasktttrained by SGLD forRtR\_\{t\}steps, we have

\|genW\(t\)\|≤\(t−1\)​2​σ2​𝒦t\(l\)\+2​σ2Neff,t​\(𝒞t\(l\)\+∑s=1t∑r=1Rs𝔼⁡\[12​log​det\(I\+ηs,r2τs,r2​Ms,r\(l\)\)\]\),\\displaystyle\\big\|\\mathrm\{gen\}\_\{W^\{\(t\)\}\}\\big\|\\leq\(t\-1\)\\sqrt\{2\\sigma^\{2\}\\mathcal\{K\}\_\{t\}^\{\(l\)\}\}\+\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\},t\}\}\\Big\(\\mathcal\{C\}\_\{t\}^\{\(l\)\}\+\\sum\_\{s=1\}^\{t\}\\sum\_\{r=1\}^\{R\_\{s\}\}\\mathbb\{E\}\\\!\\left\[\\frac\{1\}\{2\}\\log\\det\\\!\\Big\(I\+\\frac\{\\eta\_\{s,r\}^\{2\}\}\{\\tau\_\{s,r\}^\{2\}\}M^\{\(l\)\}\_\{s,r\}\\Big\)\\right\]\\Big\)\},whereMs,r\(l\)=𝔼\[Gs,rGs,r⊤\|Θr−1\(s\),W1:l\]M^\{\(l\)\}\_\{s,r\}=\\mathbb\{E\}\\\!\\big\[G\_\{s,r\}G\_\{s,r\}^\{\\top\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\\big\], and𝒦t\(l\)\\mathcal\{K\}\_\{t\}^\{\(l\)\},𝒞t\(l\)\\mathcal\{C\}\_\{t\}^\{\(l\)\}, andNeff,tN\_\{\\mathrm\{eff\},t\}denote the corresponding quantities of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)with the task horizonTTreplaced bytt\.

Theorem[V\.3](https://arxiv.org/html/2608.11690#S5.Thmtheorem3)characterizes the generalization dynamics of SGLD\-based CL algorithms from a hierarchical perspective\. It operationalizes only the optimization branch: it reduces the suffix information complexity to a cumulative trajectory budget of log\-determinants\. This budget is governed by the ratio of learning rate to injected noiseηs,r2/τs,r2\\eta^\{2\}\_\{s,r\}/\\tau^\{2\}\_\{s,r\}and the conditional second\-moment matrixMs,r\(l\)M^\{\(l\)\}\_\{s,r\}\. Unlike standard SGLD bounds that depend only on gradient covariance\[[30](https://arxiv.org/html/2608.11690#bib.bib39)\],Ms,r\(l\)M^\{\(l\)\}\_\{s,r\}explicitly captures the systematic direction of the gradient update\. The exact budget is high\-dimensional; the next subsection expands it to first order to expose what about replay drives it and to obtain measurable diagnostics\.

### V\-BGradient\-Moment Diagnostics of Task Interaction

While Theorem[V\.3](https://arxiv.org/html/2608.11690#S5.Thmtheorem3)bounds the optimization term through a cumulative log\-determinant budget, the mechanism by which replay affects this budget remains implicit\. We therefore perform a second\-order expansion that separates per\-step cost into two orthogonal forces: \(i\) Optimization Instability, driven by gradient covariance; and \(ii\) Interaction Cost, driven by the alignment of old\-task and new\-task mean gradients in a sensitivity\-aware metric\. The resulting alignment quantities are motivated proxies for task interaction in the SGLD model\. This reveals the physical meaning of “Synergy” in SGLD: it is not merely a reduction in variance, but a constructive interaction of mean descent directions in a curvature\-induced metric\.

Setup\.Throughout this subsection, we fix a split layerlland condition on the current\(Θr−1\(s\),W1:l\)\(\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\)\. For each taskssand steprr, letαs,r:=ηs,r2/τs,r2\\alpha\_\{s,r\}:=\\eta\_\{s,r\}^\{2\}/\\tau\_\{s,r\}^\{2\}be the signal\-to\-noise ratio\. Recall that the second moment matrixMs,r\(l\)M^\{\(l\)\}\_\{s,r\}is defined over the gradient mixtureGs,r=λs,old​Gs,rold\+λs,new​Gs,rnewG\_\{s,r\}=\\lambda\_\{s,\\mathrm\{old\}\}G^\{\\mathrm\{old\}\}\_\{s,r\}\+\\lambda\_\{s,\\mathrm\{new\}\}G^\{\\mathrm\{new\}\}\_\{s,r\}\.

###### Lemma V\.4\(Second\-moment split\)\.

Define the conditional mean and covariance

μs,r\(l\):=𝔼\[Gs,r∣Θr−1\(s\),W1:l\],Vs,r\(l\):=Cov\(Gs,r∣Θr−1\(s\),W1:l\)\.\\mu^\{\(l\)\}\_\{s,r\}\\\!:=\\\!\\mathbb\{E\}\[G\_\{s,r\}\\mid\\\!\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\],\\\!V^\{\(l\)\}\_\{s,r\}\\\!:=\\\!\\mathrm\{Cov\}\(G\_\{s,r\}\\mid\\\!\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\)\.Then

Ms,r\(l\)=Vs,r\(l\)\+μs,r\(l\)​μs,r\(l\)⊤\.M^\{\(l\)\}\_\{s,r\}=V^\{\(l\)\}\_\{s,r\}\+\\mu^\{\(l\)\}\_\{s,r\}\\mu^\{\(l\)\\top\}\_\{s,r\}\.

Lemma[V\.4](https://arxiv.org/html/2608.11690#S5.Thmtheorem4)separates stochasticity from signal: the covariance termVs,r\(l\)V^\{\(l\)\}\_\{s,r\}quantifies gradient instability when learning taskss, while the rank\-one termμs,r\(l\)​μs,r\(l\)⊤\\mu^\{\(l\)\}\_\{s,r\}\\mu^\{\(l\)\\top\}\_\{s,r\}encodes the systematic descent direction, representing the task interaction\.

###### Proposition V\.5\(Instability–mean decomposition\)\.

LetAs,r\(l\):=I\+αs,r​Vs,r\(l\)≻0A\_\{s,r\}^\{\(l\)\}:=I\+\\alpha\_\{s,r\}V\_\{s,r\}^\{\(l\)\}\\succ 0withαs,r=ηs,r2/τs,r2\\alpha\_\{s,r\}=\\eta\_\{s,r\}^\{2\}/\\tau\_\{s,r\}^\{2\}\. Then

12​log​det\(I\+αs,r​Ms,r\(l\)\)=12​log​det\(As,r\(l\)\)⏟instability\+12​log⁡\(1\+αs,r​μs,r\(l\)⊤​\(As,r\(l\)\)−1​μs,r\(l\)\)⏟interaction cost\.\\displaystyle\\frac\{1\}\{2\}\\log\\det\\\!\\big\(I\+\\alpha\_\{s,r\}M\_\{s,r\}^\{\(l\)\}\\big\)=\\underbrace\{\\frac\{1\}\{2\}\\log\\det\\\!\\big\(A\_\{s,r\}^\{\(l\)\}\\big\)\}\_\{\\textbf\{instability\}\}\+\\underbrace\{\\frac\{1\}\{2\}\\log\\\!\\Big\(1\+\\alpha\_\{s,r\}\\,\\mu\_\{s,r\}^\{\(l\)\\top\}\\big\(A\_\{s,r\}^\{\(l\)\}\\big\)^\{\-1\}\\mu\_\{s,r\}^\{\(l\)\}\\Big\)\}\_\{\\textbf\{interaction cost\}\}\.

Proposition[V\.5](https://arxiv.org/html/2608.11690#S5.Thmtheorem5)splits each step’s budget into two sources\. The instability term is governed by the gradient covarianceVs,r\(l\)V\_\{s,r\}^\{\(l\)\}; a largeVs,r\(l\)V\_\{s,r\}^\{\(l\)\}marks a sharp region of the landscape, which in CL typically means the replay gradients are inconsistent with the current parameters, resulting in catastrophic forgetting\. The interaction cost is the mean\-gradient energy, which captures the task alignment but is filtered by\(As,r\(l\)\)−1\(A\_\{s,r\}^\{\(l\)\}\)^\{\-1\}; this matrix suppresses the mean gradientμs,r\(l\)\\mu\_\{s,r\}^\{\(l\)\}in directions where gradients are most unstable\. This filtering has a consequence worth stating: when the instability term12​log​det\(As,r\(l\)\)\\frac\{1\}\{2\}\\log\\det\\\!\\big\(A\_\{s,r\}^\{\(l\)\}\\big\)is large, the filter shrinks and the second term falls, yet the reduction is not a gain\. It implies that the algorithm discards usable gradient signal along sharp directions rather than making progress, remaining stuck in a sharp region\. Consequently, synergistic interaction requires that the two tasks’ gradients align within flat, stable directions, which the Euclidean inner product does not distinguish\. The next definition supplies a metric that does\.

###### Definition V\.6\(Sensitivity metric\)\.

Define the sensitivity metric

Hs,r\(l\):=\(As,r\(l\)\)−1≻0\.H^\{\(l\)\}\_\{s,r\}:=\\big\(A^\{\(l\)\}\_\{s,r\}\\big\)^\{\-1\}\\succ 0\.For vectorsu,v≠0u,v\\neq 0, define theHs,r\(l\)H^\{\(l\)\}\_\{s,r\}\-inner product and norm by⟨u,v⟩H:=u⊤​Hs,r\(l\)​v\\langle u,v\\rangle\_\{H\}:=u^\{\\top\}H^\{\(l\)\}\_\{s,r\}vand‖u‖H:=⟨u,u⟩H\\\|u\\\|\_\{H\}:=\\sqrt\{\\langle u,u\\rangle\_\{H\}\}, and the induced cosine

cosH⁡\(u,v\):=⟨u,v⟩H‖u‖H​‖v‖H∈\[−1,1\]\.\\cos\_\{H\}\(u,v\):=\\frac\{\\langle u,v\\rangle\_\{H\}\}\{\\\|u\\\|\_\{H\}\\\|v\\\|\_\{H\}\}\\in\[\-1,1\]\.

Geometrically,Hs,r\(l\)H\_\{s,r\}^\{\(l\)\}induces a stability\-aware Mahalanobis metric: it discounts directions of high curvature while prioritizing flat subspaces\. This is consistent with the intuition above\.

###### Corollary V\.7\(Alignment of old/new gradients\)\.

Letμs,rold,\(l\):=𝔼\[Gs,rold∣Θr−1\(s\),W1:l\],μs,rnew,\(l\):=𝔼\[Gs,rnew∣Θr−1\(s\),W1:l\]\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}:=\\mathbb\{E\}\[G^\{\\mathrm\{old\}\}\_\{s,r\}\\mid\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\],\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}:=\\mathbb\{E\}\[G^\{\\mathrm\{new\}\}\_\{s,r\}\\mid\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\], so thatμs,r\(l\)=λs,old​μs,rold,\(l\)\+λs,new​μs,rnew,\(l\)\.\\mu^\{\(l\)\}\_\{s,r\}=\\lambda\_\{s,\\mathrm\{old\}\}\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\+\\lambda\_\{s,\\mathrm\{new\}\}\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\.Then under the sensitivity metricHs,r\(l\)H^\{\(l\)\}\_\{s,r\}in Definition[V\.6](https://arxiv.org/html/2608.11690#S5.Thmtheorem6),

μs,r\(l\)⊤​Hs,r\(l\)​μs,r\(l\)=λs,old2​‖μs,rold,\(l\)‖H2\+λs,new2​‖μs,rnew,\(l\)‖H2\+2​λs,old​λs,new​‖μs,rold,\(l\)‖H​‖μs,rnew,\(l\)‖H​cosH⁡\(μs,rold,\(l\),μs,rnew,\(l\)\),\\displaystyle\\mu^\{\(l\)\\top\}\_\{s,r\}H^\{\(l\)\}\_\{s,r\}\\mu^\{\(l\)\}\_\{s,r\}=\\lambda\_\{s,\\mathrm\{old\}\}^\{2\}\\big\\\|\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\big\\\|\_\{H\}^\{2\}\+\\lambda\_\{s,\\mathrm\{new\}\}^\{2\}\\big\\\|\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\big\\\|\_\{H\}^\{2\}\+2\\lambda\_\{s,\\mathrm\{old\}\}\\lambda\_\{s,\\mathrm\{new\}\}\\big\\\|\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\big\\\|\_\{H\}\\big\\\|\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\big\\\|\_\{H\}\\cos\_\{H\}\(\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\},\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\),where the information\-geometric alignment

Align=cosH⁡\(μs,rold,\(l\),μs,rnew,\(l\)\)\.\\mathrm\{Align\}=\\cos\_\{H\}\\Big\(\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\},\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\Big\)\.

Corollary[V\.7](https://arxiv.org/html/2608.11690#S5.Thmtheorem7)gives the interaction termℛ\(l\)\\mathcal\{R\}^\{\(l\)\}of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)to a concrete geometric reading: whether old and new task gradients align determines whether the per\-step budget is small or large\. When the two gradients align within the stable subspaces identified byHH\(constructive synergy\), they reinforce a common descent direction, the mean\-gradient energy‖μ‖H\\\|\\mu\\\|\_\{H\}is spent on progress, and the budget stays low\. When they conflict \(interference\), the updates fight rather than reinforce:‖μ‖H\\\|\\mu\\\|\_\{H\}and the instability factorAs,r\(l\)A\_\{s,r\}^\{\(l\)\}remain large, the optimization fails to settle, and the cumulative budget surges\. Alignment and interference are thus the two regimes that drive the generalization cost down or up, respectively\.

This sensitivity\-aware alignment is related to the gradient\-projection family—GEM, A\-GEM, and orthogonal projection\[[27](https://arxiv.org/html/2608.11690#bib.bib32),[5](https://arxiv.org/html/2608.11690#bib.bib33),[38](https://arxiv.org/html/2608.11690#bib.bib34)\]: both address the conflict between old\- and current\-task gradients, but ours differs in three concrete ways\. First, the gradients compared here are the replay and current mini\-batch gradients of the running optimizer, not gradients stored on a fixed past subspace\. Second, alignment is read in the sensitivity metricHs,r\(l\)=\(I\+αs,r​Vs,r\(l\)\)−1H\_\{s,r\}^\{\(l\)\}=\(I\+\\alpha\_\{s,r\}V\_\{s,r\}^\{\(l\)\}\)^\{\-1\}rather than the Euclidean inner product, so unstable directions are discounted instead of all coordinates being weighted equally\. Third, the quantity is diagnostic and feeds the soft, per\-step mixing rule of Corollary[V\.9](https://arxiv.org/html/2608.11690#S5.Thmtheorem9), not a hard orthogonality constraint on the update\. The alignment cosine is therefore a reading of the bound rather than a projection imposed on training\.

Log\-Det Budget

AlignmentcosH\\cos\_\{H\}

Forgetting Dynamics

Statistics over 20 seeds

\(a\)\(b\)\(c\)\(d\)\(e\)\(f\)\(g\)\(h\)
Fig\. 1:Decoupling Generalization Dynamics: Synergy vs\. Interference on MNIST\. Vertical dashed lines indicate task transitions\. Top \(Rotate\-MNIST\): Gradient alignment \(cosH\\cos\_\{H\}\) stays well above its interference\-regime level \(\), accompanied by stable retention \(\) and a low trajectory budget \(\)\. Bottom \(Flip\-MNIST\): Persistent anti\-alignment \(cosH≈−1\\cos\_\{H\}\\approx\-1,\) coincides with a sharp growth of the trajectory budget \(\) and the onset of catastrophic forgetting \(\)\. \(\)cosH\\cos\_\{H\}reliably distinguishes synergy from interference\. \(\) The one\-step mixing coefficientλ⋆\\lambda^\{\\star\}\(Corollary[V\.9](https://arxiv.org/html/2608.11690#S5.Thmtheorem9)\) reduces the local per\-step budget ratio relative to fixed mixing, with a larger effect under interference; we presentλ⋆\\lambda^\{\\star\}only as a local diagnostic/control signal, not as a global replay policy\.
### V\-CInterpretable Diagnostics

We conclude our theoretical analysis by translating the trajectory budget into practical scalar signals and a local sensitivity\-weighted mixing rule\. First, a first\-order relaxation of the budget yields decomposable diagnostics:

###### Corollary V\.8\(Scalar diagnostic decomposition\)\.

Conditioning on\(Θr−1\(s\),W1:l\)\(\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\), the per\-step trajectory budget admits the scalar upper bound

∑s=1t∑r=1Rs𝔼⁡\[12​log​det\(I\+αs,r​Ms,r\(l\)\)\]≤12​∑s=1t∑r=1Rsαs,r​𝔼​\[tr⁡\(Vs,r\(l\)\)\+‖μs,r\(l\)‖2\]\.\\displaystyle\\sum\_\{s=1\}^\{t\}\\sum\_\{r=1\}^\{R\_\{s\}\}\\mathbb\{E\}\\left\[\\frac\{1\}\{2\}\\log\\det\\Big\(I\+\\alpha\_\{s,r\}M^\{\(l\)\}\_\{s,r\}\\Big\)\\right\]\\leq\\frac\{1\}\{2\}\\sum\_\{s=1\}^\{t\}\\sum\_\{r=1\}^\{R\_\{s\}\}\\alpha\_\{s,r\}\\,\\mathbb\{E\}\\Big\[\\mathrm\{tr\}\\big\(V^\{\(l\)\}\_\{s,r\}\\big\)\+\\\|\\mu^\{\(l\)\}\_\{s,r\}\\\|^\{2\}\\Big\]\.Moreover, the two scalar components admit the explicit decompositionstr⁡\(Vs,r\(l\)\)=λs,old2​tr​\(Σs,rold,\(l\)\)\+λs,new2​tr​\(Σs,rnew,\(l\)\),\\mathrm\{tr\}\\big\(V^\{\(l\)\}\_\{s,r\}\\big\)=\\lambda\_\{s,\\mathrm\{old\}\}^\{2\}\\mathrm\{tr\}\\big\(\\Sigma^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\big\)\+\\lambda\_\{s,\\mathrm\{new\}\}^\{2\}\\mathrm\{tr\}\\big\(\\Sigma^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\big\),and‖μs,r\(l\)‖2=λs,old2​‖μs,rold,\(l\)‖2\+λs,new2​‖μs,rnew,\(l\)‖2\+2​λs,old​λs,new​⟨μs,rold,\(l\),μs,rnew,\(l\)⟩\\\|\\mu^\{\(l\)\}\_\{s,r\}\\\|^\{2\}=\\lambda\_\{s,\\mathrm\{old\}\}^\{2\}\\\|\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\\|^\{2\}\+\\lambda\_\{s,\\mathrm\{new\}\}^\{2\}\\\|\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\\|^\{2\}\+2\\lambda\_\{s,\\mathrm\{old\}\}\\lambda\_\{s,\\mathrm\{new\}\}\\langle\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\},\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\rangle\.

Corollary[V\.8](https://arxiv.org/html/2608.11690#S5.Thmtheorem8)yields three monitorable scalar quantities\.tr⁡\(Σs,rold,\(l\)\)\\mathrm\{tr\}\(\\Sigma^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\)measures gradient inconsistency within the replay buffer under the current representation, with elevated values serving as an early indicator of catastrophic forgetting when replay no longer provides stable optimization anchors\.tr⁡\(Σs,rnew,\(l\)\)\\mathrm\{tr\}\(\\Sigma^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\)indicates the intrinsic stochasticity of the new task gradients\.⟨μs,rold,\(l\),μs,rnew,\(l\)⟩\\langle\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\},\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\rangledirectly quantifies cross\-task alignment, distinguishing positive transfer from destructive interference\.

###### Corollary V\.9\(Sensitivity\-weighted local mixing\)\.

Fix any sensitivity metricHs,r\(l\)H^\{\(l\)\}\_\{s,r\}\. Forλ∈\[0,1\]\\lambda\\in\[0,1\], consider the mixed vectorμs,r\(l\)​\(λ\):=λ​μs,rold,\(l\)\+\(1−λ\)​μs,rnew,\(l\)\\mu\_\{s,r\}^\{\(l\)\}\(\\lambda\):=\\lambda\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\+\(1\-\\lambda\)\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}and its metric energy‖μs,r\(l\)​\(λ\)‖H2\\\|\\mu\_\{s,r\}^\{\(l\)\}\(\\lambda\)\\\|\_\{H\}^\{2\}\. Then‖μs,r\(l\)​\(λ\)‖H2\\\|\\mu^\{\(l\)\}\_\{s,r\}\(\\lambda\)\\\|\_\{H\}^\{2\}is minimized overλ∈\[0,1\]\\lambda\\in\[0,1\]by

λ⋆=clip\[0,1\]​\(μs,rnew,\(l\)⊤​H​\(μs,rnew,\(l\)−μs,rold,\(l\)\)\(μs,rold,\(l\)−μs,rnew,\(l\)\)⊤​H​\(μs,rold,\(l\)−μs,rnew,\(l\)\)\),\\lambda^\{\\star\}\\\!=\\\!\\mathrm\{clip\}\_\{\[0,1\]\}\\\!\\left\(\\\!\\frac\{\{\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\}^\{\\top\}H\(\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\-\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\)\}\{\(\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\\!\-\\\!\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\)^\{\\top\}H\(\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\\!\-\\\!\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\)\}\\\!\\right\),with the conventionλ⋆=0\\lambda^\{\\star\}=0ifμs,rold,\(l\)=μs,rnew,\(l\)\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}=\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\.

Corollary[V\.9](https://arxiv.org/html/2608.11690#S5.Thmtheorem9)derives a local mixing coefficientλ⋆\\lambda^\{\\star\}that minimizes theHH\-metric energy of the mixed mean update\. It should therefore be interpreted as a one\-step geometric diagnostic/control rule rather than as a universal globally optimal replay policy\. Without an explicit current\-task progress constraint, minimizing this local energy need not optimize the overall continual\-learning objective\. Under task alignment, the energy is relatively insensitive to the mixing ratio, permitting flexible interpolation between replay and current\-task gradients\. Under conflict,λ⋆\\lambda^\{\\star\}identifies the mixture that least excites the sensitive directions encoded inHH; replay\-heavy values should therefore be read as warning signals of severe interference rather than as standalone prescriptions\.

## VIExperiments

In this section, we empirically evaluate the predictions made by Theorems[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1),[IV\.4](https://arxiv.org/html/2608.11690#S4.Thmtheorem4)and[V\.3](https://arxiv.org/html/2608.11690#S5.Thmtheorem3)together with their corollaries\. Our goal is to determine whether the bounds and diagnostics derived in Sections[IV](https://arxiv.org/html/2608.11690#S4)–[V](https://arxiv.org/html/2608.11690#S5)retain predictive power across \(i\) controlled settings in which ground\-truth distributional objects are accessible, and \(ii\) standard replay benchmarks in which these objects must be estimated from finite samples\.

### VI\-AExperimental Details

##### Datasets, architectures, and replay methods

Controlled experiments use a synthetic Gaussian\-mixture stream and a binary MNIST stream optimized by SGLD, both of which permit closed\-form or large\-sample estimation of the replay centroidQAl,Y\|i,W1:lQ\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}and of the task populationPAl,Y\|i,W1:lP\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\. Standard\-benchmark experiments use Split\-CIFAR\-100 \(10 disjoint 10\-class tasks\) and Split\-TinyImageNet \(20 disjoint 10\-class tasks\)\. Three replay methods are evaluated as instantiations of the framework: experience replay \(ER\)\[[36](https://arxiv.org/html/2608.11690#bib.bib13)\], DER\+\+\[[3](https://arxiv.org/html/2608.11690#bib.bib24)\], and iCaRL\[[33](https://arxiv.org/html/2608.11690#bib.bib14)\]\. All benchmark configurations use buffer capacity2,0002\{,\}000,2020epochs per task, mini\-batch size6464, replay batch size6464, and SGD with learning rate0\.030\.03, momentum0\.90\.9, and weight decay5×10−45\\times 10^\{\-4\}\. For every \(benchmark, method\) pair, we report22class orderings×3\\times\\,3random seeds=6=6independent runs\.

##### Depth family and the suffix amplification factor

The funnel study of Section[IV\-B](https://arxiv.org/html/2608.11690#S4.SS2)additionally trains a family of residual networks, ResNet\-\{10,18,26,34\}\\\{10,18,26,34\\\}, on Split\-CIFAR\-100 and Split\-TinyImageNet with both ER and DER\+\+, spanning66to1818representation cuts; the1212\(method, dataset, depth\) cells with ER and DER\+\+ enter the analysis below\. Theory and experiment here concern a single quantity, the Lipschitz constant of the suffix mapaℓ↦y^a\_\{\\ell\}\\mapsto\\hat\{y\}, which sets how strongly the network above layerℓ\\ellamplifies a representation mismatch\. Theorem[IV\.4](https://arxiv.org/html/2608.11690#S4.Thmtheorem4)bounds it in closed form byρ¯ℓ​\(W\)\\bar\{\\rho\}\_\{\\ell\}\(W\), a bound covered by Remark[IV\.6](https://arxiv.org/html/2608.11690#S4.Thmtheorem6)once‖Wh‖op\\\|W\_\{h\}\\\|\_\{\\mathrm\{op\}\}is read as the per\-block Lipschitz constant\. Because that worst\-case product is many orders of magnitude loose in deep networks, we additionally estimate, for a depth\-resolved location, a data\-dependent surrogate for the same Lipschitz constant, the on\-distribution Jacobian norm ofaℓ↦y^a\_\{\\ell\}\\mapsto\\hat\{y\}, which is a lower bound on it\.

##### Performance and diagnostic observables

LetAt,iA\_\{t,i\}denote the test accuracy on taskiievaluated immediately after training on taskttcompletes\. The standard class\-incremental metrics are

AAt≜1t​∑i≤tAt,i,AFt≜1t−1​∑i<t\(maxs<t⁡As,i−At,i\),t≥2\.\\mathrm\{AA\}\_\{t\}\\triangleq\\frac\{1\}\{t\}\\sum\_\{i\\leq t\}A\_\{t,i\},\\qquad\\mathrm\{AF\}\_\{t\}\\triangleq\\frac\{1\}\{t\-1\}\\sum\_\{i<t\}\\left\(\\max\_\{s<t\}A\_\{s,i\}\-A\_\{t,i\}\\right\),\\quad t\\geq 2\.HereAAt\\mathrm\{AA\}\_\{t\}is the average task accuracy, andAFt\\mathrm\{AF\}\_\{t\}represents the average forgetting over all previously learned tasks\.

After each task transitiont→t\+1t\\\!\\to\\\!t\{\+\}1we additionally compute, for every prior taski≤ti\\leq t, the diagnostic quantities: the sensitivity\-aware gradient alignmentcosH⁡\(i,t\)\\cos\_\{H\}\(i,t\)derived from the SGLD analysis of Section[V](https://arxiv.org/html/2608.11690#S5), the layer\-ℓ\\elldrift estimateDℓ​\(i,t\)D\_\{\\ell\}\(i,t\)from the geometric relaxation of Section[IV\-B](https://arxiv.org/html/2608.11690#S4.SS2), the suffix sensitivityρ¯ℓ​\(Wt\)\\bar\{\\rho\}\_\{\\ell\}\(W\_\{t\}\)from Theorem[IV\.4](https://arxiv.org/html/2608.11690#S4.Thmtheorem4), and the gradient\-covariance tracestr​\(Σold​\(i,t\)\)\\mathrm\{tr\}\\big\(\\Sigma\_\{\\mathrm\{old\}\}\(i,t\)\\big\)andtr​\(Σnew​\(t\)\)\\mathrm\{tr\}\\big\(\\Sigma\_\{\\mathrm\{new\}\}\(t\)\\big\)\. The central response variable is the pairwise forgetting

Δ​Fi,t≜At,i−At\+1,i\.\\Delta F\_\{i,t\}\\triangleq A\_\{t,i\}\-A\_\{t\+1,i\}\.

##### Statistical methodology

Two choices make the benchmark analysis robust to the small\-sample artifacts that affect transition\-level studies\. First, the unit of analysis is the \(old\-task, new\-task\) pair rather than the task boundary: each Split\-TinyImageNet run contributes190190pairs and each Split\-CIFAR\-100 run4545, for3,6903\{,\}690pairwise observations in total, which is two orders of magnitude more than a one\-point\-per\-transition analysis\. Second, because any quantity that drifts monotonically with task index would appear spuriously predictive, for each diagnostic we report the task\-controlled partial Pearson correlation with pairwise forgetting after partialling out the new\-task index and the old\-task age\. To respect the dependence structure of the data, the partial correlation is computed within each run \(over its4545or190190pairs\); the reported95%95\\%bootstrap interval and sign consistency are then taken over the six runs of each cell, so the interval reflects between\-run variability rather than treating non\-independent pairs as independent samples\. We compare each partial correlation against two naive monotone baselines: the raw correlation of forgetting with task index and with old\-task age\. Unless stated otherwise, the sensitivity metricHs,r\(l\)H^\{\(l\)\}\_\{s,r\}is used in its diagonal approximation, so that all diagnostics are computed inO⁡\(d\)O\(d\)time and require nod×dd\\times dmatrix\.

\(a\)Gap scaling\(b\)plateau scaling
Fig\. 2:Scaling laws of the generalization gap for fixedmmornn\.

### VI\-BControlled Test of the Effective\-Sample\-Size Term

This subsection isolates the optimization\-variance branch of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1), in which the dependence on memory sizemm, current\-task sizenn, and sequence lengthTTenters through the effective sample sizeNeff=1\(T−1\)/m\+1/n\.N\_\{\\mathrm\{eff\}\}=\\frac\{1\}\{\(T\-1\)/m\+1/n\}\.Two scaling predictions inmmandnnfollow directly from the prefactor2​σ2/Neff\\sqrt\{2\\sigma^\{2\}/N\_\{\\mathrm\{eff\}\}\}in the bound\. First, for fixed\(m,T\)\(m,T\), increasing the current\-task sizenndoes not drive the variance term to zero but instead yields a finite plateaulimn→∞Neff=mT−1,\\lim\_\{n\\to\\infty\}N\_\{\\mathrm\{eff\}\}=\\frac\{m\}\{T\-1\},which we refer to as the finite\-memory variance floor\. Second, in the regimem≪nm\\ll n,Neff≈m/\(T−1\)N\_\{\\mathrm\{eff\}\}\\approx m/\(T\-1\), so the variance term scales asm−1/2m^\{\-1/2\}in the buffer size\. The empirical gap reported below is the task\-averaged discrepancy, equal to the total\-risk gap of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)divided by the fixed task countTT\.

We construct aTT\-task Gaussian\-mixture stream in which each task𝒟i\\mathcal\{D\}\_\{i\}defines a binary classification problem between two isotropic Gaussians with task\-specific means and a shared covariance\. The learner is a linear\-Gaussian model, for which the population riskL⁡\(W\)L\(W\)and the empirical riskL^​\(W\)\\hat\{L\}\(W\)are available in closed form, as are all KL terms appearing in Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)\. The empirical generalization gap\|L​\(W\)−L^​\(W\)\|\|L\(W\)\-\\hat\{L\}\(W\)\|is averaged overK=200K=200independent replications of the buffer construction and the current\-task draw\. Figure[2\(a\)](https://arxiv.org/html/2608.11690#S6.F2.sf1)plots the empirical gap as a function ofnnforT=8T=8and three values ofm∈\{40,80,160,320\}m\\in\\\{40,80,160,320\\\}\. Each curve flattens asnngrows, and the asymptote agrees with the prediction: doublingnnin the saturated regime produces a change in the gap that is statistically indistinguishable from zero, while doublingmmshifts the asymptote downward by a factor close to2\\sqrt\{2\}\. Figure[2\(b\)](https://arxiv.org/html/2608.11690#S6.F2.sf2)reports the same data in the log\-log regime, together with an ordinary\-least\-squares fit oflog⁡gap\\log\\mathrm\{gap\}versuslog⁡m\\log m\. The fitted slope isβ^=−0\.443±0\.027\\widehat\{\\beta\}=\-0\.443\\pm 0\.027\(95% bootstrap confidence interval\[−0\.498,−0\.391\]\[\-0\.498,\-0\.391\],K=200K=200replications per point\), so the interval lies just above−1/2\-1/2, and the estimate drifts systematically from−0\.454\-0\.454to−0\.423\-0\.423as the tail threshold used to define the plateau grows\. The prediction under test is them−1/2m^\{\-1/2\}scaling of the variance*prefactor*2​σ2/Neff\\sqrt\{2\\sigma^\{2\}/N\_\{\\mathrm\{eff\}\}\}in isolation\. However, what we fit is the full empirical gap, not the prefactor in isolation: it consists of the variance prefactor2​σ2/Neff\\sqrt\{2\\sigma^\{2\}/N\_\{\\mathrm\{eff\}\}\}multiplied by the square root of the information content𝒮\(ℓ\)\+𝒫\(ℓ\)−ℛ\(ℓ\)\+𝒞\(ℓ\)\\mathcal\{S\}^\{\(\\ell\)\}\+\\mathcal\{P\}^\{\(\\ell\)\}\-\\mathcal\{R\}^\{\(\\ell\)\}\+\\mathcal\{C\}^\{\(\\ell\)\}, with a residual drift contribution that persists at finitenn\. Any mildmm\-dependence of that information content shifts the composite exponent of the gap off the prefactor’s−1/2\-1/2; the measured value sitting slightly above−1/2\-1/2, and stable across tail thresholds, is consistent with such a small perturbation\.

The closed\-form setting lets us test a sharper claim: that the bound tracks the gap as its controlling quantities vary, not merely that it attains the correct asymptote\. Writing the task\-averaged variance proxy as1/Neff1/\\sqrt\{N\_\{\\mathrm\{eff\}\}\}, we first sweep the memory and sample budgets on a4×44\\times 4grid \(m∈\{40,80,160,320\}m\\in\\\{40,80,160,320\\\},n∈\{500,1000,2000,5000\}n\\in\\\{500,1000,2000,5000\\\},T=8T=8\)\. Across the1616cells the proxy and the measured gap are not merely rank\-correlated \(Spearmanρ=0\.94\\rho=0\.94at the cell level\) but essentially linear: a straight\-line fit of the gap against1/Neff1/\\sqrt\{N\_\{\\mathrm\{eff\}\}\}givesR2=0\.97R^\{2\}=0\.97\(Pearsonr=0\.98r=0\.98\), so the proxy reproduces the functional form of the\(m,n\)\(m,n\)dependence, not only its ordering\. The variance branch is thus quantitatively faithful in\(m,n\)\(m,n\)up to its loose absolute constant\. A complementaryTT\-sweep \(fixedm=160,n=1000m\{=\}160,\\,n\{=\}1000,T∈\{2,4,8,16\}T\\in\\\{2,4,8,16\\\},2020seeds\) isolates theNeffN\_\{\\mathrm\{eff\}\}prefactor from the information content of the bound\. We stress at the outset that the variance prefactor2​σ2/Neff=2​σ2​\(\(T−1\)/m\+1/n\)\\sqrt\{2\\sigma^\{2\}/N\_\{\\mathrm\{eff\}\}\}=\\sqrt\{2\\sigma^\{2\}\\big\(\(T\-1\)/m\+1/n\\big\)\}grows withTT, asT/m\\sqrt\{T/m\}in the total\-risk gap; the quantity plotted in Fig\.[3\(b\)](https://arxiv.org/html/2608.11690#S6.F3.sf2)is the task\-averaged proxy1T​2​σ2/Neff∼1/T​m\\tfrac\{1\}\{T\}\\sqrt\{2\\sigma^\{2\}/N\_\{\\mathrm\{eff\}\}\}\\sim 1/\\sqrt\{Tm\}, which decreases only because the fixed1/T1/Tnormalization \(the reported gap is the total\-risk gap of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)divided byTT\) rescales every term alike\. In either convention the prefactor’sTT\-growth is at mostT\\sqrt\{T\}, whereas the measured task\-averaged gap rises from0\.00450\.0045atT=2T\{=\}2to0\.01560\.0156atT=16T\{=\}16—faster than the prefactor and in the opposite direction to its task\-averaged proxy\. The prefactor alone therefore cannot account for theTT\-dependence\. The growth is carried by the bound’s information content: the replay\-centroid drift𝒦\(ℓ\)\\mathcal\{K\}^\{\(\\ell\)\}, which enters the drift term with an explicit\(T−1\)\(T\-1\)factor, and the optimization information𝒮\(ℓ\)\+𝒫\(ℓ\)−ℛ\(ℓ\)\\mathcal\{S\}^\{\(\\ell\)\}\+\\mathcal\{P\}^\{\(\\ell\)\}\-\\mathcal\{R\}^\{\(\\ell\)\}, whose heritage component is non\-decreasing in the number of tasks by Corollary[V\.2](https://arxiv.org/html/2608.11690#S5.Thmtheorem2)\. TheTT\-sweep thus separates the finite\-memory rate prefactorNeffN\_\{\\mathrm\{eff\}\}\(validated by the\(m,n\)\(m,n\)grid above\) from the information content that accumulates withTT\.

\(a\)Variance proxy vs\. gap over the\(m,n\)\(m,n\)grid\(b\)TT\-sweep: gap rises while the variance proxy falls
Fig\. 3:The bound is loose in constant but order\-correct, and its rate prefactor separates cleanly from its information content\.
### VI\-CThe Drift–Sensitivity Trade\-off and the Empirical Funnel

We now turn to the geometric drift component developed in Section[IV\-B](https://arxiv.org/html/2608.11690#S4.SS2)\. This subsection examines whether such a funnel is empirically observable, and whether its location is consistent across controlled and standard\-benchmark settings\.

Three quantities are tracked at each layer indexℓ\\ellover the course of training\. The driftDℓD\_\{\\ell\}is estimated as a sliced Wasserstein\-1 distance between the buffer\-induced and the population\-induced empirical distributions of\(Aℓ,Y\)\(A\_\{\\ell\},Y\), averaged over old tasks\. Suffix sensitivity is calculated in two ways\. The first option is the worst\-case factorρ¯ℓ​\(W\)\\bar\{\\rho\}\_\{\\ell\}\(W\)from Theorem[IV\.4](https://arxiv.org/html/2608.11690#S4.Thmtheorem4)\. We construct this factor by taking the product of layer\-wise spectral norms across layersℓ\+1\\ell\+1throughLL, and every operator norm is computed via power iteration\. This is a certified Lipschitz bound that we use for the controlled Gaussian setting in Fig\.[4\(a\)](https://arxiv.org/html/2608.11690#S6.F4.sf1)\. The second option relies on the on\-distribution Jacobian norm of the suffix mapaℓ↦y^a\_\{\\ell\}\\mapsto\\hat\{y\}\. This data\-driven quantity acts as a lower bound for the Lipschitz constant\. We apply it to find the funnel in the deep residual networks of Fig\.[4\(c\)](https://arxiv.org/html/2608.11690#S6.F4.sf3), since the first method produces results that are too loose to be useful\. We next compute the product proxyGenℓ=Dℓ⋅\(suffix sensitivity\)\\mathrm\{Gen\}\_\{\\ell\}=D\_\{\\ell\}\\cdot\(\\text\{suffix sensitivity\}\)for each layer, and the empirical funnelℓ^⋆\\hat\{\\ell\}^\{\\star\}is determined by taking the argmin of this proxy\.

\(a\)Controlled Gaussian \(closed\-form populations\)\(b\)Funnel location moves deeper with network depth\(c\)ResNet\-\{10,18,26,34\}\\\{10,18,26,34\\\}: drift, suffix amplification, and their product across depth
Fig\. 4:The drift–sensitivity trade\-off and the empirical funnel\. \(a\) On the controlled Gaussian setting with closed\-form populations, the product is minimized at an interior layer \(ℓ⋆=2\\ell^\{\\star\}=2\)\. \(b\) The same trade\-off holds in deep residual networks, where suffix amplification is measured on\-distribution, with the minimum appearing in a broad interior region at every depth\. \(c\) This interior region moves systematically deeper as the network depth increases\.Figure[4\(a\)](https://arxiv.org/html/2608.11690#S6.F4.sf1)reports the three curves on the Controlled Gaussian, where ground\-truth populationsPAℓ,Y∣i,W1:ℓP\_\{A\_\{\\ell\},Y\\mid i,W\_\{1:\\ell\}\}are accessible\. The driftDℓD\_\{\\ell\}decreases first and then increases, whileρ¯ℓ​\(W\)\\bar\{\\rho\}\_\{\\ell\}\(W\)falls; their product is convex with minimizer concentrated atℓ⋆=2\\ell^\{\\star\}=2on a four\-layer architecture, and the two endsℓ∈\{0,L\}\\ell\\in\\\{0,L\\\}are never selected\. In low dimensions, where the worst\-case factorρ¯ℓ\\bar\{\\rho\}\_\{\\ell\}has small dynamic range, the predicted interior minimum is directly observable\.

Figure[4\(c\)](https://arxiv.org/html/2608.11690#S6.F4.sf3)extends the construction to deep residual networks—ResNet\-\{10,18,26,34\}\\\{10,18,26,34\\\}—across both replay methods and both benchmarks, for1212\(method, dataset, depth\) cells\. Two findings hold uniformly\. First, the drift forms a pronounced “bathtub”:DℓD\_\{\\ell\}is large at both the input and the final representation and small across a broad interior, so the drift–amplification product is minimized strictly inside the network in every one of the1212cells\. Second, the interior minimizer moves systematically deeper as the network deepens: its absolute index increases withLL, and the Spearman’sρ\\rhobetween funnel index and depth falls between0\.80\.8and1\.01\.0across all three method–dataset pairs \(Fig\.[4\(b\)](https://arxiv.org/html/2608.11690#S6.F4.sf2)\)\. Because the product is flat across a broad basin where values within20%20\\%of the minimum span3030–47%47\\%of the interior, we report the funnel as an interior region rather than a single sharply\-resolved layer\.

A feature\-distillation probe that stabilizes individual blocks, consistent with Corollary[IV\.5](https://arxiv.org/html/2608.11690#S4.Thmtheorem5), confirms that stabilizing an interior block improves over the no\-distillation baseline and over stabilizing either end, but the two interior blocks nearest the basin are not statistically separated under our seed budget—as expected of a minimizer lying in a flat interior region\. On real benchmarks the funnel proxy is the weakest of our diagnostics \(Table[I](https://arxiv.org/html/2608.11690#S6.T1)\): its partial correlation with forgetting is significant and correctly signed only for ER \(−0\.37\-0\.37and−0\.30\-0\.30\), and is weak or absent for DER\+\+ and iCaRL\. We therefore present the funnel as a structural result about the drift–sensitivity trade\-off in controlled and depth\-resolved settings, and not on the same footing as the alignment diagnostic of Section[VI\-D](https://arxiv.org/html/2608.11690#S6.SS4), whose benchmark transfer is both stronger and consistent across the gradient\-replay methods\.

### VI\-DSGLD Optimization\-Branch Diagnostics

Theorem[V\.3](https://arxiv.org/html/2608.11690#S5.Thmtheorem3)reduces the optimization branch of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)to a trajectory\-level log\-determinant budget, and Lemma[V\.4](https://arxiv.org/html/2608.11690#S5.Thmtheorem4)and Proposition[V\.5](https://arxiv.org/html/2608.11690#S5.Thmtheorem5)decompose the per\-step cost into a covariance\-driven instability term and a mean\-driven interaction cost governed by the sensitivity\-aware alignmentcosH\\cos\_\{H\}\. We evaluate these diagnostics at two levels: first on a controlled MNIST stream where the gradient moments admit clean estimation and the mechanism can be examined directly, and then on standard replay benchmarks where every quantity must be estimated under the noise of practical pipelines\.

#### VI\-D1Controlled MNIST, the mechanism in isolation

On the controlled MNIST stream the gradient moments are estimable, so the decomposition of Proposition[V\.5](https://arxiv.org/html/2608.11690#S5.Thmtheorem5)can be examined mechanism by mechanism \(Fig\.[1](https://arxiv.org/html/2608.11690#S5.F1)\)\. In the positive\-transfer regime \(Rotate\-MNIST\), the per\-run mean alignmentcosH\\cos\_\{H\}sits well above its interference\-regime level \(≈−0\.5\\approx\-0\.5versus≈−0\.85\\approx\-0\.85on the box\-plots of Fig\.[1](https://arxiv.org/html/2608.11690#S5.F1)\(d\)\), the per\-step interaction cost and the gradient covariance remain small, and retention is stable\. In the interference regime \(Flip\-MNIST\),cosH\\cos\_\{H\}collapses to strongly negative values \(per\-run mean≈−0\.85\\approx\-0\.85\), the log\-determinant trajectory budget grows sharply, and forgetting rises\. The alignment signal cleanly separates the two regimes across2020seeds \(Fig\.[1](https://arxiv.org/html/2608.11690#S5.F1)\(d\)\), consistent with Corollary[V\.7](https://arxiv.org/html/2608.11690#S5.Thmtheorem7)read as a statement about relative alignment: the synergy regime sustains a markedly higher \(less negative\)cosH\\cos\_\{H\}than the interference regime\. We read the corollary’s sign prediction in this relative sense, since under the diagonal\-HHapproximation the per\-run meancosH\\cos\_\{H\}is negative in both regimes; it is their separation, not an absolute sign, that the corollary ties to constructive versus destructive interaction and that predicts forgetting\. We note one diagnostic caveat\. Under severe interference the model collapses toward chance, so once train and test risks reach the same entropy floor the terminal forgetting metric saturates and becomes uninformative; the trajectory budget andcosH\\cos\_\{H\}remain informative precisely in this regime, which is why we treat them as the primary diagnostics rather than terminal forgetting alone\.

##### The curvature\-aware metric is not the Euclidean inner product

A natural objection is thatcosH\\cos\_\{H\}merely re\-expresses the Euclidean gradient cosinecosE\\cos\_\{E\}used by orthogonal\-projection methods \(GEM/OGD\)\. The controlled streams refute this directly: because both quantities are computed at every step from the same replay and current\-task gradients, the comparison is not a post\-hoc re\-instrumentation\. On Rotate/Flip\-MNIST,cosH\\cos\_\{H\}separates the synergy and interference regimes more sharply thancosE\\cos\_\{E\}\(standardized between\-regime mean difference3\.83\.8vs\.3\.03\.0\) and predicts per\-run forgetting more strongly \(−0\.89\-0\.89vs\.−0\.84\-0\.84\); the same ordering holds on the synthetic Gaussian stream \(cosH\\cos\_\{H\}separation2\.92\.9vs\.2\.02\.0, forgetting correlation−0\.64\-0\.64vs\.−0\.55\-0\.55\)\. These gaps are statistically reliable rather than artifacts of the seed budget: a paired bootstrap over runs places thecosH−cosE\\cos\_\{H\}\-\\cos\_\{E\}difference above zero with a95%95\\%CI that excludes zero for both the regime separation \(Δ​d\\Delta dCI\[0\.7,1\.4\]\[0\.7,1\.4\]on MNIST\) and the forgetting correlation \(Δ​\|r\|\\Delta\|r\|CI\[0\.02,0\.09\]\[0\.02,0\.09\]on MNIST and\[0\.01,0\.16\]\[0\.01,0\.16\]on the synthetic stream\)\. The improvement is modest but consistent and significant, just as the theory predicts\. The curvature reweightingH=\(I\+α​V\)−1H=\(I\+\\alpha V\)^\{\-1\}discounts alignment along unstable directions \(Proposition[V\.5](https://arxiv.org/html/2608.11690#S5.Thmtheorem5)\)\. Consequently,cosH\\cos\_\{H\}is an intrinsic measure derived from our bound, not merely a hidden variant of the Euclidean OGD cosine\.

\(a\)Regime separation:cosH\\cos\_\{H\}vs\.cosE\\cos\_\{E\}\(b\)Forgetting prediction:cosH\\cos\_\{H\}vs\.cosE\\cos\_\{E\}
Fig\. 5:The sensitivity\-aware alignmentcosH\\cos\_\{H\}improves on the Euclidean gradient cosinecosE\\cos\_\{E\}\(the OGD/GEM quantity\), both logged from identical gradients\. \(a\)cosH\\cos\_\{H\}separates synergy from interference more sharply \(larger between\-regime effect size\)\. \(b\)cosH\\cos\_\{H\}is a stronger predictor of per\-run forgetting\. Rotate/Flip\-MNIST; the synthetic Gaussian stream shows the same ordering\.

#### VI\-D2Standard benchmarks: do the diagnostics track forgetting?

We now test the central empirical claim of the optimization branch: that the abstract interaction termℛ\(l\)\\mathcal\{R\}^\{\(l\)\}, made operational through the alignmentcosH\\cos\_\{H\}of Corollary[V\.7](https://arxiv.org/html/2608.11690#S5.Thmtheorem7), tracks forgetting once estimated on standard replay benchmarks\. Table[I](https://arxiv.org/html/2608.11690#S6.T1)and Fig\.[6](https://arxiv.org/html/2608.11690#S6.F6)report the task\-controlled partial correlations of the diagnostics with pairwise forgetting across all six \(benchmark, method\) cells\.

TABLE I:Task\-controlled partial Pearson correlation between theory\-derived diagnostics and pairwise forgettingΔ​Fi,t\\Delta F\_\{i,t\}, after partialling out the new\-task index and the old\-task age\. Each cell aggregates66runs \(22class orders×3\\times\\,3seeds\);†marks a95%95\\%bootstrap CI that excludes zero\. The last two columns are naive monotone baselines \(raw correlation of forgetting with task index / old\-task age\)\. Negative values are theory\-consistent: alignment of old/new gradients reduces forgetting\. “–” denotes a quantity not recorded for iCaRL under our diagnostic instrumentation, whose distillation\-based objective does not expose the same replay/current gradient split as ER and DER\+\+\.Benchmark / MethodcosH\\cos\_\{H\}⟨μold,μnew⟩\\langle\\mu\_\{\\mathrm\{old\}\},\\mu\_\{\\mathrm\{new\}\}\\rangletr⁡\(Σold\)\\mathrm\{tr\}\(\\Sigma\_\{\\mathrm\{old\}\}\)funnelDℓ​ρ¯ℓD\_\{\\ell\}\\bar\{\\rho\}\_\{\\ell\}task\-idx \(null\)old\-age \(null\)S\-CIFAR100 / ER−0\.948†\-0\.948^\{\\dagger\}−0\.915†\-0\.915^\{\\dagger\}−0\.840†\-0\.840^\{\\dagger\}−0\.372†\-0\.372^\{\\dagger\}−0\.29\-0\.29−0\.62\-0\.62S\-CIFAR100 / DER\+\+−0\.794†\-0\.794^\{\\dagger\}−0\.561†\-0\.561^\{\\dagger\}−0\.185\-0\.185−0\.043\-0\.043−0\.04\-0\.04−0\.48\-0\.48S\-CIFAR100 / iCaRL\+0\.076\+0\.076\+0\.062\+0\.062–−0\.088\-0\.088−0\.43\-0\.43−0\.11\-0\.11S\-TinyImageNet / ER−0\.796†\-0\.796^\{\\dagger\}−0\.281†\-0\.281^\{\\dagger\}−0\.261†\-0\.261^\{\\dagger\}−0\.296†\-0\.296^\{\\dagger\}−0\.23\-0\.23−0\.44\-0\.44S\-TinyImageNet / DER\+\+−0\.645†\-0\.645^\{\\dagger\}−0\.230†\-0\.230^\{\\dagger\}−0\.098\-0\.098−0\.127†\-0\.127^\{\\dagger\}−0\.13\-0\.13−0\.41\-0\.41S\-TinyImageNet / iCaRL\+0\.132†\+0\.132^\{\\dagger\}−0\.156†\-0\.156^\{\\dagger\}–\+0\.074†\+0\.074^\{\\dagger\}−0\.28\-0\.28−0\.10\-0\.10Fig\. 6:Task\-controlled partial Pearson correlation ofcosH\\cos\_\{H\}with pairwise forgetting on the six \(benchmark, method\) cells, with95%95\\%bootstrap intervals\. Blue: significant and theory\-consistent \(CI excludes zero, negative sign\); orange: weak or sign\-inconsistent \(iCaRL\)\. Grey ticks are the two naive monotone baselines \(raw correlation of forgetting with task index and with old\-task age\)\. For ER and DER\+\+ the alignment signal is several times stronger than either monotone baseline, on both benchmarks\.##### Alignment predicts forgetting for gradient\-replay methods

For the two methods whose updates are literally a mixture of replay and current\-task gradients—ER and DER\+\+—the alignment term is a strong and stable predictor of forgetting on both benchmarks\. The partial correlation ofcosH\\cos\_\{H\}with pairwise forgetting is−0\.948\-0\.948\(ER\) and−0\.794\-0\.794\(DER\+\+\) on Split\-CIFAR\-100, and−0\.796\-0\.796\(ER\) and−0\.645\-0\.645\(DER\+\+\) on Split\-TinyImageNet, with confidence intervals that exclude zero and the same sign in all six runs of every cell\. The sign is exactly the one the theory predicts: when old\- and new\-task gradients align in the stability\-aware metricHH, forgetting is small; when they anti\-align, forgetting is large\. Crucially, this is not an artifact of generic task\-order drift\. The task\-controlledcosH\\cos\_\{H\}signal is several times stronger than the two naive monotone baselines \(the raw correlation of forgetting with task index never exceeds0\.290\.29in magnitude, and with old\-task age never exceeds0\.620\.62\), confirming that the predictive content survives after the dominant monotone trends are removed\. The raw alignment inner product⟨μold,μnew⟩\\langle\\mu\_\{\\mathrm\{old\}\},\\mu\_\{\\mathrm\{new\}\}\\rangleand, for ER, the replay\-gradient inconsistencytr⁡\(Σold\)\\mathrm\{tr\}\(\\Sigma\_\{\\mathrm\{old\}\}\)\(−0\.840\-0\.840on Split\-CIFAR\-100\) are theory\-consistent as well, while the geometric funnel proxy transfers more weakly but with the correct sign for ER on both benchmarks \(−0\.372\-0\.372and−0\.296\-0\.296\)\.

##### A scope boundary, stated honestly

The iCaRL columns are included deliberately as a boundary of the framework rather than omitted\. iCaRL classifies by nearest\-mean\-of\-exemplars and is trained with knowledge distillation, so its forgetting mechanism departs from the replay\-plus\-current gradient mixture that the SGLD model of Section[V](https://arxiv.org/html/2608.11690#S5)assumes\. Consistently,cosH\\cos\_\{H\}is no longer a reliable predictor for iCaRL: its CI includes zero on Split\-CIFAR\-100 and the sign even reverses on Split\-TinyImageNet\. The raw alignment⟨μold,μnew⟩\\langle\\mu\_\{\\mathrm\{old\}\},\\mu\_\{\\mathrm\{new\}\}\\ranglestill retains the theory\-consistent negative sign after partialling \(−0\.156\-0\.156on Split\-TinyImageNet\), which is what one would expect if only the part of iCaRL’s update that resembles a gradient step carries the alignment signal\. This is the honest reading of the framework: the gradient\-alignment diagnostic is most faithful precisely for methods whose optimization matches the modeled dynamics, and weaker for exemplar\-classifier methods—a delimitation the theory itself anticipates, rather than a failure to be hidden\.

##### Correlation with actual old\-task test error

Finally, the geometric branch also transfers to actual test error rather than only to the forgetting increment\. Merging the diagnostic logs with the accuracy logs on Split\-CIFAR\-100 with iCaRL, the layer\-wise drift and the funnel proxy correlate with current old\-task test error at\+0\.262\+0\.262\(95%95\\%CI\[0\.132,0\.392\]\[0\.132,0\.392\]\) and\+0\.257\+0\.257\(\[0\.128,0\.385\]\[0\.128,0\.385\]\) respectively\. Taken together, the benchmark study addresses its core motivating question of whether the quantities appearing in the bound retain meaning when estimated under realistic conditions\. It delivers affirmative conclusions for the alignment term and gradient\-replay methods, while clearly defining the limitations of the diagnostics\.

## VIILimitations

We do not claim that the bounds are numerically tight for deep networks: as with all mutual\-information generalization bounds, the absolute multiplicative constants are loose\. What our experiments validate is sharper and, we argue, more useful—three falsifiable predictions about the structure of the gap\. First, the predicted functional form: the variance floor scales asm−1/2m^\{\-1/2\}, the measured exponent is−0\.443\-0\.443\(CI\[−0\.498,−0\.391\]\[\-0\.498,\-0\.391\]\), and the bound is order\-correct beyond this asymptote: across a memory–sample grid, its variance proxy ranks the measured gap with Spearman0\.940\.94\. Second, the predicted geometry: the drift–sensitivity trade\-off, with suffix sensitivity measured on\-distribution, is minimized strictly inside the network in all1212\(method, dataset, depth\) cells we examine, and the interior region moves systematically deeper as depth grows; we claim an interior region and its depth trend, located empirically rather than certified by the worst\-case Lipschitz factor, which is itself too loose to localize a funnel in deep networks\. Third, the predicted monotone relationship: the alignment term that the bound makes responsible for the interaction component negatively tracks forgetting on real benchmarks, with partial correlations from−0\.65\-0\.65to−0\.95\-0\.95for gradient\-replay methods\. A bound can be loose in its constants yet correct in the dependencies it predicts; our experiments target the latter\.

## VIIIConclusion

We developed a layer\-wise information\-theoretic theory of replay generalization in continual learning\. The central message is that replay generalization is governed by two coupled mechanisms: a replay\-induced representation drift created by finite memory, and an optimization\-complexity term created by shared downstream parameters\. The main Synergy–Drift theorem isolates these two mechanisms in a single decomposition\. The Wasserstein result revisits the first branch when KL\-based drift is ill\-posed under support mismatch, while the SGLD result instantiates the second branch through trajectory\-level gradient moments\. Together, these results provide a common language for studying replay bias, task interaction, and layer\-wise sensitivity in continual learning\.

## References

- \[1\]R\. Aljundi, F\. Babiloni, M\. Elhoseiny, M\. Rohrbach, and T\. Tuytelaars\(2018\)Memory aware synapses: learning what \(not\) to forget\.InProceedings of the European conference on computer vision \(ECCV\),pp\. 139–154\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[2\]R\. Aljundi, M\. Lin, B\. Goujaud, and Y\. Bengio\(2019\)Gradient based sample selection for online continual learning\.Advances in neural information processing systems32\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[3\]P\. Buzzega, M\. Boschini, A\. Porrello, D\. Abati, and S\. Calderara\(2020\)Dark experience for general continual learning: a strong, simple baseline\.Advances in Neural Information Processing Systems33,pp\. 15920–15930\.Cited by:[§VI\-A](https://arxiv.org/html/2608.11690#S6.SS1.SSS0.Px1.p1.1)\.
- \[4\]L\. Caccia, R\. Aljundi, N\. Asadi, T\. Tuytelaars, J\. Pineau, and E\. Belilovsky\(2022\)New insights on reducing abrupt representation change in online continual learning\.InInternational Conference on Learning Representations,Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[5\]A\. Chaudhry, M\. Ranzato, M\. Rohrbach, and M\. Elhoseiny\(2018\)Efficient lifelong learning with a\-gem\.arXiv preprint arXiv:1812\.00420\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1),[§V\-B](https://arxiv.org/html/2608.11690#S5.SS2.p7.1)\.
- \[6\]Q\. Chen, C\. Shui, L\. Han, and M\. Marchand\(2023\)On the stability\-plasticity dilemma in continual meta\-learning: theory and algorithm\.Advances in Neural Information Processing Systems36,pp\. 27414–27468\.Cited by:[§V](https://arxiv.org/html/2608.11690#S5.p1.1)\.
- \[7\]A\. Chrysakis and M\. Moens\(2020\)Online continual learning from imbalanced data\.InInternational Conference on Machine Learning,pp\. 1952–1961\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[8\]M\. De Lange, R\. Aljundi, M\. Masana, S\. Parisot, X\. Jia, A\. Leonardis, G\. Slabaugh, and T\. Tuytelaars\(2021\)A continual learning survey: defying forgetting in classification tasks\.IEEE transactions on pattern analysis and machine intelligence44\(7\),pp\. 3366–3385\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p1.1)\.
- \[9\]M\. Ding, K\. Ji, D\. Wang, and J\. Xu\(2024\)Understanding forgetting in continual learning with linear regression\.arXiv preprint arXiv:2405\.17583\.Cited by:[§II\-B](https://arxiv.org/html/2608.11690#S2.SS2.p1.1)\.
- \[10\]Y\. Dong, T\. Gong, H\. Chen, Z\. He, M\. Li, S\. Song, and C\. Li\(2024\)Towards generalization beyond pointwise learning: a unified information\-theoretic perspective\.InForty\-first International Conference on Machine Learning,Cited by:[Lemma A\.9](https://arxiv.org/html/2608.11690#A1.Thmtheorem9)\.
- \[11\]Y\. Dong, T\. Gong, H\. Chen, and C\. Li\(2023\)Understanding the generalization ability of deep learning algorithms: a kernelized rényi’s entropy perspective\.arXiv preprint arXiv:2305\.01143\.Cited by:[§II\-B](https://arxiv.org/html/2608.11690#S2.SS2.p1.1)\.
- \[12\]A\. Douillard, M\. Cord, C\. Ollion, T\. Robert, and E\. Valle\(2020\)Podnet: pooled outputs distillation for small\-tasks incremental learning\.InEuropean Conference on Computer Vision,pp\. 86–102\.Cited by:[§IV\-A](https://arxiv.org/html/2608.11690#S4.SS1.p7.1)\.
- \[13\]R\. M\. French\(1999\)Catastrophic forgetting in connectionist networks\.Trends in cognitive sciences3\(4\),pp\. 128–135\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[14\]I\. J\. Goodfellow, M\. Mirza, D\. Xiao, A\. Courville, and Y\. Bengio\(2013\)An empirical investigation of catastrophic forgetting in gradient\-based neural networks\.ArXiv Preprint ArXiv:1312\.6211\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p1.1)\.
- \[15\]R\. M\. Gray\(2011\)Entropy and information theory\.Springer Science & Business Media\.Cited by:[Lemma A\.6](https://arxiv.org/html/2608.11690#A1.Thmtheorem6)\.
- \[16\]Y\. Guo, B\. Liu, and D\. Zhao\(2022\)Online continual learning through mutual information maximization\.InInternational Conference on Machine Learning,pp\. 8109–8126\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p1.1)\.
- \[17\]M\. Haghifam, J\. Negrea, A\. Khisti, D\. M\. Roy, and G\. K\. Dziugaite\(2020\)Sharpened generalization bounds based on conditional mutual information and an application to noisy, iterative algorithms\.Advances in Neural Information Processing Systems33,pp\. 9925–9935\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p4.1),[§II\-B](https://arxiv.org/html/2608.11690#S2.SS2.p1.1)\.
- \[18\]H\. Harutyunyan, M\. Raginsky, G\. Ver Steeg, and A\. Galstyan\(2021\)Information\-theoretic generalization bounds for black\-box learning algorithms\.Advances in Neural Information Processing Systems34,pp\. 24670–24682\.Cited by:[Lemma A\.7](https://arxiv.org/html/2608.11690#A1.Thmtheorem7)\.
- \[19\]H\. He, C\. L\. Yu, and Z\. Goldfeld\(2024\)Hierarchical generalization bounds for deep neural networks\.In2024 IEEE International Symposium on Information Theory \(ISIT\),pp\. 2688–2693\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p4.1),[§II\-B](https://arxiv.org/html/2608.11690#S2.SS2.p1.1)\.
- \[20\]T\. Kao, K\. Jensen, G\. van de Ven, A\. Bernacchia, and G\. Hennequin\(2021\)Natural continual learning: success is a journey, not \(just\) a destination\.Advances in Neural Information Processing Systems34,pp\. 28067–28079\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p1.1)\.
- \[21\]J\. Kirkpatrick, R\. Pascanu, N\. Rabinowitz, J\. Veness, G\. Desjardins, A\. A\. Rusu, K\. Milan, J\. Quan, T\. Ramalho, A\. Grabska\-Barwinska,et al\.\(2017\)Overcoming catastrophic forgetting in neural networks\.Proceedings of the national academy of sciences114\(13\),pp\. 3521–3526\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[22\]J\. Knoblauch, H\. Husain, and T\. Diethe\(2020\)Optimal continual learning has perfect memory and is np\-hard\.InInternational Conference on Machine Learning,pp\. 5327–5337\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p4.1)\.
- \[23\]Y\. Li and E\. van Kampen\(2024\)Stability analysis for incremental adaptive dynamic programming with approximation errors\.Journal of Aerospace Engineering37\(1\),pp\. 04023097\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p4.1)\.
- \[24\]Y\. Liang and W\. Li\(2023\)Adaptive plasticity improvement for continual learning\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 7816–7825\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p1.1)\.
- \[25\]S\. Lin, P\. Ju, Y\. Liang, and N\. Shroff\(2023\)Theory on forgetting and generalization of continual learning\.InInternational Conference on Machine Learning,pp\. 21078–21100\.Cited by:[§II\-B](https://arxiv.org/html/2608.11690#S2.SS2.p1.1)\.
- \[26\]A\. T\. Lopez and V\. Jog\(2018\)Generalization error bounds using wasserstein distances\.In2018 ieee information theory workshop \(ITW\),pp\. 1–5\.Cited by:[§II\-C](https://arxiv.org/html/2608.11690#S2.SS3.p1.1)\.
- \[27\]D\. Lopez\-Paz and M\. Ranzato\(2017\)Gradient episodic memory for continual learning\.Advances in neural information processing systems30\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1),[§V\-B](https://arxiv.org/html/2608.11690#S5.SS2.p7.1)\.
- \[28\]A\. Mallya and S\. Lazebnik\(2018\)Packnet: adding multiple tasks to a single network by iterative pruning\.InProceedings of the IEEE conference on Computer Vision and Pattern Recognition,pp\. 7765–7773\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[29\]M\. McCloskey and N\. J\. Cohen\(1989\)Catastrophic interference in connectionist networks: the sequential learning problem\.InPsychology of learning and motivation,Vol\.24,pp\. 109–165\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[30\]J\. Negrea, M\. Haghifam, G\. K\. Dziugaite, A\. Khisti, and D\. M\. Roy\(2019\)Information\-theoretic generalization bounds for sgld via data\-dependent estimates\.Advances in Neural Information Processing Systems32\.Cited by:[§II\-B](https://arxiv.org/html/2608.11690#S2.SS2.p1.1),[§V\-A](https://arxiv.org/html/2608.11690#S5.SS1.p5.1)\.
- \[31\]G\. I\. Parisi, R\. Kemker, J\. L\. Part, C\. Kanan, and S\. Wermter\(2019\)Continual lifelong learning with neural networks: a review\.Neural networks113,pp\. 54–71\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[32\]G\. Peyré M\. Cuturiet al\.\(2019\)Computational optimal transport: with applications to data science\.Foundations and Trends® in Machine Learning11\(5\-6\),pp\. 355–607\.Cited by:[§II\-C](https://arxiv.org/html/2608.11690#S2.SS3.p1.1)\.
- \[33\]S\. Rebuffi, A\. Kolesnikov, G\. Sperl, and C\. H\. Lampert\(2017\)Icarl: incremental classifier and representation learning\.InProceedings of the IEEE conference on Computer Vision and Pattern Recognition,pp\. 2001–2010\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1),[§VI\-A](https://arxiv.org/html/2608.11690#S6.SS1.SSS0.Px1.p1.1)\.
- \[34\]I\. Redko, A\. Habrard, and M\. Sebban\(2017\)Theoretical analysis of domain adaptation with optimal transport\.InJoint European Conference on Machine Learning and Knowledge Discovery in Databases,pp\. 737–753\.Cited by:[§II\-C](https://arxiv.org/html/2608.11690#S2.SS3.p1.1)\.
- \[35\]M\. Riemer, I\. Cases, R\. Ajemian, M\. Liu, I\. Rish, Y\. Tu, and G\. Tesauro\(2019\)Learning to learn without forgetting by maximizing transfer and minimizing interference\.InInternational Conference on Learning Representations,Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p4.1)\.
- \[36\]D\. Rolnick, A\. Ahuja, J\. Schwarz, T\. Lillicrap, and G\. Wayne\(2019\)Experience replay for continual learning\.Advances in neural information processing systems32\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1),[§VI\-A](https://arxiv.org/html/2608.11690#S6.SS1.SSS0.Px1.p1.1)\.
- \[37\]A\. A\. Rusu, N\. C\. Rabinowitz, G\. Desjardins, H\. Soyer, J\. Kirkpatrick, K\. Kavukcuoglu, R\. Pascanu, and R\. Hadsell\(2016\)Progressive neural networks\.arXiv preprint arXiv:1606\.04671\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[38\]G\. Saha, I\. Garg, and K\. Roy\(2021\)Gradient projection memory for continual learning\.arXiv preprint arXiv:2103\.09762\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1),[§V\-B](https://arxiv.org/html/2608.11690#S5.SS2.p7.1)\.
- \[39\]Y\. Shen, S\. Dasgupta, and S\. Navlakha\(2023\)Reducing catastrophic forgetting with associative learning: a lesson from fruit flies\.Neural Computation35\(11\),pp\. 1797–1819\.Cited by:[§IV\-B](https://arxiv.org/html/2608.11690#S4.SS2.p4.1)\.
- \[40\]H\. Shi and H\. Wang\(2023\)A unified approach to domain incremental learning with memory: theory and algorithm\.Advances in Neural Information Processing Systems36,pp\. 15027–15059\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p4.1)\.
- \[41\]H\. Shin, J\. K\. Lee, J\. Kim, and J\. Kim\(2017\)Continual learning with deep generative replay\.Advances in neural information processing systems30\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[42\]S\. Sonoda, Y\. Hashimoto, I\. Ishikawa, and M\. Ikeda\(2025\)Why and when deep is better than shallow: an implementation\-agnostic state\-transition view of depth supremacy\.External Links:2505\.15064,[Link](https://arxiv.org/abs/2505.15064)Cited by:[§II\-B](https://arxiv.org/html/2608.11690#S2.SS2.p1.1)\.
- \[43\]A\. Sorrenti, G\. Bellitto, F\. P\. Salanitri, M\. Pennisi, C\. Spampinato, and S\. Palazzo\(2023\)Selective freezing for efficient continual learning\.In2023 IEEE/CVF International Conference on Computer Vision Workshops \(ICCVW\),pp\. 3542–3551\.Cited by:[§IV\-B](https://arxiv.org/html/2608.11690#S4.SS2.p4.1)\.
- \[44\]T\. Steinke and L\. Zakynthinou\(2020\)Reasoning about generalization via conditional mutual information\.InConference on Learning Theory,pp\. 3437–3452\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p4.1),[§II\-B](https://arxiv.org/html/2608.11690#S2.SS2.p1.1),[§IV\-A](https://arxiv.org/html/2608.11690#S4.SS1.p5.1)\.
- \[45\]S\. Sun, D\. Calandriello, H\. Hu, A\. Li, and M\. Titsias\(2022\)Information\-theoretic online memory selection for continual learning\.arXiv preprint arXiv:2204\.04763\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[46\]G\. M\. van de Ven, N\. Soures, and D\. Kudithipudi\(2024\)Continual learning and catastrophic forgetting\.arXiv preprint arXiv:2403\.05175\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.
- \[47\]H\. Wang, M\. Diaz, J\. C\. S\. Santos Filho, and F\. P\. Calmon\(2019\)An information\-theoretic view of generalization via wasserstein distance\.In2019 IEEE international symposium on information theory \(ISIT\),pp\. 577–581\.Cited by:[§II\-C](https://arxiv.org/html/2608.11690#S2.SS3.p1.1)\.
- \[48\]L\. Wang, X\. Zhang, H\. Su, and J\. Zhu\(2024\)A comprehensive survey of continual learning: theory, method and application\.IEEE transactions on pattern analysis and machine intelligence46\(8\),pp\. 5362–5383\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1),[§IV\-A](https://arxiv.org/html/2608.11690#S4.SS1.p1.1)\.
- \[49\]W\. Wen, T\. Gong, Z\. Gao, Y\. Zhang, W\. Zhang, and Y\. Liu\(2026\)Information\-theoretic generalization bounds of replay\-based continual learning\.External Links:2507\.12043,[Link](https://arxiv.org/abs/2507.12043)Cited by:[§II\-B](https://arxiv.org/html/2608.11690#S2.SS2.p1.1)\.
- \[50\]A\. Xu and M\. Raginsky\(2017\)Information\-theoretic analysis of generalization capability of learning algorithms\.Advances in neural information processing systems30\.Cited by:[§I](https://arxiv.org/html/2608.11690#S1.p4.1),[§II\-B](https://arxiv.org/html/2608.11690#S2.SS2.p1.1),[§IV\-A](https://arxiv.org/html/2608.11690#S4.SS1.p5.1)\.
- \[51\]F\. Zenke, B\. Poole, and S\. Ganguli\(2017\)Continual learning through synaptic intelligence\.InInternational conference on machine learning,pp\. 3987–3995\.Cited by:[§II\-A](https://arxiv.org/html/2608.11690#S2.SS1.p1.1)\.

## Appendix APrerequisite Definitions and Lemmas

###### Definition A\.1\(σ\\sigma\-sub\-Gaussian\)\.

A random variableXXisσ\\sigma\-sub\-Gaussian if for anyλ\\lambda,ln⁡𝔼⁡\[eλ⁡\(X−𝔼​X\)\]≤λ2​σ2/2\\ln\\mathbb\{E\}\[e^\{\\lambda\(X\-\\mathbb\{E\}X\)\}\]\\leq\\lambda^\{2\}\\sigma^\{2\}/2\.

###### Definition A\.2\(Kullback\-Leibler Divergence\)\.

LetPPandQQbe probability measures on a measurable space\(𝒳,Σ\)\(\\mathcal\{X\},\\Sigma\)\. IfPPis absolutely continuous with respect toQQ\(denotedP≪QP\\ll Q\), the Kullback\-Leibler \(KL\) divergence fromPPtoQQis defined as:DKL\(P∥Q\)≜∫𝒳log\(d​Pd​Q\)dPD\_\{\\mathrm\{KL\}\}\(P\\\|Q\)\\triangleq\\int\_\{\\mathcal\{X\}\}\\log\\left\(\\frac\{\\mathrm\{d\}P\}\{\\mathrm\{d\}Q\}\\right\)\\mathrm\{d\}P, whered​Pd​Q\\frac\{\\mathrm\{d\}P\}\{\\mathrm\{d\}Q\}denotes the Radon\-Nikodym derivative\.

###### Definition A\.3\(Mutual Information\)\.

LetXXandYYbe random variables taking values in measurable spaces𝒳\\mathcal\{X\}and𝒴\\mathcal\{Y\}, respectively\. LetPX,YP\_\{X,Y\}denote their joint distribution on the product space𝒳×𝒴\\mathcal\{X\}\\times\\mathcal\{Y\}, and letPXP\_\{X\}andPYP\_\{Y\}denote the corresponding marginal distributions\. The mutual information betweenXXandYYis defined as the KL divergence between the joint distribution and the product of the marginals:I\(X;Y\)≜DKL\(PX,Y∥PX⊗PY\)\.I\(X;Y\)\\triangleq D\_\{\\mathrm\{KL\}\}\(P\_\{X,Y\}\\\|P\_\{X\}\\otimes P\_\{Y\}\)\.

###### Lemma A\.4\(Interaction information identity\)\.

For any random variablesA,B,C,DA,B,C,D,

I⁡\(A,B;C∣D\)=I⁡\(A;C∣D\)\+I⁡\(B;C∣D\)−I⁡\(A;B;C∣D\)\.I\(A,B;C\\mid D\)=I\(A;C\\mid D\)\+I\(B;C\\mid D\)\-I\(A;B;C\\mid D\)\.

###### Definition A\.5\(Wasserstein Distance\)\.

LetPPandQQbe probability measures on𝒳⊆ℝk\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{k\}\. The Wasserstein\-ppdistance \(p≥1p\\geq 1\) is defined as:𝒲p​\(P,Q\)≜\(infγ∈Γ⁡\(P,Q\)𝔼\(x,x′\)∼γ​\[‖x−x′‖p\]\)1/p,\\mathcal\{W\}\_\{p\}\(P,Q\)\\triangleq\\left\(\\inf\_\{\\gamma\\in\\Gamma\(P,Q\)\}\\mathbb\{E\}\_\{\(x,x^\{\\prime\}\)\\sim\\gamma\}\\left\[\\\|x\-x^\{\\prime\}\\\|^\{p\}\\right\]\\right\)^\{1/p\},whereΓ⁡\(P,Q\)\\Gamma\(P,Q\)is the set of all joint distributions on𝒳×𝒳\\mathcal\{X\}\\times\\mathcal\{X\}with marginalsPPandQQ, and∥⋅∥\\\|\\cdot\\\|denotes the Euclidean norm\.

###### Lemma A\.6\(Donsker\-Varadhan formula\(Theorem 5\.2\.1 in\[[15](https://arxiv.org/html/2608.11690#bib.bib47)\]\)\)\.

LetPPandQQbe probability measures on𝒳\\mathcal\{X\}such thatP≪QP\\ll Q\. Then:

DKL\(P∥Q\)=supf∈L∞​\(𝒳\)\{𝔼P\[f\]−log𝔼Q\[ef\]\},D\_\{\\mathrm\{KL\}\}\(P\\\|Q\)=\\sup\_\{f\\in L\_\{\\infty\}\(\\mathcal\{X\}\)\}\\left\\\{\\mathbb\{E\}\_\{P\}\[f\]\-\\log\\mathbb\{E\}\_\{Q\}\\left\[e^\{f\}\\right\]\\right\\\},where the supremum is taken over the set of all bounded measurable functionsf:𝒳→ℝf:\\mathcal\{X\}\\to\\mathbb\{R\}\. IfXXisσ\\sigma\-sub\-Gaussian underQQ, then

\|𝔼P​\[X\]−𝔼Q​\[X\]\|≤σ​2DKL\(P∥Q\)\.\\left\|\\mathbb\{E\}\_\{P\}\[X\]\-\\mathbb\{E\}\_\{Q\}\[X\]\\right\|\\leq\\sigma\\sqrt\{2D\_\{\\mathrm\{KL\}\}\(P\\\|Q\)\}\.

###### Lemma A\.7\(Lemma 1 in\[[18](https://arxiv.org/html/2608.11690#bib.bib48)\]\)\.

Let\(X,Y\)∼PX​Y\(X,Y\)\\sim P\_\{XY\}and letY¯\\bar\{Y\}be an independent copy ofYYsuch that\(X,Y¯\)∼PX⊗PY\(X,\\bar\{Y\}\)\\sim P\_\{X\}\\otimes P\_\{Y\}\. If the functionf⁡\(X,Y¯\)f\(X,\\bar\{Y\}\)isσ\\sigma\-sub\-gaussian with respect to the distributionPX⊗PYP\_\{X\}\\otimes P\_\{Y\}, then:

\|𝔼⁡\[f⁡\(X,Y\)\]−𝔼⁡\[f⁡\(X,Y¯\)\]\|≤2​σ2​I​\(X,Y\)\.\\left\|\\mathbb\{E\}\[f\(X,Y\)\]\-\\mathbb\{E\}\[f\(X,\\bar\{Y\}\)\]\\right\|\\leq\\sqrt\{2\\sigma^\{2\}I\(X;Y\)\}\.

###### Lemma A\.8\(Kantorovich\-Rubinstein Duality\)\.

LetPPandQQbe probability measures defined on a metric space\(𝒳,c\)\(\\mathcal\{X\},c\)\. Then, the11\-Wasserstein distance is given by

𝒲1​\(P,Q\)=supf∈Lip1\{∫𝒳f​𝑑P−∫𝒳f​𝑑Q\},\\mathcal\{W\}\_\{1\}\(P,Q\)=\\sup\_\{f\\in\\mathrm\{Lip\}\_\{1\}\}\\left\\\{\\int\_\{\\mathcal\{X\}\}f\\,\\mathrm\{d\}P\-\\int\_\{\\mathcal\{X\}\}f\\,\\mathrm\{d\}Q\\right\\\},whereLip1\\mathrm\{Lip\}\_\{1\}denotes the set of 1\-Lipschitz functions with respect to the metriccc, i\.e\.,\|f⁡\(x\)−f⁡\(x′\)\|≤c⁡\(x,x′\)\|f\(x\)\-f\(x^\{\\prime\}\)\|\\leq c\(x,x^\{\\prime\}\)for anyf∈Lip1f\\in\\mathrm\{Lip\}\_\{1\}andx,x′∈𝒳x,x^\{\\prime\}\\in\\mathcal\{X\}\.

###### Lemma A\.9\(Lemma A\.11 in\[[10](https://arxiv.org/html/2608.11690#bib.bib49)\]\)\.

LetX∼𝒩⁡\(0,Σ\)X\\sim\\mathcal\{N\}\(0,\\Sigma\)and letYYbe any zero\-mean random vector withCov⁡\[Y\]=Σ\\mathrm\{Cov\}\[Y\]=\\Sigma\. Thenh⁡\(Y\)≤h⁡\(X\)h\(Y\)\\leq h\(X\)\.

###### Lemma A\.10\(Entropy upper bound under second\-moment constraint\)\.

LetX∈ℝdX\\in\\mathbb\{R\}^\{d\}be any random vector with finite second moment matrixM:=𝔼⁡\[X​X⊤\]∈ℝd×dM:=\\mathbb\{E\}\[XX^\{\\top\}\]\\in\\mathbb\{R\}^\{d\\times d\}\. Then the differential entropy satisfies

h⁡\(X\)≤12​log⁡\(\(2​π​e\)d​det\(M\)\)\.h\(X\)\\leq\\frac\{1\}\{2\}\\log\\Big\(\(2\\pi e\)^\{d\}\\det\(M\)\\Big\)\.

###### Proof\.

Letμ:=𝔼⁡\[X\]\\mu:=\\mathbb\{E\}\[X\]andX~:=X−μ\\widetilde\{X\}:=X\-\\mu\. Since entropy is translation\-invariant,h⁡\(X\)=h⁡\(X~\)h\(X\)=h\(\\widetilde\{X\}\)\. LetΣ:=Cov⁡\(X\)=𝔼⁡\[X~​X~⊤\]\\Sigma:=\\mathrm\{Cov\}\(X\)=\\mathbb\{E\}\[\\widetilde\{X\}\\widetilde\{X\}^\{\\top\}\]\. By Lemma[A\.9](https://arxiv.org/html/2608.11690#A1.Thmtheorem9), we have

h⁡\(X~\)≤12​log⁡\(\(2​π​e\)d​det\(Σ\)\)\.h\(\\widetilde\{X\}\)\\leq\\frac\{1\}\{2\}\\log\\Big\(\(2\\pi e\)^\{d\}\\det\(\\Sigma\)\\Big\)\.Next note that

M−Σ=𝔼⁡\[X​X⊤\]−𝔼⁡\[X~​X~⊤\]=μ​μ⊤⪰0,M\-\\Sigma=\\mathbb\{E\}\[XX^\{\\top\}\]\-\\mathbb\{E\}\[\\widetilde\{X\}\\widetilde\{X\}^\{\\top\}\]=\\mu\\mu^\{\\top\}\\succeq 0,henceM⪰Σ⪰0M\\succeq\\Sigma\\succeq 0\. Since determinant is monotone under the Loewner order on PSD matrices,det\(M\)≥det\(Σ\)\\det\(M\)\\geq\\det\(\\Sigma\), so

h⁡\(X\)=h⁡\(X~\)≤12​log⁡\(\(2​π​e\)d​det\(Σ\)\)≤12​log⁡\(\(2​π​e\)d​det\(M\)\)\.h\(X\)=h\(\\widetilde\{X\}\)\\leq\\frac\{1\}\{2\}\\log\\Big\(\(2\\pi e\)^\{d\}\\det\(\\Sigma\)\\Big\)\\leq\\frac\{1\}\{2\}\\log\\Big\(\(2\\pi e\)^\{d\}\\det\(M\)\\Big\)\.∎

## Appendix BOmitted Proofs

### B\-AProof of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)

###### Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)\(Restate\)\.

LetWWbe the output of a learner minimizing the empirical riskℒ^​\(W\)\\hat\{\\mathcal\{L\}\}\(W\)\. Assume the loss functionℓ⁡\(W,Z\)\\ell\(W,Z\)isσ\\sigma\-subgaussian for allZ∈𝒵Z\\in\\mathcal\{Z\}andW∈𝒲W\\in\\mathcal\{W\}\. For any split layerl∈\{0,…,L\}l\\in\\\{0,\\dots,L\\\}, we have

\|genW\|≤\(T−1\)​2​σ2​𝒦\(l\)⏟Replay\-centroid Drift\+2​σ2Neff​\(𝒮\(l\)\+𝒫\(l\)−ℛ\(l\)\+𝒞\(l\)\)⏟Optimization Variance,\|\\mathrm\{gen\}\_\{W\}\|\\leq\\underbrace\{\(T\-1\)\\sqrt\{2\\sigma^\{2\}\\,\\mathcal\{K\}^\{\(l\)\}\}\}\_\{\\text\{\\bf Replay\-centroid Drift\}\}\+\\underbrace\{\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\\Big\(\\mathcal\{S\}^\{\(l\)\}\+\\mathcal\{P\}^\{\(l\)\}\-\\mathcal\{R\}^\{\(l\)\}\+\\mathcal\{C\}^\{\(l\)\}\\Big\)\}\}\_\{\\text\{\\bf Optimization Variance\}\},whereNeff≜1\(T−1\)/m\+1/nN\_\{\\mathrm\{eff\}\}\\triangleq\\frac\{1\}\{\(T\-1\)/m\+1/n\}is the effective sample size and𝒦\(l\)≜1T−1∑i=1T−1𝔼W1:l\[DKL\(QAl,Y\|i,W1:l∥PAl,Y\|i,W1:l\)\]\\mathcal\{K\}^\{\(l\)\}\\triangleq\\frac\{1\}\{T\-1\}\\sum\_\{i=1\}^\{T\-1\}\\mathbb\{E\}\_\{W\_\{1:l\}\}\\left\[D\_\{\\mathrm\{KL\}\}\\Big\(Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\big\\\|P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\Big\)\\right\], and𝒞\(l\)≜𝔼W1:l\[DKL\(P𝐔\(l\)∣W1:l∥Q~𝐔\(l\)\|W1:l\)\]\\mathcal\{C\}^\{\(l\)\}\\triangleq\\mathbb\{E\}\_\{W\_\{1:l\}\}\\left\[D\_\{\\mathrm\{KL\}\}\\Big\(P\_\{\\mathbf\{U\}^\{\(l\)\}\\mid W\_\{1:l\}\}\\big\\\|\\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(l\)\}\|W\_\{1:l\}\}\\Big\)\\right\]\.

###### Proof\.

Fix a layerl∈\{0,…,L\}l\\in\\\{0,\\dots,L\\\}\. Recall thata0​\(x\)=xa\_\{0\}\(x\)=xand, forh∈\{1,…,L\}h\\in\\\{1,\\dots,L\\\},ah​\(x\)=ϕh​\(Wh​ah−1​\(x\)\)a\_\{h\}\(x\)=\\phi\_\{h\}\(W\_\{h\}a\_\{h\-1\}\(x\)\)\. We also adopt the boundary convention thatgWL\+1:Lg\_\{W\_\{L\+1:L\}\}is the identity map whenl=Ll=L\. To make the layer\-llrepresentation parametrization unambiguous, we explicitly decompose the network at layerll\. Define the top sub\-network mapping from the layer\-llrepresentation to logits as

gWl\+1:L:ℝdl→ℝK,fW\(x\)=gWl\+1:L\(al\(x\)\)\.g\_\{W\_\{l\+1:L\}\}:\\mathbb\{R\}^\{d\_\{l\}\}\\to\\mathbb\{R\}^\{K\},\\qquad f\_\{W\}\(x\)=g\_\{W\_\{l\+1:L\}\}\(a\_\{l\}\(x\)\)\.Accordingly, define the loss as a function of the layer\-llrepresentation–label pair:

FW\(t,y\)≜ℓ\(gWl\+1:L\(t\),y\),so thatℓ\(W,\(X,Y\)\)=FW\(al\(X\),Y\)\.F\_\{W\}\(t,y\)\\triangleq\\ell\\big\(g\_\{W\_\{l\+1:L\}\}\(t\),y\\big\),\\qquad\\text\{so that\}\\qquad\\ell\(W,\(X,Y\)\)=F\_\{W\}\(a\_\{l\}\(X\),Y\)\.By definition,

genW=𝔼W,D1:T\[ℒ\(W\)−ℒ^\(W\)\]\.\\mathrm\{gen\}\_\{W\}=\\mathbb\{E\}\_\{W,\\,D^\{1:T\}\}\[\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)\]\.Expandℒ⁡\(W\)\\mathcal\{L\}\(W\)andℒ^​\(W\)\\hat\{\\mathcal\{L\}\}\(W\):

ℒ⁡\(W\)=∑t=1T𝔼Z∼𝒟t​\[ℓ⁡\(W,Z\)\],ℒ^​\(W\)=∑i=1T−11m​∑j∈ℐiℓ⁡\(W,Zji\)\+1n​∑j=1nℓ⁡\(W,ZjT\)\.\\mathcal\{L\}\(W\)=\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{t\}\}\[\\ell\(W,Z\)\],\\qquad\\hat\{\\mathcal\{L\}\}\(W\)=\\sum\_\{i=1\}^\{T\-1\}\\frac\{1\}\{m\}\\sum\_\{j\\in\\mathcal\{I\}\_\{i\}\}\\ell\(W,Z\_\{j\}^\{i\}\)\+\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}\\ell\(W,Z\_\{j\}^\{T\}\)\.Therefore,

ℒ​\(W\)−ℒ^​\(W\)\\displaystyle\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)=∑i=1T−1\(𝔼Z∼𝒟i​\[ℓ⁡\(W,Z\)\]−1m​∑j∈ℐiℓ⁡\(W,Zji\)\)\+\(𝔼Z∼𝒟T​\[ℓ⁡\(W,Z\)\]−1n​∑j=1nℓ⁡\(W,ZjT\)\)\.\\displaystyle=\\sum\_\{i=1\}^\{T\-1\}\\left\(\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{i\}\}\[\\ell\(W,Z\)\]\-\\frac\{1\}\{m\}\\sum\_\{j\\in\\mathcal\{I\}\_\{i\}\}\\ell\(W,Z\_\{j\}^\{i\}\)\\right\)\\,\+\\,\\left\(\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{T\}\}\[\\ell\(W,Z\)\]\-\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}\\ell\(W,Z\_\{j\}^\{T\}\)\\right\)\.\(8\)For each old taski∈\[T−1\]i\\in\[T\-1\], define the layer\-llpopulation risk and replay empirical risk

Li\(W\)≜𝔼\(Al,Y\)∼PAl,Y\|i,W1:l\[FW\(Al,Y\)\],L^i\(W\)≜1m∑j∈ℐiFW\(Aj,li,Yji\)\.L\_\{i\}\(W\)\\triangleq\\mathbb\{E\}\_\{\(A\_\{l\},Y\)\\sim P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\}\\\!\\big\[F\_\{W\}\(A\_\{l\},Y\)\\big\],\\qquad\\widehat\{L\}\_\{i\}\(W\)\\triangleq\\frac\{1\}\{m\}\\sum\_\{j\\in\\mathcal\{I\}\_\{i\}\}F\_\{W\}\(A\_\{j,l\}^\{i\},Y\_\{j\}^\{i\}\)\.Then the old\-task block in \([8](https://arxiv.org/html/2608.11690#A2.E8)\) can be written as

Δold\\displaystyle\\Delta\_\{\\mathrm\{old\}\}≜∑i=1T−1\(𝔼Z∼𝒟i​\[ℓ⁡\(W,Z\)\]−1m​∑j∈ℐiℓ⁡\(W,Zji\)\)\\displaystyle\\triangleq\\sum\_\{i=1\}^\{T\-1\}\\left\(\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{i\}\}\[\\ell\(W,Z\)\]\-\\frac\{1\}\{m\}\\sum\_\{j\\in\\mathcal\{I\}\_\{i\}\}\\ell\(W,Z\_\{j\}^\{i\}\)\\right\)=∑i=1T−1\(Li​\(W\)−L^i​\(W\)\)\.\\displaystyle=\\sum\_\{i=1\}^\{T\-1\}\\Big\(L\_\{i\}\(W\)\-\\widehat\{L\}\_\{i\}\(W\)\\Big\)\.\(9\)For each old taski∈\[T−1\]i\\in\[T\-1\], recall the task\-wise replay proxyP^Al,Y\|Si,W1:l\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\}and its conditional centroid law

QAl,Y\|i,W1:l≜𝔼\[P^Al,Y\|Si,W1:l\|W1:l\]\.Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\;\\triangleq\\;\\mathbb\{E\}\\\!\\left\[\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\}\\ \\middle\|\\ W\_\{1:l\}\\right\]\.\(10\)Add and subtract𝔼\(Al,Y\)∼QAl,Y\|i,W1:l\[FW\(Al,Y\)\]\\mathbb\{E\}\_\{\(A\_\{l\},Y\)\\sim Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\}\[F\_\{W\}\(A\_\{l\},Y\)\]inside eachLi​\(W\)−L^i​\(W\)L\_\{i\}\(W\)\-\\widehat\{L\}\_\{i\}\(W\)in \([9](https://arxiv.org/html/2608.11690#A2.E9)\) to obtain

Δold=∑i=1T−1\(𝔼PAl,Y\|i,W1:l\[FW\]−𝔼QAl,Y\|i,W1:l\[FW\]\)⏟≜Δdrift\(l\)\+∑i=1T−1\(𝔼QAl,Y\|i,W1:l\[FW\]−L^i\(W\)\)⏟≜Δold,var\(l\)\.\\Delta\_\{\\mathrm\{old\}\}=\\underbrace\{\\sum\_\{i=1\}^\{T\-1\}\\Big\(\\mathbb\{E\}\_\{P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\}\[F\_\{W\}\]\-\\mathbb\{E\}\_\{Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\}\[F\_\{W\}\]\\Big\)\}\_\{\\triangleq\\ \\Delta\_\{\\mathrm\{drift\}\}^\{\(l\)\}\}\+\\underbrace\{\\sum\_\{i=1\}^\{T\-1\}\\Big\(\\mathbb\{E\}\_\{Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\}\[F\_\{W\}\]\-\\widehat\{L\}\_\{i\}\(W\)\\Big\)\}\_\{\\triangleq\\ \\Delta\_\{\\mathrm\{old,var\}\}^\{\(l\)\}\}\.\(11\)For the new task block in \([8](https://arxiv.org/html/2608.11690#A2.E8)\), define

Δnew,var\(l\)≜\(𝔼Z∼𝒟T\[ℓ\(W,Z\)\]−1n∑j=1nℓ\(W,ZjT\)\)=\(𝔼PAl,Y\|T,W1:l\[FW\(Al,Y\)\]−1n∑j=1nFW\(Aj,lT,YjT\)\),\\Delta\_\{\\mathrm\{new,var\}\}^\{\(l\)\}\\triangleq\\left\(\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{T\}\}\[\\ell\(W,Z\)\]\-\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}\\ell\(W,Z\_\{j\}^\{T\}\)\\right\)=\\left\(\\mathbb\{E\}\_\{P\_\{A\_\{l\},Y\|T,W\_\{1:l\}\}\}\[F\_\{W\}\(A\_\{l\},Y\)\]\-\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}F\_\{W\}\(A\_\{j,l\}^\{T\},Y\_\{j\}^\{T\}\)\\right\),wherePAl,Y\|T,W1:lP\_\{A\_\{l\},Y\|T,W\_\{1:l\}\}is the law of\(al​\(X\),Y\)\(a\_\{l\}\(X\),Y\)underZ=\(X,Y\)∼𝒟TZ=\(X,Y\)\\sim\\mathcal\{D\}\_\{T\}\. Combining \([8](https://arxiv.org/html/2608.11690#A2.E8)\) and \([11](https://arxiv.org/html/2608.11690#A2.E11)\) yields the exact decomposition

ℒ⁡\(W\)−ℒ^​\(W\)=Δdrift\(l\)\+Δold,var\(l\)\+Δnew,var\(l\)\.\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)=\\Delta\_\{\\mathrm\{drift\}\}^\{\(l\)\}\+\\Delta\_\{\\mathrm\{old,var\}\}^\{\(l\)\}\+\\Delta\_\{\\mathrm\{new,var\}\}^\{\(l\)\}\.\(12\)LetΔvar\(l\)≜Δold,var\(l\)\+Δnew,var\(l\)\\Delta\_\{\\mathrm\{var\}\}^\{\(l\)\}\\triangleq\\Delta\_\{\\mathrm\{old,var\}\}^\{\(l\)\}\+\\Delta\_\{\\mathrm\{new,var\}\}^\{\(l\)\}\. Taking expectation and using\|a\+b\|≤\|a\|\+\|b\|\|a\+b\|\\leq\|a\|\+\|b\|gives

\|genW\|=\|𝔼⁡\[ℒ⁡\(W\)−ℒ^​\(W\)\]\|≤𝔼⁡\[\|Δdrift\(l\)\|\]\+\|𝔼⁡\[Δvar\(l\)\]\|\.\|\\mathrm\{gen\}\_\{W\}\|=\\left\|\\mathbb\{E\}\[\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)\]\\right\|\\leq\\mathbb\{E\}\\big\[\|\\Delta\_\{\\mathrm\{drift\}\}^\{\(l\)\}\|\\big\]\+\|\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(l\)\}\\big\]\|\.\(13\)Fixi∈\[T−1\]i\\in\[T\-1\]and condition onW1:lW\_\{1:l\}\. LetQi≜PAl,Y\|i,W1:lQ\_\{i\}\\triangleq P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}andQ~i≜QAl,Y\|i,W1:l\\tilde\{Q\}\_\{i\}\\triangleq Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\. If the loss functionℓ⁡\(W,\(X,Y\)\)\\ell\(W,\(X,Y\)\)isσ\\sigma\-sub\-Gaussian for allW∈𝒲W\\in\\mathcal\{W\}andZ∈𝒵Z\\in\\mathcal\{Z\}, it follows that for any fixedWW, the random variableFW​\(Al,Y\)F\_\{W\}\(A\_\{l\},Y\)isσ\\sigma\-subgaussian underQiQ\_\{i\}\(and also underQ~i\\tilde\{Q\}\_\{i\}\)\. By Lemma[A\.6](https://arxiv.org/html/2608.11690#A1.Thmtheorem6), we can obtain

\|𝔼Qi​\[FW​\(Al,Y\)\]−𝔼Q~i​\[FW​\(Al,Y\)\]\|≤2σ2DKL\(Q~i∥Qi\)\.\\left\|\\mathbb\{E\}\_\{Q\_\{i\}\}\[F\_\{W\}\(A\_\{l\},Y\)\]\-\\mathbb\{E\}\_\{\\tilde\{Q\}\_\{i\}\}\[F\_\{W\}\(A\_\{l\},Y\)\]\\right\|\\leq\\sqrt\{2\\sigma^\{2\}\\,D\_\{\\mathrm\{KL\}\}\(\\tilde\{Q\}\_\{i\}\\\|Q\_\{i\}\)\}\.Substitute intoΔdrift\(l\)\\Delta\_\{\\mathrm\{drift\}\}^\{\(l\)\}

\|Δdrift\(l\)\|≤∑i=1T−12σ2DKL\(QAl,Y\|i,W1:l∥PAl,Y\|i,W1:l\)\.\|\\Delta\_\{\\mathrm\{drift\}\}^\{\(l\)\}\|\\leq\\sum\_\{i=1\}^\{T\-1\}\\sqrt\{2\\sigma^\{2\}\\,D\_\{\\mathrm\{KL\}\}\\\!\\big\(Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\ \\big\\\|\\ P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\big\)\}\.Take the expectation and use Jensen and Cauchy–Schwarz’s inequality,

𝔼⁡\[\|Δdrift\(l\)\|\]\\displaystyle\\mathbb\{E\}\\\!\\big\[\|\\Delta\_\{\\mathrm\{drift\}\}^\{\(l\)\}\|\\big\]≤\(T−1\)∑i=1T−1𝔼\[2σ2DKL\(QAl,Y\|i,W1:l∥PAl,Y\|i,W1:l\)\]\\displaystyle\\leq\\sqrt\{\(T\-1\)\\sum\_\{i=1\}^\{T\-1\}\\mathbb\{E\}\\\!\\left\[2\\sigma^\{2\}\\,D\_\{\\mathrm\{KL\}\}\\\!\\big\(Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\ \\big\\\|\\ P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\big\)\\right\]\}=\(T−1\)2σ2⋅1T−1∑i=1T−1𝔼\[DKL\(QAl,Y\|i,W1:l∥PAl,Y\|i,W1:l\)\]\.\\displaystyle=\(T\-1\)\\sqrt\{2\\sigma^\{2\}\\cdot\\frac\{1\}\{T\-1\}\\sum\_\{i=1\}^\{T\-1\}\\mathbb\{E\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\big\(Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\ \\big\\\|\\ P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\big\)\\right\]\}\.\(14\)For each fixedW1:lW\_\{1:l\}and eachii, take expectation overW1:lW\_\{1:l\}and define

𝒦\(l\)≜1T−1∑i=1T−1𝔼\[DKL\(QAl,Y\|i,W1:l∥PAl,Y\|i,W1:l\)\]\.\\mathcal\{K\}^\{\(l\)\}\\triangleq\\frac\{1\}\{T\-1\}\\sum\_\{i=1\}^\{T\-1\}\\mathbb\{E\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\big\(Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\ \\big\\\|\\ P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\big\)\\right\]\.Combining with \([14](https://arxiv.org/html/2608.11690#A2.E14)\) that

𝔼⁡\[\|Δdrift\(l\)\|\]≤\(T−1\)​2​σ2​𝒦\(l\)\.\\mathbb\{E\}\\\!\\big\[\|\\Delta\_\{\\mathrm\{drift\}\}^\{\(l\)\}\|\\big\]\\leq\(T\-1\)\\,\\sqrt\{2\\sigma^\{2\}\\,\\mathcal\{K\}^\{\(l\)\}\}\.\(15\)Next, we bound\|𝔼⁡\[Δvar\(l\)\]\|\|\\mathbb\{E\}\[\\Delta\_\{\\mathrm\{var\}\}^\{\(l\)\}\]\|\. Recall the weights

a≜1m\(each old replay sample\),b≜1n\(each new\-task sample\)\.a\\triangleq\\frac\{1\}\{m\}\\quad\\text\{\(each old replay sample\)\},\\qquad b\\triangleq\\frac\{1\}\{n\}\\quad\\text\{\(each new\-task sample\)\}\.LetN=\(T−1\)​m\+nN=\(T\-1\)m\+n\. Enumerate all representation–label pairs used in the variance term as a single sequence

U1\(l\),…,UN\(l\),where\{Uk\(l\)\}k=1\(T−1\)​m=𝐔old\(l\),\{Uk\(l\)\}k=\(T−1\)​m\+1N=𝐔new\(l\)\.U\_\{1\}^\{\(l\)\},\\ldots,U\_\{N\}^\{\(l\)\},\\quad\\text\{where\}\\quad\\\{U\_\{k\}^\{\(l\)\}\\\}\_\{k=1\}^\{\(T\-1\)m\}=\\mathbf\{U\}\_\{\\mathrm\{old\}\}^\{\(l\)\},\\ \\ \\\{U\_\{k\}^\{\(l\)\}\\\}\_\{k=\(T\-1\)m\+1\}^\{N\}=\\mathbf\{U\}\_\{\\mathrm\{new\}\}^\{\(l\)\}\.Define the corresponding weights

wk≜\{a,1≤k≤\(T−1\)​m,b,\(T−1\)​m<k≤N\.w\_\{k\}\\triangleq\\begin\{cases\}a,&1\\leq k\\leq\(T\-1\)m,\\\\ b,&\(T\-1\)m<k\\leq N\.\\end\{cases\}For each indexkk, define its centering marginal \(givenW1:lW\_\{1:l\}\) as

Qk\(⋅\)≜PU\(l\)∣𝗌𝗋𝖼\(k\),W1:l=\{QAl,Y\|i,W1:l,𝗌𝗋𝖼⁡\(k\)=i∈\[T−1\],PAl,Y\|T,W1:l,𝗌𝗋𝖼⁡\(k\)=new\.Q\_\{k\}\(\\cdot\)\\triangleq P\_\{U^\{\(l\)\}\\mid\\mathsf\{src\}\(k\),\\,W\_\{1:l\}\}=\\begin\{cases\}Q\_\{A\_\{l\},Y\|i,W\_\{1:l\}\},&\\mathsf\{src\}\(k\)=i\\in\[T\-1\],\\\\ P\_\{A\_\{l\},Y\|T,W\_\{1:l\}\},&\\mathsf\{src\}\(k\)=\\mathrm\{new\}\.\\end\{cases\}where𝗌𝗋𝖼⁡\(k\)∈\{1,…,T−1,new\}\\mathsf\{src\}\(k\)\\in\\\{1,\\dots,T\-1,\\mathrm\{new\}\\\}indicates the task source ofUk\(l\)U\_\{k\}^\{\(l\)\}: fork≤\(T−1\)​mk\\leq\(T\-1\)m,𝗌𝗋𝖼⁡\(k\)=i\\mathsf\{src\}\(k\)=iifUk\(l\)U\_\{k\}^\{\(l\)\}is a replayed pair from old taskii, and fork\>\(T−1\)​mk\>\(T\-1\)m,𝗌𝗋𝖼⁡\(k\)=new\\mathsf\{src\}\(k\)=\\mathrm\{new\}\.

Define the random centering functionalF¯k​\(W\)≜𝔼U∼Qk​\[FW​\(U\)\],\\bar\{F\}\_\{k\}\(W\)\\triangleq\\mathbb\{E\}\_\{U\\sim Q\_\{k\}\}\[F\_\{W\}\(U\)\],which is a random variable because it still depends on the top parametersWl\+1:LW\_\{l\+1:L\}throughFWF\_\{W\}\. With this notation, the variance term can be written as

Δvar\(l\)=∑k=1Nwk​\(F¯k​\(W\)−FW​\(Uk\(l\)\)\)\.\\Delta\_\{\\mathrm\{var\}\}^\{\(l\)\}=\\sum\_\{k=1\}^\{N\}w\_\{k\}\\Big\(\\bar\{F\}\_\{k\}\(W\)\-F\_\{W\}\(U\_\{k\}^\{\(l\)\}\)\\Big\)\.\(16\)
Introduce the filtrationℱ0⊂ℱ1⊂⋯⊂ℱN\\mathcal\{F\}\_\{0\}\\subset\\mathcal\{F\}\_\{1\}\\subset\\cdots\\subset\\mathcal\{F\}\_\{N\}with

ℱk≜σ\(W1:l,U1\(l\),…,Uk\(l\)\),k=0,1,…,N,\\mathcal\{F\}\_\{k\}\\triangleq\\sigma\\\!\\big\(W\_\{1:l\},U\_\{1\}^\{\(l\)\},\\ldots,U\_\{k\}^\{\(l\)\}\\big\),\\qquad k=0,1,\\ldots,N,\(whereℱ0=σ\(W1:l\)\\mathcal\{F\}\_\{0\}=\\sigma\(W\_\{1:l\}\)\)\.

Fix anyk∈\[N\]k\\in\[N\]and condition onℱk−1\\mathcal\{F\}\_\{k\-1\}\. Define the conditional joint law

μk≜PWl\+1:L,Uk\(l\)∣ℱk−1\.\\mu\_\{k\}\\triangleq P\_\{W\_\{l\+1:L\},\\,U\_\{k\}^\{\(l\)\}\\mid\\mathcal\{F\}\_\{k\-1\}\}\.Define the corresponding reference conditional law

νk⋆≜PWl\+1:L∣ℱk−1⊗Qk\.\\nu\_\{k\}^\{\\star\}\\triangleq P\_\{W\_\{l\+1:L\}\\mid\\mathcal\{F\}\_\{k\-1\}\}\\ \\otimes\\ Q\_\{k\}\.\(17\)Taking expectation of \([16](https://arxiv.org/html/2608.11690#A2.E16)\) and using the tower property and the definition ofΔvar\(l\)\\Delta\_\{\\mathrm\{var\}\}^\{\(l\)\}, we have

𝔼⁡\[Δvar\(l\)\]\\displaystyle\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(l\)\}\\big\]=∑k=1Nwk​𝔼​\[𝔼⁡\[F¯k​\(W\)∣ℱk−1\]−𝔼⁡\[FW​\(Uk\(l\)\)∣ℱk−1\]\],\\displaystyle=\\sum\_\{k=1\}^\{N\}w\_\{k\}\\,\\mathbb\{E\}\\\!\\left\[\\mathbb\{E\}\\\!\\big\[\\bar\{F\}\_\{k\}\(W\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\-\\mathbb\{E\}\\\!\\big\[F\_\{W\}\(U\_\{k\}^\{\(l\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\\right\],\(18\)Now we identify the two conditional expectations in \([18](https://arxiv.org/html/2608.11690#A2.E18)\) with expectations underνk⋆\\nu\_\{k\}^\{\\star\}andμk\\mu\_\{k\}\. By definitionF¯k​\(W\)=𝔼U∼Qk​\[FW​\(U\)\]\\bar\{F\}\_\{k\}\(W\)=\\mathbb\{E\}\_\{U\\sim Q\_\{k\}\}\[F\_\{W\}\(U\)\],ℱk−1\\mathcal\{F\}\_\{k\-1\}averages over the remaining randomness inWl\+1:LW\_\{l\+1:L\}\. Usingνk⋆\\nu\_\{k\}^\{\\star\}in \([17](https://arxiv.org/html/2608.11690#A2.E17)\),

𝔼⁡\[F¯k​\(W\)∣ℱk−1\]=𝔼νk⋆​\[FW​\(U\)∣ℱk−1\],\\mathbb\{E\}\\left\[\\bar\{F\}\_\{k\}\(W\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]=\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\left\[F\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\],\(19\)and explicitly

𝔼νk⋆\[FW\(U\)∣ℱk−1\]=∫\(∫F\(W1:l,w\)\(u\)Qk\(du\)\)PWl\+1:L\|ℱk−1\(dw\),\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\[F\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\]=\\int\\left\(\\int F\_\{\(W\_\{1:l\},w\)\}\(u\)Q\_\{k\}\(du\)\\right\)P\_\{W\_\{l\+1:L\}\|\\mathcal\{F\}\_\{k\-1\}\}\(dw\),which is indeedℱk−1\\mathcal\{F\}\_\{k\-1\}\-measurable\. By definition ofμk\\mu\_\{k\}in \([17](https://arxiv.org/html/2608.11690#A2.E17)\),

𝔼⁡\[FW​\(Uk\(l\)\)∣ℱk−1\]=𝔼μk​\[FW​\(Uk\(l\)\)∣ℱk−1\]\.\\mathbb\{E\}\{\\left\[F\_\{W\}\(U\_\{k\}^\{\(l\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]\}=\\mathbb\{E\}\_\{\\mu\_\{k\}\}\{\\left\[F\_\{W\}\(U\_\{k\}^\{\(l\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]\}\.\(20\)Substituting \([19](https://arxiv.org/html/2608.11690#A2.E19)\) and \([20](https://arxiv.org/html/2608.11690#A2.E20)\) into \([18](https://arxiv.org/html/2608.11690#A2.E18)\) gives the exact bridge identity,

𝔼⁡\[Δvar\(l\)\]=∑k=1Nwk​𝔼​\[𝔼νk⋆​\[FW​\(U\)∣ℱk−1\]−𝔼μk​\[FW​\(Uk\(l\)\)∣ℱk−1\]\]\.\\mathbb\{E\}\[\\Delta\_\{\\mathrm\{var\}\}^\{\(l\)\}\]=\\sum\_\{k=1\}^\{N\}w\_\{k\}\\mathbb\{E\}\\left\[\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\left\[F\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]\-\\mathbb\{E\}\_\{\\mu\_\{k\}\}\\left\[F\_\{W\}\(U\_\{k\}^\{\(l\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]\\right\]\.\(21\)Applying\|𝔼⁡\[X\]\|≤𝔼⁡\[\|X\|\]\|\\,\\mathbb\{E\}\[X\]\\,\|\\leq\\mathbb\{E\}\[\|X\|\]and the triangle inequality gives

\|𝔼⁡\[Δvar\(l\)\]\|≤∑k=1Nwk​𝔼​\[\|𝔼νk⋆​\[FW​\(U\)∣ℱk−1\]−𝔼μk​\[FW​\(Uk\(l\)\)∣ℱk−1\]\|\]\.\\Big\|\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(l\)\}\\big\]\\Big\|\\leq\\sum\_\{k=1\}^\{N\}w\_\{k\}\\,\\mathbb\{E\}\\\!\\left\[\\Big\|\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\\!\\big\[F\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\-\\mathbb\{E\}\_\{\\mu\_\{k\}\}\\\!\\big\[F\_\{W\}\(U\_\{k\}^\{\(l\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\\Big\|\\right\]\.\(22\)Under the assumption that the loss isσ\\sigma\-subgaussian, for any fixedWWthe random variableFW​\(U\)F\_\{W\}\(U\)isσ\\sigma\-subgaussian underQkQ\_\{k\}\. Hence, conditioning onℱk−1\\mathcal\{F\}\_\{k\-1\}, the same subgaussian proxy holds underνk⋆\\nu^\{\\star\}\_\{k\}, because it is a mixture overWl\+1:LW\_\{l\+1:L\}with the sameσ\\sigma\. Therefore Lemma[A\.6](https://arxiv.org/html/2608.11690#A1.Thmtheorem6)applied conditionally onℱk−1\\mathcal\{F\}\_\{k\-1\}yields

\|𝔼νk⋆​\[FW​\(U\)∣ℱk−1\]−𝔼μk​\[FW​\(Uk\(l\)\)∣ℱk−1\]\|≤2σ2DKL\(μk∥νk⋆\)\.\\Big\|\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\\!\\big\[F\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\-\\mathbb\{E\}\_\{\\mu\_\{k\}\}\\\!\\big\[F\_\{W\}\(U\_\{k\}^\{\(l\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\\Big\|\\leq\\sqrt\{2\\sigma^\{2\}\\,D\_\{\\mathrm\{KL\}\}\(\\mu\_\{k\}\\\|\\nu\_\{k\}^\{\\star\}\)\}\.\(23\)Letνk≜PWl\+1:L∣ℱk−1⊗PUk\(l\)\|ℱk−1\\nu\_\{k\}\\triangleq P\_\{W\_\{l\+1:L\}\\mid\\mathcal\{F\}\_\{k\-1\}\}\\otimes P\_\{U\_\{k\}^\{\(l\)\}\\mid\\mathcal\{F\}\_\{k\-1\}\}be the product of the true marginals underℱk−1\\mathcal\{F\}\_\{k\-1\}\. Then by the KL chain rule,

DKL\(μk∥νk⋆\)\\displaystyle D\_\{\\mathrm\{KL\}\}\(\\mu\_\{k\}\\\|\\nu\_\{k\}^\{\\star\}\)=DKL\(μk∥νk\)\+𝔼μk\[logd​νkd​νk⋆\(Wl\+1:L,Uk\(l\)\)\]\\displaystyle=D\_\{\\mathrm\{KL\}\}\(\\mu\_\{k\}\\\|\\nu\_\{k\}\)\+\\mathbb\{E\}\_\{\\mu\_\{k\}\}\\\!\\left\[\\log\\frac\{d\\nu\_\{k\}\}\{d\\nu\_\{k\}^\{\\star\}\}\(W\_\{l\+1:L\},U\_\{k\}^\{\(l\)\}\)\\right\]=I\(Uk\(l\);Wl\+1:L∣ℱk−1\)\+DKL\(PUk\(l\)\|ℱk−1∥Qk\)\.\\displaystyle=I\\\!\\big\(U\_\{k\}^\{\(l\)\};\\,W\_\{l\+1:L\}\\mid\\mathcal\{F\}\_\{k\-1\}\\big\)\+D\_\{\\mathrm\{KL\}\}\\\!\\Big\(P\_\{U\_\{k\}^\{\(l\)\}\\mid\\mathcal\{F\}\_\{k\-1\}\}\\ \\big\\\|\\ Q\_\{k\}\\Big\)\.\(24\)
Define

Ik≜I\(Uk\(l\);Wl\+1:L∣ℱk−1\),Mk≜DKL\(PUk\(l\)\|ℱk−1∥Qk\)\.I\_\{k\}\\triangleq I\\\!\\big\(U\_\{k\}^\{\(l\)\};\\,W\_\{l\+1:L\}\\mid\\mathcal\{F\}\_\{k\-1\}\\big\),\\qquad M\_\{k\}\\triangleq D\_\{\\mathrm\{KL\}\}\\\!\\Big\(P\_\{U\_\{k\}^\{\(l\)\}\\mid\\mathcal\{F\}\_\{k\-1\}\}\\ \\big\\\|\\ Q\_\{k\}\\Big\)\.Then by \([22](https://arxiv.org/html/2608.11690#A2.E22)\)–\([24](https://arxiv.org/html/2608.11690#A2.E24)\),

\|𝔼νk⋆​\[FW​\(U\)∣ℱk−1\]−𝔼μk​\[FW​\(Uk\(l\)\)∣ℱk−1\]\|≤2​σ2​\(Ik\+Mk\)\.\\Big\|\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\\!\\big\[F\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\-\\mathbb\{E\}\_\{\\mu\_\{k\}\}\\\!\\big\[F\_\{W\}\(U\_\{k\}^\{\(l\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\\Big\|\\leq\\sqrt\{2\\sigma^\{2\}\\,\(I\_\{k\}\+M\_\{k\}\)\}\.\(25\)
Substituting \([25](https://arxiv.org/html/2608.11690#A2.E25)\) into \([22](https://arxiv.org/html/2608.11690#A2.E22)\) yields

\|𝔼⁡\[Δvar\(l\)\]\|≤2​σ2​𝔼​\[∑k=1Nwk​Ik\+Mk\]\.\\Big\|\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(l\)\}\\big\]\\Big\|\\leq\\sqrt\{2\\sigma^\{2\}\}\\,\\mathbb\{E\}\\\!\\left\[\\sum\_\{k=1\}^\{N\}w\_\{k\}\\sqrt\{I\_\{k\}\+M\_\{k\}\}\\right\]\.\(26\)
Applying Cauchy–Schwarz to the inner weighted sum gives

∑k=1Nwk​Ik\+Mk≤∑k=1Nwk2⋅∑k=1N\(Ik\+Mk\)\.\\sum\_\{k=1\}^\{N\}w\_\{k\}\\sqrt\{I\_\{k\}\+M\_\{k\}\}\\leq\\sqrt\{\\sum\_\{k=1\}^\{N\}w\_\{k\}^\{2\}\}\\cdot\\sqrt\{\\sum\_\{k=1\}^\{N\}\(I\_\{k\}\+M\_\{k\}\)\}\.\(27\)Using Jensen’s inequality𝔼⁡\[X\]≤𝔼⁡\[X\]\\mathbb\{E\}\[\\sqrt\{X\}\]\\leq\\sqrt\{\\mathbb\{E\}\[X\]\}and noting,

∑k=1Nwk2=\(T−1\)​m​\(1m\)2\+n​\(1n\)2=\(T−1m\+1n\)=1Neff,\\sum\_\{k=1\}^\{N\}w\_\{k\}^\{2\}=\(T\-1\)m\{\\left\(\\frac\{1\}\{m\}\\right\)\}^\{2\}\+n\{\\left\(\\frac\{1\}\{n\}\\right\)\}^\{2\}=\{\\left\(\\frac\{T\-1\}\{m\}\+\\frac\{1\}\{n\}\\right\)\}=\\frac\{1\}\{N\_\{\\mathrm\{eff\}\}\},we obtain

\|𝔼⁡\[Δvar\(l\)\]\|≤2​σ2Neff​𝔼⁡\[∑k=1NIk\]\+𝔼⁡\[∑k=1NMk\]\.\\Big\|\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(l\)\}\\big\]\\Big\|\\leq\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\}\\,\\sqrt\{\\mathbb\{E\}\\\!\\left\[\\sum\_\{k=1\}^\{N\}I\_\{k\}\\right\]\+\\mathbb\{E\}\\\!\\left\[\\sum\_\{k=1\}^\{N\}M\_\{k\}\\right\]\}\.\(28\)
By the chain rule for mutual information and Lemma[A\.4](https://arxiv.org/html/2608.11690#A1.Thmtheorem4),

𝔼\[∑k=1NIk\]=I\(𝐔\(l\);Wl\+1:L∣W1:l\)=𝒮\(l\)\+𝒫\(l\)−ℛ\(l\)\.\\mathbb\{E\}\\\!\\left\[\\sum\_\{k=1\}^\{N\}I\_\{k\}\\right\]=I\(\\mathbf\{U\}^\{\(l\)\};\\,W\_\{l\+1:L\}\\mid W\_\{1:l\}\)=\\mathcal\{S\}^\{\(l\)\}\+\\mathcal\{P\}^\{\(l\)\}\-\\mathcal\{R\}^\{\(l\)\}\.Moreover, by the chain rule for KL under the task\-wise block product centeringQ~𝐔\(l\)\|W1:l=⨂k=1NQk\\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(l\)\}\|W\_\{1:l\}\}=\\bigotimes\_\{k=1\}^\{N\}Q\_\{k\}, we have

𝔼\[∑k=1NMk\]=𝔼\[DKL\(P𝐔\(l\)∣W1:l∥Q~𝐔\(l\)\|W1:l\)\]=𝒞\(l\)\.\\mathbb\{E\}\\\!\\left\[\\sum\_\{k=1\}^\{N\}M\_\{k\}\\right\]=\\mathbb\{E\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\Big\(P\_\{\\mathbf\{U\}^\{\(l\)\}\\mid W\_\{1:l\}\}\\ \\big\\\|\\ \\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(l\)\}\|W\_\{1:l\}\}\\Big\)\\right\]=\\mathcal\{C\}^\{\(l\)\}\.Combining \([28](https://arxiv.org/html/2608.11690#A2.E28)\), we can get

\|𝔼⁡\[Δvar\(l\)\]\|≤2​σ2Neff​\(𝒮\(l\)\+𝒫\(l\)−ℛ\(l\)\+𝒞\(l\)\)\.\\Big\|\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(l\)\}\\big\]\\Big\|\\leq\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\\Big\(\\mathcal\{S\}^\{\(l\)\}\+\\mathcal\{P\}^\{\(l\)\}\-\\mathcal\{R\}^\{\(l\)\}\+\\mathcal\{C\}^\{\(l\)\}\\Big\)\}\.\(29\)Finally, substituting \([15](https://arxiv.org/html/2608.11690#A2.E15)\) and \([29](https://arxiv.org/html/2608.11690#A2.E29)\) into \([13](https://arxiv.org/html/2608.11690#A2.E13)\) completes the proof\. ∎

### B\-BProof of Corollary[IV\.2](https://arxiv.org/html/2608.11690#S4.Thmtheorem2)

###### Corollary[IV\.2](https://arxiv.org/html/2608.11690#S4.Thmtheorem2)\(Restate\)\.

At the input layerl=0l=0, the hierarchical generalization bound simplifies to:

\|genW\|≤2​σ2Neff​I​\(S,W\),\\displaystyle\|\\mathrm\{gen\}\_\{W\}\|\\;\\leq\\;\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\\,I\\\!\\big\(S;W\\big\)\},whereSSis the original\(X,Y\)\(X,Y\)pair in the replay buffer and the current task dataset\.

###### Proof\.

The proof adopts the same framework as Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)\. Atl=0l=0, the representationA0=XA\_\{0\}=XandW1:0W\_\{1:0\}is empty\. We use the same notationFW​\(Z\)=ℓ⁡\(W,Z\)F\_\{W\}\(Z\)=\\ell\(W,Z\)forZ=\(X,Y\)∈𝒵Z=\(X,Y\)\\in\\mathcal\{Z\}\. By definition,

genW=𝔼W,D1:T\[ℒ\(W\)−ℒ^\(W\)\],\\mathrm\{gen\}\_\{W\}=\\mathbb\{E\}\_\{W,D^\{1:T\}\}\\big\[\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)\\big\],whereℒ\\mathcal\{L\}andℒ^\\hat\{\\mathcal\{L\}\}are population and empirical risks\. Expandingℒ⁡\(W\)\\mathcal\{L\}\(W\)andℒ^​\(W\)\\hat\{\\mathcal\{L\}\}\(W\)as in \([8](https://arxiv.org/html/2608.11690#A2.E8)\) yields the exact old and new split

ℒ​\(W\)−ℒ^​\(W\)\\displaystyle\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)=∑i=1T−1\(𝔼Z∼𝒟i​\[FW​\(Z\)\]−1m​∑j∈ℐiFW​\(Zji\)\)\+\(𝔼Z∼𝒟T​\[FW​\(Z\)\]−1n​∑j=1nFW​\(ZjT\)\)\.\\displaystyle=\\sum\_\{i=1\}^\{T\-1\}\\left\(\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{i\}\}\[F\_\{W\}\(Z\)\]\-\\frac\{1\}\{m\}\\sum\_\{j\\in\\mathcal\{I\}\_\{i\}\}F\_\{W\}\(Z\_\{j\}^\{i\}\)\\right\)\+\\left\(\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{T\}\}\[F\_\{W\}\(Z\)\]\-\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}F\_\{W\}\(Z\_\{j\}^\{T\}\)\\right\)\.\(30\)RecallPAl,Y\|i,W1:lP\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}denote the conditional law of the representation–label pair\(Al,Y\)\(A\_\{l\},Y\)induced by\(X,Y\)∼𝒟i\(X,Y\)\\sim\\mathcal\{D\}\_\{i\}under the current representation parametersW1:lW\_\{1:l\}\. Therefore, whenl=0l=0, for each old taski∈\[T−1\]i\\in\[T\-1\], the input\-layer population law isPX,Y\|i≡𝒟iP\_\{X,Y\|i\}\\equiv\\mathcal\{D\}\_\{i\}\. Similarly, the replay proxyP^A0,Y\|Si=PSi=1m​∑j∈ℐiδZji​\(⋅\)\\hat\{P\}\_\{A\_\{0\},Y\|S^\{i\}\}=P\_\{S^\{i\}\}=\\frac\{1\}\{m\}\\sum\_\{j\\in\\mathcal\{I\}\_\{i\}\}\\delta\_\{Z\_\{j\}^\{i\}\}\(\\cdot\), which is uniform on the buffer subset\. Define its unconditional centroid law atl=0l=0by

QA0,Y\|i≜𝔼⁡\[P^Si\]\.Q\_\{A\_\{0\},Y\|i\}\\triangleq\\mathbb\{E\}\\big\[\\hat\{P\}\_\{S^\{i\}\}\\big\]\.\(31\)Using the same add\-and\-subtract step as \([11](https://arxiv.org/html/2608.11690#A2.E11)\) \(withW1:0W\_\{1:0\}void\), we obtain the exact decomposition

ℒ⁡\(W\)−ℒ^​\(W\)=Δdrift\(0\)\+Δvar\(0\),\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)=\\Delta\_\{\\mathrm\{drift\}\}^\{\(0\)\}\+\\Delta\_\{\\mathrm\{var\}\}^\{\(0\)\},\(32\)where

Δdrift\(0\)\\displaystyle\\Delta\_\{\\mathrm\{drift\}\}^\{\(0\)\}≜∑i=1T−1\(𝔼Z∼𝒟i​\[FW​\(Z\)\]−𝔼Z∼QA0,Y\|i​\[FW​\(Z\)\]\),\\displaystyle\\triangleq\\sum\_\{i=1\}^\{T\-1\}\\Big\(\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{i\}\}\[F\_\{W\}\(Z\)\]\-\\mathbb\{E\}\_\{Z\\sim Q\_\{A\_\{0\},Y\|i\}\}\[F\_\{W\}\(Z\)\]\\Big\),\(33\)Δvar\(0\)\\displaystyle\\Delta\_\{\\mathrm\{var\}\}^\{\(0\)\}≜∑i=1T−1\(𝔼Z∼QA0,Y\|i​\[FW​\(Z\)\]−1m​∑j∈ℐiFW​\(Zji\)\)\+\(𝔼Z∼𝒟T​\[FW​\(Z\)\]−1n​∑j=1nFW​\(ZjT\)\)\.\\displaystyle\\triangleq\\sum\_\{i=1\}^\{T\-1\}\\Big\(\\mathbb\{E\}\_\{Z\\sim Q\_\{A\_\{0\},Y\|i\}\}\[F\_\{W\}\(Z\)\]\-\\frac\{1\}\{m\}\\sum\_\{j\\in\\mathcal\{I\}\_\{i\}\}F\_\{W\}\(Z\_\{j\}^\{i\}\)\\Big\)\+\\Big\(\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{T\}\}\[F\_\{W\}\(Z\)\]\-\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}F\_\{W\}\(Z\_\{j\}^\{T\}\)\\Big\)\.\(34\)At the input layer, the replay centroid coincides with the task distribution: indeed, by symmetry of uniform subsampling from i\.i\.d\. samples, a uniformly chosen element from the random subsetMiM^\{i\}has marginal law𝒟i\\mathcal\{D\}\_\{i\}, henceQA0,Y\|i=𝒟i=PX,Y\|iQ\_\{A\_\{0\},Y\|i\}=\\mathcal\{D\}\_\{i\}=P\_\{X,Y\|i\}\. Therefore,Δdrift\(0\)=0\\Delta\_\{\\mathrm\{drift\}\}^\{\(0\)\}=0in \([33](https://arxiv.org/html/2608.11690#A2.E33)\), and

\|genW\|=\|𝔼⁡\[ℒ⁡\(W\)−ℒ^​\(W\)\]\|=\|𝔼⁡\[Δvar\(0\)\]\|\.\|\\mathrm\{gen\}\_\{W\}\|=\\big\|\\mathbb\{E\}\[\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)\]\\big\|=\\big\|\\mathbb\{E\}\[\\Delta\_\{\\mathrm\{var\}\}^\{\(0\)\}\]\\big\|\.\(35\)LetN=\(T−1\)​m\+nN=\(T\-1\)m\+n\. Enumerate all examples used in \([34](https://arxiv.org/html/2608.11690#A2.E34)\) as a single sequenceU1\(0\),…,UN\(0\)∈𝒵U\_\{1\}^\{\(0\)\},\\dots,U\_\{N\}^\{\(0\)\}\\in\\mathcal\{Z\}, where the first\(T−1\)​m\(T\-1\)mterms come from the replay buffers and the lastnnterms come from the current task dataset\. Define the weights

a≜1m,b≜1n,wk≜\{a,1≤k≤\(T−1\)​m,b,\(T−1\)​m<k≤N,a\\triangleq\\frac\{1\}\{m\},\\qquad b\\triangleq\\frac\{1\}\{n\},\\qquad w\_\{k\}\\triangleq\\begin\{cases\}a,&1\\leq k\\leq\(T\-1\)m,\\\\ b,&\(T\-1\)m<k\\leq N,\\end\{cases\}and let𝗌𝗋𝖼⁡\(k\)∈\{1,…,T−1,new\}\\mathsf\{src\}\(k\)\\in\\\{1,\\dots,T\-1,\\mathrm\{new\}\\\}denote the source task ofUk\(0\)U\_\{k\}^\{\(0\)\}\. For eachkk, let

Qk​\(⋅\)≜PU\(0\)\|𝗌𝗋𝖼⁡\(k\)=\{𝒟i,𝗌𝗋𝖼⁡\(k\)=i∈\[T−1\],𝒟T,𝗌𝗋𝖼⁡\(k\)=new,Q\_\{k\}\(\\cdot\)\\triangleq P\_\{U^\{\(0\)\}\\mid\\mathsf\{src\}\(k\)\}=\\begin\{cases\}\\mathcal\{D\}\_\{i\},&\\mathsf\{src\}\(k\)=i\\in\[T\-1\],\\\\ \\mathcal\{D\}\_\{T\},&\\mathsf\{src\}\(k\)=\\mathrm\{new\},\\end\{cases\}Define the random centering functionalF¯k​\(W\)≜𝔼U∼Qk​\[FW​\(U\)\]\\bar\{F\}\_\{k\}\(W\)\\triangleq\\mathbb\{E\}\_\{U\\sim Q\_\{k\}\}\[F\_\{W\}\(U\)\], with this notation, the variance term can be written as

Δvar\(0\)=∑k=1Nwk​\(F¯k​\(W\)−FW​\(Uk\(0\)\)\)\.\\Delta\_\{\\mathrm\{var\}\}^\{\(0\)\}=\\sum\_\{k=1\}^\{N\}w\_\{k\}\\Big\(\\bar\{F\}\_\{k\}\(W\)\-F\_\{W\}\(U\_\{k\}^\{\(0\)\}\)\\Big\)\.\(36\)Define the filtration

ℱk≜σ\(U1\(0\),…,Uk\(0\)\),k=0,1,…,N\.\\mathcal\{F\}\_\{k\}\\triangleq\\sigma\\\!\\big\(U\_\{1\}^\{\(0\)\},\\dots,U\_\{k\}^\{\(0\)\}\\big\),\\qquad k=0,1,\\dots,N\.For eachkk, let the joint law beμk≜PW,Uk\(0\)\|ℱk−1\\mu\_\{k\}\\triangleq P\_\{W,\\,U\_\{k\}^\{\(0\)\}\\mid\\mathcal\{F\}\_\{k\-1\}\}and the corresponding reference conditional lawνk⋆≜PW\|ℱk−1⊗Qk\\nu\_\{k\}^\{\\star\}\\triangleq P\_\{W\\mid\\mathcal\{F\}\_\{k\-1\}\}\\ \\otimes\\ Q\_\{k\}\. Taking expectation of \([36](https://arxiv.org/html/2608.11690#A2.E36)\) and using the tower property and the definition ofΔvar\(0\)\\Delta\_\{\\mathrm\{var\}\}^\{\(0\)\}, we have

𝔼⁡\[Δvar\(0\)\]\\displaystyle\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(0\)\}\\big\]=∑k=1Nwk​𝔼​\[𝔼⁡\[F¯k​\(W\)∣ℱk−1\]−𝔼⁡\[FW​\(Uk\(0\)\)∣ℱk−1\]\],\\displaystyle=\\sum\_\{k=1\}^\{N\}w\_\{k\}\\,\\mathbb\{E\}\\\!\\left\[\\mathbb\{E\}\\\!\\big\[\\bar\{F\}\_\{k\}\(W\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\-\\mathbb\{E\}\\\!\\big\[F\_\{W\}\(U\_\{k\}^\{\(0\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\\right\],\(37\)Now we identify the two conditional expectations in \([37](https://arxiv.org/html/2608.11690#A2.E37)\) with expectations underνk⋆\\nu\_\{k\}^\{\\star\}andμk\\mu\_\{k\}\. By definitionF¯k​\(W\)=𝔼U∼Qk​\[FW​\(U\)\]\\bar\{F\}\_\{k\}\(W\)=\\mathbb\{E\}\_\{U\\sim Q\_\{k\}\}\[F\_\{W\}\(U\)\],ℱk−1\\mathcal\{F\}\_\{k\-1\}averages over the remaining randomness inWW\. Usingνk⋆\\nu\_\{k\}^\{\\star\},

𝔼⁡\[F¯k​\(W\)∣ℱk−1\]=𝔼νk⋆​\[FW​\(U\)∣ℱk−1\],\\mathbb\{E\}\\left\[\\bar\{F\}\_\{k\}\(W\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]=\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\left\[F\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\],\(38\)and explicitly

𝔼νk⋆​\[FW​\(U\)∣ℱk−1\]=∫\(∫Fw​\(u\)​Qk​\(𝑑u\)\)​PW\|ℱk−1​\(𝑑w\),\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\[F\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\]=\\int\\left\(\\int F\_\{w\}\(u\)\\,Q\_\{k\}\(du\)\\right\)P\_\{W\|\\mathcal\{F\}\_\{k\-1\}\}\(dw\),which is indeedℱk−1\\mathcal\{F\}\_\{k\-1\}\-measurable\. By definition ofμk\\mu\_\{k\},

𝔼⁡\[FW​\(Uk\(0\)\)∣ℱk−1\]=𝔼μk​\[FW​\(Uk\(0\)\)∣ℱk−1\]\.\\mathbb\{E\}\{\\left\[F\_\{W\}\(U\_\{k\}^\{\(0\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]\}=\\mathbb\{E\}\_\{\\mu\_\{k\}\}\{\\left\[F\_\{W\}\(U\_\{k\}^\{\(0\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]\}\.\(39\)Substituting \([38](https://arxiv.org/html/2608.11690#A2.E38)\) and \([39](https://arxiv.org/html/2608.11690#A2.E39)\) into \([37](https://arxiv.org/html/2608.11690#A2.E37)\) gives the exact bridge identity,

𝔼⁡\[Δvar\(0\)\]=∑k=1Nwk​𝔼​\[𝔼νk⋆​\[FW​\(U\)∣ℱk−1\]−𝔼μk​\[FW​\(Uk\(0\)\)∣ℱk−1\]\]\.\\mathbb\{E\}\[\\Delta\_\{\\mathrm\{var\}\}^\{\(0\)\}\]=\\sum\_\{k=1\}^\{N\}w\_\{k\}\\mathbb\{E\}\\left\[\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\left\[F\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]\-\\mathbb\{E\}\_\{\\mu\_\{k\}\}\\left\[F\_\{W\}\(U\_\{k\}^\{\(0\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]\\right\]\.\(40\)Applying\|𝔼⁡\[X\]\|≤𝔼⁡\[\|X\|\]\|\\,\\mathbb\{E\}\[X\]\\,\|\\leq\\mathbb\{E\}\[\|X\|\]and the triangle inequality gives

\|𝔼⁡\[Δvar\(0\)\]\|≤∑k=1Nwk​𝔼​\[\|𝔼νk⋆​\[FW​\(U\)∣ℱk−1\]−𝔼μk​\[FW​\(Uk\(0\)\)∣ℱk−1\]\|\]\.\\Big\|\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(0\)\}\\big\]\\Big\|\\leq\\sum\_\{k=1\}^\{N\}w\_\{k\}\\,\\mathbb\{E\}\\\!\\left\[\\Big\|\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\\!\\big\[F\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\-\\mathbb\{E\}\_\{\\mu\_\{k\}\}\\\!\\big\[F\_\{W\}\(U\_\{k\}^\{\(0\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\\Big\|\\right\]\.\(41\)Under the assumption that the loss isσ\\sigma\-subgaussian, for any fixedWWthe random variableFW​\(U\)F\_\{W\}\(U\)isσ\\sigma\-subgaussian underQkQ\_\{k\}\. Hence, conditioning onℱk−1\\mathcal\{F\}\_\{k\-1\}, the same subgaussian proxy holds underνk⋆\\nu^\{\\star\}\_\{k\}, because it is a mixture overWWwith the sameσ\\sigma\. Therefore Lemma[A\.6](https://arxiv.org/html/2608.11690#A1.Thmtheorem6)applied conditionally onℱk−1\\mathcal\{F\}\_\{k\-1\}yields

\|𝔼νk⋆​\[FW​\(U\)∣ℱk−1\]−𝔼μk​\[FW​\(Uk\(0\)\)∣ℱk−1\]\|≤2σ2DKL\(μk∥νk⋆\)\.\\Big\|\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\\!\\big\[F\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\-\\mathbb\{E\}\_\{\\mu\_\{k\}\}\\\!\\big\[F\_\{W\}\(U\_\{k\}^\{\(0\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\\Big\|\\leq\\sqrt\{2\\sigma^\{2\}\\,D\_\{\\mathrm\{KL\}\}\(\\mu\_\{k\}\\\|\\nu\_\{k\}^\{\\star\}\)\}\.\(42\)Letνk≜PW\|ℱk−1⊗PUk\(0\)\|ℱk−1\\nu\_\{k\}\\triangleq P\_\{W\\mid\\mathcal\{F\}\_\{k\-1\}\}\\otimes P\_\{U\_\{k\}^\{\(0\)\}\\mid\\mathcal\{F\}\_\{k\-1\}\}be the product of the true marginals underℱk−1\\mathcal\{F\}\_\{k\-1\}\. Then by the KL chain rule,

DKL\(μk∥νk⋆\)=I\(Uk\(0\);W∣ℱk−1\)\+DKL\(PUk\(0\)\|ℱk−1∥Qk\),D\_\{\\mathrm\{KL\}\}\(\\mu\_\{k\}\\\|\\nu\_\{k\}^\{\\star\}\)=I\\\!\\big\(U\_\{k\}^\{\(0\)\};\\,W\\mid\\mathcal\{F\}\_\{k\-1\}\\big\)\+D\_\{\\mathrm\{KL\}\}\\\!\\Big\(P\_\{U\_\{k\}^\{\(0\)\}\\mid\\mathcal\{F\}\_\{k\-1\}\}\\ \\big\\\|\\ Q\_\{k\}\\Big\),\(43\)
Define

Ik≜I\(Uk\(0\);W∣ℱk−1\),Mk≜DKL\(PUk\(0\)\|ℱk−1∥Qk\)\.I\_\{k\}\\triangleq I\\\!\\big\(U\_\{k\}^\{\(0\)\};\\,W\\mid\\mathcal\{F\}\_\{k\-1\}\\big\),\\qquad M\_\{k\}\\triangleq D\_\{\\mathrm\{KL\}\}\\\!\\Big\(P\_\{U\_\{k\}^\{\(0\)\}\\mid\\mathcal\{F\}\_\{k\-1\}\}\\ \\big\\\|\\ Q\_\{k\}\\Big\)\.Substituting into \([41](https://arxiv.org/html/2608.11690#A2.E41)\) and following exactly the Cauchy–Schwarz \+ Jensen aggregation in \([26](https://arxiv.org/html/2608.11690#A2.E26)\)–\([28](https://arxiv.org/html/2608.11690#A2.E28)\) yields

\|𝔼⁡\[Δvar\(0\)\]\|≤2​σ2Neff​𝔼⁡\[∑k=1NIk\]\+𝔼⁡\[∑k=1NMk\]\.\\Big\|\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(0\)\}\\big\]\\Big\|\\leq\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\}\\,\\sqrt\{\\mathbb\{E\}\\\!\\left\[\\sum\_\{k=1\}^\{N\}I\_\{k\}\\right\]\+\\mathbb\{E\}\\\!\\left\[\\sum\_\{k=1\}^\{N\}M\_\{k\}\\right\]\}\.\(44\)
By the mutual information chain rule,

𝔼⁡\[∑k=1NIk\]=I⁡\(S,W\)\.\\mathbb\{E\}\\\!\\left\[\\sum\_\{k=1\}^\{N\}I\_\{k\}\\right\]=I\(S;W\)\.Moreover, by the chain rule for KL under the task\-wise block product centeringQ~𝐔\(0\)≜\(⨂i=1T−1\(𝒟i\)⊗m\)⊗\(𝒟T\)⊗n\\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(0\)\}\}\\triangleq\\left\(\\bigotimes\_\{i=1\}^\{T\-1\}\(\\mathcal\{D\}\_\{i\}\)^\{\\otimes m\}\\right\)\\otimes\(\\mathcal\{D\}\_\{T\}\)^\{\\otimes n\}, we have

𝔼\[∑k=1NMk\]=DKL\(PS∥Q~𝐔\(0\)\)\.\\mathbb\{E\}\\\!\\left\[\\sum\_\{k=1\}^\{N\}M\_\{k\}\\right\]=D\_\{\\mathrm\{KL\}\}\\\!\\big\(P\_\{S\}\\ \\big\\\|\\ \\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(0\)\}\}\\big\)\.Atl=0l=0, the sequenceSSis block\-wise i\.i\.d\. by construction \(each task dataset is i\.i\.d\., tasks are mutually independent, and each buffer contributesmmdistinct coordinates from an i\.i\.d\. task dataset\); hencePS=Q~𝐔\(0\)P\_\{S\}=\\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(0\)\}\}and the above KL term is zero\. Combining with \([35](https://arxiv.org/html/2608.11690#A2.E35)\) and \([44](https://arxiv.org/html/2608.11690#A2.E44)\) yields

\|genW\|≤2​σ2Neff​I​\(S,W\)\.\|\\mathrm\{gen\}\_\{W\}\|\\leq\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\\,I\(S;W\)\}\.∎

### B\-CProof of Corollary[IV\.3](https://arxiv.org/html/2608.11690#S4.Thmtheorem3)

###### Corollary[IV\.3](https://arxiv.org/html/2608.11690#S4.Thmtheorem3)\(Restate\)\.

At the output layerl=Ll=L, the hierarchical generalization bound simplifies to:

\|genW\|≤\(T−1\)​2​σ2​𝒦\(L\)\+2​σ2Neff​𝒞\(L\),\\displaystyle\|\\mathrm\{gen\}\_\{W\}\|\\;\\leq\\;\(T\-1\)\\sqrt\{2\\sigma^\{2\}\\,\\mathcal\{K\}^\{\(L\)\}\}\\;\+\\;\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\\,\\mathcal\{C\}^\{\(L\)\}\},where𝒦\(L\)=1T−1∑i=1T−1𝔼W\[DKL\(QAL,Y\|i,W∥PAL,Y\|i,W\)\]\\mathcal\{K\}^\{\(L\)\}=\\frac\{1\}\{T\-1\}\\sum\_\{i=1\}^\{T\-1\}\\mathbb\{E\}\_\{W\}\[D\_\{\\mathrm\{KL\}\}\(Q\_\{A\_\{L\},Y\|i,W\}\\\|P\_\{A\_\{L\},Y\|i,W\}\)\],𝒞\(L\)=𝔼W\[DKL\(P𝐔\(L\)\|W∥Q~𝐔\(L\)\|W\)\]\\mathcal\{C\}^\{\(L\)\}=\\mathbb\{E\}\_\{W\}\[D\_\{\\mathrm\{KL\}\}\(P\_\{\\mathbf\{U\}^\{\(L\)\}\\mid W\}\\\|\\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(L\)\}\|W\}\)\]\.

###### Proof\.

We follow the proof structure of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)and highlight the simplifications atl=Ll=L\. Atl=Ll=L, the representation is the network outputAL=fW​\(X\)A\_\{L\}=f\_\{W\}\(X\)\. The top subnetwork above layerLLis empty, so we may regardgWL\+1:Lg\_\{W\_\{L\+1:L\}\}as the identity map and write the loss as a function of the output–label pair:

FW​\(t,y\)≜ℓ⁡\(t,y\),so thatℓ⁡\(W,\(X,Y\)\)=FW​\(AL,Y\)\.F\_\{W\}\(t,y\)\\triangleq\\ell\(t,y\),\\qquad\\text\{so that\}\\qquad\\ell\(W,\(X,Y\)\)=F\_\{W\}\(A\_\{L\},Y\)\.All layer\-LLdistributions below are understood as conditional laws given the full parameterW=W1:LW=W\_\{1:L\}\. As in \([8](https://arxiv.org/html/2608.11690#A2.E8)\), the population–empirical gap admits the exact split into old and new tasks\. Introducing the layer\-LLpopulation lawPAL,Y\|i,W1:LP\_\{A\_\{L\},Y\|i,W\_\{1:L\}\}of\(AL,Y\)\(A\_\{L\},Y\)under\(X,Y\)∼𝒟i\(X,Y\)\\sim\\mathcal\{D\}\_\{i\}and the replay proxyP^AL,Y\|Si,W1:L\\hat\{P\}\_\{A\_\{L\},Y\|S^\{i\},W\_\{1:L\}\}which is uniform over\{\(Aj,Li,Yji\):j∈ℐi\}\\\{\(A\_\{j,L\}^\{i\},Y\_\{j\}^\{i\}\):j\\in\\mathcal\{I\}\_\{i\}\\\}, recall the replay centroid

QAL,Y\|i,W1:L=𝔼\[P^AL,Y\|Si,W1:L\|W\]\.Q\_\{A\_\{L\},Y\|i,W\_\{1:L\}\}=\\mathbb\{E\}\\\!\\left\[\\hat\{P\}\_\{A\_\{L\},Y\|S^\{i\},W\_\{1:L\}\}\\ \\middle\|\\ W\\right\]\.\(45\)Adding and subtracting𝔼QAL,Y\|i,W1:L\[FW\(AL,Y\)\]\\mathbb\{E\}\_\{Q\_\{A\_\{L\},Y\|i,W\_\{1:L\}\}\}\[F\_\{W\}\(A\_\{L\},Y\)\]inside each old\-task gap yields an exact decomposition

ℒ⁡\(W\)−ℒ^​\(W\)=Δdrift\(L\)\+Δvar\(L\),\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)=\\Delta\_\{\\mathrm\{drift\}\}^\{\(L\)\}\+\\Delta\_\{\\mathrm\{var\}\}^\{\(L\)\},\(46\)where

Δdrift\(L\)\\displaystyle\\Delta\_\{\\mathrm\{drift\}\}^\{\(L\)\}≜∑i=1T−1\(𝔼Z∼𝒟i\[FW\(Z\)\]−𝔼Z∼QAL,Y\|i,W1:L\[FW\(Z\)\]\),\\displaystyle\\triangleq\\sum\_\{i=1\}^\{T\-1\}\\Big\(\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{i\}\}\[F\_\{W\}\(Z\)\]\-\\mathbb\{E\}\_\{Z\\sim Q\_\{A\_\{L\},Y\|i,W\_\{1:L\}\}\}\[F\_\{W\}\(Z\)\]\\Big\),\(47\)Δvar\(L\)\\displaystyle\\Delta\_\{\\mathrm\{var\}\}^\{\(L\)\}≜∑i=1T−1\(𝔼Z∼QAL,Y\|i,W1:L\[FW\(Z\)\]−1m∑j∈ℐiFW\(Zji\)\)\+\(𝔼Z∼𝒟T\[FW\(Z\)\]−1n∑j=1nFW\(ZjT\)\)\.\\displaystyle\\triangleq\\sum\_\{i=1\}^\{T\-1\}\\Big\(\\mathbb\{E\}\_\{Z\\sim Q\_\{A\_\{L\},Y\|i,W\_\{1:L\}\}\}\[F\_\{W\}\(Z\)\]\-\\frac\{1\}\{m\}\\sum\_\{j\\in\\mathcal\{I\}\_\{i\}\}F\_\{W\}\(Z\_\{j\}^\{i\}\)\\Big\)\+\\Big\(\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{T\}\}\[F\_\{W\}\(Z\)\]\-\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}F\_\{W\}\(Z\_\{j\}^\{T\}\)\\Big\)\.\(48\)By definition,

\|genW\|=\|𝔼⁡\[ℒ⁡\(W\)−ℒ^​\(W\)\]\|≤𝔼⁡\[\|Δdrift\(L\)\|\]\+\|𝔼⁡\[Δvar\(L\)\]\|\.\|\\mathrm\{gen\}\_\{W\}\|=\\left\|\\mathbb\{E\}\[\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)\]\\right\|\\leq\\mathbb\{E\}\\big\[\|\\Delta\_\{\\mathrm\{drift\}\}^\{\(L\)\}\|\\big\]\+\|\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(L\)\}\\big\]\|\.\(49\)Fix an old taski∈\[T−1\]i\\in\[T\-1\]and condition onWW\. SinceFW​\(AL,Y\)F\_\{W\}\(A\_\{L\},Y\)isσ\\sigma\-subgaussian under bothPAL,Y\|i,W1:LP\_\{A\_\{L\},Y\|i,W\_\{1:L\}\}andQAL,Y\|i,W1:LQ\_\{A\_\{L\},Y\|i,W\_\{1:L\}\}, by Lemma[A\.6](https://arxiv.org/html/2608.11690#A1.Thmtheorem6), we can obtain

\|𝔼PAL,Y\|i,W1:L\[FW\]−𝔼QAL,Y\|i,W1:L\[FW\]\|≤2σ2DKL\(QAL,Y\|i,W1:L∥PAL,Y\|i,W1:L\)\.\\Big\|\\mathbb\{E\}\_\{P\_\{A\_\{L\},Y\|i,W\_\{1:L\}\}\}\[F\_\{W\}\]\-\\mathbb\{E\}\_\{Q\_\{A\_\{L\},Y\|i,W\_\{1:L\}\}\}\[F\_\{W\}\]\\Big\|\\leq\\sqrt\{2\\sigma^\{2\}\\,D\_\{\\mathrm\{KL\}\}\\\!\\Big\(Q\_\{A\_\{L\},Y\|i,W\_\{1:L\}\}\\ \\big\\\|\\ P\_\{A\_\{L\},Y\|i,W\_\{1:L\}\}\\Big\)\}\.Averaging overiiand applying Cauchy–Schwarz exactly as in the main proof gives

𝔼⁡\[\|Δdrift\(L\)\|\]≤\(T−1\)​2​σ2​𝒦\(L\),\\mathbb\{E\}\\\!\\big\[\|\\Delta\_\{\\mathrm\{drift\}\}^\{\(L\)\}\|\\big\]\\leq\(T\-1\)\\sqrt\{2\\sigma^\{2\}\\,\\mathcal\{K\}^\{\(L\)\}\},\(50\)with𝒦\(L\)\\mathcal\{K\}^\{\(L\)\}as stated\. LetN=\(T−1\)​m\+nN=\(T\-1\)m\+nand enumerate the layer\-LLtraining pairs as a single sequenceU1\(L\),…,UN\(L\)U\_\{1\}^\{\(L\)\},\\dots,U\_\{N\}^\{\(L\)\}, with weights

a≜1m,b≜1n,wk≜\{a,1≤k≤\(T−1\)​m,b,\(T−1\)​m<k≤N,a\\triangleq\\frac\{1\}\{m\},\\qquad b\\triangleq\\frac\{1\}\{n\},\\qquad w\_\{k\}\\triangleq\\begin\{cases\}a,&1\\leq k\\leq\(T\-1\)m,\\\\ b,&\(T\-1\)m<k\\leq N,\\end\{cases\}and let𝗌𝗋𝖼⁡\(k\)∈\{1,…,T−1,new\}\\mathsf\{src\}\(k\)\\in\\\{1,\\dots,T\-1,\\mathrm\{new\}\\\}denote the source task ofUk\(L\)U\_\{k\}^\{\(L\)\}\. For eachkk, let

Qk\(⋅\)≜PU\(L\)\|𝗌𝗋𝖼⁡\(k\),W=\{QAL,Y\|i,W1:L,𝗌𝗋𝖼⁡\(k\)=i∈\[T−1\],PAL,Y\|T,W1:L,𝗌𝗋𝖼⁡\(k\)=new\.Q\_\{k\}\(\\cdot\)\\triangleq P\_\{U^\{\(L\)\}\\mid\\mathsf\{src\}\(k\),\\,W\}=\\begin\{cases\}Q\_\{A\_\{L\},Y\|i,W\_\{1:L\}\},&\\mathsf\{src\}\(k\)=i\\in\[T\-1\],\\\\ P\_\{A\_\{L\},Y\|T,W\_\{1:L\}\},&\\mathsf\{src\}\(k\)=\\mathrm\{new\}\.\\end\{cases\}With this notation, the variance term can be written as

Δvar\(L\)=∑k=1Nwk​\(𝔼U∼Qk​\[FW​\(U\)\]−FW​\(Uk\(L\)\)\)\.\\Delta\_\{\\mathrm\{var\}\}^\{\(L\)\}=\\sum\_\{k=1\}^\{N\}w\_\{k\}\\Big\(\\mathbb\{E\}\_\{U\\sim Q\_\{k\}\}\[F\_\{W\}\(U\)\]\-F\_\{W\}\(U\_\{k\}^\{\(L\)\}\)\\Big\)\.\(51\)Define the filtration \(ℱ0=σ⁡\(W\)\\mathcal\{F\}\_\{0\}=\\sigma\(W\)\)

ℱk≜σ\(W,U1\(L\),…,Uk\(L\)\),k=0,1,…,N\.\\mathcal\{F\}\_\{k\}\\triangleq\\sigma\\\!\\big\(W,U\_\{1\}^\{\(L\)\},\\dots,U\_\{k\}^\{\(L\)\}\\big\),\\qquad k=0,1,\\dots,N\.For eachkk, define the true conditional law of thekk\-th coordinate

μk≜PUk\(L\)\|ℱk−1\.\\mu\_\{k\}\\triangleq P\_\{U\_\{k\}^\{\(L\)\}\\mid\\mathcal\{F\}\_\{k\-1\}\}\.Since atl=Ll=Lthe “top parameter”WL\+1:LW\_\{L\+1:L\}is empty \(a deterministic constant\), the reference lawνk⋆\\nu\_\{k\}^\{\\star\}reduces to the centering marginal:

νk⋆≜Qk\.\\nu\_\{k\}^\{\\star\}\\triangleq Q\_\{k\}\.Taking expectation of \([51](https://arxiv.org/html/2608.11690#A2.E51)\) and using the tower property, we have

𝔼⁡\[Δvar\(L\)\]\\displaystyle\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(L\)\}\\big\]=∑k=1Nwk​𝔼​\[𝔼⁡\[FW​\(U\)∣ℱk−1\]−𝔼⁡\[FW​\(Uk\(L\)\)∣ℱk−1\]\],\\displaystyle=\\sum\_\{k=1\}^\{N\}w\_\{k\}\\,\\mathbb\{E\}\\\!\\left\[\\mathbb\{E\}\\\!\\big\[\{F\}\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\-\\mathbb\{E\}\\\!\\big\[F\_\{W\}\(U\_\{k\}^\{\(L\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\\right\],\(52\)SinceWWisℱk−1\\mathcal\{F\}\_\{k\-1\}\-measurable,𝔼⁡\[FW​\(U\)∣ℱk−1\]=𝔼νk⋆​\[FW​\(U\)\]\\mathbb\{E\}\\big\[\{F\}\_\{W\}\(U\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]=\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\left\[F\_\{W\}\(U\)\\right\]\. By definition ofμk\\mu\_\{k\},𝔼⁡\[FW​\(Uk\(L\)\)∣ℱk−1\]=𝔼μk​\[FW​\(Uk\(L\)\)∣ℱk−1\]\\mathbb\{E\}\[F\_\{W\}\(U\_\{k\}^\{\(L\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\]=\\mathbb\{E\}\_\{\\mu\_\{k\}\}\[F\_\{W\}\(U\_\{k\}^\{\(L\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\]\. Substituting into \([52](https://arxiv.org/html/2608.11690#A2.E52)\) gives the bridge identity,

𝔼⁡\[Δvar\(L\)\]=∑k=1Nwk​𝔼​\[𝔼νk⋆​\[FW​\(U\)\]−𝔼μk​\[FW​\(Uk\(L\)\)∣ℱk−1\]\]\.\\mathbb\{E\}\[\\Delta\_\{\\mathrm\{var\}\}^\{\(L\)\}\]=\\sum\_\{k=1\}^\{N\}w\_\{k\}\\mathbb\{E\}\\left\[\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\left\[F\_\{W\}\(U\)\\right\]\-\\mathbb\{E\}\_\{\\mu\_\{k\}\}\\left\[F\_\{W\}\(U\_\{k\}^\{\(L\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\right\]\\right\]\.\(53\)Applying\|𝔼⁡\[X\]\|≤𝔼⁡\[\|X\|\]\|\\,\\mathbb\{E\}\[X\]\\,\|\\leq\\mathbb\{E\}\[\|X\|\]and the triangle inequality gives

\|𝔼⁡\[Δvar\(L\)\]\|≤∑k=1Nwk​𝔼​\[\|𝔼νk⋆​\[FW​\(U\)\]−𝔼μk​\[FW​\(Uk\(L\)\)∣ℱk−1\]\|\]\.\\Big\|\\mathbb\{E\}\\big\[\\Delta\_\{\\mathrm\{var\}\}^\{\(L\)\}\\big\]\\Big\|\\leq\\sum\_\{k=1\}^\{N\}w\_\{k\}\\,\\mathbb\{E\}\\\!\\left\[\\Big\|\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\\\!\\big\[F\_\{W\}\(U\)\\big\]\-\\mathbb\{E\}\_\{\\mu\_\{k\}\}\\\!\\big\[F\_\{W\}\(U\_\{k\}^\{\(L\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\\big\]\\Big\|\\right\]\.\(54\)By Lemma[A\.6](https://arxiv.org/html/2608.11690#A1.Thmtheorem6), we can obtain

\|𝔼νk⋆​\[FW​\(U\)\]−𝔼μk​\[FW​\(Uk\(L\)\)∣ℱk−1\]\|≤2σ2DKL\(μk∥Qk\)\.\\Big\|\\mathbb\{E\}\_\{\\nu\_\{k\}^\{\\star\}\}\[F\_\{W\}\(U\)\]\-\\mathbb\{E\}\_\{\\mu\_\{k\}\}\[F\_\{W\}\(U\_\{k\}^\{\(L\)\}\)\\mid\\mathcal\{F\}\_\{k\-1\}\]\\Big\|\\leq\\sqrt\{2\\sigma^\{2\}\\,D\_\{\\mathrm\{KL\}\}\(\\mu\_\{k\}\\\|Q\_\{k\}\)\}\.Aggregating overkkwith Cauchy–Schwarz and Jensen gives

\|𝔼⁡\[Δvar\(L\)\]\|≤2​σ2Neff​𝔼\[∑k=1NDKL\(μk∥Qk\)\]\.\\Big\|\\mathbb\{E\}\[\\Delta\_\{\\mathrm\{var\}\}^\{\(L\)\}\]\\Big\|\\leq\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\}\\sqrt\{\\mathbb\{E\}\\\!\\left\[\\sum\_\{k=1\}^\{N\}D\_\{\\mathrm\{KL\}\}\\\!\\big\(\\mu\_\{k\}\\\|Q\_\{k\}\\big\)\\right\]\}\.Finally, by the chain rule for KL under the product referenceQ~𝐔\(L\)\|W=⨂k=1NQk\\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(L\)\}\|W\}=\\bigotimes\_\{k=1\}^\{N\}Q\_\{k\}, we have

𝔼\[∑k=1NDKL\(μk∥Qk\)\]=𝔼\[DKL\(P𝐔\(L\)\|W∥Q~𝐔\(L\)\|W\)\]=𝒞\(L\)\.\\mathbb\{E\}\\\!\\left\[\\sum\_\{k=1\}^\{N\}D\_\{\\mathrm\{KL\}\}\\\!\\big\(\\mu\_\{k\}\\\|Q\_\{k\}\\big\)\\right\]=\\mathbb\{E\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\\\!\\Big\(P\_\{\\mathbf\{U\}^\{\(L\)\}\\mid W\}\\ \\big\\\|\\ \\widetilde\{Q\}\_\{\\mathbf\{U\}^\{\(L\)\}\|W\}\\Big\)\\right\]=\\mathcal\{C\}^\{\(L\)\}\.Hence

\|𝔼⁡\[Δvar\(L\)\]\|≤2​σ2Neff​𝒞\(L\)\.\\Big\|\\mathbb\{E\}\[\\Delta\_\{\\mathrm\{var\}\}^\{\(L\)\}\]\\Big\|\\leq\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\}\}\}\\,\\mathcal\{C\}^\{\(L\)\}\}\.Then𝒮\(L\),𝒫\(L\),ℛ\(L\)\\mathcal\{S\}^\{\(L\)\},\\mathcal\{P\}^\{\(L\)\},\\mathcal\{R\}^\{\(L\)\}are mutual/interaction information involvingWL\+1:LW\_\{L\+1:L\}conditioned onW1:L=WW\_\{1:L\}=W\. SinceWL\+1:LW\_\{L\+1:L\}is deterministic \(empty\), all these quantities are zero:

𝒮\(L\)=𝒫\(L\)=ℛ\(L\)=0\.\\mathcal\{S\}^\{\(L\)\}=\\mathcal\{P\}^\{\(L\)\}=\\mathcal\{R\}^\{\(L\)\}=0\.
Combining the drift and variance bounds completes the proof\. ∎

### B\-DProof of Theorem[IV\.4](https://arxiv.org/html/2608.11690#S4.Thmtheorem4)

###### Theorem[IV\.4](https://arxiv.org/html/2608.11690#S4.Thmtheorem4)\(Restate\)\.

Assume the loss functionℓ⁡\(⋅,y\)\\ell\(\\cdot,y\)isρ0\\rho\_\{0\}\-Lipschitz and the activation functionsϕl\\phi\_\{l\}areρl\\rho\_\{l\}\-Lipschitz\. Then, at timeTT, we have:

\|genW\|≤minl∈\{0,…,L\}𝔼\[ρ¯l\(W\)⋅∑i=1T𝒲1\(P^Al,Y\|Si,W1:l,PAl,Y\|i,W1:l\)\],\\displaystyle\\big\|\\mathrm\{gen\}\_\{W\}\\big\|\\leq\\min\_\{l\\in\\\{0,\\dots,L\\\}\}\\mathbb\{E\}\\Bigg\[\\bar\{\\rho\}\_\{l\}\(W\)\\cdot\\sum\_\{i=1\}^\{T\}\\mathcal\{W\}\_\{1\}\\\!\\Big\(\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\},\\,P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\Big\)\\Bigg\],whereρ¯l​\(W\):=ρ0​\(1∨∏h=l\+1Lρh​‖Wh‖op\)\.\\bar\{\\rho\}\_\{l\}\(W\):=\\rho\_\{0\}\\left\(1\\vee\\prod\_\{h=l\+1\}^\{L\}\\rho\_\{h\}\\\|W\_\{h\}\\\|\_\{\\mathrm\{op\}\}\\right\)\.

###### Proof\.

Fix an arbitraryl∈\{0,1,…,L\}l\\in\\\{0,1,\\dots,L\\\}\. Define the upper sub\-network mapping from layerllto the output:GWl\+1:L\(a\):=gWL∘gWL−1∘⋯∘gWl\+1\(a\)G\_\{W\_\{l\+1:L\}\}\(a\):=g\_\{W\_\{L\}\}\\circ g\_\{W\_\{L\-1\}\}\\circ\\cdots\\circ g\_\{W\_\{l\+1\}\}\(a\), and define the re\-parameterized loss on\(a,y\)∈𝒜l×𝒴\(a,y\)\\in\\mathcal\{A\}\_\{l\}\\times\\mathcal\{Y\}:fW\(l\)\(a,y\):=ℓ\(GWl\+1:L\(a\),y\)f\_\{W\}^\{\(l\)\}\(a,y\):=\\ell\\big\(G\_\{W\_\{l\+1:L\}\}\(a\),\\,y\\big\)\. We first bound the Lipschitz constant offW\(l\)f\_\{W\}^\{\(l\)\}\. First, for any layerh∈\{l\+1,…,L\}h\\in\\\{l\+1,\\dots,L\\\}and anyu,u′u,u^\{\\prime\},

‖gWh​\(u\)−gWh​\(u′\)‖2\\displaystyle\\\|g\_\{W\_\{h\}\}\(u\)\-g\_\{W\_\{h\}\}\(u^\{\\prime\}\)\\\|\_\{2\}=‖ϕh​\(Wh​u\)−ϕh​\(Wh​u′\)‖2\\displaystyle=\\\|\\phi\_\{h\}\(W\_\{h\}u\)\-\\phi\_\{h\}\(W\_\{h\}u^\{\\prime\}\)\\\|\_\{2\}≤ρh​‖Wh​u−Wh​u′‖2\\displaystyle\\leq\\rho\_\{h\}\\\|W\_\{h\}u\-W\_\{h\}u^\{\\prime\}\\\|\_\{2\}≤ρh​‖Wh‖op​‖u−u′‖2\.\\displaystyle\\leq\\rho\_\{h\}\\\|W\_\{h\}\\\|\_\{\\mathrm\{op\}\}\\,\\\|u\-u^\{\\prime\}\\\|\_\{2\}\.HencegWhg\_\{W\_\{h\}\}isρh​‖Wh‖op\\rho\_\{h\}\\\|W\_\{h\}\\\|\_\{\\mathrm\{op\}\}\-Lipschitz\. Next, define for integersp≤qp\\leq qthe partial compositionGp:q:=gWq∘gWq−1∘⋯∘gWpG\_\{p:q\}:=g\_\{W\_\{q\}\}\\circ g\_\{W\_\{q\-1\}\}\\circ\\cdots\\circ g\_\{W\_\{p\}\}\. Then for anyu,u′u,u^\{\\prime\},

∥Gp:q\(u\)−Gp:q\(u′\)∥2\\displaystyle\\\|G\_\{p:q\}\(u\)\-G\_\{p:q\}\(u^\{\\prime\}\)\\\|\_\{2\}=∥gWq\(Gp:q−1\(u\)\)−gWq\(Gp:q−1\(u′\)\)∥2\\displaystyle=\\\|g\_\{W\_\{q\}\}\(G\_\{p:q\-1\}\(u\)\)\-g\_\{W\_\{q\}\}\(G\_\{p:q\-1\}\(u^\{\\prime\}\)\)\\\|\_\{2\}≤Lip\(gWq\)⋅∥Gp:q−1\(u\)−Gp:q−1\(u′\)∥2\\displaystyle\\leq\\mathrm\{Lip\}\(g\_\{W\_\{q\}\}\)\\cdot\\\|G\_\{p:q\-1\}\(u\)\-G\_\{p:q\-1\}\(u^\{\\prime\}\)\\\|\_\{2\}≤\(ρq∥Wq∥op\)⋅Lip\(Gp:q−1\)⋅∥u−u′∥2\\displaystyle\\leq\\big\(\\rho\_\{q\}\\\|W\_\{q\}\\\|\_\{\\mathrm\{op\}\}\\big\)\\cdot\\mathrm\{Lip\}\(G\_\{p:q\-1\}\)\\cdot\\\|u\-u^\{\\prime\}\\\|\_\{2\}≤\(ρq​‖Wq‖op\)​\(∏h=pq−1ρh​‖Wh‖op\)​‖u−u′‖2,\\displaystyle\\leq\\left\(\\rho\_\{q\}\\\|W\_\{q\}\\\|\_\{\\mathrm\{op\}\}\\right\)\\left\(\\prod\_\{h=p\}^\{q\-1\}\\rho\_\{h\}\\\|W\_\{h\}\\\|\_\{\\mathrm\{op\}\}\\right\)\\\|u\-u^\{\\prime\}\\\|\_\{2\},Takingp=l\+1p=l\+1andq=Lq=Lgives

∥GWl\+1:L\(a\)−GWl\+1:L\(a′\)∥2≤\(∏h=l\+1Lρh∥Wh∥op\)∥a−a′∥2\.\\\|G\_\{W\_\{l\+1:L\}\}\(a\)\-G\_\{W\_\{l\+1:L\}\}\(a^\{\\prime\}\)\\\|\_\{2\}\\leq\\left\(\\prod\_\{h=l\+1\}^\{L\}\\rho\_\{h\}\\\|W\_\{h\}\\\|\_\{\\mathrm\{op\}\}\\right\)\\\|a\-a^\{\\prime\}\\\|\_\{2\}\.Denoteαl​\(W\):=∏h=l\+1Lρh​‖Wh‖op\\alpha\_\{l\}\(W\):=\\prod\_\{h=l\+1\}^\{L\}\\rho\_\{h\}\\\|W\_\{h\}\\\|\_\{\\mathrm\{op\}\}for short\.

Using thatℓ⁡\(y^,y\)\\ell\(\\hat\{y\},y\)isρ0\\rho\_\{0\}\-Lipschitz in\(y^,y\)\(\\hat\{y\},y\)under the Euclidean norm, for any\(a,y\),\(a′,y′\)\(a,y\),\(a^\{\\prime\},y^\{\\prime\}\)we have

\|fW,l​\(a,y\)−fW,l​\(a′,y′\)\|\\displaystyle\|f\_\{W,l\}\(a,y\)\-f\_\{W,l\}\(a^\{\\prime\},y^\{\\prime\}\)\|=\|ℓ\(GWl\+1:L\(a\),y\)−ℓ\(GWl\+1:L\(a′\),y′\)\|\\displaystyle=\\big\|\\ell\(G\_\{W\_\{l\+1:L\}\}\(a\),y\)\-\\ell\(G\_\{W\_\{l\+1:L\}\}\(a^\{\\prime\}\),y^\{\\prime\}\)\\big\|≤ρ0∥GWl\+1:L\(a\)−GWl\+1:L\(a′\)∥22\+∥y−y′∥22\\displaystyle\\leq\\rho\_\{0\}\\sqrt\{\\\|G\_\{W\_\{l\+1:L\}\}\(a\)\-G\_\{W\_\{l\+1:L\}\}\(a^\{\\prime\}\)\\\|\_\{2\}^\{2\}\+\\\|y\-y^\{\\prime\}\\\|\_\{2\}^\{2\}\}≤ρ0​αl​\(W\)2​‖a−a′‖22\+‖y−y′‖22\.\\displaystyle\\leq\\rho\_\{0\}\\sqrt\{\\alpha\_\{l\}\(W\)^\{2\}\\\|a\-a^\{\\prime\}\\\|\_\{2\}^\{2\}\+\\\|y\-y^\{\\prime\}\\\|\_\{2\}^\{2\}\}\.Ifαl​\(W\)≤1\\alpha\_\{l\}\(W\)\\leq 1, thenαl​\(W\)2​‖a−a′‖22\+‖y−y′‖22≤‖a−a′‖22\+‖y−y′‖22\\alpha\_\{l\}\(W\)^\{2\}\\\|a\-a^\{\\prime\}\\\|\_\{2\}^\{2\}\+\\\|y\-y^\{\\prime\}\\\|\_\{2\}^\{2\}\\leq\\\|a\-a^\{\\prime\}\\\|\_\{2\}^\{2\}\+\\\|y\-y^\{\\prime\}\\\|\_\{2\}^\{2\}\. Ifαl​\(W\)≥1\\alpha\_\{l\}\(W\)\\geq 1, thenαl​\(W\)2​‖a−a′‖22\+‖y−y′‖22≤αl​\(W\)2​\(‖a−a′‖22\+‖y−y′‖22\)\\alpha\_\{l\}\(W\)^\{2\}\\\|a\-a^\{\\prime\}\\\|\_\{2\}^\{2\}\+\\\|y\-y^\{\\prime\}\\\|\_\{2\}^\{2\}\\leq\\alpha\_\{l\}\(W\)^\{2\}\(\\\|a\-a^\{\\prime\}\\\|\_\{2\}^\{2\}\+\\\|y\-y^\{\\prime\}\\\|\_\{2\}^\{2\}\)\. In both cases,

αl​\(W\)2​‖a−a′‖22\+‖y−y′‖22≤\(1∨αl​\(W\)\)​‖a−a′‖22\+‖y−y′‖22=\(1∨αl​\(W\)\)​d​\(\(a,y\),\(a′,y′\)\)\.\\sqrt\{\\alpha\_\{l\}\(W\)^\{2\}\\\|a\-a^\{\\prime\}\\\|\_\{2\}^\{2\}\+\\\|y\-y^\{\\prime\}\\\|\_\{2\}^\{2\}\}\\leq\(1\\vee\\alpha\_\{l\}\(W\)\)\\sqrt\{\\\|a\-a^\{\\prime\}\\\|\_\{2\}^\{2\}\+\\\|y\-y^\{\\prime\}\\\|\_\{2\}^\{2\}\}=\(1\\vee\\alpha\_\{l\}\(W\)\)\\,d\\big\(\(a,y\),\(a^\{\\prime\},y^\{\\prime\}\)\\big\)\.Therefore,

\|fW,l​\(a,y\)−fW,l​\(a′,y′\)\|≤ρ0​\(1∨αl​\(W\)\)​d​\(\(a,y\),\(a′,y′\)\)=ρ¯l​\(W\)​d​\(\(a,y\),\(a′,y′\)\),\|f\_\{W,l\}\(a,y\)\-f\_\{W,l\}\(a^\{\\prime\},y^\{\\prime\}\)\|\\leq\\rho\_\{0\}\(1\\vee\\alpha\_\{l\}\(W\)\)\\,d\\big\(\(a,y\),\(a^\{\\prime\},y^\{\\prime\}\)\\big\)=\\bar\{\\rho\}\_\{l\}\(W\)\\,d\\big\(\(a,y\),\(a^\{\\prime\},y^\{\\prime\}\)\\big\),which showsLip⁡\(fW,l\)≤ρ¯l​\(W\)\\mathrm\{Lip\}\(f\_\{W,l\}\)\\leq\\bar\{\\rho\}\_\{l\}\(W\)withρ¯l​\(W\)\\bar\{\\rho\}\_\{l\}\(W\)defined

ρ¯l​\(W\):=ρ0​\(1∨∏h=l\+1Lρh​‖Wh‖op\)\.\\bar\{\\rho\}\_\{l\}\(W\):=\\rho\_\{0\}\\left\(1\\vee\\prod\_\{h=l\+1\}^\{L\}\\rho\_\{h\}\\\|W\_\{h\}\\\|\_\{\\mathrm\{op\}\}\\right\)\.\(55\)For each taski∈\[T\]i\\in\[T\], by definition ofPAl,Y\|i,W1:lP\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}we have

𝔼Z∼𝒟iℓ\(W,Z\)=𝔼\(Al,Y\)∼PAl,Y\|i,W1:lfW,l\(Al,Y\)\.\\mathbb\{E\}\_\{Z\\sim\\mathcal\{D\}\_\{i\}\}\\ell\(W,Z\)=\\mathbb\{E\}\_\{\(A\_\{l\},Y\)\\sim P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\}f\_\{W,l\}\(A\_\{l\},Y\)\.Moreover, by definition ofP^Al,Y\|Si,W1:l\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\},

𝔼U∼P^Al,Y\|Si,W1:lfW,l\(U\)=\{1m​∑j∈ℐifW,l​\(Aj,li,Yji\)=1m​∑j∈ℐiℓ⁡\(W,Zji\),i∈\[T−1\],1n​∑j=1nfW,l​\(Aj,lT,YjT\)=1n​∑j=1nℓ⁡\(W,ZjT\),i=T\.\\mathbb\{E\}\_\{U\\sim\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\}\}f\_\{W,l\}\(U\)=\\begin\{cases\}\\frac\{1\}\{m\}\\sum\_\{j\\in\\mathcal\{I\}\_\{i\}\}f\_\{W,l\}\(A\_\{j,l\}^\{i\},Y\_\{j\}^\{i\}\)=\\frac\{1\}\{m\}\\sum\_\{j\\in\\mathcal\{I\}\_\{i\}\}\\ell\(W,Z\_\{j\}^\{i\}\),&i\\in\[T\-1\],\\\\\[4\.0pt\] \\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}f\_\{W,l\}\(A\_\{j,l\}^\{T\},Y\_\{j\}^\{T\}\)=\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}\\ell\(W,Z\_\{j\}^\{T\}\),&i=T\.\\end\{cases\}Substituting these identities into populationℒ⁡\(W\)\\mathcal\{L\}\(W\)and empirical risksℒ^​\(W\)\\hat\{\\mathcal\{L\}\}\(W\)yields, for the fixed realization,

ℒ\(W\)−ℒ^\(W\)=∑i=1T\(𝔼PAl,Y\|i,W1:l\[fW,l\]−𝔼P^Al,Y\|Si,W1:l\[fW,l\]\)\.\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)=\\sum\_\{i=1\}^\{T\}\\left\(\\mathbb\{E\}\_\{P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\}\\big\[f\_\{W,l\}\\big\]\-\\mathbb\{E\}\_\{\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\}\}\\big\[f\_\{W,l\}\\big\]\\right\)\.By Kantorovich–Rubinstein Lemma[A\.8](https://arxiv.org/html/2608.11690#A1.Thmtheorem8)for the11\-Wasserstein distance, for any Lipschitz functionffon\(𝒜l×𝒴,d\)\(\\mathcal\{A\}\_\{l\}\\times\\mathcal\{Y\},d\),

\|𝔼μ​f−𝔼ν​f\|≤Lip⁡\(f\)​𝒲1​\(μ,ν\)\.\\big\|\\mathbb\{E\}\_\{\\mu\}f\-\\mathbb\{E\}\_\{\\nu\}f\\big\|\\leq\\mathrm\{Lip\}\(f\)\\,\\mathcal\{W\}\_\{1\}\(\\mu,\\nu\)\.Applying this withf=fW,lf=f\_\{W,l\},μ=PAl,Y\|i,W1:l\\mu=P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}, andν=P^Al,Y\|Si,W1:l\\nu=\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\}, and usingLip⁡\(fW,l\)≤ρ¯l​\(W\)\\mathrm\{Lip\}\(f\_\{W,l\}\)\\leq\\bar\{\\rho\}\_\{l\}\(W\)from the previous steps, we obtain for eachi∈\[T\]i\\in\[T\]:

\|𝔼PAl,Y\|i,W1:lfW,l−𝔼P^Al,Y\|Si,W1:lfW,l\|≤ρ¯l\(W\)𝒲1\(P^Al,Y\|Si,W1:l,PAl,Y\|i,W1:l\)\.\\left\|\\mathbb\{E\}\_\{P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\}f\_\{W,l\}\-\\mathbb\{E\}\_\{\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\}\}f\_\{W,l\}\\right\|\\leq\\bar\{\\rho\}\_\{l\}\(W\)\\,\\mathcal\{W\}\_\{1\}\\\!\\Big\(\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\},\\,P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\Big\)\.Hence, by the triangle inequality,

\|ℒ\(W\)−ℒ^\(W\)\|≤ρ¯l\(W\)⋅∑i=1T𝒲1\(P^Al,Y\|Si,W1:l,PAl,Y\|i,W1:l\)\.\\big\|\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)\\big\|\\leq\\bar\{\\rho\}\_\{l\}\(W\)\\cdot\\sum\_\{i=1\}^\{T\}\\mathcal\{W\}\_\{1\}\\\!\\Big\(\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\},\\,P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\Big\)\.Now take expectation over all randomness, including the generally dependent training sequence lawP𝐔\(l\)∣W1:lP\_\{\\mathbf\{U\}^\{\(l\)\}\\mid W\_\{1:l\}\}:

\|genW\|=\|𝔼\[ℒ\(W\)−ℒ^\(W\)\]\|≤𝔼\|ℒ\(W\)−ℒ^\(W\)\|≤𝔼\[ρ¯l\(W\)⋅∑i=1T𝒲1\(P^Al,Y\|Si,W1:l,PAl,Y\|i,W1:l\)\]\.\\big\|\\mathrm\{gen\}\_\{W\}\\big\|=\\left\|\\mathbb\{E\}\\big\[\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)\\big\]\\right\|\\leq\\mathbb\{E\}\\big\|\\mathcal\{L\}\(W\)\-\\hat\{\\mathcal\{L\}\}\(W\)\\big\|\\leq\\mathbb\{E\}\\left\[\\bar\{\\rho\}\_\{l\}\(W\)\\cdot\\sum\_\{i=1\}^\{T\}\\mathcal\{W\}\_\{1\}\\\!\\Big\(\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\},\\,P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\Big\)\\right\]\.\(56\)Since the bound \([56](https://arxiv.org/html/2608.11690#A2.E56)\) holds for everyl∈\{0,…,L\}l\\in\\\{0,\\dots,L\\\}, taking the minimum overllgives,

\|genW\|≤minl∈\{0,…,L\}𝔼\[ρ¯l\(W\)⋅∑i=1T𝒲1\(P^Al,Y\|Si,W1:l,PAl,Y\|i,W1:l\)\]\.\\big\|\\mathrm\{gen\}\_\{W\}\\big\|\\;\\leq\\;\\min\_\{l\\in\\\{0,\\dots,L\\\}\}\\mathbb\{E\}\\left\[\\bar\{\\rho\}\_\{l\}\(W\)\\cdot\\sum\_\{i=1\}^\{T\}\\mathcal\{W\}\_\{1\}\\\!\\Big\(\\hat\{P\}\_\{A\_\{l\},Y\|S^\{i\},W\_\{1:l\}\},\\,P\_\{A\_\{l\},Y\|i,W\_\{1:l\}\}\\Big\)\\right\]\.\(57\)∎

### B\-EProof of Proposition[V\.1](https://arxiv.org/html/2608.11690#S5.Thmtheorem1)

###### Proposition[V\.1](https://arxiv.org/html/2608.11690#S5.Thmtheorem1)\(Restate\)\.

Fix a layerlland condition onW1:lW\_\{1:l\}\. At taskTT, we have

I\(𝐔T\(l\);ΘR\(T\)\|W1:l\)\\displaystyle I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta\_\{R\}^\{\(T\)\}\|W\_\{1:l\}\)≤I\(𝐔T\(l\);Θ0\(T\)\|W1:l\)\+I\(𝐔T\(l\);ΘR\(T\)\|Θ0\(T\),W1:l\)\\displaystyle\\leq I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta\_\{0\}^\{\(T\)\}\|W\_\{1:l\}\)\+I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta\_\{R\}^\{\(T\)\}\|\\Theta\_\{0\}^\{\(T\)\},W\_\{1:l\}\)=:HT\(l\)\+ΔT\(l\)\.\\displaystyle=:H\_\{T\}^\{\(l\)\}\+\\Delta\_\{T\}^\{\(l\)\}\.

###### Proof\.

Condition onW1:lW\_\{1:l\}and suppress it in the notation\. By data processing,I⁡\(𝐔T\(l\),ΘR\(T\)\)≤I⁡\(𝐔T\(l\),\(Θ0\(T\),ΘR\(T\)\)\)I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta^\{\(T\)\}\_\{R\}\)\\leq I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\(\\Theta^\{\(T\)\}\_\{0\},\\Theta^\{\(T\)\}\_\{R\}\)\)\. Applying the chain rule for mutual information yields

I⁡\(𝐔T\(l\),\(Θ0\(T\),ΘR\(T\)\)\)=I⁡\(𝐔T\(l\),Θ0\(T\)\)\+I⁡\(𝐔T\(l\);ΘR\(T\)∣Θ0\(T\)\),I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\(\\Theta^\{\(T\)\}\_\{0\},\\Theta^\{\(T\)\}\_\{R\}\)\)~=~I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta^\{\(T\)\}\_\{0\}\)\+I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta^\{\(T\)\}\_\{R\}\\mid\\Theta^\{\(T\)\}\_\{0\}\),which proves the Proposition\. ∎

### B\-FProof of Corollary[V\.2](https://arxiv.org/html/2608.11690#S5.Thmtheorem2)

###### Corollary[V\.2](https://arxiv.org/html/2608.11690#S5.Thmtheorem2)\(Restate\)\.

Fix a layerlland condition onW1:lW\_\{1:l\}\. For anyT≥2T\\geq 2, the heritage term admits the following cumulative\-budget upper bound:

HT\(l\)=I\(𝐔T\(l\);Θ0\(T\)∣W1:l\)≤∑t=1T−1Δt\(l\),H\_\{T\}^\{\(l\)\}=I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta\_\{0\}^\{\(T\)\}\\mid W\_\{1:l\}\)\\leq\\sum\_\{t=1\}^\{T\-1\}\\Delta\_\{t\}^\{\(l\)\},whereΔt\(l\):=I\(𝐔t\(l\);ΘR\(t\)∣Θ0\(t\),W1:l\)\\Delta\_\{t\}^\{\(l\)\}:=I\(\\mathbf\{U\}\_\{t\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\\mid\\Theta\_\{0\}^\{\(t\)\},W\_\{1:l\}\)is the within\-task information increment in Proposition[V\.1](https://arxiv.org/html/2608.11690#S5.Thmtheorem1)applied to tasktt\.

###### Proof\.

Throughout the proof we implicitly condition onW1:lW\_\{1:l\}to simplify notation\. By the recursionΘ0\(T\)=ΘR\(T−1\)\\Theta\_\{0\}^\{\(T\)\}=\\Theta\_\{R\}^\{\(T\-1\)\}, we haveHT\(l\)=I⁡\(𝐔T\(l\),ΘR\(T−1\)\)\.H\_\{T\}^\{\(l\)\}=I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta\_\{R\}^\{\(T\-1\)\}\)\.The task\-tttraining sequence𝐔t\(l\)\\mathbf\{U\}\_\{t\}^\{\(l\)\}is generated by sampling from \(i\) the current\-task dataset and \(ii\) a replay memoryℳ1:t−1\\mathcal\{M\}^\{1:t\-1\}built from the past, and this sampling procedure does not depend onΘR\(t−1\)\\Theta\_\{R\}^\{\(t\-1\)\}onceℳ1:t−1\\mathcal\{M\}^\{1:t\-1\}is fixed\. Then,

𝐔t\(l\)⟂ΘR\(t−1\)∣ℳ1:t−1\.\\mathbf\{U\}\_\{t\}^\{\(l\)\}\\perp\\Theta\_\{R\}^\{\(t\-1\)\}\\mid\\mathcal\{M\}^\{1:t\-1\}\.Therefore, at taskTT, we have𝐔T\(l\)⟂ΘR\(T−1\)∣ℳ1:T−1\\mathbf\{U\}\_\{T\}^\{\(l\)\}\\perp\\Theta\_\{R\}^\{\(T\-1\)\}\\mid\\mathcal\{M\}^\{1:T\-1\}, which implies the Markov chain

𝐔T\(l\)→ℳ1:T−1→ΘR\(T−1\)\.\\mathbf\{U\}\_\{T\}^\{\(l\)\}\\to\\mathcal\{M\}^\{1:T\-1\}\\to\\Theta\_\{R\}^\{\(T\-1\)\}\.By the data processing inequality,

HT\(l\)=I\(𝐔T\(l\);ΘR\(T−1\)\)≤I\(ℳ1:T−1;ΘR\(T−1\)\)\.H\_\{T\}^\{\(l\)\}=I\(\\mathbf\{U\}\_\{T\}^\{\(l\)\};\\Theta\_\{R\}^\{\(T\-1\)\}\)\\;\\leq\\;I\(\\mathcal\{M\}^\{1:T\-1\};\\Theta\_\{R\}^\{\(T\-1\)\}\)\.\(58\)Moreover, sinceℳ1:T−1\\mathcal\{M\}^\{1:T\-1\}is measurable with respect to𝐔1:T−1\(l\)\\mathbf\{U\}\_\{1:T\-1\}^\{\(l\)\}, we have Markov chains again,

ΘR\(T−1\)→𝐔1:T−1\(l\)→ℳ1:T−1\.\\Theta\_\{R\}^\{\(T\-1\)\}\\to\\mathbf\{U\}\_\{1:T\-1\}^\{\(l\)\}\\to\\mathcal\{M\}^\{1:T\-1\}\.By the data processing inequality,

I\(ℳ1:T−1;ΘR\(T−1\)\)≤I\(𝐔1:T−1\(l\);ΘR\(T−1\)\)\.I\(\\mathcal\{M\}^\{1:T\-1\};\\Theta\_\{R\}^\{\(T\-1\)\}\)\\;\\leq\\;I\(\\mathbf\{U\}\_\{1:T\-1\}^\{\(l\)\};\\Theta\_\{R\}^\{\(T\-1\)\}\)\.\(59\)
For eacht≥1t\\geq 1, defineJt:=I\(𝐔1:t\(l\);ΘR\(t\)\)\.J\_\{t\}:=I\(\\mathbf\{U\}\_\{1:t\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\)\.We claim that for everyt≥1t\\geq 1,

Jt≤Jt−1\+Δt\(l\)\.J\_\{t\}\\;\\leq\\;J\_\{t\-1\}\+\\Delta\_\{t\}^\{\(l\)\}\.\(60\)Next, our goal is to prove \([60](https://arxiv.org/html/2608.11690#A2.E60)\)\. Based on the monotonicity of mutual information sinceΘR\(t\)\\Theta\_\{R\}^\{\(t\)\}is a component of\(Θ0\(t\),ΘR\(t\)\)\(\\Theta\_\{0\}^\{\(t\)\},\\Theta\_\{R\}^\{\(t\)\}\), we have

I\(𝐔1:t\(l\);ΘR\(t\)\)≤I\(𝐔1:t\(l\);Θ0\(t\),ΘR\(t\)\)\.I\(\\mathbf\{U\}\_\{1:t\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\)\\leq I\(\\mathbf\{U\}\_\{1:t\}^\{\(l\)\};\\Theta\_\{0\}^\{\(t\)\},\\Theta\_\{R\}^\{\(t\)\}\)\.\(61\)Apply the chain rule to the RHS:

I\(𝐔1:t\(l\);Θ0\(t\),ΘR\(t\)\)\\displaystyle I\(\\mathbf\{U\}\_\{1:t\}^\{\(l\)\};\\Theta\_\{0\}^\{\(t\)\},\\Theta\_\{R\}^\{\(t\)\}\)=I\(𝐔1:t\(l\);Θ0\(t\)\)\+I\(𝐔1:t\(l\);ΘR\(t\)∣Θ0\(t\)\)\.\\displaystyle=I\(\\mathbf\{U\}\_\{1:t\}^\{\(l\)\};\\Theta\_\{0\}^\{\(t\)\}\)\+I\(\\mathbf\{U\}\_\{1:t\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\\mid\\Theta\_\{0\}^\{\(t\)\}\)\.\(62\)
We handle the two terms in \([62](https://arxiv.org/html/2608.11690#A2.E62)\) separately\.

SinceΘ0\(t\)=ΘR\(t−1\)\\Theta\_\{0\}^\{\(t\)\}=\\Theta\_\{R\}^\{\(t\-1\)\}is fully determined by the past training procedure, conditioning on𝐔1:t−1\(l\)\\mathbf\{U\}\_\{1:t\-1\}^\{\(l\)\}leaves no additional dependence on𝐔t\(l\)\\mathbf\{U\}\_\{t\}^\{\(l\)\}, thus:

I\(𝐔t\(l\);Θ0\(t\)∣𝐔1:t−1\(l\)\)=0\.I\(\\mathbf\{U\}\_\{t\}^\{\(l\)\};\\Theta\_\{0\}^\{\(t\)\}\\mid\\mathbf\{U\}\_\{1:t\-1\}^\{\(l\)\}\)=0\.Then, by the chain rule,

I\(𝐔1:t\(l\);Θ0\(t\)\)\\displaystyle I\(\\mathbf\{U\}\_\{1:t\}^\{\(l\)\};\\Theta\_\{0\}^\{\(t\)\}\)=I\(𝐔1:t−1\(l\);Θ0\(t\)\)\+I\(𝐔t\(l\);Θ0\(t\)∣𝐔1:t−1\(l\)\)\\displaystyle=I\(\\mathbf\{U\}\_\{1:t\-1\}^\{\(l\)\};\\Theta\_\{0\}^\{\(t\)\}\)\+I\(\\mathbf\{U\}\_\{t\}^\{\(l\)\};\\Theta\_\{0\}^\{\(t\)\}\\mid\\mathbf\{U\}\_\{1:t\-1\}^\{\(l\)\}\)=I\(𝐔1:t−1\(l\);Θ0\(t\)\)=I\(𝐔1:t−1\(l\);ΘR\(t−1\)\)=Jt−1\.\\displaystyle=I\(\\mathbf\{U\}\_\{1:t\-1\}^\{\(l\)\};\\Theta\_\{0\}^\{\(t\)\}\)=I\(\\mathbf\{U\}\_\{1:t\-1\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\-1\)\}\)=J\_\{t\-1\}\.\(63\)Decompose𝐔1:t\(l\)=\(𝐔1:t−1\(l\),𝐔t\(l\)\)\\mathbf\{U\}\_\{1:t\}^\{\(l\)\}=\(\\mathbf\{U\}\_\{1:t\-1\}^\{\(l\)\},\\mathbf\{U\}\_\{t\}^\{\(l\)\}\)and apply the chain rule:

I\(𝐔1:t\(l\);ΘR\(t\)∣Θ0\(t\)\)\\displaystyle I\(\\mathbf\{U\}\_\{1:t\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\\mid\\Theta\_\{0\}^\{\(t\)\}\)=I\(𝐔t\(l\);ΘR\(t\)∣Θ0\(t\)\)\+I\(𝐔1:t−1\(l\);ΘR\(t\)∣Θ0\(t\),𝐔t\(l\)\)\.\\displaystyle=I\(\\mathbf\{U\}\_\{t\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\\mid\\Theta\_\{0\}^\{\(t\)\}\)\+I\(\\mathbf\{U\}\_\{1:t\-1\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\\mid\\Theta\_\{0\}^\{\(t\)\},\\mathbf\{U\}\_\{t\}^\{\(l\)\}\)\.\(64\)For eachtt, the final parameterΘR\(t\)\\Theta\_\{R\}^\{\(t\)\}is conditionally independent of the past training sequences𝐔1:t−1\(l\)\\mathbf\{U\}\_\{1:t\-1\}^\{\(l\)\}given the initializationΘ0\(t\)\\Theta\_\{0\}^\{\(t\)\}and the current training sequence𝐔t\(l\)\\mathbf\{U\}\_\{t\}^\{\(l\)\}:

ΘR\(t\)⟂𝐔1:t−1\(l\)∣\(Θ0\(t\),𝐔t\(l\),W1:l\)\.\\Theta\_\{R\}^\{\(t\)\}\\perp\\mathbf\{U\}\_\{1:t\-1\}^\{\(l\)\}\\mid\(\\Theta\_\{0\}^\{\(t\)\},\\mathbf\{U\}\_\{t\}^\{\(l\)\},W\_\{1:l\}\)\.Therefore the second termI\(𝐔1:t−1\(l\);ΘR\(t\)∣Θ0\(t\),𝐔t\(l\)\)I\(\\mathbf\{U\}\_\{1:t\-1\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\\mid\\Theta\_\{0\}^\{\(t\)\},\\mathbf\{U\}\_\{t\}^\{\(l\)\}\)on the RHS of \([64](https://arxiv.org/html/2608.11690#A2.E64)\) is zero, hence

I\(𝐔1:t\(l\);ΘR\(t\)∣Θ0\(t\)\)=I\(𝐔t\(l\);ΘR\(t\)∣Θ0\(t\)\)=Δt\(l\)\.I\(\\mathbf\{U\}\_\{1:t\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\\mid\\Theta\_\{0\}^\{\(t\)\}\)=I\(\\mathbf\{U\}\_\{t\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\\mid\\Theta\_\{0\}^\{\(t\)\}\)=\\Delta\_\{t\}^\{\(l\)\}\.\(65\)
Plugging \([63](https://arxiv.org/html/2608.11690#A2.E63)\) and \([65](https://arxiv.org/html/2608.11690#A2.E65)\) into \([62](https://arxiv.org/html/2608.11690#A2.E62)\), and then using \([61](https://arxiv.org/html/2608.11690#A2.E61)\), yields

Jt=I\(𝐔1:t\(l\);ΘR\(t\)\)≤I\(𝐔1:t\(l\);Θ0\(t\),ΘR\(t\)\)=Jt−1\+Δt\(l\),J\_\{t\}=I\(\\mathbf\{U\}\_\{1:t\}^\{\(l\)\};\\Theta\_\{R\}^\{\(t\)\}\)\\leq I\(\\mathbf\{U\}\_\{1:t\}^\{\(l\)\};\\Theta\_\{0\}^\{\(t\)\},\\Theta\_\{R\}^\{\(t\)\}\)=J\_\{t\-1\}\+\\Delta\_\{t\}^\{\(l\)\},which establishes \([60](https://arxiv.org/html/2608.11690#A2.E60)\)\. Iterating \([60](https://arxiv.org/html/2608.11690#A2.E60)\) fromt=1t=1tot=T−1t=T\-1gives

JT−1≤J0\+∑t=1T−1Δt\(l\)\.J\_\{T\-1\}\\leq J\_\{0\}\+\\sum\_\{t=1\}^\{T\-1\}\\Delta\_\{t\}^\{\(l\)\}\.Because parameters are randomly initialized before network training, thereforeΘ0\(1\)\\Theta\_\{0\}^\{\(1\)\}is independent of𝐔1\(l\)\\mathbf\{U\}\_\{1\}^\{\(l\)\}, which impliesJ0=I\(𝐔1:0\(l\);ΘR\(0\)\)=0J\_\{0\}=I\(\\mathbf\{U\}\_\{1:0\}^\{\(l\)\};\\Theta\_\{R\}^\{\(0\)\}\)=0, we obtain

JT−1≤∑t=1T−1Δt\(l\)\.J\_\{T\-1\}\\leq\\sum\_\{t=1\}^\{T\-1\}\\Delta\_\{t\}^\{\(l\)\}\.\(66\)
From \([58](https://arxiv.org/html/2608.11690#A2.E58)\)–\([59](https://arxiv.org/html/2608.11690#A2.E59)\) we have

HT\(l\)≤I\(𝐔1:T−1\(l\);ΘR\(T−1\)\)=JT−1\.H\_\{T\}^\{\(l\)\}\\leq I\(\\mathbf\{U\}\_\{1:T\-1\}^\{\(l\)\};\\Theta\_\{R\}^\{\(T\-1\)\}\)=J\_\{T\-1\}\.Combining with \([66](https://arxiv.org/html/2608.11690#A2.E66)\) yields the desired result:

HT\(l\)≤∑t=1T−1Δt\(l\)\.H\_\{T\}^\{\(l\)\}\\leq\\sum\_\{t=1\}^\{T\-1\}\\Delta\_\{t\}^\{\(l\)\}\.∎

### B\-GProof of Theorem[V\.3](https://arxiv.org/html/2608.11690#S5.Thmtheorem3)

###### Theorem[V\.3](https://arxiv.org/html/2608.11690#S5.Thmtheorem3)\(Restate\)\.

Fix a split layerl∈\{0,…,L−1\}l\\in\\\{0,\\ldots,L\-1\\\}and condition on the base parametersW1:lW\_\{1:l\}, so that the suffix network remains nonempty\. Assume the loss isσ\\sigma\-subgaussian as in Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)\. For tasktttrained by SGLD forRtR\_\{t\}steps, we have

\|genW\(t\)\|≤\(t−1\)​2​σ2​𝒦t\(l\)\+2​σ2Neff,t​\(𝒞t\(l\)\+∑s=1t∑r=1Rs𝔼⁡\[12​log​det\(I\+ηs,r2τs,r2​Ms,r\(l\)\)\]\),\\displaystyle\\big\|\\mathrm\{gen\}\_\{W^\{\(t\)\}\}\\big\|\\leq\(t\-1\)\\sqrt\{2\\sigma^\{2\}\\mathcal\{K\}\_\{t\}^\{\(l\)\}\}\+\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\},t\}\}\\Big\(\\mathcal\{C\}\_\{t\}^\{\(l\)\}\+\\sum\_\{s=1\}^\{t\}\\sum\_\{r=1\}^\{R\_\{s\}\}\\mathbb\{E\}\\\!\\left\[\\frac\{1\}\{2\}\\log\\det\\\!\\Big\(I\+\\frac\{\\eta\_\{s,r\}^\{2\}\}\{\\tau\_\{s,r\}^\{2\}\}M^\{\(l\)\}\_\{s,r\}\\Big\)\\right\]\\Big\)\},whereMs,r\(l\)=𝔼\[Gs,rGs,r⊤\|Θr−1\(s\),W1:l\]M^\{\(l\)\}\_\{s,r\}=\\mathbb\{E\}\\\!\\big\[G\_\{s,r\}G\_\{s,r\}^\{\\top\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\\big\], and𝒦t\(l\)\\mathcal\{K\}\_\{t\}^\{\(l\)\},𝒞t\(l\)\\mathcal\{C\}\_\{t\}^\{\(l\)\}, andNeff,tN\_\{\\mathrm\{eff\},t\}denote the corresponding quantities of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)with the task horizonTTreplaced bytt\.

###### Proof\.

Throughout the proof we implicitly fix the split layerlland condition onW1:lW\_\{1:l\}to simplify notation\. Apply Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)to thett\-task scenario by substitutingttforTTin the statement of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)\. Then for any fixed split layerll,

\|genW\(t\)\|≤\(t−1\)​2​σ2​𝒦t\(l\)\+2​σ2Neff,t​\(𝒮t\(l\)\+𝒫t\(l\)−ℛt\(l\)\+𝒞t\(l\)\)\.\\big\|\\mathrm\{gen\}\_\{W^\{\(t\)\}\}\\big\|\\leq\(t\-1\)\\sqrt\{2\\sigma^\{2\}\\mathcal\{K\}\_\{t\}^\{\(l\)\}\}\+\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\},t\}\}\\Big\(\\mathcal\{S\}\_\{t\}^\{\(l\)\}\+\\mathcal\{P\}\_\{t\}^\{\(l\)\}\-\\mathcal\{R\}\_\{t\}^\{\(l\)\}\+\\mathcal\{C\}\_\{t\}^\{\(l\)\}\\Big\)\}\.\(67\)By the interaction\-information identity used in Section[V\-A](https://arxiv.org/html/2608.11690#S5.SS1), the SPS combination equals a conditional mutual information between the task\-tttraining sequence and the final upper parameters:

𝒮t\(l\)\+𝒫t\(l\)−ℛt\(l\)=I\(𝐔t\(l\);ΘRt\(t\)∣W1:l\)\.\\mathcal\{S\}\_\{t\}^\{\(l\)\}\+\\mathcal\{P\}\_\{t\}^\{\(l\)\}\-\\mathcal\{R\}\_\{t\}^\{\(l\)\}=I\\Big\(\\mathbf\{U\}\_\{t\}^\{\(l\)\};\\Theta^\{\(t\)\}\_\{R\_\{t\}\}\\mid W\_\{1:l\}\\Big\)\.\(68\)Thus, it suffices to upper boundI\(𝐔t\(l\);ΘRt\(t\)∣W1:l\)I\\Big\(\\mathbf\{U\}\_\{t\}^\{\(l\)\};\\Theta^\{\(t\)\}\_\{R\_\{t\}\}\\mid W\_\{1:l\}\\Big\)by a cumulative, algorithmic quantity\. For each task indexs≥1s\\geq 1, define the within\-task increment

Δs\(l\):=I\(𝐔s\(l\);ΘRs\(s\)∣Θ0\(s\),W1:l\),\\Delta\_\{s\}^\{\(l\)\}:=I\\left\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta\_\{R\_\{s\}\}^\{\(s\)\}\\mid\\Theta\_\{0\}^\{\(s\)\},W\_\{1:l\}\\right\),and the heritage term

Hs\(l\):=I\(𝐔s\(l\);Θ0\(s\)∣W1:l\),whereΘ0\(s\)=ΘRs−1\(s−1\)fors≥2\.H\_\{s\}^\{\(l\)\}:=I\\left\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta\_\{0\}^\{\(s\)\}\\mid W\_\{1:l\}\\right\),\\qquad\\text\{where \}\\Theta^\{\(s\)\}\_\{0\}=\\Theta^\{\(s\-1\)\}\_\{R\_\{s\-1\}\}\\ \\text\{for\}\\ s\\geq 2\.Applying Proposition[V\.1](https://arxiv.org/html/2608.11690#S5.Thmtheorem1)at tasksswe have decomposition,

I\(𝐔s\(l\);ΘRs\(s\)∣W1:l\)≤Hs\(l\)\+Δs\(l\)\.I\\Big\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta^\{\(s\)\}\_\{R\_\{s\}\}\\mid W\_\{1:l\}\\Big\)\\leq H\_\{s\}^\{\(l\)\}\+\\Delta\_\{s\}^\{\(l\)\}\.\(69\)Applying Corollary[V\.2](https://arxiv.org/html/2608.11690#S5.Thmtheorem2), we can obtain a sum\-to\-ttrecursion,

I\(𝐔t\(l\);ΘRt\(t\)∣W1:l\)≤∑s=1t−1Δs\(l\)\+Δt\(l\)=∑s=1tΔs\(l\)\.I\\Big\(\\mathbf\{U\}\_\{t\}^\{\(l\)\};\\Theta^\{\(t\)\}\_\{R\_\{t\}\}\\mid W\_\{1:l\}\\Big\)\\leq\\sum\_\{s=1\}^\{t\-1\}\\Delta\_\{s\}^\{\(l\)\}\+\\Delta\_\{t\}^\{\(l\)\}=\\sum\_\{s=1\}^\{t\}\\Delta\_\{s\}^\{\(l\)\}\.\(70\)Fix a tasks∈\{1,…,t\}s\\in\\\{1,\\dots,t\\\}\. LetΘ0:Rs\(s\):=\(Θ0\(s\),Θ1\(s\),…,ΘRs\(s\)\)\\Theta^\{\(s\)\}\_\{0:R\_\{s\}\}:=\(\\Theta^\{\(s\)\}\_\{0\},\\Theta^\{\(s\)\}\_\{1\},\\dots,\\Theta^\{\(s\)\}\_\{R\_\{s\}\}\)denote the full within\-task trajectory\. SinceΘRs\(s\)\\Theta^\{\(s\)\}\_\{R\_\{s\}\}is a coordinate ofΘ\(s\)0:Rs\\Theta^\{\(s\)\}\_\{0:R\_\{s\}\}, by monotonicity of mutual information,

I\(𝐔s\(l\);ΘRs\(s\)\|Θ0\(s\),W1:l\)≤I\(𝐔s\(l\);Θ0:Rs\(s\)\|Θ0\(s\),W1:l\)\.I\\\!\\Big\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta^\{\(s\)\}\_\{R\_\{s\}\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\\leq I\\\!\\Big\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta^\{\(s\)\}\_\{0:R\_\{s\}\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\.Apply the chain rule:

I\(𝐔s\(l\);Θ0:Rs\(s\)\|Θ0\(s\),W1:l\)=∑r=1RsI\(𝐔s\(l\);Θr\(s\)\|Θ0:r−1\(s\),Θ0\(s\),W1:l\),I\\\!\\Big\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta^\{\(s\)\}\_\{0:R\_\{s\}\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)=\\sum\_\{r=1\}^\{R\_\{s\}\}I\\\!\\Big\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{0:r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\),where there is nor=0r=0term becauseΘ0\(s\)\\Theta^\{\(s\)\}\_\{0\}is already conditioned on\. For eachr≥1r\\geq 1, write the conditional mutual information as a differential entropy difference:

I\(𝐔s\(l\);Θr\(s\)\|Θ0:r−1\(s\),Θ0\(s\),W1:l\)\\displaystyle I\\\!\\Big\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{0:r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)=h\(Θr\(s\)\|Θ0:r−1\(s\),Θ0\(s\),W1:l\)−h\(Θr\(s\)\|𝐔s\(l\),Θ0:r−1\(s\),Θ0\(s\),W1:l\)\.\\displaystyle=h\\\!\\Big\(\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{0:r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\-h\\\!\\Big\(\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\mathbf\{U\}\_\{s\}^\{\(l\)\},\\Theta^\{\(s\)\}\_\{0:r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\.First, conditioning reduces entropy, hence

h\(Θr\(s\)\|Θ0:r−1\(s\),Θ0\(s\),W1:l\)≤h\(Θr\(s\)\|Θr−1\(s\),W1:l\)\.h\\\!\\Big\(\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{0:r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\\leq h\\\!\\Big\(\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\\Big\)\.Second, by the Markov structure of the update rule, given\(𝐔s\(l\),Θr−1\(s\),W1:l\)\(\\mathbf\{U\}\_\{s\}^\{\(l\)\},\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\), the next iterateΘr\(s\)\\Theta^\{\(s\)\}\_\{r\}is generated using only the freshly sampled minibatchBs,rB\_\{s,r\}\(sampled from𝐔s\(l\)\\mathbf\{U\}\_\{s\}^\{\(l\)\}\) and the fresh Gaussian noiseNs,rN\_\{s,r\}, so it is conditionally independent ofΘ\(s\)0:r−2\\Theta^\{\(s\)\}\_\{0:r\-2\}\. Therefore

h\(Θr\(s\)\|𝐔s\(l\),Θ0:r−1\(s\),Θ0\(s\),W1:l\)=h\(Θr\(s\)\|𝐔s\(l\),Θr−1\(s\),W1:l\)\.h\\\!\\Big\(\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\mathbf\{U\}\_\{s\}^\{\(l\)\},\\Theta^\{\(s\)\}\_\{0:r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)=h\\\!\\Big\(\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\mathbf\{U\}\_\{s\}^\{\(l\)\},\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\\Big\)\.Combining these two displays gives, for eachrr,

I\(𝐔s\(l\);Θr\(s\)\|Θ0:r−1\(s\),Θ0\(s\),W1:l\)≤I\(𝐔s\(l\);Θr\(s\)\|Θr−1\(s\),Θ0\(s\),W1:l\)\.I\\\!\\Big\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{0:r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\\leq I\\\!\\Big\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\.Summing overrryields

Δs\(l\)=I\(𝐔s\(l\);ΘRs\(s\)\|Θ0\(s\),W1:l\)≤∑r=1RsI\(𝐔s\(l\);Θr\(s\)\|Θr−1\(s\),Θ0\(s\),W1:l\)\.\\Delta\_\{s\}^\{\(l\)\}=I\\\!\\Big\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta^\{\(s\)\}\_\{R\_\{s\}\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\\leq\\sum\_\{r=1\}^\{R\_\{s\}\}I\\\!\\Big\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\.Now fixrrand apply data processing from the full data𝐔s\(l\)\\mathbf\{U\}\_\{s\}^\{\(l\)\}to the sampled minibatchBs,rB\_\{s,r\}: given\(Θr−1\(s\),W1:l\)\(\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\), the algorithm drawsBs,rB\_\{s,r\}by randomized sampling from𝐔s\(l\)\\mathbf\{U\}\_\{s\}^\{\(l\)\}, and the update mapping fromBs,rB\_\{s,r\}toΘr\(s\)\\Theta^\{\(s\)\}\_\{r\}depends on𝐔s\(l\)\\mathbf\{U\}\_\{s\}^\{\(l\)\}only throughBs,rB\_\{s,r\}\. Thus we have the Markov chain

𝐔s\(l\)⟶Bs,r⟶Θr\(s\)conditioned on\(Θr−1\(s\),Θ0\(s\),W1:l\),\\mathbf\{U\}\_\{s\}^\{\(l\)\}\\longrightarrow B\_\{s,r\}\\longrightarrow\\Theta^\{\(s\)\}\_\{r\}\\quad\\text\{conditioned on \}\(\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\),and hence

I\(𝐔s\(l\);Θr\(s\)\|Θr−1\(s\),Θ0\(s\),W1:l\)≤I\(Bs,r;Θr\(s\)\|Θr−1\(s\),Θ0\(s\),W1:l\)\.I\\\!\\Big\(\\mathbf\{U\}\_\{s\}^\{\(l\)\};\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\\leq I\\\!\\Big\(B\_\{s,r\};\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\.Therefore,

Δs\(l\)≤∑r=1RsI\(Bs,r;Θr\(s\)\|Θr−1\(s\),Θ0\(s\),W1:l\)\.\\Delta\_\{s\}^\{\(l\)\}\\leq\\sum\_\{r=1\}^\{R\_\{s\}\}I\\\!\\Big\(B\_\{s,r\};\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\.\(71\)It remains to upper bound each per\-step term\. Using the SGLD updateΘr\(s\)=Θr−1\(s\)\+ηs,r​Gs,r\+Ns,r\\Theta^\{\(s\)\}\_\{r\}=\\Theta^\{\(s\)\}\_\{r\-1\}\+\\eta\_\{s,r\}G\_\{s,r\}\+N\_\{s,r\}withNs,r∼𝒩⁡\(0,τs,r2​I\)N\_\{s,r\}\\sim\\mathcal\{N\}\(0,\\tau\_\{s,r\}^\{2\}I\), translation invariance of mutual information gives

I\(Bs,r;Θr\(s\)\|Θr−1\(s\),Θ0\(s\),W1:l\)=I\(Bs,r;ηs,rGs,r\+τs,r\|Θr−1\(s\),Θ0\(s\),W1:l\)\.I\\\!\\Big\(B\_\{s,r\};\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)=I\\\!\\Big\(B\_\{s,r\};\\eta\_\{s,r\}G\_\{s,r\}\+\\tau\_\{s,r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\.Write this as a differential entropy difference:

I\(Bs,r;ηs,rGs,r\+τs,r\|Θr−1\(s\),Θ0\(s\),W1:l\)\\displaystyle I\\\!\\Big\(B\_\{s,r\};\\eta\_\{s,r\}G\_\{s,r\}\+\\tau\_\{s,r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)=h\(ηs,rGs,r\+τs,r\|Θr−1\(s\),Θ0\(s\),W1:l\)−h\(Ns,r\),\\displaystyle=h\\\!\\Big\(\\eta\_\{s,r\}G\_\{s,r\}\+\\tau\_\{s,r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\-h\(N\_\{s,r\}\),becauseGs,rG\_\{s,r\}is a deterministic function of\(Bs,r,Θr−1\(s\),W1:l\)\(B\_\{s,r\},\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\)andNs,rN\_\{s,r\}is independent ofBs,rB\_\{s,r\}and Gaussian\. LetYs,r:=ηs,r​Gs,r\+Ns,rY\_\{s,r\}:=\\eta\_\{s,r\}G\_\{s,r\}\+N\_\{s,r\}and denoted:=dl\+1:Ld:=d\_\{l\+1:L\}\. By the Lemma[A\.10](https://arxiv.org/html/2608.11690#A1.Thmtheorem10)bound applied conditionally,

h\(Ys,r\|Θr−1\(s\),Θ0\(s\),W1:l\)≤12log\(\(2πe\)ddet\(𝔼\[Ys,rYs,r⊤\|Θr−1\(s\),Θ0\(s\),W1:l\]\)\)\.h\\\!\\Big\(Y\_\{s,r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\\leq\\frac\{1\}\{2\}\\log\\Big\(\(2\\pi e\)^\{d\}\\det\\Big\(\\mathbb\{E\}\\\!\\big\[Y\_\{s,r\}Y\_\{s,r\}^\{\\top\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\big\]\\Big\)\\Big\)\.Compute the conditional second moment using independence ofNs,rN\_\{s,r\}andGs,rG\_\{s,r\}given\(Θr−1\(s\),W1:l\)\(\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\)and𝔼⁡\[Ns,r\]=0\\mathbb\{E\}\[N\_\{s,r\}\]=0:

𝔼\[Ys,rYs,r⊤\|Θr−1\(s\),Θ0\(s\),W1:l\]\\displaystyle\\mathbb\{E\}\\\!\\big\[Y\_\{s,r\}Y\_\{s,r\}^\{\\top\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\big\]=ηs,r2𝔼\[Gs,rGs,r⊤\|Θr−1\(s\),W1:l\]\+𝔼\[Ns,rNs,r⊤\]\\displaystyle=\\eta\_\{s,r\}^\{2\}\\mathbb\{E\}\\\!\\big\[G\_\{s,r\}G\_\{s,r\}^\{\\top\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\\big\]\+\\mathbb\{E\}\[N\_\{s,r\}N\_\{s,r\}^\{\\top\}\]=ηs,r2​Ms,r\(l\)\+τs,r2​I\.\\displaystyle=\\eta\_\{s,r\}^\{2\}M^\{\(l\)\}\_\{s,r\}\+\\tau\_\{s,r\}^\{2\}I\.Alsoh⁡\(Ns,r\)=12​log⁡\(\(2​π​e\)d​det\(τs,r2​I\)\)h\(N\_\{s,r\}\)=\\frac\{1\}\{2\}\\log\\big\(\(2\\pi e\)^\{d\}\\det\(\\tau\_\{s,r\}^\{2\}I\)\\big\)sinceNs,r∼𝒩⁡\(0,τs,r2​I\)N\_\{s,r\}\\sim\\mathcal\{N\}\(0,\\tau\_\{s,r\}^\{2\}I\)\. Subtracting these two entropies yields

I\(Bs,r;Θr\(s\)\|Θr−1\(s\),Θ0\(s\),W1:l\)≤12logdet\(I\+ηs,r2τs,r2Ms,r\(l\)\)I\\\!\\Big\(B\_\{s,r\};\\Theta^\{\(s\)\}\_\{r\}\\,\\big\|\\,\\Theta^\{\(s\)\}\_\{r\-1\},\\Theta^\{\(s\)\}\_\{0\},W\_\{1:l\}\\Big\)\\leq\\frac\{1\}\{2\}\\log\\det\\\!\\Big\(I\+\\frac\{\\eta\_\{s,r\}^\{2\}\}\{\\tau\_\{s,r\}^\{2\}\}M^\{\(l\)\}\_\{s,r\}\\Big\)Taking expectation and summing overrrin \([71](https://arxiv.org/html/2608.11690#A2.E71)\), we obtain

Δs\(l\)≤∑r=1Rs𝔼⁡\[12​log​det\(I\+ηs,r2τs,r2​Ms,r\(l\)\)\]\.\\Delta\_\{s\}^\{\(l\)\}\\leq\\sum\_\{r=1\}^\{R\_\{s\}\}\\mathbb\{E\}\\\!\\left\[\\frac\{1\}\{2\}\\log\\det\\\!\\Big\(I\+\\frac\{\\eta\_\{s,r\}^\{2\}\}\{\\tau\_\{s,r\}^\{2\}\}M^\{\(l\)\}\_\{s,r\}\\Big\)\\right\]\.This holds for eachss, and hence

I\(𝐔t\(l\);ΘRt\(t\)\|W1:l\)≤∑s=1tΔs\(l\)≤∑s=1t∑r=1Rs𝔼\[12logdet\(I\+ηs,r2τs,r2Ms,r\(l\)\)\]\.I\\\!\\Big\(\\mathbf\{U\}\_\{t\}^\{\(l\)\};\\Theta^\{\(t\)\}\_\{R\_\{t\}\}\\,\\big\|\\,W\_\{1:l\}\\Big\)\\leq\\sum\_\{s=1\}^\{t\}\\Delta\_\{s\}^\{\(l\)\}\\leq\\sum\_\{s=1\}^\{t\}\\sum\_\{r=1\}^\{R\_\{s\}\}\\mathbb\{E\}\\\!\\left\[\\frac\{1\}\{2\}\\log\\det\\\!\\Big\(I\+\\frac\{\\eta\_\{s,r\}^\{2\}\}\{\\tau\_\{s,r\}^\{2\}\}M^\{\(l\)\}\_\{s,r\}\\Big\)\\right\]\.\(72\)Substitute \([68](https://arxiv.org/html/2608.11690#A2.E68)\) and \([72](https://arxiv.org/html/2608.11690#A2.E72)\) into \([67](https://arxiv.org/html/2608.11690#A2.E67)\) to obtain

\|genW\(t\)\|≤\(t−1\)​2​σ2​𝒦t\(l\)\+2​σ2Neff,t​\(𝒞t\(l\)\+∑s=1t∑r=1Rs𝔼⁡\[12​log​det\(I\+ηs,r2τs,r2​Ms,r\(l\)\)\]\)\.\\big\|\\mathrm\{gen\}\_\{W^\{\(t\)\}\}\\big\|\\leq\(t\-1\)\\sqrt\{2\\sigma^\{2\}\\mathcal\{K\}\_\{t\}^\{\(l\)\}\}\+\\sqrt\{\\frac\{2\\sigma^\{2\}\}\{N\_\{\\mathrm\{eff\},t\}\}\\Big\(\\mathcal\{C\}\_\{t\}^\{\(l\)\}\+\\sum\_\{s=1\}^\{t\}\\sum\_\{r=1\}^\{R\_\{s\}\}\\mathbb\{E\}\\\!\\left\[\\frac\{1\}\{2\}\\log\\det\\\!\\Big\(I\+\\frac\{\\eta\_\{s,r\}^\{2\}\}\{\\tau\_\{s,r\}^\{2\}\}M^\{\(l\)\}\_\{s,r\}\\Big\)\\right\]\\Big\)\}\.This completes the proof\. ∎

### B\-HProof of Proposition[V\.5](https://arxiv.org/html/2608.11690#S5.Thmtheorem5)

###### Proposition[V\.5](https://arxiv.org/html/2608.11690#S5.Thmtheorem5)\(Restate\)\.

LetAs,r\(l\):=I\+αs,r​Vs,r\(l\)≻0A\_\{s,r\}^\{\(l\)\}:=I\+\\alpha\_\{s,r\}V\_\{s,r\}^\{\(l\)\}\\succ 0withαs,r=ηs,r2/τs,r2\\alpha\_\{s,r\}=\\eta\_\{s,r\}^\{2\}/\\tau\_\{s,r\}^\{2\}\. Then

12​log​det\(I\+αs,r​Ms,r\(l\)\)=12​log​det\(As,r\(l\)\)⏟instability\+12​log⁡\(1\+αs,r​μs,r\(l\)⊤​\(As,r\(l\)\)−1​μs,r\(l\)\)⏟interaction cost\.\\displaystyle\\frac\{1\}\{2\}\\log\\det\\\!\\big\(I\+\\alpha\_\{s,r\}M\_\{s,r\}^\{\(l\)\}\\big\)=\\underbrace\{\\frac\{1\}\{2\}\\log\\det\\\!\\big\(A\_\{s,r\}^\{\(l\)\}\\big\)\}\_\{\\textbf\{instability\}\}\+\\underbrace\{\\frac\{1\}\{2\}\\log\\\!\\Big\(1\+\\alpha\_\{s,r\}\\,\\mu\_\{s,r\}^\{\(l\)\\top\}\\big\(A\_\{s,r\}^\{\(l\)\}\\big\)^\{\-1\}\\mu\_\{s,r\}^\{\(l\)\}\\Big\)\}\_\{\\textbf\{interaction cost\}\}\.

###### Proof\.

We condition on\(Θr−1\(s\),W1:l\)\(\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\)throughout\. By definition,

Ms,r\(l\)=𝔼⁡\[Gs,r​Gs,r⊤\]withGs,r=λs,old​Gs,rold\+λs,new​Gs,rnew\.M^\{\(l\)\}\_\{s,r\}=\\mathbb\{E\}\[G\_\{s,r\}G\_\{s,r\}^\{\\top\}\]\\quad\\text\{with\}\\quad G\_\{s,r\}=\\lambda\_\{s,\\mathrm\{old\}\}G^\{\\mathrm\{old\}\}\_\{s,r\}\+\\lambda\_\{s,\\mathrm\{new\}\}G^\{\\mathrm\{new\}\}\_\{s,r\}\.Since the old and new minibatches are sampled independently given the state,Gs,roldG^\{\\mathrm\{old\}\}\_\{s,r\}andGs,rnewG^\{\\mathrm\{new\}\}\_\{s,r\}are conditionally independent, hence

Cov⁡\(Gs,r\)=λs,old2​Cov​\(Gs,rold\)\+λs,new2​Cov​\(Gs,rnew\)=Vs,r\(l\)\.\\mathrm\{Cov\}\(G\_\{s,r\}\)=\\lambda\_\{s,\\mathrm\{old\}\}^\{2\}\\mathrm\{Cov\}\(G^\{\\mathrm\{old\}\}\_\{s,r\}\)\+\\lambda\_\{s,\\mathrm\{new\}\}^\{2\}\\mathrm\{Cov\}\(G^\{\\mathrm\{new\}\}\_\{s,r\}\)=V^\{\(l\)\}\_\{s,r\}\.Also𝔼⁡\[Gs,r\]=λs,old​μs,rold,\(l\)\+λs,new​μs,rnew,\(l\)=μs,r\(l\)\\mathbb\{E\}\[G\_\{s,r\}\]=\\lambda\_\{s,\\mathrm\{old\}\}\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\+\\lambda\_\{s,\\mathrm\{new\}\}\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}=\\mu^\{\(l\)\}\_\{s,r\}\. Using the identity𝔼⁡\[X​X⊤\]=Cov⁡\(X\)\+𝔼⁡\[X\]​𝔼​\[X\]⊤\\mathbb\{E\}\[XX^\{\\top\}\]=\\mathrm\{Cov\}\(X\)\+\\mathbb\{E\}\[X\]\\mathbb\{E\}\[X\]^\{\\top\}, we can obtain

I\+αs,r​Ms,r\(l\)=I\+αs,r​Vs,r\(l\)\+αs,r​μs,r\(l\)​μs,r\(l\)⊤=As,r\(l\)\+u​u⊤,u:=αs,r​μs,r\(l\)\.I\+\\alpha\_\{s,r\}M^\{\(l\)\}\_\{s,r\}=I\+\\alpha\_\{s,r\}V^\{\(l\)\}\_\{s,r\}\+\\alpha\_\{s,r\}\\mu^\{\(l\)\}\_\{s,r\}\\mu^\{\(l\)\\top\}\_\{s,r\}=A^\{\(l\)\}\_\{s,r\}\+uu^\{\\top\},\\quad u:=\\sqrt\{\\alpha\_\{s,r\}\}\\,\\mu^\{\(l\)\}\_\{s,r\}\.SinceAs,r\(l\)≻0A^\{\(l\)\}\_\{s,r\}\\succ 0, the matrix determinant lemmadet\(I\+A\+u​u⊤\)=det\(I\+A\)​\(1\+u⊤​\(I\+A\)−1​u\)\\det\(I\+A\+uu^\{\\top\}\)=\\det\(I\+A\)\\left\(1\+u^\{\\top\}\(I\+A\)^\{\-1\}u\\right\)gives

det\(As,r\(l\)\+u​u⊤\)=det\(As,r\(l\)\)​\(1\+u⊤​\(As,r\(l\)\)−1​u\)=det\(As,r\(l\)\)​\(1\+αs,r​μs,r\(l\)⊤​\(As,r\(l\)\)−1​μs,r\(l\)\)\.\\det\\\!\\big\(A^\{\(l\)\}\_\{s,r\}\+uu^\{\\top\}\\big\)=\\det\\\!\\big\(A^\{\(l\)\}\_\{s,r\}\\big\)\\Big\(1\+u^\{\\top\}\\big\(A^\{\(l\)\}\_\{s,r\}\\big\)^\{\-1\}u\\Big\)=\\det\\\!\\big\(A^\{\(l\)\}\_\{s,r\}\\big\)\\Big\(1\+\\alpha\_\{s,r\}\\mu^\{\(l\)\\top\}\_\{s,r\}\\big\(A^\{\(l\)\}\_\{s,r\}\\big\)^\{\-1\}\\mu^\{\(l\)\}\_\{s,r\}\\Big\)\.Taking12​log\\frac\{1\}\{2\}\\logof both sides completes the proof\. ∎

### B\-IProof of Corollary[V\.7](https://arxiv.org/html/2608.11690#S5.Thmtheorem7)

###### Corollary[V\.7](https://arxiv.org/html/2608.11690#S5.Thmtheorem7)\(Restate\)\.

Letμs,rold,\(l\):=𝔼\[Gs,rold∣Θr−1\(s\),W1:l\],μs,rnew,\(l\):=𝔼\[Gs,rnew∣Θr−1\(s\),W1:l\]\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}:=\\mathbb\{E\}\[G^\{\\mathrm\{old\}\}\_\{s,r\}\\mid\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\],\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}:=\\mathbb\{E\}\[G^\{\\mathrm\{new\}\}\_\{s,r\}\\mid\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\], so thatμs,r\(l\)=λs,old​μs,rold,\(l\)\+λs,new​μs,rnew,\(l\)\.\\mu^\{\(l\)\}\_\{s,r\}=\\lambda\_\{s,\\mathrm\{old\}\}\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\+\\lambda\_\{s,\\mathrm\{new\}\}\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\.Then under the sensitivity metricHs,r\(l\)H^\{\(l\)\}\_\{s,r\}in Definition[V\.6](https://arxiv.org/html/2608.11690#S5.Thmtheorem6),

μs,r\(l\)⊤​Hs,r\(l\)​μs,r\(l\)=λs,old2​‖μs,rold,\(l\)‖H2\+λs,new2​‖μs,rnew,\(l\)‖H2\+2​λs,old​λs,new​‖μs,rold,\(l\)‖H​‖μs,rnew,\(l\)‖H​cosH⁡\(μs,rold,\(l\),μs,rnew,\(l\)\),\\displaystyle\\mu^\{\(l\)\\top\}\_\{s,r\}H^\{\(l\)\}\_\{s,r\}\\mu^\{\(l\)\}\_\{s,r\}=\\lambda\_\{s,\\mathrm\{old\}\}^\{2\}\\big\\\|\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\big\\\|\_\{H\}^\{2\}\+\\lambda\_\{s,\\mathrm\{new\}\}^\{2\}\\big\\\|\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\big\\\|\_\{H\}^\{2\}\+2\\lambda\_\{s,\\mathrm\{old\}\}\\lambda\_\{s,\\mathrm\{new\}\}\\big\\\|\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\big\\\|\_\{H\}\\big\\\|\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\big\\\|\_\{H\}\\cos\_\{H\}\\\!\\Big\(\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\},\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\Big\),where the information\-geometric alignment

Align=cosH⁡\(μs,rold,\(l\),μs,rnew,\(l\)\)\.\\mathrm\{Align\}=\\cos\_\{H\}\\\!\\Big\(\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\},\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\Big\)\.

###### Proof\.

Sinceμs,r\(l\)=λs,old​μs,rold,\(l\)\+λs,new​μs,rnew,\(l\)\\mu\_\{s,r\}^\{\(l\)\}=\\lambda\_\{s,\\mathrm\{old\}\}\\mu\_\{s,r\}^\{\\mathrm\{old\},\(l\)\}\+\\lambda\_\{s,\\mathrm\{new\}\}\\mu\_\{s,r\}^\{\\mathrm\{new\},\(l\)\}, then

μs,r\(l\)⊤​Hs,r\(l\)​μs,r\(l\)\\displaystyle\\mu\_\{s,r\}^\{\(l\)\\top\}H\_\{s,r\}^\{\(l\)\}\\mu\_\{s,r\}^\{\(l\)\}=‖λs,old​μs,rold,\(l\)\+λs,new​μs,rnew,\(l\)‖H2\\displaystyle=\\big\\\|\\lambda\_\{s,\\mathrm\{old\}\}\\mu\_\{s,r\}^\{\\mathrm\{old\},\(l\)\}\+\\lambda\_\{s,\\mathrm\{new\}\}\\mu\_\{s,r\}^\{\\mathrm\{new\},\(l\)\}\\big\\\|^\{2\}\_\{H\}=λs,old2​‖μs,rold,\(l\)‖H2\+λs,new2​‖μs,rnew,\(l\)‖H2\+2​λs,old​λs,new​⟨μs,rold,\(l\),μs,rnew,\(l\)⟩H\\displaystyle=\\lambda\_\{s,\\mathrm\{old\}\}^\{2\}\\\|\\mu\_\{s,r\}^\{\\mathrm\{old\},\(l\)\}\\\|^\{2\}\_\{H\}\+\\lambda\_\{s,\\mathrm\{new\}\}^\{2\}\\\|\\mu\_\{s,r\}^\{\\mathrm\{new\},\(l\)\}\\\|^\{2\}\_\{H\}\+2\\lambda\_\{s,\\mathrm\{old\}\}\\lambda\_\{s,\\mathrm\{new\}\}\\langle\\mu\_\{s,r\}^\{\\mathrm\{old\},\(l\)\},\\mu\_\{s,r\}^\{\\mathrm\{new\},\(l\)\}\\rangle\_\{H\}=λs,old2​‖μs,rold,\(l\)‖H2\+λs,new2​‖μs,rnew,\(l\)‖H2\+2​λs,old​λs,new​‖μs,rold,\(l\)‖H​‖μs,rnew,\(l\)‖H​cosH⁡\(μs,rold,\(l\),μs,rnew,\(l\)\)\.\\displaystyle=\\lambda\_\{s,\\mathrm\{old\}\}^\{2\}\\\|\\mu\_\{s,r\}^\{\\mathrm\{old\},\(l\)\}\\\|^\{2\}\_\{H\}\+\\lambda\_\{s,\\mathrm\{new\}\}^\{2\}\\\|\\mu\_\{s,r\}^\{\\mathrm\{new\},\(l\)\}\\\|^\{2\}\_\{H\}\+2\\lambda\_\{s,\\mathrm\{old\}\}\\lambda\_\{s,\\mathrm\{new\}\}\\\|\\mu\_\{s,r\}^\{\\mathrm\{old\},\(l\)\}\\\|\_\{H\}\\,\\\|\\mu\_\{s,r\}^\{\\mathrm\{new\},\(l\)\}\\\|\_\{H\}\\cos\_\{H\}\\\!\\big\(\\mu\_\{s,r\}^\{\\mathrm\{old\},\(l\)\},\\mu\_\{s,r\}^\{\\mathrm\{new\},\(l\)\}\\big\)\.This completes the proof\. ∎

### B\-JProof of Corollary[V\.8](https://arxiv.org/html/2608.11690#S5.Thmtheorem8)

###### Corollary[V\.8](https://arxiv.org/html/2608.11690#S5.Thmtheorem8)\(Restate\)\.

Conditioning on\(Θr−1\(s\),W1:l\)\(\\Theta^\{\(s\)\}\_\{r\-1\},W\_\{1:l\}\), the per\-step trajectory budget admits the scalar upper bound

∑s=1t∑r=1Rs𝔼⁡\[12​log​det\(I\+αs,r​Ms,r\(l\)\)\]≤12​∑s=1t∑r=1Rsαs,r​𝔼​\[tr⁡\(Vs,r\(l\)\)\+‖μs,r\(l\)‖2\]\.\\displaystyle\\sum\_\{s=1\}^\{t\}\\sum\_\{r=1\}^\{R\_\{s\}\}\\mathbb\{E\}\\\!\\left\[\\frac\{1\}\{2\}\\log\\det\\\!\\Big\(I\+\\alpha\_\{s,r\}M^\{\(l\)\}\_\{s,r\}\\Big\)\\right\]\\leq\\frac\{1\}\{2\}\\sum\_\{s=1\}^\{t\}\\sum\_\{r=1\}^\{R\_\{s\}\}\\alpha\_\{s,r\}\\,\\mathbb\{E\}\\\!\\Big\[\\mathrm\{tr\}\\\!\\big\(V^\{\(l\)\}\_\{s,r\}\\big\)\+\\\|\\mu^\{\(l\)\}\_\{s,r\}\\\|^\{2\}\\Big\]\.Moreover, the two scalar components admit the explicit decompositionstr⁡\(Vs,r\(l\)\)=λs,old2​tr​\(Σs,rold,\(l\)\)\+λs,new2​tr​\(Σs,rnew,\(l\)\),\\mathrm\{tr\}\\\!\\big\(V^\{\(l\)\}\_\{s,r\}\\big\)=\\lambda\_\{s,\\mathrm\{old\}\}^\{2\}\\mathrm\{tr\}\\\!\\big\(\\Sigma^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\big\)\+\\lambda\_\{s,\\mathrm\{new\}\}^\{2\}\\mathrm\{tr\}\\\!\\big\(\\Sigma^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\big\),and‖μs,r\(l\)‖2=λs,old2​‖μs,rold,\(l\)‖2\+λs,new2​‖μs,rnew,\(l\)‖2\+2​λs,old​λs,new​⟨μs,rold,\(l\),μs,rnew,\(l\)⟩\\\|\\mu^\{\(l\)\}\_\{s,r\}\\\|^\{2\}=\\lambda\_\{s,\\mathrm\{old\}\}^\{2\}\\\|\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\\|^\{2\}\+\\lambda\_\{s,\\mathrm\{new\}\}^\{2\}\\\|\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\\|^\{2\}\+2\\lambda\_\{s,\\mathrm\{old\}\}\\lambda\_\{s,\\mathrm\{new\}\}\\langle\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\},\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\rangle\.

###### Proof\.

LetB⪰0B\\succeq 0have eigenvalues\{λi\}i=1d⊂ℝ\+\\\{\\lambda\_\{i\}\\\}\_\{i=1\}^\{d\}\\subset\\mathbb\{R\}\_\{\+\}\. Thenlogdet\(I\+B\)=∑i=1dlog\(1\+λi\)\\log\\det\(I\+B\)=\\sum\_\{i=1\}^\{d\}\\log\(1\+\\lambda\_\{i\}\)andtr⁡\(B\)=∑i=1dλi\\mathrm\{tr\}\(B\)=\\sum\_\{i=1\}^\{d\}\\lambda\_\{i\}\.

Sincelog⁡\(1\+x\)≤x\\log\(1\+x\)\\leq xfor allx≥0x\\geq 0, summing overiiyieldslogdet\(I\+B\)≤tr\(B\)\\log\\det\(I\+B\)\\leq\\mathrm\{tr\}\(B\)\.

Apply toB=αs,r​Ms,r\(l\)⪰0B=\\alpha\_\{s,r\}M^\{\(l\)\}\_\{s,r\}\\succeq 0:

12​log​det\(I\+αs,r​Ms,r\(l\)\)≤12​αs,r​tr​\(Ms,r\(l\)\)\.\\frac\{1\}\{2\}\\log\\det\(I\+\\alpha\_\{s,r\}M^\{\(l\)\}\_\{s,r\}\)\\leq\\frac\{1\}\{2\}\\,\\alpha\_\{s,r\}\\,\\mathrm\{tr\}\(M^\{\(l\)\}\_\{s,r\}\)\.Taking expectation and summing over\(s,r\)\(s,r\)gives12​∑s=1t∑r=1Rsαs,r​𝔼​\[tr⁡\(Ms,r\(l\)\)\]\\frac\{1\}\{2\}\\sum\_\{s=1\}^\{t\}\\sum\_\{r=1\}^\{R\_\{s\}\}\\alpha\_\{s,r\}\\,\\mathbb\{E\}\\\!\\big\[\\mathrm\{tr\}\(M^\{\(l\)\}\_\{s,r\}\)\\big\]\.

By the identity𝔼⁡\[X​X⊤\]=Cov⁡\(X\)\+𝔼⁡\[X\]​𝔼​\[X\]⊤\\mathbb\{E\}\[XX^\{\\top\}\]=\\mathrm\{Cov\}\(X\)\+\\mathbb\{E\}\[X\]\\mathbb\{E\}\[X\]^\{\\top\}, we haveMs,r\(l\)=Vs,r\(l\)\+μs,r\(l\)​μs,r\(l\)⊤M^\{\(l\)\}\_\{s,r\}=V^\{\(l\)\}\_\{s,r\}\+\\mu^\{\(l\)\}\_\{s,r\}\\mu^\{\(l\)\\top\}\_\{s,r\}, hencetr⁡\(Ms,r\(l\)\)=tr⁡\(Vs,r\(l\)\)\+tr⁡\(μs,r\(l\)​μs,r\(l\)⊤\)=tr⁡\(Vs,r\(l\)\)\+‖μs,r\(l\)‖2\\mathrm\{tr\}\(M^\{\(l\)\}\_\{s,r\}\)=\\mathrm\{tr\}\(V^\{\(l\)\}\_\{s,r\}\)\+\\mathrm\{tr\}\(\\mu^\{\(l\)\}\_\{s,r\}\\mu^\{\(l\)\\top\}\_\{s,r\}\)=\\mathrm\{tr\}\(V^\{\(l\)\}\_\{s,r\}\)\+\\\|\\mu^\{\(l\)\}\_\{s,r\}\\\|^\{2\}, which yields12​∑s=1t∑r=1Rsαs,r​𝔼​\[tr⁡\(Vs,r\(l\)\)\+‖μs,r\(l\)‖2\]\\frac\{1\}\{2\}\\sum\_\{s=1\}^\{t\}\\sum\_\{r=1\}^\{R\_\{s\}\}\\alpha\_\{s,r\}\\,\\mathbb\{E\}\\\!\\Big\[\\mathrm\{tr\}\\\!\\big\(V^\{\(l\)\}\_\{s,r\}\\big\)\+\\\|\\mu^\{\(l\)\}\_\{s,r\}\\\|^\{2\}\\Big\]\.

Finally, applying linearity of the trace to the definition ofVs,r\(l\)V^\{\(l\)\}\_\{s,r\}, we obtain

tr⁡\(Vs,r\(l\)\)=λs,old2​tr​\(Σs,rold,\(l\)\)\+λs,new2​tr​\(Σs,rnew,\(l\)\)\.\\mathrm\{tr\}\\\!\\big\(V^\{\(l\)\}\_\{s,r\}\\big\)=\\lambda\_\{s,\\mathrm\{old\}\}^\{2\}\\mathrm\{tr\}\\\!\\big\(\\Sigma^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\big\)\+\\lambda\_\{s,\\mathrm\{new\}\}^\{2\}\\mathrm\{tr\}\\\!\\big\(\\Sigma^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\big\)\.The second identity follows by expanding the squared Euclidean norm

‖μs,r\(l\)‖2=λs,old2​‖μs,rold,\(l\)‖2\+λs,new2​‖μs,rnew,\(l\)‖2\+2​λs,old​λs,new​⟨μs,rold,\(l\),μs,rnew,\(l\)⟩\.\\\|\\mu^\{\(l\)\}\_\{s,r\}\\\|^\{2\}=\\lambda\_\{s,\\mathrm\{old\}\}^\{2\}\\\|\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\\|^\{2\}\+\\lambda\_\{s,\\mathrm\{new\}\}^\{2\}\\\|\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\\|^\{2\}\+2\\lambda\_\{s,\\mathrm\{old\}\}\\lambda\_\{s,\\mathrm\{new\}\}\\langle\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\},\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\\rangle\.∎

### B\-KProof of Corollary[V\.9](https://arxiv.org/html/2608.11690#S5.Thmtheorem9)

###### Corollary[V\.9](https://arxiv.org/html/2608.11690#S5.Thmtheorem9)\(Restate\)\.

Fix any sensitivity metricHs,r\(l\)H^\{\(l\)\}\_\{s,r\}\. Forλ∈\[0,1\]\\lambda\\in\[0,1\], consider the mixed vectorμs,r\(l\)​\(λ\):=λ​μs,rold,\(l\)\+\(1−λ\)​μs,rnew,\(l\)\\mu\_\{s,r\}^\{\(l\)\}\(\\lambda\):=\\lambda\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\+\(1\-\\lambda\)\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}and its metric energy‖μs,r\(l\)​\(λ\)‖H2\\\|\\mu\_\{s,r\}^\{\(l\)\}\(\\lambda\)\\\|\_\{H\}^\{2\}\. Then‖μs,r\(l\)​\(λ\)‖H2\\\|\\mu^\{\(l\)\}\_\{s,r\}\(\\lambda\)\\\|\_\{H\}^\{2\}is minimized overλ∈\[0,1\]\\lambda\\in\[0,1\]by

λ⋆=clip\[0,1\]​\(μs,rnew,\(l\)⊤​H​\(μs,rnew,\(l\)−μs,rold,\(l\)\)\(μs,rold,\(l\)−μs,rnew,\(l\)\)⊤​H​\(μs,rold,\(l\)−μs,rnew,\(l\)\)\),\\lambda^\{\\star\}=\\mathrm\{clip\}\_\{\[0,1\]\}\\\!\\left\(\\frac\{\{\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\}^\{\\top\}H\(\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\-\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\)\}\{\(\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\-\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\)^\{\\top\}H\(\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\-\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\)\}\\right\),with the conventionλ⋆=0\\lambda^\{\\star\}=0ifμs,rold,\(l\)=μs,rnew,\(l\)\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}=\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\.

###### Proof\.

Letd:=μs,rold,\(l\)−μs,rnew,\(l\)d:=\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\-\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\. Thenμ⁡\(λ\)=μs,rnew,\(l\)\+λ​d\\mu\(\\lambda\)=\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\+\\lambda dand

‖μ⁡\(λ\)‖H2=\(μs,rnew,\(l\)\+λ​d\)⊤​H​\(μs,rnew,\(l\)\+λ​d\)=μs,rnew,\(l\)⊤​H​μs,rnew,\(l\)\+2​λ​μs,rnew,\(l\)⊤​H​d\+λ2​d⊤​H​d\.\\\|\\mu\(\\lambda\)\\\|\_\{H\}^\{2\}=\(\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\+\\lambda d\)^\{\\top\}H\(\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\+\\lambda d\)=\{\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\}^\{\\top\}H\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\+2\\lambda\\,\{\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\}^\{\\top\}Hd\+\\lambda^\{2\}\\,d^\{\\top\}Hd\.Ifμs,rold,\(l\)=μs,rnew,\(l\)\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}=\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}, thend=0d=0and‖μ⁡\(λ\)‖H2≡μs,rnew,\(l\)⊤​H​μs,rnew,\(l\)\\\|\\mu\(\\lambda\)\\\|\_\{H\}^\{2\}\\equiv\{\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\}^\{\\top\}H\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}is constant; we defineλ⋆=0\\lambda^\{\\star\}=0\. Ifμs,rold,\(l\)≠μs,rnew,\(l\)\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\\neq\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}, thend≠0d\\neq 0and sinceH≻0H\\succ 0, we haved⊤​H​d\>0d^\{\\top\}Hd\>0\. Therefore the scalar functionφ⁡\(λ\):=‖μ⁡\(λ\)‖H2\\varphi\(\\lambda\):=\\\|\\mu\(\\lambda\)\\\|\_\{H\}^\{2\}is a strictly convex quadratic\. Differentiate:

φ′​\(λ\)=2​μs,rnew,\(l\)⊤​H​d\+2​λ​d⊤​H​d\.\\varphi^\{\\prime\}\(\\lambda\)=2\\,\{\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\}^\{\\top\}Hd\+2\\lambda\\,d^\{\\top\}Hd\.Settingφ′​\(λ\)=0\\varphi^\{\\prime\}\(\\lambda\)=0yields the unique unconstrained minimizer

λunc⋆=−μs,rnew,\(l\)⊤​H​dd⊤​H​d=μs,rnew,\(l\)⊤​H​\(μs,rnew,\(l\)−μs,rold,\(l\)\)\(μs,rold,\(l\)−μs,rnew,\(l\)\)⊤​H​\(μs,rold,\(l\)−μs,rnew,\(l\)\)\.\\lambda^\{\\star\}\_\{\\mathrm\{unc\}\}=\-\\frac\{\{\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\}^\{\\top\}Hd\}\{d^\{\\top\}Hd\}=\\frac\{\{\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\}^\{\\top\}H\(\{\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\}\-\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\)\}\{\(\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\-\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\)^\{\\top\}H\(\\mu^\{\\mathrm\{old\},\(l\)\}\_\{s,r\}\-\\mu^\{\\mathrm\{new\},\(l\)\}\_\{s,r\}\)\}\.The constrained minimizer overλ∈\[0,1\]\\lambda\\in\[0,1\]is the Euclidean projection ofλunc⋆\\lambda^\{\\star\}\_\{\\mathrm\{unc\}\}onto\[0,1\]\[0,1\], i\.e\.,λ⋆=clip\[0,1\]​\(λunc⋆\)\\lambda^\{\\star\}=\\mathrm\{clip\}\_\{\[0,1\]\}\(\\lambda^\{\\star\}\_\{\\mathrm\{unc\}\}\)\. The proof is complete\. ∎

## Appendix CDetails of the Synthetic Gaussian Experiments

We construct a sequence ofTTbinary classification tasks\{𝒯t\}t=0T−1\\\{\\mathcal\{T\}\_\{t\}\\\}\_\{t=0\}^\{T\-1\}inℝd\\mathbb\{R\}^\{d\}\. Each task is a balanced two\-class Gaussian mixture:

x\|\(y=0,𝒯t\)\\displaystyle x\\mid\(y=0,\\mathcal\{T\}\_\{t\}\)∼𝒩⁡\(μt−s​wt,Id\),\\displaystyle\\sim\\mathcal\{N\}\(\\mu\_\{t\}\-sw\_\{t\},I\_\{d\}\),x\|\(y=1,𝒯t\)\\displaystyle x\\mid\(y=1,\\mathcal\{T\}\_\{t\}\)∼𝒩⁡\(μt\+s​wt,Id\),\\displaystyle\\sim\\mathcal\{N\}\(\\mu\_\{t\}\+sw\_\{t\},I\_\{d\}\),wherewtw\_\{t\}is a unit direction controlling the class\-discriminative axis,μt\\mu\_\{t\}is a task\-specific mean shift, ands\>0s\>0is the class separation\. To control task relatedness, we restrict\{wt\}\\\{w\_\{t\}\\\}to a random 2D subspace and vary its evolution across tasks\. We consider two modes: \(i\) positive synergy, wherewtw\_\{t\}rotates smoothly by a fixed angle incrementΔ\\Deltaso adjacent tasks remain aligned; \(ii\) negative synergy, wherewtw\_\{t\}alternates sign \(wt\+1=−wtw\_\{t\+1\}=\-w\_\{t\}\), which is equivalent to a label\-flip across tasks and induces strong interference\. We control cross\-task drift viaμt=C⋅ξt\\mu\_\{t\}=C\\cdot\\xi\_\{t\}withξt∼𝒩⁡\(0,Id\)\\xi\_\{t\}\\sim\\mathcal\{N\}\(0,I\_\{d\}\)\. For each task𝒯t\\mathcal\{T\}\_\{t\}, we sample an i\.i\.d\. training setDtD\_\{t\}of sizennand an i\.i\.d\. test setDttestD\_\{t\}^\{\\mathrm\{test\}\}of sizentestn\_\{\\mathrm\{test\}\}\.

After completing tasktt, we store exactlymmexamples from its training set into a per\-task bufferℬt\\mathcal\{B\}\_\{t\}\. The total replay memory thus scales asO⁡\(T​m\)O\(Tm\)\. When training on taskt≥1t\\geq 1, each update step samples a mini\-batchBnewB^\{\\mathrm\{new\}\}from the current task and a replay mini\-batchBoldB^\{\\mathrm\{old\}\}from replay buffer∪i<tℬi\\cup\_\{i<t\}\\mathcal\{B\}\_\{i\}\. We minimize the mixture loss

ℓt​\(θ\)=λnew​\(t\)​ℒ^​\(Bnew,θ\)\+λold​\(t\)​ℒ^​\(Bold,θ\),λnew​\(t\)=1t\+1,λold​\(t\)=tt\+1\.\\ell\_\{t\}\(\\theta\)=\\lambda\_\{\\mathrm\{new\}\}\(t\)\\,\\widehat\{\\mathcal\{L\}\}\(B^\{\\mathrm\{new\}\};\\theta\)\+\\lambda\_\{\\mathrm\{old\}\}\(t\)\\,\\widehat\{\\mathcal\{L\}\}\(B^\{\\mathrm\{old\}\};\\theta\),\\qquad\\lambda\_\{\\mathrm\{new\}\}\(t\)=\\frac\{1\}\{t\+1\},\\;\\;\\lambda\_\{\\mathrm\{old\}\}\(t\)=\\frac\{t\}\{t\+1\}\.We use SGD with fixed learning rate and batch size\. After finishing allTTtasks, we report the test risk

L⁡\(θ\)=1T​∑t=0T−1ℒ^​\(Dttest,θ\),L\(\\theta\)=\\frac\{1\}\{T\}\\sum\_\{t=0\}^\{T\-1\}\\widehat\{\\mathcal\{L\}\}\(D\_\{t\}^\{\\mathrm\{test\}\};\\theta\),and the empirical risk \(old tasks via replay buffers, the last task via its training set\)

L^​\(θ\)=T−1T​ℒ^​\(⋃t=0T−2ℬt,θ\)\+1T​ℒ^​\(DT−1,θ\)\.\\widehat\{L\}\(\\theta\)=\\frac\{T\-1\}\{T\}\\widehat\{\\mathcal\{L\}\}\\Big\(\\bigcup\_\{t=0\}^\{T\-2\}\\mathcal\{B\}\_\{t\};\\theta\\Big\)\+\\frac\{1\}\{T\}\\widehat\{\\mathcal\{L\}\}\(D\_\{T\-1\};\\theta\)\.The generalization gap isgap=L​\(θ\)−L^​\(θ\)\\mathrm\{gap\}=L\(\\theta\)\-\\widehat\{L\}\(\\theta\)\. We report this task\-averaged gap; it equals the total\-risk gap of Theorem[IV\.1](https://arxiv.org/html/2608.11690#S4.Thmtheorem1)divided by the task countTT\. BecauseTTis fixed in this experiment, the constant1/T1/Trescales the gap but leaves the predictedm−1/2m^\{\-1/2\}scaling and the location of the finite\-memory plateau unchanged\. Unless stated otherwise, all curves are averaged overKKindependent runs; we report95%95\\%confidence intervals for the mean \(Student\-tt\)\.

Similar Articles

Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss

Hugging Face Daily Papers

This paper derives a unified token- and layer-wise decomposition of how learning from one token changes another prediction, connecting data attribution, catastrophic forgetting, and plasticity loss as distinct regimes of the same evolving update-behavior interaction. It yields forward-computable approximations enabling data selection, interference controls, and diagnostics for future learnability across models and training regimes.

Rethinking Transfer in Continual Learning: A Replay-Based Realisation

arXiv cs.LG

This paper introduces a framework for when transfer should be expected in continual learning and proposes Transfer-Selective Replay (TSR), which selects replay data predicted to benefit the incoming task rather than indiscriminately replaying past examples. TSR improves forward transfer while maintaining stability, outperforming existing replay baselines.

When Does Continual Learning Require Learning

arXiv cs.LG

This paper proposes a unified framework for continual learning in LLMs, disentangling change along space (new domains) and time (data drift). It evaluates various methods including prompting, supervised learning, reinforcement learning, and context compression under realistic sequential settings.