Evidence Before Expansion: Reuse, Spawn, or Defer in Lifelong Expert Pools

arXiv cs.LG Papers

Summary

This paper introduces a statistically defined decision layer for continual-learning systems, allowing expert pools to decide whether to reuse existing models, spawn new ones, or defer based on accumulated evidence, with theoretical guarantees and system contributions for managing nonstationary data streams.

arXiv:2608.19888v1 Announce Type: new Abstract: Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn a new one, or defer. We present a decision layer that makes all three outcomes statistically meaningful. Reuse and spawn are posed as one-sided sequential hypotheses on a conditional (mechanism-level) discrepancy, separated by an indifference zone; defer is exactly the state in which neither betting e-process has accumulated sufficient evidence. We prove finite-time anytime validity for the observable surrogate discrepancy of a predictable discriminator sequence, and an unconditional one-sided transfer to the population quantity in which each side's slack is the excess risk of a single discriminator; an empirically observed downward-bias regularity makes the spawn side exactly conservative. Recency without sacrificing the guarantee is obtained by a restarted e-detector: a bank of unwindowed betting supermartingales at geometrically spaced restart times (O(log t) memory), with the error budget spent over restart instances, which preserves lifetime anytime validity; spending over expert-creation order likewise controls multiplicity for unboundedly many experts. On synthetic multi-concept streams, Electricity, Covertype, and the recurrence-heavy INSECTS benchmark, the instance-accounted restarted bank achieves zero false spawns and zero false reuses after switches and matches or exceeds the retired windowed heuristic (INSECTS-reoccurring accuracy 0.675), making the deployed algorithm and the guaranteed algorithm one and the same.
Original Article
View Cached Full Text

Cached at: 08/21/26, 10:29 AM

# Evidence Before Expansion:Reuse, Spawn, or Defer in Lifelong Expert Pools
Source: [https://arxiv.org/html/2608.19888](https://arxiv.org/html/2608.19888)
Kentaro OdaAffiliation:Center for Management of Information Technologies, Kagoshima UniversityEmail:[odaken@cc\.kagoshima\-u\.ac\.jp](mailto:)

###### Abstract

Continual\-learning systems have treated uncertainty about a new batch as a nuisance to be resolved immediately; this paper makes “do not decide yet” a statistically defined action\.*Defer*is not a heuristic: it is exactly the region between accumulated evidence for reuse and accumulated evidence for spawn\. Systems that maintain a pool of expert models over a nonstationary stream must repeatedly decide whether an incoming batch should be absorbed by an existing expert, spawn a new one, or wait for more evidence\. We present a complete decision layer built on a two\-axis task comparison \(the conditional Jensen–Shannon discrepancy and its covariate companion\) with three system contributions\. \(1\)*Decision semantics*: reuse/spawn tests are posed as one\-sided sequential hypotheses separated by an indifference zone\[τ,3​τ\]\[\\tau,3\\tau\];*defer*is the state in which neither betting e\-process has accumulated enough evidence, giving the abstention a precise statistical meaning\. \(2\)*Sequential evidence*: per\-expert betting e\-processes on per\-point loss\-difference increments scored by*predictable*\(frozen\-before\-use\) discriminators gate the decisions; we prove finite\-time anytime validity for the*observable*surrogate discrepancy of the predictable discriminator sequence, and an*unconditional*one\-sided transfer to the population quantity \(each side’s slack is the excess risk of a single discriminator; a stated downward\-bias regularity, observed throughout, makes the spawn side exactly conservative\); the deployed evidence process is a*restarted e\-detector*: a bank of unwindowed supermartingales with geometrically spaced restarts and the level spent over restart*instances*, giving bounded\-memory recency*inside*a lifetime anytime\-validity guarantee \(a single unwindowed process mis\-reuses at rate0\.50\.5after concept switches; the restarted bank at0\.000\.00, with the best accuracy of any evidence variant on recurrence\-heavy streams\)\. \(3\)*Systems mechanics*: expert shortlisting by recent loss bounds per\-chunk cost; mini\-batch test\-then\-train routing removes the switch lag that otherwise dominates accuracy differences; merge closes the loop for recurring concepts\. On a four\-regime synthetic stream the batch gate attains zero false spawns and zero missed concepts with the ideal expert count, and the default streaming configuration \(restarted e\-detector \+ spending\) holds false\-spawn0\.000\.00/ false\-reuse0\.000\.00; on INSECTS \(documented drifts\) it exploits recurrence to hold1313experts where exchange\-based decisions hold1818–5252; on covertype it correctly maintains11–22experts\. We characterize the regimes where expert pools pay off \(discrete, recurring concepts\) and where they cannot \(continuous drift\), and release all code\.

## 1Introduction

Adaptive systems answer nonstationarity with one of three primitive actions: adapt an existing model, create a new one, or wait\. Existing criteria collapse this decision into a scalar trigger—input novelty \(AGE/SEMA\-style\), loss jumps \(DDM\-style\), or model\-exchange regret \(CLS\-style\)—each of which confounds at least two of the three underlying situations \(covariate shift, mechanism change, insufficient evidence\)\. This paper treats the decision layer itself as the object of design and evaluation\.

We build on the conditional Jensen–Shannon discrepancy \(CJSD\), which decomposes task discrepancy exactly into a covariate axisIxI\_\{x\}and a functional axisDCJSD\_\{\\mathrm\{CJS\}\}, both estimated from two discriminators; a companion paper develops its theory\. Here we contribute the*system*: sequential decision semantics, streaming validity, and the mechanics that make the layer run at stream rate, together with a benchmark across four stream regimes and seven decision policies\.

## 2The decision layer

DEFERaccumulate evidenceREUSEabsorb into expertkkSPAWNnew expert,αc\\alpha\_\{c\}budgetINCOMPARABLEIx\>I\_\{x\}\>ceiling: no reuse claimEreuse\(k\)≥thrE^\{\(k\)\}\_\{\\mathrm\{reuse\}\}\\geq\\mathrm\{thr\}\(evidenceDCJS<3​τD\_\{\\mathrm\{CJS\}\}<3\\tau\)allEspawn\(k\)≥thrE^\{\(k\)\}\_\{\\mathrm\{spawn\}\}\\geq\\mathrm\{thr\}\(evidenceDCJS\>τD\_\{\\mathrm\{CJS\}\}\>\\tau\)comparability gatereset evidencenew monitorFigure 1:The decision layer as a state machine\. Every chunk starts indefer; the two betting e\-processes must*earn*a transition \(thresholdKmax/αK\_\{\\max\}/\\alphaor theαc\\alpha\_\{c\}spending schedule\), the indifference zone\[τ,3​τ\]\[\\tau,3\\tau\]separates the two exits, and the comparability gate blocks reuse claims when supports barely overlap\.### Setting\.

Chunks\(Xt,yt\)\(X\_\{t\},y\_\{t\}\)arrive; a pool of experts\{Ek\}\\\{E\_\{k\}\\\}each hold a training reservoir, a held\-out reservoir, and a model\. Prequential accuracy is measured before learning\. Figure[1](https://arxiv.org/html/2608.19888#S2.F1)summarizes the layer\.

### Two\-axis gate\.

For a candidate chunk and expertkk, estimate\(Ix\(k\),DCJS\(k\)\)\(I\_\{x\}^\{\(k\)\},D\_\{\\mathrm\{CJS\}\}^\{\(k\)\}\)with confidence intervals\. IfIx\(k\)I\_\{x\}^\{\(k\)\}exceeds a comparability bound the pair is vacuous \(the functional axis is honestly zero off\-overlap\) and cannot justify reuse\. The batch and sequential gates test the*same*two hypotheses,H0sp:DCJS≤τH\_\{0\}^\{\\mathrm\{sp\}\}:D\_\{\\mathrm\{CJS\}\}\\leq\\tau\(against spawn\) andH0re:DCJS≥3​τH\_\{0\}^\{\\mathrm\{re\}\}:D\_\{\\mathrm\{CJS\}\}\\geq 3\\tau\(against reuse\), separated by the indifference zone\[τ,3​τ\]\[\\tau,3\\tau\]; the batch gate is their single\-look version: spawn when every comparable expert hasL⁡\(DCJS\)\>τL\(D\_\{\\mathrm\{CJS\}\}\)\>\\tau, reuse when some expert hasU⁡\(DCJS\)<3​τU\(D\_\{\\mathrm\{CJS\}\}\)<3\\tau—our implementation uses the stricter cutU⁡\(DCJS\)≤τU\(D\_\{\\mathrm\{CJS\}\}\)\\leq\\tau, which rejectsH0reH\_\{0\}^\{\\mathrm\{re\}\}*a fortiori*and only makes reuse more conservative—and defer otherwise\.

### Indifference zone and e\-processes\.

Sequentially, reuse and spawn are one\-sided tests ofH0sp:DCJS≤τH\_\{0\}^\{\\mathrm\{sp\}\}:D\_\{\\mathrm\{CJS\}\}\\leq\\tauandH0re:DCJS≥3​τH\_\{0\}^\{\\mathrm\{re\}\}:D\_\{\\mathrm\{CJS\}\}\\geq 3\\tau\. Per expert we maintain two betting e\-processes on the per\-point incrementsuiu\_\{i\}, scored by discriminators frozen before the chunk arrives \(a predictable scoring rule; the increments themselves are the new randomness\):EspawnE\_\{\\mathrm\{spawn\}\}bets upward,EreuseE\_\{\\mathrm\{reuse\}\}downward; an action fires at thresholdKmax/αK\_\{\\max\}/\\alpha, whereKmaxK\_\{\\max\}is a*declared design capacity*on the number of simultaneously monitored experts, enforced by the merge/prune layer \(we useKmax=16K\_\{\\max\}=16; observed pools stay below it\)\. The union bound must be overKmaxK\_\{\\max\}, not the data\-dependent pool size: a threshold that grows with each spawn does not control the family level when experts are created indefinitely\. Two caveats delimit whatKmax/αK\_\{\\max\}/\\alphabuys: it bounds*simultaneous*monitors, so if experts are pruned and replaced without limit, the lifetime family of tested hypotheses can exceedKmaxK\_\{\\max\}\. For unbounded lifetimes we implement an*α\\alpha\-spending*variant: thecc\-th created expert receivesαc=6​α/\(π2​c2\)\\alpha\_\{c\}=6\\alpha/\(\\pi^\{2\}c^\{2\}\)\(so∑cαc=α\\sum\_\{c\}\\alpha\_\{c\}=\\alpha\), split across its two one\-sided processes, giving valid familywise control over*arbitrarily many*creation events with no capacity cap\. Re\-running the full stream benchmark under spending, lifetime validity turns out to cost nothing measurable: decision quality is unchanged \(false\-spawn0\.020\.02vs0\.010\.01, false\-reuse0\.000\.00\), pools stay comparable \(6\.3→7\.06\.3\\to 7\.0experts on INSECTS\-reoccurring\), and prequential accuracy is equal or slightly higher on every stream \(0\.86→0\.880\.86\\to 0\.88synthetic,0\.77→0\.790\.77\\to 0\.79Covertype\)—early experts face*lower*thresholds thanKmax/αK\_\{\\max\}/\\alphaand the quadratically growing late thresholds never bind at the pool sizes these streams induce\. We therefore recommend spending as the default whenever expert lifetimes are unbounded; combined with the restarted e\-detector of the next paragraph, the entire deployed configuration—recency, multiplicity, and unbounded lifetimes—now sits inside the validity guarantee\.

### Sensitivity \(ablation of the retired windowed variant\)\.

On the synthetic stream aW×τW\\times\\taugrid \(W∈\{4,8,16,32\}W\\in\\\{4,8,16,32\\\},τ∈\{0\.01,0\.03,0\.05,0\.10\}\\tau\\in\\\{0\.01,0\.03,0\.05,0\.10\\\}, 3 seeds\) localizes the sensitivity entirely inτ\\tau: atτ=0\.01\\tau=0\.01the system holds false\-spawn\+\+false\-reuse at0\.010\.01with1\.31\.3experts, whileτ≥0\.03\\tau\\geq 0\.03widens the indifference zone\[τ,3​τ\]\[\\tau,3\\tau\]past the concept gap and mis\-reuses the new concept in half the runs \(combined error0\.500\.50, single\-expert collapse\)\. The window length is*inert*across the entire\[4,32\]\[4,32\]range \(identical numbers to three decimals\): decisive evidence accumulates within≤4\{\\leq\}4chunks here, soWWbinds only through the post\-switch recency mechanism of Sec\. 3\. Practical guidance: setτ\\taubelow the smallest drift mass worth reacting to;WWis not a tuning burden\. The zone\[τ,3​τ\]\[\\tau,3\\tau\]makes the reuse\-side test well posed \(without it the reuse boundary is statistically unreachable\)\.*Defer is exactly the state where neither process has crossed\.*

### What the e\-process actually tests\.

The increments are computed from*learned*discriminators, so the guarantee must be stated for the observable score, not assumed for the population quantity\. The pair scoring chunktt,\(T1,t,T2,t\)\(T\_\{1,t\},T\_\{2,t\}\), is*predictable*: updated only between chunks and frozen before the chunk arrives, so an e\-process that survives several chunks is scored by a predictable, possibly time\-varying sequence of pairs \(incremental discriminators are covered; within each chunk the pair is fixed\)\. Define thesurrogate discrepancyof the pair scoring chunktt,

D~t=𝔼⁡\[ℓ1​\(T1,t,X,Z\)−ℓ2​\(T2,t,X,Y,Z\)\|𝒢t−1\]=DCJS\+ε1,t−ε2,t,\\widetilde\{D\}\_\{t\}\\;=\\;\\mathbb\{E\}\\\!\\left\[\\ell\_\{1\}\(T\_\{1,t\};X,Z\)\-\\ell\_\{2\}\(T\_\{2,t\};X,Y,Z\)\\,\\middle\|\\,\\mathcal\{G\}\_\{t\-1\}\\right\]\\;=\\;D\_\{\\mathrm\{CJS\}\}\+\\varepsilon\_\{1,t\}\-\\varepsilon\_\{2,t\},the*conditional*population log\-loss gap achieved by the pair scoring chunktt, where𝒢t−1\\mathcal\{G\}\_\{t\-1\}is theσ\\sigma\-field of everything observed before chunktt\(which fixes\(T1,t,T2,t\)\(T\_\{1,t\},T\_\{2,t\}\)\) andεj,t≥0\\varepsilon\_\{j,t\}\\geq 0are the pair’s \(conditional\) excess risks\. Letℱi−1\\mathcal\{F\}\_\{i\-1\}be theσ\\sigma\-field generated by everything observed before pointii\(including the pair scoringiiand the bets\)\. The discriminators andλi\\lambda\_\{i\}areℱi−1\\mathcal\{F\}\_\{i\-1\}\-measurable \(predictable\); the incrementui=\(\(ℓ1,i−ℓ2,i\)\+B\)/2​B∈\[0,1\]u\_\{i\}=\(\(\\ell\_\{1,i\}\-\\ell\_\{2,i\}\)\+B\)/2B\\in\[0,1\]is the newly observed quantity, and what the null delivers is the conditional\-mean inequality𝔼⁡\[ui∣ℱi−1\]≤m0\\mathbb\{E\}\[u\_\{i\}\\mid\\mathcal\{F\}\_\{i\-1\}\]\\leq m\_\{0\}\(spawn side;≥m0\\geq m\_\{0\}for reuse\)\. Throughout,τ\\taudenotes the*normalized*threshold, so the surrogate null readsD~t/ln⁡2≤τ\\widetilde\{D\}\_\{t\}/\\ln 2\\leq\\tauandm0sp=\(τ​ln⁡2\+B\)/2​Bm\_\{0\}^\{\\mathrm\{sp\}\}=\(\\tau\\ln 2\+B\)/2Bis dimensionally consistent \(the reuse side usesm0re=\(3​τ​ln⁡2\+B\)/2​Bm\_\{0\}^\{\\mathrm\{re\}\}=\(3\\tau\\ln 2\+B\)/2B\)\.

###### Proposition 1\(Finite\-time validity for the observable score\)\.

Under clipping \(bounded losses\), for any betting strategy withλi\\lambda\_\{i\}ℱi−1\\mathcal\{F\}\_\{i\-1\}\-measurable,λi≥0\\lambda\_\{i\}\\geq 0, and the capital constraint1\+λi​σ​\(ui−m0\)≥01\+\\lambda\_\{i\}\\,\\sigma\(u\_\{i\}\-m\_\{0\}\)\\geq 0\(a negativeλi\\lambda\_\{i\}would reverse the defining inequality\), the processEn=∏i≤n\(1\+λi​σ​\(ui−m0\)\)E\_\{n\}=\\prod\_\{i\\leq n\}\\bigl\(1\+\\lambda\_\{i\}\\,\\sigma\(u\_\{i\}\-m\_\{0\}\)\\bigr\)is a nonnegative supermartingale with respect to\(ℱi\)\(\\mathcal\{F\}\_\{i\}\)whenever the surrogate null holds*pointwise over the process’ lifetime*—for every chunkttwhose increments enter the product—namelyD~t/ln⁡2≤τ\\widetilde\{D\}\_\{t\}/\\ln 2\\leq\\taufor the spawn side withm0sp=\(τ​ln⁡2\+B\)/2​Bm\_\{0\}^\{\\mathrm\{sp\}\}=\(\\tau\\ln 2\+B\)/2B,σ=\+1\\sigma=\+1, andD~t/ln⁡2≥3​τ\\widetilde\{D\}\_\{t\}/\\ln 2\\geq 3\\taufor the reuse side withm0re=\(3​τ​ln⁡2\+B\)/2​Bm\_\{0\}^\{\\mathrm\{re\}\}=\(3\\tau\\ln 2\+B\)/2B,σ=−1\\sigma=\-1\(a composite null over the predictable pair sequence\), and Ville’s inequality givesPr\[supnEn≥Kmax/α\]≤α/Kmax\\Pr\[\\sup\_\{n\}E\_\{n\}\\geq K\_\{\\max\}/\\alpha\]\\leq\\alpha/K\_\{\\max\}at any stopping time\. The guarantee is unconditional and finite\-time—but its null isD~t\\widetilde\{D\}\_\{t\}, notDCJSD\_\{\\mathrm\{CJS\}\}\.

###### Proposition 2\(One\-sided transfer to the population quantity\)\.

Unconditionally—for any frozen pair, however misspecified—the excess risks bound one direction each:DCJS−ε2,t≤D~t≤DCJS\+ε1,tD\_\{\\mathrm\{CJS\}\}\-\\varepsilon\_\{2,t\}\\leq\\widetilde\{D\}\_\{t\}\\leq D\_\{\\mathrm\{CJS\}\}\+\\varepsilon\_\{1,t\}\(companion paper, one\-sided misspecification control\)\. Hence a spawn\-side rejection ofD~t/ln⁡2≤τ\\widetilde\{D\}\_\{t\}/\\ln 2\\leq\\taucertifiesDCJS/ln⁡2\>τ−ε1,t/ln⁡2D\_\{\\mathrm\{CJS\}\}/\\ln 2\>\\tau\-\\varepsilon\_\{1,t\}/\\ln 2with*no*condition onT2T\_\{2\}, and a reuse\-side rejection certifiesDCJS/ln⁡2<3​τ\+ε2,t/ln⁡2D\_\{\\mathrm\{CJS\}\}/\\ln 2<3\\tau\+\\varepsilon\_\{2,t\}/\\ln 2with*no*condition onT1T\_\{1\}\. If moreover the pair satisfies the downward\-bias regularityε1,t≤ε2,t\\varepsilon\_\{1,t\}\\leq\\varepsilon\_\{2,t\}\(i\.e\.D~t≤DCJS\\widetilde\{D\}\_\{t\}\\leq D\_\{\\mathrm\{CJS\}\}\), the spawn slack vanishes—a rejection certifiesDCJS/ln⁡2\>τD\_\{\\mathrm\{CJS\}\}/\\ln 2\>\\tauoutright—and the reuse slack sharpens to\(ε2,t−ε1,t\)/ln⁡2\(\\varepsilon\_\{2,t\}\-\\varepsilon\_\{1,t\}\)/\\ln 2\.

The unconditional part shifts what must be controlled: slack\-aware population transfer on the spawn side involves onlyε1,t\\varepsilon\_\{1,t\}—the excess risk of the*simple*xx\-discriminator, the quantity that held\-out model selection already minimizes and that admits standard approximation\-plus\-complexity bounds \(companion paper, Prop\. 4\)—not a sign comparison between the two discriminators\. Exact level\-τ\\taupopulation conservativeness follows either by inflating the surrogate spawn threshold by a valid high\-probability upper bound onε1,t/ln⁡2\\varepsilon\_\{1,t\}/\\ln 2\(on that bound’s1−δ1\-\\deltaevent; the sequential levelα\\alphaand the bound’sδ\\deltacompose additively\), or, as the zero\-slack special case, under the downward\-bias regularity\. That regularity is an empirical refinement: it held in every lifecycle benchmark reported in this paper \(the companion paper exhibits an engineered misspecified\-marginal exception, within the proven slack\), and it sharpens the slacks but no longer carries the validity claim\. The reuse\-side slackε2,t\\varepsilon\_\{2,t\}is further limited in practice by the comparability gate, which excludes low\-overlap comparisons, one important regime in whichε2,t\\varepsilon\_\{2,t\}grows\. In one sentence: surrogate\-level sequential validity is exact; population\-CJSD decisions inherit one\-sided, discriminator\-specific slacks unconditionally, and exact zero\-slack conservativeness is the special case obtained either by threshold correction with a valid excess\-risk bound or under the empirically observed downward\-bias regularity\. The deployed recency mechanism \(next paragraph\) sits*inside*these guarantees\.

### Recency without windows: a restarted e\-detector\.

An e\-process accumulated over a long stationary stretch can absorb a concept switch: the stale product outvotes fresh contradicting evidence and mis\-fires reuse \(measured mis\-reuse rate0\.50\.5\)\. A sliding window of the lastWWper\-chunk factors with a freshness guard restores correct behavior \(mis\-reuse0\.10\.1\)—but truncating the product breaks the supermartingale property, so the windowed heuristic sits outside the validity theorem\. We resolve this with a*restarted e\-detector*: each monitor keeps a bank of*unwindowed*betting processes with geometrically spaced restart times \(slotjjholds the surviving process of age≈2j\\approx 2^\{j\}chunks;O⁡\(log⁡t\)O\(\\log t\)memory\)\. The error budget must be spent over restart*instances*, not over active slots: unboundedly many distinct processes successively occupy the sameSSslots over an unbounded stream, so a slot\-only union bound would not control the lifetime error\. Therr\-th restart instance created over the monitor’s lifetime therefore receives budgetαr=αside⋅6/\(π2​r2\)\\alpha\_\{r\}=\\alpha\_\{\\mathrm\{side\}\}\\cdot 6/\(\\pi^\{2\}r^\{2\}\)and alarms only above its own threshold1/αr1/\\alpha\_\{r\}; discarded instances can no longer alarm and consume no memory\.

###### Proposition 3\(Lifetime validity of the restarted bank\)\.

Each restart instance is a nonnegative supermartingale under the surrogate null from its own restart time \(Proposition 1 applies verbatim from that time\), and∑r≥1αr≤αside\\sum\_\{r\\geq 1\}\\alpha\_\{r\}\\leq\\alpha\_\{\\mathrm\{side\}\}, so the union bound over*all instances ever created*preserves the family level at every time; the construction is anytime\-valid withO⁡\(log⁡t\)O\(\\log t\)memory\. Recency is structural: after a switch the youngest instances contain no pre\-switch evidence, and the geometric grid keeps some restart within a factor22of the switch point\. The price of the instance accounting is logarithmic: instancerr’s log\-threshold exceeds the uniform one by2​ln⁡r\+O⁡\(1\)2\\ln r\+O\(1\), anO⁡\(log⁡r\)O\(\\log r\)additional evidence requirement recovered inO⁡\(log⁡r/g\)O\(\\log r/g\)chunks under any alternative with log\-evidence growth rateg\>0g\>0\.

Empirically the correct accounting costs almost nothing: re\-running the full benchmark with the instance\-accounted bank \(spending multiplicity\) gives post\-switch mis\-reuse0\.000\.00and false\-spawn0\.000\.00on the synthetic stream, prequential accuracy within one point of the windowed heuristic there \(0\.8460\.846vs0\.8560\.856\) and equal or better everywhere else—including the*best accuracy of any evidence variant*on the recurrence\-heavy stream \(0\.616→0\.6750\.616\\to 0\.675on INSECTS\-reoccurring\) and on Covertype \(0\.770→0\.7900\.770\\to 0\.790\)\. The windowed heuristic is therefore retired from the default configuration: the deployed system and the guarantee now coincide\.

### Mechanics\.

Shortlisting: only the top\-kkexperts by error on the*previous*chunk are compared, so candidate selection is predictable and does not touch the labels that subsequently enter the evidence \(5\.9×5\.9\\timesspeedup on 33\-dimensional INSECTS with no measured decision change\)\. Routing: mini\-batch test\-then\-train \(labels used only for routing\) cuts the switch lag from one chunk to one mini\-batch and lifts every adaptive policy to the same accuracy ceiling, isolating decision quality as the differentiator\. Merge: expert pairs with high overlap and mutually lowDCJSD\_\{\\mathrm\{CJS\}\}are merged; reservoir recency caps prevent regime pollution, which we identify as the true cause of over\-spawning on fast\-mixing streams\.

## 3Benchmark

![Refer to caption](https://arxiv.org/html/2608.19888v1/figs/fig16_stream.png)Figure 2:Improved streaming benchmark: all seven policies share online routing and pruning; a gradual\-drift phase \(shaded\) and recurrences are included\. Bottom: expert counts\.Policies\.single, spawn\-always, input\-novelty \(AGE/SEMA\-style\), loss\-jump \(DDM\-style\), exchange score \(CLS\-style\), CPD\-family gate, CJSD gate \(batch and e\-process variants\)\.Streams\.A four\-regime synthetic stream \(abrupt switches, a covariate\-only phase, gradual drift, recurrences; ground\-truth mapping ids\); Electricity; Covertype; INSECTS abrupt and incremental\-reoccurring \(documented change points\)\.

Findings\.\(1\) With routing equalized, accuracy differences between adaptive policies nearly vanish \(synthetic:0\.9390\.939–0\.9450\.945\); the true differentiators are decision quality and expert economy, where the CJSD gate is the only policy with zero false spawns and zero missed concepts at the ideal expert count\. \(2\) The gradual phase separates policies sharply: spawn rates during gradual drift are0\.100\.10\(CJSD\) vs0\.160\.16–0\.220\.22\(loss/CPD\) vs1\.01\.0\(spawn\-always\)\. \(3\) On INSECTS the pair\-level anatomy shows segments0≈2≈50\\approx 2\\approx 5recur; the CJSD gate exploits this \(13 experts vs 18–52\) and its non\-spawn at recurrent change points is correct reuse, not a miss \(Fig\.[3](https://arxiv.org/html/2608.19888#S3.F3)\)\. \(4\) On continuous\-drift Electricity no expert pool helps \(all0\.740\.74–0\.780\.78\): a boundary of applicability, diagnosed by the same machinery \(reservoir\-vs\-chunkDCJSD\_\{\\mathrm\{CJS\}\}stays permanently high\)\. \(5\) The streaming\-native variant in its default configuration \(incremental discriminators \+ restarted e\-detector \+α\\alpha\-spending\) is the most conservative policy in the pool: false\-spawn0\.000\.00and false\-reuse0\.000\.00on the synthetic stream at1\.01\.0experts, with accuracy0\.850\.85there \(batch gate:0\.910\.91; the gap is the explicit price of family\-level error control\) and the best accuracy of all evidence variants on the recurrence\-heavy INSECTS stream \(0\.6750\.675\)\.

![Refer to caption](https://arxiv.org/html/2608.19888v1/figs/fig17b_streams.png)Figure 3:INSECTS \(abrupt and reoccurring\): prequential accuracy \(top\) and expert counts \(bottom\) per decision policy; dotted lines mark documented change points\.
## 4Sequential validity in isolation

![Refer to caption](https://arxiv.org/html/2608.19888v1/figs/figE5_cs.png)Figure 4:Repeated CI peeking commits 64% of borderline cases to a near\-coin\-flip decision; anytime\-valid monitoring keeps them deferred and pays only∼2×\{\\sim\}2\\timesdelay on clear cases\.On a controlled drift boundary \(DCJS≈τD\_\{\\mathrm\{CJS\}\}\\approx\\tau\), naive repeated confidence intervals commit 64% of runs to a decision that is effectively a coin flip; valid schemes defer 90–97% of them, while on clear cases they pay a delay factor of only1\.81\.8–2\.62\.6\(Fig\.[4](https://arxiv.org/html/2608.19888#S4.F4)\)\. The e\-process variant adds a false\-alarm rate of0\.000\.00at detection delays20%20\\%above the \(invalid\) naive monitor\.

## 5Related work

Expert/adapter expansion by input novelty \(SEMA\), loss\-based drift response and model merging in federated streams \(FedDrift\), continual learning of mixed task sequences \(CAT\), and drift detectors \(ADWIN, DDM\) each implement a one\-axis trigger; exchange\-based scores \(CLS\) confound covariate shift with mechanism change\. Our layer differs in \(i\) the two\-axis gate, \(ii\) the indifference\-zone sequential semantics of defer, and \(iii\) validity under continuous monitoring\.

## 6Limitations

Continuous\-drift streams remain out of scope for any discrete\-concept pool; discriminator cost, though bounded by shortlisting, exceeds loss\-trigger baselines by∼3×\{\\sim\}3\\times\. Bounded\-memory recency is now provided*inside*the validity guarantee by the restarted e\-detector with restart\-instance spending \(Proposition 3\); what remains open is sharper\-than\-union\-bound multiplicity \(mixture e\-values\) and the finite\-sample theory of the underlying estimator \(companion paper\)\.

## References

- \[1\]K\. Oda\. Separating covariate shift from mechanism change with two discriminators\. Preprint, 2026 \(companion paper, posted concurrently\)\.
- \[2\]S\. Sun, H\. H\. Zhang, J\. C\. Watkins\. Quantifying data similarity using cross learning\. arXiv:2510\.10866, 2025\.
- \[3\]W\. Wang et al\. Self\-expansion of pre\-trained models with mixture of adapters for continual learning\. CVPR 2025\.
- \[4\]E\. Jothimurugesan et al\. Federated learning under distributed concept drift\. AISTATS 2023\.
- \[5\]Z\. Ke, B\. Liu, X\. Huang\. Continual learning of a mixed sequence of similar and dissimilar tasks\. NeurIPS 2020\.
- \[6\]V\. Souza et al\. Challenges in benchmarking stream learning algorithms with real\-world data\. DMKD 2020\.

Similar Articles

Rethinking Experience Utilization in Self-Evolving Language Model Agents

arXiv cs.CL

This paper introduces ExpWeaver, a framework that optimizes how self-evolving language model agents utilize past experiences during runtime decision-making. It demonstrates that selectively invoking experience based on reasoning uncertainty improves performance across various environments and models.