DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink Accounting
摘要
Presents DG-FedReuse, a federated learning mechanism that reuses age-decayed cached client updates under a proxy-gradient threshold to reduce uplink communication, achieving significant modeled savings with minimal accuracy loss on image classification benchmarks.
查看缓存全文
缓存时间: 2026/08/07 07:49
# DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink Accounting
Source: [https://arxiv.org/html/2608.05358](https://arxiv.org/html/2608.05358)
India rahilaftab12@gmail\.comVineet Kumar Rakesh[https://orcid.org/0009-0000-7102-6564](https://orcid.org/0009-0000-7102-6564) Engineering Science Homi Bhabha National Institute AnushaktinagarMumbai 400094MaharashtraIndia Computer and Informatics Group Variable Energy Cyclotron Centre 1/AFBidhannagarKolkata 700064West BengalIndia vineet@vecc\.gov\.inSoumya Mazumdar[https://orcid.org/0009-0006-3521-9557](https://orcid.org/0009-0006-3521-9557) Department of Computer Science and Business Systems Gargi Memorial Institute of Technology Affiliated to Maulana Abul Kalam Azad University of Technology BalarampurMouza BeraliaBaruipurKolkata 700144West BengalIndia reachme@soumyamazumdar\.comTapas Samanta[https://orcid.org/0000-0003-0521-0747](https://orcid.org/0000-0003-0521-0747) Engineering Science Homi Bhabha National Institute AnushaktinagarMumbai 400094MaharashtraIndia Computer and Informatics Group Variable Energy Cyclotron Centre 1/AFBidhannagarKolkata 700064West BengalIndia tsamanta@vecc\.gov\.in
###### Abstract
Federated learning repeatedly incurs local optimization and model\-update transmission\. We study DG\-FedReuse, a simulator\-level mechanism that allows selected clients to contribute age\-decayed cached updates when a stochastic head\-gradient discrepancy proxy remains below a round\-dependent threshold\. A hard cache\-age limit and minimum fresh\-client quota constrain reuse, while fresh updates use an adaptive per\-tensor Top\-K numerical\-field representation\. Experiments cover six image\-classification datasets, 50 virtual clients, Dirichlet label heterogeneity \(α=0\.5\\alpha=0\.5\), and three seeds\. At a common 90\-round budget, DG\-FedReuse yields 83\.36–85\.42% modeled update\-data\-field uplink saving, compared with 76\.88% for matched Top\-K FedAvg; the seed\-aligned accuracy differences range from−5\.29\-5\.29to−0\.14\-0\.14percentage points\. Best\-observed test accuracies obtained under test\-controlled checkpointing are retained only as exploratory archival evidence and range from−2\.38\-2\.38to\+0\.45\+0\.45percentage points relative to matched FedAvg\. A symmetric dense\-model\-downlink sensitivity reduces the headline saving to 41\.68–42\.71% and the incremental gain over Top\-K FedAvg to 3\.24–4\.27 percentage points, demonstrating the dependence of communication conclusions on the accounting boundary\. The study characterizes the proposed reuse rule in the implemented simulator; it does not establish unbiased generalization, end\-to\-end bandwidth reduction, runtime or energy savings, faster convergence, or superiority over existing stale\-update and lazy\-aggregation methods\.
Keywords:Federated learning; communication\-efficient learning; non\-IID data; update reuse; lazy aggregation; Top\-*K*sparsification\.
## 1Introduction
Federated learning \(FL\) trains a shared model from decentralized data by alternating local optimization and server aggregation\[[16](https://arxiv.org/html/2608.05358#bib.bib1),[8](https://arxiv.org/html/2608.05358#bib.bib2)\]\. Although raw training examples remain local, each selected client ordinarily downloads a global model, performs local optimization, and uploads an update\. Repeating this process can make communication and client computation important constraints, especially under heterogeneous data and partial participation\.
Communication\-efficient FL includes fewer communication rounds, quantization, coordinate sparsification, error\-feedback mechanisms, and adaptive server optimization\[[11](https://arxiv.org/html/2608.05358#bib.bib7),[18](https://arxiv.org/html/2608.05358#bib.bib8),[21](https://arxiv.org/html/2608.05358#bib.bib9),[17](https://arxiv.org/html/2608.05358#bib.bib5),[14](https://arxiv.org/html/2608.05358#bib.bib25)\]\. Most such methods reduce or transform a fresh update after local optimization\. A complementary design choice is whether every selected client must recompute a fresh update in every round\. Several prior methods already exploit memorized, stale, lazily transmitted, or recycled information\[[5](https://arxiv.org/html/2608.05358#bib.bib22),[19](https://arxiv.org/html/2608.05358#bib.bib23),[20](https://arxiv.org/html/2608.05358#bib.bib24),[10](https://arxiv.org/html/2608.05358#bib.bib14)\]\. The relevant novelty question is therefore not whether update reuse exists, but whether the particular gate, safeguards, sparse\-update accounting, and empirical behavior studied here add a distinct and useful mechanism\.
DG\-FedReuse uses a stochastic proxy\-gradient discrepancy to choose between a fresh local\-training path and a cached\-update path\. A round\-dependent threshold, a hard cache\-age limit, and a minimum fresh\-client quota constrain reuse\. Fresh updates refresh client\-specific state and pass through an adaptive Top\-K numerical\-field representation; reused updates are attenuated according to cache age\. The proxy is not treated as a direct measurement of concept drift because it also reflects mini\-batch sampling, data augmentation, and model\-state effects\.
The paper asks the following bounded question:
> *At a fixed communication\-round budget, when the same sparse\-update rule and coordinate\-retention ratio are applied to FedAvg\-family controls, how much additional modeled update\-data\-field uplink saving is associated with the implemented cached\-update reuse rule, and what accuracy differences accompany it?*
The contribution is an implementation and evidence study rather than a convergence or networking claim\. Specifically, the paper provides:
- •an implementation\-faithful specification of client\-indexed cached\-update reuse with a stochastic proxy\-gradient gate, age forcing, a minimum fresh\-client quota, staleness decay, and sample\-weighted aggregation;
- •deterministic numerical\-field accounting for an adaptive dense/value–index/bitmap Top\-K representation, with explicit separation of update uplink from excluded model downlink and protocol traffic;
- •a six\-dataset, three\-seed matched sparse\-control study whose primary summary uses a common 90\-round budget, plus clearly labelled archival test\-controlled results and mechanism\-oriented secondary analyses; and
- •an explicit audit of claim boundaries, including missing closest\-method baselines, test\-controlled selection, stochastic gate reliability, error feedback, complete communication measurement, and convergence analysis\.
## 2Related work and positioning
### 2\.1Federated optimization and partial participation
FedAvg established the local\-SGD and sample\-weighted aggregation template\[[16](https://arxiv.org/html/2608.05358#bib.bib1)\]\. FedProx adds a proximal term to mitigate client drift under heterogeneous data and systems\[[13](https://arxiv.org/html/2608.05358#bib.bib3)\]; SCAFFOLD instead uses client and server control variates\[[9](https://arxiv.org/html/2608.05358#bib.bib4)\]\. FedOpt applies adaptive server optimization, including FedAdam\-like moment updates\[[17](https://arxiv.org/html/2608.05358#bib.bib5)\]\. These methods primarily modify optimization or aggregation rather than deciding whether a selected client should execute a fresh local\-training path\.
Partial participation motivates server memory\. FedVARP stores client\-indexed information to reduce variance from partial participation\[[7](https://arxiv.org/html/2608.05358#bib.bib11)\]\. MIFA memorizes the latest updates from unavailable devices and uses them to correct participation\-related bias\[[5](https://arxiv.org/html/2608.05358#bib.bib22)\]\. FedStale combines fresh and stale updates through a tunable interpolation and analyzes how participation and data heterogeneity affect stale\-update utility\[[20](https://arxiv.org/html/2608.05358#bib.bib24)\]\. These methods are especially close because they establish that client\-indexed historical updates can be algorithmically useful, although their participation models and update rules differ from DG\-FedReuse\.
### 2\.2Compression, error feedback, and bidirectional accounting
Quantization and sparsification reduce the representation size of fresh updates\[[11](https://arxiv.org/html/2608.05358#bib.bib7),[18](https://arxiv.org/html/2608.05358#bib.bib8),[21](https://arxiv.org/html/2608.05358#bib.bib9),[22](https://arxiv.org/html/2608.05358#bib.bib10)\]\. Direct Top\-K is a biased compressor\. Error feedback can compensate for discarded information, and its behavior under partial participation requires specific analysis\[[14](https://arxiv.org/html/2608.05358#bib.bib25)\]\. The present implementation does not use error feedback, so matched Top\-K controls isolate the implemented codec but do not represent the strongest known sparse baseline\.
Most reported communication metrics in this study concern update uplink\. Downlink can be a distinct systems bottleneck, and methods such as DoCoFL explicitly target model broadcast compression\[[4](https://arxiv.org/html/2608.05358#bib.bib26)\]\. Accordingly, we report a simple dense\-downlink sensitivity in addition to the original uplink\-only accounting, but do not substitute that sensitivity for measured bidirectional traffic\.
### 2\.3Lazy aggregation and update recycling
Lazy or recycled\-gradient methods form the closest conceptual family\. LAQ suppresses quantized gradient transmission based on innovation\[[23](https://arxiv.org/html/2608.05358#bib.bib12)\]\. LBGM exploits a low\-rank gradient subspace and recycles update information through compact coefficients\[[1](https://arxiv.org/html/2608.05358#bib.bib13)\]\. The 3PC framework provides a general compressor class and a theory connecting lazy aggregation and error feedback\[[19](https://arxiv.org/html/2608.05358#bib.bib23)\]\. GradSkip studies conditional local computation and communication in distributed optimization\[[15](https://arxiv.org/html/2608.05358#bib.bib15)\]\. FedLUAR recycles selected layer updates at the server\[[10](https://arxiv.org/html/2608.05358#bib.bib14)\]\. These results rule out a broad first\-use or first\-recycling claim\.
Table[1](https://arxiv.org/html/2608.05358#S2.T1)positions mechanisms rather than reporting an empirical ranking\. The closest methods were not rerun under the present client partitions, models, stopping rules, and accounting boundary; their absence from the experimental baseline set is a central limitation\.
Table 1:Mechanism\-level positioning\. “Not evaluated” means that no head\-to\-head performance claim is made\.
### 2\.4Role of the evaluated controls
FedAvg and FedProx are mechanism controls because they preserve the FedAvg\-family aggregation structure and use the same Top\-K rule\. FedProx uses the same proximal\-objective form as the fresh path of DG\-FedReuse, but not the same coefficient\. The separately tuned FedAdam suite checks a different server optimizer\. These controls help interpret the implemented mechanism, but they are not substitutes for MIFA, FedStale, 3PC\-derived lazy aggregation, FedLUAR, or Top\-K with error feedback\. Consequently, the paper does not claim superiority over the stale\-update, lazy\-aggregation, or compression literature\.
## 3Problem formulation and claim boundary
LetNNclients hold local objectivesFi\(w\)F\_\{i\}\(w\)with nonnegative sample weightspip\_\{i\}satisfying∑ipi=1\\sum\_\{i\}p\_\{i\}=1\. The global objective is
minwF\(w\),F\(w\)=∑i=1NpiFi\(w\)\.\\min\_\{w\}F\(w\),\\qquad F\(w\)=\\sum\_\{i=1\}^\{N\}p\_\{i\}F\_\{i\}\(w\)\.\(1\)At roundtt, the server samplesmmclients without replacement, forming𝒮t\\mathcal\{S\}\_\{t\}\. A conventional selected client starts fromwtw\_\{t\}, performsEElocal epochs, and returnsΔi,t=wi,t\(E\)−wt\\Delta\_\{i,t\}=w\_\{i,t\}^\{\(E\)\}\-w\_\{t\}\. DG\-FedReuse partitions𝒮t\\mathcal\{S\}\_\{t\}into an active set𝒜t\\mathcal\{A\}\_\{t\}, which computes fresh updates, and a reuse setℛt=𝒮t∖𝒜t\\mathcal\{R\}\_\{t\}=\\mathcal\{S\}\_\{t\}\\setminus\\mathcal\{A\}\_\{t\}, which contributes client\-indexed cached updates\.
The evaluated implementation is a single\-process simulator with explicit client objects and server state, not a distributed network deployment\. The server invokes proxy routines and reads cached state in process\. Figure[1](https://arxiv.org/html/2608.05358#S3.F1)gives one logical mapping consistent with the quota rule: a client retains its signature and returns a scalar discrepancy score, while the server retains the cached model update, ranks scores when quota promotion is required, and applies age and reuse rules\. A design that transmits proxy vectors or stores update caches at clients would have different communication and storage costs\.
The primary communication outcome is therefore a numerical\-field ratio for client\-to\-server model updates, not measured bandwidth\. The metric excludes model broadcast, tensor identifiers, representation\-mode identifiers, framing, acknowledgements, retransmissions, and most gate/control information\. A separate sensitivity adds one dense model download per selected client per round, but still does not represent a serialized protocol\.
Figure 1:Logical interpretation of the single\-process mechanism\. A client computes a stochastic proxy\-gradient discrepancy and returns a rankable scalar score; the server retains cached model updates and applies age, quota, and decay rules\. Only numerical fields for fresh model updates and a conventional 16\-byte reused\-event charge enter the primary counter\. Model downlink, representation metadata, transport framing, and most control traffic are excluded\.
## 4DG\-FedReuse
### 4\.1Stochastic proxy\-gradient discrepancy
For a selected client with a valid cache, the implementation loads the round\-start global statewtw\_\{t\}and accumulates gradients for parameters whose names containclassifierorfcover up to five shuffled proxy mini\-batches\. The model remains in training mode; CIFAR clients can therefore include random augmentation and batch\-normalization state behavior\. Letpi,tp\_\{i,t\}denote the resulting flattened proxy and lethih\_\{i\}denote the cached signature\. The gate uses clipped cosine distance
Di,t=1−clip\(⟨pi,t,hi⟩‖pi,t‖2‖hi‖2\+10−12,−1,1\)\.D\_\{i,t\}=1\-\\operatorname\{clip\}\\\!\\left\(\\frac\{\\left\\langle p\_\{i,t\},h\_\{i\}\\right\\rangle\}\{\\left\\lVert p\_\{i,t\}\\right\\rVert\_\{2\}\\left\\lVert h\_\{i\}\\right\\rVert\_\{2\}\+10^\{\-12\}\},\-1,1\\right\)\.\(2\)A cacheless client receivesDi,t=1D\_\{i,t\}=1and is activated\. BecauseDi,tD\_\{i,t\}depends on the global state, sampled proxy batches, augmentation, and model\-mode effects, we call it a stochastic proxy\-gradient discrepancy rather than a direct measurement of distributional or concept drift\. Its repeatability was not independently evaluated\.
### 4\.2Threshold, age, and forced freshness
The round\-dependent threshold is
τt=max\(τmin,τ0e−γt\)\.\\tau\_\{t\}=\\max\\\!\\left\(\\tau\_\{\\min\},\\tau\_\{0\}e^\{\-\\gamma t\}\\right\)\.\(3\)Because the active condition isDi,t≥τtD\_\{i,t\}\\geq\\tau\_\{t\}, decreasingτt\\tau\_\{t\}makes fresh training easier to trigger\. Ifai,t=t−ria\_\{i,t\}=t\-r\_\{i\}is cache age andSSis maximum staleness, the preliminary active set contains every selected client satisfying
cacheless∨Di,t≥τt∨ai,t≥S\.\\text\{cacheless\}\\quad\\lor\\quad D\_\{i,t\}\\geq\\tau\_\{t\}\\quad\\lor\\quad a\_\{i,t\}\\geq S\.\(4\)If fewer than⌈qm⌉\\lceil qm\\rceilclients satisfy Eq\. \([4](https://arxiv.org/html/2608.05358#S4.E4)\), the server promotes non\-active selected clients in decreasingDi,tD\_\{i,t\}order until the quota is met\. The server therefore requires a rankable scalar score, not only a binary eligibility flag\.
### 4\.3Fresh and reused paths
An active client performsEElocal epochs of SGD\. The DG\-FedReuse fresh path uses the proximal\-objective form
Fi\(w\)\+μ2‖w−wt‖22\.F\_\{i\}\(w\)\+\\frac\{\\mu\}\{2\}\\left\\lVert w\-w\_\{t\}\\right\\rVert\_\{2\}^\{2\}\.\(5\)The trained\-model delta is sparsified, represented as dense\-shaped sparse tensors in the current cache implementation, and stored in client\-indexed server state\. Separately, the implementation draws another proxy mini\-batch sequence and recomputes a head\-gradient signature at the same round\-start global statewtw\_\{t\}\. Denoting this stochastic refresh proxy bypi,trefp\_\{i,t\}^\{\\mathrm\{ref\}\}, the signature cache is updated by
hi←0\.9hi\+0\.1pi,tref,h\_\{i\}\\leftarrow 0\.9h\_\{i\}\+0\.1p\_\{i,t\}^\{\\mathrm\{ref\}\},\(6\)with direct initialization when no signature exists\. Thus, a cacheful active event computes both a gating proxy and a separately sampled refresh proxy\.
A non\-active selected client contributes
Δ~i,t=ρai,tΔ^i,\\widetilde\{\\Delta\}\_\{i,t\}=\\rho^\{a\_\{i,t\}\}\\widehat\{\\Delta\}\_\{i\},\(7\)whereΔ^i\\widehat\{\\Delta\}\_\{i\}is the cached sparse update and0<ρ<10<\\rho<1\. Both fresh and reused contributions retain the selected client’s sample weight\. Withnin\_\{i\}local examples, the round update is
Δ¯t=∑i∈𝒮tni∑j∈𝒮tnj\{Δ^i,t,i∈𝒜t,ρai,tΔ^i,i∈ℛt\.\\overline\{\\Delta\}\_\{t\}=\\sum\_\{i\\in\\mathcal\{S\}\_\{t\}\}\\frac\{n\_\{i\}\}\{\\sum\_\{j\\in\\mathcal\{S\}\_\{t\}\}n\_\{j\}\}\\begin\{cases\}\\widehat\{\\Delta\}\_\{i,t\},&i\\in\\mathcal\{A\}\_\{t\},\\\\ \\rho^\{a\_\{i,t\}\}\\widehat\{\\Delta\}\_\{i\},&i\\in\\mathcal\{R\}\_\{t\}\.\\end\{cases\}\(8\)The primary FedAvg\-family server applieswt\+1=wt\+ηsΔ¯tw\_\{t\+1\}=w\_\{t\}\+\\eta\_\{s\}\\overline\{\\Delta\}\_\{t\}withηs=1\\eta\_\{s\}=1\.
### 4\.4Adaptive Top\-K numerical\-field model
For each tensor withddcoordinates and element widthbbbytes, magnitude Top\-K retains
k=min\{d,max\(1,⌊rd⌋\)\}k=\\min\\\{d,\\max\(1,\\lfloor rd\\rfloor\)\\\}\(9\)coordinates at ratiorr\. The accounting code selects the smallest numerical\-field size among
Bdense\\displaystyle B\_\{\\mathrm\{dense\}\}=db,\\displaystyle=db,\(10\)Bpairs\\displaystyle B\_\{\\mathrm\{pairs\}\}=k\(b\+4\),\\displaystyle=k\(b\+4\),\(11\)Bbitmap\\displaystyle B\_\{\\mathrm\{bitmap\}\}=kb\+⌈d8⌉\.\\displaystyle=kb\+\\left\\lceil\\frac\{d\}\{8\}\\right\\rceil\.\(12\)The 4\-byte term represents an int32 coordinate index\. These expressions count tensor values, indices, and bitmap bits\. They do not include tensor identity, dimensions, dtype, selected\-mode identifiers, object serialization, alignment, checksums, or transport framing\. The current compressor also discards unselected coordinates without error feedback\[[14](https://arxiv.org/html/2608.05358#bib.bib25)\]\.
### 4\.5Communication accounting
LetBi,tB\_\{i,t\}be the sum of Eq\. \([12](https://arxiv.org/html/2608.05358#S4.E12)\)’s selected numerical fields over tensors for a fresh client, and letB0B\_\{0\}be the dense state\-update size\. The simulator charges 16 bytes for each reused\-client event\. This is a fixed convention, not a measured wire representation\. It records
Mt\\displaystyle M\_\{t\}=∑i∈𝒜tBi,t\+16\|ℛt\|,\\displaystyle=\\sum\_\{i\\in\\mathcal\{A\}\_\{t\}\}B\_\{i,t\}\+16\|\\mathcal\{R\}\_\{t\}\|,\(13\)Mtdense\\displaystyle M\_\{t\}^\{\\mathrm\{dense\}\}=mB0,\\displaystyle=mB\_\{0\},\(14\)st\\displaystyle s\_\{t\}=1−MtMtdense\.\\displaystyle=1\-\\frac\{M\_\{t\}\}\{M\_\{t\}^\{\\mathrm\{dense\}\}\}\.\(15\)Run summaries averagests\_\{t\}, whereas the fixed\-round table uses cumulative bytes through round 90\. The phrase “modeled update\-data\-field uplink saving” always refers to this restricted boundary\.
For a transparent sensitivity, suppose each selected client also receives one dense model of sizeB0B\_\{0\}in every round and compare against a dense bidirectional reference of2mB02mB\_\{0\}\. Then
stsym=1−mB0\+Mt2mB0=st2\.s\_\{t\}^\{\\mathrm\{sym\}\}=1\-\\frac\{mB\_\{0\}\+M\_\{t\}\}\{2mB\_\{0\}\}=\\frac\{s\_\{t\}\}\{2\}\.\(16\)Equation \([16](https://arxiv.org/html/2608.05358#S4.E16)\) is not a measured system result; it only illustrates how an uncompressed downlink changes the ratio\.
Table 2:Logical events and treatment in the reported accounting\. No network serialization was executed\.
### 4\.6Operational invariants
The implementation supplies safeguards, not a convergence theorem\. After quota promotion,\|𝒜t\|≥⌈qm⌉\|\\mathcal\{A\}\_\{t\}\|\\geq\\lceil qm\\rceil\. A selected client can be reused only whenai,t<Sa\_\{i,t\}<S; Eq\. \([4](https://arxiv.org/html/2608.05358#S4.E4)\) forces a refresh at or beyond the age limit\. Because0<ρ<10<\\rho<1,
‖Δ~i,t‖2=ρai,t‖Δ^i‖2≤‖Δ^i‖2\.\\left\\lVert\\widetilde\{\\Delta\}\_\{i,t\}\\right\\rVert\_\{2\}=\\rho^\{a\_\{i,t\}\}\\left\\lVert\\widehat\{\\Delta\}\_\{i\}\\right\\rVert\_\{2\}\\leq\\left\\lVert\\widehat\{\\Delta\}\_\{i\}\\right\\rVert\_\{2\}\.\(17\)This magnitude bound does not guarantee alignment with the current descent direction\.
Algorithmic roundttof DG\-FedReuse1\.Samplemmclients and providewtw\_\{t\}\.2\.For each cacheful selected client, compute the stochastic proxy\-gradient discrepancyDi,tD\_\{i,t\}; activate cacheless, threshold\-triggered, or age\-forced clients\.3\.Promote the largest remaining discrepancy scores until at least⌈qm⌉\\lceil qm\\rceilclients are active\.4\.Each active client performs proximal local SGD, applies Top\-K, contributes the fresh sparse update, and refreshes update and signature caches\.5\.Each reusable client contributes its cached update scaled byρai,t\\rho^\{a\_\{i,t\}\}\.6\.Aggregate all selected\-client contributions with Eq\. \([8](https://arxiv.org/html/2608.05358#S4.E8)\), record the restricted byte counter, and execute the configured evaluation and stopping logic\.
*The procedure omits unexecuted networking, secure aggregation, privacy, and adversarial\-robustness layers\.*
## 5Experimental methodology
### 5\.1Datasets, partitioning, and models
We use six image\-classification tasks \(Table[3](https://arxiv.org/html/2608.05358#S5.T3)\)\. The internal keyfemnistloads EMNIST Balanced; it is not a writer\-partitioned federated character dataset\. Dirichlet allocation withα=0\.5\\alpha=0\.5partitions training labels over 50 virtual clients\. MNIST, FashionMNIST, EMNIST Balanced, and PathMNIST use a shallow CNN; CIFAR\-10 and CIFAR\-100 use torchvision ResNet\-18\[[6](https://arxiv.org/html/2608.05358#bib.bib16)\]\. CIFAR training uses random cropping and horizontal flipping\.
PathMNIST provides train, validation, and test partitions, but the historical final configurations setevaluation\_split: test\. The other final tasks likewise use their canonical test sets for periodic evaluation\. Consequently, final\-test observations were used for early stopping and best\-checkpoint selection; this design creates selection bias and prevents the reported maxima from being interpreted as unbiased held\-out estimates\[[2](https://arxiv.org/html/2608.05358#bib.bib27)\]\.
Table 3:Executed dataset and model scope\. Counts are the train/test examples used by the loaders\.
### 5\.2Executed protocol
Each primary run uses 50 clients, 10 selected per round, three local epochs, batch size 128, SGD learning rate 0\.05, momentum 0\.9, weight decay10−410^\{\-4\}, and server learning rate 1\.0\. Proxy batch size is 64\. The maximum is 1,000 rounds; evaluation occurs at round 1 and every five rounds thereafter\. Early stopping has patience eight evaluation events\. Mixed precision is enabled on CUDA, while Top\-K preparation uses a CPU path\. Seeds are 101, 202, and 303\.
The code seeds Python, NumPy, and PyTorch, but enables cuDNN benchmarking and does not enforce deterministic algorithms\. Moreover, different methods consume random numbers through different execution paths\. Thus, equal seed labels do not guarantee bitwise determinism or identical client\-selection and augmentation sequences across methods; reported within\-seed differences are described as seed\-aligned rather than fully paired experimental blocks\.
For DG\-FedReuse, the frozen gate isτ0=0\.9116988\\tau\_\{0\}=0\.9116988,τmin=0\.8134931\\tau\_\{\\min\}=0\.8134931,γ=0\.01\\gamma=0\.01,S=4S=4,ρ=0\.7279635\\rho=0\.7279635,q=0\.30q=0\.30, andμ=3\.4156×10−4\\mu=3\.4156\\times 10^\{\-4\}\. All primary controls user=0\.2r=0\.2\. FedProx usesμ=5×10−4\\mu=5\\times 10^\{\-4\}; it therefore shares the proximal form, not the coefficient, with DG\-FedReuse\.
### 5\.3Hyperparameter provenance
Table[4](https://arxiv.org/html/2608.05358#S5.T4)distinguishes historical searches that used a test\-labelled stream from later validation\-only selection\. Disjoint final seeds do not repair repeated use of the same canonical test examples for model selection\[[2](https://arxiv.org/html/2608.05358#bib.bib27)\]\. The displayed final maxima and trajectories are therefore exploratory archival evidence\. The common\-round analysis reduces one source of bias—maximization over evaluation time—but remains post hoc and test\-observed\.
Table 4:Recorded hyperparameter\-selection stages\. “Dev\-test” denotes use of a canonical test split during development\.The selected values were applied globally rather than tuned per dataset\. Fixed design choices includedγ=0\.01\\gamma=0\.01,q=0\.30q=0\.30, the head\-parameter naming rule, signature EMA coefficient, five\-batch proxy cap, reuse\-event charge, and codec field model\. They were not independently ablated\.
### 5\.4Top\-K selection
The initial Top\-K grid coveredr=1\.0,0\.95,…,0\.40r=1\.0,0\.95,\\ldots,0\.40\(52 runs\)\. A documented post\-hoc extension evaluatedr∈\{1\.0,0\.35,0\.30,0\.25,0\.20,0\.15,0\.10,0\.05\}r\\in\\\{1\.0,0\.35,0\.30,0\.25,0\.20,0\.15,0\.10,0\.05\\\}\(32 runs\)\. All candidates used 150 fixed rounds and no early stopping\. The rule selected the largest mean modeled update\-data\-field uplink saving subject to at least 98% validation\-accuracy retention relative to reuse\-onlyr=1\.0r=1\.0on both CIFAR datasets\.
Table 5:Decision boundary from the two\-stage Top\-K validation study\. Values belowr=0\.4r=0\.4belong to the post\-hoc extension\.Figure 2:Validation\-accuracy retention versus restricted update\-uplink saving in the two\-stage Top\-K study\. The selectedr=0\.2r=0\.2candidate has 98\.0003% worst\-dataset retention and therefore essentially no margin above the prespecified threshold\.
### 5\.5Outcomes and statistical reporting
The primary descriptive outcome uses accuracy at round 90, the earliest terminal round among all 54 FedAvg/FedProx/DG\-FedReuse runs, and cumulative restricted bytes through the same round\. This fixes the communication\-round budget and avoids maximizing accuracy over evaluation times\. The historical maximum test accuracy before stopping is reported separately as archival evidence\.
Tables report arithmetic mean±\\pmsample standard deviation over three seeds\. Withn=3n=3, no null\-hypothesis significance test, confidence claim, or equivalence claim is made\. Seed\-aligned differences use the same seed labels but may not share identical client schedules or augmentation draws\. The primary archive contains6×3×3=546\\times 3\\times 3=54publication\-table runs\. The reuse\-only suite contains 36 FedAvg/DG\-FedReuse runs, and the FedAdam final suite contains 18 runs after a 16\-run validation search\.
## 6Results
### 6\.1Fixed\-round matched sparse\-control comparison
Table[6](https://arxiv.org/html/2608.05358#S6.T6)is the primary descriptive comparison\. Round 90 is the earliest terminal round among the 54 primary runs, so every trajectory contains the evaluation and no missing value is imputed\. Matched Top\-K FedAvg records 76\.88% cumulative restricted uplink saving\. DG\-FedReuse records 83\.36–85\.42%, an additional 6\.48–8\.54 percentage points under the same field model\. Accuracy is lower for DG\-FedReuse on every dataset at this budget, with seed\-aligned mean differences from−0\.14\-0\.14to−5\.29\-5\.29percentage points\.
Table 6:Primary fixed\-budget comparison at round 90\. Accuracy and cumulative restricted update\-uplink saving are mean±\\pmsample SD over seeds 101/202/303\. Differences are seed\-aligned, not guaranteed to use identical random trajectories\.The fixed\-budget evidence therefore shows a communication–accuracy trade\-off, not utility preservation\. The deficit is small on MNIST, FashionMNIST, and EMNIST Balanced at round 90, but is substantial on PathMNIST and both CIFAR tasks\. No result establishes lower bytes to a common target accuracy because target attainment was not prespecified and several trajectories have different convergence rates\.
### 6\.2Dense\-downlink sensitivity
Equation \([16](https://arxiv.org/html/2608.05358#S4.E16)\) adds one uncompressed model download per selected client per round\. Under this deliberately simple symmetric reference, Top\-K FedAvg’s saving becomes 38\.44%, and DG\-FedReuse becomes 41\.68–42\.71% \(Table[7](https://arxiv.org/html/2608.05358#S6.T7)\)\. The incremental advantage is then 3\.24–4\.27 percentage points\. Control traffic and representation metadata remain excluded, so these values are still model\-based rather than measured\.
Table 7:Round\-90 communication sensitivity with one dense model downlink plus the modeled uplink, relative to dense downlink plus dense uplink\.The difference between Tables[6](https://arxiv.org/html/2608.05358#S6.T6)and[7](https://arxiv.org/html/2608.05358#S6.T7)is consequential\. A headline based only on update uplink overstates the fraction of a bidirectional dense reference removed when model broadcast remains uncompressed\. This sensitivity is consistent with the broader observation that downlink requires separate treatment in cross\-device FL\[[4](https://arxiv.org/html/2608.05358#bib.bib26)\]\.
### 6\.3Archival test\-controlled checkpoint summary
Table[8](https://arxiv.org/html/2608.05358#S6.T8)preserves the originally reported best\-observed test results for auditability\. Because the test stream controlled checkpoint selection and early stopping, the values are not unbiased final\-test estimates\. They should not be used to assert equivalence, superiority, or generalization\.
Table 8:Exploratory archival summary under test\-controlled checkpointing\. All methods user=0\.2r=0\.2and seeds 101/202/303\. Saving is the run mean of the restricted update\-uplink counter\.The archival maxima are more favorable than the common\-round values, particularly for CIFAR\-10 and CIFAR\-100\. This divergence illustrates why maximizing repeatedly observed test accuracy can materially change interpretation\[[2](https://arxiv.org/html/2608.05358#bib.bib27)\]\. Figure[3](https://arxiv.org/html/2608.05358#S6.F3)is retained as a diagnostic visualization of these archival values\.
Figure 3:Exploratory archival best\-observed test accuracy versus restricted update\-uplink saving\. Error bars are sample standard deviations over three seeds\. The plot is diagnostic because checkpoint selection used the test stream\.
### 6\.4FedAdam secondary analysis
FedAdam was selected on CIFAR\-10/CIFAR\-100 validation data with development seeds 61 and 73; server learning rate 0\.01,β1=0\.9\\beta\_\{1\}=0\.9,β2=0\.99\\beta\_\{2\}=0\.99, andτ=0\.001\\tau=0\.001were then frozen\. Its final suite nevertheless uses the same test\-controlled stopping design as the primary archive, so Table[9](https://arxiv.org/html/2608.05358#S6.T9)is also exploratory\.
Table 9:Exploratory FedAdam summary under test\-controlled checkpointing\. Accuracy and rounds are mean±\\pmsample SD over seeds 101/202/303\.Figure 4:Test\-observed trajectories for the three FedAvg\-family methods and separately selected FedAdam\. Lines are three\-seed means and bands are sample standard deviations\. Curves terminate at the earliest stopping round represented in all three corresponding runs\. They are diagnostic test trajectories, not validation curves or confidence bands\.
### 6\.5Descriptive mechanism decomposition
Table[10](https://arxiv.org/html/2608.05358#S6.T10)juxtaposes dense FedAvg, reuse without Top\-K, Top\-K without reuse, and the combined mechanism\. The conditions come from separately executed suites; they do not form a randomized2×22\\times 2factorial experiment with identical schedules and stopping rules\. The table is therefore descriptive and cannot identify an independent causal effect or interaction\.
Table 10:Descriptive cross\-suite mechanism comparison\. Each cell is exploratory best\-observed test accuracy / restricted update\-uplink saving, mean±\\pmsample SD\.Figure 5:Restricted update\-uplink saving in separately executed dense, reuse\-only, Top\-K\-only, and combined suites\. The visualization is not a factorial causal decomposition or a runtime result\.Reuse\-only saving ranges from 32\.43% to 37\.28%, which is consistent with cached\-update reuse contributing additional field\-count reduction beyond sparsification\. However, reuse\-only CIFAR\-10 accuracy is 5\.26 percentage points below dense FedAvg, and the separate\-suite design does not estimate an interaction between reuse and Top\-K\.
### 6\.6Evidence supported by the executed study
The executed evidence supports the following bounded observations:
1. 1\.At round 90, the implemented gate increases restricted update\-uplink saving over Top\-K FedAvg from 76\.88% to 83\.36–85\.42%, while reducing mean accuracy by 0\.14–5\.29 percentage points across the six tasks\.
2. 2\.Adding one dense model downlink per selected client reduces the corresponding communication ratio to 41\.68–42\.71%, with a 3\.24–4\.27 percentage\-point incremental gain over Top\-K FedAvg\.
3. 3\.Separately executed reuse\-only runs produce 32\.43–37\.28% restricted uplink saving, but do not constitute a factorial attribution experiment\.
4. 4\.Test\-controlled maxima and trajectories are reproducible archival outputs, not unbiased held\-out performance estimates\.
The study does not establish superiority over MIFA, FedStale, 3PC\-derived lazy aggregation, FedLUAR, or error\-feedback compression; faster convergence; lower communication to a target accuracy; reduced client FLOPs; or end\-to\-end network, runtime, energy, privacy, or robustness gains\.
## 7Discussion
### 7\.1Interpretation of the matched sparse controls
Applying the samer=0\.2r=0\.2field\-count rule to FedAvg, FedProx, and DG\-FedReuse prevents the full 76\.88% Top\-K reduction from being attributed to reuse\. The remaining uplink difference is associated with the implemented active/reuse decisions under the simulator’s 16\-byte reused\-event convention\. This is a useful mechanism control, but it is narrower than a state\-of\-the\-art comparison because Top\-K lacks error feedback and the closest stale\-update and lazy\-aggregation methods are absent\.
### 7\.2Accuracy–communication behavior
The fixed\-round analysis exposes slower transient progress on PathMNIST and both CIFAR tasks\. The archival maxima close part of this gap only after method\-dependent stopping and repeated test observation\. Therefore, the method should be viewed as a tunable communication–accuracy mechanism rather than an accuracy\-preserving replacement for fresh training\. A clean study should choose accuracy tolerances or target levels on validation data before observing the final test set\.
### 7\.3What is distinctive and what is not
Historical\-update use is not new: MIFA, FedStale, 3PC\-style lazy aggregation, and FedLUAR already establish memorized, stale, lazy, or recycled update mechanisms\[[5](https://arxiv.org/html/2608.05358#bib.bib22),[19](https://arxiv.org/html/2608.05358#bib.bib23),[20](https://arxiv.org/html/2608.05358#bib.bib24),[10](https://arxiv.org/html/2608.05358#bib.bib14)\]\. The distinct element evaluated here is the combination of a client\-specific stochastic head\-proxy score, an explicit age cap, a minimum fresh\-client quota, age decay, and a matched sparse numerical\-field accounting study\. Whether that combination improves the accuracy–communication frontier relative to the closest methods remains unresolved\.
## 8Threats to validity and limitations
##### Test\-controlled selection\.
The historical final configurations use the canonical test stream for checkpoint selection and early stopping\. Early shared and gate searches also used a test\-labelled development stream\. The common\-round table removes maximization over evaluation time but remains a post\-hoc analysis of repeatedly observed test data\. A confirmatory submission requires validation\-derived selection and once\-only final test evaluation\[[2](https://arxiv.org/html/2608.05358#bib.bib27)\]\.
##### Stochastic gate validity\.
The proxy uses shuffled batches, training\-mode model behavior, and random CIFAR augmentation\. Its repeatability, score variance, and decision\-flip rate at a fixed global/client state were not measured\. Consequently, the manuscript does not interpretDi,tD\_\{i,t\}as pure client drift\. A deterministic\-proxy ablation and repeated\-measurement reliability analysis are required to establish that the gate responds to the intended signal\.
##### Compression baseline\.
Direct Top\-K discards residual coordinates without error feedback\. Because biased compression can alter convergence and error\-feedback behavior changes under partial participation, a matched Top\-K\-with\-error\-feedback baseline is required for a stronger compression comparison\[[14](https://arxiv.org/html/2608.05358#bib.bib25)\]\. Client residual\-buffer memory and any additional metadata must be included\.
##### Modeled rather than measured communication\.
No distributed transport is executed\. The primary counter excludes model downlink, tensor and mode metadata, control scores, framing, acknowledgements, and retransmissions\. The symmetric sensitivity adds a dense downlink but still omits those terms\. End\-to\-end communication requires serialized uplink and downlink measurement under an explicit protocol; downlink should be treated separately rather than assumed negligible\[[4](https://arxiv.org/html/2608.05358#bib.bib26)\]\.
##### Client computation and storage\.
Reusable clients still compute a proxy, and cacheful active clients compute both a gating and refresh proxy\. The shallow\-CNN head is large\. No FLOP, latency, energy, or device\-memory reduction is measured\. Cached sparse updates are stored as dense\-shaped CPU tensors, so server memory scales with dense model size per populated client cache\.
##### Randomization and statistical power\.
Equal seed labels do not guarantee identical method\-specific random trajectories, and cuDNN deterministic algorithms are not enforced\. Three seeds permit descriptive standard deviations but provide weak tail and uncertainty characterization\. The study therefore avoids significance and equivalence claims\. A rerun should freeze client partitions and selection schedules, separate random streams by function, record environment hashes, and use more seeds with prespecified intervals or equivalence margins\.
##### Mechanism attribution\.
The dense, reuse\-only, Top\-K\-only, and combined conditions were executed in separate suites\. They do not form a common2×22\\times 2factorial design\. Independent effects and interactions therefore cannot be claimed\. A factorial experiment must use identical partitions, schedules, budgets, model\-selection rules, and final evaluation\.
##### Closest\-method comparison\.
MIFA, FedStale, 3PC\-derived lazy aggregation, FedLUAR, LAQ, LBGM, and GradSkip were not implemented under the current experimental contract\. The present controls isolate an internal mechanism but cannot support a broad novelty or superiority conclusion\.
##### Theory, privacy, and robustness\.
The minimum\-freshness and bounded\-age properties are implementation invariants, not convergence guarantees\. No analysis bounds stale\-direction error jointly with Top\-K error\. Differential privacy, secure aggregation, poisoning, gradient leakage, and adversarial robustness were not evaluated; no privacy or security guarantee is made\[[26](https://arxiv.org/html/2608.05358#bib.bib21)\]\.
## 9Required confirmatory work
A confirmatory evaluation should first create validation partitions exclusively from training data, freeze all hyperparameters and round/checkpoint rules, and evaluate each canonical test set once\. It should then compare FedAvg, FedProx, Top\-K with and without error feedback, MIFA, FedStale, a 3PC\-derived lazy method, and FedLUAR under identical partitions, client schedules, budgets, and accounting\. At least five independent seeds should be reported with seed\-level observations, confidence intervals, and prespecified accuracy\-retention or equivalence criteria\.
The mechanism study should use a2×22\\times 2reuse×\\timesTop\-K factorial design and ablate the proxy, age cap, quota, decay, signature EMA, proxy\-batch count, and threshold schedule\. Gate reliability should be measured by repeated proxy evaluations at fixed states, including score variance and decision\-flip rates\. System validation should serialize both model downlink and update uplink, including schema, tensor and mode identifiers, control scores, framing, acknowledgements, and retransmissions; it should also report bytes and wall\-clock time to a validation\-defined target accuracy, proxy and training FLOPs, energy, and cache memory\.
## 10Reproducibility and evidence governance
The supplied evidence distinguishes a 54\-run FedAvg/FedProx/DG\-FedReuse primary archive, a 36\-run dense/reuse\-only suite, a complete 84\-run two\-stage Top\-K validation study, a 16\-run FedAdam validation search, and an 18\-run FedAdam final archive\. The larger raw matched and uncompressed archives contain additional algorithm labels and artifacts; run presence does not make every label eligible for publication claims\. The manuscript tables use only the stated subsets\.
Each run directory records resolved configuration, metrics, summary, log, and checkpoints where available\. The compact evidence manifest reports 1,558 entries and 412 non\-empty metric trajectories\. The corrected release test path points toconfigs/baseline\_exp/smoke\_synthetic\.yaml; the packaged repository passes 22 tests and 14 subtests\. These checks establish artifact consistency, not scientific validity of test\-selected estimates\.
## 11Conclusion
DG\-FedReuse combines a stochastic proxy\-gradient gate, bounded cache age, a minimum fresh\-client quota, staleness decay, and matched Top\-K numerical\-field accounting\. At a common 90\-round budget, the implemented simulator records 83\.36–85\.42% restricted update\-uplink saving versus 76\.88% for Top\-K FedAvg, accompanied by seed\-aligned accuracy differences from−0\.14\-0\.14to−5\.29\-5\.29percentage points\. Adding one dense model downlink per selected client reduces the corresponding ratio to 41\.68–42\.71% and the incremental gain to 3\.24–4\.27 percentage points\. These observations show that the mechanism changes the modeled communication–accuracy trade\-off under the evaluated simulator\. They do not establish unbiased final\-test performance, end\-to\-end communication or computation savings, convergence superiority, or an advantage over existing stale\-update, lazy\-aggregation, recycling, and error\-feedback methods\. Those questions require the validation\-controlled, closest\-baseline, factorial, and distributed\-system experiments specified above\.
## Appendix AExact search spaces and fixed choices
The shared Optuna search used learning rate\{0\.001,0\.005,0\.01,0\.05\}\\\{0\.001,0\.005,0\.01,0\.05\\\}, local epochs\{1,2,3\}\\\{1,2,3\\\}, batch size\{128,256\}\\\{128,256\\\}, weight decay\{10−4,5×10−4,10−3\}\\\{10^\{\-4\},5\\times 10^\{\-4\},10^\{\-3\}\\\}, momentum\{0,0\.9\}\\\{0,0\.9\\\}, and proxy batch\{32,64,128\}\\\{32,64,128\\\}\. The DG search sampledτ0∈\[0\.85,0\.99\]\\tau\_\{0\}\\in\[0\.85,0\.99\],τmin∈\[0\.40,min\(τ0−0\.01,0\.85\)\]\\tau\_\{\\min\}\\in\[0\.40,\\min\(\\tau\_\{0\}\-0\.01,0\.85\)\],S∈\{2,…,15\}S\\in\\\{2,\\ldots,15\\\},ρ∈\[0\.5,0\.95\]\\rho\\in\[0\.5,0\.95\], and log\-uniformμ∈\[10−5,10−2\]\\mu\\in\[10^\{\-5\},10^\{\-2\}\]\. The search constrained mean active fraction to at most 0\.70\. Fixed DG choices wereγ=0\.01\\gamma=0\.01,q=0\.30q=0\.30, and a head signature\. The secondary FedOpt validation fixedβ1=0\.9\\beta\_\{1\}=0\.9,β2=0\.99\\beta\_\{2\}=0\.99, and stabilizerτ=0\.001\\tau=0\.001while searching server learning rate\.
## Appendix BModel\-state and proxy sizes
Table[11](https://arxiv.org/html/2608.05358#A2.T11)makes the simulator boundary concrete\. Dense update bytes use actual tensor element widths in the state dictionary\. Proxy bytes show an fp32 representation of the implemented head\-gradient signature; they are not charged in Eq\. \([15](https://arxiv.org/html/2608.05358#S4.E15)\)\. The logical mapping in Fig\.[1](https://arxiv.org/html/2608.05358#S3.F1)keeps that vector client\-local\. Because the current sparse cache is stored in dense\-shaped tensors, its tensor storage forNcN\_\{c\}populated client caches is exactlyNcB0N\_\{c\}B\_\{0\}before framework overhead\.
Table 11:Model, proxy, and dense\-shaped update\-cache sizes used to audit the accounting boundary\. GB uses10910^\{9\}bytes and assumes one populated update cache per client; framework overhead and signature storage are excluded from the last two columns\.
## Appendix CSeed\-aligned archival accuracy differences
Table[12](https://arxiv.org/html/2608.05358#A3.T12)reports the three seed\-aligned observations underlying the exploratory archival DG–FedAvg best\-observed test comparison\. Equal seed labels use the corresponding partition protocol, but method\-specific execution can consume different random\-number sequences and need not produce identical client schedules\. The table increases transparency but is not a paired randomized experiment, equivalence test, or significance test\.
Table 12:Seed\-aligned exploratory best\-observed test differences, DG\-FedReuse minus Top\-K FedAvg, in percentage points\. The final column is mean±\\pmsample SD across seed labels\.
## Appendix DBaseline evidence disposition
The larger archived suite contains six algorithm labels, but run completeness alone is not sufficient for scientific comparison\. FedAvg and FedProx are the matched sparse mechanism controls\. Historical FedOpt rows are excluded; only the separately selected 18\-run FedAdam suite in Table[9](https://arxiv.org/html/2608.05358#S6.T9)is summarized\. SCAFFOLD would require verified control\-variate evolution and complete traffic accounting, while Ditto would require personalized\-client outcomes\. No archived label substitutes for a correctly implemented closest\-method baseline\.
## Appendix EArtifact lineage
The numerical claims in the manuscript trace to four frozen evidence roles:
1. 1\.the validation\-only CIFAR Top\-K study for ratio selection;
2. 2\.the uncompressed three\-seed suite for descriptive reuse\-only evidence;
3. 3\.the three\-seed matchedr=0\.2r=0\.2suite for the sparse\-control comparison; and
4. 4\.the validation\-selected 18\-run FedAdam suite for the secondary optimizer analysis\.
Resolved run configurations override template YAML and narrative documentation when they differ\. Development trials are never counted as final seeds\. A result row is admitted only after the expected dataset–method–seed coverage and required artifacts are complete\.
## Acknowledgments
The authors gratefully acknowledge the support of the Variable Energy Cyclotron Centre \(VECC\), the Department of Atomic Energy \(DAE\), Government of India, for providing the infrastructure and technical environment that supported this research\. The authors also thank the staff of the VECC library for their assistance during the course of this study\.
## Declarations
#### Consent to Publish
All authors have read and approved the final manuscript and consent to its submission and publication\.
#### Conflict of Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper\.
## References
- \[1\]S\. S\. Azam, S\. Hosseinalipour, Q\. Qiu, and C\. Brinton\(2022\)Recycling model updates in federated learning: are gradient subspaces low\-rank?\.International Conference on Learning Representations\.Cited by:[§2\.3](https://arxiv.org/html/2608.05358#S2.SS3.p1.1),[Table 1](https://arxiv.org/html/2608.05358#S2.T1.1.6.5.1.1.1)\.
- \[2\]G\. C\. Cawley and N\. L\. C\. Talbot\(2010\)On over\-fitting in model selection and subsequent selection bias in performance evaluation\.Journal of Machine Learning Research11\(70\),pp\. 2079–2107\.Cited by:[§5\.1](https://arxiv.org/html/2608.05358#S5.SS1.p2.1),[§5\.3](https://arxiv.org/html/2608.05358#S5.SS3.p1.1),[§6\.3](https://arxiv.org/html/2608.05358#S6.SS3.p2.1),[§8](https://arxiv.org/html/2608.05358#S8.SS0.SSS0.Px1.p1.1)\.
- \[3\]G\. Cohen, S\. Afshar, J\. Tapson, and A\. van Schaik\(2017\)EMNIST: extending mnist to handwritten letters\.In2017 International Joint Conference on Neural Networks \(IJCNN\),Vol\.,pp\. 2921–2926\.External Links:[Document](https://dx.doi.org/10.1109/IJCNN.2017.7966217)Cited by:[Table 3](https://arxiv.org/html/2608.05358#S5.T3.1.4.3.1)\.
- \[4\]R\. Dorfman, S\. Vargaftik, Y\. Ben\-Itzhak, and K\. Y\. Levy\(2023\)DoCoFL: downlink compression for cross\-device federated learning\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 8356–8388\.Cited by:[§2\.2](https://arxiv.org/html/2608.05358#S2.SS2.p2.1),[§6\.2](https://arxiv.org/html/2608.05358#S6.SS2.p2.1),[§8](https://arxiv.org/html/2608.05358#S8.SS0.SSS0.Px4.p1.1)\.
- \[5\]X\. Gu, K\. Huang, J\. Zhang, and L\. Huang\(2021\)Fast federated learning in the presence of arbitrary device unavailability\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 12052–12064\.Cited by:[§1](https://arxiv.org/html/2608.05358#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.05358#S2.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.05358#S2.T1.1.2.1.1.1.1),[§7\.3](https://arxiv.org/html/2608.05358#S7.SS3.p1.1)\.
- \[6\]K\. He, X\. Zhang, S\. Ren, and J\. Sun\(2016\)Deep residual learning for image recognition\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,pp\. 770–778\.Cited by:[§5\.1](https://arxiv.org/html/2608.05358#S5.SS1.p1.1)\.
- \[7\]D\. Jhunjhunwala, S\. Wang, and G\. Joshi\(2022\)FedVARP: tackling the variance due to partial client participation in federated learning\.InProceedings of the Thirty\-Eighth Conference on Uncertainty in Artificial Intelligence,Proceedings of Machine Learning Research, Vol\.180,pp\. 906–916\.Cited by:[§2\.1](https://arxiv.org/html/2608.05358#S2.SS1.p2.1)\.
- \[8\]P\. Kairouzet al\.\(2021\)Advances and open problems in federated learning\.Foundations and Trends in Machine Learning14\(1–2\),pp\. 1–210\.External Links:[Document](https://dx.doi.org/10.1561/2200000083)Cited by:[§1](https://arxiv.org/html/2608.05358#S1.p1.1)\.
- \[9\]S\. P\. Karimireddy, S\. Kale, M\. Mohri, S\. Reddi, S\. Stich, and A\. T\. Suresh\(2020\)SCAFFOLD: stochastic controlled averaging for federated learning\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 5132–5143\.Cited by:[§2\.1](https://arxiv.org/html/2608.05358#S2.SS1.p1.1)\.
- \[10\]J\. Kim, S\. Kang, and S\. Lee\(2025\)Layer\-wise update aggregation with recycling for communication\-efficient federated learning\.InAdvances in Neural Information Processing Systems,Vol\.38\.Cited by:[§1](https://arxiv.org/html/2608.05358#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.05358#S2.SS3.p1.1),[Table 1](https://arxiv.org/html/2608.05358#S2.T1.1.8.7.1.1.1),[§7\.3](https://arxiv.org/html/2608.05358#S7.SS3.p1.1)\.
- \[11\]J\. Konečný, H\. B\. McMahan, F\. X\. Yu, P\. Richtárik, A\. T\. Suresh, and D\. Bacon\(2016\)Federated learning: strategies for improving communication efficiency\.29th Conference on Neural Information Processing Systems \(NIPS 2016\), Barcelona, Spain\.\.Cited by:[§1](https://arxiv.org/html/2608.05358#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.05358#S2.SS2.p1.1)\.
- \[12\]A\. Krizhevsky\(2009\)Learning multiple layers of features from tiny images\.Technical reportUniversity of Toronto\.Cited by:[Table 3](https://arxiv.org/html/2608.05358#S5.T3.1.6.5.1),[Table 3](https://arxiv.org/html/2608.05358#S5.T3.1.7.6.1)\.
- \[13\]T\. Li, A\. K\. Sahu, M\. Zaheer, M\. Sanjabi, A\. Talwalkar, and V\. Smith\(2020\)Federated optimization in heterogeneous networks\.InProceedings of Machine Learning and Systems,Vol\.2,pp\. 429–450\.Cited by:[§2\.1](https://arxiv.org/html/2608.05358#S2.SS1.p1.1)\.
- \[14\]X\. Li and P\. Li\(2023\)Analysis of error feedback in federated non\-convex optimization with biased compression: fast convergence and partial participation\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 19638–19688\.Cited by:[§1](https://arxiv.org/html/2608.05358#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.05358#S2.SS2.p1.1),[§4\.4](https://arxiv.org/html/2608.05358#S4.SS4.p1.4),[§8](https://arxiv.org/html/2608.05358#S8.SS0.SSS0.Px3.p1.1)\.
- \[15\]A\. Maranjyan, M\. Safaryan, and P\. Richtárik\(2025\)GradSkip: communication\-accelerated local gradient methods with better computational complexity\.Transactions on Machine Learning Research\.Cited by:[§2\.3](https://arxiv.org/html/2608.05358#S2.SS3.p1.1),[Table 1](https://arxiv.org/html/2608.05358#S2.T1.1.7.6.1.1.1)\.
- \[16\]H\. B\. McMahan, E\. Moore, D\. Ramage, S\. Hampson, and B\. A\. y Arcas\(2017\)Communication\-efficient learning of deep networks from decentralized data\.InProceedings of the 20th International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.54,pp\. 1273–1282\.Cited by:[§1](https://arxiv.org/html/2608.05358#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.05358#S2.SS1.p1.1)\.
- \[17\]S\. J\. Reddi, Z\. Charles, M\. Zaheer, Z\. Garrett, K\. Rush, J\. Konečný, S\. Kumar, and H\. B\. McMahan\(2021\)Adaptive federated optimization\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.05358#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.05358#S2.SS1.p1.1)\.
- \[18\]A\. Reisizadeh, A\. Mokhtari, H\. Hassani, A\. Jadbabaie, and R\. Pedarsani\(2020\)FedPAQ: a communication\-efficient federated learning method with periodic averaging and quantization\.InProceedings of the 23rd International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.108,pp\. 2021–2031\.Cited by:[§1](https://arxiv.org/html/2608.05358#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.05358#S2.SS2.p1.1)\.
- \[19\]P\. Richtárik, I\. Sokolov, E\. Gasanov, I\. Fatkhullin, Z\. Li, and E\. Gorbunov\(2022\)3PC: three point compressors for communication\-efficient distributed training and a better theory for lazy aggregation\.InProceedings of the 39th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.162,pp\. 18596–18648\.Cited by:[§1](https://arxiv.org/html/2608.05358#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.05358#S2.SS3.p1.1),[Table 1](https://arxiv.org/html/2608.05358#S2.T1.1.4.3.1.1.1),[§7\.3](https://arxiv.org/html/2608.05358#S7.SS3.p1.1)\.
- \[20\]A\. Rodio and G\. Neglia\(2024\)FedStale: leveraging stale updates in federated learning\.InECAI 2024,Frontiers in Artificial Intelligence and Applications, Vol\.392,pp\. 3071–3078\.External Links:[Document](https://dx.doi.org/10.3233/FAIA240849)Cited by:[§1](https://arxiv.org/html/2608.05358#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.05358#S2.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.05358#S2.T1.1.3.2.1.1.1),[§7\.3](https://arxiv.org/html/2608.05358#S7.SS3.p1.1)\.
- \[21\]F\. Sattler, S\. Wiedemann, K\. Müller, and W\. Samek\(2020\)Robust and communication\-efficient federated learning from non\-IID data\.IEEE Transactions on Neural Networks and Learning Systems31\(9\),pp\. 3400–3413\.External Links:[Document](https://dx.doi.org/10.1109/TNNLS.2019.2944481)Cited by:[§1](https://arxiv.org/html/2608.05358#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.05358#S2.SS2.p1.1)\.
- \[22\]S\. U\. Stich, J\. Cordonnier, and M\. Jaggi\(2018\)Sparsified SGD with memory\.InAdvances in Neural Information Processing Systems,Vol\.31\.Cited by:[§2\.2](https://arxiv.org/html/2608.05358#S2.SS2.p1.1)\.
- \[23\]J\. Sun, T\. Chen, G\. B\. Giannakis, Q\. Yang, and Z\. Yang\(2022\)Lazily aggregated quantized gradient innovation for communication\-efficient federated learning\.IEEE Transactions on Pattern Analysis and Machine Intelligence44\(4\),pp\. 2031–2044\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2020.3033286)Cited by:[§2\.3](https://arxiv.org/html/2608.05358#S2.SS3.p1.1),[Table 1](https://arxiv.org/html/2608.05358#S2.T1.1.5.4.1.1.1)\.
- \[24\]H\. Xiao, K\. Rasul, and R\. Vollgraf\(2017\)Fashion\-MNIST: a novel image dataset for benchmarking machine learning algorithms\.arXiv preprint arXiv:1708\.07747\.Cited by:[Table 3](https://arxiv.org/html/2608.05358#S5.T3.1.3.2.1)\.
- \[25\]J\. Yang, R\. Shi, and B\. Ni\(2023\)MedMNIST v2—a large\-scale lightweight benchmark for 2d and 3d biomedical image classification\.Scientific Data10\(1\),pp\. 41\.External Links:[Document](https://dx.doi.org/10.1038/s41597-022-01721-8)Cited by:[Table 3](https://arxiv.org/html/2608.05358#S5.T3.1.5.4.1)\.
- \[26\]L\. Zhu, Z\. Liu, and S\. Han\(2019\)Deep leakage from gradients\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[§8](https://arxiv.org/html/2608.05358#S8.SS0.SSS0.Px9.p1.1)\.
## Author Biographies
![[Uncaptioned image]](https://arxiv.org/html/2608.05358v1/RA.png)Rahil Aftabis an undergraduate student pursuing a Bachelor of Technology \(B\.Tech\.\) in Computer Science and Engineering at Jamia Hamdard, New Delhi, India\. His research interests include artificial intelligence, computer vision, federated learning, privacy\-preserving machine learning, and large language models\. He has contributed to research projects involving federated learning, benchmark evaluation, and AI\-driven applications\. He completed a research internship at the Variable Energy Cyclotron Centre \(VECC\), Department of Atomic Energy, India, where he conducted research on federated learning for privacy\-preserving AI systems\. His broader interests include trustworthy AI and the real\-world deployment of intelligent systems\.![[Uncaptioned image]](https://arxiv.org/html/2608.05358v1/VKR.jpg)Vineet Kumar Rakeshis a Technical Officer \(Scientific Category\) at the Variable Energy Cyclotron Centre \(VECC\), Department of Atomic Energy, India, with over 23 years of experience in software engineering, database systems, and artificial intelligence\. His research focuses on talking head generation, lip reading, and ultra\-low\-bitrate video compression for real\-time teleconferencing\. He is currently pursuing a Ph\.D\. at Homi Bhabha National Institute, Mumbai\. Mr\. Rakesh has contributed to office automation, OCR systems, and digital transformation projects at VECC\. He is an Associate Member of the Institution of Engineers \(India\) and a recipient of the DAE Group Achievement Award\.![[Uncaptioned image]](https://arxiv.org/html/2608.05358v1/SM.png)Soumya Mazumdaris a student researcher pursuing a B\.S\. in Data Science and Applications at the Indian Institute of Technology Madras and B\.Tech\. in Computer Science and Business Systems at West Bengal University of Technology \(GMIT campus\), India\. His work focuses on temporal generative modeling, geometry\-aware computer vision, and controllable video synthesis, with particular interest in diffusion\-based methods for talking\-head generation and temporal consistency\. He has served as a Research Trainee at the Variable Energy Cyclotron Centre \(VECC\), where he worked on pose\- and landmark\-conditioned video generation, benchmarking, and efficient deployment pipelines\. He has contributed to research publications in journals, conference proceedings, and edited volumes, and is also associated with an Indian patent in neural network\-based real\-time analysis\.![[Uncaptioned image]](https://arxiv.org/html/2608.05358v1/TS.png)Dr\. Tapas Samantais a senior scientist and Head of the Computer and Informatics Group at the Variable Energy Cyclotron Centre \(VECC\), Department of Atomic Energy, India\. With over two decades of experience, his work spans artificial intelligence, industrial automation, embedded systems, high\-performance computing, and accelerator control systems\. He also leads technology transfer initiatives and public scientific outreach at VECC\.相似文章
准确且资源高效的联邦持续学习
FedRAN是一种资源感知的分析型联邦持续学习框架,用紧凑的随机特征统计量替代基于梯度的更新,在显著降低通信与计算成本的同时实现高精度。
从局部失配到全局影响:优化高效扩散的缓存复用策略
本文提出Global-ImpactCache(GCache),一种双层优化框架,通过学习扩散模型的缓存复用策略,将误差加权与最终生成质量对齐,而非依赖局部相似性启发式。它在图像和视频生成任务上实现了显著的加速和质量提升,包括在Wan2.1上实现2.17倍加速且LPIPS更低。
Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients
This paper proposes FedSLM, a parameter-centric framework for federated fine-tuning of foundation models with heterogeneous compressed clients, using SVD-based decomposition and a weak-to-strong elicitation step to handle resource asymmetry. Experiments show it outperforms existing federated baselines while reducing client GPU memory by ~50%.
联邦MLLM微调中基于弹性正则化与合成回放的持续学习
提出FedCMM,一个面向多模态大模型的联邦持续学习框架,利用模态感知的弹性权重巩固、本地生成式回放以及任务相似性感知的梯度聚合来缓解灾难性遗忘。
联邦轻量级微调
本文介绍了FLITE(联邦低秩迭代训练引擎),一种联邦微调方法,通过使用冻结的仿射映射网络,从一个小型可训练潜变量和低秩可种子重生的因子分解生成权重,将每轮每客户端的通信量降至每轮1280个浮点数(约5KB)——相比于全权重FedAvg减少了8718倍。在CIFAR-100数据集上使用ResNet-18进行测试,准确率与全权重FedAvg相差在0.5个百分点以内。