Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients
摘要
This paper proposes FedSLM, a parameter-centric framework for federated fine-tuning of foundation models with heterogeneous compressed clients, using SVD-based decomposition and a weak-to-strong elicitation step to handle resource asymmetry. Experiments show it outperforms existing federated baselines while reducing client GPU memory by ~50%.
查看缓存全文
缓存时间: 2026/08/03 07:36
# Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients
Source: [https://arxiv.org/html/2607.29071](https://arxiv.org/html/2607.29071)
Shengkun Zhu†, Jinshan Zeng†, Zhihua Allen\-Zhao†, Mayi Xu, Quanqing Xu∗, Wei Ren, Qiang Yang, Yang Liu∗†Equal contribution\.∗Corresponding authors: Q\. Xu and Y\. Liu\.S\. Zhu is with The Hong Kong Polytechnic University, and also with OceanBase, Ant Group\. E\-mail: zhushengkun\.zsk@antgroup\.comJ\. Zeng is with the School of Management, Xi’an Jiaotong University\. E\-mail: jsh\.zeng@gmail\.comZ\. Allen\-Zhao is with the School of Mathematics and Statistics, Xidian University\. E\-mail: allenzhaozh@gmail\.comM\. Xu is with the School of Computer Science, Wuhan University\.Q\. Xu is with OceanBase, Ant Group\. E\-mail: xuquanqing\.xqq@oceanbase\.comW\. Ren is with the School of Computer Science, China University of Geosciences\. E\-mail: weirencs@cug\.edu\.cnY\. Liu and Q\. Yang are with The Hong Kong Polytechnic University\. E\-mail: yang\-veronica\.liu@polyu\.edu\.hk, profqiang\.yang@polyu\.edu\.hk
###### Abstract
Federated learning of foundation models faces a fundamental resource\-asymmetry challenge: the institutions holding the most valuable domain\-specific data cannot host billion\-parameter models\. Existing heterogeneous federated approaches attempt to bridge this gap through parameter\-efficient tuning, model pruning, or knowledge distillation, yet each trades away a critical property, whether full\-model memory reduction, architectural self\-containedness, or representational fidelity, leaving the core tension unresolved\. We propose FedSLM, a parameter\-centric framework for federated fine\-tuning with heterogeneous compressed clients\. FedSLM uses SVD\-based decomposition to produce self\-contained client models, whose low\-rank subspaces form nested manifolds that are structurally compatible for aggregation\. It then applies a two\-stage protocol that synchronizes lightweight adapters within compression groups and fuses full\-rank reconstructions across groups via structural alignment\. Finally, a weak\-to\-strong elicitation step with auxiliary confidence loss transfers the aggregated knowledge to the full\-scale server, while an explicit bias–variance trade\-off mitigates compression artifacts\. We provide theoretical guarantees for adapter\-level aggregation, subspace\-alignment bounds for cross\-group fusion, and a characterization of how the confidence loss mitigates weak\-supervision noise\. Experiments on natural language and vision–language benchmarks show that FedSLM outperforms existing federated baselines under both IID and non\-IID partitions, while client models operate at roughly 50% of the GPU memory required by the full model\.
## IIntroduction
Foundation Models \(FMs\) derive their capabilities from scaling model parameters and training data\[[1](https://arxiv.org/html/2607.29071#bib.bib79),[61](https://arxiv.org/html/2607.29071#bib.bib78),[31](https://arxiv.org/html/2607.29071#bib.bib80)\]\. As these models are increasingly deployed in specialized domains, the demand for domain\-specific training data has grown correspondingly\. However, much of this data resides in private silos, such as medical records\[[50](https://arxiv.org/html/2607.29071#bib.bib76),[41](https://arxiv.org/html/2607.29071#bib.bib77)\]and financial documents\[[60](https://arxiv.org/html/2607.29071#bib.bib75)\], where privacy regulations and commercial confidentiality preclude centralized collection\.
Federated learning \(FL\)\[[37](https://arxiv.org/html/2607.29071#bib.bib41),[62](https://arxiv.org/html/2607.29071#bib.bib42)\]offers a principled solution by enabling multiple participants to collaboratively train a shared model without exposing raw data\. Yet applying FL to billion\-parameter FMs exposes a*resource\-asymmetry paradox*: the institutions that hold the most valuable domain\-specific data \(hospitals, financial firms, and enterprise departments\) often operate under GPU memory budgets that preclude hosting or fine\-tuning the full model\. This paradox is not merely an engineering inconvenience; it creates a fundamental tension between*inference independence*\(each client must run a self\-contained model\) and*aggregation compatibility*\(heterogeneous client updates must be fusible into a coherent global model\)\. Existing approaches to heterogeneous FL, while effective in their respective settings, face three open challenges when applied to billion\-parameter FMs\.
Parameter\-efficient fine\-tuning methods such as LoRA\[[21](https://arxiv.org/html/2607.29071#bib.bib20)\]reduce trainable parameters to mere megabytes, and recent federated variants\[[71](https://arxiv.org/html/2607.29071#bib.bib22),[68](https://arxiv.org/html/2607.29071#bib.bib21),[3](https://arxiv.org/html/2607.29071#bib.bib23),[9](https://arxiv.org/html/2607.29071#bib.bib24),[56](https://arxiv.org/html/2607.29071#bib.bib25)\]further allow adapters of varying ranks to accommodate client heterogeneity\. However, this efficiency is confined to the*update*dimension: every client must still load the complete pre\-trained model for forward computation\. For Llama\-2\-7B\[[51](https://arxiv.org/html/2607.29071#bib.bib55)\], the base weights alone occupy over 13 GB of VRAM, so although the number of trainable parameters shrinks, the*deployment footprint*remains that of the full model\.
An alternative line of work enables model heterogeneity through architectural extraction\. HeteroFL\[[12](https://arxiv.org/html/2607.29071#bib.bib1)\]and FjORD\[[20](https://arxiv.org/html/2607.29071#bib.bib2)\]allow clients to train sub\-networks of varying widths via channel\-wise slicing or ordered dropout, while FedRolex\[[2](https://arxiv.org/html/2607.29071#bib.bib4)\], DepthFL\[[26](https://arxiv.org/html/2607.29071#bib.bib5)\], and ScaleFL\[[22](https://arxiv.org/html/2607.29071#bib.bib6)\]explore rolling, depth\-wise, and resource\-adaptive partitioning\. These methods are designed for CNNs where channels are relatively independent\. In Transformer architectures, however, reducing hidden dimensions can disrupt the tightly coupled head\-wise attention patterns formed during pre\-training, leading to performance degradation in the extracted sub\-models\. Moreover, because each sub\-model’s parameters are a subset of the global model, updates from clients of different capacities may interfere during aggregation, making it difficult to balance client\-side heterogeneity with server\-side model quality\.
Data\-centric approaches\[[58](https://arxiv.org/html/2607.29071#bib.bib14),[14](https://arxiv.org/html/2607.29071#bib.bib15),[67](https://arxiv.org/html/2607.29071#bib.bib18)\]sidestep architectural heterogeneity by exchanging synthetic samples or output probability distributions\. However, this requires projecting rich hidden\-layer representations onto the output vocabulary\[[8](https://arxiv.org/html/2607.29071#bib.bib46),[70](https://arxiv.org/html/2607.29071#bib.bib47)\], and when client and server models use different tokenizers, where vocabulary overlap can be as low as 6%\[[69](https://arxiv.org/html/2607.29071#bib.bib48)\], aligning the resulting logit spaces remains an open problem despite recent efforts in optimal transport\[[5](https://arxiv.org/html/2607.29071#bib.bib49)\], approximate likelihood matching\[[40](https://arxiv.org/html/2607.29071#bib.bib50)\], and token co\-occurrence alignment\[[29](https://arxiv.org/html/2607.29071#bib.bib52)\]\. Furthermore, capacity\-limited client models tend to introduce systematic approximation errors that, once aggregated, may propagate to the server through the*imitation fallacy*\[[6](https://arxiv.org/html/2607.29071#bib.bib19)\]\.
The common thread across these three challenges is the absence of a mechanism that simultaneously \(i\) produces client models compact enough for deployment on resource\-constrained clients, \(ii\) maintains structural alignment for parameter\-level aggregation across heterogeneous clients, and \(iii\) enables the full\-scale server to absorb federated domain knowledge without inheriting compression artifacts\. We proposeFedSLMto address all three requirements within a unified optimization framework\. Rather than requiring clients to host the full model or relying on lossy logit\-level communication, FedSLM embraces lightweight compressed client models and focuses on extracting and amplifying the domain\-specific knowledge they acquire through FL\.
FedSLM formulates the problem as a coupled optimization over three variables: client\-side adapters, server\-side reconstructed global weights, and the final strong server model\. First, we derive a family of compressed client models from the server FM via SVD\-based decomposition\[[65](https://arxiv.org/html/2607.29071#bib.bib30),[54](https://arxiv.org/html/2607.29071#bib.bib31),[52](https://arxiv.org/html/2607.29071#bib.bib32)\]at multiple compression ratios\. Since all compressed models are derived from the same server weights via rank truncation, they retain a shared low\-rank structure that naturally supports parameter\-level aggregation across different compression ratios\. Second, we design a two\-stage aggregation protocol: the first stage synchronizes lightweight adapters within groups of clients sharing the same compression ratio; the second stage reconstructs full\-parameter representations and fuses them across groups into a unified global model via structural alignment in the common full\-rank space\. Third, to bridge the capacity gap between compressed client models and the full\-scale server, we introduce weak\-to\-strong elicitation with an auxiliary confidence loss\[[6](https://arxiv.org/html/2607.29071#bib.bib19)\]that treats the aggregated model as a noisy but informative supervisor, transferring federated domain knowledge while mitigating compression artifacts through an explicit bias–variance trade\-off\.
Our contributions are as follows:
- •We propose FedSLM, a unified federated framework comprising three tightly coupled components: SVD\-based decomposition that produces self\-contained compressed models via nested low\-rank manifolds at variable compression ratios, a two\-stage aggregation protocol that achieves structural alignment within compression groups and fuses full\-rank reconstructions across groups, and weak\-to\-strong elicitation with an auxiliary confidence loss that transfers federated knowledge to the server while mitigating compression artifacts\.
- •We provide theoretical guarantees from three complementary perspectives\. We establish an𝒪\(T−1/2\)\\mathcal\{O\}\(T^\{\-1/2\}\)convergence rate for adapter\-level aggregation with an explicit client\-drift term that accounts for data heterogeneity\. We derive subspace\-alignment bounds that characterize the geometric compatibility of heterogeneous compression groups and bound the cross\-group aggregation error\. We further show that the auxiliary confidence loss mitigates compression noise during weak\-to\-strong elicitation through a bias–variance trade\-off, where the self\-anchoring term provides genuine noise suppression and the mixing coefficientα\\alphacontrols the optimal operating point\.
- •We conduct extensive experiments on natural language and vision–language benchmarks under both IID and non\-IID partitions\. FedSLM consistently outperforms existing federated baselines across all settings\. Compressed client models operate at roughly 50% of the GPU memory required by the full model while achieving strong task\-specific performance through federated adapter training, recovering and exceeding the zero\-shot accuracy of the uncompressed model on multiple benchmarks\.
## IIBackground and Related Work
We organize existing work by the*knowledge carrier*used for cross\-client transfer, which exposes the limitation that each paradigm inherits and motivates the design of FedSLM\.
### II\-AData\-Centric Knowledge Transfer
The earliest approaches to heterogeneous FL transfer knowledge through output\-level representations, such as prediction probabilities, synthetic samples, or soft labels, rather than model parameters\. Knowledge distillation\[[19](https://arxiv.org/html/2607.29071#bib.bib10)\]forms the foundation: a teacher’s soft targets serve as the carrier for transferring capabilities to a student of arbitrary architecture\. In federated settings, FedGEMS\[[7](https://arxiv.org/html/2607.29071#bib.bib12)\]enables a larger server model to selectively learn from smaller client models via logit exchange, while FedGKT\[[18](https://arxiv.org/html/2607.29071#bib.bib13)\]alternates knowledge transfer between edge CNNs and a server\-side large model\. FedKD\[[58](https://arxiv.org/html/2607.29071#bib.bib14)\]reduces communication by up to 94\.89% through adaptive mutual distillation\. For FM\-specific collaboration, FedMKT\[[14](https://arxiv.org/html/2607.29071#bib.bib15)\]addresses tokenizer heterogeneity through minimum edit distance alignment and selective knowledge transfer\. Generation\-based methods like CrossLM\[[11](https://arxiv.org/html/2607.29071#bib.bib16)\]leverage the FM’s generative capability to synthesize training data guided by small\-model feedback, and TAKFL\[[42](https://arxiv.org/html/2607.29071#bib.bib17)\]addresses knowledge dilution from heterogeneous devices through task arithmetic integration\.
Despite their architectural flexibility, data\-centric methods face two key challenges in the FM regime\. First, projecting rich hidden\-layer representations onto the output vocabulary discards much of the knowledge encoded in intermediate geometry\[[8](https://arxiv.org/html/2607.29071#bib.bib46),[70](https://arxiv.org/html/2607.29071#bib.bib47)\]\. Second, when client and server models use different tokenizers, where vocabulary overlap can be as low as 6%\[[69](https://arxiv.org/html/2607.29071#bib.bib48)\], aligning the resulting logit spaces remains an open problem despite recent efforts\[[5](https://arxiv.org/html/2607.29071#bib.bib49),[40](https://arxiv.org/html/2607.29071#bib.bib50),[30](https://arxiv.org/html/2607.29071#bib.bib51),[29](https://arxiv.org/html/2607.29071#bib.bib52)\]\.
### II\-BArchitecture\-Centric Heterogeneity
A second paradigm accommodates heterogeneous clients by extracting sub\-models of varying sizes from a global model, using*weight slices*as the knowledge carrier\. HeteroFL\[[12](https://arxiv.org/html/2607.29071#bib.bib1)\]pioneers this approach by allowing clients to train submodels of varying widths through channel selection\. FjORD\[[20](https://arxiv.org/html/2607.29071#bib.bib2)\]introduces Ordered Dropout to create nested representations with importance\-based pruning, while FedDrop\[[57](https://arxiv.org/html/2607.29071#bib.bib3)\]applies random dropout to generate heterogeneous subnets\. FedRolex\[[2](https://arxiv.org/html/2607.29071#bib.bib4)\]proposes rolling sub\-model extraction, where the extraction window advances each round\. Beyond width\-based scaling, DepthFL\[[26](https://arxiv.org/html/2607.29071#bib.bib5)\]explores depth\-based scaling through layer pruning, and ScaleFL\[[22](https://arxiv.org/html/2607.29071#bib.bib6)\]combines both dimensions for two\-dimensional scaling\. These methods are designed for CNNs where channels are relatively independent; in Transformer architectures, however, reducing hidden dimensions can disrupt the coupled head\-wise attention patterns formed during pre\-training, leading to performance degradation\.
### II\-CPEFT\-Based Federated Fine\-Tuning
A third paradigm uses*incremental low\-rank matrices*\(e\.g\., LoRA adapters\) as the knowledge carrier\. LoRA\[[21](https://arxiv.org/html/2607.29071#bib.bib20)\]freezes pre\-trained weights and learns low\-rank decomposition matrices𝐁∈ℝm×r\\mathbf\{B\}\\in\\mathbb\{R\}^\{m\\times r\}and𝐀∈ℝr×n\\mathbf\{A\}\\in\\mathbb\{R\}^\{r\\times n\}wherer≪min\(m,n\)r\\ll\\min\(m,n\), reducing trainable parameters by orders of magnitude\. FedPETuning\[[71](https://arxiv.org/html/2607.29071#bib.bib22)\]systematically studied PEFT methods in federated settings, and FedIT\[[68](https://arxiv.org/html/2607.29071#bib.bib21)\]introduced federated instruction tuning for FMs\. FlexLoRA\[[3](https://arxiv.org/html/2607.29071#bib.bib23)\]enables dynamic rank adjustment with SVD\-based weight redistribution, HetLoRA\[[9](https://arxiv.org/html/2607.29071#bib.bib24)\]proposes zero\-padding aggregation with sparsity\-weighted combination, and FLoRA\[[56](https://arxiv.org/html/2607.29071#bib.bib25)\]introduces stacking\-based aggregation to eliminate bias\. FFA\-LoRA\[[48](https://arxiv.org/html/2607.29071#bib.bib26)\]identifies LoRA’s vulnerabilities under differential privacy and proposes freezing one factor matrix\. Industrial frameworks including FATE\-LLM\[[13](https://arxiv.org/html/2607.29071#bib.bib27)\], FederatedScope\-LLM\[[28](https://arxiv.org/html/2607.29071#bib.bib28)\], and OpenFedLLM\[[63](https://arxiv.org/html/2607.29071#bib.bib29)\]provide comprehensive infrastructure for federated FM fine\-tuning\. However, PEFT methods only reduce the trainable parameter count; every client must still load the complete base model for forward computation, so the deployment footprint remains that of the full model\.
### II\-DSVD\-Based Model Compression
Deploying FMs on heterogeneous clients necessitates aggressive compression\. Singular Value Decomposition \(SVD\) constitutes a mathematically principled approach by decomposing𝐖∈ℝm×n\\mathbf\{W\}\\in\\mathbb\{R\}^\{m\\times n\}into low\-rank factors𝐔∈ℝm×r\\mathbf\{U\}\\in\\mathbb\{R\}^\{m\\times r\}and𝐕∈ℝr×n\\mathbf\{V\}\\in\\mathbb\{R\}^\{r\\times n\}, reducing parameters frommnmntor\(m\+n\)r\(m\+n\)\. ASVD\[[65](https://arxiv.org/html/2607.29071#bib.bib30)\]introduced activation\-aware decomposition to preserve perceptually salient weight components\. SVD\-LLM\[[54](https://arxiv.org/html/2607.29071#bib.bib31)\]proposed truncation\-aware data whitening, SVD\-LLM V2\[[53](https://arxiv.org/html/2607.29071#bib.bib53)\]extended this with adaptive layer\-wise truncation, and Dobi\-SVD\[[52](https://arxiv.org/html/2607.29071#bib.bib32)\]reformulated truncation position selection as a differentiable optimization problem\. QSVD\[[55](https://arxiv.org/html/2607.29071#bib.bib54)\]demonstrated effective combination of SVD with low\-precision quantization\. Compression alone, however, does not solve the federated aggregation problem\. Clients at different compression ratios produce weight matrices of incompatible dimensions, and compression artifacts can accumulate during aggregation and propagate to the server through the*imitation fallacy*\[[6](https://arxiv.org/html/2607.29071#bib.bib19)\]\.
## IIIProposed Method
We use boldface uppercase letters \(𝐖\\mathbf\{W\},𝐔\\mathbf\{U\},𝐕\\mathbf\{V\},𝚺\\boldsymbol\{\\Sigma\},𝐒\\mathbf\{S\}\) for matrices, lowercase and Greek letters \(rr,mm,nn,η\\eta,α\\alpha\) for scalars, and calligraphic letters \(𝒮\\mathcal\{S\},𝒞i\\mathcal\{C\}\_\{i\},𝒟c\\mathcal\{D\}\_\{c\},𝒴\\mathcal\{Y\}\) for sets or spaces\. The symbol≜\\triangleqmeans “defined as,”𝐀†\\mathbf\{A\}^\{\\dagger\}denotes the Moore–Penrose pseudoinverse, and∥⋅∥F\\\|\\cdot\\\|\_\{F\},∥⋅∥2\\\|\\cdot\\\|\_\{2\},∥⋅∥1\\\|\\cdot\\\|\_\{1\}denote the Frobenius, spectral, and entry\-wiseℓ1\\ell\_\{1\}norms, respectively\. Table[I](https://arxiv.org/html/2607.29071#S3.T1)lists the key symbols\.
TABLE I:Summary of key notation\.We formulate FedSLM as a constrained optimization problem over three coupled variables: client\-side adapter parameters, server\-side reconstructed global weights, and the final strong server model\. The key idea is to restrict heterogeneous client models to a shared low\-rank feasible set defined by nested SVD subspaces, then solve the resulting problem through alternating local optimization, structural alignment via server\-side aggregation, and weak\-to\-strong refinement\.
### III\-AProblem Formulation
LetℳLM\\mathcal\{M\}\_\{LM\}denote the pre\-trained server\-side large model with layer\-wise weights\{𝐖\(l\)\}l=1L\\\{\\mathbf\{W\}^\{\(l\)\}\\\}\_\{l=1\}^\{L\}, and let𝒮⊆\[L\]\\mathcal\{S\}\\subseteq\[L\]denote the subset of layers selected for low\-rank decomposition\. For each compression ratiorir\_\{i\}and each selected layerl∈𝒮l\\in\\mathcal\{S\}, the server applies an SVD\-based compression operator to obtain a family of structurally compatible small models,
𝐖\(l\)≈𝐔i\(l\)𝐕i\(l\),𝐔i\(l\)∈ℝm×ri,𝐕i\(l\)∈ℝri×n,∀l∈𝒮,\\hskip\-8\.00003pt\\mathbf\{W\}^\{\(l\)\}\\approx\\mathbf\{U\}\_\{i\}^\{\(l\)\}\\mathbf\{V\}\_\{i\}^\{\(l\)\},\\;\\mathbf\{U\}\_\{i\}^\{\(l\)\}\\in\\mathbb\{R\}^\{m\\times r\_\{i\}\},\\;\\mathbf\{V\}\_\{i\}^\{\(l\)\}\\in\\mathbb\{R\}^\{r\_\{i\}\\times n\},\\;\\forall l\\in\\mathcal\{S\},\(1\)whereri≪min\(m,n\)r\_\{i\}\\ll\\min\(m,n\)\. For a clientc∈𝒞ic\\in\\mathcal\{C\}\_\{i\}, we insert a trainable adapter𝚺i,c\(l\)∈ℝri×ri\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(l\)\}\\in\\mathbb\{R\}^\{r\_\{i\}\\times r\_\{i\}\}only at layers in𝒮\\mathcal\{S\}, so that the effective local layer is
𝐖~i,c\(l\)=𝐔i\(l\)𝚺i,c\(l\)𝐕i\(l\),∀l∈𝒮\.\\widetilde\{\\mathbf\{W\}\}\_\{i,c\}^\{\(l\)\}=\\mathbf\{U\}\_\{i\}^\{\(l\)\}\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(l\)\}\\mathbf\{V\}\_\{i\}^\{\(l\)\},\\qquad\\forall l\\in\\mathcal\{S\}\.\(2\)Accordingly, each client is constrained to optimize only adapter parameters on the selected layers, while the low\-rank bases\{𝐔i\(l\),𝐕i\(l\)\}l∈𝒮\\\{\\mathbf\{U\}\_\{i\}^\{\(l\)\},\\mathbf\{V\}\_\{i\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}remain fixed\. LetFc\(\{𝚺i,c\(l\)\}l∈𝒮;𝒟c\)F\_\{c\}\(\\\{\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\};\\mathcal\{D\}\_\{c\}\)denote the empirical loss on client dataset𝒟c\\mathcal\{D\}\_\{c\}with sizencn\_\{c\}\. The client\-side optimization problem for group𝒞i\\mathcal\{C\}\_\{i\}is
min\{𝚺i,c\(l\)\}c∈𝒞i,l∈𝒮∑c∈𝒞incniFc\(\{𝚺i,c\(l\)\}l∈𝒮;𝒟c\),\\begin\{split\}\\min\_\{\\\{\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(l\)\}\\\}\_\{c\\in\\mathcal\{C\}\_\{i\},\\,l\\in\\mathcal\{S\}\}\}&\\sum\_\{c\\in\\mathcal\{C\}\_\{i\}\}\\frac\{n\_\{c\}\}\{n\_\{i\}\}F\_\{c\}\(\\\{\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\};\\mathcal\{D\}\_\{c\}\),\\end\{split\}\(3\)whereni=∑c∈𝒞incn\_\{i\}=\\sum\_\{c\\in\\mathcal\{C\}\_\{i\}\}n\_\{c\}\. After intra\-group local updates, the server reconstructs one full\-parameter model for each compression group and solves a consensus fusion problem in the common full\-rank space over the selected layers:
min\{𝐖g\(l\)\}l∈𝒮∑i=1KniN∑l∈𝒮‖𝐖g\(l\)−𝐖i\(l\)‖F2s\.t\.𝐖i\(l\)=𝐔i\(l\)𝚺i\(T1,l\)𝐕i\(l\),∀i∈\[K\],∀l∈𝒮,\\begin\{split\}\\min\_\{\\\{\\mathbf\{W\}\_\{\\mathrm\{g\}\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}\}&\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\\sum\_\{l\\in\\mathcal\{S\}\}\\left\\\|\\mathbf\{W\}\_\{\\mathrm\{g\}\}^\{\(l\)\}\-\\mathbf\{W\}\_\{i\}^\{\(l\)\}\\right\\\|\_\{F\}^\{2\}\\\\ \\text\{s\.t\.\}\\quad&\\mathbf\{W\}\_\{i\}^\{\(l\)\}=\\mathbf\{U\}\_\{i\}^\{\(l\)\}\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\},l\)\}\\mathbf\{V\}\_\{i\}^\{\(l\)\},\\;\\forall i\\in\[K\],\\forall l\\in\\mathcal\{S\},\\end\{split\}\(4\)where𝚺i\(T1,l\)\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\},l\)\}is the aggregated adapter of group𝒞i\\mathcal\{C\}\_\{i\}andN=∑i=1KniN=\\sum\_\{i=1\}^\{K\}n\_\{i\}\. Although each𝐖i\(l\)\\mathbf\{W\}\_\{i\}^\{\(l\)\}has rank at mostrir\_\{i\}, we do not treat them as points on different low\-rank manifolds: after reconstruction they are all elements of the same Euclidean spaceℝm×n\\mathbb\{R\}^\{m\\times n\}, all fine\-tuned from the same pre\-trained anchor𝐖\(l\)\\mathbf\{W\}^\{\(l\)\}, and thus differ only by small ambient perturbations\. In this common frame, each𝐖i\(l\)\\mathbf\{W\}\_\{i\}^\{\(l\)\}is a data\-noisy estimate of a shared consensus layer, and the closed\-form minimizer𝐖g\(l\)=∑i\(ni/N\)𝐖i\(l\)\\mathbf\{W\}\_\{\\mathrm\{g\}\}^\{\(l\)\}=\\sum\_\{i\}\(n\_\{i\}/N\)\\mathbf\{W\}\_\{i\}^\{\(l\)\}is simply the weighted barycenter, recovering FedAvg semantics in a rank\-heterogeneous setting\. Finally, let𝐖g\\mathbf\{W\}\_\{\\mathrm\{g\}\}denote the full weak model obtained by replacing the selected layers with\{𝐖g\(l\)\}l∈𝒮\\\{\\mathbf\{W\}\_\{\\mathrm\{g\}\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}and keeping the unselected layers\{𝐖\(l\)\}l∉𝒮\\\{\\mathbf\{W\}^\{\(l\)\}\\\}\_\{l\\notin\\mathcal\{S\}\}unchanged\. Inspired by the weak\-to\-strong generalization framework\[[6](https://arxiv.org/html/2607.29071#bib.bib19)\], we use this aggregated weak model as a supervisor to elicit the full capacity of the server FM\. Specifically, the server refines the full FM by solving
min𝐖s𝔼x∼𝒟server\[\(1−α\)ℓCE\(fw\(x;𝐖g\),fs\(x;𝐖s\)\)\+αℓCE\(f^s\(x;𝐖s\),fs\(x;𝐖s\)\)\],\\begin\{split\}\\min\_\{\\mathbf\{W\}\_\{\\mathrm\{s\}\}\}\\;\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\_\{\\text\{server\}\}\}\\Big\[&\(1\\\!\-\\\!\\alpha\)\\,\\ell\_\{\\mathrm\{CE\}\}\\\!\\big\(f\_\{\\text\{w\}\}\(x;\\mathbf\{W\}\_\{\\text\{g\}\}\),\\,f\_\{\\mathrm\{s\}\}\(x;\\mathbf\{W\}\_\{\\mathrm\{s\}\}\)\\big\)\\\\ \+\\;&\\alpha\\,\\ell\_\{\\mathrm\{CE\}\}\\\!\\big\(\\hat\{f\}\_\{\\mathrm\{s\}\}\(x;\\mathbf\{W\}\_\{\\mathrm\{s\}\}\),\\,f\_\{\\mathrm\{s\}\}\(x;\\mathbf\{W\}\_\{\\mathrm\{s\}\}\)\\big\)\\Big\],\\end\{split\}\(5\)whereℓCE\\ell\_\{\\mathrm\{CE\}\}is the cross\-entropy loss,𝐖s\\mathbf\{W\}\_\{\\mathrm\{s\}\}denotes the strong server model weights,𝐖g\\mathbf\{W\}\_\{\\mathrm\{g\}\}is the reconstructed global weak model,fwf\_\{\\text\{w\}\}andfsf\_\{\\mathrm\{s\}\}are the forward functions of the weak and strong models respectively,f^s\\hat\{f\}\_\{\\mathrm\{s\}\}is the strong model’s output with stopped gradients, andα∈\[0,1\]\\alpha\\in\[0,1\]balances imitation of the weak aggregated model against the strong model’s own confident predictions\. Overall, FedSLM solves Eqs\. \([3](https://arxiv.org/html/2607.29071#S3.E3)\)–\([5](https://arxiv.org/html/2607.29071#S3.E5)\) by decomposing the coupled objective into three subproblems: low\-rank model derivation, heterogeneous federated aggregation, and server\-side weak\-to\-strong refinement\.
Threat Model\.We follow the standard honest\-but\-curious assumption of FL\[[37](https://arxiv.org/html/2607.29071#bib.bib41)\]: server and clients faithfully execute the protocol but may attempt to infer information from exchanged messages, and only the lightweight adapters\{𝚺i,c\(l\)\}l∈𝒮\\\{\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}are transmitted\. Raw data, gradients, and the fixed subspace factors\{𝐔i\(l\),𝐕i\(l\)\}\\\{\\mathbf\{U\}\_\{i\}^\{\(l\)\},\\mathbf\{V\}\_\{i\}^\{\(l\)\}\\\}never leave their originating party, so orthogonal privacy\-preserving primitives such as secure aggregation or differential privacy can be layered on adapter exchanges without altering the algorithm\. The complete optimization procedure is summarized in Algorithm[1](https://arxiv.org/html/2607.29071#algorithm1), and Figure[1](https://arxiv.org/html/2607.29071#S3.F1)illustrates the overall framework\.
Input:Weights
\{𝐖\(l\)\}l=1L\\\{\\mathbf\{W\}^\{\(l\)\}\\\}\_\{l=1\}^\{L\}, layers
𝒮\\mathcal\{S\}, ratios
\{ri\}i=1K\\\{r\_\{i\}\\\}\_\{i=1\}^\{K\}, groups
\{𝒞i\}i=1K\\\{\\mathcal\{C\}\_\{i\}\\\}\_\{i=1\}^\{K\}, client data
𝒟c\\mathcal\{D\}\_\{\\mathrm\{c\}\}, rounds
T1,T2T\_\{1\},T\_\{2\}, weight
α\\alpha
Output:Strong model
𝐖s\\mathbf\{W\}\_\{\\mathrm\{s\}\}, Weak supervisor
𝐖g\\mathbf\{W\}\_\{\\mathrm\{g\}\}, Adapters
\{𝚺i\(l\)\}i,l\\\{\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(l\)\}\\\}\_\{i,l\}, Compressed models:
\{𝐔i\(l\),𝐕i\(l\),𝚺i\(l\)\}l∈𝒮\\\{\\mathbf\{U\}\_\{i\}^\{\(l\)\},\\mathbf\{V\}\_\{i\}^\{\(l\)\},\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}
1
2Step 1: SVD Compression & Adapter Initialization
3\[2pt\]for*i=1,…,Ki=1,\\ldots,K*do
4foreach*l∈𝒮l\\in\\mathcal\{S\}*do
5
𝐔i\(l\),𝐕i\(l\)←SVD\-Compress\(𝐖\(l\),ri\)\\mathbf\{U\}\_\{i\}^\{\(l\)\},\\mathbf\{V\}\_\{i\}^\{\(l\)\}\\leftarrow\\texttt\{SVD\-Compress\}\(\\mathbf\{W\}^\{\(l\)\},r\_\{i\}\)
6Initialize
𝚺i\(l\)∈ℝri×ri\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(l\)\}\\in\\mathbb\{R\}^\{r\_\{i\}\\times r\_\{i\}\}
7
8Send compressed model to
𝒞i\\mathcal\{C\}\_\{i\}
9
10Step 2: Heterogeneous Clients Local Fine\-Tuning
11\[2pt\]for*t=1,…,T1t=1,\\ldots,T\_\{1\}*do
12for*i=1,…,Ki=1,\\ldots,K*do
13foreach*c∈𝒞ic\\in\\mathcal\{C\}\_\{i\}*do
14Local Optimization:
\{𝚺i,c\(t,l\)\}l←LocalTrain\(𝐔i,𝐕i,𝚺i\(t−1\),𝒟c\)\\\{\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,l\)\}\\\}\_\{l\}\\leftarrow\\texttt\{LocalTrain\}\(\\mathbf\{U\}\_\{i\},\\mathbf\{V\}\_\{i\},\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\-1\)\},\\mathcal\{D\}\_\{c\}\)
15
16Step 3: Two\-Stage Aggregation
17\[2pt\]Stage 1: Intra\-Group Adapter Aggregation:
18
𝚺i\(t,l\)←∑c∈𝒞incni𝚺i,c\(t,l\),∀l∈𝒮\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t,l\)\}\\leftarrow\\sum\_\{c\\in\\mathcal\{C\}\_\{i\}\}\\frac\{n\_\{c\}\}\{n\_\{i\}\}\\,\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,l\)\},\\;\\;\\forall\\,l\\in\\mathcal\{S\}
19Stage 2: Cross\-Group Model Fusion:
𝐖i\(l\)←𝐔i\(l\)𝚺i\(t,l\)𝐕i\(l\),∀l∈𝒮\\mathbf\{W\}\_\{i\}^\{\(l\)\}\\leftarrow\\mathbf\{U\}\_\{i\}^\{\(l\)\}\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t,l\)\}\\mathbf\{V\}\_\{i\}^\{\(l\)\},\\;\\;\\forall\\,l\\in\\mathcal\{S\}
20
21
𝐖g\(l\)←∑i=1KniN𝐖i\(l\),∀l∈𝒮\\mathbf\{W\}\_\{\\mathrm\{g\}\}^\{\(l\)\}\\leftarrow\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\\mathbf\{W\}\_\{i\}^\{\(l\)\},\\;\\;\\forall\\,l\\in\\mathcal\{S\}
22
23Step 4: Weak\-to\-Strong Elicitation
24\[2pt\] Weak Supervisor:
𝐖g\\mathbf\{W\}\_\{\\mathrm\{g\}\}, Strong Student:
𝐖s\\mathbf\{W\}\_\{\\mathrm\{s\}\}
25for*t=1,…,T2t=1,\\ldots,T\_\{2\}*do
26
ℒ←\(1−α\)ℓCE\(fw,fs\)\+αℓCE\(f^s,fs\)\\mathcal\{L\}\\leftarrow\(1\\\!\-\\\!\\alpha\)\\,\\ell\_\{\\mathrm\{CE\}\}\(f\_\{\\mathrm\{w\}\},\\,f\_\{\\mathrm\{s\}\}\)\+\\alpha\\,\\ell\_\{\\mathrm\{CE\}\}\(\\hat\{f\}\_\{\\mathrm\{s\}\},\\,f\_\{\\mathrm\{s\}\}\)
27Update
𝐖s\\mathbf\{W\}\_\{\\mathrm\{s\}\}via
∇𝐖sℒ\\nabla\_\{\\mathbf\{W\}\_\{\\mathrm\{s\}\}\}\\mathcal\{L\}
28
return
𝐖s\\mathbf\{W\}\_\{\\mathrm\{s\}\},
𝐖g\\mathbf\{W\}\_\{\\mathrm\{g\}\},
\{𝚺i\(T1,l\)\}i,l\\\{\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\},l\)\}\\\}\_\{i,l\}
Algorithm 1FedSLMFigure 1:Overview of the FedSLM framework\. Step 1: the server derives a family of structurally compatible compressed models at heterogeneous compression ratios via SVD\-based decomposition and distributes the fixed low\-rank factors\(𝐔i,𝐕i\)\(\\mathbf\{U\}\_\{i\},\\mathbf\{V\}\_\{i\}\)along with trainable adapters𝚺i\\boldsymbol\{\\Sigma\}\_\{i\}to each client group\. Step 2 and Step 3 \(Stage 1\): clients optimize adapters locally, and the server aggregates them within each group\. Step 2 \(Stage 2\): the server reconstructs full\-rank models from each group and fuses them into a unified global model via weighted averaging in the common ambient space\. Step 4: the server refines the full\-scale model through weak\-to\-strong elicitation, using the aggregated model as a supervisor and an auxiliary confidence loss to mitigate compression artifacts\.
### III\-BLow\-Rank Model Derivation
The first subproblem concerns the construction of a feasible model family that is simultaneously compatible with heterogeneous client resources and structurally aligned for downstream aggregation\. A naive alternative would be to randomly initialize a collection of small models and train them from scratch\. Such a strategy is fundamentally ill\-suited to our setting, as resource\-constrained clients do not have access to either the large\-scale corpora or the computational budget required to recover the knowledge already embedded in the pre\-trained server large model \(LM\)\.
We therefore solve this subproblem by deriving client models directly from the server LM through SVD\-based compression on a selected layer subset𝒮\\mathcal\{S\}\. Crucially, SVD operates as a*basis transformation*that compresses the information rank of each layer without altering its dimensional interface\. Concretely, for each selected layerl∈𝒮l\\in\\mathcal\{S\}, we construct low\-rank factors𝐔i\(l\)\\mathbf\{U\}\_\{i\}^\{\(l\)\}and𝐕i\(l\)\\mathbf\{V\}\_\{i\}^\{\(l\)\}at compression ratiorir\_\{i\}and represent each client\-specific effective layer:
𝐖\(l\)≈𝐔i\(l\)𝐕i\(l\),𝐖~i,c\(l\)=𝐔i\(l\)𝚺i,c\(l\)𝐕i\(l\),∀c∈𝒞i,∀l∈𝒮\.\\begin\{split\}\\mathbf\{W\}^\{\(l\)\}&\\approx\\mathbf\{U\}\_\{i\}^\{\(l\)\}\\mathbf\{V\}\_\{i\}^\{\(l\)\},\\\\ \\widetilde\{\\mathbf\{W\}\}\_\{i,c\}^\{\(l\)\}&=\\mathbf\{U\}\_\{i\}^\{\(l\)\}\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(l\)\}\\mathbf\{V\}\_\{i\}^\{\(l\)\},\\quad\\forall c\\in\\mathcal\{C\}\_\{i\},\\forall l\\in\\mathcal\{S\}\.\\end\{split\}\(6\)This factorization preserves the ambient dimensional structure of each selected layer while reducing its parameterization frommnmntori\(m\+n\)r\_\{i\}\(m\+n\)\. More importantly, it defines the nested low\-rank subspace structure that appears explicitly in Eq\. \([3](https://arxiv.org/html/2607.29071#S3.E3)\) and constrains all subsequent client updates to a structurally compatible form\. A direct truncated SVD, however, often introduces substantial approximation error and leads to severe performance degradation\. We therefore instantiate this step with a suitable SVD\-based compression method selected from the family of techniques\[[65](https://arxiv.org/html/2607.29071#bib.bib30),[55](https://arxiv.org/html/2607.29071#bib.bib54),[52](https://arxiv.org/html/2607.29071#bib.bib32),[53](https://arxiv.org/html/2607.29071#bib.bib53),[54](https://arxiv.org/html/2607.29071#bib.bib31)\]discussed in Section[II](https://arxiv.org/html/2607.29071#S2)\. The resulting factors\{𝐔i\(l\),𝐕i\(l\)\}l∈𝒮\\\{\\mathbf\{U\}\_\{i\}^\{\(l\)\},\\mathbf\{V\}\_\{i\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}are frozen throughout federated optimization and serve solely as the fixed subspace basis\. Client\-specific task adaptation is absorbed entirely by the adapters\{𝚺i,c\(l\)\}l∈𝒮\\\{\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}, each initialized to the identity matrix𝐈ri\\mathbf\{I\}\_\{r\_\{i\}\}so that the initial reconstruction𝐔i\(l\)𝚺i,c\(l\)𝐕i\(l\)\\mathbf\{U\}\_\{i\}^\{\(l\)\}\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(l\)\}\\mathbf\{V\}\_\{i\}^\{\(l\)\}coincides with the compressed model itself\. During training, only𝚺\\boldsymbol\{\\Sigma\}is updated;𝐔\\mathbf\{U\}and𝐕\\mathbf\{V\}remain fixed\.
This design is particularly important under system heterogeneity\. In practice, clients exhibit markedly different memory and compute budgets, so a single model size is generally suboptimal\. We therefore instantiate a family of compressed models with heterogeneous compression ratios\{r1,r2,⋯,rk\}\\\{r\_\{1\},r\_\{2\},\\cdots,r\_\{k\}\\\}on the shared selected layer set𝒮\\mathcal\{S\}, where eachrir\_\{i\}corresponds to a distinct feasible set indexed by𝒞i\\mathcal\{C\}\_\{i\}\. Because all variants are derived from the same server model via nested SVD truncation, they inherit structurally compatible low\-rank subspaces, which is precisely the property needed for the reconstruction constraints in Eq\. \([4](https://arxiv.org/html/2607.29071#S3.E4)\) to remain meaningful across heterogeneous groups\.
### III\-CHeterogeneous Federated Aggregation
Having specified the feasible model family, the second subproblem is to optimize client\-specific adapters and fuse the resulting updates into a single global model\. Although the decomposed models are sufficiently compact for inference, full\-parameter fine\-tuning remains infeasible on resource\-constrained clients because backpropagation requires storing activations and gradients in addition to model weights\. In practice, training memory is often several times larger than inference memory, making direct end\-to\-end optimization of the compressed model prohibitive\. Therefore, we restrict local optimization to the lightweight adapters\{𝚺i,c\(l\)\}l∈𝒮\\\{\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}inserted between the fixed factors\{𝐔i\(l\),𝐕i\(l\)\}l∈𝒮\\\{\\mathbf\{U\}\_\{i\}^\{\(l\)\},\\mathbf\{V\}\_\{i\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}\. Under this parameterization, each client in group𝒞i\\mathcal\{C\}\_\{i\}solves Eq\. \([3](https://arxiv.org/html/2607.29071#S3.E3)\) on the selected layers, while the server subsequently solves Eq\. \([4](https://arxiv.org/html/2607.29071#S3.E4)\) after mapping all heterogeneous updates back to the common full\-rank space\. This subproblem is solved in two stages:
Stage 1: Intra\-Group Adapter Optimization\.For each compression group𝒞i\\mathcal\{C\}\_\{i\}, clients independently optimize the variables appearing in Eq\. \([3](https://arxiv.org/html/2607.29071#S3.E3)\), namely\{𝚺i,c\(l\)\}l∈𝒮\\\{\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}, while keeping\{𝐔i\(l\),𝐕i\(l\)\}l∈𝒮\\\{\\mathbf\{U\}\_\{i\}^\{\(l\)\},\\mathbf\{V\}\_\{i\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}fixed\. Because all models inside a group share the same rankrir\_\{i\}, the server can aggregate the resulting adapter updates on the selected layers via standard FedAvg\[[37](https://arxiv.org/html/2607.29071#bib.bib41)\]:
𝚺i\(t,l\)=∑c∈𝒞incni𝚺i,c\(t,l\),∀l∈𝒮\.\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t,l\)\}=\\sum\_\{c\\in\\mathcal\{C\}\_\{i\}\}\\frac\{n\_\{c\}\}\{n\_\{i\}\}\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,l\)\},\\qquad\\forall l\\in\\mathcal\{S\}\.\(7\)
This first stage is therefore the explicit solver of the local group\-level problem in Eq\. \([3](https://arxiv.org/html/2607.29071#S3.E3)\), while preserving communication efficiency because only the lightweight parameters\{𝚺i,c\(t,l\)\}l∈𝒮\\\{\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}are exchanged\.
Stage 2: Inter\-Group Reconstruction and Global Fusion\.The central difficulty in heterogeneous FL is that adapters associated with different ranks cannot be aggregated directly\. To remove this dimensional mismatch, we reconstruct the selected layers of one full\-parameter model for each compression group\. Specifically, after the final Stage\-1 round, the server integrates the aggregated adapter𝚺i\(T1,l\)\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\},l\)\}into the decomposed structure and recovers each selected layer via
𝐖i\(l\)=𝐔i\(l\)𝚺i\(T1,l\)𝐕i\(l\),∀i∈\[K\],∀l∈𝒮\.\\mathbf\{W\}\_\{i\}^\{\(l\)\}=\\mathbf\{U\}\_\{i\}^\{\(l\)\}\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\},l\)\}\\mathbf\{V\}\_\{i\}^\{\(l\)\},\\qquad\\forall i\\in\[K\],\\forall l\\in\\mathcal\{S\}\.\(8\)Note that the average is taken in ambient space, not on any single rank\-rir\_\{i\}manifold; because all groups share a nested SVD basis, the aggregated𝐖g\(l\)\\mathbf\{W\}\_\{\\mathrm\{g\}\}^\{\(l\)\}may attain rank up tomaxiri\\max\_\{i\}r\_\{i\}, an intentional enrichment beyond any individual group that carries more information for the weak\-to\-strong supervisor in Sec\.[III\-D](https://arxiv.org/html/2607.29071#S3.SS4)\. Once all groups are lifted into the same ambient space on the selected layers, the global fusion subproblem in Eq\. \([4](https://arxiv.org/html/2607.29071#S3.E4)\) admits the weighted averaging solution
𝐖g\(l\)=∑i=1KniN𝐖i\(l\),∀l∈𝒮\.\\mathbf\{W\}\_\{\\text\{g\}\}^\{\(l\)\}=\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\\mathbf\{W\}\_\{i\}^\{\(l\)\},\\qquad\\forall l\\in\\mathcal\{S\}\.\(9\)This solution is the closed\-form minimizer of the consensus objective above, so it inherits a clear statistical meaning: groups with more private data exert proportionally larger influence on the fused layer, while the shared pre\-trained initialization keeps the optimization in a locally comparable parameter region\. We then form the full weak model by replacing the selected layers of the original server LM with\{𝐖g\(l\)\}l∈𝒮\\\{\\mathbf\{W\}\_\{\\text\{g\}\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}and leaving the remaining layers unchanged\.
### III\-DWeak\-to\-Strong Refinement
After aggregation, the server obtains a weak global model𝐖g\\mathbf\{W\}\_\{\\text\{g\}\}whose selected layers encode the aggregated client knowledge and whose remaining layers are inherited directly from the original server LM\. Yet this weak model remains intrinsically constrained by the low\-rank structure imposed on the selected layers during compression\. By contrast, the original server\-side LM, now parameterized by𝐖s\\mathbf\{W\}\_\{\\text\{s\}\}, preserves substantially richer representational capacity and broader world knowledge from pretraining, but it does not directly absorb the client\-specific adaptations learned during federated optimization\. The goal is therefore to transfer the federated knowledge encoded in𝐖g\\mathbf\{W\}\_\{\\text\{g\}\}back into the full\-capacity server LM without inheriting the compression artifacts of the weak model\.
Formally, letfw\(⋅;𝐖g\)f\_\{\\text\{w\}\}\(\\cdot;\\mathbf\{W\}\_\{\\text\{g\}\}\)denote the predictive function induced by the reconstructed weak model𝐖g\\mathbf\{W\}\_\{\\text\{g\}\}, and letfs\(⋅;𝐖s\)f\_\{\\text\{s\}\}\(\\cdot;\\mathbf\{W\}\_\{\\text\{s\}\}\)denote the predictive function of the server\-resident LM\. Given an unlabeled server\-side dataset𝒟server\\mathcal\{D\}\_\{\\text\{server\}\}, the server solves Eq\. \([5](https://arxiv.org/html/2607.29071#S3.E5)\) to refine the strong model under weak supervision\.
A naive objective that merely matches the weak model’s soft labels is inadequate, because it would force the strong LM to imitate the systematic approximation errors induced by low\-rank truncation\. As noted by Burns et al\.\[[6](https://arxiv.org/html/2607.29071#bib.bib19)\], this leads to the*imitation fallacy*: a strong model may regress toward the weak supervisor instead of exploiting its own latent reasoning capabilities\. To overcome this difficulty, we adopt theAuxiliary Confidence Loss\(ℒconf\\mathcal\{L\}\_\{\\text\{conf\}\}\), written in notation as
ℒconf\(𝐖s\)=\(1−α\)⋅ℓCE\(fw\(x;𝐖g\),fs\(x;𝐖s\)\)\+α⋅ℓCE\(f^s\(x;𝐖s\),fs\(x;𝐖s\)\)\\begin\{split\}\\mathcal\{L\}\_\{\\text\{conf\}\}\(\\mathbf\{W\}\_\{\\mathrm\{s\}\}\)=\\;&\(1\\\!\-\\\!\\alpha\)\\cdot\\ell\_\{\\mathrm\{CE\}\}\\\!\\big\(f\_\{\\text\{w\}\}\(x;\\mathbf\{W\}\_\{\\text\{g\}\}\),\\,f\_\{\\mathrm\{s\}\}\(x;\\mathbf\{W\}\_\{\\mathrm\{s\}\}\)\\big\)\\\\ \+\\;&\\alpha\\cdot\\ell\_\{\\mathrm\{CE\}\}\\\!\\big\(\\hat\{f\}\_\{\\mathrm\{s\}\}\(x;\\mathbf\{W\}\_\{\\mathrm\{s\}\}\),\\,f\_\{\\mathrm\{s\}\}\(x;\\mathbf\{W\}\_\{\\mathrm\{s\}\}\)\\big\)\\end\{split\}\(10\)
whereℓCE\(p,q\)\\ell\_\{\\mathrm\{CE\}\}\(p,q\)denotes the cross\-entropy loss between a target distributionppand a predictive distributionqq\. We define the hardened self\-prediction of the strong model as
f^s\(x;𝐖s\)≜𝕀\[fs\(x;𝐖s\)\>τ\],\\hat\{f\}\_\{\\mathrm\{s\}\}\(x;\\mathbf\{W\}\_\{\\mathrm\{s\}\}\)\\triangleq\\mathbb\{I\}\[f\_\{\\mathrm\{s\}\}\(x;\\mathbf\{W\}\_\{\\mathrm\{s\}\}\)\>\\tau\],\(11\)where𝕀\\mathbb\{I\}is the indicator function and the thresholdτ\\tauis set adaptively within each batch such thatfs\(x;𝐖s\)\>τf\_\{\\mathrm\{s\}\}\(x;\\mathbf\{W\}\_\{\\mathrm\{s\}\}\)\>\\tauholds for exactly half of the samples\. The hyperparameterα∈\[0,1\]\\alpha\\in\[0,1\]controls the trade\-off between inheriting domain knowledge and trusting the strong model’s own confident predictions\.
In effect, this refinement stage serves as a capacity\-recovery step: it transfers the domain specialization learned through federated optimization while allowing the full\-scale LM to mitigate low\-rank artifacts through the self\-anchoring term and re\-express that knowledge in a richer hypothesis class\. The output is therefore not a mere expansion of the compressed model, but a refined strong model that combines specialization with the generalization capacity of the server LM\.
## IVTheoretical Analysis
We provide theoretical justifications for FedSLM from three complementary perspectives: \(i\) a*convergence analysis*that establishes the𝒪\(T−1/2\)\\mathcal\{O\}\(T^\{\-1/2\}\)rate of the Stage\-1 adapter\-level aggregation \(Section[IV\-A](https://arxiv.org/html/2607.29071#S4.SS1)\); \(ii\) a*subspace\-alignment theory*that characterizes the geometric compatibility of heterogeneous compression groups and bounds the cross\-group aggregation error \(Section[IV\-B](https://arxiv.org/html/2607.29071#S4.SS2)\); and \(iii\) an*artifact\-propagation analysis*that quantifies how the auxiliary confidence loss attenuates compression noise during weak\-to\-strong knowledge elicitation \(Section[IV\-C](https://arxiv.org/html/2607.29071#S4.SS3)\)\. All proofs and supporting lemmas are deferred to Appendix[A](https://arxiv.org/html/2607.29071#A1)\.
We fix a single selected layer𝐖∈ℝm×n\\mathbf\{W\}\\in\\mathbb\{R\}^\{m\\times n\}; the analysis extends to the set𝒮\\mathcal\{S\}of selected layers by summation\. Let the singular values of𝐖\\mathbf\{W\}beλ1≥λ2≥⋯≥λmin\(m,n\)≥0\\lambda\_\{1\}\\geq\\lambda\_\{2\}\\geq\\cdots\\geq\\lambda\_\{\\min\(m,n\)\}\\geq 0, with associated left singular vectors𝐮1,…,𝐮min\(m,n\)∈ℝm\\mathbf\{u\}\_\{1\},\\ldots,\\mathbf\{u\}\_\{\\min\(m,n\)\}\\in\\mathbb\{R\}^\{m\}and right singular vectors𝐯1,…,𝐯min\(m,n\)∈ℝn\\mathbf\{v\}\_\{1\},\\ldots,\\mathbf\{v\}\_\{\\min\(m,n\)\}\\in\\mathbb\{R\}^\{n\}\. For a rankrr, the server stores the low\-rank factorization
𝐖≈𝐔r𝐕r,𝐔r∈ℝm×r,𝐕r∈ℝr×n,\\mathbf\{W\}\\;\\approx\\;\\mathbf\{U\}\_\{r\}\\,\\mathbf\{V\}\_\{r\},\\qquad\\mathbf\{U\}\_\{r\}\\in\\mathbb\{R\}^\{m\\times r\},\\;\\mathbf\{V\}\_\{r\}\\in\\mathbb\{R\}^\{r\\times n\},\(12\)whose factors are obtained from the rank\-rrtruncated SVD of𝐖\\mathbf\{W\}by splitting the retained spectrum𝚲r≜diag\(λ1,…,λr\)\\boldsymbol\{\\Lambda\}\_\{r\}\\triangleq\\mathrm\{diag\}\(\\lambda\_\{1\},\\ldots,\\lambda\_\{r\}\)evenly across the two sides,
𝐔r=𝐔r∘𝚲r1/2,𝐕r=𝚲r1/2𝐕r∘,\\mathbf\{U\}\_\{r\}\\;=\\;\\mathbf\{U\}\_\{r\}^\{\\circ\}\\,\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\},\\qquad\\qquad\\mathbf\{V\}\_\{r\}\\;=\\;\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\,\\mathbf\{V\}\_\{r\}^\{\\circ\},\(13\)where𝐔r∘∈ℝm×r\\mathbf\{U\}\_\{r\}^\{\\circ\}\\in\\mathbb\{R\}^\{m\\times r\}and𝐕r∘∈ℝr×n\\mathbf\{V\}\_\{r\}^\{\\circ\}\\in\\mathbb\{R\}^\{r\\times n\}collect the leadingrrleft/right singular vectors and are column\-/row\-orthonormal \(𝐔r∘⊤𝐔r∘=𝐈r\\mathbf\{U\}\_\{r\}^\{\\circ\\,\\top\}\\mathbf\{U\}\_\{r\}^\{\\circ\}=\\mathbf\{I\}\_\{r\},𝐕r∘𝐕r∘⊤=𝐈r\\mathbf\{V\}\_\{r\}^\{\\circ\}\\mathbf\{V\}\_\{r\}^\{\\circ\\,\\top\}=\\mathbf\{I\}\_\{r\}\), so that𝐔r𝐕r=𝐔r∘𝚲r𝐕r∘\\mathbf\{U\}\_\{r\}\\mathbf\{V\}\_\{r\}=\\mathbf\{U\}\_\{r\}^\{\\circ\}\\boldsymbol\{\\Lambda\}\_\{r\}\\mathbf\{V\}\_\{r\}^\{\\circ\}exactly recovers the rank\-rrtruncation\[𝐖\]r\[\\mathbf\{W\}\]\_\{r\}\. We write the adapter reconstruction at clientccof group𝒞i\\mathcal\{C\}\_\{i\}as
𝐖~i,c=𝐔ri𝚺i,c𝐕ri,𝚺i,c∈ℝri×ri,\\widetilde\{\\mathbf\{W\}\}\_\{i,c\}\\;=\\;\\mathbf\{U\}\_\{r\_\{i\}\}\\,\\boldsymbol\{\\Sigma\}\_\{i,c\}\\,\\mathbf\{V\}\_\{r\_\{i\}\},\\qquad\\boldsymbol\{\\Sigma\}\_\{i,c\}\\in\\mathbb\{R\}^\{r\_\{i\}\\times r\_\{i\}\},\(14\)where the trainable core𝚺i,c\\boldsymbol\{\\Sigma\}\_\{i,c\}\(initialized to𝐈ri\\mathbf\{I\}\_\{r\_\{i\}\}\) is the only quantity optimized by clients; the factors𝐔ri,𝐕ri\\mathbf\{U\}\_\{r\_\{i\}\},\\mathbf\{V\}\_\{r\_\{i\}\}are frozen\.
A truncated SVD picks𝐔r,𝐕r\\mathbf\{U\}\_\{r\},\\mathbf\{V\}\_\{r\}to minimize the weight reconstruction error‖𝐖−𝐔r𝐕r‖F\\\|\\mathbf\{W\}\-\\mathbf\{U\}\_\{r\}\\mathbf\{V\}\_\{r\}\\\|\_\{F\}\. For an FM layer this is the wrong objective: accuracy is governed by the*output*error‖\(𝐖−𝐔r𝐕r\)𝐗‖F\\\|\(\\mathbf\{W\}\-\\mathbf\{U\}\_\{r\}\\mathbf\{V\}\_\{r\}\)\\mathbf\{X\}\\\|\_\{F\}on representative activations𝐗\\mathbf\{X\}, and because FM activations are correlated and contain outlier channels, the smallest singular values of𝐖\\mathbf\{W\}need not correspond to the least important output directions\. Truncating them discards activation\-critical components and renders the compressed model unusable\. Activation\-aware and truncation\-aware variants\[[65](https://arxiv.org/html/2607.29071#bib.bib30),[54](https://arxiv.org/html/2607.29071#bib.bib31)\]resolve this by reweighting the decomposition with an invertible scaling𝐒∈ℝr×r\\mathbf\{S\}\\in\\mathbb\{R\}^\{r\\times r\}\(an activation\-magnitude diagonal or a whitening matrix from the activation covariance\), so that singular\-value truncation in the scaled coordinates aligns with the output error:
𝐖≈\(𝐔r∘𝚲r1/2𝐒−1\)⏟𝐔r\(𝐒𝚲r1/2𝐕r∘\)⏟𝐕r,\\mathbf\{W\}\\;\\approx\\;\\underbrace\{\(\\mathbf\{U\}\_\{r\}^\{\\circ\}\\,\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\,\\mathbf\{S\}^\{\-1\}\)\}\_\{\\mathbf\{U\}\_\{r\}\}\\underbrace\{\(\\mathbf\{S\}\\,\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\,\\mathbf\{V\}\_\{r\}^\{\\circ\}\)\}\_\{\\mathbf\{V\}\_\{r\}\},\(15\)with the plain SVD recovered at𝐒=𝐈r\\mathbf\{S\}=\\mathbf\{I\}\_\{r\}\. In either case the deployed factors equal the orthonormal singular vectors𝐔r∘,𝐕r∘\\mathbf\{U\}\_\{r\}^\{\\circ\},\\mathbf\{V\}\_\{r\}^\{\\circ\}up to the invertible scalings𝚲r1/2\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}and𝐒\\mathbf\{S\}, and are therefore non\-orthonormal; our analysis absorbs this spectral scaling into a single effective smoothness constant\.
###### Definition 1\(Effective smoothness constant\)\.
Letκ\(𝐒\)≜‖𝐒‖2‖𝐒−1‖2\\kappa\(\\mathbf\{S\}\)\\triangleq\\\|\\mathbf\{S\}\\\|\_\{2\}\\,\\\|\\mathbf\{S\}^\{\-1\}\\\|\_\{2\}denote the condition number of𝐒\\mathbf\{S\}in \([15](https://arxiv.org/html/2607.29071#S4.E15)\); setκ\(𝐒\)=1\\kappa\(\\mathbf\{S\}\)=1for the plain factorization \([13](https://arxiv.org/html/2607.29071#S4.E13)\)\. Writingλ1\\lambda\_\{1\}for the largest retained singular value, and given a task\-loss smoothness constantLL\(Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)\(a\)\), we define
Leff≜Lλ12κ\(𝐒\)2\.L\_\{\\mathrm\{eff\}\}\\;\\triangleq\\;L\\,\\lambda\_\{1\}^\{2\}\\,\\kappa\(\\mathbf\{S\}\)^\{2\}\.\(16\)
The factorλ12κ\(𝐒\)2\\lambda\_\{1\}^\{2\}\\,\\kappa\(\\mathbf\{S\}\)^\{2\}inLeffL\_\{\\mathrm\{eff\}\}is the exact price of optimizing the core𝚺\\boldsymbol\{\\Sigma\}through the frozen, non\-orthonormal factors\(𝐔r,𝐕r\)\(\\mathbf\{U\}\_\{r\},\\mathbf\{V\}\_\{r\}\)\. Sincefc\(𝚺\)=ℒc\(𝐔r𝚺𝐕r\)f\_\{c\}\(\\boldsymbol\{\\Sigma\}\)=\\mathcal\{L\}\_\{c\}\(\\mathbf\{U\}\_\{r\}\\boldsymbol\{\\Sigma\}\\mathbf\{V\}\_\{r\}\)composes theLL\-smooth loss with the linear map𝚺↦𝐔r𝚺𝐕r\\boldsymbol\{\\Sigma\}\\mapsto\\mathbf\{U\}\_\{r\}\\boldsymbol\{\\Sigma\}\\mathbf\{V\}\_\{r\}, its smoothness constant is amplified by the squared operator norm‖𝐔r‖22‖𝐕r‖22\\\|\\mathbf\{U\}\_\{r\}\\\|\_\{2\}^\{2\}\\,\\\|\\mathbf\{V\}\_\{r\}\\\|\_\{2\}^\{2\}of that map\. For the split \([13](https://arxiv.org/html/2607.29071#S4.E13)\), the leading singular valueλ1\\lambda\_\{1\}enters each factor asλ1\\sqrt\{\\lambda\_\{1\}\}, so‖𝐔r‖2=‖𝐕r‖2=λ1\\\|\\mathbf\{U\}\_\{r\}\\\|\_\{2\}=\\\|\\mathbf\{V\}\_\{r\}\\\|\_\{2\}=\\sqrt\{\\lambda\_\{1\}\}and the amplification equals‖𝐔r‖22‖𝐕r‖22=λ12\\\|\\mathbf\{U\}\_\{r\}\\\|\_\{2\}^\{2\}\\,\\\|\\mathbf\{V\}\_\{r\}\\\|\_\{2\}^\{2\}=\\lambda\_\{1\}^\{2\}; activation\-aware whitening contributes the extra factorκ\(𝐒\)2\\kappa\(\\mathbf\{S\}\)^\{2\}, which collapses to11for a plain SVD\. Crucially, bothλ1=‖𝐖‖2\\lambda\_\{1\}=\\\|\\mathbf\{W\}\\\|\_\{2\}andκ\(𝐒\)\\kappa\(\\mathbf\{S\}\)are fixed by the one\-time SVD of the pretrained weight and remain constant throughout federated training; they neither grow with the communication round nor depend on the optimization trajectory\. For pretrained transformer layers the spectral normλ1\\lambda\_\{1\}is a moderate, layer\-dependent constant \(weight matrices stay spectrally bounded under standard initialization and normalization\), and the whitening conditioningκ\(𝐒\)\\kappa\(\\mathbf\{S\}\)remains finite \(and is regularized to be well\-conditioned in practice\[[54](https://arxiv.org/html/2607.29071#bib.bib31)\]\)\.
### IV\-AConvergence Analysis
We focus on Stage\-1 adapter\-level aggregation \(Algorithm[1](https://arxiv.org/html/2607.29071#algorithm1)\), the iterative federated optimization component of FedSLM\. Stage 2 is a single\-step closed\-form weighted average whose approximation quality is characterized by the subspace\-alignment theory in Section[IV\-B](https://arxiv.org/html/2607.29071#S4.SS2)\. The non\-trivial challenge arises in Stage 1, where, within each compression group𝒞i\\mathcal\{C\}\_\{i\}, clients optimize in the reparameterized adapter space𝚺i,c∈ℝri×ri\\boldsymbol\{\\Sigma\}\_\{i,c\}\\in\\mathbb\{R\}^\{r\_\{i\}\\times r\_\{i\}\}: the factorization𝐖i≈𝐔ri𝚺i,c𝐕ri\\mathbf\{W\}\_\{i\}\\approx\\mathbf\{U\}\_\{r\_\{i\}\}\\boldsymbol\{\\Sigma\}\_\{i,c\}\\mathbf\{V\}\_\{r\_\{i\}\}distorts the loss landscape through the factor scalingλ1κ\(𝐒\)\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\), and the analysis must account for how this spectral distortion interacts with FedAvg\[[37](https://arxiv.org/html/2607.29071#bib.bib41)\]and client drift\. Without loss of generality, we fix an arbitrary compression group𝒞i\\mathcal\{C\}\_\{i\}operating on𝐖i∈ℝm×n\\mathbf\{W\}\_\{i\}\\in\\mathbb\{R\}^\{m\\times n\}\. Each clientc∈𝒞ic\\in\\mathcal\{C\}\_\{i\}holds a local dataset𝒟c\\mathcal\{D\}\_\{c\}of sizencn\_\{c\}\(with group massni=∑c∈𝒞incn\_\{i\}=\\sum\_\{c\\in\\mathcal\{C\}\_\{i\}\}n\_\{c\}\) and optimizes its local adapter𝚺i,c∈ℝri×ri\\boldsymbol\{\\Sigma\}\_\{i,c\}\\in\\mathbb\{R\}^\{r\_\{i\}\\times r\_\{i\}\}with𝐖~i,c\(𝚺i,c\)=𝐔ri𝚺i,c𝐕ri\\widetilde\{\\mathbf\{W\}\}\_\{i,c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}\)=\\mathbf\{U\}\_\{r\_\{i\}\}\\boldsymbol\{\\Sigma\}\_\{i,c\}\\mathbf\{V\}\_\{r\_\{i\}\}as in \([14](https://arxiv.org/html/2607.29071#S4.E14)\)\. The local adapter loss isfc\(𝚺\)≜ℒc\(𝐔ri𝚺𝐕ri\)f\_\{c\}\(\\boldsymbol\{\\Sigma\}\)\\triangleq\\mathcal\{L\}\_\{c\}\(\\mathbf\{U\}\_\{r\_\{i\}\}\\boldsymbol\{\\Sigma\}\\mathbf\{V\}\_\{r\_\{i\}\}\)and the group\-level adapter loss isfi\(𝚺\)≜∑c∈𝒞incnifc\(𝚺\)f\_\{i\}\(\\boldsymbol\{\\Sigma\}\)\\triangleq\\sum\_\{c\\in\\mathcal\{C\}\_\{i\}\}\\frac\{n\_\{c\}\}\{n\_\{i\}\}f\_\{c\}\(\\boldsymbol\{\\Sigma\}\)\. We require the following standard assumptions\.
###### Assumption 1\.
The following conditions hold for all clientsc∈𝒞ic\\in\\mathcal\{C\}\_\{i\}and all adapter parameters𝚺\\boldsymbol\{\\Sigma\}:
1. \(a\)LL\-Smoothness\.For every clientcc, the map𝐖↦ℒc\(𝐖\)\\mathbf\{W\}\\mapsto\\mathcal\{L\}\_\{c\}\(\\mathbf\{W\}\)is differentiable with Lipschitz\-continuous gradient:‖∇𝐖ℒc\(𝐖\)−∇𝐖ℒc\(𝐖′\)‖F≤L‖𝐖−𝐖′‖F\\bigl\\\|\\nabla\_\{\\mathbf\{W\}\}\\mathcal\{L\}\_\{c\}\(\\mathbf\{W\}\)\-\\nabla\_\{\\mathbf\{W\}\}\\mathcal\{L\}\_\{c\}\(\\mathbf\{W\}^\{\\prime\}\)\\bigr\\\|\_\{F\}\\leq L\\\|\\mathbf\{W\}\-\\mathbf\{W\}^\{\\prime\}\\\|\_\{F\}for all𝐖,𝐖′\\mathbf\{W\},\\mathbf\{W\}^\{\\prime\}\.
2. \(b\)Unbiased Stochastic Gradients with Bounded Variance\.𝔼\[𝐠~c\(𝚺\)∣𝚺\]=∇𝚺fc\(𝚺\)\\mathbb\{E\}\[\\widetilde\{\\mathbf\{g\}\}\_\{c\}\(\\boldsymbol\{\\Sigma\}\)\\mid\\boldsymbol\{\\Sigma\}\]=\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{c\}\(\\boldsymbol\{\\Sigma\}\), and there existsσ2≥0\\sigma^\{2\}\\geq 0such that𝔼‖𝐠~c\(𝚺\)−∇𝚺fc\(𝚺\)‖F2≤σ2\\mathbb\{E\}\\bigl\\\|\\widetilde\{\\mathbf\{g\}\}\_\{c\}\(\\boldsymbol\{\\Sigma\}\)\-\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{c\}\(\\boldsymbol\{\\Sigma\}\)\\bigr\\\|\_\{F\}^\{2\}\\leq\\sigma^\{2\}\.
3. \(c\)Bounded Data Heterogeneity\.There existsGi≥0G\_\{i\}\\geq 0such that∑c∈𝒞incni‖∇𝚺fc\(𝚺\)−∇𝚺fi\(𝚺\)‖F2≤Gi2\\sum\_\{c\\in\\mathcal\{C\}\_\{i\}\}\\frac\{n\_\{c\}\}\{n\_\{i\}\}\\bigl\\\|\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{c\}\(\\boldsymbol\{\\Sigma\}\)\-\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{i\}\(\\boldsymbol\{\\Sigma\}\)\\bigr\\\|\_\{F\}^\{2\}\\leq G\_\{i\}^\{2\}\.
###### Theorem 1\(Stage\-1 convergence with client drift\)\.
Let Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)hold, and letLeff=Lλ12κ\(𝐒\)2L\_\{\\mathrm\{eff\}\}=L\\,\\lambda\_\{1\}^\{2\}\\,\\kappa\(\\mathbf\{S\}\)^\{2\}be the effective smoothness constant from Definition[1](https://arxiv.org/html/2607.29071#Thmdefinition1)\. Suppose each client in𝒞i\\mathcal\{C\}\_\{i\}performsE≥1E\\geq 1local SGD steps per round with constant step sizeη\\etasatisfyingη≤18LeffE\\eta\\leq\\dfrac\{1\}\{8\\,L\_\{\\mathrm\{eff\}\}\\,E\}\. AfterT1T\_\{1\}rounds of FedAvg, let𝚺i\(t\)∈ℝri×ri\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\\in\\mathbb\{R\}^\{r\_\{i\}\\times r\_\{i\}\}denote the aggregated group adapter at the start of roundtt\(with𝚺i\(0\)=𝐈ri\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(0\)\}=\\mathbf\{I\}\_\{r\_\{i\}\}, the initialization in \([14](https://arxiv.org/html/2607.29071#S4.E14)\)\)\. The iterates\{𝚺i\(t\)\}t=0T1\\\{\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\\\}\_\{t=0\}^\{T\_\{1\}\}produced by Algorithm 1 satisfy
1T1∑t=0T1−1𝔼‖∇fi\(𝚺i\(t\)\)‖F2≤4Δ0ηET1⏟optimization\+4Leffησ2⏟stochastic\+8Leff2η2E2\(σ2\+6EGi2\)⏟client drift\+2Gi2⏟heterogeneity floor,\\begin\{split\}\\frac\{1\}\{T\_\{1\}\}\\sum\_\{t=0\}^\{T\_\{1\}\-1\}&\\mathbb\{E\}\\bigl\\\|\\nabla f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\)\\bigr\\\|\_\{F\}^\{2\}\\;\\leq\\;\\underbrace\{\\frac\{4\\Delta\_\{0\}\}\{\\eta E\\,T\_\{1\}\}\}\_\{\\text\{optimization\}\}\+\\underbrace\{4\\,L\_\{\\mathrm\{eff\}\}\\,\\eta\\,\\sigma^\{2\}\}\_\{\\text\{stochastic\}\}\\\\ &\+\\underbrace\{8\\,L\_\{\\mathrm\{eff\}\}^\{\\,2\}\\,\\eta^\{2\}E^\{2\}\\bigl\(\\sigma^\{2\}\+6E\\,G\_\{i\}^\{2\}\\bigr\)\}\_\{\\text\{client drift\}\}\+\\underbrace\{2\\,G\_\{i\}^\{2\}\}\_\{\\text\{heterogeneity floor\}\},\\end\{split\}\(17\)whereΔ0≜fi\(𝚺i\(0\)\)−inf𝚺fi\(𝚺\)\\Delta\_\{0\}\\triangleq f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(0\)\}\)\-\\inf\_\{\\boldsymbol\{\\Sigma\}\}f\_\{i\}\(\\boldsymbol\{\\Sigma\}\)\. Choosingη=1LeffT1E\\eta=\\dfrac\{1\}\{L\_\{\\mathrm\{eff\}\}\\sqrt\{T\_\{1\}E\}\}\(which meets the condition wheneverT1≥64ET\_\{1\}\\geq 64E\) yields
1T1∑t=0T1−1𝔼‖∇fi\(𝚺i\(t\)\)‖F2=𝒪\(LeffΔ0\+σ2T1E\)\+𝒪\(E\(σ2\+EGi2\)T1\)\+2Gi2\.\\begin\{split\}&\\frac\{1\}\{T\_\{1\}\}\\sum\_\{t=0\}^\{T\_\{1\}\-1\}\\mathbb\{E\}\\bigl\\\|\\nabla f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\)\\bigr\\\|\_\{F\}^\{2\}\\\\ =&\\;\\mathcal\{O\}\\\!\\left\(\\frac\{L\_\{\\mathrm\{eff\}\}\\Delta\_\{0\}\+\\sigma^\{2\}\}\{\\sqrt\{T\_\{1\}E\}\}\\right\)\+\\mathcal\{O\}\\\!\\left\(\\frac\{E\(\\sigma^\{2\}\+EG\_\{i\}^\{2\}\)\}\{T\_\{1\}\}\\right\)\+2\\,G\_\{i\}^\{2\}\.\\end\{split\}\(18\)
Theorem[1](https://arxiv.org/html/2607.29071#Thmtheorem1)shows that the adapter\-level FedAvg in FedSLM converges at the same𝒪\(T1−1/2\)\\mathcal\{O\}\(T\_\{1\}^\{\-1/2\}\)rate as standard FedAvg, up to the spectral penaltyλ12κ\(𝐒\)2\\lambda\_\{1\}^\{2\}\\,\\kappa\(\\mathbf\{S\}\)^\{2\}absorbed intoLeffL\_\{\\mathrm\{eff\}\}\. Theλ12\\lambda\_\{1\}^\{2\}factor is the price of folding the singular values into the frozen factors\(𝐔ri,𝐕ri\)\(\\mathbf\{U\}\_\{r\_\{i\}\},\\mathbf\{V\}\_\{r\_\{i\}\}\)so that𝐔ri𝐕ri\\mathbf\{U\}\_\{r\_\{i\}\}\\mathbf\{V\}\_\{r\_\{i\}\}reconstructs𝐖i\\mathbf\{W\}\_\{i\}directly; it merely rescales the natural step sizeη=1/\(LeffT1E\)\\eta=1/\(L\_\{\\mathrm\{eff\}\}\\sqrt\{T\_\{1\}E\}\)and leaves the convergence rate intact\. When the server uses a pure SVD,κ\(𝐒\)=1\\kappa\(\\mathbf\{S\}\)=1and only theλ12\\lambda\_\{1\}^\{2\}factor remains; when SVD variants are used with a well\-conditioned whitening matrix, the extra penaltyκ\(𝐒\)2\\kappa\(\\mathbf\{S\}\)^\{2\}is modest\. The client\-drift term𝒪\(E\(σ2\+EGi2\)/T1\)\\mathcal\{O\}\(E\(\\sigma^\{2\}\+EG\_\{i\}^\{2\}\)/T\_\{1\}\)is the standard artifact of local SGD\[[25](https://arxiv.org/html/2607.29071#bib.bib43)\]and shows that increasingEEdoes*not*uniformly improve the rate; the residual2Gi22G\_\{i\}^\{2\}is the irreducible heterogeneity floor of group𝒞i\\mathcal\{C\}\_\{i\}\. Since group𝒞i\\mathcal\{C\}\_\{i\}was arbitrary, the bound holds for everyi∈\[K\]i\\in\[K\], with the constantsGi,Δ0G\_\{i\},\\Delta\_\{0\}specialized to each group\.
### IV\-BSubspace Alignment Theory
A key challenge in FedSLM’s two\-stage aggregation is that different compression groups operate in subspaces of different ranks: group𝒞i\\mathcal\{C\}\_\{i\}with rankrir\_\{i\}can only represent weight matrices within a rank\-rir\_\{i\}feasible set\. This raises a natural question: when the server reconstructs full\-rank weights from heterogeneous groups and averages them in Stage 2, how much error does this cross\-group fusion introduce? We answer this in three steps\. We first establish that the SVD\-derived subspaces possess a nested structure \(Proposition[1](https://arxiv.org/html/2607.29071#Thmproposition1)\)\. We then derive a general upper bound on the cross\-group aggregation error in terms of three interpretable residuals \(Theorem[2](https://arxiv.org/html/2607.29071#Thmtheorem2)\), and show how a Polyak–Łojasiewicz inequality and a common\-projection\-target condition each control one of those residuals \(Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\)\. Finally, the closed\-form specialization is collected in Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3)\.
###### Definition 2\(Compression\-group subspaces\)\.
For a weight matrix𝐖∈ℝm×n\\mathbf\{W\}\\in\\mathbb\{R\}^\{m\\times n\}, let𝐮1,…,𝐮min\(m,n\)\\mathbf\{u\}\_\{1\},\\ldots,\\mathbf\{u\}\_\{\\min\(m,n\)\}and𝐯1,…,𝐯min\(m,n\)\\mathbf\{v\}\_\{1\},\\ldots,\\mathbf\{v\}\_\{\\min\(m,n\)\}be its left and right singular vectors, ordered by non\-increasing singular valuesλ1≥λ2≥⋯≥λmin\(m,n\)≥0\\lambda\_\{1\}\\geq\\lambda\_\{2\}\\geq\\cdots\\geq\\lambda\_\{\\min\(m,n\)\}\\geq 0\. We define the rank\-rrleft and right subspaces ofℝm\\mathbb\{R\}^\{m\}andℝn\\mathbb\{R\}^\{n\}respectively as
𝒰r≜span\{𝐮1,…,𝐮r\},𝒱r≜span\{𝐯1,…,𝐯r\}\.\\begin\{split\}\\mathcal\{U\}\_\{r\}&\\triangleq\\mathrm\{span\}\\\{\\mathbf\{u\}\_\{1\},\\ldots,\\mathbf\{u\}\_\{r\}\\\},\\\\ \\mathcal\{V\}\_\{r\}&\\triangleq\\mathrm\{span\}\\\{\\mathbf\{v\}\_\{1\},\\ldots,\\mathbf\{v\}\_\{r\}\\\}\.\\end\{split\}\(19\)The associated orthogonal projectors are𝐏r≜𝐔r∘𝐔r∘⊤∈ℝm×m\\mathbf\{P\}\_\{r\}\\triangleq\\mathbf\{U\}\_\{r\}^\{\\circ\}\\,\\mathbf\{U\}\_\{r\}^\{\\circ\\,\\top\}\\in\\mathbb\{R\}^\{m\\times m\}and𝐐r≜𝐕r∘⊤𝐕r∘∈ℝn×n\\mathbf\{Q\}\_\{r\}\\triangleq\\mathbf\{V\}\_\{r\}^\{\\circ\\,\\top\}\\,\\mathbf\{V\}\_\{r\}^\{\\circ\}\\in\\mathbb\{R\}^\{n\\times n\}, where𝐔r∘\\mathbf\{U\}\_\{r\}^\{\\circ\}and𝐕r∘\\mathbf\{V\}\_\{r\}^\{\\circ\}are the orthonormal factors from \([13](https://arxiv.org/html/2607.29071#S4.E13)\)–\([15](https://arxiv.org/html/2607.29071#S4.E15)\) \(related to the deployed𝐔r,𝐕r\\mathbf\{U\}\_\{r\},\\mathbf\{V\}\_\{r\}by the spectral scaling𝚲r1/2\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}and, for whitened variants, by𝐒\\mathbf\{S\}\)\.
Note that𝐏r\\mathbf\{P\}\_\{r\}and𝐐r\\mathbf\{Q\}\_\{r\}are defined using the*orthonormal*factors\(𝐔r∘,𝐕r∘\)\(\\mathbf\{U\}\_\{r\}^\{\\circ\},\\mathbf\{V\}\_\{r\}^\{\\circ\}\)of𝒰r\\mathcal\{U\}\_\{r\}and𝒱r\\mathcal\{V\}\_\{r\}, not the scaled factors𝐔r,𝐕r\\mathbf\{U\}\_\{r\},\\mathbf\{V\}\_\{r\}seen by clients\. This is essential because orthogonal projectors depend only on the subspace, not on the choice of basis\. The column space of𝐔r\\mathbf\{U\}\_\{r\}coincides with that of𝐔r∘\\mathbf\{U\}\_\{r\}^\{\\circ\}and the row space of𝐕r\\mathbf\{V\}\_\{r\}coincides with that of𝐕r∘\\mathbf\{V\}\_\{r\}^\{\\circ\}\(as𝐒\\mathbf\{S\}is invertible\), so𝒰r\\mathcal\{U\}\_\{r\}and𝒱r\\mathcal\{V\}\_\{r\}are well\-defined regardless of the factorization variant used\.
###### Proposition 1\(Nested subspace structure\)\.
For any two compression ratiosr1<r2r\_\{1\}<r\_\{2\}, the corresponding subspaces satisfy𝒰r1⊊𝒰r2\\mathcal\{U\}\_\{r\_\{1\}\}\\subsetneq\\mathcal\{U\}\_\{r\_\{2\}\}and𝒱r1⊊𝒱r2\\mathcal\{V\}\_\{r\_\{1\}\}\\subsetneq\\mathcal\{V\}\_\{r\_\{2\}\}, and the projectors satisfy the absorption identities
𝐏r1𝐏r2=𝐏r2𝐏r1=𝐏r1,𝐐r1𝐐r2=𝐐r2𝐐r1=𝐐r1\.\\begin\{split\}\\mathbf\{P\}\_\{r\_\{1\}\}\\mathbf\{P\}\_\{r\_\{2\}\}&=\\mathbf\{P\}\_\{r\_\{2\}\}\\mathbf\{P\}\_\{r\_\{1\}\}=\\mathbf\{P\}\_\{r\_\{1\}\},\\\\ \\mathbf\{Q\}\_\{r\_\{1\}\}\\mathbf\{Q\}\_\{r\_\{2\}\}&=\\mathbf\{Q\}\_\{r\_\{2\}\}\\mathbf\{Q\}\_\{r\_\{1\}\}=\\mathbf\{Q\}\_\{r\_\{1\}\}\.\\end\{split\}\(20\)
Proposition[1](https://arxiv.org/html/2607.29071#Thmproposition1)establishes that the representation spaces of different compression groups form a nested hierarchy: every low\-rank client’s feasible set is geometrically contained within that of any higher\-rank client\. This nesting is a direct consequence of all client factors being derived from the*same*server weight𝐖\\mathbf\{W\}via truncated SVD\. The absorption identities further imply that projecting onto a lower\-rank subspace and then onto a higher\-rank one is equivalent to the lower\-rank projection alone, which is central to the two\-stage aggregation protocol\.
We now fixKKcompression groups with ranksr1,…,rKr\_\{1\},\\ldots,r\_\{K\}and per\-group data massesn1,…,nKn\_\{1\},\\ldots,n\_\{K\}withN=∑iniN=\\sum\_\{i\}n\_\{i\}\. Let𝚺i\(T1\)∈ℝri×ri\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\}\)\}\\in\\mathbb\{R\}^\{r\_\{i\}\\times r\_\{i\}\}denote the adapter of groupiiafter Stage 1, and define the group\-specific reconstruction
𝐖i≜𝐔ri𝚺i\(T1\)𝐕ri,𝐖g≜∑i=1KniN𝐖i\.\\mathbf\{W\}\_\{i\}\\;\\triangleq\\;\\mathbf\{U\}\_\{r\_\{i\}\}\\,\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\}\)\}\\,\\mathbf\{V\}\_\{r\_\{i\}\},\\qquad\\mathbf\{W\}\_\{\\mathrm\{g\}\}\\;\\triangleq\\;\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\\,\\mathbf\{W\}\_\{i\}\.\(21\)
To obtain a meaningful target for the aggregation error, we fix a reference weight𝐖^∈ℝm×n\\widehat\{\\mathbf\{W\}\}\\in\\mathbb\{R\}^\{m\\times n\}against which each group is compared\. The reference is a free design parameter; setting𝐖^=𝐖\\widehat\{\\mathbf\{W\}\}=\\mathbf\{W\}\(the pre\-trained weight\) yields the closed\-form Eckart–Young expression of Definition[3](https://arxiv.org/html/2607.29071#Thmdefinition3)below, while setting𝐖^\\widehat\{\\mathbf\{W\}\}to the centralized minimizerargmin𝐖∑i\(ni/N\)ℒi\(𝐖\)\\arg\\min\_\{\\mathbf\{W\}\}\\sum\_\{i\}\(n\_\{i\}/N\)\\mathcal\{L\}\_\{i\}\(\\mathbf\{W\}\)recovers the comparison with centralized training\. For each groupi∈\[K\]i\\in\[K\], define the geometric projection adapter
𝚺^i≜𝐔ri†𝐖^𝐕ri†,\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\;\\triangleq\\;\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\dagger\}\\,\\widehat\{\\mathbf\{W\}\}\\,\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\dagger\},\(22\)which by Lemma[3](https://arxiv.org/html/2607.29071#Thmlemma3)\(Appendix[A\-A2](https://arxiv.org/html/2607.29071#A1.SS1.SSS2)\) satisfies the unconditional identity
𝐔ri𝚺^i𝐕ri=𝐏ri𝐖^𝐐ri\.\\mathbf\{U\}\_\{r\_\{i\}\}\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\mathbf\{V\}\_\{r\_\{i\}\}\\;=\\;\\mathbf\{P\}\_\{r\_\{i\}\}\\widehat\{\\mathbf\{W\}\}\\mathbf\{Q\}\_\{r\_\{i\}\}\.\(23\)Here𝐀†\\mathbf\{A\}^\{\\dagger\}denotes the Moore–Penrose pseudoinverse, which for a column\-full\-rank matrix𝐀∈ℝm×r\\mathbf\{A\}\\in\\mathbb\{R\}^\{m\\times r\}withm≥rm\\geq radmits the explicit form𝐀†=\(𝐀⊤𝐀\)−1𝐀⊤\\mathbf\{A\}^\{\\dagger\}=\(\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\)^\{\-1\}\\mathbf\{A\}^\{\\top\}, and analogously for a row\-full\-rank𝐁∈ℝr×n\\mathbf\{B\}\\in\\mathbb\{R\}^\{r\\times n\}withn≥rn\\geq r,𝐁†=𝐁⊤\(𝐁𝐁⊤\)−1\\mathbf\{B\}^\{\\dagger\}=\\mathbf\{B\}^\{\\top\}\(\\mathbf\{B\}\\mathbf\{B\}^\{\\top\}\)^\{\-1\}; see Appendix[A\-A](https://arxiv.org/html/2607.29071#A1.SS1)for a self\-contained treatment\. The corresponding identities𝐔ri𝐔ri†=𝐏ri\\mathbf\{U\}\_\{r\_\{i\}\}\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\dagger\}=\\mathbf\{P\}\_\{r\_\{i\}\}and𝐕ri†𝐕ri=𝐐ri\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\dagger\}\\mathbf\{V\}\_\{r\_\{i\}\}=\\mathbf\{Q\}\_\{r\_\{i\}\}hold independently of whether the factorization is pure SVD or scaled\. Throughout the remainder of this subsection we let𝚺i∗\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}denote an element ofargmin𝚺fi\\arg\\min\_\{\\boldsymbol\{\\Sigma\}\}f\_\{i\}\(the one visited by Stage\-1 SGD if it is non\-unique\); the bounds below depend on𝚺i∗\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}only through Frobenius distances, so the choice is immaterial\. Define the*misalignment quantity*
Δi\(𝐖^\)≜fi\(𝚺^i\)−fi∗≥0,fi∗≜min𝚺fi\(𝚺\),\\Delta\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\\;\\triangleq\\;f\_\{i\}\(\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\)\-f\_\{i\}^\{\*\}\\;\\geq\\;0,\\qquad f\_\{i\}^\{\*\}\\triangleq\\min\_\{\\boldsymbol\{\\Sigma\}\}f\_\{i\}\(\\boldsymbol\{\\Sigma\}\),\(24\)which measures the loss gap between the geometric projection𝚺^i\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}and the data\-driven minimizer; it is local and data\-estimable from each group’s training trajectory once Stage\-1 has stabilized\.
###### Definition 3\(Compression residual\)\.
Defineδr\(𝐖^\)≜‖𝐖^−𝐏r𝐖^𝐐r‖F\\delta\_\{r\}\(\\widehat\{\\mathbf\{W\}\}\)\\triangleq\\\|\\widehat\{\\mathbf\{W\}\}\-\\mathbf\{P\}\_\{r\}\\widehat\{\\mathbf\{W\}\}\\mathbf\{Q\}\_\{r\}\\\|\_\{F\}\. When𝐖^=𝐖\\widehat\{\\mathbf\{W\}\}=\\mathbf\{W\}\(the pre\-trained layer itself\), the Eckart–Young–Mirsky theorem gives the closed formδr\(𝐖\)=\(∑k\>rλk2\)1/2\\delta\_\{r\}\(\\mathbf\{W\}\)=\\bigl\(\\sum\_\{k\>r\}\\lambda\_\{k\}^\{2\}\\bigr\)^\{1/2\}\.
We organize the cross\-group aggregation error into three interpretable residuals, each capturing a distinct source of approximation:
εiopt\\displaystyle\\varepsilon\_\{i\}^\{\\mathrm\{opt\}\}≜𝐔ri\(𝚺i\(T1\)−𝚺i∗\)𝐕ri\\displaystyle\\triangleq\\mathbf\{U\}\_\{r\_\{i\}\}\\\!\\bigl\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\}\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}\\bigr\)\\mathbf\{V\}\_\{r\_\{i\}\}\(optimization residual\),εimis\\displaystyle\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}≜𝐔ri\(𝚺i∗−𝚺^i\)𝐕ri\\displaystyle\\triangleq\\mathbf\{U\}\_\{r\_\{i\}\}\\\!\\bigl\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}\-\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\bigr\)\\mathbf\{V\}\_\{r\_\{i\}\}\(misalignment residual\),εisub\\displaystyle\\varepsilon\_\{i\}^\{\\mathrm\{sub\}\}≜𝐏ri𝐖^𝐐ri−𝐖^\\displaystyle\\triangleq\\mathbf\{P\}\_\{r\_\{i\}\}\\widehat\{\\mathbf\{W\}\}\\mathbf\{Q\}\_\{r\_\{i\}\}\-\\widehat\{\\mathbf\{W\}\}\(subspace residual\),so that𝐖i−𝐖^=εiopt\+εimis\+εisub\\mathbf\{W\}\_\{i\}\-\\widehat\{\\mathbf\{W\}\}=\\varepsilon\_\{i\}^\{\\mathrm\{opt\}\}\+\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}\+\\varepsilon\_\{i\}^\{\\mathrm\{sub\}\}by direct telescoping\.
We now state the general cross\-group aggregation bound of the three residuals\. The proof is given in Appendix[A\-E](https://arxiv.org/html/2607.29071#A1.SS5)\.
###### Theorem 2\(Cross\-group aggregation error\)\.
Let Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)hold\. For any reference weight𝐖^∈ℝm×n\\widehat\{\\mathbf\{W\}\}\\in\\mathbb\{R\}^\{m\\times n\},
𝔼∥𝐖g−𝐖^∥F≤∑i=1KniN\[𝔼‖εiopt‖F\+𝔼‖εimis‖F\+δri\(𝐖^\)\]\.\\begin\{split\}\\mathbb\{E\}\\bigl\\\|\\mathbf\{W\}\_\{\\mathrm\{g\}\}\-\\widehat\{\\mathbf\{W\}\}\\bigr\\\|\_\{F\}\\;\\leq\\;\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\\\!\\Bigl\[\\,&\\mathbb\{E\}\\\|\\varepsilon\_\{i\}^\{\\mathrm\{opt\}\}\\\|\_\{F\}\+\\mathbb\{E\}\\\|\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}\\\|\_\{F\}\\\\ \+\\;&\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\)\\Bigr\]\.\\end\{split\}\(25\)The bound separates the cross\-group aggregation error into one term per residual, weighted by data mass and aggregated by triangle inequality\.
The three terms in \([25](https://arxiv.org/html/2607.29071#S4.E25)\) are conceptually orthogonal:εiopt\\varepsilon\_\{i\}^\{\\mathrm\{opt\}\}measures how far Stage\-1 SGD has progressed,εimis\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}measures whether the data\-driven minimizer aligns with the geometric projection of𝐖^\\widehat\{\\mathbf\{W\}\}, andεisub\\varepsilon\_\{i\}^\{\\mathrm\{sub\}\}measures the intrinsic cost of compressing𝐖^\\widehat\{\\mathbf\{W\}\}to rankrir\_\{i\}\. The next proposition shows how two standard regularity conditions sharpen the first two residuals: a Polyak–Łojasiewicz inequality drivesεiopt\\varepsilon\_\{i\}^\{\\mathrm\{opt\}\}to zero with the number of rounds, and a common\-projection\-target condition drivesεimis\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}to zero outright\.
###### Proposition 2\(Residual control\)\.
Fix any reference weight𝐖^∈ℝm×n\\widehat\{\\mathbf\{W\}\}\\in\\mathbb\{R\}^\{m\\times n\}\.
1. \(a\)\(PŁ condition⇒\\Rightarrowvanishing optimization residual\.\)If, for everyi∈\[K\]i\\in\[K\], the adapter lossfif\_\{i\}satisfies the Polyak–Łojasiewicz inequality ‖∇𝚺fi\(𝚺\)‖F2≥2μ\(fi\(𝚺\)−fi∗\),∀𝚺∈ℝri×ri,\\hskip\-10\.00002pt\\bigl\\\|\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{i\}\(\\boldsymbol\{\\Sigma\}\)\\bigr\\\|\_\{F\}^\{\\,2\}\\;\\geq\\;2\\mu\\bigl\(f\_\{i\}\(\\boldsymbol\{\\Sigma\}\)\-f\_\{i\}^\{\*\}\\bigr\),\\,\\forall\\boldsymbol\{\\Sigma\}\\in\\mathbb\{R\}^\{r\_\{i\}\\times r\_\{i\}\},\(26\)with constantμ\>0\\mu\>0, then under Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)the optimization residual obeys 𝔼‖εiopt‖F≤λ1κ\(𝐒\)⋅\(𝒪\(T1−1/4\)\+𝒪\(G/μ\)\),\\mathbb\{E\}\\\|\\varepsilon\_\{i\}^\{\\mathrm\{opt\}\}\\\|\_\{F\}\\;\\leq\\;\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\\\!\\cdot\\\!\\Bigl\(\\mathcal\{O\}\(T\_\{1\}^\{\-1/4\}\)\+\\mathcal\{O\}\(G/\\sqrt\{\\mu\}\)\\Bigr\),\(27\)and in particular𝔼‖εiopt‖F→0\\mathbb\{E\}\\\|\\varepsilon\_\{i\}^\{\\mathrm\{opt\}\}\\\|\_\{F\}\\to 0asT1→∞T\_\{1\}\\to\\inftyin the IID regime \(G=0G=0\)\.
2. \(b\)\(Common projection target⇒\\Rightarrowzero misalignment residual\.\)If, for everyi∈\[K\]i\\in\[K\], the geometric projection adapter𝚺^i\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}defined in \([22](https://arxiv.org/html/2607.29071#S4.E22)\) satisfies 𝚺^i∈argmin𝚺fi,\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\;\\in\\;\\arg\\min\_\{\\boldsymbol\{\\Sigma\}\}f\_\{i\},\(28\)then𝚺^i=𝚺i∗\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}=\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}andεimis=𝟎\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}=\\mathbf\{0\}for everyii\(the special caseΔi\(𝐖^\)=0\\Delta\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)=0of the bound below\)\. Without assuming \([28](https://arxiv.org/html/2607.29071#S4.E28)\), if part \(a\) holds, then 𝔼‖εimis‖F≤λ1κ\(𝐒\)2Δi\(𝐖^\)μ\.\\mathbb\{E\}\\\|\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}\\\|\_\{F\}\\;\\leq\\;\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\\,\\sqrt\{\\tfrac\{2\\,\\Delta\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\}\{\\mu\}\}\.\(29\)
The PŁ inequality is a standard regularity assumption\[[24](https://arxiv.org/html/2607.29071#bib.bib44)\]satisfied by strong convexity, overparameterized neural losses near interpolating minima\[[32](https://arxiv.org/html/2607.29071#bib.bib83)\], low\-rank matrix recovery, and many compositional objectives; PŁ\-based analyses underlie much of FL theory\[[17](https://arxiv.org/html/2607.29071#bib.bib81),[64](https://arxiv.org/html/2607.29071#bib.bib86),[47](https://arxiv.org/html/2607.29071#bib.bib85)\]\. The common\-projection\-target condition \([28](https://arxiv.org/html/2607.29071#S4.E28)\) is a non\-trivial geometric requirement: it asks that the rank\-rir\_\{i\}minimizer of the*loss*coincide with the*Frobenius*projection of𝐖^\\widehat\{\\mathbf\{W\}\}onto the SVD subspace\. To expose the theorem’s core mechanism cleanly, Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3)below analyzes the case in which the common\-projection\-target condition \([28](https://arxiv.org/html/2607.29071#S4.E28)\) holds, soεimis=𝟎\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}=\\mathbf\{0\}and the aggregation error reduces to pure subspace geometry\. When the condition fails,εimis\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}is non\-zero but admits explicit bounds via either a curvature\-side or a data\-side analysis; the analytical structure is the same, and the full treatment is given in Appendix[A\-E](https://arxiv.org/html/2607.29071#A1.SS5)\.
###### Theorem 3\(Common\-projection\-target specialization\)\.
Suppose Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)and Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\(b\) both hold, and the optimization residual is negligible\. Then the global reconstructed weight satisfies
‖𝐖g−𝐖^‖F≤∑i=1KniNδri\(𝐖^\)≤δrmin\(𝐖^\),\\bigl\\\|\\mathbf\{W\}\_\{\\mathrm\{g\}\}\-\\widehat\{\\mathbf\{W\}\}\\bigr\\\|\_\{F\}\\;\\leq\\;\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\\,\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\)\\;\\leq\\;\\delta\_\{r\_\{\\min\}\}\(\\widehat\{\\mathbf\{W\}\}\),\(30\)wherermin=minirir\_\{\\min\}=\\min\_\{i\}r\_\{i\}\. If, in addition,𝐖^\\widehat\{\\mathbf\{W\}\}coincides with the pre\-trained weight𝐖\\mathbf\{W\}, then
‖𝐖g−𝐖^‖F≤∑i=1KniN\(∑k\>riλk2\)1/2\.\\bigl\\\|\\mathbf\{W\}\_\{\\mathrm\{g\}\}\-\\widehat\{\\mathbf\{W\}\}\\bigr\\\|\_\{F\}\\;\\leq\\;\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\\,\\biggl\(\\sum\_\{k\>r\_\{i\}\}\\lambda\_\{k\}^\{2\}\\biggr\)^\{\\\!1/2\}\.\(31\)
Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3)is Theorem[2](https://arxiv.org/html/2607.29071#Thmtheorem2)specialized toεiopt≈𝟎\\varepsilon\_\{i\}^\{\\mathrm\{opt\}\}\\approx\\mathbf\{0\}\(vanishing rounds\) andεimis=𝟎\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}=\\mathbf\{0\}\(Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\(b\)\): the bound collapses to the data\-weighted average of subspace residuals\. The*operational*bound is therefore∑i\(ni/N\)δri\(𝐖^\)\\sum\_\{i\}\(n\_\{i\}/N\)\\,\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\); the loose majorantδrmin\(𝐖^\)\\delta\_\{r\_\{\\min\}\}\(\\widehat\{\\mathbf\{W\}\}\)should be read as a*worst\-case*guarantee, not a design target\. Allocating more data to higher\-rank groups reduces the operational bound even whenrminr\_\{\\min\}is fixed, and whenever𝐖^≈𝐖\\widehat\{\\mathbf\{W\}\}\\approx\\mathbf\{W\}, the server can*pre\-compute*the subspace residual from the singular spectrum of𝐖\\mathbf\{W\}before any training occurs, enabling spectrum\-informed rank allocation\.
### IV\-CArtifact Mitigation via Confidence Loss
The weak\-to\-strong refinement stage \(Phase 3\) poses a subtle risk: when the strong model is supervised by the aggregated weak model, it may not only acquire the federated domain knowledge but also replicate the systematic errors introduced by SVD truncation\[[6](https://arxiv.org/html/2607.29071#bib.bib19)\]\. The auxiliary confidence lossℒconf\\mathcal\{L\}\_\{\\mathrm\{conf\}\}\(Eq\. \([10](https://arxiv.org/html/2607.29071#S3.E10)\)\) is designed to mitigate this effect\. We analyze its role in two steps: first, we derive a gradient\-level decomposition that exposes how the compression artifact enters the update of the strong model \(Proposition[3](https://arxiv.org/html/2607.29071#Thmproposition3)\); second, we show that the mixing coefficientα\\alphacan be tuned to minimize an explicit bias–variance trade\-off \(Proposition[4](https://arxiv.org/html/2607.29071#Thmproposition4)\)\.
###### Definition 4\(Ideal vs\. actual weak predictions\)\.
Let𝒴\\mathcal\{Y\}denote the label set \(vocabulary for language models\) withC=\|𝒴\|C=\|\\mathcal\{Y\}\|classes, and letΔC−1\\Delta^\{C\-1\}denote the\(C−1\)\(C\\\!\-\\\!1\)\-dimensional probability simplex\. For any inputxx, write𝐳\(x;𝐖\)∈ℝC\\mathbf\{z\}\(x;\\,\\mathbf\{W\}\)\\in\\mathbb\{R\}^\{C\}for the pre\-softmax logit vector produced by the model with weight configuration𝐖\\mathbf\{W\}\. Define the*ideal*weak prediction as the output the aggregated model would produce if𝐖g\\mathbf\{W\}\_\{\\mathrm\{g\}\}were replaced by the reference weight𝐖^\\widehat\{\\mathbf\{W\}\}from Section[IV\-B](https://arxiv.org/html/2607.29071#S4.SS2):
f∗\(x\)≜softmax\(𝐳\(x;𝐖^\)\)∈ΔC−1\.f^\{\*\}\(x\)\\;\\triangleq\\;\\mathrm\{softmax\}\\\!\\bigl\(\\mathbf\{z\}\(x;\\,\\widehat\{\\mathbf\{W\}\}\)\\bigr\)\\;\\in\\;\\Delta^\{C\-1\}\.The weak prediction then admits the additive decomposition
fw\(x;𝐖g\)=f∗\(x\)\+ξ\(x\),f\_\{\\mathrm\{w\}\}\(x;\\,\\mathbf\{W\}\_\{\\mathrm\{g\}\}\)\\;=\\;f^\{\*\}\(x\)\+\\xi\(x\),\(32\)whereξ\(x\)∈ℝC\\xi\(x\)\\in\\mathbb\{R\}^\{C\}is the compression artifact\. Sincefwf\_\{\\mathrm\{w\}\}andf∗f^\{\*\}are both valid probability distributions \(i\.e\., their entries are non\-negative and sum to one\), the entries ofξ\(x\)\\xi\(x\)must sum to zero:𝟏⊤ξ\(x\)=0\\mathbf\{1\}^\{\\top\}\\xi\(x\)=0\.
Intuitively,f∗\(x\)f^\{\*\}\(x\)represents the “clean” prediction that the weak model would make if the federated optimization had converged perfectly and no approximation error were introduced by low\-rank truncation\. In practice, the aggregated global model𝐖g\\mathbf\{W\}\_\{\\mathrm\{g\}\}deviates from𝐖^\\widehat\{\\mathbf\{W\}\}due to both finite\-round optimization error and the intrinsic subspace approximation error analyzed in Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3)\. The artifactξ\(x\)\\xi\(x\)captures the combined effect of these two sources of error in the output space: it shifts probability mass across classes in a zero\-sum manner, systematically distorting the weak model’s predictions\. The magnitude ofξ\(x\)\\xi\(x\)is controlled by the weight\-space aggregation error‖𝐖g−𝐖^‖F\\\|\\mathbf\{W\}\_\{\\mathrm\{g\}\}\-\\widehat\{\\mathbf\{W\}\}\\\|\_\{F\}through the Lipschitz forward\-pass assumption below, which provides the bridge between the weight\-space bounds of Section[IV\-B](https://arxiv.org/html/2607.29071#S4.SS2)and the output\-space analysis that follows\.
###### Assumption 2\(Lipschitz forward pass\)\.
There existsLf≥0L\_\{f\}\\geq 0such that for every inputxxand every pair of parameter configurations\(𝐖,𝐖′\)\(\\mathbf\{W\},\\mathbf\{W\}^\{\\prime\}\)differing only in the selected layer,
‖f\(x;𝐖\)−f\(x;𝐖′\)‖1≤Lf‖𝐖−𝐖′‖F\.\\bigl\\\|f\(x;\\,\\mathbf\{W\}\)\-f\(x;\\,\\mathbf\{W\}^\{\\prime\}\)\\bigr\\\|\_\{1\}\\;\\leq\\;L\_\{f\}\\,\\\|\\mathbf\{W\}\-\\mathbf\{W\}^\{\\prime\}\\\|\_\{F\}\.\(33\)
Having established that the artifactξ\(x\)\\xi\(x\)exists as a well\-defined perturbation in the output space, we now quantify its magnitude\. The following lemma combines the Lipschitz forward\-pass assumption with the weight\-space aggregation error from Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3)to obtain a uniform bound on‖ξ\(x\)‖1\\\|\\xi\(x\)\\\|\_\{1\}across all inputs\.
###### Lemma 1\(Artifact magnitude bound\)\.
Under Assumption[2](https://arxiv.org/html/2607.29071#Thmassumption2)and Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3), for∀x\\forall\\,x
‖ξ\(x\)‖1≤Lf⋅∑i=1KniNδri\(𝐖^\)≤Lfδrmin\(𝐖^\)=:Bξ\.\\begin\{split\}\\\|\\xi\(x\)\\\|\_\{1\}&\\;\\leq\\;L\_\{f\}\\\!\\cdot\\\!\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\\,\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\)\\\\ &\\;\\leq\\;L\_\{f\}\\,\\delta\_\{r\_\{\\min\}\}\(\\widehat\{\\mathbf\{W\}\}\)\\;=:\\;B\_\{\\xi\}\.\\end\{split\}\(34\)
The boundBξB\_\{\\xi\}is determined entirely by the compression configuration: it grows with the Lipschitz constantLfL\_\{f\}of the forward pass and with the tail singular\-value energyδrmin\\delta\_\{r\_\{\\min\}\}of the most aggressively compressed group\. This quantity serves as the noise budget in the subsequent gradient\-level analysis\.
We now turn to the central question: how does the auxiliary confidence lossℒconf\\mathcal\{L\}\_\{\\mathrm\{conf\}\}interact with this artifact when updating the strong model? Recall from Eq\. \([10](https://arxiv.org/html/2607.29071#S3.E10)\) that the loss combines imitation of the weak model and self\-anchoring of the strong model’s own confident predictions\. The next proposition decomposes the logit\-level gradient of this loss into three interpretable components, revealing where the compression artifact enters and howα\\alphamodulates its influence\.
###### Proposition 3\(Gradient\-level artifact attenuation\)\.
Let𝐳s\(x\)∈ℝC\\mathbf\{z\}\_\{s\}\(x\)\\in\\mathbb\{R\}^\{C\}denote the pre\-softmax logits of the strong model\. The logit gradient ofℒconf\\mathcal\{L\}\_\{\\mathrm\{conf\}\}admits the following exact decomposition:
∇𝐳sℒconf=\(1−α\)\(fs−f∗\)⏟\(I\) clean signal−\(1−α\)ξ\(x\)⏟\(II\) attenuated artifact\+α\(fs−f^s\)⏟\(III\) self\-anchoring\.\\begin\{split\}\\nabla\_\{\\mathbf\{z\}\_\{s\}\}\\mathcal\{L\}\_\{\\mathrm\{conf\}\}=\\;&\\underbrace\{\(1\\\!\-\\\!\\alpha\)\(f\_\{\\mathrm\{s\}\}\-f^\{\*\}\)\}\_\{\\text\{\(I\) clean signal\}\}\\\\ &\-\\underbrace\{\(1\\\!\-\\\!\\alpha\)\\,\\xi\(x\)\}\_\{\\text\{\(II\) attenuated artifact\}\}\+\\underbrace\{\\alpha\(f\_\{\\mathrm\{s\}\}\-\\hat\{f\}\_\{\\mathrm\{s\}\}\)\}\_\{\\text\{\(III\) self\-anchoring\}\}\.\\end\{split\}\(35\)Under Assumption[2](https://arxiv.org/html/2607.29071#Thmassumption2), the artifact term \(II\) is bounded by
‖\(1−α\)ξ\(x\)‖2≤\(1−α\)Lfδrmin\(𝐖^\)\.\\bigl\\\|\(1\\\!\-\\\!\\alpha\)\\,\\xi\(x\)\\bigr\\\|\_\{2\}\\;\\leq\\;\(1\\\!\-\\\!\\alpha\)\\,L\_\{f\}\\,\\delta\_\{r\_\{\\min\}\}\(\\widehat\{\\mathbf\{W\}\}\)\.\(36\)
The decomposition \([35](https://arxiv.org/html/2607.29071#S4.E35)\) reveals three forces acting on the strong model at each gradient step\. Term \(I\) drives the strong model toward the ideal predictionf∗f^\{\*\}, which is the useful federated knowledge\. Term \(II\) is the compression artifact that pushes the model in a spurious direction; its magnitude is scaled down by\(1−α\)\(1\-\\alpha\)\. Term \(III\) is the self\-anchoring regularizer that pulls the strong model toward its own high\-confidence predictions, providing an independent reference signal that does not depend on the noisy weak supervisor\. The factor\(1−α\)\(1\-\\alpha\)in \([36](https://arxiv.org/html/2607.29071#S4.E36)\) is an algebraic consequence of the convex combination definingℒconf\\mathcal\{L\}\_\{\\mathrm\{conf\}\}: the same\(1−α\)\(1\-\\alpha\)scales the*clean*signal \(I\)\. Proposition[3](https://arxiv.org/html/2607.29071#Thmproposition3)therefore does*not*claim that the confidence loss selectively filters artifacts\. Its role is to make the bias–variance trade\-off explicit and to*couple*the weight\-space error of Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3)with the output\-space error of the refined strong model\. A genuine noise\-suppression argument must rely on the self\-anchoring term \(III\), which is analyzed in Proposition[4](https://arxiv.org/html/2607.29071#Thmproposition4)below\.
###### Proposition 4\(Bias–variance trade\-off ofα\\alpha\)\.
Suppose that on a random samplex∼𝒟serverx\\sim\\mathcal\{D\}\_\{\\mathrm\{server\}\}, the self\-anchoring directionfs−f^sf\_\{\\mathrm\{s\}\}\-\\hat\{f\}\_\{\\mathrm\{s\}\}has bounded second moment𝔼‖fs−f^s‖22≤Vself\\mathbb\{E\}\\\|f\_\{\\mathrm\{s\}\}\-\\hat\{f\}\_\{\\mathrm\{s\}\}\\\|\_\{2\}^\{2\}\\leq V\_\{\\mathrm\{self\}\}, and that the artifactξ\(x\)\\xi\(x\)and the self\-anchoring residualfs−f^sf\_\{\\mathrm\{s\}\}\-\\hat\{f\}\_\{\\mathrm\{s\}\}are uncorrelated, i\.e\.,𝔼⟨ξ\(x\),fs−f^s⟩=0\\mathbb\{E\}\\langle\\xi\(x\),\\,f\_\{\\mathrm\{s\}\}\-\\hat\{f\}\_\{\\mathrm\{s\}\}\\rangle=0\. We define the excess risk of the logit\-gradient estimator against the clean targetfs−f∗f\_\{\\mathrm\{s\}\}\-f^\{\*\}as follows:
ℛ\(α\)≜\(1−α\)2𝔼‖ξ\(x\)‖22⏟bias from artifact\+α2Vself⏟variance from self\-anchor\.\\mathcal\{R\}\(\\alpha\)\\;\\triangleq\\;\\underbrace\{\(1\-\\alpha\)^\{2\}\\,\\mathbb\{E\}\\\|\\xi\(x\)\\\|\_\{2\}^\{2\}\}\_\{\\text\{bias from artifact\}\}\\;\+\\;\\underbrace\{\\alpha^\{2\}\\,V\_\{\\mathrm\{self\}\}\}\_\{\\text\{variance from self\-anchor\}\}\.\(37\)Thenℛ\\mathcal\{R\}is minimized at
α∗=𝔼‖ξ\(x\)‖22Vself\+𝔼‖ξ\(x\)‖22≤Bξ2Vself\+Bξ2,\\alpha^\{\*\}\\;=\\;\\frac\{\\mathbb\{E\}\\\|\\xi\(x\)\\\|\_\{2\}^\{2\}\}\{V\_\{\\mathrm\{self\}\}\+\\mathbb\{E\}\\\|\\xi\(x\)\\\|\_\{2\}^\{2\}\}\\;\\leq\\;\\frac\{B\_\{\\xi\}^\{\\,2\}\}\{V\_\{\\mathrm\{self\}\}\+B\_\{\\xi\}^\{\\,2\}\},\(38\)andα∗\\alpha^\{\*\}is monotone non\-decreasing inBξB\_\{\\xi\}: heavier compression calls for a larger mixing weight on the self\-anchor\.
Propositions[3](https://arxiv.org/html/2607.29071#Thmproposition3)and[4](https://arxiv.org/html/2607.29071#Thmproposition4)together furnish an end\-to\-end guarantee: Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3)bounds the*weight\-space*aggregation error, Lemma[1](https://arxiv.org/html/2607.29071#Thmlemma1)transports this bound into*output space*, and Proposition[4](https://arxiv.org/html/2607.29071#Thmproposition4)prescribes how to trade off the resulting bias against self\-anchor variance viaα\\alpha\. The monotone dependence ofα∗\\alpha^\{\*\}onBξB\_\{\\xi\}yields a spectrum\-based heuristic: as the compression ratio increases and tail singular\-value energy decreases,Bξ→0B\_\{\\xi\}\\to 0andα∗→0\\alpha^\{\*\}\\to 0, so the strong model should defer more to the weak supervisor\. Conversely, at high compression,α∗→1\\alpha^\{\*\}\\to 1and the strong model should rely on its own confident predictions\.
## VExperiments
Goals\.Our experiments evaluate FedSLM along four axes: \(1\) server\-side performance against existing federated baselines on language and multimodal tasks; \(2\) client\-side performance of SVD\-compressed models; \(3\) convergence behavior across heterogeneous compression ratios; and \(4\) the effect of adapter placement strategy on downstream performance\.
For SVD\-compressed clients \(DobiSVD and QSVD\), trainable adapters are inserted only into the three MLP projections \(gate\_proj,up\_proj,down\_proj\)\. Each adapter is a square matrix𝚺\\boldsymbol\{\\Sigma\}placed between the two frozen SVD factors, with𝚺\\boldsymbol\{\\Sigma\}initialized to the identity; the SVD\-decomposed attention layers remain fully frozen\. For full\-model clients \(LLaMA\-2\-7B, LLaMA\-2\-13B, and LLaVA\-NeXT\-7B\), we apply LoRA\[[21](https://arxiv.org/html/2607.29071#bib.bib20)\]adapters to all linear projections per transformer block: the attention projections \(q\_proj,k\_proj,v\_proj,o\_proj\) and the MLP projections above\.
Datasets\.For training, we use AI2 Reasoning Challenge \(ARC\-Easy and ARC\-Challenge\)\[[10](https://arxiv.org/html/2607.29071#bib.bib57)\], PIQA\[[4](https://arxiv.org/html/2607.29071#bib.bib58)\], WinoGrande\[[45](https://arxiv.org/html/2607.29071#bib.bib59)\], Social IQA\[[46](https://arxiv.org/html/2607.29071#bib.bib60)\], HellaSwag\[[66](https://arxiv.org/html/2607.29071#bib.bib61)\], and COPA\[[44](https://arxiv.org/html/2607.29071#bib.bib63)\], together with Medical\-Flashcards\[[38](https://arxiv.org/html/2607.29071#bib.bib64)\]as a domain\-specific medical corpus\. For medical evaluation, we report zero\-shot transfer accuracy on PubMedQA\[[23](https://arxiv.org/html/2607.29071#bib.bib65)\]and MedMCQA\[[43](https://arxiv.org/html/2607.29071#bib.bib66)\]\. For multimodal evaluation, we use ScienceQA\[[34](https://arxiv.org/html/2607.29071#bib.bib67)\]and VizWiz\[[16](https://arxiv.org/html/2607.29071#bib.bib68)\]\.
Baselines\.We compare against eight federated learning methods: \(1\) FedAvg\+LoRA\[[37](https://arxiv.org/html/2607.29071#bib.bib41),[21](https://arxiv.org/html/2607.29071#bib.bib20)\], which applies standard federated averaging to LoRA adapters; \(2\) FFA\-LoRA\[[48](https://arxiv.org/html/2607.29071#bib.bib26)\], a frozen\-weight variant of federated LoRA; \(3\) HetLoRA\[[9](https://arxiv.org/html/2607.29071#bib.bib24)\], which supports heterogeneous LoRA ranks across clients; \(4\) FlexLoRA\[[3](https://arxiv.org/html/2607.29071#bib.bib23)\], which dynamically allocates LoRA ranks based on client capacity; \(5\) Fed\-RAC\-LoRA\[[35](https://arxiv.org/html/2607.29071#bib.bib69)\], which randomly freezes one LoRA factor per step and iteratively merges the update into the base weights to form a convergent chain; \(6\) FedMKT\[[14](https://arxiv.org/html/2607.29071#bib.bib15)\], a mutual knowledge transfer approach for heterogeneous models; \(7\) FedProto\[[49](https://arxiv.org/html/2607.29071#bib.bib70)\], which exchanges class prototypes instead of gradients to tolerate client heterogeneity; and \(8\) FedBiOT\[[59](https://arxiv.org/html/2607.29071#bib.bib71)\], which has clients fine\-tune a lightweight adapter on a server\-compressed LLM via bi\-level optimization\. We also report two ablated variants of our method: FedSLM \(α=0\\alpha\\\!=\\\!0\), the server model refined via weak\-to\-strong elicitation using only the weak\-imitation loss; and FedSLM \(α=0\.5\\alpha\\\!=\\\!0\.5\), which additionally activates the auxiliary confidence loss that balances weak supervision with the strong model’s confident predictions\.
TABLE II:Performance comparison on natural language understanding benchmarks with LLaMA\-2\-7B as the server model \(accuracy %\)\. We report results under IID and non\-IID \(Dirichletβ=1\.0\\beta=1\.0\) data partitions\. The best result in each column isboldfaced\.Figure 2:Client\-side efficiency of SVD\-compressed LLaMA\-2\-7B models, with a Full\+LoRA reference\. Left: GPU memory footprint before and after attaching adapters \(or LoRA modules for the full model\)\. Middle: inference throughput versus batch size \(input length 128, 32 generated tokens\)\. Right: inference throughput versus input sequence length \(batch size 8, 32 generated tokens\)\.Federated Setup\.We partition training data across clients under two settings: IID \(uniform\) and non\-IID \(Dirichletβ=1\.0\\beta=1\.0\)\. For commonsense reasoning tasks, we simulate a federation of 10 clients comprising three compression groups \(SVD\-0\.4×\\times3, SVD\-0\.6×\\times4, full server model×\\times3\) and runT1=5T\_\{1\}=5rounds of adapter\-level aggregation\. For the medical task, we scale to 20 clients \(SVD\-0\.4×\\times7, SVD\-0\.6×\\times7, full server model×\\times6\) withT1=10T\_\{1\}=10rounds\. For multimodal tasks, we deploy a federation of 10 clients with three QSVD compression groups \(QSVD\-0\.6×\\times3, QSVD\-0\.9×\\times4, full LLaVA\-NeXT\-7B×\\times3\) under the sameT1=5T\_\{1\}=5rounds of adapter\-level aggregation\. The heterogeneous client composition above is shared by FedSLM and the heterogeneity\-aware baselines FedMKT, FedProto, and FedBiOT\. The remaining baselines assume a homogeneous client pool by design; for fairness, we evaluate them on the subset of clients that natively host the full server model, with all other hyperparameters \(number of clients, data partition, and training rounds\) kept consistent\. In all settings, weak\-to\-strong elicitation runs forT2=5T\_\{2\}=5epochs\.
TABLE III:Performance comparison on natural language understanding benchmarks with LLaMA\-2\-13B as the server model \(accuracy %\)\. We report results under IID and non\-IID \(Dirichletβ=1\.0\\beta=1\.0\) data partitions\. The best result in each column isboldfaced\.### V\-AMain Results
Tables[II](https://arxiv.org/html/2607.29071#S5.T2)and[III](https://arxiv.org/html/2607.29071#S5.T3)report accuracy on nine benchmarks under IID and non\-IID partitions, with LLaMA\-2\-7B and LLaMA\-2\-13B serving as the server model respectively\.
On LLaMA\-2\-7B \(Table[II](https://arxiv.org/html/2607.29071#S5.T2)\), FedSLM \(α=0\.5\\alpha\\\!=\\\!0\.5\) achieves 71\.0% average accuracy under IID and 68\.2% under non\-IID, surpassing FedMKT \(66\.9% / 66\.4%\) by 4\.1% / 1\.8% and FedAvg\+LoRA \(64\.0% / 62\.2%\) by 7\.0% / 6\.0%\. The gains are pronounced on benchmarks requiring deeper reasoning: under IID, HellaSwag reaches 80\.2% versus 60\.5% for FedAvg\+LoRA and ARC\_c reaches 54\.8% versus 40\.8%; under non\-IID, ARC\_c reaches 57\.2% versus 52\.8% for FedMKT\. On the medical benchmarks, FedSLM \(α=0\.5\\alpha\\\!=\\\!0\.5\) reaches 71\.5% on PubMedQA and 35\.1% on MedMCQA under IID, exceeding FedAvg\+LoRA by 3\.4% and 2\.5%\.
On LLaMA\-2\-13B \(Table[III](https://arxiv.org/html/2607.29071#S5.T3)\), the trend scales up: FedSLM \(α=0\.5\\alpha\\\!=\\\!0\.5\) attains 71\.4% under IID and 68\.1% under non\-IID, outperforming FedMKT \(68\.3% / 67\.0%\) by 3\.1% / 1\.1% and FedAvg\+LoRA \(66\.1% / 64\.3%\) by 5\.3% / 3\.8%\. The advantage on reasoning\-heavy benchmarks persists, with HellaSwag reaching 68\.0% under IID versus 62\.8% for FedAvg\+LoRA and ARC\_c reaching 53\.2% versus 43\.2%\. On the medical benchmarks, FedSLM \(α=0\.5\\alpha\\\!=\\\!0\.5\) reaches 73\.9% on PubMedQA and 36\.6% on MedMCQA under IID, exceeding FedAvg\+LoRA by 3\.4% and 2\.6%\.
Comparing FedSLM \(α=0\\alpha\\\!=\\\!0\) with FedSLM \(α=0\.5\\alpha\\\!=\\\!0\.5\) isolates the contribution of the auxiliary confidence loss across both model scales: under IID, average accuracy improves from 69\.9% to 71\.0% on 7B and from 70\.6% to 71\.4% on 13B; under non\-IID, from 67\.8% to 68\.2% on 7B and from 67\.4% to 68\.1% on 13B\. The confidence loss is most valuable on benchmarks where the weak supervisor is noisier\. For example, HellaSwag improves from 75\.1% to 80\.2% under IID on 7B, and ARC\_c improves by 2\.4% on 7B and 1\.7% on 13B under non\-IID\. This matches the theoretical prediction in Proposition[3](https://arxiv.org/html/2607.29071#Thmproposition3)that the\(1−α\)\(1\-\\alpha\)factor attenuates compression artifacts when the weak supervisor is noisy\. In summary, FedSLM consistently ranks first or second on all nine benchmarks under both data partitions and at both model scales, demonstrating that the two\-stage aggregation combined with weak\-to\-strong elicitation effectively transfers federated knowledge to the server model, with the auxiliary confidence loss providing a reliable additional gain\.
### V\-BEfficiency Analysis
Figure 3:Client\-side efficiency of SVD\-compressed LLaMA\-2\-13B models, with a Full\+LoRA reference\. Left: GPU memory footprint before and after attaching adapters \(or LoRA modules for the full model\)\. Middle: inference throughput versus batch size \(input length 128, 32 generated tokens\)\. Right: inference throughput versus input sequence length \(batch size 8, 32 generated tokens\)\.A practical federated system must ensure that compressed client models fit within the memory and latency budgets of participating devices\. We therefore profile the GPU memory footprint and inference throughput of the compressed models deployed on clients, benchmarked against the respective full models at both 7B and 13B scales \(Figure[2](https://arxiv.org/html/2607.29071#S5.F2)and Figure[3](https://arxiv.org/html/2607.29071#S5.F3)\)\.
Figure[2](https://arxiv.org/html/2607.29071#S5.F2)\(left\) compares the GPU memory occupied by DobiSVD\-compressed LLaMA\-2\-7B models before and after inserting the adapter𝚺\\boldsymbol\{\\Sigma\}\. The full LLaMA\-2\-7B occupies 13\.5 GB, whereas DobiSVD\-0\.4 and DobiSVD\-0\.6 require roughly 6\.8 GB and 7\.8 GB respectively\. After attaching adapters, the memory overhead is negligible: DobiSVD\-0\.4\+Adapter and DobiSVD\-0\.6\+Adapter remain close to their base counterparts, confirming that the lightweight adapter parameterization adds minimal cost to the compressed model\. Overall, FedSLM clients require only about 54% \(0\.4\-ratio\) to 62% \(0\.6\-ratio\) of the GPU memory needed by the full model\.
Figure[3](https://arxiv.org/html/2607.29071#S5.F3)\(left\) shows the corresponding memory profile for LLaMA\-2\-13B, which additionally includes a Full\+LoRA reference for comparison against parameter\-efficient fine\-tuning of the uncompressed model\. The full LLaMA\-2\-13B occupies 24\.2 GB, while DobiSVD\-0\.4 and DobiSVD\-0\.6 require only 10\.1 GB \(42%\) and 14\.8 GB \(61%\) respectively\. After attaching the trainable adapter, DobiSVD\-0\.4\+Adapter rises to 10\.5 GB \(43%\) and DobiSVD\-0\.6\+Adapter to 15\.8 GB \(65%\), an overhead of≤1\\leq 1GB in both cases\. By contrast, attaching LoRA modules to the full model raises memory only marginally \(24\.4 GB\), confirming that the dominant memory cost is the base weights themselves\.
Figures[2](https://arxiv.org/html/2607.29071#S5.F2)and[3](https://arxiv.org/html/2607.29071#S5.F3)\(middle\) show inference throughput as a function of batch size\. At both scales, the compressed models consistently achieve higher throughput than the full model across all batch sizes, with the 0\.4\-ratio model delivering the largest speedup\. For LLaMA\-2\-13B at batch size 64, DobiSVD\-0\.4 reaches 522 tokens/s and DobiSVD\-0\.6 reaches 444 tokens/s, compared to 416 tokens/s for the full model and 384 tokens/s for the LoRA\-tuned full model—speedups of1\.25×1\.25\\timesand1\.07×1\.07\\timesover the Full baseline, and1\.36×1\.36\\timesand1\.16×1\.16\\timesover Full\+LoRA\. The throughput advantage grows with batch size, as the smaller weight matrices benefit from more efficient matrix multiplications\. Adapter attachment incurs only a small throughput penalty \(e\.g\., 512→\\to497 tokens/s for SVD\-0\.4 at batch 64\), preserving most of the compression benefit\.
Figures[2](https://arxiv.org/html/2607.29071#S5.F2)and[3](https://arxiv.org/html/2607.29071#S5.F3)\(right\) vary the input sequence length over\{32,64,128,256\}\\\{32,64,128,256\\\}with batch size fixed at 8\. Throughput decreases for all configurations as the sequence length grows, reflecting the quadratic cost of self\-attention, but the relative ordering is preserved across all tested lengths: at sequence length 256 on LLaMA\-2\-13B, DobiSVD\-0\.4\+Adapter sustains 160 tokens/s versus 122 tokens/s for the full model and 108 tokens/s for Full\+LoRA, a1\.31×1\.31\\timesand1\.48×1\.48\\timesadvantage, respectively\. The gap between Full and Full\+LoRA widens slightly at long sequences, indicating that LoRA’s overhead becomes more visible when attention dominates the wall\-clock budget; the compressed variants are insulated from this effect because their compression also shrinks the projection matrices that drive attention computation\.
Taken together, these results confirm that FedSLM clients operate with roughly half the GPU memory of the full model while achieving higher inference throughput at both 7B and 13B scales\. Critically, at the 13B scale FedSLM\-0\.4 outperforms not only the full model but also the parameter\-efficient Full\+LoRA baseline on both memory and throughput simultaneously, making federated participation feasible on GPUs that cannot host the full model\.
### V\-CConvergence Analysis
Figure 4:Training loss convergence of LLaMA\-2\-7B clients over communication rounds on all eight benchmarks under the non\-IID \(β=1\.0\\beta=1\.0\) partition\.Figure 5:Training loss convergence of LLaMA\-2\-13B clients over communication rounds on all eight benchmarks under the non\-IID \(β=1\.0\\beta=1\.0\) partition\.Figure[4](https://arxiv.org/html/2607.29071#S5.F4)and Figure[5](https://arxiv.org/html/2607.29071#S5.F5)report the training loss of FedSLM clients at compression ratios 0\.4 and 0\.6 over 10 communication rounds on all eight training benchmarks under the non\-IID partition, for LLaMA\-2\-7B and LLaMA\-2\-13B, respectively\.
Across all benchmarks, both compression variants exhibit smooth and monotonic convergence without divergence or oscillation, indicating that the adapter\-level aggregation in FedSLM remains stable under heterogeneous model capacities and data distributions\. This is a desirable property for practical deployment, where clients with different hardware constraints must coexist within the same federation\.
Comparing the two compression ratios, the 0\.6\-ratio model consistently converges to lower final loss values than the 0\.4\-ratio model, which is expected given its greater representational capacity\. However, the gap between the two curves narrows as training progresses on most benchmarks, suggesting that continued federated aggregation partially compensates for the capacity disparity\. Scaling from 7B to 13B preserves this convergence pattern: the 13B model achieves lower absolute loss values across all benchmarks while maintaining the same stable convergence trajectory, confirming that FedSLM’s adapter\-level aggregation scales gracefully with model size\. The consistent behavior across both scales validates the theoretical𝒪\(T1−1/2\)\\mathcal\{O\}\(T\_\{1\}^\{\-1/2\}\)convergence rate established in Theorem[1](https://arxiv.org/html/2607.29071#Thmtheorem1), which is independent of the underlying model dimension\.
TABLE IV:Client\-side model accuracy \(%\) after Stage 1 federated aggregation for both the LLaMA\-2\-7B and LLaMA\-2\-13B server models\. The upper sub\-block of each server model reports pre\-trained baselines without any federated training\. The best federated result in each column within each server\-model block isboldfaced\.TABLE V:Effect of adapter placement on server\-side accuracy \(%\)\. The best result in each row isboldfaced\.
### V\-DClient\-Side Model Performance
Table[IV](https://arxiv.org/html/2607.29071#S5.T4)reports client\-side accuracy after Stage 1 adapter\-level aggregation on both server models, alongside the pre\-trained baselines without federated training\.
On LLaMA\-2\-7B, federated aggregation yields consistent improvements\. Without training, SVD\-0\.4 and SVD\-0\.6 average 46\.3% and 51\.9% respectively, reflecting the information loss from low\-rank decomposition\. After 10 rounds, FedSLM\-0\.4 reaches 62\.2% \(IID\) and 59\.7% \(non\-IID\), while FedSLM\-0\.6 reaches 64\.0% and 61\.5%\. Both variants surpass the zero\-shot LLaMA\-2\-7B baseline \(57\.3%\) despite containing fewer parameters, with the largest FedSLM\-0\.6 vs\. FedSLM\-0\.4 gap appearing on ARC\_c \(46\.5% vs\. 40\.5% under IID\)\. On the medical benchmarks, pre\-trained baselines cluster around 55% on PubMedQA regardless of model size, while federated training lifts FedSLM\-0\.6 to 68\.8% \(IID\) and MedMCQA from 30\.6% \(pre\-trained\) to 32\.9% \(non\-IID\)\. We note that this comparison is between task\-specific adapter\-tuned models and a zero\-shot pre\-trained model; the improvement therefore reflects the combined effect of federated task adaptation and the adapter’s ability to recover capacity lost during compression, rather than a claim that compression is free\. Nevertheless, the result demonstrates that SVD\-compressed models are effective knowledge carriers in the federated setting: even at 40% compression, adapter training recovers and exceeds the zero\-shot performance of the uncompressed model\.
On LLaMA\-2\-13B, the client\-side trend scales consistently\. On the seven commonsense benchmarks, FedSLM\-0\.6 reaches 71\.2% \(IID\) and 68\.2% \(non\-IID\), exceeding FedSLM\-0\.4 \(63\.7% / 60\.9%\) by 7\.5% / 7\.3%\. Both compressed variants surpass the zero\-shot LLaMA\-2\-13B baseline \(63\.0%\) despite containing fewer parameters, with FedSLM\-0\.6 closing the gap most sharply on reasoning\-heavy tasks such as HellaSwag \(63\.6% vs\. 54\.8% zero\-shot\), ARC\_c \(49\.2% vs\. 43\.8%\), and WinoGrande \(80\.1% vs\. 50\.7%\)\. Across both scales, the non\-IID partition introduces a visible drop on WinoGrande \(7B FedSLM\-0\.4: 74\.7%→\\to56\.3%; 13B FedSLM\-0\.6: 80\.1%→\\to60\.1%\), which relies on fine\-grained coreference resolution sensitive to distributional skew, whereas benchmarks such as COPA and PIQA remain relatively stable\. Together these results confirm that SVD\-compressed models can achieve strong client\-side performance through federated adapter training at both scales, providing a reliable weak supervision signal for Stage 2 that explains the effectiveness of the subsequent weak\-to\-strong elicitation reported in Tables[II](https://arxiv.org/html/2607.29071#S5.T2)and[III](https://arxiv.org/html/2607.29071#S5.T3)\.
### V\-EAdapter Placement Ablation
To understand which weight matrices benefit most from adapter\-level fine\-tuning, we vary the set of linear layers to which adapters are attached while keeping all other hyperparameters fixed\. We consider four configurations:Atten\-QVadapts only the query and value projections;Atten\-Allextends to all attention projections \(Q, K, V, O\);FFN\-Onlytargets the feed\-forward sub\-layers \(gate, up, down projections\); andAlladapts all linear layers per transformer block\. For SVD\-compressed clients, the configuration name indicates which projection matrices receive low\-rank adapters; for the LLaMA clients, the corresponding LoRA modules target the same layers\.
Table[V](https://arxiv.org/html/2607.29071#S5.T5)summarizes the results\. Adapting only the attention query and value projections \(Atten\-QV\) yields the lowest average accuracy \(56\.0%\), and extending to all four attention projections \(Atten\-All\) provides only a modest improvement of 1\.5%\. In contrast, targeting the feed\-forward sub\-layers alone \(FFN\-Only\) raises the average to 61\.6%, a gain of 5\.6% overAtten\-QV\. This gap suggests that the feed\-forward network stores a larger share of the task\-relevant knowledge than the attention mechanism, consistent with recent findings on knowledge localization in transformer models\[[15](https://arxiv.org/html/2607.29071#bib.bib72),[39](https://arxiv.org/html/2607.29071#bib.bib73)\]\.
TABLE VI:Performance on vision–language benchmarks \(accuracy %\)\. SVD\-0\.6, SVD\-0\.9, and Full\-model are pre\-trained baselines without federated training\. The best result isboldfaced\.The best performance is achieved by theAllconfiguration \(63\.4%\), which adapts all linear layers per transformer block\. The incremental gain of 1\.8% overFFN\-Onlyindicates that attention projections still contribute complementary information once the feed\-forward layers are already adapted\. The improvement is most pronounced on ARC\_c \(\+4\.6%\) and HellaS \(\+3\.5%\), both of which require multi\-hop reasoning that benefits from jointly adapting the attention routing and the knowledge retrieval pathways\.
### V\-FVision–Language Experiments
To evaluate FedSLM beyond language\-only tasks, we apply it to LLaVA\-NeXT\-7B\[[33](https://arxiv.org/html/2607.29071#bib.bib56)\]on two vision–language benchmarks: ScienceQA\[[34](https://arxiv.org/html/2607.29071#bib.bib67)\]and VizWiz\[[16](https://arxiv.org/html/2607.29071#bib.bib68)\]\. Client\-side compressed models are obtained via QSVD\[[55](https://arxiv.org/html/2607.29071#bib.bib54)\]at compression ratios 0\.6 and 0\.9\. We simulate a federation of 10 clients\.
Table[VI](https://arxiv.org/html/2607.29071#S5.T6)reports the results\. The pre\-trained baselines reveal a clear capacity hierarchy: SVD\-0\.6 scores 57\.5% on ScienceQA and 44\.8% on VizWiz, while the full LLaVA\-NeXT\-7B reaches 66\.2% and 58\.4%\. FedSLM \(α=0\.5\\alpha\\\!=\\\!0\.5\) achieves 71\.8% on ScienceQA and 70\.1% on VizWiz under IID, surpassing the full LLaVA\-NeXT\-7B by 5\.6% and 11\.7% respectively, confirming that federated adapter training effectively recovers and exceeds the capacity lost to compression\. Under non\-IID, FedSLM \(α=0\.5\\alpha\\\!=\\\!0\.5\) attains 72\.7% on ScienceQA and 69\.9% on VizWiz\. Notably, the non\-IID results on ScienceQA are slightly higher than IID, likely because the distributional diversity across clients provides complementary visual reasoning patterns that benefit the server model after aggregation\.
These results demonstrate that FedSLM generalizes beyond language tasks to the multimodal setting, where the SVD\-compressed vision–language models serve as effective federated participants despite operating at a fraction of the full model’s capacity\.
## VIConclusion and Future Work
We presented FedSLM, an FL framework that enables clients with heterogeneous SVD\-compressed models to collaboratively improve a server\-side foundation model\. By integrating SVD\-based model derivation, two\-stage heterogeneous aggregation, and weak\-to\-strong elicitation with an auxiliary confidence loss, FedSLM achieves theoretical guarantees on convergence, subspace alignment, and artifact mitigation\. Experiments on natural language and vision–language benchmarks show consistent improvements over existing baselines under both IID and non\-IID partitions, with client\-side memory reduced to roughly half that of the full model\.
Several directions remain for future work, including adaptive compression\-ratio selection based on client resource availability, extension to other compression paradigms such as pruning and quantization, and integration of secure aggregation for deployment in regulated domains\.
## References
- \[1\]J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya, F\. L\. Aleman, D\. Almeida, J\. Altenschmidt, S\. Altman, S\. Anadkat,et al\.\(2023\)Gpt\-4 technical report\.arXiv preprint arXiv:2303\.08774\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p1.1)\.
- \[2\]\(2022\)FedRolex: model\-heterogeneous federated learning with rolling sub\-model extraction\.InNeurIPS,Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p4.1),[§II\-B](https://arxiv.org/html/2607.29071#S2.SS2.p1.1)\.
- \[3\]J\. Bai, D\. Chen, B\. Qian, L\. Yao, and Y\. Li\(2024\)Federated fine\-tuning of large language models under heterogeneous tasks and client resources\.InNeurIPS,Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p3.1),[§II\-C](https://arxiv.org/html/2607.29071#S2.SS3.p1.3),[§V](https://arxiv.org/html/2607.29071#S5.p5.2)\.
- \[4\]Y\. Bisk, R\. Zellers, J\. Gao, Y\. Choi,et al\.\(2020\)Piqa: reasoning about physical commonsense in natural language\.InAAAI,Vol\.34,pp\. 7432–7439\.Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p4.1)\.
- \[5\]N\. Boizard, K\. E\. Haddad, C\. Hudelot, and P\. Colombo\(2024\)Towards cross\-tokenizer distillation: the universal logit distillation loss for llms\.arXiv preprint arXiv:2402\.12030\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p5.1),[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p2.1)\.
- \[6\]C\. Burns, P\. Izmailov, J\. H\. Kirchner, B\. Baker, L\. Gao, L\. Aschenbrenner, Y\. Chen, A\. Ecoffet, M\. Joglekar, J\. Leike, I\. Sutskever, and J\. Wu\(2024\)Weak\-to\-strong generalization: eliciting strong capabilities with weak supervision\.InICML,Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p5.1),[§I](https://arxiv.org/html/2607.29071#S1.p7.1),[§II\-D](https://arxiv.org/html/2607.29071#S2.SS4.p1.5),[§III\-A](https://arxiv.org/html/2607.29071#S3.SS1.p1.27),[§III\-D](https://arxiv.org/html/2607.29071#S3.SS4.p3.1),[§IV\-C](https://arxiv.org/html/2607.29071#S4.SS3.p1.2)\.
- \[7\]S\. Cheng, J\. Wu, Y\. Xiao, and Y\. Liu\(2021\)Fedgems: federated learning of larger server models via selective knowledge fusion\.arXiv preprint arXiv:2110\.11027\.Cited by:[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p1.1)\.
- \[8\]X\. Cheng, Z\. Rao, Y\. Chen, and Q\. Zhang\(2020\)Explaining knowledge distillation by quantifying the knowledge\.InCVPR,pp\. 12922–12932\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p5.1),[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p2.1)\.
- \[9\]Y\. J\. Cho, L\. Liu, Z\. Xu, A\. Fahrezi, and G\. Joshi\(2024\)Heterogeneous lora for federated fine\-tuning of on\-device foundation models\.InEMNLP,pp\. 12903–12913\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p3.1),[§II\-C](https://arxiv.org/html/2607.29071#S2.SS3.p1.3),[§V](https://arxiv.org/html/2607.29071#S5.p5.2)\.
- \[10\]P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. Tafjord\(2018\)Think you have solved question answering? try arc, the ai2 reasoning challenge\.arXiv preprint arXiv:1803\.05457\.Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p4.1)\.
- \[11\]Y\. Deng, Z\. Qiao, Y\. Zhang, Z\. Ma, Y\. Liu, and J\. Ren\(2025\)CrossLM: A data\-free collaborative fine\-tuning framework for large and small language models\.InMobiSys,pp\. 124–137\.Cited by:[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p1.1)\.
- \[12\]E\. Diao, J\. Ding, and V\. Tarokh\(2021\)HeteroFL: computation and communication efficient federated learning for heterogeneous clients\.InICLR,Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p4.1),[§II\-B](https://arxiv.org/html/2607.29071#S2.SS2.p1.1)\.
- \[13\]T\. Fan, Y\. Kang, G\. Ma, W\. Chen, W\. Wei, L\. Fan, and Q\. Yang\(2023\)FATE\-llm: a industrial grade federated learning framework for large language models\.arXiv preprint arXiv:2310\.10049\.Cited by:[§II\-C](https://arxiv.org/html/2607.29071#S2.SS3.p1.3)\.
- \[14\]T\. Fan, G\. Ma, Y\. Kang, H\. Gu, Y\. Song, L\. Fan, K\. Chen, and Q\. Yang\(2025\)FedMKT: federated mutual knowledge transfer for large and small language models\.InCOLING,pp\. 243–255\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p5.1),[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p1.1),[§V](https://arxiv.org/html/2607.29071#S5.p5.2)\.
- \[15\]M\. Geva, R\. Schuster, J\. Berant, and O\. Levy\(2021\)Transformer feed\-forward layers are key\-value memories\.InEMNLP,pp\. 5484–5495\.Cited by:[§V\-E](https://arxiv.org/html/2607.29071#S5.SS5.p2.1)\.
- \[16\]D\. Gurari, Q\. Li, A\. J\. Stangl, A\. Guo, C\. Lin, K\. Grauman, J\. Luo, and J\. P\. Bigham\(2018\)VizWiz grand challenge: answering visual questions from blind people\.InCVPR,pp\. 3608–3617\.Cited by:[§V\-F](https://arxiv.org/html/2607.29071#S5.SS6.p1.1),[§V](https://arxiv.org/html/2607.29071#S5.p4.1)\.
- \[17\]F\. Haddadpour and M\. Mahdavi\(2019\)On the convergence of local descent methods in federated learning\.arXiv preprint arXiv:1910\.14425\.Cited by:[§IV\-B](https://arxiv.org/html/2607.29071#S4.SS2.p9.4)\.
- \[18\]C\. He, M\. Annavaram, and S\. Avestimehr\(2020\)Group knowledge transfer: federated learning of large cnns at the edge\.InNeurIPS,Cited by:[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p1.1)\.
- \[19\]G\. Hinton, O\. Vinyals, and J\. Dean\(2015\)Distilling the knowledge in a neural network\.arXiv preprint arXiv:1503\.02531\.Cited by:[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p1.1)\.
- \[20\]S\. Horváth, S\. Laskaridis, M\. Almeida, I\. Leontiadis, S\. I\. Venieris, and N\. D\. Lane\(2021\)FjORD: fair and accurate federated learning under heterogeneous targets with ordered dropout\.InNeurIPS,pp\. 12876–12889\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p4.1),[§II\-B](https://arxiv.org/html/2607.29071#S2.SS2.p1.1)\.
- \[21\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2022\)LoRA: low\-rank adaptation of large language models\.InICLR,Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p3.1),[§II\-C](https://arxiv.org/html/2607.29071#S2.SS3.p1.3),[§V](https://arxiv.org/html/2607.29071#S5.p3.2),[§V](https://arxiv.org/html/2607.29071#S5.p5.2)\.
- \[22\]F\. Ilhan, G\. Su, and L\. Liu\(2023\)ScaleFL: resource\-adaptive federated learning with heterogeneous clients\.InCVPR,pp\. 24532–24541\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p4.1),[§II\-B](https://arxiv.org/html/2607.29071#S2.SS2.p1.1)\.
- \[23\]Q\. Jin, B\. Dhingra, Z\. Liu, W\. W\. Cohen, and X\. Lu\(2019\)PubMedQA: A dataset for biomedical research question answering\.InEMNLP\-IJCNLP,pp\. 2567–2577\.Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p4.1)\.
- \[24\]H\. Karimi, J\. Nutini, and M\. Schmidt\(2016\)Linear convergence of gradient and proximal\-gradient methods under the polyak\-Łojasiewicz condition\.InECML PKDD,pp\. 795–811\.Cited by:[§A\-A3](https://arxiv.org/html/2607.29071#A1.SS1.SSS3.p1.8),[§IV\-B](https://arxiv.org/html/2607.29071#S4.SS2.p9.4),[Lemma 5](https://arxiv.org/html/2607.29071#Thmlemma5)\.
- \[25\]S\. P\. Karimireddy, S\. Kale, M\. Mohri, S\. Reddi, S\. Stich, and A\. T\. Suresh\(2020\)SCAFFOLD: stochastic controlled averaging for federated learning\.InICML,pp\. 5132–5143\.Cited by:[§A\-B1](https://arxiv.org/html/2607.29071#A1.SS2.SSS1.3.p1.6),[§IV\-A](https://arxiv.org/html/2607.29071#S4.SS1.p2.18)\.
- \[26\]M\. Kim, S\. Yu, S\. Kim, and S\. Moon\(2023\)DepthFL : depthwise federated learning for heterogeneous clients\.InICLR,Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p4.1),[§II\-B](https://arxiv.org/html/2607.29071#S2.SS2.p1.1)\.
- \[27\]A\. Koloskova, N\. Loizou, S\. Boreiri, M\. Jaggi, and S\. U\. Stich\(2020\)A unified theory of decentralized SGD with changing topology and local updates\.InICML,pp\. 5381–5393\.Cited by:[§A\-B1](https://arxiv.org/html/2607.29071#A1.SS2.SSS1.3.p1.6)\.
- \[28\]W\. Kuang, B\. Qian, Z\. Li, D\. Chen, D\. Gao, X\. Pan, Y\. Xie, Y\. Li, B\. Ding, and J\. Zhou\(2024\)FederatedScope\-llm: A comprehensive package for fine\-tuning large language models in federated learning\.InSIGKDD,pp\. 5260–5271\.Cited by:[§II\-C](https://arxiv.org/html/2607.29071#S2.SS3.p1.3)\.
- \[29\]C\. Li, J\. Zhang, and C\. Zong\(2025\)TokAlign: efficient vocabulary adaptation via token alignment\.InACL,pp\. 4109–4126\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p5.1),[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p2.1)\.
- \[30\]M\. Li, F\. Zhou, and X\. Song\(2025\)BiLD: bi\-directional logits difference loss for large language model distillation\.InCOLING,pp\. 1168–1182\.Cited by:[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p2.1)\.
- \[31\]A\. Liu, B\. Feng, B\. Xue, B\. Wang, B\. Wu, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan,et al\.\(2024\)Deepseek\-v3 technical report\.arXiv preprint arXiv:2412\.19437\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p1.1)\.
- \[32\]C\. Liu, L\. Zhu, and M\. Belkin\(2022\)Loss landscapes and optimization in over\-parameterized non\-linear systems and neural networks\.Appl\. Comput\. Harmon\. A\.59,pp\. 85–116\.Cited by:[§A\-E](https://arxiv.org/html/2607.29071#A1.SS5.p7.20),[§IV\-B](https://arxiv.org/html/2607.29071#S4.SS2.p9.4)\.
- \[33\]H\. Liu, C\. Li, Y\. Li, B\. Li, Y\. Zhang, S\. Shen, and Y\. J\. Lee\(2024\)Llava\-next: improved reasoning, ocr, and world knowledge\.arXiv preprint\.Cited by:[§V\-F](https://arxiv.org/html/2607.29071#S5.SS6.p1.1),[§V](https://arxiv.org/html/2607.29071#S5.p2.2)\.
- \[34\]P\. Lu, S\. Mishra, T\. Xia, L\. Qiu, K\. Chang, S\. Zhu, O\. Tafjord, P\. Clark, and A\. Kalyan\(2022\)Learn to explain: multimodal reasoning via thought chains for science question answering\.InNeurIPS,Cited by:[§V\-F](https://arxiv.org/html/2607.29071#S5.SS6.p1.1),[§V](https://arxiv.org/html/2607.29071#S5.p4.1)\.
- \[35\]G\. Malinovsky, U\. Michieli, H\. A\. A\. K\. Hammoud, T\. Ceritli, H\. Elesedy, M\. Ozay, and P\. Richtárik\(2024\)Randomized asymmetric chain of lora: the first meaningful theoretical framework for low\-rank adaptation\.arXiv preprint arXiv:2410\.08305\.Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p5.2)\.
- \[36\]J\. Martens and R\. B\. Grosse\(2015\)Optimizing neural networks with kronecker\-factored approximate curvature\.InICML,Vol\.37,pp\. 2408–2417\.Cited by:[§A\-E](https://arxiv.org/html/2607.29071#A1.SS5.p5.4)\.
- \[37\]B\. McMahan, E\. Moore, D\. Ramage, S\. Hampson, and B\. A\. y Arcas\(2017\)Communication\-efficient learning of deep networks from decentralized data\.InAISTATS,pp\. 1273–1282\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p2.1),[§III\-A](https://arxiv.org/html/2607.29071#S3.SS1.p2.2),[§III\-C](https://arxiv.org/html/2607.29071#S3.SS3.p2.4),[§IV\-A](https://arxiv.org/html/2607.29071#S4.SS1.p1.14),[§V](https://arxiv.org/html/2607.29071#S5.p5.2)\.
- \[38\]\(2023\)Medical\-flashcards\.Note:[https://huggingface\.co/datasets/medalpaca/medical\_meadow\_medical\_flashcards](https://huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards)Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p4.1)\.
- \[39\]K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov\(2022\)Locating and editing factual associations in GPT\.InNeurIPS,Cited by:[§V\-E](https://arxiv.org/html/2607.29071#S5.SS5.p2.1)\.
- \[40\]B\. Minixhofer, I\. Vulić, and E\. M\. Ponti\(2025\)Universal cross\-tokenizer distillation via approximate likelihood matching\.InNeurIPS,Vol\.38,pp\. 79297–79326\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p5.1),[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p2.1)\.
- \[41\]M\. Moor, O\. Banerjee, Z\. S\. H\. Abad, H\. M\. Krumholz, J\. Leskovec, E\. J\. Topol, and P\. Rajpurkar\(2023\)Foundation models for generalist medical artificial intelligence\.Nature616\(7956\),pp\. 259–265\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p1.1)\.
- \[42\]M\. Morafah, V\. Kungurtsev, H\. Chang, C\. Chen, and B\. Lin\(2024\)Towards diverse device heterogeneous federated learning via task arithmetic knowledge integration\.InNeurIPS,Cited by:[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p1.1)\.
- \[43\]A\. Pal, L\. K\. Umapathi, and M\. Sankarasubbu\(2022\)MedMCQA: A large\-scale multi\-subject multi\-choice dataset for medical domain question answering\.InCHIL,pp\. 248–260\.Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p4.1)\.
- \[44\]M\. Roemmele, C\. A\. Bejan, and A\. S\. Gordon\(2011\)Choice of plausible alternatives: an evaluation of commonsense causal reasoning\.InAAAI,Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p4.1)\.
- \[45\]K\. Sakaguchi, R\. L\. Bras, C\. Bhagavatula, and Y\. Choi\(2021\)WinoGrande: an adversarial winograd schema challenge at scale\.Commun\. ACM64\(9\),pp\. 99–106\.Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p4.1)\.
- \[46\]M\. Sap, H\. Rashkin, D\. Chen, R\. L\. Bras, and Y\. Choi\(2019\)Social iqa: commonsense reasoning about social interactions\.InEMNLP\-IJCNLP,pp\. 4462–4472\.Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p4.1)\.
- \[47\]T\. Sun, D\. Li, and B\. Wang\(2023\)Decentralized federated averaging\.IEEE Trans\. Pattern Anal\. Mach\. Intell\.45\(4\),pp\. 4289–4301\.Cited by:[§IV\-B](https://arxiv.org/html/2607.29071#S4.SS2.p9.4)\.
- \[48\]Y\. Sun, Z\. Li, Y\. Li, and B\. Ding\(2024\)Improving lora in privacy\-preserving federated learning\.InICLR,Cited by:[§II\-C](https://arxiv.org/html/2607.29071#S2.SS3.p1.3),[§V](https://arxiv.org/html/2607.29071#S5.p5.2)\.
- \[49\]Y\. Tan, G\. Long, L\. Liu, T\. Zhou, Q\. Lu, J\. Jiang, and C\. Zhang\(2022\)Fedproto: federated prototype learning across heterogeneous clients\.InAAAI,Vol\.36,pp\. 8432–8440\.Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p5.2)\.
- \[50\]A\. J\. Thirunavukarasu, D\. S\. J\. Ting, K\. Elangovan, L\. Gutierrez, T\. F\. Tan, and D\. S\. W\. Ting\(2023\)Large language models in medicine\.Nature medicine29\(8\),pp\. 1930–1940\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p1.1)\.
- \[51\]H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale,et al\.\(2023\)Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p3.1),[§V](https://arxiv.org/html/2607.29071#S5.p2.2)\.
- \[52\]Q\. Wang, J\. Ke, M\. Tomizuka, K\. Keutzer, and C\. Xu\(2025\)Dobi\-svd: differentiable svd for llm compression and some new perspectives\.InICLR,pp\. 12561–12590\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p7.1),[§II\-D](https://arxiv.org/html/2607.29071#S2.SS4.p1.5),[§III\-B](https://arxiv.org/html/2607.29071#S3.SS2.p2.14),[§V](https://arxiv.org/html/2607.29071#S5.p2.2)\.
- \[53\]X\. Wang, S\. Alam, Z\. Wan, H\. Shen, and M\. Zhang\(2025\)SVD\-LLM V2: optimizing singular value truncation for large language model compression\.InNAACL,pp\. 4287–4296\.Cited by:[§II\-D](https://arxiv.org/html/2607.29071#S2.SS4.p1.5),[§III\-B](https://arxiv.org/html/2607.29071#S3.SS2.p2.14)\.
- \[54\]X\. Wang, Y\. Zheng, Z\. Wan, and M\. Zhang\(2025\)SVD\-LLM: truncation\-aware singular value decomposition for large language model compression\.InICLR,Vol\.2025,pp\. 19299–19319\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p7.1),[§II\-D](https://arxiv.org/html/2607.29071#S2.SS4.p1.5),[§III\-B](https://arxiv.org/html/2607.29071#S3.SS2.p2.14),[§IV](https://arxiv.org/html/2607.29071#S4.p3.6),[§IV](https://arxiv.org/html/2607.29071#S4.p4.18)\.
- \[55\]Y\. Wang, H\. Wang, and S\. Q\. Zhang\(2025\)QSVD: efficient low\-rank approximation for unified query\-key\-value weight compression in low\-precision vision\-language models\.InNeurIPS,Vol\.38,pp\. 1789–1820\.Cited by:[§II\-D](https://arxiv.org/html/2607.29071#S2.SS4.p1.5),[§III\-B](https://arxiv.org/html/2607.29071#S3.SS2.p2.14),[§V\-F](https://arxiv.org/html/2607.29071#S5.SS6.p1.1),[§V](https://arxiv.org/html/2607.29071#S5.p2.2)\.
- \[56\]Z\. Wang, Z\. Shen, Y\. He, G\. Sun, H\. Wang, L\. Lyu, and A\. Li\(2024\)FLoRA: federated fine\-tuning large language models with heterogeneous low\-rank adaptations\.InNeurIPS,Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p3.1),[§II\-C](https://arxiv.org/html/2607.29071#S2.SS3.p1.3)\.
- \[57\]D\. Wen, K\. J\. Jeon, and K\. Huang\(2022\)Federated dropout \- A simple approach for enabling federated learning on resource constrained devices\.IEEE Wirel\. Commun\. Lett\.11\(5\),pp\. 923–927\.Cited by:[§II\-B](https://arxiv.org/html/2607.29071#S2.SS2.p1.1)\.
- \[58\]C\. Wu, F\. Wu, L\. Lyu, Y\. Huang, and X\. Xie\(2022\)Communication\-efficient federated learning via knowledge distillation\.Nature communications13\(1\),pp\. 2032\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p5.1),[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p1.1)\.
- \[59\]F\. Wu, Z\. Li, Y\. Li, B\. Ding, and J\. Gao\(2024\)Fedbiot: llm local fine\-tuning in federated learning without full model\.InKDD,pp\. 3345–3355\.Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p5.2)\.
- \[60\]S\. Wu, O\. Irsoy, S\. Lu, V\. Dabravolski, M\. Dredze, S\. Gehrmann, P\. Kambadur, D\. Rosenberg, and G\. Mann\(2023\)Bloomberggpt: a large language model for finance\.arXiv preprint arXiv:2303\.17564\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p1.1)\.
- \[61\]A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p1.1)\.
- \[62\]Q\. Yang, Y\. Liu, T\. Chen, and Y\. Tong\(2019\)Federated machine learning: concept and applications\.ACM Transactions on Intelligent Systems and Technology10\(2\),pp\. 1–19\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p2.1)\.
- \[63\]R\. Ye, W\. Wang, J\. Chai, D\. Li, Z\. Li, Y\. Xu, Y\. Du, Y\. Wang, and S\. Chen\(2024\)OpenFedLLM: training large language models on decentralized private data via federated learning\.InKDD,pp\. 6137–6147\.Cited by:[§II\-C](https://arxiv.org/html/2607.29071#S2.SS3.p1.3)\.
- \[64\]B\. Ying, Z\. Li, and H\. Yang\(2025\)Exact and linear convergence for federated learning under arbitrary client participation is attainable\.InNeurIPS,Vol\.38,pp\. 40156–40201\.Cited by:[§IV\-B](https://arxiv.org/html/2607.29071#S4.SS2.p9.4)\.
- \[65\]Z\. Yuan, Y\. Shang, Y\. Song, D\. Yang, Q\. Wu, Y\. Yan, and G\. Sun\(2023\)ASVD: activation\-aware singular value decomposition for compressing large language models\.arXiv preprint arXiv:2312\.05821\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p7.1),[§II\-D](https://arxiv.org/html/2607.29071#S2.SS4.p1.5),[§III\-B](https://arxiv.org/html/2607.29071#S3.SS2.p2.14),[§IV](https://arxiv.org/html/2607.29071#S4.p3.6)\.
- \[66\]R\. Zellers, A\. Holtzman, Y\. Bisk, A\. Farhadi, and Y\. Choi\(2019\)HellaSwag: can a machine really finish your sentence?\.InACL,pp\. 4791–4800\.Cited by:[§V](https://arxiv.org/html/2607.29071#S5.p4.1)\.
- \[67\]J\. Zhang, Y\. Liu, J\. Fu, Y\. Hua, T\. Zou, J\. Cao, and Q\. Yang\(2025\)PCEvolve: private contrastive evolution for synthetic dataset generation via few\-shot private data and generative APIs\.InICML,Vol\.267,pp\. 75575–75590\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p5.1)\.
- \[68\]J\. Zhang, S\. Vahidian, M\. Kuo, C\. Li, R\. Zhang, T\. Yu, G\. Wang, and Y\. Chen\(2024\)Towards building the federatedgpt: federated instruction tuning\.InICASSP,pp\. 6915–6919\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p3.1),[§II\-C](https://arxiv.org/html/2607.29071#S2.SS3.p1.3)\.
- \[69\]S\. Zhang, X\. Zhang, Z\. Sun, Y\. Chen, and J\. Xu\(2024\)Dual\-space knowledge distillation for large language models\.InEMNLP,pp\. 18164–18181\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p5.1),[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p2.1)\.
- \[70\]Y\. Zhang, D\. Long, Z\. Li, and P\. Xie\(2023\)Text representation distillation via information bottleneck principle\.InEMNLP,pp\. 14372–14383\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p5.1),[§II\-A](https://arxiv.org/html/2607.29071#S2.SS1.p2.1)\.
- \[71\]Z\. Zhang, Y\. Yang, Y\. Dai, Q\. Wang, Y\. Yu, L\. Qu, and Z\. Xu\(2023\)FedPETuning: when federated learning meets the parameter\-efficient tuning methods of pre\-trained language models\.InACL,pp\. 9963–9977\.Cited by:[§I](https://arxiv.org/html/2607.29071#S1.p3.1),[§II\-C](https://arxiv.org/html/2607.29071#S2.SS3.p1.3)\.
## Appendix AProofs and Supporting Lemmas
Throughout the appendix we adopt the notation of the main text\. In particular,𝐔r∈ℝm×r\\mathbf\{U\}\_\{r\}\\in\\mathbb\{R\}^\{m\\times r\}and𝐕r∈ℝr×n\\mathbf\{V\}\_\{r\}\\in\\mathbb\{R\}^\{r\\times n\}denote the \(non\-orthonormal\) factors delivered to clients, as in \([13](https://arxiv.org/html/2607.29071#S4.E13)\)–\([15](https://arxiv.org/html/2607.29071#S4.E15)\), while𝐔r∘∈ℝm×r\\mathbf\{U\}\_\{r\}^\{\\circ\}\\in\\mathbb\{R\}^\{m\\times r\}and𝐕r∘∈ℝr×n\\mathbf\{V\}\_\{r\}^\{\\circ\}\\in\\mathbb\{R\}^\{r\\times n\}denote the orthonormal singular\-vector factors \(𝐔r∘⊤𝐔r∘=𝐈r\\mathbf\{U\}\_\{r\}^\{\\circ\\,\\top\}\\mathbf\{U\}\_\{r\}^\{\\circ\}=\\mathbf\{I\}\_\{r\},𝐕r∘𝐕r∘⊤=𝐈r\\mathbf\{V\}\_\{r\}^\{\\circ\}\\mathbf\{V\}\_\{r\}^\{\\circ\\,\\top\}=\\mathbf\{I\}\_\{r\}\)\. The retained spectrum is𝚲r=diag\(λ1,…,λr\)\\boldsymbol\{\\Lambda\}\_\{r\}=\\mathrm\{diag\}\(\\lambda\_\{1\},\\ldots,\\lambda\_\{r\}\), and𝐒∈ℝr×r\\mathbf\{S\}\\in\\mathbb\{R\}^\{r\\times r\}is the additional invertible scaling matrix, so that𝐔r=𝐔r∘𝚲r1/2𝐒−1\\mathbf\{U\}\_\{r\}=\\mathbf\{U\}\_\{r\}^\{\\circ\}\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\mathbf\{S\}^\{\-1\}and𝐕r=𝐒𝚲r1/2𝐕r∘\\mathbf\{V\}\_\{r\}=\\mathbf\{S\}\\,\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\mathbf\{V\}\_\{r\}^\{\\circ\}\(the balanced factorization is𝐒=𝐈r\\mathbf\{S\}=\\mathbf\{I\}\_\{r\}\)\. The condition numberκ\(𝐒\)=‖𝐒‖2‖𝐒−1‖2\\kappa\(\\mathbf\{S\}\)=\\\|\\mathbf\{S\}\\\|\_\{2\}\\,\\\|\\mathbf\{S\}^\{\-1\}\\\|\_\{2\}equals11in the balanced case\. The effective smoothness constantLeff=Lλ12κ\(𝐒\)2L\_\{\\mathrm\{eff\}\}=L\\,\\lambda\_\{1\}^\{2\}\\,\\kappa\(\\mathbf\{S\}\)^\{2\}is from Definition[1](https://arxiv.org/html/2607.29071#Thmdefinition1)\.
### A\-AMathematical Preliminaries
This subsection collects the standard tools invoked in the main analysis: the Moore–Penrose pseudoinverse \(used to define the geometric projection adapter𝚺^i\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}in Section[IV\-B](https://arxiv.org/html/2607.29071#S4.SS2)\), orthogonal projectors and their algebraic properties \(used throughout the subspace alignment theory\), and the PŁ condition together with its quadratic growth consequence \(used in Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\)\. Readers familiar with these tools may skip ahead\.
#### A\-A1Moore–Penrose pseudoinverse
For a column\-full\-rank matrix𝐀∈ℝm×r\\mathbf\{A\}\\in\\mathbb\{R\}^\{m\\times r\}withm≥rm\\geq r, the Moore–Penrose pseudoinverse is defined by
𝐀†≜\(𝐀⊤𝐀\)−1𝐀⊤∈ℝr×m,\\mathbf\{A\}^\{\\dagger\}\\;\\triangleq\\;\(\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\)^\{\-1\}\\mathbf\{A\}^\{\\top\}\\;\\in\\;\\mathbb\{R\}^\{r\\times m\},\(39\)where𝐀⊤𝐀\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}is invertible precisely because𝐀\\mathbf\{A\}has full column rank\. When𝐀\\mathbf\{A\}is itself square and invertible,𝐀†\\mathbf\{A\}^\{\\dagger\}coincides with the ordinary inverse𝐀−1\\mathbf\{A\}^\{\-1\}\. For a row\-full\-rank matrix𝐁∈ℝr×n\\mathbf\{B\}\\in\\mathbb\{R\}^\{r\\times n\}withn≥rn\\geq r, the analogous formula is𝐁†=𝐁⊤\(𝐁𝐁⊤\)−1\\mathbf\{B\}^\{\\dagger\}=\\mathbf\{B\}^\{\\top\}\(\\mathbf\{B\}\\mathbf\{B\}^\{\\top\}\)^\{\-1\}\.
The pseudoinverse is the matrix form of the least\-squares\-solution operator\. For any𝐛∈ℝm\\mathbf\{b\}\\in\\mathbb\{R\}^\{m\}, the unique minimizer of‖𝐀𝐱−𝐛‖22\\\|\\mathbf\{A\}\\mathbf\{x\}\-\\mathbf\{b\}\\\|\_\{2\}^\{\\,2\}over𝐱∈ℝr\\mathbf\{x\}\\in\\mathbb\{R\}^\{r\}satisfies the normal equations𝐀⊤𝐀𝐱=𝐀⊤𝐛\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\\mathbf\{x\}=\\mathbf\{A\}^\{\\top\}\\mathbf\{b\}, yielding𝐱∗=𝐀†𝐛\\mathbf\{x\}^\{\*\}=\\mathbf\{A\}^\{\\dagger\}\\mathbf\{b\}\. The corresponding fitted value𝐀𝐀†𝐛\\mathbf\{A\}\\mathbf\{A\}^\{\\dagger\}\\mathbf\{b\}has a geometric interpretation as an orthogonal projection, which we develop after introducing the projector formalism in Appendix[A\-A2](https://arxiv.org/html/2607.29071#A1.SS1.SSS2)below\.
#### A\-A2Orthogonal projectors
We begin with the formal definition\.
###### Definition 5\(Orthogonal projector\)\.
A matrix𝐏∈ℝm×m\\mathbf\{P\}\\in\\mathbb\{R\}^\{m\\times m\}is an*orthogonal projector*onto a subspaceℳ⊆ℝm\\mathcal\{M\}\\subseteq\\mathbb\{R\}^\{m\}if it satisfies the following three equivalent properties:
1. \(a\)𝐏2=𝐏\\mathbf\{P\}^\{2\}=\\mathbf\{P\},𝐏⊤=𝐏\\mathbf\{P\}^\{\\top\}=\\mathbf\{P\}, andrange\(𝐏\)=ℳ\\mathrm\{range\}\(\\mathbf\{P\}\)=\\mathcal\{M\};
2. \(b\)𝐏\\mathbf\{P\}acts as the identity onℳ\\mathcal\{M\}and as zero on the orthogonal complementℳ⟂\\mathcal\{M\}^\{\\perp\};
3. \(c\)for every𝐯∈ℝm\\mathbf\{v\}\\in\\mathbb\{R\}^\{m\},𝐏𝐯\\mathbf\{P\}\\mathbf\{v\}is the unique element ofℳ\\mathcal\{M\}minimizing‖𝐯−𝐰‖2\\\|\\mathbf\{v\}\-\\mathbf\{w\}\\\|\_\{2\}over𝐰∈ℳ\\mathbf\{w\}\\in\\mathcal\{M\}\.
Property \(c\) in Definition[5](https://arxiv.org/html/2607.29071#Thmdefinition5)shows that an orthogonal projector is determined by the subspace alone, independently of any particular basis chosen to represent it\. A constructive realization uses the Moore–Penrose pseudoinverse \([39](https://arxiv.org/html/2607.29071#A1.E39)\): for any matrix whose columns span the target subspace, the product𝐀𝐀†\\mathbf\{A\}\\mathbf\{A\}^\{\\dagger\}delivers the projector explicitly\. Concretely, the fitted value𝐛^=𝐀𝐀†𝐛\\hat\{\\mathbf\{b\}\}=\\mathbf\{A\}\\mathbf\{A\}^\{\\dagger\}\\mathbf\{b\}is the element of the column spacecol\(𝐀\)≜\{𝐀𝐱:𝐱∈ℝr\}\\mathrm\{col\}\(\\mathbf\{A\}\)\\triangleq\\\{\\mathbf\{A\}\\mathbf\{x\}:\\mathbf\{x\}\\in\\mathbb\{R\}^\{r\}\\\}closest to𝐛\\mathbf\{b\}in Euclidean norm, so the map𝐛↦𝐀𝐀†𝐛\\mathbf\{b\}\\mapsto\\mathbf\{A\}\\mathbf\{A\}^\{\\dagger\}\\mathbf\{b\}acts as an orthogonal projection ontocol\(𝐀\)\\mathrm\{col\}\(\\mathbf\{A\}\)\. The following lemma makes this precise\.
###### Lemma 2\(Pseudoinverse–projector identity\)\.
For a column\-full\-rank𝐀∈ℝm×r\\mathbf\{A\}\\in\\mathbb\{R\}^\{m\\times r\}, the matrix𝐀𝐀†\\mathbf\{A\}\\mathbf\{A\}^\{\\dagger\}is the orthogonal projector ontocol\(𝐀\)\\mathrm\{col\}\(\\mathbf\{A\}\)\.
###### Proof\.
We verify the three properties of Definition[5](https://arxiv.org/html/2607.29071#Thmdefinition5)\(a\): idempotence, symmetry, and the fixing property\.
*Idempotence:*
\(𝐀𝐀†\)2=𝐀\(𝐀⊤𝐀\)−1\(𝐀⊤𝐀\)\(𝐀⊤𝐀\)−1𝐀⊤=𝐀\(𝐀⊤𝐀\)−1𝐀⊤=𝐀𝐀†\.\\begin\{split\}\(\\mathbf\{A\}\\mathbf\{A\}^\{\\dagger\}\)^\{2\}&=\\mathbf\{A\}\(\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\)^\{\-1\}\(\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\)\(\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\)^\{\-1\}\\mathbf\{A\}^\{\\top\}\\\\ &=\\mathbf\{A\}\(\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\)^\{\-1\}\\mathbf\{A\}^\{\\top\}=\\mathbf\{A\}\\mathbf\{A\}^\{\\dagger\}\.\\end\{split\}
*Symmetry:*
\(𝐀𝐀†\)⊤=𝐀\(\(𝐀⊤𝐀\)−1\)⊤𝐀⊤=𝐀\(𝐀⊤𝐀\)−1𝐀⊤=𝐀𝐀†,\\begin\{split\}\(\\mathbf\{A\}\\mathbf\{A\}^\{\\dagger\}\)^\{\\top\}&=\\mathbf\{A\}\(\(\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\)^\{\-1\}\)^\{\\top\}\\mathbf\{A\}^\{\\top\}\\\\ &=\\mathbf\{A\}\(\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\)^\{\-1\}\\mathbf\{A\}^\{\\top\}=\\mathbf\{A\}\\mathbf\{A\}^\{\\dagger\},\\end\{split\}since\(𝐀⊤𝐀\)−1\(\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\)^\{\-1\}inherits symmetry from𝐀⊤𝐀\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\.
*Fixingcol\(𝐀\)\\mathrm\{col\}\(\\mathbf\{A\}\):*for𝐯=𝐀𝐱\\mathbf\{v\}=\\mathbf\{A\}\\mathbf\{x\},
𝐀𝐀†𝐯=𝐀\(𝐀⊤𝐀\)−1\(𝐀⊤𝐀\)𝐱=𝐀𝐱=𝐯\.\\begin\{split\}\\mathbf\{A\}\\mathbf\{A\}^\{\\dagger\}\\mathbf\{v\}&=\\mathbf\{A\}\(\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\)^\{\-1\}\(\\mathbf\{A\}^\{\\top\}\\mathbf\{A\}\)\\mathbf\{x\}\\\\ &=\\mathbf\{A\}\\mathbf\{x\}=\\mathbf\{v\}\.\\end\{split\}Idempotence and symmetry together imply that𝐀𝐀†\\mathbf\{A\}\\mathbf\{A\}^\{\\dagger\}is an orthogonal projector, and the fixing property identifies the projection range withcol\(𝐀\)\\mathrm\{col\}\(\\mathbf\{A\}\)\. ∎
In our setting, the rank\-rrsubspaces𝒰r⊆ℝm\\mathcal\{U\}\_\{r\}\\subseteq\\mathbb\{R\}^\{m\}and𝒱r⊆ℝn\\mathcal\{V\}\_\{r\}\\subseteq\\mathbb\{R\}^\{n\}\(Definition[2](https://arxiv.org/html/2607.29071#Thmdefinition2)\) have orthogonal projectors
𝐏r=𝐔r∘𝐔r∘⊤,𝐐r=𝐕r∘⊤𝐕r∘\.\\mathbf\{P\}\_\{r\}=\\mathbf\{U\}\_\{r\}^\{\\circ\}\\,\\mathbf\{U\}\_\{r\}^\{\\circ\\,\\top\},\\qquad\\mathbf\{Q\}\_\{r\}=\\mathbf\{V\}\_\{r\}^\{\\circ\\,\\top\}\\,\\mathbf\{V\}\_\{r\}^\{\\circ\}\.\(40\)
###### Lemma 3\(Pseudoinverse formulas for𝐏ri\\mathbf\{P\}\_\{r\_\{i\}\}and𝐐ri\\mathbf\{Q\}\_\{r\_\{i\}\}\)\.
For the deployed factors𝐔ri=𝐔ri∘𝐌U\\mathbf\{U\}\_\{r\_\{i\}\}=\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\circ\}\\mathbf\{M\}\_\{U\}and𝐕ri=𝐌V𝐕ri∘\\mathbf\{V\}\_\{r\_\{i\}\}=\\mathbf\{M\}\_\{V\}\\,\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\circ\}with𝐌U,𝐌V∈ℝri×ri\\mathbf\{M\}\_\{U\},\\mathbf\{M\}\_\{V\}\\in\\mathbb\{R\}^\{r\_\{i\}\\times r\_\{i\}\}invertible \(this covers both the balanced factorization with𝐌U=𝐌V=𝚲ri1/2\\mathbf\{M\}\_\{U\}=\\mathbf\{M\}\_\{V\}=\\boldsymbol\{\\Lambda\}\_\{r\_\{i\}\}^\{1/2\}and the whitened variant \([15](https://arxiv.org/html/2607.29071#S4.E15)\) with𝐌U=𝚲ri1/2𝐒−1\\mathbf\{M\}\_\{U\}=\\boldsymbol\{\\Lambda\}\_\{r\_\{i\}\}^\{1/2\}\\mathbf\{S\}^\{\-1\}and𝐌V=𝐒𝚲ri1/2\\mathbf\{M\}\_\{V\}=\\mathbf\{S\}\\,\\boldsymbol\{\\Lambda\}\_\{r\_\{i\}\}^\{1/2\}\), we have
𝐔ri𝐔ri†=𝐏ri,𝐕ri†𝐕ri=𝐐ri\.\\mathbf\{U\}\_\{r\_\{i\}\}\\,\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\dagger\}\\;=\\;\\mathbf\{P\}\_\{r\_\{i\}\},\\qquad\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\dagger\}\\,\\mathbf\{V\}\_\{r\_\{i\}\}\\;=\\;\\mathbf\{Q\}\_\{r\_\{i\}\}\.\(41\)
###### Proof\.
Since𝐔ri∘⊤𝐔ri∘=𝐈\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\circ\\,\\top\}\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\circ\}=\\mathbf\{I\}, we have𝐔ri⊤𝐔ri=𝐌U⊤𝐌U\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\top\}\\mathbf\{U\}\_\{r\_\{i\}\}=\\mathbf\{M\}\_\{U\}^\{\\top\}\\mathbf\{M\}\_\{U\}, hence\(𝐔ri⊤𝐔ri\)−1=𝐌U−1𝐌U−⊤\(\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\top\}\\mathbf\{U\}\_\{r\_\{i\}\}\)^\{\-1\}=\\mathbf\{M\}\_\{U\}^\{\-1\}\\mathbf\{M\}\_\{U\}^\{\-\\top\}and
𝐔ri†=\(𝐔ri⊤𝐔ri\)−1𝐔ri⊤=𝐌U−1𝐌U−⊤𝐌U⊤𝐔ri∘⊤=𝐌U−1𝐔ri∘⊤,\\begin\{split\}\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\dagger\}&=\(\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\top\}\\mathbf\{U\}\_\{r\_\{i\}\}\)^\{\-1\}\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\top\}\\\\ &=\\mathbf\{M\}\_\{U\}^\{\-1\}\\mathbf\{M\}\_\{U\}^\{\-\\top\}\\mathbf\{M\}\_\{U\}^\{\\top\}\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\circ\\,\\top\}=\\mathbf\{M\}\_\{U\}^\{\-1\}\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\circ\\,\\top\},\\end\{split\}\(42\)so that
𝐔ri𝐔ri†=𝐔ri∘𝐌U𝐌U−1𝐔ri∘⊤=𝐔ri∘𝐔ri∘⊤=𝐏ri\.\\mathbf\{U\}\_\{r\_\{i\}\}\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\dagger\}=\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\circ\}\\,\\mathbf\{M\}\_\{U\}\\mathbf\{M\}\_\{U\}^\{\-1\}\\,\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\circ\\,\\top\}=\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\circ\}\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\circ\\,\\top\}=\\mathbf\{P\}\_\{r\_\{i\}\}\.\(43\)The invertible scaling𝐌U\\mathbf\{M\}\_\{U\}cancels completely\. The verification for𝐐ri\\mathbf\{Q\}\_\{r\_\{i\}\}is analogous: writing𝐕ri=𝐌V𝐕ri∘\\mathbf\{V\}\_\{r\_\{i\}\}=\\mathbf\{M\}\_\{V\}\\,\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\circ\}, row\-fullness of𝐕ri\\mathbf\{V\}\_\{r\_\{i\}\}gives𝐕ri†=𝐕ri⊤\(𝐕ri𝐕ri⊤\)−1=𝐕ri∘⊤𝐌V−1\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\dagger\}=\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\top\}\(\\mathbf\{V\}\_\{r\_\{i\}\}\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\top\}\)^\{\-1\}=\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\circ\\,\\top\}\\mathbf\{M\}\_\{V\}^\{\-1\}, and𝐕ri†𝐕ri=𝐕ri∘⊤𝐌V−1𝐌V𝐕ri∘=𝐕ri∘⊤𝐕ri∘=𝐐ri\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\dagger\}\\mathbf\{V\}\_\{r\_\{i\}\}=\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\circ\\,\\top\}\\mathbf\{M\}\_\{V\}^\{\-1\}\\mathbf\{M\}\_\{V\}\\,\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\circ\}=\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\circ\\,\\top\}\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\circ\}=\\mathbf\{Q\}\_\{r\_\{i\}\}\. ∎
Lemma[3](https://arxiv.org/html/2607.29071#Thmlemma3)equates𝐔ri𝚺^i𝐕ri\\mathbf\{U\}\_\{r\_\{i\}\}\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\mathbf\{V\}\_\{r\_\{i\}\}with𝐏ri𝐖^𝐐ri\\mathbf\{P\}\_\{r\_\{i\}\}\\widehat\{\\mathbf\{W\}\}\\mathbf\{Q\}\_\{r\_\{i\}\}: substituting𝚺^i=𝐔ri†𝐖^𝐕ri†\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}=\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\dagger\}\\widehat\{\\mathbf\{W\}\}\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\dagger\}and applying the identities in Lemma[3](https://arxiv.org/html/2607.29071#Thmlemma3)gives
𝐔ri𝚺^i𝐕ri=𝐔ri𝐔ri†⏟=𝐏ri𝐖^𝐕ri†𝐕ri⏟=𝐐ri=𝐏ri𝐖^𝐐ri\.\\mathbf\{U\}\_\{r\_\{i\}\}\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\mathbf\{V\}\_\{r\_\{i\}\}=\\underbrace\{\\mathbf\{U\}\_\{r\_\{i\}\}\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\dagger\}\}\_\{=\\mathbf\{P\}\_\{r\_\{i\}\}\}\\widehat\{\\mathbf\{W\}\}\\underbrace\{\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\dagger\}\\mathbf\{V\}\_\{r\_\{i\}\}\}\_\{=\\mathbf\{Q\}\_\{r\_\{i\}\}\}=\\mathbf\{P\}\_\{r\_\{i\}\}\\widehat\{\\mathbf\{W\}\}\\mathbf\{Q\}\_\{r\_\{i\}\}\.\(44\)
###### Lemma 4\(Absorption identity for nested projectors\)\.
Let𝐏1,𝐏2\\mathbf\{P\}\_\{1\},\\mathbf\{P\}\_\{2\}be orthogonal projectors onto subspacesℳ1⊆ℳ2\\mathcal\{M\}\_\{1\}\\subseteq\\mathcal\{M\}\_\{2\}\. Then𝐏1𝐏2=𝐏2𝐏1=𝐏1\\mathbf\{P\}\_\{1\}\\mathbf\{P\}\_\{2\}=\\mathbf\{P\}\_\{2\}\\mathbf\{P\}\_\{1\}=\\mathbf\{P\}\_\{1\}\.
###### Proof\.
For any𝐯\\mathbf\{v\},𝐏2𝐯\\mathbf\{P\}\_\{2\}\\mathbf\{v\}decomposes orthogonally as𝐏2𝐯=𝐏1𝐯\+\(𝐏2−𝐏1\)𝐯\\mathbf\{P\}\_\{2\}\\mathbf\{v\}=\\mathbf\{P\}\_\{1\}\\mathbf\{v\}\+\(\\mathbf\{P\}\_\{2\}\-\\mathbf\{P\}\_\{1\}\)\\mathbf\{v\}, where the second term lies inℳ2∩ℳ1⟂\\mathcal\{M\}\_\{2\}\\cap\\mathcal\{M\}\_\{1\}^\{\\perp\}\. Applying𝐏1\\mathbf\{P\}\_\{1\}, which is zero onℳ1⟂\\mathcal\{M\}\_\{1\}^\{\\perp\}, gives𝐏1𝐏2𝐯=𝐏1𝐯\\mathbf\{P\}\_\{1\}\\mathbf\{P\}\_\{2\}\\mathbf\{v\}=\\mathbf\{P\}\_\{1\}\\mathbf\{v\}\. The symmetric identity𝐏2𝐏1=𝐏1\\mathbf\{P\}\_\{2\}\\mathbf\{P\}\_\{1\}=\\mathbf\{P\}\_\{1\}follows by transposing\. ∎
Lemma[4](https://arxiv.org/html/2607.29071#Thmlemma4)is the abstract fact underlying Proposition[1](https://arxiv.org/html/2607.29071#Thmproposition1); the main\-text proof simply computes it in coordinates using the explicit block structure of𝐔r1∘\\mathbf\{U\}\_\{r\_\{1\}\}^\{\\circ\}relative to𝐔r2∘\\mathbf\{U\}\_\{r\_\{2\}\}^\{\\circ\}\.
#### A\-A3Polyak–Łojasiewicz condition and quadratic growth
A differentiable functionf:ℝd→ℝf:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}with nonempty minimizer setargminf\\arg\\min fand infimumf∗=infff^\{\*\}=\\inf fsatisfies the*Polyak–Łojasiewicz \(PŁ\) inequality*with constantμ\>0\\mu\>0if
‖∇f\(𝐱\)‖2≥2μ\(f\(𝐱\)−f∗\),∀𝐱\.\\\|\\nabla f\(\\mathbf\{x\}\)\\\|^\{2\}\\;\\geq\\;2\\mu\\bigl\(f\(\\mathbf\{x\}\)\-f^\{\*\}\\bigr\),\\qquad\\forall\\,\\mathbf\{x\}\.\(45\)Inequality \([45](https://arxiv.org/html/2607.29071#A1.E45)\) is weaker than strong convexity: it does not require convexity offfand permits non\-unique minimizers\. A standard sufficient condition isf\(𝐱\)=g\(𝐀𝐱\)f\(\\mathbf\{x\}\)=g\(\\mathbf\{A\}\\mathbf\{x\}\)withggstrongly convex, where𝐀\\mathbf\{A\}may be rank\-deficient\[[24](https://arxiv.org/html/2607.29071#bib.bib44)\]; compositions of this form arise naturally in overparameterized learning\.
###### Lemma 5\(Quadratic growth;\[[24](https://arxiv.org/html/2607.29071#bib.bib44)\]\)\.
Ifffsatisfies the PŁ inequality \([45](https://arxiv.org/html/2607.29071#A1.E45)\) with constantμ\\mu, then it satisfies the*quadratic\-growth \(QG\) inequality*
f\(𝐱\)−f∗≥μ2‖𝐱−Πargminf\(𝐱\)‖2,∀𝐱,f\(\\mathbf\{x\}\)\-f^\{\*\}\\;\\geq\\;\\tfrac\{\\mu\}\{2\}\\,\\bigl\\\|\\mathbf\{x\}\-\\Pi\_\{\\arg\\min f\}\(\\mathbf\{x\}\)\\bigr\\\|^\{2\},\\qquad\\forall\\,\\mathbf\{x\},\(46\)whereΠargminf\\Pi\_\{\\arg\\min f\}denotes Euclidean projection onto the \(closed\) minimizer set\.
The implication PŁ⇒\\RightarrowQG is obtained by integrating the PŁ inequality along the gradient flow offfstarting at𝐱\\mathbf\{x\}: the flow converges toargminf\\arg\\min fwith total path length bounded by2\(f\(𝐱\)−f∗\)/μ\\sqrt\{2\(f\(\\mathbf\{x\}\)\-f^\{\*\}\)/\\mu\}, which in turn upper\-bounds‖𝐱−Πargminf\(𝐱\)‖\\\|\\mathbf\{x\}\-\\Pi\_\{\\arg\\min f\}\(\\mathbf\{x\}\)\\\|\. We use Lemma[5](https://arxiv.org/html/2607.29071#Thmlemma5)in the proof of Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)to pass from an objective\-gap bound𝔼\[fi\(𝚺i\(T1\)\)−fi∗\]\\mathbb\{E\}\[f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\}\)\}\)\-f\_\{i\}^\{\*\}\]to an iterate\-distance bound𝔼‖𝚺i\(T1\)−𝚺i∗‖F\\mathbb\{E\}\\\|\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\}\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}\\\|\_\{F\}\.
### A\-BProofs for Convergence Analysis
#### A\-B1Supporting Lemmas
###### Lemma 6\(SVD isometry with scaling\)\.
Let𝐔r=𝐔r∘𝚲r1/2𝐒−1\\mathbf\{U\}\_\{r\}=\\mathbf\{U\}\_\{r\}^\{\\circ\}\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\mathbf\{S\}^\{\-1\}and𝐕r=𝐒𝚲r1/2𝐕r∘\\mathbf\{V\}\_\{r\}=\\mathbf\{S\}\\,\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\mathbf\{V\}\_\{r\}^\{\\circ\}with𝐔r∘⊤𝐔r∘=𝐈r\\mathbf\{U\}\_\{r\}^\{\\circ\\,\\top\}\\mathbf\{U\}\_\{r\}^\{\\circ\}=\\mathbf\{I\}\_\{r\},𝐕r∘𝐕r∘⊤=𝐈r\\mathbf\{V\}\_\{r\}^\{\\circ\}\\mathbf\{V\}\_\{r\}^\{\\circ\\,\\top\}=\\mathbf\{I\}\_\{r\},𝚲r=diag\(λ1,…,λr\)≻𝟎\\boldsymbol\{\\Lambda\}\_\{r\}=\\mathrm\{diag\}\(\\lambda\_\{1\},\\ldots,\\lambda\_\{r\}\)\\succ\\mathbf\{0\}, and𝐒∈ℝr×r\\mathbf\{S\}\\in\\mathbb\{R\}^\{r\\times r\}invertible\. For any𝐀∈ℝr×r\\mathbf\{A\}\\in\\mathbb\{R\}^\{r\\times r\},
‖𝐔r𝐀𝐕r‖F=‖𝚲r1/2𝐒−1𝐀𝐒𝚲r1/2‖F,\\\|\\mathbf\{U\}\_\{r\}\\,\\mathbf\{A\}\\,\\mathbf\{V\}\_\{r\}\\\|\_\{F\}\\;=\\;\\\|\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\,\\mathbf\{S\}^\{\-1\}\\mathbf\{A\}\\,\\mathbf\{S\}\\,\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\\|\_\{F\},\(47\)and consequently
λrκ\(𝐒\)−1‖𝐀‖F≤‖𝐔r𝐀𝐕r‖F≤λ1κ\(𝐒\)‖𝐀‖F\.\\lambda\_\{r\}\\,\\kappa\(\\mathbf\{S\}\)^\{\-1\}\\,\\\|\\mathbf\{A\}\\\|\_\{F\}\\;\\leq\\;\\\|\\mathbf\{U\}\_\{r\}\\,\\mathbf\{A\}\\,\\mathbf\{V\}\_\{r\}\\\|\_\{F\}\\;\\leq\\;\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\\,\\\|\\mathbf\{A\}\\\|\_\{F\}\.\(48\)In the balanced case𝐒=𝐈r\\mathbf\{S\}=\\mathbf\{I\}\_\{r\}, the bound readsλr‖𝐀‖F≤‖𝐔r𝐀𝐕r‖F≤λ1‖𝐀‖F\\lambda\_\{r\}\\,\\\|\\mathbf\{A\}\\\|\_\{F\}\\leq\\\|\\mathbf\{U\}\_\{r\}\\mathbf\{A\}\\mathbf\{V\}\_\{r\}\\\|\_\{F\}\\leq\\lambda\_\{1\}\\,\\\|\\mathbf\{A\}\\\|\_\{F\}, with equality throughout iff the retained spectrum is flat \(λ1=λr\\lambda\_\{1\}=\\lambda\_\{r\}\)\.
###### Proof\.
Write𝐁≜𝚲r1/2𝐒−1𝐀𝐒𝚲r1/2\\mathbf\{B\}\\triangleq\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\,\\mathbf\{S\}^\{\-1\}\\mathbf\{A\}\\,\\mathbf\{S\}\\,\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}, so that𝐔r𝐀𝐕r=𝐔r∘𝐁𝐕r∘\\mathbf\{U\}\_\{r\}\\mathbf\{A\}\\mathbf\{V\}\_\{r\}=\\mathbf\{U\}\_\{r\}^\{\\circ\}\\mathbf\{B\}\\,\\mathbf\{V\}\_\{r\}^\{\\circ\}\. Left multiplication by the column\-orthonormal𝐔r∘\\mathbf\{U\}\_\{r\}^\{\\circ\}and right multiplication by the row\-orthonormal𝐕r∘\\mathbf\{V\}\_\{r\}^\{\\circ\}both preserve the Frobenius norm, hence‖𝐔r𝐀𝐕r‖F=‖𝐁‖F\\\|\\mathbf\{U\}\_\{r\}\\mathbf\{A\}\\mathbf\{V\}\_\{r\}\\\|\_\{F\}=\\\|\\mathbf\{B\}\\\|\_\{F\}\. For the upper bound, using‖𝚲r1/2‖2=λ11/2\\\|\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\\|\_\{2\}=\\lambda\_\{1\}^\{1/2\},
‖𝐁‖F≤‖𝚲r1/2‖22‖𝐒−1𝐀𝐒‖F≤λ1‖𝐒−1‖2‖𝐒‖2‖𝐀‖F=λ1κ\(𝐒\)‖𝐀‖F\.\\begin\{split\}\\\|\\mathbf\{B\}\\\|\_\{F\}&\\;\\leq\\;\\\|\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\\|\_\{2\}^\{2\}\\,\\\|\\mathbf\{S\}^\{\-1\}\\mathbf\{A\}\\mathbf\{S\}\\\|\_\{F\}\\\\ &\\;\\leq\\;\\lambda\_\{1\}\\,\\\|\\mathbf\{S\}^\{\-1\}\\\|\_\{2\}\\,\\\|\\mathbf\{S\}\\\|\_\{2\}\\,\\\|\\mathbf\{A\}\\\|\_\{F\}\\;=\\;\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\\,\\\|\\mathbf\{A\}\\\|\_\{F\}\.\\end\{split\}\(49\)The lower bound follows analogously from‖𝚲r1/2𝐂𝚲r1/2‖F≥λmin\(𝚲r1/2\)2‖𝐂‖F=λr‖𝐂‖F\\\|\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\mathbf\{C\}\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\\\|\_\{F\}\\geq\\lambda\_\{\\min\}\(\\boldsymbol\{\\Lambda\}\_\{r\}^\{1/2\}\)^\{2\}\\,\\\|\\mathbf\{C\}\\\|\_\{F\}=\\lambda\_\{r\}\\,\\\|\\mathbf\{C\}\\\|\_\{F\}with𝐂=𝐒−1𝐀𝐒\\mathbf\{C\}=\\mathbf\{S\}^\{\-1\}\\mathbf\{A\}\\mathbf\{S\}, together with‖𝐒−1𝐀𝐒‖F≥κ\(𝐒\)−1‖𝐀‖F\\\|\\mathbf\{S\}^\{\-1\}\\mathbf\{A\}\\mathbf\{S\}\\\|\_\{F\}\\geq\\kappa\(\\mathbf\{S\}\)^\{\-1\}\\\|\\mathbf\{A\}\\\|\_\{F\}\. ∎
###### Lemma 7\(LeffL\_\{\\mathrm\{eff\}\}\-smoothness of the adapter loss\)\.
Under Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)\(a\) and with the factor structure of Lemma[6](https://arxiv.org/html/2607.29071#Thmlemma6), for any clientc∈𝒞ic\\in\\mathcal\{C\}\_\{i\}, the local lossfc\(𝚺\)=ℒc\(𝐔ri𝚺𝐕ri\)f\_\{c\}\(\\boldsymbol\{\\Sigma\}\)=\\mathcal\{L\}\_\{c\}\(\\mathbf\{U\}\_\{r\_\{i\}\}\\boldsymbol\{\\Sigma\}\\mathbf\{V\}\_\{r\_\{i\}\}\)isLeffL\_\{\\mathrm\{eff\}\}\-smooth in𝚺\\boldsymbol\{\\Sigma\}:
‖∇𝚺fc\(𝚺\)−∇𝚺fc\(𝚺′\)‖F≤Leff‖𝚺−𝚺′‖F,\\\|\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{c\}\(\\boldsymbol\{\\Sigma\}\)\-\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{c\}\(\\boldsymbol\{\\Sigma\}^\{\\prime\}\)\\\|\_\{F\}\\;\\leq\\;L\_\{\\mathrm\{eff\}\}\\,\\\|\\boldsymbol\{\\Sigma\}\-\\boldsymbol\{\\Sigma\}^\{\\prime\}\\\|\_\{F\},\(50\)whereLeff=Lλ12κ\(𝐒\)2L\_\{\\mathrm\{eff\}\}=L\\,\\lambda\_\{1\}^\{2\}\\,\\kappa\(\\mathbf\{S\}\)^\{2\}\. The same constant applies tofi\(𝚺\)f\_\{i\}\(\\boldsymbol\{\\Sigma\}\)by convexity of the weighted sum\.
###### Proof\.
By the chain rule,
∇𝚺fc\(𝚺\)=𝐔ri⊤∇𝐖ℒc\(𝐔ri𝚺𝐕ri\)𝐕ri⊤\.\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{c\}\(\\boldsymbol\{\\Sigma\}\)=\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\top\}\\,\\nabla\_\{\\mathbf\{W\}\}\\mathcal\{L\}\_\{c\}\(\\mathbf\{U\}\_\{r\_\{i\}\}\\boldsymbol\{\\Sigma\}\\mathbf\{V\}\_\{r\_\{i\}\}\)\\,\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\top\}\.\(51\)Set𝐃=∇𝐖ℒc\(𝐔ri𝚺𝐕ri\)−∇𝐖ℒc\(𝐔ri𝚺′𝐕ri\)\\mathbf\{D\}=\\nabla\_\{\\mathbf\{W\}\}\\mathcal\{L\}\_\{c\}\(\\mathbf\{U\}\_\{r\_\{i\}\}\\boldsymbol\{\\Sigma\}\\mathbf\{V\}\_\{r\_\{i\}\}\)\-\\nabla\_\{\\mathbf\{W\}\}\\mathcal\{L\}\_\{c\}\(\\mathbf\{U\}\_\{r\_\{i\}\}\\boldsymbol\{\\Sigma\}^\{\\prime\}\\mathbf\{V\}\_\{r\_\{i\}\}\)\. Then
‖∇𝚺fc\(𝚺\)−∇𝚺fc\(𝚺′\)‖F=‖𝐔ri⊤𝐃𝐕ri⊤‖F≤\(a\)‖𝐔ri⊤‖2‖𝐃‖F‖𝐕ri⊤‖2≤\(b\)‖𝐔ri‖2‖𝐕ri‖2L‖𝐔ri\(𝚺−𝚺′\)𝐕ri‖F≤\(c\)‖𝐔ri‖2‖𝐕ri‖2⋅L⋅λ1κ\(𝐒\)‖𝚺−𝚺′‖F,\\begin\{split\}&\\\|\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{c\}\(\\boldsymbol\{\\Sigma\}\)\-\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{c\}\(\\boldsymbol\{\\Sigma\}^\{\\prime\}\)\\\|\_\{F\}=\\\|\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\top\}\\mathbf\{D\}\\,\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\top\}\\\|\_\{F\}\\\\ &\\;\\stackrel\{\{\\scriptstyle\(a\)\}\}\{\{\\leq\}\}\\\|\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\top\}\\\|\_\{2\}\\,\\\|\\mathbf\{D\}\\\|\_\{F\}\\,\\\|\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\top\}\\\|\_\{2\}\\\\ &\\;\\stackrel\{\{\\scriptstyle\(b\)\}\}\{\{\\leq\}\}\\\|\\mathbf\{U\}\_\{r\_\{i\}\}\\\|\_\{2\}\\,\\\|\\mathbf\{V\}\_\{r\_\{i\}\}\\\|\_\{2\}\\,L\\,\\\|\\mathbf\{U\}\_\{r\_\{i\}\}\(\\boldsymbol\{\\Sigma\}\-\\boldsymbol\{\\Sigma\}^\{\\prime\}\)\\mathbf\{V\}\_\{r\_\{i\}\}\\\|\_\{F\}\\\\ &\\;\\stackrel\{\{\\scriptstyle\(c\)\}\}\{\{\\leq\}\}\\\|\\mathbf\{U\}\_\{r\_\{i\}\}\\\|\_\{2\}\\,\\\|\\mathbf\{V\}\_\{r\_\{i\}\}\\\|\_\{2\}\\\!\\cdot\\\!L\\\!\\cdot\\\!\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\\,\\\|\\boldsymbol\{\\Sigma\}\-\\boldsymbol\{\\Sigma\}^\{\\prime\}\\\|\_\{F\},\\end\{split\}\(52\)where\(a\)\(a\)uses the operator\-norm bound on sandwiching by matrices \(and‖𝐀⊤‖2=‖𝐀‖2\\\|\\mathbf\{A\}^\{\\top\}\\\|\_\{2\}=\\\|\\mathbf\{A\}\\\|\_\{2\}\),\(b\)\(b\)invokes Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)\(a\), and\(c\)\(c\)applies the upper bound of \([48](https://arxiv.org/html/2607.29071#A1.E48)\)\. Finally,‖𝐔ri‖2=‖𝐔ri∘𝚲ri1/2𝐒−1‖2≤λ11/2‖𝐒−1‖2\\\|\\mathbf\{U\}\_\{r\_\{i\}\}\\\|\_\{2\}=\\\|\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\circ\}\\boldsymbol\{\\Lambda\}\_\{r\_\{i\}\}^\{1/2\}\\mathbf\{S\}^\{\-1\}\\\|\_\{2\}\\leq\\lambda\_\{1\}^\{1/2\}\\\|\\mathbf\{S\}^\{\-1\}\\\|\_\{2\}and‖𝐕ri‖2=‖𝐒𝚲ri1/2𝐕ri∘‖2≤λ11/2‖𝐒‖2\\\|\\mathbf\{V\}\_\{r\_\{i\}\}\\\|\_\{2\}=\\\|\\mathbf\{S\}\\,\\boldsymbol\{\\Lambda\}\_\{r\_\{i\}\}^\{1/2\}\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\circ\}\\\|\_\{2\}\\leq\\lambda\_\{1\}^\{1/2\}\\\|\\mathbf\{S\}\\\|\_\{2\}, so‖𝐔ri‖2‖𝐕ri‖2≤λ1κ\(𝐒\)\\\|\\mathbf\{U\}\_\{r\_\{i\}\}\\\|\_\{2\}\\,\\\|\\mathbf\{V\}\_\{r\_\{i\}\}\\\|\_\{2\}\\leq\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\. Combining,
‖∇𝚺fc\(𝚺\)−∇𝚺fc\(𝚺′\)‖F≤Lλ12κ\(𝐒\)2‖𝚺−𝚺′‖F=Leff‖𝚺−𝚺′‖F\.∎\\begin\{split\}&\\\|\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{c\}\(\\boldsymbol\{\\Sigma\}\)\-\\nabla\_\{\\boldsymbol\{\\Sigma\}\}f\_\{c\}\(\\boldsymbol\{\\Sigma\}^\{\\prime\}\)\\\|\_\{F\}\\\\ &\\quad\\leq L\\,\\lambda\_\{1\}^\{2\}\\,\\kappa\(\\mathbf\{S\}\)^\{2\}\\,\\\|\\boldsymbol\{\\Sigma\}\-\\boldsymbol\{\\Sigma\}^\{\\prime\}\\\|\_\{F\}=L\_\{\\mathrm\{eff\}\}\\,\\\|\\boldsymbol\{\\Sigma\}\-\\boldsymbol\{\\Sigma\}^\{\\prime\}\\\|\_\{F\}\.\\qed\\end\{split\}\(53\)
###### Lemma 8\(Local update drift bound\)\.
Let clientc∈𝒞ic\\in\\mathcal\{C\}\_\{i\}runE≥1E\\geq 1local SGD steps with constant step sizeη≤18LeffE\\eta\\leq\\dfrac\{1\}\{8\\,L\_\{\\mathrm\{eff\}\}\\,E\}starting from the round\-ttaggregated group adapter𝚺i\(t\)\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\. Denote the local iterates by𝚺i,c\(t,0\)=𝚺i\(t\)\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,0\)\}=\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}and𝚺i,c\(t,e\+1\)=𝚺i,c\(t,e\)−η𝐠~c\(𝚺i,c\(t,e\)\)\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\+1\)\}=\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\-\\eta\\,\\widetilde\{\\mathbf\{g\}\}\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)fore=0,…,E−1e=0,\\ldots,E\-1\. Then for every0≤e≤E0\\leq e\\leq E,
∑c∈𝒞incni𝔼‖𝚺i,c\(t,e\)−𝚺i\(t\)‖F2≤4η2E\(σ2\+6EGi2\)\+8η2E2‖∇fi\(𝚺i\(t\)\)‖F2\.\\begin\{split\}\\sum\_\{c\\in\\mathcal\{C\}\_\{i\}\}\\frac\{n\_\{c\}\}\{n\_\{i\}\}\\mathbb\{E\}\\bigl\\\|\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\\bigr\\\|\_\{F\}^\{2\}&\\leq 4\\eta^\{2\}E\\,\(\\sigma^\{2\}\+6E\\,G\_\{i\}^\{2\}\)\\\\ &\+8\\eta^\{2\}E^\{2\}\\,\\bigl\\\|\\nabla f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\)\\bigr\\\|\_\{F\}^\{2\}\.\\end\{split\}\(54\)
###### Proof\.
The bound is a standard result from the analysis of local SGD\[[25](https://arxiv.org/html/2607.29071#bib.bib43),[27](https://arxiv.org/html/2607.29071#bib.bib45)\]\. Set𝚫c\(e\)≜𝚺i,c\(t,e\)−𝚺i\(t\)\\mathbf\{\\Delta\}\_\{c\}^\{\(e\)\}\\triangleq\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}and decompose𝐠~c\(𝚺i,c\(t,e\)\)=∇fi\(𝚺i\(t\)\)\+\[∇fc\(𝚺i,c\(t,e\)\)−∇fi\(𝚺i\(t\)\)\]\+\[𝐠~c\(𝚺i,c\(t,e\)\)−∇fc\(𝚺i,c\(t,e\)\)\]\\widetilde\{\\mathbf\{g\}\}\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)=\\nabla f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\)\+\\bigl\[\\nabla f\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)\-\\nabla f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\)\\bigr\]\+\\bigl\[\\widetilde\{\\mathbf\{g\}\}\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)\-\\nabla f\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)\\bigr\]\. Using‖a\+b\+c‖2≤3\(‖a‖2\+‖b‖2\+‖c‖2\)\\\|a\+b\+c\\\|^\{2\}\\leq 3\(\\\|a\\\|^\{2\}\+\\\|b\\\|^\{2\}\+\\\|c\\\|^\{2\}\),LeffL\_\{\\mathrm\{eff\}\}\-smoothness \(Lemma[7](https://arxiv.org/html/2607.29071#Thmlemma7)\), and Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)\(b\)–\(c\), an induction oneewith the step\-size conditionη≤18LeffE\\eta\\leq\\frac\{1\}\{8L\_\{\\mathrm\{eff\}\}E\}produces \([54](https://arxiv.org/html/2607.29071#A1.E54)\)\. ∎
#### A\-B2Proof of Theorem[1](https://arxiv.org/html/2607.29071#Thmtheorem1)
###### Proof\.
We begin with a one\-round descent\. By Lemma[7](https://arxiv.org/html/2607.29071#Thmlemma7),
fi\(𝚺i\(t\+1\)\)≤fi\(𝚺i\(t\)\)\+⟨∇fi\(t\),𝚺i\(t\+1\)−𝚺i\(t\)⟩\+Leff2‖𝚺i\(t\+1\)−𝚺i\(t\)‖F2\.\\begin\{split\}f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\+1\)\}\)\\leq\\;&f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\)\+\\langle\\nabla f\_\{i\}^\{\(t\)\},\\;\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\+1\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\\rangle\\\\ &\+\\frac\{L\_\{\\mathrm\{eff\}\}\}\{2\}\\\|\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\+1\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\\\|\_\{F\}^\{2\}\.\\end\{split\}\(55\)We next bound the inner\-product term\. The server update is
𝚺i\(t\+1\)−𝚺i\(t\)=−η∑c∈𝒞incni∑e=0E−1𝐠~c\(𝚺i,c\(t,e\)\)\.\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\+1\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}=\-\\eta\\sum\_\{c\\in\\mathcal\{C\}\_\{i\}\}\\frac\{n\_\{c\}\}\{n\_\{i\}\}\\sum\_\{e=0\}^\{E\-1\}\\widetilde\{\\mathbf\{g\}\}\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)\.\(56\)Taking expectations and using Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)\(b\),
𝔼⟨∇fi\(t\),𝚺i\(t\+1\)−𝚺i\(t\)⟩=−η∑c,encni𝔼⟨∇fi\(t\),∇fc\(𝚺i,c\(t,e\)\)⟩=−ηE‖∇fi\(t\)‖F2−η∑c,encni𝔼⟨∇fi\(t\),∇fc\(𝚺i,c\(t,e\)\)−∇fi\(t\)⟩\.\\begin\{split\}&\\mathbb\{E\}\\langle\\nabla f\_\{i\}^\{\(t\)\},\\,\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\+1\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\\rangle\\\\ \\;=&\-\\eta\\\!\\sum\_\{c,e\}\\frac\{n\_\{c\}\}\{n\_\{i\}\}\\mathbb\{E\}\\langle\\nabla f\_\{i\}^\{\(t\)\},\\,\\nabla f\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)\\rangle\\\\ \\;=&\-\\eta E\\,\\\|\\nabla f\_\{i\}^\{\(t\)\}\\\|\_\{F\}^\{2\}\\\\ &\-\\eta\\\!\\sum\_\{c,e\}\\frac\{n\_\{c\}\}\{n\_\{i\}\}\\mathbb\{E\}\\langle\\nabla f\_\{i\}^\{\(t\)\},\\,\\nabla f\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)\-\\nabla f\_\{i\}^\{\(t\)\}\\rangle\.\\end\{split\}\(57\)Applying Young’s inequality⟨a,b⟩≥−12‖a‖2−12‖b‖2\\langle a,b\\rangle\\geq\-\\tfrac\{1\}\{2\}\\\|a\\\|^\{2\}\-\\tfrac\{1\}\{2\}\\\|b\\\|^\{2\}and decomposing∇fc\(𝚺i,c\(t,e\)\)−∇fi\(t\)=\[∇fc\(𝚺i,c\(t,e\)\)−∇fc\(𝚺i\(t\)\)\]\+\[∇fc\(𝚺i\(t\)\)−∇fi\(𝚺i\(t\)\)\]\\nabla f\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)\-\\nabla f\_\{i\}^\{\(t\)\}=\[\\nabla f\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)\-\\nabla f\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\)\]\+\[\\nabla f\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\)\-\\nabla f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\)\], the second term above is controlled byLeffL\_\{\\mathrm\{eff\}\}\-smoothness \(Lemma[7](https://arxiv.org/html/2607.29071#Thmlemma7)\), Lemma[8](https://arxiv.org/html/2607.29071#Thmlemma8), and Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)\(c\), giving
𝔼⟨∇fi\(t\),𝚺i\(t\+1\)−𝚺i\(t\)⟩≤−ηE2‖∇fi\(t\)‖F2\+ηEGi2\+Leff2η3E3\(σ2\+6EGi2\)\.\\begin\{split\}&\\mathbb\{E\}\\langle\\nabla f\_\{i\}^\{\(t\)\},\\,\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\+1\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\\rangle\\\\ \\leq&\-\\tfrac\{\\eta E\}\{2\}\\\|\\nabla f\_\{i\}^\{\(t\)\}\\\|\_\{F\}^\{2\}\+\\eta E\\,G\_\{i\}^\{2\}\+L\_\{\\mathrm\{eff\}\}^\{\\,2\}\\,\\eta^\{3\}E^\{3\}\(\\sigma^\{2\}\+6EG\_\{i\}^\{2\}\)\.\\end\{split\}\(58\)For the quadratic term, Jensen’s inequality and convexity of the squared norm give
𝔼‖𝚺i\(t\+1\)−𝚺i\(t\)‖F2≤η2E∑c,encni𝔼‖𝐠~c\(𝚺i,c\(t,e\)\)‖F2\.\\begin\{split\}&\\mathbb\{E\}\\\|\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\+1\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\\\|\_\{F\}^\{2\}\\leq\\eta^\{2\}E\\sum\_\{c,e\}\\frac\{n\_\{c\}\}\{n\_\{i\}\}\\mathbb\{E\}\\\|\\widetilde\{\\mathbf\{g\}\}\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)\\\|\_\{F\}^\{2\}\.\\end\{split\}\(59\)Decomposing each stochastic gradient as𝐠~c=∇fi\(t\)\+\[∇fc\(𝚺i,c\(t,e\)\)−∇fi\(t\)\]\+\[𝐠~c−∇fc\(𝚺i,c\(t,e\)\)\]\\widetilde\{\\mathbf\{g\}\}\_\{c\}=\\nabla f\_\{i\}^\{\(t\)\}\+\[\\nabla f\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)\-\\nabla f\_\{i\}^\{\(t\)\}\]\+\[\\widetilde\{\\mathbf\{g\}\}\_\{c\}\-\\nabla f\_\{c\}\(\\boldsymbol\{\\Sigma\}\_\{i,c\}^\{\(t,e\)\}\)\]and applying‖a\+b\+c‖2≤3\(‖a‖2\+‖b‖2\+‖c‖2\)\\\|a\+b\+c\\\|^\{2\}\\leq 3\(\\\|a\\\|^\{2\}\+\\\|b\\\|^\{2\}\+\\\|c\\\|^\{2\}\), Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)\(b\)–\(c\),LeffL\_\{\\mathrm\{eff\}\}\-smoothness \(Lemma[7](https://arxiv.org/html/2607.29071#Thmlemma7)\), and \(Lemma[8](https://arxiv.org/html/2607.29071#Thmlemma8)\) yield
𝔼‖𝚺i\(t\+1\)−𝚺i\(t\)‖F2≤3η2E2‖∇fi\(t\)‖F2\+3η2E2Gi2\+3η2Eσ2\+12Leff2η4E3\(σ2\+6EGi2\)\+24Leff2η4E4‖∇fi\(t\)‖F2\.\\begin\{split\}&\\mathbb\{E\}\\\|\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\+1\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\\\|\_\{F\}^\{2\}\\\\ \\leq&3\\eta^\{2\}E^\{2\}\\,\\\|\\nabla f\_\{i\}^\{\(t\)\}\\\|\_\{F\}^\{2\}\+3\\eta^\{2\}E^\{2\}G\_\{i\}^\{2\}\+3\\eta^\{2\}E\\,\\sigma^\{2\}\\\\ &\+12\\,L\_\{\\mathrm\{eff\}\}^\{\\,2\}\\,\\eta^\{4\}E^\{3\}\(\\sigma^\{2\}\+6E\\,G\_\{i\}^\{2\}\)\\\\ &\+24\\,L\_\{\\mathrm\{eff\}\}^\{\\,2\}\\,\\eta^\{4\}E^\{4\}\\,\\\|\\nabla f\_\{i\}^\{\(t\)\}\\\|\_\{F\}^\{2\}\.\\end\{split\}\(60\)Substituting \([58](https://arxiv.org/html/2607.29071#A1.E58)\) and \([60](https://arxiv.org/html/2607.29071#A1.E60)\) into \([55](https://arxiv.org/html/2607.29071#A1.E55)\), taking expectations, and usingη≤18LeffE\\eta\\leq\\frac\{1\}\{8L\_\{\\mathrm\{eff\}\}E\}to absorb the higher\-order terms, we obtain the per\-round inequality
𝔼\[fi\(t\+1\)\]≤𝔼\[fi\(t\)\]−ηE4𝔼‖∇fi\(t\)‖F2\+ηE2Gi2\+Leffη2Eσ2\+2Leff2η3E3\(σ2\+6EGi2\)\.\\begin\{split\}\\mathbb\{E\}\[f\_\{i\}^\{\(t\+1\)\}\]&\\leq\\;\\mathbb\{E\}\[f\_\{i\}^\{\(t\)\}\]\-\\tfrac\{\\eta E\}\{4\}\\,\\mathbb\{E\}\\\|\\nabla f\_\{i\}^\{\(t\)\}\\\|\_\{F\}^\{2\}\+\\tfrac\{\\eta E\}\{2\}\\,G\_\{i\}^\{2\}\\\\ &\+L\_\{\\mathrm\{eff\}\}\\,\\eta^\{2\}E\\,\\sigma^\{2\}\+2\\,L\_\{\\mathrm\{eff\}\}^\{\\,2\}\\,\\eta^\{3\}E^\{3\}\(\\sigma^\{2\}\+6EG\_\{i\}^\{2\}\)\.\\end\{split\}\(61\)
Finally, summing \([61](https://arxiv.org/html/2607.29071#A1.E61)\) fromt=0t=0toT1−1T\_\{1\}\-1, dividing byηET1/4\\eta ET\_\{1\}/4, and usingfi\(T1\)≥inf𝚺fif\_\{i\}^\{\(T\_\{1\}\)\}\\geq\\inf\_\{\\boldsymbol\{\\Sigma\}\}f\_\{i\}yields
1T1∑t=0T1−1𝔼‖∇fi\(t\)‖F2≤4Δ0ηET1\+2Gi2\+4Leffησ2\+8Leff2η2E2\(σ2\+6EGi2\),\\begin\{split\}\\frac\{1\}\{T\_\{1\}\}\\sum\_\{t=0\}^\{T\_\{1\}\-1\}\\mathbb\{E\}\\\|\\nabla f\_\{i\}^\{\(t\)\}\\\|\_\{F\}^\{2\}\\;&\\leq\\;\\frac\{4\\Delta\_\{0\}\}\{\\eta ET\_\{1\}\}\+2G\_\{i\}^\{2\}\+4L\_\{\\mathrm\{eff\}\}\\,\\eta\\,\\sigma^\{2\}\\\\ &\\quad\+8L\_\{\\mathrm\{eff\}\}^\{\\,2\}\\,\\eta^\{2\}E^\{2\}\(\\sigma^\{2\}\+6EG\_\{i\}^\{2\}\),\\end\{split\}\(62\)which matches \([17](https://arxiv.org/html/2607.29071#S4.E17)\)\. Settingη=1LeffT1E\\eta=\\frac\{1\}\{L\_\{\\mathrm\{eff\}\}\\sqrt\{T\_\{1\}E\}\}\(which satisfiesη≤18LeffE\\eta\\leq\\frac\{1\}\{8L\_\{\\mathrm\{eff\}\}E\}wheneverT1≥64ET\_\{1\}\\geq 64E\) produces the rate \([18](https://arxiv.org/html/2607.29071#S4.E18)\)\. ∎
### A\-CRate Conversion for the Optimization Residual
We give the full derivation of the rate \([27](https://arxiv.org/html/2607.29071#S4.E27)\) stated in Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\(a\)\. By Lemma[5](https://arxiv.org/html/2607.29071#Thmlemma5), the PŁ inequality \([26](https://arxiv.org/html/2607.29071#S4.E26)\) implies the quadratic\-growth \(QG\) inequality
fi\(𝚺\)−fi∗≥μ2‖𝚺−Πargminfi\(𝚺\)‖F2,f\_\{i\}\(\\boldsymbol\{\\Sigma\}\)\-f\_\{i\}^\{\*\}\\;\\geq\\;\\tfrac\{\\mu\}\{2\}\\,\\bigl\\\|\\boldsymbol\{\\Sigma\}\-\\Pi\_\{\\arg\\min f\_\{i\}\}\(\\boldsymbol\{\\Sigma\}\)\\bigr\\\|\_\{F\}^\{\\,2\},\(63\)whereΠargminfi\\Pi\_\{\\arg\\min f\_\{i\}\}denotes the Euclidean projection onto the minimizer set\. We first convert the gradient\-norm bound into an objective gap\. Combining Theorem[1](https://arxiv.org/html/2607.29071#Thmtheorem1)with \([26](https://arxiv.org/html/2607.29071#S4.E26)\) gives
1T1∑t=0T1−1𝔼\[fi\(𝚺i\(t\)\)−fi∗\]≤12μ⋅1T1∑t=0T1−1𝔼‖∇fi\(𝚺i\(t\)\)‖F2=𝒪\(LeffΔ0\+σ2μT1E\)⏟vanishing\+G2μ⏟floor\.\\begin\{split\}&\\frac\{1\}\{T\_\{1\}\}\\sum\_\{t=0\}^\{T\_\{1\}\-1\}\\mathbb\{E\}\\bigl\[f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\)\-f\_\{i\}^\{\*\}\\bigr\]\\\\ \\leq&\\frac\{1\}\{2\\mu\}\\\!\\cdot\\\!\\frac\{1\}\{T\_\{1\}\}\\sum\_\{t=0\}^\{T\_\{1\}\-1\}\\mathbb\{E\}\\bigl\\\|\\nabla f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(t\)\}\)\\bigr\\\|\_\{F\}^\{\\,2\}\\\\ =&\\underbrace\{\\mathcal\{O\}\\\!\\Bigl\(\\tfrac\{L\_\{\\mathrm\{eff\}\}\\Delta\_\{0\}\+\\sigma^\{2\}\}\{\\mu\\sqrt\{T\_\{1\}E\}\}\\Bigr\)\}\_\{\\text\{vanishing\}\}\+\\;\\underbrace\{\\tfrac\{G^\{2\}\}\{\\mu\}\}\_\{\\text\{floor\}\}\.\\end\{split\}\(64\)Applying \([63](https://arxiv.org/html/2607.29071#A1.E63)\) to𝚺i\(T1\)\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\}\)\}and Jensen’s inequality \(t↦tt\\mapsto\\sqrt\{t\}is concave\) then turns the objective gap into an iterate\-distance bound:
𝔼‖𝚺i\(T1\)−𝚺^i‖F≤2μ𝔼\[fi\(𝚺i\(T1\)\)−fi∗\]=𝒪\(T1−1/4\)\+𝒪\(G/μ\)\.\\begin\{split\}&\\mathbb\{E\}\\bigl\\\|\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\}\)\}\-\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\bigr\\\|\_\{F\}\\\\ \\;\\leq\\;&\\sqrt\{\\tfrac\{2\}\{\\mu\}\\,\\mathbb\{E\}\\bigl\[f\_\{i\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\}\)\}\)\-f\_\{i\}^\{\*\}\\bigr\]\}\\\\ \\;=\\;&\\mathcal\{O\}\\\!\\bigl\(T\_\{1\}^\{\-1/4\}\\bigr\)\+\\mathcal\{O\}\\\!\\bigl\(G/\\sqrt\{\\mu\}\\bigr\)\.\\end\{split\}\(65\)Finally, lifting from adapter space to weight space via Lemma[6](https://arxiv.org/html/2607.29071#Thmlemma6),‖𝐔riA𝐕ri‖F≤λ1κ\(𝐒\)‖A‖F\\\|\\mathbf\{U\}\_\{r\_\{i\}\}A\\mathbf\{V\}\_\{r\_\{i\}\}\\\|\_\{F\}\\leq\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\\,\\\|A\\\|\_\{F\}, so
𝔼‖εiopt‖F≤λ1κ\(𝐒\)⋅\(𝒪\(T1−1/4\)\+𝒪\(G/μ\)\)\.\\mathbb\{E\}\\\|\\varepsilon\_\{i\}^\{\\mathrm\{opt\}\}\\\|\_\{F\}\\;\\leq\\;\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\\\!\\cdot\\\!\\Bigl\(\\mathcal\{O\}\(T\_\{1\}^\{\-1/4\}\)\+\\mathcal\{O\}\(G/\\sqrt\{\\mu\}\)\\Bigr\)\.\(66\)
The exponent−1/4\-1/4is the exact price of converting a second\-moment gradient bound into a first\-moment iterate bound: one factor of1/21/2from the square root in \([63](https://arxiv.org/html/2607.29071#A1.E63)\), another from Jensen’s inequality\. In the non\-IID regime, the additive𝒪\(λ1κ\(𝐒\)G/μ\)\\mathcal\{O\}\(\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\\,G/\\sqrt\{\\mu\}\)term should be absorbed intoδri\(𝐖^\)\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\); the qualitative conclusions of Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3)and Proposition[4](https://arxiv.org/html/2607.29071#Thmproposition4)are unchanged\.
### A\-DProofs for Subspace Alignment Theory
###### Proof of Proposition[1](https://arxiv.org/html/2607.29071#Thmproposition1)\.
Since the singular vectors are ordered by decreasing singular value,\{𝐮1,…,𝐮r1\}⊂\{𝐮1,…,𝐮r2\}\\\{\\mathbf\{u\}\_\{1\},\\ldots,\\mathbf\{u\}\_\{r\_\{1\}\}\\\}\\subset\\\{\\mathbf\{u\}\_\{1\},\\ldots,\\mathbf\{u\}\_\{r\_\{2\}\}\\\}forr1<r2r\_\{1\}<r\_\{2\}, so𝒰r1⊊𝒰r2\\mathcal\{U\}\_\{r\_\{1\}\}\\subsetneq\\mathcal\{U\}\_\{r\_\{2\}\}follows immediately\. The same argument applies to the right factors\. In matrix form, with the orthonormal factors,
𝐔r1∘=𝐔r2∘\[𝐈r1𝟎\],𝐕r1∘=\[𝐈r1𝟎\]𝐕r2∘,\\mathbf\{U\}\_\{r\_\{1\}\}^\{\\circ\}\\;=\\;\\mathbf\{U\}\_\{r\_\{2\}\}^\{\\circ\}\\begin\{bmatrix\}\\mathbf\{I\}\_\{r\_\{1\}\}\\\\ \\mathbf\{0\}\\end\{bmatrix\},\\qquad\\mathbf\{V\}\_\{r\_\{1\}\}^\{\\circ\}\\;=\\;\\begin\{bmatrix\}\\mathbf\{I\}\_\{r\_\{1\}\}&\\mathbf\{0\}\\end\{bmatrix\}\\mathbf\{V\}\_\{r\_\{2\}\}^\{\\circ\},\(67\)so
𝐏r1𝐏r2=𝐔r1∘𝐔r1∘⊤𝐔r2∘𝐔r2∘⊤=𝐔r1∘\[𝐈r1;𝟎\]𝐔r2∘⊤=𝐔r1∘𝐔r1∘⊤=𝐏r1\.\\begin\{split\}\\mathbf\{P\}\_\{r\_\{1\}\}\\mathbf\{P\}\_\{r\_\{2\}\}&=\\mathbf\{U\}\_\{r\_\{1\}\}^\{\\circ\}\\mathbf\{U\}\_\{r\_\{1\}\}^\{\\circ\\,\\top\}\\mathbf\{U\}\_\{r\_\{2\}\}^\{\\circ\}\\mathbf\{U\}\_\{r\_\{2\}\}^\{\\circ\\,\\top\}=\\mathbf\{U\}\_\{r\_\{1\}\}^\{\\circ\}\[\\mathbf\{I\}\_\{r\_\{1\}\};\\mathbf\{0\}\]\\mathbf\{U\}\_\{r\_\{2\}\}^\{\\circ\\,\\top\}\\\\ &=\\mathbf\{U\}\_\{r\_\{1\}\}^\{\\circ\}\\mathbf\{U\}\_\{r\_\{1\}\}^\{\\circ\\,\\top\}=\\mathbf\{P\}\_\{r\_\{1\}\}\.\\end\{split\}\(68\)The argument for𝐐\\mathbf\{Q\}is analogous\. ∎
### A\-EProof of Theorem[2](https://arxiv.org/html/2607.29071#Thmtheorem2)
We give the full proof of the general cross\-group aggregation bound \(Theorem[2](https://arxiv.org/html/2607.29071#Thmtheorem2)\) and Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\(residual control under PŁ and the common\-projection\-target conditions\)\. The companion identity
𝐔ri𝚺^i𝐕ri=𝐔ri𝐔ri†𝐖^𝐕ri†𝐕ri=𝐏ri𝐖^𝐐ri\\mathbf\{U\}\_\{r\_\{i\}\}\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\mathbf\{V\}\_\{r\_\{i\}\}\\;=\\;\\mathbf\{U\}\_\{r\_\{i\}\}\\mathbf\{U\}\_\{r\_\{i\}\}^\{\\dagger\}\\,\\widehat\{\\mathbf\{W\}\}\\,\\mathbf\{V\}\_\{r\_\{i\}\}^\{\\dagger\}\\mathbf\{V\}\_\{r\_\{i\}\}\\;=\\;\\mathbf\{P\}\_\{r\_\{i\}\}\\widehat\{\\mathbf\{W\}\}\\mathbf\{Q\}\_\{r\_\{i\}\}\(69\)holds unconditionally by Lemma[3](https://arxiv.org/html/2607.29071#Thmlemma3); we use it freely throughout\.
###### Proof of Theorem[2](https://arxiv.org/html/2607.29071#Thmtheorem2)\.
Adding and subtracting𝐔ri𝚺i∗𝐕ri\\mathbf\{U\}\_\{r\_\{i\}\}\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}\\mathbf\{V\}\_\{r\_\{i\}\}and𝐏ri𝐖^𝐐ri\\mathbf\{P\}\_\{r\_\{i\}\}\\widehat\{\\mathbf\{W\}\}\\mathbf\{Q\}\_\{r\_\{i\}\}inside𝐖i−𝐖^\\mathbf\{W\}\_\{i\}\-\\widehat\{\\mathbf\{W\}\}produces the telescoping identity
𝐖i−𝐖^=εiopt\+εimis\+εisub,\\begin\{split\}\\mathbf\{W\}\_\{i\}\-\\widehat\{\\mathbf\{W\}\}\\;=\\;\\varepsilon\_\{i\}^\{\\mathrm\{opt\}\}\+\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}\+\\varepsilon\_\{i\}^\{\\mathrm\{sub\}\},\\end\{split\}\(70\)with the three residuals defined as in the main text\. Each adjacent pair cancels by direct subtraction\. Apply the triangle inequality and convexity of the Frobenius norm to𝐖g−𝐖^=∑i\(ni/N\)\(𝐖i−𝐖^\)\\mathbf\{W\}\_\{\\mathrm\{g\}\}\-\\widehat\{\\mathbf\{W\}\}=\\sum\_\{i\}\(n\_\{i\}/N\)\(\\mathbf\{W\}\_\{i\}\-\\widehat\{\\mathbf\{W\}\}\), and bound each‖𝐖i−𝐖^‖F\\\|\\mathbf\{W\}\_\{i\}\-\\widehat\{\\mathbf\{W\}\}\\\|\_\{F\}using \([70](https://arxiv.org/html/2607.29071#A1.E70)\) and Definition[3](https://arxiv.org/html/2607.29071#Thmdefinition3); this gives \([25](https://arxiv.org/html/2607.29071#S4.E25)\)\. ∎
###### Proof of Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\(a\)\.
Under Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)and the PŁ inequality \([26](https://arxiv.org/html/2607.29071#S4.E26)\), the quadratic\-growth inequality \(Lemma[5](https://arxiv.org/html/2607.29071#Thmlemma5)\) gives, for each groupiiand each iterate𝚺\\boldsymbol\{\\Sigma\},
fi\(𝚺\)−fi∗≥μ2‖𝚺−Πargminfi\(𝚺\)‖F2\.f\_\{i\}\(\\boldsymbol\{\\Sigma\}\)\-f\_\{i\}^\{\*\}\\;\\geq\\;\\tfrac\{\\mu\}\{2\}\\bigl\\\|\\boldsymbol\{\\Sigma\}\-\\Pi\_\{\\arg\\min f\_\{i\}\}\(\\boldsymbol\{\\Sigma\}\)\\bigr\\\|\_\{F\}^\{\\,2\}\.\(71\)Combining with the Stage\-1 squared\-gradient\-norm bound of Theorem[1](https://arxiv.org/html/2607.29071#Thmtheorem1)and applying Jensen’s inequality yields𝔼‖𝚺i\(T1\)−𝚺i∗‖F=𝒪\(T1−1/4\)\+𝒪\(G/μ\)\\mathbb\{E\}\\\|\\boldsymbol\{\\Sigma\}\_\{i\}^\{\(T\_\{1\}\)\}\-\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}\\\|\_\{F\}=\\mathcal\{O\}\(T\_\{1\}^\{\-1/4\}\)\+\\mathcal\{O\}\(G/\\sqrt\{\\mu\}\)\(full derivation in Appendix[A\-C](https://arxiv.org/html/2607.29071#A1.SS3)\)\. Lifting from adapter space to weight space via the upper isometry of Lemma[6](https://arxiv.org/html/2607.29071#Thmlemma6)introduces a factorλ1κ\(𝐒\)\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)and produces \([27](https://arxiv.org/html/2607.29071#S4.E27)\)\. ∎
###### Proof of Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\(b\)\.
If𝚺^i∈argminfi\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\in\\arg\\min f\_\{i\}for everyii, thenfi\(𝚺^i\)=fi∗f\_\{i\}\(\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\)=f\_\{i\}^\{\*\}, soΔi\(𝐖^\)=0\\Delta\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)=0\. Whenargminfi\\arg\\min f\_\{i\}is a singleton \(the regime guaranteed by part \(a\) plus uniqueness; otherwise pick the component visited by Stage\-1 SGD\), this forces𝚺^i=𝚺i∗\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}=\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}and henceεimis=𝐔ri\(𝚺i∗−𝚺^i\)𝐕ri=𝟎\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}=\\mathbf\{U\}\_\{r\_\{i\}\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}\-\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\)\\mathbf\{V\}\_\{r\_\{i\}\}=\\mathbf\{0\}\. The general bound \([29](https://arxiv.org/html/2607.29071#S4.E29)\) \(under part \(a\)\) follows from the quadratic\-growth inequality applied at𝚺^i\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}rather than at the iterate,
‖𝚺i∗−𝚺^i‖F≤2Δi\(𝐖^\)μ,\\bigl\\\|\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}\-\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\bigr\\\|\_\{F\}\\;\\leq\\;\\sqrt\{\\tfrac\{2\\,\\Delta\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\}\{\\mu\}\},\(72\)combined with \([69](https://arxiv.org/html/2607.29071#A1.E69)\) \(soεimis=𝐔ri\(𝚺i∗−𝚺^i\)𝐕ri\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}=\\mathbf\{U\}\_\{r\_\{i\}\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}\-\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\)\\mathbf\{V\}\_\{r\_\{i\}\}\) and the upper isometry of Lemma[6](https://arxiv.org/html/2607.29071#Thmlemma6), which contributes a factorλ1κ\(𝐒\)\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\. ∎
When the common\-projection\-target condition fails, the behaviour ofεimis\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}is characterized through three complementary results\.*\(i\) Sufficient condition*\(Proposition[5](https://arxiv.org/html/2607.29071#Thmproposition5)\): under a positive\-definite quadratic surrogate, two structural properties at𝐖^\\widehat\{\\mathbf\{W\}\}together make𝚺^i\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}a loss minimizer: the Hessian commutes with the SVD projector, and the loss gradient has no retained\-subspace component\. When both hold exactly,εimis=𝟎\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}=\\mathbf\{0\}\.*\(ii\) Curvature\-side perturbation bound*\(Proposition[6](https://arxiv.org/html/2607.29071#Thmproposition6)\): when the two structural properties hold only approximately, with explicit budgetsεG,i\\varepsilon\_\{G,i\}\(gradient leakage\),εH,i\\varepsilon\_\{H,i\}\(curvature–projector commutator\), andρi\\rho\_\{i\}\(Hessian variation\), restricted strong convexity yields a closed\-form bound on‖εimis‖F\\\|\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}\\\|\_\{F\}scaling linearly withεG,i\+εH,iδri\(𝐖^\)\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\)\. The bound recoversεimis=𝟎\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}=\\mathbf\{0\}when all budgets vanish\.*\(iii\) Data\-side decomposition*\(Proposition[7](https://arxiv.org/html/2607.29071#Thmproposition7)\): an independent route decomposes the loss gapΔi\(𝐖^\)\\Delta\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)into a statistical heterogeneity termHi\(𝐖^\)H\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)and a per\-group compression termδri\(𝐖iloc\)\\delta\_\{r\_\{i\}\}\(\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\), using onlyLL\-smoothness and group\-optimum statistics without curvature information\. The two routes are complementary: neither subsumes the other\. Remark[2](https://arxiv.org/html/2607.29071#Thmremark2)below discusses when each is the operative tool\.
To state the curvature\-aligned condition in plain geometric terms we use a curvature\-weighted norm\. For a self\-adjoint positive\-definite operatorℋ:ℝm×n→ℝm×n\\mathcal\{H\}:\\mathbb\{R\}^\{m\\times n\}\\to\\mathbb\{R\}^\{m\\times n\}, write‖𝐌‖ℋ2≜⟨𝐌,ℋ\[𝐌\]⟩F\\\|\\mathbf\{M\}\\\|\_\{\\mathcal\{H\}\}^\{2\}\\triangleq\\langle\\mathbf\{M\},\\,\\mathcal\{H\}\[\\mathbf\{M\}\]\\rangle\_\{F\}\. Intuitively this measures the size of𝐌\\mathbf\{M\}with each direction weighted by how sharply the loss bends along it; the plain Frobenius norm \(ℋ=Id\\mathcal\{H\}=\\mathrm\{Id\}\) weights all directions equally\. It is the matrix analogue of the weighted norm∥⋅∥H\\\|\\cdot\\\|\_\{H\}that turns the scalar descent lemma into its anisotropic form\.
###### Proposition 5\(Sufficient condition: aligned curvature and subspace\-stationary reference\)\.
Fix the reference weight𝐖^\\widehat\{\\mathbf\{W\}\}\(e\.g\. the pre\-trained weight\)\. Suppose that, for each groupi∈\[K\]i\\in\[K\], the ambient loss is a positive\-definite quadratic, written as its second\-order Taylor expansion*around the reference*𝐖^\\widehat\{\\mathbf\{W\}\},
ℒi\(𝐖\)=ℒi\(𝐖^\)\+⟨𝐆i,𝐖−𝐖^⟩F\+12‖𝐖−𝐖^‖ℋi2,\\mathcal\{L\}\_\{i\}\(\\mathbf\{W\}\)\\;=\\;\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\+\\langle\\mathbf\{G\}\_\{i\},\\,\\mathbf\{W\}\-\\widehat\{\\mathbf\{W\}\}\\rangle\_\{F\}\+\\tfrac\{1\}\{2\}\\,\\bigl\\\|\\mathbf\{W\}\-\\widehat\{\\mathbf\{W\}\}\\bigr\\\|\_\{\\mathcal\{H\}\_\{i\}\}^\{2\},\(73\)where𝐆i≜∇ℒi\(𝐖^\)\\mathbf\{G\}\_\{i\}\\triangleq\\nabla\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)andℋi\\mathcal\{H\}\_\{i\}is the Hessian\. Both are evaluated at the*known*point𝐖^\\widehat\{\\mathbf\{W\}\}and are estimable from the first gradient/curvature probe at initialization; in particular𝐖^\\widehat\{\\mathbf\{W\}\}is*not*assumed to minimizeℒi\\mathcal\{L\}\_\{i\}, so𝐆i\\mathbf\{G\}\_\{i\}is generally nonzero\. The curvature acts as a*direction\-dependent stiffness*, generalizing the scalar smoothness constantLLto an operator recording how sharply the loss bends along each direction\. LetΠri\[⋅\]≜𝐏ri\(⋅\)𝐐ri\\Pi\_\{r\_\{i\}\}\[\\cdot\]\\triangleq\\mathbf\{P\}\_\{r\_\{i\}\}\(\\cdot\)\\mathbf\{Q\}\_\{r\_\{i\}\}be the orthogonal projector ontoℳri≜\{𝐀:𝐏ri𝐀𝐐ri=𝐀\}\\mathcal\{M\}\_\{r\_\{i\}\}\\triangleq\\\{\\mathbf\{A\}:\\mathbf\{P\}\_\{r\_\{i\}\}\\mathbf\{A\}\\mathbf\{Q\}\_\{r\_\{i\}\}=\\mathbf\{A\}\\\}, the set of weights the rank\-rir\_\{i\}compressed model can represent\. If
1. \(a\)the curvature keeps the retained SVD directions decoupled from the discarded ones, equivalently it commutes with the projector, ℋi∘Πri=Πri∘ℋi,\\mathcal\{H\}\_\{i\}\\circ\\Pi\_\{r\_\{i\}\}\\;=\\;\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\},\(74\)
2. \(b\)the loss gradient at the reference has no component inside the retained subspace, Πri\[𝐆i\]=𝐏ri∇ℒi\(𝐖^\)𝐐ri=𝟎,\\Pi\_\{r\_\{i\}\}\[\\mathbf\{G\}\_\{i\}\]=\\mathbf\{P\}\_\{r\_\{i\}\}\\,\\nabla\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\\,\\mathbf\{Q\}\_\{r\_\{i\}\}=\\mathbf\{0\},\(75\)
then the common\-projection\-target condition \([28](https://arxiv.org/html/2607.29071#S4.E28)\) holds:𝚺^i∈argmin𝚺fi\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\in\\arg\\min\_\{\\boldsymbol\{\\Sigma\}\}f\_\{i\}for everyi∈\[K\]i\\in\[K\], and consequentlyεimis=𝟎\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}=\\mathbf\{0\}\. Both hypotheses are testable at initialization from𝐆i\\mathbf\{G\}\_\{i\}andℋi\\mathcal\{H\}\_\{i\}alone, without knowledge of any optimizer\.
###### Proof of Proposition[5](https://arxiv.org/html/2607.29071#Thmproposition5)\.
Throughout,Πri\[𝐌\]=𝐏ri𝐌𝐐ri\\Pi\_\{r\_\{i\}\}\[\\mathbf\{M\}\]=\\mathbf\{P\}\_\{r\_\{i\}\}\\mathbf\{M\}\\mathbf\{Q\}\_\{r\_\{i\}\}is the orthogonal projection \(in the Frobenius inner product\) onto the subspace of representable weightsℳri=\{𝐀:𝐏ri𝐀𝐐ri=𝐀\}\\mathcal\{M\}\_\{r\_\{i\}\}=\\\{\\mathbf\{A\}:\\mathbf\{P\}\_\{r\_\{i\}\}\\mathbf\{A\}\\mathbf\{Q\}\_\{r\_\{i\}\}=\\mathbf\{A\}\\\}\. Recall it is idempotent,Πri2=Πri\\Pi\_\{r\_\{i\}\}^\{2\}=\\Pi\_\{r\_\{i\}\}, and symmetric,⟨Πri𝐀,𝐁⟩F=⟨𝐀,Πri𝐁⟩F\\langle\\Pi\_\{r\_\{i\}\}\\mathbf\{A\},\\mathbf\{B\}\\rangle\_\{F\}=\\langle\\mathbf\{A\},\\Pi\_\{r\_\{i\}\}\\mathbf\{B\}\\rangle\_\{F\}\.
By hypothesisℒi\(𝐀\)=ℒi\(𝐖^\)\+⟨𝐆i,𝐀−𝐖^⟩F\+12⟨𝐀−𝐖^,ℋi\[𝐀−𝐖^\]⟩F\\mathcal\{L\}\_\{i\}\(\\mathbf\{A\}\)=\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\+\\langle\\mathbf\{G\}\_\{i\},\\mathbf\{A\}\-\\widehat\{\\mathbf\{W\}\}\\rangle\_\{F\}\+\\tfrac\{1\}\{2\}\\langle\\mathbf\{A\}\-\\widehat\{\\mathbf\{W\}\},\\,\\mathcal\{H\}\_\{i\}\[\\mathbf\{A\}\-\\widehat\{\\mathbf\{W\}\}\]\\rangle\_\{F\}with𝐆i=∇ℒi\(𝐖^\)\\mathbf\{G\}\_\{i\}=\\nabla\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)andℋi\\mathcal\{H\}\_\{i\}symmetric positive definite, soℒi\\mathcal\{L\}\_\{i\}is a strictly convex quadratic with gradient∇ℒi\(𝐀\)=𝐆i\+ℋi\[𝐀−𝐖^\]\\nabla\\mathcal\{L\}\_\{i\}\(\\mathbf\{A\}\)=\\mathbf\{G\}\_\{i\}\+\\mathcal\{H\}\_\{i\}\[\\mathbf\{A\}\-\\widehat\{\\mathbf\{W\}\}\]and a unique minimizer onℳri\\mathcal\{M\}\_\{r\_\{i\}\}\. A point𝐀∈ℳri\\mathbf\{A\}\\in\\mathcal\{M\}\_\{r\_\{i\}\}is that minimizer exactly when the gradient is orthogonal to all ofℳri\\mathcal\{M\}\_\{r\_\{i\}\}, i\.e\. when its projection onto the subspace vanishes:
Πri\[𝐆i\+ℋi\[𝐀−𝐖^\]\]=𝟎\.\\Pi\_\{r\_\{i\}\}\\bigl\[\\mathbf\{G\}\_\{i\}\+\\mathcal\{H\}\_\{i\}\[\\mathbf\{A\}\-\\widehat\{\\mathbf\{W\}\}\]\\bigr\]=\\mathbf\{0\}\.\(76\)
We now check that the Frobenius projection of the reference,𝐀=Πri\[𝐖^\]\\mathbf\{A\}=\\Pi\_\{r\_\{i\}\}\[\\widehat\{\\mathbf\{W\}\}\], satisfies \([76](https://arxiv.org/html/2607.29071#A1.E76)\)\. By linearity ofΠri\\Pi\_\{r\_\{i\}\}the left\-hand side splits into a gradient term and a curvature term,
Πri\[𝐆i\+ℋi\[𝐀−𝐖^\]\]=Πri\[𝐆i\]⏟=0by \([75](https://arxiv.org/html/2607.29071#A1.E75)\)\+Πri\[ℋi\[𝐀−𝐖^\]\],\\Pi\_\{r\_\{i\}\}\\bigl\[\\mathbf\{G\}\_\{i\}\+\\mathcal\{H\}\_\{i\}\[\\mathbf\{A\}\-\\widehat\{\\mathbf\{W\}\}\]\\bigr\]=\\underbrace\{\\Pi\_\{r\_\{i\}\}\[\\mathbf\{G\}\_\{i\}\]\}\_\{=\\,\\mathbf\{0\}\\ \\text\{by~\\eqref\{eq:hetero\_confined\}\}\}\+\\Pi\_\{r\_\{i\}\}\\bigl\[\\mathcal\{H\}\_\{i\}\[\\mathbf\{A\}\-\\widehat\{\\mathbf\{W\}\}\]\\bigr\],so it remains to show the curvature term also vanishes\. The residual of the candidate is
𝐀−𝐖^=Πri\[𝐖^\]−𝐖^=−\(Id−Πri\)\[𝐖^\],\\mathbf\{A\}\-\\widehat\{\\mathbf\{W\}\}=\\Pi\_\{r\_\{i\}\}\[\\widehat\{\\mathbf\{W\}\}\]\-\\widehat\{\\mathbf\{W\}\}=\-\(\\mathrm\{Id\}\-\\Pi\_\{r\_\{i\}\}\)\[\\widehat\{\\mathbf\{W\}\}\],which lies in the orthogonal complement ofℳri\\mathcal\{M\}\_\{r\_\{i\}\}: projecting it back gives\(Πri∘\(Id−Πri\)\)\[𝐖^\]=\(Πri−Πri2\)\[𝐖^\]=𝟎\\bigl\(\\Pi\_\{r\_\{i\}\}\\circ\(\\mathrm\{Id\}\-\\Pi\_\{r\_\{i\}\}\)\\bigr\)\[\\widehat\{\\mathbf\{W\}\}\]=\(\\Pi\_\{r\_\{i\}\}\-\\Pi\_\{r\_\{i\}\}^\{2\}\)\[\\widehat\{\\mathbf\{W\}\}\]=\\mathbf\{0\}by idempotence\. Applying the alignment hypothesis \([74](https://arxiv.org/html/2607.29071#A1.E74)\),ℋi∘Πri=Πri∘ℋi\\mathcal\{H\}\_\{i\}\\circ\\Pi\_\{r\_\{i\}\}=\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\}, to pass the projection through the curvature,
Πri\[ℋi\[𝐀−𝐖^\]\]=−\(Πri∘ℋi∘\(Id−Πri\)\)\[𝐖^\]=−ℋi\[\(Πri∘\(Id−Πri\)\)\[𝐖^\]⏟=0\]=𝟎,\\begin\{split\}\\Pi\_\{r\_\{i\}\}\\bigl\[\\mathcal\{H\}\_\{i\}\[\\mathbf\{A\}\-\\widehat\{\\mathbf\{W\}\}\]\\bigr\]&=\-\\bigl\(\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\}\\circ\(\\mathrm\{Id\}\-\\Pi\_\{r\_\{i\}\}\)\\bigr\)\[\\widehat\{\\mathbf\{W\}\}\]\\\\ &=\-\\,\\mathcal\{H\}\_\{i\}\\Bigl\[\\,\\underbrace\{\\bigl\(\\Pi\_\{r\_\{i\}\}\\circ\(\\mathrm\{Id\}\-\\Pi\_\{r\_\{i\}\}\)\\bigr\)\[\\widehat\{\\mathbf\{W\}\}\]\}\_\{=\\,\\mathbf\{0\}\}\\,\\Bigr\]=\\mathbf\{0\},\\end\{split\}the second equality movingΠri\\Pi\_\{r\_\{i\}\}pastℋi\\mathcal\{H\}\_\{i\}by the hypothesis\. Both terms of \([76](https://arxiv.org/html/2607.29071#A1.E76)\) thus vanish, soΠri\[𝐖^\]\\Pi\_\{r\_\{i\}\}\[\\widehat\{\\mathbf\{W\}\}\]is the minimizer ofℒi\\mathcal\{L\}\_\{i\}onℳri\\mathcal\{M\}\_\{r\_\{i\}\}by uniqueness\.
It remains to translate this back to adapters\. By \([69](https://arxiv.org/html/2607.29071#A1.E69)\) the minimizerΠri\[𝐖^\]=𝐏ri𝐖^𝐐ri\\Pi\_\{r\_\{i\}\}\[\\widehat\{\\mathbf\{W\}\}\]=\\mathbf\{P\}\_\{r\_\{i\}\}\\widehat\{\\mathbf\{W\}\}\\mathbf\{Q\}\_\{r\_\{i\}\}equals𝐔ri𝚺^i𝐕ri\\mathbf\{U\}\_\{r\_\{i\}\}\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\mathbf\{V\}\_\{r\_\{i\}\}, so it is realized by the geometric adapter𝚺^i\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\. Pulling back through the bijection,𝚺^i\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}minimizesfif\_\{i\}, which is the common\-projection\-target condition \([28](https://arxiv.org/html/2607.29071#S4.E28)\); as the minimizer is unique,𝚺i∗=𝚺^i\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}=\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}and henceεimis=𝐔ri\(𝚺i∗−𝚺^i\)𝐕ri=𝟎\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}=\\mathbf\{U\}\_\{r\_\{i\}\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}\-\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\)\\mathbf\{V\}\_\{r\_\{i\}\}=\\mathbf\{0\}\.
For the separable curvatureℋi\[𝐌\]=𝐀i𝐌𝐂i\\mathcal\{H\}\_\{i\}\[\\mathbf\{M\}\]=\\mathbf\{A\}\_\{i\}\\mathbf\{M\}\\mathbf\{C\}\_\{i\}of the main text, the alignment hypothesis \([74](https://arxiv.org/html/2607.29071#A1.E74)\) holds as soon as𝐀i𝐏ri=𝐏ri𝐀i\\mathbf\{A\}\_\{i\}\\mathbf\{P\}\_\{r\_\{i\}\}=\\mathbf\{P\}\_\{r\_\{i\}\}\\mathbf\{A\}\_\{i\}and𝐂i𝐐ri=𝐐ri𝐂i\\mathbf\{C\}\_\{i\}\\mathbf\{Q\}\_\{r\_\{i\}\}=\\mathbf\{Q\}\_\{r\_\{i\}\}\\mathbf\{C\}\_\{i\}\(equivalently, the left and right singular vectors of𝐖^\\widehat\{\\mathbf\{W\}\}are eigenvectors of the matrices𝐀i\\mathbf\{A\}\_\{i\}and𝐂i\\mathbf\{C\}\_\{i\}\), since then for every𝐌\\mathbf\{M\}
\(ℋi∘Πri\)\[𝐌\]=𝐀i𝐏ri𝐌𝐐ri𝐂i=𝐏ri𝐀i𝐌𝐂i𝐐ri=\(Πri∘ℋi\)\[𝐌\],\\begin\{split\}\(\\mathcal\{H\}\_\{i\}\\circ\\Pi\_\{r\_\{i\}\}\)\[\\mathbf\{M\}\]&=\\mathbf\{A\}\_\{i\}\\mathbf\{P\}\_\{r\_\{i\}\}\\mathbf\{M\}\\mathbf\{Q\}\_\{r\_\{i\}\}\\mathbf\{C\}\_\{i\}=\\mathbf\{P\}\_\{r\_\{i\}\}\\mathbf\{A\}\_\{i\}\\mathbf\{M\}\\mathbf\{C\}\_\{i\}\\mathbf\{Q\}\_\{r\_\{i\}\}\\\\ &=\(\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\}\)\[\\mathbf\{M\}\],\\end\{split\}using𝐀i𝐏ri=𝐏ri𝐀i\\mathbf\{A\}\_\{i\}\\mathbf\{P\}\_\{r\_\{i\}\}=\\mathbf\{P\}\_\{r\_\{i\}\}\\mathbf\{A\}\_\{i\}and𝐂i𝐐ri=𝐐ri𝐂i\\mathbf\{C\}\_\{i\}\\mathbf\{Q\}\_\{r\_\{i\}\}=\\mathbf\{Q\}\_\{r\_\{i\}\}\\mathbf\{C\}\_\{i\}in the middle step\. The isotropic loss \(𝐀i=ai𝐈m\\mathbf\{A\}\_\{i\}=a\_\{i\}\\mathbf\{I\}\_\{m\},𝐂i=𝐈n\\mathbf\{C\}\_\{i\}=\\mathbf\{I\}\_\{n\}\) is the trivial case, where every matrix commutes with the projection matrices\. ∎
The two hypotheses control the two ways the loss optimum can drift away from the geometric projectionΠri\[𝐖^\]\\Pi\_\{r\_\{i\}\}\[\\widehat\{\\mathbf\{W\}\}\], and both are read off at the reference𝐖^\\widehat\{\\mathbf\{W\}\}\. Condition \([75](https://arxiv.org/html/2607.29071#A1.E75)\) is a*subspace\-stationarity*requirement: at𝐖^\\widehat\{\\mathbf\{W\}\}, the loss has no first\-order incentive to move inside the retained subspace, so all of its gradient pressure points along the discarded directions\. Equivalently, writing𝐆i=ℋi\[𝐖^−𝐖iloc\]\\mathbf\{G\}\_\{i\}=\\mathcal\{H\}\_\{i\}\[\\widehat\{\\mathbf\{W\}\}\-\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\]for the unconstrained minimizer𝐖iloc\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}, the condition says the group heterogeneity𝐖^−𝐖iloc\\widehat\{\\mathbf\{W\}\}\-\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}lies entirely along the discarded directions \(where compression removes it anyway\), without forcing𝐖iloc=𝐖^\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}=\\widehat\{\\mathbf\{W\}\}\(the no\-heterogeneity caseHi=0H\_\{i\}=0of Proposition[7](https://arxiv.org/html/2607.29071#Thmproposition7)\)\. Condition \([74](https://arxiv.org/html/2607.29071#A1.E74)\) controls the metric: it annihilates the off\-block coupling\(Πri∘ℋi∘\(Id−Πri\)\)\\bigl\(\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\}\\circ\(\\mathrm\{Id\}\-\\Pi\_\{r\_\{i\}\}\)\\bigr\)through which the curvature would otherwise pull the loss optimum off the Frobenius projection\. It constrains only the curvature*directions*, not magnitudes, so it permits arbitrary, group\-dependent, ill\-conditioned anisotropy along the singular directions of𝐖^\\widehat\{\\mathbf\{W\}\}; whitened/isotropic curvature is the trivial case where the coupling vanishes automatically\.
A concrete example is the Gauss–Newton/Fisher curvature of a layer, which K\-FAC\[[36](https://arxiv.org/html/2607.29071#bib.bib87)\]approximates in the separable formℋi\[𝐌\]=𝐀i𝐌𝐂i\\mathcal\{H\}\_\{i\}\[\\mathbf\{M\}\]=\\mathbf\{A\}\_\{i\}\\mathbf\{M\}\\mathbf\{C\}\_\{i\}with𝐀i\\mathbf\{A\}\_\{i\}and𝐂i\\mathbf\{C\}\_\{i\}the output\-gradient and input\-activation covariances\. Condition \([74](https://arxiv.org/html/2607.29071#A1.E74)\) then holds whenever these K\-FAC factors are co\-diagonalizable with the layer’s SVD subspaces, i\.e\. the curvature shares its principal axes with𝐖^\\widehat\{\\mathbf\{W\}\}, while its magnitude along each axis may vary arbitrarily and unequally across groups\. The isotropic loss is the trivial case in which this alignment is automatic\.
For curvature that genuinely couples the retained subspace to its complement \(the typical regime in federated fine\-tuning, whereℋi\\mathcal\{H\}\_\{i\}has no reason to respect the SVD geometry of𝐖^\\widehat\{\\mathbf\{W\}\}\), condition \([74](https://arxiv.org/html/2607.29071#A1.E74)\) fails and the loss\-projection departs from the Frobenius\-projection\. Proposition[5](https://arxiv.org/html/2607.29071#Thmproposition5)therefore serves as a clean limiting case rather than a typical operating point, and on its own provides no quantitative handle on the misalignment when its hypotheses are violated\. Two further questions remain\. First, the conditions \([74](https://arxiv.org/html/2607.29071#A1.E74)\)–\([75](https://arxiv.org/html/2607.29071#A1.E75)\) almost never hold exactly: how should the misalignment grow as they are violated by small but non\-zero amounts? Second, the loss landscape of foundation\-model fine\-tuning is non\-convex, so the quadratic form \([73](https://arxiv.org/html/2607.29071#A1.E73)\) is at best a second\-order Taylor expansion about𝐖^\\widehat\{\\mathbf\{W\}\}; the cubic remainder must be controlled before the conclusion is meaningful for the true loss\. The next proposition addresses both: it relaxes the two structural conditions to approximate forms with explicit error budgets, augments them with a Hessian\-variation bound that absorbs the truncation error from the quadratic surrogate, and yields a perturbation bound onεimis\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}that recovers Proposition[5](https://arxiv.org/html/2607.29071#Thmproposition5)exactly when all errors vanish\.
###### Proposition 6\(Perturbation\-robust misalignment bound\)\.
Fix the reference weight𝐖^\\widehat\{\\mathbf\{W\}\}, and letℒi:ℝm×n→ℝ\\mathcal\{L\}\_\{i\}:\\mathbb\{R\}^\{m\\times n\}\\to\\mathbb\{R\}be twice continuously differentiable\. LetΠri\\Pi\_\{r\_\{i\}\}denote the orthogonal projector ontoℳri\\mathcal\{M\}\_\{r\_\{i\}\}, write𝐖^⟂≜\(Id−Πri\)\[𝐖^\]\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\\triangleq\(\\mathrm\{Id\}\-\\Pi\_\{r\_\{i\}\}\)\[\\widehat\{\\mathbf\{W\}\}\]andRi≜‖𝐖^⟂‖F=δri\(𝐖^\)R\_\{i\}\\triangleq\\\|\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\\\|\_\{F\}=\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\), and set𝐀^i≜Πri\[𝐖^\]\\widehat\{\\mathbf\{A\}\}\_\{i\}\\triangleq\\Pi\_\{r\_\{i\}\}\[\\widehat\{\\mathbf\{W\}\}\]\. Define the perturbation budgets
εG,i\\displaystyle\\varepsilon\_\{G,i\}≜‖Πri\[∇ℒi\(𝐖^\)\]‖F,\\displaystyle\\triangleq\\bigl\\\|\\Pi\_\{r\_\{i\}\}\\bigl\[\\nabla\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\\bigr\]\\bigr\\\|\_\{F\},\(77\)εH,i\\displaystyle\\varepsilon\_\{H,i\}≜‖ℋi∘Πri−Πri∘ℋi‖op,\\displaystyle\\triangleq\\bigl\\\|\\mathcal\{H\}\_\{i\}\\circ\\Pi\_\{r\_\{i\}\}\-\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\}\\bigr\\\|\_\{\\mathrm\{op\}\},\(78\)whereℋi≜∇2ℒi\(𝐖^\)\\mathcal\{H\}\_\{i\}\\triangleq\\nabla^\{2\}\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)and∥⋅∥op\\\|\\cdot\\\|\_\{\\mathrm\{op\}\}denotes the operator norm induced by∥⋅∥F\\\|\\cdot\\\|\_\{F\}\. Suppose:
1. \(c1\)\(Restricted strong convexity onℳri\\mathcal\{M\}\_\{r\_\{i\}\}\.\) There existsμi\>0\\mu\_\{i\}\>0such that for all𝐀,𝐁∈ℳri\\mathbf\{A\},\\mathbf\{B\}\\in\\mathcal\{M\}\_\{r\_\{i\}\}, ⟨∇ℒi\(𝐀\)−∇ℒi\(𝐁\),𝐀−𝐁⟩F≥μi‖𝐀−𝐁‖F2\.\\bigl\\langle\\nabla\\mathcal\{L\}\_\{i\}\(\\mathbf\{A\}\)\-\\nabla\\mathcal\{L\}\_\{i\}\(\\mathbf\{B\}\),\\,\\mathbf\{A\}\-\\mathbf\{B\}\\bigr\\rangle\_\{F\}\\;\\geq\\;\\mu\_\{i\}\\,\\\|\\mathbf\{A\}\-\\mathbf\{B\}\\\|\_\{F\}^\{\\,2\}\.\(79\)
2. \(c2\)\(Approximate subspace\-stationarity\.\) The reference gradient has small retained\-subspace component,εG,i<∞\\varepsilon\_\{G,i\}<\\infty\.
3. \(c3\)\(Approximate curvature alignment\.\) The Hessian at𝐖^\\widehat\{\\mathbf\{W\}\}approximately commutes with the projector,εH,i<∞\\varepsilon\_\{H,i\}<\\infty\.
4. \(c4\)\(Hessian variation / cubic remainder control\.\) There existsρi≥0\\rho\_\{i\}\\geq 0and a neighbourhood𝒩i⊃\{𝐖^\+s\(𝐀−𝐖^\):𝐀∈ℳri,‖𝐀−𝐀^i‖F≤Ri,s∈\[0,1\]\}\\mathcal\{N\}\_\{i\}\\supset\\\{\\widehat\{\\mathbf\{W\}\}\+s\(\\mathbf\{A\}\-\\widehat\{\\mathbf\{W\}\}\):\\mathbf\{A\}\\in\\mathcal\{M\}\_\{r\_\{i\}\},\\ \\\|\\mathbf\{A\}\-\\widehat\{\\mathbf\{A\}\}\_\{i\}\\\|\_\{F\}\\leq R\_\{i\},\\ s\\in\[0,1\]\\\}such that for every𝐗∈𝒩i\\mathbf\{X\}\\in\\mathcal\{N\}\_\{i\}, ‖∇2ℒi\(𝐗\)−ℋi‖op≤ρi‖𝐗−𝐖^‖F\.\\bigl\\\|\\nabla^\{2\}\\mathcal\{L\}\_\{i\}\(\\mathbf\{X\}\)\-\\mathcal\{H\}\_\{i\}\\bigr\\\|\_\{\\mathrm\{op\}\}\\;\\leq\\;\\rho\_\{i\}\\,\\\|\\mathbf\{X\}\-\\widehat\{\\mathbf\{W\}\}\\\|\_\{F\}\.\(80\)
Let𝐀i∗≜argmin𝐀∈ℳriℒi\(𝐀\)\\mathbf\{A\}\_\{i\}^\{\*\}\\triangleq\\arg\\min\_\{\\mathbf\{A\}\\in\\mathcal\{M\}\_\{r\_\{i\}\}\}\\mathcal\{L\}\_\{i\}\(\\mathbf\{A\}\)\(unique by \([79](https://arxiv.org/html/2607.29071#A1.E79)\)\), and assume the stability condition4ρi\(εG,i\+εH,iRi\)<μi24\\rho\_\{i\}\\bigl\(\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}R\_\{i\}\\bigr\)<\\mu\_\{i\}^\{2\}\. Then
‖𝐀i∗−𝐀^i‖F≤εG,i\+εH,iRiμi\+ρiRi22μi\+𝒪\(ρiμi3\(εG,i\+εH,iRi\)2\),\\begin\{split\}\\bigl\\\|\\mathbf\{A\}\_\{i\}^\{\*\}\-\\widehat\{\\mathbf\{A\}\}\_\{i\}\\bigr\\\|\_\{F\}\\;\\leq\\;&\\frac\{\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}\\,R\_\{i\}\}\{\\mu\_\{i\}\}\\;\+\\;\\frac\{\\rho\_\{i\}\\,R\_\{i\}^\{2\}\}\{2\\,\\mu\_\{i\}\}\\\\ &\\;\+\\;\\mathcal\{O\}\\\!\\left\(\\frac\{\\rho\_\{i\}\}\{\\mu\_\{i\}^\{3\}\}\\bigl\(\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}R\_\{i\}\\bigr\)^\{2\}\\right\),\\end\{split\}\(81\)and consequently the misalignment residual is bounded by
‖εimis‖F≤εG,i\+εH,iRiμi\+ρiRi22μi\+𝒪\(ρiμi3\(εG,i\+εH,iRi\)2\)\.\\begin\{split\}\\bigl\\\|\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}\\bigr\\\|\_\{F\}\\;\\leq\\;&\\frac\{\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}\\,R\_\{i\}\}\{\\mu\_\{i\}\}\\;\+\\;\\frac\{\\rho\_\{i\}\\,R\_\{i\}^\{2\}\}\{2\\,\\mu\_\{i\}\}\\\\ &\\;\+\\;\\mathcal\{O\}\\\!\\left\(\\frac\{\\rho\_\{i\}\}\{\\mu\_\{i\}^\{3\}\}\\bigl\(\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}R\_\{i\}\\bigr\)^\{2\}\\right\)\.\\end\{split\}\(82\)WhenεG,i=εH,i=ρi=0\\varepsilon\_\{G,i\}=\\varepsilon\_\{H,i\}=\\rho\_\{i\}=0, the bound reads‖𝐀i∗−𝐀^i‖F=0\\\|\\mathbf\{A\}\_\{i\}^\{\*\}\-\\widehat\{\\mathbf\{A\}\}\_\{i\}\\\|\_\{F\}=0, recovering Proposition[5](https://arxiv.org/html/2607.29071#Thmproposition5)as the exact\-equality limit\.
###### Proof of Proposition[6](https://arxiv.org/html/2607.29071#Thmproposition6)\.
We work directly in the matrix spaceℝm×n\\mathbb\{R\}^\{m\\times n\}equipped with the Frobenius inner product\. Let𝐀i∗∈ℳri\\mathbf\{A\}\_\{i\}^\{\*\}\\in\\mathcal\{M\}\_\{r\_\{i\}\}be the unique minimizer \(uniqueness is immediate from \([79](https://arxiv.org/html/2607.29071#A1.E79)\)\) and writeΔ≜𝐀i∗−𝐀^i∈ℳri\\Delta\\triangleq\\mathbf\{A\}\_\{i\}^\{\*\}\-\\widehat\{\\mathbf\{A\}\}\_\{i\}\\in\\mathcal\{M\}\_\{r\_\{i\}\}, soΠri\[Δ\]=Δ\\Pi\_\{r\_\{i\}\}\[\\Delta\]=\\Delta\.
Since𝐀i∗\\mathbf\{A\}\_\{i\}^\{\*\}minimizesℒi\\mathcal\{L\}\_\{i\}onℳri\\mathcal\{M\}\_\{r\_\{i\}\}, the gradient ofℒi\\mathcal\{L\}\_\{i\}at𝐀i∗\\mathbf\{A\}\_\{i\}^\{\*\}is Frobenius\-orthogonal to every direction inℳri\\mathcal\{M\}\_\{r\_\{i\}\}, equivalently
Πri\[∇ℒi\(𝐀i∗\)\]=𝟎\.\\Pi\_\{r\_\{i\}\}\\bigl\[\\nabla\\mathcal\{L\}\_\{i\}\(\\mathbf\{A\}\_\{i\}^\{\*\}\)\\bigr\]=\\mathbf\{0\}\.\(83\)
For any𝐗∈ℝm×n\\mathbf\{X\}\\in\\mathbb\{R\}^\{m\\times n\}, the integral form of Taylor’s theorem applied to∇ℒi\\nabla\\mathcal\{L\}\_\{i\}along the segment𝐖^→𝐗\\widehat\{\\mathbf\{W\}\}\\to\\mathbf\{X\}gives
∇ℒi\(𝐗\)=∇ℒi\(𝐖^\)\+ℋi\[𝐗−𝐖^\]\+𝐑\(𝐗\),\\nabla\\mathcal\{L\}\_\{i\}\(\\mathbf\{X\}\)\\;=\\;\\nabla\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\+\\mathcal\{H\}\_\{i\}\[\\mathbf\{X\}\-\\widehat\{\\mathbf\{W\}\}\]\+\\mathbf\{R\}\(\\mathbf\{X\}\),\(84\)with remainder𝐑\(𝐗\)≜∫01\(∇2ℒi\(𝐖^\+s\(𝐗−𝐖^\)\)−ℋi\)\[𝐗−𝐖^\]𝑑s\\mathbf\{R\}\(\\mathbf\{X\}\)\\triangleq\\int\_\{0\}^\{1\}\\bigl\(\\nabla^\{2\}\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\+s\(\\mathbf\{X\}\-\\widehat\{\\mathbf\{W\}\}\)\)\-\\mathcal\{H\}\_\{i\}\\bigr\)\[\\mathbf\{X\}\-\\widehat\{\\mathbf\{W\}\}\]\\,ds\. Applying \([80](https://arxiv.org/html/2607.29071#A1.E80)\) pointwise on the segment, which lies in𝒩i\\mathcal\{N\}\_\{i\}for𝐗∈𝐀^i\+ball\(Ri\)\\mathbf\{X\}\\in\\widehat\{\\mathbf\{A\}\}\_\{i\}\+\\mathrm\{ball\}\(R\_\{i\}\)onℳri\\mathcal\{M\}\_\{r\_\{i\}\}, yields the size bound
‖𝐑\(𝐗\)‖F≤ρi∫01s‖𝐗−𝐖^‖F𝑑s⋅‖𝐗−𝐖^‖F=ρi2‖𝐗−𝐖^‖F2\.\\begin\{split\}\\\|\\mathbf\{R\}\(\\mathbf\{X\}\)\\\|\_\{F\}&\\;\\leq\\;\\rho\_\{i\}\\int\_\{0\}^\{1\}s\\,\\\|\\mathbf\{X\}\-\\widehat\{\\mathbf\{W\}\}\\\|\_\{F\}\\;ds\\cdot\\\|\\mathbf\{X\}\-\\widehat\{\\mathbf\{W\}\}\\\|\_\{F\}\\\\ &\\;=\\;\\tfrac\{\\rho\_\{i\}\}\{2\}\\,\\\|\\mathbf\{X\}\-\\widehat\{\\mathbf\{W\}\}\\\|\_\{F\}^\{\\,2\}\.\\end\{split\}\(85\)Setting𝐗=𝐀i∗=𝐀^i\+Δ\\mathbf\{X\}=\\mathbf\{A\}\_\{i\}^\{\*\}=\\widehat\{\\mathbf\{A\}\}\_\{i\}\+\\Deltaand using𝐀i∗−𝐖^=\(𝐀^i−𝐖^\)\+Δ=−𝐖^⟂\+Δ\\mathbf\{A\}\_\{i\}^\{\*\}\-\\widehat\{\\mathbf\{W\}\}=\(\\widehat\{\\mathbf\{A\}\}\_\{i\}\-\\widehat\{\\mathbf\{W\}\}\)\+\\Delta=\-\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\+\\Delta:
∇ℒi\(𝐀i∗\)=𝐆i−ℋi\[𝐖^⟂\]\+ℋi\[Δ\]\+𝐑\(𝐀i∗\),\\nabla\\mathcal\{L\}\_\{i\}\(\\mathbf\{A\}\_\{i\}^\{\*\}\)\\;=\\;\\mathbf\{G\}\_\{i\}\\;\-\\;\\mathcal\{H\}\_\{i\}\[\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\]\\;\+\\;\\mathcal\{H\}\_\{i\}\[\\Delta\]\\;\+\\;\\mathbf\{R\}\(\\mathbf\{A\}\_\{i\}^\{\*\}\),\(86\)where𝐆i=∇ℒi\(𝐖^\)\\mathbf\{G\}\_\{i\}=\\nabla\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)and‖𝐑\(𝐀i∗\)‖F≤ρi2‖−𝐖^⟂\+Δ‖F2≤ρi2\(Ri\+‖Δ‖F\)2\\\|\\mathbf\{R\}\(\\mathbf\{A\}\_\{i\}^\{\*\}\)\\\|\_\{F\}\\leq\\tfrac\{\\rho\_\{i\}\}\{2\}\\\|\{\-\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\}\+\\Delta\\\|\_\{F\}^\{\\,2\}\\leq\\tfrac\{\\rho\_\{i\}\}\{2\}\(R\_\{i\}\+\\\|\\Delta\\\|\_\{F\}\)^\{2\}\.
ApplyingΠri\\Pi\_\{r\_\{i\}\}to \([86](https://arxiv.org/html/2607.29071#A1.E86)\) and combining with \([83](https://arxiv.org/html/2607.29071#A1.E83)\):
𝟎=Πri\[𝐆i\]−Πri\[ℋi\[𝐖^⟂\]\]\+Πri\[ℋi\[Δ\]\]\+Πri\[𝐑\(𝐀i∗\)\]\.\\begin\{split\}\\mathbf\{0\}\\;=\\;&\\Pi\_\{r\_\{i\}\}\[\\mathbf\{G\}\_\{i\}\]\\;\-\\;\\Pi\_\{r\_\{i\}\}\\\!\\bigl\[\\mathcal\{H\}\_\{i\}\[\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\]\\bigr\]\\\\ &\+\\;\\Pi\_\{r\_\{i\}\}\\\!\\bigl\[\\mathcal\{H\}\_\{i\}\[\\Delta\]\\bigr\]\\;\+\\;\\Pi\_\{r\_\{i\}\}\[\\mathbf\{R\}\(\\mathbf\{A\}\_\{i\}^\{\*\}\)\]\.\\end\{split\}\(87\)The middle term is the only place the alignment defect can enter\. Using the commutator identity,Πri∘ℋi=ℋi∘Πri\+\(Πri∘ℋi−ℋi∘Πri\)\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\}=\\mathcal\{H\}\_\{i\}\\circ\\Pi\_\{r\_\{i\}\}\+\(\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\}\-\\mathcal\{H\}\_\{i\}\\circ\\Pi\_\{r\_\{i\}\}\), and noting thatΠri\[𝐖^⟂\]=𝟎\\Pi\_\{r\_\{i\}\}\[\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\]=\\mathbf\{0\}\(since𝐖^⟂\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}is in the orthogonal complement ofℳri\\mathcal\{M\}\_\{r\_\{i\}\}\),
Πri\[ℋi\[𝐖^⟂\]\]=ℋi\[Πri\[𝐖^⟂\]\]\+\(Πri∘ℋi−ℋi∘Πri\)\[𝐖^⟂\]=\(Πri∘ℋi−ℋi∘Πri\)\[𝐖^⟂\],\\begin\{split\}\\Pi\_\{r\_\{i\}\}\\\!\\bigl\[\\mathcal\{H\}\_\{i\}\[\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\]\\bigr\]&=\\mathcal\{H\}\_\{i\}\\\!\\bigl\[\\Pi\_\{r\_\{i\}\}\[\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\]\\bigr\]\\\\ &\\quad\+\(\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\}\-\\mathcal\{H\}\_\{i\}\\circ\\Pi\_\{r\_\{i\}\}\)\[\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\]\\\\ &=\(\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\}\-\\mathcal\{H\}\_\{i\}\\circ\\Pi\_\{r\_\{i\}\}\)\[\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\],\\end\{split\}\(88\)whose Frobenius norm is at mostεH,i‖𝐖^⟂‖F=εH,iRi\\varepsilon\_\{H,i\}\\,\\\|\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\\\|\_\{F\}=\\varepsilon\_\{H,i\}\\,R\_\{i\}by definition of the operator norm and \([78](https://arxiv.org/html/2607.29071#A1.E78)\)\. Substituting back:
Πri\[ℋi\[Δ\]\]=−Πri\[𝐆i\]−Πri\[𝐑\(𝐀i∗\)\]\+\(Πri∘ℋi−ℋi∘Πri\)\[𝐖^⟂\]\.\\begin\{split\}\\Pi\_\{r\_\{i\}\}\\\!\\bigl\[\\mathcal\{H\}\_\{i\}\[\\Delta\]\\bigr\]\\;=\\;&\-\\Pi\_\{r\_\{i\}\}\[\\mathbf\{G\}\_\{i\}\]\\;\-\\;\\Pi\_\{r\_\{i\}\}\[\\mathbf\{R\}\(\\mathbf\{A\}\_\{i\}^\{\*\}\)\]\\\\ &\+\\;\(\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\}\-\\mathcal\{H\}\_\{i\}\\circ\\Pi\_\{r\_\{i\}\}\)\[\\widehat\{\\mathbf\{W\}\}\_\{\\perp\}\]\.\\end\{split\}\(89\)
The condition \([79](https://arxiv.org/html/2607.29071#A1.E79)\) applied to𝐀=𝐀^i\+tΔ\\mathbf\{A\}=\\widehat\{\\mathbf\{A\}\}\_\{i\}\+t\\Deltaand𝐁=𝐀^i\\mathbf\{B\}=\\widehat\{\\mathbf\{A\}\}\_\{i\}inℳri\\mathcal\{M\}\_\{r\_\{i\}\}, divided byt2t^\{2\}and passed to the limitt→0t\\to 0, yields the infinitesimal form
⟨Δ,∇2ℒi\(𝐀^i\)\[Δ\]⟩F≥μi‖Δ‖F2\.\\bigl\\langle\\Delta,\\,\\nabla^\{2\}\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{A\}\}\_\{i\}\)\[\\Delta\]\\bigr\\rangle\_\{F\}\\;\\geq\\;\\mu\_\{i\}\\,\\\|\\Delta\\\|\_\{F\}^\{\\,2\}\.\(90\)Replacing∇2ℒi\(𝐀^i\)\\nabla^\{2\}\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{A\}\}\_\{i\}\)byℋi=∇2ℒi\(𝐖^\)\\mathcal\{H\}\_\{i\}=\\nabla^\{2\}\\mathcal\{L\}\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)costs at mostρi‖𝐀^i−𝐖^‖F=ρiRi\\rho\_\{i\}\\\|\\widehat\{\\mathbf\{A\}\}\_\{i\}\-\\widehat\{\\mathbf\{W\}\}\\\|\_\{F\}=\\rho\_\{i\}R\_\{i\}in operator norm by \([80](https://arxiv.org/html/2607.29071#A1.E80)\), so
⟨Δ,ℋi\[Δ\]⟩F≥\(μi−ρiRi\)‖Δ‖F2\.\\bigl\\langle\\Delta,\\,\\mathcal\{H\}\_\{i\}\[\\Delta\]\\bigr\\rangle\_\{F\}\\;\\geq\\;\(\\mu\_\{i\}\-\\rho\_\{i\}R\_\{i\}\)\\,\\\|\\Delta\\\|\_\{F\}^\{\\,2\}\.\(91\)UsingΠri\[Δ\]=Δ\\Pi\_\{r\_\{i\}\}\[\\Delta\]=\\Deltaand the self\-adjointness ofΠri\\Pi\_\{r\_\{i\}\},
⟨Δ,ℋi\[Δ\]⟩F=⟨Πri\[Δ\],ℋi\[Δ\]⟩F=⟨Δ,Πri\[ℋi\[Δ\]\]⟩F\.\\hskip\-9\.79996pt\\langle\\Delta,\\,\\mathcal\{H\}\_\{i\}\[\\Delta\]\\rangle\_\{F\}=\\langle\\Pi\_\{r\_\{i\}\}\[\\Delta\],\\mathcal\{H\}\_\{i\}\[\\Delta\]\\rangle\_\{F\}=\\bigl\\langle\\Delta,\\,\\Pi\_\{r\_\{i\}\}\\\!\\bigl\[\\mathcal\{H\}\_\{i\}\[\\Delta\]\\bigr\]\\bigr\\rangle\_\{F\}\.\(92\)Cauchy–Schwarz then converts \([91](https://arxiv.org/html/2607.29071#A1.E91)\) into the operator\-norm lower bound
‖Πri\[ℋi\[Δ\]\]‖F≥\(μi−ρiRi\)‖Δ‖F\.\\bigl\\\|\\Pi\_\{r\_\{i\}\}\\\!\\bigl\[\\mathcal\{H\}\_\{i\}\[\\Delta\]\\bigr\]\\bigr\\\|\_\{F\}\\;\\geq\\;\(\\mu\_\{i\}\-\\rho\_\{i\}R\_\{i\}\)\\,\\\|\\Delta\\\|\_\{F\}\.\(93\)
Taking Frobenius norms of \([89](https://arxiv.org/html/2607.29071#A1.E89)\), applying the triangle inequality with \([77](https://arxiv.org/html/2607.29071#A1.E77)\), \([88](https://arxiv.org/html/2607.29071#A1.E88)\), and the remainder bound \([85](https://arxiv.org/html/2607.29071#A1.E85)\),
‖Πri\[ℋi\[Δ\]\]‖F≤εG,i\+εH,iRi\+ρi2\(Ri\+‖Δ‖F\)2\.\\bigl\\\|\\Pi\_\{r\_\{i\}\}\\\!\\bigl\[\\mathcal\{H\}\_\{i\}\[\\Delta\]\\bigr\]\\bigr\\\|\_\{F\}\\;\\leq\\;\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}R\_\{i\}\+\\tfrac\{\\rho\_\{i\}\}\{2\}\(R\_\{i\}\+\\\|\\Delta\\\|\_\{F\}\)^\{2\}\.\(94\)Combining with the lower bound \([93](https://arxiv.org/html/2607.29071#A1.E93)\):
\(μi−ρiRi\)‖Δ‖F≤εG,i\+εH,iRi\+ρi2\(Ri\+‖Δ‖F\)2\.\(\\mu\_\{i\}\-\\rho\_\{i\}R\_\{i\}\)\\,\\\|\\Delta\\\|\_\{F\}\\;\\leq\\;\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}R\_\{i\}\+\\tfrac\{\\rho\_\{i\}\}\{2\}\(R\_\{i\}\+\\\|\\Delta\\\|\_\{F\}\)^\{2\}\.\(95\)Expanding the square and abbreviatingc≜εG,i\+εH,iRi\+ρi2Ri2c\\triangleq\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}R\_\{i\}\+\\tfrac\{\\rho\_\{i\}\}\{2\}R\_\{i\}^\{2\},β≜μi−ρiRi−ρiRi=μi−2ρiRi\\beta\\triangleq\\mu\_\{i\}\-\\rho\_\{i\}R\_\{i\}\-\\rho\_\{i\}R\_\{i\}=\\mu\_\{i\}\-2\\rho\_\{i\}R\_\{i\}\(the coefficient of‖Δ‖F\\\|\\Delta\\\|\_\{F\}on the left after moving the cross termρiRi‖Δ‖F\\rho\_\{i\}R\_\{i\}\\\|\\Delta\\\|\_\{F\}to the left\):
β‖Δ‖F≤c\+ρi2‖Δ‖F2\.\\beta\\,\\\|\\Delta\\\|\_\{F\}\\;\\leq\\;c\\;\+\\;\\tfrac\{\\rho\_\{i\}\}\{2\}\\,\\\|\\Delta\\\|\_\{F\}^\{\\,2\}\.\(96\)Under the stability condition4ρi\(εG,i\+εH,iRi\)<μi24\\rho\_\{i\}\(\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}R\_\{i\}\)<\\mu\_\{i\}^\{2\}, which \(usingRi≤‖𝐖^‖FR\_\{i\}\\leq\\\|\\widehat\{\\mathbf\{W\}\}\\\|\_\{F\}and the assumption that the perturbation budgets are small relative toμi\\mu\_\{i\}\) impliesρi‖Δ‖F≤μi/2\\rho\_\{i\}\\\|\\Delta\\\|\_\{F\}\\leq\\mu\_\{i\}/2self\-consistently, the quadratic term is sub\-dominant and the smaller root of the inequality satisfies
‖Δ‖F≤cβ\+𝒪\(ρic2β3\)=εG,i\+εH,iRiμi\+ρiRi22μi\+𝒪\(ρiμi3\(εG,i\+εH,iRi\)2\),\\begin\{split\}\\\|\\Delta\\\|\_\{F\}&\\;\\leq\\;\\frac\{c\}\{\\beta\}\+\\mathcal\{O\}\\\!\\left\(\\frac\{\\rho\_\{i\}\\,c^\{2\}\}\{\\beta^\{3\}\}\\right\)\\\\ &\\;=\\;\\frac\{\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}R\_\{i\}\}\{\\mu\_\{i\}\}\+\\frac\{\\rho\_\{i\}R\_\{i\}^\{2\}\}\{2\\mu\_\{i\}\}\\\\ &\\quad\+\\mathcal\{O\}\\\!\\left\(\\frac\{\\rho\_\{i\}\}\{\\mu\_\{i\}^\{3\}\}\\bigl\(\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}R\_\{i\}\\bigr\)^\{2\}\\right\),\\end\{split\}\(97\)which is exactly \([81](https://arxiv.org/html/2607.29071#A1.E81)\)\. Recall that𝐀^i=𝐏ri𝐖^𝐐ri=𝐔ri𝚺^i𝐕ri\\widehat\{\\mathbf\{A\}\}\_\{i\}=\\mathbf\{P\}\_\{r\_\{i\}\}\\widehat\{\\mathbf\{W\}\}\\mathbf\{Q\}\_\{r\_\{i\}\}=\\mathbf\{U\}\_\{r\_\{i\}\}\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\mathbf\{V\}\_\{r\_\{i\}\}by \([69](https://arxiv.org/html/2607.29071#A1.E69)\), and that the bijection𝚺↦𝐔ri𝚺𝐕ri\\boldsymbol\{\\Sigma\}\\mapsto\\mathbf\{U\}\_\{r\_\{i\}\}\\boldsymbol\{\\Sigma\}\\mathbf\{V\}\_\{r\_\{i\}\}fromℝri×ri\\mathbb\{R\}^\{r\_\{i\}\\times r\_\{i\}\}toℳri\\mathcal\{M\}\_\{r\_\{i\}\}identifies𝚺i∗\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}with the matrix\-space minimizer𝐀i∗\\mathbf\{A\}\_\{i\}^\{\*\}\. Consequently
εimis=𝐔ri\(𝚺i∗−𝚺^i\)𝐕ri=𝐀i∗−𝐀^i=Δ,\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}\\;=\\;\\mathbf\{U\}\_\{r\_\{i\}\}\(\\boldsymbol\{\\Sigma\}\_\{i\}^\{\*\}\-\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\)\\mathbf\{V\}\_\{r\_\{i\}\}\\;=\\;\\mathbf\{A\}\_\{i\}^\{\*\}\-\\widehat\{\\mathbf\{A\}\}\_\{i\}\\;=\\;\\Delta,\(98\)so‖εimis‖F=‖Δ‖F\\\|\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}\\\|\_\{F\}=\\\|\\Delta\\\|\_\{F\}in weight space, and \([97](https://arxiv.org/html/2607.29071#A1.E97)\) is exactly \([82](https://arxiv.org/html/2607.29071#A1.E82)\)\. SettingεG,i=εH,i=ρi=0\\varepsilon\_\{G,i\}=\\varepsilon\_\{H,i\}=\\rho\_\{i\}=0gives‖Δ‖F=0\\\|\\Delta\\\|\_\{F\}=0, recovering Proposition[5](https://arxiv.org/html/2607.29071#Thmproposition5)\. ∎
\([81](https://arxiv.org/html/2607.29071#A1.E81)\) resolves both questions raised above\.*\(i\) Graceful degradation\.*The misalignment grows linearly with the gradient leakageεG,i\\varepsilon\_\{G,i\}and with the curvature couplingεH,i\\varepsilon\_\{H,i\}, so there is no longer a discontinuity between the idealized regime \(εG,i=εH,i=0\\varepsilon\_\{G,i\}=\\varepsilon\_\{H,i\}=0, exact alignment\) and the typical regime \(small but non\-zero violations\)\.*\(ii\) Compression\-modulated curvature term\.*The curvature coupling enters only through the productεH,iRi\\varepsilon\_\{H,i\}\\,R\_\{i\}: when compression discards little energy \(Ri=δri\(𝐖^\)R\_\{i\}=\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\)small\), even strong off\-block coupling has limited effect, because the subspace\(Id−Πri\)\[𝐖^\]\(\\mathrm\{Id\}\-\\Pi\_\{r\_\{i\}\}\)\[\\widehat\{\\mathbf\{W\}\}\]on which the commutator acts is itself small\.*\(iii\) Quadratic\-truncation control\.*TheρiRi2/\(2μi\)\\rho\_\{i\}R\_\{i\}^\{2\}/\(2\\mu\_\{i\}\)term is exactly the contribution of the cubic Taylor remainder of the true \(non\-quadratic\) loss; settingρi=0\\rho\_\{i\}=0recovers the purely quadratic case\.*\(iv\) Estimable from a single probe\.*εG,i\\varepsilon\_\{G,i\}requires one gradient evaluation at𝐖^\\widehat\{\\mathbf\{W\}\}and one projection\.εH,i\\varepsilon\_\{H,i\}can be estimated by Hutchinson\-style probes ofℋi\\mathcal\{H\}\_\{i\}alongΠri\\Pi\_\{r\_\{i\}\}\- and\(Id−Πri\)\(\\mathrm\{Id\}\-\\Pi\_\{r\_\{i\}\}\)\-aligned directions\.μi\\mu\_\{i\}is lower\-bounded by the smallest eigenvalue ofΠri∘ℋi∘Πri\\Pi\_\{r\_\{i\}\}\\circ\\mathcal\{H\}\_\{i\}\\circ\\Pi\_\{r\_\{i\}\}restricted toℳri\\mathcal\{M\}\_\{r\_\{i\}\}; in overparameterized regimes near interpolating minima it is justified by\[[32](https://arxiv.org/html/2607.29071#bib.bib83)\]\.*\(v\) Pre\-training proximity\.*When𝐖^\\widehat\{\\mathbf\{W\}\}is the pre\-trained weight and federated fine\-tuning is a small perturbation of pre\-training, the condition4ρi\(εG,i\+εH,iRi\)<μi24\\rho\_\{i\}\(\\varepsilon\_\{G,i\}\+\\varepsilon\_\{H,i\}R\_\{i\}\)<\\mu\_\{i\}^\{2\}is mild becauseεG,i\\varepsilon\_\{G,i\}is bounded by the magnitude of the fine\-tuning gradient at initialization, which is itself small in this regime\.
Proposition[6](https://arxiv.org/html/2607.29071#Thmproposition6)therefore upgrades Proposition[5](https://arxiv.org/html/2607.29071#Thmproposition5)from a idealized guarantee to a quantitative diagnostic: the practitioner reads offεG,i,εH,i,ρi\\varepsilon\_\{G,i\},\\varepsilon\_\{H,i\},\\rho\_\{i\}from a single curvature/gradient probe at the reference and obtains an explicit a\-priori bound onεimis\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}, which Theorem[2](https://arxiv.org/html/2607.29071#Thmtheorem2)then injects into the global aggregation error\.
Propositions[5](https://arxiv.org/html/2607.29071#Thmproposition5)–[6](https://arxiv.org/html/2607.29071#Thmproposition6)boundεimis\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}from the*curvature side*: they read off a single gradient/Hessian probe at the reference𝐖^\\widehat\{\\mathbf\{W\}\}and trade restricted strong convexity for a closed\-form bound\. We now present a complementary,*data\-side*bound that arrives atεimis\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}through a different route, namely Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\(b\)’s PŁ/QG inequality applied to the loss gapΔi\(𝐖^\)\\Delta\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\), and exposes a structurally different decomposition\. The two are not redundant: Proposition[6](https://arxiv.org/html/2607.29071#Thmproposition6)requires Hessian information about𝐖^\\widehat\{\\mathbf\{W\}\}but no knowledge of group\-iioptima, while the bound below requiresLL\-smoothness and information about the group optima𝐖iloc\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}but no curvature alignment; we discuss when each is the operative tool in Remark[2](https://arxiv.org/html/2607.29071#Thmremark2)below\.
The misalignment residual along this PŁ route is driven by a single quantityΔi\(𝐖^\)\\Delta\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\), butΔi\\Delta\_\{i\}itself bundles two conceptually distinct effects: \(i\) the gap between the global reference𝐖^\\widehat\{\\mathbf\{W\}\}and groupii’s own task\-optimal weight, which is the genuine*statistical heterogeneity*, and \(ii\) how lossy it is to compress groupii’s task\-optimal weight to rankrir\_\{i\}, which is a per\-group*compression*effect\. We now disentangle them\. For each groupii, let
𝐖iloc∈argmin𝐖∈ℝm×nℒi\(𝐖\)\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\\;\\in\\;\\arg\\min\_\{\\mathbf\{W\}\\in\\mathbb\{R\}^\{m\\times n\}\}\\mathcal\{L\}\_\{i\}\(\\mathbf\{W\}\)\(99\)denote an unconstrained ambient minimizer of groupii’s population loss, and define the per\-group*heterogeneity*relative to the reference𝐖^\\widehat\{\\mathbf\{W\}\}as
Hi\(𝐖^\)≜‖𝐖^−𝐖iloc‖F\.H\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\\;\\triangleq\\;\\bigl\\\|\\widehat\{\\mathbf\{W\}\}\-\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\\bigr\\\|\_\{F\}\.\(100\)The quantityHi\(𝐖^\)=0H\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)=0for everyiiexactly when all groups share a common ambient task\-optimal weight; in standard FL terminology,\{Hi\}\\\{H\_\{i\}\\\}measures distribution shift across groups in the weight space, independently of the rank constraint\.
###### Proposition 7\(Heterogeneity decomposition of the misalignment\)\.
Suppose each ambient task loss𝐖↦ℒi\(𝐖\)\\mathbf\{W\}\\mapsto\\mathcal\{L\}\_\{i\}\(\\mathbf\{W\}\)isLL\-smooth \(Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)\(a\)\) and Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\(a\) holds\. Then for every groupii,
Δi\(𝐖^\)≤L2\(Hi\(𝐖^\)\+δri\(𝐖iloc\)\)2,\\Delta\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\\;\\leq\\;\\tfrac\{L\}\{2\}\\bigl\(H\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\+\\delta\_\{r\_\{i\}\}\(\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\)\\bigr\)^\{2\},\(101\)and consequently the misalignment residual obeys the data\-heterogeneity\-vs\-compression decomposition
𝔼‖εimis‖F≤λ1κ\(𝐒\)Lμ\(Hi\(𝐖^\)⏟heterogeneity\+δri\(𝐖iloc\)⏟rank\-ricompression of𝐖iloc\)\.\\begin\{split\}\\mathbb\{E\}\\\|\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}\\\|\_\{F\}\\;\\leq\\;&\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\\sqrt\{\\tfrac\{L\}\{\\mu\}\}\\,\\Bigl\(\\underbrace\{H\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\}\_\{\\text\{heterogeneity\}\}\\\\ &\\;\+\\;\\underbrace\{\\delta\_\{r\_\{i\}\}\(\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\)\}\_\{\\text\{rank\-\}r\_\{i\}\\text\{ compression of \}\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\}\\Bigr\)\.\\end\{split\}\(102\)
###### Proof of Proposition[7](https://arxiv.org/html/2607.29071#Thmproposition7)\.
Let𝐖~i∗≜𝐏ri𝐖^𝐐ri=𝐔ri𝚺^i𝐕ri\\widetilde\{\\mathbf\{W\}\}\_\{i\}^\{\*\}\\triangleq\\mathbf\{P\}\_\{r\_\{i\}\}\\widehat\{\\mathbf\{W\}\}\\mathbf\{Q\}\_\{r\_\{i\}\}=\\mathbf\{U\}\_\{r\_\{i\}\}\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\\mathbf\{V\}\_\{r\_\{i\}\}, so that the geometric projection𝚺^i\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}is the adapter realization of𝐖~i∗\\widetilde\{\\mathbf\{W\}\}\_\{i\}^\{\*\}\. By definition ofΔi\\Delta\_\{i\}and the relationfi\(𝚺\)=ℒi\(𝐔ri𝚺𝐕ri\)f\_\{i\}\(\\boldsymbol\{\\Sigma\}\)=\\mathcal\{L\}\_\{i\}\(\\mathbf\{U\}\_\{r\_\{i\}\}\\boldsymbol\{\\Sigma\}\\mathbf\{V\}\_\{r\_\{i\}\}\),
Δi\(𝐖^\)=fi\(𝚺^i\)−fi∗≤ℒi\(𝐖~i∗\)−ℒi\(𝐖iloc\),\\begin\{split\}\\Delta\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)&=f\_\{i\}\(\\widehat\{\\boldsymbol\{\\Sigma\}\}\_\{i\}\)\-f\_\{i\}^\{\*\}\\\\ &\\leq\\mathcal\{L\}\_\{i\}\(\\widetilde\{\\mathbf\{W\}\}\_\{i\}^\{\*\}\)\-\\mathcal\{L\}\_\{i\}\(\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\),\\end\{split\}\(103\)where the inequality usesfi∗=min𝚺ℒi\(𝐔ri𝚺𝐕ri\)≥min𝐖ℒi\(𝐖\)=ℒi\(𝐖iloc\)f\_\{i\}^\{\*\}=\\min\_\{\\boldsymbol\{\\Sigma\}\}\\mathcal\{L\}\_\{i\}\(\\mathbf\{U\}\_\{r\_\{i\}\}\\boldsymbol\{\\Sigma\}\\mathbf\{V\}\_\{r\_\{i\}\}\)\\geq\\min\_\{\\mathbf\{W\}\}\\mathcal\{L\}\_\{i\}\(\\mathbf\{W\}\)=\\mathcal\{L\}\_\{i\}\(\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\)\. ByLL\-smoothness ofℒi\\mathcal\{L\}\_\{i\}\(Assumption[1](https://arxiv.org/html/2607.29071#Thmassumption1)\(a\)\) and∇ℒi\(𝐖iloc\)=𝟎\\nabla\\mathcal\{L\}\_\{i\}\(\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\)=\\mathbf\{0\},
ℒi\(𝐖~i∗\)−ℒi\(𝐖iloc\)≤L2‖𝐖~i∗−𝐖iloc‖F2\.\\mathcal\{L\}\_\{i\}\(\\widetilde\{\\mathbf\{W\}\}\_\{i\}^\{\*\}\)\-\\mathcal\{L\}\_\{i\}\(\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\)\\;\\leq\\;\\tfrac\{L\}\{2\}\\bigl\\\|\\widetilde\{\\mathbf\{W\}\}\_\{i\}^\{\*\}\-\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\\bigr\\\|\_\{F\}^\{\\,2\}\.\(104\)Apply the triangle inequality to split the right\-hand side along the geometric projection of𝐖iloc\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}:
‖𝐖~i∗−𝐖iloc‖F≤‖𝐖~i∗−𝐏ri𝐖iloc𝐐ri‖F\+‖𝐏ri𝐖iloc𝐐ri−𝐖iloc‖F=‖𝐏ri\(𝐖^−𝐖iloc\)𝐐ri‖F\+δri\(𝐖iloc\)≤Hi\(𝐖^\)\+δri\(𝐖iloc\),\\begin\{split\}\\bigl\\\|\\widetilde\{\\mathbf\{W\}\}\_\{i\}^\{\*\}\-\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\\bigr\\\|\_\{F\}&\\leq\\bigl\\\|\\widetilde\{\\mathbf\{W\}\}\_\{i\}^\{\*\}\-\\mathbf\{P\}\_\{r\_\{i\}\}\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\\mathbf\{Q\}\_\{r\_\{i\}\}\\bigr\\\|\_\{F\}\\\\ &\\quad\+\\bigl\\\|\\mathbf\{P\}\_\{r\_\{i\}\}\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\\mathbf\{Q\}\_\{r\_\{i\}\}\-\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\\bigr\\\|\_\{F\}\\\\ &=\\bigl\\\|\\mathbf\{P\}\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\-\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\)\\mathbf\{Q\}\_\{r\_\{i\}\}\\bigr\\\|\_\{F\}\+\\delta\_\{r\_\{i\}\}\(\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\)\\\\ &\\leq H\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\+\\delta\_\{r\_\{i\}\}\(\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\),\\end\{split\}\(105\)where the last step uses‖𝐏ri𝐀𝐐ri‖F≤‖𝐀‖F\\\|\\mathbf\{P\}\_\{r\_\{i\}\}\\mathbf\{A\}\\mathbf\{Q\}\_\{r\_\{i\}\}\\\|\_\{F\}\\leq\\\|\\mathbf\{A\}\\\|\_\{F\}\(orthogonal projectors are non\-expansive in Frobenius norm\)\. Substituting \([105](https://arxiv.org/html/2607.29071#A1.E105)\) into \([104](https://arxiv.org/html/2607.29071#A1.E104)\) and chaining with \([103](https://arxiv.org/html/2607.29071#A1.E103)\) yields \([101](https://arxiv.org/html/2607.29071#A1.E101)\)\. The bound \([102](https://arxiv.org/html/2607.29071#A1.E102)\) now follows from \([29](https://arxiv.org/html/2607.29071#S4.E29)\) of Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\(b\):2Δi/μ≤L/μ\(Hi\(𝐖^\)\+δri\(𝐖iloc\)\)\\sqrt\{2\\Delta\_\{i\}/\\mu\}\\leq\\sqrt\{L/\\mu\}\\,\\bigl\(H\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)\+\\delta\_\{r\_\{i\}\}\(\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\)\\bigr\), multiplied byλ1κ\(𝐒\)\\lambda\_\{1\}\\,\\kappa\(\\mathbf\{S\}\)\. ∎
Proposition[7](https://arxiv.org/html/2607.29071#Thmproposition7)cleanly separates the two effects that the misalignment residual confounds\. The first term,Hi\(𝐖^\)H\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\), measures statistical heterogeneity: it is the ambient distance from groupii’s task optimum to the reference𝐖^\\widehat\{\\mathbf\{W\}\}and is*independent ofrir\_\{i\}*, so it cannot be reduced by raising the rank\. The second term,δri\(𝐖iloc\)\\delta\_\{r\_\{i\}\}\(\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}\), is purely a low\-rank compression cost evaluated at groupii’s own optimum and*vanishes asrir\_\{i\}grows*\. Two practical readings follow\. \(i\) Choosing𝐖^\\widehat\{\\mathbf\{W\}\}to minimize∑i\(ni/N\)Hi\(𝐖^\)\\sum\_\{i\}\(n\_\{i\}/N\)\\,H\_\{i\}\(\\widehat\{\\mathbf\{W\}\}\)recovers the data\-weighted ambient barycenter and gives the heterogeneity term its smallest value; this is the centralized minimizer when theℒi\\mathcal\{L\}\_\{i\}’s are quadratic\. \(ii\) When𝐖^\\widehat\{\\mathbf\{W\}\}is taken as the pre\-trained weight𝐖\\mathbf\{W\}and federated fine\-tuning is a small perturbation of pre\-training,Hi\(𝐖\)H\_\{i\}\(\\mathbf\{W\}\)is small because each𝐖iloc\\mathbf\{W\}\_\{i\}^\{\\mathrm\{loc\}\}is close to𝐖\\mathbf\{W\}, so the bound is dominated by the rank\-rir\_\{i\}compression terms, which the server can pre\-compute from per\-group spectra\.
Propositions[6](https://arxiv.org/html/2607.29071#Thmproposition6)and[7](https://arxiv.org/html/2607.29071#Thmproposition7)offer two complementary diagnostics for quantity‖εimis‖F\\\|\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}\\\|\_\{F\}, each requiring different inputs and decomposing the error along different axes\. The following remark details when each is the operative tool\.
The reference weight𝐖^\\widehat\{\\mathbf\{W\}\}is a free design parameter; setting𝐖^=𝐖\\widehat\{\\mathbf\{W\}\}=\\mathbf\{W\}, the pre\-trained weight, yields the closed\-form Eckart–Young expressionδri\(𝐖\)=\(∑k\>riλk2\)1/2\\delta\_\{r\_\{i\}\}\(\\mathbf\{W\}\)=\(\\sum\_\{k\>r\_\{i\}\}\\lambda\_\{k\}^\{2\}\)^\{1/2\}and makesΔi\(𝐖\)\\Delta\_\{i\}\(\\mathbf\{W\}\)the loss reduction achieved by fine\-tuning relative to the pre\-trained initialization, which is small whenever federated fine\-tuning is a perturbation of pre\-training\. Setting𝐖^\\widehat\{\\mathbf\{W\}\}to the centralized minimizerargmin𝐖∑i\(ni/N\)ℒi\(𝐖\)\\arg\\min\_\{\\mathbf\{W\}\}\\sum\_\{i\}\(n\_\{i\}/N\)\\mathcal\{L\}\_\{i\}\(\\mathbf\{W\}\)instead turnsΔ¯\(𝐖^\)\\overline\{\\Delta\}\(\\widehat\{\\mathbf\{W\}\}\)into a measure of the irreducible heterogeneity gap between federated and centralized training\.
The weak\-to\-strong analysis of Section[IV\-C](https://arxiv.org/html/2607.29071#S4.SS3)inherits the bound transparently: it relies on Theorem[2](https://arxiv.org/html/2607.29071#Thmtheorem2)only through‖𝐖g−𝐖^‖F\\\|\\mathbf\{W\}\_\{\\mathrm\{g\}\}\-\\widehat\{\\mathbf\{W\}\}\\\|\_\{F\}, which the Lipschitz forward pass \(Assumption[2](https://arxiv.org/html/2607.29071#Thmassumption2)\) transports into the artifact bound of Lemma[1](https://arxiv.org/html/2607.29071#Thmlemma1)\. Specializing to the common\-projection\-target condition simply zeroes out the misalignment residual in the artifact budget, leaving the bias–variance trade\-off of Proposition[4](https://arxiv.org/html/2607.29071#Thmproposition4)structurally unchanged\.
### A\-FProof of Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3)
###### Proof of Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3)\.
Under the common\-projection\-target condition \(Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\(b\)\),εimis=𝟎\\varepsilon\_\{i\}^\{\\mathrm\{mis\}\}=\\mathbf\{0\}for everyii\. Neglecting the vanishing optimization residual \(which by Proposition[2](https://arxiv.org/html/2607.29071#Thmproposition2)\(a\) and Theorem[1](https://arxiv.org/html/2607.29071#Thmtheorem1)satisfies𝔼‖εiopt‖F=𝒪\(T1−1/4\)\\mathbb\{E\}\\\|\\varepsilon\_\{i\}^\{\\mathrm\{opt\}\}\\\|\_\{F\}=\\mathcal\{O\}\(T\_\{1\}^\{\-1/4\}\)in the IID regime\), Theorem[2](https://arxiv.org/html/2607.29071#Thmtheorem2)reduces to
‖𝐖g−𝐖^‖F=‖∑i=1KniN\(𝐖i−𝐖^\)‖F≤∑i=1KniN‖εisub‖F=∑i=1KniNδri\(𝐖^\)\.\\begin\{split\}\\bigl\\\|\\mathbf\{W\}\_\{\\mathrm\{g\}\}\-\\widehat\{\\mathbf\{W\}\}\\bigr\\\|\_\{F\}&=\\Bigl\\\|\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\(\\mathbf\{W\}\_\{i\}\-\\widehat\{\\mathbf\{W\}\}\)\\Bigr\\\|\_\{F\}\\\\ &\\leq\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\\,\\\|\\varepsilon\_\{i\}^\{\\mathrm\{sub\}\}\\\|\_\{F\}=\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\\,\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\)\.\\end\{split\}\(106\)Sinceδr\(𝐖^\)\\delta\_\{r\}\(\\widehat\{\\mathbf\{W\}\}\)is non\-increasing inrr,δri\(𝐖^\)≤δrmin\(𝐖^\)\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\)\\leq\\delta\_\{r\_\{\\min\}\}\(\\widehat\{\\mathbf\{W\}\}\)for allii, and the weighted average is bounded by the maximum, giving the second inequality \([30](https://arxiv.org/html/2607.29071#S4.E30)\)\. When𝐖^\\widehat\{\\mathbf\{W\}\}coincides with the pre\-trained weight𝐖\\mathbf\{W\}, the Eckart–Young–Mirsky theorem givesδr\(𝐖\)2=∑k\>rλk2\\delta\_\{r\}\(\\mathbf\{W\}\)^\{2\}=\\sum\_\{k\>r\}\\lambda\_\{k\}^\{2\}, which yields the spectrum form \([31](https://arxiv.org/html/2607.29071#S4.E31)\)\. ∎
### A\-GProofs for Weak\-to\-Strong Theory
###### Proof of Lemma[1](https://arxiv.org/html/2607.29071#Thmlemma1)\.
By Assumption[2](https://arxiv.org/html/2607.29071#Thmassumption2)applied to the pair\(𝐖g,𝐖^\)\(\\mathbf\{W\}\_\{\\mathrm\{g\}\},\\,\\widehat\{\\mathbf\{W\}\}\),
‖ξ\(x\)‖1=‖fweak\(x;𝐖g\)−f∗\(x\)‖1≤Lf‖𝐖g−𝐖^‖F\.\\begin\{split\}\\\|\\xi\(x\)\\\|\_\{1\}&=\\\|f\_\{\\mathrm\{weak\}\}\(x;\\,\\mathbf\{W\}\_\{\\mathrm\{g\}\}\)\-f^\{\*\}\(x\)\\\|\_\{1\}\\\\ &\\leq L\_\{f\}\\,\\\|\\mathbf\{W\}\_\{\\mathrm\{g\}\}\-\\widehat\{\\mathbf\{W\}\}\\\|\_\{F\}\.\\end\{split\}\(107\)Note that Assumption[2](https://arxiv.org/html/2607.29071#Thmassumption2)is stated directly in terms of the full weight matrix𝐖\\mathbf\{W\}, so the bound above holds regardless of the factorization variant; no additionalκ\(𝐒\)\\kappa\(\\mathbf\{S\}\)factor is needed\. Substituting the aggregation\-error from Theorem[3](https://arxiv.org/html/2607.29071#Thmtheorem3)gives
‖ξ\(x\)‖1≤Lf∑i=1KniNδri\(𝐖^\)≤Lfδrmin\(𝐖^\)=Bξ,\\hskip\-10\.00002pt\\\|\\xi\(x\)\\\|\_\{1\}\\leq L\_\{f\}\\sum\_\{i=1\}^\{K\}\\frac\{n\_\{i\}\}\{N\}\\,\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\)\\leq L\_\{f\}\\,\\delta\_\{r\_\{\\min\}\}\(\\widehat\{\\mathbf\{W\}\}\)=B\_\{\\xi\},\(108\)where the second inequality usesδri\(𝐖^\)≤δrmin\(𝐖^\)\\delta\_\{r\_\{i\}\}\(\\widehat\{\\mathbf\{W\}\}\)\\leq\\delta\_\{r\_\{\\min\}\}\(\\widehat\{\\mathbf\{W\}\}\)and the fact that\{ni/N\}\\\{n\_\{i\}/N\\\}form a probability distribution\. ∎
#### A\-G1Gradient Decomposition
Let𝐳s\(x\)∈ℝ\|𝒴\|\\mathbf\{z\}\_\{s\}\(x\)\\in\\mathbb\{R\}^\{\|\\mathcal\{Y\}\|\}denote the pre\-softmax logits of the strong model, so thatfstrong\(x;𝐖s\)=softmax\(𝐳s\(x\)\)f\_\{\\mathrm\{strong\}\}\(x;\\,\\mathbf\{W\}\_\{\\mathrm\{s\}\}\)=\\mathrm\{softmax\}\(\\mathbf\{z\}\_\{s\}\(x\)\)\. Using the standard identity∇𝐳ℓCE\(p,softmax\(𝐳\)\)=softmax\(𝐳\)−p\\nabla\_\{\\mathbf\{z\}\}\\,\\ell\_\{\\mathrm\{CE\}\}\(p,\\,\\mathrm\{softmax\}\(\\mathbf\{z\}\)\)=\\mathrm\{softmax\}\(\\mathbf\{z\}\)\-p, the gradient ofℒconf\\mathcal\{L\}\_\{\\mathrm\{conf\}\}with respect to𝐳s\\mathbf\{z\}\_\{s\}is
∇𝐳sℒconf=\(1−α\)\(fs−fw\)\+α\(fs−f^s\)\.\\begin\{split\}\\nabla\_\{\\mathbf\{z\}\_\{s\}\}\\mathcal\{L\}\_\{\\mathrm\{conf\}\}&=\(1\\\!\-\\\!\\alpha\)\(f\_\{\\mathrm\{s\}\}\-f\_\{\\mathrm\{w\}\}\)\\\\ &\\quad\+\\alpha\(f\_\{\\mathrm\{s\}\}\-\\hat\{f\}\_\{\\mathrm\{s\}\}\)\.\\end\{split\}\(109\)Substitutingfw=f∗\+ξ\(x\)f\_\{\\mathrm\{w\}\}=f^\{\*\}\+\\xi\(x\)from Definition[4](https://arxiv.org/html/2607.29071#Thmdefinition4)and rearranging yields Eq\. \([35](https://arxiv.org/html/2607.29071#S4.E35)\)\.
#### A\-G2Proof of Proposition[3](https://arxiv.org/html/2607.29071#Thmproposition3)
###### Proof\.
The bound follows from Lemma[1](https://arxiv.org/html/2607.29071#Thmlemma1):
‖\(1−α\)ξ\(x\)‖2≤\(1−α\)‖ξ\(x\)‖2≤\(1−α\)‖ξ\(x\)‖1≤\(1−α\)Bξ,\\begin\{split\}\\bigl\\\|\(1\\\!\-\\\!\\alpha\)\\,\\xi\(x\)\\bigr\\\|\_\{2\}&\\leq\(1\\\!\-\\\!\\alpha\)\\,\\\|\\xi\(x\)\\\|\_\{2\}\\\\ &\\leq\(1\\\!\-\\\!\\alpha\)\\,\\\|\\xi\(x\)\\\|\_\{1\}\\leq\(1\\\!\-\\\!\\alpha\)\\,B\_\{\\xi\},\\end\{split\}\(110\)where the second inequality uses∥⋅∥2≤∥⋅∥1\\\|\\cdot\\\|\_\{2\}\\leq\\\|\\cdot\\\|\_\{1\}\. ExpandingBξ=Lfδrmin\(𝐖^\)B\_\{\\xi\}=L\_\{f\}\\,\\delta\_\{r\_\{\\min\}\}\(\\widehat\{\\mathbf\{W\}\}\)yields \([36](https://arxiv.org/html/2607.29071#S4.E36)\)\. ∎
#### A\-G3Proof of Proposition[4](https://arxiv.org/html/2607.29071#Thmproposition4)
###### Proof\.
The logit\-gradient error decomposes as\(1−α\)ξ\(x\)−α\(fstrong−f^strong\)\(1\-\\alpha\)\\xi\(x\)\-\\alpha\(f\_\{\\mathrm\{strong\}\}\-\\hat\{f\}\_\{\\mathrm\{strong\}\}\)\. By the assumed uncorrelatedness ofξ\(x\)\\xi\(x\)andfstrong−f^strongf\_\{\\mathrm\{strong\}\}\-\\hat\{f\}\_\{\\mathrm\{strong\}\}, the expected squared norm separates: the excess riskℛ\(α\)\\mathcal\{R\}\(\\alpha\)is a quadratic inα\\alpha:
ℛ\(α\)=\(1−α\)2A\+α2B\\mathcal\{R\}\(\\alpha\)=\(1\-\\alpha\)^\{2\}A\+\\alpha^\{2\}B\(111\)withA=𝔼‖ξ\(x\)‖22A=\\mathbb\{E\}\\\|\\xi\(x\)\\\|\_\{2\}^\{2\}andB=VselfB=V\_\{\\mathrm\{self\}\}\. Settingdℛdα=−2\(1−α\)A\+2αB=0\\tfrac\{\\mathrm\{d\}\\mathcal\{R\}\}\{\\mathrm\{d\}\\alpha\}=\-2\(1\-\\alpha\)A\+2\\alpha B=0givesα∗=AA\+B\\alpha^\{\*\}=\\dfrac\{A\}\{A\+B\}, which is the first equality in \([38](https://arxiv.org/html/2607.29071#S4.E38)\)\. The upper boundA≤Bξ2A\\leq B\_\{\\xi\}^\{\\,2\}follows from Lemma[1](https://arxiv.org/html/2607.29071#Thmlemma1), and the monotone dependence ofα∗\\alpha^\{\*\}onBξB\_\{\\xi\}follows from the fact thatx↦x/\(Vself\+x\)x\\mapsto x/\(V\_\{\\mathrm\{self\}\}\+x\)is monotone non\-decreasing on\[0,∞\)\[0,\\infty\)\. ∎
## Appendix BExperimental Details
All experiments are conducted on a single server equipped with 8×\\timesNVIDIA L20 GPUs \(48 GB\)\. All methods use AdamW with weight decay10−210^\{\-2\}, learning rate5×10−55\\times 10^\{\-5\}, a linear warmup overmin\(50,⌊Tsteps/5⌋\)\\min\(50,\\lfloor T\_\{\\mathrm\{steps\}\}/5\\rfloor\)steps, and gradient clipping atℓ2\\ell\_\{2\}\-norm 1\. Each client trains for 1 local epoch per round with batch size 2, gradient accumulation over 4 steps \(effective batch size 8\), and maximum sequence length 256\.
For 13B experiments we reduce the learning rate to2×10−52\\times 10^\{\-5\}, lower the per\-device batch size to 1, and increase gradient accumulation to 8\. We enable gradient checkpointing \(use\_reentrant=False\) and use 8\-bit AdamW to reduce peak GPU memory\. Computations are in bfloat16 where supported, falling back to float16 otherwise\.
LoRA adapters use rankr=16r=16,α=32\\alpha=32, and dropout 0\.05\. Fed\-RAC\-LoRA uses a lower learning rate1×10−51\\times 10^\{\-5\}for numerical stability in bfloat16\. FedMKT uses rankr=8r=8,α=16\\alpha=16, client lr5×10−55\\times 10^\{\-5\}, server lr3×10−53\\times 10^\{\-5\}, KD mixing weightλ=0\.9\\lambda=0\.9, and 20% of each client’s data as a public pool\. FedBiOT’s bottleneck dimension is 64 withWdownW\_\{\\mathrm\{down\}\}zero\-initialized\.
The weak teacher \(Stage 1 output\) is frozen throughout Stage 2\. For each sample the teacher’s per\-choice log\-likelihoods are softmax\-normalized to produce soft targets; the strong model minimizes cross\-entropy against these labels, optionally augmented by entropy regularizationα⋅H\(p^strong\)\\alpha\\cdot H\(\\hat\{p\}\_\{\\mathrm\{strong\}\}\)withα=0\.5\\alpha=0\.5\. Stage 2 runs forT2=5T\_\{2\}=5epochs at learning rate5×10−55\\times 10^\{\-5\}with the same LoRA configuration as Stage 1\.相似文章
联邦轻量级微调
本文介绍了FLITE(联邦低秩迭代训练引擎),一种联邦微调方法,通过使用冻结的仿射映射网络,从一个小型可训练潜变量和低秩可种子重生的因子分解生成权重,将每轮每客户端的通信量降至每轮1280个浮点数(约5KB)——相比于全权重FedAvg减少了8718倍。在CIFAR-100数据集上使用ResNet-18进行测试,准确率与全权重FedAvg相差在0.5个百分点以内。
FedSubMuon: 通过结构化子空间Muon的通信高效联邦LLM微调
FedSubMuon是一种通信高效的联邦LLM微调方法,它在结构化子空间内优化紧凑系数矩阵,以减少上传成本同时保持强劲性能。
面向联邦多模态图基础模型:一种拓扑感知的多模态对齐框架
提出了FedGAMMA,一种联邦多模态图基础学习框架,通过两阶段预训练和基于提示的微调,对齐多模态属性和图拓扑,在多个数据集上取得了显著提升。
HASA:面向计算受限的异构模型联邦学习的子网分配
本文提出了HASA,一种面向异构模型联邦学习的异构感知子网分配方法,该方法在固定计算预算下根据客户端异构性分数分配子网宽度,从而提升平均准确率和最差客户端准确率。
DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink Accounting
Presents DG-FedReuse, a federated learning mechanism that reuses age-decayed cached client updates under a proxy-gradient threshold to reduce uplink communication, achieving significant modeled savings with minimal accuracy loss on image classification benchmarks.