Dual Attention Heads for Personalized Federated Learning in ECG Classification
Summary
This paper proposes FedDualAtt, a personalized federated learning approach for ECG classification that splits transformer attention heads into globally aggregated and locally private branches to handle data heterogeneity across clinical sites. Experiments on the FedCVD benchmark show improved performance over existing methods.
View Cached Full Text
Cached at: 07/09/26, 07:42 AM
# This paper was accepted to IEEE MWSCAS 2026 Dual Attention Heads for Personalized Federated Learning in ECG Classification
Source: [https://arxiv.org/html/2607.06653](https://arxiv.org/html/2607.06653)
###### Abstract
Federated learning \(FL\) enables collaborative model training across institutions without sharing sensitive patient data\. However, the inherent heterogeneity of electrocardiogram \(ECG\) data across healthcare providers presents significant technical challenges for robust classification\. We propose FedDualAtt, a personalized federated learning approach that splits transformer attention heads into global and local branches\. Global heads are aggregated via FedAvg to capture shared cross\-site patterns, while local heads remain client\-specific to adapt to institution\-level recording characteristics\. Experiments on FedCVD, an FL benchmark for cardiovascular disease detection, demonstrate that FedDualAtt outperforms existing FL and personalized FL methods in ECG classification tasks\. Analysis of global\-local head ratios reveals that different clients benefit from varying levels of architectural personalization\.
## IIntroduction
Cardiovascular diseases account for approximately 19\.8 million deaths annually, making them the leading cause of global mortality\[[9](https://arxiv.org/html/2607.06653#bib.bib1)\]\. Although the 12\-lead electrocardiogram \(ECG\) remains the primary non\-invasive tool for detecting cardiac pathologies, errors in manual ECG analysis can lead to misdiagnosis and delayed treatment\[[7](https://arxiv.org/html/2607.06653#bib.bib2)\]\. Deep learning models, particularly transformer\-based architectures, have shown remarkable performance in automated multi\-label ECG classification\[[6](https://arxiv.org/html/2607.06653#bib.bib6)\], yet their deployment requires large, diverse training sets drawn from multiple clinical sites to ensure generalizability\.
Federated learning \(FL\)\[[5](https://arxiv.org/html/2607.06653#bib.bib3)\]addresses these challenges by enabling collaborative model training across healthcare providers without centralizing sensitive patient records\. Nevertheless, ECG data exhibits significant heterogeneity across healthcare providers due to variations in recording equipment, patient demographics, and local disease prevalence\[[11](https://arxiv.org/html/2607.06653#bib.bib5)\]\. These non\-independently and identically distributed \(non\-IID\) data distributions cause performance degradation in standard FL approaches like FedAvg\[[12](https://arxiv.org/html/2607.06653#bib.bib7)\], which attempt to enforce a single global model across divergent client distributions\. While algorithm\-level personalized FL methods, such as Ditto\[[2](https://arxiv.org/html/2607.06653#bib.bib9)\]and FedALA\[[10](https://arxiv.org/html/2607.06653#bib.bib10)\], mitigate this through local fine\-tuning or adaptive aggregation, they introduce additional training objectives without addressing heterogeneity at the representation level\.
We observe that transformer self\-attention is particularly distribution\-sensitive: the patterns a head attends to naturally reflect the statistics of its training data\. This motivates an architectural personalization strategy that partitions attention heads into two functional groups\. Global heads are aggregated via FedAvg to capture universal temporal ECG patterns, while local heads remain client\-specific to adapt to site\-specific recording characteristics\. This partitioning introduces no additional training objectives or hyperparameters beyond the head\-split ratio\.
Our work introduces the FedDualAtt framework with the following contributions:
- •We introduce a dual\-attention transformer module appended to the FedCVD\[[11](https://arxiv.org/html/2607.06653#bib.bib5)\]ResNet1D\-34 backbone that partitions attention heads into a global branch \(FedAvg\-aggregated\) and a local branch \(per\-client\), enabling simultaneous cross\-site generalization and site\-specific adaptation within a single forward pass;
- •We design a federated training protocol with strict parameter separation: global and local parameters are stored, transmitted, and aggregated independently;
- •We conduct empirical analysis on the FedCVD benchmark over all nine head\-ratio configurations, characterizing the stability\-performance tradeoff in the global\-local attention split\.
The remaining sections are structured as follows: Section[II](https://arxiv.org/html/2607.06653#S2)reviews related work, Section[III](https://arxiv.org/html/2607.06653#S3)presents the proposed method, Section[IV](https://arxiv.org/html/2607.06653#S4)describes experimental evaluation, and Section[V](https://arxiv.org/html/2607.06653#S5)concludes\.
## IIBackground
### II\-AFederated learning for ECG
The FedCVD benchmark\[[11](https://arxiv.org/html/2607.06653#bib.bib5)\]establishes a standardized evaluation for FL methods on multi\-center ECG classification using four real\-world datasets with 20 diagnostic labels\. Evaluating seven FL algorithms, the best\-reported result was Scaffold\[[1](https://arxiv.org/html/2607.06653#bib.bib8)\]at 70\.1% Global Micro\-F1\. All evaluated methods operate on a ResNet1D\-34 backbone without a temporal attention component, leaving sequential ECG dependencies to convolutional layers alone\. Our work addresses this gap with FedDualAtt by augmenting the ResNet1D\-34 backbone with dual\-attention transformer blocks designed to disentangle global and local temporal representations\.
### II\-BPersonalized federated learning
Parameter\-level personalization includes Ditto \(dual\-objective local fine\-tuning with a proximal term\), FedALA \(local adaptive aggregation weights\), FedBN\[[4](https://arxiv.org/html/2607.06653#bib.bib11)\]\(per\-client batch normalization statistics\), and FedProx\[[3](https://arxiv.org/html/2607.06653#bib.bib13)\]\(proximal regularization to the global model\)\. These approaches treat personalization as a post\-aggregation adaptation step and can be applied to any architecture\. FedDualAtt instead enforces the global/local boundary at design time: the parameter partition is fixed at construction, and the FL protocol applies standard FedAvg to the global partition without any additional objectives or gradient manipulation\. Head specialization has precedent in natural language processing, where analyses show that individual transformer attention heads learn distinct syntactic and semantic functions\[[8](https://arxiv.org/html/2607.06653#bib.bib12)\]\. FedDualAtt builds on this by making the specialization architecturally explicit, assigning heads to global or local partitions at design time\.
## IIIProposed Method
Figure 1:FedDualAtt framework\. \(a\) DualAttentionResNet1D with parallel global \(θg\\theta^\{g\}, FedAvg\-aggregated\) and local \(ϕk\\phi\_\{k\}, per\-client\) attention branches\. \(b\) Federated training protocol with separate parameter stores forθg\\theta^\{g\}and\{ϕk\}k=1K\\\{\\phi\_\{k\}\\\}\_\{k=1\}^\{K\}\.We consider aK=4K=4healthcare providers, each with a private datasetDkD\_\{k\}of 12\-lead ECG recordingsxikx^\{k\}\_\{i\}and 20\-class multi\-label binary targetsyiky^\{k\}\_\{i\}\. Clients collaborate forT=50T=50communication rounds\.
### III\-AModel Architecture
Fig\.[1](https://arxiv.org/html/2607.06653#S3.F1)illustrates our proposed architecture\. We partition the model parameters into two sets: global parametersθg\\theta^\{g\}\(aggregated across clients\) and per\-client local parametersϕk\\phi\_\{k\}\(never aggregated\)\.
We integrate dual attention into a hybrid CNN\-Transformer architecture, with a ResNet1D\-34 that extracts features from 12\-lead ECG signals, followed by two stacked dual attention transformer blocks, and a classification head for multi\-label prediction\.
ResNet1D\-34 feature extractor:We adopt the ResNet1D\-34 convolutional feature extractor from FedCVD without modification, ensuring direct comparability with all baseline methods\. It maps each inputxik∈ℝ12×Lx\_\{i\}^\{k\}\\in\\mathbb\{R\}^\{12\\times L\}to a feature sequence𝐗∈ℝL′×d\\mathbf\{X\}\\in\\mathbb\{R\}^\{L^\{\\prime\}\\times d\}\(d=512d=512,L′=156L^\{\\prime\}=156\) via strided convolutions, pooling, and sinusoidal positional encoding\.
Dual attention transformer block:The core design insight is that full FedAvg destroys site\-specific attention patterns, while purely local training cannot exploit cross\-site data\. We resolve this by splitting theH=8H\\\!=\\\!8attention heads intoHgH\_\{g\}global heads andHlH\_\{l\}local heads \(Hg\+Hl=HH\_\{g\}\+H\_\{l\}=H\): the global heads are aggregated via FedAvg and learn patterns transferable across healthcare providers, while the local heads are kept per\-client and adapt to each site’s own distribution\. Each block processes𝐗\\mathbf\{X\}through these two parallel multi\-head attention \(MHA\) branches with layer normalization \(LN\) and fixed head dimensiondh=64d\_\{h\}=64\. Each branch uses an input projection𝐖in\\mathbf\{W\}\_\{in\}and output projection𝐖out\\mathbf\{W\}\_\{out\}to map between model dimensionddand the branch’s attention space\.
*Global branch*:
𝐆^=LN\(𝐗\+MHA\(𝐗𝐖ing,Hg\)𝐖outg\)∈ℝL′×d\\hat\{\\mathbf\{G\}\}=\\mathrm\{LN\}\\\!\\left\(\\mathbf\{X\}\+\\mathrm\{MHA\}\\\!\\left\(\\mathbf\{X\}\\mathbf\{W\}\_\{in\}^\{g\},\\,H\_\{g\}\\right\)\\mathbf\{W\}\_\{out\}^\{g\}\\right\)\\;\\in\\mathbb\{R\}^\{L^\{\\prime\}\\times d\}\(1\)
*Local branch*:
𝐋^=LN\(𝐗\+MHA\(𝐗𝐖inl,Hl\)𝐖outl\)∈ℝL′×d\\hat\{\\mathbf\{L\}\}=\\mathrm\{LN\}\\\!\\left\(\\mathbf\{X\}\+\\mathrm\{MHA\}\\\!\\left\(\\mathbf\{X\}\\mathbf\{W\}\_\{in\}^\{l\},\\,H\_\{l\}\\right\)\\mathbf\{W\}\_\{out\}^\{l\}\\right\)\\;\\in\\mathbb\{R\}^\{L^\{\\prime\}\\times d\}\(2\)
The branches are concatenated and projected to dimensiondd, then refined through a feed\-forward network \(FFN\):
𝐗′\\displaystyle\\mathbf\{X\}^\{\\prime\}=\[𝐆^;𝐋^\]𝐖c\\displaystyle=\[\\hat\{\\mathbf\{G\}\};\\hat\{\\mathbf\{L\}\}\]\\mathbf\{W\}\_\{c\}\(3\)𝐘\\displaystyle\\mathbf\{Y\}=LN\(𝐗′\+FFN\(𝐗′\)\)\\displaystyle=\\mathrm\{LN\}\\\!\\left\(\\mathbf\{X\}^\{\\prime\}\+\\mathrm\{FFN\}\(\\mathbf\{X\}^\{\\prime\}\)\\right\)\(4\)Intuitively, the global branch learns which ECG time steps are mutually relevant across all participating healthcare providers, and because its parameters are FedAvg\-aggregated, it is able to capture universal patterns to the entire client population\. The local branch performs the same operation with institution\-specific parametersϕk\\phi\_\{k\}by learning which temporal patterns are relevant for the patient population of the clientkkand the recording conditions, and is never shared with other clients\. The combine step fuses both representations into a single sequence, and the FFN applies a position\-wise nonlinear transformation to each time step independently\. The combined projection𝐖c\\mathbf\{W\}\_\{c\}, FFN, and all three LN layers are global parametersθg\\theta^\{g\}\. Keeping𝐖c\\mathbf\{W\}\_\{c\}global ensures that global and local representations are fused in a consistent coordinate space across clients as a per\-client𝐖c\\mathbf\{W\}\_\{c\}would allow each site to reinterpret the shared global features arbitrarily, thus undermining cross\-client alignment\.
Classification head:Global average pooling reduces the sequence dimension of𝐘\\mathbf\{Y\}, and a fully\-connected layer with sigmoid activation produces 20\-class multi\-label predictions trained with binary cross\-entropy loss\.
### III\-BFederated Training Protocol
The server maintains global parametersθg\\theta^\{g\}\(global attention heads, combine layer, FFN, and classification head\)\. Each clientkkindependently stores its own local parametersϕk\\phi\_\{k\}\(local attention heads\) without sharing them with the server or other clients\. Algorithm[1](https://arxiv.org/html/2607.06653#alg1)describes one communication round of the protocol\. Sinceϕk\\phi\_\{k\}is never aggregated, each client’s local attention heads adapt exclusively to its own data distribution while the sharedθg\\theta^\{g\}benefits from all clients via FedAvg\.
Algorithm 1FedDualAtt: Communication Roundtt0:Server:
θtg\\theta^\{g\}\_\{t\}; Client
kk:
ϕk,t\\phi\_\{k,t\},
DkD\_\{k\}, epochs
EE
0:Server: updated
θt\+1g\\theta^\{g\}\_\{t\+1\}; Client
kk: updated
ϕk,t\+1\\phi\_\{k,t\+1\}
1:foreach client
k=1,…,Kk=1,\\ldots,Kin paralleldo
2:Downlink: receive
θtg\\theta^\{g\}\_\{t\}from server
3:Initialize model: load
θtg\\theta^\{g\}\_\{t\}for global params,
ϕk,t\\phi\_\{k,t\}for local attention
4:Training: run SGD for
EEepochs on
DkD\_\{k\}, updating all params jointly
5:Uplink: send
\(θk,t\+1g,nk\)\(\\theta^\{g\}\_\{k,t\+1\},\\;n\_\{k\}\)to server
6:Retain
ϕk,t\+1\\phi\_\{k,t\+1\}in local storage
7:endfor
8:Aggregation:
θt\+1g←∑k=1Knknθk,t\+1g\\theta^\{g\}\_\{t\+1\}\\leftarrow\\sum\_\{k=1\}^\{K\}\\frac\{n\_\{k\}\}\{n\}\\,\\theta^\{g\}\_\{k,t\+1\},
n=∑k=1Knkn=\\sum\_\{k=1\}^\{K\}n\_\{k\}
9:return
θt\+1g\\theta^\{g\}\_\{t\+1\}
## IVExperimental Evaluation
### IV\-ASetup
Our proposed framework is evaluated using the FedCVD multi\-center ECG classification benchmark\. This benchmark includes four distinct clinical datasets: Shandong Provincial Hospital \(SPH\), Physikalisch\-Technische Bundesanstalt \(PTB\-XL\), Shaoxing People’s Hospital \(SXPH\), and the PhysioNet 2020 Challenge \(G12EC\)\. The four sites differ substantially in dataset size, recording hardware, and diagnostic label prevalence, creating significant non\-IID conditions\[[11](https://arxiv.org/html/2607.06653#bib.bib5)\]\. The task is 20\-class multi\-label ECG classification\. We report per\-client Micro\-F1 and mean average precision \(mAP\), which measure site\-level adaptation, and global Micro\-F1 and mAP, which aggregate predictions over all test samples across all four sites and reflect the ability to generalize to new healthcare providers\. We compare FedDualAtt against standard federated algorithms, including FedAvg, FedProx, and Scaffold, as well as personalized methods such as Ditto and FedALA\. For FedDualAtt, we evaluated across all nine head\-ratio configurations \(8Hg:0Hl8H\_\{g\}\\\!:\\\!0H\_\{l\}to0Hg:8Hl0H\_\{g\}\\\!:\\\!8H\_\{l\}\)\. All FedDualAtt results are mean±\\pmstd over 5 seeds\. Baseline FL results \(Micro\-F1, mAP\) are taken from the FedCVD paper\.
### IV\-BMain Results
Table[I](https://arxiv.org/html/2607.06653#S4.T1)reports per\-client and global Micro\-F1 and mAP for all methods\. Among all configurations, FedDualAtt with8Hg:0Hl8H\_\{g\}\\\!:\\\!0H\_\{l\}\(all global heads\) achieves the highest Global Micro\-F1 at72\.7%, surpassing the previous best \(Scaffold, 70\.1%\) by2\.62\.6percentage points \(pp\)\. This improvement comes entirely from the transformer attention architecture augmenting the ResNet1D\-34 backbone, not from personalization\. The local\-only extreme \(0Hg:8Hl0H\_\{g\}\\\!:\\\!8H\_\{l\}\) achieves competitive per\-client F1 but collapses global Micro\-F1 to 50\.8%, confirming that cross\-client aggregation is essential for generalization across sites\.
Introducing local heads \(any ratio withHl≥1H\_\{l\}\\geq 1\) consistently improves per\-client Micro\-F1: SPH reaches 86\.6–87\.8% \(vs\. FedALA 84\.4%\), PTB\-XL reaches 70\.1–75\.2% \(vs\. FedALA 71\.7%\), and G12EC reaches 70\.7–74\.3% \(vs\. Ditto 73\.4%\)\. The exception is SXPH, where FedALA \(88\.2%\) remains the strongest, suggesting that adaptive aggregation is more effective for that distribution\. Notably, Ditto shows high variance on G12EC \(±\\pm6\.7 F1\), a known failure mode of proximal\-term methods on highly heterogeneous clients\. FedDualAtt remains stable across all four sites\.
TABLE I:ECG benchmark: per\-client and global Micro\-F1 / mAP \(%\)\. Baseline FL rows report mean±\\pmstd from the FedCVD paper\. FedDualAtt rows report mean±\\pmstd over 5 seeds\.Bold= best,underline= second\-best, ranked jointly\.
### IV\-CHead\-Ratio Ablation
Fig\.[2](https://arxiv.org/html/2607.06653#S4.F2)shows per\-client and global Micro\-F1 and mAP*delta*relative to the8Hg:0Hl8H\_\{g\}\\\!:\\\!0H\_\{l\}global\-only baseline, with one line per client across all nine head\-ratio configurations\. We can see three patterns emerge\. First, introducing even a single local head \(7Hg:1Hl7H\_\{g\}\\\!:\\\!1H\_\{l\}\) produces an immediate and substantial per\-client Micro\-F1 gain for most clients \(\+5\+5to\+10\+10pp on SPH, PTB\-XL, and G12EC\), confirming that local attention heads rapidly capture site\-specific patterns that global aggregation suppresses\. Second, per\-client gains are largely stable from7Hg:1Hl7H\_\{g\}\\\!:\\\!1H\_\{l\}onward, with the per\-client curves plateauing across the middle configurations \(6Hg:2Hl6H\_\{g\}\\\!:\\\!2H\_\{l\}through1Hg:7Hl1H\_\{g\}\\\!:\\\!7H\_\{l\}\), indicating that a small number of local heads is sufficient to capture most of the personalization benefit\. Third, global Micro\-F1 \(dashed purple\) declines monotonically as local heads increase, dropping−21\.9\-21\.9pp at0Hg:8Hl0H\_\{g\}\\\!:\\\!8H\_\{l\}relative to8Hg:0Hl8H\_\{g\}\\\!:\\\!0H\_\{l\}, because fully local attention cannot leverage cross\-client information during aggregation\. The same global–local trade\-off is visible in the mAP panel, with per\-client mAP also improving under moderate personalization while global mAP falls at high local\-head counts\. Taken together, the head ratioHg:HlH\_\{g\}\\\!:\\\!H\_\{l\}provides an interpretable knob for controlling this trade\-off, requiring no additional training objectives or hyperparameters\.
Figure 2:Per\-client and global Micro\-F1 \(left\) and mAP \(right\) delta relative to FedDualAtt8Hg:0Hl8H\_\{g\}\\\!:\\\!0H\_\{l\}across all head\-ratio configurations \(n=5n\\\!=\\\!5seeds\)\. Each colored line is one client; the dashed purple line is the global metric\. Positive values indicate gain over the global\-only baseline\.
## VConclusion
In this paper, we propose FedDualAtt, which addresses data heterogeneity in federated ECG classification by splitting transformer attention heads into a globally aggregated branch and a per\-client local branch\. Experiments on the FedCVD benchmark demonstrate two complementary benefits: the dual\-attention architecture alone \(8Hg:0Hl8H\_\{g\}\\\!:\\\!0H\_\{l\}\) raises Global Micro\-F1 to 72\.68%, surpassing all FL baselines, while introducing local heads consistently improves per\-client performance on three of four clients with gains that plateau after a single local head\. Our ablation reveals a monotonic trade\-off between global generalization and per\-client adaptation that is controlled entirely by the head ratioHg:HlH\_\{g\}\\\!:\\\!H\_\{l\}, without additional objectives or training phases\. These results suggest that architectural personalization at the attention level is an effective complement to algorithm\-level FL methods for heterogeneous clinical data\. The dual\-head partitioning principle extends naturally to other federated medical domains where clients share backbone features but differ in local signal characteristics, such as multi\-site electroencephalography\-based seizure detection or distributed radiology\. A promising direction for future work is automatic ratio selection, where each client learns a soft assignment of heads to the global or local pool, eliminating the need to manually tuneHg:HlH\_\{g\}\\\!:\\\!H\_\{l\}\.
## Acknowledgment
This research was supported by the National Science Foundation under the Directorate for Computer and Information Science and Engineering/Office of Advanced Cyberinfrastructure \(CISE/OAC\) Grant No\. 2600417, and was partially supported by the Florida State University Undergraduate Research Opportunity Program – Research Mentor Materials Grant\.
## References
- \[1\]S\. P\. Karimireddy, S\. Kale, M\. Mohri, S\. Reddi, S\. Stich, and A\. T\. Suresh\(2020\-07\)SCAFFOLD: stochastic controlled averaging for federated learning\.In37th International Conference on Machine Learning,Virtual,pp\. 5132–5143\.Cited by:[§II\-A](https://arxiv.org/html/2607.06653#S2.SS1.p1.1)\.
- \[2\]T\. Li, S\. Hu, A\. Beirami, and V\. Smith\(2021\-07\)Ditto: fair and robust federated learning through personalization\.In38th International Conference on Machine Learning,Virtual,pp\. 6357–6368\.Cited by:[§I](https://arxiv.org/html/2607.06653#S1.p2.1)\.
- \[3\]T\. Li, A\. K\. Sahu, M\. Zaheer, M\. Sanjabi, A\. Talwalkar, and V\. Smith\(2020\-03\)Federated optimization in heterogeneous networks\.In3rd Conference on Machine Learning and Systems,Austin, TX, USA,pp\. 429–450\.Cited by:[§II\-B](https://arxiv.org/html/2607.06653#S2.SS2.p1.1)\.
- \[4\]X\. Li, M\. Jiang, X\. Zhang, M\. Kamp, and Q\. Dou\(2021\-05\)FedBN: federated learning on non\-IID features via local batch normalization\.In9th International Conference on Learning Representations,Virtual\.Cited by:[§II\-B](https://arxiv.org/html/2607.06653#S2.SS2.p1.1)\.
- \[5\]B\. McMahan, E\. Moore, D\. Ramage, S\. Hampson, and B\. Aguera y Arcas\(2017\-04\)Communication\-efficient learning of deep networks from decentralized data\.In20th International Conference on Artificial Intelligence and Statistics,Fort Lauderdale, FL, USA,pp\. 1273–1282\.Cited by:[§I](https://arxiv.org/html/2607.06653#S1.p2.1)\.
- \[6\]A\. Natarajan, Y\. Chang, S\. Mariani, A\. Rahman, G\. Boverman, S\. Vij, and J\. Rubin\(2020\-09\)A wide and deep transformer neural network for 12\-lead ECG classification\.In2020 Computing in Cardiology,Rimini, Italy,pp\. 1–4\.Cited by:[§I](https://arxiv.org/html/2607.06653#S1.p1.1)\.
- \[7\]Y\. Sattar and L\. Chhabra\(2026\-01\)Electrocardiogram\.InStatPearls,Note:Updated June 5, 2023\. PMID: 31747210External Links:[Link](https://www.ncbi.nlm.nih.gov/books/NBK549803)Cited by:[§I](https://arxiv.org/html/2607.06653#S1.p1.1)\.
- \[8\]E\. Voita, D\. Talbot, F\. Moiseev, R\. Sennrich, and I\. Titov\(2019\)Analyzing multi\-head self\-attention: specialized heads do the heavy lifting, the rest can be pruned\.InProceedings of the 57th annual meeting of the association for computational linguistics,pp\. 5797–5808\.Cited by:[§II\-B](https://arxiv.org/html/2607.06653#S2.SS2.p1.1)\.
- \[9\]World Health Organization\(2025\-07\)Cardiovascular diseases \(CVDs\)\(Website\)Note:Accessed: 2026\-02\-17External Links:[Link](https://www.who.int/news-room/fact-sheets/detail/cardiovascular-diseases-(cvds))Cited by:[§I](https://arxiv.org/html/2607.06653#S1.p1.1)\.
- \[10\]J\. Zhang, Y\. Hua, H\. Wang, T\. Song, Z\. Xue, R\. Ma, and H\. Guan\(2023\-02\)FedALA: adaptive local aggregation for personalized federated learning\.In37th AAAI Conference on Artificial Intelligence,Washington, DC, USA,pp\. 11237–11244\.Cited by:[§I](https://arxiv.org/html/2607.06653#S1.p2.1)\.
- \[11\]Y\. Zhang, G\. Chen, Z\. Xu, J\. Wang, D\. Zeng, J\. Li, J\. Wang, Y\. Qi, and I\. King\(2024\)FedCVD: the first real\-world federated learning benchmark on cardiovascular disease data\.External Links:2411\.07050,[Link](https://arxiv.org/abs/2411.07050)Cited by:[1st item](https://arxiv.org/html/2607.06653#S1.I1.i1.p1.1),[§I](https://arxiv.org/html/2607.06653#S1.p2.1),[§II\-A](https://arxiv.org/html/2607.06653#S2.SS1.p1.1),[§IV\-A](https://arxiv.org/html/2607.06653#S4.SS1.p1.3)\.
- \[12\]Y\. Zhao, M\. Li, L\. Lai, N\. Suda, D\. Civin, and V\. Chandra\(2022\)Federated learning with non\-IID data\.External Links:1806\.00582,[Link](https://arxiv.org/abs/1806.00582)Cited by:[§I](https://arxiv.org/html/2607.06653#S1.p2.1)\.Similar Articles
Towards Federated Long-Tailed Graph Learning: An Energy-Guided Dual Decoupling Approach
This paper introduces FedEPD, a framework for federated graph learning under long-tailed data distributions. It uses an energy-guided dual decoupling approach to separate topological purification from semantic recalibration, achieving state-of-the-art performance on benchmarks with up to 4.97% accuracy improvement.
Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning
This paper introduces a hybrid quantum-inspired Kolmogorov-Arnold network for privacy-aware federated learning of ECG data, demonstrating reduced parameters and communication costs while improving classification metrics compared to traditional MLP.
One Round Is All You Need: Analytic Federated Learning for Task-Heterogeneous Multi-Label Medical Image Classification
Proposes an analytic federated learning framework that requires only one or two communication rounds for multi-label medical image classification under task heterogeneity, outperforming existing methods on ChestXray14 by up to 18.44 BACC and 13.24 AUC points.
Embedding-Based Federated Learning with Runtime Governance for Iron Deficiency Prediction
This paper presents an embedding-based federated learning pipeline for predicting iron deficiency from routine blood count data, deployed across two clinical sites with non-IID distributions. It demonstrates that personalized aggregation (FedMAP) outperforms standard FedAvg and local-only training, achieving higher ROC-AUC at both sites.
Federated Learning over Human-Body Communication for On-Body Edge Intelligence: A Survey, Taxonomy, and BODYFED-HBC Scheduling Vignette
This paper presents a comprehensive survey and taxonomy of federated learning over human-body communication for on-body edge intelligence, including a scheduling vignette called BODYFED-HBC.