基于 Serverless 的 gossip 训练 LSTM 故障检测器:在 NASA C-MAPSS 上与联邦、本地及集中式学习的同协议对比

arXiv cs.LG 论文

摘要

在 NASA C-MAPSS 涡扇发动机基准数据集上,对用于 LSTM 故障检测器的同步环形 gossip、FedAvg、本地训练和集中式训练进行同协议对比。结果表明,gossip 是联邦平均的一种实用 serverless 替代方案:无需协调器即可达到与 FedAvg 相当的性能。

arXiv:2609.35792v1 Announce Type: new Abstract: Industrial predictive maintenance increasingly depends on learning from equipment spread across sites whose sensor data cannot easily be pooled. Federated averaging (FedAvg) solves this with a central aggregation server; gossip learning removes the server, but its behaviour for recurrent failure-detection models has not been measured under controlled conditions. We compare synchronous ring gossip with FedAvg, isolated local training and a centralized reference for a stacked LSTM that detects imminent failure on the NASA C-MAPSS turbofan benchmark. All methods share one open implementation, architecture, initialization, optimizer, data split and training budget, and the primary endpoint uses one terminal window per test engine to avoid the statistical dependence of overlapping windows. On FD001 (five seeds), gossip reached a terminal-window F1 of 89.6 +/- 1.3%, compared with 89.9 +/- 1.1% for FedAvg, 83.6 +/- 6.7% for local training and 93.5 +/- 2.1% for centralized training, while transmitting the same payload as FedAvg without a coordinator. Node models agreed closely but not exactly (1.8% pairwise decision disagreement versus 5.6% without communication). Across FD002-FD004, peer communication improved terminal-window F1 over local training by 13-28 points; gossip matched FedAvg on FD003 and FD004 but was 4.3 points lower on the multi-condition FD002 subset. Simulated message loss, node failure and server outage changed neither method appreciably, whereas larger rings degraded gossip faster. Ring gossip is therefore a practical serverless alternative when data heterogeneity is moderate, and faster-mixing topologies become important as heterogeneity grows.
查看原文
查看缓存全文

缓存时间: 2026/09/30 09:36

# Serverless gossip training of LSTM failure detectors: A matched-protocol comparison with federated, local and centralized learning on NASA C-MAPSS
Source: [https://arxiv.org/html/2609.35792](https://arxiv.org/html/2609.35792)
Journal:Future Generation Computer SystemsYusuf ÖztürkAffiliation:Department of Electrical and Electronics Engineering, Antalya Bilim University, Antalya, TürkiyeEnes GöktekinBengisu AtlıAffiliation:Department of Electrical and Electronics Engineering, Antalya Bilim University, Antalya, TürkiyeAffiliation:Department of Computer Engineering, Antalya Bilim University, Antalya, TürkiyeAkın ÖztürkAffiliation:Graduate School of Natural and Applied Sciences, Ankara University, Ankara, TürkiyeZhixiang WangEmail:[zhixiang\.wang@northwestern\.edu](mailto:[email protected])Corresponding author:Co\-corresponding authors\.Affiliation:Department of Radiology, Feinberg School of Medicine, Northwestern University, Chicago, IL, 60611, USAUlas BagciEmail:[ulas\.bagci@northwestern\.edu](mailto:[email protected])Corresponding author:Co\-corresponding authors\.Affiliation:Department of Radiology, Feinberg School of Medicine, Northwestern University, Chicago, IL, 60611, USA

###### Abstract

Industrial predictive maintenance increasingly depends on learning from equipment spread across sites whose sensor data cannot easily be pooled\. Federated averaging \(FedAvg\) solves this with a central aggregation server; gossip learning removes the server, but its behaviour for recurrent failure\-detection models has not been measured under controlled conditions\. We compare synchronous ring gossip with FedAvg, isolated local training and a centralized reference for a stacked LSTM that detects imminent failure on the NASA C\-MAPSS turbofan benchmark\. All methods share one open implementation, architecture, initialization, optimizer, data split and training budget, and the primary endpoint uses one terminal window per test engine to avoid the statistical dependence of overlapping windows\. On FD001 \(five seeds\), gossip reached a terminal\-window F1 of 89\.6±\\pm1\.3%, compared with 89\.9±\\pm1\.1% for FedAvg, 83\.6±\\pm6\.7% for local training and 93\.5±\\pm2\.1% for centralized training, while transmitting the same payload as FedAvg without a coordinator\. Node models agreed closely but not exactly \(1\.8% pairwise decision disagreement versus 5\.6% without communication\)\. Across FD002–FD004, peer communication improved terminal\-window F1 over local training by 13–28 points; gossip matched FedAvg on FD003 and FD004 but was 4\.3 points lower on the multi\-condition FD002 subset\. Simulated message loss, node failure and server outage changed neither method appreciably, whereas larger rings degraded gossip faster\. Ring gossip is therefore a practical serverless alternative when data heterogeneity is moderate, and faster\-mixing topologies become important as heterogeneity grows\.

###### Keywords:

Predictive maintenance , Gossip learning , Decentralized learning , Federated learning , Edge computing , Long short\-term memory , C\-MAPSS

## 1Introduction

Predictive maintenance \(PdM\) uses condition\-monitoring data to anticipate failures and schedule interventions before breakdowns occur, reducing unplanned downtime and maintenance cost\[[15](https://arxiv.org/html/2609.35792#bib.bib1),[9](https://arxiv.org/html/2609.35792#bib.bib19)\]\. Deep sequence models have become standard tools for this task because they learn degradation patterns directly from multivariate sensor streams\[[31](https://arxiv.org/html/2609.35792#bib.bib4),[34](https://arxiv.org/html/2609.35792#bib.bib23),[16](https://arxiv.org/html/2609.35792#bib.bib24),[32](https://arxiv.org/html/2609.35792#bib.bib25)\]\. Their accuracy, however, depends on the amount and diversity of run\-to\-failure data, which in practice is distributed over plants, fleets or operators that are often unwilling or unable to centralize it\[[30](https://arxiv.org/html/2609.35792#bib.bib5),[23](https://arxiv.org/html/2609.35792#bib.bib2),[18](https://arxiv.org/html/2609.35792#bib.bib3)\]\.

Federated learning \(FL\) addresses this by training a shared model while raw data remain on the participating devices\[[19](https://arxiv.org/html/2609.35792#bib.bib7),[10](https://arxiv.org/html/2609.35792#bib.bib8)\]\. It has been applied to fault diagnosis\[[20](https://arxiv.org/html/2609.35792#bib.bib10)\], to anomaly detection under distribution shift\[[1](https://arxiv.org/html/2609.35792#bib.bib11)\], and to remaining\-useful\-life \(RUL\) prognostics across airlines\[[14](https://arxiv.org/html/2609.35792#bib.bib12)\], and a federated benchmark on C\-MAPSS was recently released\[[29](https://arxiv.org/html/2609.35792#bib.bib13)\]\. All of these systems rely on a server that collects and redistributes models in every round\. In edge and industrial settings this coordinator is a single point of failure, a communication hub whose load grows with the number of participants, and an organizational obstacle when no party is trusted to host it\. Gossip learning removes the server: each node averages its model with a few neighbours, and information spreads through the network over successive rounds\[[11](https://arxiv.org/html/2609.35792#bib.bib29),[3](https://arxiv.org/html/2609.35792#bib.bib30),[24](https://arxiv.org/html/2609.35792#bib.bib33)\]\. Decentralized stochastic gradient descent can, under suitable conditions, match the convergence of its centralized counterpart\[[17](https://arxiv.org/html/2609.35792#bib.bib14),[13](https://arxiv.org/html/2609.35792#bib.bib15)\], and gossip learning has been shown to be competitive with FL on several benchmarks\[[6](https://arxiv.org/html/2609.35792#bib.bib16)\]\.

For practitioners designing a serverless maintenance system, three questions remain open\. First, does peer\-to\-peer communication actually improve on what each site could learn alone, and by how much? Second, how much accuracy does removing the server cost relative to FedAvg when both are implemented and trained identically? Third, how do the answers change under faults, heterogeneous data and larger networks? Answering them requires a controlled comparison: many PdM studies compare methods implemented in different code bases with different budgets, and many evaluate on every sliding window of a small number of test engines, which treats strongly overlapping and therefore dependent windows as independent observations\.

This paper provides such a comparison for LSTM\-based imminent\-failure detection on the NASA C\-MAPSS turbofan benchmark\[[26](https://arxiv.org/html/2609.35792#bib.bib17)\]\. Our contributions are as follows:

- 1\.A matched\-protocol comparison of synchronous ring gossip, FedAvg, isolated local training and centralized training, in which all methods share one implementation, architecture, initialization, optimizer, data split and training budget, evaluated over repeated seeds with paired tests\.
- 2\.An evaluation design that uses one terminal window per test engine as the primary endpoint and resamples whole engines for all confidence intervals, avoiding inflated precision from overlapping windows\.
- 3\.Direct measurements of how closely gossip node models agree, together with an exact communication ledger that includes an event\-triggered gossip variant\.
- 4\.Simulations of message loss, node failure, server outage, non\-IID partitioning and network size, and a replication on all four C\-MAPSS subsets, which identify where ring gossip matches FedAvg and where its slower information mixing becomes costly\.
- 5\.Open code, data and per\-run results from which every reported number can be regenerated\.

Section[2](https://arxiv.org/html/2609.35792#S2)reviews related work\. Section[3](https://arxiv.org/html/2609.35792#S3)describes the learning task, the training protocols and the communication model, and Section[4](https://arxiv.org/html/2609.35792#S4)the data and evaluation design\. Section[5](https://arxiv.org/html/2609.35792#S5)reports the results, Section[6](https://arxiv.org/html/2609.35792#S6)discusses their implications and limitations, and Section[7](https://arxiv.org/html/2609.35792#S7)concludes\.

## 2Related work

### 2\.1Deep learning for failure prognostics

Data\-driven prognostics estimate either the remaining useful life of an asset or the probability that it will fail within a maintenance horizon\[[28](https://arxiv.org/html/2609.35792#bib.bib21),[7](https://arxiv.org/html/2609.35792#bib.bib6),[4](https://arxiv.org/html/2609.35792#bib.bib20)\]\. On C\-MAPSS, LSTM networks\[[8](https://arxiv.org/html/2609.35792#bib.bib26),[5](https://arxiv.org/html/2609.35792#bib.bib27),[34](https://arxiv.org/html/2609.35792#bib.bib23)\]and convolutional networks\[[16](https://arxiv.org/html/2609.35792#bib.bib24)\]are established baselines, and recent surveys summarize the wide range of deep architectures that have since been proposed\[[32](https://arxiv.org/html/2609.35792#bib.bib25)\]\. Edge computing platforms increasingly perform part of this processing close to the monitored equipment\[[21](https://arxiv.org/html/2609.35792#bib.bib22)\]\. Our focus is not a new predictor: we deliberately use a conventional two\-layer LSTM so that differences between conditions can be attributed to how models are trained and combined across sites\.

### 2\.2Federated learning for maintenance

FedAvg alternates local training on each client with server\-side averaging of model parameters\[[19](https://arxiv.org/html/2609.35792#bib.bib7)\]; its extensions address statistical heterogeneity, communication efficiency and privacy\[[10](https://arxiv.org/html/2609.35792#bib.bib8),[25](https://arxiv.org/html/2609.35792#bib.bib9)\]\. In manufacturing and prognostics, FL has been used for mixed fault diagnosis in rotating machinery\[[20](https://arxiv.org/html/2609.35792#bib.bib10)\], predictive maintenance and anomaly detection under data\-distribution shifts\[[1](https://arxiv.org/html/2609.35792#bib.bib11)\], and collaborative RUL prognostics among airlines on N\-CMAPSS with robust aggregation and decentralized validation\[[14](https://arxiv.org/html/2609.35792#bib.bib12)\]\. The FedCMAPSS benchmark standardizes federated RUL tasks on C\-MAPSS from IID to strongly heterogeneous client settings\[[29](https://arxiv.org/html/2609.35792#bib.bib13)\]\. These works establish that collaborative training is valuable for prognostics, but they all assume a central aggregator\.

### 2\.3Gossip and decentralized learning

Gossip protocols compute network\-wide aggregates through repeated local exchanges\[[11](https://arxiv.org/html/2609.35792#bib.bib29),[3](https://arxiv.org/html/2609.35792#bib.bib30)\]\. For averaging with a fixed doubly stochastic mixing matrix, disagreement contracts at a rate governed by the second\-largest eigenvalue modulus \(SLEM\) of that matrix\[[33](https://arxiv.org/html/2609.35792#bib.bib31),[3](https://arxiv.org/html/2609.35792#bib.bib30)\], and distributed subgradient methods combine such averaging with local optimization steps\[[22](https://arxiv.org/html/2609.35792#bib.bib32)\]\. Decentralized parallel SGD can match centralized SGD when the network mixes well enough\[[17](https://arxiv.org/html/2609.35792#bib.bib14)\], and a unified analysis covers local updates and changing topologies\[[13](https://arxiv.org/html/2609.35792#bib.bib15)\]\. Gossip learning applies these ideas to machine\-learning models without any coordinator\[[24](https://arxiv.org/html/2609.35792#bib.bib33)\], and a large empirical study found it competitive with FL across several tasks\[[6](https://arxiv.org/html/2609.35792#bib.bib16)\]\. Theory and general benchmarks therefore suggest that a well\-connected gossip network can approach FedAvg; how a sparse ring behaves for recurrent failure detectors, under realistic data fragmentation and faults, is what the present study measures\.

## 3Methods

### 3\.1Learning task

Each engine produces a multivariate time series of operational settings and sensor readings, one vector per operating cycle\. For an engine that fails at cycleTT, the remaining useful life at cyclettisRUL⁡\(t\)=T−t\\mathrm\{RUL\}\(t\)=T\-t\. An input window𝐗t∈ℝ50×25\\mathbf\{X\}\_\{t\}\\in\\mathbb\{R\}^\{50\\times 25\}contains the 50 consecutive cycles ending attt, and its label is

yt=𝟙\[RUL\(t\)≤H\],H=30cycles,y\_\{t\}=\\mathbb\{1\}\\left\[\\mathrm\{RUL\}\(t\)\\leq H\\right\],\\qquad H=30\\ \\text\{cycles\},\(1\)so that a positive prediction is an alarm that failure is expected within the maintenance horizonHH\. The horizon follows common practice for this benchmark and was fixed before any experiment\. A modelfθf\_\{\\theta\}outputs the probabilityy^t=fθ​\(𝐗t\)\\hat\{y\}\_\{t\}=f\_\{\\theta\}\(\\mathbf\{X\}\_\{t\}\)and is trained with the binary cross\-entropy loss; an alarm is raised wheny^t≥0\.5\\hat\{y\}\_\{t\}\\geq 0\.5\.

The classifier is a two\-layer LSTM\[[8](https://arxiv.org/html/2609.35792#bib.bib26),[5](https://arxiv.org/html/2609.35792#bib.bib27)\]with 100 and 50 hidden units, dropout 0\.2 after each layer, and a sigmoid output unit, giving 80,651 trainable parameters\. The first layer returns its full hidden\-state sequence and the second only its final state\. The cell equations are given in Supplementary Section S1\.

### 3\.2Training protocols

Training data are distributed acrossNNnodes \(edge sites\), each holding the complete histories of a disjoint set of training engines\. All protocols start from the same seed\-specific initial parameters, copied to every node, and proceed forK=50K=50rounds; in each round every node performs one epoch of local training on its own windows \(Fig\.[1](https://arxiv.org/html/2609.35792#S3.F1)\)\. The protocols differ only in what happens after local training\.

Figure 1:Training protocols and evaluation design\. \(a\)–\(c\) Local\-only training, FedAvg and ring gossip onN=10N=10nodes\. \(d\) Each communication round consists of one local epoch per node, a synchronization barrier and one mixing step\. \(e\) For a truncated test engine with official terminal RULRRobserved up to cycleTobsT\_\{\\mathrm\{obs\}\}, labels follow fromRUL⁡\(t\)=Tobs\+R−t\\mathrm\{RUL\}\(t\)=T\_\{\\mathrm\{obs\}\}\+R\-t; the primary endpoint scores only the terminal window of each engine, whereas the secondary analysis scores all overlapping windows\.#### Local\-only

Nodes never communicate\. This condition measures what each site can learn on its own and therefore quantifies the value of collaboration\.

#### FedAvg

After local training, nodes upload their parameters to a server, which returns the average weighted by the number of training windows at each node\[[19](https://arxiv.org/html/2609.35792#bib.bib7)\]\.

#### Ring gossip

Nodes are arranged in a ring, and nodeiiexchanges parameters only with nodesi−1i\-1andi\+1i\+1\(indices moduloNN\)\. After local training producesW~i\(k\)\\tilde\{W\}\_\{i\}^\{\(k\)\}, all nodes take a snapshot and mix synchronously:

Wi\(k\+1\)=13​\(W~i−1\(k\)\+W~i\(k\)\+W~i\+1\(k\)\)\.W\_\{i\}^\{\(k\+1\)\}=\\tfrac\{1\}\{3\}\\left\(\\tilde\{W\}\_\{i\-1\}^\{\(k\)\}\+\\tilde\{W\}\_\{i\}^\{\(k\)\}\+\\tilde\{W\}\_\{i\+1\}^\{\(k\)\}\\right\)\.\(2\)The mixed parameters are the starting point of the next round \(Algorithm[1](https://arxiv.org/html/2609.35792#alg1)\)\. No coordinator is involved at any stage\.

#### Centralized

A single model is trained on the union of all training windows\. It is not deployable when data cannot be pooled and serves as a reference\.

#### Variants

We additionally evaluate \(i\) class\-weighted FedAvg and gossip, which scale the positive\-class loss by the training negative\-to\-positive ratio \(capped at 20\), and \(ii\) event\-triggered gossip, in which a node initiates an exchange only if the relativeℓ2\\ell\_\{2\}change of its parameters since its last exchange is at least 2% or it has been silent for five rounds; an edge is used if either endpoint triggers, and a 16\-byte trigger message is counted on every directed link in every round\.

Algorithm 1Synchronous ring gossip training1:nodes

i=1,…,Ni=1,\\dots,Nwith local windows

𝒟i\\mathcal\{D\}\_\{i\}; common initialization

W\(0\)W^\{\(0\)\}; rounds

KK
2:

Wi\(0\)←W\(0\)W\_\{i\}^\{\(0\)\}\\leftarrow W^\{\(0\)\}for all

ii
3:for

k=0,…,K−1k=0,\\dots,K\-1do

4:foreach node

iiin paralleldo

5:

W~i\(k\)←\\tilde\{W\}\_\{i\}^\{\(k\)\}\\leftarrowone epoch of Adam on

𝒟i\\mathcal\{D\}\_\{i\}starting from

Wi\(k\)W\_\{i\}^\{\(k\)\}
6:endfor

7:wait until all nodes have finished⊳\\trianglerightround barrier

8:foreach node

iido

9:send

W~i\(k\)\\tilde\{W\}\_\{i\}^\{\(k\)\}to nodes

i−1i\-1and

i\+1i\+1
10:

Wi\(k\+1\)←13​\(W~i−1\(k\)\+W~i\(k\)\+W~i\+1\(k\)\)W\_\{i\}^\{\(k\+1\)\}\\leftarrow\\frac\{1\}\{3\}\\big\(\\tilde\{W\}\_\{i\-1\}^\{\(k\)\}\+\\tilde\{W\}\_\{i\}^\{\(k\)\}\+\\tilde\{W\}\_\{i\+1\}^\{\(k\)\}\\big\)
11:endfor

12:endfor

13:returnnode models

W1\(K\),…,WN\(K\)W\_\{1\}^\{\(K\)\},\\dots,W\_\{N\}^\{\(K\)\}

### 3\.3Mixing and consensus

Stacking the node parameters, Eq\. \([2](https://arxiv.org/html/2609.35792#S3.E2)\) is𝐖\(k\+1\)=A​𝐖~\(k\)\\mathbf\{W\}^\{\(k\+1\)\}=A\\,\\tilde\{\\mathbf\{W\}\}^\{\(k\)\}with a symmetric, doubly stochastic circulant matrixAAwhose rows contain1/31/3at positionsi−1i\-1,iiandi\+1i\+1\. Its eigenvalues areλj=13​\(1\+2​cos⁡\(2​π​j/N\)\)\\lambda\_\{j\}=\\frac\{1\}\{3\}\\left\(1\+2\\cos\(2\\pi j/N\)\\right\),j=0,…,N−1j=0,\\dots,N\-1\. For pure averaging without local training, the squared deviation from the network mean contracts by at leastSLEM​\(A\)2\\mathrm\{SLEM\}\(A\)^\{2\}per round\[[33](https://arxiv.org/html/2609.35792#bib.bib31),[3](https://arxiv.org/html/2609.35792#bib.bib30)\]; forN=10N=10,SLEM⁡\(A\)=0\.873\\mathrm\{SLEM\}\(A\)=0\.873, and forN=40N=40it rises to 0\.992\. Because every round also applies local gradient steps on different data, which move the models apart again, this contraction does not imply that trained networks reach consensus\[[17](https://arxiv.org/html/2609.35792#bib.bib14),[13](https://arxiv.org/html/2609.35792#bib.bib15)\]\. We therefore measure agreement directly \(Section[5\.2](https://arxiv.org/html/2609.35792#S5.SS2)\): as the root\-mean\-square distance of node parameters from their mean, and as the fraction of node pairs whose final models make different decisions on a fixed probe of validation windows \(eight per validation engine\)\.

### 3\.4Communication accounting

Communication is counted per message\. A model message carries 80,651 float32 parameters, i\.e\. 322,604 bytes\. WithN=10N=10, gossip sends2​N=202N=20directed messages per round and FedAvg sendsNNuploads andNNdownloads, so both transmit20×50×322,604=322,604,00020\\times 50\\times 322\{,\}604=322\{,\}604\{,\}000bytes \(307\.66 MiB\) over 50 rounds\. Total payload is therefore equal by construction; the protocols differ in its distribution, since each gossip node talks to two peers whereas the FedAvg server terminates all2​N2Ntransfers\. Headers, serialization, acknowledgements, initial model distribution and the exchange of normalization statistics are not counted\.

## 4Experimental setup

### 4\.1Data and labels

C\-MAPSS contains simulated run\-to\-failure trajectories of turbofan engines with three operational settings and 21 sensors per cycle\[[26](https://arxiv.org/html/2609.35792#bib.bib17)\]\. Each of its four subsets provides complete training trajectories, test trajectories truncated at an unknown point before failure, and the true RUL at the last observed test cycle \(Table[1](https://arxiv.org/html/2609.35792#S4.T1)\)\. FD001 is used for all main analyses; FD002–FD004 add multiple operating conditions and a second fault mode\.

Table 1:C\-MAPSS subsets and evaluation populations\. Test engines shorter than the 50\-cycle window are excluded; positives are windows with RUL≤\\leq30 cycles\.For a test engine observed up to cycleTobsT\_\{\\mathrm\{obs\}\}with official terminal RULRR, the failure cycle isTobs\+RT\_\{\\mathrm\{obs\}\}\+R, so every cycle of its truncated history receivesRUL⁡\(t\)=Tobs\+R−t\\mathrm\{RUL\}\(t\)=T\_\{\\mathrm\{obs\}\}\+R\-t\(Fig\.[1](https://arxiv.org/html/2609.35792#S3.F1)e\)\. Labels are aligned to the last cycle of each window and RUL values are never used as inputs\. Each time step has 25 features: the three operational settings, the cycle index and the 21 sensors\. Within each seed, 80% of the training engines are used for training and 20% for validation, split at the engine level\. Min–max scaling is fitted on the training engines only and applied unchanged to validation and test data\.

Training engines are assigned at random toN=10N=10nodes, so that each node holds eight FD001 training engines\. Validation and test engines are assigned to nodes independently at random; in the decentralized protocols each test engine is scored by the final model of its assigned node, without ensembling\. Pooled metrics are computed over all nodes’ predictions\.

### 4\.2Evaluation design and statistics

Consecutive windows of the same engine share up to 49 of 50 cycles and are strongly dependent\. The*primary endpoint*therefore scores exactly one window per test engine, the terminal window \(93 windows on FD001\)\. As a secondary analysis, every valid window is scored with stride 1, describing behaviour across early, mid and late degradation\. For both populations we report positive\-class F1 and average precision \(AP\) at the fixed threshold of 0\.5; accuracy is uninformative because only 4% of FD001 windows are positive, and precision, recall and accuracy are listed in Supplementary Table S1\.

Results are mean±\\pmstandard deviation over seeds\. A seed jointly determines the train/validation split, the node assignment, the initialization and the minibatch order, and is shared across methods, so comparisons are paired\. We report paired per\-seed differences with exact two\-sided sign\-flip tests; with five seeds the smallest attainablepp\-value is 0\.0625, so we interpret effect sizes rather than significance\. Within\-seed uncertainty is quantified with 95% bootstrap intervals that resample whole test engines \(500 replicates\), never individual windows\.

### 4\.3Implementation

All protocols use Adam\[[12](https://arxiv.org/html/2609.35792#bib.bib28)\]with learning rate10−310^\{\-3\}, batch size 200, gradient\-norm clipping at 5 and 50 rounds without early stopping; the final models are evaluated\. The optimizer state is reset at the start of each round for every protocol, including centralized training, so that all methods restart from their \(possibly mixed\) parameters in the same way\. Input weights use Xavier initialization, recurrent weights orthogonal initialization and forget\-gate biases one\. Two differences between protocols are unavoidable: centralized training takes fewer, larger\-population optimizer steps than the sum of local steps, and FedAvg weights nodes by sample count whereas ring mixing weights neighbours equally\. The network is simulated synchronously in a single process on CPU with deterministic algorithms \(PyTorch 2\)\. FD001 experiments were run on one machine; FD002–FD004 were run on a separate Linux server \(Python 3\.9, PyTorch 2\.5\.1\), and methods are compared only within a subset\.

### 4\.4Additional scenarios

With three seeds \(11, 22, 33\) on FD001, FedAvg and ring gossip were further compared under: independent loss of 20% of directed messages, where a gossip exchange is applied only if both directions arrive; permanent failure of one node from round 25, whose last model continues to score its test engines; a server outage between rounds 20 and 35, during which FedAvg nodes continue training locally; a non\-IID partition in which training engines are sorted by lifetime and assigned to nodes in contiguous blocks; removal of the cycle\-index input; andN=5N=5, 20 and 40 nodes with the total training data held fixed\. Finally, centralized, FedAvg, gossip and local\-only training were repeated on FD002, FD003 and FD004 with three seeds and no change to the protocol or hyperparameters\.

As an exploratory privacy diagnostic, a loss\-threshold membership test scores each training engine \(member\) and validation engine \(non\-member\) by the mean loss of the final model over its last 50 windows; the area under the ROC curve \(AUC\) measures how well members are distinguished \(0\.5 corresponds to chance\)\[[27](https://arxiv.org/html/2609.35792#bib.bib35)\]\.

## 5Results

### 5\.1Main comparison on FD001

Table[2](https://arxiv.org/html/2609.35792#S5.T2)and Fig\.[2](https://arxiv.org/html/2609.35792#S5.F2)summarize the matched comparison\. On the primary endpoint, centralized training reached 93\.5±\\pm2\.1% F1, FedAvg 89\.9±\\pm1\.1% and ring gossip 89\.6±\\pm1\.3%\. The paired gossip−\-FedAvg difference was−0\.3\-0\.3percentage points \(pp; SD 2\.3;p=0\.75p=0\.75\), smaller than the variation between seeds\. Local\-only training was lower and much less stable \(83\.6±\\pm6\.7%\), and gossip exceeded it by 6\.0 pp on average\. On all windows the ordering was the same: 85\.7±\\pm1\.2% \(centralized\), 80\.1±\\pm1\.5% \(FedAvg\), 79\.8±\\pm2\.9% \(gossip\) and 69\.1±\\pm5\.8% \(local\-only\)\. Here gossip exceeded local\-only training in every seed \(\+10\.7 pp\), and the threshold\-free AP showed the same pattern \(90\.6% for gossip, 90\.9% for FedAvg, 74\.0% for local\-only\)\.

Table 2:Main results on FD001 \(five seeds, mean±\\pmSD, threshold 0\.5\)\. Terminal: one window per test engine \(93 engines, 25 positive\)\. All windows: 8,255 windows \(332 positive\)\. Disagreement: fraction of node pairs whose final models give different decisions on the validation probe\. Payload: model bytes exchanged over 50 rounds\.Figure 2:Main comparison on FD001\. \(a\) Terminal\-window and \(b\) all\-window F1 for each seed \(dots\), with the mean \(black tick\) and±\\pm1 SD \(shaded bar\)\. \(c\) Paired per\-seed differences between ring gossip and each comparator on the terminal \(circles\) and all\-window \(squares\) populations\.The test set limits how finely strong methods can be separated\. With 25 positive terminal windows, one missed engine changes recall by 4 pp, and engine\-bootstrap intervals within a single seed were correspondingly wide: for seed 11, terminal F1 was 88\.0% \(95% CI 76\.0–96\.1%\) for gossip and 91\.7% \(81\.5–98\.0%\) for FedAvg, and the mean interval width across seeds was 19\.0 pp for gossip and 19\.7 pp for FedAvg\. All\-window intervals were narrower \(14\.7 and 13\.8 pp\) but still substantial, which illustrates why treating 8,255 overlapping windows as independent would overstate precision\.

Class weighting moved the operating point toward higher recall on terminal windows but did not improve all\-window F1, indicating that it mainly traded missed failures for false alarms at the fixed threshold\. When the threshold was instead selected on the validation engines to maximize F1, the ordering of the unweighted methods was unchanged \(centralized 95\.5±\\pm1\.8%, FedAvg 91\.6±\\pm1\.6%, gossip 89\.7±\\pm2\.2%, local\-only 85\.7±\\pm3\.7%\) and the advantage of class weighting shrank or disappeared \(FedAvg 93\.5±\\pm0\.8%, gossip 89\.2±\\pm3\.9%\)\.

### 5\.2Agreement between node models

Without communication, the parameter spread between nodes grew steadily throughout training, and the final local models disagreed on 5\.6% of probe decisions \(Fig\.[3](https://arxiv.org/html/2609.35792#S5.F3)\)\. Ring gossip kept the spread about an order of magnitude smaller and reduced disagreement to 1\.8±\\pm0\.5% after 50 rounds; FedAvg overwrites all node models with the server average and has no disagreement by construction\. Gossip therefore produces closely agreeing but not identical models: the residual disagreement reflects local steps taken after the last mixing step and the slow mixing of a ten\-node ring \(SLEM=0\.873\\mathrm\{SLEM\}=0\.873\)\. In deployment, the same engine could receive a slightly different risk score depending on which node evaluates it\.

Figure 3:Agreement between node models on FD001 \(mean over five seeds; shaded band: range across seeds\)\. \(a\) Root\-mean\-square distance of node parameters from the network mean \(log scale\)\. \(b\) Pairwise decision disagreement of node models on a fixed probe of validation windows, evaluated every five rounds\.
### 5\.3Communication

FedAvg and ring gossip each exchanged 1,000 model messages \(307\.66 MiB\) over 50 rounds, as derived in Section[3\.4](https://arxiv.org/html/2609.35792#S3.SS4)\. The event\-triggered variant transmitted 307\.1±\\pm0\.9 MiB including trigger messages and reached F1 similar to standard gossip\. With the change threshold fixed in advance at 2%, the relative parameter change after one local epoch remained above the threshold in almost every round, so nearly every node triggered every round and no communication was saved\. Larger thresholds or several local epochs between exchanges would be needed for savings; we did not tune the threshold on test data\. Because all protocols were simulated sequentially on one CPU, wall\-clock time reflects the number of local passes rather than deployment speed and is not used as a cost measure\.

### 5\.4Faults, heterogeneity and network size

Table[3](https://arxiv.org/html/2609.35792#S5.T3)and Fig\.[4](https://arxiv.org/html/2609.35792#S5.F4)report the additional FD001 scenarios\. Neither protocol degraded appreciably under 20% message loss \(all\-window F1 81\.2% for FedAvg and 79\.8% for gossip, compared with 81\.0% and 80\.3% without faults\), after the permanent failure of one node \(81\.2% and 80\.3%\), or during a temporary server outage, in which FedAvg nodes continued training locally and re\-synchronized afterwards \(81\.6%\)\. With the lifetime\-sorted non\-IID partition, both collaborative protocols lost 7–9 pp relative to the IID split; gossip \(73\.6%\) was not worse than FedAvg \(72\.4%\), and both remained well above local\-only training \(65\.0%\)\. Removing the cycle\-index input reduced F1 slightly \(80\.3% and 78\.3%\), showing that the models do not rely primarily on elapsed time\.

Table 3:FedAvg versus ring gossip under additional scenarios on FD001 \(three seeds, mean±\\pmSD, %\)\. Payload in MiB over 50 rounds\.Figure 4:FedAvg and ring gossip under additional scenarios on FD001 \(three seeds\)\. \(a\) All\-window F1 under faults, a lifetime\-sorted non\-IID partition and removal of the cycle input withN=10N=10\(dots: seeds; black tick: mean; bar:±\\pm1 SD\)\. \(b\) All\-window F1 when the same training data are split acrossN=5N=5to 40 nodes\.Splitting the same data across more nodes reduced F1 for both protocols, but faster for gossip: from 81\.8% atN=5N=5to 61\.2% atN=40N=40, compared with 84\.3% to 68\.2% for FedAvg \(Fig\.[4](https://arxiv.org/html/2609.35792#S5.F4)b\)\. AtN=40N=40each node holds only two training engines, and the ring’s SLEM of 0\.992 means that information from one node needs many rounds to reach distant nodes\. These experiments fragment a fixed dataset and model message delivery rather than a physical network; they probe data fragmentation, not the behaviour of large deployments\.

### 5\.5Replication on FD002–FD004

Fig\.[5](https://arxiv.org/html/2609.35792#S5.F5)and Table[4](https://arxiv.org/html/2609.35792#S5.T4)repeat the main comparison on the other three subsets without changing the protocol\. The benefit of communication held on every subset: gossip exceeded local\-only training on the primary endpoint by 13\.3 pp on FD002, 14\.1 pp on FD003 and 28\.4 pp on FD004, and did so in every seed\. The gap between gossip and FedAvg depended on the subset\. On FD003 and FD004 the two were close on terminal windows \(94\.9±\\pm2\.6% versus 94\.9±\\pm2\.6%, and 70\.0±\\pm5\.2% versus 71\.2±\\pm1\.7%\)\. On FD002, which combines six operating conditions with the largest number of engines, gossip was lower in every seed \(74\.5±\\pm1\.2% versus 78\.8±\\pm1\.2%;−4\.3\-4\.3pp\)\. On all windows, gossip was also below FedAvg on FD003 in every seed \(82\.5±\\pm1\.2% versus 86\.2±\\pm1\.0%\) and on FD002 \(50\.3±\\pm1\.2% versus 52\.5±\\pm1\.8%\)\.

Figure 5:Replication across all four C\-MAPSS subsets with an unchanged protocol\. Bars show the mean terminal\-window \(top\) and all\-window \(bottom\) F1; white dots are individual seeds\.Table 4:Results on FD002–FD004 \(three seeds, mean±\\pmSD, %\)\. Evaluation populations are listed in Table[1](https://arxiv.org/html/2609.35792#S4.T1)\.The six\-condition subsets FD002 and FD004 were considerably harder for every method\. The global min–max scaling used throughout does not normalize sensors per operating condition, and condition\-aware preprocessing, which is common for these subsets, would be expected to raise absolute performance\. We kept the FD001 protocol unchanged because the question was whether the relative ordering of the training protocols transfers, not how to maximize accuracy on each subset\.

### 5\.6Membership\-inference diagnostic

On FD001, the loss\-threshold membership test reached an AUC of 0\.82±\\pm0\.04 for local\-only models, 0\.76±\\pm0\.06 for the centralized model, 0\.60±\\pm0\.07 for gossip and 0\.57±\\pm0\.09 for FedAvg\. Membership was detectable above chance for every protocol and least so for the collaboratively trained models, whose parameters average information from many engines\. The diagnostic is confounded by differences between engines and is not a privacy guarantee\.

## 6Discussion

### 6\.1Collaboration matters most

The largest and most consistent effect in this study is the value of communication itself\. A node that sees only eight FD001 engines learns a markedly worse and less stable detector than any collaborative protocol, and on the harder subsets the gap widens to 13–28 pp of terminal\-window F1\. For a maintenance operator deciding whether to join a collaborative scheme, this is the first\-order consideration: the choice between federated and serverless aggregation is secondary to the decision to collaborate at all\.

### 6\.2When a ring is enough

On FD001, FD003 and FD004, ring gossip reached the same primary\-endpoint accuracy as FedAvg while transmitting the same payload without a coordinator\. This agrees with general empirical comparisons of gossip learning and FL\[[6](https://arxiv.org/html/2609.35792#bib.bib16)\]and with decentralized SGD theory, in which a sufficiently well\-mixing topology approaches centralized behaviour\[[17](https://arxiv.org/html/2609.35792#bib.bib14),[13](https://arxiv.org/html/2609.35792#bib.bib15)\]\. The limits of the ring became visible in two situations\. On FD002, with six operating conditions spread randomly over nodes, gossip trailed FedAvg by 4\.3 pp in every seed, and on all windows of FD003 it trailed by 3\.7 pp\. As the same data were fragmented over 20 and 40 nodes, gossip degraded faster than FedAvg\. Both observations are consistent with slow mixing: in a ring ofNNnodes the SLEM approaches one asNNgrows, so knowledge from one node reaches distant nodes only after many rounds, and heterogeneous local updates keep pulling the models apart in the meantime\.

These results suggest a practical rule for serverless PdM systems\. A sparse ring is adequate when node data are moderately heterogeneous and networks are small\. When operating regimes differ strongly across sites or many sites participate, the topology should mix faster, for example through additional chords, several gossip steps per round or time\-varying peer selection, all of which trade extra communication for faster agreement\[[3](https://arxiv.org/html/2609.35792#bib.bib30),[13](https://arxiv.org/html/2609.35792#bib.bib15)\]\. Our released code supports these variants, and quantifying this trade\-off is a direct next step\.

### 6\.3Robustness, communication and privacy in context

The simulated faults did not separate the protocols: both tolerated message loss and a single node failure, and FedAvg nodes simply trained locally during a server outage\. The robustness benefit of gossip in these settings therefore lies in not needing a coordinator at all, which matters for organizational trust and system design, rather than in higher accuracy under faults\. Total payload was identical atN=10N=10, but its distribution differs: each gossip node exchanges data with two peers, whereas the FedAvg server must terminate every transfer\. Communication\-efficient variants deserve further study, because the pre\-specified event trigger saved nothing when parameters changed by more than 2% per epoch\. Finally, keeping raw data on the nodes is not formal privacy\. Shared parameters can leak information about training data\[[35](https://arxiv.org/html/2609.35792#bib.bib34),[27](https://arxiv.org/html/2609.35792#bib.bib35)\], and the membership diagnostic confirms leakage above chance for all protocols; differential privacy or secure aggregation would be required for formal guarantees\.

### 6\.4Limitations

All experiments use simulated C\-MAPSS data\. The ten\-node partitions are constructed rather than observed, and the main analyses use the single\-condition FD001 subset; FD002–FD004 were evaluated with three seeds and without condition\-specific preprocessing\. Results may not transfer to real fleets with site\-specific operating regimes, sensor faults or label noise, and evaluation on N\-CMAPSS\[[2](https://arxiv.org/html/2609.35792#bib.bib18)\]and industrial multi\-site data is needed\. The network is simulated synchronously on one CPU, so latency, asynchronous operation and energy use on edge hardware were not measured\. The primary endpoint of FD001 contains only 93 test engines, which limits the resolution of comparisons between the stronger protocols\. The fixed RUL horizon of 30 cycles and decision threshold of 0\.5 define a single operating point, and cost\-sensitive threshold selection was not studied\.

## 7Conclusion

We compared serverless ring gossip, FedAvg, isolated local training and centralized training of an LSTM imminent\-failure detector on all four NASA C\-MAPSS subsets under a single matched protocol with repeated seeds and an engine\-level primary endpoint\. Peer communication consistently and substantially improved on local training\. Ring gossip matched FedAvg on FD001, FD003 and FD004 with the same payload and no coordinator, and its models agreed closely but not exactly\. On the heterogeneous FD002 subset and in larger rings, slower mixing made gossip measurably less accurate than FedAvg\. Within the limits of simulated data and a simulated network, ring gossip is a workable serverless option for collaborative failure detection when heterogeneity is moderate, and faster\-mixing topologies should be preferred as heterogeneity and network size grow\. All code, data and per\-run results are released to support verification and extension\.

## CRediT authorship contribution statement

Yusuf Öztürk:Conceptualization, Methodology, Formal analysis, Investigation, Data curation, Writing – original draft, Supervision, Project administration\.Enes Göktekin:Conceptualization, Methodology, Software, Validation, Investigation, Writing – original draft, Visualization\.Bengisu Atlı:Resources, Data curation, Software\.Akın Öztürk:Methodology, Formal analysis, Investigation, Data curation, Writing – original draft\.Zhixiang Wang:Software, Validation, Formal analysis, Investigation, Visualization, Writing – review & editing\.Ulas Bagci:Conceptualization, Supervision, Project administration, Writing – review & editing\.

## Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper\.

## Acknowledgements

This work was supported by the National Institutes of Health \(NIH\) under grants R01\-HL171376 and U01\-CA268808\. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health\.

## Declaration of generative AI and AI\-assisted technologies in the manuscript preparation process

During the preparation of this work the authors used ChatGPT \(OpenAI\) and Claude \(Anthropic\) in order to improve language clarity, restructure and edit the manuscript, check reference metadata, and assist with code review and analysis scripts\. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article\.

## Data availability

The NASA C\-MAPSS dataset is publicly available\[[26](https://arxiv.org/html/2609.35792#bib.bib17)\]\. The code, the C\-MAPSS data files used, experiment configurations, per\-run metrics and training histories, and the scripts that regenerate all predictions, tables and figures are available at[https://github\.com/ZhixiangWang\-CN/gossip\-lstm\-cmapss](https://github.com/ZhixiangWang-CN/gossip-lstm-cmapss)\.

## References

- \[1\]J\. Ahn, Y\. Lee, N\. Kim, C\. Park, and J\. Jeong\(2023\)Federated learning for predictive maintenance and anomaly detection using time series data distribution shifts in manufacturing processes\.Sensors23\(17\),pp\. 7331\.External Links:[Document](https://dx.doi.org/10.3390/s23177331)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.35792#S2.SS2.p1.1)\.
- \[2\]M\. Arias Chao, C\. Kulkarni, K\. Goebel, and O\. Fink\(2021\)Aircraft engine run\-to\-failure dataset under real flight conditions for prognostics and diagnostics\.Data6\(1\),pp\. 5\.External Links:[Document](https://dx.doi.org/10.3390/data6010005)Cited by:[§6\.4](https://arxiv.org/html/2609.35792#S6.SS4.p1.1)\.
- \[3\]S\. Boyd, A\. Ghosh, B\. Prabhakar, and D\. Shah\(2006\)Randomized gossip algorithms\.IEEE Trans\. Inf\. Theory52\(6\),pp\. 2508–2530\.External Links:[Document](https://dx.doi.org/10.1109/TIT.2006.874516)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.35792#S2.SS3.p1.1),[§3\.3](https://arxiv.org/html/2609.35792#S3.SS3.p1.1),[§6\.2](https://arxiv.org/html/2609.35792#S6.SS2.p2.1)\.
- \[4\]T\. P\. Carvalho, F\. A\. A\. M\. N\. Soares, R\. Vita, R\. d\. P\. Francisco, J\. P\. Basto, and S\. G\. S\. Alcalá\(2019\)A systematic literature review of machine learning methods applied to predictive maintenance\.Comput\. Ind\. Eng\.137,pp\. 106024\.External Links:[Document](https://dx.doi.org/10.1016/j.cie.2019.106024)Cited by:[§2\.1](https://arxiv.org/html/2609.35792#S2.SS1.p1.1)\.
- \[5\]F\. A\. Gers, J\. Schmidhuber, and F\. Cummins\(2000\)Learning to forget: continual prediction with LSTM\.Neural Comput\.12\(10\),pp\. 2451–2471\.External Links:[Document](https://dx.doi.org/10.1162/089976600300015015)Cited by:[§2\.1](https://arxiv.org/html/2609.35792#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.35792#S3.SS1.p2.1)\.
- \[6\]I\. Hegedűs, G\. Danner, and M\. Jelasity\(2021\)Decentralized learning works: an empirical comparison of gossip learning and federated learning\.J\. Parallel Distrib\. Comput\.148,pp\. 109–124\.External Links:[Document](https://dx.doi.org/10.1016/j.jpdc.2020.10.006)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.35792#S2.SS3.p1.1),[§6\.2](https://arxiv.org/html/2609.35792#S6.SS2.p1.1)\.
- \[7\]A\. Heng, S\. Zhang, A\. C\. C\. Tan, and J\. Mathew\(2009\)Rotating machinery prognostics: state of the art, challenges and opportunities\.Mech\. Syst\. Signal Process\.23\(3\),pp\. 724–739\.External Links:[Document](https://dx.doi.org/10.1016/j.ymssp.2008.06.009)Cited by:[§2\.1](https://arxiv.org/html/2609.35792#S2.SS1.p1.1)\.
- \[8\]S\. Hochreiter and J\. Schmidhuber\(1997\)Long short\-term memory\.Neural Comput\.9\(8\),pp\. 1735–1780\.External Links:[Document](https://dx.doi.org/10.1162/neco.1997.9.8.1735)Cited by:[§2\.1](https://arxiv.org/html/2609.35792#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.35792#S3.SS1.p2.1)\.
- \[9\]A\. K\. S\. Jardine, D\. Lin, and D\. Banjevic\(2006\)A review on machinery diagnostics and prognostics implementing condition\-based maintenance\.Mech\. Syst\. Signal Process\.20\(7\),pp\. 1483–1510\.External Links:[Document](https://dx.doi.org/10.1016/j.ymssp.2005.09.012)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p1.1)\.
- \[10\]P\. Kairouz H\. B\. McMahanet al\.\(2021\)Advances and open problems in federated learning\.Found\. Trends Mach\. Learn\.14\(1–2\),pp\. 1–210\.External Links:[Document](https://dx.doi.org/10.1561/2200000083)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.35792#S2.SS2.p1.1)\.
- \[11\]D\. Kempe, A\. Dobra, and J\. Gehrke\(2003\)Gossip\-based computation of aggregate information\.In44th Annual IEEE Symposium on Foundations of Computer Science,pp\. 482–491\.External Links:[Document](https://dx.doi.org/10.1109/SFCS.2003.1238221)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.35792#S2.SS3.p1.1)\.
- \[12\]D\. P\. Kingma and J\. Ba\(2015\)Adam: a method for stochastic optimization\.In3rd International Conference on Learning Representations \(ICLR\),External Links:[Document](https://dx.doi.org/10.48550/arXiv.1412.6980)Cited by:[§4\.3](https://arxiv.org/html/2609.35792#S4.SS3.p1.1)\.
- \[13\]A\. Koloskova, N\. Loizou, S\. Boreiri, M\. Jaggi, and S\. U\. Stich\(2020\)A unified theory of decentralized SGD with changing topology and local updates\.InProceedings of the 37th International Conference on Machine Learning,PMLR, Vol\.119,pp\. 5381–5393\.Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.35792#S2.SS3.p1.1),[§3\.3](https://arxiv.org/html/2609.35792#S3.SS3.p1.1),[§6\.2](https://arxiv.org/html/2609.35792#S6.SS2.p1.1),[§6\.2](https://arxiv.org/html/2609.35792#S6.SS2.p2.1)\.
- \[14\]D\. Landau, I\. de Pater, M\. Mitici, and N\. Saurabh\(2026\)Federated learning framework for collaborative remaining useful life prognostics: an aircraft engine case study\.Future Gener\. Comput\. Syst\.174,pp\. 107945\.External Links:[Document](https://dx.doi.org/10.1016/j.future.2025.107945)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.35792#S2.SS2.p1.1)\.
- \[15\]J\. Lee, F\. Wu, W\. Zhao, M\. Ghaffari, L\. Liao, and D\. Siegel\(2014\)Prognostics and health management design for rotary machinery systems—reviews, methodology and applications\.Mech\. Syst\. Signal Process\.42\(1–2\),pp\. 314–334\.External Links:[Document](https://dx.doi.org/10.1016/j.ymssp.2013.06.004)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p1.1)\.
- \[16\]X\. Li, Q\. Ding, and J\. Sun\(2018\)Remaining useful life estimation in prognostics using deep convolution neural networks\.Reliab\. Eng\. Syst\. Saf\.172,pp\. 1–11\.External Links:[Document](https://dx.doi.org/10.1016/j.ress.2017.11.021)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.35792#S2.SS1.p1.1)\.
- \[17\]X\. Lian, C\. Zhang, H\. Zhang, C\. Hsieh, W\. Zhang, and J\. Liu\(2017\)Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent\.InAdvances in Neural Information Processing Systems 30 \(NIPS 2017\),pp\. 5330–5340\.Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.35792#S2.SS3.p1.1),[§3\.3](https://arxiv.org/html/2609.35792#S3.SS3.p1.1),[§6\.2](https://arxiv.org/html/2609.35792#S6.SS2.p1.1)\.
- \[18\]P\. Mallioris, E\. Aivazidou, and D\. Bechtsis\(2024\)Predictive maintenance in Industry 4\.0: a systematic multi\-sector mapping\.CIRP J\. Manuf\. Sci\. Technol\.50,pp\. 80–103\.External Links:[Document](https://dx.doi.org/10.1016/j.cirpj.2024.02.003)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p1.1)\.
- \[19\]B\. McMahan, E\. Moore, D\. Ramage, S\. Hampson, and B\. A\. y Arcas\(2017\)Communication\-efficient learning of deep networks from decentralized data\.InProceedings of the 20th International Conference on Artificial Intelligence and Statistics \(AISTATS\),PMLR, Vol\.54,pp\. 1273–1282\.Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.35792#S2.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.35792#S3.SS2.SSS0.Px2.p1.1)\.
- \[20\]M\. Mehta, S\. Chen, H\. Tang, and C\. Shao\(2023\)A federated learning approach to mixed fault diagnosis in rotating machinery\.J\. Manuf\. Syst\.68,pp\. 687–694\.External Links:[Document](https://dx.doi.org/10.1016/j.jmsy.2023.05.012)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.35792#S2.SS2.p1.1)\.
- \[21\]D\. Mourtzis, J\. Angelopoulos, and N\. Panopoulos\(2022\)Design and development of an edge\-computing platform towards 5G technology adoption for improving equipment predictive maintenance\.Procedia Comput\. Sci\.200,pp\. 611–619\.External Links:[Document](https://dx.doi.org/10.1016/j.procs.2022.01.259)Cited by:[§2\.1](https://arxiv.org/html/2609.35792#S2.SS1.p1.1)\.
- \[22\]A\. Nedić and A\. Ozdaglar\(2009\)Distributed subgradient methods for multi\-agent optimization\.IEEE Trans\. Autom\. Control54\(1\),pp\. 48–61\.External Links:[Document](https://dx.doi.org/10.1109/TAC.2008.2009515)Cited by:[§2\.3](https://arxiv.org/html/2609.35792#S2.SS3.p1.1)\.
- \[23\]P\. Nunes, J\. Santos, and E\. Rocha\(2023\)Challenges in predictive maintenance – a review\.CIRP J\. Manuf\. Sci\. Technol\.40,pp\. 53–67\.External Links:[Document](https://dx.doi.org/10.1016/j.cirpj.2022.11.004)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p1.1)\.
- \[24\]R\. Ormándi, I\. Hegedűs, and M\. Jelasity\(2013\)Gossip learning with linear models on fully distributed data\.Concurr\. Comput\. Pract\. Exp\.25\(4\),pp\. 556–571\.External Links:[Document](https://dx.doi.org/10.1002/cpe.2858)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.35792#S2.SS3.p1.1)\.
- \[25\]A\. Rauniyar, D\. H\. Hagos, D\. Jha, J\. E\. Håkegård, U\. Bagci, D\. B\. Rawat, and V\. Vlassov\(2024\)Federated learning for medical applications: a taxonomy, current trends, challenges, and future research directions\.IEEE Internet Things J\.11\(5\),pp\. 7374–7398\.External Links:[Document](https://dx.doi.org/10.1109/JIOT.2023.3329061)Cited by:[§2\.2](https://arxiv.org/html/2609.35792#S2.SS2.p1.1)\.
- \[26\]A\. Saxena, K\. Goebel, D\. Simon, and N\. Eklund\(2008\)Damage propagation modeling for aircraft engine run\-to\-failure simulation\.In2008 International Conference on Prognostics and Health Management,Denver, CO,pp\. 1–9\.External Links:[Document](https://dx.doi.org/10.1109/PHM.2008.4711414)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p4.1),[§4\.1](https://arxiv.org/html/2609.35792#S4.SS1.p1.1),[Data availability](https://arxiv.org/html/2609.35792#Sx5.p1.1)\.
- \[27\]R\. Shokri, M\. Stronati, C\. Song, and V\. Shmatikov\(2017\)Membership inference attacks against machine learning models\.In2017 IEEE Symposium on Security and Privacy \(SP\),pp\. 3–18\.External Links:[Document](https://dx.doi.org/10.1109/SP.2017.41)Cited by:[§4\.4](https://arxiv.org/html/2609.35792#S4.SS4.p2.1),[§6\.3](https://arxiv.org/html/2609.35792#S6.SS3.p1.1)\.
- \[28\]X\. Si, W\. Wang, C\. Hu, and D\. Zhou\(2011\)Remaining useful life estimation – a review on the statistical data driven approaches\.Eur\. J\. Oper\. Res\.213\(1\),pp\. 1–14\.External Links:[Document](https://dx.doi.org/10.1016/j.ejor.2010.11.018)Cited by:[§2\.1](https://arxiv.org/html/2609.35792#S2.SS1.p1.1)\.
- \[29\]A\. Sorrenti, M\. Pennisi, C\. Spampinato, and S\. Palazzo\(2026\)FedCMAPSS: a benchmark for federated learning in remaining useful life estimation\.Note:arXiv preprint arXiv:2608\.26433External Links:[Document](https://dx.doi.org/10.48550/arXiv.2608.26433)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.35792#S2.SS2.p1.1)\.
- \[30\]F\. Tao, Q\. Qi, A\. Liu, and A\. Kusiak\(2018\)Data\-driven smart manufacturing\.J\. Manuf\. Syst\.48,pp\. 157–169\.External Links:[Document](https://dx.doi.org/10.1016/j.jmsy.2018.01.006)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p1.1)\.
- \[31\]J\. Wang, Y\. Ma, L\. Zhang, R\. X\. Gao, and D\. Wu\(2018\)Deep learning for smart manufacturing: methods and applications\.J\. Manuf\. Syst\.48,pp\. 144–156\.External Links:[Document](https://dx.doi.org/10.1016/j.jmsy.2018.01.003)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p1.1)\.
- \[32\]F\. Wu, Q\. Wu, Y\. Tan, and X\. Xu\(2024\)Remaining useful life prediction based on deep learning: a survey\.Sensors24\(11\),pp\. 3454\.External Links:[Document](https://dx.doi.org/10.3390/s24113454)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.35792#S2.SS1.p1.1)\.
- \[33\]L\. Xiao and S\. Boyd\(2004\)Fast linear iterations for distributed averaging\.Syst\. Control Lett\.53\(1\),pp\. 65–78\.External Links:[Document](https://dx.doi.org/10.1016/j.sysconle.2004.02.022)Cited by:[§2\.3](https://arxiv.org/html/2609.35792#S2.SS3.p1.1),[§3\.3](https://arxiv.org/html/2609.35792#S3.SS3.p1.1)\.
- \[34\]S\. Zheng, K\. Ristovski, A\. Farahat, and C\. Gupta\(2017\)Long short\-term memory network for remaining useful life estimation\.In2017 IEEE International Conference on Prognostics and Health Management \(ICPHM\),pp\. 88–95\.External Links:[Document](https://dx.doi.org/10.1109/ICPHM.2017.7998311)Cited by:[§1](https://arxiv.org/html/2609.35792#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.35792#S2.SS1.p1.1)\.
- \[35\]L\. Zhu, Z\. Liu, and S\. Han\(2019\)Deep leakage from gradients\.InAdvances in Neural Information Processing Systems 32 \(NeurIPS 2019\),pp\. 14774–14784\.Cited by:[§6\.3](https://arxiv.org/html/2609.35792#S6.SS3.p1.1)\.

相似文章

联邦学习用于分布式CNC刀具磨损预测

arXiv cs.LG

本文研究了联邦学习在CNC刀具磨损预测中的应用,表明在分布式制造环境中,联邦模型的性能接近集中式学习,并显著优于本地客户端基线。

良性及对抗性客户端异质性下航空发动机预测的鲁棒与个性化联邦学习

arXiv cs.LG

本文针对良性及对抗性客户端异质性下的航空发动机剩余使用寿命预测,对联邦学习进行了受控研究,评估了个性化和拜占庭鲁棒聚合方法。研究发现,共享表示个性化缩小了本地与集中式准确率之间的绝大部分差距;使用Krum的鲁棒聚合能有效缓解后门攻击;将两者结合可形成组合防御,在攻击成功率较低的同时,仅付出较小的准确率代价。