面向半导体制造的不确定性与业务感知剩余使用寿命估计
摘要
本文提出了一种预测性维护框架,通过深度学习和不确定性估计来改进半导体制造中剩余使用寿命的估计,以降低维护成本。
arXiv:2609.22160v1 Announce Type: new
Abstract: Semiconductor manufacturing relies on tightly interconnected components, so early identification of the assets most likely to fail is essential to prevent a single breakdown from disrupting the entire production pipeline. Maintenance planning must therefore balance unexpected failures against prematurely interrupted operating life. We present a Predictive Maintenance (PdM) framework combining Deep Learning (DL) sequence models and Simoultaneous Quantile Regression (SQR)
for uncertainty-aware Remaining Useful Life (RUL) estimation and risk-aware maintenance decisions. Several architectures are compared on ion-milling data from the 2018 PHM Data Challenge (PHM18), including architectures based on State Space Models (SSM), using prediction and business metrics: Unexpected Breaks (UB), Unexploited Lifetime (UL), and a cost-weighted objective. Diagonal State Spaces (S4D) delivers the best Remaining Useful Life (RUL) estimates across quantiles and, relative to Preventive Maintenance (PvM) baselines, substantially lowers business cost by avoiding systematically early interventions. The results support uncertainty-aware, cost-sensitive maintenance planning in semiconductor production.
查看缓存全文
缓存时间: 2026/09/22 09:15
# Uncertainty and Business-Aware Remaining Useful Life Estimation for Semiconductor Manufacturing
Source: [https://arxiv.org/html/2609.22160](https://arxiv.org/html/2609.22160)
AIArtificial IntelligenceMLMachine LearningDLDeep LearningIoTInternet of ThingsIIoTIndustrial Internet of ThingsCMCondition MonitoringHIHealth IndexRULRemaining Useful LifePdMPredictive MaintenancePvMPreventive MaintenanceRMReactive MaintenanceSSMState Space ModelsSQRSimultaneous Quantile RegressionPHMPrognostics and Health ManagementPHM182018 PHM Data ChallengeUBUnexpected BreaksULUnexploited LifetimePINNPhysics\-Informed Neural NetworkBNNBayesian Neural NetworkMLPMultilayer PerceptronCNNConvolutional Neural NetworkLSTMLong Short\-Term MemoryRNNRecurrent Neural NetworkGRUGated Recurrent UnitS4Structured State SpacesS4DDiagonal State SpacesS5Simplified Structured State SpacesTCNTemporal Convolutional NetworkQRQuantile RegressionMSEMean Squared ErrorRMSERoot Mean Squared ErrorMAEMean Absolute ErrorCDFCumulative Distribution FunctionIBEIon Beam EtchingIMEIon Milling EquipmentPBNParticle Beam NeutralizerFCPFlowcool Pressure Dropped Below Limit
## Uncertainty and Business\-Aware Remaining Useful Life Estimation for Semiconductor ManufacturingThanks:Corresponding author: davide\.frizzo\.1@studenti\.unipd\.itThanks:This work was funded by the European Union in the context of the Horizon Europe project “AIMS5\.0—Artificial Intelligence in Manufacturing Leading to Sustainability and Industry 5\.0”\. Grant Agreement ID: 101112089\.
Francesco BorsattiGian Antonio SustoAffiliation:Department of Information Engineering, University of Padova
###### Abstract
Semiconductor manufacturing relies on tightly interconnected components, so early identification of the assets most likely to fail is essential to prevent a single breakdown from disrupting the entire production pipeline\. Maintenance planning must therefore balance unexpected failures against prematurely interrupted operating life\. We present a[PdM](https://arxiv.org/html/2609.22160#id9)\([PdM](https://arxiv.org/html/2609.22160#id9)\) framework combining[DL](https://arxiv.org/html/2609.22160#id3)\([DL](https://arxiv.org/html/2609.22160#id3)\) sequence models and[SQR](https://arxiv.org/html/2609.22160#id13)\([SQR](https://arxiv.org/html/2609.22160#id13)\) for uncertainty\-aware[RUL](https://arxiv.org/html/2609.22160#id8)\([RUL](https://arxiv.org/html/2609.22160#id8)\) estimation and risk\-aware maintenance decisions\. Several architectures are compared on ion\-milling data from the[PHM18](https://arxiv.org/html/2609.22160#id15)\([PHM18](https://arxiv.org/html/2609.22160#id15)\), including architectures based on[SSM](https://arxiv.org/html/2609.22160#id12)\([SSM](https://arxiv.org/html/2609.22160#id12)\), using prediction and business metrics:[UB](https://arxiv.org/html/2609.22160#id16)\([UB](https://arxiv.org/html/2609.22160#id16)\),[UL](https://arxiv.org/html/2609.22160#id17)\([UL](https://arxiv.org/html/2609.22160#id17)\), and a cost\-weighted objective\.[S4D](https://arxiv.org/html/2609.22160#id26)\([S4D](https://arxiv.org/html/2609.22160#id26)\) delivers the best[RUL](https://arxiv.org/html/2609.22160#id8)estimates across quantiles and, relative to[PvM](https://arxiv.org/html/2609.22160#id10)\([PvM](https://arxiv.org/html/2609.22160#id10)\) baselines, substantially lowers business cost by avoiding systematically early interventions\. The results support uncertainty\-aware, cost\-sensitive maintenance planning in semiconductor production\.
*K*eywordsPredictive Maintenance⋅\\cdotDeep Learning⋅\\cdotUncertainty Estimation⋅\\cdotSemiconductor Manufacturing
## 1Introduction
Industry 4\.0 introduced theSmart Factory\[[1](https://arxiv.org/html/2609.22160#bib.bib2)\]: a connected manufacturing system intended to improve productivity, energy efficiency, sustainability, and maintenance responsiveness\[[2](https://arxiv.org/html/2609.22160#bib.bib3)\]\. These benefits are especially relevant to skill\-intensive semiconductor manufacturing, which depends on multiple tightly interconnected components, so the failure of a single asset can disrupt the entire production pipeline\. Identifying the components most likely to fail in the near term is therefore essential to enable timely maintenance, preserve pipeline continuity, and avoid costly repairs\[[3](https://arxiv.org/html/2609.22160#bib.bib4)\]\. Traditional approaches include[RM](https://arxiv.org/html/2609.22160#id11)\([RM](https://arxiv.org/html/2609.22160#id11)\) and[PvM](https://arxiv.org/html/2609.22160#id10), in which maintenance is performed after failure or at fixed intervals, respectively\. On the other hand, smart\-manufacturing[IoT](https://arxiv.org/html/2609.22160#id4)\([IoT](https://arxiv.org/html/2609.22160#id4)\) sensors provide data that[AI](https://arxiv.org/html/2609.22160#id1)\([AI](https://arxiv.org/html/2609.22160#id1)\),[ML](https://arxiv.org/html/2609.22160#id2)\([ML](https://arxiv.org/html/2609.22160#id2)\), and[DL](https://arxiv.org/html/2609.22160#id3)models can use for industrial prediction\[[4](https://arxiv.org/html/2609.22160#bib.bib6),[5](https://arxiv.org/html/2609.22160#bib.bib5)\]\. In[PdM](https://arxiv.org/html/2609.22160#id9), these models identify when maintenance is required, balancing[RM](https://arxiv.org/html/2609.22160#id11)against the conservatism of[PvM](https://arxiv.org/html/2609.22160#id10)\. This can lower costs and downtime while improving dependability and operational efficiency\[[6](https://arxiv.org/html/2609.22160#bib.bib7),[7](https://arxiv.org/html/2609.22160#bib.bib8)\]\.[PdM](https://arxiv.org/html/2609.22160#id9)usually estimates a machine’s[RUL](https://arxiv.org/html/2609.22160#id8)with[ML](https://arxiv.org/html/2609.22160#id2)or[DL](https://arxiv.org/html/2609.22160#id3)models \(Section[2](https://arxiv.org/html/2609.22160#S2)\)\. Overestimation can cause failures \([UB](https://arxiv.org/html/2609.22160#id16)\), whereas underestimation triggers unnecessarily early maintenance \([UL](https://arxiv.org/html/2609.22160#id17)\)\.
In this study we augment[DL](https://arxiv.org/html/2609.22160#id3)\-based[RUL](https://arxiv.org/html/2609.22160#id8)estimation with[SQR](https://arxiv.org/html/2609.22160#id13)uncertainty estimates\[[8](https://arxiv.org/html/2609.22160#bib.bib34)\]and business metrics \(Section[4\.3](https://arxiv.org/html/2609.22160#S4.SS3)\) to select an appropriate maintenance trade\-off\. We evaluate the framework on ion\-milling etching data, a critical stage of integrated\-circuit production\.
The paper is organised as follows\. Section[2](https://arxiv.org/html/2609.22160#S2)provides background on[RUL](https://arxiv.org/html/2609.22160#id8)estimation, and Section[3](https://arxiv.org/html/2609.22160#S3)describes the proposed approach\. Section[4](https://arxiv.org/html/2609.22160#S4)presents the experimental setup, including the dataset description and preprocessing steps, and Section[5](https://arxiv.org/html/2609.22160#S5)reports the experimental evaluation\. Finally, Section[6](https://arxiv.org/html/2609.22160#S6)discusses conclusions and future work\.
## 2Related Work
The aim of[PdM](https://arxiv.org/html/2609.22160#id9)is to model machine degradation by using signals from multiple[CM](https://arxiv.org/html/2609.22160#id6)\([CM](https://arxiv.org/html/2609.22160#id6)\) sensors as inputs and producing a so\-called[HI](https://arxiv.org/html/2609.22160#id7)\([HI](https://arxiv.org/html/2609.22160#id7)\) that indicates when the equipment reaches a low health level and requires maintenance\. When this threshold is reached, the machine raises an alarm and maintenance scheduling begins\. One of the most commonly used health indicators is[RUL](https://arxiv.org/html/2609.22160#id8), a decreasing signal that represents the amount of useful life left in the equipment\. Consequently,[PdM](https://arxiv.org/html/2609.22160#id9)is normally framed as a[RUL](https://arxiv.org/html/2609.22160#id8)estimation problem\.
Different kinds of predictive models can be used to estimate[RUL](https://arxiv.org/html/2609.22160#id8)\. Traditionally, model\-based approaches that represent the physics of the equipment under analysis were employed\[[9](https://arxiv.org/html/2609.22160#bib.bib24)\]\. Their knowledge of the system dynamics enables accurate simulation even when the system configuration changes\. However, their computational requirements and deployment costs often prevent economical scaling\. Owing to advances in[AI](https://arxiv.org/html/2609.22160#id1), most recent[RUL](https://arxiv.org/html/2609.22160#id8)frameworks instead use data\-driven approaches\. These methods exploit the large volumes of data produced by[CM](https://arxiv.org/html/2609.22160#id6)sensors to model relationships between sensor measurements and machine degradation\. They can produce[RUL](https://arxiv.org/html/2609.22160#id8)estimates within fractions of a second and are therefore more readily deployed on devices with limited memory and compute capabilities\. Their main limitation is the lack of physical knowledge, which can reduce their effectiveness when the machine configuration changes\. Consequently, hybrid models, usually based on[PINN](https://arxiv.org/html/2609.22160#id18)\([PINN](https://arxiv.org/html/2609.22160#id18)\)s, have recently been developed\[[10](https://arxiv.org/html/2609.22160#bib.bib26),[11](https://arxiv.org/html/2609.22160#bib.bib25)\]\. This study focuses on data\-driven[RUL](https://arxiv.org/html/2609.22160#id8)estimation, for which both[ML](https://arxiv.org/html/2609.22160#id2)\- and[DL](https://arxiv.org/html/2609.22160#id3)\-based approaches can be used\.[ML](https://arxiv.org/html/2609.22160#id2)models generally require an additional preprocessing step to extract meaningful features from the raw signals\[[12](https://arxiv.org/html/2609.22160#bib.bib17)\]\. These typically include statistical \(i\.e\., mean, root mean square, skewness, kurtosis, …\), temporal, and frequency \(i\.e\., power spectrum\) features, which are computed over sliding windows to better highlight the equipment’s life degradation trend\. The enlarged feature set is then normally used to train[ML](https://arxiv.org/html/2609.22160#id2)models\[[13](https://arxiv.org/html/2609.22160#bib.bib18)\]but can also provide additional context to[DL](https://arxiv.org/html/2609.22160#id3)\-based approaches\[[14](https://arxiv.org/html/2609.22160#bib.bib10)\]\.[DL](https://arxiv.org/html/2609.22160#id3)models are much more expressive than[ML](https://arxiv.org/html/2609.22160#id2)architectures and their hierarchical structure, made up of multiple hidden layers, enables automatic extraction of the features needed for accurate[RUL](https://arxiv.org/html/2609.22160#id8)estimation\. This comes at the cost of larger models in terms of memory and compute\. Moreover these models normally require a significant amount of data to achieve optimal performance\. Nevertheless, modern[IIoT](https://arxiv.org/html/2609.22160#id5)\([IIoT](https://arxiv.org/html/2609.22160#id5)\) systems produce large volumes of data, enabling the development of numerous[DL](https://arxiv.org/html/2609.22160#id3)\-based approaches to[RUL](https://arxiv.org/html/2609.22160#id8)estimation\[[15](https://arxiv.org/html/2609.22160#bib.bib20)\]\. Common architectures include[LSTM](https://arxiv.org/html/2609.22160#id22)\([LSTM](https://arxiv.org/html/2609.22160#id22)\)s\[[16](https://arxiv.org/html/2609.22160#bib.bib21)\],[RNN](https://arxiv.org/html/2609.22160#id23)\([RNN](https://arxiv.org/html/2609.22160#id23)\)s\[[17](https://arxiv.org/html/2609.22160#bib.bib11)\],[CNN](https://arxiv.org/html/2609.22160#id21)\([CNN](https://arxiv.org/html/2609.22160#id21)\)–[LSTM](https://arxiv.org/html/2609.22160#id22)combinations\[[18](https://arxiv.org/html/2609.22160#bib.bib22)\], and attention layers\[[19](https://arxiv.org/html/2609.22160#bib.bib23)\], which model multivariate[IoT](https://arxiv.org/html/2609.22160#id4)time series\.
The[RUL](https://arxiv.org/html/2609.22160#id8)estimation task is not new in the context of semiconductor manufacturing\. It gained popularity through[PHM18](https://arxiv.org/html/2609.22160#id15)\[[20](https://arxiv.org/html/2609.22160#bib.bib1)\], which introduced a high\-quality, challenging dataset on ion\-milling machines, an important type of equipment in chip fabrication that requires careful monitoring\. Several studies, including this one, have used this data source to train and validate their[PdM](https://arxiv.org/html/2609.22160#id9)approaches for wafer fabrication\[[2](https://arxiv.org/html/2609.22160#bib.bib3)\]\. The authors of\[[21](https://arxiv.org/html/2609.22160#bib.bib12)\]combine a[TCN](https://arxiv.org/html/2609.22160#id28)\([TCN](https://arxiv.org/html/2609.22160#id28)\), an[LSTM](https://arxiv.org/html/2609.22160#id22), and attention to predict[RUL](https://arxiv.org/html/2609.22160#id8)on[PHM18](https://arxiv.org/html/2609.22160#id15)\. The study in\[[22](https://arxiv.org/html/2609.22160#bib.bib13)\]addresses an important problem in[PHM](https://arxiv.org/html/2609.22160#id14)\([PHM](https://arxiv.org/html/2609.22160#id14)\): varying machine operating conditions and fault types\. In particular, a[TCN](https://arxiv.org/html/2609.22160#id28)\-based model is trained on the base fault in[PHM18](https://arxiv.org/html/2609.22160#id15)and then fine\-tuned on other fault types\. Finally,\[[23](https://arxiv.org/html/2609.22160#bib.bib14)\]considers a multiscale Transformer\[[24](https://arxiv.org/html/2609.22160#bib.bib29)\]\. An[LSTM](https://arxiv.org/html/2609.22160#id22)layer is integrated into the attention block to combine its features with those extracted by the Transformer through a Hadamard product\.
Although[DL](https://arxiv.org/html/2609.22160#id3)\-based approaches can accurately determine the end of a machine’s life, their predictions involve uncertainty arising both from the data \(i\.e\., aleatoric uncertainty\) and from the model itself \(i\.e\., epistemic uncertainty\)\. Quantifying uncertainty in[PHM](https://arxiv.org/html/2609.22160#id14)is crucial because a missed maintenance intervention can have catastrophic economic consequences\. Several approaches to uncertainty estimation exist in the[DL](https://arxiv.org/html/2609.22160#id3)literature\[[25](https://arxiv.org/html/2609.22160#bib.bib27)\]; most are based on Bayesian inference and[BNN](https://arxiv.org/html/2609.22160#id19)\([BNN](https://arxiv.org/html/2609.22160#id19)\)s\[[26](https://arxiv.org/html/2609.22160#bib.bib28)\]\. More recently, Quantile Regression approaches have been introduced as lightweight, efficient alternatives to[BNN](https://arxiv.org/html/2609.22160#id19)s\[[27](https://arxiv.org/html/2609.22160#bib.bib33),[28](https://arxiv.org/html/2609.22160#bib.bib35)\]\.
## 3Proposed Method
This section presents the proposed[PdM](https://arxiv.org/html/2609.22160#id9)framework for uncertainty\-aware[RUL](https://arxiv.org/html/2609.22160#id8)estimation\.
### 3\.1[RUL](https://arxiv.org/html/2609.22160#id8)Estimation
A[RUL](https://arxiv.org/html/2609.22160#id8)model receives sensor signals from monitored equipment\. Data are split into intervals between maintenance interventions, denoted asrun\-to\-failure cycles; each cyclei∈\{1,…,N\}i\\in\\\{1,\\dots,N\\\}is processed independently\. Sensor readings are𝒳∈ℝn×m\\mathcal\{X\}\\in\\mathbb\{R\}^\{n\\times m\}, wherennis the number of samples andmmis the number of sensors\. The unmeasurable[RUL](https://arxiv.org/html/2609.22160#id8)target is𝒴∈ℝn\\mathcal\{Y\}\\in\\mathbb\{R\}^\{n\}and is defined by a degradation model, most commonly linear or piecewise linear\. The linear degradation model is the simplest of the two and defines the target signal as follows:
RUL=\{n−1,n−2,…,0\}\.\\text\{RUL\}=\\\{n\-1,n\-2,\\dots,0\\\}\.\(1\)
This model assumes that the machine’s health begins to degrade as soon as it starts operating and continues to degrade until the final sample, at which point the health status reaches zero and the equipment reaches the end of its life\. By contrast, the piecewise\-linear model is defined asRUL=\{MAX\_RUL,…,MAX\_RUL,MAX\_RUL−1,…,0\}\\text\{RUL\}=\\\{\\text\{MAX\\\_\{RUL\}\},\\dots,\\text\{MAX\\\_\{RUL\}\},\\text\{MAX\\\_\{RUL\}\}\-1,\\dots,0\\\}\. The health status starts atMAX\_RUL, remains constant forn−MAX\_RULn\-\\text\{MAX\\\_\{RUL\}\}samples, and then decreases linearly towards zero as specified by \([1](https://arxiv.org/html/2609.22160#S3.E1)\)\. This model is generally preferred because it is reasonable to assume that the machine remains fully healthy for some time before declining towards failure\. The valueMAX\_RUL, also known as theelboworkneepoint, is a hyperparameter that can be defined using reasonable assumptions about machine life\-cycle durations or statistics computed from the training data\[[29](https://arxiv.org/html/2609.22160#bib.bib19)\]\. We use a piecewise\-linear degradation model withMAX\_RUL=500\\text\{MAX\\\_\{RUL\}\}=500, following\[[17](https://arxiv.org/html/2609.22160#bib.bib11),[23](https://arxiv.org/html/2609.22160#bib.bib14)\]\. Once the target signal is defined,[RUL](https://arxiv.org/html/2609.22160#id8)estimation can be framed as a regression task on multivariate time\-series data\. The learning objective is therefore to find a functionf:𝒳→𝒴f:\\mathcal\{X\}\\rightarrow\\mathcal\{Y\}such thatyi≈f\(xi\)y\_\{i\}\\approx f\(x\_\{i\}\)for every paired sample\(xi,yi\)\(x\_\{i\},y\_\{i\}\)\.
We adopt the encoder\-decoder architecture of\[[27](https://arxiv.org/html/2609.22160#bib.bib33)\]:
- •Encoder: projects inputs into a higher\-dimensional feature space\.
- •Feature Extractor: learns spatio\-temporal representations using an[RNN](https://arxiv.org/html/2609.22160#id23),[GRU](https://arxiv.org/html/2609.22160#id24)\([GRU](https://arxiv.org/html/2609.22160#id24)\),[LSTM](https://arxiv.org/html/2609.22160#id22), Transformer\[[24](https://arxiv.org/html/2609.22160#bib.bib29)\],[S4](https://arxiv.org/html/2609.22160#id25)\([S4](https://arxiv.org/html/2609.22160#id25)\)\[[30](https://arxiv.org/html/2609.22160#bib.bib30)\],[S4D](https://arxiv.org/html/2609.22160#id26)\[[31](https://arxiv.org/html/2609.22160#bib.bib32)\], or[S5](https://arxiv.org/html/2609.22160#id27)\([S5](https://arxiv.org/html/2609.22160#id27)\)\[[32](https://arxiv.org/html/2609.22160#bib.bib31)\]\.
- •Decoder: maps the extracted features to the[RUL](https://arxiv.org/html/2609.22160#id8)estimate through a linear head\.
### 3\.2Weighted Loss Function
Given the long life\-cycle durations in the[PHM18](https://arxiv.org/html/2609.22160#id15)dataset \(Section[4\.1](https://arxiv.org/html/2609.22160#S4.SS1)\), it is impractical to feed an entire signal to the model\. Therefore, similarly to\[[27](https://arxiv.org/html/2609.22160#bib.bib33)\], we use a sliding\-window approach to divide inputs into non\-overlapping subsequences on which predictions are performed\. In particular, this approach divides sequences into mini\-batches that are processed in parallel during training and validation\. Zero\-padding ensures that all windows share the same length\. The model outputs the[RUL](https://arxiv.org/html/2609.22160#id8)signal for an input window and the predictions for the windows representing a run\-to\-failure cycle are concatenated to obtain the final estimate for the entire cycle\. Because a piecewise\-linear degradation model is used for[RUL](https://arxiv.org/html/2609.22160#id8), the target signal is constant atMAX\_RULfor most windows\. A decreasing trend is observable only in the windows after the elbow point \(i\.e\., in the finalMAX\_RULsamples\)\. During training, the model processes one window at a time; because most windows have a constant target signal, the model may learn to always predict a constant[RUL](https://arxiv.org/html/2609.22160#id8), making its predictions practically useless\. To address this challenge, which resembles class imbalance in classification tasks, we employ a weighted loss function\. Consider a generic loss functionℒ\(f\(x\),y^\)\\mathcal\{L\}\(f\(x\),\\hat\{y\}\)wheref\(x\)f\(x\)is the model prediction on an input window andy^\\hat\{y\}is the corresponding[RUL](https://arxiv.org/html/2609.22160#id8)target signal\. The loss function employed in the proposed approach can be formalized as follows:
\{1Wconstant⋅ℒ\(f\(x\),y^\)ify^is constant1Wdecreasing⋅ℒ\(f\(x\),y^\)otherwise,\\begin\{cases\}\\frac\{1\}\{W\_\{constant\}\}\\cdot\\mathcal\{L\}\(f\(x\),\\hat\{y\}\)&\\text\{if\}\\ \\hat\{y\}\\ \\text\{is constant\}\\\\ \\frac\{1\}\{W\_\{decreasing\}\}\\cdot\\mathcal\{L\}\(f\(x\),\\hat\{y\}\)&\\text\{otherwise\}\\end\{cases\},\(2\)
whereWconstantW\_\{constant\}andWdecreasingW\_\{decreasing\}are the numbers of constant and decreasing windows, respectively\. With this definition, errors on decreasing windows are penalized more heavily than errors on constant windows\. Accurate prediction of the constant region of the[RUL](https://arxiv.org/html/2609.22160#id8)signal is not the priority, because the primary aim of a[PdM](https://arxiv.org/html/2609.22160#id9)framework is to detect the so\-calledmaintenance point, i\.e\., the time at which a maintenance intervention is most appropriate\.
### 3\.3Uncertainty Estimation through Quantile Regression
As stated in Section[1](https://arxiv.org/html/2609.22160#S1), quantifying predictive uncertainty is fundamental to finding the optimal trade\-off between overestimating and underestimating[RUL](https://arxiv.org/html/2609.22160#id8), thereby avoiding[UB](https://arxiv.org/html/2609.22160#id16)while limiting[UL](https://arxiv.org/html/2609.22160#id17)\. In semiconductor manufacturing, where equipment is operated by specialized personnel, uncertainty\-aware estimates are particularly important for enabling operators to make informed, risk\-aware maintenance decisions\. Bayesian uncertainty methods require priors and computationally costly sampling procedures, which can be difficult to apply in high\-dimensional settings\. We instead use[SQR](https://arxiv.org/html/2609.22160#id13)\[[8](https://arxiv.org/html/2609.22160#bib.bib34)\], which adds uncertainty quantification to a regression model with minimal architectural and training changes\.[QR](https://arxiv.org/html/2609.22160#id29)\([QR](https://arxiv.org/html/2609.22160#id29)\) estimates quantiles ofp\(y\|x\)p\(y\|x\), wherexxdenotes sensor readings andyydenotes[RUL](https://arxiv.org/html/2609.22160#id8)\. By contrast, classical regression models trained with[MSE](https://arxiv.org/html/2609.22160#id30)\([MSE](https://arxiv.org/html/2609.22160#id30)\) or[MAE](https://arxiv.org/html/2609.22160#id32)\([MAE](https://arxiv.org/html/2609.22160#id32)\) estimate the mean and median ofp\(y\|x\)p\(y\|x\), respectively\. The selected quantile encodes risk preference\. High quantiles produce optimistic[RUL](https://arxiv.org/html/2609.22160#id8)estimates, favouring production at greater failure risk; low quantiles are conservative and protect equipment but can waste useful life\. Given the[CDF](https://arxiv.org/html/2609.22160#id33)\([CDF](https://arxiv.org/html/2609.22160#id33)\) of the target variable,F\(y\)=P\(Y≤y\)F\(y\)=P\(Y\\leq y\), thequantile functionfor a quantile levelτ∈\[0,1\]\\tau\\in\[0,1\]isF−1\(τ\)=inf\{y∣F\(y\)≥τ\}F^\{\-1\}\(\\tau\)=\\inf\\\{y\\mid F\(y\)\\geq\\tau\\\}\. To estimate quantileτ\\tauofYYgivenxx, wherex∈ℝnx\\in\\mathbb\{R\}^\{n\}, the modely^=f^τ\(x\)\\hat\{y\}=\\hat\{f\}\_\{\\tau\}\(x\)approximates theconditional quantile function\. The model is trained using thepinball loss, which is appropriate for this task:
ℒτ\(y,y^\)=\{τ\(y−y^\)ify≥y^\(1−τ\)\(y^−y\)otherwise\.\\mathcal\{L\}\_\{\\tau\}\(y,\\hat\{y\}\)=\\begin\{cases\}\\tau\(y\-\\hat\{y\}\)\\quad\\text\{if\}\\ y\\geq\\hat\{y\}\\\\ \(1\-\\tau\)\(\\hat\{y\}\-y\)\\quad\\text\{otherwise\}\\end\{cases\}\.\(3\)
This training criterion alone limits estimation to the single quantileτ\\tau\. To estimate all quantiles ofp\(y\|x\)p\(y\|x\),[SQR](https://arxiv.org/html/2609.22160#id13)augments each input samplexix\_\{i\}with a randomly sampled quantileτi\\tau\_\{i\}, yielding\(xi,τi\)\(x\_\{i\},\\tau\_\{i\}\)\. The pinball loss forτi\\tau\_\{i\}is then used to update the model parameters:ℒτi\(yi,y^i\)\\mathcal\{L\}\_\{\\tau\_\{i\}\}\(y\_\{i\},\\hat\{y\}\_\{i\}\)\. The full training objective is thereforef^∈argminf1n∑i=1n𝔼τ∼U\(0,1\)\[ℒτ\(f\(xi,τ\),yi\)\]\\hat\{f\}\\in\\text\{argmin\}\_\{f\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathbb\{E\}\_\{\\tau\\sim U\(0,1\)\}\[\\mathcal\{L\}\_\{\\tau\}\(f\(x\_\{i\},\\tau\),y\_\{i\}\)\], whereU\(0,1\)U\(0,1\)denotes the uniform distribution over the interval\[0,1\]\[0,1\]\. At inference time, a quantile levelτ\\tauis selected and provided to the model together with the sensor readingsxxto obtain the corresponding conditional[RUL](https://arxiv.org/html/2609.22160#id8)estimatef^\(x,τ\)\\hat\{f\}\(x,\\tau\)\.
## 4Experimental Setup
This section describes the setup of the evaluation experiments reported in Section[5](https://arxiv.org/html/2609.22160#S5)\. Section[4\.1](https://arxiv.org/html/2609.22160#S4.SS1)presents the problem domain \(Section[4\.1](https://arxiv.org/html/2609.22160#S4.SS1.SSS0.Px1)\) and the[PHM18](https://arxiv.org/html/2609.22160#id15)benchmark dataset \(Section[4\.1](https://arxiv.org/html/2609.22160#S4.SS1.SSS0.Px2)\)\. Section[4\.2](https://arxiv.org/html/2609.22160#S4.SS2)then describes data preprocessing, and Section[4\.3](https://arxiv.org/html/2609.22160#S4.SS3)introduces the business metrics used to evaluate the models in practical terms\.
### 4\.1[PHM](https://arxiv.org/html/2609.22160#id14)Dataset
The benchmark dataset adopted for this study is[PHM18](https://arxiv.org/html/2609.22160#id15), introduced in the[PHM18](https://arxiv.org/html/2609.22160#id15)to enable comparisons among[PHM](https://arxiv.org/html/2609.22160#id14)solutions for semiconductor manufacturing\. It focuses on ion\-milling etching tools and faults associated with their use\.
#### Ion\-Milling Etching Process
Ion milling \([IBE](https://arxiv.org/html/2609.22160#id34)\([IBE](https://arxiv.org/html/2609.22160#id34)\)\) precisely removes wafer material\. A wafer is processed in a vacuum chamber through a multi\-step recipe specifying beam voltage and current, incidence angle, rotation speed, and duration\. An ion source accelerates inert\-gas ions, typically argon, into a collimated beam; impact ejects surface atoms through sputtering\. A rotating, tiltable stage promotes uniform removal, a shutter blocks the beam until conditions are reached, and the[PBN](https://arxiv.org/html/2609.22160#id36)\([PBN](https://arxiv.org/html/2609.22160#id36)\) controls beam profile and charge distribution\. Helium\-assisted water cooling prevents damaging temperature increases\. Grid and chamber wear and flowcool leaks can reduce quality, scrap wafers, and cause downtime\. Accurate health\-state and[RUL](https://arxiv.org/html/2609.22160#id8)estimates therefore support maintenance that avoids failures while preserving availability\.
#### Dataset Overview
[PHM18](https://arxiv.org/html/2609.22160#id15)contains readings from 20 etching tools under two operating modes\. Following the challenge split, 15 tools train the models and five test them:01M02, 02M02, 03M01, 04M01, 06M01\. The run\-to\-failure cycles in the dataset are computed based on three different failure modes\. In particular, the end of life is defined as the time at which an operator stops the machine to perform maintenance \(i\.e\., this generally differs from the time at which the fault is observed\):
- •F1F\_\{1\}:[FCP](https://arxiv.org/html/2609.22160#id37)\([FCP](https://arxiv.org/html/2609.22160#id37)\) low
- •F2F\_\{2\}:[FCP](https://arxiv.org/html/2609.22160#id37)high
- •F3F\_\{3\}: Flowcool leak
We consider onlyF1F\_\{1\}, the most represented fault type, following\[[22](https://arxiv.org/html/2609.22160#bib.bib13)\]; all subsequent descriptions and results use this setting\. This yields 703 training cycles and five test cycles\. Table[1](https://arxiv.org/html/2609.22160#S4.T1)reports training\-cycle statistics used by the[PvM](https://arxiv.org/html/2609.22160#id10)baselines\.
Table 1:Statistics of the[PHM](https://arxiv.org/html/2609.22160#id14)training\-cycle durationsThe mean exceeds the median and lies between the 0\.75 and 0\.9 quantiles, indicating strong right skew and exceptionally long upper\-tail cycles\. This heterogeneity makes the task harder\. Table[2](https://arxiv.org/html/2609.22160#S4.T2)further illustrates this variability: three test cycles exceed the training 0\.9 quantile, whereas Life 1 is exceptionally short\.
Table 2:Durations of the[PHM](https://arxiv.org/html/2609.22160#id14)test cycles
### 4\.2Data Preprocessing
We apply the following preprocessing steps before evaluation\. Inputs are normalised to\[0,1\]\[0,1\]withMinMaxScaler\. Of the 24 sensor signals, we select nine following\[[17](https://arxiv.org/html/2609.22160#bib.bib11)\]\.FIXTURESHUTTERPOSITIONis a binary operational\-state signal that switches from 0 to 1\[[23](https://arxiv.org/html/2609.22160#bib.bib14)\]; we retain all input features only for samples in which this signal equals 1, thereby considering only useful machine life\. We use the piecewise\-linear target of Section[3\.1](https://arxiv.org/html/2609.22160#S3.SS1), clipping[RUL](https://arxiv.org/html/2609.22160#id8)atMAX\_RUL=500\\text\{MAX\\\_\{RUL\}\}=500\. Finally, targets are normalised to\[0,1\]\[0,1\]during training and denormalised for metric computation\.
### 4\.3Business Metrics
As described in Section[3\.1](https://arxiv.org/html/2609.22160#S3.SS1), the primary objective of the proposed model is to accurately estimate the target[RUL](https://arxiv.org/html/2609.22160#id8)signal\. However, predictive accuracy alone is insufficient in practical industrial scenarios, where the estimated[RUL](https://arxiv.org/html/2609.22160#id8)must ultimately support maintenance decision\-making\. In particular, each prediction is used to determine the*maintenance point*, i\.e\., the time instant at which the system generates a maintenance notification indicating that the asset is approaching the end of its useful life; this point is denoted byTmT\_\{m\}\. Typically,TmT\_\{m\}is defined as the time at whichRUL=0\\text\{RUL\}=0\. However, because the maintenance notification must be issued sufficiently in advance to allow maintenance planners to schedule the intervention, procure the required resources, and complete the maintenance operation before a failure occurs\[[33](https://arxiv.org/html/2609.22160#bib.bib16)\], a more practical definition is needed\. Business metrics introduced in\[[34](https://arxiv.org/html/2609.22160#bib.bib9),[14](https://arxiv.org/html/2609.22160#bib.bib10)\]help define the maintenance point by accounting for the business and economic costs associated with different types of errors made by the[RUL](https://arxiv.org/html/2609.22160#id8)estimation model\. From an operational perspective, prediction errors have different consequences depending on their direction\. An overestimation of the[RUL](https://arxiv.org/html/2609.22160#id8)postpones the maintenance notification and may result in the asset failing before maintenance can be performed, leading to a so\-called[UB](https://arxiv.org/html/2609.22160#id16)\(i\.e\.,ρUB\\rho\_\{UB\}\)\. Conversely, an underestimation anticipates the maintenance notification, reducing the risk of failure but potentially causing components to be replaced while useful operating life remains\. This situation is termed[UL](https://arxiv.org/html/2609.22160#id17)\(i\.e\.,ρUL\\rho\_\{UL\}\)\. To account for this trade\-off, amaintenance windowmmis introduced as a safety margin\. Rather than scheduling maintenance when the predicted[RUL](https://arxiv.org/html/2609.22160#id8)reaches zero, the maintenance notification is triggered when the predicted[RUL](https://arxiv.org/html/2609.22160#id8)falls below the thresholdmm\. In other words, if the model predicts a maintenance point atTm∗T\_\{m\}^\{\*\}, the adjusted point isTm∗−mT\_\{m\}^\{\*\}\-m\. The maintenance window is subtracted from the maintenance point because a more preventive approach is generally preferred: an unexpected break typically costs more than not fully exploiting the machine’s lifetime\. Increasingmmdecreases the likelihood of unexpected failures by anticipating the maintenance intervention\. However, an excessively conservative value results in maintenance being performed too early, thereby reducing the utilisation of the asset and increasing operational costs\. The objective of a[PdM](https://arxiv.org/html/2609.22160#id9)system is therefore not only to minimise the prediction error, but also to identify the best compromise between[UB](https://arxiv.org/html/2609.22160#id16)and[UL](https://arxiv.org/html/2609.22160#id17)\. This trade\-off can be evaluated through the business metricJJproposed in\[[14](https://arxiv.org/html/2609.22160#bib.bib10)\], which combines[UB](https://arxiv.org/html/2609.22160#id16)and[UL](https://arxiv.org/html/2609.22160#id17)into a single score and enables the selection of the optimal maintenance windowmm\. MetricJJis defined as follows:
J=ρUB⋅cUB\+ρUL⋅cUL,J=\\rho\_\{UB\}\\cdot c\_\{UB\}\+\\rho\_\{UL\}\\cdot c\_\{UL\},\(4\)
whereρUB\\rho\_\{UB\}is the percentage of unexpected breaks across the test cycles,ρUL\\rho\_\{UL\}is the amount of unexploited life, andcUBc\_\{UB\}andcULc\_\{UL\}are hyperparameters representing the costs associated with[UB](https://arxiv.org/html/2609.22160#id16)and[UL](https://arxiv.org/html/2609.22160#id17), respectively\. The values and relative magnitudes of these costs depend on the specific application, but in most casescUB≫cULc\_\{UB\}\\gg c\_\{UL\}\.
## 5Experimental Results
This section reports the experimental evaluation of the proposed approach\. Section[5\.1](https://arxiv.org/html/2609.22160#S5.SS1)compares different sequence\-learning architectures across five quantile levels; their business metrics are compared at quantile 0\.5 \(i\.e\., the median\)\. Section[5\.2](https://arxiv.org/html/2609.22160#S5.SS2)then examines the best\-performing model across all five quantiles to demonstrate their effect on model predictions and business metrics\. To obtain robust results, we employ five\-fold cross\-validation on the training cycles while holding out the test cycles for evaluation\. All results in this section are averaged across folds\.
### 5\.1Comparative Evaluation of Benchmark Models
This section benchmarks different[DL](https://arxiv.org/html/2609.22160#id3)\-based backbones for[RUL](https://arxiv.org/html/2609.22160#id8)estimation within the general pipeline described in Section[3\.1](https://arxiv.org/html/2609.22160#S3.SS1)\. As outlined in Section[3\.3](https://arxiv.org/html/2609.22160#S3.SS3), the[QR](https://arxiv.org/html/2609.22160#id29)\-based approach enables evaluation at different quantile levels\. Table[3](https://arxiv.org/html/2609.22160#S5.T3)compares the models across five quantiles\. Quantile 0\.5 \(i\.e\., the median\) corresponds to the well\-known[MAE](https://arxiv.org/html/2609.22160#id32)loss and is used below for the business\-metric and accuracy–complexity comparisons\. A detailed quantile\-wise evaluation of the best model is reported in Section[5\.2](https://arxiv.org/html/2609.22160#S5.SS2)\.
Table 3:[RMSE](https://arxiv.org/html/2609.22160#id31)of the benchmark models at different evaluation quantiles\. The best result is shown in bold and the second\-best in italics\.Table[3](https://arxiv.org/html/2609.22160#S5.T3)shows that the[SSM](https://arxiv.org/html/2609.22160#id12)\-based architectures consistently provide the most accurate[RUL](https://arxiv.org/html/2609.22160#id8)predictions across the evaluated quantiles, with[S4D](https://arxiv.org/html/2609.22160#id26)achieving the strongest overall performance\. In contrast, classical sequence\-learning architectures and the non\-sequential[MLP](https://arxiv.org/html/2609.22160#id20)\([MLP](https://arxiv.org/html/2609.22160#id20)\) and Linear baselines struggle to model the[RUL](https://arxiv.org/html/2609.22160#id8)dynamics effectively\. Figure[1](https://arxiv.org/html/2609.22160#S5.F1)complements the aggregate metrics by showing the predicted[RUL](https://arxiv.org/html/2609.22160#id8)trajectories for the last 5000 samples of each test life\. The[SSM](https://arxiv.org/html/2609.22160#id12)models are the only approaches that consistently respond to the degradation phases:[S4D](https://arxiv.org/html/2609.22160#id26)closely follows the decline in Life 1 and captures the late\-life decreases in Lives 0 and 3, while[S4](https://arxiv.org/html/2609.22160#id25)and[S5](https://arxiv.org/html/2609.22160#id27)also reproduce parts of these trends\. Their errors are nevertheless apparent on Life 2, where the predicted trajectories depart markedly from the target after the decline begins\. In contrast, the[MLP](https://arxiv.org/html/2609.22160#id20), Linear,[LSTM](https://arxiv.org/html/2609.22160#id22),[GRU](https://arxiv.org/html/2609.22160#id24),[RNN](https://arxiv.org/html/2609.22160#id23), and Transformer baselines remain nearly constant for substantial portions of Lives 1 and 3 or collapse prematurely in Life 2, and therefore fail to represent the observed[RUL](https://arxiv.org/html/2609.22160#id8)evolution\. This trajectory\-level comparison is consistent with the superior overall accuracy of[S4D](https://arxiv.org/html/2609.22160#id26)in Table[3](https://arxiv.org/html/2609.22160#S5.T3)\.
Figure 1:[RUL](https://arxiv.org/html/2609.22160#id8)estimates from all benchmark models at quantile 0\.5\(a\)Unexpected Breaks\(b\)Unexploited Lifetime\(c\)Metric J
Figure 2:Business\-metric comparison of all benchmark models at quantile 0\.5Figure[2](https://arxiv.org/html/2609.22160#S5.F2)evaluates the maintenance decisions obtained from the[DL](https://arxiv.org/html/2609.22160#id3)models as the maintenance\-window size increases\. In this evaluation,[PdM](https://arxiv.org/html/2609.22160#id9)models are compared with[PvM](https://arxiv.org/html/2609.22160#id10)baselines, denoted bymeanandmedianbecause their maintenance points are computed from the mean and median training\-cycle durations, respectively\. The left panel \(Fig\.[2a](https://arxiv.org/html/2609.22160#S5.F2.sf1)\) reportsρUB\\rho\_\{UB\}, namely the fraction of test lives that experience an unexpected break before the planned intervention; the centre panel \(Fig\.[2b](https://arxiv.org/html/2609.22160#S5.F2.sf2)\) reportsρUL\\rho\_\{UL\}, i\.e\., the[UL](https://arxiv.org/html/2609.22160#id17)associated with an early intervention\. Finally, the right panel \(Fig\.[2c](https://arxiv.org/html/2609.22160#S5.F2.sf3)\) reports a sensitivity analysis ofJJ, defined by \([4](https://arxiv.org/html/2609.22160#S4.E4)\): its horizontal axis is the cost ratiocUB/cULc\_\{UB\}/c\_\{UL\}, while its vertical axis shows the correspondingJJvalues\. This ratio expresses how many minutes or cycles of[UL](https://arxiv.org/html/2609.22160#id17)have the same cost as one[UB](https://arxiv.org/html/2609.22160#id16)\. We use this ratio\-based representation because the absolute values ofcUBc\_\{UB\}andcULc\_\{UL\}must be determined for the specific application through a careful cost study, which should identify the pair that minimisesJJ\. Lower values ofJJindicate a more favourable maintenance trade\-off\. As the maintenance window increases, the[PdM](https://arxiv.org/html/2609.22160#id9)curves reduce the risk of unexpected breaks because maintenance is requested earlier; this reduction is accompanied by the expected increase in unexploited lifetime\. In contrast, the preventive baselines maintain nearly constant, exceptionally large[UL](https://arxiv.org/html/2609.22160#id17)values, reflecting their systematically early interventions induced by the right\-skewed training\-life distribution\. This behaviour is also evident in theJJpanel: despite their low or null breakage risk, the large amount of discarded life makes the baseline costs orders of magnitude higher than those of the[PdM](https://arxiv.org/html/2609.22160#id9)models\.
\(a\)Model parameters
\(b\)GFLOPs
Figure 3:Predictive accuracy–complexity trade\-off across benchmark modelsFigure[3](https://arxiv.org/html/2609.22160#S5.F3)compares predictive accuracy with model complexity \(i\.e\., the number of model parameters in the left panel\) and computational cost \(i\.e\., GFLOPs in the right panel\) at quantile 0\.5\.[S4D](https://arxiv.org/html/2609.22160#id26)achieves the lowest[RMSE](https://arxiv.org/html/2609.22160#id31)with fewer than one million parameters\. More importantly, it lies on the Pareto front in both panels: no competing model simultaneously improves its prediction error and its parameter count or computational cost\. This confirms that[S4D](https://arxiv.org/html/2609.22160#id26)combines high predictive accuracy with both a small memory footprint and high throughput\. The closely related[S4](https://arxiv.org/html/2609.22160#id25)and[S5](https://arxiv.org/html/2609.22160#id27)models also yield low errors, whereas the[MLP](https://arxiv.org/html/2609.22160#id20),[RNN](https://arxiv.org/html/2609.22160#id23),[GRU](https://arxiv.org/html/2609.22160#id24), and[LSTM](https://arxiv.org/html/2609.22160#id22)have substantially higher errors despite comparable or larger resource requirements\. The Linear model is also present on both Pareto fronts because its extremely small parameter count and computational cost cannot be matched by the other models\. Its prediction error, however, is too high for practical application; its Pareto optimality therefore reflects its exceptionally low resource usage rather than a useful accuracy\-efficiency compromise\. At the opposite extreme, the Transformer achieves accuracy comparable to that of the[SSM](https://arxiv.org/html/2609.22160#id12)models but at a prohibitively high complexity\. Overall, the[SSM](https://arxiv.org/html/2609.22160#id12)\-based models, especially[S4D](https://arxiv.org/html/2609.22160#id26), provide the most favourable and practically relevant accuracy\-complexity trade\-off in this benchmark\.
### 5\.2Quantile\-wise Analysis of the Best\-Performing Model
This section considers the performance of the best\-performing model,[S4D](https://arxiv.org/html/2609.22160#id26)according to Table[3](https://arxiv.org/html/2609.22160#S5.T3), at all evaluation quantiles to demonstrate how modelling uncertainty through[QR](https://arxiv.org/html/2609.22160#id29)shapes model behaviour\.
Table 4:[RMSE](https://arxiv.org/html/2609.22160#id31)of[S4D](https://arxiv.org/html/2609.22160#id26)at different evaluation quantilesTable[4](https://arxiv.org/html/2609.22160#S5.T4)shows how the effect of the evaluation quantile varies across test lives\. The metrics generally improve as the quantile increases for Lives 0, 2, and 3, while the results for Life 4 remain nearly constant across quantiles\. In contrast, the model performs less accurately on Life 1 at every quantile\. As discussed in Section[4](https://arxiv.org/html/2609.22160#S4), its exceptionally short duration differs substantially from the predominantly long and highly variable training life cycles, making it difficult for the model to generalise to this case\.
Figure 4:[S4D](https://arxiv.org/html/2609.22160#id26)[RUL](https://arxiv.org/html/2609.22160#id8)estimates across evaluation quantilesFigure[4](https://arxiv.org/html/2609.22160#S5.F4)shows that the quantile estimates generally capture the declining trend of the[RUL](https://arxiv.org/html/2609.22160#id8)for Lives 0 and 1, with lower quantiles providing more conservative predictions and upper quantiles yielding progressively larger[RUL](https://arxiv.org/html/2609.22160#id8)estimates\. The separation between quantile trajectories becomes more pronounced as degradation progresses, reflecting increased predictive uncertainty near failure\. Performance is less consistent for Lives 2 and 3: the model does not reproduce the abrupt terminal decrease in Life 2 and exhibits substantial quantile\-dependent bias for Life 3\. These results are probably due to the skewed life distribution described in Section[4\.1](https://arxiv.org/html/2609.22160#S4.SS1.SSS0.Px2), particularly the exceptionally short Life 1\.
\(a\)Unexpected Breaks\(b\)Unexploited Lifetime\(c\)Metric J
Figure 5:Business\-metric comparison for[S4D](https://arxiv.org/html/2609.22160#id26)across evaluation quantilesFigure[5](https://arxiv.org/html/2609.22160#S5.F5)isolates the effect of the evaluation quantile on the maintenance trade\-off for[S4D](https://arxiv.org/html/2609.22160#id26)\. For the higher quantiles, 0\.75 and 0\.9, the risk of[UB](https://arxiv.org/html/2609.22160#id16)initially remains high, but decreases sharply as the maintenance window is enlarged\. As expected, the[UL](https://arxiv.org/html/2609.22160#id17)increases with the maintenance window for every quantile, since interventions are scheduled earlier\. However, the 0\.75 and 0\.9 quantiles consistently yield lower[UL](https://arxiv.org/html/2609.22160#id17)than the lower quantiles\. This reduced[UL](https://arxiv.org/html/2609.22160#id17)is also reflected in theJJcurves, where these two quantiles achieve the lowest costs over the examined cost ratios, despite their higher initial[UB](https://arxiv.org/html/2609.22160#id16)risk\. Thus, for the cost scenarios considered here, the savings associated with avoiding excessively early maintenance outweigh the corresponding increase in breakage risk\.
## 6Conclusion
We presented a[PdM](https://arxiv.org/html/2609.22160#id9)framework for ion\-milling[RUL](https://arxiv.org/html/2609.22160#id8)estimation that combines[DL](https://arxiv.org/html/2609.22160#id3)sequence models and[SQR](https://arxiv.org/html/2609.22160#id13)to estimate conditional quantiles\. Quantile selection adjusts maintenance decisions to operational risk, while[UB](https://arxiv.org/html/2609.22160#id16),[UL](https://arxiv.org/html/2609.22160#id17), and cost\-weightedJJassess their business consequences and identify a failure\-utilisation trade\-off\. This solution is particularly useful in the contex of semiconductor manufacturing where taking risk\-informed decisions is of paramount importance to avoid single failures to disrupt the entire production pipeline\. On theF1F\_\{1\}[PHM18](https://arxiv.org/html/2609.22160#id15)fault mode,[SSM](https://arxiv.org/html/2609.22160#id12)architectures are the most memory and compute efficient and produced the most accurate[RUL](https://arxiv.org/html/2609.22160#id8)estimates, with[S4D](https://arxiv.org/html/2609.22160#id26)best overall across quantiles\. Quantile and maintenance\-window choices controlled the[PdM](https://arxiv.org/html/2609.22160#id9)trade\-off: larger windows reduced[UB](https://arxiv.org/html/2609.22160#id16)for[PdM](https://arxiv.org/html/2609.22160#id9)models while gradually increasing[UL](https://arxiv.org/html/2609.22160#id17)\. By contrast,[PvM](https://arxiv.org/html/2609.22160#id10)baselines incurred very large[UL](https://arxiv.org/html/2609.22160#id17)because their decisions used summary statistics of a strongly right\-skewed training\-life distribution\. TheirJJvalues were therefore much higher despite low breakage risk, showing that quantile\-based[RUL](https://arxiv.org/html/2609.22160#id8)estimation can support more efficient, risk\-aware policies\.
Future research may extend the approach beyond theF1F\_\{1\}fault mode by transferring knowledge learned from the available data to other failure types in[PHM18](https://arxiv.org/html/2609.22160#id15)\(i\.e\.,F2F\_\{2\}andF3F\_\{3\}\), for example through domain\-adaptation techniques\[[35](https://arxiv.org/html/2609.22160#bib.bib15)\]\. Continual learning strategies could also further enable the model to adapt to new fault conditions and evolving operating regimes without repeatedly retraining it from scratch\. Additional directions include evaluating the framework on larger and more diverse industrial datasets, calibrating the quantile estimates and maintenance\-cost parameters with plant\-specific operational data, and studying adaptive policies that jointly select the quantile level and maintenance window online\.
## References
- \[1\]P\. Osterrieder, L\. Budde, and T\. Friedli\(2020\)The smart factory as a key construct of industry 4\.0: a systematic literature review\.International Journal of Production Economics221,pp\. 107476\.External Links:ISSN 0925\-5273,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.ijpe.2019.08.011),[Link](https://www.sciencedirect.com/science/article/pii/S0925527319302865)Cited by:[§1](https://arxiv.org/html/2609.22160#S1.p1.1)\.
- \[2\]J\. Park, B\. Yoo, S\. Yi Baek, C\. Youn, S\. Kim, D\. Kim, S\. Roh, S\. Jun Park, J\. Kim, C\. Lee, and C\. Choi\(2025\)Advancing condition\-based maintenance in the semiconductor industry: innovations, challenges and future directions for predictive maintenance\.IEEE Transactions on Semiconductor Manufacturing38\(1\),pp\. 96–105\.External Links:[Document](https://dx.doi.org/10.1109/TSM.2025.3530964)Cited by:[§1](https://arxiv.org/html/2609.22160#S1.p1.1),[§2](https://arxiv.org/html/2609.22160#S2.p3.1)\.
- \[3\]K\. E\. Chong and K\. C\. Ng\(2016\)Relationship between overall equipment effectiveness, throughput and production part cost in semiconductor manufacturing industry\.In2016 IEEE International Conference on Industrial Engineering and Engineering Management \(IEEM\),Vol\.,pp\. 75–79\.External Links:[Document](https://dx.doi.org/10.1109/IEEM.2016.7797839)Cited by:[§1](https://arxiv.org/html/2609.22160#S1.p1.1)\.
- \[4\]S\. H\. Shah and I\. Yaqoob\(2016\)A survey: internet of things \(iot\) technologies, applications and challenges\.In2016 IEEE Smart Energy Grid Engineering \(SEGE\),Vol\.,pp\. 381–385\.External Links:[Document](https://dx.doi.org/10.1109/SEGE.2016.7589556)Cited by:[§1](https://arxiv.org/html/2609.22160#S1.p1.1)\.
- \[5\]D\. Mazzei and R\. Ramjattan\(2022\)Machine learning for industry 4\.0: a systematic review using deep learning\-based topic modelling\.Sensors22\(22\)\.External Links:[Link](https://www.mdpi.com/1424-8220/22/22/8641),ISSN 1424\-8220,[Document](https://dx.doi.org/10.3390/s22228641)Cited by:[§1](https://arxiv.org/html/2609.22160#S1.p1.1)\.
- \[6\]S\. Putteti, G\. Santhi, G\. R\. Mittoor, C\. Nagamani, and P\. Udayaraju\(2025\)Intelligent industrial iot: a data\-driven approach for smart manufacturing and predictive maintenance\.In2025 Third International Conference on Augmented Intelligence and Sustainable Systems \(ICAISS\),Vol\.,pp\. 1032–1040\.External Links:[Document](https://dx.doi.org/10.1109/ICAISS61471.2025.11041978)Cited by:[§1](https://arxiv.org/html/2609.22160#S1.p1.1)\.
- \[7\]K\. Patel, D\. Kolla, G\. Yadav, U\. Waghmode, S\. K\. Mannava, and A\. Narayanan\(2025\)IoT\-enabled predictive maintenance system for smart manufacturing plants\.In2025 10th International Conference on Communication and Electronics Systems \(ICCES\),Vol\.,pp\. 309–314\.External Links:[Document](https://dx.doi.org/10.1109/ICCES67310.2025.11336293)Cited by:[§1](https://arxiv.org/html/2609.22160#S1.p1.1)\.
- \[8\]N\. Tagasovska and D\. Lopez\-Paz\(2019\)Single\-model uncertainties for deep learning\.External Links:1811\.00908,[Link](https://arxiv.org/abs/1811.00908)Cited by:[§1](https://arxiv.org/html/2609.22160#S1.p2.1),[§3\.3](https://arxiv.org/html/2609.22160#S3.SS3.p1.1)\.
- \[9\]N\. Bolander, H\. Qiu, N\. Eklund, E\. Hindle, and T\. Rosenfeld\(2009\)Physics\-based remaining useful life prediction for aircraft engine bearing prognosis\.Annual Conference of the Prognostics and Health Management Society\.Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p2.1)\.
- \[10\]M\. Raissi, P\. Perdikaris, and G\.E\. Karniadakis\(2019\)Physics\-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations\.Journal of Computational Physics378,pp\. 686–707\.External Links:ISSN 0021\-9991,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jcp.2018.10.045),[Link](https://www.sciencedirect.com/science/article/pii/S0021999118307125)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p2.1)\.
- \[11\]M\. Arias Chao, C\. Kulkarni, K\. Goebel, and O\. Fink\(2022\)Fusing physics\-based and deep learning models for prognostics\.Reliability Engineering & System Safety217,pp\. 107961\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.ress.2021.107961),ISSN 0951\-8320,[Link](https://www.sciencedirect.com/science/article/pii/S0951832021004725)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p2.1)\.
- \[12\]J\. Sim, S\. Kim, H\. J\. Park, and J\. Choi\(2020\)A tutorial for feature engineering in the prognostics and health management of gears and bearings\.Applied Sciences10\(16\)\.External Links:[Link](https://www.mdpi.com/2076-3417/10/16/5639),ISSN 2076\-3417,[Document](https://dx.doi.org/10.3390/app10165639)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p2.1)\.
- \[13\]M\. Assafo and P\. Langendoerfer\(2025\)Tool remaining useful life prediction using feature extraction and machine learning\-based sensor fusion\.Results in Engineering28,pp\. 107297\.External Links:ISSN 2590\-1230,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.rineng.2025.107297),[Link](https://www.sciencedirect.com/science/article/pii/S2590123025033523)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p2.1)\.
- \[14\]L\. Lorenti, D\. D\. Pezze, J\. Andreoli, C\. Masiero, N\. Gentner, Y\. Yang, and G\. A\. Susto\(2023\)Predictive maintenance in the industry: a comparative study on deep learning\-based remaining useful life estimation\.In2023 IEEE 21st International Conference on Industrial Informatics \(INDIN\),Vol\.,pp\. 1–9\.External Links:[Document](https://dx.doi.org/10.1109/INDIN51400.2023.10218065)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p2.1),[§4\.3](https://arxiv.org/html/2609.22160#S4.SS3.p1.1)\.
- \[15\]S\. Vollert and A\. Theissler\(2021\)Challenges of machine learning\-based rul prognosis: a review on nasa’s c\-mapss data set\.In2021 26th IEEE International Conference on Emerging Technologies and Factory Automation \(ETFA \),Vol\.,pp\. 1–8\.External Links:[Document](https://dx.doi.org/10.1109/ETFA45728.2021.9613682)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p2.1)\.
- \[16\]J\. Wang, G\. Wen, S\. Yang, and Y\. Liu\(2018\)Remaining useful life estimation in prognostics using deep bidirectional lstm neural network\.In2018 Prognostics and System Health Management Conference \(PHM\-Chongqing\),Vol\.,pp\. 1037–1042\.External Links:[Document](https://dx.doi.org/10.1109/PHM-Chongqing.2018.00184)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p2.1)\.
- \[17\]V\. TV, P\. Gupta, P\. Malhotra, L\. Vig, and G\. Shroff\(2018\)Recurrent neural networks for online remaining useful life estimation in ion mill etching system\.Annual Conference of the PHM Society10\(1\)\.External Links:[Document](https://dx.doi.org/10.36001/phmconf.2018.v10i1.589)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p2.1),[§3\.1](https://arxiv.org/html/2609.22160#S3.SS1.p3.1),[§4\.2](https://arxiv.org/html/2609.22160#S4.SS2.p1.1)\.
- \[18\]J\. Gao, Y\. Wang, and Z\. Sun\(2024\)An interpretable rul prediction method of aircraft engines under complex operating conditions using spatio\-temporal features\.Measurement Science and Technology35\(7\),pp\. 076003\.External Links:[Document](https://dx.doi.org/10.1088/1361-6501/ad3b2c),[Link](https://dx.doi.org/10.1088/1361-6501/ad3b2c)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p2.1)\.
- \[19\]F\. Wang, A\. Liu, C\. Qu, R\. Xiong, and L\. Chen\(2025\)A deep\-learning method for remaining useful life prediction of power machinery via dual\-attention mechanism\.Sensors25\(2\)\.External Links:[Link](https://www.mdpi.com/1424-8220/25/2/497),ISSN 1424\-8220,[Document](https://dx.doi.org/10.3390/s25020497)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p2.1)\.
- \[20\]A\. C\. J\. Bonatakis and N\. Propes\.\(2018\)Phm data challenge 2018\.PHM Society\.External Links:[Link](https://c3.ndc.nasa.gov/dashlink/resources/139/)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p3.1)\.
- \[21\]C\. Hsu, Y\. Lu, and J\. Yan\(2022\)Temporal convolution\-based long\-short term memory network with attention mechanism for remaining useful life prediction\.IEEE Transactions on Semiconductor Manufacturing35\(2\),pp\. 220–228\.External Links:[Document](https://dx.doi.org/10.1109/TSM.2022.3164578)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p3.1)\.
- \[22\]C\. Liu, L\. Zhang, J\. Li, J\. Zheng, and C\. Wu\(2021\)Two\-stage transfer learning for fault prognosis of ion mill etching process\.IEEE Transactions on Semiconductor Manufacturing34\(2\),pp\. 185–193\.External Links:[Document](https://dx.doi.org/10.1109/TSM.2021.3059025)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p3.1),[§4\.1](https://arxiv.org/html/2609.22160#S4.SS1.SSS0.Px2.p3.1)\.
- \[23\]Z\. Yuan and R\. Wang\(2024\)Multi\-scale and multi\-branch transformer network for remaining useful life prediction in ion mill etching process\.IEEE Transactions on Semiconductor Manufacturing37\(1\),pp\. 67–75\.External Links:[Document](https://dx.doi.org/10.1109/TSM.2023.3324057)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p3.1),[§3\.1](https://arxiv.org/html/2609.22160#S3.SS1.p3.1),[§4\.2](https://arxiv.org/html/2609.22160#S4.SS2.p1.1)\.
- \[24\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.External Links:1706\.03762,[Link](https://arxiv.org/abs/1706.03762)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p3.1),[2nd item](https://arxiv.org/html/2609.22160#S3.I1.i2.p1.1)\.
- \[25\]C\. Zhao, S\. Xiang, S\. Hao, F\. Niu, and K\. Li\(2025\)Quantification of uncertainty information in remaining useful life estimation\.Applied Mathematical Modelling,pp\. 115992\.External Links:ISSN 0307\-904X,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.apm.2025.115992),[Link](https://www.sciencedirect.com/science/article/pii/S0307904X25000678)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p4.1)\.
- \[26\]L\. D\. Libera, J\. Andreoli, D\. D\. Pezze, M\. Ravanelli, and G\. A\. Susto\(2024\)Bayesian deep learning for remaining useful life estimation via stein variational gradient descent\.External Links:2402\.01098,[Link](https://arxiv.org/abs/2402.01098)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p4.1)\.
- \[27\]D\. Frizzo, F\. Borsatti, and G\. A\. Susto\(2025\)A quantile regression approach for remaining useful life estimation with state space models\.IFAC\-PapersOnLine59\(26\),pp\. 347–352\.Note:7th IFAC Conference on Intelligent Control and Automation Sciences ICONS 2025External Links:ISSN 2405\-8963,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.ifacol.2025.12.059),[Link](https://www.sciencedirect.com/science/article/pii/S2405896325027351)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p4.1),[§3\.1](https://arxiv.org/html/2609.22160#S3.SS1.p4.1),[§3\.2](https://arxiv.org/html/2609.22160#S3.SS2.p1.1)\.
- \[28\]M\. Zhang, D\. Wang, N\. Amaitik, and Y\. Xu\(2022\)A distributional perspective on remaining useful life prediction with deep learning and quantile regression\.IEEE Open Journal of Instrumentation and Measurement1\(\),pp\. 1–13\.External Links:[Document](https://dx.doi.org/10.1109/OJIM.2022.3205649)Cited by:[§2](https://arxiv.org/html/2609.22160#S2.p4.1)\.
- \[29\]M\. Rigamonti, P\. Baraldi, E\. Zio, I\. Roychoudhury, K\. Goebel, and S\. Poll\(2016\)Echo state network for the remaining useful life prediction of a turbofan engine\.PHM Society European Conference\.External Links:[Link](https://doi.org/10.36001/phme.2016.v3i1.1623)Cited by:[§3\.1](https://arxiv.org/html/2609.22160#S3.SS1.p3.1)\.
- \[30\]A\. Gu, K\. Goel, and C\. Ré\(2022\)Efficiently modeling long sequences with structured state spaces\.External Links:2111\.00396,[Link](https://arxiv.org/abs/2111.00396)Cited by:[2nd item](https://arxiv.org/html/2609.22160#S3.I1.i2.p1.1)\.
- \[31\]A\. Gupta, A\. Gu, and J\. Berant\(2022\)Diagonal state spaces are as effective as structured state spaces\.External Links:2203\.14343,[Link](https://arxiv.org/abs/2203.14343)Cited by:[2nd item](https://arxiv.org/html/2609.22160#S3.I1.i2.p1.1)\.
- \[32\]J\. T\. H\. Smith, A\. Warrington, and S\. W\. Linderman\(2023\)Simplified state space layers for sequence modeling\.External Links:2208\.04933,[Link](https://arxiv.org/abs/2208.04933)Cited by:[2nd item](https://arxiv.org/html/2609.22160#S3.I1.i2.p1.1)\.
- \[33\]S\. Umeda, K\. Tamaki, M\. Sumiya, and Y\. Kamaji\(2021\)Planned maintenance schedule update method for predictive maintenance of semiconductor plasma etcher\.IEEE Transactions on Semiconductor Manufacturing34\(3\),pp\. 296–300\.External Links:[Document](https://dx.doi.org/10.1109/TSM.2021.3071487)Cited by:[§4\.3](https://arxiv.org/html/2609.22160#S4.SS3.p1.1)\.
- \[34\]G\. A\. Susto, A\. Schirru, S\. Pampuri, S\. McLoone, and A\. Beghi\(2015\)Machine learning for predictive maintenance: a multiple classifier approach\.IEEE Transactions on Industrial Informatics11\(3\),pp\. 812–820\.External Links:[Document](https://dx.doi.org/10.1109/TII.2014.2349359)Cited by:[§4\.3](https://arxiv.org/html/2609.22160#S4.SS3.p1.1)\.
- \[35\]M\. Azamfar, X\. Li, and J\. Lee\(2020\)Deep learning\-based domain adaptation method for fault diagnosis in semiconductor manufacturing\.IEEE Transactions on Semiconductor Manufacturing33\(3\),pp\. 445–453\.External Links:[Document](https://dx.doi.org/10.1109/TSM.2020.2995548)Cited by:[§6](https://arxiv.org/html/2609.22160#S6.p2.1)\.相似文章
用于发动机健康管理与剩余使用寿命预测的科学机器学习
本文提出了一种用于涡轮机预测的多任务科学机器学习框架,该框架使用共享序列编码器和任务特定头,联合预测发动机健康指标和剩余使用寿命,并量化不确定性。
量子退火增强强化学习用于精确剩余使用寿命预测
本文提出了一种量子退火增强的Q-learning框架,用于剩余使用寿命预测,利用D-Wave系统求解QUBO公式以进行动作选择。在NASA C-MAPSS和预测维护数据集上,它优于经典和量子基线。
面向复杂系统中可解释预测性维护的语义特征分割
本文提出了一种用于预测性维护的语义特征分割框架,将监测信号分解为规范成分和残差成分,以提高可解释性,同时保持预测性能。
基于时间序列基础模型嵌入的剩余使用寿命估计
本文介绍了一种轻量级方法,利用Chronos-2时间序列基础模型的冻结嵌入,结合一个简单的回归头,进行剩余使用寿命估计,在工业传感器数据上相比基线方法取得了更优的性能。
针对联合故障诊断与剩余使用寿命估计的注意力增强多任务学习的泄漏鲁棒评估与数据规模敏感性
本文证明了在滑动窗口序列上使用简单的训练/测试分割会严重夸大或缩小预测性维护中多任务学习的性能指标,并提出了一种泄漏鲁棒的评估协议。