Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review
Summary
A systematic literature review of 212 studies investigates how Physics-Informed Machine Learning (PIML) is applied in Prognostics and Health Management (PHM), introducing a four-class classification scheme and finding that PIML consistently improves predictive performance over conventional baselines, though the literature is skewed toward batteries and bearings and lacks strong evidence for claims regarding generalization and interpretability.
View Cached Full Text
Cached at: 08/12/26, 08:27 AM
# Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review
Source: [https://arxiv.org/html/2608.10047](https://arxiv.org/html/2608.10047)
\[1,2\]\\fnmChristopher\\surBraun\\orcidhttps://orcid\.org/0009\-0006\-4153\-4772\\equalcontThese authors contributed equally to this work\.
\\equalcont
These authors contributed equally to this work\.
1\]\\orgdivInstitute of Industrial Manufacturing and Management IFF,\\orgnameUniversity of Stuttgart,\\orgaddress\\streetAllmandring 35,\\cityStuttgart,\\postcode70569,\\stateBaden\-Württemberg,\\countryGermany
2\]\\orgnameFraunhofer Institute for Manufacturing Engineering and Automation IPA,\\orgaddress\\streetNobelstraße 12,\\cityStuttgart,\\postcode70569,\\stateBaden\-Württemberg,\\countryGermany
###### Abstract
In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and Health Management \(PHM\)\. Machine Learning \(ML\) has largely driven advancements in diagnostics and prognostics, yet purely data\-driven models face inherent limitations, such as poor generalization, an inability to infer causal relationships, and a lack of interpretability\. Physics\-Informed Machine Learning \(PIML\) helps mitigate these limitations by incorporating prior physical knowledge directly into the ML pipeline, thereby fostering growing interest in its application to PHM\. This work investigates how PIML is being leveraged in the context of PHM through a systematic literature review of 212 studies\. The review introduces a four\-class classification scheme, consisting of observational bias, inductive bias, learning bias, and hybrid approaches, and further categorizes studies by PHM task\. Across all four classes, the reviewed studies consistently demonstrate improved predictive performance over conventional baselines across a broad range of assets, although the literature is heavily skewed toward lithium\-ion batteries and bearings, and dominated by problem\-specific solutions\. Overall, the review indicates that physics\-informed approaches already provide tangible benefits, whereas claims of improvements concerning some of the aforementioned limitations lack sufficient supporting evidence\. Future research should prioritize transferable design patterns, benchmarks comparing integration strategies, and uncertainty\-aware models that are lightweight and robust enough for online deployment in real\-world settings\.
###### keywords:
Systematic literature review, Physics\-informed machine learning, Prognostics and health management, Prior physical knowledge, Hybrid approaches
\{strip\}
The version of record of this article, first published in Journal of Intelligent Manufacturing, is available online at Publisher’s website:[https://dx\.doi\.org/10\.1007/s10845\-026\-02930\-3](https://dx.doi.org/10.1007/s10845-026-02930-3)\. This arXiv version is content\-equivalent to the version of record but differs in four respects: \(i\) the list of studies excluded after full\-text analysis, available as supplementary material \(referred to as Online Resource 1 in the version of record\), is included here as an appendix; \(ii\) citations are consistently disambiguated, resolving cases in which several distinct references share an identical in\-text citation string; \(iii\) section headings are numbered, so that the numbered cross\-references used throughout the text can be resolved; and \(iv\) all figures are embedded such that the text they contain remains selectable and searchable\. No claims, results, or conclusions have been altered\. Please cite the version of record\.
Abbreviations \(Technical Terms\)
AEAutoencoderBiLSTMBidirectional Long Short\-Term MemoryBPFIBall Pass Frequency Inner RaceBPFOBall Pass Frequency Outer RaceBSFBall Spin FrequencyCAEConvolutional AutoencoderCMCondition MonitoringCNNConvolutional Neural NetworkCWTContinuous Wavelet TransformDLDeep LearningDOFDegrees of FreedomECAEfficient Channel AttentionECMEquivalent Circuit ModelEKFExtended Kalman FilterFACFrequency\-Aware ConvolutionFEFinite ElementFLOPFloating Point OperationGANGenerative Adversarial NetworkGNNGraph Neural NetworkGPGaussian ProcessGPRGaussian Process RegressionGRUGated Recurrent UnitHIHealth IndexIMLInformed Machine LearningkNNK\-Nearest NeighborLSTMLong Short\-Term MemoryMLMachine LearningMLPMultilayer PerceptronNNNeural NetworkODEOrdinary Differential EquationPDEPartial Differential EquationPdMPredictive MaintenancePFParticle FilterPHMPrognostics and Health ManagementPIMLPhysics\-Informed Machine LearningPINNPhysics\-Informed Neural NetworkReLURectified Linear UnitResNetResidual NetworkRFRandom ForestRFRRandom Forest RegressionRULRemaining Useful LifeRLReinforcement LearningRNNRecurrent Neural NetworkSEISolid Electrolyte InterphaseSOCState of ChargeSOHState of HealthSPMSingle\-Particle ModelSVRSupport Vector RegressionSVMSupport Vector MachineTGDSTheory\-Guided Data ScienceTLTransfer LearningTRLTechnology Readiness Level
## 1Introduction
[Prognostics and Health Management](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\([PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\) has emerged as a cornerstone of modern industrial operations, driven by the increasing need for reliability, safety, and efficiency in complex engineering systems\[vogl2019AReviewOfDiagnostic\]\. Specifically,[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)provides the methodological foundation for[Predictive Maintenance](https://arxiv.org/html/2608.10047#p6.32.32.32.32)\([PdM](https://arxiv.org/html/2608.10047#p6.32.32.32.32)\), enabling organizations to transition from reactive or schedule\-based strategies through condition\-based monitoring to predictive measures, supporting informed decision\-making\[huang2024prognostics\]\. This transition has been accelerated by advances in sensing technologies, connectivity, and industrial digitalization—central pillars of the Industry 4\.0 paradigm\. As industrial assets become more interconnected and operational demands intensify,[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)plays a critical role in reducing unplanned downtime, optimizing maintenance costs, and ensuring continuous production\.
[Machine Learning](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\([ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\) has significantly influenced[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)research, with its adoption growing substantially in recent years\[sajjadi2025machine\]\.[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models have demonstrated strong performance in tasks such as anomaly detection\[cannizzaro2025MachineLearningEnabled\], fault diagnosis\[zhang2025JointDistributionDomain\], and[Remaining Useful Life](https://arxiv.org/html/2608.10047#p6.41.41.41.41)\([RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)\) prediction\[yin2025RemainingUsefulLife\]\. Despite these achievements, purely data\-driven approaches face inherent limitations, particularly when deployed in real\-world industrial settings\. In such settings, sensor measurements are frequently noisy, incomplete, or inconsistent, operational conditions vary widely across assets and environments, and failure data remain sparse due to the rarity of catastrophic events\. Moreover,[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models often lack interpretability and struggle to extrapolate beyond the state space covered by the training data\[hagmeyer2022integration\]\. These limitations are particularly critical in industrial sectors where reliability, safety, and trustworthiness are paramount, driving the need for more advanced modeling strategies\.
[Physics\-Informed Machine Learning](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\([PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\)\[karniadakis2021physics\]constitutes a promising paradigm for addressing these challenges by incorporating prior physical knowledge into the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)pipeline\. By combining the scalability and predictive capabilities of ML with the structure and interpretability of physics\-based modeling,[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)offers a pathway toward robust, generalizable, and physically consistent[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)solutions\. These characteristics align strongly with the stringent reliability and transparency demands of industrial settings, where decisions based on model outputs often carry significant operational or safety implications\[zio2022prognostics\]\.
Beyond methodological motivations, the industrial context itself amplifies the relevance of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\. Modern industrial systems operate under harsh, dynamic, and heterogeneous conditions\. Assets may experience variable loads, nonlinear degradation, or rapid transitions between operating regimes\. Sensor availability and quality can differ across machines and sites, and data sharing is frequently impeded by privacy and security concerns\. Fleet\-level variability introduces further complexity\. Models must generalize across equipment units that share design principles but exhibit different usage patterns or environmental exposures, often driving costly model reengineering\[zeng2025ApplicationofFrequency\]\. These practical challenges underscore the need for models that do not rely solely on data, but instead leverage physical insights to ensure robustness and adaptability—precisely the strengths that physics\-informed approaches promise to provide\. Notably, the growing adoption of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)extends beyond[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)into adjacent fields such as structural health monitoring\[rizvi2023data\], underscoring its cross\-domain relevance for systems subject to degradation\.
Despite the growing body of research at the intersection of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)and[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), the field remains fragmented across ways of integrating physics, across[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks, and across application domains\. Existing surveys tend to focus on specific assets, such as lithium\-ion batteries\[meng2019Areviewon\], gas turbines\[farhat2025physics\], or bridges\[mammeri2025traditional\]\. Moreover, they tend to adopt a more exploratory approach, providing detailed insights in specific areas while leaving some aspects less systematically addressed\. Given that[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)began attracting significant attention following the seminal work by\[karniadakis2021physics\], the field has been developing rapidly, making it challenging for recent surveys to fully capture the current state of research\. As a result, researchers lack a unified classification that enables systematic comparison of how physics is integrated across different[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks and application domains\. Practitioners, in turn, have no consolidated evidence base to guide the selection of an appropriate integration strategy for a given industrial setting\. This limits both the cumulative advancement of methods and their translation into operational practice\.
At its core, this work addresses the question of how[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)is being leveraged in the context of[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), and what the key challenges and opportunities are\. To answer this systematically and ensure both reproducibility and completeness, the most comprehensive systematic literature review of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)to date is conducted, covering 212 studies\. The outcome provides researchers, developers, and practitioners with the necessary insights to advance data\-driven[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)applications by effectively incorporating prior physical knowledge\. Specifically, the research questions this work seeks to answer are:
1. 1\.Knowledge\(a\) What types of prior physical knowledge are being leveraged, and \(b\) what forms of representation are employed?
2. 2\.Incorporation\(a\) How can prior physical knowledge be incorporated, and \(b\) how does the form of representation influence which approaches to incorporation are feasible?
3. 3\.Practice\(a\) How does incorporating prior physical knowledge help overcome limitations of purely data\-driven methods, and \(b\) what are the primary challenges in developing and applying physics\-informed approaches?
The remainder of this work is organized as follows\. Section[2](https://arxiv.org/html/2608.10047#S2)provides background on[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), physics\-informed learning, and a brief description of related work\. Section[3](https://arxiv.org/html/2608.10047#S3)describes the methodology underlying the systematic literature review, including the search strategy, as well as the screening and quality assessment procedures\. Section[4](https://arxiv.org/html/2608.10047#S4)presents the four\-class classification scheme, including observational bias, inductive bias, learning bias, and hybrid approaches\. Subsequently, the identified studies are summarized\. Section[5](https://arxiv.org/html/2608.10047#S5)analyzes methodological trends and discusses current challenges and opportunities for future advancements in industrial[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\-based[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. Finally, Section[6](https://arxiv.org/html/2608.10047#S6)concludes with a synthesis of key insights and implications, concisely answering the outlined research questions\.
## 2Background
The following section outlines the theoretical foundation necessary to understand the key concepts in this review\.[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)is introduced first, covering its core tasks, as well as various approaches to its implementation\. Among these, hybrid approaches are particularly promising, with[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)at the forefront\. This paradigm is defined and explored, along with related research fields\. Finally, an overview of related work is provided to position this review within the broader literature\.
### 2\.1Prognostics and Health Management
[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)is an engineering discipline that focuses on detecting, isolating, and diagnosing potential faults in a system, assessing its current[State of Health](https://arxiv.org/html/2608.10047#p6.46.46.46.46)\([SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)\), and predicting its[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)with the aim of preventing unplanned downtimes, enhancing reliability, and supporting additional system\-level objectives\. These diagnostic and prognostic measures play a crucial role in ensuring continuous operation by monitoring system health and predicting incipient faults before they progress into catastrophic failures\. Hence,[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)not only improves operational efficiency but also enhances the safety of the monitored system\. However, implementing[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)presents several challenges that vary depending on the chosen approach for modeling the system\. While the adoption of[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)in industrial settings holds significant potential, it also requires addressing various technical and organizational obstacles to fully realize its benefits\. With[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)encompassing both diagnostics and prognostics, four essential tasks are typically delineated for implementation\[hagmeyer2022integration,jia2018review\]:
1. 1\.Fault detectionis aimed at determining the presence or absence of a fault\.
2. 2\.Diagnosisis aimed at attributing observed faults to their root causes\.
3. 3\.Health assessmentis aimed at estimating the system’s current[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)or risk of failure\.
4. 4\.Prognosisis aimed at predicting the future development of the[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)or[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)\.
Table 1:Key strengths and limitations of physical model\-based and purely data\-driven approaches\[karniadakis2021physics,baur2020review\]\.To carry out these tasks, a range of methodological approaches is used, varying in the extent to which they rely on physical knowledge, data\-driven insights, or a combination of both\. These approaches are generally subdivided into three distinct categories\[kim2017prognostics,atamuradov2017prognostics\]:
- •Physical model\-based approaches
- •Purely data\-driven approaches
- •Hybrid approaches
According to\[gouriveau2016prognostics\], physical model\-based approaches“require the construction of a dynamic model representing the behavior of the system and integrating the degradation mechanism \(mainly by models of fatigue, wear, or corrosion\), whose evolution is modeled by a deterministic law or by a stochastic process\.”In contrast, purely data\-driven approaches \(including statistical and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)methods\) leverage monitoring data—either directly or via extracted features—to model a system’s behavior and health state\[goodman2019prognostics\]\. Hybrid approaches, which integrate aspects of the aforementioned methods, seek to leverage their strengths while mitigating their individual limitations \(see Tab\.[1](https://arxiv.org/html/2608.10047#S2.T1)\)\. Combining the accuracy and robustness of physical models with the flexibility and adaptability of data\-driven models enables enhancing the overall performance of the respective solution\. The aim is to capitalize on the synergies between these approaches, ultimately providing a more comprehensive and effective solution that can accommodate diverse operational scenarios\. Nonetheless, this does not preclude scenarios in which a physics\-based or data\-driven approach is preferable\.
Definitions of hybrid approaches may vary, however, particularly in terms of the extent to which physics is included\. Certain hybrid approaches employ complete physical models, while others rely on prior physical knowledge that may be insufficient for holistic modeling\. The choice of approach is generally dictated by the underlying physics of the problem, the availability of relevant data, and specific requirements imposed by the intended solution\. The continuous automation of modern industrial machinery leads to increasingly complex degradation processes, which are often poorly understood, dynamic, and highly nonlinear\[zio2022prognostics\]\. As a consequence, high\-fidelity physics\-based modeling is becoming increasingly difficult, if not impossible\. Accordingly, the integration of \(partial\) physical knowledge into data\-driven methods is gaining momentum\.
### 2\.2Physics\-Informed Machine Learning
The incorporation of prior knowledge alongside empirical data constitutes a central paradigm in[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)research \(see Fig\.[1](https://arxiv.org/html/2608.10047#S2.F1)\)\.\[vonrueden2021informed\]provide a general conceptualization of learning from such hybrid information sources, thereby establishing the foundational framework of[Informed Machine Learning](https://arxiv.org/html/2608.10047#p6.24.24.24.24)\([IML](https://arxiv.org/html/2608.10047#p6.24.24.24.24)\)\. According to this framework, prior knowledge is expected to originate from independent sources, be formally represented, and be explicitly integrated into the learning process\.\[karpatne2017theory\]propose a related framework with a specific focus on scientific knowledge, referred to as[Theory\-Guided Data Science](https://arxiv.org/html/2608.10047#p6.50.50.50.50)\([TGDS](https://arxiv.org/html/2608.10047#p6.50.50.50.50)\)\. Narrowing the focus further,\[karniadakis2021physics\]introduce[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)as a means to improve the modeling of physical systems by embedding prior physical knowledge directly into[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models\.
Figure 1:[IML](https://arxiv.org/html/2608.10047#p6.24.24.24.24)\[vonrueden2021informed\]provides a foundational framework for integrating various forms of prior knowledge into[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\. Within this context,[TGDS](https://arxiv.org/html/2608.10047#p6.50.50.50.50)\[karpatne2017theory\]focuses specifically on incorporating scientific knowledge, while[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\[karniadakis2021physics\]narrows the scope even further to leverage prior physical knowledge\.Instead of relying solely on data,[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)incorporates physics such as governing differential equations, conservation laws, or symmetries to ensure that model predictions remain physically consistent, which is especially valuable in regimes where data are sparse or noisy\[karniadakis2021physics\]\. This makes[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)particularly attractive in scientific and engineering domains, where high\-fidelity measurements can be expensive or limited by experimental feasibility\. Operationally,[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)can be implemented through a variety of complementary strategies, including data\-centric approaches, design\-level interventions, and regularization\-based constraints\. Yet developing[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)remains a nontrivial task, involving challenges such as balancing data fidelity with physics constraints, managing computational costs, and ensuring efficient training\[jahani2024enhancing\]\. Nonetheless,[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)has become a prominent paradigm for developing models that are both data\-driven and physically informed, balancing accuracy, interpretability, and robustness\.
Three main pathways have been defined for embedding physics into[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models, following the principles outlined by\[karniadakis2021physics\]\. Each is characterized by the introduction of an appropriate bias:
- •Observational biascan be introduced directly through the training data\.
- •Inductive biascan be introduced by tailored interventions to the model design\.
- •Learning biascan be introduced through modifications to the learning algorithm\.
The hypothesis space provides a useful lens for understanding how these biases affect learning\. It denotes the set of all functions a model could, in principle, choose to map inputs to outputs; for example, all functions realizable by a[Neural Network](https://arxiv.org/html/2608.10047#p6.29.29.29.29)\([NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)\) with a given architecture\. Observational bias, stemming from the training data, does not alter the space itself\. It merely biases the training process toward functions that better fit the observed data, leaving the set of representable functions intact\. Inductive bias, in contrast, acts directly on the model design and explicitly shapes the hypothesis space\. Embedding symmetry constraints, conservation laws, or other physical principles into the model restricts the hypothesis space to functions consistent with these principles, actively enforcing physical plausibility\. Learning bias exerts a more subtle influence: choices in optimization algorithms, regularization, or training strategies such as early stopping guide the search toward certain solutions, making some regions more likely to be explored while leaving the space itself unchanged\. In summary, observational bias influences the selection of hypotheses within the existing space \(not imposing constraints\), inductive bias modifies the structure of the space itself \(imposing hard constraints\), and learning bias guides the learning process toward particular regions, thereby effectively prioritizing certain solutions over others \(imposing soft constraints\)\.
### 2\.3Related Work
Extensive reviews have emerged around both[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\[tsui2015prognostics,atamuradov2017prognostics,hu2022prognostics,zio2022prognostics\]and[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\[karniadakis2021physics,cai2021physics,cuomo2022scientific\]individually\. However, while both fields have advanced significantly, the synergy between them is increasingly recognized as crucial for addressing the intricate challenges of effectively managing system health\. Accordingly, several reviews have sought to synthesize the state of research at the intersection of[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)and[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35), among which the following were known to the authors prior to conducting this systematic literature review:\[deng2023Physicsinformedmachinelearning\],\[kundu2020Areviewon\], and the review by\[meng2019Areviewon\]—all of which were also identified through the systematic approach employed in this work\. Hence, a detailed description is omitted here, since these \(along with eight additional reviews\) will be examined in detail in Section[4\.2](https://arxiv.org/html/2608.10047#S4.SS2)\. Furthermore, a broader perspective on[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)in the context of intelligent manufacturing is provided by\[leng2026physics\]\.
## 3Methodology
Table 2:Keywords used to identify the initial set of records\. A wildcard operator \(\*\) accounts for morphological variations\.The following section describes the methodology employed to conduct the systematic literature review on[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. Starting from the research questions formulated in Section[1](https://arxiv.org/html/2608.10047#S1), the process encompasses the selection of relevant keywords and databases, a rigorous screening procedure to identify all relevant studies, and a quality\-based refinement to distill the final selection\. By systematically analyzing the literature, this work provides a solid foundation for synthesizing current physics\-informed approaches to diagnostics and prognostics, identifying methodological gaps, and guiding future research directions\.
### 3\.1Keywords
The keywords used to identify the initial set of records are listed in Table[2](https://arxiv.org/html/2608.10047#S3.T2)\. In total, 105 unique terms were defined across four categories:prior physical knowledge,machine learning,prognostics and health management, andhealthcare\. The first three categories ensure relevance to the scope of this work\. The last category serves as an implicit filter to rule out irrelevant healthcare\-related records resulting from ambiguities in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\-related keywords\.
Relevant keywords were identified by examining the research questions’ core concepts and corresponding synonyms\. This also included screening prominent studies explicitly addressing the scope of this review for their author\-assigned keywords, employing large language models to generate additional keyword suggestions, and consulting peers to review and validate the keyword collection\. Furthermore, keywords were iteratively refined by performing exploratory searches in the selected databases, enabling the identification and exclusion of keywords of lower importance or higher ambiguity \(e\.g\.,“energy”or“force”within the category ofprior physical knowledge;“fault”or“failure”within the category ofprognostics and health management\)\. This approach ensured that the final keyword set was both precise and comprehensive\.
### 3\.2Databases
Scopus and Web of Science \(WoS\) were selected as the primary databases due to their comprehensive coverage of peer\-reviewed literature across various disciplines, ensuring a robust and thorough search\. Additionally, the preprint server arXiv was included to capture the most recent research developments\. This decision was made to provide a more accurate depiction of the current state of research, acknowledging that many cutting\-edge studies are first disseminated through preprints before formal publication, especially in the realm of[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\.
[TGDS](https://arxiv.org/html/2608.10047#p6.50.50.50.50), a precursor to[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35), was formally established around 2017 \(as stated in Sec\.[2\.2](https://arxiv.org/html/2608.10047#S2.SS2)\), marking a significant milestone in integrating scientific knowledge with[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)techniques\. However, the search covered the period from 2012 onward—a year widely recognized within the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)community as a pivotal moment due to the breakthrough results of AlexNet\[AlexNet\]in the ImageNet competition\. This choice reflects the consensus that 2012 represents the inception of modern[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)and[Deep Learning](https://arxiv.org/html/2608.10047#p6.10.10.10.10)\([DL](https://arxiv.org/html/2608.10047#p6.10.10.10.10)\)\[mienye2024comprehensive\]\. Yet the search on arXiv was limited to the period from 2023 to the present, under the assumption that high\-quality studies submitted to arXiv before 2023 would have already been published in peer\-reviewed journals or conference proceedings\.
The complete search string was constructed by connecting the keywords of the various categories with the appropriate logical operators \(see Tab\.[2](https://arxiv.org/html/2608.10047#S3.T2)\)\. The syntax for the logical operators was adapted to the specific requirements of each database\. The fields examined for relevant records included the title, abstract, and author\-assigned keywords\. For arXiv, which does not provide searching capabilities for author\-assigned keywords, only the title and abstract were considered\. All keywords were enclosed in quotation marks in order to ensure exact phrases in the search queries\. Moreover, both hyphenated and non\-hyphenated variants of keywords were automatically retrieved by the databases\.
### 3\.3Screening Procedure
The screening procedure was aided by ASReview\[asreview\], which is an[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\-based tool designed to streamline the systematic review process\. It leverages active learning algorithms to prioritize the most relevant records from a large corpus, significantly reducing the manual effort required for initial screening\. Records are presented iteratively, with their order continually updated based on reviewer feedback \(researcher\-in\-the\-loop\) to facilitate the identification of relevant records\. Additionally, ASReview supports bias\-free screening with respect to author names, affiliations, and other details by providing only the title and abstract of each record\.
ASReview was employed to manage the substantial number of records initially retrieved from the searched databases\. A structured approach to screening was implemented, consisting of the following steps:
1. 1\.Constructing a dataset containing the records to screen\.
2. 2\.Specifying inclusion and exclusion criteria to determine whether records are considered relevant\.
3. 3\.Defining a stopping criterion to determine at which point screening will conclude\.
4. 4\.Conducting screening in two cycles: 1. 4\.1Configuring the active learning algorithm and providing a subset of records for its initial training\. 2. 4\.2Performing the screening based on the specified criteria until the stopping criterion is met\.
5. 5\.Extracting the records identified as relevant to compile the subset of records for further analysis\.
The set of records initially retrieved from the designated databases constitutes the dataset, with all duplicates removed\. Only records containing both a title and an abstract are retained\. To determine the relevance of identified records, inclusion and exclusion criteria are specified\. The inclusion criterion requires that a record substantively addresses all three areas: prior physical knowledge,[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27), and[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. While the search string \(see Tab\.[2](https://arxiv.org/html/2608.10047#S3.T2)\) ensures that matching terms appear in the title, abstract, or author\-assigned keywords of every retrieved record, syntactic matching alone is insufficient to guarantee genuine topical relevance—records may, for instance, reference related concepts in a negating or merely peripheral context\. The reviewer therefore objectively assesses whether all three areas are adequately addressed in substance, rather than relying solely on keyword occurrence\. Regarding the exclusion criterion, a record is excluded if the studied system, component, material, or process is not investigated with respect to degradation during operation, as diagnostics and prognostics fundamentally rely on the asset’s current health state\. The scope of eligible assets further excludes transportation systems and their infrastructure, unmanned aerial vehicles, consumer electronics, and other technical systems used in residential settings\.
\[asreview\]state that, based on their simulation studies, reviewing 8–33 % of the total number of records using ASReview is typically sufficient to identify 95 % of the relevant records\. Based on these findings, the stopping criterion is defined as follows: at least 8 % of the records must be reviewed, and screening is terminated at the latest once 33 % have been screened\. Within these bounds, screening is considered complete when 50 consecutive irrelevant records are encountered—a heuristically set threshold\.
Screening comprises two cycles, with the first cycle leveraging a rather simple but fast configuration of the active learning algorithm\. The minimum requirement for labeled training data is to include at least one record labeled as relevant and one as irrelevant\. The aim during this cycle is to gather all relevant records that are more easily distinguishable from the rest of the dataset\. At a certain point, however, this configuration reaches its limitations, and inevitably, the stopping criterion is met\. Nevertheless, it can be assumed that the dataset still contains relevant records, though identifying them will require a more nuanced approach\. This necessitates a second cycle of screening, where a more sophisticated model is employed to capture all remaining relevant records that are more difficult to detect\. These records often elude initial screening due to subtle semantic nuances and complex contextual variations that a simpler model might overlook\. The result \(i\.e\., all labeled records\) of the first cycle serves as the initial training data for the more complex model\. Upon reaching the stopping criterion for the second time, it is assumed that a representative subset of relevant records has been identified\.
Two reviewers screened records independently in alternating one\-hour sessions\. Records of debatable relevance were discussed and decided upon through mutual consultation\. Periodic discussions ensured a shared understanding and consistent application of the criteria outlined earlier\.
### 3\.4Quality\-Based Refinement
Table 3:Conferences considered for the quality\-based refinement of conference papers \(in alphabetical order\)\.After the screening procedure is completed, the collection of identified records is refined by retaining only those that meet specified quality criteria\. With respect to journal articles, the quality\-based refinement depends on both the impact factor \(taken from the Journal Citation Reports provided by\[clarivate\]\) and the SCImago Journal Rank \(SJR\) \(provided by\[scimago\]\)\. The former is required to be 3 or greater, while the SJR is required to be Q2 or better\. If a journal is assigned to multiple categories within the SJR, it must maintain a minimum ranking of Q2 across all categories\. Both the impact factor and SJR are referenced according to the publication year of the article, with the most recent available values used when the corresponding year’s metrics are not yet released\. An article will still be considered if only one of the two metrics is available, provided the journal meets the required standard\. Articles from journals for which neither metric is available are excluded from the final selection\.
For conference papers, no established metric or ranking system with broad interdisciplinary applicability comparable to those used for journals exists\. Consequently, the quality\-based refinement for these records is applied differently\. Only papers presented at conferences recognized for their relevance and impact in the field of[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)are considered \(see Tab\.[3](https://arxiv.org/html/2608.10047#S3.T3)\)\.
Contributions from the preprint server arXiv cannot undergo quality\-based refinement\. Accordingly, arXiv papers are incorporated into the final selection of relevant records without additional assessment if they are identified as relevant during the screening procedure\. In contrast, other forms of publication, such as book chapters or technical reports, are excluded due to the lack of a reliable strategy for assessing their academic rigor and impact\.
### 3\.5Results
The methodology was applied twice \(referred to as two rounds\), with database searches conducted on each occasion to systematically gather relevant records and provide an up\-to\-date depiction of the state of research\. This approach revealed a notable increase in publications over time, reflecting heightened research activity and growing interest at the intersection of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)and[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. Among the sources, Scopus yielded the highest number of records, followed by WoS, while arXiv produced the fewest due to the more restricted time frame applied to that search\. While the period from January 2012 to August 27, 2024, yielded 6,586 records, the subsequent period from August 27, 2024, to May 4, 2025, produced 1,874 records\. This indicates a marked acceleration in contributions, with nearly 30 % of the previous 12\.5 years’ output occurring within this brief interval\. After retrieval, these records were preprocessed, resulting in a reduction to 3,956 and 1,026 records, respectively\. Thereafter, systematic screening and filtering yielded the final set, which formed the basis for analyzing the most current and relevant literature on[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\(see Fig\.[2](https://arxiv.org/html/2608.10047#S3.F2)\)\.
The two rounds comprised a total of three cycles, each utilizing a specifically configured active learning algorithm within the ASReview framework\. In the first cycle of round one, a comparatively simple yet computationally efficient configuration was adopted\. Feature extraction was performed usingterm frequency\-inverse document frequency, and a Naive Bayes classifier served as the underlying model for prioritizing records\. The query strategy and the balancing strategy were set tomixedanddynamic resampling, respectively\. Upon reaching the stopping criterion, a more sophisticated configuration was introduced for the second cycle\. Specifically, feature extraction was switched tosBERT, which produces contextually richer embeddings, and the classifier was replaced by an[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29), while the query and balancing strategies remained unchanged\. The second round of screening \(cycle three\) adopted the same advanced configuration as cycle two, leveraging all previously labeled records from the first round as initial training data\.
In preparation for the first cycle, 20 labeled records were provided for initialization of the active learning algorithm\. Although the minimum requirement is only two labeled records, this provided the model with a more informative starting point\. Records initially labeled as relevant were not guaranteed inclusion in the final review, as some may have been excluded during quality\-based refinement or full\-text analysis\. The initial labeled set comprised the following records:
- •Ten relevant records:\[sun2018AHybridApproach,ma2024Accurateandefficient,ellis2022Ahybridframework,chen2022PhysicsInformedLSTMhyperparameters,badora2023Usingphysicsinformedneural,yucesan2019Windturbinemain,garpelli2023Physicsguidedneuralnetworks,gareev2021Improvedfaultdiagnosis,gurgen2022Developmentandassessment,zhou2023PHYSICSINFORMEDMACHINELEARNING\]
- •Ten irrelevant records:\[gao2021sensing,he2024training,gong2022deep,pan2019aqlearning,sajedi2023twin,mcmahon2024river,chen2020augmenting,hu2017semi,zjavka2017nwp,wanasundara2023detecting\]
Figure 2:This flowchart illustrates the key stages of the methodology and the corresponding number of records at each stage for both rounds\.Two reviewers assessed a total of 1,189 \(30\.1 %\) and 339 \(33\.0 %\) records for relevance during the two rounds of screening, respectively \(see Tab\.[4](https://arxiv.org/html/2608.10047#S3.T4)\)\. Screening was performed in an alternating fashion, requiring 43 iterations overall\. During round one \(i\.e\., cycle one and two\), screening was terminated upon encountering 50 consecutive irrelevant records, in accordance with the stopping criterion, whereas round two concluded automatically upon reaching the maximum screening threshold of 33 %\. The difference in the total number of records screened per reviewer can be attributed to several factors, such as encountering longer abstracts, varying complexity in determining whether all three areas \(prior physical knowledge,[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27), and[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\) were substantively addressed, and varying cognitive load\. The screening yielded a total of 384 records across both rounds, a number considered excessive even for a thorough review\.
The application of the quality\-based refinement effectively reduced the number of records\. The final set comprised 212 records, including 180 journal articles, 18 conference papers \(ESREL, ICPHM, PHMSC, and RAMS\), and 14 preprints sourced from arXiv\. Accordingly, this work constitutes the most comprehensive review to date in this area of research\. All 212 records were subsequently subjected to full\-text analysis, after which 83 records were excluded due to insufficient alignment with the scope of this review\. All exclusions were transparently documented with their corresponding rationale \(see[Appendix](https://arxiv.org/html/2608.10047#Sx2)\)\. Hence, 129 records will be presented and discussed in this review\.
Table 4:This table provides a detailed overview of the screening results across all three cycles, where the first two authors correspond to reviewer A and B \(in no particular order\)\.
### 3\.6Limitations
Given that a reproducible systematic literature review depends on transparency, all potential limitations related to the outlined methodology are clearly articulated, mainly concerning the aided screening procedure and the subsequent quality\-based refinement\. Screening a large number of records required incorporating an additional tool—ASReview—aimed at facilitating the identification of those that are actually relevant\. Although ASReview relies on iterative reviewer feedback, semantic nuances may still pose challenges for the underlying active learning algorithms\. As a result, some relevant records may not have been surfaced by the algorithm before the stopping criterion was met\. This constitutes an inherent limitation of active learning\-based screening that cannot be fully eliminated without exhaustive manual review of the entire corpus of nearly 5,000 records\. However, the stopping criterion, defined in accordance with\[asreview\], inherently implies that approximately 5 % of relevant studies may remain unidentified by design\.
The decision to apply a quality\-based refinement was partly driven by the substantial number of records identified as relevant during screening\. The set of 384 records was considered impractical for full review, necessitating a reduction to a more manageable number\. Retaining only high\-quality studies not only reduced the number of records but also enhanced both the relevance of the findings and the reliability of the conclusions drawn\. Nonetheless, the thresholds for the impact factor and SJR were established based on the authors’ expert judgment, as was the selection of high\-impact conferences, which may be considered a limitation of the review\.
## 4Literature Review
The following section provides a synthesis of all studies included in the systematic literature review, classified according to their methodological approach to combining physics and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\. The classification scheme adopts and further refines the three pathways outlined by\[karniadakis2021physics\], imposing a more stringent framework\. Additionally, hybrid approaches are recognized as a distinct class, with the rationale discussed in detail below\. After outlining the classification scheme, a summary of the included review studies is presented, emphasizing the motivation for conducting the current review\. This is followed by a detailed depiction of the current state of research based on the remaining studies\.
### 4\.1Classification
Figure 3:The classification scheme underlying this review comprises four classes, with the first three classes representing genuine physics\-informed approaches and the fourth capturing hybrid approaches\.Approaches to implementing[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)are generally subdivided into model\-based, data\-driven, and hybrid approaches \(see Sec\.[2\.1](https://arxiv.org/html/2608.10047#S2.SS1)\)\. In this context,[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)is generally regarded as a paradigm within hybrid approaches\. Nonetheless, a key observation motivated distinguishing[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)from hybrid approaches, resulting in a classification scheme comprising four classes: observational bias, inductive bias, learning bias, and hybrid approaches \(see Fig\.[3](https://arxiv.org/html/2608.10047#S4.F3)\)\. While all four approaches involve physics and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27), the last category is conceptually distinct with respect to the role that physics plays within the overall model\. Crucially, it departs from the core principle of[IML](https://arxiv.org/html/2608.10047#p6.24.24.24.24)that prior knowledge is explicitly integrated into the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)pipeline\[vonrueden2021informed\], since hybrid approaches couple two independent models\. In light of this, studies in the fourth class could technically be considered outside the scope of this review\. However, to foster understanding of these conceptual differences and to provide a comprehensive overview of the intersection between physics and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), this class has been deliberately included\. In doing so, the distinct role of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)within the broader landscape becomes more evident\.
The first three classes are based on the well\-established three pathways for[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\. While inductive and learning bias are defined sufficiently to allow accurate classification, observational bias requires a stricter interpretation to avoid conflating physics\-informed approaches with conventional[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\. In the original work,\[karniadakis2021physics\]describe two ways of introducing observational bias, namely“through data that embody the underlying physics or carefully crafted data augmentation procedures\.”While the former applies to virtually any[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)problem involving real\-world systems whose behavior follows physical principles, the latter provides little practical guidance, as“carefully crafted”is not further defined\. To resolve the former ambiguity, the definition of prior knowledge as formulated by\[vonrueden2021informed\]is adopted, which specifies that it“exist\[s\] in an external, separated way from the learning problem and the usual training data\.”Hence, observational bias is not introduced by the mere use of empirical data obtained from physical systems\. As a result, the inclusion of additional \(physics\-informed\) data becomes obligatory to incorporate this form of bias\. Primary approaches include the use of simulated data, either to enrich the training data or as the source domain data within a[Transfer Learning](https://arxiv.org/html/2608.10047#p6.51.51.51.51)\([TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)\) setting—both strategies essentially addressing data scarcity\. This also implies, however, that training \(and testing\) exclusively on simulated data does not qualify as[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35), as it fails to meet the requirement that prior knowledge is separate from the training data\.
The ambiguity surrounding carefully crafted data augmentation procedures is resolved as follows\. On the one hand, it is virtually impossible to determine what qualifies as carefully crafted, which is why conventional data augmentation techniques are categorically excluded from the scope of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\. By the same reasoning, feature engineering \(regardless of its complexity\) does not constitute[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)either\. Ultimately, such measures remain conventional[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)in that they operate on or derive from existing training data\. On the other hand, the integration of entire physics\-based models in conjunction with an[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)model forms the basis of the fourth class\. While some studies employing a hybrid approach could, in principle, have been classified under observational bias according to the broad definition given by\[karniadakis2021physics\], it was established earlier that they violate the requirement of explicit integration of prior knowledge into the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)pipeline\. Instead, the physics\-based and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)components interact largely independently\.
Figure 4:This Sankey diagram shows three categories—approach,[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)taskandasset—each containing multiple items, with connections between them representing the state of research based on the 118 studies discussed in Sections[4\.3](https://arxiv.org/html/2608.10047#S4.SS3)–[4\.6](https://arxiv.org/html/2608.10047#S4.SS6)\. The width of each path connecting two items is proportional to the number of studies supporting that link, reflecting the relative strength of evidence\. To highlight areas of focus, paths supported by five or more studies across all three categories \(i\.e\., studies that employ the sameapproach, target the samePHM taskand address the sameasset\) are shown in darker gray\. Within theassetcategory, only those studied by five or more contributions are represented as individual items, while all remaining assets are aggregated under“Other”to preserve visual clarity\. Moreover, the number of studies corresponding to each item is indicated in brackets\. As some studies contribute to multiple items within a category \(e\.g\., by employing different approaches\), category totals may exceed 118\.Based on a thorough examination of studies in this fourth class, two types of approaches emerged, referred to as in\-parallel and in\-series, respectively\. Drawing on terminology from electrical engineering, these labels aptly capture the relationships between the physics\-based and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models\. Primary in\-parallel approaches include residual learning and ensemble methods\. In\-series approaches involve integrating the physics\-based and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models sequentially, so that the output of one model directly informs the input of the other\. Conceptually, this sequential integration can occur in two variants, with the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)model directing the physics\-based model, or vice versa\. Importantly, these hybrid approaches differ in two key aspects with respect to observational bias and feature engineering, respectively\. Unlike approaches that incorporate observational bias, in which physics\-based simulators are used only offline and are no longer required once the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)model is trained, hybrid approaches retain a physics\-based model as an integral component of the prediction pipeline during both training and inference\. Additionally, they go beyond feature engineering, where fixed algebraic relations are used to transform input variables but do not constitute a model with meaningful physical parameters\. Hence, studies following a hybrid approach involve a physics\-based component corresponding to a parameterized model, combined either in parallel or in series with the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)model\. In addition, a clear distinction is drawn with respect to inductive\-bias approaches that appear to employ an in\-series coupling\. If a physics\-based model is inserted as a fixed, differentiable module within the computational graph \(i\.e\., the model output is passed through the physics\-based model prior to computing the loss, and gradients are propagated through this module\), it is referred to as being a part of the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)model\. This links back to the hypothesis space: all admissible predictions are, by construction, consistent with the employed physics module\.
All studies \(excluding reviews\) are classified and presented following the classification scheme outlined above\. Each class is further subdivided by the four[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks: fault detection, diagnosis, health assessment, and prognosis \(see Sec\.[2\.1](https://arxiv.org/html/2608.10047#S2.SS1)\)\. Within these task\-specific sections, studies are grouped and reported according to related use cases, where possible\. Studies that incorporate multiple mechanisms for embedding physics into[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27), or that target multiple[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks, are addressed as follows\. If mechanisms or tasks are clearly separable, they are reported individually and may therefore appear multiple times across the following sections\. Studies with tightly coupled mechanisms are assigned to the class that best reflects the primary driving mechanism\. The same logic applies to[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks\. For each class of approaches, a representative example is presented in a dedicated figure, highlighting the prior knowledge employed and its incorporation\. The schematics illustrating each method are adapted from the corresponding study and simplified to highlight the essential information, where blue denotes the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)components and orange the physics components\. Lastly, the Sankey diagram in Figure[4](https://arxiv.org/html/2608.10047#S4.F4)offers a visual overview, capturing overall trends regarding the approaches employed for specific[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks across various assets, thereby offering context for subsequent analysis of methodological patterns\.
### 4\.2Reviews
The relevant literature comprises eleven review studies published between 2019 and 2025, exploring how the integration of physics into data\-driven methods has advanced the broader field of[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. These reviews offer valuable insights into the field’s evolution and present unique perspectives\. The following summaries are presented in order of relevance to the current work, highlighting each review’s contributions, methodologies, and individual strengths and limitations\.
Table 5:Overview of all identified reviews, listed in order of relevance\. Dashes indicate omitted information\.\[deng2023Physicsinformedmachinelearning\]present a broad overview of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), without focusing on any specific domain or task\. The review is well\-motivated, highlighting the advantages of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)over physical model\-based and purely data\-driven methods, respectively\. While the methodology is the most detailed among the reviews discussed here, it lacks specifics on the screening procedure, only stating that it was done manually without explaining the criteria used to differentiate relevant from irrelevant studies\. The identified studies are organized in a manner reminiscent of the three pathways outlined by\[karniadakis2021physics\], albeit using different terminology\. However, a significant concern is that, despite identifying 122 relevant studies, only 69 are categorized into the three approaches, leaving 53 unaddressed without any explanation from the authors\.
The comprehensive review by\[wu2024Physicsinformedmachinelearning\]focuses on the application of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)to anomaly detection and condition monitoring, but fails to establish a connection to[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. With the methodology only briefly outlined, key details on screening and selection are missing\. Supported by informative figures, the various strategies for integrating physics into ML provide a clear structure for the review\. Overall, it offers a valuable contribution with a thorough analysis of recent developments, though the level of detail can be excessive\.
\[li2024Areviewon\]thoroughly review methods for predicting[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41), focusing on[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)while also identifying and discussing the fusion of physics\-based and data\-driven models and the development of stochastic degradation models\. The review demonstrates technical depth and includes striking figures to clarify various approaches, but the repeated subdivision of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)methods with inconsistent terminology impedes comprehension\. Although the methodology for identifying relevant literature is briefly outlined, the authors themselves acknowledge that an exhaustive collection of studies remains elusive\.
\[khan2024Areviewof\]review physics\-based learning for system health management but fail to establish a clear connection to[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. Although the study provides a transparent background, the methodology is vaguely described, offering only examples of databases and keywords and failing to detail the screening procedure\. The authors reference\[karniadakis2021physics\]when introducing[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35), yet merely distinguish between physics\-based loss functions and“various architectures”—a category that lacks clear definition and fails to meaningfully differentiate between the diverse approaches emerging in the field\. Rather than offering a comprehensive analysis of the identified literature, the review discusses some[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)approaches in general terms, with only a limited selection of examples provided\.
The review by\[yan2025KnowledgeDrivenMachine\]sets out to introduce a“universal concept, knowledge driven machine learning, for integrating diverse knowledge into machine learning pipeline \[sic\] in PHM domain\.”However, it largely reproduces the[IML](https://arxiv.org/html/2608.10047#p6.24.24.24.24)taxonomy by\[vonrueden2021informed\]without demonstrable novelty\. While the authors attempt to synthesize their findings systematically and provide tangible case studies, the article’s contribution is weakened by conceptual ambiguity, methodological omissions, and linguistic inaccuracies, limiting its value as a rigorous or original synthesis of the field\.
\[fassi2024TowardPhysicsInformedMachineLearningBased\]provide a comprehensive overview of[PdM](https://arxiv.org/html/2608.10047#p6.32.32.32.32)for power converters, structured around model\-based, data\-driven, and hybrid approaches, with notable emphasis on[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\. However, the review overlooks[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)as the foundational framework for[PdM](https://arxiv.org/html/2608.10047#p6.32.32.32.32)and follows a narrative rather than systematic methodology\. Moreover, the discussion of studies employing[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)partly diverges from power converters, drawing substantially on adjacent domains\. Lastly, the absence of an in\-depth discussion of the reviewed studies limits broader insights, leading to a brief and generic outlook on future work\.
\[zhu2023Physicsinformedmachinelearning\]provide a brief review of the application of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)in structural integrity, including failure mechanism modeling and[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. However, the review’s completeness is undermined by the lack of a reported methodology\. Instead of offering a critical analysis, it mainly reports the identified studies, missing opportunities to extract valuable insights, such as the prior physical knowledge used, resulting in a weak foundation for discussing current challenges\.
The overview of[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\-based battery safety by\[zhao2024Batterysafety:Machine\]serves as a solid entry point, outlining the core mechanisms driving battery faults and failures\. Yet it provides no information on the methodology, rendering the review narrative rather than systematic and leaving its completeness uncertain\. Despite the title emphasizing prognostics, studies spanning the entire[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)spectrum are covered\. With respect to[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35), the structure is somewhat diffuse:[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)in conjunction with battery models is first presented \(mainly as a source of domain\-specific features\), yet several of these approaches effectively fall under[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35), with a later section specifically dedicated to[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)further blurring conceptual boundaries\.
The review by\[cuesta2025Areviewof\]provides a comprehensive overview of[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)in wind energy, focusing on[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)estimation for key turbine components such as gearboxes, generators, blades, and bearings\. It outlines three main degradation modeling approaches: physics\-based, data\-driven, and hybrid models\. In discussing hybrid models, the authors emphasize the integration of physical knowledge into[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)frameworks but do not explicitly situate this discussion within the emerging field of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\. The review further identifies challenges related to uncertainty quantification, integration of physical knowledge, environmental variability, and system complexity\. However, the absence of a clear methodology for study selection raises concerns about the completeness and representativeness of the reviewed literature\.
\[kundu2020Areviewon\]present a review of diagnostic and prognostic approaches for gears\. Offering both technical depth and clarity, this review is an excellent starting point for researchers and practitioners developing or applying[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)in this area\. However, the review lacks information on the process used to gather the relevant literature\. Additionally, the authors report only a few studies on hybrid approaches, explaining that few such methods have been developed, and attempt to cover all major approaches up to and including 2020—none of which employ[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\.
The review on[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)of lithium\-ion batteries by\[meng2019Areviewon\]is structured into physics\-based, data\-driven, and hybrid approaches, with the latter focusing primarily on[Particle Filter](https://arxiv.org/html/2608.10047#p6.33.33.33.33)\([PF](https://arxiv.org/html/2608.10047#p6.33.33.33.33)\)\-based and Kalman filter\-based methods\. Given the lack of methodological rigor, the work constitutes a descriptive survey rather than a systematic analysis\. Moreover, the absence of studies employing[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)as a tool for[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), likely due to the review’s 2019 publication date, underscores the rapid progress of the field\.
While acknowledging the significant contributions of previous reviews at the intersection of[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)and[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35), it is concluded that a more comprehensive and systematic review is needed to fully address the research questions outlined in Section[1](https://arxiv.org/html/2608.10047#S1)\. Although these reviews provide valuable insights, they often offer only partial coverage, focusing on particular domains, e\.g\., mechanical components\[kundu2020Areviewon\], batteries\[meng2019Areviewon,zhao2024Batterysafety:Machine\], power converters\[fassi2024TowardPhysicsInformedMachineLearningBased\], or wind energy\[cuesta2025Areviewof\], or specific[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks, such as[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction\[li2024Areviewon\], and in certain instances, extending more broadly to related areas, such as condition monitoring\[wu2024Physicsinformedmachinelearning\]or[PdM](https://arxiv.org/html/2608.10047#p6.32.32.32.32)\[fassi2024TowardPhysicsInformedMachineLearningBased\]\. Moreover, in some cases,[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)is not the central focus but is instead addressed only as a peripheral topic\. Furthermore, a recurring issue across all reviews is the lack of a systematic approach and detailed descriptions, which undermines scientific rigor, as transparency and reproducibility are essential to research integrity\. Notably, most reviews fail to report the number of included studies or the time period they cover \(see Tab\.[5](https://arxiv.org/html/2608.10047#S4.T5)\)\. In response, this review seeks to bridge existing gaps and advance the understanding and application of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)by providing a comprehensive, focused exploration of the state of the art, guided by a rigorous methodology\.
### 4\.3Observational Bias
Observational bias is characterized by introducing prior physical knowledge directly through the training data\. The identified studies are organized by PHM task in Tables[6](https://arxiv.org/html/2608.10047#S4.T6)–[8](https://arxiv.org/html/2608.10047#S4.T8)\. A representative example is illustrated in Figure[5](https://arxiv.org/html/2608.10047#S4.F5)\.
Figure 5:A representative example of observational\-bias approaches is proposed by\[dong2022Anewdynamic\], demonstrating how prior physical knowledge can be introduced as additional training data\. Own illustration based on the corresponding study\. See the original work for full technical details\.#### 4\.3\.1Fault Detection
No studies have been identified that introduce observational bias for the purpose of fault detection\.
#### 4\.3\.2Diagnosis
Table 6:All studies employingobservational biasto addressdiagnosis, listed in alphabetical order\.Observational Bias for DiagnosisReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[dong2022Anewdynamic\]4\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)dynamic bearing model[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)systemBearing\[li2024Asimulationdatadriven\]Numerical simulation of bearing fault vibration, based on impulse\-response signal modelAlgebraic equationBearing\[liu2024Enhancingmultitypefault\]Discretized first\-order[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13)Recursive algebraic equationLithium\-ion battery\[ma2025Aphysicsbasedsample\]Multi\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)rotor\-bearing dynamic modelCoupled nonlinear[ODEs](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Bearing\[pettorossi2024AddressingDataScarcity\]Proton exchange membrane fuel cell simulation[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)systemFuel cell\[qin2024Inversephysicsinformed\]2\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)dynamic bearing model[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)systemBearing\[qin2025SimulationdataDrivenGeneralized\]2\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)dynamic bearing model[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)systemBearing\[song2022Researchonfault\]Rigid\-flexible multibody dynamics simulation of a planetary gearboxDifferential\-algebraic equation systemPlanetary gearbox\[zhang2024AdversarialDomainAdaptation\]System level axial piston pump model \(rotor, bearings, fluid\), bearing dynamic model[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)systemAxial piston pump, bearing\[zhu2025DataGenerationApproach\]4\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)dynamic bearing model[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)systemBearingA total of ten studies employing observational bias for diagnosis have been identified \(see Tab\.[6](https://arxiv.org/html/2608.10047#S4.T6)\)\.\[liu2024Enhancingmultitypefault\]aim to solve the problem of high similarity among different faults in lithium\-ion battery voltage signatures\. Considering faults, such as internal short circuits, capacity anomaly, and[State of Charge](https://arxiv.org/html/2608.10047#p6.45.45.45.45)\([SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45)\) anomaly occurring in a series\-connected battery system, the authors calculate fault features by applying a mean difference model to terminal voltages\. These serve as input to a vision Transformer, followed by a[Multilayer Perceptron](https://arxiv.org/html/2608.10047#p6.28.28.28.28)\([MLP](https://arxiv.org/html/2608.10047#p6.28.28.28.28)\) head for fault diagnosis\. They use a discretized first\-order[Equivalent Circuit Model](https://arxiv.org/html/2608.10047#p6.13.13.13.13)\([ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13)\) to generate simulation data for pretraining the model, while fine\-tuning is performed on experimental data, leading to improved accuracy\.
\[qin2024Inversephysicsinformed\]focus on the challenge of accurately diagnosing bearing faults given imbalanced fault samples\. An inverse[Physics\-Informed Neural Network](https://arxiv.org/html/2608.10047#p6.36.36.36.36)\([PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)\) is employed to estimate the parameters of a 2\-[Degrees of Freedom](https://arxiv.org/html/2608.10047#p6.11.11.11.11)\([DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)\) dynamic bearing model, specifically stiffness and damping ratio, from real vibration data\. This enables the generation of simulated fault data whose frequency\-domain characteristics closely match those of real measurements\. The simulated data are then used to supplement imbalanced datasets, significantly enhancing the accuracy of a[Convolutional Neural Network](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\([CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\)\-based diagnostic model for bearing fault detection\. In another work,\[qin2025SimulationdataDrivenGeneralized\]address the challenge of unseen compound faults in bearing fault diagnosis\. The approach employs a composed model consisting of a[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\-based feature extractor, a Cycle\-[Generative Adversarial Network](https://arxiv.org/html/2608.10047#p6.18.18.18.18)\([GAN](https://arxiv.org/html/2608.10047#p6.18.18.18.18)\)\-based semantic mapping model \(trained on simulated data\) and a Cycle\-[GAN](https://arxiv.org/html/2608.10047#p6.18.18.18.18)\-based feature generator \(trained on measured single\-fault data\), and a multi\-agent deep[Reinforcement Learning](https://arxiv.org/html/2608.10047#p6.42.42.42.42)\([RL](https://arxiv.org/html/2608.10047#p6.42.42.42.42)\) model for the final diagnosis\. Simulated vibration signals for various single and compound fault types are generated with a 2\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)bearing dynamic model\. These signals are employed in conjunction with their respective envelope spectra as semantics to train the semantic mapping model\. Hence, zero\-shot learning regarding real\-world compound faults can be performed\. The proposed approach demonstrates superior performance in comparison to other state\-of\-the\-art methods for three bearing datasets\.\[dong2022Anewdynamic\]present a fault diagnosis framework for bearing race faults that addresses the small\-sample problem using a dynamic model and[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)strategies \(see Fig\.[5](https://arxiv.org/html/2608.10047#S4.F5)\)\. Simulated data generated from a 4\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)bearing dynamic model are used to pretrain a[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8), and parameter transfer strategies \(i\.e\., selectively freezing or fine\-tuning different parts of the network\) adapt the model to limited real\-world data\. By aligning simulation and real\-world conditions, the approach enables effective feature transfer and reduces distribution mismatch\. Compared to conventional[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models \(i\.e\.,[Support Vector Machine](https://arxiv.org/html/2608.10047#p6.49.49.49.49)\([SVM](https://arxiv.org/html/2608.10047#p6.49.49.49.49)\),[K\-Nearest Neighbor](https://arxiv.org/html/2608.10047#p6.25.25.25.25)\([kNN](https://arxiv.org/html/2608.10047#p6.25.25.25.25)\),[Random Forest](https://arxiv.org/html/2608.10047#p6.39.39.39.39)\([RF](https://arxiv.org/html/2608.10047#p6.39.39.39.39)\), and[MLP](https://arxiv.org/html/2608.10047#p6.28.28.28.28)\), both with and without simulation\-based[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51), the proposed[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\-based[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)approach consistently outperforms all alternatives in bearing fault diagnosis under small\-sample conditions\.\[li2024Asimulationdatadriven\]address the challenge of bearing fault diagnosis with unlabeled data\. Both completely unlabeled real data and labeled simulation data from a bearing failure simulation are combined in a semi\-supervised method\. The simulation covers time\-domain vibration signals for faults occurring in the outer race, inner race, and rolling elements\. Specifically, a multi\-kernel[kNN](https://arxiv.org/html/2608.10047#p6.25.25.25.25)graph connecting both data sources is built and edge weights are refined with both layer attention and dot\-product attention\. The resulting node embeddings are fed into fully connected layers for fault classification\. The proposed approach outperforms advanced[Graph Neural Network](https://arxiv.org/html/2608.10047#p6.19.19.19.19)\([GNN](https://arxiv.org/html/2608.10047#p6.19.19.19.19)\)\-based methods\. The domain gap problem regarding simulation and real data in bearing small\-sample fault diagnosis is studied by\[zhu2025DataGenerationApproach\]\. A bearing dynamic model is used to generate time\-series vibration data and corresponding fault labels, which are used as inputs for the generator of a conditional deep convolutional[GAN](https://arxiv.org/html/2608.10047#p6.18.18.18.18)instead of random noise\. The discriminator is trained to distinguish real vibration data from[GAN](https://arxiv.org/html/2608.10047#p6.18.18.18.18)\-generated data, encouraging the generator to produce synthetic signals that both respect the underlying fault mechanism and closely match the distribution of real measurements\. This generated data can then be used to train a fault diagnostic model alongside real data\. The presented method achieves higher accuracy and lower variance compared to a scenario in which the original simulation and real data are used directly to train the fault diagnostic model\.\[ma2025Aphysicsbasedsample\]address few\-shot bearing fault diagnosis by using a nonlinear rotor\-bearing system model to simulate mechanistic vibration signals for different fault types and severities\. These signals are combined with equipment\-specific characteristics extracted from healthy data, denoised via wavelet packet decomposition, and further refined using a cosine\-similarity\-guided noise injection, yielding a mixed dataset that comprises simulated and scarce real fault samples\. Across multiple[DL](https://arxiv.org/html/2608.10047#p6.10.10.10.10)architectures, training on the mixed dataset results in superior diagnostic accuracy compared to training solely on measured, simulated, or[GAN](https://arxiv.org/html/2608.10047#p6.18.18.18.18)\-generated data\.
The small\-sample problem, resulting from the limited availability of fault data for mechanical components, is addressed by\[zhang2024AdversarialDomainAdaptation\]through the use of simulation models\. Time\-series data from the simulation are used as the source domain, along with unlabeled experimental data as the target domain, to train an adversarial domain adaptation model with a feature extractor, healthy pattern recognizer, and domain discriminator\. Experiments on both axial piston pump and bearing datasets demonstrate that the proposed method significantly enhances fault\-diagnosis accuracy and generalization under small\-sample conditions, outperforming other domain\-adaptation methods\.
\[song2022Researchonfault\]present a method for planetary gearbox fault diagnosis based on[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)\. In this study, gearbox dynamics are simulated to generate labeled training data, addressing the challenge of limited real\-world fault samples\. A deep[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)framework is then applied to generalize the fault detection model across different operating conditions\. The framework employs a[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\-based feature extractor with shared parameters for both simulated source domain data and experimental target domain data, as well as a classifier\. Besides the classification loss, a domain adaptation loss is employed to guide domain\-invariant features\. Experimental results for different[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)tasks demonstrate the effectiveness of this approach in identifying cracked gear and missing tooth faults\. The proposed method performs slightly better than comparable[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)baselines\.
Both the small\-sample problem present in proton exchange membrane fuel cell fault diagnosis and the distribution mismatch between simulated and real data are studied by\[pettorossi2024AddressingDataScarcity\]\. They use a supervised domain\-adversarial adaptation model, combining a[Long Short\-Term Memory](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\([LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\)\-based feature extractor, a fully connected fault classifier, and a gradient\-reversal\-based domain discriminator\. A simulated dataset produced by a calibrated physics\-based proton exchange membrane fuel cell model is used as the source domain, whereas measurements from a real stack serve as the target domain\. In experiments incorporating varying amounts of real data into the training process, the proposed approach consistently outperforms a baseline trained exclusively on real data without the domain discriminator\. Notably, these performance gains become more pronounced as the quantity of real data decreases\.
#### 4\.3\.3Health Assessment
Table 7:All studies employingobservational biasto addresshealth assessment, listed in alphabetical order\.Observational Bias for Health AssessmentReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[bachar2024AMultidisciplinaryFramework\]Gear vibration modelNonlinear second\-order[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Gear\[kohtz2022Physicsinformedmachinelearning\]1D[FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)\-growth modelCoupled[PDEs](https://arxiv.org/html/2608.10047#p6.31.31.31.31)Lithium\-ion battery\[li2024PhysicsGuidedDeepLearning\]Empirical flank wear evolution model, mechanistic milling force modelAlgebraic equationsMilling\[yishengliu2024HybridFusionfor\]Extended[SPM](https://arxiv.org/html/2608.10047#p6.47.47.47.47)Nonlinear state\-space system derived from coupled[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)\-based electrochemical equationsLithium\-ion battery\[matania2023Onefaultshotlearningfor\]Gear vibration modelNonlinear second\-order[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Gear\[mei2024Ahybridphysicsinformed\]Physics\-of\-failure\-based relay degradation modelCoupled first\-order[ODEs](https://arxiv.org/html/2608.10047#p6.30.30.30.30)with time\-varying stochastic input parameters, evaluated via[FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)\-based numerical simulationElectromagnetic relays\[navidi2024PhysicsInformedMachineLearning\]Half\-cell degradation modelAlgebraic equationsLithium\-ion battery\[ren2025Healthassessmentof\]Electromagnetic and electromechanical modelCoupled[ODEs](https://arxiv.org/html/2608.10047#p6.30.30.30.30)implemented via 2D[FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)simulationBrushless direct\-current motor\[sun2022MicrocrackDefectQuantification\]Ultrasonic guided\-wave propagation and scattering model[PDEs](https://arxiv.org/html/2608.10047#p6.31.31.31.31)solved via[FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)simulationMicrocrack quantification in aluminum plate\[thelen2022Integratingphysicsbasedmodeling\]Half\-cell degradation modelAlgebraic equationsLithium\-ion batteryAs shown in Table[7](https://arxiv.org/html/2608.10047#S4.T7), ten studies employ observational bias for health assessment\. The study by\[yishengliu2024HybridFusionfor\]focuses on accurately estimating the[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)of lithium\-ion batteries under realistic fast\-charging conditions using minimal labeled data\. A reduced\-order electrochemical model is calibrated on laboratory cycling tests, expanded via stochastic perturbation of aging parameters, and then used to generate partial\-charge profiles for pretraining a[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\. The pretrained network is subsequently adapted to specific batteries via[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)with only a small number of real partial\-charge segments\. Ablation studies \(pre\- and post\-[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51), varying real\-data volume and distribution\) demonstrate that incorporating simulated data significantly improves generalization and enables extrapolation from early\-life data to mid\- and late\-life degradation states\.\[kohtz2022Physicsinformedmachinelearning\]introduce a multi\-fidelity framework for estimating the[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)of lithium\-ion batteries from a single short partial charging segment\. A 1D[Finite Element](https://arxiv.org/html/2608.10047#p6.16.16.16.16)\([FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)\)[Solid Electrolyte Interphase](https://arxiv.org/html/2608.10047#p6.44.44.44.44)\([SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)\)\-growth model is first used to simulate capacity fade and[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)thickness, which are subsequently fused with experimental data to train a co\-kriging multi\-fidelity model that links[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)thickness to capacity \([SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)\)\. Separately, nested[Gaussian Process Regression](https://arxiv.org/html/2608.10047#p6.21.21.21.21)\([GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)\) models are trained on[FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)data to map operating conditions and[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)levels to voltage points and charging time, which are then used to infer[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)thickness—and thus[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)—from a single partial charging segment without any prior usage history\. The authors validate their proposed framework by comparing the[FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)model, single\-fidelity surrogates, and the co\-kriging multi\-fidelity model, with the latter attaining the smallest capacity estimation errors\. However, it is not quantitatively benchmarked against alternative[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation methods\.\[thelen2022Integratingphysicsbasedmodeling\]address the online health assessment of lithium\-ion batteries by estimating both capacity \([SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)\) and three internal degradation modes\. Prior physical knowledge is incorporated via a half\-cell model, which is used to generate degradation data and subsequently combined with limited early\-life experimental data\. Two approaches are compared: data augmentation, in which simulated and experimental data are merged to train a single model, and delta learning, in which an estimator trained on simulated data is corrected by a second model trained on experimental data\. Across both scenarios, data augmentation consistently outperforms delta learning for all four lightweight[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models tested, with the elastic net \(a linear regression model with combined L1/L2 regularization\) achieving the highest overall performance\. As a follow\-up study,\[navidi2024PhysicsInformedMachineLearning\]compare four approaches, including the elastic net\-based data augmentation and delta learning methods from the previous study\[thelen2022Integratingphysicsbasedmodeling\]\. Additionally, delta learning employing[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)is examined, alongside another method that introduces a learning bias \(see Sec\.[4\.5\.3](https://arxiv.org/html/2608.10047#S4.SS5.SSS3)\)\. Among the approaches incorporating observational bias, the[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)\-based method yields the lowest errors for capacity estimation\. A qualitative comparison of all four methods, considering aspects such as model flexibility, data requirements, and ease of implementation, provides a broader perspective on their applicability\.
\[li2024PhysicsGuidedDeepLearning\]propose a method for online tool condition monitoring in milling, where the goal is to estimate flank wear from cutting\-force signals under varying operating conditions\. A mechanistic tool wear model and a milling force model are first calibrated on a small set of offline measurements and then used to generate synthetic data\. Simulated force signals are used to pretrain a[DL](https://arxiv.org/html/2608.10047#p6.10.10.10.10)model that combines a residual network with learnable soft\-threshold attention \(for denoising and feature extraction\) and a[Bidirectional Long Short\-Term Memory](https://arxiv.org/html/2608.10047#p6.2.2.2.2)\([BiLSTM](https://arxiv.org/html/2608.10047#p6.2.2.2.2)\) \(for mapping force features to simulated wear labels\)\. With a strong initialization provided by pretraining on simulated data, the model is subsequently fine\-tuned on a limited set of real samples\. In rigorous comparisons with its purely physics\-based and purely data\-driven counterparts, as well as with[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)approaches and several standard[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)baselines, the proposed method consistently achieves lower errors and more robust wear estimation across multiple milling conditions, while requiring substantially fewer real labels\.
By quantifying existing microcracks from ultrasonic guided\-wave measurements obtained via a specially designed, unidirectionally focusing electromagnetic acoustic transducer and a circular array of receivers around the suspected defect location,\[sun2022MicrocrackDefectQuantification\]tackle the health assessment of metallic plate structures\. To compensate for the small set of experimental measurements, synthetic guided\-wave signals are generated using an[FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)simulation that spans a broad range of crack geometries\. The architecture includes two branches with shared weights that process experimental and simulated signals in parallel, enabling the model to compare them and infer crack length\. The method additionally incorporates a learning bias \(see Sec\.[4\.5\.3](https://arxiv.org/html/2608.10047#S4.SS5.SSS3)\)\. While the proposed model achieves substantially lower errors compared to conventional[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)variants, the study does not include ablation studies to infer the corresponding contribution of the various biases incorporated\.
\[matania2023Onefaultshotlearningfor\]address fault severity estimation for spur gears when only one faulty experimental sample is available\. A dynamic gear model is used to generate vibration signals\. The vibration data undergo several preprocessing steps, such as angular resampling, synchronous averaging, propagation through an estimated transfer function, and extraction of time\-domain features\. The authors then mix the preprocessed simulation data with real data to serve as input for a[kNN](https://arxiv.org/html/2608.10047#p6.25.25.25.25)regressor\. The ablation studies conducted do not specifically study the effect of incorporating simulation data or comparisons against state\-of\-the\-art baselines\. Building on this work, some of the authors propose an unsupervised framework with a broader scope that can perform health assessment for multiple fault types\[bachar2024AMultidisciplinaryFramework\]\.
In their study on health assessment of brushless direct\-current motor stators,\[ren2025Healthassessmentof\]use a stator damage matrix derived from torque, speed, and load signals, which is then mapped to a scalar[Health Index](https://arxiv.org/html/2608.10047#p6.23.23.23.23)\([HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)\)\. To this end, an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)is trained on a mixed dataset of experimental and[FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)\-simulated torque sequences to predict a damage matrix describing turn\-level open\-circuit and short\-circuit faults\. Based on the predicted damage matrix, a scalar[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)is computed via cosine similarity between the current damage state and a health reference state, where a bootstrap procedure provides confidence intervals for the[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)\. Compared with four model\-based methods and a conventional[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26), the non\-invasive approach improves damage estimation performance, especially under small\-sample conditions\.
\[mei2024Ahybridphysicsinformed\]employ a variational[Autoencoder](https://arxiv.org/html/2608.10047#p6.1.1.1.1)\([AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1)\) to model degradation of the operating voltage signal of electromagnetic relays and perform time\-dependent reliability \(failure\-probability\) assessment\. A high\-fidelity physics\-of\-failure[FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)simulation of the relay’s electromagnetic and elastic subsystems is used to generate synthetic degradation trajectories, which are combined with a small set of experimental trajectories to train the variational[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1)\. This then functions as a generative degradation model: operating voltage trajectories are sampled over time, and at each time point, the fraction of generated samples exceeding a voltage threshold is used to estimate the time\-dependent failure probability of a batch of relays\. From a data perspective, training on the combined simulation and experimental trajectories outperforms training on either source alone\. From a modeling perspective, the variational[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1)achieves lower reliability\-assessment error and computational cost than both[Gaussian Process](https://arxiv.org/html/2608.10047#p6.20.20.20.20)\([GP](https://arxiv.org/html/2608.10047#p6.20.20.20.20)\)\- and[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\-based variants\.
#### 4\.3\.4Prognosis
Table 8:All studies employingobservational biasto addressprognosis, listed in alphabetical order\.Observational Bias for PrognosisReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[hervedebeaulieu2024RemainingUsefulLife\]Stiction model combined with a generic exponential degradation modelNonlinear difference equation with a time\-varying parameter prescribed by an exponential functionAir distribution system\[deng2023ACalibrationBasedHybrid\]5\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)dynamic bearing model[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)systemBearing\[zhang2023DynamicModelAssistedBearing\]5\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)dynamic bearing model[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)systemBearing\[zhang2024Modeldatahybriddriven\]Cutting model \(accounting for geometry, material, friction, and wear\)[FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)\-based numerical simulationMilling\[zhu2024RemainingUsefulLife\]Vibration degradation modelAnalytic model defined by algebraic equationsBearingFive studies employing observational bias for prognosis have been identified, which are listed in Table[8](https://arxiv.org/html/2608.10047#S4.T8)\.\[deng2023ACalibrationBasedHybrid\]leverage a 5\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)dynamic bearing model for[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction across different machines\. The dynamic model describes vibration under crack propagation and spall growth, with the associated damage parameters inferred via[PF](https://arxiv.org/html/2608.10047#p6.33.33.33.33)\-based calibration\. Both calibrated simulated data and real measurements are fed into a Bayesian[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29), which processes one\-dimensional time\-series features through a[Gated Recurrent Unit](https://arxiv.org/html/2608.10047#p6.22.22.22.22)\([GRU](https://arxiv.org/html/2608.10047#p6.22.22.22.22)\) branch and two\-dimensional time\-frequency features through a[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)branch, jointly predicting the[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)\. Within an adversarial[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)setup, multiple domain discriminators are weighted according to the similarity of calibrated physical parameters, so that source bearings with degradation dynamics closer to the target bearing have a stronger influence during adaptation\. Experiments on two benchmark bearing datasets demonstrate more accurate[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)predictions than purely data\-driven[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)baselines\. Similarly,\[zhang2023DynamicModelAssistedBearing\]generate full life\-cycle vibration data using a 5\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)dynamic bearing model\. The resulting simulated data serve as the source domain for a multilayer Transformer\-based network with maximum mean discrepancy\-based domain alignment, which is trained to transfer degradation knowledge to measured vibration signals\. Experimental results show that the proposed model substantially reduces[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction errors compared to several state\-of\-the\-art methods, particularly when only limited measured run\-to\-failure data are available\.\[zhu2024RemainingUsefulLife\]tackle[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction for rolling bearings, proposing a framework that incorporates physics in three ways\. In terms of observational bias, a physics\-based degradation model is used to simulate full life\-cycle vibration data\. The simulated signals are subsequently fused with sensed data for feature extraction, thereby enriching the training set used to train the[BiLSTM](https://arxiv.org/html/2608.10047#p6.2.2.2.2)\. In combination with the other two mechanisms \(see Sec\.[4\.5\.4](https://arxiv.org/html/2608.10047#S4.SS5.SSS4)and Sec\.[4\.6\.4](https://arxiv.org/html/2608.10047#S4.SS6.SSS4), respectively\), the proposed framework outperforms purely data\-driven baselines by achieving lower prediction errors\. However, due to the absence of ablation studies, it is not possible to isolate how much of the improvement is attributable to this particular strategy\.
Cutting\-tool[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction under scarce degradation data is studied by\[zhang2024Modeldatahybriddriven\], who propose combining a physics\-based cutting model with an improved inverse[GP](https://arxiv.org/html/2608.10047#p6.20.20.20.20)\. An[FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)cutting model is used offline to generate full\-life wear trajectories, which are used for initial parameter estimation of the inverse[GP](https://arxiv.org/html/2608.10047#p6.20.20.20.20)\. During operation, its parameters are updated online with measured wear via Bayesian inference\. Experiments on three tool\-wear datasets indicate improved predictive accuracy over a purely data\-driven baseline\.
In their study,\[hervedebeaulieu2024RemainingUsefulLife\]focus on predicting the[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)of industrial systems without relying on labeled run\-to\-failure data\. Linear closed\-loop models of an aircraft cockpit temperature control system are identified from flight data and coupled with an exponential valve\-stiction degradation model to generate nominal and degraded time series\. An[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1)is trained only on nominal data to derive an unsupervised[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)from reconstruction error, which is then forecast using an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\. The approach is validated through a reliability\-based assessment using Weibull and Kolmogorov\-Smirnov tests, showing internal consistency of predictions\. However, the study does not report experimental comparisons to alternative methods\.
### 4\.4Inductive Bias
Inductive bias is characterized by tailored interventions to the model design\. The identified studies are organized by PHM task in Tables[9](https://arxiv.org/html/2608.10047#S4.T9)–[12](https://arxiv.org/html/2608.10047#S4.T12)\. A representative example is illustrated in Figure[6](https://arxiv.org/html/2608.10047#S4.F6)\.
#### 4\.4\.1Fault Detection
Table 9:All studies employinginductive biasto addressfault detection, listed in alphabetical order\.Two studies have been identified that employ inductive bias to address fault detection \(see Tab\.[9](https://arxiv.org/html/2608.10047#S4.T9)\)\.\[liu2025Graphembeddedpatchsense\]address anomaly detection in large\-scale multi\-component systems with complex interdependencies\. The authors propose a model based on an[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1)and graphs that learns unit\-level patch embeddings and models global relationships in multivariate time series in an unsupervised manner\. Prior knowledge of system structure is encoded as a graph and incorporated via an adjacency matrix, guiding the model to produce representations consistent with known system topology for both anomaly detection and component\-level localization\. Experiments on liquid rocket engine and train transmission datasets demonstrate improved accuracy and reliability, at a higher computational cost\. In a closely related study also considering the liquid rocket engine system,\[feng2022FullGraphAutoencoder\]encode prior knowledge into the adjacency matrix as well\. Topological relationships of physical sensors in terms of hardware layout or shared subsystems are leveraged\. Improvements in both accuracy and generalizability are achieved, compared to strong non\-graph\-based and graph\-based baselines\.
#### 4\.4\.2Diagnosis
Figure 6:A representative example of inductive\-bias approaches is proposed by\[zeng2025ApplicationofFrequency\], demonstrating how prior physical knowledge can be introduced into the model design\. Own illustration based on the corresponding study\. See the original work for full technical details\.Table 10:All studies employinginductive biasto addressdiagnosis, listed in alphabetical order\.Inductive Bias for DiagnosisReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[gao2024MPINet:MultiscalePhysicsInformed\]Vibration model for a localized single\-point defect, bearing fault characteristic frequenciesAnalytical time\-domain signal model, algebraic equationsBearing\[jin2024GraphSpatioTemporalNetworks\]Structural topology of the wind turbine and qualitative causal relationshipsDirected graph over selected sensor signals \(adjacency matrix\)Wind turbine\[zeng2025ApplicationofFrequency\]Fault characteristic frequenciesAlgebraic equationsBearingAs shown in Table[10](https://arxiv.org/html/2608.10047#S4.T10), three studies employ inductive bias to target diagnosis\.\[zeng2025ApplicationofFrequency\]tackle rolling bearing fault diagnosis by embedding bearing fault frequency information into a[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)via a novel[Frequency\-Aware Convolution](https://arxiv.org/html/2608.10047#p6.15.15.15.15)\([FAC](https://arxiv.org/html/2608.10047#p6.15.15.15.15)\) layer \(see Fig\.[6](https://arxiv.org/html/2608.10047#S4.F6)\)\. Vibration signals are first transformed into time\-frequency maps via[Continuous Wavelet Transform](https://arxiv.org/html/2608.10047#p6.9.9.9.9)\([CWT](https://arxiv.org/html/2608.10047#p6.9.9.9.9)\), which serve as inputs to the network\. The architecture is built from multiple stacked units, each comprising a residual block, the[FAC](https://arxiv.org/html/2608.10047#p6.15.15.15.15)layer, and an[Efficient Channel Attention](https://arxiv.org/html/2608.10047#p6.12.12.12.12)\([ECA](https://arxiv.org/html/2608.10047#p6.12.12.12.12)\) module\[wang2020eca\]\. Using the characteristic fault frequencies of the inner ring \([BPFI](https://arxiv.org/html/2608.10047#p6.3.3.3.3)\), outer ring \([BPFO](https://arxiv.org/html/2608.10047#p6.4.4.4.4)\), and rolling body \([BSF](https://arxiv.org/html/2608.10047#p6.5.5.5.5)\), the[FAC](https://arxiv.org/html/2608.10047#p6.15.15.15.15)layer generates a set of sinusoidal waveforms to form frequency response kernels \(one per frequency\) with a bandpass profile\. A learnable shift parameter, initialized to zero and constrained within a physically meaningful range, allows these kernels to dynamically adapt to varying vibration signals across different operating conditions while remaining close to the analytical fault frequencies\. Each frequency response kernel is multiplied element\-wise with a standard convolutional kernel to form enhanced kernels that selectively emphasize the corresponding fault\-related frequency bands while suppressing irrelevant components\. The proposed model demonstrates superior accuracy, noise robustness, and generalization across varying loads compared with conventional[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\-based models, as validated on two datasets\. Additionally, feature\-space visualization \(based on t\-distributed stochastic neighbor embedding\) reveals physically consistent clusters aligned with bearing fault mechanisms\. A conceptually similar, yet structurally different approach is proposed by\[gao2024MPINet:MultiscalePhysicsInformed\], who use a set of subnetworks to target small\-sample learning\. Each subnetwork is tailored to a specific failure mode \(class\)\. Based on the corresponding bearing fault characteristic frequency, a kernel that simulates an ideal bearing fault vibration replaces the data\-driven kernel of the first convolutional layer of an otherwise conventional[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8), while the healthy\-condition subnetwork retains a standard data\-driven first layer\. These subnetworks are trained independently as binary classifiers to specialize in fault\-specific feature extraction from vibration signals\. Fusing the extracted features and feeding them to a unifying classifier enables multi\-class diagnosis of rolling bearings\. Validated on two datasets, the proposed model improves small\-sample accuracy compared to both classical baselines and[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\-based models\.
In their work,\[jin2024GraphSpatioTemporalNetworks\]propose a spatio\-temporal[GNN](https://arxiv.org/html/2608.10047#p6.19.19.19.19)for wind turbine condition monitoring\. They construct a directed graph based on prior knowledge of turbine structure and inter\-component causal relationships, with nodes representing sensor signals and edges encoding structural and causal connections\. On this prior graph, a graph attention network layer captures spatial dependencies between signals, while global\-local attention and recurrent layers model their temporal evolution under healthy operation\. The network is trained to predict the next\-step values of all signals, and deviations between predicted and actual values are monitored at both the graph and node levels\. By doing node\-level anomaly detection, it does perform component\-level fault localization, which is a practically meaningful form of diagnosis\. A multi\-node fault propagation chain, constrained by the prior graph, is used to distinguish true faults from scattered false alarms and to identify the origin of abnormal behavior\. Case studies on a real turbine with a generator bearing failure and several false alarms show comparable early warning capability, fewer false alarms, and more interpretable monitoring compared with purely data\-driven[DL](https://arxiv.org/html/2608.10047#p6.10.10.10.10)baselines\.
#### 4\.4\.3Health Assessment
Table 11:All studies employinginductive biasto addresshealth assessment, listed in alphabetical order\.Inductive Bias for Health AssessmentReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[kristupasbajarunas2024HealthIndexEstimation\]Causal relationship \(sensors, operating conditions, degradation\)Architectural constraint \(and corresponding regularization terms\)Turbofan engine, lithium\-ion battery\[cheng2024Researchongas\]Thermodynamic and structural relations of the gas turbineAdjacency matrixGas turbine\[ellis2022Ahybridframework\]Model relating crack length and first natural frequencyNon\-stationary[GP](https://arxiv.org/html/2608.10047#p6.20.20.20.20)priorTurbomachine rotor blade\[fu2024PhysicsInformedNeuralNetwork\]Second\-order[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13)Parametric state\-space modelLithium\-ion battery\[hao2023Anoveldeep\]Monotonicity assumptionActivation functionMilling\[huang2022AnEnhancedDataDriven\]Empirical degradation modelAlgebraic equationLithium\-ion battery\[lehmann2024LearningtheAgeing\]Cyclic stress model, calendric stress model and aging curve modelAlgebraic equationsLithium\-ion battery\[li2025Applicationofphysicsguided\]Monotonicity assumptionArchitectural constraintsMilling\[ma2023Ahybriddrivenprobabilistic\]Degradation dynamics \(Wiener process\)Stochastic discrete\-time state\-space modelMilling\[qin2025ManagingBatteryPerformance\]First\-order[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13)with hysteresis and Coulomb counting model[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)\-based layersLithium\-ion battery\[xie2024DegradationStateAssessment\]Monotonically increasing one\-dimensional degradation state in the range\[0,1\]\[0,1\]Architectural constraint \(with associated loss terms\)Semiconductor \(insulated gate bipolar transistor\)\[yucesan2019Windturbinemain,yucesan2020Ahybridmodel,yucesan2021Hybridphysicsinformedneural,yucesan2022Ahybridphysicsinformed,yucesan2023Physicsinformeddigitaltwin\]Cumulative fatigue damage model with lubricant influenceFirst\-order[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30), algebraic equationsBearing\[zhu2023PhysicsinformedGaussianprocess\]Empirical tool wear modelsAlgebraic equationsMillingInductive bias is used in 17 studies on health assessment \(see Tab\.[11](https://arxiv.org/html/2608.10047#S4.T11)\)\. To estimate the[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)of lithium\-ion batteries,\[huang2022AnEnhancedDataDriven\]design a custom kernel for[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)based on an empirical degradation model that captures both linear and exponential decay over cycling\. The kernel hyperparameters are initialized using the parameters of the empirical model \(identified via least squares\), facilitating subsequent optimization\. Benchmarking against standard Gaussian and Matérn kernels shows that the proposed kernel consistently achieves lower prediction errors on two experimental datasets\.\[fu2024PhysicsInformedNeuralNetwork\]propose a[Recurrent Neural Network](https://arxiv.org/html/2608.10047#p6.43.43.43.43)\([RNN](https://arxiv.org/html/2608.10047#p6.43.43.43.43)\) whose architecture is derived from a discretized second\-order[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13), in which the open\-circuit voltage is modeled by an[MLP](https://arxiv.org/html/2608.10047#p6.28.28.28.28)\. Treating the[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13)parameters as trainable variables enables accurate online parameter identification and terminal\-voltage prediction with low computational complexity\. Experimental results demonstrate that the proposed method outperforms both a white\-box neural circuit model and an enhanced[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13)\. Further analysis reveals a linear relationship between the identified ohmic resistance and remaining capacity, enabling accurate[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation\. Also addressing[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation of lithium\-ion batteries,\[lehmann2024LearningtheAgeing\]propose a stress\-factor\-based aging model that is parameterized using both laboratory aging tests and data from an electric bus fleet\. The model consists of cyclic and calendric stress maps and a square\-root aging curve\. The stress maps are represented by[NNs](https://arxiv.org/html/2608.10047#p6.29.29.29.29)that share this fixed aging\-curve structure with a separately fitted empirical model\. Capacities predicted by the empirical model at collocation points are included in the training objective, which ties the learned stress maps to the established functional relationships in regions that are sparsely covered by data\. Using[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)to adapt the aging\-curve parameters to different cell types and to the fleet, the approach improves capacity and[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimates at check\-up tests compared with an equal\-stress baseline and with the coupled model without adaptation\. For lithium\-ion battery health management,\[qin2025ManagingBatteryPerformance\]propose a predictor\-estimator framework that jointly handles[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45), internal resistance, and maximum available capacity\. A linear\-exponential two\-stage degradation model with recursive least\-squares updating and Bayesian knee\-point detection predicts the evolution of several health indicators \([SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46), a resistance\-based[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23), efficiency, and[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45)at discharge start\), from which future capacity and resistance trajectories are derived\. Three lightweight neural estimators then infer[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45), resistance, and capacity from operational time series, where physics is embedded by hard\-wiring Coulomb counting dynamics with learnable efficiency and initial\-[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45)corrections for[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45)estimation\. The second estimator couples an[MLP](https://arxiv.org/html/2608.10047#p6.28.28.28.28)with an equivalent\-circuit[Ordinary Differential Equation](https://arxiv.org/html/2608.10047#p6.30.30.30.30)\([ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)\) model whose parameters are learned under physical bounds for resistance estimation, while the capacity estimator exploits these latent variables via channel\-attention and temporal convolutions\. Experiments on two public battery datasets show that this framework provides more accurate degradation prediction and state estimation than state\-of\-the\-art baselines, generalizes well across different chemistries, temperatures, and loading conditions, and achieves these benefits with very compact estimator networks and low computational overhead\.
\[yucesan2019Windturbinemain\]propose a framework for estimating wind\-turbine main bearing fatigue life by combining a physics\-based cumulative\-damage model for bearing fatigue with a data\-driven model for grease degradation, organized within a recurrent model that models damage accumulation over time\. Bearing fatigue is determined incrementally at each timestep by computing the fatigue damage rate from the current load and speed using an L10\-based bearing life relation and accumulating damage via Palmgren\-Miner’s rule\. Simultaneously, an[MLP](https://arxiv.org/html/2608.10047#p6.28.28.28.28)predicts increments in grease degradation—a hidden variable affecting fatigue calculations through viscosity and contamination factors—but is only indirectly supervised through periodic grease observations\. By embedding these components within a custom recurrent cell, the model captures both the well\-understood physics of bearing fatigue and the complex dynamics of grease degradation that are difficult to model from first principles\. While promising, the results provide no quantitative performance metrics and lack baseline comparisons, limiting the ability to fully assess the approach’s effectiveness\. Even though the model is elaborated upon in subsequent works\[yucesan2020Ahybridmodel,yucesan2021Hybridphysicsinformedneural,yucesan2022Ahybridphysicsinformed,yucesan2023Physicsinformeddigitaltwin\], the fundamental approach remains unchanged from a[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)perspective\.
Studying tool wear estimation in high\-speed milling,\[li2025Applicationofphysicsguided\]propose embedding an architectural constraint into a[GRU](https://arxiv.org/html/2608.10047#p6.22.22.22.22)network\. To account for the irreversible nature of tool wear, monotonicity is enforced via constrained hidden\-state updates and using[Rectified Linear Unit](https://arxiv.org/html/2608.10047#p6.37.37.37.37)\([ReLU](https://arxiv.org/html/2608.10047#p6.37.37.37.37)\) activations\. This ensures that predicted wear values cannot decrease over time, which improves both physical plausibility and predictive accuracy compared to unconstrained data\-driven models\. See Section[4\.5\.3](https://arxiv.org/html/2608.10047#S4.SS5.SSS3)for a description of the learning bias incorporated by\[li2025Applicationofphysicsguided\]\. A similar approach is adopted by\[hao2023Anoveldeep\], where a softplus activation function is incorporated into the network architecture prior to the output layer\. This design choice explicitly accounts for the monotonic degradation behavior of milling tool wear, ensuring nondecreasing wear predictions over time\. As a result, the proposed model likewise achieves improved predictive accuracy compared to purely data\-driven baseline approaches\.\[zhu2023PhysicsinformedGaussianprocess\]propose a[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)\-based approach for predicting tool wear\. Prior knowledge is incorporated through three physics\-based tool wear models \(generalized Taylor formula, a cubic polynomial wear\-time law, and a generic flank wear model\) that describe the evolution of tool flank wear over cutting time\. By adopting these wear laws as the prior mean function, the model supports small\-sample training, continual online updating, and substantially improved extrapolation\. Comparative experiments confirm that the proposed approach significantly reduces prediction error and improves robustness relative to both standalone physics\-based models and purely data\-driven approaches\.\[ma2023Ahybriddrivenprobabilistic\]also address tool wear monitoring in milling\. A physics\-based state\-space model based on a Wiener process\-inspired wear law serves as prior knowledge within a[GP](https://arxiv.org/html/2608.10047#p6.20.20.20.20)\-based probabilistic state\-space model\. A[PF](https://arxiv.org/html/2608.10047#p6.33.33.33.33)estimates the unknown posterior distribution of the degradation state\. Compared with data\-driven baselines such as[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8),[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26), and[Support Vector Regression](https://arxiv.org/html/2608.10047#p6.48.48.48.48)\([SVR](https://arxiv.org/html/2608.10047#p6.48.48.48.48)\), as well as purely physics\-based approaches, the method yields a better predictive performance and tighter confidence intervals for tool replacement decisions\. The method is also suitable for prognosis tasks and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction\.
\[kristupasbajarunas2024HealthIndexEstimation\]aim to leverage general knowledge about degradation to broaden the applicability of unsupervised[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)estimation across various systems\. They use assumed causal relationships between sensor readings, operating conditions, and the degradation to design the architecture of a convolutional[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1): the encoder maps sensor readings to a scalar latent variable, while the decoder reconstructs sensor readings from this latent variable and operating conditions, thereby forcing the latent variable to encode degradation information\. To further shape this latent variable into a suitable[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23), they complement this inductive\-bias approach with loss terms that guide[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)properties desired in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), such as monotonicity and trendability, as well as an optional term that encourages consistency with degradation trends derived from reliability theory\. The approach is validated on both turbofan engine data and lithium\-ion battery data, leading to improvements in prediction performance and out\-of\-distribution robustness compared to residual\-based baselines\. With a subsequent[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8), the approach is extended for[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction, yet the[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)aspects in this paper solely correspond to health assessment\.
To assess the degradation state of insulated gate bipolar transistor modules under varying operating conditions,\[xie2024DegradationStateAssessment\]propose an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\-based[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1)designed to disentangle degradation from operating conditions\. The encoder compresses inputs into a low\-dimensional latent vector, reserving a single scalar for the degradation index while the remaining dimensions capture variations in operating conditions\. Two decoders are then employed: one reconstructs non\-degraded behavior from the condition\-related latents alone, and the other reconstructs degraded behavior using the full latent vector\. This design forces degradation information to flow through a single neuron and yields an interpretable[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)suitable for online monitoring\. Complementing this inductive\-bias approach, the authors employ additional loss terms that guide this single health\-indicator neuron to satisfy both monotonicity and range constraints\. This leads to better results in terms of prediction performance and physical consistency compared to the inductive\-bias approach alone\.
\[cheng2024Researchongas\]assess gas turbine health via a spatio\-temporal[GNN](https://arxiv.org/html/2608.10047#p6.19.19.19.19)\. Thermodynamic and structural knowledge is encoded by constructing a temporal graph over key monitored parameters, whose edges combine[kNN](https://arxiv.org/html/2608.10047#p6.25.25.25.25)\-based data correlations with links derived from small\-deviation compressor, combustor, and turbine equations\. This topology effectively constrains how information propagates between variables, enabling accurate health assessment by mapping multivariate time series to discrete health stages, as corroborated by corresponding ablation studies\.
As part of a broader framework,\[ellis2022Ahybridframework\]address the health assessment of turbomachine rotor blades by estimating root crack length from blade tip timing\-derived natural frequencies under scarce inspection data\. A physics\-based model \([FE](https://arxiv.org/html/2608.10047#p6.16.16.16.16)simulations with an unscented transform\) is first built to map natural frequency to crack length, and this ensemble is then used as the prior mean and covariance of a[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)model\. The latter is conditioned on crack\-length measurements obtained via non\-destructive testing during routine maintenance\. The proposed model preserves physically plausible behavior in data\-sparse regions while correcting systematic errors near observed non\-destructive testing points, and it consistently outperforms both the pure physics\-based model and several purely data\-driven regressors\.
#### 4\.4\.4Prognosis
Table 12:All studies employinginductive biasto addressprognosis, listed in alphabetical order\.Inductive Bias for PrognosisReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[abiria2025Highcycleandveryhighcycle\]Basquin’s law, nonnegativity assumptionAlgebraic equations, differential equations, architectural constraintsAdditive manufacturing\[badora2023Usingphysicsinformedneural\]Paris’ lawDifferential equationHigh\-pressure nozzle of an industrial gas turbine\[bai2023PrognosticsofLithiumIon\]Bounded and monotonically decreasing capacity fade over cyclesInequality constraintsLithium\-ion battery\[cai2025Knowledgeembeddedspatial\]System/sensor topologyKnowledge graphTurbofan engine, milling\[dourado2019Physicsinformedneuralnetworks,dourado2022Ensembleofhybrid\]Walker model for fatigue crack propagationRecurrent algebraic equationAircraft fuselage panels\[jiang2025PhysicsinformedGaussianprocess\]Paris’ law for fatigue crack growth[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)discretized to a damage accumulation modelAluminum specimens\[nascimento2021Hybridphysicsinformedneural\]Nernst and Butler\-Volmer equationsAlgebraic equationsLithium\-ion battery\[nguyen2023Physicsinfusedfuzzygenerative\]Spall\-growth model, modified Eyring modelAlgebraic equationBearing, turbofan engine\[qiang2023Integratingphysicsinformedrecurrent\]Empirical linear tool wear modelAlgebraic equationMilling\[qin2024AnInterpretableNeuroDynamic\]Power equations of low\- and high\-pressure compressor, high\-speed shaft dynamicsAlgebraic equation,[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Turbofan engine\[yin2025Physicsguideddegradationtrajectory\]Monotonicity assumptionArchitectural constraintsBearing\[zhang2025Applicationofphysicsinformed\]Diamond\-shaped wear particle model, Archard wear model and wear model by\[zou1996abrasivewearmodel\]Algebraic equationsAxial piston pump\[zhou2023Timevaryingtrajectorymodeling\]Nonnegativity assumptionArchitectural constraintTurbofan engine, bearing\[zhou2025Physicsinformedspatiotemporalhybrid\]Thermal cycle modelFixed sensor association graph \(adjacency matrix\)Turbofan engineTable[12](https://arxiv.org/html/2608.10047#S4.T12)provides an overview of the 15 identified studies regarding inductive bias for prognosis\.\[nascimento2021Hybridphysicsinformedneural\]propose a method for lithium\-ion batteries that embeds core electrochemical relations to predict voltage discharge curves \(and thereby end\-of\-discharge time\) under varying loads and to forecast aging\-induced capacity fade and resistance growth\. A reduced\-order model based on the Nernst and Butler\-Volmer equations is implemented as a recurrent cell, while[MLPs](https://arxiv.org/html/2608.10047#p6.28.28.28.28)replace the non\-ideal voltage \(activity\) terms that are difficult to capture analytically\. The framework treats lumped internal resistance and maximum available charge as cell\- and age\-dependent parameters, and models their evolution with cumulative discharged energy using variational ensemble learning to obtain quantitative aging indicators and uncertainty\-aware forecasts\. Despite extensive experiments, no quantitative results are reported that demonstrate improvements over purely data\-driven baselines\. Methodologically, this approach is closely related to the work of\[yucesan2019Windturbinemain\]and its follow\-up studies \(see Sec\.[4\.4\.3](https://arxiv.org/html/2608.10047#S4.SS4.SSS3)\)\.\[bai2023PrognosticsofLithiumIon\]address lithium\-ion battery capacity prognostics using a two\-stage framework\. In the first stage, an[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)combined with a dual[Extended Kalman Filter](https://arxiv.org/html/2608.10047#p6.14.14.14.14)\([EKF](https://arxiv.org/html/2608.10047#p6.14.14.14.14)\) is employed for online estimation of the[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45)and capacity from voltage and current measurements, producing capacity trajectories\. In the second stage, these trajectories are used as inputs to a[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)\-based degradation model to forecast future capacity evolution\. Two inequality constraints are imposed on the[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21), ensuring that capacity predictions remain bounded and monotonically decreasing with respect to the cycle number\. The constrained[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)outperforms its unconstrained version, as well as three additional baseline methods in terms of both predictive accuracy and reduced uncertainty\.
Addressing[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction for rolling bearings,\[yin2025Physicsguideddegradationtrajectory\]focus on physically consistent modeling of degradation\. The paper proposes a method that leverages phase space reconstruction to transform vibration signals into trajectories, turning[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction into a variation estimation problem\. Prior knowledge about the monotonic nature of bearing degradation is integrated into a 1D\-[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)’s final activation function, ensuring that the predicted[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)cannot increase unrealistically\. This improves robustness, smoothness, and physical plausibility of predictions compared to purely data\-driven models\. Comparative experiments under varying working conditions demonstrate that the approach outperforms state\-of\-the\-art methods in both predictive accuracy and generalizability\.\[nguyen2023Physicsinfusedfuzzygenerative\]address the limitations of purely data\-driven[GANs](https://arxiv.org/html/2608.10047#p6.18.18.18.18)for[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction, including instability, sample inefficiency, and lack of physical consistency\. Their proposed architecture attaches a differentiable fuzzy logic module to the output of a conditional[GAN](https://arxiv.org/html/2608.10047#p6.18.18.18.18)’s generator\. Within this module, the standard product aggregation operator is replaced by a dataset\-specific physics model \(a spall\-growth model for bearings; a modified Eyring model for turbofan engines\)\. The generator thus learns fuzzy implications whose values serve as parameters of the physics model, constraining predictions to physically realistic solutions while the adversarial training signal still flows end\-to\-end through both the fuzzy and physics layers\. Experiments on two different datasets \(bearings and turbofan engines\) show reduced prediction errors\. Additional experiments, in which the dataset size is iteratively reduced, demonstrate improved data efficiency compared to a conventional[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\.\[zhou2023Timevaryingtrajectorymodeling\]target[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction as a time\-varying trajectory modeling problem rather than a point\-wise estimation\. Building on Neural ODE\[chen2018Neuralordinarydifferentialequations\], the proposed approach integrates additional prior physical knowledge about smooth degradation behavior\. Physically meaningful[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)trends are enforced by incorporating a nonnegative bounded function prior to providing the final prediction\. Furthermore, a dynamic learning scheme utilizing a super\-network\[wu2021neuralarchitecturesearchassparsesupernet\]and deep[RL](https://arxiv.org/html/2608.10047#p6.42.42.42.42)enables adaptive time\-dependent network architectures, improving the model’s ability to capture underlying degradation dynamics\. This approach results in more stable and interpretable[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)predictions that align with real\-world degradation processes\. The proposed method shows smooth and accurate prediction results compared to common data\-driven methods such as[Residual Network](https://arxiv.org/html/2608.10047#p6.38.38.38.38)\([ResNet](https://arxiv.org/html/2608.10047#p6.38.38.38.38)\) and[BiLSTM](https://arxiv.org/html/2608.10047#p6.2.2.2.2), as demonstrated through experiments on both bearing and turbofan engine datasets\.
\[qiang2023Integratingphysicsinformedrecurrent\]study tool wear prediction in milling under varying cutting parameters, where new operating conditions offer only limited labeled wear data\. To tackle this challenge, an instance\-based regression transfer algorithm \(Two\-stage TrAdaBoost\.R2, proposed by\[pardoe2010boosting\]\) is combined with a recurrent[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)base learner\. The[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)’s mean function encodes empirical physical relations between flank wear, cutting power, and previous wear to capture time\-accumulation and degradation trends, while the kernel models residual nonlinearities\. Using only about 30 % of early\-life wear data, the framework accurately extrapolates the full wear trajectory and yields tight confidence intervals\. The proposed framework is evaluated against three alternative approaches: an otherwise identical framework employing a standard \(uninformed\)[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21), a recurrent[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21), and an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26), with the latter two not incorporating[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)\. Experiments show that the proposed framework achieves substantially lower prediction errors, better tracking of late\-life wear growth, and more stable performance across different cutting\-parameter combinations\.
Unlike many studies that rely only on temporal sensor sequences,\[cai2025Knowledgeembeddedspatial\]take spatial interactions among multiple sensors into account\. Two use cases \(milling and turbofan engines\) are considered to leverage system topology and sensor placement information\. In a first step, embeddings are learned that represent the real\-world topological structure\. This is done by employing an energy\-based knowledge embedding algorithm\. Distances between embeddings are used to construct graph edges and an initial weighted adjacency matrix, whose weights are then dynamically updated via an attention mechanism\. These are fed into a model consisting of spatial modules \(graph convolutional network and attention mechanism\) as well as temporal modules \([LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\) and a final fully connected layer for direct[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction\. In ablation scenarios considering the turbofan use case, it is shown that temporal modules, spatial modules, and the knowledge encoding in the adjacency matrix contribute to the model performance\. In comparison with multiple baselines \(e\.g\.,[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8),[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26), and Bayesian models\), the proposed method achieves superior performance in direct[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction tasks in the majority of scenarios\. Comparable results are also reported for the milling use case\.\[qin2024AnInterpretableNeuroDynamic\]address the inadequate interpretability prevalent in[DL](https://arxiv.org/html/2608.10047#p6.10.10.10.10)methods applied to[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)estimation, with a particular focus on turbofan engines\. The proposed model comprises an augmenter for noise filtering, interpolation, and unobservable state estimation, and an estimator for[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction\. The augmenter is based on a Neural[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)framework that integrates physical models, data\-driven models, and a Runge\-Kutta[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)solver, all trained end\-to\-end\. The estimator combines[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\-based encoding with feature and temporal attention\. As prior knowledge, power equations of the low\-pressure compressor and high\-pressure compressor, shaft\-speed dynamics, and stall\-margin and efficiency\-modifier equations are embedded into the augmenter\. Compared to a variety of data\-driven baselines, the proposed method proves to be superior in terms of predictive performance\. Also targeting the[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction of turbofan engines,\[zhou2025Physicsinformedspatiotemporalhybrid\]embed thermodynamic prior knowledge into a spatio\-temporal[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)via graph construction\. In the spatial branch, thermodynamic cycle and engine\-structure knowledge are used to build an association graph between sensors, which is fused with a gray\-relation graph to form a fixed adjacency matrix for a multilayer graph attention network with pooling\. In the temporal branch, an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)extracts temporal features, while a temporal\-pattern attention module derives time\-invariant features from its hidden states, which together form the temporal\-domain features\. An attention module then fuses temporal and spatial features for[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction\. Experiments on C\-MAPSS show that the proposed model achieves consistently lower prediction errors than standard[DL](https://arxiv.org/html/2608.10047#p6.10.10.10.10)and other spatio\-temporal baselines\.
\[zhang2025Applicationofphysicsinformed\]address the degradation of hydraulic piston pumps, which they characterize as arising from the interplay of several factors, most notably the progressive wear of internal friction pairs together with fluctuating external load and operating conditions\. To capture this behavior, they propose a framework for predicting the[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)of hydraulic piston pumps built around an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\. Wear laws for three key friction pairs \(valve plate, piston, slipper\) are embedded in an end\-to\-end design, with wear model parameters updated jointly with the[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)weights\. These adaptive wear models convert monitoring data into degradation indicators, which the LSTM uses to predict return oil flow\.[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)is estimated as the time until the predicted flow exceeds a critical threshold\. The proposed method outperforms conventional baselines \(including[SVR](https://arxiv.org/html/2608.10047#p6.48.48.48.48),[RF](https://arxiv.org/html/2608.10047#p6.39.39.39.39), and[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\), though more advanced prognostics methods were not evaluated\.
Aging aircraft fleets are studied in the works of\[dourado2019Physicsinformedneuralnetworks,dourado2022Ensembleofhybrid\], with a specific focus on fatigue crack propagation in aircraft fuselage panels\. In both contributions, a physical fatigue crack propagation model \(the Walker model, an adapted version of the well\-known Paris law\) is embedded directly into a custom recurrent cell, while a data\-driven branch within the cell learns a correction term that accounts for corrosion effects not captured by the Walker model\. However, the authors do not benchmark their method against purely data\-driven or purely physics\-based baselines, making it difficult to quantitatively assess the added value of the proposed approach\. Similar to\[nascimento2021Hybridphysicsinformedneural\], this approach is also methodologically related to the work of\[yucesan2019Windturbinemain\]and its follow\-up studies\.\[jiang2025PhysicsinformedGaussianprocess\]study probabilistic prognosis of fatigue crack growth in metallic specimens\. Monte Carlo simulations of a physics\-based model \(Paris’ law\) are truncated using a standard[GP](https://arxiv.org/html/2608.10047#p6.20.20.20.20)fitted to current observations, whereas the resulting trajectories are used to construct a non\-stationary prior mean and covariance for the final[GP](https://arxiv.org/html/2608.10047#p6.20.20.20.20)\. Experiments show that incorporating these priors markedly improves extrapolation performance and predictive accuracy compared to both a standard[GP](https://arxiv.org/html/2608.10047#p6.20.20.20.20)and a[PF](https://arxiv.org/html/2608.10047#p6.33.33.33.33)\.\[badora2023Usingphysicsinformedneural\]introduce a custom[RNN](https://arxiv.org/html/2608.10047#p6.43.43.43.43)cell designed to model fatigue crack growth in a gas turbine nozzle\. An[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)estimates the stress intensity factor range at shutdown, while a physics\-based part then applies Paris’ law to compute the crack length increment due to fatigue\. The model accurately predicts crack growth over multiple cycles, even with limited observed data, and outperforms standard regression models in terms of predictive performance\. In their work,\[abiria2025Highcycleandveryhighcycle\]tackle the challenge of predicting fatigue life in additively manufactured alloys, where cyclic loading leads to microscopic damage accumulation and eventual failure\. The authors address this problem by integrating prior knowledge into a conventional[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)through modified activation functions derived from Basquin’s law, a modified Paris law, and a nonnegativity condition\. The proposed model shows better generalization capabilities compared to other physics\-informed and purely physics\-based variants\.
### 4\.5Learning Bias
Learning bias is characterized by introducing prior physical knowledge into the learning algorithm, frequently realized via a composite loss function\. Such a composite loss typically combines a data\-fidelity term, which minimizes the discrepancy between predictions and observations, with one or more physics\-informed terms that penalize violations of the governing equations, boundary conditions, or other physical constraints, each scaled by a weighting coefficient that controls its relative influence on training\. The identified studies are organized by PHM task in Tables[13](https://arxiv.org/html/2608.10047#S4.T13)–[16](https://arxiv.org/html/2608.10047#S4.T16)\. A representative example is illustrated in Figure[7](https://arxiv.org/html/2608.10047#S4.F7), which includes a composite loss function\.
#### 4\.5\.1Fault Detection
Table 13:All studies employinglearning biasto addressfault detection, listed in alphabetical order\.Learning Bias for Fault DetectionReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[wang2025Physicallyinformedhierarchical\]Wave\-type vibration model of spline shaftNonhomogeneous 1D wave[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)Aero\-engine involute spline coupling\[wang2024Adigitaltwin\]Multi\-energy model of a robot joint, multi\-joint rigid\-body dynamicsAlgebraic energy\-balance equations and recursive Newton\-Euler dynamic equationsIndustrial multi\-axis robot\[xu2024Physicsguideddeeplearning\]Damage index model, monotonic stress\-strain relationAlgebraic equationsCarbon fiber reinforced polymer laminatesTable[13](https://arxiv.org/html/2608.10047#S4.T13)lists the three studies that employ learning bias for fault detection\.\[xu2024Physicsguideddeeplearning\]target Lamb\-wave\-based fatigue damage detection in carbon fiber reinforced polymer laminates by augmenting a[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)with an additional branch that estimates global stiffness degradation from time\-frequency images of guided\-wave signals obtained via[CWT](https://arxiv.org/html/2608.10047#p6.9.9.9.9)\. Their approach incorporates additional loss terms based on a damage index model and a stress\-strain constraint, thereby regularizing training toward physically consistent progressive degradation\. The damage index model relates stiffness degradation and off\-axis angle to normalized power spectral density changes of Lamb\-wave responses, converting the network’s predicted stiffness into pseudo\-damage labels\. The stress\-strain constraint enforces monotonically increasing strain under constant load, penalizing non\-monotonic strain trajectories derived from the predicted stiffness\. Using only data from one carbon fiber reinforced polymer layup, the method generalizes to unseen layups with substantially improved cross\-structure detection performance compared to a standard[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8), and the resulting path\-level damage predictions support accurate delamination localization\.
\[wang2025Physicallyinformedhierarchical\]propose a soft sensing framework for estimating difficult\-to\-measure aero\-engine variables\. The approach extends[PINNs](https://arxiv.org/html/2608.10047#p6.36.36.36.36)to nonhomogeneous[Partial Differential Equations](https://arxiv.org/html/2608.10047#p6.31.31.31.31)with unknown, unmeasurable driving terms by training two coupled[NNs](https://arxiv.org/html/2608.10047#p6.29.29.29.29)in a hierarchical \(alternating\) optimization scheme: one approximates the[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)solution, the other the unknown source term\. By employing a joint loss in which the learned source is embedded in the[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)residual, the solution is regularized toward[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)\-consistent behavior, while the source is simultaneously constrained to produce driving terms that are compatible with both the measurements and the governing equation\. A recurrent\-prediction term further refines the solution using delayed hard\-sensor and soft\-sensor outputs to mitigate information loss due to sparse sampling and unmeasured sources\. Applied as a virtual vibration sensor on an aero\-engine spline\-coupling test rig, the method achieves substantially lower prediction errors than standard[PINNs](https://arxiv.org/html/2608.10047#p6.36.36.36.36)and supports a proof\-of\-concept anomaly detection example for spline\-coupling health monitoring\.
By using a convolutional[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1),\[wang2024Adigitaltwin\]estimate joint electrical current from multivariate sensor data \(e\.g\., motion, temperature, and vibration\) for anomaly detection in industrial robot systems\. Physical knowledge from energy conservation in the joints \(multi\-energy model\) and Newton\-Euler multi\-joint dynamics is incorporated via additional loss terms, enforcing energy\-flow consistency and torque coupling between joints\. Using the Kullback\-Leibler divergence between estimated and measured current as a health indicator, the proposed method detects injected motor and reducer faults on real factory robots with superior accuracy, outperforming several state\-of\-the\-art time\-series anomaly detection methods\.
#### 4\.5\.2Diagnosis
Table 14:All studies employinglearning biasto addressdiagnosis, listed in alphabetical order\.Learning Bias for DiagnosisReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[chao2025Physicsinformedneural\]Discharge\-pressure and internal leakage model, volumetric efficiency definitionNonlinear[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30), algebraic equationAxial piston pump\[dong2025Innovativefaultdiagnosis\]Mass and momentum conservation in fluid dynamics, valve and periodic boundary conditions[PDEs](https://arxiv.org/html/2608.10047#p6.31.31.31.31), algebraic equationsAxial piston pump\[huang2025Physicsinformedcausallearning\]Structural causal model for gearbox vibration dataAlgebraic equationPlanetary gearbox, wind turbine gearbox\[li2024Hybridphysicsembeddedrecurrent\]Robot dynamic modelSecond\-order[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Industrial robot\[qiao2024APriorKnowledge\]Bearing fault frequenciesNumerical valuesBearing\[sun2024Contrastivelearningand\]4\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)multi\-body bearing dynamics modelSecond\-order[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)systemBearing\[tang2024Apriorknowledgeenhanced\]Demand for consistent time and frequency domain latent representationsInvariance lossBearing, gearbox\[xu2024Physicsinformedprobabilisticdeep\]Bearing fault frequenciesAlgebraic equationsBearing\[zhu2024PhysiCausalNet:ACausaland\]Bearing dynamic model, domain\-invariant features per fault caseSecond\-order[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30), progressive consistency causal factorization lossBearingNine studies employing learning bias for diagnosis have been identified, as summarized in Table[14](https://arxiv.org/html/2608.10047#S4.T14)\.\[qiao2024APriorKnowledge\]focus on the small\-sample problem in bearing fault diagnosis under variable operating conditions\. A 1D\-[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)with a sequential temporal attention module is trained within a contrastive learning framework to obtain discriminative fault representations from vibration signals\. Prior knowledge enters as analytically derived fault characteristic frequencies, which are provided as additional targets and predicted from the learned embedding via an auxiliary fully connected head\. The loss between predicted and known characteristic frequencies, weighted within a composite loss alongside cross\-entropy and contrastive loss, softly enforces that the latent representation encodes these physically meaningful frequency features, thereby guiding the network toward fault\-relevant structure in the data\. The resulting model outperforms purely data\-driven baselines on two publicly available bearing datasets, with ablation studies confirming the benefit of the embedded learning bias, particularly in small\-sample scenarios\. To address distribution shifts arising from variations in bearing structure and operating conditions,\[zhu2024PhysiCausalNet:ACausaland\]propose a domain generalization method capable of extracting domain\-invariant features for fault diagnosis\. The model consists of a Fourier\-based low\-pass filtering module with learnable parameters, a[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\-based feature extractor and a fully connected classifier\. Two kinds of regularization terms are incorporated: a dynamic embedding loss to guide features that account for the state of the target machine’s bearing dynamics and a progressive consistency causal factorization loss\. The latter guides correlation for features of the same fault case across different machines and operating conditions, while discouraging correlation between features of different fault cases\. Therefore, domain\-invariant features can be realized\. Experiments show that the proposed approach demonstrates superior performance in terms of accuracy, generalizability, and interpretability in comparison to common domain generalization methods\.\[xu2024Physicsinformedprobabilisticdeep\]also incorporate fault characteristic frequencies of bearings into their approach\. They use a dual\-branch[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1)to reconstruct bearing vibration data in the frequency domain for both real and imaginary parts of the spectrum\. A modified loss function that makes use of masked target data is used for reconstruction\. The mask highlights bands relevant to fault frequencies\. Thus, a robust latent space that is biased toward fault frequencies relevant to fault cases is obtained as input for a subsequent classifier\. Their approach outperforms state\-of\-the\-art methods in terms of accuracy and well\-separated feature spaces\. To effectively tackle label\-free fault diagnosis for rolling bearings, the proposed framework\[sun2024Contrastivelearningand\]combines a contrastive\-learning backbone with a dynamics\-embedding network based on sparse identification of nonlinear dynamics\[champion2019sindy\]: a coordinate encoder reconstructs 4\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)latent states from 1\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)acceleration via delay embedding, and a physics\-based equation library derived from a 4\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)multi\-body bearing model is used together with a sparse coefficient matrix to infer the fault type\. Physics\-based constraints are imposed via the loss function, which enforces consistency between latent accelerations, reconstructed measurement signals and the accelerations predicted by the equation library, enabling the network to learn both the latent dynamics and a sparse, interpretable governing equation from raw signals\. Experiments on both simulated and experimental bearing data show that the proposed framework can correctly distinguish between inner\-race, outer\-race and roller faults without labels while providing physically meaningful diagnostic explanations\.
\[tang2024Apriorknowledgeenhanced\]propose a self\-supervised learning framework that addresses the small\-sample problem in rotating machinery fault diagnosis, studying both bearings and gearboxes\. Two[CNNs](https://arxiv.org/html/2608.10047#p6.8.8.8.8)are used as feature extractors for the time domain and frequency domain, respectively\. During pretraining, an additional loss function incorporating a distance metric is employed to obtain time\-frequency domain invariant latent embeddings\. In the downstream task, the network is fine\-tuned on the limited amount of labeled data to learn to diagnose the fault underlying the rotating part, showing superiority over alternative approaches both in terms of accuracy and data efficiency\. Aiming to incorporate prior knowledge into domain generalization methods,\[huang2025Physicsinformedcausallearning\]propose a causal learning network based on ResNet18\. It is used to extract independent and causal features regarding the relation between the sensor data of gearboxes and distinct fault cases for diagnosis\. Losses for an adversarial mask and an autocorrelation matrix, both founded on a structural causal model for vibration data, are incorporated to favor the aforementioned properties\. Compared to other domain generalization methods, the approach demonstrates superior performance in terms of accuracy and class\-distinguishability\.
The integration of model\-based knowledge into diagnostic methods proves challenging in the field of axial piston pumps\. Therefore,\[dong2025Innovativefaultdiagnosis\]present a[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)framework serving as a high\-frequency virtual dynamic flow meter in axial piston pumps by predicting pump flow ripple\. The framework integrates fundamental physical principles of hydraulic systems, such as mass and momentum conservation, boundary conditions relevant to pump operation, and periodic characteristics of pump behavior into the loss function\. The method facilitates robust pump fault diagnosis\. Overall, the study validates the effectiveness of the proposed approach through numerical simulations, demonstrating close agreement with reference solutions, and through experimental investigations, showing that predicted flow ripples consistently reflect expected fault characteristics\. Building on[PINNs](https://arxiv.org/html/2608.10047#p6.36.36.36.36),\[chao2025Physicsinformedneural\]tackle wear detection in axial piston pumps by reconstructing the discharge pressure while simultaneously inferring the fluid film thicknesses at the pump’s main friction pairs\. Following the standard[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)framework, an analytically derived[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)for the time derivative of discharge pressure is incorporated into the loss function, with internal leakage flows scaling cubically with film thickness\. To stabilize the joint estimation of multiple wear\-related parameters with different scales, these parameters are learned as bounded variables via sigmoid\-based range constraints\. In a sequential step, the identified film thicknesses are used in analytical formulas for volumetric efficiency and Cohen’sddeffect size, providing physically interpretable wear indicators and enabling localization of the worn friction pair\. Experiments on a real pump with naturally worn components demonstrate accurate pressure reconstruction, physically plausible thickness estimates, and correct identification of the valve plate and cylinder block pair as the worn pair, although comparisons with alternative methods are not reported\.
Using available proprioceptive signals instead of external measurements,\[li2024Hybridphysicsembeddedrecurrent\]propose an approach for fault diagnosis of industrial robots\. The framework consists of an encoder based on[GRU](https://arxiv.org/html/2608.10047#p6.22.22.22.22)and fully connected layers, while the decoder combines a robot dynamic model and data\-driven residual model\. The decoder is solely used in training for reconstruction purposes, while a classifier leverages the latent variable for fault diagnosis both in training and inference\. Fault cases of increased joint friction, partial loss of actuator effectiveness, and drivetrain mechanical faults are considered\. In ablation studies, the approach performed superiorly to non\-[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)variants, evaluated with both a simulated UR5 dataset and a real industrial robot in\-situ dataset\.
#### 4\.5\.3Health Assessment
Table 15:All studies employinglearning biasto addresshealth assessment, listed in alphabetical order\.Learning Bias for Health AssessmentReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[deng2025ANovelMethod\]Empirical aging trendMonotonicity constraint \(loss term\)Lithium\-ion battery\[freeman2022Physicsinformedturbulenceintensity\]Turbulence intensity as analytical and empirical definitionAlgebraic equationOcean current turbines\[jang2025Stateofhealth\]Lumped energy\-balance model for battery heat generationFirst\-order[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Lithium\-ion battery\[li2025Applicationofphysicsguided\]Empirical flank tool wear modelAlgebraic equationMilling\[liu2025Aphysicsguidedapproach\]Mechanistic[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)\-growth capacity\-fade modelNonlinear[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Lithium\-ion battery\[navidi2024PhysicsInformedMachineLearning\]Half\-cell degradation modelAlgebraic equationsLithium\-ion battery\[pan2025Inservicefatiguecrack\]Paris’ law[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Fatigue crack propagation \(aluminum specimens\)\[singh2023HybridModelingof\]Fick’s second law of diffusion, initial/boundary conditions[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31), algebraic equationsLithium\-ion battery\[sun2022MicrocrackDefectQuantification\]Analytical relationships between signal\-derived features and crack geometriesClosed\-form algebraic equations and inequality constraintsMicrocrack quantification in aluminum plate\[yonastefera2025ConstraintGuidedLearningof\]Monotonicity assumption, boundary constraintAlgebraic equationBearing\[wang2025ABatteryState\]Semi\-empirical Verhulst degradation model with Arrhenius temperature factorNonlinear logistic\-type[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Lithium\-ion battery\[wangz\.2023Physicsinformedneural\]Discharge pressure model describing leakageFirst\-order[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Axial piston pump\[wang2024Physicalknowledgeguided\]Continuous degradation\-trendRank\-N\-contrast lossLithium\-ion battery\[wang2024Physicsinformedneuralnetwork\]Monotonic degradation, multivariate degradation trendRegularization term, learnable dynamic model for[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)lossLithium\-ion battery\[wang2025PhysicsInformedNeuralNetwork\]Multivariate degradation trendLearnable dynamic model for[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)lossLithium\-ion battery\[pengfeiwen2023PhysicsInformedNeuralNetworks\]Semi\-empirical Verhulst degradation modelNonlinear logistic\-type[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Lithium\-ion battery\[xu2024PhysicsConstraintVariationalNeural\]Pressure pulsation response modelAlgebraic equationGear pump\[zhang2025AnElectrochemicalAgingInformed\]Electrochemical aging modelCoupled[ODEs](https://arxiv.org/html/2608.10047#p6.30.30.30.30), algebraic equationsLithium\-ion batteryAn overview of all studies \(18\) employing learning bias for health assessment is provided in Table[15](https://arxiv.org/html/2608.10047#S4.T15)\. Focusing on[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation in lithium\-ion batteries,\[deng2025ANovelMethod\]train an[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)on features extracted from charge, discharge, and incremental\-capacity curves\. An additional loss term penalizes deviations from a monotonic relationship between the peak of the incremental\-capacity curve and[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46), reflecting their consistently one\-directional trend over aging\. Trained on two public aging datasets with different chemistries and operating conditions, the resulting model achieves lower[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation errors than a standard[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)and a[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\. The challenge of accurately estimating the[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)of lithium\-ion batteries under dynamic operating conditions is studied by\[wang2024Physicalknowledgeguided\]\. Building on[ResNet](https://arxiv.org/html/2608.10047#p6.38.38.38.38), the authors use a constraint that guides relative distances and ranking between the embedding space and the output space \([SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)\) as prior knowledge\. This battery degradation property was integrated by means of the Rank\-N\-Contrast loss\. Validated on two datasets, the approach demonstrates superior performance in terms of predictive accuracy and structured latent representations compared to traditional and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\-based methods\. Additional aspects of this work that fall into the class of hybrid approaches are reported in Section[4\.6\.3](https://arxiv.org/html/2608.10047#S4.SS6.SSS3)\.
\[navidi2024PhysicsInformedMachineLearning\]compare four approaches for estimating the capacity \([SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)\) and three internal degradation modes of lithium\-ion batteries\. As a follow\-up study to\[thelen2022Integratingphysicsbasedmodeling\], it includes three approaches that incorporate observational bias, which are described in Section[4\.3\.3](https://arxiv.org/html/2608.10047#S4.SS3.SSS3)\. Furthermore, the authors introduce an approach in which a shallow[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)predicts half\-cell model parameters\. A differentiable surrogate of the half\-cell model is embedded via additional loss terms that weakly enforces consistency of the predicted parameters and the degradation behavior implied by the half\-cell model\. Trained on early\-life experimental data together with simulation data, the regularized network outperforms an otherwise identical purely data\-driven network and all approaches incorporating observational bias\. Explicitly accounting for electrochemical parameter inconsistencies across cells,\[zhang2025AnElectrochemicalAgingInformed\]present an approach to estimating the[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)of lithium\-ion batteries\. Dedicated subnetworks jointly estimate lithium\-ion concentration dynamics and cell\-specific electrochemical parameters by processing initial\-state features from an early\-life discharge to encode parameter variability, and a capacity\-difference sequence and sampled time coordinates to capture degradation behavior\. The estimated internal states and parameters are then passed through a reduced electrochemical aging model \(enhanced[Single\-Particle Model](https://arxiv.org/html/2608.10047#p6.47.47.47.47)\([SPM](https://arxiv.org/html/2608.10047#p6.47.47.47.47)\) with polynomial solid\-phase diffusion, Butler\-Volmer kinetics, Padé\-approximated electrolyte diffusion, and electrode stoichiometry shifts\) to reconstruct terminal voltage and capacity, with its governing equations and boundary conditions enforced via additional loss terms during training\. Across multiple battery chemistries and operating conditions, the method attains very low[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)\-prediction errors even with scarce training data and supports fast inference, outperforming baselines such as[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36),[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8), and enhanced[SPM](https://arxiv.org/html/2608.10047#p6.47.47.47.47), though at the cost of higher training complexity\.
Figure 7:A representative example of learning\-bias approaches is proposed by\[singh2023HybridModelingof\], demonstrating how prior physical knowledge can be introduced in the loss function\. Own illustration based on the corresponding study\. See the original work for full technical details\.\[singh2023HybridModelingof\]address the joint estimation of[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45)and[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)for lithium\-ion cells operating under varying temperatures and limited measurement data \(see Fig\.[7](https://arxiv.org/html/2608.10047#S4.F7)\)\. Their method incorporates the[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)governing solid\-phase lithium diffusion from an[SPM](https://arxiv.org/html/2608.10047#p6.47.47.47.47), along with Neumann flux boundary conditions driven by the applied current, into the loss function\. To balance the data and physics\-informed loss terms, the authors employ gradient normalization for adaptive loss balancing based on the work of\[chen2018gradnorm\], although no further details on its implementation are given\. The[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)predicts spatio\-temporal lithium concentration fields within the anode, from which SOC and capacity\-based SOH are inferred via established concentration\-to\-SOC and capacity relations\. Results on three cells cycled at different temperatures show that the method accurately tracks both[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45)and[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)over time\. Additionally, an ablation study on one cell, where the model is trained on early\-life check\-up cycles and then applied to subsequent cycles, provides a limited demonstration of short\-horizon[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)prognosis\. Given that the heat generation rate varies significantly with the[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46), a[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)is adopted by\[jang2025Stateofhealth\], incorporating a lumped energy\-balance equation into its loss function to infer both temperature and the time\-varying heat generation rate in a manner consistent with battery thermodynamics\. Subsequently, three methods are proposed that leverage the resulting[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)\-derived profiles to estimate the[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46): an[MLP](https://arxiv.org/html/2608.10047#p6.28.28.28.28), a 1D\-[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8), and an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\-[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\-based approach\. However, the evaluation is limited to comparisons among these three methods, leaving their comparative performance relative to established electrical\-based or data\-driven[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation approaches unclear\.\[pengfeiwen2023PhysicsInformedNeuralNetworks\]also adopt a[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)\-based approach by embedding a semi\-empirical Verhulst degradation model\[xian2013prognosticsverhulstmodel\]into the loss function of an[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)for lithium\-ion battery[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation\. An uncertainty\-based weighting scheme adaptively balances the losses during training, allowing the[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)to exploit the Verhulst model without manual tuning of loss weights\. Experiments demonstrate that this approach achieves lower[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation errors compared to both a purely data\-driven counterpart and a[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)\-based method\. Adopting a similar approach,\[wang2025ABatteryState\]leverage the same Verhulst model\[xian2013prognosticsverhulstmodel\]\(while additionally integrating the Arrhenius equation\) as a soft constraint in the loss function, resulting in a[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)\-based approach\.[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)labels are derived from multi\-stage constant\-current charging via an incremental capacity\-based method\. Driving\-style\-related features, together with key operating variables, are used by the[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)to estimate the[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)of lithium\-ion batteries in electric vehicles\. Compared with both purely data\-driven baselines and alternative[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)\-based approaches, the proposed model achieves the lowest[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation errors\. Addressing[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation for lithium\-ion batteries via capacity loss,\[liu2025Aphysicsguidedapproach\]aim to obtain accurate and physically plausible degradation trajectories from charge\-discharge data\. They first compute entropy\-based health indicators from voltage and current, and use these features together with cycle count as inputs to a feedforward[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)\. A mechanistic[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)\-growth model for capacity fade is then embedded as an additional physics\-based loss term that penalizes mismatches between the network’s time derivative of capacity loss and the[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)model, thereby regularizing the network toward smooth, mechanistically consistent[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)evolution and suppressing non\-physical fluctuations\. Across three datasets with different chemistries and cycling regimes, this physics\-regularized model achieves substantially lower errors than purely data\-driven[MLP](https://arxiv.org/html/2608.10047#p6.28.28.28.28)\- and[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\-based models\. Also covering[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation for lithium\-ion batteries,\[wang2024Physicsinformedneuralnetwork\]employ a[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)\-based approach, in which one[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)estimates the[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)based on the current cycle and features generated from sensor readings in the charge phase, and a second[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)estimates the battery degradation dynamic behavior solely needed for loss construction\. From a[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)perspective, besides a monotonicity loss, a[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)\-based loss is also integrated\. The latter accounts for the requirement that the degradation trajectory depends on charging rate, discharging rate, temperature, etc\., rather than being merely a time\-dependent univariate function\. In comparison with an[MLP](https://arxiv.org/html/2608.10047#p6.28.28.28.28)and a[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8), the proposed method performs better in terms of predictive performance, data efficiency and generalizability, as demonstrated in experiments covering[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)scenarios and data\-sparse scenarios\. Except for the monotonicity loss, the main aspects of this work are also adopted by\[wang2025PhysicsInformedNeuralNetwork\], where they are specialized for 2\-minute data segments obtained from laboratory experiments simulating satellite batteries\.
\[yonastefera2025ConstraintGuidedLearningof\]address the problem of loss balancing in regularization approaches\. Considering a convolutional[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1), the latent variable is fed into multiple fully connected layers for extraction of a[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)suitable for bearing health assessment\. Three constraints are directly integrated into the gradient descent algorithm instead of loss terms\. Besides a monotonicity constraint, a boundary constraint leading to a normalized[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)and a constraint that guides consistency between the signal energy and the[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)are employed\. The approach shows better predictive performance in the majority of experiments compared to other convolutional[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1)baselines\.
In this study,\[li2025Applicationofphysicsguided\]leverage an empirical tool flank wear model relating wear to milling time as prior knowledge for health assessment in high\-speed milling\. This knowledge is incorporated via a loss function that penalizes deviations between the model predictions and the empirical model during the training of a[GRU](https://arxiv.org/html/2608.10047#p6.22.22.22.22)\. The regularization enforces physically consistent degradation behavior during training\. As a result, the model achieves improved physical consistency while maintaining low prediction errors compared to baselines\. See Section[4\.4\.3](https://arxiv.org/html/2608.10047#S4.SS4.SSS3)for a description of the inductive bias incorporated by\[li2025Applicationofphysicsguided\]\.
\[wangz\.2023Physicsinformedneural\]incorporate a discharge pressure model for axial piston pumps as a dynamic loss into a[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)\. Hence, geometrical parameters related to leakage that represent the current health state of the piston\-cylinder interfaces can be fitted\. Determining these parameters is challenging using traditional methods\. As the leakage of the aforementioned intersections can be quantified for multiple piston\-cylinder pairs independently, this can also be considered a diagnosis approach\. Studying gear pumps,\[xu2024PhysicsConstraintVariationalNeural\]tackle the black\-box nature of[DL](https://arxiv.org/html/2608.10047#p6.10.10.10.10)models\. The output of a compound[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)\(consisting of[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8),[BiLSTM](https://arxiv.org/html/2608.10047#p6.2.2.2.2)and attention\) is guided to match the parameters of a physical pressure model regarding the outlet pressure pulsation\. The physics\-based model can be used to reconstruct the input signal to perform unsupervised training\. The estimated physical parameters are used to construct a health indicator based on distance metrics\. The proposed method is superior to purely data\-driven approaches in terms of interpretability \(the authors show this qualitatively by means of t\-distributed stochastic neighbor embedding\), though not necessarily in terms of reconstruction capabilities\.
\[pan2025Inservicefatiguecrack\]address both the small\-sample problem and the poor generalization ability in[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\-based fatigue crack quantification\. Using aluminum specimens representative of aircraft structures, the authors model fatigue crack propagation with an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)network\. A Paris law\-based crack\-growth model is utilized to formulate an additional loss term aiming to improve the predictive accuracy for crack growth under complex environmental conditions\. The positive contribution of the incorporated physics to the model’s overall performance is demonstrated through dedicated ablation studies\. With the incorporated observational bias described in Section[4\.3\.3](https://arxiv.org/html/2608.10047#S4.SS3.SSS3),\[sun2022MicrocrackDefectQuantification\]also introduce a learning bias by augmenting the training objective with additional physics\-based loss terms that penalize inconsistencies between the network outputs and analytical relationships derived from guided\-wave scattering theory\. For crack length, an additional loss term penalizes pairs of samples whose predicted lengths violate the required monotonic relationship between a width\-like feature \(constructed from neighboring reflection amplitudes\) and the angular spread of the reflected lobe\. For depth and direction, further penalty terms enforce analytical formulas that link the ratio of reflected to transmitted energy and the corrected reflection angle, respectively, to the corresponding crack parameters given the current length prediction\. A single weighting factor controls the influence of all physics\-based penalties; sensitivity studies show that choosing this weight appropriately yields substantially lower quantification errors than training without these constraints\.
The method of\[freeman2022Physicsinformedturbulenceintensity\]enables the classification of rotor blade pitch imbalance faults in ocean current turbines according to their level of severity\. An[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)pipeline with the single\-phase power output of the turbine as input consists of principal component analysis and multinomial logistic regression for flow\-speed classification, an[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)with an augmented loss for turbulence\-intensity classification and a final[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)for fault severity classification\. Both an empirical and an analytical expression of turbulence intensity are used to define a monotonicity constraint as an additional loss\. In ablation studies, the proposed method outperforms its purely data\-driven counterpart in terms of predictive performance\.
#### 4\.5\.4Prognosis
Table 16:All studies employinglearning biasto addressprognosis, listed in alphabetical order\.Learning Bias for PrognosisReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[badora2023Usingphysicsinformedneural\]Fracture\-mechanics relation between load ratio of thermal stresses and corresponding stress intensity factorsAlgebraic ratio constraintHigh\-pressure nozzle of an industrial gas turbine\[e2025Aphysicsinformedneural\]Empirical aging law relating[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)to cycle numberAlgebraic equationSupercapacitors\[fassi2024PhysicsInformedMachineLearning\]Empirical aging trendMonotonicity and boundedness constraints \(loss terms\)Metal\-oxide\-semiconductor field\-effect transistor\[he2025Physicsinformedneuralnetwork\]Impedance\-based degradation mechanism, stochastic degradation behavior \(Wiener process\)Intermediate physical variables, Gaussian increment modelLithium\-ion battery\[najeraflores2023APhysicsConstrainedBayesian\]Empirical aging trendMonotonicity constraint \(loss term\)Lithium\-ion battery\[pugalenthi2024RemainingUsefulLife\]Semi\-empirical[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)\-film\-based capacity fade modelTwo\-term exponential equationLithium\-ion battery\[ramirez2024ResidualbasedAttentionPhysicsinformed\]Heat\-diffusion model of transformer oil, standard thermal\-aging model for winding insulation1D heat\-diffusion[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31), algebraic equationsTransformer\[wang2024Phyformer:Adegradation\]Monotonicity assumptionSingle\-term exponential functionBearing, transformer, electromechanical servo system\[wang2025KoopmanInformedNeuralNetwork\]Koopman operator theoryLinear operator with algebraic consistency lossTurbofan engine, bearing\[wang2025Aremaininguseful\]Reliability model of bearing failure processWeibull cumulative distribution functionBearing\[xu2022Aphysicsinformeddynamic\]First\-order Thevenin model[ODEs](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Lithium\-ion battery\[zhang2025APhysicsInformedHybrid\]Enhanced[SPM](https://arxiv.org/html/2608.10047#p6.47.47.47.47)[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)system with algebraic voltage relationLithium\-ion battery\[zhu2024RemainingUsefulLife\]Inverse monotonic relationship between crack surface area and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)Monotonicity constraint \(loss term\)BearingThirteen studies employ learning bias for prognosis \(see Tab\.[16](https://arxiv.org/html/2608.10047#S4.T16)\), several of which focus on lithium\-ion battery prognostics\.[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)prognosis of lithium\-ion batteries is addressed by\[xu2022Aphysicsinformeddynamic\]\. A[ResNet](https://arxiv.org/html/2608.10047#p6.38.38.38.38)\-based encoder\-decoder model uses capacity and secondary variables \(i\.e\., temperature, voltage, and current\) from the current cycle to predict capacity and full secondary\-variable profiles for the subsequent cycle\. This one\-step\-ahead predictor is iterated from the first discharge cycle to generate long\-horizon degradation trajectories over the battery’s life\. Physics is incorporated through additional loss terms that penalize violations of a first\-order Thevenin\-based state equation and a capacity\-balance equation, which are combined into an[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)system\. The proposed method achieves substantially lower[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)prediction errors than several[GP](https://arxiv.org/html/2608.10047#p6.20.20.20.20)\-based baselines using only the first discharge cycle, with ablation studies attributing the performance gains to the added regularization based on the underlying battery model\.\[pugalenthi2024RemainingUsefulLife\]present a prognostic framework for lithium\-ion batteries, where an[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)is first identified from a single run\-to\-failure cell using a[PF](https://arxiv.org/html/2608.10047#p6.33.33.33.33)and then used to predict the capacity trajectories of other cells\. A semi\-empirical[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)\-film\-based capacity fade model is incorporated as an additional loss term to improve physical plausibility and reduce errors in predicted capacity, especially when only limited training data are available\. The trade\-off is higher computational cost due to the extra overhead of evaluating the physics\-based loss on top of the[PF](https://arxiv.org/html/2608.10047#p6.33.33.33.33)\-based parameter estimation\.\[he2025Physicsinformedneuralnetwork\]propose a degradation modeling framework that combines[DL](https://arxiv.org/html/2608.10047#p6.10.10.10.10)with a Wiener process for lithium\-ion batteries\. The network takes degradation features \(i\.e\., constant\-current charging time, incremental\-capacity peak, and temperature peak\) and maps them through intermediate physical variables that represent impedance\-related quantities, before producing a nonlinear degradation path that serves as the drift of the Wiener process\. A composite loss combining mean squared error on these intermediate physical variables with a likelihood\-based term for degradation increments is used to train both the network and stochastic\-process parameters jointly from historical data\. The latter are further updated online via Bayesian inference using real\-time[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)measurements\. Ablation and comparative studies on two battery datasets show that adding the impedance\-informed latent layer and the Wiener process module improves not only[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)trajectory prediction but also the fidelity of the resulting reliability curves and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)and lifetime distributions, compared with purely data\-driven and purely stochastic\-process baselines\. To effectively target[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction of lithium\-ion batteries,\[zhang2025APhysicsInformedHybrid\]propose a two\-stage approach\. In the first stage, an electrochemical\-informed generative model, constrained by a reduced\-order enhanced[SPM](https://arxiv.org/html/2608.10047#p6.47.47.47.47), reconstructs electrode\-level states\. The model is trained with a composite loss over the electrochemical governing equations, initial and boundary conditions, and terminal\-voltage mismatch, whose weights are adaptively balanced via gradient\-norm ratios\. In the second stage, incremental\-capacity and differential\-voltage curves derived from the reconstructed states serve as electrode\-level features\. These are combined with cell\-level features \(capacity\-difference curves and charging protocols\) in an[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29): two[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)branches encode the cell\-level inputs, while a[GRU](https://arxiv.org/html/2608.10047#p6.22.22.22.22)with self\-attention encodes the electrode\-level inputs, and the concatenated representations are decoded to predict the[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)\. Across four datasets with different chemistries and operating conditions, the method outperforms both mechanistic and purely data\-driven baselines in terms of prediction error, shows better robustness with limited training data, and offers faster inference, while also enabling identification of electrode\-level degradation modes\. Early\-stage[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction is studied by\[najeraflores2023APhysicsConstrainedBayesian\]\. To this end, a neural differential operator is learned for the discharge capacity rate from early\-life cycling data\. The proposed architecture draws inspiration from DeepONet\[lu2021learning\], featuring a Bayesian branch network that encodes cell\-specific early\-life features and a deterministic trunk network that encodes time\. Multiple loss terms are employed to promote accurate modeling of the discharge capacity, including a monotonicity constraint that weakly enforces negative self\-acceleration of the capacity trajectory\. At inference, the learned operator is sampled from the Bayesian branch, integrated forward in time to reconstruct the capacity curve, and the[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)is obtained as the difference between the predicted end\-of\-life time \(at a capacity threshold\) and the current cycle\. Experiments show that, given the same early\-life training data, the proposed physics\-constrained Bayesian operator achieves smaller[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction errors and more reliable uncertainty quantification than a simplified physics\-based failure forecast model, a[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)\-based method, and a similarity\-based method\.
\[wang2025Aremaininguseful\]propose an approach for bearing[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction using acoustic emission signals\. A novel health indicator is introduced to quantify acoustic emission signal complexity and is shown to outperform standard time\-domain and entropy features in monotonicity, robustness, and trendability\. This health indicator is fed to an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)whose loss combines mean squared error with a Weibull\-based term derived from reliability engineering, thereby constraining the learned degradation trajectory to be consistent with the expected failure behavior\. Experiments demonstrate that the proposed method yields substantially lower[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction errors than a conventional[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\. Already discussed in Section[4\.3\.4](https://arxiv.org/html/2608.10047#S4.SS3.SSS4),\[zhu2024RemainingUsefulLife\]propose three mechanisms for incorporating physics to tackle bearing[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction\. In terms of learning bias, the authors design an inconsistency loss that penalizes pairs of predicted[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)values violating the inverse monotonic relation between crack surface area and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)—imposed via a[ReLU](https://arxiv.org/html/2608.10047#p6.37.37.37.37)\-based penalty\. Again, the complete framework proves superior, but without ablation experiments, the specific contribution of the additional loss term cannot be quantified\. See Section[4\.6\.4](https://arxiv.org/html/2608.10047#S4.SS6.SSS4)for a synopsis focusing on the hybrid aspect of the framework\.
To facilitate prognostics of complex industrial machinery,\[wang2025KoopmanInformedNeuralNetwork\]present a novel approach grounded in Koopman operator theory\. The proposed Koopman\-informed[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)enables accurate[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction by learning nonlinear system dynamics through a linear representation in a latent eigenfunction space\. Within an encoder\-decoder architecture, the encoder maps data spanning the entire operational lifespan into an eigenfunction space, where the system’s evolution is approximated by a finite\-dimensional Koopman matrix: multi\-step temporal forecasting is performed via repeated applications of the learned forward and backward Koopman operators\. The decoder reconstructs future system states from the propagated latent representation\. Multiple loss terms \(including a consistency loss that penalizes discrepancies between forward and backward evolution\) regularize the learned dynamics and promote stable, physically coherent temporal behavior\. Lastly, a nonlinear regression head leverages the learned high\-level features to produce precise[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)estimates\. The proposed Koopman\-informed[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)is evaluated on both a bearing and a turbofan engine dataset, where the model consistently outperforms several strong[DL](https://arxiv.org/html/2608.10047#p6.10.10.10.10)baselines in terms of predictive performance\.
\[wang2024Phyformer:Adegradation\]propose a Transformer\-based prognostics model designed for scenarios lacking reliable degradation physics\. The approach hinges on decomposing time\-series data into a slowly varying trend component \(capturing the overall degradation trend\) and a residual component \(capturing higher\-frequency fluctuations\)\. Thereafter, a simple monotonic parametric curve is fitted to the trend component in a sliding\-window procedure\. Rather than constructing accurate, domain\-specific degradation models, these local fits \(here, single\-term exponentials\) are treated as“simple, general and imperfect”approximations\. An additional loss term penalizes large deviations from these fitted curves, thereby regularizing the model toward predictions that reflect the assumed irreversibility of degradation\. Given its domain\-agnostic prior knowledge, the model is applied to three different use cases \(bearings, transformers, and an electromechanical servo system\), where it consistently reduces long\-horizon prediction errors compared to several state\-of\-the\-art[DL](https://arxiv.org/html/2608.10047#p6.10.10.10.10)models\.
Given that failure of power semiconductor devices poses a significant reliability challenge for power converter systems,\[fassi2024PhysicsInformedMachineLearning\]address[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction for power metal\-oxide\-semiconductor field\-effect transistors under thermal aging\. To improve predictive accuracy and physical consistency, the authors incorporate additional loss terms that weakly enforce monotonic, bounded[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)trajectories via[ReLU](https://arxiv.org/html/2608.10047#p6.37.37.37.37)\-based penalties\. A series of experiments across various recurrent network architectures demonstrates that it achieves lower or comparable mean squared error than purely data\-driven counterparts, while simultaneously enabling faster convergence\.
\[e2025Aphysicsinformedneural\]present a method for predicting the[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)of commercial supercapacitors by embedding an empirical aging law that models[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)as a logarithmic function of cycle number into the loss function of an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\. A scalar weighting factor between the data and physics losses is tuned via Bayesian optimization, leading to stronger regularization under scarce data and reduced reliance on the physical loss term as more data become available\. Experiments demonstrate that the proposed model substantially outperforms its purely data\-driven counterpart, with prediction errors comparable to more advanced data\-driven methods while using only a fraction of the full life cycle as training data\.
In addition to the inductive bias reported in Section[4\.4\.4](https://arxiv.org/html/2608.10047#S4.SS4.SSS4),\[badora2023Usingphysicsinformedneural\]incorporate a learning bias independent of the former\. An additional loss term regularizes training with respect to the fracture\-mechanics load\-ratio relation between applied thermal stresses and the corresponding stress intensity factors, using a large set of synthetically generated stress\-crack\-length combinations\. A dynamically adjusted weighting factor gradually shifts emphasis from this physics term toward the empirical crack\-length error as training progresses, ensuring that the final model both respects the underlying fracture mechanics and fits the sparse inspection data\. While the need for a dynamic weighting between physics and empirical loss terms is well motivated, the specific piecewise schedule for the weighting coefficient is introduced without justification, making this part of the approach heuristic and potentially hard to generalize or reproduce\. Combined with the inductive encoding of Paris’ law, the proposed method effectively targets fatigue crack growth modeling in a gas turbine nozzle\.
\[ramirez2024ResidualbasedAttentionPhysicsinformed\]propose an efficient spatio\-temporal model for predicting transformer winding temperature, including the local hotspot, and insulation aging\. Embedding a simplified one\-dimensional heat\-diffusion[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)with uniform heating into a[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)enables estimating oil temperatures, which are then used to calculate winding temperatures, aging acceleration factors, and the associated loss of life\. Tested on a distribution transformer in a floating photovoltaic power plant, the method closely matches numerical[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)solutions and improves hotspot and aging estimation compared to a standard analytic hotspot model, with validation against fiber optic sensor measurements\. Additionally, a residual\-based attention scheme improves convergence and training stability of the[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)\.
### 4\.6Hybrid Approaches
Hybrid approaches are characterized by combining independent physics\-based and data\-driven models, either in parallel or in series\. The identified studies are organized by[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)task in Tables[17](https://arxiv.org/html/2608.10047#S4.T17)–[20](https://arxiv.org/html/2608.10047#S4.T20)\. A representative example is illustrated in Figure[8](https://arxiv.org/html/2608.10047#S4.F8)\.
Figure 8:A representative example for hybrid approaches is proposed by\[firoozi2022CylindricalBatteryFault\], demonstrating how physics\-based and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models can be combined\. Own illustration based on the corresponding study\. For brevity, time indices are omitted, with the prime denoting the next state\. See the original work for full technical details\.#### 4\.6\.1Fault Detection
Table 17:All studies employinghybrid approachesto addressfault detection, listed in alphabetical order\.Two studies employing hybrid approaches for fault detection have been identified \(see Tab\.[17](https://arxiv.org/html/2608.10047#S4.T17)\), both targeting lithium\-ion batteries\.\[firoozi2022CylindricalBatteryFault\]target real\-time fault detection in cylindrical lithium\-ion batteries under extreme fast charging, aiming for early detection of voltage and thermal faults \(see Fig\.[8](https://arxiv.org/html/2608.10047#S4.F8)\)\. Two detection observers are built on an experimentally identified reduced\-order electrochemical\-thermal model; a[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)is used in parallel to learn additive voltage and temperature uncertainty terms from residuals between model predictions and measurements in no\-fault cycles\. The learned uncertainty corrections are fed back into the observers to suppress non\-fault deviations, and residuals are evaluated against calibrated thresholds to detect voltage and thermal faults\. Comparative results with a model\-only observer indicate that incorporating the[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)can enable detection of smaller faults that the purely physics\-based counterpart misses\. Building on a similar concept,\[zhang2024Adaptivefaultdetection\]\(who also reference\[firoozi2022CylindricalBatteryFault\]as related work\) propose an adaptive fault detection framework for lithium\-ion batteries that combines a thermoelectric model\-based[EKF](https://arxiv.org/html/2608.10047#p6.14.14.14.14)observer with a[BiLSTM](https://arxiv.org/html/2608.10047#p6.2.2.2.2)\. A second\-order[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13)coupled with a simplified thermal model is used in the[EKF](https://arxiv.org/html/2608.10047#p6.14.14.14.14)to estimate voltage, temperature, and[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45), while the[BiLSTM](https://arxiv.org/html/2608.10047#p6.2.2.2.2)is trained on healthy data to learn the voltage observation error and subsequently compensates the observer to suppress uncertainty\-induced residuals\. The corrected residuals are compared against a threshold calibrated from healthy operation to detect soft internal short circuit faults\. Experiments on driving\-cycle data show that the proposed framework yields substantially lower voltage estimation errors than the standalone[EKF](https://arxiv.org/html/2608.10047#p6.14.14.14.14), thereby enabling more reliable fault detection with fewer false alarms\.
#### 4\.6\.2Diagnosis
Table 18:All studies employinghybrid approachesto addressdiagnosis, listed in alphabetical order\.Hybrid Approaches for DiagnosisReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[pettorossi2025Physicsguidedfaultdiagnosis\]Proton exchange membrane fuel cell simulation[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)systemFuel cell\[singh2024Hybridphysicsinfused1DCNN\]0D high\-fidelity physics\-based engine modelCoupled algebraic equations,[ODEs](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Diesel engine\[xu2023Physicsguideddatarefinedfault\]Fault hierarchies and statistical fault evolution/propagation mechanisms \(time\-to\-failure behavior and correlations between failure modes\)Probability density functions and cumulative distributions of failure modes, algebraic update formulasOffshore wind turbineTable[18](https://arxiv.org/html/2608.10047#S4.T18)lists the three studies that employ hybrid approaches for diagnosis\.\[pettorossi2025Physicsguidedfaultdiagnosis\]address fault diagnosis for proton exchange membrane fuel cells, focusing on identifying and isolating four different fault types: flooding, drying, air starvation, and hydrogen starvation\. To tackle this, the authors propose an approach that combines a physics\-based proton exchange membrane fuel cell model with an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\. The physics\-based model provides estimates of unmeasured process variables, including membrane resistance, water content, and current density distribution\. These variables are combined with measured stack signals and fed to the[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26), both in training and inference\. Hence, more informed and robust fault classification is enabled\. This integration improves diagnostic accuracy, reduces detection time, and enhances generalizability compared to purely data\-driven approaches\.
\[singh2024Hybridphysicsinfused1DCNN\]propose a fault diagnosis framework for a diesel engine that combines a 0D \(lumped, time\-dependent\) high\-fidelity physics\-based engine model with an[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1)and a 1D\-[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)\. The[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1)compresses high\-dimensional sensor data into a latent feature vector, while the physics\-based engine model provides additional simulated variables\. These are then concatenated and passed to a 1D\-[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)that classifies four conditions: nominal operation, and faults due to injection pressure, injection duration, and start of injection\. The engine model thus acts as an in\-parallel source of physics\-based features that complement the data\-driven latent representation for fault classification\. Experiments on test\-bed data show that the proposed model achieves higher diagnostic accuracy than a purely data\-driven 1D\-[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)and exhibits improved robustness to sensor noise and to extrapolation across unseen engine speeds and operating conditions, with notably fewer false positives in nominal conditions\.
Fault root cause tracing in complex electromechanical systems is studied by\[xu2023Physicsguideddatarefinedfault\], demonstrated on an offshore wind turbine experiencing an unscheduled power drop\. Common physics\- and statistics\-based knowledge about fault mechanisms is first encoded into a static hierarchical fault root cause tracing network—a probabilistic graph whose nodes and edges represent functional units, fault modes, and their propagation relationships\. This model is then refined with operation data: anomalies detected via a Wasserstein[GAN](https://arxiv.org/html/2608.10047#p6.18.18.18.18), together with statistical laws describing how faults evolve over operating time, update the weights of fault nodes and edges\. Finally, a bidirectional probabilistic reasoning scheme combines forward fault propagation and backward tracing information across the hierarchy to rank fault nodes and identify the most likely root cause and fault paths\.
#### 4\.6\.3Health Assessment
Table[19](https://arxiv.org/html/2608.10047#S4.T19)summarizes the five studies that employ hybrid approaches for health assessment, all but one of which target lithium\-ion batteries\. Addressing[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation,\[feng2024comprehensive\]leverage a second\-order resistor\-capacitor[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13), whose parameters are identified from post\-charge relaxation voltage via nonlinear least squares and then used as inputs to various[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models, including[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21), XGBoost,[SVR](https://arxiv.org/html/2608.10047#p6.48.48.48.48), elastic net, and a simple[MLP](https://arxiv.org/html/2608.10047#p6.28.28.28.28)\. For each regressor, the hybrid approach is compared with its purely data\-driven counterpart trained either on raw relaxation voltage samples or on statistical features extracted from the relaxation curve\. The results indicate that the hybrid[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)generally achieves the lowest prediction error, with its advantage particularly evident when training data or relaxation time are limited\. Essentially following the same approach,\[lin2025Physicsinformedmachinelearning\]use a fractional\-order[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13), whose parameters are identified via recursive least squares and then passed to an[RF](https://arxiv.org/html/2608.10047#p6.39.39.39.39)regressor for[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation\. Although not investigating further[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models, they also benchmark the hybrid[RF](https://arxiv.org/html/2608.10047#p6.39.39.39.39)against a standard[RF](https://arxiv.org/html/2608.10047#p6.39.39.39.39)\(either trained on raw relaxation data or on statistical features\), with the former generally outperforming its counterparts in terms of prediction error\.\[kohtz2022PhysicsbasedMachineLearning\]present a hybrid approach for the online joint[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45)and[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46)estimation of lithium\-ion batteries\. A dual[EKF](https://arxiv.org/html/2608.10047#p6.14.14.14.14)is built on a simple empirical voltage measurement function \(polynomial open\-circuit voltage\-type relation in[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45), capacity, and current\), and an[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)is trained offline as a residual model to learn the error between this physics\-based measurement function and the measured voltage\. In online operation, the[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)correction is added to the empirical measurement function within the dual[EKF](https://arxiv.org/html/2608.10047#p6.14.14.14.14)measurement equation\. Results show that embedding the residual model significantly reduces capacity estimation error, demonstrating more accurate battery health assessment\. In addition to the learning bias discussed in Section[4\.5\.3](https://arxiv.org/html/2608.10047#S4.SS5.SSS3), the method proposed by\[wang2024Physicalknowledgeguided\]also falls into the class of in\-series hybrid approaches\. An[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13)’s identified parameters are concatenated with data\-driven features extracted from current measurements, both in training and inference\. The concatenated feature vector forms the input for a subsequent encoder\.
Table 19:All studies employinghybrid approachesto addresshealth assessment, listed in alphabetical order\.Hybrid Approaches for Health AssessmentReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[feng2024comprehensive\]Second\-order RC[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13)Nonlinear algebraic expressions of exponentials \(derived from first\-order linear[ODEs](https://arxiv.org/html/2608.10047#p6.30.30.30.30)\)Lithium\-ion battery\[kohtz2022PhysicsbasedMachineLearning\]Empirical voltage\-[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45)\-capacity relation \(open\-circuit voltage\-type measurement model\)Algebraic equationLithium\-ion battery\[lin2025Physicsinformedmachinelearning\]Fractional\-order[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13)Fractional\-order state\-space modelLithium\-ion battery\[wang2024Physicalknowledgeguided\]Second\-order RC[ECM](https://arxiv.org/html/2608.10047#p6.13.13.13.13)[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)systemLithium\-ion battery\[xu2024Wearstateassessment\]System\-level lumped\-parameter dynamic modelCoupled nonlinear[ODEs](https://arxiv.org/html/2608.10047#p6.30.30.30.30)Gear pump\[xu2024Wearstateassessment\]present a digital twin\-based framework for wear state assessment of gear pumps in fuel control systems\. A first\-principles dynamic model of the fuel system is built in Simulink, while a deep[RL](https://arxiv.org/html/2608.10047#p6.42.42.42.42)agent adaptively updates flow correction coefficients so that simulated pressures match measured ones\. These learned coefficients, together with normalized operating conditions and model error, form an interpretable wear feature vector whose distance to a healthy reference indicates wear severity\.
#### 4\.6\.4Prognosis
Table 20:All studies employinghybrid approachesto addressprognosis, listed in alphabetical order\.Hybrid Approaches for PrognosisReferencePrior Physical KnowledgeUse CaseTypeRepresentation\[aizpurua2023Integratedmachinelearning\]Arrhenius\-based thermal\-stress model with Miner’s rule for stator winding insulationAlgebraic equationElectric motor\[kundu2024Developmentofdatadriven\]Pit\-growth model inspired by Paris’ lawAlgebraic equationGearbox\[li2024Particlefilterbasedfatiguedamageprognosisusingprognosticaidedmodelupdating\]Fatigue crack growth model \(Paris’ law\)Algebraic equationFatigue crack growth in aluminum lug joint\[liang2024Ahybridapproach\]Double exponential model for capacity predictionAlgebraic equationLithium\-ion battery\[ma2024Accurateandefficient\]Incremental capacity curves expressed as sum of Lorentzian functionsAlgebraic equationLithium\-ion battery\[shi2022Batteryhealthmanagement\]Semi\-empirical calendar and cyclic aging modelAlgebraic equationLithium\-ion battery\[sun2023Adaptiveevolutionenhanced\]Electrochemical\-thermal\-[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)modelCoupled[PDEs](https://arxiv.org/html/2608.10047#p6.31.31.31.31)plus[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)capacity\-fade[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30), solved as a high\-fidelity numerical simulationLithium\-ion battery\[xu2023ANovelHybrid\]Pseudo\-two\-dimensional electrochemical model \(Doyle\-Fuller\-Newman model\)[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)Lithium\-ion battery\[zhu2024PhysicsInformedDeepLearning\]Flank wear modelAlgebraic equationHigh\-speed milling\[zhu2024RemainingUsefulLife\]Vibration degradation modelAnalytic model defined by algebraic equationsBearingA total of ten studies employ hybrid approaches for prognosis, as listed in Table[20](https://arxiv.org/html/2608.10047#S4.T20)\.\[sun2023Adaptiveevolutionenhanced\]propose a prognostics framework for lithium\-ion batteries, in which a multi\-physics simulation model \(electrochemical\-thermal with[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)\-coupling\) and an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)operate in series\. The simulation model uses measured current, voltage and temperature to estimate[SOH](https://arxiv.org/html/2608.10047#p6.46.46.46.46), which is then used both as training targets and as a slower, high\-fidelity reference to periodically recalibrate the[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)during operation\. The[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)employs a dynamically sized input window whose length is adapted based on the Kullback\-Leibler divergence between consecutive windows, so that the network alternately emphasizes long\-term trends and short\-term fluctuations\. An adaptive evolution mechanism retrains and updates the[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)whenever its predictions deviate too much from the simulation, improving long\-term prediction performance\. Compared to a conventional[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26), the proposed framework consistently achieves lower[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction errors across two datasets and laboratory experiments conducted under varying operating conditions\. Specifically tailored for data\-scarce scenarios,\[liang2024Ahybridapproach\]propose a hybrid method for[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction and uncertainty quantification in lithium\-ion batteries\. An ensemble learning approach is adopted to integrate an empirical degradation model and a data\-driven component\. While the former, a double exponential degradation model, captures the nonlinear degradation trend, the[GRU](https://arxiv.org/html/2608.10047#p6.22.22.22.22)\-[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)network learns to predict short\-term fluctuations\. The output of these models, along with the preprocessed data, is fed into a Bayesian[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)that predicts[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)and provides uncertainty quantification, showing comparable performance to several alternative approaches and increased data efficiency in ablation studies\.\[ma2024Accurateandefficient\]present an approach for predicting battery[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)from a single constant\-current charging curve\. The method also addresses the black\-box nature of data\-driven methods by combining battery physics and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\. Characteristic peaks from incremental capacity curves are approximated with Lorentzian functions and integrated to form a smooth, parametric capacity\-voltage model\. From this model, the peak centers, widths, and areas are extracted as features sensitive to degradation for a lightweight[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)that maps them to the[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)\. Across different chemistries and operating conditions, the proposed approach provides more stable and accurate predictions than purely data\-driven methods\. The framework of\[xu2023ANovelHybrid\]enables predicting lithium\-ion battery degradation trajectories and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)\. Early\-life features are extracted in three forms: the variance of the difference of discharge capacity\-voltage curves between cycles 10 and 100; the anode state\-of\-lithiation change between cycles 90 and 100 obtained from a pseudo\-2\-dimensional electrochemical model; and a final feature that multiplies the aforementioned to capture both observable and internal degradation signals\. Battery cells are clustered using k\-means clustering based on the third feature mentioned to group similar aging patterns\. Data augmentation techniques are applied to enrich cluster\-specific datasets\. For each cluster, an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)encoder\-decoder model predicts the full capacity degradation trajectory from limited early cycles\. The proposed method enables accurate and early prediction of battery life across different aging conditions and chemistries and performs better than[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21),[SVR](https://arxiv.org/html/2608.10047#p6.48.48.48.48)and an autoregressive[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)baseline\.\[shi2022Batteryhealthmanagement\]present a physics\-informed framework for lithium\-ion battery degradation estimation and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction\. Their method combines a calendar\-and\-cycle\-aging model, expressed through five semi\-empirical operating stress\-factor formulations, with an[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\. The physics\-based component represents degradation associated with operating conditions \(including cycle duration, rest time, temperature,[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45), and load\), while the[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)uses this modeled degradation together with cycle\-level monitoring signals to learn degradation behavior that is not captured by the former alone\. The estimated degradation trajectory is then passed to a separate[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)that forecasts future capacity loss;[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)is obtained from the predicted point at which the end\-of\-life threshold is reached\. In the reported experiments, the method achieves better capacity\-fade modeling performance than the[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)and[BiLSTM](https://arxiv.org/html/2608.10047#p6.2.2.2.2)baselines under the tested conditions\.
Introducing three distinct mechanisms to inform a[BiLSTM](https://arxiv.org/html/2608.10047#p6.2.2.2.2),\[zhu2024RemainingUsefulLife\]propose a framework for bearing[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction\. One of these mechanisms represents an in\-series approach: a parametric physics degradation model generates an explicit degradation trajectory, which is subsequently passed to the recurrent model to guide training by a mechanistic representation of health progression\. Combined with the observational and learning biases \(see Sec\.[4\.3\.4](https://arxiv.org/html/2608.10047#S4.SS3.SSS4)and Sec\.[4\.5\.4](https://arxiv.org/html/2608.10047#S4.SS5.SSS4), respectively\), the framework consistently outperforms purely data\-driven baselines in terms of[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction errors\. Yet the lack of ablation studies prevents determining which of the three mechanisms yields the greatest benefit\.
\[zhu2024PhysicsInformedDeepLearning\]target tool wear monitoring and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction in high\-speed milling\. A[BiLSTM](https://arxiv.org/html/2608.10047#p6.2.2.2.2)\-based architecture processes cutting force signals to extract temporal features, while a physics\-based wear model \(driven by the milling parameters\) runs in parallel and provides a flank\-wear estimate that conditions an attention mechanism aggregating the learned features into a shared representation used for prediction\. Assuming that tool wear and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)are two manifestations of the same underlying tool state, the network jointly predicts both quantities from this shared representation under a single multi\-task loss with two separate regression heads\. Experiments over different milling conditions show that the proposed model yields clearly improved predictive accuracy compared with an otherwise identical purely data\-driven model\.
To address fatigue crack prognosis in an aluminum lug joint monitored by Lamb waves,\[li2024Particlefilterbasedfatiguedamageprognosisusingprognosticaidedmodelupdating\]targets[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction within a hybrid,[PF](https://arxiv.org/html/2608.10047#p6.33.33.33.33)\-based framework\. Based on Paris’ law, a nonlinear state\-space model is constructed to represent the evolution of fatigue crack length\. Additionally, a[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)model maps the Lamb\-wave feature to lifetime percentage, which is subsequently used to modify[PF](https://arxiv.org/html/2608.10047#p6.33.33.33.33)state and parameter samples \(prognostic\-aided model updating\) so that they are consistent with both past measurements and prognostic information\. Applied to five specimens across five testing scenarios, the method yields more accurate[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)and lifetime\-percentage predictions with generally tighter uncertainty bounds, and achieves more reliable prognosis than both the approach without prognostic\-aided updating and the standalone[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)model\.
In order to perform[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction for gearboxes,\[kundu2024Developmentofdatadriven\]propose multiple approaches\. In one approach,[Random Forest Regression](https://arxiv.org/html/2608.10047#p6.40.40.40.40)\([RFR](https://arxiv.org/html/2608.10047#p6.40.40.40.40)\) is used to determine the current pitting area based on a correlation coefficient\-based[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)constructed from both healthy and faulty gearbox vibration data\. With the current pitting area known, a physical pit\-growth model can be employed to estimate the number of remaining cycles until failure that directly relates to the[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)\. Additionally, Bayesian inference techniques are used to update the parameters of the physical model during inference\.
\[aizpurua2023Integratedmachinelearning\]address prognostics of permanent magnet motors in a maritime context\. With the winding insulation as the degradation quantity of interest, features derived from wind speed and vessel speed are used as inputs to two[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models connected in series\. The first model predicts torque from operational and meteorological data, and the second predicts winding temperature using the predicted torque along with the same input features\. The resulting winding\-temperature trajectory is then fed into an Arrhenius\-based thermal\-stress degradation model with Miner’s rule and Monte Carlo simulation to obtain a probabilistic[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)estimate for the insulation\. The authors benchmark several alternative[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models \(including linear regression, gradient boosting,[RF](https://arxiv.org/html/2608.10047#p6.39.39.39.39), and[MLP](https://arxiv.org/html/2608.10047#p6.28.28.28.28)\) for the torque and temperature prediction tasks and select the best\-performing configurations, but the hybrid[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)framework itself is only validated on a single case study and is not quantitatively compared to other prognostics approaches\.
## 5Discussion
This section discusses the main findings of the review, focusing on recurring methodological patterns, the observed effects of incorporating physics into[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\-based[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), and implications for different[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks\. It then highlights contextual limitations, barriers to deployment, and terminological issues in the current literature and concludes by outlining promising directions for future research\.
### 5\.1Methodological Patterns
Beyond the classification of individual methods \(see Fig\.[4](https://arxiv.org/html/2608.10047#S4.F4)\), examining recurring methodological patterns provides insight into conceptual maturity of the field, the emergence of shared modeling principles, and the extent to which physics integration has become methodologically standardized\. The following analysis examines each class with respect to the diversity of implementation strategies, the strength of physics enforcement, and the implications for transferability and practical adoption\.
#### 5\.1\.1Observational Bias
Methods incorporating observational bias rely on physics solely as a data source, while the model and training objectives remain conventional and agnostic to the underlying physics\. Two dominant strategies emerge across the reviewed literature: using physics\-based simulators to generate simulated data that enrich limited or imbalanced datasets\[li2024Asimulationdatadriven,qin2024Inversephysicsinformed,qin2025SimulationdataDrivenGeneralized,ren2025Healthassessmentof,mei2024Ahybridphysicsinformed,kohtz2022Physicsinformedmachinelearning,deng2023ACalibrationBasedHybrid,zhu2024RemainingUsefulLife\], and pretraining on simulated data followed by fine\-tuning on scarce real\-world data, typically in a[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)setting\[dong2022Anewdynamic,liu2024Enhancingmultitypefault,pettorossi2024AddressingDataScarcity,song2022Researchonfault,zhang2024AdversarialDomainAdaptation,li2024PhysicsGuidedDeepLearning,yishengliu2024HybridFusionfor,zhang2023DynamicModelAssistedBearing\]\. Delta learning, a less common variant in which a model trained on simulated data is corrected by a secondary model trained on experimental data, has also been explored\[thelen2022Integratingphysicsbasedmodeling,navidi2023PHYSICSINFORMEDNEURALNETWORKS\]\.
This apparent methodological uniformity is not a sign of community consensus but rather a consequence of the definitional boundary\. Because physics can enter only through additional training data, the set of feasible implementation strategies is inherently small, regardless of the specific[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)task or asset\. This constraint simultaneously explains both the accessibility and the limitations of observational\-bias approaches: they are straightforward to apply, provided that a simulator of sufficient fidelity can be constructed\. Once training is complete, however, the simulator is discarded and the deployed model behaves like a conventional black box, retaining no mechanism to enforce physical plausibility at inference time\.
Ultimately, observational\-bias approaches are well\-positioned to address data scarcity within regimes covered by the simulated data, yet structurally ill\-suited to deliver on stronger promises such as physical consistency and reliable extrapolation to unseen regimes—properties that require explicit structural or algorithmic enforcement\.
#### 5\.1\.2Inductive Bias
Methods incorporating inductive bias embed physical principles directly into the model through tailored architectural interventions\. Consequently, they achieve the closest alignment between model structure and prior physical knowledge, albeit at the expense of being highly asset\-specific\. The resulting methodological landscape is correspondingly heterogeneous\. Characteristic implementation patterns include: custom recurrent cells informed by specific degradation laws\[yucesan2019Windturbinemain,dourado2019Physicsinformedneuralnetworks,dourado2022Ensembleofhybrid,nascimento2021Hybridphysicsinformedneural,badora2023Usingphysicsinformedneural\], graph topologies reflecting asset structure\[liu2025Graphembeddedpatchsense,feng2022FullGraphAutoencoder,jin2024GraphSpatioTemporalNetworks,cheng2024Researchongas,zhou2025Physicsinformedspatiotemporalhybrid\], layers tailored to known frequency content\[zeng2025ApplicationofFrequency,gao2024MPINet:MultiscalePhysicsInformed\], informed[GP](https://arxiv.org/html/2608.10047#p6.20.20.20.20)priors\[huang2022AnEnhancedDataDriven,zhu2023PhysicsinformedGaussianprocess,qiang2023Integratingphysicsinformedrecurrent\], or activation and output constraints enforcing monotonicity\[hao2023Anoveldeep,li2025Applicationofphysicsguided,yin2025Physicsguideddegradationtrajectory,zhou2023Timevaryingtrajectorymodeling\]\.
While these patterns share broad conceptual goals, their implementations are fundamentally different and tightly coupled to the specific asset under study\. Graph\-based encodings naturally align with diagnostics in multi\-component systems\[liu2025Graphembeddedpatchsense,feng2022FullGraphAutoencoder,jin2024GraphSpatioTemporalNetworks\], whereas custom recurrent cells and informed[GP](https://arxiv.org/html/2608.10047#p6.20.20.20.20)priors predominantly serve prognostics\[yucesan2019Windturbinemain,badora2023Usingphysicsinformedneural,zhu2023PhysicsinformedGaussianprocess\], where degradation dynamics must be captured temporally\. Monotonicity constraints represent a notable exception: because they encode a domain\-agnostic property of irreversible degradation, they transfer readily across assets\. Yet they provide only a weak inductive bias compared to the mechanistic models embedded in asset\-specific architectures\.
Overall, research on inductive bias largely constitutes a proliferation of problem\-specific solutions—the product of tailored model engineering rather than mature, reusable modeling strategies\. While such tailoring yields strong performance and structurally enforces adherence to the embedded physical principles, it contributes little to developing reusable design patterns, leaving this class fragmented\.
#### 5\.1\.3Learning Bias
Learning\-bias approaches are structurally more uniform, with physics incorporated almost exclusively as additional terms in the loss function, regularizing the model toward physically plausible solutions\. Three main strategies emerge for introducing learning bias: the first relies on applying[PINNs](https://arxiv.org/html/2608.10047#p6.36.36.36.36), as demonstrated across a broad range of assets, such as aero\-engine spline couplings\[wang2025Physicallyinformedhierarchical\], power transformers\[ramirez2024ResidualbasedAttentionPhysicsinformed\], axial piston pumps\[dong2025Innovativefaultdiagnosis,chao2025Physicsinformedneural,wangz\.2023Physicsinformedneural\], and lithium\-ion batteries\[singh2023HybridModelingof,jang2025Stateofhealth,pengfeiwen2023PhysicsInformedNeuralNetworks,wang2025ABatteryState,liu2025Aphysicsguidedapproach,wang2024Physicsinformedneuralnetwork\]\. The second strategy hinges on embedding mechanistic models as soft constraints in otherwise standard[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models, including crack growth and wear laws\[pan2025Inservicefatiguecrack,badora2023Usingphysicsinformedneural\], half\-cell or electrochemical aging models and equivalent\-circuit dynamics\[navidi2024PhysicsInformedMachineLearning,zhang2025AnElectrochemicalAgingInformed,xu2022Aphysicsinformeddynamic,pugalenthi2024RemainingUsefulLife,zhang2025APhysicsInformedHybrid\], and dynamic models of industrial robots derived from joint multi\-energy and rigid\-body dynamics\[wang2024Adigitaltwin,li2024Hybridphysicsembeddedrecurrent\]\. Lastly, a third set of strategies relies on comparatively simple constraints, with the two most prominent patterns either imposing generic degradation properties such as monotonicity on health indicators or degradation trajectories\[xu2024Physicsguideddeeplearning,deng2025ANovelMethod,freeman2022Physicsinformedturbulenceintensity,najeraflores2023APhysicsConstrainedBayesian,wang2024Phyformer:Adegradation,fassi2024PhysicsInformedMachineLearning\]or encoding prior knowledge about fault signatures in the frequency domain\[qiao2024APriorKnowledge,tang2024Apriorknowledgeenhanced,xu2024Physicsinformedprobabilisticdeep\]\.
Among these, the[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)formulation has become ade factostandard for learning\-bias integration, offering a unified way to weakly enforce the governing dynamics of a system, whether mechanical, electrochemical, or otherwise\. Its abstract, problem\-agnostic formulation makes it straightforward to implement across diverse domains, which has contributed to their widespread adoption—particularly in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. Yet this convenience introduces a characteristic challenge: the relative weighting between data\-fidelity and physics\-residual losses\. Across the reviewed literature, loss balancing is predominantly controlled by scalar weighting coefficients that are either fixed or tuned empirically, with only a few exceptions proposing more systematic schemes such as uncertainty\-based weighting\[pengfeiwen2023PhysicsInformedNeuralNetworks\], Bayesian optimization\-based weighting\[wang2025PhysicsInformedNeuralNetwork\], or gradient\-norm\-based balancing\[singh2023HybridModelingof,zhang2025APhysicsInformedHybrid\]\. When these weights are poorly calibrated, the model may overfit the data while failing to satisfy physical constraints, or vice versa—a tension that remains an open practical challenge\.
Notably, among all studies incorporating a learning bias, only a single study does so through the optimization procedure itself:\[yonastefera2025ConstraintGuidedLearningof\]embed monotonicity, boundary, and energy\-consistency constraints directly into the gradient\-descent updates\. Whether optimization\-level integration of physics offers practical advantages over loss\-based regularization \(e\.g\., in settings with multiple competing constraints\) remains an open question, given that only a single study has explored this pathway\. However, the near\-complete dominance of regularization\-based approaches likely reflects practical considerations\. Adding penalty terms is trivial in modern[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)frameworks, whereas implementing custom optimizers demands specialized expertise and is harder to generalize\.
Ultimately, a key trade\-off emerges within this class: excluding qualitative degradation properties \(e\.g\., monotonicity or boundedness\), introducing learning bias inherently couples the[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)formulation to the specific asset, trading the method’s transferability across different scenarios for stronger constraints within the intended domain\. An example of a more transferable design is the Koopman\-informed[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)by\[wang2025KoopmanInformedNeuralNetwork\], which learns an approximately linear evolution in a latent eigenfunction space and regularizes forward\-backward consistency of the dynamics without relying on asset\-specific mechanistic models\.
#### 5\.1\.4Hybrid Approaches
Hybrid approaches couple independent physics\-based and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models either in parallel or in series, without embedding physical knowledge within the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)pipeline itself\. Across all four classes, hybrid approaches exhibit the most pronounced methodological uniformity, arising from the limited structural design space: one model feeds the other, or both run concurrently with combined outputs\.
The dominant pattern is in\-series coupling in which physics informs[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\. Several studies leverage calibrated physics\-based models, namely[ECMs](https://arxiv.org/html/2608.10047#p6.13.13.13.13)\[lin2025Physicsinformedmachinelearning,feng2024comprehensive\], electrochemical and aging models for batteries\[ma2024Accurateandefficient,sun2023Adaptiveevolutionenhanced,shi2022Batteryhealthmanagement,xu2023ANovelHybrid\], fuel\-cell models\[pettorossi2025Physicsguidedfaultdiagnosis\], and wear and degradation laws\[zhu2024PhysicsInformedDeepLearning,zhu2024RemainingUsefulLife\]to derive intermediate quantities \(e\.g\., model parameters or health indicators\) that serve as inputs to an[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)model performing the actual[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)task\. In effect, the physics\-based models act asmechanism\-alignedencoders that compress raw sensor data into compact, physically interpretable features\. By reducing the input dimensionality to as few as six\[lin2025Physicsinformedmachinelearning\]or ten parameters\[ma2024Accurateandefficient\], this in\-series coupling enables simplifying the downstream learning task, justifying the use of simple models such as[RFs](https://arxiv.org/html/2608.10047#p6.39.39.39.39)and shallow fully connected networks\.
Moreover, in three other cases, the physics\-based model runs alongside a data\-driven encoder, with their outputs concatenated and fed to a single[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)model\. From the viewpoint of the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)models performing fault diagnosis\[singh2024Hybridphysicsinfused1DCNN\], health assessment\[wang2024Physicalknowledgeguided\], and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction\[liang2024Ahybridapproach\], respectively, this configuration effectively retains an in\-series coupling\. This approach arises when the physics\-based model is acknowledged as too low\-fidelity to serve as the sole encoder \(unlike the in\-series approaches discussed above\), yet still contributes structured information absent from raw data\. In each case, the architecture implicitly decomposes the prediction task: the physics branch supplies trend\-level or steady\-state features while the data\-driven branch captures residual dynamics or high\-frequency fluctuations\. Ablation studies in two of the three studies confirm that the combined approach outperforms either branch in isolation, while the third\[singh2024Hybridphysicsinfused1DCNN\]demonstrates clear performance gains over the data\-driven branch alone\.
The reverse coupling \([ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)informing physics\) is less common but mostly follows a consistent logic: the physics\-based model requires as input a quantity that is not directly observable from available sensors, and the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)model’s sole task is to estimate precisely this quantity\. Identified examples include a[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)model that infers lifetime percentage to correct particle states in a Paris law\-based crack\-growth model\[li2024Particlefilterbasedfatiguedamageprognosisusingprognosticaidedmodelupdating\], an[RFR](https://arxiv.org/html/2608.10047#p6.40.40.40.40)that maps vibration features to the current pitting area for a Paris law\-inspired pit\-growth model\[kundu2024Developmentofdatadriven\], a deep[RL](https://arxiv.org/html/2608.10047#p6.42.42.42.42)agent that updates flow correction coefficients of a first\-principles system\-level digital twin of a fuel control system\[xu2024Wearstateassessment\], and a Wasserstein[GAN](https://arxiv.org/html/2608.10047#p6.18.18.18.18)that infers local fault evidence which is used to update node weights of the probabilistic fault root cause tracing network\[xu2023Physicsguideddatarefinedfault\]\. Crucially, the physics\-based model retains responsibility for the mechanistic inference that constitutes the actual[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)task, thereby preserving physically consistent predictions\.
In\-parallel approaches are the least studied, with all identified examples addressing lithium\-ion batteries, where data\-driven residual models run alongside electrochemical and equivalent\-circuit battery models\[zhang2024Adaptivefaultdetection,kohtz2022PhysicsbasedMachineLearning,firoozi2022CylindricalBatteryFault\]\. All three studies share a common residual\-correction scheme: a physics\-based model \(equivalent\-circuit or electrochemical\) provides a structured baseline prediction, while the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)component \([NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29),[BiLSTM](https://arxiv.org/html/2608.10047#p6.2.2.2.2)or[GPR](https://arxiv.org/html/2608.10047#p6.21.21.21.21)\) explicitly learns the discrepancy between that prediction and the measured signal\. The concentration of in\-parallel approaches in battery applications likely reflects the availability of compact, well\-understood physics models whose outputs are directly comparable to sensor measurements—a prerequisite for meaningfully defining a learnable residual\.
Hybrid approaches offer clear practical advantages: they require comparatively low engineering effort provided a physics\-based model of sufficient fidelity exists\. However, in configurations where the physics model informs a downstream[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)model, the latter remains a black box, inheriting the limitations of purely data\-driven techniques—most notably the risk of physically implausible predictions and unreliable extrapolation\. Nevertheless, hybrid approaches are still relatively common in the[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)literature \(see Fig\.[4](https://arxiv.org/html/2608.10047#S4.F4)\)\.
#### 5\.1\.5Multi\-Class Approaches
Although the four classes are conceptually distinct, the literature confirms that they are complementary in practice\. Although studies have been assigned to multiple classes only if their mechanisms for incorporation are clearly separable, five works nevertheless span multiple classes\[sun2022MicrocrackDefectQuantification,li2025Applicationofphysicsguided,badora2023Usingphysicsinformedneural,wang2024Physicalknowledgeguided,zhu2024RemainingUsefulLife\], and in all cases a learning\-bias component is present\. This pattern likely reflects the comparatively low implementation barrier of adding physics\-informed regularization, which naturally facilitates its combination with the other approaches\. However, none of these multi\-class studies report ablation experiments that isolate the contribution of each embedded mechanism\.
To conclude, the multi\-class cases should primarily be interpreted as evidence that the four classes are practically compatible, rather than as proof that specific combinations outperform carefully designed single\-class approaches\. Establishing when and how multi\-mechanism designs offer systematic advantages remains an open question\.
### 5\.2The Effects of Incorporating Physics
Having characterized the methodological landscape in Section[5\.1](https://arxiv.org/html/2608.10047#S5.SS1), this section examines the empirical evidence for the effects of incorporating prior physical knowledge into[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\-based[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. The analysis is organized along six dimensions: predictive performance, generalizability, robustness, data efficiency, interpretability, and physical consistency\. For each dimension, the strength and scope of available evidence is assessed, and—where possible—the observed effects are traced back to the class of incorporation\. The section concludes by identifying systematic gaps and biases in how effects are currently reported\.
#### 5\.2\.1Predictive Performance
Improved predictive performance constitutes the most consistently reported benefit and is supported across all four classes of approaches and all four[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks\. For fault detection and diagnosis, numerous studies report higher classification accuracy relative to baselines\[qiao2024APriorKnowledge,ma2025Aphysicsbasedsample,dong2022Anewdynamic,xu2024Physicsinformedprobabilisticdeep,zeng2025ApplicationofFrequency,gao2024MPINet:MultiscalePhysicsInformed,huang2025Physicsinformedcausallearning,zhu2024PhysiCausalNet:ACausaland\]\. For health assessment and prognosis, reduced regression errors are typical\[yucesan2022Ahybridphysicsinformed,wang2025ABatteryState,mei2024Ahybridphysicsinformed,najeraflores2023APhysicsConstrainedBayesian\]\. While this breadth of evidence is encouraging, its interpretation requires caution for two reasons\.
First, the strength of the evidence varies substantially with experimental design\. A minority of studies include clear ablation experiments comparing the proposed model against an architecturally identical counterpart from which only the physics component has been removed\. The majority instead compare against simple baselines \(e\.g\., standalone[SVR](https://arxiv.org/html/2608.10047#p6.48.48.48.48),[RF](https://arxiv.org/html/2608.10047#p6.39.39.39.39), or[LSTM](https://arxiv.org/html/2608.10047#p6.26.26.26.26)\), even when the proposed[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)model is architecturally far more complex\. Under such conditions, it is difficult to disentangle performance gains attributable to the embedded physics from those arising from increased model capacity, additional engineering effort, or more sophisticated training procedures\. Comparisons against strong, state\-of\-the\-art baselines are rare\[wang2025KoopmanInformedNeuralNetwork,liu2025Graphembeddedpatchsense\], further limiting the conclusiveness of reported performance gains\.
Second, the magnitude of reported improvements is almost never contextualized with respect to practical significance\. In[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), a small reduction in the[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction error may or may not alter maintenance decisions depending on the asset’s failure consequences, the planning horizon, and the associated prediction uncertainty\. Yet no study in the reviewed literature connects reported accuracy gains to downstream decision quality or maintenance cost savings\. This omission limits the practical value of reported performance improvements for deployment contexts\.
Moreover, the prevalence of highly problem\-specific solutions precludes drawing overarching conclusions about which type of prior knowledge or integration pathway yields the greatest performance gains in general\.
#### 5\.2\.2Generalizability
Generalization, the ability to maintain performance under distributional shift, is a less frequently but repeatedly reported benefit\. Evidence emerges predominantly from two experimental protocols: either in[TL](https://arxiv.org/html/2608.10047#p6.51.51.51.51)settings\[wang2024Physicsinformedneuralnetwork,zhu2024PhysiCausalNet:ACausaland,zhang2023DynamicModelAssistedBearing,zhang2024AdversarialDomainAdaptation\], or across different operating and environmental conditions \(such as varying loads, speeds, or temperatures\)\[zhu2024PhysicsInformedDeepLearning,abiria2025Highcycleandveryhighcycle,feng2022FullGraphAutoencoder,zeng2025ApplicationofFrequency,wang2025KoopmanInformedNeuralNetwork\]\.
A critical distinction, seldom made explicit in the reviewed literature, is that between in\-distribution and out\-of\-distribution generalization\. The former refers to conditions spanned by the training data but held out for evaluation, whereas the latter refers to genuinely novel conditions, environments, or assets\. Most reported generalization evidence pertains to the former\. True out\-of\-distribution generalization, which represents the stronger and practically more relevant claim, is almost never targeted directly and remains largely unsubstantiated\.
Observational\-bias approaches are particularly well\-represented among generalization claims\. However, their purported generalization advantage must be interpreted carefully: by augmenting the training distribution with simulated data that spans a broader operating regime, the effective training distribution is expanded\. Improvements relative to a data\-driven baseline trained on less data therefore partly reflect the additional information injected via simulation, rather than a structural capacity to generalize\. Without controlling for training\-set coverage, such claims risk conflating data augmentation with genuine generalization ability\.
#### 5\.2\.3Robustness
Robustness encompasses a broad range of notions in the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)literature, including stability under input perturbations, resilience to label noise, and resistance to distributional shift\.\[yin2025Physicsguideddegradationtrajectory\],\[zeng2025ApplicationofFrequency\], and\[singh2024Hybridphysicsinfused1DCNN\], for example, investigate robustness against noisy data\.\[kristupasbajarunas2024HealthIndexEstimation\]evaluate robustness to distributional shift by comparing[HI](https://arxiv.org/html/2608.10047#p6.23.23.23.23)quality and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction performance across methods\. Other studies interpret robustness differently, including low variance in[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction errors across varying battery chemistries and operating conditions\[ma2024Accurateandefficient\]and the model’s ability to avoid false alarms\[jin2024GraphSpatioTemporalNetworks\]\. In the latter case, however, evaluation was conducted on a single turbine, with only one documented true fault and a few false\-alarm episodes, without broader tests across turbines, fault types, or operating conditions—limiting the generalizability of the finding\. The inconsistent operationalization of robustness across studies precludes aggregating evidence into a coherent assessment\. While[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)shows promise in this dimension, results are fragmented and often context\-specific, leaving the overall evidence inconclusive\.
#### 5\.2\.4Data Efficiency
Data efficiency, the ability to achieve a given performance level with fewer real observations, is rarely evaluated through dedicated experimental protocols\. Where evidence exists, it typically derives from iteratively reducing dataset sizes and observing performance degradation\[nguyen2023Physicsinfusedfuzzygenerative,wang2024Physicsinformedneuralnetwork,tang2024Apriorknowledgeenhanced,zhang2025APhysicsInformedHybrid\], or from limiting the available history or operational horizon\[fu2024PhysicsInformedNeuralNetwork,zhu2024PhysicsInformedDeepLearning,e2025Aphysicsinformedneural,liang2024Ahybridapproach\]\. The available results generally suggest that incorporating physics largely preserves performance as data volume decreases\.
Improved extrapolation capability can be regarded as a manifestation of data efficiency: by imposing physical constraints, a model can accurately recover solutions with fewer observations in poorly explored regions of the input space, as demonstrated by\[xu2022Aphysicsinformeddynamic\],\[singh2023HybridModelingof\], and\[singh2024Hybridphysicsinfused1DCNN\]\. However, the same qualification noted above applies to observational\-bias methods: the claim of requiring little real data is only partially justified when the underlying data is effectively augmented with simulated observations\. The actual data budget \(including simulation data generation, physics\-model calibration, and real data collection\) is rarely reported transparently, obscuring the true resource requirements\.
#### 5\.2\.5Interpretability
Improved interpretability is among the most frequently claimed yet least substantiated benefits\. In several instances, studies assert interpretability solely because prior physical knowledge has been incorporated, without providing any empirical evidence\. Incorporating prior knowledge does not inherently make a model interpretable\. Indeed, many[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)models remain functional black boxes that offer negligible insight into how predictions are derived\.
A few exceptions demonstrate that at least parts of the model’s inner workings may exhibit a degree of interpretability, as shown by\[xu2024Physicsinformedprobabilisticdeep\], where internal representations qualitatively align with known fault frequencies and uncertainty patterns\. Several studies illustrate how integrating prior physical knowledge enhances the separability of learned representations between regimes, classes, or operating conditions, often visualizing this effect using methods such as t\-distributed stochastic neighbor embedding\[zhu2024PhysiCausalNet:ACausaland,cheng2024Researchongas,li2024PhysicsGuidedDeepLearning,huang2025Physicsinformedcausallearning,zeng2025ApplicationofFrequency,dong2022Anewdynamic,liu2025Graphembeddedpatchsense\]\. While such observations suggest that physics guides the model toward more structured feature spaces, improved cluster separability does not constitute interpretability in a rigorous sense\. Consequently, interpretability remains an aspiration rather than a demonstrated outcome of current[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)methods in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\.
#### 5\.2\.6Physical Consistency
Physical consistency is frequently claimed but lacks a shared definition, rendering it context\-dependent\. As used in the reviewed literature, it encompasses: \(i\) adherence to fundamental physical laws \(e\.g\., conservation principles, thermodynamics, or electrochemistry\); \(ii\) conformity with established degradation properties \(e\.g\., monotonicity, irreversibility, or bounded ranges\); and \(iii\) admissibility of internal model variables \(e\.g\.,[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45)∈\[0,1\]\\in\[0,1\], crack length≥0\\geq 0, or temperature within certain limits\)\. As established earlier, the four classes offer fundamentally different structural guarantees that must be considered when formulating or interpreting physical consistency claims \(see Sec\.[5\.1](https://arxiv.org/html/2608.10047#S5.SS1)\)\.
Beyond these class\-level differences, several structural limitations deserve emphasis\. First, many constraints enforce local behavior \(e\.g\., step\-to\-step monotonicity\), without guaranteeing globally realistic trajectories \(e\.g\., correct knee behavior in battery aging\)\. Second,[PINNs](https://arxiv.org/html/2608.10047#p6.36.36.36.36)enforce governing equations only at a finite set of collocation points\. Third, constraints typically cover only a subset of the physics \(e\.g\., simple wear law\), while ignoring multi\-physics coupling effects that may dominate in certain operating regimes\. Consequently, partial enforcement of physical constraints can create a misleading impression of physically consistent behavior when the unmodeled physics becomes dominant\.
Ultimately, the field lacks a common metric for quantifying the degree of physical consistency of a model, which in turn presupposes consensus on what this notion precisely entails\. Without such metrics, the claims of physical consistency remain qualitative and largely unverifiable\.
#### 5\.2\.7Caveats on Reported Effects
Beyond the dimension\-specific observations above, several cross\-cutting issues affect the reliability of the overall evidence base\. First, the absence of ablation experiments in many studies prevents isolating the contribution of incorporated physics from confounding factors such as architecture changes, additional hyperparameter tuning, or increased training data\. This is particularly acute in studies spanning multiple classes\[sun2022MicrocrackDefectQuantification,li2025Applicationofphysicsguided,badora2023Usingphysicsinformedneural,wang2024Physicalknowledgeguided,zhu2024RemainingUsefulLife\], where no ablation experiment disentangles the individual mechanisms\.
Second,[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)introduces a unique and under\-recognized risk of evaluation leakage: when physics\-model parameters are calibrated on data that overlaps with the test set—as in\[abiria2025Highcycleandveryhighcycle\], where Basquin’s and Paris’ law parameters are fitted to the complete dataset before train\-test splitting—the physics prior becomes partially informed by the test data itself\. This systematically favors the physics\-informed model over purely data\-driven baselines and renders generalization claims overly optimistic\. The vulnerability extends beyond this single instance: any[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)approach that calibrates embedded physics\-model parameters from data is susceptible to this form of leakage unless the calibration is strictly confined to the training partition\.
Third, potential drawbacks of incorporating physics are almost never quantified\. Computational cost \(during training and inference\), convergence behavior, sensitivity to loss\-weight selection, and implementation overhead relative to purely data\-driven alternatives are consistently omitted from evaluations\. Among the reviewed studies, only\[liu2025Graphembeddedpatchsense\]report[Floating Point Operations](https://arxiv.org/html/2608.10047#p6.17.17.17.17)for all compared models\. The widespread absence of such information represents a significant barrier to informed method selection and yields a strongly benefit\-skewed evidence base that is insufficient as a foundation for balanced deployment decisions\.
### 5\.3Contextual Limitations
Having discussed the effects of incorporating physics into ML in Section[5\.2](https://arxiv.org/html/2608.10047#S5.SS2), this section highlights the contextual limitations that must be considered when assessing the maturity and generalizability of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. It first documents a pronounced concentration of the reviewed literature around a narrow set of assets, then traces this imbalance to two reinforcing drivers, namely data availability and the availability of formalized prior physical knowledge, and finally discusses the resulting implications for maturity assessment\.
Across the reviewed literature, lithium\-ion batteries and bearings clearly emerge as the primary focus, accounting for approximately half of the total studies \(see Fig\.[4](https://arxiv.org/html/2608.10047#S4.F4)\)\. Only a small subset of other assets \(i\.e\., cutting tools, pumps, turbofan engines, and metal specimens\) has been investigated in five or more studies, highlighting their relatively limited attention\. The remaining assets are examined sporadically, often in a single study, underscoring a significant gap in research coverage\. This concentration likely reflects the convergence of multiple factors: the industrial and commercial significance of batteries and bearings, the maturity of their respective research communities, and \(as discussed below\) the favorable availability of both public datasets and well\-characterized physics\-based models for these assets\. These factors are mutually reinforcing rather than independent\.
[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)research is strongly shaped by data availability, which in turn influences both the problems studied and the methods developed\. The work of\[mauthe2025overview\_esrel\]presents the most comprehensive overview and analysis of publicly available degradation datasets for[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), covering 98 datasets in total\. The ongoing updating of this overview, along with complete documentation for each dataset, is available online\[mauthe2024overview\_arxiv\]\. Notably, batteries and bearings form the two largest asset categories in terms of dataset count, with 15 each\. By contrast, most other asset types are represented by only one or two datasets\. This imbalance means that researchers reliant on publicly available data encounter a markedly richer landscape for batteries and bearings than for other assets\. In addition, specific datasets \(such as the battery dataset provided by the NASA Prognostics Center of Excellence\[saha2007battery\], XJTU\-SY\[wang2020hybrid\]and FEMTO\[nectoux2012pronostia\]for bearings, or C\-MAPSS for turbofan engines\[saxena2008turbofan\]\) have acquired the status ofde factocommunity benchmarks, further concentrating research activity around the assets they represent\. The reviewed literature reflects this pattern, drawing heavily on publicly available datasets and exhibiting a similar asset distribution\. Consequently, the predominance of studies on lithium\-ion batteries and bearings is, at least in part, reinforced by an availability bias\.
The review reveals a dependency between the degree of formalization of available prior physical knowledge and the range of incorporation strategies that become feasible\. Lithium\-ion batteries and bearings are particularly amenable to[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)because their underlying physics is well\-characterized and formalized\. For batteries, electrochemical models—including[SPM](https://arxiv.org/html/2608.10047#p6.47.47.47.47)variants\[yishengliu2024HybridFusionfor,singh2023HybridModelingof,zhang2025AnElectrochemicalAgingInformed,zhang2025APhysicsInformedHybrid\], half\-cell\[navidi2024PhysicsInformedMachineLearning,thelen2022Integratingphysicsbasedmodeling\]and[SEI](https://arxiv.org/html/2608.10047#p6.44.44.44.44)\-growth models\[kohtz2022PhysicsbasedMachineLearning,liu2025Aphysicsguidedapproach\]—are widely employed, alongside[ECMs](https://arxiv.org/html/2608.10047#p6.13.13.13.13)for cell voltage and[SOC](https://arxiv.org/html/2608.10047#p6.45.45.45.45)dynamics\[liu2024Enhancingmultitypefault,fu2024PhysicsInformedNeuralNetwork,qin2025ManagingBatteryPerformance\]\. For bearings, multi\-[DOF](https://arxiv.org/html/2608.10047#p6.11.11.11.11)dynamic models\[dong2022Anewdynamic,qin2024Inversephysicsinformed,qin2025SimulationdataDrivenGeneralized,deng2023ACalibrationBasedHybrid,zhang2023DynamicModelAssistedBearing,sun2024Contrastivelearningand\], analytical vibration signal models\[li2024Asimulationdatadriven,zhu2024RemainingUsefulLife,gao2024FaultDiagnosisof\], and fault characteristic frequencies\[zeng2025ApplicationofFrequency,qiao2024APriorKnowledge,xu2024Physicsinformedprobabilisticdeep,gao2024MPINet:MultiscalePhysicsInformed\]provide a rich repository of embeddable prior knowledge\. Accordingly, the literature appears to be influenced, at least in part, by a methodological selection bias\. This creates a self\-reinforcing pattern in which assets whose physics is already well\-formalized attract disproportionate research attention, while those whose degradation involves poorly understood or multi\-physics mechanisms remain underrepresented\.
These contextual limitations must inform any evaluation of the current maturity of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. While the mere volume of studies covered in this review may suggest that[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)has matured into an established standard for industrial[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), such an interpretation would be misleading\. The observed imbalance necessitates a more differentiated assessment\. For lithium\-ion batteries and bearings, a certain level of methodological maturity can reasonably be claimed\. However, even for these assets, existing studies focus predominantly on specific[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks: health assessment and prognosis for batteries; diagnosis and prognosis for bearings\. Thus, the apparent maturity is confined to a few asset\-task combinations rather than the full[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)spectrum\. Beyond these focal assets,[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)research remains at an early, exploratory stage with only scattered and often isolated evidence\. The skew in asset coverage is particularly consequential because[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)methods are generally tied to the specific asset under study, owing to the use case\-specific nature of the embedded prior knowledge\. This entanglement constrains applicability to the studied context, and, by extension, raises the question of how to leverage[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)without sacrificing generalizability\.
### 5\.4Barriers to Deployment
The concentration of the literature around a narrow set of assets and tasks \(see Sec\.[5\.3](https://arxiv.org/html/2608.10047#S5.SS3)\) already constitutes a barrier to the broader deployment of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\-based[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. Even within these well\-studied domains, however, a substantial gap separates current proofs of concept from industrially deployable solutions\. This section examines the practical barriers that collectively account for this gap\. These barriers are not independent but cumulative: constructing a[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)solution demands significant expertise and engineering effort; even where such effort is invested, the resulting models rarely provide decision\-relevant uncertainty estimates; and even if both of the former barriers were overcome, insufficient evidence exists regarding whether these models can operate under real\-time and resource constraints\.
#### 5\.4\.1Implementation Effort and Expertise
A prerequisite for any[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)solution is the successful identification, formalization, and embedding of appropriate prior physical knowledge—a process that remains largely undocumented and unquantified across the reviewed literature\. Unlike purely data\-driven pipelines, which can often be constructed by[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)practitioners with general domain familiarity,[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)demands expertise at the intersection of two traditionally separate disciplines: the physics of the asset’s degradation mechanisms and the engineering of[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)architectures and training procedures\. This dual requirement manifests at multiple stages: selecting which physics to embed \(and, equally importantly, which to omit\); translating qualitative physical understanding into a formal, computable representation amenable to integration; choosing an appropriate pathway \(observational, inductive, or learning bias\); and calibrating the interplay between data\-driven and physics\-informed components, such as loss\-weight tuning or simulator fidelity\.
No study among those reviewed reports the human effort, development time, or iterative design cycles required to arrive at the final physics\-informed model\. Yet this engineering overhead is arguably the most immediate practical barrier, particularly when seeking to deploy[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)at scale across heterogeneous asset fleets\. The observation from Sections[5\.1](https://arxiv.org/html/2608.10047#S5.SS1)and[5\.3](https://arxiv.org/html/2608.10047#S5.SS3)that nearly all solutions are tightly coupled to the specific asset under study is, in part, a downstream consequence of this barrier: each new asset demands the aforementioned integration effort, which cannot be easily amortized\.
Incorporating more general prior knowledge, such as simple monotonicity constraints, can broaden applicability, although the resulting performance gains may often be modest\. Developing methods that are both applicable across diverse contexts \(or even assets\) and capable of delivering substantial improvements thus remains a key challenge\. Nevertheless, three patterns in the literature suggest pathways toward reducing this overhead by circumventing the need for intricate physical models: \(i\) structural and topological knowledge describing an asset’s component layout or the relative placement of sensors is particularly suitable for graph\-based approaches\[feng2022FullGraphAutoencoder,liu2025Graphembeddedpatchsense,jin2024GraphSpatioTemporalNetworks,cheng2024Researchongas,zhou2025Physicsinformedspatiotemporalhybrid\]; \(ii\) qualitative degradation properties \(e\.g\., monotonicity\) require no system\-specific physical model and can be imposed via architectural constraints, such as monotone activations, constrained hidden\-state updates, and bounded output layers\[hao2023Anoveldeep,li2025Applicationofphysicsguided,abiria2025Highcycleandveryhighcycle,bai2023PrognosticsofLithiumIon,yin2025Physicsguideddegradationtrajectory,zhou2023Timevaryingtrajectorymodeling\], or inequality\-type loss terms that penalize local violations\[deng2025ANovelMethod,wang2024Physicsinformedneuralnetwork,fassi2024PhysicsInformedMachineLearning,najeraflores2023APhysicsConstrainedBayesian\]; and \(iii\) causal relationships between operating conditions, sensor signals, and the underlying degradation state, although studied less frequently\[kristupasbajarunas2024HealthIndexEstimation\]\. Beyond this, very few studies actually demonstrate applicability across multiple use cases without requiring modifications\[tang2024Apriorknowledgeenhanced,kristupasbajarunas2024HealthIndexEstimation,wang2025KoopmanInformedNeuralNetwork,wang2024Phyformer:Adegradation,zhou2023Timevaryingtrajectorymodeling\], and in each case separate training is still required\.
Taken together, these structural, qualitative, and causal forms of prior knowledge represent the most accessible entry points for transferable[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. However, transferable[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)solutions remain an aspiration rather than an established practice, with the high implementation cost per asset being a principal inhibitor\.
#### 5\.4\.2Uncertainty Quantification
Uncertainty quantification is widely recognized as a core requirement for prognostics in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), where single\-point estimates are“usually considered meaningless”for industrial applications\[kundu2020Areviewon\]\. Yet the vast majority of reviewed studies produce exclusively deterministic predictions, offering no calibrated confidence information to inform decision\-making\.
A smaller subset of studies combines physics with probabilistic models\[bai2023PrognosticsofLithiumIon,jiang2025PhysicsinformedGaussianprocess,ellis2022Ahybridframework\]\. In addition, Bayesian[NNs](https://arxiv.org/html/2608.10047#p6.29.29.29.29)are used, though to a lesser extent\[deng2023ACalibrationBasedHybrid,liang2024Ahybridapproach,najeraflores2023APhysicsConstrainedBayesian\]\. In another study, a Wiener process\-based stochastic degradation model is embedded in an[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)\[he2025Physicsinformedneuralnetwork\]\. However, among these, very few explicitly investigate how the incorporation of prior physical knowledge improves the quality of uncertainty estimates relative to purely data\-driven probabilistic models\[bai2023PrognosticsofLithiumIon,xu2024Physicsinformedprobabilisticdeep\]\. This represents a missed opportunity, because physics\-informed constraints \(including inductive and learning bias\) have the potential to sharpen predictive distributions, e\.g\., by ruling out physically implausible predictions and thus yielding tighter confidence bounds\.
Conversely, this same mechanism introduces a distinctive risk: in operating regimes where the embedded physics becomes inaccurate \(e\.g\., due to unmodeled multi\-physics coupling, degradation\-mode transitions, or environmental conditions outside the model’s validity range\), overly constrained physics\-informed predictions may produce dangerously miscalibrated uncertainty estimates\. However, neither improvements nor deteriorations in uncertainty estimates resulting from the incorporation of prior physical knowledge have been sufficiently studied in the current literature\. As a result,[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)models in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)rarely provide the reliable, calibrated uncertainty information needed for decision\-making, representing a critical barrier to deployment in settings where decisions carry safety or financial consequences\.
#### 5\.4\.3Computational Cost and Operational Readiness
The transition from offline validation to operational deployment introduces requirements that the current literature leaves largely unaddressed: real\-time inference under latency constraints, execution on resource\-limited hardware, adaptation to evolving conditions, and integration with existing monitoring and maintenance infrastructure\. Assessing the feasibility of meeting these requirements presupposes insight into computational costs, yet such information is largely absent from the reviewed studies\.
Among all reviewed studies, only\[liu2025Graphembeddedpatchsense\]report[FLOPs](https://arxiv.org/html/2608.10047#p6.17.17.17.17)for all compared models, thereby enabling fully transparent computational assessment\. Partial reporting is more common but insufficient:\[fu2024PhysicsInformedNeuralNetwork\]and\[xie2024DegradationStateAssessment\]report[FLOPs](https://arxiv.org/html/2608.10047#p6.17.17.17.17)for their proposed physics\-informed models but not for baselines\.\[zeng2025ApplicationofFrequency\]and\[pugalenthi2024RemainingUsefulLife\]evaluate efficiency solely in terms of training time without reporting[FLOPs](https://arxiv.org/html/2608.10047#p6.17.17.17.17)or parameter counts, limiting the comparability of their efficiency claims across studies\.\[wang2025ABatteryState\]additionally report inference times, revealing that their[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)\-based approach incurs training times approximately two orders of magnitude above the fastest baseline, while inference times remain comparable across models\. These scattered observations do not enable systematic comparison of performance gains versus computational overhead across the field\.
Nevertheless, the inherent characteristics of each class of approaches permit a qualitative assessment of computational trade\-offs\. Observational\-bias approaches shift the computational burden to an offline simulation phase: once the ML model is trained, inference cost is identical to a purely data\-driven model\. Inductive\-bias approaches embed physics directly in the architecture, where the added inference cost depends on the specific mechanism—ranging from negligible \(e\.g\., constrained activations\) to non\-trivial \(e\.g\., embedded[ODE](https://arxiv.org/html/2608.10047#p6.30.30.30.30)solvers\)\. Learning\-bias approaches, particularly[PINNs](https://arxiv.org/html/2608.10047#p6.36.36.36.36), increase training cost through additional loss evaluations and automatic differentiation but typically incur no overhead at inference, unless online retraining is required\. Hybrid approaches are unique in retaining the physics\-based model during inference, which can become a bottleneck for real\-time deployment when the physics model is computationally expensive\. These qualitative distinctions provide initial guidance for practitioners facing real\-time engineering decisions, while underscoring the urgent need for future studies to systematically report computational metrics alongside predictive performance\.
Beyond computational cost in isolation, several interrelated deployment requirements remain entirely unaddressed\. First, virtually no study reports inference latency, throughput, or memory footprint under conditions representative of industrial monitoring systems\. Second, experiments investigating whether current[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)models can be executed on resource\-constrained hardware such as embedded edge devices are missing\. Third, the coupling of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)models with operational infrastructure \(e\.g\., programmable logic controllers, supervisory control and data acquisition systems, cloud\- and edge\-based data pipelines, or maintenance management systems\) is never discussed\. Fourth, the question of model maintenance after deployment, including adaptation when operating conditions drift, physics assumptions degrade, or new failure modes emerge, is entirely unstudied\. Based on the experimental settings reported \(i\.e\., predominantly laboratory datasets, offline evaluations, and controlled conditions\), the reviewed methods appear to correspond broadly to early[Technology Readiness Levels](https://arxiv.org/html/2608.10047#p6.52.52.52.52)\(approximately[TRL](https://arxiv.org/html/2608.10047#p6.52.52.52.52)3–4\), with no study demonstrating operational deployment\.
### 5\.5Terminology
The review highlights a notable inconsistency in the terminology regarding[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)across[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)studies\. Within this paradigm, many researchers introduce novel contributions by combining the descriptor“physics\-informed”with a term denoting their specific model or approach\. As a result, the descriptor“physics\-informed”emerges as the most frequently used term among the reviewed studies, effectively signaling that prior physical knowledge is incorporated into the[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)pipeline in some form\. This dominance, however, may have prompted some researchers to use other descriptors better reflecting the particular characteristics of their approach, such as“physics\-guided”\[li2025Applicationofphysicsguided\]or“physics\-constrained”\[najeraflores2023APhysicsConstrainedBayesian\]—a phenomenon known to the authors prior to conducting this review, which also influenced the keyword selection \(see Tab\.[2](https://arxiv.org/html/2608.10047#S3.T2)\)\. While the choice of the respective descriptor may be appropriate in some instances, it is rarely accompanied by an explicit explanation\. The wide variety of terms used is, on the one hand, likely a consequence of the relatively recent emergence of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)as a research field, particularly in older studies\. On the other hand, in certain cases, the descriptor appears to reflect an arbitrary choice of wording\. In milder instances, this results in descriptors other than“physics\-informed”being used consistently within a single study—individually unproblematic, yet collectively contributing to terminological fragmentation across the field\. In more extreme cases, however, multiple variants may appear within the same study\. For instance,\[freeman2022Physicsinformedturbulenceintensity\]refer to their loss function as“physics\-informed,”“physics\-guided,”and“physics\-based,”which likely reflects an attempt to employ synonyms for stylistic variation\. While inconsistent terminology across studies is understandable, given that the seminal work by\[karniadakis2021physics\]was published in 2021, inconsistencies within individual studies remain problematic, as they can obscure the intended message and compromise clarity for the reader\.
When the descriptor“physics\-informed”is used consistently, an ambiguity naturally arises with the term[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)—a pattern repeatedly observed across the reviewed studies\. Since\[raissi2019physics\]coined the term to describe[NNs](https://arxiv.org/html/2608.10047#p6.29.29.29.29)that incorporate known physical laws, expressed as differential equations, into their loss function, the term[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)is somewhat constrained in its usage\. Several studies make use of this term, yet their approaches do not align with its original definition\[badora2023Usingphysicsinformedneural,deng2025ANovelMethod,navidi2023PHYSICSINFORMEDNEURALNETWORKS\]\. In light of the highly influential work of\[raissi2019physics\], many readers may reasonably expect the term[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)to imply the incorporation of differential equations for regularization, creating potential confusion when used differently\. It is acknowledged, however, that establishing clear terminology for novel approaches that are both physics\-informed and employ[NNs](https://arxiv.org/html/2608.10047#p6.29.29.29.29), yet do not align with the concept of[PINNs](https://arxiv.org/html/2608.10047#p6.36.36.36.36), remains a challenge\.
Since approaches to implementing[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)are typically subdivided into model\-based, data\-driven, or hybrid, it is understandable that some studies use the terms“physics\-informed”and“hybrid”interchangeably \(e\.g\.,\[lehmann2024LearningtheAgeing\]\), given that the latter is generally regarded as a broader category encompassing[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\. While this may be technically correct, the classification scheme adopted in this review makes an explicit distinction between[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)and hybrid approaches to further enhance clarity in this regard, thereby facilitating more precise communication regarding the methodology employed in each study\. In this context, the combined use of these terms within a single phrase can be misleading, such as“hybrid physics\-informed neural network”\[dourado2022Ensembleofhybrid\],“physics\-informed spatio\-temporal hybrid neural network”\[zhou2025Physicsinformedspatiotemporalhybrid\], or“hybrid physics\-embedded recurrent neural network”\[li2024Hybridphysicsembeddedrecurrent\]\. This does not pose a problem when the corresponding study actually combines a hybrid approach that incorporates bias via one of the three[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)pathways\.
Apart from the use of descriptors to specify the proposed approach, some studies even apply them to the entire field of research, disregarding the foundational definition of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)by\[karniadakis2021physics\]\. While these alternative terms still refer to research that effectively incorporates prior physical knowledge into[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27), this practice appears to be related to the earlier\-discussed issue of unnecessary lexical variation\. Although there is some flexibility in naming a novel approach differently, using alternative terms for the research field may give the impression of a subtle yet notable distinction regarding[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\. For example,\[kohtz2022PhysicsbasedMachineLearning\]propose employing“physics\-based machine learning techniques\.”\[bai2023PrognosticsofLithiumIon\]deviate even further, referring to it as“knowledge\-constrained machine learning\.”\[yan2025KnowledgeDrivenMachine\]claim to introduce a novel taxonomy, named“knowledge driven machine learning,”intended to synthesize the current state of research at the intersection of prior knowledge and[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\. They explicitly adopt the taxonomy of\[vonrueden2021informed\], yet contribute no novelty and could therefore have been referred to simply as[IML](https://arxiv.org/html/2608.10047#p6.24.24.24.24)\.\[yin2025Physicsguideddegradationtrajectory\]reference\[karniadakis2021physics\]to outline the three pathways for introducing physics into[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\. In doing so, they fully adopt the conceptualization of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35), yet deliberately use the term“physics\-guided”throughout their entire study—a choice for which no clear rationale is provided\.
### 5\.6Future Research
The preceding analysis reveals that[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)already delivers tangible performance benefits for[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), yet persistent methodological gaps \(see Sec\.[5\.1](https://arxiv.org/html/2608.10047#S5.SS1)\), insufficiently substantiated claims \(see Sec\.[5\.2](https://arxiv.org/html/2608.10047#S5.SS2)\), narrow asset coverage \(see Sec\.[5\.3](https://arxiv.org/html/2608.10047#S5.SS3)\), practical deployment barriers \(see Sec\.[5\.4](https://arxiv.org/html/2608.10047#S5.SS4)\), and terminological inconsistencies \(see Sec\.[5\.5](https://arxiv.org/html/2608.10047#S5.SS5)\) collectively constrain the field’s advancement\. The following directions are organized to address each of these gaps in turn\.
The review demonstrates the potential of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)for[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), yet the analyzed studies also reveal persistent methodological gaps and open challenges, as discussed in Section[5\.1](https://arxiv.org/html/2608.10047#S5.SS1)\. Addressing these issues will be essential to translate current approaches into robust, deployable solutions for industrial practice\. Building on the synthesis and critical discussion presented above, one key opportunity for future research is the design of benchmark studies that systematically compare observational\-, inductive\-, and learning\-bias approaches under controlled conditions\. Wherever feasible, these benchmarks ought to reuse the same prior physical knowledge so that differences in performance can be attributed to the pathway of integration rather than to the physics itself\. The resulting evidence would enable well\-founded guidelines for selecting an integration strategy based on the type and fidelity of available prior knowledge, the volume and quality of data, and application\-level requirements such as robustness and interpretability\.
Tackling the highly heterogeneous landscape of inductive\-bias approaches \(see Sec\.[5\.1\.2](https://arxiv.org/html/2608.10047#S5.SS1.SSS2)\) requires moving from problem\-specific designs toward modular building blocks that can be reused with minimal adaptation—analogous to the[PINN](https://arxiv.org/html/2608.10047#p6.36.36.36.36)framework for learning\-bias approaches\. Ultimately, the goal should be to reduce the often\-overlooked engineering overhead of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35), making inductive\-bias approaches more practical and scalable\. Moreover, regularization\-based methods dominate the landscape of learning\-bias approaches \(see Sec\.[5\.1\.3](https://arxiv.org/html/2608.10047#S5.SS1.SSS3)\), yet reliance on empirically tuned loss balancing suggests a shift toward adaptive methods\. While aiming to stabilize training and promote convergence, such methods should ultimately ensure that the physical loss contributes appropriately, preventing the model from over\-prioritizing the minimization of data loss\. To gain more precise control over both training dynamics and the influence of physics on model training, future research should additionally explore integrating physics directly into the optimizer\. Although challenging to implement, physics\-informed optimizers can produce solutions that are both accurate and physically plausible by incorporating constraints directly into the optimization step, providing stronger adherence than loss\-based regularization\.
While promising, the claimed improvements from incorporating prior physical knowledge largely remain unsubstantiated beyond predictive performance, and require more rigorous experimental design \(see Sec\.[5\.2](https://arxiv.org/html/2608.10047#S5.SS2)\)\. To enable fair and informative comparisons, models built on simpler architectures should be benchmarked against conventional baselines, whereas more powerful architectures should be compared to state\-of\-the\-art alternatives\. In all cases, ablation studies are essential to isolate the specific contribution of the embedded physics\. Moreover, experiments need to encompass diverse operating and environmental conditions to rigorously evaluate the validity of the claims—an aspect particularly crucial in the context of industrial[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\. In addition, physical consistency often lacks a precise formulation\. To facilitate meaningful comparisons, it is necessary to establish a formal definition and derive corresponding metrics\. These could quantify the frequency and magnitude of constraint violations, deviations from physically plausible ranges, and cumulative errors over time, enabling a nuanced assessment of how well models respect physical laws while maintaining predictive accuracy\. Finally, to counter the prevailing benefit\-skewed view, evaluations should routinely quantify the trade\-offs and costs of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\(e\.g\., computational load, convergence behavior, training stability, and engineering effort\) alongside any performance gains\. This would facilitate a balanced assessment of whether physics integration justifies its associated overhead in a given deployment context\.
Although deeper study of different asset types may mitigate the contextual limitations identified in Section[5\.3](https://arxiv.org/html/2608.10047#S5.SS3), prioritizing generalizable solutions over asset\-specific studies represents a more productive allocation of research effort\. By systematically analyzing and abstracting the mechanisms that have proven most effective, researchers can derive models applicable across entire classes of assets\. This challenge hinges on shared fundamental principles governing diverse degradation phenomena, highlighting the need to study trade\-offs between asset\-specific and general prior knowledge to understand how performance is gained or sacrificed in pursuit of broader applicability\. A key enabler in this regard is modular design patterns that facilitate adaptation to new assets without requiring substantial engineering overhead\. Taken to its logical conclusion, this points to the development of physics\-informed foundational degradation models that are trained in a multi\-task fashion to simultaneously address all core[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks across multiple assets and operating conditions\. Such models would essentially function as universal backbones, obviating laborious problem\-specific development\.
Translating[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)from research to industry faces the deployment barriers discussed in Section[5\.4](https://arxiv.org/html/2608.10047#S5.SS4), necessitating progress along multiple interrelated axes\. Complementing the efforts toward transferable solutions outlined above, automated or semi\-automated methods for selecting, calibrating, and embedding prior physical knowledge would lower the entry barrier for practitioners who possess domain expertise but lack specialized[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)engineering skills\.
Moreover, uncertainty quantification must shift from an optional addition to an integral design objective\. Future work should both systematically investigate how incorporated physics affects the quality of uncertainty estimates and develop diagnostic indicators that signal when a model’s physics assumptions are being violated, providing actionable safeguards against overconfident predictions in safety\-critical scenarios\.
Beyond this, future reporting should adopt full transparency regarding computational costs, while research should concurrently investigate lightweight physics\-informed architectures suitable for execution on resource\-constrained edge hardware, including model compression and pruning techniques\.
Moreover, while generalization from simulation data to test bench data is frequently studied, the subsequent transfer to field data remains unaddressed\. Future studies must demonstrate generalization to customized industrial machines, especially in settings where idealized physical laws may no longer apply because environmental influences or multi\-component interactions are superimposed on the embedded prior knowledge\. By extension, models must be capable of updating as operating conditions drift or new failure modes emerge, pointing to research on online adaptation mechanisms\.
Finally, demonstrating closed\-loop integration with industrial monitoring and maintenance infrastructure \(e\.g\., programmable logic controllers, supervisory control and data acquisition systems, or cloud\-edge data pipelines\) would constitute a critical step toward elevating[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\-based[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)from laboratory validation to operational maturity\.
In future research, reaching a consensus on core terminology would facilitate clearer communication, more consistent methodology, and more meaningful synthesis of findings across studies\. Addressing this challenge \(identified in Sec\.[5\.5](https://arxiv.org/html/2608.10047#S5.SS5)\) entails both conceptual and practical efforts: conceptually, by establishing precise definitions—such as that of physical consistency—which can underpin the development of relevant metrics, and practically, by adhering to a consistent naming convention when introducing novel methods\. With respect to the latter, the results suggest adhering to“physics\-informed”as the descriptor, in line with[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\. If deviation from this term is justified, either an explicit explanation should be provided or the alternative term should be unambiguous by default, as exemplified by the“Koopman\-informed neural network”\[wang2025KoopmanInformedNeuralNetwork\]\. In both cases, it is crucial that the chosen term be used consistently throughout the study\.
## 6Conclusion
[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)is increasingly expected to provide reliable, trustworthy, and data\-efficient decision support from sparse, noisy, and heterogeneous data\. Given these demands, this review set out to examine how[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)is currently leveraged in[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)by systematically addressing several research questions that explore the prior knowledge employed, its incorporation, and corresponding implications for practice\. By conducting the most comprehensive systematic literature review to date at the intersection of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)and[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), covering 212 studies, this work provides comprehensive answers to these questions\.
1. 1\.Knowledge\(a\) What types of prior physical knowledge are being leveraged?
In terms of prior physical knowledge, the field relies predominantly on mechanistic models and explicit degradation laws, whereas structural and causal relationships and qualitative degradation properties see less frequent application\.
1. 1\.Knowledge\(b\) What forms of representation are employed?
Consequently, the corresponding forms of representation span from highly formalized expressions, through empirical and phenomenological formulations, down to implicit forms that capture principled assumptions\. While the current focus capitalizes on well\-established physical understanding, evidence from the reviewed literature indicates that prior knowledge across the full spectrum of types and representations can be productively leveraged, with each type contributing differently to the trade\-off between physical fidelity and transferability\.
1. 2\.Incorporation\(a\) How can prior physical knowledge be incorporated?
Methodologically, three overarching pathways constitute distinct yet complementary classes of approaches to incorporating prior physical knowledge: observational bias, inductive bias, and learning bias\. Among these, learning\-bias approaches are the most widely adopted, followed by inductive\-bias and observational\-bias approaches\. This pattern likely reflects the advantages of learning\-bias approaches in balancing flexibility regarding the types of prior physical knowledge that can be incorporated with practical implementation feasibility, while still imposing soft constraints that effectively regularize model behavior\. In contrast, observational\-bias approaches require a simulator of sufficient fidelity, whereas inductive\-bias approaches face inherent difficulties in their tailored implementation—both presenting practical challenges that limit, to some extent, their adoption\.
In terms of methodological maturity, observational\-bias approaches leave little room \(by definition\) for alternative ways of incorporating physics\. The remaining two pathways encompass broader design spaces, which account for the observed heterogeneity in how the corresponding methods are implemented\. Although distinct schemes for introducing inductive bias have emerged, they share only broad conceptual similarities\. A similar pattern is observed regarding learning bias, where prior knowledge is almost exclusively embedded via additional loss terms, though the form and weighting of these constraints vary widely\. In parallel to these three pathways, hybrid approaches form a conceptually distinct fourth class, where a physics\-based model and an[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)model are either coupled in parallel or in series\. Although[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)is generally subsumed underhybridin the context of[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), distinguishing physics\-informed from hybrid approaches elucidates the distinct ways in which prior knowledge is applied: the former integrate prior knowledge directly into the ML pipeline itself, whereas the latter retain an explicit stand\-alone physics\-based model that remains an integral part of the prediction pipeline during inference\. Yet hybrid approaches constitute the least frequently adopted class\.
1. 2\.Incorporation\(b\) How does the form of representation influence which approaches to incorporation are feasible?
In terms of feasibility, the form in which prior physical knowledge is represented is not a minor implementation detail but the primary design lever that determines which pathways are viable\. Mechanistic models are flexible enough to support all four classes of approaches, while being most readily incorporated as observational bias or in a hybrid setting\. Structural and topological information naturally maps to inductive bias via graph\-based approaches, whereas qualitative properties are predominantly expressed as learning bias and, in some cases, as simple architectural constraints\. As indicated earlier, this creates a systematic trade\-off: richer, more formal representations permit closer alignment with physics but require substantial modeling effort and are often asset\-specific, whereas qualitative properties are inherently limited in their physical fidelity yet typically retain broad applicability\. Accordingly, with development largely being shaped by the form of representation, the effort to transform prior knowledge into suitable representations becomes instrumental in opening previously inaccessible pathways—an aspect that has received little attention\.
1. 3\.Practice\(a\) How does incorporating prior physical knowledge help overcome limitations of purely data\-driven methods?
Purely data\-driven[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)applications continue to face several challenges in real\-world settings, including scarce and imbalanced degradation data, sensitivity to distributional shift across operating conditions and assets, poor extrapolation beyond the training regime, and predictions that can be physically implausible\. Although[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)is explicitly intended to mitigate these inherent shortcomings, empirical evidence only partially supports improvements in these areas\. This does not imply that[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)is incapable of delivering improvements, but that the number of systematic experiments specifically targeting these areas remains limited across the literature\. Nevertheless, across all four classes of approaches, the reviewed studies consistently demonstrate improved predictive performance for each[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)task and across a broad range of assets\.
There is also accumulating, though less robust, evidence that incorporating prior physical knowledge can enhance data efficiency, support more stable generalization, and enhance extrapolation\. Furthermore, claims regarding improved interpretability, robustness, and physical consistency remain weakly substantiated: interpretability gains are typically inferred solely from the integration of prior knowledge but not quantified, robustness is defined inconsistently and is rarely evaluated across a sufficiently broad range of operating and environmental conditions, and physical consistency lacks a shared definition and corresponding metrics, thereby leaving reported improvements largely qualitative\.
Lastly, some approaches are by definition structurally ill\-suited to deliver on certain promises\. Observational\-bias methods, for example, can plausibly tackle data scarcity and improve in\-distribution generalization, but they leave the hypothesis space unchanged and are therefore poorly equipped to enforce physical consistency or principled extrapolation beyond the regimes spanned by the \(simulated\) training data\. Similarly, hybrid approaches that employ an in\-series coupling, in which a physics\-based model feeds an[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)model, largely inherit the shortcomings attributed earlier to purely data\-driven methods\.
Overall, the current evidence indicates that physics\-informed approaches already provide tangible advantages, yet broader benefits frequently ascribed to[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)in prior work are only partially substantiated and will require more rigorous, multi\-dimensional evaluation before they can be regarded as confirmed\.
1. 3\.Practice\(b\) What are the primary challenges in developing and applying physics\-informed approaches?
Despite the promising outlook suggested by the reviewed literature, efforts to develop and apply[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)approaches within the field of[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)remain subject to major challenges\. With the current landscape being highly fragmented, the field largely lacks standardized design patterns\.
Furthermore, each method is characterized by a certain degree of entanglement between the embedded physics and the specific asset addressed—an entanglement that determines its applicability to other settings\. Given that the literature has largely focused on a limited set of assets \(where the same problem\-specific limitations apply\), the field is still regarded as nascent in terms of empirically grounded development of transferable methods\. The evidence suggests that continued incremental work within existing silos is unlikely to produce cumulative progress\. Instead, advancement requires moving from problem\-specific solutions toward foundational degradation models that facilitate physics\-informed modeling irrespective of the asset in question\.
Furthermore, uncertainty quantification, which is widely regarded as indispensable for risk\-aware maintenance planning and safety\-critical decision\-making, is largely absent\. Likewise, as noted earlier, the lack of evidence for interpretability impedes real\-world adoption, since practitioners demand predictions that are not only accurate but also transparent and trustworthy\.
Lastly, online capability is rarely addressed explicitly, and closed\-loop integration with existing monitoring and maintenance workflows is virtually never demonstrated\. Consequently, from a deployment perspective, most methods remain at the level of offline prototypes rather than operational tools\. Thus, the practical feasibility of large\-scale industrial deployment has not yet been explored\.
## Declarations
FundingThe research leading to these results received funding from Deutsche Forschungsgemeinschaft \(DFG, German Research Foundation\) under grant number 514247199 as part of the research projectTheoMation\.
Conflict of InterestThe authors have no competing interests to declare that are relevant to the content of this article\.
## References
## Appendix
### Excluded Studies
After full\-text analysis of the 212 eligible studies, 83 were excluded, most commonly due to a small number of recurring reasons\. Many contributions focused on domains or assets outside the[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)scope adopted for this review \(e\.g\., civil infrastructure or buildings\), or did not perform any[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)task \(e\.g\., quality\-control or design\-phase fatigue life prediction only\)\. A substantial share of studies did not satisfy the operational definition of[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)used here: they were purely data‑driven, relied exclusively on simulated data without real measurements, or used domain knowledge only in the form of conventional feature engineering\. Additional studies were excluded because they did not employ[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)at all, were abstract‑only or doctoral‑symposium contributions, or had been retracted\. TableLABEL:tab:excluded\_recordsdetails the excluded records \(totaling 83\) and the rationale for their exclusion to ensure completeness and enhance the transparency of the review process\.
Table 21:This table lists all studies deemed outside the scope of this review following full\-text analysis\. For each study, the reference, the rationale for exclusion, and the reviewer are provided\. Reviewers are identified as“A”and“B,”corresponding to the two equally contributing authors \(in no particular order\)\. The rationale is not a summary of the respective work but concisely explains the basis for exclusion, i\.e\., some rationales may require consultation of the original study for full context\.ReferenceRationaleRev\.\[carter2025Imperfectphysicsguidedneural\]The use of abstract simulated systems \(Tinkerbell attractor, Rössler attractor, and a continuous stirred tank reactor\) leads to a scope mismatch, as the work is detached from industrial settings and lacks[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)relevance \(see Sec\. 3; Fig\. 5\)\.B\[che2025UnlockingInterpretablePrediction\]The“physics extractor”[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)is trained to learn the electrochemical impedance spectroscopy from earlyQQ\-VVcurves, after which“the learned physical features were augmented to the measured features”for capacity prediction, representing standard deep learning \(see Fig\. 1; Sec\. 4; Eq\. 1\)\.A\[han2025Physicsinformedsymbolicregression\]The selection of a recursive model due to the dynamic nature of tool wear as well as the selection of input features based on domain knowledge is insufficient to qualify as[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\(see Sec\. 4\.5; Fig\. 3\)\.B\[hong2025Physicsinformedmachinelearning\]A model is derived that can be used to quantify valve flow rate, but no[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)task is addressed, i\.e\., the work rather serves as a“reference for research on the design and control methods of hydraulic control systems”\(see Sec\. 5\)\.B\[kadiwala2025Decodingdegradation:The\]The[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)is identified via sparse regression on the same empirical dataset and the“solution of the[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)is then integrated into the feature set”for a standard Gaussian process regression model, i\.e\., feature engineering \(see Sec\. 2\.3/2\.4; Fig\. 4\)\.A\[li2025Ahybridphysics\]The proposed approach is designed to predict the fatigue life of alloy samples under multiaxial loading\. Since the current state of the system is neglected by not taking live sensor data into account, this cannot be considered[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\(see Sec\. 2\.2\.5; Fig\. 4\)\.B\[li2025Remainingusefullife\]Alongside end\-to\-end feature extraction from raw data, time\-, frequency\-, and time\-frequency\-domain features are computed, misleadingly framed as leveraging prior knowledge, i\.e\., standard feature engineering \(see Fig\. 3; Sec\. 3\.1\.2; Tab\. 1\)\.B\[liao2025Classifierguidedneuralblind\]Although labeled“physics\-informed”, the method employs a conventional multi\-term loss function that comprises kurtosis,l2l\_\{2\}/l4l\_\{4\}norm and cross\-entropy, without incorporating any explicit physical laws or constraints \(see Tab\. 1; Sec\. 3; Eq\. 29\)\.B\[lu2025Priorknowledgeembedding\]Although they are integrated into the latent space of an[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1), the prior information corresponds to common features \(such as root mean square, standard deviation, kurtosis\), which are therefore insufficient to qualify as[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\(see Sec\. 3\.1\.1; Tab\. 1\)\.B\[luo2025Amethodfor\]Although the approach is physics\-informed—using a loss term to regularize the monotonic relationship between membrane resistance and capacity loss—it is excluded by definition because the model is trained solely on simulated data \(see Sec\. 2\.2; Fig\. 4\)\.A\[stoyanov2025Modellingthefatigue\]By definition, using only so\-called“physics\-informed datasets”generated from high\-fidelity thermo\-mechanical finite\-element simulations, without any real data, does not constitute[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\(see Fig\. 1; Sec\. 3/4\.1\)\.B\[sun2025Amethodfor\]This work is excluded as it was retracted at the request by the Editor\-in\-Chief due to plagiarism \(see[A method for estimating lithium\-ion battery state of health based on physics\-informed machine learning](https://www.sciencedirect.com/science/article/pii/S0378775324017191), accessed on December 23, 2025\)\.A\[wang2025DigitalTwinDrivenPhysically\]Since real\-world data are used solely to validate the digital twin, and the proposed approach inherently relies on generated data to train the fault diagnosis model, this work is excluded by definition \(see Fig\. 9; Sec\. III\-D/IV\-C\)\.A\[wang2025Fewshotfaultdiagnosis\]The additional loss term designed to preserve distance information in the embedding space of continuous wavelet transformation snippets that are input to a vision Transformer cannot be considered as physically meaningful prior knowledge \(see Sec\. 3\.2\)\.B\[yuwang2025MetaLearningandKnowledge\]The method assumes degradation follows an unknown[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)in a latent state space, inferring the governing dynamics from data, thereby learning a[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)\-like relationship between the hidden state and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)without physical priors \(see Sec\. 3\.2\.3; Fig\. 2\)\.A\[zhang2025Priorknowledgeinformedmultitask\]By learning ten“signal feature indicators”\(such as max, min, standard deviation\) through an auxiliary task in a shared[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8), the method provides only feature self\-supervision and does not incorporate prior physical knowledge \(see Tab\. 2; Sec\. 2\.4; Fig\. 2\)\.B\[zheng2025Predictionmodeloptimization\]Although claiming to introduce prior knowledge, the method does not incorporate domain knowledge, relying instead on pre\-trained weights and teacher outputs for transfer learning and distillation, i\.e\., a purely data\-driven approach \(see Sec\. 3\.1; Fig\. 3\)\.A\[zhong2025MIPISincNet:Anexplainable\]A“physics\-informed convolutional layer”with analytically designed, fixed kernels based on bearing fault orders and Sinc\-based bandpass filters is used as the first network layer, which effectively amounts to feature engineering \(see Sec\. 3\.2; Tab\. 1\)\.B\[chen2024OpticalSpectralPhysicsInformed\]The“physics\-informed”part fuses raw optical spectra \(grayscale images\) with a second input channel \(selected emission lines, statistical/time\-frequency features, and operating parameters\), effectively performing feature engineering \(see Sec\. III\-C; Fig\. 6\)\.B\[chen2024KnowledgeInformedWheelWear\]The authors refer to their method as“interpretable feature engineering,”relying on spectral denoising and extraction of the wheel perimeter\-related dominant frequency as the interpretable feature, which does not qualify as[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\(see Sec\. III\-A/B; Fig\. 8\)\.B\[chen2024KnowledgeEmbeddedAutoencoder\]Although integrated into the latent space of a convolutional[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1), the prior knowledge used solely consists of trivial statistical and frequency\-domain features, which does not represent prior physical knowledge \(see Sec\. III; Fig\. 1\)\.B\[johannesexenberger2024GeneralizableTemperatureNowcasting\]Although the work is framed as relevant to predictive maintenance, it is limited to now\-casting bearing temperature and does not perform any actual[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks, leaving the connection to[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)superficial \(see Sec\. 1/2\)\.A\[fernandez2024Trainingofphysicsinformed\]Although technically representing a physics\-informed approach, the work is excluded as it solely targets end\-of\-discharge prediction, where“further research should \[be\] undertaken about the aging effect in Li\-ion batteries”\(see Sec\. 4\)\.A\[fernandez2024Physicsguidedrecurrentneural\]While the proposed approach is indeed physics\-informed, it specifically addresses accelerations in concrete buildings under seismic events, thereby illustrating a scope mismatch \(see Sec\. 4\.2\)\.B\[ge2024Domainadaptationfor\]While physics\-informed, the work falls outside the scope of this review, addressing structural health monitoring of civil infrastructure, i\.e\., steel beams and truss bridges \(see Sec\. 3/4\)\.A\[han2024AnInterpretableCNN\]A[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)with a built\-in wavelet feature\-extraction layer and reinforcement learning\-guided selection is built, where“the validation set accuracy \[…\] is taken as a priori knowledge,”obtained during pretraining, making the approach purely data\-driven \(see Sec\. II\-A; Fig\. 1; Eq\. 9\)\.A\[jia2024KneePointConsciousBatteryAging\]Extracting 17 statistical features from discharge, incremental capacity, and differential voltage curves, and augmenting them with degradation indicators from open\-circuit voltage reconstruction and hybrid pulse power characterization testing, constitutes mere feature engineering \(see Sec\. 3; Fig\. 4/5\)\.A\[feilongjiang2024SpatiotemporalAttentionbasedHidden\]While the approach may appear physics\-informed, it assumes \(without theoretical or domain\-specific justification\) a generic[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)on a learned latent state with empirically chosen derivative order, resulting in an arbitrary modeling choice \(see Sec\. 2\.3/3\.3\)\.A\[kayedpour2024WindTurbineHybrid\]The proposed“hybrid physics\-based deep learning framework”for wind turbine diagnosis is solely trained on data generated from a multiphysics simulation and is therefore excluded from this review by definition \(see Sec\. I\)\.B\[kim2024Singledomaingeneralizable\]Using“prior knowledge that the bearing fault signals are impulse excitation signals”for signal processing, followed by a vanilla[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)and an explainable artificial intelligence technique, the method combines conventional feature engineering with a post\-hoc explanation \(see Fig\. 2; Sec\. 3\.2\)\.A\[kumari2024Efficientstochasticparametric\]By definition, training solely on simulated data, with no real\-world data incorporated—as exemplified here by a stochastic battery degradation model—does not qualify as[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\(see Fig\. 1\)\.B\[lai2024PhysicsInformeddeepAutoencoder\]Due to the fact that the method for fault detection in electro\-hydraulic servo actuators used in turbofan engine fuel systems involves solely simulated data for the development of the reconstruction model, the work is excluded by definition \(see Sec\. 3\.1\)\.A\[li2024Particlefilterbasedfatiguedamageprognosisbyfusingmultipledegradationmodels\]Since the data\-driven components are limited to fixed low\-order polynomial regressions identified from experimental and simulated data, with no explicit ML model being employed, the study ultimately lies outside the[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)scope \(see Tab\. 3\)\.A\[shang2024Anoveldata\]By aligning run\-to\-failure sequences with dynamic time warping and averaging them via weighted barycenter to create additional time series, the approach is an interpolative, correlation\-aware yet standard data augmentation method \(see Fig\. 2/4; Sec\. 3\.1\)\.A\[su2024Knowledgeinformeddeepnetworks\]Although labeled“knowledge\-informed”, the so\-called knowledge\-based features are merely standard time\- and frequency\-domain statistics, effectively amounting to conventional feature engineering \(see Sec\. 2\.1; Fig\. 2\)\.B\[shengyutao2024NondestructiveDegradationPattern\]The proposed“ultra\-early prototype verification method”focuses on post\-production quality assessment of lithium\-ion batteries and is therefore excluded, as it does not implement operational[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\(see Fig\. 1; Discussions\)\.A\[xie2024Aknowledgedistillation\]The teacher Transformer is pre\-trained to learn degradation patterns and provide“prior knowledge of degradation laws”to the student[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)—knowledge that is merely learned from data \(see Fig\. 1/2; Methodology\)\.A\[yan2024DiscriminationandSparsityDrivenWeightOriented\]The method weights spectral lines using a generalized Rayleigh quotient eigenproblem, fuses each spectrum into a single“degradation feature,”and classifies them by Euclidean distance—reflecting classical linear algebra rather than[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\(see Sec\. II\-C/D\)\.A\[yan2024NovelAnchorDiscrimination\]With the method being“formulated as a generalized Rayleigh quotient, which can be conveniently and easily solved by a maximum eigenvalue problem,”it exemplifies classical linear algebra rather than[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\(see Sec\. II\-C; Eq\. 11\)\.B\[yan2024PhysicsEnhancedNMFToward\]This study is excluded because it investigates high\-speed train carriages, which are considered transportation systems and are therefore outside the scope of this review \(see Fig\. 1; Sec\. VII\)\.A\[yang2024Detectionofwind\]The proposed method, embedding a vanillaNeural ODE\(a purely data\-driven model that learns dynamics from data\) into an[AE](https://arxiv.org/html/2608.10047#p6.1.1.1.1)for anomaly detection, does not incorporate prior knowledge and is therefore not physics\-informed \(see Sec\. 2\.2/3\.1\.2\)\.A\[ye2024Amethodfor\]Ambiguous“secondary training”protocols \(testing/validation\), altering the physics constraint mid\-work, and misreporting physical loss for the baseline neural network undermine the study’s scientific rigor and comparability \(see Sec\. 2\.3/3; Eq\. 8/13; Fig\. 8\)\.A\[zhou2024PriorKnowledgeAugmentedMetaLearning\]While a“prior knowledge\-augmented”meta\-learning framework is proposed, it merely uses generic time\-domain indicators as pseudolabels and heatmaps \(gradient\-weighted class activation mapping\) fused at test time, not constituting a physics\-informed approach \(see Fig\. 1/2; Tab\. 1; Sec\. III\-A\)\.A\[abadi2023PhysicsInformedDeepLearningBased\]While appearing in the proceedings of the annual[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)Society conference, it constitutes a PhD research outline for the doctoral symposium track rather than a full technical paper, and is therefore excluded \(see Sec\. 3\)\.A\[cvijic2023NeedforAI\]Although the approach is basically suitable for diagnosis, prognosis, and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)prediction of transformers, no quantitative results incorporating ground truth values are reported for the full framework \(see Sec\. 4\)\.B\[divyanshidwivedi2023DynamoPMU:APhysics\]Despite being labeled“physics\-inspired,”the method for detecting anomalous events is entirely data\-driven, relying solely on“μ\\muPMU measurement data without any information about the network model or prior labeling of the events”\(see Abstract; Sec\. I‑B\)\.B\[fricke2023MissionSpecificPrognosisof\]Even though it appears in the proceedings of the annual[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)Society conference, it represents a PhD research outline submitted to the doctoral symposium track rather than a full technical paper, and is therefore excluded \(see Sec\. 3\)\.A\[furlong2023APhysicsinformedTransfer\]Although published in the proceedings of the annual[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)Society conference, it constitutes a PhD research outline submitted to the doctoral symposium track rather than a technical paper, and is therefore excluded \(see Sec\. 3\)\.A\[garpelli2023Physicsguidedneuralnetworks\]Although the approach embeds physics via a residual loss, it is excluded by definition because“simulated data is used during the learning process of the neural network, while experimental data is employed for the testing phase”\(see Conclusion\)\.B\[alfonsogijon2023Predictionofwind\]The approach targets wind turbine power modeling, explicitly framed as a“previous step to developing optimal controllers and applying failure detection methods,”resulting in a scope mismatch, as no actual[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)tasks are performed \(see Sec\. 1\)\.A\[he2023Asystematicmethod\]While presented as a“novel physics\-informed loss function,”the additional term simply penalizes overestimation of[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)more than underestimation for safety or economic reasons, which does not constitute prior physical knowledge \(see Sec\. 2\.2\)\.B\[he2023Multiaxialfatiguelife\]This approach is designed for offline fatigue life prediction, mapping from load parameters \(stress/strain\) to the cycle\-based fatigue life\. The current system health state is not considered\. Therefore, this method doesn’t qualify as[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\(see Sec\. 1; Fig\. 3\)\.B\[koutsoupakis2023Machinelearningbased\]By definition, the method does not qualify as[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35), since the[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)“is trained on numerical data alone,”generated via simulation for the purpose of identifying damage across different health states \(see Sec\. 5\.1\)\.B\[kumar2023EstimatingRemainingUseful\]Contrary to the claim of proposing a novel approach, ”which integrates machine learning techniques with electrochemical modeling,” the implemented method is a conventional neural network, reflecting standard deep learning \(see Sec\. 2/4\)\.A\[lee2023Developmentofthe\]While motivated by[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), the work does not directly address[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), as the fault detection and diagnosis methodology“that uses the developed estimation model will be proposed in the future,”resulting in a scope mismatch \(see Sec\. 5\)\.A\[lei2023Priorknowledgeembeddedmetatransfer\]“Embedding prior knowledge”is limited to computed order tracking\-based data augmentation—resampling vibration signals via pseudo‑speed ratios with amplitude scaling and Gaussian noise—followed by a standard metric‑based meta‑learner \(see Sec\. 3\.1\)\.A\[liao2023Remainingusefullife\]With the underlying physics “completely unknown, a deepHPM can be used to approximate it,” the method effectively infers a[PDE](https://arxiv.org/html/2608.10047#p6.31.31.31.31)\-like relationship between the hidden state and[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41)from data, rather than integrating prior physical knowledge \(see Sec\. 2\.2; Fig\. 2\)\.B\[liu2023KnowledgeEmbeddedLightweight\]A method is proposed,“where the prior knowledge of machining parameters is fused with features extracted from multiple sensor information”to construct“hybrid texture data”fed into a vision Transformer, i\.e\., feature engineering \(see Sec\. 2\.2; Fig\. 1/2\)\.A\[ma2023PhysicsInformedMachineLearning\]While both“physics‑informed”feature extraction and parameter tuning are claimed, the former reduces to selecting leakage‑related inputs and adopting rise time as the health indicator, and the latter is merely a vanilla hyperparameter search \(see Sec\. 2\.1/2\.2/3\.2\)\.A\[mochammad2023EnhancingRealisticRemaining\]The work lacks key methodological details \(e\.g\., low\-fidelity model identification, datasets, training procedure, used loss function, and questionable baseline comparisons\), preventing a rigorous assessment of its contribution \(see Sec\. 2\)\.B\[tu2023Integratingphysicsbasedmodeling\]Although a series of hybrid models are proposed, with the sole focus of enabling highly accurate voltage prediction for lithium\-batteries, the work does not directly address[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)and is therefore excluded \(see Sec\. 2\.1; Fig\. 2\)\.A\[wang2023InherentlyInterpretablePhysicsInformed\]The approach is indeed[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\. As it is used for end\-of\-discharge prediction of lithium\-ion batteries, hence not taking aging effects into account, it cannot be seen as[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\(see Sec\. IV\-F\)\.B\[wang2023Interpretableconvolutionalneural\]The method is a standard supervised[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)trained with cross\-entropy, merely embedding discrete wavelet transform and attention as signal\-processing layers to improve feature learning and noise robustness \(see Fig\. 2; Sec\. 3\.5\)\.A\[weddle2023Batterystateofhealthdiagnostics\]By definition, training \(in this case\) VGG\-16 solely on simulated battery degradation does not qualify as[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)due to the absence of real\-world measurements \(see Sec\. 3\.4; Fig\. 5\)\.A\[wu2023Physicsinformedgatedrecurrent\]Given that the“Secure Water Treatment testbed \(SWaT\) and Water Distribution testbed \(WADI\)”datasets are designed for cybersecurity research, the anomaly detection task does not target asset health, placing the work outside the scope of[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)\(see Sec\. 4\.1\)\.B\[zhang2023DuAK:ReinforcementLearningBased\]While industrial, the work focuses on quality inspection of manufactured steel products rather than the health of industrial assets, which is the core concern of[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34), resulting in an inherent scope mismatch \(see Sec\. I/II\-A\)\.A\[ariaschao2022Fusingphysicsbasedand\]While theoretically physics\-informed, the use of“a discrete\-time counterpart of the physics\-based model F in the form of a deep neural network”effectively yields a purely data\-driven approach \(see Sec\. 3\.1; Eq\. 5; Fig\. 3\)\.A\[chen2022PhysicsInformedLSTMhyperparameters\]The optimal parameters of the gearbox fault detection model considered in this study are determined by superimposing crack fault signatures onto validation data for hyperparameter optimization\. This is a data augmentation technique \(see Sec\. 3\.1; Fig\. 3\)\.B\[hajiha2022Aphysicsregularizeddatadriven\]The proposed“physics\-regularized data\-driven approach”relies, upon closer examination, exclusively on classical probabilistic modeling rather than traditional[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27), revealing a scope mismatch \(see Sec\. 2\)\.A\[russell2022Physicsinformeddeeplearning\]While stating that“including a loss term during AE training that is sensitive to frequency content introduces a physically informed objective,”additional autocorrelation and Fast Fourier transform\-based losses are effectively equivalent to standard multi\-task learning \(see Sec\. 2\.4\)\.A\[zgraggen2022PhysicsInformedDeep\]Focusing on tracker faults that usually occur“when the tracker gets stuck at a certain orientation instead of tracking the sun,”this work addresses operational anomaly detection rather than asset health, and is therefore excluded \(see Sec\. 1\)\.A\[guo2021Parameteridentificationof\]In this approach, solely a[NN](https://arxiv.org/html/2608.10047#p6.29.29.29.29)is used to estimate battery health indicators\. A fractional\-order model reconstructs measured quantities based on the health indicators for validation purposes, yet not directly relevant to health assessment \(see Sec\. 4\.2/4\.3\)\.B\[keizers2021UnscentedKalmanFiltering\]Although the method aims to predict[RUL](https://arxiv.org/html/2608.10047#p6.41.41.41.41), its reliance on an Unscented Kalman Filter for online parameter and state estimation rather than on[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)leads to a fundamental mismatch in scope \(see Sec\. 3\.3/3\.4/4\.3\)\.A\[li2021Particlefilterbasedhybriddamageprognosisconsideringmeasurementbias\]The data\-driven part is limited to an offline fitted, low\-order polynomial measurement equation whose parameters remain fixed during prognosis, resulting in a scope mismatch due to the lack of[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)\(see Sec\. 3\.2; Eq\. 18\)\.A\[lyathakula2021Aprobabilisticfatigue\]The work targets probabilistic fatigue life prediction based on offline test data, intended to replace/reduce expensive fatigue testing in the design phase of adhesively bonded joints, essentially not performing any[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)task using in\-situ monitoring data \(see Sec\. 2\)\.A\[wang2021Digitaltwinenhanced\]By definition, an approach that relies entirely on simulated data from a digital twin to train a[CNN](https://arxiv.org/html/2608.10047#p6.8.8.8.8)for autoclave fault prediction, without using actual operational data, does not qualify as[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)\(see Sec\. 6\)\.B\[cofremartel2020Aphysicsinformeddeep\]This work is an abstract\-only contribution and is therefore excluded for providing insufficient methodological and empirical detail \(see[A physics\-informed deep learning approach for fatigue crack propagation](https://www.rpsonline.com.sg/proceedings/esrel2020/html/4973.xml), accessed on December 23, 2025\)\.B\[kobrich2020Physicsbaseddeep\]Targeting crack growth prediction, the work proposes training a deep learning model solely on data simulated using extended finite element method, which by definition precludes it from being a[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)approach due to the absence of real\-world measurements \(see Sec\. 4\.2\)\.A\[akkad2019Aphysicsbased\]Although it appears in the proceedings of the annual[PHM](https://arxiv.org/html/2608.10047#p6.34.34.34.34)Society conference, it reflects a PhD research outline submitted to the doctoral symposium track rather than a full technical paper, and is therefore excluded \(see Sec\. 3\)\.A\[sadoughi2019Adeeplearning\]Using“physics\-based feature extraction, which is based on conventional signal processing techniques in time and frequency domain”to obtain features such as root mean square, peak\-to\-peak, and kurtosis does not constitute a[PIML](https://arxiv.org/html/2608.10047#p6.35.35.35.35)approach \(see Sec\. I/II\-B\)\.A\[neerukatti2018Ahybridprognosis\]The proposed“hybrid prognosis model”for predicting crack propagation is solely trained on data generated from finite element simulations, and is therefore excluded from this review by definition \(see Prognosis model\)\.A\[liu2015Faultdiagnosisfor\]The work focuses on a solar‑assisted heat pump system and proposes a fault diagnosis method which is solely trained on incomplete simulation data, and is therefore excluded by definition \(see Abstract; Sec\. 5\)\.A\[kulkarni2013Physicsbaseddegradation\]The development of physics\-based degradation models is proposed to predict electrolytic capacitor aging under thermal overstress, using experimental data for calibration without employing any[ML](https://arxiv.org/html/2608.10047#p6.27.27.27.27)techniques \(see Sec\. 3\.2; Fig\. 2\)\.ASimilar Articles
Physics-Informed Machine Learning for Short-Term Flood Prediction
Researchers propose a Physics-Informed Machine Learning (PIML) framework that integrates hydrological constraints into an LSTM loss function to improve short-term flood forecasting, particularly in data-scarce regimes. A 'Trend Alignment' constraint enforcing consistency between precipitation and discharge trends improves Nash-Sutcliffe Efficiency and eliminates unphysical predictions during extreme events.
A PAC-Bayesian View of Generalisation for Physics-Informed Machine Learning
This paper develops a PAC-Bayesian framework for physics-informed machine learning, providing high-probability generalization guarantees for unbounded losses. It proposes a multi-task perspective that jointly handles data fidelity, PDE residuals, and boundary conditions, and introduces a self-bounding learning algorithm.
Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap
This review paper surveys the application of large models in battery prognostics and health management, addressing long-standing challenges and proposing a roadmap for future research in this domain.
Physics-Informed Machine Learning Under Small-Data Constraints: Lessons from Abrasive Waterjet Milling
This paper presents methodological contributions for physics-informed machine learning under small-data constraints, using an abrasive waterjet milling dataset of 155 points. It shows that data curation choices, evaluation design, and physics integration form matter significantly, with Gaussian Process variants outperforming other models.
PiDDM: Physics-Informed Differentiable Degradation Modeling for Lithium-Ion Battery State-of-Health Prediction
This paper introduces PiDDM, a physics-informed differentiable degradation modeling framework that embeds battery degradation kinetics into neural network training to improve lithium-ion battery state-of-health prediction accuracy and physical consistency across diverse cycling protocols.