Predicting Steel Fatigue Life from Micrographs Using Physics-Informed Deep Learning

arXiv cs.LG Papers

Summary

This paper introduces FatigueCV, a physics-informed deep learning framework that predicts steel fatigue life from optical micrographs in under 65ms, using a CNN with uncertainty estimation. Validation on synthetic micrographs shows strong performance (R²=0.93), though real-world validation is noted as future work.

arXiv:2607.28695v1 Announce Type: new Abstract: Here is the plain text version optimized for arXiv's submission form. Custom macros (like \CV and \SI) have been converted to standard text/math so they render correctly on the webpage: Evaluating the fatigue life of structural steels conventionally requires mechanical testing lasting tens to hundreds of hours, making it impractical for rapid quality control. We present CV, a computer vision framework that estimates the fatigue life ($\log N_f$) of lightweight alloy steels directly from optical micrographs without physical testing.The pipeline features a seven-stage OpenCV preprocessing routine to remove artifacts, a 28-dimensional physics-informed feature extractor (quantifying crack morphology, grain structure, porosity, and texture), and a CNN regression model trained with a Gaussian negative log-likelihood (GNLL) loss to jointly predict $\log N_f$ and sample-specific uncertainty $\hat{\sigma}$.Evaluating three architectures (SE-CNN, ResNet-50, VGG-16) on a synthetic micrograph benchmark, ResNet-50 achieves $R^2 = 0.93$, RMSE = 0.18 log-cycles, and macro-F1 = 0.91. The GNLL objective reduces Expected Calibration Error by 76% compared to a mean-squared-error baseline (ECE: $0.089 \rightarrow 0.021$). Grad-CAM maps confirm the network attends to metallurgically meaningful microstructural features.Running in under 65 ms per image, the pipeline and synthetic dataset generator are open-sourced. Because validation relies entirely on synthetic micrographs, these results demonstrate methodological soundness under simulated conditions; a domain-transfer study on real field samples is the immediate next step.
Original Article
View Cached Full Text

Cached at: 08/03/26, 07:31 AM

# Predicting Steel Fatigue Life from Micrographs Using Physics-Informed Deep Learning
Source: [https://arxiv.org/html/2607.28695](https://arxiv.org/html/2607.28695)
###### Abstract

Evaluating the fatigue life of structural steels conventionally requires mechanical testing lasting tens to hundreds of hours, making it impractical for rapid quality control\. We presentFatigueCV, a computer\-vision framework estimating the fatigue life \(log10⁡\(Nf\)\\log\_\{10\}\(N\_\{f\}\)\{\}\) of lightweight alloy steels directly from optical micrographs without physical testing\.

The pipeline features a seven\-stage OpenCV preprocessing routine to remove artifacts, a 28\-dimensional physics\-informed feature extractor \(quantifying crack morphology, grain structure, porosity, and texture\), and a CNN regression model trained with a Gaussian negative log\-likelihood \(GNLL\) loss to jointly predictlog10⁡\(Nf\)\\log\_\{10\}\(N\_\{f\}\)\{\}and sample\-specific uncertaintyσ^\\hat\{\\sigma\}\.

Evaluating three architectures \(SE\-CNN, ResNet\-50, VGG\-16\) on a synthetic micrograph benchmark, ResNet\-50 achievesR2=0\.93R^\{2\}=0\.93, RMSE = 0\.18 log\-cycles, and macro\-F1 = 0\.91\. The GNLL objective reduces Expected Calibration Error by76%/76\\text\{\\,\}\\mathrm\{\\char 37\\relax\}\\text\{/\}compared to a mean\-squared\-error baseline \(ECE:0\.089→0\.0210\.089\\rightarrow 0\.021\)\. Grad\-CAM maps confirm the network attends to metallurgically meaningful microstructural features\.

Running in under65ms/65\\text\{\\,\}\\mathrm\{ms\}\\text\{/\}per image, the pipeline and synthetic dataset generator are open\-sourced\. Because validation relies entirely on synthetic micrographs, these results demonstrate methodological soundness under simulated conditions; a domain\-transfer study on real field samples is the immediate next step\.

## IIntroduction

Fatigue accounts for an estimated 50–90 % of all in\-service structural failures\[[1](https://arxiv.org/html/2607.28695#bib.bib1)\]\. Accurate knowledge of the remaining fatigue lifeNfN\_\{f\}\{\}is therefore essential for both safety assurance and lifecycle\-cost management in lightweight\-critical sectors such as rail, automotive, and aerospace engineering\. Despite this importance, the load\-controlled S–N test remains the dominant means of determiningNfN\_\{f\}\{\}, even though it is destructive, slow \(50–500 hours per specimen\), and blind to the spatial microstructural variability that ultimately governs failure\.

Optical microscopy, in contrast, is already a routine step in steel quality control\. Every micrograph implicitly encodes much of the microstructural state that drives the fatigue response: grain\-size distribution, crack morphology, porosity, and precipitate density\. What remains unresolved is how to convert this rich but qualitative visual information into a quantitative, calibrated estimate of fatigue life\.

*This paper addresses that gap\.*We presentFatigueCV, an end\-to\-end system that predictslog10⁡\(Nf\)\\log\_\{10\}\(N\_\{f\}\)\{\}with calibrated uncertainty directly from a steel micrograph in under65ms/65\\text\{\\,\}\\mathrm\{ms\}\\text\{/\}\. Because no public dataset pairs micrographs with ground\-truth fatigue life, the system is developed and evaluated on a physics\-constrained synthetic benchmark, and every result in this paper should be read in that light\. Our contributions are:

1. C1\.A seven\-stage OpenCV preprocessing pipeline engineered specifically for optical metallography artifacts\.
2. C2\.A 28\-dimensional physics\-informed feature vector whose components map directly to known fatigue damage mechanisms\.
3. C3\.GNLL\-based heteroscedastic CNN regression that produces calibrated per\-sample 95 % confidence intervals\.
4. C4\.A systematic Grad\-CAM validation procedure that correlates network attention with fatigue damage stages\.
5. C5\.A configurable, physics\-labeled synthetic microscopy dataset generator, released as open\-source software, together with an explicit discussion of the synthetic\-to\-real domain gap and the validation steps required before field deployment\.

## IIRelated Work

### II\-AClassical Fatigue Life Models

The S–N \(Wöhler\) framework\[[2](https://arxiv.org/html/2607.28695#bib.bib2)\]underpins classical fatigue design\. The Morrow mean\-stress correction\[[3](https://arxiv.org/html/2607.28695#bib.bib3)\]and the Smith–Watson–Topper \(SWT\) parameter\[[4](https://arxiv.org/html/2607.28695#bib.bib4)\]extend it to non\-zero mean stress and multiaxial loading, respectively\. These models assume microstructural homogeneity and require material\-specific calibration constants, an assumption that breaks down for cast alloys and additively manufactured components in which porosity and grain heterogeneity dominate\. Microstructure\-sensitive crystal\-plasticity models\[[6](https://arxiv.org/html/2607.28695#bib.bib6),[5](https://arxiv.org/html/2607.28695#bib.bib5)\]address this limitation, but each prediction requires days of computation and detailed 3\-D grain\-orientation maps\.

### II\-BMachine Learning for Fatigue

Data\-driven approaches generally achieve lower prediction error than classical closed\-form models\. Liuet al\.\[[7](https://arxiv.org/html/2607.28695#bib.bib7)\]applied support vector regression to composition and mechanical properties, reportingR2=0\.87R^\{2\}=0\.87on aluminum\-alloy S–N data\. Agrawalet al\.\[[8](https://arxiv.org/html/2607.28695#bib.bib8)\]applied gradient boosting to elemental composition vectors, and DeCost and Holm\[[9](https://arxiv.org/html/2607.28695#bib.bib9)\]developed CNN\-based microstructure classification\. Both lines of work rely on*tabular*inputs and do not exploit image information directly\. Azimiet al\.\[[10](https://arxiv.org/html/2607.28695#bib.bib10)\]demonstrated pixel\-wise steel\-phase segmentation\. None of these studies quantify predictive uncertainty or connect visual microstructural features to a continuous mechanical\-property estimate\. To our knowledge, no prior work provides calibrated uncertainty bounds for end\-to\-end regression from micrograph to fatigue life\.

### II\-CUncertainty Quantification in Deep Learning

Kendall and Gal\[[11](https://arxiv.org/html/2607.28695#bib.bib11)\]formalized the distinction between epistemic \(model\) and aleatoric \(data\) uncertainty in deep networks, and showed that a GNLL objective can be used to learn per\-sample aleatoric uncertainty for regression tasks such as monocular depth estimation\. To our knowledge, this technique has not previously been applied to materials fatigue prediction\.

### II\-DPosition of This Work

FatigueCVis, to our knowledge, the first system to address image\-to\-life regression, calibrated uncertainty, and metallurgically grounded explainability jointly, within a single reproducible pipeline\.

## IIIMethodology

### III\-AOverview

Raw Micrograph\(steel sample\)7\-Stage Preprocessing\(image cleaning and enhancement\)Feature Extraction\(28\-D physics\-informed vector\)Dual\-Head CNN Backbone\(regression model\)Fatigue Life Estimateμ^±1\.96​σ^\\hat\{\\mu\}\\pm 1\.96\\,\\hat\{\\sigma\}Risk Category\(four\-tier classification\)Figure 1:Overview of theFatigueCVpipeline\. A raw steel micrograph passes through a seven\-stage preprocessing stack, a physics\-informed feature extractor, and a dual\-head CNN regression backbone to produce a calibrated fatigue life estimateμ^±1\.96​σ^\\hat\{\\mu\}\\pm 1\.96\\hat\{\\sigma\}together with a four\-tier risk classification\.TheFatigueCVpipeline \([fig\.˜1](https://arxiv.org/html/2607.28695#S3.F1)\) is organized as four sequential layers\.L1\(preprocessing\) standardizes raw micrographs;L2\(feature extraction\) distills them into a physically interpretable vector;L3\(CNN regression\) maps image and feature information to\(μ^,s^\)\(\\hat\{\\mu\},\\,\\hat\{s\}\); andL4mapsμ^\\hat\{\\mu\}to a discrete risk tier\.

### III\-BSynthetic Dataset Generation

No public dataset pairs steel fatigue micrographs with ground\-truthNfN\_\{f\}\{\}values\. We therefore built a physics\-constrained synthetic generator that produces labeled images as follows:

- •Voronoi tessellationfor grain structure \(30–120 grains per image, log\-normal size distribution\);
- •Crack simulation: a correlated random walk with angular varianceσθ2=0\.3​rad2\\sigma\_\{\\theta\}^\{2\}=0\.3\\,\\text\{rad\}^\{2\}and branching probabilitypb=0\.15p\_\{b\}=0\.15per step;
- •Void simulation: Poisson\-distributed circular pores with radiusr∼𝒰​\[1,8\]r\\sim\\mathcal\{U\}\[1,8\]px;
- •Second\-phase inclusions: ellipses with aspect ratio∼𝒰​\[1\.2,3\.0\]\\sim\\mathcal\{U\}\[1\.2,\\,3\.0\]\.

Labels are assigned by the physics\-motivated regression

log10⁡\(Nf\)=7\.0⏟reference−1\.2​c~⏟cracking−0\.8​ϕ⏟porosity−0\.5​log10⁡\(d50\+1\)⏟grain size \(Hall–Petch\)−0\.3​n~⏟inclusions\+ε⏟scatter,\\begin\{split\}\\log\_\{10\}\(N\_\{f\}\)&=\\underbrace\{7\.0\}\_\{\\text\{reference\}\}\-\\underbrace\{1\.2\\,\\tilde\{c\}\}\_\{\\text\{cracking\}\}\-\\underbrace\{0\.8\\,\\phi\}\_\{\\text\{porosity\}\}\\\\ &\\quad\-\\underbrace\{0\.5\\log\_\{10\}\\\!\\left\(\\tfrac\{d\}\{50\}\+1\\right\)\}\_\{\\text\{grain size \(Hall\-\-Petch\)\}\}\-\\underbrace\{0\.3\\,\\tilde\{n\}\}\_\{\\text\{inclusions\}\}\+\\underbrace\{\\varepsilon\}\_\{\\text\{scatter\}\},\\end\{split\}\(1\)wherec~\\tilde\{c\}is a normalized crack\-severity index \(c~=Nc​s¯/A\\tilde\{c\}=N\_\{c\}\\bar\{s\}/A, withs¯\\bar\{s\}the mean crack severity\),ϕ\\phiis the areal void fraction,ddis the mean grain diameter inµ​m/\\mathrm\{\\SIUnitSymbolMicro m\}\\text\{/\},n~\\tilde\{n\}is the normalized inclusion density, andε∼𝒩​\(0,0\.15\)\\varepsilon\\sim\\mathcal\{N\}\(0,0\.15\)represents inherent material scatter\[[3](https://arxiv.org/html/2607.28695#bib.bib3)\]\. The Hall–Petch grain\-boundary term\[[19](https://arxiv.org/html/2607.28695#bib.bib19)\]is embedded directly in the grain\-size contribution\. The resulting dataset spanslog10⁡\(Nf\)∈\[3\.5,8\.5\]\\log\_\{10\}\(N\_\{f\}\)\\in\[3\.5,\\,8\.5\]\.

Because Eq\. \([1](https://arxiv.org/html/2607.28695#S3.E1)\) defines the ground truth used to train and evaluate the network, the reported metrics in[section˜IV](https://arxiv.org/html/2607.28695#S4)measure how well the CNN recovers a known, simulated physics relationship from images, rather than how well it predicts fatigue life on physical steel specimens\. We treat this synthetic benchmark as a controlled testbed for the modeling and uncertainty\-quantification methodology, and we return to this point in[section˜VI](https://arxiv.org/html/2607.28695#S6)\.

### III\-CPreprocessing Pipeline

Raw steel micrographs are affected by illumination gradients, scale\-bar borders, sensor noise, and low crack contrast\. The seven\-stage pipeline in[table˜I](https://arxiv.org/html/2607.28695#S3.T1)addresses each artifact class independently before any learned model is applied\.

TABLE I:Seven\-stage OpenCV preprocessing pipeline\.
### III\-DPhysics\-Informed Feature Extraction

Each preprocessed image is mapped to a 28\-dimensional feature vector𝒙∈ℝ28\\bm\{x\}\\in\\mathbb\{R\}^\{28\}\([table˜II](https://arxiv.org/html/2607.28695#S3.T2)\)\. Features are grouped into five physically motivated categories corresponding to established fatigue damage mechanisms\[[1](https://arxiv.org/html/2607.28695#bib.bib1),[6](https://arxiv.org/html/2607.28695#bib.bib6)\]\.

TABLE II:Feature categories, dimensionality, and physical basis\.CategoryddPhysical BasisCrack morphology5NcN\_\{c\},LcL\_\{c\},ρc\\rho\_\{c\},w¯\\bar\{w\}, and branching indexβ\\beta: directly govern crack\-initiation life\[[1](https://arxiv.org/html/2607.28695#bib.bib1),[16](https://arxiv.org/html/2607.28695#bib.bib16)\]Grain structure5Count, mean diameterd¯\\bar\{d\}, standard deviation, aspect ratio, and coefficient of variationCVd\\text\{CV\}\_\{d\}: Hall–Petch grain\-boundary strengthening\[[19](https://arxiv.org/html/2607.28695#bib.bib19)\]Porosity4Pore count, void fractionϕ\\phi, mean pore area, anddmaxd\_\{\\max\}: pores act as stress concentrators and initiation sitesTexture \(GLCM\)6Contrast, energy, homogeneity, entropy, mean intensity, andσI\\sigma\_\{I\}: encode phase\-boundary sharpnessGradient and fractal8Edge density, Sobel\-magnitude mean and standard deviation, LBP entropy and uniformity, and box\-counting fractal dimensionDfD\_\{f\}NcN\_\{c\}: crack count;LcL\_\{c\}: total length;ρc\\rho\_\{c\}: areal fraction;w¯\\bar\{w\}: mean width;DfD\_\{f\}: fractal dimension\.Fractal dimension is estimated via box\-counting on the binarized crack mapℳc\\mathcal\{M\}\_\{c\}:

Df=−limϵ→0log⁡N​\(ϵ\)log⁡ϵ,D\_\{f\}=\-\\lim\_\{\\epsilon\\to 0\}\\frac\{\\log N\(\\epsilon\)\}\{\\log\\epsilon\},\(2\)approximated over box sizesϵ∈\{2,4,8,16,32\}\\epsilon\\in\\\{2,4,8,16,32\\\}px\. A higherDfD\_\{f\}indicates a more branched crack network and, in this synthetic setting, correlates strongly by construction with reduced fatigue life\[[1](https://arxiv.org/html/2607.28695#bib.bib1)\]\.

### III\-ECNN Architectures

Backbone Network\(ResNet\-50 / SE\-CNN / VGG\-16\)Initial layers frozenSE attention blocksGlobal Average Poolingreduces spatial dimensionsShared FC Layerextracts common featuresMean Head\(μ^\\hat\{\\mu\}\)predictslog10⁡\(Nf\)\\log\_\{10\}\(N\_\{f\}\)\{\}Log\-Variance Head\(s^\\hat\{s\}\)predictss^=log⁡σ^2\\hat\{s\}=\\log\\hat\{\\sigma\}^\{2\}Joint Training via GNLL Loss\(Gaussian negative log\-likelihood\)

Figure 2:Dual\-head CNN architecture shared across all three backbones\. The mean headμ^\\hat\{\\mu\}predictslog10⁡\(Nf\)\\log\_\{10\}\(N\_\{f\}\)\{\}; the log\-variance heads^=log⁡σ^2\\hat\{s\}=\\log\\hat\{\\sigma\}^\{2\}encodes per\-sample aleatoric uncertainty\. Both heads share a common feature trunk and are trained jointly via the GNLL loss \([eq\.˜6](https://arxiv.org/html/2607.28695#S3.E6)\)\.All three backbones share the dual\-head design shown in[fig\.˜2](https://arxiv.org/html/2607.28695#S3.F2): amean headfμ:𝒳→ℝf\_\{\\mu\}:\\mathcal\{X\}\\to\\mathbb\{R\}predictinglog10⁡\(Nf\)\\log\_\{10\}\(N\_\{f\}\)\{\}, and alog\-variance headfs:𝒳→ℝf\_\{s\}:\\mathcal\{X\}\\to\\mathbb\{R\}encoding per\-sample aleatoric uncertaintys^=log⁡σ^2\\hat\{s\}=\\log\\hat\{\\sigma\}^\{2\}\.

#### III\-E1SE\-CNN \(FatigueCNN\)

Five convolutional blocks with Squeeze\-and\-Excitation attention\[[15](https://arxiv.org/html/2607.28695#bib.bib15)\]after blocks 3–5:

𝒛k=𝑭k⊙σg​\(𝑾2​δ​\(𝑾1​GAP​\(𝑭k\)\)\),\\bm\{z\}\_\{k\}=\\bm\{F\}\_\{k\}\\odot\\sigma\_\{g\}\\\!\\left\(\\bm\{W\}\_\{2\}\\,\\delta\\\!\\left\(\\bm\{W\}\_\{1\}\\,\\text\{GAP\}\(\\bm\{F\}\_\{k\}\)\\right\)\\right\),\(3\)where𝑭k\\bm\{F\}\_\{k\}is thekk\-th block feature map,σg\\sigma\_\{g\}is the sigmoid,δ\\deltais ReLU, and𝑾1,𝑾2\\bm\{W\}\_\{1\},\\bm\{W\}\_\{2\}are the excitation weights with reduction ratior=16r=16\. Feature fusion uses parallel global average and global max pooling followed by concatenation\.8\.2 M parameters,28ms/28\\text\{\\,\}\\mathrm\{ms\}\\text\{/\}inference\.

#### III\-E2ResNet\-50

Pretrained on ImageNet\[[12](https://arxiv.org/html/2607.28695#bib.bib12)\]\. Stages 1–2 are frozen to retain low\-level texture transfer from natural images, while stages 3–4 and a custom regression head are fine\-tuned:

y^=𝑾4​LayerNorm​\(δ​\(𝑾3​GAP​\(𝑭L​4\)\)\),\\hat\{y\}=\\bm\{W\}\_\{4\}\\,\\text\{LayerNorm\}\\\!\\left\(\\delta\\\!\\left\(\\bm\{W\}\_\{3\}\\,\\text\{GAP\}\(\\bm\{F\}\_\{L4\}\)\\right\)\\right\),\(4\)with𝑾3∈ℝ512×2048\\bm\{W\}\_\{3\}\\in\\mathbb\{R\}^\{512\\times 2048\}and𝑾4∈ℝ2×128\\bm\{W\}\_\{4\}\\in\\mathbb\{R\}^\{2\\times 128\}\.23\.5 M parameters;44ms/44\\text\{\\,\}\\mathrm\{ms\}\\text\{/\}inference; bestR2=0\.93R^\{2\}=0\.93\.

#### III\-E3VGG\-16

Pretrained on ImageNet\[[13](https://arxiv.org/html/2607.28695#bib.bib13)\]; convolutional blocks 1–3 are frozen\. Adaptive average pooling to4×44\\times 4yields an 8192\-dimensional embedding fed to a batch\-normalized regression stack\.138 M parameters;98ms/98\\text\{\\,\}\\mathrm\{ms\}\\text\{/\}inference;R2=0\.91R^\{2\}=0\.91\.

### III\-FGaussian Negative Log\-Likelihood Training

The GNLL loss trains the network to predict the parameters of a Gaussian distribution overlog10⁡\(Nf\)\\log\_\{10\}\(N\_\{f\}\)\{\}:

p​\(y∣μ^,σ^2\)\\displaystyle p\\\!\\left\(y\\mid\\hat\{\\mu\},\\hat\{\\sigma\}^\{2\}\\right\)=𝒩​\(y;μ^,σ^2\),\\displaystyle=\\mathcal\{N\}\\\!\\left\(y;\\,\\hat\{\\mu\},\\,\\hat\{\\sigma\}^\{2\}\\right\),\(5\)ℒGNLL​\(μ^,s^,y\)\\displaystyle\\mathcal\{L\}\_\{\\text\{GNLL\}\}\(\\hat\{\\mu\},\\hat\{s\},y\)=12​\[e−s^​\(y−μ^\)2\+s^\],\\displaystyle=\\frac\{1\}\{2\}\\\!\\left\[e^\{\-\\hat\{s\}\}\\,\\left\(y\-\\hat\{\\mu\}\\right\)^\{2\}\+\\hat\{s\}\\right\],\(6\)wheres^=log⁡σ^2\\hat\{s\}=\\log\\hat\{\\sigma\}^\{2\}\. The two terms in this loss are in tension: the first term penalizes over\-confidence by rewarding a larger predicted variance when the residual is large, while the second term penalizes an unnecessarily large predicted variance\. At the optimum, the network learns an input\-dependent aleatoric uncertainty that is larger for more heavily damaged specimens\. This behavior is consistent with the well\-documented increase in fatigue\-life scatter at short lives\[[3](https://arxiv.org/html/2607.28695#bib.bib3)\]\.

The predicted 95 % confidence interval is

μ^±1\.96​σ^,\\hat\{\\mu\}\\pm 1\.96\\,\\hat\{\\sigma\},\(7\)expressed in log\-cycle units\. An MSE\-trained baseline, by contrast, yields only a single globalσ^global\\hat\{\\sigma\}\_\{\\text\{global\}\}estimated from the residual distribution and cannot adapt this estimate on a per\-sample basis\.

### III\-GTraining Protocol

TABLE III:Training hyperparameters\.
### III\-HHardware and Software

All experiments were run in a simulated training environment using a single GPU\-backed instance with 16 GB of device memory\. The codebase is implemented in PyTorch, with OpenCV 4\.x for preprocessing and scikit\-learn for the SHAP and GBM baselines reported in[section˜IV](https://arxiv.org/html/2607.28695#S4)\. Total wall\-clock time for training all three backbones on the synthetic dataset was under six GPU\-hours\. Exact package versions and a pinnedrequirements\.txtare provided with the released codebase to support exact reproduction of the results reported here\.

## IVResults and Discussion

All results in this section were obtained on the synthetic benchmark described in[section˜III\-B](https://arxiv.org/html/2607.28695#S3.SS2); see[section˜VI](https://arxiv.org/html/2607.28695#S6)for a discussion of what these results do and do not establish about performance on real steel micrographs\.

### IV\-APreprocessing Quality

[Table˜IV](https://arxiv.org/html/2607.28695#S4.T4)quantifies the improvement across six image quality metrics after the full seven\-stage pipeline, measured on a held\-out set of 200 simulated micrographs with synthetically injected illumination and noise artifacts\.

TABLE IV:Image quality before and after preprocessing\.The largest single gain is in crack\-detection recall \(61%→89%61\\,\\%\\rightarrow 89\\,\\%\), which we attribute to the multi\-orientation black\-hat morphology capturing cracks at all angles and to the Scharr operator’s stronger response to diagonal edges relative to the Sobel kernel\.

### IV\-BFeature Importance

SHAP analysis on a hybrid gradient\-boosted machine \(GBM\) baseline trained on the 28\-dimensional feature vector \([table˜V](https://arxiv.org/html/2607.28695#S4.T5)\) shows that crack\-related features account for more than 50 % of total predictive importance, consistent with fatigue\-mechanics theory\[[1](https://arxiv.org/html/2607.28695#bib.bib1)\]\. This is, by construction, expected given Eq\. \([1](https://arxiv.org/html/2607.28695#S3.E1)\), and serves primarily as a sanity check that the extracted features are consistent with the physics used to generate the labels\.

TABLE V:Top eight features by SHAP importance value\.FeatureSHAP \(%\)Mechanistic Linkcrack\_density23\.4DirectNfN\_\{f\}reductionfractal\_dimDfD\_\{f\}16\.8Crack network complexitytexture\_entropy12\.1Phase heterogeneityvoid\_fraction11\.7Initiation\-site densitygrain\_CVdd9\.3Stress concentration sitesbranch\_idxβ\\beta8\.9Advanced damage stagegradient\_std7\.2Boundary roughnesstexture\_contrast6\.4Phase\-boundary sharpnessRemaining 4\.2 % distributed among 20 features\.
### IV\-CRegression Performance

[Table˜VI](https://arxiv.org/html/2607.28695#S4.T6)compares all architectures on the held\-out test split\. ResNet\-50 achieves the best trade\-off between accuracy and parameter count\.

TABLE VI:Architecture comparison on held\-out test set \(200 images, 80/10/10 split\)\. Metrics inlog10\\log\_\{10\}\-cycle units\. Bold indicates the best value in each column\.ModelR2R^\{2\}RMSEMAEMBEParam\.msSE\-CNN \(ours\)0\.880\.240\.19\+0\.02\+0\.028\.2 M28ResNet\-50 \(ours\)0\.930\.180\.14−0\.01\-0\.0123\.5 M44VGG\-160\.910\.210\.17\+0\.03\+0\.03138 M98Hybrid CV\+GBM0\.860\.270\.22\+0\.05\+0\.05—8SVR\[[7](https://arxiv.org/html/2607.28695#bib.bib7)\]†0\.87—————†\\daggerLiuet al\.on aluminum\-alloy tabular data; included for reference only\.[Figure˜3](https://arxiv.org/html/2607.28695#S4.F3)visualizes the predicted\-versus\-actual relationship for ResNet\-50 on the synthetic test set\.

![Refer to caption](https://arxiv.org/html/2607.28695v1/fig_scatter.png)

Figure 3:Predicted vs\. actuallog10⁡\(Nf\)\\log\_\{10\}\(N\_\{f\}\)\{\}for ResNet\-50 on the test set\. Marker shape encodes risk category\. Error bars represent predicted 95 % confidence intervals\. The dashed line is the identity \(μ^=y\\hat\{\\mu\}=y\); the shaded band shows±\\pmRMSE\.R2=0\.93R^\{2\}=0\.93\.
### IV\-DUncertainty Calibration

A well\-calibrated model satisfiesP​\(\|y−μ^\|≤zα/2​σ^\)≈1−αP\(\|y\-\\hat\{\\mu\}\|\\leq z\_\{\\alpha/2\}\\,\\hat\{\\sigma\}\)\\approx 1\-\\alphafor allα\\alpha\. We evaluate this using the regression analogue of Expected Calibration Error:

ECE=∑b=1B\|Bb\|N​\|cov​\(Bb\)−conf​\(Bb\)\|,\\text\{ECE\}=\\sum\_\{b=1\}^\{B\}\\frac\{\|B\_\{b\}\|\}\{N\}\\left\|\\,\\text\{cov\}\(B\_\{b\}\)\-\\text\{conf\}\(B\_\{b\}\)\\right\|,\(8\)whereBbB\_\{b\}are equally spaced confidence bins andconf​\(Bb\)\\text\{conf\}\(B\_\{b\}\)is the nominal coverage level for that bin\.

GNLL vs\. MSE:ECEMSE=0\.089→ECEGNLL=0\.021\\text\{ECE\}\_\{\\text\{MSE\}\}=0\.089\\rightarrow\\text\{ECE\}\_\{\\text\{GNLL\}\}=0\.021, a76%/76\\text\{\\,\}\\mathrm\{\\char 37\\relax\}\\text\{/\}improvement\. The nominal 95 % confidence interval empirically contains 93\.8 % of test samples, close to the 95\.0 % target\.

Risk\-stratified uncertainty:CriticalandHighspecimens yieldσ^∈\[0\.22,0\.28\]\\hat\{\\sigma\}\\in\[0\.22,\\,0\.28\], versus\[0\.13,0\.16\]\[0\.13,\\,0\.16\]forLow\-risk specimens\. This pattern is consistent with the well\-known increase in fatigue scatter at short lives\[[3](https://arxiv.org/html/2607.28695#bib.bib3)\], though we note again that this consistency is partly guaranteed by the synthetic label model in Eq\. \([1](https://arxiv.org/html/2607.28695#S3.E1)\)\.[Figure˜4](https://arxiv.org/html/2607.28695#S4.F4)shows the corresponding calibration curve\.

![Refer to caption](https://arxiv.org/html/2607.28695v1/fig_calibration.png)

Figure 4:Calibration reliability diagram\. The GNLL\-trained model \(solid line\) closely tracks the ideal diagonal across confidence levels, while the MSE baseline \(dashed line\) systematically over\- or under\-covers\. ECE: 0\.021 \(GNLL\) vs\. 0\.089 \(MSE\)\.
### IV\-ERisk Classification

TABLE VII:Risk\-tier classification metrics \(ResNet\-50\)\.ACritical\-tier recall of 0\.90 is the most important figure for safety\-relevant deployment, since missed critical cases carry the highest cost\. Roughly 73 % of remaining classification errors occur between the adjacentMediumandHightiers, whoselog10⁡\(Nf\)\\log\_\{10\}\(N\_\{f\}\)ranges overlap within the predicted 95 % confidence interval; we regard this as a direct and expected consequence of calibrated uncertainty near a tier boundary, rather than a classification deficiency\.

### IV\-FGrad\-CAM Explainability

Grad\-CAM\[[14](https://arxiv.org/html/2607.28695#bib.bib14)\]computes spatial saliency as

LGradCAMc=ReLU​\(∑k1Z​∑i,j∂yc∂Ai​jk⏟αkc⋅Ak\),L^\{c\}\_\{\\text\{GradCAM\}\}=\\text\{ReLU\}\\\!\\left\(\\sum\_\{k\}\\underbrace\{\\frac\{1\}\{Z\}\\sum\_\{i,j\}\\frac\{\\partial y^\{c\}\}\{\\partial A^\{k\}\_\{ij\}\}\}\_\{\\alpha^\{c\}\_\{k\}\}\\cdot A^\{k\}\\right\),\(9\)whereAi​jkA^\{k\}\_\{ij\}is the\(i,j\)\(i,j\)\-th activation of thekk\-th feature map in the final convolutional layer andZZis the spatial dimension used for averaging\.

[Figure˜5](https://arxiv.org/html/2607.28695#S4.F5)shows representative activation maps for two risk tiers\. Across the four tiers we observe a saliency progression that is broadly consistent with the fatigue damage stages described by Suresh\[[1](https://arxiv.org/html/2607.28695#bib.bib1)\]:

![Refer to caption](https://arxiv.org/html/2607.28695v1/INSYN_0001.png)

\(a\)Medium: Stage I, nucleation
![Refer to caption](https://arxiv.org/html/2607.28695v1/Medium.png)

\(b\)Critical: Stage III, pervasive

Figure 5:Grad\-CAM activation maps overlaid on preprocessed micrographs for representative medium\- and critical\-risk samples \(ResNet\-50\)\. Darker overlay regions indicate higher saliency\. Damage\-stage annotations follow Suresh\[[1](https://arxiv.org/html/2607.28695#bib.bib1)\]\.- •Low: saliency concentrates on intact grain boundaries; the network associates an undamaged polycrystalline structure with long fatigue life \(Stage 0\)\.
- •Medium\([fig\.˜5\(a\)](https://arxiv.org/html/2607.28695#S4.F5.sf1)\): mixed saliency at grain boundaries and early crack\-nucleation sites, consistent with Stage I initiation at persistent slip bands\.
- •High: strong, localized activation at crack tips and branching junctions, consistent with the Stage II propagation front governed by Paris\-law crack growth\[[16](https://arxiv.org/html/2607.28695#bib.bib16)\]\.
- •Critical\([fig\.˜5\(b\)](https://arxiv.org/html/2607.28695#S4.F5.sf2)\): near\-uniform activation across the full crack network, consistent with Stage III \(fast fracture\) damage\.

This concordance between Grad\-CAM attention and metallurgically established damage stages is, on the synthetic benchmark, consistent with the network having learned representations aligned with the physics encoded in Eq\. \([1](https://arxiv.org/html/2607.28695#S3.E1)\), rather than spurious correlations specific to the rendering pipeline\. Whether this alignment transfers to real micrographs, where damage cues are visually noisier, remains to be tested\.

## VContributions and Novelty

C1 — Physics\-informed feature engineering:To our knowledge, no prior machine\-learning fatigue study extracts image\-based features derived explicitly from fatigue\-mechanics theory\. Our 28\-dimensional vector directly maps to crack initiation, Hall–Petch grain\-boundary strengthening, void\-induced stress concentration, and fractal crack\-propagation complexity\[[1](https://arxiv.org/html/2607.28695#bib.bib1),[6](https://arxiv.org/html/2607.28695#bib.bib6),[16](https://arxiv.org/html/2607.28695#bib.bib16),[19](https://arxiv.org/html/2607.28695#bib.bib19)\]\.

C2 — Heteroscedastic uncertainty for materials fatigue:This is, to our knowledge, the first application of per\-sample GNLL uncertainty to fatigue\-life prediction from microscopy\. The physically consistent, risk\-stratifiedσ^\\hat\{\\sigma\}values and the 76 % ECE improvement over the MSE baseline \(0\.089→0\.0210\.089\\rightarrow 0\.021\) support both the statistical and the physical validity of this approach on the synthetic benchmark\.

C3 — Metallography\-specific preprocessing:three components tailored to optical metallography: \(i\) multi\-orientation black\-hat morphology for near\-isotropic crack detection; \(ii\) rolling\-ball illumination correction for microscope vignetting; and \(iii\) Laplacian/SNR quality gating that prevents corrupted images from reaching the model\.

C4 — Fatigue\-mechanics Grad\-CAM validation:a systematic comparison of CNN saliency against four fatigue damage stages \(Stages 0–III\), offered as a reusable validation procedure for explainability studies in materials informatics\.

C5 — Physics\-labeled synthetic generator:an open, configurable synthetic\-data engine combining Voronoi grain tessellation, parametric crack and void modeling, and the labeling rule in Eq\. \([1](https://arxiv.org/html/2607.28695#S3.E1)\), intended to support repeatable benchmarking and pretraining/transfer studies pending validation on real micrographs\.

## VIDiscussion

On a harder task than prior work – image input rather than tabular features, and a broader material scope –FatigueCVoutperforms the strongest tabular ML baseline we are aware of \(R2=0\.87R^\{2\}=0\.87, Liuet al\.\[[7](https://arxiv.org/html/2607.28695#bib.bib7)\]\), reachingR2=0\.93R^\{2\}=0\.93\. We interpret this primarily as evidence that the GNLL objective and the physics\-guided feature set are a well\-matched inductive bias for this problem class, rather than as a claim of superiority on real\-world data, since the two studies use different materials and different data modalities\. The achieved calibration,ECE=0\.021\\text\{ECE\}=0\.021, is below the 0\.03–0\.05 range reported by Kendall and Gal\[[11](https://arxiv.org/html/2607.28695#bib.bib11)\]for depth estimation, though the two tasks are not directly comparable\.

We highlight four limitations, in order of importance:

1. 1\.*Synthetic\-to\-real domain gap\.*This is the central limitation of the present study\. All training and evaluation data are generated by the physics\-motivated simulator in[section˜III\-B](https://arxiv.org/html/2607.28695#S3.SS2); the model has not been exposed to real steel micrographs, real imaging noise, or real specimen preparation artifacts\. Direct application to field micrographs will likely require domain\-adaptive fine\-tuning \(for example, CycleGAN\-based style transfer or few\-shot calibration on a small real\-labeled set\) and re\-validation of both accuracy and calibration\.
2. 2\.*Material scope\.*The preprocessing pipeline and feature set are tuned for carbon and low\-alloy steels; titanium and aluminum alloys, which exhibit different microstructural signatures, would require retraining and likely some redesign of the feature vector\.
3. 3\.*2\-D projection\.*Optical microscopy captures only a surface section and cannot detect subsurface voids or cracks; integration with X\-ray computed tomography is planned as a complementary 3\-D input modality\.
4. 4\.*Life decomposition\.*NfN\_\{f\}is predicted holistically; damage\-tolerant design in practice often requires separating initiation life from propagation life, which the current formulation does not provide\.

Given these limitations, we position the results in[section˜IV](https://arxiv.org/html/2607.28695#S4)as a methodological proof of concept: they show that the proposed preprocessing, feature\-engineering, and heteroscedastic\-regression pipeline can recover a known, physics\-consistent fatigue relationship from images with high accuracy and well\-calibrated uncertainty\. Establishing predictive validity on physical steel specimens is left as future work and would require, at minimum, a paired dataset of real micrographs and experimentally measuredNfN\_\{f\}values\.

## VIIConclusion

We introducedFatigueCV, a computer\-vision pipeline for predicting the fatigue life of lightweight alloy steels from optical micrographs under simulated, physics\-constrained conditions\. Combining a metallography\-specific preprocessing stack, physics\-informed features, GNLL\-based heteroscedastic CNN regression, and fatigue\-mechanics\-validated Grad\-CAM analysis, the system reachesR2=0\.93R^\{2\}=0\.93, RMSE=0\.18=0\.18log\-cycles, ECE=0\.021=0\.021, and macro\-F1=0\.91=0\.91on a synthetic steel benchmark, outperforming the tabular machine\-learning baselines we compared against\. Taken together, the five contributions — physics\-informed features, calibrated heteroscedastic uncertainty, metallography\-specific preprocessing, fatigue\-mechanics Grad\-CAM validation, and an open synthetic\-data generator — provide a reproducible foundation for image\-based fatigue assessment\. The principal open question is whether these results transfer to real steel micrographs; validating the pipeline on experimentally measured fatigue data is the natural next step and is the direction we intend to pursue\.

## References

- \[1\]S\. Suresh,Fatigue of Materials, 2nd ed\. Cambridge Univ\. Press, 1998\.
- \[2\]A\. Wöhler, “Über die Festigkeitsversuche mit Eisen und Stahl,”Z\. Bauwesen, vol\. 20, pp\. 73–106, 1870\.
- \[3\]J\. Morrow, “Cyclic plastic strain energy and fatigue of metals,”ASTM STP 378, pp\. 45–87, 1965\.
- \[4\]K\. N\. Smith, P\. Watson, and T\. H\. Topper, “A stress\-strain function for the fatigue of metals,”J\. Mater\., vol\. 5, no\. 4, pp\. 767–778, 1970\.
- \[5\]Y\. Guilhemet al\., “Investigation of grain cluster effects on fatigue crack initiation in polycrystals,”Int\. J\. Fatigue, vol\. 32, no\. 11, pp\. 1748–1763, 2010\.
- \[6\]D\. L\. McDowell and F\. P\. E\. Dunne, “Microstructure\-sensitive computational modeling of fatigue crack formation,”Int\. J\. Fatigue, vol\. 32, no\. 9, pp\. 1521–1542, 2010\.
- \[7\]Z\. Liuet al\., “A machine learning approach to fatigue life prediction for Al alloys,”Mater\. Sci\. Eng\. A, vol\. 689, pp\. 211–219, 2017\.
- \[8\]A\. Agrawalet al\., “Deep materials informatics: Applications of deep learning in materials science,”MRS Commun\., vol\. 9, no\. 3, pp\. 779–792, 2019\.
- \[9\]B\. L\. DeCost and E\. A\. Holm, “A computer vision approach for automated analysis of microstructural image data,”Comput\. Mater\. Sci\., vol\. 110, pp\. 126–133, 2015\.
- \[10\]S\. M\. Azimiet al\., “Advanced steel microstructural classification by deep learning methods,”Sci\. Rep\., vol\. 8, p\. 2128, 2018\.
- \[11\]A\. Kendall and Y\. Gal, “What uncertainties do we need in Bayesian deep learning for computer vision?” inProc\. NeurIPS, 2017, pp\. 5574–5584\.
- \[12\]K\. He, X\. Zhang, S\. Ren, and J\. Sun, “Deep residual learning for image recognition,” inProc\. CVPR, 2016, pp\. 770–778\.
- \[13\]K\. Simonyan and A\. Zisserman, “Very deep convolutional networks for large\-scale image recognition,” inProc\. ICLR, 2015\.
- \[14\]R\. R\. Selvarajuet al\., “Grad\-CAM: Visual explanations from deep networks via gradient\-based localization,” inProc\. ICCV, 2017, pp\. 618–626\.
- \[15\]J\. Hu, L\. Shen, and G\. Sun, “Squeeze\-and\-excitation networks,” inProc\. CVPR, 2018, pp\. 7132–7141\.
- \[16\]P\. Paris and F\. Erdogan, “A critical analysis of crack propagation laws,”J\. Basic Eng\., vol\. 85, no\. 4, pp\. 528–533, 1963\.
- \[17\]I\. Loshchilov and F\. Hutter, “Decoupled weight decay regularization,” inProc\. ICLR, 2019\.
- \[18\]A\. Buslaevet al\., “Albumentations: Fast and flexible image augmentations,”Information, vol\. 11, no\. 2, p\. 125, 2020\.
- \[19\]E\. O\. Hall, “The deformation and ageing of mild steel: III,”Proc\. Phys\. Soc\. B, vol\. 64, no\. 9, pp\. 747–753, 1951\.

Similar Articles