A Physics-Chemistry-Informed Neural Network (PCINN) for Real-Time Spatial-ALD Coverage Prediction and Reliable Kinetics Inversion
Summary
This paper introduces PCINN, a physics-chemistry-informed neural network that acts as a hybrid AI surrogate for real-time spatial atomic layer deposition (SALD) coverage prediction, achieving CFD-level accuracy in ~77ms and enabling reliable kinetics inversion via identifiability analysis.
View Cached Full Text
Cached at: 08/04/26, 07:37 AM
# A Physics–Chemistry-Informed Neural Network (PCINN) for Real-Time Spatial-ALD Coverage Prediction and Reliable Kinetics Inversion A hybrid neural-ODE surrogate with hard-coded chemistry and a built-in identifiability diagnostic
Source: [https://arxiv.org/html/2608.00212](https://arxiv.org/html/2608.00212)
Ning HuCorresponding author\. E\-mail:jluhooning@163\.comSchool of Mechanical Engineering, Hangzhou Dianzi University, Hangzhou, ChinaChang LiuE\-mail:laiurenty@163\.comInformation Engineering School, Hangzhou Dianzi University, Hangzhou, ChinaYunlei JiangE\-mail:2142010020@hdu\.edu\.cnSchool of Mechanical Engineering, Hangzhou Dianzi University, Hangzhou, ChinaYuan DongE\-mail:Dongy@hdu\.edu\.cnSchool of Mechanical Engineering, Hangzhou Dianzi University, Hangzhou, China
###### Abstract
Spatial atomic layer deposition \(SALD\) is a leading atmospheric\-pressure, high\-throughput route to industrial ALD, yet its design and control are limited by the cost of predicting substrate surface coverage: high\-fidelity CFD is accurate but far too slow for operating\-window scans or online inference, while classical analytic models cannot capture transport modulation such as the gas curtain\. We present a physics–chemistry\-informed neural network \(PCINN\), a hybrid AI surrogate that delivers CFD\-level accuracy at real\-time speed\. A trained query returns coverage in about77ms—roughly5×1045\\times 10^\{4\}times faster than a CFD solve*per query*, the one\-time CFD training cost being amortised over many queries—reaching a testRlog2=0\.998R^\{2\}\_\{\\log\}=0\.998\(Rraw2=0\.989R^\{2\}\_\{\\mathrm\{raw\}\}=0\.989; leave\-one\-outRraw2=0\.974R^\{2\}\_\{\\mathrm\{raw\}\}=0\.974\) from only 30 training cases across coverage spanning about four orders of magnitude, with a known model\-form bias that under\-predicts the highest\-coverage cases by1313–36%36\\%\(compressed in log space\)\. This makes PCINN directly usable for SALD process design, operating\-window mapping, and control\-loop deployment\.
The architecture is deliberately not a black box: PCINN splits degrees of freedom by what is known versus unknown\. A small neural network learns only the operating\-condition→\\,\\to\\,effective near\-wall concentration transport closure \(the physics branch\), while the known surface kinetics is encoded as a hard\-coded, trainable chemistry layer and integrated along the substrate trajectory\. Compressing the data\-driven freedom to a single scalar preserves accuracy under sparse data while keeping the model interpretable and invertible\.
Because engineering deployment requires knowing the limits of an inference, we equip PCINN with a rigorous identifiability analysis\. Using initial\-value drift, supervision\-weight sweeps, the Fisher information matrix and profile likelihood, we find the adsorption energyEadsE\_\{\\mathrm\{ads\}\}and desorption ratekdesk\_\{\\mathrm\{des\}\}robustly identifiable, whereaskadsk\_\{\\mathrm\{ads\}\}is not separately identifiable at a single temperature \(onlykads⋅cwallk\_\{\\mathrm\{ads\}\}\\\!\\cdot\\\!c\_\{\\mathrm\{wall\}\}is\)\. Across four temperatures the pre\-exponential factorν\\nuandEadsE\_\{\\mathrm\{ads\}\}do*not*separate; they bind along a*weakly identifiable degeneracy valley*\(slope≈0\.065\\approx 0\.065eV/decade, ground truth on the line\), so multiple temperatures compress the pair into a weakly identifiable interval \(χ12\\chi^\{2\}\_\{1\}95%95\\%width≈0\.6\\approx 0\.6decade\) rather than breaking it—conditioning onν\\nuthen recoversEadsE\_\{\\mathrm\{ads\}\}to within0\.3%0\.3\\%\. We derive this slope analytically askBTeffln10k\_\{B\}T\_\{\\mathrm\{eff\}\}\\ln 10\(TeffT\_\{\\mathrm\{eff\}\}the harmonic mean of the sampled temperatures\), show it is independent of the prefactor parameterisation, and turn it into an operational reliability diagnostic: a seven\-chemistry, three\-mechanism mismatch matrix—extended by an energy\-split calibration sweep and a non\-Fickian transport\-layer mismatch, with bootstrap resampling and coarse\-grid sampling—confirms the slope is invariant under any single\-Arrhenius \(and the transport\) mismatch and shifts only when a second thermally activated process is introduced, so a slope departure is a falsifiable flag for unmodelled site heterogeneity\.
We emphasise that the data are generated by simulation from*known*ground truth and inverted with the*same*kinetic form; the identifiability study is therefore first a concept verification of pipeline self\-consistency and of the*identifiability boundary*, rather than a discovery of parameters of a real system\. PCINN thus provides an accurate, sample\-efficient, real\-time SALD coverage surrogate coupled to a kinetic\-inversion framework with clearly delineated identifiability limits and a transferable, physics\-derived reliability diagnostic\.
Keywords:engineering AI surrogate modeling; physics\-informed neural networks; hybrid gray\-box modeling; spatial atomic layer deposition; surface coverage prediction; adsorption\-kinetics inversion; parameter identifiability
## 1\. Introduction
### 1\.1\. Background
Atomic layer deposition \(ALD\) grows films layer by layer through the self\-limiting surface reactions of alternately dosed precursors, giving atomic\-scale thickness control and excellent conformality\. Spatial ALD \(SALD\) replaces temporal separation by spatial separation: the substrate moves between precursor zones separated by inert gas curtains, enabling atmospheric\-pressure, high\-throughput operation and is regarded as an important route to the industrialisation of ALD\. The film quality and process window of SALD are controlled by the substrate surface coverage and by the adsorption/desorption kinetics\. Predicting coverage as a function of operating conditions, and inverting the underlying kinetics from coverage observations, are therefore two central tasks for SALD design and control\.
### 1\.2\. Limitations of existing methods
Two families of models are commonly used\. High\-fidelity reaction–transport CFD couples flow, diffusion and surface reactions and accurately resolves the concentration and coverage fields, but the per\-case cost is high, which makes dense operating\-window scans or iterative inversion against data impractical\. Classical analytic ALD models are extremely cheap and useful for scaling analysis, but they rest on a sufficient\-supply assumption and cannot represent transport modulation such as the gas curtain\. A method that combines the accuracy of CFD, the efficiency of analytic models, and the interpretability required for kinetic inversion is therefore desirable\.
### 1\.3\. Approach and contributions
We propose a physics–chemistry\-informed neural network \(PCINN\)\. The key idea is to split degrees of freedom by what is known versus unknown: the expensive and hard\-to\-derive operating\-condition→\\,\\to\\,effective near\-wall concentration transport closure is assigned to a small neural network \(the physics branch\), while the known surface kinetics is written as a hard\-coded, trainable chemistry layer; the two are coupled through the near\-wall concentration and integrated along the substrate trajectory to yield coverage\. The data\-driven degrees of freedom are thereby compressed to a single scalar, preserving accuracy under sparse data while retaining an interpretable, invertible kinetic structure\.
The unifying thesis is thata physics\-constrained AI surrogate can be both fast enough for real\-time engineering use and self\-aware of its own inference limits: the precision boundary of its hybrid inversion is set by the degeneracy geometry of the parameter space, not by fitting power, and this degeneracy can be predicted analytically and tested statistically—giving the surrogate a built\-in, transferable reliability diagnostic\. The main contributions follow this line\.
1. 1\.A systematic, honest identifiability\-analysis framework \(the methodological core\)\.While invertingEadsE\_\{\\mathrm\{ads\}\}andkdesk\_\{\\mathrm\{des\}\}, we separate identifiable from non\-identifiable quantities using initial\-value drift, supervision\-weight sweeps, the Fisher information matrix and profile likelihood\. We are deliberately careful not to over\-claim novelty for the degeneracies themselves: the single\-temperature non\-identifiability ofkadsk\_\{\\mathrm\{ads\}\}is an*exact reparameterisation identity*\(kads→αkadsk\_\{\\mathrm\{ads\}\}\\\!\\to\\\!\\alpha k\_\{\\mathrm\{ads\}\},cwall→cwall/αc\_\{\\mathrm\{wall\}\}\\\!\\to\\\!c\_\{\\mathrm\{wall\}\}/\\alpha\), which we prove analytically, and the single\-temperature non\-separation ofν\\nuandEadsE\_\{\\mathrm\{ads\}\}is a textbook property of the Arrhenius form\. The contribution is rather to*confirm and diagnose*these known structures on a genuine multiphysics inverse problem, distinguishing the three distinct mechanisms behind the evidence \(structural degeneracy, initial\-value dependence, and a supervision\-anchored biased minimum\)—a workflow transferable to other hybrid\-model inversions\.
2. 2\.An analytic slope law for the multi\-temperature degeneracy and a mismatch\-diagnostic boundary \(the main novel result\)\.Extending to four temperatures \(96 cases\),ν\\nuandEadsE\_\{\\mathrm\{ads\}\}remain bound along a weakly identifiable one\-dimensional valley \(ground truth on the line;χ12\\chi^\{2\}\_\{1\}95%95\\%width≈0\.6\\approx 0\.6decade; conditioning onν\\nurecoversEadsE\_\{\\mathrm\{ads\}\}to0\.3%0\.3\\%\)\. Beyond confirming the valley, we*derive*its slope from the Arrhenius law,dEads/dlog10ν=kBTeffln10dE\_\{\\mathrm\{ads\}\}/d\\log\_\{10\}\\nu=k\_\{B\}T\_\{\\mathrm\{eff\}\}\\ln 10withTeff=⟨1/T⟩−1T\_\{\\mathrm\{eff\}\}=\\langle 1/T\\rangle^\{\-1\}the harmonic mean of the sampled temperatures, predicting0\.06520\.0652eV/decade in agreement with the observed0\.06470\.0647to0\.7%0\.7\\%\. This law predicts—and the seven\-chemistry, three\-mechanism mismatch matrix below confirms—that the slope is*invariant*under any mismatch preserving a single Arrhenius process and*shifts*only when a second thermally activated process is added \(dual\-site\); an energy\-split calibration sweep then bounds the diagnostic’s detection window\. The slope therefore becomes a usable diagnostic for unmodelled site heterogeneity—the genuinely new increment over the known sloppy\-model phenomenology\.
3. 3\.Enabling results underpinning the analysis\(treated as prerequisites, not primary claims\)\. The identifiability results rest on three supporting elements, established first so that the inversion can be trusted: \(i\) a SALD reaction–transport model whose operating window is*verified saturable*, and a self\-consistent Langmuir–Arrhenius synthetic benchmark generated from known ground truth; \(ii\) the demonstration that PCINN is an*accurate, sample\-efficient*coverage surrogate \(testRlog2=0\.9975±0\.0005R^\{2\}\_\{\\log\}=0\.9975\\pm 0\.0005over 8 seeds from only 30 cases, no appreciable over\-fitting gap; a query costs∼0\.3\\sim 0\.3–77ms,∼5×104\\sim 5\\times 10^\{4\}–1\.1×1061\.1\\times 10^\{6\}times faster than a per\-case CFD solve, amortised over many queries\)—which is what licenses treating its inversion as meaningful; and \(iii\) a delineation of the*action boundary*of the embedded physics: extrapolation gains along the structurally known \(residence\-time\) axis, versus a classical analytic baseline that over\-predicts by several orders of magnitude across the curtain\-dominated window\.
Positioning, emphasised up front\.The data are generated by high\-fidelity simulation from*known*Langmuir–Arrhenius ground truth and inverted with the*same*kinetic form\. The primary goal is therefore*not*to discover unknown kinetic parameters of a real system, but to verify the self\-consistency of the pipeline and to characterise its*identifiability boundary*faithfully: which parameters are recoverable, which are not, and to what precision\. The inverted adsorption energy \(≈0\.77\\approx 0\.77eV\) is consistent in magnitude with DFT values reported for the real TMA/Al2O3system \(0\.50\.5–1\.11\.1eV\), but serves only as a physical\-plausibility reference, not as a calibration\. To verify that the conclusions are not confined to the self\-consistent loop, Section 7 further reports a*seven\-chemistry, three\-mechanism model\-mismatch matrix*\(with an energy\-split calibration sweep\) \(desorption\-side Temkin, adsorption\-side Freundlich, and a site\-heterogeneous dual\-site model\), together with bootstrap resampling and coarse\-grid Latin\-hypercube sampling, showing that the identifiability conclusions persist under mismatch\. Quantitative calibration against and validation with real surface chemistry are left to future work\.
### 1\.4\. Organisation
Section 2 reviews SALD self\-limiting chemistry, ALD modeling, PINN inverse problems, and identifiability\. Section 3 formalises the problem\. Section 4 builds the COMSOL multiphysics model and generates the dataset\. Section 5 details the PCINN architecture, two\-segment trajectory coupling, and the loss\. Section 6 reports single\-temperature surrogate accuracy, baseline comparisons, extrapolation boundaries, and identifiability\. Section 7 extends the pipeline to multiple temperatures and to the model\-mismatch test\. Section 8 concludes\.
## 2\. Background and related work
### 2\.1\. ALD/SALD and self\-limiting surface chemistry
ALD grows films through the self\-limiting surface reactions of alternately exposed precursors\. Within each half\-cycle, precursor molecules adsorb on available surface sites and the reaction terminates once the sites are exhausted, giving atomic\-scale thickness control\. SALD converts the temporal separation into a spatial one, with the substrate moving between precursor zones and inert curtains\. In a continuum description the coverageθ\\thetais commonly described by Langmuir\-type kinetics, with the net adsorption fluxJnet=kadscwall\(1−θ\)−kdesΓsθJ\_\{\\mathrm\{net\}\}=k\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}\(1\-\\theta\)\-k\_\{\\mathrm\{des\}\}\\Gamma\_\{s\}\\theta, wherecwallc\_\{\\mathrm\{wall\}\}is the near\-wall precursor concentration andΓs\\Gamma\_\{s\}the saturation site density; the desorption rate constant follows an Arrhenius lawkdes=νexp\(−Eads/kBT\)k\_\{\\mathrm\{des\}\}=\\nu\\exp\(\-E\_\{\\mathrm\{ads\}\}/k\_\{B\}T\)\. For the representative TMA/H2O process producing Al2O3, first\-principles \(DFT\) calculations place the TMA adsorption energy and the methane\-elimination barrier at about0\.50\.5–1\.11\.1eV, providing a physical reference for the inverted adsorption energy here\.
### 2\.2\. Continuum and analytic modeling of ALD
Models from analytic approximations to full CFD have been developed to predict dose times, conformality and utilisation\. Yanguas\-Gil and Elam derived analytic expressions for coverage, throughput and precursor utilisation under plug\-flow and well\-mixed approximations, showing that the utilisation of cross\-flow, particle coating, and spatial ALD can be unified by a single residence\-time\-related parameter; such models are extremely cheap but rest on a sufficient\-supply assumption\. Finite\-element/finite\-volume reaction–transport CFD, by contrast, resolves the coupled flow, diffusion and surface reaction and accurately captures the concentration and coverage fields, at a high per\-case cost that hinders dense operating\-window scans or iterative inversion\. Such reactor\-scale CFD has been applied specifically to SALD—for example the COMSOL flow/species/surface\-reaction model of Masse de la Huerta et al\. for a close\-proximity SALD head, and the atmospheric\-SALD process\-optimisation studies of Deng et al\. and Pan et al\.—establishing the forward modelling on which the present inverse study builds\. We use high\-fidelity CFD to generate synthetic “measurements” and use the analytic model as a baseline, delineating the validity boundary of each\.
### 2\.3\. Physics\-informed neural networks and inverse problems
Physics\-informed neural networks \(PINN\) embed governing equations as soft constraints in the training loss and can address both forward and inverse problems\. Relative to purely data\-driven surrogates, physical priors typically improve generalisation and sample efficiency under sparse data; relative to purely numerical solvers, the inversion capability makes them suitable for parameter identification—though the credibility of inversion depends on parameter identifiability rather than on fit accuracy itself \(Section 2\.4\)\. The specific form matters: when part of the physics is known and part \(e\.g\. a complex transport closure\) is hard to derive, hard\-coding the known part and learning only the unknown closure is often more robust and interpretable than imposing all equations as soft constraints\. This “known physics\+\+learned closure” strategy has a well\-established lineage: universal differential equations \(UDE\) embed a neural network in the unknown terms of a differential equation and train it jointly with known mechanistic terms; gray\-box/hybrid modeling and neural closure models are likewise widely used to complete mechanistic models with data\-driven closures\. PCINN is a concrete instance of this lineage—with hard\-coded, trainable chemistry and a learned transport closure\. We stress that the architecture itself is not the primary novelty; the core contribution is to use this construction*as a controlled probe of identifiability*: by compressing the data\-driven freedom to a single scalar it enables a systematic, transparent characterisation of the parameter identifiability boundary of a genuine multiphysics inverse problem \(Sections 6 and 7\)\.
### 2\.4\. Parameter identifiability
Whether parameters can be uniquely determined from observations—identifiability—is a prerequisite for inversion credibility\. It is distinguished into structural \(model\-level\) and practical \(given data and noise\) identifiability\. Common diagnostics include the sensitivity\-based Fisher information matrix, whose condition number and eigenstructure reveal the degree to which parameter directions are distinguishable, and the profile likelihood, which fixes a target parameter, re\-optimises the rest, and examines the shape of the loss curve\. In surface\-reaction inversion, when the observable’s sensitivity to certain kinetic parameters degenerates along a particular combination direction, those parameters are not separately identifiable and only a product or ratio can be determined\. We apply these diagnostics to the kinetic parameters of PCINN to characterise the identifiability ofkadsk\_\{\\mathrm\{ads\}\},kdesk\_\{\\mathrm\{des\}\}, andEadsE\_\{\\mathrm\{ads\}\}, avoiding over\-claiming\.
## 3\. Problem formulation
We consider a single precursor species \(S1\) in a 2\-D steady SALD reaction zone\. The state consists of the incompressible laminar flow field, the gas\-phase precursor concentrationcc, and the substrate surface coverageθ\\theta\. The flow satisfies the steady incompressible Navier–Stokes equations with prescribed inlet velocities and a moving substrate wall \(velocityvsubv\_\{\\mathrm\{sub\}\}\)\. The dilute species transport obeys
∇⋅\(−D∇c\+𝐮c\)=0,\\nabla\\\!\\cdot\\\!\(\-D\\nabla c\+\\mathbf\{u\}\\,c\)=0,\(1\)withc=C0c=C\_\{0\}at the A inlet andc=0c=0elsewhere; at the substrate the gas\-phase flux equals the net surface reaction rate,−𝐧⋅\(−D∇c\)=Jnet\-\\mathbf\{n\}\\\!\\cdot\\\!\(\-D\\nabla c\)=J\_\{\\mathrm\{net\}\}\. The substrate motion advects adsorbed species, soθ\(x\)\\theta\(x\)obeys a spatial advection equation,
Γsvsub∂θ∂x=Jnet=kadscwall\(1−θ\)−kdesΓsθ,kdes=νe−Eads/kBT,\\Gamma\_\{s\}\\,v\_\{\\mathrm\{sub\}\}\\,\\frac\{\\partial\\theta\}\{\\partial x\}=J\_\{\\mathrm\{net\}\}=k\_\{\\mathrm\{ads\}\}\\,c\_\{\\mathrm\{wall\}\}\(1\-\\theta\)\-k\_\{\\mathrm\{des\}\}\\,\\Gamma\_\{s\}\\,\\theta,\\hskip 18\.49988ptk\_\{\\mathrm\{des\}\}=\\nu\\,e^\{\-E\_\{\\mathrm\{ads\}\}/k\_\{B\}T\},\(2\)withθ=0\\theta=0at the substrate inlet edge\. Temperature enters only throughkdesk\_\{\\mathrm\{des\}\}\(all ofkadsk\_\{\\mathrm\{ads\}\},DD,μ\\mu,ρ\\rho,Γs\\Gamma\_\{s\},C0C\_\{0\}and the geometry areTT\-independent\)\. The primary observable \(label\) is the substrate\-averaged A coverageθ¯A\\bar\{\\theta\}\_\{A\}\. The inverse problem is to recover the kinetic parameters in Eq\. \([2](https://arxiv.org/html/2608.00212#S3.E2)\) from coverage observations across operating conditions, and to determine which of\{kads,kdes,Eads,ν\}\\\{k\_\{\\mathrm\{ads\}\},k\_\{\\mathrm\{des\}\},E\_\{\\mathrm\{ads\}\},\\nu\\\}are identifiable\.
The strong\-adsorption ratiokads/\(kdesΓs\)≈2000k\_\{\\mathrm\{ads\}\}/\(k\_\{\\mathrm\{des\}\}\\Gamma\_\{s\}\)\\approx 2000at 300 K ensures that coverage can saturate even when the near\-wall concentration is strongly depleted; the gap height and inlet velocities are chosen so that a sweep overvsubv\_\{\\mathrm\{sub\}\}crosses the saturated–undersaturated transition, providing the dynamic range required for kinetic inversion\. We note that the surface Damköhler numberDah=kadsH/D≈0\.33Da\_\{h\}=k\_\{\\mathrm\{ads\}\}H/D\\approx 0\.33and the supply numberRsup=UAH2/\(DwA\)≈0\.67R\_\{\\mathrm\{sup\}\}=U\_\{A\}H^\{2\}/\(D\\,w\_\{A\}\)\\approx 0\.67are used here as*design*criteria to exclude extreme transport\-limited regimes, not as a measure of saturability; the actual saturability is verified directly through the coverage dynamic range \(Section 4\)\.
## 4\. COMSOL multiphysics model and dataset generation
### 4\.1\. Geometry and computational domain
The domain is a 2\-D rectangle of lengthL=30L=30mm and gap heightH=0\.2H=0\.2mm, representing the narrow slit between the showerhead and the substrate\. The top edge carries, from left to right, the precursor\-A inlet, the inert curtain inlet, and the precursor\-B inlet, separated by solid wall segments; the bottom edge is a substrate moving atvsubv\_\{\\mathrm\{sub\}\}; the two ends are outlets\. The A exposure zone has widthwA=5w\_\{A\}=5mm with its trailing edge atxA,end=7\.5x\_\{A,\\mathrm\{end\}\}=7\.5mm\. The gap height0\.20\.2mm \(rather than0\.50\.5mm\) improves cross\-gap diffusive supply and is closer to realistic showerhead dimensions\.
### 4\.2\. Governing equations and interfaces
Laminar flow \(Taylor–Hood P2\+\+P1 to satisfy LBB stability\) is solved together with dilute species transport and a coefficient\-form boundary PDE for the surface coverage, as formalised in Section 3\. Temperature enters only through the Arrheniuskdesk\_\{\\mathrm\{des\}\}, and the measured wall shear rates are identical across temperatures, confirming that the flow is decoupled from temperature and that temperature information enters the observations solely through the surface kinetics\.
### 4\.3\. Model parameters
Key parameters:L=30L=30mm,H=0\.2H=0\.2mm,wA=5w\_\{A\}=5mm; precursor inlet velocityUA=0\.5U\_\{A\}=0\.5m/s; curtain velocity swept over0\.50\.5–1616m/s; substrate velocity swept over0\.050\.05–2\.02\.0m/s;C0=0\.2C\_\{0\}=0\.2mol/m3;D=6×10−6D=6\\times 10^\{\-6\}m2/s;Γs=5×10−6\\Gamma\_\{s\}=5\\times 10^\{\-6\}mol/m2; ground\-truthkads=0\.01k\_\{\\mathrm\{ads\}\}=0\.01m/s,kdes=1\.0k\_\{\\mathrm\{des\}\}=1\.0s\-1at 300 K viaν=1013\\nu=10^\{13\}s\-1andEads=0\.774E\_\{\\mathrm\{ads\}\}=0\.774eV;kB=8\.617×10−5k\_\{B\}=8\.617\\times 10^\{\-5\}eV/K\. The strong\-adsorption ratiokads/\(kdesΓs\)≈2000k\_\{\\mathrm\{ads\}\}/\(k\_\{\\mathrm\{des\}\}\\Gamma\_\{s\}\)\\approx 2000guarantees saturability despite near\-wall depletion\.
### 4\.4\. Feasibility trial
Before the 48\-case scan, a single trial at the most favourable condition \(slowest substrate, weakest curtain:vsub=0\.05v\_\{\\mathrm\{sub\}\}=0\.05,Ucurtain=0\.5U\_\{\\mathrm\{curtain\}\}=0\.5\) verifies that the working region is saturable with adequate dynamic range\. The trial gives a peak coverageθpeak=0\.97\\theta\_\{\\mathrm\{peak\}\}=0\.97and an exit coverageθexit=0\.91\\theta\_\{\\mathrm\{exit\}\}=0\.91, confirming the saturated end\. The saturability is thus established by the directly measured coverage dynamic range, not inferred fromDahDa\_\{h\}orRsupR\_\{\\mathrm\{sup\}\}\.
### 4\.5\. Mesh and grid independence
A boundary layer is placed on the substrate \(first layer0\.5μ0\.5~\\mum, 10 layers, stretch 1\.2\); the bulk maximum element size is40μ40~\\mum, all scaled by a uniform refinement factormfm\_\{f\}\. The reference case is solved manually at successivemfm\_\{f\}\. Table[1](https://arxiv.org/html/2608.00212#S4.T1)lists three refinement levels\. The fitted integral labelθ¯A\\bar\{\\theta\}\_\{A\}is converged to±0\.1%\\pm 0\.1\\%\(it varies only within0\.77420\.7742–0\.77590\.7759acrossmf=2,4,8m\_\{f\}=2,4,8\), whereas the near\-wall concentration⟨c⟩/C0\\langle c\\rangle/C\_\{0\}does*not*converge—even as a surface average it swings non\-monotonically \(0\.0148→0\.0123→0\.02060\.0148\\\!\\to\\\!0\.0123\\\!\\to\\\!0\.0206, a67%67\\%jump frommf=4m\_\{f\}=4to88\), because it samples the extremely steep leading\-edge gradient in the strong\-adsorption regime\. The production mesh ismf=2m\_\{f\}=2\(≈3\.5×104\\approx 3\.5\\times 10^\{4\}elements\)\.
Table 1:Mesh convergence at the reference case \(vsub=0\.05v\_\{\\mathrm\{sub\}\}=0\.05,Ucurtain=0\.5U\_\{\\mathrm\{curtain\}\}=0\.5,T=300T=300K\)\. The integral labelθ¯A\\bar\{\\theta\}\_\{A\}converges; the near\-wall concentration⟨c⟩/C0\\langle c\\rangle/C\_\{0\}does not\.This near\-wall pathology does*not*, however, contaminate the kinetic inversion\. The adsorption energyEadsE\_\{\\mathrm\{ads\}\}is recovered from the*temperature dependence*of the converged integralθ¯A\\bar\{\\theta\}\_\{A\}, not from the amplitude of⟨c⟩/C0\\langle c\\rangle/C\_\{0\}; and the productkadscwallk\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}is in any case structurally non\-identifiable \(Section 6\.5\), so the mesh sensitivity ofcwallc\_\{\\mathrm\{wall\}\}’s scale only translates the solution along that already\-degenerate direction without moving the identifiableEadsE\_\{\\mathrm\{ads\}\}or the degeneracy slope\. Two facts make this concrete rather than merely plausible\. First, because temperature enters only throughkdesk\_\{\\mathrm\{des\}\}and the flow is decoupled from temperature \(the measured wall shear rates are identical across all temperatures, Section 4\.2\), any residual mesh bias inθ¯A\\bar\{\\theta\}\_\{A\}is*common\-mode across temperatures*and therefore largely cancels in the Arrhenius slopedlnkdes/d\(1/T\)d\\ln k\_\{\\mathrm\{des\}\}/d\(1/T\)from whichEadsE\_\{\\mathrm\{ads\}\}is read; the±0\.1%\\pm 0\.1\\%spread ofθ¯A\\bar\{\\theta\}\_\{A\}acrossmf=2,4,8m\_\{f\}=2,4,8\(Table[1](https://arxiv.org/html/2608.00212#S4.T1)\) thus propagates to a negligible shift inEadsE\_\{\\mathrm\{ads\}\}\. Second, this is confirmed*empirically*by the independent coarse\-grid experiment of Section 7\.4: regenerating the dataset on a much coarser mesh \(maximum element80μ80~\\mum, a\>2×\>2\\timescoarsening that exercises exactly the near\-wall pathology reported here\) shifts the empirical degeneracy slope only to0\.0630\.063\(versus0\.06470\.0647on the fine grid\) and the conditionalEadsE\_\{\\mathrm\{ads\}\}to0\.7770\.777eV \(versus0\.7760\.776\)—a∼1\\sim 1meV change—so theEadsE\_\{\\mathrm\{ads\}\}profile\-likelihood minimum is mesh\-stable to well within its own uncertainty\. The single\-point wall concentration does not converge under refinement \(extremely steep near\-wall gradient in the strong\-adsorption regime\), so near\-wall concentration observations use the stable surface/line average⟨c⟩/C0\\langle c\\rangle/C\_\{0\}only as an auxiliary supervision label anchoring the*magnitude*ofCs∗C\_\{s\}^\{\*\}\(validated quantitatively in Sections 6\.1 and 6\.6\); the integral quantitiesθ¯A\\bar\{\\theta\}\_\{A\}andJnetJ\_\{\\mathrm\{net\}\}converge well\.
### 4\.6\. Case scan and solving
The main scan takesvsub∈\{0\.05,0\.1,0\.2,0\.3,0\.5,0\.8,1\.2,2\.0\}v\_\{\\mathrm\{sub\}\}\\in\\\{0\.05,0\.1,0\.2,0\.3,0\.5,0\.8,1\.2,2\.0\\\}m/s \(8 values\) andUcurtain∈\{0\.5,1,2,4,8,16\}U\_\{\\mathrm\{curtain\}\}\\in\\\{0\.5,1,2,4,8,16\\\}m/s \(6 values\), a full cross of 48 cases\. The scan ranges are chosen by physical purpose rather than arbitrarily: the lowervsubv\_\{\\mathrm\{sub\}\}bound ensures a saturable end, the upper bound covers high\-throughput roll\-to\-roll/rotary operation, and the curtain range covers from “almost no isolation” to “strong curtain strongly suppressing near\-wall concentration”; together they guarantee the full saturated–transition–undersaturated dynamic range \(about four orders of magnitude\) required for kinetic invertibility\. With the segregated\-step damping lowered from0\.70\.7to0\.50\.5and the maximum iterations raised from300300to500500, all 48 cases converge\. Each case exports the substrate\-averaged coverageθ¯A\\bar\{\\theta\}\_\{A\}\(primary label\), the exit/peak coverage \(diagnostics\), the surface\-averaged near\-wall concentrationCs,AmidC\_\{s,A\}^\{mid\}\(auxiliary supervision\), and conservation diagnostics\.
### 4\.7\. Dataset and split
The dataset has 48 rows, each recording\(vsub,Ucurtain\)\(v\_\{\\mathrm\{sub\}\},U\_\{\\mathrm\{curtain\}\}\), their normalisationsv∗=vsub/2\.0v^\{\*\}=v\_\{\\mathrm\{sub\}\}/2\.0andU∗=Ucurtain/16U^\{\*\}=U\_\{\\mathrm\{curtain\}\}/16, and the derived scalars\. Here the divisors2\.02\.0and1616are simply the upper bounds of the respective parameter scans \(map\-to\-\[0,1\]\[0,1\]feature scaling\); they are not physical constants\. It is stratified into3030train /99validation /99test, with near\-saturated cases placed mainly in the training set so that the desorption information can be learned\. Training and validation labels carryσ=5%\\sigma=5\\%relative noise; the test set is clean\. The substrate\-averaged coverageθ¯A\\bar\{\\theta\}\_\{A\}spans about four orders of magnitude \(from the saturated end0\.780\.78to the strong\-curtain/fast\-substrate end∼10−4\\sim 10^\{\-4\}\), with the curtain velocity as the dominant control, forming a clear operating window in the\(vsub,Ucurtain\)\(v\_\{\\mathrm\{sub\}\},U\_\{\\mathrm\{curtain\}\}\)plane\.
## 5\. Physics–chemistry\-informed neural network \(PCINN\)
### 5\.1\. Overall architecture
PCINN couples two branches in series\. The physics branch \(a neural network\) approximates the operating\-condition→\\,\\to\\,effective near\-wall concentration closure; the chemistry branch \(a hard\-coded Langmuir model with trainable kinetics\) integrates coverage along the substrate trajectory\. The data\-driven freedom is compressed to a single scalarCs∗C\_\{s\}^\{\*\}, with everything else carried by known physics:
\(v∗,U∗\)→physics branchCs∗→cwall=Cs∗C0chemistry branch \(trajectory integration\)→θ¯A\.\(v^\{\*\},U^\{\*\}\)\\;\\xrightarrow\{\\text\{physics branch\}\}\\;C\_\{s\}^\{\*\}\\;\\xrightarrow\{c\_\{\\mathrm\{wall\}\}=C\_\{s\}^\{\*\}C\_\{0\}\}\\;\\text\{chemistry branch \(trajectory integration\)\}\\;\\to\\;\\bar\{\\theta\}\_\{A\}\.\(3\)
Figure 1:PCINN architecture\. A physics branch \(small MLP\) maps the operating condition\(v∗,U∗\[,T∗\]\)\(v^\{\*\},U^\{\*\}\[,T^\{\*\}\]\)to an effective near\-wall concentrationCs∗C\_\{s\}^\{\*\}; a hard\-coded chemistry layer with trainable kinetics integrates coverage along the substrate trajectory in two segments \(A\-zone exposure and downstream desorption\) to yieldθ¯A\\bar\{\\theta\}\_\{A\}\. The data\-driven freedom is compressed to a single scalar\.
### 5\.2\. Physics branch
A fully connected network with inputs\(v∗,U∗\)\(v^\{\*\},U^\{\*\}\)\(plusT∗T^\{\*\}for the multi\-temperature case\), four hidden layers of 64 units, SiLU activation, and a softplus output ensuringCs∗≥0C\_\{s\}^\{\*\}\\geq 0\. The output is not capped atC0C\_\{0\}, allowing the network to learn an effective driving concentration; the learned effectivecwallc\_\{\\mathrm\{wall\}\}is about5\.5×5\.5\\timesthe surface\-averagedCs,AmidC\_\{s,A\}^\{mid\}, which is physically reasonable \(Section 6\.1\)\.
### 5\.3\. Chemistry branch and two\-segment trajectory coupling
The chemistry branch reuses the same spatial advection equation, with the near\-wall concentration replaced by the physics\-branch output, integrated forward along the substrate\. Kinetic parameters are parameterised inlog10\\log\_\{10\}space\. To make the output consistent with the primary labelθ¯A\\bar\{\\theta\}\_\{A\}\(whole\-substrate average\), the integration is split into two segments: an A\-zone exposure segment \(lengthLA=wA=5L\_\{A\}=w\_\{A\}=5mm,cwall=Cs∗C0c\_\{\\mathrm\{wall\}\}=C\_\{s\}^\{\*\}C\_\{0\}\) and a downstream desorption segment \(lengthLdown=22\.5L\_\{\\mathrm\{down\}\}=22\.5mm,cwall=0c\_\{\\mathrm\{wall\}\}=0\), with the substrate average
θ¯A=1L\(∫Aθ𝑑x\+∫downθ𝑑x\),\\bar\{\\theta\}\_\{A\}=\\frac\{1\}\{L\}\\Big\(\\int\_\{A\}\\theta\\,dx\+\\int\_\{\\mathrm\{down\}\}\\theta\\,dx\\Big\),\(4\)discretised by forward Euler \(NA=120N\_\{A\}=120,Ndown=200N\_\{\\mathrm\{down\}\}=200steps\), fully differentiable\.
The trajectory integrator is an ordinary differential equation in the coverageθ\\theta\(not an advection PDE\), so it is free of numerical diffusion; within each segment the kinetics are*linear*inθ\\theta\(dθ/dx=A−\(A\+B\)θd\\theta/dx=A\-\(A\+B\)\\thetain the A zone withA=kadscwall/\(Γsvsub\)A=k\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}/\(\\Gamma\_\{s\}v\_\{\\mathrm\{sub\}\}\),B=kdes/vsubB=k\_\{\\mathrm\{des\}\}/v\_\{\\mathrm\{sub\}\}, anddθ/dx=−Bθd\\theta/dx=\-B\\thetadownstream\), which admits an exact solution against which the forward\-Euler discretisation is checked\. The explicit scheme is unconditionally usable in this regime: at the stiffest operating point the dimensionless step parameter\(A\+B\)dx≈0\.04≪2\(A\+B\)\\,dx\\approx 0\.04\\ll 2, far from the stability limit\. Step\-independence is verified against this analytic reference \(Table[2](https://arxiv.org/html/2608.00212#S5.T2)\): at the production resolution the relative error inθ¯A\\bar\{\\theta\}\_\{A\}is0\.17%0\.17\\%\(median over the operating window\) and at most1\.4%1\.4\\%\(at the extreme low\-coverage cornervsub=0\.05v\_\{\\mathrm\{sub\}\}=0\.05, highTT,θ∼5×10−4\\theta\\sim 5\\times 10^\{\-4\}\), both well below theσ=5%\\sigma=5\\%label noise; the error halves under each doubling of the step count \(clean first\-order convergence\), confirming thatNA=120N\_\{A\}=120,Ndown=200N\_\{\\mathrm\{down\}\}=200lies on the converged plateau\.
Table 2:Forward\-Euler step independence\. Relative error inθ¯A\\bar\{\\theta\}\_\{A\}against the exact per\-segment linear\-ODE solution, at the worst\-case \(extreme low\-coverage\) operating point\. The production resolution \(NA=120N\_\{A\}=120,Ndown=200N\_\{\\mathrm\{down\}\}=200\) already sits well below theσ=5%\\sigma=5\\%label noise; the error halves under each step doubling\.A known model\-form bias follows from approximating the entire A\-zone exposure by a single effectivecwallc\_\{\\mathrm\{wall\}\}and the downstream by pure desorption: at the saturated end the high\-coverage cases are systematically under\-estimated by∼13%\\sim 13\\%–36%36\\%\. Physically, this is the expressivity limit of the single\-scalar near\-wall\-concentration closure in the strong\-depletion / near\-saturation regime, where one effectivecwallc\_\{\\mathrm\{wall\}\}cannot reproduce the full streamwise concentration profile\. This bias is compressed in log space \(Rlog2R^\{2\}\_\{\\log\}remains∼0\.998\\sim 0\.998\) and is the main cause of theRraw2≈0\.989R^\{2\}\_\{\\mathrm\{raw\}\}\\approx 0\.989deduction; its contribution to the multi\-temperatureEadsE\_\{\\mathrm\{ads\}\}inversion is later quantified at≈0\.016\\approx 0\.016eV \(Section 7\), so it does*not*threaten the identifiability conclusions\. A bias\-versus\-coverage diagnostic \(Fig\.[2](https://arxiv.org/html/2608.00212#S5.F2)\) makes this explicit: the raw relative bias reaches tens of percent only at the highest coverage, yet the log\-space residual stays within theσ=5%\\sigma=5\\%label\-noise band across all four decades of coverage \(RMS0\.090\.09\)\.
Figure 2:Model\-form bias versus coverage\. \(a\) Raw relative prediction bias\(θ^−θ\)/θ\(\\hat\{\\theta\}\-\\theta\)/\\theta: a systematic under\-prediction of1313–36%36\\%appears only at the saturated \(high\-coverage\) end, the expressivity limit of the single\-scalar near\-wall closure\. \(b\) The same residuals in log spacelog10\(θ^/θ\)\\log\_\{10\}\(\\hat\{\\theta\}/\\theta\)stay within theσ=5%\\sigma=5\\%label\-noise band across four decades of coverage \(RMS0\.090\.09\); the bias is therefore real but log\-compressed, and perturbs the log\-space kinetic inversion by only≈0\.016\\approx 0\.016eV \(EadsE\_\{\\mathrm\{ads\}\}\), leaving the identifiability conclusions intact\.
### 5\.4\. Loss function
The loss combines the data term, the near\-wall concentration supervision, a monotonicity prior, and a weak kinetic prior:
ℒ=ℒdata\+wCℒCs\+wmℒmono\+wpℒprior,wC=0\.1,wm=0\.1,wp=0\.01,\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{data\}\}\+w\_\{C\}\\mathcal\{L\}\_\{C\_\{s\}\}\+w\_\{m\}\\mathcal\{L\}\_\{\\mathrm\{mono\}\}\+w\_\{p\}\\mathcal\{L\}\_\{\\mathrm\{prior\}\},\\hskip 18\.49988ptw\_\{C\}=0\.1,\\ w\_\{m\}=0\.1,\\ w\_\{p\}=0\.01,\(5\)with all terms inlog10\\log\_\{10\}space\. The data term in log space is essential becauseθ\\thetaspans four orders of magnitude \(a raw\-value MSE is dominated by large values and collapses on the low\-coverage cases,Rlog2≈0\.68R^\{2\}\_\{\\log\}\\approx 0\.68, versus∼0\.998\\sim 0\.998in log space\)\. The concentration supervisionℒCs\\mathcal\{L\}\_\{C\_\{s\}\}uses the surface/line average; the monotonicity prior penalises∂Cs∗/∂U∗\>0\\partial C\_\{s\}^\{\*\}/\\partial U^\{\*\}\>0\(a stronger curtain lowers near\-wall concentration\); the weak kinetic prior\(log10kdes−log10kdesprior\)2\(\\log\_\{10\}k\_\{\\mathrm\{des\}\}\-\\log\_\{10\}k\_\{\\mathrm\{des\}\}^\{\\mathrm\{prior\}\}\)^\{2\}is used at a single temperature and removed for the multi\-temperature inversion \(where the Arrhenius relation across temperatures provides the constraint\)\. The Adam optimiser \(10−310^\{\-3\}\) with ReduceLROnPlateau is used, with early stopping on the validation log data\-loss; headline results are reported as mean±\\pms\.d\. over 8 random seeds, each perturbing both the initialisation and the label\-noise realisation\.
## 6\. Results and discussion
### 6\.1\. Surrogate accuracy
Using only 30 training cases, PCINN reproducesθ¯A\\bar\{\\theta\}\_\{A\}across about four orders of magnitude\. Over 8 random seeds \(each perturbing both initialisation and label noise\), the testRlog2=0\.9975±0\.0005R^\{2\}\_\{\\log\}=0\.9975\\pm 0\.0005andRraw2=0\.989±0\.003R^\{2\}\_\{\\mathrm\{raw\}\}=0\.989\\pm 0\.003; the train/val/testRlog2R^\{2\}\_\{\\log\}are0\.991/0\.996/0\.99750\.991/0\.996/0\.9975\(cross\-seed s\.d\.≤0\.001\\leq 0\.001\)\. The testRlog2R^\{2\}\_\{\\log\}is slightly above the training value on every seed—not because there is no over\-fitting, but because the worst\-fit high\-coverage cases happen to fall mostly in the training set; the robust statement is “no appreciable over\-fitting gap”\.
The effective near\-wall concentration learned by the physics branch agrees in magnitude with the COMSOL surface\-averagedCs,AmidC\_\{s,A\}^\{mid\}but is systematically about5\.5×5\.5\\timeshigher \(range3\.33\.3–7\.9×7\.9\\times\)\. This is*not*a supervision failure but reflects the network learning the physically correct effective driving concentration, which can be verified quantitatively from the COMSOL field\. At the saturated case \(vsub=0\.05v\_\{\\mathrm\{sub\}\}=0\.05,Ucurtain=0\.5U\_\{\\mathrm\{curtain\}\}=0\.5,T=300T=300K\), the near\-wallc/C0c/C\_\{0\}along the A zone decays sharply from the leading edge:0\.1240\.124atx=2\.5x=2\.5mm,0\.0590\.059at the quarter point,0\.0150\.015at the midpoint,0\.00110\.0011at the three\-quarter point,0\.00030\.0003at the trailing edge—about400×400\\timesfrom leading edge to end \(the precursor is rapidly consumed on entering the A zone\)\. The leading\-edge concentration0\.1240\.124nearly coincides with the learnedCs∗≈0\.021×5\.5≈0\.115C\_\{s\}^\{\*\}\\approx 0\.021\\times 5\.5\\approx 0\.115, whereas the surface averageCs,Amid≈0\.021C\_\{s,A\}^\{mid\}\\approx 0\.021mixes the high leading\-edge value with the near\-zero trailing values and severely under\-estimates the effective concentration\. The auxiliary supervisionℒCs\\mathcal\{L\}\_\{C\_\{s\}\}therefore*anchorsCs∗C\_\{s\}^\{\*\}at the correct magnitude*\(preventingCs∗C\_\{s\}^\{\*\}from drifting freely under thekads⋅cwallk\_\{\\mathrm\{ads\}\}\\\!\\cdot\\\!c\_\{\\mathrm\{wall\}\}degeneracy\), while within that constraint the network learns the effective leading\-edge concentration; the two are not in conflict\. The spread of the ratio \(3\.33\.3–7\.9×7\.9\\times\) is itself systematic rather than scatter: it tracks how sharply the near\-wall concentration is depleted along the A zone, set by the balance between adsorption consumption and convective/diffusive resupply\. Strongly consuming conditions \(lowvsubv\_\{\\mathrm\{sub\}\}or near\-saturation\) produce a steeply peaked profile and a large leading\-edge\-to\-average ratio, whereas weakly consuming conditions \(highvsubv\_\{\\mathrm\{sub\}\}, a strong curtain, or highTTwith fast desorption\) flatten the profile and push the ratio toward unity\. Because the auxiliary loss anchors only the order of magnitude while the network learns this condition\-dependent effective concentration within that constraint, the spread does not degrade the anchoring: the surrogate accuracy and, in particular, theEadsE\_\{\\mathrm\{ads\}\}inversion \(read from the temperature slope, not the absolutecwallc\_\{\\mathrm\{wall\}\}scale\) are unaffected\. This also explains the∼7×\\sim 7\\timesoff\-truthkadsk\_\{\\mathrm\{ads\}\}minimum in the profile likelihood of Section 6\.5: supervision anchors only the magnitude, not the absolute value, so the residual concentration\-scale freedom can still be absorbed bykadsk\_\{\\mathrm\{ads\}\}\.
Figure 3:Surrogate accuracy \(representative run\)\. Predicted versus trueθ¯A\\bar\{\\theta\}\_\{A\}across about four orders of magnitude on the clean test set; PCINN reaches testRlog2=0\.998R^\{2\}\_\{\\log\}=0\.998using only 30 training cases\.
### 6\.2\. Sample efficiency: physics versus a data\-only baseline
The comparisons in this and the following two subsections are not an end in themselves: their purpose is to establish that the embedded physics buys*invertibility and a trustworthy identifiability analysis*, not merely surrogate accuracy—so the value of the physics constraint is relocated from “fitting accuracy” to “invertibility\+\+extrapolation\+\+identifiability\-boundary characterisation”, which the rest of the paper develops\. To quantify the gain from the physics constraint, we compare PCINN with a structure\-free multilayer perceptron \(MLP\)\. Both comparisons \(sample efficiency and extrapolation\) are performed in the*primary caliber*\(θ¯A\\bar\{\\theta\}\_\{A\}, two\-segment\), consistent with the main results; the MLP directly regresseslog10θ¯A\\log\_\{10\}\\bar\{\\theta\}\_\{A\}from\(v∗,U∗\)\(v^\{\*\},U^\{\*\}\)with comparable capacity and identical loss, optimiser and early\-stopping\. At the sparsestn=6n=6, PCINN attainsRlog2=0\.915±0\.005R^\{2\}\_\{\\log\}=0\.915\\pm 0\.005with high seed consistency, whereas the MLP gives only0\.78±0\.130\.78\\pm 0\.13with an order of magnitude larger variance \(0\.650\.65–0\.920\.92\); PCINN saturates at0\.997±0\.0010\.997\\pm 0\.001byn≈15n\\approx 15, while the MLP catches up to0\.9900\.990only atn=30n=30\.
To calibrate the magnitude of the “physics gain”, we add a Gaussian\-process \(GP\) regression baseline \(anisotropic Matérn kernel,\(v∗,U∗\)→log10θ¯A\(v^\{\*\},U^\{\*\}\)\\to\\log\_\{10\}\\bar\{\\theta\}\_\{A\}\), a comparison deliberately favourable to the baseline since a 2\-D dense sampling is the GP’s home ground \(Table[3](https://arxiv.org/html/2608.00212#S6.T3)\)\. At full data the GP reachesRlog2=0\.983R^\{2\}\_\{\\log\}=0\.983\(Rraw2=0\.936R^\{2\}\_\{\\mathrm\{raw\}\}=0\.936\), below PCINN’s0\.9980\.998but already high, confirming that a pure interpolator suffices for the surrogate task under dense sampling\. In sample efficiency, however, the gap widens sharply at the sparse end: atn=6n=6the GP gives only0\.58±0\.420\.58\\pm 0\.42\(nearly unusable\) versus PCINN’s0\.92±0\.010\.92\\pm 0\.01, and stabilises above0\.960\.96only byn≈18n\\approx 18\. This clarifies the true positioning of the PCINN gain: when data are abundant the GP interpolation accuracy approaches PCINN, so the value of the physics constraint is*not*mainly in interpolation accuracy; its irreplaceable advantages are \(i\) stable high accuracy under data scarcity, \(ii\) invertible physical parameters \(EadsE\_\{\\mathrm\{ads\}\},kdesk\_\{\\mathrm\{des\}\}, which a GP cannot provide\), and \(iii\) extrapolation along the residence\-time axis and identifiability analysis \(Sections 6\.4, 6\.5 and 7\)\. The comparison with the GP thus does not weaken the contribution but relocates it from “surrogate accuracy” to “invertibility\+\+extrapolation\+\+identifiability\-boundary characterisation”\.
Table 3:GP baseline versus PCINN \(primary caliberθ¯A\\bar\{\\theta\}\_\{A\}\)\.Figure 4:Sample efficiency \(primary caliberθ¯A\\bar\{\\theta\}\_\{A\}, two\-segment\)\. PCINN attainsRlog2≈0\.92R^\{2\}\_\{\\log\}\\approx 0\.92already atn=6n=6training cases with small seed variance, whereas a structure\-free MLP needs∼30\\sim 30cases to catch up and has an order\-of\-magnitude larger variance at the sparse end\.
### 6\.3\. Comparison with an analytic baseline
We compare against the Yanguas\-Gil–Elam plug\-flow self\-limiting model \(which assumes sufficient supplyc=C0c=C\_\{0\}and integrates along residence time, depending only onvsubv\_\{\\mathrm\{sub\}\}\)\. Of the 48 cases, only 5 \(slow\-substrate/low\-curtain saturated corner\) fall within2×2\\timesof the analytic value, with a median over\-prediction of1\.551\.55decades and up to3\.943\.94decades at high curtain—clearly delineating the “analytic\-fails, PCINN/resolved\-transport\-required” regime\. \(This comparison usesθexit\\theta\_\{\\mathrm\{exit\}\}since the analytic expression only yields an exit coverage, as noted\.\) We stress that this is*not*a fair accuracy contest and we do not use it to inflate PCINN: the plug\-flow model fails here*by construction*, because its sufficient\-supply assumption \(c=C0c=C\_\{0\}\) is exactly what the gas curtain violates, so it is structurally incapable of representing the curtain\-dominated bulk of the window\. The comparison should therefore be read as an*operating\-domain partition*—identifying where a cheap analytic estimate remains usable \(the supply\-rich saturated corner\) versus where resolved transport is mandatory—rather than as evidence that PCINN is ”more accurate” than a model built for a different regime\.
Figure 5:Analytic baseline \(Yanguas\-Gil–Elam plug\-flow self\-limiting model\)\. Only the slow\-substrate/low\-curtain saturated corner falls within2×2\\timesof the analytic value; the median over\-prediction is1\.551\.55decades \(up to3\.943\.94at high curtain\), delineating the regime where resolved transport \(PCINN\) is required\.
### 6\.4\. Extrapolation: the action boundary of the physics constraint
Extrapolating along the substrate velocityvsubv\_\{\\mathrm\{sub\}\}\(train onvsub≤0\.5v\_\{\\mathrm\{sub\}\}\\leq 0\.5, predictvsub≥0\.8v\_\{\\mathrm\{sub\}\}\\geq 0\.8\), PCINN stays robustly positive across all88seeds \(Rlog2=0\.90±0\.13R^\{2\}\_\{\\log\}=0\.90\\pm 0\.13, range0\.610\.61–0\.9980\.998\), whereas the MLP is unstable \(−0\.16±1\.01\-0\.16\\pm 1\.01, range−2\.01\-2\.01to0\.950\.95, an order\-of\-magnitude larger spread and frequently negative\)\. The reason is that the coverage dependence onvsubv\_\{\\mathrm\{sub\}\}enters through the residence timeτ=wA/vsub\\tau=w\_\{A\}/v\_\{\\mathrm\{sub\}\}, which is*known physics*in the chemistry\-branch integration, while the effective near\-wall concentration grows only mildly withvsubv\_\{\\mathrm\{sub\}\}\. Conversely, extrapolating along the curtain velocityUcurtainU\_\{\\mathrm\{curtain\}\}\(train onUcurtain≤2U\_\{\\mathrm\{curtain\}\}\\leq 2, predictUcurtain≥4U\_\{\\mathrm\{curtain\}\}\\geq 4\), both models fail and are statistically indistinguishable \(88\-seed mean±\\pms\.d\. PCINN−7\.9±5\.1\-7\.9\\pm 5\.1, MLP−7\.8±3\.3\-7\.8\\pm 3\.3\), becauseUcurtainU\_\{\\mathrm\{curtain\}\}drives precisely the operating\-condition→\\,\\to\\,Cs∗C\_\{s\}^\{\*\}mapping that is learned by the network and unconstrained by the chemistry layer\. Thus the physics constraint provides extrapolation gains*only*along the structurally known \(residence\-time\) axis, and confers no advantage along the data\-driven transport axis—a deliberately honest negative control\.
Figure 6:Extrapolation to held\-out operating regions \(88seeds, mean±\\pms\.d\.; dots are individual seeds\)\. Along the substrate velocityvsubv\_\{\\mathrm\{sub\}\}\(residence time, carried by the chemistry branch\) PCINN stays robustly positive \(Rlog2=0\.90±0\.13R^\{2\}\_\{\\log\}=0\.90\\pm 0\.13, all seeds\>0\>0\) while the MLP is unstable; along the curtain velocityUcurtainU\_\{\\mathrm\{curtain\}\}\(the data\-driven transport axis\) both fail and are indistinguishable\. The physics constraint helps only along the structurally known axis\. \(Rlog2<0R^\{2\}\_\{\\log\}<0means the prediction is worse than a constant\-mean baseline, i\.e\. a failed extrapolation\.\)
### 6\.5\. Kinetic inversion and identifiability
Identifiable\.EadsE\_\{\\mathrm\{ads\}\}andkdesk\_\{\\mathrm\{des\}\}are robustly identifiable: over 8 seeds,Eads=0\.773±0\.001E\_\{\\mathrm\{ads\}\}=0\.773\\pm 0\.001eV \(truth0\.7740\.774\) andkdes=1\.04±0\.03k\_\{\\mathrm\{des\}\}=1\.04\\pm 0\.03s\-1\(truth1\.01\.0\)\. Note thatEadsE\_\{\\mathrm\{ads\}\}andkdesk\_\{\\mathrm\{des\}\}are two expressions of the same identifiable quantity \(Eads=kBTln\(ν/kdes\)E\_\{\\mathrm\{ads\}\}=k\_\{B\}T\\ln\(\\nu/k\_\{\\mathrm\{des\}\}\)at fixedν,T\\nu,T\); the narrower\-lookingEadsE\_\{\\mathrm\{ads\}\}interval is purely log compression\.kdesk\_\{\\mathrm\{des\}\}can be anchored because the dataset contains near\-saturated cases \(where the desorption term has appreciable magnitude\)\.
Not separately identifiable\.kadsk\_\{\\mathrm\{ads\}\}: the data constrain only the adsorption\-flux productkadscwallk\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}, sokadsk\_\{\\mathrm\{ads\}\}is reported as an effective value \(8 seeds0\.0116±0\.00050\.0116\\pm 0\.0005\)\. The small cross\-seed s\.d\. is not evidence of identifiability \(all runs share the initial value\); the non\-identifiability is established by three lines of evidence\.
Three mechanisms\.\(i\)*Structural \(Fisher\)\.*With the learned near\-wall concentration closure fixed, the Fisher information matrix of\(log10kads,log10kdes\)\(\\log\_\{10\}k\_\{\\mathrm\{ads\}\},\\log\_\{10\}k\_\{\\mathrm\{des\}\}\)has condition number≈1\.4×102\\approx 1\.4\\times 10^\{2\}; once the concentration scale of the closure is allowed to vary, the condition number jumps to≈1\.25×109\\approx 1\.25\\times 10^\{9\}with the flattest eigendirection\(\+1,−1\)/2\(\+1,\-1\)/\\sqrt\{2\}, i\.e\. information vanishes along thekadscwallk\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}\-invariant direction\. The Fisher matrix is constructed as follows: with the parameter vector𝜷\\boldsymbol\{\\beta\}and the model log\-coverage predictionfi\(𝜷\)=log10θ¯A,if\_\{i\}\(\\boldsymbol\{\\beta\}\)=\\log\_\{10\}\\bar\{\\theta\}\_\{A,i\}at each caseii, we take the autodiff gradients∂fi/∂βj\\partial f\_\{i\}/\\partial\\beta\_\{j\}at the converged solution and assemble the Gauss–Newton approximationFjk=σ−2∑i\(∂fi/∂βj\)\(∂fi/∂βk\)F\_\{jk\}=\\sigma^\{\-2\}\\sum\_\{i\}\(\\partial f\_\{i\}/\\partial\\beta\_\{j\}\)\(\\partial f\_\{i\}/\\partial\\beta\_\{k\}\)\. HereFFis taken from the*data term only*, excluding the auxiliary supervision and regularisers, in order to isolate the constraint of the data itself\. \(ii\)*Initial\-value drift\.*Varying thekadsk\_\{\\mathrm\{ads\}\}initial value \(0\.001/0\.01/0\.10\.001/0\.01/0\.1\) drives the converged value monotonically \(0\.0104/0\.0376/0\.05850\.0104/0\.0376/0\.0585\), not to the truth\. \(iii\)*Profile likelihood\.*Thekadsk\_\{\\mathrm\{ads\}\}profile has a clear minimum but offset from the truth by∼7×\\sim 7\\times\(kads≈0\.07k\_\{\\mathrm\{ads\}\}\\approx 0\.07vs0\.010\.01\), i\.e\. the fit actively prefers a wrongkadsk\_\{\\mathrm\{ads\}\}; theEadsE\_\{\\mathrm\{ads\}\}profile minimum lies near the truth\. The three differ in whether the minimum sits at the truth, not in “flat versus sharp”\.
Deepened structural analysis\.Beyond the condition number, we examine the full eigenstructure of the closure\-free Fisher matrix in\(log10kads,log10kdes,log10c\-scale\)\(\\log\_\{10\}k\_\{\\mathrm\{ads\}\},\\log\_\{10\}k\_\{\\mathrm\{des\}\},\\log\_\{10\}c\\text\{\-scale\}\)\. Its eigenvalues spanλ1≈3\.3×104\\lambda\_\{1\}\\approx 3\.3\\times 10^\{4\},λ2≈1\.1×102\\lambda\_\{2\}\\approx 1\.1\\times 10^\{2\}, andλ3\\lambda\_\{3\}at the double\-precision floor \(\|λ3\|/λ1≈2×10−16\|\\lambda\_\{3\}\|/\\lambda\_\{1\}\\approx 2\\times 10^\{\-16\}, an*exact structural zero*contaminated only by round\-off\), with null eigenvector\(\+1,0,−1\)/2\(\+1,0,\-1\)/\\sqrt\{2\}—thekadscwallk\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}\-invariant direction\. This structural \(model\-level\) non\-identifiability ofkadsk\_\{\\mathrm\{ads\}\}admits an analytic proof: the adsorption flux enters Eq\. \(2\) only as the productkadscwallk\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}, so the reparameterisationkads→αkadsk\_\{\\mathrm\{ads\}\}\\\!\\to\\\!\\alpha k\_\{\\mathrm\{ads\}\},cwall→cwall/αc\_\{\\mathrm\{wall\}\}\\\!\\to\\\!c\_\{\\mathrm\{wall\}\}/\\alphaleaves*every*prediction unchanged; the likelihood is exactly flat along it and the corresponding Fisher eigenvalue is identically zero\. Because the matrix is structurally singular we do not quote a finite condition number for it \(the1\.25×1091\.25\\times 10^\{9\}above is the value when the concentration scale is only*partially*constrained by the supervision, i\.e\. a practical\-identifiability estimate\)\. Displacing the parameters by±10%\\pm 10\\%along this null direction changes the model output by less than the label noise \(a direct visualisation of the sloppy manifold\), whereas the same displacement alongλ1\\lambda\_\{1\}changes it by orders of magnitude; the singular\-value spectrum of the prediction\-sensitivity matrix∂fi/∂βj\\partial f\_\{i\}/\\partial\\beta\_\{j\}mirrors this hierarchy\. We therefore distinguish*structural*non\-identifiability \(kadsk\_\{\\mathrm\{ads\}\}alone, exactly unrecoverable for any data\) from*practical*non\-identifiability \(ν\\nuat multiple temperatures, recoverable only to a sub\-decade interval \(χ12\\chi^\{2\}\_\{1\}95%95\\%width≈0\.6\\approx 0\.6decade\) given finite, noisy data; Section 7\.2\)\.


Figure 7:Single\-temperature identifiability\. Top: Fisher confidence ellipses—once the concentration scale is freed, the likelihood is nearly flat along thekadscwallk\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}\-invariant direction\. Bottom: profile likelihood—thekadsk\_\{\\mathrm\{ads\}\}profile has a clear minimum offset from the truth by∼7×\\sim 7\\times\(minimum off\-true\), whereas theEadsE\_\{\\mathrm\{ads\}\}profile minimum sits at the truth\.


Figure 8:Deepened structural analysis\.\(a\)Eigenvalue spectrum of the closure\-free Fisher matrix—λ3\\lambda\_\{3\}is an exact structural zero at the double\-precision floor \(\|λ3\|/λ1≈2×10−16\|\\lambda\_\{3\}\|/\\lambda\_\{1\}\\\!\\approx\\\!2\\times 10^\{\-16\}\)\.\(b\)Singular\-value spectrum of the prediction\-sensitivity matrix∂fi/∂βj\\partial f\_\{i\}/\\partial\\beta\_\{j\}, mirroring the hierarchy\.\(c\)Sloppy\-manifold visualisation—moving±10%\\pm 10\\%along the null direction leaves the model output essentially unchanged, while the same step along the stiff direction changes it by orders of magnitude\.
### 6\.6\. Prior ablation: identification does not depend on prior leakage
The single\-temperature inversion includes a weakkdesk\_\{\\mathrm\{des\}\}prior whose centre is set at the truth by default, raising the concern that the prior \(and a same\-magnitude initialisation\) might leak the truth\. We therefore re\-run the otherwise identical inversion while varying only the centre and presence of this prior \(Table[4](https://arxiv.org/html/2608.00212#S6.T4)\)\. Two facts rule out leakage: \(i\) with the prior*removed*,Eads=0\.772E\_\{\\mathrm\{ads\}\}=0\.772eV \(within0\.3%0\.3\\%of truth\), essentially unchanged; and \(ii\) even with the prior centre mis\-placed by±2\\pm 2decades,EadsE\_\{\\mathrm\{ads\}\}stays within0\.7620\.762–0\.7880\.788eV, a spread of only2626meV across all settings\.kdesk\_\{\\mathrm\{des\}\}itself is mildly pulled by a mis\-placed prior \(0\.580\.58–1\.581\.58\), but after the log compression ofEads=kBTln\(ν/kdes\)E\_\{\\mathrm\{ads\}\}=k\_\{B\}T\\ln\(\\nu/k\_\{\\mathrm\{des\}\}\)this is only±1\.5%\\pm 1\.5\\%inEadsE\_\{\\mathrm\{ads\}\}\. The testRlog2R^\{2\}\_\{\\log\}is0\.9970\.997for all settings\. HenceEadsE\_\{\\mathrm\{ads\}\}is data\-driven, not prior\-leaked; the prior acts only as numerical regularisation\.
Table 4:kdesk\_\{\\mathrm\{des\}\}prior ablation \(wp=0\.01w\_\{p\}=0\.01, 8\-seed mean±\\pms\.d\.\)\.
### 6\.7\. Leave\-one\-out cross\-validation
Because the fixed split places near\-saturated cases mostly in training, one may ask whether the 9\-point testR2R^\{2\}is favourably biased\. We therefore perform leave\-one\-out cross\-validation \(LOOCV\) over all 48 cases: each fold holds out one case as a clean test and trains on the other 47\. The overall LOOCVRlog2=0\.996R^\{2\}\_\{\\log\}=0\.996andRraw2=0\.974R^\{2\}\_\{\\mathrm\{raw\}\}=0\.974, essentially identical to the fixed split \(0\.998/0\.9890\.998/0\.989\); the invertedEads=0\.776±0\.003E\_\{\\mathrm\{ads\}\}=0\.776\\pm 0\.003eV andkdes=0\.94±0\.09k\_\{\\mathrm\{des\}\}=0\.94\\pm 0\.09s\-1also agree with the main results\. The∼10%\\sim\\\!10\\%shift inkdesk\_\{\\mathrm\{des\}\}relative to the fixed\-split1\.04±0\.031\.04\\pm 0\.03s\-1\(Section 6\.5\) lies within the leave\-one\-out fold\-to\-fold spread \(±0\.09\\pm 0\.09\) and reflects the subset composition of the held\-out folds, not an instability; the identifiable quantityEadsE\_\{\\mathrm\{ads\}\}agrees to within0\.4%0\.4\\%\(0\.7760\.776vs0\.7730\.773eV\), as expected since the weak\-constraintkdesk\_\{\\mathrm\{des\}\}variation is compressed to±1\.5%\\pm 1\.5\\%inEadsE\_\{\\mathrm\{ads\}\}throughEads=kBTln\(ν/kdes\)E\_\{\\mathrm\{ads\}\}=k\_\{B\}T\\ln\(\\nu/k\_\{\\mathrm\{des\}\}\)\. The original split therefore does not embellish the metrics\. The LOOCVRraw2R^\{2\}\_\{\\mathrm\{raw\}\}and the maximum log error \(0\.1870\.187, at thevsub=1\.2v\_\{\\mathrm\{sub\}\}=1\.2high\-speed saturated corner\) further expose the saturated\-corner model\-form bias of Section 6\.1, which LOOCV honestly reveals rather than hides\. The anchoring necessity of “saturated cases in training” is reconciled with evaluation fairness because each training fold still contains all other saturated cases\. As a further check at coarser granularity,55\- and1010\-fold cross\-validation give the same picture, with the invertedEadsE\_\{\\mathrm\{ads\}\}varying by only∼1\.5\\sim 1\.5meV across folds, confirming that the inversion is stable to the train/test partition\.
### 6\.8\. Computational cost
The 48\-case COMSOL solve \(production mesh, 16\-core CPU\) takes about17,11717\{,\}117s in total, i\.e\.≈357\\approx 357s per case\. The one\-off PCINN training takes≈1\.6\\approx 1\.6min \(≈13\\approx 13min for 8 seeds\); a trained query needs a single forward pass:≈0\.34\\approx 0\.34ms amortised per case in batch,≈7\\approx 7ms standalone\. The model is deliberately tiny: the physics branch is a2→64→64→64→64→12\\\!\\to\\\!64\\\!\\to\\\!64\\\!\\to\\\!64\\\!\\to\\\!64\\\!\\to\\\!1MLP \(12,73712\{,\}737weights and biases\) plus two trainable kinetic scalars,12,73912\{,\}739parameters in total \(≈12,800\\approx 12\{,\}800for the three\-input multi\-temperature pipeline\); a single coverage evaluation costs only≈3×104\\approx 3\\times 10^\{4\}FLOPs \(≈2\.5×104\\approx 2\.5\\times 10^\{4\}for the MLP and≈4×103\\approx 4\\times 10^\{3\}for the320320\-step trajectory integration\)\. The millisecond inference time is therefore dominated by the serial integration loop rather than by arithmetic, and the small parameter count and FLOPs are a direct consequence of hard\-coding the chemistry and compressing the data\-driven freedom to a single scalar\. All computations run on a1616\-core CPU \(an RTX 4070 Laptop GPU is available but the serial chemistry integration does not benefit from it\)\. Thus, once the dataset and model are ready, predicting a new case is∼5×104\\sim 5\\times 10^\{4\}\(standalone\) to∼1\.1×106\\sim 1\.1\\times 10^\{6\}\(amortised\) times faster than CFD\. Honestly, generating the training set still incurs the CFD cost, so the surrogate does not eliminate the up\-front investment but amortises it: once the number of required evaluations exceeds the dataset size \(e\.g\. thousands of evaluations for operating\-window optimisation\), the advantage is realised—scanning10410^\{4\}cases with PCINN takes≈3\.4\\approx 3\.4s versus≈42\\approx 42days of case\-by\-case CFD\.
## 7\. Multi\-temperature extension and a model\-mismatch test
At a single temperature,EadsE\_\{\\mathrm\{ads\}\}andkdesk\_\{\\mathrm\{des\}\}are the same identifiable quantity, so the single\-temperature “EadsE\_\{\\mathrm\{ads\}\}” is converted from an*assumed*ν\\nu\. Separatingν\\nuandEadsE\_\{\\mathrm\{ads\}\}intuitively calls for multi\-temperature data, with the slope and intercept of the Arrhenius linelnkdes\(T\)\\ln k\_\{\\mathrm\{des\}\}\(T\)versus1/T1/TgivingEadsE\_\{\\mathrm\{ads\}\}andν\\nu\. We test this intuition over four temperatures and characterise the resulting identifiability faithfully\.
### 7\.1\. Multi\-temperature dataset and free inversion
We sweepT∈\{300,320,340,360\}T\\in\\\{300,320,340,360\\\}K×\\timesvsubv\_\{\\mathrm\{sub\}\}\(8\)×\\timesUcurtain∈\{0\.5,2,8\}U\_\{\\mathrm\{curtain\}\}\\in\\\{0\.5,2,8\\\}\(=96=96cases\), withkdes≈\{1\.0,6\.5,33\.7,146\}k\_\{\\mathrm\{des\}\}\\approx\\\{1\.0,6\.5,33\.7,146\\\}s\-1spanning two decades and remaining saturable throughout\. \(Higher temperatures, e\.g\.450450K, would collapse the coverage to∼10−5\\sim 10^\{\-5\}and lose the saturated cases that anchorkdesk\_\{\\mathrm\{des\}\}\.\) The wall shear rates are identical across temperatures, confirming flow–temperature decoupling\. The physics\-branch input is augmented to\(v∗,U∗,T∗\)\(v^\{\*\},U^\{\*\},T^\{\*\}\); the chemistry branch trainslog10ν\\log\_\{10\}\\nuandEadsE\_\{\\mathrm\{ads\}\}directly \(dropping thekdesk\_\{\\mathrm\{des\}\}prior\), synthesisingkdes\(T\)k\_\{\\mathrm\{des\}\}\(T\)per sample\.
Withθ¯A\\bar\{\\theta\}\_\{A\}over 8 seeds, the surrogate reaches testRlog2=0\.9991±0\.0003R^\{2\}\_\{\\log\}=0\.9991\\pm 0\.0003withEadsE\_\{\\mathrm\{ads\}\}cross\-seed scatter of only±0\.002\\pm 0\.002eV—highly reproducible\. Yet the point estimateEads=0\.711±0\.002E\_\{\\mathrm\{ads\}\}=0\.711\\pm 0\.002eV is systematically8%8\\%below truth, andlog10ν=11\.998±0\.006\\log\_\{10\}\\nu=11\.998\\pm 0\.006barely leaves its initial value of12\.012\.0, i\.e\.ν\\nuis not data\-driven but parked at the initialisation, whilekadsk\_\{\\mathrm\{ads\}\}\(effective\) drifts to0\.0530\.053to absorb the amplitude slack \(Table[5](https://arxiv.org/html/2608.00212#S7.T5)\)\.It must be stated plainly that, under multiple temperatures, the free\-inversion point estimates ofν\\nuandEadsE\_\{\\mathrm\{ads\}\}are not trustworthy and should not be reported as inversion results; the identifiability conclusion must be given via the profile likelihood interval or an independent prior \(Section 7\.3\)\.TheEads=0\.711E\_\{\\mathrm\{ads\}\}=0\.711in Table[5](https://arxiv.org/html/2608.00212#S7.T5)is therefore to be read as “the parked value of the optimiser in the degeneracy valley \(a non\-unique solution\)”, not as an estimate ofEadsE\_\{\\mathrm\{ads\}\}\.
Table 5:Multi\-temperature free inversion \(labelθ¯A\\bar\{\\theta\}\_\{A\}, 8\-seed mean±\\pms\.d\.\)\.
### 7\.2\. A degeneracy valley containing the ground truth
Initial\-value drift\.Fixing theEadsE\_\{\\mathrm\{ads\}\}andkadsk\_\{\\mathrm\{ads\}\}initial values and placinglog10ν\\log\_\{10\}\\nuat\{11,12,13\}\\\{11,12,13\\\}\(3 seeds each\), the convergedlog10ν\\log\_\{10\}\\nufollows its initial value \(regression slope0\.970\.97\), while the re\-optimisedEadsE\_\{\\mathrm\{ads\}\}takes0\.652/0\.711/0\.7690\.652/0\.711/0\.769, all with testR2\>0\.998R^\{2\}\>0\.998\. The three points are collinear, giving the degeneracy lineEads=0\.064log10ν−0\.061E\_\{\\mathrm\{ads\}\}=0\.064\\log\_\{10\}\\nu\-0\.061; extrapolating tolog10ν=13\\log\_\{10\}\\nu=13givesEads=0\.775E\_\{\\mathrm\{ads\}\}=0\.775eV, so the ground truth\(1013,0\.774\)\(10^\{13\},0\.774\)lies on the line\.
Profile likelihood\.Freezinglog10ν\\log\_\{10\}\\nuon a grid\[11\.0,13\.5\]\[11\.0,13\.5\]and re\-optimising the rest, the data\-misfit curve is a shallow one\-sided valley with a floor near13\.2513\.25and a curvature of only2\.2×10−42\.2\\times 10^\{\-4\}/decade2\. The practical identifiable interval, placed on a principled footing by aχ12\\chi^\{2\}\_\{1\}likelihood\-ratio criterion rather than an ad\-hoc misfit multiple \(the profile misfit is a log\-space mean\-squared error; with theσ=5%\\sigma=5\\%label noise the95%95\\%level corresponds to a misfit rise ofχ1,0\.952σ2/N\\chi^\{2\}\_\{1,0\.95\}\\,\\sigma^\{2\}/Nabove the floor\), islog10ν∈\[12\.85,13\.43\]\\log\_\{10\}\\nu\\in\[12\.85,13\.43\]\(≈0\.6\\approx 0\.6decade, containing the truthlog10ν=13\\log\_\{10\}\\nu=13\), giving a lower boundν≳1012\.85\\nu\\gtrsim 10^\{12\.85\}\. The previously reported “2×2\\timesfloor” band\[12\.5,13\.5\]\[12\.5,13\.5\]\(≈1\\approx 1decade\) in fact corresponds to a much higher confidence \(∼99\.99%\\sim 99\.99\\%\) and is thus a conservative over\-estimate of the width; theχ12\\chi^\{2\}\_\{1\}interval is the defensible one\.Conditioning onν=νtrue\\nu=\\nu\_\{\\mathrm\{true\}\}recoversEads=0\.776±0\.001E\_\{\\mathrm\{ads\}\}=0\.776\\pm 0\.001eV \(0\.3%0\.3\\%error\): givenν\\nu, the multi\-temperature data recoverEadsE\_\{\\mathrm\{ads\}\}almost exactly; the free\-inversion0\.7110\.711is purely the optimiser parking low in the valley \(small gradient on the shallow slope; classic sloppy behaviour\)\. The two independent experiments yield the same degeneracy line\.
Theoretical slope\.The valley is not a numerical artefact but an exact structural consequence of the Arrhenius form\. Fixing the desorption rate at a reference temperatureT0T\_\{0\},kdes\(T0\)=νe−Eads/kBT0k\_\{\\mathrm\{des\}\}\(T\_\{0\}\)=\\nu\\,e^\{\-E\_\{\\mathrm\{ads\}\}/k\_\{B\}T\_\{0\}\}implieslnν=lnkdes\(T0\)\+Eads/\(kBT0\)\\ln\\nu=\\ln k\_\{\\mathrm\{des\}\}\(T\_\{0\}\)\+E\_\{\\mathrm\{ads\}\}/\(k\_\{B\}T\_\{0\}\); holding the data\-constrainedkdes\(T0\)k\_\{\\mathrm\{des\}\}\(T\_\{0\}\)fixed therefore bindsEadsE\_\{\\mathrm\{ads\}\}andlog10ν\\log\_\{10\}\\nualong a line of slope
dEadsdlog10ν=kBTeffln10\.\\frac\{dE\_\{\\mathrm\{ads\}\}\}\{d\\log\_\{10\}\\nu\}=k\_\{B\}\\,T\_\{\\mathrm\{eff\}\}\\,\\ln 10\.\(6\)With multi\-temperature data the Arrhenius constraint weights1/T1/T, so the effective temperature is the harmonic mean of the sampled temperatures,Teff=⟨1/T⟩−1=328\.5T\_\{\\mathrm\{eff\}\}=\\langle 1/T\\rangle^\{\-1\}=328\.5K for\{300,320,340,360\}\\\{300,320,340,360\\\}K, giving a*predicted*slope of0\.06520\.0652eV/decade—within0\.7%0\.7\\%of the observed0\.06470\.0647\. Three consequences follow: \(i\) the slope is fixed analytically by the Arrhenius law and the temperature window, independent of the ground truth, the network, or noise, so the valley is a structural property rather than a numerical coincidence; \(ii\) it predicts that the slope is nearly invariant under any model mismatch that preserves a*single*Arrhenius process, and shifts only when a*second*thermally activated process is introduced \(both confirmed in Section 7\.3\); and \(iii\) breaking the degeneracy requires enlargingTeffT\_\{\\mathrm\{eff\}\}\(a wider temperature lever, which loses the saturated anchoring cases\) or supplying an independentν\\nuprior\. The observed identifiable interval \(χ12\\chi^\{2\}\_\{1\}95%95\\%width≈0\.6\\approx 0\.6decade\) is thus a direct quantitative consequence of the6060K temperature window: withTeff=328\.5T\_\{\\mathrm\{eff\}\}=328\.5K the slopekBTeffln10=0\.0652k\_\{B\}T\_\{\\mathrm\{eff\}\}\\ln 10=0\.0652eV/decade sets how far alonglog10ν\\log\_\{10\}\\nuthe valley extends before the data\-misfit doubles, so the interval width and the window are two views of the same quantity—the geometry is therefore usable as a design tool \(Section 8\.4\), telling in advance how much wider a temperature range would be needed to separateν\\nuandEadsE\_\{\\mathrm\{ads\}\}\. A partial remedy that requires no new data is to re\-parameterise the Arrhenius law about a reference temperature inside the window, writinglnkdes\(T\)=lnkdes\(Tref\)−EadskB\(1T−1Tref\)\\ln k\_\{\\mathrm\{des\}\}\(T\)=\\ln k\_\{\\mathrm\{des\}\}\(T\_\{\\mathrm\{ref\}\}\)\-\\tfrac\{E\_\{\\mathrm\{ads\}\}\}\{k\_\{B\}\}\\\!\\left\(\\tfrac\{1\}\{T\}\-\\tfrac\{1\}\{T\_\{\\mathrm\{ref\}\}\}\\right\): this re\-centres the fit on the well\-constrained combinationkdes\(Tref\)k\_\{\\mathrm\{des\}\}\(T\_\{\\mathrm\{ref\}\}\)and largely decorrelates the estimated slope \(EadsE\_\{\\mathrm\{ads\}\}\) from the intercept \(lnν\\ln\\nu\), tightening the conditional interval without wideningTeffT\_\{\\mathrm\{eff\}\}\. It does not break the structural degeneracy—the flat direction persists—but it improves its numerical conditioning, and is the natural finite\-data companion to the manifold\-boundary approximation method of Transtrum and Qiu, in which such degenerate directions are removed by boundary \(limiting\) reparameterisations rather than by more data\.
Generality beyond the strict Arrhenius prefactor\.The slope law depends only on the activated \(exponential\) factor and the sampled temperature set, not on the form of the prefactor\. For any generalized ratekdes\(T\)=νg\(T\)e−Eads/kBTk\_\{\\mathrm\{des\}\}\(T\)=\\nu\\,g\(T\)\\,e^\{\-E\_\{\\mathrm\{ads\}\}/k\_\{B\}T\}in whichg\(T\)g\(T\)is a*known*temperature modulation—modified Arrheniusg\(T\)=Tng\(T\)=T^\{n\}, Eyring/transition\-stateg\(T\)=kBT/hg\(T\)=k\_\{B\}T/h, or any fixed dependence—the logarithmlnkdes=lnν\+lng\(T\)−Eads/\(kBT\)\\ln k\_\{\\mathrm\{des\}\}=\\ln\\nu\+\\ln g\(T\)\-E\_\{\\mathrm\{ads\}\}/\(k\_\{B\}T\)placesggin a term that is known and hence does not enter theν\\nu–EadsE\_\{\\mathrm\{ads\}\}trade\-off\. The degeneracy direction is set by the−Eads/\(kBT\)\-E\_\{\\mathrm\{ads\}\}/\(k\_\{B\}T\)term alone, so the slope remainskBTeffln10k\_\{B\}T\_\{\\mathrm\{eff\}\}\\ln 10with the*same*harmonic\-meanTeffT\_\{\\mathrm\{eff\}\}\. The law is thus a property of single\-channel activated\-rate inversion in general, not of the strict Arrhenius prefactor or of ALD; it is altered only by a genuinely different activation structure—a second exponential \(Section 7\.3\) or a temperature\-dependent barrierEads\(T\)E\_\{\\mathrm\{ads\}\}\(T\)—which is exactly what makes the slope a diagnostic for such structure\.
Figure 9:Multi\-temperatureν\\nu–EadsE\_\{\\mathrm\{ads\}\}degeneracy valley\.\(a\)Re\-optimizedEadsE\_\{\\mathrm\{ads\}\}versus fixedlog10ν\\log\_\{10\}\\nu: the solutions slide along one degeneracy line with the ground truth on it, and conditioning onνtrue\\nu\_\{\\mathrm\{true\}\}recoversEads≈0\.776E\_\{\\mathrm\{ads\}\}\\approx 0\.776eV\.\(b\)Profile likelihood inlog10ν\\log\_\{10\}\\nu—a shallow one\-sided valley; the shaded band is the heuristic “2×2\\timesfloor” regionlog10ν∈\[12\.5,13\.5\]\\log\_\{10\}\\nu\\in\[12\.5,13\.5\], which corresponds to a very high \(∼99\.99%\\sim 99\.99\\%\) confidence; the defensibleχ12\\chi^\{2\}\_\{1\}likelihood\-ratio95%95\\%interval is the tighter\[12\.85,13\.43\]\[12\.85,13\.43\]\(≈0\.6\\approx 0\.6decade,ν≳1012\.85\\nu\\gtrsim 10^\{12\.85\}; Section 7\.2\)\.\(c\)Arrhenius plotlnkdes\\ln k\_\{\\mathrm\{des\}\}vs1/T1/T: the free\-inversion line \(parked low in the valley\) versus the ground truth, which theν\\nu\-conditioned estimate reproduces\. The vertical dotted line marks the truthlog10ν=13\\log\_\{10\}\\nu=13; the dotted horizontal line in \(a\) marksEads=0\.774E\_\{\\mathrm\{ads\}\}=0\.774eV\.
### 7\.3\. Model\-mismatch matrix: seven chemistries across three mechanisms
Since the data are generated and inverted with the same Langmuir–Arrhenius form, we test whether the conclusions persist when the generating chemistry differs from the inversion assumption\. Rather than a single mismatch, we build a matrix of*seven*generating chemistries spanning*three orthogonal mechanism families*—desorption\-side coverage dependence \(Temkin\), adsorption\-side non\-linearity \(Freundlich\), and site heterogeneity \(dual\-site\)—each generated by COMSOL with the*same*underlying Arrhenius truth and inverted with the*unchanged*standard single\-site Langmuir PCINN \(profile likelihood, best\-seed\)\. We first detail the canonical desorption\-side case \(Temkinβ=2\\beta=2\) and then summarise the full matrix \(Table[7](https://arxiv.org/html/2608.00212#S7.T7)\)\. The Temkin desorption reads
Jnetgen=kadscwall\(1−θ\)−kdes,0eβθΓsθ,kdes,0=νe−Eads/kBT,β=2\.0,J\_\{\\mathrm\{net\}\}^\{\\mathrm\{gen\}\}=k\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}\(1\-\\theta\)\-k\_\{\\mathrm\{des\},0\}\\,e^\{\\beta\\theta\}\\,\\Gamma\_\{s\}\\,\\theta,\\hskip 18\.49988ptk\_\{\\mathrm\{des\},0\}=\\nu e^\{\-E\_\{\\mathrm\{ads\}\}/k\_\{B\}T\},\\ \\beta=2\.0,\(7\)wherekdes,0k\_\{\\mathrm\{des\},0\}keeps the same Arrhenius truth and the mismatch enters only througheβθe^\{\\beta\\theta\}\(motivated by the fact that the Langmuir assumption of coverage\-independent heat of adsorption generally fails; the Temkin isotherm is the most common correction\)\. Withβ=2\\beta=2, the saturated\-corner coverage drops by36%36\\%–60%60\\%relative to Langmuir\. We re\-run the 96\-case scan with this mismatched chemistry \(all converged\) and invert with the*unchanged*standard Langmuir PCINN\. Three findings emerge \(Table[6](https://arxiv.org/html/2608.00212#S7.T6)\)\.
First, the surrogate is robust:Rraw2R^\{2\}\_\{\\mathrm\{raw\}\}drops only from0\.9920\.992to0\.9870\.987—the Langmuir form still absorbs most of the Temkin data\. Second, the free\-inversionEadsE\_\{\\mathrm\{ads\}\}barely moves \(0\.711→0\.7050\.711\\to 0\.705, only66meV\) despite the strong mismatch, which strongly corroborates that the free\-inversion point estimate is dominated by the parking position in the valley and is not trustworthy; the mismatch instead leaves a diagnosable signature in the systematic temperature structure of the per\-temperaturekdesk\_\{\\mathrm\{des\}\}bias \(the pivot shifts from∼320\\sim 320K to∼340\\sim 340K and the low\-temperature positive bias amplifies,\+14\.7%→\+39\.2%\+14\.7\\%\\to\+39\.2\\%\)\. Third, and central to the diagnostic, thedegeneracy valley persists under mismatch: the profile degeneracy slope is0\.06450\.0645\(versus0\.06470\.0647self\-consistent\), the conditional inversion givesEads=0\.768E\_\{\\mathrm\{ads\}\}=0\.768eV \(versus0\.7760\.776, only0\.7%0\.7\\%degradation\), and the valley merely flattens further \(curvature1\.5×10−41\.5\\times 10^\{\-4\}, interval widening to2\.52\.5decades\) so thatν\\nubecomes harder to pin down, without breaking the identifiability geometry\.
Table 6:Self\-consistent \(Langmuir\) versus mismatched \(Temkin\) profile/degeneracy\.Table 7:Model\-mismatch matrix: seven core generating chemistries across three mechanism families, plus two additional dual\-site variants \(the energy\-split calibration sweep, Section 7\.3\) and one*transport\-layer*mismatch \(non\-Fickian, concentration\-dependent diffusivity\)—each inverted with the*unchanged*single\-site Langmuir PCINN \(profile likelihood, best\-seed\)\. The degeneracy slope is nearly invariant \(0\.0630\.063–0\.0700\.070, theory0\.06520\.0652, Section 7\.2\) across all six*single\-Arrhenius*chemistries*and*under the transport perturbation; among the dual\-site \(double\-Arrhenius\) cases it is shifted above the band only at the intermediate energy split \(ΔE=0\.083\\Delta E=0\.083\), while the small and large splits stay within it and are instead flagged by fit degradation \(Section 7\.3, Fig\.[11](https://arxiv.org/html/2608.00212#S7.F11)\)\.The full matrix and a slope\-invariance boundary\.Extending beyond the canonical Temkin case, we add a stronger desorption dependence \(Temkinβ=3,4\\beta=3,4\), an adsorption\-side non\-linearity \(Freundlich, with the adsorption flux∝\(1−θ\)n\\propto\(1\-\\theta\)^\{n\},n=0\.5,1\.5n=0\.5,1\.5—the opposite mechanism family\), and site heterogeneity \(a dual\-site model with two*independent*Arrhenius desorption channels,Eads\(1\)=0\.732E\_\{\\mathrm\{ads\}\}^\{\(1\)\}=0\.732andEads\(2\)=0\.815E\_\{\\mathrm\{ads\}\}^\{\(2\)\}=0\.815eV at equal site density\)\. Table[7](https://arxiv.org/html/2608.00212#S7.T7)reveals a sharp structural distinction\. All six chemistries that preserve a*single*Arrhenius process—across both the desorption \(Temkin\) and adsorption \(Freundlich\) families—yield a degeneracy slope clustered in0\.0630\.063–0\.0700\.070\(theory0\.06520\.0652, Section 7\.2\) and a conditionalEadsE\_\{\\mathrm\{ads\}\}within3%3\\%of the truth: the slope is invariant not only to the mismatch strength but to the mechanism family\. The case that moves the slope is the dual\-site model, the only one that superposes a*second*Arrhenius process: at an intermediate energy split \(ΔE=0\.083\\Delta E=0\.083\) its slope rises to0\.07550\.0755and its conditionalEadsE\_\{\\mathrm\{ads\}\}to0\.8110\.811eV \(close to the slow, sticky site0\.8150\.815rather than the two\-site mean0\.7740\.774\), exactly as theTeffT\_\{\\mathrm\{eff\}\}theory of Section 7\.2 predicts \(an apparentTeff≈380T\_\{\\mathrm\{eff\}\}\\approx 380K, beyond the data window\)\. This both confirms the theory and*delineates its boundary*: the degeneracy slope is invariant under any single\-Arrhenius mismatch and shifts when a second thermally activated process is introduced—though, as the energy\-split sweep below shows, the magnitude of the shift is itselfΔE\\Delta E\-dependent, peaking when the two channels are comparably weighted over the temperature window\.
An operational threshold with a controlled false\-positive rate\.This turns the slope into a concrete, statistically calibrated test rather than a qualitative observation\. The six single\-Arrhenius chemistries \(plus the non\-Fickian transport mismatch\), spanning both the desorption and adsorption mechanism families, form a tight reference cluster: mean slopeμ=0\.0655\\mu=0\.0655eV/decade with scatterσ=0\.0024\\sigma=0\.0024\(range0\.0630\.063–0\.0700\.070; the collapsed\-fit Temkinβ=4\\beta=4row is excluded as non\-quantitative\)\. The dual\-site signal at0\.07550\.0755lies\+4\.2σ\+4\.2\\sigmaabove this cluster mean—a separation corresponding to a single\-process false\-positive probability of∼1\.5×10−5\\sim 1\.5\\times 10^\{\-5\}, not a marginal∼2×\\sim 2\\timesgap \(Fig\.[10](https://arxiv.org/html/2608.00212#S7.F10)a\)\. We therefore replace the earlier “∼10%\\sim 10\\%” rule with a threshold that controls the false\-positive rate directly: flag a second thermally activated process when the empirical slope exceedsμ\+1\.64σ=0\.0695\\mu\+1\.64\\sigma=0\.0695eV/decade \(5%5\\%false\-positive rate\) or, more conservatively,μ\+2\.33σ=0\.0711\\mu\+2\.33\\sigma=0\.0711eV/decade \(1%1\\%\)—consistent with, and now giving a statistical basis to, the previously stated≳0\.072\\gtrsim 0\.072eV/decade\. The diagnostic is also robust to noise: adding10%10\\%/20%20\\%/30%30\\%label noise \(Section 7\.4\) leaves the single\-process slope at0\.0660\.066/0\.0670\.067/0\.0680\.068, below the5%5\\%threshold up to20%20\\%noise while the dual\-site signal stays clearly above it \(Fig\.[10](https://arxiv.org/html/2608.00212#S7.F10)b\), so the test retains its discriminating power over a realistic noise budget\. Because the warning is computed from the profile geometry rather than the fit residual, it fires precisely in the insidious case where the surrogateR2R^\{2\}remains excellent and no other symptom is visible\.
An energy\-split calibration sweep, and a multi\-signal detection panel\.To test the slope beyond the single dual\-Arrhenius realisation, we generated two further equal\-density two\-site datasets bracketing it, with energy splitsΔE=Eads\(2\)−Eads\(1\)=0\.04\\Delta E=E\_\{\\rm ads\}^\{\(2\)\}\-E\_\{\\rm ads\}^\{\(1\)\}=0\.04and0\.1480\.148eV \(mean fixed at the0\.7740\.774eV truth\), and inverted each with the unchanged single\-site PCINN\. The result is a*detection window*\(Fig\.[11](https://arxiv.org/html/2608.00212#S7.F11)a\): the slope excursion above the single\-process band is maximal at the intermediate split \(ΔE=0\.083\\Delta E=0\.083,0\.07550\.0755,\+4\.2σ\+4\.2\\sigma\) and returns*into*the band at both the small split \(0\.04→0\.06630\.04\\\!\\to\\\!0\.0663\) and the large split \(0\.148→0\.06730\.148\\\!\\to\\\!0\.0673\)\. This is exactly what theTeffT\_\{\\mathrm\{eff\}\}theory predicts: the apparentTeffT\_\{\\mathrm\{eff\}\}\(and hence the slope\) is raised only when the two channels are*comparably weighted*over the sampled window; near degeneracy \(ΔE→0\\Delta E\\\!\\to\\\!0\) there is little heterogeneity to detect, and at largeΔE\\Delta Eone channel dominates the observable and the apparent behaviour reverts toward single\-Arrhenius\. The slope is therefore a*specific but not universally sensitive*flag—a departure implies heterogeneity, but the absence of one does not exclude it\.
The two splits the slope does not flag are nonetheless not missed by the inversion as a whole, because mismatch is diagnosed by a*panel*of signals, not the slope alone \(Fig\.[11](https://arxiv.org/html/2608.00212#S7.F11)b\)\. The small split \(ΔE=0\.04\\Delta E=0\.04\) is near\-degenerate and genuinely benign: a single\-site fit recovers the conditionalEadsE\_\{\\mathrm\{ads\}\}to within1\.3%1\.3\\%\(0\.7640\.764vs0\.7740\.774\), with a profile\-misfit floor \(2\.3×10−42\.3\\times 10^\{\-4\}\) and valley location \(log10ν=13\.25\\log\_\{10\}\\nu\\\!=\\\!13\.25\) indistinguishable from the self\-consistent case—so no signal fires, correctly\. The large split \(ΔE=0\.148\\Delta E=0\.148\), which the slope misses, is caught instead by the*fit geometry*: its profile\-misfit floor is∼22×\\sim\\\!22\\timesthe self\-consistent value \(2\.3×10−32\.3\\times 10^\{\-3\}vs1\.0×10−41\.0\\times 10^\{\-4\}\) and its valley floor is pushed to the edge of theν\\nugrid \(log10ν=11\.0\\log\_\{10\}\\nu\\\!=\\\!11\.0, two decades from the truth\), both strong departures absent in the self\-consistent and small\-split cases\. Thus the slope excursion covers the intermediate\-split regime where the fit stays excellent \(the insidious case\), while fit degradation and valley\-floor displacement cover the large\-split regime; together the panel flags site heterogeneity across theΔE\\Delta Erange wherever it produces a non\-negligible bias, and stays appropriately silent where a single\-site description is adequate\.
Figure 10:Statistical validation of the slope diagnostic\. \(a\) The six single\-Arrhenius generating chemistries \(plus the non\-Fickian transport mismatch\) cluster atμ=0\.0655±0\.0024\\mu=0\.0655\\pm 0\.0024eV/decade; the dual\-site \(double\-Arrhenius\) signal at0\.07550\.0755lies\+4\.2σ\+4\.2\\sigmaabove the cluster \(false\-positive probability∼1\.5×10−5\\sim 1\.5\\times 10^\{\-5\}\), well clear of theα=5%\\alpha=5\\%\(μ\+1\.64σ\\mu\+1\.64\\sigma\) andα=1%\\alpha=1\\%\(μ\+2\.33σ\\mu\+2\.33\\sigma\) flag thresholds\. \(b\) Noise robustness: the single\-process slope stays below the5%5\\%threshold up to20%20\\%label noise while the dual\-site signal remains separable, degrading only at30%30\\%\.Figure 11:Energy\-split calibration and multi\-signal detection\. \(a\) Degeneracy slope versus the dual\-site energy splitΔE\\Delta E: the excursion above the single\-process band \(μ±σ\\mu\\pm\\sigma\) peaks at the intermediate split \(ΔE=0\.083\\Delta E=0\.083, slope\-flagged\) and returns into the band at the small \(0\.040\.04\) and large \(0\.1480\.148\) splits—a detection window predicted by theTeffT\_\{\\mathrm\{eff\}\}theory\. \(b\) Complementary signal: the profile\-misfit floor\. The large split \(ΔE=0\.148\\Delta E=0\.148\), which the slope does not flag, raises the misfit floor∼22×\\sim\\\!22\\timesover the self\-consistent value and pushes theν\\nu\-valley to the grid edge \(two decades from truth\), while the near\-degenerate small split \(ΔE=0\.04\\Delta E=0\.04\) is benign on every signal\. The slope and the fit geometry are thus complementary, covering intermediate and large splits respectively\.The matrix also exposes two complementary mismatch signatures\. Strong*single\-process*mismatch \(Temkinβ=4\\beta=4, Freundlichn=0\.5n=0\.5\) is self\-reporting through a collapse of the surrogate fit \(Rlog2R^\{2\}\_\{\\log\}falling from0\.9970\.997to0\.690\.69–0\.860\.86\): the inversion warns of its own invalidity\. For these collapsed\-fit rows the tabulated slope and conditionalEadsE\_\{\\mathrm\{ads\}\}\(e\.g\.0\.0680\.068and0\.7820\.782eV for Temkinβ=4\\beta=4atRlog2=0\.69R^\{2\}\_\{\\log\}=0\.69\) should be read only qualitatively—they enter Table[7](https://arxiv.org/html/2608.00212#S7.T7)to corroborate that a strong single\-process mismatch self\-reports through fit collapse, not as reliable slope estimates on par with the high\-R2R^\{2\}rows\. Site heterogeneity, by contrast, is the most insidious—it retains an excellent fit \(R2\>0\.99R^\{2\}\>0\.99, no warning\) yet biases the conditionalEadsE\_\{\\mathrm\{ads\}\}by\+4\.8%\+4\.8\\%and steepens the slope by16%16\\%\. Together these bracket when an inversion result can be trusted: a highR2R^\{2\}is necessary but not sufficient, and the residual/slope structure \(not the fit quality\) is the reliable mismatch diagnostic\.
A transport\-layer mismatch\.The chemistry matrix above perturbs only the surface\-kinetic closure; the transport model \(Fickian dilute\-species diffusion\) is shared between the data\-generating CFD and the PCINN forward map\. To probe this remaining shared assumption, we regenerate the9696\-case multi\-temperature dataset with a*non\-Fickian*, concentration\-dependent diffusivityDeff\(c\)=DTMA\(1−12c/C0\)D\_\{\\mathrm\{eff\}\}\(c\)=D\_\{\\mathrm\{TMA\}\}\(1\-\\tfrac\{1\}\{2\}c/C\_\{0\}\)—a structural transport\-model error that alters the operating\-condition→\\tonear\-wall\-concentration mapping the physics branch must learn—and re\-run the*unchanged*Langmuir inversion\. Under this transport perturbation the profile degeneracy slope is0\.06510\.0651eV/decade, essentially unchanged from the0\.06470\.0647of the self\-consistent case at the same full precision \(both round to the0\.0650\.065quoted in Tables[6](https://arxiv.org/html/2608.00212#S7.T6)–[7](https://arxiv.org/html/2608.00212#S7.T7)and agree with the theoretical0\.06520\.0652\), and the conditionalEadsE\_\{\\mathrm\{ads\}\}at the trueν\\nuis0\.7780\.778eV \(versus0\.7760\.776\), with an identical practical interval \(log10ν∈\[12\.75,13\.50\]\\log\_\{10\}\\nu\\in\[12\.75,13\.50\]\) and surrogate quality \(Rlog2≥0\.998R^\{2\}\_\{\\log\}\\geq 0\.998\); of the9696cases,8787converged, the nine failures clustering at the highest\-shear corner \(Ucurtain=8U\_\{\\mathrm\{curtain\}\}=8\), where the non\-Fickian term most stiffens the near\-wall layer\. The learned closure thus*absorbs*the transport error: because the physics branch fits the operating\-condition→cwall\\to c\_\{\\mathrm\{wall\}\}map directly from data, it simply learns the new \(non\-Fickian\) mapping, and the kinetic inversion—which readsEadsE\_\{\\mathrm\{ads\}\}from the*temperature slope*ofθ¯A\\bar\{\\theta\}\_\{A\}—is left unbiased\. This outcome is structural rather than fortuitous, and it fixes the precise scope of the test: the perturbation is*temperature\-independent*, so it rescales the near\-wall concentration amplitude without touching its Arrhenius temperature scaling\. The experiment should therefore be read as a*sufficiency proof for amplitude\-type \(temperature\-independent\) transport mismatch*—it establishes that the physics\-branch closure absorbs any transport error that leaves the temperature dependence intact, and hence cannot, by construction, move the degeneracy slope\. It is deliberately not a test of the harder case: a*temperature\-dependent*transport error—whose physical instances include a temperature\-dependent gas diffusivityD\(T\)D\(T\)or viscosityμ\(T\)\\mu\(T\), both realistic in a heated SALD gap—would introduce an additional temperature channel and, exactly like a second thermally activated process \(Section 7\.2\),*would*be expected to shift the slope—the same falsifiable signature the diagnostic is designed to catch\. That case delineates the boundary of the present test and is left to future work alongside validation on real data\. Within its stated scope the result is nonetheless a genuine extension beyond the chemistry matrix: it demonstrates that the identifiability geometry and slope diagnostic survive a mismatch in the transport layer, not only the chemistry, without substituting for validation against experimental data\.
The structural conclusions of the multi\-temperature identifiability analysis—the existence of the degeneracy valley, its slope, and the “condition onν\\nuthenEadsE\_\{\\mathrm\{ads\}\}is recoverable” reporting convention—thus hold under moderate model mismatch\. The effect of mismatch is a further flattening of the valley and a mild \(0\.7%0\.7\\%\) degradation of the conditionalEadsE\_\{\\mathrm\{ads\}\}, not a breakdown of the identifiability geometry\.
Figure 12:Degeneracy valley under model mismatch \(profile likelihood inlog10ν\\log\_\{10\}\\nu\)\. Across the seven\-chemistry matrix the valley and its slope persist; the conditionalEadsE\_\{\\mathrm\{ads\}\}atνtrue\\nu\_\{\\mathrm\{true\}\}degrades by at most a few percent for single\-Arrhenius mismatches, and only the dual\-site \(double\-Arrhenius\) case shifts the slope \(Table[7](https://arxiv.org/html/2608.00212#S7.T7)\)\.
### 7\.4\. Robustness: noise, bootstrap resampling, and coarse\-grid sampling
Three further perturbations confirm that the degeneracy geometry is not an artefact of the specific regular grid or of the clean labels\. \(i\)*Label noise\.*Adding10%/20%/30%10\\%/20\\%/30\\%Gaussian relative noise to the coverage labels leaves the slope at0\.066/0\.067/0\.0680\.066/0\.067/0\.068up to20%20\\%; only at30%30\\%does the inversion degrade, bounding the noise tolerance\. \(ii\)*Bootstrap resampling\.*Drawing3030training cases at random, repeated2020times, gives a conditionalEads=0\.785±0\.012E\_\{\\mathrm\{ads\}\}=0\.785\\pm 0\.012eV \(median0\.7770\.777\), so the result is insensitive to the training\-subset composition rather than to one favourable split\. \(iii\)*Coarse\-grid, irregular sampling\.*To rule out that the result depends on the fine regular grid—and to address the concern that a dense regular scan resembles interpolation—we coarsen the CFD mesh \(maximum element80μ80~\\mum, tolerance10−310^\{\-3\},≈74\\approx 74s per case, about5×5\\timescheaper\) and draw150150Latin\-hypercube samples over\(vsub,Ucurtain,T\)\(v\_\{\\mathrm\{sub\}\},U\_\{\\mathrm\{curtain\}\},T\)\. The degeneracy slope is0\.0630\.063\(versus0\.0650\.065on the fine regular grid\), the conditionalEadsE\_\{\\mathrm\{ads\}\}extrapolates to0\.7770\.777eV atνtrue\\nu\_\{\\mathrm\{true\}\}, and the valley remains flat inν\\nu\. The identifiability geometry is thus independent of mesh resolution and of sampling regularity\. We deliberately did*not*augment the data with a machine\-learning generator: synthesising data from a model trained on the same cases would form a circular validation loop and add no independent information, whereas bootstrap and coarse\-grid CFD provide genuine robustness checks without that methodological flaw\.
## 8\. Conclusions and outlook
### 8\.1\. Conclusions
We established a saturable\-window SALD model and a physics–chemistry\-informed neural network that, under sparse data, achieves a high\-accuracy coverage surrogate \(testRlog2=0\.998R^\{2\}\_\{\\log\}=0\.998with 30 cases\) and a robust adsorption\-energy inversion, while transparently delineating the identifiability and extrapolation boundaries\. The methodological core is thatthe precision boundary of PCINN inversion is set by the degeneracy structure of parameter space, not by fitting power: at a single temperaturekadsk\_\{\\mathrm\{ads\}\}is not separately identifiable \(onlykadscwallk\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}is\); under multiple temperaturesν\\nuandEadsE\_\{\\mathrm\{ads\}\}are bound along a weakly identifiable degeneracy valley \(slope≈0\.065\\approx 0\.065eV/decade, truth on the line\), withEadsE\_\{\\mathrm\{ads\}\}recoverable to0\.3%0\.3\\%conditional onν\\nu\. We state as a main result, not a caveat, thatmultiple temperatures do not, in this setting, separateν\\nuandEadsE\_\{\\mathrm\{ads\}\}: they only compress the pair into a weakly identifiable interval \(χ12\\chi^\{2\}\_\{1\}95%95\\%width≈0\.6\\approx 0\.6decade\), and genuine separation requires either an independentν\\nuprior or spatially resolved coverage supervision—a deliberately negative finding that we consider as informative as a positive one\. A seven\-chemistry, three\-mechanism mismatch matrix, together with three independent lines of evidence for the slope—its analytic derivationkBTeffln10k\_\{B\}T\_\{\\mathrm\{eff\}\}\\ln 10\(Section 7\.2\), bootstrap resampling, and coarse\-grid regeneration \(Sections 7\.3–7\.4\)—show these conclusions persist, the slope shifting only when a second Arrhenius process is added\. A prior ablation \(mis\-placed or removed\) and LOOCV confirm that the identification does not depend on prior leakage or on the split\.
### 8\.2\. Limitations
\(ii\)kadsk\_\{\\mathrm\{ads\}\}is an effective value under a weak prior; onlykadscwallk\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}andEadsE\_\{\\mathrm\{ads\}\}are identifiable at a single temperature\. \(iiii\) The near\-wall concentration supervision uses a surface/line average rather than a single\-point value \(the latter does not converge in the strong\-adsorption regime\)\. \(iiiiii\) Multiple temperatures do not separately calibrateν\\nuandEadsE\_\{\\mathrm\{ads\}\}but narrow them to a sub\-decade \(≈0\.6\\approx 0\.6\-decade,χ12\\chi^\{2\}\_\{1\}95%95\\%\) weakly identifiable valley; the absolute calibration ofν\\nustill requires an independent prior or spatially resolved profile supervision\. \(iviv\) The two\-segment integration assumescwall=0c\_\{\\mathrm\{wall\}\}=0downstream of the A zone; under a weak curtain \(lowUcurtainU\_\{\\mathrm\{curtain\}\}\), precursor cross\-talk downstream may slightly under\-estimate the coverage, partly co\-sourced with the saturated\-corner under\-estimation; likewise the monotonicity prior assumes a monotone near\-wall concentration inUcurtainU\_\{\\mathrm\{curtain\}\}and would bias the result if real flow features \(e\.g\. recirculation\-induced re\-adsorption\) were non\-monotone\. \(vv\) The model is single\-species \(S1\) and 2\-D steady; the data are synthetic “measurements” without real process data or experimental uncertainty\. Regarding the choice of synthetic ground truth \(Eads=0\.774E\_\{\\mathrm\{ads\}\}=0\.774eV,ν=1013\\nu=10^\{13\}s\-1\): these lie within the DFT range and the transition\-state\-theory magnitude respectively, and the identifiability\-related conclusions do not depend on the specific values—changing the truth only translates the valley in parameter space without altering its existence or orientation \(the mismatch test confirms the slope is nearly unchanged even under a different generating chemistry\)\.
### 8\.3\. Outlook
Pinning downν\\nuabsolutely and breaking thekadscwallk\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}degeneracy require new information sources—spatially resolved coverage/concentration\-profile supervision, or an independent physical prior forν\\nu—rather than widening the temperature window \(which would lose the saturated anchoring cases\)\. A two\-species \(S2\) extension would yield the true bidirectional cross\-talk and operating window\. Conceptually, the degeneracy slope law is expected to carry over to such coupled half\-reactions with an important qualification: as long as each half\-reaction desorbs through a*single*Arrhenius channel, its ownν\\nu–EadsE\_\{\\mathrm\{ads\}\}valley retains the slopekBTeffln10k\_\{B\}T\_\{\\mathrm\{eff\}\}\\ln 10set by the shared temperature window, and the inversion factorises per species\. Coupling between successive adsorption–desorption cycles \(e\.g\. a coverage\-dependent sticking probability of one precursor on sites conditioned by the other\) enters as an effective coverage dependence of the rate—formally the same structure as the Temkin/Freundlich mismatches of Section 7\.3, which leave the slope invariant—so it would perturb the valley floor and the conditionalEadsE\_\{\\mathrm\{ads\}\}without moving the slope,*unless*the coupling introduces a second thermally activated timescale, in which case the slope shifts exactly as in the dual\-site case \(an apparentTeffT\_\{\\mathrm\{eff\}\}departure\)\. The single\-species diagnostic therefore extends to the two\-species setting as a per\-channel test, and a slope anomaly would flag genuinely coupled \(double\-Arrhenius\) inter\-cycle chemistry\. The present study is deliberately scoped as a*concept verification*on a compact, self\-consistent benchmark \(30 training cases; leave\-one\-outRraw2=0\.974R^\{2\}\_\{\\mathrm\{raw\}\}=0\.974controls over\-fitting\), so that the identifiability analysis is not confounded by data volume\. Scaling to substantially larger datasets and to*field\-level*generalisation—predicting full coverage/concentration fields rather than the trajectory\-averagedθ¯A\\bar\{\\theta\}\_\{A\}, e\.g\. via a Fourier neural operator \(FNO\) surrogate—is left to a dedicated follow\-up, of which this paper is the identifiability\-focused first part\. Applying the pipeline to real SALD coverage/thickness data, under experimental noise and model mismatch, is the key next step towards engineering use\.
### 8\.4\. Implications for the methodological community
The identifiability findings have significance beyond SALD\. Thekadscwallk\_\{\\mathrm\{ads\}\}c\_\{\\mathrm\{wall\}\}degeneracy \(single\-temperature\) and theν\\nu–EadsE\_\{\\mathrm\{ads\}\}degeneracy valley \(multi\-temperature\) are instances of a widely known structure: multi\-parameter nonlinear models often exhibit very weak sensitivity along a few “sloppy” directions, rendering parameter combinations along them practically non\-identifiable even when the overall fit is excellent\.
*Relation to prior work, and our increment\.*The sloppy\-model geometry of Transtrum, Machta and Sethna and the profile\-likelihood framework of Raue, Kreutz and co\-workers already established the vocabulary of degeneracy valleys, profile intervals and conditional identifiability that we use here; our analysis lives largely within that framework rather than replacing it\. Rather than claiming priority for applying identifiability analysis*per se*—profile likelihood and Fisher analysis are long established on mechanistic ODE models in systems biology—our increment is a specific*combination*that this literature does not address\. First, we use these diagnostics to locate the*action boundary of a learned transport closure*feeding a hard\-coded kinetic integrator: the negative control of Section 6\.4 shows that the embedded physics confers extrapolation power only along the structurally known \(residence\-time\) axis and none along the data\-driven transport axis—a delineation specific to hybrid \(known\-physics\-plus\-learned\-closure\) models and absent from the generic sloppy\-model picture\. Second, we turn the degeneracy slope from a descriptive feature into a*predictive, falsifiable diagnostic*via the analytic lawkBTeffln10k\_\{B\}T\_\{\\mathrm\{eff\}\}\\ln 10and its invariance boundary\. We are explicit that the derivation of the slope itself is elementary—it is a one\-line differentiation of the Arrhenius relation, not a new result—so the contribution is not the algebra but its conversion into a*statistically validated*test: a threshold with a controlled false\-positive rate, a\+4\.2σ\+4\.2\\sigmasignal\-to\-cluster separation, and a characterised noise budget \(Section 7\.3, Fig\.[10](https://arxiv.org/html/2608.00212#S7.F10)\), together with the delineation of where the hybrid closure’s embedded physics does and does not help\. It is the pairing of these two—an action\-boundary localisation for a learned closure and a physics\-derived,*empirically calibrated*degeneracy test—rather than the underlying algebra, that constitutes the novelty\. We give three cautions for PINN/scientific\-ML inversion\. First,*a high goodness of fit does not imply parameter identifiability*: here the surrogateRlog2R^\{2\}\_\{\\log\}reaches0\.9990\.999yet some parameters \(kadsk\_\{\\mathrm\{ads\}\}at a single temperature,ν\\nuat multiple temperatures\) are fundamentally non\-identifiable\. Second,*a free\-inversion point estimate may merely be the optimiser’s parking position in a degeneracy valley*, drifting with the initialisation and statistically meaningless; the responsible practice is to characterise the identifiable directions via profile likelihood/Fisher analysis and report parameters as intervals or conditional on a stated prior\. Third,*model mismatch need not manifest as a drop in goodness of fit but may hide in the residual structure*\(here, the systematic temperature dependence of the per\-temperaturekdesk\_\{\\mathrm\{des\}\}bias\); diagnosing mismatch should examine the residual structure, not only the overallR2R^\{2\}\.
A practitioner checklist\.These cautions translate into a concrete procedure for anyone inverting kinetics from coverage \(or analogous\) data with a hybrid model\. \(i\) Before trusting any point estimate, run a profile\-likelihood scan over each target parameter and report the interval where the misfit stays within a small multiple of its floor, not a single number\. \(ii\) If a parameter is only weakly identifiable \(a shallow valley\), do not report it as a calibrated value; instead report it*conditional on*an explicitly stated independent prior \(e\.g\. a DFT/TST estimate ofν\\nu\), and propagate that prior’s uncertainty through the conditional estimate\. \(iii\) Use the geometry as a design tool: the slope lawkBTeffln10k\_\{B\}T\_\{\\mathrm\{eff\}\}\\ln 10tells how much a wider temperature window \(largerTeffT\_\{\\mathrm\{eff\}\}\) would tighten the valley, quantifying the experimental effort needed to separateν\\nuandEadsE\_\{\\mathrm\{ads\}\}\. \(iv\) Treat a highR2R^\{2\}as necessary but not sufficient, and screen for model mismatch by inspecting the*structure*of the residuals—here, the temperature dependence of the per\-temperaturekdesk\_\{\\mathrm\{des\}\}bias and the value of the degeneracy slope \(a slope departing fromkBTeffln10k\_\{B\}T\_\{\\mathrm\{eff\}\}\\ln 10flags an unmodelled second activated process, e\.g\. site heterogeneity\), rather than the aggregate fit metric alone\. \(v\) Do not augment a sparse training set with samples drawn from the same model class being validated: this manufactures a circular validation loop in which the generator injects no independent information\. We deliberately avoided ML\-based data augmentation for this reason, using instead bootstrap resampling and a coarse\-grid Latin\-hypercube redraw \(Section 7\.4\) to test robustness without compromising the methodology\.
What transfers to the computational\-physics reader\.Stripped of the ALD specifics, this paper contributes a reusable workflow for*any*inverse problem that couples a known governing law to a learned closure \(a hybrid neural\-ODE / UDE or gray\-box model\)\. The workflow is: \(1\) invert with the known\-physics\-plus\-learned\-closure forward map; \(2\) map each suspect parameter’s degeneracy by profile likelihood \(and its local curvature by the Fisher information\); \(3\) where the governing law admits it,*derive*the analytic degeneracy direction—herekBTeffln10k\_\{B\}T\_\{\\mathrm\{eff\}\}\\ln 10, a quantity fixed by the physics and the sampling design alone—and compare it to the empirical slope; \(4\) read a match as confirmation that the degeneracy is the expected structural one \(report conditional estimates and use the geometry to design a degeneracy\-breaking measurement\), and a departure beyond a stated tolerance as a falsifiable flag for unmodelled structure, even at excellent fit\. Steps \(1\)–\(2\) are generic; the value added here is step \(3\)—an*analytic*, physics\-derived degeneracy direction that upgrades the usual empirical sloppy\-model picture into a predictive, testable diagnostic\. Any activated\-rate or otherwise structurally degenerate inversion—not only surface kinetics—inherits the same construction\.
## Appendix A\. Nomenclature
SymbolMeaning \(unit\)θ\(x\)\\theta\(x\),θ¯A\\bar\{\\theta\}\_\{A\}local / substrate\-averaged A\-zone coverage \(–\)θpeak,θexit\\theta\_\{\\mathrm\{peak\}\},\\theta\_\{\\mathrm\{exit\}\}peak / exit coverage along the trajectory \(–\)kadsk\_\{\\mathrm\{ads\}\}adsorption rate constant \(m s\-1\); truth0\.010\.01kdesk\_\{\\mathrm\{des\}\}desorption rate constant \(s\-1\); truth1\.01\.0at300300KEadsE\_\{\\mathrm\{ads\}\}adsorption energy \(eV\); truth0\.7740\.774ν\\nuArrhenius pre\-exponential factor \(s\-1\); truth101310^\{13\}cwallc\_\{\\mathrm\{wall\}\}near\-wall precursor concentration \(mol m\-3\)Cs∗C\_\{s\}^\{\*\}learned effective near\-wall concentration scale \(–\)CsmidC\_\{s\}^\{\\mathrm\{mid\}\}COMSOL surface\-/line\-averaged near\-wall conc\. \(mol m\-3\)Γs\\Gamma\_\{s\}saturation site density \(mol m\-2\);5×10−65\\times 10^\{\-6\}C0C\_\{0\}inlet precursor concentration \(mol m\-3\);0\.20\.2DDgas\-phase diffusivity \(m2s\-1\);6×10−66\\times 10^\{\-6\}JnetJ\_\{\\mathrm\{net\}\}net adsorption flux \(mol m\-2s\-1\)T,T∗T,\\,T^\{\*\}temperature \(K\); normalisedT∗=T/300T^\{\*\}=T/300TeffT\_\{\\mathrm\{eff\}\}effective \(harmonic\-mean\) temperature,⟨1/T⟩−1\\langle 1/T\\rangle^\{\-1\}\(K\)vsub,v∗v\_\{\\mathrm\{sub\}\},\\,v^\{\*\}substrate velocity \(m s\-1\);v∗=vsub/2\.0v^\{\*\}=v\_\{\\mathrm\{sub\}\}/2\.0Ucurtain,U∗U\_\{\\mathrm\{curtain\}\},\\,U^\{\*\}curtain velocity \(m s\-1\);U∗=Ucurtain/16U^\{\*\}=U\_\{\\mathrm\{curtain\}\}/16kBk\_\{B\}Boltzmann constant;8\.617×10−58\.617\\times 10^\{\-5\}eV K\-1β\\betaTemkin coverage\-dependence exponent \(–\)nnFreundlich exponent in\(1−θ\)n\(1\-\\theta\)^\{n\}\(–\)Rlog2,Rraw2R^\{2\}\_\{\\log\},\\,R^\{2\}\_\{\\mathrm\{raw\}\}coefficient of determination inlog10θ\\log\_\{10\}\\theta/ rawθ\\thetawC,wm,wpw\_\{C\},w\_\{m\},w\_\{p\}loss weights \(concentration, monotonicity, prior\)
## Appendix B\. PCINN forward pass and training \(pseudocode\)
```
# --- Forward: operating condition -> coverage (fully differentiable) ---
# globals: C0, Gamma_s, kB, L_A, L_down, N_A, N_down; per-case: v_sub, T
def coverage(v_star, U_star, T_star, params): # params = (log10_kads,
C_s = softplus(MLP(v_star, U_star, T_star)) # log10_nu, E_ads)
c_wall = C_s * C0
k_ads = 10**params.log10_kads
k_des = params.nu * exp(-params.E_ads / (kB * T)) # Arrhenius (multi-T)
theta, integral = 0.0, 0.0
dx = L_A / N_A # A-zone: c_wall = C_s*C0
for _ in range(N_A):
J = k_ads*c_wall*(1 - theta) - k_des*Gamma_s*theta
theta += (dx / (Gamma_s*v_sub)) * J # forward Euler
integral += theta * dx
dx = L_down / N_down # downstream: c_wall = 0
for _ in range(N_down):
J = -k_des*Gamma_s*theta
theta += (dx / (Gamma_s*v_sub)) * J
integral += theta * dx
return integral / (L_A + L_down) # -> theta_bar_A
# --- Training ---
for seed in range(8):
init MLP, params; optimizer = Adam(lr=1e-3) + ReduceLROnPlateau
repeat until early-stop on val log-data-loss:
pred = coverage(batch);
loss = MSE(log10 pred, log10 label) # data term (log space)
+ w_C * Cs_supervision_loss # anchor C_s magnitude
+ w_m * monotonicity_penalty(dC_s/dU>0)
+ w_p * (log10 k_des - log10 k_des_prior)**2 # single-T only
loss.backward(); optimizer.step()
# report mean +/- s.d. over seeds; identifiability via profile likelihood / Fisher
```
## Data Availability
The data supporting the findings of this study were generated by simulation from the Langmuir–Arrhenius ground truth and the reaction–transport model fully specified in Section 4; all governing equations, boundary conditions, kinetic parameters and sampling ranges are reported in the manuscript, so the datasets \(the single\-temperature and multi\-temperature CFD cases\) can be regenerated independently\. The generated datasets and the associated PCINN training/inference, identifiability\-analysis \(Fisher information, profile likelihood, initial\-value drift\) and figure\-generation code are available from the corresponding author upon reasonable request\.
## Declaration of Generative AI and AI\-Assisted Technologies in the Writing Process
During the preparation of this work, the authors used AI\-assisted tools to support language editing, organization of manuscript drafts, and consistency checks across the manuscript and supplementary files\. The authors reviewed, edited, and verified all AI\-assisted output, including numerical claims and references, and take full responsibility for the content of the publication\.
## Funding
This work was supported by the State Key Laboratory of Ocean Engineering, Shanghai Jiao Tong University \(Grant No\. GKZD010089\), and the State Key Laboratory of Acoustics, Chinese Academy of Sciences \(Grant No\. SKLA202406\)\. The funders had no role in the study design; collection, analysis, and interpretation of data; writing of the manuscript; or decision to submit the article for publication\.
## Declaration of Competing Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper\.
## CRediT Authorship Contribution Statement
Ning Hu: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Data curation, Visualization, Writing – original draft, Writing – review and editing\.
Chang Liu: Conceptualization, Methodology, Supervision, Resources, Funding acquisition, Project administration, Writing – review and editing\.
Yunlei Jiang: Validation, Investigation, Writing – review and editing\.
Yuan Dong: Resources, Funding acquisition, Writing – review and editing\.
## References
1. \[1\]S\. M\. George, Atomic layer deposition: an overview,Chem\. Rev\.110 \(2010\) 111–131\. DOI:10\.1021/cr900056b\.
2. \[2\]P\. Poodt et al\., Spatial atomic layer deposition: a route towards further industrialization,J\. Vac\. Sci\. Technol\. A30 \(2012\) 010802\. DOI:10\.1116/1\.3670745\.
3. \[3\]R\. L\. Puurunen, Surface chemistry of atomic layer deposition: the trimethylaluminum/water process,J\. Appl\. Phys\.97 \(2005\) 121301\. DOI:10\.1063/1\.1940727\.
4. \[4\]I\. Langmuir, The adsorption of gases on plane surfaces of glass, mica and platinum,J\. Am\. Chem\. Soc\.40 \(1918\) 1361–1403\. DOI:10\.1021/ja02242a004\.
5. \[5\]A\. Yanguas\-Gil, J\. W\. Elam, Analytic expressions for atomic layer deposition: coverage, throughput, and materials utilization,J\. Vac\. Sci\. Technol\. A32 \(2014\) 031504\. DOI:10\.1116/1\.4867441\.
6. \[6\]M\. Raissi, P\. Perdikaris, G\. E\. Karniadakis, Physics\-informed neural networks,J\. Comput\. Phys\.378 \(2019\) 686–707\. DOI:10\.1016/j\.jcp\.2018\.10\.045\.
7. \[7\]G\. E\. Karniadakis et al\., Physics\-informed machine learning,Nat\. Rev\. Phys\.3 \(2021\) 422–440\. DOI:10\.1038/s42254\-021\-00314\-5\.
8. \[8\]A\. Raue et al\., Structural and practical identifiability analysis of partially observed dynamical models by exploiting the profile likelihood,Bioinformatics25 \(2009\) 1923–1929\. DOI:10\.1093/bioinformatics/btp358\.
9. \[9\]C\. Rackauckas et al\., Universal differential equations for scientific machine learning, arXiv:2001\.04385 \(2020\)\.
10. \[10\]M\. K\. Transtrum, B\. B\. Machta, J\. P\. Sethna, Geometry of nonlinear least squares with applications to sloppy models and optimization,Phys\. Rev\. E83 \(2011\) 036701\. DOI:10\.1103/PhysRevE\.83\.036701\.
11. \[11\]R\. T\. Q\. Chen, Y\. Rubanova, J\. Bettencourt, D\. Duvenaud, Neural ordinary differential equations, in:Advances in Neural Information Processing Systems \(NeurIPS\), 2018, pp\. 6571–6583\.
12. \[12\]R\. N\. Gutenkunst, J\. J\. Waterfall, F\. P\. Casey, K\. S\. Brown, C\. R\. Myers, J\. P\. Sethna, Universally sloppy parameter sensitivities in systems biology models,PLoS Comput\. Biol\.3 \(2007\) e189\. DOI:10\.1371/journal\.pcbi\.0030189\.
13. \[13\]M\. K\. Transtrum, B\. B\. Machta, K\. S\. Brown, B\. C\. Daniels, C\. R\. Myers, J\. P\. Sethna, Perspective: sloppiness and emergent theories in physics, biology, and beyond,J\. Chem\. Phys\.143 \(2015\) 010901\. DOI:10\.1063/1\.4923066\.
14. \[14\]A\. Raue, M\. Schilling, J\. Bachmann, et al\., Lessons learned from quantitative dynamical modeling in systems biology,PLoS ONE8 \(2013\) e74335\. DOI:10\.1371/journal\.pone\.0074335\.
15. \[15\]A\. F\. Villaverde, A\. Barreiro, A\. Papachristodoulou, Structural identifiability of dynamic systems biology models,PLoS Comput\. Biol\.12 \(2016\) e1005153\. DOI:10\.1371/journal\.pcbi\.1005153\.
16. \[16\]O\. Chiş, J\. R\. Banga, E\. Balsa\-Canto, Structural identifiability of systems biology models: a critical comparison of methods,PLoS ONE6 \(2011\) e27755\. DOI:10\.1371/journal\.pone\.0027755\.
17. \[17\]M\. Temkin, V\. Pyzhev, Kinetics of ammonia synthesis on promoted iron catalysts,Acta Physicochim\. URSS12 \(1940\) 327–356\.
18. \[18\]H\. Freundlich, Über die Adsorption in Lösungen,Z\. Phys\. Chem\.57 \(1907\) 385–470\.
19. \[19\]C\. E\. Rasmussen, C\. K\. I\. Williams,Gaussian Processes for Machine Learning, MIT Press, Cambridge, MA, 2006\.
20. \[20\]M\. D\. McKay, R\. J\. Beckman, W\. J\. Conover, A comparison of three methods for selecting values of input variables in the analysis of output from a computer code,Technometrics21 \(1979\) 239–245\. DOI:10\.1080/00401706\.1979\.10489755\.
21. \[21\]B\. Efron, Bootstrap methods: another look at the jackknife,Ann\. Statist\.7 \(1979\) 1–26\. DOI:10\.1214/aos/1176344552\.
22. \[22\]D\. P\. Kingma, J\. Ba, Adam: a method for stochastic optimization, in:Int\. Conf\. on Learning Representations \(ICLR\), 2015\. arXiv:1412\.6980\.
23. \[23\]S\. Elfwing, E\. Uchibe, K\. Doya, Sigmoid\-weighted linear units for neural network function approximation in reinforcement learning,Neural Networks107 \(2018\) 3–11\. DOI:10\.1016/j\.neunet\.2017\.12\.012\.
24. \[24\]L\. Lu, X\. Meng, Z\. Mao, G\. E\. Karniadakis, DeepXDE: a deep learning library for solving differential equations,SIAM Review63 \(2021\) 208–228\. DOI:10\.1137/19M1274067\.
25. \[25\]S\. L\. Brunton, J\. L\. Proctor, J\. N\. Kutz, Discovering governing equations from data by sparse identification of nonlinear dynamical systems,Proc\. Natl\. Acad\. Sci\. USA113 \(2016\) 3932–3937\. DOI:10\.1073/pnas\.1517384113\.
26. \[26\]K\. Champion, B\. Lusch, J\. N\. Kutz, S\. L\. Brunton, Data\-driven discovery of coordinates and governing equations,Proc\. Natl\. Acad\. Sci\. USA116 \(2019\) 22445–22451\. DOI:10\.1073/pnas\.1906995116\.
27. \[27\]M\. Raissi, Deep hidden physics models: deep learning of nonlinear partial differential equations,J\. Mach\. Learn\. Res\.19 \(2018\) 1–24\.
28. \[28\]O\. Owoyele, P\. Pal, ChemNODE: a neural ordinary differential equations framework for efficient chemical kinetic solvers,Energy and AI7 \(2022\) 100118\. DOI:10\.1016/j\.egyai\.2021\.100118\.
29. \[29\]C\. Kreutz, A\. Raue, D\. Kaschek, J\. Timmer, Profile likelihood in systems biology,FEBS J\.280 \(2013\) 2564–2571\. DOI:10\.1111/febs\.12276\.
30. \[30\]G\. Pang, L\. Lu, G\. E\. Karniadakis, fPINNs: fractional physics\-informed neural networks,SIAM J\. Sci\. Comput\.41 \(2019\) A2603–A2626\. DOI:10\.1137/18M1229845\.
31. \[31\]S\. Cai, Z\. Mao, Z\. Wang, M\. Yin, G\. E\. Karniadakis, Physics\-informed neural networks \(PINNs\) for fluid mechanics: a review,Acta Mech\. Sin\.37 \(2021\) 1727–1738\. DOI:10\.1007/s10409\-021\-01148\-1\.
32. \[32\]R\. G\. Patel, I\. Manickam, N\. A\. Trask, et al\., Thermodynamically consistent physics\-informed neural networks for hyperbolic systems,J\. Comput\. Phys\.449 \(2022\) 110754\. DOI:10\.1016/j\.jcp\.2021\.110754\.
33. \[33\]E\. Kharazmi, Z\. Zhang, G\. E\. Karniadakis, hp\-VPINNs: variational physics\-informed neural networks with domain decomposition,Comput\. Methods Appl\. Mech\. Engrg\.374 \(2021\) 113547\. DOI:10\.1016/j\.cma\.2020\.113547\.
34. \[34\]M\. Raissi, A\. Yazdani, G\. E\. Karniadakis, Hidden fluid mechanics: learning velocity and pressure fields from flow visualizations,Science367 \(2020\) 1026–1030\. DOI:10\.1126/science\.aaw4741\.
35. \[35\]A\. Yazdani, L\. Lu, M\. Raissi, G\. E\. Karniadakis, Systems biology informed deep learning for inferring parameters and hidden dynamics,PLoS Comput\. Biol\.16 \(2020\) e1007575\. DOI:10\.1371/journal\.pcbi\.1007575\.
36. \[36\]S\. Wang, X\. Yu, P\. Perdikaris, When and why PINNs fail to train: a neural tangent kernel perspective,J\. Comput\. Phys\.449 \(2022\) 110768\. DOI:10\.1016/j\.jcp\.2021\.110768\.
37. \[37\]A\. D\. Jagtap, K\. Kawaguchi, G\. E\. Karniadakis, Adaptive activation functions accelerate convergence in deep and physics\-informed neural networks,J\. Comput\. Phys\.404 \(2020\) 109136\. DOI:10\.1016/j\.jcp\.2019\.109136\.
38. \[38\]F\.\-G\. Wieland, A\. L\. Hauber, M\. Rosenblatt, C\. Tönsing, J\. Timmer, On structural and practical identifiability,Curr\. Opin\. Syst\. Biol\.25 \(2021\) 60–69\. DOI:10\.1016/j\.coisb\.2021\.03\.005\.
39. \[39\]D\. Pan, L\. Ma, Y\. Xie, T\. C\. Jen, C\. Yuan, On the physical and chemical details of alumina atomic layer deposition: a combined experimental and numerical approach,J\. Vac\. Sci\. Technol\. A33 \(2015\) 021511\. DOI:10\.1116/1\.4905726\.
40. \[40\]M\. K\. Transtrum, P\. Qiu, Model reduction by manifold boundaries,Phys\. Rev\. Lett\.113 \(2014\) 098701\. DOI:10\.1103/PhysRevLett\.113\.098701\.
41. \[41\]C\. Masse de la Huerta, V\. H\. Nguyen, J\.\-M\. Dedulle, D\. Bellet, C\. Jiménez, D\. Muñoz\-Rojas, Influence of the geometric parameters on the deposition mode in spatial atomic layer deposition: a novel approach to area\-selective deposition,Coatings9 \(2019\) 5\. DOI:10\.3390/coatings9010005\.
42. \[42\]Z\. Deng, W\. He, C\. Duan, R\. Chen, B\. Shan, Mechanistic modeling study on process optimization and precursor utilization with atmospheric spatial atomic layer deposition,J\. Vac\. Sci\. Technol\. A34 \(2016\) 01A108\. DOI:10\.1116/1\.4932564\.
43. \[43\]D\. Pan, T\. C\. Jen, C\. Yuan, Effects of gap size, temperature and pumping pressure on the fluid dynamics and chemical kinetics of in\-line spatial atomic layer deposition of Al2O3,Int\. J\. Heat Mass Transf\.96 \(2016\) 189–198\. DOI:10\.1016/j\.ijheatmasstransfer\.2016\.01\.034\.Similar Articles
EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks
EvoPINN is an agentic framework that reformulates PINN development as an execution-grounded algorithm discovery problem, using an LLM agent to propose programmatic modifications. It autonomously discovers PDE-specialized learning algorithms, including a novel architecture called SLRC-PINN, which outperforms baselines across diverse PDE regimes.
Adaptive Quantum Physics-Informed Neural Networks for Differential Equations with Applications to Fluid Dynamics
The paper introduces a hybrid quantum-classical framework enhancing Quantum Physics-Informed Neural Networks (QPINNs) with adaptive collocation point sampling and loss-aware attention for solving differential equations, achieving significant accuracy improvements in fluid dynamics and reaction-diffusion benchmarks.
Curriculum Learning of Physics-Informed Neural Networks based on Spatial Correlation
This paper proposes a spatially correlated curriculum learning framework for Physics-Informed Neural Networks (PINNs) that improves training stability and solution accuracy by leveraging spatial correlations among subregions, addressing issues like high-dimensional non-convex loss landscapes and imbalanced multi-objective constraints.
Physics-Informed Neural Networks with Learnable Loss Balancing and Transfer Learning
This paper proposes a self-supervised physics-informed neural network (PINN) framework with a learnable blending neuron to adaptively balance physics-based and data-driven losses, and integrates transfer learning to improve efficiency under data scarcity. It is validated on liquid-metal miniature heat sink CFD data with only 87 datapoints, achieving under 8% error.
Alternating Levenberg-Marquardt Training of Physics-Informed Neural Networks with Fourier-Enhanced Features
This paper proposes FALM-PINN, an alternating Levenberg-Marquardt training framework for physics-informed neural networks that uses Fourier-enhanced features to address spectral bias and representation-coefficient coupling, achieving up to two orders of magnitude lower errors on high-frequency and nonlinear PDEs.