CAHR-Net: Condition-Adaptive Hysteresis Reconstruction for Compact and Interpretable Magnetic Core Loss Modeling

arXiv cs.LG Papers

Summary

CAHR-Net proposes a condition-adaptive hysteresis reconstruction network that improves magnetic core loss modeling by injecting operating conditions into intermediate representations, achieving lower errors with fewer parameters compared to existing methods.

arXiv:2609.01991v1 Announce Type: new Abstract: Magnetic core loss originates in the hysteresis loop: the energy dissipated per excitation cycle equals the loop area, and frequency, temperature, and waveform shape set the loss by reshaping the loop geometry. Most existing models let these conditions act only on a terminal scalar - empirical equations fold them into fitted exponents, and data-driven predictors append them to encoded features - so no intermediate hysteresis representation remains for the conditions to reshape. This paper proposes CAHR-Net, a condition-adaptive hysteresis reconstruction network that injects the operating conditions where they physically act. It preserves the interpretable chain from flux density waveform to magnetic field reconstruction, loop-area integration, and power loss estimation, and uses feature-wise linear modulation to inject frequency, temperature, and waveform statistics into the intermediate reconstruction representation. A matched large-batch training protocol based on AdamW, cosine scheduling, and a staged reconstruction-to-power-loss objective is also reported, because the modulation pathway takes effect only within it. On the MagNet final A-E material protocol, CAHR-Net attains an average p95 relative error of 6.89% with only 1874 parameters, the lowest among all compared methods, together with a lower worst-material p95 than the strongest black-box solution at about 48x fewer parameters; it reduces the average p95 of the physical reconstruction backbone from 7.47% to 6.89% and the p95 of material D, the most difficult material, from 16.40% to 14.87%. Ablation and condition-slice analyses attribute the improvement to the coupling of physical loop reconstruction, structured condition modulation, and the matched optimization trajectory.
Original Article
View Cached Full Text

Cached at: 09/03/26, 06:13 AM

# CAHR-Net: Condition-Adaptive Hysteresis Reconstruction for Compact and Interpretable Magnetic Core Loss Modeling
Source: [https://arxiv.org/html/2609.01991](https://arxiv.org/html/2609.01991)
Chunye Gong and Cong Yao††thanks:Manuscript received Month˜xx, 2026; revised Month˜xx, 2026\. This work was supported in part by the National Key Research and Development Program of China under Grant 2025YFB3003605 and in part by the National Natural Science Foundation of China under Grant U2530207\.*\(†˜Chunye Gong and Cong Yao contributed equally to this work\.\)**\(Corresponding author: Cong Yao\.\)*††thanks:Chunye Gong and Cong Yao are with the College of Computing, National University of Defense Technology, Changsha 410073, China, also with the National Supercomputer Center in Tianjin, Tianjin 300457, China, and also with the Laboratory of Digitizing Software for Frontier Equipment, National University of Defense Technology, Changsha 410073, China \(e\-mail: gongchunye@nudt\.edu\.cn; ycong744@gmail\.com\)\.

###### Abstract

Magnetic core loss originates in the hysteresis loop: the energy dissipated per excitation cycle equals the loop area, and frequency, temperature, and waveform shape set the loss by reshaping the loop geometry\. Most existing models, however, let these operating conditions act only on a terminal scalar—empirical equations fold them into fitted exponents, while data\-driven predictors append them to encoded features—so no intermediate hysteresis representation remains for the conditions to reshape\. This paper proposes a condition\-adaptive hysteresis reconstruction network \(CAHR\-Net\) that injects the operating conditions where they physically act\. The proposed method preserves the physically interpretable chain from flux density waveform to magnetic field reconstruction, hysteresis\-loop area integration, and power loss estimation\. Different from scalar concatenation or terminal correction, CAHR\-Net uses feature\-wise linear modulation to inject frequency, temperature, and waveform statistics into the intermediate reconstruction representation\. A matched large\-batch training protocol based on AdamW, cosine learning\-rate scheduling, and a staged reconstruction\-loss\-to\-power\-loss objective is further reported, because the modulation pathway takes effect only within it\. Experiments on the MagNet final A–E material protocol show that CAHR\-Net attains an averagep​95p95relative error of6\.89%6\.89\\%with only 1874 parameters—the lowest averagep​95p95among all compared methods, together with a lower worst\-materialp​95p95than the strongest black\-box solution at about48×48\\timesfewer parameters—and reduces the averagep​95p95of the physical reconstruction backbone from7\.47%7\.47\\%to6\.89%6\.89\\%and thep​95p95of material D, the most difficult material, from16\.40%16\.40\\%to14\.87%14\.87\\%\. Ablation and condition\-slice analyses indicate that the improvement comes from the coupling of physical loop reconstruction, structured condition modulation, and the matched optimization trajectory\.

###### Index Terms:

Condition modulation, core loss, feature\-wise linear modulation, hysteresis reconstruction, MagNet, power magnetics\.

## IIntroduction

Magnetic core loss is not an abstract scalar response: in each excitation cycle, the energy dissipated in the core equals the area of theBB–HHhysteresis loop, and this loss largely dictates the thermal design, power density, and efficiency of high\-frequency converters\. Operating conditions govern the loss through the geometry of this loop—frequency broadens it, temperature rescales its amplitude and local slope, and the flux\-density waveform sets the trajectory along which it is traversed\. Every core loss model is therefore, implicitly, an answer to the question of where the operating conditions should enter\. The classical Steinmetz equation\[[1](https://arxiv.org/html/2609.01991#bib.bib1)\]answers it by collapsing the loop into a compact power law whose fitted exponents absorb the conditions, and successive refinements have extended this answer to nonsinusoidal excitation by adding parameters to the same scalar law\[[2](https://arxiv.org/html/2609.01991#bib.bib2),[3](https://arxiv.org/html/2609.01991#bib.bib3),[4](https://arxiv.org/html/2609.01991#bib.bib4),[5](https://arxiv.org/html/2609.01991#bib.bib5),[6](https://arxiv.org/html/2609.01991#bib.bib6)\]\. The collapse, however, is lossy: fixed functional forms cannot track how arbitrary waveforms, wide temperature ranges, and material\-specific nonlinearities jointly deform the loop\.

This difficulty is fundamental rather than parametric\. Maxwell’s equations and finite\-element field solvers already describe the linear behavior of conductors and the geometric and thermal aspects of a component with high fidelity; what resists compact modeling is the strongly nonlinear, history\-dependent magnetization that generates the loop itself, compounded by the dispersion that material composition and manufacturing introduce at the component level\. Hysteresis formalisms such as the Preisach model\[[7](https://arxiv.org/html/2609.01991#bib.bib7)\]and the Jiles–Atherton model\[[8](https://arxiv.org/html/2609.01991#bib.bib8)\]keep the loop as the modeling object, but their semi\-empirical parameters are difficult to identify consistently across wide frequency, temperature, and waveform ranges\. Core loss modeling has therefore long faced a dilemma: preserve the loop and struggle to fit it, or collapse the loop and lose the very geometry through which the operating conditions act\.

Public datasets and modeling challenges\[[9](https://arxiv.org/html/2609.01991#bib.bib9),[10](https://arxiv.org/html/2609.01991#bib.bib10),[11](https://arxiv.org/html/2609.01991#bib.bib11),[12](https://arxiv.org/html/2609.01991#bib.bib12)\]have recently made a third route practical: learning the waveform\-to\-loss map directly from large\-scale measurements\. Neural predictors built on waveform encoding, multimodal fusion, and physics\-inspired structures\[[13](https://arxiv.org/html/2609.01991#bib.bib13),[14](https://arxiv.org/html/2609.01991#bib.bib14),[15](https://arxiv.org/html/2609.01991#bib.bib15),[16](https://arxiv.org/html/2609.01991#bib.bib16)\], together with knowledge\-aware networks\[[17](https://arxiv.org/html/2609.01991#bib.bib17)\], history\-dependent hysteresis operators\[[18](https://arxiv.org/html/2609.01991#bib.bib18)\], adversarial generation\[[19](https://arxiv.org/html/2609.01991#bib.bib19)\], and cross\-material transfer\[[20](https://arxiv.org/html/2609.01991#bib.bib20)\], have pushed accuracy well beyond the empirical family\. Yet most of these models achieve accuracy by abandoning the loop altogether, and they inherit a subtler form of the same misplacement: frequency, temperature, and flux\-density statistics are appended to a feature vector or concatenated at the regression head, so the model can observe which condition produced a sample, while no intermediate hysteresis representation remains for the conditions to reshape\.

Gray\-box reconstruction restores the missing object\. The HARDCORE framework\[[21](https://arxiv.org/html/2609.01991#bib.bib21)\]first reconstructs the magnetic field waveform and then integrates theB​HBHloop area, showing that keeping the loop inside the model yields a favorable accuracy–complexity tradeoff under arbitrary waveforms\. Even here, however, the operating conditions still enter as appended scalars: the object is right, but the conditioning pathway is not\. We observe that the leading effects of frequency and temperature on the loop—broadening, tilting, and amplitude scaling—are naturally expressed as a scaling and shifting of the representation from which the loop is generated\. This suggests encoding the operating conditions into feature\-wise scale and shift parameters that act directly on the intermediate reconstruction representation\. CAHR\-Net, the condition\-adaptive hysteresis reconstruction network proposed in this paper, realizes this principle: a dilated temporal\-convolutional encoder\[[22](https://arxiv.org/html/2609.01991#bib.bib22)\]extracts the waveform skeleton, feature\-wise linear modulation \(FiLM\)\[[23](https://arxiv.org/html/2609.01991#bib.bib23)\]lets the conditions reshape the hidden representation before the magnetic field waveform is reconstructed, and a staged reconstruction\-to\-loss objective with matched large\-batch optimization allows the modulation branch to leave the near\-identity regime and become effective\.

The contributions of this paper follow from this single design question—where should the operating conditions enter the model—and are summarized as follows\.

- •*The object: a physically interpretable reconstruction backbone\.*A lightweight architecture keeps the hysteresis loop as the modeling object for arbitrary\-waveform core loss estimation, following the interpretable pathB⁡\(t\)→H^​\(t\)→B​H→P^vB\(t\)\\rightarrow\\hat\{H\}\(t\)\\rightarrow BH\\rightarrow\\hat\{P\}\_\{\\mathrm\{v\}\}rather than direct terminal regression\.
- •*The pathway: condition\-adaptive FiLM modulation\.*Frequency, temperature, and waveform statistics are encoded into feature\-wise scaling and shifting parameters that act directly on the intermediate hysteresis reconstruction representation\. This pathway admits an electromagnetic reading—scaling matches loop broadening and amplitude rescaling, shifting matches bias and remanence drift—that scalar concatenation and terminal correction do not provide\.
- •*The evidence: a favorable accuracy–parameter–interpretability tradeoff\.*Under the unified MagNet final A–E protocol, CAHR\-Net attains an averagep​95p95relative error of6\.89%6\.89\\%with only 1874 parameters, the lowest averagep​95p95among all compared methods—and a lower worst\-materialp​95p95than the strongest black\-box entry at about48×48\\timesfewer parameters—and it reduces the averagep​95p95of the physical reconstruction backbone from7\.47%7\.47\\%to6\.89%6\.89\\%and thep​95p95of material D, the most difficult material, from16\.40%16\.40\\%to14\.87%14\.87\\%\.

The rest of this paper is organized as follows\.[SectionII](https://arxiv.org/html/2609.01991#S2)develops CAHR\-Net, from the problem formulation to the loop reconstruction chain, the condition modulation pathway, and the deployment\-oriented inference optimization\.[SectionIII](https://arxiv.org/html/2609.01991#S3)describes the experimental setup, including the network configuration and training protocol,[SectionIV](https://arxiv.org/html/2609.01991#S4)presents the results and analyses, and[SectionV](https://arxiv.org/html/2609.01991#S5)concludes the paper\.

## IIMethodology

### II\-AProblem Formulation and Method Overview

Given a single\-period flux density waveform

𝐁=\{Bt\}t=1N,N=1024,\\mathbf\{B\}=\\\{B\_\{t\}\\\}\_\{t=1\}^\{N\},\\quad N=1024,\(1\)and a scalar operating vector𝐬\\mathbf\{s\}containing frequency, temperature, and waveform\-level statistics, the goal is to predict the volumetric core lossPvP\_\{\\mathrm\{v\}\}\. A direct map

P^v=fθ​\(𝐁,𝐬\)\\hat\{P\}\_\{\\mathrm\{v\}\}=f\_\{\\theta\}\(\\mathbf\{B\},\\mathbf\{s\}\)\(2\)treats the loss as an isolated scalar output and leaves no intermediate representation for the operating conditions to act on\. CAHR\-Net instead uses a physically structured intermediate target that keeps such a pathway open: it reconstructs the magnetic field waveformH^​\(t\)\\hat\{H\}\(t\), computes a loop\-area\-based initial estimate, and then applies a lightweight residual correction in the logarithmic loss domain\.

[Fig\.1](https://arxiv.org/html/2609.01991#S2.F1)shows the resulting architecture\. CAHR\-Net contains a temporal convolutional branch for waveform encoding, a scalar multilayer perceptron for condition encoding, a FiLM block\[[23](https://arxiv.org/html/2609.01991#bib.bib23)\]that couples the two, a magnetic\-field reconstruction head withB​HBHarea integration, and a residual power\-loss correction head\.

![Refer to caption](https://arxiv.org/html/2609.01991v1/figures/pic-best.png)Fig\. 1:Overall architecture of CAHR\-Net\. The operating vector𝐬\\mathbf\{s\}is encoded into the FiLM pair\[γ,β\]\[\\gamma,\\beta\], which rescales \(⊙\\odot, by1\+α​γ1\+\\alpha\\gamma\) and then shifts \(\+\+, byβ\\beta\) eleven of the twelve encoder channels before the magnetic\-field reconstruction head; one raw channel bypasses modulation\. The reconstructed fieldH^​\(t\)\\hat\{H\}\(t\)drives theB​HBH\-area integration, and a residual multilayer perceptron corrects the loop\-area estimate in the logarithmic loss domain to produceP^v\\hat\{P\}\_\{\\mathrm\{v\}\}\.
### II\-BPhysical Loop Reconstruction Chain

The temporal encoder first extracts a waveform skeleton from the normalized 1024\-point flux density sequence,

𝐅0=ℰB​\(B⁡\(t\)\),\\mathbf\{F\}\_\{0\}=\\mathcal\{E\}\_\{B\}\\\!\\left\(B\(t\)\\right\),\(3\)whereℰB​\(⋅\)\\mathcal\{E\}\_\{B\}\(\\cdot\)denotes a stack of dilated temporal convolution blocks\[[22](https://arxiv.org/html/2609.01991#bib.bib22)\]\. The scalar branch maps the operating vector to modulation parameters\. After modulation, the reconstruction head generates

H^​\(t\)=𝒢⁡\(B⁡\(t\),𝐬,γ,β\)\.\\hat\{H\}\(t\)=\\mathcal\{G\}\\\!\\left\(B\(t\),\\mathbf\{s\};\\gamma,\\beta\\right\)\.\(4\)
The reconstructed field waveform is converted to an initial power\-loss estimate through discrete loop\-area integration,

P^B​H=f​blim​hlim​trapz⁡\(H^,B\),\\hat\{P\}\_\{BH\}=f\\,b\_\{\\lim\}h\_\{\\lim\}\\operatorname\{trapz\}\\\!\\left\(\\hat\{H\},B\\right\),\(5\)whereffis the excitation frequency,blimb\_\{\\lim\}andhlimh\_\{\\lim\}are the scale factors used to restore physical magnitudes, andtrapz⁡\(⋅\)\\operatorname\{trapz\}\(\\cdot\)denotes trapezoidal integration\.

The final prediction is obtained in the logarithmic domain:

y^P=log⁡\(P^B​H\)\+Δ⁡\(\[𝐬,log⁡\(P^B​H\)\]\),\\hat\{y\}\_\{P\}=\\log\\\!\\left\(\\hat\{P\}\_\{BH\}\\right\)\+\\Delta\\\!\\left\(\[\\mathbf\{s\},\\log\(\\hat\{P\}\_\{BH\}\)\]\\right\),\(6\)P^v=exp⁡\(y^P\)\.\\hat\{P\}\_\{\\mathrm\{v\}\}=\\exp\(\\hat\{y\}\_\{P\}\)\.\(7\)In implementation, the residual head uses a standardized form oflog⁡\(P^B​H\)\\log\(\\hat\{P\}\_\{BH\}\)for numerical stability\.[Equations5](https://arxiv.org/html/2609.01991#S2.E5),[6](https://arxiv.org/html/2609.01991#S2.E6)and[7](https://arxiv.org/html/2609.01991#S2.E7)keep the prediction chain anchored to the physical relation between the hysteresis loop and core loss while allowing a small learned correction for scale and integration bias\.

Both the reconstruction target and the scale factors in[Equation5](https://arxiv.org/html/2609.01991#S2.E5)are anchored in measured data\. For every record, the MagNet database provides the measured magnetic field waveformH⁡\(t\)H\(t\), obtained from the sensed excitation current through Ampère’s law, alongside the flux density waveform obtained from the induced voltage\[[9](https://arxiv.org/html/2609.01991#bib.bib9)\]\. The reconstruction loss defined in[SectionIII\-D](https://arxiv.org/html/2609.01991#S3.SS4)therefore supervisesH^​\(t\)\\hat\{H\}\(t\)with measured waveforms rather than with model\-generated references, following the same protocol as HARDCORE\[[21](https://arxiv.org/html/2609.01991#bib.bib21)\]\. The factorblimb\_\{\\lim\}is the maximum absolute flux density over the training records of the material, andhlimh\_\{\\lim\}is the corresponding field normalization bound, capped at150​A/m150~\\mathrm\{A/m\}and rescaled per record by the record’s relative flux amplitude, so that[Equation5](https://arxiv.org/html/2609.01991#S2.E5)restores physical magnitudes from normalized waveforms\.

### II\-CCondition\-Adaptive FiLM Modulation

The key difference between CAHR\-Net and scalar\-concatenation models is the location and form of condition injection\. Letℰs​\(⋅\)\\mathcal\{E\}\_\{s\}\(\\cdot\)be the scalar encoder\. The modulation parameters are generated by

\[γ,β\]=Ψ⁡\(ℰs​\(𝐬\)\),\[\\gamma,\\beta\]=\\Psi\\\!\\left\(\\mathcal\{E\}\_\{s\}\(\\mathbf\{s\}\)\\right\),\(8\)whereΨ⁡\(⋅\)\\Psi\(\\cdot\)projects the condition embedding into the feature modulation space\. For an intermediate representation𝐅\\mathbf\{F\}, CAHR\-Net applies

𝐅~=\(1\+α​γ\)⊙𝐅\+β,\\widetilde\{\\mathbf\{F\}\}=\(1\+\\alpha\\gamma\)\\odot\\mathbf\{F\}\+\\beta,\(9\)where⊙\\odotis channel\-wise multiplication andα\\alphacontrols modulation strength\. The default setting usesα=0\.1\\alpha=0\.1\.

The form of[Equation9](https://arxiv.org/html/2609.01991#S2.E9)mirrors the two dominant ways in which operating conditions deform the measured hysteresis loop\. As frequency rises, eddy\-current and excess contributions widen the loop and rescale its local slope, which is a multiplicative deformation of the trajectory amplitude; temperature shifts the permeability and coercivity operating point, which acts closer to a translation of the loop\. The FiLM pair reproduces these two degrees of freedom on the hidden representation:γ\\gammaapplies channel\-wise rescaling, matching loop broadening and amplitude scaling, whileβ\\betaapplies channel\-wise shifting, matching bias and remanence drift\. The same waveform pattern can therefore be amplified, suppressed, or translated under different operating conditions beforeH^​\(t\)\\hat\{H\}\(t\)is reconstructed\. Additive bias injection, as used in the backbone, covers only the translational degree of freedom, and input\-level concatenation provides neither in a structured form\. We do not claim a one\-to\-one identification between individual FiLM parameters and specific loop features; the correspondence holds at the level of transformation families, and it is consistent with the ablation in[Fig\.6](https://arxiv.org/html/2609.01991#S4.F6)\(a\), where replacing FiLM with bias\-only injection removes most of the gain under the matched optimization protocol\.

### II\-DDeployment\-Oriented Inference Optimization

The 1874\-parameter budget suggests that CAHR\-Net is inexpensive to deploy, but parameter count alone does not determine the latency realized in practice\. In the intended deployment scenarios—loss evaluation inside iterative converter design loops, large design\-space sweeps, and online estimation on general\-purpose processors without an accelerator—the model is queried one operating point at a time on a CPU, and for a network of this size the per\-query cost is dominated by framework dispatch rather than by arithmetic\. CAHR\-Net therefore admits an accuracy\-neutral inference optimization that lowers this per\-query cost by changing only the execution backend, without modifying weights, architecture, or numerical precision\.

The optimization exports the trained model to the ONNX format and serves it with ONNX Runtime 1\.18\.1\[[24](https://arxiv.org/html/2609.01991#bib.bib24)\]\. The export is not a push\-button step: the loop\-area integration of[Equation5](https://arxiv.org/html/2609.01991#S2.E5)relies on a trapezoidal\-integration operator that has no ONNX equivalent\. A thin export wrapper therefore rewrites the integration as an explicit trapezoidal sum built from elementary slicing, addition, multiplication, and reduction operations, and bakes the per\-material normalization constants into the graph, so that the exported graph is self\-contained\. Two numerical gates guard the export\. First, the wrapper is verified to be bit\-identical to the original TorchScript model, with a maximum absolute output difference of exactly zero\. Second, the exported model under ONNX Runtime is compared against the original over all 7651 material\-A test samples: the maximum deviation is\|Δ​log⁡P^v\|<8×10−6\|\\Delta\\log\\hat\{P\}\_\{\\mathrm\{v\}\}\|<8\\times 10^\{\-6\}, and the test\-setp​95p95relative error remains5\.86%5\.86\\%, unchanged to two decimals\. The optimization is thus accuracy\-neutral both by construction and by measurement; its latency benefit is quantified in[SectionIV\-F](https://arxiv.org/html/2609.01991#S4.SS6)\.

## IIIExperimental Setup

### III\-ADataset, Protocol, and Metrics

Experiments use the MagNet final\-stage A–E material protocol\[[9](https://arxiv.org/html/2609.01991#bib.bib9),[12](https://arxiv.org/html/2609.01991#bib.bib12)\]\. Each material is trained and tested independently on its corresponding final split\. The input waveform contains 1024 samples per period, and the scalar vector contains frequency, temperature, and waveform statistics as specified in[SectionIII\-C](https://arxiv.org/html/2609.01991#S3.SS3)\.

Prediction quality is measured by the relative error

Rel\.Err\.=\|P^v−PvPv\|×100%,\\mathrm\{Rel\.\\ Err\.\}=\\left\|\\frac\{\\hat\{P\}\_\{\\mathrm\{v\}\}\-P\_\{\\mathrm\{v\}\}\}\{P\_\{\\mathrm\{v\}\}\}\\right\|\\times 100\\%,\(10\)and reported as the average relative error, the averagep​95p95andp​99p99percentiles across the five materials, and the worst\-materialp​95p95\[[25](https://arxiv.org/html/2609.01991#bib.bib25)\]\. In all experiments reported below, the worst\-materialp​95p95occurs on material D, so “worstp​95p95” and “material\-Dp​95p95” denote the same quantity\. Thep​95p95\-oriented view is adopted because thermal design is usually more sensitive to high\-risk tail errors than to mean behavior; an averagep​95p95relative error below10%10\\%is generally regarded as strong under the MagNet evaluation protocol\[[12](https://arxiv.org/html/2609.01991#bib.bib12)\]\. All reported results follow the single\-model reporting convention of the MagNet Challenge and HARDCORE\[[12](https://arxiv.org/html/2609.01991#bib.bib12),[21](https://arxiv.org/html/2609.01991#bib.bib21)\], using one trained model per method under the released protocol\.

### III\-BCompared Methods

Three groups of baselines are considered\. First, empirical models, including SE, MSE, GSE, and iGSE, are re\-fitted under the same A–E protocol\. SE is fitted by logarithmic least squares, whereas MSE, GSE, and iGSE use bounded nonlinear least squares with five multi\-start restarts\. Second, representative public challenge methods are summarized from the MagNet Challenge report\[[12](https://arxiv.org/html/2609.01991#bib.bib12)\]\. Third, internal models are evaluated under the same experimental chain, including the HARDCORE physical reconstruction backbone\[[21](https://arxiv.org/html/2609.01991#bib.bib21)\], the same backbone trained with AdamW and cosine scheduling \(HARDCORE \+ AWC\), SE\-TCN, FiLM\-TCN, CAHR\-Net, and width or combination variants\.

### III\-CNetwork Configuration

[TableI](https://arxiv.org/html/2609.01991#S3.T1)lists the complete layer\-by\-layer configuration of CAHR\-Net; the parameter subtotals add up to the 1874 parameters reported throughout the paper\. The waveform encoder receives five input channels—the globally normalized flux density waveform, a per\-record normalized copy, its first and second time derivatives, and a saturation\-emphasizing tangent transform\[[21](https://arxiv.org/html/2609.01991#bib.bib21)\]\. The operating vector𝐬∈ℝ11\\mathbf\{s\}\\in\\mathbb\{R\}^\{11\}collects the logarithmic frequency, the temperature, four waveform\-shape indicator variables, the peak\-to\-peak flux density and the mean absolute flux slew rate together with their logarithms, and the fundamental period\. All convolutions use circular padding, consistent with the periodic single\-period input, and the three temporal blocks use kernel size 9 with dilation rates 4, 8, and 16\. FiLM acts at the interface between the waveform encoder and the reconstruction head: eleven of the twelve encoder channels are modulated byγ,β∈ℝ11\\gamma,\\beta\\in\\mathbb\{R\}^\{11\}, while the first channel passes through unmodulated so that a raw waveform pathway is always preserved\.

TABLE I:Layer\-by\-Layer Configuration of CAHR\-NetStageConfigurationParamsWaveform encoderℰB\\mathcal\{E\}\_\{B\}Conv1d5→125\{\\rightarrow\}12,k=9k\{=\}9,d=4d\{=\}4, tanh552Condition encoderℰs,Ψ\\mathcal\{E\}\_\{s\},\\PsiLinear11→2211\{\\rightarrow\}22, tanh,\[γ,β\]\[\\gamma,\\beta\]264FiLM modulation[Equation9](https://arxiv.org/html/2609.01991#S2.E9), 11 of 12 channels0Reconstruction head𝒢\\mathcal\{G\}Conv1d12→812\{\\rightarrow\}8,k=9k\{=\}9,d=8d\{=\}8, tanh872Conv1d8→18\{\\rightarrow\}1,k=9k\{=\}9,d=16d\{=\}1673B​HBHintegrationTrapezoidal,[Equation5](https://arxiv.org/html/2609.01991#S2.E5)0Residual headΔ\\DeltaLinear12→812\{\\rightarrow\}8, tanh; Linear8→18\{\\rightarrow\}1113Total1874CAHR\-Net is fully specified by[TableI](https://arxiv.org/html/2609.01991#S3.T1)and the training protocol below; HARDCORE is included in[TableIII](https://arxiv.org/html/2609.01991#S4.T3)only as a reconstruction baseline under the same A–E evaluation protocol\.

### III\-DTraining Protocol

Introducing FiLM alone does not guarantee a stable gain under large\-batch training\. CAHR\-Net therefore uses a matched optimization protocol\. The magnetic\-field reconstruction loss and logarithmic power\-loss loss are

ℒH=MSE⁡\(H^,H\),ℒP=MSE⁡\(y^P,log⁡Pv\)\.\\mathcal\{L\}\_\{H\}=\\mathrm\{MSE\}\\\!\\left\(\\hat\{H\},H\\right\),\\quad\\mathcal\{L\}\_\{P\}=\\mathrm\{MSE\}\\\!\\left\(\\hat\{y\}\_\{P\},\\log P\_\{\\mathrm\{v\}\}\\right\)\.\(11\)At epocheeofEEtotal epochs, the total loss is

ℒ⁡\(e\)=\(1−eE\)​ℒH\+eE​ℒP\.\\mathcal\{L\}\(e\)=\\left\(1\-\\frac\{e\}\{E\}\\right\)\\mathcal\{L\}\_\{H\}\+\\frac\{e\}\{E\}\\mathcal\{L\}\_\{P\}\.\(12\)This schedule prioritizes stable loop reconstruction at the beginning and gradually shifts the emphasis toward loss prediction\. The ordering matters for interpretability: if the power\-loss objective dominated from the start, the residual head could compensate for a physically meaningless loop, and the reconstruction chain would lose its electromagnetic reading\.

The optimizer is AdamW\[[26](https://arxiv.org/html/2609.01991#bib.bib26)\]with cosine learning\-rate scheduling\[[27](https://arxiv.org/html/2609.01991#bib.bib27)\], learning rate2×10−32\\times 10^\{\-3\}, weight decay10−410^\{\-4\}, batch size 512, and 10000 epochs; the modulation strengthα=0\.1\\alpha=0\.1in[Equation9](https://arxiv.org/html/2609.01991#S2.E9)keeps the FiLM branch close to identity at initialization\. These settings are stated in full because of an empirical finding documented in[Fig\.6](https://arxiv.org/html/2609.01991#S4.F6)\(a\): under the original NAdam optimizer with step decay, FiLM behaves close to bias injection, whereas AdamW with cosine scheduling and the enlarged learning rate allows the modulation branch to leave the near\-identity regime and take effect\. We regard this coupling between the conditioning structure and the optimization trajectory as an empirical finding rather than a standalone contribution\.[Fig\.2](https://arxiv.org/html/2609.01991#S3.F2)summarizes the resulting training\-to\-inference protocol\.

Per\-material train setA–E, independentCAHR\-Netforward passB,𝐬→H^→P^vB,\\mathbf\{s\}\\\!\\to\\\!\\hat\{H\}\\\!\\to\\\!\\hat\{P\}\_\{\\mathrm\{v\}\}\(Fig\.[1](https://arxiv.org/html/2609.01991#S2.F1)\)Staged lossℒ⁡\(e\)\\mathcal\{L\}\(e\)AdamW\+\+cosineH^,P^v\\hat\{H\},\\hat\{P\}\_\{\\mathrm\{v\}\}∇θℒ\\nabla\_\{\\theta\}\\mathcal\{L\}updateθ\\thetaTraining phaseTest inputB⁡\(t\),𝐬B\(t\),\\mathbf\{s\}held\-out, per materialCAHR\-Net\(trained\)H^​\(t\),P^v\\hat\{H\}\(t\),\\ \\hat\{P\}\_\{\\mathrm\{v\}\}loop & core lossRelative errorAvg /p95p\_\{95\}/p99p\_\{99\}deployθ⋆\\theta^\{\\star\}Inference & evaluation phase

Fig\. 2:Training\-to\-inference protocol of CAHR\-Net\. Each material A–E is trained independently\. During training, the forward pass of[Fig\.1](https://arxiv.org/html/2609.01991#S2.F1)producesH^\\hat\{H\}andP^v\\hat\{P\}\_\{\\mathrm\{v\}\}, which are scored by the staged lossℒ⁡\(e\)\\mathcal\{L\}\(e\)of[Equation12](https://arxiv.org/html/2609.01991#S3.E12)and optimized with AdamW and cosine scheduling; the curriculum schedule shifts the emphasis from the magnetic\-field reconstruction lossℒH\\mathcal\{L\}\_\{H\}early in training to the logarithmic core\-loss objectiveℒP\\mathcal\{L\}\_\{P\}late in training\. The trained parametersθ⋆\\theta^\{\\star\}are then deployed on held\-out inputs, and the average,p95p\_\{95\}, andp99p\_\{99\}relative\-error metrics are computed per material\.

## IVResults and Discussion

### IV\-AComparison: Empirical Equations, Challenge Entries, and the Accuracy–Parameter Frontier

[TableII](https://arxiv.org/html/2609.01991#S4.T2)consolidates the external comparison into three views\.*Empirical equations\.*Even when SE, MSE, GSE, and iGSE are fitted independently for each material, they remain in the56\.84%56\.84\\%–78\.30%78\.30\\%average\-p​95p95band; their limitation is therefore structural rather than a matter of parameter fitting, because a terminal scalar law cannot represent arbitrary\-waveform, wide\-condition effects\.*MagNet Challenge entries\.*The public field spans a wide range, from the strongest black\-box solution, Bristol, at7\.78%7\.78\\%averagep​95p95with90,65390\{,\}653parameters, to gray\-box entries that stay above13%13\\%; these numbers were obtained under different implementations and possibly different evaluation protocols, so they serve as an external reference rather than a controlled ranking\.*Accuracy–parameter frontier\.*Isolating the four most competitive operating points shows where the proposed method leads: Bristol, which reports the best accuracy yet placed third officially; Fuzhou, the official runner\-up; HARDCORE, the method of the official winner Paderborn, evaluated here as the reconstruction backbone; and CAHR\-Net\. Note that the official final ranking weighted model size jointly with accuracy\[[12](https://arxiv.org/html/2609.01991#bib.bib12)\], which is why the 8914\-parameter Fuzhou entry outranked Bristol despite Bristol’s lower error\. CAHR\-Net attains the lowest averagep​95p95in the whole table,6\.89%6\.89\\%against Bristol’s7\.78%7\.78\\%and Fuzhou’s7\.94%7\.94\\%, together with the lowest worst\-materialp​95p95,14\.87%14\.87\\%against Bristol’s15\.90%15\.90\\%, yet reaches this with only 1874 parameters—about1/481/48of the Bristol budget and1/4\.81/4\.8of Fuzhou—while the gray\-box models near its size stay above13%13\\%averagep​95p95\. CAHR\-Net thus occupies a region of the accuracy–parameter plane that neither the black\-box nor the existing gray\-box group covers, while keeping the prediction chain physically readable\.[Fig\.3](https://arxiv.org/html/2609.01991#S4.F3)visualizes this relationship\.

TABLE II:Unified Comparison: Empirical Equations, MagNet Challenge Entries, and the Accuracy–Parameter FrontierMethodCore Expression / CategoryParamsAvgp​95p95\(%\)Worstp​95p95\(%\)\(a\) Empirical equations \(per\-material fit, unified A–E protocol\)SEk​fα​Bpkβkf^\{\\alpha\}B\_\{\\mathrm\{pk\}\}^\{\\beta\}—78\.3085\.84MSEk​fα​Bpkβ​feqα−1kf^\{\\alpha\}B\_\{\\mathrm\{pk\}\}^\{\\beta\}f\_\{\\mathrm\{eq\}\}^\{\\alpha\-1\}—56\.8462\.31GSEk​fα​Δ​Bβ−α​∫\|B˙\|αkf^\{\\alpha\}\\Delta B^\{\\beta\-\\alpha\}\\int\|\\dot\{B\}\|^\{\\alpha\}—57\.6664\.52iGSEk​fα​∑iΔ​Biβ​Δ​ti1−αkf^\{\\alpha\}\\sum\_\{i\}\\Delta B\_\{i\}^\{\\beta\}\\Delta t\_\{i\}^\{1\-\\alpha\}—61\.2372\.42\(b\) MagNet Challenge entries \(as reported\[[12](https://arxiv.org/html/2609.01991#bib.bib12)\]\)BristolBlack\-box906537\.7815\.90FuzhouBlack\-box89147\.9420\.70TsinghuaBlack\-box11606116\.8829\.90XJTUBlack\-box1734214\.2030\.00MMINN\[[15](https://arxiv.org/html/2609.01991#bib.bib15)\]Gray\-box108413\.8630\.70PI\-MFF\-CN\[[13](https://arxiv.org/html/2609.01991#bib.bib13)\]Gray\-box13993830\.5679\.10\(c\) Accuracy–parameter frontierBristol \(best reported accuracy\)Black\-box90653 \(48×48\\times\)7\.78 \(\+12\.9%\)15\.90 \(\+6\.9%\)Fuzhou \(official 2nd place\)Black\-box8914 \(4\.8×4\.8\\times\)7\.94 \(\+15\.2%\)20\.70 \(\+39\.2%\)HARDCORE \(official 1st place; backbone\)Gray\-box1742 \(0\.93×0\.93\\times\)7\.47 \(\+8\.4%\)16\.40 \(\+10\.3%\)CAHR\-NetGray\-box18746\.8914\.87

Blocks \(a\) and \(c, backbone/CAHR\-Net rows\) are evaluated under the unified protocol of this work \([SectionIII](https://arxiv.org/html/2609.01991#S3)\); block \(b\) and the Bristol/Fuzhou rows of \(c\) are as reported in the original sources\[[12](https://arxiv.org/html/2609.01991#bib.bib12)\], so the cross\-group comparison is an external reference rather than a controlled ranking\. Official placements refer to the challenge’s final ranking, which jointly weighted accuracy and model size\. Parenthesized values in block \(c\) are relative to CAHR\-Net: the parameter count as a multiple of 1874, and the twop​95p95columns as the relative increase over6\.89%6\.89\\%and14\.87%14\.87\\%, respectively\. The Bristol entry used 90653\-parameter models for materials A–C and 16449\-parameter transfer\-learned models for D–E, so its worst\-materialp​95p95\(15\.90, material D\) was obtained with the 16449\-parameter model\. Empirical models are analytic per\-material fits, whose coefficient count is not comparable to a network parameter budget \(“—”\)\.

![Refer to caption](https://arxiv.org/html/2609.01991v1/figures/ch6_cahr_public_scatter.png)Fig\. 3:Parameter\-count and average\-p​95p95relation between public representative methods and CAHR\-Net\.
### IV\-BAblation Study

[TableIII](https://arxiv.org/html/2609.01991#S4.T3)isolates the contribution of each design choice under a shared data pipeline and evaluation script, with rows marked AWC further sharing the optimization protocol of[SectionIII\-D](https://arxiv.org/html/2609.01991#S3.SS4), moving from the HARDCORE backbone toward CAHR\-Net one axis at a time\. Three axes are examined\.*Optimization protocol\.*Applying the matched large\-batch protocol \(AWC\) to the unchanged backbone lowers the averagep​95p95from7\.47%7\.47\\%to7\.27%7\.27\\%and the worst\-materialp​95p95from16\.40%16\.40\\%to15\.64%15\.64\\%, so a better\-conditioned training region already helps before any architectural change\.*Condition\-injection module\.*Compared under the same protocol, replacing the injection with SE or a plain FiLM branch reaches7\.14%7\.14\\%and7\.85%7\.85\\%averagep​95p95, whereas the condition\-adaptive injection of CAHR\-Net attains6\.89%6\.89\\%together with the lowest averagep​99p99of11\.82%11\.82\\%and the lowest worst\-materialp​95p95of14\.87%14\.87\\%; the gap to FiLM\-TCN trained at the original learning rate is consistent with the finding of[SectionIII\-D](https://arxiv.org/html/2609.01991#S3.SS4)that the modulation pathway requires the matched optimization region to become fully effective\.*Capacity\.*Widening the network through CAHR\-W12, SE\+CAHR\-W12, and SE\+CAHR\-W16 raises the parameter count from18741874to as many as30573057yet degrades the averagep​95p95to the7\.69%7\.69\\%–8\.47%8\.47\\%range and pushes the worst\-materialp​95p95above20%20\\%, indicating that the improvement stems from how condition information is injected and trained rather than from added capacity\. CAHR\-Net therefore gives the best accuracy of the chain at a near\-minimal parameter budget\.

TABLE III:Ablation Study Under the Unified Experimental Chain\.MethodParametersAvg\. Rel\. Err\. \(%\)Avg\.p​95p95\(%\)Avg\.p​99p99\(%\)Worstp​95p95\(%\)Rel\.Δ​p​95\\Delta p95vs\. CAHR\-Net \(%\)HARDCORE17422\.637\.4712\.5316\.40\+8\.4\+8\.4HARDCORE \+ AWC17422\.627\.2711\.9515\.64\+5\.5\+5\.5SE\-TCN \+ AWC \+2×10−32\\times 10^\{\-3\}18752\.487\.1411\.8316\.23\+3\.6\+3\.6FiLM\-TCN \+ AWC \+10−310^\{\-3\}18742\.707\.8513\.1917\.61\+13\.9\+13\.9CAHR\-Net18742\.426\.8911\.8214\.870\.0CAHR\-W1223462\.507\.7413\.2420\.68\+12\.3\+12\.3SE \+ CAHR\-W1225242\.467\.6912\.9420\.04\+11\.6\+11\.6SE \+ CAHR\-W1630572\.708\.4715\.8624\.41\+22\.9\+22\.9

### IV\-CMaterial\-Level Tail Error

[Fig\.4](https://arxiv.org/html/2609.01991#S4.F4)\(a\) shows that the main gain is concentrated on material D, the most difficult material\. CAHR\-Net lowers the material\-Dp​95p95from16\.40%16\.40\\%to14\.87%14\.87\\%, while materials A, B, C, and E do not show obvious degradation\. This behavior is desirable because it improves the high\-risk tail without transferring error to easier materials\.

![Refer to caption](https://arxiv.org/html/2609.01991v1/figures/ch6_cahr_material_breakdown.png)

![Refer to caption](https://arxiv.org/html/2609.01991v1/figures/ch6_cahr_material_d_cdf.png)

Fig\. 4:Material\-level tail analysis\. \(a\) Material\-levelp​95p95comparison between HARDCORE and CAHR\-Net\. \(b\) Empirical cumulative distribution of relative error on material D; the inset magnifies the tail region around the 95th percentile, where CAHR\-Net shifts thep​95p95from16\.40%16\.40\\%to14\.87%14\.87\\%\.[Fig\.4](https://arxiv.org/html/2609.01991#S4.F4)\(b\) further verifies that the error distribution on material D shifts left and contracts in the high\-error region\. Therefore, the proposed method does not merely reduce the mean error; it directly compresses samples that would otherwise dominate design risk\.

### IV\-DCondition\-Slice Analysis

To locate where the improvement appears, sample errors are pooled across the five materials and stratified by temperature, frequency, peak flux density, and true loss magnitude under shared quartile boundaries\.[Fig\.5](https://arxiv.org/html/2609.01991#S4.F5)summarizes these four condition\-slice views\. The gains are broad but modest: CAHR\-Net lowers thep​95p95in most slices by a few tenths to about two percentage points, and holds the temperature\-slicep​95p95within7\.7%7\.7\\%–10\.7%10\.7\\%versus8\.5%8\.5\\%–10\.8%10\.8\\%for the backbone\. The clearest improvements fall on the lowest peak\-flux\-density quartile, from10\.91%10\.91\\%to8\.53%8\.53\\%, and on the highest\-error, lowest\-loss\-magnitude quartile, from12\.62%12\.62\\%to11\.00%11\.00\\%, where relative error is intrinsically harder to control; a small number of slices, such as the second flux\-density quartile, are essentially unchanged\.

![Refer to caption](https://arxiv.org/html/2609.01991v1/figures/ch6_cahr_temp_slices.png)

![Refer to caption](https://arxiv.org/html/2609.01991v1/figures/ch6_cahr_freq_slices.png)

![Refer to caption](https://arxiv.org/html/2609.01991v1/figures/ch6_cahr_bpk_slices.png)

![Refer to caption](https://arxiv.org/html/2609.01991v1/figures/ch6_cahr_loss_slices.png)

Fig\. 5:Condition\-slice analysis ofp​95p95relative error, shown from top to bottom for temperature, frequency, peak flux densityBpkB\_\{\\mathrm\{pk\}\}, and true loss magnitude quartiles\.
### IV\-EInjection–Optimization Interaction and Parameter Efficiency

[Fig\.6](https://arxiv.org/html/2609.01991#S4.F6)\(a\) presents a two\-dimensional sweep over condition injection and optimization strategy\. Under NAdam with step decay, bias, SE, and FiLM injection are close\. After switching to AdamW with cosine scheduling, FiLM starts to show a stable advantage\. Increasing the learning rate to2×10−32\\times 10^\{\-3\}moves FiLM into the best region, corresponding to the CAHR\-Net configuration reported in[TableIII](https://arxiv.org/html/2609.01991#S4.T3)\.

![Refer to caption](https://arxiv.org/html/2609.01991v1/figures/ch6_cahr_ablation_heatmap.png)

![Refer to caption](https://arxiv.org/html/2609.01991v1/figures/ch6_cahr_pareto.png)

Fig\. 6:Injection–optimization interaction and parameter efficiency\. \(a\) Interaction between condition injection strategy and optimization protocol\. \(b\) Parameter\-count and average\-p​95p95Pareto distribution of internal model variants\.[Fig\.6](https://arxiv.org/html/2609.01991#S4.F6)\(b\) shows that widening the model or adding SE\-style recalibration does not reliably outperform CAHR\-Net under a similar parameter budget\. The dominant improvement therefore comes from how condition information enters the reconstruction process and how this structure is trained, rather than from a simple increase in capacity\.

### IV\-FInference Latency Optimization

We now quantify the latency benefit of the accuracy\-neutral deployment optimization of[SectionII\-D](https://arxiv.org/html/2609.01991#S2.SS4)\.[TableIV](https://arxiv.org/html/2609.01991#S4.T4)reports the per\-sample latency of four execution backends on a single thread of a server\-class Xeon Platinum 8360Y CPU, using the material\-A test set and the trained checkpoint; each entry is the median over repeated timed runs, with 300 repetitions at batch size 1\. Three observations follow\. First, at batch size 1—the regime that matters for per\-operating\-point queries—switching from TorchScript to ONNX Runtime reduces the latency from0\.820\.82ms to0\.190\.19ms per sample, a4\.4×4\.4\\timesspeedup, and5\.6×5\.6\\timesover eager execution, at zero accuracy cost\. Second, the table itself diagnoses where the gain comes from: at batch size 256 all four backends converge to approximately150​μ150\\,\\mus per sample, which is the arithmetic floor of the model on this core; consequently, about82%82\\%of the batch\-1 TorchScript time is per\-invocation framework overhead rather than computation, and the runtime switch removes precisely this overhead instead of approximating the model\. Third, at batch size 2048 the PyTorch\-based backends degrade to roughly370​μ370\\,\\mus per sample whereas ONNX Runtime remains at174​μ174\\,\\mus, indicating better operator fusion and memory locality at large batches\.

To assess whether these gains are tied to the proposed architecture, the identical export\-and\-verification procedure is applied to the bias\-injection backbone HARDCORE, and[TableV](https://arxiv.org/html/2609.01991#S4.T5)contrasts the two models before and after the backend switch\. The backbone exhibits the same4\.4×4\.4\\timesreduction in single\-sample latency, from798798to182​μ182\\,\\mus, which indicates that the optimization exploits a property shared by this class of compact loop\-reconstruction models—their inference time is dominated by framework dispatch rather than by computation—and hence generalizes beyond CAHR\-Net itself\.[TableV](https://arxiv.org/html/2609.01991#S4.T5)also shows that the latency of the two models remains within5%5\\%of each other at every batch size under the optimized backend—186186versus182​μ182\\,\\mus at batch size 1—an absolute difference of a few microseconds per sample\. Condition\-adaptive modulation thus incurs no practically relevant serving cost: the accuracy improvements of[TableIII](https://arxiv.org/html/2609.01991#S4.T3)are obtained at essentially the latency of the unmodulated backbone\.

TABLE IV:Single\-Thread CPU Inference Latency of CAHR\-Net Across Execution Backends \(μ\\mus per Sample, Material\-A Test Set\)Batch SizePyTorch EagerTorchScriptTorchScript \(Opt\.\)ONNX Runtime11051822788186162001851691492561611541521522048366372372174

TABLE V:Model\-Level Latency Before and After Backend Optimization \(μ\\mus per Sample, Single CPU Thread\)Batch SizeHARDCORE \(TorchScript\)CAHR\-Net Before Opt\. \(TorchScript\)CAHR\-Net After Opt\. \(ONNX Runtime\)1798822186161671851492561471541522048354372174

From a deployment perspective,0\.190\.19ms per sample on one CPU thread corresponds to more than50005000loss evaluations per second per core, which is sufficient to embed the model directly in interactive design iteration or in the inner loop of a circuit\-level optimization sweep over thousands of operating points, with no GPU in the serving path\. The exported ONNX graph, together with the roughly7\.57\.5KB of FP32 weights, is also the standard entry format for embedded inference toolchains, so it constitutes a ready artifact for microcontroller\-class deployment\. Two caveats bound this result: the reported latencies are measured on a server\-class CPU core rather than on a microcontroller, and further reductions through integer quantization or sequence downsampling are left as future work\.

## VConclusion

This paper returns to a fundamental question: where, in a magnetic core loss model, should operating conditions enter? CAHR\-Net answers by placing them in the hysteresis\-loop reconstruction process rather than at the model output\. The network predicts along the physicalB⁡\(t\)→H^​\(t\)→B​H→P^vB\(t\)\\rightarrow\\hat\{H\}\(t\)\\rightarrow BH\\rightarrow\\hat\{P\}\_\{\\mathrm\{v\}\}chain, keeping the hysteresis loop as an explicit modeling object throughout\. On this basis, FiLM converts frequency, temperature, and waveform statistics into channel\-wise scales and shifts that act directly on the intermediate representation from which the magnetic\-field waveform is reconstructed—the point at which these conditions physically act—instead of merely correcting a scalar output at the end of the model\. The training protocol is presented together with the network architecture because the two are inseparable: under the original NAdam and step\-decay configuration, the modulation pathway remains in a near\-identity regime; only under the matched protocol combining AdamW, cosine scheduling, and a larger learning rate does it leave that regime and become effective\. CAHR\-Net therefore derives its value not from any single dimension, but from the coupling of physical loop reconstruction, structured condition modulation, and matched optimization: within a model budget on the order of10310^\{3\}parameters, it retains intermediate quantities that engineers can compare directly with measured hysteresis loops while achieving the tail accuracy attainable by black\-box models\.

Under the MagNet A–E final protocol, CAHR\-Net attains the lowest averagep​95p95among the compared methods—6\.89%6\.89\\%with 1874 parameters, versus7\.78%7\.78\\%for the strongest black\-box entry at about48×48\\timesthe parameter count, which it also surpasses on worst\-materialp​95p95—and offers an operating point that neither black\-box nor existing gray\-box models cover: a material\-Dp​95p95compressed from16\.40%16\.40\\%to14\.87%14\.87\\%relative to the reconstruction backbone, and a prediction chain whose intermediate quantities remain electromagnetically readable\. Ablations attribute this gain to the synergy of loop reconstruction, structured condition modulation, and matched optimization\. The accuracy–interpretability combination is also cheap to serve: after an accuracy\-neutral export to a dedicated inference runtime, single\-sample CPU latency drops from0\.820\.82ms to0\.190\.19ms on one thread, so the model can run inside design loops without an accelerator\.

This study remains limited to per\-material training on the five MagNet ferrites and single\-period steady\-state excitation without DC bias\. Future work will address these issues together with material recommendation, uncertainty\-aware design margins, and deployment\-oriented model compression\.

## References

- \[1\]C\. P\. Steinmetz, “On the law of hysteresis,”*Transactions of the American Institute of Electrical Engineers*, vol\. 9, pp\. 1–64, 1892\.
- \[2\]J\. Reinert, A\. Brockmeyer, and R\. W\. A\. A\. De Doncker, “Calculation of losses in ferro\- and ferrimagnetic materials based on the modified Steinmetz equation,”*IEEE Transactions on Industry Applications*, vol\. 37, no\. 4, pp\. 1055–1061, 2001\.
- \[3\]K\. Venkatachalam, C\. R\. Sullivan, T\. Abdallah, and H\. Tacca, “Accurate prediction of ferrite core loss with nonsinusoidal waveforms using only Steinmetz parameters,” in*2002 IEEE Workshop on Computers in Power Electronics*, 2002, pp\. 36–41\.
- \[4\]J\. Mühlethaler, J\. Biela, J\. W\. Kolar, and A\. Ecklebe, “Improved core\-loss calculation for magnetic components employed in power electronic systems,”*IEEE Transactions on Power Electronics*, vol\. 27, no\. 2, pp\. 964–973, 2012\.
- \[5\]T\. Guillod, J\. S\. Lee, H\. Li, S\. Wang, M\. Chen, and C\. R\. Sullivan, “Calculation of ferrite core losses with arbitrary waveforms using the composite waveform hypothesis,” in*2023 IEEE Applied Power Electronics Conference and Exposition*, 2023, pp\. 1586–1593\.
- \[6\]A\. Arruti, J\. Anzola, F\. J\. Perez\-Cebolla, I\. Aizpuru, and M\. Mazuela, “The composite improved generalized Steinmetz equation \(ciGSE\): An accurate model combining the composite waveform hypothesis with classical approaches,”*IEEE Transactions on Power Electronics*, vol\. 39, no\. 1, pp\. 1162–1173, 2024\.
- \[7\]F\. Preisach, “Über die magnetische nachwirkung,”*Zeitschrift für Physik*, vol\. 94, no\. 5–6, pp\. 277–302, 1935\.
- \[8\]D\. C\. Jiles and D\. L\. Atherton, “Theory of ferromagnetic hysteresis,”*Journal of Magnetism and Magnetic Materials*, vol\. 61, no\. 1–2, pp\. 48–60, 1986\.
- \[9\]H\. Li, D\. Serrano, T\. Guillod, S\. Wang, E\. Dogariu, A\. Nadler, M\. Luo, V\. Bansal, N\. K\. Jha, Y\. Chen, C\. R\. Sullivan, and M\. Chen, “How MagNet: Machine learning framework for modeling power magnetic material characteristics,”*IEEE Transactions on Power Electronics*, vol\. 38, no\. 12, pp\. 15 829–15 853, 2023\.
- \[10\]H\. Li, D\. Serrano, S\. Wang, and M\. Chen, “MagNet\-AI: Neural network as datasheet for magnetics modeling and material recommendation,”*IEEE Transactions on Power Electronics*, vol\. 38, no\. 12, pp\. 15 854–15 869, 2023\.
- \[11\]D\. Serrano, H\. Li, S\. Wang, T\. Guillod, M\. Luo, V\. Bansal, N\. K\. Jha, Y\. Chen, C\. R\. Sullivan, and M\. Chen, “Why MagNet: Quantifying the complexity of modeling power magnetic material characteristics,”*IEEE Transactions on Power Electronics*, vol\. 38, no\. 11, pp\. 14 292–14 316, 2023\.
- \[12\]M\. Chen, H\. Li, S\. Wang, T\. Guillod, D\. Serrano, N\. Forster, W\. Kirchgassner, T\. Piepenbrock, O\. Schweins, O\. Wallscheid, Q\. Huang, Y\. Li, X\. Shen, H\. Wouters, W\. Martinez*et al\.*, “MagNet challenge for data\-driven power magnetics modeling,”*IEEE Open Journal of Power Electronics*, vol\. 6, pp\. 883–898, 2025\.
- \[13\]Y\. Hu, J\. Xu, J\. Wang, and W\. Xu, “Physics\-inspired multimodal feature fusion cascaded networks for data\-driven magnetic core loss modeling,”*IEEE Transactions on Power Electronics*, vol\. 39, no\. 9, pp\. 11 356–11 367, 2024\.
- \[14\]N\. Rajput, H\. B\. Sandhibigraha, N\. Agrawal, and V\. M\. Iyer, “An empirical model informed neural network core loss predictor for soft magnetic materials,”*IEEE Transactions on Power Electronics*, vol\. 40, no\. 8, pp\. 11 257–11 267, 2025\.
- \[15\]Q\. Huang, Y\. Li, J\. Zhu, and S\. Li, “Magnetization mechanism\-inspired neural networks for core loss estimation,”*IEEE Transactions on Power Electronics*, vol\. 39, no\. 12, pp\. 16 382–16 390, 2024\.
- \[16\]Y\. Xiao, C\. Li, and Z\. Zheng, “A magnetic core loss model based on physics\-informed neural network with cross\-attention,”*IEEE Transactions on Power Electronics*, vol\. 41, no\. 1, pp\. 92–96, 2026\.
- \[17\]J\. Deng, W\. Wang, Z\. Ning, P\. Venugopal, J\. Popovic, and G\. Rietveld, “High\-frequency core loss modeling based on knowledge\-aware artificial neural network,”*IEEE Transactions on Power Electronics*, vol\. 39, no\. 2, pp\. 1968–1973, 2024\.
- \[18\]Q\. Huang, Y\. Li, Y\. Dou, Y\. Li, J\. Zhu, and S\. Li, “History\-dependent Prandtl\-Ishlinskii neural network for quasi\-static core loss prediction under arbitrary excitation waveforms,”*IEEE Transactions on Power Electronics*, 2025, early access\.
- \[19\]X\. Shen, Y\. Zuo, and W\. Martinez, “Conditional generative adversarial network aided iron loss prediction for high\-frequency magnetic components,”*IEEE Transactions on Power Electronics*, vol\. 39, no\. 8, pp\. 9953–9964, 2024\.
- \[20\]J\. Chen, Y\. Liang, X\. Li, X\. Liu, H\. Jiang, X\. Miao, and L\. Zhang, “A magnetic core loss modeling method based on domain adaptation using multiple kernel maximum mean discrepancy,”*IEEE Transactions on Power Electronics*, 2025, early access\.
- \[21\]W\. Kirchgassner, N\. Forster, T\. Piepenbrock, O\. Schweins, and O\. Wallscheid, “HARDCORE: H\-field and power loss estimation for arbitrary waveforms with residual, dilated convolutional neural networks in ferrite cores,”*IEEE Transactions on Power Electronics*, vol\. 40, no\. 2, pp\. 3326–3335, 2025\.
- \[22\]S\. Bai, J\. Z\. Kolter, and V\. Koltun, “An empirical evaluation of generic convolutional and recurrent networks for sequence modeling,”*arXiv preprint arXiv:1803\.01271*, 2018\.
- \[23\]E\. Perez, F\. Strub, H\. de Vries, V\. Dumoulin, and A\. Courville, “FiLM: Visual reasoning with a general conditioning layer,” in*Proceedings of the AAAI Conference on Artificial Intelligence*, vol\. 32, no\. 1, 2018\.
- \[24\]ONNX Runtime developers, “ONNX Runtime,”https://onnxruntime\.ai, 2021, version 1\.18\.1\.
- \[25\]C\. Tofallis, “A better measure of relative prediction accuracy for model selection and model estimation,”*Journal of the Operational Research Society*, vol\. 66, no\. 8, pp\. 1352–1362, 2015\.
- \[26\]I\. Loshchilov and F\. Hutter, “Decoupled weight decay regularization,” in*International Conference on Learning Representations*, 2019\.
- \[27\]——, “SGDR: Stochastic gradient descent with warm restarts,” in*International Conference on Learning Representations*, 2017\.

Similar Articles

Neural Network-Assisted CLEAN for Channel Modeling in Low-SNR Regimes

arXiv cs.LG

This paper proposes NN-CLEAN, a hybrid framework that embeds a multi-head residual network into the iterative CLEAN extraction loop for efficient multipath parameter estimation in low-SNR channel modeling. It matches traditional Grid-Search CLEAN accuracy while greatly reducing computational complexity and enabling parallelization for real-time MIMO systems.