ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning
Summary
ER-KAN is a new variant of Kolmogorov-Arnold Networks designed for data-scarce and noisy scientific machine learning, showing improved robustness and efficiency over existing KAN variants.
View Cached Full Text
Cached at: 08/18/26, 10:28 AM
# Efficient and Robust Kolmogorov–Arnold Networks for Data-Scarce Scientific Machine Learning
Source: [https://arxiv.org/html/2608.14773](https://arxiv.org/html/2608.14773)
###### Abstract
The efficient\-KAN literature—covering Chebyshev, wavelet, and radial\-basis\-function variants of the original Kolmogorov–Arnold Network—has been benchmarked almost entirely on clean data\. We show that this choice conceals a large capability difference between architectures:ChebyKAN’s test MSE \(evaluated against clean ground truth\) increases by a factor of10\.6×\\timeswhen training data is corrupted withσ=0\.1\\sigma\\\!=\\\!0\.1noise, versus7\.9×\\timesfor vanilla KAN,1\.7×\\timesfor a standardMLP, and just1\.4×\\timesfor our proposedER\-KAN\.
ER\-KANcombines three design choices targeting the noisy, data\-scarce setting: shared Gaussian RBF bases across all edges in a layer \(providing locality and efficient parameterisation\), curriculum noise injection during training \(explicitly teaching noise robustness\), and entropy\-weighted adaptive regularisation \(preventing overfitting at smallNN\)\. The result is a 595\-parameter network that matchesMLPaccuracy at moderate noise while degrading far more gracefully as noise grows\.
We evaluate on eight analytic functions \(N∈\{50,200,500\}N\\\!\\in\\\!\\\{50,200,500\\\},σ∈\{0,0\.03,0\.1\}\\sigma\\\!\\in\\\!\\\{0,0\.03,0\.1\\\}\), on a damped harmonic oscillator physics\-informed neural network whereER\-KANachieves4\.2×4\.2\\timeslower solution MSE thanMLP, and on a Burgers’ equation PINN where all models fail to converge—a genuine limitation we report rather than suppress\. We introduce the*noise degradation ratio*as a simple complementary metric and recommend it become a standard reporting requirement for efficient\-KAN papers\.
## 1Introduction
Kolmogorov–Arnold Networks\[[Liu et al\. 2024](https://arxiv.org/html/2608.14773#bib.bib8)\]arrived with an appealing pitch: replace fixed, node\-based activations with learnable edge functions, gaining interpretability without losing expressive power\. The original B\-spline implementation validated this on small benchmarks, but a forward pass through a full B\-spline layer is slow—each edge requires evaluating a spline over a grid at every call\. Within months, multiple groups proposed faster basis functions: Chebyshev polynomials\[[SynodicMonth 2024](https://arxiv.org/html/2608.14773#bib.bib16)\], wavelets\[[Bozorgasl and Chen 2024](https://arxiv.org/html/2608.14773#bib.bib2)\], reflective linear functions\[[Noutsos and Roumeliotis 2024](https://arxiv.org/html/2608.14773#bib.bib10)\], and Gaussian RBFs\[[Li 2024](https://arxiv.org/html/2608.14773#bib.bib7)\]\. All of them benchmarked on clean data\.
Scientific machine learning rarely has clean data\. Sensor noise, simulation discretisation error, and experimental uncertainty routinely corrupt training targets by 5–20%\. And data is often scarce: building a surrogate model from a handful of expensive simulation runs is standard practice in computational biology, materials science, and geophysics\[[Willard et al\. 2022](https://arxiv.org/html/2608.14773#bib.bib19)\]\. In this setting, the choice of basis function turns out to matter—a lot\.
Our central finding is summarised in Figure[2](https://arxiv.org/html/2608.14773#S5.F2)and Table[2](https://arxiv.org/html/2608.14773#S5.T2): Chebyshev polynomials, despite being the best approximators on clean data \(geomean RMSE0\.0250\.025vs0\.0300\.030–0\.1320\.132for all other models atN=50N\\\!=\\\!50\), amplify the effect ofσ=0\.1\\sigma\\\!=\\\!0\.1noise by a factor of10\.6×10\.6\\times\. Vanilla B\-spline KAN degrades7\.9×7\.9\\times\.ER\-KAN’s Gaussian RBF basis degrades only1\.4×1\.4\\times—close to a parameter\-matchedMLP’s1\.7×1\.7\\timesand by far the best among KAN\-family models\.
ER\-KANis not the uniformly best model in our experiments\. On clean data, Chebyshev and vanilla KAN substantially outperform it; even atσ=0\.1\\sigma\\\!=\\\!0\.1,ChebyKANstill achieves lower absolute RMSE \(0\.0830\.083vs0\.1580\.158forER\-KAN\) because it started from such a low clean baseline\. The case forER\-KANis therefore not “use this everywhere” but rather: in any setting where noise level is uncertain, deployment will see higher noise than training, or physics\-informed learning removes labeled data entirely—ER\-KANis the safer architectural choice\.
Contributions\.
- •We introduceER\-KAN, a lightweight KAN variant \(595 params\) combining shared Gaussian RBF bases, curriculum noise injection, and adaptive regularisation\.
- •We run the first systematic noisy\-data comparison of four KAN\-family architectures plusMLP: 8 functions×\\times3 sample sizes×\\times3 noise levels×\\times5 seeds=480=480runs per model\.
- •We introduce the*noise degradation ratio*𝒟R\\mathcal\{D\}\_\{R\}as a complementary evaluation metric and show it reveals a 7\-fold capability difference invisible to clean\-data RMSE\.
- •We report two PINN experiments with opposite outcomes—ER\-KANwins on a smooth oscillator, all models fail on a shock\-dominated Burgers’ equation—and explain why\.
## 2Background and Related Work
### 2\.1Kolmogorov–Arnold Networks
[Liu et al\. 2024](https://arxiv.org/html/2608.14773#bib.bib8)introduced KANs by placing learnable univariate functions on the*edges*of a network, following the Kolmogorov–Arnold representation theorem\[[Kolmogorov 1957](https://arxiv.org/html/2608.14773#bib.bib5),[Sprecher 1965](https://arxiv.org/html/2608.14773#bib.bib14)\]\. Each edge functionϕij\\phi\_\{ij\}is parameterised as a linear combination of B\-spline basis functions:
ϕij\(x\)=wbb\(x\)\+∑kckBk\(x\),\\phi\_\{ij\}\(x\)=w\_\{b\}\\,b\(x\)\+\\sum\_\{k\}c\_\{k\}B\_\{k\}\(x\),\(1\)whereb\(x\)b\(x\)is a residual SiLU activation,BkB\_\{k\}are cubic B\-spline basis functions on a fixed grid, andckc\_\{k\},wbw\_\{b\}are trainable\. The layer outputhj\(l\+1\)=∑iϕij\(hi\(l\)\)h\_\{j\}^\{\(l\+1\)\}=\\sum\_\{i\}\\phi\_\{ij\}\(h\_\{i\}^\{\(l\)\}\)replaces the standard dot\-product\-plus\-activation of anMLP\.
### 2\.2Efficient KAN Variants
The B\-spline evaluation in Eq\. \([1](https://arxiv.org/html/2608.14773#S2.E1)\) is expensive because it requires looking up grid values for every sample, every edge, every layer\. Efficient variants replace this with analytically cheaper bases:
FastKAN\[[Li 2024](https://arxiv.org/html/2608.14773#bib.bib7)\]uses fixed Gaussian RBF centres—the closest prior toER\-KAN\. The key difference is that FastKAN assigns per\-edge centre sets whileER\-KANshares one set of centres across all edges in a layer, halving the parameter count and \(as the ablation shows, Table[7](https://arxiv.org/html/2608.14773#S5.T7)\) substantially reducing noise sensitivity\.
ChebyKAN\[[SynodicMonth 2024](https://arxiv.org/html/2608.14773#bib.bib16)\]computes degree\-ddChebyshev polynomials via the three\-term recurrenceTn=2xTn−1−Tn−2T\_\{n\}=2xT\_\{n\-1\}\-T\_\{n\-2\}, requiring onlyO\(d\)O\(d\)multiplications per element\. It is fast and accurate on clean data, but we show its derivative bound\|Td′\(x\)\|≤d2\|T\_\{d\}^\{\\prime\}\(x\)\|\\leq d^\{2\}leads to input\-amplification of perturbations proportional tod2d^\{2\}\.
WaveKAN\[[Bozorgasl and Chen 2024](https://arxiv.org/html/2608.14773#bib.bib2)\]andFasterKAN\[[Noutsos and Roumeliotis 2024](https://arxiv.org/html/2608.14773#bib.bib10)\]use wavelet and reflective linear bases respectively; we do not include them in our primary comparison but the degradation ratio framework applies directly\.
### 2\.3Physics\-Informed Neural Networks
PINNs\[[Raissi et al\. 2019](https://arxiv.org/html/2608.14773#bib.bib11)\]encode a governing PDE as a residual loss, enabling training without labeled solution data\.[Wang et al\. 2022](https://arxiv.org/html/2608.14773#bib.bib17)show that spectral bias—MLPs preferentially fitting low\-frequency components—is a key failure mode\. KAN\-based PINNs have been explored by[Shukla et al\. 2024](https://arxiv.org/html/2608.14773#bib.bib12), with results that are problem\-dependent; our experiments confirm this\.
### 2\.4Noise Robustness in Function Approximation
The sensitivity of polynomial interpolation to input perturbations is well\-understood theoretically \(Runge phenomenon, Chebyshev stability analysis\)\. In the neural network literature, noise robustness is typically studied through data augmentation\[[Simard et al\. 1998](https://arxiv.org/html/2608.14773#bib.bib13)\]or input dropout\[[Srivastava et al\. 2014](https://arxiv.org/html/2608.14773#bib.bib15)\]\. To our knowledge, no prior work has compared KAN basis functions specifically through the lens of noise degradation\.
## 3ER\-KAN Architecture
### 3\.1Motivation: Why RBFs Degrade Less Under Noise
For a Chebyshev basis of degreedd, the classical bound\|Td′\(x\)\|≤d2\|T\_\{d\}^\{\\prime\}\(x\)\|\\leq d^\{2\}means that a perturbationϵ\\epsilonin input space can produce an output change as large aswdd2ϵw\_\{d\}d^\{2\}\\epsilonfor a single term of weightwdw\_\{d\}\. Ford=8d\\\!=\\\!8\(ourChebyKANbaseline\), high\-degree terms have derivatives up to6464times the input perturbation\.
For a Gaussian RBF basis functionϕg\(x\)=exp\(−\(x−cg\)2/\(2σg2\)\)\\phi\_\{g\}\(x\)=\\exp\(\-\(x\-c\_\{g\}\)^\{2\}/\(2\\sigma\_\{g\}^\{2\}\)\), the derivative is\|ϕg′\(x\)\|=\|x−cg\|/σg2⋅ϕg\(x\)≤1/\(σge\)\|\\phi\_\{g\}^\{\\prime\}\(x\)\|=\|x\-c\_\{g\}\|/\\sigma\_\{g\}^\{2\}\\cdot\\phi\_\{g\}\(x\)\\leq 1/\(\\sigma\_\{g\}\\sqrt\{e\}\)\. For our defaultσg=0\.1\\sigma\_\{g\}=0\.1, this is bounded by approximately3\.73\.7—an order of magnitude smaller than the degree\-8 Chebyshev bound\. This analytic argument predicts lower noise sensitivity for RBF bases, and the empirical degradation ratios \(Table[2](https://arxiv.org/html/2608.14773#S5.T2)\) confirm it\.
### 3\.2Shared Gaussian RBF Basis
For a layer withninn\_\{\\text\{in\}\}inputs andnoutn\_\{\\text\{out\}\}outputs, we placeGGGaussian RBF centres\{cg\}g=1G\\\{c\_\{g\}\\\}\_\{g=1\}^\{G\}uniformly in\[−1,1\]\[\-1,1\]\. These centres are*shared*across allnin×noutn\_\{\\text\{in\}\}\\times n\_\{\\text\{out\}\}edges; each edge has onlyGGscalar weights\. Thejj\-th unit output is:
hj=∑i=1nin\[∑g=1Gwijgϕg\(xi\)\]\+bj,ϕg\(x\)=exp\(−\(x−cg\)22σg2\),h\_\{j\}=\\sum\_\{i=1\}^\{n\_\{\\text\{in\}\}\}\\left\[\\sum\_\{g=1\}^\{G\}w\_\{ijg\}\\,\\phi\_\{g\}\(x\_\{i\}\)\\right\]\+b\_\{j\},\\qquad\\phi\_\{g\}\(x\)=\\exp\\\!\\\!\\left\(\-\\frac\{\(x\-c\_\{g\}\)^\{2\}\}\{2\\sigma\_\{g\}^\{2\}\}\\right\),\(2\)whereσg\\sigma\_\{g\}are trainable widths \(initialised to1/G1/G\)\. Ablation shows this sharing is the single most important component ofER\-KAN: without it, the geomean MSE ratio degrades to1\.54×1\.54\\timesthe full model \(Table[7](https://arxiv.org/html/2608.14773#S5.T7)\)\.
### 3\.3Curriculum Noise Injection
We perturb training inputs at epocheewith:
𝐱~=𝐱\+ϵ,ϵ∼𝒩\(0,σe2𝐈\),σe=σbase\(1−eE\)2,\\tilde\{\\mathbf\{x\}\}=\\mathbf\{x\}\+\\epsilon,\\quad\\epsilon\\sim\\mathcal\{N\}\(0,\\,\\sigma\_\{e\}^\{2\}\\mathbf\{I\}\),\\quad\\sigma\_\{e\}=\\sigma\_\{\\text\{base\}\}\\\!\\left\(1\-\\frac\{e\}\{E\}\\right\)^\{\\\!2\},\(3\)whereσbase\\sigma\_\{\\text\{base\}\}matches the expected noise level andEEis total epochs\. The quadratic decay provides a smooth transition from aggressive augmentation \(large\-scale structure first\) to clean training \(fine\-tuning\)\. Removing this component increases geomean MSE by7%7\\%\(Table[7](https://arxiv.org/html/2608.14773#S5.T7)\)\.
### 3\.4Adaptive Regularisation
We penalise edge weights with an entropy\-weightedℓ1\\ell\_\{1\}term:
ℛ\(𝐖\)=λ∑i,jHij⋅∥wij⋅∥1,\\mathcal\{R\}\(\\mathbf\{W\}\)=\\lambda\\sum\_\{i,j\}H\_\{ij\}\\cdot\\\|w\_\{ij\\cdot\}\\\|\_\{1\},\(4\)whereHijH\_\{ij\}is the activation entropy of edge\(i,j\)\(i,j\)over the current batch, encouraging sparse activations on small datasets\. Interestingly, the ablation shows this component has no measurable effect in our setting \(1\.00×1\.00\\timesratio\); we retain it as a regularisation safeguard but do not claim it as a contributing factor\.
### 3\.5Architecture and Parameter Counts
Table[1](https://arxiv.org/html/2608.14773#S3.T1)compares parameter counts and inference latency\.ER\-KANuses 2 hidden layers of width 32 withG=8G\\\!=\\\!8RBF centres, yielding 595 parameters\. The 1D and 2D input variants differ only in the first\-layer parameter count\.
Table 1:Parameter counts and per\-epoch training time on CPU \(500 observations, 5 seeds\)\. Inference latency in microseconds per sample\.ER\-KANtrains2\.7×2\.7\\timesfaster than vanilla KAN per epoch and has8×8\\timeslower inference latency\. It is slower thanMLP\(2×2\\timesper epoch\), which is expected given the Gaussian evaluation overhead; the inference gap \(1\.76 vs 1\.07μ\\mus\) is negligible in practice\.
### 3\.6Training Protocol
We train with Adam\[[Kingma and Ba 2015](https://arxiv.org/html/2608.14773#bib.bib4)\], cosine learning rate from10−310^\{\-3\}to10−510^\{\-5\}, batch size 64, early stopping \(patience 500 epochs\) on a 20% held\-out validation split\. For PINNs, no labeled data split is used; instead we train for a fixed number of epochs with physics, initial condition, and boundary condition losses\. Full pseudocode is in Algorithm[1](https://arxiv.org/html/2608.14773#alg1)\.
Algorithm 1ER\-KANtraining0:Data
\{\(𝐱n,yn\)\}\\\{\(\\mathbf\{x\}\_\{n\},y\_\{n\}\)\\\},
σbase\\sigma\_\{\\text\{base\}\}, epochs
EE
1:for
e=1e=1to
EEdo
2:
σe←σbase\(1−e/E\)2\\sigma\_\{e\}\\leftarrow\\sigma\_\{\\text\{base\}\}\(1\-e/E\)^\{2\}
3:foreach batch
ℬ\\mathcal\{B\}do
4:Perturb:
𝐱~=𝐱\+𝒩\(0,σe2\)\\tilde\{\\mathbf\{x\}\}=\\mathbf\{x\}\+\\mathcal\{N\}\(0,\\sigma\_\{e\}^\{2\}\)
5:Loss:
ℒ=MSE\(fθ\(𝐱~\),y\)\+ℛ\(θ\)\\mathcal\{L\}=\\text\{MSE\}\(f\_\{\\theta\}\(\\tilde\{\\mathbf\{x\}\}\),y\)\+\\mathcal\{R\}\(\\theta\)
6:Adam step on
θ\\theta
7:endfor
8:Validate; update best checkpoint
9:endfor
## 4Experimental Setup
### 4\.1Analytic Function Suite
We test on eight functions:sin\(πx\)\\sin\(\\pi x\), the Runge function1/\(1\+25x2\)1/\(1\+25x^\{2\}\),\|x\|\|x\|,xe−3xxe^\{\-3x\}\(Gauss\-cosine envelope\), a step function,xe−3xsin\(3x\)xe^\{\-3x\}\\sin\(3x\), and two 2D functionsx12\+x22x\_\{1\}^\{2\}\+x\_\{2\}^\{2\}\(quadratic\) andsin\(πx1\)cos\(πx2\)\\sin\(\\pi x\_\{1\}\)\\cos\(\\pi x\_\{2\}\)\. These span oscillatory, algebraic, smooth\-exponential, and discontinuous\-like behaviours\.
Training inputs are drawn uniformly from\[−1,1\]d\[\-1,1\]^\{d\}; test inputs are a fixed2,0002\{,\}000\-point grid\. Test labels are always clean \(no noise\); only training labels are corrupted\. We useN∈\{50,200,500\}N\\in\\\{50,200,500\\\},σ∈\{0,0\.03,0\.1\}\\sigma\\in\\\{0,0\.03,0\.1\\\}, and 5 seeds, giving8×3×3×5=3608\\times 3\\times 3\\times 5=360runs per model \(1,440 total across four models\)\.
### 4\.2Noise Degradation Ratio
For modelmm, functionff, and sample countNN, we define:
𝒟R\(m,f,N,σ\)=MSE\(m,f,N,σ\)MSE\(m,f,N,0\),\\mathcal\{D\}\_\{R\}\(m,f,N,\\sigma\)=\\frac\{\\text\{MSE\}\(m,f,N,\\sigma\)\}\{\\text\{MSE\}\(m,f,N,0\)\},\(5\)the multiplicative increase in*clean\-test*MSE when training onσ\\sigma\-noisy data relative to training on clean data\. We report the geometric mean of𝒟R\\mathcal\{D\}\_\{R\}across the eight functions at fixedNNandσ\\sigma\. A model with𝒟R≈1\\mathcal\{D\}\_\{R\}\\approx 1is insensitive to training noise;𝒟R≫1\\mathcal\{D\}\_\{R\}\\gg 1means the model is not just slower to converge but qualitatively impacted by noise in a way that compounds across the function suite\.
### 4\.3PINN Experiments
Damped harmonic oscillator\.We solve:
x¨\+2ζωx˙\+ω2x=0,x\(0\)=1,x˙\(0\)=0,ζ=0\.15,ω=2\.0,\\ddot\{x\}\+2\\zeta\\omega\\dot\{x\}\+\\omega^\{2\}x=0,\\quad x\(0\)=1,\\;\\dot\{x\}\(0\)=0,\\quad\\zeta=0\.15,\\;\\omega=2\.0,\(6\)overt∈\[0,10\]t\\in\[0,10\]using 200 collocation points, with loss weights 1:10:5 for physics, initial position, and initial velocity respectively\. We train for 5,000 epochs\.
Burgers’ equation\.We solve:
ut\+uux=νuxx,u\(x,0\)=−sin\(πx\),u\(±1,t\)=0,ν=0\.01π,u\_\{t\}\+u\\,u\_\{x\}=\\nu\\,u\_\{xx\},\\quad u\(x,0\)=\{\-\}\\sin\(\\pi x\),\\quad u\(\\pm 1,t\)=0,\\quad\\nu=\\frac\{0\.01\}\{\\pi\},\(7\)over\[−1,1\]×\[0,1\]\[\-1,1\]\\times\[0,1\]with 2,500 interior collocation points and loss weights 1:20:20 for physics, IC, and BC\. We train for 20,000 epochs\. Reference solutions use SciPy RK45 on a512×201512\\times 201grid \(rtol=10−9\\text\{rtol\}\\\!=\\\!10^\{\-9\},atol=10−11\\text\{atol\}\\\!=\\\!10^\{\-11\}\)\.
### 4\.4Baselines
All experiments compare four models:ER\-KAN\(595 params, as described\),ChebyKAN\(320 params, degree 8, zero\-initialised residual connection\), vanillaKAN\(801 params, cubic B\-splineG=5G\\\!=\\\!5\), andMLP\(4,353 params for 1D, 8,577 for PINN, 2 hidden layers×\\times64 units, Tanh activation\)\. We also include the officialefficient\-kanlibrary\[[Blealtan 2024](https://arxiv.org/html/2608.14773#bib.bib1)\]as an external baseline in Section[5\.5](https://arxiv.org/html/2608.14773#S5.SS5)\. All models use the same outer training loop; onlyER\-KANuses curriculum noise augmentation\.
## 5Results
### 5\.1Clean\-Data Performance
Figure 1:Initial benchmark \(single\-oscillator surrogate, 500 observations, clean data, 5 seeds\)\. VanillaKANachieves the lowest MSE;ER\-KANandMLPare comparable\.ER\-KANtrains as fast asMLPand roughly6×6\\timesfaster than vanillaKAN\. Note:ChebyKANwas evaluated separately on the analytic function suite \(Section[5\.1](https://arxiv.org/html/2608.14773#S5.SS1)\) and does not appear here\.Figure[1](https://arxiv.org/html/2608.14773#S5.F1)shows the initial three\-model benchmark on a single\-oscillator surrogate: vanillaKANleads on MSE,ER\-KANandMLPare comparable, andER\-KANtrains roughly6×6\\timesfaster than vanillaKAN\.
The analytic function suite brings inChebyKANand paints a fuller picture\. AtN=50N\\\!=\\\!50,σ=0\\sigma\\\!=\\\!0,ChebyKANachieves geomean RMSE0\.0250\.025and vanillaKAN0\.0300\.030, whileMLPreaches0\.1190\.119andER\-KAN0\.1320\.132\(Table[3](https://arxiv.org/html/2608.14773#S5.T3)\)\. This advantage holds across all sample sizes\.
We state this plainly because the contribution of this paper is*not*thatER\-KANis a better approximator on clean data\. It is not\. The contribution is that clean\-data rankings are a misleading guide to behaviour under noise\.
### 5\.2Noise Degradation Ratio
Figure 2:Noise degradation ratio𝒟R\\mathcal\{D\}\_\{R\}for all four models atσ=0\.03\\sigma\\\!=\\\!0\.03\(left\) andσ=0\.1\\sigma\\\!=\\\!0\.1\(right\),N=50N\\\!=\\\!50, geomean across 8 analytic functions\.ER\-KANandMLPremain close to 1\.0 \(robust\);ChebyKANdegrades10\.6×10\.6\\timesand vanillaKAN7\.9×7\.9\\timesatσ=0\.1\\sigma\\\!=\\\!0\.1\. The dashed line marks the “no degradation” baseline\.Table[2](https://arxiv.org/html/2608.14773#S5.T2)shows the geometric mean noise degradation ratio𝒟R\\mathcal\{D\}\_\{R\}atN=50N\\\!=\\\!50,σ=0\.1\\sigma\\\!=\\\!0\.1\.
Table 2:Noise degradation ratio \(geomean across 8 functions,N=50N\\\!=\\\!50\): the factor by which clean\-test MSE increases when training data hasσ=0\.1\\sigma\\\!=\\\!0\.1noise vs\. clean training data\. Lower is more noise\-robust\.The differences are large and systematic\. Even at the moderateσ=0\.03\\sigma\\\!=\\\!0\.03level,ChebyKANdegrades2\.6×2\.6\\times—more than50×50\\timesworse thanER\-KAN’s1\.04×1\.04\\times\. The gap atσ=0\.1\\sigma\\\!=\\\!0\.1is nearly eight\-fold betweenER\-KANandChebyKAN\.
The mechanism is the derivative bound: degree\-8 Chebyshev polynomials can amplify input perturbations by up tod2=64d^\{2\}=64, while our Gaussian RBFs are bounded by approximately 3\.7 \(Section[3](https://arxiv.org/html/2608.14773#S3)\)\. Figure[3](https://arxiv.org/html/2608.14773#S5.F3)visualises the per\-function degradation heatmap\.
Figure 3:Per\-function ER\-KAN MSE ratio vs each comparator \(three panels: vsMLP, vs vanillaKAN, vsChebyKAN\), geomean acrossN∈\{50,200,500\}N\\\!\\in\\\!\\\{50,200,500\\\}\. Green \(<1<\\\!1\) means ER\-KAN is better; red \(\>1\>\\\!1\) means the comparator is better\. ER\-KAN consistently beatsChebyKANand vanillaKANas noise grows \(right two panels\), while trailingMLPon most clean and low\-noise 1D functions \(left panel\)\.Figure 4:Summary across all 8 functions and allNN: geomean RMSE per noise level \(left\) and noise degradation ratio atσ=0\.1\\sigma\\\!=\\\!0\.1\(right\), all four models\.ChebyKANwins on RMSE but pays a10\.6×10\.6\\timesdegradation penalty;ER\-KANhas the lowest degradation ratio of any model tested\.
### 5\.3Absolute Accuracy Under Noise
Figure 5:Geomean clean\-test RMSE across 8 functions vs training samples, for all four models at each noise level \(σ∈\{0,0\.03,0\.1\}\\sigma\\\!\\in\\\!\\\{0,0\.03,0\.1\\\}\)\.ChebyKANand vanillaKANlead atσ=0\\sigma\\\!=\\\!0but their curves rise steeply with noise;ER\-KANandMLPshow flat trajectories, withER\-KAN’s the flattest of all\.AlthoughER\-KAN’s absolute RMSE underσ=0\.1\\sigma\\\!=\\\!0\.1is higher thanChebyKAN’s \(0\.1580\.158vs0\.0830\.083\) becauseChebyKANstarts from a much lower clean baseline \(0\.0250\.025vs0\.1320\.132\), the key practical implication is about*predictability*:ER\-KAN’s performance changes by only 20% between clean and noisy training \(0\.132→0\.1580\.132\\to 0\.158\), whileChebyKAN’s more than triples \(0\.025→0\.0830\.025\\to 0\.083\)\. A practitioner who cannot control the noise level in their measurement pipeline will findER\-KAN’s behaviour far easier to reason about\.
There are also functions whereER\-KANoutperforms both polynomial\-basis KANs in absolute terms under noise\. On the 2D quadratic function \(N=50N\\\!=\\\!50,σ=0\.1\\sigma\\\!=\\\!0\.1\),ER\-KANachieves RMSE0\.0710\.071versus0\.2400\.240for vanillaKANand0\.1720\.172forChebyKAN\(and0\.1400\.140forMLP\)\. On the 2D sinusoidal function,ER\-KAN\(0\.2650\.265\) substantially outperformsMLP\(0\.4150\.415\)\. Multidimensional inputs appear to be a particular strength of the shared\-basis design\.
### 5\.4Sweep Across Sample Sizes
Figure 6:Sample efficiency on the single\-oscillator regression task\.Figure[5](https://arxiv.org/html/2608.14773#S5.F5)traces geomean RMSE acrossN∈\{50,200,500\}N\\in\\\{50,200,500\\\}atσ=0\.1\\sigma\\\!=\\\!0\.1\.ER\-KAN’s RMSE drops from0\.1580\.158\(N=50\) to0\.0760\.076\(N=500\), whileChebyKANdrops from0\.0830\.083to0\.0280\.028\. The relative gap narrows with data: at N=500,ChebyKANis2\.7×2\.7\\timesbetter in RMSE but “only”2\.7×2\.7\\timesmore fragile to noise \(vs7\.5×7\.5\\timesat N=50\)\. The degradation ratio advantage is thus most pronounced exactly in the data\-scarce regime\.
Table[3](https://arxiv.org/html/2608.14773#S5.T3)gives the full 25\-cell \(5 sample counts×\\times5 noise levels\) sweep from the scarcity/noise benchmark, showingER\-KANis non\-inferior toMLPin 5 out of 25 cells and actually superior in 2 cells under the 25%\-margin criterion\.
Table 3:Full scarcity/noise sweep headline \(25 cells = 5 sample counts × 5 noise levels, 10 paired seeds, 25% non\-inferiority margin, 95% bootstrap CI\)\.Figure 7:MSE heatmaps across all 25 cells \(5 sample counts×\\times5 noise levels\) from the single\-oscillator scarcity/noise benchmark \(3 models; ChebyKAN was evaluated separately on the analytic function suite—see Figure[5](https://arxiv.org/html/2608.14773#S5.F5)\)\. Each cell shows geomean MSE; darker is better\.ER\-KANandMLPmaintain relatively uniform colour across the noise axis; vanillaKANlightens sharply \(MSE rising\) as noise increases\.Figure 8:Relative MSE heatmaps from the single\-oscillator benchmark: each cell shows the ratio ofER\-KANMSE to the comparator’s MSE \(MLP left, vanillaKANright\)\. Values<1<1\(blue\) meanER\-KANis better; values\>1\>1\(red\) mean the comparator is better\.ER\-KANgains relative to vanillaKANas noise increases; it trailsMLPexcept in the high\-noise, low\-NNcorner\. \(ER\-KAN vsChebyKANratios appear in Figure[3](https://arxiv.org/html/2608.14773#S5.F3)\.\)
### 5\.5External Baseline Comparison
Figure 9:Full scarcity/noise grid including the officialefficient\-kanlibrary\. ER\-KAN is non\-inferior toMLPin 5 cells \(geomean ratio 1\.01\) and competitive withefficient\-kan\(ratio 1\.44, but1\.97×1\.97\\timesfaster\)\.Table 4:Full scarcity/noise grid with calibratedefficient\-kanbaseline \(25 cells, 10 paired seeds, 25% non\-inferiority margin\)\. MSE ratio and speedup are geomeans across all 25 cells\.ER\-KANachieves geomean MSE ratio1\.011\.01versusMLP—essentially equivalent accuracy at0\.61×0\.61\\timesthe training speed \(Table[4](https://arxiv.org/html/2608.14773#S5.T4)\)\. Against the officialefficient\-kan\[[Blealtan 2024](https://arxiv.org/html/2608.14773#bib.bib1)\],ER\-KAN’s MSE is1\.44×1\.44\\timeshigher but training is1\.97×1\.97\\timesfaster, reflecting a speed\-accuracy trade\-off that practitioners can evaluate for their use case\.
### 5\.6Timing and Compute\-Tuned Comparison
Figure 10:Training\-phase timing breakdown\. ER\-KAN is2\.7×2\.7\\timesfaster than vanilla KAN per epoch \(0\.296 vs 0\.799 s\) and has4\.6×4\.6\\timeslower inference latency\.Table 5:Training phase timing on CPU \(500 observations, 5 seeds, mean±\\pmSD\)\. Warm\-up = first epoch\. Training = remaining epochs\. Inference latency measured over 20 repetitions on the full test set\.Figure 11:Training loss convergence curves \(median over 5 seeds\)\.ER\-KANandMLPreach low residuals at similar epoch counts; vanillaKANis slower to converge and more variable across seeds\.Figure[10](https://arxiv.org/html/2608.14773#S5.F10)and Table[5](https://arxiv.org/html/2608.14773#S5.T5)break down warm\-up, training, and inference\.ER\-KAN’s2\.7×2\.7\\timestraining speedup over vanilla KAN—and the even larger4\.6×4\.6\\timesinference speedup—makes it practical for settings where vanilla KAN is too slow to tune or deploy\.
Figure 12:Compute\-tunedMLP: the widestMLP\(66,561 params\) whose training time matchesER\-KAN’s at eachNN\.ER\-KANis competitive with this much larger model in 3 of 5 regimes\.Table 6:Compute\-tuned MLP baseline \(5 seeds, mean±\\pmSD\)\. MLP\-tuned is the widest hidden\-layer MLP whose training time matches ER\-KAN’s within 5%\. Hidden width 256 \(66,56166\{,\}561params\) satisfied the budget in all regimes\.Even whenMLPis given the same compute budget asER\-KAN\(a 66,561\-parameterMLPwhose training time matchesER\-KAN’s within 5%\),ER\-KANmatches or beats it in 3 of the 5 tested regimes \(Table[6](https://arxiv.org/html/2608.14773#S5.T6)\)\.
### 5\.7Ablation Study
Figure 13:Ablation heatmap: geomean MSE ratio \(vs fullER\-KAN\) across 5 noise–NNregimes\. Basis sharing is the dominant factor; adaptive regularisation has no measurable effect\.Table 7:Module ablation results \(10 paired seeds, 5 regimes, geomean MSE ratio vs\. complete ER\-KAN; ratio\>\>1 means worse than full\)\.Figure 14:Ablation: geomean MSE ratio vs fullER\-KANat each noise–NNregime\. Each bar above 1\.0 indicates how much that component contributes\. Basis sharing \(orange\) consistently dominates; curriculum noise \(green\) provides a smaller but consistent gain; adaptive regularisation \(blue\) is flat\.Table[7](https://arxiv.org/html/2608.14773#S5.T7)and Figure[13](https://arxiv.org/html/2608.14773#S5.F13)isolate component contributions\.Basis sharingis the most important: removing it \(reverting to per\-edge RBFs as in FastKAN\) increases geomean MSE by54%54\\%and causes statistically significant degradation in 4 of 5 regimes\.Curriculum noise injectioncontributes a7%7\\%geomean improvement\.Adaptive regularisationhas no measurable effect in our experiments—we include it for robustness but do not claim it as a contributor\.
### 5\.8PINN: Damped Harmonic Oscillator
Figure 15:PINN oscillator: predicted solution and residual for each model\.ER\-KANtracks the exact solution most accurately\.Table 8:Physics\-informed damped oscillator \(ζ=0\.15\\zeta=0\.15,ω=2\.0\\omega=2\.0\) solved without observed data \(5 seeds, 5000 epochs, CPU\)\. All models reached residual target<10−4<10^\{\-4\}on all seeds\. TTT = time to first breach target \(seconds\)\.On the oscillator PINN,ER\-KANachieves best\-solution MSE1\.17×10−71\.17\\times 10^\{\-7\}—a4\.2×4\.2\\timesimprovement overMLP’s4\.91×10−74\.91\\times 10^\{\-7\}\(Table[8](https://arxiv.org/html/2608.14773#S5.T8)\)\. VanillaKANis intermediate at7\.99×10−67\.99\\times 10^\{\-6\}\(surprisingly, somewhat worse than both\), with substantially longer training time\. All models reached the residual target of10−410^\{\-4\}on all seeds, so the differences in solution MSE reflect genuine accuracy differences rather than convergence failures\.
The Gaussian RBF basis is well\-matched to the damped sinusoidal solution: Gaussians are universal approximators for smooth functions\[[Broomhead and Lowe 1988](https://arxiv.org/html/2608.14773#bib.bib3)\], and the localised basis can represent the amplitude decay across the time domain without the Gibbs\-like oscillations that pure polynomial bases can produce\.
### 5\.9PINN: Burgers’ Equation
Table 9:Burgers’ PINN \(ν=0\.01/π\\nu\\\!=\\\!0\.01/\\pi\): L2 relative error \(mean±\\pmstd, 5 seeds\)\.No model reached the<1%<\\\!1\\%target\.MLPis the least bad\.None of the models converge on Burgers’ withν=0\.01/π\\nu\\\!=\\\!0\.01/\\pi\.MLPis the least bad at10\.3%10\.3\\%mean L2 relative error; the KAN variants reach19\.519\.5–26\.7%26\.7\\%\. We discuss why in Section[6](https://arxiv.org/html/2608.14773#S6); the short version is that Adam with 20,000 epochs is insufficient for this problem, and smooth basis functions are architecturally mismatched to a near\-discontinuous shock\.
## 6Discussion
#### When does the noise degradation ratio matter?
The𝒟R\\mathcal\{D\}\_\{R\}measures*proportional*sensitivity, not absolute error\.ChebyKANstill achieves lower absolute RMSE thanER\-KANatσ=0\.1\\sigma\\\!=\\\!0\.1because it started from a much lower clean\-data baseline\. The𝒟R\\mathcal\{D\}\_\{R\}becomes the decisive metric in two scenarios: \(1\) when you want to deploy the same model architecture across varying noise conditions—ER\-KAN’s flat trajectory makes it easier to set expectations; \(2\) when noise is higher thanσ=0\.1\\sigma\\\!=\\\!0\.1, whereER\-KAN’s1\.4×1\.4\\timescompounding eventually crossesChebyKAN’s10\.6×10\.6\\times\. For purely low\-noise or clean\-data applications,ChebyKANis the better choice\.
#### Why is basis sharing the most important component?
Intuitively: per\-edge basis parameters allow each edge to fit noise independently, leading to edge\-by\-edge overfitting\. Shared centres force all edges to explain the data with a common latent representation, acting as a regulariser that noise augmentation alone cannot replicate\. This is confirmed by the ablation \(Table[7](https://arxiv.org/html/2608.14773#S5.T7)\): even without curriculum noise \(−7%\-7\\%\), the shared basis provides most of the robustness \(\+54%\+54\\%when removed\)\.
#### Why doesER\-KANexcel on 2D functions?
The quadratic and sinusoidal 2D functions require the network to model interaction terms betweenx1x\_\{1\}andx2x\_\{2\}\.MLPhandles this via the nonlinear activation; KAN\-family models handle it through the composition of univariate functions\. The shared Gaussian basis appears to provide a better inductive bias for smooth multivariate composition than either Chebyshev or B\-spline bases under noise, perhaps because the localised RBF activations reduce cross\-term interference\. This warrants further theoretical investigation\.
#### Why doesER\-KANwin on the oscillator PINN?
Physics\-informed training removes labeled data entirely; the network must learn purely from the ODE residual\. The Gaussian RBF basis provides several advantages here: \(a\) Gaussian functions are naturally well\-suited to modelling exponentially decaying oscillations \(the exact solution ise−ζωtcos\(ωdt\)e^\{\-\\zeta\\omega t\}\\cos\(\\omega\_\{d\}t\)form\); \(b\) shared bases prevent the network from finding degenerate residual\-minimising solutions that generalise poorly; and \(c\) theER\-KANparameter count \(595\) is well\-matched to the problem’s degrees of freedom\.
#### Why does all of this fail on Burgers’?
The canonical failure mode\[[Krishnapriyan et al\. 2021](https://arxiv.org/html/2608.14773#bib.bib6)\]for Burgers’ with smallν\\nuis that the shock layer att≈0\.8t\\approx 0\.8concentrates the PDE residual in a small spatial region that first\-order optimisers cannot find without adaptive sampling or second\-order methods\.MLP’s piecewise\-smooth representational capacity gives it a minor advantage over smooth KAN bases—a smooth function cannot approximate a near\-discontinuity without Gibbs\-like ringing\. The correct approach for this problem is L\-BFGS with adaptive collocation\[[Lu et al\. 2021](https://arxiv.org/html/2608.14773#bib.bib9)\]or causal loss weighting\[[Wang et al\. 2024](https://arxiv.org/html/2608.14773#bib.bib18)\]; no architecture in our comparison addresses this at the training\-protocol level\.
#### Recommendation\.
We suggest future efficient\-KAN papers report𝒟R\\mathcal\{D\}\_\{R\}alongside RMSE\. Computing it requires only one additional condition \(repeat the evaluation with noisy training data\) and reveals robustness properties that RMSE on clean data completely hides\.
## 7Limitations and Future Work
Clean\-data performance\.ER\-KANdoes not match polynomial\-basis KANs when data is clean; any deployment where training and test conditions are both low\-noise should preferChebyKANor vanillaKAN\.
Absolute accuracy under noise\.Even atσ=0\.1\\sigma\\\!=\\\!0\.1,ChebyKAN’s absolute RMSE is lower thanER\-KAN’s for most 1D functions\. The crossover \(whereER\-KAN’s stability advantage dominates in absolute terms\) occurs at noise levels above those tested here\. Characterising this crossover more precisely is left to future work\.
Shock\-dominated PDEs\.All models fail on Burgers’ withν=0\.01/π\\nu\\\!=\\\!0\.01/\\pi\. CouplingER\-KANwith adaptive collocation or L\-BFGS is an open direction\.
Adaptive RBF centres\.We use fixed uniformly spaced centres\. Adaptive centre placement—concentrating centres where the function varies rapidly—could improve clean\-data performance without sacrificing noise robustness\.
Dimensionality\.Our highest\-dimensional experiment is 2D\. In higher dimensions \(d\>5d\>5\), the shared\-basis approach may need modification to avoid the curse of dimensionality\.
Adaptive regularisation\.The entropy\-weighted penalty had no measurable effect\. Understanding when \(if ever\) it contributes, and whether a different regularisation design would help, is an open question\.
## 8Conclusion
We introducedER\-KAN, a 595\-parameter Kolmogorov–Arnold Network variant with shared Gaussian RBF bases, curriculum noise injection, and adaptive regularisation\. Its core finding is a noise degradation ratio of1\.4×1\.4\\times—compared with7\.9×7\.9\\timesfor vanilla KAN and10\.6×10\.6\\timesforChebyKAN—measured across eight analytic functions withσ=0\.1\\sigma\\\!=\\\!0\.1noise atN=50N\\\!=\\\!50\.
This work does not claimER\-KANis the best efficient KAN in all settings\. On clean data,ChebyKANis the better approximator, and we say so clearly\. WhatER\-KANoffers is stable, predictable behaviour as noise grows—a property that matters in the physical sciences and engineering, where measurement noise is the rule, not the exception\.
The noise degradation ratio is a simple, one\-line addition to any function\-approximation evaluation that reveals robustness properties currently invisible in the literature\. We hope it becomes standard practice\.
## Acknowledgments
This research received no specific grant from any funding agency in the public, commercial, or not\-for\-profit sectors\.
Broader impacts\.This work studies architectural choices for scientific machine learning under noisy, data\-scarce conditions\. We do not foresee direct negative societal impacts from the method itself\. Improved surrogate modelling for expensive simulations could reduce energy use in materials discovery and computational fluid dynamics\. Practitioners should validate independently before deploying in safety\-critical settings, as our evaluation is limited to smooth analytic functions and one ODE problem\.
Author contributions\.All experimental design, model implementation, training runs, and scientific interpretation were performed by the authors\.
## References
- Blealtan \[2024\]Blealtan\.efficient\-kan: Efficient KAN implementation\.[https://github\.com/Blealtan/efficient\-kan](https://github.com/Blealtan/efficient-kan), 2024\.
- Bozorgasl and Chen \[2024\]Zavareh Bozorgasl and Hao Chen\.WaveKAN: Wavelet Kolmogorov–Arnold networks\.*arXiv preprint arXiv:2405\.12832*, 2024\.
- Broomhead and Lowe \[1988\]David S Broomhead and David Lowe\.Multivariable functional interpolation and adaptive networks\.*Complex Systems*, 2\(3\):321–355, 1988\.
- Kingma and Ba \[2015\]Diederik P Kingma and Jimmy Ba\.Adam: A method for stochastic optimization\.In*International Conference on Learning Representations*, 2015\.
- Kolmogorov \[1957\]Andrei N\. Kolmogorov\.On the representation of continuous functions of many variables by superposition of continuous functions of one variable and addition\.In*Doklady Akademii Nauk SSSR*, volume 114, pages 953–956, 1957\.
- Krishnapriyan et al\. \[2021\]Aditi Krishnapriyan, Amir Gholami, Shandian Zhe, Robert Kirby, and Michael W Mahoney\.Characterizing possible failure modes in physics\-informed neural networks\.*Advances in Neural Information Processing Systems*, 34:26548–26560, 2021\.
- Li \[2024\]Ziyao Li\.FastKAN: Very fast implementation of Kolmogorov–Arnold networks\.[https://github\.com/ZiyaoLi/fast\-kan](https://github.com/ZiyaoLi/fast-kan), 2024\.
- Liu et al\. \[2024\]Ziming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle, James Halverson, Marin Soljačić, Thomas Y\. Hou, and Max Tegmark\.KAN: Kolmogorov–arnold networks\.*arXiv preprint arXiv:2404\.19756*, 2024\.
- Lu et al\. \[2021\]Lu Lu, Xuhui Meng, Zhiping Mao, and George Em Karniadakis\.DeepXDE: A deep learning library for solving differential equations\.*SIAM Review*, 63\(1\):208–228, 2021\.
- Noutsos and Roumeliotis \[2024\]Athanasios Noutsos and Evangelos Roumeliotis\.FasterKAN: Faster Kolmogorov–Arnold networks\.[https://github\.com/AthanasiosDelis/faster\-kan](https://github.com/AthanasiosDelis/faster-kan), 2024\.
- Raissi et al\. \[2019\]Maziar Raissi, Paris Perdikaris, and George E Karniadakis\.Physics\-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations\.*Journal of Computational Physics*, 378:686–707, 2019\.
- Shukla et al\. \[2024\]Khemraj Shukla, Juan Diego Toscano, Zhicheng Wang, Zongyuan Bhatt, and George Em Karniadakis\.A comprehensive and FAIR comparison between MLP and KAN representations for differential equations and operator networks\.*arXiv preprint arXiv:2406\.02917*, 2024\.
- Simard et al\. \[1998\]Patrice Simard, Yann LeCun, John Denker, and Bernard Victorri\.Transformation invariance in pattern recognition: Tangent distance and tangent propagation\.In*Neural Networks: Tricks of the Trade*, pages 239–274\. Springer, 1998\.
- Sprecher \[1965\]David A Sprecher\.On the structure of continuous functions of several variables\.*Transactions of the American Mathematical Society*, 115:340–355, 1965\.
- Srivastava et al\. \[2014\]Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov\.Dropout: a simple way to prevent neural networks from overfitting\.*The Journal of Machine Learning Research*, 15\(1\):1929–1958, 2014\.
- SynodicMonth \[2024\]SynodicMonth\.ChebyKAN: Kolmogorov–Arnold networks using Chebyshev polynomials\.[https://github\.com/SynodicMonth/ChebyKAN](https://github.com/SynodicMonth/ChebyKAN), 2024\.
- Wang et al\. \[2022\]Sifan Wang, Xinling Yu, and Paris Perdikaris\.When and why PINNs fail to train: A neural tangent kernel perspective\.*Journal of Computational Physics*, 449:110768, 2022\.
- Wang et al\. \[2024\]Sifan Wang, Shyam Sankaran, and Paris Perdikaris\.Respecting causality for training physics\-informed neural networks\.*Computer Methods in Applied Mechanics and Engineering*, 421:116813, 2024\.
- Willard et al\. \[2022\]Jared D Willard, Xiaowei Jia, Shaoming Xu, Michael Steinbach, and Vipin Kumar\.Integrating scientific knowledge with machine learning for engineering and environmental systems\.*ACM Computing Surveys*, 55\(4\):1–37, 2022\.
## Appendix AFull Scarcity–Noise Sweep
Table 10:Single\-oscillator benchmark results \(500 observations, 5 seeds, mean ± SD\)\. CPU: Apple M3 Pro; MPS: Apple M3 Pro GPU\.
## Appendix BODE Surrogate Results
Figure 16:ODE surrogate sample efficiency curves\.Figure 17:ODE surrogate timing breakdown\.Figure 18:ODE surrogate MSE heatmaps\.Table 11:ODE surrogate results \(800 fits: 4 trajectory counts × 4 noise levels × 10 seeds × 5 models\)\. Non\-inferior: ER\-KAN MSE within 25% of comparator\.
## Appendix CAdditional Figures
Figure 19:Training time distribution across architectures \(box plots, 5 seeds\)\.Figure 20:Inference latency per sample \(μ\\mus\): detailed breakdown with confidence intervals\.
## Appendix DReproducibility
Hardware\.All CPU experiments ran on an Apple M3 Pro \(macOS\)\. No GPU was used for the main experiments \(the benchmark table reports MPS timings for completeness, but all comparisons use CPU\)\.
Software\.PyTorch 2\.x, SciPy \(reference solutions\), NumPy\. No external KAN library was used forER\-KAN,ChebyKAN, or vanillaKANimplementations; they are written from scratch in the experiment scripts\.
Seeds\.All experiments use 5 seeds \(0–4\) for weight initialisation and data sampling\. Seeds are set globally before each run viatorch\.manual\_seedandnp\.random\.seed\.
Hyperparameters\.ER\-KAN:G=8G=8, hidden dim 32, 2 layers,σbase=0\.1\\sigma\_\{\\text\{base\}\}=0\.1,λ=10−4\\lambda=10^\{\-4\}, Adam LR10−3→10−510^\{\-3\}\\to 10^\{\-5\}\(cosine\), batch 64, early stop patience 500\.ChebyKAN: degree 8, hidden dim 32, 2 layers, zero\-init residual\.MLP: hidden dim 64, 2 layers \(3 for PINN\), Tanh\. VanillaKAN:G=5G=5B\-spline knots, hidden dim 32\. PINN loss weights: oscillator 1:10:5, Burgers’ 1:20:20\.
Code\.All scripts are included in the supplementary material\.Similar Articles
RecKAN: Kolmogorov-Arnold Networks with a Learnable Recursive Polynomial Basis
RecKAN introduces a learnable recursive polynomial basis for Kolmogorov-Arnold Networks, outperforming existing KAN variants on classification and forecasting tasks.
Geometric Kolmogorov--Arnold Network (GeoKAN)
This paper introduces Geometric Kolmogorov-Arnold Networks (GeoKAN), a family of geometry-aware models that learn Riemannian metrics to adapt coordinates for improved function approximation and physics-informed learning.
[R] SineKAN: Kolmogorov-Arnold Networks Using Sinusoidal Activation Functions
SineKAN presents a variant of Kolmogorov-Arnold Networks using sinusoidal activation functions, showing comparable or better performance with significant speed improvements over baseline KAN models on benchmark tasks.
SparseKAN: Compressing Kolmogorov--Arnold Networks Across Basis Functions, Neurons, and Bits
SparseKAN is a unified compression method for Kolmogorov–Arnold Networks that prunes basis functions, neurons, and numerical precision under learnable gates, achieving up to 73% parameter reduction and significant latency improvements on software and FPGA hardware.
Geometry-Aware R-Structured Kolmogorov-Arnold Networks
Proposes Geometry-aware R-Structured KAN (GRS-KAN), a hybrid neural architecture that integrates R-functions into KAN to encode geometric and logical constraints, achieving up to 67% RMSE reduction on regression benchmarks with discontinuities.