SW-KAN: Kolmogorov-Arnold Networks with Stieltjes-Wigert q-Orthogonal Polynomials

arXiv cs.LG Papers

Summary

该论文提出 SW-KAN,一种基于 Stieltjes-Wigert q-正交多项式的 Kolmogorov-Arnold 网络,通过指数-tanh 域映射和 O(N) 三递推计算,在图像分类与函数逼近任务中实现更优的精度-效率权衡。

arXiv:2610.00050v1 Announce Type: new Abstract: Kolmogorov-Arnold Networks (KANs) represent a paradigmatic shift in deep learning by replacing fixed node activations with learnable univariate functions on edges, offering enhanced interpretability and parameter efficiency. While recent polynomial-based KAN variants have addressed the computational overhead of original B-spline implementations, they introduce a fundamental yet underexplored challenge: the domain mismatch between unbounded real-valued inputs and the bounded or semi-infinite support of orthogonal polynomial bases. To address this limitation, we propose the Stieltjes-Wigert Kolmogorov-Arnold Network (SW-KAN), a novel architecture that employs Stieltjes-Wigert q-orthogonal polynomials defined on the semi-infinite domain (0, infinity). We introduce a smooth exponential-of-tanh mapping that stably bridges the domain gap while preserving well-conditioned gradients, and leverage a numerically stable three-term recurrence that evaluates polynomial expansions in O(N) operations without special-function calls. Through comprehensive experiments spanning image classification and continuous function approximation, we demonstrate that SW-KAN achieves superior accuracy-efficiency trade-offs across diverse tasks. The log-normal weight structure and learnable q-parameter of Stieltjes-Wigert polynomials provide a distinct inductive bias that enables robust performance under resource-constrained conditions, including reduced feature dimensionality and limited training data. The proposed architecture not only outperforms established polynomial KAN baselines on standard benchmarks but also exhibits strong representational capacity for approximating complex multivariate functions with remarkably few parameters, making it a compelling alternative for efficient function approximation and classification in resource-constrained settings.
Original Article
View Cached Full Text

Cached at: 10/03/26, 09:50 AM

# Kolmogorov–Arnold Networks with Stieltjes–Wigert 𝑞-Orthogonal Polynomials
Source: [https://arxiv.org/html/2610.00050](https://arxiv.org/html/2610.00050)
Amirhosein Azarpour[https://orcid.org/0009-0004-1738-1919](https://orcid.org/0009-0004-1738-1919)††thanks:These authors contributed equally to this work\.Affiliation:Faculty of Mathematical Sciences, Department of Computer Science, Shahid Beheshti University, TehranSeyyed Moein Kazemi[https://orcid.org/0009-0001-1860-7236](https://orcid.org/0009-0001-1860-7236)††thanks:These authors contributed equally to this work\.\{am\.azarpour, seyy\.kazemi\}@mail\.sbu\.ac\.ir

###### Abstract

Kolmogorov\-Arnold Networks \(KANs\) represent a paradigmatic shift in deep learning by replacing fixed node activations with learnable univariate functions on edges, offering enhanced interpretability and parameter efficiency\. While recent polynomial\-based KAN variants have addressed the computational overhead of original B\-spline implementations, they introduce a fundamental yet underexplored challenge: the domain mismatch between unbounded real\-valued inputs and the bounded or semi\-infinite support of orthogonal polynomial bases\. To address this limitation, we propose the Stieltjes\-Wigert Kolmogorov\-Arnold Network \(SW\-KAN\), a novel architecture that employs Stieltjes\-Wigertqq\-orthogonal polynomials defined on the semi\-infinite domain\(0,∞\)\(0,\\infty\)\. We introduce a smooth exponential\-of\-tanh mapping that stably bridges the domain gap while preserving well\-conditioned gradients, and leverage a numerically stable three\-term recurrence that evaluates polynomial expansions in𝒪⁡\(N\)\\mathcal\{O\}\(N\)operations without special\-function calls\. Through comprehensive experiments spanning image classification and continuous function approximation, we demonstrate that SW\-KAN consistently achieves superior accuracy\-efficiency trade\-offs across diverse tasks\. Our results show that the log\-normal weight structure and learnableqq\-parameter of Stieltjes\-Wigert polynomials provide a distinct inductive bias that enables robust performance under resource\-constrained conditions, including reduced feature dimensionality and limited training data\. The proposed architecture not only outperforms established polynomial KAN baselines on standard benchmarks but also exhibits strong representational capacity for approximating complex multivariate functions with remarkably few parameters\. These findings establish SW\-KAN as a compelling and versatile alternative for efficient function approximation and classification, with particular promise for real\-world applications where computational resources and labeled data are limited\.

11footnotetext:The source code and pretrained models are available at:[https://github\.com/amirhoseinazarpour/SW\-KAN](https://github.com/amirhoseinazarpour/SW-KAN)Keywords:Kolmogorov\-Arnold Networks, Stieltjes\-Wigert polynomials,qq\-orthogonal polynomials, polynomial basis functions, domain mapping, function approximation, deep learning, parameter efficiency, resource\-constrained learning\.

## 1Introduction

Multi\-Layer Perceptrons \(MLPs\) have served as the foundational building blocks of modern deep learning, with their representational capacity grounded in the universal approximation theorem\[[1](https://arxiv.org/html/2610.00050#bib.bib13)\]\. Despite their broad adoption, MLPs assign fixed activation functions to nodes and rely on learned scalar weights on edges, a design that can require large parameter counts to capture intricate nonlinear relationships and offers limited mathematical transparency\. A notable paradigm shift was introduced by Liuet al\.\[[2](https://arxiv.org/html/2610.00050#bib.bib22)\], who proposed Kolmogorov\-Arnold Networks \(KANs\) as a principled alternative to MLPs\. Motivated by the Kolmogorov\-Arnold Representation Theorem, which asserts that any continuous multivariate functionf:\[0,1\]n→ℝf:\[0,1\]^\{n\}\\to\\mathbb\{R\}can be expressed as a finite superposition of continuous univariate functions\[[3](https://arxiv.org/html/2610.00050#bib.bib26)\], KANs replace fixed nodal activations with learnable univariate functions placed on the network’s*edges*\. This architectural choice decouples the representational role of activations from the aggregation role of nodes, enabling more compact models with favorable accuracy\-parameter trade\-offs and improved structural interpretability\[[2](https://arxiv.org/html/2610.00050#bib.bib22)\]\.

Following the introduction of the original KAN, which parametrizes edge activations with B\-spline curves, a rapidly growing body of research has explored alternative basis functions that preserve the KAN framework’s interpretability while addressing practical limitations of splines\. Liuet al\.\[[4](https://arxiv.org/html/2610.00050#bib.bib23)\]subsequently extended the framework in KAN 2\.0, introducing multiplication nodes and a symbolic compiler to further bridge the gap between neural computation and analytical scientific expression\. Seydi\[[5](https://arxiv.org/html/2610.00050#bib.bib27)\]conducted a systematic comparative study of 18 polynomial families as KAN basis functions encompassing orthogonal polynomials, hypergeometric polynomials,qq\-polynomials, and Fibonacci\-related polynomials and evaluated their performance on the MNIST handwritten digit classification benchmark\. Concurrently, Sidharthet al\.\[[6](https://arxiv.org/html/2610.00050#bib.bib8)\]proposed the Chebyshev KAN \(ChebyKAN\), which leverages the minimax optimality and three\-term recurrence of Chebyshev polynomials of the first kind to parametrize edge activations, demonstrating improved parameter efficiency over the original B\-spline formulation\. The fractional KAN \(fKAN\) of Aghaei\[[7](https://arxiv.org/html/2610.00050#bib.bib2)\]further extended this line of inquiry by incorporating trainable fractional\-orthogonal Jacobi basis functions, exploiting their non\-polynomial behavior and analytically simple derivative formulas to improve accuracy across both supervised and physics\-informed tasks\. Aghaei\[[8](https://arxiv.org/html/2610.00050#bib.bib1)\]additionally investigated rational basis functions for KANs via Padé approximation and rational Jacobi functions, establishing the rational KAN \(rKAN\) and reporting competitive classification accuracy on standard benchmarks\. Collectively, these works establish a clear trajectory: classical orthogonal and rational polynomial families are increasingly preferred over B\-splines as the primary parametrization strategy for KAN edge activations\.

However, several substantive limitations persist in the current landscape of polynomial KAN variants\. The original B\-spline formulation incurs significant computational overhead, as maintaining knot vectors and evaluating piecewise polynomial segments complicates both the forward pass and gradient computation\[[2](https://arxiv.org/html/2610.00050#bib.bib22),[8](https://arxiv.org/html/2610.00050#bib.bib1)\]\. While classical polynomial KANs attenuate these costs, they frequently exhibit sub\-optimal accuracy\-parameter trade\-offs: increasing polynomial degree to improve approximation fidelity inflates the parameter count without proportional gains in task accuracy, as demonstrated by the degree\-accuracy ablation reported for ChebyKAN\[[6](https://arxiv.org/html/2610.00050#bib.bib8)\]\. Furthermore, a structural challenge that receives insufficient attention in the existing literature is the*domain mismatch problem*\. Classical orthogonal polynomial families such as Chebyshev and Legendre are defined on the bounded interval\[−1,1\]\[\-1,1\], while layer inputs in deep networks are, in general, unbounded real\-valued quantities\. Existing approaches commonly address this by applying atanh\\tanhnormalization to compress inputs into\[−1,1\]\[\-1,1\]; however, this compression can saturate in the tails and introduce gradient attenuation for large\-magnitude activations\[[6](https://arxiv.org/html/2610.00050#bib.bib8)\]\. For polynomial families defined on semi\-infinite or otherwise non\-standard domains whose weight structures may offer distinct inductive biases no principled and numerically stable domain\-bridging strategy has been systematically proposed or evaluated within the KAN framework\. Recent work on learnable recursive polynomial bases for KANs\[[9](https://arxiv.org/html/2610.00050#bib.bib21)\]has shown the promise of adapting the basis\-generating rule itself\.

To address these limitations, we propose theStieltjes\-Wigert KAN \(SW\-KAN\), a novel KAN architecture that replaces B\-spline edge activations with Stieltjes\-Wigert \(SW\)qq\-orthogonal polynomials, which are orthogonal with respect to a log\-normal weight function on the semi\-infinite support\(0,∞\)\(0,\\infty\), providing a mathematically distinct inductive structure relative to classical bounded\-domain polynomial families\. To resolve the mismatch between unbounded real\-valued layer inputsx∈ℝx\\in\\mathbb\{R\}and this strictly positive support, we introduce a smooth, invertible*exponential\-of\-tanh*domain mappingϕ⁡\(x\)=etanh⁡\(x\)\\phi\(x\)=e^\{\\tanh\(x\)\}, which bijectively mapsℝ\\mathbb\{R\}onto the compact subinterval\(e−1,e\)⊂\(0,∞\)\(e^\{\-1\},e\)\\subset\(0,\\infty\)while maintaining bounded, well\-conditioned gradients\. The SW polynomial expansion along each edge is then computed via a three\-term recurrence, evaluating a degree\-NNbasis in only𝒪⁡\(N\)\\mathcal\{O\}\(N\)operations per edge and bypassing costly hypergeometric special\-function evaluations\.

The main contributions of this work are:

- •SW\-KAN Architecture:learnable edge activations parametrized by Stieltjes\-Wigertqq\-orthogonal polynomials, introducing a semi\-infinite log\-normal\-support polynomial family into the KAN design space\.
- •Exponential\-of\-Tanh Domain Mapping:a smooth, invertible transformation that stably projects unbounded real\-valued layer inputs into the semi\-infinite support\(0,∞\)\(0,\\infty\)of the SW polynomials, preserving well\-conditioned gradients and mitigating numerical saturation\.
- •Numerically Stable Recurrence\-Based Evaluation:an efficient, numerically stable recurrence that computes the full degree\-NNSW expansion per edge in𝒪⁡\(N\)\\mathcal\{O\}\(N\)operations, without reliance on hypergeometric function libraries\.
- •Empirical Validation Across Diverse Tasks:on MNIST\[[5](https://arxiv.org/html/2610.00050#bib.bib27)\], SW\-KAN achieves a competitive test accuracy of98\.24%98\.24\\%while outperforming all 18 polynomial KAN baselines evaluated therein, including Gottlieb\-KAN and Vieta\-Pell\-KAN, at a comparable parameter count\. We further validate SW\-KAN under resource\-constrained conditions on Fashion\-MNIST \(reduced feature dimensionality, limited training data\) and on continuous multivariate function approximation, showing its advantages generalize beyond a single benchmark and task type\.

The remainder of this paper is organized as follows: Section[2](https://arxiv.org/html/2610.00050#S2)reviews the mathematical foundations of KART, Stieltjes\-Wigertqq\-orthogonal polynomials, and related KAN architectures; Section[3](https://arxiv.org/html/2610.00050#S3)details the SW\-KAN architecture, including the domain mapping and recurrence\-based forward pass; Section[4](https://arxiv.org/html/2610.00050#S4)presents results on MNIST, Fashion\-MNIST, and continuous function approximation against polynomial\-KAN and B\-spline\-KAN baselines; Section[5](https://arxiv.org/html/2610.00050#S5)concludes with findings and future directions\.

## 2Related Work

### 2\.1Theoretical Foundations and the Limits of Multi\-Layer Perceptrons

The theoretical grounding of neural\-network function approximation rests on the Universal Approximation Theorem, which establishes that Multi\-Layer Perceptrons \(MLPs\) with a single hidden layer of sufficient width can approximate any continuous function on a compact domain to arbitrary precision\[[10](https://arxiv.org/html/2610.00050#bib.bib15)\]\. Despite this expressive guarantee, the MLP architecture fixes its non\-linearities at the*nodes*, forcing representational capacity to be purchased exclusively through increased width or depth\. This design choice imposes well\-documented parameter inefficiency: for a given approximation accuracy, MLPs often require substantially more trainable parameters than the target function’s intrinsic complexity would suggest, while simultaneously obscuring the mathematical structure of the learned mapping\[[2](https://arxiv.org/html/2610.00050#bib.bib22)\]\. A parallel theoretical tradition, originating with Kolmogorov’s resolution of Hilbert’s thirteenth problem, offers a structurally distinct viewpoint: the Kolmogorov\-Arnold Representation Theorem \(KART\) asserts that any continuous functionf:\[0,1\]n→ℝf\\colon\[0,1\]^\{n\}\\\!\\to\\\!\\mathbb\{R\}can be expressed as a finite superposition of continuous univariate functions\[[11](https://arxiv.org/html/2610.00050#bib.bib34)\]\. Montanelli and Du\[[12](https://arxiv.org/html/2610.00050#bib.bib25)\]demonstrated that a constructive proof of KART directly yields sharp approximation\-error bounds for deep ReLU networks, providing a principled alternative to empirical depth\-width heuristics\. The representational analysis was further refined in\[[13](https://arxiv.org/html/2610.00050#bib.bib9)\], where it was shown that a modified form of the Kolmogorov\-Arnold decomposition transfers the smoothness of the target function to its univariate components a property that is essential for stable numerical optimisation\. Together, these results recast KART from a pure\-mathematical theorem into a tractable blueprint for interpretable network design\.

### 2\.2Kolmogorov\-Arnold Networks: Activations on Edges

Motivated by KART, Liu et al\.\[[2](https://arxiv.org/html/2610.00050#bib.bib22)\]introduced Kolmogorov\-Arnold Networks \(KANs\), which relocate learnable non\-linearities from the nodes to the*edges*: every scalar weight in a conventional MLP is replaced by a trainable univariate B\-spline function, so each connection itself performs a non\-linear transformation before the outputs are summed at the receiving node\. This seemingly localised change has two empirically validated consequences\. First, for equivalent approximation accuracy, KANs require substantially fewer parameters than comparable MLPs on function\-regression and partial differential equation \(PDE\) benchmarks\[[2](https://arxiv.org/html/2610.00050#bib.bib22)\]\. Second, the learned edge functions can be directly inspected and symbolically extracted, yielding models that are interpretable in a mathematically meaningful sense rather than merely post\-hoc explainable\. Liu et al\. extended the framework in KAN 2\.0\[[4](https://arxiv.org/html/2610.00050#bib.bib23)\]by introducing multiplication nodes, a KAN\-to\-symbolic compiler, and tree\-graph converters that enable the automated discovery of physical conservation laws and Lagrangian structures\. The applied utility of the edge\-activation paradigm was corroborated by Wang et al\.\[[14](https://arxiv.org/html/2610.00050#bib.bib32)\], whose Kolmogorov\-Arnold Informed Neural Network \(KINN\) integrated KAN layers into a physics\-informed training loop and demonstrated that KINNs outperform MLP\-based PINNs in accuracy and convergence speed across a range of solid\-mechanics PDEs\. Furthermore, Bozorgasl and Chen\[[15](https://arxiv.org/html/2610.00050#bib.bib5)\]explored wavelet\-parameterised edge functions, showing that the flexibility of the KAN edge is not restricted to splines but can accommodate any sufficiently regular univariate family, thereby generalising the architectural principle introduced in\[[2](https://arxiv.org/html/2610.00050#bib.bib22)\]\.

### 2\.3Polynomial KAN Variants: Efficiency Gains and Persistent Bottlenecks

The computational characteristics of B\-spline edge functions, however, present scalability challenges that have motivated a sustained research effort in polynomial KAN variants\. Specifically, evaluating B\-splines of degreekkover a grid ofGGknots requires𝒪⁡\(k​G\)\\mathcal\{O\}\(kG\)recursive operations per edge that resist straightforward parallelisation, as demonstrated in the MatrixKAN analysis of Atwell et al\.\[[16](https://arxiv.org/html/2610.00050#bib.bib24)\]\. Furthermore, the grid\-extension procedure which redistributes knots to track the empirical input distribution during training couples the network’s parameter structure to the data statistics in a manner that complicates mini\-batch training and requires careful scheduling\[[17](https://arxiv.org/html/2610.00050#bib.bib31)\]\. Free\-Knots KAN\[[18](https://arxiv.org/html/2610.00050#bib.bib11)\]partially addressed training instability by analytically bounding the number of active knots, yet the dependency of approximation quality on input\-domain coverage remained a structural constraint\.

In response to these limitations, a body of work has pursued globally defined orthogonal polynomial bases as drop\-in replacements for B\-splines on KAN edges\. Seydi\[[5](https://arxiv.org/html/2610.00050#bib.bib27)\]conducted a systematic comparative evaluation of eighteen distinct polynomial families including classical orthogonal polynomials \(Chebyshev, Legendre, Jacobi, Gegenbauer\), hypergeometric polynomials,qq\-polynomials, and combinatorial families as KAN edge\-basis functions, finding that while certain families achieve competitive accuracy on image classification benchmarks, no single family consistently attains an optimal parameter\-accuracy trade\-off across heterogeneous tasks\. The Chebyshev KAN \(Cheb\-KAN\) proposed in\[[6](https://arxiv.org/html/2610.00050#bib.bib8)\]demonstrated that Chebyshev polynomials of the first kind, by virtue of their orthogonality on\[−1,1\]\[\-1,1\]with respect to the Chebyshev weight and their minimax approximation property, yield parameter\-efficient edge activations with improved function\-approximation fidelity relative to B\-spline KANs of comparable size\. This orthogonality advantage was further exploited by Li et al\.\[[19](https://arxiv.org/html/2610.00050#bib.bib7)\]in a physics\-informed Chebyshev\-KAN for fluid\-mechanics PDEs, achieving accuracy improvements at reduced parameter counts\. Concurrently, Aghaei\[[7](https://arxiv.org/html/2610.00050#bib.bib2)\]introduced the Fractional KAN \(fKAN\), which employs trainable fractional\-orthogonal Jacobi functions as edge activations; the fractional parameterisation confers non\-polynomial behaviour and activity for both positive and negative inputs, resulting in measurable training\-speed gains over integer\-order polynomial variants\. Rational extensions were explored by He et al\.\[[8](https://arxiv.org/html/2610.00050#bib.bib1)\], whose rKAN substitutes Padé\-approximant and rational Jacobi basis functions for B\-splines, and by Aghaei and colleagues in the context of 3D point\-cloud processing\[[20](https://arxiv.org/html/2610.00050#bib.bib17)\], where Jacobi polynomials with tunable shape parameters\(α,β\)\(\\alpha,\\beta\)were evaluated across multiple special cases \(Legendre, Chebyshev, Gegenbauer\)\.

Collectively, these polynomial KAN variants address the computational overhead of B\-spline recursion and grid management\. However, they share a common structural challenge that has received insufficient systematic treatment in the existing literature: classical orthogonal polynomials are orthogonal with respect to weight functions supported on*bounded*intervals, most commonly\[−1,1\]\[\-1,1\]\. When such architectures process inputs drawn from the unbounded real lineℝ\\mathbb\{R\}, an explicit domain\-mapping step is required to project each input onto the polynomial’s natural domain\. Commonly employed strategies—linear rescaling, sigmoid, ortanh\\tanhnormalisation—may saturate the polynomial’s approximation range for inputs with heavy\-tailed or skewed distributions, or introduce boundary artefacts near±1\\pm 1where polynomial oscillations concentrate\[[5](https://arxiv.org/html/2610.00050#bib.bib27),[6](https://arxiv.org/html/2610.00050#bib.bib8)\]\. In the absence of a numerically principled mapping, evaluations near the boundary of\[−1,1\]\[\-1,1\]may exhibit gradient vanishing or amplification, degrading both training convergence and out\-of\-distribution generalisation\. This domain\-mismatch problem, compounded by residual parameter\-accuracy trade\-offs, represents the principal unresolved bottleneck in the polynomial\-KAN paradigm\.

### 2\.4q\-Orthogonal Polynomials and the Semi\-Infinite Domain Opportunity

A class of polynomial families that sidesteps the domain\-mismatch problem entirely consists of orthogonal polynomials defined on semi\-infinite or fully unbounded domains\. Within the classical theory ofqq\-orthogonal polynomials\[[21](https://arxiv.org/html/2610.00050#bib.bib18)\], the Stieltjes\-Wigert polynomials occupy a distinguished position: they are orthogonal with respect to a log\-normal weight function supported on the semi\-infinite interval\(0,∞\)\(0,\\infty\), rendering them structurally compatible with positive real\-valued inputs without any domain truncation\. Moreover, Stieltjes\-Wigert polynomials satisfy a three\-term recurrence relation that admits𝒪⁡\(N\)\\mathcal\{O\}\(N\)evaluation of the firstNNbasis functions without invoking special\-function libraries, a computational profile that compares favourably with B\-spline grids of comparable polynomial degree and with the trigonometric evaluations required by Chebyshev recurrences for largeNN\[[5](https://arxiv.org/html/2610.00050#bib.bib27)\]\. Theqq\-deformation parameter provides an additional mechanism for adjusting the polynomial’s concentration and oscillatory frequency to match the empirical input distribution, a degree of adaptability absent from fixed classical families\. Furthermore, the underlyingqq\-analytic framework connects Stieltjes\-Wigert polynomials to the broader Askeyqq\-scheme\[[21](https://arxiv.org/html/2610.00050#bib.bib18)\], providing a well\-developed approximation\-theoretic foundation\. Despite these properties, the integration ofqq\-orthogonal polynomial families into KAN edge activations, together with a principled and numerically stable mechanism for mapping arbitrary real\-valued inputs onto\(0,∞\)\(0,\\infty\), has not been addressed in the existing literature\. The present work closes this gap by proposing SW\-KAN, which replaces B\-spline edge activations with Stieltjes\-Wigertqq\-orthogonal basis functions, employs an exponential\-of\-tanh\\tanhtransformation for bijective and numerically stable domain matching fromℝ\\mathbb\{R\}to\(0,∞\)\(0,\\infty\), and exploits the three\-term recurrence for𝒪⁡\(N\)\\mathcal\{O\}\(N\)basis evaluation simultaneously resolving the grid\-extension overhead of B\-spline KANs and the domain\-mismatch instability inherent in bounded\-polynomial KANs\.

## 3Methodology

### 3\.1Kolmogorov–Arnold Theorem

The Kolmogorov–Arnold Representation Theorem \(KART\) is a fundamental result in approximation theory that provides the theoretical foundation for Kolmogorov–Arnold Networks\[[11](https://arxiv.org/html/2610.00050#bib.bib34),[22](https://arxiv.org/html/2610.00050#bib.bib3)\]\.

#### 3\.1\.1Statement of the Theorem

In 1957, Kolmogorov proved that any continuous multivariate function on a bounded domain can be exactly decomposed into a finite composition of continuous univariate functions and addition, thereby resolving Hilbert’s 13th problem in the affirmative\[[11](https://arxiv.org/html/2610.00050#bib.bib34)\]\. The result was subsequently refined by his student Arnold\[[22](https://arxiv.org/html/2610.00050#bib.bib3)\]\.

###### Theorem 1\(Kolmogorov, 1957; Arnold, 1957\)\.

For any continuous functionf:\[0,1\]n→ℝf:\[0,1\]^\{n\}\\to\\mathbb\{R\}, there exist continuous univariate functionsΦq:ℝ→ℝ\\Phi\_\{q\}:\\mathbb\{R\}\\to\\mathbb\{R\}andφq,p:\[0,1\]→ℝ\\varphi\_\{q,p\}:\[0,1\]\\to\\mathbb\{R\}such that

f⁡\(x1,x2,…,xn\)=∑q=02​nΦq​\(∑p=1nφq,p​\(xp\)\)\.f\(x\_\{1\},x\_\{2\},\\ldots,x\_\{n\}\)=\\sum\_\{q=0\}^\{2n\}\\Phi\_\{q\}\\\!\\left\(\\sum\_\{p=1\}^\{n\}\\varphi\_\{q,p\}\(x\_\{p\}\)\\right\)\.\(1\)

Equation \([1](https://arxiv.org/html/2610.00050#S3.E1)\) shows that any continuous mapping innnvariables can be constructed entirely from univariate functions and summation no multiplication between input variables is required\[[23](https://arxiv.org/html/2610.00050#bib.bib6)\]\. The representation is exact, not an approximation\. The original construction uses a two\-layer structure: an inner layer ofφq,p\\varphi\_\{q,p\}functions and an outer layer ofΦq\\Phi\_\{q\}functions, with widths\(2​n\+1\)\(2n\+1\)and11, respectively\.

#### 3\.1\.2Significance and Limitations

The essence of the theorem lies in its ability to transform a function of many variables into a sum of functions, each depending on only one variable\. This decomposition demonstrates that multivariate continuous functions are not fundamentally more complex than univariate ones, as they can always be expressed through compositions of the latter\[[11](https://arxiv.org/html/2610.00050#bib.bib34),[3](https://arxiv.org/html/2610.00050#bib.bib26)\]\.

Despite its theoretical elegance, the KART was long considered impractical for machine learning due to the potential non\-smoothness of the inner functionsφq,p\\varphi\_\{q,p\}and the lack of constructive algorithms\[[24](https://arxiv.org/html/2610.00050#bib.bib12),[23](https://arxiv.org/html/2610.00050#bib.bib6)\]\. Early attempts to connect the KART to neural network architectures were explored by Hecht\-Nielsen\[[25](https://arxiv.org/html/2610.00050#bib.bib14)\], who reinterpreted the theorem as an existence result for mapping neural networks\. This interpretation was later placed on firmer theoretical footing by Kůrková\[[26](https://arxiv.org/html/2610.00050#bib.bib20)\], who demonstrated that the representation could be realized within standard neural network frameworks\.

#### 3\.1\.3From KART to KAN

Recently, Liu et al\.\[[2](https://arxiv.org/html/2610.00050#bib.bib22)\]revived this line of research by proposing the Kolmogorov–Arnold Network \(KAN\), a neural architecture that replaces fixed activation functions at nodes with learnable univariate functions on edges\. In a KAN, each edge carries a learnable univariate function and each node performs only summation\. For a KAN layer mapping𝐱\(ℓ\)∈ℝnℓ\\mathbf\{x\}^\{\(\\ell\)\}\\in\\mathbb\{R\}^\{n\_\{\\ell\}\}to𝐱\(ℓ\+1\)∈ℝnℓ\+1\\mathbf\{x\}^\{\(\\ell\+1\)\}\\in\\mathbb\{R\}^\{n\_\{\\ell\+1\}\}, the output of thejj\-th neuron is

xj\(ℓ\+1\)=∑i=1nℓφj,i\(ℓ\)\(xi\(ℓ\)\),j=1,…,nℓ\+1,x\_\{j\}^\{\(\\ell\+1\)\}=\\sum\_\{i=1\}^\{n\_\{\\ell\}\}\\varphi\_\{j,i\}^\{\(\\ell\)\}\\\!\\left\(x\_\{i\}^\{\(\\ell\)\}\\right\),\\quad j=1,\\ldots,n\_\{\\ell\+1\},\(2\)where eachφj,i\(ℓ\):ℝ→ℝ\\varphi\_\{j,i\}^\{\(\\ell\)\}:\\mathbb\{R\}\\to\\mathbb\{R\}is a learnable univariate activation function placed on the edge connecting nodeiiin layerℓ\\ellto nodejjin layerℓ\+1\\ell\+1\[[2](https://arxiv.org/html/2610.00050#bib.bib22)\]\.

While the original KAN parametrizes eachφj,i\(ℓ\)\\varphi\_\{j,i\}^\{\(\\ell\)\}using B\-splines, the framework is inherently modular: any family of univariate functions can serve as the basis\[[2](https://arxiv.org/html/2610.00050#bib.bib22),[6](https://arxiv.org/html/2610.00050#bib.bib8)\]\. This modularity has prompted a wave of polynomial KAN variants, including those based on Chebyshev\[[6](https://arxiv.org/html/2610.00050#bib.bib8)\], Jacobi\[[7](https://arxiv.org/html/2610.00050#bib.bib2)\], and Legendre\[[13](https://arxiv.org/html/2610.00050#bib.bib9)\]polynomials\. In this work, we replace B\-splines with the Stieltjes–Wigert polynomials, aqq\-orthogonal family with distinctive computational and theoretical properties \(Section[3\.2](https://arxiv.org/html/2610.00050#S3.SS2)\)\.

### 3\.2Stieltjes–Wigert Polynomials

The Stieltjes–Wigert polynomials are a family ofqq\-orthogonal polynomials first studied by Stieltjes\[[27](https://arxiv.org/html/2610.00050#bib.bib29)\]and Wigert\[[28](https://arxiv.org/html/2610.00050#bib.bib33)\], residing at the lower level of theqq\-Askey scheme\[[29](https://arxiv.org/html/2610.00050#bib.bib19)\]\. They are distinguished by their connection to the log\-normal distribution and the remarkable property that their associated moment problem is indeterminate\[[30](https://arxiv.org/html/2610.00050#bib.bib10)\]\. Named after the mathematicians Thomas Jan Stieltjes and Carl Severin Wigert, these polynomials are especially important in the context ofqq\-series and approximation theory due to their single\-parameter simplicity and efficient recurrence structure\[[31](https://arxiv.org/html/2610.00050#bib.bib16)\]\.

#### 3\.2\.1Definition and Equations

The Stieltjes–Wigert polynomialsSn​\(x,q\)S\_\{n\}\(x;q\)are defined for a parameter0<q<10<q<1andx∈\(0,∞\)x\\in\(0,\\infty\)\.

#### 3\.2\.2Hypergeometric Form

The polynomials are given in terms of the basic hypergeometric functionϕ11\{\}\_\{1\}\\phi\_\{1\}as\[[29](https://arxiv.org/html/2610.00050#bib.bib19)\]:

Sn​\(x,q\)\\displaystyle S\_\{n\}\(x;q\)=1\(q,q\)n​ϕ11​\(q−n0,q,−qn\+1​x\)=1\(q,q\)n​∑k=0n\(nk\)q​qk2​\(−x\)k,\\displaystyle=\\frac\{1\}\{\(q;\\,q\)\_\{n\}\}\\;\{\}\_\{1\}\\phi\_\{1\}\\\!\\left\(\\frac\{q^\{\-n\}\}\{0\}\\;;\\;q,\\;\-q^\{n\+1\}x\\right\)=\\frac\{1\}\{\(q;\\,q\)\_\{n\}\}\\sum\_\{k=0\}^\{n\}\\binom\{n\}\{k\}\_\{\\\!q\}\\,q^\{k^\{2\}\}\\,\(\-x\)^\{k\},\(3\)where\(q,q\)n=∏j=1n\(1−qj\)\(q;\\,q\)\_\{n\}=\\prod\_\{j=1\}^\{n\}\(1\-q^\{j\}\)is theqq\-shifted factorial \(Pochhammer symbol\), and\(nk\)q=\(q,q\)n\(q,q\)k​\(q,q\)n−k\\displaystyle\\binom\{n\}\{k\}\_\{\\\!q\}=\\frac\{\(q;\\,q\)\_\{n\}\}\{\(q;\\,q\)\_\{k\}\\,\(q;\\,q\)\_\{n\-k\}\}is theqq\-binomial coefficient\.

The parameterqqis related to the weight\-function parameterk\>0k\>0by

q=exp⁡\(−12​k2\)\.q=\\exp\\\!\\left\(\-\\frac\{1\}\{2k^\{2\}\}\\right\)\.\(4\)
##### 3\.2\.2\.1 Explicit Low\-Degree Forms

The first few Stieltjes–Wigert polynomials are:

S0​\(x,q\)\\displaystyle S\_\{0\}\(x;\\,q\)=1,\\displaystyle=1,\(5\)S1​\(x,q\)\\displaystyle S\_\{1\}\(x;\\,q\)=11−q​\(1−q2​x\),\\displaystyle=\\frac\{1\}\{1\-q\}\\bigl\(1\-q^\{2\}x\\bigr\),\(6\)S2​\(x,q\)\\displaystyle S\_\{2\}\(x;\\,q\)=1\(1−q\)​\(1−q2\)​\(1−\(1\+q\)​q2​x\+q5​x2\),\\displaystyle=\\frac\{1\}\{\(1\-q\)\(1\-q^\{2\}\)\}\\bigl\(1\-\(1\+q\)\\,q^\{2\}x\+q^\{5\}x^\{2\}\\bigr\),\(7\)S3​\(x,q\)\\displaystyle S\_\{3\}\(x;\\,q\)=1\(q,q\)3​\(1−CLOSEOPEN\(1\+q\+q2\)​q2​x\+\(1\+q\+q2\)​q5​x2−q9​x3\)\.\\displaystyle=\\begin\{aligned\} \\frac\{1\}\{\(q;\\,q\)\_\{3\}\}\\bigl\(1\-&\(1\+q\+q^\{2\}\)\\,q^\{2\}x\+\(1\+q\+q^\{2\}\)\\,q^\{5\}x^\{2\}\-q^\{9\}x^\{3\}\\bigr\)\.\\end\{aligned\}\(8\)

#### 3\.2\.3Recurrence Relations

The Stieltjes–Wigert polynomials satisfy a three\-term recurrence relation, which is central for their efficient computation\. The recurrence reads\[[29](https://arxiv.org/html/2610.00050#bib.bib19),[21](https://arxiv.org/html/2610.00050#bib.bib18)\]:

−q2​n\+1​x​Sn​\(x,q\)=\(1−qn\+1\)​Sn\+1​\(x,q\)−\(1\+q−qn\+1\)​Sn​\(x,q\)\+q​Sn−1​\(x,q\),\\displaystyle\-q^\{2n\+1\}\\,x\\,S\_\{n\}\(x;\\,q\)=\(1\-q^\{n\+1\}\)\\,S\_\{n\+1\}\(x;\\,q\)\-\(1\+q\-q^\{n\+1\}\)\\,S\_\{n\}\(x;\\,q\)\+q\\,S\_\{n\-1\}\(x;\\,q\),\(9\)
with initial conditionsS−1​\(x,q\)=0S\_\{\-1\}\(x;\\,q\)=0andS0​\(x,q\)=1S\_\{0\}\(x;\\,q\)=1\.

##### 3\.2\.3\.1 Monic form\.

WritingSn​\(x,q\)=\(−1\)n​qn2\(q,q\)n​pn​\(x\)S\_\{n\}\(x;\\,q\)=\\frac\{\(\-1\)^\{n\}\\,q^\{n^\{2\}\}\}\{\(q;\\,q\)\_\{n\}\}\\,p\_\{n\}\(x\), the monic polynomialspn​\(x\)p\_\{n\}\(x\)satisfy\[[29](https://arxiv.org/html/2610.00050#bib.bib19)\]:

x​pn​\(x\)=pn\+1​\(x\)\+bn​\(q\)​pn​\(x\)\+an2​\(q\)​pn−1​\(x\),x\\,p\_\{n\}\(x\)=p\_\{n\+1\}\(x\)\+b\_\{n\}\(q\)\\,p\_\{n\}\(x\)\+a\_\{n\}^\{2\}\(q\)\\,p\_\{n\-1\}\(x\),\(10\)
where the recurrence coefficients are

bn​\(q\)=q−2​n−32​\(1\+q−qn\+1\),an2​\(q\)=q−4​n​\(1−qn\)\.b\_\{n\}\(q\)=q^\{\-2n\-\\frac\{3\}\{2\}\}\\bigl\(1\+q\-q^\{n\+1\}\\bigr\),\\qquad a\_\{n\}^\{2\}\(q\)=q^\{\-4n\}\\bigl\(1\-q^\{n\}\\bigr\)\.\(11\)
A key computational advantage is that the coefficientsbnb\_\{n\}andan2a\_\{n\}^\{2\}involve only powers ofqq, enabling evaluation ofS0,S1,…,SNS\_\{0\},S\_\{1\},\\ldots,S\_\{N\}in𝒪⁡\(N\)\\mathcal\{O\}\(N\)arithmetic operations per input sample, without special\-function calls\[[29](https://arxiv.org/html/2610.00050#bib.bib19)\]\.

#### 3\.2\.4Orthogonality

The Stieltjes–Wigert polynomials are orthogonal on the semi\-infinite interval\(0,∞\)\(0,\\infty\)\. A distinguishing feature is that the associated moment problem is*indeterminate*\[[30](https://arxiv.org/html/2610.00050#bib.bib10),[27](https://arxiv.org/html/2610.00050#bib.bib29)\], meaning that multiple distinct weight functions yield the same orthogonality relation\.

##### 3\.2\.4\.1 Discrete Weight

With respect to the discrete measure involving theqq\-Pochhammer theta product\[[29](https://arxiv.org/html/2610.00050#bib.bib19)\]:

∫0∞\\displaystyle\\int\_\{0\}^\{\\infty\}Sm​\(x,q\)​Sn​\(x,q\)​d​x\(−x,−q/x;q\)∞=−lnq⋅q−n\(q,q\)∞​\(q,q\)n​δm​n\.\\displaystyle S\_\{m\}\(x;\\,q\)\\,S\_\{n\}\(x;\\,q\)\\,\\frac\{dx\}\{\(\-x,\\,\-q/x;\\,q\)\_\{\\infty\}\}=\\frac\{\-\\ln q\\cdot q^\{\-n\}\}\{\(q;\\,q\)\_\{\\infty\}\\,\(q;\\,q\)\_\{n\}\}\\;\\delta\_\{mn\}\.\(12\)

##### 3\.2\.4\.2 Log\-Normal Weight

An equivalent orthogonality holds with respect to the continuous log\-normal weight function\[[28](https://arxiv.org/html/2610.00050#bib.bib33),[30](https://arxiv.org/html/2610.00050#bib.bib10)\]:

w\(x\)=kπx−1/2exp\(−k2ln2x\),x\>0,w\(x\)=\\frac\{k\}\{\\sqrt\{\\pi\}\}\\,x^\{\-1/2\}\\,\\exp\\\!\\bigl\(\-k^\{2\}\\ln^\{2\}x\\bigr\),\\qquad x\>0,\(13\)whereq=e−1/\(2k2\)q=e^\{\-1/\(2k^\{2\}\)\}\. Under this weight:

∫0∞Sm​\(x,q\)​Sn​\(x,q\)​w​\(x\)​𝑑x=hn​δm​n,\\int\_\{0\}^\{\\infty\}S\_\{m\}\(x;\\,q\)\\,S\_\{n\}\(x;\\,q\)\\,w\(x\)\\,dx=h\_\{n\}\\,\\delta\_\{mn\},\(14\)wherehn\>0h\_\{n\}\>0is the normalization constant\. The moments of the log\-normal weight are given by\[[30](https://arxiv.org/html/2610.00050#bib.bib10)\]:

μn=∫0∞xnw\(x\)dx=q−n\(n\+1\)/2,n=0,1,2,…\\mu\_\{n\}=\\int\_\{0\}^\{\\infty\}x^\{n\}\\,w\(x\)\\,dx=q^\{\-n\(n\+1\)/2\},\\qquad n=0,1,2,\\ldots\(15\)
The indeterminacy of the moment problem means that infinitely many distinct positive measures on\(0,∞\)\(0,\\infty\)share the same moment sequence\{μn\}n≥0\\\{\\mu\_\{n\}\\\}\_\{n\\geq 0\}\[[30](https://arxiv.org/html/2610.00050#bib.bib10)\]\. This is a consequence of the failure of Carleman’s condition\[[31](https://arxiv.org/html/2610.00050#bib.bib16),[32](https://arxiv.org/html/2610.00050#bib.bib30)\]:

∑n=1∞μn−1/\(2n\)<∞\.\\sum\_\{n=1\}^\{\\infty\}\\mu\_\{n\}^\{\-1/\(2n\)\}<\\infty\.\(16\)

### 3\.3Justification for the Stieltjes–Wigert Basis vs\. Classical Alternatives

While classical orthogonal families such as Hermite \(defined onℝ\\mathbb\{R\}\) or Laguerre \(defined on\(0,∞\)\(0,\\infty\)\) might appear as intuitive alternatives that either bypass or simplify the domain\-mapping requirement, the Stieltjes–Wigert family offers three critical mathematical advantages for parameterizing network edges\.

First, the underlying weight structures differ significantly\. Hermite polynomials are orthogonal with respect to a rapidly decaying Gaussian weight, and Laguerre polynomials with respect to an exponential weight\. In contrast, Stieltjes–Wigert polynomials are orthogonal with respect to a continuous log\-normal weight function \(see Equation[13](https://arxiv.org/html/2610.00050#S3.E13)\)\. This heavy\-tailed weight structure is inherently better suited to capture the skewed, heavy\-tailed activation distributions frequently observed in deep neural networks, mitigating the risk of premature saturation for large\-magnitude inputs\.

Second, unlike rigid classical families, the Stieltjes–Wigert basis is aqq\-deformed family\. The parameterq∈\(0,1\)q\\in\(0,1\)acts as an additional, learnable degree of freedom\. As demonstrated in Equation[9](https://arxiv.org/html/2610.00050#S3.E9),qqdynamically tunes the spacing of roots and the oscillatory frequency of the basis functions\. This allows the network to adapt its local representation resolution to the empirical data distribution during gradient descent, a level of geometric adaptability entirely absent in standard Hermite or Laguerre polynomials\.

Finally, the indeterminate nature of the Stieltjes–Wigert moment problem implies that the basis does not strictly commit to a single unique measure\. From a learning perspective, this indeterminacy acts as an implicit regularizer, providing a smoother optimization landscape and reducing the risk of overfitting \(Runge’s phenomenon\) which is a known vulnerability when deploying high\-degree classical polynomials on unbounded domains\.

### 3\.4Domain Handling

The Stieltjes–Wigert polynomials are naturally defined on the semi\-infinite domain\(0,∞\)\(0,\\infty\)\[[27](https://arxiv.org/html/2610.00050#bib.bib29),[28](https://arxiv.org/html/2610.00050#bib.bib33),[29](https://arxiv.org/html/2610.00050#bib.bib19)\]\. However, in a neural network setting, layer inputsξ\\xitypically span the entire real lineℝ\\mathbb\{R\}, or at least an unbounded subset thereof\. To bridge this mismatch, we employ an invertible mappingϕ:ℝ→\(0,∞\)\\phi:\\mathbb\{R\}\\to\(0,\\infty\)that transforms arbitrary real\-valued inputs into the support of the polynomials before evaluation\[[33](https://arxiv.org/html/2610.00050#bib.bib28),[34](https://arxiv.org/html/2610.00050#bib.bib4)\]\.

Several classes of smooth, invertible mappings fromℝ\\mathbb\{R\}\(or a subset\) to\(0,∞\)\(0,\\infty\)have been explored in the spectral methods literature\. For mappings that first target the interval\(−1,1\)\(\-1,1\)as an intermediate step, common choices with a scale parameterι\>0\\iota\>0include:

logarithmic:ψlog\(ξ;ι\)\\displaystyle\\text\{logarithmic:\}\\quad\\psi\_\{\\log\}\(\\xi;\\iota\)=2​tanh⁡\(ξι\)−1,\\displaystyle=2\\tanh\\\!\\left\(\\frac\{\\xi\}\{\\iota\}\\right\)\-1,\(17\)algebraic:ψalg\(ξ;ι\)\\displaystyle\\text\{algebraic:\}\\quad\\psi\_\{\\mathrm\{alg\}\}\(\\xi;\\iota\)=ξ−ιξ\+ι,\\displaystyle=\\frac\{\\xi\-\\iota\}\{\\xi\+\\iota\},\(18\)exponential:ψexp\(ξ;ι\)\\displaystyle\\text\{exponential:\}\\quad\\psi\_\{\\exp\}\(\\xi;\\iota\)=1−2​exp⁡\(−ξι\)\.\\displaystyle=1\-2\\exp\\\!\\left\(\-\\frac\{\\xi\}\{\\iota\}\\right\)\.\(19\)For the infinite domainΩ=\(−∞,∞\)\\Omega=\(\-\\infty,\\infty\), typical nonlinear mappings directly to\(−1,1\)\(\-1,1\)are:

logarithmic:ψlog∞\(ξ;ι\)\\displaystyle\\text\{logarithmic:\}\\quad\\psi\_\{\\log\}^\{\\infty\}\(\\xi;\\iota\)=tanh⁡\(ξι\),\\displaystyle=\\tanh\\\!\\left\(\\frac\{\\xi\}\{\\iota\}\\right\),\(20\)algebraic:ψalg∞\(ξ;ι\)\\displaystyle\\text\{algebraic:\}\\quad\\psi\_\{\\mathrm\{alg\}\}^\{\\infty\}\(\\xi;\\iota\)=ξξ2\+ι2\.\\displaystyle=\\frac\{\\xi\}\{\\sqrt\{\\xi^\{2\}\+\\iota^\{2\}\}\}\.\(21\)A simple way to obtain a mapping into\(0,∞\)\(0,\\infty\)is then to compose one of the above transforms with an exponential lift, e\.g\.

ϕ⁡\(ξ\)=exp⁡\(γ​ψlog∞​\(ξ,ι\)\),γ\>0,\\phi\(\\xi\)=\\exp\\\!\\bigl\(\\gamma\\,\\psi\_\{\\log\}^\{\\infty\}\(\\xi;\\iota\)\\bigr\),\\qquad\\gamma\>0,\(22\)which mapsℝ\\mathbb\{R\}smoothly and monotonically onto the bounded interval\(e−γ,eγ\)⊂\(0,∞\)\(e^\{\-\\gamma\},e^\{\\gamma\}\)\\subset\(0,\\infty\)\. This choice guarantees strict positivity, bounds the effective domain to avoid numerical overflow in the recurrence, and provides well\-behaved gradients during training\[[33](https://arxiv.org/html/2610.00050#bib.bib28)\]\. In our experimental configuration, we set the domain mapping hyperparameters toγ=2\\gamma=2andι=1\\iota=1\. While these parameters could theoretically be treated as learnable weights, this specific static configuration is mathematically motivated by the root distribution of the Stieltjes–Wigert polynomials\. Settingγ=2\\gamma=2explicitly bounds the transformed inputs within the compact interval\(e−2,e2\)≈\(0\.135,7\.389\)\(e^\{\-2\},e^\{2\}\)\\approx\(0\.135,7\.389\)\. This interval is critical because it strategically encapsulates the primary region of oscillation \(i\.e\., the location of the roots\) for low\-to\-moderate degree Stieltjes–Wigert polynomials\. By keeping the inputs strictly bounded away from zero and preventing them from reaching extreme positive values, we prevent the polynomial evaluations from entering asymptotic tails where gradients typically vanish or explode\. Furthermore, settingι=1\\iota=1perfectly aligns the active linear regime of thetanh\\tanhfunction with standard normalized layer inputs, ensuring well\-conditioned gradient flow and maximum expressivity during backpropagation without the need for additional grid\-extension overheads\.

You can see Stieltjes\-Wigert polynomials with different values of the parameterqqin Figure[1](https://arxiv.org/html/2610.00050#S3.F1)\.

![Refer to caption](https://arxiv.org/html/2610.00050v1/sw_poly_q03.png)\(a\)q=0\.3q=0\.3
![Refer to caption](https://arxiv.org/html/2610.00050v1/sw_poly_q05.png)\(b\)q=0\.5q=0\.5
![Refer to caption](https://arxiv.org/html/2610.00050v1/sw_poly_q07.png)\(c\)q=0\.7q=0\.7
![Refer to caption](https://arxiv.org/html/2610.00050v1/sw_poly_q08.png)\(d\)q=0\.8q=0\.8
![Refer to caption](https://arxiv.org/html/2610.00050v1/sw_poly_q09.png)\(e\)q=0\.9q=0\.9

Figure 1:Stieltjes–Wigert polynomialsS0S\_\{0\}–S5S\_\{5\}for different values ofqq\.
### 3\.5Stieltjes–Wigert KAN

Kolmogorov–Arnold Networks \(KANs\) replace fixed activation functions on nodes with learnable univariate functions on edges, yielding compact and interpretable models for function approximation, symbolic regression, and scientific computing tasks\[[2](https://arxiv.org/html/2610.00050#bib.bib22),[6](https://arxiv.org/html/2610.00050#bib.bib8),[7](https://arxiv.org/html/2610.00050#bib.bib2),[13](https://arxiv.org/html/2610.00050#bib.bib9)\]\. In this section we introduce the Stieltjes–Wigert KAN \(SW–KAN\), a KAN variant whose edge activations are expanded in the Stieltjes–Wigert polynomials discussed in Section[3\.2](https://arxiv.org/html/2610.00050#S3.SS2)\[[29](https://arxiv.org/html/2610.00050#bib.bib19),[31](https://arxiv.org/html/2610.00050#bib.bib16),[30](https://arxiv.org/html/2610.00050#bib.bib10)\]\.

#### 3\.5\.1Edge Activation via Stieltjes–Wigert Polynomials

Consider a KAN layer mapping𝐱\(ℓ\)∈ℝnℓ\\mathbf\{x\}^\{\(\\ell\)\}\\in\\mathbb\{R\}^\{n\_\{\\ell\}\}to𝐱\(ℓ\+1\)∈ℝnℓ\+1\\mathbf\{x\}^\{\(\\ell\+1\)\}\\in\\mathbb\{R\}^\{n\_\{\\ell\+1\}\}\. For each edge\(ℓ,i\)→\(ℓ\+1,j\)\(\\ell,i\)\\to\(\\ell\+1,j\), we parameterize the learnable univariate activation function using the Stieltjes–Wigert basis as

φj,i\(ℓ\)​\(x\)\\displaystyle\\varphi\_\{j,i\}^\{\(\\ell\)\}\(x\)=∑n=0Ncj,i,n\(ℓ\)Sn\(ϕ\(x\);q\(ℓ\)\),i=1,…,nℓ,j=1,…,nℓ\+1,\\displaystyle=\\sum\_\{n=0\}^\{N\}c\_\{j,i,n\}^\{\(\\ell\)\}\\,S\_\{n\}\\\!\\bigl\(\\phi\(x\);\\,q^\{\(\\ell\)\}\\bigr\),\\qquad i=1,\\dots,n\_\{\\ell\},\\;j=1,\\dots,n\_\{\\ell\+1\},\(23\)where:

- •Sn​\(⋅,q\)S\_\{n\}\(\\cdot;q\)are the Stieltjes–Wigert polynomials of degreennwith parameter0<q<10<q<1\(Section[3\.2](https://arxiv.org/html/2610.00050#S3.SS2)\)\[[27](https://arxiv.org/html/2610.00050#bib.bib29),[28](https://arxiv.org/html/2610.00050#bib.bib33),[29](https://arxiv.org/html/2610.00050#bib.bib19)\];
- •ϕ:ℝ→\(0,∞\)\\phi:\\mathbb\{R\}\\to\(0,\\infty\)is a smooth mapping that matches the natural domain of the Stieltjes–Wigert polynomials \(Section[3\.4](https://arxiv.org/html/2610.00050#S3.SS4)\)\[[33](https://arxiv.org/html/2610.00050#bib.bib28),[34](https://arxiv.org/html/2610.00050#bib.bib4)\];
- •cj,i,n\(ℓ\)∈ℝc\_\{j,i,n\}^\{\(\\ell\)\}\\in\\mathbb\{R\}are learnable expansion coefficients, one set per edge\(j,i\)\(j,i\)and per degreenn;
- •q\(ℓ\)∈\(0,1\)q^\{\(\\ell\)\}\\in\(0,1\)is a single learnable parameter*shared across all edges of layerℓ\\ell*, controlling the shape of the basis functions used by that entire layer\.

A convenient choice of input mapping is the exponential\-of\-tanh transform\[[33](https://arxiv.org/html/2610.00050#bib.bib28)\]:

ϕ⁡\(ξ\)=exp⁡\(γ​tanh⁡\(ξ/ι\)\),γ\>0,ι\>0,\\phi\(\\xi\)=\\exp\\\!\\bigl\(\\gamma\\tanh\(\\xi/\\iota\)\\bigr\),\\qquad\\gamma\>0,\\;\\iota\>0,\(24\)which mapsℝ\\mathbb\{R\}smoothly and monotonically onto\(e−γ,eγ\)⊂\(0,∞\)\(e^\{\-\\gamma\},e^\{\\gamma\}\)\\subset\(0,\\infty\), ensuring strict positivity, boundedness, and stable gradients during backpropagation\.

#### 3\.5\.2Recurrence\-Based Evaluation

Directly evaluating the explicit sum in the hypergeometric definition ofSn​\(x,q\)S\_\{n\}\(x;q\)for alln=0,…,Nn=0,\\dots,Nat each edge would be computationally expensive\. Instead, we exploit the three\-term recurrence relation for the monic Stieltjes–Wigert polynomials from Koekoek–Lesky–Swarttouw\[[29](https://arxiv.org/html/2610.00050#bib.bib19),[21](https://arxiv.org/html/2610.00050#bib.bib18)\]:

x​Sn​\(x,q\)\\displaystyle x\\,S\_\{n\}\(x;q\)=Sn\+1​\(x,q\)\+bn​\(q\)​Sn​\(x,q\)\+an2​\(q\)​Sn−1​\(x,q\),\\displaystyle=S\_\{n\+1\}\(x;q\)\+b\_\{n\}\(q\)\\,S\_\{n\}\(x;q\)\+a\_\{n\}^\{2\}\(q\)\\,S\_\{n\-1\}\(x;q\),\(25\)n≥0,\\displaystyle n\\geq 0,with initial conditionsS−1​\(x,q\)=0S\_\{\-1\}\(x;q\)=0andS0​\(x,q\)=1S\_\{0\}\(x;q\)=1, and recurrence coefficients

bn​\(q\)=q−2​n−3​1\+q−qn\+12,an2​\(q\)=q−4​n​\(1−qn\)\.b\_\{n\}\(q\)=q^\{\-2n\-3\}\\,\\frac\{1\+q\-q^\{n\+1\}\}\{2\},\\qquad a\_\{n\}^\{2\}\(q\)=q^\{\-4n\}\\,\(1\-q^\{n\}\)\.\(26\)Equivalently,

Sn\+1​\(x,q\)=x​Sn​\(x,q\)−bn​\(q\)​Sn​\(x,q\)−an2​\(q\)​Sn−1​\(x,q\)\.S\_\{n\+1\}\(x;q\)=x\\,S\_\{n\}\(x;q\)\-b\_\{n\}\(q\)\\,S\_\{n\}\(x;q\)\-a\_\{n\}^\{2\}\(q\)\\,S\_\{n\-1\}\(x;q\)\.\(27\)
For each edge and each input valuez=ϕ⁡\(x\)z=\\phi\(x\), we computeS0​\(z,q\),…,SN​\(z,q\)S\_\{0\}\(z;q\),\\dots,S\_\{N\}\(z;q\)by iterating \([27](https://arxiv.org/html/2610.00050#S3.E27)\), which requires only𝒪⁡\(N\)\\mathcal\{O\}\(N\)arithmetic operations and powers ofqq, without special\-function calls\. This makes SW–KAN significantly cheaper than KAN variants based on heavier hypergeometric families such as Askey–Wilson polynomials\[[29](https://arxiv.org/html/2610.00050#bib.bib19),[5](https://arxiv.org/html/2610.00050#bib.bib27)\]\.

#### 3\.5\.3Numerically Stable Evaluation via an Orthonormal Recurrence

While \([27](https://arxiv.org/html/2610.00050#S3.E27)\) is algebraically exact, iterating the*monic*recurrence directly in finite\-precision arithmetic is numerically unusable beyond a modest degree\. The coefficientan2​\(q\)=q−4​n​\(1−qn\)a\_\{n\}^\{2\}\(q\)=q^\{\-4n\}\(1\-q^\{n\}\)in \([26](https://arxiv.org/html/2610.00050#S3.E26)\) grows asq−4​nq^\{\-4n\}, so for anyqqnot extremely close to11the monic polynomialsSn​\(z,q\)S\_\{n\}\(z;q\)inflate by many orders of magnitude within a handful of recursion steps \(e\.g\.,q=0\.35q=0\.35andn=6n=6already yieldsq−4​n≈3×1010q^\{\-4n\}\\approx 3\\times 10^\{10\}\)\. Becausean2​\(q\)a\_\{n\}^\{2\}\(q\)enters the recurrence multiplicatively, this growth compounds across degrees and routinely producesinf/NaNactivations during training well before the target polynomial degreeNNis reached, regardless of how the input is scaled byϕ\\phiin \([24](https://arxiv.org/html/2610.00050#S3.E24)\)\.

We resolve this by evaluating the mathematically equivalent*orthonormal*form of the same three\-term recurrence rather than the monic form\. Writingan​\(q\)=an2​\(q\)a\_\{n\}\(q\)=\\sqrt\{a\_\{n\}^\{2\}\(q\)\}witha0​\(q\):=0a\_\{0\}\(q\):=0, we defineTn​\(z,q\):=Sn​\(z,q\)/hnT\_\{n\}\(z;q\):=S\_\{n\}\(z;q\)/\\sqrt\{h\_\{n\}\}, wherehn\>0h\_\{n\}\>0is the norm constant implied by the orthogonality relation in \([14](https://arxiv.org/html/2610.00050#S3.E14)\)\. The sequence\{Tn\}\\\{T\_\{n\}\\\}satisfies

Tn\+1​\(z,q\)=\(z−bn​\(q\)\)​Tn​\(z,q\)−an​\(q\)​Tn−1​\(z,q\)an\+1​\(q\),T\_\{n\+1\}\(z;q\)=\\frac\{\\bigl\(z\-b\_\{n\}\(q\)\\bigr\)\\,T\_\{n\}\(z;q\)\-a\_\{n\}\(q\)\\,T\_\{n\-1\}\(z;q\)\}\{a\_\{n\+1\}\(q\)\},\(28\)with initial conditionsT−1​\(z,q\)=0T\_\{\-1\}\(z;q\)=0andT0​\(z,q\)=1T\_\{0\}\(z;q\)=1, using the same coefficientsbn​\(q\)b\_\{n\}\(q\)andan​\(q\)a\_\{n\}\(q\)as in \([26](https://arxiv.org/html/2610.00050#S3.E26)\)\. This is the standard normalization used to generate orthogonal polynomials stably from Jacobi recurrence coefficients, underlying, e\.g\., the node/weight computation in Gaussian quadrature\[[31](https://arxiv.org/html/2610.00050#bib.bib16),[29](https://arxiv.org/html/2610.00050#bib.bib19)\]\. Rather than multiplying by the unboundedan2​\(q\)a\_\{n\}^\{2\}\(q\)at each step, \([28](https://arxiv.org/html/2610.00050#S3.E28)\)*divides*byan\+1​\(q\)a\_\{n\+1\}\(q\), which keepsTn​\(z,q\)T\_\{n\}\(z;q\)of order𝒪⁡\(1\)\\mathcal\{O\}\(1\)across the working domain\(e−γ,eγ\)\(e^\{\-\\gamma\},e^\{\\gamma\}\)for the full range of admissibleqqand degreeNN, analogous to how Chebyshev polynomials remain bounded in\[−1,1\]\[\-1,1\]\.

Crucially, replacingSnS\_\{n\}withTnT\_\{n\}in the edge activation \([23](https://arxiv.org/html/2610.00050#S3.E23)\) does not alter the expressiveness of the model:TnT\_\{n\}differs fromSnS\_\{n\}only by the fixed positive scalar1/hn1/\\sqrt\{h\_\{n\}\}, so\{Tn\}n=0N\\\{T\_\{n\}\\\}\_\{n=0\}^\{N\}spans exactly the same degree\-NNpolynomial subspace as\{Sn\}n=0N\\\{S\_\{n\}\\\}\_\{n=0\}^\{N\}\. Since the expansion coefficientscj,i,n\(ℓ\)c\_\{j,i,n\}^\{\(\\ell\)\}are learned freely, this rescaling is absorbed automatically during training and the edge activation retains the form

φj,i\(ℓ\)​\(x\)=∑n=0Ncj,i,n\(ℓ\)​Tn​\(ϕ⁡\(x\),q\(ℓ\)\)\.\\varphi\_\{j,i\}^\{\(\\ell\)\}\(x\)=\\sum\_\{n=0\}^\{N\}c\_\{j,i,n\}^\{\(\\ell\)\}\\,T\_\{n\}\\\!\\bigl\(\\phi\(x\);\\,q^\{\(\\ell\)\}\\bigr\)\.\(29\)In practice we additionally add a smallε=10−8\\varepsilon=10^\{\-8\}under the square root,an​\(q\)=an2​\(q\)\+εa\_\{n\}\(q\)=\\sqrt\{a\_\{n\}^\{2\}\(q\)\+\\varepsilon\}, to avoid division by zero whenqqis close to11andan2​\(q\)→0a\_\{n\}^\{2\}\(q\)\\to 0\. All SW–KAN results reported in Section[4](https://arxiv.org/html/2610.00050#S4)use \([28](https://arxiv.org/html/2610.00050#S3.E28)\) rather than the raw monic recurrence \([27](https://arxiv.org/html/2610.00050#S3.E27)\); the latter is presented in Section[3\.5\.2](https://arxiv.org/html/2610.00050#S3.SS5.SSS2)for consistency with the classical definition of the Stieltjes–Wigert polynomials, but is not what is executed at training or inference time\.

#### 3\.5\.4Learnable Parameters

For a single SW–KAN layer withnℓn\_\{\\ell\}input nodes andnℓ\+1n\_\{\\ell\+1\}output nodes, and maximum polynomial degreeNN, the total number of learnable parameters is

nℓ​nℓ\+1​\(N\+1\)​\+⏟layer\-wise basis parameter​1=nℓ​nℓ\+1​\(N\+1\)\+1,n\_\{\\ell\}n\_\{\\ell\+1\}\(N\+1\)\\;\\;\\underbrace\{\+\}\_\{\\text\{layer\-wise basis parameter\}\}\\;\\;1\\;=\\;n\_\{\\ell\}n\_\{\\ell\+1\}\(N\+1\)\+1,\(30\)where:

- •nℓ​nℓ\+1​\(N\+1\)n\_\{\\ell\}n\_\{\\ell\+1\}\(N\+1\)parameters correspond to the expansion coefficientscj,i,n\(ℓ\)c\_\{j,i,n\}^\{\(\\ell\)\}, i\.e\.N\+1N\+1coefficients for each of thenℓ​nℓ\+1n\_\{\\ell\}n\_\{\\ell\+1\}edges;
- •the additional single parameterq\(ℓ\)q^\{\(\\ell\)\}is*shared across all edges of the layer*rather than defined per edge, controlling the shape of the Stieltjes–Wigert basis functions for that layer as a whole\.

To enforce the constraintq\(ℓ\)∈\(0,1\)q^\{\(\\ell\)\}\\in\(0,1\)during training,q\(ℓ\)q^\{\(\\ell\)\}is obtained from a single unconstrained parameterq^\(ℓ\)∈ℝ\\hat\{q\}^\{\(\\ell\)\}\\in\\mathbb\{R\}per layer via a sigmoid:

q\(ℓ\)=σ⁡\(q^\(ℓ\)\),σ⁡\(t\)=11\+e−t\.q^\{\(\\ell\)\}=\\sigma\\\!\\bigl\(\\hat\{q\}^\{\(\\ell\)\}\\bigr\),\\qquad\\sigma\(t\)=\\frac\{1\}\{1\+e^\{\-t\}\}\.\(31\)This simple parameterization keepsq\(ℓ\)q^\{\(\\ell\)\}in the admissible range and allows gradient\-based optimization to adjust it jointly with the coefficientscj,i,n\(ℓ\)c\_\{j,i,n\}^\{\(\\ell\)\}\. Sharing a singleq\(ℓ\)q^\{\(\\ell\)\}per layer, rather than instantiating one per edge, keeps the parameter overhead of theqq\-deformation negligible \(one extra scalar per layer\) while still allowing each layer to adapt the concentration and oscillatory frequency of its Stieltjes–Wigert basis to the empirical distribution of its inputs\.

#### 3\.5\.5Network Computation

Given an input vector𝐱\(ℓ\)∈ℝnℓ\\mathbf\{x\}^\{\(\\ell\)\}\\in\\mathbb\{R\}^\{n\_\{\\ell\}\}, a SW–KAN layer first maps each componentxi\(ℓ\)x\_\{i\}^\{\(\\ell\)\}to the semi\-infinite domain viaϕ\\phiin \([24](https://arxiv.org/html/2610.00050#S3.E24)\), then evaluates the Stieltjes–Wigert expansion on every edge\. The output of neuronjjin layerℓ\+1\\ell\+1is

xj\(ℓ\+1\)\\displaystyle x\_\{j\}^\{\(\\ell\+1\)\}=∑i=1nℓφj,i\(ℓ\)​\(xi\(ℓ\)\)=∑i=1nℓ∑n=0Ncj,i,n\(ℓ\)​Sn​\(ϕ⁡\(xi\(ℓ\)\),q\(ℓ\)\),\\displaystyle=\\sum\_\{i=1\}^\{n\_\{\\ell\}\}\\varphi\_\{j,i\}^\{\(\\ell\)\}\\\!\\bigl\(x\_\{i\}^\{\(\\ell\)\}\\bigr\)=\\sum\_\{i=1\}^\{n\_\{\\ell\}\}\\sum\_\{n=0\}^\{N\}c\_\{j,i,n\}^\{\(\\ell\)\}\\,S\_\{n\}\\\!\\bigl\(\\phi\(x\_\{i\}^\{\(\\ell\)\}\);\\,q^\{\(\\ell\)\}\\bigr\),\(32\)j=1,…,nℓ\+1,\\displaystyle j=1,\\dots,n\_\{\\ell\+1\},which is a direct specialization of the general KAN layer formula with a polynomial basis\[[2](https://arxiv.org/html/2610.00050#bib.bib22),[6](https://arxiv.org/html/2610.00050#bib.bib8),[5](https://arxiv.org/html/2610.00050#bib.bib27)\]\.

StackingLLsuch layers yields the full SW–KAN:

SW\-KAN\(𝐱\)=Φ\(L−1\)∘Φ\(L−2\)∘⋯∘Φ\(0\)\(𝐱\),\\mathrm\{SW\\text\{\-\}KAN\}\(\\mathbf\{x\}\)=\\Phi^\{\(L\-1\)\}\\circ\\Phi^\{\(L\-2\)\}\\circ\\cdots\\circ\\Phi^\{\(0\)\}\(\\mathbf\{x\}\),\(33\)where eachΦ\(ℓ\)\\Phi^\{\(\\ell\)\}is the function\-valued matrix\(φj,i\(ℓ\)\)j,i\\bigl\(\\varphi\_\{j,i\}^\{\(\\ell\)\}\\bigr\)\_\{j,i\}parameterized via Stieltjes–Wigert polynomials as in \([23](https://arxiv.org/html/2610.00050#S3.E23)\)\. This architecture can be trained using standard stochastic gradient methods, with the recurrence\-based evaluation ensuring that the additional expressivity of the Stieltjes–Wigert basis does not incur prohibitive computational cost\[[29](https://arxiv.org/html/2610.00050#bib.bib19),[2](https://arxiv.org/html/2610.00050#bib.bib22)\]\.

### 3\.6Theoretical Computational Complexity and Polynomial Degree

To rigorously establish the computational efficiency of SW–KAN, we must quantify the cost of the three\-term recurrence relation compared to the original B\-spline formulation\. Evaluating a B\-spline of degreekkover a grid ofGGknots via the Cox–de Boor recursion requires𝒪⁡\(k​G\)\\mathcal\{O\}\(kG\)arithmetic operations per edge\. Moreover, B\-splines necessitate dynamic grid\-extension algorithms during training to ensure the knots dynamically cover the shifting empirical activation distributions\. This injects significant non\-differentiable overhead and memory fragmentation into the training pipeline\.

In stark contrast, the Stieltjes–Wigert forward pass relies entirely on the recurrence relation in Equation \([27](https://arxiv.org/html/2610.00050#S3.E27)\) \(evaluated in its numerically stable orthonormal form, Equation \([28](https://arxiv.org/html/2610.00050#S3.E28)\)\)\. Since the coefficientsbn​\(q\)b\_\{n\}\(q\)andan​\(q\)a\_\{n\}\(q\)depend solely on the parameterqqand can be updated in𝒪⁡\(1\)\\mathcal\{O\}\(1\)time, evaluating the basis up to degreeNNstrictly requires𝒪⁡\(N\)\\mathcal\{O\}\(N\)multiplications and additions per edge\. This yields a deterministic𝒪⁡\(N\)\\mathcal\{O\}\(N\)time complexity that trivially parallelizes on modern GPU architectures, completely bypassing knot management and grid updates\.

For our medium SW–KAN configuration yielding105,856105\{,\}856parameters, the maximum polynomial degree is constrained toN=3N=3\. From an approximation\-theoretic standpoint,N=3N=3provides sufficient representational capacity \(i\.e\., cubic non\-linearity with up to two inflection points\) to capture complex, non\-monotonic feature interactions\. Simultaneously, restricting the basis to a low degree explicitly prevents the high\-frequency oscillatory artifacts \(Runge’s phenomenon\) that typically degrade the generalization of higher\-degree polynomial networks, thereby explaining the robust accuracy achieved without requiring massive parameter counts\.

## 4Experiments and Results

In this section, we evaluate the proposed Stieltjes–Wigert Kolmogorov–Arnold Network \(SW–KAN\) on the MNIST handwritten digit classification benchmark and compare it against a broad family of polynomial\-based KAN models\. Our goal is to assess whether replacing B\-spline edge activations with Stieltjes–Wigertqq\-orthogonal polynomials yields a better trade\-off between accuracy, model size, and training time than existing polynomial KAN variants\. We further investigate SW\-KAN’s performance on the more challenging Fashion\-MNIST dataset and its capability in approximating continuous multivariate functions, demonstrating its generalization across diverse tasks\.

### 4\.1Digit Classification on MNIST

The MNIST dataset consists of60,00060\{,\}000training and10,00010\{,\}000test grayscale images of handwritten digits00–99, each of size28×2828\\times 28\. We use the standard train/test split without data augmentation\. Input images are flattened to784784\-dimensional vectors and linearly normalized to\[0,1\]\[0,1\]\. The task is to predict the correct digit label for each image\.

#### 4\.1\.1Experimental Setup

We adopt a fully connected KAN backbone with three hidden KAN layers followed by a linear classification layer, closely following the architecture used in prior work on polynomial\-based KANs\[[2](https://arxiv.org/html/2610.00050#bib.bib22),[5](https://arxiv.org/html/2610.00050#bib.bib27)\]\. Each hidden layer is implemented as a KAN layer in which every edge\(ℓ,i\)→\(ℓ\+1,j\)\(\\ell,i\)\\to\(\\ell\+1,j\)carries a learnable univariate activation function\. In the SW–KAN model, this activation is parameterized by a Stieltjes–Wigert expansion

ϕj,i\(ℓ\)​\(x\)=∑n=0Ncj,i,n\(ℓ\)​Sn​\(φ⁡\(x\),q\(ℓ\)\),\\phi^\{\(\\ell\)\}\_\{j,i\}\(x\)=\\sum\_\{n=0\}^\{N\}c^\{\(\\ell\)\}\_\{j,i,n\}\\,S\_\{n\}\\\!\\big\(\\varphi\(x\);q^\{\(\\ell\)\}\\big\),whereSn​\(⋅,q\)S\_\{n\}\(\\cdot;q\)are Stieltjes–Wigert polynomials,φ:ℝ→\(0,∞\)\\varphi:\\mathbb\{R\}\\to\(0,\\infty\)is the exponential\-of\-tanh mapping from Section[3\.5](https://arxiv.org/html/2610.00050#S3.SS5),cj,i,n\(ℓ\)c^\{\(\\ell\)\}\_\{j,i,n\}are trainable coefficients, andq\(ℓ\)∈\(0,1\)q^\{\(\\ell\)\}\\in\(0,1\)is a single trainable shape parameter shared across all edges of layerℓ\\ell, enforced via a sigmoid reparameterization\.

To study the effect of model size, we consider two configurations:

- •SW–KAN \(medium\)with105,856105\{,\}856trainable parameters, chosen to be comparable to the majority of polynomial KAN baselines\.
- •SW–KAN \(large\)with219,907219\{,\}907parameters, matching the parameter count of the widest baseline \(Gottlieb–KAN\)\.

Both configurations use the same depth and layer widths; the parameter budget is controlled by the maximal polynomial degreeNNand the number of channels per edge\.

All models are trained with the Adam optimizer and cross\-entropy loss, using a batch size of128128, an initial learning rate of10−310^\{\-3\}, and a maximum of5050epochs\. We track the test accuracy at the end of each epoch and report the best test accuracy achieved during training\. Training is performed on a single GPU in a PyTorch implementation\.

#### 4\.1\.2Baselines

As baselines, we include the family of polynomial Kolmogorov–Arnold Networks introduced by Seydi\[[5](https://arxiv.org/html/2610.00050#bib.bib27)\], where the edge activations are expanded in a variety of classical and combinatorial polynomial bases\. Specifically, we consider Fermat–KAN, AlSalam–Carlitz–KAN, Bannai–Ito–KAN, Boas–Buck–KAN, Boubaker–KAN, Charlier–KAN, Gottlieb–KAN, Heptanacci–KAN, Hexanacci–KAN, Meixner–Pollaczek–KAN, Narayana–KAN, Octanacci–KAN, Pado–KAN, Pentanacci–KAN, Tetranacci–KAN, Tribo–KAN, Vieta–Pell–KAN, and Askey–Wilson–KAN\. All of these models are configured to use a similar KAN backbone, differing only in the choice of polynomial basis on each edge and minor adjustments in polynomial degree to keep the parameter count close to10510^\{5\}\(except Gottlieb–KAN\)\. Their MNIST test accuracies and parameter counts are summarized in Table[1](https://arxiv.org/html/2610.00050#S4.T1)\.

Input\(28×\\times28\)SW0\(L1\)SW1\(L1\)SW2\(L1\)LN1\(64\)SW0\(L2\)SW1\(L2\)SW2\(L2\)LN2\(10\)Output\(10\)Figure 2:SW\-KAN architecture for MNIST\.Table 1:Test accuracy, number of parameters, and total training time \(until best epoch\) for polynomial\-based KAN models and SW\-KAN on MNIST\. \(SW\-KAN results are reported at the epoch of best validation accuracy, achieved within 50 epochs\.\)

### 4\.2Fashion\-MNIST Classification

To evaluate the generalization capability and resource efficiency of SW\-KAN beyond the MNIST dataset, we conducted a comparative experiment on the Fashion\-MNIST benchmark\. Fashion\-MNIST consists of 60,000 training and 10,000 test grayscale images of 10 fashion categories, with the same dimensionality as MNIST \(28×2828\\times 28pixels\)\. Unlike the MNIST experiments, we designed this experiment specifically to assess SW\-KAN’s performance under resource constrained conditions that are common in real world applications\. We applied PCA to reduce the input dimensionality from 784 to 40 components \(a 95% reduction\) and subsampled the training set to only 10,000 samples \(from the original 60,000\), simulating scenarios where computational resources or labeled data are limited\. This setup tests the model’s data efficiency and robustness to dimensionality reduction\. We compared SW\-KAN against the standard B\-spline KAN baseline under identical conditions\. Both models share the same architecture: a single hidden layer with 48 neurons followed by a 10\-class output layer\. SW\-KAN uses polynomial degreeN=3N=3, while the B\-spline KAN uses grid size 5 and order 3\. For a fair comparison with the B\-spline baseline, which includes a residual base\-activation path, we likewise equip the SW\-KAN layer here with a SiLU\-based residual path\. All models were trained with the Adam optimizer for 30 epochs with learning rate10−210^\{\-2\}and weight decay10−410^\{\-4\}\.

![Refer to caption](https://arxiv.org/html/2610.00050v1/fashion_mnist_comparison.png)Figure 3:Comparison of SW\-KAN and standard B\-spline KAN on Fashion\-MNIST\. Left: test accuracy and parameter counts\. Right: validation accuracy during training\.As shown in Figure[3](https://arxiv.org/html/2610.00050#S4.F3), SW\-KAN achieves a test accuracy of 83\.10%, outperforming the standard B\-spline KAN, which attains 82\.65%\. Notably, SW\-KAN accomplishes this with significantly fewer parameters \(14,634 vs\. 21,658\), representing a 32% reduction in model size while still delivering superior accuracy\. This result is consistent with our MNIST findings and further confirms that Stieltjes\-Wigert polynomial basis functions offer a more parameter\-efficient representation compared to B\-splines\. The validation accuracy curves in Figure[3](https://arxiv.org/html/2610.00050#S4.F3)\(right\) show that both models converge within approximately 15 epochs, with SW\-KAN reaching its best validation accuracy slightly earlier than the B\-spline KAN\. The B\-spline KAN achieves its highest validation accuracy of 83\.13% at epoch 9, while SW\-KAN reaches 82\.60% at epoch 12\. However, SW\-KAN maintains competitive performance throughout training and achieves superior final test accuracy despite a slightly lower best validation accuracy, suggesting better generalization to the test set\. Most importantly, these results demonstrate SW\-KAN’s robustness under three critical resource\-constrained conditions:

- •Data efficiency: With only 10,000 training samples \(17% of the full dataset\), SW\-KAN maintains competitive performance, indicating strong generalization from limited labeled data\.
- •Dimensionality reduction tolerance: Despite reducing input features from 784 to 40 dimensions \(95% reduction via PCA\), SW\-KAN preserves its representational power, making it suitable for high\-dimensional problems where feature compression is necessary\.
- •Parameter efficiency: With 32% fewer parameters than the B\-spline KAN, SW\-KAN achieves higher accuracy, demonstrating superior parameter utilization\.

### 4\.3Function Approximation on 2D Surfaces

To evaluate SW\-KAN’s capability in learning continuous multivariate functions, we conducted a function approximation experiment on three distinct 2D surfaces on the domain\[−2,2\]2\[\-2,2\]^\{2\}:

f1​\(x,y\)\\displaystyle f\_\{1\}\(x,y\)=sin⁡\(x\)​cos⁡\(y\)\+0\.5​x​y\\displaystyle=\\sin\(x\)\\cos\(y\)\+0\.5xy\(34\)f2​\(x,y\)\\displaystyle f\_\{2\}\(x,y\)=exp⁡\(−x2\+y22\)​\(1\+0\.3​sin⁡\(3​x\)​cos⁡\(2​y\)\)\\displaystyle=\\exp\\left\(\-\\frac\{x^\{2\}\+y^\{2\}\}\{2\}\\right\)\\left\(1\+0\.3\\sin\(3x\)\\cos\(2y\)\\right\)\(35\)f3​\(x,y\)\\displaystyle f\_\{3\}\(x,y\)=x2\+y2\+0\.5​sin⁡\(4​x\)\+0\.3​cos⁡\(4​y\)\\displaystyle=x^\{2\}\+y^\{2\}\+0\.5\\sin\(4x\)\+0\.3\\cos\(4y\)\(36\)
The SW\-KAN architecture consists of a single hidden layer with 16 neurons and polynomial degreeN=4N=4, resulting in only 389 trainable parameters\. Training used 5,000 randomly sampled points for 5,000 epochs with learning rate10−210^\{\-2\}\.

![Refer to caption](https://arxiv.org/html/2610.00050v1/screenshot_879.png)Figure 4:Function approximation results for SW\-KAN on three 2D surfaces\. Each row corresponds to a different target function: \(left\) ground truth, \(middle\) SW\-KAN prediction, \(right\) absolute error\.Figure[4](https://arxiv.org/html/2610.00050#S4.F4)presents the approximation results\. For all three functions, SW\-KAN achieves near\-perfect reconstruction with minimal errors\. The test MSE values are2\.5×10−52\.5\\times 10^\{\-5\}forf1f\_\{1\},2\.9×10−52\.9\\times 10^\{\-5\}forf2f\_\{2\}, and3\.5×10−33\.5\\times 10^\{\-3\}forf3f\_\{3\}\. The error plots confirm that the approximation errors are uniformly distributed across the domain\.

These results demonstrate SW\-KAN’s strong representational power and parameter efficiency, achieving high accuracy across diverse function landscapes with only 389 parameters\. The model successfully captures both smooth sinusoidal surfaces and multi\-scale oscillatory structures, validating its suitability for scientific computing applications requiring accurate function approximation\.

## 5Conclusion and Discussion

In this paper, we introduced the Stieltjes–Wigert Kolmogorov–Arnold Network \(SW–KAN\), a novel architecture that integratesqq\-orthogonal polynomials into the KAN framework to address the computational bottlenecks associated with traditional B\-spline edge activations\. To overcome the structural challenge of mapping unbounded real\-valued inputs \(ℝ\\mathbb\{R\}\) to the semi\-infinite domain\(0,∞\)\(0,\\infty\)of the Stieltjes–Wigert polynomials, we proposed a smooth exponential\-of\-tanh transformation\. Furthermore, by leveraging a numerically stable, orthonormalized three\-term recurrence, we formulated a highly efficient evaluation strategy that computes polynomial expansions in𝒪⁡\(N\)\\mathcal\{O\}\(N\)operations without relying on computationally expensive special\-function calls\.

### 5\.1Discussion of Experimental Findings

Our empirical evaluations demonstrate that SW–KAN achieves a superior accuracy\-efficiency trade\-off compared to existing polynomial\-based KANs, and that this advantage generalizes across classification and function\-approximation tasks rather than being an artifact of a single benchmark\.

On the MNIST handwritten digit classification benchmark, the medium configuration of SW–KAN reached97\.53%97\.53\\%test accuracy with105,856105\{,\}856parameters, already surpassing the best\-performing baseline of comparable size, Vieta–Pell\-KAN \(97\.49%97\.49\\%with105,856105\{,\}856parameters\)\. The large configuration further improved the accuracy to98\.24%98\.24\\%using219,907219\{,\}907parameters, exactly matching Gottlieb\-KAN’s parameter count yet achieving a higher accuracy \(97\.59%97\.59\\%\)\. Taken together, both the medium and large SW–KAN configurations outperform all 18 established polynomial KAN baselines at their respective parameter budgets\.

On Fashion\-MNIST, under a deliberately resource\-constrained setting \(input features reduced from 784 to 40 dimensions via PCA and training limited to10,00010\{,\}000samples\), SW\-KAN reached83\.10%83\.10\\%test accuracy with14,63414\{,\}634parameters, outperforming a standard B\-spline KAN of the same architecture \(82\.65%82\.65\\%with21,65821\{,\}658parameters, a32%32\\%larger model\)\. This indicates that the parameter efficiency observed on MNIST persists, and if anything becomes more pronounced, when training data and feature dimensionality are scarce\.

On continuous multivariate function approximation over three qualitatively different 2D surfaces, a SW\-KAN model with only389389parameters achieved test MSE values as low as2\.5×10−52\.5\\times 10^\{\-5\}–2\.9×10−52\.9\\times 10^\{\-5\}for the smoother targets and3\.5×10−33\.5\\times 10^\{\-3\}for the more oscillatory target, confirming that the Stieltjes–Wigert basis is expressive enough for accurate scientific\-computing\-style regression at a very small parameter budget\.

Taken together, the performance delta over baselines cannot be attributed solely to parameter capacity, since the large SW–KAN matches Gottlieb\-KAN’s parameter count exactly and the Fashion\-MNIST SW\-KAN model uses fewer parameters than its B\-spline counterpart\. Instead, this superiority is rooted in the mathematical properties of the basis functions and their interaction with gradient\-based optimization\. Classical orthogonal polynomials are typically defined on strictly bounded intervals \(e\.g\.,\[−1,1\]\[\-1,1\]\) or discrete domains\. When standard normalizations \(liketanh\\tanhor sigmoid\) are used to map unbounded network activations into these fixed domains, the inputs frequently saturate at the boundaries, leading to severe gradient attenuation \(vanishing gradients\) for heavy\-tailed distributions\. Conversely, the Stieltjes–Wigert polynomials are naturally supported on\(0,∞\)\(0,\\infty\)\. By employing the smooth exponential\-of\-tanh mapping, inputs are projected into a safe, strictly positive sub\-interval that completely avoids zero and extreme asymptotic tails\.

Furthermore, unlike rigid classical families, the learnableqqparameter in SW–KAN acts as a dynamic frequency modulator during backpropagation\. This geometric flexibility, combined with the log\-normal weight structure, yields a smoother and more well\-conditioned loss landscape, which is consistent with the model’s competitive results across image classification and continuous function approximation despite comparable or smaller parameter budgets than its baselines\.

### 5\.2Limitations and Future Work

Despite the promising results, this study exhibits certain limitations that provide avenues for future research\. First, while our empirical validation now spans MNIST, Fashion\-MNIST under resource\-constrained conditions, and continuous function approximation on 2D surfaces, it remains focused on relatively small\-scale, low\-to\-moderate\-dimensional tasks\. Evaluating SW–KAN on larger, higher\-dimensional problems, such as ImageNet\-scale image classification, dense time\-series modeling, and physics\-informed neural network \(PINN\) applications, is necessary to fully assess its scalability\.

Second, the domain mapping parameters \(e\.g\., the scaling factors in the exponential\-of\-tanh function\) were treated as fixed hyperparameters\. Future iterations of SW–KAN could benefit from making these domain\-mapping parameters fully learnable, enabling dynamic, layer\-specific distribution matching\. A related direction is to relax the current layer\-wise sharing ofq\(ℓ\)q^\{\(\\ell\)\}\(Section[3\.5\.4](https://arxiv.org/html/2610.00050#S3.SS5.SSS4)\) toward finer\-grained, e\.g\. edge\-wise or group\-wise,qqparameterizations, and to characterize the resulting accuracy\-versus\-parameter\-count trade\-off relative to the layer\-shared design studied here\.

Finally, a distinguishing mathematical feature of the Stieltjes–Wigert polynomials is their indeterminate moment problem, meaning multiple distinct weight measures yield the same orthogonality relation\. The theoretical implications of this indeterminacy on the loss landscape, implicit regularization, and generalization bounds of deep neural networks remain an open question\. Future theoretical work will focus on understanding how this indeterminacy influences the optimization dynamics ofqq\-orthogonal KANs\.

## Data and Code Availability

The source code, configuration files, and pretrained models required to reproduce the experimental results presented in this paper are publicly available at[https://github\.com/amirhoseinazarpour/SW\-KAN](https://github.com/amirhoseinazarpour/SW-KAN)\.

## References

- \[1\]I\. Goodfellow, Y\. Bengio, and A\. Courville\(2016\)Deep learning\.MIT Press,Cambridge, MA, USA\.External Links:[Link](http://www.deeplearningbook.org/)Cited by:[§1](https://arxiv.org/html/2610.00050#S1.p1.1)\.
- \[2\]Z\. Liu, Y\. Wang, S\. Vaidya, F\. Ruehle, J\. Halverson, M\. Soljačić, T\. Y\. Hou, and M\. Tegmark\(2025\)KAN: Kolmogorov\-Arnold Networks\.arXiv preprint arXiv:2404\.19756\(arXiv:2404\.19756\)\.External Links:2404\.19756,[Document](https://dx.doi.org/10.48550/arXiv.2404.19756)Cited by:[§1](https://arxiv.org/html/2610.00050#S1.p1.1),[§1](https://arxiv.org/html/2610.00050#S1.p3.1),[§2\.1](https://arxiv.org/html/2610.00050#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2610.00050#S2.SS2.p1.1),[§3\.1\.3](https://arxiv.org/html/2610.00050#S3.SS1.SSS3.p1.1),[§3\.1\.3](https://arxiv.org/html/2610.00050#S3.SS1.SSS3.p1.2),[§3\.1\.3](https://arxiv.org/html/2610.00050#S3.SS1.SSS3.p2.1),[§3\.5\.5](https://arxiv.org/html/2610.00050#S3.SS5.SSS5.p1.3),[§3\.5\.5](https://arxiv.org/html/2610.00050#S3.SS5.SSS5.p2.2),[§3\.5](https://arxiv.org/html/2610.00050#S3.SS5.p1.1),[§4\.1\.1](https://arxiv.org/html/2610.00050#S4.SS1.SSS1.p1.1)\.
- \[3\]J\. Schmidt\-Hieber\(2021\)The Kolmogorov–Arnold representation theorem revisited\.Neural Networks137,pp\. 119–126\.External Links:ISSN 0893\-6080,[Document](https://dx.doi.org/10.1016/j.neunet.2021.01.020)Cited by:[§1](https://arxiv.org/html/2610.00050#S1.p1.1),[§3\.1\.2](https://arxiv.org/html/2610.00050#S3.SS1.SSS2.p1.1)\.
- \[4\]Z\. Liu, P\. Ma, Y\. Wang, W\. Matusik, and M\. Tegmark\(2024\)KAN 2\.0: Kolmogorov\-Arnold Networks Meet Science\.arXiv preprint arXiv:2408\.10205\(arXiv:2408\.10205\)\.External Links:2408\.10205,[Document](https://dx.doi.org/10.48550/arXiv.2408.10205)Cited by:[§1](https://arxiv.org/html/2610.00050#S1.p2.1),[§2\.2](https://arxiv.org/html/2610.00050#S2.SS2.p1.1)\.
- \[5\]S\. T\. Seydi\(2024\)Exploring the Potential of Polynomial Basis Functions in Kolmogorov\-Arnold Networks: A Comparative Study of Different Groups of Polynomials\.arXiv preprint arXiv:2406\.02583\(arXiv:2406\.02583\)\.External Links:2406\.02583,[Document](https://dx.doi.org/10.48550/arXiv.2406.02583)Cited by:[4th item](https://arxiv.org/html/2610.00050#S1.I1.i4.p1.1),[§1](https://arxiv.org/html/2610.00050#S1.p2.1),[§2\.3](https://arxiv.org/html/2610.00050#S2.SS3.p2.1),[§2\.3](https://arxiv.org/html/2610.00050#S2.SS3.p3.1),[§2\.4](https://arxiv.org/html/2610.00050#S2.SS4.p1.1),[§3\.5\.2](https://arxiv.org/html/2610.00050#S3.SS5.SSS2.p2.1),[§3\.5\.5](https://arxiv.org/html/2610.00050#S3.SS5.SSS5.p1.3),[§4\.1\.1](https://arxiv.org/html/2610.00050#S4.SS1.SSS1.p1.1),[§4\.1\.2](https://arxiv.org/html/2610.00050#S4.SS1.SSS2.p1.1)\.
- \[6\]S\. SS, K\. AR, G\. R, and A\. KP\(2024\)Chebyshev Polynomial\-Based Kolmogorov\-Arnold Networks: An Efficient Architecture for Nonlinear Function Approximation\.arXiv preprint arXiv:2405\.07200\(arXiv:2405\.07200\)\.External Links:2405\.07200,[Document](https://dx.doi.org/10.48550/arXiv.2405.07200)Cited by:[§1](https://arxiv.org/html/2610.00050#S1.p2.1),[§1](https://arxiv.org/html/2610.00050#S1.p3.1),[§2\.3](https://arxiv.org/html/2610.00050#S2.SS3.p2.1),[§2\.3](https://arxiv.org/html/2610.00050#S2.SS3.p3.1),[§3\.1\.3](https://arxiv.org/html/2610.00050#S3.SS1.SSS3.p2.1),[§3\.5\.5](https://arxiv.org/html/2610.00050#S3.SS5.SSS5.p1.3),[§3\.5](https://arxiv.org/html/2610.00050#S3.SS5.p1.1)\.
- \[7\]A\. Afzal Aghaei\(2025\)fKAN: Fractional Kolmogorov–Arnold Networks with trainable Jacobi basis functions\.Neurocomputing623,pp\. 129414\.External Links:ISSN 0925\-2312,[Document](https://dx.doi.org/10.1016/j.neucom.2025.129414)Cited by:[§1](https://arxiv.org/html/2610.00050#S1.p2.1),[§2\.3](https://arxiv.org/html/2610.00050#S2.SS3.p2.1),[§3\.1\.3](https://arxiv.org/html/2610.00050#S3.SS1.SSS3.p2.1),[§3\.5](https://arxiv.org/html/2610.00050#S3.SS5.p1.1)\.
- \[8\]A\. A\. Aghaei\(2024\)rKAN: Rational Kolmogorov\-Arnold Networks\.arXiv preprint arXiv:2406\.14495\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2406.14495)Cited by:[§1](https://arxiv.org/html/2610.00050#S1.p2.1),[§1](https://arxiv.org/html/2610.00050#S1.p3.1),[§2\.3](https://arxiv.org/html/2610.00050#S2.SS3.p2.1)\.
- \[9\]A\. Azarpour\(2026\)RecKAN: kolmogorov\-arnold networks with a learnable recursive polynomial basis\.arXiv preprintarXiv:2609\.01729\.Cited by:[§1](https://arxiv.org/html/2610.00050#S1.p3.1)\.
- \[10\]K\. Hornik, M\. Stinchcombe, and H\. White\(1989\)Multilayer feedforward networks are universal approximators\.Neural Networks2\(5\),pp\. 359–366\.Cited by:[§2\.1](https://arxiv.org/html/2610.00050#S2.SS1.p1.1)\.
- \[11\]A\. N\. Kolmogorov\(1957\)On the representation of continuous functions of many variables by superposition of continuous functions of one variable and addition\.Doklady Akademii Nauk SSSR114,pp\. 953–956\.Cited by:[§2\.1](https://arxiv.org/html/2610.00050#S2.SS1.p1.1),[§3\.1\.1](https://arxiv.org/html/2610.00050#S3.SS1.SSS1.p1.1),[§3\.1\.2](https://arxiv.org/html/2610.00050#S3.SS1.SSS2.p1.1),[§3\.1](https://arxiv.org/html/2610.00050#S3.SS1.p1.1)\.
- \[12\]H\. Montanelli and H\. Yang\(2020\)Error bounds for deep ReLU networks using the Kolmogorov–Arnold superposition theorem\.arXiv preprint arXiv:1906\.11945\(arXiv:1906\.11945\)\.External Links:1906\.11945,[Document](https://dx.doi.org/10.48550/arXiv.1906.11945)Cited by:[§2\.1](https://arxiv.org/html/2610.00050#S2.SS1.p1.1)\.
- \[13\]W\. Chen, Y\. Liu, and Q\. Xia\(2027\)Legendre\-KAN: High Accuracy KA Network Based on Legendre Polynomials\.InPattern Recognition,M\. De Marsico, T\. K\. Ho, F\. Jurie, C\. Liu, D\. Lopresti, I\. Nyström, J\. Ogier, A\. Ross, and L\. Wang \(Eds\.\),Vol\.16817,pp\. 659–673\.External Links:[Document](https://dx.doi.org/10.1007/978-3-032-31673-8%5F44),ISBN 978\-3\-032\-31672\-1 978\-3\-032\-31673\-8Cited by:[§2\.1](https://arxiv.org/html/2610.00050#S2.SS1.p1.1),[§3\.1\.3](https://arxiv.org/html/2610.00050#S3.SS1.SSS3.p2.1),[§3\.5](https://arxiv.org/html/2610.00050#S3.SS5.p1.1)\.
- \[14\]Y\. Wang, J\. Sun, J\. Bai, C\. Anitescu, M\. S\. Eshaghi, X\. Zhuang, T\. Rabczuk, and Y\. Liu\(2025\)Kolmogorov Arnold Informed neural network: A physics\-informed deep learning framework for solving forward and inverse problems based on Kolmogorov Arnold Networks\.Computer Methods in Applied Mechanics and Engineering433,pp\. 117518\.External Links:2406\.11045,ISSN 00457825,[Document](https://dx.doi.org/10.1016/j.cma.2024.117518)Cited by:[§2\.2](https://arxiv.org/html/2610.00050#S2.SS2.p1.1)\.
- \[15\]Z\. Bozorgasl and H\. Chen\(2024\)Wav\-KAN: Wavelet Kolmogorov\-Arnold Networks\.arXiv preprint arXiv:2405\.12832\(arXiv:2405\.12832\)\.External Links:2405\.12832,[Document](https://dx.doi.org/10.48550/arXiv.2405.12832)Cited by:[§2\.2](https://arxiv.org/html/2610.00050#S2.SS2.p1.1)\.
- \[16\]C\. Coffman and L\. Chen\(2025\)MatrixKAN: Parallelized Kolmogorov\-Arnold Network\.arXiv preprint arXiv:2502\.07176\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2502.07176)Cited by:[§2\.3](https://arxiv.org/html/2610.00050#S2.SS3.p1.1)\.
- \[17\]A\. Moradzadeh, L\. Wawrzyniak, M\. Macklin, and S\. G\. Paliwal\(2024\)UKAN: Unbound Kolmogorov\-Arnold Network Accompanied with Accelerated Library\.arXiv preprint arXiv:2408\.11200\(arXiv:2408\.11200\)\.External Links:2408\.11200,[Document](https://dx.doi.org/10.48550/arXiv.2408.11200)Cited by:[§2\.3](https://arxiv.org/html/2610.00050#S2.SS3.p1.1)\.
- \[18\]L\. N\. Zheng, W\. E\. Zhang, L\. Yue, M\. Xu, O\. Maennel, and W\. Chen\(2025\)Free\-Knots Kolmogorov\-Arnold Network: On the Analysis of Spline Knots and Advancing Stability\.arXiv preprint arXiv:2501\.09283\(arXiv:2501\.09283\)\.External Links:2501\.09283,[Document](https://dx.doi.org/10.48550/arXiv.2501.09283)Cited by:[§2\.3](https://arxiv.org/html/2610.00050#S2.SS3.p1.1)\.
- \[19\]C\. Guo, L\. Sun, S\. Li, Z\. Yuan, and C\. Wang\(2024\)Physics\-informed Kolmogorov\-Arnold Network with Chebyshev Polynomials for Fluid Mechanics\.arXiv preprint arXiv:2411\.04516\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2411.04516)Cited by:[§2\.3](https://arxiv.org/html/2610.00050#S2.SS3.p2.1)\.
- \[20\]A\. Kashefi\(2025\)PointNet with KAN versus PointNet with MLP for 3D classification and segmentation of point sets\.Computers & Graphics131,pp\. 104319\.External Links:ISSN 0097\-8493,[Document](https://dx.doi.org/10.1016/j.cag.2025.104319)Cited by:[§2\.3](https://arxiv.org/html/2610.00050#S2.SS3.p2.1)\.
- \[21\]R\. Koekoek and R\. F\. Swarttouw\(1996\)The Askey\-scheme of hypergeometric orthogonal polynomials and its q\-analogue\.Technical reportTechnical ReportarXiv:math/9602214,arXiv,Delft University of Technology\.External Links:math/9602214,[Document](https://dx.doi.org/10.48550/arXiv.math/9602214)Cited by:[§2\.4](https://arxiv.org/html/2610.00050#S2.SS4.p1.1),[§3\.2\.3](https://arxiv.org/html/2610.00050#S3.SS2.SSS3.p1.1),[§3\.5\.2](https://arxiv.org/html/2610.00050#S3.SS5.SSS2.p1.2)\.
- \[22\]V\. I\. Arnold\(2009\)On functions of three variables\.InCollected Works: Representations of Functions, Celestial Mechanics and KAM Theory, 1957–1965,V\. I\. Arnold, A\. B\. Givental, B\. A\. Khesin, J\. E\. Marsden, A\. N\. Varchenko, V\. A\. Vassiliev, O\. Ya\. Viro, and V\. M\. Zakalyukin \(Eds\.\),pp\. 5–8\.External Links:[Document](https://dx.doi.org/10.1007/978-3-642-01742-1%5F2),ISBN 978\-3\-642\-01742\-1Cited by:[§3\.1\.1](https://arxiv.org/html/2610.00050#S3.SS1.SSS1.p1.1),[§3\.1](https://arxiv.org/html/2610.00050#S3.SS1.p1.1)\.
- \[23\]J\. Braun and M\. Griebel\(2009\)On a Constructive Proof of Kolmogorov’s Superposition Theorem\.Constructive Approximation30\(3\),pp\. 653–675\.External Links:ISSN 1432\-0940,[Document](https://dx.doi.org/10.1007/s00365-009-9054-2)Cited by:[§3\.1\.1](https://arxiv.org/html/2610.00050#S3.SS1.SSS1.p2.1),[§3\.1\.2](https://arxiv.org/html/2610.00050#S3.SS1.SSS2.p2.1)\.
- \[24\]F\. Girosi and T\. Poggio\(1989\)Representation Properties of Networks: Kolmogorov’s Theorem Is Irrelevant\.Neural Computation1\(4\),pp\. 465–469\.External Links:ISSN 0899\-7667,[Document](https://dx.doi.org/10.1162/neco.1989.1.4.465)Cited by:[§3\.1\.2](https://arxiv.org/html/2610.00050#S3.SS1.SSS2.p2.1)\.
- \[25\]R\. Hecht\-Nielsen\(1987\)Kolmogorov’s mapping neural network existence theorem\.InProc\. IEEE International Conference on Neural Networks,Vol\.3,pp\. 11–14\.Cited by:[§3\.1\.2](https://arxiv.org/html/2610.00050#S3.SS1.SSS2.p2.1)\.
- \[26\]V\. Kůrková\(1991\)Kolmogorov’s Theorem Is Relevant\.Neural Computation3\(4\),pp\. 617–622\.External Links:ISSN 1530\-888X,[Document](https://dx.doi.org/10.1162/neco.1991.3.4.617)Cited by:[§3\.1\.2](https://arxiv.org/html/2610.00050#S3.SS1.SSS2.p2.1)\.
- \[27\]T\.\-J\. Stieltjes\(1894\)Recherches sur les fractions continues\.Annales de la Faculté des sciences de Toulouse : Mathématiques8\(4\),pp\. J1–J122\.External Links:ISSN 0240\-2963Cited by:[1st item](https://arxiv.org/html/2610.00050#S3.I1.i1.p1.1),[§3\.2\.4](https://arxiv.org/html/2610.00050#S3.SS2.SSS4.p1.1),[§3\.2](https://arxiv.org/html/2610.00050#S3.SS2.p1.1),[§3\.4](https://arxiv.org/html/2610.00050#S3.SS4.p1.1)\.
- \[28\]S\. Wigert\(1923\)Sur les polynômes orthogonaux et l’approximation des fonctions continues\.Arkiv för matematik, astronomi och fysik17\(18\),pp\. 1–15\.Note:JFM 49\.0296\.01Cited by:[1st item](https://arxiv.org/html/2610.00050#S3.I1.i1.p1.1),[¶3\.2\.4\.2](https://arxiv.org/html/2610.00050#S3.SS2.SSS4.P2.p1.1),[§3\.2](https://arxiv.org/html/2610.00050#S3.SS2.p1.1),[§3\.4](https://arxiv.org/html/2610.00050#S3.SS4.p1.1)\.
- \[29\]R\. Koekoek, P\. A\. Lesky, and R\. F\. Swarttouw\(2010\)Hypergeometric Orthogonal Polynomials and Their q\-Analogues\.Springer Monographs in Mathematics,Springer,Berlin, Heidelberg\.External Links:[Document](https://dx.doi.org/10.1007/978-3-642-05014-5),ISBN 978\-3\-642\-05013\-8 978\-3\-642\-05014\-5Cited by:[1st item](https://arxiv.org/html/2610.00050#S3.I1.i1.p1.1),[§3\.2\.2](https://arxiv.org/html/2610.00050#S3.SS2.SSS2.p1.1),[¶3\.2\.3\.1](https://arxiv.org/html/2610.00050#S3.SS2.SSS3.P1.p1.1),[¶3\.2\.3\.1](https://arxiv.org/html/2610.00050#S3.SS2.SSS3.P1.p5.1),[§3\.2\.3](https://arxiv.org/html/2610.00050#S3.SS2.SSS3.p1.1),[¶3\.2\.4\.1](https://arxiv.org/html/2610.00050#S3.SS2.SSS4.P1.p1.2),[§3\.2](https://arxiv.org/html/2610.00050#S3.SS2.p1.1),[§3\.4](https://arxiv.org/html/2610.00050#S3.SS4.p1.1),[§3\.5\.2](https://arxiv.org/html/2610.00050#S3.SS5.SSS2.p1.2),[§3\.5\.2](https://arxiv.org/html/2610.00050#S3.SS5.SSS2.p2.1),[§3\.5\.3](https://arxiv.org/html/2610.00050#S3.SS5.SSS3.p2.2),[§3\.5\.5](https://arxiv.org/html/2610.00050#S3.SS5.SSS5.p2.2),[§3\.5](https://arxiv.org/html/2610.00050#S3.SS5.p1.1)\.
- \[30\]J\. S\. Christiansen\(2003\)The moment problem associated with the Stieltjes–Wigert polynomials\.Journal of Mathematical Analysis and Applications277\(1\),pp\. 218–245\.External Links:ISSN 0022\-247X,[Document](https://dx.doi.org/10.1016/S0022-247X%2802%2900534-6)Cited by:[¶3\.2\.4\.2](https://arxiv.org/html/2610.00050#S3.SS2.SSS4.P2.p1.1),[¶3\.2\.4\.2](https://arxiv.org/html/2610.00050#S3.SS2.SSS4.P2.p1.3),[¶3\.2\.4\.2](https://arxiv.org/html/2610.00050#S3.SS2.SSS4.P2.p2.1),[§3\.2\.4](https://arxiv.org/html/2610.00050#S3.SS2.SSS4.p1.1),[§3\.2](https://arxiv.org/html/2610.00050#S3.SS2.p1.1),[§3\.5](https://arxiv.org/html/2610.00050#S3.SS5.p1.1)\.
- \[31\]M\. E\. H\. Ismail\(2005\)Classical and Quantum Orthogonal Polynomials in One Variable\.Encyclopedia of Mathematics and Its Applications,Cambridge University Press,Cambridge\.External Links:[Document](https://dx.doi.org/10.1017/CBO9781107325982),ISBN 978\-0\-521\-14347\-9Cited by:[¶3\.2\.4\.2](https://arxiv.org/html/2610.00050#S3.SS2.SSS4.P2.p2.1),[§3\.2](https://arxiv.org/html/2610.00050#S3.SS2.p1.1),[§3\.5\.3](https://arxiv.org/html/2610.00050#S3.SS5.SSS3.p2.2),[§3\.5](https://arxiv.org/html/2610.00050#S3.SS5.p1.1)\.
- \[32\]G\. Szegő\(1939\)Orthogonal Polynomials\.Colloquium Publications, Vol\.23,American Mathematical Society,Providence, Rhode Island\.External Links:[Document](https://dx.doi.org/10.1090/coll/023),ISBN 978\-0\-8218\-1023\-1 978\-0\-8218\-8952\-7 978\-1\-4704\-3171\-6Cited by:[¶3\.2\.4\.2](https://arxiv.org/html/2610.00050#S3.SS2.SSS4.P2.p2.1)\.
- \[33\]J\. Shen, T\. Tang, and L\. Wang\(2011\)Spectral Methods: Algorithms, Analysis and Applications\.Springer Series in Computational Mathematics, Vol\.41,Springer,Berlin, Heidelberg\.External Links:[Document](https://dx.doi.org/10.1007/978-3-540-71041-7),ISBN 978\-3\-540\-71040\-0 978\-3\-540\-71041\-7Cited by:[2nd item](https://arxiv.org/html/2610.00050#S3.I1.i2.p1.1),[§3\.4](https://arxiv.org/html/2610.00050#S3.SS4.p1.1),[§3\.4](https://arxiv.org/html/2610.00050#S3.SS4.p2.4),[§3\.5\.1](https://arxiv.org/html/2610.00050#S3.SS5.SSS1.p3.1)\.
- \[34\]J\. P\. Boyd\(2001\)Chebyshev and fourier spectral methods\.2nd edition,Dover Publications,Mineola, New York\.External Links:ISBN 978\-0486411835Cited by:[2nd item](https://arxiv.org/html/2610.00050#S3.I1.i2.p1.1),[§3\.4](https://arxiv.org/html/2610.00050#S3.SS4.p1.1)\.

Similar Articles

SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions

arXiv cs.LG

SechKAN is a novel Kolmogorov-Arnold Network architecture that uses hyperbolic secant functions as basis functions, achieving competitive performance in function fitting, PDE problems, and image classification tasks while maintaining parameter efficiency comparable to MLPs.

Geometric Kolmogorov--Arnold Network (GeoKAN)

arXiv cs.LG

This paper introduces Geometric Kolmogorov-Arnold Networks (GeoKAN), a family of geometry-aware models that learn Riemannian metrics to adapt coordinates for improved function approximation and physics-informed learning.

STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting

arXiv cs.LG

This paper introduces STKAN, a spatio-temporal forecasting architecture that integrates Taylor-polynomial Kolmogorov-Arnold Network modules for spatial and temporal token mixing. Experiments on five traffic benchmarks show competitive performance, suggesting nonlinear function approximators can complement architectural design.