Product units in gated recurrent units improve nuclear-mass prediction

arXiv cs.LG Papers

Summary

This paper proposes a novel complex-valued gated recurrent unit (GRU) architecture with multiplicative product units (AM-PU-GRU) for predicting nuclear masses, achieving state-of-the-art interpolation and extrapolation accuracy on the Atomic Mass Evaluation datasets.

arXiv:2606.06866v1 Announce Type: new Abstract: The prediction of masses of atomic nuclei using machine learning can complement theoretical models and advance the exploration of poorly known domains of the nuclear chart. We propose a machine learning technique based on gated recurrent units (GRU), which have demonstrated competitive performance in nuclear-mass prediction by exploiting long-term dependencies. By integrating multiplicative interactions and product-unit transformations within recurrent units, we report significant improvements in nuclear-mass prediction. Computations are performed in the complex domain to jointly capture amplitude and phase dynamics. For interpolation and temporal-extrapolation tasks based on the atomic mass evaluation (AME2016 and AME2020), the complex additive-multiplicative product-unit gated recurrent unit (AM-PU-GRU) model consistently achieves the lowest prediction errors, with an interpolation RMSE of 0.227 $\pm$ 0.004 MeV and an extrapolation RMSE of 0.179 $\pm$ 0.015 MeV. These results surpass other state-of-the-art machine learning models and also outperform the real-valued GRU baseline and product-unit ablation variants, while remaining robust to different theoretical priors, including WS4 and SEMF. Our findings establish complex-valued product-unit recurrent networks as a new benchmark for sequence-based nuclear-mass prediction.
Original Article
View Cached Full Text

Cached at: 06/08/26, 09:18 AM

# Product units in gated recurrent units improve nuclear-mass prediction
Source: [https://arxiv.org/html/2606.06866](https://arxiv.org/html/2606.06866)
11institutetext:Department of Mathematics, Informatics and Technology, University of Applied Sciences Koblenz, Joseph\-Rovan\-Allee 2, 53424 Remagen, Germany
11email:dellen@hs\-koblenz\.de22institutetext:Technical University of Munich, Munich, Germany33institutetext:Department of Mathematics, University of Madeira, Campus Universitário da Penteada, Funchal, 9020\-105, Portugal44institutetext:Department of Physics, Washington University in St\. Louis, 1 Brookings Drive, St\. Louis, 63130, MO, USA###### Abstract

The prediction of masses of atomic nuclei using machine learning can complement theoretical models and advance the exploration of poorly known domains of the nuclear chart\. We propose a machine learning technique based on gated recurrent units \(GRU\), which have demonstrated competitive performance in nuclear\-mass prediction by exploiting long\-term dependencies\. By integrating multiplicative interactions and product\-unit transformations within recurrent units, we report significant improvements in nuclear\-mass prediction\. Computations are performed in the complex domain to jointly capture amplitude and phase dynamics\. For interpolation and temporal\-extrapolation tasks based on the atomic mass evaluation \(AME2016 and AME2020\), the complex additive\-multiplicative product\-unit gated recurrent unit \(AM\-PU\-GRU\) model consistently achieves the lowest prediction errors, with an interpolation RMSE of 0\.227±\\pm0\.004 MeV and an extrapolation RMSE of 0\.179±\\pm0\.015 MeV\. These results surpass other state\-of\-the\-art machine learning models and also outperform the real\-valued GRU baseline and product\-unit ablation variants, while remaining robust to different theoretical priors, including WS4 and SEMF\. Our findings establish complex\-valued product\-unit recurrent networks as a new benchmark for sequence\-based nuclear\-mass prediction\.

## 1Introduction

Accurate prediction of nuclear masses is a fundamental problem in nuclear physics, with important implications for nuclear–structure theory, nucleosynthesis pathways, and applications in nuclear energy and astrophysics\. Experimental evaluations such as AME2016\[[9](https://arxiv.org/html/2606.06866#bib.bib17),[10](https://arxiv.org/html/2606.06866#bib.bib18)\]and AME2020\[[6](https://arxiv.org/html/2606.06866#bib.bib19),[21](https://arxiv.org/html/2606.06866#bib.bib20)\]provide precise measurements for many nuclides, yet large regions of the nuclear chart remain inaccessible\. This motivates theoretical and data\-driven approaches capable of both interpolation in known regions and extrapolation to unknown nuclides\.

Traditional mass models, such as the Weizsäcker–Skyrme Model \(Version 4\) \(WS4\)\[[22](https://arxiv.org/html/2606.06866#bib.bib16)\], achieve high accuracy through physics\-motivated parametrizations, but their predictive power in unexplored regions is limited\[[8](https://arxiv.org/html/2606.06866#bib.bib6)\]\. In parallel, machine learning methods have emerged to complement nuclear\-mass prediction, notably recurrent neural network \(RNN\)\[[8](https://arxiv.org/html/2606.06866#bib.bib6)\], gated recurrent unit \(GRU\)\[[1](https://arxiv.org/html/2606.06866#bib.bib21),[8](https://arxiv.org/html/2606.06866#bib.bib6)\], and mixture density network \(MDN\)\[[18](https://arxiv.org/html/2606.06866#bib.bib10),[17](https://arxiv.org/html/2606.06866#bib.bib8),[12](https://arxiv.org/html/2606.06866#bib.bib1)\]have achieved strong interpolation accuracy within experimentally known regions, sometimes comparable to physics\-based models\. These approaches benefit from their ability to learn complex nonlinear relationships directly from data and have been shown to reduce prediction errors when sufficient experimental measurements are available\. However, many of these architectures remain limited in extrapolation capability, as they primarily rely on additive representations and struggle to capture the higher\-order nonlinear dependencies inherent in nuclear\-mass systematics\.

The product\-unit \(PU\)\[[4](https://arxiv.org/html/2606.06866#bib.bib22),[11](https://arxiv.org/html/2606.06866#bib.bib23),[3](https://arxiv.org/html/2606.06866#bib.bib12),[2](https://arxiv.org/html/2606.06866#bib.bib5),[13](https://arxiv.org/html/2606.06866#bib.bib13),[15](https://arxiv.org/html/2606.06866#bib.bib15),[14](https://arxiv.org/html/2606.06866#bib.bib24)\]approach to machine learning was introduced as an alternative to summation\-based formulations, enabling multiplicative interactions that provide compact representations of polynomial and power\-law relations\. More recently, complex\-valued PU extensions\[[2](https://arxiv.org/html/2606.06866#bib.bib5),[13](https://arxiv.org/html/2606.06866#bib.bib13)\]have been explored, enabling joint modeling of amplitude and phase dynamics\.

Building on these advances, we propose complex\-valued GRU extensions for nuclear mass prediction\. By integrating multiplicative interactions and product\-unit transformations into recurrent frameworks, we obtain two novel architectures, the multiplicative\-interaction product\-unit GRU \(MI\-PU\-GRU\) and the additive\-multiplicative product\-unit GRU \(AM\-PU\-GRU\)\.

We next evaluate these architectures on interpolation and temporal extrapolation based on AME2016 and AME2020, comparing against real\-valued baselines, ablations, and prior\-informed setups\. The complex AM\-PU\-GRU model achieves the lowest error on both tasks, setting a new benchmark for sequence\-based nuclear\-mass prediction\.

## 2Related Work

Product units \(PUs\) were introduced as an alternative to summation\-based neurons in neural networks\. Instead of computing a weighted sum followed by a nonlinear activation, a PU models multiplicative interactions by raising each input to a learnable exponent and taking their product according to

y=∏i=1nxiwi=exp⁡\(∑i=1nwi​log⁡\|xi\|\)\.y=\\prod\_\{i=1\}^\{n\}x\_\{i\}^\{w\_\{i\}\}=\\exp\\left\(\\sum\_\{i=1\}^\{n\}w\_\{i\}\\log\|x\_\{i\}\|\\right\)\.\(1\)This formulation substantially enhances the expressive capacity of neural networks, enabling compact representation of polynomial, power\-law, and rational relationships\. As a result, PU\-based architectures have been shown to improve extrapolation in scientific prediction tasks such as image classification and function approximation\[[3](https://arxiv.org/html/2606.06866#bib.bib12),[2](https://arxiv.org/html/2606.06866#bib.bib5)\]\.

More recently, complex\-valued PU networks have been explored for signals with amplitude and phase\. In Eq\. \([1](https://arxiv.org/html/2606.06866#S2.E1)\), the complex extension combineslog⁡\|xi\|\\log\|x\_\{i\}\|witharg⁡\(xi\)\\arg\(x\_\{i\}\)in the exponent, introducing both amplitude scaling and phase rotation via the exponential mapping\. This richer inductive bias has improved robustness in MRI reconstruction and nuclear\-mass prediction\[[2](https://arxiv.org/html/2606.06866#bib.bib5),[13](https://arxiv.org/html/2606.06866#bib.bib13),[14](https://arxiv.org/html/2606.06866#bib.bib24)\]\.

## 3Methodology

### 3\.1Baseline: Gated Recurrent Unit

The GRU is a recurrent neural network variant designed to capture long\-term dependencies by mitigating the vanishing gradient problem\. Two gating mechanisms are introduced: an update gate and a reset gate, which regulate the information flow and memory update within each recurrent unit\.

Given an input vectorxt∈ℝdx\_\{t\}\\in\\mathbb\{R\}^\{d\}and the previous hidden stateht−1∈ℝHh\_\{t\-1\}\\in\\mathbb\{R\}^\{H\}at time steptt, a GRU computes the update gatezt∈ℝHz\_\{t\}\\in\\mathbb\{R\}^\{H\}, reset gatert∈ℝHr\_\{t\}\\in\\mathbb\{R\}^\{H\}, candidate hidden stateh~t∈ℝH\\tilde\{h\}\_\{t\}\\in\\mathbb\{R\}^\{H\}, and the new hidden stateht∈ℝHh\_\{t\}\\in\\mathbb\{R\}^\{H\}as follows:

zt\\displaystyle z\_\{t\}=σ​\(Wz​xt\+Uz​ht−1\)\\displaystyle=\\sigma\(W\_\{z\}x\_\{t\}\+U\_\{z\}h\_\{t\-1\}\)update gate\(2\)rt\\displaystyle r\_\{t\}=σ​\(Wr​xt\+Ur​ht−1\)\\displaystyle=\\sigma\(W\_\{r\}x\_\{t\}\+U\_\{r\}h\_\{t\-1\}\)reset gate\(3\)h~t\\displaystyle\\tilde\{h\}\_\{t\}=tanh⁡\(Wh​xt\+Uh​\(rt⊙ht−1\)\)\\displaystyle=\\tanh\(W\_\{h\}x\_\{t\}\+U\_\{h\}\(r\_\{t\}\\odot h\_\{t\-1\}\)\)candidate hidden state\(4\)ht\\displaystyle h\_\{t\}=\(1−zt\)⊙ht−1\+zt⊙h~t\\displaystyle=\(1\-z\_\{t\}\)\\odot h\_\{t\-1\}\+z\_\{t\}\\odot\\tilde\{h\}\_\{t\}new hidden state\.\\displaystyle\\text\{new hidden state\}\.\(5\)Here,σ​\(⋅\)\\sigma\(\\cdot\)denotes the sigmoid activation function,tanh⁡\(⋅\)\\tanh\(\\cdot\)is the hyperbolic tangent, and⊙\\odotrepresents element\-wise multiplication\. The learnable parameters satisfyW\{z,r,h\}∈ℝH×dW\_\{\\\{z,r,h\\\}\}\\in\\mathbb\{R\}^\{H\\times d\}andU\{z,r,h\}∈ℝH×HU\_\{\\\{z,r,h\\\}\}\\in\\mathbb\{R\}^\{H\\times H\}\. Bias terms are used in our implementation but omitted for notational simplicity\.

![Refer to caption](https://arxiv.org/html/2606.06866v1/fig/gru.png)Figure 1:Computational graph of a standard GRU cell\.Figure[1](https://arxiv.org/html/2606.06866#S3.F1)illustrates the internal structure of a standard GRU cell\. The diagram is consistent with the above equations: The inputxtx\_\{t\}and previous hidden stateht−1h\_\{t\-1\}are jointly used to compute the reset gatertr\_\{t\}and update gateztz\_\{t\}\. The reset gatertr\_\{t\}controls how much ofht−1h\_\{t\-1\}contributes to the candidate activationh~t\\tilde\{h\}\_\{t\}\. The final hidden statehth\_\{t\}is computed as a convex combination ofht−1h\_\{t\-1\}andh~t\\tilde\{h\}\_\{t\}, governed by the update gateztz\_\{t\}\. The parameters of the model are defined through the weight matrices associated with the update gate, reset gate, and candidate hidden state\.

### 3\.2Multiplicative\-Interaction Product\-Unit GRU \(MI\-PU\-GRU\)

To enhance the nonlinear modeling capacity of GRUs, we introduce the MI\-PU\-GRU cell, which incorporates two major innovations: a multiplicative interaction \(MI\) branch and a product\-unit \(PU\) transformation in the candidate\-state computation\. Figure[2](https://arxiv.org/html/2606.06866#S3.F2)shows the MI\-PU\-GRU computational graph\.

![Refer to caption](https://arxiv.org/html/2606.06866v1/fig/mi-pu-gru.png)Figure 2:Computational graph of the MI\-PU\-GRU cell\.The MI\-PU\-GRU case maintains the standard GRU gating structure but modifies the candidate hidden stateh~t\\tilde\{h\}\_\{t\}as follows:

zt\\displaystyle z\_\{t\}=σ​\(Wz​xt\+Uz​ht−1\)\\displaystyle=\\sigma\(W\_\{z\}x\_\{t\}\+U\_\{z\}h\_\{t\-1\}\)update gate\(6\)rt\\displaystyle r\_\{t\}=σ​\(Wr​xt\+Ur​ht−1\)\\displaystyle=\\sigma\(W\_\{r\}x\_\{t\}\+U\_\{r\}h\_\{t\-1\}\)reset gate\(7\)MIt\\displaystyle\\text\{MI\}\_\{t\}=MI​\(xt,rt⊙ht−1\)\\displaystyle=\\text\{MI\}\(x\_\{t\},r\_\{t\}\\odot h\_\{t\-1\}\)multiplicative interaction\(8\)h~t\\displaystyle\\tilde\{h\}\_\{t\}=PU​\(\[xt,rt⊙ht−1,MIt\]\)\\displaystyle=\\text\{PU\}\(\[x\_\{t\},r\_\{t\}\\odot h\_\{t\-1\},\\text\{MI\}\_\{t\}\]\)PU\-transformed candidate\(9\)ht\\displaystyle h\_\{t\}=\(1−zt\)⊙ht−1\+zt⊙h~t\\displaystyle=\(1\-z\_\{t\}\)\\odot h\_\{t\-1\}\+z\_\{t\}\\odot\\tilde\{h\}\_\{t\}new hidden state\.\\displaystyle\\text\{new hidden state\}\.\(10\)The MI term is computed using an element\-wise exponential interaction between the input and hidden representations:

MI​\(x,h\)=exp⁡\(\(Wmi​x\)⊙\(Umi​h\)\+bmi\),\\text\{MI\}\(x,h\)=\\exp\\left\(\(W\_\{\\text\{mi\}\}x\)\\odot\(U\_\{\\text\{mi\}\}h\)\+b\_\{\\text\{mi\}\}\\right\),\(11\)whereWmiW\_\{\\text\{mi\}\}andUmiU\_\{\\text\{mi\}\}are learnable projection matrices andbmib\_\{\\text\{mi\}\}is a bias term\. The MI branch computes a feature\-wise exponential interaction between the inputxtx\_\{t\}and the reset\-gated previous statert⊙ht−1r\_\{t\}\\odot h\_\{t\-1\}, producing a multiplicative representationMIt\\text\{MI\}\_\{t\}\. The exponential function ensures that the output is strictly positive, which is necessary for the subsequent PU transformation\. This representation, along withxtx\_\{t\}andrt⊙ht−1r\_\{t\}\\odot h\_\{t\-1\}, is fed into a Product\-Unit \(PU\) transformation:

PU​\(x\)=exp⁡\(Wpu⋅log⁡\(max⁡\(x,softplus​\(θ\)\+10−7\)\)\+bpu\),\\begin\{split\}\\mathrm\{PU\}\(x\)=\\exp\\Big\(W\_\{\\mathrm\{pu\}\}\\cdot\\log\\\!\\big\(\\max\(x,\\,\\mathrm\{softplus\}\(\\theta\)\+10^\{\-7\}\)\\big\)\+b\_\{\\mathrm\{pu\}\}\\Big\),\\end\{split\}\(12\)whereθ\\thetais a learnable threshold that ensures numerical stability of the logarithm\.

Compared with standard GRUs, MI\-PU\-GRU significantly enhances model expressivity by introducing higher\-order multiplicative feature interactions and a PU\-based nonlinear transformation\.

### 3\.3Additive\-Multiplicative Product\-Unit GRU \(AM\-PU\-GRU\)

The AM\-PU\-GRU architecture further extends the MI\-PU\-GRU strategy by explicitly modeling both additive and multiplicative candidate paths\. It introduces a learnable fusion gate to adaptively combine the two contributions, thereby unifying the strengths of traditional GRU\-style linearity and PU\-based multiplicative expressiveness\. Figure[3](https://arxiv.org/html/2606.06866#S3.F3)shows the two\-path candidate computation and fusion\. The overall update equations are as follows:

zt\\displaystyle z\_\{t\}=σ​\(Wz​xt\+Uz​ht−1\)\\displaystyle=\\sigma\(W\_\{z\}x\_\{t\}\+U\_\{z\}h\_\{t\-1\}\)update gate\(13\)rt\\displaystyle r\_\{t\}=σ​\(Wr​xt\+Ur​ht−1\)\\displaystyle=\\sigma\(W\_\{r\}x\_\{t\}\+U\_\{r\}h\_\{t\-1\}\)reset gate\(14\)gt\\displaystyle g\_\{t\}=σ​\(Wg​xt\+Ug​ht−1\)\\displaystyle=\\sigma\(W\_\{g\}x\_\{t\}\+U\_\{g\}h\_\{t\-1\}\)fusion gate\(15\)hadd\\displaystyle h\_\{\\text\{add\}\}=tanh⁡\(Wadd​xt\+Uadd​\(rt⊙ht−1\)\)\\displaystyle=\\tanh\(W\_\{\\text\{add\}\}x\_\{t\}\+U\_\{\\text\{add\}\}\(r\_\{t\}\\odot h\_\{t\-1\}\)\)additive candidate\(16\)MIt\\displaystyle\\text\{MI\}\_\{t\}=MI​\(xt,rt⊙ht−1\)\\displaystyle=\\text\{MI\}\(x\_\{t\},r\_\{t\}\\odot h\_\{t\-1\}\)multiplicative interaction\(17\)hmul\\displaystyle h\_\{\\text\{mul\}\}=PU​\(\[xt,rt⊙ht−1,MIt\]\)\\displaystyle=\\text\{PU\}\(\[x\_\{t\},r\_\{t\}\\odot h\_\{t\-1\},\\text\{MI\}\_\{t\}\]\)multiplicative candidate\(18\)h~t\\displaystyle\\tilde\{h\}\_\{t\}=gt⊙hmul\+\(1−gt\)⊙hadd\\displaystyle=g\_\{t\}\\odot h\_\{\\text\{mul\}\}\+\(1\-g\_\{t\}\)\\odot h\_\{\\text\{add\}\}fused candidate state\(19\)ht\\displaystyle h\_\{t\}=\(1−zt\)⊙ht−1\+zt⊙h~t\\displaystyle=\(1\-z\_\{t\}\)\\odot h\_\{t\-1\}\+z\_\{t\}\\odot\\tilde\{h\}\_\{t\}new hidden state\.\\displaystyle\\text\{new hidden state\}\.\(20\)Here, the additive candidate path retains a GRU\-style formulation, while the multiplicative candidate is constructed using the MI and PU modules\. The fusion gategtg\_\{t\}dynamically balances the two components\. This gating mechanism enables the model to flexibly adapt between linear and nonlinear dynamics depending on the temporal context\.

![Refer to caption](https://arxiv.org/html/2606.06866v1/fig/am-pu-gru.png)Figure 3:Computational graph of the AM\-PU\-GRU cell\.Compared to MI\-PU\-GRU, the AM\-PU\-GRU architecture introduces an additional level of flexibility and interpretability by learning whether additive or multiplicative dynamics are more appropriate at each time step\.

### 3\.4Complex\-Valued MI\-PU\-GRU

To extend MI\-PU\-GRU to the complex domain, we adopt a fully complex\-valued formulation\. All learnable weights, hidden states, and intermediate activations are complex\-valued\. The real\-valued inputxt∈ℝdx\_\{t\}\\in\\mathbb\{R\}^\{d\}is embedded into the complex vector space via a zero\-imaginary extensionx~t=xt\+i0∈ℂd\\tilde\{x\}\_\{t\}=x\_\{t\}\+\\mathrm\{i\}0\\in\\mathbb\{C\}^\{d\}, and the hidden state satisfiesht−1∈ℂHh\_\{t\-1\}\\in\\mathbb\{C\}^\{H\}\.

The update and reset gates are kept real\-valued to ensure stable and interpretable gating; specifically, we apply the sigmoid to the real part of complex pre\-activations, yieldingzt,rt∈ℝHz\_\{t\},r\_\{t\}\\in\\mathbb\{R\}^\{H\}:

zt\\displaystyle z\_\{t\}=σ​\(ℜ⁡\(Wz​x~t\+Uz​ht−1\)\),\\displaystyle=\\sigma\\\!\\left\(\\Re\\\!\\left\(W\_\{z\}\\tilde\{x\}\_\{t\}\+U\_\{z\}h\_\{t\-1\}\\right\)\\right\),\(21\)rt\\displaystyle r\_\{t\}=σ​\(ℜ⁡\(Wr​x~t\+Ur​ht−1\)\),\\displaystyle=\\sigma\\\!\\left\(\\Re\\\!\\left\(W\_\{r\}\\tilde\{x\}\_\{t\}\+U\_\{r\}h\_\{t\-1\}\\right\)\\right\),\(22\)MIt\\displaystyle\\mathrm\{MI\}\_\{t\}=MIℂ​\(x~t,rt⊙ht−1\),\\displaystyle=\\mathrm\{MI\}\_\{\\mathbb\{C\}\}\\\!\\left\(\\tilde\{x\}\_\{t\},\\,r\_\{t\}\\odot h\_\{t\-1\}\\right\),\(23\)h~t\\displaystyle\\tilde\{h\}\_\{t\}=PUℂ​\(\[x~t,rt⊙ht−1,MIt\]\),\\displaystyle=\\mathrm\{PU\}\_\{\\mathbb\{C\}\}\\\!\\left\(\[\\tilde\{x\}\_\{t\},\\,r\_\{t\}\\odot h\_\{t\-1\},\\,\\mathrm\{MI\}\_\{t\}\]\\right\),\(24\)ht\\displaystyle h\_\{t\}=\(1−zt\)⊙ht−1\+zt⊙h~t\.\\displaystyle=\(1\-z\_\{t\}\)\\odot h\_\{t\-1\}\+z\_\{t\}\\odot\\tilde\{h\}\_\{t\}\.\(25\)Using real\-valued gates \(bounded in\[0,1\]\[0,1\]\) provides a numerically stable and interpretable mixing of complex hidden states without introducing additional phase rotations\.

The MI branch performs feature\-wise exponential interactions in the complex domain:

MIℂ​\(x,h\)=exp⁡\(\(Wmi​x\)⊙\(Umi​h\)\+bmi\),\\mathrm\{MI\}\_\{\\mathbb\{C\}\}\(x,h\)=\\exp\\\!\\left\(\(W\_\{\\mathrm\{mi\}\}x\)\\odot\(U\_\{\\mathrm\{mi\}\}h\)\+b\_\{\\mathrm\{mi\}\}\\right\),\(26\)where all parameters are complex\-valued\. Sinceexp⁡\(a\+i​b\)=exp⁡\(a\)​\(cos⁡b\+i​sin⁡b\)\\exp\(a\+\\mathrm\{i\}b\)=\\exp\(a\)\\big\(\\cos b\+\\mathrm\{i\}\\sin b\\big\), the interaction introduces both amplitude scaling and phase rotation, enhancing expressiveness\.

The PU transformation in the complex domain is defined as

PUℂ​\(x\)=exp⁡\(Wpu​\(log⁡r​\(x\)\+i​ϕ​\(x\)\)\+bpu\),\\mathrm\{PU\}\_\{\\mathbb\{C\}\}\(x\)=\\exp\\\!\\left\(W\_\{\\mathrm\{pu\}\}\\big\(\\log r\(x\)\+\\mathrm\{i\}\\,\\phi\(x\)\\big\)\+b\_\{\\mathrm\{pu\}\}\\right\),\(27\)wherer​\(x\)=max⁡\(\|x\|,softplus​\(θ\)\+ϵ\)r\(x\)=\\max\\\!\\big\(\|x\|,\\mathrm\{softplus\}\(\\theta\)\+\\epsilon\\big\),ϕ​\(x\)=arg⁡\(x\)\\phi\(x\)=\\arg\(x\), andϵ=10−7\\epsilon=10^\{\-7\}\. Here, the logarithm is applied to the stabilized magnituder​\(x\)r\(x\)and the phase termϕ​\(x\)\\phi\(x\)is handled explicitly\.

Since our downstream task requires real\-valued predictions, the final output of the complex\-valued recurrent network is mapped back to the real domain by extracting the real part of the hidden state\. This mapping ensures compatibility with standard loss functions defined overℝ\\mathbb\{R\}\.

### 3\.5Complex\-Valued AM\-PU\-GRU

As with standard AM\-PU\-GRU, complex AM\-PU\-GRU builds upon complex MI\-PU\-GRU by integrating an additive path and a fusion gate to balance linear and nonlinear dynamics inℂ\\mathbb\{C\}\. The update procedure is given by:

zt\\displaystyle z\_\{t\}=σ​\(ℜ⁡\(Wz​x~t\+Uz​ht−1\)\),\\displaystyle=\\sigma\\\!\\left\(\\Re\\\!\\left\(W\_\{z\}\\tilde\{x\}\_\{t\}\+U\_\{z\}h\_\{t\-1\}\\right\)\\right\),\(28\)rt\\displaystyle r\_\{t\}=σ​\(ℜ⁡\(Wr​x~t\+Ur​ht−1\)\),\\displaystyle=\\sigma\\\!\\left\(\\Re\\\!\\left\(W\_\{r\}\\tilde\{x\}\_\{t\}\+U\_\{r\}h\_\{t\-1\}\\right\)\\right\),\(29\)gt\\displaystyle g\_\{t\}=σ​\(ℜ⁡\(Wg​x~t\+Ug​ht−1\)\),\\displaystyle=\\sigma\\\!\\left\(\\Re\\\!\\left\(W\_\{g\}\\tilde\{x\}\_\{t\}\+U\_\{g\}h\_\{t\-1\}\\right\)\\right\),\(30\)hadd\\displaystyle h\_\{\\mathrm\{add\}\}=tanh⁡\(ℜ⁡\(Wadd​x~t\+Uadd​\(rt⊙ht−1\)\)\),\\displaystyle=\\tanh\\\!\\left\(\\Re\\\!\\left\(W\_\{\\mathrm\{add\}\}\\tilde\{x\}\_\{t\}\+U\_\{\\mathrm\{add\}\}\\left\(r\_\{t\}\\odot h\_\{t\-1\}\\right\)\\right\)\\right\),\(31\)MIt\\displaystyle\\mathrm\{MI\}\_\{t\}=MIℂ​\(x~t,rt⊙ht−1\),\\displaystyle=\\mathrm\{MI\}\_\{\\mathbb\{C\}\}\\\!\\left\(\\tilde\{x\}\_\{t\},\\,r\_\{t\}\\odot h\_\{t\-1\}\\right\),\(32\)hmul\\displaystyle h\_\{\\mathrm\{mul\}\}=PUℂ​\(\[x~t,rt⊙ht−1,MIt\]\),\\displaystyle=\\mathrm\{PU\}\_\{\\mathbb\{C\}\}\\\!\\left\(\[\\tilde\{x\}\_\{t\},\\,r\_\{t\}\\odot h\_\{t\-1\},\\,\\mathrm\{MI\}\_\{t\}\]\\right\),\(33\)h~tadd\\displaystyle\\tilde\{h\}^\{\\mathrm\{add\}\}\_\{t\}=hadd\+i0,\\displaystyle=h\_\{\\mathrm\{add\}\}\+\\mathrm\{i\}0,\(34\)h~t\\displaystyle\\tilde\{h\}\_\{t\}=gt⊙hmul\+\(1−gt\)⊙h~tadd,\\displaystyle=g\_\{t\}\\odot h\_\{\\mathrm\{mul\}\}\+\(1\-g\_\{t\}\)\\odot\\tilde\{h\}^\{\\mathrm\{add\}\}\_\{t\},\(35\)ht\\displaystyle h\_\{t\}=\(1−zt\)⊙ht−1\+zt⊙h~t\.\\displaystyle=\(1\-z\_\{t\}\)\\odot h\_\{t\-1\}\+z\_\{t\}\\odot\\tilde\{h\}\_\{t\}\.\(36\)Here,hadd∈ℝHh\_\{\\mathrm\{add\}\}\\in\\mathbb\{R\}^\{H\}is recast toℂH\\mathbb\{C\}^\{H\}before fusion\. The use of real\-valued gates ensures numerically stable and interpretable mixing while allowing the hidden state to evolve in the full complex plane\.

### 3\.6Task Formulation

Our goal is to predict the mass excess of an atomic nucleus from sequential nuclear structure data, which we formulate as a sequence\-to\-one regression problem\. For each target nucleus with proton numberZ⋆Z^\{\\star\}and neutron numberN⋆N^\{\\star\}, we construct a short ordered sequence ofT=5T=5neighboring nuclides based on proximity in the\(Z,N\)\(Z,N\)chart\.

For a target nucleus\(Z⋆,N⋆\)\(Z^\{\\star\},N^\{\\star\}\)with mass numberA⋆=Z⋆\+N⋆A^\{\\star\}=Z^\{\\star\}\+N^\{\\star\}, the firstT−1=4T\-1=4elements are selected from the corresponding reference set of experimentally measured nuclei\. Specifically, excluding the target nucleus itself, we consider nuclei satisfyingA≤A⋆A\\leq A^\{\\star\}and select the four nearest ones in the\(Z,N\)\(Z,N\)plane as context nuclei; their measured mass excess values are used as input features\. The final elementxTx\_\{T\}corresponds to the target nucleus; its mass excess value in the input is initialized by the WS4 mass model \(i\.e\.,MET=MEWS4​\(Z⋆,N⋆\)\\mathrm\{ME\}\_\{T\}=\\mathrm\{ME\}^\{\\mathrm\{WS4\}\}\(Z^\{\\star\},N^\{\\star\}\)\) to provide a physically informed prior\.

GivenXX, the model learns a nonlinear mappingf:ℝT×3→ℝf:\\mathbb\{R\}^\{T\\times 3\}\\rightarrow\\mathbb\{R\}to predict the ground\-truth mass excessy=MEAME​\(Z⋆,N⋆\)y=\\mathrm\{ME\}^\{\\mathrm\{AME\}\}\(Z^\{\\star\},N^\{\\star\}\)of the target nucleus\. This design allows the model to exploit local continuity over neighboring nuclides while combining empirical measurements and theoretical priors\.

To capture the complex nonlinear dependencies inherent in nuclear\-mass systematics, we propose complex\-valued extensions of the MI\-PU\-GRU and AM\-PU\-GRU architectures, which integrate multiplicative interactions and product\-unit transformations into the recurrent framework\.

In addition to these primary models, we also implement their real\-valued counterparts \(MI\-PU\-GRU and AM\-PU\-GRU\) to provide a controlled comparison between real and complex formulations\. Furthermore, we design ablation variants in the complex domain, including PU\-GRU, MI\-GRU, AM\-GRU, and the baseline GRU, by selectively disabling specific architectural components\. These ablations allow us to isolate and quantify the contribution of multiplicative interactions, product\-unit transformations, and additive–multiplicative fusion\. Finally, to evaluate robustness with respect to theoretical priors, we perform prior stability experiments where the target nucleus is initialized using the semi\-empirical mass formula \(SEMF\), which incorporates volume, surface, Coulomb, asymmetry, and pairing terms\. The SEMF\-based mass excess estimate is computed using a simplified version of the Weizsäcker formula\[[23](https://arxiv.org/html/2606.06866#bib.bib14)\]:

M​E=\(Z​Mp\+N​Mn\+Z​Me−B−A\)⋅931\.5×103​keV,ME=\(ZM\_\{p\}\+NM\_\{n\}\+ZM\_\{e\}\-B\-A\)\\cdot 931\.5\\times 10^\{3\}\\text\{ keV\},\(37\)whereBBis the total binding energy, which is calculated as

B=av​A−as​A2/3−ac​Z​\(Z−1\)A1/3−aa​\(A−2​Z\)2A\+δ,B=a\_\{v\}A\-a\_\{s\}A^\{2/3\}\-a\_\{c\}\\frac\{Z\(Z\-1\)\}\{A^\{1/3\}\}\-a\_\{a\}\\frac\{\(A\-2Z\)^\{2\}\}\{A\}\+\\delta,\(38\)withA=Z\+NA=Z\+Nand the pairing termδ\\deltadefined asδ=\+ap/A\\delta=\+a\_\{p\}/\\sqrt\{A\}for even–even nuclei,δ=−ap/A\\delta=\-a\_\{p\}/\\sqrt\{A\}for odd–odd nuclei, andδ=0\\delta=0otherwise\.

## 4Experiments and Results

### 4\.1Experimental Setup

#### Experiment I: Interpolation\.

This experiment assesses the model’s interpolation capability within the domain of experimentally known nuclides\. We randomly split the AME2020 dataset into training, validation, and test subsets using a 70%\-15%\-15% ratio\. During sequence construction for the validation and test sets, only the experimentally measured mass\-excess values from the training set are used to provide input features for precursor nuclei\. That is, all non\-target entries in a sequence are filled using training set information, ensuring no data leakage\. Model selection is based on validation performance: the version of the model that achieves the lowest validation loss during training is retained and used for final evaluation on the test set\.

#### Experiment II: Temporal Extrapolation\.

To test the extrapolation capability of our models, we simulate a temporal generalization scenario by training on AME2016\-era nuclides and testing on nuclides added in the subsequent AME2020 release\. This setup allows us to evaluate the model’s ability to generalize to newly measured nuclei that were previously unknown\. In addition, we reserve 5% of the AME2016 data as an interpolation test set to provide a comprehensive assessment of the model’s performance on both interpolation and extrapolation tasks within the same experimental framework\. As in Experiment I, all sequences for both extrapolation and interpolation test sets are constructed using only nuclei from the training set to provide precursor mass excess values\.

#### Training Setup\.

All models are trained in PyTorch for 3000 epochs using MSE loss and the RAdam optimizer \(initial learning rate10−310^\{\-3\}\)\. We use a step learning\-rate schedule that halves the learning rate every 500 epochs, and set the batch size to 64\.

### 4\.2Main Results

To evaluate the predictive capability of the proposed architectures, we first compare the baseline real\-valued GRU with the complex\-valued MI\-PU\-GRU and AM\-PU\-GRU under the WS4 prior\. For a fair comparison, we control the parameter scale by adjusting network depth: the baseline GRU is implemented with 120 stacked recurrent layers, while the complex MI\-PU\-GRU and AM\-PU\-GRU are implemented with 60 and 50 layers, respectively\. This configuration results in a comparable number of trainable parameters across all models\. For complex\-valued architectures, each learnable weight is represented asa\+i​ba\+ib, whereaais the real andbbthe imaginary part of the complex weight, and thus counts as two trainable parameters when reporting parameter sizes\.

Each network is trained independently five times without fixing the random seed, in order to account for stochasticity in initialization and training dynamics\. We report the mean and standard deviation of the RMSE over these runs\.

The interpolation results of Experiment I are summarized in Table[1](https://arxiv.org/html/2606.06866#S4.T1)\. As can be seen, both complex\-valued PU\-GRU variants consistently outperform the real\-valued GRU baseline\. In particular, the complex AM\-PU\-GRU model achieves the lowest interpolation error \(0\.227 MeV\) and the lowest sample standard deviation \(0\.004 MeV\), highlighting the effectiveness of combining additive and multiplicative candidate paths within the complex domain\.

Table 1:Principal results ofExperiment Iusing WS4 estimates\. Here, RV denotes the real\-valued model, CV the complex\-valued model\. Results are reported as the mean±\\pmsample standard deviation \(MeV\)\.NparamN\_\{\\mathrm\{param\}\}indicates the number of trainable parameters\. IntRMSE denotes interpolation RMSE \(and ExtRMSE denotes extrapolation RMSE, when applicable\)\. The best result in each column is highlighted in bold\.The results of Experiment II are summarized in Table[2](https://arxiv.org/html/2606.06866#S4.T2)\. Compared to the real\-valued GRU baseline, both complex PU\-GRU variants yield lower extrapolation errors, with the complex AM\-PU\-GRU again achieving the best overall performance\. In particular, AM\-PU\-GRU attains the lowest interpolation RMSE of 0\.253 ± 0\.003 MeV and the lowest extrapolation RMSE of 0\.179 ± 0\.015 MeV\. These findings indicate that the proposed complex architectures improve both interpolation and temporal\-extrapolation performance under the present experimental setup\. In particular, the extrapolation capability of the complex AM\-PU\-GRU is significantly stronger, indicating its effectiveness in capturing long\-range structural dependencies in nuclear mass systematics\. Moreover, AM\-PU\-GRU exhibits the most stable extrapolation performance, as reflected by the lowest standard deviation among all models\.

Table 2:Main results ofExperiment IIusing WS4 estimates\.The lower extrapolation RMSE in Table 2 does not mean that extrapolation is generally easier\. In Experiment II, the interpolation set is a 5% hold\-out from AME2016, whereas the extrapolation set consists of nuclei newly added in AME2020, so the two test sets differ in composition\. Since prediction depends on local neighborhood structure in the\(Z,N\)\(Z,N\)chart and the WS4 prior, some AME2020\-added nuclei may be easier to predict than the held\-out AME2016 subset\.

We further compare our best\-performing models with existing approaches from the literature, as summarized in the Discussion section\.

### 4\.3Comparative Analysis: Real vs\. Complex

To assess the effect of complex\-valued modeling, we compare the performance of real\-valued PU\-GRU variants with their complex\-valued counterparts under the WS4 prior\. Specifically, we implement a 90\-layer real\-valued MI\-PU\-GRU and a 75\-layer real\-valued AM\-PU\-GRU, such that their parameter counts are comparable to those of the 60\-layer complex MI\-PU\-GRU and the 50\-layer complex AM\-PU\-GRU, respectively\. Each real\-valued network is trained independently three times, and the reported performance is the average RMSE across runs\.

The results, summarized in Tables[3](https://arxiv.org/html/2606.06866#S4.T3)and[4](https://arxiv.org/html/2606.06866#S4.T4), indicate that the complex\-valued PU\-GRU variants consistently outperform their real\-valued counterparts in both interpolation and extrapolation settings\. This demonstrates that extending PU\-GRU architectures into the complex domain yields richer representational capacity by jointly modeling amplitude and phase dynamics, leading to improved predictive accuracy\.

Table 3:Results ofExperiment Ifor real\-valued variants\.Table 4:Results ofExperiment IIfor real\-valued variants\.
### 4\.4Ablation Study

To isolate the contributions of individual architectural components, we have performed an ablation study in the complex domain\. We implement four variants: a complex\-valued GRU with 80 layers, a 70\-layer complex MI\-GRU \(without PU and additive path\), a 60\-layer complex AM\-GRU \(without PU but with additive multiplicative fusion\), and an 80\-layer complex PU\-GRU \(retaining PU but without explicit MI or AM mechanisms\)\. All models are configured to have comparable parameter counts for a fair comparison\.

The results are presented in Tables[5](https://arxiv.org/html/2606.06866#S4.T5)and[6](https://arxiv.org/html/2606.06866#S4.T6)\. We observe that neither the complex MI\-GRU nor the complex AM\-GRU improves upon the baseline complex GRU, and in fact both exhibit degraded performance\. By contrast, the complex PU\-GRU model achieves a substantial reduction in error, clearly outperforming the baseline GRU\. This indicates that the product\-unit transformation is the key factor driving improvements, while multiplicative interactions or additive–multiplicative fusion alone are insufficient to enhance predictive accuracy\.

Table 5:Ablation results ofExperiment Iwith complex\-valued GRU, MI\-GRU, AM\-GRU, and PU\-GRU implementation\.Table 6:Ablation results ofExperiment IIwith complex\-valued GRU, MI\-GRU, AM\-GRU and PU\-GRU strategies\.
### 4\.5Prior Stability Analysis

Finally, we evaluate robustness to the choice of the theoretical baseline estimate used to initialize the target nucleus \(WS4 vs\. SEMF\)\. In this experiment, the network configurations remain the same as in the main results, i\.e\., a 120\-layer real\-valued GRU baseline, a 60\-layer complex MI\-PU\-GRU, and a 50\-layer complex AM\-PU\-GRU\. The only difference lies in the initialization of the target nucleus, where we replace the WS4 estimate with the SEMF, which incorporates volume, surface, Coulomb, asymmetry, and pairing terms\.

The results are summarized in Tables[7](https://arxiv.org/html/2606.06866#S4.T7)and[8](https://arxiv.org/html/2606.06866#S4.T8)\. Compared to the WS4\-based experiments, the relative ranking of models is unchanged: both complex PU\-GRU variants outperform the real\-valued GRU, and the complex AM\-PU\-GRU version continues to achieve the best overall accuracy\. These findings confirm that the proposed architectures are robust to the choice of baseline initialization model \(WS4 vs\. SEMF\) and maintain stable capability under this change\.

Table 7:Prior stability results ofExperiment Iunder SEMF initialization\.Table 8:Prior stability results ofExperiment IIunder SEMF initialization\.

## 5Discussion

Beyond comparisons among our proposed variants, it is important to contextualize the performance of the PU\-GRU architectures with respect to existing approaches in the literature\. Numerous machine\-learning methods have been previously applied to nuclear mass prediction, including MDN, categorical gradient boosting trees \(CatBoost\), and fully connected neural networks \(FCNN\)\. We summarize representative results reported in prior studies and compare them against our best\-performing model, the complex AM\-PU\-GRU\. Table[9](https://arxiv.org/html/2606.06866#S5.T9)summarizes the interpolation performance of various machine learning models and our models on the AME2020 dataset \(Experiment I\)\. All models are evaluated on the task of predicting nuclear mass excess values using various input feature sets and network architectures\.

Table 9:Comparison with existing approaches inExperiment I\. All RMSE values are reported in MeV\. The column “Selection criterion” specifies the nuclide selection rule,NfeatN\_\{\\mathrm\{feat\}\}denotes the number of input features per nuclide, andEmeasureE\_\{\\mathrm\{measure\}\}refers to the experimental uncertainty of the measured mass\-excess values\.ModelSelection criterionNfeatN\_\{\\mathrm\{feat\}\}IntRMSE \(MeV\)XGBoost\[[20](https://arxiv.org/html/2606.06866#bib.bib3)\]3\.702MLP Regressor\[[20](https://arxiv.org/html/2606.06866#bib.bib3)\]43\.128RFR\[[20](https://arxiv.org/html/2606.06866#bib.bib3)\]3\.089MISR\[[19](https://arxiv.org/html/2606.06866#bib.bib7)\]12≤Z≤5012\\leq Z\\leq 50100\.99MDN\[[17](https://arxiv.org/html/2606.06866#bib.bib8)\]Z,N≥5Z,N\\geq 590\.395MDN\[[12](https://arxiv.org/html/2606.06866#bib.bib1)\]Z≥10Z\\geq 10110\.246CatBoost\[[5](https://arxiv.org/html/2606.06866#bib.bib4)\]8≤Z≤1088\\leq Z\\leq 1088≤N≤1628\\leq N\\leq 16270\.189PI\-FCNN\[[7](https://arxiv.org/html/2606.06866#bib.bib2)\]Z,N\>20Z,N\>20Em​e​a​s​u​r​e<100​k​e​VE\_\{measure\}<100\\,keV130\.122CPUN\[[2](https://arxiv.org/html/2606.06866#bib.bib5)\]100\.438RNN\[[8](https://arxiv.org/html/2606.06866#bib.bib6)\]0\.601LSTM\[[8](https://arxiv.org/html/2606.06866#bib.bib6)\]Z,N≥8Z,N\\geq 8110\.557GRU\[[8](https://arxiv.org/html/2606.06866#bib.bib6)\]0\.459CV AM\-PU\-GRU \[ours\]Z,N≥8Z,N\\geq 830\.227As shown in Table[9](https://arxiv.org/html/2606.06866#S5.T9), Jalili et al\.\[[8](https://arxiv.org/html/2606.06866#bib.bib6)\]reported results for several recurrent architectures, including GRU, RNN, and LSTM, but their models utilize broader input features, as well as standard recurrent designs without product\-unit augmentation\. Our complex\-valued AM\-PU\-GRU model achieves superior performance with a more compact and physics\-informed input representation, achieving an RMSE of 0\.227 MeV for mass\-excess prediction\.

We also compare with the CPUN proposed by Dellen et al\.\[[2](https://arxiv.org/html/2606.06866#bib.bib5)\], which yields an RMSE of 0\.438 MeV\. While their model also incorporates product units and complex\-valued representations, it is based on a fully connected feedforward architecture and lacks the temporal modeling capability of our recurrent design\. In contrast, our PU\-GRU framework integrates multiplicative interactions into a recurrent paradigm and explicitly exploits sequence\-level dependencies among neighboring nuclides, leading to improved performance over CPUN\[[2](https://arxiv.org/html/2606.06866#bib.bib5)\]\.

Nevertheless, our model does not outperform the best results reported in the literature, notably CatBoost\[[5](https://arxiv.org/html/2606.06866#bib.bib4)\]\(0\.189 MeV\) and Physics\-Informed FCNN\[[7](https://arxiv.org/html/2606.06866#bib.bib2)\]\(0\.122 MeV\)\. This gap is partly due to differences in data selection and domain\-specific constraints\. For example,\[[7](https://arxiv.org/html/2606.06866#bib.bib2)\]evaluates only nuclides withZ,N\>20Z,N\>20and measured error<100<100keV, excluding many difficult edge cases, while\[[5](https://arxiv.org/html/2606.06866#bib.bib4)\]imposes upper bounds ofZ≤108Z\\leq 108andN≤162N\\leq 162, which also simplifies the task\. By contrast, our model is evaluated over a broader nuclide range without such filtering, making the problem more challenging\. In addition, unlike tree\-based methods such as CatBoost\[[5](https://arxiv.org/html/2606.06866#bib.bib4)\], which are not end\-to\-end and often depend on feature engineering, our PU\-GRU follows an end\-to\-end paradigm that maps raw nuclear sequences directly to mass excess values while retaining modular interpretability\.

Table[10](https://arxiv.org/html/2606.06866#S5.T10)summarizes the interpolation performance on AME2016 and temporal extrapolation on AME2020 of various machine\-learning models and our models \(Experiment II\)\. Across both interpolation and extrapolation tasks, our complex\-valued AM\-PU\-GRU model consistently achieves the lowest prediction error \(0\.253 MeV and 0\.179 MeV\) among all compared methods\. Furthermore, its performance even surpasses certain non\-end\-to\-end approaches, such as physics\-informed FCNN\[[7](https://arxiv.org/html/2606.06866#bib.bib2)\]and CNN\-WS4\[[16](https://arxiv.org/html/2606.06866#bib.bib11)\], highlighting the effectiveness of learning nuclear mass systematics directly from sequential data\.

Table 10:Comparison with existing approaches inExperiment II\. All RMSE values are reported in MeV\.
## 6Conclusion

We have presented complex\-valued extensions of the MI\-PU\-GRU and AM\-PU\-GRU architectures for nuclear mass prediction\. By embedding multiplicative interactions and product\-unit transformations into recurrent units and extending them to the complex domain, our models capture nonlinear dependencies and long\-range structural correlations that are difficult to represent with standard additive recurrent architectures\.

Through systematic evaluation on interpolation and extrapolation tasks using AME2016 and AME2020 datasets, we have shown that the complex AM\-PU\-GRU model achieves the best overall performance, consistently yielding the lowest RMSE across all scenarios\. Comparative analyses further reveal that complex\-valued models outperform their real\-valued counterparts, while ablation studies identify the product\-unit transformation as the critical component driving accuracy improvements\. Moreover, prior stability experiments confirm that the proposed architectures maintain consistent predictive capability under both WS4 and SEMF initialization, demonstrating robustness with respect to theoretical priors\.

Beyond outperforming existing end\-to\-end machine learning approaches, our models also surpass certain non\-end\-to\-end correction schemes, underscoring the potential of product\-unit recurrent architectures as powerful tools for scientific prediction tasks\. Future work will explore extending complex PU\-based recurrent networks to other nuclear observables and to broader domains where extrapolation capability is essential\.

## References

- \[1\]K\. Cho, B\. Van Merriënboer, C\. Gulcehre, D\. Bahdanau, F\. Bougares, H\. Schwenk, and Y\. Bengio\(2014\)Learning phrase representations using rnn encoder\-decoder for statistical machine translation\.arXiv preprint arXiv:1406\.1078\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p2.1)\.
- \[2\]B\. Dellen, U\. Jaekel, P\. S\. Freitas, and J\. W\. Clark\(2024\)Predicting nuclear masses with product\-unit networks\.Physics Letters B852,pp\. 138608\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p3.1),[§2](https://arxiv.org/html/2606.06866#S2.p1.2),[§2](https://arxiv.org/html/2606.06866#S2.p2.2),[Table 9](https://arxiv.org/html/2606.06866#S5.T9.14.14.4.1),[§5](https://arxiv.org/html/2606.06866#S5.p3.1)\.
- \[3\]B\. Dellen, U\. Jaekel, and M\. Wolnitza\(2019\)Function and pattern extrapolation with product\-unit networks\.InComputational Science–ICCS 2019: 19th International Conference, Faro, Portugal, June 12–14, 2019, Proceedings, Part II 19,pp\. 174–188\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p3.1),[§2](https://arxiv.org/html/2606.06866#S2.p1.2)\.
- \[4\]R\. Durbin and D\. E\. Rumelhart\(1989\)Product units: a computationally powerful and biologically plausible extension to backpropagation networks\.Neural Computation1\(1\),pp\. 133–142\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p3.1)\.
- \[5\]J\. Guo, H\. Wang, Z\. Zhang, and M\. Liu\(2025\)Probing the refined performance of the categorical\-boosting algorithm to the hartree\-fock\-bogoliubov mass model with several skyrme forces\.Physical Review C111\(5\),pp\. 054322\.Cited by:[Table 9](https://arxiv.org/html/2606.06866#S5.T9.10.6.3),[§5](https://arxiv.org/html/2606.06866#S5.p4.4)\.
- \[6\]W\. Huang, M\. Wang, F\. G\. Kondev, G\. Audi, and S\. Naimi\(2021\)The ame 2020 atomic mass evaluation \(i\)\. evaluation of input data, and adjustment procedures\.Chinese Physics C45\(3\),pp\. 030002\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p1.1)\.
- \[7\]Y\. Huang, J\. Chen, J\. Jia, L\. Liu, Y\. Ma, and C\. Zhang\(2025\)Validation and extrapolation of atomic masses with a physics\-informed fully connected neural network\.Physical Review C111\(3\),pp\. 034329\.Cited by:[Table 10](https://arxiv.org/html/2606.06866#S5.T10.6.6.3),[Table 9](https://arxiv.org/html/2606.06866#S5.T9.12.8.3),[§5](https://arxiv.org/html/2606.06866#S5.p4.4),[§5](https://arxiv.org/html/2606.06866#S5.p5.1)\.
- \[8\]A\. Jalili, F\. Pan, A\. X\. Chen, and J\. P\. Draayer\(2025\)Deep learning approaches for nuclear binding energy prediction: a comparative study of rnn, gru and lstm models\.arXiv preprint arXiv:2503\.19348\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p2.1),[Table 9](https://arxiv.org/html/2606.06866#S5.T9.13.9.2),[Table 9](https://arxiv.org/html/2606.06866#S5.T9.14.15.5.1),[Table 9](https://arxiv.org/html/2606.06866#S5.T9.14.16.6.1),[§5](https://arxiv.org/html/2606.06866#S5.p2.1)\.
- \[9\]F\. Kondev and S\. Naimi\(2017\)The ame2016 atomic mass evaluation \(i\)\. evaluation of input data; and adjustment procedures\.Chinese physics C41\(3\),pp\. 030002\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p1.1)\.
- \[10\]F\. Kondev and S\. Naimi\(2017\)The ame2016 atomic mass evaluation \(ii\)\. tables, graphs and references\.Chinese Physics C41\(3\),pp\. 030003\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p1.1)\.
- \[11\]L\. Leerink, C\. Giles, B\. Horne, and M\. Jabri\(1994\)Learning with product units\.Advances in neural information processing systems7\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p3.1)\.
- \[12\]M\. Li, T\. M\. Sprouse, B\. S\. Meyer, and M\. R\. Mumpower\(2024\)Atomic masses with machine learning for the astrophysical r process\.Physics Letters B848,pp\. 138385\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p2.1),[Table 9](https://arxiv.org/html/2606.06866#S5.T9.8.4.2)\.
- \[13\]Z\. Li, U\. Jaekel, and B\. Dellen\(2024\)Data\-driven 3d shape completion with product units\.InInternational Conference on Computational Science,pp\. 302–315\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p3.1),[§2](https://arxiv.org/html/2606.06866#S2.p2.2)\.
- \[14\]Z\. Li, U\. Jaekel, and B\. Dellen\(2025\)Advancing complex\-valued neural networks with product units for mri reconstruction\.InInternational Conference on Neural Information Processing,pp\. 540–554\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p3.1),[§2](https://arxiv.org/html/2606.06866#S2.p2.2)\.
- \[15\]Z\. Li, U\. Jaekel, and B\. Dellen\(2025\)Deep residual learning with product units\.arXiv preprint arXiv:2505\.04397\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p3.1)\.
- \[16\]Y\. Lu, T\. Shang, P\. Du, J\. Li, H\. Liang, and Z\. Niu\(2025\)Nuclear mass predictions based on a convolutional neural network\.Physical Review C111\(1\),pp\. 014325\.Cited by:[Table 10](https://arxiv.org/html/2606.06866#S5.T10.4.4.2),[§5](https://arxiv.org/html/2606.06866#S5.p5.1)\.
- \[17\]M\. Mumpower, M\. Li, T\. M\. Sprouse, B\. S\. Meyer, A\. E\. Lovell, and A\. T\. Mohan\(2023\)Bayesian averaging for ground state masses of atomic nuclei in a machine learning approach\.Frontiers in physics11,pp\. 1198572\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p2.1),[Table 9](https://arxiv.org/html/2606.06866#S5.T9.7.3.2)\.
- \[18\]M\. R\. Mumpower, T\. M\. Sprouse, A\. E\. Lovell, and A\. T\. Mohan\(2022\)Physically interpretable machine learning for nuclear masses\.Physical Review C106\(2\),pp\. L021301\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p2.1),[Table 10](https://arxiv.org/html/2606.06866#S5.T10.3.3.2)\.
- \[19\]J\. M\. Munoz, S\. M\. Udrescu, and R\. F\. Garcia Ruiz\(2025\)Discovering nuclear models from symbolic machine learning\.Communications Physics8\(1\),pp\. 101\.Cited by:[Table 9](https://arxiv.org/html/2606.06866#S5.T9.6.2.2)\.
- \[20\]B\. Pandey, S\. Giri, R\. D\. Pant, M\. Jalan, A\. Chaudhary, and N\. P\. Adhikari\(2024\)Prediction of binding energy using machine learning approach\.AIP Advances14\(10\)\.Cited by:[Table 9](https://arxiv.org/html/2606.06866#S5.T9.14.11.1.1),[Table 9](https://arxiv.org/html/2606.06866#S5.T9.14.12.2.1),[Table 9](https://arxiv.org/html/2606.06866#S5.T9.14.13.3.1)\.
- \[21\]M\. Wang, W\. J\. Huang, F\. G\. Kondev, G\. Audi, and S\. Naimi\(2021\)The ame 2020 atomic mass evaluation \(ii\)\. tables, graphs and references\.Chinese Physics C45\(3\),pp\. 030003\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p1.1)\.
- \[22\]N\. Wang, M\. Liu, X\. Wu, and J\. Meng\(2014\)Surface diffuseness correction in global mass formula\.Physics Letters B734,pp\. 215–219\.Cited by:[§1](https://arxiv.org/html/2606.06866#S1.p2.1)\.
- \[23\]C\. v\. Weizsäcker\(1935\)Zur theorie der kernmassen\.Zeitschrift für Physik96\(7\),pp\. 431–458\.Cited by:[§3\.6](https://arxiv.org/html/2606.06866#S3.SS6.p5.7)\.
- \[24\]E\. Yüksel, D\. Soydaner, and H\. Bahtiyar\(2024\)Nuclear mass predictions using machine learning models\.Physical Review C109\(6\),pp\. 064322\.Cited by:[Table 10](https://arxiv.org/html/2606.06866#S5.T10.2.2.2),[Table 10](https://arxiv.org/html/2606.06866#S5.T10.7.8.1.1)\.

Similar Articles