A matched-integrator evaluation of Hamiltonian neural networks on pendulum and Kepler dynamics

arXiv cs.LG 论文

摘要

The paper presents a controlled matched-integrator evaluation of Hamiltonian Neural Networks against a parameter-matched baseline on pendulum and Kepler dynamics, showing significant reductions in energy drift and trajectory error. It also examines the behavior of Störmer–Verlet-style rollouts with learned non-separable Hamiltonians.

arXiv:2608.10235v1 Announce Type: new Abstract: Hamiltonian Neural Networks (HNNs) parameterize conservative dynamics through a learned scalar Hamiltonian, providing an architectural prior that is absent from generic vector-field neural networks. We evaluate this prior under a controlled protocol in which an HNN and a parameter-matched feedforward baseline are trained on the same RK4-generated trajectories, use the same central-difference derivative targets and optimization settings, and are integrated at inference with the same RK4 scheme. Results are reported over five independent training seeds. On the nonlinear pendulum, the HNN reduces mean energy drift by 42-fold and mean trajectory MSE by 15.8-fold at T = 100, approximately 16 pendulum periods. Its energy drift also remains bounded and exhibits substantially lower seed-to-seed variability than the standard-network baseline. An energy-stratified analysis shows that the difference becomes more pronounced as trajectories explore more nonlinear regions of phase space. As an additional diagnostic, we examine an explicit St\"ormer--Verlet-style rollout of the learned HNN. Because the learned Hamiltonian is not constrained to the separable form H(q,p) = T(p) + V(q), the standard symplecticity guarantee of velocity Verlet does not directly apply. We further apply the same matched-integrator protocol to the three-dimensional Kepler two-body problem. The HNN again exhibits lower trajectory, energy, and angular-momentum drift than the parameter-matched baseline. These experiments provide a controlled study of how Hamiltonian parameterization affects long-horizon prediction and physical consistency across two conservative dynamical systems.
查看原文
查看缓存全文

缓存时间: 2026/08/12 08:28

# A matched-integrator evaluation of Hamiltonian neural networks on pendulum and Kepler dynamics
Source: [https://arxiv.org/html/2608.10235](https://arxiv.org/html/2608.10235)
Nyabuto Lenick Kemunto Yaé Ulrich Gaba Birahim Tewe African Institute for Mathematical Sciences \(AIMS\) Senegal

###### Abstract

Hamiltonian Neural Networks \(HNNs\) parameterize conservative dynamics through a learned scalar Hamiltonian, providing an architectural prior that is absent from generic vector\-field neural networks\. We evaluate this prior under a controlled protocol in which an HNN and a parameter\-matched feedforward baseline are trained on the same RK4\-generated trajectories, use the same central\-difference derivative targets and optimization settings, and are integrated at inference with the same RK4 scheme\. Results are reported over five independent training seeds\.

On the nonlinear pendulum, the HNN reduces mean energy drift by42×42\\timesand mean trajectory MSE by15\.8×15\.8\\timesatT=100T=100, approximately 16 pendulum periods\. Its energy drift also remains bounded and exhibits substantially lower seed\-to\-seed variability than the standard\-network baseline\. An energy\-stratified analysis shows that the difference becomes more pronounced as trajectories explore more nonlinear regions of phase space\.

As an additional diagnostic, we examine an explicit Störmer–Verlet\-style rollout of the learned HNN\. Because the learned Hamiltonian is not constrained to the separable formH​\(q,p\)=T​\(p\)\+V​\(q\)H\(q,p\)=T\(p\)\+V\(q\), the standard symplecticity guarantee of velocity Verlet does not directly apply\.

We further apply the same matched\-integrator protocol to the three\-dimensional Kepler two\-body problem\. The HNN again exhibits lower trajectory, energy, and angular\-momentum drift than the parameter\-matched baseline\. These experiments provide a controlled study of how Hamiltonian parameterization affects long\-horizon prediction and physical consistency across two conservative dynamical systems\.

## 1Introduction

Physics\-informed machine learning augments or replaces a numerical integrator with a learned surrogate whose architecture reflects known structure of the target system\[[7](https://arxiv.org/html/2608.10235#bib.bib7)\]\. Among the architectural priors proposed for conservative dynamics, Hamiltonian Neural Networks \(HNNs\)\[[4](https://arxiv.org/html/2608.10235#bib.bib4)\]occupy a central position: they learn a scalar HamiltonianH^​\(q,p\)\\widehat\{H\}\(q,p\)from data and derive the vector field through Hamilton’s equationsq˙=∂H^/∂p\\dot\{q\}=\\partial\\widehat\{H\}/\\partial p,p˙=−∂H^/∂q\\dot\{p\}=\-\\partial\\widehat\{H\}/\\partial q, so the learned dynamics are Hamiltonian by construction\. Related architectures embed similar priors for Lagrangian systems\[[3](https://arxiv.org/html/2608.10235#bib.bib3)\], symplectic maps\[[6](https://arxiv.org/html/2608.10235#bib.bib6)\], and controlled Hamiltonian systems\[[11](https://arxiv.org/html/2608.10235#bib.bib11)\]\.

This paper asks a controlled version of the HNN\-versus\-baseline question: how much improvement remains when the learned architectures are compared using the same data, derivative targets, optimization settings, comparable parameter counts, and the same numerical integrator at inference? Rather than attributing any observed difference to the integrator, we hold the RK4 rollout procedure fixed and study the effect of the learned parameterization itself\. We additionally investigate how an explicit Störmer–Verlet\-style rollout behaves when applied to a Hamiltonian that is learned as a generic, potentially non\-separable scalar function\.

We study these questions first on the nonlinear pendulum, which provides a controlled benchmark with known dynamics and a closed analytical energy function, and then on the three\-dimensional Kepler two\-body problem\. Neither benchmark is presented as a stand\-in for molecular dynamics or general many\-body systems; they are used to isolate the architectural mechanism before considering more complex settings\.

#### Contributions\.

We evaluate the models under a controlled matched\-integrator protocol and report the following\.

1. 1\.Controlled matched\-integrator comparison\.Parameter\-matched standard NN and HNN are trained on identical central\-difference derivative targets from the same RK4\-generated data and integrated at inference by the same RK4 scheme, overNseed=5N\_\{\\mathrm\{seed\}\}=5independent training seeds\. Under this control the HNN reduces mean long\-horizon energy drift by42×42\\timesand mean trajectory MSE by15\.8×15\.8\\timesatT=100T=100, with tight seed\-to\-seed variance on the HNN column and broad variance on the standard\-NN column \(Table[3](https://arxiv.org/html/2608.10235#S4.T3)\)\.
2. 2\.Bounded\-versus\-growing signature\.The HNN’s energy drift saturates byT=5T=5near3×10−43\\times 10^\{\-4\}and stays there throughT=100T=100; the standard NN’s drift grows monotonically from5×10−45\\times 10^\{\-4\}to2×10−22\\times 10^\{\-2\}\. The bounded\-versus\-growing behavior provides qualitative evidence that the Hamiltonian parameterization improves long\-horizon physical stability, in addition to its lower numerical error\.
3. 3\.Energy\-stratified analysis\.On trajectories stratified by initial energyH0H\_\{0\}, the standard NN’s drift grows∼5×\\sim 5\\timesfrom low\- to high\-energy terciles while the HNN’s grows only∼2×\\sim 2\\timesand never exceeds3×10−43\\times 10^\{\-4\}\. The architectural advantage widens in more nonlinear regions of phase space\.
4. 4\.Computational\-cost comparison\.We report wall\-clock per\-step and per\-rollout cost for all four inference paths \(Section[4\.7](https://arxiv.org/html/2608.10235#S4.SS7)\)\. On this 1\-DOF benchmark, the learned surrogates are slower than RK4 with the analytical vector field; the scaling argument for learned models is with respect to the number of interacting bodies and is not settled by the pendulum\.
5. 5\.Negative result on HNN \+ Verlet\.A Störmer–Verlet rollout of the trained HNN does not improve on HNN \+ RK4 at long horizon\. The learnedH^\\widehat\{H\}is a generic scalar network and is not constrained to the separable form required by the standard velocity\-Verlet splitting, so that symplecticity guarantee does not directly apply\. We discuss two possible remedies: explicitly constraining the learned Hamiltonian to a separable form, or using a symplectic integrator applicable to general Hamiltonians\. We also relate this diagnostic to alternative structure\-preserving architectures such as SympNets\[[6](https://arxiv.org/html/2608.10235#bib.bib6)\]\.
6. 6\.Transfer to 3D\.Applying the same protocol to the Kepler two\-body problem inℝ3\\mathbb\{R\}^\{3\}\(6D phase space, rotational symmetry\) atNseed=5N\_\{\\mathrm\{seed\}\}=5seeds shows the architectural advantage transferring: atT=10T=10, the HNN’s energy drift is8\.0±1\.0×8\.0\\pm 1\.0\\timessmaller, its angular\-momentum drift4\.2±0\.5×4\.2\\pm 0\.5\\timessmaller, and its trajectory MSE5\.5±0\.8×5\.5\\pm 0\.8\\timessmaller than the standard\-NN baseline\. Angular momentum is not included explicitly in the training loss; the smaller drift therefore suggests that the learned Hamiltonian approximately captures part of the rotational structure of the Kepler dynamics\.

The scientific contribution is not the pendulum numbers themselves, which are consistent with the finding first reported byGreydanus et al\. \[[4](https://arxiv.org/html/2608.10235#bib.bib4)\]\. The contribution is the combination of a controlled capacity and integrator\-matched comparison, multi\-seed evaluation, horizon\- and energy\-stratified diagnostics, computational cost analysis, the explicit\-Verlet compatibility diagnostic, and a six\-dimensional Kepler extension\.

## 2Background and related work

#### Hamiltonian mechanics\.

A conservative system with generalised coordinateqqand conjugate momentumppis described by a scalar functionH​\(q,p\)H\(q,p\), the Hamiltonian, through

q˙=∂H∂p,p˙=−∂H∂q\.\\dot\{q\}=\\frac\{\\partial H\}\{\\partial p\},\\qquad\\dot\{p\}=\-\\frac\{\\partial H\}\{\\partial q\}\.\(1\)The flow generated by \([1](https://arxiv.org/html/2608.10235#S2.E1)\) preserves the symplectic 2\-formd​q∧d​pdq\\wedge dpand, along any trajectory,d​H/d​t=0dH/dt=0\. Numerically integrating \([1](https://arxiv.org/html/2608.10235#S2.E1)\) with a generic scheme such as RK4 introduces a small energy error at each step that accumulates over time\. Symplectic integrators preserve a modified HamiltonianH~\\widetilde\{H\}close toHH, which bounds the true\-energy drift over exponentially long times\[[5](https://arxiv.org/html/2608.10235#bib.bib5)\]\.

#### The pendulum\.

We use the normalised pendulumH​\(q,p\)=12​p2\+\(1−cos⁡q\)H\(q,p\)=\\tfrac\{1\}\{2\}p^\{2\}\+\(1\-\\cos q\)\(mass, length, gravity all unit\)\. Below the separatrix \(H0<2H\_\{0\}<2\) the system librates; above, it rotates\. We restrict to librational initial conditions so all trajectories are closed orbits\. The exact solution is expressible in Jacobi elliptic functions; we do not use the exact form and instead take an RK4 solution atΔ​t=10−2\\Delta t=10^\{\-2\}as reference solution, whose numerical energy drift is at the double\-precision noise floor \(∼10−11\\sim 10^\{\-11\}overT=100T=100\)\.

#### HNN and related architectures\.

Greydanus et al\. \[[4](https://arxiv.org/html/2608.10235#bib.bib4)\]introduced the HNN with the loss‖∇⟂H^−\(q˙,p˙\)𝖳‖2\\\|\\nabla\_\{\\perp\}\\widehat\{H\}\-\(\\dot\{q\},\\dot\{p\}\)^\{\\mathsf\{T\}\}\\\|^\{2\}where∇⟂=\(∂p,−∂q\)\\nabla\_\{\\perp\}=\(\\partial\_\{p\},\-\\partial\_\{q\}\)and the derivative targets are estimated from data\. Subsequent work extended this in several directions\. Lagrangian NNs\[[3](https://arxiv.org/html/2608.10235#bib.bib3)\]lift the same structural\-prior idea to systems specified by a Lagrangian and generalise to constrained coordinates\. Symplectic ODE\-Net\[[11](https://arxiv.org/html/2608.10235#bib.bib11)\]combines the Hamiltonian prior with explicit control inputs\. Hamiltonian generative networks\[[10](https://arxiv.org/html/2608.10235#bib.bib10)\]bring HNNs into a variational\-inference setting for learning from images\. SympNets\[[6](https://arxiv.org/html/2608.10235#bib.bib6)\]take a different approach: rather than learning a scalar Hamiltonian and subsequently integrating its vector field, they construct neural maps that are symplectic by design\. Their formulation can represent dynamics arising from both separable and non\-separable Hamiltonian systems\. This distinction is important for the explicit\-Verlet diagnostic considered later in this paper\.

#### Physics\-informed learning and Neural ODEs\.

A parallel line of work embeds physical constraints as loss terms rather than as architectural priors\[[9](https://arxiv.org/html/2608.10235#bib.bib9)\]\. Neural ODEs\[[2](https://arxiv.org/html/2608.10235#bib.bib2)\]learn the vector fieldfθf\_\{\\theta\}directly and integrate it with an off\-the\-shelf solver; our Model A is closely related to the vector\-field parameterization used in Neural ODEs, but is trained directly on derivative targets and is evaluated using a fixed RK4 rollout scheme\. The distinction we draw is between an*architectural*prior \(HNN, LNN, SympNets\) and a*soft*prior \(PINN loss terms\), the former of which is closer in spirit to the equivariant\-network literature\[[1](https://arxiv.org/html/2608.10235#bib.bib1)\]\.

#### Controlled\-comparison rationale\.

Using the same numerical integrator for the compared architectures is not itself new\. Our focus is instead the combination of matched model capacity, identical derivative supervision, a common inference integrator, multi\-seed evaluation, horizon\-dependent analysis, and a separate integrator\-compatibility diagnostic\. Together these controls make the architectural comparison easier to interpret\.

## 3Method

### 3\.1Data generation

Reference trajectories are generated by RK4 withΔ​t=10−2\\Delta t=10^\{\-2\}overT=10T=10fromN=100N=100initial conditions\(q0,p0\)∼𝒰​\(\[−1,1\]×\[−1,1\]\)\(q\_\{0\},p\_\{0\}\)\\sim\\mathcal\{U\}\(\[\-1,1\]\\times\[\-1,1\]\)restricted toH0<2H\_\{0\}<2\(librational regime\)\. Each trajectory contains10011001state points; discarding the endpoints for the central\-difference target leaves999999state–derivative pairs per trajectory\. The100100trajectories are split70/15/1570/15/15into train / validation / test sets, giving69 93069\\,930training pairs,14 98514\\,985validation pairs, and14 98514\\,985test pairs\.

The derivative target for supervision is the second\-order central\-difference estimate

x˙nCD=xn\+1−xn−12​Δ​t=\(qn\+1−qn−12​Δ​t,pn\+1−pn−12​Δ​t\)\.\\dot\{x\}\_\{n\}^\{\\mathrm\{CD\}\}=\\frac\{x\_\{n\+1\}\-x\_\{n\-1\}\}\{2\\Delta t\}=\\left\(\\frac\{q\_\{n\+1\}\-q\_\{n\-1\}\}\{2\\Delta t\},\\;\\frac\{p\_\{n\+1\}\-p\_\{n\-1\}\}\{2\\Delta t\}\\right\)\.\(2\)We use \([2](https://arxiv.org/html/2608.10235#S3.E2)\) rather than the analytical vector field to keep the comparison to standard NNs fair: neither model sees reference solution derivatives at training time, only what a data\-driven pipeline could compute from observed trajectories\.

### 3\.2Models

#### Model A: standard feedforward NN\.

A fully connected networkfϕ:ℝ2→ℝ2f\_\{\\phi\}:\\mathbb\{R\}^\{2\}\\to\\mathbb\{R\}^\{2\}mapping\(q,p\)→\(q˙,p˙\)\(q,p\)\\to\(\\dot\{q\},\\dot\{p\}\)directly\. Architecture2→64→64→64→22\\to 64\\to 64\\to 64\\to 2withtanh\\tanhactivations and8 6428\\,642trainable parameters \(Figure[1](https://arxiv.org/html/2608.10235#S3.F1)\)\. The loss is

ℒA​\(ϕ\)=1N​∑n=1N‖fϕ​\(xn\)−x˙nCD‖22\.\\mathcal\{L\}\_\{A\}\(\\phi\)=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}\\\|f\_\{\\phi\}\(x\_\{n\}\)\-\\dot\{x\}\_\{n\}^\{\\mathrm\{CD\}\}\\\|\_\{2\}^\{2\}\.\(3\)
qnq\_\{n\}pnp\_\{n\}tanh\\tanh6464tanh\\tanh6464tanh\\tanh6464q˙^\\hat\{\\dot\{q\}\}p˙^\\hat\{\\dot\{p\}\}ℒA​\(ϕ\)=1N​∑n‖fϕ​\(xn\)−x˙nCD‖22\\mathcal\{L\}\_\{A\}\(\\phi\)=\\tfrac\{1\}\{N\}\\sum\_\{n\}\\\|f\_\{\\phi\}\(x\_\{n\}\)\-\\dot\{x\}\_\{n\}^\{\\mathrm\{CD\}\}\\\|\_\{2\}^\{2\}Figure 1:Model A: standard feedforward network\. Two inputs\(qn,pn\)\(q\_\{n\},p\_\{n\}\), three hidden layers of6464tanh units, two outputs\(q˙^,p˙^\)\(\\hat\{\\dot\{q\}\},\\hat\{\\dot\{p\}\}\)predicted directly and8 6428\\,642trainable parameters\. Loss is MSE against the central\-difference derivative target; no structural prior on the learned dynamics\.
#### Model B: Hamiltonian NN\.

A scalar networkH^θ:ℝ2→ℝ\\widehat\{H\}\_\{\\theta\}:\\mathbb\{R\}^\{2\}\\to\\mathbb\{R\}mapping\(q,p\)→H^θ​\(q,p\)\(q,p\)\\to\\widehat\{H\}\_\{\\theta\}\(q,p\)\. Architecture2→64→64→64→12\\to 64\\to 64\\to 64\\to 1withtanh\\tanhactivations and8 5778\\,577trainable parameters \(6565fewer than Model A, due to the single scalar output head; Figure[2](https://arxiv.org/html/2608.10235#S3.F2)\)\. The predicted vector field is derived from Hamilton’s equations by automatic differentiation:

x˙^n,θ=\(∂H^θ∂p​\(qn,pn\),−∂H^θ∂q​\(qn,pn\)\)\.\\widehat\{\\dot\{x\}\}\_\{n,\\theta\}=\\left\(\\frac\{\\partial\\widehat\{H\}\_\{\\theta\}\}\{\\partial p\}\(q\_\{n\},p\_\{n\}\),\\;\-\\frac\{\\partial\\widehat\{H\}\_\{\\theta\}\}\{\\partial q\}\(q\_\{n\},p\_\{n\}\)\\right\)\.\(4\)The loss is

ℒB​\(θ\)=1N​∑n=1N‖x˙^n,θ−x˙nCD‖22,\\mathcal\{L\}\_\{B\}\(\\theta\)=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}\\\|\\widehat\{\\dot\{x\}\}\_\{n,\\theta\}\-\\dot\{x\}\_\{n\}^\{\\mathrm\{CD\}\}\\\|\_\{2\}^\{2\},\(5\)identical in form to Model A’s, but on the derivative field produced by the Hamiltonian gradient\.

qnq\_\{n\}pnp\_\{n\}tanh\\tanh6464tanh\\tanh6464tanh\\tanh6464H^\\widehat\{H\}autogradJ​∇H^J\\,\\nabla\\widehat\{H\}q˙^\\hat\{\\dot\{q\}\}p˙^\\hat\{\\dot\{p\}\}ℒB​\(θ\)=1N​∑n‖\(∂pH^θ,−∂qH^θ\)​\(xn\)−x˙nCD‖22\\mathcal\{L\}\_\{B\}\(\\theta\)=\\tfrac\{1\}\{N\}\\sum\_\{n\}\\big\\\|\\big\(\\partial\_\{p\}\\widehat\{H\}\_\{\\theta\},\\;\-\\partial\_\{q\}\\widehat\{H\}\_\{\\theta\}\\big\)\(x\_\{n\}\)\-\\dot\{x\}\_\{n\}^\{\\mathrm\{CD\}\}\\big\\\|\_\{2\}^\{2\}Figure 2:Model B: Hamiltonian NN\. The scalar outputH^θ\\widehat\{H\}\_\{\\theta\}is passed through an autograd step returningJ​∇H^θ=\(∂pH^θ,−∂qH^θ\)J\\nabla\\widehat\{H\}\_\{\\theta\}=\(\\partial\_\{p\}\\widehat\{H\}\_\{\\theta\},\\,\-\\partial\_\{q\}\\widehat\{H\}\_\{\\theta\}\), i\.e\. Hamilton’s equations applied to the learned Hamiltonian\. Three hidden layers of6464tanh units,8 5778\\,577trainable parameters \(same width and depth as Model A\)\. Loss is on the derivative field, identical in form to Model A’s\.

### 3\.3Training protocol

Both models are trained with Adam\[[8](https://arxiv.org/html/2608.10235#bib.bib8)\]at learning rate10−310^\{\-3\}, batch size512512, for300300epochs, in float64 precision\. The best\-validation\-loss checkpoint is retained\. Both models see the same train/validation/test split, the same targets, and the same optimiser state schedule\. Within each seed, the same random seed is used before training each architecture\. Because the models have different output parameterizations, this does not make their initial weights identical; rather, it provides a reproducible control of the random\-number generation and data\-order randomness\. Across seeds, the dataset, targets, splits, and hyperparameters remain fixed\. The principal architectural difference is the parameterization of the dynamics: direct vector\-field prediction for Model A versus the symplectic gradient of a learned scalar Hamiltonian for Model B\. The autograd derivative computation required by Model B is treated as a consequence of that architecture\.

Figure[3](https://arxiv.org/html/2608.10235#S3.F3)sketches the training loop and the Adam update applied at every mini\-batch iteration\.

batch\(xn,x˙nCD\)\(x\_\{n\},\\dot\{x\}\_\{n\}^\{\\mathrm\{CD\}\}\)forwardfϕ​\(x\)f\_\{\\phi\}\(x\)orJ​∇H^θ​\(x\)J\\nabla\\widehat\{H\}\_\{\\theta\}\(x\)MSEℒ\\mathcal\{L\}backwardg=∇θℒg=\\nabla\_\{\\theta\}\\mathcal\{L\}Adam stepθ←θ−η​m^/\(v^\+ε\)\\theta\\leftarrow\\theta\-\\eta\\,\\hat\{m\}/\(\\sqrt\{\\hat\{v\}\}\+\\varepsilon\)gk=∇θℒ​\(θk\)mk=β1​mk−1\+\(1−β1\)​gkvk=β2​vk−1\+\(1−β2\)​gk⊙2g\_\{k\}=\\nabla\_\{\\theta\}\\mathcal\{L\}\(\\theta\_\{k\}\)\\qquad m\_\{k\}=\\beta\_\{1\}m\_\{k\-1\}\+\(1\-\\beta\_\{1\}\)g\_\{k\}\\qquad v\_\{k\}=\\beta\_\{2\}v\_\{k\-1\}\+\(1\-\\beta\_\{2\}\)g\_\{k\}^\{\\odot 2\}m^k=mk/\(1−β1k\)v^k=vk/\(1−β2k\)θk\+1=θk−η​m^k/\(v^k\+ε\)\\hat\{m\}\_\{k\}=m\_\{k\}/\(1\-\\beta\_\{1\}^\{k\}\)\\hskip 18\.49988pt\\hat\{v\}\_\{k\}=v\_\{k\}/\(1\-\\beta\_\{2\}^\{k\}\)\\hskip 18\.49988pt\\theta\_\{k\+1\}=\\theta\_\{k\}\-\\eta\\,\\hat\{m\}\_\{k\}/\(\\sqrt\{\\hat\{v\}\_\{k\}\}\+\\varepsilon\)Figure 3:Training loop and Adam update\[[8](https://arxiv.org/html/2608.10235#bib.bib8)\]applied identically to both models\. At each mini\-batch of size\|B\|=512\|B\|=512we compute the model’s derivative prediction \(directly for Model A, via autograd onH^θ\\widehat\{H\}\_\{\\theta\}for Model B\), its MSE against the central\-difference target, the parameter gradient by backpropagation, and an Adam step withη=10−3\\eta=10^\{\-3\},β1=0\.9\\beta\_\{1\}=0\.9,β2=0\.999\\beta\_\{2\}=0\.999,ε=10−8\\varepsilon=10^\{\-8\}\. Training runs for300300epochs, and the best\-validation\-loss checkpoint is retained\.
### 3\.4Matched\-integrator protocol

At inference, both models are rolled out from the same initial conditions using the*same*RK4 stepper atΔ​t=10−2\\Delta t=10^\{\-2\}\. Model A supplies the vector field directly; Model B supplies the autograd\-derived vector field fromH^θ\\widehat\{H\}\_\{\\theta\}\. This ensures that differences observed in our primary comparison cannot be attributed to using different numerical integrators at inference\.

An additional*diagnostic*rollout is included: the same trained Model B is evaluated with an explicit Störmer–Verlet\-style \(velocity\-Verlet\) stepper\. The standard velocity\-Verlet splitting is symplectic for separable Hamiltonians\. This is not treated as Model B’s primary rollout, its purpose is to test whether an off\-the\-shelf symplectic scheme inherits its guarantees on a generic learned Hamiltonian\. Section[5](https://arxiv.org/html/2608.10235#S5)returns to this\.

### 3\.5Seed protocol

A single training run can give a misleading picture of the architectural comparison, particularly at small data budgets where seed variance dominates\. All main\-experiment numbers in Section[4](https://arxiv.org/html/2608.10235#S4)are reported as the mean±\\pmstandard deviation overNseed=5N\_\{\\mathrm\{seed\}\}=5independent training seeds drawn from\{42,43,44,45,46\}\\\{42,43,44,45,46\\\}\. Each seed re\-initialises the PyTorch and NumPy random number generators but leaves the data, the train/validation/test split, and every other hyperparameter unchanged\. Seed\-to\-seed variation therefore reflects only initialisation and stochastic\-optimisation noise\. Tables report mean and standard deviation asmean±std\\text\{mean\}\\pm\\text\{std\}; figures show mean traces with±1​σ\\pm 1\\sigmashaded bands or error bars, as appropriate\. The sample\-efficiency sweep \(Section[4\.6](https://arxiv.org/html/2608.10235#S4.SS6)\) uses the same seed count per data budget to allow honest comparison across budgets; the earlier single\-seed sweep is retained in the supplement as an illustration of why the multi\-seed protocol matters\.

Table 1:Controlled comparison\. The only intentional difference is the architectural prior on the learned dynamics\.

## 4Results

### 4\.1Fit quality

Both models reach comparable but distinct final\-fit quality\. Averaged overNseed=5N\_\{\\mathrm\{seed\}\}=5seeds, the best\-validation derivative MSE isMSEAval=\(1\.35±0\.43\)×10−7\\mathrm\{MSE\}\_\{A\}^\{\\mathrm\{val\}\}=\(1\.35\\pm 0\.43\)\\times 10^\{\-7\}andMSEBval=\(7\.69±2\.06\)×10−8\\mathrm\{MSE\}\_\{B\}^\{\\mathrm\{val\}\}=\(7\.69\\pm 2\.06\)\\times 10^\{\-8\}; test derivative MSE is\(1\.97±0\.43\)×10−7\(1\.97\\pm 0\.43\)\\times 10^\{\-7\}\(A\) versus\(1\.09±0\.35\)×10−7\(1\.09\\pm 0\.35\)\\times 10^\{\-7\}\(B\)\. The HNN is1\.8×1\.8\\timesmore accurate on the supervised objective \(ratio of means\) despite having6565fewer parameters\. Learning curves are provided in Fig\.[4](https://arxiv.org/html/2608.10235#S4.F4)\.

![Refer to caption](https://arxiv.org/html/2608.10235v1/x1.png)\(a\)Model A \(standard NN\)\.
![Refer to caption](https://arxiv.org/html/2608.10235v1/x2.png)\(b\)Model B \(HNN\)\.

Figure 4:Training and validation loss over300300epochs \(seed4242representative curves; multi\-seed pattern is stable\)\. Both models fit the derivative target smoothly; the HNN reaches a lower validation MSE\.
### 4\.2Short\-horizon accuracy

AtT=2T=2\(roughly1/31/3of a period, well inside the training horizon\) both models track the RK4 reference visually\. Fig\.[5](https://arxiv.org/html/2608.10235#S4.F5)overlays three representative test trajectories for each model\.

![Refer to caption](https://arxiv.org/html/2608.10235v1/x3.png)\(a\)Standard NN vs RK4\.
![Refer to caption](https://arxiv.org/html/2608.10235v1/x4.png)\(b\)HNN vs RK4\.

Figure 5:Short\-horizon test rollouts on three test initial conditions\. Curves are nearly indistinguishable at this horizon; the differences show up as trajectory MSE grows with time \(Table[2](https://arxiv.org/html/2608.10235#S4.T2)\)\.Quantitatively, mean trajectory MSE atT=2T=2is\(3\.12±1\.36\)×10−7\(3\.12\\pm 1\.36\)\\times 10^\{\-7\}\(A\) versus\(1\.42±0\.97\)×10−7\(1\.42\\pm 0\.97\)\\times 10^\{\-7\}\(B\), a ratio of means of2\.2×2\.2\\times\. This ratio grows with the rollout horizon \(Table[2](https://arxiv.org/html/2608.10235#S4.T2)\)\.

### 4\.3Long\-horizon stability and energy conservation

Table 2:Trajectory mean squared error across prediction horizons on the held\-out test set, mean±\\pmstd overNseed=5N\_\{\\mathrm\{seed\}\}=5seeds\. Both models integrated with the same RK4 stepper\. The ratio is computed on means rather than per\-seed to keep it robust to Model A’s large seed\-to\-seed variance at long horizon\.Table 3:Maximum absolute energy driftmaxt∈\[0,T\]⁡\|H​\(t\)−H​\(0\)\|\\max\_\{t\\in\[0,T\]\}\|H\(t\)\-H\(0\)\|across prediction horizons, mean±\\pmstd overNseed=5N\_\{\\mathrm\{seed\}\}=5seeds\. The RK4\-reference column sits at the double\-precision noise floor and is deterministic\. Model A’s drift grows monotonically and shows large seed\-to\-seed variance \(σ/μ≈85%\\sigma/\\mu\\approx 85\\%atT=100T=100\)\. Model B’s drift saturates byT=5T=5and stays flat throughT=100T=100with much smaller variance \(σ/μ≈42%\\sigma/\\mu\\approx 42\\%\): the HNN’s drift is not only smaller but tighter\. Ratio\-of\-means atT=100T=100is42×42\\times\.†HNN\+Verlet is included as an auxiliary diagnostic using the representative seed 42; all primary Model A versus Model B pendulum results are reported overNseed=5N\_\{\\mathrm\{seed\}\}=5independent training seeds\. See Section[5](https://arxiv.org/html/2608.10235#S5)\.Table[2](https://arxiv.org/html/2608.10235#S4.T2)extends the short\-horizon result toT=100T=100\(approximately1616pendulum periods\)\. The ratio of mean trajectory MSE grows non\-monotonically from2\.2×2\.2\\timesatT=2T=2to15\.8×15\.8\\timesatT=100T=100: at short horizons the two networks are within a factor of22on trajectory error, but Model A’s error accumulates faster because its dynamics lack the structural prior\. The apparent dip atT=10T=10\(ratio1\.2×1\.2\\times\) is not an inversion of the effect: at that horizon both models sit within their per\-seed standard deviations of each other, so the ratio\-of\-means is dominated by initialisation noise rather than by any architectural difference; the ratio only becomes a meaningful architectural signal once Model A’s error has accumulated past that noise floor, which happens aroundT=20T=20–5050\. Seed\-to\-seed variability on Model A becomes very large at long horizon \(σ/μ\>100%\\sigma/\\mu\>100\\%atT=100T=100\), indicating that its long\-horizon behavior is more sensitive to the particular training realization\.

The energy story is cleaner \(Table[3](https://arxiv.org/html/2608.10235#S4.T3)\)\. Under matched RK4, mean Model A energy drift grows monotonically from2\.75×10−42\.75\\times 10^\{\-4\}atT=2T=2to9\.08×10−39\.08\\times 10^\{\-3\}atT=100T=100, a factor of33×33\\timesover the window\. Model B’s mean drift saturates byT=5T=5at2\.16×10−42\.16\\times 10^\{\-4\}and stays there throughT=100T=100\(Fig\.[6\(a\)](https://arxiv.org/html/2608.10235#S4.F6.sf1)\)\. The ratio of means atT=100T=100is42×42\\times\. The contrasting curve shapes are consistent with the expected effect of the Hamiltonian inductive bias: Model A exhibits accumulating energy error, whereas Model B remains within a bounded energy\-error band over the tested interval\. The experiment therefore provides evidence of improved long\-horizon physical consistency, without implying exact conservation of the true Hamiltonian\. The HNN’s drift is also*tighter*across seeds \(σ/μ≈42%\\sigma/\\mu\\approx 42\\%\) than the standard NN’s \(σ/μ≈85%\\sigma/\\mu\\approx 85\\%\), so the structural prior buys both mean\-error and predictability\.

Fig\.[7](https://arxiv.org/html/2608.10235#S4.F7)shows theT=100T=100long\-horizon rollout\. The standard\-NN trajectory has visibly separated from the closed orbit in phase space; HNN\+RK4 and HNN\+Verlet remain on it, and the energy panels match Table[3](https://arxiv.org/html/2608.10235#S4.T3)\.

![Refer to caption](https://arxiv.org/html/2608.10235v1/x5.png)\(a\)Max energy drift atT=100T=100, four rollouts \(log scale\)\.
![Refer to caption](https://arxiv.org/html/2608.10235v1/x6.png)\(b\)Energy and energy drift vs\. time,T=10T=10\.

Figure 6:Energy conservation\. Under matched RK4, the ratio of mean HNN drift to mean standard\-NN drift atT=100T=100is42×42\\times\(5 seeds\); the HNN drift is bounded, the standard\-NN drift growing\. HNN\+Verlet is included as a diagnostic \(Section[5](https://arxiv.org/html/2608.10235#S5)\)\.![Refer to caption](https://arxiv.org/html/2608.10235v1/x7.png)Figure 7:Long\-horizon rollout atT=100T=100\(∼16\\sim 16pendulum periods\)\. Left: phase\-space trajectory\. Middle: Hamiltonian vs\. time\. Right:log10⁡\|H​\(t\)−H​\(0\)\|\\log\_\{10\}\|H\(t\)\-H\(0\)\|\. The standard NN separates from the closed orbit and its energy drifts linearly; the HNN rollouts remain on the orbit and the energy is bounded\.
### 4\.4Learned Hamiltonian

Fig\.[8](https://arxiv.org/html/2608.10235#S4.F8)compares the learned scalarH^θ​\(q,p\)\\widehat\{H\}\_\{\\theta\}\(q,p\)against the analytical HamiltonianH​\(q,p\)H\(q,p\)on the training domain \(both mean\-centred, since dynamics are invariant to an additive constant\)\. Level sets match qualitatively across the domain, confirming that the HNN has learned a scalar function whose gradients reproduce the pendulum vector field rather than fitting the vector field directly\. Quantitatively, on a200×200200\\times 200grid over\(q,p\)∈\[−1,1\]2\(q,p\)\\in\[\-1,1\]^\{2\}, the mean\-centred learned Hamiltonian matches the analytical one to RMSE7\.3×10−57\.3\\times 10^\{\-5\}andL∞L\_\{\\infty\}error2\.0×10−42\.0\\times 10^\{\-4\}\(0\.01%0\.01\\%and0\.02%0\.02\\%of the Hamiltonian range on the domain\); the gradient error‖∇H^−∇H‖2\\\|\\nabla\\widehat\{H\}\-\\nabla H\\\|\_\{2\}has mean3\.0×10−43\.0\\times 10^\{\-4\}\(0\.04%0\.04\\%of the true gradient norm\)\. The learned\-Hamiltonian approximation error is on the same order of magnitude as the saturated HNN energy drift \(2\.16×10−42\.16\\times 10^\{\-4\}, Table[3](https://arxiv.org/html/2608.10235#S4.T3)\)\. This agreement is consistent with the hypothesis that long\-horizon energy error is strongly influenced by approximation error in the learned Hamiltonian, although it does not by itself establish a causal relationship\.

![Refer to caption](https://arxiv.org/html/2608.10235v1/x8.png)Figure 8:Analytical vs\. learned Hamiltonian on the librational domain, both mean\-centred\. The agreement in contour structure is consistent with the bounded\-drift behavior observed in the HNN rollout\.
### 4\.5Energy\-stratified analysis

Test trajectories are bucketed by initial energyH0H\_\{0\}into low, medium, and high terciles \(Table[4](https://arxiv.org/html/2608.10235#S4.T4)\)\. The standard NN’s mean drift grows by a factor of4\.3×4\.3\\timesfrom low\- \(4\.16×10−44\.16\\times 10^\{\-4\}\) to high\-energy \(1\.78×10−31\.78\\times 10^\{\-3\}\) buckets; the HNN’s mean drift grows only by2\.5×2\.5\\times\(1\.26×10−41\.26\\times 10^\{\-4\}to3\.12×10−43\.12\\times 10^\{\-4\}\) and never exceeds3\.2×10−43\.2\\times 10^\{\-4\}\. The architectural advantage widens with orbital nonlinearity\. This is a second axis on which the structural prior pays off, distinct from the horizon axis of the previous section\.

Table 4:Energy\-stratified test analysis atT=10T=10\(mean±\\pmstd overNseed=5N\_\{\\mathrm\{seed\}\}=5seeds\)\. Model A’s mean drift grows by4\.3×4\.3\\timesfrom low\- to high\-energy buckets; Model B’s grows by only2\.5×2\.5\\timesand never exceeds3\.2×10−43\.2\\times 10^\{\-4\}\. The architectural advantage widens in more nonlinear regions of phase space\.
### 4\.6Sample\-efficiency sweep

Retraining both models atNtraj∈\{8,16,32,64\}N\_\{\\mathrm\{traj\}\}\\in\\\{8,16,32,64\\\}withNseed=5N\_\{\\mathrm\{seed\}\}=5seeds per data budget gives the test derivative MSE in Table[5](https://arxiv.org/html/2608.10235#S4.T5)\. Two findings are worth stating\.

*A low\-data crossover appears in this experiment\.*AtNtraj=8N\_\{\\mathrm\{traj\}\}=8the standard NN is approximately2×2\\timesbetter than the HNN, whereas the HNN outperforms it at the three larger tested budgets\. This suggests a possible low\-data regime in which estimating a Hamiltonian\-gradient field is more difficult, although the present sweep is not intended to identify a universal crossover point\.

*Beyond that budget the HNN wins consistently, plateauing around a factor of5×5\\times\.*AtNtraj=16N\_\{\\mathrm\{traj\}\}=16the HNN pulls ahead3×3\\times, atNtraj=32N\_\{\\mathrm\{traj\}\}=32by5\.3×5\.3\\times, and atNtraj=64N\_\{\\mathrm\{traj\}\}=64by2\.7×2\.7\\times\. The ratio atNtraj=64N\_\{\\mathrm\{traj\}\}=64is smaller than at3232because both models are effectively saturated at that budget; the interesting regime is the low\- and medium\-data zone where the structural prior earns its keep once it has enough data to be identified\.

Table 5:Sample\-efficiency sweep, mean±\\pmstd of test derivative MSE overNseed=5N\_\{\\mathrm\{seed\}\}=5seeds per data budget\. AtNtraj=8N\_\{\\mathrm\{traj\}\}=8the standard NN outperforms the HNN, whereas the HNN performs better at the three larger tested budgets\.![Refer to caption](https://arxiv.org/html/2608.10235v1/x9.png)Figure 9:Sample\-efficiency curve \(log–log\), mean over 5 seeds with error bars\. BelowNtraj≈12N\_\{\\mathrm\{traj\}\}\\approx 12the standard NN wins; above, the HNN does\.
### 4\.7Computational cost

We measure per\-step wall\-clock cost for the four inference paths on a single CPU core with float64 precision\. Timings are averaged over10 00010\\,000calls after a1 0001\\,000\-call warm\-up usingtime\.perf\_counter\(\)\. Table[6](https://arxiv.org/html/2608.10235#S4.T6)gives the per\-step cost and the total wall\-clock for aT=100T=100rollout atΔ​t=10−2\\Delta t=10^\{\-2\}\(10 00010\\,000steps\)\.

Table 6:Wall\-clock cost per rollout step and totalT=100T=100rollout on a single CPU core, mean±\\pmstd across10 00010\\,000timing repetitions\. RK4 with the analytical vector field is fastest by an order of magnitude; Model A is∼9×\\sim 9\\timesslower; Model B is∼26×\\sim 26\\timesslower under RK4 \(four autograd calls per step\) and∼16×\\sim 16\\timesslower under Verlet \(three autograd calls per step\)\.#### Interpretation\.

On the 1\-DOF pendulum, the analytical vector field is substantially faster than the learned models\. The present benchmark therefore provides no evidence of a computational advantage for the learned surrogates\. Whether a structured learned model becomes computationally advantageous at higher dimension depends on the architecture, interaction representation, hardware, and number of bodies\. That scaling question is outside the scope of the present experiment and requires a dedicated many\-body benchmark\.

## 5Explicit Verlet\-style rollout: an informative diagnostic

The rightmost column of Table[3](https://arxiv.org/html/2608.10235#S4.T3)reports a representative seed\-42 diagnostic obtained by rolling out the same trained HNN with an explicit Störmer–Verlet\-style stepper rather than RK4\. AtT=100T=100, the diagnostic energy drift is7\.71×10−47\.71\\times 10^\{\-4\}, compared with2\.89×10−42\.89\\times 10^\{\-4\}for HNN\+RK4 on the same seed, a factor of approximately2\.72\.7\. This auxiliary comparison is separate from the primary pendulum evaluation, for which both Model A and Model B are reported overNseed=5N\_\{\\mathrm\{seed\}\}=5independent training seeds\. It is interpreted only as an integrator\-compatibility diagnostic\.

The standard velocity\-Verlet splitting is derived for separable Hamiltonians of the form

H​\(q,p\)=T​\(p\)\+V​\(q\),H\(q,p\)=T\(p\)\+V\(q\),with the updates

pn\+1/2\\displaystyle p\_\{n\+1/2\}=pn−Δ​t2​∇qV​\(qn\),\\displaystyle=p\_\{n\}\-\\tfrac\{\\Delta t\}\{2\}\\nabla\_\{q\}V\(q\_\{n\}\),\(6\)qn\+1\\displaystyle q\_\{n\+1\}=qn\+Δ​t​∇pT​\(pn\+1/2\),\\displaystyle=q\_\{n\}\+\\Delta t\\,\\nabla\_\{p\}T\(p\_\{n\+1/2\}\),\(7\)pn\+1\\displaystyle p\_\{n\+1\}=pn\+1/2−Δ​t2​∇qV​\(qn\+1\)\.\\displaystyle=p\_\{n\+1/2\}\-\\tfrac\{\\Delta t\}\{2\}\\nabla\_\{q\}V\(q\_\{n\+1\}\)\.\(8\)For the analytical pendulum this decomposition is exact:T​\(p\)=12​p2T\(p\)=\\tfrac\{1\}\{2\}p^\{2\}andV​\(q\)=1−cos⁡qV\(q\)=1\-\\cos q\.

For the learned HNN, however,H^θ​\(q,p\)\\widehat\{H\}\_\{\\theta\}\(q,p\)is represented by a generic scalar neural network and is not constrained to the separable formT​\(p\)\+V​\(q\)T\(p\)\+V\(q\)\. Therefore, the assumptions under which the standard velocity\-Verlet splitting is symplectic are not guaranteed to hold for the learned Hamiltonian\. We consequently interpret the HNN\+Verlet experiment as an integrator\-compatibility diagnostic rather than as a genuinely symplectic integration ofH^θ\\widehat\{H\}\_\{\\theta\}\.

Two distinct alternatives are relevant\. First, one may explicitly parameterize

H^​\(q,p\)=T^​\(p\)\+V^​\(q\),\\widehat\{H\}\(q,p\)=\\widehat\{T\}\(p\)\+\\widehat\{V\}\(q\),in which case a velocity\-Verlet splitting is compatible with the learned Hamiltonian by construction\. Second, one may retain a genericH^​\(q,p\)\\widehat\{H\}\(q,p\)and use a symplectic method applicable to general Hamiltonians, such as the implicit midpoint rule\. SympNets\[[6](https://arxiv.org/html/2608.10235#bib.bib6)\]provide a different structure\-preserving strategy: they construct symplectic neural maps directly and are designed to handle both separable and non\-separable Hamiltonian systems\.

This diagnostic therefore motivates a more precise transfer question: which structural assumptions required by a numerical integrator are actually enforced by the learned model?

## 6Extension to a 3D Hamiltonian system: the Kepler two\-body problem

To test whether the matched\-integrator result transfers beyond the 1\-DOF pendulum, we apply the same protocol to the Kepler two\-body problem in three\-dimensional configuration space\. The Hamiltonian is

H​\(q,p\)=12​‖p‖2−1‖q‖,q,p∈ℝ3,H\(q,p\)=\\frac\{1\}\{2\}\\\|p\\\|^\{2\}\-\\frac\{1\}\{\\\|q\\\|\},\\qquad q,p\\in\\mathbb\{R\}^\{3\},\(9\)withqqthe reduced\-body position,ppthe conjugate momentum, and reduced mass and gravitational constant set to unity\. Phase space is six\-dimensional\. The system conserves not only energy but the full angular\-momentum vectorL=q×pL=q\\times p, giving*four*scalar conserved quantities as opposed to the pendulum’s one\. Bound orbits \(H<0H<0\) are closed ellipses lying in the plane orthogonal toLL\.

#### Setup\.

Both models are widened to accommodate the higher input dimension:6→128→128→128→\{6,1\}6\\to 128\\to 128\\to 128\\to\\\{6,1\\\}withtanh\\tanhactivations \(34 69434\\,694and34 04934\\,049trainable parameters respectively\)\. Training data arentraj=60n\_\{\\mathrm\{traj\}\}=60trajectories of lengthT=6T=6atΔ​t=10−2\\Delta t=10^\{\-2\}with near\-circular initial conditions:q0q\_\{0\}uniform on spheres of radiusr0∈\[0\.9,1\.4\]r\_\{0\}\\in\[0\.9,1\.4\], andp0p\_\{0\}perpendicular toq0q\_\{0\}with‖p0‖\\\|p\_\{0\}\\\|equal to the circular\-orbit speed1/r0\\sqrt\{1/r\_\{0\}\}perturbed by up to±25%\\pm 25\\%\. This produces a near\-circular\-to\-moderately\-eccentric training set\. Random ICs with arbitraryppdirection can produce highly eccentric orbits that approach the origin, where the1/r1/rpotential is singular and the vector field becomes very large, making derivative targets poorly conditioned\. The train/validation/test split, optimiser, batch size, and epoch count match the pendulum protocol\. The Kepler comparison is likewise evaluated overNseed=5N\_\{\\mathrm\{seed\}\}=5independent training seeds, and both models are rolled out at inference by the same RK4 stepper\.

#### Metrics\.

In addition to trajectory MSE and energy drift, we report the maximum drift in the angular\-momentum vector,

Δ​L​\(T\)=maxt∈\[0,T\]⁡‖L​\(t\)−L​\(0\)‖2\.\\Delta L\(T\)=\\max\_\{t\\in\[0,T\]\}\\\|L\(t\)\-L\(0\)\\\|\_\{2\}\.\(10\)Angular\-momentum conservation provides an additional structural diagnostic beyond energy\. If the learned HamiltonianH^\\widehat\{H\}approximately inherits the rotational symmetry of the training dynamics, its induced flow should exhibit approximate angular\-momentum conservation\. Rotational invariance is not explicitly enforced by the generic HNN architecture used here, however, so this property is evaluated empirically\. Neither architecture explicitly enforces rotational equivariance in this experiment\.

#### Results\.

Table[7](https://arxiv.org/html/2608.10235#S6.T7)reports the matched\-integrator comparison at four horizons for trajectory MSE, energy drift, and angular\-momentum drift, averaged overNseed=5N\_\{\\mathrm\{seed\}\}=5independent training seeds with300300epochs each\. As a sanity check, RK4 on the analytical vector field conserves energy to∼7×10−12\\sim 7\\times 10^\{\-12\}and angular momentum to∼1×10−12\\sim 1\\times 10^\{\-12\}atT=10T=10\(double\-precision floor\)\. Figure[10](https://arxiv.org/html/2608.10235#S6.F10)shows the 3D orbit and the log\-scale conservation diagnostics for the first test trajectory of seed4242\.

Table 7:3D Kepler matched\-integrator comparison, mean±\\pmstd overNseed=5N\_\{\\mathrm\{seed\}\}=5independent training seeds \(ntraj=60n\_\{\\mathrm\{traj\}\}=60,300300epochs each\)\. The angular\-momentum\-drift ratio is an additional observable because angular momentum is not included explicitly in the training loss\. It tests whether the scalar\-Hamiltonian parameterization empirically captures the rotational structure of the Kepler trajectories more effectively under the present experimental conditions\.![Refer to caption](https://arxiv.org/html/2608.10235v1/x10.png)\(a\)3D orbit trajectory atT=10T=10\.
![Refer to caption](https://arxiv.org/html/2608.10235v1/x11.png)\(b\)Energy and angular\-momentum drift vs\. time \(log scale\)\.

Figure 10:3D Kepler rollouts\. Left: RK4 truth \(black\) vs Model A \(orange\) vs Model B, HNN \(green\); the central mass is at the origin\. Right: two\-panel conservation diagnostic showing energy drift\|H​\(t\)−H​\(0\)\|\|H\(t\)\-H\(0\)\|and angular\-momentum drift‖L​\(t\)−L​\(0\)‖2\\\|L\(t\)\-L\(0\)\\\|\_\{2\}on log scale\.
#### Interpretation\.

Three observations, averaged over five seeds\.

*The architectural advantage transfers to 3D at slightly smaller magnitude than the pendulum’s\.*AtT=10T=10\(about1\.61\.6orbital periods for the mean initial radius\), matched\-integrator energy drift is8\.0±1\.0×8\.0\\pm 1\.0\\timessmaller for Model B \(4\.20±1\.4×10−24\.20\\pm 1\.4\\times 10^\{\-2\}\) than for Model A \(3\.26±0\.80×10−13\.26\\pm 0\.80\\times 10^\{\-1\}\)\. The pendulum result atT=100T=100\(about1616periods\) was42×42\\times\(ratio of means, 5 seeds\); the Kepler ratio atT=10T=10is a factor of∼5\\sim 5smaller, which is consistent with a shorter test horizon relative to the period, a higher configuration\-space dimension against a fixed training\-data budget, and Kepler’s larger dynamic range of derivatives near perihelion\. The qualitative claim matched\-integrator HNN beats matched\-integrator standard NN on long\-horizon conservation holds, and the standard deviation across seeds \(1\.01\.0on the ratio\) is small enough that the effect is well outside noise\.

*Angular\-momentum drift provides an additional structural diagnostic\.*AtT=10T=10, Model B’s‖L​\(t\)−L​\(0\)‖2\\\|L\(t\)\-L\(0\)\\\|\_\{2\}is4\.2±0\.5×4\.2\\pm 0\.5\\timessmaller than Model A’s\. Because angular momentum is not included explicitly in the training loss, this result suggests that the learned scalar Hamiltonian captures the rotational structure of the training dynamics more accurately than the standard vector\-field baseline under the present experimental conditions\. The result should be interpreted empirically rather than as an automatic consequence of Hamiltonian parameterization, since rotational invariance is not explicitly built into either architecture\.

*Trajectory error grows with horizon at both networks; the ratio holds\.*Trajectory MSE atT=10T=10is5\.5±0\.8×5\.5\\pm 0\.8\\timessmaller for Model B\. The lower trajectory error is consistent with the improved conservation behavior, although the present experiment does not establish a direct causal relationship between the two quantities\.

*Scope\.*Two deliberate scoping decisions frame this section\. First, ICs are restricted to near\-circular\-to\-moderately\-eccentric orbits\. Random\-direction momenta can produce highly eccentric orbits that approach the1/r1/rsingularity, where the vector field becomes very large and the derivative targets become poorly conditioned\. The restriction isolates the transfer question from a data\-conditioning problem; extending to arbitrary eccentricity is a separate study that requires a soft\-core potentialH=12​‖p‖2−1/‖q‖2\+ϵ2H=\\tfrac\{1\}\{2\}\\\|p\\\|^\{2\}\-1/\\sqrt\{\\\|q\\\|^\{2\}\+\\epsilon^\{2\}\}, Kustaanheimo–Stiefel regularisation of the coordinates, or an orbital\-elements parametrisation that removes the singularity by construction\. Second, the training\-data budget \(6060trajectories, roughly25 00025\\,000state–derivative pairs\) is smaller per\-configuration\-space\-dimension than the pendulum run \(7070trajectories,∼70 000\\sim 70\\,000pairs on a 2D phase space\); if the transfer ratio attenuates further at fixed budget, data scale is the first knob to turn\.

## 7Discussion

#### What the matched\-integrator protocol shows\.

The42×42\\timesmean\-energy\-drift ratio and15\.8×15\.8\\timesmean\-trajectory\-MSE ratio atT=100T=100\(5 seeds\) are architectural\. The two networks were trained on identical data with identical targets, optimisers, capacities, and inference integrators; the only intentional difference is that Model B’s vector field is the symplectic gradient of a learned scalar rather than a direct prediction\. Under that control the structural prior on the dynamics buys1\.51\.5to22orders of magnitude on energy conservation and more than an order on trajectory error, and as important makes the HNN’s drift*predictable*across seeds where the standard NN’s drift varies by nearly an order of magnitude\. The bounded\-versus\-growing curve shapes make the argument qualitatively clear before any ratio is computed\. This is the argument that HNN papers often make; the contribution of this study is to make it under a clean control\.

#### Where the prior helps most\.

The energy\-stratified analysis shows the advantage widening with orbital nonlinearity: the standard NN leaks more energy on high\-energy orbits \(which sample more of the phase plane and the pendulum’s nonlinearities\), while the HNN’s drift is roughly constant\. This is consistent with a picture in which the standard NN learns a locally\-good vector field that fails to compose globally, while the HNN learns a globally\-good scalar and derives the vector field from it\.

#### Limitations\.

Three are worth naming\. \(i\)*Limited benchmark diversity\.*We evaluate two integrable conservative systems: the one\-degree\-of\-freedom nonlinear pendulum and the three\-dimensional Kepler two\-body problem\. Although the Kepler experiment increases phase\-space dimension from two to six and introduces rotational symmetry, neither benchmark tests chaotic, non\-integrable, dissipative, stochastic, or many\-body dynamics\. \(ii\)*Baseline choice\.*The comparison is against a standard MLP\. Comparison with Lagrangian neural networks, Neural ODE baselines, equivariant models, and symplectic\-map architectures such as SympNets under matched data budgets, parameter budgets, step sizes, and evaluation horizons would further strengthen the study\. \(iii\)*Compute setting\.*All timings are reported single\-CPU, float64\. GPU behavior, mixed precision, and the compute cost of autograd at higher configuration\-space dimension are not characterized\.

#### Future work\.

Four follow\-up directions fall directly out of the present results\.

*\(a\) Separable\-HNN ablation\.*An explicitly two\-headedT^θ​\(p\)\+V^ϕ​\(q\)\\widehat\{T\}\_\{\\theta\}\(p\)\+\\widehat\{V\}\_\{\\phi\}\(q\)architecture can test whether a learned separable Hamiltonian improves the explicit\-Verlet diagnostic under the same data and training controls\. SympNets provide a distinct symplectic\-map baseline rather than a separable\-HNN implementation\.

*\(b\) General\-Hamiltonian symplectic integration\.*An implicit midpoint or another symplectic Runge–Kutta rollout of the generic HNN would test the integrator side of the compatibility question without changing the learned Hamiltonian parameterization\.

*\(c\) Chaotic and non\-integrable systems\.*Natural next benchmarks include the Hénon–Heiles system, the double pendulum, and the restricted three\-body problem\. These systems would test whether the bounded\-drift advantage observed here persists when long\-term trajectory prediction becomes intrinsically sensitive or non\-integrable\.

*\(d\) Many\-body scaling\.*A dedicatedNN\-body benchmark is needed to determine whether the computational cost of a structured learned model can become competitive with direct force evaluation as system size increases\. Genuinely non\-separable Hamiltonians would also provide a useful test of architecture–integrator compatibility\.

## 8Conclusion

Under a controlled matched\-integrator protocol, a Hamiltonian neural network reduces mean long\-horizon energy drift by42×42\\timesand mean trajectory error by15\.8×15\.8\\timesover a parameter\-matched standard NN on the nonlinear pendulum \(five seeds\), with bounded rather than growing drift and substantially lower seed\-to\-seed variability\. The structural advantage also widens in more nonlinear pendulum regimes\.

The explicit\-Verlet diagnostic further shows that structure preservation cannot be inferred merely from combining a learned Hamiltonian with a method usually described as symplectic\. The assumptions required by the numerical method must also be compatible with the learned Hamiltonian\. Testing this point with an explicitly separable HNN and with an implicit symplectic method for general Hamiltonians is a direct next step\.

We do not claim that these results transfer directly to molecular dynamics or general many\-body systems\. The pendulum and three\-dimensional Kepler problem are controlled integrable benchmarks chosen to isolate the architectural mechanism\. Whether similar advantages persist for chaotic, non\-integrable, noisy, dissipative, and many\-body dynamics remains an empirical question\.

## References

- Bronstein et al\. \[2021\]M\. M\. Bronstein, J\. Bruna, T\. Cohen, and P\. Veličković\.Geometric deep learning: Grids, groups, graphs, geodesics, and gauges\.arXiv:2104\.13478, 2021\.
- Chen et al\. \[2018\]R\. T\. Q\. Chen, Y\. Rubanova, J\. Bettencourt, and D\. Duvenaud\.Neural Ordinary Differential Equations\.In*Advances in Neural Information Processing Systems 31*, 2018\. arXiv:1806\.07366\.
- Cranmer et al\. \[2020\]M\. Cranmer, S\. Greydanus, S\. Hoyer, P\. Battaglia, D\. Spergel, and S\. Ho\.Lagrangian Neural Networks\.ICLR 2020 Workshop on Deep Differential Equations\. arXiv:2003\.04630\.
- Greydanus et al\. \[2019\]S\. Greydanus, M\. Dzamba, and J\. Yosinski\.Hamiltonian Neural Networks\.In*Advances in Neural Information Processing Systems 32*, 2019\. arXiv:1906\.01563\.
- Hairer et al\. \[2006\]E\. Hairer, C\. Lubich, and G\. Wanner\.*Geometric Numerical Integration: Structure\-Preserving Algorithms for Ordinary Differential Equations*, 2nd ed\.Springer, 2006\.
- Jin et al\. \[2020\]P\. Jin, Z\. Zhang, A\. Zhu, Y\. Tang, and G\. E\. Karniadakis\.SympNets: Intrinsic structure\-preserving symplectic networks for identifying Hamiltonian systems\.*Neural Networks*, 132:166–179, 2020\. arXiv:2001\.03750\.
- Karniadakis et al\. \[2021\]G\. E\. Karniadakis, I\. G\. Kevrekidis, L\. Lu, P\. Perdikaris, S\. Wang, and L\. Yang\.Physics\-informed machine learning\.*Nature Reviews Physics*, 3:422–440, 2021\.
- Kingma and Ba \[2015\]D\. P\. Kingma and J\. Ba\.Adam: A method for stochastic optimization\.In*International Conference on Learning Representations*, 2015\. arXiv:1412\.6980\.
- Raissi et al\. \[2019\]M\. Raissi, P\. Perdikaris, and G\. E\. Karniadakis\.Physics\-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations\.*Journal of Computational Physics*, 378:686–707, 2019\.
- Toth et al\. \[2020\]P\. Toth, D\. J\. Rezende, A\. Jaegle, S\. Racanière, A\. Botev, and I\. Higgins\.Hamiltonian Generative Networks\.In*International Conference on Learning Representations*, 2020\. arXiv:1909\.13789\.
- Zhong et al\. \[2020\]Y\. D\. Zhong, B\. Dey, and A\. Chakraborty\.Symplectic ODE\-Net: Learning Hamiltonian dynamics with control\.In*International Conference on Learning Representations*, 2020\. arXiv:1909\.12077\.

相似文章

从微分几何视角看哈密顿神经网络

Reddit r/artificial

一篇博客文章,通过微分几何解释哈密顿神经网络,使用简单的质量-弹簧系统演示如何通过网络架构施加守恒定律以实现更高效的学习。作者从基础微积分开始逐步构建了辛流形和泊松括号等数学工具。

使用端口-哈密顿神经网络识别非线性弦动力学

arXiv cs.LG

本文扩展了端口-哈密顿神经网络(PHNNs)到偏微分方程(PDEs)中,用于从数据中学习非线性弦动力学。该方法能够同时恢复哈密顿量和耗散,在准确性和可解释性方面优于非物理信息基线方法。

扩散Fitzhugh-Nagumo模型中的均衡传播与哈密顿推断

arXiv cs.LG

本文将均衡传播扩展到斜梯度系统,并展示了深度能量模型与哈密顿神经网络之间的等价性,重点关注扩散耦合的Fitzhugh-Nagumo神经元。它还推导了此类网络中用于推理的逐层哈密顿递归关系。

解锁哈密顿视频动力学模型中的时间泛化

arXiv cs.LG

本文识别了哈密顿生成网络(HGN)在非保守环境中阻止时间泛化到不同步长的失败模式,并提出了针对性的修复方案,以实现可变时间分辨率下的稳定动力学预测。