Neural Non-Equilibrium Hamiltonian Monte Carlo for Corrected Boltzmann Sampling

arXiv cs.LG Papers

Summary

This paper introduces Neural Non-Equilibrium Hamiltonian Monte Carlo (NHMC), a train-then-correct method for sampling from unnormalized Boltzmann densities by learning stochastic Hamiltonian-style paths and correcting them using non-equilibrium work.

arXiv:2607.15682v1 Announce Type: new Abstract: Sampling from an unnormalized Boltzmann density requires proposals that move probability mass globally while retaining enough path-probability information for statistical correction. We introduce Neural Non-Equilibrium Hamiltonian Monte Carlo (NHMC), a train-then-correct learned Hamiltonian sampler. Starting from a tractable base distribution, NHMC learns stochastic Hamiltonian-style paths toward the target. Once training is complete, the learned proposal parameters are fixed; the proposal then generates complete paths and endpoint configurations, which are statistically corrected using the recorded non-equilibrium work. This dimensionless generalized work is determined by the probability ratio between the forward proposal path and a reverse reference path. During training, minimizing its mean reduces a path-space KL divergence and controls an upper bound on endpoint mismatch. During evaluation, the same quantity defines weights for self-normalized importance sampling on paths (path-SNIS), estimates normalizing constants or free-energy differences, and gives the acceptance ratio for path-space independent Metropolis-Hastings (path-IMH). The same forward-reverse laws also define a shared-bridge round-trip Metropolis kernel that acts directly on configurations and preserves the Boltzmann target. On double-well and finite-volume lattice $\phi^4$ targets, the NHMC construction gives corrected estimates when path overlap is sufficient; when overlap is poor, weight degeneracy, low acceptance, and long autocorrelation expose proposal failure. We additionally report a molecular internal-coordinate feasibility study using an MD prior and learned-force path proposal.
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:30 AM

# Neural Non-Equilibrium Hamiltonian Monte Carlo for Corrected Boltzmann Sampling
Source: [https://arxiv.org/html/2607.15682](https://arxiv.org/html/2607.15682)
Moxian Qian Helmholtz Institute Mainz Johannes Gutenberg University Mainz Institute of Molecular Biology \(IMB\) Mainz, Germany mqian@students\.uni\-mainz\.de

###### Abstract

Sampling from an unnormalized Boltzmann density requires proposals that move probability mass globally while retaining enough path\-probability information for statistical correction\. We introduce Neural Non\-Equilibrium Hamiltonian Monte Carlo \(NHMC\), a train\-then\-correct learned Hamiltonian sampler\. Starting from a tractable base distribution, NHMC learns stochastic Hamiltonian\-style paths toward the target\. Once training is complete, the learned proposal parameters are fixed; the proposal then generates complete paths and endpoint configurations, which are statistically corrected using the recorded non\-equilibrium work\. This dimensionless generalized work is determined by the probability ratio between the forward proposal path and a reverse reference path\. During training, minimizing its mean reduces a path\-space KL divergence and controls an upper bound on endpoint mismatch\. During evaluation, the same quantity defines weights for self\-normalized importance sampling on paths \(path\-SNIS\), estimates normalizing constants or free\-energy differences, and gives the acceptance ratio for path\-space independent Metropolis–Hastings \(path\-IMH\)\. The same forward–reverse laws also define a shared\-bridge round\-trip Metropolis kernel that acts directly on configurations and preserves the Boltzmann target\. On double\-well and finite\-volume latticeϕ4\\phi^\{4\}targets, the NHMC construction gives corrected estimates when path overlap is sufficient; when overlap is poor, weight degeneracy, low acceptance, and long autocorrelation expose proposal failure\. We additionally report a molecular internal\-coordinate feasibility study using an MD prior and learned\-force path proposal\.

## 1Introduction

Sampling from a Boltzmann distribution

π​\(x\)=1Z​exp⁡\[−E​\(x\)\]\\pi\(x\)=\\frac\{1\}\{Z\}\\exp\[\-E\(x\)\]\(1\)is a core problem in machine learning, physical simulation, and lattice field theory\. Classical Markov\-chain methods provide asymptotically valid equilibrium estimates under standard conditions, but local transitions can mix slowly between separated modes or metastable regions\. Neural samplers offer a complementary route by learning global proposals that move probability mass across configuration space\. A finite\-time learned transport, however, generally does not preserve the target distribution exactly\. A useful neural Boltzmann sampler should therefore combine global proposal moves with sufficient probability information to correct the resulting estimates or transitions\.

#### Nonequilibrium path correction\.

Nonequilibrium work identities provide one route to such a correction\. Jarzynski’s equality relates the exponential average of finite\-time work to an equilibrium free\-energy difference, whereas Crooks’ fluctuation relation identifies the corresponding forward–reverse trajectory probability ratio\(Jarzynski,[1997](https://arxiv.org/html/2607.15682#bib.bib23); Crooks,[1999](https://arxiv.org/html/2607.15682#bib.bib24)\)\. Bennett’s acceptance\-ratio method estimates free\-energy differences using samples from both directions\(Bennett,[1976](https://arxiv.org/html/2607.15682#bib.bib25)\)\. Annealed importance sampling \(AIS\) and Hamiltonian AIS construct importance weights on extended annealing paths, while sequential Monte Carlo combines related path weights with resampling\(Neal,[2001](https://arxiv.org/html/2607.15682#bib.bib26); Sohl\-Dickstein and Culpepper,[2012](https://arxiv.org/html/2607.15682#bib.bib27); Del Moralet al\.,[2006](https://arxiv.org/html/2607.15682#bib.bib33)\)\. Nonequilibrium candidate Monte Carlo instead uses nonequilibrium work in a Metropolis acceptance rule\(Nilmeieret al\.,[2011](https://arxiv.org/html/2607.15682#bib.bib13)\)\. Together, these methods show that finite\-time proposal paths can be statistically corrected when their probability ratios are retained\.

#### Learned samplers\.

Another line of work learns the sampling mechanism itself\. Hamiltonian Monte Carlo \(HMC\) augments the configuration with momenta and uses Hamiltonian integration to construct long\-range Metropolis proposals\(Duaneet al\.,[1987](https://arxiv.org/html/2607.15682#bib.bib1)\)\. Learned Markov chain Monte Carlo methods, including A\-NICE\-MC, L2HMC, and related Hamiltonian constructions, parameterize transition operators for an existing chain while retaining target invariance through a Metropolis correction\(Songet al\.,[2017](https://arxiv.org/html/2607.15682#bib.bib4); Levyet al\.,[2018](https://arxiv.org/html/2607.15682#bib.bib2); Foremanet al\.,[2021](https://arxiv.org/html/2607.15682#bib.bib3)\)\. Boltzmann Generators and lattice normalizing flows instead learn endpoint transports from a tractable reference distribution to target\-like configurations; the change\-of\-variables density supports importance weighting or flow\-based Metropolis correction\(Noéet al\.,[2019](https://arxiv.org/html/2607.15682#bib.bib5); Albergoet al\.,[2019](https://arxiv.org/html/2607.15682#bib.bib6); Del Debbioet al\.,[2021](https://arxiv.org/html/2607.15682#bib.bib7); Abbottet al\.,[2023](https://arxiv.org/html/2607.15682#bib.bib8)\)\.

#### Learned stochastic paths\.

Learned stochastic\-path methods connect endpoint transport with path\-space correction\. Stochastic normalizing flows combine deterministic invertible maps with stochastic sampling blocks and compute exact importance weights on the complete generative path\(Wuet al\.,[2020](https://arxiv.org/html/2607.15682#bib.bib12)\)\. Path Integral Sampler and Controlled Monte Carlo Diffusions learn controlled diffusion processes for unnormalized targets using stochastic\-control or path\-space objectives\(Zhang and Chen,[2022](https://arxiv.org/html/2607.15682#bib.bib30); Vargaset al\.,[2024](https://arxiv.org/html/2607.15682#bib.bib31)\)\. AFT, CRAFT, and Sequential Boltzmann Generators combine normalizing flows with annealed or sequential refinement\(Arbelet al\.,[2021](https://arxiv.org/html/2607.15682#bib.bib10); Matthewset al\.,[2022](https://arxiv.org/html/2607.15682#bib.bib14); Tanet al\.,[2025](https://arxiv.org/html/2607.15682#bib.bib16)\)\. FAB uses AIS\-bootstrapped samples to train a mass\-covering endpoint flow and reduce importance\-weight variance\(Midgleyet al\.,[2023](https://arxiv.org/html/2607.15682#bib.bib9)\)\.

Data\-driven equivariant flow matching learns transports from target samples\(Kleinet al\.,[2023](https://arxiv.org/html/2607.15682#bib.bib15)\)\. Energy\- and score\-based methods instead learn transports, reverse processes, likelihood surrogates, or stochastic controls from target energies or their gradients\(Woo and Ahn,[2024](https://arxiv.org/html/2607.15682#bib.bib34); Dernet al\.,[2025](https://arxiv.org/html/2607.15682#bib.bib17); Aggarwalet al\.,[2025](https://arxiv.org/html/2607.15682#bib.bib35); Vargaset al\.,[2023](https://arxiv.org/html/2607.15682#bib.bib18); Akhound\-Sadeghet al\.,[2024](https://arxiv.org/html/2607.15682#bib.bib19); OuYanget al\.,[2026](https://arxiv.org/html/2607.15682#bib.bib20); Heet al\.,[2025](https://arxiv.org/html/2607.15682#bib.bib21)\)\. NETS learns drift corrections along an annealed transport process while retaining work\-based weights for unbiasing\(Albergo and Vanden\-Eijnden,[2025](https://arxiv.org/html/2607.15682#bib.bib11)\)\. Adjoint Sampling learns a controlled diffusion toward the target and emphasizes the reuse of energy evaluations during training, rather than a fixed\-proposal SNIS or IMH correction after training\(Havenset al\.,[2025](https://arxiv.org/html/2607.15682#bib.bib22)\)\. NHMC occupies a Hamiltonian version of this space: its stochastic construction learns conditional forward and reverse momentum laws interleaved with reversible, volume\-preserving kick–drift maps\. The DW maps use the target energy gradient; the lattice implementations use learned position\-dependent kick fields\.

#### Neural non\-equilibrium Hamiltonian Monte Carlo\.

We introduce Neural Non\-Equilibrium Hamiltonian Monte Carlo \(NHMC\), a train\-then\-correct learned Hamiltonian sampler for unnormalized Boltzmann targets\. Starting from a tractable base distribution, NHMC learns a stochastic Hamiltonian\-style path toward the target\. Each proposal stage samples a momentum from a learned conditional density, applies an invertible volume\-preserving leapfrog\-type map, and records the corresponding forward and reverse log probabilities\. A stage may also include a recorded discrete involution, such as the sign flips used in the DW runs\. In the DW realization, the learned stochastic component is the conditional momentum law attached to each stage\. The lattice realizations additionally learn position\-dependent kick fields while retaining a reversible, volume\-preserving kick–drift composition\. These recorded quantities make the complete path likelihood ratio tractable even when an endpoint proposal density is unavailable\.

LetQθQ\_\{\\theta\}denote the learned forward law over complete pathsΓ\\Gamma, and letPθP\_\{\\theta\}denote a reverse\-reference path law whose endpoint marginal is the Boltzmann target\. With the normalized base convention,Δ​F=−log⁡Z\\Delta F=\-\\log Z, and the recorded work satisfies

log⁡d​Qθd​Pθ​\(Γ\)=Wθ​\(Γ\)−Δ​F\.\\log\\frac\{dQ\_\{\\theta\}\}\{dP\_\{\\theta\}\}\(\\Gamma\)=W\_\{\\theta\}\(\\Gamma\)\-\\Delta F\.\(2\)We refer toWθW\_\{\\theta\}as the dimensionless generalized work associated with the forward–reverse path\-density ratio; it need not coincide with mechanical work along a trajectory\. This identity yields the unnormalized path weightω~θ​\(Γ\)=exp⁡\[−Wθ​\(Γ\)\]\\tilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)=\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]and the Jarzynski normalization\. It also gives𝔼Qθ​\[Wθ\]−Δ​F=DKL​\(Qθ∥Pθ\)\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[W\_\{\\theta\}\]\-\\Delta F=D\_\{\\mathrm\{KL\}\}\(Q\_\{\\theta\}\\\|P\_\{\\theta\}\), so reducing average excess work reduces path\-space irreversibility\.

After training, the proposal parameters are fixed\. The recorded work then defines path\-space self\-normalized importance sampling \(path\-SNIS\), estimates normalizing constants or free\-energy differences, and gives the acceptance ratio for path\-space independent Metropolis–Hastings \(path\-IMH\)\. The objective is designed to improve endpoint overlap by reducing a path\-space KL divergence; the recorded path ratio supplies the statistical correction\.

#### Closest related methods\.

The closest related approaches are NETS, stochastic path sampling, and counterdiabatic Hamiltonian Monte Carlo\. NETS learns an additional drift along an annealed transport process and retains nonequilibrium weights for unbiasing\(Albergo and Vanden\-Eijnden,[2025](https://arxiv.org/html/2607.15682#bib.bib11)\)\. A stochastic path sampler learns forward and backward Langevin\-type path laws and applies path reweighting or extended\-space independent Metropolis–Hastings\(Chenet al\.,[2026b](https://arxiv.org/html/2607.15682#bib.bib28)\)\. Counterdiabatic HMC learns an auxiliary Hamiltonian term and uses nonequilibrium weights inside an SMC construction\(Cohn\-Gordonet al\.,[2026](https://arxiv.org/html/2607.15682#bib.bib29)\)\. NHMC instead learns bidirectional conditional momentum laws and, in its lattice realizations, position\-dependent kick fields within reversible, volume\-preserving kick–drift maps\. Its correction factor is the probability ratio of the realized stochastic Hamiltonian path, rather than an endpoint change\-of\-variables density\. A concurrent diffusion\-path construction uses an augmented\-space Metropolis correction to remove bias from approximate scores and discretization\(Chenet al\.,[2026a](https://arxiv.org/html/2607.15682#bib.bib32)\)\.

NHMC therefore provides a learned stochastic Hamiltonian proposal whose recorded work connects training, free\-energy estimation, path\-SNIS, and path\-IMH\. In addition, we derive a Metropolis\-adjusted round\-trip NHMC kernel: the reverse NHMC law pulls the current configuration back to a shared bridge state, and the forward law generates a candidate from that bridge\. A path\-swap Metropolis correction makes the resulting configuration\-space kernel reversible with respect to the Boltzmann target\. We evaluate the learned NHMC path proposal and its path\-SNIS, path\-IMH, and round\-trip corrections on analytically tractable many\-well targets and finite\-volume latticeϕ4\\phi^\{4\}\. A separate molecular internal\-coordinate study evaluates learned\-force prior\-action scores initialized from an MD prior\. The numerical evaluations report weighted observables together with path\-ESS, work spread, Metropolis acceptance, autocorrelation, and mode coverage\. Appendix[C](https://arxiv.org/html/2607.15682#A3)also reports a compactU​\(1\)U\(1\)round\-trip pilot as a gauge\-target stress test\.

## 2Neural Non\-Equilibrium Hamiltonian Monte Carlo

![Refer to caption](https://arxiv.org/html/2607.15682v1/x1.png)Figure 1:NHMC path construction and recorded\-work corrections\. After training, the learned proposal parameters are fixed, and the resulting proposal generates complete stochastic Hamiltonian paths\. The recorded work combines endpoint target/base terms with per\-stage forward–reverse momentum log\-ratios and, when present, discrete\-move log\-ratios\. The stage formula shown in the schematic suppresses this optional discrete term\. The same quantity supports path\-SNIS, normalizer/free\-energy estimation, and independent path\-IMH\. The round\-trip branch instead acts directly on configurations: starting from the currentxx, it traverses the forward\-oriented pathΓx:u→x\\Gamma\_\{x\}:u\\to xin reverse to a shared bridgeuu, followsΓy:u→y\\Gamma\_\{y\}:u\\to yto a candidate, and then makes one Metropolis decision\.NHMC defines a tractable probability law over complete stochastic paths\. Starting from a normalized base distributionq0q\_\{0\}, it keeps the auxiliary randomness inside the pathΓ\\Gamma\. The forward\-to\-reverse path ratio then remains computable even when the endpoint proposal density is unavailable\.

### 2\.1Stochastic Hamiltonian Path Proposal

Letγ​\(x\)=exp⁡\[−E​\(x\)\]\\gamma\(x\)=\\exp\[\-E\(x\)\],Z=∫γ​\(x\)​𝑑xZ=\\int\\gamma\(x\)\\,dx, andπ​\(x\)=γ​\(x\)/Z\\pi\(x\)=\\gamma\(x\)/Z\. A stage may first draw a discrete auxiliary variable

st∼rθ,tF\(⋅∣xt\),x~t=gst\(xt\),s\_\{t\}\\sim r^\{F\}\_\{\\theta,t\}\(\\cdot\\mid x\_\{t\}\),\\qquad\\widetilde\{x\}\_\{t\}=g\_\{s\_\{t\}\}\(x\_\{t\}\),\(3\)wheregstg\_\{s\_\{t\}\}is an invertible, volume\-preserving transformation\. When no discrete move is used,x~t=xt\\widetilde\{x\}\_\{t\}=x\_\{t\}and the corresponding probability factor is one\. With auxiliary momentumppand kinetic energyK​\(p\)=12​p⊤​M−1​pK\(p\)=\\frac\{1\}\{2\}p^\{\\top\}M^\{\-1\}p, NHMC then draws a fresh input momentum

pt\\displaystyle p\_\{t\}∼qθ,tF\(⋅∣x~t\),\\displaystyle\\sim q^\{F\}\_\{\\theta,t\}\(\\cdot\\mid\\widetilde\{x\}\_\{t\}\),qθ,tF\\displaystyle q^\{F\}\_\{\\theta,t\}=𝒩​\(μθ,tF​\(x~t\),diag​\(σθ,tF​\(x~t\)\)2\),\\displaystyle=\\mathcal\{N\}\\\!\\left\(\\mu^\{F\}\_\{\\theta,t\}\(\\widetilde\{x\}\_\{t\}\),\\mathrm\{diag\}\\,\(\\sigma^\{F\}\_\{\\theta,t\}\(\\widetilde\{x\}\_\{t\}\)\)^\{2\}\\right\),\(4\)and applies an invertible, reversible, volume\-preserving kick–drift map

\(xt\+1,p¯t\+1\)=Φt​\(x~t,pt\)\.\(x\_\{t\+1\},\\bar\{p\}\_\{t\+1\}\)=\\Phi\_\{t\}\(\\widetilde\{x\}\_\{t\},p\_\{t\}\)\.\(5\)The symbolptp\_\{t\}denotes the newly sampled input momentum;p¯t\+1\\bar\{p\}\_\{t\+1\}denotes the output momentum of the deterministic map\. The next stage samples a new momentum, so these two quantities are distinct\. In the DW realization,Φt\\Phi\_\{t\}is a leapfrog composition using the target\-energy gradient\. The lattice realizations use the same reversible, volume\-preserving kick–drift composition with learned position\-dependent kick fields\. These lattice fields are not assumed to be conservative or symplectic; the correction argument uses only invertibility, reversibility, and volume preservation\. All reported maps use state\-independent step sizes\. The complete forward path isΓ=\(x0,\{st,x~t,pt,xt\+1,p¯t\+1\}t=0L−1\)\\Gamma=\(x\_\{0\},\\\{s\_\{t\},\\widetilde\{x\}\_\{t\},p\_\{t\},x\_\{t\+1\},\\bar\{p\}\_\{t\+1\}\\\}\_\{t=0\}^\{L\-1\}\), with absentsts\_\{t\}variables omitted andx0∼q0x\_\{0\}\\sim q\_\{0\}\. Each stage is part of proposal construction, not a target\-preserving Markov transition by itself\. Target\-corrected outputs arise from path\-SNIS or path\-IMH after a complete path is generated, or from the paired reverse–forward construction of round\-trip NHMC\-MH\.

### 2\.2Reverse\-Reference Path and Recorded Work

For each forward stochastic law, NHMC also defines normalized reverse lawsqθ,tR​\(p∣x\)q^\{R\}\_\{\\theta,t\}\(p\\mid x\)and, when needed,rθ,tR​\(s∣x\)r^\{R\}\_\{\\theta,t\}\(s\\mid x\)\. The reverse\-reference law starts fromxL∼πx\_\{L\}\\sim\\piand is used to define the path probability ratio\. For a realized forward stage satisfying

Φt​\(x~t,pt\)=\(xt\+1,p¯t\+1\),\\Phi\_\{t\}\(\\widetilde\{x\}\_\{t\},p\_\{t\}\)=\(x\_\{t\+1\},\\bar\{p\}\_\{t\+1\}\),\(6\)time reversibility gives

Φt​\(xt\+1,−p¯t\+1\)=\(x~t,−pt\)\.\\Phi\_\{t\}\(x\_\{t\+1\},\-\\bar\{p\}\_\{t\+1\}\)=\(\\widetilde\{x\}\_\{t\},\-p\_\{t\}\)\.\(7\)Thus the reverse input momentum associated with forward stagettis−p¯t\+1\-\\bar\{p\}\_\{t\+1\}\. The earlier configuration is then recovered withxt=gst−1​\(x~t\)x\_\{t\}=g\_\{s\_\{t\}\}^\{\-1\}\(\\widetilde\{x\}\_\{t\}\)\. We writest†s\_\{t\}^\{\\dagger\}for the reverse label associated with this inverse operation, so thatgst†=gst−1g\_\{s\_\{t\}^\{\\dagger\}\}=g\_\{s\_\{t\}\}^\{\-1\}\. In the DW implementation,gstg\_\{s\_\{t\}\}is the recorded coordinate\-wise sign flip, hencest†=sts\_\{t\}^\{\\dagger\}=s\_\{t\}\.

LetQθQ\_\{\\theta\}be the forward path law and letPθP\_\{\\theta\}be the reverse\-reference law written on the same path coordinates\. The complete pathΓ\\Gammais the deterministic augmentation of the coordinates\(x0,s0,p0,…,sL−1,pL−1\)\(x\_\{0\},s\_\{0\},p\_\{0\},\\ldots,s\_\{L\-1\},p\_\{L\-1\}\)by the mapsΦt\\Phi\_\{t\}\. All path densities below are taken with respect to the common path\-coordinate measured​λ​\(Γ\)d\\lambda\(\\Gamma\)induced by these maps\. Volume preservation gives

Qθ​\(d​Γ\)\\displaystyle Q\_\{\\theta\}\(d\\Gamma\)=q0​\(x0\)​∏t=0L−1rθ,tF​\(st∣xt\)​qθ,tF​\(pt∣x~t\)​d​λ​\(Γ\),\\displaystyle=q\_\{0\}\(x\_\{0\}\)\\prod\_\{t=0\}^\{L\-1\}r^\{F\}\_\{\\theta,t\}\(s\_\{t\}\\mid x\_\{t\}\)q^\{F\}\_\{\\theta,t\}\(p\_\{t\}\\mid\\widetilde\{x\}\_\{t\}\)\\,d\\lambda\(\\Gamma\),\(8\)Pθ​\(d​Γ\)\\displaystyle P\_\{\\theta\}\(d\\Gamma\)=π​\(xL\)​∏t=0L−1rθ,tR​\(st†∣xt\+1\)​qθ,tR​\(−p¯t\+1∣xt\+1\)​d​λ​\(Γ\),\\displaystyle=\\pi\(x\_\{L\}\)\\prod\_\{t=0\}^\{L\-1\}r^\{R\}\_\{\\theta,t\}\(s\_\{t\}^\{\\dagger\}\\mid x\_\{t\+1\}\)q^\{R\}\_\{\\theta,t\}\\\!\\left\(\-\\bar\{p\}\_\{t\+1\}\\mid x\_\{t\+1\}\\right\)\\,d\\lambda\(\\Gamma\),\(9\)and the configuration endpoint marginal ofPθP\_\{\\theta\}isπ\\pi\. We call the following path\-density log ratio the dimensionless generalized recorded work:

Wθ​\(Γ\)\\displaystyle W\_\{\\theta\}\(\\Gamma\)=log⁡q0​\(x0\)−log⁡γ​\(xL\)\\displaystyle=\\log q\_\{0\}\(x\_\{0\}\)\-\\log\\gamma\(x\_\{L\}\)\(10\)\+∑t=0L−1\[logrθ,tF\(st∣xt\)−logrθ,tR\(st†∣xt\+1\)\\displaystyle\\quad\+\\sum\_\{t=0\}^\{L\-1\}\\left\[\\log r^\{F\}\_\{\\theta,t\}\(s\_\{t\}\\mid x\_\{t\}\)\-\\log r^\{R\}\_\{\\theta,t\}\(s\_\{t\}^\{\\dagger\}\\mid x\_\{t\+1\}\)\\right\.\+logqθ,tF\(pt∣x~t\)−logqθ,tR\(−p¯t\+1∣xt\+1\)\]\.\\displaystyle\\qquad\\qquad\\left\.\+\\log q^\{F\}\_\{\\theta,t\}\(p\_\{t\}\\mid\\widetilde\{x\}\_\{t\}\)\-\\log q^\{R\}\_\{\\theta,t\}\(\-\\bar\{p\}\_\{t\+1\}\\mid x\_\{t\+1\}\)\\right\]\.Thus

log⁡d​Qθd​Pθ​\(Γ\)\\displaystyle\\log\\frac\{dQ\_\{\\theta\}\}\{dP\_\{\\theta\}\}\(\\Gamma\)=Wθ​\(Γ\)\+log⁡Z\\displaystyle=W\_\{\\theta\}\(\\Gamma\)\+\\log Z=Wθ​\(Γ\)−Δ​F,Δ​F=−log⁡Z,\\displaystyle=W\_\{\\theta\}\(\\Gamma\)\-\\Delta F,\\qquad\\Delta F=\-\\log Z,\(11\)or equivalentlyd​Pθ/d​Qθ=exp⁡\[−Wθ​\(Γ\)\]/ZdP\_\{\\theta\}/dQ\_\{\\theta\}=\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]/Z\. Henceω~θ​\(Γ\)=exp⁡\[−Wθ​\(Γ\)\]\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)=\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]is an unnormalized path correction weight\. Extra stochastic moves contribute their forward\-minus\-reverse log probabilities toWθW\_\{\\theta\}, and non\-volume\-preserving deterministic maps contribute their Jacobian terms\. In concrete systems, the generation process can record the complete path\-probability information; thenexp⁡\[−Wθ​\(Γ\)\]\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]is computed from the recorded forward/reverse probabilities, Jacobian terms, and endpoint energies without requiring an explicit endpoint proposal density\.

### 2\.3Learning by Reducing Path\-Space Irreversibility

The theoretical core objective is the average recorded work,

ℒwork​\(θ\)=𝔼Γ∼Qθ​\[Wθ​\(Γ\)\]\.\\mathcal\{L\}\_\{\\mathrm\{work\}\}\(\\theta\)=\\mathbb\{E\}\_\{\\Gamma\\sim Q\_\{\\theta\}\}\[W\_\{\\theta\}\(\\Gamma\)\]\.\(12\)The reported DW and lattice runs optimize the finite\-batch objective

ℒtrain​\(θ\)=ℒwork​\(θ\)\+λvar​VarQθ⁡\[Wθ​\(Γ\)\],\\mathcal\{L\}\_\{\\mathrm\{train\}\}\(\\theta\)=\\mathcal\{L\}\_\{\\mathrm\{work\}\}\(\\theta\)\+\\lambda\_\{\\mathrm\{var\}\}\\,\\operatorname\{Var\}\_\{Q\_\{\\theta\}\}\[W\_\{\\theta\}\(\\Gamma\)\],\(13\)with coefficients given in Appendix[B](https://arxiv.org/html/2607.15682#A2)\. The variance term is an optimization regularizer; the path\-ratio identity and the post\-training corrections depend only on the recorded work\. Taking expectations in Eq\. \([11](https://arxiv.org/html/2607.15682#S2.E11)\) gives

ℒwork​\(θ\)−Δ​F=DKL​\(Qθ∥Pθ\)\.\\mathcal\{L\}\_\{\\mathrm\{work\}\}\(\\theta\)\-\\Delta F=D\_\{\\mathrm\{KL\}\}\(Q\_\{\\theta\}\\\|P\_\{\\theta\}\)\.\(14\)The single\-path workWθW\_\{\\theta\}is distinct from dissipated work,Wθ−Δ​FW\_\{\\theta\}\-\\Delta F\. SinceΔ​F\\Delta Fis independent ofθ\\theta, minimizing average work minimizes the path\-space irreversibility and, by marginalization, controls an upper bound on the endpoint proposal\-to\-target KLDKL​\(qθ,L∥π\)D\_\{\\mathrm\{KL\}\}\(q\_\{\\theta,L\}\\\|\\pi\)\. Gaussian momentum laws are trained by reparameterizingpt=μθ,tF​\(x~t\)\+σθ,tF​\(x~t\)⊙εtp\_\{t\}=\\mu^\{F\}\_\{\\theta,t\}\(\\widetilde\{x\}\_\{t\}\)\+\\sigma^\{F\}\_\{\\theta,t\}\(\\widetilde\{x\}\_\{t\}\)\\odot\\varepsilon\_\{t\}withεt∼𝒩​\(0,I\)\\varepsilon\_\{t\}\\sim\\mathcal\{N\}\(0,I\); the reverse law is evaluated at−p¯t\+1\-\\bar\{p\}\_\{t\+1\}\. The variance regularizer changes optimization behavior but not the correction identity\. Endpoint overlap is assessed separately through ESS, acceptance, and observable agreement\.

### 2\.4Evaluation with Path\-SNIS and Path\-IMH

#### Evaluation after training\.

Once training is complete, the learned proposal parameters are fixed\. Independent\-path and endpoint evaluations use four correction modes: endpoint SNIS, endpoint IMH, path\-SNIS, or path\-IMH\. The shared\-bridge construction in Section[3\.4](https://arxiv.org/html/2607.15682#S3.SS4)gives an additional configuration\-space Markov kernel\. Here SNIS denotes self\-normalized importance sampling, and IMH denotes independent Metropolis–Hastings\. For independent pathsΓi∼Qθ\\Gamma\_\{i\}\\sim Q\_\{\\theta\}, Jarzynski estimates the normalizing constant byZ^N=N−1​∑iexp⁡\[−Wθ​\(Γi\)\]\\widehat\{Z\}\_\{N\}=N^\{\-1\}\\sum\_\{i\}\\exp\[\-W\_\{\\theta\}\(\\Gamma\_\{i\}\)\]\. For an endpoint observableOO, path\-SNIS reports

𝔼π​\[O\]^=∑iω~i​O​\(xL\(i\)\)∑iω~i,ω~i=e−Wθ​\(Γi\)\.\\widehat\{\\mathbb\{E\}\_\{\\pi\}\[O\]\}=\\frac\{\\sum\_\{i\}\\widetilde\{\\omega\}\_\{i\}O\(x\_\{L\}^\{\(i\)\}\)\}\{\\sum\_\{i\}\\widetilde\{\\omega\}\_\{i\}\},\\qquad\\widetilde\{\\omega\}\_\{i\}=e^\{\-W\_\{\\theta\}\(\\Gamma\_\{i\}\)\}\.\(15\)Path\-IMH treats the complete path as the Markov\-chain state\. Given current pathΓ\\Gammaand proposalΓ′∼Qθ\\Gamma^\{\\prime\}\\sim Q\_\{\\theta\}, it accepts with probability

α​\(Γ,Γ′\)=1∧exp⁡\[−Wθ​\(Γ′\)\+Wθ​\(Γ\)\]\.\\alpha\(\\Gamma,\\Gamma^\{\\prime\}\)=1\\wedge\\exp\[\-W\_\{\\theta\}\(\\Gamma^\{\\prime\}\)\+W\_\{\\theta\}\(\\Gamma\)\]\.\(16\)The resulting path chain preservesPθP\_\{\\theta\}, whose endpoint marginal isπ\\pi\. SNIS gives consistent weighted estimates; ESS and work spread summarize finite\-sample efficiency\. IMH gives target\-invariant Markov kernels; acceptance and autocorrelation measure mixing\. When the endpoint densityqθ,L​\(x\)q\_\{\\theta,L\}\(x\)is tractable, endpoint SNIS and endpoint IMH are recovered by replacingω~θ​\(Γ\)\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)withγ​\(x\)/qθ,L​\(x\)\\gamma\(x\)/q\_\{\\theta,L\}\(x\)\. Direct comparisons to literature baselines use matched targets, budgets, correction modes, and observables\. Appendix[A](https://arxiv.org/html/2607.15682#A1)gives the endpoint, path, and general current\-state Metropolis ratios\.

## 3Correction from Recorded Path Ratios

The key correction step is the probability\-ratio relation between the forward proposal path lawQθQ\_\{\\theta\}and a reverse\-reference path lawPθP\_\{\\theta\}\. This relation ties together the recorded workWθ​\(Γ\)W\_\{\\theta\}\(\\Gamma\), the normalizing constantZZ, and the reverse\-reference law used for correction; Jarzynski normalization, path\-SNIS weighting, and path\-IMH acceptance all follow from it\. We use this relation to state the target of each estimator or chain, and then return to why average work is a useful training objective\.

#### Standing assumptions\.

Throughout,θ\\thetais fixed,0<Z=∫γ​\(x\)​dx<∞0<Z=\\int\\gamma\(x\)\\,\\mathrm\{d\}x<\\infty, andπ​\(x\)=γ​\(x\)/Z\\pi\(x\)=\\gamma\(x\)/Z\. Endpoint corrections require a tractable normalized endpoint densityqθ,Lq\_\{\\theta,L\}together with the usual support and moment conditions\. Path corrections use the forward/reverse construction of Section[2](https://arxiv.org/html/2607.15682#S2), with finite positive path weights\. When an endpoint density is unavailable, the path\-weight estimators and path\-IMH kernel are used instead\.

The distinction used below is simple\. SNIS statements concern weighted endpoint averages\. IMH statements concern target invariance of a chain whose state is either an endpoint or a complete path\. These are post\-training statements: training changes proposal overlap and finite\-sample efficiency; once the learned proposal is fixed, the SNIS weight or IMH acceptance rule determines the target law represented by the estimator or Markov chain\.

### 3\.1Forward–Reverse Path Ratio

Section[2\.2](https://arxiv.org/html/2607.15682#S2.SS2)defines the pathΓ\\Gamma, the forward path lawQθQ\_\{\\theta\}, the reverse\-reference path lawPθP\_\{\\theta\}, and the recorded workWθW\_\{\\theta\}in Eq\. \([10](https://arxiv.org/html/2607.15682#S2.E10)\)\. The reverse\-reference law starts fromxL∼πx\_\{L\}\\sim\\pi, so its endpoint marginal is the desired Boltzmann distribution\. This identity is the NHMC analogue of a Crooks\-type forward–reverse trajectory relation on the extended path space: the recorded work is the log ratio of the realized forward and reverse path laws, up to the free\-energy term\(Crooks,[1999](https://arxiv.org/html/2607.15682#bib.bib24); Nilmeieret al\.,[2011](https://arxiv.org/html/2607.15682#bib.bib13); Wuet al\.,[2020](https://arxiv.org/html/2607.15682#bib.bib12)\)\.

###### Theorem 1\(Path probability ratio\)\.

Assume thatq0q\_\{0\},qθ,tFq^\{F\}\_\{\\theta,t\},qθ,tRq^\{R\}\_\{\\theta,t\},rθ,tFr^\{F\}\_\{\\theta,t\}, andrθ,tRr^\{R\}\_\{\\theta,t\}are normalized densities or mass functions with respect to the chosen base measures\. Assume that each optionalgstg\_\{s\_\{t\}\}is invertible and volume preserving, and that eachΦt\\Phi\_\{t\}is invertible, volume preserving, and reversible under the momentum flipℛ​\(x,p\)=\(x,−p\)\\mathcal\{R\}\(x,p\)=\(x,\-p\), so thatΦt−1=ℛ∘Φt∘ℛ\\Phi\_\{t\}^\{\-1\}=\\mathcal\{R\}\\circ\\Phi\_\{t\}\\circ\\mathcal\{R\}, and thatQθQ\_\{\\theta\}andPθP\_\{\\theta\}are mutually absolutely continuous\. Then the reverse\-reference lawPθP\_\{\\theta\}is normalized, has endpoint marginalπ\\pi, and, almost everywhere on their common path support,

log⁡d​Qθd​Pθ​\(Γ\)\\displaystyle\\log\\frac\{dQ\_\{\\theta\}\}\{dP\_\{\\theta\}\}\(\\Gamma\)=Wθ​\(Γ\)\+log⁡Z,\\displaystyle=W\_\{\\theta\}\(\\Gamma\)\+\\log Z,\(17\)d​Pθd​Qθ​\(Γ\)\\displaystyle\\frac\{dP\_\{\\theta\}\}\{dQ\_\{\\theta\}\}\(\\Gamma\)=exp⁡\[−Wθ​\(Γ\)\]Z\.\\displaystyle=\\frac\{\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]\}\{Z\}\.\(18\)

###### Proof\.

The reverse construction first drawsxL∼πx\_\{L\}\\sim\\piand then draws normalized reverse discrete variables and momenta while reconstructing earlier states withΦt−1\\Phi\_\{t\}^\{\-1\}andgst−1g\_\{s\_\{t\}\}^\{\-1\}\. The reverse law is therefore normalized and has endpoint marginalπ\\pi\. Dividing Eqs\. \([8](https://arxiv.org/html/2607.15682#S2.E8)\) and \([9](https://arxiv.org/html/2607.15682#S2.E9)\) with respect to their common path\-coordinate measure gives

d​Qθd​Pθ​\(Γ\)\\displaystyle\\frac\{dQ\_\{\\theta\}\}\{dP\_\{\\theta\}\}\(\\Gamma\)=q0​\(x0\)​∏t=0L−1rθ,tF​\(st∣xt\)​qθ,tF​\(pt∣x~t\)π​\(xL\)​∏t=0L−1rθ,tR​\(st†∣xt\+1\)​qθ,tR​\(−p¯t\+1∣xt\+1\)\\displaystyle=\\frac\{q\_\{0\}\(x\_\{0\}\)\\prod\_\{t=0\}^\{L\-1\}r^\{F\}\_\{\\theta,t\}\(s\_\{t\}\\mid x\_\{t\}\)q^\{F\}\_\{\\theta,t\}\(p\_\{t\}\\mid\\widetilde\{x\}\_\{t\}\)\}\{\\pi\(x\_\{L\}\)\\prod\_\{t=0\}^\{L\-1\}r^\{R\}\_\{\\theta,t\}\(s\_\{t\}^\{\\dagger\}\\mid x\_\{t\+1\}\)q^\{R\}\_\{\\theta,t\}\(\-\\bar\{p\}\_\{t\+1\}\\mid x\_\{t\+1\}\)\}\(19\)=Z​exp⁡\[Wθ​\(Γ\)\],\\displaystyle=Z\\exp\[W\_\{\\theta\}\(\\Gamma\)\],\(20\)whereπ​\(xL\)=γ​\(xL\)/Z\\pi\(x\_\{L\}\)=\\gamma\(x\_\{L\}\)/Zand Eq\. \([10](https://arxiv.org/html/2607.15682#S2.E10)\) was used in the second line\. Taking logarithms gives the first identity, and inversion gives the second\. An expanded coordinate\-level derivation is given in Appendix[A\.8](https://arxiv.org/html/2607.15682#A1.SS8)\. ∎

#### Jacobian convention\.

Volume preservation removes deterministic\-map Jacobian terms\. For a non\-volume\-preserving map, the corresponding log\-Jacobian contribution must be included inWθW\_\{\\theta\}\. The deterministic maps used in the DW and lattice NHMC runs are volume preserving, so these terms vanish\. Thus subsequent corrections can be written using the same unnormalized path weight,

ω~θ​\(Γ\)=exp⁡\[−Wθ​\(Γ\)\]\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)=\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]\(21\)as shown below\.

### 3\.2Weighted Correction

###### Corollary 1\(Jarzynski normalization and path reweighting\)\.

Under the assumptions of Theorem[1](https://arxiv.org/html/2607.15682#Thmtheorem1),

𝔼Qθ​\[exp⁡\[−Wθ​\(Γ\)\]\]=Z\.\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]\]=Z\.\(22\)For any endpoint observableOOwith𝔼Qθ​\[ω~θ​\|O​\(xL\)\|\]<∞\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\widetilde\{\\omega\}\_\{\\theta\}\|O\(x\_\{L\}\)\|\]<\\infty,

𝔼π​\[O\]=𝔼Qθ​\[ω~θ​\(Γ\)​O​\(xL\)\]𝔼Qθ​\[ω~θ​\(Γ\)\]\.\\mathbb\{E\}\_\{\\pi\}\[O\]=\\frac\{\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)O\(x\_\{L\}\)\]\}\{\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)\]\}\.\(23\)

###### Proof\.

Integratingd​Pθ/d​Qθ=exp⁡\[−Wθ\]/ZdP\_\{\\theta\}/dQ\_\{\\theta\}=\\exp\[\-W\_\{\\theta\}\]/Zwith respect toQθQ\_\{\\theta\}gives the first identity\. Since the endpoint marginal ofPθP\_\{\\theta\}isπ\\pi,

𝔼π​\[O\]=𝔼Pθ​\[O​\(xL\)\]=1Z​𝔼Qθ​\[exp⁡\[−Wθ​\(Γ\)\]​O​\(xL\)\]\.\\mathbb\{E\}\_\{\\pi\}\[O\]=\\mathbb\{E\}\_\{P\_\{\\theta\}\}\[O\(x\_\{L\}\)\]=\\frac\{1\}\{Z\}\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]O\(x\_\{L\}\)\]\.\(24\)Substituting the first identity gives the stated ratio\. ∎

The empirical path\-SNIS estimator derived from Corollary[1](https://arxiv.org/html/2607.15682#Thmcorollary1)is strongly consistent by the strong law\. Reported SNIS observables are computed from this weighted empirical measure; unweighted endpoints remain proposal endpoints\. Ifqθ,Lq\_\{\\theta,L\}is tractable, the same argument gives endpoint SNIS withwθ​\(x\)=γ​\(x\)/qθ,L​\(x\)w\_\{\\theta\}\(x\)=\\gamma\(x\)/q\_\{\\theta,L\}\(x\)\.

### 3\.3Independent Metropolis–Hastings Correction

The independent Metropolis–Hastings correction is defined on the same state space as the proposal\. For path\-IMH, the Markov\-chain state is the complete trajectoryΓ\\Gamma, and observables use the endpointxLx\_\{L\}\. Each transition generates one full candidate pathΓ′∼Qθ\\Gamma^\{\\prime\}\\sim Q\_\{\\theta\}and then makes a single accept/reject decision fromWθW\_\{\\theta\}; there is no leapfrog\-level Metropolis step\. Related augmented\-trajectory Metropolis constructions appear in NCMC, stochastic normalizing flows, and SPS\(Nilmeieret al\.,[2011](https://arxiv.org/html/2607.15682#bib.bib13); Wuet al\.,[2020](https://arxiv.org/html/2607.15682#bib.bib12); Chenet al\.,[2026b](https://arxiv.org/html/2607.15682#bib.bib28)\)\.

###### Theorem 2\(Path\-IMH target invariance\)\.

Under the assumptions of Theorem[1](https://arxiv.org/html/2607.15682#Thmtheorem1), consider the independent path proposalΓ′∼Qθ\\Gamma^\{\\prime\}\\sim Q\_\{\\theta\}, accepted with probability

αθ​\(Γ,Γ′\)=1∧ω~θ​\(Γ′\)ω~θ​\(Γ\)=1∧exp⁡\[−Wθ​\(Γ′\)\+Wθ​\(Γ\)\]\.\\alpha\_\{\\theta\}\(\\Gamma,\\Gamma^\{\\prime\}\)=1\\wedge\\frac\{\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma^\{\\prime\}\)\}\{\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)\}=1\\wedge\\exp\[\-W\_\{\\theta\}\(\\Gamma^\{\\prime\}\)\+W\_\{\\theta\}\(\\Gamma\)\]\.\(25\)The resulting Markov kernel preservesPθP\_\{\\theta\}\. Hence, at stationarity, the endpointxLx\_\{L\}has marginal distributionπ\\pi\.

###### Proof\.

By Theorem[1](https://arxiv.org/html/2607.15682#Thmtheorem1),Pθ​\(d​Γ\)=Z−1​ω~θ​\(Γ\)​Qθ​\(d​Γ\)P\_\{\\theta\}\(d\\Gamma\)=Z^\{\-1\}\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)Q\_\{\\theta\}\(d\\Gamma\)\. For two distinct paths, the accepted transition measure is proportional to

Qθ​\(d​Γ\)​Qθ​\(d​Γ′\)​ω~θ​\(Γ\)​min⁡\{1,ω~θ​\(Γ′\)ω~θ​\(Γ\)\}\\displaystyle Q\_\{\\theta\}\(d\\Gamma\)Q\_\{\\theta\}\(d\\Gamma^\{\\prime\}\)\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)\\min\\left\\\{1,\\frac\{\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma^\{\\prime\}\)\}\{\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)\}\\right\\\}\(26\)=Qθ​\(d​Γ\)​Qθ​\(d​Γ′\)​min⁡\{ω~θ​\(Γ\),ω~θ​\(Γ′\)\}\.\\displaystyle\\qquad=Q\_\{\\theta\}\(d\\Gamma\)Q\_\{\\theta\}\(d\\Gamma^\{\\prime\}\)\\min\\\{\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\),\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma^\{\\prime\}\)\\\}\.\(27\)The last expression is symmetric inΓ\\GammaandΓ′\\Gamma^\{\\prime\}, so accepted moves satisfy detailed balance with respect toPθP\_\{\\theta\}\. Rejections are self\-loops and preserve detailed balance as well\. ThusPθP\_\{\\theta\}is invariant, and its endpoint marginal isπ\\pi\. Appendix[A\.10](https://arxiv.org/html/2607.15682#A1.SS10)gives the expanded path\-space derivation\. ∎

Endpoint IMH target invariance is the special case with statexx, proposalqθ,Lq\_\{\\theta,L\}, and weightwθ​\(x\)=γ​\(x\)/qθ,L​\(x\)w\_\{\\theta\}\(x\)=\\gamma\(x\)/q\_\{\\theta,L\}\(x\)\. The theorem establishes invariance; convergence from an arbitrary initialization additionally requires the usual irreducibility and aperiodicity conditions for the independent Metropolis kernel\.

### 3\.4Metropolis\-Adjusted Round\-Trip NHMC

The learned forward and reverse laws also define a Markov kernel whose state is the current configuration rather than a complete path\. Letfθ​\(a∣u\)f\_\{\\theta\}\(a\\mid u\)andrθ​\(b∣x\)r\_\{\\theta\}\(b\\mid x\)denote the normalized densities of the forward and reverse auxiliary path records\. The reversible, volume\-preserving path construction induces a bijection

Ψθ:\(u,a\)⟼\(x,b\),\\Psi\_\{\\theta\}:\(u,a\)\\longmapsto\(x,b\),\(28\)whereuuis a bridge configuration,xxis the endpoint, andbbis the reverse record associated with the realized forward path\.

Starting from a current configurationxx, drawb∼rθ\(⋅∣x\)b\\sim r\_\{\\theta\}\(\\cdot\\mid x\)and recover\(u,a\)=Ψθ−1​\(x,b\)\(u,a\)=\\Psi\_\{\\theta\}^\{\-1\}\(x,b\)\. Independently drawa′∼fθ\(⋅∣u\)a^\{\\prime\}\\sim f\_\{\\theta\}\(\\cdot\\mid u\)and set\(y,b′\)=Ψθ​\(u,a′\)\(y,b^\{\\prime\}\)=\\Psi\_\{\\theta\}\(u,a^\{\\prime\}\)\. Writing the two realized paths in the forward orientation asΓx:u→x\\Gamma\_\{x\}:u\\to xandΓy:u→y\\Gamma\_\{y\}:u\\to y, acceptyywith probability

αRT\\displaystyle\\alpha\_\{\\mathrm\{RT\}\}=1∧γ​\(y\)​rθ​\(b′∣y\)​fθ​\(a∣u\)γ​\(x\)​rθ​\(b∣x\)​fθ​\(a′∣u\)\\displaystyle=1\\wedge\\frac\{\\gamma\(y\)r\_\{\\theta\}\(b^\{\\prime\}\\mid y\)f\_\{\\theta\}\(a\\mid u\)\}\{\\gamma\(x\)r\_\{\\theta\}\(b\\mid x\)f\_\{\\theta\}\(a^\{\\prime\}\\mid u\)\}\(29\)=1∧exp⁡\[−Wθ​\(Γy\)\+Wθ​\(Γx\)\]\.\\displaystyle=1\\wedge\\exp\[\-W\_\{\\theta\}\(\\Gamma\_\{y\}\)\+W\_\{\\theta\}\(\\Gamma\_\{x\}\)\]\.\(30\)The second equality follows because both paths share the same bridge state, so the twoq0​\(u\)q\_\{0\}\(u\)terms in the recorded works cancel\.

###### Proposition 1\(Round\-trip target reversibility\)\.

Fixθ\\theta\. Supposefθf\_\{\\theta\}andrθr\_\{\\theta\}are normalized on their generated supports andΨθ\\Psi\_\{\\theta\}is a measure\-preserving bijection\. Then the round\-trip transition in Eqs\. \([29](https://arxiv.org/html/2607.15682#S3.E29)\)– \([30](https://arxiv.org/html/2607.15682#S3.E30)\) is reversible with respect toπ\\piand hence preserves the Boltzmann target\.

###### Proof\.

Forz=\(x,b,a′\)z=\(x,b,a^\{\\prime\}\), let\(u,a\)=Ψθ−1​\(x,b\)\(u,a\)=\\Psi\_\{\\theta\}^\{\-1\}\(x,b\)and defineξθ​\(z\)=π​\(x\)​rθ​\(b∣x\)​fθ​\(a′∣u\)\\xi\_\{\\theta\}\(z\)=\\pi\(x\)r\_\{\\theta\}\(b\\mid x\)f\_\{\\theta\}\(a^\{\\prime\}\\mid u\)\. Itsxx\-marginal isπ\\pi\. The path\-swap map

Tθ=\(Ψθ×I\)∘Swap∘\(Ψθ−1×I\)T\_\{\\theta\}=\(\\Psi\_\{\\theta\}\\times I\)\\circ\\operatorname\{Swap\}\\circ\(\\Psi\_\{\\theta\}^\{\-1\}\\times I\)\(31\)is a measure\-preserving involution and maps\(x,b,a′\)\(x,b,a^\{\\prime\}\)to\(y,b′,a\)\(y,b^\{\\prime\},a\)\. Consequently the first ratio in Eq\. \([29](https://arxiv.org/html/2607.15682#S3.E29)\) isξθ​\(Tθ​z\)/ξθ​\(z\)\\xi\_\{\\theta\}\(T\_\{\\theta\}z\)/\\xi\_\{\\theta\}\(z\)\. The accepted augmented flux ismin⁡\{ξθ​\(z\),ξθ​\(Tθ​z\)\}\\min\\\{\\xi\_\{\\theta\}\(z\),\\xi\_\{\\theta\}\(T\_\{\\theta\}z\)\\\}, which is invariant underz↔Tθ​zz\\leftrightarrow T\_\{\\theta\}z; rejected moves are self\-loops\. Projecting this detailed\-balance identity onto the configuration coordinate proves the claim\. Appendix[A\.11](https://arxiv.org/html/2607.15682#A1.SS11)gives the coordinate\-level argument\. ∎

The acceptance formula resembles path\-IMH, but the kernels are different: path\-IMH independently proposes a complete path and preservesPθP\_\{\\theta\}on path space, whereas round\-trip NHMC\-MH uses an actual reverse pullback and a shared bridge to preserveπ\\pidirectly on configuration space\. The path\-swap proof is an instance of the general involutive\-MCMC construction\(Neklyudovet al\.,[2020](https://arxiv.org/html/2607.15682#bib.bib41)\)\. The NHMC\-specific element is the shared\-bridge involution, for which the acceptance ratio reduces to a difference of recorded works\. It is also related to augmented nonequilibrium\-path Metropolis correction\(Nilmeieret al\.,[2011](https://arxiv.org/html/2607.15682#bib.bib13); Chenet al\.,[2026a](https://arxiv.org/html/2607.15682#bib.bib32)\)\. Convergence from an arbitrary initialization additionally requires irreducibility and aperiodicity of the projected configuration\-space kernel\.

### 3\.5Why Minimizing Work Helps

The same identity explains the training objective\. WritingΔ​F=−log⁡Z\\Delta F=\-\\log Z,

DKL​\(Qθ∥Pθ\)=𝔼Qθ​\[Wθ\]−Δ​F\.D\_\{\\mathrm\{KL\}\}\(Q\_\{\\theta\}\\\|P\_\{\\theta\}\)=\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[W\_\{\\theta\}\]\-\\Delta F\.\(32\)The recorded work differs from dissipated work by the free\-energy term: the dissipated work isWθ−Δ​FW\_\{\\theta\}\-\\Delta F\. SinceΔ​F\\Delta Fis independent ofθ\\theta, minimizing average work minimizes path\-space irreversibility\. Ifqθ,Lq\_\{\\theta,L\}is the endpoint marginal ofQθQ\_\{\\theta\}, data processing givesDKL​\(qθ,L∥π\)≤DKL​\(Qθ∥Pθ\)D\_\{\\mathrm\{KL\}\}\(q\_\{\\theta,L\}\\\|\\pi\)\\leq D\_\{\\mathrm\{KL\}\}\(Q\_\{\\theta\}\\\|P\_\{\\theta\}\)\. This explains the learning principle by upper\-bounding the endpoint proposal\-to\-target KL, while ESS, work dispersion, acceptance, and autocorrelation still measure finite\-sample efficiency\. Related path\-space objectives are used in stochastic normalizing flows, controlled Monte Carlo diffusions, and SPS\(Wuet al\.,[2020](https://arxiv.org/html/2607.15682#bib.bib12); Vargaset al\.,[2024](https://arxiv.org/html/2607.15682#bib.bib31); Chenet al\.,[2026b](https://arxiv.org/html/2607.15682#bib.bib28)\)\.

The results above apply only when the implemented path satisfies the stated support, invertibility, and Jacobian assumptions\. The molecular feasibility study uses boundary projections and is therefore reported separately rather than as evidence for Theorem[1](https://arxiv.org/html/2607.15682#Thmtheorem1)\. Onceθ⋆\\theta^\{\\star\}is fixed, path\-SNIS and path\-IMH apply toQθ⋆Q\_\{\\theta^\{\\star\}\}andPθ⋆P\_\{\\theta^\{\\star\}\}\. Appendix[A](https://arxiv.org/html/2607.15682#A1)gives the detailed balance and consistency arguments\.

## 4Numerical Results and Diagnostics

The numerical evaluations separate three roles of the recorded work\. Double\-well targets provide analytic normalizers, target samples, and known sign modes, enabling direct checks of normalizer estimation, path\-SNIS weighting, and path\-IMH correction\. The finite\-volume latticeϕ4\\phi^\{4\}benchmark reports corrected physical observables and compares independent path\-IMH with the shared\-bridge round\-trip kernel near a susceptibility peak\. The molecular study separately evaluates prior\-action scores for a learned\-force internal\-coordinate proposal initialized from an MD prior\.

### 4\.1Work, Weighting, and Path\-IMH Correction

The double\-well family used here provides analytic normalizing constants, exact target samples, and known sign\-mode structure\. This setting isolates the three evaluation\-time uses of path work: Jarzynski log\-normalizer estimation, path\-SNIS weighting, and path\-IMH acceptance\. For

log⁡γd​\(x\)=∑j=0d/2−1\[−x2​j4\+6​x2​j2\+0\.5​x2​j−12​x2​j\+12\],\\log\\gamma\_\{d\}\(x\)=\\sum\_\{j=0\}^\{d/2\-1\}\\left\[\-x\_\{2j\}^\{4\}\+6x\_\{2j\}^\{2\}\+0\.5x\_\{2j\}\-\\frac\{1\}\{2\}x\_\{2j\+1\}^\{2\}\\right\],\(33\)the continuous neural NHMC sampler starts from a Gaussian base, uses learned momentum laws with target\-force leapfrog updates, and uses SNIS or path\-IMH to correct the generated paths and endpoints\.

Table 1:Path correction on analytically tractable many\-well targets\. DW4 is the main learned path setting; DW8 gives a higher\-dimensional scaling result\. These selected main evaluations are distinct from the three\-checkpoint ensembles used for the round\-trip comparison below\. Normalizer errors are single\-evaluation point estimates and do not include across\-training uncertainty\.On DW4, the complete NHMC proposal, including its recorded stochastic sign\-flip move, reaches all four sign modes\. The path work yields\|log⁡Z^−log⁡Z\|=6\.91×10−3\|\\log\\widehat\{Z\}\-\\log Z\|=6\.91\\times 10^\{\-3\}, path\-SNIS ESS18\.30%18\.30\\%, and path\-IMH acceptance30\.15%30\.15\\%\. Path\-IMH shifts the empirical mode masses toward the asymmetric exact target induced by the linear tilt, with mode\-massL1L\_\{1\}error6\.46×10−46\.46\\times 10^\{\-4\}\. DW8 retains coverage of all sixteen sign modes and a small observed normalizer error in this evaluation, with path\-SNIS ESS2\.65%2\.65\\%, path\-IMH acceptance12\.70%12\.70\\%, and mode\-massL1L\_\{1\}error3\.02×10−23\.02\\times 10^\{\-2\}\. The round\-trip comparison uses three separately trained checkpoints per target; for DW4 these are 32\-stage, five\-leapfrog\-per\-stage proposals rather than the selected main evaluation in Table[1](https://arxiv.org/html/2607.15682#S4.T1)\. On DW4, the configuration\-space round\-trip kernel raises the mean acceptance from32\.78%32\.78\\%to35\.37%35\.37\\%and reducesτint​\(x0\)\\tau\_\{\\rm int\}\(x\_\{0\}\)per transition from2\.192\.19to1\.741\.74across three trained checkpoints\. Each chain is initialized by a Gaussian base draw followed by one forward NHMC path\. Because each round\-trip transition evaluates two paths, its force\-normalized effective\-sample rate is0\.630\.63times that of independent path\-IMH; Appendix[C](https://arxiv.org/html/2607.15682#A3)reports both views of the comparison\. On DW8, round\-trip NHMC\-MH raises mean acceptance from12\.62%12\.62\\%to15\.87%15\.87\\%and reducesτint​\(x0\)\\tau\_\{\\rm int\}\(x\_\{0\}\)per transition from9\.089\.08to4\.464\.46\. After accounting for its two paths per transition, its force\-normalized effective\-sample rate is1\.011\.01times that of independent path\-IMH, and both kernels cover all sixteen modes for every checkpoint\.

![Refer to caption](https://arxiv.org/html/2607.15682v1/figures/dw4_corrected_sampling.png)Figure 2:Neural NHMC on the DW4 target\. The complete learned path proposal, including the recorded stochastic sign\-flip move, covers all four sign modes, while path\-IMH correction shifts the empirical mode masses to the asymmetric target induced by the linear tilt\. White contours give the exact target reference\.
### 4\.2Finite\-Volume Latticeϕ4\\phi^\{4\}

The two\-dimensional8×88\\times 8latticeϕ4\\phi^\{4\}benchmark reports corrected physical observables\. At zero external field the finite\-volume target has the globalZ2Z\_\{2\}symmetryϕx↦−ϕx\\phi\_\{x\}\\mapsto\-\\phi\_\{x\}\. HenceM=V−1​∑xϕxM=V^\{\-1\}\\sum\_\{x\}\\phi\_\{x\}has target mean zero when both sign sectors are represented\. We useχ=V​\(⟨M2⟩−⟨\|M\|⟩2\)\\chi=V\(\\langle M^\{2\}\\rangle\-\\langle\|M\|\\rangle^\{2\}\)andUB=1−⟨M4⟩/\(3​⟨M2⟩2\)U\_\{B\}=1\-\\langle M^\{4\}\\rangle/\(3\\langle M^\{2\}\\rangle^\{2\}\)together with\|M\|\|M\|to avoid sign cancellation\. TheZ2Z\_\{2\}sign is kept as a sector\-balance check; the observable estimates come from SNIS or IMH correction after training\.

![Refer to caption](https://arxiv.org/html/2607.15682v1/x2.png)Figure 3:Two\-dimensional latticeϕ4\\phi^\{4\}benchmark on the8×88\\times 8κ\\kappascan using NHMC proposal evaluations with SNIS correction\. Curves compare raw proposal observables, SNIS\-corrected observables, and the high\-statistics HMC reference for⟨\|M\|⟩\\langle\|M\|\\rangle, susceptibility, and the Binder cumulant\. The SNIS ESS fraction ranges from0\.0540\.054to0\.5990\.599across the scan\.Across theκ\\kappascan, path\-SNIS estimates of⟨\|M\|⟩\\langle\|M\|\\rangletrack the HMC reference, while the ESS fraction ranges from0\.5990\.599atκ=0\.20\\kappa=0\.20to0\.0650\.065atκ=0\.30\\kappa=0\.30, with its smallest value0\.0540\.054near the finite\-volume susceptibility peak\. The corrected observable stays close to HMC even where the raw proposal deviates; the falling ESS indicates increasing proposal mismatch\. The susceptibility and Binder\-cumulant panels use the same SNIS weights and the same HMC reference ensemble\. In a separate single\-chain autocorrelation check atκ=0\.2705\\kappa=0\.2705, HMC, independent path\-IMH, and round\-trip NHMC\-MH are compared with 4800 post\-burn values of\|M\|\|M\|\. HMC and path\-IMH discard the first 200 of 5000 saved states\. The round\-trip chain starts from a Gaussian base draw followed by one forward NHMC path, runs for 10000 transitions, discards 1000, and contributes the first 4800 remaining states to the matched\-length comparison\. The integrated autocorrelation times per saved transition are19\.9519\.95,8\.948\.94, and2\.572\.57, respectively\. A round\-trip transition evaluates two NHMC paths, giving a path\-cost\-normalized value of5\.155\.15; this remains below the independent path\-IMH value in this run\. The independent path\-IMH and round\-trip acceptance rates are36\.67%36\.67\\%and38\.92%38\.92\\%\. For HMC, path\-IMH, and round\-trip NHMC\-MH, respectively, the matched traces givep\+=0\.507,0\.502,0\.465p\_\{\+\}=0\.507,0\.502,0\.465and 80, 839, and 782 signed\-sector changes\. This is a representative single\-chain comparison rather than a multi\-chain performance estimate\.

### 4\.3Molecular Boltzmann Targets

The molecular feasibility study uses alanine peptide targets whose OpenMM systems are provided by bgmol MiniPeptide systems\. States are represented in internal coordinates and initialized from stored300​K300\\,\\mathrm\{K\}MD prior ensembles\. The proposal uses learned position\-dependent kick fields rather than OpenMM forces inside the leapfrog steps\. OpenMM supplies endpoint energies and first\-order forces for the internal\-coordinate energy and its training surrogate, without differentiating through the OpenMM force evaluations; the internal\-to\-Cartesian Jacobian is included in that energy\. Evaluation uses the prior\-action log score−EIC​\(zL\)\+EIC​\(z0\)−log⁡Ppath\-E\_\{\\rm IC\}\(z\_\{L\}\)\+E\_\{\\rm IC\}\(z\_\{0\}\)\-\\log P\_\{\\rm path\}\. The Ala4 and Ala6 rows in Table[2](https://arxiv.org/html/2607.15682#S4.T2)have finite scores for every reported evaluation sample\. Prior\-action score ESS is11\.00%11\.00\\%–39\.53%39\.53\\%, log\-score spread is0\.9660\.966–1\.4341\.434, and weighted energy\-overlap summaries are92\.76%92\.76\\%–93\.30%93\.30\\%\. Because the current implementation projects bond and angle coordinates back into their numerical domains, these rows are an empirical feasibility study and are not used as evidence for the invertible\-map claim of Theorem[1](https://arxiv.org/html/2607.15682#Thmtheorem1)\.

Table 2:Molecular learned\-force prior\-action score summaries\. All rows use NHMC proposal evaluations after training, with finite scores for all reported samples\. These empirical scores are not exact Radon–Nikodym correction weights for the continuous molecular target\.#### Compact gauge pilot\.

For compactU​\(1\)U\(1\), the inverse\-map and path\-ratio identities hold to numerical precision\. Acceptance is1\.05%1\.05\\%atβ=4\\beta=4and0\.43%0\.43\\%atβ=8\\beta=8, showing poor gauge\-field overlap for this trained proposal\. The pilot tests the implemented round\-trip ratio and local observables, not efficient equilibration or topology sampling\.

## 5Discussion

#### What the numerical results show\.

NHMC combines a learned global Boltzmann proposal with recorded\-work correction\. The learned Hamiltonian path moves probability mass between separated modes or regions\. Once training is complete, the recorded work determines the path\-SNIS weights or the path\-IMH acceptance rule, and therefore the target law represented by the final estimator or chain\. DW4 gives the cleanest validation: all sign modes are reached, the log\-normalizer error is small, and the ESS and path\-IMH acceptance are in a usable range\. Theϕ4\\phi^\{4\}results extend the same exact path correction to finite\-volume physical observables\. In the representative single\-chain check atκ=0\.2705\\kappa=0\.2705, the shared\-bridge round\-trip kernel gives the smallest autocorrelation per saved transition; this is not a multi\-chain performance estimate, and its two paths per transition must be included in cost comparisons\. The molecular study instead evaluates a learned\-force, MD\-initialized implementation empirically; its current boundary projections place it outside the invertible\-map theorem\.

#### Limitations and outlook\.

DW8 exhibits concentrated path weights and lower path\-IMH acceptance, whereas Ala6 has a more concentrated prior\-action score distribution\. Promising directions are stronger reverse momentum distributions, parameterizations that use known symmetries, and training objectives that reduce work variance while preserving support coverage\. The low\-acceptance compact\-U​\(1\)U\(1\)round\-trip pilot shows that formal target invariance alone does not overcome poor path overlap\. The currentϕ4\\phi^\{4\}study remains a finite\-volume test of corrected observables; larger\-scale field\-theory sampling requires separate study\.

## Acknowledgments

The author thanks Gert Aarts, Elia Cellini, Shiyang Chen, Chuan Liu, Miaoxin Liu, Biagio Lucini, Alessandro Nada, Jingkang Ouyang, Lukas Stelzl, Hartmut Wittig, Jianhui Zhang, Rui Zhang, and Kai Zhou for useful discussions\.

## References

- R\. Abbott, M\. S\. Albergo, A\. Botev, D\. Boyda, K\. Cranmer, D\. C\. Hackett, G\. Kanwar, A\. G\. D\. G\. Matthews, S\. Racanière, A\. Razavi, D\. J\. Rezende, F\. Romero\-López, P\. E\. Shanahan, and J\. M\. Urban \(2023\)Normalizing flows for Lattice Gauge Theory in arbitrary space\-time dimension\.arXiv preprint arXiv:2305\.02402\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2305.02402),[Link](https://arxiv.org/abs/2305.02402)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px2.p1.1)\.
- BoltzNCE: learning likelihoods for Boltzmann generation with stochastic interpolants and noise contrastive estimation\.InAdvances in Neural Information Processing Systems,Vol\.38\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2507.00846),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/f455010f4ff54a5d857fde29f14bdd3e-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p2.1)\.
- T\. Akhound\-Sadegh, J\. Rector\-Brooks, A\. J\. Bose, S\. Mittal, P\. Lemos, C\. Liu, M\. Sendera, S\. Ravanbakhsh, G\. Gidel, Y\. Bengio, N\. Malkin, and A\. Tong \(2024\)Iterated denoising energy matching for sampling from Boltzmann densities\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 760–786\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2402.06121),[Link](https://proceedings.mlr.press/v235/akhound-sadegh24a.html)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p2.1)\.
- M\. S\. Albergo, G\. Kanwar, and P\. E\. Shanahan \(2019\)Flow\-based generative models for Markov Chain Monte Carlo in Lattice Field Theory\.Physical Review D100\(3\),pp\. 034515\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevD.100.034515),[Link](https://doi.org/10.1103/PhysRevD.100.034515)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px2.p1.1)\.
- M\. S\. Albergo and E\. Vanden\-Eijnden \(2025\)NETS: a non\-equilibrium transport sampler\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 1026–1055\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.02711),[Link](https://proceedings.mlr.press/v267/albergo25a.html)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p2.1),[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px5.p1.1)\.
- M\. Arbel, A\. G\. D\. G\. Matthews, and A\. Doucet \(2021\)Annealed flow transport Monte Carlo\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 318–330\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2102.07501),[Link](https://proceedings.mlr.press/v139/arbel21a.html)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p1.1)\.
- C\. H\. Bennett \(1976\)Efficient estimation of free energy differences from monte carlo data\.Journal of Computational Physics22\(2\),pp\. 245–268\.External Links:[Document](https://dx.doi.org/10.1016/0021-9991%2876%2990078-4),[Link](https://doi.org/10.1016/0021-9991(76)90078-4)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px1.p1.1)\.
- H\. Chen, S\. Liu, and J\. Yang \(2026a\)Markov chain monte carlo with diffusion paths\.arXiv preprint arXiv:2607\.11631\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2607.11631),[Link](https://arxiv.org/abs/2607.11631)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px5.p1.1),[§3\.4](https://arxiv.org/html/2607.15682#S3.SS4.p3.2)\.
- S\. Chen, M\. Qian, G\. Aarts, B\. Lucini, and K\. Zhou \(2026b\)Stochastic path sampler for Lattice Field Theory\.arXiv preprint arXiv:2606\.13790\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2606.13790),[Link](https://arxiv.org/abs/2606.13790)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px5.p1.1),[§3\.3](https://arxiv.org/html/2607.15682#S3.SS3.p1.4),[§3\.5](https://arxiv.org/html/2607.15682#S3.SS5.p1.7)\.
- R\. Cohn\-Gordon, U\. Seljak, and D\. Sels \(2026\)Counterdiabatic Hamiltonian Monte Carlo\.arXiv preprint arXiv:2602\.21272\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.21272),[Link](https://arxiv.org/abs/2602.21272)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px5.p1.1)\.
- G\. E\. Crooks \(1999\)Entropy production fluctuation theorem and the nonequilibrium work relation for free energy differences\.Physical Review E60\(3\),pp\. 2721–2726\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevE.60.2721),[Link](https://arxiv.org/abs/cond-mat/9901352)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2607.15682#S3.SS1.p1.5)\.
- L\. Del Debbio, J\. Marsh Rossney, and M\. Wilson \(2021\)Efficient modeling of trivializing maps for latticeϕ4\\phi^\{4\}theory using normalizing flows: a first look at scalability\.Physical Review D104\(9\),pp\. 094507\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevD.104.094507),[Link](https://doi.org/10.1103/PhysRevD.104.094507)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px2.p1.1)\.
- P\. Del Moral, A\. Doucet, and A\. Jasra \(2006\)Sequential monte carlo samplers\.Journal of the Royal Statistical Society Series B: Statistical Methodology68\(3\),pp\. 411–436\.External Links:[Document](https://dx.doi.org/10.1111/j.1467-9868.2006.00553.x),[Link](https://doi.org/10.1111/j.1467-9868.2006.00553.x)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px1.p1.1)\.
- N\. Dern, L\. Redl, S\. Pfister, M\. Kollovieh, D\. Lüdke, and S\. Günnemann \(2025\)Energy\-weighted flow matching: unlocking continuous normalizing flows for efficient and scalable boltzmann sampling\.arXiv preprint arXiv:2509\.03726\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2509.03726),[Link](https://arxiv.org/abs/2509.03726)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p2.1)\.
- S\. Duane, A\. D\. Kennedy, B\. J\. Pendleton, and D\. Roweth \(1987\)Hybrid Monte Carlo\.Physics Letters B195\(2\),pp\. 216–222\.External Links:[Document](https://dx.doi.org/10.1016/0370-2693%2887%2991197-X),[Link](https://doi.org/10.1016/0370-2693(87)91197-X)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px2.p1.1)\.
- P\. Eastman, J\. Swails, J\. D\. Chodera, R\. T\. McGibbon, Y\. Zhao, K\. A\. Beauchamp, L\. Wang, A\. C\. Simmonett, M\. P\. Harrigan, C\. D\. Stern, R\. P\. Wiewiora, B\. R\. Brooks, and V\. S\. Pande \(2017\)OpenMM 7: rapid development of high performance algorithms for molecular dynamics\.PLOS Computational Biology13\(7\),pp\. e1005659\.External Links:[Document](https://dx.doi.org/10.1371/journal.pcbi.1005659),[Link](https://doi.org/10.1371/journal.pcbi.1005659)Cited by:[§B\.1](https://arxiv.org/html/2607.15682#A2.SS1.p4.4)\.
- S\. Foreman, X\. Jin, and J\. C\. Osborn \(2021\)Deep learning Hamiltonian Monte Carlo\.InICLR 2021 Workshop on Deep Learning for Simulation \(SimDL\),External Links:[Document](https://dx.doi.org/10.48550/arXiv.2105.03418),[Link](https://arxiv.org/abs/2105.03418)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px2.p1.1)\.
- A\. J\. Havens, B\. K\. Miller, B\. Yan, C\. Domingo\-Enrich, A\. Sriram, D\. S\. Levine, B\. M\. Wood, B\. Hu, B\. Amos, B\. Karrer, X\. Fu, G\. Liu, and R\. T\. Q\. Chen \(2025\)Adjoint sampling: highly scalable diffusion samplers via adjoint matching\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 22204–22237\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2504.11713),[Link](https://proceedings.mlr.press/v267/havens25a.html)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p2.1)\.
- J\. He, W\. Chen, M\. Zhang, D\. Barber, and J\. M\. Hernández\-Lobato \(2025\)Training neural samplers with reverse diffusive KL divergence\.InProceedings of the 28th International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.258,pp\. 5167–5175\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.12456),[Link](https://proceedings.mlr.press/v258/he25a.html)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p2.1)\.
- C\. Jarzynski \(1997\)Nonequilibrium equality for free energy differences\.Physical Review Letters78\(14\),pp\. 2690–2693\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevLett.78.2690),[Link](https://doi.org/10.1103/PhysRevLett.78.2690)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px1.p1.1)\.
- L\. Klein, A\. Krämer, and F\. Noé \(2023\)Equivariant flow matching\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2306.15030),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/bc827452450356f9f558f4e4568d553b-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p2.1)\.
- D\. Levy, M\. D\. Hoffman, and J\. Sohl\-Dickstein \(2018\)Generalizing Hamiltonian Monte Carlo with neural networks\.InInternational Conference on Learning Representations,External Links:[Document](https://dx.doi.org/10.48550/arXiv.1711.09268),[Link](https://arxiv.org/abs/1711.09268)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px2.p1.1)\.
- K\. Lindorff\-Larsen, S\. Piana, K\. Palmo, P\. Maragakis, J\. L\. Klepeis, R\. O\. Dror, and D\. E\. Shaw \(2010\)Improved side\-chain torsion potentials for the AMBER ff99sb protein force field\.Proteins: Structure, Function, and Bioinformatics78\(8\),pp\. 1950–1958\.External Links:[Document](https://dx.doi.org/10.1002/prot.22711),[Link](https://doi.org/10.1002/prot.22711)Cited by:[§B\.1](https://arxiv.org/html/2607.15682#A2.SS1.p4.4)\.
- M\. Lüscher \(1982\)Topology of Lattice Gauge Fields\.Communications in Mathematical Physics85,pp\. 39–48\.External Links:[Document](https://dx.doi.org/10.1007/BF02029132),[Link](https://doi.org/10.1007/BF02029132)Cited by:[§B\.1](https://arxiv.org/html/2607.15682#A2.SS1.p3.9)\.
- A\. G\. D\. G\. Matthews, M\. Arbel, D\. J\. Rezende, and A\. Doucet \(2022\)Continual repeated annealed flow transport Monte Carlo\.InProceedings of the 39th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.162,pp\. 15196–15219\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2201.13117),[Link](https://proceedings.mlr.press/v162/matthews22a.html)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p1.1)\.
- L\. I\. Midgley, V\. Stimper, G\. N\. C\. Simm, B\. Schölkopf, and J\. M\. Hernández\-Lobato \(2023\)Flow annealed importance sampling bootstrap\.InInternational Conference on Learning Representations,External Links:[Document](https://dx.doi.org/10.48550/arXiv.2208.01893),[Link](https://openreview.net/forum?id=XCTVFJwS9LJ)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p1.1)\.
- R\. M\. Neal \(2001\)Annealed importance sampling\.Statistics and Computing11\(2\),pp\. 125–139\.External Links:[Document](https://dx.doi.org/10.1023/A%3A1008923215028),[Link](https://doi.org/10.1023/A:1008923215028)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px1.p1.1)\.
- K\. Neklyudov, M\. Welling, E\. Egorov, and D\. Vetrov \(2020\)Involutive MCMC: a unifying framework\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 7273–7282\.External Links:[Link](https://proceedings.mlr.press/v119/neklyudov20a.html)Cited by:[§3\.4](https://arxiv.org/html/2607.15682#S3.SS4.p3.2)\.
- J\. P\. Nilmeier, G\. E\. Crooks, D\. D\. L\. Minh, and J\. D\. Chodera \(2011\)Nonequilibrium Candidate Monte Carlo is an efficient tool for equilibrium simulation\.Proceedings of the National Academy of Sciences108\(45\),pp\. E1009–E1018\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1106094108),[Link](https://doi.org/10.1073/pnas.1106094108)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2607.15682#S3.SS1.p1.5),[§3\.3](https://arxiv.org/html/2607.15682#S3.SS3.p1.4),[§3\.4](https://arxiv.org/html/2607.15682#S3.SS4.p3.2)\.
- F\. Noé, S\. Olsson, J\. Köhler, and H\. Wu \(2019\)Boltzmann Generators: sampling equilibrium states of many\-body systems with deep learning\.Science365\(6457\),pp\. eaaw1147\.External Links:[Document](https://dx.doi.org/10.1126/science.aaw1147),[Link](https://doi.org/10.1126/science.aaw1147)Cited by:[§B\.1](https://arxiv.org/html/2607.15682#A2.SS1.p4.4),[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px2.p1.1)\.
- A\. Onufriev, D\. Bashford, and D\. A\. Case \(2004\)Exploring protein native states and large\-scale conformational changes with a modified generalized born model\.Proteins: Structure, Function, and Bioinformatics55\(2\),pp\. 383–394\.External Links:[Document](https://dx.doi.org/10.1002/prot.20033),[Link](https://doi.org/10.1002/prot.20033)Cited by:[§B\.1](https://arxiv.org/html/2607.15682#A2.SS1.p4.4)\.
- R\. OuYang, B\. Qiang, and J\. M\. Hernández\-Lobato \(2026\)BNEM: a Boltzmann sampler based on bootstrapped noised energy matching\.Transactions on Machine Learning Research\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2409.09787),[Link](https://openreview.net/forum?id=ZZktU0U6Pu)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p2.1)\.
- J\. Sohl\-Dickstein and B\. J\. Culpepper \(2012\)Hamiltonian annealed importance sampling for partition function estimation\.arXiv preprint arXiv:1205\.1925\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.1205.1925),[Link](https://arxiv.org/abs/1205.1925)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px1.p1.1)\.
- J\. Song, S\. Zhao, and S\. Ermon \(2017\)A\-NICE\-MC: adversarial training for MCMC\.InAdvances in Neural Information Processing Systems,External Links:[Document](https://dx.doi.org/10.48550/arXiv.1706.07561),[Link](https://arxiv.org/abs/1706.07561)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px2.p1.1)\.
- C\. B\. Tan, A\. J\. Bose, C\. Lin, L\. Klein, M\. M\. Bronstein, and A\. Tong \(2025\)Scalable equilibrium sampling with sequential Boltzmann Generators\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 58467–58498\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2502.18462),[Link](https://proceedings.mlr.press/v267/tan25a.html)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p1.1)\.
- F\. Vargas, W\. Grathwohl, and A\. Doucet \(2023\)Denoising diffusion samplers\.InInternational Conference on Learning Representations,External Links:[Document](https://dx.doi.org/10.48550/arXiv.2302.13834),[Link](https://arxiv.org/abs/2302.13834)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p2.1)\.
- F\. Vargas, S\. Padhy, D\. Blessing, and N\. Nüsken \(2024\)Transport meets variational inference: controlled monte carlo diffusions\.InInternational Conference on Learning Representations,External Links:[Document](https://dx.doi.org/10.48550/arXiv.2307.01050),[Link](https://openreview.net/forum?id=PP1rudnxiW)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p1.1),[§3\.5](https://arxiv.org/html/2607.15682#S3.SS5.p1.7)\.
- D\. Woo and S\. Ahn \(2024\)Iterated energy\-based flow matching for sampling from boltzmann densities\.arXiv preprint arXiv:2408\.16249\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2408.16249),[Link](https://arxiv.org/abs/2408.16249)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p2.1)\.
- H\. Wu, J\. Köhler, and F\. Noé \(2020\)Stochastic normalizing flows\.InAdvances in Neural Information Processing Systems,Vol\.33\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2002.06707),[Link](https://proceedings.neurips.cc/paper/2020/hash/41d80bfc327ef980528426fc810a6d7a-Abstract.html)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2607.15682#S3.SS1.p1.5),[§3\.3](https://arxiv.org/html/2607.15682#S3.SS3.p1.4),[§3\.5](https://arxiv.org/html/2607.15682#S3.SS5.p1.7)\.
- Q\. Zhang and Y\. Chen \(2022\)Path integral sampler: a stochastic control approach for sampling\.InInternational Conference on Learning Representations,External Links:[Document](https://dx.doi.org/10.48550/arXiv.2111.15141),[Link](https://openreview.net/forum?id=_uCb2ynRu7Y)Cited by:[§1](https://arxiv.org/html/2607.15682#S1.SS0.SSS0.Px3.p1.1)\.

## Appendix APath Measures and Correction Arguments

### A\.1Forward and Reverse Path Laws

Throughout this appendix the configuration\-space target is a normalizable unnormalized densityγ​\(x\)=exp⁡\[−E​\(x\)\]\\gamma\(x\)=\\exp\[\-E\(x\)\]with0<Z=∫γ​\(x\)​dx<∞0<Z=\\int\\gamma\(x\)\\,\\mathrm\{d\}x<\\infty, and the learned parameterθ\\thetais fixed during evaluation\. The arguments depend on the fixed proposal law used at evaluation time, not on the particular training objective\. Endpoint corrections require a known endpoint proposal density andπ≪qθ\\pi\\ll q\_\{\\theta\}; path corrections require a computable Radon–Nikodym weight between the fixed forward path lawQθQ\_\{\\theta\}and a reverse\-reference path lawPθP\_\{\\theta\}whose configuration endpoint marginal isπ\\pi\. Deterministic maps are invertible and either volume preserving or accompanied by their exact Jacobian determinants; every stochastic choice appears through its forward and reverse transition probabilities\.

The main construction is the forward–reverse path ratio\. Endpoint SNIS and endpoint IMH are the corresponding special cases when the endpoint proposal density is available\. Table[3](https://arxiv.org/html/2607.15682#A1.T3)summarizes the estimators and chains before the detailed derivations\.

Table 3:Endpoint and path corrections used in the paper\.Endpoint\-SNIS and path\-SNIS supply consistent weighted estimates through their weighted empirical measures\. Endpoint\-IMH and path\-IMH define kernels that preserve their respective target laws, while finite\-chain efficiency still depends on proposal overlap, acceptance rate, and autocorrelation\. Round\-trip NHMC\-MH directly preserves the configuration\-space target by Proposition[1](https://arxiv.org/html/2607.15682#Thmproposition1)\.

### A\.2Boltzmann Marginal of the Extended State

For a configurationxxand momentumpp, definez=\(x,p\)z=\(x,p\),

Π​\(z\)=π​\(x\)​ρ​\(p\)=1Zx​Zp​exp⁡\[−E​\(x\)−K​\(p\)\]\.\\Pi\(z\)=\\pi\(x\)\\rho\(p\)=\\frac\{1\}\{Z\_\{x\}Z\_\{p\}\}\\exp\[\-E\(x\)\-K\(p\)\]\.\(34\)Sinceρ\\rhois normalized,

∫Π​\(x,p\)​𝑑p=π​\(x\)\.\\int\\Pi\(x,p\)\\,dp=\\pi\(x\)\.\(35\)Therefore any Markov kernel preservingΠ\\Piand refreshing or marginalizing momenta preserves the desired Boltzmann marginal\.

### A\.3Endpoint Independent Metropolis Step

Letπ​\(x\)=γ​\(x\)/Z\\pi\(x\)=\\gamma\(x\)/Zandwθ​\(x\)=γ​\(x\)/qθ​\(x\)w\_\{\\theta\}\(x\)=\\gamma\(x\)/q\_\{\\theta\}\(x\)\. The endpoint\-IMH accepted transition density is

qθ​\(y\)​αθ​\(x,y\),αθ​\(x,y\)=min⁡\{1,wθ​\(y\)wθ​\(x\)\}\.q\_\{\\theta\}\(y\)\\alpha\_\{\\theta\}\(x,y\),\\qquad\\alpha\_\{\\theta\}\(x,y\)=\\min\\left\\\{1,\\frac\{w\_\{\\theta\}\(y\)\}\{w\_\{\\theta\}\(x\)\}\\right\\\}\.\(36\)Forx≠yx\\neq y,

π​\(x\)​qθ​\(y\)​αθ​\(x,y\)\\displaystyle\\pi\(x\)q\_\{\\theta\}\(y\)\\alpha\_\{\\theta\}\(x,y\)=qθ​\(x\)​wθ​\(x\)Z​qθ​\(y\)​min⁡\{1,wθ​\(y\)wθ​\(x\)\}\\displaystyle=\\frac\{q\_\{\\theta\}\(x\)w\_\{\\theta\}\(x\)\}\{Z\}q\_\{\\theta\}\(y\)\\min\\left\\\{1,\\frac\{w\_\{\\theta\}\(y\)\}\{w\_\{\\theta\}\(x\)\}\\right\\\}\(37\)=qθ​\(x\)​qθ​\(y\)Z​min⁡\{wθ​\(x\),wθ​\(y\)\}\\displaystyle=\\frac\{q\_\{\\theta\}\(x\)q\_\{\\theta\}\(y\)\}\{Z\}\\min\\\{w\_\{\\theta\}\(x\),w\_\{\\theta\}\(y\)\\\}\(38\)=π​\(y\)​qθ​\(x\)​αθ​\(y,x\)\.\\displaystyle=\\pi\(y\)q\_\{\\theta\}\(x\)\\alpha\_\{\\theta\}\(y,x\)\.\(39\)The rejection term is a self\-loop and therefore also satisfies detailed balance\. Henceπ​Kθ=π\\pi K\_\{\\theta\}=\\pi\.

### A\.4Endpoint Weighted Estimates

Letxi∼qθx\_\{i\}\\sim q\_\{\\theta\}iid\. Sincewθ=γ/qθw\_\{\\theta\}=\\gamma/q\_\{\\theta\},

𝔼qθ​\[wθ​\(X\)\]=∫γ​\(x\)​dx=Z,𝔼qθ​\[wθ​\(X\)​O​\(X\)\]=Z​𝔼π​\[O\]\.\\mathbb\{E\}\_\{q\_\{\\theta\}\}\[w\_\{\\theta\}\(X\)\]=\\int\\gamma\(x\)\\,\\mathrm\{d\}x=Z,\\qquad\\mathbb\{E\}\_\{q\_\{\\theta\}\}\[w\_\{\\theta\}\(X\)O\(X\)\]=Z\\mathbb\{E\}\_\{\\pi\}\[O\]\.\(40\)The strong law of large numbers gives

∑iwθ​\(xi\)​O​\(xi\)∑iwθ​\(xi\)→a\.s\.𝔼π​\[O\],\\frac\{\\sum\_\{i\}w\_\{\\theta\}\(x\_\{i\}\)O\(x\_\{i\}\)\}\{\\sum\_\{i\}w\_\{\\theta\}\(x\_\{i\}\)\}\\xrightarrow\{\\mathrm\{a\.s\.\}\}\\mathbb\{E\}\_\{\\pi\}\[O\],\(41\)provided the required first moments exist\. The self\-normalized estimator is therefore consistent\. As usual for self\-normalized importance sampling, this is the property used for the weighted estimates in the empirical tables\.

### A\.5Endpoint Normalizer Estimates

The same endpoint weights give

Z^N=1N​∑i=1Nwθ​\(xi\)\.\\widehat\{Z\}\_\{N\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}w\_\{\\theta\}\(x\_\{i\}\)\.\(42\)Because𝔼qθ​\[wθ​\(X\)\]=Z\\mathbb\{E\}\_\{q\_\{\\theta\}\}\[w\_\{\\theta\}\(X\)\]=Z, this estimator is unbiased forZZat finiteNNand, by the strong law of large numbers,

Z^N→a\.s\.Z\.\\widehat\{Z\}\_\{N\}\\xrightarrow\{\\mathrm\{a\.s\.\}\}Z\.\(43\)SinceZ\>0Z\>0, the continuous\-mapping theorem gives

log⁡Z^N→a\.s\.log⁡Z\\log\\widehat\{Z\}\_\{N\}\\xrightarrow\{\\mathrm\{a\.s\.\}\}\\log Z\(44\)on the event thatZ^N\\widehat\{Z\}\_\{N\}is eventually positive\. Thus endpoint log\-normalizer errors used in the DW/toy summaries are consistent estimates when compared with an exact or reference value\. The log transform is generally biased at finiteNN, just as in the Jarzynski free\-energy estimator below\.

### A\.6Weight Concentration

For unnormalized nonnegative weightswiw\_\{i\}, define normalized weights

w¯i=wi∑j=1Nwj\.\\bar\{w\}\_\{i\}=\\frac\{w\_\{i\}\}\{\\sum\_\{j=1\}^\{N\}w\_\{j\}\}\.\(45\)The SNIS effective sample size reported in the paper is

ESSSNIS=1∑iw¯i2=\(∑iwi\)2∑iwi2,ESSSNISN=\(∑iwi\)2N​∑iwi2\.\\mathrm\{ESS\}\_\{\\mathrm\{SNIS\}\}=\\frac\{1\}\{\\sum\_\{i\}\\bar\{w\}\_\{i\}^\{2\}\}=\\frac\{\\left\(\\sum\_\{i\}w\_\{i\}\\right\)^\{2\}\}\{\\sum\_\{i\}w\_\{i\}^\{2\}\},\\qquad\\frac\{\\mathrm\{ESS\}\_\{\\mathrm\{SNIS\}\}\}\{N\}=\\frac\{\\left\(\\sum\_\{i\}w\_\{i\}\\right\)^\{2\}\}\{N\\sum\_\{i\}w\_\{i\}^\{2\}\}\.\(46\)This quantity is invariant to multiplying all weights by a positive constant, so unknown normalizing constants and free\-energy offsets do not change it\. It lies in\[1,N\]\[1,N\], with equality atNNonly when all finite weights are equal\. It measures finite\-sample weight concentration and is distinct from the autocorrelation\-based effective chain countN/\(2​τint\)N/\(2\\tau\_\{\\mathrm\{int\}\}\)used for Markov chains; unweighted proposal samples remain proposal samples\.

### A\.7Weighted Histograms and Mode Masses

Several figures use weighted mode masses, histograms, and action\-overlap panels\. These are all functions of the same weighted empirical measure

μ^Nw=∑i=1Nw¯i​δxi,w¯i=wi∑jwj\.\\widehat\{\\mu\}\_\{N\}^\{w\}=\\sum\_\{i=1\}^\{N\}\\bar\{w\}\_\{i\}\\delta\_\{x\_\{i\}\},\\qquad\\bar\{w\}\_\{i\}=\\frac\{w\_\{i\}\}\{\\sum\_\{j\}w\_\{j\}\}\.\(47\)For any bounded measurable bin, mode indicator, energy\-window indicator, or smoothed histogram test functionff,

∫f​\(x\)​μ^Nw​\(d​x\)=∑i=1Nw¯i​f​\(xi\)→a\.s\.𝔼π​\[f​\(X\)\]\\int f\(x\)\\,\\widehat\{\\mu\}\_\{N\}^\{w\}\(dx\)=\\sum\_\{i=1\}^\{N\}\\bar\{w\}\_\{i\}f\(x\_\{i\}\)\\xrightarrow\{\\mathrm\{a\.s\.\}\}\\mathbb\{E\}\_\{\\pi\}\[f\(X\)\]\(48\)under the same first\-moment assumptions as endpoint SNIS\. The path\-weight version follows by replacingwiw\_\{i\}andxix\_\{i\}withω~θ​\(Γi\)\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\_\{i\}\)and the endpointzL​\(Γi\)z\_\{L\}\(\\Gamma\_\{i\}\)\.

If a figure displays a SNIS\-resampled visual set, the resampling is from the discrete measureμ^Nw\\widehat\{\\mu\}\_\{N\}^\{w\}\. Conditional on the weighted sample, the resampled points are draws from this empirical approximation; at finiteNN, weighted histograms and resampled panels visualize the weighted empirical measure and its coverage of target features, while consistency or stationarity comes from the corresponding SNIS or IMH/path\-IMH result\.

### A\.8Coordinate\-Level Proof of Theorem[1](https://arxiv.org/html/2607.15682#Thmtheorem1)

The main theorem follows from the probability ratio between the explicit forward proposal path density and the reverse\-reference path density\. All path densities below are taken with respect to the common path\-coordinate measure induced by the deterministic maps\. Equivalently, a forward path is parameterized by\(x0,s0,p0,…,sL−1,pL−1\)\(x\_\{0\},s\_\{0\},p\_\{0\},\\ldots,s\_\{L\-1\},p\_\{L\-1\}\), with absentsts\_\{t\}variables omitted and all later configurations and output momenta obtained deterministically\. Write the displayed path as

Γ=\(x0,\{st,x~t,pt,xt\+1,p¯t\+1\}t=0L−1\),\\Gamma=\\left\(x\_\{0\},\\\{s\_\{t\},\\widetilde\{x\}\_\{t\},p\_\{t\},x\_\{t\+1\},\\bar\{p\}\_\{t\+1\}\\\}\_\{t=0\}^\{L\-1\}\\right\),wherex0∼q0x\_\{0\}\\sim q\_\{0\},st∼rθ,tF\(⋅∣xt\)s\_\{t\}\\sim r^\{F\}\_\{\\theta,t\}\(\\cdot\\mid x\_\{t\}\),x~t=gst​\(xt\)\\widetilde\{x\}\_\{t\}=g\_\{s\_\{t\}\}\(x\_\{t\}\),pt∼qθ,tF\(⋅∣x~t\)p\_\{t\}\\sim q^\{F\}\_\{\\theta,t\}\(\\cdot\\mid\\widetilde\{x\}\_\{t\}\), and\(xt\+1,p¯t\+1\)=Φt​\(x~t,pt\)\(x\_\{t\+1\},\\bar\{p\}\_\{t\+1\}\)=\\Phi\_\{t\}\(\\widetilde\{x\}\_\{t\},p\_\{t\}\)\. Hereptp\_\{t\}is the sampled input momentum at stagett, whilep¯t\+1\\bar\{p\}\_\{t\+1\}is the deterministic map output\. The optionalgstg\_\{s\_\{t\}\}maps are invertible and volume preserving; the DW sign flip is an involution with unit absolute Jacobian determinant\. The mapsΦt\\Phi\_\{t\}are invertible and volume preserving for the leapfrog updates used in the reported NHMC paths:

###### Lemma 1\(Kick–drift–kick invertibility and volume preservation\)\.

For a state\-independent step size, each kick or drift update is an invertible shear with unit Jacobian determinant\. Their compositionΦt\\Phi\_\{t\}is therefore invertible and volume preserving\. With momentum reversalℛ​\(x,p\)=\(x,−p\)\\mathcal\{R\}\(x,p\)=\(x,\-p\), the inverse map isΦt−1=ℛ∘Φt∘ℛ\\Phi\_\{t\}^\{\-1\}=\\mathcal\{R\}\\circ\\Phi\_\{t\}\\circ\\mathcal\{R\}\.

###### Proof\.

The momentum half\-step keepsxxfixed and translatesppby a differentiable function ofxx; the position step keepsppfixed and translatesxxby a function ofpp\. Each map is triangular with determinant one and has an inverse given by the same shear with the step sign reversed\. Composing the two half momentum steps and the position step preserves invertibility and determinant one\. The standard time\-reversal identity follows by applying the same sequence after flipping the momentum\. ∎

For non\-volume\-preserving deterministic maps, the Jacobian terms enter the density ratio described in the deterministic\-map discussion below\. With the volume\-preserving maps used here, the forward path density is

Qθ​\(d​Γ\)=q0​\(x0\)​∏t=0L−1rθ,tF​\(st∣xt\)​qθ,tF​\(pt∣x~t\)​d​λ​\(Γ\),Q\_\{\\theta\}\(d\\Gamma\)=q\_\{0\}\(x\_\{0\}\)\\prod\_\{t=0\}^\{L\-1\}r^\{F\}\_\{\\theta,t\}\(s\_\{t\}\\mid x\_\{t\}\)q^\{F\}\_\{\\theta,t\}\(p\_\{t\}\\mid\\widetilde\{x\}\_\{t\}\)\\,d\\lambda\(\\Gamma\),\(49\)whered​λd\\lambdadenotes the common path\-coordinate base measure induced by the deterministic maps\.

Now define a reverse reference law by first drawing the endpoint from the Boltzmann target,xL∼π​\(x\)=γ​\(x\)/Zx\_\{L\}\\sim\\pi\(x\)=\\gamma\(x\)/Z\. For a realized forward stage satisfying

Φt​\(x~t,pt\)=\(xt\+1,p¯t\+1\),\\Phi\_\{t\}\(\\widetilde\{x\}\_\{t\},p\_\{t\}\)=\(x\_\{t\+1\},\\bar\{p\}\_\{t\+1\}\),\(50\)time reversibility gives

Φt​\(xt\+1,−p¯t\+1\)=\(x~t,−pt\)\.\\Phi\_\{t\}\(x\_\{t\+1\},\-\\bar\{p\}\_\{t\+1\}\)=\(\\widetilde\{x\}\_\{t\},\-p\_\{t\}\)\.\(51\)Thus the reverse input momentum associated with forward stagettis−p¯t\+1\-\\bar\{p\}\_\{t\+1\}; applyinggst−1g\_\{s\_\{t\}\}^\{\-1\}after the inverse leapfrog reconstructsxtx\_\{t\}\. The reverse\-reference density on the same path coordinates is

Pθ​\(d​Γ\)=γ​\(xL\)Z​∏t=0L−1rθ,tR​\(st†∣xt\+1\)​qθ,tR​\(−p¯t\+1∣xt\+1\)​d​λ​\(Γ\)\.P\_\{\\theta\}\(d\\Gamma\)=\\frac\{\\gamma\(x\_\{L\}\)\}\{Z\}\\prod\_\{t=0\}^\{L\-1\}r^\{R\}\_\{\\theta,t\}\(s\_\{t\}^\{\\dagger\}\\mid x\_\{t\+1\}\)q^\{R\}\_\{\\theta,t\}\(\-\\bar\{p\}\_\{t\+1\}\\mid x\_\{t\+1\}\)\\,d\\lambda\(\\Gamma\)\.\(52\)It is normalized becauseπ\\pi, the reverse discrete laws, and the reverse momentum laws are normalized and the inverse deterministic reconstructions preserve volume\. Since the construction first drawsxL∼πx\_\{L\}\\sim\\piand then conditions all other path variables on this endpoint, integrating over the non\-endpoint variables leaves endpoint marginalπ​\(d​xL\)\\pi\(dx\_\{L\}\)\.

The ratio of the two path densities is therefore

log⁡d​Qθd​Pθ​\(Γ\)\\displaystyle\\log\\frac\{dQ\_\{\\theta\}\}\{dP\_\{\\theta\}\}\(\\Gamma\)=log⁡q0​\(x0\)−log⁡γ​\(xL\)\+log⁡Z\\displaystyle=\\log q\_\{0\}\(x\_\{0\}\)\-\\log\\gamma\(x\_\{L\}\)\+\\log Z\(53\)\+∑t=0L−1\[log⁡rθ,tF​\(st∣xt\)−log⁡rθ,tR​\(st†∣xt\+1\)\]\\displaystyle\\quad\+\\sum\_\{t=0\}^\{L\-1\}\\left\[\\log r^\{F\}\_\{\\theta,t\}\(s\_\{t\}\\mid x\_\{t\}\)\-\\log r^\{R\}\_\{\\theta,t\}\(s\_\{t\}^\{\\dagger\}\\mid x\_\{t\+1\}\)\\right\]\+∑t=0L−1\[log⁡qθ,tF​\(pt∣x~t\)−log⁡qθ,tR​\(−p¯t\+1∣xt\+1\)\]\\displaystyle\\quad\+\\sum\_\{t=0\}^\{L\-1\}\\left\[\\log q^\{F\}\_\{\\theta,t\}\(p\_\{t\}\\mid\\widetilde\{x\}\_\{t\}\)\-\\log q^\{R\}\_\{\\theta,t\}\(\-\\bar\{p\}\_\{t\+1\}\\mid x\_\{t\+1\}\)\\right\]\(54\)=Wθ​\(Γ\)\+log⁡Z\.\\displaystyle=W\_\{\\theta\}\(\\Gamma\)\+\\log Z\.\(55\)Equivalently,

d​Pθd​Qθ​\(Γ\)=exp⁡\[−Wθ​\(Γ\)\]Z\.\\frac\{dP\_\{\\theta\}\}\{dQ\_\{\\theta\}\}\(\\Gamma\)=\\frac\{\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]\}\{Z\}\.\(56\)Thusω~θ​\(Γ\)=exp⁡\[−Wθ​\(Γ\)\]\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)=\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]is an unnormalized path correction weight, and integrating the last display underQθQ\_\{\\theta\}gives𝔼Qθ​\[exp⁡\[−Wθ\]\]=Z\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\exp\[\-W\_\{\\theta\}\]\]=Z\. This completes the coordinate\-level proof of Theorem[1](https://arxiv.org/html/2607.15682#Thmtheorem1)\.

### A\.9Path Reweighting Identity

LetQθ​\(d​Γ\)Q\_\{\\theta\}\(d\\Gamma\)be the learned forward path measure and letPθ​\(d​Γ\)P\_\{\\theta\}\(d\\Gamma\)be the reverse reference path law constructed above\. Defineω~θ​\(Γ\)=exp⁡\[−Wθ​\(Γ\)\]\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)=\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]\. Assume0<𝔼Qθ​\[ω~θ\]<∞0<\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\widetilde\{\\omega\}\_\{\\theta\}\]<\\inftyand𝔼Qθ​\[ω~θ​\|O​\(xL\)\|\]<∞\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\widetilde\{\\omega\}\_\{\\theta\}\|O\(x\_\{L\}\)\|\]<\\infty\. Since

Pθ​\(d​Γ\)=ω~θ​\(Γ\)𝔼Qθ​\[ω~θ\]​Qθ​\(d​Γ\),P\_\{\\theta\}\(d\\Gamma\)=\\frac\{\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)\}\{\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\widetilde\{\\omega\}\_\{\\theta\}\]\}Q\_\{\\theta\}\(d\\Gamma\),\(57\)then

𝔼π​\[O\]\\displaystyle\\mathbb\{E\}\_\{\\pi\}\[O\]=𝔼Pθ​\[O​\(xL\)\]\\displaystyle=\\mathbb\{E\}\_\{P\_\{\\theta\}\}\[O\(x\_\{L\}\)\]\(58\)=𝔼Qθ​\[ω~θ​\(Γ\)​O​\(xL\)\]𝔼Qθ​\[ω~θ​\(Γ\)\]\.\\displaystyle=\\frac\{\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)O\(x\_\{L\}\)\]\}\{\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)\]\}\.\(59\)The empirical SNIS estimator follows by applying the strong law to numerator and denominator\.

### A\.10Detailed\-Balance Proof of Theorem[2](https://arxiv.org/html/2607.15682#Thmtheorem2)

LetPθ​\(d​Γ\)∝ω~θ​\(Γ\)​Qθ​\(d​Γ\)P\_\{\\theta\}\(d\\Gamma\)\\propto\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)Q\_\{\\theta\}\(d\\Gamma\)with0<𝔼Qθ​\[ω~θ\]<∞0<\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\widetilde\{\\omega\}\_\{\\theta\}\]<\\inftyand positive finite weights on the sampled support\. ProposeΓ′∼Qθ\\Gamma^\{\\prime\}\\sim Q\_\{\\theta\}\. The independent path acceptance probability is

αθ​\(Γ,Γ′\)=min⁡\{1,ω~θ​\(Γ′\)ω~θ​\(Γ\)\}\.\\alpha\_\{\\theta\}\(\\Gamma,\\Gamma^\{\\prime\}\)=\\min\\left\\\{1,\\frac\{\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma^\{\\prime\}\)\}\{\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)\}\\right\\\}\.\(60\)The accepted path flux is

Pθ​\(d​Γ\)​Qθ​\(d​Γ′\)​αθ​\(Γ,Γ′\)\\displaystyle P\_\{\\theta\}\(d\\Gamma\)Q\_\{\\theta\}\(d\\Gamma^\{\\prime\}\)\\alpha\_\{\\theta\}\(\\Gamma,\\Gamma^\{\\prime\}\)∝Qθ​\(d​Γ\)​Qθ​\(d​Γ′\)​min⁡\{ω~θ​\(Γ\),ω~θ​\(Γ′\)\},\\displaystyle\\propto Q\_\{\\theta\}\(d\\Gamma\)Q\_\{\\theta\}\(d\\Gamma^\{\\prime\}\)\\min\\\{\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\),\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma^\{\\prime\}\)\\\},\(61\)which is symmetric inΓ\\GammaandΓ′\\Gamma^\{\\prime\}\. Self\-loops handle rejections, soPθP\_\{\\theta\}is invariant\. Since the constructedPθP\_\{\\theta\}has configuration endpoint marginalπ\\pi, stationaryxLx\_\{L\}endpoints are Boltzmann distributed\. This argument proves invariance\. Convergence from an arbitrary initialization requires the usual irreducibility and aperiodicity conditions for the independent Metropolis kernel\.

### A\.11Round\-Trip NHMC\-MH Proof

Let𝖷\\mathsf\{X\}be configuration space with reference measureμ\\mu, and let𝖠\\mathsf\{A\}be the product space of all stochastic labels and momenta along a path, with product reference measuremm\. A forward record is denoted byaaand the associated reverse record bybb\.

For one stage, an optional discrete, measure\-preserving mapgt,stg\_\{t,s\_\{t\}\}is followed by a reversible, volume\-preserving kick–drift mapΦt\\Phi\_\{t\}:

x~t\\displaystyle\\widetilde\{x\}\_\{t\}=gt,st​\(xt\),\\displaystyle=g\_\{t,s\_\{t\}\}\(x\_\{t\}\),\(62\)\(xt\+1,p¯t\+1\)\\displaystyle\(x\_\{t\+1\},\\bar\{p\}\_\{t\+1\}\)=Φt​\(x~t,pt\),\\displaystyle=\\Phi\_\{t\}\(\\widetilde\{x\}\_\{t\},p\_\{t\}\),\(63\)pt†\\displaystyle p\_\{t\}^\{\\dagger\}=−p¯t\+1\.\\displaystyle=\-\\bar\{p\}\_\{t\+1\}\.\(64\)The inverse label satisfiesgt,st†=gt,st−1g\_\{t,s\_\{t\}^\{\\dagger\}\}=g\_\{t,s\_\{t\}\}^\{\-1\}\. SinceΦt−1=ℛ∘Φt∘ℛ\\Phi\_\{t\}^\{\-1\}=\\mathcal\{R\}\\circ\\Phi\_\{t\}\\circ\\mathcal\{R\}forℛ​\(x,p\)=\(x,−p\)\\mathcal\{R\}\(x,p\)=\(x,\-p\), the inverse stage is explicit:

\(x~t,−pt\)\\displaystyle\(\\widetilde\{x\}\_\{t\},\-p\_\{t\}\)=Φt​\(xt\+1,pt†\),\\displaystyle=\\Phi\_\{t\}\(x\_\{t\+1\},p\_\{t\}^\{\\dagger\}\),\(65\)xt\\displaystyle x\_\{t\}=gt,st−1​\(x~t\)\.\\displaystyle=g\_\{t,s\_\{t\}\}^\{\-1\}\(\\widetilde\{x\}\_\{t\}\)\.\(66\)Every factor in this stage map is invertible and preserves its reference measure\. Composition over the stages therefore gives a bimeasurable, measure\-preserving bijection

Ψθ:\(u,a\)⟼\(x,b\)\\Psi\_\{\\theta\}:\(u,a\)\\longmapsto\(x,b\)\(67\)from\(𝖷×𝖠,μ⊗m\)\(\\mathsf\{X\}\\times\\mathsf\{A\},\\mu\\otimes m\)to itself\. This is the coordinate statement needed below; a non\-volume\-preserving implementation would instead require its exact Jacobian in the acceptance ratio\.

Letfθ​\(a∣u\)f\_\{\\theta\}\(a\\mid u\)andrθ​\(b∣x\)r\_\{\\theta\}\(b\\mid x\)be the normalized forward and reverse auxiliary densities\. Given the currentxx, the implemented transition draws

b∼rθ\(⋅∣x\),\(u,a\)=Ψθ−1\(x,b\),a′∼fθ\(⋅∣u\),\(y,b′\)=Ψθ\(u,a′\)\.b\\sim r\_\{\\theta\}\(\\cdot\\mid x\),\\qquad\(u,a\)=\\Psi\_\{\\theta\}^\{\-1\}\(x,b\),\\qquad a^\{\\prime\}\\sim f\_\{\\theta\}\(\\cdot\\mid u\),\\qquad\(y,b^\{\\prime\}\)=\\Psi\_\{\\theta\}\(u,a^\{\\prime\}\)\.\(68\)Figure[4](https://arxiv.org/html/2607.15682#A1.F4)summarizes this shared\-bridge construction and the single acceptance decision made after both path maps have been evaluated\.

![Refer to caption](https://arxiv.org/html/2607.15682v1/x3.png)Figure 4:Shared\-bridge round\-trip NHMC\-MH transition\. Starting fromxkx\_\{k\}, the algorithm traverses the forward\-oriented pathΓx:u→xk\\Gamma\_\{x\}:u\\to x\_\{k\}in reverse to recoveruu, then followsΓy:u→y\\Gamma\_\{y\}:u\\to yto the candidate\. One Metropolis decision acceptsyyor retainsxkx\_\{k\}\. The curves are realized invertible path maps conditional on the sampled auxiliary records; their forward orientations define the two work values\.Forz=\(x,b,a′\)z=\(x,b,a^\{\\prime\}\), write\(u​\(x,b\),a​\(x,b\)\)=Ψθ−1​\(x,b\)\(u\(x,b\),a\(x,b\)\)=\\Psi\_\{\\theta\}^\{\-1\}\(x,b\)and define the augmented density, with respect toν=μ⊗m⊗m\\nu=\\mu\\otimes m\\otimes m, by

ξθ​\(z\)=π​\(x\)​rθ​\(b∣x\)​fθ​\(a′∣u​\(x,b\)\)\.\\xi\_\{\\theta\}\(z\)=\\pi\(x\)r\_\{\\theta\}\(b\\mid x\)f\_\{\\theta\}\(a^\{\\prime\}\\mid u\(x,b\)\)\.\(69\)Successive integration overa′a^\{\\prime\}andbbshows thatξθ\\xi\_\{\\theta\}is normalized and that itsxx\-marginal isπ\\pi\. No assumption on the induced bridge distribution is needed\.

Define the path\-swap map

Tθ​\(x,b,a′\)=\(y,b′,a\)=\(Ψθ×I\)∘Swap∘\(Ψθ−1×I\)​\(x,b,a′\),T\_\{\\theta\}\(x,b,a^\{\\prime\}\)=\(y,b^\{\\prime\},a\)=\(\\Psi\_\{\\theta\}\\times I\)\\circ\\operatorname\{Swap\}\\circ\(\\Psi\_\{\\theta\}^\{\-1\}\\times I\)\(x,b,a^\{\\prime\}\),\(70\)whereSwap⁡\(u,a,a′\)=\(u,a′,a\)\\operatorname\{Swap\}\(u,a,a^\{\\prime\}\)=\(u,a^\{\\prime\},a\)\. This conjugation immediately givesTθ2=IT\_\{\\theta\}^\{2\}=I\. Moreover,Ψθ\\Psi\_\{\\theta\}, its inverse, and the swap all preserve their respective product reference measures, so\(Tθ\)\#​ν=ν\(T\_\{\\theta\}\)\_\{\\\#\}\\nu=\\nu\.

Bijectivity givesΨθ−1​\(y,b′\)=\(u,a′\)\\Psi\_\{\\theta\}^\{\-1\}\(y,b^\{\\prime\}\)=\(u,a^\{\\prime\}\)\. Hence

ξθ​\(Tθ​z\)ξθ​\(z\)\\displaystyle\\frac\{\\xi\_\{\\theta\}\(T\_\{\\theta\}z\)\}\{\\xi\_\{\\theta\}\(z\)\}=π​\(y\)​rθ​\(b′∣y\)​fθ​\(a∣u\)π​\(x\)​rθ​\(b∣x\)​fθ​\(a′∣u\)\\displaystyle=\\frac\{\\pi\(y\)r\_\{\\theta\}\(b^\{\\prime\}\\mid y\)f\_\{\\theta\}\(a\\mid u\)\}\{\\pi\(x\)r\_\{\\theta\}\(b\\mid x\)f\_\{\\theta\}\(a^\{\\prime\}\\mid u\)\}\(71\)=γ​\(y\)​rθ​\(b′∣y\)​fθ​\(a∣u\)γ​\(x\)​rθ​\(b∣x\)​fθ​\(a′∣u\)\.\\displaystyle=\\frac\{\\gamma\(y\)r\_\{\\theta\}\(b^\{\\prime\}\\mid y\)f\_\{\\theta\}\(a\\mid u\)\}\{\\gamma\(x\)r\_\{\\theta\}\(b\\mid x\)f\_\{\\theta\}\(a^\{\\prime\}\\mid u\)\}\.\(72\)The unknown target normalizer cancels\. With the Metropolis probabilityαθ​\(z\)=1∧ξθ​\(Tθ​z\)/ξθ​\(z\)\\alpha\_\{\\theta\}\(z\)=1\\wedge\\xi\_\{\\theta\}\(T\_\{\\theta\}z\)/\\xi\_\{\\theta\}\(z\), the accepted augmented flux is

ξθ​\(z\)​αθ​\(z\)=min⁡\{ξθ​\(z\),ξθ​\(Tθ​z\)\}\.\\xi\_\{\\theta\}\(z\)\\alpha\_\{\\theta\}\(z\)=\\min\\\{\\xi\_\{\\theta\}\(z\),\\xi\_\{\\theta\}\(T\_\{\\theta\}z\)\\\}\.\(73\)For measurableA,B⊆𝖷A,B\\subseteq\\mathsf\{X\}, the accepted configuration\-space flow is therefore

J​\(A,B\)=∫𝟏A​\(x​\(z\)\)​𝟏B​\(x​\(Tθ​z\)\)​min⁡\{ξθ​\(z\),ξθ​\(Tθ​z\)\}​ν​\(d​z\)\.J\(A,B\)=\\int\\mathbf\{1\}\_\{A\}\(x\(z\)\)\\mathbf\{1\}\_\{B\}\(x\(T\_\{\\theta\}z\)\)\\min\\\{\\xi\_\{\\theta\}\(z\),\\xi\_\{\\theta\}\(T\_\{\\theta\}z\)\\\}\\,\\nu\(\\,\\mathrm\{d\}z\)\.\(74\)The change of variablesw=Tθ​zw=T\_\{\\theta\}z, together withTθ2=IT\_\{\\theta\}^\{2\}=Iand\(Tθ\)\#​ν=ν\(T\_\{\\theta\}\)\_\{\\\#\}\\nu=\\nu, givesJ​\(A,B\)=J​\(B,A\)J\(A,B\)=J\(B,A\)\. Rejections remain at the current configuration and are symmetric on the diagonal\. Thus the projected kernel satisfies detailed balance with respect toπ\\pi, proving Proposition[1](https://arxiv.org/html/2607.15682#Thmproposition1)\.

Finally, for the two forward\-oriented realized paths define

Wx\\displaystyle W\_\{x\}=log⁡q0​\(u\)−log⁡γ​\(x\)\+log⁡fθ​\(a∣u\)−log⁡rθ​\(b∣x\),\\displaystyle=\\log q\_\{0\}\(u\)\-\\log\\gamma\(x\)\+\\log f\_\{\\theta\}\(a\\mid u\)\-\\log r\_\{\\theta\}\(b\\mid x\),\(75\)Wy\\displaystyle W\_\{y\}=log⁡q0​\(u\)−log⁡γ​\(y\)\+log⁡fθ​\(a′∣u\)−log⁡rθ​\(b′∣y\)\.\\displaystyle=\\log q\_\{0\}\(u\)\-\\log\\gamma\(y\)\+\\log f\_\{\\theta\}\(a^\{\\prime\}\\mid u\)\-\\log r\_\{\\theta\}\(b^\{\\prime\}\\mid y\)\.\(76\)Because both paths share the sameuu, subtracting these expressions cancelslog⁡q0​\(u\)\\log q\_\{0\}\(u\)and turns Eq\. \([72](https://arxiv.org/html/2607.15682#A1.E72)\) into

αθ=1∧exp⁡\[−Wy\+Wx\]\.\\alpha\_\{\\theta\}=1\\wedge\\exp\[\-W\_\{y\}\+W\_\{x\}\]\.\(77\)This equality explains why independent path\-IMH and round\-trip NHMC\-MH have similar\-looking work\-difference formulas even though they act on different state spaces\.

### A\.12Work, Free Energy, and Path KL

The constructed path laws give

log⁡d​Qθd​Pθ​\(Γ\)=Wθ​\(Γ\)\+log⁡Z\.\\log\\frac\{dQ\_\{\\theta\}\}\{dP\_\{\\theta\}\}\(\\Gamma\)=W\_\{\\theta\}\(\\Gamma\)\+\\log Z\.\(78\)Equivalently, withΔ​F=−log⁡Z\\Delta F=\-\\log Zunder the normalized\-base convention,log⁡\(d​Qθ/d​Pθ\)=Wθ−Δ​F\\log\(dQ\_\{\\theta\}/dP\_\{\\theta\}\)=W\_\{\\theta\}\-\\Delta F\. Then

DKL​\(Qθ∥Pθ\)=𝔼Qθ​\[log⁡d​Qθd​Pθ\]=𝔼Qθ​\[Wθ\]−Δ​F\.D\_\{\\mathrm\{KL\}\}\(Q\_\{\\theta\}\\\|P\_\{\\theta\}\)=\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\\left\[\\log\\frac\{dQ\_\{\\theta\}\}\{dP\_\{\\theta\}\}\\right\]=\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[W\_\{\\theta\}\]\-\\Delta F\.\(79\)Thus dissipated work satisfies𝔼Qθ​\[Wθ\]−Δ​F=DKL​\(Qθ∥Pθ\)\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[W\_\{\\theta\}\]\-\\Delta F=D\_\{\\mathrm\{KL\}\}\(Q\_\{\\theta\}\\\|P\_\{\\theta\}\)exactly; minimizing𝔼Qθ​\[Wθ\]\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[W\_\{\\theta\}\]is equivalent to minimizing this path\-space KL becauseΔ​F\\Delta Fis fixed by the target and base convention\. Ifqθ,Lq\_\{\\theta,L\}andπ\\pidenote the configuration endpoint marginals ofQθQ\_\{\\theta\}andPθP\_\{\\theta\}, the chain rule for KL gives

DKL​\(qθ,L∥π\)≤DKL​\(Qθ∥Pθ\)\.D\_\{\\mathrm\{KL\}\}\(q\_\{\\theta,L\}\\\|\\pi\)\\leq D\_\{\\mathrm\{KL\}\}\(Q\_\{\\theta\}\\\|P\_\{\\theta\}\)\.\(80\)Therefore minimizing average dissipated work controls an upper bound on the endpoint proposal\-to\-target KLDKL​\(qθ,L∥π\)D\_\{\\mathrm\{KL\}\}\(q\_\{\\theta,L\}\\\|\\pi\)\.

### A\.13Path Weights, Work Weights, and Free Energy

The work convention also fixes how the path weight used for correction relates to the free\-energy estimate\. From the constructed density ratio

d​Pθd​Qθ​\(Γ\)=exp⁡\[−Wθ​\(Γ\)\]Z,\\frac\{dP\_\{\\theta\}\}\{dQ\_\{\\theta\}\}\(\\Gamma\)=\\frac\{\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]\}\{Z\},\(81\)Thus the correction weight may be taken as either the normalized ratiod​Pθ/d​QθdP\_\{\\theta\}/dQ\_\{\\theta\}or the empirical work weightexp⁡\[−Wθ\]\\exp\[\-W\_\{\\theta\}\]; the two differ by the positive constantZ−1Z^\{\-1\}\. More generally,

ω~θ​\(Γ\)=c​exp⁡\[−Wθ​\(Γ\)\]\\widetilde\{\\omega\}\_\{\\theta\}\(\\Gamma\)=c\\,\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]\(82\)for anyc\>0c\>0gives the same SNIS estimator and the same path\-IMH acceptance probability, because the constant cancels between numerator and denominator or between proposed and current path weights\. Integratingd​Pθ/d​QθdP\_\{\\theta\}/dQ\_\{\\theta\}underQθQ\_\{\\theta\}gives

1=𝔼Qθ​\[exp⁡\[−Wθ\]Z\],𝔼Qθ​\[exp⁡\[−Wθ\]\]=Z=e−Δ​F,1=\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\\left\[\\frac\{\\exp\[\-W\_\{\\theta\}\]\}\{Z\}\\right\],\\qquad\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[\\exp\[\-W\_\{\\theta\}\]\]=Z=e^\{\-\\Delta F\},\(83\)whereΔ​F=−log⁡Z\\Delta F=\-\\log Z\. This is the Jarzynski identity under the sign convention used in the paper\.

### A\.14Jacobian and Volume Preservation

For a deterministic map used in a path proposal, unit Jacobian determinant gives

\|det∂zL∂z0\|=1\\left\|\\det\\frac\{\\partial z\_\{L\}\}\{\\partial z\_\{0\}\}\\right\|=1\(84\)and contributes no additional density term\. Ifzk\+1=Tk,θ​\(zk\)z\_\{k\+1\}=T\_\{k,\\theta\}\(z\_\{k\}\)andJk,θ=\|det∂Tk,θ​\(zk\)/∂zk\|J\_\{k,\\theta\}=\|\\det\\partial T\_\{k,\\theta\}\(z\_\{k\}\)/\\partial z\_\{k\}\|, a deterministic work convention is

Wθ​\(Γ\)=∑k=0L−1\[Hk\+1​\(zk\+1\)−Hk​\(zk\)−log⁡Jk,θ\]\.W\_\{\\theta\}\(\\Gamma\)=\\sum\_\{k=0\}^\{L\-1\}\\left\[H\_\{k\+1\}\(z\_\{k\+1\}\)\-H\_\{k\}\(z\_\{k\}\)\-\\log J\_\{k,\\theta\}\\right\]\.\(85\)This is a generic deterministic\-map convention\. In the NHMC runs reported here, the deterministic kick–drift maps are volume preserving, soJk,θ=1J\_\{k,\\theta\}=1and the Jacobian term vanishes; the stochastic momentum refreshments contribute instead through the forward\-minus\-reverse log\-density terms in Eq\. \([10](https://arxiv.org/html/2607.15682#S2.E10)\)\.

### A\.15Jarzynski Estimator

If the path convention gives

𝔼Qθ​\[e−Wθ\]=Z=e−Δ​F,\\mathbb\{E\}\_\{Q\_\{\\theta\}\}\[e^\{\-W\_\{\\theta\}\}\]=Z=e^\{\-\\Delta F\},\(86\)then

Δ​F^=−log⁡\(1N​∑i=1Ne−Wi\)\\widehat\{\\Delta F\}=\-\\log\\left\(\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}e^\{\-W\_\{i\}\}\\right\)\(87\)is consistent forΔ​F\\Delta Funder standard moment conditions\. The estimator ofe−Δ​Fe^\{\-\\Delta F\}is unbiased; the log\-transformed free\-energy estimator is generally biased at finiteNNbut consistent\.

### A\.16Train Then Evaluate with Fixed Parameters

The invariance and consistency results above are for a fixed parameterθ\\theta\. After training produces a fixed parameterθ⋆\\theta^\{\\star\}, the resulting endpoint\-IMH, path\-IMH, or round\-trip NHMC\-MH kernel preserves the corresponding target law by applying the same argument atθ⋆\\theta^\{\\star\}\. Online adaptation would require additional adaptive MCMC conditions\.

### A\.17Symmetry\-Structured Proposal Densities

LetGGbe a finite group of differentiable, volume\-preserving transformations on configuration space\. Given any proposal densityqθq\_\{\\theta\}, define the group\-averaged density

qθG​\(x\)=1\|G\|​∑g∈Gqθ​\(g−1​x\)\.q^\{G\}\_\{\\theta\}\(x\)=\\frac\{1\}\{\|G\|\}\\sum\_\{g\\in G\}q\_\{\\theta\}\(g^\{\-1\}x\)\.\(88\)This is a normalized density because eachggpreserves volume\. If the target is invariant,γ​\(g​x\)=γ​\(x\)\\gamma\(gx\)=\\gamma\(x\), thenqθGq^\{G\}\_\{\\theta\}is a natural symmetry\-respecting proposal\. The endpoint weight becomes

wθG​\(x\)=γ​\(x\)qθG​\(x\)\.w^\{G\}\_\{\\theta\}\(x\)=\\frac\{\\gamma\(x\)\}\{q^\{G\}\_\{\\theta\}\(x\)\}\.\(89\)Providedπ\\piis absolutely continuous with respect toqθGq^\{G\}\_\{\\theta\}, the detailed\-balance proof for endpoint IMH applies withqθGq^\{G\}\_\{\\theta\}andwθGw^\{G\}\_\{\\theta\}in place ofqθq\_\{\\theta\}andwθw\_\{\\theta\}\. The SNIS consistency proof likewise applies after replacing the sampling law byqθGq^\{G\}\_\{\\theta\}\. Thus symmetry changes finite\-sample efficiency and weight variance, while the correction argument remains unchanged\.

If the physical target breaks the candidate symmetry, a symmetrized proposal remains valid when its density is computed correctly and has sufficient support, although its statistical efficiency can be poor\. For this reason, the tilted public DW target is evaluated with its asymmetric target masses instead of enforced equal sign\-mode probabilities\.

## Appendix BNumerical Setup and Implementation Details

### B\.1Target Definitions

All numerical studies use unnormalized target densitiesγ​\(x\)=exp⁡\[−E​\(x\)\]\\gamma\(x\)=\\exp\[\-E\(x\)\]\. For the double\-well family, the target is

log⁡γd​\(x\)=∑j=0d/2−1\[−x2​j4\+6​x2​j2\+0\.5​x2​j−12​x2​j\+12\]\.\\log\\gamma\_\{d\}\(x\)=\\sum\_\{j=0\}^\{d/2\-1\}\\left\[\-x\_\{2j\}^\{4\}\+6x\_\{2j\}^\{2\}\+0\.5x\_\{2j\}\-\\frac\{1\}\{2\}x\_\{2j\+1\}^\{2\}\\right\]\.\(90\)The Gaussian coordinate in each pair factorizes, while the even coordinate is a tilted double well\. The base law isq0=𝒩​\(0,I\)q\_\{0\}=\\mathcal\{N\}\(0,I\)for the neural NHMC path runs\. Analytic normalizers and exact target samples are obtained from the one\-dimensional factorization; sign modes are defined by the signs of the double\-well coordinates\.

For the two\-dimensional scalarϕ4\\phi^\{4\}benchmark we use the finite\-volume lattice action convention

S​\[ϕ\]=∑x\[−2​κ​∑μ=12ϕx​ϕx\+μ^\+\(1−2​λ\)​ϕx2\+λ​ϕx4\],S\[\\phi\]=\\sum\_\{x\}\\left\[\-2\\kappa\\sum\_\{\\mu=1\}^\{2\}\\phi\_\{x\}\\phi\_\{x\+\\hat\{\\mu\}\}\+\(1\-2\\lambda\)\\phi\_\{x\}^\{2\}\+\\lambda\\phi\_\{x\}^\{4\}\\right\],\(91\)with periodic boundary conditions,λ=0\.022\\lambda=0\.022, and theκ\\kappavalues listed in the main and appendix tables\. The magnetization isM=V−1​∑xϕxM=V^\{\-1\}\\sum\_\{x\}\\phi\_\{x\}and the ordering observable shown is\|M\|\|M\|\. We use

χ=V​\(⟨M2⟩−⟨\|M\|⟩2\),UB=1−⟨M4⟩3​⟨M2⟩2\.\\chi=V\\left\(\\langle M^\{2\}\\rangle\-\\langle\|M\|\\rangle^\{2\}\\right\),\\qquad U\_\{B\}=1\-\\frac\{\\langle M^\{4\}\\rangle\}\{3\\langle M^\{2\}\\rangle^\{2\}\}\.\(92\)The same definitions are applied to the raw proposal, weighted NHMC, and HMC reference samples\.

For compactU​\(1\)U\(1\)we parameterize links by anglesUx,μ=exp⁡\(i​θx,μ\)U\_\{x,\\mu\}=\\exp\(i\\theta\_\{x,\\mu\}\)on a periodic two\-dimensional lattice\. The Wilson action used by the local\-chain baselines is

S​\[U\]=β​∑p\(1−cos⁡θp\),S\[U\]=\\beta\\sum\_\{p\}\(1\-\\cos\\theta\_\{p\}\),\(93\)equivalent to theexp⁡\[β​cos⁡θp\]\\exp\[\\beta\\cos\\theta\_\{p\}\]convention up to the constantβ​V\\beta V\. Plaquette observables usecos⁡θp\\cos\\theta\_\{p\}\. The geometric topological charge is computed from wrapped plaquette angles,

Q​\[U\]=12​π​∑pArg​\{exp⁡\(i​θp\)\},Q\[U\]=\\frac\{1\}\{2\\pi\}\\sum\_\{p\}\\mathrm\{Arg\}\\\{\\exp\(i\\theta\_\{p\}\)\\\},\(94\)and rounded sectors are used only for finite\-chain trajectory summaries\. This wrapped\-plaquette convention follows the geometric lattice topological\-charge construction\(Lüscher,[1982](https://arxiv.org/html/2607.15682#bib.bib39)\)\. Finite\-volume plaquette and square Wilson\-loop references are evaluated from the character expansionZV∝∑n∈ℤIn​\(β\)VZ\_\{V\}\\propto\\sum\_\{n\\in\\mathbb\{Z\}\}I\_\{n\}\(\\beta\)^\{V\}, with an area\-AAloop replacingInVI\_\{n\}^\{V\}byInV−A​In\+1AI\_\{n\}^\{V\-A\}I\_\{n\+1\}^\{A\}in the numerator\.

For the molecular targets, OpenMM 8\.4 is used as the energy and force engine\(Eastmanet al\.,[2017](https://arxiv.org/html/2607.15682#bib.bib36)\)\. The Ala4 and Ala6 systems are bgmol MiniPeptide systems associated with the Boltzmann\-generator benchmark suite\(Noéet al\.,[2019](https://arxiv.org/html/2607.15682#bib.bib5)\)\. The OpenMM systems use the AMBER99SB\-ILDN protein force field\(Lindorff\-Larsenet al\.,[2010](https://arxiv.org/html/2607.15682#bib.bib37)\)with OBC/GBSA implicit solvent\(Onufrievet al\.,[2004](https://arxiv.org/html/2607.15682#bib.bib38)\)\. Energies are evaluated in units ofkB​Tk\_\{B\}TatT=300​KT=300\\,\\mathrm\{K\}\. The proposal state is a vector of bond, angle, and torsion internal coordinateszz, mapped to Cartesian coordinates byT​\(z\)T\(z\)\. The dimensionless energy used by the implementation is

EIC​\(z\)=Ecart​\(T​\(z\)\)−log⁡\|detJT​\(z\)\|\.E\_\{\\mathrm\{IC\}\}\(z\)=E\_\{\\mathrm\{cart\}\}\(T\(z\)\)\-\\log\\left\|\\det J\_\{T\}\(z\)\\right\|\.\(95\)OpenMM supplies Cartesian endpoint energies and first\-order forces\. The force evaluations are treated as fixed first\-order target information during training; the implementation does not differentiate through OpenMM or evaluate Hessian–vector products\. Proposal kicks are generated by learned position\-dependent fields\. Initial states are sampled uniformly from a stored molecular dynamics prior pool\. For an evaluated path, the implementation records the empirical prior\-action log score

log⁡sprior=−EIC​\(zL\)\+EIC​\(z0\)−log⁡Ppath,\\log s\_\{\\mathrm\{prior\}\}=\-E\_\{\\mathrm\{IC\}\}\(z\_\{L\}\)\+E\_\{\\mathrm\{IC\}\}\(z\_\{0\}\)\-\\log P\_\{\\mathrm\{path\}\},\(96\)wherelog⁡Ppath\\log P\_\{\\mathrm\{path\}\}is the accumulated forward\-minus\-reverse path log\-probability\. Bond and angle coordinates are projected back into their numerical domains after proposal updates, while torsions are wrapped periodically\. These projections are not invertible, so the molecular results are reported as a learned\-force prior\-action feasibility study rather than as an application of Theorem[1](https://arxiv.org/html/2607.15682#Thmtheorem1)\.

### B\.2NHMC Proposal Families and Training

For each learned proposal we fix the trained parameters before evaluation\. The neural DW runs use learned forward and reverse conditional momentum laws around target\-force leapfrog updates; the implementation clips only extremely large force values for numerical stability\. The lattice runs use learned position\-dependent kicks in the same reversible, volume\-preserving composition\. The DW runs also apply a learned Bernoulli sign flip to the double\-well coordinates at each stage, and include its forward and reverse log probabilities in the recorded work\. Theϕ4\\phi^\{4\}implementation similarly records its globalZ2Z\_\{2\}move\. These Bernoulli outcomes are sampled as hard discrete variables without a continuous relaxation\. During backpropagation the realized outcome is held fixed, while gradients reach the stage logits through its recorded forward and reverse Bernoulli log\-probability terms; no Gumbel or straight\-through estimator is used\. The DW and lattice objectives combine mean recorded work with a recorded\-work variance regularizer\. Its coefficient is0\.050\.05for the reported DW runs and0\.010\.01for theϕ4\\phi^\{4\}and compact\-U​\(1\)U\(1\)runs\. Forward and reverse momentum networks are parameterized separately and use SiLU activations\. DW checkpoints are selected by normalized ESS on fixed validation paths; theϕ4\\phi^\{4\}checkpoint is selected by mean normalized ESS across the validationκ\\kappagrid\. The mass matrix is the identity and the learned stage step sizes are state independent\. The8×88\\times 8ϕ4\\phi^\{4\}scan uses oneκ\\kappa\-conditional proposal trained overκ∈\[0\.20,0\.30\]\\kappa\\in\[0\.20,0\.30\], rather than a separate model at each grid point\. Stage time andκ\\kappaenter the momentum and learned\-kick networks through a joint conditioning embedding\. The selected DW4 main result and the DW4 round\-trip comparison use different trained proposals\. The latter uses three 32\-stage, five\-leapfrog\-per\-stage checkpoints, as listed separately below\.

Table 4:NHMC training and evaluation budgets\.
### B\.3Statistical Quantities

For weighted SNIS estimates with unnormalized weightswiw\_\{i\}, the normalized effective sample size is

ESS/N=\(∑iwi\)2N​∑iwi2\.\\mathrm\{ESS\}/N=\\frac\{\(\\sum\_\{i\}w\_\{i\}\)^\{2\}\}\{N\\sum\_\{i\}w\_\{i\}^\{2\}\}\.\(97\)This quantity is invariant to multiplying all weights by a positive constant and is separate from autocorrelation\-based Markov\-chain effective counts\. The log\-weight spread is the sample standard deviation of the log weights used for the weighted estimator\. Free\-energy or log\-normalizer error is computed fromlog⁡Z^=log⁡\(N−1​∑iwi\)\\log\\widehat\{Z\}=\\log\(N^\{\-1\}\\sum\_\{i\}w\_\{i\}\)against the exact or named reference for that result\. IMH chains are summarized by acceptance and integrated autocorrelation time\. For multimodal targets we record observed mode coverage and mode\-mass error\. For theϕ4\\phi^\{4\}single\-chain comparison, standard errors of linear trace averages use the autocorrelation\-based effective count\. Susceptibility and Binder\-cumulant uncertainties use contiguous\-block delete\-one jackknife estimates; the high\-statistics HMC reference is analyzed with the same observable definitions\.

For any weighted histogram, mode\-mass, or action\-overlap panel, the plotted finite\-sample object is the weighted empirical measureμ^Nw=∑iw¯i​δxi\\widehat\{\\mu\}\_\{N\}^\{w\}=\\sum\_\{i\}\\bar\{w\}\_\{i\}\\delta\_\{x\_\{i\}\}withw¯i=wi/∑jwj\\bar\{w\}\_\{i\}=w\_\{i\}/\\sum\_\{j\}w\_\{j\}\. If a panel displays an SNIS\-resampled visual set, the resampling is from this empirical measure\. Weighted histograms and resampled panels are therefore visualizations of the same finite\-sample weighted measure\.

For the rotated DW8 action\-overlap comparison, the overlap is the common histogram mass between the reference action distribution and the weighted NHMC action distribution\. With normalized histogram massesp^b\\widehat\{p\}\_\{b\}andq^b\\widehat\{q\}\_\{b\}over common bins, this is

overlap​\(p,q\)=∑bmin⁡\{p^b,q^b\}\.\\mathrm\{overlap\}\(p,q\)=\\sum\_\{b\}\\min\\\{\\widehat\{p\}\_\{b\},\\widehat\{q\}\_\{b\}\\\}\.\(98\)
The molecular energy overlap uses the same common\-mass definition with 100 common bins spanning the joint energy range of the weighted NHMC and molecular dynamics reference ensembles\. Torsion summaries use periodic\(φ,ψ\)∈\[−π,π\)2\(\\varphi,\\psi\)\\in\[\-\\pi,\\pi\)^\{2\}histograms\. The displayed free\-energy surfaces use72×7272\\times 72bins, whereas the reported overlap uses36×3636\\times 36bins and is

𝒪FES=∑bmin⁡\{p^b,q^b\}=1−TV​\(p^,q^\)\.\\mathcal\{O\}\_\{\\mathrm\{FES\}\}=\\sum\_\{b\}\\min\\\{\\widehat\{p\}\_\{b\},\\widehat\{q\}\_\{b\}\\\}=1\-\\mathrm\{TV\}\(\\widehat\{p\},\\widehat\{q\}\)\.\(99\)For Ala4 and Ala6, the FES overlap is the mean of this common mass over the available residue\-level torsion pairs\. The reference files contain 50000 saved molecular\-dynamics configurations; this count is not interpreted as 50000 independent effective samples\.

For the compactU​\(1\)U\(1\)round\-trip pilot, uncertainties are standard errors across the 16 chains\. Topological sectors are assigned by rounding the measured charge to the nearest integer after burn\-in, and only the aggregate number of adjacent sector changes is reported\. These traces are not used to estimate topological susceptibility or autocorrelation\. The maximum inverse\-map, momentum\-reversal, and work\-ratio residuals are evaluated directly from the recorded paths\.

### B\.4Latticeϕ4\\phi^\{4\}Settings

Table 5:ϕ4\\phi^\{4\}settings for the main finite\-volume observable comparison\.

## Appendix CSupplementary Numerical Results

### C\.1Neural NHMC on DW / Many\-Well

The learned\-momentum Hamiltonian DW run uses the many\-well density

log⁡γd​\(x\)=∑j=0d/2−1\[−x2​j4\+6​x2​j2\+0\.5​x2​j−12​x2​j\+12\],\\log\\gamma\_\{d\}\(x\)=\\sum\_\{j=0\}^\{d/2\-1\}\\left\[\-x\_\{2j\}^\{4\}\+6x\_\{2j\}^\{2\}\+0\.5x\_\{2j\}\-\\frac\{1\}\{2\}x\_\{2j\+1\}^\{2\}\\right\],\(100\)with analyticlog⁡Z\\log Zand exact target sampling available by the one\-dimensional factorization\. The neural NHMC proposal starts fromq0=𝒩​\(0,I\)q\_\{0\}=\\mathcal\{N\}\(0,I\), samples learned momenta, and applies leapfrog updates with the target force\. With forward and reverse learned momentum lawsqθ,tFq^\{F\}\_\{\\theta,t\}andqθ,tRq^\{R\}\_\{\\theta,t\}, and withp¯t\+1\\bar\{p\}\_\{t\+1\}denoting the output momentum of the Hamiltonian map at stagett, the recorded work is

Wθ​\(Γ\)=\\displaystyle W\_\{\\theta\}\(\\Gamma\)=\{\}log⁡q0​\(x0\)−log⁡γ​\(xL\)\\displaystyle\\log q\_\{0\}\(x\_\{0\}\)\-\\log\\gamma\(x\_\{L\}\)\(101\)\+∑t=0L−1\[logrθ,tF\(st∣xt\)−logrθ,tR\(st†∣xt\+1\)\\displaystyle\+\\sum\_\{t=0\}^\{L\-1\}\\bigl\[\\log r^\{F\}\_\{\\theta,t\}\(s\_\{t\}\\mid x\_\{t\}\)\-\\log r^\{R\}\_\{\\theta,t\}\(s\_\{t\}^\{\\dagger\}\\mid x\_\{t\+1\}\)\+logqθ,tF\(pt∣x~t\)−logqθ,tR\(−p¯t\+1∣xt\+1\)\]\.\\displaystyle\\hskip 50\.00008pt\+\\log q^\{F\}\_\{\\theta,t\}\(p\_\{t\}\\mid\\widetilde\{x\}\_\{t\}\)\-\\log q^\{R\}\_\{\\theta,t\}\(\-\\bar\{p\}\_\{t\+1\}\\mid x\_\{t\+1\}\)\\bigr\]\.and the correction weight is proportional toexp⁡\[−Wθ​\(Γ\)\]\\exp\[\-W\_\{\\theta\}\(\\Gamma\)\]\. A learned stochastic sign flip over double\-well coordinates is included in the recorded path\. The move is an involution, but its stage\-indexed forward and reverse log probabilities are retained explicitly in the recorded work above\.

Table 6:Neural NHMC learned\-momentum leapfrog with path\-work correction on the many\-well DW4/DW8 targets\.Separate DW4 and DW8 checkpoint ensembles support the configuration\-space round\-trip kernel of Proposition[1](https://arxiv.org/html/2607.15682#Thmproposition1)\. Table[7](https://arxiv.org/html/2607.15682#A3.T7)uses three independently trained checkpoints per target and fixes the total number of evaluated NHMC paths within each comparison\. Round\-trip NHMC\-MH improves acceptance and autocorrelation per transition on both targets\. After accounting for its two paths per transition, it is less force\-efficient on DW4 and approximately matched to independent path\-IMH on DW8\.

Table 7:DW4 and DW8 independent path\-IMH and round\-trip NHMC\-MH across three trained checkpoints per target\. These checkpoint ensembles are separate from the selected main evaluations in Table[1](https://arxiv.org/html/2607.15682#S4.T1); the DW4 ensemble uses 32 proposal stages and five leapfrog steps per stage\. Values are means±\\pmsample standard deviations\. Within each target, both kernels use the same total number of evaluated NHMC paths; a round\-trip transition uses two paths\. Chains start from endpoints generated by a Gaussian base draw and one forward NHMC path\. For every checkpoint, 32 chains retain 2000 path\-IMH transitions after discarding 400, or 1000 round\-trip transitions after discarding 200; both protocols therefore evaluate the same total number of NHMC paths\. Mode masses use all retained endpoints\. The reportedτint​\(x0\)\\tau\_\{\\rm int\}\(x\_\{0\}\)is the median across chains, using an FFT autocorrelation estimate with a self\-consistent window and a maximum lag of one quarter of the retained chain length\. The force\-cost column multiplies this value by the number of target\-force evaluations per transition\.DW2 makes the path dynamics explicit: the best training ESS is about43%43\\%and both sign modes are covered\. DW4 is the main neural Hamiltonian\-path DW result\. DW8 still covers all1616sign modes and has a small observed log\-normalizer error in the selected evaluation; its low ESS makes it a scaling and hyperparameter\-sensitivity probe\.

![Refer to caption](https://arxiv.org/html/2607.15682v1/x4.png)Figure 5:DW2 neural NHMC time trace\. The left panel follows real forward trajectories from the Gaussian base through the learned Hamiltonian path, with proposal stagetton the horizontal axis and the double\-well coordinatex0x\_\{0\}on the vertical axis\. The right panel applies the learned reverse path to exact target samples and traces the pullback toward the base\. For this fixed proposal, path\-SNIS ESS is43\.4%43\.4\\%and the reverse pullback standard deviation is1\.161\.16\.![Refer to caption](https://arxiv.org/html/2607.15682v1/figures/dw2_weighted_correction.png)Figure 6:DW2 correction on the asymmetric two\-well target\. SNIS and IMH correction recover the intentionally asymmetric target masses; the right mode is heavier because the target factor includes the linear tilt0\.5​x0\.5x\.
### C\.2Rotated DW8 Endpoint\-Density Ablation

The native many\-well target factorizes across coordinate pairs, so a factorized proposal can be unusually favorable\. To remove that coordinate advantage, we rotate the DW8 target by a fixed orthogonal mapRR,

γR​\(x\)=γ​\(x​R\),\\gamma\_\{R\}\(x\)=\\gamma\(xR\),\(102\)which preserves the analytic log normalizer and target sampler but destroys coordinate\-wise factorization in the observed coordinates\. We then train a global affine proposal map on top of the many\-well base proposal with a rotation curriculum\. This endpoint\-density ablation asks whether proposal training restores concentrated correction weights on a coupled version of the same Boltzmann target\.

![Refer to caption](https://arxiv.org/html/2607.15682v1/x5.png)Figure 7:Rotated DW8 pairwise marginal coverage for the endpoint\-density ablation\. Contours are empirical pairwise target marginals from the analytic target sampler\. The untrained factorized proposal is misaligned in the rotated coordinates; the trained proposal and its SNIS\-resampled visual set track the target marginals under the same correction weights\.![Refer to caption](https://arxiv.org/html/2607.15682v1/x6.png)Figure 8:Rotated DW8 endpoint\-density training trajectory\. The curriculum reaches the final rotated target at step 1500; after training the batch ESS stabilizes near one and the log\-weight spread collapses\.
### C\.3Action Overlap for Rotated DW8

The rotated\-DW8 action\-overlap check uses one fixedN=50000N=50000evaluation\. The factorized proposal provides the pre\-training baseline and the trained affine proposal the learned map\. The target action is−log⁡γR​\(x\)\-\\log\\gamma\_\{R\}\(x\)\. Training changes the action overlap from1\.46%1\.46\\%to99\.56%99\.56\\%, collapses the log\-weight spread from19\.9719\.97to0\.1570\.157, and gives96\.67%96\.67\\%ESS and91\.58%91\.58\\%stationary IMH acceptance\. The SNIS\-resampled visual set has99\.38%99\.38\\%action overlap and small mode\-mass error\.

Table 8:Unified rotated\-DW8 endpoint\-density evaluation withN=50,000N=50\{,\}000\. The untrainedq0q\_\{0\}row is a before\-training control\. Action overlap is the common histogram mass between proposal and exact target actions; the SNIS\-resampled row visualizes the weighted empirical measure\.![Refer to caption](https://arxiv.org/html/2607.15682v1/figures/rotated_dw_overlap.png)Figure 9:Action overlap for the rotated DW8 endpoint\-density ablation\. The untrained proposal has poor target\-action overlap and distorted sign\-mode masses; after training, the proposal action distribution overlaps the reference target action distribution, the log weights concentrate, and the SNIS\-resampled estimates track the target action and mode masses\.
### C\.4Molecular Prior\-Action Score Summaries

The molecular rows in Table[2](https://arxiv.org/html/2607.15682#S4.T2)use learned\-force proposal evaluations and the prior\-action scores defined in Appendix[B](https://arxiv.org/html/2607.15682#A2)\. All5000050000scores are finite for both Ala4 and Ala6\. Table[2](https://arxiv.org/html/2607.15682#S4.T2)gives score ESS, log\-score spread, and common mass between the weighted energy histogram and the molecular\-dynamics reference histogram\. Score ESS measures concentration in the full recorded prior\-action score, whereas FES overlap measures only low\-dimensional torsion marginals; the two summaries need not vary monotonically\.

For Ala4 and Ala6, torsion records are available for the sameN=50000N=50000NHMC evaluations used in Table[2](https://arxiv.org/html/2607.15682#S4.T2)\. The panels use the corresponding normalized prior\-action scores and compare weighted two\-dimensional torsion histograms againstN=50000N=50000molecular\-dynamics reference ensembles\.

![Refer to caption](https://arxiv.org/html/2607.15682v1/x7.png)Figure 10:Ala6 torsion free\-energy surfaces from the prior\-action\-score\-weighted NHMC ensemble and the molecular\-dynamics reference\. The prior\-action score ESS is11\.00%11\.00\\%, the score\-weighted energy overlap is93\.30%93\.30\\%, and the aggregate FES overlap is0\.8890\.889\. Each ensemble usesN=50000N=50000configurations\.![Refer to caption](https://arxiv.org/html/2607.15682v1/x8.png)Figure 11:Ala4 torsion free\-energy surfaces from the prior\-action\-score\-weighted NHMC ensemble and the molecular\-dynamics reference\. The prior\-action score ESS is39\.53%39\.53\\%, the score\-weighted energy overlap is92\.76%92\.76\\%, and the aggregate FES overlap is0\.8050\.805\. Each ensemble usesN=50000N=50000configurations\.
### C\.5Finite\-Volumeϕ4\\phi^\{4\}Details

Table[9](https://arxiv.org/html/2607.15682#A3.T9)summarizes the two\-dimensional8×88\\times 8κ\\kappascan and the smaller4×44\\times 4correction regime\.

Table 9:Finite\-volumeϕ4\\phi^\{4\}NHMC\-SNIS estimates\.The zero\-field scalar action is invariant under the global sign flipϕx↦−ϕx\\phi\_\{x\}\\mapsto\-\\phi\_\{x\}\. In finite volume this means the signed magnetizationM=V−1​∑xϕxM=V^\{\-1\}\\sum\_\{x\}\\phi\_\{x\}cancels under sector\-balanced target sampling, giving⟨M⟩=0\\langle M\\rangle=0\. The main text therefore uses\|M\|\|M\|,M2M^\{2\}, susceptibility, and Binder\-cumulant quantities\. Table[10](https://arxiv.org/html/2607.15682#A3.T10)gives sign\-sector balance for the8×88\\times 8scan\. It pairs the proposal sign fraction with the SNIS\-corrected observable estimates and ESS, quantifying finite\-budget sector coverage\.

Table 10:Z2Z\_\{2\}sign\-sector summary for representative points in theϕ4\\phi^\{4\}8×88\\times 8scan\. The columnp^\+\\widehat\{p\}\_\{\+\}is the SNIS\-weighted positive\-sign fraction; corrected observables and ESS use the same SNIS weights\.The8×88\\times 8\|M\|\|M\|autocorrelation comparison uses one HMC chain, one independent NHMC path\-IMH chain, and one round\-trip NHMC\-MH chain atκ=0\.2705\\kappa=0\.2705\. The HMC and path\-IMH chains each have 5000 saved states and discard the first 200\. The round\-trip chain is initialized by drawing from the Gaussian base and running one forward NHMC path; it then runs for 10000 transitions and discards the first 1000\. The comparison uses the first 4800 post\-burn values from each chain with the same dot\-normalized autocorrelation estimator\. The integrated autocorrelation times per saved transition are19\.9519\.95,8\.948\.94, and2\.572\.57, respectively\. Because a round\-trip transition uses two paths, its path\-cost\-normalized value is5\.155\.15\. Using all 9000 post\-burn round\-trip states gives⟨\|M\|⟩=0\.9066±0\.0100\\langle\|M\|\\rangle=0\.9066\\pm 0\.0100,χ=9\.98±0\.32\\chi=9\.98\\pm 0\.32, andUB=0\.5099±0\.0060U\_\{B\}=0\.5099\\pm 0\.0060\. The corresponding 198000\-state HMC reference values, computed with the same observable definitions, are0\.9068±0\.00680\.9068\\pm 0\.0068,9\.71±0\.129\.71\\pm 0\.12, and0\.5097±0\.00360\.5097\\pm 0\.0036\. This 198000\-state reference is a separate long single\-chain HMC run from the 60\-chain grid used in Fig\.[3](https://arxiv.org/html/2607.15682#S4.F3)\. Table[11](https://arxiv.org/html/2607.15682#A3.T11)gives sector and autocorrelation summaries for all three matched traces\.

Table 11:ϕ4\\phi^\{4\}8×88\\times 8matched single\-chain summaries atκ=0\.2705\\kappa=0\.2705\. Each row uses 4800 post\-burn states\. Acceptance is measured per method\-specific transition and is not a force\-cost\-matched comparison\.![Refer to caption](https://arxiv.org/html/2607.15682v1/x9.png)Figure 12:ϕ4\\phi^\{4\}8×88\\times 8single\-chain comparison for\|M\|\|M\|atκ=0\.2705\\kappa=0\.2705\. HMC, independent path\-IMH, and round\-trip NHMC\-MH each use 4800 post\-burn states\. The round\-trip chain begins from an NHMC forward endpoint; the HMC chain is used only as an independent reference\. The left panel compares autocorrelation per saved transition; the right panel shows running means against the high\-statistics HMC reference\. A round\-trip transition evaluates two NHMC paths\.
### C\.6CompactU​\(1\)U\(1\)Round\-Trip Pilot

We applied the shared\-bridge round\-trip construction to the trained compactU​\(1\)U\(1\)proposal on an8×88\\times 8lattice\. Each chain is initialized by an independent draw from the training base law followed by one forward NHMC path; no HMC or target\-state warm start is used\. The inverse\-map and momentum\-reversal residuals are below4×10−154\\times 10^\{\-15\}, and the direct augmented\-density ratio agrees with the work\-difference expression to2\.3×10−132\.3\\times 10^\{\-13\}\. These checks verify the implemented transition and acceptance ratio independently of its sampling efficiency\.

Table 12:CompactU​\(1\)U\(1\)round\-trip NHMC\-MH pilot on an8×88\\times 8lattice\. Each chain starts from an independent draw from the training base law followed by one forward NHMC path\. Theβ=4\\beta=4row uses 2000 transitions per chain and discards 400; theβ=8\\beta=8row uses 5000 transitions per chain and discards 1000\. Uncertainties are standard errors across chains\. The final column counts rounded\-sector changes over all post\-discard traces\.At both couplings, the plaquette,W22W\_\{22\}, andW33W\_\{33\}estimates are within one standard error of their finite\-volume character\-expansion values: respectively\(0\.86353,0\.55610,0\.26724\)\(0\.86353,0\.55610,0\.26724\)atβ=4\\beta=4and\(0\.93548,0\.76809,0\.55942\)\(0\.93548,0\.76809,0\.55942\)atβ=8\\beta=8\. Acceptance is nevertheless only1\.05%1\.05\\%and0\.43%0\.43\\%\. The post\-discard traces contain 172 and 83 adjacent rounded\-sector changes; atβ=8\\beta=8, 14 of 16 chains change sector at least once\. The low acceptance and limited number of accepted burn\-in moves prevent interpreting these runs as evidence of efficient equilibration or topology sampling\.

## Appendix DReproducibility Statement

#### Data and code\.

Code, evaluation scripts, and numerical outputs will be released through[https://github\.com/qxxmax/lattice\-ml](https://github.com/qxxmax/lattice-ml), subject to redistribution constraints\. The release will include target definitions, budgets, random seeds, correction modes, and the values used in the tables and figures\. Proposal parameters are fixed before evaluation; observable definitions and uncertainty estimators are given in Appendix[B](https://arxiv.org/html/2607.15682#A2)\. Restricted external data will be identified by source, with only the derived quantities used here redistributed\.

#### Use of AI tools\.

OpenAI ChatGPT and Codex assisted with language editing, literature and citation checks, LaTeX, and analysis scripts\. The author verified the mathematical arguments, calculations, numerical results, and citations and takes responsibility for the manuscript\.

Similar Articles

Hamiltonian Neural Networks from a Differential Geometry Perspective

Reddit r/artificial

A blog post explaining Hamiltonian Neural Networks through differential geometry, using a simple mass-spring system to demonstrate how imposing conservation laws via network architecture can lead to more efficient learning. The author builds up mathematical tools like symplectic manifolds and Poisson brackets from basic calculus.

Show HN: Neural Particle Automata

Hacker News Top

Introduces Neural Particle Automata, a method for learning self-organizing particle dynamics using smooth particle hydrodynamics perception, enabling particles to have local perception vectors for an update rule, analogous to Neural Cellular Automata but on continuous particle positions.