Hybrid Probabilistic Zonotopes for Identifiable and Refinable Predictive Uncertainty
摘要
This paper introduces Hybrid Probabilistic Zonotopes (HProbZ), a neural network output head that jointly models discrete modes, bounded drift, and stochastic noise with closed-form likelihood, enabling identifiable uncertainty decomposition, observation-driven contraction, and multi-modal conformal coverage.
查看缓存全文
缓存时间: 2026/08/07 07:50
# Hybrid Probabilistic Zonotopes for Identifiable and Refinable Predictive Uncertainty
Source: [https://arxiv.org/html/2608.05454](https://arxiv.org/html/2608.05454)
Zhen Zhang Amr Alanwar School of Computation, Information and Technology Technical University of Munich, Germany \{zhenzhang\.zhang, alanwar\}@tum\.de
###### Abstract
Probabilistic prediction heads in neural networks typically output either a Gaussian mixture or a single conformal region\. Neither separates the distinct sources of uncertainty often present in real prediction tasks: a discrete choice among modes, bounded systematic drift within the chosen mode, and irreducible stochastic noise\. We introduce the Hybrid Probabilistic Zonotope \(HProbZ\), an output head that represents these three sources as binary, bounded, and stochastic generators of a zonotope, and admits a closed\-form likelihood by convolution\. Sharing the bounded generator across prediction steps couples future predictions algebraically, so observing one step refines the predictive distribution at every remaining step in a single forward pass\. We establish that the three generators are identifiable from the likelihood up to permutation, and that an HProbZ density is representationally distinct from any finite Gaussian mixture\. The same shared structure provides analytic per\-mode risk and distribution\-free multi\-modal conformal sets at inference time\. Empirical analysis on representative prediction benchmarks supports the effectiveness of the design relative to same\-encoder mixture baselines, while offering structural properties that mixture or convex\-conformal predictors do not jointly provide\.
## 1Introduction
Many predictive tasks involve three distinct uncertainty sources: discrete modal choice, bounded systematic drift, and irreducible stochastic noise\. Trajectory prediction is a canonical instance, arising whenever a model must commit to a discrete behaviour*and*report within\-mode uncertainty that sharpens as observations arrive\. Existing heads address one source or the other: mixture density networks commit to sharp modes but their within\-mode covariances are network outputs that observations can only re\-weight; conformal methods give distribution\-free coverage but as a single convex set that inflates on multi\-modal targets\. We introduce a predictive distribution offering four properties simultaneously — identifiable modal/bounded/stochastic decomposition, observation\-driven contraction, analytic per\-mode risk, and multi\-modal conformal coverage — that prior approaches provide individually but not jointly \(Tab\.[1](https://arxiv.org/html/2608.05454#S1.T1)\)\.


Figure 1:Overview\.Left:ETH/UCY LOO: revealingk=8k\{=\}8observations contracts HProbZ’s predictive std by23\.6%23\.6\\%through joint posterior updates over shared\(β,α\)\(\\beta,\\alpha\), vs\.3\.6%3\.6\\%for a same\-capacity Gaussian mixture \(weight reweighting only\); mean over 3 seeds\.Right:schematic atk=0k\{=\}0vsk=6k\{=\}6: each HProbZ hexagon isc\+Gbβ\+Gdαc\{\+\}G\_\{b\}\\beta\{\+\}G\_\{d\}\\alphainflated by the2σ2\\sigmahalo ofGsνG\_\{s\}\\nu\. Downstream \(nuScenes,N=39,829N\{=\}39\{,\}829,66s, static obstacle\): at near\-matched10%10\\%FAR HProbZ produces𝟑𝟎%\\mathbf\{30\\%\}fewer trajectory\-level collisions than same\-encoder MDN\-K2 \(Tab\.[5](https://arxiv.org/html/2608.05454#S5.T5)\); per\-step reverses\.Hybrid Probabilistic Zonotopes \(HProbZ\) add three generators to a centre point as a single neural output: a binary generator for discrete modes, a bounded generator \(shared across all future timesteps\) for continuous drift, and a Gaussian generator for noise\. Because the bounded generator is shared, revealing one step pins it down and sharpens every other step in one forward pass\. Fig\.[1](https://arxiv.org/html/2608.05454#S1.F1)contrasts this with a same\-capacity Gaussian mixture: HProbZ’s predictive support shrinks under observation, while the mixture can only re\-weight components and leaves within\-mode covariance frozen\. A shared\-latent CVAE contracts but collapses modal, drift, and noise into one Gaussian, erasing the separation downstream planners and conformal\-set users need; HProbZ keeps all three sources identifiable from one closed\-form likelihood\.
Contributions\.\(i\)We introduce the*Hybrid Probabilistic Zonotope*\(HProbZ\): a structured\-output head whose binary, bounded, and stochastic generators are trained jointly under a single closed\-form convolution likelihood \(§[3](https://arxiv.org/html/2608.05454#S3)–§[4](https://arxiv.org/html/2608.05454#S4)\)\.\(ii\)We establish three properties of this representation: HProbZ generators are uniquely identifiable from the exact likelihood; HProbZ densities lie strictly outside the finite Gaussian\-mixture family; and the entropy cost of compact support is bounded by a universal constant per dimension\.\(iii\)We show that sharing the bounded generator across prediction steps yields single\-pass observation\-driven uncertainty contraction, analytic per\-mode risk, and distribution\-free multi\-modal conformal sets — a combination prior heads provide individually but not jointly \(Tab\.[1](https://arxiv.org/html/2608.05454#S1.T1)\)\.\(iv\)We validate the representation on three trajectory\-forecasting benchmarks and a closed\-loop stress test \(§[5](https://arxiv.org/html/2608.05454#S5)\)\.
Table 1:Design\-space positioning\. HProbZ targets a combination of structural properties absent from same\-encoder mixtures, conformal predictors, and recent SOTA forecasters; minADE is one axis, not the optimisation target\.Decomp\.: identifiable modal/bounded/stochastic factorisation\.Multi\-modal CP: distribution\-free non\-convex conformal set\.Refinement:O\(1/k\)O\(1/k\)contraction \(Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)\)\.Raw point accuracy is one axis and HProbZ does not occupy its extreme; we occupy an axis on which no existing method we tested provides all four structural properties simultaneously \(Tab\.[1](https://arxiv.org/html/2608.05454#S1.T1)\)\.
## 2Preliminaries
A zonotope\{c\+Gβ∣β∈\[−1,1\]m\}\\\{c\+G\\beta\\mid\\beta\\in\[\-1,1\]^\{m\}\\\}is a centrally symmetric polytope with centerc∈ℝdc\\in\\mathbb\{R\}^\{d\}and generator matrixG∈ℝd×mG\\in\\mathbb\{R\}^\{d\\times m\}\. A constrained zonotope adds linear equality constraintsAβ=bA\\beta=b, restricting to a subset of the zonotope’s interior\. The hybrid zonotope\[[5](https://arxiv.org/html/2608.05454#bib.bib15)\]augments with binary factorsξ∈\{−1,1\}nb\\xi\\in\\\{\-1,1\\\}^\{n\_\{b\}\}and represents a non\-convex region as a union of up to2nb2^\{n\_\{b\}\}constrained zonotopes\. Stochastic\-reachability work, originating in safety analysis of automated traffic\[[2](https://arxiv.org/html/2608.05454#bib.bib16)\], instead augments zonotopes with Gaussian factorsη∼𝒩\(0,I\)\\eta\\sim\\mathcal\{N\}\(0,I\), inducing a distribution over the zonotope’s interior\. An MDN\[[6](https://arxiv.org/html/2608.05454#bib.bib14)\]parameterizesp\(y\|x\)=∑k=1Kπk\(x\)𝒩\(y;μk\(x\),Σk\(x\)\)p\(y\|x\)=\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}\(x\)\\mathcal\{N\}\(y;\\mu\_\{k\}\(x\),\\Sigma\_\{k\}\(x\)\)via network outputs; the within\-component covariancesΣk\\Sigma\_\{k\}are fixed outputs that cannot be refined by observations, and trajectory\-prediction MDNs\[[8](https://arxiv.org/html/2608.05454#bib.bib8)\]are known to suffer mode\-collapse failures\[[20](https://arxiv.org/html/2608.05454#bib.bib9)\]that motivate the structured\-decomposition alternative we develop below\.
#### The gap\.
The hybrid form carries discrete modes but no likelihood; the probabilistic form carries a likelihood but only a single mode; an MDN supplies both, but its independent\-Gaussian components let observations only re\-weight modes, never tighten within\-component covariance\. A predictive distribution that keeps binary modes separable from bounded drift and Gaussian noise, admits a closed\-form likelihood, and tightens its bounded support under observation through shared algebraic structure exists in neither lineage; this is the gap HProbZ fills in §[3](https://arxiv.org/html/2608.05454#S3)\.
## 3Hybrid Probabilistic Zonotopes
###### Definition 1\(Hybrid Probabilistic Zonotope \(HProbZ\)\)\.
An HProbZ is a random set defined by
𝒵=\{c\+Gbβ\+Gdα\+Gsν\|β∈\{−1,1\}nb,α∈\[−1,1\]nd,ν∼𝒩\(0,Inq\),Abβ\+Adα\+Asν=b\}\\mathcal\{Z\}=\\left\\\{c\+G\_\{b\}\\beta\+G\_\{d\}\\alpha\+G\_\{s\}\\nu\\;\\middle\|\\;\\begin\{matrix\}\\beta\\in\\\{\-1,1\\\}^\{n\_\{b\}\},\\;\\alpha\\in\[\-1,1\]^\{n\_\{d\}\},\\;\\nu\\sim\\mathcal\{N\}\(0,I\_\{n\_\{q\}\}\),\\\\ A\_\{b\}\\beta\+A\_\{d\}\\alpha\+A\_\{s\}\\nu=b\\end\{matrix\}\\right\\\}\(1\)wherec∈ℝdc\\in\\mathbb\{R\}^\{d\}is the center,Gb∈ℝd×nbG\_\{b\}\\in\\mathbb\{R\}^\{d\\times n\_\{b\}\},Gd∈ℝd×ndG\_\{d\}\\in\\mathbb\{R\}^\{d\\times n\_\{d\}\}, andGs∈ℝd×nqG\_\{s\}\\in\\mathbb\{R\}^\{d\\times n\_\{q\}\}are the binary, bounded, and stochastic generator matrices, respectively, andAb∈ℝp×nbA\_\{b\}\\in\\mathbb\{R\}^\{p\\times n\_\{b\}\},Ad∈ℝp×ndA\_\{d\}\\in\\mathbb\{R\}^\{p\\times n\_\{d\}\},As∈ℝp×nqA\_\{s\}\\in\\mathbb\{R\}^\{p\\times n\_\{q\}\},b∈ℝpb\\in\\mathbb\{R\}^\{p\}encodepplinear equality constraints\.
The single\-mode slice \(nb=0n\_\{b\}\{=\}0,p=0p\{=\}0\) recovers the probabilistic zonotope of stochastic reachability\[[2](https://arxiv.org/html/2608.05454#bib.bib16)\]; HProbZ extends it with binary modes and observation\-activated constraints\. Paralleling the hybrid zonotope\[[5](https://arxiv.org/html/2608.05454#bib.bib15)\], for each binary assignmentβ\\betathe constraintAbβ\+Adα\+Asν=bA\_\{b\}\\beta\+A\_\{d\}\\alpha\+A\_\{s\}\\nu=breduces to a Constrained Probabilistic Zonotope𝒞\(β\)\\mathcal\{C\}\(\\beta\)with centerc\+Gbβc\+G\_\{b\}\\betaand residual constraintsAdα\+Asν=b−AbβA\_\{d\}\\alpha\+A\_\{s\}\\nu=b\-A\_\{b\}\\beta, so the HProbZ is a union of2nb2^\{n\_\{b\}\}CPZs:
𝒵=⋃β∈\{−1,1\}nb𝒞\(β\)\.\\mathcal\{Z\}=\\bigcup\_\{\\beta\\in\\\{\-1,1\\\}^\{n\_\{b\}\}\}\\mathcal\{C\}\(\\beta\)\.\(2\)Whenp=0p\{=\}0all factors are independent andz=c\+Gbβ\+Gdα\+Gsνz=c\+G\_\{b\}\\beta\+G\_\{d\}\\alpha\+G\_\{s\}\\nu; this is the head used by the network in §[4](https://arxiv.org/html/2608.05454#S4)\. Observations introduce new constraint rows into\(Ab,Ad,As,b\)\(A\_\{b\},A\_\{d\},A\_\{s\},b\)for sequential refinement \(§[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px3)\)\.
Setting all generators to zero recovers a point prediction; retaining onlyGdG\_\{d\}yields a zonotope, onlyGsG\_\{s\}a Gaussian,Gb\+GsG\_\{b\}\{\+\}G\_\{s\}a2nb2^\{n\_\{b\}\}\-component shared\-covariance GMM, and diagonalGdG\_\{d\}alone an axis\-aligned conformal box\. In the unconstrained case the induced measure is a Minkowski\-sum convolution \(Proposition[8](https://arxiv.org/html/2608.05454#Thmtheorem8)\):
μ𝒵=δc∗\(12nb∑βδGbβ\)⏟discrete modal∗\(Gd\)∗𝒰\(\[−1,1\]nd\)⏟bounded zonoid∗𝒩\(0,GsGs⊤\)⏟Gaussian,\\mu\_\{\\mathcal\{Z\}\}=\\delta\_\{c\}\\ast\\underbrace\{\\Big\(\\tfrac\{1\}\{2^\{n\_\{b\}\}\}\\sum\_\{\\beta\}\\delta\_\{G\_\{b\}\\beta\}\\Big\)\}\_\{\\text\{discrete modal\}\}\\ast\\underbrace\{\(G\_\{d\}\)\_\{\*\}\\mathcal\{U\}\(\[\-1,1\]^\{n\_\{d\}\}\)\}\_\{\\text\{bounded zonoid\}\}\\ast\\underbrace\{\\mathcal\{N\}\(0,G\_\{s\}G\_\{s\}^\{\\top\}\)\}\_\{\\text\{Gaussian\}\},\(3\)where∗\\astdenotes convolution\. The bounded factor has compact support and is not absolutely continuous w\.r\.t\. any Gaussian — an inductive bias giving HProbZ algebraic properties Gaussian mixtures lack, at a cost of density flexibility\.
###### Theorem 2\(HProbZ density is not a finite Gaussian mixture\)\.
IfGdG\_\{d\}has at least one nonzero entry, then no finite Gaussian mixture∑k=1Kπk𝒩\(μk,Σk\)\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}\\mathcal\{N\}\(\\mu\_\{k\},\\Sigma\_\{k\}\), for anyKK, has the same distribution asμ𝒵\\mu\_\{\\mathcal\{Z\}\}\.
###### Proof sketch\.
Project both CFs onto a direction whereGdG\_\{d\}is nonzero and continue tos=iys=iy: the HProbZ CF picks up asinh\(ay\)/\(ay\)\\sinh\(ay\)/\(ay\)factor with an algebraic1/y1/ycorrection that no finite\-GMM CF \(a sum of pure exponentials\) can match\. Full argument in App\.[B](https://arxiv.org/html/2608.05454#A2); empirical fingerprint in Fig\.[2](https://arxiv.org/html/2608.05454#S3.F2)\. ∎
Figure 2:Empirical fingerprint of Theorem[2](https://arxiv.org/html/2608.05454#Thmtheorem2)\(1\-D HProbZ witha=2,σ=0\.3a\{=\}2,\\sigma\{=\}0\.3,10510^\{5\}samples\)\.Left:GMMs \(K=1,…,64K\{=\}1\{,\}\\ldots\{,\}64\) visually approximate the density\.Middle:at the sinc\-zero frequenciestk=kπ/at\_\{k\}=k\\pi/a,\|ϕGMM\|\|\\phi\_\{\\text\{GMM\}\}\|plateaus at∼2−5×10−3\\sim 2\{\-\}5\{\\times\}10^\{\-3\}regardless ofKK— the structural signature the proof predicts no finite GMM can erase\.Right:matching HProbZ’s held\-out NLL with a finite GMM requiresK=8K\{=\}8\(2323free density parameters\) vs HProbZ’s22, an11\.5×11\.5\\timesratio\. Full setup in App\.[C](https://arxiv.org/html/2608.05454#A3)\.A natural question is whether the richer parameterisation is identifiable\.
###### Theorem 3\(Identifiability under exact likelihood\)\.
SupposeGdG\_\{d\}andGsG\_\{s\}are diagonal \(the parameterisation used in §[4](https://arxiv.org/html/2608.05454#S4)\), and \(i\) for each axisjjthe bounded half\-widthaj=\|Gd,jj\|\>0a\_\{j\}=\|G\_\{d,jj\}\|\>0and noise scaleσj=\|Gs,jj\|\>0\\sigma\_\{j\}=\|G\_\{s,jj\}\|\>0, and \(ii\) for each axisjjthe projected mode centers\{\(Gbβ\)j\}β∈\{−1,1\}nb\\\{\(G\_\{b\}\\beta\)\_\{j\}\\\}\_\{\\beta\\in\\\{\-1,1\\\}^\{n\_\{b\}\}\}are pairwise distinct\. Then\(c,Gb,Gd,Gs\)\(c,G\_\{b\},G\_\{d\},G\_\{s\}\)are identifiable from the exact convolution likelihood, up to column\-permutations and global column sign\-flips ofGbG\_\{b\}\.
###### Proof sketch\.
Per\-axis projection yields a 1\-D mixture of uniform\-normal convolutions: sinc zeros of the 1\-D CF give the bounded half\-widthsaja\_\{j\}, the Gaussian envelope givesσj\\sigma\_\{j\}, and classical location\-mixture identifiability\[[41](https://arxiv.org/html/2608.05454#bib.bib40)\]pins down the mode centers\. Cross\-axis consistency reduces the labeling ambiguity to the cube symmetry above\. See App\.[B](https://arxiv.org/html/2608.05454#A2)\. ∎
#### Practical identifiability\.
The distinct\-centers condition is generic for randomGbG\_\{b\}; SGD on softplus outputs reliably keepsaj,σja\_\{j\},\\sigma\_\{j\}positive and distinct, and 10 seeds recover ground\-truth parameters within0\.0080\.008\(App\.[C](https://arxiv.org/html/2608.05454#A3)\)\.
## 4Training HProbZ as a Neural Network Output
#### Output parameterization\.
We attach an unconstrained HProbZ head \(p=0p\{=\}0\) to a transformer encoder:h=Encoder\(x\)h=\\text\{Encoder\}\(x\),c=Wchc=W\_\{c\}handGb=WbhG\_\{b\}=W\_\{b\}hare linear projections,Gd,GsG\_\{d\},G\_\{s\}diagonal softplus outputs\. The head adds\(nb\+nd\+nq\+1\)×d\(n\_\{b\}\+n\_\{d\}\+n\_\{q\}\+1\)\\times dparameters, matching MDN\-K=2nb2^\{n\_\{b\}\}at the same count\. Constraints from Def\.[1](https://arxiv.org/html/2608.05454#Thmtheorem1)activate only at inference \(§[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px3)\); non\-diagonalGd,GsG\_\{d\},G\_\{s\}left to future work\. We use a uniform mode priorπk=2−nb\\pi\_\{k\}\{=\}2^\{\-n\_\{b\}\}; learnableπk\(x\)\>0\\pi\_\{k\}\(x\)\{\>\}0preserves identifiability/contraction with statistically tied accuracy \(14\-seed nuScenes, App\.[D\.28](https://arxiv.org/html/2608.05454#A4.SS28)\)\.
#### The Gaussian approximation and its failure\.
The naïve aggregate\-covariance lossℒGauss=−log∑k2−nb𝒩\(y;μk,GsGs⊤\+GdGd⊤/3\)\\mathcal\{L\}\_\{\\text\{Gauss\}\}=\-\\log\\sum\_\{k\}2^\{\-n\_\{b\}\}\\mathcal\{N\}\(y;\\mu\_\{k\},G\_\{s\}G\_\{s\}^\{\\top\}\+G\_\{d\}G\_\{d\}^\{\\top\}/3\)withμk=c\+Gbβ\(k\)\\mu\_\{k\}=c\+G\_\{b\}\\beta^\{\(k\)\}is non\-identifiable:
###### Proposition 4\(Gaussian Collapse\)\.
UnderℒGauss\\mathcal\{L\}\_\{\\text\{Gauss\}\},GdG\_\{d\}andGsG\_\{s\}are interchangeable: for any\(Gd,Gs\)\(G\_\{d\},G\_\{s\}\)achieving lossℓ\\ell, there existsGs′G\_\{s\}^\{\\prime\}withGs′\(Gs′\)⊤=GsGs⊤\+GdGd⊤/3G\_\{s\}^\{\\prime\}\(G\_\{s\}^\{\\prime\}\)^\{\\top\}=G\_\{s\}G\_\{s\}^\{\\top\}\+G\_\{d\}G\_\{d\}^\{\\top\}/3such that\(0,Gs′\)\(0,G\_\{s\}^\{\\prime\}\)achieves the sameℓ\\ell, and vice versa\. Gradient\-based optimization therefore cannot distinguish bounded from stochastic uncertainty\.
###### Proof\.
ℒGauss\\mathcal\{L\}\_\{\\text\{Gauss\}\}depends on\(Gd,Gs\)\(G\_\{d\},G\_\{s\}\)only through the aggregate covarianceΣapprox=GsGs⊤\+GdGd⊤/3\\Sigma\_\{\\text\{approx\}\}=G\_\{s\}G\_\{s\}^\{\\top\}\+G\_\{d\}G\_\{d\}^\{\\top\}/3\. ChoosingGs′G\_\{s\}^\{\\prime\}withGs′\(Gs′\)⊤=ΣapproxG\_\{s\}^\{\\prime\}\(G\_\{s\}^\{\\prime\}\)^\{\\top\}=\\Sigma\_\{\\text\{approx\}\}\(e\.g\., the Cholesky factor\) andGd=0G\_\{d\}=0leavesΣapprox\\Sigma\_\{\\text\{approx\}\}and henceℒGauss\\mathcal\{L\}\_\{\\text\{Gauss\}\}unchanged\. ∎
The resolution is to train with the exact within\-mode likelihood\. Conditioned on modekk, the residualrj=yj−μk,jr\_\{j\}=y\_\{j\}\-\\mu\_\{k,j\}in each dimension follows the convolutionUniform\(−aj,aj\)∗𝒩\(0,σj2\)\\text\{Uniform\}\(\-a\_\{j\},a\_\{j\}\)\\ast\\mathcal\{N\}\(0,\\sigma\_\{j\}^\{2\}\), whose density admits the closed form
f\(r;a,σ\)=12a\[Φ\(r\+aσ\)−Φ\(r−aσ\)\],f\(r;a,\\sigma\)=\\frac\{1\}\{2a\}\\Big\[\\Phi\\\!\\Big\(\\frac\{r\+a\}\{\\sigma\}\\Big\)\-\\Phi\\\!\\Big\(\\frac\{r\-a\}\{\\sigma\}\\Big\)\\Big\],\(4\)whereΦ\\Phiis the standard normal CDF \(numerical clamp and Gaussian\-limit fallback in App\.[A](https://arxiv.org/html/2608.05454#A1)\)\. The exact mixture NLL isℒexact=−log∑k2−nb∏jf\(yj−μk,j;aj,σj\)\\mathcal\{L\}\_\{\\text\{exact\}\}=\-\\log\\sum\_\{k\}2^\{\-n\_\{b\}\}\\prod\_\{j\}f\(y\_\{j\}\-\\mu\_\{k,j\};a\_\{j\},\\sigma\_\{j\}\)\. Its CF has zeros at integer multiples ofπ/a\\pi/a, a compact\-support spectral signature absent from any Gaussian; this resolves the collapse in Prop\.[4](https://arxiv.org/html/2608.05454#Thmtheorem4)and enables Theorem[3](https://arxiv.org/html/2608.05454#Thmtheorem3)\.
#### The structure–density trade\-off\.
Introducing bounded generators witha\>0a\>0imposes compact support, which necessarily reduces entropy relative to a Gaussian with the same variance\. However, this cost is bounded:
###### Proposition 5\(Structure–density trade\-off\)\.
Letfκf\_\{\\kappa\}denote the 1\-D HProbZ density with half\-widthaaand noiseσ\\sigma, and letgκ=𝒩\(0,σ2\+a2/3\)g\_\{\\kappa\}=\\mathcal\{N\}\(0,\\sigma^\{2\}\+a^\{2\}/3\)be the variance\-matched Gaussian\. Define the structure ratioκ=a/σ\\kappa=a/\\sigma\. Then:
\(i\)Bounded entropy cost\.The entropy gaph\(gκ\)−h\(fκ\)h\(g\_\{\\kappa\}\)\-h\(f\_\{\\kappa\}\)is monotonically increasing inκ\\kappaand satisfies
h\(gκ\)−h\(fκ\)≤12log\(πe/6\)≈0\.176nats,h\(g\_\{\\kappa\}\)\-h\(f\_\{\\kappa\}\)\\;\\leq\\;\\tfrac\{1\}\{2\}\\log\\\!\\big\(\\pi e/6\\big\)\\;\\approx\\;0\.176\\text\{ nats\},\(5\)for allκ≥0\\kappa\\geq 0, with equality in the limitκ→∞\\kappa\\to\\infty\.
\(ii\)Unbounded constraint benefit\.In the constrained setting of Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7), the per\-observation Fisher information for the shared bounded factor isρ≥κ2\\rho\\geq\\kappa^\{2\}per dimension, yielding bounded\-variance decayO\(1/\(κ2k\)\)O\(1/\(\\kappa^\{2\}k\)\)afterkkobservations; the dimensionless derivation is in Appendix[B\.5](https://arxiv.org/html/2608.05454#A2.SS5)\.
\(iii\)Exchange rate\.The entropy cost is bounded by a universal constant independent ofκ\\kappa, while the constraint contraction rate grows without bound asκ\\kappaincreases\. In particular, for anyκ\>0\\kappa\>0, the per\-observation information gainρ\\rhocan be made arbitrarily large at a per\-dimension NLL cost of at most0\.1760\.176nats\.
Proposition[5](https://arxiv.org/html/2608.05454#Thmtheorem5)bounds the entropy gap betweenfκf\_\{\\kappa\}and its variance\-matched Gaussian by0\.1760\.176nats/dim; it does*not*bound the empirical NLL gap, which also contains a data\-dependent misspecification term\. Direct computation on three benchmarks shows zero violations across13,000\+13\{,\}000\+samples \(App\.[C](https://arxiv.org/html/2608.05454#A3)\)\.
#### Split\-conformal calibration\.
To obtain distribution\-free coverage guarantees even under model misspecification, we define the HProbZ nonconformity score
s\(x,y\)=minβ∈\{−1,1\}nbmaxj∈\[d\]\|yj−cj\(x\)−\(Gb\(x\)β\)j\|−‖Gd\(x\)\[j,:\]‖1‖Gs\(x\)\[j,:\]‖2,s\(x,y\)=\\min\_\{\\beta\\in\\\{\-1,1\\\}^\{n\_\{b\}\}\}\\max\_\{j\\in\[d\]\}\\frac\{\|y\_\{j\}\-c\_\{j\}\(x\)\-\(G\_\{b\}\(x\)\\beta\)\_\{j\}\|\-\\\|G\_\{d\}\(x\)\[j,:\]\\\|\_\{1\}\}\{\\\|G\_\{s\}\(x\)\[j,:\]\\\|\_\{2\}\},\(6\)which selects the closest binary mode and measures the worst\-dimension standardized residual after exhausting the bounded generator’s reach\.
###### Theorem 6\(Distribution\-free HProbZ coverage\)\.
Assume‖Gs\(x\)\[j,:\]‖2\>0\\\|G\_\{s\}\(x\)\[j,:\]\\\|\_\{2\}\>0for everyxxandjj\(in practice ensured by softplus\)\. Letqαq\_\{\\alpha\}be the⌈\(1−α\)\(n\+1\)⌉/n\\lceil\(1\{\-\}\\alpha\)\(n\{\+\}1\)\\rceil/nquantile of calibration scores\{s\(xi,yi\)\}i=1n\\\{s\(x\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{n\}\. The prediction setS^α\(x\)=⋃β\{z:\|zj−cj−\(Gbβ\)j\|≤‖Gd\[j,:\]‖1\+qα‖Gs\[j,:\]‖2,∀j\}\\hat\{S\}\_\{\\alpha\}\(x\)=\\bigcup\_\{\\beta\}\\\{z:\|z\_\{j\}\-c\_\{j\}\-\(G\_\{b\}\\beta\)\_\{j\}\|\\leq\\\|G\_\{d\}\[j,:\]\\\|\_\{1\}\+q\_\{\\alpha\}\\\|G\_\{s\}\[j,:\]\\\|\_\{2\},\\,\\forall j\\\}satisfiesℙ\(ytest∈S^α\)≥1−α\\mathbb\{P\}\(y\_\{\\text\{test\}\}\\in\\hat\{S\}\_\{\\alpha\}\)\\geq 1\{\-\}\\alphaunder only exchangeability\.
S^α\\hat\{S\}\_\{\\alpha\}is a union of up to2nb2^\{n\_\{b\}\}axis\-aligned boxes rather than a single convex region; the per\-dimension edge is22–3×3\\timesshorter than Standard CP and∼1\.5×\\sim 1\.5\\timesshorter than MDN\-CP\. Empirical volumes, coverage, and OOD robustness are reported in §[5\.5](https://arxiv.org/html/2608.05454#S5.SS5); a parallel distributional coverage boundP\(y∈Sγ\)≥1−2de−γ2/2P\(y\\in S\_\{\\gamma\}\)\\geq 1\-2de^\{\-\\gamma^\{2\}/2\}wheny∣xy\\mid xfollows HProbZ is in App\.[B](https://arxiv.org/html/2608.05454#A2)\.
## 5Experiments
We test four questions: \(i\) does the exact likelihood recover the modal/bounded/stochastic decomposition; \(ii\) is HProbZ competitive on accuracy benchmarks; \(iii\) does structured uncertainty improve downstream decisions; \(iv\) does the algebraic constraint realise theO\(1/k\)O\(1/k\)contraction of Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)\. HProbZ and MDN share encoder, schedule, and sampling budget throughout \(App\.[A\.4](https://arxiv.org/html/2608.05454#A1.SS4),[A](https://arxiv.org/html/2608.05454#A1)\); metrics are minADE/minFDE for accuracy, CRPS and conformal coverage for calibration, per\-mode risk and closed\-loop collision for decisions; raw NLL in App\.[C\.6](https://arxiv.org/html/2608.05454#A3.SS6)\.
### 5\.1Controlled Decomposition
We construct a trajectory task with three independently controllable uncertainty sources — branchesKK, driftδdrift\\delta\_\{\\mathrm\{drift\}\}, noiseσnoise\\sigma\_\{\\mathrm\{noise\}\}— and train HProbZ under both exact and Gaussian\-approximate likelihoods, measuring whether each generator selectively tracks its matching source\. The full sweep figure and per\-value tables \(Fig\.[5](https://arxiv.org/html/2608.05454#A3.F5), App\.[C\.1](https://arxiv.org/html/2608.05454#A3.SS1)\) match the theoretical predictions: under exact likelihood,‖Gb‖\\\|G\_\{b\}\\\|grows∼9×\\sim 9\\timesacross the mode sweep while drift is routed toGdG\_\{d\}\(Δ‖Gd‖=2\.90≫Δ‖Gs‖=2\.27\\Delta\\\|G\_\{d\}\\\|\{=\}2\.90\\gg\\Delta\\\|G\_\{s\}\\\|\{=\}2\.27\); the Gaussian approximation reverses the drift dominance \(Δ‖Gs‖=2\.70\\Delta\\\|G\_\{s\}\\\|\{=\}2\.70wins\), a direct witness of Prop\.[4](https://arxiv.org/html/2608.05454#Thmtheorem4)\. Theorem[2](https://arxiv.org/html/2608.05454#Thmtheorem2)’s CF fingerprint is in Fig\.[2](https://arxiv.org/html/2608.05454#S3.F2); Theorem[3](https://arxiv.org/html/2608.05454#Thmtheorem3)’s multi\-seed parameter recovery \(within0\.0080\.008across 10 seeds\) is in App\.[C](https://arxiv.org/html/2608.05454#A3)\.
#### Decomposition on real driving data\.
The same modal/bounded/stochastic split is recovered without supervision on nuScenes vehicle trajectories\. Stratifying the trainedd=256d\{=\}256model by past\-observed vehicle speed \(Fig\.[3](https://arxiv.org/html/2608.05454#S5.F3)\),‖Gb‖\\\|G\_\{b\}\\\|scales by∼370×\\sim 370\\timesfrom stationary to high\-speed agents \(modal ambiguity grows with speed\),‖Gs‖\\\|G\_\{s\}\\\|scales by∼94×\\sim 94\\times\(stochastic noise\), and‖Gd‖\\\|G\_\{d\}\\\|stays within5\.4×5\.4\\times\(systematic drift is less speed\-dependent\)\. The three generators therefore track interpretable, qualitatively distinct sources of uncertainty on real trajectories, not only on the controlled task above; setup details in App\.[D\.17](https://arxiv.org/html/2608.05454#A4.SS17)\.



Figure 3:Structural\-property validation on trained HProbZ models\.Left:nuScenes speed\-stratified generator norms \(log\-yy\) —‖Gb‖\\\|G\_\{b\}\\\|scales∼370×\\sim 370\\timesstationary→\\tohigh\-speed,‖Gs‖\\\|G\_\{s\}\\\|scales∼94×\\sim 94\\times,‖Gd‖\\\|G\_\{d\}\\\|stays within5\.4×5\.4\\times, so the three generators track interpretable, qualitatively distinct uncertainty sources \(Theorem[3](https://arxiv.org/html/2608.05454#Thmtheorem3); setup App\.[D\.17](https://arxiv.org/html/2608.05454#A4.SS17)\)\.Middle:entropy\-gap CDF across13,000\+13\{,\}000\{\+\}samples on three benchmarks; all curves lie strictly below the theoretical bound12log\(πe/6\)≈0\.1765\\tfrac\{1\}\{2\}\\log\(\\pi e/6\)\\\!\\approx\\\!0\.1765nats/dim \(red dashed\)\.Right:per\-dataset distribution of the same gap \(symmetric log scale\)\. Zero violations on any benchmark for Proposition[5](https://arxiv.org/html/2608.05454#Thmtheorem5)\(App\.[C\.4](https://arxiv.org/html/2608.05454#A3.SS4)\)\.
### 5\.2Real\-World Pedestrian Trajectories
ETH/UCY leave\-one\-out \(8\-step past, 12\-step future\)\. Base HProbZ: single\-agent transformer encoder, exact likelihood,nb=3,nd=1,nq=1n\_\{b\}\{=\}3,n\_\{d\}\{=\}1,n\_\{q\}\{=\}1\. Social variant adds 8\-neighbour attention\. Baselines: 8\-component MDN and CVAE with identical input features\.
Table 2:ETH/UCY: minADE@20 / minFDE@20 \(m\), 5\-scene leave\-one\-out averages, no maps\. Base =d=64d\{=\}64, 3 seeds \(CVAE wins this row\); Social =d=128d\{=\}128\+ 8\-neighbour attention, 5 seeds\. Social\-CVAE not included due to compute\. Per\-scene table in App\.[D\.13](https://arxiv.org/html/2608.05454#A4.SS13)\.Social HProbZ reaches0\.197/0\.3320\.197/0\.332minADE/minFDE, beating same\-encoder MDN\-K=8K\{=\}8by39\.3%\\mathbf\{39\.3\\%\}ADE /20\.7%\\mathbf\{20\.7\\%\}FDE \(Tab\.[2](https://arxiv.org/html/2608.05454#S5.T2), 5 seeds\), matching LED \(0\.180\.18\) and EqMotion \(0\.190\.19\) on minADE and trailing them by∼0\.06\\sim 0\.06m on minFDE \(full vs\-published table in App\.[D\.16](https://arxiv.org/html/2608.05454#A4.SS16)\); goal\-conditioned methods like Y\-Net/LED sharpen endpoints by conditioning on predicted endpoints but provide no decomposition, conformal coverage, or refinement\.
At based=64d\{=\}64HProbZ beats MDN\-K=8K\{=\}8on ADE but trails CVAE; the social variant recovers and surpasses CVAE\. The∼21%\\sim 21\\%ADE advantage over MDN\-K=20K\{=\}20is stable across encoder sizes 110K–2\.4M \(App\.[D\.6](https://arxiv.org/html/2608.05454#A4.SS6)\), ruling out a capacity confound on minADE;*on minFDE MDN\-K=20K\{=\}20wins at all three capacities*, so the social\-row FDE advantage comes from the social\-attention encoder, not the head\. Raw NLL exceeds MDN’s, a compact\-support misspecification penalty \(App\.[C\.6](https://arxiv.org/html/2608.05454#A3.SS6)\); we lead with ADE/FDE and downstream metrics\. A controlled cross\-method comparison \(App\.[D\.12](https://arxiv.org/html/2608.05454#A4.SS12)\) gives HProbZ FDE@200\.390\.39m vs MDN\-K2/K40\.410\.41–0\.420\.42and a lightly\-tuned Diffusion0\.430\.43/2\.242\.24\(DDIM\-5050/\-55\) at∼530×\\sim 530\\timesslower inference \(diffusion baseline not swept over schedule/depth/target/CFG\)\.
### 5\.3nuScenes Vehicle Trajectory Benchmark
nuScenes\[[7](https://arxiv.org/html/2608.05454#bib.bib39)\]: 40k test windows at 2 Hz\. Samed=128d\{=\}128encoder; HProbZ \(nb=1n\_\{b\}\{=\}1, 278k params\) vs MDN\-K2 \(281k\), MDN\-K4 under identical training\. Structured head runtime overhead1\.4%1\.4\\%on H200 at batch 256\.
Table 3:nuScenes minADE \(m\): published map\-aware baselines, trajectory\-onlyd=128d\{=\}128controlled sweep \(3 seeds\), andd=256d\{=\}256matched\-encoder\. MDN\-K2†: best\-tuned \(App\.[D\.22](https://arxiv.org/html/2608.05454#A4.SS22)\)\.*Seed asymmetry in thed=256d\{=\}256block*: HProbZ is mean±\\pmstd over 14 seeds \(App\.[D\.24](https://arxiv.org/html/2608.05454#A4.SS24)\); MDN\-K2 row is single\-seed \(no±\\pmstd\), so the46%46\\%gap is HProbZ’s 14\-seed mean vs one MDN\-K2 seed\.Atd=128d\{=\}128HProbZ beats best\-tuned MDN\-K2 by23%/19%23\\%/19\\%atK=5/K=20K\{=\}5/K\{=\}20and reaches1\.46±0\.021\.46\{\\pm\}0\.02minADE@10 \(3 seeds\), matching MID’s1\.441\.44within rounding\. Scaling MDN toK=8K\{=\}8regresses to3\.05±0\.233\.05\{\\pm\}0\.23\(mode collapse, App\.[D\.27](https://arxiv.org/html/2608.05454#A4.SS27)\)\. Scaling the shared encoder tod=256d\{=\}256\(2\.132\.13M params, sym\-brokenGbG\_\{b\}init, App\.[D\.24](https://arxiv.org/html/2608.05454#A4.SS24)\), HProbZ reaches1\.29±0\.07\\mathbf\{1\.29\{\\pm\}0\.07\}minADE@10 over 14 seeds, approaching MID’s published1\.441\.44without maps\. The sym\-break is load\-bearing: with random init the same architecture trains to1\.48±0\.101\.48\{\\pm\}0\.10over 5 seeds \(above MID; full ablation Tab\.[33](https://arxiv.org/html/2608.05454#A4.T33)\)\. Single\-seed same\-encoder MDN\-K2 reaches2\.412\.41atd=256d\{=\}256\(a46%46\\%gap to HProbZ’s 14\-seed mean; multi\-seed MDN d=256 reproduction left to future work\), and same\-encoder DDPM reaches2\.992\.99at5\.2×5\.2\\timesparams and500×500\\timesslower \(App\.[D\.11](https://arxiv.org/html/2608.05454#A4.SS11)\)\.
#### Calibration: CRPS\.
Onn=5n\{=\}5paired matched\-architecture seeds, HProbZ and MDN\-K2 are statistically tied on mean CRPS; per\-seed spread analysis is in App\.[C\.5](https://arxiv.org/html/2608.05454#A3.SS5)\.
### 5\.4Argoverse 2 Motion Forecasting
Argoverse 2\[[37](https://arxiv.org/html/2608.05454#bib.bib13)\]: 200k scenarios, 6 s horizon\. Samed=128d\{=\}128encoder, 2 Hz subsampling, 618k training windows, identical conditions across rows\.
Table 4:Argoverse 2 minADE \(m\), 3 seeds, same\-encoder trajectory\-only models \(d=128d\{=\}128\)\.HProbZ wins minADE at everyKK, beating MDN\-K2 by∼7%\\sim 7\\%across 3 seeds \(Tab\.[4](https://arxiv.org/html/2608.05454#S5.T4)\); seed\-level std is4−5×4\{\-\}5\\timestighter \(σHProbZ≤0\.02\\sigma\_\{\\mathrm\{HProbZ\}\}\\\!\\leq\\\!0\.02m vs\. MDN’s0\.05−0\.100\.05\{\-\}0\.10, directional atn=3n\{=\}3\)\. The Prop\.[5](https://arxiv.org/html/2608.05454#Thmtheorem5)entropy\-gap bound is confirmed across three benchmarks \(0 violations,13,000\+13\{,\}000\{\+\}samples; Fig\.[3](https://arxiv.org/html/2608.05454#S5.F3), App\.[C\.4](https://arxiv.org/html/2608.05454#A3.SS4)\)\.
#### Downstream risk\-aware planning\.
On a pedestrian\-crossing scenario \(App\.[D\.23](https://arxiv.org/html/2608.05454#A4.SS23)\), the structured planner that enumerates2nb2^\{n\_\{b\}\}binary modes and evaluates per\-mode risk analytically recovers oracle cost0on Modal\-Safe scenes vs\. the Gaussian\-collapsed planner’s1\.51\.5, with the advantage emerging at mode separationΔy≳4\\Delta y\\gtrsim 4\(App\.[D\.18](https://arxiv.org/html/2608.05454#A4.SS18)\)\. Across5,0005\{,\}000nuScenes scenarios×20\\times 20seeds, HProbZ’s analytic per\-mode probability is deterministic \(zero variance, zero flips\); MDN’s sample estimator decays from33%33\\%flips atS=2S\{=\}2to0\.9%0\.9\\%atS=100S\{=\}100but still flips2\.7%2\.7\\%of*ambiguous*cases \(full sample\-budget sweep in App\.[D\.19](https://arxiv.org/html/2608.05454#A4.SS19)\)\.
#### Closed\-loop collision avoidance\.
On a66s static\-obstacle stress test over nuScenes \(N=39,829N\{=\}39\{,\}829, App\.[D\.20](https://arxiv.org/html/2608.05454#A4.SS20)\), HProbZ produces𝟑𝟎%\\mathbf\{30\\%\}fewer trajectory\-level joint\-event collisions than same\-encoder MDN\-K2 at near\-matched10%10\\%FAR \(Tab\.[5](https://arxiv.org/html/2608.05454#S5.T5); FARs9\.5%9\.5\\%vs10\.6%10\.6\\%\)\.*The per\-step variant reverses*\(HProbZ8\.01%8\.01\\%vs MDN6\.32%6\.32\\%\): the shared\(β,α\)\(\\beta,\\alpha\)pays off through correlated per\-step events that MDN’s per\-step independent Gaussians underestimate at the joint\-tail; without that correlation MDN’s per\-step flexibility dominates\. A planner conditions on the joint event\.
Table 5:Closed\-loop collision avoidance on nuScenes \(N=39,829N\{=\}39\{,\}829,66s,r=2r\{=\}2m, 3 seeds\)\. HProbZ wins trajectory\-level by30%30\\%at near\-matched10%10\\%FAR \(HProbZ9\.5%9\.5\\%, MDN10\.6%10\.6\\%\); MDN\-K2 wins per\-step\. Bold = column winner\.
#### Constrained sequential refinement\.
Withβ,α\\beta,\\alphashared across steps viazt=ct\+Gb,tβ\+Gd,tα\+Gs,tνtz\_\{t\}=c\_\{t\}\+G\_\{b,t\}\\beta\+G\_\{d,t\}\\alpha\+G\_\{s,t\}\\nu\_\{t\}, each observation adds a constraint row that simultaneously reweights the2nb2^\{n\_\{b\}\}mode likelihoods and tightensα\\alpha’s feasible set; the tightening propagates to all unobserved steps\.
###### Theorem 7\(Contraction of constrained HProbZ, informal\)\.
Under the shared\-generator model with the variance\-matched Gaussian relaxationα∼𝒩\(0,13I\)\\alpha\\sim\\mathcal\{N\}\(0,\\tfrac\{1\}\{3\}I\), conditioning onkkobservations\(i\)identifies the true modeβ∗\\beta^\{\*\}with posterior probability→1\\to 1;\(ii\)contracts the bounded predictive variance at any unobserved stepssastr\(Cov\(Gd,sα∣obs,β∗\)\)≤‖Gd,s‖F2/\(13\+kρmin\)=O\(1/k\)\\mathrm\{tr\}\(\\mathrm\{Cov\}\(G\_\{d,s\}\\alpha\\mid\\mathrm\{obs\},\\beta^\{\*\}\)\)\\leq\\\|G\_\{d,s\}\\\|\_\{F\}^\{2\}/\(\\tfrac\{1\}\{3\}\+k\\rho\_\{\\min\}\)=O\(1/k\), while the stochastic termtr\(Gs,sGs,s⊤\)\\mathrm\{tr\}\(G\_\{s,s\}G\_\{s,s\}^\{\\top\}\)is irreducible;\(iii\)an MDN updates only its mixing weights, leaving within\-component covariancesΘ\(1\)\\Theta\(1\)inkk\. Formal statement, proof, and discussion of the relaxation gap in Appendix[B\.6](https://arxiv.org/html/2608.05454#A2.SS6)\.
The relaxation is variance\-matched to the trainedUniform\(\[−1,1\]\)\\mathrm\{Uniform\}\(\[\-1,1\]\)prior; a 1\-D numerical check bounds the relative gap between relaxed and true posterior variances by≤9\.03%\\leq 9\.03\\%on the\(κ,k\)\(\\kappa,k\)grid we test \(App\.[B\.6](https://arxiv.org/html/2608.05454#A2.SS6)\)\. Observable contraction is controlled byκ=‖Gd‖/σ\\kappa=\\\|G\_\{d\}\\\|/\\sigma: ETH/UCY learnsκ≈2\.1\\kappa\\\!\\approx\\\!2\.1\(ceiling∼\\sim32%32\\%\); on nuScenes, sweeping theGdG\_\{d\}\-bias overκ∈\{1\.5,2\.0,2\.8,3\.7\}\\kappa\\in\\\{1\.5,2\.0,2\.8,3\.7\\\}gives spread reductions\{12\.3±3\.6,17\.7±4\.9,25\.8±4\.2,33\.2±5\.9\}%\\\{12\.3\{\\pm\}3\.6,\\,17\.7\{\\pm\}4\.9,\\,25\.8\{\\pm\}4\.2,\\,33\.2\{\\pm\}5\.9\\\}\\%\(5 seeds; per\-seed values and 3\-seed reference in Tab\.[21](https://arxiv.org/html/2608.05454#A4.T21)\)\. We instantiate on ETH/UCY withnb=1,nd=1n\_\{b\}\{=\}1,n\_\{d\}\{=\}1shared acrossT=12T\{=\}12steps, revealing positions0,…,80,\\ldots,8\.
Table 6:Sequential refinement on ETH/UCY \(same\-encoder, 5\-scene aggregate, 3 seeds\)\. HProbZ contracts spread23\.6%23\.6\\%and improves FDE16\.5%16\.5\\%; MDN’s FDE degrades\.Table[6](https://arxiv.org/html/2608.05454#S5.T6)shows simultaneous spread\-and\-mean refinement: HProbZ’s std drops23\.6%23\.6\\%and FDE improves16\.5%16\.5\\%, while MDN re\-weights toward a mis\-located mode \(FDE\+9\.6%\+9\.6\\%, std−3\.6%\-3\.6\\%\) — exactly Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)\(iii\)\.
### 5\.5Conformal Coverage with Multi\-Modal Sets
Split conformal atα=0\.1\\alpha\{=\}0\.1\(Theorem[6](https://arxiv.org/html/2608.05454#Thmtheorem6)\), comparing Standard CP \(singleL∞L\_\{\\infty\}box\), MDN\-CP \(KK\-component union\), and HProbZ\-CP \(2nb2^\{n\_\{b\}\}per\-mode union via Eq\. \([6](https://arxiv.org/html/2608.05454#S4.E6)\)\);D=24D\{=\}24\. All three reach∼90%\\sim 90\\%coverage; HProbZ\-CP achieves it with substantially smaller volumes on motion\-rich scenes \(Tab\.[7](https://arxiv.org/html/2608.05454#S5.T7)\)\.
Table 7:Conformal set volume comparison \(α=0\.1\\alpha\{=\}0\.1,D=24D\{=\}24, within\-scene calibration\)\. All methods reach∼90%\\sim 90\\%coverage; HProbZ\-CP is5\.55\.5/3\.93\.9orders smaller than StdCP/MDN\-CP on the ETH/UCY average and16\.316\.3/13\.913\.9orders smaller on Argoverse 2\. HOTEL is a counter\-example: MDN\-CP is3\.53\.5OOM smaller than HProbZ\-CP there \(App\.[D\.4](https://arxiv.org/html/2608.05454#A4.SS4)discusses the mechanism\)\. Bold marks the smallest volume per row\.Volume gaps compound from22–3×3\\timesper\-dimension edge ratios overD=24D\{=\}24\. The largest improvements are on open\-walkway scenes \(UNIV6\.56\.5, ZARA19\.29\.2, ZARA29\.59\.5OOM vs StdCP\) and Argoverse 2 \(16\.316\.3OOM\); the gap shrinks on ETH \(2\.02\.0\) and reverses on HOTEL where MDN\-CP wins by3\.53\.5OOM — HProbZ\-CP wins the average and 4/5 scenes\. OOD coverage holds:90\.9%90\.9\\%ETH/UCY leave\-one\-out average,89\.7%/99\.8%89\.7\\%/99\.8\\%IID/OOD on nuScenes speed split \(App\.[D\.5](https://arxiv.org/html/2608.05454#A4.SS5)\);α\\alpha\-sweep and protocol in App\.[D\.4](https://arxiv.org/html/2608.05454#A4.SS4)\.
## 6Related Work
#### Trajectory forecasting\.
Recurrent\[[1](https://arxiv.org/html/2608.05454#bib.bib34),[12](https://arxiv.org/html/2608.05454#bib.bib29)\], CVAE\[[27](https://arxiv.org/html/2608.05454#bib.bib22),[22](https://arxiv.org/html/2608.05454#bib.bib30),[21](https://arxiv.org/html/2608.05454#bib.bib31)\], transformer\[[42](https://arxiv.org/html/2608.05454#bib.bib32),[39](https://arxiv.org/html/2608.05454#bib.bib33),[31](https://arxiv.org/html/2608.05454#bib.bib35),[47](https://arxiv.org/html/2608.05454#bib.bib3),[43](https://arxiv.org/html/2608.05454#bib.bib44),[35](https://arxiv.org/html/2608.05454#bib.bib45)\], language\-modeling\[[30](https://arxiv.org/html/2608.05454#bib.bib2)\], diffusion\[[11](https://arxiv.org/html/2608.05454#bib.bib24),[23](https://arxiv.org/html/2608.05454#bib.bib41),[15](https://arxiv.org/html/2608.05454#bib.bib46)\], flow\[[29](https://arxiv.org/html/2608.05454#bib.bib5)\], equivariant\[[40](https://arxiv.org/html/2608.05454#bib.bib43)\], and refinement\[[46](https://arxiv.org/html/2608.05454#bib.bib6)\]forecasters emit unstructured samples or independent\-Gaussian mixtures without algebraic constraints across steps; HProbZ’s shared bounded factor adds such a constraint under a closed\-form output\.
#### Uncertainty decomposition and conformal prediction\.
Epistemic–aleatoric decomposition\[[16](https://arxiv.org/html/2608.05454#bib.bib28),[18](https://arxiv.org/html/2608.05454#bib.bib20)\]is orthogonal and coarser; Theorem[3](https://arxiv.org/html/2608.05454#Thmtheorem3)gives a likelihood\-identifiable within\-aleatoric one\. Split conformal\[[36](https://arxiv.org/html/2608.05454#bib.bib17),[26](https://arxiv.org/html/2608.05454#bib.bib18),[4](https://arxiv.org/html/2608.05454#bib.bib19),[3](https://arxiv.org/html/2608.05454#bib.bib38)\]and trajectory variants\[[19](https://arxiv.org/html/2608.05454#bib.bib4),[14](https://arxiv.org/html/2608.05454#bib.bib7),[45](https://arxiv.org/html/2608.05454#bib.bib47)\]return single convex regions; Eq\. \([6](https://arxiv.org/html/2608.05454#S4.E6)\) returns a union of2nb2^\{n\_\{b\}\}axis\-aligned boxes from the trained model\.
#### Zonotopes in machine learning\.
Zonotopes serve as inputs/intermediate sets for NN verification\[[32](https://arxiv.org/html/2608.05454#bib.bib25),[44](https://arxiv.org/html/2608.05454#bib.bib26),[24](https://arxiv.org/html/2608.05454#bib.bib27),[5](https://arxiv.org/html/2608.05454#bib.bib15),[17](https://arxiv.org/html/2608.05454#bib.bib37)\]with the network fixed; HProbZ inverts the role, making the zonotope the trained*output*via Eq\. \([4](https://arxiv.org/html/2608.05454#S4.E4)\)\.
## 7Conclusion
HProbZ is a structured output head that keeps modal, bounded, and stochastic uncertainty separately identifiable\. It approaches map\-aware MID’s published minADE@10 on nuScenes \(1\.29±0\.071\.29\\\!\\pm\\\!0\.07, 14 seeds\) and is on par with LED/EqMotion on ETH/UCY \(0\.1970\.197\)\. Structural gains: observation\-driven contraction up to33±6%33\{\\pm\}6\\%\(5 seeds\),30%30\\%fewer trajectory\-level closed\-loop collisions \(per\-step reverses\), and conformal sets3\.93\.9/13\.913\.9orders smaller than MDN\-CP on ETH/UCY/Argoverse 2\.
#### Limitations\.
*Cross\-dataset transfer is largely negative*: zero\-shot nuScenes→\\tonuPlan AUC0\.5190\.519on the full agent mix \(within22pp of chance,14,65314\{,\}653pairs\); only the close\-proximity subset preserves\+0\.159\+0\.159AUC \(App\.[D\.26](https://arxiv.org/html/2608.05454#A4.SS26)\)\.*Closed\-loop ordering is metric\-dependent*: the30%30\\%headline is trajectory\-level; per\-step reverses \(HProbZ8\.01%8\.01\\%vs MDN6\.32%6\.32\\%\)\.*The SOTA\-chase number depends on a method\-specific init*:1\.291\.29uses the symmetry\-breakingGbG\_\{b\}\-bias \(App\.[D\.24](https://arxiv.org/html/2608.05454#A4.SS24)\); without it, the same architecture trains to1\.48±0\.101\.48\{\\pm\}0\.10, above MID’s1\.441\.44\. Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)is proven under a Gaussian\-relaxed prior with a≤9\.03%\\leq 9\.03\\%numerical gap to the trained uniform prior\. Raw NLL exceeds MDN’s; the bared=64d\{=\}64head trails CVAE on ADE before social attention, and at vanilla encoders MDN\-K=20 wins minFDE@20 \(Tab\.[17](https://arxiv.org/html/2608.05454#A4.T17)\)\. Other caveats \(mode prior,Gd/GsG\_\{d\}/G\_\{s\}diagonality, baseline tuning, deployment\) in App\.[E](https://arxiv.org/html/2608.05454#A5)\.
## References
- \[1\]\(2016\)Social LSTM: human trajectory prediction in crowded spaces\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[2\]M\. Althoff, O\. Stursberg, and M\. Buss\(2007\)Safety assessment of autonomous cars using verification techniques\.In2007 American Control Conference,pp\. 4154–4159\.Cited by:[§2](https://arxiv.org/html/2608.05454#S2.p1.9),[§3](https://arxiv.org/html/2608.05454#S3.p1.8)\.
- \[3\]A\. N\. Angelopoulos and S\. Bates\(2023\)Conformal prediction: a gentle introduction\.Foundations and Trends in Machine Learning16\(4\),pp\. 494–591\.External Links:[Document](https://dx.doi.org/10.1561/2200000101)Cited by:[§B\.3](https://arxiv.org/html/2608.05454#A2.SS3.p1.5),[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px2.p1.1)\.
- \[4\]R\. F\. Barber, E\. J\. Candès, A\. Ramdas, and R\. J\. Tibshirani\(2021\)Predictive inference with the jackknife\+\.Annals of Statistics49\(1\),pp\. 486–507\.Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px2.p1.1)\.
- \[5\]T\. J\. Bird, H\. C\. Pangborn, N\. Jain, and J\. P\. Koeln\(2023\)Hybrid zonotopes: a new set representation for reachability analysis of mixed logical dynamical systems\.Automatica154,pp\. 111107\.External Links:[Document](https://dx.doi.org/10.1016/j.automatica.2023.111107)Cited by:[§2](https://arxiv.org/html/2608.05454#S2.p1.9),[§3](https://arxiv.org/html/2608.05454#S3.p1.8),[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px3.p1.1)\.
- \[6\]C\. M\. Bishop\(1994\)Mixture density networks\.Technical reportTechnical ReportNCRG/94/004,Aston University\.Cited by:[Table 1](https://arxiv.org/html/2608.05454#S1.T1.16.14.14.4),[§2](https://arxiv.org/html/2608.05454#S2.p1.9)\.
- \[7\]H\. Caesar, V\. Bankiti, A\. H\. Lang, S\. Vora, V\. E\. Liong, Q\. Xu, A\. Krishnan, Y\. Pan, G\. Baldan, and O\. Beijbom\(2020\)NuScenes: a multimodal dataset for autonomous driving\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,Cited by:[§A\.1](https://arxiv.org/html/2608.05454#A1.SS1.SSS0.Px2.p1.4),[§D\.26](https://arxiv.org/html/2608.05454#A4.SS26.p1.9),[§5\.3](https://arxiv.org/html/2608.05454#S5.SS3.p1.3)\.
- \[8\]H\. Cui, V\. Radosavljevic, F\. Chou, T\. Lin, T\. Nguyen, T\. Huang, J\. Schneider, and N\. Djuric\(2019\)Multimodal trajectory predictions for autonomous driving using deep convolutional networks\.InIEEE International Conference on Robotics and Automation \(ICRA\),pp\. 2090–2096\.Cited by:[§2](https://arxiv.org/html/2608.05454#S2.p1.9)\.
- \[9\]Y\. Gal and Z\. Ghahramani\(2016\)Dropout as a bayesian approximation: representing model uncertainty in deep learning\.InInternational Conference on Machine Learning,Cited by:[§D\.1](https://arxiv.org/html/2608.05454#A4.SS1.p1.8)\.
- \[10\]S\. Ghosal and A\. van der Vaart\(2017\)Fundamentals of nonparametric Bayesian inference\.Cambridge University Press\.Cited by:[§B\.6](https://arxiv.org/html/2608.05454#A2.SS6.p1.11)\.
- \[11\]T\. Gu, G\. Chen, J\. Li, C\. Lin, Y\. Rao, J\. Zhou, and J\. Lu\(2022\)Stochastic trajectory prediction via motion indeterminacy diffusion\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,Cited by:[Table 28](https://arxiv.org/html/2608.05454#A4.T28.8.7.6.1),[Table 33](https://arxiv.org/html/2608.05454#A4.T33.24.18.2),[Table 3](https://arxiv.org/html/2608.05454#S5.T3.26.17.5.1),[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[12\]A\. Gupta, J\. Johnson, L\. Fei\-Fei, S\. Savarese, and A\. Alahi\(2018\)Social GAN: socially acceptable trajectories with generative adversarial networks\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[13\]J\. Ho, A\. Jain, and P\. Abbeel\(2020\)Denoising diffusion probabilistic models\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 6840–6851\.Cited by:[§D\.11](https://arxiv.org/html/2608.05454#A4.SS11.p1.4)\.
- \[14\]H\. Huang, S\. He, and F\. Miao\(2025\)CUQDS: conformal uncertainty quantification under distribution shift for trajectory prediction\.InProceedings of the AAAI Conference on Artificial Intelligence \(Vol\. 39, No\. 16\),pp\. 17422–17430\.Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px2.p1.1)\.
- \[15\]C\. M\. Jiang, A\. Cornman, C\. Park, B\. Sapp, Y\. Zhou, and D\. Anguelov\(2023\)MotionDiffuser: controllable multi\-agent motion prediction using diffusion\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 9644–9653\.Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[16\]A\. Kendall and Y\. Gal\(2017\)What uncertainties do we need in Bayesian deep learning for computer vision?\.InAdvances in Neural Information Processing Systems,Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px2.p1.1)\.
- \[17\]N\. Kochdumper and M\. Althoff\(2023\)Constrained polynomial zonotopes\.Acta Informatica60\(3\),pp\. 279–316\.External Links:[Document](https://dx.doi.org/10.1007/s00236-023-00437-5)Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px3.p1.1)\.
- \[18\]B\. Lakshminarayanan, A\. Pritzel, and C\. Blundell\(2017\)Simple and scalable predictive uncertainty estimation using deep ensembles\.InAdvances in Neural Information Processing Systems,Cited by:[§D\.1](https://arxiv.org/html/2608.05454#A4.SS1.p1.8),[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px2.p1.1)\.
- \[19\]L\. Lindemann, M\. Cleaveland, G\. Shim, and G\. J\. Pappas\(2023\)Safe planning in dynamic environments using conformal prediction\.IEEE Robotics and Automation Letters8\(8\),pp\. 5116–5123\.External Links:[Document](https://dx.doi.org/10.1109/LRA.2023.3292071)Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px2.p1.1)\.
- \[20\]O\. Makansi, E\. Ilg, Ö\. Çiçek, and T\. Brox\(2019\)Overcoming limitations of mixture density networks: a sampling and fitting framework for multimodal future prediction\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 7144–7153\.Cited by:[§2](https://arxiv.org/html/2608.05454#S2.p1.9)\.
- \[21\]K\. Mangalam, Y\. An, H\. Girase, and J\. Malik\(2021\)From goals, waypoints & paths to long term human trajectory forecasting\.InProceedings of the IEEE/CVF International Conference on Computer Vision,Cited by:[Table 28](https://arxiv.org/html/2608.05454#A4.T28.8.5.4.1),[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[22\]K\. Mangalam, H\. Girase, S\. Agarwal, K\. Lee, E\. Adeli, J\. Malik, and A\. Gaidon\(2020\)It is not the journey but the destination: endpoint conditioned trajectory prediction\.InEuropean Conference on Computer Vision,Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[23\]W\. Mao, C\. Xu, Q\. Zhu, S\. Chen, and Y\. Wang\(2023\)Leapfrog diffusion model for stochastic trajectory prediction\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 5517–5526\.Cited by:[Table 28](https://arxiv.org/html/2608.05454#A4.T28.8.8.7.1),[Table 1](https://arxiv.org/html/2608.05454#S1.T1.6.4.4.5),[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[24\]M\. Mirman, T\. Gehr, and M\. Vechev\(2018\)Differentiable abstract interpretation for provably robust neural networks\.InInternational Conference on Machine Learning,Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px3.p1.1)\.
- \[25\]A\. Q\. Nichol and P\. Dhariwal\(2021\)Improved denoising diffusion probabilistic models\.InInternational Conference on Machine Learning,pp\. 8162–8171\.Cited by:[§D\.11](https://arxiv.org/html/2608.05454#A4.SS11.p1.4)\.
- \[26\]Y\. Romano, E\. Patterson, and E\. Candès\(2019\)Conformalized quantile regression\.InAdvances in Neural Information Processing Systems,Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px2.p1.1)\.
- \[27\]T\. Salzmann, B\. Ivanovic, P\. Chakravarty, and M\. Pavone\(2020\)Trajectron\+\+: dynamically\-feasible trajectory forecasting with heterogeneous data\.InEuropean Conference on Computer Vision,Cited by:[Table 28](https://arxiv.org/html/2608.05454#A4.T28.8.2.1.1),[Table 1](https://arxiv.org/html/2608.05454#S1.T1.13.11.11.4),[Table 3](https://arxiv.org/html/2608.05454#S5.T3.26.15.3.1),[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[28\]R\. Schneider\(2014\)Convex bodies: the Brunn–Minkowski theory\.2nd edition,Cambridge University Press\.Cited by:[Appendix B](https://arxiv.org/html/2608.05454#A2.1.p1.6)\.
- \[29\]C\. Schöller and A\. Knoll\(2021\)FloMo: tractable motion prediction with normalizing flows\.InIEEE/RSJ International Conference on Intelligent Robots and Systems \(IROS\),pp\. 7977–7984\.Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[30\]A\. Seff, B\. Cera, D\. Chen, M\. Ng, A\. Zhou, N\. Nayakanti, K\. S\. Refaat, R\. Al\-Rfou, and B\. Sapp\(2023\)MotionLM: multi\-agent motion forecasting as language modeling\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 8579–8590\.Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[31\]S\. Shi, L\. Jiang, D\. Dai, and B\. Schiele\(2022\)MTR: motion transformer with global intention localization and local movement refinement\.InAdvances in Neural Information Processing Systems,Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[32\]G\. Singh, T\. Gehr, M\. Mirman, M\. Püschel, and M\. Vechev\(2018\)Fast and effective robustness certification\.InAdvances in Neural Information Processing Systems,Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px3.p1.1)\.
- \[33\]K\. Sohn, H\. Lee, and X\. Yan\(2015\)Learning structured output representation using deep conditional generative models\.InAdvances in Neural Information Processing Systems,Cited by:[§D\.1](https://arxiv.org/html/2608.05454#A4.SS1.p1.8)\.
- \[34\]J\. Song, C\. Meng, and S\. Ermon\(2021\)Denoising diffusion implicit models\.InInternational Conference on Learning Representations,Cited by:[§D\.11](https://arxiv.org/html/2608.05454#A4.SS11.p1.4)\.
- \[35\]N\. Song, B\. Zhang, X\. Zhu, and L\. Zhang\(2024\)Motion forecasting in continuous driving\.InAdvances in Neural Information Processing Systems,Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[36\]V\. Vovk, A\. Gammerman, and G\. Shafer\(2005\)Algorithmic learning in a random world\.Springer\.Cited by:[§B\.3](https://arxiv.org/html/2608.05454#A2.SS3.p1.5),[§D\.4](https://arxiv.org/html/2608.05454#A4.SS4.p1.8),[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px2.p1.1)\.
- \[37\]B\. Wilson, W\. Qi, T\. Agarwal, J\. Lambert, J\. Singh, S\. Khandelwal, B\. Pan, R\. Kumar, A\. Hartnett, J\. K\. Pontes, D\. Ramanan, P\. Carr, and J\. Hays\(2021\)Argoverse 2: next generation datasets for self\-driving perception and forecasting\.InAdvances in Neural Information Processing Systems Datasets and Benchmarks Track \(Round 2\),Cited by:[§A\.1](https://arxiv.org/html/2608.05454#A1.SS1.SSS0.Px3.p1.2),[§5\.4](https://arxiv.org/html/2608.05454#S5.SS4.p1.1)\.
- \[38\]C\. Xu, M\. Li, Z\. Ni, Y\. Zhang, and S\. Chen\(2022\)GroupNet: multiscale hypergraph neural networks for trajectory prediction with relational reasoning\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 6498–6507\.Cited by:[Table 28](https://arxiv.org/html/2608.05454#A4.T28.8.6.5.1)\.
- \[39\]C\. Xu, W\. Mao, W\. Zhang, and S\. Chen\(2022\)Remember intentions: retrospective\-memory\-based trajectory prediction\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,Cited by:[Table 28](https://arxiv.org/html/2608.05454#A4.T28.8.4.3.1),[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[40\]C\. Xu, R\. T\. Tan, Y\. Tan, S\. Chen, Y\. G\. Wang, X\. Wang, and Y\. Wang\(2023\)EqMotion: equivariant multi\-agent motion prediction with invariant interaction reasoning\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 1410–1420\.Cited by:[Table 28](https://arxiv.org/html/2608.05454#A4.T28.8.9.8.1),[Table 1](https://arxiv.org/html/2608.05454#S1.T1.10.8.8.5),[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[41\]S\. J\. Yakowitz and J\. D\. Spragins\(1968\)On the identifiability of finite mixtures\.The Annals of Mathematical Statistics39\(1\),pp\. 209–214\.Cited by:[§B\.2](https://arxiv.org/html/2608.05454#A2.SS2.p3.2),[§3](https://arxiv.org/html/2608.05454#S3.2.p1.2)\.
- \[42\]Y\. Yuan, X\. Weng, Y\. Ou, and K\. M\. Kitani\(2021\)AgentFormer: agent\-aware transformers for socio\-temporal multi\-agent forecasting\.InProceedings of the IEEE/CVF International Conference on Computer Vision,Cited by:[Table 28](https://arxiv.org/html/2608.05454#A4.T28.8.3.2.1),[Table 3](https://arxiv.org/html/2608.05454#S5.T3.26.16.4.1),[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[43\]B\. Zhang, N\. Song, and L\. Zhang\(2024\)DeMo: decoupling motion forecasting into directional intentions and dynamic states\.InAdvances in Neural Information Processing Systems,Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[44\]H\. Zhang, T\. Weng, P\. Chen, C\. Hsieh, and L\. Daniel\(2018\)Efficient neural network robustness certification with general activation functions\.InAdvances in Neural Information Processing Systems,Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px3.p1.1)\.
- \[45\]Y\. Zhou, L\. Lindemann, and M\. Sesia\(2024\)Conformalized adaptive forecasting of heterogeneous trajectories\.InProceedings of the 41st International Conference on Machine Learning,pp\. 62002–62056\.Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px2.p1.1)\.
- \[46\]Y\. Zhou, H\. Shao, L\. Wang, S\. L\. Waslander, H\. Li, and Y\. Liu\(2024\)SmartRefine: a scenario\-adaptive refinement framework for efficient motion prediction\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 15281–15290\.Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
- \[47\]Z\. Zhou, J\. Wang, Y\. Li, and Y\. Huang\(2023\)Query\-centric trajectory prediction\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 17863–17873\.Cited by:[§6](https://arxiv.org/html/2608.05454#S6.SS0.SSS0.Px1.p1.1)\.
Appendix outline\.AImplementation Details and Reproducibility \(datasets, hyperparameters, compute, per\-experiment protocol\)\.BFull Proofs \(Theorems 1–5, Propositions 1–3\)\.CEmpirical Theorem Validation \(per\-value sweep tables; CF fingerprint; parameter recovery; entropy\-gap bound; raw\-NLL note\)\.DAdditional Experiments grouped as: D\.1 Ablations and Hyperparameter Sweeps; D\.2 Calibration Extensions; D\.3 Downstream Task Details; D\.4 Baseline Comparisons \(Extended\); D\.5 Per\-Scene, Per\-Modality, Per\-Seed Analyses; D\.6 Visualisations\.EBroader Impact and Ethical Considerations\.
## Appendix AImplementation Details and Reproducibility
### A\.1Datasets
#### ETH/UCY pedestrian\.
Five scenes \(ETH, HOTEL, UNIV, ZARA1, ZARA2\) recorded at2\.52\.5Hz, evaluated under the standard leave\-one\-out protocol: train on 4 scenes, test on the held\-out scene\. Trajectories are observed for88frames \(3\.23\.2s\) and predicted for1212frames \(4\.84\.8s\)\. Coordinates are in metres, world frame\. Source: original ETH/UCY recordings released for research use\.
#### nuScenes vehicle\.
1,000 driving scenes from Boston and Singapore, 40 k test windows of vehicle trajectories at22Hz\[[7](https://arxiv.org/html/2608.05454#bib.bib39)\]\. Past horizon11s, future horizon66s \(Tf=12T\_\{f\}\{=\}12\)\. Ego\-centric coordinate frame\. License: nuScenes non\-commercial research use\.
#### Argoverse 2 motion forecasting\.
200 k training scenarios, 25 k validation, 24 k test \(97 k test windows\) at1010Hz\[[37](https://arxiv.org/html/2608.05454#bib.bib13)\]; we subsample to22Hz for matched comparison\. License: CC\-BY\-NC\-SA 4\.0\.
#### nuPlan \(zero\-shot evaluation only\)\.
The nuPlan mini split \(Boston/Pittsburgh/Singapore/Las Vegas logs\) is used only for cross\-dataset collision\-warning evaluation in §[D\.26](https://arxiv.org/html/2608.05454#A4.SS26); no training\. License: nuPlan non\-commercial research use\.
### A\.2HProbZ Architecture and Training
Table 8:HProbZ master hyperparameter table\. ETH/UCY entry covers both single\-agent \(d=64d\{=\}64\) and Social \(d=128d\{=\}128, 8\-neighbour attention\) variants\. Optimiser is AdamW with cosine LR decay; warmup is linear over the first1010epochs\. “Sym\-breakb0b\_\{0\}” is the positive bias added to theGbG\_\{b\}output channels at init \(Appendix[D\.24](https://arxiv.org/html/2608.05454#A4.SS24)\)\.
### A\.3Baseline Architectures and Hyperparameters
MDN\.Linear\(d,dd,d\)→\\toGELU→\\toLinear\(d,Tf⋅K⋅5d,T\_\{f\}\\cdot K\\cdot 5\)producing per\-component means, variances, and mixing logits\. Trained with the same encoder, optimiser, batch, and schedule as the matched HProbZ row\. Full hyperparameter sweep in Appendix[D\.22](https://arxiv.org/html/2608.05454#A4.SS22); the pathological MDN\-K=8 collapse sweep is in Appendix[D\.27](https://arxiv.org/html/2608.05454#A4.SS27)\.
CVAE\.32\-dim latent with shared\-encoder amortized posterior; KL warmup over2020epochs\. Used in Table[2](https://arxiv.org/html/2608.05454#S5.T2)and as a baseline in the controlled experiments of Appendix[D\.1](https://arxiv.org/html/2608.05454#A4.SS1)\.
Diffusion \(DDPM\)\.Temporal transformer denoiser with adaLN\-style timestep conditioning,44blocks, cosineT=200T\{=\}200schedule, DDIM5050\-step sampling,η=0\.5\\eta\{=\}0\.5; details and step\-count sweep in Appendix[D\.11](https://arxiv.org/html/2608.05454#A4.SS11)\.
### A\.4Per\-Experiment Protocol Index
Every sub\-experiment in §[5](https://arxiv.org/html/2608.05454#S5)is a matched\-encoder HProbZ vs\. MDN comparison; Table[9](https://arxiv.org/html/2608.05454#A1.T9)lists the controlled variables in each\.
Table 9:Summary of experimental protocols\. Within each row, HProbZ and the MDN baseline share the same encoder architecture, training schedule, and sampling budget\.
### A\.5Compute Resources
All training experiments use single H100/H200\-class GPUs; small\-scale ablations and runtime measurements use a consumer GPU \(RTX\-class\)\. Singled=256d\{=\}256HProbZ on nuScenes:∼18\\sim 18minutes per seed end\-to\-end \(150 epochs, batch 1024, 162k training windows\); the 14\-seed sweep used≈252\\approx 252GPU\-minutes \(≈4\.2\\approx 4\.2GPU\-hours\)\. Argoverse 2 and ETH/UCY runs are smaller; the 600\-epoch Social re\-run used a 4\-GPU cluster for the 5\-seed×\\times5\-scene grid\. Inference benchmarks \(throughput, refinement runtime\) are measured on a single consumer GPU\. Code, per\-seed checkpoints, and per\-experiment compute logs are released to avoid repeated retraining\.
### A\.6Algorithmic Specification
We provide pseudocode for the two HProbZ\-specific procedures\. Algorithm[1](https://arxiv.org/html/2608.05454#alg1)is the per\-batch training step \(gradient through the closed\-form convolution likelihood of Eq\. \([4](https://arxiv.org/html/2608.05454#S4.E4)\)\); Algorithm[2](https://arxiv.org/html/2608.05454#alg2)is the constrained inference procedure that activates the algebraic constraint machinery of Definition[1](https://arxiv.org/html/2608.05454#Thmtheorem1)whenkkground\-truth positions become available\.*Refinement is performed only at inference; training \(Algorithm[1](https://arxiv.org/html/2608.05454#alg1)\) sees no constraint rows \(p=0p\{=\}0\)\.*The Gaussian relaxation ofα\\alphain Algorithm[2](https://arxiv.org/html/2608.05454#alg2)is therefore an inference\-time approximation to the conditional posterior over the bounded latent, used solely to obtain a closed\-form posterior covariance and not present in the training objective\.
Algorithm 1HProbZ training step \(one batch\)\.1:batch
\{\(xi,yi\)\}i=1B\\\{\(x\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{B\}, encoder
Encϕ\\mathrm\{Enc\}\_\{\\phi\}, 2\-layer MLP head
HeadW\(h\)=W2GELU\(W1h\)\\mathrm\{Head\}\_\{W\}\(h\)=W\_\{2\}\\,\\mathrm\{GELU\}\(W\_\{1\}h\)producing
\(c,Gb,loga,logσ\)\(c,G\_\{b\},\\log a,\\log\\sigma\)
2:initial bias on head output:
GbG\_\{b\}entries
=b0=b\_\{0\}\(symmetry\-break, benchmark\-specific;
b0=0\.05b\_\{0\}\{=\}0\.05for ETH/UCY,
b0=1\.0b\_\{0\}\{=\}1\.0for nuScenes
d=256d\{=\}256, see Tab\.[8](https://arxiv.org/html/2608.05454#A1.T8)and App\.[D\.24](https://arxiv.org/html/2608.05454#A4.SS24)\),
loga0=0\.7\\log a\_\{0\}=0\.7,
logσ0=1\.1\\log\\sigma\_\{0\}=1\.1; weight init
𝒩\(0,0\.012\)\\mathcal\{N\}\(0,0\.01^\{2\}\)
3:loss flag
ℒ∈\{exact,Gauss\}\\mathcal\{L\}\\in\\\{\\text\{exact\},\\text\{Gauss\}\\\}, gradient\-clip threshold
τ=1\.0\\tau=1\.0
4:for
i=1,…,Bi=1,\\ldots,Bdo
5:
hi←Encϕ\(xi\)h\_\{i\}\\leftarrow\\mathrm\{Enc\}\_\{\\phi\}\(x\_\{i\}\)
6:
\(ci,Gb,i,logai,logσi\)←HeadW\(hi\)\(c\_\{i\},\\,G\_\{b,i\},\\,\\log a\_\{i\},\\,\\log\\sigma\_\{i\}\)\\leftarrow\\mathrm\{Head\}\_\{W\}\(h\_\{i\}\)
7:
ai←exp\(clamp\(logai,−4,4\)\);σi←exp\(clamp\(logσi,−4,4\)\)a\_\{i\}\\leftarrow\\exp\(\\mathrm\{clamp\}\(\\log a\_\{i\},\\,\-4,\\,4\)\);\\ \\sigma\_\{i\}\\leftarrow\\exp\(\\mathrm\{clamp\}\(\\log\\sigma\_\{i\},\\,\-4,\\,4\)\)
8:foreach binary
β∈\{−1,1\}nb\\beta\\in\\\{\-1,1\\\}^\{n\_\{b\}\}do
9:
μi,β←ci\+Gb,iβ\\mu\_\{i,\\beta\}\\leftarrow c\_\{i\}\+G\_\{b,i\}\\beta
10:endfor
11:if
ℒ=exact\\mathcal\{L\}=\\text\{exact\}then
12:
ℓi←−log∑β2−nb∏jf\(yi,j−μi,β,j;ai,j,σi,j\)\\ell\_\{i\}\\leftarrow\-\\log\\sum\_\{\\beta\}2^\{\-n\_\{b\}\}\\prod\_\{j\}f\(y\_\{i,j\}\-\\mu\_\{i,\\beta,j\};\\,a\_\{i,j\},\\sigma\_\{i,j\}\)⊳\\trianglerightfffrom Eq\. \([4](https://arxiv.org/html/2608.05454#S4.E4)\)
13:else
14:
ℓi←−log∑β2−nb𝒩\(yi;μi,β,σi2\+ai2/3\)\\ell\_\{i\}\\leftarrow\-\\log\\sum\_\{\\beta\}2^\{\-n\_\{b\}\}\\,\\mathcal\{N\}\(y\_\{i\};\\,\\mu\_\{i,\\beta\},\\,\\sigma\_\{i\}^\{2\}\+a\_\{i\}^\{2\}/3\)⊳\\trianglerightGaussian surrogate
15:endif
16:endfor
17:
g←∇ϕ,W1B∑iℓig\\leftarrow\\nabla\_\{\\phi,W\}\\,\\tfrac\{1\}\{B\}\\sum\_\{i\}\\ell\_\{i\}
18:Clip:
g←g⋅min\(1,τ/‖g‖2\)g\\leftarrow g\\cdot\\min\\\!\\big\(1,\\ \\tau/\\\|g\\\|\_\{2\}\\big\)
19:Update
ϕ,W\\phi,Wby AdamW\(
gg\)\.
Algorithm 2HProbZ constrained inference \(refinement afterkkrevealed positions\)\.1:trained head outputs at all steps
\{\(ct,Gb,t,Gd,t,Gs,t\)\}t=1T\\\{\(c\_\{t\},G\_\{b,t\},G\_\{d,t\},G\_\{s,t\}\)\\\}\_\{t=1\}^\{T\}, observed positions
\{zti\}i=1k\\\{z\_\{t\_\{i\}\}\\\}\_\{i=1\}^\{k\}
2:Gaussian relaxation
α∼𝒩\(0,13Ind\)\\alpha\\sim\\mathcal\{N\}\(0,\\tfrac\{1\}\{3\}I\_\{n\_\{d\}\}\)
3:foreach binary
β∈\{−1,1\}nb\\beta\\in\\\{\-1,1\\\}^\{n\_\{b\}\}do
4:
rti←zti−cti−Gb,tiβr\_\{t\_\{i\}\}\\leftarrow z\_\{t\_\{i\}\}\-c\_\{t\_\{i\}\}\-G\_\{b,t\_\{i\}\}\\betafor
i=1,…,ki=1,\\ldots,k
5:
Σs,ti←Gs,tiGs,ti⊤\\Sigma\_\{s,t\_\{i\}\}\\leftarrow G\_\{s,t\_\{i\}\}G\_\{s,t\_\{i\}\}^\{\\top\}
6:
Ik←∑i=1kGd,ti⊤Σs,ti−1Gd,tiI\_\{k\}\\leftarrow\\sum\_\{i=1\}^\{k\}G\_\{d,t\_\{i\}\}^\{\\top\}\\Sigma\_\{s,t\_\{i\}\}^\{\-1\}G\_\{d,t\_\{i\}\}⊳\\trianglerightFisher information
7:
Σpost\(β\)←\(3Ind\+Ik\)−1\\Sigma\_\{\\mathrm\{post\}\}\(\\beta\)\\leftarrow\(3I\_\{n\_\{d\}\}\+I\_\{k\}\)^\{\-1\}
8:
μpost\(β\)←Σpost\(β\)∑i=1kGd,ti⊤Σs,ti−1rti\\mu\_\{\\mathrm\{post\}\}\(\\beta\)\\leftarrow\\Sigma\_\{\\mathrm\{post\}\}\(\\beta\)\\,\\sum\_\{i=1\}^\{k\}G\_\{d,t\_\{i\}\}^\{\\top\}\\Sigma\_\{s,t\_\{i\}\}^\{\-1\}r\_\{t\_\{i\}\}
9:
w\(β\)←2−nb∏i=1kp\(zti∣β\)w\(\\beta\)\\leftarrow 2^\{\-n\_\{b\}\}\\prod\_\{i=1\}^\{k\}p\(z\_\{t\_\{i\}\}\\mid\\beta\)⊳\\trianglerightmode posterior weight
10:endfor
11:Normalize
\{w\(β\)\}\\\{w\(\\beta\)\\\}to a posterior over
β\\beta\.
12:foreach unobserved step
ssdo
13:Tightened mean:
z^s\(β\)←cs\+Gb,sβ\+Gd,sμpost\(β\)\\hat\{z\}\_\{s\}\(\\beta\)\\leftarrow c\_\{s\}\+G\_\{b,s\}\\beta\+G\_\{d,s\}\\,\\mu\_\{\\mathrm\{post\}\}\(\\beta\)
14:Tightened covariance:
Σ^s\(β\)←Gd,sΣpost\(β\)Gd,s⊤\+Σs,s\\widehat\{\\Sigma\}\_\{s\}\(\\beta\)\\leftarrow G\_\{d,s\}\\,\\Sigma\_\{\\mathrm\{post\}\}\(\\beta\)\\,G\_\{d,s\}^\{\\top\}\+\\Sigma\_\{s,s\}
15:endfor
16:return
\{w\(β\),z^s\(β\),Σ^s\(β\)\}\\\{w\(\\beta\),\\hat\{z\}\_\{s\}\(\\beta\),\\widehat\{\\Sigma\}\_\{s\}\(\\beta\)\\\}for all
β\\betaand unobserved
ss\.
Total cost of Algorithm[2](https://arxiv.org/html/2608.05454#alg2)isO\(2nb\(knd2\+nd3\)\)O\(2^\{n\_\{b\}\}\(k\\,n\_\{d\}^\{2\}\+n\_\{d\}^\{3\}\)\)per scene; this is dominated by thend×ndn\_\{d\}\\times n\_\{d\}posterior solve and is negligible compared to a single encoder forward pass for thenb,nd∈\{1,2,3\}n\_\{b\},n\_\{d\}\\in\\\{1,2,3\\\}regime used in this paper\. The Fisher\-information form here is the formal statement; in the body we wrote the simplifiedρmin\\rho\_\{\\min\}bound \(Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7), Appendix[B\.6](https://arxiv.org/html/2608.05454#A2.SS6)\)\.
## Appendix BFull Proofs
###### Proposition 8\(Minkowski decomposition of the HProbZ measure\)\.
Letμ𝒵\\mu\_\{\\mathcal\{Z\}\}denote the distribution ofy=c\+Gbβ\+Gdα\+Gsνy=c\+G\_\{b\}\\beta\+G\_\{d\}\\alpha\+G\_\{s\}\\nuwhereβ∼Uniform\(\{−1,1\}nb\)\\beta\\sim\\mathrm\{Uniform\}\(\\\{\-1,1\\\}^\{n\_\{b\}\}\),α∼Uniform\(\[−1,1\]nd\)\\alpha\\sim\\mathrm\{Uniform\}\(\[\-1,1\]^\{n\_\{d\}\}\), andν∼𝒩\(0,Inq\)\\nu\\sim\\mathcal\{N\}\(0,I\_\{n\_\{q\}\}\)are independent in the unconstrained casep=0p\{=\}0\. Thenμ𝒵=δc∗\(12nb∑βδGbβ\)∗\(Gd\)∗𝒰\(\[−1,1\]nd\)∗𝒩\(0,GsGs⊤\)\\mu\_\{\\mathcal\{Z\}\}=\\delta\_\{c\}\\ast\(\\frac\{1\}\{2^\{n\_\{b\}\}\}\\sum\_\{\\beta\}\\delta\_\{G\_\{b\}\\beta\}\)\\ast\(G\_\{d\}\)\_\{\*\}\\mathcal\{U\}\(\[\-1,1\]^\{n\_\{d\}\}\)\\ast\\mathcal\{N\}\(0,G\_\{s\}G\_\{s\}^\{\\top\}\)\.
###### Proof\.
Independence ofβ,α,ν\\beta,\\alpha,\\nuimplies the distribution of the sum is the convolution of the marginals\. The distribution ofGbβG\_\{b\}\\betais the uniform measure on the finite setGb\{−1,1\}nbG\_\{b\}\\\{\-1,1\\\}^\{n\_\{b\}\}; that ofGdαG\_\{d\}\\alphais the pushforward of the uniform measure on the cube underGdG\_\{d\}, i\.e\., the zonoid measure\[[28](https://arxiv.org/html/2608.05454#bib.bib36)\]; andGsνG\_\{s\}\\nucontributes the Gaussian factor\. ∎
### B\.1Proof of Theorem[2](https://arxiv.org/html/2608.05454#Thmtheorem2)
Setup\.The characteristic function ofμ𝒵\\mu\_\{\\mathcal\{Z\}\}att∈ℝdt\\in\\mathbb\{R\}^\{d\}factorises as
ϕ𝒵\(t\)=eic⊤t⋅ϕB\(t\)⋅∏m=1ndsinc\(Gd\[:,m\]⊤t\)⋅exp\(−12t⊤GsGs⊤t\),\\phi\_\{\\mathcal\{Z\}\}\(t\)=e^\{ic^\{\\top\}t\}\\cdot\\phi\_\{B\}\(t\)\\cdot\\prod\_\{m=1\}^\{n\_\{d\}\}\\mathrm\{sinc\}\(G\_\{d\}\[:,m\]^\{\\top\}t\)\\cdot\\exp\\\!\\left\(\-\\tfrac\{1\}\{2\}t^\{\\top\}G\_\{s\}G\_\{s\}^\{\\top\}t\\right\),whereϕB\(t\)=2−nb∑βei\(Gbβ\)⊤t\\phi\_\{B\}\(t\)=2^\{\-n\_\{b\}\}\\sum\_\{\\beta\}e^\{i\(G\_\{b\}\\beta\)^\{\\top\}t\}is bounded in magnitude by11\. Suppose for contradiction that there is a finite Gaussian mixture with CFϕGMM\(t\)=∑k=1Kπkeiμk⊤t−12t⊤Σkt\\phi\_\{\\text\{GMM\}\}\(t\)=\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}e^\{i\\mu\_\{k\}^\{\\top\}t\-\\tfrac\{1\}\{2\}t^\{\\top\}\\Sigma\_\{k\}t\}such thatϕ𝒵\(t\)=ϕGMM\(t\)\\phi\_\{\\mathcal\{Z\}\}\(t\)=\\phi\_\{\\text\{GMM\}\}\(t\)for allt∈ℝdt\\in\\mathbb\{R\}^\{d\}\.
Reduction to 1D\.SinceGdG\_\{d\}has a nonzero entry, choose any unit vectoru∈ℝdu\\in\\mathbb\{R\}^\{d\}witha:=\|Gd\[:,m∗\]⊤u\|\>0a:=\|G\_\{d\}\[:,m^\{\*\}\]^\{\\top\}u\|\>0for somem∗m^\{\*\}\. By Cramér–Wold, the two distributions project to the same 1\-D law alonguu, so the 1\-D CFs match:ϕ𝒵,u\(s\)=ϕGMM,u\(s\)\\phi\_\{\\mathcal\{Z\},u\}\(s\)=\\phi\_\{\\text\{GMM\},u\}\(s\)for alls∈ℝs\\in\\mathbb\{R\}, where
ϕ𝒵,u\(s\)\\displaystyle\\phi\_\{\\mathcal\{Z\},u\}\(s\)=eicus⋅ϕB,u\(s\)⋅∏m:Gd\[:,m\]⊤u≠0sinc\(Gd\[:,m\]⊤u⋅s\)⋅e−12σu2s2,\\displaystyle=e^\{ic\_\{u\}s\}\\cdot\\phi\_\{B,u\}\(s\)\\cdot\\prod\_\{m:G\_\{d\}\[:,m\]^\{\\top\}u\\neq 0\}\\mathrm\{sinc\}\(G\_\{d\}\[:,m\]^\{\\top\}u\\cdot s\)\\cdot e^\{\-\\tfrac\{1\}\{2\}\\sigma\_\{u\}^\{2\}s^\{2\}\},ϕGMM,u\(s\)\\displaystyle\\phi\_\{\\text\{GMM\},u\}\(s\)=∑k=1Kπkeiμk,us−12σk,u2s2,\\displaystyle=\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}e^\{i\\mu\_\{k,u\}s\-\\tfrac\{1\}\{2\}\\sigma\_\{k,u\}^\{2\}s^\{2\}\},withσu2=u⊤GsGs⊤u\\sigma\_\{u\}^\{2\}=u^\{\\top\}G\_\{s\}G\_\{s\}^\{\\top\}u,σk,u2=u⊤Σku\>0\\sigma\_\{k,u\}^\{2\}=u^\{\\top\}\\Sigma\_\{k\}u\>0,cu=u⊤cc\_\{u\}=u^\{\\top\}c,μk,u=u⊤μk\\mu\_\{k,u\}=u^\{\\top\}\\mu\_\{k\}, andϕB,u\(s\)=2−nb∑βei\(Gbβ\)⊤us\\phi\_\{B,u\}\(s\)=2^\{\-n\_\{b\}\}\\sum\_\{\\beta\}e^\{i\(G\_\{b\}\\beta\)^\{\\top\}u\\,s\}a trigonometric polynomial \(bounded by11onℝ\\mathbb\{R\}\)\.
Analytic continuation to the imaginary axis\.Bothϕ𝒵,u\\phi\_\{\\mathcal\{Z\},u\}andϕGMM,u\\phi\_\{\\text\{GMM\},u\}extend to entire functions onℂ\\mathbb\{C\}\(the 1\-D formulas above define holomorphic extensions\)\. Evaluating ats=iys=iywith realy\>0y\>0:
sinc\(α⋅iy\)=sin\(iαy\)iαy=sinh\(αy\)αy,e−12σu2\(iy\)2=e12σu2y2,\\mathrm\{sinc\}\(\\alpha\\cdot iy\)=\\frac\{\\sin\(i\\alpha y\)\}\{i\\alpha y\}=\\frac\{\\sinh\(\\alpha y\)\}\{\\alpha y\},\\quad e^\{\-\\tfrac\{1\}\{2\}\\sigma\_\{u\}^\{2\}\(iy\)^\{2\}\}=e^\{\\tfrac\{1\}\{2\}\\sigma\_\{u\}^\{2\}y^\{2\}\},andϕB,u\(iy\)=2−nb∑βe−\(Gbβ\)⊤uy\\phi\_\{B,u\}\(iy\)=2^\{\-n\_\{b\}\}\\sum\_\{\\beta\}e^\{\-\(G\_\{b\}\\beta\)^\{\\top\}u\\,y\}is a finite sum of positive real exponentials\. Chooseuugeneric so that exactly oneβ⋆∈\{−1,1\}nb\\beta^\{\\star\}\\in\\\{\-1,1\\\}^\{n\_\{b\}\}achievesγu:=maxβ\{−\(Gbβ\)⊤u\}\\gamma\_\{u\}:=\\max\_\{\\beta\}\\\{\-\(G\_\{b\}\\beta\)^\{\\top\}u\\\}\(this fails only on a measure\-zero set of directions in the unit sphere\)\. ThenϕB,u\(iy\)∼2−nbeγuy\\phi\_\{B,u\}\(iy\)\\sim 2^\{\-n\_\{b\}\}\\,e^\{\\gamma\_\{u\}y\}asy→∞y\\to\\inftywith no polynomial pre\-factor; ties at the maximizer would contribute a constant multiplicity but noyy\-dependent factor, and the contradiction below is unaffected\.
Asymptotic magnitude on the imaginary axis\.LetSm=\|Gd\[:,m\]⊤u\|S\_\{m\}=\|G\_\{d\}\[:,m\]^\{\\top\}u\|formmwithSm\>0S\_\{m\}\>0, andA=∑mSmA=\\sum\_\{m\}S\_\{m\}\. Fory→∞y\\to\\infty,sinh\(Smy\)/\(Smy\)=eSmy/\(2Smy\)\(1\+O\(e−2Smy\)\)\\sinh\(S\_\{m\}y\)/\(S\_\{m\}y\)=e^\{S\_\{m\}y\}/\(2S\_\{m\}y\)\(1\+O\(e^\{\-2S\_\{m\}y\}\)\), so
\|ϕ𝒵,u\(iy\)\|=Cynd′⋅e\(A\+γu\)y\+12σu2y2⋅\(1\+o\(1\)\)\|\\phi\_\{\\mathcal\{Z\},u\}\(iy\)\|\\;=\\;\\frac\{C\}\{y^\{n\_\{d\}^\{\\prime\}\}\}\\cdot e^\{\(A\+\\gamma\_\{u\}\)y\+\\tfrac\{1\}\{2\}\\sigma\_\{u\}^\{2\}y^\{2\}\}\\cdot\(1\+o\(1\)\)wherend′n\_\{d\}^\{\\prime\}is the number of directions withGd\[:,m\]⊤u≠0G\_\{d\}\[:,m\]^\{\\top\}u\\neq 0andC\>0C\>0is constant \(independent ofyy\)\. The key feature is the algebraic factory−nd′y^\{\-n\_\{d\}^\{\\prime\}\},nd′≥1n\_\{d\}^\{\\prime\}\\geq 1\.
For the Gaussian mixture ats=iys=iy:
\|ϕGMM,u\(iy\)\|=\|∑k=1Kπkeμk,uy\+12σk,u2y2\|\.\|\\phi\_\{\\text\{GMM\},u\}\(iy\)\|=\\left\|\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}e^\{\\mu\_\{k,u\}y\+\\tfrac\{1\}\{2\}\\sigma\_\{k,u\}^\{2\}y^\{2\}\}\\right\|\.Asy→∞y\\to\\infty, this sum is dominated by the term\(s\) with maximum exponent12σk,u2y2\+μk,uy\\tfrac\{1\}\{2\}\\sigma\_\{k,u\}^\{2\}y^\{2\}\+\\mu\_\{k,u\}y\. Letσ⋆=maxkσk,u\\sigma^\{\\star\}=\\max\_\{k\}\\sigma\_\{k,u\}andμ⋆=max\{μk,u:σk,u=σ⋆\}\\mu^\{\\star\}=\\max\\\{\\mu\_\{k,u\}:\\sigma\_\{k,u\}=\\sigma^\{\\star\}\\\}; then
\|ϕGMM,u\(iy\)\|=D⋅eμ⋆y\+12\(σ⋆\)2y2⋅\(1\+o\(1\)\)\|\\phi\_\{\\text\{GMM\},u\}\(iy\)\|\\;=\\;D\\cdot e^\{\\mu^\{\\star\}y\+\\tfrac\{1\}\{2\}\(\\sigma^\{\\star\}\)^\{2\}y^\{2\}\}\\cdot\(1\+o\(1\)\)whereD=∑k:\(σk,u,μk,u\)=\(σ⋆,μ⋆\)πk\>0D=\\sum\_\{k:\\,\(\\sigma\_\{k,u\},\\mu\_\{k,u\}\)=\(\\sigma^\{\\star\},\\mu^\{\\star\}\)\}\\pi\_\{k\}\>0is the sum of weights of all components attaining the dominant exponent\. Since everyπk\>0\\pi\_\{k\}\>0by the GMM definition and at least one component attains the maximum,D≥πk⋆\>0D\\geq\\pi\_\{k^\{\\star\}\}\>0; in particularDDis bounded away from zero independently ofyy, so noyy\-dependent cancellation can collapse the leading constant\.
Contradiction\.The equalityϕ𝒵,u\(iy\)=ϕGMM,u\(iy\)\\phi\_\{\\mathcal\{Z\},u\}\(iy\)=\\phi\_\{\\text\{GMM\},u\}\(iy\)forces the asymptotic magnitudes to agree\. Matching they2y^\{2\}andyycoefficients in the exponent yieldsσu2=\(σ⋆\)2\\sigma\_\{u\}^\{2\}=\(\\sigma^\{\\star\}\)^\{2\}andA\+γu=μ⋆A\+\\gamma\_\{u\}=\\mu^\{\\star\}\. After dividing both sides bye\(A\+γu\)y\+12σu2y2e^\{\(A\+\\gamma\_\{u\}\)y\+\\tfrac\{1\}\{2\}\\sigma\_\{u\}^\{2\}y^\{2\}\}, the LHS tends to0likey−nd′y^\{\-n\_\{d\}^\{\\prime\}\}\(algebraically\), while the RHS tends to the strictly positive constantDD\. This contradictsnd′≥1n\_\{d\}^\{\\prime\}\\geq 1, soϕ𝒵≠ϕGMM\\phi\_\{\\mathcal\{Z\}\}\\neq\\phi\_\{\\text\{GMM\}\}\. By Lévy’s inversion theorem \(uniqueness of characteristic functions\),μ𝒵≠μGMM\\mu\_\{\\mathcal\{Z\}\}\\neq\\mu\_\{\\text\{GMM\}\}\.∎
#### Remark on why the “entire\-vanishing” shortcut fails\.
The sinc zeros ofϕ𝒵\\phi\_\{\\mathcal\{Z\}\}form an unbounded discrete set with no finite accumulation point, so the identity theorem for holomorphic functions does not forceϕGMM\\phi\_\{\\text\{GMM\}\}to vanish identically from agreement on that set alone \(a counter\-intuition:cost\\cos tis entire, has unbounded zero set\{\(k\+12\)π\}\\\{\(k\+\\tfrac\{1\}\{2\}\)\\pi\\\}, andcos0=1\\cos 0=1\)\. The above growth\-rate argument avoids this subtlety by separating algebraic \(y−nd′y^\{\-n\_\{d\}^\{\\prime\}\}\) and exponential contributions along the imaginary axis\.
### B\.2Proof of Theorem[3](https://arxiv.org/html/2608.05454#Thmtheorem3)
Projectingμ𝒵\\mu\_\{\\mathcal\{Z\}\}onto axisjjgives a 1\-D mixture of uniform\-normal convolutions sharing shape\(aj,σj\)\(a\_\{j\},\\sigma\_\{j\}\)with distinct locations\.
Shape recovery\.The 1\-D CF isϕj\(t\)=eicjt⋅2−nb∑βei\(Gbβ\)jt⋅sinc\(ajt\)⋅e−σj2t2/2\\phi\_\{j\}\(t\)=e^\{ic\_\{j\}t\}\\cdot 2^\{\-n\_\{b\}\}\\sum\_\{\\beta\}e^\{i\(G\_\{b\}\\beta\)\_\{j\}t\}\\cdot\\mathrm\{sinc\}\(a\_\{j\}t\)\\cdot e^\{\-\\sigma\_\{j\}^\{2\}t^\{2\}/2\}\. The smallest positive zero att=π/ajt=\\pi/a\_\{j\}determinesaja\_\{j\}; dividing out the sinc factor reveals the Gaussian envelope, recoveringσj\\sigma\_\{j\}\.
Location recovery\.With the kernel known, standard identifiability of location mixtures\[[41](https://arxiv.org/html/2608.05454#bib.bib40)\]recovers the2nb2^\{n\_\{b\}\}mode centers\{cj\+\(Gbβ\)j\}\\\{c\_\{j\}\+\(G\_\{b\}\\beta\)\_\{j\}\\\}\.
Cross\-dimension consistency\.Per\-axis identification yields, for eachjj, the multiset\{cj\+\(Gbβ\)j\}β∈\{−1,1\}nb\\\{c\_\{j\}\+\(G\_\{b\}\\beta\)\_\{j\}\\\}\_\{\\beta\\in\\\{\-1,1\\\}^\{n\_\{b\}\}\}\. Hypothesis \(ii\) of Theorem[3](https://arxiv.org/html/2608.05454#Thmtheorem3)—axis\-wise distinct projected centers—makes each multiset a set of2nb2^\{n\_\{b\}\}points\. Stacking acrossjjrecovers the rows of\[c\|Gb\]\[c\\,\|\\,G\_\{b\}\]up to a permutationπ\\piof\{−1,1\}nb\\\{\-1,1\\\}^\{n\_\{b\}\}that is consistent across all axes \(since the sameπ\\pimust label every per\-axis multiset\)\. The consistency reducesπ\\pito the symmetries of thenbn\_\{b\}\-cube acting onGbG\_\{b\}’s columns, namely column\-permutation composed with global column sign\-flip\. The remaining ambiguity is therefore precisely as stated\.∎
### B\.3Proof of Theorem[6](https://arxiv.org/html/2608.05454#Thmtheorem6)
The scores\(x,y\)s\(x,y\)is a measurable function; split\-conformal prediction\[[36](https://arxiv.org/html/2608.05454#bib.bib17),[3](https://arxiv.org/html/2608.05454#bib.bib38)\]givesℙ\(ytest∈\{y:s\(xtest,y\)≤qα\}\)≥1−α\\mathbb\{P\}\(y\_\{\\text\{test\}\}\\in\\\{y:s\(x\_\{\\text\{test\}\},y\)\\leq q\_\{\\alpha\}\\\}\)\\geq 1\-\\alpha\. The sublevel set unfolds ass\(x,y\)≤q⇔∃β:∀j,\|yj−cj−\(Gbβ\)j\|≤‖Gd\[j,:\]‖1\+q‖Gs\[j,:\]‖2s\(x,y\)\\leq q\\iff\\exists\\beta:\\forall j,\\,\|y\_\{j\}\-c\_\{j\}\-\(G\_\{b\}\\beta\)\_\{j\}\|\\leq\\\|G\_\{d\}\[j,:\]\\\|\_\{1\}\+q\\\|G\_\{s\}\[j,:\]\\\|\_\{2\}, matchingS^α\(x\)\\hat\{S\}\_\{\\alpha\}\(x\)\.∎
### B\.4Theorem[9](https://arxiv.org/html/2608.05454#Thmtheorem9): Distributional Coverage Statement and Proof
###### Theorem 9\(Distributional coverage\)\.
Wheny\|xy\|xfollows the HProbZ distribution, theγ\\gamma\-expanded setSγ\(x\)=⋃β\{z:\|zj−cj−\(Gbβ\)j\|≤‖Gd\[j,:\]‖1\+γ‖Gs\[j,:\]‖2,∀j\}S\_\{\\gamma\}\(x\)=\\bigcup\_\{\\beta\}\\\{z:\|z\_\{j\}\-c\_\{j\}\-\(G\_\{b\}\\beta\)\_\{j\}\|\\leq\\\|G\_\{d\}\[j,:\]\\\|\_\{1\}\+\\gamma\\\|G\_\{s\}\[j,:\]\\\|\_\{2\},\\,\\forall j\\\}satisfiesP\(y∈Sγ\)≥1−2d⋅e−γ2/2P\(y\\in S\_\{\\gamma\}\)\\geq 1\-2d\\cdot e^\{\-\\gamma^\{2\}/2\}; settingγ=2ln\(2d/α\)\\gamma=\\sqrt\{2\\ln\(2d/\\alpha\)\}yields coverage≥1−α\\geq 1\{\-\}\\alpha\. This is a parallel parametric coverage bound to the distribution\-free conformal Theorem[6](https://arxiv.org/html/2608.05454#Thmtheorem6)of the body\.
###### Proof\.
Conditioned on the true modeβ∗\\beta^\{\*\}, the residual in dimensionjjsatisfies\|yj−cj−\(Gbβ∗\)j\|≤‖Gd\[j,:\]‖1\+\|\(Gsν∗\)j\|\|y\_\{j\}\-c\_\{j\}\-\(G\_\{b\}\\beta^\{\*\}\)\_\{j\}\|\\leq\\\|G\_\{d\}\[j,:\]\\\|\_\{1\}\+\|\(G\_\{s\}\\nu^\{\*\}\)\_\{j\}\|\. The marginal\(Gsν∗\)j∼𝒩\(0,‖Gs\[j,:\]‖22\)\(G\_\{s\}\\nu^\{\*\}\)\_\{j\}\\sim\\mathcal\{N\}\(0,\\\|G\_\{s\}\[j,:\]\\\|\_\{2\}^\{2\}\)holds for eachjjregardless of cross\-dimension correlations inGsν∗G\_\{s\}\\nu^\{\*\}, so the per\-dimension Gaussian tail boundP\(\|\(Gsν∗\)j\|\>γ‖Gs\[j,:\]‖2\)≤2e−γ2/2P\(\|\(G\_\{s\}\\nu^\{\*\}\)\_\{j\}\|\>\\gamma\\\|G\_\{s\}\[j,:\]\\\|\_\{2\}\)\\leq 2e^\{\-\\gamma^\{2\}/2\}applies independently\. The two\-sided union bound over the2d2devents\{\(Gsν∗\)j\>γ‖Gs\[j,:\]∥2\}∪\{\(Gsν∗\)j<−γ∥Gs\[j,:\]∥2\}j=1d\\\{\(G\_\{s\}\\nu^\{\*\}\)\_\{j\}\>\\gamma\\\|G\_\{s\}\[j,:\]\\\|\_\{2\}\\\}\\cup\\\{\(G\_\{s\}\\nu^\{\*\}\)\_\{j\}<\-\\gamma\\\|G\_\{s\}\[j,:\]\\\|\_\{2\}\\\}\_\{j=1\}^\{d\}yieldsP\(y∉Sγ∣β=β∗\)≤2d⋅e−γ2/2P\(y\\notin S\_\{\\gamma\}\\mid\\beta=\\beta^\{\*\}\)\\leq 2d\\cdot e^\{\-\\gamma^\{2\}/2\}— valid whether the marginals are independent or correlated, since the union bound only requires the marginal tail probabilities\. HereSγS\_\{\\gamma\}is the set defined in the theorem statement:Sγ=⋃βB\(β\)S\_\{\\gamma\}=\\bigcup\_\{\\beta\}B\(\\beta\)whereB\(β\)=\{z:\|zj−cj−\(Gbβ\)j\|≤‖Gd\[j,:\]‖1\+γ‖Gs\[j,:\]‖2∀j\}B\(\\beta\)=\\\{z:\|z\_\{j\}\-c\_\{j\}\-\(G\_\{b\}\\beta\)\_\{j\}\|\\leq\\\|G\_\{d\}\[j,:\]\\\|\_\{1\}\+\\gamma\\\|G\_\{s\}\[j,:\]\\\|\_\{2\}\\,\\forall j\\\}\. The above bound establishesP\(y∈B\(β∗\)∣β=β∗\)≥1−2de−γ2/2P\(y\\in B\(\\beta^\{\*\}\)\\mid\\beta=\\beta^\{\*\}\)\\geq 1\-2d\\,e^\{\-\\gamma^\{2\}/2\}\. SinceB\(β∗\)⊆SγB\(\\beta^\{\*\}\)\\subseteq S\_\{\\gamma\}, the same lower bound applies toP\(y∈Sγ∣β=β∗\)P\(y\\in S\_\{\\gamma\}\\mid\\beta=\\beta^\{\*\}\)\. Marginalising overβ∗\\beta^\{\*\}preserves the bound because the conditional bound is the same for every value ofβ∗\\beta^\{\*\}\(the constant2de−γ2/22d\\,e^\{\-\\gamma^\{2\}/2\}does not depend onβ∗\\beta^\{\*\}\):P\(y∈Sγ\)=𝔼β∗\[P\(y∈Sγ∣β∗\)\]≥1−2de−γ2/2P\(y\\in S\_\{\\gamma\}\)=\\mathbb\{E\}\_\{\\beta^\{\*\}\}\[P\(y\\in S\_\{\\gamma\}\\mid\\beta^\{\*\}\)\]\\geq 1\-2d\\,e^\{\-\\gamma^\{2\}/2\}\. ∎
### B\.5Proof of Proposition[5](https://arxiv.org/html/2608.05454#Thmtheorem5)
Part \(i\): Bounded entropy cost\.LetU∼Uniform\(−1,1\)U\\sim\\mathrm\{Uniform\}\(\-1,1\)andZ∼𝒩\(0,1\)Z\\sim\\mathcal\{N\}\(0,1\)be independent, sofκf\_\{\\kappa\}is the density ofa\(U\+Z/κ\)=aU\+σZa\(U\+Z/\\kappa\)=aU\+\\sigma Z\. By scaling,h\(fκ\)=loga\+h\(U\+Z/κ\)h\(f\_\{\\kappa\}\)=\\log a\+h\(U\+Z/\\kappa\)\.
The matched Gaussian has varianceσ2\+a2/3=a2\(1/κ2\+1/3\)\\sigma^\{2\}\+a^\{2\}/3=a^\{2\}\(1/\\kappa^\{2\}\+1/3\), soh\(gκ\)=12log\(2πe⋅a2\(1/κ2\+1/3\)\)=loga\+12log\(2πe\(1/κ2\+1/3\)\)h\(g\_\{\\kappa\}\)=\\frac\{1\}\{2\}\\log\(2\\pi e\\cdot a^\{2\}\(1/\\kappa^\{2\}\+1/3\)\)=\\log a\+\\frac\{1\}\{2\}\\log\(2\\pi e\(1/\\kappa^\{2\}\+1/3\)\)\.
The gap isΔ\(κ\)=12log\(2πe\(1/κ2\+1/3\)\)−h\(U\+Z/κ\)\\Delta\(\\kappa\)=\\frac\{1\}\{2\}\\log\(2\\pi e\(1/\\kappa^\{2\}\+1/3\)\)\-h\(U\+Z/\\kappa\)\. Atκ→0\\kappa\{\\to\}0:Z/κZ/\\kappadominatesUU, soU\+Z/κU\+Z/\\kappaapproaches a Gaussian with the matched variance andΔ\(κ\)→0\\Delta\(\\kappa\)\{\\to\}0\. Asκ→∞\\kappa\\to\\infty:h\(U\+Z/κ\)→h\(U\)=log2h\(U\+Z/\\kappa\)\\to h\(U\)=\\log 2, givingΔ\(∞\)=12log\(2πe/3\)−log2=12log\(πe/6\)\\Delta\(\\infty\)=\\frac\{1\}\{2\}\\log\(2\\pi e/3\)\-\\log 2=\\frac\{1\}\{2\}\\log\(\\pi e/6\)\.*Monotonicity\.*Differentiating inκ\\kappaand using De Bruijn’s identity \(ddth\(U\+tZ\)=t⋅J\(U\+tZ\)\\frac\{d\}\{dt\}h\(U\+tZ\)=t\\cdot J\(U\+tZ\)forZ∼𝒩\(0,1\)Z\\sim\\mathcal\{N\}\(0,1\), whereJJis Fisher information\) gives
Δ′\(κ\)=1κ3\[J\(U\+Z/κ\)−1Var\(U\+Z/κ\)\]\.\\Delta^\{\\prime\}\(\\kappa\)\\;=\\;\\frac\{1\}\{\\kappa^\{3\}\}\\Big\[J\(U\+Z/\\kappa\)\\;\-\\;\\frac\{1\}\{\\mathrm\{Var\}\(U\+Z/\\kappa\)\}\\Big\]\.By the Cramér–Rao inequalityJ\(X\)≥1/Var\(X\)J\(X\)\\geq 1/\\mathrm\{Var\}\(X\)with equality iffXXis Gaussian;U\+Z/κU\+Z/\\kappais non\-Gaussian for any finiteκ\>0\\kappa\>0\(the bounded uniform component never vanishes\), soΔ′\(κ\)\>0\\Delta^\{\\prime\}\(\\kappa\)\>0on\(0,∞\)\(0,\\infty\)\. HenceΔ\\Deltais monotone increasing inκ\\kappawith supremum12log\(πe/6\)\\frac\{1\}\{2\}\\log\(\\pi e/6\)\.
Part \(ii\)follows directly from the Fisher information computation in the proof of Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)\(ii\): for a single observed step withGd=aIG\_\{d\}=aIand noiseσ\\sigma, the Fisher information matrix forα\\alphaisGd⊤Gd/σ2=\(a2/σ2\)I=κ2IG\_\{d\}^\{\\top\}G\_\{d\}/\\sigma^\{2\}=\(a^\{2\}/\\sigma^\{2\}\)I=\\kappa^\{2\}I\.
Part \(iii\)is immediate from \(i\) and \(ii\):Δ≤0\.18\\Delta\\leq 0\.18nats regardless ofκ\\kappa, whileρ=κ2→∞\\rho=\\kappa^\{2\}\\to\\infty\.∎
### B\.6Proof of Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)
###### Theorem\(Contraction rate of constrained HProbZ, formal restatement\)\.
Consider the shared\-generator modelzt=ct\+Gb,tβ\+Gd,tα\+Gs,tνtz\_\{t\}=c\_\{t\}\+G\_\{b,t\}\\beta\+G\_\{d,t\}\\alpha\+G\_\{s,t\}\\nu\_\{t\}withβ∈\{−1,1\}nb\\beta\\in\\\{\-1,1\\\}^\{n\_\{b\}\},α∈\[−1,1\]nd\\alpha\\in\[\-1,1\]^\{n\_\{d\}\}, andνt∼iid𝒩\(0,Inq\)\\nu\_\{t\}\\stackrel\{\{\\scriptstyle iid\}\}\{\{\\sim\}\}\\mathcal\{N\}\(0,I\_\{n\_\{q\}\}\)\. The bound below is stated under the variance\-matched Gaussian relaxationα∼𝒩\(0,13Ind\)\\alpha\\sim\\mathcal\{N\}\(0,\\tfrac\{1\}\{3\}I\_\{n\_\{d\}\}\), which is also the relaxation used at inference for the closed\-form posterior solve\. Suppose we observe ground\-truth values at stepst1,…,tkt\_\{1\},\\ldots,t\_\{k\}\.\(i\)Mode identification\.Letβ∗\\beta^\{\*\}denote the true binary assignment\. Under the priorUniform\(\{−1,1\}nb\)\\mathrm\{Uniform\}\(\\\{\-1,1\\\}^\{n\_\{b\}\}\),P\(β=β∗∣zt1,…,ztk\)→1P\(\\beta=\\beta^\{\*\}\\mid z\_\{t\_\{1\}\},\\ldots,z\_\{t\_\{k\}\}\)\\to 1ask→∞k\\to\\infty, provided the mode centersGb,tβG\_\{b,t\}\\betaare distinct for at least one observed step\.\(ii\)Bounded variance contraction\.Define the Fisher informationIk=∑i=1kGd,ti⊤Gd,ti/σti2I\_\{k\}=\\sum\_\{i=1\}^\{k\}G\_\{d,t\_\{i\}\}^\{\\top\}G\_\{d,t\_\{i\}\}/\\sigma\_\{t\_\{i\}\}^\{2\}andρmin=miniλmin\(Gd,ti⊤Gd,ti\)/σti2\\rho\_\{\\min\}=\\min\_\{i\}\\lambda\_\{\\min\}\(G\_\{d,t\_\{i\}\}^\{\\top\}G\_\{d,t\_\{i\}\}\)/\\sigma\_\{t\_\{i\}\}^\{2\}\. Conditioned onβ∗\\beta^\{\*\},
Cov\(α∣zt1,…,ztk,β∗\)⪯\(13Ind\+Ik\)−1,\\mathrm\{Cov\}\(\\alpha\\mid z\_\{t\_\{1\}\},\\ldots,z\_\{t\_\{k\}\},\\beta^\{\*\}\)\\;\\preceq\\;\\Big\(\\tfrac\{1\}\{3\}I\_\{n\_\{d\}\}\+I\_\{k\}\\Big\)^\{\-1\},\(7\)and for any unobserved stepss,tr\(Cov\(Gd,sα∣obs,β∗\)\)≤‖Gd,s‖F2/\(13\+kρmin\)=O\(1/k\)\\mathrm\{tr\}\(\\mathrm\{Cov\}\(G\_\{d,s\}\\alpha\\mid\\mathrm\{obs\},\\beta^\{\*\}\)\)\\leq\\\|G\_\{d,s\}\\\|\_\{F\}^\{2\}/\(\\tfrac\{1\}\{3\}\+k\\rho\_\{\\min\}\)=O\(1/k\), whiletr\(Gs,sGs,s⊤\)\\mathrm\{tr\}\(G\_\{s,s\}G\_\{s,s\}^\{\\top\}\)is irreducible\.\(iii\)MDN structural invariance\.For an MDN\{\(πj,μj,Σj\)\}j=1J\\\{\(\\pi\_\{j\},\\mu\_\{j\},\\Sigma\_\{j\}\)\\\}\_\{j=1\}^\{J\}, conditioning updates only the weights \(πj↦πj∏ip\(zti∣j\)/Z\\pi\_\{j\}\\mapsto\\pi\_\{j\}\\prod\_\{i\}p\(z\_\{t\_\{i\}\}\\mid j\)/Z\);tr\(Σj,s\)\\mathrm\{tr\}\(\\Sigma\_\{j,s\}\)at any unobserved step isΘ\(1\)\\Theta\(1\)inkk\.
Part \(i\): Mode identification\.Conditional on the latentα\\alpha, the per\-step log\-likelihood ratios
Li\(β∗,β;α\)=log𝒩\(zti;cti\+Gb,tiβ∗\+Gd,tiα,Gs,tiGs,ti⊤\)𝒩\(zti;cti\+Gb,tiβ\+Gd,tiα,Gs,tiGs,ti⊤\)L\_\{i\}\(\\beta^\{\*\},\\beta;\\alpha\)\\;=\\;\\log\\frac\{\\mathcal\{N\}\\\!\\left\(z\_\{t\_\{i\}\};\\,c\_\{t\_\{i\}\}\+G\_\{b,t\_\{i\}\}\\beta^\{\*\}\+G\_\{d,t\_\{i\}\}\\alpha,\\,G\_\{s,t\_\{i\}\}G\_\{s,t\_\{i\}\}^\{\\top\}\\right\)\}\{\\mathcal\{N\}\\\!\\left\(z\_\{t\_\{i\}\};\\,c\_\{t\_\{i\}\}\+G\_\{b,t\_\{i\}\}\\beta\+G\_\{d,t\_\{i\}\}\\alpha,\\,G\_\{s,t\_\{i\}\}G\_\{s,t\_\{i\}\}^\{\\top\}\\right\)\}are independent acrossii\(the only random source given\(α,β∗\)\(\\alpha,\\beta^\{\*\}\)is the iid noiseνti\\nu\_\{t\_\{i\}\}\), each with positive mean wheneverGb,tiβ∗≠Gb,tiβG\_\{b,t\_\{i\}\}\\beta^\{\*\}\\neq G\_\{b,t\_\{i\}\}\\beta\. By the strong law of large numbers conditional onα\\alpha,∑iLi\(β∗,β;α\)→\+∞\\sum\_\{i\}L\_\{i\}\(\\beta^\{\*\},\\beta;\\alpha\)\\to\+\\inftya\.s\. ask→∞k\\to\\inftyprovided distinct mode centers are realized at infinitely many observed steps\. The marginal posteriorP\(β∣zt1,…,ztk\)∝2−nb∏ip\(zti∣β\)=2−nb∏i∫\[−1,1\]nd𝒩\(zti;cti\+Gb,tiβ\+Gd,tiα,Gs,tiGs,ti⊤\)dα/Vol\(\[−1,1\]nd\)P\(\\beta\\mid z\_\{t\_\{1\}\},\\ldots,z\_\{t\_\{k\}\}\)\\propto 2^\{\-n\_\{b\}\}\\prod\_\{i\}p\(z\_\{t\_\{i\}\}\\mid\\beta\)=2^\{\-n\_\{b\}\}\\prod\_\{i\}\\int\_\{\[\-1,1\]^\{n\_\{d\}\}\}\\mathcal\{N\}\(z\_\{t\_\{i\}\};\\,c\_\{t\_\{i\}\}\+G\_\{b,t\_\{i\}\}\\beta\+G\_\{d,t\_\{i\}\}\\alpha,\\,G\_\{s,t\_\{i\}\}G\_\{s,t\_\{i\}\}^\{\\top\}\)\\mathrm\{d\}\\alpha/\\mathrm\{Vol\}\(\[\-1,1\]^\{n\_\{d\}\}\)inherits the same limit by standard posterior consistency for identifiable parametric mixtures\[[10](https://arxiv.org/html/2608.05454#bib.bib10), Ch\. 6\]: the conditional ratio diverges Lebesgue\-a\.e\. inα\\alpha, and the dominated convergence theorem propagates the divergence to the marginal\. The technical regularity conditions \(compactness ofα∈\[−1,1\]nd\\alpha\\in\[\-1,1\]^\{n\_\{d\}\}, smoothness of the Gaussian likelihood, finite mode set\) are all satisfied\.
Part \(ii\): Bounded variance contraction \(Gaussian relaxation\)\.Conditioned onβ∗\\beta^\{\*\}, the observation model becomeszti−cti−Gb,tiβ∗=Gd,tiα\+Gs,tiνtiz\_\{t\_\{i\}\}\-c\_\{t\_\{i\}\}\-G\_\{b,t\_\{i\}\}\\beta^\{\*\}=G\_\{d,t\_\{i\}\}\\alpha\+G\_\{s,t\_\{i\}\}\\nu\_\{t\_\{i\}\}, a linear regression inα\\alphawith known Gaussian noise covarianceΣs,ti:=Gs,tiGs,ti⊤\\Sigma\_\{s,t\_\{i\}\}:=G\_\{s,t\_\{i\}\}G\_\{s,t\_\{i\}\}^\{\\top\}\. We replace the true uniform priorα∼Uniform\(\[−1,1\]nd\)\\alpha\\sim\\mathrm\{Uniform\}\(\[\-1,1\]^\{n\_\{d\}\}\)\(covariance13Ind\\tfrac\{1\}\{3\}I\_\{n\_\{d\}\}\) with the variance\-matched Gaussianα∼𝒩\(0,13Ind\)\\alpha\\sim\\mathcal\{N\}\(0,\\tfrac\{1\}\{3\}I\_\{n\_\{d\}\}\), which is the relaxation used at inference for the closed\-form posterior solve\. Under this relaxation the posterior isα∣obs,β∗∼𝒩\(μpost,Σpost\)\\alpha\\mid\\mathrm\{obs\},\\beta^\{\*\}\\sim\\mathcal\{N\}\(\\mu\_\{\\mathrm\{post\}\},\\Sigma\_\{\\mathrm\{post\}\}\)withΣpost=\(3Ind\+Ik\)−1\\Sigma\_\{\\mathrm\{post\}\}=\(3I\_\{n\_\{d\}\}\+I\_\{k\}\)^\{\-1\}whereIk=∑i=1kGd,ti⊤Σs,ti−1Gd,tiI\_\{k\}=\\sum\_\{i=1\}^\{k\}G\_\{d,t\_\{i\}\}^\{\\top\}\\Sigma\_\{s,t\_\{i\}\}^\{\-1\}G\_\{d,t\_\{i\}\}, giving Eq\. \([7](https://arxiv.org/html/2608.05454#A2.E7)\) as a bound on the*relaxed*posterior covariance\. The empirical refinement results in §[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px3)are measured under the true uniform prior and exhibit the sameO\(1/k\)O\(1/k\)rate; we do not formally bound the relaxation gap here\. For the predictive variance at stepss:tr\(Cov\(Gd,sα∣obs\)\)=tr\(Gd,sΣpostGd,s⊤\)≤‖Gd,s‖F2⋅λmax\(Σpost\)≤‖Gd,s‖F2/\(13\+kρmin\)\\mathrm\{tr\}\(\\mathrm\{Cov\}\(G\_\{d,s\}\\alpha\\mid\\mathrm\{obs\}\)\)=\\mathrm\{tr\}\(G\_\{d,s\}\\Sigma\_\{\\mathrm\{post\}\}G\_\{d,s\}^\{\\top\}\)\\leq\\\|G\_\{d,s\}\\\|\_\{F\}^\{2\}\\cdot\\lambda\_\{\\max\}\(\\Sigma\_\{\\mathrm\{post\}\}\)\\leq\\\|G\_\{d,s\}\\\|\_\{F\}^\{2\}/\(\\tfrac\{1\}\{3\}\+k\\rho\_\{\\min\}\), whereρmin=miniλmin\(Gd,ti⊤Σs,ti−1Gd,ti\)\\rho\_\{\\min\}=\\min\_\{i\}\\lambda\_\{\\min\}\(G\_\{d,t\_\{i\}\}^\{\\top\}\\Sigma\_\{s,t\_\{i\}\}^\{\-1\}G\_\{d,t\_\{i\}\}\)and we usedλmax\(\(A\+B\)−1\)≤1/λmin\(A\+B\)\\lambda\_\{\\max\}\(\(A\+B\)^\{\-1\}\)\\leq 1/\\lambda\_\{\\min\}\(A\+B\)andλmin\(Ik\)≥kρmin\\lambda\_\{\\min\}\(I\_\{k\}\)\\geq k\\rho\_\{\\min\}\. The stochastic componentGs,sνsG\_\{s,s\}\\nu\_\{s\}is independent ofα\\alphawith covarianceΣs,s\\Sigma\_\{s,s\}regardless of any conditioning\.
Part \(iii\): MDN structural invariance\.An MDN parameterizesp\(zs∣x\)=∑jπj𝒩\(zs;μj,s\(x\),Σj,s\(x\)\)p\(z\_\{s\}\\mid x\)=\\sum\_\{j\}\\pi\_\{j\}\\mathcal\{N\}\(z\_\{s\};\\mu\_\{j,s\}\(x\),\\Sigma\_\{j,s\}\(x\)\), where all parameters are deterministic outputs of the network\. Conditioning on observed steps updates the mixture weights via Bayes’ rule,πj↦πj∏i𝒩\(zti;μj,ti,Σj,ti\)/Z\\pi\_\{j\}\\mapsto\\pi\_\{j\}\\prod\_\{i\}\\mathcal\{N\}\(z\_\{t\_\{i\}\};\\mu\_\{j,t\_\{i\}\},\\Sigma\_\{j,t\_\{i\}\}\)/Z, but the conditional distribution within each component remains𝒩\(μj,s,Σj,s\)\\mathcal\{N\}\(\\mu\_\{j,s\},\\Sigma\_\{j,s\}\)since the components are independent across steps given the component index\. Hencetr\(Σj,s\)\\mathrm\{tr\}\(\\Sigma\_\{j,s\}\)is invariant tokk\.∎
#### Numerical check of the relaxation gap\.
The bound above is for the Gaussian\-relaxed posteriorα∼𝒩\(0,13I\)\\alpha\\sim\\mathcal\{N\}\(0,\\tfrac\{1\}\{3\}I\), while the empirical23\.6%/31\.5%23\.6\\%/31\.5\\%contractions of §[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px3)and Tab\.[21](https://arxiv.org/html/2608.05454#A4.T21)are measured under the true uniform priorα∼Uniform\(\[−1,1\]\)\\alpha\\sim\\mathrm\{Uniform\}\(\[\-1,1\]\)\. To check that the relaxation is a faithful proxy in the regime our experiments operate in, we run a 1\-D synthetic check: sampleα∗∼Uniform\(\[−1,1\]\)\\alpha^\{\*\}\\sim\\mathrm\{Uniform\}\(\[\-1,1\]\), generatek∈\{1,2,4,8\}k\\in\\\{1,2,4,8\\\}Gaussian observations under eachκ∈\{1\.5,2\.0,2\.8,3\.7\}\\kappa\\in\\\{1\.5,2\.0,2\.8,3\.7\\\}matching Tab\.[21](https://arxiv.org/html/2608.05454#A4.T21), then compare the true uniform\-prior posterior varianceVtrue\(α∣obs\)V\_\{\\text\{true\}\}\(\\alpha\\mid\\mathrm\{obs\}\)\(computed by fine\-grid integration over\[−1,1\]\[\-1,1\]\) against the closed\-form Gaussian\-relaxed varianceVrelaxed=1/\(3\+kκ2\)V\_\{\\text\{relaxed\}\}=1/\(3\+k\\kappa^\{2\}\)\. Averaging over1,0001\{,\}000Monte\-Carlo realisations per\(κ,k\)\(\\kappa,k\)cell, the maximum relative gap\|Vrelaxed−Vtrue\|/Vtrue\|V\_\{\\text\{relaxed\}\}\-V\_\{\\text\{true\}\}\|/V\_\{\\text\{true\}\}across all1616cells is9\.03%\\mathbf\{9\.03\\%\}\(Fig\.[4](https://arxiv.org/html/2608.05454#A2.F4)\); the gap is positive throughout, soVrelaxedV\_\{\\text\{relaxed\}\}is a small over\-estimate ofVtrueV\_\{\\text\{true\}\}and the contraction reported in the body is, if anything, slightly conservative relative to the true uniform\-prior posterior\. Code and reproduction script in the supplementary archive\.
Figure 4:Relative gap\(Vrelaxed−Vtrue\)/Vtrue\(V\_\{\\mathrm\{relaxed\}\}\-V\_\{\\mathrm\{true\}\}\)/V\_\{\\mathrm\{true\}\}between the Gaussian\-relaxed posterior variance \(closed\-form\) and the true uniform\-prior posterior variance \(grid integration,1,0001\{,\}000MC reps\), as a function of revealed observationskkand Fisher\-information ratioκ\\kappa\. All1616cells lie inside the±10%\\pm 10\\%band; the relaxation slightly over\-estimates the true variance, so contraction reported in the body is conservative relative to the true uniform\-prior posterior\.
#### Implementation of constrained inference\.
At inference, whenkkground\-truth positionsztiz\_\{t\_\{i\}\}become available, we form the constraint rowsAbβ\+Adα\+Asν=bA\_\{b\}\\beta\+A\_\{d\}\\alpha\+A\_\{s\}\\nu=bof Definition[1](https://arxiv.org/html/2608.05454#Thmtheorem1)by stackingAb=\[Gb,t1;…;Gb,tk\]A\_\{b\}=\[G\_\{b,t\_\{1\}\};\\dots;G\_\{b,t\_\{k\}\}\],Ad=\[Gd,t1;…;Gd,tk\]A\_\{d\}=\[G\_\{d,t\_\{1\}\};\\dots;G\_\{d,t\_\{k\}\}\],As=blkdiag\(σtiI\)A\_\{s\}=\\mathrm\{blkdiag\}\(\\sigma\_\{t\_\{i\}\}I\), andb=\[zti−cti\]i=1kb=\[z\_\{t\_\{i\}\}\-c\_\{t\_\{i\}\}\]\_\{i=1\}^\{k\}\. We then enumerate all2nb2^\{n\_\{b\}\}binary assignmentsβ\\beta, evaluate the closed\-form Gaussian posterior over\(α,ν\)\(\\alpha,\\nu\)from Part \(ii\), and propagate the resulting tightened generators to all unobserved steps in one matrix\-vector pass\. Total cost isO\(2nb\(k⋅nd2\+nd3\)\)O\(2^\{n\_\{b\}\}\(k\\cdot n\_\{d\}^\{2\}\+n\_\{d\}^\{3\}\)\)per scene—dominated by thend×ndn\_\{d\}\\times n\_\{d\}posterior solve—and is negligible compared to the encoder forward pass for thenb,nd∈\{1,2,3\}n\_\{b\},n\_\{d\}\\in\\\{1,2,3\\\}regime used throughout this paper\.
## Appendix CEmpirical Theorem Validation
### C\.1Per\-Value Tables for Controlled Sweeps
Figure 5:Controlled generator sweeps under exact vs\. Gaussian\-approximate likelihood; mean±\\pmstd over 5 seeds\.Left:mode sweep \(K=1→7K\{=\}1\{\\to\}7\): both likelihoods identify modal structure viaGbG\_\{b\}\.Middle:drift sweep: the exact likelihood routes drift toGdG\_\{d\}\(Δ‖Gd‖=2\.90≫Δ‖Gs‖=2\.27\\Delta\\\|G\_\{d\}\\\|\{=\}2\.90\\gg\\Delta\\\|G\_\{s\}\\\|\{=\}2\.27,Δ‖Gb‖=0\.73\\Delta\\\|G\_\{b\}\\\|\{=\}0\.73\); the Gaussian approximation reverses the dominance \(Δ‖Gs‖=2\.70\\Delta\\\|G\_\{s\}\\\|\{=\}2\.70wins\), confirming Prop\.[4](https://arxiv.org/html/2608.05454#Thmtheorem4)\.Right:noise sweep:‖Gs‖\\\|G\_\{s\}\\\|tracksσ\\sigmamonotonically; residualGd/GsG\_\{d\}/G\_\{s\}degeneracy is halved by the kurtosis regulariser \(App\.[D\.7](https://arxiv.org/html/2608.05454#A4.SS7)\)\. Per\-value tables below\.This section gives the controlled\-sweep figure summarising the body §[5\.1](https://arxiv.org/html/2608.05454#S5.SS1)narrative, and the per\-value tables behind each panel\.
Table 10:Mode sweep \(exact likelihood, fixed drift0\.30\.3, noise0\.150\.15, mean±\\pmstd over 5 random seeds\)\.‖Gb‖\\\|G\_\{b\}\\\|grows monotonically with mode count and dominates the per\-mode delta; the elevated std atK∈\{3,5\}K\\in\\\{3,5\\\}reflects the bimodal optimisation landscape on this synthetic task \(see Appendix[D\.24](https://arxiv.org/html/2608.05454#A4.SS24)for the same effect on the drift sweep, where sym\-broken initialisation suppresses it\)\.Table 11:Drift sweep: exact vs\. Gaussian approximation \(fixed 3 modes, noise0\.150\.15,55random seeds∈\{42,123,456,789,2026\}\\in\\\{42,123,456,789,2026\\\}, mean±\\pmstd, sym\-brokenGbG\_\{b\}initialisation per Appendix[D\.24](https://arxiv.org/html/2608.05454#A4.SS24)\)\. Under the exact likelihood,‖Gb‖\\\|G\_\{b\}\\\|remains essentially flat \(Δ=0\.73\\Delta\{=\}0\.73\) whileΔ‖Gd‖=2\.90\\Delta\\\|G\_\{d\}\\\|\{=\}2\.90dominates—drift is correctly routed to the bounded generator\. Under the Gaussian approximation,Δ‖Gs‖=2\.70\\Delta\\\|G\_\{s\}\\\|\{=\}2\.70dominates instead, confirming Proposition[4](https://arxiv.org/html/2608.05454#Thmtheorem4): the surrogate misattributes bounded variance to the stochastic component\.Table 12:Noise sweep: exact vs\. Gaussian approximation \(fixed 3 modes, drift0\.30\.3, 200 epochs, mean±\\pmstd over 5 seeds∈\{42,123,456,789,2026\}\\in\\\{42,123,456,789,2026\\\}\)\.‖Gs‖\\\|G\_\{s\}\\\|rises monotonically withσ\\sigmaunder both likelihoods\. ResidualGd/GsG\_\{d\}/G\_\{s\}degeneracy is addressed separately by the kurtosis regulariser \(App\.[D\.7](https://arxiv.org/html/2608.05454#A4.SS7), Tab\.[19](https://arxiv.org/html/2608.05454#A4.T19)\)\.
### C\.2Theorem[2](https://arxiv.org/html/2608.05454#Thmtheorem2): Characteristic\-Function Fingerprint Setup
The body figure \(Fig\.[2](https://arxiv.org/html/2608.05454#S3.F2)\) summarises the empirical witness; this subsection records the full setup\. We sample10510^\{5\}points from a 1\-D HProbZ witha=2,σ=0\.3a\{=\}2,\\sigma\{=\}0\.3and fit GMMs withK∈\{1,…,64\}K\\in\\\{1,\\ldots,64\\\}\. At the predicted sinc\-zero frequenciestk=kπ/at\_\{k\}=k\\pi/a, the GMM’s CF magnitude plateaus at≈2−5×10−3\\approx 2\{\-\}5\\times 10^\{\-3\}and ceases to decrease withKK, the theorem\-predicted signature that no finite GMM can replicate these structural zeros\. HProbZ matches its own held\-out NLL with 2 free density parameters\(a,σ\)\(a,\\sigma\)\(the mean is fixed at0\); the smallest matching GMM requiresK=8K\{=\}8with3K−1=233K\-1=23free parameters \(means, variances, andK−1K\-1mixing weights\), an11\.5×11\.5\\timesratio\. Counting all 4 HProbZ scalars\(c,Gb,Gd,Gs\)\(c,G\_\{b\},G\_\{d\},G\_\{s\}\)in 1\-D against3K=243K\{=\}24GMM parameters gives6×6\\times, still well outside the GMM’s reach for any finiteKK\.
### C\.3Theorem[3](https://arxiv.org/html/2608.05454#Thmtheorem3): Multi\-Seed Parameter Recovery
Training from 10 random seeds on5×1045\\times 10^\{4\}samples from a known HProbZ, all seeds recover every parameter to within0\.0080\.008of ground truth, with held\-out NLL identical to four decimal places \(Figure[6](https://arxiv.org/html/2608.05454#A3.F6)\)\.
Figure 6:All 10 seeds recover ground\-truth HProbZ parameters; per\-seed error below0\.0040\.004\.
### C\.4Proposition[5](https://arxiv.org/html/2608.05454#Thmtheorem5): Entropy\-Gap Bound Verification
We distinguish Proposition[5](https://arxiv.org/html/2608.05454#Thmtheorem5)’s claim—the density\-level entropy gaph\(gκ\)−h\(fκ\)h\(g\_\{\\kappa\}\)\-h\(f\_\{\\kappa\}\)is at most12log\(πe/6\)≈0\.1765\\tfrac\{1\}\{2\}\\log\(\\pi e/6\)\\approx 0\.1765nats/dim—from the empirical NLL gapNLL\(fκ\|data\)−NLL\(g∗\|data\)\\text\{NLL\}\(f\_\{\\kappa\}\|\\text\{data\}\)\-\\text\{NLL\}\(g^\{\*\}\|\\text\{data\}\), which also contains data\-dependent misspecification orthogonal to the theoretical claim\. To check the theorem as stated, we computeh\(gκ\)−h\(fκ\)h\(g\_\{\\kappa\}\)\-h\(f\_\{\\kappa\}\)directly from the per\-dimension\(a,σ\)\(a,\\sigma\)parameters output by trained HProbZ models on all three benchmarks\. The visual summary is in the body \(Figure[3](https://arxiv.org/html/2608.05454#S5.F3)\); per\-dataset table follows\. The observed maxima \(0\.140\.14on nuScenes;≤1\.3×10−4\\leq 1\.3\{\\times\}10^\{\-4\}on ETH/UCY and Argoverse 2\) sit strictly below0\.17650\.1765\. This validates Proposition[5](https://arxiv.org/html/2608.05454#Thmtheorem5)as stated, and locates the larger raw NLL gaps entirely within the separately\-known misspecification excess that any bounded\-support family incurs on unbounded\-tail data\.
Table 13:Empirical verification of Proposition[5](https://arxiv.org/html/2608.05454#Thmtheorem5): the entropy gaph\(gκ\)−h\(fκ\)h\(g\_\{\\kappa\}\)\-h\(f\_\{\\kappa\}\)computed on trained HProbZ models across three benchmarks is universally bounded by12log\(πe/6\)≈0\.1765\\tfrac\{1\}\{2\}\\log\(\\pi e/6\)\\approx 0\.1765nats/dim, with zero violations across 13,000\+ samples\.
### C\.5nuScenes CRPS: Per\-Seed Stability
The body argues that HProbZ’s structural advantage on CRPS is best read as*seed\-level stability*, not mean reduction\. Table[14](https://arxiv.org/html/2608.05454#A3.T14)reports the per\-seed CRPS for55paired \(HProbZ, MDN\-K2\) trainings with matched encoder and matched architecture \(the original paper\-submission MDN\-K2 head:Linear\(d,d\)→\\toGELU→\\toLinear\(d,Tf⋅K⋅5T\_\{\\\!f\}\\\!\\cdot\\\!K\\\!\\cdot\\\!5\)\), evaluated on the full39,82939\{,\}829\-sample nuScenes test set withK=100K\{=\}100trajectory samples per estimate\. Seed pairs\{7,13,21,35\}\\\{7,13,21,35\\\}are freshly trained; thev5\_paperpair uses the original released checkpoints\.
Table 14:Per\-seed CRPS on nuScenes for matched \(HProbZ, MDN\-K2\) pairs\. HProbZ stays in\[0\.80,0\.86\]\[0\.80,0\.86\]across all five seeds; MDN\-K2 spans\[0\.66,0\.92\]\[0\.66,0\.92\]\. Means are statistically tied \(95%95\\%paired\-bootstrap CI on the gap\[−0\.07,\+0\.07\]\[\-0\.07,\+0\.07\], contains0\); HProbZ is tighter on every spread metric we computed on these 5 seeds \(sample\-std4\.5×4\.5\\times, IQR2\.6×2\.6\\times, MAD9\.6×9\.6\\times, range5\.0×5\.0\\times\), but withn=5n\{=\}5the std\-ratio bootstrap CI is wide \(\[0\.9×,45×\]\[0\.9\\times,45\\times\], contains unity\), so this is a directional observation rather than an established gap\. Robustness diagnostics in the paragraph below\. All numbers reproducible via the CRPS evaluation script in the supplementary code\.The seed\-level std difference \(σHProbZ=0\.020\\sigma\_\{\\text\{HProbZ\}\}\{=\}0\.020,σMDN=0\.092\\sigma\_\{\\text\{MDN\}\}\{=\}0\.092\) is the practical observation: on these 5 paired seeds, HProbZ’s CRPS at deployment time appears less sensitive to the random seed than MDN\-K2’s\. The bootstrap CI on the std ratio is wide \(\[0\.9×,45×\]\[0\.9\\times,45\\times\]\) and one MDN seed \(seed 21, CRPS=0\.662=0\.662\) is responsible for a meaningful share of the spread gap; we therefore do*not*claim an established calibration advantage from thisn=5n\{=\}5pairing\. For safety\-critical use, calibration that depends on training luck is a deployment risk in principle, but verifying that this risk differs significantly between heads requires a larger paired sweep that we leave to future work\.
#### Bootstrap and robustness analysis\.
Becausen=5n\{=\}5paired seeds is small for std\-of\-std inference, we report three diagnostics\. \(a\)*Paired bootstrap\.*Resampling the 5 seed indices with replacement \(B=10,000B\{=\}10\{,\}000\): mean\-CRPS gap95%95\\%CI\[−0\.07,\+0\.07\]\[\-0\.07,\+0\.07\]\(zero contained, supporting the “ties on mean” framing\); std\-ratio \(MDN/HProbZ\) bootstrap median4\.6×4\.6\\times,95%95\\%percentile CI\[0\.9×,45×\]\[0\.9\\times,45\\times\]\. The CI is wide becausen=5n\{=\}5, but the directional findingσMDN\>σHProbZ\\sigma\_\{\\text\{MDN\}\}\>\\sigma\_\{\\text\{HProbZ\}\}holds in95\.4%95\.4\\%of resamples\. \(b\)*Spread\-metric robustness\.*Sample\-std ratio4\.5×4\.5\\times; IQR ratio0\.079/0\.030=2\.6×0\.079/0\.030=2\.6\\times; MAD ratio0\.048/0\.005=9\.6×0\.048/0\.005=9\.6\\times; range ratio0\.260/0\.052=5\.0×0\.260/0\.052=5\.0\\times\. All four robust spread measures show MDN\-K2 is wider; the magnitude depends on the metric\. \(c\)*Leave\-one\-out\.*Removing each seed in turn yields std\-ratios\{4\.80,4\.55,2\.04,6\.59,4\.70\}×\\\{4\.80,4\.55,2\.04,6\.59,4\.70\\\}\\times\. The2\.04×2\.04\\timesentry corresponds to dropping seed 21 \(MDN=0\.662=0\.662\); the remaining MDN range0\.8080\.808–0\.9220\.922is still2\.2×2\.2\\timeswider than HProbZ’s, so the effect survives outlier removal but with diminished magnitude\. We report “consistently tighter” rather than “4\.5×4\.5\\timestighter” in the body to reflect this fragility\.
### C\.6Why HProbZ’s Raw NLL Exceeds MDN’s
HProbZ’s per\-axis marginal is12fUG\(y−c−gb;a,σ\)\+12fUG\(y−c\+gb;a,σ\)\\tfrac\{1\}\{2\}f\_\{\\text\{UG\}\}\(y\-c\-g\_\{b\};\\,a,\\sigma\)\+\\tfrac\{1\}\{2\}f\_\{\\text\{UG\}\}\(y\-c\+g\_\{b\};\\,a,\\sigma\)withfUG\(z;a,σ\)=12a\[Φ\(z\+aσ\)−Φ\(z−aσ\)\]f\_\{\\text\{UG\}\}\(z;a,\\sigma\)=\\tfrac\{1\}\{2a\}\[\\Phi\(\\tfrac\{z\+a\}\{\\sigma\}\)\-\\Phi\(\\tfrac\{z\-a\}\{\\sigma\}\)\]\. The density is near\-flat at≈1/\(2a\)\\approx 1/\(2a\)on the plateau\|z\|≤a\|z\|\\leq a, whereas an MDN component𝒩\(μk,σk2\)\\mathcal\{N\}\(\\mu\_\{k\},\\sigma\_\{k\}^\{2\}\)peaks at1/\(σk2π\)1/\(\\sigma\_\{k\}\\sqrt\{2\\pi\}\)—typically larger whena≫σka\\gg\\sigma\_\{k\}\. NLL, a sum of pointwise log\-densities, rewards this sharper peak; the cross\-time shared\-latent coupling that makes Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)possible is invisible to a pointwise score\. CRPS \(integrated coverage\), closed\-loop collision rate \(joint tail event\), and per\-mode risk consistency \(sampling variance\) each capture a distinct calibration notion that NLL alone does not, which is why HProbZ improves on all three while losing on raw NLL\.
## Appendix DAdditional Experiments
The subsections below cover six logical groups:
Ablations and hyperparameter sweeps\.Encoder Scaling \(§[D\.6](https://arxiv.org/html/2608.05454#A4.SS6)\); Kurtosis Regulariser \(§[D\.7](https://arxiv.org/html/2608.05454#A4.SS7)\);nbn\_\{b\}Sweep \(§[D\.8](https://arxiv.org/html/2608.05454#A4.SS8)\); Structural Ablation \(§[D\.10](https://arxiv.org/html/2608.05454#A4.SS10)\); MDN Sweep \(§[D\.22](https://arxiv.org/html/2608.05454#A4.SS22)\); MDN\-K=8 Collapse \(§[D\.27](https://arxiv.org/html/2608.05454#A4.SS27)\); Learnable Mode Prior \(§[D\.28](https://arxiv.org/html/2608.05454#A4.SS28)\); Symmetry\-Broken Init \(§[D\.24](https://arxiv.org/html/2608.05454#A4.SS24)\)\.
Calibration extensions\.Split\-Conformal Calibration on the controlled task \(§[D\.3](https://arxiv.org/html/2608.05454#A4.SS3)\); Conformal Set Volume on Real Benchmarks \(§[D\.4](https://arxiv.org/html/2608.05454#A4.SS4)\); Out\-of\-Distribution Conformal Robustness \(§[D\.5](https://arxiv.org/html/2608.05454#A4.SS5)\)\.
Downstream task details\.Pedestrian Crossing Planning \(§[D\.23](https://arxiv.org/html/2608.05454#A4.SS23)\); Planning Separation Sweep \(§[D\.18](https://arxiv.org/html/2608.05454#A4.SS18)\); Risk\-Estimate Consistency Full Sweep \(§[D\.19](https://arxiv.org/html/2608.05454#A4.SS19)\); Trajectory\-Level Closed\-Loop Collision \(§[D\.20](https://arxiv.org/html/2608.05454#A4.SS20)\); Rolling Closed\-Loop \(§[D\.25](https://arxiv.org/html/2608.05454#A4.SS25)\); nuPlan Cross\-Dataset \(§[D\.26](https://arxiv.org/html/2608.05454#A4.SS26)\); Refinement Runtime \(§[D\.21](https://arxiv.org/html/2608.05454#A4.SS21)\)\.
Baseline comparisons \(extended\)\.Varying\-Uncertainty Baselines \(§[D\.1](https://arxiv.org/html/2608.05454#A4.SS1)\); Controlled Diffusion Comparison \(§[D\.11](https://arxiv.org/html/2608.05454#A4.SS11)\); ETH/UCY Diffusion Baseline \(§[D\.12](https://arxiv.org/html/2608.05454#A4.SS12)\)\. Published\-method comparison \(Tab\.[28](https://arxiv.org/html/2608.05454#A4.T28)\) and design\-space positioning \(Tab\.[1](https://arxiv.org/html/2608.05454#S1.T1)\) are in the body\.
Per\-scene, per\-modality, and per\-seed analyses\.nuScenes Constraint Refinement \(§[D\.9](https://arxiv.org/html/2608.05454#A4.SS9)\); ETH/UCY Per\-Scene \(§[D\.13](https://arxiv.org/html/2608.05454#A4.SS13)\); Per\-Scene Generator Decomposition \(§[D\.15](https://arxiv.org/html/2608.05454#A4.SS15)\); nuScenes Speed\-Regime \(§[D\.17](https://arxiv.org/html/2608.05454#A4.SS17)\); nuScenes CRPS Per\-Seed Stability \(§[C\.5](https://arxiv.org/html/2608.05454#A3.SS5), in §[C](https://arxiv.org/html/2608.05454#A3)\)\.
Visualisations\.Active Learning via Modal Uncertainty \(next subsection\); Concept Illustration and Trajectory Visualisations \(§[D\.14](https://arxiv.org/html/2608.05454#A4.SS14)\)\.
### D\.1Baseline Comparison under Varying Uncertainty
We compare HProbZ against MDN, Deep Ensembles\[[18](https://arxiv.org/html/2608.05454#bib.bib20)\], MC\-Dropout\[[9](https://arxiv.org/html/2608.05454#bib.bib21)\], and CVAE\[[33](https://arxiv.org/html/2608.05454#bib.bib23)\]across three uncertainty regimes \(Low:22modes / drift0\.10\.1/ noise0\.10\.1; Medium:4/0\.3/0\.24/0\.3/0\.2; High:7/0\.5/0\.37/0\.5/0\.3\)\. When the data\-generating uncertainty is non\-trivial \(Medium and High regimes\), HProbZ achieves the highest empirical coverage \(≥0\.997\\geq 0\.997\) while matching MDN\-level FDE \(Figure[7](https://arxiv.org/html/2608.05454#A4.F7)\)\. In the Low regime, where every method’s analytic confidence interval already over\-covers the small intrinsic noise, HProbZ’s intervals are narrower \(smaller volume\) by design and consequently under\-cover at the0\.950\.95target \(0\.630\.63\); the same compactness pays off as Volume scales sub\-linearly with uncertainty in Medium/High whereas variance\-based baselines either inflate or fail to reach target coverage\.
Figure 7:HProbZ achieves the highest coverage when the data\-generating uncertainty is non\-trivial \(Medium and High regimes,≥0\.997\\geq 0\.997\) while maintaining competitive FDE\. In the Low regime, HProbZ’s tight intervals trade coverage for compactness; variance\-based baselines reach the95%95\\%target there only by inflating volume far beyond what the small intrinsic noise warrants\.
### D\.2Active Learning via Modal Uncertainty
Binary generators isolate modal uncertainty from bounded and stochastic components, giving a direct acquisition signal for active learning:‖Gb‖\\\|G\_\{b\}\\\|targets samples with high modal ambiguity rather than high overall variance\. Over ten acquisition rounds \(100 queries each from a 10,000\-sample pool,200200epochs of fresh training per round with deterministic per\-round seeding\), HProbZ\-‖Gb‖\\\|G\_\{b\}\\\|selection yields a net test\-NLL improvement of\+0\.88\+0\.88, whereas MDN total\-variance acquisition fails \(−1\.80\-1\.80\) and random selection lies in between \(\+0\.52\+0\.52\); per\-round trajectories in Figure[8](https://arxiv.org/html/2608.05454#A4.F8); reproduction archive in the supplementary code\.
Figure 8:HProbZ’s modal acquisition signal improves test NLL; MDN total\-variance acquisition and random selection both degrade\. Scope: comparison is HProbZ\-‖Gb‖\\\|G\_\{b\}\\\|vs MDN\-total\-variance vs random under a single acquisition recipe; standard Bayesian\-acquisition baselines \(BALD, BatchBALD, ensemble disagreement\) are*not*included, and the negative MDN result is specific to the total\-variance signal\. We report this as a structural\-decomposition demonstration, not a horse\-race against the AL literature\.
### D\.3Split\-Conformal Calibration
Validating Theorem[6](https://arxiv.org/html/2608.05454#Thmtheorem6)on the controlled task, conformal calibration produces prediction sets whose volume is roughly half that of the distributional bound while retaining target coverage within sampling noise \(Table[15](https://arxiv.org/html/2608.05454#A4.T15)\)\.
Table 15:Conformal vs\. distributional coverage and volume\.
### D\.4Conformal Set Volume on Real Benchmarks
The body table \(Tab\.[7](https://arxiv.org/html/2608.05454#S5.T7)\) reports per\-scene volumes; this section records the calibration protocol\. For each held\-out scene, we train HProbZ \(nb=3,nd=1,nq=1n\_\{b\}\{=\}3,n\_\{d\}\{=\}1,n\_\{q\}\{=\}1\) and MDN \(K=8K\{=\}8\) on the remaining four scenes, then split the held\-out scene’s test data 50/50 into calibration and evaluation sets \(α=0\.1\\alpha\{=\}0\.1\)\. This within\-scene calibration ensures exchangeability between calibration and test samples, yielding valid marginal coverage guarantees\[[36](https://arxiv.org/html/2608.05454#bib.bib17)\]\. Three conformal methods are compared:HProbZ\-CPuses the nonconformity score from Eq\. \([6](https://arxiv.org/html/2608.05454#S4.E6)\), producing a union of2nb=82^\{n\_\{b\}\}\{=\}8axis\-aligned boxes;Standard CPwraps a singleL∞L\_\{\\infty\}box around the HProbZ center prediction;MDN\-CPselects the closest MDN component per sample and uses the standardized residual, producing a union ofKKcomponent\-specific boxes\. The prediction dimension isD=24D\{=\}24\(12 steps×\\times2 coordinates\), so even modest per\-dimension differences compound into large volume gaps\.
The volume advantage stems from HProbZ’s multi\-modal structure: the union of tight boxes centered on each binary mode covers the same target region as a single inflated box, but with far less wasted volume in the high\-dimensional \(D=24D\{=\}24\) prediction space\. Coverage tracks the target across multipleα\\alphalevels \(α=0\.05\\alpha\{=\}0\.05:94\.5%94\.5\\%;α=0\.10\\alpha\{=\}0\.10:89\.9%89\.9\\%;α=0\.15\\alpha\{=\}0\.15:85\.2%85\.2\\%\), confirming calibration validity\. The HOTEL counter\-example reflects HOTEL’s tightly concentrated entrance trajectories: MDN\-CP’s per\-component covariance fits the narrow target region closely enough that itsK=8K\{=\}8component union beats HProbZ\-CP’s2nb=82^\{n\_\{b\}\}\{=\}8binary\-mode union; on the more dispersed ETH/UNIV/ZARA1/ZARA2 scenes the structural gap reasserts itself\.
### D\.5Out\-of\-Distribution Conformal Robustness
The split\-conformal guarantee \(Theorem[6](https://arxiv.org/html/2608.05454#Thmtheorem6)\) requires only exchangeability, not identical distributions\. We test this robustness under two realistic distribution shifts\.
#### ETH/UCY leave\-one\-out \(cross\-scene OOD\)\.
We use the pre\-trained leave\-one\-out HProbZ checkpoints \(nb=3,nd=1,nq=1n\_\{b\}\{=\}3,n\_\{d\}\{=\}1,n\_\{q\}\{=\}1, trained on 4 scenes with the 5th held out\)\. For each held\-out scene, we split its test set 50/50 into calibration and evaluation subsets and compute coverage atα=0\.1\\alpha\{=\}0\.1\. Results in Table[16](https://arxiv.org/html/2608.05454#A4.T16)show coverage remains near 90% across all 5 held\-out scenes \(avg90\.9%90\.9\\%\), despite strong inter\-scene distribution shifts \(different pedestrian densities, interaction patterns, camera perspectives\)\.
#### nuScenes speed\-stratified \(within\-scene OOD\)\.
We compute the past\-observed mean speed for each nuScenes test trajectory \(finite differences at1010Hz\) and split into low\-speed \(<5<5m/s;31,18531\{,\}185samples\) and high\-speed \(≥5\\geq 5m/s;8,6448\{,\}644samples\)\. The trained nuScenes HProbZ \(nb=1n\_\{b\}\{=\}1\) is calibrated on half of the low\-speed subset, then evaluated on \(i\) the remaining low\-speed \(IID\) and \(ii\) the entire high\-speed subset \(OOD\)\. Coverage is89\.7%89\.7\\%IID and99\.8%99\.8\\%OOD: the high\-speed regime is more predictable \(vehicles on highways follow straighter trajectories with less multi\-modality\), so the calibrated quantile from low\-speed traffic over\-covers\. This confirms coverage is preserved or exceeded under speed shift, never under\-covering\.
Table 16:OOD conformal coverage: HProbZ maintains calibration under distribution shift\. ETH/UCY uses leave\-one\-out \(train on 4, test on 1\)\. nuScenes calibrates on low\-speed, evaluates on both low\-speed \(IID\) and high\-speed \(OOD\)\.SettingCoverageTargetETH/UCY LOO: ETH held out94\.5%90%ETH/UCY LOO: HOTEL held out90\.3%90%ETH/UCY LOO: UNIV held out89\.8%90%ETH/UCY LOO: ZARA1 held out89\.5%90%ETH/UCY LOO: ZARA2 held out90\.4%90%Average \(5 scenes\)90\.9%90%nuScenes IID \(low\-speed eval\)89\.7%90%nuScenes OOD \(high\-speed eval\)99\.8%90%Full per\-scene/per\-regime numbers are in Tab\.[16](https://arxiv.org/html/2608.05454#A4.T16)above\.
### D\.6Encoder Scaling: HProbZ vs\. MDN
Table[17](https://arxiv.org/html/2608.05454#A4.T17)tests whether HProbZ’s structured uncertainty head benefits more from increased encoder capacity than an unstructured MDN head\. We sweep the transformer encoder across three sizes—Small \(d=64d\{=\}64, 2 layers, 110K params\), Base \(d=128d\{=\}128, 3 layers, 615K params\), and Large \(d=256d\{=\}256, 4 layers, 2\.4M params\)—while keeping the HProbZ head \(nb=3,nd=1,nq=1n\_\{b\}\{=\}3,n\_\{d\}\{=\}1,n\_\{q\}\{=\}1\) and the MDN head \(K=20K\{=\}20\) fixed\. Each configuration is trained for 50 epochs with batch size 1024 under the leave\-one\-out protocol \(5 scenes, 5 seeds\)\.
Table 17:Encoder scaling on ETH/UCY \(mean±\\pmstd over 5 seeds\)\. Across all three encoder sizes \(110K–2\.4M parameters\), HProbZ’s minADE is∼21%\\sim 21\\%lower than MDN\-K20’s; the ADE gap is stable across capacity\.*On minFDE the comparison reverses*: MDN\-K20 wins by0\.060\.06–0\.090\.09m at every capacity \(bolded entries\)\. Net read: in this vanilla protocol HProbZ is sample\-efficient on full\-trajectory accuracy and MDN\-K20 is sharper on endpoint; the social\-attention encoder used for the SOTA\-chase row of Tab\.[26](https://arxiv.org/html/2608.05454#A4.T26)closes the FDE gap\.HProbZ achieves21%21\\%lower minADE than MDN across all three encoder sizes \(e\.g\.,0\.2660\.266vs\.0\.3370\.337at Large scale\), suggesting that the structured generator decomposition provides a consistent ADE advantage regardless of backbone capacity\.*On minFDE the comparison reverses*: MDN\-K20 wins by0\.060\.06–0\.090\.09m at every capacity \(0\.436/0\.422/0\.4160\.436/0\.422/0\.416m vs HProbZ0\.526/0\.505/0\.4760\.526/0\.505/0\.476\)\. Both methods improve with larger encoders \(HProbZ ADE:0\.290→0\.2660\.290\\to 0\.266, MDN ADE:0\.369→0\.3370\.369\\to 0\.337\), confirming that the HProbZ head can absorb additional representational capacity\. The MDN\-K20 FDE advantage we attribute to MDN’s freedom to placeK=20K\{=\}20independent sharp Gaussian endpoints \(the network has no incentive to share parameters across timesteps\), whereas HProbZ’s shared bounded generator distributes probability mass along the full prediction horizon\. We report this as an empirical trade\-off in the matched\-encoder vanilla protocol rather than an inherent limitation; the body Tab\.[26](https://arxiv.org/html/2608.05454#A4.T26)\(Social HProbZ at MDN\-K8\) shows HProbZ winning FDE under the social\-attention configuration we use for the SOTA\-chase number, so the Social\-row FDE advantage is contributed by the social encoder rather than the head alone\. Notably, HProbZ’s ADE advantage \(21%21\\%\) exceeds MDN’s FDE advantage \(14%14\\%\), and minADE measures accuracy over the entire predicted trajectory, which is more relevant for safety\-critical applications such as collision avoidance and path planning where the full predicted path—not just the endpoint—determines the decision boundary\.
### D\.7Kurtosis Regularizer Ablation
#### Drift sweep\.
On the drift\-sweep experiment \(§[5\.1](https://arxiv.org/html/2608.05454#S5.SS1)\), adding an empirical\-kurtosis\-matching regularizer actually degrades identification due to interaction withGbG\_\{b\}mode assignment \(Table[18](https://arxiv.org/html/2608.05454#A4.T18)\)\. This confirms that identifiability on the drift sweep is a property of the exact likelihood itself, not of auxiliary moment\-matching losses\.
Table 18:Kurtosis regularizer ablation on drift sweep: the exact likelihood alone achieves correct identification\.
#### Noise sweep \(addressingGd/GsG\_\{d\}/G\_\{s\}degeneracy\)\.
On the noise sweep whereGd/GsG\_\{d\}/G\_\{s\}degeneracy is the primary challenge \(drift=0\.3\{\}=0\.3, noise∈\[0\.05,0\.50\]\\in\[0\.05,0\.50\]\), the kurtosis regularizer substantially improves separation\. We match empirical excess kurtosis of residuals to the closed\-form targetκ=−65a4/\(a2/3\+σ2\)2\\kappa=\-\\frac\{6\}\{5\}a^\{4\}/\(a^\{2\}/3\+\\sigma^\{2\}\)^\{2\}from the uniform–Gaussian convolution\.
Table 19:Kurtosis regularizer on noise sweep: reducingGd/GsG\_\{d\}/G\_\{s\}degeneracy\. Avg\. ratio = mean‖Gd‖/‖Gs‖\\\|G\_\{d\}\\\|/\\\|G\_\{s\}\\\|across noise levels \(lower is better\);ΔGd,ΔGs\\Delta G\_\{d\},\\Delta G\_\{s\}= response from lowest to highest noise\. Mean±\\pmstd over 5 seeds\.The kurtosis regularizer reduces theGd/GsG\_\{d\}/G\_\{s\}ratio from0\.920\.92to0\.470\.47\(50% improvement atλ=1\.0\\lambda\{=\}1\.0, mean over 5 seeds\), roughly halvingGdG\_\{d\}’s spurious response to noise \(ΔGd\\Delta G\_\{d\}:2\.44→1\.482\.44\\to 1\.48\) while preservingGsG\_\{s\}’s correct noise tracking \(ΔGs≈4\.2\\Delta G\_\{s\}\\approx 4\.2across all settings\)\. Performance is broadly stable forλ∈\[0\.1,1\.0\]\\lambda\\in\[0\.1,1\.0\]withλ=1\.0\\lambda\{=\}1\.0slightly preferred; full per\-σ\\sigmavalues for theλ=1\.0\\lambda\{=\}1\.0column are in Table[19](https://arxiv.org/html/2608.05454#A4.T19)\. This indicates that theGd/GsG\_\{d\}/G\_\{s\}degeneracy is an addressable optimization challenge, not a fundamental limitation of the representation\.
### D\.8Binary Generator Ablation \(nbn\_\{b\}Sweep\)
We sweepnb∈\{1,2,3,4\}n\_\{b\}\\in\\\{1,2,3,4\\\}with fixednd=1,nq=1n\_\{d\}\{=\}1,n\_\{q\}\{=\}1on ETH/UCY under the standard leave\-one\-out protocol, comparing against an MDN withK=2K\{=\}2components sharing the same encoder\. All models use a transformer encoder withd=64d\{=\}64and are trained for 80 epochs with automatic batch sizing\.
Table 20:Effect of binary generator countnbn\_\{b\}on ETH/UCY \(aggregate over 5 scenes, best\-mode ADE/FDE in meters, mean±\\pmstd over 3 seeds\)\. Increasingnbn\_\{b\}monotonically improves displacement metrics while adding negligible parameters\.The results reveal two key findings\. First, increasingnbn\_\{b\}monotonically improves both ADE and FDE, withnb=4n\_\{b\}\{=\}4achieving20%20\\%lower ADE and28%28\\%lower FDE than the MDN baseline\. The parameter overhead is negligible: going fromnb=1n\_\{b\}\{=\}1tonb=4n\_\{b\}\{=\}4adds only55k parameters \(∼5%\{\\sim\}5\\%\)\. Second, the MDN\-K2 baseline—which has the same number of mixture components as HProbZ withnb=1n\_\{b\}\{=\}1—achieves worse ADE \(0\.6040\.604vs\.0\.5490\.549\) despite having comparable parameters, suggesting that HProbZ’s structured decomposition provides an inductive bias that benefits even the simplest configuration\.
### D\.9nuScenes Constraint Refinement
We replicate the constrained sequential refinement experiment of §[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px3)on nuScenes vehicle trajectories\. A Constrained HProbZ withnb=1,nd=1n\_\{b\}\{=\}1,n\_\{d\}\{=\}1and shared generators acrossT=12T\{=\}12future steps is trained alongside an MDN\-K2 baseline on 163k training windows\. To probe the predictedκ\\kappa\-dependence of the contraction rate \(Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)\), we sweep the bounded\-generator log\-bias initialisation over\{0\.7,1\.0,1\.3,1\.6\}\\\{0\.7,1\.0,1\.3,1\.6\\\}, which seedsκ=‖Gd‖/σ\\kappa\{=\}\\\|G\_\{d\}\\\|/\\sigmaat approximately\{1\.5,2\.0,2\.8,3\.7\}\\\{1\.5,2\.0,2\.8,3\.7\\\}\.
Table 21:Constraint refinement on nuScenes \(40k test windows\) across fourκ\\kappainitialisations, mean±\\pmstd over 3 seeds\. Both HProbZ FDE improvement and predictive\-spread reduction grow monotonically withκ\\kappa, exactly as Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)predicts \(O\(1/\(κ2k\)\)O\(1/\(\\kappa^\{2\}k\)\)rate\)\. At the largestκ\\kappa, the spread reduction \(31\.5%31\.5\\%\) exceeds the ETH/UCY headline \(23\.6%23\.6\\%\)\. The MDN row entries are identical across allκ\\kapparows because MDN has noκ\\kappaparameter to sweep; we re\-display the single MDN\-K2 baseline \(\+83\.4±58\.3%\+83\.4\\\!\\pm\\\!58\.3\\%FDE drift, mean±\\pmstd over 3 MDN seeds\) in each row purely for visual contrast against theκ\\kappa\-controlled HProbZ trajectory\. Body §[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px3)cites a 5\-seed update \(3 original \+ 2 added\) with spread reductions12\.312\.3,17\.717\.7,25\.825\.8,33\.2%33\.2\\%; full per\-seed values in the supplementary archive\.The structural advantage observed on ETH/UCY pedestrian data is reproduced on nuScenes vehicle trajectories withκ\\kappa\-matched configuration\. HProbZ’s predictive spread reduction grows monotonically withκ\\kappa, from10\.3%10\.3\\%atκ=1\.5\\kappa\{=\}1\.5to31\.5%31\.5\\%atκ=3\.7\\kappa\{=\}3\.7\(mean over 3 seeds\)—a3\.1×3\.1\\timesamplification controlled by a single initialisation choice and tracking the1/κ21/\\kappa^\{2\}rate of Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)\. The largest setting \(κ≈3\.7\\kappa\{\\approx\}3\.7\) exceeds the ETH/UCY23\.6%23\.6\\%headline\. The single shared MDN\-K2 baseline \(re\-displayed in eachκ\\kapparow of Tab\.[21](https://arxiv.org/html/2608.05454#A4.T21)\) shows no improvement under refinement: its FDE worsens substantially \(\+83%\+83\\%\), since revealing observations shifts the conditioning context but cannot reweight a mixture’s frozen component covariances\. The structural distinction therefore reproduces on a vehicle\-trajectory dataset under the sameκ\\kappa\-controlled protocol, rather than being specific to pedestrian dynamics\.
#### MDN refinement protocol\.
For the MDN\-K2 row in Tab\.[21](https://arxiv.org/html/2608.05454#A4.T21), refinement is implemented by re\-running the encoder with thekkrevealed positions appended to the past context \(i\.e\., conditioning the mixing weights via a fresh forward pass\), then re\-evaluating FDE on the remaining unobserved horizonT−kT\{\-\}k\. The MDN’s within\-component covariances are deterministic outputs of the encoder and so cannot tighten under additional context; what does change is the mixing weight assignment, which can re\-distribute mass toward components whose prior trajectories are inconsistent with the late\-horizon ground truth, producing the\+83%\+83\\%FDE drift\. The 3 seeds give per\-seed FDE drifts of\+111\.4%\+111\.4\\%,\+136\.7%\+136\.7\\%, and\+2\.2%\+2\.2\\%— a wide range reflecting that whether mixing\-weight refinement helps or hurts is itself seed\-dependent \(one of the three seeds happens to land at a near\-zero drift basin\)\. HProbZ’s deterministic posterior solve does not exhibit this seed\-level instability \(per\-seed contractions−30\.8%,−38\.3%,−25\.3%\-30\.8\\%,\-38\.3\\%,\-25\.3\\%forκ=3\.7\\kappa\{=\}3\.7\)\. This is precisely the mechanism flagged in Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)\(iii\) and is consistent with the standard MDN trajectory\-forecasting recipe; we are not claiming the MDN is broken, only that weight\-only refinement is not a substitute for within\-mode contraction\. The same encoder code path is used for HProbZ\.
### D\.10Structural Ablation: Which Generators Matter?
To isolate the contribution of each generator type to HProbZ’s sample efficiency advantage, we train ablated variants on nuScenes with matched parameter counts \(∼295\{\\sim\}295k\): “−Gb\-\\,G\_\{b\}” removes binary generators \(nb=0,nd=1,nq=2n\_\{b\}\{=\}0,n\_\{d\}\{=\}1,n\_\{q\}\{=\}2\), “−Gd\-\\,G\_\{d\}” removes bounded generators \(nb=1,nd=0,nq=2n\_\{b\}\{=\}1,n\_\{d\}\{=\}0,n\_\{q\}\{=\}2\), and “pureGsG\_\{s\}” removes both \(nb=0,nd=0,nq=2n\_\{b\}\{=\}0,n\_\{d\}\{=\}0,n\_\{q\}\{=\}2\)\. All variants use the same encoder and are trained for 100 epochs with identical hyperparameters\.
Table 22:Structural ablation on nuScenes \(minADE in meters, lower is better; all rows trained atd=256d\{=\}256,200200epochs with symmetry\-brokenGbG\_\{b\}init, seed=42\)\.*This table tests sample\-efficiency at fixedKK, not refinement\-under\-observation\.*GbG\_\{b\}is the primary driver of sample efficiency atK≥5K\{\\geq\}5: removing it degradesK≥5K\{\\geq\}5performance by20−26%20\{\-\}26\\%\(K=5K\{=\}5:20%20\\%,K=10K\{=\}10:25%25\\%,K=20K\{=\}20:26%26\\%\)\. RemovingGdG\_\{d\}has minimal sample\-efficiency effect \(≤1%\\leq 1\\%\);GdG\_\{d\}’s contribution is realised only when its shared\-latent constraint is activated by observations — i\.e\., refinement and conformal sets \(Tab\.[6](https://arxiv.org/html/2608.05454#S5.T6)and Tab\.[21](https://arxiv.org/html/2608.05454#A4.T21)\)\. Downstream pipelines that consume only single\-shot best\-of\-KKsamples and never condition on partial observations therefore see no benefit from includingGdG\_\{d\}; we report it as a structural option, not a universal recommendation\. AtK=1K\{=\}1the−Gb\-G\_\{b\}variant beats the full model \(1\.801\.80vs\.2\.112\.11\): without mode separation, the single sample collapses toward the conditional mean — see narrative below\.The ablation reveals a clear division of labor\.GbG\_\{b\}drives sample efficiency: the full model achieves0\.970\.97m atK=20K\{=\}20vs\.1\.221\.22m withoutGbG\_\{b\}\(21%21\\%degradation\), while removingGdG\_\{d\}causes less than1%1\\%degradation \(0\.980\.98m\)\. The mechanism is thatGbG\_\{b\}’s binary modes create well\-separated trajectory clusters, so evenK=2K\{=\}2samples cover both dominant modes\. WithoutGbG\_\{b\}, the model compensates with larger stochastic variance, producing diffuse samples that require many draws to cover the prediction space\. RemovingGbG\_\{b\}helps atK=1K\{=\}1\(1\.801\.80vs\.2\.112\.11\), because the single sample from a model without mode separation is closer to the conditional mean; this reverses forK≥2K\{\\geq\}2once mode coverage matters\.GdG\_\{d\}contributes to constraint refinement\(§[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px3)\), not to sample efficiency, consistent with its role as a bounded drift term that enablesO\(1/k\)O\(1/k\)variance contraction per observation\.
### D\.11Controlled Diffusion Comparison
To directly compare HProbZ against diffusion\-based trajectory prediction under controlled conditions, we train a DDPM model\[[13](https://arxiv.org/html/2608.05454#bib.bib1)\]sharing the same TrajEncoder as HProbZ and MDN\. The denoiser is a temporal transformer with self\-attention and cross\-attention to encoder output,44blocks with adaptive LayerNorm \(DiT\-style\), and sinusoidal timestep conditioning \(1\.44M parameters\)\. We use a cosine noise schedule\[[25](https://arxiv.org/html/2608.05454#bib.bib11)\]withT=200T\{=\}200diffusion steps and DDIM\[[34](https://arxiv.org/html/2608.05454#bib.bib12)\]sampling with5050steps and stochasticityη=0\.5\\eta\{=\}0\.5\.
#### Hyperparameter details and a sampling\-step sweep\.
We trained the DDPM with AdamW \(lr3×10−43\{\\times\}10^\{\-4\}, weight decay1×10−41\{\\times\}10^\{\-4\}, batch10241024\) for200200epochs under cosine LR decay, matching the v3 controlled\-comparison recipe; trainingϵ\\epsilon\-MSE decreased monotonically from0\.0910\.091to0\.0740\.074over the schedule\. To verify that the published5050\-step DDIM number is not an artefact of an arbitrary sampling\-step choice, we re\-evaluated the same trained model at three step counts:
Table 23:DDPM sampling\-step sweep on nuScenes \(single trained model,ϵ\\epsilon\-prediction\)\. minADE/minFDE in meters\. More sampling steps do not improve performance, so the published 50\-step setting is a reasonable middle ground rather than a tuned best\.The three settings differ by less than0\.130\.13m at everyKK, with2525\-step marginally best; the chosen5050\-step configuration sits within0\.030\.03m of the optimum\. We did not separately sweep the noise schedule, denoiser depth, prediction target, or classifier\-free\-guidance scale—additional tuning could narrow the gap to HProbZ, but the∼1\.6\\sim 1\.6–2\.4×2\.4\\timesminADE gap and500×500\\timesinference\-cost gap documented below are too large to be explained by hyperparameter sensitivity alone\.
Table 24:Controlled comparison on nuScenes \(same encoder, same data\)\. Bold marks the column winner\. HProbZ achieves the best minADE at everyKK, with5×5\\timesfewer parameters and500×500\\timesfaster inference than the lightly\-tuned same\-encoder Diffusion baseline; the diffusion noise schedule, denoiser depth, prediction target, and CFG are not separately swept \(App\.[D\.11](https://arxiv.org/html/2608.05454#A4.SS11)\)\. On FDE@20, MDN\-K2 \(1\.151\.15\) wins the column; HProbZ \(1\.401\.40\) ties MDN\-K4 \(1\.401\.40\)\. The diffusion FDE \(6\.406\.40\) reflects itsϵ\\epsilon\-MSE training objective rather than a metric\-optimised baseline\.The diffusion model with temporal transformer denoising underperforms HProbZ on ADE by2\.6×2\.6\\timesatK=20K\{=\}20\(2\.992\.99vs\.1\.131\.13m\) at4\.9×4\.9\\timesmore parameters\. Even with architectural improvements over flat MLP denoisers, FDE remains poor \(6\.406\.40vs\. HProbZ’s1\.401\.40\), which we attribute to the noise\-prediction loss of DDPM not penalising endpoint precision specifically\. The gap fromK=1K\{=\}1toK=20K\{=\}20is only0\.880\.88m for Diffusion, whereas HProbZ converges byK=5K\{=\}5\(0\.180\.18m gap toK=20K\{=\}20\)\. HProbZ’s inference is500×500\\timesfaster \(124k vs\. 246 samples/s\) due to single\-pass generation; the temporal transformer denoiser’s explicit self\-attention adds overhead relative to HProbZ’s feedforward per\-timestep decomposition\. Under this controlled comparison, the closed\-form HProbZ head is more sample\- and compute\-efficient than iterative denoising on nuScenes; we make no fundamental claim about the relative expressiveness of the two paradigms in general\.
### D\.12ETH/UCY Vanilla Cross\-Method Comparison
To validate the controlled comparison framework on pedestrian trajectory prediction, we train HProbZ \(d=128d\{=\}128,nb=2n\_\{b\}\{=\}2, sym\-broken init, 2\-layer MLP head,200200epochs\), matched\-encoder MDN baselines, and a temporal\-transformer DDPM \(port of the nuScenes Diffusion model in Tab\.[24](https://arxiv.org/html/2608.05454#A4.T24), retrained on ETH/UCY\) on 5\-scene leave\-one\-out \(8\-frame past, 12\-frame future\)\. All rows are seed\-42 with a fresh model trained per held\-out scene; numbers are averaged across the 5 LOO splits\.Note onnbn\_\{b\}:this controlled\-comparison HProbZ usesnb=2n\_\{b\}\{=\}2\(44binary modes\) to match the effective binary budget of the matched\-parameter Diffusion denoiser; the body\-text Social HProbZ \(Tab\.[2](https://arxiv.org/html/2608.05454#S5.T2), Tab\.[26](https://arxiv.org/html/2608.05454#A4.T26)\) usesnb=3n\_\{b\}\{=\}3\(88binary modes\) with social\-attention encoding, which is the configuration we report as our SOTA\-chase number\. To probe the speed/quality trade\-off of iterative denoising, we report Diffusion at the standard5050\-step DDIM regime and at a matched single\-pass inference budget \(55\-step DDIM\)\.
Table 25:Controlled cross\-method comparison on vanilla ETH/UCY \(no Social attention; mean across 5 leave\-one\-out scenes\)\. HProbZ achieves the best FDE@20 and inference latency among the four same\-encoder configurations tested\. Diffusion’s lower ADE at5050DDIM steps is bought at∼530×\\sim 530\\timesHProbZ’s inference cost; at matched single\-pass inference budget \(55DDIM steps\), Diffusion ADE rises by an order of magnitude\. The diffusion baseline is not separately swept over noise schedule, denoiser depth, prediction target, or CFG \(App\.[D\.11](https://arxiv.org/html/2608.05454#A4.SS11)\)\.HProbZ wins on FDE@20 across the four same\-encoder baselines tested \(0\.390\.39vs\. MDN0\.410\.41–0\.420\.42, vs\. Diffusion0\.430\.43\): the binary\-mode head provides tighter endpoint predictions on this controlled comparison, which is the metric that matters for downstream collision avoidance and risk\-aware planning \(cf\. §[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px1)\)\. On ADE, Diffusion at5050\-step DDIM achieves a lower absolute number than HProbZ \(0\.260\.26vs\.0\.310\.31\), but at∼530×\\sim 530\\timesslower inference \(18\.918\.9vs\.10\.110\.1k samples/s\)\. When Diffusion is constrained to a matched single\-pass inference budget \(55\-step DDIM\), ADE rises to2\.522\.52m and FDE to2\.242\.24m—over8×8\\timesand5×5\\timesworse than HProbZ on this configuration\. Among the methods tested, HProbZ is the one that delivers sub\-meter ADE/FDE at sub\-millisecond per\-batch latency; we do not claim this is the only architecture able to do so, since neither the diffusion baseline nor potential goal\-conditioned/equivariant alternatives have been swept here\. The vanilla HProbZ here \(no Social attention\) is roughly1\.6×1\.6\\timesbehind the Social variant in Table[26](https://arxiv.org/html/2608.05454#A4.T26)\(Avg\. ADE0\.1970\.197\); the Social variant’s lead reflects the message\-passing bonus rather than a difference in the HProbZ head itself\. Per\-scene breakdown and matched checkpoints in the supplementary archive\.
### D\.13ETH/UCY Per\-Scene Results
Per\-scene breakdown of Table[2](https://arxiv.org/html/2608.05454#S5.T2)\(main body reports only the average across 5 leave\-one\-out scenes\):
Table 26:Per\-scene minADE@20 / minFDE@20 on ETH/UCY, leave\-one\-out protocol\. Base block \(d=64d\{=\}64, 2\-layer encoder, no social attention, AdamW lr=3×10−43\{\\times\}10^\{\-4\}with 10\-ep linear warmup \+ cosine decay, batch=1024, 600 epochs, mean±\\pmstd over 3 seeds∈\{42,123,456\}\\in\\\{42,123,456\\\}\) is a small\-encoder ablation\. Social block \(d=128d\{=\}128, 3\-layer encoder with 8\-neighbour attention, AdamW lr=3×10−43\{\\times\}10^\{\-4\}, batch=2048, 600 epochs, 5 seeds∈\{42,123,456,789,2026\}\\in\\\{42,123,456,789,2026\\\}\) is the headline matched\-encoder comparison: Social HProbZ vs same\-encoder Social MDN\-K=8\. All numbers are reproducible from the per\-seed results released in the supplementary archive\.#### Reproducibility and seed coverage\.
The Base block \(d=64d\{=\}64\) is a fresh 3\-seed reproduction \(seeds∈\{42,123,456\}\\in\\\{42,123,456\\\},batch=1024\\textsc\{batch\}\{=\}1024, 600 epochs, AdamW lr3×10−43\{\\times\}10^\{\-4\}with 10\-epoch linear warmup and cosine decay\)\. The Social HProbZ row uses 5 seeds∈\{42,123,456,789,2026\}\\in\\\{42,123,456,789,2026\\\},batch=2048\\textsc\{batch\}\{=\}2048, 600 epochs,nb=3n\_\{b\}\{=\}3,nd=1n\_\{d\}\{=\}1,nq=1n\_\{q\}\{=\}1, 8\-neighbour attention\. The Social MDN\-K=8 row uses an identical encoder, optimiser, and schedule, only swapping the HProbZ head for aK=8K\{=\}8Gaussian\-mixture head\. Per\-seed std is≤0\.009\\leq 0\.009on every \(method, scene\) cell, so the matched\-encoder gap \(HProbZ39\.3%\\mathbf\{39\.3\\%\}ADE /20\.7%\\mathbf\{20\.7\\%\}FDE below Social MDN\-K=8\) is well outside seed\-level noise\. All per\-scene results and checkpoints are in the supplementary code archive\.
### D\.14Concept Illustration and Trajectory Visualizations
Figure 9:HProbZ unifies existing prediction set types as special cases of a single parameterization\.Figure 10:ETH/UCY trajectory predictions with structurally interpretable binary\-generator mode centers\.Figure 11:Qualitative visualization of HProbZ’s2nb=82^\{n\_\{b\}\}\{=\}8binary\-mode trajectories on three UNIV leave\-one\-out scenarios \(nb=3,nd=1,nq=1n\_\{b\}\{=\}3,n\_\{d\}\{=\}1,n\_\{q\}\{=\}1\)\. Solid colored lines are mode centers; faint lines are per\-mode samples revealing the stochastic envelope; black dashed is ground truth\. The structured decomposition produces well\-separated modal branches that a Gaussian\-mixture baseline would have to cover with a single inflated covariance\.Figure 12:HProbZ mode visualization on three HOTEL leave\-one\-out scenarios\. The same binary generators produce qualitatively different mode branchings across scenes without supervision, demonstrating input\-dependent structural discovery\.
### D\.15Per\-Scene Generator Decomposition
Table 27:Per\-scene learned decomposition of generator Frobenius norms on ETH/UCY \(multi\-seed mean over 5 seeds, large\-batch protocol\)\. Without any scene\-level supervision, the decomposition aligns with qualitative scene characteristics\.ETH, a busy intersection, exhibits the largest generators across all three types; HOTEL, with well\-defined entrance patterns, shows the lowest noise; and the open walkways \(UNIV, ZARA\) are dominated by stochastic perturbation with minimal bounded drift\.
### D\.16Comparison with Published ETH/UCY Methods
Table 28:ETH/UCY vs published methods \(Avg minADE / minFDE @20, best\-of\-20, meters\)\. Social HProbZ is competitive with recent SOTA on minADE; trails LED/Y\-Net by∼0\.06\\sim 0\.06m on minFDE\. Among the listed methods, only HProbZ provides structured decomposition, conformal coverage, and refinement \(this is a property of the listed set, not a literature\-wide claim\)\.Goal\-conditioned methods \(Y\-Net, LED\) achieve sharper minFDE by conditioning on predicted endpoints but provide neither decomposition nor coverage guarantees; the\(Gb,Gd,Gs\)\(G\_\{b\},G\_\{d\},G\_\{s\}\)head is orthogonal to encoder architecture and can replace any encoder’s Gaussian\-mixture output, including these architectures\.
### D\.17nuScenes Speed\-Regime Decomposition: Setup
The body figure \(Fig\.[3](https://arxiv.org/html/2608.05454#S5.F3)\) summarises the result; this subsection records the protocol\. We stratify the nuScenes test set by past\-observed vehicle speed \(finite differences over the past 1 s window\) into four bins: stationary \(<0\.5<0\.5m/s\), low \(0\.50\.5–55m/s\), high \(55–1515m/s\), and very\-high \(≥15\\geq 15m/s\)\. For each bin we compute the Frobenius norms‖Gb‖,‖Gd‖,‖Gs‖\\\|G\_\{b\}\\\|,\\\|G\_\{d\}\\\|,\\\|G\_\{s\}\\\|averaged over the trainedd=256d\{=\}256HProbZ model’s per\-sample outputs, then plot on log scale\.
For high\-speed vehicles,‖Gb‖\\\|G\_\{b\}\\\|increases by370×370\\timesand‖Gs‖\\\|G\_\{s\}\\\|by94×94\\timesrelative to stationary agents, reflecting the greater modal ambiguity and stochastic noise inherent in fast\-moving traffic\. The bounded generator‖Gd‖\\\|G\_\{d\}\\\|shows a more moderate5\.4×5\.4\\timesincrease, consistent with systematic drift being less speed\-dependent\.
### D\.18Planning Separation Sweep
Figure 13:Planning cost vs\. mode separationΔy\\Delta y\. The structured HProbZ planner transitions sharply atΔy≈4\\Delta y\\approx 4and matches the oracle forΔy≥5\\Delta y\\geq 5, while scalar and MDN planners remain at the conservative SWERVE cost across all separations\.
### D\.19Risk\-Estimate Consistency: Full Sweep
The complete sweep over MDN sample budgets, extending the body’s risk consistency comparison \(§[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px1)\)\. HProbZ’s analytic computation produces deterministic risk estimates regardless of budget; MDN flip rates decrease monotonically withSSbut remain non\-trivial \(0\.9%0\.9\\%\) even atS=100S\{=\}100samples per evaluation\.
Table 29:Risk\-estimate consistency on nuScenes \(5,000 scenarios, 20 seeds\)\. Full sample\-budget sweep\.
### D\.20Trajectory\-Level Closed\-Loop Collision Avoidance: Details
This appendix gives the full protocol behind Tab\.[5](https://arxiv.org/html/2608.05454#S5.T5)\(body\)\.
#### Protocol\.
We use the full nuScenes test split \(N=39,829N\{=\}39\{,\}829vehicle trajectories,T=12T\{=\}12steps at22Hz,66s horizon\) in ego\-centric coordinates\. For each trajectory, a stationary obstacle is placed atobs=z¯\+ξ\\mathrm\{obs\}=\\bar\{z\}\+\\xi, wherez¯\\bar\{z\}is the mean of the ground\-truth trajectory over the evaluation window\[tlo,thi\]=\[3,11\]\[t\_\{\\text\{lo\}\},t\_\{\\text\{hi\}\}\]=\[3,11\]\(i\.e\.1\.51\.5–5\.55\.5s ahead\) andξ∼𝒩\(0,σ2I2\)\\xi\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\_\{2\}\)withσ=2\.5\\sigma=2\.5m\. A ground\-truth collision is defined asmint∈\[tlo,thi\]‖zt−obs‖≤r\\min\_\{t\\in\[t\_\{\\text\{lo\}\},t\_\{\\text\{hi\}\}\]\}\\\|z\_\{t\}\-\\mathrm\{obs\}\\\|\\leq rwithr=2r=2m\. Theσ=2\.5\\sigma=2\.5m jitter produces a33\.6%33\.6\\%empirical positive rate with most samples genuinely borderline \(‖ξ‖∈\[1,4\]\\\|\\xi\\\|\\in\[1,4\]m\), so classification is not trivial\.
A conservative planner brakes when the predictor’s estimated probability of a collision event exceeds thresholdτ\\tau\. Both methods estimate this probability by Monte\-Carlo samplingns=500n\_\{s\}=500full trajectories per test instance and computing the empirical rate at which a sample enters the collision radius at any step in the window:
p^collide=1ns∑k=1ns𝟙\[mint∈\[tlo,thi\]‖z~t\(k\)−obs‖≤r\]\.\\hat\{p\}\_\{\\text\{collide\}\}=\\frac\{1\}\{n\_\{s\}\}\\sum\_\{k=1\}^\{n\_\{s\}\}\\mathbb\{1\}\\\!\\left\[\\min\_\{t\\in\[t\_\{\\text\{lo\}\},t\_\{\\text\{hi\}\}\]\}\\\|\\tilde\{z\}\_\{t\}^\{\(k\)\}\-\\mathrm\{obs\}\\\|\\leq r\\right\]\.\(8\)HProbZ samples are drawn respecting the shared\-latent structure:βk∈\{−1,\+1\}\\beta\_\{k\}\\in\\\{\-1,\+1\\\}andαk∈\[−1,1\]nd\\alpha\_\{k\}\\in\[\-1,1\]^\{n\_\{d\}\}are each sampled once per trajectory \(shared acrosstt\), whileηk,t∼𝒩\(0,Inq\)\\eta\_\{k,t\}\\sim\\mathcal\{N\}\(0,I\_\{n\_\{q\}\}\)is independent per step\. MDN samples draw a mixture component once per trajectory \(shared acrossttby convention\) and then add independent per\-step Gaussian noise\. The planner sweepsτ∈\[0\.01,0\.50\]\\tau\\in\[0\.01,0\.50\]on a2020\-point grid; we report the operating point closest to10%10\\%realized false\-alarm rate\.
#### Why HProbZ wins on trajectory\-level but not pointwise\.
The quantity being estimated isP\(∃t∈W:∥zt−obs∥≤r\)P\(\\exists t\\in W\\\!:\\\!\\\|z\_\{t\}\-\\mathrm\{obs\}\\\|\\leq r\)with\|W\|=9\|W\|=9\. For HProbZ, the nine events are strongly positively correlated through the shared\(β,α\)\(\\beta,\\alpha\): once a trajectory enters the obstacle vicinity at one step it is much more likely to be there at nearby steps, so the per\-step probabilities cannot be multiplied as if independent\. MDN’s per\-step Gaussian noise is independent given the component choice, so joint tail events are underestimated through implicit independence \(the component\-sharing component provides only partial correlation through sharedμt\\mu\_\{t\}\)\. For the single\-step variant\|W\|=1\|W\|=1, no correlation applies and MDN’s greater per\-step flexibility \(K=2K\{=\}2Gaussians with per\-dimension variance\) dominates\.
#### Reproducibility\.
The experiment runs in under1010s on a single consumer GPU for the full test set; the headline numbers are mean±\\pmstd over 3 seeds \(42,123,45642,123,456\), withσAUC≤0\.001\\sigma\_\{\\mathrm\{AUC\}\}\\leq 0\.001on every metric—i\.e\. HProbZ’s5\.75\.7percentage\-point AUC margin over MDN\-K2 is roughly50×50\\timesthe seed\-level standard deviation\. Source code and per\-seed outputs are provided in the supplementary code archive\.
### D\.21Refinement Runtime Analysis
Table[30](https://arxiv.org/html/2608.05454#A4.T30)reports wall\-clock times for forward inference and constraint propagation on an RTX 5090 GPU, mean overB∈\{1,64,512\}B\\in\\\{1,64,512\\\}\. HProbZ’s algebraic constraint update \(mode reweighting \+ precision accumulation per Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)\) runs in∼0\.51\{\\sim\}0\.51ms per observation step, comparable to the full forward pass \(∼0\.59\{\\sim\}0\.59ms\)\. A complete 8\-step refinement adds∼4\.1\{\\sim\}4\.1ms total overhead—modest relative to a single forward pass\. MDN’s Bayesian weight update is faster \(∼0\.15\{\\sim\}0\.15ms\) but provides only weight rebalancing without within\-component covariance contraction \(Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7), part iii\)\.
Table 30:Inference runtime \(ms per call, RTX 5090\)\. Constraint propagation is in the same order as a forward pass, with 8 observations adding∼4\{\\sim\}4ms total\.
### D\.22MDN Baseline Hyperparameter Sweep
To ensure fair comparison, we swept MDN hyperparameters on nuScenes: covariance type \(diagonal vs\. full Cholesky\), learning rate∈\{10−4,3×10−4,10−3\}\\in\\\{10^\{\-4\},3\{\\times\}10^\{\-4\},10^\{\-3\}\\\}, componentsK∈\{2,4,8\}K\\in\\\{2,4,8\\\}, and weight decay∈\{0,10−4,10−2\}\\in\\\{0,10^\{\-4\},10^\{\-2\}\\\}\. Table[31](https://arxiv.org/html/2608.05454#A4.T31)reports the best configuration per\(K,cov\)\(K,\\text\{cov\}\)combination\. Two findings stand out: \(i\) the best MDN \(diagonal,K=2K\{=\}2, ADE@5==2\.05\) does not close the gap to HProbZ \(ADE@5==1\.58\), confirming that the advantage stems from the structured output parameterization rather than baseline under\-tuning; \(ii\) full Cholesky covariance consistently degrades ADE relative to diagonal, indicating that naively adding cross\-dimensional correlation parameters hurts optimization without improving mode quality—in contrast, HProbZ captures cross\-dimensional structure implicitly throughGbG\_\{b\}without extra per\-component parameters\.
Table 31:MDN hyperparameter sweep on nuScenes \(best\-seed\-of\-3 per config, minADE in meters\)\. Even the best\-tuned MDN with full covariance does not match HProbZ\.∗HProbZ uses a discrete\-uniform\-Gaussian convolution likelihood; its NLL is not on the same scale as the MDN mixture\-density NLLs and is therefore omitted from this table \(see Appendix[C\.6](https://arxiv.org/html/2608.05454#A3.SS6)\)\.†TheK=8K\{=\}8entry shows the non\-collapsed seed; under mode collapse the 3\-seed mean is3\.134±0\.2253\.134\\pm 0\.225ADE@10—see the dedicated K=8 sweep in Appendix[D\.27](https://arxiv.org/html/2608.05454#A4.SS27)\.
### D\.23Pedestrian Crossing Planning Scenario: Details
Full costs by regime\. The structured HProbZ planner recovers the oracle’s decision on Modal\-Safe because it resolves two clearly separated safe modes that the Gaussian\-collapsed planner cannot distinguish from lane\-overlapping mass\.
Table 32:Average cost per episode on 2 000 scenarios \(lower is better\)\.
### D\.24Scaled HProbZ with Symmetry\-Broken Initialisation
Thenb=1n\_\{b\}\{=\}1HProbZ negative log\-likelihood is invariant under the sign flipGb↦−GbG\_\{b\}\\mapsto\-G\_\{b\}\(both produce the same two\-mode mixture\)\. Standard random initialisation therefore places the model in one of two equivalent basins at random; empirically we observe high seed variance in the resulting minADE—one seed in our initial 5\-seed sweep landed in an exceptionally good basin \(1\.291\.29\) while four others clustered near1\.521\.52–1\.561\.56\(mean1\.48±0\.101\.48\\pm 0\.10\)\. Because the symmetry is structural, scaling the encoder alone does not fix it\.
#### Fix\.
We break the symmetry deterministically at initialisation by adding a positive biasb0b\_\{0\}to theGbG\_\{b\}output channels of the prediction head:
bias\[t,d,chan\(Gb\)\]\+=b0,\\text\{bias\}\[t,d,\\text\{chan\}\(G\_\{b\}\)\]\\mathrel\{\+\}=b\_\{0\},which commits early forward passes toGb\>0G\_\{b\}\>0before gradient descent sees the data\. We sweepb0∈\{0\.5,1\.0\}b\_\{0\}\\in\\\{0\.5,1\.0\\\}; both work, withb0=1\.0b\_\{0\}\{=\}1\.0giving lower mean and tighter std \(Table[33](https://arxiv.org/html/2608.05454#A4.T33)\)\. Combined with a slower learning rate \(2×10−42\\\!\\times\\\!10^\{\-4\}vs\.3×10−43\\\!\\times\\\!10^\{\-4\}\), longer warmup \(1010epochs\), and longer training \(150150epochs\), the 14\-seed minADE@10 on the full nuScenes test set converges to1\.29±0\.07\\mathbf\{1\.29\\pm 0\.07\}atb0=1\.0b\_\{0\}\{=\}1\.0\. Ablation:b0=2\.0b\_\{0\}\{=\}2\.0over\-biases good seeds and is worse on average; we therefore reportb0=1\.0b\_\{0\}\{=\}1\.0in the body\.
Table 33:Effect of sym\-broken init strength on nuScenes minADE@10 \(d=256,4d\{=\}256,4\-layer encoder\)\. Both sym\-break sweeps use identical architecture and optimiser; only theGbG\_\{b\}bias init differs\. Theb0=1\.0b\_\{0\}\{=\}1\.0row is what the body reports\.
#### Per\-seed breakdown \(b0=1\.0b\_\{0\}\{=\}1\.0\)\.
Fourteen seeds give the per\-seed minADE@10 values listed in Table[34](https://arxiv.org/html/2608.05454#A4.T34); the body reports the corresponding1414\-seed mean±\\pmstd \(1\.29±0\.071\.29\\pm 0\.07\) without exclusions\. Multi\-seed results are released alongside the code in the supplementary archive\.
Table 34:Per\-seed minADE@10 on nuScenes,b0=1\.0b\_\{0\}\{=\}1\.0symmetry\-broken initialisation, 14 seeds\.
#### Why it works\.
The positive bias addsb0⋅𝟏b\_\{0\}\\cdot\\bm\{1\}to theGbG\_\{b\}channels at init, so the pre\-training forward pass producesGb≈b0𝟏\>0G\_\{b\}\\approx b\_\{0\}\\bm\{1\}\>0regardless of encoder output; subsequent gradient updates moveGbG\_\{b\}away but remain in the positive\-GbG\_\{b\}basin with high probability\. This is a structural \(not stochastic\) symmetry break: no seed dependence in the direction chosen\. The remaining seed\-to\-seed variance reflects ordinary optimisation noise within the fixed basin; the1414\-seed mean±\\pmstd \(1\.29±0\.071\.29\\pm 0\.07\) we report includes all seeds\.
### D\.25Rolling Closed\-Loop Collision Test on nuScenes
The single\-shot test of §[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px1)measures one brake/no\-brake decision per scenario\. A real planner makes this decision repeatedly as time advances\. We extend the test to a rolling rollout: at each stepτ∈\{0,1,2,3\}\\tau\\in\\\{0,1,2,3\\\}the planner observes a sliding55\-step window\[past\[τ:\]∥future\[:τ\]\]\[\\mathrm\{past\}\[\\tau\{:\}\\,\]\\,\\\|\\,\\mathrm\{future\}\[\{:\}\\tau\]\], runs the predictor on the rest of the future, computes the collision probability against the same fixed obstacle, and makes a brake/no\-brake decision\. We track two new metrics that the single\-shot test cannot measure: \(i\) the number of brake/no\-brake decision flips per scenario over the rollout, and \(ii\) the false\-alarm rate at the chosen operating point\.
Table 35:Rolling closed\-loop collision test on5,0005\{,\}000nuScenes scenarios over33seeds, brake thresholdτ=0\.10\\tau\{=\}0\.10\. Decision flips are the unique contribution of the rolling formulation; HProbZ flips less than half as often as MDN\-K2, consistent with its shared\(β,α\)\(\\beta,\\alpha\)inducing temporally coherent predictions\.At a fixed brake threshold of0\.100\.10, HProbZ matches MDN\-K2’s miss rate \(32\.5%32\.5\\%vs33\.3%33\.3\\%\) at half the false\-alarm rate \(6\.5%6\.5\\%vs13\.0%13\.0\\%\) and—most relevant for the rolling formulation—less than half the decision flips \(0\.200\.20vs0\.490\.49per scenario\)\. The flip metric is what the single\-shot test cannot capture: it reflects the temporal coherence of HProbZ’s predictions as the rollout window slides, which directly determines how often a real planner would oscillate between braking and proceeding\. The matched\-FAR comparison from the single\-shot test \(HProbZ6\.73%6\.73\\%vs MDN9\.70%9\.70\\%collisions at10%10\\%FAR, §[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px1)\) remains the headline; this rolling test adds the temporal\-stability dimension\. Per\-scenario results archive provided in the supplementary code\.
### D\.26nuPlan Zero\-Shot Cross\-Dataset Probe \(Safety\-Critical Subset\)
*Scope statement\.*This appendix tests*zero\-shot transfer*of nuScenes\-trained predictors to nuPlan logs, restricted to safety\-critical close\-proximity scenarios where the structural near\-field claim is operationally relevant\. We do*not*claim general cross\-dataset performance: full\-mix AUC sits near chance \(0\.5190\.519, see below\), and the structural advantage emerges only on the safety\-critical subset\. We tested the nuScenes\-trained predictors on500500scenarios from the nuPlan mini split\[[7](https://arxiv.org/html/2608.05454#bib.bib39)\]: real logs from Boston, Pittsburgh, Singapore, and Las Vegas\. At every11s decision step we extract each vehicle’s22s past in its own agent frame, run the predictor, and aggregate into a per\-step collision probability against a straight\-line ego reference \(66s horizon,2\.52\.5m buffer\)\. The ground\-truth label is whether any agent’s logged trajectory enters the buffer in the same window\. After filtering to agents within66m of the ego reference, we collect14,65314\{,\}653decision\-agent pairs \(3,4013\{,\}401positives\)\.
Table 36:nuPlan zero\-shot collision\-warning AUC over500500scenarios\.*Full\-mix transfer is essentially negative*: HProbZ AUC0\.5190\.519is within22pp of chance, MDN\-K20\.4920\.492is below chance\. The\+0\.159\+0\.159gap is recovered only on the safety\-critical subset \(1,967 / 14,653 pairs,13\.4%13\.4\\%\) whose nuPlan scenario tag indicates close\-proximity vehicle interaction \(stationary\_in\_traffic,near\_long\_vehicle,near\_high\_speed\_vehicle\)\. The pattern is consistent with a structural near\-field claim, but readers should not interpret these numbers as evidence of general cross\-dataset generalisation\. AUCs are trapezoidal\-integrated ROC over the per\-decision predicted collision probabilities\.#### Interpretation\.
On the full mix of nuPlan driving*both predictors are essentially uninformative*: HProbZ AUC0\.5190\.519is within22pp of chance, MDN\-K2 AUC0\.4920\.492is below chance\. The\+0\.027\+0\.027AUC gap on14,65314\{,\}653pairs is small and not, on its own, evidence that HProbZ generalises cross\-dataset\. Restricting to the safety\-critical subset where another vehicle is physically close to the ego, HProbZ’s AUC reaches0\.5470\.547\(just above chance\) and the gap grows to\+0\.159\\mathbf\{\+0\.159\}\. The pattern is consistent with the paper’s structural\-decomposition thesis—the bounded generator produces its sharpest distributions on near\-field interactions where a safety planner needs tight calibration—but the absolute AUC of0\.5470\.547is modest and the subset is post\-hoc, so we read this as suggestive evidence that the structural advantage survives on the slice where it should, not as a cross\-dataset generalisation result\. Per\-step results archive provided in the supplementary code; a larger2,0002\{,\}000\-scene replication reproduces the same pattern \(all\-scenes gap\+0\.054\+0\.054, safety\-critical gap\+0\.123\+0\.123\)\.
Scope caveat: we use nuPlan as the scene source only; our kinematic ego model does not include reactive steering or lane\-change controllers, so we do not report end\-to\-end closed\-loop collision rates\. Full integration with nuPlan’s reactive\-agents simulator is left to future work\.
### D\.27MDN\-K=8K\{=\}8Baseline: Mode\-Collapse Regression
To assess whether scaling MDN to a largerKKmight close the gap to HProbZ, we trained a flat MDN withK=8K\{=\}8Gaussian components on nuScenes using the same encoder, optimiser, and schedule as the MDN\-K2/K4 entries in Table[3](https://arxiv.org/html/2608.05454#S5.T3)\(d=128d\{=\}128,100100epochs, AdamW, batch10241024\)\. To guard against under\-tuning, we additionally swept three learning rates \(1×10−41\{\\times\}10^\{\-4\},3×10−43\{\\times\}10^\{\-4\},1×10−31\{\\times\}10^\{\-3\}\) with three seeds each,99runs total:
Table 37:MDN\-K=8K\{=\}8on nuScenes across three learning rates \(3 seeds each\), versus the existing MDN\-K2 and HProbZ entries from Table[3](https://arxiv.org/html/2608.05454#S5.T3)\. Even the best\-tuned LR regresses by∼1\\sim 1m below MDN\-K2, consistent with classical mode\-collapse: more components compete for limited data and produce dead modes\.The best\-tuned MDN\-K8 run \(lr=10−3=10^\{\-3\}\) reaches3\.045±0\.2323\.045\\pm 0\.232minADE@10—a0\.050\.05m improvement over the LR=3×10−4=3\{\\times\}10^\{\-4\}entry but still1\.211\.21m worse than the matched\-encoder MDN\-K2 baseline \(1\.831\.83\)\. The46%46\\%regression vs\.K=2K\{=\}2is robust across the LR sweep, ruling out under\-tuning as an explanation\. Inspecting the trained mixture weights confirms the diagnosis: in99–1111of the1212future timesteps, more than half of the88component weights collapse to<10−3<10^\{\-3\}, leaving22–44effective components\. This is the classical MDN pathology under limited data and a shared encoder\. By contrast, HProbZ’s2nb2^\{n\_\{b\}\}binary modes share generators by construction and cannot collapse in this way: the binary head’s positive\-bias initialisation \(Appendix[D\.24](https://arxiv.org/html/2608.05454#A4.SS24)\) keeps every mode active throughout training\. Per\-seed results across all three learning rates are provided in the supplementary code\.
#### Scope caveat\.
We do*not*deploy known mode\-collapse mitigations \(e\.g\., mode dropout, entropy regularisers, annealed mixing\) in thisK=8K\{=\}8sweep\. Whether such mitigations would close the gap is an open empirical question; the structural distinction we focus on is that HProbZ’s2nb2^\{n\_\{b\}\}binary modes share generators by construction and cannot collapse irrespective of mitigation\. TheK=2K\{=\}2entries elsewhere in the paper use the same MDN configuration without any mode\-collapse mitigation, so the within\-paper comparison is consistent\.
### D\.28Learnable Mode Prior Ablation
The main paper uses a uniform mixture priorπk=2−nb\\pi\_\{k\}=2^\{\-n\_\{b\}\}over the2nb2^\{n\_\{b\}\}binary configurations ofβ\\beta\. To check that this design choice does not silently leave headline accuracy on the table, we add a per\-sample logistic gating head,πk\(x\)=softmax\(Wπh\)k\\pi\_\{k\}\(x\)=\\mathrm\{softmax\}\(W\_\{\\pi\}h\)\_\{k\}, on top of the same encoder, train under the same recipe, and re\-run the 14\-seed nuScenes sweep\. Identifiability and contraction \(Theorems[3](https://arxiv.org/html/2608.05454#Thmtheorem3)and[7](https://arxiv.org/html/2608.05454#Thmtheorem7)\) carry over verbatim as long asπk\(x\)\>0\\pi\_\{k\}\(x\)\>0, which softmax guarantees\.
Table 38:14\-seed nuScenes minADE@10 with learnable input\-dependent mixture priorπk\(x\)\\pi\_\{k\}\(x\)versus the uniformπk=2−nb\\pi\_\{k\}=2^\{\-n\_\{b\}\}used in the main paper\. Same encoder \(d=256d\{=\}256, 4 layers,2\.132\.13M parameters\), same training recipe \(150150epochs, lr2×10−42\{\\times\}10^\{\-4\},b0=1\.0b\_\{0\}\{=\}1\.0symmetry break\)\. Both still beat the map\-aware diffusion baseline MID at1\.441\.44\.The learnable prior shifts mean minADE by\+0\.025\+0\.025m, well within one standard deviation of the uniform\-prior baseline, and all 14 seeds still beat MID\. Two findings stand out\. First, the design choice is empirically validated: the uniform prior is not leaving accuracy on the table\. Second, the per\-seed standard deviation halves \(0\.068→0\.0380\.068\\to 0\.038\), suggesting that learnable gating absorbs some of the basin\-selection variance otherwise handled by the symmetry\-breaking initialisation\. Per\-seed results are released in the supplementary code\.
## Appendix EBroader Impact and Ethical Considerations
Predictive distributions over agent trajectories are used by downstream planners that take real\-world actions, making the calibration and structure of those distributions safety\-relevant\. We discuss the intended benefits, the downstream risks of deployment, and dataset\-level caveats that a practitioner should weigh before using HProbZ in a real system\.
#### Detailed limitations \(extending body §[7](https://arxiv.org/html/2608.05454#S7)\)\.
\(a\) Mode prior and generator parameterisation: the main paper uses a uniform mode priorπk=2−nb\\pi\_\{k\}\{=\}2^\{\-n\_\{b\}\}and diagonalGd,GsG\_\{d\},G\_\{s\}matrices\. Replacingπk\\pi\_\{k\}with a learnable input\-dependent gateπk\(x\)\>0\\pi\_\{k\}\(x\)\{\>\}0preserves identifiability and contraction, with statistically tied accuracy on a 14\-seed nuScenes sweep \(App\.[D\.28](https://arxiv.org/html/2608.05454#A4.SS28)\); non\-diagonalGd,GsG\_\{d\},G\_\{s\}are left to future work\. \(b\) Theorem[7](https://arxiv.org/html/2608.05454#Thmtheorem7)relaxation: the proof uses a Gaussian relaxationα∼𝒩\(0,13I\)\\alpha\\sim\\mathcal\{N\}\(0,\\tfrac\{1\}\{3\}I\)variance\-matched to the trainedUniform\(\[−1,1\]\)\\mathrm\{Uniform\}\(\[\-1,1\]\)prior; a 1\-D numerical check \(Fig\.[4](https://arxiv.org/html/2608.05454#A2.F4)\) bounds the relative gap between relaxed and true uniform\-prior posterior variances by≤9\.03%\\leq 9\.03\\%on the\(κ,k\)\(\\kappa,k\)grid we test, but a closed\-form bound is open\. \(c\) Raw NLL: HProbZ’s per\-axis compact\-support density incurs a misspecification penalty on data with unbounded tails, separate from the entropy\-bound of Prop\.[5](https://arxiv.org/html/2608.05454#Thmtheorem5)\(App\.[C\.6](https://arxiv.org/html/2608.05454#A3.SS6)\); we therefore lead with ADE/FDE and downstream metrics\. \(d\) Closed\-loop scope: trajectory\-level30%30\\%collision reduction is on a static\-obstacle, non\-reactive\-ego stress test \(N=39,829N\{=\}39\{,\}829\), and reverses on the per\-step variant since MDN’s per\-step independent Gaussian wins single\-step events while HProbZ’s shared\(β,α\)\(\\beta,\\alpha\)wins joint\-event tail coverage\. \(e\) Cross\-dataset transfer: scope is by design \(close\-proximity safety\-critical prediction\); on nuPlan’s full agent mix that includes static, parked, and far\-field agents whose behaviour is not bounded\-drift, zero\-shot AUC settles near chance, while the close\-proximity safety\-critical subset preserves the structural advantage \(App\.[D\.26](https://arxiv.org/html/2608.05454#A4.SS26)\)\. \(f\) Encoder coupling: atd=64d\{=\}64single\-agent the bare HProbZ head trails CVAE on ADE; the advantage emerges with the social\-attention encoder used by recent SOTA, leaving head–encoder interaction open \(App\.[D\.12](https://arxiv.org/html/2608.05454#A4.SS12)\)\.
#### Intended benefits\.
Any planner that conditions on a learned predictive distribution must reason about where other agents are likely to be, and under what uncertainty\. The current standard output—a Gaussian mixture—entangles modal ambiguity, bounded drift, and stochastic noise, which complicates both risk computation and constraint satisfaction in downstream planners\. HProbZ separates these three sources in a single forward pass, makes analytic per\-mode collision probability available for free, and provides distribution\-free multi\-modal conformal sets that remain valid under the cross\-scene distribution shifts tested in Appendix[D\.5](https://arxiv.org/html/2608.05454#A4.SS5)\. In the closed\-loop evaluation of §[5\.4](https://arxiv.org/html/2608.05454#S5.SS4.SSS0.Px1), this translated into a30%30\\%reduction in realised collisions at matched false\-alarm rate against an MDN baseline\. We do*not*claim that this number transfers to deployed risk in production safety\-critical stacks: the test uses a synthetic obstacle configuration rather than a reactive multi\-agent simulator, and is a necessary but not sufficient condition for safer downstream behaviour\. Integration with reactive simulators and full verification & validation pipelines remains future work\.
#### Deployment risks and required safeguards\.
None of our results should be taken as a certificate of safety in a real system\. Three caveats apply\.\(i\) Evaluation gap\.Our closed\-loop test uses a synthetic obstacle placed from ground\-truth trajectories, not a reactive multi\-agent simulator such as nuPlan or CARLA\. The qualitative structural win—joint tail coverage via shared latents—is likely to transfer; the specific30%30\\%number is not a guarantee\.\(ii\) OOD behaviour\.Conformal coverage degrades gracefully under the distribution shifts we tested, namely ETH/UCY leave\-one\-out and nuScenes speed strata, but coverage is not guaranteed under adversarial or out\-of\-support inputs such as construction zones, unusual road geometry, or rare agent classes\. HProbZ is not a substitute for validated OOD detection\.\(iii\) Calibration drift\.The sym\-broken initialisation we recommend in Appendix[D\.24](https://arxiv.org/html/2608.05454#A4.SS24)improves seed\-level stability; it does not address sensor drift or calibration changes between training and deployment\. We recommend conformal recalibration on a rolling held\-out window when deployed\.
#### Dataset composition and demographic caveats\.
Both ETH/UCY and nuScenes are urban, Western, daylight\-biased\. ETH/UCY over\-represents university campus pedestrian behaviour; nuScenes captures Boston and Singapore mostly under dry, daytime conditions\. Argoverse 2 is more geographically diverse but still U\.S\.\-centric\. A model trained on any of these risks under\-performing on pedestrian or vehicle behaviour patterns under\-represented in training \(e\.g\., pedestrians using mobility aids, vehicle types common in other regions, inclement weather\)\. Before deploying HProbZ in a new region, we recommend re\-training or at minimum re\-calibrating the conformal threshold on a representative sample of that region’s trajectories\.
#### Dual\-use and surveillance\.
Trajectory prediction models can also be used to track and anticipate individuals’ movements, and stronger predictors enable stronger surveillance\. HProbZ’s core technical contributions—structured uncertainty decomposition, analytic per\-mode risk, and observation\-driven contraction—do not specifically strengthen surveillance capabilities beyond what a well\-tuned MDN already provides\. We encourage practitioners and platform operators to treat trajectory forecasting systems as personal data processing and to scope data retention, access control, and purpose\-limitation accordingly\.
#### Environmental cost\.
See Appendix[A\.5](https://arxiv.org/html/2608.05454#A1.SS5)for compute disclosure \(∼4\.2\\sim 4\.2GPU\-hours for the 14\-seed nuScenes sweep on one H200; smaller for Argoverse 2 and ETH/UCY\)\. We consider this a negligible fraction of typical large\-model training budgets; code and checkpoints are released to avoid repeated retraining\.
#### Human subjects and consent\.
We use pre\-existing public benchmarks—nuScenes, Argoverse 2, and ETH/UCY—under their respective research licences\. No additional human data were collected for this work\. The ETH/UCY pedestrian recordings pre\-date contemporary consent norms; we follow the research community’s standard practice of aggregate evaluation only and publish no individual trajectories\.相似文章
不确定性的一致风险视角:用于解耦和超越代理评估的后验风险
本文提出了一种将不确定性统一定义为逐点后验风险的方法,并引入了一个有理论支持的基准,利用半合成数据集直接计算神谕认知不确定性和偶然不确定性,从而超越代理任务实现细粒度评估。
几何感知的神经算子事后不确定性量化
提出REEF-GP,一种事后不确定性量化框架,通过将高斯过程拟合到冻结神经算子的残差上并利用其内部嵌入,以低成本实现几何感知且校准的不确定性。
具有可处理不确定性量化的结构保持神经替代模型
本文提出了针对偏微分方程的结构保持神经替代模型,该模型集成了Gaussian process regression以提供可处理的不确定性量化,从而能够实现具有闭式误差估计的实时仿真。
神经算子的共形预测:物理模拟中的无分布不确定性量化
提出了将分裂共形预测首次应用于基于神经算子的物理模拟,提供了具有有限样本覆盖保证的无分布预测区间,并利用MC Dropout不确定性生成自适应宽度的区间。
从混合机理-数据驱动建模到神经符号人工智能:是什么、为什么以及如何实现
本文介绍了Hybrid-to-NeSy (H2N)框架,该框架系统地将混合机理-数据驱动模型转化为神经符号人工智能设计,从而能够推导出结构违规率和信念离散度等指标,作为机理部分认知不确定性的度量。