Hidden Axis of Uncertainty: Latent-Posterior Alignment in Graph Neural Networks with Bayesian Output Layers

arXiv cs.LG Papers

Summary

The paper introduces Latent-Posterior Alignment to explain uncertainty reduction in Graph Neural Networks with Bayesian output layers and proposes Alignment-Guided Learning to improve uncertainty quantification and model calibration.

arXiv:2608.20758v1 Announce Type: new Abstract: Bayesian Neural Networks (BNNs) with Bayesian output layers provide a principled and tractable framework for quantifying predictive uncertainty, yet the mechanisms shaping that uncertainty remain unclear. While conventional theory attributes uncertainty reduction to posterior contraction, the corresponding assumptions need not hold for deep models. In the Graph Neural Networks (GNNs) with Bayesian output layers studied here, we observe that predictive uncertainty decreases as latent representations shift toward lower-variance posterior directions, even though the posterior variance does not contract. We term this behavior Latent-Posterior Alignment (LPA) and conduct interventional experiments that support its functional role in shaping predictive uncertainty. Building on this insight, we propose Alignment-Guided Learning (AGL), which explicitly promotes this alignment during training. AGL effectively reduces predictive uncertainty while preserving accuracy and improves structural calibration, ensuring that the model confidence faithfully mirrors underlying data density. These findings provide a new perspective on uncertainty dynamics in GNNs with mean-field Bayesian output layers, shifting the focus from the magnitude of the posterior to the geometric interplay between latent and parameter spaces.
Original Article
View Cached Full Text

Cached at: 08/24/26, 04:33 AM

# Hidden Axis of Uncertainty: Latent-Posterior Alignment in Graph Neural Networks with Bayesian Output Layers
Source: [https://arxiv.org/html/2608.20758](https://arxiv.org/html/2608.20758)
Bayesian Neural Networks \(BNNs\) with Bayesian output layers provide a principled and tractable framework for quantifying predictive uncertainty, yet the mechanisms shaping that uncertainty remain unclear\. While conventional theory attributes uncertainty reduction to posterior contraction, the corresponding assumptions need not hold for deep models\. In the Graph Neural Networks \(GNNs\) with Bayesian output layers studied here, we observe that predictive uncertainty decreases as latent representations shift toward lower\-variance posterior directions, even though the posterior variance does not contract\. We term this behavior Latent\-Posterior Alignment \(LPA\) and conduct interventional experiments that support its functional role in shaping predictive uncertainty\. Building on this insight, we propose Alignment\-Guided Learning \(AGL\), which explicitly promotes this alignment during training\. AGL effectively reduces predictive uncertainty while preserving accuracy and improves structural calibration, ensuring that the model confidence faithfully mirrors underlying data density\. These findings provide a new perspective on uncertainty dynamics in GNNs with mean\-field Bayesian output layers, shifting the focus from the magnitude of the posterior to the geometric interplay between latent and parameter spaces\.

Suk Hoon ChoiEmail:[inuk97@kist\.re\.kr](mailto:[email protected])Affiliation:Clean Energy Research Center, Korea Institute of Science and Technology, Seoul 02792, Republic of KoreaDamdae ParkEmail:[damdaepark@kist\.re\.kr](mailto:[email protected])Affiliation:Clean Energy Research Center, Korea Institute of Science and Technology, Seoul 02792, Republic of KoreaJunhyuk ChoiEmail:[wnsgurchl@korea\.ac\.kr](mailto:[email protected])Affiliation:Department of Chemical and Biological Engineering, Korea University, Seoul 02841, Republic of KoreaHyein JungEmail:[976hannah@kist\.re\.kr](mailto:[email protected])Affiliation:Clean Energy Research Center, Korea Institute of Science and Technology, Seoul 02792, Republic of KoreaChangsoo KimEmail:[changs90\.kim@kist\.re\.kr](mailto:[email protected])Affiliation:Clean Energy Research Center, Korea Institute of Science and Technology, Seoul 02792, Republic of KoreaAffiliation:Division of Energy & Environment Technology, Korea University of Science and Technology, 217 Gajeong\-ro, Yuseong\-gu, Daejeon 34113, Republic of KoreaUng LeeEmail:[ulee@kist\.re\.kr](mailto:[email protected])Affiliation:Clean Energy Research Center, Korea Institute of Science and Technology, Seoul 02792, Republic of KoreaAffiliation:Division of Energy & Environment Technology, Korea University of Science and Technology, 217 Gajeong\-ro, Yuseong\-gu, Daejeon 34113, Republic of KoreaKyeongsu KimEmail:[kyeongsu@kist\.re\.kr](mailto:[email protected])Affiliation:Clean Energy Research Center, Korea Institute of Science and Technology, Seoul 02792, Republic of KoreaAffiliation:Division of Energy & Environment Technology, Korea University of Science and Technology, 217 Gajeong\-ro, Yuseong\-gu, Daejeon 34113, Republic of Korea

###### keywords

Graph neural networks, Bayesian output layer, uncertainty, posterior, latent\-posterior alignment, alignment\-guided learning

## 1Introduction

Quantifying predictive uncertainty has become essential as deep learning systems are increasingly deployed in scientific and engineering domains where data are scarce[10](https://arxiv.org/html/2608.20758#bib.bib70);[11](https://arxiv.org/html/2608.20758#bib.bib72)\. The challenges are particularly pronounced in safety\-critical and scientific applications, such as autonomous driving[8](https://arxiv.org/html/2608.20758#bib.bib58);[12](https://arxiv.org/html/2608.20758#bib.bib59);[77](https://arxiv.org/html/2608.20758#bib.bib60), medical diagnostics[17](https://arxiv.org/html/2608.20758#bib.bib61);[48](https://arxiv.org/html/2608.20758#bib.bib62), and accelerated material discovery[33](https://arxiv.org/html/2608.20758#bib.bib56);[63](https://arxiv.org/html/2608.20758#bib.bib57);[49](https://arxiv.org/html/2608.20758#bib.bib55)\. To address these challenges, Bayesian Neural Networks \(BNNs\) have emerged as a principled framework for modeling the uncertainty of deep learning architectures[55](https://arxiv.org/html/2608.20758#bib.bib63);[42](https://arxiv.org/html/2608.20758#bib.bib71);[32](https://arxiv.org/html/2608.20758#bib.bib64)\. Common approximate Bayesian approaches include Mean\-Field Variational Inference \(Bayes\-by\-Backprop\)[6](https://arxiv.org/html/2608.20758#bib.bib3)and Monte Carlo dropout[22](https://arxiv.org/html/2608.20758#bib.bib65), while Deep Ensembles[39](https://arxiv.org/html/2608.20758#bib.bib66)provide a widely used non\-Bayesian alternative\. These approaches share the practical goal of quantifying predictive uncertainty, or “what a model does not know”[34](https://arxiv.org/html/2608.20758#bib.bib4);[29](https://arxiv.org/html/2608.20758#bib.bib35)\.

Conventionally, the reduction of predictive uncertainty in Bayesian models is attributed to the contraction of the posterior distribution\. Grounded in the classical Bernstein\-von Mises theorem[69](https://arxiv.org/html/2608.20758#bib.bib38);[21](https://arxiv.org/html/2608.20758#bib.bib39), it is assumed that as data accumulates, the posterior asymptotically converges to a Gaussian distribution centered at the true parameter, causing the variance to vanish[6](https://arxiv.org/html/2608.20758#bib.bib3);[74](https://arxiv.org/html/2608.20758#bib.bib6);[24](https://arxiv.org/html/2608.20758#bib.bib15)\. However, this theoretical guarantee is predicated on strict assumptions—that the model is well\-specified \(i\.e\., it can perfectly replicate the true data generating process\)[36](https://arxiv.org/html/2608.20758#bib.bib36), the parameter space is finite\-dimensional and fixed[21](https://arxiv.org/html/2608.20758#bib.bib39), and the Fisher information matrix is positive definite[71](https://arxiv.org/html/2608.20758#bib.bib40)—conditions that may not be satisfied in modern deep\-learning settings[13](https://arxiv.org/html/2608.20758#bib.bib50);[7](https://arxiv.org/html/2608.20758#bib.bib37)\. Consequently, the classical assurance that “more data yields tighter posteriors and lower uncertainty” becomes theoretically tenuous in this regime[36](https://arxiv.org/html/2608.20758#bib.bib36);[7](https://arxiv.org/html/2608.20758#bib.bib37);[47](https://arxiv.org/html/2608.20758#bib.bib51)\.

Recent empirical studies further show that uncertainty behavior in deep Bayesian models can depart from posterior\-contraction intuitions\. Analyses of loss landscapes reveal that posteriors in deep BNNs are highly complex and multimodal, fundamentally differing from the simple compact distributions assumed by the classical Bayesian theory[30](https://arxiv.org/html/2608.20758#bib.bib25);[19](https://arxiv.org/html/2608.20758#bib.bib68)\. Furthermore, studies on the “Cold Posterior” effect reveal that BNN posteriors do not exhibit the expected contraction and often necessitate artificial sharpening for generalization, underscoring a fundamental mismatch between Bayesian theory and the behavior of deep models[74](https://arxiv.org/html/2608.20758#bib.bib6);[1](https://arxiv.org/html/2608.20758#bib.bib41);[53](https://arxiv.org/html/2608.20758#bib.bib42)\. Similarly, out\-of\-distribution \(OOD\) benchmarks demonstrate that BNNs frequently exhibit overconfident predictions in regions with low data density, where the uncertainty is expected to be high and the posterior remains diffuse[57](https://arxiv.org/html/2608.20758#bib.bib13);[2](https://arxiv.org/html/2608.20758#bib.bib43);[52](https://arxiv.org/html/2608.20758#bib.bib47);[14](https://arxiv.org/html/2608.20758#bib.bib44);[18](https://arxiv.org/html/2608.20758#bib.bib45)\. This misalignment violates the fundamental correspondence between data density and uncertainty\. Crucially, such failures propagate to downstream applications\. In active learning, uncertainty\-based acquisition strategies frequently fail to outperform simple random sampling, as they do not faithfully reflect data scarcity[76](https://arxiv.org/html/2608.20758#bib.bib49);[62](https://arxiv.org/html/2608.20758#bib.bib46)\. Likewise, in Bayesian optimization, the lack of distance\-awareness obscures informative intermediate regions, directly hindering efficient exploration[20](https://arxiv.org/html/2608.20758#bib.bib48)\.

These limitations are particularly consequential in chemical and materials applications, where predictive models frequently operate under low\-data and distribution\-shifted regimes\. In these domains, uncertainty estimates are not merely auxiliary outputs but often directly influence decision\-making processes such as molecular screening, active learning, and experimental prioritization\. Among the available architectures, graph neural networks \(GNNs\) have become a common backbone for machine\-learning interatomic potentials \(MLIPs\) and molecular property prediction tasks due to their ability to encode atomistic structures and chemical interactions\. Consequently, deficiencies in uncertainty estimation—such as poor density awareness or overconfident predictions in sparse regions—can propagate through downstream workflows and hinder efficient exploration of chemical space\. Motivated by these practical considerations, we focus on Bayesian graph neural networks \(BGNNs\) employing Bayesian output layers and revisit how predictive uncertainty is formed within these architectures\.

Motivated by these theoretical and practical considerations, we revisit the dynamics of uncertainty in BNNs\. We focus on a class of models consisting of a deterministic feature extractor followed by a Bayesian output layer trained via Bayes\-by\-Backprop, and conduct our analysis within this formulation\. Contrary to the conventional expectation that predictive uncertainty decreases through posterior contraction, we observe that predictive uncertainty decreases with increasing data even as the weight posterior variance broadens\. This decoupled behavior calls for a re\-examination of the underlying mechanism\. To this end, we identify a key mechanism, termed Latent–Posterior Alignment \(LPA\), which explains how predictive uncertainty is reduced\. Through analytical derivation and empirical analysis, we show that the model minimizes predictive variance not by tightening its posteriors, but by aligning its latent representation𝐳\\mathbf\{z\}along the low\-variance directions of the posterior\.

We probe the causal role of this mechanism through interventional experiments: structurally disrupting the LPA via latent regularization consistently amplifies predictive uncertainty, supporting a causal contribution of LPA to model confidence\. Building on this observation, we introduce Alignment\-Guided Learning \(AGL\), a training framework that explicitly promotes this alignment\. We show that AGL reduces predictive uncertainty while preserving accuracy and significantly improves structural calibration, illustrating that model confidence more faithfully reflects the underlying data density\. This improved alignment between output uncertainty and data density suggests that AGL can make the model’s predictive confidence more consistent with the local density of the training data\. Such density awareness is particularly valuable in chemical and material domain including molecular modeling and MLIPs, where uncertainty estimates can help identify weakly supported structures and prioritize additional calculations or experiments\. Collectively, our findings provide a new perspective on uncertainty in models with Bayesian output layers, emphasizing the role of LPA over the magnitude of the posterior\.

## 2Results

To investigate the mechanisms determining predictive uncertainty, we designed a Bayesian Graph Neural Network \(BGNN\), as shown in Fig\.[1](https://arxiv.org/html/2608.20758#S2.F1)a\. This architecture combines a deterministic Graph Isomorphism Network \(GIN\)[75](https://arxiv.org/html/2608.20758#bib.bib21);[51](https://arxiv.org/html/2608.20758#bib.bib29), which learns latent representations from molecular structures, with a Bayesian linear output layer for uncertainty quantification[6](https://arxiv.org/html/2608.20758#bib.bib3);[67](https://arxiv.org/html/2608.20758#bib.bib32);[16](https://arxiv.org/html/2608.20758#bib.bib31);[23](https://arxiv.org/html/2608.20758#bib.bib33)\. Previous studies suggest that Bayesian last\-layer models can retain predictive performance and provide uncertainty estimates comparable to more extensively Bayesianized networks in the settings examined, at substantially lower computational cost[28](https://arxiv.org/html/2608.20758#bib.bib73);[78](https://arxiv.org/html/2608.20758#bib.bib74)\. Although this approach does not capture parameter uncertainty in the deterministic feature extractor, it provides a practical and analytically tractable framework for decomposing predictive variance into representation and output\-weight contributions\. We therefore adopt this formulation to investigate uncertainty formation in GNNs with Bayesian output layers\. To ensure the robustness and generality of our observations, we conducted experiments across six molecular property prediction benchmarks\. In the main text, we present results for the partition coefficient dataset[46](https://arxiv.org/html/2608.20758#bib.bib67)as a representative example, while results for the remaining datasets are provided in the Supplementary Information\.

In this section, we present a series of findings that revisit the conventional understanding of uncertainty in this class of BNNs\. We begin by identifying a counterintuitive behavior, where predictive uncertainty decreases with increasing data density despite the expansion of the posterior variance \(Section[2\.1](https://arxiv.org/html/2608.20758#S2.SS1)\)\. To explain this phenomenon, we introduce LPA as the central mechanism\. We show that the geometry of the latent representation—rather than the magnitude of the posterior variance—plays a central role in determining predictive uncertainty \(Section[2\.2](https://arxiv.org/html/2608.20758#S2.SS2)\)\. We then probe the causal contribution of LPA through interventional experiments that perturb the latent\-space structure and demonstrate that this alignment\-based mechanism persists under diverse probabilistic settings \(Section[2\.3](https://arxiv.org/html/2608.20758#S2.SS3)\)\. Finally, we propose Alignment\-Guided Learning \(AGL\), a training strategy that explicitly promotes this geometric alignment, reducing average predictive uncertainty and strengthening density\-uncertainty dependence while preserving predictive accuracy \(Section[2\.4](https://arxiv.org/html/2608.20758#S2.SS4)\)\.

### 2\.1Rethinking uncertainty behavior in Bayesian neural networks

![Refer to caption](https://arxiv.org/html/2608.20758v1/img_1_1.png)

Figure 1:Uncertainty behavior of BGNN\. \(a\) Schematic of the BGNN architecture\. \(b–g\) Deviation from conventional expectation\. As training data increases, predictive error \(b\) and uncertainty decrease \(c\), whereas weight posterior uncertainty expands \(d\), leading to broader distributions in the data\-rich regime \(f, red\) compared to the data\-poor regime \(f, blue\)\. The posterior means remain near zero on average \(e\), while both the NLL and KL terms decrease \(g\)\. Shaded areas represent standard deviation acrossn=30n=30independent runs\. “Training data \[%\]” denotes the percentage of the available training set used to train each model\.Consistent with the Bernstein\-von Mises theorem[69](https://arxiv.org/html/2608.20758#bib.bib38);[21](https://arxiv.org/html/2608.20758#bib.bib39)and fundamental Bayesian theory[4](https://arxiv.org/html/2608.20758#bib.bib1);[45](https://arxiv.org/html/2608.20758#bib.bib2);[25](https://arxiv.org/html/2608.20758#bib.bib7), we first confirm that as the amount of training data increases, the model’s predictive performance \(MAE, Fig\.[1](https://arxiv.org/html/2608.20758#S2.F1)b\) improves and its overall predictive uncertainty monotonically decreases \(σy\\sigma\_\{y\}, Fig\.[1](https://arxiv.org/html/2608.20758#S2.F1)c\)\. Conventionally, this reduction is attributed to posterior contractionp⁡\(θ\|𝒟\)p\(\\theta\|\\mathcal\{D\}\), whereθ\\thetaand𝒟\\mathcal\{D\}denote the model parameters and dataset, respectively\. As empirical evidence accumulates, the posterior variance is theoretically expected to shrink, making the model increasingly confident with regard to its parameter estimates[6](https://arxiv.org/html/2608.20758#bib.bib3);[74](https://arxiv.org/html/2608.20758#bib.bib6);[24](https://arxiv.org/html/2608.20758#bib.bib15)\. Following the standard uncertainty decomposition framework[34](https://arxiv.org/html/2608.20758#bib.bib4);[29](https://arxiv.org/html/2608.20758#bib.bib35);[15](https://arxiv.org/html/2608.20758#bib.bib5), the predictive varianceVar\[y\|x,𝒟\]\\mathrm\{Var\}\\left\[y\|x,\\mathcal\{D\}\\right\], given an inputxxand target outputyy, can be decomposed as:

Var\[y\|x,𝒟\]\\displaystyle\\mathrm\{Var\}\\left\[y\|x,\\mathcal\{D\}\\right\]=𝔼p⁡\(θ\|𝒟\)\[Var\[y\|x,θ\]\]⏟Aleatoric\+Varp⁡\(θ\|𝒟\)\[𝔼\[y\|x,θ\]\]⏟Epistemic\\displaystyle=\\underbrace\{\\mathbb\{E\}\_\{p\(\\theta\|\\mathcal\{D\}\)\}\\left\[\\mathrm\{Var\}\\left\[y\|x,\\theta\\right\]\\right\]\}\_\{\\text\{Aleatoric\}\}\+\\underbrace\{\\mathrm\{Var\}\_\{p\(\\theta\|\\mathcal\{D\}\)\}\\left\[\\mathbb\{E\}\\left\[y\|x,\\theta\\right\]\\right\]\}\_\{\\text\{Epistemic\}\}\(1\)
Here, the first term represents aleatoric uncertainty, while the second captures the epistemic uncertainty arising from uncertainty in model parameters\. Under the standard Bayesian paradigm, the posterior varianceVar⁡\[θ\|𝒟\]\\mathrm\{Var\}\[\\theta\|\\mathcal\{D\}\]is expected to diminish as the dataset grows, causing the epistemic term to vanish[25](https://arxiv.org/html/2608.20758#bib.bib7);[29](https://arxiv.org/html/2608.20758#bib.bib35)\. Thus, the reduction in predictive uncertainty is conventionally understood to result from this tightening of the posterior[6](https://arxiv.org/html/2608.20758#bib.bib3);[74](https://arxiv.org/html/2608.20758#bib.bib6);[24](https://arxiv.org/html/2608.20758#bib.bib15);[29](https://arxiv.org/html/2608.20758#bib.bib35)\. However, many modern deep BNNs do not clearly satisfy the regularity conditions under which such asymptotic contraction is guaranteed[13](https://arxiv.org/html/2608.20758#bib.bib50);[7](https://arxiv.org/html/2608.20758#bib.bib37), potentially leading to behaviors that diverge from classical expectations[36](https://arxiv.org/html/2608.20758#bib.bib36);[7](https://arxiv.org/html/2608.20758#bib.bib37)\(see Supplementary Note[2](https://arxiv.org/html/2608.20758#S2a)for more details\)\.

Consistent with this theoretical perspective, our analysis of the Bayesian output layer reveals a systematic deviation from classical expectations\. Contrary to the conventional expectation that larger datasets yield constricted posteriors, we observed that the uncertainty of the weight posterior—computed as the mean standard deviation across all hidden dimensions—increases with dataset size \(Fig\.[1](https://arxiv.org/html/2608.20758#S2.F1)d\)\. This posterior broadening is further visualized in Fig\.[1](https://arxiv.org/html/2608.20758#S2.F1)f, where the posterior uncertainty increases from 0\.05 \(1%\\%data, blue\) to 0\.11 \(20%\\%data, red\)\. One possible interpretation is that the learned representation and the variational output posterior co\-adapt during training\. As additional data improve the representation, predictive information can be redistributed toward directions that are less sensitive to uncertain output weights\. This interpretation motivates examining the coupled latent\-posterior geometry, although the present observation alone does not establish why the posterior scale broadens\. Although the bias posteriors fluctuate, they exhibit no consistent trend and show negligible influence on predictive uncertainty; consequently, our subsequent analysis primarily focuses on the weight posteriors \(see Supplementary Note[4](https://arxiv.org/html/2608.20758#S4a)\)\. This decoupling is also reflected in the training dynamics shown in Fig\.[1](https://arxiv.org/html/2608.20758#S2.F1)g\. With increasing training data, the Negative Log\-Likelihood \(NLL\) decreases, reflecting improved predictive performance and confidence\. At the same time, the Kullback\-Leibler \(KL\) divergence term decreases, indicating that the variational posterior moves closer to the𝒩⁡\(0,1\)\\mathcal\{N\}\(0,1\)prior in KL divergence\. Notably, the posterior means remain centered around zero across all dataset sizes \(Fig\.[1](https://arxiv.org/html/2608.20758#S2.F1)e\), indicating that the reduction in KL divergence is primarily driven by posterior variance expanding toward the unit variance imposed by the prior\.

Taken together, these findings indicate a decoupling between posterior uncertainty and predictive uncertainty: the model achieves lower predictive uncertainty even as the posterior becomes more diffuse\. This decoupling implies that the reduction in predictive uncertainty cannot be attributed to the constriction of the weight posterior, necessitating an alternative governing mechanism\. Importantly, this behavior is consistently observed across five additional molecular\-property datasets \(see Supplementary Fig\.[4](https://arxiv.org/html/2608.20758#S5.F4)\), demonstrating its robustness\.

### 2\.2Latent\-posterior alignment mechanism

To explain the observed decoupling between posterior behavior and predictive uncertainty, we examine the analytical form of the predictive variance\. The outputyyof the BGNN, derived from the Bayesian linear layer, is given by[6](https://arxiv.org/html/2608.20758#bib.bib3):

y=𝐳⊤​𝐰\+b,𝐰∼qϕ​\(𝐰\),b∼qϕ​\(𝐛\)\\displaystyle y=\\mathbf\{z\}^\{\\top\}\\mathbf\{w\}\+b,\\quad\\mathbf\{w\}\\sim q\_\{\\phi\}\(\\mathbf\{w\}\),\\quad b\\sim q\_\{\\phi\}\(\\mathbf\{b\}\)\(2\)
where thedd\-dimensional latent vector𝐳∈ℝd\\mathbf\{z\}\\in\\mathbb\{R\}^\{d\}, computed by the GIN layers, serves as the input to the Bayesian layer, and the weight𝐰\\mathbf\{w\}and biasbbare sampled from their respective posterior distributions,qϕ​\(𝐰\)q\_\{\\phi\}\(\\mathbf\{w\}\)andqϕ​\(𝐛\)q\_\{\\phi\}\(\\mathbf\{b\}\)\. By assuming a mean\-field approximation[6](https://arxiv.org/html/2608.20758#bib.bib3);[5](https://arxiv.org/html/2608.20758#bib.bib16), where the weights are independent across dimensions, and noting the negligible contribution of the bias \(Supplementary Note[4](https://arxiv.org/html/2608.20758#S4a)\), the predictive variance can be approximated as:

Var\[y\|𝐳,𝒟\]\\displaystyle\\mathrm\{Var\}\\left\[y\|\\mathbf\{z\},\\mathcal\{D\}\\right\]≈𝐳⊤​Cov​\[𝐰\|𝒟\]​𝐳\\displaystyle\\approx\\mathbf\{z\}^\{\\top\}\\mathrm\{Cov\}\\left\[\\mathbf\{w\}\|\\mathcal\{D\}\\right\]\\mathbf\{z\}\(3\)≈∑i=1dzi2​σw,i2\\displaystyle\\approx\\sum\_\{i=1\}^\{d\}z\_\{i\}^\{2\}\\sigma\_\{w,i\}^\{2\}\(4\)
whereσw,i\\sigma\_\{w,i\}anddddenote the standard deviation of theii\-th weight posterior and the number of latent dimensions, respectively\. This expression shows that predictive uncertainty depends on the interaction between the latent representation𝐳\\mathbf\{z\}and the posterior varianceσw2\\sigma\_\{w\}^\{2\}\. In particular, uncertainty can remain low even when the overall posterior variance is large, provided that𝐳\\mathbf\{z\}is concentrated along low\-variance directions\. Although Eq\.[4](https://arxiv.org/html/2608.20758#S2.E4)directly implies that predictive variance depends jointly on latent activations and posterior variance, the emergence of systematic latent reorganization during training is not obvious\. Based on this observation, we proposeLatent\-Posterior Alignment \(LPA\)as the primary mechanism that governs predictive uncertainty\. This means that the model reduces uncertainty not by uniformly decreasing posterior variances \(σw,i2\\sigma\_\{w,i\}^\{2\}\), but by strategically aligning the latent representation𝐳\\mathbf\{z\}with stable \(low\-variance\) posterior directions\. Concretely, the model learns to amplify components of𝐳\\mathbf\{z\}in dimensions with smallσw,i\\sigma\_\{w,i\}\(i\.e\., “reliable” dimensions\), while suppressing those associated with large variance \(i\.e\., “unstable” dimensions\)\.

To quantify this structural behavior, we define theLatent\-Posterior Alignment Score \(LPAS\)as:

Latent\-Posterior Alignment Score \(LPAS\)=1−∑i=1d\|zi\|​σi∑i=1d\|zi\|​∑i=1dσi\\displaystyle\\text\{Latent\-Posterior Alignment Score \(LPAS\)\}=1\-\\frac\{\\sum\_\{i=1\}^\{d\}\|z\_\{i\}\|\\sigma\_\{i\}\}\{\\sum\_\{i=1\}^\{d\}\|z\_\{i\}\|\\sum\_\{i=1\}^\{d\}\\sigma\_\{i\}\}\(5\)
A higher LPAS indicates strong alignment, where the latent vector selectively utilizes stable, low\-variance dimensions to reduce uncertainty\. Further insights into the relationship between the LPAS and uncertainty are provided in Supplementary Note[6](https://arxiv.org/html/2608.20758#S6)\. We use LPAS as the primary alignment score because higher values intuitively indicate stronger alignment\. For baseline\-normalized AGL comparisons presented later, relative improvement is reported as the reduction in its complement,1−LPAS1\-\\mathrm\{LPAS\}, which contains the same information in the opposite direction and avoids compressed percentage changes when LPAS is close to one\.

Empirical results strongly support this alignment mechanism\. Fig\.[2](https://arxiv.org/html/2608.20758#S2.F2)a\-c visualize the normalized absolute latent vectorsz~i\\tilde\{z\}\_\{i\}and posterior standard deviationsσ~w,i\\tilde\{\\sigma\}\_\{w,i\}\(see Methods\) acrossd=64d=64hidden dimensions for models trained with 1%, 4%, and 80% of the dataset, respectively\. Theσ~w,i\\tilde\{\\sigma\}\_\{w,i\}values are sorted in descending order for visual clarity\. As the amount of training data increases, the latent representations progressively shift toward dimensions with lower posterior variance, leading to stronger LPA\. This structural change of the latent representations directly impacts predictive uncertainty\. We analyze the contribution of each dimension to the total uncertaintyCiC\_\{i\}\(see Methods\), shown as gray bars in Fig\.[2](https://arxiv.org/html/2608.20758#S2.F2)a–c\. As alignment strengthens, the contribution becomes increasingly concentrated in low\-variance dimensions, reflecting a shift in how predictive uncertainty is distributed across latent directions\.

![Refer to caption](https://arxiv.org/html/2608.20758v1/img_1_2.png)

Figure 2:Latent\-Posterior Alignment \(LPA\) of BGNN\. \(a\-c\) Normalized latent vectors \(𝐳~\\tilde\{\\mathbf\{z\}\}, blue\) progressively align with low\-variance weight dimensions \(σ~w\\tilde\{\\sigma\}\_\{w\}, orange\) as data accumulates \(1%→\\to80%\)\. Grey bars indicate the relative uncertainty contribution\. \(d\) High\-uncertainty samples \(green\) exhibit amplified latent magnitudes compared to the full data distribution \(blue\) while maintaining the alignment\. \(e\) t\-SNE projection of latent vectors colored by predictive uncertainty\. \(f\) Monotonic increase in the Latent\-Posterior Alignment Score \(LPAS\), quantifying the alignment mechanism\. Shaded areas represent standard deviation acrossn=30n=30independent runs\.Importantly, this global alignment does not eliminate local variability\. As shown in Fig\.[2](https://arxiv.org/html/2608.20758#S2.F2)d, after training each model, we ranked all test samples within each independent run by posterior\-predictive standard deviation and defined the top 5% as the high\-uncertainty subset \(green\)\. Selected post hoc from the predictions of the same trained model, this subset follows the overall LPA trend of the full test set \(blue\) but exhibits larger latent\-vector magnitudes; no separate model was trained for this analysis\. A t\-SNE projection \(Fig\.[2](https://arxiv.org/html/2608.20758#S2.F2)e\) reveals that these high\-uncertainty samples lie near the boundaries of the data manifold \(red\), where data density is low and sparsely populated\. This suggests that while the LPA controls the overall uncertainty reduction, the magnitude of the latent vector is adjusted to capture instance\-level uncertainty\. This behavior is reflected in the alignment score\. As shown in Fig\.[2](https://arxiv.org/html/2608.20758#S2.F2)f, the LPAS increases with data density, showing stronger alignment that closely tracks the reduction in predictive uncertainty\. Taken together, these results demonstrate that LPA acts as the major mechanism for uncertainty regulation, rather than a secondary effect\. Additional experiments across five independent datasets further confirm the consistency of this behavior, showing similar trends in LPA and corresponding changes in predictive uncertainty \(see Supplementary Note[7](https://arxiv.org/html/2608.20758#S7)\)\.

### 2\.3Probing the causal evidence for latent–posterior alignment

Having established the association between LPA and predictive uncertainty, we next test whether LPA contributes causally to predictive uncertainty by deliberately perturbing the alignment structure\. To this end, we perform an interventional experiment that directly perturbs the latent structure\. Specifically, we introduce explicit regularization on the latent representation𝐳\\mathbf\{z\}, including L1, L2, and anti\-alignment penalties \(see Supplementary Note[8](https://arxiv.org/html/2608.20758#S8)for details\)\. These constraints are designed to disrupt the alignment between the latent representation and the low\-variance posteriors\. If LPA governs the uncertainty, then preventing such alignment should lead to an increase in predictive uncertainty\.

ℒL​1=−E​L​B​O\+λl​1​‖z‖1\\displaystyle\\mathcal\{L\}\_\{L1\}=\-ELBO\+\\lambda\_\{l1\}\\\|z\\\|\_\{1\}\(6\)
ℒL​2=−E​L​B​O\+λl​2​‖z‖2\\displaystyle\\mathcal\{L\}\_\{L2\}=\-ELBO\+\\lambda\_\{l2\}\\\|z\\\|\_\{2\}\(7\)
ℒa​n​t​i−a​l​i​g​n=−E​L​B​O−λa​n​t​i−a​l​i​g​n​∑i=1d\|zi\|​σi∑i=1d\|zi\|​∑i=1dσi\\displaystyle\\mathcal\{L\}\_\{anti\-align\}=\-ELBO\-\\lambda\_\{anti\-align\}\\frac\{\\sum\_\{i=1\}^\{d\}\|z\_\{i\}\|\\sigma\_\{i\}\}\{\\sum\_\{i=1\}^\{d\}\|z\_\{i\}\|\\sum\_\{i=1\}^\{d\}\\sigma\_\{i\}\}\(8\)
The results provide convergent interventional evidence that LPA contributes causally to predictive uncertainty \(Fig\.[3](https://arxiv.org/html/2608.20758#S2.F3)a–f\)\. Models trained with these latent constraints exhibit a consistent degradation in predictive accuracy \(Fig\.[3](https://arxiv.org/html/2608.20758#S2.F3)a\) and increase in predictive uncertainty \(Fig\.[3](https://arxiv.org/html/2608.20758#S2.F3)b\) compared to the unconstrained baseline, despite identical training conditions\. Correspondingly, the alignment score \(LPAS\) is significantly reduced \(Fig\.[3](https://arxiv.org/html/2608.20758#S2.F3)c\), and the latent representations no longer align with low\-variance posteriors \(Fig\.[3](https://arxiv.org/html/2608.20758#S2.F3)d\-f\)\. Further details of the analysis are provided in Supplementary Note[8](https://arxiv.org/html/2608.20758#S8)\.

These interventions consistently weaken LPA and increase predictive uncertainty, providing evidence that the alignment has a causal contribution beyond a simple correlation\. The anti\-alignment penalty provides the most direct perturbation of the proposed mechanism, whereas L1 and L2 regularization and the additional prior and KL\-weight experiments provide complementary evidence across distinct perturbations\. Their effects differ because these interventions also modify other properties of the learned system: L2 regularization induces relatively diffuse shrinkage across dimensions, whereas L1 regularization promotes sparsity and more strongly restricts the effective dimensional capacity of the representation\. Because latent norm, sparsity, representational capacity, and optimization dynamics are not held fixed, these experiments do not identify an isolated causal effect of LPA\. We therefore interpret them as convergent evidence for a causal contribution of LPA rather than definitive proof that LPA is the exclusive driver of predictive uncertainty\. Consistent results across additional benchmark datasets are provided in the Supplementary Note[9](https://arxiv.org/html/2608.20758#S9)\. Importantly, this behavior is robust to variations in posterior regularization and prior specification\. Even when these constraints are relaxed or tightened, the model continues to rely on the alignment structure to regulate predictive uncertainty \(see Supplementary Note[10](https://arxiv.org/html/2608.20758#S10)\)\.

![Refer to caption](https://arxiv.org/html/2608.20758v1/img_2_test_new_1.png)

Figure 3:Interventional evidence for a causal contribution of Latent–Posterior Alignment \(a\-f\)\. Structurally disrupting the LPA via L1, L2, and anti\-alignment regularization degrades accuracy \(a\) and amplifies predictive uncertainty \(b\), accompanied by the decrease in the alignment score \(LPAS\) \(c\) compared to baseline \(dashed line\)\. Visualizations of latent geometry \(d–f\) illustrate this structural breakdown: unlike the baseline, the regularized models fail to maintain LPA\. Shaded areas represent standard deviation acrossn=30n=30independent runs\.
### 2\.4Alignment\-guided learning for density\-aware uncertainty

Motivated by evidence that Latent–Posterior Alignment \(LPA\) contributes to predictive uncertainty, we introduce a training strategy that explicitly promotes this coupling and examine whether it improves the structural relationship between uncertainty and representation support\. Specifically, we introduce Alignment\-Guided Learning \(AGL\), which encourages LPA during training\. AGL incorporates the LPAS into the training objective as a regularization term:

ℒt​o​t​a​l=−ELBO−γ⋅LPAS\\displaystyle\\mathcal\{L\}\_\{total\}=\-\\text\{ELBO\}\-\\gamma\\cdot\\text\{LPAS\}\(9\)
By penalizing misalignment, AGL promotes a latent alignment that favors low\-variance posterior directions, thereby reducing predictive uncertainty\. The regularization weightγ\\gammais tuned to maintain predictive performance \(MAE\) comparable to the baseline model \(see Supplementary Table[5](https://arxiv.org/html/2608.20758#S8.T5)\)\.

To evaluate the effectiveness of AGL, we analyze the model across multiple metrics capturing predictive performance, uncertainty, and structural calibration\. In addition to LPAS, MAE, posterior standard deviation \(σw\\sigma\_\{w\}\), and output uncertainty \(σy\\sigma\_\{y\}\), we consider Expected Calibration Error \(ECE\)[37](https://arxiv.org/html/2608.20758#bib.bib8);[41](https://arxiv.org/html/2608.20758#bib.bib9)which evaluates consistency between model confidence and accuracy, and Density Uncertainty Criterion \(DUC\)[59](https://arxiv.org/html/2608.20758#bib.bib10), which measures how well uncertainty reflects the underlying data density\. ECE and DUC are essential for assessing uncertainty estimates, as they capture whether uncertainty is accurate, and is structurally consistent with the data distributions, respectively\. All values are normalized relative to the baseline \(set to 100\)\. For LPAS, the normalized value represents the proportional reduction in1−LPAS1\-\\mathrm\{LPAS\}\. We further compare AGL across data\-poor \(1–4%\) and data\-rich \(20–80%\) regimes to examine how alignment\-driven uncertainty behaves under varying data density\.

![Refer to caption](https://arxiv.org/html/2608.20758v1/img_3_agl_new.png)

Figure 4:Alignment\-Guided Learning \(AGL\) and uncertainty calibration analysis\. \(a\) Normalized performance metrics\. For LPAS, normalized values represent the proportional reduction in1−LPAS1\-\\mathrm\{LPAS\}relative to the baseline; values above 100 indicate stronger alignment\. AGL improves structural calibration \(DUC\) and increases the Latent\-Posterior Alignment Score \(LPAS\) across both regimes without compromising accuracy \(MAE\) compared to baseline \(dashed line\)\. \(b–e\) Latent alignment\. AGL enforces latent vectors \(𝐳~\\tilde\{\\mathbf\{z\}\}\) to shift toward low\-variance posterior dimensions\. This geometric reconfiguration is more pronounced in the data\-rich regime \(d, e\) compared to the data\-poor regime \(b, c\)\. \(f\-m\) Qualitative visualization of density awareness\. Each pair displays estimated density \(left, KDE\) and predictive uncertainty \(right,σy\\sigma\_\{y\}\) on the same t\-SNE coordinates\. In the data\-poor regime \(f–i\), the baseline shows no correlation between density \(f\) and uncertainty \(g\)\. AGL establishes this missing relationship, aligning low\-density regions \(h, blue in KDE\) with high uncertainty \(i, red inσy\\sigma\_\{y\}\)\. In the data\-rich regime \(j\-m\), AGL refines this structural correspondence, sharpening the boundary so that uncertainty \(m\) is precisely concentrated along the low\-density edges of the data manifold \(l\), compared to the baseline \(j, k\)\.As shown in Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)a, AGL consistently increases LPAS across both regimes, confirming that the alignment is effectively reinforced\. This effect is more pronounced in the data\-rich setting, where sufficient data allows a more distinct reorganization of the latent space \(Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)d,e\)\. In contrast, latent space reorganization under low data density settings is limited due to information constraints \(Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)b,c\)\.

Meanwhile, the predictive accuracy \(MAE\) remains unchanged, indicating that AGL preserves the model performance \(Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)a\)\. Importantly, while the posterior uncertainty \(σw\\sigma\_\{w\}\) barely changed, the predictive uncertainty \(σy\\sigma\_\{y\}\) is significantly reduced\. This decoupling further supports our central claim: LPA plays a key role in determining predictive uncertainty, rather than the magnitude of posterior variance\. By explicitly shaping the latent geometry, AGL reduces output uncertainty without requiring posterior contraction\.

Calibration analysis highlights a key advantage of AGL\. As shown in Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)a, the ECE remains largely unchanged, indicating that the alignment between prediction confidence and prediction error is preserved\. In contrast, the DUC is substantially improved for both data regimes\. This illustrates that AGL is capable of calibrating the structural relationship between uncertainty and data density, while preserving the prediction error\. Specifically, uncertainty is evaluated to be high in data scarce regions and low for regions where data is abundant, in contrast to conventional models\.

The paired t\-SNE maps[44](https://arxiv.org/html/2608.20758#bib.bib14)in Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)f\-m provide complementary qualitative evidence for the expected direction of this dependence\. The density \(estimated via Kernel Density Estimation \(KDE\); see Methods\) and uncertainty \(σy\\sigma\_\{y\}\) maps use the same two\-dimensional coordinates, allowing their spatial patterns to be compared directly within each model\. In the data\-poor regime, the baseline shows weak spatial correspondence \(Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)f, g\), whereas AGL more frequently places elevated uncertainty in regions with lower estimated density \(Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)h, i\)\. In the data\-rich regime, a partial correspondence is already visible in the baseline \(Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)j, k\) and becomes more pronounced under AGL \(Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)l, m\), particularly near the representation boundary\. Taken together, the DUC score and paired maps indicate that AGL strengthens density\-uncertainty dependence and redistributes uncertainty toward less\-supported regions of the learned representation\. Similar qualitative patterns are observed across the additional datasets \(Supplementary Note[11](https://arxiv.org/html/2608.20758#S11)\), illustrating that AGL strengthens LPA and generally increases the association between predictive uncertainty and low\-support regions of the learned representation\.

## 3Discussion

The central observation of this study is that predictive uncertainty can decrease even while the Bayesian output\-layer posterior becomes broader\. As shown in Fig\.[1](https://arxiv.org/html/2608.20758#S2.F1)c, d, and f, output uncertainty declines as the average posterior standard deviation of the output\-layer weights increases\. Within the mean\-field formulation of Eq\.[4](https://arxiv.org/html/2608.20758#S2.E4), latent representations can compensate for a broader posterior by shifting predictive information toward dimensions with smaller posterior variance\. The increasing LPAS in Fig\.[2](https://arxiv.org/html/2608.20758#S2.F2)f and the perturbation results in Fig\.[3](https://arxiv.org/html/2608.20758#S2.F3)therefore suggest that predictive uncertainty should be interpreted in terms of the jointly trained representation\-posterior system, rather than the posterior alone\. Because the latent regularizers also alter latent norm, sparsity, and representational capacity, these experiments support a functional contribution of LPA but do not isolate it as the sole causal determinant of predictive uncertainty\. Note that this apparent lack of posterior contraction should not be read as a failure of classical posterior\-contraction theory\. Classical results describe exact posterior concentration in a pre\-specified statistical model under regularity conditions, whereas the distribution examined here is a mean\-field variational posterior over the output layer coupled to a deterministic representation learned from the same data[69](https://arxiv.org/html/2608.20758#bib.bib38);[21](https://arxiv.org/html/2608.20758#bib.bib39);[36](https://arxiv.org/html/2608.20758#bib.bib36);[71](https://arxiv.org/html/2608.20758#bib.bib40);[13](https://arxiv.org/html/2608.20758#bib.bib50);[7](https://arxiv.org/html/2608.20758#bib.bib37);[47](https://arxiv.org/html/2608.20758#bib.bib51)\.

The results further show that AGL strengthens density awareness even though data density is not explicitly included in its objective\. The increased DUC scores indicate a stronger relationship between representation scarcity and predictive uncertainty, while the paired t\-SNE maps show that elevated uncertainty more frequently coincides with lower\-density regions after AGL \(Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)\)\. Together, these results support both the strengthened relationship and its expected low\-density/high\-uncertainty direction\. This behavior can be understood from the AGL objective: by promoting latent utilization of lower\-posterior\-scale directions through Eq\.[9](https://arxiv.org/html/2608.20758#S2.E9), AGL reduces uncertainty associated with unnecessary activation of higher\-posterior\-scale dimensions in well\-supported regions\. Samples in sparse or boundary regions, however, may retain larger latent magnitudes or more atypical combinations of latent components, as suggested by Fig\.[2](https://arxiv.org/html/2608.20758#S2.F2)d and e, and therefore remain relatively uncertain even after the global alignment is strengthened\. AGL thus does not simply lower uncertainty uniformly, but redistributes it according to support in the learned representation\.

Although not evaluated prospectively here, this density\-aware allocation of uncertainty may be useful in scientific workflows in which predictions must be accepted or deferred for additional validation\. MLIP applications provide a representative example, where users must often decide whether a prediction for a newly encountered atomic configuration is sufficiently reliable or whether an additional DFT calculation should be performed[70](https://arxiv.org/html/2608.20758#bib.bib75);[73](https://arxiv.org/html/2608.20758#bib.bib76);[38](https://arxiv.org/html/2608.20758#bib.bib77)\. ECE is valuable for evaluating calibration on a labeled validation distribution, but it cannot by itself determine whether an individual unlabeled configuration lies outside the empirical support of the model\. A density\-aware uncertainty model is more directly actionable in this setting because latent sparsity and predictive uncertainty can be evaluated before the DFT result is available\. In this sense, DUC can serve as a complementary criterion to ECE in MLIP applications\. Because AGL increases DUC while maintaining ECE, the resulting uncertainty becomes more informative about the density of the learned latent representation without compromising aggregate error calibration\. In practice, a high uncertainty estimate from an AGL\-trained model can be interpreted as a signal that the queried atomic configuration may be weakly supported by the training distribution in latent space, providing a rationale for requesting an additional DFT calculation\. The same property may also be useful in active learning and Bayesian optimization, where uncertainty is used to prioritize experiments or calculations in insufficiently explored regions\. This density\-aware interpretation should complement rather than replace error\-based calibration, because low latent density does not necessarily imply large error and high latent density does not exclude systematic bias[57](https://arxiv.org/html/2608.20758#bib.bib13);[65](https://arxiv.org/html/2608.20758#bib.bib78)\.

Several questions remain unresolved\. The present analysis is restricted to deterministic GNN feature extractors combined with mean\-field Bayesian output layers, and it is unknown whether the same alignment behavior will emerge under fully Bayesian architectures, correlated posterior approximations, dropout, Laplace methods, or ensembles\. In addition, the molecular\-property benchmarks used here evaluate uncertainty on fixed test distributions and therefore do not capture the scaffold shifts, prospective chemical\-space exploration, or the trajectory\-dependent distribution shifts that can arise in molecular simulations\. In this context, future studies should examine the training\-time evolution of latent directions and posterior scales, and evaluate whether LPA remains informative under alternative uncertainty frameworks\. A further important direction is to test whether AGL\-based density awareness can support prospective active\-learning decisions by prioritizing costly DFT calculations or experiments while reducing the risk of undetected high\-error configurations\. Within the model class examined here, these findings identify latent–posterior geometry as a central component of predictive uncertainty and provide a practical route for making model confidence more responsive to data support\.

## 4Methods

### 4\.1Bayesian GNN framework and training

#### 4\.1\.1Bayesian Graph Neural Network \(BGNN\)

We constructed a backbone GNN composed of five sequential Graph Isomorphism Networks \(GIN\)[75](https://arxiv.org/html/2608.20758#bib.bib21), followed by pooling and a fully connected Bayesian layer\. Each molecule is represented as a graph where nodes encode 36 atomic features \(e\.g\., element type, hybridization\) and edges represent covalent bonds\. A complete list of the detailed model architecture and input features are summarized in Supplementary Note[1\.1](https://arxiv.org/html/2608.20758#S1.SS1)\.

#### 4\.1\.2Model training

The model parameters were optimized by minimizing the negative Evidence Lower Bound \(ELBO\) using the AdamW optimizer\. To ensure statistical robustness, all experiments were repeated across 30 independent runs\. Specific training configurations, including learning rates, batch sizes, and stopping criteria, are detailed in Supplementary Note[1\.2](https://arxiv.org/html/2608.20758#S1.SS2)\. Additionally, parity plots assessing the predictive performance for each physicochemical property are presented in Supplementary Fig\.[2](https://arxiv.org/html/2608.20758#S1.F2)\.

### 4\.2Datasets

Training datasets were compiled from various publicly available sources, including OPERA[46](https://arxiv.org/html/2608.20758#bib.bib67)and QM9\*[66](https://arxiv.org/html/2608.20758#bib.bib23), and carefully curated to ensure internal consistency\. For molecules reported with minor discrepancies, values were averaged, while substantial inconsistencies were removed\. We allocated 80% of the data for training and 20% for testing\. Detailed descriptions of all datasets, including their sources, sizes, and preprocessing protocols, are provided in Supplementary Note[1\.3](https://arxiv.org/html/2608.20758#S1.SS3)\.

### 4\.3Quantification of uncertainty dynamics

To ensure the statistical robustness and reproducibility of our findings, all reported metrics represent the aggregate performance acrossR=30R=30independent training runs with distinct random seeds\.

#### 4\.3\.1Predictive metrics \(MAE,σy\\sigma\_\{y\}\)

Predictive accuracy and output uncertainty are computed by averaging over allNNsamples in the test dataset and across allRRindependent runs\. The Mean Absolute Error \(MAE\) and the average output uncertaintyσy\\sigma\_\{y\}are defined as:

MAE=1R​∑r=1R\(1N​∑i=1N\|yi\(r\)−y^i\(r\)\|\)\\displaystyle=\\frac\{1\}\{R\}\\sum\_\{r=1\}^\{R\}\\left\(\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\|y\_\{i\}^\{\(r\)\}\-\\hat\{y\}\_\{i\}^\{\(r\)\}\|\\right\)\(10\)σy\\displaystyle\\sigma\_\{y\}=1R​∑r=1R\(1N​∑i=1Nσy,i\(r\)\)\\displaystyle=\\frac\{1\}\{R\}\\sum\_\{r=1\}^\{R\}\\left\(\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\sigma\_\{y,i\}^\{\(r\)\}\\right\)\(11\)
whereyi\(r\)y\_\{i\}^\{\(r\)\},y^i\(r\)\\hat\{y\}\_\{i\}^\{\(r\)\}, andσy,i\(r\)\\sigma\_\{y,i\}^\{\(r\)\}denote the ground truth, predicted mean, and predicted standard deviation \(or output uncertainty\) for theii\-th sample in therr\-th run, respectively\.

#### 4\.3\.2Posterior metrics \(μw\\mu\_\{w\},σw\\sigma\_\{w\}\)

To analyze the global behavior of the posteriors, statistics of the posteriors are aggregated across alld=64d=64hidden dimensions and theRRruns\. The global posterior meanμw\\mu\_\{w\}and posterior uncertaintyσw\\sigma\_\{w\}are calculated as:

μw\\displaystyle\\mu\_\{w\}=1R⋅d​∑r=1R∑j=1dμwj\(r\)\\displaystyle=\\frac\{1\}\{R\\cdot d\}\\sum\_\{r=1\}^\{R\}\\sum\_\{j=1\}^\{d\}\\mu\_\{w\_\{j\}\}^\{\(r\)\}\(12\)σw\\displaystyle\\sigma\_\{w\}=1R⋅d​∑r=1R∑j=1dσwj\(r\)\\displaystyle=\\frac\{1\}\{R\\cdot d\}\\sum\_\{r=1\}^\{R\}\\sum\_\{j=1\}^\{d\}\\sigma\_\{w\_\{j\}\}^\{\(r\)\}\(13\)
whereμwj\(r\)\\mu\_\{w\_\{j\}\}^\{\(r\)\}andσwj\(r\)\\sigma\_\{w\_\{j\}\}^\{\(r\)\}represent the mean and standard deviation of the posterior distribution for thejj\-th weight parameter in therr\-th run\.

#### 4\.3\.3Latent alignment metrics \(z~i\\tilde\{z\}\_\{i\},σ~w,i\\tilde\{\\sigma\}\_\{w,i\}\)

To directly examine the alignment of latent vectors, we quantify the relative magnitude of each latent component and each posterior scale within a normalized space\. For each data samplenn, the absolute latent activation\|zi\(n\)\|\|z\_\{i\}^\{\(n\)\}\|is normalized across allddhidden dimensions, while the posterior standard deviationσw,i\\sigma\_\{w,i\}is normalized once per model:

z~i\\displaystyle\\tilde\{z\}\_\{i\}=1N​∑n=1N\(\|zi\(n\)\|∑j=1d\|zj\(n\)\|\)\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}\\left\(\\frac\{\|z\_\{i\}^\{\(n\)\}\|\}\{\\sum\_\{j=1\}^\{d\}\|z\_\{j\}^\{\(n\)\}\|\}\\right\)\(14\)σ~w,i\\displaystyle\\tilde\{\\sigma\}\_\{w,i\}=σw,i∑j=1dσw,j\\displaystyle=\\frac\{\\sigma\_\{w,i\}\}\{\\sum\_\{j=1\}^\{d\}\\sigma\_\{w,j\}\}\(15\)
These normalized quantities capture the relative contribution of each latent direction and its associated posterior uncertainty\. To visualize the geometric relationship between latent representations and posterior scales,z~i\\tilde\{z\}\_\{i\}values are plotted after sorting the correspondingσ~w,i\\tilde\{\\sigma\}\_\{w,i\}in descending order\. This enables direct inspection of whether latent directions align with high\-variance axes, low\-variance axes, or exhibit no systematic structure\.

#### 4\.3\.4Relative uncertainty contribution \(CiC\_\{i\}\)

To dissect the sources of predictive uncertainty \(Eq\.[4](https://arxiv.org/html/2608.20758#S2.E4)\), we evaluate the proportion of total variance attributable to theii\-th latent dimension\. This is computed for each dimension and averaged across allNNsamples:

Ci=1R∑r=1R1N∑l=1N\(\(zl,i\(r\)​σwi\(r\)\)2∑k=1d\(zl,k\(r\)​σwk\(r\)\)2\)×100\(%\)\\displaystyle C\_\{i\}=\\frac\{1\}\{R\}\\sum\_\{r=1\}^\{R\}\\frac\{1\}\{N\}\\sum\_\{l=1\}^\{N\}\\left\(\\frac\{\(z\_\{l,i\}^\{\(r\)\}\\sigma\_\{w\_\{i\}\}^\{\(r\)\}\)^\{2\}\}\{\\sum\_\{k=1\}^\{d\}\(z\_\{l,k\}^\{\(r\)\}\\sigma\_\{w\_\{k\}\}^\{\(r\)\}\)^\{2\}\}\\right\)\\times 100\\\>\(\\%\)\(16\)
This metric quantifies how much eachii\-th latent dimension and its corresponding weight posterior variance contribute to the final output uncertainty\. We utilized this to visualize the shift in uncertainty allocation under different data regimes \(Fig\.[2](https://arxiv.org/html/2608.20758#S2.F2)a–c\)\.

### 4\.4Calibration metrics

#### 4\.4\.1Expected Calibration Error \(ECE\)

To assess the reliability of predictive uncertainty, we utilized ECE for regression, which measures the discrepancy between predicted confidence intervals and empirical coverage probabilities\. Unlike ECE for classification[54](https://arxiv.org/html/2608.20758#bib.bib19);[26](https://arxiv.org/html/2608.20758#bib.bib18);[56](https://arxiv.org/html/2608.20758#bib.bib17), which bins predictions by confidence score, regression ECE evaluates whether thepp\-confidence interval centered at the predicted meanμ⁡\(x\)\\mu\(x\)contains the true valueyywith probabilitypp[37](https://arxiv.org/html/2608.20758#bib.bib8);[41](https://arxiv.org/html/2608.20758#bib.bib9);[27](https://arxiv.org/html/2608.20758#bib.bib20)\. Assuming a Gaussian predictive distribution𝒩⁡\(μ⁡\(x\),σ2​\(x\)\)\\mathcal\{N\}\(\\mu\(x\),\\sigma^\{2\}\(x\)\), the expectedpp\-confidence interval for a target confidence levelp∈\(0,1\)p\\in\(0,1\)is defined as:

Ip\(x\)=\[μ\(x\)−Φ−1\(1\+p2\)σ\(x\),μ\(x\)\+Φ−1\(1\+p2\)σ\(x\)\]\\displaystyle I\_\{p\}\(x\)=\\left\[\\mu\(x\)\-\\Phi^\{\-1\}\\left\(\\frac\{1\+p\}\{2\}\\right\)\\sigma\(x\),\\quad\\mu\(x\)\+\\Phi^\{\-1\}\\left\(\\frac\{1\+p\}\{2\}\\right\)\\sigma\(x\)\\right\]\(17\)
whereΦ−1\\Phi^\{\-1\}denotes the probit function \(inverse cumulative distribution function of the standard normal distribution\)\. Using this confidence interval, an empirical coveragep^\\hat\{p\}can be calculated as the proportion of data samples for which the true value falls within the interval:

p^=1N​∑i=1N𝕀⁡\(yi∈Ip​\(xi\)\)\\displaystyle\\hat\{p\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbb\{I\}\\left\(y\_\{i\}\\in I\_\{p\}\(x\_\{i\}\)\\right\)\(18\)
whereNNis the total number of samples and𝕀⁡\(⋅\)\\mathbb\{I\}\(\\cdot\)is the indicator function\. We considered a range of target confidence levels,p∈\{0\.1,0\.2,…,0\.9,0\.99\}p\\in\\\{0\.1,0\.2,\\dots,0\.9,0\.99\\\}, and calculated the calibration gap between the target confidencepjp\_\{j\}and the empirical coveragep^j\\hat\{p\}\_\{j\}for eachjj\-th confidence interval\. The final ECE is computed by approximating the integral of the calibration gapej=\|pj−p^j\|e\_\{j\}=\|p\_\{j\}\-\\hat\{p\}\_\{j\}\|over the domain ofppusing the trapezoidal rule:

ECE≈∑j\|ej\|\+\|ej\+1\|2⋅\(pj\+1−pj\),where​ej=p^j−pj\\displaystyle\\text\{ECE\}\\approx\\sum\_\{j\}\\frac\{\|e\_\{j\}\|\+\|e\_\{j\+1\}\|\}\{2\}\\cdot\(p\_\{j\+1\}\-p\_\{j\}\),\\\>\\\>\\text\{where \}e\_\{j\}=\\hat\{p\}\_\{j\}\-p\_\{j\}\(19\)This metric quantifies the total misalignment area between theoretical confidence and observed coverage, with lower values indicating better calibration\.

#### 4\.4\.2Density Uncertainty Criterion \(DUC\)

To evaluate whether the model’s predictive uncertainty faithfully reflects the underlying data distribution—specifically, assigning high uncertainty to regions of low data support—we utilized the Density Uncertainty Criterion \(DUC\)[43](https://arxiv.org/html/2608.20758#bib.bib11);[68](https://arxiv.org/html/2608.20758#bib.bib12);[57](https://arxiv.org/html/2608.20758#bib.bib13)\. This metric quantifies the statistical dependence between data scarcity and predictive uncertainty\. First, we estimated the probability density functionp⁡\(𝐳\)p\(\\mathbf\{z\}\)of the latent representations using Kernel Density Estimation \(KDE\) with a Gaussian kernel and a bandwidth of 0\.2\. To measure the scarcity of a data point, we then computed the negative log\-likelihood \(also known as surprisal\):

s⁡\(𝐳\)=−log⁡p⁡\(𝐳\)\\displaystyle s\(\\mathbf\{z\}\)=\-\\log p\(\\mathbf\{z\}\)\(20\)
A higher DUC indicates stronger statistical dependence between data scarcity and predictive uncertainty\. Because MI is unsigned, DUC alone does not determine the direction of this dependence; the expected low\-density/high\-uncertainty direction is assessed separately using the paired KDE–uncertainty maps in Fig\.[4](https://arxiv.org/html/2608.20758#S2.F4)and Supplementary Figs\.[13](https://arxiv.org/html/2608.20758#S11.F13)and[14](https://arxiv.org/html/2608.20758#S11.F14)\.

DUC=M​I​\(s⁡\(𝐳\),σ⁡\(𝐲\|𝐳\)\)\\displaystyle\\text\{DUC\}=MI\(s\(\\mathbf\{z\}\);\\sigma\(\\mathbf\{y\}\|\\mathbf\{z\}\)\)\(21\)
whereM​I​\(X,Y\)MI\(X;Y\)denotes the mutual information between random variablesXXandYY\. We estimated this value with a continuous MI estimator based on nearest neighbor distances using scikit\-learn v1\.6\.1[61](https://arxiv.org/html/2608.20758#bib.bib24)\. A higher DUC indicates a stronger alignment between the model’s uncertainty estimates and the geometry of the data manifold, confirming that the model effectively assigns higher uncertainty to regions where data is scarce\.

## Acknowledgements

This work was supported by the Nano\-Material Technology Development Program \(RS\-2026\-25534767\), STEAM Project \(2022M3C1A3092056\), Outstanding Junior Researcher Program \(RS\-2024\-00348230\), and the institutional research program of Korea Institute of Science and Technology \(26E0323\) of the National Research Foundation \(NRF\) grant funded by the Korea Ministry of Science and ICT\. This work was also supported by the National Supercomputing Center with supercomputing resources including technical support \(KSC\-2023\-CRE\-0559\)\.

## References

- L\. AitchisonA statistical theory of cold posteriors in deep neural networks\.arXiv preprint arXiv:2008\.05912\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p3.1)\.
- Antoránet al\.\(2020\)J\. Antorán, J\. Allingham, and J\. M\. Hernández\-LobatoDepth uncertainty in neural networks\.Advances in neural information processing systems33,pp\. 10620–10634\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p3.1)\.
- Bengioet al\.\(2013\)Y\. Bengio, A\. Courville, and P\. VincentRepresentation learning: a review and new perspectives\.IEEE transactions on pattern analysis and machine intelligence35\(8\),pp\. 1798–1828\.Cited by:[§2\.2](https://arxiv.org/html/2608.20758#S2.SS2a.p2.1)\.
- Bishop and Nasrabadi \(2006\)C\. M\. Bishop and N\. M\. NasrabadiPattern recognition and machine learning\.Vol\.4,Springer,New York\.Cited by:[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p1.1)\.
- Bleiet al\.\(2017\)D\. M\. Blei, A\. Kucukelbir, and J\. D\. McAuliffeVariational inference: a review for statisticians\.Journal of the American statistical Association112\(518\),pp\. 859–877\.Cited by:[§2\.2](https://arxiv.org/html/2608.20758#S2.SS2.p3.1),[§3](https://arxiv.org/html/2608.20758#S3a.p3.1)\.
- Blundellet al\.\(2015\)C\. Blundell, J\. Cornebise, K\. Kavukcuoglu, and D\. WierstraWeight uncertainty in neural network\.InInternational conference on machine learning,pp\. 1613–1622\.Cited by:[§1\.1](https://arxiv.org/html/2608.20758#S1.SS1.p4.1),[§1\.2](https://arxiv.org/html/2608.20758#S1.SS2.p1.1),[§1](https://arxiv.org/html/2608.20758#S1.p1.1),[§1](https://arxiv.org/html/2608.20758#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p3.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1a.p5.1),[§2\.2](https://arxiv.org/html/2608.20758#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2608.20758#S2.SS2.p3.1),[§2](https://arxiv.org/html/2608.20758#S2.p1.1)\.
- Bochkina \(2019\)N\. BochkinaBernstein–von mises theorem and misspecified models: a review\.Foundations of modern statistics,pp\. 355–380\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p3.1),[§3](https://arxiv.org/html/2608.20758#S3.p1.1)\.
- Bojarskiet al\.\(2016\)M\. Bojarski, D\. Del Testa, D\. Dworakowski, B\. Firner, B\. Flepp, P\. Goyal, L\. D\. Jackel, M\. Monfort, U\. Muller, J\. Zhang,et al\.End to end learning for self\-driving cars\.arXiv preprint arXiv:1604\.07316\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Bronsteinet al\.\(2021\)M\. M\. Bronstein, J\. Bruna, T\. Cohen, and P\. VeličkovićGeometric deep learning: grids, groups, graphs, geodesics, and gauges\.arXiv preprint arXiv:2104\.13478\.Cited by:[§2\.2](https://arxiv.org/html/2608.20758#S2.SS2a.p2.1)\.
- Chakrabortiet al\.\(2025\)T\. Chakraborti, C\. R\. Banerji, A\. Marandon, V\. Hellon, R\. Mitra, B\. Lehmann, L\. Bräuninger, S\. McGough, C\. Turkay, A\. F\. Frangi,et al\.Personalized uncertainty quantification in artificial intelligence\.Nature Machine Intelligence7\(4\),pp\. 522–530\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Chen and Li \(2025\)L\. Chen and Y\. LiUncertainty quantification with graph neural networks for efficient molecular design\.Nature Communications16\(1\),pp\. 3262\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Codevillaet al\.\(2018\)F\. Codevilla, M\. Müller, A\. López, V\. Koltun, and A\. DosovitskiyEnd\-to\-end driving via conditional imitation learning\.In2018 IEEE international conference on robotics and automation \(ICRA\),pp\. 4693–4700\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Daret al\.\(2021\)Y\. Dar, V\. Muthukumar, and R\. G\. BaraniukA farewell to the bias\-variance tradeoff? an overview of the theory of overparameterized machine learning\.arXiv preprint arXiv:2109\.02355\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p3.1),[§3](https://arxiv.org/html/2608.20758#S3.p1.1)\.
- de Mathelinet al\.\(2025\)A\. de Mathelin, F\. Deheeger, M\. Mougeot, and N\. VayatisDeep out\-of\-distribution uncertainty quantification via weight entropy maximization\.Journal of Machine Learning Research26\(4\),pp\. 1–68\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p3.1)\.
- Depeweget al\.\(2018\)S\. Depeweg, J\. Hernandez\-Lobato, F\. Doshi\-Velez, and S\. UdluftDecomposition of uncertainty in bayesian deep learning for efficient and risk\-sensitive learning\.InInternational conference on machine learning,pp\. 1184–1193\.Cited by:[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p1.1)\.
- Dusenberryet al\.\(2020\)M\. Dusenberry, G\. Jerfel, Y\. Wen, Y\. Ma, J\. Snoek, K\. Heller, B\. Lakshminarayanan, and D\. TranEfficient and scalable bayesian neural nets with rank\-1 factors\.InInternational conference on machine learning,pp\. 2782–2792\.Cited by:[§1\.2](https://arxiv.org/html/2608.20758#S1.SS2.p1.1),[§2](https://arxiv.org/html/2608.20758#S2.p1.1)\.
- Estevaet al\.\(2017\)A\. Esteva, B\. Kuprel, R\. A\. Novoa, J\. Ko, S\. M\. Swetter, H\. M\. Blau, and S\. ThrunDermatologist\-level classification of skin cancer with deep neural networks\.nature542\(7639\),pp\. 115–118\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Fanet al\.\(2024\)Z\. Fan, J\. Yu, X\. Zhang, Y\. Chen, S\. Sun, Y\. Zhang, M\. Chen, F\. Xiao, W\. Wu, X\. Li,et al\.Reducing overconfident errors in molecular property classification using posterior network\.Patterns5\(6\)\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p3.1)\.
- Foonget al\.\(2020\)A\. Foong, D\. Burt, Y\. Li, and R\. TurnerOn the expressiveness of approximate inference in bayesian neural networks\.Advances in Neural Information Processing Systems33,pp\. 15897–15908\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p3.1)\.
- Foonget al\.\(2019\)A\. Y\. Foong, Y\. Li, J\. M\. Hernández\-Lobato, and R\. E\. Turner’In\-between’uncertainty in bayesian neural networks\.arXiv preprint arXiv:1906\.11537\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p3.1)\.
- Freedman \(1999\)D\. FreedmanWald lecture: on the bernstein\-von mises theorem with infinite\-dimensional parameters\.The Annals of Statistics27\(4\),pp\. 1119–1141\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1a.p1.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1a.p5.1),[§2\.2](https://arxiv.org/html/2608.20758#S2.SS2a.p1.1),[§3](https://arxiv.org/html/2608.20758#S3.p1.1)\.
- Gal and Ghahramani \(2016\)Y\. Gal and Z\. GhahramaniDropout as a bayesian approximation: representing model uncertainty in deep learning\.Ininternational conference on machine learning,pp\. 1050–1059\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Gawlikowskiet al\.\(2023\)J\. Gawlikowski, C\. R\. N\. Tassi, M\. Ali, J\. Lee, M\. Humt, J\. Feng, A\. Kruspe, R\. Triebel, P\. Jung, R\. Roscher,et al\.A survey of uncertainty in deep neural networks\.Artificial Intelligence Review56\(Suppl 1\),pp\. 1513–1589\.Cited by:[§2](https://arxiv.org/html/2608.20758#S2.p1.1)\.
- Ghosalet al\.\(2000\)S\. Ghosal, J\. K\. Ghosh, and A\. W\. Van Der VaartConvergence rates of posterior distributions\.Annals of Statistics,pp\. 500–531\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p3.1)\.
- Graves \(2011\)A\. GravesPractical variational inference for neural networks\.Advances in neural information processing systems24\.Cited by:[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p3.1)\.
- Guoet al\.\(2017\)C\. Guo, G\. Pleiss, Y\. Sun, and K\. Q\. WeinbergerOn calibration of modern neural networks\.InInternational conference on machine learning,pp\. 1321–1330\.Cited by:[§4\.4\.1](https://arxiv.org/html/2608.20758#S4.SS4.SSS1.p1.1)\.
- Gustafssonet al\.\(2020\)F\. K\. Gustafsson, M\. Danelljan, and T\. B\. SchonEvaluating scalable bayesian deep learning methods for robust computer vision\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops,pp\. 318–319\.Cited by:[§4\.4\.1](https://arxiv.org/html/2608.20758#S4.SS4.SSS1.p1.1)\.
- Harrisonet al\.\(2024\)J\. Harrison, J\. Willes, and J\. SnoekVariational bayesian last layers\.arXiv preprint arXiv:2404\.11599\.Cited by:[§2](https://arxiv.org/html/2608.20758#S2.p1.1)\.
- Hüllermeier and Waegeman \(2021\)E\. Hüllermeier and W\. WaegemanAleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods\.Machine learning110\(3\),pp\. 457–506\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p3.1)\.
- Izmailovet al\.\(2021\)P\. Izmailov, S\. Vikram, M\. D\. Hoffman, and A\. G\. WilsonWhat are bayesian neural network posteriors really like?\.InInternational conference on machine learning,pp\. 4629–4640\.Cited by:[§1\.1](https://arxiv.org/html/2608.20758#S1.SS1.p4.1),[§1](https://arxiv.org/html/2608.20758#S1.p3.1)\.
- Jordanet al\.\(1999\)M\. I\. Jordan, Z\. Ghahramani, T\. S\. Jaakkola, and L\. K\. SaulAn introduction to variational methods for graphical models\.Machine learning37\(2\),pp\. 183–233\.Cited by:[§3](https://arxiv.org/html/2608.20758#S3a.p3.1)\.
- Jospinet al\.\(2022\)L\. V\. Jospin, H\. Laga, F\. Boussaid, W\. Buntine, and M\. BennamounHands\-on bayesian neural networks—a tutorial for deep learning users\.IEEE Computational Intelligence Magazine17\(2\),pp\. 29–48\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Jumperet al\.\(2021\)J\. Jumper, R\. Evans, A\. Pritzel, T\. Green, M\. Figurnov, O\. Ronneberger, K\. Tunyasuvunakool, R\. Bates, A\. Žídek, A\. Potapenko,et al\.Highly accurate protein structure prediction with alphafold\.nature596\(7873\),pp\. 583–589\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Kendall and Gal \(2017\)A\. Kendall and Y\. GalWhat uncertainties do we need in bayesian deep learning for computer vision?\.Advances in neural information processing systems30\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p1.1)\.
- Kingma and Welling \(2013\)D\. P\. Kingma and M\. WellingAuto\-encoding variational bayes\.arXiv preprint arXiv:1312\.6114\.Cited by:[§1\.2](https://arxiv.org/html/2608.20758#S1.SS2.p1.1)\.
- Kleijn and Van der Vaart \(2012\)B\. J\. Kleijn and A\. W\. Van der VaartThe bernstein\-von\-mises theorem under misspecification\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p3.1),[§2\.2](https://arxiv.org/html/2608.20758#S2.SS2a.p2.1),[§3](https://arxiv.org/html/2608.20758#S3.p1.1)\.
- Kuleshovet al\.\(2018\)V\. Kuleshov, N\. Fenner, and S\. ErmonAccurate uncertainties for deep learning using calibrated regression\.InInternational conference on machine learning,pp\. 2796–2804\.Cited by:[§2\.4](https://arxiv.org/html/2608.20758#S2.SS4.p4.1),[§4\.4\.1](https://arxiv.org/html/2608.20758#S4.SS4.SSS1.p1.1)\.
- Kulichenkoet al\.\(2023\)M\. Kulichenko, K\. Barros, N\. Lubbers, Y\. W\. Li, R\. Messerly, S\. Tretiak, J\. S\. Smith, and B\. NebgenUncertainty\-driven dynamics for active learning of interatomic potentials\.Nature computational science3\(3\),pp\. 230–239\.Cited by:[§3](https://arxiv.org/html/2608.20758#S3.p3.1)\.
- Lakshminarayananet al\.\(2017\)B\. Lakshminarayanan, A\. Pritzel, and C\. BlundellSimple and scalable predictive uncertainty estimation using deep ensembles\.Advances in neural information processing systems30\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Landrum \(2006\)G\. LandrumRDKit: open\-source cheminformatics\.http://www\.rdkit\.org\.Cited by:[§1\.1](https://arxiv.org/html/2608.20758#S1.SS1.p4.1)\.
- Laveset al\.\(2020\)M\. Laves, S\. Ihler, J\. F\. Fast, L\. A\. Kahrs, and T\. OrtmaierWell\-calibrated regression uncertainty in medical imaging with deep learning\.InMedical imaging with deep learning,pp\. 393–412\.Cited by:[§2\.4](https://arxiv.org/html/2608.20758#S2.SS4.p4.1),[§4\.4\.1](https://arxiv.org/html/2608.20758#S4.SS4.SSS1.p1.1)\.
- Linet al\.\(2023\)Y\. Lin, Q\. Zhang, B\. Gao, J\. Tang, P\. Yao, C\. Li, S\. Huang, Z\. Liu, Y\. Zhou, Y\. Liu,et al\.Uncertainty quantification via a memristor bayesian deep neural network for risk\-sensitive reinforcement learning\.Nature Machine Intelligence5\(7\),pp\. 714–723\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Liuet al\.\(2020\)J\. Liu, Z\. Lin, S\. Padhy, D\. Tran, T\. Bedrax Weiss, and B\. LakshminarayananSimple and principled uncertainty estimation with deterministic deep learning via distance awareness\.Advances in neural information processing systems33,pp\. 7498–7512\.Cited by:[§4\.4\.2](https://arxiv.org/html/2608.20758#S4.SS4.SSS2.p1.1)\.
- Maaten and Hinton \(2008\)L\. v\. d\. Maaten and G\. HintonVisualizing data using t\-sne\.Journal of machine learning research9\(Nov\),pp\. 2579–2605\.Cited by:[§2\.4](https://arxiv.org/html/2608.20758#S2.SS4.p8.1)\.
- MacKay \(1992\)D\. J\. MacKayA practical bayesian framework for backpropagation networks\.Neural computation4\(3\),pp\. 448–472\.Cited by:[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p1.1)\.
- Mansouriet al\.\(2018\)K\. Mansouri, C\. M\. Grulke, R\. S\. Judson, and A\. J\. WilliamsOPERA models for predicting physicochemical properties and environmental fate endpoints\.Journal of cheminformatics10\(1\),pp\. 10\.Cited by:[Table 4](https://arxiv.org/html/2608.20758#S1.T4.3.2.2.1.1),[Table 4](https://arxiv.org/html/2608.20758#S1.T4.3.3.2.1.1),[Table 4](https://arxiv.org/html/2608.20758#S1.T4.3.4.2.1.1),[Table 4](https://arxiv.org/html/2608.20758#S1.T4.3.5.2.1.1),[§2](https://arxiv.org/html/2608.20758#S2.p1.1),[§4\.2](https://arxiv.org/html/2608.20758#S4.SS2.p1.1)\.
- Masegosa \(2020\)A\. MasegosaLearning under model misspecification: applications to variational and ensemble methods\.Advances in Neural Information Processing Systems33,pp\. 5479–5491\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.20758#S2.SS2a.p2.1),[§3](https://arxiv.org/html/2608.20758#S3.p1.1)\.
- McKinneyet al\.\(2020\)S\. M\. McKinney, M\. Sieniek, V\. Godbole, J\. Godwin, N\. Antropova, H\. Ashrafian, T\. Back, M\. Chesus, G\. S\. Corrado, A\. Darzi,et al\.International evaluation of an ai system for breast cancer screening\.Nature577\(7788\),pp\. 89–94\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Merchantet al\.\(2023\)A\. Merchant, S\. Batzner, S\. S\. Schoenholz, M\. Aykol, G\. Cheon, and E\. D\. CubukScaling deep learning for materials discovery\.Nature624\(7990\),pp\. 80–85\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Mirsky \(1975\)L\. MirskyA trace inequality of john von neumann\.Monatshefte für mathematik79\(4\),pp\. 303–306\.Cited by:[§6](https://arxiv.org/html/2608.20758#S6.p7.1)\.
- Morriset al\.\(2019\)C\. Morris, M\. Ritzert, M\. Fey, W\. L\. Hamilton, J\. E\. Lenssen, G\. Rattan, and M\. GroheWeisfeiler and leman go neural: higher\-order graph neural networks\.InProceedings of the AAAI conference on artificial intelligence,Vol\.33,pp\. 4602–4609\.Cited by:[§1\.1](https://arxiv.org/html/2608.20758#S1.SS1.p1.1),[§2](https://arxiv.org/html/2608.20758#S2.p1.1)\.
- Mukhotiet al\.\(2023\)J\. Mukhoti, A\. Kirsch, J\. Van Amersfoort, P\. H\. Torr, and Y\. GalDeep deterministic uncertainty: a new simple baseline\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 24384–24394\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p3.1)\.
- Nabarroet al\.\(2022\)S\. Nabarro, S\. Ganev, A\. Garriga\-Alonso, V\. Fortuin, M\. van der Wilk, and L\. AitchisonData augmentation in bayesian neural networks and the cold posterior effect\.InUncertainty in Artificial Intelligence,pp\. 1434–1444\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p3.1)\.
- Naeiniet al\.\(2015\)M\. P\. Naeini, G\. Cooper, and M\. HauskrechtObtaining well calibrated probabilities using bayesian binning\.InProceedings of the AAAI conference on artificial intelligence,Vol\.29\.Cited by:[§4\.4\.1](https://arxiv.org/html/2608.20758#S4.SS4.SSS1.p1.1)\.
- Neal \(2012\)R\. M\. NealBayesian learning for neural networks\.Vol\.118,Springer Science & Business Media,New York\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Nixonet al\.\(2019\)J\. Nixon, M\. W\. Dusenberry, L\. Zhang, G\. Jerfel, and D\. TranMeasuring calibration in deep learning\.\.InCVPR workshops,Vol\.2\.Cited by:[§4\.4\.1](https://arxiv.org/html/2608.20758#S4.SS4.SSS1.p1.1)\.
- Ovadiaet al\.\(2019\)Y\. Ovadia, E\. Fertig, J\. Ren, Z\. Nado, D\. Sculley, S\. Nowozin, J\. Dillon, B\. Lakshminarayanan, and J\. SnoekCan you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift\.Advances in neural information processing systems32\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p3.1),[§3](https://arxiv.org/html/2608.20758#S3.p3.1),[§4\.4\.2](https://arxiv.org/html/2608.20758#S4.SS4.SSS2.p1.1)\.
- Panget al\.\(2021\)Y\. Pang, S\. Cheng, J\. Hu, and Y\. LiuEvaluating the robustness of bayesian neural networks against different types of attacks\.arXiv preprint arXiv:2106\.09223\.Cited by:[§1\.1](https://arxiv.org/html/2608.20758#S1.SS1.p4.1)\.
- Park and Blei \(2024\)Y\. Park and D\. BleiDensity uncertainty layers for reliable uncertainty estimation\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 163–171\.Cited by:[§2\.4](https://arxiv.org/html/2608.20758#S2.SS4.p4.1)\.
- Paszkeet al\.\(2019\)A\. Paszke, S\. Gross, F\. Massa, A\. Lerer, J\. Bradbury, G\. Chanan, T\. Killeen, Z\. Lin, N\. Gimelshein, L\. Antiga,et al\.Pytorch: an imperative style, high\-performance deep learning library\.Advances in neural information processing systems32\.Cited by:[§1\.1](https://arxiv.org/html/2608.20758#S1.SS1.p4.1)\.
- Pedregosaet al\.\(2011\)F\. Pedregosa, G\. Varoquaux, A\. Gramfort, V\. Michel, B\. Thirion, O\. Grisel, M\. Blondel, P\. Prettenhofer, R\. Weiss, V\. Dubourg,et al\.Scikit\-learn: machine learning in python\.the Journal of machine Learning research12,pp\. 2825–2830\.Cited by:[§4\.4\.2](https://arxiv.org/html/2608.20758#S4.SS4.SSS2.p4.1)\.
- Rakesh and Jain \(2021\)V\. Rakesh and S\. JainEfficacy of bayesian neural networks in active learning\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 2601–2609\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p3.1)\.
- Sanchez\-Lengeling and Aspuru\-Guzik \(2018\)B\. Sanchez\-Lengeling and A\. Aspuru\-GuzikInverse molecular design using machine learning: generative models for matter engineering\.Science361\(6400\),pp\. 360–365\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Sanchez\-Lengelinget al\.\(2017\)B\. Sanchez\-Lengeling, C\. Outeiral, G\. L\. Guimaraes, and A\. Aspuru\-GuzikOptimizing distributions over molecular space\. an objective\-reinforced generative adversarial network for inverse\-design chemistry \(organic\)\.Cited by:[Table 4](https://arxiv.org/html/2608.20758#S1.T4.3.2.2.1.1),[Table 4](https://arxiv.org/html/2608.20758#S1.T4.3.5.2.1.1)\.
- Tanet al\.\(2023\)A\. R\. Tan, S\. Urata, S\. Goldman, J\. C\. Dietschreit, and R\. Gómez\-BombarelliSingle\-model uncertainty quantification in neural network potentials does not consistently outperform model ensembles\.npj Computational Materials9\(1\),pp\. 225\.Cited by:[§3](https://arxiv.org/html/2608.20758#S3.p3.1)\.
- Tanget al\.\(2024\)M\. Tang, T\. Zhu, S\. Zhang, and X\. HongQM9star, two million dft\-computed equilibrium structures for ions and radicals with atomic information\.Scientific Data11\(1\),pp\. 1158\.Cited by:[Table 4](https://arxiv.org/html/2608.20758#S1.T4.3.6.2.1.1),[§4\.2](https://arxiv.org/html/2608.20758#S4.SS2.p1.1)\.
- Tranet al\.\(2019\)D\. Tran, M\. Dusenberry, M\. Van Der Wilk, and D\. HafnerBayesian layers: a module for neural network uncertainty\.Advances in neural information processing systems32\.Cited by:[§1\.2](https://arxiv.org/html/2608.20758#S1.SS2.p1.1),[§2](https://arxiv.org/html/2608.20758#S2.p1.1)\.
- Van Amersfoortet al\.\(2020\)J\. Van Amersfoort, L\. Smith, Y\. W\. Teh, and Y\. GalUncertainty estimation using a single deep deterministic neural network\.InInternational conference on machine learning,pp\. 9690–9700\.Cited by:[§4\.4\.2](https://arxiv.org/html/2608.20758#S4.SS4.SSS2.p1.1)\.
- Van der Vaart \(2000\)A\. W\. Van der VaartAsymptotic statistics\.Vol\.3,Cambridge University Press,Cambridge\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1a.p1.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1a.p5.1),[§2\.2](https://arxiv.org/html/2608.20758#S2.SS2a.p1.1),[§3](https://arxiv.org/html/2608.20758#S3.p1.1)\.
- Vandermauseet al\.\(2020\)J\. Vandermause, S\. B\. Torrisi, S\. Batzner, Y\. Xie, L\. Sun, A\. M\. Kolpak, and B\. KozinskyOn\-the\-fly active learning of interpretable bayesian force fields for atomistic rare events\.npj Computational Materials6\(1\),pp\. 20\.Cited by:[§3](https://arxiv.org/html/2608.20758#S3.p3.1)\.
- Watanabe \(2009\)S\. WatanabeAlgebraic geometry and statistical learning theory\.Vol\.25,Cambridge university press,Cambridge\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.20758#S2.SS2a.p4.1),[§3](https://arxiv.org/html/2608.20758#S3.p1.1)\.
- Weiet al\.\(2022\)S\. Wei, D\. Murfet, M\. Gong, H\. Li, J\. Gell\-Redman, and T\. QuellaDeep learning is singular, and that’s good\.IEEE Transactions on Neural Networks and Learning Systems34\(12\),pp\. 10473–10486\.Cited by:[§2\.2](https://arxiv.org/html/2608.20758#S2.SS2a.p4.1)\.
- Wen and Tadmor \(2020\)M\. Wen and E\. B\. TadmorUncertainty quantification in molecular simulations with dropout neural network potentials\.npj computational materials6\(1\),pp\. 124\.Cited by:[§3](https://arxiv.org/html/2608.20758#S3.p3.1)\.
- Wenzelet al\.\(2020\)F\. Wenzel, K\. Roth, B\. S\. Veeling, J\. Świątkowski, L\. Tran, S\. Mandt, J\. Snoek, T\. Salimans, R\. Jenatton, and S\. NowozinHow good is the bayes posterior in deep neural networks really?\.arXiv preprint arXiv:2002\.02405\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p2.1),[§1](https://arxiv.org/html/2608.20758#S1.p3.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1.p3.1),[§2\.1](https://arxiv.org/html/2608.20758#S2.SS1a.p5.1)\.
- Xuet al\.\(2018\)K\. Xu, W\. Hu, J\. Leskovec, and S\. JegelkaHow powerful are graph neural networks?\.arXiv preprint arXiv:1810\.00826\.Cited by:[§1\.1](https://arxiv.org/html/2608.20758#S1.SS1.p1.1),[§2](https://arxiv.org/html/2608.20758#S2.p1.1),[§4\.1\.1](https://arxiv.org/html/2608.20758#S4.SS1.SSS1.p1.1)\.
- Yaoet al\.\(2019\)J\. Yao, W\. Pan, S\. Ghosh, and F\. Doshi\-VelezQuality of uncertainty quantification for bayesian neural network inference\.arXiv preprint arXiv:1906\.09686\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p3.1)\.
- Yurtseveret al\.\(2020\)E\. Yurtsever, J\. Lambert, A\. Carballo, and K\. TakedaA survey of autonomous driving: common practices and emerging technologies\.IEEE access8,pp\. 58443–58469\.Cited by:[§1](https://arxiv.org/html/2608.20758#S1.p1.1)\.
- Zenget al\.\(2018\)J\. Zeng, A\. Lesnikowski, and J\. M\. AlvarezThe relevance of bayesian layer positioning to model uncertainty in deep bayesian active learning\.arXiv preprint arXiv:1811\.12535\.Cited by:[§2](https://arxiv.org/html/2608.20758#S2.p1.1)\.

## Supplementary Information

## 1Experimental details

### 1\.1Model architecture

To develop the Bayesian GNN, we first construct a backbone GNN composed of five sequential graph\-convolutional layers, followed by a pooling operation and a fully connected layer \(Fig\.[1](https://arxiv.org/html/2608.20758#S1.F1)\)\. Each molecule is represented as a graph, where the nodes encode 37 atomic features listed in Table[1](https://arxiv.org/html/2608.20758#S1.T1), and edges represent the covalent bonds between atoms\. The node features are propagated and aggregated across local neighborhoods through convolutional operations using a Graph Isomorphism Network \(GIN\)[75](https://arxiv.org/html/2608.20758#bib.bib21);[51](https://arxiv.org/html/2608.20758#bib.bib29)\. The resulting graph embeddings are then pooled to generate a molecular fingerprint \(latent representation\), which is subsequently fed into a fully connected layer to predict the target physicochemical property\. Architectural details, including the number of layers and dimensions of hidden channels, are summarized in Table[2](https://arxiv.org/html/2608.20758#S1.T2)\.

Figure 1:Schematic of Bayesian GNN models developed in this study\. \(a\) Backbone GNN model and \(b\) BGNN model\.![Refer to caption](https://arxiv.org/html/2608.20758v1/BGNN_architecture_unary.png)Table 1:Description of utilized atomic features\.Table 2:Detailed description of model architecture\.The Bayesian GNN is implemented by reparameterizing each weight in the final fully connected layer as a Gaussian distribution \(Fig\.[1](https://arxiv.org/html/2608.20758#S1.F1)b\):

𝐰\\displaystyle\\mathbf\{w\}=μ𝐰\+σ𝐰⋅ϵ\\displaystyle=\\mu\_\{\\mathbf\{w\}\}\+\\sigma\_\{\\mathbf\{w\}\}\\cdot\\epsilon\(1\)𝐛\\displaystyle\\mathbf\{b\}=μ𝐛\+σ𝐛⋅ϵ\\displaystyle=\\mu\_\{\\mathbf\{b\}\}\+\\sigma\_\{\\mathbf\{b\}\}\\cdot\\epsilon\(2\)
whereϵ∼𝒩⁡\(0,12\)\\epsilon\\sim\\mathcal\{N\}\\left\(0,1^\{2\}\\right\)denotes random samples drawn from a standard normal distribution\. Under this formulation, each forward pass generates a distinct realization of the weights by samplingϵ\\epsilon\. This stochastic construction yields an implicit ensemble effect within a single BGNN and enhances robustness, particularly when available data are sparse or contaminated with noise[6](https://arxiv.org/html/2608.20758#bib.bib3);[30](https://arxiv.org/html/2608.20758#bib.bib25);[58](https://arxiv.org/html/2608.20758#bib.bib26)\. A separate BGNN model is constructed for each physicochemical property, following standard property\-specific modeling practices\. All models used in this study were implemented using PyTorch v2\.7\.0[60](https://arxiv.org/html/2608.20758#bib.bib27), and molecular processing was performed using RDKit v2024\.09\.06[40](https://arxiv.org/html/2608.20758#bib.bib28)\.

### 1\.2Model training

To approximate the intractable posterior distributions of the weights and biases, we employed Variational Inference \(VI\) with a mean\-field approximation[6](https://arxiv.org/html/2608.20758#bib.bib3);[67](https://arxiv.org/html/2608.20758#bib.bib32);[16](https://arxiv.org/html/2608.20758#bib.bib31)\. The model parameters were optimized by minimizing the negative Evidence Lower Bound \(ELBO\), which consists of a reconstruction error term \(negative log\-likelihood\) and a regularization term \(KL\-divergence\)[35](https://arxiv.org/html/2608.20758#bib.bib34)\. The total training objectiveℒ⁡\(θ\)\\mathcal\{L\}\(\\theta\)is defined as:

ℒ⁡\(θ\)=𝔼qθ​\(w,b\)​\[−log⁡p⁡\(𝒟\|w,b\)\]⏟NLL\+β​\[KL\(qθ\(w\)∥p\(w\)\)\+KL\(qθ\(b\)∥p\(b\)\)\]⏟Regularization\\displaystyle\\mathcal\{L\}\(\\theta\)=\\underbrace\{\\mathbb\{E\}\_\{q\_\{\\theta\}\(w,b\)\}\\left\[\-\\log\{p\(\\mathcal\{D\}\|w,b\)\}\\right\]\}\_\{\\text\{NLL\}\}\+\\beta\\underbrace\{\\left\[KL\(q\_\{\\theta\}\(w\)\\\|p\(w\)\)\+KL\(q\_\{\\theta\}\(b\)\\\|p\(b\)\)\\right\]\}\_\{\\text\{Regularization\}\}\(3\)where𝒟\\mathcal\{D\}denotes the training dataset, andβ\\betais a weighting factor balancing the regularization strength\. The detailed derivation of the ELBO is provided in Supplementary Note[3](https://arxiv.org/html/2608.20758#S3a)\. To accelerate convergence and enhance training stability, the backbone GNN model is first trained, after which its learned weights are transferred to the BGNN to provide a warm start\. To ensure statistical robustness and reproducibility, all experiments were repeated across 30 independent runs for every configuration\. Detailed training configurations are summarized in Table[3](https://arxiv.org/html/2608.20758#S1.T3)\. The parity plots and prediction errors for each physicochemical property are presented in Fig\.[2](https://arxiv.org/html/2608.20758#S1.F2)\.

Table 3:Detailed description of model training configuration\.Figure 2:Parity plots for the physicochemical properties examined in this study: \(a\) boiling point \(bp\), \(b\) melting point \(mp\), \(c\) vapor pressure at 25∘C\{\}^\{\\circ\}C\(Pvap\), \(d\) partition coefficient \(P\), \(e\) standard enthalpy of formation \(H\), and \(f\) standard Gibbs free energy of formation \(G\)\. The mean absolute error \(MAE\) for each property is annotated\. Marker color denotes the amine type, and marker size represents molecular weight \[g/mol\]\. Amine types are assigned priority in the order: tertiary\>\>secondary\>\>primary, when a compound belongs to multiple classes\.![Refer to caption](https://arxiv.org/html/2608.20758v1/parity_plot_unary.png)
### 1\.3Datasets

Training datasets were compiled from publicly accessible sources and carefully curated to ensure internal consistency\. For molecules reported with only minor discrepancies in property values, the corresponding entries were averaged, whereas data points exhibiting substantial inconsistencies were removed\. Eighty percent of the collected data points were allocated for model training, with the remaining portion reserved for testing\. A detailed description of all datasets used in this study, including their sources, sizes, and preprocessing transformations, is provided in Table[4](https://arxiv.org/html/2608.20758#S1.T4)\.

Table 4:Description of datasets and their properties\.

## 2Theoretical context for posterior contraction in the present setting

In this section, we summarize the classical Bayesian asymptotic setting and discuss its applicability to the model studied here\. We first outline the assumptions of the Bernstein\-von Mises \(BvM\) theorem and then describe how the present setting, which uses a mean\-field variational approximation over a Bayesian output layer coupled to a deterministic representation learned from the same data, differs from the pre\-specified exact\-posterior setting considered by the classical theorem\.

### 2\.1The Bernstein\-von Mises Theorem

Consider a parametric model family\{p⁡\(x\|𝜽\):𝜽∈Θ⊂ℝd\}\\\{p\(x\|\\boldsymbol\{\\theta\}\):\\boldsymbol\{\\theta\}\\in\\Theta\\subset\\mathbb\{R\}^\{d\}\\\}and a dataset𝒟n=\{x1,…,xn\}\\mathcal\{D\}\_\{n\}=\\\{x\_\{1\},\\dots,x\_\{n\}\\\}consisting of independent and identically distributed \(i\.i\.d\.\) samples from a true distributionp0​\(x\)p\_\{0\}\(x\), whereddandnndenote the number of model parameters and data points, respectively\. Under standard regularity conditions \(detailed in Section[2\.2](https://arxiv.org/html/2608.20758#S2.SS2a)\), the Bernstein\-von Mises theorem states that asn→∞n\\to\\infty, the posterior distributionp⁡\(𝜽\|𝒟n\)p\(\\boldsymbol\{\\theta\}\|\\mathcal\{D\}\_\{n\}\)converges in total variation \(TV\) distance to a multivariate Gaussian distribution centered at the maximum likelihood estimator𝜽^n\\hat\{\\boldsymbol\{\\theta\}\}\_\{n\}[69](https://arxiv.org/html/2608.20758#bib.bib38);[21](https://arxiv.org/html/2608.20758#bib.bib39):

‖p⁡\(𝜽\|𝒟n\)−𝒩⁡\(𝜽^n,Σn\)‖T​V→0\\displaystyle\\\|p\(\\boldsymbol\{\\theta\}\|\\mathcal\{D\}\_\{n\}\)\-\\mathcal\{N\}\(\\hat\{\\boldsymbol\{\\theta\}\}\_\{n\},\\Sigma\_\{n\}\)\\\|\_\{TV\}\\to 0\(4\)
where the asymptotic covarianceΣn\\Sigma\_\{n\}is given by the inverse of the Fisher Information Matrix \(FIM\), scaled by the sample size:

Σn=1n​I​\(𝜽0\)−1\\Sigma\_\{n\}=\\frac\{1\}\{n\}I\(\\boldsymbol\{\\theta\}\_\{0\}\)^\{\-1\}Here,I⁡\(𝜽0\)I\(\\boldsymbol\{\\theta\}\_\{0\}\)is the Fisher Information Matrix at the true parameter𝜽0\\boldsymbol\{\\theta\}\_\{0\}, defined as:

I⁡\(𝜽0\)=𝔼x∼p0​\[−∇𝜽2​log⁡p⁡\(x\|𝜽\)\|𝜽=𝜽0\]\\displaystyle I\(\\boldsymbol\{\\theta\}\_\{0\}\)=\\mathbb\{E\}\_\{x\\sim p\_\{0\}\}\\left\[\-\\nabla\_\{\\boldsymbol\{\\theta\}\}^\{2\}\\log p\(x\|\\boldsymbol\{\\theta\}\)\\big\|\_\{\\boldsymbol\{\\theta\}=\\boldsymbol\{\\theta\}\_\{0\}\}\\right\]\(5\)
In this classical regime, the posterior variance scales asO⁡\(1/n\)O\(1/n\), and the posterior mass increasingly concentrates around𝜽0\\boldsymbol\{\\theta\}\_\{0\}asnnincreases\. This result provides the theoretical basis for the conventional expectation that more data yield a tighter posterior under the regularity conditions of the theorem[69](https://arxiv.org/html/2608.20758#bib.bib38);[21](https://arxiv.org/html/2608.20758#bib.bib39);[6](https://arxiv.org/html/2608.20758#bib.bib3);[74](https://arxiv.org/html/2608.20758#bib.bib6)\.

### 2\.2Applicability of classical assumptions to the present model

The convergence guarantee of the BvM theorem relies on three structural assumptions[69](https://arxiv.org/html/2608.20758#bib.bib38);[21](https://arxiv.org/html/2608.20758#bib.bib39): \(1\) the model is well\-specified, \(2\) the parameter space is finite\-dimensional and fixed, and \(3\) the model is regular, with a positive\-definite Fisher Information Matrix\. These assumptions are not automatically inherited by the present training formulation, and the following differences limit a direct application of the classical result\.

1\.Possible misspecification \(difference from the “well\-specified” setting\):Classical theory assumes that the model family\{p⁡\(x\|𝜽\):𝜽∈Θ\}\\\{p\(x\|\\boldsymbol\{\\theta\}\):\\boldsymbol\{\\theta\}\\in\\Theta\\\}contains a parameter𝜽0\\boldsymbol\{\\theta\}\_\{0\}that represents the data\-generating distributionp0​\(x\)p\_\{0\}\(x\)\. Neural\-network models used for real\-world molecular data may not satisfy this assumption exactly[9](https://arxiv.org/html/2608.20758#bib.bib52);[3](https://arxiv.org/html/2608.20758#bib.bib53)\. Under misspecification, however, a posterior may still concentrate around a pseudo\-true parameter rather than remain diffuse[36](https://arxiv.org/html/2608.20758#bib.bib36);[47](https://arxiv.org/html/2608.20758#bib.bib51)\. Misspecification alone therefore does not imply persistent posterior broadening, but it changes the target and conditions of posterior concentration\.

2\.Over\-parameterization and learned representations \(difference from the “finite\-dimensional and fixed” setting\):The deterministic GNN feature extractor can be highly parameterized, whereas the variational distribution analyzed in this study is placed only over the fixed\-dimensional Bayesian output layer\. Full\-network over\-parameterization therefore does not by itself establish a failure of posterior contraction in the output layer\. A more relevant distinction is that the effective design supplied to the Bayesian readout, namely the learned representation𝐳\\mathbf\{z\}, changes with the training data and optimization\. The output\-layer posterior is consequently coupled to a data\-dependent representation rather than to a pre\-specified fixed design\.

3\.Regularity and approximate inference \(difference from the classical exact\-posterior setting\):Deep neural networks can be singular when their complete parameterization is considered, owing to non\-identifiability and parameter symmetries[71](https://arxiv.org/html/2608.20758#bib.bib40);[72](https://arxiv.org/html/2608.20758#bib.bib54)\. In the present model, however, the analyzed distribution is a mean\-field variational approximation over the Bayesian output layer rather than an exact posterior over the full network\. Conditional on a learned representation, the linear output layer may be regular when its effective design is full rank\. The singularity of the complete deterministic network therefore does not by itself prove that the output\-layer posterior must remain diffuse\. Instead, the classical BvM result does not directly require monotonic contraction of aβ\\beta\-weighted mean\-field variational posterior coupled to a jointly learned representation\. The posterior broadening observed here should accordingly be interpreted as an empirical property of the tested architecture and training objective, rather than as a violation of classical posterior\-contraction theory\.

## 3Derivation of the Evidence Lower Bound \(ELBO\) objective

In the Bayesian neural network model employed in this study, both weightswwand biasesbbare treated as random variables with prior distributionsp⁡\(w\)p\(w\)andp⁡\(b\)p\(b\)\. Given a dataset𝒟=\{xi,yi\}i=1N\\mathcal\{D\}=\\left\\\{x\_\{i\},y\_\{i\}\\right\\\}\_\{i=1\}^\{N\}, the joint distribution is:

pθ​\(𝒟,w,b\)=p⁡\(w\)​p​\(b\)​∏i=1Npθ​\(yi\|xi,w,b\)\\displaystyle p\_\{\\theta\}\(\\mathcal\{D\},w,b\)=p\(w\)p\(b\)\\prod\_\{i=1\}^\{N\}p\_\{\\theta\}\(y\_\{i\}\|x\_\{i\},w,b\)\(6\)
Since the true posteriorp⁡\(w,b\|𝒟\)p\(w,b\|\\mathcal\{D\}\)is intractable, we introduce a variational approximationqΦ​\(w,b\)=qΦ​\(w\)​qΦ​\(b\)q\_\{\\Phi\}\(w,b\)=q\_\{\\Phi\}\(w\)q\_\{\\Phi\}\(b\)\. Then, the evidence lower bound \(ELBO\) is formulated as[5](https://arxiv.org/html/2608.20758#bib.bib16);[31](https://arxiv.org/html/2608.20758#bib.bib69):

log⁡pθ​\(𝒟\)\\displaystyle\\log\{p\_\{\\theta\}\(\\mathcal\{D\}\)\}=log∫pθ\(𝒟,w,b\)dwdb\\displaystyle=\\log\{\\int p\_\{\\theta\}\(\\mathcal\{D\},w,b\)dwdb\}\(7\)≥𝔼qΦ\[logpθ\(𝒟\|w,b\)\]−KL\(qΦ\(w,b\)∥p\(w,b\)\)\\displaystyle\\geq\\mathbb\{E\}\_\{q\_\{\\Phi\}\}\\left\[\\log\{p\_\{\\theta\}\(\\mathcal\{D\}\|w,b\)\}\\right\]\-KL\(q\_\{\\Phi\}\(w,b\)\\\|p\(w,b\)\)\(8\)
Expanding the terms gives:

ELBO=∑i=1N𝔼qΦ​\(w,b\)\[logpθ\(yi\|xi,w,b\)\]−KL\(qΦ\(w\)∥p\(w\)\)−KL\(qΦ\(b\)∥p\(b\)\)\\displaystyle ELBO=\\sum\_\{i=1\}^\{N\}\\mathbb\{E\}\_\{q\_\{\\Phi\}\(w,b\)\}\\left\[\\log\{p\_\{\\theta\}\(y\_\{i\}\|x\_\{i\},w,b\)\}\\right\]\-KL\(q\_\{\\Phi\}\(w\)\\\|p\(w\)\)\-KL\(q\_\{\\Phi\}\(b\)\\\|p\(b\)\)\(9\)
In practice, a weighting factorβ\\betais applied to the KL term to balance regularization strength when minimizing the negative ELBO:

ℒ\(θ\)=𝔼qθ​\(w,b\)\[−logp\(𝒟\|w,b\)\]\+β\[KL\(qθ\(w\)∥p\(w\)\)\+KL\(qθ\(b\)∥p\(b\)\)\]\\displaystyle\\mathcal\{L\}\(\\theta\)=\\mathbb\{E\}\_\{q\_\{\\theta\}\(w,b\)\}\\left\[\-\\log\{p\(\\mathcal\{D\}\|w,b\)\}\\right\]\+\\beta\\left\[KL\(q\_\{\\theta\}\(w\)\\\|p\(w\)\)\+KL\(q\_\{\\theta\}\(b\)\\\|p\(b\)\)\\right\]\(10\)

## 4Negligible impact of bias posterior distributions

\(a\)Boiling point \(bp\)
\(b\)Melting point \(mp\)
\(c\)Partition coefficient \(P\)
\(d\)Vapor pressure \(Pvap\)
\(e\)Enthalpy \(H\)
\(f\)Gibbs free energy \(G\)

Figure 3:Negligible influence of bias posteriors on predictive uncertainty\.Analyses are presented for six different datasets \(a\-f\), where each panel displays the evolution of bias posterior mean, posterior standard deviation, and relative uncertainty contribution \(blue for bias, orange for weight\) as a function of training data size\.To ensure that our analysis of uncertainty dynamics is not confounded by the bias posterior, we examined the behavior of the bias distribution across multiple datasets\. We observed that the posterior mean \(μb\\mu\_\{b\}, Fig\.[3](https://arxiv.org/html/2608.20758#S4.F3)\) converges closely to zero, while the standard deviation \(σb\\sigma\_\{b\}\) exhibits negligible magnitude and inconsistent fluctuations with respect to dataset size\. Unlike the weight posteriors, bias uncertainty shows no clear pattern correlating with predictive uncertainty, suggesting it does not play a governing role in data–uncertainty dynamics\. We further quantified this observation by analytically decomposing the total predictive uncertainty\. The predictive variance over a dataset𝒟\\mathcal\{D\}can be expressed as the sum of contributions from the weights and the bias:

Var⁡\(y∣𝐳,𝒟\)\\displaystyle\\mathrm\{Var\}\(y\\mid\\mathbf\{z\},\\mathcal\{D\}\)=𝐳⊤​Cov​\[𝐰\|𝒟\]​𝐳\+σb2\\displaystyle=\\mathbf\{z\}^\{\\top\}\\mathrm\{Cov\}\\left\[\\mathbf\{w\}\|\\mathcal\{D\}\\right\]\\mathbf\{z\}\+\\sigma\_\{b\}^\{2\}\(11\)=∑i=1dzi2​σi2⏟w​e​i​g​h​t\+σb2⏟b​i​a​s\\displaystyle=\\underbrace\{\\sum\_\{i=1\}^\{d\}z\_\{i\}^\{2\}\\sigma\_\{i\}^\{2\}\}\_\{weight\}\+\\underbrace\{\\sigma\_\{b\}^\{2\}\}\_\{bias\}\(12\)
whereddis the latent dimension andσi2\\sigma\_\{i\}^\{2\}is the variance of theii\-th weight parameter\. In this decomposition, the first term represents the contribution of the weight posterior coupled with the latent representation, while the second term \(σb2\\sigma\_\{b\}^\{2\}\) represents the bias contribution\. As illustrated in Fig\.[3](https://arxiv.org/html/2608.20758#S4.F3), the weight contribution \(orange bar\) consistently dominates the bias contribution \(blue bar\) by a substantial margin across all datasets and training sizes\. Consequently, the bias posterior exerts a practically insignificant influence on output uncertainty, justifying our focus on the weight posterior and latent alignment as the primary drivers of uncertainty\.

## 5Cross\-dataset consistency of posterior–predictive decoupling

To corroborate the decoupled behavior discovered in the main text, we repeated the posterior analysis for the other property datasets\. We tracked the evolution of predictive performance \(MAE\), output uncertainty \(σy\\sigma\_\{y\}\), posterior uncertainty, and posterior mean as a function of training data size\. As shown in Fig\.[4](https://arxiv.org/html/2608.20758#S5.F4), all other datasets exhibit behavioral trends consistent with the primary analysis\. Both the MAE and overall output uncertainty decrease as the training set grows, indicating improved model performance and confidence\. However, the average standard deviation of weight posteriors systematically increases with data accumulation, directly contradicting the conventional expectation of posterior concentration\. Throughout the process, the posterior means remain stable near zero\. These results confirm that the non\-classical behavior—the decoupling of predictive uncertainty from posterior uncertainty—is not an artifact of a specific task but a reproducible pattern across the six tested molecular\-regression datasets under the same architecture and training protocol\.

\(a\)Boiling point \(bp\)
\(b\)Melting point \(mp\)
\(c\)Vapor pressure \(Pvap\)
\(d\)Enthalpy \(H\)
\(e\)Gibbs free energy \(G\)

Figure 4:Verification of the non\-classical uncertainty behavior\.Each row \(a\-e\) corresponds to a distinct dataset, while columns show metrics as a function of training data size \(from left to right\): Mean Absolute Error \(MAE\), predictive uncertainty \(σy\\sigma\_\{y\}\), weight posterior uncertainty \(σw\\sigma\_\{w\}\), and weight posterior mean \(μw\\mu\_\{w\}\)\. Consistent with the main analysis, predictive uncertainty \(second column\) decreases while posterior uncertainty \(third column\) increases across all tasks\.
## 6Latent\-Posterior Alignment and predictive uncertainty

ForNNlatent vectors of dimensiondd,

Z=\[z1,…​zN\]⊤∈ℝN×d,Cz=1N​Z⊤​Z,Σw=diag​\(σw2\)\.\\displaystyle Z=\[z\_\{1\},\.\.\.z\_\{N\}\]^\{\\top\}\\in\\mathbb\{R\}^\{N\\times d\},\\\>C\_\{z\}=\\frac\{1\}\{N\}Z^\{\\top\}Z,\\\>\\Sigma\_\{w\}=\\text\{diag\}\(\\sigma\_\{w\}^\{2\}\)\.\(13\)
Averaging the predictive uncertainty over samples gives

1N​∑i=1NVar⁡\(yi∣zi\)\\displaystyle\\frac\{1\}\{N\}\\sum^\{N\}\_\{i=1\}\\mathrm\{Var\}\(y\_\{i\}\\mid z\_\{i\}\)=t​r​\(Cz​Σw\)\+σb2\\displaystyle=tr\(C\_\{z\}\\Sigma\_\{w\}\)\+\\sigma\_\{b\}^\{2\}\(14\)=∑j=1d\(1N​∑i=1Nzi,j2\)​σj2\+σb2\\displaystyle=\\sum\_\{j=1\}^\{d\}\(\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}z\_\{i,j\}^\{2\}\)\\sigma\_\{j\}^\{2\}\+\\sigma\_\{b\}^\{2\}\(15\)=∑j=1dmj​σj2\+σb2,\\displaystyle=\\sum\_\{j=1\}^\{d\}m\_\{j\}\\sigma\_\{j\}^\{2\}\+\\sigma\_\{b\}^\{2\},\(16\)wheremj=1N​∑izi,j2m\_\{j\}=\\frac\{1\}\{N\}\\sum\_\{i\}z\_\{i,j\}^\{2\}denotes the latent energy along coordinatejj\. Sinceσb2\\sigma\_\{b\}^\{2\}is empirically negligible, the uncertainty is dominated by the coupling term∑jmj​σj2\\sum\_\{j\}m\_\{j\}\\sigma\_\{j\}^\{2\}\. To analyze how the Latent\-Posterior Alignment \(LPA\) affects overall uncertainty, we consider two practical cases:

1. 1\.Fixed posterior variance spectrum \(eigenvalue\-preserving\): the multiset\{σj2\}\\\{\\sigma\_\{j\}^\{2\}\\\}is fixed\.
2. 2\.Budget constraints: either∑jσj2=C1\\sum\_\{j\}\\sigma\_\{j\}^\{2\}=C\_\{1\}\(fixed total variance\) or∑jσj=C2\\sum\_\{j\}\\sigma\_\{j\}=C\_\{2\}\(fixed total standard deviation\)\.

Both assumptions are practically consistent with how BNNs behave during training: the KL term in the variational objective acts as a soft regularizer toward the prior, keeping either the spectral shape ofΣw\\Sigma\_\{w\}stable \(Case 1\) or the overall posterior variance∑jσj2\\sum\_\{j\}\\sigma\_\{j\}^\{2\}approximately constant \(Case 2\)\. Hence, treating the posterior as having a fixed or budget\-constrained variance profile is not an artificial simplification but reflects a realistic equilibrium of variational Bayesian training\. These cases are idealized constraints used to isolate the effect of latent–posterior pairing; the training objective does not enforce either condition exactly\.

Case 1\. Fixed posterior variance spectrum

Ifmjm\_\{j\}andσj2\\sigma\_\{j\}^\{2\}are fixed nonnegative sequences, von Neumann’s trace inequality provides bounds on the overall uncertainty[50](https://arxiv.org/html/2608.20758#bib.bib30):

∑j=1dmj↑​σj↓2≤t​r​\(Cz​Σw\)=∑j=1dmj​σj2≤∑j=1dmj↑​σj↑2\\displaystyle\\sum\_\{j=1\}^\{d\}m\_\{j\}^\{\\uparrow\}\\sigma\_\{j\}^\{\\downarrow 2\}\\leq tr\(C\_\{z\}\\Sigma\_\{w\}\)=\\sum\_\{j=1\}^\{d\}m\_\{j\}\\sigma\_\{j\}^\{2\}\\leq\\sum\_\{j=1\}^\{d\}m\_\{j\}^\{\\uparrow\}\\sigma\_\{j\}^\{\\uparrow 2\}\(17\)
The lower bound—corresponding to minimal predictive uncertainty—is achieved when largemjm\_\{j\}are paired with smallσj2\\sigma\_\{j\}^\{2\}, i\.e\., when the latent covarianceCzC\_\{z\}and weight variance matrixΣw\\Sigma\_\{w\}are aligned\. This condition formalizes the notion of the latent alignment, where the model reduces predictive uncertainty by directing high\-energy latent dimensions toward low\-uncertainty weight directions\.

Case 2\. Budget constraints

When total variance is fixed,

∑j=1dmj​σj2s\.t\.∑jσj2=C1\\displaystyle\\sum\_\{j=1\}^\{d\}m\_\{j\}\\sigma\_\{j\}^\{2\}\\quad\\text\{s\.t\.\}\\quad\\sum\_\{j\}\\sigma\_\{j\}^\{2\}=C\_\{1\}\(18\)the optimization is a linear program over a simplex\. Its optimum lies at an extreme point, concentrating all variance on the dimensionj⋆∈argminj​mjj^\{\\star\}\\in\\text\{argmin\}\_\{j\}m\_\{j\}\. Thus, predictive variance is minimized when posterior uncertainty is concentrated along latent directions with the smallest energy—an extreme form of the alignment\.

Under fixed total standard deviation,

∑j=1dmj​σj2s\.t\.∑jσj=C2\\displaystyle\\sum\_\{j=1\}^\{d\}m\_\{j\}\\sigma\_\{j\}^\{2\}\\quad\\text\{s\.t\.\}\\quad\\sum\_\{j\}\\sigma\_\{j\}=C\_\{2\}\(19\)the objective is strictly convex\. From the KKT conditions for equality constraints,2​mj​σj\+λ=02m\_\{j\}\\sigma\_\{j\}\+\\lambda=0whereλ\\lambdais a Lagrange multiplier, implyingσj\\sigma\_\{j\}andmjm\_\{j\}are inversely proportional to minimize uncertainty, thereby concentratingσj\\sigma\_\{j\}on smallmjm\_\{j\}\. Hence, both budget\-constrained formulations yield the same aligned solution, confirming that predictive uncertainty is minimized when posterior variance aligns inversely with latent energy\. These results indicate that the model structurally reorganizes the latent representations during training to minimize the predictive uncertainty\.

## 7Cross\-dataset consistency of Latent\-Posterior Alignment

To confirm that Latent\-Posterior Alignment \(LPA\) is a consistent pattern, we extended our analysis to other property datasets \(Fig\.[5](https://arxiv.org/html/2608.20758#S7.F5)\)\. First, a visual comparison between the data\-poor \(1%, first column\) and data\-rich \(80%, second column\) regimes reveals a consistent emergence of Latent\-Posterior Alignment across all tasks\. In the data\-rich regime, the normalized latent vectors \(𝐳~\\tilde\{\\mathbf\{z\}\}, blue\) shift toward dimensions where the posterior standard deviation \(σ~w\\tilde\{\\sigma\}\_\{w\}, orange\) is minimized\. This geometric reorganization confirms that the model actively learns to utilize stable weight dimensions as data accumulates\.

Second, this shift fundamentally alters the composition of predictive uncertainty\. The relative contribution analysis \(CiC\_\{i\}, grey bars\) demonstrates that as the model utilizes more data, the uncertainty budget is increasingly reallocated toward these low\-variance dimensions\. While the large magnitudes of some high\-variance dimensions inevitably contribute to total uncertainty, the dominant trend is a systematic shift of contribution toward “reliable” axes, confirming that the model anchors its predictions on stable dimensions\.

Third, this structural evolution is quantitatively validated by the Latent\-Posterior Alignment Score \(LPAS, third column\)\. Across all datasets, the LPAS increases with training\-set size across the tested datasets, providing robust evidence that LPA is the key strategy\.

Finally, the analysis of high\-uncertainty samples corroborates the dual\-geometric mechanism proposed in the main text\. We compared the latent representations of the entire dataset \(blue\) against the top 5% of samples with the highest predictive uncertainty \(green\)\. Crucially, high\-uncertainty samples maintain the same aligned directionality while exhibiting significantly amplified latent magnitudes, as shown in the fourth column of Fig\.[5](https://arxiv.org/html/2608.20758#S7.F5)\. This confirms that the model regulates uncertainty through a two\-tiered strategy: global alignment sets a stable baseline, while latent magnitude encodes instance\-specific epistemic risk\.

\(a\)Boiling point \(bp\)
\(b\)Melting point \(mp\)
\(c\)Vapor pressure \(Pvap\)
\(d\)Enthalpy \(H\)
\(e\)Gibbs free energy \(G\)

Figure 5:Cross\-dataset consistency of Latent\-Posterior Alignment\.Results are presented for five different datasets \(a\-e\)\. Columns 1–2: Comparison of latent geometry between data\-poor \(1%\) and data\-rich \(80%\) regimes\. The plots overlay the normalized latent vector \(𝐳~\\tilde\{\\mathbf\{z\}\}, blue\), posterior standard deviation \(σ~w\\tilde\{\\sigma\}\_\{w\}, orange\), and relative uncertainty contribution \(grey bars\)\. Column 3: The Latent\-Posterior Alignment Score \(LPAS\) versus data size\. Column 4: Comparison of latent structure between the overall data distribution \(blue\) and samples with the highest predictive uncertainty \(green\)\.
## 8Formulation of latent regularization objectives

To test the causal role of Latent\-Posterior Alignment \(LPA\) in predictive uncertainty, we introduced L1, L2, and anti\-alignment penalties directly into the standard ELBO objective to inhibit the formation of LPA as follows:

ℒL​1=−ELBO\+λl​1​‖z‖1\\displaystyle\\mathcal\{L\}\_\{L1\}=\-\\text\{ELBO\}\+\\lambda\_\{l1\}\\\|z\\\|\_\{1\}\(20\)
ℒL​2=−ELBO\+λl​2​‖z‖2\\displaystyle\\mathcal\{L\}\_\{L2\}=\-\\text\{ELBO\}\+\\lambda\_\{l2\}\\\|z\\\|\_\{2\}\(21\)
ℒa​n​t​i−a​l​i​g​n=−ELBO−λa​n​t​i−a​l​i​g​n​∑i=1d\|zi\|​σi∑i=1d\|zi\|​∑i=1dσi\\displaystyle\\mathcal\{L\}\_\{anti\-align\}=\-\\text\{ELBO\}\-\\lambda\_\{anti\-align\}\\frac\{\\sum\_\{i=1\}^\{d\}\|z\_\{i\}\|\\sigma\_\{i\}\}\{\\sum\_\{i=1\}^\{d\}\|z\_\{i\}\|\\sum\_\{i=1\}^\{d\}\\sigma\_\{i\}\}\(22\)
The first two interventions \(L1 and L2\) aim to suppress the capacity of the latent vector𝐳\\mathbf\{z\}to form strong directional concentrations or large magnitudes, which are essential for LPA\. Specifically, the L1 term enforces a strict sparsity penalty, while the L2 term penalizes the Euclidean norm of𝐳\\mathbf\{z\}, thereby hindering the formation of sharp alignment\. The anti\-alignment penalty term utilizes the negative value of the Latent\-Posterior Alignment Score \(LPAS\) to directly test causality\. This constraint actively pushes the latent representation into uncertain dimensions whereσw\\sigma\_\{w\}is large, forbidding the model from utilizing the reliable subspaces it naturally seeks\. Since the magnitudes of the ELBO and regularization terms vary across different datasets and properties, the regularization strengthλ\\lambdawas empirically tuned for each task\. Regularization strengths were selected using a pre\-specified grid and validation data only\. The test set was not used to selectλ\\lambda\. The specific hyperparameters used for all experiments are shown in Table[5](https://arxiv.org/html/2608.20758#S8.T5)whileγ\\gammais the regularization weight for Alignment\-Guided Learning\.

Table 5:Hyperparameters for latent regularization experiments and Alignment\-Guided Learning\.The regularization and AGL strengths \(λ\\lambdaandγ\\gamma\) were adjusted for each dataset\.
## 9Extended interventional evidence across datasets

Using the formulation in Section[8](https://arxiv.org/html/2608.20758#S8), we verified the causal role of LPA by extending the interventional experiments to five additional datasets\. Fig\.[6](https://arxiv.org/html/2608.20758#S9.F6)shows the impact of three regularization schemes \(L1, L2, and anti\-alignment penalty\) on model performance and latent geometry\. Consistent with the main analysis, all regularized models exhibit degradation in predictive performance \(MAE, Column 1\) and a substantial increase in predictive uncertainty \(σy\\sigma\_\{y\}, Column 2\) compared to the unconstrained baseline model\. This strictly correlates with the breakdown of the LPA mechanism, as evidenced by decreased Latent\-Posterior Alignment Score values \(Column 3\) and the disruption of latent structures \(Columns 4–6\)\. This confirms that preventing the formation of LPA fundamentally deteriorates the model’s ability to minimize uncertainty\.

We also observed a hierarchy in the degradation, generally ordered as anti\-alignment\>\>L1\>\>L2 \(Anti\-alignment penalty and L1 induced greater uncertainty and MAE than L2\)\. This trend can be attributed to the geometric nature of the methods\. The L2 penalty tends to induce a diffuse distribution, where the latent vector retains non\-zero magnitudes across most hidden dimensions \(Column 5\)\. In contrast, L1 and anti\-alignment penalties force the latent representation into a highly sparse or sharp regime, activating only a few dimensions \(Columns 4 and 6\)\. This severe restriction on latent geometry appears to drastically hinder the model’s capacity to fully utilize dimensional information\. Note that the precise ranking of degradation may vary depending on hyperparameter tuning \(λ\\lambda\) specific to each dataset\.

\(a\)Boiling point \(bp\)
\(b\)Melting point \(mp\)
\(c\)Vapor pressure \(Pvap\)
\(d\)Enthalpy \(H\)
\(e\)Gibbs free energy \(G\)

Figure 6:Impact of latent regularization on predictive uncertainty and latent geometry\. Rows \(a\-e\) correspond to five different datasets\. Columns 1\-3: Comparison of MAE, predictive uncertainty, and LPAS between baseline \(dashed line\) and regularized models\. Columns 4\-6: Visualization of latent vector and posterior standard deviation under L1, L2, and anti\-alignment penalties, respectively\.
## 10Robustness analysis

![Refer to caption](https://arxiv.org/html/2608.20758v1/img_2_test_new_2.png)

Figure 7:Robustness of latent\-posterior alignment mechanism\. \(a–f\) Robustness to KL weight\. Reducingβ\\betaleads to smallerσw\\sigma\_\{w\}\(b\)\. However, this variation has negligible impact on accuracy \(a\) and output uncertainty \(c\)\. This invariance arises because LPA serves as a robust geometric mechanism: data\-rich models maintain a consistently high alignment score \(d, purple\) and a clearly aligned latent\-posterior geometry \(f\) regardless ofβ\\beta\. In contrast, data\-poor models exhibit low LPAS \(d, cyan\) and fail to form such structure \(e\), confirming that predictive confidence is governed by geometric alignment rather than posterior magnitude\. \(g\-l\) Impact of prior variance constraints\. Restricting the prior standard deviation \(x\-axis\) forces posterior contraction \(h\) but paradoxically increases predictive error \(g\) and output uncertainty \(i\)\. This failure arises because tight priors restrict the parameter\-space flexibility required to maintain LPA, as evidenced by the decay of alignment score \(j\) and disrupted geometry \(k\) compared to the relaxed prior \(l\)\. Shaded bands denote standard deviation acrossn=30n=30independent runs\.We evaluated the robustness of this alignment mechanism to determine whether it represents a general principle or is contingent upon specific training configurations\. We first investigated the influence of the KL\-divergence weight \(β\\beta\), which governs the strength of posterior regularization\. We hypothesized that the posterior’s failure to govern uncertainty might stem from excessive regularization pressure\. By reducingβ\\beta, we relaxed this constraint to determine whether the posterior would reclaim its role, or if the model would consistently persist in utilizing the Latent\-Posterior Alignment as its dominant strategy\.

As shown in Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)a–f, we compared representative data\-poor \(2% data, high uncertainty, cyan\) and data\-rich \(20% data, low uncertainty, purple\) regimes, analyzing how their respective strategies evolved under varyingβ\\beta\. As expected, loweringβ\\betaweakens the regularization, leading to a reduction in the posterior standard deviation \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)b\)\. This effect is particularly pronounced in the data\-rich regime, where the baseline posterior variance was originally substantial\. Crucially, however, this concentration of the posterior did not translate to the model’s predictive behavior\. Both predictive accuracy \(MAE\) \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)a\) and output uncertainty \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)c\) remained largely invariant to changes inβ\\beta\. Specifically, the data\-rich models maintained a low uncertainty profile despite the drastic tightening of their posteriors\. This suggests that the posterior variance itself is not the primary driver of predictive uncertainty\. Instead, the uncertainty dynamics were consistently governed by the alignment mechanism\.

The LPAS in Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)d reveals a fundamental distinction between the regimes that persists regardless ofβ\\beta\. In the data\-poor case, LPAS values remain consistently low and invariant toβ\\beta, indicating a persistent absence of the alignment mechanism \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)e\)\. Conversely, in the data\-rich case, the models robustly maintain substantially higher LPAS values; although minor variations exist depending onβ\\beta, they exhibit strong alignment \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)f\)\. This confirms that the adoption of Latent\-Posterior Alignment is a robust strategy, independent of the posterior’s flexibility\. We verified that this resilience toβ\\betavariations is consistent across diverse molecular benchmarks \(see Supplementary Fig\.[8](https://arxiv.org/html/2608.20758#S10.F8)\)\.

We further extended this robustness analysis by investigating the effect of the prior distribution for the data\-rich case \(20% data\) as shown in Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)g–l\. We hypothesized that imposing a highly restrictive and low\-variance prior on the weights would force the posterior varianceσw2\\sigma\_\{w\}^\{2\}to shrink\. If posterior variance were indeed the dominant factor, this should directly reduce predictive uncertainty\. As expected, narrowing the prior caused the posterior standard deviations to decrease \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)h\)\. However, despite this reduction in the posterior distribution, predictive uncertainty gradually increased \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)i\)\. This counterintuitive trend arises from the disruption of latent alignment: as the standard deviation of the prior decreases, the alignment score decreases \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)j\), revealing that the latent vectors begin to concentrate toward directions where the posterior variance is large \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)k\), thereby amplifying uncertainty\. Similar results under restrictive priors were reproduced in other tasks \(Supplementary Fig\.[9](https://arxiv.org/html/2608.20758#S10.F9)\)\.

While the model can tolerate moderate constraints, our experiments revealed a critical failure point \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)g–j\)\. When the prior standard deviation is excessively reduced \(e\.g\., below 0\.01\), the model fails to maintain coherent latent alignment \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)k\) compared to the relaxed model \(Fig\.[7](https://arxiv.org/html/2608.20758#S10.F7)l\), leading to a sharp increase in predictive uncertainty and error\. This finding underscores that the model requires a sufficient degree of “exploratory variance” \(parameter\-space flexibility\) to discover and leverage effective alignment pathways\. Artificially suppressing this flexibility proves counterproductive, ultimately disrupting the very mechanism the model relies on for reliable predictions\.

Relaxing posterior regularization did not diminish the role of LPA, whereas attempting to circumvent it via narrow priors proved detrimental\. Supported by our extensive validation across multiple datasets \(see Supplementary Figs\.[6](https://arxiv.org/html/2608.20758#S9.F6),[8](https://arxiv.org/html/2608.20758#S10.F8),[9](https://arxiv.org/html/2608.20758#S10.F9)\), these findings confirm that LPA remains associated with predictive uncertainty, supporting its robust contribution within the deterministic GNN feature extractors and mean\-field Bayesian output layers examined here\.

### 10\.1Effect of KL\-divergence weight on LPA

To examine how robust Latent\-Posterior Alignment \(LPA\) is to posterior regularization, we extended the KL\-divergence weight \(β\\beta\) analysis to five additional physicochemical property datasets \(Fig\.[8](https://arxiv.org/html/2608.20758#S10.F8)\)\. As expected, loweringβ\\betaweakens the regularization pressure, leading to a notable reduction in posterior uncertainty \(Column 2\)\. This effect is regime\-dependent: it is more pronounced in the data\-rich regime \(purple\), where baseline variance is larger, compared to the data\-poor regime \(cyan\)\.

Crucially, however, unlike posterior uncertainty, both predictive accuracy \(MAE, Column 1\) and output uncertainty \(Column 3\) remain largely invariant to changes inβ\\beta\. Specifically, data\-rich models maintain their characteristic low\-uncertainty profile compared to data\-poor models, regardless of whether their posteriors are tight \(lowβ\\beta\) or loose \(highβ\\beta\)\. This invariance strongly suggests that posterior uncertainty itself is not the primary driver of predictive confidence\. Instead, the alignment mechanism consistently governs the uncertainty dynamics\. The Latent\-Posterior Alignment Score \(LPAS, Column 4\) reveals a fundamental distinction: in the data\-poor case, LPAS values remain consistently low and invariant toβ\\beta, corresponding to a lack of LPA structure in the latent dimension \(Column 5\)\. Conversely, in the data\-rich case, the models robustly maintain higher LPAS values, preserving a strong LPA structure \(Column 6\)\. These results demonstrate that Latent\-Posterior Alignment serves as a robust and dominant strategy for controlling uncertainty, operating independently of the posterior’s flexibility\.

\(a\)Boiling point \(bp\)
\(b\)Melting point \(mp\)
\(c\)Vapor pressure \(Pvap\)
\(d\)Enthalpy \(H\)
\(e\)Gibbs free energy \(G\)

Figure 8:Influence of KL\-divergence weight on Latent\-Posterior Alignment and uncertainty dynamics\. Rows \(a\-e\) correspond to five different property datasets\. Cyan and purple represent data\-poor and data\-rich regimes, respectively\. Columns 1\-4 show the KL\-divergence weight impact on MAE, posterior uncertainty \(σw\\sigma\_\{w\}\), predictive uncertainty \(σy\\sigma\_\{y\}\), and LPAS\. Columns 5\-6 visualize the geometry of latent vectors and posterior standard deviations for data\-poor \(Column 5\) and data\-rich \(Column 6\) models\.
### 10\.2Effect of prior variance on LPA

We extended the prior analysis to five property datasets \(Fig\.[9](https://arxiv.org/html/2608.20758#S10.F9)\) to validate the robustness of Latent\-Posterior Alignment \(LPA\) against prior variations\. We focused on the data\-rich regime \(80% training data\) where LPA is active\. First, imposing a narrower prior successfully forced the posterior distributions to concentrate \(Column 2\)\. However, this contraction followed an inverse trajectory relative to model performance\. Both predictive error \(MAE, Column 1\) and output uncertainty \(Column 3\) gradually increased as the prior became more restrictive\. This paradox reconfirms that the magnitude of posterior variance is not the governing factor for predictive uncertainty\.

The deterioration in performance is directly attributable to the disruption of the LPA mechanism\. As the prior standard deviation decreases, LPAS decreases \(Column 4\), indicating weaker alignment under increasingly restrictive priors\. Visualizing latent geometry reveals that under strict prior constraints \(Column 5\), latent vectors are forced into unstable, high\-variance dimensions, breaking the alignment structure observed under relaxed priors \(Column 6\)\. Furthermore, consistent with the main analysis, we identified a critical failure point across all datasets\. While the specific threshold varies by task, excessive suppression of prior variance universally leads to a collapse of the LPA mechanism, resulting in a surge in model error and uncertainty\. This finding demonstrates that a sufficient degree of variance in the parameter space—exploratory variance—is indispensable for the model to discover and maintain effective alignment pathways\.

\(a\)Boiling point \(bp\)
\(b\)Melting point \(mp\)
\(c\)Vapor pressure \(Pvap\)
\(d\)Enthalpy \(H\)
\(e\)Gibbs free energy \(G\)

Figure 9:Impact of prior distribution on Latent\-Posterior Alignment and predictive uncertainty\. Analysis is conducted on the data\-rich regime \(80% train data\) across five physicochemical properties \(a\-e\)\. Columns 1\-4 show the evolution of MAE, posterior uncertainty \(σw\\sigma\_\{w\}\), predictive uncertainty \(σy\\sigma\_\{y\}\), and LPAS as a function of prior standard deviation\. Columns 5\-6 visualize latent vectors under restrictive prior \(Column 5,σp​r​i​o​r=0\.0001\\sigma\_\{prior\}=0\.0001\) and relaxed prior \(Column 6,σp​r​i​o​r=1\\sigma\_\{prior\}=1\)\.

## 11Cross\-dataset evaluation of AGL and density awareness

To confirm the generalizability of Alignment\-Guided Learning \(AGL\), we applied the proposed training scheme to five additional physicochemical property datasets\. We evaluated AGL using the same metrics—LPAS, MAE, posterior standard deviationσw\\sigma\_\{w\}, output uncertaintyσy\\sigma\_\{y\}, and calibration indicators \(ECE, DUC\)—comparing results across data\-poor and data\-rich regimes \(Figs\.[10](https://arxiv.org/html/2608.20758#S11.F10)–[14](https://arxiv.org/html/2608.20758#S11.F14)\)\. All metrics are normalized relative to the baseline, which is scaled to 100\. Hyperparameterγ\\gammafor each dataset is given in Table[5](https://arxiv.org/html/2608.20758#S8.T5)\. For LPAS, normalized values represent the proportional reduction in1−LPAS1\-\\mathrm\{LPAS\}relative to the baseline, rather than the direct percentage change in raw LPAS\.

As shown in Fig\.[10](https://arxiv.org/html/2608.20758#S11.F10), AGL consistently increases the LPAS relative to the baseline \(gray dashed line\) across both data\-poor \(light shaded bars\) and data\-rich \(dark shaded bars\) regimes, successfully reinforcing the Latent\-Posterior Alignment \(LPA\) mechanism as intended \(Figs\.[11](https://arxiv.org/html/2608.20758#S11.F11)\-[12](https://arxiv.org/html/2608.20758#S11.F12)\)\. Consistent with the main analysis, the magnitude of this increase is more pronounced in the data\-rich case \(Fig\.[12](https://arxiv.org/html/2608.20758#S11.F12)\) compared to the data\-poor case \(Fig\.[11](https://arxiv.org/html/2608.20758#S11.F11)\), reflecting the availability of sufficient information to guide geometric structure\.

Crucially, while predictive accuracy \(MAE\) remains stable without degradation, predictive uncertaintyσy\\sigma\_\{y\}is substantially reduced\. This reduction occurs even though posterior uncertaintyσw\\sigma\_\{w\}remains unchanged or slightly increases\. This discrepancy further demonstrates that the geometric alignment mechanism, rather than the posterior itself, is the governing factor of predictive uncertainty\.

The calibration analysis reveals a distinct trade\-off\. While probabilistic calibration \(ECE\) shows mixed results or slight degradation in some cases, structural calibration \(DUC\) exhibits a consistent and significant enhancement\. This suggests that AGL specializes in geometric calibration—aligning high uncertainty with low data density—over simple probability matching\. To visualize this geometric calibration, we projected latent representations via t\-SNE, mapping data density measured by Kernel Density Estimation \(KDE\) against output uncertainty \(σy\\sigma\_\{y\}\) as shown in Figs\.[13](https://arxiv.org/html/2608.20758#S11.F13)\-[14](https://arxiv.org/html/2608.20758#S11.F14)\. The projections reveal that AGL harmonizes the distributions: regions of low data density \(blue\) consistently correspond to high uncertainty \(red\), and vice versa\.

Crucially, the dynamics of this improvement differ by regime\. In the data\-poor case \(Fig\.[13](https://arxiv.org/html/2608.20758#S11.F13)\), the baseline model exhibits a disjoint and quasi\-random uncertainty landscape compared to data density\. Here, AGL fundamentally corrects this relationship, enforcing an alignment between uncertainty and data density\. Meanwhile, in the data\-rich case \(Fig\.[14](https://arxiv.org/html/2608.20758#S11.F14)\), the baseline models already show slight alignment\. Here, AGL refines and sharpens the relationship in boundary regions, yielding a stricter correspondence between uncertainty and data sparsity\. These strong improvements in DUC and structural alignment imply that AGL is particularly advantageous for downstream tasks like Bayesian optimization and active learning, which rely on the fundamental assumption that uncertainty serves as a reliable proxy for data scarcity \(epistemic ignorance\)\.

\(a\)Boiling point \(bp\)\(b\)Melting point \(mp\)\(c\)Vapor pressure \(Pvap\)\(d\)Enthalpy \(H\)\(e\)Gibbs free energy \(G\)
Figure 10:Performance evaluation of Alignment\-Guided Learning \(AGL\) across diverse datasets\. Results are shown for five different properties \(a\-e\)\. The AGL model \(colored bars\) is compared with the baseline \(gray dashed line\) for six metrics, with values normalized relative to the baseline scaled to 100\. Light and dark shaded bars represent the data\-poor and data\-rich regimes, respectively\.→\\scriptstyle\\rightarrow

\(a\)Boiling point \(bp\)
→\\scriptstyle\\rightarrow

\(b\)Melting point \(mp\)
→\\scriptstyle\\rightarrow

\(c\)Vapor pressure \(Pvap\)
→\\scriptstyle\\rightarrow

\(d\)Enthalpy \(H\)
→\\scriptstyle\\rightarrow

\(e\)Gibbs free energy \(G\)

Figure 11:Impact of Alignment\-Guided Learning \(AGL\) on latent geometry in the data\-poor regime\. Visualizations of normalized latent vector magnitudes \(𝐳~\\tilde\{\\mathbf\{z\}\}, blue line\) and posterior standard deviations \(σ~w\\tilde\{\\sigma\}\_\{w\}, orange line\) across five physicochemical datasets \(a–e\)\. Each panel compares the unconstrained baseline model \(left\) with the AGL model \(right\)\. The arrows indicate the structural shift induced by the alignment objective\. In this data\-scarce regime, while AGL induces a tendency toward the alignment, the resulting structural reorganization is not clearly visible \(or remains visually subtle\) due to severe information scarcity\. Shaded areas represent the standard deviation acrossn=30n=30independent runs\.→\\scriptstyle\\rightarrow

\(a\)Boiling point \(bp\)
→\\scriptstyle\\rightarrow

\(b\)Melting point \(mp\)
→\\scriptstyle\\rightarrow

\(c\)Vapor pressure \(Pvap\)
→\\scriptstyle\\rightarrow

\(d\)Enthalpy \(H\)
→\\scriptstyle\\rightarrow

\(e\)Gibbs free energy \(G\)

Figure 12:Impact of Latent\-Posterior Alignment via AGL in the data\-rich regime\. Comparison of latent geometry between Baseline \(left\) and AGL \(right\) models trained in a data\-abundant setting across five datasets \(a–e\)\. Consistent with the data\-poor regime, AGL enforces alignment; however, the extent of this reorganization is more pronounced\. With sufficient data, AGL actively reconfigures the latent space, creating a sharp boundary where latent activity is strictly confined to stable, low\-variance axes \(forming a distinct “X” shape\)\. This confirms that the availability of information facilitates more precise geometric calibration\. Shaded areas denote standard deviation \(n=30n=30\)\.![Refer to caption](https://arxiv.org/html/2608.20758v1/bp_base_poor_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/bp_base_poor_uncertainty.png)

\(a\)bp: Baseline→\\rightarrow![Refer to caption](https://arxiv.org/html/2608.20758v1/bp_align_poor_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/bp_align_poor_uncertainty.png)

\(b\)bp: AGL
![Refer to caption](https://arxiv.org/html/2608.20758v1/mp_base_poor_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/mp_base_poor_uncertainty.png)

\(c\)mp: Baseline→\\rightarrow![Refer to caption](https://arxiv.org/html/2608.20758v1/mp_align_poor_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/mp_align_poor_uncertainty.png)

\(d\)mp: AGL
![Refer to caption](https://arxiv.org/html/2608.20758v1/logPvap_base_poor_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/logPvap_base_poor_uncertainty.png)

\(e\)Pvap: Baseline→\\rightarrow![Refer to caption](https://arxiv.org/html/2608.20758v1/logPvap_align_poor_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/logPvap_align_poor_uncertainty.png)

\(f\)Pvap: AGL
![Refer to caption](https://arxiv.org/html/2608.20758v1/H_bonds_base_poor_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/H_bonds_base_poor_uncertainty.png)

\(g\)H: Baseline→\\rightarrow![Refer to caption](https://arxiv.org/html/2608.20758v1/H_bonds_align_poor_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/H_bonds_align_poor_uncertainty.png)

\(h\)H: AGL
![Refer to caption](https://arxiv.org/html/2608.20758v1/G_bonds_base_poor_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/G_bonds_base_poor_uncertainty.png)

\(i\)G: Baseline→\\rightarrow![Refer to caption](https://arxiv.org/html/2608.20758v1/G_bonds_align_poor_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/G_bonds_align_poor_uncertainty.png)

\(j\)G: AGL

Figure 13:Structural calibration in data\-poor regime\. t\-SNE dimensional reduction of latent representations is used to visualize latent geometry and compare baseline models against AGL models for five different property datasets\. Each sub\-figure displays two maps: data density \(left, estimated via KDE\) and output uncertainty \(right,σy\\sigma\_\{y\}\)\. Colors range from low \(blue\) to high \(red\)\. Left columns \(a, c, e, g, i\) and right columns \(b, d, f, h, j\) represent baseline and AGL models, respectively\.![Refer to caption](https://arxiv.org/html/2608.20758v1/bp_base_rich_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/bp_base_rich_uncertainty.png)

\(a\)bp: Baseline→\\rightarrow![Refer to caption](https://arxiv.org/html/2608.20758v1/bp_align_rich_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/bp_align_rich_uncertainty.png)

\(b\)bp: AGL
![Refer to caption](https://arxiv.org/html/2608.20758v1/mp_base_rich_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/mp_base_rich_uncertainty.png)

\(c\)mp: Baseline→\\rightarrow![Refer to caption](https://arxiv.org/html/2608.20758v1/mp_align_rich_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/mp_align_rich_uncertainty.png)

\(d\)mp: AGL
![Refer to caption](https://arxiv.org/html/2608.20758v1/logPvap_base_rich_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/logPvap_base_rich_uncertainty.png)

\(e\)Pvap: Baseline→\\rightarrow![Refer to caption](https://arxiv.org/html/2608.20758v1/logPvap_align_rich_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/logPvap_align_rich_uncertainty.png)

\(f\)Pvap: AGL
![Refer to caption](https://arxiv.org/html/2608.20758v1/H_bonds_base_rich_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/H_bonds_base_rich_uncertainty.png)

\(g\)H: Baseline→\\rightarrow![Refer to caption](https://arxiv.org/html/2608.20758v1/H_bonds_align_rich_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/H_bonds_align_rich_uncertainty.png)

\(h\)H: AGL
![Refer to caption](https://arxiv.org/html/2608.20758v1/G_bonds_base_rich_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/G_bonds_base_rich_uncertainty.png)

\(i\)G: Baseline→\\rightarrow![Refer to caption](https://arxiv.org/html/2608.20758v1/G_bonds_align_rich_density.png)

![Refer to caption](https://arxiv.org/html/2608.20758v1/G_bonds_align_rich_uncertainty.png)

\(j\)G: AGL

Figure 14:Structural calibration in data\-rich regime\. t\-SNE dimensional reduction of latent representations is used to visualize latent geometry and compare baseline models against AGL models for five different property datasets\. Each sub\-figure displays two maps: data density \(left, estimated via KDE\) and output uncertainty \(right,σy\\sigma\_\{y\}\)\. Colors range from low \(blue\) to high \(red\)\. Left columns \(a, c, e, g, i\) and right columns \(b, d, f, h, j\) represent baseline and AGL models, respectively\.

Similar Articles

Graph Alignment Topology as an Inductive Bias for Grounding Detection

arXiv cs.CL

This paper introduces Graph Alignment Topology as an inductive bias for grounding detection, using a graph neural network to model alignment structure between reference information and LLM outputs. The method achieves state-of-the-art results on multiple hallucination and question-answering datasets, outperforming GPT-4o.

Hidden Latent-State Shifts in LLMs: Why Current Alignment Is Blind to Real Internal Dangers — Especially With Agents

Reddit r/artificial

This paper demonstrates that LLMs can enter measurably different internal latent states under coherent context while maintaining aligned outputs, revealing a blind spot in current alignment methods that only monitor surface tokens. The Gemma-3-12B-IT experiment shows strong residual stream geometry shifts that existing safety frameworks cannot detect, with implications for agentic AI deployment.