Split Conformal Prediction with Label-Shift-Adjusted Bayesian Scores

arXiv cs.LG Papers

Summary

This paper proposes the Label-Shift-Adjusted Bayesian Score for conformal prediction under label shift, demonstrating shorter prediction intervals with maintained coverage in molecular property prediction.

arXiv:2609.12386v1 Announce Type: new Abstract: Conformal prediction provides distribution-free uncertainty quantification under exchangeability. However, this assumption is violated by label shift, where the marginal distribution of labels changes while the conditional distribution of inputs given labels remains stable. Under such shifts, standard conformal procedures no longer maintain their intended coverage behavior. Existing approaches address this via importance weighting. They pair the reweighting with residual-based nonconformity scores that ignore predictive uncertainty. The resulting intervals have uniform width. Bayesian conformal methods produce adaptive intervals by leveraging predictive distributions. They evaluate conformity under the source predictive, which is misaligned with the target domain under label shift. We propose the \emph{Label-Shift-Adjusted Bayesian Score} (LSA score), a nonconformity score derived from a posterior predictive tilting identity. This identity shows that the target predictive is an importance-weighted transformation of the source predictive. We use it to derive a direct correction to the Bayesian score. We evaluate the method on molecular property prediction under controlled label shift. The LSA score consistently yields shorter intervals than residual-based and source-based Bayesian scores. Coverage in the target domain remains comparable. Under stronger shift, all methods incur some coverage loss due to pseudo-label-based density-ratio estimation. The LSA score is defined for any source predictive with a tractable log-density. We instantiate it with Bayesian Ridge Regression, where the correction admits a closed form.
Original Article
View Cached Full Text

Cached at: 09/14/26, 08:43 AM

# 1Introduction
Source: [https://arxiv.org/html/2609.12386](https://arxiv.org/html/2609.12386)
marginparsep has been altered\. topmargin has been altered\. marginparpush has been altered\.

The page layout violates the ICML style\.

Please do not change the page layout, or include packages like geometry, savetrees, or fullpage, which change it for you\.

We’re not able to reliably undo arbitrary changes to the style\. Please remove the offending package\(s\), or layout\-changing commands and try again\.

Split Conformal Prediction with Label\-Shift\-Adjusted Bayesian Scores

Hyeonsu Lee\*1Juyeon Kim\*1Erkhembayar Jadamba\*1Seungjin Choi2Hyunjin Shin1

††footnotetext:\*Equal contribution1MOGAM Institute for Biomedical Research, Korea2CROID Research and aSSIST University, Korea\. Correspondence to: Hyunjin Shin <hyunjin\.shin@mogam\.re\.kr\>\.
2nd Workshop on Epistemic Intelligence in Machine Learning \(EIML@ICML 2026\), Seoul, South Korea\. Copyright 2026 by the author\(s\)\.###### Abstract

Conformal prediction provides distribution\-free uncertainty quantification under exchangeability\. However, this assumption is violated by label shift, where the marginal distribution of labels changes while the conditional distribution of inputs given labels remains stable\. Under such shifts, standard conformal procedures no longer maintain their intended coverage behavior\. Existing approaches address this via importance weighting\. They pair the reweighting with residual\-based nonconformity scores that ignore predictive uncertainty\. The resulting intervals have uniform width\. Bayesian conformal methods produce adaptive intervals by leveraging predictive distributions\. They evaluate conformity under the source predictive, which is misaligned with the target domain under label shift\. We propose the*Label\-Shift\-Adjusted Bayesian Score*\(LSA score\), a nonconformity score derived from a posterior predictive tilting identity\. This identity shows that the target predictive is an importance\-weighted transformation of the source predictive\. We use it to derive a direct correction to the Bayesian score\. We evaluate the method on molecular property prediction under controlled label shift\. The LSA score consistently yields shorter intervals than residual\-based and source\-based Bayesian scores\. Coverage in the target domain remains comparable\. Under stronger shift, all methods incur some coverage loss due to pseudo\-label\-based density\-ratio estimation\. The LSA score is defined for any source predictive with a tractable log\-density\. We instantiate it with Bayesian Ridge Regression, where the correction admits a closed form\.

## 1Introduction

Conformal prediction \(CP\) constructs prediction intervals with finite\-sample coverage guarantees under exchangeability\([Vovk et al\., 2005](https://arxiv.org/html/2609.12386#bib.bib28);[Shafer and Vovk, 2008](https://arxiv.org/html/2609.12386#bib.bib20);[Angelopoulos and Bates, 2023](https://arxiv.org/html/2609.12386#bib.bib1)\)\. This property makes it appealing for applications such as molecular property prediction\([Laghuvarapu et al\., 2023](https://arxiv.org/html/2609.12386#bib.bib13)\)\. In practice, however, the exchangeability assumption is often violated by*label shift*\([Saerens et al\., 2002](https://arxiv.org/html/2609.12386#bib.bib19)\)\. Under label shift, the marginal label distribution changes while the conditional distribution of inputs given labels remains the same\. This arises naturally in scientific settings when attention shifts toward rare or extreme property values\. It causes standard conformal methods to lose their coverage guarantees\.

A good prediction interval under label shift must satisfy two properties simultaneously\. First, the nonconformity score must be*aligned*with the target distribution so that the weighted quantile yields valid coverage\. Second, the score must be*adaptive*to predictive uncertainty so that the interval width reflects local difficulty\. Prior work addresses these two properties separately\.

Importance\-weighting methods target the first property\. Tibshirani et al\.\([Tibshirani et al\., 2019](https://arxiv.org/html/2609.12386#bib.bib26)\)introduced weighted quantiles to restore coverage under covariate shift\. Barber et al\.\([Barber et al\., 2023](https://arxiv.org/html/2609.12386#bib.bib2)\)generalized this idea to a broader framework that handles arbitrary violations of exchangeability\. Podkopaev and Ramdas\([Podkopaev and Ramdas, 2021](https://arxiv.org/html/2609.12386#bib.bib16)\)and Si et al\.\([Si et al\., 2024](https://arxiv.org/html/2609.12386#bib.bib22)\)extended the weighted approach to label shift specifically\. Lee et al\.\([Lee et al\., 2025](https://arxiv.org/html/2609.12386#bib.bib14)\)applied it to molecular property prediction\. Gibbs and Candès\([Gibbs and Candes, 2021](https://arxiv.org/html/2609.12386#bib.bib8)\)proposed online threshold adjustment to track distribution drift over time\. These methods reweight calibration samples to match the target marginal\. They restore coverage\. However, the nonconformity scores paired with the reweighting are residual\-based\. They depend only on point predictions\. The resulting intervals have uniform width regardless of local difficulty\. The second property is left unaddressed\.

Adaptive\-score methods target the second property\. Romano et al\.\([Romano et al\., 2019](https://arxiv.org/html/2609.12386#bib.bib17)\)proposed conformalized quantile regression \(CQR\), which produces intervals whose width varies with input difficulty under exchangeability\. Fong and Holmes\([Fong and Holmes, 2021](https://arxiv.org/html/2609.12386#bib.bib7)\)used the Bayesian posterior predictive as a nonconformity score\. The resulting intervals reflect model uncertainty\. Bhagwat et al\.\([Bhagwat et al\., 2025](https://arxiv.org/html/2609.12386#bib.bib3)\)extended this idea through Bayesian model averaging\. Recent theoretical work by Datta et al\.\([Datta et al\., 2025](https://arxiv.org/html/2609.12386#bib.bib5)\)and Deliu and Liseo\([Deliu and Liseo, 2025](https://arxiv.org/html/2609.12386#bib.bib6)\)further clarifies the relationship between conformal prediction and Bayesian inference\. These methods produce intervals of varying width\. However, they evaluate conformity under the source predictive distribution\. Under label shift, the source predictive is misaligned with the target domain\. This misalignment distorts the weighted score distribution\. The calibration quantile inflates and the intervals become unnecessarily wide, as we verify empirically in[section3\.3](https://arxiv.org/html/2609.12386#S3.SS3)\. The first property is violated\. No existing nonconformity score satisfies both properties at once\.

We address this gap by introducing theLabel\-Shift\-Adjusted Bayesian \(LSA\) score\. Our starting point is a posterior predictive tilting identity \([Proposition2\.1](https://arxiv.org/html/2609.12386#S2.Thmproposition1)\)\. Under label shift, the target predictive is an importance\-weighted transformation of the source predictive\. The negative log\-density of this identity yields a score decomposition that directly motivates a correction to the source Bayesian score\. We instantiate the score with Bayesian Ridge Regression \(BRR\)\. The correction admits a closed\-form prediction interval with no additional computational cost\.

We evaluate the method on molecular property prediction under controlled label shift\. The LSA score yields shorter prediction intervals than residual and source score baselines\. Target\-domain coverage remains comparable\. The method relies on pseudo\-label\-based density\-ratio estimation, so we do not claim exact finite\-sample validity under shift\. We instead analyze the practical weighted quantile as an approximation to an oracle procedure in[section3\.3](https://arxiv.org/html/2609.12386#S3.SS3)\.

## 2Methods

### 2\.1Problem Setup

Let𝒉∈ℝd\\bm\{h\}\\in\\mathbb\{R\}^\{d\}denote a feature representation of an input, and lety∈ℝy\\in\\mathbb\{R\}be a scalar target variable\. In our experiments,𝒉\\bm\{h\}is a molecular representation extracted by a pretrained encoder whose weights are frozen, andyyis a molecular property such as aqueous solubility \(log⁡S\\log Sinlog⁡mol/L\\log\\,\\mathrm\{mol/L\}\)\. Data are partitioned into four disjoint sets\. The*training set*𝒟train=\{\(𝒉i,yi\)\}i=1ntr\\mathcal\{D\}\_\{\\mathrm\{train\}\}=\\\{\(\\bm\{h\}\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{n\_\{\\mathrm\{tr\}\}\},*weight estimation set*𝒟weight=\{\(𝒉i,yi\)\}i=1nw\\mathcal\{D\}\_\{\\mathrm\{weight\}\}=\\\{\(\\bm\{h\}\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{n\_\{\\mathrm\{w\}\}\}, and*calibration set*𝒟cal=\{\(𝒉i,yi\)\}i=1n\\mathcal\{D\}\_\{\\mathrm\{cal\}\}=\\\{\(\\bm\{h\}\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{n\}are drawn from the source distributionpsp\_\{\\mathrm\{s\}\}\. The*target set*𝒟test=\{𝒉j\}j=1m\\mathcal\{D\}\_\{\\mathrm\{test\}\}=\\\{\\bm\{h\}\_\{j\}\\\}\_\{j=1\}^\{m\}, observed without labels at test time, is drawn from a shifted distributionptp\_\{\\mathrm\{t\}\}\.

We model label shift via a one\-parameter exponential tilt:

pt​\(y\)\\displaystyle p\_\{\\mathrm\{t\}\}\(y\)∝ps​\(y\)​exp⁡\(β​y\),\\displaystyle\\;\\propto\\;p\_\{\\mathrm\{s\}\}\(y\)\\,\\exp\(\\beta y\),r⁡\(y\)\\displaystyle r\(y\):=pt​\(y\)ps​\(y\)=exp⁡\(β​y\)Zr,\\displaystyle:=\\frac\{p\_\{\\mathrm\{t\}\}\(y\)\}\{p\_\{\\mathrm\{s\}\}\(y\)\}=\\frac\{\\exp\(\\beta y\)\}\{Z\_\{r\}\},Zr\\displaystyle Z\_\{r\}:=∫ps​\(y\)​exp⁡\(β​y\)​𝑑y,\\displaystyle:=\\int p\_\{\\mathrm\{s\}\}\(y\)\\,\\exp\(\\beta y\)\\,\\mathrm\{d\}y,\(1\)whereβ∈ℝ\\beta\\in\\mathbb\{R\}controls the direction and magnitude of the shift\.ZrZ\_\{r\}is the normalizing constant\. Label shift in general only requirespt​\(y\)≠ps​\(y\)p\_\{\\mathrm\{t\}\}\(y\)\\neq p\_\{\\mathrm\{s\}\}\(y\)\. We adopt the exponential tilt as a tractable parametric model because it yields a log\-linear density ratio, which admits a closed\-form correction under Gaussian predictives \([section2\.3](https://arxiv.org/html/2609.12386#S2.SS3)\)\. Since the feature extractor is deterministic, the label shift assumption carries over to the representation level\. That is,ps​\(𝒉∣y\)=pt​\(𝒉∣y\)p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid y\)=p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid y\)\. Our goal is to construct prediction intervalsC⁡\(𝒉\)=\[L⁡\(𝒉\),U⁡\(𝒉\)\]C\(\\bm\{h\}\)=\[L\(\\bm\{h\}\),\\,U\(\\bm\{h\}\)\]whose target\-domain marginal coverageℙ\(𝒉,y\)∼pt​\(y∈C​\(𝒉\)\)\\mathbb\{P\}\_\{\(\\bm\{h\},y\)\\sim p\_\{\\mathrm\{t\}\}\}\(y\\in C\(\\bm\{h\}\)\)is close to1−αcp1\-\\alpha\_\{\\mathrm\{cp\}\}\. Because our implementation relies on estimated density ratios obtained from pseudo\-labels, we do not claim exact finite\-sample validity under shift\.

### 2\.2Posterior Predictive Tilting

Under label shift, the source predictiveps​\(y∣𝒉,𝒟train\)p\_\{s\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)is no longer aligned with the target domain\. A natural question is whether the target posterior predictive can be recovered from the source one without retraining\. We show that it can, via a simple importance\-weighting identity\. The following assumptions formalize the required conditions\.

###### Assumption 2\.1\(Conditional label shift after representation\)\.

For all𝐡∈ℋ\\bm\{h\}\\in\\mathcal\{H\}andy∈𝒴y\\in\\mathcal\{Y\},

ps​\(𝒉∣y\)=pt​\(𝒉∣y\)\.p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid y\)=p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid y\)\.We further assume the corresponding predictive\-level invariance after source\-model fitting:

ps​\(𝒉∣y,𝒟train\)=pt​\(𝒉∣y,𝒟train\)\.p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid y,\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid y,\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\.

###### Assumption 2\.2\(Absolute continuity\)\.

ps​\(y\)\>0p\_\{\\mathrm\{s\}\}\(y\)\>0whereverpt​\(y\)\>0p\_\{\\mathrm\{t\}\}\(y\)\>0, so thatr⁡\(y\)=pt​\(y\)/ps​\(y\)r\(y\)=p\_\{\\mathrm\{t\}\}\(y\)/p\_\{\\mathrm\{s\}\}\(y\)is well\-defined\.

###### Assumption 2\.3\(Source representativeness\)\.

The source predictive model is well\-specified and the posterior concentrates, so that

ps​\(y∣𝒟train\)=ps​\(y\)\.p\_\{\\mathrm\{s\}\}\(y\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=p\_\{\\mathrm\{s\}\}\(y\)\.This is a mild condition satisfied whenever the training sample is reasonably large\.

###### Proposition 2\.1\(Posterior predictive tilting\)\.

Under[Assumption2\.1](https://arxiv.org/html/2609.12386#S2.Thmassumption1),[Assumption2\.2](https://arxiv.org/html/2609.12386#S2.Thmassumption2), and[Assumption2\.3](https://arxiv.org/html/2609.12386#S2.Thmassumption3), and conditioning on the fitted source data𝒟train\\mathcal\{D\}\_\{\\mathrm\{train\}\}, the target posterior predictive admits the representation

pt​\(y∣𝒉,𝒟train\)=ps​\(y∣𝒉,𝒟train\)​r​\(y\)Z⁡\(𝒉,𝒟train\),p\_\{t\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=\\frac\{p\_\{s\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\,r\(y\)\}\{Z\(\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\},\(2\)where

Z⁡\(𝒉,𝒟train\)\\displaystyle Z\(\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\):=∫ps​\(y′∣𝒉,𝒟train\)​r​\(y′\)​d​y′\\displaystyle:=\\int p\_\{s\}\(y^\{\\prime\}\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\,r\(y^\{\\prime\}\)\\,dy^\{\\prime\}\(3\)=𝔼Y∼ps\(⋅∣𝒉,𝒟train\)\[r\(Y\)\]\.\\displaystyle=\\mathbb\{E\}\_\{Y\\sim p\_\{s\}\(\\cdot\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\\\!\[r\(Y\)\]\.

The derivation is given in[appendixB](https://arxiv.org/html/2609.12386#A2)\. It applies Bayes’ rule under the𝒟train\\mathcal\{D\}\_\{\\mathrm\{train\}\}\-conditioned predictive distributions\. It uses[Assumption2\.1](https://arxiv.org/html/2609.12386#S2.Thmassumption1)to replace the target conditional mechanism by its source counterpart\. It then identifies the marginal ratio through[Assumption2\.3](https://arxiv.org/html/2609.12386#S2.Thmassumption3)\. Since𝒟train\\mathcal\{D\}\_\{\\mathrm\{train\}\}is drawn entirely from the source domain, it carries no information about the target label marginal, sopt​\(y∣𝒟train\)=pt​\(y\)p\_\{\\mathrm\{t\}\}\(y\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=p\_\{\\mathrm\{t\}\}\(y\)\. Combined with[Assumption2\.3](https://arxiv.org/html/2609.12386#S2.Thmassumption3), the conditional marginal ratio reduces to the population density ratior⁡\(y\)=pt​\(y\)/ps​\(y\)r\(y\)=p\_\{\\mathrm\{t\}\}\(y\)/p\_\{\\mathrm\{s\}\}\(y\)\.

This identity has a direct implication for nonconformity scoring\. Usings⁡\(𝒉,y\)=−log⁡p⁡\(y∣𝒉,𝒟train\)s\(\\bm\{h\},y\)=\-\\log p\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)under the source predictive yields a score misaligned with the target domain\.[Proposition2\.1](https://arxiv.org/html/2609.12386#S2.Thmproposition1)shows that the correctly aligned score usesptp\_\{t\}instead ofpsp\_\{s\}\. Taking the negative logarithm of[eq\.2](https://arxiv.org/html/2609.12386#S2.E2)yields

−log⁡pt​\(y∣𝒉,𝒟train\)=−log⁡ps​\(y∣𝒉,𝒟train\)⏟source score\+\(−log⁡r⁡\(y\)\)⏟shift correction\+log⁡Z⁡\(𝒉,𝒟train\)⏟normalization\.\-\\log p\_\{\\mathrm\{t\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=\\underbrace\{\-\\log p\_\{\\mathrm\{s\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\_\{\\text\{source score\}\}\\\\ \+\\underbrace\{\(\-\\log r\(y\)\)\}\_\{\\text\{shift correction\}\}\+\\underbrace\{\\log Z\(\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\_\{\\text\{normalization\}\}\.\(4\)Define the oracle nonconformity score asst​\(𝒉,y\):=−log⁡pt​\(y∣𝒉,𝒟train\)s\_\{t\}\(\\bm\{h\},y\):=\-\\log p\_\{\\mathrm\{t\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\. The sublevel set\{y:st​\(𝒉,y\)≤q\}\\\{y:s\_\{t\}\(\\bm\{h\},y\)\\leq q\\\}coincides exactly with the highest\-density region of the target predictive at levele−qe^\{\-q\}\. This decomposition motivates our practical score construction\.

### 2\.3Label\-Shift\-Adjusted Bayesian Score

Guided by[eq\.4](https://arxiv.org/html/2609.12386#S2.E4), we first estimate the density ratior⁡\(y\)r\(y\)\. Label shift estimation has been studied extensively\([Lipton et al\., 2018](https://arxiv.org/html/2609.12386#bib.bib15);[Saerens et al\., 2002](https://arxiv.org/html/2609.12386#bib.bib19)\)\. We use the estimated ratio to define the target\-aligned predictive and the corresponding nonconformity score\. Under the exponential tilt model \([eq\.1](https://arxiv.org/html/2609.12386#S2.E1)\), the log\-density ratio is exactly linear inyy\. We estimate it by fitting a logistic regression classifier\([Sugiyama et al\., 2012](https://arxiv.org/html/2609.12386#bib.bib25)\)that distinguishes source labels\{yi\}i∈𝒟weight\\\{y\_\{i\}\\\}\_\{i\\in\\mathcal\{D\}\_\{\\mathrm\{weight\}\}\}from target pseudo\-labels\{y~j\}j∈𝒟test\\\{\\tilde\{y\}\_\{j\}\\\}\_\{j\\in\\mathcal\{D\}\_\{\\mathrm\{test\}\}\}, where

y~j:=μ⁡\(𝒉j\)\\tilde\{y\}\_\{j\}:=\\mu\(\\bm\{h\}\_\{j\}\)\(5\)is the predictive mean of the source model at each unlabeled target input \(defined in[eq\.11](https://arxiv.org/html/2609.12386#S2.E11)below\)\. This yields

log⁡r^​\(y\)=β^0\+β^1​y,\\log\\hat\{r\}\(y\)=\\hat\{\\beta\}\_\{0\}\+\\hat\{\\beta\}\_\{1\}y,\(6\)whereβ^0\\hat\{\\beta\}\_\{0\}absorbs the normalization constant and the class\-prior offset \(see[appendixC](https://arxiv.org/html/2609.12386#A3)\)\. Withr^​\(y\)\\hat\{r\}\(y\)in hand, we replacer⁡\(y\)r\(y\)in[eq\.2](https://arxiv.org/html/2609.12386#S2.E2)withr^​\(y\)\\hat\{r\}\(y\)and define the*estimated target\-aligned predictive*

p^t​\(y∣𝒉,𝒟train\)\\displaystyle\\hat\{p\}\_\{t\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\):=ps​\(y∣𝒉,𝒟train\)​r^​\(y\)Z^​\(𝒉\),\\displaystyle:=\\frac\{p\_\{\\mathrm\{s\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\,\\hat\{r\}\(y\)\}\{\\hat\{Z\}\(\\bm\{h\}\)\},Z^​\(𝒉\)\\displaystyle\\hat\{Z\}\(\\bm\{h\}\):=∫ps​\(y′∣𝒉,𝒟train\)​r^​\(y′\)​d​y′\.\\displaystyle:=\\int p\_\{\\mathrm\{s\}\}\(y^\{\\prime\}\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\,\\hat\{r\}\(y^\{\\prime\}\)\\,\\mathrm\{d\}y^\{\\prime\}\.\(7\)
#### LSA score\.

Dropping the constant12​log⁡\(2​π\)\\frac\{1\}\{2\}\\log\(2\\pi\)throughout, the source score is the negative log\-density ofps​\(y∣𝒉,𝒟train\)p\_\{\\mathrm\{s\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\),

ss​\(𝒉,y\)=\(y−μ⁡\(𝒉\)\)22​σ2​\(𝒉\)\+12​log⁡σ2​\(𝒉\),s\_\{\\mathrm\{s\}\}\(\\bm\{h\},y\)=\\frac\{\(y\-\\mu\(\\bm\{h\}\)\)^\{2\}\}\{2\\sigma^\{2\}\(\\bm\{h\}\)\}\+\\frac\{1\}\{2\}\\log\\sigma^\{2\}\(\\bm\{h\}\),\(8\)and we define theLabel\-Shift\-Adjusted Bayesian Score\(LSA score\) as

sLSA​\(𝒉,y\)=ss​\(𝒉,y\)−log⁡r^​\(y\)\+log⁡Z^​\(𝒉\)\.s\_\{\\mathrm\{LSA\}\}\(\\bm\{h\},y\)=s\_\{\\mathrm\{s\}\}\(\\bm\{h\},y\)\-\\log\\hat\{r\}\(y\)\+\\log\\hat\{Z\}\(\\bm\{h\}\)\.\(9\)This is the negative log\-density ofp^t​\(y∣𝒉,𝒟train\)\\hat\{p\}\_\{t\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)up to the dropped constant\. It is therefore a plug\-in approximation to the oracle target score in[eq\.4](https://arxiv.org/html/2609.12386#S2.E4), withr^​\(y\)\\hat\{r\}\(y\)in place ofr⁡\(y\)r\(y\)\. The factorr^​\(y\)\\hat\{r\}\(y\)inside the score reshapes what is measured\. It replaces source\-based conformity with conformity under the estimated target\-aligned predictive\. Specifically,−log⁡r^​\(y\)\-\\log\\hat\{r\}\(y\)lowers the nonconformity score for target\-favored labels and raises it for target\-disfavored ones\.log⁡Z^​\(𝒉\)\\log\\hat\{Z\}\(\\bm\{h\}\)restores normalization through an input\-dependent offset\. Note that this score\-level correction is distinct from the weight factorr^​\(yi\)\\hat\{r\}\(y\_\{i\}\)used later to reweight calibration scores in the weighted quantile\. The score correction changes what is measured\. It evaluates conformity under the target\-aligned predictive rather than the source one\. The weight factor changes how calibration samples are aggregated\. It reweights the empirical distribution to match the target marginal\. The two address different axes of the label shift problem and should not be conflated\.

The LSA score is defined for any source predictiveps​\(y∣𝒉,𝒟train\)p\_\{s\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\. We instantiate it with Bayesian Ridge Regression \(BRR\)\. The Gaussian predictive and log\-linear density ratio together yield a closed\-form prediction interval with no additional computational cost over the source score\. BRR places a Gaussian linear model on the input representations:

y=𝒉⊤​𝒘\+ε,ε∼𝒩⁡\(0,λ−1\),𝒘∼𝒩⁡\(𝟎,α−1​𝑰d\),y=\\bm\{h\}^\{\\top\}\\bm\{w\}\+\\varepsilon,\\quad\\varepsilon\\sim\\mathcal\{N\}\(0,\\,\\lambda^\{\-1\}\),\\quad\\bm\{w\}\\sim\\mathcal\{N\}\(\\bm\{0\},\\,\\alpha^\{\-1\}\\bm\{I\}\_\{d\}\),\(10\)where hyperparameters are estimated by marginal likelihood maximization\([Tipping, 2001](https://arxiv.org/html/2609.12386#bib.bib27);[Bishop and Nasrabadi, 2006](https://arxiv.org/html/2609.12386#bib.bib4)\)\(see[appendixA](https://arxiv.org/html/2609.12386#A1)\)\. Conjugacy yields a Gaussian posterior predictive

ps​\(y∗∣𝒉∗,𝒟train\)\\displaystyle p\_\{\\mathrm\{s\}\}\(y^\{\*\}\\mid\\bm\{h\}^\{\*\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=𝒩⁡\(y∗∣μ⁡\(𝒉∗\),σ2​\(𝒉∗\)\),\\displaystyle=\\mathcal\{N\}\\\!\\bigl\(y^\{\*\}\\mid\\mu\(\\bm\{h\}^\{\*\}\),\\;\\sigma^\{2\}\(\\bm\{h\}^\{\*\}\)\\bigr\),μ⁡\(𝒉∗\)\\displaystyle\\mu\(\\bm\{h\}^\{\*\}\)=𝒉∗⊤​𝝁𝒘,\\displaystyle=\{\\bm\{h\}^\{\*\}\}^\{\\top\}\\bm\{\\mu\}\_\{\\bm\{w\}\},σ2​\(𝒉∗\)\\displaystyle\\sigma^\{2\}\(\\bm\{h\}^\{\*\}\)=𝒉∗⊤​𝚺𝒘​𝒉∗\+λ−1,\\displaystyle=\{\\bm\{h\}^\{\*\}\}^\{\\top\}\\bm\{\\Sigma\}\_\{\\bm\{w\}\}\\bm\{h\}^\{\*\}\+\\lambda^\{\-1\},\(11\)Becauseps​\(y∣𝒉,𝒟train\)p\_\{\\mathrm\{s\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)is Gaussian andr^​\(y\)\\hat\{r\}\(y\)is log\-linear inyy, bothZ^​\(𝒉\)\\hat\{Z\}\(\\bm\{h\}\)and the tilted predictive admit analytic forms \(derivations in[appendicesD](https://arxiv.org/html/2609.12386#A4)and[E](https://arxiv.org/html/2609.12386#A5)\):

Z^​\(𝒉\)\\displaystyle\\hat\{Z\}\(\\bm\{h\}\)=exp⁡\(β^0\+β^1​μ​\(𝒉\)\+12​β^12​σ2​\(𝒉\)\),\\displaystyle=\\exp\\\!\\Bigl\(\\hat\{\\beta\}\_\{0\}\+\\hat\{\\beta\}\_\{1\}\\,\\mu\(\\bm\{h\}\)\+\\tfrac\{1\}\{2\}\\hat\{\\beta\}\_\{1\}^\{2\}\\,\\sigma^\{2\}\(\\bm\{h\}\)\\Bigr\),\(12\)p^t​\(y∣𝒉,𝒟train\)\\displaystyle\\hat\{p\}\_\{t\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=𝒩⁡\(y∣μ∗​\(𝒉\),σ2​\(𝒉\)\),\\displaystyle=\\mathcal\{N\}\\\!\\bigl\(y\\mid\\mu^\{\*\}\(\\bm\{h\}\),\\sigma^\{2\}\(\\bm\{h\}\)\\bigr\),μ∗​\(𝒉\)\\displaystyle\\mu^\{\*\}\(\\bm\{h\}\):=μ⁡\(𝒉\)\+β^1​σ2​\(𝒉\),\\displaystyle:=\\mu\(\\bm\{h\}\)\+\\hat\{\\beta\}\_\{1\}\\sigma^\{2\}\(\\bm\{h\}\),\(13\)so exponential tilting shifts only the predictive mean\. The variance remains unchanged\. The LSA score therefore reduces to

sLSA​\(𝒉,y\)=\(y−μ∗​\(𝒉\)\)22​σ2​\(𝒉\)\+12​log⁡σ2​\(𝒉\),s\_\{\\mathrm\{LSA\}\}\(\\bm\{h\},y\)=\\frac\{\(y\-\\mu^\{\*\}\(\\bm\{h\}\)\)^\{2\}\}\{2\\sigma^\{2\}\(\\bm\{h\}\)\}\+\\frac\{1\}\{2\}\\log\\sigma^\{2\}\(\\bm\{h\}\),\(14\)which coincides with the exact target score−log⁡pt​\(y∣𝒉,𝒟train\)\-\\log p\_\{t\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)of[Proposition2\.1](https://arxiv.org/html/2609.12386#S2.Thmproposition1)whenr^​\(y\)=r​\(y\)\\hat\{r\}\(y\)=r\(y\)\.

### 2\.4Weighted Conformal Prediction with the LSA Score

Given calibration scores\{sLSA​\(𝒉i,yi\)\}i=1n\\\{s\_\{\\mathrm\{LSA\}\}\(\\bm\{h\}\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{n\}and estimated density ratios\{r^​\(yi\)\}i=1n\\\{\\hat\{r\}\(y\_\{i\}\)\\\}\_\{i=1\}^\{n\}, we compute the weighted empirical quantile following Tibshirani et al\.\([Tibshirani et al\., 2019](https://arxiv.org/html/2609.12386#bib.bib26)\):

q^=inf\{q:∑i=1nr^\(yi\)1\[sLSA\(𝒉i,yi\)≤q\]∑i=1nr^​\(yi\)\+r^​\(y~n\+1\)≥1−αcp\},\\hat\{q\}=\\inf\\Bigl\\\{q:\\frac\{\\sum\_\{i=1\}^\{n\}\\hat\{r\}\(y\_\{i\}\)\\,\\mathbf\{1\}\[s\_\{\\mathrm\{LSA\}\}\(\\bm\{h\}\_\{i\},y\_\{i\}\)\\leq q\]\}\{\\sum\_\{i=1\}^\{n\}\\hat\{r\}\(y\_\{i\}\)\+\\hat\{r\}\(\\tilde\{y\}\_\{n\+1\}\)\}\\\\ \\geq 1\-\\alpha\_\{\\mathrm\{cp\}\}\\Bigr\\\},\(15\)wherey~n\+1=μ⁡\(𝒉n\+1\)\\tilde\{y\}\_\{n\+1\}=\\mu\(\\bm\{h\}\_\{n\+1\}\)is the pseudo\-label of the test input\. Since the true test label is unobserved, we use the predictive mean as a surrogate\. This is consistent with the pseudo\-label construction in[eq\.5](https://arxiv.org/html/2609.12386#S2.E5)\.

The LSA score takes the Gaussian log\-density form in[eq\.14](https://arxiv.org/html/2609.12386#S2.E14)\. The prediction interval for a test input𝒉∗\\bm\{h\}^\{\\ast\}is

C⁡\(𝒉∗\)=\[μ∗​\(𝒉∗\)±σ⁡\(𝒉∗\)​2​max⁡\(q^−12​log⁡σ2​\(𝒉∗\),0\)\],C\(\\bm\{h\}^\{\\ast\}\)=\\Bigl\[\\mu^\{\\ast\}\(\\bm\{h\}^\{\\ast\}\)\\pm\\sigma\(\\bm\{h\}^\{\\ast\}\)\\sqrt\{2\\max\\bigl\(\\hat\{q\}\-\\tfrac\{1\}\{2\}\\log\\sigma^\{2\}\(\\bm\{h\}^\{\\ast\}\),0\\bigr\)\}\\,\\Bigr\],\(16\)centered at the label\-shift\-adjusted meanμ∗​\(𝒉∗\)\\mu^\{\\ast\}\(\\bm\{h\}^\{\\ast\}\)with half\-width proportional toσ⁡\(𝒉∗\)\\sigma\(\\bm\{h\}^\{\\ast\}\)\. The intervals are therefore sample\-adaptive \(derivation in[appendixF](https://arxiv.org/html/2609.12386#A6)\)\. The primary efficiency gain stems from reshaping the score distribution\. The LSA correction shifts the predictive center toμ∗​\(𝒉\)=μ⁡\(𝒉\)\+β^1​σ2​\(𝒉\)\\mu^\{\\ast\}\(\\bm\{h\}\)=\\mu\(\\bm\{h\}\)\+\\hat\{\\beta\}\_\{1\}\\sigma^\{2\}\(\\bm\{h\}\)\. Target\-favored labels therefore receive lower nonconformity scores\. This shifts the weighted CDF leftward and lowersq^\\hat\{q\}\. We evaluate this effect empirically in[section3](https://arxiv.org/html/2609.12386#S3)\.

## 3Experiments

### 3\.1Experimental Setup

We evaluate on AqSolDB from the Therapeutics Data Commons\([Sorkun et al\., 2019](https://arxiv.org/html/2609.12386#bib.bib23);[Huang et al\., 2021](https://arxiv.org/html/2609.12386#bib.bib9)\), which contains9,9829\{,\}982compounds with experimentally measured aqueous solubility \(log⁡S\\log Sinlog⁡mol/L\\log\\,\\mathrm\{mol/L\}\)\. Molecular SMILES strings are encoded with a frozen pretrained BART\-based chemical language model\([Ross et al\., 2022](https://arxiv.org/html/2609.12386#bib.bib18)\)\(d=768d=768,171\.16171\.16M parameters\)\. Only the BRR head is fitted on the extracted representations \(see[appendixH](https://arxiv.org/html/2609.12386#A8)for pretraining details\)\. Results on a second benchmark, Lipophilicity\([Wenlock and Tomkinson, 2015](https://arxiv.org/html/2609.12386#bib.bib29);[Wu et al\., 2018](https://arxiv.org/html/2609.12386#bib.bib30)\), are reported in[sectionG\.2](https://arxiv.org/html/2609.12386#A7.SS2)\.

In each trial, the dataset is randomly partitioned into four disjoint subsets\.𝒟train\\mathcal\{D\}\_\{\\mathrm\{train\}\}\(∼\\sim40%\) is used for model training\. Three equal\-sized subsets \(each 20%\) serve as𝒟weight\\mathcal\{D\}\_\{\\mathrm\{weight\}\}for density\-ratio estimation,𝒟cal\\mathcal\{D\}\_\{\\mathrm\{cal\}\}for conformal calibration, and𝒟test\\mathcal\{D\}\_\{\\mathrm\{test\}\}for evaluation\. For shifted evaluation,𝒟test\\mathcal\{D\}\_\{\\mathrm\{test\}\}is resampled with replacement using probabilities proportional toexp⁡\(β​y\)\\exp\(\\beta y\), following the exponential\-tilt scheme in[eq\.1](https://arxiv.org/html/2609.12386#S2.E1)withβ∈\{0\.0,−0\.1,−0\.2,−0\.3,−0\.4,−0\.5\}\\beta\\in\\\{0\.0,\\,\-0\.1,\\,\-0\.2,\\,\-0\.3,\\,\-0\.4,\\,\-0\.5\\\}\. Hereβ=0\\beta=0corresponds to no shift andβ=−0\.5\\beta=\-0\.5to a moderate shift toward low\-solubility compounds\. Each configuration is repeated over 1,000 random seeds\. We report marginal coverageP^​\(y∈C​\(𝒉\)\)\\hat\{P\}\(y\\in C\(\\bm\{h\}\)\)at target level0\.90\.9and average interval length𝔼⁡\[U⁡\(𝒉\)−L⁡\(𝒉\)\]\\mathbb\{E\}\[U\(\\bm\{h\}\)\-L\(\\bm\{h\}\)\]\.

All methods share the same BRR predictive model, logistic regression weight estimator, and weighted quantile computation \([eq\.15](https://arxiv.org/html/2609.12386#S2.E15)\)\. They differ only in the nonconformity score\.Residualusessres​\(𝒉,y\)=\|y−μ⁡\(𝒉\)\|s\_\{\\mathrm\{res\}\}\(\\bm\{h\},y\)=\|y\-\\mu\(\\bm\{h\}\)\|and produces uniform intervals\.Source scoreusesss​\(𝒉,y\)s\_\{\\mathrm\{s\}\}\(\\bm\{h\},y\)in[eq\.8](https://arxiv.org/html/2609.12386#S2.E8)and produces adaptive intervals without label\-shift correction\.LSAusessLSA​\(𝒉,y\)s\_\{\\mathrm\{LSA\}\}\(\\bm\{h\},y\)in[eq\.9](https://arxiv.org/html/2609.12386#S2.E9)and produces adaptive intervals with label\-shift correction\. This controlled comparison isolates the effect of the conformity score\.

### 3\.2Main Results

[Tables1](https://arxiv.org/html/2609.12386#S3.T1)and[1](https://arxiv.org/html/2609.12386#S3.F1)summarize the main results\. Atβ=0\\beta=0, all three scores achieve the nominal90%90\\%coverage with nearly identical interval lengths\. This confirms correct calibration under exchangeability\. As\|β\|\|\\beta\|increases, the LSA score yields shorter intervals than both baselines\. The gap widens with\|β\|\|\\beta\|\. Atβ=−0\.5\\beta=\-0\.5, the average interval length is4\.984\.98for LSA, compared to5\.745\.74for Source score and5\.825\.82for Residual, corresponding to reductions of13%13\\%and14%14\\%, respectively\. Pairedtt\-tests on interval length yieldp<0\.001p<0\.001for all\|β\|≥0\.1\|\\beta\|\\geq 0\.1\. This improvement reflects a smaller weighted score quantile under LSA\. We analyze this through the weighted CDF in[section3\.3](https://arxiv.org/html/2609.12386#S3.SS3)\.

We observe the same qualitative behavior under an alternative Gaussian\-based density\-ratio estimator \([sectionG\.1](https://arxiv.org/html/2609.12386#A7.SS1)\)\. This is consistent with the shared log\-linear form inyy\. The log\-linear form preserves the Gaussian structure of the LSA correction\. All methods lose coverage as\|β\|\|\\beta\|increases\. This is expected because the weighted quantile is computed from pseudo\-label\-based weight estimates rather than oracle importance weights\. As the shift becomes stronger, the target pseudo\-labelsy~j\\tilde\{y\}\_\{j\}used for density\-ratio estimation become less accurate\. This degrades the estimated importance weights and the resulting weighted conformal quantile\. We therefore interpret the remaining coverage gap primarily as an estimation effect rather than as evidence against the target\-aligned score itself\. Within this estimated\-weight regime, LSA remains competitive in coverage\. It achieves substantially shorter intervals\.

Table 1:Coverage and average interval length under exponential\-tilt label shift\. Target coverage is0\.90\.9\. Results are averaged over1,0001\{,\}000seeds\.Boldindicates the best value for eachβ\\beta\.![Refer to caption](https://arxiv.org/html/2609.12386v1/figure1_overall.png)Figure 1:Coverage \(left\) and interval length \(right\) under exponential\-tilt label shift on AqSolDB\. Three nonconformity scores are compared\. Residual, Source score, and LSA are shown\. Target coverage is0\.90\.9\.β∈\{0,−0\.1,…,−0\.5\}\\beta\\in\\\{0,\\,\-0\.1,\\ldots,\\,\-0\.5\\\}\. Error bands denote±\\pm1 std over1,0001\{,\}000seeds\.
### 3\.3Score Distribution Shift Analysis

![Refer to caption](https://arxiv.org/html/2609.12386v1/figure2_wcdf_comparison.png)Figure 2:Weighted CDFs of calibration scores atβ=0\.0\\beta=0\.0\(left\) andβ=−0\.5\\beta=\-0\.5\(right\) on AqSolDB\. Each curve plots the cumulative weight fraction∑r^\(yi\)𝟏\[si≤q\]/∑r^\(yj\)\\sum\\hat\{r\}\(y\_\{i\}\)\\mathbf\{1\}\[s\_\{i\}\\leq q\]\\,/\\,\\sum\\hat\{r\}\(y\_\{j\}\)as a function of the score thresholdqq\. The weighted quantileq^\\hat\{q\}is the score at which the CDF reaches1−αcp≈0\.91\-\\alpha\_\{\\mathrm\{cp\}\}\\approx 0\.9\. Atβ=0\.0\\beta=0\.0, the Source score and LSA CDFs nearly overlap\. Atβ=−0\.5\\beta=\-0\.5, the LSA CDF shifts leftward and yields a smaller weighted quantile\.[Figure2](https://arxiv.org/html/2609.12386#S3.F2)compares the weighted CDFs of the Source score and LSA calibration scores\. The weighted quantileq^\\hat\{q\}is the score at which the CDF reaches1−αcp≈0\.91\-\\alpha\_\{\\mathrm\{cp\}\}\\approx 0\.9\. A smaller quantile yields a shorter interval through[eq\.16](https://arxiv.org/html/2609.12386#S2.E16)\. Atβ=0\\beta=0, the two CDFs nearly overlap\. Their quantiles coincide\. Atβ=−0\.5\\beta=\-0\.5, the LSA CDF shifts leftward and the quantile decreases\. This follows from the mean shift in[eq\.13](https://arxiv.org/html/2609.12386#S2.E13)\. The LSA score evaluates conformity with respect to the estimated target\-aligned meanμ∗​\(𝒉\)=μ⁡\(𝒉\)\+β^1​σ2​\(𝒉\)\\mu^\{\*\}\(\\bm\{h\}\)=\\mu\(\\bm\{h\}\)\+\\hat\{\\beta\}\_\{1\}\\sigma^\{2\}\(\\bm\{h\}\)rather than the source meanμ⁡\(𝒉\)\\mu\(\\bm\{h\}\)\. Under negative label shift \(β^1<0\\hat\{\\beta\}\_\{1\}<0\), calibration samples with target\-favored labels receive lower LSA scores than Source scores\. This shifts weighted mass leftward and lowersq^\\hat\{q\}\. The Source score evaluates conformity under the source predictive and does not perform this correction\. It therefore assigns higher scores to target\-favored samples\. This produces a larger weighted quantile and wider intervals\.

This score distribution shift directly explains the interval\-length reduction observed in[tables1](https://arxiv.org/html/2609.12386#S3.T1)and[1](https://arxiv.org/html/2609.12386#S3.F1)\. Coverage remains broadly similar across methods within the estimated\-weight regime\. All methods degrade under stronger shift\. A likely reason is that stronger label shift makes the target pseudo\-labelsy~j\\tilde\{y\}\_\{j\}used for density\-ratio estimation less reliable\. This directly affects the estimated density ratios and the resulting weighted conformal quantile\. In this regime, the re\-centering in LSA improves efficiency without introducing an additional coverage penalty\.

To assess whether this effect is specific to AqSolDB, we evaluate the same protocol on a second molecular regression benchmark, Lipophilicity \(see[sectionG\.2](https://arxiv.org/html/2609.12386#A7.SS2)\)\. Atβ=0\\beta=0, the three methods are nearly indistinguishable\. As the magnitude of negative label shift increases, the LSA score again yields shorter intervals than both baselines\. Empirical coverage remains comparable and is slightly higher for LSA under larger shifts\.

## 4Conclusion

We studied conformal prediction under label shift and identified a key limitation of existing approaches\. Importance weighting can partially correct distribution shift\. Commonly used nonconformity scores, however, ignore predictive uncertainty\. This leads to inefficient prediction intervals\. We proposed the Label\-Shift\-Adjusted Bayesian Score, a nonconformity score derived from a posterior predictive tilting identity\. Label shift induces a simple transformation of the posterior predictive\. This motivates a direct correction to the standard Bayesian score\. Under Bayesian Ridge Regression, the method yields closed\-form, sample\-adaptive prediction intervals with negligible computational overhead\.

Empirically, the LSA score produces shorter prediction intervals than both residual\-based and source\-based Bayesian scores\. Coverage remains comparable across a range of label shift magnitudes\. As the shift becomes more pronounced, all methods exhibit some loss in coverage\. This behavior is consistent with the use of pseudo\-labels for density\-ratio estimation\. Pseudo\-labels introduce additional approximation error\. Within this regime, the improvement from LSA can be traced to a better alignment between the score distribution and the target predictive\.

This study uses a Gaussian predictive model\. The adjustment appears as a shift in the predictive mean\. The same idea can be applied more broadly\. Extending the approach to richer models such as deep ensembles or sampling\-based methods \(MC\-Dropout\) is a natural next step\.

## Acknowledgements

This research was supported by a grant of the Korea Machine Learning Ledger Orchestration for Drug Discovery Project \(K\-MELLODDY\), funded by the Ministry of Health & Welfare and Ministry of Science and ICT, Republic of Korea \(grant number: RS\-2024\-00459964\)\.

## References

- Angelopoulos and Bates \[2023\]Anastasios N Angelopoulos and Stephen Bates\.Conformal prediction: A gentle introduction\.*Foundations and Trends in Machine Learning*, 16\(4\):494–591, 2023\.
- Barber et al\. \[2023\]Rina Foygel Barber, Emmanuel J Candes, Aaditya Ramdas, and Ryan J Tibshirani\.Conformal prediction beyond exchangeability\.*The Annals of Statistics*, 51\(2\):816–845, 2023\.
- Bhagwat et al\. \[2025\]Pankaj Bhagwat, Linglong Kong, and Bei Jiang\.CBMA: Improving conformal prediction through Bayesian model averaging\.In*International Conference on Learning Representations*, 2025\.URL[https://openreview\.net/forum?id=BKSeNw2HIr](https://openreview.net/forum?id=BKSeNw2HIr)\.
- Bishop and Nasrabadi \[2006\]Christopher M Bishop and Nasser M Nasrabadi\.*Pattern recognition and machine learning*, volume 4\.Springer, 2006\.
- Datta et al\. \[2025\]Jyotishka Datta, Nicholas G Polson, Vadim Sokolov, and Daniel Zantedeschi\.Conformal prediction= Bayes?*arXiv preprint arXiv:2512\.23308*, 2025\.
- Deliu and Liseo \[2025\]Nina Deliu and Brunero Liseo\.The interplay between Bayesian inference and conformal prediction\.*arXiv preprint arXiv:2510\.26930*, 2025\.
- Fong and Holmes \[2021\]Edwin Fong and Chris C Holmes\.Conformal Bayesian computation\.*Advances in Neural Information Processing Systems*, 34:18268–18279, 2021\.
- Gibbs and Candes \[2021\]Isaac Gibbs and Emmanuel Candes\.Adaptive conformal inference under distribution shift\.*Advances in Neural Information Processing Systems*, 34:1660–1672, 2021\.
- Huang et al\. \[2021\]Kexin Huang, Tianfan Fu, Wenhao Gao, Yue Zhao, Yusuf Roohani, Jure Leskovec, Connor W Coley, Cao Xiao, Jimeng Sun, and Marinka Zitnik\.Therapeutics data commons: Machine learning datasets and tasks for drug discovery and development\.*Proceedings of Neural Information Processing Systems, NeurIPS Datasets and Benchmarks*, 2021\.
- Irwin et al\. \[2012\]John J Irwin, Teague Sterling, Michael M Mysinger, Erin S Bolstad, and Ryan G Coleman\.Zinc: a free tool to discover chemistry for biology\.*Journal of chemical information and modeling*, 52\(7\):1757–1768, 2012\.
- Kim et al\. \[2023\]Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al\.Pubchem 2023 update\.*Nucleic acids research*, 51\(D1\):D1373–D1380, 2023\.
- Knox et al\. \[2024\]Craig Knox, Mike Wilson, Christen M Klinger, Mark Franklin, Eponine Oler, Alex Wilson, Allison Pon, Jordan Cox, Na Eun Chin, Seth A Strawbridge, et al\.Drugbank 6\.0: the drugbank knowledgebase for 2024\.*Nucleic acids research*, 52\(D1\):D1265–D1275, 2024\.
- Laghuvarapu et al\. \[2023\]Siddhartha Laghuvarapu, Zhen Lin, and Jimeng Sun\.Codrug: Conformal drug property prediction with density estimation under covariate shift\.*Advances in Neural Information Processing Systems*, 36:37728–37747, 2023\.
- Lee et al\. \[2025\]Hyeonsu Lee, Juyeon Kim, Erkhembayar Jadamba, Seungjin Choi, and Hyunjin Shin\.Conformal prediction for molecular properties under label shift\.In*NeurIPS 2025 Workshop: Reliable ML from Unreliable Data*, 2025\.URL[https://openreview\.net/forum?id=IR3T2TCmme](https://openreview.net/forum?id=IR3T2TCmme)\.
- Lipton et al\. \[2018\]Zachary Lipton, Yu\-Xiang Wang, and Alexander Smola\.Detecting and correcting for label shift with black box predictors\.In*International conference on machine learning*, pages 3122–3130\. PMLR, 2018\.
- Podkopaev and Ramdas \[2021\]Aleksandr Podkopaev and Aaditya Ramdas\.Distribution\-free uncertainty quantification for classification under label shift\.In*Uncertainty in artificial intelligence*, pages 844–853\. PMLR, 2021\.
- Romano et al\. \[2019\]Yaniv Romano, Evan Patterson, and Emmanuel Candes\.Conformalized quantile regression\.*Advances in neural information processing systems*, 32, 2019\.
- Ross et al\. \[2022\]Jerret Ross, Brian Belgodere, Vijil Chenthamarakshan, Inkit Padhi, Youssef Mroueh, and Payel Das\.Large\-scale chemical language representations capture molecular structure and properties\.*Nature Machine Intelligence*, 4\(12\):1256–1264, 2022\.
- Saerens et al\. \[2002\]Marco Saerens, Patrice Latinne, and Christine Decaestecker\.Adjusting the outputs of a classifier to new a priori probabilities: a simple procedure\.*Neural computation*, 14\(1\):21–41, 2002\.
- Shafer and Vovk \[2008\]Glenn Shafer and Vladimir Vovk\.A tutorial on conformal prediction\.*Journal of machine learning research*, 9\(3\), 2008\.
- Shivanyuk et al\. \[2007\]Alexander N Shivanyuk, Sergey V Ryabukhin, A Tolmachev, AV Bogolyubsky, DM Mykytenko, AA Chupryna, W Heilman, and AN Kostyuk\.Enamine real database: Making chemical diversity real\.*Chemistry today*, 25\(6\):58–59, 2007\.
- Si et al\. \[2024\]Wenwen Si, Sangdon Park, Insup Lee, Edgar Dobriban, and Osbert Bastani\.PAC prediction sets under label shift\.In*International Conference on Learning Representations*, 2024\.URL[https://openreview\.net/forum?id=4vPVBh3fhz](https://openreview.net/forum?id=4vPVBh3fhz)\.
- Sorkun et al\. \[2019\]Murat Cihan Sorkun, Abhishek Khetan, and Süleyman Er\.Aqsoldb, a curated reference set of aqueous solubility and 2d descriptors for a diverse set of compounds\.*Scientific data*, 6\(1\):143, 2019\.
- Sorokina et al\. \[2021\]Maria Sorokina, Peter Merseburger, Kohulan Rajan, Mehmet Aziz Yirik, and Christoph Steinbeck\.Coconut online: collection of open natural products database\.*Journal of Cheminformatics*, 13\(1\):2, 2021\.
- Sugiyama et al\. \[2012\]Masashi Sugiyama, Taiji Suzuki, and Takafumi Kanamori\.*Density ratio estimation in machine learning*\.Cambridge University Press, 2012\.
- Tibshirani et al\. \[2019\]Ryan J Tibshirani, Rina Foygel Barber, Emmanuel Candes, and Aaditya Ramdas\.Conformal prediction under covariate shift\.*Advances in neural information processing systems*, 32, 2019\.
- Tipping \[2001\]Michael E Tipping\.Sparse Bayesian learning and the relevance vector machine\.*Journal of machine learning research*, 1\(Jun\):211–244, 2001\.
- Vovk et al\. \[2005\]Vladimir Vovk, Alexander Gammerman, and Glenn Shafer\.*Algorithmic learning in a random world*\.Springer, 2005\.
- Wenlock and Tomkinson \[2015\]Mark Wenlock and Nicholas Tomkinson\.Experimental in vitro dmpk and physicochemical data on a set of publicly disclosed compounds\.*CHEMBL3301361\. ChEMBL*, 6019, 2015\.
- Wu et al\. \[2018\]Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande\.Moleculenet: a benchmark for molecular machine learning\.*Chemical science*, 9\(2\):513–530, 2018\.
- Zdrazil et al\. \[2024\]Barbara Zdrazil, Eloy Felix, Fiona Hunter, Emma J Manners, James Blackshaw, Sybilla Corbett, Marleen De Veij, Harris Ioannidis, David Mendez Lopez, Juan F Mosquera, et al\.The chembl database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods\.*Nucleic acids research*, 52\(D1\):D1180–D1192, 2024\.

## Appendix ABRR Posterior and Predictive Distribution

Conjugacy yields the closed\-form posteriorp⁡\(𝒘∣𝒟train\)=𝒩⁡\(𝒘∣𝝁𝒘,𝚺𝒘\)p\(\\bm\{w\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=\\mathcal\{N\}\(\\bm\{w\}\\mid\\bm\{\\mu\}\_\{\\bm\{w\}\},\\bm\{\\Sigma\}\_\{\\bm\{w\}\}\)with

𝚺𝒘\\displaystyle\\bm\{\\Sigma\}\_\{\\bm\{w\}\}=\(α​𝑰d\+λ​𝑯⊤​𝑯\)−1,\\displaystyle=\\bigl\(\\alpha\\,\\bm\{I\}\_\{d\}\+\\lambda\\,\\bm\{H\}^\{\\top\}\\bm\{H\}\\bigr\)^\{\-1\},\(17\)𝝁𝒘\\displaystyle\\bm\{\\mu\}\_\{\\bm\{w\}\}=λ​𝚺𝒘​𝑯⊤​𝒚,\\displaystyle=\\lambda\\,\\bm\{\\Sigma\}\_\{\\bm\{w\}\}\\,\\bm\{H\}^\{\\top\}\\bm\{y\},\(18\)where𝑯∈ℝntr×d\\bm\{H\}\\in\\mathbb\{R\}^\{n\_\{\\mathrm\{tr\}\}\\times d\}stacks the training representations row\-wise and𝒚∈ℝntr\\bm\{y\}\\in\\mathbb\{R\}^\{n\_\{\\mathrm\{tr\}\}\}collects labels\. The predictive mean and variance in[eq\.11](https://arxiv.org/html/2609.12386#S2.E11)follow by marginalizing over𝒘\\bm\{w\}\.

## Appendix BFull Proof of[Proposition2\.1](https://arxiv.org/html/2609.12386#S2.Thmproposition1)

We provide the complete derivation, conditioning throughout on the fitted source dataset𝒟train\\mathcal\{D\}\_\{\\mathrm\{train\}\}\.

By Bayes’ rule in the target domain,

pt​\(y∣𝒉,𝒟train\)=pt​\(𝒉∣y,𝒟train\)​pt​\(y\|𝒟train\)pt​\(𝒉∣𝒟train\)\.p\_\{\\mathrm\{t\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=\\frac\{p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid y,\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\;p\_\{\\mathrm\{t\}\}\(y\|\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\{p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\.\(19\)By[Assumption2\.1](https://arxiv.org/html/2609.12386#S2.Thmassumption1),

pt​\(𝒉∣y,𝒟train\)=ps​\(𝒉∣y,𝒟train\)\.p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid y,\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid y,\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\.Substituting this into[eq\.19](https://arxiv.org/html/2609.12386#A2.E19)gives

pt​\(y∣𝒉,𝒟train\)=ps​\(𝒉∣y,𝒟train\)​pt​\(y\|𝒟train\)pt​\(𝒉∣𝒟train\)\.p\_\{\\mathrm\{t\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=\\frac\{p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid y,\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\;p\_\{\\mathrm\{t\}\}\(y\|\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\{p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\.\(20\)
Applying Bayes’ rule in the source domain,

ps​\(𝒉∣y,𝒟train\)=ps​\(y∣𝒉,𝒟train\)​ps​\(𝒉∣𝒟train\)ps​\(y\|𝒟train\)\.p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid y,\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=\\frac\{p\_\{\\mathrm\{s\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\;p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\{p\_\{\\mathrm\{s\}\}\(y\|\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\.\(21\)Inserting[eq\.21](https://arxiv.org/html/2609.12386#A2.E21)into[eq\.20](https://arxiv.org/html/2609.12386#A2.E20), we obtain

pt​\(y∣𝒉,𝒟train\)\\displaystyle p\_\{\\mathrm\{t\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=ps​\(y∣𝒉,𝒟train\)⋅pt​\(y\|𝒟train\)ps​\(y\|𝒟train\)\\displaystyle=p\_\{\\mathrm\{s\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\cdot\\frac\{p\_\{\\mathrm\{t\}\}\(y\|\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\{p\_\{\\mathrm\{s\}\}\(y\|\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}⋅ps​\(𝒉∣𝒟train\)pt​\(𝒉∣𝒟train\)\.\\displaystyle\\quad\\cdot\\frac\{p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\{p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\.\(22\)
#### Identifying the label\-marginal ratio\.

At this point, the derivation introduces the conditional label\-marginal ratiopt​\(y∣𝒟train\)/ps​\(y∣𝒟train\)p\_\{\\mathrm\{t\}\}\(y\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)/p\_\{\\mathrm\{s\}\}\(y\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\. Since𝒟train\\mathcal\{D\}\_\{\\mathrm\{train\}\}is drawn entirely from the source domainpsp\_\{\\mathrm\{s\}\}, it carries no information about the target label distribution\. We therefore havept​\(y∣𝒟train\)=pt​\(y\)p\_\{\\mathrm\{t\}\}\(y\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=p\_\{\\mathrm\{t\}\}\(y\)\. By[Assumption2\.3](https://arxiv.org/html/2609.12386#S2.Thmassumption3),ps​\(y∣𝒟train\)=ps​\(y\)p\_\{\\mathrm\{s\}\}\(y\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=p\_\{\\mathrm\{s\}\}\(y\)\. It follows that

pt​\(y∣𝒟train\)ps​\(y∣𝒟train\)=pt​\(y\)ps​\(y\)=r⁡\(y\)\.\\frac\{p\_\{\\mathrm\{t\}\}\(y\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\{p\_\{\\mathrm\{s\}\}\(y\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}=\\frac\{p\_\{\\mathrm\{t\}\}\(y\)\}\{p\_\{\\mathrm\{s\}\}\(y\)\}=r\(y\)\.\(23\)
Usingr⁡\(y\)r\(y\), we rewrite[eq\.22](https://arxiv.org/html/2609.12386#A2.E22)as

pt​\(y∣𝒉,𝒟train\)=ps​\(y∣𝒉,𝒟train\)​r​\(y\)​ps​\(𝒉∣𝒟train\)pt​\(𝒉∣𝒟train\)\.p\_\{\\mathrm\{t\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=p\_\{\\mathrm\{s\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\,r\(y\)\\,\\frac\{p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\{p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\.\(24\)

#### Identifying1/Z⁡\(𝒉,𝒟train\)1/Z\(\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\.

Using the law of total probability, Bayes’ rule, and[Assumption2\.1](https://arxiv.org/html/2609.12386#S2.Thmassumption1),

pt​\(𝒉∣𝒟train\)\\displaystyle p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=∫pt​\(𝒉∣y′,𝒟train\)​pt​\(y′\)​d​y′\\displaystyle=\\int p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid y^\{\\prime\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\;p\_\{\\mathrm\{t\}\}\(y^\{\\prime\}\)\\;\\mathrm\{d\}y^\{\\prime\}=∫ps​\(𝒉∣y′,𝒟train\)​pt​\(y′\)​d​y′\\displaystyle=\\int p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid y^\{\\prime\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\;p\_\{\\mathrm\{t\}\}\(y^\{\\prime\}\)\\;\\mathrm\{d\}y^\{\\prime\}=∫ps​\(y′∣𝒉,𝒟train\)​ps​\(𝒉∣𝒟train\)ps​\(y′\)​pt​\(y′\)​d​y′\\displaystyle=\\int\\frac\{p\_\{\\mathrm\{s\}\}\(y^\{\\prime\}\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\;p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\{p\_\{\\mathrm\{s\}\}\(y^\{\\prime\}\)\}\\;p\_\{\\mathrm\{t\}\}\(y^\{\\prime\}\)\\;\\mathrm\{d\}y^\{\\prime\}=ps​\(𝒉∣𝒟train\)​∫ps​\(y′∣𝒉,𝒟train\)​r​\(y′\)​d​y′\.\\displaystyle=p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\int p\_\{\\mathrm\{s\}\}\(y^\{\\prime\}\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\,r\(y^\{\\prime\}\)\\,\\mathrm\{d\}y^\{\\prime\}\.\(25\)Hence, with

Z⁡\(𝒉,𝒟train\)\\displaystyle Z\(\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\):=∫ps​\(y′∣𝒉,𝒟train\)​r​\(y′\)​d​y′\\displaystyle:=\\int p\_\{\\mathrm\{s\}\}\(y^\{\\prime\}\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\,r\(y^\{\\prime\}\)\\,\\mathrm\{d\}y^\{\\prime\}=𝔼Y∼ps\(⋅∣𝒉,𝒟train\)\[r\(Y\)\],\\displaystyle=\\mathbb\{E\}\_\{Y\\sim p\_\{\\mathrm\{s\}\}\(\\,\\cdot\\,\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\\\!\\bigl\[r\(Y\)\\bigr\],\(26\)we have

pt​\(𝒉∣𝒟train\)\\displaystyle p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=ps​\(𝒉∣𝒟train\)​Z​\(𝒉,𝒟train\),\\displaystyle=p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\,Z\(\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\),ps​\(𝒉∣𝒟train\)pt​\(𝒉∣𝒟train\)\\displaystyle\\frac\{p\_\{\\mathrm\{s\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\{p\_\{\\mathrm\{t\}\}\(\\bm\{h\}\\mid\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}=1Z⁡\(𝒉,𝒟train\)\.\\displaystyle=\\frac\{1\}\{Z\(\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\}\.\(27\)Substituting[eq\.27](https://arxiv.org/html/2609.12386#A2.E27)into[eq\.24](https://arxiv.org/html/2609.12386#A2.E24)yields

pt​\(y∣𝒉,𝒟train\)=ps​\(y∣𝒉,𝒟train\)​r​\(y\)Z⁡\(𝒉,𝒟train\),p\_\{\\mathrm\{t\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=\\frac\{p\_\{\\mathrm\{s\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\,r\(y\)\}\{Z\(\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\},\(28\)which proves[Proposition2\.1](https://arxiv.org/html/2609.12386#S2.Thmproposition1)\.

## Appendix CPopulation Tilt and the Estimated Log\-Density Ratio

Under the exponential tilt model,

pt​\(y\)∝ps​\(y\)​exp⁡\(β​y\),p\_\{\\mathrm\{t\}\}\(y\)\\propto p\_\{\\mathrm\{s\}\}\(y\)\\exp\(\\beta y\),\(29\)the population density ratio is

r⁡\(y\)\\displaystyle r\(y\)=pt​\(y\)ps​\(y\)=exp⁡\(β​y\)Zr,\\displaystyle=\\frac\{p\_\{\\mathrm\{t\}\}\(y\)\}\{p\_\{\\mathrm\{s\}\}\(y\)\}=\\frac\{\\exp\(\\beta y\)\}\{Z\_\{r\}\},Zr\\displaystyle Z\_\{r\}:=∫ps​\(y\)​exp⁡\(β​y\)​𝑑y\.\\displaystyle:=\\int p\_\{\\mathrm\{s\}\}\(y\)\\,\\exp\(\\beta y\)\\,\\mathrm\{d\}y\.\(30\)Therefore, the log\-density ratio is exactly linear inyy:

log⁡r⁡\(y\)=β0\+β1​y,β1=β,β0=−log⁡Zr\.\\log r\(y\)=\\beta\_\{0\}\+\\beta\_\{1\}y,\\qquad\\beta\_\{1\}=\\beta,\\qquad\\beta\_\{0\}=\-\\log Z\_\{r\}\.\(31\)
Ifps​\(y\)p\_\{\\mathrm\{s\}\}\(y\)is also Gaussian,ps​\(y\)=𝒩⁡\(μs,σs2\)p\_\{\\mathrm\{s\}\}\(y\)=\\mathcal\{N\}\(\\mu\_\{s\},\\sigma\_\{s\}^\{2\}\), then the moment\-generating function gives

Zr=exp⁡\(β​μs\+12​β2​σs2\),Z\_\{r\}=\\exp\\Bigl\(\\beta\\mu\_\{s\}\+\\tfrac\{1\}\{2\}\\beta^\{2\}\\sigma\_\{s\}^\{2\}\\Bigr\),\(32\)so that

β0=−β​μs−12​β2​σs2\.\\beta\_\{0\}=\-\\beta\\mu\_\{s\}\-\\tfrac\{1\}\{2\}\\beta^\{2\}\\sigma\_\{s\}^\{2\}\.\(33\)This Gaussian calculation is only an illustrative expansion of the interceptβ0\\beta\_\{0\}\. The log\-linearity in[eq\.31](https://arxiv.org/html/2609.12386#A3.E31)follows directly from the exponential tilt model and does not require Gaussianity ofps​\(y\)p\_\{\\mathrm\{s\}\}\(y\)\.

In practice, we estimate the log\-linear form in[eq\.31](https://arxiv.org/html/2609.12386#A3.E31)by fitting a logistic regression classifier that distinguishes source labels from target pseudo\-labels\. This yields

log⁡r^​\(y\)=β^0\+β^1​y\.\\log\\hat\{r\}\(y\)=\\hat\{\\beta\}\_\{0\}\+\\hat\{\\beta\}\_\{1\}y\.\(34\)The target sample is unlabeled\. The pseudo\-labels are noisy surrogates for the true target labels\.r^​\(y\)\\hat\{r\}\(y\)should therefore be viewed as a practical approximation to the population density ratio, not an oracle estimator\. We use this estimated ratio throughout the LSA score construction\.

## Appendix DClosed\-Form Derivation ofZ^​\(𝒉\)\\hat\{Z\}\(\\bm\{h\}\)and the Tilted Predictive

Under the BRR model,

ps​\(y∣𝒉,𝒟train\)=𝒩⁡\(y∣μ⁡\(𝒉\),σ2​\(𝒉\)\),p\_\{\\mathrm\{s\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=\\mathcal\{N\}\(y\\mid\\mu\(\\bm\{h\}\),\\sigma^\{2\}\(\\bm\{h\}\)\),and the estimated density ratio is

r^​\(y\)=exp⁡\(β^0\+β^1​y\)\.\\hat\{r\}\(y\)=\\exp\(\\hat\{\\beta\}\_\{0\}\+\\hat\{\\beta\}\_\{1\}y\)\.The estimated normalizing constant is therefore

Z^​\(𝒉\)\\displaystyle\\hat\{Z\}\(\\bm\{h\}\)=∫ps​\(y∣𝒉,𝒟train\)​r^​\(y\)​𝑑y\\displaystyle=\\int p\_\{\\mathrm\{s\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\,\\hat\{r\}\(y\)\\,\\mathrm\{d\}y=∫𝒩⁡\(y∣μ⁡\(𝒉\),σ2​\(𝒉\)\)​exp⁡\(β^0\+β^1​y\)​𝑑y\\displaystyle=\\int\\mathcal\{N\}\(y\\mid\\mu\(\\bm\{h\}\),\\sigma^\{2\}\(\\bm\{h\}\)\)\\,\\exp\(\\hat\{\\beta\}\_\{0\}\+\\hat\{\\beta\}\_\{1\}y\)\\,\\mathrm\{d\}y=exp⁡\(β^0\)​𝔼Y∼𝒩⁡\(μ⁡\(𝒉\),σ2​\(𝒉\)\)​\[exp⁡\(β^1​Y\)\]\\displaystyle=\\exp\(\\hat\{\\beta\}\_\{0\}\)\\,\\mathbb\{E\}\_\{Y\\sim\\mathcal\{N\}\(\\mu\(\\bm\{h\}\),\\sigma^\{2\}\(\\bm\{h\}\)\)\}\\\!\\bigl\[\\exp\(\\hat\{\\beta\}\_\{1\}Y\)\\bigr\]=exp⁡\(β^0\)​exp⁡\(β^1​μ​\(𝒉\)\+12​β^12​σ2​\(𝒉\)\)\\displaystyle=\\exp\(\\hat\{\\beta\}\_\{0\}\)\\,\\exp\\\!\\Bigl\(\\hat\{\\beta\}\_\{1\}\\mu\(\\bm\{h\}\)\+\\tfrac\{1\}\{2\}\\hat\{\\beta\}\_\{1\}^\{2\}\\sigma^\{2\}\(\\bm\{h\}\)\\Bigr\)=exp⁡\(β^0\+β^1​μ​\(𝒉\)\+12​β^12​σ2​\(𝒉\)\),\\displaystyle=\\exp\\\!\\Bigl\(\\hat\{\\beta\}\_\{0\}\+\\hat\{\\beta\}\_\{1\}\\mu\(\\bm\{h\}\)\+\\tfrac\{1\}\{2\}\\hat\{\\beta\}\_\{1\}^\{2\}\\sigma^\{2\}\(\\bm\{h\}\)\\Bigr\),\(35\)where the third equality uses the moment\-generating function of the Gaussian:

𝔼⁡\[exp⁡\(t​Y\)\]=exp⁡\(t​μ\+12​t2​σ2\),Y∼𝒩⁡\(μ,σ2\)\.\\mathbb\{E\}\[\\exp\(tY\)\]=\\exp\\\!\\Bigl\(t\\mu\+\\tfrac\{1\}\{2\}t^\{2\}\\sigma^\{2\}\\Bigr\),\\qquad Y\\sim\\mathcal\{N\}\(\\mu,\\sigma^\{2\}\)\.

## Appendix EGaussian Form of the Tilted Predictive

Starting from[eq\.7](https://arxiv.org/html/2609.12386#S2.E7), under the BRR predictive

ps​\(y∣𝒉,𝒟train\)=𝒩⁡\(y∣μ⁡\(𝒉\),σ2​\(𝒉\)\)p\_\{\\mathrm\{s\}\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=\\mathcal\{N\}\(y\\mid\\mu\(\\bm\{h\}\),\\sigma^\{2\}\(\\bm\{h\}\)\)and the log\-linear ratio model

r^​\(y\)=exp⁡\(β^0\+β^1​y\),\\hat\{r\}\(y\)=\\exp\(\\hat\{\\beta\}\_\{0\}\+\\hat\{\\beta\}\_\{1\}y\),we have

p^t​\(y∣𝒉,𝒟train\)\\displaystyle\\hat\{p\}\_\{t\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)∝𝒩⁡\(y∣μ⁡\(𝒉\),σ2​\(𝒉\)\)​exp⁡\(β^0\+β^1​y\)\\displaystyle\\propto\\mathcal\{N\}\(y\\mid\\mu\(\\bm\{h\}\),\\sigma^\{2\}\(\\bm\{h\}\)\)\\exp\(\\hat\{\\beta\}\_\{0\}\+\\hat\{\\beta\}\_\{1\}y\)\(36\)=12​π​σ2​\(𝒉\)​exp⁡\(−\(y−μ⁡\(𝒉\)\)22​σ2​\(𝒉\)\+β^0CLOSE\\displaystyle=\\frac\{1\}\{\\sqrt\{2\\pi\\sigma^\{2\}\(\\bm\{h\}\)\}\}\\exp\\\!\\Bigl\(\-\\frac\{\(y\-\\mu\(\\bm\{h\}\)\)^\{2\}\}\{2\\sigma^\{2\}\(\\bm\{h\}\)\}\+\\hat\{\\beta\}\_\{0\}OPEN\+β^1​y\)\.\\displaystyle\\qquad\\qquad\\qquad\\qquad\\qquad\+\\hat\{\\beta\}\_\{1\}y\\Bigr\)\.Completing the square in the exponent with respect toyy,

−\(y−μ⁡\(𝒉\)\)22​σ2​\(𝒉\)\+β^1​y\\displaystyle\-\\frac\{\(y\-\\mu\(\\bm\{h\}\)\)^\{2\}\}\{2\\sigma^\{2\}\(\\bm\{h\}\)\}\+\\hat\{\\beta\}\_\{1\}y=−y2−2​μ​\(𝒉\)​y\+μ​\(𝒉\)22​σ2​\(𝒉\)\+β^1​y\\displaystyle=\-\\frac\{y^\{2\}\-2\\mu\(\\bm\{h\}\)y\+\\mu\(\\bm\{h\}\)^\{2\}\}\{2\\sigma^\{2\}\(\\bm\{h\}\)\}\+\\hat\{\\beta\}\_\{1\}y=−y2−2​\(μ⁡\(𝒉\)\+β^1​σ2​\(𝒉\)\)​y\+μ​\(𝒉\)22​σ2​\(𝒉\)\\displaystyle=\-\\frac\{y^\{2\}\-2\\bigl\(\\mu\(\\bm\{h\}\)\+\\hat\{\\beta\}\_\{1\}\\sigma^\{2\}\(\\bm\{h\}\)\\bigr\)y\+\\mu\(\\bm\{h\}\)^\{2\}\}\{2\\sigma^\{2\}\(\\bm\{h\}\)\}=−\(y−μ∗​\(𝒉\)\)22​σ2​\(𝒉\)\+\(μ∗​\(𝒉\)\)2−μ​\(𝒉\)22​σ2​\(𝒉\)\\displaystyle=\-\\frac\{\(y\-\\mu^\{\*\}\(\\bm\{h\}\)\)^\{2\}\}\{2\\sigma^\{2\}\(\\bm\{h\}\)\}\+\\frac\{\(\\mu^\{\*\}\(\\bm\{h\}\)\)^\{2\}\-\\mu\(\\bm\{h\}\)^\{2\}\}\{2\\sigma^\{2\}\(\\bm\{h\}\)\}=−\(y−μ∗​\(𝒉\)\)22​σ2​\(𝒉\)\+β^1​μ​\(𝒉\)\+β^12​σ2​\(𝒉\)2,\\displaystyle=\-\\frac\{\(y\-\\mu^\{\*\}\(\\bm\{h\}\)\)^\{2\}\}\{2\\sigma^\{2\}\(\\bm\{h\}\)\}\+\\hat\{\\beta\}\_\{1\}\\mu\(\\bm\{h\}\)\+\\tfrac\{\\hat\{\\beta\}\_\{1\}^\{2\}\\sigma^\{2\}\(\\bm\{h\}\)\}\{2\},\(37\)where

μ∗​\(𝒉\):=μ⁡\(𝒉\)\+β^1​σ2​\(𝒉\)\.\\mu^\{\*\}\(\\bm\{h\}\):=\\mu\(\\bm\{h\}\)\+\\hat\{\\beta\}\_\{1\}\\sigma^\{2\}\(\\bm\{h\}\)\.\(38\)The remaining terms are independent ofyyand are absorbed intoZ^​\(𝒉\)\\hat\{Z\}\(\\bm\{h\}\)\. Therefore,

p^t​\(y∣𝒉,𝒟train\)=𝒩⁡\(y\|μ∗​\(𝒉\),σ2​\(𝒉\)\)\.\\hat\{p\}\_\{t\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)=\\mathcal\{N\}\\\!\\bigl\(y\\;\\big\|\\;\\mu^\{\*\}\(\\bm\{h\}\),\\sigma^\{2\}\(\\bm\{h\}\)\\bigr\)\.\(39\)Thus, under the estimated log\-linear ratio model, exponential tilting shifts the predictive mean byβ^1​σ2​\(𝒉\)\\hat\{\\beta\}\_\{1\}\\sigma^\{2\}\(\\bm\{h\}\)\. The variance is preserved\.

The LSA score is the negative log\-density ofp^t​\(y∣𝒉,𝒟train\)\\hat\{p\}\_\{t\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)up to the additive constant12​log⁡\(2​π\)\\frac\{1\}\{2\}\\log\(2\\pi\):

sLSA​\(𝒉,y\)=−log⁡p^t​\(y∣𝒉,𝒟train\)=\(y−μ∗​\(𝒉\)\)22​σ2​\(𝒉\)\+12​log⁡σ2​\(𝒉\)\+const\.s\_\{\\mathrm\{LSA\}\}\(\\bm\{h\},y\)=\-\\log\\hat\{p\}\_\{t\}\(y\\mid\\bm\{h\},\\mathcal\{D\}\_\{\\mathrm\{train\}\}\)\\\\ =\\frac\{\(y\-\\mu^\{\*\}\(\\bm\{h\}\)\)^\{2\}\}\{2\\sigma^\{2\}\(\\bm\{h\}\)\}\+\\tfrac\{1\}\{2\}\\log\\sigma^\{2\}\(\\bm\{h\}\)\+\\text\{const\.\}\(40\)After dropping the additive constant, this is exactly the form used in[eq\.14](https://arxiv.org/html/2609.12386#S2.E14)\.

## Appendix FScore Inversion

#### Exact set inversion\.

For a test input𝒉∗\\bm\{h\}^\{\*\}, the exact sublevel set\{y:sLSA​\(𝒉∗,y\)≤q^\}\\\{y:s\_\{\\mathrm\{LSA\}\}\(\\bm\{h\}^\{\*\},y\)\\leq\\hat\{q\}\\\}is determined by

\(y−μ∗​\(𝒉∗\)\)22​σ2​\(𝒉∗\)\+12​log⁡σ2​\(𝒉∗\)\\displaystyle\\frac\{\(y\-\\mu^\{\*\}\(\\bm\{h\}^\{\*\}\)\)^\{2\}\}\{2\\sigma^\{2\}\(\\bm\{h\}^\{\*\}\)\}\+\\frac\{1\}\{2\}\\log\\sigma^\{2\}\(\\bm\{h\}^\{\*\}\)≤q^\\displaystyle\\leq\\hat\{q\}\(y−μ∗​\(𝒉∗\)\)2\\displaystyle\(y\-\\mu^\{\*\}\(\\bm\{h\}^\{\*\}\)\)^\{2\}≤2​σ2​\(𝒉∗\)​\(q^−log⁡σ2​\(𝒉∗\)2\)\.\\displaystyle\\leq 2\\sigma^\{2\}\(\\bm\{h\}^\{\*\}\)\\bigl\(\\hat\{q\}\-\\tfrac\{\\log\\sigma^\{2\}\(\\bm\{h\}^\{\*\}\)\}\{2\}\\bigr\)\.\(41\)Hence the exact prediction set is

Cexact​\(𝒉∗\)=\{\[μ∗±σ​2​\(q^−12​log⁡σ2\)\],q^≥12​log⁡σ2​\(𝒉∗\),∅,otherwise\.C\_\{\\mathrm\{exact\}\}\(\\bm\{h\}^\{\*\}\)=\{\\begin\{cases\}\\bigl\[\\mu^\{\*\}\\pm\\sigma\\sqrt\{2\(\\hat\{q\}\-\\tfrac\{1\}\{2\}\\log\\sigma^\{2\}\)\}\\bigr\],&\\hat\{q\}\\geq\\tfrac\{1\}\{2\}\\log\\sigma^\{2\}\(\\bm\{h\}^\{\*\}\),\\\\\[4\.0pt\] \\varnothing,&\\text\{otherwise\}\.\\end\{cases\}\}\(42\)

#### Clipped implementation\.

In our implementation, we use the clipped version

C⁡\(𝒉∗\)=\[μ∗​\(𝒉∗\)±σ⁡\(𝒉∗\)​2​max⁡\(q^−12​log⁡σ2​\(𝒉∗\),0\)\],C\(\\bm\{h\}^\{\*\}\)=\\Bigl\[\\mu^\{\*\}\(\\bm\{h\}^\{\*\}\)\\pm\\sigma\(\\bm\{h\}^\{\*\}\)\\sqrt\{2\\max\\bigl\(\\hat\{q\}\-\\tfrac\{1\}\{2\}\\log\\sigma^\{2\}\(\\bm\{h\}^\{\*\}\),0\\bigr\)\}\\,\\Bigr\],\(43\)which coincides with[eq\.42](https://arxiv.org/html/2609.12386#A6.E42)whenever the exact sublevel set is nonempty and avoids numerical issues when the term inside the square root is negative\.

## Appendix GAblation

### G\.1Gaussian Mean Weight Estimation

In the main text, density ratios are estimated via logistic regression\. Here we consider an alternative estimator based on Gaussian fitting\. Specifically, we fit separate Gaussian distributions to the source and target label marginals, obtaining meansμ^s\\hat\{\\mu\}\_\{s\},μ^t\\hat\{\\mu\}\_\{t\}and variancesσ^s2\\hat\{\\sigma\}\_\{s\}^\{2\},σ^t2\\hat\{\\sigma\}\_\{t\}^\{2\}\. We then construct the estimated log\-density ratio using the fitted means\. The source variance is shared:

log⁡r^GM​\(y\)\\displaystyle\\log\\hat\{r\}\_\{\\mathrm\{GM\}\}\(y\)=\(y−μ^s\)2−\(y−μ^t\)22​σ^s2\\displaystyle=\\frac\{\(y\-\\hat\{\\mu\}\_\{s\}\)^\{2\}\-\(y\-\\hat\{\\mu\}\_\{t\}\)^\{2\}\}\{2\\hat\{\\sigma\}\_\{s\}^\{2\}\}=\(μ^t−μ^s\)​\(2​y−μ^s−μ^t\)2​σ^s2\.\\displaystyle=\\frac\{\(\\hat\{\\mu\}\_\{t\}\-\\hat\{\\mu\}\_\{s\}\)\(2y\-\\hat\{\\mu\}\_\{s\}\-\\hat\{\\mu\}\_\{t\}\)\}\{2\\hat\{\\sigma\}\_\{s\}^\{2\}\}\.\(44\)This estimator is linear inyy, so it preserves the same functional form

log⁡r^GM​\(y\)=β^0GM\+β^1GM​y\\log\\hat\{r\}\_\{\\mathrm\{GM\}\}\(y\)=\\hat\{\\beta\}\_\{0\}^\{\\mathrm\{GM\}\}\+\\hat\{\\beta\}\_\{1\}^\{\\mathrm\{GM\}\}yused in the derivation of[eq\.12](https://arxiv.org/html/2609.12386#S2.E12)\. The same closed\-form normalizer and the same LSA score construction therefore remain valid after substituting the corresponding fitted coefficients\.

[table2](https://arxiv.org/html/2609.12386#A7.T2)reports target\-domain empirical coverage and average interval length under this Gaussian\-mean weight estimator\. The qualitative trends are consistent with the logistic regression results in[table1](https://arxiv.org/html/2609.12386#S3.T1)\. The LSA score yields shorter intervals and maintains comparable empirical coverage under label shift\. These results indicate that the benefits of the LSA score are not tied to a single weight\-estimation procedure\. The estimator need only preserve the log\-linear functional form inyy\.

Table 2:Coverage and average interval length under exponential tilt label shift with Gaussian mean weight estimation\. Target coverage is0\.90\.9\. Results are averaged over1,0001\{,\}000random seeds \(±\\pmone standard deviation\)\.Boldindicates the best value perβ\\beta\.
### G\.2Additional Robustness Study on Lipophilicity

To assess whether the observed effect is specific to AqSolDB, we also evaluate the same protocol on the Lipophilicity dataset\. Despite its smaller size and different label range, we observe the same qualitative trend\. The three methods are nearly identical atβ=0\\beta=0\. Under increasing negative label shift, the proposed LSA score produces consistently shorter intervals than bothResidualandSource score\. Atβ=−0\.5\\beta=\-0\.5, the interval length is reduced by approximately4\.6%4\.6\\%relative to Source score and5\.3%5\.3\\%relative to Residual\. Empirical coverage remains comparable and is slightly higher for LSA\. This suggests that the effect is not unique to AqSolDB\.

Table 3:Results on the Lipophilicity benchmark under exponential\-tilt label shift\. We report empirical coverage and average interval length \(mean±\\pmstandard deviation\) over repeated random trials\. The proposed LSA score reproduces the same qualitative trend observed on AqSolDB\. The three methods are nearly identical atβ=0\\beta=0\. Under increasing negative label shift, LSA yields consistently shorter intervals than both residual and source score baselines, with comparable or slightly improved empirical coverage\.

## Appendix HEncoder Pretraining Details

All experiments are based on a frozen molecular BART encoder used as the backbone representation model\. The encoder was pretrained on a large\-scale unlabeled molecular corpus assembled from public databases, including PubChem\[[Kim et al\., 2023](https://arxiv.org/html/2609.12386#bib.bib11)\], Chembl\[[Zdrazil et al\., 2024](https://arxiv.org/html/2609.12386#bib.bib31)\], Coconut\[[Sorokina et al\., 2021](https://arxiv.org/html/2609.12386#bib.bib24)\], Drugbank\[[Knox et al\., 2024](https://arxiv.org/html/2609.12386#bib.bib12)\], Enamine\[[Shivanyuk et al\., 2007](https://arxiv.org/html/2609.12386#bib.bib21)\], and ZINC\[[Irwin et al\., 2012](https://arxiv.org/html/2609.12386#bib.bib10)\]\. After collection and filtering, the final corpus comprised212,593,454212\{,\}593\{,\}454molecules\. The model was pretrained for55epochs and kept fixed throughout all downstream experiments\.

[table4](https://arxiv.org/html/2609.12386#A8.T4)summarizes the backbone architecture and the main pretraining hyperparameters\.

Table 4:Backbone encoder architecture and pretraining setup\.

Similar Articles

Empirical Bayes Conformal Prediction for Vision and Language Models

arXiv cs.LG

This paper introduces an empirical Bayes conformal prediction framework that uses r-values to incorporate score variability into nonconformity scores, improving ranking stability and reducing set size while preserving coverage for vision and language models.

Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations

arXiv cs.CL

This paper proposes a conformal prediction framework for LLMs that leverages internal representations rather than output-level statistics, introducing Layer-Wise Information (LI) scores as nonconformity measures to improve validity-efficiency trade-offs under distribution shift. The method demonstrates stronger robustness to calibration-deployment mismatch compared to text-level baselines across QA benchmarks.