I-SAFE: Wasserstein Coherence Metrics for Structural Auditing of Scientific AI Models

arXiv cs.LG Papers

Summary

This paper introduces I-SAFE, a post-hoc distributional auditing framework for scientific AI models using Wasserstein Coherence Metrics, which reveals structural differences in model outputs that accuracy-based evaluation fails to capture. Demonstrated on drug-target interaction prediction, the framework is model-agnostic and applicable to any domain with structured inputs and external priors.

arXiv:2605.21731v1 Announce Type: new Abstract: Deep learning models are increasingly used in scientific prediction tasks where strong benchmark performance is often interpreted as evidence of scientifically meaningful behavior. This interpretation is fragile, as models may exploit shortcut features, dataset-specific regularities, or distributional biases that are predictive on held-out data but not aligned with domain-relevant structure. To address this limitation, we introduce the \textsc{I-SAFE} (Interventional Secure, Accurate, Fair and Explainable) framework, a post-hoc distributional auditing framework for scientific AI models centered on the Wasserstein Coherence Metric (WCM). Given a trained black-box predictor and an external structural prior encoding domain knowledge about task-relevant input structure, \textsc{I-SAFE} evaluates raw model outputs under structurally guided perturbations of the input. The proposed audit measures output-distribution coherence through three complementary metrics: a Quantile-Based Metric (QBM) for location-level coherence, the WCM for ordinal coherence, and a translation-invariant WCM variant for shape coherence. We instantiate \textsc{I-SAFE} on drug--target interaction (DTI) prediction using the Davis kinase benchmark, KLIFS (Kinase--Ligand Interaction Fingerprints and Structures) binding-pocket annotations, and three sequence-based DTI models: DeepConvDTI, DeepDTA, and TAPB. Although the models operate in a comparable predictive regime, \textsc{I-SAFE} reveals substantially different distributional response profiles, a distinction invisible to accuracy-based evaluation. The framework is model-agnostic and applicable to any domain where inputs admit a structured decomposition and an external prior is available.
Original Article
View Cached Full Text

Cached at: 05/22/26, 08:51 AM

# I-SAFE: Wasserstein Coherence Metrics for Structural Auditing of Scientific AI Models
Source: [https://arxiv.org/html/2605.21731](https://arxiv.org/html/2605.21731)
Barbara TarantinoDepartment of Economics, University of Pavia, Via S\. Felice Al Monastero, 5, Pavia, 27100, Italy, email: barbara\.tarantino@unipv\.itGennaro AuricchioDepartment of Mathematics, University of Padua, Via Trieste, 63, Padua, 35131, Italy, email: gennaro\.auricchio@unipd\.itPaolo GiudiciDepartment of Economics, University of Pavia, Via S\. Felice Al Monastero, 5, Pavia, 27100, Italy, email: paolo\.giudici@unipv\.it

###### Abstract

Deep learning models are increasingly used in scientific prediction tasks where strong benchmark performance is often interpreted as evidence of scientifically meaningful behavior\. This interpretation is fragile, as models may exploit shortcut features, dataset\-specific regularities, or distributional biases that are predictive on held\-out data but not aligned with domain\-relevant structure\. To address this limitation, we introduce theI\-SAFE\(Interventional Secure, Accurate, Fair and Explainable\) framework, a post\-hoc distributional auditing framework for scientific AI models centered on the Wasserstein Coherence Metric \(WCM\)\. Given a trained black\-box predictor and an external structural prior encoding domain knowledge about task\-relevant input structure,I\-SAFEevaluates raw model outputs under structurally guided perturbations of the input\. The proposed audit measures output\-distribution coherence through three complementary metrics: a Quantile\-Based Metric \(QBM\) for location\-level coherence, the WCM for ordinal coherence, and a translation\-invariant WCM variant for shape coherence\. We instantiateI\-SAFEon drug–target interaction \(DTI\) prediction using the Davis kinase benchmark, KLIFS \(Kinase–Ligand Interaction Fingerprints and Structures\) binding\-pocket annotations, and three sequence\-based DTI models: DeepConvDTI, DeepDTA, and TAPB\. Although the models operate in a comparable predictive regime,I\-SAFEreveals substantially different distributional response profiles, a distinction invisible to accuracy\-based evaluation\. The framework is model\-agnostic and applicable to any domain where inputs admit a structured decomposition and an external prior is available\.

## 1Introduction

Held\-out predictive performance remains the dominant criterion for evaluating deep learning models in scientific machine learning, including molecular property prediction, drug–target interaction \(DTI\), and gene perturbation response\[[37](https://arxiv.org/html/2605.21731#bib.bib37),[14](https://arxiv.org/html/2605.21731#bib.bib14),[21](https://arxiv.org/html/2605.21731#bib.bib21)\]\. In these settings, strong benchmark performance is often interpreted as evidence of scientifically meaningful behavior\. This interpretation is fragile, as predictive accuracy measures are affected by statistical regularities that are present in the benchmark distribution, and not only on whether the model rightly identified the relevant structures of the input\[[25](https://arxiv.org/html/2605.21731#bib.bib25)\]\. This gap is central to modern machine learning, where high\-capacity models can achieve strong performance by exploiting shortcuts that do not necessarily align with the scientific structure of interest\[[12](https://arxiv.org/html/2605.21731#bib.bib12),[15](https://arxiv.org/html/2605.21731#bib.bib15)\]\.

In scientific applications, this limitation is empirical as well as conceptual\. For example, high\-performing protein–ligand affinity models have been shown to rely substantially on ligand memorisation rather than interaction\-specific information\[[23](https://arxiv.org/html/2605.21731#bib.bib23)\]\. Specifically, in DTI, apparent progress on standard benchmarks can be influenced by target bias, scaffold effects, leakage, and other forms of distributional bias\[[20](https://arxiv.org/html/2605.21731#bib.bib20),[35](https://arxiv.org/html/2605.21731#bib.bib35),[13](https://arxiv.org/html/2605.21731#bib.bib13)\], which leads the model to fail on structurally similar compounds with markedly different functional behaviour\[[33](https://arxiv.org/html/2605.21731#bib.bib33)\]\.

These findings expose a mismatch between the predictive success of a model and its ability to capture the structural properties of the problem\. Standard benchmark evaluation can establish that a model predicts well on a given distribution, but cannot characterize how model outputs reorganize when scientifically relevant input structure is perturbed\.

A principled response to this problem is to move from observational to interventional evaluation, where the question is not only whether a model predicts correctly, but how its predictions change under controlled perturbations of the input\. In causal terms, mechanistic claims concern responses to interventions rather than associations observed under a fixed distribution\[[25](https://arxiv.org/html/2605.21731#bib.bib25),[26](https://arxiv.org/html/2605.21731#bib.bib26)\]\. This perspective has informed several post\-hoc lines of work, including causal abstraction of internal representations\[[10](https://arxiv.org/html/2605.21731#bib.bib10),[11](https://arxiv.org/html/2605.21731#bib.bib11)\], model\-randomisation tests for explanation faithfulness\[[1](https://arxiv.org/html/2605.21731#bib.bib1)\], and perturbation probes of reasoning behaviour in language models\[[38](https://arxiv.org/html/2605.21731#bib.bib38),[40](https://arxiv.org/html/2605.21731#bib.bib40)\]\. In scientific prediction, domain knowledge provides a structural prior over inputs, enabling comparison of model responses to perturbations on prior\-selected versus non\-prior selected components\. This yields a post\-hoc, prior\-relative audit of whether the model response is organized with respect to meaningful input structure, as formalized in\[[32](https://arxiv.org/html/2605.21731#bib.bib32)\]\.

#### Our Contribution\.

In this paper, we introduceI\-SAFE\(Interventional Secure, Accurate, Fair and Explainable\), a post\-hoc distributional auditing framework for trained black\-box scientific predictors\. While existing intervention\-based audits rely on scalar summaries, theI\-SAFEframework evaluates the full output distribution induced by the model, capturing how output ranks reorder under prior\-guided perturbations of selected input components\. Our main theoretical contribution thus consists in defining a set of metrics that capture different levels of ranking coherence: the Quantile\-Based Metric for location\-level coherence, the Wasserstein Coherence Metric for ordinal coherence, and a translation\-invariant Wasserstein Metric for distributional shape coherence\. For each metric, the prior\-relative contrast compares the ranking\-coherence under outside\-prior controls and prior\-selected perturbations\. Positive contrasts indicate that perturbations of prior\-selected components induce more coherent output responses than their controls outside the prior\.

We then instantiateI\-SAFEon the Davis kinase benchmark\[[8](https://arxiv.org/html/2605.21731#bib.bib8)\], auditing three sequence\-based DTI models, DeepDTA\[[24](https://arxiv.org/html/2605.21731#bib.bib24)\], DeepConvDTI\[[18](https://arxiv.org/html/2605.21731#bib.bib18)\], and TAPB\[[20](https://arxiv.org/html/2605.21731#bib.bib20)\], using Kinase–Ligand Interaction Fingerprints and Structures \(KLIFS\) binding\-pocket annotations as an external structural prior\[[17](https://arxiv.org/html/2605.21731#bib.bib17),[16](https://arxiv.org/html/2605.21731#bib.bib16)\]\. While the audited models operate in a comparable predictive regime, we show that their interventional behaviour differs substantially\. In particular, TAPB is the only model for which KLIFS\-aligned pocket perturbations induce significantly more coherent quantile\-level and ordinal responses than non\-pocket controls\. Our results show that predictive performance and distributional coherence capture distinct aspects of scientific model behaviour, indicating thatI\-SAFEprovides a statistically grounded, prior\-relative test of how trained scientific predictors operate under structurally meaningful perturbations\. We stress that the audit does not establish causal validity of the model, nor does it require access to the model itself, it rather provides reusable, model\-agnostic metrics for evaluating black\-box model responses under structurally guided perturbations, applicable to any setting where inputs admit a structured decomposition and a domain prior is available\.

## 2Related Work

#### From explanation to interventional model analysis\.

Attribution and explanation methods provide tools for inspecting trained models\. Saliency maps, local surrogate models, Shapley\-based explanations, and integrated gradients\[[28](https://arxiv.org/html/2605.21731#bib.bib28),[22](https://arxiv.org/html/2605.21731#bib.bib22),[31](https://arxiv.org/html/2605.21731#bib.bib31),[29](https://arxiv.org/html/2605.21731#bib.bib29)\]identify input features that influence predictions, supporting transparency and debugging\[[9](https://arxiv.org/html/2605.21731#bib.bib9)\]\. However, such explanations are mostly associational: they indicate the influential features under the observed input distribution, without establishing whether predictions depend on scientifically relevant structure\. Their fragility has been documented empirically: Adebayo et al\.\[[1](https://arxiv.org/html/2605.21731#bib.bib1)\]show that saliency methods can be invariant to model parameter randomisation, calling into question their mechanistic faithfulness\. Intervention\-based approaches address this limitation by asking how the model changes under controlled modifications\. Causal abstraction and related methods\[[10](https://arxiv.org/html/2605.21731#bib.bib10),[11](https://arxiv.org/html/2605.21731#bib.bib11)\]compare neural representations with interpretable causal variables through interchange interventions on internal activations\. These methods provide a formal route to mechanistic comparison, but usually require access to internal representations and a target causal structure\. Input\-level perturbation methods avoid this requirement by probing the input–output matching induced by the model and have been applied to language models\[[38](https://arxiv.org/html/2605.21731#bib.bib38),[40](https://arxiv.org/html/2605.21731#bib.bib40)\]\. In scientific prediction, post\-hoc structural auditing under an external prior has been proposed to contrast mechanistic and spurious perturbations through scalar summaries\[[32](https://arxiv.org/html/2605.21731#bib.bib32)\]\.I\-SAFEretains the post\-hoc black\-box setting, but shifts the object of comparison from scalar aggregated values to the distributional changes in the model output\.

#### Shortcut learning and structural robustness\.

The need for auditing frameworks is reinforced by extensive evidence that predictive success can be driven by shortcuts or unstable correlations rather than task\-relevant structures\[[12](https://arxiv.org/html/2605.21731#bib.bib12),[15](https://arxiv.org/html/2605.21731#bib.bib15)\]\. In molecular prediction and DTI, this issue appears as ligand memorisation\[[23](https://arxiv.org/html/2605.21731#bib.bib23)\], target prior bias\[[20](https://arxiv.org/html/2605.21731#bib.bib20)\], benchmark artifacts related to leakage and split design\[[35](https://arxiv.org/html/2605.21731#bib.bib35),[13](https://arxiv.org/html/2605.21731#bib.bib13)\], and failures on activity cliffs, where structurally similar compounds have different functional effects\[[33](https://arxiv.org/html/2605.21731#bib.bib33)\]\. These findings motivate the search for predictors relying on stable relations\. Invariant causal prediction\[[26](https://arxiv.org/html/2605.21731#bib.bib26)\]and Invariant Risk Minimization\[[3](https://arxiv.org/html/2605.21731#bib.bib3)\], formalize stability across environments as a route toward more robust prediction\. The objective ofI\-SAFEis complementary\. Rather than modifying the learning procedure or requiring environment annotations, it asks whether a trained model responds coherently when input components identified by external scientific knowledge are perturbed\. This distinction is important in scientific AI, where models are used as black\-box predictors and retraining is impractical, unavailable, or insufficient to diagnose the predictor\. A parallel limitation affects intervention\-based evaluation itself, where the richness of the auditing framework is bounded by the statistical resolution at which model responses are characterized\.

#### Beyond aggregate performance\.

Single aggregate metrics are often too coarse to characterize behaviour relevant for downstream use\. In language model evaluation, DecodingTrust\[[36](https://arxiv.org/html/2605.21731#bib.bib36)\], HELM\[[19](https://arxiv.org/html/2605.21731#bib.bib19)\], and BIG\-bench\[[30](https://arxiv.org/html/2605.21731#bib.bib30)\]address this issue by organizing assessment across multiple capabilities, risks, and scenarios\. Likewise, in scientific model evaluation, held\-out performance establishes predictive adequacy on a benchmark distribution, but not how predictions reorganize under perturbations of scientifically meaningful input structure\. A similar limitation arises within intervention\-based evaluation itself\. Model responses are often reduced to average effects, scalar sensitivity scores, or a small set of moments\[[10](https://arxiv.org/html/2605.21731#bib.bib10),[40](https://arxiv.org/html/2605.21731#bib.bib40),[32](https://arxiv.org/html/2605.21731#bib.bib32)\]\. These summaries detect location\-level contrasts, but they might miss whether output distributions shift coherently, preserve ordinal structure, or change shape\. A standard approach to overcome single aggregate metrics relies on using two\-sample distributional comparison\[[7](https://arxiv.org/html/2605.21731#bib.bib7),[2](https://arxiv.org/html/2605.21731#bib.bib2)\]and optimal transport theory\[[34](https://arxiv.org/html/2605.21731#bib.bib34),[27](https://arxiv.org/html/2605.21731#bib.bib27)\], as these metrics capture the geometry of the underlying space\.I\-SAFEbrings this perspective to structural auditing by decomposing the distributional response to perturbations into three complementary axes: location, ordinal structure, and shape\. Each axis captures aspects of output reorganization invisible to scalar summaries\.

## 3The I\-SAFE Framework

In this section, we formalize theI\-SAFEframework as a post\-hoc auditing procedure for fixed black\-box predictors\.I\-SAFEleverages prior\-guided perturbations to induce paired profiles of raw model outputs and evaluates the coherence of their distributional reorganization after the intervention\. We consider a predictorfBB:𝒳→ℝf\_\{\\texttt\{BB\}\}:\\mathcal\{X\}\\to\\mathbb\{R\}accessed only through its input–output map, where𝒳=∏m=1M𝒳m\\mathcal\{X\}=\\prod\_\{m=1\}^\{M\}\\mathcal\{X\}\_\{m\}admits a decomposition inMMidentifiable components\. For an inputx=\(x\(1\),…,x\(M\)\)x=\(x^\{\(1\)\},\\ldots,x^\{\(M\)\}\), the components ofxxdefine the units on which interventions act\. Throughout the audit,fBB​\(x\)f\_\{\\texttt\{BB\}\}\(x\)denotes the raw model output on inputx∈𝒳x\\in\\mathcal\{X\}, before thresholding, calibration, or downstream decision rules\.

### 3\.1Problem formulation

Given a black\-box predictorfBBf\_\{\\texttt\{BB\}\}, the audit is performed on𝒜=\{xi\}i=1N⊂𝒳\\mathcal\{A\}=\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}\\subset\\mathcal\{X\}which is disjoint from the data used to trainfBBf\_\{\\texttt\{BB\}\}\. The set𝒜\\mathcal\{A\}contains the elements of𝒳\\mathcal\{X\}on which all perturbations are applied and all output distributions are compared\. To determine how to perform a perturbation, we have access to a structural prior over𝒳\\mathcal\{X\}, derived from domain knowledge that does not depend onfBBf\_\{\\texttt\{BB\}\}\.

###### Definition 1\(Structural prior\)\.

We say that a map𝒫:𝒳→2\[M\]\\mathcal\{P\}:\\mathcal\{X\}\\to 2^\{\[M\]\}is a*structural prior*if it satisfies∅⊊𝒫​\(x\)⊊\[M\]\\emptyset\\subsetneq\\mathcal\{P\}\(x\)\\subsetneq\[M\]for allx∈𝒳x\\in\\mathcal\{X\}\.

For each inputxx, the set𝒫​\(x\)⊂\[M\]\\mathcal\{P\}\(x\)\\subset\[M\]indexes the components that domain knowledge regards as relevant for the prediction task\. The strict inclusions ensure we avoid cases in which no components are relevant \(i\.e\.𝒫​\(x\)=∅\\mathcal\{P\}\(x\)=\\emptyset\) and the case in which all components are relevant \(i\.e\.𝒫​\(x\)=\[M\]\\mathcal\{P\}\(x\)\\\!=\\\!\[M\]\)\.

### 3\.2Prior\-relative intervention design

Given𝒜\\mathcal\{A\}and a structural prior𝒫\\mathcal\{P\}, the audit compares two classes of interventions defined relative to𝒫\\mathcal\{P\}: the interventions acting on𝒫​\(x\)\\mathcal\{P\}\(x\), and the interventions acting on the complement of𝒫​\(x\)\\mathcal\{P\}\(x\)\. The comparison is thus prior\-dependent and evaluates whether the model responds differently when the same perturbation is applied to prior\-selected components rather than to components outside the prior\.

###### Definition 2\(Mechanistic and spurious perturbations\)\.

A perturbationφ𝒫:𝒳→𝒳\\varphi\_\{\\mathcal\{P\}\}:\\mathcal\{X\}\\to\\mathcal\{X\}is*mechanistic*if it acts only on components selected by the prior, that is

\(φ𝒫​\(x\)\)i=xi∀i∉𝒫​\(x\)\.\\bigl\(\\varphi\_\{\\mathcal\{P\}\}\(x\)\\bigr\)\_\{i\}=x\_\{i\}\\qquad\\forall\\,i\\notin\\mathcal\{P\}\(x\)\.\(1\)Likewise,φ\\varphiis said to be*spurious*if it leaves all prior\-selected components unchanged, that is

\(φ𝒫​\(x\)\)i=xi∀i∈𝒫​\(x\)\.\\bigl\(\\varphi\_\{\\mathcal\{P\}\}\(x\)\\bigr\)\_\{i\}=x\_\{i\}\\qquad\\forall\\,i\\in\\mathcal\{P\}\(x\)\.\(2\)

The term*spurious*is used only relative to the specified prior𝒫​\(x\)\\mathcal\{P\}\(x\)\. Notice that this does not imply that such components are irrelevant in an absolute scientific sense\. To make the mechanistic\-spurious comparison interpretable, we distinguish the perturbation rule from the intervention support\.

###### Definition 3\(Perturbation operator\)\.

A*perturbation operator*is a mapΦ:𝒳×2\[M\]→𝒳\\Phi:\\mathcal\{X\}\\times 2^\{\[M\]\}\\to\\mathcal\{X\}such that, for everyx∈𝒳x\\in\\mathcal\{X\}andS⊆\[M\]S\\subseteq\[M\],Φ​\(x,S\)\\Phi\(x,S\)differs fromxxonly on components indexed bySS\.

Givenx∈𝒳x\\in\\mathcal\{X\}, the mechanistic support ofxxis𝒫​\(x\)\\mathcal\{P\}\(x\)\. A spurious support is a setSspur​\(x\)⊆\[M\]∖𝒫​\(x\)S^\{\\mathrm\{spur\}\}\(x\)\\subseteq\[M\]\\setminus\\mathcal\{P\}\(x\)\. Both supports are modified using the same operatorΦ\\Phi\. Thus, the two interventions share the perturbation rule and differ only in the position of their support relative to the prior\.

###### Definition 4\(Matched perturbation pair\)\.

A pair\(φmech,φspur\)\(\\varphi^\{\\mathrm\{mech\}\},\\varphi^\{\\mathrm\{spur\}\}\)is a*matched perturbation pair*on𝒜\\mathcal\{A\}if there exists a perturbation operatorΦ\\Phisuch that, for everyx∈𝒜x\\in\\mathcal\{A\},

φmech​\(x\)=Φ​\(x,𝒫​\(x\)\),φspur​\(x\)=Φ​\(x,Sspur​\(x\)\),\\varphi^\{\\mathrm\{mech\}\}\(x\)=\\Phi\(x,\\mathcal\{P\}\(x\)\),\\qquad\\varphi^\{\\mathrm\{spur\}\}\(x\)=\\Phi\(x,S^\{\\mathrm\{spur\}\}\(x\)\),\(3\)whereSspur​\(x\)⊆\[M\]∖𝒫​\(x\)S^\{\\mathrm\{spur\}\}\(x\)\\subseteq\[M\]\\setminus\\mathcal\{P\}\(x\)and\|Sspur​\(x\)\|=\|𝒫​\(x\)\|\|S^\{\\mathrm\{spur\}\}\(x\)\|=\|\\mathcal\{P\}\(x\)\|\.

A matched perturbation pair controls the two intervention degrees of the method: the perturbation rule and the number of perturbed components\. Acting on the same number of entries enables the isolation of the prior alignment role in the model response, according to the prior𝒫\\mathcal\{P\}\.

###### Definition 5\(Interventional response profile\)\.

Letφ:𝒳→𝒳\\varphi:\\mathcal\{X\}\\to\\mathcal\{X\}be either a mechanistic or a spurious perturbation, and let𝒜=\{xi\}i=1N⊂𝒳\\mathcal\{A\}=\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}\\subset\\mathcal\{X\}be the auditing set\. Then, we define

V𝒜=\(fBB​\(xi\)\)i=1NandVφ​\(𝒜\)=\(fBB​\(φ​\(xi\)\)\)i=1N\.V\_\{\\mathcal\{A\}\}=\\bigl\(f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\\bigr\)\_\{i=1\}^\{N\}\\qquad\\qquad\\qquad\\text\{and\}\\qquad\\qquad\\qquad V\_\{\\varphi\(\\mathcal\{A\}\)\}=\\bigl\(f\_\{\\texttt\{BB\}\}\(\\varphi\(x\_\{i\}\)\)\\bigr\)\_\{i=1\}^\{N\}\.\(4\)The pairℛfBB​\(φ;𝒜\)=\(V𝒜,Vφ​\(𝒜\)\)\\mathcal\{R\}\_\{f\_\{\\texttt\{BB\}\}\}\(\\varphi;\\mathcal\{A\}\)=\(V\_\{\\mathcal\{A\}\},V\_\{\\varphi\(\\mathcal\{A\}\)\}\)is the*interventional response profile*offBBf\_\{\\texttt\{BB\}\}underφ\\varphion𝒜\\mathcal\{A\}\.

The response profile contains the raw model outputs before and after intervention on the same auditing population, preserving both the empirical output distributions and the input\-wise pairing induced by the perturbation\. For each matched pair,I\-SAFEcompares the mechanistic and spurious response profiles through the coherence metrics introduced in the following section\.

### 3\.3The I\-SAFE Auditing Metrics

In this section, we introduce a set of three metrics to quantify how the raw output distribution offBBf\_\{\\texttt\{BB\}\}changes in terms of relative ranking coherence under a perturbationφ\\varphi\. In line with Secure, Accurate, Fair and Explainable \(SAFE\) metrics\[[4](https://arxiv.org/html/2605.21731#bib.bib4),[6](https://arxiv.org/html/2605.21731#bib.bib6)\], we define our metrics to represent the percentage of coherence explainable by input differences\. In other words, the metrics should attain values close to11when the coherence is small and values close to0when the coherence is high\.

First, we assess the location\-level distributional change by comparing empirical quantiles of the original and perturbed output profiles,i\.e\.,\{fBB​\(x\)\}x∈𝒜\\\{f\_\{\\texttt\{BB\}\}\(x\)\\\}\_\{x\\in\\mathcal\{A\}\}and\{fBB​\(φ​\(x\)\)\}x∈𝒜\\\{f\_\{\\texttt\{BB\}\}\(\\varphi\(x\)\)\\\}\_\{x\\in\\mathcal\{A\}\}\. Given a vector of target quantile levelsq=\(q1,…,qK\)q=\(q\_\{1\},\\ldots,q\_\{K\}\), we denote byqk\(𝒜\)q^\{\(\\mathcal\{A\}\)\}\_\{k\}the empiricalqkq\_\{k\}\-quantile of\{fBB​\(xi\)\}i=1N\\\{f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\\\}\_\{i=1\}^\{N\}, and byqk\(φ​\(𝒜\)\)q^\{\(\\varphi\(\\mathcal\{A\}\)\)\}\_\{k\}ithe empiricalqkq\_\{k\}\-quantile of\{fBB​\(φ​\(xi\)\)\}i=1N\\\{f\_\{\\texttt\{BB\}\}\(\\varphi\(x\_\{i\}\)\)\\\}\_\{i=1\}^\{N\}\. Finally, we setq\(𝒜\)=\(q1\(𝒜\),…,qK\(𝒜\)\)q^\{\(\\mathcal\{A\}\)\}=\(q^\{\(\\mathcal\{A\}\)\}\_\{1\},\\ldots,q^\{\(\\mathcal\{A\}\)\}\_\{K\}\)and byq\(φ​\(𝒜\)\)=\(q1\(φ​\(𝒜\)\),…,qK\(φ​\(𝒜\)\)\)q^\{\(\\varphi\(\\mathcal\{A\}\)\)\}=\(q^\{\(\\varphi\(\\mathcal\{A\}\)\)\}\_\{1\},\\ldots,q^\{\(\\varphi\(\\mathcal\{A\}\)\)\}\_\{K\}\)\. TheQuantile\-based Metric\(QBM\) compares the displacement of these representative locations with the total Mean Square Error of the values\.

###### Definition 6\.

Given an auditing set𝒜\\mathcal\{A\}, a perturbationφ\\varphi, a black\-box modelfBBf\_\{\\texttt\{BB\}\}, and a quantile vectorq=\(q1,…,qK\)q=\(q\_\{1\},\\ldots,q\_\{K\}\), the Quantile\-Based Metric \(QBM\) is defined as

Q​B​M​\(φ;fBB\)=\(1−1K​∑i=1K\|qiφ​\(𝒜\)−qi\(𝒜\)\|21\|𝒜\|​∑i=1\|𝒜\|\|fBB​\(φ​\(xi\)\)−fBB​\(xi\)\|2\)\+,QBM\(\\varphi;f\_\{\\texttt\{BB\}\}\)=\\Bigg\(1\-\\sqrt\{\\frac\{\\frac\{1\}\{K\}\\sum\_\{i=1\}^\{K\}\|q^\{\\varphi\(\\mathcal\{A\}\)\}\_\{i\}\-q^\{\(\\mathcal\{A\}\)\}\_\{i\}\|^\{2\}\}\{\\frac\{1\}\{\|\\mathcal\{A\}\|\}\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\|f\_\{\\texttt\{BB\}\}\(\\varphi\(x\_\{i\}\)\)\-f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\|^\{2\}\}\}\\Bigg\)\_\{\+\},\(5\)

where\(x\)\+\(x\)\_\{\+\}denotes the positive part ofxx\.

By definition, the lower the values of QBM, the larger the shift in the quantile displacement\. Conversely, larger values of QBM indicate that the quantile structure remains mostly unaffected by the perturbation despite potentially large pointwise changes\.

###### Proposition 1\.

For every auditing set𝒜\\mathcal\{A\}, perturbationφ\\varphi, black\-box modelfBBf\_\{\\texttt\{BB\}\}, and quantile vectorqqit holdsQ​B​M​\(φ;fBB\)∈\[0,1\]QBM\(\\varphi;f\_\{\\texttt\{BB\}\}\)\\in\[0,1\]\.

Notice that QBM depends on both the output distributions and the number of quantiles we consider\. Coarser quantile grids summarize location\-level changes, whereas finer grids retain more information about the empirical distributions of\{fBB​\(x\)\}x∈𝒜\\\{f\_\{\\texttt\{BB\}\}\(x\)\\\}\_\{x\\in\\mathcal\{A\}\}and\{fBB​\(φ​\(x\)\)\}x∈𝒜\\\{f\_\{\\texttt\{BB\}\}\(\\varphi\(x\)\)\\\}\_\{x\\in\\mathcal\{A\}\}\. When the grid is refined to the empirical order statistics, the QBM becomes a normalized optimal\-transport metric, which we nameWasserstein Coherence Metric\(WCM\) and does not depend on the quantile choice\.

###### Definition 7\.

Given an auditing set𝒜\\mathcal\{A\}, a perturbationφ\\varphi, and a black\-box modelfBBf\_\{\\texttt\{BB\}\}, we define the Wasserstein Coherence Metric \(WCM\) as follows

W​C​M​\(φ;fBB\):=1−minπ∈Πn​\(𝒜\)​∑x∈𝒜\|fBB​\(φ​\(π​\(x\)\)\)−fBB​\(x\)\|2∑x∈𝒜\|fBB​\(φ​\(x\)\)−fBB​\(x\)\|2,WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\):=1\-\\sqrt\{\\frac\{\\min\_\{\\pi\\in\\Pi\_\{n\}\(\\mathcal\{A\}\)\}\\sum\_\{x\\in\\mathcal\{A\}\}\|f\_\{\\texttt\{BB\}\}\(\\varphi\(\\pi\(x\)\)\)\-f\_\{\\texttt\{BB\}\}\(x\)\|^\{2\}\}\{\\sum\_\{x\\in\\mathcal\{A\}\}\|f\_\{\\texttt\{BB\}\}\(\\varphi\(x\)\)\-f\_\{\\texttt\{BB\}\}\(x\)\|^\{2\}\}\},\(6\)whereΠ​\(𝒜\)\\Pi\(\\mathcal\{A\}\)denotes the set of all permutations of the auditing inputs in𝒜\\mathcal\{A\}\.

The denominator of \([6](https://arxiv.org/html/2605.21731#S3.E6)\) measures the paired output displacement under the natural correspondencex↦φ​\(x\)x\\mapsto\\varphi\(x\), while the numerator measures the smallest displacement attainable after optimally reordering the perturbed outputs\. WCM is therefore small when the natural input\-wise pairing is close to this optimal reordering, and large when the perturbed outputs can be substantially better matched only after reordering\.

Definition[7](https://arxiv.org/html/2605.21731#Thmdefinition7)yields two useful invariance properties\(i\)WCM is normalized and\(ii\)WCM is invariant under positive rescaling of the model output\.These properties make WCM an easy to interpret and scale\-free measure of ordinal coherence\.

###### Proposition 2\.

For every auditing set𝒜\\mathcal\{A\}, perturbationφ\\varphi, and black\-box modelfBBf\_\{\\texttt\{BB\}\}, it holdsW​C​M​\(φ;fBB\)∈\[0,1\]WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)\\in\[0,1\]\. Moreover,\(i\)W​C​M​\(φ;fBB\)=0WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)=0if and only if the pairing\{\(fBB​\(x\),fBB​\(φ​\(x\)\)\)\}x∈𝒜\\\{\(f\_\{\\texttt\{BB\}\}\(x\),f\_\{\\texttt\{BB\}\}\(\\varphi\(x\)\)\)\\\}\_\{x\\in\\mathcal\{A\}\}is a coupling between the perturbed and original output that maximizes the covariance; and\(ii\)W​C​M​\(φ;fBB\)=1WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)=1if and only if there exists a permutation, namelyπ\\pi, such thatπ\\piis not the identity and for which it holdsfBB​\(x\)=fBB​\(φ​\(π​\(x\)\)\)f\_\{\\texttt\{BB\}\}\(x\)=f\_\{\\texttt\{BB\}\}\(\\varphi\(\\pi\(x\)\)\)for everyx∈𝒜x\\in\\mathcal\{A\}\.

Lastly, theW​C​MWCMis invariant under change of scales, meaning that its value does not change if we alter unit of measure of the output layer, i\.e\.W​C​M​\(φ;fBB\)=W​C​M​\(φ;λ​fBB\)WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)=WCM\(\\varphi;\\lambda f\_\{\\texttt\{BB\}\}\)for anyλ\>0\\lambda\>0\.

From Proposition[2](https://arxiv.org/html/2605.21731#Thmproposition2), we infer that the WCM measures the covariance alignment between the natural coupling induced by the data and the perturbation, that is\(fBB​\(x\),fBB​\(φ​\(x\)\)\)\(f\_\{\\texttt\{BB\}\}\(x\),f\_\{\\texttt\{BB\}\}\(\\varphi\(x\)\)\), and the one that maximises the covariance\. This induces a natural connection between the WCM and the Wasserstein Distance between the empirical distribution induced by the black\-box modelfBBf\_\{\\texttt\{BB\}\}output before and after the perturbation\. Consequently, we can express the minimum in the numerator of \([6](https://arxiv.org/html/2605.21731#S3.E6)\) as the euclidean norm of the vectors\(fBB​\(φ​\(x\)\)\)x∈𝒜\(f\_\{\\texttt\{BB\}\}\(\\varphi\(x\)\)\)\_\{x\\in\\mathcal\{A\}\}and\(fBB​\(x\)\)x∈𝒜\(f\_\{\\texttt\{BB\}\}\(x\)\)\_\{x\\in\\mathcal\{A\}\}rearranged in non\-decreasing order\.

###### Proposition 3\.

Given a perturbationφ\\varphi, a black\-box modelfBBf\_\{\\texttt\{BB\}\}, and an auditing set𝒜\\mathcal\{A\}, then

W​C​M​\(φ;fBB\):=1−∑i=1\|𝒜\|\(\(V𝒜\)ri−\(Vφ​\(𝒜\)\)ri\(φ\)\)2∑i=1\|𝒜\|\(\(V𝒜\)i−\(Vφ​\(𝒜\)\)i\)2WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\):=1\-\\sqrt\{\\frac\{\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\\big\(\(V\_\{\\mathcal\{A\}\}\)\_\{r\_\{i\}\}\-\(V\_\{\\varphi\(\\mathcal\{A\}\)\}\)\_\{r\_\{i\}^\{\(\\varphi\)\}\}\\big\)^\{2\}\}\{\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\\big\(\(V\_\{\\mathcal\{A\}\}\)\_\{i\}\-\(V\_\{\\varphi\(\\mathcal\{A\}\)\}\)\_\{i\}\\big\)^\{2\}\}\}\(7\)where\(i\)V𝒜V\_\{\\mathcal\{A\}\}andVφ​\(𝒜\)V\_\{\\varphi\(\\mathcal\{A\}\)\}are defined as in \([4](https://arxiv.org/html/2605.21731#S3.E4)\) and\(ii\)rir\_\{i\}is a monotone non\-decreasing reordering ofV𝒜V\_\{\\mathcal\{A\}\}andri\(φ\)r\_\{i\}^\{\(\\varphi\)\}is a non\-decreasing reordering ofVϕ​\(𝒜\)V\_\{\\phi\(\\mathcal\{A\}\)\}\.

Lastly, we introduce the third I\-SAFE metric, the Translation\-Invariant WCM, which measures only differences in distributional shape\.

###### Definition 8\.

Given an auditing set𝒜\\mathcal\{A\}, a perturbationφ\\varphi, and a black\-box modelfBBf\_\{\\texttt\{BB\}\}, we define the translation invariant Wasserstein Coherence Metric \(TI\-WCM\) as follows

T​I−W​C​M​\(φ,fBB\):=1−W2​\(μVφ​\(𝒜\),μV𝒜\)2−\(m𝒜−mφ​\(𝒜\)\)2ℓ2​\(Vφ​\(𝒜\),V𝒜\),TI\-WCM\(\\varphi,f\_\{\\texttt\{BB\}\}\):=1\-\\frac\{\\sqrt\{W\_\{2\}\(\\mu\_\{V\_\{\\varphi\(\\mathcal\{A\}\)\}\},\\mu\_\{V\_\{\\mathcal\{A\}\}\}\)^\{2\}\-\(m\_\{\\mathcal\{A\}\}\-m\_\{\\varphi\(\\mathcal\{A\}\)\}\)^\{2\}\}\}\{\\ell\_\{2\}\(V\_\{\\varphi\(\\mathcal\{A\}\)\},V\_\{\\mathcal\{A\}\}\)\},\(8\)wherem𝒜m\_\{\\mathcal\{A\}\}is the mean value of\{fBB​\(x\)\}x∈𝒜\\\{f\_\{\\texttt\{BB\}\}\(x\)\\\}\_\{x\\in\\mathcal\{A\}\}andmφ​\(𝒜\)m\_\{\\varphi\(\\mathcal\{A\}\)\}is the mean value of\{fBB\(φ\(x\)\}x∈𝒜\\\{f\_\{\\texttt\{BB\}\}\(\\varphi\(x\)\\\}\_\{x\\in\\mathcal\{A\}\}\.

The TI\-WCM is invariant under model bias, meaning that if we add a constant value to the output of the modelfBBf\_\{\\texttt\{BB\}\}, the TI\-WCM does not change\. This makes the TI\-WCM a stricter metric than the WCM introduced in Definition[7](https://arxiv.org/html/2605.21731#Thmdefinition7), as shown in\[[5](https://arxiv.org/html/2605.21731#bib.bib5)\]\. Lastly, we notice that the TI\-WCM is still a normalized value between0and11that measures the coherence between the ranking of the two output sets\. The full discussion is deferred to Appendix[A\.2](https://arxiv.org/html/2605.21731#A1.SS2)\.

## 4Experiments

We evaluateI\-SAFEon DTI prediction, a scientific task in which binding\-pocket annotations provide a natural structural prior over target residues\. In kinase–inhibitor prediction, these annotations identify protein regions involved in molecular recognition, while sequence\-based DTI models may also exploit global sequence statistics, target bias, or dataset\-specific regularities\[[20](https://arxiv.org/html/2605.21731#bib.bib20),[35](https://arxiv.org/html/2605.21731#bib.bib35),[13](https://arxiv.org/html/2605.21731#bib.bib13)\]\. This makes DTI a suitable setting for testing whether models with comparable predictive performance exhibit different prior\-relative interventional behaviour\. The goal is not to rank models by accuracy, but to evaluate how raw output distributions reorganize under matched perturbations of mechanistic and spurious input components\.

### 4\.1Task and framework instantiation

Each input is a drug–target pairx=\(d,t\)x=\(d,t\), whereddis a drug represented by a Simplified Molecular Input Line Entry System \(SMILES\) string andt=\(t1,…,tM\)t=\(t\_\{1\},\\ldots,t\_\{M\}\)is a protein target represented as an amino\-acid sequence of lengthM=\|t\|M=\|t\|\. The audited model returns a raw scorefBB​\(d,t\)∈ℝf\_\{\\texttt\{BB\}\}\(d,t\)\\in\\mathbb\{R\}, used directly for theI\-SAFEmetrics before any thresholding, calibration, or downstream decision rule\. Labelsy∈\{0,1\}y\\in\\\{0,1\\\}define the underlying interaction task and are used only for predictive\-regime verification; theI\-SAFEmetrics are computed exclusively from raw model scores, without label information at any stage of the audit\. In this instantiation, the identifiable input components are the residues of the target sequence, so thatx\(m\)=tmx^\{\(m\)\}=t\_\{m\}form=1,…,Mm=1,\\ldots,M\. Interventions act exclusively on the protein targettt, while the drug representationddis held fixed\. Thus, the audit characterizes target\-side interventional behaviour under a fixed drug context\.

#### Structural prior\.

The structural prior𝒫​\(t\)\\mathcal\{P\}\(t\)is derived from KLIFS\[[17](https://arxiv.org/html/2605.21731#bib.bib17),[16](https://arxiv.org/html/2605.21731#bib.bib16)\], a curated resource of kinase–ligand structural annotations\. For each kinase target,𝒫​\(t\)⊆\{1,…,M\}\\mathcal\{P\}\(t\)\\subseteq\\\{1,\\ldots,M\\\}denotes the residue indices corresponding to the annotated binding pocket\. This prior is specified independently of the audited models and is used only to define prior\-aligned intervention scopes\. In the auditing set, the median pocket size is 85 residues\.

### 4\.2Dataset and auditing set

Experiments are conducted on the Davis benchmark\[[8](https://arxiv.org/html/2605.21731#bib.bib8)\], a standard kinase inhibitor dataset for DTI prediction\. We use the train/validation/test splits released with the TAPB benchmark\[[20](https://arxiv.org/html/2605.21731#bib.bib20)\]and apply them consistently across all audited models\. These splits are used for model training and predictive\-regime verification, while interventional auditing is performed post\-hoc on the structurally valid subset of the test split\. The auditing set is obtained by deterministic filtering\. We retain targets with KLIFS binding\-pocket annotations and for which annotated pocket residues can be unambiguously mapped to the amino\-acid sequence\. This yields208208kinase targets from the379379test targets and3,0443\{,\}044drug–target pairs, each with a unique residue\-level interventional representation\. The median KLIFS pocket size is8585residues, and exact cardinality matching is achieved for all audited interventions\. Full structural coverage statistics are reported in Appendix Table[2](https://arxiv.org/html/2605.21731#A1.T2)\.

For each audited pair\(d,t\)∈𝒜\(d,t\)\\in\\mathcal\{A\}and perturbation operatorφ∈Φ\\varphi\\in\\Phi, we construct a matched pair of intervention scopes\. The prior\-aligned scope is drawn from the KLIFS binding\-pocket residues𝒫​\(t\)\\mathcal\{P\}\(t\), whereas the prior\-misaligned scope is sampled from\{1,…,M\}∖𝒫​\(t\)\\\{1,\\ldots,M\\\}\\setminus\\mathcal\{P\}\(t\)with identical cardinality\. The same operator is applied to both scopes, ensuring that the resulting intervention classes differ only in their alignment with the structural prior while preserving the perturbation rule and intervention size\. In our experiments,Φ\\Phicontains masking, which replaces selected residues with a dedicated mask token, and physicochemically constrained substitution, which replaces each selected residue with an amino acid from the same biochemical class\.

### 4\.3Audited models and predictive regime

We audit three sequence\-based DTI models that differ in their target\-side inductive biases\.DeepDTA\[[24](https://arxiv.org/html/2605.21731#bib.bib24)\]provides a convolutional baseline over SMILES and amino\-acid sequences\.DeepConvDTI\[[18](https://arxiv.org/html/2605.21731#bib.bib18)\]emphasizes protein\-sequence representation through target\-side convolutions and global pooling\.TAPB\[[20](https://arxiv.org/html/2605.21731#bib.bib20)\]incorporates target\-aware attention together with an interventional debiasing objective for target\-prior bias\. All models are audited post\-hoc from frozen checkpoints trained with their original protocols, without model\-specific tuning or adjustment during theI\-SAFEaudit\.

#### Predictive regime verification\.

Before auditing, we use the area under the receiver operating characteristic curve \(AUROC\) on the Davis auditing subset𝒜\\mathcal\{A\}only to verify that the models operate in a comparable predictive regime\. Appendix Table[3](https://arxiv.org/html/2605.21731#A1.T3)reports mean AUROC across five training seeds, with 95 % confidence intervals obtained by non\-parametric bootstrap \(B=100B=100\) over𝒜\\mathcal\{A\}: DeepConvDTI0\.8760\.876\[0\.875,0\.878\]\[0\.875,0\.878\], TAPB0\.8820\.882\[0\.851,0\.899\]\[0\.851,0\.899\], and DeepDTA0\.9070\.907\[0\.902,0\.913\]\[0\.902,0\.913\]\. The largest inter\-model difference is approximately three AUROC points\. We therefore treat subsequentI\-SAFEcontrasts as comparisons of interventional response within a common predictive regime, rather than as differences in baseline predictive performance\.

### 4\.4Interventional response structure

We analyze the interventional response profiles induced by mechanistic and spurious perturbations\. For each metricM∈\{QBM,WCM,TI​\-​WCM\}M\\in\\\{\\mathrm\{QBM\},\\mathrm\{WCM\},\\mathrm\{TI\\text\{\-\}WCM\}\\\}, we report both class\-specific values and the prior\-relative contrastΔ​M=Mspurious−Mmechanistic\\Delta M=M\_\{\\mathrm\{spurious\}\}\-M\_\{\\mathrm\{mechanistic\}\}\. Since lower values indicate more coherent output responses, positiveΔ​M\\Delta Mdenotes greater coherence under mechanistic perturbations than under matched spurious controls; negative values denote the reverse\. QBM is computed usingq=\(0\.25,0\.50,0\.75\)q=\(0\.25,0\.50,0\.75\), the lower quartile, median, and upper quartile of the output distribution; sensitivity to the quantile grid is reported in Appendix Table[5](https://arxiv.org/html/2605.21731#A1.T5)\.

#### Absolute coherence by perturbation class\.

Table[1](https://arxiv.org/html/2605.21731#S4.T1)reports theI\-SAFEmetrics separately for mechanistic and spurious perturbations\. These values are diagnostic rather than ranking metrics\. They identify which aspect of the raw output profile is reorganized by a given intervention, separating changes in distributional location, output ordering, and residual shape\.

Under mechanistic perturbations,TAPBshows the strongest coherence on the two axes that retain location information, with the lowest QBM \(0\.5150\.515,\[0\.476,0\.553\]\[0\.476,0\.553\]\) and WCM \(0\.4430\.443,\[0\.419,0\.467\]\[0\.419,0\.467\]\)\. Thus, perturbing KLIFS pocket residues induces an organized change inTAPBraw scores, both in distributional location and output ranking\.DeepDTAshows the least coherent response under the same perturbations, with the highest QBM \(0\.8420\.842,\[0\.784,0\.897\]\[0\.784,0\.897\]\) and WCM \(0\.7680\.768,\[0\.735,0\.799\]\[0\.735,0\.799\]\), indicating that predictive adequacy does not imply coherent behaviour under structurally targeted intervention\.DeepConvDTIlies between these regimes, with high QBM \(0\.7370\.737,\[0\.678,0\.799\]\[0\.678,0\.799\]\) but lower WCM \(0\.5120\.512,\[0\.488,0\.537\]\[0\.488,0\.537\]\), suggesting partial preservation of score ordering without an equally coherent quantile\-level shift\.

The comparison with spurious perturbations already reveals distinct response profiles\. ForTAPB, QBM and WCM are lower under mechanistic than spurious perturbations, whereasDeepConvDTIshows the reverse pattern on WCM andDeepDTAchanges little across classes\. TI\-WCM further qualifies the interpretation:TAPBhas high mechanistic TI\-WCM \(0\.7850\.785,\[0\.751,0\.812\]\[0\.751,0\.812\]\) despite favourable QBM and WCM, indicating that its response is not primarily a residual shape\-preserving effect after removing the mean shift\. The absolute metrics therefore localize the main signal to location and ranking\. The contrast analysis below tests whether this structure is selective for KLIFS\-aligned pocket perturbations relative to non\-pocket controls\.

Table 1:I\-SAFEabsolute metric values under mechanistic and spurious perturbations \(operator=all, 5 seeds, 95 % CI\)\.
#### Prior\-relative coherence contrasts\.

Figure[1](https://arxiv.org/html/2605.21731#S4.F1)\(and Table[4](https://arxiv.org/html/2605.21731#A1.T4)in Appendix\) reports the prior\-relative contrasts\. These contrasts test whether perturbing KLIFS pocket residues induces a more coherent distributional response than applying the same operator, with the same cardinality, outside the pocket\.

TAPBshows the clearest prior\-relative pattern on the two metrics that retain location information\. It achievesΔ​QBM=\+0\.093\\Delta\\mathrm\{QBM\}=\+0\.093\(\[0\.039,0\.147\]\[0\.039,0\.147\]\) andΔ​WCM=\+0\.063\\Delta\\mathrm\{WCM\}=\+0\.063\(\[0\.028,0\.099\]\[0\.028,0\.099\]\)\. This indicates that its coherent response is selectively associated with KLIFS\-aligned pocket perturbations rather than with generic target\-sequence perturbation, and is expressed both through quantile\-level displacement and output ordering\. The other architectures follow a different pattern\. ForDeepConvDTI,Δ​QBM=−0\.011\\Delta\\mathrm\{QBM\}=\-0\.011\(\[−0\.090,0\.068\]\[\-0\.090,0\.068\]\) andΔ​WCM=−0\.046\\Delta\\mathrm\{WCM\}=\-0\.046\(\[−0\.081,−0\.012\]\[\-0\.081,\-0\.012\]\), indicating that the ordinal component is more coherent under non\-pocket controls than under pocket perturbations\. ForDeepDTA,Δ​QBM=−0\.021\\Delta\\mathrm\{QBM\}=\-0\.021\(\[−0\.099,0\.057\]\[\-0\.099,0\.057\]\) andΔ​WCM=−0\.013\\Delta\\mathrm\{WCM\}=\-0\.013\(\[−0\.057,0\.031\]\[\-0\.057,0\.031\]\), showing no comparable prior\-aligned coherence advantage\.

The translation\-invariant component further refines the interpretation\.Δ​TI​\-​WCM\\Delta\\mathrm\{TI\\text\{\-\}WCM\}is negative across all models:−0\.043\-0\.043\(\[−0\.078,−0\.009\]\[\-0\.078,\-0\.009\]\) forDeepConvDTI,−0\.014\-0\.014\(\[−0\.054,0\.026\]\[\-0\.054,0\.026\]\) forDeepDTA, and−0\.038\-0\.038\(\[−0\.082,0\.006\]\[\-0\.082,0\.006\]\) forTAPB\. Hence, the positive TAPB signal is localized to location and ranking rather than to residual shape preservation after removing the mean shift\. Overall, among models with comparable AUROC, onlyTAPBexhibits a prior\-selective distributional response to the binding\-pocket prior\.

![Refer to caption](https://arxiv.org/html/2605.21731v1/x1.png)Figure 1:I\-SAFEprior\-relative coherence contrasts on the Davis benchmark:Δ\\DeltaQBM \(a\),Δ\\DeltaWCM \(b\), andΔ\\DeltaTI\-WCM \(c\), computed as spurious minus mechanistic coherence\. The dashed line marks no differential coherence; positive values indicate greater coherence under mechanistic perturbations\. Error bars denote 95 % confidence intervals across five seeds\.

## 5Discussion

I\-SAFEintroduces a distributional layer for post\-hoc structural auditing of scientific predictors\. It moves the audit beyond asking whether a model is sensitive to prior\-aligned perturbations, toward asking how its raw output distribution reorganizes under intervention\. This distinction matters because benchmark performance does not establish that model behaviour is organized around the structures that domain knowledge identifies as relevant\.

The Davis DTI case study illustrates this separation\. The audited models operate in a comparable predictive regime, yet exhibit different interventional response profiles\.TAPBis the only architecture with positive contrasts on both the quantile\-level and ordinal axes, while the translation\-invariant component does not provide a positive differential signal\. The signal detected byI\-SAFEis therefore not a generic distributional effect\. It is expressed through location\-level and ranking\-level organization of raw output scores under KLIFS binding\-pocket perturbations\. These results separate predictive performance from distributional coherence as distinct dimensions of model behaviour\.

#### Prior\-relative interpretation

The contrast metrics are the prior\-relative component of the audit\. They compare the coherence induced by perturbations of prior\-selected components with the coherence induced by controls outside the prior under the same intervention design\. A positive contrast is therefore a behavioural statement about a fixed trained model relative to an independently specified structural prior\. It does not imply that the prior is complete or that the model has recovered the causal mechanism of the data\-generating process\. It shows that the model’s output distribution is more coherently organized under perturbations of the prior\-selected region\.

#### Beyond scalar auditing\.

Scalar auditing\[[32](https://arxiv.org/html/2605.21731#bib.bib32)\]summarizes relative interventional sensitivity as a single magnitude\-based contrast\.I\-SAFEretains more of the response structure\. QBM measures coherent movement of representative output locations, WCM measures preservation of output ordering, and TI\-WCM separates residual shape coherence from global translation\. In our experiments,TAPB’s positive signal appears on QBM and WCM, but not on TI\-WCM, indicating that its prior\-relative behaviour is carried by location and ranking rather than by shape preservation\. This conclusion cannot be obtained from accuracy or scalar sensitivity alone\.

#### Assumptions, scope, and generalization\.

The structural prior defines the audited contrast\. Here, KLIFS binding\-pocket annotations provide an external, biologically grounded prior over kinase target residues, specifying a meaningful axis of comparison rather than a ground\-truth causal mechanism\. Accordingly, the empirical claims are specific to target\-side interventions on Davis, with the drug held fixed\. Outside\-prior controls are matched in operator and cardinality, but not in all local sequence or geometric properties\. Natural extensions include context\-matched controls, drug\-side or joint interventions, and alternative structural priors\.

Overall,I\-SAFEcontributes an evaluation methodology for scientific AI systems whose benchmark performance is insufficient to characterize their structural behaviour\. Predictive metrics assess whether a model performs well; scalar audits summarize whether it responds to prior\-guided perturbations;I\-SAFEevaluates how the output distribution reorganizes under those perturbations\. By turning structurally guided interventions into interpretable distributional evidence,I\-SAFEprovides a reusable, model\-agnostic evaluation protocol that supports more precise evaluative claims about black\-box scientific predictors across scientific domains\.

## References

- \[1\]Julius Adebayo, Justin Gilmer, Michael Muelly, Ian J\. Goodfellow, Moritz Hardt, and Been Kim\.Sanity checks for saliency maps\.InAdvances in Neural Information Processing Systems, volume 31, pages 9525–9536, 2018\.
- \[2\]T\. W\. Anderson\.On the distribution of the two\-sample Cramér–von Mises criterion\.The Annals of Mathematical Statistics, 33\(3\):1148–1159, 1962\.
- \[3\]Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez\-Paz\.Invariant risk minimization\.arXiv preprint arXiv:1907\.02893, 2019\.
- \[4\]Gennaro Auricchio, Adelaide Emma Bernardelli, Paolo Giudici, and Giuseppe Toscani\.On rank graduation metrics for high\-dimensional ordinal data\.Mathematical Models and Methods in Applied Sciences, pages 1–35, 2026\.
- \[5\]Gennaro Auricchio, Andrea Codegoni, Stefano Gualandi, Giuseppe Toscani, and Marco Veneroni\.The equivalence of fourier\-based and wasserstein metrics on imaging problems\.Rendiconti Lincei, 31\(3\):627–649, 2020\.
- \[6\]Golnoosh Babaei, Paolo Giudici, and Emanuela Raffinetti\.A rank graduation box for safe ai\.Expert systems with applications, 259:125239, 2025\.
- \[7\]Harald Cramér\.On the composition of elementary errors\.Scandinavian Actuarial Journal, 1928\(1\):13–74, 1928\.
- \[8\]Mindy I\. Davis, Jeremy P\. Hunt, Sanna Herrgard, Pietro Ciceri, Lisa M\. Wodicka, Gabriel Pallares, Michael Hocker, Daniel K\. Treiber, and Patrick P\. Zarrinkar\.Comprehensive analysis of kinase inhibitor selectivity\.Nature Biotechnology, 29\(11\):1046–1051, 2011\.
- \[9\]Finale Doshi\-Velez and Been Kim\.Towards a rigorous science of interpretable machine learning\.arXiv preprint arXiv:1702\.08608, 2017\.
- \[10\]Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts\.Causal abstractions of neural networks\.InAdvances in Neural Information Processing Systems, volume 34, pages 9574–9586, 2021\.
- \[11\]Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D\. Goodman\.Finding alignments between interpretable causal variables and distributed neural representations\.InProceedings of the Third Conference on Causal Learning and Reasoning, volume 236 ofProceedings of Machine Learning Research, pages 160–187, 2024\.
- \[12\]Robert Geirhos, Jörn\-Henrik Jacobsen, Claudio Michaelis, Richard S\. Zemel, Wieland Brendel, Matthias Bethge, and Felix A\. Wichmann\.Shortcut learning in deep neural networks\.Nature Machine Intelligence, 2:665–673, 2020\.
- \[13\]Dennis Graber, Patrick Stockinger, Fabian Meyer, Siddharth Mishra, Christopher Horn, and Rebecca Buller\.Resolving data bias improves generalization in binding affinity prediction\.Nature Machine Intelligence, 7\(10\):1713–1725, 2025\.
- \[14\]Kexin Huang, Tianfan Fu, Wenhao Gao, Yue Zhao, Yusuf Roohani, Jure Leskovec, Connor W\. Coley, Cao Xiao, Jimeng Sun, and Marinka Zitnik\.Therapeutics data commons: machine learning datasets and tasks for drug discovery and development\.InAdvances in Neural Information Processing Systems, 2021\.Datasets and Benchmarks Track\.
- \[15\]Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Mądry\.Adversarial examples are not bugs, they are features\.InAdvances in Neural Information Processing Systems, volume 32, 2019\.
- \[16\]Georgi K\. Kanev, Chris de Graaf, Bart A\. Westerman, Iwan J\. P\. de Esch, and Albert J\. Kooistra\.KLIFS: an overhaul after the first 5 years of supporting kinase research\.Nucleic Acids Research, 49\(D1\):D562–D569, 2021\.
- \[17\]Albert J\. Kooistra, Georgi K\. Kanev, Oscar P\. J\. van Linden, Rob Leurs, Iwan J\. P\. de Esch, and Chris de Graaf\.KLIFS: a structural kinase–ligand interaction database\.Nucleic Acids Research, 44\(D1\):D365–D371, 2016\.
- \[18\]Ingoo Lee, Jongsoo Keum, and Hojung Nam\.DeepConv\-DTI: prediction of drug\-target interactions via deep learning with convolution on protein sequences\.PLOS Computational Biology, 15\(6\):e1007129, 2019\.
- \[19\]Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian D\. Cosgrove, Christopher D\. Manning, Christopher Ré, Diana Acosta\-Navas, Drew A\. Hudson, Eric Zeiler, Dan Jurafsky, Tatsunori Hashimoto, Peter Henderson, and Christopher Potts\.Holistic evaluation of language models\.Transactions on Machine Learning Research, 2023\.
- \[20\]Guanxing Lin, Xinyi Zhang, Zhen Ren, Quan Zou, Prayag Tiwari, Cheng Zhou, and Yi Ding\.TAPB: an interventional debiasing framework for alleviating target prior bias in drug–target interaction prediction\.Nature Communications, 16:10867, 2025\.
- \[21\]Mohammad Lotfollahi, Anna Klimovskaia Susmelj, Carlo De Donno, et al\.Predicting cellular responses to complex perturbations in high\-throughput screens\.Molecular Systems Biology, 19:e11517, 2023\.
- \[22\]Scott M\. Lundberg and Su\-In Lee\.A unified approach to interpreting model predictions\.InAdvances in Neural Information Processing Systems, volume 30, pages 4766–4777, 2017\.
- \[23\]Andrea Mastropietro, Giuseppe Pasculli, and Jürgen Bajorath\.Learning characteristics of graph neural networks predicting protein–ligand affinities\.Nature Machine Intelligence, 5:1427–1436, 2023\.
- \[24\]Hakime Öztürk, Arzucan Özgür, and Elif Ozkirimli\.DeepDTA: deep drug–target binding affinity prediction\.Bioinformatics, 34\(17\):i821–i829, 2018\.
- \[25\]Judea Pearl\.Causality: Models, Reasoning, and Inference\.Cambridge University Press, 2nd edition, 2009\.
- \[26\]Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen\.Causal inference by using invariant prediction: identification and confidence intervals\.Journal of the Royal Statistical Society: Series B, 78\(5\):947–1012, 2016\.
- \[27\]Gabriel Peyré and Marco Cuturi\.Computational optimal transport\.Foundations and Trends in Machine Learning, 11\(5–6\):355–607, 2019\.
- \[28\]Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin\.“Why should I trust you?”: explaining the predictions of any classifier\.InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1135–1144, 2016\.
- \[29\]Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman\.Deep inside convolutional networks: visualising image classification models and saliency maps\.arXiv preprint arXiv:1312\.6034, 2014\.
- \[30\]Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R\. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga\-Alonso, et al\.Beyond the imitation game: quantifying and extrapolating the capabilities of language models\.Transactions on Machine Learning Research, 2023\.
- \[31\]Mukund Sundararajan, Ankur Taly, and Qiqi Yan\.Axiomatic attribution for deep networks\.InProceedings of the 34th International Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Research, pages 3319–3328, 2017\.
- \[32\]Barbara Tarantino, Sun Kim, Yijingxiu Lu, and Paolo Giudici\.Isaac: Auditing causal reasoning in deep models for drug\-target interaction, 2026\.
- \[33\]Derek van Tilborg, Alisa Alenicheva, and Francesca Grisoni\.Exposing the limitations of molecular machine learning with activity cliffs\.Journal of Chemical Information and Modeling, 62\(23\):5938–5951, 2022\.
- \[34\]Cédric Villani\.Optimal Transport: Old and New\.Springer, Berlin, 2009\.
- \[35\]Izhar Wallach and Abraham Heifets\.Most ligand\-based classification benchmarks reward memorization rather than generalization\.Journal of Chemical Information and Modeling, 58\(5\):916–932, 2018\.
- \[36\]Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Ryan Schaeffer, Sang T\. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li\.DecodingTrust: a comprehensive assessment of trustworthiness in GPT models\.InAdvances in Neural Information Processing Systems, volume 36, 2023\.
- \[37\]Zhenqin Wu, Bharath Ramsundar, Evan N\. Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S\. Pappu, Karl Leswing, and Vijay Pande\.MoleculeNet: a benchmark for molecular machine learning\.Chemical Science, 9\(2\):513–530, 2018\.
- \[38\]Xiao Xu, Robert Lawrence, Kumar Dubey, Ayush Pandey, Ryo Ueno, Fabian Falck, Aditya V\. Nori, Rishabh Sharma, Abhay Sharma, and Javier González\.RE\-IMAGINE: symbolic benchmark synthesis for reasoning evaluation\.InProceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, 2025\.
- \[39\]Bin Yu and Karl Kumbier\.Veridical data science\.Proceedings of the National Academy of Sciences, 117\(8\):3920–3929, 2020\.
- \[40\]Meng Zhang, Keng Kiat Goh, Ping Zhang, Jingwei Sun, Ronald Lok Xin, and Huan Zhang\.LLMScan: causal scan for LLM misbehavior detection\.InProceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, 2025\.

## Appendix AAppendix

In this appendix we report the missing proof and all the technical discussion omitted from the main body of the paper\.

### A\.1Proofs

First, we report the missing proofs\.

###### Proof of Proposition[1](https://arxiv.org/html/2605.21731#Thmproposition1)\.

It trivially follows from the fact that

0≤∑i=1K\|qi\(𝒜\)−qi\(φ​\(𝒜\)\)\|∑i=1\|𝒜\|\|fBB​\(xi\)−fBB​\(φ​\(π​\(xi\)\)\)\|2\.0\\leq\\frac\{\\sum\_\{i=1\}^\{K\}\|q\_\{i\}^\{\(\\mathcal\{A\}\)\}\-q\_\{i\}^\{\(\\varphi\(\\mathcal\{A\}\)\)\}\|\}\{\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\|f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-f\_\{\\texttt\{BB\}\}\(\\varphi\(\\pi\(x\_\{i\}\)\)\)\|^\{2\}\}\.\(9\)∎

###### Proof of Proposition[2](https://arxiv.org/html/2605.21731#Thmproposition2)\.

Sinceπ=I​d\\pi=Idis a feasible point for the minimization problem

minπ∈Πn​\(𝒜\)⁡\|fBB​\(xi\)−fBB​\(φ​\(π​\(xi\)\)\)\|2\\min\_\{\\pi\\in\\Pi\_\{n\}\(\\mathcal\{A\}\)\}\|f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-f\_\{\\texttt\{BB\}\}\(\\varphi\(\\pi\(x\_\{i\}\)\)\)\|^\{2\}\(10\)we infer

minπ∈Πn​\(𝒜\)​∑i=1\|𝒜\|\|fBB​\(xi\)−fBB​\(φ​\(π​\(xi\)\)\)\|2≤∑i=1\|𝒜\|\|fBB​\(xi\)−fBB​\(φ​\(xi\)\)\|2,\\min\_\{\\pi\\in\\Pi\_\{n\}\(\\mathcal\{A\}\)\}\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\|f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-f\_\{\\texttt\{BB\}\}\(\\varphi\(\\pi\(x\_\{i\}\)\)\)\|^\{2\}\\leq\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\|f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-f\_\{\\texttt\{BB\}\}\(\\varphi\(x\_\{i\}\)\)\|^\{2\},\(11\)thus0≤minπ∈Πn​\(𝒜\)​∑i=1\|𝒜\|\|fBB​\(xi\)−fBB​\(φ​\(π​\(xi\)\)\)\|2∑i=1\|𝒜\|\|fBB​\(xi\)−fBB​\(φ​\(xi\)\)\|2≤10\\leq\\frac\{\\min\_\{\\pi\\in\\Pi\_\{n\}\(\\mathcal\{A\}\)\}\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\|f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-f\_\{\\texttt\{BB\}\}\(\\varphi\(\\pi\(x\_\{i\}\)\)\)\|^\{2\}\}\{\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\|f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-f\_\{\\texttt\{BB\}\}\(\\varphi\(x\_\{i\}\)\)\|^\{2\}\}\\leq 1, hence the first part of the proof\.

Let us now assume thatW​C​M​\(φ;fBB\)=0WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)=0, we then have that

minπ∈Πn​\(𝒜\)​∑i=1\|𝒜\|\|fBB​\(xi\)−fBB​\(φ​\(π​\(xi\)\)\)\|2=∑i=1\|𝒜\|\|fBB​\(xi\)−fBB​\(φ​\(xi\)\)\|2,\\min\_\{\\pi\\in\\Pi\_\{n\}\(\\mathcal\{A\}\)\}\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\|f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-f\_\{\\texttt\{BB\}\}\(\\varphi\(\\pi\(x\_\{i\}\)\)\)\|^\{2\}=\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\|f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-f\_\{\\texttt\{BB\}\}\(\\varphi\(x\_\{i\}\)\)\|^\{2\},\(12\)meaning that the identity permutation is an optimal solution to the minimization problem in \([10](https://arxiv.org/html/2605.21731#A1.E10)\)

Let us assume thatW​C​M​\(φ;fBB\)=1WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)=1andV𝒜≠Vφ​\(𝒜\)V\_\{\\mathcal\{A\}\}\\neq V\_\{\\varphi\(\\mathcal\{A\}\)\}\. We then have thatminπ∈Πn​\(𝒜\)​∑i=1\|𝒜\|\|fBB​\(xi\)−fBB​\(φ​\(π​\(xi\)\)\)\|2=0\\min\_\{\\pi\\in\\Pi\_\{n\}\(\\mathcal\{A\}\)\}\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\|f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-f\_\{\\texttt\{BB\}\}\(\\varphi\(\\pi\(x\_\{i\}\)\)\)\|^\{2\}=0, meaning that there exists a permutationπ\\pithat is not equal to the identity for which it holds

\(V𝒜\)i=\(Vφ​\(𝒜\)\)π​\(i\)\(V\_\{\\mathcal\{A\}\}\)\_\{i\}=\(V\_\{\\varphi\(\\mathcal\{A\}\)\}\)\_\{\\pi\(i\)\}\(13\)for everyi=1,…,\|𝒜\|i=1,\\dots,\|\\mathcal\{A\}\|\.

To conclude, we notice that for everyλ\\lambda, it holds

minπ∈Πn​\(𝒜\)⁡\|λ​fBB​\(xi\)−λ​fBB​\(φ​\(π​\(xi\)\)\)\|2=λ2​minπ∈Πn​\(𝒜\)⁡\|fBB​\(xi\)−fBB​\(φ​\(π​\(xi\)\)\)\|2\\min\_\{\\pi\\in\\Pi\_\{n\}\(\\mathcal\{A\}\)\}\|\\lambda f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-\\lambda f\_\{\\texttt\{BB\}\}\(\\varphi\(\\pi\(x\_\{i\}\)\)\)\|^\{2\}=\\lambda^\{2\}\\min\_\{\\pi\\in\\Pi\_\{n\}\(\\mathcal\{A\}\)\}\|f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-f\_\{\\texttt\{BB\}\}\(\\varphi\(\\pi\(x\_\{i\}\)\)\)\|^\{2\}\(14\)Likewise,

∑i=1\|𝒜\|\|λ​fBB​\(xi\)−λ​fBB​\(φ​\(xi\)\)\|2=λ2​∑i=1\|𝒜\|\|fBB​\(xi\)−fBB​\(φ​\(xi\)\)\|2,\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\|\\lambda f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-\\lambda f\_\{\\texttt\{BB\}\}\(\\varphi\(x\_\{i\}\)\)\|^\{2\}=\\lambda^\{2\}\\sum\_\{i=1\}^\{\|\\mathcal\{A\}\|\}\|f\_\{\\texttt\{BB\}\}\(x\_\{i\}\)\-f\_\{\\texttt\{BB\}\}\(\\varphi\(x\_\{i\}\)\)\|^\{2\},\(15\)hence the scalar invariant property\. ∎

###### Proof of Proposition[3](https://arxiv.org/html/2605.21731#Thmproposition3)\.

It follows from the fact that

\(A−D\)2\+\(B−C\)2\>\(A−C\)2\+\(B−D\)2\(A\-D\)^\{2\}\+\(B\-C\)^\{2\}\>\(A\-C\)^\{2\}\+\(B\-D\)^\{2\}\(16\)wheneverA<BA<BandC<DC<D\. By iteratively reordering the entries of the vectorsV𝒜V\_\{\\mathcal\{A\}\}andVφ​\(𝒜\)V\_\{\\varphi\(\\mathcal\{A\}\)\}we are able to show that the minimum of the problem \([10](https://arxiv.org/html/2605.21731#A1.E10)\) is given by the increasing reordering, thus the thesis\. ∎

### A\.2Translation\-Invariant WCM: Properties and Discussion

Given a probability distributionμ\\muwith finite averagemμm\_\{\\mu\}, we denote byμ^\\hat\{\\mu\}as the probability distribution shifted by−mμ\-m\_\{\\mu\}, so that the average value ofμ^\\hat\{\\mu\}is equal to0\.

First, we recall that, given any couple of probability distributions with finite average, namelyμ\\muandν\\nu, it holds

W22​\(μ,ν\)−\(mμ−mν\)2=W22​\(μ^,ν^\),W\_\{2\}^\{2\}\(\\mu,\\nu\)\-\(m\_\{\\mu\}\-m\_\{\\nu\}\)^\{2\}=W\_\{2\}^\{2\}\(\\hat\{\\mu\},\\hat\{\\nu\}\),\(17\)thus we can rewrite the Ti\-WCM as follows

T​I−W​C​M​\(φ,fBB\):=1−W2​\(μ^Vφ​\(𝒜\),μ^V𝒜\)ℓ2​\(Vφ​\(𝒜\),V𝒜\)\.TI\-WCM\(\\varphi,f\_\{\\texttt\{BB\}\}\):=1\-\\frac\{W\_\{2\}\(\\hat\{\\mu\}\_\{V\_\{\\varphi\(\\mathcal\{A\}\)\}\},\\hat\{\\mu\}\_\{V\_\{\\mathcal\{A\}\}\}\)\}\{\\ell\_\{2\}\(V\_\{\\varphi\(\\mathcal\{A\}\)\},V\_\{\\mathcal\{A\}\}\)\}\.\(18\)As a consequence, we have that0≤T​I−W​C​M​\(φ,fBB\)0\\leq TI\-WCM\(\\varphi,f\_\{\\texttt\{BB\}\}\)\. Moreover, since\(mμ−mν\)2≥0\(m\_\{\\mu\}\-m\_\{\\nu\}\)^\{2\}\\geq 0, we have

W2​\(μ^Vφ​\(𝒜\),μ^V𝒜\)≤W2​\(μVφ​\(𝒜\),μV𝒜\)W\_\{2\}\(\\hat\{\\mu\}\_\{V\_\{\\varphi\(\\mathcal\{A\}\)\}\},\\hat\{\\mu\}\_\{V\_\{\\mathcal\{A\}\}\}\)\\leq W\_\{2\}\(\\mu\_\{V\_\{\\varphi\(\\mathcal\{A\}\)\}\},\\mu\_\{V\_\{\\mathcal\{A\}\}\}\)\(19\)which, used in conjunction with Proposition[2](https://arxiv.org/html/2605.21731#Thmproposition2)allows us to conclude thatT​I−W​C​M​\(φ,fBB\)∈\[0,1\]TI\-WCM\(\\varphi,f\_\{\\texttt\{BB\}\}\)\\in\[0,1\]\. As a byproduct, we infer that

W2​\(μ^Vφ​\(𝒜\),μ^V𝒜\)ℓ2​\(Vφ​\(𝒜\),V𝒜\)≤W2​\(μVφ​\(𝒜\),μV𝒜\)ℓ2​\(Vφ​\(𝒜\),V𝒜\)\\frac\{W\_\{2\}\(\\hat\{\\mu\}\_\{V\_\{\\varphi\(\\mathcal\{A\}\)\}\},\\hat\{\\mu\}\_\{V\_\{\\mathcal\{A\}\}\}\)\}\{\\ell\_\{2\}\(V\_\{\\varphi\(\\mathcal\{A\}\)\},V\_\{\\mathcal\{A\}\}\)\}\\leq\\frac\{W\_\{2\}\(\\mu\_\{V\_\{\\varphi\(\\mathcal\{A\}\)\}\},\\mu\_\{V\_\{\\mathcal\{A\}\}\}\)\}\{\\ell\_\{2\}\(V\_\{\\varphi\(\\mathcal\{A\}\)\},V\_\{\\mathcal\{A\}\}\)\}\(20\)meaning that0≤W​C​M​\(φ;fBB\)≤T​I−W​C​M​\(φ;fBB\)≤10\\leq WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)\\leq TI\-WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)\\leq 1for everyφ\\varphiand every black\-box modelfBBf\_\{\\texttt\{BB\}\}\. In particular, wheneverW​C​M​\(φ;fBB\)=1WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)=1thenT​I−W​C​M​\(φ;fBB\)=1TI\-WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)=1likewise, ifT​I−W​C​M​\(φ;fBB\)=0TI\-WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)=0thenW​C​M​\(φ;fBB\)=0WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)=0\.

We then notice that the TI\-WCM is a stricter metric than the WCM\. Indeed, ifT​I−W​C​M​\(φ;fBB\)=1TI\-WCM\(\\varphi;f\_\{\\texttt\{BB\}\}\)=1it must be the case that

1. \(i\)the model output distribution after the perturbation has the same average value as the model output distribution before the permutation and
2. \(ii\)it induces the same ranking ordering\.

Lastly, it is easy to adapt the argument used to prove Proposition[2](https://arxiv.org/html/2605.21731#Thmproposition2)to show that also TI\-WCM is scale invariant\.

### A\.3Additional Tables

Table 2:Structural coverage and auditing\-set composition for the Davis benchmark\. Rows report the successive requirements needed to define well\-posed KLIFS\-based interventions under the matched cardinality design\.CriterionCountFraction of test splitTargets in test split379With KLIFS annotation32184\.7%With realizable interventions20854\.9%Drug–target pairs in𝒜\\mathcal\{A\}3,044of whichy=1y=11545\.1%of whichy=0y=02,89094\.9%Median pocket size\|𝒫​\(t\)\|\|\\mathcal\{P\}\(t\)\|85 residuesExact cardinality matching100%Table 3:Predictive performance on the Davis auditing subset\(208\(208targets,3,0443\{,\}044drug–target pairs\)\. AUROC is reported as mean with 95% confidence intervals across five training seeds and is used only to verify a comparable predictive regime before interventional auditing\.Table 4:I\-SAFEprior\-relative coherence contrasts on the Davis benchmark \(operator=all, five seeds, 95 % CI\)\. Contrasts are defined as spurious minus mechanistic\. Positive values indicate greater coherence under mechanistic perturbations than under spurious controls\.Table 5:QBM sensitivity to the quantile grid \(operator=all, five seeds, 95 % CI\)\.Δ​QBM=QBMspur−QBMmech\\Delta\\text\{QBM\}=\\text\{QBM\}\_\{\\text\{spur\}\}\-\\text\{QBM\}\_\{\\text\{mech\}\}; positive values indicate greater coherence under mechanistic perturbations\.
### A\.4Computational Details

AllI\-SAFEanalyses are performed post hoc on fixed model checkpoints, with no retraining, fine\-tuning, or audit\-specific model optimization\. To support reproducibility, the supplemental archive includes the per\-seed audit outputs from which all reported results are derived\. Starting from these outputs, the reproduction pipeline, comprising AUROC verification,I\-SAFEmetric estimation, confidence\-interval computation, figure generation, and QBM sensitivity analysis, completes in approximately 17 minutes and does not require dedicated GPU acceleration\. This runtime was measured on a Windows 11 laptop with an AMD Ryzen AI 7 350 processor, 31\.3 GB RAM, and Python 3\.10\.19 from conda\-forge\. Model training is separate from the released audit pipeline and followed the protocols and original implementations described in prior work\[[24](https://arxiv.org/html/2605.21731#bib.bib24),[18](https://arxiv.org/html/2605.21731#bib.bib18),[20](https://arxiv.org/html/2605.21731#bib.bib20)\]\.

Similar Articles

Adaptive auditing of AI systems with anytime-valid guarantees

arXiv cs.AI

This paper introduces a statistical framework for adaptively auditing AI systems using Safe Anytime-Valid Inference (SAVI) to draw rigorous conclusions with limited data. It proposes a 'testing by betting' approach to validate model robustness while controlling type-I errors during adaptive sampling.