Pathology Transport: Optimal-Transport Explanations for Clinical Data, and When Their Heatmaps (Fail to) Localize Disease

arXiv cs.LG Papers

Summary

This paper investigates optimal-transport explanations for clinical data, showing that while heatmaps can localize synthetic lesions, they fail to localize real disease, highlighting a synthetic-to-real gap in explainable AI for healthcare.

arXiv:2608.17370v1 Announce Type: new Abstract: Generative models promise a route to explainable clinical AI: rather than probe a classifier, model the distributions of healthy and diseased patients and read explanations off the geometry between them. We build such a system - an optimal-transport rectified flow trained between two clinical distributions - and use it to ask a pointed question the field too rarely tests: do the resulting explanation heatmaps actually localize disease? On tabular tumour biomarkers (Breast Cancer Wisconsin) a single flow yields per-patient counterfactuals, an unsupervised malignancy score (AUROC 0.91; 0.93 +/- 0.01 across five seeds), and a label-free attribution that agrees with a supervised classifier (r ~ 0.5) - a compact, honest interpretability engine, though it never out-predicts logistic regression. Moving to chest X-rays, we show the transport heatmap is a population-level signal, not a localiser; a reconstruction-based, identity-preserving variant does localize synthetic lesions (pointing game 0.52), yet on real RSNA radiologist boxes it collapses to chance while only supervised Grad-CAM stays above it. The central result is a synthetic-to-real gap: label-free heatmaps that look compelling on planted lesions are not evidence of real localisation. We contribute a reusable optimal-transport recipe for generative explanations and a controlled benchmark for stress-testing whether they localize.
Original Article
View Cached Full Text

Cached at: 08/19/26, 10:27 AM

# Optimal-Transport Explanations for Clinical Data,and When Their Heatmaps (Fail to) Localize Disease
Source: [https://arxiv.org/html/2608.17370](https://arxiv.org/html/2608.17370)
Lalit KumarAffiliation:Department of Computer Science The University of Texas at Austin lalit\.kumar@utexas\.eduemail:[mailto:](mailto:)

###### Abstract\.

Generative models promise a route to explainable clinical AI: rather than probe a classifier, model the distributions of*healthy*and*diseased*patients and read explanations off the geometry between them\. We build such a system—an optimal\-transport*rectified flow*trained between two clinical distributions—and use it to ask a pointed question the field too rarely tests: do the resulting explanation*heatmaps actually localize disease*? On tabular tumour biomarkers \(Breast Cancer Wisconsin\) a single flow yields per\-patient counterfactuals, an unsupervised malignancy score \(AUROC 0\.91;0\.93±0\.010\.93\\pm 0\.01across five seeds\), and a label\-free attribution that agrees with a supervised classifier \(r≈0\.5r\{\\approx\}0\.5\)—a compact, honest interpretability engine, though it never out\-predicts logistic regression\. Moving to chest X\-rays, we show the transport heatmap is a*population\-level*signal, not a localiser; a reconstruction\-based, identity\-preserving variant*does*localize*synthetic*lesions \(pointing game0\.520\.52\), yet on*real*RSNA radiologist boxes it collapses to chance while only supervised Grad\-CAM stays above it\. The central result is a*synthetic\-to\-real gap*: label\-free heatmaps that look compelling on planted lesions are not evidence of real localisation\. We contribute a reusable optimal\-transport recipe for generative explanations and a controlled benchmark for stress\-testing whether they localize\.

###### Keywords:

rectified flow, optimal transport, counterfactual explanation, explainable AI, generative models, breast cancer, clinical decision support

## 1\.Introduction

Clinical adoption of predictive models is limited less by accuracy than by*trust and actionability*\. A model that outputs “87% malignant” gives a clinician a number but not a rationale: which measurements drove the decision, and what would have to be different for the verdict to change? Explainable\-AI methods such as SHAP and saliency answer the first question with feature\-importance weights, but they do not produce a concrete, on\-distribution example of the counterfactual patient\. Counterfactual explanations—“the smallest change to the inputs that flips the prediction”—answer exactly this question and are increasingly seen as the form of explanation clinicians and regulators actually want\([Wachter et al\. 2017](https://arxiv.org/html/2608.17370#bib.bib15)\)\.

The dominant way to obtain counterfactuals is to perturb the input against a fixed classifier\. We ask a different, higher\-risk question:*what if we never train a classifier at all, and instead learn the geometry that separates health from disease directly?*Concretely, we treat the benign and malignant patient populations as two probability distributions in biomarker space and learn the*optimal\-transport map*that carries one onto the other\. Recent advances in generative modelling make this practical:*rectified flow*\([Liu et al\. 2023](https://arxiv.org/html/2608.17370#bib.bib6)\)and*flow matching*\([Lipman et al\. 2023](https://arxiv.org/html/2608.17370#bib.bib5)\)learn a velocity field whose ordinary differential equation transports one distribution into another by simple regression, and mini\-batch optimal\-transport coupling\([Tong et al\. 2024](https://arxiv.org/html/2608.17370#bib.bib13)\)makes those trajectories nearly straight and stable\.

Our thesis is that a*single*such transport model is a surprisingly complete clinical explanation engine\. The map itself is a per\-patient counterfactual generator; the length a patient must travel is an unsupervised risk score; and the average displacement is a global attribution\. We test this on the Breast Cancer Wisconsin \(Diagnostic\) dataset\([Street et al\. 1993](https://arxiv.org/html/2608.17370#bib.bib12)\)—569 biopsies, 30 nuclear biomarkers—chosen because it is real clinical data, is fully offline, and has a known ground\-truth biomarker signature against which our label\-free attribution can be validated\.

This is a high\-risk design and we treat it as such\. Learned distribution transport can overshoot and fabricate off\-manifold “patients”; an unsupervised score has no guarantee of matching a supervised classifier; and near\-linearly\-separable data is a regime where simple linear methods are notoriously hard to beat\. We report where the method succeeds and, just as clearly, where it does not\. The imaging half then presses a sharper question—do these generative heatmaps genuinely*localize*disease?—which we answer with a controlled synthetic benchmark and a real RSNA annotation test\.

### Contributions\.

- •We reframe tumour diagnosis as optimal transport between clinical distributions and implement it with an OT\-coupled rectified flow \(Section[4](https://arxiv.org/html/2608.17370#S4)\)\.
- •From one model we derive three interpretability artefacts—counterfactuals, an unsupervised risk score, and population attribution—and evaluate each quantitatively \(Section[5](https://arxiv.org/html/2608.17370#S5)\)\.
- •We give an honest account of failure modes: no accuracy gain over logistic regression, non\-sparse edits, and a score whose discrimination partly reflects off\-manifold drift\.
- •We show the recipe is modality\-agnostic and, via a controlled synthetic\-lesion benchmark and a*real*RSNA bounding\-box test, deliver a cautionary finding: a label\-free heatmap that localises synthetic lesions fails to transfer to real pathology\.

## 2\.Related Work

### Counterfactual explanations\.

Wachter et al\.\([Wachter et al\. 2017](https://arxiv.org/html/2608.17370#bib.bib15)\)formalised counterfactual explanations as an optimisation that finds the nearest input flipping a model’s decision, and argued they satisfy legal “right to explanation” requirements without exposing model internals\. Subsequent work adds validity, sparsity and plausibility constraints\([Guidotti 2024](https://arxiv.org/html/2608.17370#bib.bib3)\)\. Our approach differs in that the counterfactual is produced by a*generative transport map*rather than by gradient descent against a classifier, so it is defined even when no classifier exists and is naturally biased toward the real data manifold\.

### Explainable AI in healthcare\.

Feature\-attribution methods such as SHAP\([Lundberg and Lee 2017](https://arxiv.org/html/2608.17370#bib.bib7)\)have become the default lens for interpreting clinical risk models, assigning each input a contribution to the prediction\. They are, however, tied to a trained predictor and explain*a model’s decision*rather than*the disease*: they cannot synthesise the patient who would receive a different verdict, nor localise pathology in an image without a supervised detector\. Our transport\-based view is complementary—it explains the*data geometry*separating health from disease, and a single object yields counterfactuals, a score, and spatial attribution at once\.

### Generative and diffusion counterfactuals\.

A growing line of work generates counterfactuals with deep generative models, especially in medical imaging, e\.g\. diffusion\-based counterfactuals that morph a diseased scan into its healthy version to localise pathology\([Jeanneret et al\. 2022](https://arxiv.org/html/2608.17370#bib.bib4)\)\. These methods still condition on an external classifier for guidance\. We instead let the transport between two*unconditional*class distributions define the edit, which is closer in spirit to population\-dynamics models that use optimal transport to interpolate biological state, such as TrajectoryNet for single\-cell trajectories\([Tong et al\. 2020](https://arxiv.org/html/2608.17370#bib.bib14)\)\.

### Saliency, localisation, and faithfulness\.

For images we compare against*Grad\-CAM*\([Selvaraju et al\. 2017](https://arxiv.org/html/2608.17370#bib.bib10)\), the standard supervised saliency method, and quantify map quality with deletion/insertion faithfulness\([Petsiuk et al\. 2018](https://arxiv.org/html/2608.17370#bib.bib8)\)and a pointing\-game/IoU protocol against ground\-truth boxes\. Our label\-free localiser is a*reconstruction\-based*anomaly detector in the spirit of f\-AnoGAN\([Schlegl et al\. 2019](https://arxiv.org/html/2608.17370#bib.bib9)\)and autoencoder anomaly segmentation\([Baur et al\. 2021](https://arxiv.org/html/2608.17370#bib.bib2)\)—a model of healthy anatomy whose reconstruction error flags the abnormal, an approach whose strongest variants now use diffusion restoration\([Wyatt et al\. 2022](https://arxiv.org/html/2608.17370#bib.bib16)\)—which we evaluate on real annotated pneumonia from the RSNA Pneumonia Detection Challenge\([Shih et al\. 2019](https://arxiv.org/html/2608.17370#bib.bib11)\)\.

### Flow matching and optimal transport\.

Rectified flow\([Liu et al\. 2023](https://arxiv.org/html/2608.17370#bib.bib6)\)and flow matching\([Lipman et al\. 2023](https://arxiv.org/html/2608.17370#bib.bib5)\)learn a velocity field by regressing onto the direction of straight\-line interpolations between paired samples, avoiding the simulation of stochastic diffusion\. Tong et al\.\([Tong et al\. 2024](https://arxiv.org/html/2608.17370#bib.bib13)\)show that choosing the pairing via mini\-batch optimal transport straightens trajectories and improves sample quality; this coupling is central to our method’s stability\. To our knowledge, using an OT\-coupled rectified flow*between two clinical class distributions*to jointly yield counterfactuals, a risk score, and attribution has not been reported\.

## 3\.Background

### Optimal transport\.

Given a source distributionμ\\muand a targetν\\nu, optimal transport seeks the map \(or plan\) that morphs one into the other at minimum total cost\. Under the squared\-Euclidean cost the Monge problem isminT:T\#​μ=ν∫∥x−T\(x\)∥2dμ\(x\)\\min\_\{T:\\,T\_\{\\\#\}\\mu=\\nu\}\\int\\\|x\-T\(x\)\\\|^\{2\}\\,d\\mu\(x\), and its Kantorovich relaxation optimises over couplingsπ∈Π⁡\(μ,ν\)\\pi\\in\\Pi\(\\mu,\\nu\), minimising∫‖x0−x1‖2​𝑑π​\(x0,x1\)\\int\\\|x\_\{0\}\-x\_\{1\}\\\|^\{2\}\\,d\\pi\(x\_\{0\},x\_\{1\}\)\. Intuitively, OT pairs each source point with the target point it can reach most cheaply, so the “movement” it prescribes is the most economical—and therefore most interpretable—transformation of one population into the other\. We never form the full continuous plan; we approximate it per mini\-batch with a discrete assignment \(Section[4](https://arxiv.org/html/2608.17370#S4)\)\.

### Rectified flow\.

A rectified flow\([Liu et al\. 2023](https://arxiv.org/html/2608.17370#bib.bib6)\)represents transport as an ordinary differential equationx˙=vθ​\(x,t\)\\dot\{x\}=v\_\{\\theta\}\(x,t\)that carries samples ofμ\\muatt=0t\{=\}0to samples ofν\\nuatt=1t\{=\}1\. Rather than simulating a stochastic process, it*regresses*the velocity onto the direction of a straight line between paired endpoints—the flow\-matching objective\([Lipman et al\. 2023](https://arxiv.org/html/2608.17370#bib.bib5)\)\. When the endpoints are paired by optimal transport\([Tong et al\. 2024](https://arxiv.org/html/2608.17370#bib.bib13)\), the target velocities have low variance and the learned trajectories are nearly straight, so a coarse Euler integrator suffices and the net displacementx−F⁡\(x\)x\-F\(x\)is a faithful estimate of the transport\. That displacement is the quantity we mine for every downstream artefact\.

## 4\.Methodology

![Refer to caption](https://arxiv.org/html/2608.17370v1/workflow.png)Figure 1\.Workflow\. Standardised biomarkers are split into benign \(source\) and malignant \(target\) sets\. Within each mini\-batch a Hungarian optimal\-transport assignment pairs source and target points; a time\-conditioned velocity fieldvθ​\(xt,t\)v\_\{\\theta\}\(x\_\{t\},t\)is regressed onto the pairing direction \(rectified flow\) and integrated with an Euler ODE under EMA weights\. Two flows,Fb→mF\_\{b\\to m\}andFm→bF\_\{m\\to b\}, yield a risk score, per\-patient counterfactuals, and population attribution, each evaluated against supervised or naive baselines\.### Data\.

The Breast Cancer Wisconsin \(Diagnostic\) dataset contains 569 fine\-needle\-aspiration biopsies, each described by 30 real\-valued nuclear biomarkers \(the mean, standard error and “worst” of ten cell\-nucleus measurements\)\. We label malignant as the positive class, standardise features using statistics from the training split only, and hold out 30% for evaluation\. Working inzz\-score space makes a unit of transport displacement comparable across biomarkers\.

### Rectified flow\.

We learn a velocity fieldvθ​\(x,t\):ℝ30×\[0,1\]→ℝ30v\_\{\\theta\}\(x,t\):\\mathbb\{R\}^\{30\}\\times\[0,1\]\\to\\mathbb\{R\}^\{30\}whose ODEx˙=vθ​\(x,t\)\\dot\{x\}=v\_\{\\theta\}\(x,t\)transports the benign distribution att=0t\{=\}0into the malignant distribution att=1t\{=\}1\. Following rectified flow, we train on straight\-line interpolations between paired endpoints:

\(1\)xt=\(1−t\)​x0\+t​x1,ℒ⁡\(θ\)=𝔼x0,x1,t​‖vθ​\(xt,t\)−\(x1−x0\)‖2\.x\_\{t\}=\(1\-t\)\\,x\_\{0\}\+t\\,x\_\{1\},\\quad\\mathcal\{L\}\(\\theta\)=\\mathbb\{E\}\_\{x\_\{0\},x\_\{1\},t\}\\big\\\|v\_\{\\theta\}\(x\_\{t\},t\)\-\(x\_\{1\}\-x\_\{0\}\)\\big\\\|^\{2\}\.vθv\_\{\\theta\}is a four\-layer SiLU multilayer perceptron with a sinusoidal time embedding\.

### Optimal\-transport coupling\.

The choice of pairing\(x0,x1\)\(x\_\{0\},x\_\{1\}\)is decisive\. Independent random pairing yields extremely high\-variance regression targets, so the integrated ODE overshoots and lands off\-manifold\. We instead use mini\-batch optimal\-transport coupling: within each batch we solve the assignmentmin⁡∑iπ⁡‖x0\(i\)−x1\(π⁡\(i\)\)‖2\\min\_\{\\pi\}\\sum\_\{i\}\\\|x\_\{0\}^\{\(i\)\}\-x\_\{1\}^\{\(\\pi\(i\)\)\}\\\|^\{2\}with the Hungarian algorithm and pair points by their transport\-optimal match\([Tong et al\. 2024](https://arxiv.org/html/2608.17370#bib.bib13)\)\. In our experiments this alone cut the counterfactual edit size by2\.6×2\.6\\timesand turned an incoherent attribution \(r=−0\.03r\{=\}\-0\.03\) into a clinically sensible one \(r=0\.49r\{=\}0\.49\)\. We keep an exponential\-moving\-average copy of the weights for stable integration and train both directions,Fb→mF\_\{b\\to m\}andFm→bF\_\{m\\to b\}, with a 100\-step Euler integrator \(Figure[1](https://arxiv.org/html/2608.17370#S4.F1)\)\.

Algorithm 1OT\-coupled rectified\-flow training1:source set

AA, target set

BB, steps

NN, batch size

bb
2:initialise

vθv\_\{\\theta\}; EMA copy

θ¯←θ\\bar\{\\theta\}\\leftarrow\\theta
3:for

i=1i=1to

NNdo

4:sample

\{x0k\}∼A\\\{x\_\{0\}^\{k\}\\\}\\sim A,

\{x1k\}∼B\\\{x\_\{1\}^\{k\}\\\}\\sim B,

k=1\.\.bk=1\.\.b
5:

Ck​l←‖x0k−x1l‖2C\_\{kl\}\\leftarrow\\\|x\_\{0\}^\{k\}\-x\_\{1\}^\{l\}\\\|^\{2\}⊳\\trianglerightcost matrix

6:

π←Hungarian​\(C\)\\pi\\leftarrow\\textsc\{Hungarian\}\(C\)⊳\\trianglerightmini\-batch OT pairing

7:

x1←x1​\[π\]x\_\{1\}\\leftarrow x\_\{1\}\[\\pi\]
8:

t∼𝒰⁡\(0,1\)t\\sim\\mathcal\{U\}\(0,1\);

xt←\(1−t\)​x0\+t​x1x\_\{t\}\\leftarrow\(1\-t\)x\_\{0\}\+t\\,x\_\{1\}
9:

ℒ←‖vθ​\(xt,t\)−\(x1−x0\)‖2\\mathcal\{L\}\\leftarrow\\\|v\_\{\\theta\}\(x\_\{t\},t\)\-\(x\_\{1\}\-x\_\{0\}\)\\\|^\{2\}
10:

θ←Adam​\(∇θℒ\)\\theta\\leftarrow\\text\{Adam\}\(\\nabla\_\{\\theta\}\\mathcal\{L\}\);

θ¯←0\.999​θ¯\+0\.001​θ\\bar\{\\theta\}\\leftarrow 0\.999\\,\\bar\{\\theta\}\+0\.001\\,\\theta
11:endfor

12:returnEMA model

F=θ¯F=\\bar\{\\theta\}

### Three artefacts\.

*\(1\) Risk score\.*We score a patient by the magnitude of the benign→\\tomalignant transport applied to it,s⁡\(x\)=‖x−Fb→m​\(x\)‖s\(x\)=\\\|x\-F\_\{b\\to m\}\(x\)\\\|\. A genuinely malignant patient lies off the benign source support, so the ODE drifts it farther, giving a larger score\.*\(2\) Counterfactuals\.*For a benign patient we integrateFb→mF\_\{b\\to m\}to synthesise the nearest malignant version of that same tumour; the displacement names the responsible biomarkers\.*\(3\) Attribution\.*Averaging the displacement over the benign cohort gives a global picture of the benign→\\tomalignant transition\.

### Baselines and metrics\.

We compare the risk score against supervised logistic regression and a naive distance\-to\-benign\-mean, by AUROC\. Counterfactuals are assessed by*validity*\(does an independent classifier flip?\),*plausibility*\(distance to the nearest real malignant patient\), and*individualisation*\(mean cosine similarity of edit directions across patients\), against a mean\-shift baseline that adds the constant class\-mean difference to everyone\. Attribution is compared to logistic\-regression coefficients\. All experiments use a fixed seed and are reproducible\.

### Implementation\.

Table[1](https://arxiv.org/html/2608.17370#S4.T1)lists the settings for both modalities\. The tabular velocity field is a four\-layer SiLU MLP; the image field is a small time\-conditioned U\-Net\. Each flow trains in a few minutes on a single Apple\-silicon GPU\.

Table 1\.Architecture and training settings\.
### Ablation: the coupling choice\.

Table[2](https://arxiv.org/html/2608.17370#S4.T2)isolates the single most important design decision\. Replacing independent \(random\) pairing with mini\-batch optimal\-transport pairing—changing nothing else—shrinks the mean counterfactual edit from25\.325\.3to9\.69\.6and flips the population attribution from anti\-correlated noise \(r=−0\.03r\{=\}\-0\.03\) to a clinically coherent signal \(r=0\.49r\{=\}0\.49\)\. This is the difference between a model that fabricates off\-manifold “patients” and one whose displacements carry meaning; it is the crux of the whole method\.

Table 2\.Effect of the coupling on the tabular flow \(all else fixed\)\.

## 5\.Results

![Refer to caption](https://arxiv.org/html/2608.17370v1/risk_roc_auroc.png)Figure 2\.Left: ROC curves for malignancy detection\. Right: AUROC\. The unsupervised transport score \(0\.91\) beats the naive distance baseline \(0\.88\) but not supervised logistic regression \(0\.99\)\.### An unsupervised risk score \(Figure[2](https://arxiv.org/html/2608.17370#S5.F2)\)\.

The transport magnitudes⁡\(x\)s\(x\)reaches AUROC0\.9120\.912*without ever seeing a label during flow training*, above the naive baseline \(0\.8820\.882\) and far above chance\. It does not beat supervised logistic regression \(0\.9920\.992\)\. This is the expected high\-risk outcome: on a near\-linearly\-separable cohort a simple supervised classifier is very hard to beat, and part of the transport score’s discrimination reflects off\-manifold drift of malignant inputs rather than a calibrated likelihood\. The value of the flow lies in the explanations it produces, not raw discrimination\.

Table 3\.Counterfactual quality \(benign→\\tomalignant\)\. Lower plausibility is better; cosine similarity of 1\.0 means every patient receives an identical edit\.![Refer to caption](https://arxiv.org/html/2608.17370v1/counterfactual_waterfall.png)Figure 3\.Counterfactual for the most malignant\-looking benign patient\. The flow pushes exactly the biomarkers a pathologist expects—larger, more concave, more irregular nuclei—flipping the classifier’s probability\.
### Per\-patient counterfactuals \(Table[3](https://arxiv.org/html/2608.17370#S5.T3), Figure[3](https://arxiv.org/html/2608.17370#S5.F3)\)\.

The flow’s counterfactuals are 84% valid and, unlike mean\-shift,*individualised*: edit directions differ across patients \(cosine similarity 0\.64 versus 1\.00\)\. The worked example raises concavity, concave points, perimeter and area, matching clinical intuition\. The honest trade\-off is that mean\-shift trivially flips 100% of a linear classifier and lands slightly closer to the malignant centroid—on near\-linear data a constant translation is hard to beat—but it applies the*same*edit to every patient and ignores each tumour’s geometry\.

![Refer to caption](https://arxiv.org/html/2608.17370v1/attribution.png)Figure 4\.Left: mean transport displacement per biomarker\. Right: flow attribution versus logistic\-regression coefficients \(r=0\.49r\{=\}0\.49\)\. The two label\-agnostic and supervised views agree on the malignancy signature\.
### Population attribution \(Figure[4](https://arxiv.org/html/2608.17370#S5.F4)\)\.

The mean displacement correlates with logistic\-regression coefficients \(r=0\.49r\{=\}0\.49\) and agrees on the drivers of malignancy—concavity, concave points, perimeter, radius and area\. The flow recovers this signature*without labels*, purely from moving one distribution onto another, a non\-trivial validation that the transport captures real structure\.

![Refer to caption](https://arxiv.org/html/2608.17370v1/pca_trajectories.png)Figure 5\.Benign patients \(green\) transported into the malignant cloud \(red\) along the learned flow, projected to two principal components—a direct picture of “pathology transport”\.Figure[5](https://arxiv.org/html/2608.17370#S5.F5)visualises the transport: benign patients are carried smoothly into the malignant region, confirming that the OT\-coupled flow produces coherent trajectories rather than erratic jumps\.

### Multi\-seed confidence intervals\.

The headline numbers above come from a single split\. Repeating the tabular pipeline over five random train/test splits, the transport risk score attains AUROC0\.934±0\.0120\.934\\pm 0\.012\(mean±\\pm95% CI\), versus0\.894±0\.0160\.894\\pm 0\.016for the naive baseline and0\.992±0\.0040\.992\\pm 0\.004for logistic regression; the attribution correlation is0\.62±0\.070\.62\\pm 0\.07\. The gaps are stable and the ranking never changes across seeds, so the single\-split results are representative\.

### Counterfactual baselines\.

Table[4](https://arxiv.org/html/2608.17370#S5.T4)compares our flow counterfactual with the classic gradient method of Wachter et al\.\([Wachter et al\. 2017](https://arxiv.org/html/2608.17370#bib.bib15)\)and the mean\-shift baseline\. No method dominates\. Wachter produces the sparsest, smallest edits—but it*requires*a differentiable classifier and optimises directly against it\. Mean\-shift trivially flips every case yet applies an identical, generic edit to all patients\. The flow is the only method that is simultaneously*classifier\-free*and*individualised*, at the cost of larger, denser edits: its niche is generating per\-patient counterfactuals when no classifier is available\.

Table 4\.Counterfactual methods \(benign→\\tomalignant, seed 0\)\. Proximity/\#feat measure edit size and sparsity; plausibility is 1\-NN distance to real malignant data \(lower is better\)\.†needs class means; applies the identical edit to every patient\.

## 6\.Extension to Chest X\-ray Images

Nothing in the method is specific to tabular data\. To test generality we apply the*identical*recipe to raw pixels: OT\-coupled rectified flows between*normal*and*pneumonia*chest X\-rays \(PneumoniaMNIST\([Yang et al\. 2023](https://arxiv.org/html/2608.17370#bib.bib17)\),28×2828\{\\times\}28, 4708 train / 624 test\), replacing the MLP velocity field with a small time\-conditioned convolutional U\-Net\. We trainFn→pF\_\{n\\to p\}\(normal→\\topneumonia, for synthesis\) andFp→nF\_\{p\\to n\}\(pneumonia→\\tonormal, for scoring and localisation\)\.

![Refer to caption](https://arxiv.org/html/2608.17370v1/img_risk_auroc.png)Figure 6\.Image risk score‖x−Fp→n​\(x\)‖\\\|x\-F\_\{p\\to n\}\(x\)\\\|\(edit\-distance\-to\-normal\)\. The unsupervised transport score \(AUROC 0\.67\) beats naive mean\-intensity \(0\.58\) but not a supervised pixel classifier \(0\.93\)—the same pattern as the tabular cohort\.### Risk score \(Figure[6](https://arxiv.org/html/2608.17370#S6.F6)\)\.

Scoring an image by the edit\-distance required to make it look normal gives AUROC0\.670\.67, above the naive mean\-intensity baseline \(0\.580\.58\) and below a supervised pixel\-level logistic regression \(0\.930\.93\)—echoing Part A\.

![Refer to caption](https://arxiv.org/html/2608.17370v1/img_morph.png)

![Refer to caption](https://arxiv.org/html/2608.17370v1/img_heatmap.png)

Figure 7\.Top: integratingFn→pF\_\{n\\to p\}synthesises disease—a healthy lung progressively fills with haze \(consolidation\)\. Bottom: for real pneumonia inputs,Fp→n​\(x\)−xF\_\{p\\to n\}\(x\)\-xshows what the flow removes to normalise the lung; the signal \(blue\) concentrates in the lung fields \(its faithfulness is examined in Section[7](https://arxiv.org/html/2608.17370#S7)\)\.
### Synthesis and spatial attribution \(Figure[7](https://arxiv.org/html/2608.17370#S6.F7)\)\.

With no pixel\-level labels, the flow*synthesises*plausible disease progression and produces a spatial map that appears to highlight the lung fields\. Whether this is a*faithful*localisation—a strong claim—we test directly in Section[7](https://arxiv.org/html/2608.17370#S7); the honest answer at28×2828\{\\times\}28is that it is a qualitative visualisation, not a validated detector\. The synthesised images are also blurry and the risk score only moderately discriminative\.

## 7\.Do Generative Heatmaps Localize Disease?

This section is the paper’s core empirical question\. Having established the transport model as a tabular interpretability engine, we now ask whether its*image*explanations genuinely localise pathology—first on controlled synthetic lesions, then on real radiologist annotations\.

### Does the heatmap localise pathology? \(A ground\-truth test\.\)

We stress\-tested the paper’s most eye\-catching claim with a controlled benchmark: we insert a soft Gaussian opacity into a*normal*lung at a*known*location and ask whether a heatmap lands on the lesion, measuring pointing\-game accuracy, IoU, and the heatmap energy inside the lesion \(Table[5](https://arxiv.org/html/2608.17370#S7.T5)\)\. The population transport heatmap\|x−Fp→n​\(x\)\|\|x\-F\_\{p\\to n\}\(x\)\|barely exceeds a random map \(pointing game0\.170\.17\) and trails supervised Grad\-CAM \(0\.570\.57\)\. The reason is instructive: the flow performs a*global*distribution shift—nudging the whole lung toward the normal manifold—so its displacement spreads across the image instead of concentrating on the lesion\. Corroborating probes agree: deletion/insertion faithfulness against a CNN \(test AUROC0\.9140\.914\) cannot separate the transport map from random at28×2828\{\\times\}28, and its rank\-correlation with Grad\-CAM is−0\.09\-0\.09\. The population transport heatmap is thus the image analogue of our tabular attribution—a*population\-level*signal, not a per\-patient localiser\.

### A lesion\-focused fix\.

This diagnosis suggests the remedy: replace the population map with an*identity\-preserving*model that changes only the anomaly\. We test two label\-free or sparse variants on the same benchmark \(Table[5](https://arxiv.org/html/2608.17370#S7.T5), Figure[8](https://arxiv.org/html/2608.17370#S7.F8)\)\. A minimal\-L1L\_\{1\}counterfactual against the CNN fails \(0\.080\.08\): the classifier only weakly flags the synthetic opacity \(p=0\.61p\{=\}0\.61\), leaving little gradient to concentrate\. But a*normal\-manifold autoencoder*—a small model trained to reconstruct*healthy*lungs only, whose reconstruction error flags whatever it cannot explain—localises the lesion at pointing game0\.390\.39, IoU0\.200\.20: over2×2\\timesthe population transport and approaching supervised Grad\-CAM,*without any labels*\. Higher resolution helps further: the same autoencoder at128×128128\{\\times\}128reaches pointing game0\.520\.52and IoU0\.360\.36\(Figure[9](https://arxiv.org/html/2608.17370#S7.F9)\), cleanly isolating focal opacities, while a simple supervised Grad\-CAM does not transfer to focal localisation at that resolution\. On these*synthetic*lesions, then, localisation is recoverable with an identity\-preserving objective\. Whether it survives contact with*real*pathology is the decisive question, which we test next\.

Table 5\.Localization on synthetic lesions, ground truth known \(N=200N\{=\}200; higher is better\)\. The population transport is near\-random; a label\-free normal\-manifold autoencoder recovers most of the gap to supervised Grad\-CAM\.![Refer to caption](https://arxiv.org/html/2608.17370v1/sparse_localization.png)Figure 8\.Lesion\-focused variants vs population transport \(green circle = true lesion\)\. The identity\-preserving normal\-manifold autoencoder \(middle row\) concentrates on the lesion, whereas the sparse counterfactual \(top\) and the population transport \(bottom\) scatter and miss it\.![Refer to caption](https://arxiv.org/html/2608.17370v1/localization128.png)Figure 9\.Scaling the same label\-free normal\-manifold autoencoder to128×128128\{\\times\}128\(top\) tightly localises the synthetic lesion \(green\), reaching pointing game0\.520\.52/ IoU0\.360\.36; a simple supervised Grad\-CAM \(bottom\) does not localise focal opacities at this resolution\.
### Does it transfer to real pathology? \(An honest synthetic\-to\-real gap\.\)

Synthetic success can mislead, so we ran the decisive test on*real*annotated data: the RSNA Pneumonia Detection Challenge\([Shih et al\. 2019](https://arxiv.org/html/2608.17370#bib.bib11)\), whose radiologist bounding boxes give true lesion locations\. We downloaded a bounded subset \(200200boxed positives and400400normals, DICOM, resized to128128\), trained the same normal\-manifold autoencoder on the healthy images, and scored localisation against the real boxes \(Table[6](https://arxiv.org/html/2608.17370#S7.T6), Figure[10](https://arxiv.org/html/2608.17370#S7.F10)\)\. The result is sobering: the autoencoder that excelled on synthetic lesions now performs*at or below*a random map \(pointing game0\.110\.11vs0\.180\.18\), because on real chest X\-rays the reconstruction error is dominated by normal anatomical variation—rib edges, diaphragm, mediastinum—rather than by the diffuse consolidation\. Only the*supervised*Grad\-CAM localises above chance \(0\.300\.30\), and even it is weak on this hard task\. The lesson is methodological and, we think, the most useful finding of the image study: a heatmap that looks compelling on synthetic anomalies is not evidence of real localisation, and label\-free reconstruction does not yet solve pneumonia localisation\. Targeted attempts to close the gap—tripling the healthy training set, a denoising autoencoder, and a reconstruction\-derived lung\-focus mask—did not help \(all remained at or below a random map\), indicating the problem demands fundamentally stronger anomaly models rather than tuning\.

Table 6\.Localization on*real*RSNA lesion boxes \(8080held\-out positives; higher is better\)\. The label\-free autoencoder that worked on synthetic lesions is now no better than random; only supervised saliency exceeds chance\.![Refer to caption](https://arxiv.org/html/2608.17370v1/rsna_localization.png)Figure 10\.Real\-lesion localisation on RSNA Pneumonia \(green = radiologist bounding box\)\. On real chest X\-rays the label\-free autoencoder \(top\) no longer concentrates on the lesion; a supervised Grad\-CAM \(bottom\) does modestly better, but the task remains hard\.

## 8\.Limitations

Our study is a controlled investigation, not a clinical validation, and its scope is deliberately small\. \(i\)*Scale and cohorts\.*The tabular data is a single 569\-patient cohort; the imaging experiments use2828–128128px inputs and an∼1\.4\{\\sim\}1\.4k\-image RSNA subset with a modestly accurate classifier \(AUROC0\.680\.68–0\.750\.75\), so its Grad\-CAM is a floor, not a ceiling\. \(ii\)*No method win\.*Every artefact ties or trails a simple supervised baseline; the contribution is honest characterisation, not state of the art\. \(iii\)*Synthetic benchmark\.*The planted\-lesion benchmark isolates localisation cleanly but does not reproduce the diffuse, textured appearance of real consolidation—which is precisely why the synthetic\-to\-real gap arises\. \(iv\)*Calibration\.*The transport risk score is uncalibrated and must not gate care\. \(v\)*Sensitivity\.*Results are fixed\-seed \(and, where noted, averaged over five splits\), but a full hyper\-parameter and architecture sweep is left to future work\.

## 9\.Conclusion

We reframed diagnosis as optimal transport between clinical distributions and realised it with OT\-coupled rectified flows across*two*modalities\. On tabular tumour biomarkers a single model produced clinically coherent per\-patient counterfactuals, an unsupervised malignancy score \(AUROC 0\.91\), and a label\-free biomarker attribution that agrees with a supervised classifier \(r=0\.49r\{=\}0\.49\)\. On chest X\-rays the*same*recipe synthesised disease progression and produced a spatial pathology heatmap without any pixel labels\. Neither out\-predicted logistic regression—the expected, honest result—but both produced something a classifier cannot: an explicit, navigable path between health and disease\. A controlled study also delivered a cautionary result: a label\-free reconstruction localiser that excels on*synthetic*lesions fails to transfer to*real*RSNA bounding boxes, a synthetic\-to\-real gap that we consider the study’s most useful lesson\.

### Ethical and clinical considerations\.

Because the transport*synthesises*plausible patients and lungs, it must be framed as a hypothesis\-generating and explanatory aid, not a diagnostic device: a synthesised “malignant twin” visualises the model’s learned geometry, not medical advice, and could mislead if shown without context\. The unsupervised score is uncalibrated and should never gate care on its own\. Both datasets are small and demographically narrow—the biopsy cohort is single\-institution and the X\-rays are paediatric—so the attribution may not transfer across scanners, sites, or populations\. Any deployment would require calibration, prospective validation, and a fairness audit of the attribution across subgroups\.

### Future work\.

Several extensions could make the method clinically actionable: \(i\) anL1L\_\{1\}or immutability penalty along the ODE to yield*sparse*, actionable counterfactuals; \(ii\) exact log\-likelihood via the flow’s instantaneous change of variables for a calibrated risk score; \(iii\) a single class\- and time\-conditioned field with classifier guidance instead of two flows; \(iv\) closing the synthetic\-to\-real localisation gap with stronger anomaly models \(self\-supervised or diffusion\-based restoration\([Wyatt et al\. 2022](https://arxiv.org/html/2608.17370#bib.bib16)\)\) evaluated on real annotated boxes \(RSNA\); and \(v\) stochastic \(SDE\) transport for uncertainty and a fairness audit of attribution across patient subgroups\.

## References

- \(1\)
- Baur et al\.\(2021\)Christoph Baur, Stefan Denner, Benedikt Wiestler, Nassir Navab, and Shadi Albarqouni\. 2021\.Autoencoders for unsupervised anomaly segmentation in brain MR images: A comparative study\.*Medical Image Analysis*69 \(2021\), 101952\.
- Guidotti \(2024\)Riccardo Guidotti\. 2024\.Counterfactual explanations and how to find them: literature review and benchmarking\.*Data Mining and Knowledge Discovery*38, 5 \(2024\), 2770–2824\.
- Jeanneret et al\.\(2022\)Guillaume Jeanneret, Loïc Simon, and Frédéric Jurie\. 2022\.Diffusion models for counterfactual explanations\. In*Asian Conference on Computer Vision \(ACCV\)*\.
- Lipman et al\.\(2023\)Yaron Lipman, Ricky T\. Q\. Chen, Heli Ben\-Hamu, Maximilian Nickel, and Matt Le\. 2023\.Flow matching for generative modeling\. In*International Conference on Learning Representations \(ICLR\)*\.
- Liu et al\.\(2023\)Xingchao Liu, Chengyue Gong, and Qiang Liu\. 2023\.Flow straight and fast: Learning to generate and transfer data with rectified flow\. In*International Conference on Learning Representations \(ICLR\)*\.
- Lundberg and Lee \(2017\)Scott M\. Lundberg and Su\-In Lee\. 2017\.A unified approach to interpreting model predictions\. In*Advances in Neural Information Processing Systems \(NeurIPS\)*\.
- Petsiuk et al\.\(2018\)Vitali Petsiuk, Abir Das, and Kate Saenko\. 2018\.RISE: Randomized input sampling for explanation of black\-box models\. In*British Machine Vision Conference \(BMVC\)*\.
- Schlegl et al\.\(2019\)Thomas Schlegl, Philipp Seeböck, Sebastian M\. Waldstein, Georg Langs, and Ursula Schmidt\-Erfurth\. 2019\.f\-AnoGAN: Fast unsupervised anomaly detection with generative adversarial networks\.*Medical Image Analysis*54 \(2019\), 30–44\.
- Selvaraju et al\.\(2017\)Ramprasaath R\. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra\. 2017\.Grad\-CAM: Visual explanations from deep networks via gradient\-based localization\. In*IEEE International Conference on Computer Vision \(ICCV\)*\. 618–626\.
- Shih et al\.\(2019\)George Shih, Carol C\. Wu, Safwan S\. Halabi, Marc D\. Kohli, Luciano M\. Prevedello, Tessa S\. Cook, Arjun Sharma, Judith K\. Amorosa, Veronica Arteaga, Maya Galperin\-Aizenberg, et al\.2019\.Augmenting the National Institutes of Health chest radiograph dataset with expert annotations of possible pneumonia\.*Radiology: Artificial Intelligence*1, 1 \(2019\), e180041\.
- Street et al\.\(1993\)W\. Nick Street, William H\. Wolberg, and Olvi L\. Mangasarian\. 1993\.Nuclear feature extraction for breast tumor diagnosis\. In*IS&T/SPIE Symposium on Electronic Imaging: Biomedical Image Processing and Biomedical Visualization*, Vol\. 1905\. 861–870\.
- Tong et al\.\(2024\)Alexander Tong, Kilian Fatras, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector\-Brooks, Guy Wolf, and Yoshua Bengio\. 2024\.Improving and generalizing flow\-based generative models with minibatch optimal transport\.*Transactions on Machine Learning Research \(TMLR\)*\(2024\)\.
- Tong et al\.\(2020\)Alexander Tong, Jessie Huang, Guy Wolf, David van Dijk, and Smita Krishnaswamy\. 2020\.TrajectoryNet: A dynamic optimal transport network for modeling cellular dynamics\. In*International Conference on Machine Learning \(ICML\)*\.
- Wachter et al\.\(2017\)Sandra Wachter, Brent Mittelstadt, and Chris Russell\. 2017\.Counterfactual explanations without opening the black box: Automated decisions and the GDPR\.*Harvard Journal of Law & Technology*31, 2 \(2017\), 841–887\.
- Wyatt et al\.\(2022\)Julian Wyatt, Adam Leach, Sebastian M\. Schmon, and Chris G\. Willcocks\. 2022\.AnoDDPM: Anomaly detection with denoising diffusion probabilistic models using simplex noise\. In*IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops \(CVPRW\)*\. 650–656\.
- Yang et al\.\(2023\)Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, and Bingbing Ni\. 2023\.MedMNIST v2: A large\-scale lightweight benchmark for 2D and 3D biomedical image classification\.*Scientific Data*10, 1 \(2023\), 41\.

Similar Articles

Multimodal Routing for Interpretable, Robust, and Auditable Clinical Prediction

arXiv cs.LG

This paper proposes an explicit multimodal routing framework for clinical prediction using EHR data, enabling interpretable, robust, and auditable reasoning across structured variables, clinical notes, and chest X-rays via discrete unimodal, bimodal, and trimodal routes with inference-time route masking for missing modality simulation.