相关性引导的流匹配与退火掩码用于空间转录组学生成

arXiv cs.LG 论文

摘要

本文提出了 CorrFlow,一个相关性引导的流匹配框架,用于从组织学图像预测空间转录组学,显式建模基因间依赖关系,以提高生成谱中的生物连贯性。

arXiv:2609.22187v1 Announce Type: new Abstract: Spatial transcriptomics (ST) provides spatially resolved gene expression profiling but remains expensive, motivating the prediction of ST from histology images. Generative models have emerged as a mainstream paradigm for ST prediction due to their ability to model the conditional distribution of gene expression and capture its inherent stochasticity. However, these methods typically treat genes as independent prediction targets and overlook the intrinsic gene-gene interactions in biological systems, which limits their ability to preserve biologically meaningful co-expression patterns. We argue that gene-gene interactions, which reflect shared pathways and regulatory mechanisms, are essential for generating numerically accurate and biologically coherent ST profiles. In this paper, we propose CorrFlow, a correlation-guided flow matching framework for histology-to-ST prediction that explicitly models gene-gene dependencies through two complementary mechanisms. First, we introduce an annealed masked flow matching strategy, where subsets of genes are progressively masked following a timestep-dependent annealing schedule, encouraging the model to infer masked genes conditioned on the remaining genes and promoting joint conditional modeling beyond per-gene marginal estimation. Second, we devise a gene graph-regularized optimization scheme that integrates prior knowledge from the STRING database and data-driven co-expression estimated by WGCNA to construct a gene affinity graph, which enforces both local consistency and global smoothness in the predicted expression. Extensive experiments across 12 datasets show that CorrFlow achieves the best average PCC and HPCC among evaluated methods, leading to more biologically coherent ST predictions.
查看原文
查看缓存全文

缓存时间: 2026/09/22 09:17

# Correlation-Guided Flow Matching with Annealed Masking for Spatial Transcriptomics Generation
Source: [https://arxiv.org/html/2609.22187](https://arxiv.org/html/2609.22187)
Hao ChenDepartment of Clinical NeurosciencesUniversity of CambridgeUKEmail:[hc666@cam\.ac\.uk](mailto:)Li PanUsher InstituteThe University of EdinburghEH16 RUXEmail:[l\.pan\-13@sms\.ed\.ac\.uk](mailto:)Chao LiDepartment of Clinical NeurosciencesUniversity of CambridgeUKEmail:[cl647@cam\.ac\.uk](mailto:)Xiaohan XingDepartment of Diagnostic RadiologyNational University of SingaporeSingaporeEmail:[xhxing@nus\.edu\.sg](mailto:)

###### Abstract

Spatial transcriptomics \(ST\) provides spatially resolved gene expression profiling but remains expensive, motivating the prediction of ST from histology images\. Generative models have emerged as a mainstream paradigm for ST prediction due to their ability to model the conditional distribution of gene expression and capture its inherent stochasticity\. However, these methods typically treat genes as independent prediction targets and overlook the intrinsic gene\-gene interactions in biological systems, which limits their ability to preserve biologically meaningful co\-expression patterns\. We argue that gene\-gene interactions, which reflect shared pathways and regulatory mechanisms, are essential for generating numerically accurate and biologically coherent ST profiles\. In this paper, we proposeCorrFlow, a correlation\-guided flow matching framework for histology\-to\-ST prediction that explicitly models gene\-gene dependencies through two complementary mechanisms\. First, we introduce anannealed masked flow matchingstrategy, where subsets of genes are progressively masked following a timestep\-dependent annealing schedule, encouraging the model to infer masked genes conditioned on the remaining genes and promoting joint conditional modeling beyond per\-gene marginal estimation\. Second, we devise agene graph\-regularized optimizationscheme that integrates prior knowledge from the STRING database and data\-driven co\-expression estimated by WGCNA to construct a gene affinity graph, which enforces both local consistency and global smoothness in the predicted expression\. Extensive experiments across 12 datasets show thatCorrFlowachieves the best average PCC and HPCC among evaluated methods, leading to more biologically coherent ST predictions\.

22footnotetext:Corresponding authors\.## 1Introduction

Spatial transcriptomics \(ST\) measures gene expression at spatially resolved locations across tissue, providing a molecular complement to histology and enabling applications such as spatial domain identification and tumor microenvironment analysis\[[20](https://arxiv.org/html/2609.22187#bib.bib20),[27](https://arxiv.org/html/2609.22187#bib.bib27),[9](https://arxiv.org/html/2609.22187#bib.bib9),[2](https://arxiv.org/html/2609.22187#bib.bib2),[3](https://arxiv.org/html/2609.22187#bib.bib3),[39](https://arxiv.org/html/2609.22187#bib.bib39),[34](https://arxiv.org/html/2609.22187#bib.bib34),[7](https://arxiv.org/html/2609.22187#bib.bib7)\]\. However, current ST assays remain expensive and low\-throughput, motivating efforts to predict ST from H&E\-stained histology images\[[30](https://arxiv.org/html/2609.22187#bib.bib30)\]\. Early approaches adopt deterministic regression\[[8](https://arxiv.org/html/2609.22187#bib.bib8),[38](https://arxiv.org/html/2609.22187#bib.bib38),[23](https://arxiv.org/html/2609.22187#bib.bib23),[6](https://arxiv.org/html/2609.22187#bib.bib6)\], learning a direct mapping from image patches to expression vectors, but failing to capture the inherent uncertainty and stochasticity of gene expression\. Retrieval\-based methods\[[35](https://arxiv.org/html/2609.22187#bib.bib35),[37](https://arxiv.org/html/2609.22187#bib.bib37)\]predict expression by matching query patches against a reference set but depend heavily on its diversity and coverage\. Recently, generative approaches\[[40](https://arxiv.org/html/2609.22187#bib.bib40),[11](https://arxiv.org/html/2609.22187#bib.bib11)\]have emerged as a promising paradigm for ST prediction by modeling the conditional distribution of gene expression and capturing its inherent stochasticity\.

Motivation\.Despite their advantages, current generative approaches\[[40](https://arxiv.org/html/2609.22187#bib.bib40),[11](https://arxiv.org/html/2609.22187#bib.bib11)\]leave a fundamental issue insufficiently addressed: genes are typically treated as independent targets, despite their rich and structured biological dependencies\. In reality, gene expression is governed by regulatory networks, co\-expression modules, and coordinated pathway activity, forming highly interdependent systems\[[21](https://arxiv.org/html/2609.22187#bib.bib21),[31](https://arxiv.org/html/2609.22187#bib.bib31),[36](https://arxiv.org/html/2609.22187#bib.bib36),[17](https://arxiv.org/html/2609.22187#bib.bib17),[28](https://arxiv.org/html/2609.22187#bib.bib28)\]\. However, most existing models lack explicit mechanisms to enforce such gene\-gene relationships\. We show in Section[3\.1](https://arxiv.org/html/2609.22187#S3.SS1)that this limitation is intrinsic to the training objective\. Although generative models define a joint distribution over genes, standard diffusion \(Fig\.[1](https://arxiv.org/html/2609.22187#S1.F1)Diffusion\) and flow matching \(Fig\.[1](https://arxiv.org/html/2609.22187#S1.F1)Flow Matching\) objectives decompose across dimensions, providing no explicit incentive to capture cross\-gene dependencies\. This can lead to uneven and uncoordinated performance across genes, where some genes are accurately predicted \(e\.g\.,x11x\_\{1\}^\{1\}andx12x\_\{1\}^\{2\}\) while other highly correlated genes \(e\.g\.,x13x\_\{1\}^\{3\}\) remain poorly estimated\. As a result, the model fails to effectively capture biologically meaningful correlation patterns, limiting biological fidelity and downstream utility\. This motivates explicitly modeling gene\-gene interactions within generative frameworks to promote more coherent and biologically consistent predictions\.

Figure 1:Correlation\-guided Flow Matching enables joint trajectory inference for correlated genes\.Given correlated genes with source states\{x0i\}i=13\\\{x\_\{0\}^\{i\}\\\}\_\{i=1\}^\{3\}, existing methods such as Diffusion and Flow Matching treat each trajectory independently, failing to capture inter\-gene dependencies and producing inconsistent predictionsx~13\\tilde\{x\}\_\{1\}^\{3\}\. Our method applies an annealed maskingℳ\\mathcal\{M\}and an explicit graph regularization, enforcing dependency across trajectories of genes11and22\. This dependency\-aware guidance steersx13x\_\{1\}^\{3\}toward a coherent joint prediction consistent with the target distribution\.In this paper, we proposeCorrFlow, a correlation\-guided flow matching framework for histology\-to\-ST prediction that explicitly models gene\-gene dependencies through two complementary mechanisms\. First, we propose*annealed masked flow matching*, where subsets of genes are progressively masked according to the flow timestepttand must be inferred from the remaining genes\. As the masking ratio increases withtt, the model is increasingly encouraged to exploit cross\-gene relationships rather than rely on independent per\-gene prediction, as illustrated in Fig\.[1](https://arxiv.org/html/2609.22187#S1.F1)\(Ours\)\. Second, we introduce a*gene graph\-regularized optimization*scheme to explicitly incorporate gene\-gene correlations\. We construct a gene affinity graph from two complementary sources: the STRING database\[[29](https://arxiv.org/html/2609.22187#bib.bib29)\], which provides functional associations between genes, and WGCNA\[[24](https://arxiv.org/html/2609.22187#bib.bib24)\], which estimates data\-driven gene co\-expression\. The fused graph is then used to regularize the predicted expression with two terms: a Huber\-based term for local consistency and a Laplacian term for global smoothness\. Together, these components constrain the output space and improve the biological coherence of the generated gene expression\. The contributions of this work are summarized as:

- •We identify a key limitation of existing generative models for ST prediction, their lack of explicit modeling of intrinsic gene\-gene correlations, and propose CorrFlow, a conditional flow matching framework that explicitly models gene dependencies for histology\-to\-ST prediction\.
- •We introduce anannealed masking strategyconditioned on the timesteptt, which progressively shifts the learning objective from per\-gene marginal estimation to joint conditional modeling\.
- •We introduce agene graph\-regularized optimizationscheme that integrates prior biological knowledge and data\-driven co\-expression graphs to explicitly enforce structured gene dependencies during ST generation\.
- •Across 12 ST datasets, CorrFlow achieves the best average PCC and HPCC among evaluated methods, highlighting the importance of gene dependency modeling in ST prediction\.

## 2Related Work

##### Histology\-to\-ST Prediction\.

Early work on predicting ST from histology largely formulated the task as supervised regression from image patches to spot\-level expression values, optionally combined with spatial coordinates\. ST\-Net\[[8](https://arxiv.org/html/2609.22187#bib.bib8)\]first demonstrated that spatial gene expression can be predicted from histology\. Subsequent regression\-based methods enhanced this pipeline with stronger spatial and contextual modeling to capture local and long\-range dependencies\[[38](https://arxiv.org/html/2609.22187#bib.bib38),[23](https://arxiv.org/html/2609.22187#bib.bib23),[6](https://arxiv.org/html/2609.22187#bib.bib6),[32](https://arxiv.org/html/2609.22187#bib.bib32)\]\. However, regression\-based approaches typically produce deterministic point estimates and fail to capture the inherent uncertainty and one\-to\-many nature of gene expression\. Retrieval\-based methods\[[35](https://arxiv.org/html/2609.22187#bib.bib35),[37](https://arxiv.org/html/2609.22187#bib.bib37),[33](https://arxiv.org/html/2609.22187#bib.bib33),[10](https://arxiv.org/html/2609.22187#bib.bib10)\]instead align histology and expression in a shared embedding space, and predict ST by retrieving similar patches from a reference set\. However, their performance depends heavily on the diversity and coverage of the reference set\.

More recently, there has been a shift toward generative modeling, motivated by its ability to capture uncertainty and the inherent one\-to\-many nature of gene expression\. For example, Stem\[[40](https://arxiv.org/html/2609.22187#bib.bib40)\]formulates expression prediction as conditional diffusion, and STFlow\[[11](https://arxiv.org/html/2609.22187#bib.bib11)\]as conditional flow matching over whole\-slide spot collections, jointly modeling expression across spots with improved sampling efficiency\. GenAR\[[22](https://arxiv.org/html/2609.22187#bib.bib22)\]generates expression autoregressively over gene groups, and STPath\[[12](https://arxiv.org/html/2609.22187#bib.bib12)\]adopts mask\-based generative pretraining\. However, existing methods primarily optimize per\-gene predictions or marginal distributions without explicitly modeling gene\-gene dependencies, limiting their ability to preserve biologically meaningful co\-expression structure\. This limitation motivates our dependency\-aware generative framework for ST prediction\.

##### Gene\-Gene Interaction Modeling\.

Gene expression is organized by co\-expression modules, regulatory programs, and interaction networks\. Gene\-gene interactions can be modeled using either prior\-based or data\-driven approaches\. For instance, WGCNA\[[24](https://arxiv.org/html/2609.22187#bib.bib24)\]constructs gene co\-expression networks in a data\-driven manner by identifying modules of highly correlated genes\. STRING\[[29](https://arxiv.org/html/2609.22187#bib.bib29)\]provides curated and predicted protein\-protein interaction networks that capture functional associations between genes\. Incorporating such structures has been shown to improve molecular prediction and interpretation\[[25](https://arxiv.org/html/2609.22187#bib.bib25)\]\. This idea is especially appealing for histology\-to\-ST generation, where biologically plausible outputs should preserve meaningful cross\-gene structure rather than merely optimize per\-gene accuracy\[[10](https://arxiv.org/html/2609.22187#bib.bib10)\]\. However, such priors are typically applied in regression models or feature encoders, and are rarely integrated into the dynamics of conditional generative learning\. Our work takes a step toward incorporating structured gene dependencies into conditional generative frameworks, enabling improved biological fidelity and coherence\.

## 3Methodology

Motivated by the per\-gene decomposition limitation identified in Section[3\.1](https://arxiv.org/html/2609.22187#S3.SS1), we introduce CorrFlow, a dependency\-aware conditional flow matching framework for histology\-to\-ST prediction\. As shown in Fig\.[2](https://arxiv.org/html/2609.22187#S3.F2), CorrFlow combinesAnnealed Masked Flow Matching\(Section[3\.2](https://arxiv.org/html/2609.22187#S3.SS2)\) andGene Graph\-Regularized Optimization\(Section[3\.3](https://arxiv.org/html/2609.22187#S3.SS3)\) to encourage cross\-gene modeling\. Given histology featureszz, spatial coordinatesqq, timesteptt, and the interpolantxt=\(1−t\)​x0\+t​x1x\_\{t\}=\(1\-t\)x\_\{0\}\+tx\_\{1\}between a zero\-inflated negative binomial \(ZINB\) prior samplex0x\_\{0\}and ground\-truth expressionx1x\_\{1\}, we obtain the masked inputx^t\\hat\{x\}\_\{t\}using the timestep\-dependent masking ratiop⁡\(t\)=pmax⋅tp\(t\)=p\_\{\\max\}\\cdot t\. A spatial\-transformer denoiser\[[11](https://arxiv.org/html/2609.22187#bib.bib11)\]then predicts the target expression from\(z,q,t,x^t\)\(z,q,t,\\hat\{x\}\_\{t\}\)and is optimized with both reconstruction and gene graph\-based regularization\.

### 3\.1Theoretical Background

We begin by formalizing a structural limitation of standard flow matching\[[18](https://arxiv.org/html/2609.22187#bib.bib18),[1](https://arxiv.org/html/2609.22187#bib.bib1),[19](https://arxiv.org/html/2609.22187#bib.bib19)\]when applied to multi\-gene expression generation\.

Setup and Notation\.Letx1∈ℝGx\_\{1\}\\in\\mathbb\{R\}^\{G\}denote a gene expression vector withGGgenes at a single spatial spot, and letccdenote the conditioning information \(histology features, spatial coordinates\)\. In conditional flow matching \(CFM\), we draw a prior samplex0∼p0x\_\{0\}\\sim p\_\{0\}and form the linear interpolant

xt=\(1−t\)​x0\+t​x1,t∈\[0,1\]\.x\_\{t\}=\(1\-t\)\\,x\_\{0\}\+t\\,x\_\{1\},\\qquad t\\in\[0,1\]\.\(1\)A velocity networkvθ​\(xt,t,c\)∈ℝGv\_\{\\theta\}\(x\_\{t\},t,c\)\\in\\mathbb\{R\}^\{G\}is trained to predict the conditional velocityut=x1−x0u\_\{t\}=x\_\{1\}\-x\_\{0\}via

ℒCFM​\(θ\)=𝔼t,x0,x1,c​‖vθ​\(xt,t,c\)−\(x1−x0\)‖2\.\\mathcal\{L\}\_\{\\mathrm\{CFM\}\}\(\\theta\)=\\mathbb\{E\}\_\{t,\\,x\_\{0\},\\,x\_\{1\},\\,c\}\\bigl\\\|v\_\{\\theta\}\(x\_\{t\},t,c\)\-\(x\_\{1\}\-x\_\{0\}\)\\bigr\\\|^\{2\}\.\(2\)In practice, many methods predict endpointx1x\_\{1\}directly rather than the velocity; the two formulations are equivalent up to the affine relationx1=x0\+utx\_\{1\}=x\_\{0\}\+u\_\{t\}\. We use the endpoint\-prediction form below\.

![Refer to caption](https://arxiv.org/html/2609.22187v1/framework.png)Figure 2:Overview of the proposed CorrFlow\.We model histology\-to\-ST inference as conditional flow matching conditioned on histology features and spatial coordinates\. To promote gene\-dependent generation, we introduce \(a\) anannealed masking scheduleconditioned on timestep t and \(b\) agene graph\-regularized optimizationstrategy, which constructs gene affinity by fusing an external functional prior with data\-driven co\-expression, and imposes a Huber\-based*local smoothness*and a Laplacian*global smoothness*to improve the biological coherence of generated expression\.###### Proposition 1\(Per\-gene decomposition of the standard objective\)\.

Letvθ​\(xt,t,c\)=\(vθ1,…,vθG\)v\_\{\\theta\}\(x\_\{t\},t,c\)=\(v\_\{\\theta\}^\{1\},\\dots,v\_\{\\theta\}^\{G\}\)denote theGG\-dimensional output of the velocity network\. The objective \([2](https://arxiv.org/html/2609.22187#S3.E2)\) decomposes as

ℒCFM​\(θ\)=∑g=1G𝔼t,x0,x1,c​\(vθg​\(xt,t,c\)−utg\)2=∑g=1GℒCFM\(g\)​\(θ\),\\mathcal\{L\}\_\{\\mathrm\{CFM\}\}\(\\theta\)=\\sum\_\{g=1\}^\{G\}\\mathbb\{E\}\_\{t,\\,x\_\{0\},\\,x\_\{1\},\\,c\}\\bigl\(v\_\{\\theta\}^\{g\}\(x\_\{t\},t,c\)\-u\_\{t\}^\{g\}\\bigr\)^\{2\}=\\sum\_\{g=1\}^\{G\}\\mathcal\{L\}\_\{\\mathrm\{CFM\}\}^\{\(g\)\}\(\\theta\),\(3\)whereutg=x1g−x0gu\_\{t\}^\{g\}=x\_\{1\}^\{g\}\-x\_\{0\}^\{g\}is thegg\-th component of the target velocity\. Consequently, the optimal velocity field satisfies, for each geneg∈\{1,…,G\}g\\in\\\{1,\\dots,G\\\}independently,

v∗,g\(xt,t,c\)=𝔼\[utg∣xt,t,c\]=𝔼\[x1g∣xt,t,c\]−x0g\.v^\{\*,g\}\(x\_\{t\},t,c\)=\\mathbb\{E\}\\bigl\[u\_\{t\}^\{g\}\\mid x\_\{t\},t,c\\bigr\]=\\mathbb\{E\}\\bigl\[x\_\{1\}^\{g\}\\mid x\_\{t\},t,c\\bigr\]\-x\_\{0\}^\{g\}\.\(4\)

###### Proof\.

The squaredℓ2\\ell\_\{2\}norm decomposes over coordinates:‖vθ−ut‖2=∑g\(vθg−utg\)2\\\|v\_\{\\theta\}\-u\_\{t\}\\\|^\{2\}=\\sum\_\{g\}\(v\_\{\\theta\}^\{g\}\-u\_\{t\}^\{g\}\)^\{2\}\. Because expectation is linear,ℒCFM\\mathcal\{L\}\_\{\\mathrm\{CFM\}\}is the sum ofGGindependent per\-gene losses\. Minimizing eachℒCFM\(g\)\\mathcal\{L\}\_\{\\mathrm\{CFM\}\}^\{\(g\)\}pointwise yieldsv∗,g=𝔼\[utg∣xt,t,c\]v^\{\*,g\}=\\mathbb\{E\}\[u\_\{t\}^\{g\}\\mid x\_\{t\},t,c\], which depends onx1gx\_\{1\}^\{g\}only through its marginal conditional distribution given\(xt,t,c\)\(x\_\{t\},t,c\)\. ∎

Proposition[1](https://arxiv.org/html/2609.22187#Thmproposition1)does not claim that the network*cannot*represent cross\-gene dependencies; a shared backbone with sufficient capacity can, in principle, share intermediate representations across genes\. The proposition states something fundamental: the*training objective*never rewards such sharing\. Any improvement in genegg’s prediction that arises from modeling its correlation with genejjis equally achievable by a predictor that ignores genejjentirely, because the loss for geneggdepends only on the marginalp⁡\(x1g∣xt,t,c\)p\(x\_\{1\}^\{g\}\\mid x\_\{t\},t,c\)\. Cross\-gene coordination is therefore*unidentifiable*under the standard objective that it may emerge incidentally but is never incentivized\.

### 3\.2Annealed Masked Flow Matching

The decomposition in Proposition[1](https://arxiv.org/html/2609.22187#Thmproposition1)arises because every gene dimension of the targetutu\_\{t\}has a corresponding gene dimension in the inputxtx\_\{t\}, allowing each gene to be predicted independently\. We proposedannealed masked flow matching\(Annealed\-MFM\) to break this correspondence: at each step, a subset of gene dimensions is withheld from the input, requiring the model to infer the missing genes from the remaining ones\. We show that this simple modification provably shifts the optimal solution from per\-gene marginals to conditional joint distributions\.

Formulation\.Letℳ⊆\{1,…,G\}\\mathcal\{M\}\\subseteq\\\{1,\\dots,G\\\}denote a mask set sampled from a distributionq⁡\(ℳ\)q\(\\mathcal\{M\}\)at each training step, and let𝒪=\{1,…,G\}∖ℳ\\mathcal\{O\}=\\\{1,\\dots,G\\\}\\setminus\\mathcal\{M\}\. The model receives a partially masked inputx~t\\tilde\{x\}\_\{t\}in which the entries forℳ\\mathcal\{M\}are replaced by learnable gene\-wise mask tokensm∈ℝGm\\in\\mathbb\{R\}^\{G\}, shared across spatial locations:

x~t,g=\{mg,g∈ℳ,xt,g,g∈𝒪\.\\tilde\{x\}\_\{t,g\}=\\begin\{cases\}m\_\{g\},&g\\in\\mathcal\{M\},\\\\ x\_\{t,g\},&g\\in\\mathcal\{O\}\.\\end\{cases\}\(5\)
Crucially, masking operates exclusively on the*input*xt→x~tx\_\{t\}\\to\\tilde\{x\}\_\{t\}; the*target*utu\_\{t\}and the*supervision signal*remain unchanged across allGGgenes\. The Annealed\-MFM objective is therefore

ℒMFM​\(θ,ℳ\)=𝔼t,x0,x1,c​∑g=1G\(vθg​\(x~t,t,c\)−utg\)2\.\\mathcal\{L\}\_\{\\mathrm\{MFM\}\}\(\\theta;\\mathcal\{M\}\)=\\mathbb\{E\}\_\{t,\\,x\_\{0\},\\,x\_\{1\},\\,c\}\\;\\sum\_\{g=1\}^\{G\}\\bigl\(v\_\{\\theta\}^\{g\}\(\\tilde\{x\}\_\{t\},t,c\)\-u\_\{t\}^\{g\}\\bigr\)^\{2\}\.\(6\)
###### Proposition 2\(Joint conditional modeling under masking\)\.

Although the loss supervises allGGgenes equally, the nature of the prediction task differs between the two groups, producing the asymmetry that shifts the optimal solution for the masked genes from per\-gene marginals to the joint conditional\. Under the masked objective \([6](https://arxiv.org/html/2609.22187#S3.E6)\), the optimal predictor for*every*geneg∈\{1,…,G\}g\\in\\\{1,\\dots,G\\\}satisfies

v∗,g\(x~t,t,c\)=𝔼\[utg\|xt\(𝒪\),t,c\]\.v^\{\*,g\}\(\\tilde\{x\}\_\{t\},t,c\)=\\mathbb\{E\}\\bigl\[u\_\{t\}^\{g\}\\;\\big\|\\;x\_\{t\}^\{\(\\mathcal\{O\}\)\},\\,t,\\,c\\bigr\]\.\(7\)For observed genesg∈𝒪g\\in\\mathcal\{O\}, the conditioning setxt\(𝒪\)x\_\{t\}^\{\(\\mathcal\{O\}\)\}includesxt,gx\_\{t,g\}itself, so the predictor retains direct information aboutx1,gx\_\{1,g\}and behaves similarly to standard flow matching\. For masked genesg∈ℳg\\in\\mathcal\{M\}, the conditioning set contains no information aboutxt,gx\_\{t,g\}; the optimal prediction instead depends on thejoint conditional distributionp⁡\(x1\(ℳ\)∣x1\(𝒪\),c\)p\(x\_\{1\}^\{\(\\mathcal\{M\}\)\}\\mid x\_\{1\}^\{\(\\mathcal\{O\}\)\},c\)\. If genes inℳ\\mathcal\{M\}are correlated with those in𝒪\\mathcal\{O\}, the optimal prediction for any masked geneggdepends on the co\-variation structure between the two sets\.

We provide the proof of Proposition[2](https://arxiv.org/html/2609.22187#Thmproposition2)in the Appendix Section[A\.1](https://arxiv.org/html/2609.22187#A1.SS1)\.

#### 3\.2\.1Implementation Details

Annealed masking schedule\.Conditional flow matching suffers from a practical shortcut issue: ast→1t\\to 1, the interpolantxt→x1x\_\{t\}\\to x\_\{1\}, allowing the network to shortcut by copying near\-target values rather than learning meaningful transport dynamics\. We suppress this by coupling the masking ratio to the timestep:

p⁡\(t\)=pmax⋅t,p\(t\)=p\_\{\\max\}\\cdot t,\(8\)wherepmax∈\(0,1\)p\_\{\\max\}\\in\(0,1\)\. At smalltt,xtx\_\{t\}is far fromx1x\_\{1\}and little masking is needed \(x~t≈xt\\tilde\{x\}\_\{t\}\\approx x\_\{t\}\); at largett, heavy masking removes the shortcut signal \(x~t≈m\\tilde\{x\}\_\{t\}\\approx m\), forcing the network to reconstructutu\_\{t\}from cross\-gene context rather than copying\. In practice, we samplet∼Uniform⁡\(0,1\)t\\sim\\mathrm\{Uniform\}\(0,1\)and draw a Bernoulli mask independently per gene with probabilityp⁡\(t\)p\(t\)\.

Learnable mask tokens\.A natural masking strategy is to replace masked genes with a fixed zero token \(mg=0m\_\{g\}=0\)\. However, this introduces a systematic shift in input magnitude: during training, a fractionp⁡\(t\)p\(t\)of entries are zeroed, reducing the average input norm by a factor of approximately1−p⁡\(t\)1\-p\(t\)\. At inference, no masking is applied and all entries carry real values, resulting in a train\-inference magnitude mismatch analogous to the dropout scaling problem\[[26](https://arxiv.org/html/2609.22187#bib.bib26)\]\.

To mitigate this issue, we introduce a learnable mask tokenm∈ℝGm\\in\\mathbb\{R\}^\{G\}, initialized to zero and optimized via backpropagation\. Rather than acting as a fixed placeholder, the token adapts to preserve the input statistics expected by the network across different masking ratios\. In effect, the network jointly learns*what*to substitute for missing genes \(the token value\) and*how*to use the remaining genes \(the velocity predictor\), enabling stable optimization without explicit rescaling during inference\. We provide the convergence analysis of Annealed\-MFM in the Appendix Section[A\.2](https://arxiv.org/html/2609.22187#A1.SS2)\.

### 3\.3Gene Graph\-Regularized Optimization

Annealed\-MFM addresses the*optimization*limitation identified in Proposition[1](https://arxiv.org/html/2609.22187#Thmproposition1)by restructuring the training signal to encourage joint modeling\. However, finite model capacity and noisy gradients can still lead to predictions that violate known co\-expression structures\. To address this, we introduce a complementary mechanism that imposes constraints based on external functional prior and data\-driven co\-expression, guiding the model toward biologically coherent outputs\.

#### 3\.3\.1Gene Affinity Graph Construction

We construct a gene affinity graph𝒢=\(𝒱,ℰ,S\)\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\},S\)over the target gene set, where𝒱\\mathcal\{V\}is the set of genes,ℰ\\mathcal\{E\}is the set of edges connecting functionally related genes, andS∈ℝG×GS\\in\\mathbb\{R\}^\{G\\times G\}is the affinity matrix encoding edge weights\. The graph fuses two complementary sources: STRING\[[29](https://arxiv.org/html/2609.22187#bib.bib29)\]captures functional gene\-gene associations from existing databases, providing a stable prior that generalizes across datasets; we denote its affinity matrix byS\(s\)S^\{\(\\mathrm\{s\}\)\}\. WGCNA\-derived graph\[[24](https://arxiv.org/html/2609.22187#bib.bib24)\]computed from the training expression data, capturing the co\-expression landscape specific to the current dataset; we denote its affinity matrix byS\(w\)S^\{\(\\mathrm\{w\}\)\}\. We fuse them via weighted averaging:

S=α​S\(s\)\+\(1−α\)​S\(w\),S=\\alpha\\,S^\{\(\\mathrm\{s\}\)\}\+\(1\-\\alpha\)\\,S^\{\(\\mathrm\{w\}\)\},\(9\)whereS\(s\),S\(w\)∈ℝG×GS^\{\(\\mathrm\{s\}\)\},S^\{\(\\mathrm\{w\}\)\}\\in\\mathbb\{R\}^\{G\\times G\}, andα∈\[0,1\]\\alpha\\in\[0,1\]controls the balance\. The fused matrix is sparsified by retaining the top\-kkneighbors per gene and symmetrized to yield the final edge setℰ\\mathcal\{E\}\. FromSSwe derive the normalized adjacency and graph Laplacian:

L=I−A,A=D−1/2SD−1/2,L=I\-A,\\qquad A=D^\{\-1/2\}S\\,D^\{\-1/2\},\(10\)whereDDis the diagonal degree matrix withDi​i=∑jSi​jD\_\{ii\}=\\sum\_\{j\}S\_\{ij\},AAis the symmetrically normalized adjacency matrix, andLLis the corresponding normalized graph Laplacian\.

#### 3\.3\.2Graph\-Regularized Objective

We use the gene affinity graph𝒢\\mathcal\{G\}to regularize the predicted expression, treating co\-expression structure as a soft prior on valid outputs\. The overall objective for graph constraintsℒgraph\\mathcal\{L\}\_\{\\rm graph\}combines two complementary regularizers operating at local and global scales\.

Local graph constraint\.Connected genes in𝒢\\mathcal\{G\}should produce similar predictions\. Since genes operate at different scales, we z\-score normalize each geneiiprediction using running EMA estimates of its mean and standard deviation \(μi\\mu\_\{i\},σi\\sigma\_\{i\}\)\[[15](https://arxiv.org/html/2609.22187#bib.bib15)\]:zi=\(x^1,i−μi\)/σi\.z\_\{i\}=\(\\hat\{x\}\_\{1,i\}\-\\mu\_\{i\}\)/\\sigma\_\{i\}\.For each edge\(i,j\)∈ℰ\(i,j\)\\in\\mathcal\{E\}, we penalize the discrepancyΔn,i​j=zn​i−zn​j\\Delta\_\{n,ij\}=z\_\{ni\}\-z\_\{nj\}at every spotnn:

ℒlocal=1N​∑n=1N∑\(i,j\)∈ℰwi​j​Huberβ​\(Δn,i​j\),wi​j∝Si​j\.\\mathcal\{L\}\_\{\\mathrm\{local\}\}=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}\\sum\_\{\(i,j\)\\in\\mathcal\{E\}\}w\_\{ij\}\\;\\mathrm\{Huber\}\_\{\\beta\}\(\\Delta\_\{n,ij\}\),\\qquad w\_\{ij\}\\propto S\_\{ij\}\.\(11\)The Huber penalty\[[13](https://arxiv.org/html/2609.22187#bib.bib13)\]with thresholdβ\\betais quadratic for\|Δ\|≤β\|\\Delta\|\\leq\\betaand linear for\|Δ\|\>β\|\\Delta\|\>\\beta, promoting smoothness among co\-regulated genes while remaining robust to biological variation on weak edges\. With the constraint ofℒlocal\\mathcal\{L\}\_\{\\mathrm\{local\}\}, graph\-connected genes with larger affinitieswi​jw\_\{ij\}are more strongly coupled, encouraging more consistent standardized expression patterns\.

Global graph constraint\.The local termℒlocal\\mathcal\{L\}\_\{\\mathrm\{local\}\}operates on individual edges and cannot capture graph\-wide patterns\. We complement it with a spectral penalty via the normalized graph Laplacian \(Eq\.[10](https://arxiv.org/html/2609.22187#S3.E10)\), whereSSis the affinity matrix andDDis the diagonal degree matrix withDi​i=∑jSi​jD\_\{ii\}=\\sum\_\{j\}S\_\{ij\}:

ℒglobal=1N​∑n=1N𝐱^1,n⊤​L​𝐱^1,n\.\\mathcal\{L\}\_\{\\mathrm\{global\}\}=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}\\hat\{\\mathbf\{x\}\}\_\{1,n\}^\{\\top\}L\\;\\hat\{\\mathbf\{x\}\}\_\{1,n\}\.\(12\)Intuitively,ℒglobal\\mathcal\{L\}\_\{\\mathrm\{global\}\}penalizes predictions that oscillate rapidly between neighboring genes, encouraging the output to align with the low\-frequency eigenmodes of𝒢\\mathcal\{G\}, which correspond to large\-scale co\-expression modules\.

Combined graph constraint\.The two terms operate at complementary scales:ℒlocal\\mathcal\{L\}\_\{\\mathrm\{local\}\}enforces pairwise consistency, whileℒglobal\\mathcal\{L\}\_\{\\mathrm\{global\}\}suppresses graph\-wide oscillations\. We combine them asℒgraph=ρ​ℒlocal\+λ​ℒglobal,\\mathcal\{L\}\_\{\\mathrm\{graph\}\}=\\rho\\,\\mathcal\{L\}\_\{\\mathrm\{local\}\}\+\\lambda\\,\\mathcal\{L\}\_\{\\mathrm\{global\}\},whereρ\\rhoandλ\\lambdacontrol the local and global regularization strengths\. Detailed hyperparameter analysis is provided in Appendix Table[6](https://arxiv.org/html/2609.22187#A1.T6)\. The overall training objective combines the masked flow matching loss \(i\.e\.,ℒMFM\\mathcal\{L\}\_\{\\mathrm\{MFM\}\}in Eq\.[6](https://arxiv.org/html/2609.22187#S3.E6)\) with the graph constraintℒgraph\\mathcal\{L\}\_\{\\mathrm\{graph\}\}\.

## 4Experiments and Results

### 4\.1Experimental Setup

Datasets\.We evaluated our method on two public ST collections: HEST\-1K\[[16](https://arxiv.org/html/2609.22187#bib.bib16)\]and STImage\-1K4M\[[4](https://arxiv.org/html/2609.22187#bib.bib4)\]\. Our main evaluation is conducted on the 10 benchmark datasets from HEST\-1K, including IDC, PRAD, PAAD, SKCM, COAD, READ, CCRCC, HCC, LUNG, and LYMPH\-IDC \(denoted as LYMPH\)\. HEST\-1K\[[16](https://arxiv.org/html/2609.22187#bib.bib16)\]is used as the primary benchmark because it provides standardized patient\-stratified splits, resulting in a leakage\-controlled cross\-validation setting across multiple cancer types and tissue contexts\. In addition, we conducted supplementary experiments on the two largest human cancer categories from STImage\-1K4M\[[4](https://arxiv.org/html/2609.22187#bib.bib4)\], Brain and Breast, to assess scalability on larger ST collections\. Since STImage\-1K4M does not provide standardized patient\-level benchmark splits comparable to HEST\-1K for these categories, we followed a random train/validation/test split of 8:1:1 and report the corresponding results in Appendix Table[9](https://arxiv.org/html/2609.22187#A1.T9)\.

Implementation Details\.Following prior work\[[40](https://arxiv.org/html/2609.22187#bib.bib40)\], we selected the top 50 genes \(denoted as HMHVG\-50\) from the intersection of highly expressed and highly variable genes\. UNI\[[5](https://arxiv.org/html/2609.22187#bib.bib5)\]was selected as the histology feature extractor, following prior work\[[11](https://arxiv.org/html/2609.22187#bib.bib11)\]\. The default optimization setup used a batch size of 2, learning rate5×10−45\\times 10^\{\-4\}, 100 epochs,pmaxp\_\{\\rm max\}as 0\.75, andα\\alphaas 0\.6,ρ\\rhoas 0\.3, andλ\\lambdaas10−310^\{\-3\}\. All experiments are conducted on two NVIDIA RTX A5000 GPUs\. We evaluate ST prediction using Pearson correlation \(PCC\) and introduce Hallmark\-Gene PCC \(HPCC\), a pathway\-informed extension of gene\-wise PCC\. Given per\-gene PCCrgr\_\{g\}, target genes𝒢\\mathcal\{G\}, and MSigDB Hallmark gene sets𝒫=\{Pk\}k=1K\\mathcal\{P\}=\\\{P\_\{k\}\\\}\_\{k=1\}^\{K\}\[[14](https://arxiv.org/html/2609.22187#bib.bib14)\], we define

HPCC=1\|𝒦\|​∑k∈𝒦1\|Pk∩𝒢\|​∑g∈Pk∩𝒢rg,𝒦=\{k:\|Pk∩𝒢\|≥5\}\.\\mathrm\{HPCC\}=\\frac\{1\}\{\|\\mathcal\{K\}\|\}\\sum\_\{k\\in\\mathcal\{K\}\}\\frac\{1\}\{\|P\_\{k\}\\cap\\mathcal\{G\}\|\}\\sum\_\{g\\in P\_\{k\}\\cap\\mathcal\{G\}\}r\_\{g\},\\quad\\mathcal\{K\}=\\\{k:\\,\|P\_\{k\}\\cap\\mathcal\{G\}\|\\geq 5\\\}\.\(13\)HPCC averages gene\-wise prediction accuracy over Hallmark gene sets with sufficient coverage in the evaluated target panel, providing a biologically informed complement to standard PCC\. In addition to PCC and HPCC, we report complementary results on GGC\-Pearson and MSE\-Mean in Appendix[C\.2](https://arxiv.org/html/2609.22187#A3.SS2)\. These additional metrics assess gene\-gene correlation fidelity and numerical reconstruction error, respectively\.

Baselines\.We compared our model against representative histology\-to\-ST baselines spanning complementary modeling paradigms\. These include ST\-Net\[[8](https://arxiv.org/html/2609.22187#bib.bib8)\]as an early supervised regression\-based method, BLEEP\[[35](https://arxiv.org/html/2609.22187#bib.bib35)\]as a contrastive learning and retrieval\-based approach, TRIPLEX\[[6](https://arxiv.org/html/2609.22187#bib.bib6)\]for context\-aware modeling with global and neighborhood histology features, Stem\[[40](https://arxiv.org/html/2609.22187#bib.bib40)\]as a diffusion\-based generative model that can leverage multiple foundation\-model\-derived features, and STFlow\[[11](https://arxiv.org/html/2609.22187#bib.bib11)\]as a flow\-matching\-based generative model\. Due to the high computational cost of Stem\[[40](https://arxiv.org/html/2609.22187#bib.bib40)\], we provide a resource\-bounded reproduction in the Appendix Table[11](https://arxiv.org/html/2609.22187#A1.T11)\. Specifically, we train Stem\[[40](https://arxiv.org/html/2609.22187#bib.bib40)\]for 400 epochs and use 200 sampling steps at inference, while keeping remaining settings consistent with the original implementation\. Other baselines are reproduced using unified training with 100 epochs for a fair comparison\. Model complexity is reported in Section[C\.7](https://arxiv.org/html/2609.22187#A3.SS7)\.

Table 1:Comparison across the 10 HEST\-1K benchmark datasets\. PCC and HPCC on HMHVG\-50 genes are reported\. Values are shown asmeanstd\\mathrm\{mean\}\_\{\\mathrm\{std\}\}across folds/seeds\.DatasetPCC↑\\uparrowHPCC↑\\uparrowST\-NetBLEEPBLEEP\-UNITRIPLEXSTFlowOursST\-NetBLEEPBLEEP\-UNITRIPLEXSTFlowOursIDC0\.715\.0960\.608\.2270\.767\.0520\.753\.0590\.790\.0600\.791\.0610\.710\.1030\.605\.2250\.769\.0560\.751\.0620\.787\.0660\.788\.072PRAD0\.406\.0080\.345\.0380\.451\.0140\.412\.0420\.494\.0290\.495\.0200\.507\.0440\.420\.0010\.582\.0170\.501\.0330\.599\.0540\.610\.042PAAD0\.504\.0420\.465\.0630\.543\.0520\.516\.0560\.574\.0500\.580\.0520\.535\.0600\.503\.0770\.599\.0530\.567\.0630\.609\.0510\.622\.053SKCM0\.706\.0530\.645\.0880\.758\.0110\.748\.0020\.813\.0300\.817\.0250\.691\.0470\.648\.0070\.660\.0690\.663\.0630\.707\.1260\.732\.123COAD0\.567\.0300\.254\.0670\.648\.0310\.579\.0350\.649\.0550\.665\.0410\.423\.0020\.165\.0330\.553\.0910\.535\.0270\.574\.0600\.598\.028READ0\.190\.1140\.231\.1320\.408\.0880\.208\.1360\.428\.0510\.449\.0360\.253\.0380\.216\.0660\.329\.0550\.186\.0270\.418\.0680\.444\.100CCRCC0\.317\.0710\.306\.0580\.391\.0690\.291\.1660\.457\.0560\.470\.0590\.332\.0750\.312\.0710\.385\.0800\.317\.1790\.469\.0590\.476\.065HCC0\.291\.0440\.121\.1110\.256\.0580\.086\.0700\.382\.1040\.433\.1330\.461\.0600\.208\.1140\.398\.0750\.118\.1070\.541\.1060\.564\.087LUNG0\.734\.0170\.704\.0170\.757\.0040\.750\.0040\.782\.0080\.783\.0050\.768\.0190\.737\.0170\.779\.0140\.782\.0120\.814\.0030\.815\.000LYMPH0\.545\.1300\.511\.1090\.564\.1180\.594\.1040\.629\.1280\.644\.123––––––Average0\.4970\.4190\.5540\.4940\.6000\.6130\.5200\.4240\.5620\.4910\.6130\.628

Table 2:Ablation study of our methods on the HMHVG\-50\. Average PCC across the 10 HEST\-1K benchmark datasets is reported; per\-dataset results are provided in Appendix Table[5](https://arxiv.org/html/2609.22187#A1.T5)\.Metricw/ow/omaskingw/ow/oℒgraph\\mathcal\{L\}\_\{\\rm graph\}w/ow/oℒglobal\\mathcal\{L\}\_\{\\rm global\}w/ow/oℒlocal\\mathcal\{L\}\_\{\\rm local\}w/ow/oWGCNAw/ow/oSTRINGOursAverage PCC0\.6090\.6050\.6090\.6050\.6070\.6040\.613

### 4\.2Experimental Results

##### Quantitative Comparison with Baselines\.

Table[1](https://arxiv.org/html/2609.22187#S4.T1)compares our method with five histology\-to\-ST approaches on the 10 HEST\-1K benchmarks\. Our method achieves the best average performance on both PCC and HPCC, improving over the strongest existing baseline STFlow from0\.6000\.600to0\.6130\.613in PCC and from0\.6130\.613to0\.6280\.628in HPCC\. The gain is particularly meaningful because STFlow already substantially outperforms earlier spot\-based and slide\-level baselines, making this a relatively strong comparison setting\. The stronger gain on HPCC suggests that explicitly modeling inter\-gene dependencies is beneficial not only for individual gene prediction but also for recovering biologically coordinated expression programs\. Across expanded HMHVG\-100/150/200 settings as shown in the Appendix Table[4](https://arxiv.org/html/2609.22187#A1.T4), our method consistently outperforms STFlow in average PCC and HPCC, indicating that its dependency\-aware design remains effective as the target gene set broadens\.

Table 3:Hyperparameter study on masking strategy, maximum masking ratio, and alpha for graph fusion \(PCC on HMHVG\-50\)\.DatasetMasking Strategy \(Learnable\)Maximum Masking RatioAlpha for Graph Fusionf⁡\(t\)=0\.75f\(t\)=0\.75f⁡\(t\)=1−tf\(t\)=1\-tf⁡\(t\)=tf\(t\)=tpm​a​x=0\.15p\_\{max\}=0\.15pm​a​x=0\.3p\_\{max\}=0\.3pm​a​x=0\.5p\_\{max\}=0\.5pm​a​x=0\.75p\_\{max\}=0\.75pm​a​x=0\.9p\_\{max\}=0\.9α=0\.4\\alpha=0\.4α=0\.5\\alpha=0\.5α=0\.6\\alpha=0\.6α=0\.7\\alpha=0\.7IDC0\.791\.0620\.792\.0610\.791\.0610\.795\.0580\.788\.0610\.791\.0500\.791\.0610\.787\.0590\.791\.0610\.791\.0580\.791\.0610\.790\.054PRAD0\.499\.0210\.494\.0200\.495\.0200\.490\.0230\.493\.0160\.489\.0070\.495\.0200\.494\.0210\.499\.0240\.494\.0210\.495\.0200\.492\.028PAAD0\.580\.0500\.579\.0440\.580\.0520\.577\.0530\.582\.0520\.570\.0520\.580\.0520\.564\.0580\.569\.0550\.576\.0500\.580\.0520\.582\.052SKCM0\.817\.0260\.808\.0350\.817\.0250\.818\.0170\.807\.0340\.818\.0260\.817\.0250\.816\.0250\.816\.0250\.817\.0250\.817\.0250\.813\.029COAD0\.666\.0410\.666\.0410\.665\.0410\.662\.0430\.653\.0550\.664\.0400\.665\.0410\.665\.0400\.665\.0400\.665\.0410\.665\.0410\.670\.046READ0\.438\.0350\.416\.0440\.449\.0360\.421\.0480\.416\.0480\.422\.0490\.449\.0360\.417\.0570\.438\.0460\.430\.0380\.449\.0360\.446\.026CCRCC0\.465\.0610\.462\.0530\.470\.0590\.456\.0520\.462\.0530\.446\.0670\.470\.0590\.468\.0560\.445\.0730\.457\.0560\.470\.0590\.456\.064HCC0\.431\.1350\.431\.1340\.433\.1330\.430\.1300\.432\.1320\.433\.1340\.433\.1330\.432\.1320\.432\.1320\.432\.1330\.433\.1330\.432\.131LUNG0\.781\.0070\.781\.0090\.783\.0050\.785\.0050\.782\.0070\.782\.0070\.783\.0050\.784\.0040\.785\.0040\.784\.0040\.783\.0050\.782\.007LYMPH0\.630\.1200\.631\.1280\.644\.1230\.637\.1270\.633\.1330\.638\.1220\.644\.1230\.629\.1270\.628\.1330\.634\.1280\.644\.1230\.633\.131Average0\.6100\.6060\.6130\.6070\.6050\.6050\.6130\.6060\.6070\.6080\.6130\.609

Furthermore, statistical tests across all pairwise comparisons with competing methods yield p\-values consistently below 0\.05, indicating that our method significantly outperforms all baselines\. Details are reported in Appendix Section[C\.5](https://arxiv.org/html/2609.22187#A3.SS5)and Section[C\.6](https://arxiv.org/html/2609.22187#A3.SS6)\.

##### Ablation Study\.

Table[2](https://arxiv.org/html/2609.22187#S4.T2)summarizes the ablation study of our method, reporting average PCC for HMHVG\-50 over the 10 HEST\-1K benchmarks\. The full model achieves the best average PCC of 0\.613, suggesting that the proposed components are complementary\. Removing annealed masking reduces the average PCC to 0\.609, indicating that masking provides a useful training signal for preventing direct shortcut reconstruction and encouraging dependency\-aware prediction\. Removing the graph\-based regularization term yields a larger decrease to 0\.605, supporting the role of graph constraints in stabilizing gene\-expression prediction under structured inter\-gene dependencies\. We further examined the construction of the gene graph by removing either WGCNA\-derived co\-expression graph or STRING\-derived functional interaction priors\. Both variants underperform the full model, with average PCCs of 0\.607 and 0\.604, respectively, suggesting that dataset\-specific co\-expression structure and external biological priors provide complementary information for graph construction\. The per\-dataset results are provided in Appendix Table[5](https://arxiv.org/html/2609.22187#A1.T5)\.

##### Hyperparameter Study\.

We conducted the hyperparameter studies with respect to the masking schedule, masking intensity, and graph\-prior fusion weight, as summarized in Table[3](https://arxiv.org/html/2609.22187#S4.T3)\. Among the timestep\-dependent masking strategies, the proportional schedulef⁡\(t\)=tf\(t\)=tgives the strongest average performance, whereas the inverse schedulef⁡\(t\)=1−tf\(t\)=1\-tperforms worst\. This ordering supports our motivation that stronger masking is most effective near the expression endpoint, where the model is more prone to shortcut reconstruction due to stronger target\-expression signals in the interpolated input\. Varying the maximum masking ratio reveals a similar balance:pmax=0\.75p\_\{\\max\}=0\.75performs best, indicating that effective masking should be strong enough to expose inter\-gene dependencies but not so aggressive that it removes the contextual signals required for stable prediction\. For the hybrid gene graph, the best result is obtained withα=0\.6\\alpha=0\.6, suggesting a modest preference toward the STRING\-derived functional prior while preserving the dataset\-specific co\-expression structure captured by WGCNA\-derived graph\. The hyperparameter study for the coefficientsρ\\rhoandλ\\lambdain the combined graph constraint introduced in Section[3\.3\.2](https://arxiv.org/html/2609.22187#S3.SS3.SSS2)is reported in Appendix Table[6](https://arxiv.org/html/2609.22187#A1.T6)\.

![Refer to caption](https://arxiv.org/html/2609.22187v1/vis_lymph_idc.png)Figure 3:Visualization on LYMPH\_IDC\. \(a\) Visualization of a biologically relevant gene KRT7 on sample NCBI68, with histology image, ground\-truth ST, and predictions of CorrFlow and STFlow \(individual min\-max color scales to reflect expression intensity\)\. \(b\) Gene PCC comparison between CorrFlow and STFlow \(each point denotes a gene and points above the diagonal indicate improved prediction by CorrFlow\)\. \(c\) PCC improvements on the top\-8 genes by average absolute correlation\.
##### Qualitative Visualization\.

To qualitatively assess CorrFlow, we visualize predictions from three perspectives\. In all cases, our method is compared against STFlow under the same test splits and gene panels\. We first visualize theprediction of biologically relevant genes\. As shown in Fig\.[3](https://arxiv.org/html/2609.22187#S4.F3)\(a\), for the selected marker gene KRT7 in the LYMPH dataset, we show the histology image, the ground\-truth ST expression, and the predictions from our method and STFlow\. This comparison highlights whether the predicted ST recovers the major spatial gradients, localized high\-expression regions, and tissue\-specific heterogeneity observed in the histology image and ground truth\. As shown in Fig\.[3](https://arxiv.org/html/2609.22187#S4.F3)\(a\), our method generally produces spatial patterns that are visually closer to the ground truth, with clear spatial patterns\. More results on other datasets are presented in the Appendix Fig\.[4](https://arxiv.org/html/2609.22187#A1.F4)\. Fig\.[3](https://arxiv.org/html/2609.22187#S4.F3)\(b\) presents agene\-level PCC scatter comparison\. Each point corresponds to one gene, where the x\-axis denotes the PCC achieved by STFlow and the y\-axis denotes the PCC achieved by our method\. Points above the diagonal indicate genes for which our method outperforms STFlow\. This visualization provides a global view of gene\-level performance differences and reveals whether improvements are broadly distributed or concentrated in a subset of genes\. As shown in Fig\.[3](https://arxiv.org/html/2609.22187#S4.F3)\(b\), a large proportion of genes lie above the diagonal, consistent with the overall quantitative gains reported in the benchmark results\. Fig\.[3](https://arxiv.org/html/2609.22187#S4.F3)\(c\) further visualizes thetop\-kkgenes with high gene\-gene correlation\. To focus on genes that are strongly coupled with the overall transcriptomic program, we ranked genes using the ground\-truth gene\-gene correlation matrix from the test set\. Specifically, for each genegg, we computed the average absolute correlation with all other genes:AvgAbsCorr⁡\(g\)=1G−1​∑h≠g\|Corr⁡\(g,h\)\|\.\\mathrm\{AvgAbsCorr\}\(g\)=\\frac\{1\}\{G\-1\}\\sum\_\{h\\neq g\}\\left\|\\mathrm\{Corr\}\(g,h\)\\right\|\.We then selected the top\-kkgenes with the highestAvgAbsCorr\\mathrm\{AvgAbsCorr\}values and compared the PCC achieved by our method and STFlow on these genes\. As shown in Fig\.[3](https://arxiv.org/html/2609.22187#S4.F3)\(c\), our method consistently achieves higher PCC on many highly correlated genes, suggesting improved preservation of dependency\-relevant expression structure\. These qualitative trends are consistent with the observed improvements in PCC, further supporting the effectiveness of our method\.

## 5Conclusion

We present CorrFlow, a correlation\-guided conditional flow matching framework for histology\-to\-ST generation\. Rather than treating genes as independent prediction targets, CorrFlow explicitly promotes cross\-gene joint modeling throughannealed masked flow matchingandgene graph\-regularized optimization\. The annealed masking strategy encourages the model to infer masked gene expression from related genes, while the gene affinity graph integrates external functional priors and data\-driven co\-expression structure to impose local consistency and global smoothness on generated ST profiles\. Across 10 patient\-stratified HEST\-1K benchmarks and two supplementary STImage\-1K4M datasets, CorrFlow achieves the best average PCC and HPCC among evaluated methods\. These results suggest that modeling intrinsic gene\-gene dependencies is a promising direction for more biologically coherent histology\-to\-ST prediction\. The limitations and future work are presented in the Appendix Section[D](https://arxiv.org/html/2609.22187#A4)\.

## References

- \[1\]Albergo, M\.S\., Vanden\-Eijnden, E\.: Building normalizing flows with stochastic interpolants\. arXiv preprint arXiv:2209\.15571 \(2022\)
- \[2\]Biancalani, T\., Scalia, G\., Buffoni, L\., Avasthi, R\., Lu, Z\., Sanger, A\., Tokcan, N\., Vanderburg, C\.R\., Segerstolpe, Å\., Zhang, M\., et al\.: Deep learning and alignment of spatially resolved single\-cell transcriptomes with tangram\. Nature methods18\(11\), 1352–1362 \(2021\)
- \[3\]Chen, J\., Luo, T\., Jiang, M\., Liu, J\., Gupta, G\.P\., Li, Y\.: Cell composition inference and identification of layer\-specific spatial transcriptional profiles with polaris\. Science Advances9\(9\), eadd9818 \(2023\)
- \[4\]Chen, J\., Zhou, M\., Wu, W\., Zhang, J\., Li, Y\., Li, D\.: Stimage\-1k4m: A histopathology image\-gene expression dataset for spatial transcriptomics\. Advances in neural information processing systems37, 35796–35823 \(2024\)
- \[5\]Chen, R\.J\., Ding, T\., Lu, M\.Y\., Williamson, D\.F\., Jaume, G\., Song, A\.H\., Chen, B\., Zhang, A\., Shao, D\., Shaban, M\., et al\.: Towards a general\-purpose foundation model for computational pathology\. Nature medicine30\(3\), 850–862 \(2024\)
- \[6\]Chung, Y\., Ha, J\.H\., Im, K\.C\., Lee, J\.S\.: Accurate spatial gene expression prediction by integrating multi\-resolution features\. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition\. pp\. 11591–11600 \(2024\)
- \[7\]De Visser, K\.E\., Joyce, J\.A\.: The evolving tumor microenvironment: From cancer initiation to metastatic outgrowth\. Cancer cell41\(3\), 374–403 \(2023\)
- \[8\]He, B\., Bergenstråhle, L\., Stenbeck, L\., Abid, A\., Andersson, A\., Borg, Å\., Maaskola, J\., Lundeberg, J\., Zou, J\.: Integrating spatial gene expression and breast tumour morphology via deep learning\. Nature biomedical engineering4\(8\), 827–834 \(2020\)
- \[9\]Hu, J\., Li, X\., Coleman, K\., Schroeder, A\., Ma, N\., Irwin, D\.J\., Lee, E\.B\., Shinohara, R\.T\., Li, M\.: Spagcn: Integrating gene expression, spatial location and histology to identify spatial domains and spatially variable genes by graph convolutional network\. Nature methods18\(11\), 1342–1351 \(2021\)
- \[10\]Hu, S\., Zeng, Q\., Bhasker, N\., Kather, J\.N\., Speidel, S\.: Histoprism: Unlocking functional pathway analysis from pan\-cancer histology via gene expression prediction\. arXiv preprint arXiv:2601\.21560 \(2026\)
- \[11\]Huang, T\., Liu, T\., Babadi, M\., Jin, W\., Ying, Z\.: Scalable generation of spatial transcriptomics from histology images via whole\-slide flow matching\. In: Singh, A\., Fazel, M\., Hsu, D\., Lacoste\-Julien, S\., Berkenkamp, F\., Maharaj, T\., Wagstaff, K\., Zhu, J\. \(eds\.\) Proceedings of the 42nd International Conference on Machine Learning\. Proceedings of Machine Learning Research, vol\. 267, pp\. 25550–25565\. PMLR \(13–19 Jul 2025\),[https://proceedings\.mlr\.press/v267/huang25t\.html](https://proceedings.mlr.press/v267/huang25t.html)
- \[12\]Huang, T\., Liu, T\., Babadi, M\., Ying, R\., Jin, W\.: Stpath: a generative foundation model for integrating spatial transcriptomics and whole\-slide images\. NPJ Digital Medicine8\(1\), 659 \(2025\)
- \[13\]Huber, P\.J\.: Robust estimation of a location parameter\. In: Breakthroughs in statistics: Methodology and distribution, pp\. 492–518\. Springer \(1992\)
- \[14\]Institute, B\.: Molecular signatures database \(msigdb\) \(2025\),[https://www\.gsea\-msigdb\.org/gsea/msigdb](https://www.gsea-msigdb.org/gsea/msigdb), accessed: 2025\-09\-10
- \[15\]Ioffe, S\., Szegedy, C\.: Batch normalization: Accelerating deep network training by reducing internal covariate shift\. In: International conference on machine learning\. pp\. 448–456\. pmlr \(2015\)
- \[16\]Jaume, G\., Doucet, P\., Song, A\.H\., Lu, M\.Y\., Almagro\-Pérez, C\., Wagner, S\.J\., Vaidya, A\.J\., Chen, R\.J\., Williamson, D\.F\., Kim, A\., et al\.: Hest\-1k: A dataset for spatial transcriptomics and histology image analysis\. Advances in Neural Information Processing Systems37, 53798–53833 \(2024\)
- \[17\]Jaume, G\., Vaidya, A\., Chen, R\.J\., Williamson, D\.F\., Liang, P\.P\., Mahmood, F\.: Modeling dense multimodal interactions between biological pathways and histology for survival prediction\. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition\. pp\. 11579–11590 \(2024\)
- \[18\]Lipman, Y\., Chen, R\.T\.Q\., Ben\-Hamu, H\., Nickel, M\., Le, M\.: Flow matching for generative modeling \(2023\),[https://arxiv\.org/abs/2210\.02747](https://arxiv.org/abs/2210.02747)
- \[19\]Liu, X\., Gong, C\., Liu, Q\.: Flow straight and fast: Learning to generate and transfer data with rectified flow\. arXiv preprint arXiv:2209\.03003 \(2022\)
- \[20\]Marx, V\.: Method of the year: spatially resolved transcriptomics\. Nature methods18\(1\), 9–14 \(2021\)
- \[21\]Muzio, G\., O’Bray, L\., Borgwardt, K\.: Biological network analysis with deep learning\. Briefings in bioinformatics22\(2\), 1515–1530 \(2021\)
- \[22\]Ouyang, J\., Wang, Y\., Gao, Y\., Xu, Y\., Yang, S\., Chen, H\.: Genar: Next\-scale autoregressive generation for spatial gene expression prediction\. arXiv preprint arXiv:2510\.04315 \(2025\)
- \[23\]Pang, M\., Su, K\., Li, M\.: Leveraging information in spatial transcriptomics to predict super\-resolution gene expression from histology images in tumors\. BioRxiv pp\. 2021–11 \(2021\)
- \[24\]Rezaie, N\., Reese, F\., Mortazavi, A\.: Pywgcna: a python package for weighted gene co\-expression network analysis\. Bioinformatics39\(7\), btad415 \(2023\)
- \[25\]Shi, H\., Chi, C\., Wan, P\., Zhang, D\., Shao, W\.: Multi\-modal topology\-embedded graph learning for spatially resolved genes prediction from pathology images with prior gene similarity information\. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition\. pp\. 20810–20819 \(2025\)
- \[26\]Srivastava, N\., Hinton, G\., Krizhevsky, A\., Sutskever, I\., Salakhutdinov, R\.: Dropout: a simple way to prevent neural networks from overfitting\. The journal of machine learning research15\(1\), 1929–1958 \(2014\)
- \[27\]Ståhl, P\.L\., Salmén, F\., Vickovic, S\., Lundmark, A\., Navarro, J\.F\., Magnusson, J\., Giacomello, S\., Asp, M\., Westholm, J\.O\., Huss, M\., et al\.: Visualization and analysis of gene expression in tissue sections by spatial transcriptomics\. Science353\(6294\), 78–82 \(2016\)
- \[28\]Sun, E\.D\., Ma, R\., Zou, J\.: Sprite: improving spatial gene expression imputation with gene and cell networks\. Bioinformatics40\(Supplement\_1\), i521–i528 \(2024\)
- \[29\]Szklarczyk, D\., Nastou, K\., Koutrouli, M\., Kirsch, R\., Mehryary, F\., Hachilif, R\., Hu, D\., Peluso, M\.E\., Huang, Q\., Fang, T\., et al\.: The string database in 2025: protein networks with directionality of regulation\. Nucleic acids research53\(D1\), D730–D737 \(2025\)
- \[30\]Wang, C\., Chan, A\.S\., Fu, X\., Ghazanfar, S\., Kim, J\., Patrick, E\., Yang, J\.Y\.: Benchmarking the translational potential of spatial gene expression prediction from histology\. Nature Communications16\(1\), 1544 \(2025\)
- \[31\]Wang, H\., Sham, P\., Tong, T\., Pang, H\.: Pathway\-based single\-cell rna\-seq classification, clustering, and construction of gene\-gene interactions networks using random forests\. IEEE Journal of Biomedical and Health Informatics24\(6\), 1814–1822 \(2019\)
- \[32\]Wang, H\., Du, X\., Liu, J\., Ouyang, S\., Chen, Y\.W\., Lin, L\.: M2ost: Many\-to\-one regression for predicting spatial transcriptomics from digital pathology images\. In: Proceedings of the AAAI Conference on Artificial Intelligence\. vol\. 39, pp\. 7709–7717 \(2025\)
- \[33\]Wang, Y\., Wang, J\., Xu, Y\., Liu, N\., Liu, B\., Li, Y\., Yu, G\.: Fmh2st: foundation model\-based spatial transcriptomics generation from histological images\. Nucleic Acids Research53\(17\), gkaf865 \(2025\)
- \[34\]Xiao, Y\., Yu, D\.: Tumor microenvironment as a therapeutic target in cancer\. Pharmacology & therapeutics221, 107753 \(2021\)
- \[35\]Xie, R\., Pang, K\., Chung, S\., Perciani, C\., MacParland, S\., Wang, B\., Bader, G\.: Spatially resolved gene expression prediction from histology images via bi\-modal contrastive learning\. Advances in Neural Information Processing Systems36, 70626–70637 \(2023\)
- \[36\]Yan, R\., Zhang, X\., Jiang, Z\., Wang, B\., Bian, X\., Ren, F\., Zhou, S\.K\.: Pathway\-aware multimodal transformer \(pamt\): Integrating pathological image and gene expression for interpretable cancer survival analysis\. IEEE Transactions on Pattern Analysis and Machine Intelligence \(2025\)
- \[37\]Yang, Y\., Hossain, M\.Z\., Stone, E\.A\., Rahman, S\.: Exemplar guided deep neural network for spatial transcriptomics analysis of gene expression prediction\. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision\. pp\. 5039–5048 \(2023\)
- \[38\]Zeng, Y\., Wei, Z\., Yu, W\., Yin, R\., Yuan, Y\., Li, B\., Tang, Z\., Lu, Y\., Yang, Y\.: Spatial transcriptomics prediction from histology jointly through transformer and graph neural networks\. Briefings in Bioinformatics23\(5\), bbac297 \(2022\)
- \[39\]Zhang, D\., Schroeder, A\., Yan, H\., Yang, H\., Hu, J\., Lee, M\.Y\., Cho, K\.S\., Susztak, K\., Xu, G\.X\., Feldman, M\.D\., et al\.: Inferring super\-resolution tissue architecture by integrating spatial transcriptomics with histology\. Nature biotechnology42\(9\), 1372–1377 \(2024\)
- \[40\]Zhu, S\., Zhu, Y\., Tao, M\., Qiu, P\.: Diffusion generative modeling for spatially resolved gene expression inference from histology images\. In: Yue, Y\., Garg, A\., Peng, N\., Sha, F\., Yu, R\. \(eds\.\) International Conference on Learning Representations\. vol\. 2025, pp\. 19666–19690 \(2025\),[https://proceedings\.iclr\.cc/paper\_files/paper/2025/file/31cc93d156b4ab1514d71ca5147a6e67\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2025/file/31cc93d156b4ab1514d71ca5147a6e67-Paper-Conference.pdf)

Appendix

## Appendix AMethodology Extension

### A\.1Proof of Proposition 2\.

###### Proof\.

The masked inputx~t\\tilde\{x\}\_\{t\}is shared across all output dimensions, andx~t\\tilde\{x\}\_\{t\}restricted toℳ\\mathcal\{M\}contains only fixed mask tokens carrying no sample\-specific information, so the optimal predictor for every geneggisv∗,g=𝔼\[utg∣xt\(𝒪\),t,c\]v^\{\*,g\}=\\mathbb\{E\}\[u\_\{t\}^\{g\}\\mid x\_\{t\}^\{\(\\mathcal\{O\}\)\},t,c\]\. Forg∈ℳg\\in\\mathcal\{M\}, expanding via the law of total expectation:

𝔼\[utg∣xt\(𝒪\),t,c\]=∫utgp\(x1\(ℳ\)∣xt\(𝒪\),t,c\)dx1\(ℳ\)\.\\mathbb\{E\}\\bigl\[u\_\{t\}^\{g\}\\mid x\_\{t\}^\{\(\\mathcal\{O\}\)\},t,c\\bigr\]=\\int u\_\{t\}^\{g\}\\;p\\bigl\(x\_\{1\}^\{\(\\mathcal\{M\}\)\}\\mid x\_\{t\}^\{\(\\mathcal\{O\}\)\},t,c\\bigr\)\\;\\mathrm\{d\}x\_\{1\}^\{\(\\mathcal\{M\}\)\}\.\(14\)This integral can only be estimated if the model captures the co\-variation betweenx1\(ℳ\)x\_\{1\}^\{\(\\mathcal\{M\}\)\}andx1\(𝒪\)x\_\{1\}^\{\(\\mathcal\{O\}\)\}; without such knowledge the model has no basis for predictingutgu\_\{t\}^\{g\}beyond the unconditional prior\. ∎

### A\.2Convergence analysis\.

A natural concern is whether training under the masked objectiveℒMFM\\mathcal\{L\}\_\{\\mathrm\{MFM\}\}still yields a valid generative model, given that the network is optimized on corrupted inputsx~t\\tilde\{x\}\_\{t\}but deployed on clean inputsxtx\_\{t\}at inference\. We show that minimizingℒMFM\\mathcal\{L\}\_\{\\mathrm\{MFM\}\}guarantees convergence ofℒCFM\\mathcal\{L\}\_\{\\mathrm\{CFM\}\}, which governs inference\-time generation\.

###### Theorem 1\(Convergence of Annealed\-MFM\)\.

Letq⁡\(ℳ\)q\(\\mathcal\{M\}\)be any masking distribution over subsets of\{1,…,G\}\\\{1,\\dots,G\\\}\. Then:

1. \(i\)Lower bound\.ℒMFM​\(θ,ℳ\)≥ℒCFM​\(θ\)\\mathcal\{L\}\_\{\\mathrm\{MFM\}\}\(\\theta;\\mathcal\{M\}\)\\geq\\mathcal\{L\}\_\{\\mathrm\{CFM\}\}\(\\theta\)for allθ\\thetaand allℳ\\mathcal\{M\}\.
2. \(ii\)Shared minimizer\.There existsθ∗\\theta^\{\*\}that minimizesℒMFM\\mathcal\{L\}\_\{\\mathrm\{MFM\}\}and simultaneously drivesℒCFM\\mathcal\{L\}\_\{\\mathrm\{CFM\}\}to its global minimum\. Consequently,ℒMFM​\(θ,ℳ\)→0\\;\\mathcal\{L\}\_\{\\mathrm\{MFM\}\}\(\\theta;\\mathcal\{M\}\)\\to 0lead toℒCFM​\(θ\)→0\\mathcal\{L\}\_\{\\mathrm\{CFM\}\}\(\\theta\)\\to 0, and the learned flow converges to the target distribution under standard regularity conditions\[[18](https://arxiv.org/html/2609.22187#bib.bib18)\]\.

###### Proof\.

\(i\)Both losses supervise allGGgenes against the same targetutu\_\{t\}and differ only in the network input \(x~t\\tilde\{x\}\_\{t\}vs\.xtx\_\{t\}\)\. Sincex~t\\tilde\{x\}\_\{t\}is a deterministic degradation ofxtx\_\{t\}, the data\-processing inequality gives, for each genegg,

𝔼⁡\[\(vθg​\(x~t,t,c\)−utg\)2\]≥𝔼⁡\[\(vθg​\(xt,t,c\)−utg\)2\]\.\\mathbb\{E\}\\bigl\[\(v\_\{\\theta\}^\{g\}\(\\tilde\{x\}\_\{t\},t,c\)\-u\_\{t\}^\{g\}\)^\{2\}\\bigr\]\\;\\geq\\;\\mathbb\{E\}\\bigl\[\(v\_\{\\theta\}^\{g\}\(x\_\{t\},t,c\)\-u\_\{t\}^\{g\}\)^\{2\}\\bigr\]\.\(15\)Summing overggyields the bound\.

\(ii\)Letθ∗\\theta^\{\*\}satisfyvθ∗\(x,t,c\)=𝔼\[ut∣x,t,c\]v\_\{\\theta^\{\*\}\}\(x,t,c\)=\\mathbb\{E\}\[u\_\{t\}\\mid x,t,c\]for all inputsxxover a sufficiently expressive function\. Thenvθ∗\(x~t,t,c\)=𝔼\[ut∣x~t,t,c\]v\_\{\\theta^\{\*\}\}\(\\tilde\{x\}\_\{t\},t,c\)=\\mathbb\{E\}\[u\_\{t\}\\mid\\tilde\{x\}\_\{t\},t,c\], which is the pointwise minimizer ofℒMFM\\mathcal\{L\}\_\{\\mathrm\{MFM\}\}\. Sinceθ∗\\theta^\{\*\}simultaneously minimizes both objectives, andℒCFM​\(θ∗\)=0\\mathcal\{L\}\_\{\\mathrm\{CFM\}\}\(\\theta^\{\*\}\)=0under standard flow matching guarantees\[[18](https://arxiv.org/html/2609.22187#bib.bib18)\], the global minimum ofℒMFM\\mathcal\{L\}\_\{\\mathrm\{MFM\}\}also drivesℒCFM\\mathcal\{L\}\_\{\\mathrm\{CFM\}\}to zero\. ∎

### A\.3Exploratory Extension: Graph\-Structured Masking

Proposition[2](https://arxiv.org/html/2609.22187#Thmproposition2)establishes that masking induces joint modeling, but the quality of the learning signal depends on*which*genes are masked together\. Under independent Bernoulli masking, masked genes are scattered across unrelated pathways, and the joint conditionalp⁡\(x1\(ℳ\)∣x1\(𝒪\),c\)p\(x\_\{1\}^\{\(\\mathcal\{M\}\)\}\\mid x\_\{1\}^\{\(\\mathcal\{O\}\)\},c\)factorizes approximately, weakening the inter\-gene signal that masking is designed to exploit\.

Instead, we leverage the gene affinity graph𝒢\\mathcal\{G\}to construct structured mask sets that respect biological modularity\. At each training step, we sample a set of seed genes uniformly at random and expand each seed to a connected subgraph viarr\-hop neighborhood sampling on𝒢\\mathcal\{G\}\. This ensures co\-regulated genes are masked*together*, forcing the model to reconstruct entire expression modules from regulatory context and maximizing the joint\-modeling benefit of Proposition[2](https://arxiv.org/html/2609.22187#Thmproposition2)\.

Parameters\.The number of seeds is set to⌈p⁡\(t\)⋅G/n¯r⌉\\lceil p\(t\)\\cdot G/\\bar\{n\}\_\{r\}\\rceil, wheren¯r\\bar\{n\}\_\{r\}is the averagerr\-hop neighborhood size in𝒢\\mathcal\{G\}, so that the expected number of masked genes matchesp⁡\(t\)⋅Gp\(t\)\\cdot G\. We fixr=1r=1throughout all experiments; the masking ratiop⁡\(t\)p\(t\)then controls the number of seeds rather than the neighborhood depth\. If sampled neighborhoods overlap, the union is taken as the final mask setℳ\\mathcal\{M\}\.

Results\.Results and analysis are provided in the Appendix Table[10](https://arxiv.org/html/2609.22187#A1.T10)and Appendix Section[C\.4](https://arxiv.org/html/2609.22187#A3.SS4)\.

### A\.4Training Procedure

The overall training objective combines the masked flow matching loss \(Eq\.[6](https://arxiv.org/html/2609.22187#S3.E6)\) with the graph regularizer \(Eq\.[3\.3\.2](https://arxiv.org/html/2609.22187#S3.SS3.SSS2)\):

ℒ=ℒMFM\+ℒgraph\.\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{MFM\}\}\+\\mathcal\{L\}\_\{\\mathrm\{graph\}\}\.\(16\)ℒMFM\\mathcal\{L\}\_\{\\mathrm\{MFM\}\}supervises allGGgenes under partial observation, incentivizing joint modeling \(Proposition[2](https://arxiv.org/html/2609.22187#Thmproposition2)\);ℒgraph\\mathcal\{L\}\_\{\\mathrm\{graph\}\}constrains the output toward biologically coherent expression profiles \(Section[3\.3\.2](https://arxiv.org/html/2609.22187#S3.SS3.SSS2)\)\. The full procedure is summarized in Algorithm[1](https://arxiv.org/html/2609.22187#alg1)\.

Algorithm 1CorrFlow Training1:Gene affinity graph

𝒢\\mathcal\{G\}, max masking ratio

pmaxp\_\{\\max\}, learning rate

η\\eta,

x0∼p0x\_\{0\}\\sim p\_\{0\},

x1∼p1x\_\{1\}\\sim p\_\{1\}
2:foreach training iterationdo

3:

xt←\(1−t\)​x0\+t​x1x\_\{t\}\\leftarrow\(1\-t\)\\,x\_\{0\}\+t\\,x\_\{1\},

t∼Uniform⁡\(0,1\)t\\sim\\mathrm\{Uniform\}\(0,1\)⊳\\trianglerightInterpolation

4:

ℳ∼q⁡\(ℳ∣𝒢,p⁡\(t\)\)\\mathcal\{M\}\\sim q\(\\mathcal\{M\}\\mid\\mathcal\{G\},\\,p\(t\)\),

p⁡\(t\)←pmax⋅tp\(t\)\\leftarrow p\_\{\\max\}\\cdot t⊳\\trianglerightGraph mask \(§[A\.3](https://arxiv.org/html/2609.22187#A1.SS3)\), Eq\.[8](https://arxiv.org/html/2609.22187#S3.E8)

5:

x~t,g←\{mgg∈ℳxt,gg∉ℳ\\tilde\{x\}\_\{t,g\}\\leftarrow\\begin\{cases\}m\_\{g\}&g\\in\\mathcal\{M\}\\\\ x\_\{t,g\}&g\\notin\\mathcal\{M\}\\end\{cases\}⊳\\trianglerightMasking, Eq\.[5](https://arxiv.org/html/2609.22187#S3.E5)

6:

vt←vθ​\(x~t,t,c\)v\_\{t\}\\leftarrow v\_\{\\theta\}\(\\tilde\{x\}\_\{t\},t,c\)⊳\\trianglerightPrediction

7:

ℒ←‖vt−\(x1−x0\)‖2\+ρ​ℒlocal\+λ​ℒglobal\\mathcal\{L\}\\leftarrow\\\|v\_\{t\}\-\(x\_\{1\}\{\-\}x\_\{0\}\)\\\|^\{2\}\+\\rho\\,\\mathcal\{L\}\_\{\\mathrm\{local\}\}\+\\lambda\\,\\mathcal\{L\}\_\{\\mathrm\{global\}\}⊳\\trianglerightEq\.[6](https://arxiv.org/html/2609.22187#S3.E6), Eq\.[3\.3\.2](https://arxiv.org/html/2609.22187#S3.SS3.SSS2)

8:

θ←θ−η​∇θ\(ℒ\)\\theta\\leftarrow\\theta\-\\eta\\,\\nabla\_\{\\theta\}\(\\mathcal\{L\}\)⊳\\trianglerightUpdate

9:endfor

Table 4:Comparison between STFlow and ours on HEST\-1K benchmarks under HMHVG\-100/150/200 settings, with PCC and HPCC reported\.DatasetPCC\-100↑\\uparrowPCC\-150↑\\uparrowPCC\-200↑\\uparrowHPCC\-100↑\\uparrowHPCC\-150↑\\uparrowHPCC\-200↑\\uparrowSTFlowOursSTFlowOursSTFlowOursSTFlowOursSTFlowOursSTFlowOursIDC0\.772\.0500\.778\.0520\.735\.0480\.741\.0530\.707\.0520\.705\.0550\.777\.0400\.786\.0420\.756\.0390\.760\.0470\.722\.0480\.719\.052PRAD0\.480\.0260\.475\.0050\.475\.0280\.479\.0080\.473\.0240\.473\.0070\.497\.0300\.484\.0040\.484\.0250\.487\.0010\.470\.0190\.466\.001PAAD––––––––––––SKCM0\.793\.0200\.805\.0140\.801\.0030\.810\.0020\.788\.0090\.789\.0010\.809\.0350\.820\.0310\.812\.0040\.822\.0050\.797\.0100\.799\.004COAD0\.602\.0610\.604\.0530\.580\.0560\.581\.0640\.547\.0630\.557\.0490\.615\.0630\.607\.0510\.570\.0750\.570\.0770\.553\.0810\.560\.065READ0\.368\.0430\.372\.0430\.336\.0320\.356\.0210\.337\.0380\.349\.0530\.364\.0130\.371\.0310\.328\.0030\.351\.0250\.342\.0040\.348\.027CCRCC0\.450\.0490\.463\.0520\.436\.0620\.433\.0580\.433\.0540\.442\.0540\.456\.0370\.462\.0510\.460\.0580\.451\.0570\.447\.0560\.456\.055HCC0\.392\.0910\.425\.1130\.329\.0470\.353\.0660\.356\.0630\.378\.0820\.393\.0910\.402\.0980\.328\.0430\.336\.0550\.379\.0640\.384\.079LUNG0\.731\.0100\.730\.0120\.685\.0130\.690\.014––0\.747\.0220\.747\.0240\.716\.0190\.721\.021––LYMPH0\.668\.1180\.670\.1180\.639\.1010\.647\.0980\.661\.1020\.655\.1070\.690\.1090\.690\.1110\.665\.0950\.675\.0890\.674\.0950\.670\.099Average0\.5840\.5910\.5570\.5660\.5380\.5430\.5940\.5970\.5690\.5750\.5480\.550

Table 5:Ablation study of different graph priors and module design \(PCC on HMHVG\-50\)\.AblationIDCPRADPAADSKCMCOADREADCCRCCHCCLUNGLYMPHAvg\.w/ow/omasking0\.790\.0630\.490\.0290\.579\.0490\.813\.0210\.666\.0480\.434\.0570\.461\.0550\.428\.1220\.780\.0080\.645\.1230\.609w/ow/oℒgraph\\mathcal\{L\}\_\{\\rm graph\}0\.794\.0590\.496\.0330\.577\.0490\.818\.0290\.647\.0500\.451\.0380\.460\.0490\.398\.0980\.781\.0080\.632\.1250\.605w/ow/oℒLocal\\mathcal\{L\}\_\{\\rm Local\}0\.789\.0540\.497\.0210\.574\.0560\.811\.0260\.649\.0480\.451\.0380\.462\.0480\.408\.1070\.780\.0090\.633\.1280\.605w/ow/oℒGlobal\\mathcal\{L\}\_\{\\rm Global\}0\.792\.0600\.497\.0220\.573\.0580\.815\.0260\.663\.0420\.452\.0190\.459\.0580\.430\.1310\.782\.0070\.631\.1330\.609w/ow/oWGCNA0\.792\.0560\.491\.0230\.576\.0510\.805\.0340\.650\.0530\.448\.0300\.459\.0520\.436\.1280\.783\.0060\.630\.1230\.607w/ow/oSTRING0\.788\.0630\.502\.0200\.575\.0500\.809\.0370\.651\.0540\.425\.0450\.456\.0610\.430\.1320\.781\.0060\.626\.1270\.604Ours0\.791\.0610\.495\.0200\.580\.0520\.817\.0250\.665\.0410\.449\.0360\.470\.0590\.433\.1330\.783\.0050\.644\.1230\.613

Table 6:Hyperparameter study of the Local and Global regularization coefficient on HMHVG\-50 genes\. PCC across the 10 HEST\-1K benchmark datasets is reported asmeanstd\\mathrm\{mean\}\_\{\\mathrm\{std\}\}, where the subscript denotes the standard deviation across folds/seeds\.DatasetLocal Coefficientρ\\mathbf\{\\rho\}Global Coefficientλ\\mathbf\{\\lambda\}0\.10\.30\.5𝟏𝟎−𝟒\\mathbf\{10^\{\-4\}\}𝟏𝟎−𝟑\\mathbf\{10^\{\-3\}\}𝟏𝟎−𝟐\\mathbf\{10^\{\-2\}\}IDC0\.792\.0580\.791\.0610\.793\.0580\.794\.0570\.791\.0610\.789\.063PRAD0\.499\.0240\.495\.0200\.495\.0190\.494\.0200\.495\.0200\.497\.014PAAD0\.575\.0560\.580\.0520\.566\.0550\.565\.0500\.580\.0520\.553\.066SKCM0\.825\.0190\.817\.0250\.815\.0310\.814\.0260\.817\.0250\.821\.023COAD0\.651\.0590\.665\.0410\.665\.0480\.664\.0420\.665\.0410\.656\.039READ0\.444\.0240\.449\.0360\.447\.0420\.437\.0420\.449\.0360\.444\.020CCRCC0\.462\.0610\.470\.0590\.444\.0640\.438\.0670\.470\.0590\.451\.057HCC0\.419\.1240\.433\.1330\.428\.1270\.431\.1310\.433\.1330\.446\.136LUNG0\.781\.0080\.783\.0050\.779\.0090\.782\.0060\.783\.0050\.776\.008LYMPH0\.634\.1230\.644\.1230\.629\.1240\.640\.1250\.644\.1230\.647\.119Average0\.6080\.6130\.6060\.6060\.6130\.608

![Refer to caption](https://arxiv.org/html/2609.22187v1/vis_appendix_ST.png)Figure 4:Visualization of Marker Genes on HEST\-1K benchmark\.Table 7:Comparison across the 10 HEST\-1K benchmark datasets\. GGC\-Pearson on HMHVG\-50 genes is reported\. Values are shown asmeanstd\\mathrm\{mean\}\_\{\\mathrm\{std\}\}, where the subscript denotes the standard deviation across folds/seeds\.DatasetST\-NetBLEEPBLEEP\-UNITRIPLEXStemSTFlowOursIDC0\.742\.0550\.691\.0880\.774\.0680\.781\.0590\.711\.0620\.793\.0640\.796\.054PRAD0\.851\.0190\.843\.0030\.815\.0850\.836\.0460\.828\.0090\.802\.0630\.778\.075PAAD0\.545\.0530\.498\.0270\.553\.0510\.581\.125\-0\.025\.0120\.685\.0970\.624\.102SKCM0\.625\.0150\.696\.0440\.761\.0940\.725\.1200\.001\.0030\.749\.1100\.759\.104COAD0\.664\.0490\.539\.0010\.694\.0800\.666\.0280\.088\.1350\.673\.0630\.687\.044READ0\.830\.0150\.748\.1020\.645\.2130\.829\.0300\.011\.0120\.867\.0430\.876\.049CCRCC0\.584\.0950\.562\.1410\.643\.1010\.592\.1030\.634\.1160\.657\.1080\.669\.091HCC0\.271\.0540\.131\.0870\.276\.0340\.140\.1070\.000\.0130\.181\.0370\.240\.124LUNG0\.581\.0400\.608\.0520\.649\.0340\.578\.0070\.015\.0060\.656\.0190\.706\.002LYMPH0\.364\.1440\.434\.1020\.472\.1480\.481\.1550\.111\.1440\.261\.1780\.341\.225Average0\.6060\.5750\.6280\.6210\.2370\.6330\.648

Table 8:Comparison across the 10 HEST\-1K benchmark datasets\. MSE\-Mean on HMHVG\-50 genes is reported\. Values are shown asmeanstd\\mathrm\{mean\}\_\{\\mathrm\{std\}\}, where the subscript denotes the standard deviation across folds/seeds\.DatasetST\-NetBLEEPBLEEP\-UNITRIPLEXSTFlowOursIDC2\.3133\.2452\.7633\.3852\.2603\.3662\.1083\.2752\.2883\.1822\.0132\.799PRAD1\.5510\.7292\.0200\.9221\.6960\.7361\.5950\.8171\.4520\.8651\.4090\.713PAAD1\.3320\.8751\.4010\.8971\.2220\.7701\.2790\.8471\.1230\.7241\.1410\.774SKCM2\.8942\.1873\.3111\.8382\.5651\.8052\.8531\.8052\.2061\.7652\.1821\.620COAD3\.4362\.3475\.3703\.8333\.0062\.4053\.5452\.8243\.1842\.6803\.0612\.341READ2\.6792\.0102\.4572\.0532\.0101\.6512\.6731\.9601\.5641\.0951\.6111\.143CCRCC1\.6551\.0981\.9381\.0501\.7090\.9871\.5821\.0431\.4670\.8031\.3210\.764HCC4\.8923\.6845\.3843\.9354\.6823\.8554\.7673\.6274\.3663\.5294\.1633\.346LUNG1\.6911\.7801\.7991\.9471\.6762\.0361\.7271\.9381\.5392\.0131\.5302\.229LYMPH2\.4443\.0073\.1743\.7662\.4873\.2502\.2383\.1332\.0002\.7261\.9022\.844Average2\.4892\.9622\.3312\.4362\.1192\.033

Table 9:Supplementary evaluation on the two largest human cancer categories from STImage\-1K4M\. PCC and HPCC are reported asmeanstd\\mathrm\{mean\}\_\{\\mathrm\{std\}\}\.DatasetPearson↑\\uparrowHPCC↑\\uparrowSTFlowOursSTFlowOursBrain0\.546\.2640\.558\.245––Breast0\.513\.1730\.525\.1600\.464\.0930\.473\.081Avg\.0\.5300\.5420\.4640\.473

Table 10:Pearson correlation comparison between STFlow, a Graph\-Structured Masking variant of our method, and our main model\. Results are reported on the 10 HEST\-1K benchmarks and the two supplementary STImage\-1K4M cancer categories\. Values are shown asmeanstd\\mathrm\{mean\}\_\{\\mathrm\{std\}\}, where the subscript denotes the standard deviation across folds/seeds\.DatasetSTFlowGraph\-Structured MaskingOursHEST\-1K benchmarksIDC0\.790\.0600\.797\.0560\.791\.061PRAD0\.494\.0290\.497\.0190\.495\.020PAAD0\.574\.0500\.577\.0490\.580\.052SKCM0\.813\.0300\.805\.0300\.817\.025COAD0\.649\.0550\.664\.0410\.665\.041READ0\.428\.0510\.438\.0640\.449\.036CCRCC0\.457\.0560\.449\.0620\.470\.059HCC0\.382\.1040\.431\.1290\.433\.133LUNG0\.782\.0080\.782\.0080\.783\.005LYMPH0\.629\.1280\.642\.1210\.644\.123Average0\.6000\.6080\.613STImage\-1K4M supplementary categoriesBrain0\.546\.2640\.556\.2380\.558\.245Breast0\.513\.1730\.511\.1820\.525\.160Average0\.5300\.5340\.542

Table 11:Comparison across the 10 HEST\-1K benchmark datasets\. PCC and HPCC on the HMHVG\-50 are reported\. Values are reported asmeanstd\\mathrm\{mean\}\_\{\\mathrm\{std\}\}, where the subscript denotes the standard deviation across folds/seeds\.DatasetPCC↑\\uparrowHPCC↑\\uparrowST\-NetBLEEPBLEEP\-UNITRIPLEXStemSTFlowOursST\-NetBLEEPBLEEP\-UNITRIPLEXStemSTFlowOursIDC0\.715\.0960\.608\.2270\.767\.0520\.753\.0590\.612\.0560\.790\.0600\.791\.0610\.710\.1030\.605\.2250\.769\.0560\.751\.0620\.587\.0560\.787\.0660\.788\.072PRAD0\.406\.0080\.345\.0380\.451\.0140\.412\.0420\.227\.0380\.494\.0290\.495\.0200\.507\.0440\.420\.0010\.582\.0170\.501\.0330\.256\.0900\.599\.0540\.610\.042PAAD0\.504\.0420\.465\.0630\.543\.0520\.516\.0560\.006\.0040\.574\.0500\.580\.0520\.535\.0600\.503\.0770\.599\.0530\.567\.0630\.009\.0060\.609\.0510\.622\.053SKCM0\.706\.0530\.645\.0880\.758\.0110\.748\.0020\.003\.0010\.813\.0300\.817\.0250\.691\.0470\.648\.0070\.660\.0690\.663\.063\-0\.008\.0170\.707\.1260\.732\.123COAD0\.567\.0300\.254\.0670\.648\.0310\.579\.0350\.048\.0420\.649\.0550\.665\.0410\.423\.0020\.165\.0330\.553\.0910\.535\.0270\.054\.0470\.574\.0600\.598\.028READ0\.190\.1140\.231\.1320\.408\.0880\.208\.136\-0\.001\.0040\.428\.0510\.449\.0360\.253\.0380\.216\.0660\.329\.0550\.186\.0270\.001\.0020\.418\.0680\.444\.100CCRCC0\.317\.0710\.306\.0580\.391\.0690\.291\.1660\.257\.0890\.457\.0560\.470\.0590\.332\.0750\.312\.0710\.385\.0800\.317\.1790\.257\.1060\.469\.0590\.476\.065HCC0\.291\.0440\.121\.1110\.256\.0580\.086\.070\-0\.002\.0030\.382\.1040\.433\.1330\.461\.0600\.208\.1140\.398\.0750\.118\.1070\.004\.0060\.541\.1060\.564\.087LUNG0\.734\.0170\.704\.0170\.757\.0040\.750\.0040\.004\.0000\.782\.0080\.783\.0050\.768\.0190\.737\.0170\.779\.0140\.782\.0120\.006\.0020\.814\.0030\.815\.000LYMPH0\.545\.1300\.511\.1090\.564\.1180\.594\.1040\.253\.0800\.629\.1280\.644\.123–––––––Average0\.4970\.4190\.5540\.4940\.1410\.6000\.6130\.5200\.4240\.5620\.4910\.1290\.6130\.628

Table 12:Gene\-level PCC significance analysis on HEST\-1K\. For each dataset, gene\-wise PCC values were first averaged across cross\-validation folds, yielding 50 gene\-level PCC values per dataset\. These values were then pooled across all 10 datasets, producing 500 dataset\-gene pairs for each pairwise comparison\. We compared CorrFlow against each baseline using paired Wilcoxon signed\-rank tests, followed by Holm correction\. Mean Diff\. denotes CorrFlow minus the baseline\.BaselineMean Diff\.WinsLossesRawppHolmppST\-Net0\.1152479214\.16×10−794\.16\\times 10^\{\-79\}1\.66×10−781\.66\\times 10^\{\-78\}BLEEP0\.193549732\.33×10−832\.33\\times 10^\{\-83\}1\.17×10−821\.17\\times 10^\{\-82\}BLEEP\-UNI0\.0584424769\.15×10−559\.15\\times 10^\{\-55\}1\.83×10−541\.83\\times 10^\{\-54\}TRIPLEX0\.1189475251\.99×10−781\.99\\times 10^\{\-78\}5\.98×10−785\.98\\times 10^\{\-78\}STFlow0\.01293301701\.49×10−161\.49\\times 10^\{\-16\}1\.49×10−161\.49\\times 10^\{\-16\}

Table 13:Pooled Hallmark\-Gene PCC significance analysis on HEST\-1K\. For each dataset, pathway\-level Hallmark\-Gene PCC values were first aggregated across cross\-validation folds by matching pathway names and averaging the corresponding gene\-level pathway PCC values when a pathway was present in multiple folds\. The resulting dataset\-pathway pairs were then pooled across all datasets and compared between CorrFlow and each baseline using paired Wilcoxon signed\-rank tests, followed by Holm correction\. Mean Diff\. denotes CorrFlow minus the baseline\.BaselineMean Diff\.WinsLossesRawppHolmppST\-Net0\.11532807\.45×10−97\.45\\times 10^\{\-9\}3\.73×10−83\.73\\times 10^\{\-8\}BLEEP0\.19212807\.45×10−97\.45\\times 10^\{\-9\}3\.73×10−83\.73\\times 10^\{\-8\}BLEEP\-UNI0\.05982711\.49×10−81\.49\\times 10^\{\-8\}3\.73×10−83\.73\\times 10^\{\-8\}TRIPLEX0\.11452807\.45×10−97\.45\\times 10^\{\-9\}3\.73×10−83\.73\\times 10^\{\-8\}STFlow0\.01122269\.32×10−59\.32\\times 10^\{\-5\}9\.32×10−59\.32\\times 10^\{\-5\}

Table 14:Model complexity comparison on the HEST benchmark\. \#Params counts the trainable parameters used under the current benchmark training setup, excluding external frozen or precomputed feature extractors\. Avg\. Time / Epoch is averaged over all dataset\-fold runs of each method\.Method\#Params \(M\)Model Size \(MB\)Avg\. Time / Epoch \(min\)ST\-Net11\.20242\.730\.301BLEEP24\.17892\.230\.922BLEEP\-UNI38\.205145\.748\.342TRIPLEX53\.908205\.641\.433Stem35\.581135\.730\.230STFlow1\.1484\.380\.038CorrFlow\(Our\)1\.1484\.380\.038

## Appendix BEvaluation Metrics

We evaluate histology\-to\-ST prediction from two complementary perspectives: expression accuracy and dependency\-structure fidelity\. Expression accuracy is measured using Pearson correlation and Hallmark\-Gene PCC, while dependency\-structure fidelity is measured using a gene\-gene correlation preservation metric \(GGC\-Pearson\)\. As an auxiliary expression\-level error metric, we also report MSE\-Mean\.

##### Pearson correlation coefficient \(PCC\)\.

We first evaluate gene\-level prediction fidelity using the PCC, computed independently for each gene across spatial spots\. Given predicted expressiony^:,g\\hat\{y\}\_\{:,g\}and ground\-truth expressiony:,gy\_\{:,g\}for genegg, the PCC is defined as

rg=cov\(y^:,g,y:,g\)σ\(y^:,g\)σ\(y:,g\)\.r\_\{g\}=\\frac\{\\mathrm\{cov\}\(\\hat\{y\}\_\{:,g\},y\_\{:,g\}\)\}\{\\sigma\(\\hat\{y\}\_\{:,g\}\)\\,\\sigma\(y\_\{:,g\}\)\}\.\(17\)We report the mean PCC over all evaluated genes:

Pearson=1\|𝒢\|​∑g∈𝒢rg,\\mathrm\{Pearson\}=\\frac\{1\}\{\|\\mathcal\{G\}\|\}\\sum\_\{g\\in\\mathcal\{G\}\}r\_\{g\},\(18\)where𝒢\\mathcal\{G\}denotes the set of target genes\. This metric captures how well the model preserves spatial variation patterns for individual genes\.

##### Hallmark\-Gene PCC \(HPCC\)\.

To assess biological consistency on functionally meaningful genes, we adopt a pathway\-informed evaluation based on Hallmark gene sets from MSigDB\[[14](https://arxiv.org/html/2609.22187#bib.bib14)\]\. The equation is shown in Equation[13](https://arxiv.org/html/2609.22187#S4.E13)\. Compared with Pearson, this metric emphasizes prediction quality on genes associated with curated biological programs\.

##### Gene\-Gene Correlation \(GGC\)\.

Our method is motivated by the observation that gene expression is not independent across genes, but instead exhibits structured gene\-gene dependencies\. To evaluate whether a model preserves this dependency structure, we compare the gene\-gene correlation matrices induced by the predicted and ground\-truth transcriptomes across spatial spots\.

LetY^,Y∈ℝN×G\\hat\{Y\},Y\\in\\mathbb\{R\}^\{N\\times G\}denote the predicted and ground\-truth expression matrices overNNspots andGGtarget genes\. We compute Pearson gene\-gene correlation matrices across spatial spots using the full evaluated gene set:

C^g,h=corr\(y^:,g,y^:,h\),Cg,h=corr\(y:,g,y:,h\)\.\\hat\{C\}\_\{g,h\}=\\mathrm\{corr\}\(\\hat\{y\}\_\{:,g\},\\hat\{y\}\_\{:,h\}\),\\qquad C\_\{g,h\}=\\mathrm\{corr\}\(y\_\{:,g\},y\_\{:,h\}\)\.\(19\)We then extract the upper\-triangular entries \(excluding the diagonal\), vectorize them, and compute the Pearson correlation between the predicted and ground\-truth vectors:

GGC​\-​Pearson=corr⁡\(vec△​\(C^\),vec△​\(C\)\)\.\\mathrm\{GGC\\mbox\{\-\}Pearson\}=\\mathrm\{corr\}\\\!\\left\(\\mathrm\{vec\}\_\{\\triangle\}\(\\hat\{C\}\),\\;\\mathrm\{vec\}\_\{\\triangle\}\(C\)\\right\)\.\(20\)For numerical stability, undefined correlation values arising from constant\-expression genes are set to zero in the implementation\. This metric measures the global extent to which the predicted transcriptome preserves real gene\-gene dependency structure\.

##### Mean Squared Error \(MSE\)\.

In addition to correlation\-based metrics, we report the MSE to measure the absolute deviation between predicted and ground\-truth gene expression values in Table[8](https://arxiv.org/html/2609.22187#A1.T8)\. For a given genegg, lety^i​g\\hat\{y\}\_\{ig\}andyi​gy\_\{ig\}denote the predicted and ground\-truth expression values at spotii, respectively, and letNNbe the number of test spots\. The gene\-wise MSE is defined as

MSEg=1N​∑i=1N\(y^i​g−yi​g\)2\.\\mathrm\{MSE\}\_\{g\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\(\\hat\{y\}\_\{ig\}\-y\_\{ig\}\\right\)^\{2\}\.\(21\)We then summarize performance across genes by averaging the gene\-wise MSE values:

MSE​\-​Mean=1G​∑g=1GMSEg,\\mathrm\{MSE\\text\{\-\}Mean\}=\\frac\{1\}\{G\}\\sum\_\{g=1\}^\{G\}\\mathrm\{MSE\}\_\{g\},\(22\)whereGGis the number of evaluated genes\. Lower MSE\-Mean indicates better predictive accuracy, as it reflects smaller absolute discrepancies between predicted and observed expression profiles\.

Overall, higher values indicate better performance for PCC, HPCC, and GGC\-Pearson, whereas lower values are better for MSE\-Mean\.

## Appendix CSupplementary Results

### C\.1Evaluation on Expanded Gene Sets on the HEST\-1K Benchmark\.

To further evaluate whether the advantage persists beyond a small target\-gene set, Table[4](https://arxiv.org/html/2609.22187#A1.T4)reports results on expanded HMHVG\-100, HMHVG\-150, and HMHVG\-200 settings\. Compared with STFlow, our method maintains a higher average PCC across all three gene scopes, with average PCC improving from0\.5840\.584to0\.5910\.591for HMHVG\-100, from0\.5570\.557to0\.5660\.566for HMHVG\-150, and from0\.5380\.538to0\.5430\.543for HMHVG\-200\. A similar pattern is observed for HPCC, where our method improves the average score from0\.5940\.594to0\.5970\.597,0\.5690\.569to0\.5750\.575, and0\.5480\.548to0\.5500\.550under the HMHVG\-100, HMHVG\-150, and HMHVG\-200 settings, respectively\. Although the absolute margins are moderate, the improvements are stable across increasingly broad gene sets, supporting the effectiveness of dependency\-aware generation when the prediction target expands from highly variable marker genes to larger biologically relevant gene scopes\. For PAAD, these expanded\-gene evaluations are omitted because fewer than 100 selected genes are available after preprocessing\.

### C\.2Results with Additional Metrics on HEST\-1K Benchmark

In the main text, we report PCC and HPCC as the primary evaluation metrics\. Here, we further provide GGC\-Pearson and MSE\-Mean to assess complementary aspects of prediction quality\. GGC\-Pearson measures the agreement between the predicted and ground\-truth gene\-gene correlation structures, thereby evaluating whether a method preserves transcriptomic dependency patterns beyond per\-gene prediction\. MSE\-Mean measures the mean squared prediction error averaged over genes, reflecting the numerical fidelity of the generated expression values\.

As shown in Table[7](https://arxiv.org/html/2609.22187#A1.T7), our method achieves the highest average GGC\-Pearson of0\.6480\.648, outperforming STFlow \(0\.6330\.633\) and BLEEP\-UNI \(0\.6280\.628\)\. Although the best\-performing method varies across individual datasets, the strongest average result indicates that our dependency\-aware design better preserves global gene\-gene correlation structure across the HEST\-1K benchmark\. Table[8](https://arxiv.org/html/2609.22187#A1.T8)further shows that our method obtains the lowest average MSE\-Mean of2\.0332\.033, compared with2\.1192\.119for STFlow and2\.3312\.331for BLEEP\-UNI, and achieves the best MSE\-Mean on most datasets\. These results suggest that the CorrFlow improves not only correlation\-based prediction quality, but also the numerical reconstruction fidelity of spatial gene expression\.

### C\.3Results on STImage\-1K4M

Table[9](https://arxiv.org/html/2609.22187#A1.T9)reports supplementary evaluation on the two largest human cancer categories from STImage\-1K4M\[[4](https://arxiv.org/html/2609.22187#bib.bib4)\]\. Our method improves PCC over STFlow on both Brain and Breast, increasing the average PCC from0\.5300\.530to0\.5420\.542\. For HPCC, only Breast satisfies the Hallmark gene\-set coverage criterion, where our method improves the score from0\.4640\.464to0\.4730\.473\. These results are consistent with the HEST\-1K findings and suggest that the proposed dependency\-aware design remains beneficial on larger ST collections\.

### C\.4Results for Exploratory Extension with Graph\-Structured Masking

Table[10](https://arxiv.org/html/2609.22187#A1.T10)reports the empirical results of this variant\. Graph\-structured masking improves over STFlow on average, increasing PCC from0\.6000\.600to0\.6080\.608on the 10 HEST\-1K benchmarks and from0\.5300\.530to0\.5340\.534on the two STImage\-1K4M categories\. However, it remains below our main model, which achieves0\.6130\.613and0\.5420\.542, respectively\. We hypothesize that masking entire graph neighborhoods can make each masked block relatively large, which weakens the fine\-grained timestep\-dependent control provided by annealed random masking\. As a result, the structured variant provides useful but limited gains, further supporting the importance of the proposed annealed masking design in the main model\.

### C\.5Gene\-Level Statistical Significance on HEST\-1K

To provide a finer\-grained complementary analysis beyond dataset\-level significance testing, we performed a pooled gene\-level PCC analysis on HEST\-1K\. Specifically, for each dataset, we first averaged each gene’s PCC across cross\-validation folds, resulting in 50 gene\-level PCC values per dataset\. Pooling these values across all 10 datasets yielded 500 dataset\-gene pairs, which were then compared between CorrFlow and each baseline using paired Wilcoxon signed\-rank tests with Holm correction\. As shown in Table[12](https://arxiv.org/html/2609.22187#A1.T12), CorrFlow significantly outperforms all baselines under this pooled gene\-level analysis, with adjustedpp\-values far below 0\.05 in every comparison\. The strongest improvements are observed against BLEEP and TRIPLEX, while the comparison with STFlow remains significant but with a smaller effect size, indicating that STFlow is the most competitive baseline under this analysis\.

### C\.6Hallmark\-Gene Statistical Significance on HEST\-1K

We further performed an HPCC significance analysis on HEST\-1K\. Specifically, for each dataset, we matched Hallmark pathways across cross\-validation folds by pathway name and averaged the corresponding pathway\-level gene\-PCC values when the same pathway appeared in multiple folds\. These dataset\-pathway pairs were then pooled across datasets and compared between CorrFlow and each baseline using paired Wilcoxon signed\-rank tests with Holm correction\. As shown in Table[13](https://arxiv.org/html/2609.22187#A1.T13), CorrFlow significantly outperforms all baselines under this pooled Hallmark\-Gene analysis, with particularly large gains over BLEEP and TRIPLEX\. The comparison with STFlow remains significant, although the effect size is substantially smaller than for the other baselines\.

### C\.7Model Complexity

Table[14](https://arxiv.org/html/2609.22187#A1.T14)summarizes model complexity under the current HEST benchmark training setup\. We report the number of trainable parameters, the corresponding fp32 model size, and the average training time per epoch\. Notably, CorrFlow and STFlow have the same trainable parameter count because both methods use the same trainable denoising backbone, while differing in loss design and regularization\. Compared with the larger end\-to\-end or partial\-finetuning baselines, CorrFlow maintains a smaller trainable footprint and a very low per\-epoch training cost\.

## Appendix DLimitations and Future Work

Despite the strong empirical performance of CorrFlow, several limitations remain\. First, although HEST\-1K provides a diverse cross\-platform benchmark with standardized patient\-stratified splits, our evaluation does not include a fully independent external cohort\. Such external validation remains essential for assessing clinical robustness under variations in tissue preparation, sequencing platform, staining protocol, scanner type, and cohort composition\. Future work will therefore incorporate independently collected ST cohorts to more rigorously evaluate out\-of\-distribution generalization\.

Second, although CorrFlow achieves the best overall performance across the evaluated datasets, the current prediction quality remains insufficient for direct clinical translation\. Histology\-to\-ST generation is intrinsically challenging, as molecular states cannot always be reliably inferred from morphology alone and may be affected by cohort size, tissue heterogeneity, technical noise, and gene\-panel coverage\. Improving robustness in clinically relevant settings will require stronger dependency\-aware modeling and higher\-quality ST training data\.

Finally, our current evaluation focuses primarily on reconstruction fidelity and transcriptomic structure preservation\. A key next step is to assess whether the generated ST profiles provide measurable benefits for downstream biological and clinical analyses, such as tissue\-domain identification, pathway activity estimation, biomarker discovery, patient stratification, and prognosis modeling\. Demonstrating utility in such downstream tasks is important for establishing the practical value of histology\-conditioned ST generation beyond benchmark\-level prediction metrics\.

相似文章

基于最优传输势的多边缘流匹配

arXiv cs.LG

提出OTP-FM,一种新颖的多边缘流匹配方法,利用最优传输势来软性地引导流通过中间边缘分布,在单细胞RNA测序、海洋学和气象学数据集上实现了最先进的性能。

SDFlow:用于时间序列生成的相似性驱动流匹配

arXiv cs.AI

本文介绍了 SDFlow,这是一种用于时间序列生成的相似性驱动流匹配框架,旨在解决自回归模型中的暴露偏差问题。通过在冻结的 VQ 潜在空间中进行低秩流形分解,SDFlow 实现了最先进的性能并显著提升了推理速度。