GATE-ST:面向空间转录组学的基因感知文本-图像编码器

arXiv cs.AI 论文

摘要

GATE-ST 提出了一种基因感知的文本-图像编码器,通过交叉注意力机制将文本编码的基因摘要与病理组织图像嵌入进行整合,以提升空间基因表达预测性能;在病理影像基准测试中,其表现优于仅使用图像以及使用随机嵌入的基线方法。

arXiv:2609.38690v1 Announce Type: new Abstract: Spatial transcriptomics enables spatially resolved gene expression analysis from slide-level images while preserving morphological features, providing valuable information for studying disease mechanisms and developing treatments. However, spatial gene expression profiling typically requires expensive and time-consuming tests. While existing image-based prediction optimizations mostly revolve around including positional embeddings and further image-based changes, text-based optimizations remain relatively unexplored. We present GATE-ST, which incorporates text-based inputs into image-based spatial gene expression predictions. With this approach, generated text descriptions of genes are utilized to better spatial transcriptomics prediction results. Gene summaries are put through a text encoder, generating embeddings that integrate with image embeddings through cross-attention layers to align with morphological features. We demonstrate the effectiveness of such text inputs by benchmarking performance against random gene embeddings and multiple other image-text fusion architectures, and show that GATE-ST outperforms these alternatives. Our results demonstrate the effectiveness of GATE-ST in pathology imaging, which may greatly reduce the time and cost of accurate spatial transcriptomic predictions, proving the potential of text-guided spatial gene expression prediction.
查看原文
查看缓存全文

缓存时间: 2026/10/01 09:42

# GATE-ST: Gene-Aware Text-image Encoder for Spatial Transcriptomics
Source: [https://arxiv.org/html/2609.38690](https://arxiv.org/html/2609.38690)
Jian LuoAffiliation:Stony Brook University Stony Brook, United States jian\.luo@stonybrook\.eduAffiliation:Wentao HuangAffiliation:Stony Brook University Stony Brook, United States wenthuang@cs\.stonybrook\.eduChao ChenAffiliation:Stony Brook University Stony Brook, United States chao\.chen\.1@stonybrook\.edu

###### Abstract

Spatial transcriptomics enables spatially resolved gene expression analysis from slide\-level images while preserving morphological features, providing valuable information for studying disease mechanisms and developing treatments\. However, spatial gene expression profiling typically requires expensive and time\-consuming tests\. While existing image\-based prediction optimizations mostly revolve around including positional embeddings and further image\-based changes, text\-based optimizations remain relatively unexplored\. We present GATE\-ST, which incorporates text\-based inputs into image\-based spatial gene expression predictions\. With this approach, generated text descriptions of genes are utilized to better spatial transcriptomics prediction results\. Gene summaries are put through a text encoder, generating embeddings that integrate with image embeddings through cross\-attention layers to align with morphological features\. We demonstrate the effectiveness of such text inputs by benchmarking performance against random gene embeddings and multiple other image\-text fusion architectures, and show that GATE\-ST outperforms these alternatives\. Our results demonstrate the effectiveness of GATE\-ST in pathology imaging, which may greatly reduce the time and cost of accurate spatial transcriptomic predictions, proving the potential of text\-guided spatial gene expression prediction\.

###### Index Terms:

Gene Prediction, Spatial Transcriptomics, Histopathology Image Analysis

## IIntroduction

Spatial gene\-expression patterns often correspond to morphological features, tissue composition, and associations with disease\. Bulk and single\-cell RNA sequencing are currently the leading methodologies for measuring gene expression; however, these methods are costly and compromise morphological feature information\. Spatial transcriptomics instead measures gene expression at specific spatial tissue locations while retaining spatial context, allowing it to create associations with tissue morphology\. Compared to expensive tests, preparation for spatial transcriptomics only requires widely available and inexpensive hematoxylin and eosin \(H&E\)\-stained tissue slides\. Further development in this field would allow for cheap detection of diseases such as Alzheimer’s and cancer, giving great motivation to optimize the ability for spatial transcriptomics to accurately predict gene expression\.

Existing spatial gene\-expression prediction methods have largely focused on improving visual representations or incorporating spatial context from morphological features\. Although such optimizations have improved performance, other forms of optimization remain unutilized\. Further morphological information is limited by the information a histology patch can contain and requires much larger datasets\. Comparatively, textual information can provide knowledge about individual genes and their biological functions, contributing entirely new information\. Recent text\-to\-image matching models have surfaced, but comparatively little work has explored this\. Incorporating gene descriptions may therefore provide a complementary source of information that can improve existing image\-based models\.

In this paper, we introduce GATE\-ST, Gene\-Aware Text\-image Encoder for Spatial Transcriptomics, which incorporates gene\-text summaries into image\-based spatial gene\-expression predictions\. Given an H&E patch, we extract morphological features with an image encoder, while encoding textual summaries of target genes\. These representations are passed through cross\-attention and MLP layers to align the two\. This lets text embeddings align with relevant morphological features for model tuning\. Our model shows improvements over non text\-based models, supporting the potential of text\-based optimization\.

## IIRelated Work

Early methods predict spatial gene expression directly from individual histology patches\. ST\-Net\[[1](https://arxiv.org/html/2609.38690#bib.bib1)\]applies a pretrained convolutional network followed by a regression head, while HisToGene\[[2](https://arxiv.org/html/2609.38690#bib.bib2)\]and Hist2ST\[[3](https://arxiv.org/html/2609.38690#bib.bib3)\]further model relationships among spatial locations using Transformer\- and graph\-based architectures\. Later methods incorporate broader tissue context through multi\-resolution or long\-range modeling, including M2OST\[[4](https://arxiv.org/html/2609.38690#bib.bib13)\], TRIPLEX\[[5](https://arxiv.org/html/2609.38690#bib.bib6)\], and MERGE\[[6](https://arxiv.org/html/2609.38690#bib.bib10)\]\. Another line of work uses reference\-based prediction: BLEEP\[[7](https://arxiv.org/html/2609.38690#bib.bib5)\]aligns image and expression representations through contrastive learning, whereas EGN\[[8](https://arxiv.org/html/2609.38690#bib.bib4)\]uses retrieved exemplars to refine expression estimates\. RankByGene\[[9](https://arxiv.org/html/2609.38690#bib.bib16)\]further strengthens image–gene alignment with a cross\-modal ranking\-consistency loss that preserves the relative ordering of pairwise similarities across modalities\. Despite their effectiveness, these methods generally treat genes as a predefined set of output dimensions and do not explicitly model the semantic information associated with individual genes\. Recent studies have explored gene names, functional annotations, and phenotype descriptions as additional semantic information\. Most of these methods rely on relatively simple image–text fusion\. SGN\[[10](https://arxiv.org/html/2609.38690#bib.bib7)\]and AGP\-Net\[[11](https://arxiv.org/html/2609.38690#bib.bib11)\]predict expression through similarity matching between image and gene\-text representations, while GeneQuery\[[12](https://arxiv.org/html/2609.38690#bib.bib8)\]initially combines projected image and text features through additive fusion before further processing\. DeepSpot\-M\[[13](https://arxiv.org/html/2609.38690#bib.bib12)\]adopts a more expressive gene\-query formulation over image tokens, although textual information is only one of several biological embedding sources used by the model\.

In contrast, GATE\-ST uses gene\-description embeddings as semantic queries over patch\-level visual tokens and introduces image\-residual pathways to explicitly preserve morphology\-derived information throughout gene\-conditioned fusion\.

## IIIMethodology

![Refer to caption](https://arxiv.org/html/2609.38690v1/diagram.png)Fig\. 1:Overview of the model architecture\.The proposed method, GATE\-ST, predicts spatial gene expression from H&E\-stained images by combining morphological visual information and textual information describing individual genes, as shown in Fig\.[1](https://arxiv.org/html/2609.38690#S3.F1)\. We use cross\-attention layers to align gene descriptions with relevant morphological features, training a distinct representation for each gene\. We aim to minimize the difference between our predicted expression and ground\-truth measurements\. As many genes across the dataset have poor expression, we select the 250 most highly expressed genes as prediction targets, aligning with those in other leading models such as TRIPLEX\[[5](https://arxiv.org/html/2609.38690#bib.bib6)\]\.

An overview of GATE\-ST is shown in Fig\.[1](https://arxiv.org/html/2609.38690#S3.F1)\. The architecture consists of two modules working in parallel prior to the cross\-attention architecture: first an image encoder to convert patch images into image embeddings, and then a text encoder to convert gene summaries into text embeddings\. The two embeddings are projected into a common feature space and then combined through three cross\-attention layers, learning morphological representations for each gene\. A global image pathway preserving original patch information also feeds into each cross\-attention layer\. Afterwards, a final prediction network converts these representations into predicted gene expression values\.

### III\-AImage Encoder

The first module of the architecture, the image encoder, extracts morphological features from the H&E image patches using a pretrained UNI\[[14](https://arxiv.org/html/2609.38690#bib.bib14)\]or CONCH image encoder\[[15](https://arxiv.org/html/2609.38690#bib.bib15)\]\. LetB∈ℕB\\in\\mathbb\{N\}denote the batch size, anddi∈ℕd\_\{i\}\\in\\mathbb\{N\}denote the number of pixels along each side of an input image\. A batch of image patches is represented asI∈ℝB×3×di×di\.I\\in\\mathbb\{R\}^\{B\\times 3\\times d\_\{i\}\\times d\_\{i\}\}\.

Next, we pass the image features through an image encoderEimgE\_\{\\mathrm\{img\}\}\. Letds∈ℕd\_\{s\}\\in\\mathbb\{N\}denote the number of spatial image tokens anddt∈ℕd\_\{t\}\\in\\mathbb\{N\}denote the dimensionality of each token\. For each imageIiI\_\{i\}, the encoder produces image embeddingsXi=Eimg​\(Ii\)∈ℝds×dt\.X\_\{i\}=E\_\{\\mathrm\{img\}\}\(I\_\{i\}\)\\in\\mathbb\{R\}^\{d\_\{s\}\\times d\_\{t\}\}\.

For our implementation, UNI inputs are resized to224×224224\\times 224pixels and CONCH inputs to448×448448\\times 448pixels using bicubic interpolation\. Both start from pretrained checkpoints\. UNI consists of 24 updatable Transformer blocks and producesds=196d\_\{s\}=196spatial tokens with dimensionalitydt=1024d\_\{t\}=1024, while CONCH consists of 12 Transformer blocks and internally produces 784 tokens of dimensionality 768 before pooling them into a singledt=512d\_\{t\}=512image embedding \(ds=1d\_\{s\}=1\)\. These representations are projected into the common cross\-attention dimensiondc∈ℕd\_\{c\}\\in\\mathbb\{N\}\. Both encoders may be frozen or updated with LoRA, which we apply to the final 12 blocks of both encoders\.

### III\-BText Encoder

In parallel, we pass gene\-text summaries through a text encoder\. Letdg∈ℕd\_\{g\}\\in\\mathbb\{N\}denote the number of genes the model makes predictions on\. For each genegg, we obtain a text summarysgs\_\{g\}describing its biological function\. Each summary is processed by a pretrained CONCH tokenizerτ\\tau, producing representations for each gene\. We denote the tokenizer asτ⁡\(sg\)∈ℕLg,\\tau\(s\_\{g\}\)\\in\\mathbb\{N\}^\{L\_\{g\}\},whereLg∈ℕL\_\{g\}\\in\\mathbb\{N\}denotes the length of the tokenized sequence of the summarysgs\_\{g\}\. Next, we pass these tokenized sequences through a pretrained CONCH text encoderEtextE\_\{\\mathrm\{text\}\}, producing representationseg=Etext​\(τ⁡\(sg\)\)∈ℝde\.e\_\{g\}=E\_\{\\mathrm\{text\}\}\(\\tau\(s\_\{g\}\)\)\\in\\mathbb\{R\}^\{d\_\{e\}\}\.Here,de∈ℕd\_\{e\}\\in\\mathbb\{N\}represents the dimensionality of the text embedding\. We define the stacked embeddings of all genes asGBG\_\{B\}\. Prior to the cross\-attention layers, we pass each gene embedding through a trainable residual adapter to improve performance\. The resulting representations are projected into the common cross\-attention dimensiondcd\_\{c\}\. As CONCH was pretrained to pair pathology images with textual descriptions, these representations may be trained to align with morphological features generated by the image encoder\.

### III\-CCross\-attention Layer

The two representations are now combined in a cross\-attention layer, with the image tensor and text embeddings shaped to the common attention dimensiondcd\_\{c\}\. We experimentally found that reducing image\-embedding dimensionality does not hinder performance\. Here, the gene serves as the query and finds relevant morphological features from the spatial image representations, which serve as the keys and values\. The corresponding query, key, and value tensors are

Q\\displaystyle Q=GB​WQ∈ℝB×dg×dc\\displaystyle=G\_\{B\}W\_\{Q\}\\in\\mathbb\{R\}^\{B\\times d\_\{g\}\\times d\_\{c\}\}K\\displaystyle K=X​WK∈ℝB×ds×dc\\displaystyle=XW\_\{K\}\\in\\mathbb\{R\}^\{B\\times d\_\{s\}\\times d\_\{c\}\}V\\displaystyle V=X​WV∈ℝB×ds×dc\.\\displaystyle=XW\_\{V\}\\in\\mathbb\{R\}^\{B\\times d\_\{s\}\\times d\_\{c\}\}\.
Cross\-attention is then computed as

A=Softmax⁡\(Q​KTdk\)\.A=\\operatorname\{Softmax\}\\left\(\\frac\{QK^\{T\}\}\{\\sqrt\{d\_\{k\}\}\}\\right\)\.\(1\)
Here,dk∈ℕd\_\{k\}\\in\\mathbb\{N\}is the dimensionality of each attention head, andA∈ℝB×dg×dsA\\in\\mathbb\{R\}^\{B\\times d\_\{g\}\\times d\_\{s\}\}contains the attention weights relating each gene representation to each image token\. Each gene obtains an independent weighting over all of the image representations\. These weights are applied to the corresponding vectors, producingC=A​V∈ℝB×dg×dc\.C=AV\\in\\mathbb\{R\}^\{B\\times d\_\{g\}\\times d\_\{c\}\}\.Thus,CCcarries an image\-conditioned representation for each gene\. We use multi\-head attention withh∈ℕh\\in\\mathbb\{N\}heads and stackNc∈ℕN\_\{c\}\\in\\mathbb\{N\}cross\-attention layers, passing forward these updated embeddings as the new queries\. In our implementation, we useh=4h=4heads, stackNc=3N\_\{c\}=3layers, and a cross\-attention dimensiondc=256d\_\{c\}=256\. This combines the benefits of both existing systems with updatable or frozen input layers to align text and image features\. Additionally, we add a global image residual pathway to preserve patch information\. The image tokens for this are aggregated by mean pooling,

X¯=1ds∑j=1dsX:,j,:∈ℝB×dt\.\\bar\{X\}=\\frac\{1\}\{d\_\{s\}\}\\sum\_\{j=1\}^\{d\_\{s\}\}X\_\{:,j,:\}\\in\\mathbb\{R\}^\{B\\times d\_\{t\}\}\.\(2\)
The resulting image representation is passed through an image multilayer perceptron, projecting it into the shared cross\-attention dimensiondcd\_\{c\}\. The resulting feature is broadcast across thedgd\_\{g\}gene representations and incorporated at each cross\-attention stage\. After the final cross\-attention layer, we produceH\(Nc\)∈ℝB×dg×dc,H^\{\(N\_\{c\}\)\}\\in\\mathbb\{R\}^\{B\\times d\_\{g\}\\times d\_\{c\}\},whereHHrepresents the output of the cross\-attention layer\. This tensor contains onedcd\_\{c\}\-dimensional representation for each gene\.

### III\-DPrediction Head

After the final cross\-attention layer, each gene is represented by adcd\_\{c\}\-dimensional feature vector\. A shared prediction networkfpredf\_\{\\mathrm\{pred\}\}is applied to each gene independently\. We apply this to all genes and image patches to produce a final prediction matrix\. For each geneggand imageii, we havey^i,g=fpred\(Hi,g,:Nc\)∈ℝ\.\\hat\{y\}\_\{i,g\}=f\_\{\\mathrm\{pred\}\}\\left\(H^\{N\_\{c\}\}\_\{i,g,:\}\\right\)\\in\\mathbb\{R\}\.We apply this to all genes to generate a final prediction matrixY^∈ℝB×dg\\hat\{Y\}\\in\\mathbb\{R\}^\{B\\times d\_\{g\}\}\. In our implementation, the prediction MLP maps each representation to a hidden dimension and then a final scalar output layer\. This produces one predicted expression value for each gene and image patch\. We use Mean Squared Error \(MSE\) as the loss function:

MSEg=1N​∑i=1N\(y^i,g−μy^gσy^g−yi,g−μygσyg\)2\.\\mathrm\{MSE\}\_\{g\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\(\\frac\{\\hat\{y\}\_\{i,g\}\-\\mu\_\{\\hat\{y\}\_\{g\}\}\}\{\\sigma\_\{\\hat\{y\}\_\{g\}\}\}\-\\frac\{y\_\{i,g\}\-\\mu\_\{y\_\{g\}\}\}\{\\sigma\_\{y\_\{g\}\}\}\\right\)^\{2\}\.

## IVExperiments

Dataset & PreprocessingExperiments were conducted on the HER2\+ breast cancer dataset\[[16](https://arxiv.org/html/2609.38690#bib.bib9)\], consisting of 36 spatially profiled tissue samples from 8 patients with HER2\-positive breast tumors\. In the original dataset, donors were labelled A\-H, with four patients providing 6 samples each and the other four providing 3 samples each\. These sections were hematoxylin\-and\-eosin stained at×20\\times 20magnification, providing the images used for training and testing\. The dataset includes 13,136 total spots with a diameter of100​μ​m100\\mu m\. We smooth the expressions with 8\-neighborhood smoothing, averaging the gene expression values of each spot with its eight neighbors from a local3×33\\times 3patch collection following MERGE\[[6](https://arxiv.org/html/2609.38690#bib.bib10)\]\. This reduces technical dropouts and sparsity in spatial transcriptomic measurements\[[6](https://arxiv.org/html/2609.38690#bib.bib10)\]\. We analyze the 250 most highly expressed genes following the gene\-selection protocol used in TRIPLEX\[[5](https://arxiv.org/html/2609.38690#bib.bib6)\]\.

BaselinesWe evaluate GATE\-ST using both UNI\[[14](https://arxiv.org/html/2609.38690#bib.bib14)\]and CONCH\[[15](https://arxiv.org/html/2609.38690#bib.bib15)\]as image encoders and compare with five baseline architectures\. 1\)Multi\-layer Perceptron \(MLP\):this method uses only the image encoder by passing image embeddings through two linear layers to make predictions; 2\)Direct Feature Concatenation:this method directly concatenates the text and image embeddings, then passes the concatenated representation through two linear layers to produce final predictions; 3\)Embedding\-Image Similarity Based Predictions:this method calculates cosine similarity between the image and text embeddings, then passes the similarity scores through an MLP layer to make final predictions; 4\)Cross\-attention Fusion:We replace direct feature concatenation with cross\-attention layers between the two\. Summary embeddings serve as query vectors and image embeddings produce key and value vectors, passing through three cross\-attention layers and a final prediction MLP layer to make predictions; 5\)Cross\-attention with image\-residual MLP:This model extends the cross\-attention fusion baseline to include an image representation fed back at each layer\. Image tokens are averaged to produce a global image feature that passes through two MLP layers to add to each cross\-attention layer, with the remaining architecture unchanged\.

MetricsWe evaluate performance using Mean Squared Error \(MSE\) and Pearson correlation\. For each gene, we standardize the predicted and actual expression values across the spots to scale this metric\. Final MSE is the mean of the squared errors for each gene, reducing differences in absolute expression scale between genes\. Pearson correlation is calculated between predicted and measured expression across spots for each gene\. Optimizations are aimed at improving these two metrics\. Both metrics are averaged across the 250 genes\.

HyperparametersHyperparameters were selected by grid search\. For example, three cross\-attention layers consistently outperformed other settings across the various architectures we tested, as shown in Fig\.[3](https://arxiv.org/html/2609.38690#S4.F3)\. A comprehensive list of our hyperparameters is shown in Table[I](https://arxiv.org/html/2609.38690#S4.T1)\.

TABLE I:Final hyperparameter configurations for models using UNI and CONCH image encoders\.HyperparameterUNICONCHBatch size3232Learning rate3×10−43\\times 10^\{\-4\}3×10−43\\times 10^\{\-4\}Maximum epochs200200Prediction\-head dropout0\.40\.4LoRA rank \(rr\)88LoRA scaling \(α\\alpha\)1616

### IV\-AMain Experiment

For our experiments we used 8\-fold donor\-level cross\-validation\. Model selection was based on the highest\-performing model on the validation split\. For each gene, the measured and predicted values are standardized prior to metric calculation\. Results of our experiments are shown in Table[II](https://arxiv.org/html/2609.38690#S4.T2)\. For both the CONCH and UNI image encoders, GATE\-ST outperforms all other text\-integration methods\. GATE\-ST yields a 0\.7280 MSE and 0\.6360 Pearson coefficient with the CONCH image encoder, and 0\.6774 MSE and 0\.6613 Pearson with the UNI image encoder\. Compared with direct feature concatenation, GATE\-ST yields a 0\.1006 lower MSE and 0\.0503 higher Pearson coefficient with the CONCH image encoder, and a 0\.0194 lower MSE and 0\.0097 higher Pearson coefficient, supporting its effectiveness\.

Text encoding likely carries useful additional information that better gene predictions\. As GATE\-ST contains cross\-attention, such features can be learned to match with morphological features, providing more information and improving predictions\. Further information added by textual inputs is shown to improve results, as each of our optimizations that allows for more incorporation of textual inputs improves results\. However, stronger image encoders may benefit less from such additions, as shown with the smaller improvements when using the UNI image encoder\.

TABLE II:Comparison of gene\-expression prediction architectures using CONCH and UNI image encoders\.Qualitative ResultsWe plot gene expression for specific relevant genes to compare our predictions with the ground truth as shown in Fig\.[2](https://arxiv.org/html/2609.38690#S4.F2)\. These heatmaps show the gene expression of the IGKC gene across a particular slide in the dataset\. It is clear that the predictions show strong correspondence with the ground truth, supporting the effectiveness of GATE\-ST\.

![Refer to caption](https://arxiv.org/html/2609.38690v1/Heatmaps.png)Fig\. 2:Comparison Heatmaps of the IGKC Gene Expression\.
### IV\-BAblation

Prior experiments outlined in Table[II](https://arxiv.org/html/2609.38690#S4.T2)provide quantitative comparisons and contributions of model changes, but we also quantitatively analyze the results by freezing updatable portions of the encoder\. We first test whether improvements come from model architecture changes or if textual inputs carry true information, so for this we replace our CONCH text embeddings with random Gaussian\-distributed orthogonal vectors that carry no information, where improved performance over this baseline would imply text embeddings were helpful\. Otherwise, this follows the same architecture as the fully optimized model\. This also indicates that this method of changing the architecture with random embeddings does not improve results, attributing the improvement to textual input\. Table[III](https://arxiv.org/html/2609.38690#S4.T3)compares this ablation with the complete model\.

TABLE III:Ablation study of individual components in GATE\-ST using CONCH and UNI image encoders\.As shown in Table[III](https://arxiv.org/html/2609.38690#S4.T3), the ablation study with informationless vectors performs worse than standard experiments with true text embeddings, with the CONCH system reporting an MSE of 0\.8191 and Pearson correlation of 0\.5910 with the Gaussian distribution, and an MSE of 0\.7280 and Pearson of 0\.6360 with our final optimizations\. The UNI baseline was at 0\.7338 MSE and 0\.6311 Pearson with the Gaussian distribution, but 0\.6774 MSE and 0\.6613 Pearson with our optimizations\. There exists a significant improvement between these two models, which can be attributed to textual inputs\. This means textual inputs carry meaningful information for morphological features to align with, which suggests opportunities for future work\. This table further proves that the full model also improves with cross\-attention, LoRA adaptation, and an added global image residual, justifying such changes to our architecture\.

### IV\-CHyperparameter Tuning

Various hyperparameters such as learning rates and the number of cross\-attention layers were tested for model stability across hyperparameter changes\. Relatively stable performance across different numbers of layers and hyperparameters proves its stability and robustness to hyperparameter variation\. As shown in Fig\.[4](https://arxiv.org/html/2609.38690#S4.F4), varying the number of cross\-attention layers after two and changing learning rates have little to no effect on the model performance\. Across the tested settings, this shows the model is resistant to small hyperparameter adjustments\. Our final configurations were chosen based on the best\-performing settings, but our experiments show such changes do not meaningfully impact the results\.

Fig\. 3:Cross\-Attention Layers Fine TuningFig\. 4:Learning Rate Fine Tuning

## VConclusion

In this paper, we present GATE\-ST, a new approach to spatial transcriptomics utilizing text\-based gene inputs which remain relatively unexplored\. Previous spatial transcriptomics optimization methods focus on positional encoding or purely image or morphology\-based optimizations\. However, few studies have explored text\-based optimization for spatial transcriptomics\. We build on such models and present improvements over non\-text inputs once they are included for querying\. Text\-based inputs improve performance and outperform random embeddings, proving their information\-carrying nature\. Such descriptions will be readily and inexpensively available to medical professionals, offering very great potential for future developments on this type of optimization\. Despite such benefits, opportunities remain for improving text implementation and evaluating the approach on larger datasets, as our experiments were conducted on a relatively small dataset\. However, our paper has proved the potential of text\-based inputs in optimizing spatial transcriptomics predictions\.

## References

- \[1\]B\. He, L\. Bergenstråhle, L\. Stenbeck, A\. Abid, A\. Andersson, Å\. Borg, J\. Maaskola, J\. Lundeberg, and J\. Zou\(2020\)Integrating spatial gene expression and breast tumour morphology via deep learning\.Nature biomedical engineering4\(8\),pp\. 827–834\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1)\.
- \[2\]M\. Pang, K\. Su, and M\. Li\(2021\)Leveraging information in spatial transcriptomics to predict super\-resolution gene expression from histology images in tumors\.BioRxiv,pp\. 2021–11\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1)\.
- \[3\]Y\. Zeng, Z\. Wei, W\. Yu, R\. Yin, Y\. Yuan, B\. Li, Z\. Tang, Y\. Lu, and Y\. Yang\(2022\)Spatial transcriptomics prediction from histology jointly through transformer and graph neural networks\.Briefings in Bioinformatics23\(5\),pp\. bbac297\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1)\.
- \[4\]H\. Wang, X\. Du, J\. Liu, S\. Ouyang, Y\. Chen, and L\. Lin\(2025\)M2ost: many\-to\-one regression for predicting spatial transcriptomics from digital pathology images\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 7709–7717\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1)\.
- \[5\]Y\. Chung, J\. H\. Ha, K\. C\. Im, and J\. S\. Lee\(2024\)Accurate spatial gene expression prediction by integrating multi\-resolution features\.In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 11591–11600\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1),[§III](https://arxiv.org/html/2609.38690#S3.p1.1),[§IV](https://arxiv.org/html/2609.38690#S4.p1.1)\.
- \[6\]A\. Ganguly, D\. Chatterjee, W\. Huang, J\. Zhang, A\. Yurovsky, T\. S\. Johnson, and C\. Chen\(2025\)Merge: multi\-faceted hierarchical graph\-based gnn for gene expression prediction from whole slide histopathology images\.In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 15611–15620\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1),[§IV](https://arxiv.org/html/2609.38690#S4.p1.1)\.
- \[7\]R\. Xie, K\. Pang, S\. Chung, C\. Perciani, S\. MacParland, B\. Wang, and G\. Bader\(2023\)Spatially resolved gene expression prediction from histology images via bi\-modal contrastive learning\.Advances in Neural Information Processing Systems36,pp\. 70626–70637\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1)\.
- \[8\]Y\. Yang, M\. Z\. Hossain, E\. A\. Stone, and S\. Rahman\(2023\)Exemplar guided deep neural network for spatial transcriptomics analysis of gene expression prediction\.In2023 IEEE/CVF Winter Conference on Applications of Computer Vision \(WACV\),pp\. 5028–5037\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1)\.
- \[9\]W\. Huang, M\. Xu, X\. Hu, S\. Abousamra, A\. Ganguly, S\. Kapse, A\. Yurovsky, P\. Prasanna, T\. Kurc, J\. Saltz,et al\.\(2026\)Rankbygene: gene\-guided histopathology representation learning through cross\-modal ranking consistency\.IEEE Transactions on Medical Imaging\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1)\.
- \[10\]Y\. Yang, M\. Z\. Hossain, X\. Li, S\. Rahman, and E\. Stone\(2024\)Spatial transcriptomics analysis of zero\-shot gene expression prediction\.InInternational Conference on Medical Image Computing and Computer\-Assisted Intervention,pp\. 492–502\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1)\.
- \[11\]Y\. Yang, X\. Li, L\. Pan, G\. Zhang, L\. Liu, and E\. Stone\(2025\)Agp\-net: a universal network for gene expression prediction of spatial transcriptomics\.bioRxiv,pp\. 2025–03\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1)\.
- \[12\]Y\. Xiong, L\. Liu, Y\. Cui, S\. Wu, X\. Liu, A\. B\. Chan, and C\. J\. Xue\(2024\)GeneQuery: a general qa\-based framework for spatial gene expression predictions from histology images\.arXiv preprint arXiv:2411\.18391\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1)\.
- \[13\]K\. Nonchev, S\. Dawo, K\. Silina, V\. H\. Koelzer, and G\. Rätsch\(2026\)DeepSpot\-m: a multimodal foundation model for transcriptome\-wide virtual spatial transcriptomics from histology\.medRxiv,pp\. 2026–06\.Cited by:[§II](https://arxiv.org/html/2609.38690#S2.p1.1)\.
- \[14\]R\. J\. Chen, T\. Ding, M\. Y\. Lu, D\. F\. K\. Williamson, G\. Jaume, A\. H\. Song, B\. Chen, A\. Zhang, D\. Shao,et al\.\(2024\)Towards a general\-purpose foundation model for computational pathology\.Nature Medicine30,pp\. 850–862\.External Links:[Document](https://dx.doi.org/10.1038/s41591-024-02857-3)Cited by:[§III\-A](https://arxiv.org/html/2609.38690#S3.SS1.p1.1),[TABLE II](https://arxiv.org/html/2609.38690#S4.T2.1.1.1.3.1),[§IV](https://arxiv.org/html/2609.38690#S4.p2.1)\.
- \[15\]M\. Y\. Lu, B\. Chen, D\. F\. K\. Williamson, R\. J\. Chen, I\. Liang, T\. Ding, G\. Jaume, I\. Odintsov, L\. P\. Le, G\. Gerber,et al\.\(2024\)A visual\-language foundation model for computational pathology\.Nature Medicine30,pp\. 863–874\.External Links:[Document](https://dx.doi.org/10.1038/s41591-024-02856-4)Cited by:[§III\-A](https://arxiv.org/html/2609.38690#S3.SS1.p1.1),[TABLE II](https://arxiv.org/html/2609.38690#S4.T2.1.1.1.2),[§IV](https://arxiv.org/html/2609.38690#S4.p2.1)\.
- \[16\]G\. Jaume, P\. Doucet, A\. H\. Song, M\. Y\. Lu, C\. Almagro\-Pérez, S\. J\. Wagner, A\. J\. Vaidya, R\. J\. Chen, D\. F\. Williamson, A\. Kim,et al\.\(2024\)Hest\-1k: a dataset for spatial transcriptomics and histology image analysis\.Advances in Neural Information Processing Systems37,pp\. 53798–53833\.Cited by:[§IV](https://arxiv.org/html/2609.38690#S4.p1.1)\.

相似文章