CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models

arXiv cs.AI Papers

Summary

CellWorld introduces a latent-space predictive pretraining approach for spatial transcriptomics foundation models, predicting latent representations of masked cells instead of reconstructing gene measurements. Across held-out datasets, even small variants outperform existing baselines on all benchmarks, showing that scaling and broad biological diversity improve transferability.

arXiv:2608.06659v1 Announce Type: new Abstract: This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay-specific technical variation and limiting representation transferability. To avoid directly reconstructing such variation, we shift the prediction target from observed gene measurements to latent cell representations and introduce CellWorld, which predicts the latent representations of masked cells from visible spatial context and a limited partial-expression hint. We pretrain four CellWorld variants, spanning 5.74M to 94.56M trainable parameters, on a corpus of 46 million human cells. Our controlled scaling experiments show that performance improves with model capacity, particularly on spatial tasks, while spatial transfer depends more on sufficient optimization and broad biological source diversity than on cell count alone. Across four held-out datasets, even CellWorld-Small, with 5.74M trainable parameters, outperforms every baseline on all 11 linear-probe benchmarks and all seven fine-tuned spatial benchmarks. Most notably, a frozen CellWorld-Large pretrained on only 5\% of the corpus with broad biological source coverage outperforms every fully fine-tuned baseline across all seven spatial benchmarks. Code is available at https://github.com/UoM-HealthAI/CellWorld.
Original Article
View Cached Full Text

Cached at: 08/10/26, 07:58 AM

# CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models
Source: [https://arxiv.org/html/2608.06659](https://arxiv.org/html/2608.06659)
###### Abstract

This paper shows that latent\-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics\. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay\-specific technical variation and limiting representation transferability\. To avoid directly reconstructing such variation, we shift the prediction target from observed gene measurements to latent cell representations and introduce CellWorld, which predicts the latent representations of masked cells from visible spatial context and a limited partial\-expression hint\. We pretrain four CellWorld variants, spanning 5\.74M to 94\.56M trainable parameters, on a corpus of 46 million human cells\. Our controlled scaling experiments show that performance improves with model capacity, particularly on spatial tasks, while spatial transfer depends more on sufficient optimization and broad biological source diversity than on cell count alone\. Across four held\-out datasets, even CellWorld\-Small, with 5\.74M trainable parameters, outperforms every baseline on all 11 linear\-probe benchmarks and all seven fine\-tuned spatial benchmarks\. Most notably, a frozen CellWorld\-Large pretrained on only 5% of the corpus with broad biological source coverage outperforms every fully fine\-tuned baseline across all seven spatial benchmarks\. Code is available athttps://github\.com/UoM\-HealthAI/CellWorld\.

## Introduction

Spatial transcriptomics \(ST\) measures gene expression while preserving tissue spatial organization, enabling cells to be characterized within their local tissue context\(Ståhl et al\.[2016](https://arxiv.org/html/2608.06659#bib.bib17); Chen et al\.[2015](https://arxiv.org/html/2608.06659#bib.bib7)\)\. The rapidly growing scale and diversity of ST datasets create an opportunity to learn transferable representations across datasets and downstream tasks through large\-scale pretraining\.

The first generation of ST foundation models\(Tejada\-Lapuerta et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib18); Wen et al\.[2024](https://arxiv.org/html/2608.06659#bib.bib20); Zhang et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib21); Madhu et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib13); Wang et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib19)\)has pursued this opportunity primarily through masked reconstruction of gene\-level observations\. However, ST measurements are sparse and heterogeneous across platforms owing to differences in targeted gene panels, measurement noise, and batch effects\. Direct reconstruction may therefore encourage models to reproduce assay\-specific variation alongside biological signal\.

Latent prediction instead follows the principle that*transferable representations can be learned by predicting the abstract states of unobserved inputs rather than reconstructing raw observations*\(Baevski et al\.[2022](https://arxiv.org/html/2608.06659#bib.bib3)\)\. Joint\-embedding predictive architectures \(JEPAs\) have demonstrated the potential of this principle in vision, where models predict the latent representations of masked target regions from visible context\(Assran et al\.[2023](https://arxiv.org/html/2608.06659#bib.bib2),[2025](https://arxiv.org/html/2608.06659#bib.bib1)\)\. However, directly applying this formulation to ST is nontrivial\. Image regions typically exhibit strong local continuity with their surroundings, whereas ST slides comprise discrete cells whose identities and expression states can differ even across otherwise similar spatial neighborhoods\. Visible spatial context alone may therefore not uniquely determine the identity and molecular state of a masked cell\. We refer to this underdetermination as cell\-level target*ambiguity*\. To date, it remains unclear whether and under what formulation latent prediction can serve as a reliable and scalable pretraining objective for ST foundation models\.

To address this gap, we introduce CellWorld, a large\-scale ST foundation model pretrained through latent cell prediction\. Given a local patch of spatially neighboring cells, CellWorld maps each cell’s gene\-expression profile and metadata to a cell token and models cell–cell interactions using spatial Transformers\. It randomly designates a subset of cells as masked targets and uses the remaining cells as visible context\. A context encoder processes only the visible cell tokens, whereas an EMA\-updated target encoder processes the complete patch to generate a latent target for each masked cell\. A spatial predictor is trained to recover each latent target from the encoded visible context and the corresponding target coordinate\. To resolve cell\-level target ambiguity, the predictor is additionally conditioned on a limited partial\-expression hint from the masked cell\. After pretraining, the target encoder and predictor are discarded, while the cell tokenizer and context encoder are retained for downstream tasks\.

We pretrain four CellWorld variants, namely Small, Base, Large, and Huge, on 46 million human cells spanning three platforms and 11 organs\. As shown in Figure[1](https://arxiv.org/html/2608.06659#Sx1.F1), even CellWorld\-Small, with only 5\.74M trainable parameters, outperforms every existing method on all 11 benchmarks under linear probing and all seven spatial benchmarks under fine\-tuning, with larger variants further extending this lead\. Our controlled scaling experiments show continued gains with model capacity, particularly on spatial tasks\. Data scaling shows that transferable information about cell identity can be learned from relatively little pretraining data, whereas spatial transfer is governed less by raw cell count than by*broad biological source diversity and sufficient optimization*\. Most notably, a frozen CellWorld\-Large pretrained on only 5% of the corpus with broad biological source coverage outperforms every fully fine\-tuned baseline on all seven spatial benchmarks\.

![Refer to caption](https://arxiv.org/html/2608.06659v1/x1.png)Figure 1:Comparison with existing methods across 11 held\-out task–dataset pairs under \(a\) linear probing and \(b\) fine\-tuning\. Each axis is normalized to the full\-data CellWorld\-Large score under the corresponding protocol \(100%\); numeric annotations give its absolute scores\. The dashed contour in \(b\) shows the frozen CellWorld\-Large pretrained on a broadly sampled 5% corpus subset, normalized to the fine\-tuned full\-data model\. Higher is better\.
## Related Work

#### ST foundation models\.

Nicheformer\(Tejada\-Lapuerta et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib18)\)predicts masked gene identities from expression\-ranked gene sequences, whereas scGPT\-spatial\(Wang et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib19)\), CellPLM\(Wen et al\.[2024](https://arxiv.org/html/2608.06659#bib.bib20)\), and BrainBeacon\(Zhang et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib21)\)reconstruct masked expression values\. HEIST\(Madhu et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib13)\)additionally incorporates spatial and contrastive objectives\. SToFM\(Zhao et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib22)\)reconstructs masked, expression\-derived cell embeddings produced by a domain\-adapted single\-cell encoder, alongside a spatial reconstruction objective, and thus remains reconstruction\-based despite operating in an embedding space\. In contrast, CellWorld is pretrained to predict contextualized latent representations of masked cells\.

#### Latent predictive representation learning\.

Latent prediction has emerged as an alternative to reconstructive self\-supervision across visual modalities\. BYOL\(Grill et al\.[2020](https://arxiv.org/html/2608.06659#bib.bib9)\)trains an online network to predict the representation produced by an exponential\-moving\-average target network under a different augmentation\. I\-JEPA\(Assran et al\.[2023](https://arxiv.org/html/2608.06659#bib.bib2)\)predicts latent representations of masked image regions from visible context, while subsequent work has explored and extended this paradigm across diverse architectures, modalities, and application domains\(Assran et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib1); Mur\-Labadia et al\.[2026](https://arxiv.org/html/2608.06659#bib.bib14); Chen et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib6); Saito, Kudeshia, and Poovvancheri[2025](https://arxiv.org/html/2608.06659#bib.bib16); Balestriero and LeCun[2025](https://arxiv.org/html/2608.06659#bib.bib4); Klindt, LeCun, and Balestriero[2026](https://arxiv.org/html/2608.06659#bib.bib11); Litman et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib12); ElSheikh et al\.[2026](https://arxiv.org/html/2608.06659#bib.bib8)\)\. Concurrent work, ST\-JEPA\(Birk et al\.[2026](https://arxiv.org/html/2608.06659#bib.bib5)\), applies latent prediction to masked gene tokens in sequences constructed from local cellular neighborhoods, focusing on representation learning evaluated through clustering rather than large\-scale cross\-dataset pretraining and transfer\. CellWorld instead represents each cell as a single token, models cell–cell interactions with a spatial encoder, and pretrains at scale by predicting contextualized latent representations of masked cells for transfer across diverse datasets and downstream tasks\.

## Method

### Spatial Sampling Strategy

We construct 4,096\-cell local patches through a three\-stage spatial sampling procedure\. First, we partition each slide into spatially coherent subslides usingkk\-means on the two\-dimensional cell coordinates, withKsub=max⁡\(1,⌊n/50,000⌋\)K\_\{\\mathrm\{sub\}\}=\\max\(1,\\lfloor n/50\{,\}000\\rfloor\)subslides for a slide containingnncells\. Within each subslidess, farthest\-point sampling \(FPS\) selectsKs=2​⌈ns/4096⌉K\_\{s\}=2\\lceil n\_\{s\}/4096\\rceilpatch centres, wherensn\_\{s\}is the number of cells in the subslide\. FPS is rerun at each epoch, and the 4,096 cells nearest to each centre form a training sample\.

### CellWorld Pretraining Framework

Figure[2](https://arxiv.org/html/2608.06659#Sx3.F2)illustrates the overall architecture\. Within each sampled spatial patch, CellWorld converts gene\-expression profiles into cell tokens, models their interactions using spatial attention, and predicts the latent representations of masked target cells from the encoded visible context and limited partial\-expression hints\. We detail each component below\.

![Refer to caption](https://arxiv.org/html/2608.06659v1/Figures/CellWorld.png)Figure 2:Overview of the CellWorld pretraining architecture\.#### Cell tokenization\.

We first map each cell’s gene symbols to a unified vocabulary of 19,227 genes constructed from the pretraining corpus, with a single embedding for each gene shared across platforms\. For each cellii,𝒢i\\mathcal\{G\}\_\{i\}contains up toK=512K=512genes with the highest expression values; this retains the complete assayed panel for 95% of cells\. CellWorld then constructs the cell token by aggregating their embeddings according to their expression values and adding organ and platform embeddings:

𝐮i=∑g∈𝒢ixi​g​𝐞g\+𝐞iorgan\+𝐞iplatform\.\\mathbf\{u\}\_\{i\}=\\sum\_\{g\\in\\mathcal\{G\}\_\{i\}\}x\_\{ig\}\\mathbf\{e\}\_\{g\}\+\\mathbf\{e\}^\{\\mathrm\{organ\}\}\_\{i\}\+\\mathbf\{e\}^\{\\mathrm\{platform\}\}\_\{i\}\.\(1\)Here,xi​gx\_\{ig\}is the expression value of genegg,𝐞g∈ℝd\\mathbf\{e\}\_\{g\}\\in\\mathbb\{R\}^\{d\}is its gene embedding, and𝐞iorgan\\mathbf\{e\}^\{\\mathrm\{organ\}\}\_\{i\}and𝐞iplatform\\mathbf\{e\}^\{\\mathrm\{platform\}\}\_\{i\}are metadata embeddings\. The context encoder, target encoder, and hint\-conditioned spatial predictor share the same cell tokenizer\.

#### Spatial attention\.

Because coordinate systems and spatial scales vary across slides, we avoid absolute positional embeddings and instead encode relative spatial relationships through attention biases\. Specifically, before masking, coordinates are normalized within each spatial patch as𝐬~i=\(𝐬i−𝐬¯\)/δ\\tilde\{\\mathbf\{s\}\}\_\{i\}=\(\\mathbf\{s\}\_\{i\}\-\\bar\{\\mathbf\{s\}\}\)/\\delta, where𝐬¯\\bar\{\\mathbf\{s\}\}is the patch center andδ=maxj⁡‖𝐬j−𝐬¯‖∞\\delta=\\max\_\{j\}\\\|\\mathbf\{s\}\_\{j\}\-\\bar\{\\mathbf\{s\}\}\\\|\_\{\\infty\}is the largest absolute coordinate displacement\. CellWorld then adopts a two\-dimensional variant of ALiBi\(Press, Smith, and Lewis[2021](https://arxiv.org/html/2608.06659#bib.bib15)\), which converts pairwise distances between normalized coordinates into additive self\-attention biases using fixed, head\-specific slopes\. For attention headhh, the bias isBi​j\(h\)=−mh​‖𝐬~i−𝐬~j‖2B\_\{ij\}^\{\(h\)\}=\-m\_\{h\}\\\|\\tilde\{\\mathbf\{s\}\}\_\{i\}\-\\tilde\{\\mathbf\{s\}\}\_\{j\}\\\|\_\{2\}and is added to the standard attention logits:

Attn\(h\)=softmax⁡\(𝐐\(h\)​𝐊\(h\)⊤dh\+𝐁\(h\)\)​𝐕\(h\)\.\\operatorname\{Attn\}^\{\(h\)\}=\\operatorname\{softmax\}\\\!\\left\(\\frac\{\\mathbf\{Q\}^\{\(h\)\}\\mathbf\{K\}^\{\(h\)\\top\}\}\{\\sqrt\{d\_\{h\}\}\}\+\\mathbf\{B\}^\{\(h\)\}\\right\)\\mathbf\{V\}^\{\(h\)\}\.\(2\)Here,mhm\_\{h\}controls the effective spatial scale: larger values emphasize nearby cells, whereas smaller values preserve broader spatial context\. This mechanism is used by the context encoder, target encoder, and predictor\.

#### Random masking and context encoder\.

For each spatial patch, we apply random masking by uniformly sampling without replacement a fractionr=0\.6r=0\.6of cells as prediction targetsℳ\\mathcal\{M\}, with the remaining cells forming the visible context𝒞\\mathcal\{C\}\. The target set is resampled each time a patch is presented during pretraining\. Target cells are removed before context encoding rather than replaced with learnable mask tokens\. The context encoderfθf\_\{\\theta\}therefore processes only the visible cell tokens and their normalized coordinates:

𝐳𝒞=fθ​\(𝐮𝒞,𝐬~𝒞\)\.\\mathbf\{z\}\_\{\\mathcal\{C\}\}=f\_\{\\theta\}\\left\(\\mathbf\{u\}\_\{\\mathcal\{C\}\},\\tilde\{\\mathbf\{s\}\}\_\{\\mathcal\{C\}\}\\right\)\.\(3\)

#### Target encoder\.

The target encoderfθ¯f\_\{\\bar\{\\theta\}\}mirrors the context encoder but processes the complete patch, with its outputs at the target positions defining the latent prediction targets:

𝐳ℳ=sg⁡\[fθ¯​\(𝐮𝒞∪ℳ,𝐬~𝒞∪ℳ\)ℳ\],\\mathbf\{z\}\_\{\\mathcal\{M\}\}=\\operatorname\{sg\}\\left\[f\_\{\\bar\{\\theta\}\}\\left\(\\mathbf\{u\}\_\{\\mathcal\{C\}\\cup\\mathcal\{M\}\},\\tilde\{\\mathbf\{s\}\}\_\{\\mathcal\{C\}\\cup\\mathcal\{M\}\}\\right\)\_\{\\mathcal\{M\}\}\\right\],\(4\)wheresg\\operatorname\{sg\}denotes stop\-gradient\. These representations serve as latent targets that the predictor is trained to recover\.

#### Hint\-conditioned spatial predictor\.

To reduce cell\-level target ambiguity, we condition each target query on a limited partial\-expression hint\. For each target cellii, we define its retained non\-zero gene set as𝒢i\+=\{g∈𝒢i:xi​g\>0\}\\mathcal\{G\}\_\{i\}^\{\+\}=\\\{g\\in\\mathcal\{G\}\_\{i\}:x\_\{ig\}\>0\\\}and independently include each gene with probabilityκ=0\.10\\kappa=0\.10, usingξi​g∼Bernoulli⁡\(κ\)\\xi\_\{ig\}\\sim\\operatorname\{Bernoulli\}\(\\kappa\)\. The resulting hint representation is

𝐡i=∑g∈𝒢i\+ξi​g​xi​g​𝐞g\+𝐞iorgan\+𝐞iplatform\.\\mathbf\{h\}\_\{i\}=\\sum\_\{g\\in\\mathcal\{G\}\_\{i\}^\{\+\}\}\\xi\_\{ig\}x\_\{ig\}\\mathbf\{e\}\_\{g\}\+\\mathbf\{e\}^\{\\mathrm\{organ\}\}\_\{i\}\+\\mathbf\{e\}^\{\\mathrm\{platform\}\}\_\{i\}\.\(5\)
The hint is used exclusively by the predictor and is not passed to either encoder\. Each target query is𝐪i=𝐦\+Wcp​𝐡i\\mathbf\{q\}\_\{i\}=\\mathbf\{m\}\+W\_\{\\mathrm\{cp\}\}\\mathbf\{h\}\_\{i\}, where𝐦\\mathbf\{m\}is a learnable mask token andWcpW\_\{\\mathrm\{cp\}\}projects the hint into the predictor space\. The same projection maps the visible context representations into this space, after which they are concatenated with the target queries\. Using the normalized coordinates of both visible and target cells, the spatial predictorpϕp\_\{\\phi\}produces

𝐳^ℳ=Wout​pϕ​\(\[Wcp​𝐳𝒞;𝐪ℳ\],\[𝐬~𝒞;𝐬~ℳ\]\)ℳ,\\hat\{\\mathbf\{z\}\}\_\{\\mathcal\{M\}\}=W\_\{\\mathrm\{out\}\}\\,p\_\{\\phi\}\\left\(\[W\_\{\\mathrm\{cp\}\}\\mathbf\{z\}\_\{\\mathcal\{C\}\};\\mathbf\{q\}\_\{\\mathcal\{M\}\}\],\[\\tilde\{\\mathbf\{s\}\}\_\{\\mathcal\{C\}\};\\tilde\{\\mathbf\{s\}\}\_\{\\mathcal\{M\}\}\]\\right\)\_\{\\mathcal\{M\}\},\(6\)whereWoutW\_\{\\mathrm\{out\}\}maps the predictor outputs back to the target\-representation dimension\. The predicted representations𝐳^ℳ\\hat\{\\mathbf\{z\}\}\_\{\\mathcal\{M\}\}are matched against the target representations𝐳ℳ\\mathbf\{z\}\_\{\\mathcal\{M\}\}under the training objective described next\.

#### Training objective\.

We minimize the mean squared error between the predicted and target representations over all target cells:

ℒpred=1\|ℳ\|​d​∑i∈ℳ‖𝐳^i−𝐳i‖22,\\mathcal\{L\}\_\{\\mathrm\{pred\}\}=\\frac\{1\}\{\|\\mathcal\{M\}\|d\}\\sum\_\{i\\in\\mathcal\{M\}\}\\left\\\|\\hat\{\\mathbf\{z\}\}\_\{i\}\-\\mathbf\{z\}\_\{i\}\\right\\\|\_\{2\}^\{2\},\(7\)whereddis the target representation dimension\. The target encoder is initialized from the context encoder and updated only through an exponential moving average:θ¯←τ​θ¯\+\(1−τ\)​θ,\\bar\{\\theta\}\\leftarrow\\tau\\bar\{\\theta\}\+\(1\-\\tau\)\\theta,whereτ\\tauis the momentum coefficient\. All other components, including the cell tokenizer, context encoder, predictor, and projection layers, are optimized through backpropagation\.

### Model Configurations and Pretraining

#### Model configurations\.

We instantiate CellWorld at four scales—Small, Base, Large, and Huge—by varying the width, depth, and number of attention heads in the context encoder and spatial predictor\. The resulting models span 5\.74M to 94\.56M trainable parameters\. Detailed architectural configurations are provided in Table[1](https://arxiv.org/html/2608.06659#Sx3.T1)\.

Table 1:Architectural configurations of the CellWorld model family\. Enc\. and Pred\. denote the context encoder and spatial predictor, respectively; W/L/H denotes embedding width, number of layers, and number of attention heads\. Trainable parameter counts exclude the EMA target encoder, whereas total parameter counts include it\.
#### Pretraining recipe and compute\.

We pretrain CellWorld on the full corpus for 10,000 optimization steps with a global batch size of 256 using BF16 precision and AdamW\. The learning rate is linearly warmed up to3×10−33\\times 10^\{\-3\}over the first 1,000 steps and then decayed using a cosine schedule\. The EMA momentum of the target encoder increases linearly from 0\.9990 to 0\.9999 throughout pretraining\. We use random masking with a mask ratio of 0\.6 and a target\-cell hint ratio of 0\.10\. All four model scales are trained on AMD MI250X accelerators, with compute allocation scaled from 8 to 32 GPU compute dies \(GCDs\)\. The same recipe remains numerically stable across all model scales; additional optimization and compute details are provided in Appendix C\.

## Experiments

### Experimental Setup

To pretrain CellWorld, we assemble a corpus of 46 million human cells spanning 11 organs and three high\-resolution spatial transcriptomics platforms—MERFISH, Xenium, and CosMx\. We additionally curate four datasets exclusively for downstream evaluation: MERFISH human brain, CosMx normal liver, CosMx liver cancer, and Xenium lung fibrosis\. All four datasets are excluded from pretraining for both CellWorld and the baseline models\. Expression counts in both the pretraining corpus and downstream datasets are normalized to 10,000 counts per cell and log\-transformed aslog⁡\(1\+x\)\\log\(1\+x\)\. Each downstream dataset is additionally restricted to 300 highly variable genes selected using all of its cells; no other preprocessing is applied\. Detailed dataset statistics are provided in Appendix A\.

We consider three downstream tasks: cell annotation, which evaluates cell\-level biological identity; region prediction, which assigns cells to annotated tissue regions; and niche\-composition prediction, which estimates the local proportions of neighboring cell types\. We use macro\-F1 for cell annotation and region prediction and the Pearson correlation coefficient \(PCC\) for niche\-composition prediction\. Throughout the tables, these tasks are denoted Cell, Region, and Niche, respectively\.

Our experiments investigate CellWorld’s main design choices under controlled settings, scaling behavior with model capacity and pretraining data scale, and performance relative to existing methods\. Each task is evaluated on all held\-out datasets for which the required annotations are available\. The controlled analyses use linear probing only, whereas the scaling and comparison experiments use both linear probing and full fine\-tuning\. Under linear probing, all pretrained encoders are frozen, and the same task\-specific head architecture and evaluation protocol are used for every CellWorld variant and external baseline\. CellWorld fine\-tuning uses a single shared configuration across all model\- and data\-scaling conditions, whereas external baselines use architecture\-specific best\-effort configurations\. Complete configurations are provided in Appendices B–D\.

We split each downstream dataset by FOV rather than by individual cells, ensuring that cells from the same FOV do not appear in both training and validation sets, thereby reducing spatial information leakage\. For the controlled\-design and scaling experiments, we average each dataset’s best validation score across three downstream seeds and then compute unweighted means across applicable datasets\. Comparisons with existing methods instead report every task–dataset pair, averaged over the three downstream seeds\. Details of the FOV\-level train–validation splits, task construction, and metric computation are provided in Appendix B, with complete dataset\-level results for the controlled\-design and scaling experiments in Appendix E\.

### Controlled Design Analysis

We conduct controlled comparisons using CellWorld\-Base pretrained on the full corpus with the default recipe, varying one design factor at a time while holding all other settings fixed\. For the target\-cell\-hint analysis, we additionally include matched CellWorld\-Large runs to assess whether the effect depends on model capacity\. Additional ablations of model components, including metadata embeddings and 2D\-ALiBi, are provided in Appendix E\.

#### Masking strategy\.

Masking controls both the spatial organization of prediction targets and the visible context\. We compare uniform random masking with spatial block masking, which selects four connected target regions on the spatialkk\-nearest\-neighbour graph, following vision JEPA models\(Assran et al\.[2023](https://arxiv.org/html/2608.06659#bib.bib2),[2025](https://arxiv.org/html/2608.06659#bib.bib1)\)\. To separate masking strategy from mask ratio, we evaluate both atr∈\{0\.4,0\.6,0\.8\}r\\in\\\{0\.4,0\.6,0\.8\\\}, whererris the fraction of cells withheld from the context encoder as prediction targets\. As shown in Table[2](https://arxiv.org/html/2608.06659#Sx4.T2), random masking outperforms block masking in eight of nine task\-level comparisons; the only exception is region prediction atr=0\.6r=0\.6, with a difference below 0\.001 before rounding\. Random masking also performs better in 27 of 33 matched dataset\-level comparisons \(Appendix E\)\. We therefore adopt random masking withr=0\.6r=0\.6, as performance varies only marginally across random\-mask ratios\.

![Refer to caption](https://arxiv.org/html/2608.06659v1/x2.png)Figure 3:Illustration of \(a\) random masking and \(b\) spatial block masking at the same 60% mask ratio\. Colored cells represent visible context cells, with colors indicating different cell populations; gray cells denote masked targets, and dashed contours mark the sampled spatial blocks\.Table 2:Comparison of random and spatial block masking\.
#### Target\-cell hint\.

We next examine how partial gene\-expression hints from target cells affect latent prediction\. As shown in Table[3](https://arxiv.org/html/2608.06659#Sx4.T3), CellWorld\-Base without hints performs best on both spatial tasks, whereas a hint ratio of 0\.10 performs best on cell annotation and increasing it to 0\.50 reduces performance across all three tasks\. Predictor diagnostics in Appendix E nevertheless show that the no\-hint predictor produces nearly identical outputs for different targets within the same spatial patch, approximating a patch\-conditioned mean, while a 0\.10 hint restores target\-specific variation\. This reveals a trade\-off between distinguishing individual targets and preserving reliance on spatial context\. Because downstream evaluation uses the context encoder, the degeneration remains confined to the predictor at the Base scale and does not impair spatial transfer\.

At the Large scale, however, the same predictor degeneration is no longer benign\. Figure[4](https://arxiv.org/html/2608.06659#Sx4.F4)shows that no\-hint training collapses after approximately 3,000 steps, and Table[3](https://arxiv.org/html/2608.06659#Sx4.T3)shows substantial degradation across all three tasks\. In contrast, the 0\.10\-hint model avoids this predictor degeneration, remains stable throughout training, and outperforms its no\-hint counterpart on every task\. We therefore adopt a hint ratio of 0\.10 to stabilize scaling while limiting target information\.

Table 3:Downstream performance across target\-cell hint ratios at Base and Large model scales\. Bold indicates the best result within each scale\.![Refer to caption](https://arxiv.org/html/2608.06659v1/x3.png)Figure 4:Training and validation loss trajectories for CellWorld\-Large with hint ratios of 0 and 0\.10\.
#### Prediction objective\.

We next compare latent target prediction with a matched MAE\-style control\(He et al\.[2022](https://arxiv.org/html/2608.06659#bib.bib10)\), in which the predictor serves as a decoder that reconstructs selected expression values of masked cells\. Both objectives use a mask ratio of 0\.6 and a target\-cell hint ratio of 0\.10, with all other settings matched; we tune the peak learning rate for expression reconstruction separately to ensure stable training \(Appendix D\)\. As shown in Table[4](https://arxiv.org/html/2608.06659#Sx4.T4), expression reconstruction achieves comparable cell annotation performance but performs worse on both spatial tasks\. This advantage holds for latent prediction in six of the seven spatial task–dataset pairs \(Appendix E\), supporting it as a more effective objective for learning spatially contextual representations\.

#### Spatial context\.

We test whether CellWorld uses spatial neighborhood structure by permuting cell expression profiles across fixed coordinate slots within each subslide during both pretraining and downstream inference, while preserving the coordinate\-defined KNN graph\. As shown in Table[4](https://arxiv.org/html/2608.06659#Sx4.T4), this permutation leaves cell annotation essentially unchanged but reduces performance on both spatial tasks, indicating that CellWorld exploits the correspondence between cellular states and their spatial neighborhoods\.

Table 4:Controlled comparisons of prediction objective and spatial context\. Bold indicates the best result in each column\.

### Scaling Behavior

#### Model capacity\.

We first scale CellWorld from Small to Huge while keeping the corpus and recipe fixed\. Overall, as shown in Figure[5](https://arxiv.org/html/2608.06659#Sx4.F5), CellWorld exhibits clear scaling behavior under both evaluation protocols, with CellWorld\-Huge achieving the highest task\-level performance across all three tasks\. Under linear probing, performance improves monotonically across all three tasks, indicating that larger models learn more transferable pretrained representations\. Under fine\-tuning, cell annotation and niche\-composition prediction also improve monotonically, while region prediction is the only exception, decreasing slightly from Small to Base before improving at Large and Huge\.

![Refer to caption](https://arxiv.org/html/2608.06659v1/x4.png)Figure 5:CellWorld model scaling under linear probing and fine\-tuning versus trainable parameter count \(log scale\)\.Across tasks, the gains in cell annotation are concentrated in scaling from Small to Base, with only modest improvements thereafter\. This early saturation suggests that Base already provides sufficient capacity to capture most cell\-type\-discriminative information\. In contrast, region prediction and niche\-composition prediction show substantial gains from Base to Large\. Beyond Large, gains generally taper, while region prediction under fine\-tuning continues to improve substantially\.

#### Data scaling\.

To characterize how CellWorld scales with pretraining data, we pretrain CellWorld\-Large on 5%, 25%, and 100% of the corpus under the default recipe and 10,000\-step schedule\. The smaller subsets are constructed by distributed subsampling of local patches across source slides, reducing the total cell count while largely preserving biological source diversity across organs and slides\. Both distributed subsets retain all three platforms and 11 organs; the 25% subset includes all 72 slides, while the 5% subset retains 64\. Remarkably, Table[5](https://arxiv.org/html/2608.06659#Sx4.T5.fig1)shows that the 25% subset matches or slightly outperforms the full corpus across tasks, while the 5% subset incurs only marginal overall degradation\. At face value, these results suggest that transfer performance saturates with only a small fraction of the pretraining cells\.

However, this apparent saturation may reflect two factors: under the fixed 10,000\-step schedule, smaller subsets revisit the same cells and local contexts more frequently, while distributed subsampling preserves broad biological source diversity\. We first test repeated exposure by keeping the sampling strategy fixed and scaling the pretraining steps in proportion to data size: 500, 2,500, and 10,000 steps for the 5%, 25%, and 100% subsets, respectively\. This approximately equalizes the number of visits per local context across data scales\. Surprisingly, compared with the same subsets trained for 10,000 steps, the proportionally scaled models retain comparable cell annotation performance but show marked declines on both spatial tasks under linear probing and fine\-tuning\. These results suggest that cell annotation saturates after relatively few steps, whereas spatial transfer improves markedly with additional optimization even on the same cells and local contexts\.

We next test the second factor by replacing distributed patch subsampling with whole\-slide sampling while retaining the 10,000\-step schedule, thereby concentrating each data fraction in fewer slides and organs\. At 25% data, reducing coverage from 72 slides and 11 organs to 11 slides and six organs leaves performance largely unchanged\. Further reducing coverage to four slides and two organs preserves cell annotation but consistently lowers Region and Niche performance under both evaluation protocols\. Notably, despite containing only one\-fifth as many cells, the broadly sampled 5% subset matches or slightly outperforms the concentrated 25% subset on both spatial tasks under both protocols\. Most strikingly, at a matched 5% data budget, concentrating the data from 64 slides and 11 organs to four slides and three organs substantially reduces all three linear\-probe scores\. Fine\-tuning nearly recovers cell annotation, but substantial gaps remain on both spatial tasks\. Thus, biological source diversity becomes particularly important for spatial transfer in the low\-data regime\.

Taken together, these experiments show that cell count alone does not determine the effective pretraining scale\. Optimization and biological source diversity play complementary roles: repeated exposure allows CellWorld to extract more transferable spatial structure from the available cells, whereas source diversity determines the breadth of spatial variation available to learn\. Under a limited cell budget, distributing cells across organs and slides can therefore be more valuable than adding cells from a small number of sources, particularly for spatial transfer\.

Table 5:Effects of cell count, optimization, and biological source coverage under linear probing \(LP\) and fine\-tuning \(FT\)\. The three row blocks vary these factors in sequence; all subsets retain three platforms\.Table 6:Comparison with baselines across held\-out task–dataset pairs\. Brain, Liver\-N, Liver\-C, and Lung\-F denote MERFISH human brain, CosMx normal liver, CosMx liver cancer, and Xenium lung fibrosis, respectively\. CellWorld\-Large \(5%\) uses a broadly sampled 5% corpus subset and the full 10k\-step schedule\. Bold marks the best result per protocol and column\.

### Comparison with Existing Methods

We compare four full\-data CellWorld scales and the 5% CellWorld\-Large variant with broad biological source coverage against three popular ST foundation models under both linear probing and fine\-tuning: Nicheformer\(Tejada\-Lapuerta et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib18)\), scGPT\-spatial\(Wang et al\.[2025](https://arxiv.org/html/2608.06659#bib.bib19)\), and CellPLM\(Wen et al\.[2024](https://arxiv.org/html/2608.06659#bib.bib20)\)\. For linear probing, we additionally include PCA, a competitive non\-neural baseline that ranks among the top two external methods in 10 of 11 benchmarks\. Baseline\-specific configurations are detailed in Appendix D\. Figure[1](https://arxiv.org/html/2608.06659#Sx1.F1)summarizes the comparison, while Table[6](https://arxiv.org/html/2608.06659#Sx4.T6)reports complete results for all task–dataset pairs\.

Overall, CellWorld establishes a substantial performance lead over existing ST foundation models, achieving state\-of\-the\-art results in every linear\-probe benchmark and every fine\-tuned spatial benchmark\. Under linear probing, even CellWorld\-Small, with only 5\.74M trainable parameters, outperforms every baseline across all 11 benchmarks\. Under fine\-tuning, CellWorld\-Small likewise outperforms all baselines across the seven spatial benchmarks\. Performance generally improves with model scale\. For example, CellWorld’s margins over the strongest baselines reach 0\.157 for niche\-composition prediction on CosMx normal liver under linear probing and 0\.177 for region prediction on CosMx liver cancer under fine\-tuning\. NicheFormer remains strongest on fine\-tuned cell annotation\.

Most notably, with its encoder frozen and only a linear probe trained, CellWorld\-Large pretrained on just 5% of the corpus with broad biological source coverage outperforms every fully fine\-tuned baseline across all seven spatial task–dataset pairs\. On CosMx liver cancer, the frozen 5% model leads the strongest fine\-tuned baselines by 0\.067 in Region F1 and 0\.045 in Niche PCC\. Full\-corpus pretraining further improves six of the seven spatial benchmarks, widening these margins to 0\.078 and 0\.058, respectively\.

## Conclusion

CellWorld bridges the gap between latent predictive learning and spatial transcriptomics foundation models\. Across held\-out datasets, even CellWorld\-Small, with 5\.74M trainable parameters, achieves state\-of\-the\-art performance on every linear\-probe benchmark and every fine\-tuned spatial benchmark, while a frozen CellWorld\-Large pretrained on only 5% of the corpus surpasses all fully fine\-tuned baselines across all seven spatial benchmarks\. Controlled scaling shows that spatial transfer improves with model capacity and depends more on sufficient optimization and broad biological source diversity than on cell count\. We hope that future large\-scale perturbational and temporal ST datasets will enable CellWorld to predict how cellular states and tissue organization evolve over time and under interventions, moving toward a world model of tissue dynamics\.

## References

- Assran et al\. \(2025\)Assran, M\.; Bardes, A\.; Fan, D\.; Garrido, Q\.; Howes, R\.; Muckley, M\.; Rizvi, A\.; Roberts, C\.; Sinha, K\.; Zholus, A\.; et al\. 2025\.V\-jepa 2: Self\-supervised video models enable understanding, prediction and planning\.*arXiv preprint arXiv:2506\.09985*\.
- Assran et al\. \(2023\)Assran, M\.; Duval, Q\.; Misra, I\.; Bojanowski, P\.; Vincent, P\.; Rabbat, M\.; LeCun, Y\.; and Ballas, N\. 2023\.Self\-supervised learning from images with a joint\-embedding predictive architecture\.In*Proceedings of the IEEE/CVF conference on computer vision and pattern recognition*, 15619–15629\.
- Baevski et al\. \(2022\)Baevski, A\.; Hsu, W\.\-N\.; Xu, Q\.; Babu, A\.; Gu, J\.; and Auli, M\. 2022\.Data2vec: A general framework for self\-supervised learning in speech, vision and language\.In*International conference on machine learning*, 1298–1312\. PMLR\.
- Balestriero and LeCun \(2025\)Balestriero, R\.; and LeCun, Y\. 2025\.Lejepa: Provable and scalable self\-supervised learning without the heuristics\.*arXiv preprint arXiv:2511\.08544*\.
- Birk et al\. \(2026\)Birk, S\.; Vahidi, A\.; Sanian, M\. V\.; Merchant, A\.; and Lotfollahi, M\. 2026\.ST\-JEPA: Joint\-Embedding Predictive Architecture for Spatial Transcriptomics\.In*The 2026 Workshop on Generative and Agentic AI for Biology*\.
- Chen et al\. \(2025\)Chen, D\.; Shukor, M\.; Moutakanni, T\.; Chung, W\.; Yu, J\.; Kasarla, T\.; Bang, Y\.; Bolourchi, A\.; LeCun, Y\.; and Fung, P\. 2025\.Vl\-jepa: Joint embedding predictive architecture for vision\-language\.*arXiv preprint arXiv:2512\.10942*\.
- Chen et al\. \(2015\)Chen, K\. H\.; Boettiger, A\. N\.; Moffitt, J\. R\.; Wang, S\.; and Zhuang, X\. 2015\.Spatially resolved, highly multiplexed RNA profiling in single cells\.*Science*, 348\(6233\): aaa6090\.
- ElSheikh et al\. \(2026\)ElSheikh, A\.; Wang, R\.\-X\.; Wu, W\.; Wen, Y\.; Dibaeinia, P\.; Zhang, J\. Y\.; Hu, J\. Y\.\-C\.; Knudson, M\.; Babu, S\.; Sun, S\.\-H\.; et al\. 2026\.Cell\-JEPA: Latent Representation Learning for Single\-Cell Transcriptomics\.*arXiv preprint arXiv:2602\.02093*\.
- Grill et al\. \(2020\)Grill, J\.\-B\.; Strub, F\.; Altché, F\.; Tallec, C\.; Richemond, P\.; Buchatskaya, E\.; Doersch, C\.; Avila Pires, B\.; Guo, Z\.; Gheshlaghi Azar, M\.; et al\. 2020\.Bootstrap your own latent\-a new approach to self\-supervised learning\.*Advances in neural information processing systems*, 33: 21271–21284\.
- He et al\. \(2022\)He, K\.; Chen, X\.; Xie, S\.; Li, Y\.; Dollár, P\.; and Girshick, R\. 2022\.Masked autoencoders are scalable vision learners\.In*Proceedings of the IEEE/CVF conference on computer vision and pattern recognition*, 16000–16009\.
- Klindt, LeCun, and Balestriero \(2026\)Klindt, D\.; LeCun, Y\.; and Balestriero, R\. 2026\.When Does LeJEPA Learn a World Model?*arXiv preprint arXiv:2605\.26379*\.
- Litman et al\. \(2025\)Litman, E\.; Myers, T\.; Agarwal, V\.; Mittal, E\.; Li, O\.; Gopinath, A\.; and Kassis, T\. 2025\.GeneJepa: A Predictive World Model of the Transcriptome\.*bioRxiv*, 2025–10\.
- Madhu et al\. \(2025\)Madhu, H\.; Rocha, J\. F\.; Huang, T\.; Viswanath, S\.; Krishnaswamy, S\.; and Ying, R\. 2025\.HEIST: A Graph Foundation Model for Spatial Transcriptomics and Proteomics Data\.*arXiv preprint arXiv:2506\.11152*\.
- Mur\-Labadia et al\. \(2026\)Mur\-Labadia, L\.; Muckley, M\.; Bar, A\.; Assran, M\.; Sinha, K\.; Rabbat, M\.; LeCun, Y\.; Ballas, N\.; and Bardes, A\. 2026\.V\-jepa 2\.1: Unlocking dense features in video self\-supervised learning\.*arXiv preprint arXiv:2603\.14482*\.
- Press, Smith, and Lewis \(2021\)Press, O\.; Smith, N\. A\.; and Lewis, M\. 2021\.Train short, test long: Attention with linear biases enables input length extrapolation\.*arXiv preprint arXiv:2108\.12409*\.
- Saito, Kudeshia, and Poovvancheri \(2025\)Saito, A\.; Kudeshia, P\.; and Poovvancheri, J\. 2025\.Point\-jepa: A joint embedding predictive architecture for self\-supervised learning on point cloud\.In*2025 IEEE/CVF Winter Conference on Applications of Computer Vision \(WACV\)*, 7348–7357\. IEEE\.
- Ståhl et al\. \(2016\)Ståhl, P\. L\.; Salmén, F\.; Vickovic, S\.; Lundmark, A\.; Navarro, J\. F\.; Magnusson, J\.; Giacomello, S\.; Asp, M\.; Westholm, J\. O\.; Huss, M\.; et al\. 2016\.Visualization and analysis of gene expression in tissue sections by spatial transcriptomics\.*Science*, 353\(6294\): 78–82\.
- Tejada\-Lapuerta et al\. \(2025\)Tejada\-Lapuerta, A\.; Schaar, A\. C\.; Gutgesell, R\.; Palla, G\.; Halle, L\.; Minaeva, M\.; Vornholz, L\.; Dony, L\.; Drummer, F\.; Richter, T\.; et al\. 2025\.Nicheformer: a foundation model for single\-cell and spatial omics\.*Nature methods*, 1–14\.
- Wang et al\. \(2025\)Wang, C\.; Cui, H\.; Zhang, A\.; Xie, R\.; Goodarzi, H\.; and Wang, B\. 2025\.scGPT\-spatial: Continual pretraining of single\-cell foundation model for spatial transcriptomics\.*biorxiv*, 2025–02\.
- Wen et al\. \(2024\)Wen, H\.; Tang, W\.; Dai, X\.; Ding, J\.; Jin, W\.; Xie, Y\.; and Tang, J\. 2024\.CellPLM: Pre\-training of cell language model beyond single cells\.In*International Conference on Learning Representations*, volume 2024, 5649–5673\.
- Zhang et al\. \(2025\)Zhang, C\.; Yang, Y\.; Jiao, Y\.; Yang, Q\.; Guo, X\.; Xu, J\.; Li, J\.; Zhou, Y\.; Liu, Z\.; Wu, Y\.; et al\. 2025\.BrainBeacon: A Cross\-Species Foundation Model for Single\-cell Resolved Brain Spatial Transcriptomics\.*bioRxiv*, 2025–07\.
- Zhao et al\. \(2025\)Zhao, S\.; Luo, Y\.; Yang, G\.; Zhong, Y\.; Zhou, H\.; and Nie, Z\. 2025\.Stofm: a multi\-scale foundation model for spatial transcriptomics\.*arXiv preprint arXiv:2507\.11588*\.

Similar Articles