What's in an Earth Embedding? An Explainability Analysis of Location Encoders

arXiv cs.LG Papers

Summary

This paper introduces methods to decompose location embeddings from geographic implicit neural representations into human-interpretable features, such as sparse latent concepts, natural language concepts, and visual features, revealing geographic structures like forests and urban areas.

arXiv:2606.24997v1 Announce Type: new Abstract: Geographic implicit neural representations (INRs) learn to map any coordinate on Earth to a location embedding, implicitly encoding geospatial data into the weights of a neural network. Location embeddings are widely used off the shelf as general-purpose geospatial representations, yet users lack principled tools to audit what geographic or semantic information these embeddings capture. In this work, we analyze the information content of geographic INRs through their location embeddings. We decompose these embeddings into human-interpretable features$\unicode{x2014}$namely, (i) sparse latent concepts, (ii) natural language concepts, and (iii) visual features. The latent concept embeddings are learned using sparse autoencoders. To recover natural language concepts, we apply sparse linear concept embeddings (SpLiCE) over a predefined geospatial dictionary. Finally, visual features are extracted using saliency maps derived from CLIP Surgery. We show that location embeddings can be decomposed into human-interpretable representations while retaining high reconstruction capability, revealing interpretable geographic structures such as forests, deserts, and urban features. Across methods, sparse decompositions expose systematic differences in encoded information, ranging from urban structures to broader biome and climate signals, and pretraining-space saliency maps further highlight complementary features such as roads and landmarks. We hope this work provides a first step toward interpretable geospatial representations.
Original Article
View Cached Full Text

Cached at: 06/25/26, 05:10 AM

# What’s in an Earth Embedding? An Explainability Analysis of Location Encoders
Source: [https://arxiv.org/html/2606.24997](https://arxiv.org/html/2606.24997)
Livia Betti1,∗Sebastian Ricke1,∗Ivica Obadic2,3,∗Adam J\. Stewart2Esther Rolf11University of Colorado Boulder2Technical University of Munich3Munich Center for Machine Learning∗Equal Contribution

###### Abstract

Geographic implicit neural representations \(INRs\) learn to map any coordinate on Earth to a location embedding, implicitly encoding geospatial data into the weights of a neural network\. Location embeddings are widely used off the shelf as general\-purpose geospatial representations, yet users lack principled tools to audit what geographic or semantic information these embeddings capture\. In this work, we analyze the information content of geographic INRs through their location embeddings\. We decompose these embeddings into human\-interpretable features—namely, \(i\) sparse latent concepts, \(ii\) natural language concepts, and \(iii\) visual features\. The latent concept embeddings are learned using sparse autoencoders\. To recover natural language concepts, we apply sparse linear concept embeddings \(SpLiCE\) over a predefined geospatial dictionary\. Finally, visual features are extracted using saliency maps derived from CLIP Surgery\. We show that location embeddings can be decomposed into human\-interpretable representations while retaining high reconstruction capability, revealing interpretable geographic structures such as forests, deserts, and urban features\. Across methods, sparse decompositions expose systematic differences in encoded information, ranging from urban structures to broader biome and climate signals, and pretraining\-space saliency maps further highlight complementary features such as roads and landmarks\. We hope this work provides a first step toward interpretable geospatial representations\.

00footnotetext:Correspondence to:livia\.betti@colorado\.edu## 1Introduction

The archive of geospatial data is large and growing, including remote sensing images, meteorological measurements, and geotagged social media posts\[[58](https://arxiv.org/html/2606.24997#bib.bib82)\]\. The availability of this data across space and time enables a range of tasks such as weather forecasting\[[27](https://arxiv.org/html/2606.24997#bib.bib47)\], tree canopy mapping\[[51](https://arxiv.org/html/2606.24997#bib.bib48)\], and disaster response\[[22](https://arxiv.org/html/2606.24997#bib.bib46)\]\. The massive amount of geospatial data—generally considered a barrier\-to\-entry for those without resources—has encouraged the use of machine learning to develop compact, general\-purpose geospatial representationsknown as Earth embeddings\.These representations can be learned either directly from Earth observation data or indirectly through geographic implicit neural representations \(INRs\)\[[26](https://arxiv.org/html/2606.24997#bib.bib28)\]\.In this work, we focus on geographic INRs,which encode geospatial data in the weights of a neural network called alocation encoderthat maps geographic coordinates \(latitude, longitude\) to a representation called alocation embedding\[[35](https://arxiv.org/html/2606.24997#bib.bib3),[26](https://arxiv.org/html/2606.24997#bib.bib28)\]\. Location embeddings serve as succinct representations of the geospatial information at a given location, and are increasingly used in downstream geospatial tasks, such as temperature mapping\[[25](https://arxiv.org/html/2606.24997#bib.bib1)\], air quality prediction\[[55](https://arxiv.org/html/2606.24997#bib.bib79)\], and species presence and identification\[[10](https://arxiv.org/html/2606.24997#bib.bib23)\]\.However, the information content of location embeddings is particularly hard to interpret as these embeddings do not correspond to a single input image, but rather to a specific latitude/longitude coordinate, and currently, there is no principled way to determine what information these representations encode\.

Explainability research in remote sensing has largely focused on traditional supervised image models, leaving geographic INRs underexploreddespite the fundamental differences inherent to training such architectures\.For geographic INRs, the dominant method for assessing embedding quality relies on predictive performance on benchmark tasks, which does not addresswhylocation embeddings perform well, orwhenthey may fail on unseen prediction tasks\. For example, in applications such as poverty mapping, it can be unclear whether predictions derived from Earth embeddings use meaningful indicators of poverty or instead rely on more predictable proxies such as infrastructure, which is seen more readily in satellite imagery\[[2](https://arxiv.org/html/2606.24997#bib.bib49)\]\.To address this ambiguity, we require more direct ways of interrogating what information location embeddings contain\.

We aim to interpret the information content of location embeddings through the decomposition into human\-interpretable features, providing three perspectives through 1\) latent concepts, 2\) natural language concepts, and 3\) visual features \([Figure˜1](https://arxiv.org/html/2606.24997#S1.F1)\)\. We first extract latent concepts using sparse dictionary learning methods and find that the resulting sparse decompositions can differentiate between geographic concepts while reliably reconstructing the original location embedding \([Section˜3](https://arxiv.org/html/2606.24997#S3)\)\. We then extend these explanations to use natural language concepts in an aligned location–text space\. Qualitative analysis emphasizes fine\-grained urban signals for some location encoders, and broad geographic attributes like climate or biome for others \([Section˜4](https://arxiv.org/html/2606.24997#S4)\)\. Finally, we examine the visual features learned from the training data by using saliency maps, highlightingthe differences in attention between natural and satellite imagery\([Section˜5](https://arxiv.org/html/2606.24997#S5)\)\. Taken together, these three explainability methods provide complementary but consistent interpretations of the types of information contained within location embeddings\. We conclude in[Section˜6](https://arxiv.org/html/2606.24997#S6)with a discussion of what this means for the current landscape of explainability methods that can be readily applied to understand Earth embeddings and location encoders, and where gaps exist that future work should address\.

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/method/earth.png)Location Encoderpretrainedgeographic INRLocation Embedding\[0\.21−0\.440\.78⋮\]\\left\[\\begin\{array\}\[\]\{r\}0\.21\\\\ \-0\.44\\\\ 0\.78\\\\ \\lx@intercol\\hfil\\vdots\\hfil\\lx@intercol\\\\ \\end\{array\}\\right\]Downstreamtask\[lat lon\]Feature attributionVisual features extracted from saliency maps\.![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/saliency/1.png)Text decompositionDecomposition into a predefined set of concepts in aligned location–text space\.cityparkstreetlightvegetationconcept weightSparse Latent ConceptsLearning an embedding dictionary with a sparse autoencoder\.\[0\.21−0\.44⋮\]\\left\[\\begin\{array\}\[\]\{r\}0\.21\\\\ \-0\.44\\\\ \\lx@intercol\\hfil\\vdots\\hfil\\lx@intercol\\\\ \\end\{array\}\\right\]locationembeddingEnc\.\[0\.01\.3⋮\]\\left\[\\begin\{array\}\[\]\{r\}0\.0\\\\ \\color\[rgb\]\{1,\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{1,\.5,0\}\{1\.3\}\\\\ \\lx@intercol\\hfil\\vdots\\hfil\\lx@intercol\\\\ \\end\{array\}\\right\]sparseembeddingDec\.\(A\) Generation of location embeddings from coordinate inputs\(B\) Three complementary explainability methods

Figure 1:Overview of our decomposition framework for geographic INRs\. A geographic INR maps coordinate inputs to a location embedding \(top\)\. We aim to interpret pretrained location encoders through three complementary explainability methods \(bottom\): sparse latent concepts, text decomposition, and feature attribution\.
## 2Related Work

We situate our work within the literature of geographic INRs and explainability\. For the explainability methods, we discuss in greater depth the methods that are most relevant to this work\.

### 2\.1Location Encoders and Location Embeddings

A location encoder serves both as the model as well as a compact, queryable database of geospatial information\[[26](https://arxiv.org/html/2606.24997#bib.bib28)\]\. Location encoders typically pass a latitude–longitude coordinate to a positional encoding followed by a \(small\) neural network\[[35](https://arxiv.org/html/2606.24997#bib.bib3)\]\. Existing location encoders have leveraged various positional encodings, including sinusoidal encodings \(e\.g\., SINR and Climplicit\[[34](https://arxiv.org/html/2606.24997#bib.bib51),[10](https://arxiv.org/html/2606.24997#bib.bib23),[13](https://arxiv.org/html/2606.24997#bib.bib29)\]\), random Fourier features \(RFFs\) \(e\.g\., GeoCLIP\[[54](https://arxiv.org/html/2606.24997#bib.bib15)\]\), and spherical harmonics basis functions \(e\.g\., SatCLIP\[[45](https://arxiv.org/html/2606.24997#bib.bib52),[25](https://arxiv.org/html/2606.24997#bib.bib1)\]\)\. Popular choices for the neural network that follows the positional encoder include feedforward MLPs, SIREN\[[48](https://arxiv.org/html/2606.24997#bib.bib53)\]or RESIREN\[[13](https://arxiv.org/html/2606.24997#bib.bib29)\]architectures\.

### 2\.2Explainability Methods for Geospatial ML

The growing adoption of ML methods for geospatial data analysis has motivated increasing interest in explaining geospatial ML models\[[19](https://arxiv.org/html/2606.24997#bib.bib73)\]\. Common explanation methods for image\-based geospatial models include class activation mapping \(CAM\)\[[57](https://arxiv.org/html/2606.24997#bib.bib88)\]and Grad\-CAM\[[46](https://arxiv.org/html/2606.24997#bib.bib9)\], Local Interpretable Model\-agnostic Explanations \(LIME\)\[[44](https://arxiv.org/html/2606.24997#bib.bib86)\], SHapley Additive exPlanations \(SHAP\)\[[33](https://arxiv.org/html/2606.24997#bib.bib87)\], and the use of attention mechanisms\[[3](https://arxiv.org/html/2606.24997#bib.bib89),[53](https://arxiv.org/html/2606.24997#bib.bib90)\]\. Beyond image\-based models,Jiaet al\.\[[23](https://arxiv.org/html/2606.24997#bib.bib76)\]integrate concept\-bottleneck models to design interpretable geolocalization models\.

While image\-based remote sensing models and geolocalization frameworks have attracted some attention in explainability research, geographic INRs and their resulting location embeddings remain largely unexplored in this regard\. Some recent work has begun to examine the information of pixelwise Earth embeddings, such as Google’s AlphaEarth Foundations \(AEF\)\[[6](https://arxiv.org/html/2606.24997#bib.bib8)\]\.Benavides\-Martinezet al\.\[[4](https://arxiv.org/html/2606.24997#bib.bib77)\]apply feature selection methods to analyze the contribution of individual embedding dimensions\. Beyond AEF,Raoet al\.\[[43](https://arxiv.org/html/2606.24997#bib.bib13)\]investigate the intrinsic dimension of various geographic INRs, proposing it as a quantitative measure of the amount of information they contain, butdo not auditwhatinformation is being represented\. Our work aims to fill gaps left by these previous studies, by developing a more detailed understanding of the types of information encoded in geographic INRs\.

### 2\.3General Explainability Methods

##### Sparse Autoencoders\.

Recent work in sparse dictionary learning aims to provide interpretations through latent concepts encoded in vector representations\[[47](https://arxiv.org/html/2606.24997#bib.bib78)\]\. These methods use sparse autoencoders \(SAEs\) to learn a set of relevant concepts in neural activations\. More specifically, SAEs project embeddings into a higher\-dimensional space under a sparsity constraint, encouraging the separation of features that are entangled in the original embedding space\. SAEs have been successfully used to identify interpretable, monosemantic features in large language models\[[21](https://arxiv.org/html/2606.24997#bib.bib38)\]and vision–language models\[[40](https://arxiv.org/html/2606.24997#bib.bib14)\], but have not to our knowledge been used to expose structures in location embeddings\. Evaluation of these SAE\-based methods for explainability typically involves manual inspection\.Pachet al\.\[[40](https://arxiv.org/html/2606.24997#bib.bib14)\]additionally proposed a metric for measuring the monosemanticity of SAE activations on vision tasks, using the similarity of images that activate given neurons\.

##### Natural Language Explanations\.

In contrast to sparse dictionary learning, various methods aim to decompose representations into predefined natural language concepts\. For example,Hernandezet al\.\[[18](https://arxiv.org/html/2606.24997#bib.bib58)\]propose finding natural language concepts that maximize mutual information with an input\. Other methods rely on a CLIP\-aligned image–text space, proposing techniques such as aligning embedding to a CLIP space\[[38](https://arxiv.org/html/2606.24997#bib.bib59)\]or using CLIP to dissect a neural network\[[39](https://arxiv.org/html/2606.24997#bib.bib60)\]\. In this work, we leverage on a recent approach which posits that representations in CLIP\-aligned image–text space can be expressed as linear functions of single\-concept text embeddings\[[5](https://arxiv.org/html/2606.24997#bib.bib11)\]\. Using this hypothesis, sparse linear concept embeddings \(SpLiCE\) represent CLIP image embeddings as sparse linear combinations of a predefined dictionary of text concepts\[[5](https://arxiv.org/html/2606.24997#bib.bib11)\]\.

##### Feature Attribution and Saliency Maps\.

A key aspect of geographic INR explainability is understanding how the pretraining data—often, georeferenced natural or remote sensing imagery—contributes to the resulting location embedding\. Saliency maps identify the regions of an image that most strongly influence a model’s prediction\. Gradient\-weighted Class Activation Mapping \(Grad\-CAM\)\[[46](https://arxiv.org/html/2606.24997#bib.bib9)\]does so by computing gradients of the class score and has extensions to embedding models\[[8](https://arxiv.org/html/2606.24997#bib.bib10)\]\.As most existing location encoders are trained using CLIP\-style losses, we leverage methods specific to CLIP\. Grad\-CAM on CLIP models has been shown to produce misleading or unsatisfactory results, such as highlighting background elements and noisy activations\[[29](https://arxiv.org/html/2606.24997#bib.bib7),[30](https://arxiv.org/html/2606.24997#bib.bib2)\]\.ECLIP\[[29](https://arxiv.org/html/2606.24997#bib.bib7)\]proposes masked max pooling to avoid this semantic shift, but it requires an additional alignment step, limiting its applicability\.We utilize CLIP Surgery\[[30](https://arxiv.org/html/2606.24997#bib.bib2)\],which instead modifies theinference architecture of a CLIP\-trained model to address these visualization drawbacks, producing better qualitative results without the need for finetuning\.

## 3Unsupervised Discovery of Sparse Concepts

To uncover relevant concepts in location embeddings, our first explainability analysis uses sparse dictionary learning to disentangle these embeddings into human\-interpretable features\. Building on prior work\[[21](https://arxiv.org/html/2606.24997#bib.bib38),[40](https://arxiv.org/html/2606.24997#bib.bib14)\], we use a sparse autoencoder \(SAE\) to learn a sparse dictionary of concepts that enables accurate reconstruction of location embeddings\. The sparsity objective promotes interpretability by ensuring that only a small number of neurons in the autoencoder are activated for accurate reconstruction of location embeddings\. Uncovering the semantics of these neurons into human\-interpretable concepts can reveal the structure present in the location embeddings\.

In our analysis, we use the BatchTopK SAE\[[7](https://arxiv.org/html/2606.24997#bib.bib39)\]method commonly used for interpreting large language models and vision–language models\[[40](https://arxiv.org/html/2606.24997#bib.bib14)\]\. This method constrains at mostb×kb\\times kelements per batch of sizebbto be non\-sparse\. As such, it enables a variable, flexible number of non\-sparse elements to be activated per sample in the batch\. To interpret the semantics of sparse neurons, we evaluate their visual and geographic monosemanticity \(MS\)\. The visual MS metric defined byPachet al\.\[[40](https://arxiv.org/html/2606.24997#bib.bib14)\]measures whether a neuron in the sparse autoencoder layer activates for images with similar visual features\. We compute this by iterating over pairs of locations and for each pair computing the correlation between the shared neuron activation at those locations and the visual similarity between the corresponding images\. Higher values indicate that the locations activating a neuron share similar visual features\. Further, the geographic monosemanticity \(GeoMS\) metric introduced in our study assesses whether a sparse concept activates over a local or global geographical region\. We compute this analogously to the previous metric by replacing visual similarity with the inverse Haversine distance between two points on Earth\. Thus, neurons with higher GeoMS values activate more strongly within a local geographic region\.

### 3\.1Sparse Autoencoder Experimental Setup

We use the BatchTopK autoencoder model to interpret SatCLIP, GeoCLIP, and Climplicit location encoders in terms of sparse dictionary concepts\. To account for the geographical distribution bias in the location encoders, we train the SAE on two sets of 100K locations\. The first set is sampled uniformly at random \(UAR\) across landmasses, similar to the training distribution for the SatCLIP and Climplicit encoders\. The second set focuses on human\-visited areas and is formed by sampling locations from the Geo\-YFCC dataset\[[14](https://arxiv.org/html/2606.24997#bib.bib81)\], which has a distribution similar to that used to train the GeoCLIP model\. We follow the evaluation protocol outlined in Appendix Table[7](https://arxiv.org/html/2606.24997#S7.T7)and compute the monosemanticity of SAE models trained on the UAR dataset using satellite imagery from the S2\-100k dataset\[[25](https://arxiv.org/html/2606.24997#bib.bib1)\], and on human\-visited locations using natural imagery from the Geo\-YFCC dataset\. For training the BatchTopK autoencoder model, we set the sparse dictionary size to the embedding dimension for each location encoder and usek=20k=20to flexibly constrain the number of non\-sparse neurons per sample\. We use a learning rate of3​e−43e^\{\-4\}; for the other hyperparameters, we use the default values provided byPachet al\.\[[40](https://arxiv.org/html/2606.24997#bib.bib14)\]\.

### 3\.2Reconstruction Quality and Concept Monosemanticity

Table[1](https://arxiv.org/html/2606.24997#S3.T1)compares the reconstruction quality \(MSE,R2R^\{2\}score\) and the monosemanticity of the BatchTopK sparse autoencoders for three different location encoders on the datasets used in this study\. GeoCLIP embeddings yield the lowest reconstruction scores, achieving only about half and two\-thirds of theR2R^\{2\}variance explained for the UAR and human\-visited locations datasets, respectively\. In contrast, SatCLIP and Climplicit demonstrate superior performance, as the BatchTopK SAE achieves close\-to\-perfect reconstruction of their embeddings on both datasets\. In addition, the sparse neurons for SatCLIP and Climplicit exhibit stronger visual and geographic monosemanticity than those for GeoCLIP, indicating that these neurons activate imagery with more similar visual patterns and regions that are geographically closer\. Finally,[Table˜1](https://arxiv.org/html/2606.24997#S3.T1)highlights a trade\-off in monosemanticity scores driven by the location distribution in the training dataset and the source of imagery used for evaluation\. The models trained in UAR locations and evaluated on satellite imagery from the S2\-100k exhibit stronger visual monosemanticity\. Conversely, SAE models trained on human\-visited sites and evaluated on natural imagery from Geo\-YFCC achieve higher geographical monosemanticity, with sparse neurons that show increased specificity to local geographical regions\.

Table 1:Reconstruction quality of the location embeddings with the BatchTopK SAE method and monosemanticity of the latent neurons\.The Max columns report the monosemanticity of the neuron with the highest monosemanticity \(MS\) values in the SAE, while the Mean columns report the average monosemanticity over all neurons\. The location embeddings of both SatCLIP and Climplicit allow explaining a higher percentage of the variance after SAE reconstruction, and the visual MS values suggest that their sparse neurons activate on more coherent visual concepts than in GeoCLIP\. Further, these concepts pertain to more regional geographies\.Visual MS↑\\mathbf\{\\uparrow\}Geo MS↑\\mathbf\{\\uparrow\}DatasetModelMSE↓\\mathbf\{\\downarrow\}𝐑𝟐↑\\mathbf\{R^\{2\}\\uparrow\}MeanMaxMeanMaxUARGeoCLIP0\.00±0\.000\.00\\pm 0\.000\.52±0\.000\.52\\pm 0\.000\.03±0\.000\.03\\pm 0\.000\.16±0\.090\.16\\pm 0\.090\.03±0\.000\.03\\pm 0\.000\.12±0\.040\.12\\pm 0\.04SatCLIP0\.01±0\.000\.01\\pm 0\.000\.96±0\.000\.96\\pm 0\.000\.13±0\.000\.13\\pm 0\.000\.42±0\.010\.42\\pm 0\.010\.07±0\.000\.07\\pm 0\.000\.31±0\.060\.31\\pm 0\.06Climplicit0\.01±0\.000\.01\\pm 0\.000\.92±0\.000\.92\\pm 0\.000\.17±0\.000\.17\\pm 0\.000\.59±0\.020\.59\\pm 0\.020\.19±0\.000\.19\\pm 0\.000\.59±0\.040\.59\\pm 0\.04Human\-VisitedGeoCLIP0\.00±0\.000\.00\\pm 0\.000\.65±0\.000\.65\\pm 0\.000\.19±0\.000\.19\\pm 0\.000\.35±0\.030\.35\\pm 0\.030\.07±0\.000\.07\\pm 0\.000\.65±0\.070\.65\\pm 0\.07SatCLIP0\.00±0\.000\.00\\pm 0\.000\.99±0\.000\.99\\pm 0\.000\.20±0\.000\.20\\pm 0\.000\.33±0\.020\.33\\pm 0\.020\.18±0\.000\.18\\pm 0\.000\.87±0\.060\.87\\pm 0\.06Climplicit0\.01±0\.000\.01\\pm 0\.000\.99±0\.000\.99\\pm 0\.000\.21±0\.000\.21\\pm 0\.000\.41±0\.030\.41\\pm 0\.030\.22±0\.000\.22\\pm 0\.000\.92±0\.040\.92\\pm 0\.04
### 3\.3Interpreting Sparse Concepts

To interpret the semantics of neurons in the sparse autoencoder layer, in Figure[2](https://arxiv.org/html/2606.24997#S3.F2)we visualize the strongest activating samples from the S2\-100k and the Geo\-YFCC datasets for neurons with the top\-10 highest visual and geographic monosemanticity for the reconstruction of the Climplicit embeddings\. These samples reveal that the sparse neurons having high monosemanticity scores capture three distinct patterns: 1\)regional monosemantic concepts, 2\)regional polysemantic concepts, and 3\)recurring visual patterns over broader area\. The regional monosemantic concepts in Figure[2](https://arxiv.org/html/2606.24997#S3.F2)are represented by neurons 241 and 757 from the S2\-100k dataset, as well as by neuron 731 from the Geo\-YFCC dataset\. They capture rainforests in the Amazon basin, deserts in Australia, and archaeological sites in the Middle East, respectively\. Further, regional polysemantic neurons fire on diverse visual concepts specific to a certain region, such as neuron 102 from the S2\-100k dataset, which encodes concepts such as agricultural fields and deserts specific to the Middle East and the Horn of Africa, and neuron 823 from Geo\-YFCC, which captures events from the Middle East, such as a wedding chuppah and a dance scene\. The third type of neuron encodes recurring visual patterns over a broader area, and is exemplified in[Figure˜2](https://arxiv.org/html/2606.24997#S3.F2)by neuron 30 from the Geo\-YFCC dataset that strongly activates in natural areas in Europe, often including animal scenes\. Furthermore, we note the existence of sparse neurons with limited interpretability that have lower monosemanticity scores and activate globally across diverse landscapes\.

In addition to improving the interpretability of the location encoders, we found that the sparse autoencoders can also be useful in identifying specific artifacts in the pretraining datasets\. In particular, Figure[8](https://arxiv.org/html/2606.24997#S7.F8)in the Appendix illustrates that some of the top\-10 neurons with the highest visual monosemanticity encode visual artifacts in Sentinel\-2 imagery around Greenland from the S2\-100k dataset\.

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/sae_figures/sae_qualitative_both_datasets.png)Figure 2:Qualitative analysis of selected neurons from the top\-10 ranked by visual \(top and middle row\) and geospatial \(bottom row\) monosemanticity for SAE reconstruction on the Climplicit encoder\.The left panels display images from the top\-10 activating locations from the S2\-100k and Geo\-YFCC datasets for selected neurons, with the MS values indicating their visual monosemanticity\. The corresponding spatial distribution of these neurons is shown on the map on the right, with marker sizes indicating the strength of neurons activations\. The visualized neurons typically encode rainforests in the Amazon basin \(Neuron 241\), deserts in Australia \(Neuron 757\), and agricultural fields and deserts in the Middle East and the Horn of Africa \(Neuron 102\) for the S2\-100k dataset\. Furthermore, the sparse neurons on Geo\-YFCC activate to archaeological sites in the Middle East \(Neuron 731\), coherent visual patterns over a broader area, such as natural landscapes with animals in Europe \(Neuron 30\), and regional polysemantic visual concepts, such as wedding chuppah and dance scenes in the Middle East \(Neuron 823\)\.

## 4Decomposition into Predefined Natural Language Concepts

The next method we use to explain location embeddings also utilizes sparsity to identify embedding structure, but introduces additional inductive bias through a prespecified natural language concept dictionary, following the SpLiCE framework proposed byBhallaet al\.\[[5](https://arxiv.org/html/2606.24997#bib.bib11)\]\. In this setting, the set of one or two\-word dictionary concepts are constructed a priori through domain knowledge\.

##### Aligning Location Embeddings and Text Embeddings\.

SpLiCE was originally developed for interpreting CLIP image embeddings, which are naturally aligned with text embeddings\. Extending this approach to location embeddings requires a shared semantic space oflocationand text embeddings\. Some geographic INRs have this prerequisite by design, such as GeoCLIP\[[54](https://arxiv.org/html/2606.24997#bib.bib15)\], which uses a CLIP location encoder during pretraining\. To apply SpLiCE to other location embeddings, we construct this aligned space manually\. We achieve this alignment by treating this as a locked\-image tuning problem\[[56](https://arxiv.org/html/2606.24997#bib.bib80)\]with a fixed location encoder and a learned text encoder\.

As GeoCLIP is trained contrastively with the OpenCLIP ViT\-L\[[42](https://arxiv.org/html/2606.24997#bib.bib63)\]image encoder, we can directly replace the image encoder with the OpenCLIP ViT\-L text encoder to produce text embeddings in a space aligned with the GeoCLIP location encoder\. To develop a shared semantic space of location and text embeddings for all other location encoders, we finetune the OpenCLIP ViT\-L model using the Git\-10M dataset\[[32](https://arxiv.org/html/2606.24997#bib.bib70)\], which contains 10\.5M geolocated remote\-sensing images paired with captions generated via the GPT\-4o API\[[1](https://arxiv.org/html/2606.24997#bib.bib71)\]\. From Git\-10M, we additionally construct a vocabulary of 2000 geospatial concepts by selecting the most frequent one\-word nouns appearing in the captions\. We then apply SpLiCE\[[5](https://arxiv.org/html/2606.24997#bib.bib11)\]using this fixed concept dictionary\. In practice, CLIP’s image and text modalities are not fully aligned\[[31](https://arxiv.org/html/2606.24997#bib.bib43)\], which we find applies to the location and text alignment as well\. Therefore, we mean\-center and re\-normalize the location and text embeddings as inBhallaet al\.\[[5](https://arxiv.org/html/2606.24997#bib.bib11)\]\. Again following the methodology ofBhallaet al\.\[[5](https://arxiv.org/html/2606.24997#bib.bib11)\], we choose theℓ1\\ell\_\{1\}regularization \(λ\\lambda\) to encourage sparsity, aiming for solutions with approximately 5–20 active concepts per example\.

### 4\.1Reconstructing Location Embeddings from Sparse Decompositions

As with the sparse autoencoder explanations, we first measure how well the SpLiCE decomposition preserves the geometry of the original embedding space \([Table˜2](https://arxiv.org/html/2606.24997#S4.T2)\)\. Unlike reconstruction\-based approaches, here the goal is not to perfectly recover𝐳\\mathbf\{z\}, but to assess whether a sparse and interpretable concept representation retains directional structure in embedding space\. The results in[Table˜2](https://arxiv.org/html/2606.24997#S4.T2)show that SpLiCE achieves reasonable directional reconstruction \(cosine similarity substantially above 0\) and overall reconstruction \(low MSE\) while maintaining a substantially lower number of active concepts compared to the original location embeddings\. This indicates that the concept dictionary can be used to capture meaningful structure in the embedding space under sparsity constraints\.

### 4\.2Qualitative Concept Identification and Comparisons

[Figure˜3](https://arxiv.org/html/2606.24997#S4.F3)provides a qualitative analysis of SpLiCE decompositions by inspecting the top\-4 text concepts associated with four regions representing distinct land cover types\. This qualitative evaluation illustrates how different location encoders distribute semantic information over the geospatial concept set\.For GeoCLIP embeddings, concept decompositions tend to be more fine\-grained and specific in highly photographed locations such as Paris\. In contrast, less frequently photographed regions like Siberia demonstrate some interpretability in the top concept “tundra”, yet other terms represent general, non\-geographic concepts\. On the other hand, SatCLIP embeddings seem to capture more general trends related to the landcover of the region, as seen in the concepts relevant for Paris\. Climplicit shows mixed behavior: the SpLiCE decompositions are partially meaningful for regions such as the Sahara, where coarse land cover signals may be prevalent, but yield less coherent decompositions for Paris and Siberia\. Lastly, CSP\-fMoW produces concept decompositions that are generally less interpretable across the different regions, with top concepts appearing less consistently aligned with expected geographic attributes\. Together, these results suggest that the interpretability of these SpLiCE decompositions varies significantly across the different location embeddings\.

Table 2:Sparse concept decompositions preserve structure in location embedding space\.We compare original location embeddings𝐳\\mathbf\{z\}with their sparse representations produced by SpLiCE\. Results are obtained from SpLiCE decompositions of the same two sets of points used to evaluate SAEs:UARandHuman\-Visited\. All values are averaged over all points in the dataset and reported as mean±\\pmstandard deviation\. The sparsity regularization parameter is set toλ=0\.25\\lambda=0\.25for GeoCLIP and SatCLIP, andλ=0\.175\\lambda=0\.175for Climplicit and CSP\-fMoW\.UARHuman\-visitedModelMSE↓\\downarrowCos\. Sim\.↑\\uparrow\# ConceptsMSE↓\\downarrowCos\. Sim\.↑\\uparrow\# ConceptsGeoCLIP0\.002±\\pm0\.0000\.435±\\pm0\.0648\.2±\\pm2\.90\.002±\\pm0\.0000\.377±\\pm0\.0617\.8±\\pm3\.1SatCLIP0\.005±\\pm0\.0010\.360±\\pm0\.0789\.1±\\pm3\.50\.004±\\pm0\.0010\.462±\\pm0\.0888\.7±\\pm3\.4Climplicit0\.001±\\pm0\.0000\.299±\\pm0\.0919\.1±\\pm4\.30\.001±\\pm0\.0000\.361±\\pm0\.09211\.3±\\pm4\.8CSP\-fMoW0\.004±\\pm0\.0010\.442±\\pm0\.19215\.4±\\pm5\.40\.005±\\pm0\.0010\.331±\\pm0\.12713\.8±\\pm6\.7

ParisSaharaSiberiaBaliGeoCLIP![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/geoclip_Paris_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/geoclip_Sahara_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/geoclip_Siberia_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/geoclip_Bali_region_top_4.png)SatCLIP![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/satclip_Paris_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/satclip_Sahara_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/satclip_Siberia_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/satclip_Bali_region_top_4.png)Figure 3:SpLiCE decompositions highlight different signals captured by location embeddings\.Bar plots show the concept weights of the top 4 concepts in SpLiCE decompositions for GeoCLIP \(top\) and SatCLIP \(bottom\) across four locations representing different land cover types: Paris, Sahara, Siberia, and Bali\.SpLiCE decomposition weights are averaged over a0\.5∘0\.5^\{\\circ\}grid within each region’s boundary\. Results for the other location encoders are included in Appendix[Section˜7\.5\.2](https://arxiv.org/html/2606.24997#S7.SS5.SSS2)\.

## 5Feature Attribution to Extract Visual Features

For our third explainability analysis, we deviate from methods that explain embeddings “off\-the\-shelf” and now turn to feature attribution in input space to validate if the information content of location embeddings corresponds to meaningful visual evidence\. We leverage CLIP Surgery\[[30](https://arxiv.org/html/2606.24997#bib.bib2)\]to generate saliency maps for contrastively pretrained location encoders\. CLIP surgery has been shown to produce superior visualizations compared to traditional backpropagation\-based methods such as Grad\-CAM \(Section[2](https://arxiv.org/html/2606.24997#S2)\)\. In this approach: \(1\) all modifications are applied at inference time, requiring access to the pretrained location encoder weights but no fine\-tuning; and \(2\) the original attention computation is preserved, keeping the class token output unchanged and thus leaving general model usage unaffected\.

To adapt CLIP Surgery for location encoders, we use location as the unique class and compute the similarity between the location and image embeddings\. CLIP Surgery utilizes an empty string to approximate noisebutlocation encoders have no analogue to the null class or empty string used in language modelssince every location is valid\. As a heuristic, we choose the north pole \(lat = 90, lon = 0\) to approximate a “null” embedding, because it lies in a sparsely \(or not at all\) sampled region for most location encoder training datasets\.

### 5\.1Experimental Setup

CLIP Surgery requires access to both model weights and source code; therefore we restrict our visualization results to GeoCLIP and SatCLIP, both of which are released under permissive open\-source MIT licenses\. For GeoCLIP, we applied our method to Im2GPS\[[41](https://arxiv.org/html/2606.24997#bib.bib64)\], which contains natural images from around the world\. We manually grouped the results into four categories—Landmarks, Landscapes, Text, and Other—based on the visual cues qualitatively observed to drive the model’s attention as shown in Figure[4](https://arxiv.org/html/2606.24997#S5.F4)\.

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/Landmarks/india.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/Landscapes/california.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/Text/france_2.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/Other/england.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/Landmarks/sydney.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/Landscapes/utah.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/Text/LA.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/Other/wisconsin.png)
\(a\)Landmarks\(b\)Landscapes\(c\)Text\(d\)Other
Figure 4:Natural image saliency maps generated using GeoCLIP and CLIP Surgery, grouped into four categories based on the visual cues the model attends to for geolocation prediction, which includea\) landmarks: recognizable architectural or cultural structures, b\) landscapes: natural features such as terrain, vegetation patterns, and sky, c\) text: textual cues embedded in the scene, such as signage or place names, and d\) other: including street furniture, vegetation, etc\.For SatCLIP, we based our observations on satellite imagery from urban areas, which tend to have greater diversity of landscape elements compared to rural areas\. We selected four cities—Paris, Amsterdam, New York, and St\. Louis—to investigate cross\-location patterns, specifically whether European cities exhibit distinct structural elements compared to their American counterparts, and whether the model captures any meaningful similarities or differences between them\. The method for computing SatCLIP saliency maps was identical to that used for GeoCLIP, except that we considered all Sentinel\-2 multispectral bands rather than only RGB channels\.

We performed an additional analysis step for the satellite imagery visualizations, since direct maps on satellite imagery are more difficult to interpret than those on natural imagery\. We computed saliency maps for 50 images within a city’s radius to capture diverse land cover types \(urban, suburban, rural, etc\.\), and for each image we cropped the regions the model attended to\. To obtain these crops, we applied a binary mask to the image using the saliency map output scores, then normalized the mask so that each detected region was of roughly equal size\. A bounding box was drawn around each region and cropped accordingly\. These crops were then automatically grouped by computing their embeddings and applying dimensionality reduction and clustering via t\-SNE\[[52](https://arxiv.org/html/2606.24997#bib.bib65)\]and DBSCAN\[[15](https://arxiv.org/html/2606.24997#bib.bib66)\]respectively\. For the embeddings, we used a ResNet\-50\[[17](https://arxiv.org/html/2606.24997#bib.bib67)\]pretrained on ImageNet\[[11](https://arxiv.org/html/2606.24997#bib.bib68)\], rather than the SatCLIP image encoder, in order to capture purely RGB visual similarities without any bias from satellite\-specific training\. Figure[5](https://arxiv.org/html/2606.24997#S5.F5)visualizes our pipeline to extract and group these visual elements and Figure[6](https://arxiv.org/html/2606.24997#S5.F6)shows examples of images taken from each saliency cluster\.Additional saliency results for other land cover types, including deserts and rainforests, are provided in Appendix[7\.5\.3](https://arxiv.org/html/2606.24997#S7.SS5.SSS3)\.

![[Uncaptioned image]](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes_pipeline/saliency.png)![[Uncaptioned image]](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes_pipeline/original_mask.png)![[Uncaptioned image]](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes_pipeline/mask.png)![[Uncaptioned image]](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes_pipeline/bbox.png)![[Uncaptioned image]](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes_pipeline/cluster.png)Saliency MapBinary maskMask normalizationBounding BoxesClusteringCompute saliency maps using CLIP SurgeryUse minimum threshold to draw binary maskSplit the individual masks until max area is achievedDraw bounding boxes surrounding each areaCluster all bounding boxes from all images in a specific radius

Figure 5:Overview of the visual element extraction pipeline for satellite location embeddingsStarting from a raw image, we compute CLIP\-Surgery saliency maps, threshold them into binary masks, normalize and split regions by area, fit bounding boxes around salient regions, and finally cluster all boxes across images within a fixed spatial radius\.
![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/fields/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/fields/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/fields/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/bridges/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/bridges/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/bridges/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/stadium/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/stadium/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/stadium/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/roundabouts/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/roundabouts/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/roundabouts/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/streets/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/streets/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/streets/3.png)

\(a\)Paris
![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/fields/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/fields/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/fields/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/bridges/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/bridges/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/bridges/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/city-water/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/city-water/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/city-water/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/whites/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/whites/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/whites/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/water/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/water/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/water/3.png)

\(b\)Amsterdam
![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/water/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/water/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/water/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/bridge/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/bridge/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/bridge/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/golf/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/golf/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/golf/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/whites/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/whites/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/whites/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/houses/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/houses/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/houses/3.png)

\(c\)New YorkCity
![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/roads/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/roads/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/roads/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/houses/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/houses/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/houses/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/suburbs/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/suburbs/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/suburbs/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/whites/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/whites/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/whites/3.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/water/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/water/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/water/3.png)

\(d\)St\. LouisFigure 6:Extracted visual elements from satellite saliency maps for four cities: Paris, Amsterdam, New York City and St\. Louis\.Although clusters are derived independently per city, similar clusters were manually aligned across cities for comparison\. Original saliency map examples for each city can be found at Appendix Section[7\.5\.3](https://arxiv.org/html/2606.24997#S7.SS5.SSS3)\.

### 5\.2Comparing Feature Attributions in Natural and Satellite imagery

Our initial impressions are that saliency maps on location encoders trained on natural images, such as GeoCLIP, are easier to interpret directly\. The location encoder models tend to attend to specific architectural elements in landmarks that are highly distinctive to a particular location, when available\. Similarly, for landscapes, the model focuses on a broad range of natural features such as trees, rocks, and sky\. Particularly interesting are cases where the model attends to more nuanced image features, such as distinctive lamp posts or tree species,which may serve as implicit geolocation cues \(direct image encoders have been shown to pick up on such small differences, like building facades\[[12](https://arxiv.org/html/2606.24997#bib.bib72)\]\)\. Also noteworthy are cases where textual cues such as street signs are present\.

This motivated our additional analysis step to visualize trends in the saliency maps for satellite imagery\. In this modality, the model’s attention tends to concentrate on structural boundaries: road patterns \(such as roundabouts\), land–water edges \(rivers, ports\), and large urban structures such as bridges and stadiums\. Patterns are broadly consistent across European and American cities, yet distinctive elements emerge in each—such as roundabouts prevalent in European cities, and suburban housing clusters and large industrial or storage facilities more typical of American cities\.

## 6Discussion

We focus on geographic INRs as an especially challenging case for explainability since, unlike direct embeddings of Earth observation data, they learn continuous representations of multi\-scale spatial patterns without directly corresponding to a single input image\. We analyze the location embeddings resulting from these geographic INRs through three complementary decomposition strategies: 1\) latent concepts, 2\) natural language concepts, and 3\) visual features\. These explainability methods can help to demystify the inner workings of location encoders and support a greater understanding and fidelity in different location embeddings\. This transparency can guide users in selecting which location encoder is best suited for a given use case\. For model developers, the combination of explainability methods enables debugging and analyzing model performance\. For example, SpLiCE can be used to identify concept gaps in the location embeddings, while CLIP Surgery can be used to identify if the model is relying on irrelevant visual features or to identify low\-quality data samples and improve the dataset\. From a debugging perspective, in Appendix Figure[8](https://arxiv.org/html/2606.24997#S7.F8), we illustrate that the sparse encoder analysis can be helpful to identify visual artifacts in the dataset\.

The three decomposition methods yield different, but largely consistent interpretations of location embeddings, evidenced by monosemantic neurons representing rainforests and deserts, polysemantic neurons specific to a local area, geospatial text concepts capturing city\-level and regional climatic or terrain semantics, and landmarks or distinct visual characteristics of natural and satellite imagery in saliency maps\.The explainability methods also highlight differences between existing location encoders\. Structure in latent and natural language concepts reflect the influence of pretraining data:some explanations of location encoder structure highlight fine\-grained features such as urban information, while others capture signals of broader geographic attributes like climate zone or biome\.Consistent with these findings,saliency mapsemphasize key scene elements in pretraining data space—street furniture, trees, and signage—in natural imagery, while for satellite imagery, they tend to concentrate on structural boundaries such asroads, urban structures and water\.

Interestingly, the reconstruction results using sparse latent concepts and natural language concepts demonstrate that these decompositions retain significant information content of the original embeddings\. For example,we find that decomposition into latent concepts leads to sparse decompositions that can reconstruct the original location embedding \(yielding anR2R^\{2\}score\>0\.9\>0\.9for location embeddings from SatCLIP and Climplicit\)\.This is consistent with past work that showed the intrinsic dimension of locations embeddings is generally much lower than their ambient dimension\[[43](https://arxiv.org/html/2606.24997#bib.bib13)\], while adding a constructive approach for how one might generate sparse, interpretable concepts either during pretraining or as a post\-training step\.

##### Limitations & Future Work\.

While an important step toward explaining black\-box location encoders, our study has several limitations worth noting\. Across all three families of methods, interpretability requires some degree of manual effort—for instance, the manual inspection of neuron activations or saliency maps, the steps of aligning the embeddings space of location and text embeddings, or clustering hotspots in satellite image saliency maps\.Some of these efforts could be avoided with an interpretable\-by\-design framework, e\.g\., a pretrained text\-location embedding space would enable direct decomposition into natural concepts via SpLiCE\.Thus, while we find that level of interpretability varies noticeably across location encoder models for particular methods, this may not be a fundamental limitation of the methodology, so much as the need to very carefully tailor each type of explainability analysis for location encoders\.These analyses highlight critical areas for future work, such as using the interpretation of location embeddings for downstream tasks in the presence of polysemantic concepts\.Here, we opted for a broad analysis to understand what reasonable modifications of existing explainability techniques provide when applied to location encoders\. We hope our analyses lay a foundation for deeper tailoring of any one of these methodsmore specifically to geospatial models\(or development of new methods to address current gaps\) in future work\.

## Acknowledgements

This material is based upon work supported by the NSF Graduate Research Fellowship under Grant No\. DGE 2040434\. We would like to acknowledge use of Jetstream2 at Indiana University through allocation CIS240692 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support \(ACCESS\) program, which is supported by National Science Foundation grants \#2138259, \#2138286, \#2138307, \#2137603, and \#2138296\. Further, the work of Ivica Obadic is supported by the ML4Earth project of the German Federal Ministry for Economic Affairs and Energy under grant number 50EE2201C\.

We also note the use of LLMS to help make our text more concise and generate figure layouts in`tikz`, such as[Figure˜1](https://arxiv.org/html/2606.24997#S1.F1), and table layouts\. However, all results are manually input and verified by the authors\.

## References

- \[1\]J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya, F\. L\. Aleman, D\. Almeida, J\. Altenschmidt, S\. Altman, S\. Anadkat,et al\.\(2023\)GPT\-4 technical report\.arXiv preprint arXiv:2303\.08774\.Cited by:[§4](https://arxiv.org/html/2606.24997#S4.SS0.SSS0.Px1.p2.2)\.
- \[2\]\(2023\)Fairness and representation in satellite\-based poverty maps: evidence of urban\-rural disparities and their impacts on downstream policy\.InProceedings of the Thirty\-Second International Joint Conference on Artificial Intelligence,IJCAI ’23\.External Links:[Link](https://doi.org/10.24963/ijcai.2023/653),[Document](https://dx.doi.org/10.24963/ijcai.2023/653)Cited by:[§1](https://arxiv.org/html/2606.24997#S1.p2.1.2)\.
- \[3\]D\. Bahdanau, K\. Cho, and Y\. Bengio\(2015\)Neural machine translation by jointly learning to align and translate\.InProceedings of the 3rd International Conference on Learning Representations \(ICLR\),External Links:[Link](http://arxiv.org/abs/1409.0473)Cited by:[§2\.2](https://arxiv.org/html/2606.24997#S2.SS2.p1.1)\.
- \[4\]I\. F\. Benavides\-Martinez, J\. Guthrie, J\. E\. Arias, Y\. A\. Garces\-Gomez, A\. I\. Guzman\-Alvis, C\. V\. Portilla\-Cabrera, S\. Mondal, A\. J\. Allyn, and A\. R\. Ganguly\(2026\)What on Earth is AlphaEarth? Hierarchical structure and functional interpretability for global land cover\.arXiv preprint arXiv:2603\.16911\.Cited by:[§2\.2](https://arxiv.org/html/2606.24997#S2.SS2.p2.1)\.
- \[5\]U\. Bhalla, A\. Oesterling, S\. Srinivas, F\. P\. Calmon, and H\. Lakkaraju\(2024\)Interpreting clip with sparse linear concept embeddings \(splice\)\.37,pp\. 84298–84328\.External Links:[Document](https://dx.doi.org/10.52202/079017-2678),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/996bef37d8a638f37bdfcac2789e835d-Paper-Conference.pdf)Cited by:[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2606.24997#S4.SS0.SSS0.Px1.p2.2),[§4](https://arxiv.org/html/2606.24997#S4.p1.1)\.
- \[6\]C\. F\. Brown, M\. Kazmierski, V\. J\. Pasquarella, W\. J\. Rucklidge, M\. Samsikova, C\. Zhang, E\. Shelhamer, E\. Lahera, O\. Wiles, S\. Ilyushchenko, N\. Gorelick, L\. Zhang, S\. Alj, E\. Schechter, S\. Askay, O\. Guinan, R\. Moore, A\. Boukouvalas, and P\. Kohli\(2025\)AlphaEarth foundations: an embedding field model for accurate and efficient global mapping from sparse label data\.arXiv preprint arXiv:2507\.22291\.Cited by:[§2\.2](https://arxiv.org/html/2606.24997#S2.SS2.p2.1)\.
- \[7\]B\. Bussmann, P\. Leask, and N\. Nanda\(2024\)BatchTopK sparse autoencoders\.InNeurIPS 2024 Workshop on Scientific Methods for Understanding Deep Learning,External Links:[Link](https://openreview.net/forum?id=d4dpOCqybL)Cited by:[§3](https://arxiv.org/html/2606.24997#S3.p2.2)\.
- \[8\]L\. Chen, J\. Chen, H\. Hajimirsadeghi, and G\. Mori\(2020\)Adapting Grad\-CAM for embedding networks\.InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,pp\. 2794–2803\.Cited by:[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px3.p1.1)\.
- \[9\]G\. Christie, N\. Fendley, J\. Wilson, and R\. Mukherjee\(2018\)Functional map of the world\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,pp\. 6172–6180\.Cited by:[Table 6](https://arxiv.org/html/2606.24997#S7.T6.5.5.4)\.
- \[10\]E\. Cole, G\. V\. Horn, C\. Lange, A\. Shepard, P\. Leary, P\. Perona, S\. Loarie, and O\. Mac Aodha\(2023\-23–29 Jul\)Spatial implicit neural representations for global\-scale species mapping\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 6320–6342\.External Links:[Link](https://proceedings.mlr.press/v202/cole23a.html)Cited by:[§1](https://arxiv.org/html/2606.24997#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.24997#S2.SS1.p1.1)\.
- \[11\]J\. Deng, W\. Dong, R\. Socher, L\. Li, K\. Li, and L\. Fei\-Fei\(2009\)ImageNet: a large\-scale hierarchical image database\.In2009 IEEE Conference on Computer Vision and Pattern Recognition,Vol\.,pp\. 248–255\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2009.5206848)Cited by:[§5\.1](https://arxiv.org/html/2606.24997#S5.SS1.p3.1)\.
- \[12\]C\. Doersch, S\. Singh, A\. Gupta, J\. Sivic, and A\. A\. Efros\(2015\-11\)What makes Paris look like Paris?\.Commun\. ACM58\(12\),pp\. 103–110\.External Links:[Link](https://doi.org/10.1145/2830541),[Document](https://dx.doi.org/10.1145/2830541)Cited by:[§5\.2](https://arxiv.org/html/2606.24997#S5.SS2.p1.1)\.
- \[13\]J\. Dollinger, D\. Robert, E\. Plekhanova, L\. Drees, and J\. D\. Wegner\(2025\)Climplicit: climatic implicit embeddings for global ecological tasks\.InICLR 2025 Workshop on Tackling Climate Change with Machine Learning,External Links:[Link](https://www.climatechange.ai/papers/iclr2025/44)Cited by:[§2\.1](https://arxiv.org/html/2606.24997#S2.SS1.p1.1),[Table 6](https://arxiv.org/html/2606.24997#S7.T6.5.4.1)\.
- \[14\]A\. Dubey, V\. Ramanathan, A\. Pentland, and D\. Mahajan\(2021\-06\)Adaptive methods for real\-world domain generalization\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 14340–14349\.Cited by:[§3\.1](https://arxiv.org/html/2606.24997#S3.SS1.p1.2),[§7\.3](https://arxiv.org/html/2606.24997#S7.SS3.SSS0.Px1.p1.2),[§7\.3](https://arxiv.org/html/2606.24997#S7.SS3.SSS0.Px3.p2.1),[Table 7](https://arxiv.org/html/2606.24997#S7.T7.1.3.2)\.
- \[15\]M\. Ester, H\. Kriegel, J\. Sander, and X\. Xu\(1996\)A density\-based algorithm for discovering clusters in large spatial databases with noise\.InProceedings of the Second International Conference on Knowledge Discovery and Data Mining,KDD’96,pp\. 226–231\.Cited by:[§5\.1](https://arxiv.org/html/2606.24997#S5.SS1.p3.1)\.
- \[16\]D\. Faget, J\. L\. Lisani, and M\. Colom\(2026\)Combi\-CAM: a novel multi\-layer approach for explainable image geolocalization\.In21st International Conference on Computer Vision Theory and Applications,Vol\.1,pp\. 275–281\.Cited by:[§7\.4\.3](https://arxiv.org/html/2606.24997#S7.SS4.SSS3.p8.4)\.
- \[17\]K\. He, X\. Zhang, S\. Ren, and J\. Sun\(2016\)Deep residual learning for image recognition\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 770–778\.Cited by:[§5\.1](https://arxiv.org/html/2606.24997#S5.SS1.p3.1)\.
- \[18\]E\. Hernandez, S\. Schwettmann, D\. Bau, T\. Bagashvili, A\. Torralba, and J\. Andreas\(2022\)Natural language descriptions of deep features\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=NudBMY-tzDr)Cited by:[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px2.p1.1)\.
- \[19\]A\. Höhl, I\. Obadic, M\. Fernández\-Torres, H\. Najjar, D\. A\. B\. Oliveira, Z\. Akata, A\. Dengel, and X\. X\. Zhu\(2024\)Opening the black box: a systematic review on explainable artificial intelligence in remote sensing\.IEEE Geoscience and Remote Sensing Magazine12\(4\),pp\. 261–304\.External Links:[Document](https://dx.doi.org/10.1109/MGRS.2024.3467001)Cited by:[§2\.2](https://arxiv.org/html/2606.24997#S2.SS2.p1.1)\.
- \[20\]E\. J\. Hu, yelong shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2022\)LoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[§7\.3](https://arxiv.org/html/2606.24997#S7.SS3.SSS0.Px2.p1.1)\.
- \[21\]R\. Huben, H\. Cunningham, L\. R\. Smith, A\. Ewart, and L\. Sharkey\(2024\)Sparse autoencoders find highly interpretable features in language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=F76bwRSLeK)Cited by:[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2606.24997#S3.p1.1)\.
- \[22\]M\. Ivić\(2019\)Artificial intelligence and geospatial analysis in disaster management\.The International Archives of the Photogrammetry, Remote Sensing and Spatial Information SciencesXLII\-3/W8,pp\. 161–166\.External Links:[Link](https://isprs-archives.copernicus.org/articles/XLII-3-W8/161/2019/),[Document](https://dx.doi.org/10.5194/isprs-archives-XLII-3-W8-161-2019)Cited by:[§1](https://arxiv.org/html/2606.24997#S1.p1.1)\.
- \[23\]F\. Jia, L\. Liu, C\. Hou, F\. Zhang, X\. Liu, and Y\. Liu\(2025\)Towards interpretable geo\-localization: a concept\-aware global image\-GPS alignment framework\.arXiv preprint arXiv:2509\.01910\.Cited by:[§2\.2](https://arxiv.org/html/2606.24997#S2.SS2.p1.1)\.
- \[24\]D\. N\. Karger, O\. Conrad, J\. Böhner, T\. Kawohl, H\. Kreft, R\. W\. Soria\-Auza, N\. E\. Zimmermann, H\. P\. Linder, and M\. Kessler\(2017\)Climatologies at high resolution for the earth’s land surface areas\.Scientific data4\(1\),pp\. 1–20\.External Links:[Document](https://dx.doi.org/10.1038/sdata.2017.122),[Link](https://doi.org/10.1038/sdata.2017.122)Cited by:[Table 6](https://arxiv.org/html/2606.24997#S7.T6.5.4.4)\.
- \[25\]K\. Klemmer, E\. Rolf, C\. Robinson, L\. Mackey, and M\. Rußwurm\(2025\)SatCLIP: global, general\-purpose location embeddings with satellite imagery\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 4347–4355\.Cited by:[§1](https://arxiv.org/html/2606.24997#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.24997#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2606.24997#S3.SS1.p1.2),[§7\.3](https://arxiv.org/html/2606.24997#S7.SS3.SSS0.Px1.p1.2),[Table 6](https://arxiv.org/html/2606.24997#S7.T6.5.3.1),[Table 6](https://arxiv.org/html/2606.24997#S7.T6.5.3.4),[Table 7](https://arxiv.org/html/2606.24997#S7.T7.1.2.2),[Table 7](https://arxiv.org/html/2606.24997#S7.T7.1.2.4)\.
- \[26\]K\. Klemmer, E\. Rolf, M\. Russwurm, G\. Camps\-Valls, M\. Czerkawski, S\. Ermon, A\. Francis, N\. Jacobs, H\. R\. Kerner, L\. Mackey, G\. Mai, O\. Mac Aodha, M\. Reichstein, C\. Robinson, D\. Rolnick, E\. Shelhamer, V\. Sitzmann, D\. Tuia, and X\. Zhu\(2025\)Earth embeddings: towards AI\-centric representations of our planet\.EarthArXiv\.External Links:[Document](https://dx.doi.org/10.31223/X5HX9S),[Link](https://doi.org/10.31223/X5HX9S)Cited by:[§1](https://arxiv.org/html/2606.24997#S1.p1.1),[§1](https://arxiv.org/html/2606.24997#S1.p1.1.2),[§2\.1](https://arxiv.org/html/2606.24997#S2.SS1.p1.1)\.
- \[27\]R\. Lam, A\. Sanchez\-Gonzalez, M\. Willson, P\. Wirnsberger, M\. Fortunato, F\. Alet, S\. Ravuri, T\. Ewalds, Z\. Eaton\-Rosen, W\. Hu, A\. Merose, S\. Hoyer, G\. Holland, O\. Vinyals, J\. Stott, A\. Pritzel, S\. Mohamed, and P\. Battaglia\(2023\)Learning skillful medium\-range global weather forecasting\.Science382\(6677\),pp\. 1416–1421\.External Links:[Document](https://dx.doi.org/10.1126/science.adi2336),[Link](https://www.science.org/doi/abs/10.1126/science.adi2336)Cited by:[§1](https://arxiv.org/html/2606.24997#S1.p1.1)\.
- \[28\]M\. Larson, M\. Soleymani, G\. Gravier, B\. Ionescu, and G\. J\.F\. Jones\(2017\)The benchmarking initiative for multimedia evaluation: mediaeval 2016\.IEEE MultiMedia24\(1\),pp\. 93–96\.External Links:[Document](https://dx.doi.org/10.1109/MMUL.2017.9)Cited by:[Table 6](https://arxiv.org/html/2606.24997#S7.T6.5.2.4)\.
- \[29\]Y\. Li, H\. Wang, Y\. Duan, H\. Xu, and X\. Li\(2022\)Exploring visual interpretability for contrastive language\-image pre\-training\.arXiv preprint arXiv:2209\.07046\.Cited by:[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px3.p1.1.1),[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px3.p1.1.2)\.
- \[30\]Y\. Li, H\. Wang, Y\. Duan, J\. Zhang, and X\. Li\(2025\)A closer look at the explainability of contrastive language\-image pre\-training\.Pattern Recognition162,pp\. 111409\.External Links:ISSN 0031\-3203,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.patcog.2025.111409),[Link](https://www.sciencedirect.com/science/article/pii/S003132032500069X)Cited by:[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px3.p1.1),[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px3.p1.1.1),[§5](https://arxiv.org/html/2606.24997#S5.p1.1)\.
- \[31\]V\. W\. Liang, Y\. Zhang, Y\. Kwon, S\. Yeung, and J\. Zou\(2022\)Mind the gap: understanding the modality gap in multi\-modal contrastive representation learning\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 17612–17625\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/702f4db7543a7432431df588d57bc7c9-Paper-Conference.pdf)Cited by:[§4](https://arxiv.org/html/2606.24997#S4.SS0.SSS0.Px1.p2.2)\.
- \[32\]C\. Liu, K\. Chen, R\. Zhao, Z\. Zou, and Z\. Shi\(2025\)Text2Earth: unlocking text\-driven remote sensing image generation with a global\-scale dataset and a foundation model\.IEEE Geoscience and Remote Sensing Magazine13\(3\),pp\. 238–259\.External Links:[Document](https://dx.doi.org/10.1109/MGRS.2025.3560455)Cited by:[§4](https://arxiv.org/html/2606.24997#S4.SS0.SSS0.Px1.p2.2)\.
- \[33\]S\. M\. Lundberg and S\. Lee\(2017\)A unified approach to interpreting model predictions\.InAdvances in Neural Information Processing Systems,Vol\.30\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2017/file/8a20a8621978632d76c43dfd28b67767-Paper.pdf)Cited by:[§2\.2](https://arxiv.org/html/2606.24997#S2.SS2.p1.1)\.
- \[34\]O\. Mac Aodha, E\. Cole, and P\. Perona\(2019\)Presence\-only geographical priors for fine\-grained image classification\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 9596–9606\.Cited by:[§2\.1](https://arxiv.org/html/2606.24997#S2.SS1.p1.1)\.
- \[35\]G\. Mai, K\. Janowicz, Y\. Hu, S\. Gao, B\. Yan, R\. Zhu, L\. Cai, and N\. Lao\(2022\)A review of location encoding for geoai: methods and applications\.International Journal of Geographical Information Science36\(4\),pp\. 639–673\.External Links:[Document](https://dx.doi.org/10.1080/13658816.2021.2004602),[Link](https://doi.org/10.1080/13658816.2021.2004602)Cited by:[§1](https://arxiv.org/html/2606.24997#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.24997#S2.SS1.p1.1)\.
- \[36\]G\. Mai, K\. Janowicz, B\. Yan, R\. Zhu, L\. Cai, and N\. Lao\(2020\)Multi\-scale representation learning for spatial feature distributions using grid cells\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=rJljdh4KDH)Cited by:[Table 6](https://arxiv.org/html/2606.24997#S7.T6.5.5.2)\.
- \[37\]L\. McInnes, J\. Healy, N\. Saul, and L\. Großberger\(2018\)UMAP: uniform manifold approximation and projection\.The Journal of Open Source Software3\(29\),pp\. 861\.Cited by:[§7\.3](https://arxiv.org/html/2606.24997#S7.SS3.SSS0.Px3.p1.1)\.
- \[38\]M\. Moayeri, K\. Rezaei, M\. Sanjabi, and S\. Feizi\(2023\)Text\-to\-concept \(and back\) via cross\-model alignment\.InInternational Conference on Machine Learning,pp\. 25037–25060\.Cited by:[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px2.p1.1)\.
- \[39\]T\. Oikarinen and T\. Weng\(2023\)CLIP\-dissect: automatic description of neuron representations in deep vision networks\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=iPWiwWHc1V)Cited by:[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px2.p1.1)\.
- \[40\]M\. Pach, S\. Karthik, Q\. Bouniot, S\. Belongie, and Z\. Akata\(2025\)Sparse autoencoders learn monosemantic features in vision\-language models\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 95706–95742\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/89e83382abeee53b932a6df62edbf9cc-Paper-Conference.pdf)Cited by:[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2606.24997#S3.SS1.p1.2),[§3](https://arxiv.org/html/2606.24997#S3.p1.1),[§3](https://arxiv.org/html/2606.24997#S3.p2.2),[§7\.4\.1](https://arxiv.org/html/2606.24997#S7.SS4.SSS1.p2.3)\.
- \[41\]S\. Pramanick, E\. M\. Nowara, J\. Gleason, C\. D\. Castillo, and R\. Chellappa\(2022\)Where in the world is this image? transformer\-based geo\-localization in the wild\.InEuropean Conference on Computer Vision,pp\. 196–215\.External Links:[Link](https://doi.org/10.1007/978-3-031-19839-7_12)Cited by:[§5\.1](https://arxiv.org/html/2606.24997#S5.SS1.p1.1)\.
- \[42\]A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark, G\. Krueger, and I\. Sutskever\(2021\-18–24 Jul\)Learning transferable visual models from natural language supervision\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 8748–8763\.External Links:[Link](https://proceedings.mlr.press/v139/radford21a.html)Cited by:[§4](https://arxiv.org/html/2606.24997#S4.SS0.SSS0.Px1.p2.2),[§7\.3](https://arxiv.org/html/2606.24997#S7.SS3.SSS0.Px2.p1.1)\.
- \[43\]A\. Rao, M\. Rußwurm, K\. Klemmer, and E\. Rolf\(2026\)Measuring the intrinsic dimension of earth representations\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=gQPD83DrGp)Cited by:[§2\.2](https://arxiv.org/html/2606.24997#S2.SS2.p2.1),[§6](https://arxiv.org/html/2606.24997#S6.p3.2.2)\.
- \[44\]M\. T\. Ribeiro, S\. Singh, and C\. Guestrin\(2016\)“Why should I trust you?": explaining the predictions of any classifier\.InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,KDD ’16,New York, NY, USA,pp\. 1135–1144\.External Links:[Link](https://doi.org/10.1145/2939672.2939778),[Document](https://dx.doi.org/10.1145/2939672.2939778)Cited by:[§2\.2](https://arxiv.org/html/2606.24997#S2.SS2.p1.1)\.
- \[45\]M\. Rußwurm, K\. Klemmer, E\. Rolf, R\. Zbinden, and D\. Tuia\(2024\)Geographic location encoding with spherical harmonics and sinusoidal representation networks\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 1746–1759\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/073c8584ef86bee26fe9d639ec648e28-Paper-Conference.pdf)Cited by:[§2\.1](https://arxiv.org/html/2606.24997#S2.SS1.p1.1),[Table 6](https://arxiv.org/html/2606.24997#S7.T6.5.3.2)\.
- \[46\]R\. R\. Selvaraju, M\. Cogswell, A\. Das, R\. Vedantam, D\. Parikh, and D\. Batra\(2017\)Grad\-CAM: visual explanations from deep networks via gradient\-based localization\.InProceedings of the IEEE International Conference on Computer Vision,pp\. 618–626\.Cited by:[§2\.2](https://arxiv.org/html/2606.24997#S2.SS2.p1.1),[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px3.p1.1)\.
- \[47\]D\. Shu, X\. Wu, H\. Zhao, D\. Rai, Z\. Yao, N\. Liu, and M\. Du\(2025\)A survey on sparse autoencoders: interpreting the internal mechanisms of large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 1690–1712\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.89/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.89)Cited by:[§2\.3](https://arxiv.org/html/2606.24997#S2.SS3.SSS0.Px1.p1.1)\.
- \[48\]V\. Sitzmann, J\. Martel, A\. Bergman, D\. Lindell, and G\. Wetzstein\(2020\)Implicit neural representations with periodic activation functions\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 7462–7473\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/53c04118df112c13a8c34b38343b9c10-Paper.pdf)Cited by:[§2\.1](https://arxiv.org/html/2606.24997#S2.SS1.p1.1)\.
- \[49\]A\. J\. Stewart, C\. Robinson, I\. A\. Corley, A\. Ortiz, J\. M\. Lavista Ferres, and A\. Banerjee\(2025\)TorchGeo: deep learning with geospatial data\.ACM Trans\. Spatial Algorithms Syst\.11\(4\)\.External Links:ISSN 2374\-0353,[Link](https://doi.org/10.1145/3707459),[Document](https://dx.doi.org/10.1145/3707459)Cited by:[§7\.1](https://arxiv.org/html/2606.24997#S7.SS1.p1.1)\.
- \[50\]M\. Tancik, P\. Srinivasan, B\. Mildenhall, S\. Fridovich\-Keil, N\. Raghavan, U\. Singhal, R\. Ramamoorthi, J\. Barron, and R\. Ng\(2020\)Fourier features let networks learn high frequency functions in low dimensional domains\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 7537–7547\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/55053683268957697aa39fba6f231c68-Paper.pdf)Cited by:[Table 6](https://arxiv.org/html/2606.24997#S7.T6.5.2.2)\.
- \[51\]J\. Tolan, H\. Yang, B\. Nosarzewski, G\. Couairon, H\. V\. Vo, J\. Brandt, J\. Spore, S\. Majumdar, D\. Haziza, J\. Vamaraju, T\. Moutakanni, P\. Bojanowski, T\. Johns, B\. White, T\. Tiecke, and C\. Couprie\(2024\)Very high resolution canopy height maps from RGB imagery using self\-supervised vision transformer and convolutional decoder trained on aerial lidar\.Remote Sensing of Environment300,pp\. 113888\.External Links:ISSN 0034\-4257,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.rse.2023.113888),[Link](https://www.sciencedirect.com/science/article/pii/S003442572300439X)Cited by:[§1](https://arxiv.org/html/2606.24997#S1.p1.1)\.
- \[52\]L\. van der Maaten and G\. Hinton\(2008\)Visualizing data using t\-SNE\.Journal of Machine Learning Research9\(86\),pp\. 2579–2605\.External Links:[Link](http://jmlr.org/papers/v9/vandermaaten08a.html)Cited by:[§5\.1](https://arxiv.org/html/2606.24997#S5.SS1.p3.1)\.
- \[53\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.InAdvances in Neural Information Processing Systems,Vol\.30,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf)Cited by:[§2\.2](https://arxiv.org/html/2606.24997#S2.SS2.p1.1)\.
- \[54\]V\. Vivanco Cepeda, G\. K\. Nayak, and M\. Shah\(2023\)GeoCLIP: clip\-inspired alignment between locations and images for effective worldwide geo\-localization\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 8690–8701\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/1b57aaddf85ab01a2445a79c9edc1f4b-Paper-Conference.pdf)Cited by:[§2\.1](https://arxiv.org/html/2606.24997#S2.SS1.p1.1),[§4](https://arxiv.org/html/2606.24997#S4.SS0.SSS0.Px1.p1.1),[§7\.3](https://arxiv.org/html/2606.24997#S7.SS3.SSS0.Px1.p1.2),[Table 6](https://arxiv.org/html/2606.24997#S7.T6.5.2.1),[Table 7](https://arxiv.org/html/2606.24997#S7.T7.1.3.4)\.
- \[55\]Z\. Wang, K\. Lane, L\. Cai, M\. Karimzadeh, and E\. Rolf\(2026\-06\)A proxy consistency loss for grounded fusion of earth observation and location encoders\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\) Workshops,pp\. 8075–8084\.Cited by:[§1](https://arxiv.org/html/2606.24997#S1.p1.1)\.
- \[56\]X\. Zhai, X\. Wang, B\. Mustafa, A\. Steiner, D\. Keysers, A\. Kolesnikov, and L\. Beyer\(2022\-06\)LiT: zero\-shot transfer with locked\-image text tuning\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 18123–18133\.Cited by:[§4](https://arxiv.org/html/2606.24997#S4.SS0.SSS0.Px1.p1.1)\.
- \[57\]B\. Zhou, A\. Khosla, A\. Lapedriza, A\. Oliva, and A\. Torralba\(2016\-06\)Learning deep features for discriminative localization\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§2\.2](https://arxiv.org/html/2606.24997#S2.SS2.p1.1)\.
- \[58\]X\. X\. Zhu, Z\. Xiong, Y\. Wang, A\. J\. Stewart, K\. Heidler, Y\. Wang, Z\. Yuan, T\. Dujardin, Q\. Xu, and Y\. Shi\(2026\-01\-08\)On the foundations of earth foundation models\.Communications Earth & Environment7\(1\),pp\. 103\.External Links:[Document](https://dx.doi.org/10.1038/s43247-025-03127-x),[Link](https://doi.org/10.1038/s43247-025-03127-x)Cited by:[§1](https://arxiv.org/html/2606.24997#S1.p1.1)\.

## 7Appendix

### 7\.1Reproducibility

All code necessary to reproduce the experiments in this paper is available at[https://github\.com/sricke/explainable\-earth\-embeddings](https://github.com/sricke/explainable-earth-embeddings)\. All datasets, model code, and model weights are publicly available and can be downloaded from the links in Tables[3](https://arxiv.org/html/2606.24997#S7.T3)–[5](https://arxiv.org/html/2606.24997#S7.T5)\. While most datasets and models are available under permissive open\-source licenses, many lack any information about licensing\. We have reached out to all authors to encourage them to make the license of their data or models clear\. Datasets and models under open\-source licenses will be contributed to TorchGeo\[[49](https://arxiv.org/html/2606.24997#bib.bib54)\]to further lower the barrier of entry of reproducibility\.

Table 3:Datasets used in this work\.DatasetLicenseURLS2\-100kMIT[https://hf\.co/datasets/kklmmr/s2\-100k](https://hf.co/datasets/kklmmr/s2-100k)Im2GPS?[https://www\.cis\.jhu\.edu/˜shraman/TransLocator](https://www.cis.jhu.edu/~shraman/TransLocator)Git\-10MCC\-BY\-NC\-ND\-4\.0[https://hf\.co/datasets/lcybuaa/Git\-10M](https://hf.co/datasets/lcybuaa/Git-10M)Table 4:Model code used in this work\.ModelLicenseURLResNet\-50BSD\-3\-Clause[https://github\.com/pytorch/vision](https://github.com/pytorch/vision)CSP\-fMoW?[https://github\.com/gengchenmai/csp](https://github.com/gengchenmai/csp)GeoCLIPMIT[https://github\.com/VicenteVivan/geo\-clip](https://github.com/VicenteVivan/geo-clip)SatCLIPMIT[https://github\.com/microsoft/satclip](https://github.com/microsoft/satclip)ClimplicitMIT[https://github\.com/ecovision\-uzh/climplicit](https://github.com/ecovision-uzh/climplicit)Table 5:Model weights used in this work\.ModelLicenseURLResNet\-50?[https://download\.pytorch\.org/models](https://download.pytorch.org/models)CSP\-fMoW?[https://gengchenmai\.github\.io/csp\-website](https://gengchenmai.github.io/csp-website)GeoCLIPMIT[https://github\.com/VicenteVivan/geo\-clip](https://github.com/VicenteVivan/geo-clip)SatCLIPMIT[https://hf\.co/microsoft/SatCLIP\-ViT16\-L10](https://hf.co/microsoft/SatCLIP-ViT16-L10)ClimplicitCC\-BY\-4\.0[https://hf\.co/Jobedo/climplicit](https://hf.co/Jobedo/climplicit)All SAE experiments were performed on an NVIDIA A40 with 46 GB of RAM and 6 parallel workers for data loading\. SatCLIP experiments ran for 2 hours, GeoCLIP experiments ran for 30 minutes, Climplicit experiments ran for 5 minutes, and xAI computation ran for 2 hours per location encoder\. Combined, these experiments ran for 275 minutes for 3 different random seeds, resulting in 825 minutes in total\.

The majority of SpLiCE and CLIP Surgery experiments were performed on an NVIDIA RTX 8000 with 48 GB of RAM\. The location–text alignment of SatCLIP and Climplicit was performed for 55k steps and took approximately 4 days\. GeoCLIP location–text alignment was not trained\. SpLiCE evaluations took around 2 minutes\. Inference of saliency maps takes approximately one second per image for the GeoCLIP and SatCLIP models\. The CSP\-fMoW experiments were performed on an NVIDIA H100 with 80 GB RAM and took 20 hours in total\.

### 7\.2Location Encoders and Embeddings

Table 6:Location encoders analyzed in this work\.ModelEncodingDim\.Pretraining DataGeoCLIP\[[54](https://arxiv.org/html/2606.24997#bib.bib15)\]RFF\[[50](https://arxiv.org/html/2606.24997#bib.bib95)\]512MP\-16 \(Flickr\)\[[28](https://arxiv.org/html/2606.24997#bib.bib91)\]SatCLIP\[[25](https://arxiv.org/html/2606.24997#bib.bib1)\]Spherical Harmonics\[[45](https://arxiv.org/html/2606.24997#bib.bib52)\]256S2\-100K \(Sentinel\-2\)\[[25](https://arxiv.org/html/2606.24997#bib.bib1)\]Climplicit\[[13](https://arxiv.org/html/2606.24997#bib.bib29)\]Sinusoidal1024CHELSA \(Climate Variables\)\[[24](https://arxiv.org/html/2606.24997#bib.bib92)\]CSP\-fMoWGrid\-Cell\[[36](https://arxiv.org/html/2606.24997#bib.bib93)\]512fMoW \(DigitalGlobe Satellite Imagery\)\[[9](https://arxiv.org/html/2606.24997#bib.bib94)\]In[Table˜6](https://arxiv.org/html/2606.24997#S7.T6), we summarize the location encoders analyzed in this work\. We report the pretraining data used to train the location encoder, which is an important consideration when considering the concepts encoded in location embeddings\. We additionally report the dimension of each embedding, which can be compared to the number of concepts enforced in the sparsity constraint of SAEs as well as the number of active concepts resulting from SpLiCE \([Table˜2](https://arxiv.org/html/2606.24997#S4.T2)\)\.

### 7\.3Additional Model Training Details

##### SAE and SpLiCE Evaluation\.

Table 7:Experimental setup for evaluating SAE and SpLiCE Reconstruction\.We randomly sampled 10% of the locations and their corresponding images from each evaluation dataset\.DatasetEval DatasetSamples \(NN\)Visual EncoderUARS2\-100K \(Sentinel\-2\)\[[25](https://arxiv.org/html/2606.24997#bib.bib1)\]9185MoCo ViT\-S/16 \(SatCLIP\[[25](https://arxiv.org/html/2606.24997#bib.bib1)\]\)Human\-visitedGeo\-YFCC \(natural\)\[[14](https://arxiv.org/html/2606.24997#bib.bib81)\]8094CLIP ViT\-L/14 \(GeoCLIP\[[54](https://arxiv.org/html/2606.24997#bib.bib15)\]\)As shown in[Table˜7](https://arxiv.org/html/2606.24997#S7.T7), we used two different datasets to evaluate SAE reconstruction \([Table˜1](https://arxiv.org/html/2606.24997#S3.T1)\) and SpLiCE reconstruction \([Table˜2](https://arxiv.org/html/2606.24997#S4.T2)\)\. The first dataset consists of100,000100,000points distributed uniformly\-at\-random \(UAR\) across the landmasses\. To evaluate the monosemanticity scores of the SAE, we sample 10% from the samples from the S2\-100K dataset, which consists of Sentinel\-2 imagery\. To compute the monosemanticity scores, we use embeddings from the SatCLIP visual encoder \(MoCo ViT\-S/16\[[25](https://arxiv.org/html/2606.24997#bib.bib1)\]\)\. The second dataset we use comprises100,000100,000points sampled from the Geo\-YFCC dataset\[[14](https://arxiv.org/html/2606.24997#bib.bib81)\], which contains natural imagery fromhuman\-visitedlocations\. We also sample 10% of the points from this dataset to evaluate the visual monosemanticity, and use the GeoCLIP image encoder \(CLIP ViT\-L/14\[[54](https://arxiv.org/html/2606.24997#bib.bib15)\]\) to obtain the image embeddings\.

##### Location–Text Alignment

To create a shared location–text space, we aligned the OpenCLIP ViT\-L text encoder\[[42](https://arxiv.org/html/2606.24997#bib.bib63)\]to each location encoder \(SatCLIP, Climplicit, CSP\-fMoW\)\. We train the text encoder using a symmetric CLIP loss while the location encoder is kept frozen\. A linear projection head maps text embeddings to the shared representation space \(with the same dimensionality of the embeddings from the location encoder\)\. The text encoder is adapted using LoRA\[[20](https://arxiv.org/html/2606.24997#bib.bib69)\]\(rank 8 applied to the final 8 layers\)\.

Optimization is performed using AdamW with a learning rate of1×10−41\\times 10^\{\-4\}and weight decay of0\.010\.01\. Training uses a batch size of 2048 with44gradient accumulation steps for an effective batch size of81928192\. We use a cosine learning rate schedule\. The logit temperature is learnable and is initialized at0\.070\.07\.

We train for 55k steps, validating every 500 steps and applying early stopping based on validation loss \(with a patience of 5 validation checks\)\. Training is conducted on NVIDIA RTX 8000s and H100s\.

##### Assessing and Visualizing Location–Text Alignment\.

We show visualizations of alignment in[Figure˜7](https://arxiv.org/html/2606.24997#S7.F7)for GeoCLIP and SatCLIP to assess the success of the location–text alignment training for SatCLIP\. UMAPs are created by first applying PCA to project down to 30 dimensions \(to reduce noise\), and then applying the UMAP algorithm\[[37](https://arxiv.org/html/2606.24997#bib.bib83)\]with 200 neighbors and minimum distance0\.10\.1to capture global alignment structure\.

For GeoCLIP, we display the UMAP results using the Geo\-YFCC concept set, which is generated by taking the 10,000 most common concepts from the Geo\-YFCC dataset\[[14](https://arxiv.org/html/2606.24997#bib.bib81)\]\. We visualize the GeoCLIP aligned space with this concept set, as the OpenCLIP ViT\-L text encoder was trained on scraped image\-text pairs, more closely mirroring the captions in Geo\-YFCC\. This concept set is overcomplete, containing non\-geospatial concepts\. For SpLiCE decompositions, we default to using a geospatial concept set \(described in the main text\)\.

In these UMAP visualizations, we can see overlap between concepts and certain locations\. Manual inspection of these figures can qualitatively show how aligned these spaces are\. For example, for the GeoCLIP UMAP with the Geo\-YFCC concepts, locations in Western North America have overlap with concepts such as “rocky mountains” and specific states in the Western US, whereas locations in South America correspond to concepts such as “andes” for the mountain range\. On the other hand, the Git\-10M concepts contain less location\-specific terms, leading to overlap in concepts and locations that’s less immediately interpretable in the SatCLIP UMAP\. Additionally, in the GeoCLIP UMAP visualization, we see a cluster of solely concepts in the center, which tend to be non\-geospatial terms\. We don’t see this for SatCLIP as we restrict ourselves to exclusively geospatial terms\.

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/umaps/geoclip_umap.png)\(a\)GeoCLIP \(Geo\-YFCC concepts\)
![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/umaps/satclip_umap.png)\(b\)SatCLIP \(Git\-10M concepts\)

Figure 7:UMAP visualizations of GeoCLIP and SatCLIP embeddings\(with respective concept sets\)\. Location embeddings are generated from a dense grid of100,000100,000points sampled uniformly at random over the landmasses\.
##### Evaluation Regions for SpLiCE Concepts\.

We define four geographic regions using publicly available shapefiles\. The Russia and Indonesia shapefiles are sourced from GADM \([https://gadm\.org/download\_country\.html](https://gadm.org/download_country.html)\)\. Siberia is constructed by dissolving 12 level\-1 Russian administrative regions \(gadm41\_RUS\_1\), including Krasnoyarsk, Irkutsk, Buryat, Altay, and Tomsk, among others\. Bali is extracted from the level\-1 Indonesian province shapefile \(gadm41\_IDN\_1\), selecting the province named “Bali\.” The Sahara shapefile is sourced from GISCarta \([https://map\.giscarta\.com](https://map.giscarta.com/)\) and reprojected to WGS84 \(EPSG:4326\)\. The Paris shapefile is derived from the French commune boundaries dataset on data\.gouv\.fr \([https://www\.data\.gouv\.fr/datasets/contours\-administratifs](https://www.data.gouv.fr/datasets/contours-administratifs)\), selecting the commune named “Paris\.”

### 7\.4Method Details

#### 7\.4\.1Sparse AutoEncoders for Reconstructing Earth Embeddings

Let𝐱∈ℝde​e\\mathbf\{x\}\\in\\mathbb\{R\}^\{d\_\{ee\}\}denote an location embedding with dimensionde​ed\_\{ee\}\(e​eeestands for “Earth embedding”\)\. The SAE first maps𝐱\\mathbf\{x\}into a higher\-dimensional latent space of dimensiondld\_\{l\}via an encoder\-decoder architecture\. The encoder computes a sparse latent representation𝐳=fenc​\(𝐱\):=σ​\(𝐖enc⊤​\(𝐱−𝐛\)\)\\mathbf\{z\}=f\_\{\\mathrm\{enc\}\}\(\\mathbf\{x\}\):=\\sigma\(\\mathbf\{W\}\_\{\\text\{enc\}\}^\{\\top\}\(\\mathbf\{x\}\-\\mathbf\{b\}\)\), where𝐖enc∈ℝde​e×dl\\mathbf\{W\}\_\{\\text\{enc\}\}\\in\\mathbb\{R\}^\{d\_\{ee\}\\times d\_\{l\}\}andσ\\sigmais applied elementwise\. The decoder reconstructs the embedding as𝐱^=fdec​\(𝐳\):=𝐖dec⊤​𝐳\+𝐛\\hat\{\\mathbf\{x\}\}=f\_\{\\mathrm\{dec\}\}\(\\mathbf\{z\}\):=\\mathbf\{W\}\_\{\\text\{dec\}\}^\{\\top\}\\mathbf\{z\}\+\\mathbf\{b\}\. The columns of𝐖dec⊤\\mathbf\{W\}\_\{\\text\{dec\}\}^\{\\top\}define a learned dictionary of feature directions\{𝐜i\}i=1dl\\\{\\mathbf\{c\}\_\{i\}\\\}\_\{i=1\}^\{d\_\{l\}\}in embedding space such that the reconstruction can be expressed as a sparse linear combination𝐱^=∑i=1dlzi​𝐜i\+𝐛\\hat\{\\mathbf\{x\}\}=\\sum\_\{i=1\}^\{d\_\{l\}\}z\_\{i\}\\mathbf\{c\}\_\{i\}\+\\mathbf\{b\}\. In this formulation, each location embedding is represented using only a small subset of dictionary elements, with sparsity enforced in the latent embedding𝐳\\mathbf\{z\}\.

The full SAE \(encoder and decoder\) is trained using a reconstruction loss combined with a sparsity penalty,ℒ​\(𝐱\):=ℛ​\(𝐱\)\+λ​𝒮​\(𝐱\)\\mathcal\{L\}\(\\mathbf\{x\}\):=\\mathcal\{R\}\(\\mathbf\{x\}\)\+\\lambda\\mathcal\{S\}\(\\mathbf\{x\}\), whereλ\\lambdacontrols the tradeoff between sparsity and reconstruction accuracy\. After training, the decoder matrix serves as a learned dictionary of features: each column corresponds to a direction in embedding space that captures a distinct, approximately monosemantic\[[40](https://arxiv.org/html/2606.24997#bib.bib14)\]feature, while the sparse latent representation𝐳=fenc​\(𝐱\)\\mathbf\{z\}=f\_\{\\mathrm\{enc\}\}\(\\mathbf\{x\}\)determines which features are active for a given input\.

#### 7\.4\.2Sparse Linear Concept Embeddings \(SpLiCE\)

To apply SpLiCE, we use apredefinedconcept dictionary,𝐂∈ℝnc×de​e\\mathbf\{C\}\\in\\mathbb\{R\}^\{n\_\{c\}\\times d\_\{ee\}\}, wherencn\_\{c\}is the number of concepts\. Each row in𝐂\\mathbf\{C\}corresponds to an embedding in the same space as the location embedding𝐱∈ℝde​e\\mathbf\{x\}\\in\\mathbb\{R\}^\{d\_\{ee\}\}\. This method uses a centered and normalized concept dictionary, given by𝐂=\[σ​\(ftext​\(𝐱1con\)−μtext\),…,σ​\(ftext​\(𝐱nccon\)−μtext\)\]\\mathbf\{C\}=\[\\sigma\(f\_\{\\text\{text\}\}\(\\mathbf\{x\}\_\{1\}^\{\\text\{con\}\}\)\-\\mu\_\{\\text\{text\}\}\),\\ldots,\\sigma\(f\_\{\\text\{text\}\}\(\\mathbf\{x\}\_\{n\_\{c\}\}^\{\\text\{con\}\}\)\-\\mu\_\{\\text\{text\}\}\)\], whereftextf\_\{\\text\{text\}\}denotes the text encoder,𝐱icon\\mathbf\{x\}\_\{i\}^\{\\text\{con\}\}a text concept,μtext\\mu\_\{\\text\{text\}\}the \(estimated\) mean of text embeddings, andσ​\(⋅\)\\sigma\(\\cdot\)denotes normalization\. Given an input embedding𝐱\\mathbf\{x\}, SpLiCE presents the following objective

min𝐰∈ℝ\+nc⁡‖𝐰‖0​s\.t\.​⟨𝐱,σ​\(𝐂𝐰\)⟩≥1−ε\.\\min\_\{\\mathbf\{w\}\\in\\mathbb\{R\}\_\{\+\}^\{n\_\{c\}\}\}\\\|\\mathbf\{w\}\\\|\_\{0\}\\,\\text\{s\.t\.\}\\,\\langle\\mathbf\{x\},\\sigma\(\\mathbf\{C\}\\mathbf\{w\}\)\\rangle\\geq 1\-\\varepsilon\.\(1\)In practice, SpLiCE computes a sparse concept representation𝐰∈ℝnc\\mathbf\{w\}\\in\\mathbb\{R\}^\{n\_\{c\}\}by solving the following relaxation:

min𝐰∈ℝd⁡‖𝐂⊤​𝐰−𝐱‖22\+λ​‖𝐰‖1\\min\_\{\\mathbf\{w\}\\in\\mathbb\{R\}^\{d\}\}\\\|\\mathbf\{C\}^\{\\top\}\\mathbf\{w\}\-\\mathbf\{x\}\\\|\_\{2\}^\{2\}\+\\lambda\\\|\\mathbf\{w\}\\\|\_\{1\}\(2\)The reconstructed embedding is then given by𝐱^=σ​\(𝐂𝐰∗\+μloc\)\\hat\{\\mathbf\{x\}\}=\\sigma\(\\mathbf\{C\}\\mathbf\{w\}^\{\*\}\+\\mu\_\{\\text\{loc\}\}\)\.

#### 7\.4\.3CLIP Surgery for Feature Attribution

CLIP Surgery changes the attention schema to what the authors callconsistent self\-attention, which replacesQQandKKwithVV:

Attn​\(V,V,V\)=softmax​\(V​V⊤d\)​V\\text\{Attn\}\(V,V,V\)=\\text\{softmax\}\\\!\\left\(\\frac\{VV^\{\\top\}\}\{\\sqrt\{d\}\}\\right\)VOnly features from this partial consistent self\-attention are used, without the feed forward network \(FFN\), arguing that FFNs push features toward negatives when identifying positives, thus leading to opposite saliency maps\. They implement these changes only to the last 6 layers of the vision model leaving other layers unchanged\.

After these changes to the vision model inference, given an input imageℐ∈ℝH×W×C\\mathcal\{I\}\\in\\mathbb\{R\}^\{H\\times W\\times C\}, the normalized output of the image encoderℱi∈ℝNt×Nd\\mathcal\{F\}\_\{i\}\\in\\mathbb\{R\}^\{N\_\{t\}\\times N\_\{d\}\}, whereNtN\_\{t\}corresponds to the number of output tokens \(excluding the class token\) andNdN\_\{d\}is the dimensionality of the aligned CLIP space\.

Giventhe normalized location encoder outputℱl∈ℝNd\\mathcal\{F\}\_\{l\}\\in\\mathbb\{R\}^\{N\_\{d\}\},a noise reduction step is applied in which a mean noise estimateℱn\\mathcal\{F\}\_\{n\}is subtracted from the location encoder output

ℱl′=ℱl−ℱn,ℱl′∈ℝNd\\mathcal\{F\}\_\{l\}^\{\\prime\}=\\mathcal\{F\}\_\{l\}\-\\mathcal\{F\}\_\{n\},\\qquad\\mathcal\{F\}\_\{l\}^\{\\prime\}\\in\\mathbb\{R\}^\{N\_\{d\}\}
ℱn\\mathcal\{F\}\_\{n\}can be obtained in one of two ways: \(1\) compute the mean over different classes or \(2\) proxied by the embedding of a “null” class \(which in the case of text is an empty string\)\.

After this,per\-patch similarity scores𝒮n\\mathcal\{S\}\_\{n\}are obtained between the image and location representations:

𝒮n=ℱi⋅ℱl′,𝒮n∈ℝNt\\mathcal\{S\}\_\{n\}=\\mathcal\{F\}\_\{i\}\\cdot\\mathcal\{F\}\_\{l\}^\{\\prime\},\\qquad\\mathcal\{S\}\_\{n\}\\in\\mathbb\{R\}^\{N\_\{t\}\}
Finally, function𝒰\\mathcal\{U\}is used to \(1\) reshape𝒮n\\mathcal\{S\}\_\{n\}to a square∈ℝNt×Nt\\in\\mathbb\{R\}^\{\\sqrt\{N\_\{t\}\}\\times\\sqrt\{N\_\{t\}\}\}; \(2\) expand to original image shapeℝH×W\\mathbb\{R\}^\{H\\times W\}through bilinear interpolation; and \(3\) min–max normalize values\.

𝒮f=𝒰​\(𝒮n\)\\mathcal\{S\}\_\{f\}=\\mathcal\{U\}\(\\mathcal\{S\}\_\{n\}\)where the final saliency map𝒮f∈ℝH×W\\mathcal\{S\}\_\{f\}\\in\\mathbb\{R\}^\{H\\times W\}can beoverlaid on theoriginal the image\.

Inspired byFagetet al\.\[[16](https://arxiv.org/html/2606.24997#bib.bib6)\], we compute a per\-layer saliency map𝒮l\\mathcal\{S\}\_\{l\}for each layerl∈ℒ′l\\in\\mathcal\{L\}^\{\\prime\}, whereℒ′\\mathcal\{L\}^\{\\prime\}denotes the same 6 layers to which CLIP Surgery is applied\. The per\-layer maps are then summed and min\-max normalized to obtain the final saliency map𝒮~f\\tilde\{\\mathcal\{S\}\}\_\{f\}:

𝒮sum=∑l∈ℒ′𝒮l\\mathcal\{S\}\_\{\\text\{sum\}\}=\\sum\_\{l\\in\\mathcal\{L\}^\{\\prime\}\}\\mathcal\{S\}\_\{l\}𝒮~f=𝒮sum−min⁡\(𝒮sum\)max⁡\(𝒮sum\)−min⁡\(𝒮sum\)\\tilde\{\\mathcal\{S\}\}\_\{f\}=\\frac\{\\mathcal\{S\}\_\{\\text\{sum\}\}\-\\min\(\\mathcal\{S\}\_\{\\text\{sum\}\}\)\}\{\\max\(\\mathcal\{S\}\_\{\\text\{sum\}\}\)\-\\min\(\\mathcal\{S\}\_\{\\text\{sum\}\}\)\}

### 7\.5Additional Results

#### 7\.5\.1Sparse Autoencoder Explanations

We provide additional visualizations for the SAE method, showing that this method can be used to identify visual artifacts \([Figure˜8](https://arxiv.org/html/2606.24997#S7.F8)\)\.

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/sae_figures/sae_visual_artifacts.png)Figure 8:Visual artifacts from the S2\-100k datasetaround the Arctic region identified by analyzing the images for the strongest activating samples for neurons ranked among the top\-10 visual monosemanticity values\.
#### 7\.5\.2SpLiCE Natural Language Explanations

ParisSaharaSiberiaBaliClimplicit![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/climplicit_Paris_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/climplicit_Sahara_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/climplicit_Siberia_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/climplicit_Bali_region_top_4.png)CSP\-fMoW![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/csp_fmow_Paris_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/csp_fmow_Sahara_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/csp_fmow_Siberia_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/csp_fmow_Bali_region_top_4.png)Figure 9:SpLiCE decompositions for Climplicit and CSP\-fMoW\.Top\-4 concepts in SpLiCE decompositions for Climplicit \(top\) and CSP\-fMoW \(bottom\) across four regions \(chosen to represent different land cover types\)\. All decompositions use a shared concept basis derived from Git\-10M\.We first provide SpLiCE decomposition results for additional location encoders Climplicit and CSP\-fMoW in[Figure˜9](https://arxiv.org/html/2606.24997#S7.F9)\. These SpLiCE results show that some of the concepts for these location embeddings seem less clearly aligned with underlying geographic semantics\.

Table 8:Sparse concept decompositions preserve structure in location embedding space even on a small geospatial concept set\.Here, we use a subset of about 200 geospatial concepts from the Git\-10M dataset\. To compute the reconstruction metrics, we use the same methods as in[Table˜2](https://arxiv.org/html/2606.24997#S4.T2)\. We useλ=0\.175\\lambda=0\.175for GeoCLIP and SatCLIP andλ=0\.125\\lambda=0\.125for Climplicit and CSP\-fMoW\.UARHuman\-visitedModelMSE↓\\downarrowCos\. Sim\.↑\\uparrow\# ConceptsMSE↓\\downarrowCos\. Sim\.↑\\uparrow\# ConceptsGeoCLIP0\.002±\\pm0\.0000\.451±\\pm0\.06111\.4±\\pm3\.10\.002±\\pm0\.0000\.384±\\pm0\.05610\.4±\\pm3\.0SatCLIP0\.005±\\pm0\.0010\.322±\\pm0\.0749\.8±\\pm3\.20\.004±\\pm0\.0010\.433±\\pm0\.08810\.0±\\pm3\.1Climplicit0\.001±\\pm0\.0000\.282±\\pm0\.08110\.2±\\pm3\.30\.001±\\pm0\.0000\.341±\\pm0\.08311\.8±\\pm4\.0CSP\-fMoW0\.005±\\pm0\.0010\.407±\\pm0\.18313\.9±\\pm4\.10\.006±\\pm0\.0010\.290±\\pm0\.11613\.4±\\pm5\.8

ParisSaharaSiberiaBaliGeoCLIP![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/geoclip_Paris_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/geoclip_Sahara_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/geoclip_Siberia_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/geoclip_Bali_small_subset_concepts_region_top_4.png)SatCLIP![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/satclip_Paris_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/satclip_Sahara_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/satclip_Siberia_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/satclip_Bali_small_subset_concepts_region_top_4.png)Climplicit![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/climplicit_Paris_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/climplicit_Sahara_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/climplicit_Siberia_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/climplicit_Bali_small_subset_concepts_region_top_4.png)CSP\-fMoW![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/csp_fmow_Paris_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/csp_fmow_Sahara_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/csp_fmow_Siberia_small_subset_concepts_region_top_4.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/splice_figures/decompositions/over_region/csp_fmow_Bali_small_subset_concepts_region_top_4.png)Figure 10:SpLiCE results on a smaller geospatial concept set, containing roughly 200 concepts,on GeoCLIP, SatCLIP, Climplicit, and CSP\-fMoW Using a smaller geospatial concept set biases the natural language decompositions to contain solely the geographic information, with tradeoffs for how much fine\-grained information can be extracted\.Next, we provide additional SpLiCE results on a small geospatial concept set consisting of 200 geospatial concepts derived from Git\-10M concepts\.[Table˜8](https://arxiv.org/html/2606.24997#S7.T8)shows that even with a small geospatial concept set, MSE and average cosine similarity show that sparse embeddings can still be used to reconstruct the original embeddings\. A smaller geospatial concept, however, introduces more inductive bias by further limiting the potential concepts that can be revealed by SpLiCE\. However, we do see that the SpLiCE decompositions in[Figure˜10](https://arxiv.org/html/2606.24997#S7.F10)align better with geographic attributes\. Interestingly, between location encoders, different geographic text concepts are expressed for the same location\.For example, the GeoCLIP location embeddings in the Sahara decompose with concepts such as “desert” and “wadi”, while SatCLIP decomposes with “orchard” and “butte”\.

#### 7\.5\.3CLIP Surgery Saliency Maps

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/SatCLIP-saliency-NP/africa.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/SatCLIP-saliency-NP/africa-saliency.png)

\(a\)Africa
![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/SatCLIP-saliency-NP/atacama.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/SatCLIP-saliency-NP/atacama-saliency.png)

\(b\)Atacama
![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/SatCLIP-saliency-NP/congo.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/SatCLIP-saliency-NP/congo-saliency.png)

\(c\)Congo
![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/SatCLIP-saliency-NP/amazon.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/SatCLIP-saliency-NP/amazon-saliency.png)

\(d\)Amazon

Figure 11:CLIP Surgery saliency maps generated from Sentinel\-2 \(S2\) imagery using the SatCLIP location and image encoders\. The first row shows desert\-like land cover from geographically distinct yet visually similar regions — the Atacama and African Deserts — while the second row shows tropical rainforest\-like land cover from the Congo Basin and Amazon\. Though attention to water edges and rivers emerges in rainforest areas, deserts are much more difficult to interpret\. This suggests that saliency maps are less reliable for satellite imagery over spatially homogeneous land cover\.![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/saliency/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Paris/saliency/2.png)

\(a\)Paris
![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/saliency/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/Amsterdam/saliency/2.png)

\(b\)Amsterdam
![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/saliency/1.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/StLouis/saliency/2.png)

\(c\)St\. Louis
![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/saliency/2.png)

![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes-NP/ImageNet/NYC/saliency/3.png)

\(d\)New York City

Figure 12:Satellite image saliency maps for four cities: Paris, Amsterdam, St\. Louis and New York City\. While the model’s attention tends to highlight structural boundaries such as roads and waterways, the maps are difficult to interpret directly due to their scale and resolution\. We therefore introduce an additional step of cropping and clustering to extract more interpretable visual patterns\.![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/NoSurgery/NoSurgery/florence_1.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/NoSurgery/NoSurgery/paris.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/NoSurgery/NoSurgery/seoul.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/NoSurgery/NoSurgery/peru.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/NoSurgery/Surgery/florence_1.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/NoSurgery/Surgery/paris.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/NoSurgery/Surgery/seoul.png)![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/GeoClip-NP/NoSurgery/Surgery/peru.png)
\(a\)Florence\(b\)Paris\(c\)Seoul\(d\)Peru
Figure 13:Natural image saliency maps taken from Im2GPS dataset from four locations: Florence, Paris, Seoul, Peru, shown without \(top\) and with \(bottom\) CLIP Surgery applied\.Confusing highlights on background elements, a known issue in CLIP\-based models saliency maps, are noticeably reduced when CLIP Surgery is applied\.![Refer to caption](https://arxiv.org/html/2606.24997v1/sec/figures/bboxes_pipeline/cluster.png)Figure 14:Spatial clustering of salient region crops within the Paris area\. Crop embeddings were computed using a ResNet\-50 pretrained on ImageNet weights, reduced to 2D via t\-SNE, and clustered using DBSCAN\. Colors denote distinct clusters\.

Similar Articles

Location-Aware Language Models via Secondary Embeddings

arXiv cs.CL

The paper proposes a lightweight, model-agnostic method to enhance language models with geo-spatial awareness by augmenting embeddings with location data, improving spatial alignment while maintaining standard NLP performance.