LE4Mob:用于人类移动性建模的归纳式、距离感知通用位置嵌入
摘要
LE4Mob 是一个归纳式、距离感知的通用位置嵌入框架,用于增强人类移动性建模中的任务,如下一位置预测和通勤流量生成。
arXiv:2609.22117v1 Announce Type: new
Abstract: Location representations provide mobility models with fundamental information about the spatial position, functional characteristics, and relationships of places. However, existing embeddings are often dependent on mobility observations, unable to represent unseen locations, and weakly constrained to retain geographic distance. This limits their reuse across datasets and mobility tasks. To address these limitations, we propose LE4Mob, an inductive, distance-aware, and geography-derived location embedding framework for mobility modelling. LE4Mob extends contrastive language-location pre-training while introducing a distance-aware regularisation objective that encourages the embedding space to preserve spatial relationships. Pre-trained from geographic context, LE4Mob can encode rich spatial-semantic information and generate embeddings for unseen locations inductively. Its independence from downstream mobility task supervision also makes it transferable across different mobility tasks. We evaluate LE4Mob on individual-level next location prediction and population-level commuter flow generation. Experiments across multiple datasets and study areas show that LE4Mob outperforms strong baselines, with particular advantages in inductive settings and when downstream models rely directly on interactions between location embeddings. These findings demonstrate the potential of distance-aware, geography-derived location representations as reusable foundations for human mobility modelling.
查看缓存全文
缓存时间: 2026/09/22 09:11
# LE4Mob: Towards Inductive, Distance-Aware and General-Purpose Location Embedding for Human Mobility Modelling
Source: [https://arxiv.org/html/2609.22117](https://arxiv.org/html/2609.22117)
Stephen Law[https://orcid.org/0000-0003-3184-572X](https://orcid.org/0000-0003-3184-572X)Zichao Zeng[https://orcid.org/0009-0002-8975-875X](https://orcid.org/0009-0002-8975-875X)Junyuan Liu[https://orcid.org/0009-0009-7194-6868](https://orcid.org/0009-0009-7194-6868)Guangsheng Dong[https://orcid.org/0000-0001-7676-497X](https://orcid.org/0000-0001-7676-497X)Tao Cheng[https://orcid.org/0000-0002-5503-9813](https://orcid.org/0000-0002-5503-9813)Thanks:This work was supported by the project “Understanding the Impact of Covid\-19 & NPIs on Mobility and Places in Singapore,” funded by The Alan Turing Institute and DSO National Laboratories, Singapore\.Thanks:Xinglei Wang, Zichao Zeng, Junyuan Liu, and Tao Cheng are with SpaceTimeLab, Department of Civil, Environmental and Geomatic Engineering, University College London, London, United Kingdom\. Corresponding author: Tao Cheng \(e\-mail: tao\.cheng@ucl\.ac\.uk\)\.Thanks:Stephen Law is with the Department of Geography, University College London, London, United Kingdom, and the Department of Geography and Resource Management, The Chinese University of Hong Kong, Hong Kong, China\.Thanks:Zichao Zeng is also affiliated with 3DIMPACT, Department of Civil, Environmental and Geomatic Engineering, University College London, London, United Kingdom\.Thanks:Guangsheng Dong is with the State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing \(LIESMARS\), Wuhan University, Wuhan, China\. He is also a visiting scholar with SpaceTimeLab, Department of Civil, Environmental and Geomatic Engineering, University College London, London, United Kingdom\.Thanks:Tao Cheng is also with The Alan Turing Institute, London, United Kingdom\.
###### Abstract
Location representations provide mobility models with fundamental information about the spatial position, functional characteristics, and relationships of places\. However, existing embeddings are often dependent on mobility observations, unable to represent unseen locations, and weakly constrained to retain geographic distance\. This limits their reuse across datasets and mobility tasks\. To address these limitations, we propose LE4Mob, an inductive, distance\-aware, and geography\-derived location embedding framework for mobility modelling\. LE4Mob extends contrastive language–location pre\-training while introducing a distance\-aware regularisation objective that encourages the embedding space to preserve spatial relationships\. Pre\-trained from geographic context, LE4Mob can encode rich spatial\-semantic information and generate embeddings for unseen locations inductively\. Its independence from downstream mobility task supervision also makes it transferable across different mobility tasks\. We evaluate LE4Mob on individual\-level next location prediction and population\-level commuter flow generation\. Experiments across multiple datasets and study areas show that LE4Mob outperforms strong baselines, with particular advantages in inductive settings and when downstream models rely directly on interactions between location embeddings\. These findings demonstrate the potential of distance\-aware, geography\-derived location representations as reusable foundations for human mobility modelling\.
###### Index Terms:
Representation learning, human mobility, location encoding, next location prediction, commuter flow generation\.
## IIntroduction
Human mobility, the movement of individuals across space and time, shapes urban dynamics, transport demand, disease transmission, economic activity, and public policy\[[1](https://arxiv.org/html/2609.22117#bib.bib18)\]\. The increasing availability of large\-scale mobility data, together with advances in deep learning, has led to rapid progress in computational mobility modelling and established data\-driven methods as an important direction for mobility science\[[24](https://arxiv.org/html/2609.22117#bib.bib19)\]\. Within this broad field, four major tasks have received particular attention: next location prediction, trajectory generation, crowd flow prediction, and flow generation\[[17](https://arxiv.org/html/2609.22117#bib.bib20)\]\. These tasks differ in their prediction targets, spatial scales, and methodological formulations, but they share a foundational requirement: models must represent the locations between, through, or towards which movement occurs\.
Locations are commonly encoded as dense vectors that provide compact inputs to downstream models\. Effective representations should describe not only location identity, but also geographic position, surrounding urban functions, and spatial relationships with other places\. These properties are central to mobility behaviour\. Individuals choose destinations partly according to the activities and opportunities available there\[[30](https://arxiv.org/html/2609.22117#bib.bib32),[26](https://arxiv.org/html/2609.22117#bib.bib33)\], while travel likelihood and aggregate interaction volumes are strongly constrained by geographic distance\[[43](https://arxiv.org/html/2609.22117#bib.bib29)\]\. Location representation can therefore be considered as a common foundation for both individual\- and population\-level mobility modelling\.
Nevertheless, location representations have largely been developed separately for different task families\. Individual\-level models commonly learn embeddings from historical check\-ins or trajectories, using co\-occurrence or transition patterns to infer relationships among destinations\[[6](https://arxiv.org/html/2609.22117#bib.bib13),[40](https://arxiv.org/html/2609.22117#bib.bib25),[42](https://arxiv.org/html/2609.22117#bib.bib23),[32](https://arxiv.org/html/2609.22117#bib.bib8),[13](https://arxiv.org/html/2609.22117#bib.bib7)\]\. Population\-level models often represent regions using handcrafted attributes, such as population, land use, and points of interest \(POIs\)\[[28](https://arxiv.org/html/2609.22117#bib.bib11)\], or learn latent representations from observed origin–destination \(OD\) networks\[[16](https://arxiv.org/html/2609.22117#bib.bib34),[38](https://arxiv.org/html/2609.22117#bib.bib35)\]\. These approaches can capture task\-relevant patterns, but the resulting embeddings are usually tied to the mobility data, spatial units, and downstream objective used during training\.
This task\-specific paradigm creates three limitations\. First, mobility\-derived embeddings may inadequately capture the functional semantics of the urban environment\. Human movement is influenced by the spatial distribution and accessibility of activity opportunities \(e\.g\., employment, education, recreation, transport services\) associated with different land uses and urban functions\[[3](https://arxiv.org/html/2609.22117#bib.bib28),[12](https://arxiv.org/html/2609.22117#bib.bib21)\]\. Although semantic attributes can be introduced as auxiliary features, they are not necessarily encoded in a reusable representation when locations are learned mainly from their occurrence in mobility sequences or OD graphs\.
Second, many existing embeddings are transductive\. They assign a trainable vector to each location observed during training and cannot directly encode unseen destinations or regions\. This constraint is problematic when new locations, spatial units, or study areas appear after deployment\. Extending the location vocabulary then requires retraining or learning additional embeddings, limiting the transfer across datasets and regions\.
Third, existing location representations rarely preserve geographic distance explicitly\. Distance is a fundamental constraint on both destination choice and aggregate flows, as reflected in distance\-decay effects and classical gravity\-\[[43](https://arxiv.org/html/2609.22117#bib.bib29)\]and radiation\-based formulations\[[29](https://arxiv.org/html/2609.22117#bib.bib30)\]\. However, embeddings optimised for mobility co\-occurrence, semantic similarity, or graph connectivity do not necessarily retain geographic proximity\. Functionally similar but distant locations may be placed close together, whereas nearby locations with different functions may become separated\. Such distortion can make it harder for downstream models to capture interaction patterns\. Some recent work has recognised the importance of geographic distance and introduced soft distance supervision\[[8](https://arxiv.org/html/2609.22117#bib.bib43)\], but it focuses on street\-view image representations for geolocalisation rather than location embeddings for mobility modelling\.
These limitations motivate distance\-aware geography\-driven location representation learning\. Instead of deriving embeddings from a particular mobility dataset, a location encoder can learn from information available independently of downstream mobility observations, such as coordinates, POIs, text, imagery\. Such an encoder should be distance\-explicit, inductive, and generalisable across mobility tasks\.
Our previous work provided initial evidence for this direction by applying CaLLiPer\[[34](https://arxiv.org/html/2609.22117#bib.bib22)\], a POI\-based spatial\-semantic location encoder, to next location prediction\[[33](https://arxiv.org/html/2609.22117#bib.bib31)\]\. CaLLiPer uses contrastive language–location pre\-training to align coordinate representations with textual descriptions of surrounding POIs\. Its embeddings improved next location prediction, particularly when individuals visited locations unseen during downstream training\. However, this study considered only individual\-level task, and CaLLiPer did not explicitly preserve geographic distance in its embedding space\.
To address these limitations, we propose LE4Mob, an inductive, semantics\-rich, and distance\-awareLocationEmbedding frameworkforhumanMobility modelling\. LE4Mob extends CaLLiPer with a distance\-preserving objective that aligns relationships in the embedding space with geographic distances between locations\. Its pre\-training relies on general geographic information rather than trajectory or OD\-flow labels, allowing the same encoder to be integrated into different mobility datasets and model architectures\.
We evaluate LE4Mob on two different mobility modelling tasks: individual\-level next location prediction and population\-level commuter flow generation\. These tasks differ in spatial units, prediction targets, and modelling assumptions, providing a stringent test of the model’s generalisability across heterogeneous mobility settings\. Across conventional and inductive next location prediction, LE4Mob achieves the strongest overall performance, with particularly clear advantages in the inductive setting and on geographic distance\-based errors\. In commuter flow generation in London, LE4Mob achieves 2\.2–11\.3% performance improvements relative to the strongest baselines\. Ablation and qualitative analyses further demonstrate that the distance\-preserving design improves the spatial awareness and downstream utility of the learned location embeddings\.
The contributions are threefold\. Conceptually, we frame location representation learning as a transferable problem spanning individual\- and population\-level mobility modelling\. Methodologically, we propose LE4Mob, a framework that addresses task dependence, transductivity, and spatial distortion by combining semantic context, inductive encoding, and distance\-aware regularisation\. Empirically, we evaluate LE4Mob on next location prediction and commuter flow generation, demonstrating its effectiveness across distinct mobility scales, and release the code and datasets to support reproducibility\.
## IIRelated work
In this study,location representationis an umbrella term for any numerical description of a geographic location, including raw attributes, handcrafted features, and learned latent representations\. Alocation embeddingrefers specifically to a dense latent vector, either assigned to a fixed location identifier through a lookup table or generated from observable geographic attributes by a location encoder\.Location representation learningdenotes the broader process of learning these embeddings or the encoder that produces them\.
The geographic entities regarded as locations vary by task\. Individual\-level models commonly represent locations as discretised spatial cells, POIs, clustered stay points or significant places\[[9](https://arxiv.org/html/2609.22117#bib.bib1)\]\. Population\-level models generally use areal units such as grids, traffic analysis zones, census tracts, output areas\[[28](https://arxiv.org/html/2609.22117#bib.bib11)\], etc\. Despite the differences, the locations in both models can be characterised by their geographic position, functional or semantic attributes, and spatial relationships with other locations\.
Accordingly, a general\-purpose location representation should satisfy four properties\. It should capture semantic and functional characteristics because travel decisions depend on the activities and opportunities available at a location\. It should preserve spatial relationships that shape destination choice and aggregate interactions\. It should also generalise inductively to locations unseen during representation learning and remain transferable across downstream objectives\. The following subsections examine how existing approaches address these requirements in location encoding, individual mobility prediction, and population flow modelling\.
### II\-ALocation Encoders
Location encoders are parameterised functions that map observable geographic information to dense representations\. Coordinate\-based encoders typically transform longitude and latitude through positional encoding and a neural network\[[18](https://arxiv.org/html/2609.22117#bib.bib5)\]\. Because they learn a mapping rather than a location\-specific lookup table, they can encode arbitrary coordinates and are inherently inductive\.
Some encoders are trained through supervised classification\[[19](https://arxiv.org/html/2609.22117#bib.bib6),[20](https://arxiv.org/html/2609.22117#bib.bib9)\], whereas multimodal methods use contrastive learning to align coordinates with images, text, POIs, or other environmental observations\[[11](https://arxiv.org/html/2609.22117#bib.bib26),[34](https://arxiv.org/html/2609.22117#bib.bib22),[14](https://arxiv.org/html/2609.22117#bib.bib27)\]\. The latter capture both geographic position and urban function and may therefore transfer beyond their original training tasks\.
CaLLiPer\[[34](https://arxiv.org/html/2609.22117#bib.bib22)\]is a multimodal geography\-derived encoder that aligns coordinate representations with textual descriptions of surrounding POIs\. Its objective is independent of trajectories and OD flows, enabling application across mobility datasets\. However, incorporating coordinates does not guarantee that the final embedding geometry preserves geographic distance\. Nonlinear transformations optimised for classification or contrastive alignment may place semantically similar but distant locations close together or separate nearby places with different functions\. Few geographic encoders explicitly supervise the correspondence between geographic and embedding space distances, although this property is particularly important for mobility modelling\.
### II\-BLocation Representations for Individual Mobility Prediction
Next location prediction estimates a person’s next destination from previous movements\. Deep models commonly represent candidate destinations using trainable lookup embeddings and combine them with temporal, user, and contextual features\[[5](https://arxiv.org/html/2609.22117#bib.bib2),[9](https://arxiv.org/html/2609.22117#bib.bib1)\]\. These embeddings are efficient but depend entirely on task\-specific observations, do not inherently encode coordinates or urban functions, and cannot represent locations outside the training vocabulary\.
Self\-supervised approaches derive richer embeddings from movement sequences\. Word2Vec\-style methods treat locations as words and trajectories as sentences\[[21](https://arxiv.org/html/2609.22117#bib.bib15)\], while extensions incorporate spatial\[[6](https://arxiv.org/html/2609.22117#bib.bib13)\], temporal\[[32](https://arxiv.org/html/2609.22117#bib.bib8)\], and spatial\-temporal information\[[40](https://arxiv.org/html/2609.22117#bib.bib25)\]\. BERT\-inspired models further produce context\-dependent representations from adjacent visits\[[13](https://arxiv.org/html/2609.22117#bib.bib7)\]\. These approaches capture behavioural regularities, but sparse or unseen locations receive unreliable or no representations\. Moreover, spatial co\-occurrence does not necessarily reflect functional semantics or geographic proximity of locations\.
A previous work instead encodes coordinates and environmental context independently of trajectories, supporting inductive prediction for unseen destinations\[[33](https://arxiv.org/html/2609.22117#bib.bib31)\]\. However, it did not explicitly preserve geographic distance or examine population\-level transfer\.
### II\-CLocation Representations for Population Flow Modelling
Population flow modelling estimates aggregate movements between regions, including commuting, migration, and travel demand\[[1](https://arxiv.org/html/2609.22117#bib.bib18)\]\. Flow prediction forecasts future flows from historical observations, whereas flow generation estimates an origin–destination \(OD\) matrix from the attributes and relationships of origins and destinations\[[28](https://arxiv.org/html/2609.22117#bib.bib11)\]\. This study focuses on the latter\.
Classical gravity models relate flows to the masses of origin and destination regions and their spatial separation\[[43](https://arxiv.org/html/2609.22117#bib.bib29)\], while radiation models emphasise the role of intervening opportunities\[[30](https://arxiv.org/html/2609.22117#bib.bib32),[29](https://arxiv.org/html/2609.22117#bib.bib30)\]\. DeepGravity replaces the rigid functions with neural networks to capture nonlinear interactions\[[28](https://arxiv.org/html/2609.22117#bib.bib11)\]\. Recent models further embed spatial interaction principles into deep architectures\[[39](https://arxiv.org/html/2609.22117#bib.bib42)\]\. These methods primarily concern downstream flow modelling rather than the pre\-training of transferable representations of the regions\.
A growing body of work instead learns latent location representations jointly with a downstream flow decoder, such as a multilayer perceptron\[[31](https://arxiv.org/html/2609.22117#bib.bib37)\]or bilinear interaction model\[[16](https://arxiv.org/html/2609.22117#bib.bib34),[36](https://arxiv.org/html/2609.22117#bib.bib38)\]\. Many such approaches use graph neural networks to integrate regional attributes with geographic, transport, or other relational structures, while optimising the resulting embeddings using observed OD flows as supervision\[[16](https://arxiv.org/html/2609.22117#bib.bib34),[38](https://arxiv.org/html/2609.22117#bib.bib35),[27](https://arxiv.org/html/2609.22117#bib.bib36),[31](https://arxiv.org/html/2609.22117#bib.bib37)\]\. However, they are commonly evaluated under a transductive protocol in which all regions are present in a fixed graph and the split is performed over OD edges or flow records\. Test region embeddings are therefore generated within the same graph used during training, and the evaluation measures prediction of unobserved flows between known regions rather than generalisation to previously unseen spatial units\. Consequently, these studies provide limited evidence that the learned representations can be applied to regions absent during representation learning\.
Geography\-derived self\-supervised methods provide an alternative by learning region representations from information available independently of OD flows, such as coordinates, POIs, land use, text, or satellite imagery\. For example, recent work pre\-trains region representations from satellite imagery and evaluates them using held\-out origin regions\[[36](https://arxiv.org/html/2609.22117#bib.bib38)\]\. Because such representations can be generated without observing the target region’s flow records, they are better suited to encoding unseen regions and to settings where historical mobility observations are unavailable or incomplete\.
Nevertheless, geographic distance is often introduced only at the downstream stage as an explicit OD\-pair feature\. Although effective, supplying distance to a decoder does not ensure that the region embeddings themselves retain spatial relationships\. Representation\-level distance supervision instead incorporates spatial structure during pre\-training, producing embeddings that remain spatially informative when used with decoders that do not receive distance separately and potentially improving transfer across downstream architectures\.
## IIINotation and Problem Formulation
This section defines the common notation, formulates the two downstream mobility tasks, and specifies the location representation learning problem addressed by LE4Mob\.
### III\-ALocations
Definition 1 \(Location\)\.Letℒ=\{li\}i=1N\\mathcal\{L\}=\\\{l\_\{i\}\\\}\_\{i=1\}^\{N\}denote a set ofNNlocations\. Each locationlil\_\{i\}is characterised byli=\(gi,ci\)l\_\{i\}=\(g\_\{i\},c\_\{i\}\), wheregig\_\{i\}denotes its geometry andcic\_\{i\}denotes its contextual information, such as surrounding points of interest \(POIs\) or land use characteristics\. We further use𝐱i\\mathbf\{x\}\_\{i\}to denote a representative coordinate derived fromgig\_\{i\}, such as the coordinate of a point location or the centroid of a polygon\.
The spatial form of a location depends on the mobility task\. In next location prediction, a location may correspond to a POI, venue, clustered stay region, or discretised spatial cell\. In flow generation, locations are areal units forming a tessellation of the study area, such as regular grids, census zones, or administrative regions\.
The geographic distance between locationslil\_\{i\}andljl\_\{j\}is denoted bydij=dgeo\(𝐱i,𝐱j\)d\_\{ij\}=d\_\{\\mathrm\{geo\}\}\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\), wheredgeo\(⋅,⋅\)d\_\{\\mathrm\{geo\}\}\(\\cdot,\\cdot\)is the selected geographic distance function\.
### III\-BNext Location Prediction
Definition 2 \(Individual Mobility Trajectory\)\.A user’s mobility history is represented as a spatio\-temporal trajectorySuS\_\{u\}\. It consists of a time\-ordered sequence of visits:Su1:Tu=\[\(luk,tuk\)\]k=1TuS\_\{u\}^\{1:T\_\{u\}\}=\{\[\(l\_\{u\}^\{k\},t\_\{u\}^\{k\}\)\]\}\_\{k=1\}^\{T\_\{u\}\}where\(luk,tuk\)\(l\_\{u\}^\{k\},t\_\{u\}^\{k\}\)means that useruuvisited locationluk∈ℒl\_\{u\}^\{k\}\\in\\mathcal\{L\}at time steptukt\_\{u\}^\{k\}\.
Problem 1 \(Next Location Prediction\)\.Given a user’s previous visits within an observation windowm:nm:n,Sum:n=\[\(lum,tum\),…,\(lun,tun\)\]S\_\{u\}^\{m:n\}=\\left\[\(l\_\{u\}^\{m\},t\_\{u\}^\{m\}\),\\ldots,\(l\_\{u\}^\{n\},t\_\{u\}^\{n\}\)\\right\], the goal is to predict the next locationlun\+1l\_\{u\}^\{n\+1\}that the user will visit\. The task is formulated as a multi\-class classification problem\. The model assigns a probability to each candidate location in the candidate location setℒc\\mathcal\{L\}^\{c\}, and the location with the highest probability is selected as the predicted next location:
l^un\+1=argmaxl∈ℒcP\(l∣Sum:n\)\.\\hat\{l\}\_\{u\}^\{n\+1\}=\\arg\\max\_\{l\\in\\mathcal\{L\}^\{c\}\}\\mathrm\{P\}\\left\(l\\mid S\_\{u\}^\{m:n\}\\right\)\.\(1\)In the inductive setting, the candidate location setℒc\\mathcal\{L\}^\{c\}may include locations that were not observed during model training\.
### III\-CFlow Generation
Problem 2 \(Flow Generation\)\.Letℒ\\mathcal\{L\}denote a set of spatial regions, and letyij≥0y\_\{ij\}\\geq 0denote the observed flow from originlil\_\{i\}to destinationljl\_\{j\}, wherei≠ji\\neq j\. The total outflow from originlil\_\{i\}isOi=∑j≠iyijO\_\{i\}=\\sum\_\{j\\neq i\}y\_\{ij\}\. GivenOiO\_\{i\}and information describing the origin and destination locations, the objective is to estimate the flowy^ij\\hat\{y\}\_\{ij\}for arbitrary location pairs\(li,lj\)\(l\_\{i\},l\_\{j\}\)\.
A flow generation model may estimate eithery^ij\\hat\{y\}\_\{ij\}directly or a destination probabilityp^ij\\hat\{p\}\_\{ij\}, from which the flow is obtained asy^ij=Oip^ij,∑j≠ip^ij=1\\hat\{y\}\_\{ij\}=O\_\{i\}\\hat\{p\}\_\{ij\},\\sum\_\{j\\neq i\}\\hat\{p\}\_\{ij\}=1\.
### III\-DLocation Representation Learning for Mobility Modelling
Problem 3 \(Location Representation Learning\)\.The objective is to learn a parameterised mapping functionFθF\_\{\\theta\}, which generates a dense embedding𝐞i=Fθ\(gi,ci\),𝐞i∈ℝd\\mathbf\{e\}\_\{i\}=F\_\{\\theta\}\(g\_\{i\},c\_\{i\}\),\\mathbf\{e\}\_\{i\}\\in\\mathbb\{R\}^\{d\}, from the geographic and contextual information of locationlil\_\{i\}\.
The mapping function is designed to satisfy three properties:
1. 1\.Inductive\.The encoderFθF\_\{\\theta\}can represent a location not observed during representation pre\-training without learning a new location\-specific parameter or retraining the encoder\.
2. 2\.Distance\-aware\.Geographic distance is explicitly incorporated into representation learning, such that relationships between embeddings retain information about the corresponding geographic distancesdijd\_\{ij\}\.
3. 3\.Downstream\-task\-agnostic\.The parametersθ\\thetaare learned independently of task\-specific trajectory labels or OD flow values\. The resulting embeddings can therefore be integrated into different downstream mobility models, including those defined in Problems 1 and 2\.
## IVMethodology
This section presents LE4Mob, a spatial\-semantic location encoder augmented with explicit distance\-aware regularisation\. We first show why the spatial structure introduced by positional encoding is not necessarily retained after neural transformation\. We then describe the semantic and distance\-aware training objectives and the integration of the pre\-trained embeddings into downstream mobility models\.
### IV\-ATheoretical Motivation
Standard location encoders commonly follow
𝐞\(𝐱\)=fθ\(PE\(𝐱\)\),\\mathbf\{e\}\(\\mathbf\{x\}\)=f\_\{\\theta\}\\bigl\(\\mathrm\{PE\}\(\\mathbf\{x\}\)\\bigr\),\(2\)where𝐱\\mathbf\{x\}is a coordinate,PE\(⋅\)\\mathrm\{PE\}\(\\cdot\)is a positional encoding, andfθ\(⋅\)f\_\{\\theta\}\(\\cdot\)is a trainable neural network\[[18](https://arxiv.org/html/2609.22117#bib.bib5)\]\. Although positional encodings introduce spatial structure, an unconstrained neural transformation is not guaranteed to retain it\.
Definition 3 \(Distance Preservation\)\.Let\(𝒳,dgeo\)\(\\mathcal\{X\},d\_\{\\mathrm\{geo\}\}\)be a geographic space and let𝐞:𝒳→ℝd\\mathbf\{e\}:\\mathcal\{X\}\\rightarrow\\mathbb\{R\}^\{d\}be a non\-zero embedding function\. We call𝐞\\mathbf\{e\}distance\-preserving if, for any anchor𝐱i\\mathbf\{x\}\_\{i\}and locations𝐱j\\mathbf\{x\}\_\{j\}and𝐱k\\mathbf\{x\}\_\{k\},
dgeo\(𝐱i,𝐱j\)<dgeo\(𝐱i,𝐱k\)⟹⟨𝐞i,𝐞j⟩\>⟨𝐞i,𝐞k⟩,d\_\{\\mathrm\{geo\}\}\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\)<d\_\{\\mathrm\{geo\}\}\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{k\}\)\\Longrightarrow\\langle\\mathbf\{e\}\_\{i\},\\mathbf\{e\}\_\{j\}\\rangle\>\\langle\\mathbf\{e\}\_\{i\},\\mathbf\{e\}\_\{k\}\\rangle,\(3\)where⟨⋅,⋅⟩\\langle\\cdot,\\cdot\\rangledenotes inner product or cosine similarity\[[18](https://arxiv.org/html/2609.22117#bib.bib5)\]\.
Proposition 1\.A positional encoding may be distance\-preserving, while its composition with an unconstrained linear transformation is not\.
Proof\.Consider the one\-dimensional positional encoding
PE\(x\)=\[cosxsinx\],x∈\[0,π\]\.\\mathrm\{PE\}\(x\)=\\begin\{bmatrix\}\\cos x\\\\ \\sin x\\end\{bmatrix\},\\qquad x\\in\[0,\\pi\]\.\(4\)Its cosine similarity satisfiesPE\(xi\)⊤PE\(xj\)=cos\(xi−xj\)\\mathrm\{PE\}\(x\_\{i\}\)^\{\\top\}\\mathrm\{PE\}\(x\_\{j\}\)=\\cos\(x\_\{i\}\-x\_\{j\}\), which decreases monotonically with\|xi−xj\|\|x\_\{i\}\-x\_\{j\}\|over this interval\.
Now apply
W=\[10001\]W=\\begin\{bmatrix\}10&0\\\\ 0&1\\end\{bmatrix\}\(5\)and define
𝐞~\(x\)=WPE\(x\)∥WPE\(x\)∥2\.\\tilde\{\\mathbf\{e\}\}\(x\)=\\frac\{W\\mathrm\{PE\}\(x\)\}\{\\lVert W\\mathrm\{PE\}\(x\)\\rVert\_\{2\}\}\.\(6\)Letxi=π3,xj=π2,xk=0x\_\{i\}=\\frac\{\\pi\}\{3\},x\_\{j\}=\\frac\{\\pi\}\{2\},x\_\{k\}=0\. Although\|xi−xj\|=π6<π3=\|xi−xk\|\|x\_\{i\}\-x\_\{j\}\|=\\frac\{\\pi\}\{6\}<\\frac\{\\pi\}\{3\}=\|x\_\{i\}\-x\_\{k\}\|, the transformed similarities aresim\(𝐞~\(xi\),𝐞~\(xj\)\)=3103≈0\.171\\operatorname\{sim\}\\left\(\\tilde\{\\mathbf\{e\}\}\(x\_\{i\}\),\\tilde\{\\mathbf\{e\}\}\(x\_\{j\}\)\\right\)=\\frac\{\\sqrt\{3\}\}\{\\sqrt\{103\}\}\\approx 0\.171,sim\(𝐞~\(xi\),𝐞~\(xk\)\)=10103≈0\.985\\operatorname\{sim\}\\left\(\\tilde\{\\mathbf\{e\}\}\(x\_\{i\}\),\\tilde\{\\mathbf\{e\}\}\(x\_\{k\}\)\\right\)=\\frac\{10\}\{\\sqrt\{103\}\}\\approx 0\.985\.
The more distant location therefore becomes more similar to the anchor than the closer location, violating distance consistency\. This failure mode is not unique to sinusoidal encodings\. Even a simple distance\-preserving PE like a constrained identity map is easily distorted by the anisotropy introduced by unconstrained linear layers\.
In a deeper encoder trained only through semantic alignment, geographically distant but functionally similar places may be pulled together, while nearby places with different functions may be separated\. To mitigate the spatial distortion, LE4Mob therefore introduces an explicit distance\-aware objective\.
### IV\-BLE4Mob Framework
Fig\. 1:Framework of LE4Mob\. \(a\) In the pre\-training phase, batches of POIs are encoded through a dual\-stream architecture that aligns coordinate\-based location embeddings with POI\-derived textual semantics using contrastive learning, while a distance\-preserving loss regularises sampled location pairs to retain geographic distance information\. \(b\) After pre\-training, the frozen location encoder generates embeddings for downstream mobility tasks, including individual\-level next location prediction and population\-level flow generation\.LE4Mob follows a “pre\-training→\\rightarrowdownstream application” paradigm, as illustrated in Fig\.[1](https://arxiv.org/html/2609.22117#S4.F1)\. It extends CaLLiPer\[[34](https://arxiv.org/html/2609.22117#bib.bib22)\]and contains three backbone components:
1. 1\.a location encoderfθf\_\{\\theta\}that maps a PE\-encoded coordinate𝐱i\\mathbf\{x\}\_\{i\}to a location embedding;
2. 2\.a frozen text encodergϕg\_\{\\phi\}that extracts semantic features from the contextual descriptioncic\_\{i\}of surrounding POIs; and
3. 3\.a trainable projection layerpψp\_\{\\psi\}that maps text features to the location embedding space\.
For each locationlil\_\{i\}, the two branches produce
𝐳i\\displaystyle\\mathbf\{z\}\_\{i\}=fθ\(PE\(𝐱i\)\)∥fθ\(PE\(𝐱i\)\)∥2,\\displaystyle=\\frac\{f\_\{\\theta\}\(\\mathrm\{PE\}\(\\mathbf\{x\}\_\{i\}\)\)\}\{\\lVert f\_\{\\theta\}\(\\mathrm\{PE\}\(\\mathbf\{x\}\_\{i\}\)\)\\rVert\_\{2\}\},\(7\)𝐪i\\displaystyle\\mathbf\{q\}\_\{i\}=pψ\(gϕ\(ci\)\)∥pψ\(gϕ\(ci\)\)∥2\.\\displaystyle=\\frac\{p\_\{\\psi\}\(g\_\{\\phi\}\(c\_\{i\}\)\)\}\{\\lVert p\_\{\\psi\}\(g\_\{\\phi\}\(c\_\{i\}\)\)\\rVert\_\{2\}\}\.\(8\)
Contextual descriptions are used only as semantic supervision during pre\-training\. Once training is complete,gϕg\_\{\\phi\}andpψp\_\{\\psi\}are discarded, andfθf\_\{\\theta\}independently generates embeddings from coordinates\.
### IV\-CSemantic Alignment
Given a mini\-batch ofBBpaired coordinates and contextual descriptions, we calculate the cross\-modal similarity
aij=𝐳i⊤𝐪jτ,a\_\{ij\}=\\frac\{\\mathbf\{z\}\_\{i\}^\{\\top\}\\mathbf\{q\}\_\{j\}\}\{\\tau\},\(9\)whereτ\\tauis a temperature parameter\.
Following CLIP\-style contrastive learning\[[25](https://arxiv.org/html/2609.22117#bib.bib16)\], the location\-to\-text and text\-to\-location losses are
ℒl2t\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{l2t\}\}=−1B∑i=1Blogexp\(aii\)∑j=1Bexp\(aij\),\\displaystyle=\-\\frac\{1\}\{B\}\\sum\_\{i=1\}^\{B\}\\log\\frac\{\\exp\(a\_\{ii\}\)\}\{\\sum\_\{j=1\}^\{B\}\\exp\(a\_\{ij\}\)\},\(10\)ℒt2l\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{t2l\}\}=−1B∑i=1Blogexp\(aii\)∑j=1Bexp\(aji\)\.\\displaystyle=\-\\frac\{1\}\{B\}\\sum\_\{i=1\}^\{B\}\\log\\frac\{\\exp\(a\_\{ii\}\)\}\{\\sum\_\{j=1\}^\{B\}\\exp\(a\_\{ji\}\)\}\.\(11\)The semantic alignment objective is
ℒsem=12\(ℒl2t\+ℒt2l\)\.\\mathcal\{L\}\_\{\\mathrm\{sem\}\}=\\frac\{1\}\{2\}\\left\(\\mathcal\{L\}\_\{\\mathrm\{l2t\}\}\+\\mathcal\{L\}\_\{\\mathrm\{t2l\}\}\\right\)\.\(12\)
This objective encourages a coordinate embedding to align with the semantic description of its surrounding urban environment\.
### IV\-DDistance\-Aware Regularisation
For a mini\-batch ofBBlocations, there are\(B2\)\\binom\{B\}\{2\}possible unordered pairs\. We randomly sampleK=⌊ρ\(B2\)⌋K=\\left\\lfloor\\rho\\binom\{B\}\{2\}\\right\\rfloorpairs, whereρ∈\(0,1\]\\rho\\in\(0,1\]is the pair\-sampling ratio\. Let𝒫B\\mathcal\{P\}\_\{B\}denote the sampled pair set\.
For each\(i,j\)∈𝒫B\(i,j\)\\in\\mathcal\{P\}\_\{B\}, we calculate the geographic distancedij=dgeo\(𝐱i,𝐱j\)d\_\{ij\}=d\_\{\\mathrm\{geo\}\}\(\\mathbf\{x\}\_\{i\},\\mathbf\{x\}\_\{j\}\)\. Exact normalisation over all training\-location pairs requiresO\(N2\)O\(N^\{2\}\)pairwise computations and is computationally expensive for datasets containing hundreds of thousands of locations\. We therefore apply min–max normalisation within each sampled mini\-batch:
d~ij=dij−dmin\(B\)dmax\(B\)−dmin\(B\)\+ϵ,\\tilde\{d\}\_\{ij\}=\\frac\{d\_\{ij\}\-d\_\{\\min\}^\{\(B\)\}\}\{d\_\{\\max\}^\{\(B\)\}\-d\_\{\\min\}^\{\(B\)\}\+\\epsilon\},\(13\)wheredmin\(B\)d\_\{\\min\}^\{\(B\)\}anddmax\(B\)d\_\{\\max\}^\{\(B\)\}are the minimum and maximum sampled distances andϵ\\epsilonensures numerical stability\.
The cosine similarity between two location embeddings is mapped to\[0,1\]\[0,1\]:
s~ij=1\+𝐳i⊤𝐳j2\.\\tilde\{s\}\_\{ij\}=\\frac\{1\+\\mathbf\{z\}\_\{i\}^\{\\top\}\\mathbf\{z\}\_\{j\}\}\{2\}\.\(14\)We define the batch\-relative target similarity as
sij∗=1−d~ijs\_\{ij\}^\{\*\}=1\-\\tilde\{d\}\_\{ij\}\(15\)and minimise
ℒdist=1K∑\(i,j\)∈𝒫B\(s~ij−sij∗\)2\.\\mathcal\{L\}\_\{\\mathrm\{dist\}\}=\\frac\{1\}\{K\}\\sum\_\{\(i,j\)\\in\\mathcal\{P\}\_\{B\}\}\\left\(\\tilde\{s\}\_\{ij\}\-s\_\{ij\}^\{\*\}\\right\)^\{2\}\.\(16\)
The linear target provides a bounded, parameter\-free spatial signal\. It is used as a soft regulariser rather than an exact geometric constraint: geographically closer locations are encouraged to be more similar, while the repulsive signal increases with geographic separation\.
### IV\-ETraining Objective
The complete LE4Mob objective is
ℒ=ℒsem\+λdistℒdist,\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{sem\}\}\+\\lambda\_\{\\mathrm\{dist\}\}\\mathcal\{L\}\_\{\\mathrm\{dist\}\},\(17\)whereλdist\\lambda\_\{\\mathrm\{dist\}\}balances semantic alignment and distance\-aware regularisation\.
During pre\-training, the text encodergϕg\_\{\\phi\}is frozen, while the location encoderfθf\_\{\\theta\}and projection layerpψp\_\{\\psi\}are optimised\. The resulting location encoder is then frozen and reused across downstream mobility tasks\.
### IV\-FIntegration with Downstream Mobility Models
#### IV\-F1Next Location Prediction
Following previous work\[[33](https://arxiv.org/html/2609.22117#bib.bib31),[9](https://arxiv.org/html/2609.22117#bib.bib1)\], we use a multi\-head self\-attention \(MHSA\) model to process the historical visitation sequence\. Let𝐡n\\mathbf\{h\}^\{n\}denote its final hidden representation\. The next location probabilities are
𝐏\(l^n\+1\)=Softmax\(MLPo\(𝐡n\)\),\{\\mathbf\{P\}\}\(\\hat\{l\}^\{n\+1\}\)=\\operatorname\{Softmax\}\\left\(\\mathrm\{MLP\}\_\{\\mathrm\{o\}\}\(\\mathbf\{h\}^\{n\}\)\\right\),\(18\)whereMLPo\\mathrm\{MLP\}\_\{\\mathrm\{o\}\}denotes the output module consisting of fully\-connected layers with residual connections\. The downstream model is trained using standard cross\-entropy loss:
ℒnext\_loc=−∑k=1\|ℒc\|ykn\+1log𝐏\(l^kn\+1\),\\mathcal\{L\}\_\{\\text\{next\\\_loc\}\}=\-\\sum\_\{k=1\}^\{\|\\mathcal\{L\}\_\{c\}\|\}y^\{n\+1\}\_\{k\}\\log\{\\mathbf\{P\}\}\(\\hat\{l\}^\{n\+1\}\_\{k\}\),\(19\)where𝐏\(l^kn\+1\)\{\\mathbf\{P\}\}\(\\hat\{l\}^\{n\+1\}\_\{k\}\)denotes the predicted probability of visiting thekkth location andykn\+1y^\{n\+1\}\_\{k\}is the one\-hot ground\-truth destination\. The same MHSA architecture is used for all representation methods, and all pre\-trained embeddings remain frozen\.
#### IV\-F2Commuter Flow Generation
For each Output Area \(OA\)lil\_\{i\}, we use its population\-weighted centroid𝐱i\\mathbf\{x\}\_\{i\}as the representative coordinate and obtain its location embedding as𝐞i=fθ\(PE\(𝐱i\)\)\\mathbf\{e\}\_\{i\}=f\_\{\\theta\}\(\\mathrm\{PE\}\(\\mathbf\{x\}\_\{i\}\)\)\.
We evaluate two downstream flow models\. The first follows DeepGravity\[[28](https://arxiv.org/html/2609.22117#bib.bib11)\], where each OD pair is represented by concatenating the origin embedding, destination embedding, and OD distance:𝐯ij=\[𝐞i;𝐞j;dij\]\\mathbf\{v\}\_\{ij\}=\[\\mathbf\{e\}\_\{i\};\\mathbf\{e\}\_\{j\};d\_\{ij\}\]\.
The vector𝐯ij\\mathbf\{v\}\_\{ij\}is passed through a shared feed\-forward neural network to produce an OD scoresijs\_\{ij\}, which is normalised over candidate destinations:
p^ij=exp\(sij\)∑k≠iexp\(sik\)\.\\hat\{p\}\_\{ij\}=\\frac\{\\exp\(s\_\{ij\}\)\}\{\\sum\_\{k\\neq i\}\\exp\(s\_\{ik\}\)\}\.\(20\)To assess the dependence on explicit distance, we also evaluate a distance\-ablated variant \(the results are presented in Section[VI\-C2](https://arxiv.org/html/2609.22117#S6.SS3.SSS2)\):𝐯ij−dist=\[𝐞i;𝐞j\]\\mathbf\{v\}\_\{ij\}^\{\-\\mathrm\{dist\}\}=\[\\mathbf\{e\}\_\{i\};\\mathbf\{e\}\_\{j\}\]\.
The second model is a bilinear decoder, which first computes a bilinear interaction between an OD pair by
𝐫ij=𝐞i⊤Wb𝐞j,\\mathbf\{r\}\_\{ij\}=\\mathbf\{e\}\_\{i\}^\{\\top\}W\_\{\\mathrm\{b\}\}\\mathbf\{e\}\_\{j\},\(21\)whereWbW\_\{\\mathrm\{b\}\}is learnable\. The resulting interaction vector is passed through an MLP that outputs one logit for each candidate destination\. The destination probabilities are similarly obtained by a Softmax function like Eq\.[20](https://arxiv.org/html/2609.22117#S4.E20)\.
Both models are trained using the cross entropy between the observed and predicted destination distributions:
ℒflow=−∑i∑j≠iyijOilogp^ij,\\mathcal\{L\}\_\{\\text\{flow\}\}=\-\\sum\_\{i\}\\sum\_\{j\\neq i\}\\frac\{y\_\{ij\}\}\{O\_\{i\}\}\\log\\hat\{p\}\_\{ij\},\(22\)Estimated flows are obtained from the predicted destination probabilities as defined above\.
## VExperimental Setup
We evaluate LE4Mob on individual\-level next location prediction and population\-level commuter flow generation\. The experiments address three questions:
- •RQ1:Can our pre\-trained location encoder support mobility tasks operating at different scales?
- •RQ2:How effectively does LE4Mob generalise when previously unseen locations appear after training?
- •RQ3:Does distance\-aware pre\-training improve the spatial structure and downstream utility of location embeddings, particularly when distance is not supplied separately to the downstream model?
### V\-ANext Location Prediction
#### V\-A1Datasets and Preprocessing
We use four public mobility datasets: Foursquare New York \(FSQ\-NYC\), Foursquare Tokyo \(FSQ\-TKY\)\[[37](https://arxiv.org/html/2609.22117#bib.bib4)\], Gowalla\-London \(Gowalla\-LD\)\[[4](https://arxiv.org/html/2609.22117#bib.bib24)\], and Geolife\[[41](https://arxiv.org/html/2609.22117#bib.bib3)\]\. The first three contain location\-based social network \(LBSN\) check\-ins, while Geolife contains GNSS trajectories\.
FSQ\-NYC and FSQ\-TKY provide venue information that can be used to construct contextual POI descriptions\. Gowalla and Geolife do not provide sufficiently detailed POI context; we therefore obtain POIs for London and Beijing from Foursquare Open Places\[[7](https://arxiv.org/html/2609.22117#bib.bib39)\]\.
TABLE I:Basic statistics of the mobility datasets after preprocessing\. The mean and standard deviation across users are reported\.FSQ\-NYCFSQ\-TKYGowalla\-LDGeolife\#Users53516436839\#Days tracked3193183671703\#Total unique locations4019695910651248\#Stays per user178\.3±206\.1178\.3\\pm 206\.1225\.1±220\.9225\.1\\pm 220\.9127\.1±123\.3127\.1\\pm 123\.3358\.6±383\.6358\.6\\pm 383\.6\#Stays per user per day2\.1±2\.12\.1\\pm 2\.12\.7±2\.72\.7\\pm 2\.72\.4±3\.02\.4\\pm 3\.02\.4±1\.52\.4\\pm 1\.5\#Unique locations per user34\.7±21\.434\.7\\pm 21\.456\.6±33\.556\.6\\pm 33\.546\.3±44\.246\.3\\pm 44\.261\.5±52\.861\.5\\pm 52\.8For the LBSN datasets, we remove POIs with fewer than ten check\-ins and users with fewer than ten records\. For Geolife, we remove users with fewer than 50 active days, extract stay points, and cluster spatially proximate stay points into discrete locations\. Additional preprocessing details and dataset statistics are reported in Table[I](https://arxiv.org/html/2609.22117#S5.T1)\.
#### V\-A2Baselines
We compare LE4Mob with the following location representation methods:
- •Vanilla\-E2E: a trainable lookup embedding optimised jointly with the downstream predictor;
- •Skip\-gram\[[21](https://arxiv.org/html/2609.22117#bib.bib15)\]: a Word2Vec\-style method that learns location co\-occurrence from mobility sequences;
- •POI2Vec\[[6](https://arxiv.org/html/2609.22117#bib.bib13)\]: a sequence\-based method incorporating spatial proximity through a geographical hierarchy;
- •Geo\-Teaser\[[40](https://arxiv.org/html/2609.22117#bib.bib25)\]: a geo\-temporal embedding method using spatial and temporal negative sampling;
- •TALE\[[32](https://arxiv.org/html/2609.22117#bib.bib8)\]: a time\-aware location embedding method based on a temporal tree structure;
- •CaLLiPer\[[33](https://arxiv.org/html/2609.22117#bib.bib31)\]: the direct predecessor of LE4Mob, trained using semantic alignment without the distance\-aware objective\.
#### V\-A3Conventional Setting
Each dataset is divided chronologically into non\-overlapping training, validation, and test sets using a 60%–20%–20% split\. Mobility sequences are constructed using a seven\-day sliding observation window, following previous work\[[9](https://arxiv.org/html/2609.22117#bib.bib1)\]\.
#### V\-A4Inductive Setting
The inductive setting simulates deployment in which a small number of previously unseen locations emerge after model training\. We first construct the complete location vocabulary and the chronological train–validation–test split\. We then randomly select 10% of locations as a held\-out setℒheld\\mathcal\{L\}^\{\\mathrm\{held\}\}and remove every training or validation sequence containing a held\-out location\. The full test set remains unchanged and therefore contains both seen and unseen locations\. To specifically assess inductive generalisation, the inductive results reported in Table[II](https://arxiv.org/html/2609.22117#S6.T2)are calculated over test samples involving at least one held\-out location\.
Coordinate\-context pairs associated withℒheld\\mathcal\{L\}^\{\\mathrm\{held\}\}are also excluded from LE4Mob and CaLLiPer pre\-training\. The test set remains unchanged and therefore contains a mixture of seen and unseen locations\.
Each experiment is repeated five times using different random seeds and independently sampled held\-out sets\. We report the mean and standard deviation\.
#### V\-A5Evaluation Metrics
We evaluate performance from three complementary perspectives:exact hit performance,ranking quality, andspatial fidelity\. Exact hit performance is measured usingAccuracy@KK, indicating whether the ground\-truth destination appears among the top\-KKpredictions; we report the accuracy metrics atK∈\{1,5\}K\\in\\\{1,5\\\}\. Ranking quality is assessed using Normalised Discounted Cumulative Gain at 10 \(nDCG@10\), which rewards placing the correct destination higher in the ranked list\. Spatial fidelity captures how geographically close the predicted locations are to the ground\-truth destination, complementing exact\-match and ranking\-based evaluation\. We measure it using Geographic Distance Error at 1 \(GDE@1\), defined as the geographic distance \(in kilometres\) between the top\-ranked prediction and the ground\-truth destination, and Mean Geographic Distance Error at 5 \(mGDE@5\), defined as the average geographic distance between the top five predictions and the ground\-truth destination\.
### V\-BCommuter Flow Generation
#### V\-B1Study Areas and Data
We evaluate commuter flow generation in London and Greater Manchester, an international global city and a major metropolitan area in northern England, respectively\. Commuting flows and population data at Output Area level are obtained from the UK 2021 Census\[[23](https://arxiv.org/html/2609.22117#bib.bib40)\]\. POIs are obtained from Foursquare Open Places\[[7](https://arxiv.org/html/2609.22117#bib.bib39)\]\.
Each OA is represented spatially by its population\-weighted centroid\. LE4Mob and the other location encoders generate an embedding at this coordinate\.
#### V\-B2Location\-Disjoint Split
For each city, OAs are randomly divided into mutually exclusive training, validation, and test sets using a 65%–15%–20% split\. For each split, we retain only flows whose origin and destination both belong to that split\. Flows crossing different subsets are excluded\. The total outflow is recomputed within each subset\. This node\-level inductive split ensures that test OAs and their associated flows are unavailable during downstream training\.
#### V\-B3Baselines
We compare LE4Mob with the following POI\-derived urban representations:
- •Hand\-Crafted: Handcrafted feature vectors consisting of POI proportions and population;
- •LDA\[[2](https://arxiv.org/html/2609.22117#bib.bib14)\]: a common baseline that learns probabilistic topic representation from regional POI distributions;
- •SPPE\[[10](https://arxiv.org/html/2609.22117#bib.bib10)\]: a POI\-based spatial representation method;
- •Urban2Vec\[[35](https://arxiv.org/html/2609.22117#bib.bib12)\]: a deep urban representation model learned from geographic context;
- •Space2Vec\[[19](https://arxiv.org/html/2609.22117#bib.bib6)\]: an inductive coordinate\-based location encoder; and
- •CaLLiPer\[[34](https://arxiv.org/html/2609.22117#bib.bib22)\]: the semantic\-focused predecessor of LE4Mob\.
Hand\-Crafted directly contains POI composition and population density, whereas learned baselines use their original pre\-trained representations without appending additional socioeconomic variables\.
Some graph learning\-based models\[[16](https://arxiv.org/html/2609.22117#bib.bib34),[38](https://arxiv.org/html/2609.22117#bib.bib35)\]are not included because they generally learn node representations from the full OD graph, violating our location\-disjoint evaluation protocol\.
#### V\-B4Downstream Flow Models
Each representation is evaluated with DeepGravity and Bilinear models\. The standard DeepGravity configuration includes distance and is used in the main performance comparison\. To isolate representation\-level distance awareness, we additionally remove the distance feature for every representation method\. This ablation is analysed separately in Section[VI\-C2](https://arxiv.org/html/2609.22117#S6.SS3.SSS2)\. The Bilinear model uses only the origin and destination representations\.
#### V\-B5Evaluation Metrics
The primary metric is the Common Part of Commuters \(CPC\), also known as the Sørensen–Dice index:
CPC=2∑i∑j≠imin\(yij,y^ij\)∑i∑j≠iyij\+∑i∑j≠iy^ij\.\\mathrm\{CPC\}=\\frac\{2\\sum\_\{i\}\\sum\_\{j\\neq i\}\\min\(y\_\{ij\},\\hat\{y\}\_\{ij\}\)\}\{\\sum\_\{i\}\\sum\_\{j\\neq i\}y\_\{ij\}\+\\sum\_\{i\}\\sum\_\{j\\neq i\}\\hat\{y\}\_\{ij\}\}\.\(23\)CPC lies in\[0,1\]\[0,1\], with larger values indicating greater overlap between observed and generated flows\.
We additionally report Mean Absolute Error \(MAE\), Root Mean Squared Error \(RMSE\), and Jensen–Shannon Divergence \(JSD\) to measure absolute error and distributional dissimilarity\.
### V\-CImplementation Details
#### V\-C1Location Representation Pre\-training
For fair comparison, we keep the architectural configuration for the location encoders consistent across Space2Vec, CaLLiPer, and LE4Mob\. We employ Grid\[[19](https://arxiv.org/html/2609.22117#bib.bib6)\]as the PE and FC\-Net as the neural network\. The mathematical formulation of Grid, withλ,ϕ\\lambda,\\phidenoting 2\-D coordinates, is as follows:
PE\(λ,ϕ\)=⋃s=0S−1\(cosλαs,sinλαs,cosϕαs,sinϕαs\)\\text\{PE\}\(\\lambda,\\phi\)=\\bigcup\_\{s=0\}^\{S\-1\}\(\\cos\\frac\{\\lambda\}\{\\alpha\_\{s\}\},\\sin\\frac\{\\lambda\}\{\\alpha\_\{s\}\},\\cos\\frac\{\\phi\}\{\\alpha\_\{s\}\},\\sin\\frac\{\\phi\}\{\\alpha\_\{s\}\}\)\(24\)
αs=rmin⋅\(rmaxrmin\)sS−1\\alpha\_\{s\}=r\_\{\\text\{min\}\}\\cdot\(\\frac\{r\_\{\\text\{max\}\}\}\{r\_\{\\text\{min\}\}\}\)^\{\\frac\{s\}\{S\-1\}\}\(25\)whererminr\_\{\\text\{min\}\}andrmaxr\_\{\\text\{max\}\}are the minimum and maximum radii, respectively, andSSis the number of scales\. These hyperparameters control the resolutions of the multi\-scale encoding of the coordinates\. We adopted Sentence Transformer\[[22](https://arxiv.org/html/2609.22117#bib.bib17)\]as the text encoder and a linear layer as the projection layer\.
The hyperparameters of Grid vary across different cities/datasets:rminr\_\{\\text\{min\}\}andrmaxr\_\{\\text\{max\}\}are set to 0\.01 and 10 for New York City \(FSQ\-NYC\) and Tokyo \(FSQ\-TKY\), 1 and 1000 for London \(Gowalla\-LD and flow generation\), 0\.01 and 10 for Beijing \(Geolife\), and 10 and 1000 for Manchester \(flow generation\)\. All other hyperparameters remain consistent: we set the number of scalesS=32S=32and the hidden dimension of the FC\-Net as 256\.
For LE4Mob, the pair\-sampling ratioρ\\rhoand loss weightλdist\\lambda\_\{\\mathrm\{dist\}\}were set as 0\.1 and 1, respectively\. For all other baselines, we tuned the hyperparameters via random search and trained them until convergence\.
All location embeddings are set to 128 dimensions except for two baselines in the flow generation task: Hand\-Crafted uses 10 dimensions, including 9 POI proportions and 1 population feature, while LDA dimensionality is selected based on perplexity score and set to 14 for London and 8 for Manchester\.
All the pre\-training has been conducted using a learning rate of 0\.001 and the Adam optimiser\. Batch sizes vary depending on the volumes of the training data\.
#### V\-C2Downstream Models
The hyperparameters of the MHSA model are kept consistent with the original paper\[[9](https://arxiv.org/html/2609.22117#bib.bib1)\]\. For the DeepGravity and Bilinear models, the number of hidden layers is set to 6 and the hidden dimension is set to 128\. We tune the dropout rate from 0, 0\.2, 0\.5, with the final value varying across models\. All downstream models are trained using the Adam optimiser with a learning rate of 0\.001\. Early stopping is applied\. The full settings are provided in the GitHub repository\.
To ensure that evaluation results are robust, and that observed improvements are consistent rather than due to random held\-out sets or spatial splits, we report the mean and standard deviation over five runs for all tasks\. For conventional next location prediction, the five runs use different MHSA initialisation seeds\. For inductive next location prediction, they use five differently sampled held\-out location sets, with MHSA initialised once\. Flow generation task uses five spatially disjoint OA splits generated with different seeds\.
## VIResults
### VI\-APerformance on Next Location Prediction
TABLE II:Performance comparison of different embedding methods on the next location prediction task\. The best and second\-best performance are marked inboldandunderlined, respectively, based on the unrounded mean values\. For better readability, the ranking metric values are scaled by a factor of10210^\{2\}, while GDE and mGDE are reported in kilometres\.ConventionalInductiveDataModelAcc@1↑\\uparrowAcc@5↑\\uparrownDCG@10↑\\uparrowGDE@1↓\\downarrowmGDE@5↓\\downarrowAcc@1↑\\uparrowAcc@5↑\\uparrownDCG@10↑\\uparrowGDE@1↓\\downarrowmGDE@5↓\\downarrowFSQ\-NYCVanilla\-E2E19\.94±\\pm0\.1946\.84±\\pm0\.4437\.52±\\pm0\.254\.45±\\pm0\.095\.36±\\pm0\.109\.80±\\pm1\.1824\.41±\\pm2\.1019\.23±\\pm1\.937\.53±\\pm0\.758\.47±\\pm0\.69Skip\-gram19\.81±\\pm0\.3445\.76±\\pm0\.1936\.85±\\pm0\.294\.43±\\pm0\.065\.20±\\pm0\.038\.68±\\pm1\.1921\.71±\\pm2\.9217\.35±\\pm2\.408\.07±\\pm0\.828\.96±\\pm1\.15POI2Vec19\.62±\\pm0\.5844\.94±\\pm0\.7136\.32±\\pm0\.664\.53±\\pm0\.165\.38±\\pm0\.128\.84±\\pm1\.3121\.80±\\pm3\.4017\.37±\\pm2\.708\.26±\\pm0\.738\.74±\\pm0\.72Geo\-Teaser19\.43±\\pm0\.7345\.46±\\pm0\.8436\.48±\\pm0\.594\.52±\\pm0\.085\.30±\\pm0\.088\.08±\\pm1\.4720\.33±\\pm3\.1316\.16±\\pm2\.748\.77±\\pm0\.489\.26±\\pm0\.61TALE20\.20±\\pm0\.4045\.80±\\pm0\.1236\.98±\\pm0\.124\.46±\\pm0\.065\.35±\\pm0\.058\.52±\\pm1\.4721\.04±\\pm3\.4116\.83±\\pm2\.768\.23±\\pm0\.999\.07±\\pm0\.99CaLLiPer20\.33±\\pm0\.3648\.38±\\pm0\.1538\.67±\\pm0\.124\.10±\\pm0\.034\.85±\\pm0\.0210\.63±\\pm0\.9726\.59±\\pm1\.6921\.13±\\pm1\.565\.44±\\pm0\.355\.97±\\pm0\.37LE4Mob20\.45±\\pm0\.3248\.43±\\pm0\.3638\.73±\\pm0\.184\.09±\\pm0\.054\.82±\\pm0\.0210\.74±\\pm1\.3327\.04±\\pm2\.4021\.37±\\pm2\.095\.20±\\pm0\.395\.65±\\pm0\.27FSQ\-TKYVanilla\-E2E21\.54±\\pm0\.0645\.54±\\pm0\.1637\.32±\\pm0\.034\.58±\\pm0\.025\.41±\\pm0\.0111\.72±\\pm0\.4828\.01±\\pm1\.0422\.45±\\pm0\.935\.86±\\pm0\.256\.56±\\pm0\.22Skip\-gram21\.85±\\pm0\.2346\.38±\\pm0\.1037\.93±\\pm0\.134\.41±\\pm0\.035\.22±\\pm0\.0212\.02±\\pm0\.5928\.74±\\pm1\.1223\.07±\\pm1\.015\.49±\\pm0\.196\.09±\\pm0\.16POI2Vec22\.03±\\pm0\.1346\.35±\\pm0\.0737\.98±\\pm0\.064\.40±\\pm0\.025\.21±\\pm0\.0012\.18±\\pm0\.7428\.71±\\pm1\.5023\.08±\\pm1\.205\.63±\\pm0\.316\.26±\\pm0\.29Geo\-Teaser21\.86±\\pm0\.1946\.70±\\pm0\.1438\.13±\\pm0\.144\.40±\\pm0\.025\.21±\\pm0\.0211\.93±\\pm0\.7028\.56±\\pm1\.2922\.93±\\pm1\.125\.55±\\pm0\.276\.15±\\pm0\.20TALE21\.44±\\pm0\.0945\.83±\\pm0\.0337\.46±\\pm0\.064\.60±\\pm0\.015\.42±\\pm0\.0211\.92±\\pm0\.5128\.78±\\pm1\.1922\.96±\\pm0\.955\.81±\\pm0\.286\.47±\\pm0\.24CaLLiPer20\.20±\\pm0\.1446\.19±\\pm0\.0937\.19±\\pm0\.074\.51±\\pm0\.025\.25±\\pm0\.0112\.07±\\pm0\.4529\.72±\\pm1\.2823\.69±\\pm0\.975\.37±\\pm0\.245\.94±\\pm0\.23LE4Mob20\.26±\\pm0\.1146\.33±\\pm0\.0937\.29±\\pm0\.064\.47±\\pm0\.025\.22±\\pm0\.0112\.04±\\pm0\.4829\.88±\\pm1\.3323\.77±\\pm0\.965\.35±\\pm0\.265\.84±\\pm0\.21Gowalla\-LDVanilla\-E2E13\.19±\\pm0\.9828\.05±\\pm2\.1422\.54±\\pm1\.436\.04±\\pm0\.297\.13±\\pm0\.435\.18±\\pm2\.1310\.79±\\pm4\.949\.10±\\pm3\.978\.62±\\pm0\.729\.03±\\pm0\.81Skip\-gram16\.67±\\pm0\.7735\.67±\\pm1\.3628\.85±\\pm0\.704\.79±\\pm0\.366\.03±\\pm0\.228\.65±\\pm3\.5318\.18±\\pm6\.2314\.55±\\pm4\.726\.73±\\pm1\.027\.59±\\pm0\.61POI2Vec15\.67±\\pm0\.8933\.06±\\pm1\.4927\.09±\\pm1\.265\.18±\\pm0\.156\.15±\\pm0\.167\.36±\\pm2\.1916\.29±\\pm5\.0013\.36±\\pm3\.938\.20±\\pm1\.038\.80±\\pm0\.95Geo\-Teaser16\.48±\\pm0\.6135\.20±\\pm0\.4428\.38±\\pm0\.354\.64±\\pm0\.295\.82±\\pm0\.278\.27±\\pm3\.0917\.41±\\pm6\.4714\.05±\\pm4\.717\.67±\\pm2\.307\.80±\\pm1\.19TALE14\.73±\\pm0\.9730\.33±\\pm1\.6524\.67±\\pm0\.955\.56±\\pm0\.396\.76±\\pm0\.246\.95±\\pm3\.2715\.44±\\pm4\.9712\.40±\\pm4\.227\.08±\\pm0\.917\.97±\\pm0\.87CaLLiPer16\.58±\\pm1\.0936\.07±\\pm2\.1829\.35±\\pm1\.354\.45±\\pm0\.255\.62±\\pm0\.215\.77±\\pm2\.7714\.33±\\pm3\.5611\.47±\\pm3\.547\.83±\\pm1\.358\.15±\\pm0\.80LE4Mob19\.02±\\pm0\.5539\.34±\\pm1\.1431\.96±\\pm0\.374\.39±\\pm0\.125\.55±\\pm0\.329\.80±\\pm3\.8522\.79±\\pm8\.2517\.90±\\pm6\.276\.90±\\pm2\.177\.54±\\pm1\.34GeolifeVanilla\-E2E44\.15±\\pm0\.5474\.72±\\pm2\.3262\.79±\\pm0\.862\.71±\\pm0\.054\.19±\\pm0\.1432\.16±\\pm3\.3453\.27±\\pm6\.0144\.66±\\pm4\.733\.86±\\pm0\.844\.93±\\pm1\.00Skip\-gram45\.32±\\pm0\.5077\.75±\\pm1\.8064\.36±\\pm0\.942\.84±\\pm0\.073\.92±\\pm0\.0833\.10±\\pm4\.9456\.19±\\pm6\.0846\.35±\\pm5\.923\.67±\\pm0\.814\.67±\\pm1\.08POI2Vec43\.08±\\pm1\.5174\.53±\\pm1\.8862\.17±\\pm1\.182\.82±\\pm0\.073\.93±\\pm0\.0329\.71±\\pm5\.1051\.11±\\pm7\.0242\.34±\\pm6\.273\.84±\\pm0\.935\.17±\\pm0\.96Geo\-Teaser46\.30±\\pm1\.1275\.01±\\pm0\.9563\.50±\\pm0\.772\.74±\\pm0\.103\.88±\\pm0\.0534\.91±\\pm4\.4355\.53±\\pm7\.7147\.51±\\pm5\.983\.54±\\pm0\.964\.64±\\pm1\.13TALE45\.09±\\pm1\.8774\.75±\\pm2\.6262\.92±\\pm1\.892\.82±\\pm0\.113\.99±\\pm0\.1834\.92±\\pm6\.2455\.81±\\pm7\.2447\.08±\\pm7\.043\.61±\\pm0\.874\.69±\\pm1\.01CaLLiPer46\.28±\\pm2\.6274\.36±\\pm0\.4563\.22±\\pm1\.692\.64±\\pm0\.153\.90±\\pm0\.0735\.18±\\pm4\.0556\.34±\\pm7\.3247\.63±\\pm5\.713\.51±\\pm1\.094\.60±\\pm1\.05LE4Mob46\.55±\\pm1\.3076\.43±\\pm0\.7764\.54±\\pm0\.322\.57±\\pm0\.123\.83±\\pm0\.1535\.35±\\pm4\.3055\.34±\\pm8\.4647\.07±\\pm6\.413\.39±\\pm0\.874\.50±\\pm0\.92Table[II](https://arxiv.org/html/2609.22117#S6.T2)compares the location representations under conventional and inductive next location prediction\.LE4Mob achieves the highest mean performance in 30 of the 40 dataset–setting–metric combinations, including 13 of the 16 spatial error comparisons\. On FSQ\-NYC, it consistently outperforms all baselines across both ranking and spatial metrics under conventional and inductive evaluation\. Strong improvements are also observed on Gowalla\-LD: across the three ranking metrics, LE4Mob improves over the strongest competing method by approximately 8\.9–14\.1% under conventional evaluation and 13\.3–25\.4% under inductive evaluation\. Its spatial errors are also generally favourable\. On FSQ\-TKY, trajectory\-derived representations remain stronger under conventional evaluation, whereas LE4Mob achieves the best performance on four of the five inductive metrics, with POI2Vec leading Acc@1\. On Geolife, LE4Mob achieves the best results in seven of the ten reported comparisons, including all four spatial error metrics\. Overall, LE4Mob shows its most consistent advantages under inductive evaluation and on metrics that directly measure the spatial proximity of predictions\.
Baseline performance varies across datasets and evaluation settings\. Trajectory\-derived embeddings perform well when their learned mobility patterns match a dataset, such as POI2Vec and Geo\-Teaser on FSQ\-TKY, and Skip\-gram on Gowalla\-LD and Geolife\. CaLLiPer is the most consistently competitive baseline, especially on FSQ\-NYC and Geolife, highlighting the value of inductive spatial\-semantic representations\. In comparison, LE4Mob achieves the best overall results across substantially more dataset–metric combinations, demonstrating greater consistency across heterogeneous mobility environments\. It also records the lowest error in 13 of the 16 spatial error comparisons, indicating that distance\-aware pre\-training improves both destination ranking and the spatial proximity of predictions\.
### VI\-BPerformance on Flow Generation
TABLE III:Performance comparison of different embedding methods on the flow generation task\. The best and second\-best performances are determined using unrounded mean values and marked inboldandunderlined, respectively\.DeepGravityBilinearDataModelCPC↑\\uparrowMAE↓\\downarrowRMSE↓\\downarrowJSD↓\\downarrowCPC↑\\uparrowMAE↓\\downarrowRMSE↓\\downarrowJSD↓\\downarrowManchesterHandcrafted0\.765±\\pm0\.0080\.693±\\pm0\.0550\.871±\\pm0\.0830\.202±\\pm0\.0050\.757±\\pm0\.0110\.717±\\pm0\.0630\.915±\\pm0\.0910\.207±\\pm0\.008LDA0\.784±\\pm0\.0170\.649±\\pm0\.0800\.835±\\pm0\.1160\.183±\\pm0\.0130\.769±\\pm0\.0210\.694±\\pm0\.0950\.904±\\pm0\.1420\.192±\\pm0\.019SPPE0\.778±\\pm0\.0120\.669±\\pm0\.0660\.866±\\pm0\.0970\.189±\\pm0\.0090\.796±\\pm0\.0090\.610±\\pm0\.0580\.795±\\pm0\.0890\.173±\\pm0\.006Urban2Vec0\.769±\\pm0\.0170\.693±\\pm0\.0830\.895±\\pm0\.1190\.196±\\pm0\.0140\.809±\\pm0\.0170\.584±\\pm0\.0780\.793±\\pm0\.1090\.157±\\pm0\.012Space2Vec0\.764±\\pm0\.0160\.705±\\pm0\.0800\.903±\\pm0\.1160\.199±\\pm0\.0130\.783±\\pm0\.0220\.660±\\pm0\.0950\.892±\\pm0\.1400\.183±\\pm0\.019CaLLiPer0\.775±\\pm0\.0150\.675±\\pm0\.0740\.869±\\pm0\.1080\.188±\\pm0\.0120\.806±\\pm0\.0170\.603±\\pm0\.0800\.821±\\pm0\.1210\.163±\\pm0\.015LE4Mob0\.779±\\pm0\.0140\.664±\\pm0\.0730\.858±\\pm0\.1070\.184±\\pm0\.0120\.814±\\pm0\.0150\.582±\\pm0\.0720\.788±\\pm0\.1070\.154±\\pm0\.012LondonHandcrafted0\.786±\\pm0\.0020\.579±\\pm0\.0220\.715±\\pm0\.0380\.186±\\pm0\.0020\.800±\\pm0\.0050\.542±\\pm0\.0250\.686±\\pm0\.0440\.175±\\pm0\.005LDA0\.809±\\pm0\.0030\.524±\\pm0\.0210\.662±\\pm0\.0350\.165±\\pm0\.0020\.797±\\pm0\.0070\.555±\\pm0\.0320\.710±\\pm0\.0520\.171±\\pm0\.006SPPE0\.812±\\pm0\.0040\.516±\\pm0\.0250\.651±\\pm0\.0410\.160±\\pm0\.0030\.810±\\pm0\.0030\.523±\\pm0\.0220\.666±\\pm0\.0400\.159±\\pm0\.003Urban2Vec0\.793±\\pm0\.0080\.564±\\pm0\.0350\.707±\\pm0\.0530\.175±\\pm0\.0070\.826±\\pm0\.0070\.487±\\pm0\.0270\.676±\\pm0\.0490\.150±\\pm0\.007Space2Vec0\.798±\\pm0\.0080\.554±\\pm0\.0360\.689±\\pm0\.0540\.173±\\pm0\.0070\.813±\\pm0\.0080\.516±\\pm0\.0350\.651±\\pm0\.0530\.157±\\pm0\.007CaLLiPer0\.807±\\pm0\.0080\.532±\\pm0\.0350\.671±\\pm0\.0530\.162±\\pm0\.0070\.817±\\pm0\.0080\.508±\\pm0\.0320\.671±\\pm0\.0540\.153±\\pm0\.007LE4Mob0\.830±\\pm0\.0100\.477±\\pm0\.0390\.630±\\pm0\.0600\.142±\\pm0\.0090\.838±\\pm0\.0100\.459±\\pm0\.0370\.613±\\pm0\.0580\.135±\\pm0\.008Table[III](https://arxiv.org/html/2609.22117#S6.T3)compares the embedding methods using DeepGravity and Bilinear downstream models\. Among all embedding–model combinations,LE4Mob with Bilinear achieves the strongest overall results, while LE4Mob records the highest individual mean performance in12 of the 16 city–model–metric combinations\. It also consistently outperforms CaLLiPer across all evaluated configurations, indicating that explicitly preserving spatial proximity improves the utility of spatial\-semantic representations for flow generation\.
In Greater Manchester, LE4Mob obtains the strongest Bilinear results across all four metrics\. The DeepGravity results present the main exception: LDA performs best, while LE4Mob remains competitive, and continues to outperform CaLLiPer\. This suggests that the effectiveness of a location representation depends partly on how the downstream architecture incorporates spatial information\. As DeepGravity already receives explicit OD distance, the spatial signal encoded by LE4Mob may become partly redundant in this configuration\.
In London, LE4Mob consistently outperforms all baseline methods under both downstream architectures\. With DeepGravity, it improves CPC by 2\.2% and reduces MAE, RMSE, and JSD by 7\.6%, 3\.2%, and 11\.3%, respectively, relative to the strongest competing method\. With Bilinear, it improves CPC by 1\.5% and reduces MAE, RMSE, and JSD by 5\.7%, 5\.8%, and 10\.0%, respectively, compared to the second\-best results\. These results demonstrate that LE4Mob can support both nonlinear feature\-based and direct origin–destination interaction models
The baseline methods exhibit different strengths across configurations: LDA perform strongly with DeepGravity in Greater Manchester, SPPE is competitive with DeepGravity in London, and Urban2Vec is the strongest Bilinear baseline\. However, none of these methods performs better consistently across all configurations\. In comparison, LE4Mob provides the strongest overall performance and the most consistent improvements across cities, downstream models, and evaluation metrics\.
### VI\-CAnalysis of Distance Awareness
#### VI\-C1Correspondence between Geographic and Embedding Similarities
TABLE IV:Correspondence between geographic similarity and embedding cosine similarity for OA pairs in the Greater London test set\. All correlations are statistically significant atp<0\.0001p<0\.0001\.ModelPearson’srrSpearman’sρ\\rhoCaLLiPer0\.2680\.315LE4Mob0\.3690\.438We next examine whether the distance\-preserving objective successfully incorporates geographic proximity into the learned location embeddings\. We compare LE4Mob with CaLLiPer, as LE4Mob extends the same spatial\-semantic representation framework by introducing distance\-aware regularisation\. London is selected as the case\-study area\.
For each pair of OAs in the test set, we define geographic similarity assijgeo=1−d~ijs^\{\\mathrm\{geo\}\}\_\{ij\}=1\-\\tilde\{d\}\_\{ij\}, whered~ij\\tilde\{d\}\_\{ij\}is the min\-max normalised Euclidean distance between their centroids\. Embedding similarity is measured using the cosine similarity between the corresponding embedding vectors\. We then calculate Pearson’s correlation coefficient to assess the linear correspondence between geographic and embedding similarities, and Spearman’s rank correlation coefficient to assess the preservation of their relative ordering\. Higher coefficient values indicate stronger preservation of geographic proximity\.
As shown in Table[IV](https://arxiv.org/html/2609.22117#S6.T4), both methods exhibit statistically significant positive correlations\. However, LE4Mob increases the Pearson correlation from 0\.268 to 0\.369 and the Spearman correlation from 0\.315 to 0\.438, representing relative improvements of approximately 38% and 39%, respectively\. The higher Pearson correlation indicates that embedding similarity varies more consistently with the magnitude of geographic proximity, while the higher Spearman correlation shows that LE4Mob better preserves the relative ordering of nearby and distant location pairs\. These results provide direct evidence that the distance\-preserving objective more effectively incorporates physical spatial relationships into the learned embeddings\.
#### VI\-C2Flow Generation without Explicit Distance Inputs
To examine whether distance information is encoded directly in the location representations, we remove the explicit origin–destination distance feature from DeepGravity and compare the resulting CPC with the standard configuration \(the left part of Table[III](https://arxiv.org/html/2609.22117#S6.T3)\)\. As shown in Figure[2](https://arxiv.org/html/2609.22117#S6.F2), removing distance consistently reduces the performance of Hand\-Crafted and LDA representations\. Their CPC decreases by1\.80%1\.80\\%–2\.91%2\.91\\%in Greater Manchester and2\.52%2\.52\\%–3\.45%3\.45\\%in London\. This is because that these representations do not explicitly preserve the spatial relationships between regions and the downstream DeepGravity model rely substantially on the separately provided distance feature\.
The downstream DeepGravity model generally produces more robust results after the removal of distance when paired with deep location embeddings, although the responses vary across methods and cities\. Most notably, LE4Mob improves from 0\.7791 to 0\.8106 in Greater Manchester and from 0\.8297 to 0\.8406 in London, achieving the highest CPC among all representations in both cities when distance is excluded\. CaLLiPer also remains robust, with improvements of0\.53%0\.53\\%and1\.44%1\.44\\%, while Urban2Vec records increases of4\.28%4\.28\\%and0\.90%0\.90\\%\. By contrast, the effects on SPPE and Space2Vec are less consistent across the two cities\.
These results indicate that explicit distance is important when the input representation itself contains limited spatial relational information\. In contrast, LE4Mob can support accurate flow generation without requiring distance as an additional handcrafted input\. The improvement after removing distance further suggests that, once spatial separation has been incorporated into the embedding space, supplying it again as a separate feature may be redundant and could introduce competing signals into the downstream model\. Nevertheless, because some other learned embeddings also remain robust without distance, this ablation should be interpreted as evidence that LE4Mob successfully encodes usable distance\-related information, rather than as evidence that such information is exclusive to LE4Mob\.
Fig\. 2:Flow generation performance with and without explicit distance inputs in the downstream DeepGravity model\. Error bars indicate standard deviations across repeated runs\. Percentages denote the relative CPC change after removing the explicit distance feature\. HC, U2V, S2V, and CaL denote Hand\-Crafted, Urban2Vec, Space2Vec, and CaLLiPer, respectively\.
## VIIDiscussion
### VII\-ARole of Distance\-Aware Representation Learning
The results show that embedding spatial structure into the location representation can benefit mobility modelling tasks\. Although many mobility models already use distance explicitly, especially in flow generation, supplying distance as a downstream OD\-pair feature is different from learning embeddings whose geometry already reflects spatial relationships\. The former is decoder\-specific; the latter makes spatial information reusable across models\.
In next location prediction, spatially structured embeddings may provide the Transformer\-based model with a more informative input space for learning distance\-related mobility regularities, as reflected in LE4Mob’s consistently lower spatial errors\. This property is also practically relevant to location\-based services and POI recommendation, where users’ destination choices are strongly distance\-sensitive; when the exact destination is not identified, a geographically closer prediction can still provide a relevant and actionable alternative\. In flow generation, the effect is more direct: the Bilinear model estimates OD compatibility from interactions between origin and destination embeddings, and therefore benefits strongly from embeddings that preserve spatial relationships\.
The DeepGravity results further support this interpretation\. When explicit OD distance is included, LE4Mob is competitive but does not always outperform all baselines, suggesting that the decoder can already exploit the supplied distance feature\. When distance is removed, however, baseline representations degrade, while LE4Mob improves and achieves the highest CPC in both London and Greater Manchester\. This indicates that LE4Mob preserves useful spatial information in the embeddings\. Its contribution is therefore not to replace explicit distance features, but to make location embeddings more spatially self\-contained and less dependent on decoder\-specific distance engineering\.
### VII\-BImplications for Mobility Modelling
LE4Mob separates the representation of places from the modelling of movement behaviour\. Instead of learning embeddings from trajectories or OD flows, it learns from geographic context and transfers the resulting representations to downstream mobility tasks\. The encoder captures where places are, what functions they serve, and how they relate spatially, while downstream models such as MHSA, DeepGravity, and Bilinear learn task\-specific behavioural patterns\.
This separation is particularly relevant because next location prediction and commuter flow generation differ in scale, data structure, and modelling objective\. The results provide evidence that geography\-derived embeddings can serve as a shared representational layer across individual\- and population\-level mobility models\. This modular design also points towards future mobility foundation models in which location representation and behavioural modelling are handled by separate but connected components\.
LE4Mob is also grounded in established mobility theory and recent advances in urban representation learning\. Classical gravity models explain spatial interaction through the masses \(population\) of origins and destinations and the distance between them\[[43](https://arxiv.org/html/2609.22117#bib.bib29)\], while DeepGravity demonstrates that incorporating richer geographic attributes through deep neural networks can substantially improve flow estimation\[[28](https://arxiv.org/html/2609.22117#bib.bib11)\]\. LE4Mob follows this broader principle by encoding place functions through POI\-derived semantic context and spatial relationships through distance\-aware regularisation\. Previous research has further shown that POI\-derived urban representations can capture information associated with population and socioeconomic characteristics\[[15](https://arxiv.org/html/2609.22117#bib.bib41)\]\. The combination of spatial semantics and distance awareness therefore provides a theoretically motivated representation of the factors shaping mobility, which may partly explain its effectiveness across the evaluated tasks\.
### VII\-CLimitations and Future Work
This study focuses on next location prediction and flow generation because both directly test location representation quality\. It is not intended as a benchmark across all mobility tasks\. Trajectory generation and crowd flow prediction involve additional challenges, such as sequential path generation and time\-series forecasting, and would require different experimental designs\.
LE4Mob’s inductiveness is also evaluated within the geographic domain covered by pre\-training\. It can encode unseen locations in the study area, but this should not be conflated with zero\-shot transfer to entirely new cities\. Cross\-city transfer would involve substantial domain shift in POI distributions, urban morphology, and semantic–spatial relationships\. Future work could explore multi\-city or national\-scale pre\-training for broader geographic generalisation\.
Finally, LE4Mob relies on POI\-derived contextual descriptions and uses a soft distance\-aware regulariser\. POIs are unevenly distributed, so representations may be better constrained in POI\-dense areas than in sparse areas\. Meanwhile, the distance objective encourages spatial consistency but does not enforce exact global distance preservation\. This is a deliberate trade\-off, since rigid distance preservation could weaken semantic organisation\. Future work could explore spatially balanced sampling, auxiliary spatial anchors, and alternative distance formulations \(e\.g\. travel time\) and distance\-aware objectives\.
## VIIIConclusion
This study proposed LE4Mob, an inductive, distance\-aware, and geography\-derived location embedding framework for human mobility modelling\. Unlike mobility\-derived embeddings learned from task\-specific trajectories or OD flows, LE4Mob pre\-trains a location encoder from coordinates and POI\-derived semantic context, while explicitly regularising the embedding space using geographic distance\. By capturing both the functional characteristics of places and their spatial relationships, LE4Mob provides a reusable representation layer for downstream mobility models\.
We evaluated LE4Mob on individual\-level next location prediction and population\-level commuter flow generation\. The results show that geography\-derived embeddings can transfer across substantially different mobility tasks and that representation\-level distance awareness improves spatial informativeness, especially when downstream models rely directly on embedding interactions or lack explicit distance inputs\. Future work will explore broader multi\-city pre\-training and alternative distance objectives to improve geographic generalisation\. Overall, LE4Mob offers a promising step towards shared geographic representations for human mobility modelling\.
## References
- \[1\]H\. Barbosa, M\. Barthelemy, G\. Ghoshal, C\. R\. James, M\. Lenormand, T\. Louail, R\. Menezes, J\. J\. Ramasco, F\. Simini, and M\. Tomasini\(2018\)Human mobility: models and applications\.Physics Reports734,pp\. 1–74\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p1.1),[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p1.1)\.
- \[2\]D\. M\. Blei, A\. Y\. Ng, and M\. I\. Jordan\(2003\)Latent dirichlet allocation\.Journal of machine Learning research3\(Jan\),pp\. 993–1022\.Cited by:[2nd item](https://arxiv.org/html/2609.22117#S5.I3.i2.p1.1)\.
- \[3\]R\. Cervero and K\. Kockelman\(1997\)Travel demand and the 3ds: density, diversity, and design\.Transportation research part D: Transport and environment2\(3\),pp\. 199–219\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p4.1)\.
- \[4\]E\. Cho, S\. A\. Myers, and J\. Leskovec\(2011\)Friendship and mobility: user movement in location\-based social networks\.InProceedings of the 17th ACM SIGKDD international conference on Knowledge discovery and data mining,pp\. 1082–1090\.Cited by:[§V\-A1](https://arxiv.org/html/2609.22117#S5.SS1.SSS1.p1.1)\.
- \[5\]J\. Feng, Y\. Li, C\. Zhang, F\. Sun, F\. Meng, A\. Guo, and D\. Jin\(2018\)Deepmove: predicting human mobility with attentional recurrent networks\.InProceedings of the 2018 world wide web conference,pp\. 1459–1468\.Cited by:[§II\-B](https://arxiv.org/html/2609.22117#S2.SS2.p1.1)\.
- \[6\]S\. Feng, G\. Cong, B\. An, and Y\. M\. Chee\(2017\)Poi2vec: geographical latent representation for predicting future visitors\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.31\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.22117#S2.SS2.p2.1),[3rd item](https://arxiv.org/html/2609.22117#S5.I2.i3.p1.1.1)\.
- \[7\]Foursquare\(2024\)Foursquare open source places: a new foundational dataset for the geospatial community\(Website\)External Links:[Link](https://location.foursquare.com/resources/blog/products/foursquare-open-source-places-a-new-foundational-dataset-for-the-geospatial-community/)Cited by:[§V\-A1](https://arxiv.org/html/2609.22117#S5.SS1.SSS1.p2.1),[§V\-B1](https://arxiv.org/html/2609.22117#S5.SS2.SSS1.p1.1)\.
- \[8\]T\. Han, F\. Li, C\. Chen, H\. Huang, Y\. Chen, and M\. Wu\(2026\)Spatially\-weighted clip for street\-view geo\-localization\.arXiv preprint arXiv:2604\.04357\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p6.1)\.
- \[9\]Y\. Hong, Y\. Zhang, K\. Schindler, and M\. Raubal\(2023\)Context\-aware multi\-head self\-attentional neural network model for next location prediction\.Transportation Research Part C: Emerging Technologies156,pp\. 104315\.Cited by:[§II\-B](https://arxiv.org/html/2609.22117#S2.SS2.p1.1),[§II](https://arxiv.org/html/2609.22117#S2.p2.1),[§IV\-F1](https://arxiv.org/html/2609.22117#S4.SS6.SSS1.p1.1),[§V\-A3](https://arxiv.org/html/2609.22117#S5.SS1.SSS3.p1.1),[§V\-C2](https://arxiv.org/html/2609.22117#S5.SS3.SSS2.p1.1)\.
- \[10\]W\. Huang, L\. Cui, M\. Chen, D\. Zhang, and Y\. Yao\(2022\)Estimating urban functional distributions with semantics preserved poi embedding\.International Journal of Geographical Information Science36\(10\),pp\. 1905–1930\.Cited by:[3rd item](https://arxiv.org/html/2609.22117#S5.I3.i3.p1.1.1)\.
- \[11\]K\. Klemmer, E\. Rolf, C\. Robinson, L\. Mackey, and M\. Rußwurm\(2025\)Satclip: global, general\-purpose location embeddings with satellite imagery\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 4347–4355\.Cited by:[§II\-A](https://arxiv.org/html/2609.22117#S2.SS1.p2.1)\.
- \[12\]M\. Lee and P\. Holme\(2015\)Relating land use and human intra\-city mobility\.PloS one10\(10\),pp\. e0140152\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p4.1)\.
- \[13\]Y\. Lin, H\. Wan, S\. Guo, and Y\. Lin\(2021\)Pre\-training context and time aware location embeddings from spatial\-temporal trajectories for user next location prediction\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.35,pp\. 4241–4248\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.22117#S2.SS2.p2.1)\.
- \[14\]J\. Liu, X\. Wang, and T\. Cheng\(2025\)Enriching Location Representation with Detailed Semantic Information\.In13th International Conference on Geographic Information Science \(GIScience 2025\),Leibniz International Proceedings in Informatics \(LIPIcs\), Vol\.346,pp\. 3:1–3:15\.Note:Keywords: Location Embedding, Contrastive Learning, Pretrained ModelExternal Links:ISBN 978\-3\-95977\-378\-2,ISSN 1868\-8969,[Document](https://dx.doi.org/10.4230/LIPIcs.GIScience.2025.3)Cited by:[§II\-A](https://arxiv.org/html/2609.22117#S2.SS1.p2.1)\.
- \[15\]J\. Liu, X\. Wang, Z\. Zeng, J\. Feng, Q\. Qin, I\. Ilyankou, G\. Dong, and T\. Cheng\(2026\)CITYREP: a unified benchmark for urban representations across cities, tasks, and modalities\.arXiv preprint arXiv:2605\.26036\.Cited by:[§VII\-B](https://arxiv.org/html/2609.22117#S7.SS2.p3.1)\.
- \[16\]Z\. Liu, F\. Miranda, W\. Xiong, J\. Yang, Q\. Wang, and C\. Silva\(2020\)Learning geo\-contextual embeddings for commuting flow prediction\.InProceedings of the AAAI conference on artificial intelligence,Vol\.34,pp\. 808–816\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p3.1),[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p3.1),[§V\-B3](https://arxiv.org/html/2609.22117#S5.SS2.SSS3.p4.1)\.
- \[17\]M\. Luca, G\. Barlacchi, B\. Lepri, and L\. Pappalardo\(2021\)A survey on deep learning for human mobility\.ACM Computing Surveys \(CSUR\)55\(1\),pp\. 1–44\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p1.1)\.
- \[18\]G\. Mai, K\. Janowicz, Y\. Hu, S\. Gao, B\. Yan, R\. Zhu, L\. Cai, and N\. Lao\(2022\)A review of location encoding for geoai: methods and applications\.International Journal of Geographical Information Science36,pp\. 639–673\.External Links:ISSN 1365\-8816, 1362\-3087,[Document](https://dx.doi.org/10.1080/13658816.2021.2004602)Cited by:[§II\-A](https://arxiv.org/html/2609.22117#S2.SS1.p1.1),[§IV\-A](https://arxiv.org/html/2609.22117#S4.SS1.p1.2),[§IV\-A](https://arxiv.org/html/2609.22117#S4.SS1.p2.2)\.
- \[19\]G\. Mai, K\. Janowicz, B\. Yan, R\. Zhu, L\. Cai, and N\. Lao\(2020\)Multi\-scale representation learning for spatial feature distributions using grid cells\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=rJljdh4KDH)Cited by:[§II\-A](https://arxiv.org/html/2609.22117#S2.SS1.p2.1),[5th item](https://arxiv.org/html/2609.22117#S5.I3.i5.p1.1.1),[§V\-C1](https://arxiv.org/html/2609.22117#S5.SS3.SSS1.p1.1)\.
- \[20\]G\. Mai, Y\. Xuan, W\. Zuo, Y\. He, J\. Song, S\. Ermon, K\. Janowicz, and N\. Lao\(2023\)Sphere2Vec: a general\-purpose location representation learning over a spherical surface for large\-scale geospatial predictions\.ISPRS Journal of Photogrammetry and Remote Sensing202,pp\. 439–462\.Cited by:[§II\-A](https://arxiv.org/html/2609.22117#S2.SS1.p2.1)\.
- \[21\]T\. Mikolov\(2013\)Efficient estimation of word representations in vector space\.arXiv preprint arXiv:1301\.3781\.Cited by:[§II\-B](https://arxiv.org/html/2609.22117#S2.SS2.p2.1),[2nd item](https://arxiv.org/html/2609.22117#S5.I2.i2.p1.1.1)\.
- \[22\]R\. Nils and G\. Iryna\(2019\)Sentence\-bert: sentence embeddings using siamese bert\-networks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),pp\. 3982–3992\.Cited by:[§V\-C1](https://arxiv.org/html/2609.22117#S5.SS3.SSS1.p3.1)\.
- \[23\]Office for National Statistics\(2023\)Origin\-destination data, England and Wales: Census 2021\.Note:Released 26 October 2023\. Census 2021 origin\-destination workplace flow data, accessed via Nomis[https://www\.nomisweb\.co\.uk/sources/census\_2021\_od](https://www.nomisweb.co.uk/sources/census_2021_od)Cited by:[§V\-B1](https://arxiv.org/html/2609.22117#S5.SS2.SSS1.p1.1)\.
- \[24\]L\. Pappalardo, E\. Manley, V\. Sekara, and L\. Alessandretti\(2023\)Future directions in human mobility science\.Nature computational science3\(7\),pp\. 588–600\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p1.1)\.
- \[25\]A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.\(2021\)Learning transferable visual models from natural language supervision\.InInternational conference on machine learning,pp\. 8748–8763\.Cited by:[§IV\-C](https://arxiv.org/html/2609.22117#S4.SS3.p2.1)\.
- \[26\]M\. Schneider\(1959\)Gravity models and trip distribution theory\.Papers in Regional Science5\(1\),pp\. 51–56\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p2.1)\.
- \[27\]Q\. Shi, L\. Zhuo, H\. Tao, and J\. Yang\(2024\)A fusion model of temporal graph attention network and machine learning for inferring commuting flow from human activity intensity dynamics\.International Journal of Applied Earth Observation and Geoinformation126,pp\. 103610\.Cited by:[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p3.1)\.
- \[28\]F\. Simini, G\. Barlacchi, M\. Luca, and L\. Pappalardo\(2021\)A deep gravity model for mobility flows generation\.Nature communications12\(1\),pp\. 6576\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p3.1),[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p1.1),[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p2.1),[§II](https://arxiv.org/html/2609.22117#S2.p2.1),[§IV\-F2](https://arxiv.org/html/2609.22117#S4.SS6.SSS2.p2.1),[§VII\-B](https://arxiv.org/html/2609.22117#S7.SS2.p3.1)\.
- \[29\]F\. Simini, M\. C\. González, A\. Maritan, and A\. Barabási\(2012\)A universal model for mobility and migration patterns\.Nature484\(7392\),pp\. 96–100\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p6.1),[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p2.1)\.
- \[30\]S\. A\. Stouffer\(1940\)Intervening opportunities: a theory relating mobility and distance\.American sociological review5\(6\),pp\. 845–867\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p2.1),[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p2.1)\.
- \[31\]Y\. Tu, P\. Wang, J\. N\. Zhu, Z\. Zhao, J\. Li, and S\. Wu\(2026\)GAGNN: a geography\-aware graph neural network for citywide commuting flows prediction\.International Journal of Applied Earth Observation and Geoinformation147,pp\. 105175\.Cited by:[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p3.1)\.
- \[32\]H\. Wan, Y\. Lin, S\. Guo, and Y\. Lin\(2021\)Pre\-training time\-aware location embeddings from spatial\-temporal trajectories\.IEEE Transactions on Knowledge and Data Engineering34\(11\),pp\. 5510–5523\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.22117#S2.SS2.p2.1),[5th item](https://arxiv.org/html/2609.22117#S5.I2.i5.p1.1.1)\.
- \[33\]X\. Wang, T\. Cheng, S\. Law, Z\. Zeng, I\. Ilyankou, J\. Liu, L\. Yin, W\. Huang, and N\. Jongwiriyanurak\(2025\)Into the unknown: applying inductive spatial\-semantic location embeddings for predicting individuals’ mobility beyond visited places\.InProceedings of the 33rd ACM International Conference on Advances in Geographic Information Systems,pp\. 1046–1055\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p8.1),[§II\-B](https://arxiv.org/html/2609.22117#S2.SS2.p3.1),[§IV\-F1](https://arxiv.org/html/2609.22117#S4.SS6.SSS1.p1.1),[6th item](https://arxiv.org/html/2609.22117#S5.I2.i6.p1.1)\.
- \[34\]X\. Wang, T\. Cheng, S\. Law, Z\. Zeng, L\. Yin, and J\. Liu\(2025\)Multi\-modal contrastive learning of urban space representations from poi data\.Computers, Environment and Urban Systems120,pp\. 102299\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p8.1),[§II\-A](https://arxiv.org/html/2609.22117#S2.SS1.p2.1),[§II\-A](https://arxiv.org/html/2609.22117#S2.SS1.p3.1),[§IV\-B](https://arxiv.org/html/2609.22117#S4.SS2.p1.1),[6th item](https://arxiv.org/html/2609.22117#S5.I3.i6.p1.1.1)\.
- \[35\]Z\. Wang, H\. Li, and R\. Rajagopal\(2020\)Urban2vec: incorporating street view imagery and pois for multi\-modal urban neighborhood embedding\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.34,pp\. 1013–1020\.Cited by:[4th item](https://arxiv.org/html/2609.22117#S5.I3.i4.p1.1.1)\.
- \[36\]Y\. Xu, S\. Gao, Q\. Huang, A\. Göçmen, Q\. Zhu, and F\. Zhang\(2025\)Predicting human mobility flows in cities using deep learning on satellite imagery\.Nature Communications16\(1\),pp\. 10372\.Cited by:[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p3.1),[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p4.1)\.
- \[37\]D\. Yang, D\. Zhang, V\. W\. Zheng, and Z\. Yu\(2014\)Modeling user activity preference by leveraging user spatial temporal characteristics in lbsns\.IEEE Transactions on Systems, Man, and Cybernetics: Systems45\(1\),pp\. 129–142\.Cited by:[§V\-A1](https://arxiv.org/html/2609.22117#S5.SS1.SSS1.p1.1)\.
- \[38\]G\. Yin, Z\. Huang, Y\. Bao, H\. Wang, L\. Li, X\. Ma, and Y\. Zhang\(2023\)ConvGCN\-rf: a hybrid learning model for commuting flow prediction considering geographical semantics and neighborhood effects\.GeoInformatica27\(2\),pp\. 137–157\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p3.1),[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p3.1),[§V\-B3](https://arxiv.org/html/2609.22117#S5.SS2.SSS3.p4.1)\.
- \[39\]M\. Zhao, D\. Zhang, Z\. Shi, and C\. Wu\(2026\)Theory\-informed and interpretable graph learning for urban commuting flows\.Sustainable Cities and Society,pp\. 107575\.Cited by:[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p2.1)\.
- \[40\]S\. Zhao, T\. Zhao, I\. King, and M\. R\. Lyu\(2017\)Geo\-teaser: geo\-temporal sequential embedding rank for point\-of\-interest recommendation\.InProceedings of the 26th international conference on world wide web companion,pp\. 153–162\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.22117#S2.SS2.p2.1),[4th item](https://arxiv.org/html/2609.22117#S5.I2.i4.p1.1.1)\.
- \[41\]Y\. Zheng, X\. Xie, W\. Ma,et al\.\(2010\)GeoLife: a collaborative social networking service among user, location and trajectory\.\.IEEE Data Eng\. Bull\.33\(2\),pp\. 32–39\.Cited by:[§V\-A1](https://arxiv.org/html/2609.22117#S5.SS1.SSS1.p1.1)\.
- \[42\]Y\. Zhou and Y\. Huang\(2018\)Deepmove: learning place representations through large scale movement data\.In2018 IEEE international conference on big data \(big data\),pp\. 2403–2412\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p3.1)\.
- \[43\]G\. K\. Zipf\(1946\)The p 1 p 2/d hypothesis: on the intercity movement of persons\.American sociological review11\(6\),pp\. 677–686\.Cited by:[§I](https://arxiv.org/html/2609.22117#S1.p2.1),[§I](https://arxiv.org/html/2609.22117#S1.p6.1),[§II\-C](https://arxiv.org/html/2609.22117#S2.SS3.p2.1),[§VII\-B](https://arxiv.org/html/2609.22117#S7.SS2.p3.1)\.相似文章
MobiDiff:面向人类移动数据生成的多通道语义感知离散扩散框架
介绍MobiDiff,一种端到端的离散扩散框架,通过对多通道语义骨架进行去噪来生成人类移动数据,在实际数据集上实现了更快的推理速度和具有竞争力的保真度。
通过二次嵌入实现位置感知的语言模型
本文提出了一种轻量级、模型无关的方法,通过利用位置数据增强嵌入来提升语言模型的地理空间感知能力,从而在保持标准NLP性能的同时改善空间对齐。
地球嵌入中有什么?位置编码器的可解释性分析
本文介绍了将地理隐式神经表示中的位置嵌入分解为人类可解释特征的方法,例如稀疏潜在概念、自然语言概念和视觉特征,揭示了森林和城市区域等地理结构。
LLM驱动的免训练位置-属性协同融合:一种用于双源加密POIs和LULC制图的闭环范式
本文提出一种LLM驱动的免训练闭环范式,用于融合双源加密POIs,改进LULC制图中的位置对齐和属性匹配,性能优于现有方法。
TrajGenAgent:一种用于人类移动轨迹生成的分层LLM智能体
TrajGenAgent提出了一种分层LLM智能体框架,将宏观活动规划与微观时空实例化解耦,用于无需微调即可生成逼真的人类移动轨迹。它还引入了一种基于异常检测的评估方法,用于行为保真度。