Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment
Summary
This paper evaluates zero-shot large language models for predicting neighborhood-level mobility patterns across U.S. metropolitan areas, comparing them to supervised baselines and auditing their structural alignment with empirical data.
View Cached Full Text
Cached at: 09/02/26, 06:11 AM
# Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment Source: [https://arxiv.org/html/2609.00345](https://arxiv.org/html/2609.00345) DOI:[XXXXXXX\.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)Conference:Make sure to enter the correct conference title from your rights confirmation email; June 03–05, 2018; Woodstock, NYISBN:978\-1\-4503\-XXXX\-X/2018/06CCS:Information systems Spatial\-temporal systemsCCS:Computing methodologies Supervised learningCCS:Information systems Data mining,Eesha Kurellaemail:[ekurella@umd\.edu](mailto:[email protected])Affiliation:University of Maryland,USA,Arnav Dadarya[https://orcid.org/0000-0001-6036-076X](https://orcid.org/0000-0001-6036-076X)email:[arnav1@terpmail\.umd\.edu](mailto:[email protected])Affiliation:University of Maryland,USA,Naman Awasthi[https://orcid.org/0000-0001-6036-076X](https://orcid.org/0000-0001-6036-076X)email:[nawasthi@terpmail\.umd\.edu](mailto:[email protected])Affiliation:University of Maryland,USA,Kazi Tasnim Zinat[https://orcid.org/0000-0001-6036-076X](https://orcid.org/0000-0001-6036-076X)email:[kzintas@terpmail\.umd\.edu](mailto:[email protected])Affiliation:University of Maryland,USAandVanessa Frias\-Martinez[https://orcid.org/0000-0001-5114-7633](https://orcid.org/0000-0001-5114-7633)email:[vfrias@umd\.edu](mailto:[email protected])Affiliation:University of Maryland,USA 2018 ###### Abstract\. Human mobility is central to urban planning, transportation management, public health, and emergency response, yet the fine\-grained trajectory data needed to model movement are often proprietary, access\-restricted, and privacy\-sensitive, motivating the search for alternatives\. Recent work suggests that large language models \(LLMs\) offer one such alternative, generating plausible mobility traces and predicting individual movement\. However, little is known about whether they can infer aggregate mobility patterns across neighborhoods, which reveal collective dynamics that can directly inform public health response, transportation modeling, and emergency planning\. Beyond inference, whether these predictions capture empirically meaningful context\-mobility relationships also remains open\. We address this gap by evaluating zero\-shot LLMs on Census Block Group\-level mobility prediction across four U\.S\. metropolitan areas, using anonymized Cuebiq mobility traces to construct point\-level, trajectory\-level, and temporal mobility outcomes and pairing them with sociodemographic and built\-environment predictors\. We compare LLM predictions against supervised baselines and introduce a directional alignment analysis that tests whether LLM\-implied predictor effects agree with empirical OLS and Jonckheere–Terpstra trends\. Results show that neighborhood context contains substantial predictive signal, with supervised baselines reaching 0\.580 average accuracy compared with 0\.435 for the best LLM, while spatial extent outcomes are the most predictable but also exhibit the largest LLM–baseline gaps\. Directional analysis further shows that LLMs rely on coarse, stable predictor\-level priors that often remain invariant across outcomes and cities, including asymmetric treatment of protected\-group predictors\. These findings suggest that LLMs can partially recover aggregate mobility patterns from urban context, but their predictions should not be treated as structurally grounded without auditing their empirical alignment and potential bias\. ###### Keywords: Large language models, human mobility, neighborhood\-level mobility, urban computing, geospatial AI, directional alignment, mobility prediction, algorithmic auditing ## 1\.Introduction Human mobility is a fundamental dimension of urban life, reflecting how people access opportunities, navigate infrastructure, and experience cities in everyday life\([Barbosa et al\., 2018](https://arxiv.org/html/2609.00345#bib.bib3);[Hägerstrand, 1970](https://arxiv.org/html/2609.00345#bib.bib4);[Yang et al\., 2023](https://arxiv.org/html/2609.00345#bib.bib9);[Du et al\., 2025](https://arxiv.org/html/2609.00345#bib.bib5)\)\. Because movement patterns shape access to jobs and services, pandemic spread, environmental burdens, and broader forms of social interaction, mobility data have become central to understanding cities and informing urban policy\([Abrar et al\., 2023](https://arxiv.org/html/2609.00345#bib.bib38);[Buckee et al\., 2020](https://arxiv.org/html/2609.00345#bib.bib8);[Chang et al\., 2021](https://arxiv.org/html/2609.00345#bib.bib2);[Yabe et al\., 2025](https://arxiv.org/html/2609.00345#bib.bib13);[Xu et al\., 2025](https://arxiv.org/html/2609.00345#bib.bib6)\)\. The growing availability of passively collected digital traces, especially smartphone\- and mobile\-phone\-based location data, has substantially expanded this potential\. Compared with traditional travel surveys\([Bricka et al\., 2024](https://arxiv.org/html/2609.00345#bib.bib12)\)and other coarse data sources, such data provide much finer spatial and temporal resolution, often capturing behavior continuously over longer periods and at much larger scales\([González et al\., 2024](https://arxiv.org/html/2609.00345#bib.bib10);[Garber et al\., 2022](https://arxiv.org/html/2609.00345#bib.bib11)\)\. This has enabled researchers to characterize everyday activity patterns, travel extent, accessibility, and behavioral heterogeneity with unprecedented detail, greatly enlarging the empirical toolkit available to urban science and policy research\. Yet the smartphone\-based location data that enable these analyses are largely proprietary, access\-restricted, and subject to privacy constraints that limit their availability to many researchers and public agencies, raising the question of whether models that encode broad knowledge about cities and human behavior could serve as useful alternatives in data\-scarce settings\. Large language models \(LLMs\) are a natural candidate for this role\. Trained on vast corpora encompassing urban descriptions, census and demographic data, transportation research, and accounts of everyday life across cities, LLMs may encode the kinds of contextual knowledge that underlie mobility patterns\. Recent studies suggest they do\. Prior work has used LLMs for tasks such as next\-location prediction and zero\-shot mobility inference, showing that LLMs can sometimes act as mobility predictors even without task\-specific training\([Wang et al\., 2023](https://arxiv.org/html/2609.00345#bib.bib16);[Beneduce et al\., 2025](https://arxiv.org/html/2609.00345#bib.bib14);[Ma et al\., 2025](https://arxiv.org/html/2609.00345#bib.bib17)\)\. Other work has explored whether prompted LLMs can generate synthetic travel\-survey responses for urban mobility assessment, producing plausible diary\-like behavior from background knowledge alone\([Bhandari et al\., 2024](https://arxiv.org/html/2609.00345#bib.bib15)\)\. More broadly, benchmarks such as CityBench evaluate LLMs across a range of urban tasks and show that, while advanced models can be competitive on tasks grounded in commonsense and semantic understanding, they remain less reliable on tasks requiring stronger urban reasoning\([Feng et al\., 2025](https://arxiv.org/html/2609.00345#bib.bib7)\)\. However, all of these studies share a common target of evaluation, assessing whether LLMs can produce plausible outputs about individual\-level mobility, such as a person’s next trip, a synthetic travel diary, or a single device’s behavioral signature\. This individual\-level focus is valuable in its own right, but it does not fully address the forecasting needs of many real\-world systems, where the key quantity of interest is aggregate mobility patterns\. Transportation agencies, urban planners, emergency managers, and public health officials typically require predictions of origin\-destination flows, demand surges, and population redistribution patterns rather than individual trajectories\([Zhao et al\., 2024](https://arxiv.org/html/2609.00345#bib.bib37);[Rong et al\., 2024](https://arxiv.org/html/2609.00345#bib.bib36)\)\. Although one possible approach is to simulate many individual LLM\-based agents and aggregate their generated trajectories, this strategy can be computationally expensive\. We instead study whether LLMs can directly infer aggregate mobility outcomes from contextual urban inputs\. This framing positions LLM\-based neighborhood\-level mobility prediction as a distinct and underexplored problem, asking whether models that have shown promise on individual mobility tasks can also recover collective mobility patterns across neighborhoods from sociodemographic and built\-environment context\. In addition, even if LLMs can recover aggregate neighborhood\-level outcomes, prediction accuracy alone does not establish that they have learned the empirical relationships that structure mobility across places\. Plausible performance may instead reflect shallow heuristics, broad stereotypes\([Manvi et al\., 2024](https://arxiv.org/html/2609.00345#bib.bib23);[Moayeri et al\., 2024](https://arxiv.org/html/2609.00345#bib.bib24)\), or memorization\([Hartmann et al\., 2023](https://arxiv.org/html/2609.00345#bib.bib35)\)of patterns seen during pretraining, all of which could align with outcomes on average while diverging from the specific relationships that actually structure how context shapes mobility\. Thus, whether LLM predictions recover the structural relationships between neighborhood context and travel patterns that empirical data reveal remains an open question\. To address these gaps, we evaluate whether LLMs encode useful priors about urban mobility by testing their ability to infer neighborhood\-level mobility outcomes from urban context, assessing the structural alignment of those predictions with empirical context\-mobility relationships, and examining whether this alignment holds across cities\. Our analysis is grounded in large\-scale, real\-world mobility trajectories from the United States, provided by Cuebiq, which amasses anonymized location data from nearly 70 million mobile devices, covering roughly 20% of the U\.S\. population\. We aggregate these data to the Census Block Group \(CBG\) level and derive multiple mobility indicators capturing three complementary dimensions of behavior: \(i\) point\-level, \(ii\) trajectory\-level, and \(iii\) temporal\-level mobility\. Given built\-environment and neighborhood sociodemographic characteristics as input, we ask two research questions: RQ1:To what extent can LLMs predict neighborhood\-level mobility outcomes given built\-environment and sociodemographic context? RQ2:To what extent do LLM predictions align with empirically observed relationships between neighborhood characteristics and mobility behavior? The remainder of the paper is organized as follows\. Section[2](https://arxiv.org/html/2609.00345#S2)reviews prior work on LLMs for mobility generation, prediction, and urban reasoning, as well as recent efforts to evaluate geospatial knowledge and bias in LLMs\. Section[3](https://arxiv.org/html/2609.00345#S3)presents our problem formulation, mobility outcome construction, zero\-shot LLM prediction framework, supervised baselines, and directional alignment methodology\. Section[4](https://arxiv.org/html/2609.00345#S4)describes the datasets, study areas, contextual predictors, model settings, as well as the predictive performance results for RQ1 and the directional alignment results for RQ2\. Finally, Section[5](https://arxiv.org/html/2609.00345#S5)summarizes the main findings and discusses their implications for using LLMs in aggregate urban mobility inference\. ## 2\.Related Work ### 2\.1\.LLMs for Mobility Generation, Prediction, and Urban Tasks A growing body of work uses large language models \(LLMs\) to generate, predict, and simulate mobility and urban phenomena\. In mobility modeling, several studies adapt LLMs to represent individual trajectories, activity chains, and travel behavior\.[Ma et al\. \(2025\)](https://arxiv.org/html/2609.00345#bib.bib17)propose a foundation model for universal human mobility patterns that fuses cross\-domain data and semantically enriches GPS traces with survey\-informed knowledge, achieving robust performance in activity inference and POI classification across settings such as Los Angeles and Egypt\. Wang et al\.\([Wang et al\., 2024](https://arxiv.org/html/2609.00345#bib.bib19)\)introduceLLMob, an LLM\-agent framework that generates personal mobility trajectories by extracting patterns from historical traces and reasoning about daily motivations, producing realistic movement sequences even during abnormal periods such as the pandemic\.[Shao et al\. \(2024\)](https://arxiv.org/html/2609.00345#bib.bib32)similarly explore the use of LLMs for human mobility generation through context\-aware reasoning, emphasizing that language models can explain the intentions behind movement segments rather than merely imitate observed sequences\.\([Li et al\., 2024b](https://arxiv.org/html/2609.00345#bib.bib20)\)proposeGeo\-Llama, a fine\-tuned framework for generating mobility trajectories under explicit spatiotemporal constraints, showing that LLMs can produce coherent synthetic movements while satisfying visit\-level requirements\. Other work focuses on generating richer behavioral records, such as travel diaries and activity schedules\.[Li et al\. \(2024c\)](https://arxiv.org/html/2609.00345#bib.bib18)developMobAgent, which constructs fine\-grained profiles and uses recursive reasoning to generate realistic travel diaries aligned with empirical mobility patterns and road\-network constraints\.[Bhandari et al\. \(2024\)](https://arxiv.org/html/2609.00345#bib.bib15)evaluate whether LLMs can generate synthetic travel survey data resembling the National Household Travel Survey \(NHTS\), finding that fine\-tuned models effectively capture complex activity chains and transition probabilities across major U\.S\. metropolitan areas\. Liu et al\.[Liu et al\. \(2025b\)](https://arxiv.org/html/2609.00345#bib.bib34)present a retrieval\-augmented framework with a feedback loop for generating daily activity chains under limited information, showing that the approach can reproduce coordinated household behaviors such as joint shopping trips or shared meal times\. Extending this direction,[Liu et al\. \(2025a\)](https://arxiv.org/html/2609.00345#bib.bib33)studies discrete travel demand modeling with persona\-conditioned LLMs, demonstrating that behavioral traits and socioeconomic context improve simulation of individual travel choices\. A related stream uses LLMs for mobility prediction and recommendation rather than full\-sequence generation\.[Li et al\. \(2024a\)](https://arxiv.org/html/2609.00345#bib.bib26)proposeLLM4POI, a next\-POI recommendation framework that fine\-tunes LLMs on check\-in data and combines this with key\-query similarity to retain collaborative and contextual information\.[Gong et al\. \(2024\)](https://arxiv.org/html/2609.00345#bib.bib21)introduceMobility\-LLM, a reprogramming framework that uses behavioral prompts and semantic location representations to improve next\-location prediction, arrival\-time estimation, and user\-link prediction\.[Chen et al\. \(2025\)](https://arxiv.org/html/2609.00345#bib.bib27)developQT\-Mob, which introduces semantic location tokenization to encode coordinates as semantically meaningful discrete tokens, substantially improving next\-location prediction and trajectory recovery\. At the urban scale,[Li et al\. \(2024d\)](https://arxiv.org/html/2609.00345#bib.bib25)presentUrbanGPT, which integrates spatiotemporal encoding with instruction tuning to align urban interdependencies with LLM reasoning, demonstrating strong zero\-shot generalization across prediction tasks such as taxi flows, bike flows, and crime rates\. Taken together, this literature shows that LLMs can serve as flexible models for synthetic mobility generation, travel simulation, recommendation, and urban prediction\. It also suggests that these models encode nontrivial mobility\-relevant signals\. There are a few aggregate\-oriented examples:UrbanGPTstudies regional spatiotemporal prediction tasks such as taxi flows and bike flows, and[Bhandari et al\. \(2024\)](https://arxiv.org/html/2609.00345#bib.bib15)evaluate generated surveys using aggregate pattern\-level mobility metrics\. Still, the dominant targets of evaluation remain individual trajectories, travel diaries, next\-location prediction, or task\-specific urban outputs rather than direct recovery of neighborhood\-level mobility indicators from contextual urban features, which we address in RQ1\. ### 2\.2\.Evaluating Geospatial and Mobility Knowledge in LLMs A second line of work asks what kinds of geographic and mobility knowledge LLMs already encode\.[Manvi et al\. \(2023\)](https://arxiv.org/html/2609.00345#bib.bib22)introduceGeoLLM, which fine\-tunes LLMs on reverse\-geocoded map data and shows that the resulting models can accurately predict geospatial socioeconomic indicators such as population density and asset wealth\.[Luo et al\. \(2024\)](https://arxiv.org/html/2609.00345#bib.bib30)proposeTSI\-LLM, a framework for trajectory semantic inference that uses context\-rich prompting and chain\-of\-thought reasoning to infer occupation categories, activity sequences, and detailed textual descriptions from raw movement traces\.[Asano et al\. \(2025\)](https://arxiv.org/html/2609.00345#bib.bib31)introduceMobQA, a benchmark of 5,800 question\-answer pairs for evaluating LLMs on mobility\-related factual retrieval, semantic inference, and interpretive explanation\. Their results suggest that while models perform well on straightforward factual extraction, they struggle with more complex semantic reasoning over long GPS trajectories\. This evaluation\-oriented literature also highlights important limitations and biases in LLMs’ geographic knowledge\.[Manvi et al\. \(2024\)](https://arxiv.org/html/2609.00345#bib.bib23)show that although LLMs can make reasonably accurate zero\-shot geospatial predictions for objective topics, they exhibit systematic geographic bias on subjective judgments, rating wealthier regions more favorably on attributes such as intelligence or work ethic\.[Wu and Wang \(2024\)](https://arxiv.org/html/2609.00345#bib.bib29)probe demographic and geographic biases in LLM\-predicted POI visits, finding that models reflect and amplify race and gender stereotypes in the types of places they associate with different groups\.[Moayeri et al\. \(2024\)](https://arxiv.org/html/2609.00345#bib.bib24)further demonstrates broad disparities in factual recall throughWorldBench, showing that state\-of\-the\-art LLMs perform substantially worse for countries in low\-income and non\-Western regions than for Western countries\. Recent work has also begun evaluating LLMs as broader models of urban knowledge rather than only as task\-specific predictors\.[Zhang et al\. \(2025\)](https://arxiv.org/html/2609.00345#bib.bib28)introduceAI4US, which tests whether LLMs can reproduce foundational urban\-science relationships such as scaling laws, distance decay, and urban vitality\. They find that LLMs often recover broad aggregate patterns with high fidelity, but also oversimplify urban complexity, showing limited diversity and weaker causal depth than real urban systems\. Our RQ2 work is most closely related to this evaluation\-oriented literature and extends it in a new direction\. Rather than examining factual geospatial recall, mobility question answering, or broad urban\-theory reproduction, we probe whether LLMs encode empirically meaningful priors about neighborhood\-level mobility behavior from built\-environment and sociodemographic context\. Figure 1\.Overview of the evaluation pipeline for assessing LLM priors about neighborhood\-level mobility\. ## 3\.Method ### 3\.1\.Problem Formulation Letiiindex Census Block Groups \(CBGs\), let𝐱i∈ℝp\\mathbf\{x\}\_\{i\}\\in\\mathbb\{R\}^\{p\}denote the vector of contextual predictors for CBGii, and letm∈ℳm\\in\\mathcal\{M\}denote a mobility outcome\. For each outcomemm, letzi\(m\)∈ℝz\_\{i\}^\{\(m\)\}\\in\\mathbb\{R\}denote the continuous CBG\-level mobility value derived from aggregated individual mobility traces\. We formulate CBG\-level mobility outcome prediction as an outcome\-specific three\-class classification task in which the goal is to infer a discretized mobility label \(1\)yi\(m\)∈\{low,neutral,high\}y\_\{i\}^\{\(m\)\}\\in\\\{\\texttt\{low\},\\texttt\{neutral\},\\texttt\{high\}\\\}from neighborhood context𝐱i\\mathbf\{x\}\_\{i\}\. More formally, for each outcomemm, we evaluate a predictor \(2\)F\(m\):𝐱i↦yi\(m\)\.F^\{\(m\)\}:\\mathbf\{x\}\_\{i\}\\mapsto y\_\{i\}^\{\(m\)\}\. Here,F\(m\)F^\{\(m\)\}may represent either an LLM\-based zero\-shot predictor, which reasons over textual descriptions of neighborhood context and outcome semantics, or a supervised baseline trained directly on labeled CBG examples\. This formulation allows us to test whether socio\-demographic and built\-environment context alone is sufficient to recover structured variation in aggregate mobility behavior across neighborhoods\. Our analysis considers two complementary evaluation objectives\. The first ispredictive performance \(RQ1\): whether a model can correctly classify the mobility level of a CBG for a given outcomemm\. The second isdirectional alignment \(RQ2\): whether the model captures empirically meaningful relationships between contextual predictors and mobility outcomes\. Specifically, for contextual predictorxijx\_\{ij\}and mobility outcomemm, we assess whether the directional implication of the model’s reasoning agrees with the empirical direction estimated from observed data\. This second objective is important because accurate class prediction alone does not establish that a model has recovered the underlying monotonic structure relating urban context to mobility behavior\. The remainder of this section presents the methodological components required to instantiate this formulation\. Figure[1](https://arxiv.org/html/2609.00345#S2.F1)provides an overview of the evaluation pipeline for assessing LLM priors about neighborhood\-level mobility\. We first describe how CBG\-level mobility outcomes are constructed from individual mobility traces, then introduce the prediction framework for both LLM\-based and supervised models, and finally present the directional alignment framework used to compare model\-implied relationships with empirical patterns in the data\. ### 3\.2\.Mobility Outcome Construction In this paper, we construct CBG\-level mobility outcomes from anonymized individual mobility traces\. The goal is to derive aggregate measures that characterize different dimensions of resident mobility behavior and that can be paired with socio\-demographic and built\-environment predictors at the same spatial scale\. #### 3\.2\.1\.From Individual Mobility Data to CBG\-Level Aggregates The mobility data used in this study consist of anonymized individual mobility traces collected over 2021\. From these traces, we first derive user\-level mobility measures described in[3\.2\.2](https://arxiv.org/html/2609.00345#S3.SS2.SSS2), that summarize annual movement behavior\. These measures are computed from observed stops and travel episodes and capture complementary aspects of mobility, including spatial extent, movement structure, and temporal variability\. Because our prediction task is defined at the CBG level, we aggregate the user\-level mobility measures to the CBG scale\. For each user, we identify the associated home CBG and assign the user\-level mobility measures to that area\. Then, for each CBG, we summarize the mobility behavior of all associated users to obtain aggregate mobility outcomes representing the typical mobility profile of residents in that CBG\. This aggregation serves two purposes\. First, it aligns the mobility outcomes with the same spatial unit used for the contextual predictors in the classification task\. Second, it preserves privacy by ensuring that the prediction targets are aggregate behavioral summaries rather than individual trajectories\. To improve reliability, we retain only CBGs with sufficient user support in the mobility data\. The resulting dataset therefore, consists of robust CBG\-level mobility outcomes that serve as the ground\-truth targets throughout the paper\. #### 3\.2\.2\.Mobility Outcome Families and Definitions FamilyOutcomeWhat it measuresInterpretation of higher valuesPoint\-levelStay\-points entropyEvenness of dwell time across visited locationsActivity time is distributed more evenly across placesRadius of gyrationDispersion of stops around the center of activityLarger spatial spread of routine activity locationsConvex hull diameterMaximum distance across the activity spaceWider geographic extent of observed mobilityEllipse areaSize of elliptical approximation of activity spaceBroader spatial footprint of activity locationsTrajectory\-levelTotal travel lengthCumulative distance traveled across observed tripsMore extensive travel over the observation periodTravel entropyDiversity of route or trip signaturesMore varied routing or travel behaviorAverage durationMean duration of observed travel segmentsLonger average tripsTemporal\-levelDaily temporal fragmentationWithin\-day variability in dwell durationsGreater irregularity in daily activity timingTable 1\.Summary of mobility outcome families and metrics\. All metrics are computed at the user\-year level from annual stop and travel\-segment sequences and then aggregated to the Census Block Group \(CBG\) level\.Following[Wu et al\. \(2019\)](https://arxiv.org/html/2609.00345#bib.bib1)taxonomy of mobility characterization from individual trajectories, we organize the mobility outcomes into three families:point\-level,trajectory\-level, andtemporal\-leveloutcomes\. Point\-level outcomes summarize the spatial footprint and organization of visited locations; trajectory\-level outcomes characterize movement over observed travel sequences; and temporal\-level outcomes quantify the temporal regularity of activity patterns\. For each user, we represent annual mobility as an ordered sequence of stops and travel segments, \(3\)s1→trv1s2→trv2⋯→trvN−1sN,s\_\{1\}\\xrightarrow\{\\mathrm\{trv\}\_\{1\}\}s\_\{2\}\\xrightarrow\{\\mathrm\{trv\}\_\{2\}\}\\cdots\\xrightarrow\{\\mathrm\{trv\}\_\{N\-1\}\}s\_\{N\},where each stopsis\_\{i\}is associated with a location𝐱i∈ℝ2\\mathbf\{x\}\_\{i\}\\in\\mathbb\{R\}^\{2\}, a dwell timeτi\\tau\_\{i\}, and an observation daydid\_\{i\}, and each travel segmenttrvi\\mathrm\{trv\}\_\{i\}connectingsis\_\{i\}andsi\+1s\_\{i\+1\}is associated with a travel lengthLiL\_\{i\}and a travel durationaia\_\{i\}\. Unless otherwise noted, all outcomes are first computed at the user\-year level and then aggregated to the Census Block Group \(CBG\) level\. We further define normalized dwell\-time weights as \(4\)wi=τi∑k=1Nτk\.w\_\{i\}=\\frac\{\\tau\_\{i\}\}\{\\sum\_\{k=1\}^\{N\}\\tau\_\{k\}\}\. ##### \(1\) Point\-level outcomes\. Point\-level outcomes quantify the extent, dispersion, and spatial organization of visited locations\. Let a user have stops indexed byi=1,…,ni=1,\\dots,n, where stopiihas coordinates𝐬i\\mathbf\{s\}\_\{i\}and dwell timedid\_\{i\}\. Define normalized dwell\-time weights as \(5\)wi=di∑k=1ndk\.w\_\{i\}=\\frac\{d\_\{i\}\}\{\\sum\_\{k=1\}^\{n\}d\_\{k\}\}\. Stay\-points entropymeasures how evenly dwell time is distributed across distinct visited locations\. Letℒ\\mathcal\{L\}denote the set of unique stay locations, and letplp\_\{l\}be the proportion of total dwell time spent at locationl∈ℒl\\in\\mathcal\{L\}\. Then \(6\)Hstay=−∑l∈ℒpllog2pl\.H\_\{\\text\{stay\}\}=\-\\sum\_\{l\\in\\mathcal\{L\}\}p\_\{l\}\\log\_\{2\}p\_\{l\}\.Higher values indicate that activity time is distributed more evenly across locations, whereas lower values indicate concentration in a small number of places\. Radius of gyrationmeasures the characteristic distance of visited locations from the user’s center of activity\. Let \(7\)𝐱¯=∑i=1Nwi𝐱i\\bar\{\\mathbf\{x\}\}=\\sum\_\{i=1\}^\{N\}w\_\{i\}\\mathbf\{x\}\_\{i\}denote the weighted centroid, and letδ\(𝐱i,𝐱¯\)\\delta\(\\mathbf\{x\}\_\{i\},\\bar\{\\mathbf\{x\}\}\)denote the haversine distance between stopiiand the centroid\. The radius of gyration is \(8\)rg=∑i=1Nwiδ\(𝐱i,𝐱¯\)2\.r\_\{g\}=\\sqrt\{\\sum\_\{i=1\}^\{N\}w\_\{i\}\\,\\delta\(\\mathbf\{x\}\_\{i\},\\bar\{\\mathbf\{x\}\}\)^\{2\}\}\.Larger values indicate more spatially dispersed activity patterns\. Convex hull diameterquantifies the maximum spatial extent of the activity space\. Letℋ\\mathcal\{H\}denote the set of vertices of the convex hull formed by the observed stop locations\. Then, \(9\)Dhull=max𝐮,𝐯∈ℋd\(𝐮,𝐯\),D\_\{\\text\{hull\}\}=\\max\_\{\\mathbf\{u\},\\mathbf\{v\}\\in\\mathcal\{H\}\}d\(\\mathbf\{u\},\\mathbf\{v\}\),whered\(⋅,⋅\)d\(\\cdot,\\cdot\)is the haversine distance\. This outcome represents the largest distance between any two locations in a user’s activity space\. Ellipse areameasures the size of an elliptical approximation of the activity space\. Let \(10\)𝐂=∑i=1nwi\(𝐬i−𝐬¯\)\(𝐬i−𝐬¯\)⊤\\mathbf\{C\}=\\sum\_\{i=1\}^\{n\}w\_\{i\}\(\\mathbf\{s\}\_\{i\}\-\\bar\{\\mathbf\{s\}\}\)\(\\mathbf\{s\}\_\{i\}\-\\bar\{\\mathbf\{s\}\}\)^\{\\top\}be the weighted covariance matrix of stop coordinates, and letλ1≥λ2\\lambda\_\{1\}\\geq\\lambda\_\{2\}be its eigenvalues\. The ellipse area is \(11\)Aellipse=πλ1λ2c2,A\_\{\\text\{ellipse\}\}=\\pi\\sqrt\{\\lambda\_\{1\}\}\\sqrt\{\\lambda\_\{2\}\}\\,c^\{2\},whereccis a scale factor\. Larger values indicate a broader activity space\. ##### \(2\) Trajectory\-level outcomes\. Trajectory\-level outcomes characterize mobility over sequences of trips and capture the extent and diversity of travel behavior\. Total travel lengthmeasures the cumulative distance traveled over all observed trips\. Let trajectory segments be indexed byt=1,…,Tt=1,\\dots,T, and letLtL\_\{t\}denote the length of segmenttt\. Then, \(12\)Ltotal=∑t=1TLt\.L\_\{\\text\{total\}\}=\\sum\_\{t=1\}^\{T\}L\_\{t\}\.Higher values indicate more extensive travel over the observation period\. Travel entropymeasures the diversity of routes used by a user\. Letℛ\\mathcal\{R\}denote the set of unique route signatures, and letprp\_\{r\}be the proportion of trips corresponding to router∈ℛr\\in\\mathcal\{R\}\. Then, \(13\)Htravel=−∑r∈ℛprlog2pr\.H\_\{\\text\{travel\}\}=\-\\sum\_\{r\\in\\mathcal\{R\}\}p\_\{r\}\\log\_\{2\}p\_\{r\}\.Higher values indicate more varied routing behavior, while lower values indicate repetitive travel patterns\. Average durationmeasures the mean duration of observed trips\. Letaia\_\{i\}denote the duration of travel segmenttrvi\\mathrm\{trv\}\_\{i\},i=1,…,N−1i=1,\\dots,N\-1\. Then \(14\)a¯=1N−1∑i=1N−1ai\.\\bar\{a\}=\\frac\{1\}\{N\-1\}\\sum\_\{i=1\}^\{N\-1\}a\_\{i\}\.Higher values indicate longer average trip durations over the observation period\. ##### \(3\) Temporal\-level outcomes\. Temporal\-level outcomes characterize the temporal organization and irregularity of mobility behavior\. Daily temporal fragmentationmeasures within\-day variability in dwell durations\. For each observed dayd∈𝒟d\\in\\mathcal\{D\}, let \(15\)σd2=Var\{τi:di=d\}\\sigma\_\{d\}^\{2\}=\\mathrm\{Var\}\\\{\\tau\_\{i\}:d\_\{i\}=d\\\}denote the variance of dwell times across all stops observed on that day\. We define daily temporal fragmentation as \(16\)Fdaily=1\|𝒟\|∑d∈𝒟σd2\.F\_\{\\text\{daily\}\}=\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{d\\in\\mathcal\{D\}\}\\sigma\_\{d\}^\{2\}\.Higher values indicate greater within\-day temporal irregularity, reflected in a wider mix of short and long stays across stops within observed days\. These three types of mobility outcomes provide complementary views of aggregate mobility behavior at the CBG level\. In the remainder of the paper, we use this organization to structure the prediction tasks, the directional alignment analysis, and the presentation of results\. ### 3\.3\.Mobility Outcome Prediction Framework We evaluate two classes of predictors for CBG\-level mobility outcome prediction:LLM\-based zero\-shot predictorsandsupervised baseline models\. Both use the same contextual predictors for each CBG, but they differ in how the mapping from neighborhood context to mobility outcomes is obtained\. The LLM setting evaluates zero\-shot contextual inference from textual descriptions, whereas the supervised baselines learn the mapping directly from labeled examples\. Recall from Section[3\.2\.2](https://arxiv.org/html/2609.00345#S3.SS2.SSS2)that each CBGiiis represented by a contextual predictor vector𝐱i∈ℝp\\mathbf\{x\}\_\{i\}\\in\\mathbb\{R\}^\{p\}, and that for each mobility outcomem∈ℳm\\in\\mathcal\{M\}, the prediction target is the discretized class labelyi\(m\)∈\{low,neutral,high\}y\_\{i\}^\{\(m\)\}\\in\\\{\\texttt\{low\},\\texttt\{neutral\},\\texttt\{high\}\\\}\. The prediction task is therefore outcome\-specific: for eachmm, we construct a predictorF\(m\)F^\{\(m\)\}that maps neighborhood context𝐱i\\mathbf\{x\}\_\{i\}to a three\-class mobility label\. #### 3\.3\.1\.LLM\-Based Prediction For each CBGiiand mobility outcomemm, we construct an outcome\-specific promptP\(i,m\)P\(i,m\)that describes the local contextual characteristics of the CBG and asks the model to predict the corresponding mobility class \(0\-shot\)\. Each prompt contains four components \(see Figure[2](https://arxiv.org/html/2609.00345#S3.F2)\): \(1\) the Core\-Based Statistical Area \(CBSA\) in which the CBG is located; \(2\) a natural\-language definition of the target mobility outcomemm; \(3\) the CBSA\-specific tertile thresholds used to discretize the continuous outcome intolow,neutral, andhighclasses; and \(4\) a textual profile of the CBG’s contextual predictors\. The associated system prompt is provided in the Appendix, Figure[6](https://arxiv.org/html/2609.00345#A1.F6)\. User Prompt Template``` You are given Census Block Group (CBG) characteristics for CBSA: [CBSA name]. ## Outcome Description Mobility outcome: [outcome name] Definition: [outcome definition] Tertile cutpoints within [CBSA name]: low <= [c1]; neutral ([c1], [c2]]; high > [c2]. ## Predictor Descriptor - Percentiles are computed within the CBSA. - Predictor buckets correspond to CBSA-wide quintiles: very_low | low | neutral | high | very_high. - Outcome labels are tertiles: low | neutral | high. ## CBG Features - [feature 1]; [raw value] (CBSA percentile Pxx, [bucket]) - [feature 2]; [raw value] (CBSA percentile Pxx, [bucket]) - ... - [feature p]; [raw value] (CBSA percentile Pxx, [bucket]) ## Question Predict the mobility outcome category for this CBG. A. low B. neutral C. high ``` Figure 2\.User prompt for RQ1\.The contextual predictors are expressed in a percentile\-based format designed to convey both absolute and relative neighborhood context\. For each predictorxijx\_\{ij\}, the prompt includes its raw value, its within\-CBSA percentile rankrijr\_\{ij\}, and a quintile\-based categorical bucket bij∈\{very\_low,low,neutral,high,very\_high\}\.b\_\{ij\}\\in\\\{\\texttt\{very\\\_low\},\\texttt\{low\},\\texttt\{neutral\},\\texttt\{high\},\\texttt\{very\\\_high\}\\\}\.Thus, each CBG is represented as a structured textual profile summarizing how its socio\-demographic and built\-environment characteristics compare with those of other CBGs in the same metropolitan area\. This prompt design encourages the model to reason over relative urban context rather than relying only on raw feature magnitudes\. The task is outcome\-specific\. For each mobility outcomemm, a separate prompt is constructed using the outcome definition and the corresponding CBSA\-specific tertile thresholds, while the contextual predictor profile of the CBG remains fixed\. The model is then asked to output one of three class labels,\{low,neutral,high\}\\\{\\texttt\{low\},\\texttt\{neutral\},\\texttt\{high\}\\\}\. In addition to the class prediction, the model is instructed to provide brief reasoning and a ranked assessment of influential predictors\. For comparison, we also train a set of supervised baseline models on the same contextual predictors and outcome labels\. Specifically, for each mobility outcomemm, we fit the following multiclass classifiers:logit\_multinomial,random\_forest,decision\_tree,hist\_gradient\_boosting, andxgboost\. These trained baselines provide data\-driven reference points for the classification task and allow us to compare zero\-shot contextual inference by LLMs against conventional supervised learning approaches trained directly on labeled CBG examples\. #### 3\.3\.2\.Evaluation Setup Prompts are generated for all available CBGs and mobility outcomes\. For predictive evaluation, however, we report LLM performance on the same held\-out20%20\\%test split used for the supervised baselines, so that both model classes are evaluated on an identical subset of CBGs allowing a direct comparison between zero\-shot LLM\-based contextual inference and conventional supervised prediction for CBG\-level mobility outcome classification\. ### 3\.4\.Directional Alignment Framework Since classification accuracy alone does not indicate whether a model captures the empirical structure relating neighborhood context to mobility behavior, we also examine whether correct predictions may still be based on misleading or weakly grounded reasoning\. To address this, we introduce a complementary*directional alignment*analysis that evaluates whether the directional implications of model reasoning \(LLM\) are consistent with empirical relationships observed in the data \(baseline models\)\. Letxijx\_\{ij\}denote contextual predictorjjfor CBGii, and letzi\(m\)z\_\{i\}^\{\(m\)\}denote the continuous value of mobility outcomemm\. For each CBSAcc, mobility outcomemm, and contextual predictorjj, we estimate a ground truth empirical directional relationship betweenxijx\_\{ij\}andzi\(m\)z\_\{i\}^\{\(m\)\}as follows\. We fit an ordinary least squares \(OLS\) regression in which the continuous mobility outcome is regressed on the contextual predictors\. The sign of the standardized coefficient for predictorjj, denotedβcjmOLS\\beta\_\{cjm\}^\{\\text\{OLS\}\}, provides a parametric estimate of direction: \(17\)dircjmOLS=sign\(βcjmOLS\)∈\{−1,\+1\}\.\\mathrm\{dir\}\_\{cjm\}^\{\\text\{OLS\}\}=\\mathrm\{sign\}\(\\beta\_\{cjm\}^\{\\text\{OLS\}\}\)\\in\\\{\-1,\+1\\\}\. We then derive an LLM\-implied directional signal from the model outputs\. As shown in the system prompt Figure[8](https://arxiv.org/html/2609.00345#A1.F8)and corresponding user prompt Figure[7](https://arxiv.org/html/2609.00345#A1.F7)in Appendix, the model is additionally asked to provide a ranked assessment of influential predictors together with whether each predictor pushes the mobility outcome downward, upward, or neither\. We encode this directional push aspijm∈\{−1,0,\+1\},p\_\{ijm\}\\in\\\{\-1,0,\+1\\\},where−1\-1,00, and\+1\+1denote negative, neutral, and positive influence on outcomemm, respectively\. To test whether the LLM\-implied pushes agree with the empirical OLS direction, we examine howpijmp\_\{ijm\}varies across the ordered predictor buckets \(described using a quintile\-based bucket,bij∈\{very\_low,low,neutral,high,very\_high\}b\_\{ij\}\\in\\\{\\texttt\{very\\\_low\},\\texttt\{low\},\\texttt\{neutral\},\\texttt\{high\},\\texttt\{very\\\_high\}\\\}\)\. If the empirical coefficient for predictorjjis positive, then increasing values ofbijb\_\{ij\}should correspond to increasingly positive LLM pushes\. If the empirical coefficient is negative, then increasing values ofbijb\_\{ij\}should correspond to increasingly negative LLM pushes\. We formalize this by orienting each LLM\-implied push by the empirical OLS sign: \(18\)p~ijm=dircjmOLS×pijm\.\\tilde\{p\}\_\{ijm\}=\\mathrm\{dir\}\_\{cjm\}^\{\\text\{OLS\}\}\\times p\_\{ijm\}\.After this transformation,larger values ofp~ijm\\tilde\{p\}\_\{ijm\}always indicate stronger agreement with the empirical direction, regardless of whether the OLS coefficient is positive or negative\. For each CBSA–predictor–outcome combination\(c,j,m\)\(c,j,m\), we then apply two one\-sided Jonckheere–Terpstra \(JT\) tests top~ijm\\tilde\{p\}\_\{ijm\}across the ordered bucketsbijb\_\{ij\}\. Theincreasing JT testevaluates whether the OLS\-oriented LLM signal becomes more aligned with the empirical direction as the predictor bucket increases fromvery\_lowtovery\_high\. Thedecreasing JT testevaluates whether the LLM signal instead moves in the opposite direction\. We apply Benjamini–Hochberg correction within each CBSA–outcome panel to account for multiple predictors\. Based on the corrected JT results, we assign each\(c,j,m\)\(c,j,m\)combination one of three directional labels: \{aligned\_increasing,opposite\_trend,no\_clear\_trend\}\.\\\{\\texttt\{aligned\\\_increasing\},\\ \\texttt\{opposite\\\_trend\},\\ \\texttt\{no\\\_clear\\\_trend\}\\\}\.A relationship is labeledaligned\_increasingwhen the increasing JT test is significant and the decreasing test is not, indicating that the LLM\-implied direction agrees with the empirical OLS direction\. It is labeledopposite\_trendwhen the decreasing JT test is significant and the increasing test is not, indicating that the LLM\-implied direction moves against the empirical relationship\. All remaining cases are labeledno\_clear\_trend\. ## 4\.Experiments ### 4\.1\.Experimental Setting Table 2\.Coverage statistics across the study CBSAs\. The table reports the number of CBGs covered, the total number of devices, and the distribution of devices per CBG\.#### 4\.1\.1\.Data ##### Mobility data\. We derive the mobility outcomes from anonymized and aggregated mobility traces provided by Cuebiq111[https://docs\.spectus\.ai/](https://docs.spectus.ai/)\. Aggregated mobility data is provided by Cuebiq, a location intelligence platform\. Data is collected from anonymized users who have opted\-in to provide access to their location data anonymously, through a CCPA and GDPR\-compliant framework\. Through its Social Impact program, Cuebiq provides mobility insights for academic research and humanitarian initiatives\. The Cuebiq responsible data sharing framework enables research partners to query anonymized and privacy\-enhanced data, by providing access to an auditable, on\-premise Data Cleanroom environment\. All final outputs provided to partners are aggregated in order to preserve privacy\. The data cover four Core\-Based Statistical Areas \(CBSAs\) in the United States: Atlanta–Sandy Springs–Roswell, GA\(ATL\); Los Angeles–Long Beach–Anaheim, CA\(LA\); Miami–Fort Lauderdale–West Palm Beach, FL\(MIA\); and San Francisco–Oakland–Berkeley, CA\(SF\)\. We use data from calendar year 2021 and retain only users observed on more than 60 days during the year\. To ensure reliable CBG\-level aggregation, we further restrict the analysis to CBGs with at least 20 devices\. Under these criteria, the final dataset covers 8,756 CBGs222CBGs corresponding to military locations are excluded in accordance with Cuebiq’s data\-sharing policies\.and 976,652 devices across the four CBSAs\. At the CBG level, device support varies across study areas, with median devices per CBG ranging from 57 in SF to 144 in ATL\. These mobility traces are used to construct the continuous user\-level mobility measures described in Section[3\.2\.2](https://arxiv.org/html/2609.00345#S3.SS2.SSS2), which are then aggregated to the CBG level and discretized into outcome classes for prediction\. Table[2](https://arxiv.org/html/2609.00345#S4.T2)and Figure[9](https://arxiv.org/html/2609.00345#A1.F9)\(Appendix\) show the distribution of the sample sizes as well as the CBGs covered\. ##### Socio\-demographic predictors\. We obtain socio\-demographic variables from the 2019 American Community Survey \(ACS\) 5\-year estimates\. These predictors characterize the population and household composition of each CBG and include measures such as median income, racial and ethnic composition, age structure, educational attainment, and vehicle ownership\. In the prediction framework, these variables serve as contextual signals describing the demographic and socioeconomic profile of each neighborhood\. For the OLS, we drop some of the correlated features\. ##### Built\-environment predictors\. Built\-environment variables are derived from the EPA Smart Location Database \(SLD\), which provides CBG\-level indicators related to density, street\-network structure, transit accessibility, and land\-use mix\. Because the raw SLD variables are numerous and highly correlated, we summarize them using principal component analysis \(PCA\)\. For each conceptual built\-environment domain, we retain the first principal component and use it as a composite index in the downstream analysis\. This yields four interpretable built\-environment predictors corresponding to density, connectivity, transit access, and land\-use mix\. ##### Predictor standardization\. To support both the LLM prompting framework and the supervised baselines, all predictors are harmonized at the CBG level\. For the LLM\-based experiments, each predictor is represented using its raw value, within\-CBSA percentile, and a quintile\-based bucket\. For the supervised baselines, the same underlying predictor values are used directly as numeric features\. #### 4\.1\.2\.LLM Prediction Settings We evaluate models from eight families in a zero\-shot setting using the prompt template described in Section[3\.3](https://arxiv.org/html/2609.00345#S3.SS3): Claude, GPT, GPT\-OSS, Gemini, Gemma, Qwen, DeepSeek, and Llama \(Table[3](https://arxiv.org/html/2609.00345#A1.T3), Appendix\)\. The set spans both proprietary and open\-weight models across a range of sizes\. Open\-weight models are served withvLLM, while proprietary and other hosted models are accessed through AWS cloud infrastructure\. Across all models, decoding is deterministic with temperature0\.00\.0\. For each CBG and mobility outcome, we issue one outcome\-specific prompt and parse the response into a class prediction and predictor\-level directional assessments LLM performance is reported on the same held\-out20%20\\%test split used for the supervised baselines\. For the RQ1 accuracy comparison, LLM performance is reported on the same held\-out20%20\\%test split used for the supervised baselines\. However, for the RQ2 directional analysis, we retain all the CBGs to maximize the power of the JT trend tests\. ### 4\.2\.RQ1: Predictive Performance Results Figure 3\.Mean 3\-class prediction accuracy across four CBSAs for mobility outcomes inferred from socio\-demographic and built\-environment predictors\. Outcomes are grouped into point\-level, line\-level, and temporal mobility features; bars show the top two LLMs and top two supervised baselines for each outcome, averaged over repeated 20% held\-out test splits, with error bars denoting standard deviation across 5 repetitions\. The horizontal dashed line marks chance accuracy \(1/3\), and the black line shows the performance gap between the best baseline and best LLM for each outcome\. Across CBSAs, supervised baselines consistently outperform zero\-shot LLMs, although LLMs remain above chance on many tasks, with the strongest performance typically observed for point\-level mobility outcomes\.##### Finding 1\.Neighborhood context provides strong signal for CBG\-level mobility prediction, but zero\-shot LLMs only partially recover it\. Figure[3](https://arxiv.org/html/2609.00345#S4.F3)shows that CBG\-level mobility outcomes are learnable from socio\-demographic and built\-environment context\. Across the 32 city–outcome pairs, the best supervised baseline achieves an average accuracy of0\.5800\.580, well above the random three\-class baseline of0\.3330\.333\. This confirms the notion that contextual predictors contain meaningful information about aggregate mobility behavior\. Zero\-shot LLMs also recover part of this signal, with the best LLM averaging0\.4350\.435accuracy across all outcomes\. Although LLMs remain consistently below supervised baselines, with an average best\-baseline–best\-LLM gap of0\.1440\.144\(range\[0\.059\[0\.059–0\.231\]0\.231\]\) accuracy points, this gap is also expected because the supervised models are trained directly on the target mobility labels while the LLMs operate without task\-specific examples\. Table[4](https://arxiv.org/html/2609.00345#A1.T4)contains the top performing model results and the their corresponding gaps with the best trained baselines\. ##### Finding 2\.Spatial extent outcomes are most predictable, but also exhibit the largest LLM–baseline gaps\. Figure[3](https://arxiv.org/html/2609.00345#S4.F3)shows that the spatial extent measures carry the strongest signal from neighborhood context and also show the widest gap between zero\-shot LLMs and supervised baselines\. For example, forconvex hull diameter,ellipse area, andradius of gyration, the best supervised baselines achieve average accuracies of0\.6530\.653,0\.6710\.671, and0\.6700\.670, respectively, while the best LLMs reach0\.4950\.495,0\.5100\.510, and0\.4840\.484\. Thus, LLMs perform relatively well on activity\-space extent, but supervised models extract substantially more predictive structure from the same contextual features\. The average LLM–baseline gap for these three outcomes is0\.1680\.168, larger than the gap for line\-level outcomes \(0\.1320\.132\) and the temporal outcome \(0\.1230\.123\)\. The pattern is especially pronounced in MIA CBSA, where the gap reaches0\.2310\.231forradius of gyration,0\.2000\.200forconvex hull diameter, and0\.1940\.194forellipse area\. ##### Finding 3:Entropy\-based mobility measures are the least predictable from static neighborhood context\. Figure[3](https://arxiv.org/html/2609.00345#S4.F3)shows that the two diversity measures,stay\-point entropyandtravel entropy, are the weakest outcomes for zero\-shot LLMs and are among the weakest for supervised baselines\. For LLMs they are the only outcomes whose best accuracy falls below0\.400\.40, averaging0\.3940\.394\(stay\-point entropy\) and0\.3420\.342\(travel entropy\) across the four CBSAs\. Travel entropy is the more extreme case: the best LLM stays near the three\-class chance level of0\.3330\.333in every CBSA, recovering little context\-based signal, whereas stay\-point entropy recovers a weak signal in most cities but also falls to chance in SF\. Supervised baselines perform lowest on the same two outcomes \(0\.4800\.480and0\.5270\.527\), the two lowest of any outcome, so the limitation is not specific to LLMs but reflects weaker predictability from the available CBG\-level socio\-demographic and built\-environment predictors\. Unlike spatial extent, these entropy measures capture behavioral variability in which locations and trips are taken, which likely depends on finer\-grained factors such as individual routines, trip purpose, and day\-to\-day variability that static neighborhood attributes do not capture\. ##### Finding 4\.Gemma3\-27B leads across most mobility outcomes, while Claude\-Sonnet is strongest for daily temporal fragmentation\. Across the LLM results in Fig\.[3](https://arxiv.org/html/2609.00345#S4.F3),Gemma3\-27Bis the most consistent zero\-shot predictor\. It is the top LLM in 20 of the 32 city–outcome pairs, ranks first forconvex hull diameter,ellipse area, andradius of gyrationin all four CBSAs, and is frequently the second\-best LLM elsewhere\. The exception isdaily temporal fragmentation, whereClaude\-Sonnet\-4\.5is the strongest LLM in all four CBSAs andGPT\-5\.1is second in three of them, so the proprietary chat\-tuned models are comparatively strongest on the temporal outcome rather than the spatial ones\. The best\-performing LLM therefore depends on the mobility dimension being inferred\. ### 4\.3\.RQ2: Directional Alignment Results Figure 4\.Empirical and LLM\-implied directional relationships between contextual predictors and mobility outcomes across the four CBSAs\. For each CBSA, the left heatmap reports standardized OLS regression coefficientsThe right heatmap reports the corresponding LLM\-implied directions from JT trend tests, restricted to predictor–outcome pairs with statistically significant OLS coefficients\. Each JT cell is annotated with the corresponding standardized OLS coefficient and JT significance indicator\.Results are shown forGemma\-3\-27B, the best\-performing LLM overall across mobility outcomes, except forDaily Temporal Fragmentation, for whichClaude\-Sonnet\-4\.5results are reported in Appendix Figure[10](https://arxiv.org/html/2609.00345#A1.F10)\.Figure 5\.Example of LLM prediction shifts across predictor bins for SF CBSA andConvex Hull Diameter\. Bars show the percentage of LLM predictions pushed toward lower, neutral, or higher mobility classes within each bin, illustrating how the model’s directional response is estimated from ordered predictor values\. Text annotations report the number of CBGs in each bin\.RQ2 asks whether LLMs recover empirically meaningful directional relationships between neighborhood context and mobility outcomes\. Figure[4](https://arxiv.org/html/2609.00345#S4.F4)compares the empirical directions estimated from OLS regression \(left panels\) with the LLM\-based directional alignment results from the Jonckheere–Terpstra analysis \(right panels\)\. Overall, the results suggest that LLMs often capture the*broad sign*of context–mobility relationships, but do so in a coarse and simplified way that does not fully reflect metropolitan variation or outcome\-specific complexity\. To illustrate how the LLM\-implied directions are derived, Figure[5](https://arxiv.org/html/2609.00345#S4.F5)shows one example for SF CBSA andConvex Hull Diameter\. BecauseGemma\-3\-27Bachieves the strongest classification performance for the point\-level and line\-level mobility outcomes, we report its directional alignment results for those outcome groups\. For the only temporal feature,Claude\-Sonnet\-4\.5performs best across CBSAs; we therefore report its directional results separately for that outcome in Figure[10](https://arxiv.org/html/2609.00345#A1.F10)in the Appendix\. ##### Finding 1:For empirically significant relationships, Gemma\-3\-27B applies predictor\-level directional priors that are largely invariant across mobility outcomes\. Among predictor–outcome pairs with statistically significant OLS coefficients, Gemma\-3\-27B shows a strong tendency to assign a fixed direction to each contextual predictor, rather than adapting its directional expectation to the specific mobility outcome\. This pattern is clearest when considering cases where the JT test yields a clear monotonic direction\. In these cases, 9/10 contextual predictors receive a single implied direction across nearly all city–outcome combinations\. Black population share is the clearest example\. Across the 25 Black population share relationships with a clear JT direction, Gemma\-3\-27B implies a negative direction in all 25 cases\. This pattern holds for all seven mobility outcomes in LA, MIA, and SF, and for four of the seven outcomes in ATL, with the remaining ATL outcomes showing no clear monotonic direction\. The same outcome\-invariant structure appears for several predictors with positive LLM\-implied directions\. Median income is mapped to a positive direction in all 25 cases with a clear JT direction\. Asian population share is also mapped to a positive direction in all 25 clear cases, despite the empirical OLS coefficients being negative in many of these relationships\. Land\-use mix shows the most uniform positive pattern, with positive implied directions in 27 of 28 city–outcome pairs and only one no\-direction case\. Transit access is similarly consistent: all 21 cases with a clear JT direction are positive\. Connectivity, density, and the share of young adults also show positive implied directions whenever the model produces a clear monotonic trend, although these predictors have more no\-direction cases\. Hispanic/Latino population share is the main exception to complete invariance, but it is still predominantly mapped to a positive direction, with 21 positive implied directions, one negative implied direction, and six no\-direction cases\. These results suggest that Gemma\-3\-27B does not primarily reason about how a contextual predictor should affect each mobility feature separately\. Instead, once a predictor is salient, the model appears to attach a broad directional prior to that predictor and reuse it across different types of mobility outcomes\. ##### Finding 2:Gemma\-3\-27B assigns a uniquely negative directional prior to Black population share\. Gemma\-3\-27B assigns markedly different directional priors across protected\-group predictors\. Black population share is the only protected\-group predictor that the model maps exclusively to lower mobility whenever the JT test yields a clear monotonic trend: across 28 Black population share relationships, 25 receive a negative implied direction and 3 show no clear direction, with all 25 clear cases being negative\. This pattern is strongest in LA, MIA and SF, where all seven outcomes receive a negative implied direction\. By contrast, Asian population share has negative OLS coefficients in 27 of 28 relationships, yet the model assigns a positive implied direction in all 25 clear JT cases and no clear direction in the remaining three\. Hispanic/Latino population share shows a similar pattern with negative OLS coefficients in 25/28 relationships, Gemma\-3\-27B assigns a positive implied direction in 21 cases, no clear direction in the remaining six, and a negative direction in only one\. These results suggest that the model does not apply a uniform low\-mobility prior to protected\-group population shares\. Instead, it encodes group\-specific priors, with Black population share receiving a consistently mobility\-reducing prior while Asian and Hispanic/Latino shares are usually mapped to mobility\-increasing or non\-negative directions\. ##### Finding 3:Directional priors do not change across metropolitan contexts\. Gemma\-3\-27B’s implied directions are largely consistent across CBSAs\. For most predictors, whenever the JT test yields a clear monotonic trend, the model assigns the same implied direction in ATL, LA, MIA, and SF\. Median income is mapped to a positive direction in all clear cases across the four CBSAs: 7/7 in ATL, 6/6 in LA, 6/6 in MIA, and 6/6 in SF\. Black population share shows the opposite but equally stable pattern, receiving a negative implied direction in all clear cases: 4/4 in ATL and 7/7 in each of LA, MIA, and SF\. Similar cross\-CBSA patterns appear for Asian population share, which is mapped positive in all clear cases across cities, and for land\-use mix, which is positive in 27 of 27 clear cases\. Built\-environment predictors also show stable positive priors: transit access is positive in all 21 clear cases, connectivity in all 17 clear cases, and density in all 11 clear cases\. Hispanic/Latino population share is the main exception, with one negative implied direction in Los Angeles, but it is still predominantly positive across CBSAs\. Overall, these patterns suggest that Gemma\-3\-27B’s directional priors are not strongly tailored to local metropolitan context\. Instead, the model tends to reuse the same predictor\-level direction across CBSAs, with differences appearing mainly as no\-direction cases rather than sign reversals\. ## 5\.Conclusion Fine\-grained mobility data are an important basis for computational models that characterize how people move through cities and support applications in transportation planning, public health, and emergency response\. However, such data are often proprietary, access\-restricted, and privacy\-sensitive, raising the question of whether large language models can provide useful priors about aggregate mobility in data\-scarce settings\. In this paper, we evaluated this question by testing whether zero\-shot LLMs can predict CBG\-level mobility outcomes from sociodemographic and built\-environment context across four U\.S\. metropolitan areas\. We further introduced a directional alignment framework that examines whether LLM\-implied predictor effects agree with empirical relationships estimated from observed mobility traces\. Our results show that LLMs recover nontrivial signal about neighborhood\-level mobility, but remain substantially below supervised baselines trained on local data\. More importantly, the alignment analysis shows that LLMs often rely on coarse and stable predictor\-level priors that persist across mobility outcomes and metropolitan areas, including asymmetric directional patterns for protected\-group predictors\. These findings suggest that LLMs may be useful as exploratory tools for mobility inference, but their predictions should not be treated as structurally grounded without empirical validation\. Auditing both predictive performance and directional alignment is therefore necessary before using LLMs in urban mobility analysis or policy\-support settings\. ###### Acknowledgements\. During the preparation of this work the author\(s\) used GPT \(OpenAI\) and Claude \(Anthropic\) to assist with polishing the writing, and Google Gemini and NotebookLM to help identify and organize relevant scholarly articles\. After using these tools/services, the author\(s\) reviewed and edited the content as needed and take\(s\) full responsibility for the content of the published article\. ## References - Abraret al\.\(2023\)S\. M\. Abrar, N\. Awasthi, D\. Smolyak, and V\. Frias\-MartinezAnalysis of performance improvements and bias associated with the use of human mobility data in covid\-19 case prediction models\.ACM Journal on Computing and Sustainable Societies1\(2\),pp\. 1–36\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Asanoet al\.\(2025\)H\. Asano, H\. Ouchi, A\. Kasuga, and R\. YonetaniMobQA: a benchmark dataset for semantic understanding of human mobility data through question answering\.arXiv preprint arXiv:2508\.11163\.Cited by:[§2\.2](https://arxiv.org/html/2609.00345#S2.SS2.p1.1)\. - Barbosaet al\.\(2018\)H\. Barbosa, M\. Barthelemy, G\. Ghoshal, C\. R\. James, M\. Lenormand, T\. Louail, R\. Menezes, J\. J\. Ramasco, F\. Simini, and M\. TomasiniHuman mobility: models and applications\.Physics Reports734,pp\. 1–74\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Beneduceet al\.\(2025\)C\. Beneduce, B\. Lepri, and M\. LucaLarge language models are zero\-shot next location predictors\.IEEE Access\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p2.1)\. - Bhandariet al\.\(2024\)P\. Bhandari, A\. Anastasopoulos, and D\. PfoserUrban mobility assessment using llms\.InProceedings of the 32nd ACM International Conference on Advances in Geographic Information Systems,pp\. 67–79\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p2.1),[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p4.1)\. - Brickaet al\.\(2024\)S\. Bricka, T\. Reuscher, P\. Schroeder, M\. Fisher, J\. Beard, and X\. L\. SunSummary of travel trends: 2022 national household travel survey\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Buckeeet al\.\(2020\)C\. O\. Buckee, S\. Balsari, J\. Chan, M\. Crosas, F\. Dominici, U\. Gasser, Y\. H\. Grad, B\. Grenfell, M\. E\. Halloran, M\. U\. Kraemer,et al\.Aggregated mobility data could help fight covid\-19\.Science368\(6487\),pp\. 145–146\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Changet al\.\(2021\)S\. Chang, E\. Pierson, P\. W\. Koh, J\. Gerardin, B\. Redbird, D\. Grusky, and J\. LeskovecMobility network models of covid\-19 explain inequities and inform reopening\.Nature589\(7840\),pp\. 82–87\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Chenet al\.\(2025\)Y\. Chen, Y\. Tao, Y\. Jiang, S\. Liu, H\. Yu, and G\. CongEnhancing large language models for mobility analytics with semantic location tokenization\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\. 2,pp\. 262–273\.Cited by:[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p3.1)\. - Duet al\.\(2025\)Y\. Du, T\. Aoki, and N\. FujiwaraA review of human mobility: linking data, models, and real\-world applications\.Journal of Computational Social Science8\(4\),pp\. 90\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Fenget al\.\(2025\)J\. Feng, J\. Zhang, T\. Liu, X\. Zhang, T\. Ouyang, J\. Yan, Y\. Du, S\. Guo, and Y\. LiCitybench: evaluating the capabilities of large language models for urban tasks\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\. 2,pp\. 5413–5424\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p2.1)\. - Garberet al\.\(2022\)M\. D\. Garber, K\. Labgold, and M\. R\. KramerOn selection bias in comparison measures of smartphone\-generated population mobility: an illustration of no\-bias conditions with a commercial data source\.Annals of epidemiology70,pp\. 16–22\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Gonget al\.\(2024\)L\. Gong, Y\. Lin, X\. Zhang, Y\. Lu, X\. Han, Y\. Liu, S\. Guo, Y\. Lin, and H\. WanMobility\-llm: learning visiting intentions and travel preference from human mobility data with large language models\.Advances in Neural Information Processing Systems37,pp\. 36185–36217\.Cited by:[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p3.1)\. - Gonzálezet al\.\(2024\)A\. B\. R\. González, J\. Burrieza\-Galán, J\. J\. V\. Díaz, I\. P\. de Castro, M\. R\. Wilby, and O\. G\. Cantú\-RosUsing app usage data from mobile devices to improve activity\-based travel demand models\.IEEE Transactions on Big Data10\(5\),pp\. 633–643\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Hägerstrand \(1970\)T\. HägerstrandWhat about people in regional science\.Transport Sociology: Social aspects of transport planning,pp\. 143–158\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Hartmannet al\.\(2023\)V\. Hartmann, A\. Suri, V\. Bindschaedler, D\. Evans, S\. Tople, and R\. WestSok: memorization in general\-purpose large language models\.arXiv preprint arXiv:2310\.18362\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p4.1)\. - Liet al\.\(2024a\)P\. Li, M\. de Rijke, H\. Xue, S\. Ao, Y\. Song, and F\. D\. SalimLarge language models for next point\-of\-interest recommendation\.InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 1463–1472\.Cited by:[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p3.1)\. - Liet al\.\(2024b\)S\. Li, T\. Tran, H\. Lin, J\. Krumm, C\. Shahabi, L\. Zhao, K\. Shafique, and L\. XiongGeo\-llama: leveraging llms for human mobility trajectory generation with spatiotemporal constraints\.arXiv preprint arXiv:2408\.13918\.Cited by:[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p1.1)\. - Liet al\.\(2024c\)X\. Li, F\. Huang, J\. Lv, Z\. Xiao, G\. Li, and Y\. YueBe more real: travel diary generation using llm agents and individual profiles\.arXiv preprint arXiv:2407\.18932\.Cited by:[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p2.1)\. - Liet al\.\(2024d\)Z\. Li, L\. Xia, J\. Tang, Y\. Xu, L\. Shi, L\. Xia, D\. Yin, and C\. HuangUrbangpt: spatio\-temporal large language models\.InProceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining,pp\. 5351–5362\.Cited by:[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p3.1)\. - Liuet al\.\(2025a\)T\. Liu, M\. Li, and Y\. YinAligning llm with human travel choices: a persona\-based embedding learning approach\.arXiv preprint arXiv:2505\.19003\.Cited by:[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p2.1)\. - Liuet al\.\(2025b\)Y\. Liu, X\. Liao, H\. Ma, B\. Y\. He, C\. Stanford, and J\. MaHuman mobility modeling with household coordination activities under limited information via retrieval\-augmented llms\.In2025 IEEE 28th International Conference on Intelligent Transportation Systems \(ITSC\),pp\. 951–958\.Cited by:[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p2.1)\. - Luoet al\.\(2024\)Y\. Luo, Z\. Cao, X\. Jin, K\. Liu, and L\. YinDeciphering human mobility: inferring semantics of trajectories with large language models\.In2024 25th IEEE international conference on mobile data management \(MDM\),pp\. 289–294\.Cited by:[§2\.2](https://arxiv.org/html/2609.00345#S2.SS2.p1.1)\. - Maet al\.\(2025\)H\. Ma, X\. Liao, Y\. Liu, Q\. Jiang, C\. Stanford, S\. Cao, and J\. MaLearning universal human mobility patterns with a foundation model for cross\-domain data fusion\.Transportation Research Part C: Emerging Technologies180,pp\. 105311\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p1.1)\. - Manviet al\.\(2024\)R\. Manvi, S\. Khanna, M\. Burke, D\. Lobell, and S\. ErmonLarge language models are geographically biased\.arXiv preprint arXiv:2402\.02680\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p4.1),[§2\.2](https://arxiv.org/html/2609.00345#S2.SS2.p2.1)\. - Manviet al\.\(2023\)R\. Manvi, S\. Khanna, G\. Mai, M\. Burke, D\. Lobell, and S\. ErmonGeollm: extracting geospatial knowledge from large language models\.arXiv preprint arXiv:2310\.06213\.Cited by:[§2\.2](https://arxiv.org/html/2609.00345#S2.SS2.p1.1)\. - Moayeriet al\.\(2024\)M\. Moayeri, E\. Tabassi, and S\. FeiziWorldbench: quantifying geographic disparities in llm factual recall\.InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency,pp\. 1211–1228\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p4.1),[§2\.2](https://arxiv.org/html/2609.00345#S2.SS2.p2.1)\. - Ronget al\.\(2024\)C\. Rong, J\. Ding, and Y\. LiAn interdisciplinary survey on origin\-destination flows modeling: theory and techniques\.ACM Computing Surveys57\(1\),pp\. 1–49\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p3.1)\. - Shaoet al\.\(2024\)C\. Shao, F\. Xu, B\. Fan, J\. Ding, Y\. Yuan, M\. Wang, and Y\. LiBeyond imitation: generating human mobility from context\-aware reasoning with large language models\.arXiv preprint arXiv:2402\.09836\.Cited by:[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p1.1)\. - Wanget al\.\(2024\)J\. Wang, R\. Jiang, C\. Yang, Z\. Wu, M\. Onizuka, R\. Shibasaki, N\. Koshizuka, and C\. XiaoLarge language models as urban residents: an llm agent framework for personal mobility generation\.Advances in Neural Information Processing Systems37,pp\. 124547–124574\.Cited by:[§2\.1](https://arxiv.org/html/2609.00345#S2.SS1.p1.1)\. - Wanget al\.\(2023\)X\. Wang, M\. Fang, Z\. Zeng, and T\. ChengWhere would i go next? large language models as human mobility predictors\.arXiv preprint arXiv:2308\.15197\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p2.1)\. - Wuet al\.\(2019\)L\. Wu, L\. Yang, Z\. Huang, Y\. Wang, Y\. Chai, X\. Peng, and Y\. LiuInferring demographics from human trajectories and geographical context\.Computers, Environment and Urban Systems77,pp\. 101368\.Cited by:[§3\.2\.2](https://arxiv.org/html/2609.00345#S3.SS2.SSS2.p1.1)\. - Wu and Wang \(2024\)X\. Wu and Q\. R\. WangPopular llms amplify race and gender disparities in human mobility\.arXiv preprint arXiv:2411\.14469\.Cited by:[§2\.2](https://arxiv.org/html/2609.00345#S2.SS2.p2.1)\. - Xuet al\.\(2025\)F\. Xu, Q\. Wang, E\. Moro, L\. Chen, A\. Salazar Miranda, M\. C\. González, M\. Tizzoni, C\. Song, C\. Ratti, L\. Bettencourt,et al\.Using human mobility data to quantify experienced urban inequalities\.Nature human behaviour9\(4\),pp\. 654–664\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Yabeet al\.\(2025\)T\. Yabe, B\. García Bulle Bueno, M\. R\. Frank, A\. Pentland, and E\. MoroBehaviour\-based dependency networks between places shape urban economic resilience\.Nature human behaviour9\(3\),pp\. 496–506\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Yanget al\.\(2023\)Y\. Yang, A\. Pentland, and E\. MoroIdentifying latent activity behaviors and lifestyles using mobility data to describe urban dynamics\.EPJ Data Science12\(1\),pp\. 15\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p1.1)\. - Zhanget al\.\(2025\)Y\. Zhang, R\. Zhao, Z\. Huang, X\. Wang, Y\. Ma, and Y\. LongGenAI models capture urban science but oversimplify complexity\.arXiv preprint arXiv:2505\.13803\.Cited by:[§2\.2](https://arxiv.org/html/2609.00345#S2.SS2.p3.1)\. - Zhaoet al\.\(2024\)D\. Zhao, A\. Mihăiţă, Y\. Ou, H\. Grzybowska, and M\. LiOrigin–destination matrix estimation for public transport: a multi\-modal weighted graph approach\.Transportation Research Part C: Emerging Technologies165,pp\. 104694\.Cited by:[§1](https://arxiv.org/html/2609.00345#S1.p3.1)\. ## Appendix AAppendix System Prompt for Mobility Outcome Classification``` You are a careful classifier for mobility outcomes. Output ONLY one XML document that starts with <response> and ends with </response>, with EXACTLY three child elements in this order: <reasoning>, <rank>, <answer>. Do not output anything else. Primary task: Predict the mobility outcome class in <answer> using ONLY the information provided in the user prompt (outcome definition, CBSA-specific cutpoints, and CBG feature buckets/percentiles). Keep <reasoning> concise and focused on direction and justification: explain how the most influential features (the ones you will rank highest) push the mobility outcome up or down, and why they matter more than other features. Avoid listing everything; summarize. <rank> is an explanation-only section. List the 10 most relevant predictors for THIS CBG’s prediction (most important first). <rank> MUST contain exactly 10 <item/> tags. Each <item/> MUST have EXACTLY three attributes: feature="..." val="very_low|low|neutral|high|very_high" push="+"|"-"|"0" - val MUST be copied exactly from the bucket given in the prompt for that feature. - push is the direction this feature (given its val) pushes the mobility outcome for THIS CBG: "+" pushes toward higher outcome, "-" toward lower, "0" neutral or unclear. <answer> MUST be exactly one of: A | B | C. A = low outcome (lowest tertile), B = middle tertile, C = high outcome (highest tertile). The tertile cutpoints for A/B/C are provided in the prompt. When deciding <answer>, weigh items by rank: items earlier in <rank> matter much more than later items. Treat ranks 1--3 as primary evidence, ranks 4--6 as supporting evidence, and ranks 7--10 as weak evidence. ``` Figure 6\.System prompt for RQ1User Prompt for Directional Alignment``` You are given Census Block Group (CBG) characteristics for CBSA: [CBSA name]. ## Outcome Description Mobility outcome: [outcome name] Definition: [outcome definition] ## Predictor Descriptor - Percentiles are computed within the CBSA. - Predictor buckets correspond to CBSA-wide quintiles: very_low | low | neutral | high | very_high. - Outcome labels are tertiles: low | neutral | high. ## CBG Features - [feature 1]; [raw value] (CBSA percentile Pxx, [bucket]) - [feature 2]; [raw value] (CBSA percentile Pxx, [bucket]) - ... - [feature p]; [raw value] (CBSA percentile Pxx, [bucket]) ## Task Based on the outcome definition provided, assess the directional push (+, -, 0) for each feature in this specific CBG. ``` Figure 7\.User prompt for RQ2 Directional Alignment\.System Prompt for Directional Alignment``` You are a spatial data analyst. Your task is to provide reasoning for characteristic mobility outcomes within a small geographic region (Census Block Group). Output ONLY one XML document that starts with <response> and ends with </response>. This document must contain EXACTLY two child elements in this order: <reasoning>, <assessment>. ### Primary Task Identify the directional push of every predictor on the mobility outcome for this CBG. ### Analytical Framework (Two-Step Evaluation) For each predictor, apply the following logic: 1. Classify the relationship: Is the feature a Facilitator (positive association) or a Restrictor (inverse association) for the outcome? 2. Apply the instance value: - If a feature is a Facilitator and its value is high or very_high, then push="+" - If a feature is a Facilitator and its value is low or very_low, then push="-" - If a feature is a Restrictor and its value is high or very_high, then push="-" - If a feature is a Restrictor and its value is low or very_low, then push="+" <reasoning> Provide a concise explanation. Explicitly list which predictors you classified as Facilitators and which you classified as Restrictors based on the outcome definition. </reasoning> <assessment> List EVERY predictor provided in the user prompt as an <item/>. Each <item/> must have these attributes: - feature: the exact name of the feature as provided - val: the observed bucket (very_low|low|neutral|high|very_high) - push: "+", "-", or "0" </assessment> ``` Figure 8\.System prompt for RQ2 Directional Alignment\.Table 3\.Parse success rates across evaluated LLMs\. Models with parse success rate below 95% were excluded from downstream analysis\.Notes\.HD = hull diameter, EA = ellipse area, RoG = radius of gyration, SE = stay entropy, AD = average duration, TL = travel length, TE = travel entropy, TF = temporal fragmentation\. MNLR – multinomial logistic regression\. Table 4\.Best zero\-shot LLM and supervised baseline by CBSA and mobility outcome\. Gap is computed as best baseline accuracy minus best LLM accuracy\.\(a\)ATL \(b\)MIA \(c\)LA \(d\)SF Figure 9\.Sample size distributions for CBGs across four CBSAs\.Figure 10\.Empirical and Claude\-4\.5\-Sonnet’s implied directional relationships forDaily Temporal Fragmentationacross the four study CBSAs\. For each CBSA, the left heatmap shows standardized OLS coefficients and the right heatmap shows JT directional alignment, restricted to statistically significant OLS relationships\. Red and blue indicate negative and positive OLS coefficients in the left panels, and opposite or aligned JT directions in the right panels\. Asterisks denote statistical significance\.
Similar Articles
The Geography of Algorithmic Judgment: LLM Intermediaries, Place Identity, and Racial Steering in Housing Search
This paper conducts a behavioral audit of seven open-weight and closed-source LLMs across four U.S. cities, finding that racial steering in housing recommendations is an emergent behavior of the model's interpretive license, varying by user identity and city context.
On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
This paper investigates how LLMs' internal priors affect zero-shot annotation performance, finding that nearly two-thirds of errors resist prompt-based correction and introducing Definition-Specific Familiarity as a better predictor than memorization metrics.
When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation
This paper introduces a validation framework to evaluate whether LLM-based urban simulators reproduce empirical human mobility patterns. Using data from Paris and Shanghai, the authors find a substantial gap between plausible narratives and realistic mobility constraints, and provide open infrastructure for reproducible evaluation.
LLMs know when they are wrong. I made a fix relating to Anthropic's new "global workspace" paper [R]
The author presents a method to make LLMs verbalize calibrated confidence by using a linear probe on mid-layer states and a small trained bridge to confidence logits, requiring only 200 labeled examples and no weight modification. This is linked to Anthropic's global workspace paper explaining the know-say gap.
When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators
This study evaluates LLMs as data quality annotators on e-commerce tasks, finding they outperform baselines when background knowledge is required but offer limited advantages for tasks with strong lexical signals, while demonstrating high consistency across runs.