Space as an Interventional Invariant: Cross-Modal Predictive Geometry for Stratified Cities and Em-Spaced Intelligence

arXiv cs.LG Papers

Summary

This paper defines space as an interventional invariant and develops a cross-modal predictive geometry to unify spatial structure across mathematics, physics, cognition, and urban science.

arXiv:2609.11959v1 Announce Type: new Abstract: Space is a foundational concept across mathematics, physics, spatial cognition, urban science, and embodied intelligence, yet these fields often treat spatial structure either as a shared geometric container or as a collection of disconnected representations. Such approaches struggle to explain how heterogeneous sensory and urban processes can jointly reveal a common spatial structure, particularly when different modalities do not share the same metric or representation. This paper addresses this gap by defining space as an interventional invariant: the minimal relational structure that preserves local compatibility and the conditional laws of future observations under admissible actions. We develop a cross-modal predictive geometry that integrates local state spaces, modality-specific observation maps, an action groupoid, and a canonical predictive-state quotient, with explicit causal conditions for identifying interventional rather than merely observational structure. The key theoretical result shows that, under joint point separation, equivariance, and interventional faithfulness, the latent space is identifiable up to the centraliser of the intervention group, thereby reducing representational ambiguity to residual coordinate freedom. The framework is further extended to stratified urban systems using sheaf-valued representations, allowing geometric, physical, mobility, social, and economic layers to coexist without being reduced to a single metric. Synthetic experiments under noise evaluate equivariance, predictive sufficiency, holonomy, restriction-map recovery, cross-scale consistency, and context saturation. The resulting framework provides a unified and falsifiable foundation for spatial cognition, urban science, embodied AI, and em-spaced intelligence.
Original Article
View Cached Full Text

Cached at: 09/14/26, 08:29 AM

# Cross-Modal Predictive Geometry for Stratified Cities and Em-Spaced Intelligence
Source: [https://arxiv.org/html/2609.11959](https://arxiv.org/html/2609.11959)
## Space as an Interventional Invariant: Cross\-Modal Predictive Geometry for Stratified Cities and Em\-Spaced Intelligence

Tao YangCorresponding author:yangtao128@tsinghua\.edu\.cnSchool of Architecture, Tsinghua University, Beijing, ChinaKunyao LiSchool of Engineering, Cardiff University, Cardiff, United KingdomHaijiang LiSchool of Engineering, Cardiff University, Cardiff, United Kingdom

###### Abstract

Space is a foundational concept across mathematics, physics, spatial cognition, urban science, and embodied intelligence, yet these fields often treat spatial structure either as a shared geometric container or as a collection of disconnected representations\. Such approaches struggle to explain how heterogeneous sensory and urban processes can jointly reveal a common spatial structure, particularly when different modalities do not share the same metric or representation\. This paper addresses this gap by defining space as an interventional invariant: the minimal relational structure that preserves local compatibility and the conditional laws of future observations under admissible actions\. We develop a cross\-modal predictive geometry that integrates local state spaces, modality\-specific observation maps, an action groupoid, and a canonical predictive\-state quotient, with explicit causal conditions for identifying interventional rather than merely observational structure\. The key theoretical result shows that, under joint point separation, equivariance, and interventional faithfulness, the latent space is identifiable up to the centraliser of the intervention group, thereby reducing representational ambiguity to residual coordinate freedom\. The framework is further extended to stratified urban systems using sheaf\-valued representations, allowing geometric, physical, mobility, social, and economic layers to coexist without being reduced to a single metric\. Synthetic experiments under noise evaluate equivariance, predictive sufficiency, holonomy, restriction\-map recovery, cross\-scale consistency, and context saturation\. The resulting framework provides a unified and falsifiable foundation for spatial cognition, urban science, embodied AI, and em\-spaced intelligence\.

Keywords\.space; intervention; causal identifiability; cross\-modal prediction; predictive state; bisimulation; sheaf; holonomy; projective limit; stratified space; urban geometry; embodied spatial intelligence; em\-spaced intelligence

> Space is not identified with a coordinate container\. It is the least relational structure that preserves local compatibility and the intervention\-conditioned future laws shared by heterogeneous modes of observation\.

## 1Introduction

The word*space*performs several jobs at once\. In mathematics it names an object equipped with specified structure and morphisms\. In physics it refers to an empirically constrained geometry within a dynamical theory\. In perception it names the organisation by which an agent anticipates what will change when it moves\. Urban practice adds a further complication\. Buildings and roads are explicit geometric objects, whereas wind, heat, electromagnetic propagation, sound, traffic, institutions and exchange generate less visible but equally consequential geometries of reachability, resistance and flow\. A satisfactory theory must relate these senses without erasing their differences\.

The starting observation is that sensory modalities need not resemble one another in order to disclose a common world\. A turn of the head changes retinal flow, binaural delay and proprioceptive state according to one movement\. Opening a door modifies light, air, heat, sound and accessibility together\. Closing a road leaves the Euclidean map nearly unchanged but transforms travel\-time, economic and social reachability\. The stable object is therefore not a shared signal format\. It is a family of transformations that remains jointly predictable under intervention\.

This paper proposes that the primary definition of space should be relational and interventional, while dimension should be treated as a derived complexity index\. Dimension alone cannot determine which states are adjacent, which transitions are admissible, which local descriptions glue, or which interventions distinguish two apparently identical situations\. Conversely, a relation becomes spatial only when it supports locality, reachability, composable change, cross\-modal covariance and stable prediction\. Mere statistical association is not enough\.

##### Contribution 1\.

A typed definition of cross\-modal predictive geometry is given in terms of probes, local states, interventions and conditional future laws, together with the causal assumptions under which the conditional laws are interventional rather than merely observational\.

##### Contribution 2\.

The central identifiability result is proved: when the observation family jointly separates points and the learned representation is equivariant, the latent space is recovered not up to an arbitrary homeomorphism but up to the centraliser of the intervention group\. When the action is simply transitive this centraliser is the group itself, so space is recovered as a torsor and coordinates are exactly the residual gauge freedom — no more and no less\. Removing the interventions collapses the statement to the known impossibility results for unsupervised disentanglement\.

##### Contribution 3\.

Space is shown to be well defined independently of any single observational context\. The context\-indexed predictive quotients form a projective system, and space is its limit; the limit is attained at a finite stage exactly when a sufficient context exists, which is an experimentally testable saturation claim\.

##### Contribution 4\.

A stratified, fibred and sheaf\-valued urban model places physical and social geometries on a common base without assuming that they share one distance function\. A holonomy obstruction is exhibited that no graph Laplacian with identity restriction maps can detect, which is what the sheaf formalism buys over a multilayer network\.

##### Contribution 5\.

The formulae are subjected to type, limit and numerical checks under noise, ablation and refinement, and the construction is translated into an experimentally testable architecture for embodied and em\-spaced intelligence\.

## 2Three senses of space

### 2\.1Mathematical space

A mathematical space is an objectXXin a category𝐂\\mathbf\{C\}, together with the additional structure that makes the intended questions meaningful\. A topological space privileges continuity; a Riemannian manifold adds a metric tensor; a metric\-measure space joins distance to mass; a graph records adjacency; a sheaf records the passage from local data to compatible global data; a Hilbert space records linear and inner\-product structure\. There is no structure\-free mathematical meaning of space\. Its identity is always relative to a chosen class of structure\-preserving maps\[[40](https://arxiv.org/html/2609.11959#bib.bib1),[22](https://arxiv.org/html/2609.11959#bib.bib2),[53](https://arxiv.org/html/2609.11959#bib.bib3)\]\.

### 2\.2Physical space

Physical space is not merely a manifold written in coordinates\. It is a model together with fields, laws, symmetries, measuring operations and error\. In general relativity a three\-dimensional space is usually obtained from a Lorentzian spacetime only after a choice of foliation; in continuum mechanics, the material and spatial descriptions are related but distinct; in thermodynamics and field theory, geometry is inferred through what instruments and bodies do\. Coordinate changes are representational\. Causal and metrical invariants are empirical\[[42](https://arxiv.org/html/2609.11959#bib.bib4),[44](https://arxiv.org/html/2609.11959#bib.bib22)\]\.

### 2\.3Perceptual space

Perceptual space is the organisation of possible sensorimotor consequences\. The sensorimotor tradition correctly insists that the spatial content of seeing, hearing or touching lies partly in lawful changes under movement\[[19](https://arxiv.org/html/2609.11959#bib.bib47),[43](https://arxiv.org/html/2609.11959#bib.bib48),[54](https://arxiv.org/html/2609.11959#bib.bib49),[35](https://arxiv.org/html/2609.11959#bib.bib50)\]\. Cognitive\-map research adds that a useful spatial representation supports flexible inference beyond immediate sensation\[[7](https://arxiv.org/html/2609.11959#bib.bib51)\]\. The present proposal sharpens these claims: two embodied histories occupy the same perceptual state precisely when every admissible future policy induces the same law of future multimodal observations\.

![Refer to caption](https://arxiv.org/html/2609.11959v1/figures/fig1.png)Figure 1:Mathematical, physical and perceptual space are linked by modelling, measurement and action, but their equivalence relations remain distinct\.###### Definition 2\.1\(interventional invariant\)\.

An interventional invariant is an equivalence class of relational models whose observable conditional laws are unchanged under an admissible change of representation, but vary lawfully under interventions on the represented system\.

### 2\.4Relation to existing predictive and causal formalisms

The construction below inherits from four literatures, and it is worth stating precisely what is taken and what is added\. From computational mechanics comes the*causal state*: the equivalence class of pasts inducing the same conditional distribution over futures, together with its minimality among sufficient statistics\[[49](https://arxiv.org/html/2609.11959#bib.bib36)\]\. From reinforcement learning come predictive state representations\[[38](https://arxiv.org/html/2609.11959#bib.bib37)\], observable operator models\[[32](https://arxiv.org/html/2609.11959#bib.bib38)\]and, in the controlled case, stochastic bisimulation and model minimisation for Markov decision processes\[[20](https://arxiv.org/html/2609.11959#bib.bib39)\], with their quantitative refinements as bisimulation metrics\[[17](https://arxiv.org/html/2609.11959#bib.bib40),[11](https://arxiv.org/html/2609.11959#bib.bib41)\]and their use as representation\-learning objectives\[[64](https://arxiv.org/html/2609.11959#bib.bib42)\]\. Definition 3\.1 and Theorem 3\.2 are the extension of these notions to a groupoid of interventions acting on a sheaf of typed local states; the extension is what allows locality, cross\-modal compatibility and scale to constrain the same quotient, but the minimality argument itself is not new and is not claimed as the contribution\.

From causal inference comes the distinction between conditioning and intervening\[[44](https://arxiv.org/html/2609.11959#bib.bib22)\], the identification of interventional laws from logged data under sequential ignorability and positivity\[[46](https://arxiv.org/html/2609.11959#bib.bib23)\], and the recent programme of causal representation learning\[[48](https://arxiv.org/html/2609.11959#bib.bib24)\]\. The specific results that make the present paper’s central claim provable are the identifiability theorems for latent variables under interventions\[[9](https://arxiv.org/html/2609.11959#bib.bib25),[1](https://arxiv.org/html/2609.11959#bib.bib26),[56](https://arxiv.org/html/2609.11959#bib.bib27),[37](https://arxiv.org/html/2609.11959#bib.bib28),[51](https://arxiv.org/html/2609.11959#bib.bib29),[57](https://arxiv.org/html/2609.11959#bib.bib30)\]; Theorem 3\.7 is their topological counterpart, with the residual ambiguity expressed as a centraliser rather than as a permutation\-and\-scaling group\. From applied topology come cellular sheaves and their Laplacians\[[47](https://arxiv.org/html/2609.11959#bib.bib15),[25](https://arxiv.org/html/2609.11959#bib.bib16),[8](https://arxiv.org/html/2609.11959#bib.bib18),[26](https://arxiv.org/html/2609.11959#bib.bib17)\], the connection Laplacian and vector diffusion maps\[[50](https://arxiv.org/html/2609.11959#bib.bib19),[2](https://arxiv.org/html/2609.11959#bib.bib20)\], and sheaf\-based sensor fusion with non\-identity restriction maps\[[33](https://arxiv.org/html/2609.11959#bib.bib21)\]\.

Two further connections are worth flagging even though they are not developed here\. The successor representation and successor features\[[14](https://arxiv.org/html/2609.11959#bib.bib43),[4](https://arxiv.org/html/2609.11959#bib.bib45)\]are the linear, discounted special case of the predictive state, and the hippocampal predictive\-map hypothesis\[[52](https://arxiv.org/html/2609.11959#bib.bib44)\]is the corresponding neuroscientific claim; the cognitive\-map literature\[[7](https://arxiv.org/html/2609.11959#bib.bib51)\]can therefore be read as an empirical instance of Definition 3\.1 rather than as a separate tradition\. Modern self\-supervised world models\[[36](https://arxiv.org/html/2609.11959#bib.bib46)\]optimise a predictive objective of the same type, but typically without an explicit action groupoid or descent constraint, which is precisely the structure the penalties of Section[6\.2](https://arxiv.org/html/2609.11959#S6.SS2)supply\.

### 3\.1Typed observational system

LetBBbe a site of probes\. A probe may be a spatial region, sensor footprint, body\-centred receptive field, cell of a complex, or finite experimental context\. Letℱ\\mathcal\{F\}be a sheaf \(or, where homotopy is essential, an∞\\infty\-sheaf\) of local system states onBB\. For each modalityα\\alphain a finite index setAA, letYαY\_\{\\alpha\}be a standard Borel observation space and lethα,U:ℱ​\(U\)→Yα​\(U\)h\_\{\\alpha,U\}\\colon\\mathcal\{F\}\(U\)\\to Y\_\{\\alpha\}\(U\)be a measurable local observation map\. LetGGbe a measurable action groupoid: its arrows include bodily motions, environmental controls and institutional interventions, and composition represents executable succession\.

𝔖=\(B,ℱ,G,\{Yα,hα\}α∈A,Φ,P\)\.\\mathfrak\{S\}\\;=\\;\\bigl\(B,\\ \\mathcal\{F\},\\ G,\\ \\\{Y\_\{\\alpha\},h\_\{\\alpha\}\\\}\_\{\\alpha\\in A\},\\ \\Phi,\\ P\\bigr\)\.\(1\)HereΦg\\Phi\_\{g\}maps a state on the source of an actionggto a state on its target, whilePPis the family of regular conditional laws for future observations\. Equation \([1](https://arxiv.org/html/2609.11959#S3.E1)\) is a type declaration, not a claim that all modalities inhabit one vector space\. Their commonality lies in their response to the same arrows ofGG\.

hα∘Φg:ℱ​\(s​\(g\)\)→Yα​\(t​\(g\)\),Tαg∘hα:ℱ​\(s​\(g\)\)→Yα​\(t​\(g\)\)\.h\_\{\\alpha\}\\circ\\Phi\_\{g\}\\colon\\mathcal\{F\}\(s\(g\)\)\\to Y\_\{\\alpha\}\(t\(g\)\),\\qquad T^\{g\}\_\{\\alpha\}\\circ h\_\{\\alpha\}\\colon\\mathcal\{F\}\(s\(g\)\)\\to Y\_\{\\alpha\}\(t\(g\)\)\.\(2\)The two sides of \([2](https://arxiv.org/html/2609.11959#S3.E2)\) are now comparable\. Approximate equivariance is measured rather than asserted by an untyped resemblance sign:

Eeq=𝔼g,x​\[∑αwα​dα​\(hα​\(Φg​x\),Tαg​\(hα​\(x\)\)\)2\]\.E\_\{\\mathrm\{eq\}\}\\;=\\;\\mathbb\{E\}\_\{g,x\}\\Bigl\[\\ \\sum\_\{\\alpha\}w\_\{\\alpha\}\\,d\_\{\\alpha\}\\bigl\(h\_\{\\alpha\}\(\\Phi\_\{g\}x\),\\,T^\{g\}\_\{\\alpha\}\(h\_\{\\alpha\}\(x\)\)\\bigr\)^\{2\}\\ \\Bigr\]\.\(3\)

### 3\.2The canonical predictive state

LetHHbe the standard Borel space of finite multimodal histories\. LetΠ\\Pibe a countable policy class andKKa countable set of horizons\. Forπ∈Π\\pi\\in\\Piandk∈Kk\\in K, writePπ,k\(⋅∣h\)P^\{\\pi,k\}\(\\cdot\\mid h\)for the regular conditional law of the nextkkobservations under policyπ\\pi\. Define

S\(h\)=\(Pπ,k\(⋅∣h\)\)π∈Π,k∈K∈∏π,k𝒫\(Y1:k\),S\(h\)\\;=\\;\\bigl\(P^\{\\pi,k\}\(\\cdot\\mid h\)\\bigr\)\_\{\\pi\\in\\Pi,\\,k\\in K\}\\ \\in\\ \\prod\_\{\\pi,k\}\\mathcal\{P\}\\bigl\(Y^\{1:k\}\\bigr\),\(4\)h∼Π,Kh′⇔Pπ,k\(⋅∣h\)=Pπ,k\(⋅∣h′\)for allπ∈Πandk∈K\.h\\sim\_\{\\Pi,K\}h^\{\\prime\}\\iff P^\{\\pi,k\}\(\\cdot\\mid h\)=P^\{\\pi,k\}\(\\cdot\\mid h^\{\\prime\}\)\\quad\\text\{for all \}\\pi\\in\\Pi\\text\{ and \}k\\in K\.\(5\)
###### Definition 3\.1\(cross\-modal predictive space\)\.

The perceptual state space induced by\(Π,K\)\(\\Pi,K\)is the imageS​\(H\)S\(H\), equivalently the quotientH/∼Π,KH/\{\\sim\_\{\\Pi,K\}\}, equipped with the smallestσ\\sigma\-algebra making every coordinate prediction measurable\.

###### Theorem 3\.2\(minimal predictive quotient\)\.

Writeσ​\(S\)=S−1​\(ℬ​\(∏π,k𝒫​\(Y1:k\)\)\)⊆ℬ​\(H\)\\sigma\(S\)=S^\{\-1\}\\bigl\(\\mathcal\{B\}\(\\prod\_\{\\pi,k\}\\mathcal\{P\}\(Y^\{1:k\}\)\)\\bigr\)\\subseteq\\mathcal\{B\}\(H\)\. Then:

1. \(i\)everyPπ,k\(⋅∣⋅\)P^\{\\pi,k\}\(\\cdot\\mid\\cdot\)isσ​\(S\)\\sigma\(S\)\-measurable;
2. \(ii\)ifz:H→Zz\\colon H\\to Zis Borel and everyPπ,k\(⋅∣h\)P^\{\\pi,k\}\(\\cdot\\mid h\)equalsrπ,k​\(z​\(h\)\)r\_\{\\pi,k\}\(z\(h\)\)for Borelrπ,kr\_\{\\pi,k\}, thenσ​\(S\)⊆σ​\(z\)\\sigma\(S\)\\subseteq\\sigma\(z\)and there is a BorelrrwithS=r∘zS=r\\circ z;
3. \(iii\)σ​\(S\)\\sigma\(S\)is countably generated;
4. \(iv\)any two minimal sufficient sub\-σ\\sigma\-algebras agree modulo the null sets of any fixed prior onHH;
5. \(v\)S​\(H\)S\(H\)is analytic, and ifS​\(H\)S\(H\)is Borel — which holds whenΠ×K\\Pi\\times Kis finite, or whenHHis compact metrisable andSSis continuous — thenSSdescends to a Borel isomorphism ofH/∼Π,KH/\{\\sim\_\{\\Pi,K\}\}ontoS​\(H\)S\(H\)by the Lusin–Souslin theorem, so the quotient is standard Borel\. Without such a hypothesis the quotient is well defined as aσ\\sigma\-algebra but need not be standard Borel\.

###### Proof\.

\(i\) Each required law is a coordinate projection ofSS, henceσ​\(S\)\\sigma\(S\)\-measurable\. \(ii\) Ifzzis sufficient, writePπ,k\(⋅∣h\)=rπ,k\(z\(h\)\)P^\{\\pi,k\}\(\\cdot\\mid h\)=r\_\{\\pi,k\}\(z\(h\)\); countability ofΠ×K\\Pi\\times Kallows the coordinate maps to be assembled into one Borel product maprrwithS=r∘zS=r\\circ z, andσ​\(S\)=σ​\(r∘z\)⊆σ​\(z\)\\sigma\(S\)=\\sigma\(r\\circ z\)\\subseteq\\sigma\(z\)\. \(iii\) Each𝒫​\(Y1:k\)\\mathcal\{P\}\(Y^\{1:k\}\)is standard Borel, hence countably generated, and a countable product of countably generatedσ\\sigma\-algebras is countably generated\. \(iv\) Two minimal sufficientσ\\sigma\-algebras are each contained in the other modulo null sets by \(ii\), and mutual almost\-sure inclusion is almost\-sure equality\. \(v\)SSis Borel from a standard Borel space, so its image is analytic\. If the image is Borel andSSis injective on the quotient — which it is by construction — then Lusin–Souslin gives that the induced map is a Borel isomorphism onto its image\. ∎

###### Remark 3\.2a\.

The reason for stating \(v\) carefully is that the image of a Borel map need not be Borel, so the frequently made claim that a predictive quotient is “a standard Borel space” is not automatic\. What is always available is theσ\\sigma\-algebraσ​\(S\)\\sigma\(S\), and every statement in this paper that treats the quotient as a space should be read as carrying the hypothesis of \(v\)\. Statement \(iv\) is the sense in which the minimal sufficient representation is unique: uniqueness holds modulo null sets, not pointwise\.

###### Remark 3\.2b\(what is and is not new\)\.

For an uncontrolled process and a single trivial policy, Theorem 3\.2 reduces to the minimality of causal states in computational mechanics\[[49](https://arxiv.org/html/2609.11959#bib.bib36)\]; for a finite Markov decision process it reduces to the minimality of the stochastic bisimulation quotient\[[20](https://arxiv.org/html/2609.11959#bib.bib39),[17](https://arxiv.org/html/2609.11959#bib.bib40)\]\. The contribution here is not the minimality argument, which is standard, but the setting:SSis defined over a groupoid of interventions acting on a sheaf of typed local states, so that the same quotient is simultaneously constrained by prediction, by cross\-modal equivariance \([3](https://arxiv.org/html/2609.11959#S3.E3)\), by local\-to\-global descent \([12](https://arxiv.org/html/2609.11959#S4.E12)\) and by cross\-scale commutation \([16](https://arxiv.org/html/2609.11959#S4.E16)\)\. A categorical treatment of sufficiency in the same spirit, though without the interventional layer, is given byFritz \[[18](https://arxiv.org/html/2609.11959#bib.bib10)\]\. Section[3\.3](https://arxiv.org/html/2609.11959#S3.SS3)shows what this buys\.

![Refer to caption](https://arxiv.org/html/2609.11959v1/figures/fig2.png)Figure 2:The common space is the minimal action\-conditioned predictive state inferred from heterogeneous local observations\.
### 3\.3Interventional semantics, identifiability and the status of coordinates

Nothing in Section 3\.2 is causal so far\. The familyPPwas introduced as a family of regular conditional laws, and a conditional law computed from logged data is not in general the law that would obtain under an executed intervention\. Since the definition of space proposed here is interventional, this gap must be closed explicitly rather than absorbed into the notation\.

#### 3\.3\.1When the predictive state is interventional

###### Assumption I1\(interventional reading\)\.

Forπ∈Π\\pi\\in\\Piand a historyhh,Pπ,k\(⋅∣h\)P^\{\\pi,k\}\(\\cdot\\mid h\)denotes the law of the nextkkmultimodal observations underdo​\(π\)\\mathrm\{do\}\(\\pi\)— that is, under the arrows ofGGthatπ\\piselects, applied to the state reached byhh— and not the conditional law of observations in data generated by some other policy\.

###### Assumption I2\(sequential ignorability\)\.

The behaviour policybbthat generated the data selects each arrowgtg\_\{t\}as a measurable function of the observed historyhth\_\{t\}together with exogenous randomness that is independent of the latent state givenhth\_\{t\}\. Equivalently, no unobserved variable simultaneously drives the choice of intervention and the subsequent observations\.

###### Assumption I3\(positivity, or interventional excitation\)\.

There isε\>0\\varepsilon\>0such thatb​\(g∣h\)≥εb\(g\\mid h\)\\geq\\varepsilonfor every admissible arrowggat every historyhhin the support of the data\.

###### Lemma 3\.5\(identification by g\-computation\)\.

Under I2 and I3, for everyπ∈Π\\pi\\in\\Piandk∈Kk\\in Kthe interventional lawPπ,k\(⋅∣h\)P^\{\\pi,k\}\(\\cdot\\mid h\)is identified from the observational distribution by the g\-formula

Pπ,k​\(y1:k∣h\)=∫∏j=1kp​\(yj∣h,g1:j,y1:j−1\)​π​\(gj∣h,y1:j−1\)​d​g1:k\.P^\{\\pi,k\}\(y\_\{1:k\}\\mid h\)=\\int\\prod\_\{j=1\}^\{k\}p\\bigl\(y\_\{j\}\\mid h,\\,g\_\{1:j\},\\,y\_\{1:j\-1\}\\bigr\)\\,\\pi\\bigl\(g\_\{j\}\\mid h,\\,y\_\{1:j\-1\}\\bigr\)\\,\\mathrm\{d\}g\_\{1:k\}\.\(6\)Consequently the canonical predictive stateSSof \([4](https://arxiv.org/html/2609.11959#S3.E4)\) is estimable from logged data\. If I3 fails on a subfamily of arrows, thenSSis identified only for the policies supported bybb, and Theorem 3\.2 returns the minimal sufficient statistic for that subfamily, which is a strictly coarser space\.

###### Proof\.

Immediate from the sequential g\-computation identity for time\-varying treatments under sequential ignorability and positivity\[[46](https://arxiv.org/html/2609.11959#bib.bib23), §6\]; the groupoid structure supplies the composition of arrows but plays no role in the identification argument\. ∎

###### Remark 3\.5a\.

Lemma 3\.5 is where the philosophical claim of this paper acquires an operational edge\. Positivity is not a technical convenience\. An intervention that was never executed contributes nothing to the geometry, and a state distinction that only such an intervention could reveal is simply not part of the space that the data define\. The frequently voiced intuition that a city “has” a geometry which sensing merely uncovers is, on this account, an assertion that the relevant interventions have positive probability under the observation regime — an empirical claim, and a checkable one\.

#### 3\.3\.2Recovery of the latent topology

At any fixed probeUU, collect the observation maps intoH=\(h1,…,hm\):X→∏αYαH=\(h\_\{1\},\\dots,h\_\{m\}\)\\colon X\\to\\prod\_\{\\alpha\}Y\_\{\\alpha\}\. A family of modalities*jointly separates points*when, for everyx≠x′x\\neq x^\{\\prime\}, someα\\alphasatisfieshα​\(x\)≠hα​\(x′\)h\_\{\\alpha\}\(x\)\\neq h\_\{\\alpha\}\(x^\{\\prime\}\)\. No individual modality need be injective\.

###### Lemma 3\.6\(joint topological embedding\)\.

IfXXis compact Hausdorff, eachYαY\_\{\\alpha\}is Hausdorff, everyhαh\_\{\\alpha\}is continuous and the family\{hα\}\\\{h\_\{\\alpha\}\\\}jointly separates points, thenHHis a topological embedding ofXXinto∏αYα\\prod\_\{\\alpha\}Y\_\{\\alpha\}\. If moreover everyhαh\_\{\\alpha\}is equivariant with respect toΦg\\Phi\_\{g\}andTαgT^\{g\}\_\{\\alpha\}, then the latent action is conjugate onH​\(X\)H\(X\)to the product observation action∏αTαg\\prod\_\{\\alpha\}T^\{g\}\_\{\\alpha\}\.

###### Proof\.

Joint point separation makesHHinjective\. A continuous injection from a compact space into a Hausdorff space is a homeomorphism onto its image\. Equivariance givesH∘Φg=\(∏αTαg\)∘HH\\circ\\Phi\_\{g\}=\(\\prod\_\{\\alpha\}T^\{g\}\_\{\\alpha\}\)\\circ H, which is conjugacy after restricting the product action toH​\(X\)H\(X\)\. ∎

This is a statement about the true observation maps, not about a learned encoder, and by itself it settles nothing statistically\. The next subsection supplies what is actually needed\.

#### 3\.3\.3What interventions buy: identifiability up to the centraliser

Assume throughout:

\(E1\)XXis a compact connected Hausdorff space andΦ:G→Homeo⁡\(X\)\\Phi\\colon G\\to\\operatorname\{Homeo\}\(X\)is an action of a topological groupGGby homeomorphisms;

\(E2\)eachhαh\_\{\\alpha\}is continuous and the family jointly separates points;

\(E3\)equivariance holds,hα∘Φg=Tαg∘hαh\_\{\\alpha\}\\circ\\Phi\_\{g\}=T^\{g\}\_\{\\alpha\}\\circ h\_\{\\alpha\};

\(E4\)*interventional faithfulness*: for everyx≠x′x\\neq x^\{\\prime\}there existπ∈Π\\pi\\in\\Piandk∈Kk\\in KwithPπ,k\(⋅∣x\)≠Pπ,k\(⋅∣x′\)P^\{\\pi,k\}\(\\cdot\\mid x\)\\neq P^\{\\pi,k\}\(\\cdot\\mid x^\{\\prime\}\)\.

###### Theorem 3\.7\(identifiability up to the centraliser of the intervention group\)\.

Letz:∏αYα→Zz\\colon\\prod\_\{\\alpha\}Y\_\{\\alpha\}\\to Zbe continuous, predictively sufficient in the sense that everyPπ,kP^\{\\pi,k\}factors throughzz, and equivariant for an actionΨ\\PsiofGGonZZ, that isz∘\(∏αTαg\)=Ψg∘zz\\circ\(\\prod\_\{\\alpha\}T^\{g\}\_\{\\alpha\}\)=\\Psi\_\{g\}\\circ z\. Then under E1–E4 the compositeφ:=z∘H:X→Z\\varphi:=z\\circ H\\colon X\\to Zis a homeomorphism onto its image and satisfiesφ∘Φg=Ψg∘φ\\varphi\\circ\\Phi\_\{g\}=\\Psi\_\{g\}\\circ\\varphifor everygg\. Ifφ′\\varphi^\{\\prime\}is any second map with the same properties, thenτ:=φ′⁣−1∘φ\\tau:=\\varphi^\{\\prime\-1\}\\circ\\varphiis a homeomorphism ofXXcommuting with everyΦg\\Phi\_\{g\}, that is

τ∈ZHomeo⁡\(X\)​\(Φ​\(G\)\),\\tau\\ \\in\\ Z\_\{\\operatorname\{Homeo\}\(X\)\}\\bigl\(\\Phi\(G\)\\bigr\),\(7\)the centraliser of the action\. The latent space is therefore determined up to this centraliser, and not merely up to an arbitrary homeomorphism\.

###### Proof\.

By E1–E2 and Lemma 3\.6,HHis an embedding\. Supposez​\(H​\(x\)\)=z​\(H​\(x′\)\)z\(H\(x\)\)=z\(H\(x^\{\\prime\}\)\)\. Since everyPπ,kP^\{\\pi,k\}factors throughzz, we getPπ,k\(⋅∣x\)=Pπ,k\(⋅∣x′\)P^\{\\pi,k\}\(\\cdot\\mid x\)=P^\{\\pi,k\}\(\\cdot\\mid x^\{\\prime\}\)for allπ\\piandkk, sox=x′x=x^\{\\prime\}by E4; henceφ=z∘H\\varphi=z\\circ His injective\. It is continuous as a composite of continuous maps, and a continuous injection from a compact space into a Hausdorff space is a homeomorphism onto its image\. Equivariance follows by composition:

φ∘Φg=z∘H∘Φg=z∘\(∏αTαg\)∘H=Ψg∘z∘H=Ψg∘φ,\\varphi\\circ\\Phi\_\{g\}=z\\circ H\\circ\\Phi\_\{g\}=z\\circ\\Bigl\(\\prod\_\{\\alpha\}T^\{g\}\_\{\\alpha\}\\Bigr\)\\circ H=\\Psi\_\{g\}\\circ z\\circ H=\\Psi\_\{g\}\\circ\\varphi,\(8\)using E3 and the equivariance ofzz\. Finally, ifφ\\varphiandφ′\\varphi^\{\\prime\}are bothGG\-equivariant homeomorphisms onto the same image, thenτ=φ′⁣−1∘φ\\tau=\\varphi^\{\\prime\-1\}\\circ\\varphisatisfiesτ∘Φg=φ′⁣−1∘Ψg∘φ=Φg∘φ′⁣−1∘φ=Φg∘τ\\tau\\circ\\Phi\_\{g\}=\\varphi^\{\\prime\-1\}\\circ\\Psi\_\{g\}\\circ\\varphi=\\Phi\_\{g\}\\circ\\varphi^\{\\prime\-1\}\\circ\\varphi=\\Phi\_\{g\}\\circ\\tau, soτ\\taulies in the centraliser\. ∎

###### Corollary 3\.8\(space as a torsor; the exact status of coordinates\)\.

If the actionΦ\\Phiis simply transitive, so thatXXis a principal homogeneousGG\-space, thenZHomeo⁡\(X\)​\(Φ​\(G\)\)≅GZ\_\{\\operatorname\{Homeo\}\(X\)\}\(\\Phi\(G\)\)\\cong Gacting by right translations\. ConsequentlyXXis recovered as aGG\-torsor: the observation family together with the intervention family determines everything about the space except the choice of an origin and a frame, and that choice is exactly the residual ambiguity — no more and no less\.

###### Proof\.

Fixx0x\_\{0\}and identifyXXwithGGbyg↦Φg​x0g\\mapsto\\Phi\_\{g\}x\_\{0\}; this is a homeomorphism because the action is simply transitive\. Ifτ\\taucommutes with everyΦg\\Phi\_\{g\}andτ​\(x0\)=Φa​x0\\tau\(x\_\{0\}\)=\\Phi\_\{a\}x\_\{0\}, thenτ​\(Φg​x0\)=Φg​τ​\(x0\)=Φg​a​x0\\tau\(\\Phi\_\{g\}x\_\{0\}\)=\\Phi\_\{g\}\\tau\(x\_\{0\}\)=\\Phi\_\{ga\}x\_\{0\}, soτ\\tauis right translation byaa\. Conversely every right translation commutes with the left action\. ∎

Corollary 3\.8 is the precise form of the claim, made informally in Section 2\.2, that coordinate changes are representational while causal and metrical invariants are empirical\. A coordinate system is a section of the torsor\. Two observers who disagree about coordinates disagree by an element ofGGand about nothing else, and this is a theorem about the observational system rather than a stipulation\.

###### Proposition 3\.9\(quantitative version\)\.

Let\(X,d\)\(X,d\)be a compact metric space and define the*interventional separation modulus*

η\(τ\):=inf\{supπ,kTV\(Pπ,k\(⋅∣x\),Pπ,k\(⋅∣x′\)\):d\(x,x′\)≥τ\}\.\\eta\(\\tau\)\\;:=\\;\\inf\\Bigl\\\{\\,\\sup\_\{\\pi,k\}\\ \\mathrm\{TV\}\\bigl\(P^\{\\pi,k\}\(\\cdot\\mid x\),\\,P^\{\\pi,k\}\(\\cdot\\mid x^\{\\prime\}\)\\bigr\)\\ :\\ d\(x,x^\{\\prime\}\)\\geq\\tau\\,\\Bigr\\\}\.\(9\)Suppose an encoderzzand decoder familyQQsatisfysupπ,kTV\(Qπ,k\(⋅∣z\(H\(x\)\)\),Pπ,k\(⋅∣x\)\)≤ϵ\\sup\_\{\\pi,k\}\\mathrm\{TV\}\\bigl\(Q^\{\\pi,k\}\(\\cdot\\mid z\(H\(x\)\)\),\\,P^\{\\pi,k\}\(\\cdot\\mid x\)\\bigr\)\\leq\\epsilonfor everyxx\. Then every fibre ofz∘Hz\\circ Hhasdd\-diameter at mostτϵ:=inf\{τ\>0:η​\(τ\)\>2​ϵ\}\\tau\_\{\\epsilon\}:=\\inf\\\{\\tau\>0:\\eta\(\\tau\)\>2\\epsilon\\\}\. Exact recovery is the limiting caseϵ→0\\epsilon\\to 0withη​\(τ\)\>0\\eta\(\\tau\)\>0for everyτ\>0\\tau\>0, which is E4\.

###### Proof\.

Ifz​\(H​\(x\)\)=z​\(H​\(x′\)\)z\(H\(x\)\)=z\(H\(x^\{\\prime\}\)\)then for allπ\\piandkkthe triangle inequality for total variation givesTV\(Pπ,k\(⋅∣x\),Pπ,k\(⋅∣x′\)\)≤2ϵ\\mathrm\{TV\}\(P^\{\\pi,k\}\(\\cdot\\mid x\),P^\{\\pi,k\}\(\\cdot\\mid x^\{\\prime\}\)\)\\leq 2\\epsilon\. Wered​\(x,x′\)≥τd\(x,x^\{\\prime\}\)\\geq\\tauwithη​\(τ\)\>2​ϵ\\eta\(\\tau\)\>2\\epsilon, the definition ofη\\etawould force the left\-hand side to exceed2​ϵ2\\epsilon, a contradiction\. ∎

Proposition 3\.9 converts the predictive dimension of Section 3\.5 into a resolution statement: an encoder that predicts withinϵ\\epsilonresolves space to withinτϵ\\tau\_\{\\epsilon\}, and the functionϵ↦τϵ\\epsilon\\mapsto\\tau\_\{\\epsilon\}is the operating characteristic of the observational system\. It is estimable, becauseη\\etacan be lower\-bounded empirically by executing pairs of policies from nearby states\.

###### Remark 3\.10\(relation to interventional causal representation learning, and a sanity check\)\.

Theorem 3\.7 is the topological counterpart of the identifiability results proved for structural causal models under interventions\[[9](https://arxiv.org/html/2609.11959#bib.bib25),[1](https://arxiv.org/html/2609.11959#bib.bib26),[56](https://arxiv.org/html/2609.11959#bib.bib27),[37](https://arxiv.org/html/2609.11959#bib.bib28),[51](https://arxiv.org/html/2609.11959#bib.bib29),[57](https://arxiv.org/html/2609.11959#bib.bib30)\], which recover latent causal variables up to permutation and elementwise reparameterisation\. Here the residual ambiguity is expressed intrinsically as the centraliser of the intervention group; it specialises to permutation\-and\-scaling when the action is generated by independent single\-coordinate interventions\. The present statement is weaker in one respect that should be stated plainly: equivariance of the learned encoder is assumed in E\-form rather than derived, and in practice it is enforced by the penalty \([3](https://arxiv.org/html/2609.11959#S3.E3)\) and verified as a residual, not guaranteed\. The corresponding sanity check is instructive\. If the intervention family is removed —GGtrivial andΠ\\Pia singleton — then E4 fails for every pair of states not already separated by a single observation, the centraliser becomes all ofHomeo⁡\(X\)\\operatorname\{Homeo\}\(X\), and Theorem 3\.7 asserts nothing\. This is exactly the impossibility of unsupervised disentanglement without inductive bias\[[39](https://arxiv.org/html/2609.11959#bib.bib35)\]and the non\-identifiability of nonlinear ICA without auxiliary structure\[[31](https://arxiv.org/html/2609.11959#bib.bib33),[34](https://arxiv.org/html/2609.11959#bib.bib34)\]\. Partial recoveries are possible from multiple sufficiently distinct views or from spatial dependence\[[21](https://arxiv.org/html/2609.11959#bib.bib31),[24](https://arxiv.org/html/2609.11959#bib.bib32)\], but each such result buys identifiability by importing exactly the kind of auxiliary structure that an intervention family supplies here\. The theorem therefore does not merely tolerate interventions; it is empty without them, which is the strongest form in which the paper’s central thesis can be stated\.

Two caveats remain\. Compactness in E1 excludes unbounded latent spaces and is used only to convert continuous injectivity into homeomorphism; it can be replaced by properness ofφ\\varphi\. Joint separation in E2 is a strong assumption in any real deployment, and Section[9](https://arxiv.org/html/2609.11959#S9)treats its failure as the principal identifiability risk rather than as an edge case\.

### 3\.4Distance as action cost

For an arrowu:x→x′u\\colon x\\to x^\{\\prime\}, letA​\(u\)≥0A\(u\)\\geq 0be its cost\. Composition is executable succession\. Define

cA​\(x,x′\)=inf\{A​\(u\):u:x→x′​in​G\},inf∅:=\+∞\.c\_\{A\}\(x,x^\{\\prime\}\)\\;=\\;\\inf\\\{\\,A\(u\)\\ :\\ u\\colon x\\to x^\{\\prime\}\\ \\text\{in \}G\\,\\\},\\qquad\\inf\\emptyset:=\+\\infty\.\(10\)
###### Proposition 3\.4\(directed action geometry\)\.

IfA​\(idx\)=0A\(\\mathrm\{id\}\_\{x\}\)=0andA​\(v∘u\)≤A​\(u\)\+A​\(v\)A\(v\\circ u\)\\leq A\(u\)\+A\(v\), thencAc\_\{A\}is an extended directed quasi\-metric:cA​\(x,x\)=0c\_\{A\}\(x,x\)=0andcA​\(x,z\)≤cA​\(x,y\)\+cA​\(y,z\)c\_\{A\}\(x,z\)\\leq c\_\{A\}\(x,y\)\+c\_\{A\}\(y,z\)\. It becomes a metric only if finite mutual reachability, definiteness and reversal symmetry are added\.

This correction matters in cities\. Walking uphill and downhill, travelling with and against congestion, gaining and losing institutional access, and buying and selling in illiquid markets are generally asymmetric\. Calling every such cost a metric conceals the very geometry one wishes to study\. Finsler and directed network geometries are often the appropriate intermediate objects\[[3](https://arxiv.org/html/2609.11959#bib.bib7)\]\.

### 3\.5Predictive dimension

Dimension is not taken as the definition of space\. It quantifies the complexity of a chosen predictive representation\. Without regularity restrictions a real number can measurably encode pathological amounts of information, so the encoder class must be fixed\. Let𝒵d,L\\mathcal\{Z\}\_\{d,L\}be an admissible class ofdd\-dimensionalLL\-Lipschitz encoders, letDDbe a divergence, and letQQdenote a decoder family\.

Dpred\(ϵ;Π,K,L\)=inf\{d:∃z∈𝒵d,L,Q,supπ,k𝔼D\(Pπ,k\(⋅∣H\)∥Qπ,k\(⋅∣z\(H\)\)\)≤ϵ\}\.D\_\{\\mathrm\{pred\}\}\(\\epsilon;\\Pi,K,L\)=\\inf\\Bigl\\\{d\\ :\\ \\exists\\,z\\in\\mathcal\{Z\}\_\{d,L\},\\,Q,\\ \\ \\sup\_\{\\pi,k\}\\ \\mathbb\{E\}\\,D\\bigl\(P^\{\\pi,k\}\(\\cdot\\mid H\)\\,\\big\\\|\\,Q^\{\\pi,k\}\(\\cdot\\mid z\(H\)\)\\bigr\)\\leq\\epsilon\\Bigr\\\}\.\(11\)Equation \([11](https://arxiv.org/html/2609.11959#S3.E11)\) makes dimension task\-, scale\-, error\- and model\-relative\. Topological, Hausdorff, spectral, information and predictive dimensions may disagree because they answer different questions\. Their comparison is informative; their identification is not\.

## 4A sheaf\-valued geometry of the city

### 4\.1Stratified base and fibres

LetBBbe a finite regular cell complex or a Whitney\-stratified space representing rooms, air volumes, walls, façades, streets, rails, pipes, junctions and sensors\[[45](https://arxiv.org/html/2609.11959#bib.bib5),[15](https://arxiv.org/html/2609.11959#bib.bib6)\]\. Letπ:E→B×ℝ\\pi\\colon E\\to B\\times\\mathbb\{R\}be a bundle\-like projection\. The fibre over a cell and time is not one homogeneous coordinate vector but a typed product of local geometric, physical, mobility, social and economic state spaces\. Different strata can have different dimensions\. Singular junctions are not defects to be smoothed away; they are where transfer, impedance and governance conditions are imposed\.

A cellular sheafℱ\\mathcal\{F\}assigns a stalkℱ​\(σ\)\\mathcal\{F\}\(\\sigma\)to every cellσ\\sigmaand a restriction mapρσ≤τ:ℱ​\(σ\)→ℱ​\(τ\)\\rho\_\{\\sigma\\leq\\tau\}\\colon\\mathcal\{F\}\(\\sigma\)\\to\\mathcal\{F\}\(\\tau\)to every incidence relation, with functorial compatibility\. A section is a choice of local states\. It is global when all incident choices agree after restriction\. This is precisely the right logic for combining sensors, models and institutions whose local descriptions overlap but are not numerically identical\[[47](https://arxiv.org/html/2609.11959#bib.bib15),[25](https://arxiv.org/html/2609.11959#bib.bib16),[8](https://arxiv.org/html/2609.11959#bib.bib18)\]; a systematic account of cellular sheaves and cosheaves is given byCurry \[[13](https://arxiv.org/html/2609.11959#bib.bib11)\]\.

![Refer to caption](https://arxiv.org/html/2609.11959v1/figures/fig3.png)Figure 3:A city is represented by a stratified base, typed local stalks, restriction maps, cross\-scale coarsening and interventions that may change both state and gluing\.
### 4\.2Descent energy, the sheaf Laplacian and holonomy

For a finite cellular sheaf with inner products on stalks, letC0​\(B;ℱ\)C^\{0\}\(B;\\mathcal\{F\}\)be the space of0\-cochains andδℱ\\delta\_\{\\mathcal\{F\}\}the sheaf coboundary\. LetWWbe positive definite on incidence residuals\.

Lℱ=δℱ∗​W​δℱ,Edesc​\(z\)=‖W1/2​δℱ​z‖2=⟨z,Lℱ​z⟩\.L\_\{\\mathcal\{F\}\}=\\delta\_\{\\mathcal\{F\}\}^\{\*\}\\,W\\,\\delta\_\{\\mathcal\{F\}\},\\qquad E\_\{\\mathrm\{desc\}\}\(z\)=\\bigl\\\|W^\{1/2\}\\delta\_\{\\mathcal\{F\}\}z\\bigr\\\|^\{2\}=\\langle z,L\_\{\\mathcal\{F\}\}z\\rangle\.\(12\)
###### Proposition 4\.1\(consistency certificate\)\.

Edesc​\(z\)=0E\_\{\\mathrm\{desc\}\}\(z\)=0if and only ifzzis a global section\. Equivalently,ker⁡Lℱ=ker⁡δℱ\\ker L\_\{\\mathcal\{F\}\}=\\ker\\delta\_\{\\mathcal\{F\}\}\.

###### Proof\.

Positive definiteness ofWWgivesEdesc​\(z\)=0E\_\{\\mathrm\{desc\}\}\(z\)=0exactly whenδℱ​z=0\\delta\_\{\\mathcal\{F\}\}z=0\. The kernel of the coboundary is the space of global compatible sections, and⟨z,Lℱ​z⟩=⟨δℱ​z,W​δℱ​z⟩\\langle z,L\_\{\\mathcal\{F\}\}z\\rangle=\\langle\\delta\_\{\\mathcal\{F\}\}z,W\\delta\_\{\\mathcal\{F\}\}z\\rangle\. ∎

The scalar graph Laplacian is recovered when every stalk is one\-dimensional and every restriction is the identity\. Nontrivial restriction maps allow neighbouring cells to translate between coordinate frames, sensor types, institutional categories or learned feature bases\. Neural sheaf diffusion exploits precisely this additional geometry\[[8](https://arxiv.org/html/2609.11959#bib.bib18)\]\.

Proposition 4\.1 as stated is an identity rather than a discovery, and on its own it would not justify the sheaf machinery: the same certificate for a graph Laplacian with identity restrictions is equally immediate\. What distinguishes a sheaf from a multilayer network is that non\-identity restriction maps carry*holonomy*, and holonomy is an obstruction that no identity\-restriction model can represent\.

###### Proposition 4\.2\(holonomy obstruction\)\.

Letℱ\\mathcal\{F\}be a cellular sheaf on a connected graphBBwith stalksℝn\\mathbb\{R\}^\{n\}and restriction maps inO​\(n\)O\(n\), so that\(δℱ​z\)e=ρj,e​zj−ρi,e​zi\(\\delta\_\{\\mathcal\{F\}\}z\)\_\{e\}=\\rho\_\{j,e\}z\_\{j\}\-\\rho\_\{i,e\}z\_\{i\}for each edgee=\(i,j\)e=\(i,j\)\. Fix a base vertexv0v\_\{0\}and letHol⁡\(ℱ,v0\)⊆O​\(n\)\\operatorname\{Hol\}\(\\mathcal\{F\},v\_\{0\}\)\\subseteq O\(n\)be the group generated by the holonomies of all cycles based atv0v\_\{0\}\. Then

dimker⁡Lℱ=dimFix⁡\(Hol⁡\(ℱ,v0\)\),\\dim\\ker L\_\{\\mathcal\{F\}\}\\;=\\;\\dim\\operatorname\{Fix\}\\bigl\(\\operatorname\{Hol\}\(\\mathcal\{F\},v\_\{0\}\)\\bigr\),\(13\)the dimension of the subspace ofℝn\\mathbb\{R\}^\{n\}fixed pointwise by the holonomy group\. In particular, with identity restrictions the holonomy is trivial anddimker⁡Lℱ=n\\dim\\ker L\_\{\\mathcal\{F\}\}=nper connected component, which is the graph Laplacian case; whereas a single cycle with nontrivial holonomy can forcedimker⁡Lℱ=0\\dim\\ker L\_\{\\mathcal\{F\}\}=0, so that the sheaf admits no nonzero global section even though the graph is connected and every local datum is individually consistent\.

###### Proof\.

A0\-cochain lies inker⁡δℱ\\ker\\delta\_\{\\mathcal\{F\}\}precisely when it is parallel, that iszj=ρj,e−1​ρi,e​ziz\_\{j\}=\\rho\_\{j,e\}^\{\-1\}\\rho\_\{i,e\}z\_\{i\}along every edge\. Parallel transport along a path is then determined by the restriction maps, so a parallel section is determined by its value atv0v\_\{0\}, and such a value extends consistently if and only if it is fixed by the holonomy of every cycle atv0v\_\{0\}\. The assignmentz↦z​\(v0\)z\\mapsto z\(v\_\{0\}\)is therefore a linear isomorphism fromker⁡δℱ\\ker\\delta\_\{\\mathcal\{F\}\}ontoFix⁡\(Hol⁡\(ℱ,v0\)\)\\operatorname\{Fix\}\(\\operatorname\{Hol\}\(\\mathcal\{F\},v\_\{0\}\)\), andker⁡Lℱ=ker⁡δℱ\\ker L\_\{\\mathcal\{F\}\}=\\ker\\delta\_\{\\mathcal\{F\}\}by Proposition 4\.1\. ∎

This is the sheaf\-theoretic form of the connection Laplacian of vector diffusion maps and of the graph connection Laplacian\[[50](https://arxiv.org/html/2609.11959#bib.bib19),[2](https://arxiv.org/html/2609.11959#bib.bib20)\], and it is what gives Proposition 4\.1 empirical content\. A numerical instance is reported in Section[5\.1](https://arxiv.org/html/2609.11959#S5.SS1): on an eight\-cycle withℝ2\\mathbb\{R\}^\{2\}stalks and rotation restrictions, the graph Laplacian has a two\-dimensional kernel regardless of the data, the sheaf with trivial holonomy also has a two\-dimensional kernel, and the sheaf with holonomy angle0\.700\.70has kernel dimension zero with spectral gap7\.7×10−37\.7\\times 10^\{\-3\}\. The same experiment exhibits the converse failure mode: a field that is genuinely compatible with the correct sheaf receives descent energy1\.0×10−161\.0\\times 10^\{\-16\}, while the identity\-restriction graph model assigns it energy1\.311\.31and so reports a spurious inconsistency\.

The urban reading is direct\. Suppose a vector\-valued layer — wind velocity, acoustic intensity gradient, pedestrian flow — is recorded in local frames aligned to street orientation, as is normal practice\. The restriction maps across an incidence are the rotations between adjacent street orientations, and a closed circuit of streets whose orientations do not compose to the identity has nontrivial holonomy\. A model that averages raw local vectors, which is what a graph Laplacian does, will then report systematic residuals at exactly those circuits, and will do so whether or not the physical field is compatible\. The sheaf model makes a falsifiable prediction here: the residual of the identity\-restriction model at a circuit should scale with the holonomy angle of that circuit, and should vanish for the sheaf model\. Section[6\.3](https://arxiv.org/html/2609.11959#S6.SS3)lists this as Experiment F\.

###### Proposition 4\.3\(restriction maps are identifiable, and the constraint removes an inconsistency\)\.

Suppose paired local states\{\(zi\(m\),zj\(m\)\)\}m=1M\\\{\(z\_\{i\}^\{\(m\)\},z\_\{j\}^\{\(m\)\}\)\\\}\_\{m=1\}^\{M\}are observed across an incidence withzj=ρ​zi\+noisez\_\{j\}=\\rho z\_\{i\}\+\\text\{noise\}andρ∈O​\(n\)\\rho\\in O\(n\)\. Then the orthogonal Procrustes estimatorρ^=U​V𝖳\\hat\{\\rho\}=UV^\{\\mathsf\{T\}\}, where∑mzj\(m\)​\(zi\(m\)\)𝖳=U​Σ​V𝖳\\sum\_\{m\}z\_\{j\}^\{\(m\)\}\(z\_\{i\}^\{\(m\)\}\)^\{\\mathsf\{T\}\}=U\\Sigma V^\{\\mathsf\{T\}\}, is consistent\. Moreover, whenziz\_\{i\}is itself measured with independent error of varianceσ2\\sigma^\{2\}— the errors\-in\-variables regime that is the norm in sensor networks — the unconstrained least\-squares estimator is*inconsistent*, converging to the attenuated mapλ​\(λ\+σ2\)−1​ρ\\lambda\(\\lambda\+\\sigma^\{2\}\)^\{\-1\}\\rhowithλ\\lambdathe per\-component signal variance, whereas the projection ontoO​\(n\)O\(n\)discards precisely this multiplicative attenuation and remains consistent\.

###### Proof\.

Consistency of the Procrustes estimator follows from the strong law applied to the cross\-moment matrix together with continuity of the polar factor at a nonsingular limit\. For the second claim, the least\-squares limit is\(Σx​x\+σ2​I\)−1​Σx​y=λ​\(λ\+σ2\)−1​ρ\(\\Sigma\_\{xx\}\+\\sigma^\{2\}I\)^\{\-1\}\\Sigma\_\{xy\}=\\lambda\(\\lambda\+\\sigma^\{2\}\)^\{\-1\}\\rhofor isotropic signal, which is a positive scalar multiple ofρ\\rhoand therefore has the same polar factor; theO​\(n\)O\(n\)projection of the attenuated map isρ\\rhoexactly\. ∎

The practical consequence is stronger than regularisation\. The equivariance and descent constraints are usually presented as priors that trade bias for variance, and would then be expected to help only in the small\-sample regime\. Proposition 4\.3 says something else: under measurement error in the inputs, the unconstrained estimator has an asymptotic bias floor that no amount of data removes, and the structural constraint removes it\. The numerical check in Section[5\.1](https://arxiv.org/html/2609.11959#S5.SS1)confirms this, with the unconstrained error plateauing at the predicted floor1\.9×10−41\.9\\times 10^\{\-4\}atσ=0\.1\\sigma=0\.1and2\.7×10−32\.7\\times 10^\{\-3\}atσ=0\.2\\sigma=0\.2while the constrained error continues to decrease with sample size\.

### 4\.3Coupled field and flow dynamics

Letztz\_\{t\}be a sheaf\-valued urban state andutu\_\{t\}an intervention\. A deliberately general evolution law is

M​\(zt\)​z˙t=−Lℱ​\(zt\)​zt\+N​\(zt\)\+B​\(zt\)​ut\+st\+ξt\.M\(z\_\{t\}\)\\,\\dot\{z\}\_\{t\}\\;=\\;\-L\_\{\\mathcal\{F\}\}\(z\_\{t\}\)\\,z\_\{t\}\+N\(z\_\{t\}\)\+B\(z\_\{t\}\)u\_\{t\}\+s\_\{t\}\+\\xi\_\{t\}\.\(14\)The block operatorLℱL\_\{\\mathcal\{F\}\}carries diffusion and compatibility terms;NNcontains nonlinear transport, reaction, behavioural response and market dynamics;BBspecifies actuators;ssandξ\\xiare forcing and uncertainty\. Physical conservation laws may occupy selected blocks\. Social and economic layers need not be forced into mass conservation\. Their locality is instead encoded by stalks, restrictions, reachability and intervention response\.

do​\(u\):\(ℱ,ρ,Φ\)⟼\(ℱu,ρu,Φu\)\.\\mathrm\{do\}\(u\)\\ \\colon\\ \(\\mathcal\{F\},\\ \\rho,\\ \\Phi\)\\ \\longmapsto\\ \(\\mathcal\{F\}\_\{u\},\\ \\rho\_\{u\},\\ \\Phi\_\{u\}\)\.\(15\)Equation \([15](https://arxiv.org/html/2609.11959#S4.E15)\) allows an intervention to change not only the current state but the coupling architecture\. Opening a door alters acoustic, thermal and aerodynamic transfer maps\. Closing a road alters mobility restrictions\. A zoning or pricing rule may alter economic accessibility without moving any wall\. This is where causal intervention and sheaf gluing meet\.

### 4\.4Cross\-scale compatibility

LetRℓ→mR\_\{\\ell\\to m\}coarsen a state from levelℓ\\elltomm\. Exact commutation is exceptional; a useful model measures the defect\.

Escale​\(ℓ,m;u\)=𝔼z​‖Rℓ→m​Φuℓ​\(z\)−Φum​\(Rℓ→m​z\)‖2\.E\_\{\\mathrm\{scale\}\}\(\\ell,m;u\)\\;=\\;\\mathbb\{E\}\_\{z\}\\bigl\\\|R\_\{\\ell\\to m\}\\,\\Phi^\{\\ell\}\_\{u\}\(z\)\\;\-\\;\\Phi^\{m\}\_\{u\}\\bigl\(R\_\{\\ell\\to m\}z\\bigr\)\\bigr\\\|^\{2\}\.\(16\)Small residual means that acting and then aggregating nearly agrees with aggregating and then acting\. Large residual identifies a scale at which omitted heterogeneity matters\. The city is therefore not assumed to possess one privileged resolution\. It is an inverse or multiresolution system whose transition maps are themselves testable\.

## 5Audit of the core formulae

The following audit separates statements that are identities, statements that require hypotheses and statements that should be estimated as residuals\. This prevents suggestive notation from being mistaken for a theorem\.

Table 1:Type and falsifiability audit of the principal formulae\.ObjectMathematical statusConditionsFailure mode / testPredictive state \([4](https://arxiv.org/html/2609.11959#S3.E4)\)Well\-typed product of probability lawsStandard Borel histories and futures; regular conditional laws existUncountable policy classes need extra measurable structure; check countability ofΠ×K\\Pi\\times K\.Predictive quotient \([5](https://arxiv.org/html/2609.11959#S3.E5)\)Minimal sufficient statisticEvery required future law is indexed inSS; null sets treated modulo a priorA restricted policy family may merge states that another intervention separates\.Joint embedding \(Lemma 3\.6\)Conditional theoremCompact Hausdorff latent space; continuous maps into Hausdorff targets; joint separationUnknown nonlinear encoders remain statistically non\-identifiable without further structure\.Equivariance \([3](https://arxiv.org/html/2609.11959#S3.E3)\)Loss / defect, not assumed equalityBoth compositions have the same source and target; modality metrics declaredCoordinate mismatch or misaligned action timestamps increase the residual\.Interventional identification \(Lemma 3\.5\)Conditional theoremSequential ignorability \(I2\) and positivity \(I3\); admissible arrows have behaviour probability bounded belowUnobserved confounding or zero\-probability interventions make the quotient statistical, not causal; test by policy\-support diagnostics and negative controls\.Centraliser identifiability \(Theorem 3\.7\)Conditional theoremCompact connected latent space; joint separation; equivariant learned encoder; interventional faithfulness \(E4\)Removing interventions makes the centraliser all ofHomeo⁡\(X\)\\operatorname\{Homeo\}\(X\)and the statement vacuous, recovering\[[39](https://arxiv.org/html/2609.11959#bib.bib35)\]; test E4 by the separation modulus of Proposition 3\.9\.Action cost \([10](https://arxiv.org/html/2609.11959#S3.E10)\)Directed extended quasi\-metricIdentity has zero cost; composition is subadditiveWithout reversal symmetry it is not a metric; test asymmetry explicitly\.Predictive dimension \([11](https://arxiv.org/html/2609.11959#S3.E11)\)Model\-relative complexity indexEncoder regularity, policy set, horizon, divergence and tolerance all declaredWithout regularity, pathological encodings make dimensional claims vacuous\.Sheaf energy \([12](https://arxiv.org/html/2609.11959#S4.E12)\)Exact consistency certificateFinite cellular sheaf; positive\-definiteWW; correct restriction mapsA wrong sheaf can report false consistency; validate restrictions first\.Holonomy obstruction \([13](https://arxiv.org/html/2609.11959#S4.E13)\)Exact theoremCellular sheaf on a connected graph withO​\(n\)O\(n\)restriction mapsMisspecified restriction maps produce spurious or missing obstructions; validate by Procrustes recovery \(Proposition 4\.3\) before interpreting the kernel\.Cross\-scale law \([16](https://arxiv.org/html/2609.11959#S4.E16)\)Empirical residualCoarsening map and norm specified; comparable interventions at both levelsHigh\-frequency or singular effects can make aggregation and action fail to commute\.Projective limit \(Theorem 8\.4\)Exact theorem; saturation clause is empiricalDirected context poset; countable cofinal chain for the Borel clauseSaturation \(iv\) may fail: a curve that never flattens means no finite instrument set determines the space\. ReportEsatE\_\{\\mathrm\{sat\}\}and separate noise\-driven from refinement\-driven decrease\.### 5\.1Reproducible synthetic study under noise

To check the algebra without disguising a toy model as urban evidence, consider a latent state\(x,y\)\(x,y\)on the two\-torus\. Vision observesv​\(x\)=\(cos⁡x,sin⁡x\)v\(x\)=\(\\cos x,\\sin x\), audition observesa​\(y\)=\(cos⁡y,sin⁡y\)a\(y\)=\(\\cos y,\\sin y\), and touch observest​\(x,y\)=\(cos⁡\(x\+y\),sin⁡\(x\+y\)\)t\(x,y\)=\(\\cos\(x\+y\),\\sin\(x\+y\)\); a fourth modalityw​\(x,y\)=\(cos⁡\(x−y\),sin⁡\(x−y\)\)w\(x,y\)=\(\\cos\(x\-y\),\\sin\(x\-y\)\)is held in reserve for the refinement test\. Each view is partial; the visual–auditory pair separates points\. Six translation actions are sampled\. For every action the observation transformation is a planar rotation, so exact equivariance is known:

v​\(x\+Δ​x\)=RΔ​x​v​\(x\),a​\(y\+Δ​y\)=RΔ​y​a​\(y\),t​\(x\+Δ​x,y\+Δ​y\)=RΔ​x\+Δ​y​t​\(x,y\)\.v\(x\+\\Delta x\)=R\_\{\\Delta x\}\\,v\(x\),\\qquad a\(y\+\\Delta y\)=R\_\{\\Delta y\}\\,a\(y\),\\qquad t\(x\+\\Delta x,\\,y\+\\Delta y\)=R\_\{\\Delta x\+\\Delta y\}\\,t\(x,y\)\.\(17\)
The earlier version of this test was run noiselessly, which made every reported figure either machine precision or an artefact of the fitting protocol, and left an inversion in the ablation ordering that we now resolve\. All experiments below therefore add isotropic Gaussian observation noise of standard deviationσ\\sigmato every modality at both the input and the target, and every entry is reported as the mean and standard deviation over twenty seeds derived from the base seed20260731\. Training uses12,00012\{,\}000states and testing5,0005\{,\}000held\-out states per seed\. Linear least squares is fitted separately for each action, and joint features include the bilinear visual–auditory products needed to reconstruct touch\.

Six checks are reported\. \(a\) Cross\-modal prediction under the five representations of Table[2](https://arxiv.org/html/2609.11959#S5.T2)\. \(b\) The value of the equivariance constraint, comparing an unconstrained least\-squares operator with its projection ontoO​\(2\)O\(2\), across sample sizes and noise levels\. \(c\) The holonomy obstruction of Proposition 4\.2 on an eight\-cycle sheaf\. \(d\) Recovery of restriction maps by Procrustes under noise, per Proposition 4\.3\. \(e\) The cross\-scale residual \([16](https://arxiv.org/html/2609.11959#S4.E16)\) as a function of the wavenumber of the underlying field\. \(f\) The saturation of the projective system of Theorem 8\.4, obtained by refining the observational context from\{v\}\\\{v\\\}to\{v,a\}\\\{v,a\\\}to\{v,a,t\}\\\{v,a,t\\\}to\{v,a,t,w\}\\\{v,a,t,w\\\}\.

Table 2:Held\-out prediction error in the synthetic cross\-modal experiment, with observation noiseσ=0\.05\\sigma=0\.05, mean±\\pmstandard deviation over 20 seeds\.![Refer to caption](https://arxiv.org/html/2609.11959v1/figures/fig4.png)Figure 4:Numerical checks under noise, twenty seeds\. \(a\) Exact recovery requires both joint observation and correctly aligned actions; the action\-blind and misaligned conditions are statistically indistinguishable\. \(b\) TheO​\(2\)O\(2\)constraint removes the errors\-in\-variables bias floor of the unconstrained estimator \(dotted lines, predicted analytically\), so the benefit persists asymptotically rather than vanishing with sample size\. \(c\) A holonomy obstruction reduces the sheaf kernel to zero while the graph Laplacian kernel is insensitive to it\. \(d\) Restriction maps are recoverable from paired data with error linear in the noise\. \(e\) Cross\-scale commutation degrades as the field approaches the coarse Nyquist wavenumber\. \(f\) The projective system of predictive quotients saturates atc0=\{v,a\}c\_\{0\}=\\\{v,a\\\}; the further decrease under noise reflects variance reduction, not refinement, as the noiseless control shows\.Two results deserve comment because they correct the earlier presentation\. First, in Table[2](https://arxiv.org/html/2609.11959#S5.T2)the action\-blind and misaligned\-action conditions are statistically indistinguishable, with a Welch statistic of−1\.94\-1\.94over twenty seeds\. This is the expected outcome and not a defect: both conditions marginalise over the action set, one by omitting it and one by destroying its correspondence with the observations, so both estimate the same average operator\. The earlier report of a small advantage for the misaligned condition was seed noise on a single run\. The informative contrast is between either of these and the correctly aligned condition, which is fourteen times better and approaches the irreducible noise floorσ2=2\.5×10−3\\sigma^\{2\}=2\.5\\times 10^\{\-3\}\.

Second, the refinement test in panel \(f\) supports the saturation clause of Theorem 8\.4 but only in the noiseless control, and the distinction matters\. Under noise the residual falls by3\.7×10−13\.7\\times 10^\{\-1\}from\{v\}\\\{v\\\}to\{v,a\}\\\{v,a\\\}and then by roughly1×10−31\\times 10^\{\-3\}at each further refinement; those later reductions are consistent with variance reduction from redundant noisy views rather than with genuine refinement of the quotient\. In the noiseless control the residual falls to below10−2910^\{\-29\}at\{v,a\}\\\{v,a\\\}and does not fall further, which is the stationarity condition of Theorem 8\.4\(iv\) with sufficient contextc0=\{v,a\}c\_\{0\}=\\\{v,a\\\}\. The methodological lesson is that a saturation curve measured under noise cannot by itself certify sufficiency, and an empirical protocol must separate the two effects by holding predictive capacity fixed while adding modalities\.

The result establishes internal consistency only\. It does not show that a city admits the chosen sheaf, that its policies are known, or that a robot will learn the correct latent state\. Those claims require field and embodied experiments designed to break the model\.

## 6From embodied to em\-spaced intelligence

### 6\.1Two meanings of ESI

Two recent uses of the initials ESI must be distinguished\.*Embodied Spatial Intelligence*places the agent inside a perception\-action loop and asks it to acquire evidence through movement\. ESI\-Bench makes this requirement explicit and reports that active exploration outperforms passive observation, while poor action selection causes cascading perceptual failure\[[30](https://arxiv.org/html/2609.11959#bib.bib63)\]\.*Em\-Spaced Intelligence*places intelligence partly in the spatial environment itself; the term is formed by analogy with*embodied*, intelligence being em\-bedded in and partly constituted by the instrumented space, and it is unrelated to the typographic em space\. In that programme, Artificially Evolved Spatial Organisms combine multimodal foundation models, graph neural networks and non\-Euclidean manifold geometry, and a Dynamic Space Protocol mediates the co\-evolution of robots and urban space\[[63](https://arxiv.org/html/2609.11959#bib.bib62)\]\. The overlap with embodied spatial intelligence is substantive, but the ontological emphasis differs, and the framework of this paper is neutral between them: both are instances of Definition 6\.1 with different allocations of sensing and actuation between the two bodies\.

###### Definition 6\.1\(em\-spaced intelligent system\)\.

An em\-spaced intelligent system is a coupled pair\(A,E\)\(A,E\)of a mobile agentAAand a persistent spatial bodyEE, together with a shared predictive stateSS, bidirectional observation maps, agent actionsπ\\pi, environmental interventionsuu, and governance constraintsΓ\\Gamma\. BothAAandEEcan alter the future field of admissible interaction\.

![Refer to caption](https://arxiv.org/html/2609.11959v1/figures/fig5.png)Figure 5:Dual embodiment\. The mobile body and the spatial body share a predictive core but retain distinct sensors, actuators, memories and constraints\.
### 6\.2Learning and control objective

A trainable model can combine prediction with structural penalties rather than treating geometry as a purely visual latent variable\.

ℒ=ℒpred\+λeq​Eeq\+λdesc​Edesc\+λscale​Escale\+λcf​ℒcf\+λsafe​CΓ\.\\mathcal\{L\}\\;=\\;\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\lambda\_\{\\mathrm\{eq\}\}E\_\{\\mathrm\{eq\}\}\+\\lambda\_\{\\mathrm\{desc\}\}E\_\{\\mathrm\{desc\}\}\+\\lambda\_\{\\mathrm\{scale\}\}E\_\{\\mathrm\{scale\}\}\+\\lambda\_\{\\mathrm\{cf\}\}\\mathcal\{L\}\_\{\\mathrm\{cf\}\}\+\\lambda\_\{\\mathrm\{safe\}\}C\_\{\\Gamma\}\.\(18\)The first term scores future multimodal prediction\.EeqE\_\{\\mathrm\{eq\}\}tests action covariance, in the sense made systematic by the geometric deep learning programme\[[10](https://arxiv.org/html/2609.11959#bib.bib52)\];EdescE\_\{\\mathrm\{desc\}\}tests local\-to\-global compatibility;EscaleE\_\{\\mathrm\{scale\}\}tests aggregation;ℒcf\\mathcal\{L\}\_\{\\mathrm\{cf\}\}scores counterfactual interventions;CΓC\_\{\\Gamma\}penalises unsafe or institutionally inadmissible acts\. A coupled planner then chooses both mobile and environmental actions:

\(π∗,u∗\)=arg⁡minπ,u⁡𝔼​\[∑τc​\(Sτ,πτ,uτ\)\|St\]subject to\(πτ,uτ\)∈Γ​\(Sτ\)\.\(\\pi^\{\*\},u^\{\*\}\)\\;=\\;\\arg\\min\_\{\\pi,u\}\\ \\mathbb\{E\}\\Bigl\[\\ \\sum\_\{\\tau\}c\(S\_\{\\tau\},\\pi\_\{\\tau\},u\_\{\\tau\}\)\\ \\Big\|\\ S\_\{t\}\\Bigr\]\\quad\\text\{subject to\}\\quad\(\\pi\_\{\\tau\},u\_\{\\tau\}\)\\in\\Gamma\(S\_\{\\tau\}\)\.\(19\)This is more than a robot using a smart building as a sensor\. The building can change illumination, ventilation, access, signage, signalling and information disclosure; the robot can move, inspect and manipulate\. Each intervention changes what the other can learn and do\. The object of intelligence is their coupled law\.

### 6\.3Experimental programme

##### Experiment A — cross\-modal intervention prediction\.

Instrument a building with vision, acoustics, temperature, airflow, occupancy and door state\. Hold out complete intervention types, not random frames\. Compare the full model with modality concatenation, action\-blind prediction and single\-layer graph baselines\.

##### Experiment B — stratified transfer\.

Train at room and building scales, then test at floor, block and district scales\. ReportEscaleE\_\{\\mathrm\{scale\}\}, held\-out likelihood, missing\-sensor reconstruction and the stability of recovered topological features\.

##### Experiment C — dual embodiment\.

Compare robot\-only control, environment\-only automation and joint\(π,u\)\(\\pi,u\)control on navigation, inspection, emergency egress and human\-robot coexistence\. Report task success, energy, intervention regret, safety violations and calibration under distribution shift\.

##### Experiment D — social and economic geometry\.

Use road closure, timetable, price or access\-rule changes as explicit interventions\. Test whether the inferred reachability geometry predicts mobility and exchange beyond Euclidean distance while remaining stable under alternative demographic or institutional partitions\.

##### Experiment E — context saturation\.

Instrument a building with an ordered family of modalities and add them one at a time, holding predictive capacity and training budget fixed\. Report the saturation residualEsat​\(c,c′\)E\_\{\\mathrm\{sat\}\}\(c,c^\{\\prime\}\)of Corollary 8\.6 as a curve\. The claim that the building has a well\-defined spatial description relative to this instrument set is the claim that the curve saturates; the modality at which it stops falling identifies a sufficient context, and any later fall exhibits a modality carrying genuinely new spatial information\. As Section[5\.1](https://arxiv.org/html/2609.11959#S5.SS1)shows, the noise\-driven and refinement\-driven components of the curve must be separated before sufficiency can be asserted\.

##### Experiment F — holonomy\.

Identify circuits in the street network whose local frames do not compose to the identity, and record a vector\-valued layer in local frames along them\. Proposition 4\.2 predicts that a graph model with identity restrictions shows residuals scaling with the circuit holonomy angle while the sheaf model with Procrustes\-estimated restrictions does not\. This is the cleanest available discriminator between the sheaf formalism and a multilayer network, because it concerns a structural obstruction rather than a difference in fitting capacity\.

Spatial\-intelligence benchmarks already reveal weaknesses in orientation, view selection and interaction\-aware 3D representation\[[30](https://arxiv.org/html/2609.11959#bib.bib63),[59](https://arxiv.org/html/2609.11959#bib.bib64),[65](https://arxiv.org/html/2609.11959#bib.bib65)\]\. The proposed experiments add two demands that current benchmarks rarely combine: compatibility across urban scales and interventions performed by the environment itself\.

## 7Euclidean, non\-Euclidean and directional geometry

Euclidean and non\-Euclidean descriptions need not be rival ontologies\. A building survey may use Euclidean coordinates while acoustic travel time, wind transport, social access and economic exchange define distinct effective geometries on the same stratified base\. Smoothly varying positive\-definite tensors interpolate among Riemannian geometries; Finsler costs express directional asymmetry; graphs and sheaves handle discontinuities and singular junctions; hyperbolic embeddings accommodate hierarchical reachability at low distortion\[[41](https://arxiv.org/html/2609.11959#bib.bib53)\]; metric\-measure and optimal\-transport structures compare distributions and evolving mass\[[22](https://arxiv.org/html/2609.11959#bib.bib2),[3](https://arxiv.org/html/2609.11959#bib.bib7),[55](https://arxiv.org/html/2609.11959#bib.bib8),[5](https://arxiv.org/html/2609.11959#bib.bib54),[12](https://arxiv.org/html/2609.11959#bib.bib9)\]\. The configurational tradition in urban morphology has long made the same methodological point in a different vocabulary, treating accessibility rather than metric distance as the primary spatial variable\[[27](https://arxiv.org/html/2609.11959#bib.bib55),[29](https://arxiv.org/html/2609.11959#bib.bib56),[28](https://arxiv.org/html/2609.11959#bib.bib59),[6](https://arxiv.org/html/2609.11959#bib.bib57),[62](https://arxiv.org/html/2609.11959#bib.bib58)\], and recent work in that tradition has begun to couple it to generative and manifold\-based models of urban form\[[60](https://arxiv.org/html/2609.11959#bib.bib60),[61](https://arxiv.org/html/2609.11959#bib.bib61)\]\. Continuity of description is possible where the structure varies continuously, but topology\-changing interventions and strata transitions may be genuinely discontinuous\.

Directional geometry deserves a brief note, stated conservatively\. Some urban layers are governed by directional rather than isotropic propagation: acoustic, electromagnetic and radiative transfer concentrate along thin tubes, and the natural obstruction theory for such layers is harmonic\-analytic\. Fefferman’s ball\-multiplier counterexample shows that directional decompositions cannot be summed naively\[[16](https://arxiv.org/html/2609.11959#bib.bib12)\], and the recent resolution of the Kakeya set conjecture in three dimensions sharpens the control of unions of tubes\[[23](https://arxiv.org/html/2609.11959#bib.bib13),[58](https://arxiv.org/html/2609.11959#bib.bib14)\]\. We record the connection as a layer\-specific prior and nothing more\. Kakeya estimates supply inequalities and obstruction mechanisms; they do not determine a sensing geometry, do not by themselves fix a minimum number of tomographic angles or beams, and do not supply a reconstruction algorithm\. Converting them into a stability guarantee requires a specified forward operator, noise model and discretisation, which we do not develop here and flag as future work rather than as a result of this paper\.

## 8What, then, is space?

Space is best defined neither as dimension alone nor as an arbitrary relation\. Dimension states how many independent degrees of freedom a chosen representation needs\. Relation states which distinctions and transitions matter\. The latter is logically prior: without adjacency, reachability, compatibility and admissible transformation, a dimension is a number without spatial content\. Yet not every relation is spatial\. A relation becomes spatial when it is localisable, composable, testable by intervention, stable enough to support counterfactual prediction and compatible across overlapping probes\.

###### Definition 8\.1\(space, provisional form\)\.

For a specified family of probes, modalities, actions and scales, space is the minimal local\-to\-global relational object that makes the intervention\-conditioned transformations of those modalities jointly representable and predictively sufficient\.

Definition 8\.1 invites an obvious objection, and the objection must be met rather than deflected\. If space is defined relative to a specified family of probes, modalities and actions, then changing the family changes the space, and the word*invariant*in the title of this paper is doing no work: what has been described is an observer\-dependent artefact\. The answer is that the family\-indexed constructions are not independent of one another\. They are functorially related, and the object that Definition 8\.1 names is properly the limit of the whole system rather than any one of its terms\.

### 8\.1Observational contexts and refinement

###### Definition 8\.2\(observational context\)\.

An observational context is a tuplec=\(Bc,Ac,Πc,Kc\)c=\(B\_\{c\},A\_\{c\},\\Pi\_\{c\},K\_\{c\}\)consisting of a subfamily of probes, a subset of modalities, a countable class of admissible policies and a set of horizons\. WriteScS\_\{c\}for the predictive state \([4](https://arxiv.org/html/2609.11959#S3.E4)\) formed from the coordinates indexed byΠc×Kc\\Pi\_\{c\}\\times K\_\{c\}using the modalities inAcA\_\{c\}at the probes inBcB\_\{c\}, and∼c\\sim\_\{c\}for the induced equivalence on histories\. Order contexts by componentwise inclusion:c≼c′c\\preccurlyeq c^\{\\prime\}whenBc⊆Bc′B\_\{c\}\\subseteq B\_\{c^\{\\prime\}\},Ac⊆Ac′A\_\{c\}\\subseteq A\_\{c^\{\\prime\}\},Πc⊆Πc′\\Pi\_\{c\}\\subseteq\\Pi\_\{c^\{\\prime\}\}andKc⊆Kc′K\_\{c\}\\subseteq K\_\{c^\{\\prime\}\}\. The poset\(𝒞,≼\)\(\\mathcal\{C\},\\preccurlyeq\)is directed, since the componentwise union of two contexts is again a context and countability is preserved by finite unions\.

###### Lemma 8\.3\(refinement\)\.

Ifc≼c′c\\preccurlyeq c^\{\\prime\}then∼c′⁣⊆⁣∼c\{\\sim\_\{c^\{\\prime\}\}\}\\subseteq\{\\sim\_\{c\}\}, and there is a unique surjectionpc′​c:Sc′​\(H\)→Sc​\(H\)p\_\{c^\{\\prime\}c\}\\colon S\_\{c^\{\\prime\}\}\(H\)\\to S\_\{c\}\(H\)withSc=pc′​c∘Sc′S\_\{c\}=p\_\{c^\{\\prime\}c\}\\circ S\_\{c^\{\\prime\}\}\. These maps satisfypc​c=idp\_\{cc\}=\\mathrm\{id\}andpc′​c∘pc′′​c′=pc′′​cp\_\{c^\{\\prime\}c\}\\circ p\_\{c^\{\\prime\\prime\}c^\{\\prime\}\}=p\_\{c^\{\\prime\\prime\}c\}wheneverc≼c′≼c′′c\\preccurlyeq c^\{\\prime\}\\preccurlyeq c^\{\\prime\\prime\}\.

###### Proof\.

The coordinates ofScS\_\{c\}form a subset of the coordinates ofSc′S\_\{c^\{\\prime\}\}, soScS\_\{c\}is the corresponding coordinate projection ofSc′S\_\{c^\{\\prime\}\}; in particularSc′​\(h\)=Sc′​\(h′\)S\_\{c^\{\\prime\}\}\(h\)=S\_\{c^\{\\prime\}\}\(h^\{\\prime\}\)impliesSc​\(h\)=Sc​\(h′\)S\_\{c\}\(h\)=S\_\{c\}\(h^\{\\prime\}\), which is both the inclusion of equivalence relations and the well\-definedness ofpc′​cp\_\{c^\{\\prime\}c\}on the image\. Uniqueness holds becausepc′​cp\_\{c^\{\\prime\}c\}is determined onSc′​\(H\)S\_\{c^\{\\prime\}\}\(H\), which is the whole of its domain, and the cocycle identities are the corresponding identities for coordinate projections\. ∎

### 8\.2Space as a projective limit

###### Theorem 8\.4\(space as a projective limit\)\.

The pairs\(\{Sc​\(H\)\},\{pc′​c\}\)\(\\\{S\_\{c\}\(H\)\\\},\\\{p\_\{c^\{\\prime\}c\}\\\}\)form a projective system over the directed poset\(𝒞,≼\)\(\\mathcal\{C\},\\preccurlyeq\)\. Let

S∞:=lim←c∈𝒞⁡Sc​\(H\)=\{\(sc\)c∈∏cSc​\(H\):pc′​c​\(sc′\)=sc​whenever​c≼c′\},S\_\{\\infty\}\\;:=\\;\\varprojlim\_\{c\\in\\mathcal\{C\}\}S\_\{c\}\(H\)\\;=\\;\\Bigl\\\{\(s\_\{c\}\)\_\{c\}\\in\\prod\_\{c\}S\_\{c\}\(H\)\\ :\\ p\_\{c^\{\\prime\}c\}\(s\_\{c^\{\\prime\}\}\)=s\_\{c\}\\ \\text\{ whenever \}c\\preccurlyeq c^\{\\prime\}\\Bigr\\\},\(20\)and letScan​\(h\):=\(Sc​\(h\)\)cS\_\{\\mathrm\{can\}\}\(h\):=\(S\_\{c\}\(h\)\)\_\{c\}\. Then:

1. \(i\)ScanS\_\{\\mathrm\{can\}\}takes values inS∞S\_\{\\infty\}andSc=prc∘ScanS\_\{c\}=\\mathrm\{pr\}\_\{c\}\\circ S\_\{\\mathrm\{can\}\}for everycc;
2. \(ii\)σ​\(Scan\)=⋁cσ​\(Sc\)\\sigma\(S\_\{\\mathrm\{can\}\}\)=\\bigvee\_\{c\}\\sigma\(S\_\{c\}\), soScanS\_\{\\mathrm\{can\}\}is the minimal statistic sufficient for all contexts simultaneously;
3. \(iii\)eachScS\_\{c\}is an interventional invariant relative tocc, whereasScanS\_\{\\mathrm\{can\}\}is invariant under change of context;
4. \(iv\)if there is a contextc0c\_\{0\}such thatpc​c0p\_\{cc\_\{0\}\}is injective for everyc≽c0c\\succcurlyeq c\_\{0\}, thenS∞≅Sc0S\_\{\\infty\}\\cong S\_\{c\_\{0\}\}, and we callc0c\_\{0\}a*sufficient context*;
5. \(v\)if𝒞\\mathcal\{C\}admits a countable cofinal chainc1≼c2≼⋯c\_\{1\}\\preccurlyeq c\_\{2\}\\preccurlyeq\\cdotswith eachScn​\(H\)S\_\{c\_\{n\}\}\(H\)standard Borel and each connecting map Borel, thenS∞S\_\{\\infty\}is a Borel subset of a Polish product and is therefore standard Borel\.

###### Proof\.

The system is projective by Lemma 8\.3 and directed by Definition 8\.2\. \(i\) The compatibility conditions definingS∞S\_\{\\infty\}are exactly the identitiesSc=pc′​c∘Sc′S\_\{c\}=p\_\{c^\{\\prime\}c\}\\circ S\_\{c^\{\\prime\}\}, which hold pointwise\. \(ii\) EachScS\_\{c\}factors throughScanS\_\{\\mathrm\{can\}\}by \(i\), so⋁cσ​\(Sc\)⊆σ​\(Scan\)\\bigvee\_\{c\}\\sigma\(S\_\{c\}\)\\subseteq\\sigma\(S\_\{\\mathrm\{can\}\}\); converselyScanS\_\{\\mathrm\{can\}\}is measurable with respect to the productσ\\sigma\-algebra generated by the coordinatesScS\_\{c\}, giving the reverse inclusion\. Minimality then follows from Theorem 3\.2\(ii\) applied coordinatewise\. \(iii\) Immediate from \(i\) and the definition of the limit, which involves no choice of context\. \(iv\) Ifpc​c0p\_\{cc\_\{0\}\}is injective for allc≽c0c\\succcurlyeq c\_\{0\}, then a compatible family is determined by itsc0c\_\{0\}\-component, because everyccis dominated by somec′≽c0c^\{\\prime\}\\succcurlyeq c\_\{0\}by directedness; the mapS∞→Sc0S\_\{\\infty\}\\to S\_\{c\_\{0\}\}is thus a bijection with inverse given by transport along the system\. \(v\) A cofinal countable chain determines the limit, which is then the set of compatible sequences in a countable product of Polish spaces — a Borel subset, since it is a countable intersection of preimages of diagonals under Borel maps, and Borel subsets of standard Borel spaces are standard Borel\. ∎

###### Definition 8\.5\(space, final form\)\.

For a class𝒞\\mathcal\{C\}of admissible observational contexts, space is the projective limitS∞=lim←c∈𝒞⁡Sc​\(H\)S\_\{\\infty\}=\\varprojlim\_\{c\\in\\mathcal\{C\}\}S\_\{c\}\(H\)of the context\-indexed predictive quotients, together with the action groupoidGGand the descent data that make eachScS\_\{c\}a compatible local description\.

###### Remark 8\.5a\(the relativism objection answered\)\.

Every individualScS\_\{c\}is context\-relative, exactly as every coordinate chart is chart\-relative\. What Theorem 8\.4\(iii\) supplies is that the family is functorial and the limit is not context\-relative; what varies between observers is which finite approximation toS∞S\_\{\\infty\}their instruments give them access to, and that is a statement about epistemic access rather than about the object\. Clause \(v\) says the limit is a legitimate measurable space, so the definition is not merely formal\. Clause \(iv\) says that in favourable cases the limit is attained at a finite stage, so a finitely instrumented city can in principle carry a complete spatial description rather than an endlessly improvable approximation\. This is the sense in which the construction is an invariant, and it is the sense the title intends\.

###### Corollary 8\.6\(saturation is falsifiable\)\.

Forc≼c′c\\preccurlyeq c^\{\\prime\}define the saturation residual

Esat​\(c,c′\):=supπ,k𝔼D\(Pπ,k\(⋅∣H\)∥Qπ,k\(⋅∣Sc\(H\)\)\)−supπ,k𝔼D\(Pπ,k\(⋅∣H\)∥Qπ,k\(⋅∣Sc′\(H\)\)\)\.\\begin\{split\}E\_\{\\mathrm\{sat\}\}\(c,c^\{\\prime\}\)\\;:=\\;\\ &\\sup\_\{\\pi,k\}\\ \\mathbb\{E\}\\,D\\bigl\(P^\{\\pi,k\}\(\\cdot\\mid H\)\\,\\big\\\|\\,Q^\{\\pi,k\}\(\\cdot\\mid S\_\{c\}\(H\)\)\\bigr\)\\\\\[2\.0pt\] &\-\\ \\sup\_\{\\pi,k\}\\ \\mathbb\{E\}\\,D\\bigl\(P^\{\\pi,k\}\(\\cdot\\mid H\)\\,\\big\\\|\\,Q^\{\\pi,k\}\(\\cdot\\mid S\_\{c^\{\\prime\}\}\(H\)\)\\bigr\)\.\\end\{split\}\(21\)The assertion thatc0c\_\{0\}is a sufficient context predictsEsat​\(c0,c′\)=0E\_\{\\mathrm\{sat\}\}\(c\_\{0\},c^\{\\prime\}\)=0within tolerance for everyc′≽c0c^\{\\prime\}\\succcurlyeq c\_\{0\}\. A nonzero residual at anyc′c^\{\\prime\}falsifies sufficiency and exhibits a modality, probe or policy carrying spatial information not present inc0c\_\{0\}\.

In the synthetic system of Section[5\.1](https://arxiv.org/html/2609.11959#S5.SS1)the residual falls by3\.7×10−13\.7\\times 10^\{\-1\}from\{v\}\\\{v\\\}to\{v,a\}\\\{v,a\\\}and thereafter, in the noiseless control, to below10−2910^\{\-29\}with no further decrease, soc0=\{v,a\}c\_\{0\}=\\\{v,a\\\}is a sufficient context there\. Section[6\.3](https://arxiv.org/html/2609.11959#S6.SS3)states the corresponding urban protocol as Experiment E\. It is worth being explicit that saturation is a substantive empirical claim about a city and not a theorem about cities: a city whose saturation curve never flattens would be one for which no finite instrument set determines a spatial description, and nothing in this framework rules that out\.

### 8\.3Time

Time enters through composition and irreversibility\. IfΦs,t\\Phi\_\{s,t\}maps states fromsstott, temporal coherence requires

Φt,u∘Φs,t=Φs,u,Φt,t=id\.\\Phi\_\{t,u\}\\circ\\Phi\_\{s,t\}=\\Phi\_\{s,u\},\\qquad\\Phi\_\{t,t\}=\\mathrm\{id\}\.\(22\)This algebra supplies order and duration only when coupled to clocks, causal cones, dissipation or action cost\. Space and time are therefore not produced merely by extending a list of dimensions\. Space records the compatibility of possible co\-existence and transition; time records the composable ordering, rate and irreversibility of change\.

## 9Limitations and boundary conditions

##### Identifiability\.

Theorem 3\.7 recovers the latent space only up to the centraliser of the intervention group, and only under joint separation, equivariance and interventional faithfulness\. Each hypothesis is a real risk in deployment\. The policy family must excite the relevant degrees of freedom, which by Lemma 3\.5 is a positivity condition on the data\-generating regime rather than a property of the city; modalities must jointly separate the states of interest, which no single instrument suite guarantees; and equivariance of the learned encoder is enforced by penalty and verified as a residual, not derived\.

##### Confounding\.

Assumption I2 is the strongest substantive assumption in the paper\. In an instrumented building, interventions are often selected by facility managers, schedules or occupants in response to conditions that the sensor suite does not record, which is precisely unobserved confounding\. Where it fails, the estimated predictive state is a statistical rather than an interventional object, and Corollary 3\.8 does not apply\.

##### Nonstationarity\.

Urban institutions and populations change the observation and intervention laws\. A fixed sheaf can become obsolete; restriction maps and strata may need online change\-point tests\. In the language of Section 8 the context class itself drifts, so the projective system is indexed by a moving poset and the limit of Theorem 8\.4 need not exist\.

##### Normativity\.

Social and economic geometries contain power, exclusion and value\. Predictability does not make a policy legitimate\. Governance constraints must be model inputs, not conclusions deduced from geometric efficiency\. This is sharpened by Lemma 3\.5: because only executed interventions enter the geometry, whoever controls which interventions are performed thereby controls which spatial distinctions the model is able to represent at all\.

##### Computation\.

Exact inference over rich sheaves, long horizons and joint robot\-environment interventions is generally intractable\. Approximation changes the effective predictive quotient and must be reported as part of the model\.

##### Evidence\.

The synthetic tests check algebraic coherence and the specific claims of Propositions 4\.2, 4\.3 and Theorem 8\.4\(iv\) in a system where the ground truth is known\. They supply no field validation\. A publishable empirical claim requires preregistered interventions, real multimodal data, ablations and external replication\.

## 10Conclusion

The central claim can be said plainly\. Vision, hearing, touch and the many fields and flows of a city do not disclose space because they look alike\. They disclose it because the same actions produce stable, mutually constraining and counterfactually testable changes in them\. Mathematical space supplies the formal object, physical space supplies empirical law, and perceptual space supplies the minimal predictive quotient needed for action\. Their unity lies in an interventional invariant, not in the erasure of their distinctions\.

A stratified base, typed fibres and sheaf descent then provide a disciplined way to place buildings, fields, mobility, institutions and exchange in one model\. The same construction clarifies em\-spaced intelligence: an intelligent agent does not merely occupy a passive environment; agent and spatial body jointly determine the future possibilities of sensing and action\. The mathematical programme is consequently falsifiable\. Its predictions fail when modalities do not align under intervention, local models do not glue, scale maps do not commute within tolerance, or a simpler predictive state performs equally well\. Those failure conditions are not weaknesses of the definition\. They are what make it scientific\.

## Appendix ANotation

Table 3:Principal notation\.SymbolMeaningBBSite of probes or stratified urban base\.ℱ\\mathcal\{F\}Sheaf of typed local states;ℱ​\(U\)\\mathcal\{F\}\(U\)is the state space over probeUU\.GGAction groupoid; arrows are executable movements or interventions\.hαh\_\{\\alpha\}Observation map for modalityα\\alpha\.Φg\\Phi\_\{g\}State transformation induced by actiongg\.TαgT^\{g\}\_\{\\alpha\}Corresponding transformation in modalityα\\alpha\.HHSpace of multimodal histories\.Π\\Pi,KKAdmissible policies and prediction horizons\.S​\(h\)S\(h\)Canonical predictive state, a product of conditional future laws\.δℱ\\delta\_\{\\mathcal\{F\}\},LℱL\_\{\\mathcal\{F\}\}Sheaf coboundary and weighted sheaf Laplacian\.Rℓ→mR\_\{\\ell\\to m\}Cross\-scale restriction or coarsening map\.cc,𝒞\\mathcal\{C\}Observational context and the directed poset of contexts\.ScS\_\{c\},pc′​cp\_\{c^\{\\prime\}c\}Context\-relative predictive quotient and the refinement projections\.S∞S\_\{\\infty\}Projective limit of the context\-indexed quotients; space in the final sense\.ZHomeo⁡\(X\)​\(Φ​\(G\)\)Z\_\{\\operatorname\{Homeo\}\(X\)\}\(\\Phi\(G\)\)Centraliser of the intervention action; the residual ambiguity group\.Hol⁡\(ℱ,v0\)\\operatorname\{Hol\}\(\\mathcal\{F\},v\_\{0\}\)Holonomy group of the cellular sheafℱ\\mathcal\{F\}at base vertexv0v\_\{0\}\.η​\(τ\)\\eta\(\\tau\),τϵ\\tau\_\{\\epsilon\}Interventional separation modulus and the resolution it implies\.bb, I1–I3Behaviour policy and the causal assumptions makingPPinterventional\.
## Appendix BAcceptance criteria for an empirical paper

##### B1\. Intervention coverage\.

The training and test sets must state which actions are observed, which are held out, and whether policy support is sufficient to distinguish the claimed spatial states\.

##### B2\. Competing geometries\.

Compare Euclidean coordinates, graph distance, learned latent distance, action cost and the sheaf model\. Report where each succeeds rather than selecting one post hoc\.

##### B3\. Structural ablations\.

Remove action conditioning, cross\-modal alignment, nontrivial restriction maps, cross\-scale penalties and environmental actuation one at a time\.

##### B4\. Calibration and shift\.

Report conditional likelihood or proper scoring rules, uncertainty calibration and performance under unseen buildings, populations, weather and institutional regimes\.

##### B5\. Governance\.

State who controls spatial interventions, whose costs enter the objective and which safety or access constraints are inviolable\.

## References

- \[1\]K\. Ahuja, D\. Mahajan, Y\. Wang, and Y\. Bengio\(2023\)Interventional causal representation learning\.InProceedings of the 40th International Conference on Machine Learning,PMLR, Vol\.202,pp\. 372–407\.Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[Remark 3\.10](https://arxiv.org/html/2609.11959#Thminnerrem4.p1.3)\.
- \[2\]A\. S\. Bandeira, A\. Singer, and D\. A\. Spielman\(2013\)A Cheeger inequality for the graph connection Laplacian\.SIAM Journal on Matrix Analysis and Applications34\(4\),pp\. 1611–1630\.External Links:[Document](https://dx.doi.org/10.1137/120875338)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[§4\.2](https://arxiv.org/html/2609.11959#S4.SS2.p4.5)\.
- \[3\]D\. Bao, S\. Chern, and Z\. Shen\(2000\)An introduction to Riemann–Finsler geometry\.Graduate Texts in Mathematics, Vol\.200,Springer\.External Links:[Document](https://dx.doi.org/10.1007/978-1-4612-1268-3)Cited by:[§3\.4](https://arxiv.org/html/2609.11959#S3.SS4.p2.1),[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[4\]A\. Barreto, W\. Dabney, R\. Munos, J\. J\. Hunt, T\. Schaul, H\. van Hasselt, and D\. Silver\(2017\)Successor features for transfer in reinforcement learning\.InAdvances in Neural Information Processing Systems,Vol\.30\.Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p3.1)\.
- \[5\]F\. Battiston, G\. Cencetti, I\. Iacopini, V\. Latora, M\. Lucas, A\. Patania, J\. Young, and G\. Petri\(2020\)Networks beyond pairwise interactions: structure and dynamics\.Physics Reports874,pp\. 1–92\.External Links:[Document](https://dx.doi.org/10.1016/j.physrep.2020.05.004)Cited by:[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[6\]M\. Batty\(2013\)The new science of cities\.MIT Press,Cambridge, MA\.Cited by:[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[7\]T\. E\. J\. Behrens, T\. H\. Muller, J\. C\. R\. Whittington, S\. Mark, A\. B\. Baram, K\. L\. Stachenfeld, and Z\. Kurth\-Nelson\(2018\)What is a cognitive map? Organizing knowledge for flexible behavior\.Neuron100\(2\),pp\. 490–509\.External Links:[Document](https://dx.doi.org/10.1016/j.neuron.2018.10.002)Cited by:[§2\.3](https://arxiv.org/html/2609.11959#S2.SS3.p1.1),[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p3.1)\.
- \[8\]C\. Bodnar, F\. Di Giovanni, B\. P\. Chamberlain, P\. Liò, and M\. M\. Bronstein\(2022\)Neural sheaf diffusion: a topological perspective on heterophily and oversmoothing in GNNs\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 18527–18541\.Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[§4\.1](https://arxiv.org/html/2609.11959#S4.SS1.p2.4),[§4\.2](https://arxiv.org/html/2609.11959#S4.SS2.p2.1)\.
- \[9\]J\. Brehmer, P\. de Haan, P\. Lippe, and T\. S\. Cohen\(2022\)Weakly supervised causal representation learning\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 38319–38331\.Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[Remark 3\.10](https://arxiv.org/html/2609.11959#Thminnerrem4.p1.3)\.
- \[10\]M\. M\. Bronstein, J\. Bruna, T\. Cohen, and P\. Veličković\(2021\)Geometric deep learning: grids, groups, graphs, geodesics, and gauges\.arXiv preprint arXiv:2104\.13478\.External Links:2104\.13478Cited by:[§6\.2](https://arxiv.org/html/2609.11959#S6.SS2.p1.5)\.
- \[11\]P\. S\. Castro\(2020\)Scalable methods for computing state similarity in deterministic Markov decision processes\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.34,pp\. 10069–10076\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v34i06.6564)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p1.1)\.
- \[12\]A\. Connes\(1994\)Noncommutative geometry\.Academic Press,San Diego\.Cited by:[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[13\]J\. M\. Curry\(2014\)Sheaves, cosheaves and applications\.Ph\.D\. Thesis,University of Pennsylvania\.External Links:1303\.3255Cited by:[§4\.1](https://arxiv.org/html/2609.11959#S4.SS1.p2.4)\.
- \[14\]P\. Dayan\(1993\)Improving generalization for temporal difference learning: the successor representation\.Neural Computation5\(4\),pp\. 613–624\.External Links:[Document](https://dx.doi.org/10.1162/neco.1993.5.4.613)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p3.1)\.
- \[15\]H\. Edelsbrunner and J\. Harer\(2010\)Computational topology: an introduction\.American Mathematical Society\.External Links:[Document](https://dx.doi.org/10.1090/mbk/069)Cited by:[§4\.1](https://arxiv.org/html/2609.11959#S4.SS1.p1.2)\.
- \[16\]C\. Fefferman\(1971\)The multiplier problem for the ball\.Annals of Mathematics94\(2\),pp\. 330–336\.Cited by:[§7](https://arxiv.org/html/2609.11959#S7.p2.1)\.
- \[17\]N\. Ferns, P\. Panangaden, and D\. Precup\(2004\)Metrics for finite Markov decision processes\.InProceedings of the 20th Conference on Uncertainty in Artificial Intelligence,pp\. 162–169\.Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p1.1),[Remark 3\.2b](https://arxiv.org/html/2609.11959#Thminnerrem2.p1.1)\.
- \[18\]T\. Fritz\(2020\)A synthetic approach to Markov kernels, conditional independence and theorems on sufficient statistics\.Advances in Mathematics370,pp\. 107239\.External Links:[Document](https://dx.doi.org/10.1016/j.aim.2020.107239)Cited by:[Remark 3\.2b](https://arxiv.org/html/2609.11959#Thminnerrem2.p1.1)\.
- \[19\]J\. J\. Gibson\(1979\)The ecological approach to visual perception\.Houghton Mifflin,Boston\.Cited by:[§2\.3](https://arxiv.org/html/2609.11959#S2.SS3.p1.1)\.
- \[20\]R\. Givan, T\. Dean, and M\. Greig\(2003\)Equivalence notions and model minimization in Markov decision processes\.Artificial Intelligence147\(1–2\),pp\. 163–223\.External Links:[Document](https://dx.doi.org/10.1016/S0004-3702%2802%2900376-4)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p1.1),[Remark 3\.2b](https://arxiv.org/html/2609.11959#Thminnerrem2.p1.1)\.
- \[21\]L\. Gresele, P\. K\. Rubenstein, A\. Mehrjou, F\. Locatello, and B\. Schölkopf\(2020\)The Incomplete Rosetta Stone problem: identifiability results for multi\-view nonlinear ICA\.InProceedings of the 35th Conference on Uncertainty in Artificial Intelligence,PMLR, Vol\.115,pp\. 217–227\.Cited by:[Remark 3\.10](https://arxiv.org/html/2609.11959#Thminnerrem4.p1.3)\.
- \[22\]M\. Gromov\(2007\)Metric structures for Riemannian and non\-Riemannian spaces\.Modern Birkhäuser Classics,Birkhäuser,Boston\.External Links:[Document](https://dx.doi.org/10.1007/978-0-8176-4583-0)Cited by:[§2\.1](https://arxiv.org/html/2609.11959#S2.SS1.p1.2),[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[23\]L\. Guth, H\. Wang, and J\. Zahl\(2026\)A streamlined proof of the Kakeya set conjecture inℝ3\\mathbb\{R\}^\{3\}\.arXiv preprint arXiv:2601\.14411\.Note:PreprintExternal Links:2601\.14411Cited by:[§7](https://arxiv.org/html/2609.11959#S7.p2.1)\.
- \[24\]H\. Hälvä, J\. So, R\. E\. Turner, and A\. Hyvärinen\(2024\)Identifiable feature learning for spatial data with nonlinear ICA\.InProceedings of the 27th International Conference on Artificial Intelligence and Statistics,PMLR, Vol\.238,pp\. 3331–3339\.Cited by:[Remark 3\.10](https://arxiv.org/html/2609.11959#Thminnerrem4.p1.3)\.
- \[25\]J\. Hansen and R\. Ghrist\(2019\)Toward a spectral theory of cellular sheaves\.Journal of Applied and Computational Topology3\(4\),pp\. 315–358\.External Links:[Document](https://dx.doi.org/10.1007/s41468-019-00038-7)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[§4\.1](https://arxiv.org/html/2609.11959#S4.SS1.p2.4)\.
- \[26\]J\. Hansen and R\. Ghrist\(2021\)Opinion dynamics on discourse sheaves\.SIAM Journal on Applied Mathematics81\(5\),pp\. 2033–2060\.External Links:[Document](https://dx.doi.org/10.1137/20M1341088)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1)\.
- \[27\]B\. Hillier and J\. Hanson\(1984\)The social logic of space\.Cambridge University Press\.Cited by:[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[28\]B\. Hillier and T\. Yang\(2005\)The art of place and the science of space\.World Architecture\(11\)\.External Links:[Document](https://dx.doi.org/10.16414/j.wa.2005.11.003)Cited by:[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[29\]B\. Hillier\(1996\)Space is the machine: a configurational theory of architecture\.Cambridge University Press\.Cited by:[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[30\]Y\. Hong, J\. Liu, H\. Yin, M\. Li, L\. Guibas, L\. Fei\-Fei, J\. Wu, and Y\. Choi\(2026\)ESI\-Bench: towards embodied spatial intelligence that closes the perception\-action loop\.arXiv preprint arXiv:2605\.18746\.External Links:2605\.18746Cited by:[§6\.1](https://arxiv.org/html/2609.11959#S6.SS1.p1.1),[§6\.3](https://arxiv.org/html/2609.11959#S6.SS3.SSS0.Px6.p2.1)\.
- \[31\]A\. Hyvärinen, H\. Sasaki, and R\. Turner\(2019\)Nonlinear ICA using auxiliary variables and generalized contrastive learning\.InProceedings of the 22nd International Conference on Artificial Intelligence and Statistics,PMLR, Vol\.89,pp\. 859–868\.Cited by:[Remark 3\.10](https://arxiv.org/html/2609.11959#Thminnerrem4.p1.3)\.
- \[32\]H\. Jaeger\(2000\)Observable operator models for discrete stochastic time series\.Neural Computation12\(6\),pp\. 1371–1398\.External Links:[Document](https://dx.doi.org/10.1162/089976600300015411)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p1.1)\.
- \[33\]C\. A\. Joslyn, L\. Charles, C\. DePerno, N\. Gould, K\. Nowak, B\. Praggastis, E\. Purvine, M\. Robinson, J\. Strules, and P\. Whitney\(2020\)A sheaf theoretical approach to uncertainty quantification of heterogeneous geolocation information\.Sensors20\(12\),pp\. 3418\.External Links:[Document](https://dx.doi.org/10.3390/s20123418)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1)\.
- \[34\]I\. Khemakhem, D\. P\. Kingma, R\. P\. Monti, and A\. Hyvärinen\(2020\)Variational autoencoders and nonlinear ICA: a unifying framework\.InProceedings of the 23rd International Conference on Artificial Intelligence and Statistics,PMLR, Vol\.108,pp\. 2207–2217\.Cited by:[Remark 3\.10](https://arxiv.org/html/2609.11959#Thminnerrem4.p1.3)\.
- \[35\]A\. Laflaquière, J\. K\. O’Regan, B\. Gas, and A\. V\. Terekhov\(2018\)Discovering space — grounding spatial topology and metric regularity in a naïve agent’s sensorimotor experience\.Neural Networks105,pp\. 371–392\.External Links:[Document](https://dx.doi.org/10.1016/j.neunet.2018.06.001)Cited by:[§2\.3](https://arxiv.org/html/2609.11959#S2.SS3.p1.1)\.
- \[36\]Y\. LeCun\(2022\)A path towards autonomous machine intelligence, version 0\.9\.2\.Note:OpenReview, 27 June 2022External Links:[Link](https://openreview.net/forum?id=BZ5a1r-kVsf)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p3.1)\.
- \[37\]P\. Lippe, S\. Magliacane, S\. Löwe, Y\. M\. Asano, T\. Cohen, and S\. Gavves\(2022\)CITRIS: causal identifiability from temporal intervened sequences\.InProceedings of the 39th International Conference on Machine Learning,PMLR, Vol\.162,pp\. 13557–13603\.Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[Remark 3\.10](https://arxiv.org/html/2609.11959#Thminnerrem4.p1.3)\.
- \[38\]M\. L\. Littman, R\. S\. Sutton, and S\. Singh\(2002\)Predictive representations of state\.InAdvances in Neural Information Processing Systems,Vol\.14,pp\. 1555–1561\.Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p1.1)\.
- \[39\]F\. Locatello, S\. Bauer, M\. Lucic, G\. Rätsch, S\. Gelly, B\. Schölkopf, and O\. Bachem\(2019\)Challenging common assumptions in the unsupervised learning of disentangled representations\.InProceedings of the 36th International Conference on Machine Learning,PMLR, Vol\.97,pp\. 4114–4124\.Cited by:[Table 1](https://arxiv.org/html/2609.11959#S5.T1.3.3.1.1.1),[Remark 3\.10](https://arxiv.org/html/2609.11959#Thminnerrem4.p1.3)\.
- \[40\]S\. Mac Lane and I\. Moerdijk\(1992\)Sheaves in geometry and logic: a first introduction to topos theory\.Universitext,Springer,New York\.External Links:[Document](https://dx.doi.org/10.1007/978-1-4612-0927-0)Cited by:[§2\.1](https://arxiv.org/html/2609.11959#S2.SS1.p1.2)\.
- \[41\]M\. Nickel and D\. Kiela\(2017\)Poincaré embeddings for learning hierarchical representations\.InAdvances in Neural Information Processing Systems,Vol\.30\.Cited by:[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[42\]B\. O’Neill\(1983\)Semi\-Riemannian geometry with applications to relativity\.Pure and Applied Mathematics, Vol\.103,Academic Press\.Cited by:[§2\.2](https://arxiv.org/html/2609.11959#S2.SS2.p1.1)\.
- \[43\]J\. K\. O’Regan and A\. Noë\(2001\)A sensorimotor account of vision and visual consciousness\.Behavioral and Brain Sciences24\(5\),pp\. 939–973\.External Links:[Document](https://dx.doi.org/10.1017/S0140525X01000115)Cited by:[§2\.3](https://arxiv.org/html/2609.11959#S2.SS3.p1.1)\.
- \[44\]J\. Pearl\(2009\)Causality: models, reasoning and inference\.2nd edition,Cambridge University Press\.Cited by:[§2\.2](https://arxiv.org/html/2609.11959#S2.SS2.p1.1),[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1)\.
- \[45\]M\. J\. Pflaum\(2001\)Analytic and geometric study of stratified spaces\.Lecture Notes in Mathematics, Vol\.1768,Springer\.External Links:[Document](https://dx.doi.org/10.1007/3-540-45436-5)Cited by:[§4\.1](https://arxiv.org/html/2609.11959#S4.SS1.p1.2)\.
- \[46\]J\. Robins\(1986\)A new approach to causal inference in mortality studies with a sustained exposure period — application to control of the healthy worker survivor effect\.Mathematical Modelling7\(9–12\),pp\. 1393–1512\.External Links:[Document](https://dx.doi.org/10.1016/0270-0255%2886%2990088-6)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[§3\.3\.1](https://arxiv.org/html/2609.11959#S3.SS3.SSS1.1.p1.1)\.
- \[47\]M\. Robinson\(2017\)Sheaves are the canonical data structure for sensor integration\.Information Fusion36,pp\. 208–224\.External Links:[Document](https://dx.doi.org/10.1016/j.inffus.2016.12.002)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[§4\.1](https://arxiv.org/html/2609.11959#S4.SS1.p2.4)\.
- \[48\]B\. Schölkopf, F\. Locatello, S\. Bauer, N\. R\. Ke, N\. Kalchbrenner, A\. Goyal, and Y\. Bengio\(2021\)Toward causal representation learning\.Proceedings of the IEEE109\(5\),pp\. 612–634\.External Links:[Document](https://dx.doi.org/10.1109/JPROC.2021.3058954)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1)\.
- \[49\]C\. R\. Shalizi and J\. P\. Crutchfield\(2001\)Computational mechanics: pattern and prediction, structure and simplicity\.Journal of Statistical Physics104\(3\),pp\. 817–879\.External Links:[Document](https://dx.doi.org/10.1023/A%3A1010388907793)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p1.1),[Remark 3\.2b](https://arxiv.org/html/2609.11959#Thminnerrem2.p1.1)\.
- \[50\]A\. Singer and H\. Wu\(2012\)Vector diffusion maps and the connection Laplacian\.Communications on Pure and Applied Mathematics65\(8\),pp\. 1067–1144\.External Links:[Document](https://dx.doi.org/10.1002/cpa.21395)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[§4\.2](https://arxiv.org/html/2609.11959#S4.SS2.p4.5)\.
- \[51\]C\. Squires, A\. Seigal, S\. S\. Bhate, and C\. Uhler\(2023\)Linear causal disentanglement via interventions\.InProceedings of the 40th International Conference on Machine Learning,PMLR, Vol\.202,pp\. 32540–32560\.Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[Remark 3\.10](https://arxiv.org/html/2609.11959#Thminnerrem4.p1.3)\.
- \[52\]K\. L\. Stachenfeld, M\. M\. Botvinick, and S\. J\. Gershman\(2017\)The hippocampus as a predictive map\.Nature Neuroscience20\(11\),pp\. 1643–1653\.External Links:[Document](https://dx.doi.org/10.1038/nn.4650)Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p3.1)\.
- \[53\]K\. Sturm\(2024\)Metric measure spaces and synthetic Ricci bounds: fundamental concepts and recent developments\.arXiv preprint arXiv:2404\.15755\.External Links:2404\.15755Cited by:[§2\.1](https://arxiv.org/html/2609.11959#S2.SS1.p1.2)\.
- \[54\]A\. V\. Terekhov and J\. K\. O’Regan\(2016\)Space as an invention of active agents\.Frontiers in Robotics and AI3,pp\. 4\.External Links:[Document](https://dx.doi.org/10.3389/frobt.2016.00004)Cited by:[§2\.3](https://arxiv.org/html/2609.11959#S2.SS3.p1.1)\.
- \[55\]C\. Villani\(2009\)Optimal transport: old and new\.Grundlehren der mathematischen Wissenschaften, Vol\.338,Springer\.Cited by:[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[56\]J\. von Kügelgen, M\. Besserve, W\. Liang, L\. Gresele, A\. Kekić, E\. Bareinboim, D\. M\. Blei, and B\. Schölkopf\(2023\)Nonparametric identifiability of causal representations from unknown interventions\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 48603–48638\.Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[Remark 3\.10](https://arxiv.org/html/2609.11959#Thminnerrem4.p1.3)\.
- \[57\]J\. von Kügelgen, Y\. Sharma, L\. Gresele, W\. Brendel, B\. Schölkopf, M\. Besserve, and F\. Locatello\(2021\)Self\-supervised learning with data augmentations provably isolates content from style\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 16451–16467\.Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p2.1),[Remark 3\.10](https://arxiv.org/html/2609.11959#Thminnerrem4.p1.3)\.
- \[58\]H\. Wang and J\. Zahl\(2025\)Volume estimates for unions of convex sets, and the Kakeya set conjecture in three dimensions\.arXiv preprint arXiv:2502\.17655\.Note:PreprintExternal Links:2502\.17655Cited by:[§7](https://arxiv.org/html/2609.11959#S7.p2.1)\.
- \[59\]W\. Wanget al\.\(2025\)SITE: towards spatial intelligence thorough evaluation\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),External Links:2505\.05456Cited by:[§6\.3](https://arxiv.org/html/2609.11959#S6.SS3.SSS0.Px6.p2.1)\.
- \[60\]T\. Yang, C\. Deng, X\. Lin, and W\. Luo\(2022\)Artificially evolving meta\-city systems: an intelligent generation of urban spatial form\.Shanghai Urban Planning\(3\),pp\. 14–22\.Note:In ChineseCited by:[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[61\]T\. Yang and X\. Lin\(2025\)Generative artificial intelligence and urban spatial morphology: a manifold approach to future cities\.Urban and Regional Planning Research17\(1\),pp\. 1–14\.Note:In ChineseCited by:[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[62\]T\. Yang\(2019\)The value of spatial networks: multi\-scale space syntax\.China Architecture & Building Press\.Note:In ChineseCited by:[§7](https://arxiv.org/html/2609.11959#S7.p1.1)\.
- \[63\]T\. Yang\(2026\)The genesis of artificially evolved spatial organisms: from static syntax to dynamic semantic generation in the “deep water” of urban regeneration\.China City Planning Review35\(2\),pp\. 49–56\.External Links:[Document](https://dx.doi.org/10.20113/j.ccpr.20260208a)Cited by:[§6\.1](https://arxiv.org/html/2609.11959#S6.SS1.p1.1)\.
- \[64\]A\. Zhang, R\. McAllister, R\. Calandra, Y\. Gal, and S\. Levine\(2021\)Learning invariant representations for reinforcement learning without reconstruction\.InInternational Conference on Learning Representations,Cited by:[§2\.4](https://arxiv.org/html/2609.11959#S2.SS4.p1.1)\.
- \[65\]H\. Zhuet al\.\(2025\)SPA: 3D spatial\-awareness enables effective embodied representation\.InInternational Conference on Learning Representations,External Links:2410\.08208Cited by:[§6\.3](https://arxiv.org/html/2609.11959#S6.SS3.SSS0.Px6.p2.1)\.

Similar Articles

Exploring Spatial Intelligence from a Generative Perspective

Hugging Face Daily Papers

Researchers introduce GSI-Bench, the first benchmark to quantify generative spatial intelligence in multimodal models by evaluating 3D spatial constraint compliance during image generation. Fine-tuning on their synthetic dataset boosts both spatial editing fidelity and downstream spatial understanding, showing generative training can strengthen spatial reasoning.