CLAM: Causal Spatial Disaggregation to Infer Local Effects From Coarse Data
摘要
This paper introduces CLAM, a method for estimating localized causal effects from coarse-resolution data by jointly learning causal mechanisms and a disaggregation mapping, with applications in public health and environmental policy.
查看缓存全文
缓存时间: 2026/08/11 08:09
# Causal Spatial Disaggregation to Infer Local Effects From Coarse Data
Source: [https://arxiv.org/html/2608.08064](https://arxiv.org/html/2608.08064)
Gerrit Großmann DFKI Kaiserslautern, Germany Gerrit\.Grossmann@dfki\.de &Sumantrak Mukherjee DFKI11footnotemark:1 Kaiserslautern, Germany Sumantrak\.Mukherjee@dfki\.de&Sebastian Vollmer DFKI11footnotemark:1, RPTU Kaiserslautern, Germany Sebastian\.Vollmer@dfki\.de DFKI: Deutsches Forschungszentrum für Künstliche Intelligenz GmbH \(German Research Center for Artificial Intelligence\)\.
###### Abstract
Learning fine\-grained spatial patterns from coarse\-resolution data is challenging, especially in causal settings where high\-resolution effects must be inferred from aggregated interventions and outcomes\. We introduceClam, a method for estimating localized causal effects from coarse observations by exploiting high\-resolution contextual covariates that modulate these effects\. By jointly learning the causal mechanism and a disaggregation mapping,Clamcaptures interactions that are missed when addressing these problems independently\. The method supports localized effect estimation, counterfactual reasoning, and principled outcome disaggregation, and reliably captures spatially varying causal effects across diverse settings\. This is particularly relevant for applications such as public health and environmental policy, where decisions are made at broad scales despite substantial local heterogeneity\. Code is available at[github\.com/gerritgr/clam](https://github.com/gerritgr/clam)\.
## 1Introduction
Figure 1:Motivating example\.Clamtakes regional treatments \(e\.g\., a vaccination awareness campaign\) and outcomes \(e\.g\., vaccination levels\), along with subregional covariates \(e\.g\., demographics\), to learn treatment–outcome relationships and estimate local causal effects or sample counterfactuals\.Going from high\-resolution \(HR\) to low\-resolution \(LR\) data is typically straightforward, as it can be done via aggregation or pooling functions such as sums or means\. These operations generally remove information\. In contrast, going from LR to HR data is fundamentally difficult because it requires restoring information that has been lost\. This generally necessitates some form of prior knowledge about the HR data\. This inverse problem is typically studied under terms such as*spatial disaggregation*\[[28](https://arxiv.org/html/2608.08064#bib.bib24)\],*downscaling*\[[39](https://arxiv.org/html/2608.08064#bib.bib28)\],*ecological inference*\[[37](https://arxiv.org/html/2608.08064#bib.bib47)\], or*super\-resolution*\[[1](https://arxiv.org/html/2608.08064#bib.bib29)\]\. It is relevant across many fields involving spatio\-temporal statistics, including earth science\[[18](https://arxiv.org/html/2608.08064#bib.bib33)\], public health\[[24](https://arxiv.org/html/2608.08064#bib.bib31)\], and the social sciences\[[42](https://arxiv.org/html/2608.08064#bib.bib32)\]\.
Causality\[[29](https://arxiv.org/html/2608.08064#bib.bib30)\], on the other hand, studies the effects of interventions on a system\. Its core insight is that statistical associations do not imply causal relationships\. In other words, observational patterns may be useful for prediction, but they do not, in general, tell us how outcomes will change under intervention\.
Answering causal questions is often attempted using LR data, but this data is frequently insufficient for this\[[4](https://arxiv.org/html/2608.08064#bib.bib13)\]\. For example, interventions, like vaccination campaigns, might be applied across broad regions, yet their impact could vary significantly at the subregional level\. Surprisingly little effort has been made to properly combine spatial disaggregation with the quest to answer causal questions\.
Our hypothesis is that causal inference and spatial disaggregation are better solved together than apart, and that doing so makes questions answerable that neither can address alone\.
##### Contribution\.
This work addresses the gap between spatial disaggregation and causal inference by introducing a unified framework grounded in structural causal models \(SCMs\) and Pearl’s do\-calculus\[[29](https://arxiv.org/html/2608.08064#bib.bib30)\]\. We argue that jointly tackling causal inference and disaggregation is natural in many real\-world settings, where interventions are implemented at coarse scales but effects manifest locally\. While our primary focus is the formulation of this framework, we illustrate its capabilities and limitations through a series of simulation studies that mimic realistic data\-generating processes, as well as a case study based on real\-world data on the impact of heat waves on gun violence\.
Clam\(*C*ausa*l*dis*a*ggregation*m*ethod\) can be summarized as follows: We learn a mechanistic function that maps an intervention indicator and contextual information to high\-resolution \(HR\) outcomes\. In doing so, we treat HR contextual data \(or*covariates*\) as an auxiliary variable that implicitly defines a prior over the HR outcome \(cf\. Figure[1](https://arxiv.org/html/2608.08064#S1.F1)\)\. The model is trained by aggregating the predicted HR outcomes and enforcing consistency with the observed low\-resolution \(LR\) measurements\. This enables causal effect estimation, counterfactual reasoning, and the computation of HR estimates from LR inputs\.
Because no existing method targets this estimand, our contribution is the problem formulation first and the algorithm second, and we prioritize making the assumptions and limits of the setting explicit over benchmarking against methods built for a different task\.
## 2Problem Setup
For notational simplicity and without loss of generality, we assume that all regions share the same subregion structure, interventions are binary, and all variables are scalar rather than vector\-valued\.
We considerNNregions, each divided intoMMsubregions\. The LRtreatment datais a vectorT^∈\{0,1\}N\\widehat\{T\}\\in\\\{0,1\\\}^\{N\}, where each entryt^i\\widehat\{t\}\_\{i\}indicates whether an intervention occurred in regionii\. We assume that, when present, the intervention affects all subregions within that region\.
The LRoutcome datais a vectorY^∈ℝN\\widehat\{Y\}\\in\\mathbb\{R\}^\{N\}, obtained by applying a row\-wise aggregation function \(e\.g\., summation or averaging\) to an unobserved HR outcome matrixY∈ℝN×MY\\in\\mathbb\{R\}^\{N\\times M\}, whereyi,jy\_\{i,j\}denotes the outcome value for subregionjjwithin regionii\. The HRcontextual data\(covariates\) is given by a matrixC∈ℝN×MC\\in\\mathbb\{R\}^\{N\\times M\}, whereci,jc\_\{i,j\}denotes the contextual value for subregionjjin regionii\.
Finally, we assume thatNNandMMare perfect squares, allowing us to represent the spatial structure as aN×N\\sqrt\{N\}\\times\\sqrt\{N\}grid of regions, each containing aM×M\\sqrt\{M\}\\times\\sqrt\{M\}grid of subregions \(see Figure[1](https://arxiv.org/html/2608.08064#S1.F1)for an example withN=M=16N=M=16andN=M=4\\sqrt\{N\}=\\sqrt\{M\}=4\)\.
##### Goal\.
Our goal is to learn a function that estimates local causal effects and to disaggregate coarse outcomes into fine\-grained predictions\. We formalize this objective in the next section\.
##### Generalizations\.
The setup admits several natural extensions\. Variables \(such ascijc\_\{ij\}\) can be vector\-valued, treatments can be continuous instead of binary, and the observed LR treatmentT^\\widehat\{T\}can be interpreted as an aggregation of an underlying HR treatment matrix that varies within a region\. The framework also extends to spatio\-temporal settings, whereTT,YY, andCCevolve over time\. More generally, additional structure such as hidden confounding, mediation, or non\-trivial treatment assignment can be incorporated; we discuss these cases in Section[4](https://arxiv.org/html/2608.08064#S4)\. Naturally, the spatial layout need not follow a regular grid, which we adopt only for illustration\.
##### Causal Assumptions\.
Depending on the experimental setting, different assumptions may hold or be relaxed\. We may assume a known causal graph, which constrains the model and reduces the risk of spurious explanations, particularly important when contextual variables are high\-dimensional and correlated, but only a subset is causally relevant\. We typically assume no hidden confounding, meaning all relevant variables affecting the outcome are observed, allowing interventions to be interpreted causally up to reasonable noise\. Causal effect identification also requires variation in interventions across regions\. If the aggregation function from subregional to regional outcomes is known \(e\.g\., a mean or sum\), disaggregation becomes more tractable\. Moreover, we typically assume that the treatment does not depend on the context values\.
ti,1t\_\{i,1\}ci,1c\_\{i,1\}ϵi,1\\epsilon\_\{i,1\}ci,1′c^\{\\prime\}\_\{i,1\}yi,1y\_\{i,1\}ti,2t\_\{i,2\}ci,2c\_\{i,2\}ϵi,2\\epsilon\_\{i,2\}ci,2′c^\{\\prime\}\_\{i,2\}yi,2y\_\{i,2\}⋯\\cdotsti,mt\_\{i,m\}ci,mc\_\{i,m\}ϵi,m\\epsilon\_\{i,m\}ci,m′c^\{\\prime\}\_\{i,m\}yi,my\_\{i,m\}AggregationY^i\\widehat\{Y\}\_\{i\}Index\(i,j\)\(i,j\)denotes regionii, subregionjjti,j∈\{0,1\}t\_\{i,j\}\\in\\\{0,1\\\}: intervention indicatorci,j∈ℝc\_\{i,j\}\\in\\mathbb\{R\}: context variableϵi,j∈ℝ\\epsilon\_\{i,j\}\\in\\mathbb\{R\}: noiseci,j′∈ℝc^\{\\prime\}\_\{i,j\}\\in\\mathbb\{R\}: unobserved contextual covariateyi,j∈ℝy\_\{i,j\}\\in\\mathbb\{R\}: subregion outcomeY^i∈ℝ\\widehat\{Y\}\_\{i\}\\in\\mathbb\{R\}: aggregated regional outcome
Figure 2:Causal graph for one region\. All functional relations are shared among subregions and regions\. Variables could also be vector\-valued\. Unobserved covariates and a causal link from context to treatment are optional\. The aggregation is typically a mean or a sum\.
## 3Our Method:Clam
Our method combines two insights\. First, we represent the causal disaggregation problem using a structural causal model formalism\. This formalism encodes assumptions and specifies the missing pieces \(functions, parameters\)\. Second, we train an ML model to infer the missing pieces where we use the LR data as a loss signal to estimate the most likely HR data and functional relationships\.
### 3\.1Structural Causal Model
Each region–subregion pair\(i,j\)\(i,j\)is associated with a treatmentti,jt\_\{i,j\}, contextci,jc\_\{i,j\}, noiseϵi,j\\epsilon\_\{i,j\}, and an optional hidden confounder \(or non\-confounding covariate\)ci,j′c^\{\\prime\}\_\{i,j\}\(cf\. Figure[2](https://arxiv.org/html/2608.08064#S2.F2)\)\. A shared structural function maps these inputs to a subregional outcomeyi,jy\_\{i,j\}, and then aggregates the subregional outcomes to a regional outcomeY^i\\widehat\{Y\}\_\{i\}:
yi,j=f𝜽\(ti,j,ci,j,ϵi,j,ci,j′\),Y^i=gϕ\(\{yi,j\}j=1M\)\.y\_\{i,j\}=f\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},\\,c\_\{i,j\},\\,\\epsilon\_\{i,j\},\\,c^\{\\prime\}\_\{i,j\}\),\\quad\\widehat\{Y\}\_\{i\}=g\_\{\\boldsymbol\{\\phi\}\}\\bigl\(\\\{y\_\{i,j\}\\\}\_\{j=1\}^\{M\}\\bigr\)\.\(1\)
Each subregion is modeled by a local functional mechanism, with uncertainty captured viaϵi,j\\epsilon\_\{i,j\}\. Although the functionsf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)andgϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\)are shared across regions, region\-specific dynamics can still be modeled through context variables \(e\.g\., region indicators or positional encodings\)\. The difference betweenci,j′c^\{\\prime\}\_\{i,j\}andϵi,j\\epsilon\_\{i,j\}is thatϵi,j\\epsilon\_\{i,j\}is i\.i\.d\. noise, whileci,j′c^\{\\prime\}\_\{i,j\}is structured, for example spatially autocorrelated or shared across observations in time\.
In general,gϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\)can be any function parameterized byϕ\\boldsymbol\{\\phi\}\. For simplicity, we typically assume the aggregation is a sum, the noise is additive, and there are no hidden confounders\. Then the model simplifies to:
Y^i=∑j=1M\(f𝜽\(ti,j,ci,j\)\+ϵi,j\)=∑j=1Mf𝜽\(ti,j,ci,j\)\+∑j=1Mϵi,j\.\\widehat\{Y\}\_\{i\}=\\sum\_\{j=1\}^\{M\}\\bigl\(f\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\}\)\+\\epsilon\_\{i,j\}\\bigr\)=\\sum\_\{j=1\}^\{M\}f\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\}\)\+\\sum\_\{j=1\}^\{M\}\\epsilon\_\{i,j\}\.\(2\)
If we want to simplify further, we can replace the flexible estimatorf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)with a linear one, where𝜽=\(α1,β1,α2,β2\)\\boldsymbol\{\\theta\}=\(\\alpha\_\{1\},\\beta\_\{1\},\\alpha\_\{2\},\\beta\_\{2\}\)\. Conceptually, we use two linear models, one for treated regions and one for untreated regions, each relating the local contextci,jc\_\{i,j\}to the outcome:
Y^i=∑j=1M\(ti,j\(α1\+β1ci,j\)\+\(1−ti,j\)\(α2\+β2ci,j\)\)\+∑j=1Mϵi,j\.\\widehat\{Y\}\_\{i\}=\\sum\_\{j=1\}^\{M\}\\bigl\(t\_\{i,j\}\\left\(\\alpha\_\{1\}\+\\beta\_\{1\}c\_\{i,j\}\\right\)\+\(1\-t\_\{i,j\}\)\\left\(\\alpha\_\{2\}\+\\beta\_\{2\}c\_\{i,j\}\\right\)\\bigr\)\+\\sum\_\{j=1\}^\{M\}\\epsilon\_\{i,j\}\.\(3\)
### 3\.2Training
The model is trained by minimizing the discrepancy between the observed LR outcomesY^\\widehat\{Y\}and the predicted aggregated outcomesμ\\mu, typically using an MSE loss:ℒ𝜽,ϕ=1N∑i=1N\(μi−Y^i\)2\.\\mathcal\{L\}\_\{\\boldsymbol\{\\theta\},\\boldsymbol\{\\phi\}\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\bigl\(\\mu\_\{i\}\-\\widehat\{Y\}\_\{i\}\\bigr\)^\{2\}\.Note that we useY^i\\widehat\{Y\}\_\{i\}for the observed regional outcome andμi\\mu\_\{i\}for the corresponding prediction, which under zero\-mean additive noise is the conditional expectation of the aggregate\.
Importantly, the loss is only observed at the aggregated \(region\) level\. The model therefore does not receive direct supervision for individual subregions; instead, it must learn a subregional mechanism that, when combined, is consistent with the observed aggregated outcomes\.
When relaxing the assumptions from Section[2](https://arxiv.org/html/2608.08064#S2), training becomes more complicated\. For example, if the aggregation function,gϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\), is unknown, it must be learned jointly with the local mechanismf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\), which leads to a coupled estimation problem\. In general,gϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\)can be implemented using an \(order\-invariant\) set neural network \(e\.g\.,\[[43](https://arxiv.org/html/2608.08064#bib.bib35)\]\)\. This is conceptually similar to settings in graph neural networks, where both the message\-passing mechanism and the aggregation function are learned\[[12](https://arxiv.org/html/2608.08064#bib.bib34)\]\.
Similarly, if treatments are not uniformly applied across subregions, the underlying HR treatment matrix must be inferred jointly during training\. More generally, additional latent variables \(e\.g\., unobserved confounders\) can be incorporated and estimated within the same framework\. This includes uneven delivery, for example a vaccination campaign reaching subregions to differing degrees\.
### 3\.3Downstream Applications
Training yields the predictive mechanismf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\), from which several causal quantities can be derived\. The most direct is what we call the*local conditional average treatment effect*\(LOCATE\), a version of the conditional average treatment effect \(CATE\) conditioned on subregion\-specific factors:
Ei,j=f𝜽\(1,ci,j\)−f𝜽\(0,ci,j\),E\_\{i,j\}=f\_\{\\boldsymbol\{\\theta\}\}\(1,\\,c\_\{i,j\}\)\-f\_\{\\boldsymbol\{\\theta\}\}\(0,\\,c\_\{i,j\}\),\(4\)i\.e\., the treatment–control difference at the subregional level\. Aggregating these quantities yields ATEs for regions, populations, or environments; with continuous treatments, the same mechanism gives conditional dose–response curves\. Note that LOCATE is noise\-marginalized\.
Beyond effect estimation, the framework supports outcome*disaggregation*—mapping coarse regional outcomes back to fine\-grained predictions—and counterfactual reasoning, e\.g\., predicting outcomes under hypothetical intervention placements or alternative contexts\. In the simplest case, this amounts to evaluatingf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)under new inputs\. For counterfactuals, when the noise term is believed to carry realization\-specific information, inferred noise values can be reused to produce sharper, sample\-specific counterfactuals rather than population\-level averages\.
The framework can also recover latent HR treatment locations by representing them as trainable variables \(cf\. Exp\. 2 in Appendix[D\.2](https://arxiv.org/html/2608.08064#A4.SS2)\)\. With time\-series data,Clamcan further infer latent variables that modulate causal effects, provided temporal variation supplies enough constraints \(cf\. Exp\. 3 in Appendix[D\.3](https://arxiv.org/html/2608.08064#A4.SS3)\)\.
### 3\.4Identifiability
We do not provide a general identifiability result, and a complete treatment is left as an open problem\. Appendix[A](https://arxiv.org/html/2608.08064#A1)distinguishes causal identification, statistical identification from aggregate observations, and finite\-sample recovery by the optimizer, and characterizes several tractable special cases\. In short, recovery of the mechanism and recovery of the high\-resolution outcome map have to be distinguished, and simple counterexamples show that neither is pinned down without further assumptions\. In a noise\-free saturated setting with known sum aggregation, the problem reduces to a linear equation system, which makes the required variability in the data explicit through a rank condition\. Beyond that, the closest existing result is the consistency analysis ofZhanget al\.\[[45](https://arxiv.org/html/2608.08064#bib.bib10)\], which applies to an idealized version of our setting\. How these results extend to spatially dependent covariates, latent variables, and more general aggregation functions, and what the optimizer actually recovers in finite samples, remain open\.
### 3\.5Relationship to Causal Deabstraction
We viewClamas an instance of*causal deabstraction*: recovering a fine\-grained causal mechanism from coarse, aggregated observations\. While prior work studies when causal structure is preserved under abstraction\[[4](https://arxiv.org/html/2608.08064#bib.bib13),[44](https://arxiv.org/html/2608.08064#bib.bib14)\], we consider the inverse direction: inferring a high\-resolution model that is consistent with observed aggregates\. Appendix[B](https://arxiv.org/html/2608.08064#A2)formalizes this view and clarifies the causal semantics\. Aggregation deletes the pairing between outcomes and subregions\. HR covariates and an invariant causal mechanism\[[31](https://arxiv.org/html/2608.08064#bib.bib36),[36](https://arxiv.org/html/2608.08064#bib.bib37)\]can partially restore it when covariate compositions vary across regions\.
## 4Exploiting Context Across Motifs in the Causal Graph
Throughout the paper, we focus on the simplest causal motif: a treatmentTTacting on an outcomeYY, with a contextual covariateCCthat modulates the effect\. Real\-world settings can exhibit a much richer variety of causal structures, and the role of high\-resolution context,CC, shifts with the motif: sometimes it must be conditioned on, sometimes ignored, sometimes leveraged for new identification strategies\. Do\-calculus tells us*whether*an effect is identifiable given the graph, this section asks*how*the LR/HR distinction interacts with that question\. Here, we provide an outlook in which we sketch how HR context can be exploited in common causal graph motifs\.
##### Baseline:CCas a Pure Effect Modifier:T→YT\\to Y,C→YC\\to Y\.
The default setting throughout this paper is the simplest motif:TTandCCare independent inputs, both pointing intoYY, with no edge between them\.CCis not a confounder, only a covariate that modulates the local effect ofTT, and the backdoor criterion holds trivially and gives rise to the LOCATE in Eq\. \([4](https://arxiv.org/html/2608.08064#S3.E4)\)\. Thus,CCprovides multiple “views” of the same regional outcome and thereby constrains the shared mechanismf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)\.
##### Confounder:C→TC\\to T,C→YC\\to Y,T→YT\\to Y\.
When the contextCCalso drives whether \(or by how much\) a region is treated, it becomes a confounder\. The backdoor criterion then requires conditioning onCC\. TreatingTTas exogenous yields biased estimates\. Modeling the treatment\-assignment mechanism at HR is what makes the adjustment correct\. Exp\. 5 \(Appendix[F](https://arxiv.org/html/2608.08064#A6)\) instantiates this motif and shows that ignoring the dependence inflates the MSE on local effects\. A concrete real\-world example to illustrate this problem is a wildfire\-prevention program: forested counties with dry summers are more likely to receive funding, and those same counties also have higher baseline fire incidence, so the program’s effect is conflated with the underlying risk\.
##### Collider:T→CT\\to C,Y→CY\\to C\.
A collider is a variable influenced by both the treatment and the outcome\. Conditioning on such a variable opens a spurious path betweenTTandYY, so naïvely passingCCas an input tof𝜽\(t,c\)f\_\{\\boldsymbol\{\\theta\}\}\(t,c\)would bias the estimated effect\. Nevertheless, HR collider information may still be useful: sinceCCis partly determined by the HR datayi,jy\_\{i,j\}, it can provide fine\-grained information for the reconstruction ofYYor serve as a diagnostic check\. For example, consider a regional vaccination campaign \(TT\), where the outcome of interest is infection rates \(YY\), and the HR variableCCis neighborhood\-level pharmacy purchases of rapid tests\. The campaign may increase public attention and test purchases through its messaging, while actual local infections also increase demand for rapid tests\.
A possible way to exploit this structure is to invert the usual roles: instead of usingCCas a covariate of the outcome, we treatCCitself as the HR target and learn a local mechanismfinv\(t,y\)=cf^\{\\mathrm\{inv\}\}\(t,y\)=cthat maps treatment and outcome to the collider\. Under this setup,Clam’s aggregation\-consistency principle still applies, but it is used to jointly learn the mechanism and the disaggregation ofTTandYYagainst the observed HR values ofCC\.
##### Mediator:T→CT\\to C,C→YC\\to Y,T→YT\\to Y\.
A mediatorCClies on the causal pathway fromTTtoYY, so part of the treatment effect is transmitted throughCCrather than acting directly\. Naïve estimation that conditions onCCcannot separate the direct effect ofTTfrom the indirect effect throughCC\. When the high\-resolution mediator is observed, one can instead jointly learn the treatment\-to\-mediator, mediator\-to\-outcome, and direct treatment\-to\-outcome mechanisms\. Only the two outcome\-side mechanisms rely on the aggregated regional observations, since the treatment\-to\-mediator mechanism is supervised directly by the observedci,jc\_\{i,j\}\. This enables a decomposition of the learned effect into a direct componentT→YT\\to Yand an indirect componentT→C→YT\\to C\\to Y\. A concrete example is a regional advertising campaign observed at low resolution, whose effect on regional sales, also observed at low resolution, is partly direct and partly mediated by word\-of\-mouth dynamics reflected in local online discourse observed at high resolution\.
##### Instrumental variable:C→TC\\to T,T→YT\\to Y,U→TU\\to T,U→YU\\to Y\.
If a hidden confounder links treatment and outcome, but the HR covariate acts on the outcome only through the treatment, that covariate may play the role of an instrument\. An HR instrument is unusual but plausible\. Consider subregional weather \(CC, HR\), which influences whether a vaccination campaign is deployed \(TT, observed at LR\), but does not affect the infection rate \(YY, observed at LR\) in any other way\. One possible extension ofClamwould then be to disaggregate treatment and outcome jointly, asyi,j=f2\(f1\(ci,j\)\)y\_\{i,j\}=f\_\{2\}\(f\_\{1\}\(c\_\{i,j\}\)\), wheref1\(⋅\)f\_\{1\}\(\\cdot\)maps weather to local deployment and is trained against the LR treatment totals, andf2\(⋅\)f\_\{2\}\(\\cdot\)maps deployment to local infections and is trained against the LR outcome totals\. In such a construction,f2\(⋅\)f\_\{2\}\(\\cdot\)would only see the deployment predicted from the weather, which may help isolate the component of treatment variation attributable to the instrument rather than the hidden confounder\.
##### Spill\-over\.
A subregional outcome may depend on neighboring treatments and contexts: interventions can spill over across space\. The local mechanism is then no longer purely local, and aggregated observations alone cannot distinguish a strong local effect from a weaker effect that diffuses outward\. HR context can help by exposing spatial structure\. Technically, one can replace the pointwisef𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)with a grid\-based model\. In principle, the model can learn both direct local effects and spill\-over patterns from aggregated outcomes\. A concrete example is a driving ban in one district that displaces traffic into adjacent ones, so regional air quality depends on both the ban location and surrounding road and vegetation patterns\.
##### Time\-Varying Confounding\.
In spatiotemporal settings, confounders may evolve over time and be shaped by past treatments\. Present context can then no longer be treated as a fixed pre\-treatment covariate, because it may already encode earlier interventions\. This makes estimation harder: naïve adjustment can block part of the causal pathway or introduce bias\[[21](https://arxiv.org/html/2608.08064#bib.bib25)\]\.Clamcan be extended by indexingf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)over time and by explicitly modeling how past treatments, outcomes, and HR covariates affect the current state\. Exp\. 3 already uses temporal variation to recover a stable latent spatial factor; a natural next experiment would make this factor dynamic and treatment\-dependent\. For example, neighborhoods with high prior infection rates may receive additional public\-health resources, which then change later vulnerability and influence future allocation decisions\.
##### Summary\.
Across all motifs, the role of HR data can be summarized as follows: it provides*variation within regions*that translates into observable differences at the aggregated level\. This variation can be exploited to constrain and identify fine\-grained causal mechanisms, but only when the causal graph permits it\.
More concretely: \(i\) for effect modifiers and confounders, HR structure is highly informative and can be directly leveraged for identification and adjustment; \(ii\) for mediators, it remains informative but requires explicitly modeling the underlying causal pathways; \(iii\) for colliders, it can be actively misleading if the framework is not changed accordingly\.
Clamdoes not replace causal identification via do\-calculus: the causal graph determines whether a target effect is identifiable in principle, whileClamestimates the corresponding fine\-grained structure from aggregated data\. Identifiability in principle, however, does not guarantee recovery in practice, which further requires sufficient variability in the HR covariates, a well\-behaved optimization problem, and effective training of the chosen function class\. See Appendix[A](https://arxiv.org/html/2608.08064#A1)and Section[7](https://arxiv.org/html/2608.08064#S7)\.
## 5Related Work
Clamsits at the intersection of multiple research strands; an extended discussion is provided in Appendix[C](https://arxiv.org/html/2608.08064#A3)\.
##### Spatial Disaggregation and Downscaling\.
A long line of work in statistical downscaling and spatial disaggregation infers high\-resolution estimates from coarse data using interpolation, regression, or learning\-based techniques\[[28](https://arxiv.org/html/2608.08064#bib.bib24),[39](https://arxiv.org/html/2608.08064#bib.bib28),[23](https://arxiv.org/html/2608.08064#bib.bib1),[5](https://arxiv.org/html/2608.08064#bib.bib11)\]\. These methods are predominantly associational, recent extensions integrate causally relevant predictors\[[40](https://arxiv.org/html/2608.08064#bib.bib15),[8](https://arxiv.org/html/2608.08064#bib.bib16)\]but still rely on relatively fine\-grained information and do not target interventional or counterfactual queries\.
##### Causal Inference under Aggregation and Abstraction\.
Reasoning causally with aggregated data is well known to be hazardous: results on the ecological fallacy show that aggregate statistics may obscure or invert relationships at the unit level\[[11](https://arxiv.org/html/2608.08064#bib.bib17),[33](https://arxiv.org/html/2608.08064#bib.bib19),[41](https://arxiv.org/html/2608.08064#bib.bib26),[20](https://arxiv.org/html/2608.08064#bib.bib27)\]\. Work on causal abstraction formalizes when mappings between models at different granularities preserve interventional or counterfactual semantics\[[4](https://arxiv.org/html/2608.08064#bib.bib13),[3](https://arxiv.org/html/2608.08064#bib.bib12),[44](https://arxiv.org/html/2608.08064#bib.bib14)\], and spatiotemporal causal inference\[[6](https://arxiv.org/html/2608.08064#bib.bib20),[46](https://arxiv.org/html/2608.08064#bib.bib42),[25](https://arxiv.org/html/2608.08064#bib.bib46),[38](https://arxiv.org/html/2608.08064#bib.bib43),[27](https://arxiv.org/html/2608.08064#bib.bib44),[26](https://arxiv.org/html/2608.08064#bib.bib45),[2](https://arxiv.org/html/2608.08064#bib.bib21)\]addresses related challenges with different data regimes\. None of these directly target the joint problem of disaggregating outcomes*and*estimating local causal effects from coarse interventions\.
##### Invariant Mechanisms, GNNs, and Hierarchical Models\.
Conceptually,Clamrelies on the principle of*invariant causal mechanisms*\[[31](https://arxiv.org/html/2608.08064#bib.bib36),[36](https://arxiv.org/html/2608.08064#bib.bib37),[16](https://arxiv.org/html/2608.08064#bib.bib38),[34](https://arxiv.org/html/2608.08064#bib.bib39)\]: the local mechanismf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)is shared across subregions while contextual covariates drive heterogeneity\. This turns aggregation into a constraint and connects our architecture to message\-passing GNNs\[[12](https://arxiv.org/html/2608.08064#bib.bib34)\]and Deep Sets\[[43](https://arxiv.org/html/2608.08064#bib.bib35)\], as well as to hierarchical Bayesian models that pool information across groups\[[10](https://arxiv.org/html/2608.08064#bib.bib40),[7](https://arxiv.org/html/2608.08064#bib.bib41)\]\. It also placesClamin the setting of learning from aggregate observations\[[45](https://arxiv.org/html/2608.08064#bib.bib10)\], whose training loss coincides with ours under mean aggregation \(cf\. Appendix[A](https://arxiv.org/html/2608.08064#A1)\)\.
## 6Experiments
Figure 3:Results Overview\.Top:Input of LR intervention locations \(red\), LR outcomes \(blue\), and HR context \(green\)\.Bottom:Selected inference results\.We evaluateClamon three synthetic studies and one real\-world study\. In the synthetic studies the ground truth mechanism is available isolate distinct capabilities of the framework under controlled conditions: heterogeneous local\-effect recovery, latent intervention localization, and the reconstruction of unobserved covariates\. The real\-world study then probes the method on U\.S\. gun\-violence data, where we assume that only coarse regional outcomes are observed\.
Full experimental details, per\-experiment figures, and additional analyses are provided in Appendices[D](https://arxiv.org/html/2608.08064#A4)and[E](https://arxiv.org/html/2608.08064#A5), with the ablation study included in Appendix[D](https://arxiv.org/html/2608.08064#A4)\. Two additional studies—covering confounded treatment allocation and an unknown aggregation function—are deferred to Appendix[F](https://arxiv.org/html/2608.08064#A6)\. Our focus is qualitative: we aim to characterize the behavior of the method\. Figure[3](https://arxiv.org/html/2608.08064#S6.F3)provides an overview of the inputs and headline results\. Code that runs off\-the\-shelf on Google Colab is provided\.
### 6\.1Synthetic Studies
##### Exp\. 1: Political Campaigning\.
We first ask whetherClamcan recover heterogeneous local causal effects from aggregated outcomes when the effect of treatment depends on subregional context\. We model a binary regional intervention \(campaign spending\) whose impact varies with subregional wealth \(z\-scored income values\)\. Regional\-level covariates largely remain uninformative by construction\. The reported outcome is the candidate’s relative improvement, observed only at the regional level\. We compare against a naïvebaseline\(Figure[3](https://arxiv.org/html/2608.08064#S6.F3), with confidence intervals in the Appendix\) that applies a uniform disaggregation mechanism to the regional outcome and then learns the causal mechanism\.
In the limiting case where all regional outcomes are identical, uniform, bicubic, and ecological\-regression\[[13](https://arxiv.org/html/2608.08064#bib.bib8)\]baselines necessarily reach the same LOCATE error floor: without informative HR context, they cannot distinguish subregions and can at best assign the regional average effect\. Exp\. 1 is constructed close to this regime, with only minimal regional outcome variation, so explicitly evaluating these additional baselines provides little additional information\.Clamoperates below this floor because the shared mechanism ties each outcome to the context of its subregion\. This is the gain hypothesized in the introduction, where treating disaggregation and causal inference separately caps what is achievable\.
We use a simple neural network to estimate the causal mechanism\.Clamaccurately recovers the underlying subregional causal effects, with the MAE to the ground\-truth LOCATE matrix \(cf\. Eq\. \([4](https://arxiv.org/html/2608.08064#S3.E4)\)\) decreasing alongside the training loss\. The learned model supports coherent counterfactual reasoning at both regional and subregional resolutions and correctly identifies regions whose outcome differences arise purely from contextual heterogeneity rather than mean effects\. Full setup and per\-subregion visualizations are given in Appendix[D\.1](https://arxiv.org/html/2608.08064#A4.SS1)\(Figure[5](https://arxiv.org/html/2608.08064#A4.F5)\)\.
##### Exp\. 2: Public School Funding\.
Next, we test whetherClamcan recover*latent intervention locations*jointly with local causal effects when only aggregated intervention totals and outcomes are observed\. Each of the30×3030\\times 30regions consists of4×44\\times 4subregions and receives binary funding in either no subregion or exactly one unobserved subregion\. Treatment effects vary with subregional socioeconomic status, and the outcome is education\-score improvement\. We implementfθ\(⋅\)f\_\{\\theta\}\(\\cdot\)using either a parametric function or a small neural network\. In both cases,Clamjointly identifies the treated subregions and accurately recovers the underlying causal\-effect function\. To sample a treated subregion, we use the common Gumbel\-Max trick\. Details are provided in Appendix[D\.2](https://arxiv.org/html/2608.08064#A4.SS2)and Figure[7](https://arxiv.org/html/2608.08064#A4.F7)\.
##### Exp\. 3: Spatiotemporal Effects of Heat Waves\.
Finally, we ask whetherClamcan reconstruct an*unobserved spatial effect modifier*from aggregated outcomes in a spatiotemporal setting\. Heat waves negatively affect school performance, with heterogeneous impact depending on parental education \(observed\) and vegetation coverage \(unobserved\), the latter mitigating heat exposure\. Only region\-level outcomes are observed over time\. Under a correctly specified linear mechanism, both the regional prediction loss and the MSE converge rapidly, and the estimated disaggregation of the vegetation is visually indistinguishable from the ground\-truth vegetation map\. With a flexible MLP,Clamstill recovers meaningful spatial structure but with substantially higher error\. This reflects an identifiability trade\-off whereby expressive models can absorb transformations of the vegetation while preserving aggregated predictions\. Details are provided in Appendix[D\.3](https://arxiv.org/html/2608.08064#A4.SS3)\(Figure[8](https://arxiv.org/html/2608.08064#A4.F8)\)\.
##### Takeaways\.
Across the three studies,Clamperforms well in the synthetic setting, particularly for estimating causal effects\. The recovery of latent variables, such as intervention locations or unobserved effect modifiers, depends on the specific experiment\. Overfitting may occur, when large amounts of data is available\.
### 6\.2Semi\-Synthetic Case Study With Real\-World Data: Heat and Gun Violence
Figure 4:Side\-by\-side comparison of the learned rate functionsf𝜽\(temp,urban\)f\_\{\\boldsymbol\{\\theta\}\}\(\\text\{temp\},\\text\{urban\}\)for the disaggregation fromClam\(left\) and the oracle \(a NN with the ground truth subregional data; right\)\. Predictions are generated across a temperature range of\[−15,45\]∘C\[\-15,45\]^\{\\circ\}\\mathrm\{C\}and an urbanization scale of\[0,1\]\[0,1\]\.We evaluateClamon U\.S\. gun\-violence data from the Gun Violence Archive \(GVA\) for 2022–2023\[[15](https://arxiv.org/html/2608.08064#bib.bib49)\], motivated by prior evidence linking higher ambient temperatures to increased firearm violence\[[22](https://arxiv.org/html/2608.08064#bib.bib51)\]\. We treat temperature as a continuous exposure\. Urbanization is the high\-resolution \(HR\) contextual covariate, and gun\-incident counts are observed at weekly resolution\. A core difficulty in evaluating disaggregation methods on real data is the absence of ground truth at the fine scale\. We work around this by starting from data that is natively HR \(incident counts at the cell level\) and*artificially*aggregating it into LR regional counts, which serve as the inputs toClam\. To construct a high\-quality referencebaseline, the original HR counts are only used for fitting an “oracle” that we use for evaluation\.
We useN=100N=100regions andM=100M=100cells per region, indexed by regionii, subregionjj, and weekww\. LetYi,wLRY^\{\\mathrm\{LR\}\}\_\{i,w\}be the aggregated regional count in regioniiduring weekw≤52w\\leq 52,ci,jc\_\{i,j\}the HR urbanization covariate, andt~i,w\\widetilde\{t\}\_\{i,w\}the standardized regional temperature\. We model subregion\-level rates asλi,j,w=f𝜽\(t~i,w,ci,j\)\\lambda\_\{i,j,w\}=f\_\{\\boldsymbol\{\\theta\}\}\(\\widetilde\{t\}\_\{i,w\},c\_\{i,j\}\)and aggregate them asμi,w=∑j=1Mλi,j,w\\mu\_\{i,w\}=\\sum\_\{j=1\}^\{M\}\\lambda\_\{i,j,w\}, withYi,wLR∼Poisson\(μi,w\)Y^\{\\mathrm\{LR\}\}\_\{i,w\}\\sim\\mathrm\{Poisson\}\(\\mu\_\{i,w\}\)\. Training minimizes the regional Poisson negative log\-likelihood\. Figure[4](https://arxiv.org/html/2608.08064#S6.F4)shows thatClamrecovers a qualitatively similar temperature–urbanization interaction to the oracle\. The spatial comparison in Appendix Figure[9](https://arxiv.org/html/2608.08064#A5.F9)shows the same pattern: plausible structure, but inflated and more diffuse local intensities, reflecting an aggregation\-induced identifiability gap\.
## 7Limitations
Clamrelies on assumptions that are necessary given the nature of inferring fine\-grained causal effects from aggregated data, but are nonetheless restrictive\. We discuss the main limitations here\.
##### Prerequisite: Do\-Calculus Identifiability\.
Clamis not a substitute for causal identification\. The method should only be applied in settings where do\-calculus \(applied to a known or expert\-specified causal graph\) establishes that the causal effect is identifiable in principle\. If the causal graph is unknown or the required identification conditions fail \(e\.g\., the backdoor criterion is violated by unmodeled confounders\),Clamwill fit a model to the data that does not carry a causal interpretation\.
##### No Identification Guarantee from Optimization Alone\.
Even when do\-calculus certifies identifiability in principle, we do not provide a formal guarantee that our optimization procedure recovers the true causal effect\. The training objective is non\-convex, and the best we can reasonably claim is that, under good initialization and sufficient covariate diversity, gradient descent finds a local minimum whose value is close to the true causal effect\. It is easy to come up with pathological cases in which different fine\-grained disaggregations are equally consistent with the observed regional outcomes\. How to detect or avoid them in practice, remains an open question\.
##### Assumptions and Their Scope\.
Successful recovery of the local mechanism generally relies on \(i\) an invariant causal mechanism across subregions, \(ii\) sufficient diversity in subregional covariate distributions across regions, \(iii\) no hidden confounding, or its explicit modeling, and \(iv\) a known or sufficiently constrained aggregation function\.
The invariance assumption \(i\) is less restrictive than it may appear: althoughf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)is shared across regions and subregions, region\-specific dynamics can still be captured through contextual covariates such as region indicators or positional encodings\. Apparent heterogeneity across regions is absorbed by the context, yielding a*conditionally*invariant model\. The stronger requirement is that the contextual covariates provide sufficient variation to constrain the shared mechanism and capture sources of heterogeneity that are correlated with the treatment\. Unobserved factors that are uncorrelated with the treatment can effectively act as additional noise: they make recovery harder but need not bias the estimated effect\. Unobserved confounders, by contrast, will generally lead to incorrect causal estimates unless they are explicitly modeled\. The aggregation function \(iv\) need not be perfectly known in advance: as demonstrated in Exp\. 4 \(Appendix[F](https://arxiv.org/html/2608.08064#A6)\), it can be jointly learned alongside the local mechanism when it is sufficiently parameterized and the available variation allows the two components to be distinguished\.
##### Identifiability Gap and Latent Variable Recovery\.
Even when all assumptions hold, aggregation induces an intrinsic identifiability gap: matching aggregated predictions to aggregated observations may admit multiple equally valid solutions\. Our real\-world study illustrates this, withClamsystematically inflating local effect magnitudes relative to the cell\-supervised oracle\. The problem is amplified when latent variables are jointly inferred \(e\.g\., intervention locations in Exp\. 2, unobserved covariates in Exp\. 3\), since different inferred values can each correspond to a mechanism that explains the data equally well\. We discuss this further in Appendix[A](https://arxiv.org/html/2608.08064#A1)and regard a rigorous treatment, together with practical diagnostics, as the most important direction for future work\.
##### Limited Real\-World Validation\.
Our evaluation is primarily qualitative and relies substantially on synthetic studies, where the data\-generating assumptions are known and can be matched to the model\. The real\-world study uses a subregion\-supervised oracle as a reference rather than a true causal benchmark\. However, the identified causal effect of temperature on gun\-violence intensity is also supported by prior literature\[[22](https://arxiv.org/html/2608.08064#bib.bib51)\]\. Robustness to noise, measurement error, and treatment–context dependence is explored in the additional experiments, but a more systematic benchmarking remains an important direction for future work\.
## 8Conclusions and Future Work
The primary contribution of this work is not a single algorithm but the articulation of a*problem class*: estimating high\-resolution causal effects from low\-resolution interventions and outcomes by exploiting high\-resolution context\. We frame this as*causal deabstraction*, inverting the standard notion of causal abstraction and recasting spatial aggregation as a constraint satisfaction problem, thereby bridging structural causal models and spatial statistics\. In many real\-world settings such as public health, environmental policy, and education, spatial disaggregation and causal estimation are inseparable, since interventions are implemented at coarse scales but their effects manifest locally\. Treating them jointly within a structural causal model makes assumptions explicit, clarifies what is and is not identifiable, and turns causal inference into a tractable optimization problem\.
Clamis one concrete instantiation of this framework\. It explicitly adjusts for subregional treatment\-context confounding using only coarse regional supervision, a challenging problem in ecological inference, and jointly learns latent high\-resolution variables such as exact intervention locations and latent spatial effect modifiers alongside the causal mechanism\. Across synthetic and real\-world studies, we recover meaningful spatial and functional structure from aggregated data, while exposing the inherent limits of the setting: relative structure and causal dependencies are recoverable, but absolute local magnitudes may be systematically overestimated without fine\-grained outcome supervision\.
Because the problem class is, to our knowledge, new in this explicit form, principled baselines are scarce and a complete identification theory does not yet exist\. We see both as consequences of stating a new problem, and hope this paper motivates methods that supersede ours\.
Although this paper primarily proposes a research framework, misuse in high\-stakes domains such as epidemiology could cause harm if poorly identified local effects guided interventions or resource allocation; outputs should be accompanied by uncertainty estimates, sensitivity analyses, and domain expertise\. Disaggregation itself may also raise privacy concerns, since inferred high\-resolution patterns can reveal sensitive local information\.
Several directions stand out\. Theoretically, the most pressing need is a rigorous identifiability analysis: which combinations of aggregation function, contextual diversity, and structural restrictions guarantee recovery of the local mechanism, and which only up to monotone or scale transformations? Methodologically, principled regularization, hierarchical priors, or Bayesian formulations could control aggregation\-induced bias and provide uncertainty quantification\. Relaxing the shared\-mechanism assumption, integrating causal discovery over contextual variables, incorporating instrumental\-variable settings, and extending to spatiotemporal aggregation are natural next steps\. We hope that framingClamas a first concrete entry in a broader problem class makes these extensions easier to pursue\.
## References
- \[1\]\(2020\)A deep journey into super\-resolution: a survey\.ACM computing surveys \(CSUR\)53\(3\),pp\. 1–34\.Cited by:[§1](https://arxiv.org/html/2608.08064#S1.p1.1)\.
- \[2\]S\. V\. Balkus, S\. W\. Delaney, and N\. S\. Hejazi\(2024\)The causal effects of modified treatment policies under network interference\.arXiv preprint arXiv:2412\.02105\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px6.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[3\]S\. Beckers, F\. Eberhardt, and J\. Y\. Halpern\(2020\)Approximate causal abstractions\.InUncertainty in artificial intelligence,pp\. 606–615\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[4\]S\. Beckers and J\. Y\. Halpern\(2019\)Abstracting causal models\.InProceedings of the aaai conference on artificial intelligence,Vol\.33,pp\. 2678–2685\.Cited by:[§B\.3](https://arxiv.org/html/2608.08064#A2.SS3.p1.1),[Appendix B](https://arxiv.org/html/2608.08064#A2.p1.1),[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2608.08064#S1.p3.1),[§3\.5](https://arxiv.org/html/2608.08064#S3.SS5.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[5\]S\. L\. Chau, S\. Bouabid, and D\. Sejdinovic\(2021\)Deconditional downscaling with gaussian processes\.Advances in Neural Information Processing Systems34,pp\. 17813–17825\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px1.p1.1)\.
- \[6\]R\. Christiansen, M\. Baumann, T\. Kuemmerle, M\. D\. Mahecha, and J\. Peters\(2022\)Toward causal inference for spatio\-temporal data: conflict and forest loss in colombia\.Journal of the American Statistical Association117\(538\),pp\. 591–601\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px6.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[7\]R\. Dias, N\. L\. Garcia, and A\. M\. Schmidt\(2013\)A hierarchical model for aggregated functional data\.Technometrics55\(3\),pp\. 321–334\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px5.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px3.p1.1)\.
- \[8\]R\. Dutta and R\. Maity\(2020\)Identification of potential causal variables for statistical downscaling models: effectiveness of graphical modeling approach\.Theoretical and Applied Climatology142\(3\),pp\. 1255–1269\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p2.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px1.p1.1)\.
- \[9\]R\. E\. Fay and R\. A\. Herriot\(1979\)Estimates of income for small places: an application of James\-Stein procedures to census data\.Journal of the American Statistical Association74\(366\),pp\. 269–277\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p1.1)\.
- \[10\]A\. Feller and A\. Gelman\(2015\)Hierarchical models for causal effects\.Emerging trends in the social and behavioral sciences1,pp\. 16\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px5.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px3.p1.1)\.
- \[11\]D\. A\. Freedman\(1999\)Ecological inference and the ecological fallacy\.International Encyclopedia of the social & Behavioral sciences6\(4027\-4030\),pp\. 1–7\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p3.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[12\]J\. Gilmer, S\. S\. Schoenholz, P\. F\. Riley, O\. Vinyals, and G\. E\. Dahl\(2017\)Neural message passing for quantum chemistry\.InInternational conference on machine learning,pp\. 1263–1272\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px3.p1.2),[§3\.2](https://arxiv.org/html/2608.08064#S3.SS2.p3.3),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px3.p1.1)\.
- \[13\]L\. A\. Goodman\(1953\)Ecological regressions and behavior of individuals\.\.American sociological review18\(6\),pp\. 663\.Cited by:[§6\.1](https://arxiv.org/html/2608.08064#S6.SS1.SSS0.Px1.p2.1)\.
- \[14\]C\. A\. Gotway and L\. J\. Young\(2002\)Combining incompatible spatial data\.Journal of the American Statistical Association97\(458\),pp\. 632–648\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p1.1)\.
- \[15\]Gun Violence Archive\(2015\)Gun violence archive gva\.Note:Web ArchiveUnited States\. Retrieved from Library of Congress: https://www\.loc\.gov/item/lcwaN0016293/External Links:[Link](https://www.loc.gov/item/lcwaN0016293/)Cited by:[§6\.2](https://arxiv.org/html/2608.08064#S6.SS2.p1.1)\.
- \[16\]C\. Heinze\-Deml, J\. Peters, and N\. Meinshausen\(2018\)Invariant causal prediction for nonlinear models\.Journal of Causal Inference6\(2\),pp\. 20170016\.Cited by:[§B\.2](https://arxiv.org/html/2608.08064#A2.SS2.p1.3),[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px4.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px3.p1.1)\.
- \[17\]G\. King\(1997\)A solution to the ecological inference problem: reconstructing individual behavior from aggregate data\.Princeton University Press\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p3.1)\.
- \[18\]B\. Kumar, K\. Atey, B\. B\. Singh, R\. Chattopadhyay, N\. Acharya, M\. Singh, R\. S\. Nanjundiah, and S\. A\. Rao\(2023\)On the modern deep learning approaches for precipitation downscaling\.Earth Science Informatics16\(2\),pp\. 1459–1472\.Cited by:[§1](https://arxiv.org/html/2608.08064#S1.p1.1)\.
- \[19\]H\. C\. L\. Law, D\. Sejdinovic, E\. Cameron, T\. C\.D\. Lucas, S\. Flaxman, K\. Battle, and K\. Fukumizu\(2018\)Variational learning on aggregate outputs with Gaussian processes\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p1.1)\.
- \[20\]X\. Li, H\. Zhang, and X\. Zhou\(2026\)Spatio\-temporal hierarchical causal models\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 23230–23238\.Cited by:[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[21\]J\. Liley, S\. Emerson, B\. Mateen, C\. Vallejos, L\. Aslett, and S\. Vollmer\(2021\)Model updating after interventions paradoxically introduces bias\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 3916–3924\.Cited by:[§4](https://arxiv.org/html/2608.08064#S4.SS0.SSS0.Px7.p1.1)\.
- \[22\]V\. H\. Lyons, E\. L\. Gause, K\. R\. Spangler, G\. A\. Wellenius, and J\. Jay\(2022\)Analysis of daily ambient temperature and firearm violence in 100 us cities\.JAMA network open5\(12\),pp\. e2247207–e2247207\.Cited by:[§6\.2](https://arxiv.org/html/2608.08064#S6.SS2.p1.1),[§7](https://arxiv.org/html/2608.08064#S7.SS0.SSS0.Px5.p1.1)\.
- \[23\]D\. Maraun\(2019\)Statistical downscaling for climate science\.InOxford Research Encyclopedia of Climate Science,Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px1.p1.1)\.
- \[24\]T\. C\. Matisziw, T\. H\. Grubesic, and H\. Wei\(2008\)Downscaling spatial structure for the analysis of epidemiological data\.Computers, Environment and Urban Systems32\(1\),pp\. 81–93\.Cited by:[§1](https://arxiv.org/html/2608.08064#S1.p1.1)\.
- \[25\]M\. Mukaigawara, K\. Imai, J\. Lyall, and G\. Papadogeorgou\(2025\)Spatiotemporal causal inference with arbitrary spillover and carryover effects\.arXiv preprint arXiv:2504\.03464\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px6.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[26\]U\. Ninad, J\. Wahl, A\. Gerhardus, and J\. Runge\(2025\)Causal discovery on vector\-valued variables and consistency\-guided aggregation\.arXiv preprint arXiv:2505\.10476\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px6.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[27\]M\. Oprescu, D\. Park, X\. Luo, S\. Yoo, and N\. Kallus\(2026\)GST\-unet: a neural framework for spatiotemporal causal inference with time\-varying confounding\.Advances in Neural Information Processing Systems38,pp\. 17295–17322\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px6.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[28\]S\. Patil, N\. Pflugradt, J\. M\. Weinand, D\. Stolten, and J\. Kropp\(2024\)A systematic review of spatial disaggregation methods for climate action planning\.Energy and AI17,pp\. 100386\.Cited by:[§1](https://arxiv.org/html/2608.08064#S1.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px1.p1.1)\.
- \[29\]J\. Pearl\(2009\)Causality\.Cambridge university press\.Cited by:[§1](https://arxiv.org/html/2608.08064#S1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.08064#S1.p2.1)\.
- \[30\]J\. Peng, A\. Loew, O\. Merlin, and N\. E\. Verhoest\(2017\)A review of spatial downscaling of satellite remotely sensed soil moisture\.Reviews of Geophysics55\(2\),pp\. 341–366\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p1.1)\.
- \[31\]J\. Peters, P\. Bühlmann, and N\. Meinshausen\(2016\)Causal inference by using invariant prediction: identification and confidence intervals\.Journal of the Royal Statistical Society Series B: Statistical Methodology78\(5\),pp\. 947–1012\.Cited by:[§B\.2](https://arxiv.org/html/2608.08064#A2.SS2.p1.3),[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px4.p1.1),[§3\.5](https://arxiv.org/html/2608.08064#S3.SS5.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px3.p1.1)\.
- \[32\]W\. S\. Robinson\(1950\)Ecological correlations and the behavior of individuals\.American Sociological Review15\(3\),pp\. 351–357\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p3.1)\.
- \[33\]D\. F\. Rogers, R\. D\. Plante, R\. T\. Wong, and J\. R\. Evans\(1991\)Aggregation and disaggregation techniques and methodology in optimization\.Operations research39\(4\),pp\. 553–582\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p3.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[34\]M\. Rojas\-Carulla, B\. Schölkopf, R\. Turner, and J\. Peters\(2018\)Invariant models for causal transfer learning\.Journal of Machine Learning Research19\(36\),pp\. 1–34\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px4.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px3.p1.1)\.
- \[35\]J\. Runge, S\. Bathiany, E\. Bollt, G\. Camps\-Valls, D\. Coumou, E\. Deyle, C\. Glymour, M\. Kretschmer, M\. D\. Mahecha, J\. Muñoz\-Marí,et al\.\(2019\)Inferring causation from time series in earth system sciences\.Nature communications10\(1\),pp\. 2553\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px2.p1.1)\.
- \[36\]B\. Schölkopf, F\. Locatello, S\. Bauer, N\. R\. Ke, N\. Kalchbrenner, A\. Goyal, and Y\. Bengio\(2021\)Toward causal representation learning\.Proceedings of the IEEE109\(5\),pp\. 612–634\.Cited by:[§B\.2](https://arxiv.org/html/2608.08064#A2.SS2.p1.3),[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px4.p1.1),[§3\.5](https://arxiv.org/html/2608.08064#S3.SS5.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px3.p1.1)\.
- \[37\]A\. A\. Schuessler\(1999\)Ecological inference\.Proceedings of the National Academy of Sciences96\(19\),pp\. 10578–10581\.Cited by:[§1](https://arxiv.org/html/2608.08064#S1.p1.1)\.
- \[38\]Z\. Song and G\. Papadogeorgou\(2024\)Bipartite causal inference with interference, time series data, and a random network\.arXiv preprint arXiv:2404\.04775\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px6.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[39\]Y\. Sun, K\. Deng, K\. Ren, J\. Liu, C\. Deng, and Y\. Jin\(2024\)Deep learning in statistical downscaling for deriving high spatial resolution gridded meteorological data: a systematic review\.ISPRS Journal of Photogrammetry and Remote Sensing208,pp\. 14–38\.Cited by:[§1](https://arxiv.org/html/2608.08064#S1.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px1.p1.1)\.
- \[40\]A\. Vázquez\-Patiño, E\. Samaniego, L\. Campozano, and A\. Avilés\(2022\)Effectiveness of causality\-based predictor selection for statistical downscaling: a case study of rainfall in an ecuadorian andes basin\.Theoretical and Applied Climatology150\(3\),pp\. 987–1013\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px1.p2.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px1.p1.1)\.
- \[41\]E\. N\. Weinstein and D\. M\. Blei\(2026\)Hierarchical causal models\.Journal of Machine Learning Research27\(37\),pp\. 1–73\.Cited by:[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[42\]X\. Ye, Q\. Huang, and W\. Li\(2016\)Integrating big social data, computing and modeling for spatial social science\.Vol\.43,Taylor & Francis\.Cited by:[§1](https://arxiv.org/html/2608.08064#S1.p1.1)\.
- \[43\]M\. Zaheer, S\. Kottur, S\. Ravanbakhsh, B\. Poczos, R\. R\. Salakhutdinov, and A\. J\. Smola\(2017\)Deep sets\.Advances in neural information processing systems30\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px3.p1.2),[§3\.2](https://arxiv.org/html/2608.08064#S3.SS2.p3.3),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px3.p1.1)\.
- \[44\]F\. M\. Zennaro\(2022\)Abstraction between structural causal models: a review of definitions and properties\.arXiv preprint arXiv:2207\.08603\.Cited by:[§B\.3](https://arxiv.org/html/2608.08064#A2.SS3.p1.1),[Appendix B](https://arxiv.org/html/2608.08064#A2.p1.1),[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px2.p1.1),[§3\.5](https://arxiv.org/html/2608.08064#S3.SS5.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
- \[45\]Y\. Zhang, N\. Charoenphakdee, Z\. Wu, and M\. Sugiyama\(2020\)Learning from aggregate observations\.Advances in Neural Information Processing Systems33,pp\. 7993–8005\.Cited by:[item 2](https://arxiv.org/html/2608.08064#A1.I1.i2.p1.1),[§A\.4](https://arxiv.org/html/2608.08064#A1.SS4.p1.1),[§A\.4](https://arxiv.org/html/2608.08064#A1.SS4.p2.2),[§A\.4](https://arxiv.org/html/2608.08064#A1.SS4.p4.2),[§A\.4](https://arxiv.org/html/2608.08064#A1.SS4.p6.1),[§A\.4](https://arxiv.org/html/2608.08064#A1.SS4.p7.1),[§3\.4](https://arxiv.org/html/2608.08064#S3.SS4.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px3.p1.1)\.
- \[46\]L\. Zhou, K\. Imai, J\. Lyall, and G\. Papadogeorgou\(2024\)Estimating heterogeneous treatment effects for spatio\-temporal causal inference: how economic assistance moderates the effects of airstrikes on insurgent violence\.arXiv preprint arXiv:2412\.15128\.Cited by:[Appendix C](https://arxiv.org/html/2608.08064#A3.SS0.SSS0.Px6.p1.1),[§5](https://arxiv.org/html/2608.08064#S5.SS0.SSS0.Px2.p1.1)\.
## Appendix AIdentifiability
Identification inClamdepends on the target of interest and on the model assumptions\. We distinguish causal identification under an assumed graph, statistical identification from aggregate observations, and finite\-sample recovery by the optimizer\. The last is an estimation problem rather than an identification question\.
A complete theory for statistical identification from aggregates is beyond the scope of this work\. Below, we characterize several tractable special cases and discuss the main remaining open problems\.
### A\.1Identifiability of What
First,*causal identification*asks whether a quantity such as LOCATE can be expressed in terms of the observed distribution under the assumed causal graph\. We treat the graph as given and do not consider causal discovery\. This is separate from whether the required fine\-grained mechanism can be recovered from aggregate observations\.
Second,*statistical identification*asks whether the local mechanismf\(⋅\)f\(\\cdot\)is uniquely determined by the population distribution of the aggregate data\. We are generally interested in the function itself rather than the parameter vector𝜽\\boldsymbol\{\\theta\}, since different parameterizations may represent the same function\.
Third, one can ask whether the high\-resolution outcome map\{yi,j\}\\\{y\_\{i,j\}\\\}is identified\. Even with a known mechanism, exact local outcomes may remain unknown because the individual noise realizationsϵi,j\\epsilon\_\{i,j\}and possibly latent inputsci,j′c^\{\\prime\}\_\{i,j\}are unobserved\. Conversely, aggregates can sometimes determine particular local outcomes without identifying the full mechanism\.
### A\.2General Non\-Identifiability
Without restrictions, neither the mechanism nor the high\-resolution outcome map is identified from aggregates\.
Consider a noise\-free setting similar to Equation \([2](https://arxiv.org/html/2608.08064#S3.E2)\) in which no region receives treatment\. Take two regionsAAandBB, each with two subregions, with context values\(0,1\)\(0,1\)and\(0\.5,0\.5\)\(0\.5,0\.5\), respectively, and suppose both observed regional aggregates equal11\.
If both the local mechanismf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)and the aggregation functiongϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\)are unknown, infinitely many pairs are compatible with these observations\. For example,f𝜽\(0,c\)=cf\_\{\\boldsymbol\{\\theta\}\}\(0,c\)=ctogether with sum aggregation yields subregional outcomes\(0,1\)\(0,1\)for regionAAand\(0\.5,0\.5\)\(0\.5,0\.5\)for regionBB\. Equally, the constant mechanismf𝜽\(0,c\)≡1f\_\{\\boldsymbol\{\\theta\}\}\(0,c\)\\equiv 1together with mean aggregation yields\(1,1\)\(1,1\)for both regions\. Both explain the observed aggregates while implying different high\-resolution outcomes\.
The ambiguity remains even when the aggregation function is known\. Supposegϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\)is a sum\. Thenf𝜽\(0,c\)=cf\_\{\\boldsymbol\{\\theta\}\}\(0,c\)=cis compatible with the data, but so isf𝜽\(0,c\)≡0\.5f\_\{\\boldsymbol\{\\theta\}\}\(0,c\)\\equiv 0\.5\. The former yields\(0,1\)\(0,1\)in regionAA, whereas the latter yields\(0\.5,0\.5\)\(0\.5,0\.5\), while both sum to11\.
### A\.3A Saturated View
A useful limiting case is to think off𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)as a dictionary lookup: each distinct observed input\(ti,j,ci,j\)\(t\_\{i,j\},c\_\{i,j\}\)is assigned an independent unknown value\. With known sum aggregation and no noise, the aggregate observations define a linear system
𝐘=A𝐯,\\mathbf\{Y\}=A\\mathbf\{v\},where𝐯\\mathbf\{v\}contains the unknown values off𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)at the observed inputs\. Identification is therefore a rank question: the local values are identified exactly whenAAhas full column rank\. Otherwise, multiple local\-value assignments produce the same aggregates\.
A neural network departs from this saturated case by tying values at different inputs together through its parameterization and inductive bias\. The linear model in Equation \([3](https://arxiv.org/html/2608.08064#S3.E3)\) imposes an even stronger restriction; there, identification analogously reduces to full rank of the corresponding low\-dimensional design matrix\.
### A\.4Aggregate Regression
The saturated view above is deliberately pessimistic because it treats the value at every distinct input as unrelated\. With a shared mechanism and sufficiently varied aggregate observations, much more can be identified\. This is closely related to learning from aggregate observations as studied byZhanget al\.\[[45](https://arxiv.org/html/2608.08064#bib.bib10)\]\.
The additive model in Equation \([2](https://arxiv.org/html/2608.08064#S3.E2)\) has a direct connection to their framework\. Consider the special case of known sum aggregation, fixedMM, additive homoscedastic Gaussian noise, no latentc′c^\{\\prime\}, and the sampling assumptions ofZhanget al\.\[[45](https://arxiv.org/html/2608.08064#bib.bib10)\]\. Since fixed\-size sum and mean aggregation differ only by a constant scaling, the resulting aggregate MSE is equivalent to their regression\-from\-mean objective\.
Let\[W0\]\[W\_\{0\}\]denote the equivalence class of all parameter values that are observationally equivalent toW0W\_\{0\}, meaning that they induce the same population likelihood\. Their Proposition 7 characterizes this class as
\[W0\]=\{W∈𝒲:f\(X;W\)=f\(X;W0\)a\.e\.\}\.\[W\_\{0\}\]=\\left\\\{W\\in\\mathcal\{W\}:f\(X;W\)=f\(X;W\_\{0\}\)\\ \\text\{a\.e\.\}\\right\\\}\.\(5\)Thus, the parameterizationW0W\_\{0\}need not be unique, but all observationally equivalent parameters represent the same function almost everywhere under the distribution ofXX\.
There is one subtlety in translating this result to our setting\. Treatment is assigned at the regional level and is therefore shared by all subregions in a region\. The result is therefore most naturally interpreted separately within each treatment arm\. For a fixedt∈\{0,1\}t\\in\\\{0,1\\\}, define
ft\(c\)=f\(t,c\)\.f\_\{t\}\(c\)=f\(t,c\)\.If the subregional contexts within that treatment arm satisfy the sampling assumptions ofZhanget al\.\[[45](https://arxiv.org/html/2608.08064#bib.bib10)\], Proposition 7 identifiesft\(⋅\)f\_\{t\}\(\\cdot\)almost everywhere under the corresponding context distribution\.
Consequently, recovering bothf0\(c\)f\_\{0\}\(c\)andf1\(c\)f\_\{1\}\(c\)does not automatically identify LOCATE everywhere\. The contrast
f1\(c\)−f0\(c\)f\_\{1\}\(c\)\-f\_\{0\}\(c\)is supported by this result only where the context distributions of the treated and untreated populations overlap\. Outside this overlap, a flexible model relies on interpolation or extrapolation induced by its function class rather than on the identification result\.
The independence assumptions are also restrictive for spatial data\. The model ofZhanget al\.\[[45](https://arxiv.org/html/2608.08064#bib.bib10)\]assumes conditional independence of the individual targets and treats the features within an aggregate as independent draws\. Their proof of Proposition 7 uses this independence to factorize the cross terms in the population loss\. Spatially structured covariate compositions need not satisfy these assumptions\.Zhanget al\.\[[45](https://arxiv.org/html/2608.08064#bib.bib10)\]themselves note that their independence assumption can fail when observations are collected group by group, explicitly naming spatially aggregated data as an example\. Their theory therefore provides a useful base case and suitable language for functional equivalence, but it does not cover the full spatial setting considered here\.
The counterexample above does not contradict this result: it uses only two fixed aggregate observations, which leave multiple mechanisms possible\. Under the sampling assumptions ofZhanget al\.\[[45](https://arxiv.org/html/2608.08064#bib.bib10)\], many independently varying aggregates provide additional constraints that identify the shared function in the population limit\.
### A\.5Open Problems
The cases above provide a general counterexample and tractable results for restricted settings\. Several important aspects of the general model remain open\.
##### Spatio\-Temporal Extensions and Latent Structure\.
Repeated spatio\-temporal observations provide different information from additional independent aggregate groups\. The same spatial units may recur over time with largely persistent spatial structure\. A latent quantityc′c^\{\\prime\}that is shared across observations therefore enters many aggregate constraints, whereas an idiosyncratic noise realizationϵi,j\\epsilon\_\{i,j\}enters only one observation\.
Repeated observations can consequently make persistent latent structure estimable while observation\-specific noise remains unrecoverable\. Exp\. 3 illustrates this empirically: the same latent vegetation field contributes to regional outcomes over many time points, allowing constraints on it to accumulate\. We do not claim a general identification theorem for this setting; recovery depends on temporal variation, the functional form, and how the latent variable enters the mechanism\.
##### Unknown Aggregation\.
Ifgϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\)is unknown and learned jointly withf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\), different combinations of the local mechanism and aggregation function may induce the same observed regional outcomes, as illustrated by the first counterexample\. Identification therefore requires restrictions on at least one of these functions, sufficient variation to distinguish them, or prior knowledge of the aggregation process\.
This is one reason why known sum or mean aggregation is substantially easier than the unrestricted model in Equation \([1](https://arxiv.org/html/2608.08064#S3.E1)\)\.
##### Nonlinear Aggregation and Noise\.
Linear aggregation has the convenient property that, with additive zero\-mean noise, aggregation commutes with expectation\. For a nonlinear aggregation function this generally fails:
gϕ\(\{𝔼\[yi,j∣T,C\]\}j\)≠𝔼\[gϕ\(\{yi,j\}j\)∣T,C\]\.g\_\{\\boldsymbol\{\\phi\}\}\\left\(\\left\\\{\\mathbb\{E\}\[y\_\{i,j\}\\mid T,C\]\\right\\\}\_\{j\}\\right\)\\neq\\mathbb\{E\}\\left\[g\_\{\\boldsymbol\{\\phi\}\}\\left\(\\\{y\_\{i,j\}\\\}\_\{j\}\\right\)\\mid T,C\\right\]\.\(6\)
Applying a nonlinear aggregator directly to noise\-free conditional\-mean predictions therefore does not generally target the conditional mean of the observed aggregate\. A statistically correct objective must propagate the noise through the aggregation, for example through the induced likelihood or a Monte Carlo approximation\. This removes the objective mismatch but does not by itself establish identification\. Iff𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\),gϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\), or the noise distribution are simultaneously unknown, further observational equivalences may remain\.
##### Causal Structure\.
The statistical arguments above concern recovery of a function from aggregate data and do not, by themselves, give that function a causal interpretation\. We do not attempt to discover the causal graph\. Instead, we assume a graph and ask whether, under that graph, the recovered function corresponds to the desired interventional quantity\.
For the baseline motif, whereCCis an observed effect modifier and treatment is unconfounded conditional on the relevant variables,f\(t,c\)f\(t,c\)can represent the conditional interventional mean required for LOCATE\. IfCCis a confounder, the treatment\-assignment mechanism must be handled appropriately\. IfCCis a mediator or collider, conditioning on it has a different causal meaning and may not correspond to the desired treatment effect\. Thus, changing the assumed causal structure need not change the aggregate regression problem mathematically, but it changes which recovered functions correspond to identified causal estimands\.
##### Identification Versus Optimization\.
Even when the population objective identifies the desired function, this does not imply that a finite\-sample non\-convex optimizer will recover it\. Neural\-network parameterization, jointly inferred latent variables, and learned aggregation functions may create multiple local optima or flat regions of the objective\. Conversely, failure of one optimizer to recover the ground truth does not itself imply non\-identifiability\.
We therefore treat optimization stability as an empirical question separate from identification\. In Exp\. 1, repeated random initializations produce similar LOCATE estimates when the aggregate loss is low, but this is evidence about the estimator in that experiment, not proof that the population problem is identified\.
### A\.6Practical Diagnostics
No finite\-sample diagnostic can prove structural identification\. Nevertheless, several checks can reveal that the aggregate observations do not sufficiently constrain the fine\-grained quantity of interest\.
We recommend comparing recovered LOCATE estimates across random restarts, architectures, regularization choices, and plausible aggregation or noise specifications\. Large disagreement is evidence of an underdetermined problem\. Agreement is reassuring but does not establish identification, since the same inductive bias may repeatedly select the same solution\.
External validation against higher\-resolution observations should be used whenever available\. Sensitivity analyses under plausible violations of the causal assumptions provide complementary information; our hidden\-confounding ablation is one example\. Bootstrap intervals quantify sampling variability but should not be interpreted as a test of structural identification\.
##### Summary of Claims\.
The discussion can be summarized in three levels\.
1. 1\.For the linear model in Equation \([3](https://arxiv.org/html/2608.08064#S3.E3)\), identification reduces to a rank condition on the aggregate design matrix\.
2. 2\.For the additive model in Equation \([2](https://arxiv.org/html/2608.08064#S3.E2)\), the i\.i\.d\. mean\-aggregation special case connects directly to the functional consistency result ofZhanget al\.\[[45](https://arxiv.org/html/2608.08064#bib.bib10)\]\. Within each treatment arm, their result identifies the conditional\-mean function almost everywhere under the corresponding context distribution\. Identification of LOCATE additionally requires sufficient overlap between the treatment arms\.
3. 3\.For the general model in Equation \([1](https://arxiv.org/html/2608.08064#S3.E1)\), including dependent spatial designs, structured latent variables, unknown or nonlinear aggregation, and extrapolation outside the observed treatment–context support, we do not claim a general identification theorem\. Recovery in these regimes is studied empirically and remains an important theoretical problem\.
Causal identification under the assumed graph is separate from all three statements: even a statistically identified fine\-grained mechanism has the desired causal interpretation only under the assumptions encoded by that graph\.
## Appendix BCausal Deabstraction Framework
In this section, we situate our work within the broader causal inference literature\. Specifically, we frame this as a novel instance of*causal deabstraction*\[[4](https://arxiv.org/html/2608.08064#bib.bib13),[44](https://arxiv.org/html/2608.08064#bib.bib14)\]\. Our methodology is informed by the framework of SCMs and is guided by the principle of ICM\.
### B\.1Structural Causal Framework
The structural causal model introduced in Section[3\.1](https://arxiv.org/html/2608.08064#S3.SS1)can be formalized in generative form as follows\. Each spatial unit \(region\)i∈\{1,…,N\}i\\in\\\{1,\\dots,N\\\}is composed ofMMsubregions indexed byj∈\{1,…,M\}j\\in\\\{1,\\dots,M\\\}\. For each subregion\(i,j\)\(i,j\), we posit the following SCM:
ci,j\\displaystyle c\_\{i,j\}observed \(exogenous\)ϵi,j\\displaystyle\\epsilon\_\{i,j\}∼Pϵ\\displaystyle\\sim P\_\{\\epsilon\}ti,j\\displaystyle t\_\{i,j\}=Ti∈\{0,1\}\\displaystyle=T\_\{i\}\\in\\\{0,1\\\}yi,j\\displaystyle y\_\{i,j\}=f𝜽\(ti,j,ci,j,ϵi,j\)\\displaystyle=f\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\},\\epsilon\_\{i,j\}\)Y^i\\displaystyle\\widehat\{Y\}\_\{i\}=gϕ\(\{yi,j\}j=1M\)\\displaystyle=g\_\{\\boldsymbol\{\\phi\}\}\(\\\{y\_\{i,j\}\\\}\_\{j=1\}^\{M\}\)Here,ci,jc\_\{i,j\}denotes a high\-resolution contextual covariate for subregionjjwithin regionii, which is observed and exogenous, meaning it is not caused byTTorYY, andϵi,j\\epsilon\_\{i,j\}represents exogenous noise drawn fromPϵP\_\{\\epsilon\}\. The binary interventionti,j=Ti∈\{0,1\}t\_\{i,j\}=T\_\{i\}\\in\\\{0,1\\\}is assigned at the regional level and shared across all subregions\. The functionf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)encodes the local causal mechanism, shared across all subregions, that maps the intervention, context, and noise to the subregional outcomeyi,jy\_\{i,j\}\. The aggregation functiongϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\)maps the collection of subregional outcomes\{yi,j\}j=1M\\\{y\_\{i,j\}\\\}\_\{j=1\}^\{M\}to the regional outcomeY^i\\widehat\{Y\}\_\{i\}\.
In the main section, we considered simplified cases, such as sum aggregation, additive noise, and the absence of hidden confounders, to provide intuition\. The formulation here generalizes that view by expressing the structural causal model in explicit stochastic form, specifying how interventions, contextual variables, and noise are generated, while making clear that bothf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)andgϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\)are shared across regions\.
### B\.2Invariant Causal Mechanism Assumption
A central assumption inClamis that the causal functionf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)is*invariant*across all subregions and regions\. That is, while the covariates\{ci,j\}\\\{c\_\{i,j\}\\\}may vary across regions due to differences in local composition, the functional formf𝜽\(t,c\)f\_\{\\boldsymbol\{\\theta\}\}\(t,c\)does not change\. This assumption reflects the principle of*invariant causal mechanisms*\[[36](https://arxiv.org/html/2608.08064#bib.bib37),[31](https://arxiv.org/html/2608.08064#bib.bib36),[16](https://arxiv.org/html/2608.08064#bib.bib38)\], which posits that causal relations are stable across environments unless directly intervened upon\.
We treat each regioniias an*environment*defined by its observed covariate compositionci,j∗j=1M\{c\_\{i,j\}\}\*\{j=1\}^\{M\}, but governed by a shared mechanismf∗𝜽\(⋅\)f\*\{\\boldsymbol\{\\theta\}\}\(\\cdot\)\. The variability in covariate composition across regions provides information for estimatingf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)from the observed aggregated outcomesY^i\\widehat\{Y\}\_\{i\}\.
### B\.3Deabstraction via Invariance
We propose to view the task of estimatingf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)from aggregated outcomes as an instance of*causal deabstraction*, that is, inferring a latent, high\-resolution causal model that is consistent with a lower\-resolution aggregated model\. This perspective is dual to recent work on causal abstraction\[[4](https://arxiv.org/html/2608.08064#bib.bib13),[44](https://arxiv.org/html/2608.08064#bib.bib14)\], which studies when a coarse model preserves the counterfactual semantics of a fine\-grained one\. Here, we reverse the direction: we seek to recover a fine\-grained causal mechanism that is*compatible*with the observed coarse\-level effects\.
This is possible due to two key properties:
1. 1\.The aggregation functiongϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\)is consistent and applied uniformly across regions\.
2. 2\.The observed covariate compositions\{ci,j\}j=1M\\\{c\_\{i,j\}\\\}\_\{j=1\}^\{M\}differ across regions, which induces identifiable variation in the aggregatesY^i\\widehat\{Y\}\_\{i\}under the shared mechanismf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)\.
Given sufficient diversity and the assumption of causal invariance, the aggregated outputs provide a supervisory signal to recover the latent causal function \(although we do not provide a general correctness guarantee\)\.
### B\.4Causal Semantics via Do\-Calculus
Under the assumption of no hidden confounding, the causal effect of the interventionTiT\_\{i\}on a subregional outcomeyi,jy\_\{i,j\}is identified via the backdoor criterion:
𝔼\[yi,j∣do\(Ti=t\),ci,j\]=∫f𝜽\(t,ci,j,ϵi,j\)p\(ϵi,j\)𝑑ϵi,j\\mathbb\{E\}\[y\_\{i,j\}\\mid do\(T\_\{i\}=t\),\\,c\_\{i,j\}\]=\\int f\_\{\\boldsymbol\{\\theta\}\}\(t,c\_\{i,j\},\\epsilon\_\{i,j\}\)\\,p\(\\epsilon\_\{i,j\}\)\\,d\\epsilon\_\{i,j\}In Exps 1–4, we assumeTi⟂⟂ϵi,j∣ci,jT\_\{i\}\\perp\\\!\\\!\\\!\\perp\\epsilon\_\{i,j\}\\mid c\_\{i,j\}, and thatTiT\_\{i\}is either randomized or independent ofCi,jC\_\{i,j\}\. Under this setting, the expectation can be approximated via observational data\. However, because only the aggregate outcomeY^i\\widehat\{Y\}\_\{i\}is observed, we must rely on the consistency constraint:
Y^i≈∑j=1Mf𝜽\(Ti,ci,j\)\\widehat\{Y\}\_\{i\}\\approx\\sum\_\{j=1\}^\{M\}f\_\{\\boldsymbol\{\\theta\}\}\(T\_\{i\},c\_\{i,j\}\)The training procedure thus finds the functionf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)that jointly explains all observed aggregates under this constraint, while enforcing thatf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)is invariant across regions\. In this sense, the estimation off𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)becomes a constraint satisfaction problem, where causal invariance and aggregation consistency define the feasible solution space\. In Exp 5, this conditional independence assumption is deliberately violated by construction, and valid estimation requires explicitly modeling the treatment allocation mechanism\.
### B\.5Implications
The assumption of an invariant subregional mechanism plays a central role in enabling causal inference in our disaggregated setting\. By treating the coarse observations as aggregates of fine\-grained causal effects and leveraging variation in covariate compositions,Clamperforms inference not merely on statistical associations but on causal mechanisms that are meaningful and robust to changes in population structure\. All in all, we contribute a formal and algorithmic instantiation of*causal deabstraction*from low\-resolution outcomes and high\-resolution covariates\.
## Appendix CExtended Related Work
Our work builds on and extends three main strands of research: spatial disaggregation, causal inference under aggregation, and causal abstraction\.
##### Statistical Downscaling and Disaggregation\.
Statistical downscaling methods aim to infer high\-resolution estimates from coarse data, using spatial interpolation, small area estimation\[[9](https://arxiv.org/html/2608.08064#bib.bib7)\], or learning\-based techniques\. These have been widely applied in climate and environmental sciences\[[23](https://arxiv.org/html/2608.08064#bib.bib1),[30](https://arxiv.org/html/2608.08064#bib.bib9),[14](https://arxiv.org/html/2608.08064#bib.bib4)\], including recent work on probabilistic downscaling using Gaussian processes\[[5](https://arxiv.org/html/2608.08064#bib.bib11),[19](https://arxiv.org/html/2608.08064#bib.bib3)\]\. However, such methods typically model associational patterns and do not support interventional or counterfactual reasoning\.
To address this, some recent approaches integrate causal reasoning into downscaling pipelines, e\.g\., by identifying causally relevant predictors\[[40](https://arxiv.org/html/2608.08064#bib.bib15),[8](https://arxiv.org/html/2608.08064#bib.bib16)\]\. Yet, these still rely on observable variables and assume access to relatively fine\-grained information\. In contrast, our approach estimates high\-resolutioncausal effectsusing only aggregated outcome data and high\-resolution covariates\.
The challenges of reasoning with aggregate data are well\-documented\. Classical results on the ecological fallacy show that aggregated statistics may obscure or even invert causal relationships\[[32](https://arxiv.org/html/2608.08064#bib.bib5),[11](https://arxiv.org/html/2608.08064#bib.bib17),[17](https://arxiv.org/html/2608.08064#bib.bib6)\]\. Earlier work on aggregation and disaggregation in optimization similarly highlights the complexity introduced when latent heterogeneity is present but unobserved\[[33](https://arxiv.org/html/2608.08064#bib.bib19)\]\.Clamexplicitly addresses these challenges by modeling the aggregation process and learning disentangled causal effects at the subregional level\.
##### Causal Deabstraction\.
Structurally, our approach is inspired by recent work on causal abstraction, which studies mappings between causal models at different levels of granularity\[[4](https://arxiv.org/html/2608.08064#bib.bib13),[3](https://arxiv.org/html/2608.08064#bib.bib12),[44](https://arxiv.org/html/2608.08064#bib.bib14)\]\. These works formalize when such mappings preserve counterfactual or interventional semantics, but do not address learning disaggregated effects from data\. Finally, our work relates to spatiotemporal causal discovery\[[35](https://arxiv.org/html/2608.08064#bib.bib18)\], but differs by operating in a setting where temporal data is limited and only aggregate observations are available\.
##### Graph Neural Networks\.
Conceptually, our architecture is related to message\-passing graph neural networks\[[12](https://arxiv.org/html/2608.08064#bib.bib34)\]\. In a GNN, pairwise messages are computed and aggregated within a single update step\. Similarly, our model learns an update functionf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)that operates on ordered inputs, and optionally an aggregation functiongϕ\(⋅\)g\_\{\\boldsymbol\{\\phi\}\}\(\\cdot\)that is permutation\-invariant, akin to a Deep Sets model\[[43](https://arxiv.org/html/2608.08064#bib.bib35)\]\. As in GNNs, we share the parameters of the structural function across nodes or regions\.
##### Invariant Causal Mechanisms\.
Our work also relates to the principle of*invariant causal mechanisms*, which posits that the functional relationships between variables remain stable across different environments or interventions\[[31](https://arxiv.org/html/2608.08064#bib.bib36),[36](https://arxiv.org/html/2608.08064#bib.bib37)\]\. This idea has been used for causal discovery\[[16](https://arxiv.org/html/2608.08064#bib.bib38)\]and domain generalisation\[[34](https://arxiv.org/html/2608.08064#bib.bib39)\], and is particularly relevant for spatial disaggregation where intervention effects may vary locally but share stable dependencies on contextual covariates\.Clamincorporates this principle by enforcing that the learned local causal mechanismf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)is shared across all subregions while allowing spatial heterogeneity to emerge through contextual variables\.
##### Hierarchical Bayesian Model\.
The design also parallels hierarchical Bayesian models\[[10](https://arxiv.org/html/2608.08064#bib.bib40),[7](https://arxiv.org/html/2608.08064#bib.bib41)\], where group\-specific effects are drawn from a shared prior distribution and information is partially pooled across groups\. InClam, the local mechanismf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)plays the role of the shared prior, and contextual covariates capture structured variation across subregions, allowing data\-sparse subregions to benefit from patterns learned in others\.
##### Spatiotemporal Causality\.
Spatiotemporal causal inference has received growing attention as researchers seek to understand how interventions and their effects evolve across both space and time\. A recent work presents a causal framework for studying the interplay between conflict and forest loss in Colombia, emphasizing the need to model spatial and temporal dependencies jointly\[[6](https://arxiv.org/html/2608.08064#bib.bib20)\]\. Extending this perspective, methods proposed in\[[46](https://arxiv.org/html/2608.08064#bib.bib42)\]estimate heterogeneous treatment effects in settings where economic assistance moderates the impact of airstrikes on insurgent violence, while approaches introduced in\[[25](https://arxiv.org/html/2608.08064#bib.bib46)\]accommodate arbitrary spill\-over and carryover effects across space and time\. Network\-based settings with temporal dynamics have also been explored\. In\[[38](https://arxiv.org/html/2608.08064#bib.bib43)\]the authors study bipartite causal inference with interference in time series data, and in\[[2](https://arxiv.org/html/2608.08064#bib.bib21)\]methods are developed to analyse the causal effects of modified treatment policies under network interference\. Recent methodological innovations include GST\-UNet\[[27](https://arxiv.org/html/2608.08064#bib.bib44)\], a deep learning architecture for spatiotemporal causal inference under time\-varying confounding, and consistency\-guided aggregation for causal discovery in multivariate spatiotemporal data introduced in\[[26](https://arxiv.org/html/2608.08064#bib.bib45)\]\. Together, these works highlight the diversity of approaches to spatiotemporal causality, spanning structural models, interference\-aware estimation, and machine learning\-driven inference\.
## Appendix DDetails of Main Experiments
### D\.1Exp\. 1: Political Campaigning
##### Context\.
We consider a simple political campaigning setting in which a candidate either spends campaign resources in a region or does not\. The treatment is binary at the regional level and is inherited by all subregions\. The outcome is the candidate’s relative improvement, but it is observed only as a regional aggregate\. The effect of campaigning depends on subregional context, here interpreted as standardized wealth\. Thus, two regions can have similar aggregate context and similar aggregate outcomes while still containing different subregional causal effects\.
Figure 5:Evaluation of the learned outcomes, causal effect estimates, counterfactual predictions, and training behavior\.
##### Setup\.
We useN=100N=100regions arranged on a10×1010\\times 10grid\. Each region containsM=16M=16subregions arranged on a4×44\\times 4grid\. The treatment matrix isT∈\{0,1\}N×MT\\in\\\{0,1\\\}^\{N\\times M\}\. Exactly half of the regions are treated, and within a treated region all subregions have treatment value11; within a control region all subregions have treatment value0\.
The context matrixC∈ℝN×MC\\in\\mathbb\{R\}^\{N\\times M\}contains one continuous contextual value per subregion\. We sample the entries from a standard Gaussian distribution and then subtract the mean within each region, so that each row ofCChas mean zero\. This makes the regional average context uninformative by construction\. We then randomly swap entries across the matrix to introduce additional variation while preserving the overall structure\. The observed outcome is generated at the subregional level and then aggregated to the region level by averaging\.
The ground\-truth outcome function is
f\(t,c\)=\{1,t=0,interp\(c\),t=1,f\(t,c\)=\\begin\{cases\}1,&t=0,\\\\ \\operatorname\{interp\}\(c\),&t=1,\\end\{cases\}whereinterp\(c\)\\operatorname\{interp\}\(c\)is the piecewise linear interpolation through the anchor points\{\(−2,2\),\(0,1\),\(1,0\)\}\\\{\(\-2,2\),\(0,1\),\(1,0\)\\\}\.
Thus, untreated subregions have outcome11, while treated subregions have an outcome that depends nonlinearly on the local contextcc\. The noisy subregional outcome is
yi,j=f\(ti,j,ci,j\)\+ϵi,j,ϵi,j∼𝒩\(0,0\.12\),y\_\{i,j\}=f\(t\_\{i,j\},c\_\{i,j\}\)\+\\epsilon\_\{i,j\},\\qquad\\epsilon\_\{i,j\}\\sim\\mathcal\{N\}\(0,0\.1^\{2\}\),and the observed regional outcome is
Y^i=1M∑j=1Myi,j\.\\widehat\{Y\}\_\{i\}=\\frac\{1\}\{M\}\\sum\_\{j=1\}^\{M\}y\_\{i,j\}\.The ground\-truth local causal effect, which we use only for evaluation, is
Ei,j=f\(ti,j,ci,j\)−f\(0,ci,j\)\.E\_\{i,j\}=f\(t\_\{i,j\},c\_\{i,j\}\)\-f\(0,c\_\{i,j\}\)\.
##### Training\.
We train a small neural networkf𝜽\(t,c\)f\_\{\\boldsymbol\{\\theta\}\}\(t,c\)with two inputs, treatment and context, and one scalar output\. The model has two hidden layers with3232hidden units each, ReLU activations, dropout with probability0\.050\.05, and is trained with AdamW using learning rate10−310^\{\-3\}and weight decay10−310^\{\-3\}for40004000epochs\. Importantly,Clamis trained only from the regional aggregate signal\. For each region, we average the model’s subregional predictions and minimize
ℒClam=1N∑i=1N\(1M∑j=1Mf𝜽\(ti,j,ci,j\)−Y^i\)2\.\\mathcal\{L\}\_\{\\textsc\{Clam\}\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\(\\frac\{1\}\{M\}\\sum\_\{j=1\}^\{M\}f\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\}\)\-\\widehat\{Y\}\_\{i\}\\right\)^\{2\}\.
##### Baseline\.
We compare against a uniform disaggregation baseline\. Since the true subregional outcomes are not observed, this baseline assigns the same regional outcome to every subregion,
y~i,j=Y^i\.\\widetilde\{y\}\_\{i,j\}=\\widehat\{Y\}\_\{i\}\.It then trains the same neural network architecture using an element\-wise loss,
ℒbase=1NM∑i=1N∑j=1M\(f𝜽\(ti,j,ci,j\)−y~i,j\)2\.\\mathcal\{L\}\_\{\\text\{base\}\}=\\frac\{1\}\{NM\}\\sum\_\{i=1\}^\{N\}\\sum\_\{j=1\}^\{M\}\\left\(f\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\}\)\-\\widetilde\{y\}\_\{i,j\}\\right\)^\{2\}\.This gives the baseline a dense subregional loss signal, but the signal is based on a uniformity assumption that removes the within\-region heterogeneity needed to recover local causal effects\.
##### Evaluation\.
We evaluate whether the learned model can recover the fine\-grained causal structure from aggregate observations \(Figure[5](https://arxiv.org/html/2608.08064#A4.F5)\)\. We report the training loss and LOCATE MAE \(cf\. Eq\. \([4](https://arxiv.org/html/2608.08064#S3.E4)\)\) over training \(Figure[5](https://arxiv.org/html/2608.08064#A4.F5)a\)\. Confidence intervals \(95%\) are bootstrapped and based on 20 runs\. As expected, both decrease forClam, while the baseline cannot learn the causal mechanism\. We also report the full LOCATE matrix, comparing the ground truth with the one inferred byClam, in Figure[5](https://arxiv.org/html/2608.08064#A4.F5)b\. The recovered causal effect is almost indistinguishable from the ground truth, while the baseline, as expected, cannot reconstruct the underlying disaggregation\. Lastly, we construct a counterfactual treatment matrix \(Figure[5](https://arxiv.org/html/2608.08064#A4.F5)c\) by inverting the observed treatment assignment,Tcf=1−TT^\{\\mathrm\{cf\}\}=1\-T, and compare the ground\-truth counterfactual outcomef\(Tcf,C\)f\(T^\{\\mathrm\{cf\}\},C\)with the neural\-network estimatef𝜽\(Tcf,C\)f\_\{\\boldsymbol\{\\theta\}\}\(T^\{\\mathrm\{cf\}\},C\)\. This tests whether the learned mechanism supports coherent counterfactual prediction beyond the observed treatment assignment\.
Figure 6:We evaluate the effect of varying the variance of regional context means and, for each fixed context mean, the inter\-regional variance\. Results are averaged over 5 random seeds for each parameter combination, and uncertainty bounds represent the corresponding variability\.
##### Ablation\.
To test when this recovery is possible, we repeat the experiment while varying two properties of the context matrix\. The first is the standard deviation of the regional context means, denotedσreg\\sigma\_\{\\text\{reg\}\}\. The second is the standard deviation of context values within each region, denotedσsubreg\\sigma\_\{\\text\{subreg\}\}\. For each pair\(σreg,σsubreg\)\(\\sigma\_\{\\text\{reg\}\},\\sigma\_\{\\text\{subreg\}\}\), we run the experiment over multiple random seeds and average the final LOCATE MAE\. This ablation shows that high variability within regions is associated with a smaller error\.
#### D\.1\.1Ablation of Exp\. 1: Political Campaigning
Our setting targets low\-variability aggregate regimes, where regional outcomes are nearly constant and the inverse problem is underdetermined: many fine\-grained causal explanations can yield indistinguishable regional observations\. The key idea is that high\-dimensional subregional covariates can restore information by exposing the shared causal mechanismf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)to diverse contextual compositions within each aggregate unit\. Exp\. 1 and the corresponding ablation support this mechanism: as subregional heterogeneity increases, recovery of the true causal effect improves, even when aggregate outcomes remain similarly uninformative\. By contrast, varying regional\-level heterogeneity has a weaker effect, indicating that the gains do not come primarily from more variable aggregate outcomes, but from richer within\-region contextual structure\. See Figure[6](https://arxiv.org/html/2608.08064#A4.F6)for details\. Thus, causal deabstraction becomes possible in otherwise weakly identified settings when subregional covariate variation, together with an invariant causal mechanism, provides enough compositional diversity to distinguish the underlying fine\-grained causal function\.
### D\.2Exp\. 2: Public School Funding vs Improvement in Outcomes
Figure 7:Performance of the training loss and MSE of the local\-effect estimation, alongside ground\-truth causal effects, estimated causal effects, and the corresponding estimation errors across subregions\.##### Context\.
We examine how public spending on schools affects educational outcomes when the exact locations of the expenditures are*unknown*\. Each administrative region contains multiple subregions, interpreted here as school districts, but only one district in a treated region receives a funding intervention\. At the regional level, we only observe an aggregated treatment indicator and the aggregated outcome, not the specific district that received the intervention\. The goal is to infer the high\-resolution intervention locations from aggregated data\.
The subregional context consists of high\-resolution demographic and economic data, represented here by a normally distributed*wealth score*\. We hypothesize that the effect of school funding depends on this wealth score through a global shift and scale\.
##### Setup\.
We assumeN=30×30=900N=30\\times 30=900regions, each withM=2×2=4M=2\\times 2=4subregions\.
We simulate:
- •A high\-resolution intervention matrixT∈ℝN×MT\\in\\mathbb\{R\}^\{N\\times M\}, whereti,jt\_\{i,j\}indicates whether subregionjjof regioniireceives the funding intervention\.
- •A regional intervention vectorT^∈ℝN\\widehat\{T\}\\in\\mathbb\{R\}^\{N\}, whereT^i=∑j=1Mti,j\\widehat\{T\}\_\{i\}=\\sum\_\{j=1\}^\{M\}t\_\{i,j\}is the absolute intervention intensity in regionii\.
- •A high\-resolution context matrixC∈ℝN×MC\\in\\mathbb\{R\}^\{N\\times M\}encoding the socioeconomic status of each subregion, where smallerci,jc\_\{i,j\}values represent lower socioeconomic status and largerci,jc\_\{i,j\}values represent higher socioeconomic status\.
- •A high\-resolution noise matrixϵ∈ℝN×M\\epsilon\\in\\mathbb\{R\}^\{N\\times M\}with i\.i\.d\. entriesϵi,j∼𝒩\(0,σ2\)\\epsilon\_\{i,j\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}\), whereσ2=0\.005\\sigma^\{2\}=0\.005\.
The intervention matrixTTis generated by:
1. 1\.Initializing all entries to zero\.
2. 2\.Randomly selecting50%50\\%of the regions for treatment\.
3. 3\.For each treated regionii, choosing exactly one subregionj∈\{1,…,M\}j\\in\\\{1,\\dots,M\\\}uniformly at random and assigningti,j=1t\_\{i,j\}=1\.
The context matrixCCis sampled i\.i\.d\. from a standard normal distribution:
ci,j∼𝒩\(0,1\)\.c\_\{i,j\}\\sim\\mathcal\{N\}\(0,1\)\.
The outcome variable is given by the ground\-truth functional relationship
yi,j=f𝜽\(ti,j,ci,j\)\+ϵi,j=\(θshift−ci,j\)⋅θscale⋅ti,j\+ϵi,j,y\_\{i,j\}=f\_\{\\boldsymbol\{\\theta\}\}\\\!\\left\(t\_\{i,j\},c\_\{i,j\}\\right\)\+\\epsilon\_\{i,j\}=\\left\(\\theta\_\{\\mathrm\{shift\}\}\-c\_\{i,j\}\\right\)\\cdot\\theta\_\{\\mathrm\{scale\}\}\\cdot t\_\{i,j\}\+\\epsilon\_\{i,j\},with ground\-truth parameters
θshift=4\.2,θscale=1\.23\.\\theta\_\{\\mathrm\{shift\}\}=4\.2,\\qquad\\theta\_\{\\mathrm\{scale\}\}=1\.23\.
The observed regional outcome is aggregated by summation:
Y^i=∑j=1Myi,j\.\\widehat\{Y\}\_\{i\}=\\sum\_\{j=1\}^\{M\}y\_\{i,j\}\.
As in Exp\. 1, we assume no hidden confounding, no time\-series effects, treatment assignment independent of context, and aggregation\-based regional observations\.
##### Training\.
We train a model to estimate both the latent subregional intervention assignments and the parameters of the causal effect function\.
The model has:
- •N×M=900×4N\\times M=900\\times 4latent treatment logits, one vector of lengthMMfor each region\.
- •Two scalar parameters\{θshift,θscale\}\\\{\\theta\_\{\\mathrm\{shift\}\},\\theta\_\{\\mathrm\{scale\}\}\\\}specifying the parametric form off𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)\.
The learned causal effect function is parametrized as
f𝜽\(ti,j,ci,j\)=\(θshift−ci,j\)⋅θscale⋅ti,j\.f\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\}\)=\\left\(\\theta\_\{\\mathrm\{shift\}\}\-c\_\{i,j\}\\right\)\\cdot\\theta\_\{\\mathrm\{scale\}\}\\cdot t\_\{i,j\}\.To ensure a positive scale parameter,θscale\\theta\_\{\\mathrm\{scale\}\}is represented through a softplus transformation\.
An auxiliary preprocessing function is applied to the latent treatment logits using the Gumbel\-Softmax relaxation\. This produces a soft row\-wise assignment over theMMsubregions\. The resulting assignment is processed so that untreated rows remain all\-zero and treated rows have approximately one active subregion\. During training, the Gumbel\-Softmax temperature is exponentially annealed from2\.02\.0to0\.10\.1\.
In the forward pass:
1. 1\.Sample a mini\-batch of region indices\.
2. 2\.Process the latent treatment logits using the Gumbel\-Softmax relaxation to obtain an estimate of the high\-resolution treatment matrixTTfor the mini\-batch\.
3. 3\.Compute predicted subregional outcomes: μi,j=f𝜽\(t^i,j,ci,j\)\.\\mu\_\{i,j\}=f\_\{\\boldsymbol\{\\theta\}\}\\\!\\left\(\\widehat\{t\}\_\{i,j\},c\_\{i,j\}\\right\)\.
4. 4\.Aggregate to regional predictions by summation: μi=∑j=1Mμi,j\.\\mu\_\{i\}=\\sum\_\{j=1\}^\{M\}\\mu\_\{i,j\}\.
5. 5\.Compare the regional prediction with the observed regional outcomeY^i\\widehat\{Y\}\_\{i\}using the MSE loss\.
We optimize all parameters jointly using AdamW with learning rate10−410^\{\-4\}\. Training is performed for200,000200\{,\}000epochs with a fixed random seed\.
##### Neural Network\.
Additionally, we train a neural\-network variant following the same treatment\-assignment procedure\. In this version, the causal effect functionf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)is implemented as a feed\-forward neural network taking\(ti,j,ci,j\)\(t\_\{i,j\},c\_\{i,j\}\)as input\.
The network has three hidden layers with hidden dimension1616andtanh\\tanhactivations:
\(ti,j,ci,j\)↦MLP𝜽\(ti,j,ci,j\)\.\(t\_\{i,j\},c\_\{i,j\}\)\\mapsto\\mathrm\{MLP\}\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\}\)\.To stabilize the local\-effect interpretation, the model uses the centered form
f𝜽\(ti,j,ci,j\)=MLP𝜽\(ti,j,ci,j\)−MLP𝜽\(0,ci,j\),f\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\}\)=\\mathrm\{MLP\}\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\}\)\-\\mathrm\{MLP\}\_\{\\boldsymbol\{\\theta\}\}\(0,c\_\{i,j\}\),which enforcesf𝜽\(0,ci,j\)≈0f\_\{\\boldsymbol\{\\theta\}\}\(0,c\_\{i,j\}\)\\approx 0\. The latent treatment indicators are processed in the same way as in the parametric model, using Gumbel\-Softmax with temperature annealing from2\.02\.0to0\.10\.1\.
The neural\-network model is also trained for200,000200\{,\}000epochs using AdamW with learning rate10−410^\{\-4\}and mini\-batches containing20%20\\%of the regions\.
##### Results\.
Figure[7](https://arxiv.org/html/2608.08064#A4.F7)a shows the evolution of the mini\-batch training loss and the quality of the treatment\-location reconstruction \(measured as MAE\)\. Both the neural network and the parametric function recover the locations reasonably well\. We however find that this depends on the noise level and on proper temperature annealing\. The final reconstructed locations are shown in Figure[7](https://arxiv.org/html/2608.08064#A4.F7)b, where incorrect predictions are shown in purple\.
### D\.3Exp\. 3: Spatiotemporal Effects of Heat Waves on School Performance
Figure 8:Ground\-truth vegetation index, the corresponding estimated latent effect\-modifier fields, and joint plots of training loss and vegetation MSE\. Results are shown for two settings: linear scaling \(left\) and MLP learning \(right\)\.##### Context\.
In this experiment, we examine the impact of extreme heat events \(“heat waves”\) on educational outcomes across both space and time\. We model a spatial grid ofN=100N=100regions, each containingM=9M=9subregions, observed overW=48W=48consecutive months\. The intervention represents the occurrence of a heat wave in a given subregion and month\. Heat waves have a detrimental effect on school performance, but the strength of this effect depends on the educational background of parents in the subregion and on an unobserved environmental factor, vegetation coverage, which mitigates the impact of heat\.
The observed context for each subregion is a categorical indicator of parents’ education level, taking one of three values \(*low*= 1,*medium*= 2,*high*= 3\), which evolves slowly over time to reflect gradual demographic changes\. The unobserved context is a fixed vegetation index between0and11for each subregion, representing the proportion of green cover and its protective effect against heat\. The intervention process is binary at the subregional level, with most regions remaining unaffected in a given month, but some experiencing heat waves simultaneously in all their subregions, corresponding to large\-scale weather patterns\.
The outcome, measured as a monthly change in school performance for each subregion, is driven by the interaction between the intervention, parents’ education level, and the vegetation coverage\. Subregions with lower parental education are more strongly affected, while vegetation dampens the negative effect\. The challenge in this experiment is to recover both the mapping from context, intervention, and unobserved effect modifier to outcomes, and the vegetation values themselves, given only aggregated regional\-level outcomes over time\.
##### Setup\.
We modelN=100N=100regions, each withM=9M=9subregions, observed overW=48W=48discrete time points \(months\)\.
For each monthw∈\{1,…,W\}w\\in\\\{1,\\dots,W\\\}, we define:
- •A high‐resolution binary treatment matrixT\(w\)∈\{0,1\}N×MT^\{\(w\)\}\\in\\\{0,1\\\}^\{N\\times M\}, whereti,j\(w\)=1t\_\{i,j\}^\{\(w\)\}=1indicates that subregionjjof regioniiis experiencing a heat wave in monthww\. In our data generation, each row ofT\(w\)T^\{\(w\)\}is either all zeros \(probability0\.70\.7\) or all ones \(probability0\.30\.3\), corresponding to large‐scale heat events affecting entire regions\.
- •An observed context matrixC\(w\)∈\{1,2,3\}N×MC^\{\(w\)\}\\in\\\{1,2,3\\\}^\{N\\times M\}encoding the categorical education level of parents in each subregion\. Atw=1w=1, these values are sampled uniformly\. Forw\>1w\>1,C\(w\)C^\{\(w\)\}is obtained fromC\(w−1\)C^\{\(w\-1\)\}by flipping each entry to a random category with probability0\.050\.05, capturing slow demographic change\.
- •An unobserved covariate context matrixU∈\[0,1\]N×MU\\in\[0,1\]^\{N\\times M\}containing fixed vegetation index values for each subregion, drawn uniformly from\[0,1\]\[0,1\]and constant across all months\.
- •A noise matrixϵ\(w\)∈ℝN×M\\epsilon^\{\(w\)\}\\in\\mathbb\{R\}^\{N\\times M\}with i\.i\.d\. entriesϵi,j\(w\)∼𝒩\(0,σ2\)\\epsilon\_\{i,j\}^\{\(w\)\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}\), withσ2=0\.02\\sigma^\{2\}=0\.02\.
The subregional outcome is generated as:
yi,j\(w\)=f𝜽\(ti,j\(w\),ci,j\(w\),ui,j\)\+ϵi,j\(w\),y\_\{i,j\}^\{\(w\)\}=f\_\{\\boldsymbol\{\\theta\}\}\\\!\\left\(t\_\{i,j\}^\{\(w\)\},c\_\{i,j\}^\{\(w\)\},u\_\{i,j\}\\right\)\+\\epsilon\_\{i,j\}^\{\(w\)\},whereci,j\(w\)c\_\{i,j\}^\{\(w\)\}is the parents’ education category,ui,ju\_\{i,j\}is the vegetation index, and:
f𝜽\(ti,j\(w\),ci,j\(w\),ui,j\)=\{0ifti,j\(w\)=0,\(10⋅𝟙\[ci,j\(w\)=1\]\+5⋅𝟙\[ci,j\(w\)=2\]\+1\[ci,j\(w\)=3\]\)⋅\(1−ui,j\)ifti,j\(w\)=1\.f\_\{\\boldsymbol\{\\theta\}\}\\\!\\left\(t\_\{i,j\}^\{\(w\)\},c\_\{i,j\}^\{\(w\)\},u\_\{i,j\}\\right\)=\\begin\{cases\}0&\\text\{if \}t\_\{i,j\}^\{\(w\)\}=0,\\\\\[6\.0pt\] \\begin\{aligned\} &\\bigl\(10\\cdot\\mathds\{1\}\[c\_\{i,j\}^\{\(w\)\}=1\]\+5\\cdot\\mathds\{1\}\[c\_\{i,j\}^\{\(w\)\}=2\]\\\\ &\\quad\+\\;\\mathds\{1\}\[c\_\{i,j\}^\{\(w\)\}=3\]\\bigr\)\\cdot\(1\-u\_\{i,j\}\)\\end\{aligned\}&\\text\{if \}t\_\{i,j\}^\{\(w\)\}=1\.\\end\{cases\}
This function specifies that the subregional outcome depends on whether a heat wave occurs and, if so, how vulnerable the subregion is based on the education level of parents and the vegetation coverage\. Ifti,j\(w\)=0t\_\{i,j\}^\{\(w\)\}=0, the effect is zero\. Ifti,j\(w\)=1t\_\{i,j\}^\{\(w\)\}=1, the impact is determined by the categorical education level with coefficients\(10,5,1\)\(10,5,1\)corresponding to low, medium, and high parental education, respectively, and is scaled by\(1−ui,j\)\(1\-u\_\{i,j\}\), which reduces the effect in proportion to the amount of vegetation present\. The indicator𝟙\[⋅\]\\mathds\{1\}\[\\cdot\]selects the appropriate coefficient for each subregion\.
The regional outcome at monthwwis obtained via mean aggregation\. This setup induces a spatiotemporal causal problem with heterogeneous treatment effects driven by both observed and unobserved context, where the latter must be recovered from aggregated outcomes across multiple time points\.
##### Training\.
We train a model to jointly estimate:
1. 1\.The mappingf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)from observed context, intervention, and vegetation index to subregional outcomes,
2. 2\.The static vegetation index valuesU∈\[0,1\]N×M\{U\}\\in\[0,1\]^\{N\\times M\}for all subregions\.
The functionf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)is implemented as a 4 layer multilayer perceptron \(MLP\) with hidden dimension1616and a dropout rate of0\.10\.1after the second hidden layer\. Its inputs for each subregion are:
zi,j\(w\)=\(𝟙\[ci,j\(w\)=1\],1\[ci,j\(w\)=2\],1\[ci,j\(w\)=3\],ti,j\(w\),ui,j\),z\_\{i,j\}^\{\(w\)\}=\\big\(\\mathds\{1\}\[c\_\{i,j\}^\{\(w\)\}=1\],\\ \\mathds\{1\}\[c\_\{i,j\}^\{\(w\)\}=2\],\\ \\mathds\{1\}\[c\_\{i,j\}^\{\(w\)\}=3\],\\ t\_\{i,j\}^\{\(w\)\},\\ u\_\{i,j\}\\big\),whereui,ju\_\{i,j\}is the vegetation index\. The output is the predicted outcomeyi,j\(w\)y\_\{i,j\}^\{\(w\)\}for that subregion and month\.
In the forward pass for a given monthww:
1. 1\.The model takes as inputT\(w\)T^\{\(w\)\},C\(w\)C^\{\(w\)\}, andU\{U\},
2. 2\.The MLP evaluatesf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)at each\(i,j\)\(i,j\)to produce predicted subregional outcomesyi,j\(w\)y\_\{i,j\}^\{\(w\)\},
3. 3\.These are aggregated to the regional level via mean aggregation\.
4. 4\.The primary loss is the MSE between predicted and observed regional outcomes\.
Training proceeds over all months in random order at each epoch\. The total parameter set consists of the MLP weights𝜽\\boldsymbol\{\\theta\}and the vegetation matrixU\{U\}\. We optimize using Adam with a learning rate of0\.0010\.001for10,00010\{,\}000epochs, with fixed random seeds for reproducibility\. Loss curves forℒregion\\mathcal\{L\}\_\{\\mathrm\{region\}\}and the MSE ofℒveg\\mathcal\{L\}\_\{\\mathrm\{veg\}\}are recorded throughout training\.
##### Results\.
1. 1\.Figure[8](https://arxiv.org/html/2608.08064#A4.F8)a shows the loss curves and vegetation MSE under the linear scaling parameterization\. Both training loss and MSE decrease rapidly, and the final error is very small, indicating near‐perfect recovery of the unobserved covariate\.
2. 2\.Figure[8](https://arxiv.org/html/2608.08064#A4.F8)b compares the ground‐truth vegetation field to the estimated vegetation under linear scaling\. The two maps are almost indistinguishable, confirming that the restricted functional form matches the data\-generating process and allows highly accurate recovery\.
3. 3\.Figures[8](https://arxiv.org/html/2608.08064#A4.F8)c and d present the same results for the non‐linear MLP parameterization\. While the model still recovers meaningful structure in the vegetation, the MSE remains orders of magnitude higher than in the linear case, and deviations from the ground truth are clearly visible in the estimated field\. This is notable since the true vegetation values are not necessarily identifiable; the MLP could instead store a transformed version and map them back during the forward process\.
4. 4\.Overall, these results demonstrate that the model can reconstruct regional outcomes while also uncovering hidden drivers of heterogeneity\. The comparison between linear and non‐linear parameterizations highlights the trade‐off: structural restrictions yield more accurate recovery when they match the true mechanism, while flexible approximators such as MLPs still learn the vegetation but with reduced precision\.
## Appendix EReal\-World Case Study Details: Effect of Heat on Gun Violence
Figure 9:Comparison of learned spatial intensities for week 51\. Both models are rendered on a shared color scale across theN⋅M×N⋅M\\sqrt\{N\\cdot M\}\\times\\sqrt\{N\\cdot M\}grid, alongside the observed events for that week\.This section provides the full details of the real\-world study summarized in Section[6\.2](https://arxiv.org/html/2608.08064#S6.SS2)\. The notation follows Section[2](https://arxiv.org/html/2608.08064#S2): regions are indexed byi∈\{1,…,N\}i\\in\\\{1,\\dots,N\\\}, subregions or cells byj∈\{1,…,M\}j\\in\\\{1,\\dots,M\\\}, and weeks byw∈\{1,…,W\}w\\in\\\{1,\\dots,W\\\}\. Temperature is used as a continuous, time\-varying treatment\-like exposure; urbanization is the HR static contextual covariate; and gun\-incident counts are the observed LR outcomes\.
##### Problem Setting\.
The target variable is observed at a coarse spatial resolution, while covariates are available at a finer resolution\. The study area is discretized into a fine grid ofN⋅M×N⋅M\\sqrt\{N\\cdot M\}\\times\\sqrt\{N\\cdot M\}cells, grouped intoN×N\\sqrt\{N\}\\times\\sqrt\{N\}coarse regions\. WithN=100N=100andM=100M=100, this gives a100×100100\\times 100cell grid grouped into a10×1010\\times 10region grid\. For each weekwwand regionii, we observe a nonnegative count outcome
Yi,wLR∈ℤ≥0,Y^\{\\mathrm\{LR\}\}\_\{i,w\}\\in\\mathbb\{Z\}\_\{\\geq 0\},\(7\)representing the number of gun\-incident events occurring in regioniiduring weekww\.
##### Data Representation and Covariates\.
Letci,j,wc\_\{i,j,w\}denote the observed urbanization measurement for cell\(i,j\)\(i,j\)and weekww\. Since urbanization is treated as time\-invariant in this study, we use its temporal mean,
ci,j=1W∑w=1Wci,j,w\.c\_\{i,j\}=\\frac\{1\}\{W\}\\sum\_\{w=1\}^\{W\}c\_\{i,j,w\}\.\(8\)Temperature is modeled at the region level\. Letti,j,wHRt^\{\\mathrm\{HR\}\}\_\{i,j,w\}denote the cell\-level temperature for cell\(i,j\)\(i,j\)and weekww\. We compute the region\-average temperature exposure as
ti,wLR=1M∑j=1Mti,j,wHR,t^\{\\mathrm\{LR\}\}\_\{i,w\}=\\frac\{1\}\{M\}\\sum\_\{j=1\}^\{M\}t^\{\\mathrm\{HR\}\}\_\{i,j,w\},\(9\)and standardize it using the global mean and standard deviation across all region–week pairs:
t~i,w=ti,wLR−μtσt\+ε,\\widetilde\{t\}\_\{i,w\}=\\frac\{t^\{\\mathrm\{LR\}\}\_\{i,w\}\-\\mu\_\{t\}\}\{\\sigma\_\{t\}\+\\varepsilon\},\(10\)whereε\>0\\varepsilon\>0is a small constant for numerical stability\. Missing temperature values are imputed using the global median temperature, and missing outcome counts are set to zero\.
##### Disaggregation Model\.
We model the latent HR intensity, or rate,λi,j,w\>0\\lambda\_\{i,j,w\}\>0as a function of the standardized regional temperature exposure and the HR urbanization covariate:
λi,j,w=f𝜽\(t~i,w,ci,j\),\\lambda\_\{i,j,w\}=f\_\{\\boldsymbol\{\\theta\}\}\\\!\\left\(\\widetilde\{t\}\_\{i,w\},\\,c\_\{i,j\}\\right\),\(11\)wheref𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)is a feed\-forward neural network with parameters𝜽\\boldsymbol\{\\theta\}\. To enforce nonnegativity, the network output is passed through a softplus transformation\.
The model\-implied LR expected count is obtained by summing the HR rates over all cells in regionii:
μi,w=∑j=1Mλi,j,w\.\\mu\_\{i,w\}=\\sum\_\{j=1\}^\{M\}\\lambda\_\{i,j,w\}\.\(12\)Thus,μi,w\\mu\_\{i,w\}denotes the predicted regional count intensity, whileYi,wLRY^\{\\mathrm\{LR\}\}\_\{i,w\}denotes the observed regional count\. This additive aggregation is consistent with the SCM in Section[3\.1](https://arxiv.org/html/2608.08064#S3.SS1)\.
##### Training Objective\.
We assume a Poisson observation model at the region–week level,
Yi,wLR∼Poisson\(μi,w\),Y^\{\\mathrm\{LR\}\}\_\{i,w\}\\sim\\mathrm\{Poisson\}\\\!\\left\(\\mu\_\{i,w\}\\right\),\(13\)and trainf𝜽f\_\{\\boldsymbol\{\\theta\}\}by minimizing the negative Poisson log\-likelihood, up to additive constants:
ℒ\(𝜽\)=1\|𝒲train\|⋅N∑w∈𝒲train∑i=1N\(μi,w−Yi,wLRlog\(μi,w\+ε\)\),\\mathcal\{L\}\(\\boldsymbol\{\\theta\}\)=\\frac\{1\}\{\|\\mathcal\{W\}\_\{\\mathrm\{train\}\}\|\\cdot N\}\\sum\_\{w\\in\\mathcal\{W\}\_\{\\mathrm\{train\}\}\}\\sum\_\{i=1\}^\{N\}\\left\(\\mu\_\{i,w\}\-Y^\{\\mathrm\{LR\}\}\_\{i,w\}\\log\(\\mu\_\{i,w\}\+\\varepsilon\)\\right\),\(14\)where𝒲train\\mathcal\{W\}\_\{\\mathrm\{train\}\}is the set of training weeks andε\>0\\varepsilon\>0is a small constant for numerical stability\.
##### Spatial Intensity Comparison\.
Figure[9](https://arxiv.org/html/2608.08064#A5.F9)visualizes observed incidents for a representative week, week 51, alongside spatial intensity estimates learned under different supervision regimes\. When trained only on region\-level counts,Clamproduces spatially structured intensity maps that concentrate mass around incident clusters but exhibit noticeably higher peak intensities and broader spatial support compared to the cell\-supervised estimate\. This overestimation reflects a fundamental ambiguity induced by aggregation: without access to fine\-grained outcomes, the model must allocate regional mass across plausible subregional locations, leading to inflated local rates in areas consistent with the observed context and covariates\. The response\-surface comparison in the main paper, Figure[4](https://arxiv.org/html/2608.08064#S6.F4), shows the same pattern in covariate space rather than physical space\.
##### Learning\.
The training procedure optimizes the function parameters𝜽\\boldsymbol\{\\theta\}\(or any other missing components of the SCM\)\. The structural functionf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)can be represented by a neural network or any other parameterized function\. Since the fine\-grained outcomesyi,jy\_\{i,j\}are unobserved, training relies solely on their aggregated counterparts\.
We optimize with the goal that the inferred aggregated predictions \(μi\\mu\_\{i\}\) match the observed aggregate:ℒ𝜽,ϕ=∑i=1N\(Y^i−μi\)2,\\mathcal\{L\}\_\{\\boldsymbol\{\\theta\},\\boldsymbol\{\\phi\}\}=\\sum\_\{i=1\}^\{N\}\\bigl\(\\widehat\{Y\}\_\{i\}\-\\mu\_\{i\}\\bigr\)^\{2\},whereY^i\\widehat\{Y\}\_\{i\}denotes the observed aggregate of regionii\. If the noise terms are not i\.i\.d\. and normally distributed, they can also be explicitly learned within the optimization loop\.
## Appendix FAdditional Experiments
Exp\. 4 considers driving bans with an unknown aggregation function between subregional and regional outcomes\. Exp\. 5 introduces treatment–context confounding, where the probability of subregional treatment depends on contextual variables, affecting both allocation and outcomes
### F\.1Exp\. 4: Unknown Aggregation Functions
Figure 10:Results Overview\.Top:Input of LR intervention locations \(red\), LR outcomes \(blue\), and HR context \(green\) for mean aggregation and max aggregation cases\.Bottom:We observe the estimated local effects and the performance of our model over multiple epochs for both cases\.##### Context\.
We study the causal effect of driving bans on air quality under the assumption that the reported region\-level air quality values are an unknown aggregation of finer\-scale \(subregional\) measurements\. Specifically, we model the aggregation as a learned combination of the subregional mean and maximum values, with the combination weights parameterized via a logit transform\.
A realistic example is a city\-wide driving ban where air quality is officially reported as a single regional index \(e\.g\., PM2\.5concentration\), while the underlying measurement network collects data from multiple monitoring stations\. Depending on reporting policies, the published index might resemble an average across stations, a worst\-case station reading, or something in between\.
The high\-resolution context for each subregion is a single continuous variable: a vegetation index, capturing the amount of green coverage in the area\. Vegetation can mitigate pollution and thus may modulate the effect of a driving ban\. We hypothesize that the causal effect of driving bans varies with vegetation and that the unknown aggregation mechanism must be learned alongside the effect parameters to estimate subregional impacts correctly\.
##### Setup\.
We assumeN=100N=100regions, each withM=100M=100subregions\.
We are given:
- •A binary treatment vectorT∈\{0,1\}NT\\in\\\{0,1\\\}^\{N\}, whereTi=1T\_\{i\}=1indicates that a driving ban was implemented in regionii\(and thus all its subregions\)\.
- •A high\-resolution context matrixC∈\[0,1\]N×MC\\in\[0,1\]^\{N\\times M\}encoding the vegetation index of each subregion, whereci,j=0c\_\{i,j\}=0represents minimal vegetation cover andci,j=1c\_\{i,j\}=1represents dense vegetation\.
- •A high\-resolution noise matrixϵ∈ℝN×M\\epsilon\\in\\mathbb\{R\}^\{N\\times M\}with i\.i\.d\. entriesϵi,j∼𝒩\(0,σ2\)\\epsilon\_\{i,j\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}\), whereσ2=0\.02\\sigma^\{2\}=0\.02\.
The treatment vectorTTis generated using a Bernoulli distribution with probability0\.50\.5, assigning each region to either the intervention or control group at random\. The vegetation index values inCCare drawn i\.i\.d\. fromUniform\(0,1\)\\mathrm\{Uniform\}\(0,1\)\.
The subregional outcome variable is given by the ground\-truth functional relationship:
yi,j=f𝜽\(ti,j,ci,j\)\+ϵi,j,y\_\{i,j\}=f\_\{\\boldsymbol\{\\theta\}\}\\\!\\left\(t\_\{i,j\},c\_\{i,j\}\\right\)\+\\epsilon\_\{i,j\},where:
- •ti,j=Tit\_\{i,j\}=T\_\{i\}is the treatment status of subregionjjin regionii\(inherited from the region\),
- •ci,jc\_\{i,j\}is the vegetation index,
- •ϵi,j\\epsilon\_\{i,j\}is Gaussian noise\.
The causal effect function is parameterized as:
f𝜽\(ti,j,ci,j\)=\{0\.0ifti,j=0,θbase\+θveg⋅ci,jifti,j=1,f\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\}\)=\\begin\{cases\}0\.0&\\text\{if \}t\_\{i,j\}=0,\\\\ \\theta\_\{\\mathrm\{base\}\}\+\\theta\_\{\\mathrm\{veg\}\}\\cdot c\_\{i,j\}&\\text\{if \}t\_\{i,j\}=1,\\end\{cases\}whereθbase\\theta\_\{\\mathrm\{base\}\}captures the baseline effect of a driving ban andθveg\\theta\_\{\\mathrm\{veg\}\}captures how this effect changes with vegetation cover\.
Unlike Exp\. 1, the regional outcomexix\_\{i\}is not a simple mean over subregions\. Instead, we model the aggregation function as:
yi=∑j=1Mpi,j\(τ\)⋅yi,j,y\_\{i\}=\\sum\_\{j=1\}^\{M\}p\_\{i,j\}\(\\tau\)\\cdot y\_\{i,j\},where:
pi,j\(τ\)=exp\(yi,j/τ\)∑k=1Mexp\(yi,k/τ\)p\_\{i,j\}\(\\tau\)=\\frac\{\\exp\\\!\\left\(y\_\{i,j\}/\\tau\\right\)\}\{\\sum\_\{k=1\}^\{M\}\\exp\\\!\\left\(y\_\{i,k\}/\\tau\\right\)\}is the softmax weight assigned to subregionjjwith temperature parameterτ\>0\\tau\>0\. Whenτ→∞\\tau\\to\\infty,pi,jp\_\{i,j\}approaches a uniform distribution \(mean aggregation\); whenτ→0\\tau\\to 0,pi,jp\_\{i,j\}concentrates on the subregion with the highest outcome \(max aggregation\)\.
##### Training\.
We train a model to estimate:
1. 1\.The parameters𝜽=\{θbase,θveg\}\\boldsymbol\{\\theta\}=\\\{\\theta\_\{\\mathrm\{base\}\},\\theta\_\{\\mathrm\{veg\}\}\\\}of the causal effect functionf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\),
2. 2\.The aggregation temperature parameterτ\>0\\tau\>0governing the softmax\-based aggregation from subregional to regional outcomes\.
The learned function takes the form:
f𝜽\(ti,j,ci,j\)=\{0\.0ifti,j=0,θbase\+θveg⋅ci,jifti,j=1\.f\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\}\)=\\begin\{cases\}0\.0&\\text\{if \}t\_\{i,j\}=0,\\\\ \\theta\_\{\\mathrm\{base\}\}\+\\theta\_\{\\mathrm\{veg\}\}\\cdot c\_\{i,j\}&\\text\{if \}t\_\{i,j\}=1\.\\end\{cases\}
In the forward pass:
1. 1\.Compute predicted subregional outcomes:yi,j=f𝜽\(ti,j,ci,j\),y\_\{i,j\}=f\_\{\\boldsymbol\{\\theta\}\}\\\!\\left\(t\_\{i,j\},c\_\{i,j\}\\right\),
2. 2\.Compute subregion\-level aggregation weights via the softmax:pi,j\(τ\)=exp\(yi,j/τ\)∑k=1Mexp\(yi,k/τ\),p\_\{i,j\}\(\\tau\)=\\frac\{\\exp\\\!\\left\(y\_\{i,j\}/\\tau\\right\)\}\{\\sum\_\{k=1\}^\{M\}\\exp\\\!\\left\(y\_\{i,k\}/\\tau\\right\)\},
3. 3\.Compute predicted regional outcomes as the expectation underpi,j\(τ\)p\_\{i,j\}\(\\tau\):yi=∑j=1Mpi,j\(τ\)⋅yi,j,y\_\{i\}=\\sum\_\{j=1\}^\{M\}p\_\{i,j\}\(\\tau\)\\cdot y\_\{i,j\},
4. 4\.Compute the loss as the mean squared error with the observedYiY\_\{i\}\.
All parameters\{θbase,θveg,τ\}\\\{\\theta\_\{\\mathrm\{base\}\},\\theta\_\{\\mathrm\{veg\}\},\\tau\\\}are optimized jointly using Adam with a learning rate of0\.0010\.001for10001000epochs\. We initializeτ\\tauto a moderate value \(e\.g\.,τ=1\.0\\tau=1\.0\) and enforceτ\>0\\tau\>0during training by optimizing its logarithm\.
##### Results\.
1. 1\.Visualizations of the aggregated causal effects show that mean aggregation preserves more variation across regions, whereas max aggregation suppresses smaller effects and produces sharper contrasts\.
2. 2\.Although the subregional causal effect loss is never used directly for training, it decreases steadily over epochs, indicating that the model recovers fine scale effects from only regional supervision\. Training loss is substantially higher under max aggregation, as weaker subregional signals are obscured\.
3. 3\.Comparing estimated and ground truth local effects confirms this bias: areas with small true effects tend to be overestimated when max aggregation dominates\.
4. 4\.The aggregation function, parameterized by the temperatureτ\\tau, is recovered with high accuracy\. In the mean aggregation case,τ\\tauis estimated within0\.010\.01of the ground truth, while in the max aggregation case, the estimate is slightly underestimated \(5\.65\.6\), consistent with the plateauing behavior of the softmax function at low temperatures\.
5. 5\.Overall, the experiment demonstrates that our approach can successfully recover both the aggregation mechanism \(within its parametric form\) and the underlying causal effects, despite only observing region\-level outcomes\.
### F\.2Exp\. 5: Covariate\-Based Confounding of Treatment Subregion
Figure 11:Results Overview\.Top:Input of LR intervention locations \(red\), LR outcomes \(blue\), and HR context \(green\)\.Bottom:Ground truth treatment locations, estimation of treatment locations based on modeling of location covariate\-based confounding, naïve estimation of treatment locations with all free parameters\.##### Context\.
In this experiment, we study how confounding between treatment allocation and contextual variables affects causal effect estimation\. We take the example of school districts receiving additional public funding and the corresponding changes in educational outcomes\.
In reality, funding allocation is rarely random: wealthier or poorer districts may have systematically different chances of receiving funding due to political priorities, lobbying power, or policy constraints\. This creates a dependency between the probability of treatment and district level context variables, violating the standard assumption of independent treatment assignment\.
We model this confounding explicitly by making the probability of treatment allocation a deterministic function of the contextual variable, passed through a logistic transform to obtain valid probabilities for each region\. The treatment effect itself is also a function of the same contextual variable, representing heterogeneous impacts across different socioeconomic conditions\.
##### Setup\.
We assumeN=100N=100regions, each withM=100M=100subregions \(school districts\)\.
We are given:
- •A high\-resolution treatment matrixT∈\{0,1\}N×MT\\in\\\{0,1\\\}^\{N\\times M\}with*exactly one*treated subregion per region, i\.e\.,∑j=1Mti,j=1\\sum\_\{j=1\}^\{M\}t\_\{i,j\}=1for allii\.
- •A context matrixC∈\[0,1\]N×MC\\in\[0,1\]^\{N\\times M\}encoding a socioeconomic/need index for each subregion, where a higherci,jc\_\{i,j\}indicates greater need\.
- •A noise matrixϵ∈ℝN×M\\epsilon\\in\\mathbb\{R\}^\{N\\times M\}with i\.i\.d\. entriesϵi,j∼𝒩\(0,σ2\)\\epsilon\_\{i,j\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}\)\(defaultσ2=0\.02\\sigma^\{2\}=0\.02\)\.
##### Confounded Treatment Allocation\.
For each regionii, we form logits from the subregional context,
ℓi,j=α0\+α1ci,j,\\ell\_\{i,j\}=\\alpha\_\{0\}\+\\alpha\_\{1\}\\,c\_\{i,j\},and convert them to a categorical distribution over subregions via a row\-wise softmax,
pi,j=exp\(ℓi,j\)∑k=1Mexp\(ℓi,k\)\.p\_\{i,j\}\\;=\\;\\frac\{\\exp\(\\ell\_\{i,j\}\)\}\{\\sum\_\{k=1\}^\{M\}\\exp\(\\ell\_\{i,k\}\)\}\.We then draw exactly one treated subregionj⋆∼Categorical\(pi,1:M\)j^\{\\star\}\\sim\\mathrm\{Categorical\}\(p\_\{i,1:M\}\)and set
ti,j=𝟙\{j=j⋆\},∑j=1Mti,j=1\.t\_\{i,j\}=\\mathds\{1\}\\\{j=j^\{\\star\}\\\},\\qquad\\sum\_\{j=1\}^\{M\}t\_\{i,j\}=1\.The parameters\(α0,α1\)\(\\alpha\_\{0\},\\alpha\_\{1\}\)control the strength and direction of confounding between context and treatment assignment \(defaultα0=0\\alpha\_\{0\}=0,α1\>0\\alpha\_\{1\}\>0so higher\-need subregions are more likely to be treated\)\.
##### Outcome Model With Context\-Dependent Effects\.
Subregional outcomes follow the ground\-truth functional relationship
yi,j=f𝜽\(ti,j,ci,j\)\+ϵi,j,y\_\{i,j\}\\;=\\;f\_\{\\boldsymbol\{\\theta\}\}\\\!\\left\(t\_\{i,j\},c\_\{i,j\}\\right\)\+\\epsilon\_\{i,j\},with the causal effect function
f𝜽\(ti,j,ci,j\)=\{0ifti,j=0,θbase\+θsocci,jifti,j=1,f\_\{\\boldsymbol\{\\theta\}\}\(t\_\{i,j\},c\_\{i,j\}\)\\;=\\;\\begin\{cases\}0&\\text\{if \}t\_\{i,j\}=0,\\\\\[2\.0pt\] \\theta\_\{\\mathrm\{base\}\}\+\\theta\_\{\\mathrm\{soc\}\}\\,c\_\{i,j\}&\\text\{if \}t\_\{i,j\}=1,\\end\{cases\}where𝜽=\{θbase,θsoc\}\\boldsymbol\{\\theta\}=\\\{\\theta\_\{\\mathrm\{base\}\},\\,\\theta\_\{\\mathrm\{soc\}\}\\\}captures the baseline funding effect and its modulation by socioeconomic need\.
##### Regional Aggregation\.
Region\-level outcomes are the sum over subregions \(as in Exp\. 1\)\.
This construction induces confounding because \(i\) treatment allocation depends onci,jc\_\{i,j\}through the softmax logits, and \(ii\) treatment effects also vary withci,jc\_\{i,j\}viaf𝜽\(⋅\)f\_\{\\boldsymbol\{\\theta\}\}\(\\cdot\)\.
##### Training\.
We compare two training regimes for estimating the causal effect parameters𝜽=\{θbase,θsoc\}\\boldsymbol\{\\theta\}=\\\{\\theta\_\{\\mathrm\{base\}\},\\theta\_\{\\mathrm\{soc\}\}\\\}and the high resolution treatment assignmentsT∈ℝN×M\{T\}\\in\\mathbb\{R\}^\{N\\times M\}in the presence of treatment context confounding\.
##### Regime 1: Naïve Training Under Independence Assumption\.
This setup directly applies the training procedure of Exp\. 2 \(Section[D](https://arxiv.org/html/2608.08064#A4)\) to the confounded data, it:
1. 1\.Treats the treatment allocation as if independent of the context,
2. 2\.LearnsT\{T\}as free parameters, constrained to be one hot per row via a differentiable argmax with temperature,
3. 3\.Learns𝜽\\boldsymbol\{\\theta\}by minimizing the MSE between observed and predicted regional outcomes\.
This ignores the fact that the treatment assignment mechanism is context dependent\.
##### Regime 2: Confounding Aware Training\.
Here, we explicitly model the context\-dependent treatment allocation mechanism\. For each regionii, we predict logits
ℓi,j=β0\+β1ci,j,\\ell\_\{i,j\}=\\beta\_\{0\}\+\\beta\_\{1\}\\,c\_\{i,j\},where\(β0,β1\)\(\\beta\_\{0\},\\beta\_\{1\}\)are learned parameters, and convert them to a categorical distribution via a row\-wise softmax with a learnable temperatureτconf\>0\\tau\_\{\\mathrm\{conf\}\}\>0:
pi,j\(τconf\)=exp\(ℓi,j/τconf\)∑k=1Mexp\(ℓi,k/τconf\)\.p\_\{i,j\}\(\\tau\_\{\\mathrm\{conf\}\}\)=\\frac\{\\exp\\\!\\left\(\\ell\_\{i,j\}/\\tau\_\{\\mathrm\{conf\}\}\\right\)\}\{\\sum\_\{k=1\}^\{M\}\\exp\\\!\\left\(\\ell\_\{i,k\}/\\tau\_\{\\mathrm\{conf\}\}\\right\)\}\.We then usepi,j\(τconf\)p\_\{i,j\}\(\\tau\_\{\\mathrm\{conf\}\}\)as the \(differentiable\) treatment assignment in the forward pass\.
In both regimes, predicted subregional outcomes are given by:
yi,j=f𝜽\(ti,j,ci,j\)=\{0ifti,j=0,θbase\+θsoc⋅ci,jifti,j=1,y\_\{i,j\}=f\_\{\\boldsymbol\{\\theta\}\}\\\!\\left\(t\_\{i,j\},c\_\{i,j\}\\right\)=\\begin\{cases\}0&\\text\{if \}t\_\{i,j\}=0,\\\\\[2\.0pt\] \\theta\_\{\\mathrm\{base\}\}\+\\theta\_\{\\mathrm\{soc\}\}\\cdot c\_\{i,j\}&\\text\{if \}t\_\{i,j\}=1,\\end\{cases\}where in Regime 1,ti,jt\_\{i,j\}comes from the naïvely learnedT^\\widehat\{T\}, and in Regime 2,ti,j=pi,j\(τconf\)t\_\{i,j\}=p\_\{i,j\}\(\\tau\_\{\\mathrm\{conf\}\}\)from the confounding model\.
Regional outcomes are summed and aggregated\. and the training loss is the mean squared error, All parameters are optimized jointly using Adam with a learning rate of0\.0010\.001for10001000epochs\. Regime 1 learns\{T^,θbase,θsoc\}\\\{\\widehat\{T\},\\theta\_\{\\mathrm\{base\}\},\\theta\_\{\\mathrm\{soc\}\}\\\}, while Regime 2 learns\{β0,β1,τconf,θbase,θsoc\}\\\{\\beta\_\{0\},\\beta\_\{1\},\\tau\_\{\\mathrm\{conf\}\},\\theta\_\{\\mathrm\{base\}\},\\theta\_\{\\mathrm\{soc\}\}\\\}\.
##### Results\.
1. 1\.We compare two cases of confounding in subregional treatment allocation, holding region\-level treatments and contextual factors fixed\. The aggregated causal effects differ markedly depending on the level of confounding, showing that the treatment context dependency directly shapes the observed regional outcomes\.
2. 2\.We visualize the ground truth intervention locations alongside the estimates obtained with and without explicitly modeling confounding\. Under strong confounding, the confounding\-aware approach recovers the true intervention locations with high fidelity, while the naïve method from Exp\. 2 retrieves only a small fraction correctly\.
3. 3\.In the low confounding setting, both methods perform similarly, and location probabilities are estimated with comparable accuracy\. This confirms that adjusting for confounding does not reduce performance when confounding is weak\.
4. 4\.For causal effect estimation, the final MSE of subregional effects under low confounding was2\.042\.04\(naïve\) versus1\.601\.60\(confounding aware\)\. Under high confounding, the respective errors were0\.740\.74versus0\.320\.32\. These results show that accounting for confounding improves both intervention location recovery and the accuracy of estimated causal effects by conditioning on the correct contextual information\.
Overall, this experiment demonstrates that explicitly modeling treatment context confounding is essential as it enables accurate recovery of intervention locations and yields more reliable causal effect estimates, particularly when treatment assignment is strongly biased by contextual factors\.相似文章
CausaLab: 面向AI科学家的可扩展交互式因果发现环境
CausaLab 是一个可扩展的环境,用于评估LLM智能体在交互式因果发现中的表现,同时衡量预测准确性和对潜在因果机制的忠实复现。实验揭示了预测与机制复现之间的差距,突显了当前LLM智能体作为实验性因果推理者的局限性。
社交媒体中因果关系提取的大型语言模型:灾害情报的验证框架
本文提出了一个验证框架,用于评估大型语言模型(LLM)在灾害期间从社交媒体帖子中提取因果关系的有效性。通过将LLM生成的结果与基于专家知识的参考图谱进行比较,评估其在识别因果关系方面的可靠性及潜在风险。
SpatialClaw: 重新思考智能体空间推理的动作接口
SpatialClaw是一个无需训练的框架,它采用代码作为动作接口,使视觉语言模型能够进行灵活、有状态的空间推理,在多种3D/4D空间推理任务上取得了卓越性能。
不确定性引导的LLM语义增强用于异质性治疗效果估计
本文提出CURL,一种即插即用适配器,利用估计器不确定性分配预训练LLM的语义能力,以改进异质性治疗效果(CATE)估计。它引入两种角色调节的提示,构建面向分配和异质性的表示,在四个基准上提升了十个宿主学习器的性能。
用于地理临界点早期预警的SpatioTemporal Causal Network Diagnostics
本文介绍了SpatioTemporal Causal Network Diagnostics (ST-CND)框架,该框架利用数据驱动的因果网络和动态模态分解,为地理临界点提供局部早期预警,在合成和观测基准上优于经典空间指标。