One Intervention per Component is Enough: Towards Identifiability in Linear Stochastic Dynamics from Steady State
Summary
This paper establishes that one intervention per strongly connected component suffices to recover parameters of linear stochastic dynamics from steady-state data, and proposes a recursive learning algorithm.
View Cached Full Text
Cached at: 09/18/26, 09:13 AM
# One Intervention per Component is Enough:Towards Identifiability in Linear Stochastic Dynamics from Steady State Source: [https://arxiv.org/html/2609.19955](https://arxiv.org/html/2609.19955) Saber SalehkaleybarAffiliation:Leiden Institute of Advanced Computer Science, Leiden University, The NetherlandsCorrespondence to:[s\.salehkaleybar@liacs\.leidenuniv\.nl](mailto:[email protected]) ###### Abstract We study the problem of recovering the parameters of a multivariate Ornstein–Uhlenbeck \(OU\) process from steady\-state observational and interventional data\. In many applications, such as large\-scale gene perturbation experiments, only stationary “snapshot” measurements are available, making standard stochastic differential equation estimation methods that rely on time\-series trajectories inapplicable\. We first establish an identifiability result: one intervention per strongly connected component \(SCC\) of the drift graph suffices to recover all OU process parameters generically up to a global scaling factor\. This holds provided that the SCC condensation graph is connected with a single root and certain spectral nondegeneracy assumptions hold\. We propose a recursive learning algorithm that orders SCCs topologically and, for each component, isolates its marginal dynamics and solves a linear system derived from the steady\-state moment equations, leveraging parameters recovered for upstream components\. Building on this theoretical foundation, we propose a regularized least\-squares estimator that jointly minimizes residuals of the steady\-state mean and covariance equations across observational and interventional data\. Experimental results validate our theoretical findings in recovering parameters of the underlying OU process\. ###### Keywords: Machine Learning, ICML ## 1Introduction Understanding the structure and parameters of dynamical systems from data is a fundamental problem in many scientific domains, including systems biology\([Bansal et al\., 2006](https://arxiv.org/html/2609.19955#bib.bib1);[Marbach et al\., 2012](https://arxiv.org/html/2609.19955#bib.bib2)\), neuroscience\([Paninski et al\., 2010](https://arxiv.org/html/2609.19955#bib.bib3)\), and economics\([Hamilton, 2020](https://arxiv.org/html/2609.19955#bib.bib4)\)\. In many settings, such as gene regulatory networks, the underlying dynamics can be modeled by a stochastic differential equation \(SDE\) whose steady\-state distribution encodes both the interaction structure and kinetic parameters\([Gardiner, 2009](https://arxiv.org/html/2609.19955#bib.bib5);[Villaverde et al\., 2016](https://arxiv.org/html/2609.19955#bib.bib6)\)\. Accurately recovering these parameters enables causal representation of interactions among variables, and predictions under unseen perturbations\([Pearl, 2009](https://arxiv.org/html/2609.19955#bib.bib7)\)\. However, in experimental biology and related areas, data are often available only in the form of “snapshot” measurements\([Cao et al\., 2019](https://arxiv.org/html/2609.19955#bib.bib9);[Schiebinger et al\., 2019](https://arxiv.org/html/2609.19955#bib.bib10)\)\. The time\-series trajectories \(required by many SDE estimation methods\([Zhang et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib11);[Oh et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib12)\)\) are expensive or infeasible to obtain at scale\. In such cases, interventions such as gene knockdown\([Datlinger et al\., 2017](https://arxiv.org/html/2609.19955#bib.bib13)\)provide additional information by selectively intervening on parts of the system, potentially resolving non\-identifiability issues that arise from observational data alone\([Peters et al\., 2017](https://arxiv.org/html/2609.19955#bib.bib18)\)\. A substantial body of work addresses causal inference in linear stochastic systems\.\([Varando and Hansen, 2020](https://arxiv.org/html/2609.19955#bib.bib14);[Dettling et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib17)\)focus on recovering drift structures from steady\-state data using the Lyapunov equation\. These approaches rely solely on observational measurements and do not leverage interventions, which can limit identifiability\. More recently,\([Lorch et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib15)\), focused on fitting the observational/interventional distributions without providing parameter recovery guarantees\. Intervention\-aware models, such as in\([Rohbeck et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib16)\), achieve good empirical prediction, but the identifiability results are very restricted \(see the related work section for more details\)\. In this paper, our aim is to incorporate interventions into the learning process, establishing identifiability results on recovering the parameters of SDEs and developing learning algorithms grounded in these theoretical results\. Our contributions are as follows: - •We establish an identifiability result for recovering the parameters of a multivariate Ornstein–Uhlenbeck \(OU\) process from steady\-state observational and interventional data\. We show that a single intervention per SCC of the drift graph is sufficient to recover the parameters up to a global scaling111In OU processes, from the stationary distribution, parameters are identifiable at most up to a global scaling, an intrinsic ambiguity that cannot be resolved from stationary data alone\.if the SCC condensation graph is connected with a single root and certain spectral/rank nondegeneracy assumptions \(Assumptions[A\.1](https://arxiv.org/html/2609.19955#A1.Thmtheorem1)and[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)\) hold\. To prove this, we first consider the case where the drift matrix is a single SCC and show that by solving the linear system of equations for the first and second moments \(based on observational data and data from one intervention\), we can recover all the parameters up to a global scaling \(Theorem[3\.1](https://arxiv.org/html/2609.19955#S3.Thmtheorem1)\)\. - •For the general case with multiple SCCs \(Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6)\), we provide a recursive algorithm that topologically orders the SCCs and, for each component, isolates its marginal dynamics\. This reduces the problem to a sequence of single\-SCC cases\. For each SCC, the algorithm solves a linear system \(derived from the steady\-state moment equations for the observational and interventional settings\) using parameters recovered for upstream components\. We also provide an example that if the SCC condensation graph has more than one root, then it is impossible to learn all the parameters with a single global scaling\. - •We show that changes in the steady\-state means under single\-node interventions reveal the SCC\-level structure of the graph\. In particular, Theorem[3\.3](https://arxiv.org/html/2609.19955#S3.Thmtheorem3)characterizes the set of nodes whose means change after an intervention as the downstream region of the intervened SCC\. Consequently, one intervention per SCC is sufficient to recover the SCC decomposition and a topological order of the DAG over SCCs \(Remark[3\.4](https://arxiv.org/html/2609.19955#S3.Thmtheorem4)\)\. This structural recovery step does not require the spectral/rank nondegeneracy assumptions used for parameter identification\. - •Building on these theoretical findings, we formulate a regularized least\-squares optimization problem that jointly minimizes the squared residuals of the steady\-state mean and covariance equations across both observational and interventional data\. Empirical results validate the identifiability results in recovering parameters and predicting unseen interventions\. ## 2Problem Formulation Ornstein–Uhlenbeck process: The Ornstein–Uhlenbeck \(OU\) process is a continuous\-time stochastic process that satisfies the stochastic differential equation \(SDE\): d𝐱=\(−𝚲𝐱\+𝐛\)dt\+𝝈d𝐖t,d\\mathbf\{x\}=\(\-\\bm\{\\Lambda\}\\mathbf\{x\}\+\\mathbf\{b\}\)\\,dt\+\\bm\{\\sigma\}\\,d\\mathbf\{W\}\_\{t\},\(1\)where𝐱∈ℝn\\mathbf\{x\}\\in\\mathbb\{R\}^\{n\}is the state vector,𝚲∈ℝn×n\\bm\{\\Lambda\}\\in\\mathbb\{R\}^\{n\\times n\}is a drift matrix \(assumed positive stable\),𝐛∈ℝn\\mathbf\{b\}\\in\\mathbb\{R\}^\{n\}is a constant vector representing external input,𝝈∈ℝn×n\\bm\{\\sigma\}\\in\\mathbb\{R\}^\{n\\times n\}is the diffusion matrix, andd𝐖td\\mathbf\{W\}\_\{t\}is annn\-dimensional Wiener process\. The stationary distribution of the OU process has a multivariate normal distribution with the following mean and covariance: - •Mean: 𝝁=𝔼\[𝐱∞\]=𝚲−1𝐛,\\bm\{\\mu\}=\\mathbb\{E\}\[\\mathbf\{x\}\_\{\\infty\}\]=\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{b\},\(2\)where𝐱∞\\mathbf\{x\}\_\{\\infty\}denotes a random vector distributed according to the stationary distribution of the OU process\. - •Covariance: 𝚲𝚺\+𝚺𝚲⊤=𝝈𝝈⊤,\\bm\{\\Lambda\}\\bm\{\\Sigma\}\+\\bm\{\\Sigma\}\\bm\{\\Lambda\}^\{\\top\}=\\bm\{\\sigma\}\\bm\{\\sigma\}^\{\\top\},\(3\)where𝚺∈ℝn×n\\bm\{\\Sigma\}\\in\\mathbb\{R\}^\{n\\times n\}is the steady\-state covariance matrix:𝚺=𝔼\[\(𝐱∞−𝝁\)\(𝐱∞−𝝁\)⊤\]\.\\bm\{\\Sigma\}=\\mathbb\{E\}\[\(\\mathbf\{x\}\_\{\\infty\}\-\\bm\{\\mu\}\)\(\\mathbf\{x\}\_\{\\infty\}\-\\bm\{\\mu\}\)^\{\\top\}\]\. We assume that the diffusion power matrix is diagonal and positive:𝐃:=𝝈𝝈⊤=diag\(d1,…,dn\)≻0\\mathbf\{D\}:=\\bm\{\\sigma\}\\bm\{\\sigma\}^\{\\\!\\top\}=\\operatorname\{diag\}\(d\_\{1\},\\dots,d\_\{n\}\)\\succ 0\. Intervention:We define an intervention onii\-th coordinate of𝐱\\mathbf\{x\}as follows, where theii\-th row of the drift matrix𝚲\\bm\{\\Lambda\}is modified to remove influence from all other variables, while preserving its self\-regulation\. Formally, we define the modified drift matrix𝚲~\(i\)\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}as: Λ~kj\(i\)=\{Λkj,i≠k,0,i=k,j≠i,Λii,i=k,j=i\.\\widetilde\{\\Lambda\}^\{\(i\)\}\_\{kj\}=\\begin\{cases\}\\Lambda\_\{kj\},&i\\neq k,\\\\ 0,&i=k,\\,j\\neq i,\\\\ \\Lambda\_\{ii\},&i=k,\\,j=i\.\\end\{cases\}That is, we zero out all off\-diagonal entries in theii\-th row, but retainΛii\\Lambda\_\{ii\}, analogous to a knockout perturbation that isolates a variable from its regulators\. This notion of intervention is often considered in the causal SDE literature, see, e\.g\.,\([Hansen and Sokol, 2014](https://arxiv.org/html/2609.19955#bib.bib27)\)and\([Boeken and Mooij, 2024](https://arxiv.org/html/2609.19955#bib.bib28)\), where some post\-intervention SDEs are formulated in this form\. The modified dynamic under intervention is:d𝐱=\(−𝚲~\(i\)𝐱\+𝐛\)dt\+𝝈d𝐖t\.d\\mathbf\{x\}=\(\-\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\\mathbf\{x\}\+\\mathbf\{b\}\)\\,dt\+\\bm\{\\sigma\}\\,d\\mathbf\{W\}\_\{t\}\.In the intervened system, the mean is denoted by𝝁\(i\)\\bm\{\\mu\}^\{\(i\)\}and satisfies: 𝝁\(i\)=𝔼\[𝐱∞\(i\)\]=\(𝚲~\(i\)\)−1𝐛,\\bm\{\\mu\}^\{\(i\)\}=\\mathbb\{E\}\[\\mathbf\{x\}\_\{\\infty\}^\{\(i\)\}\]=\\left\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\\right\)^\{\-1\}\\mathbf\{b\},\(4\)assuming𝚲~\(i\)\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}is invertible\. Similarly, the steady\-state covariance𝚺\(i\)\\bm\{\\Sigma\}^\{\(i\)\}satisfies the Lyapunov equation: 𝚲~\(i\)𝚺\(i\)\+𝚺\(i\)\(𝚲~\(i\)\)⊤=𝝈𝝈⊤\.\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\\bm\{\\Sigma\}^\{\(i\)\}\+\\bm\{\\Sigma\}^\{\(i\)\}\\left\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\\right\)^\{\\top\}=\\bm\{\\sigma\}\\bm\{\\sigma\}^\{\\top\}\.\(5\) We assume that, for every intervention considered,𝚲~\(i\)\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}is positive stable, so the interventional stationary mean and covariance are well\-defined\. Our goal is to recover the parameters of OU process, i\.e\.,𝚲,𝐛\\mathbf\{\\Lambda\},\\mathbf\{b\}, and𝑫\\bm\{D\}from the observational mean and covariance\(𝝁,𝚺\)\(\\bm\{\\mu\},\\mathbf\{\\Sigma\}\)and a collection of interventional means and covariances\{\(𝝁\(i\),𝚺\(i\)\)\}i∈ℐ\\\{\(\\bm\{\\mu\}^\{\(i\)\},\\mathbf\{\\Sigma\}^\{\(i\)\}\)\\\}\_\{i\\in\\mathcal\{I\}\}whereℐ\\mathcal\{I\}is the set of coordinates intervened on\. Please note that the parameters are identifiable only up to a global scaling because the stationary first\-moment and second\-moment equations are invariant under global scaling of parameters\. Graph definitions:In the following, we briefly review the graph\-theoretic concepts used throughout the paper: \- Strongly connected components\.For a directed graphGG, a*strongly connected component \(SCC\)*is a maximal subset of nodesCCsuch that for every pairu,v∈Cu,v\\in Cthere exists a directed path fromuutovvand a directed path fromvvtouu\. The SCCs ofGGform a partition of its nodes\. \- Condensation graph\.Given the SCCsC1,…,CKC\_\{1\},\\dots,C\_\{K\}ofGG, the*condensation graph*is a directed graph whose nodes correspond to the SCCs and which contains a directed edgeCi→CjC\_\{i\}\\to C\_\{j\}whenever there exists an edge inGGfrom some node inCiC\_\{i\}to some node inCjC\_\{j\}\. By construction, the condensation graph is always a directed acyclic graph \(DAG\)\. A root \(or source\) SCC has no incoming edge from another SCC\. Connectedness of this DAG refers to its underlying undirected graph; with finitely many SCCs, a unique root reaches every SCC\. \- Topological order over SCCs\.A*topological order*of the SCCs is any ordering of the nodes of the condensation graph such that all directed edges point from earlier to later components\. Equivalently,CiC\_\{i\}may appear beforeCjC\_\{j\}in the order if and only if there is no directed path fromCjC\_\{j\}toCiC\_\{i\}in the condensation graph\. ## 3Identifiability Results For a drift matrix𝚲\\bm\{\\Lambda\}, define the associated directed graphG\(𝚲\)G\(\\bm\{\\Lambda\}\)by considering an edgej→ij\\to iif and only ifΛij≠0\\Lambda\_\{ij\}\\neq 0\. In the following, first we consider thatG\(𝚲\)G\(\\bm\{\\Lambda\}\)is strongly connected\. Under some spectral/rank non\-degeneracy assumptions, we show that the parameters of the OU process can be recovered generically up to some global scaling by just having one intervention on any coordinate\. We then consider the case whereG\(𝚲\)G\(\\bm\{\\Lambda\}\)has multiple SCCs\. In this setting, we first show that changes in interventional means identify the SCCs and a topological ordering over the SCC condensation DAG222This structural step does not require the spectral/rank nondegeneracy assumptions\.\. Given thisSCC\-level structure, we prove generic identification up to a single global scaling in the graph\-compatible model class, under the unique\-root, intervention\-coverage, and isolated\-realization hypotheses of Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6)\. ### 3\.1Single SCC ###### Theorem 3\.1\. Consider the OU process in equation[1](https://arxiv.org/html/2609.19955#S2.E1)with true parameters𝚲,𝐛\\mathbf\{\\Lambda\},\\mathbf\{b\}, and𝐃\\bm\{D\}\. Moreover, suppose that we have access to the observational steady\-state mean and covariance\(𝛍,𝚺\)\(\\bm\{\\mu\},\\bm\{\\Sigma\}\)and an interventional mean and covariance\(𝛍\(i\),𝚺\(i\)\)\(\\bm\{\\mu\}^\{\(i\)\},\\bm\{\\Sigma\}^\{\(i\)\}\)\(iican be any coordinate\)\. If the graphG\(𝚲\)G\(\\bm\{\\Lambda\}\)is strongly connected and certain spectral/rank non\-degeneracy assumptions hold \(See Assumption[A\.1](https://arxiv.org/html/2609.19955#A1.Thmtheorem1)and Assumption[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)in Appendix[A\.6](https://arxiv.org/html/2609.19955#A1.SS6)\), generically333Please see the definition of genericity in Appendix[A\.2](https://arxiv.org/html/2609.19955#A1.SS2)\., any other parameter triple\(𝚲^,𝐛^,𝐃^\)\(\\widehat\{\\bm\{\\Lambda\}\},\\widehat\{\\mathbf\{b\}\},\\widehat\{\\mathbf\{D\}\}\)that yields the same observational and interventional moments must satisfy: 𝚲^=c𝚲,𝐛^=c𝐛,𝐃^=c𝐃,\\widehat\{\\bm\{\\Lambda\}\}=c\\bm\{\\Lambda\},\\quad\\widehat\{\\mathbf\{b\}\}=c\\mathbf\{b\},\\quad\\widehat\{\\mathbf\{D\}\}=c\\mathbf\{D\},for some scalarc\>0c\>0\. All the proofs of theorems \(if not given in the main body\) are available in the appendix\. The key idea in the proof is based on forming a system of linear equations according to equation[2](https://arxiv.org/html/2609.19955#S2.E2), equation[3](https://arxiv.org/html/2609.19955#S2.E3), equation[4](https://arxiv.org/html/2609.19955#S2.E4), and equation[5](https://arxiv.org/html/2609.19955#S2.E5), showing that any possible solution for this set of equations should be a scale of the true parameters\. In particular, let Θ=\[vec\(𝚲\)𝐛𝐝\]∈ℝp,\\Theta=\\begin\{bmatrix\}\\operatorname\{vec\}\(\\bm\{\\Lambda\}\)\\\\\[2\.0pt\] \\mathbf\{b\}\\\\\[2\.0pt\] \\mathbf\{d\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{p\},wherep:=n2\+2np:=n^\{2\}\+2n, and𝐝:=\(d1,…,dn\)⊤,\\mathbf\{d\}:=\(d\_\{1\},\\dots,d\_\{n\}\)^\{\\\!\\top\},wheredk:=Dkk\>0d\_\{k\}:=D\_\{kk\}\>0\. Moreover, the operatorvec\\operatorname\{vec\}stacks the columns of its matrix argument\. Therefore,vec\(𝚲\)∈ℝn2\\operatorname\{vec\}\(\\bm\{\\Lambda\}\)\\in\\mathbb\{R\}^\{n^\{2\}\}\. Now, we can write the linear equations in the following form:𝐀Θ=𝟎\\mathbf\{A\}\\,\\Theta=\\mathbf\{0\}where𝐀∈ℝm×p\\mathbf\{A\}\\in\\mathbb\{R\}^\{\\,m\\times p\}\(refer to Appendix[A\.1](https://arxiv.org/html/2609.19955#A1.SS1)for the definition of𝐀\\mathbf\{A\}\) andm=n2\+3nm=n^\{2\}\+3n\. In the proof, we show that every admissible solution of this linear system is a positive scaling of the true parameters\. ### 3\.2General Case In the general case, where the graphG\(𝚲\)G\(\\mathbf\{\\Lambda\}\)has multiple SCCs, assume that the intervention setℐ\\mathcal\{I\}contains at least one intervention targeting a variable inside each SCC\. Moreover, the DAG of SCCs is connected and there is exactly one root SCC\. Under these conditions and certain spectral/rank non\-degeneracy assumptions \(see Assumptions[A\.1](https://arxiv.org/html/2609.19955#A1.Thmtheorem1)and[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2), in the appendix\), we show that the parameters of the OU process can be recovered up to a global scaling\. The key idea is to first recover a topological ordering over the SCCs by inspecting which means change under each intervention; this structural recovery step does not require the spectral/rank nondegeneracy assumptions\. The following theorem allows us to identify the set of variables in each SCC, as well as a topological ordering over the SCCs\. With this structural information in hand, we then recursively identify the model parameters for each SCC by conditioning on previously resolved components, thereby reducing the multi\-SCC case to a sequence of single\-SCC problems\. ###### Theorem 3\.3\. LetCCbe an SCC inG\(𝚲\)G\(\\bm\{\\Lambda\}\)and leti∈Ci\\in C\. Consider the intervention oniiand let𝛍\\bm\{\\mu\}and𝛍\(i\)\\bm\{\\mu\}^\{\(i\)\}be the observational and interventional steady\-state means, respectively\. *\(Non\-null case\)\.*If theii\-th row is non\-null, i\.e\.,∃j≠i\\exists\\,j\\neq iwithΛij≠0\\Lambda\_\{ij\}\\neq 0, then, generically, for the set of parameters\(𝚲,𝐛\)\(\\bm\{\\Lambda\},\\mathbf\{b\}\)consistent withG\(𝚲\)G\(\\bm\{\\Lambda\}\), μk\(i\)≠μk⟺k∈Desc\(C\),\\mu^\{\(i\)\}\_\{k\}\\neq\\mu\_\{k\}\\quad\\Longleftrightarrow\\quad k\\in\\mathrm\{Desc\}\(C\),whereDesc\(C\)\\mathrm\{Desc\}\(C\)denotes the set of nodes in SCCs reachable fromCCin the DAG of SCCs \(includingCCitself\)\. *\(Null case\)\.*If theii\-th row has no off\-diagonals, i\.e\.,Λij=0\\Lambda\_\{ij\}=0for allj≠ij\\neq i, then𝝁\(i\)=𝝁\\bm\{\\mu\}^\{\(i\)\}=\\bm\{\\mu\}\. ###### Theorem 3\.6\. Consider the OU process in equation[1](https://arxiv.org/html/2609.19955#S2.E1)with true parameters𝚲,𝐛,𝐃\\bm\{\\Lambda\},\\mathbf\{b\},\\mathbf\{D\}\. Suppose we have access to the observational steady\-state mean and covariance\(𝛍,𝚺\)\(\\bm\{\\mu\},\\bm\{\\Sigma\}\), as well as to interventional means and covariances\(𝛍\(i\),𝚺\(i\)\)\(\\bm\{\\mu\}^\{\(i\)\},\\bm\{\\Sigma\}^\{\(i\)\}\)for at least one intervention in each SCC ofG\(𝚲\)G\(\\bm\{\\Lambda\}\)\.Assume that, for each non\-singleton SCC and its selected intervention, the admissible isolated SCC model admits a realization satisfying Assumptions[A\.1](https://arxiv.org/html/2609.19955#A1.Thmtheorem1)and[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)\. If the DAG over SCCs ofG\(𝚲\)G\(\\bm\{\\Lambda\}\)is connected with a single root SCC, then, generically, any other admissible parameter triple\(𝚲^,𝐛^,𝐃^\)\(\\widehat\{\\bm\{\\Lambda\}\},\\widehat\{\\mathbf\{b\}\},\\widehat\{\\mathbf\{D\}\}\)compatible with the same graphG\(𝚲\)G\(\\bm\{\\Lambda\}\)that yields the same observational and interventional moments must satisfy 𝚲^=c𝚲,𝐛^=c𝐛,𝐃^=c𝐃,\\widehat\{\\bm\{\\Lambda\}\}=c\\bm\{\\Lambda\},\\quad\\widehat\{\\mathbf\{b\}\}=c\\mathbf\{b\},\\quad\\widehat\{\\mathbf\{D\}\}=c\\mathbf\{D\},for some scalarc\>0c\>0\. ###### Proof\. Based on Theorem[3\.3](https://arxiv.org/html/2609.19955#S3.Thmtheorem3), we can infer the SCCs and also a topological ordering over them if there is at least one intervention in each SCC\. Let us denote these SCCs based on the topological ordering asC1,C2,⋯,CKC\_\{1\},C\_\{2\},\\cdots,C\_\{K\}whereKKis the number of components\. Suppose that we already learned the parameters of the OU process in the componentsC1,C2,⋯,CrC\_\{1\},C\_\{2\},\\cdots,C\_\{r\}up to some global scaling\. Now, we aim for learning the parameters inCr\+1C\_\{r\+1\}\. We partition the state vector according to the SCC decomposition ofG\(𝚲\)G\(\\bm\{\\Lambda\}\): 𝐱=𝐱P⏟C1∪⋯∪Cr⊕𝐱T⏟Cr\+1⊕𝐱F⏟Cr\+2∪⋯∪CK,\\mathbf\{x\}=\\underbrace\{\\mathbf\{x\}\_\{P\}\}\_\{C\_\{1\}\\cup\\cdots\\cup C\_\{r\}\}\\;\\oplus\\;\\underbrace\{\\mathbf\{x\}\_\{T\}\}\_\{C\_\{r\+1\}\}\\;\\oplus\\;\\underbrace\{\\mathbf\{x\}\_\{F\}\}\_\{C\_\{r\+2\}\\cup\\cdots\\cup C\_\{K\}\},where the operator⊕\\oplusdenotes concatenation of subvectors corresponding to disjoint index sets\. The steady\-state mean and covariance matrices are partitioned accordingly: 𝝁=\[𝝁P𝝁T𝝁F\],𝚺=\[𝚺PP𝚺PT𝚺PF𝚺TP𝚺TT𝚺TF𝚺FP𝚺FT𝚺FF\]\.\\bm\{\\mu\}=\\begin\{bmatrix\}\\bm\{\\mu\}\_\{P\}\\\\ \\bm\{\\mu\}\_\{T\}\\\\ \\bm\{\\mu\}\_\{F\}\\end\{bmatrix\},\\qquad\\bm\{\\Sigma\}=\\begin\{bmatrix\}\\bm\{\\Sigma\}\_\{PP\}&\\bm\{\\Sigma\}\_\{PT\}&\\bm\{\\Sigma\}\_\{PF\}\\\\ \\bm\{\\Sigma\}\_\{TP\}&\\bm\{\\Sigma\}\_\{TT\}&\\bm\{\\Sigma\}\_\{TF\}\\\\ \\bm\{\\Sigma\}\_\{FP\}&\\bm\{\\Sigma\}\_\{FT\}&\\bm\{\\Sigma\}\_\{FF\}\\end\{bmatrix\}\.Everything inside thePP\-block is assumed known, i\.e\., the blocks𝚲PP\\bm\{\\Lambda\}\_\{PP\},𝐛P\\mathbf\{b\}\_\{P\}, and𝐃P\\mathbf\{D\}\_\{P\}, up to the same scalingcc\. Because𝐱P\\mathbf\{x\}\_\{P\}and𝐱T\\mathbf\{x\}\_\{T\}are jointly Gaussian, we can write: 𝐱T=𝐁𝐱P\+𝐫,where𝐁:=𝚺TP𝚺PP−1\.\\mathbf\{x\}\_\{T\}=\\mathbf\{B\}\\,\\mathbf\{x\}\_\{P\}\+\\mathbf\{r\},\\quad\\text\{where \}\\mathbf\{B\}:=\\bm\{\\Sigma\}\_\{TP\}\\bm\{\\Sigma\}\_\{PP\}^\{\-1\}\.The regression matrix𝐁\\mathbf\{B\}is computable directly from the observed moments, with no dependence on model parameters\. Moreover, the residual term𝐫:=𝐱T−𝐁𝐱P\\mathbf\{r\}:=\\mathbf\{x\}\_\{T\}\-\\mathbf\{B\}\\mathbf\{x\}\_\{P\}is a Gaussian variable with 𝝁T\|P=𝝁T−𝐁𝝁P,𝚺T\|P=𝚺TT−𝚺TP𝚺PP−1𝚺PT\.\\bm\{\\mu\}\_\{T\\mid P\}=\\bm\{\\mu\}\_\{T\}\-\\mathbf\{B\}\\,\\bm\{\\mu\}\_\{P\},\\qquad\\bm\{\\Sigma\}\_\{T\\mid P\}=\\bm\{\\Sigma\}\_\{TT\}\-\\bm\{\\Sigma\}\_\{TP\}\\bm\{\\Sigma\}\_\{PP\}^\{\-1\}\\bm\{\\Sigma\}\_\{PT\}\.\(6\) Substituting the definitions of𝐁\\mathbf\{B\}and𝐫\\mathbf\{r\}into the OU dynamics yields: 𝐫˙=−𝚲TT𝐫\+\(−𝚲TP−𝚲TT𝐁\+𝐁𝚲PP\)𝐱P\+\(𝐛T−𝐁𝐛P\)\+\(𝝈T𝐖˙T−𝐁𝝈P𝐖˙P\)\.\\begin\{split\}\\dot\{\\mathbf\{r\}\}=\-\\bm\{\\Lambda\}\_\{TT\}\\mathbf\{r\}&\+\\bigl\(\-\\bm\{\\Lambda\}\_\{TP\}\-\\bm\{\\Lambda\}\_\{TT\}\\mathbf\{B\}\+\\mathbf\{B\}\\bm\{\\Lambda\}\_\{PP\}\\bigr\)\\mathbf\{x\}\_\{P\}\\\\ &\+\(\\mathbf\{b\}\_\{T\}\-\\mathbf\{B\}\\mathbf\{b\}\_\{P\}\)\+\(\\bm\{\\sigma\}\_\{T\}\\dot\{\\mathbf\{W\}\}\_\{T\}\-\\mathbf\{B\}\\bm\{\\sigma\}\_\{P\}\\dot\{\\mathbf\{W\}\}\_\{P\}\)\.\\end\{split\}\(7\)To identify the residual moment equations, we use the cross\-covariance block of the Lyapunov equation: 𝚲TP𝚺PP\+𝚲TT𝚺TP\+𝚺TP𝚲PP⊤=𝟎\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{TP\}\\bm\{\\Sigma\}\_\{PP\}\+\\bm\{\\Lambda\}\_\{TT\}\\bm\{\\Sigma\}\_\{TP\}\+\\bm\{\\Sigma\}\_\{TP\}\\bm\{\\Lambda\}\_\{PP\}^\{\\\!\\top\}=\\mathbf\{0\}\.\}Using𝐁=𝚺TP𝚺PP−1\\mathbf\{B\}=\\bm\{\\Sigma\}\_\{TP\}\\bm\{\\Sigma\}\_\{PP\}^\{\-1\}, this gives 𝚲TP\+𝚲TT𝐁=−𝐁𝚺PP𝚲PP⊤𝚺PP−1\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{TP\}\+\\bm\{\\Lambda\}\_\{TT\}\\mathbf\{B\}=\-\\mathbf\{B\}\\bm\{\\Sigma\}\_\{PP\}\\bm\{\\Lambda\}\_\{PP\}^\{\\\!\\top\}\\bm\{\\Sigma\}\_\{PP\}^\{\-1\}\.\}Combining this with the parent Lyapunov equation 𝚲PP𝚺PP\+𝚺PP𝚲PP⊤=𝐃P,\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{PP\}\\bm\{\\Sigma\}\_\{PP\}\+\\bm\{\\Sigma\}\_\{PP\}\\bm\{\\Lambda\}\_\{PP\}^\{\\\!\\top\}=\\mathbf\{D\}\_\{P\},\}we obtain 𝚲TP\+𝚲TT𝐁=𝐁𝚲PP−𝐁𝐃P𝚺PP−1\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{TP\}\+\\bm\{\\Lambda\}\_\{TT\}\\mathbf\{B\}=\\mathbf\{B\}\\bm\{\\Lambda\}\_\{PP\}\-\\mathbf\{B\}\\mathbf\{D\}\_\{P\}\\bm\{\\Sigma\}\_\{PP\}^\{\-1\}\.\}\(8\)For the mean, using 𝚲TP𝝁P\+𝚲TT𝝁T=𝐛T,𝝁T=𝝁T\|P\+𝐁𝝁P,\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{TP\}\\bm\{\\mu\}\_\{P\}\+\\bm\{\\Lambda\}\_\{TT\}\\bm\{\\mu\}\_\{T\}=\\mathbf\{b\}\_\{T\},\\qquad\\bm\{\\mu\}\_\{T\}=\\bm\{\\mu\}\_\{T\\mid P\}\+\\mathbf\{B\}\\bm\{\\mu\}\_\{P\},\}together with equation[8](https://arxiv.org/html/2609.19955#S3.E8), gives 𝚲TT𝝁T\|P=𝐛T−𝐁\(𝐛P−𝐃P𝚺PP−1𝝁P\)\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{TT\}\\bm\{\\mu\}\_\{T\\mid P\}=\\mathbf\{b\}\_\{T\}\-\\mathbf\{B\}\\left\(\\mathbf\{b\}\_\{P\}\-\\mathbf\{D\}\_\{P\}\\bm\{\\Sigma\}\_\{PP\}^\{\-1\}\\bm\{\\mu\}\_\{P\}\\right\)\.\}\(9\)Similarly, for the residual covariance, 𝚲TT𝚺T\|P\+𝚺T\|P𝚲TT⊤=𝐃T\+𝐁𝐃P𝐁⊤\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{TT\}\\bm\{\\Sigma\}\_\{T\\mid P\}\+\\bm\{\\Sigma\}\_\{T\\mid P\}\\bm\{\\Lambda\}\_\{TT\}^\{\\\!\\top\}=\\mathbf\{D\}\_\{T\}\+\\mathbf\{B\}\\mathbf\{D\}\_\{P\}\\mathbf\{B\}^\{\\\!\\top\}\.\}\(10\) Now, based on the first and second moments of the residual which are given above, we can identify the parameters of componentTTand also𝚲TP\\bm\{\\Lambda\}\_\{TP\}up to the same scaling\. In particular, we have: 𝚲TT\\mathbf\{\\Lambda\}\_\{TT\}:Note that the corresponding diffusion power matrix of the residual moment equation is𝐃T\+𝐁𝐃P𝐁⊤\\mathbf\{D\}\_\{T\}\+\\mathbf\{B\}\\mathbf\{D\}\_\{P\}\\mathbf\{B\}^\{\\\!\\top\}, and therefore it is not necessarily diagonal\. Nevertheless, the proof of Theorem[3\.1](https://arxiv.org/html/2609.19955#S3.Thmtheorem1)can be adapted to recover𝚲TT\\bm\{\\Lambda\}\_\{TT\}up to the same scalingccby solving the residual moment equations; see Appendix[A\.7](https://arxiv.org/html/2609.19955#A1.SS7)\. 𝚲TP\\mathbf\{\\Lambda\}\_\{TP\}:Having𝚲PP\\bm\{\\Lambda\}\_\{PP\},𝐃P\\mathbf\{D\}\_\{P\}, and𝚲TT\\bm\{\\Lambda\}\_\{TT\}up to the same scalingcc, equation[8](https://arxiv.org/html/2609.19955#S3.E8)gives 𝚲TP=𝐁𝚲PP−𝐁𝐃P𝚺PP−1−𝚲TT𝐁\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{TP\}=\\mathbf\{B\}\\bm\{\\Lambda\}\_\{PP\}\-\\mathbf\{B\}\\mathbf\{D\}\_\{P\}\\bm\{\\Sigma\}\_\{PP\}^\{\-1\}\-\\bm\{\\Lambda\}\_\{TT\}\\mathbf\{B\}\.\}Therefore,𝚲TP\\bm\{\\Lambda\}\_\{TP\}is recovered with the same scaling\. 𝐛T\\mathbf\{b\}\_\{T\}andDT\\bm\{D\}\_\{T\}:According to equation[2](https://arxiv.org/html/2609.19955#S2.E2), we have 𝐛T=𝚲TT𝝁T\+𝚲TP𝝁P\.\\mathbf\{b\}\_\{T\}=\\bm\{\\Lambda\}\_\{TT\}\\bm\{\\mu\}\_\{T\}\+\\bm\{\\Lambda\}\_\{TP\}\\bm\{\\mu\}\_\{P\}\.Since we recovered𝚲TT\\bm\{\\Lambda\}\_\{TT\}and𝚲TP\\bm\{\\Lambda\}\_\{TP\}with the same scaling factor,𝒃T\\bm\{b\}\_\{T\}is identifiable with the same scaling from the above equation\. Regarding𝝈T\\bm\{\\sigma\}\_\{T\}\(or diffusion power matrix𝐃T\\mathbf\{D\}\_\{T\}\), from the Lyapunov equation for the covariance matrix of residual \(in other words,𝚺T\|P\\bm\{\\Sigma\}\_\{T\|P\}\), we have: 𝚲TT𝚺T\|P\+𝚺T\|P𝚲TT⊤=𝐃T\+𝐁𝐃P𝐁⊤\.\\bm\{\\Lambda\}\_\{TT\}\\bm\{\\Sigma\}\_\{T\\mid P\}\+\\bm\{\\Sigma\}\_\{T\\mid P\}\\bm\{\\Lambda\}\_\{TT\}^\{\\\!\\top\}=\\mathbf\{D\}\_\{T\}\+\\mathbf\{B\}\\,\\mathbf\{D\}\_\{P\}\\,\\mathbf\{B\}^\{\\\!\\top\}\.Therefore, 𝐃T=diag\(𝚲TT𝚺T\|P\+𝚺T\|P𝚲TT⊤−𝐁𝐃P𝐁⊤\),\{\\color\[rgb\]\{0,0,0\}\\mathbf\{D\}\_\{T\}=\\operatorname\{diag\}\\\!\\left\(\\bm\{\\Lambda\}\_\{TT\}\\bm\{\\Sigma\}\_\{T\\mid P\}\+\\bm\{\\Sigma\}\_\{T\\mid P\}\\bm\{\\Lambda\}\_\{TT\}^\{\\\!\\top\}\-\\mathbf\{B\}\\mathbf\{D\}\_\{P\}\\mathbf\{B\}^\{\\\!\\top\}\\right\),\}and hence𝐃T\\mathbf\{D\}\_\{T\}is learned with the same scaling\. This completes the recursive step and the proof\.∎ Building on the above, we design a recursive algorithm that proceeds in two stages \(the pseudo\-code is given in Algorithm[1](https://arxiv.org/html/2609.19955#alg1)\): - •In the first stage, we use the changes in interventional means to identify the SCCs of the drift graphG\(𝚲\)G\(\\bm\{\\Lambda\}\), along with a topological ordering over these components\. This structural information is inferred using Theorem[3\.3](https://arxiv.org/html/2609.19955#S3.Thmtheorem3)\. - •In the second stage, we iterate over the SCCs in topological order\. For each componentTT, we treat the union of previously processed SCCs asPP, and condition on𝐱P\\mathbf\{x\}\_\{P\}to isolate the marginal dynamics of𝐱T\\mathbf\{x\}\_\{T\}\. Using the conditional moments\(𝝁T\|P,𝚺T\|P\)\(\\bm\{\\mu\}\_\{T\\mid P\},\\bm\{\\Sigma\}\_\{T\\mid P\}\), we first recover𝚲TT\\bm\{\\Lambda\}\_\{TT\}up to a scaling\. Then, leveraging the structure of the OU dynamics, we recover𝚲TP,𝒃T\\bm\{\\Lambda\}\_\{TP\},\\;\\bm\{b\}\_\{T\}, and𝑫T\\bm\{D\}\_\{T\}up to the same scaling\. The procedure continues recursively until all components have been identified\. Algorithm 1Recursive OU Learning with Interventions1:Input:Observational and interventional means and covariances \(𝝁,𝚺\),\{\(𝝁\(i\),𝚺\(i\)\)\}i∈ℐ\(\\bm\{\\mu\},\\bm\{\\Sigma\}\),\\\{\(\\bm\{\\mu\}^\{\(i\)\},\\bm\{\\Sigma\}^\{\(i\)\}\)\\\}\_\{i\\in\\mathcal\{I\}\} 2:Phase 1: Learn SCCs 3:Identify SCCs, C1,C2,…,CKC\_\{1\},C\_\{2\},\\dots,C\_\{K\}, and a topological ordering over them from mean changes using Theorem[3\.3](https://arxiv.org/html/2609.19955#S3.Thmtheorem3) 4:Phase 2: Recover parameters per SCC 5:foreach component T=CjT=C\_\{j\}in topological orderdo 6:Let P:=C1∪⋯∪Cj−1P:=C\_\{1\}\\cup\\cdots\\cup C\_\{j\-1\}be the union of previously processed SCCs 7:Compute conditional mean 𝝁T\|P\\bm\{\\mu\}\_\{T\\mid P\}and covariance 𝚺T\|P\\mathbf\{\\Sigma\}\_\{T\\mid P\}according to equation[6](https://arxiv.org/html/2609.19955#S3.E6) 8:Recover 𝚲TT\\bm\{\\Lambda\}\_\{TT\}using the interventional moments with the same scaling in PP 9:Recover 𝚲TP,𝒃T,𝑫T\\bm\{\\Lambda\}\_\{TP\},\\;\\bm\{b\}\_\{T\},\\;\\bm\{D\}\_\{T\}up to the same scaling using cross\-covariances and stationary conditions 10:endfor 11:Output:Drift matrix 𝚲\\bm\{\\Lambda\}, input vector 𝒃\\bm\{b\}, and diffusion matrix 𝑫\\bm\{D\}\(up to a global scaling\) ## 4Learning Algorithm In finite\-sample settings, plugging in empirical moments generally yields full\-column\-rank systems, which admit only the trivial solution of zero vector\. Nevertheless, the identifiability result in Theorem[3\.1](https://arxiv.org/html/2609.19955#S3.Thmtheorem1)motivates replacing true moments with empirical ones and relaxing the equations into a least\-squares objective\. Specifically, in the observational case, the mean vector𝝁\\bm\{\\mu\}and covariance matrix𝚺\\bm\{\\Sigma\}satisfy the linear system in equation[2](https://arxiv.org/html/2609.19955#S2.E2)and equation[3](https://arxiv.org/html/2609.19955#S2.E3), respectively\. Moreover, for each interventioni∈ℐi\\in\\mathcal\{I\}, the mean vector𝝁\(i\)\\bm\{\\mu\}^\{\(i\)\}and𝚺\(i\)\\bm\{\\Sigma\}^\{\(i\)\}satisfy equations in equation[4](https://arxiv.org/html/2609.19955#S2.E4)and equation[5](https://arxiv.org/html/2609.19955#S2.E5), respectively\. For the set of free parametersΘ=\(vec\(𝚲\),𝐛,𝐝\)\\Theta=\(\\operatorname\{vec\}\(\\bm\{\\Lambda\}\),\\mathbf\{b\},\\mathbf\{d\}\), where𝐝\\mathbf\{d\}is the diagonal of matrix𝐃\\mathbf\{D\}, we define the following least\-square objective: ℒ\(Θ\)=αO\(‖𝚲𝝁^−𝐛‖22\+‖𝚲𝚺^\+𝚺^𝚲⊤−𝐃‖F2\)\+αI∑i∈ℐ\(‖𝚲~\(i\)𝝁^\(i\)−𝐛‖22\)\+αI∑i∈ℐ\(‖𝚲~\(i\)𝚺^\(i\)\+𝚺^\(i\)\(𝚲~\(i\)\)⊤−𝐃‖F2\),\\begin\{split\}&\\mathcal\{L\}\(\\Theta\)=\\alpha\_\{O\}\\left\(\\left\\\|\\bm\{\\Lambda\}\\,\\hat\{\\bm\{\\mu\}\}\-\\mathbf\{b\}\\right\\\|\_\{2\}^\{2\}\+\\left\\\|\\bm\{\\Lambda\}\\,\\hat\{\\bm\{\\Sigma\}\}\+\\hat\{\\bm\{\\Sigma\}\}\\,\\bm\{\\Lambda\}^\{\\\!\\top\}\-\\mathbf\{D\}\\right\\\|\_\{F\}^\{2\}\\right\)\\\\ &\+\\alpha\_\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\left\(\\left\\\|\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\\,\\hat\{\\bm\{\\mu\}\}^\{\(i\)\}\-\\mathbf\{b\}\\right\\\|\_\{2\}^\{2\}\\right\)\\\\ &\+\\alpha\_\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\left\(\\left\\\|\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\\,\\hat\{\\bm\{\\Sigma\}\}^\{\(i\)\}\+\\hat\{\\bm\{\\Sigma\}\}^\{\(i\)\}\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\)^\{\\\!\\top\}\-\\mathbf\{D\}\\right\\\|\_\{F\}^\{2\}\\right\),\\end\{split\}\(11\)whereαO,αI\>0\\alpha\_\{O\},\\alpha\_\{I\}\>0are weighting coefficients; we often give a larger weight to observational terms since observational samples are often more abundant and their estimates are more accurate\. Moreover,𝝁^\\hat\{\\bm\{\\mu\}\},𝚺^\\hat\{\\bm\{\\Sigma\}\},𝝁^\(i\)\\hat\{\\bm\{\\mu\}\}^\{\(i\)\},𝚺^\(i\)\\hat\{\\bm\{\\Sigma\}\}^\{\(i\)\},i∈ℐi\\in\\mathcal\{I\}are the unbiased estimates of first and second moments in the observational and interventional settings and∥⋅∥F\\\|\\cdot\\\|\_\{F\}is the Frobenius norm\. To promote sparsity in the drift graph, we add anℓ1\\ell\_\{1\}penalty on the off\-diagonal elements of𝚲\\bm\{\\Lambda\}:ℛ\(𝚲\)=γ∑i≠j\|Λij\|,\\mathcal\{R\}\(\\bm\{\\Lambda\}\)=\\gamma\\sum\_\{i\\neq j\}\|\\Lambda\_\{ij\}\|,whereγ\>0\\gamma\>0controls the sparsity level\. Therefore, the optimization problem becomes minΘℒ\(Θ\)\+ℛ\(𝚲\),\\min\_\{\\Theta\}\\ \\mathcal\{L\}\(\\Theta\)\+\\mathcal\{R\}\(\\bm\{\\Lambda\}\),\(12\)subject to positive stability of𝚲\\bm\{\\Lambda\}and the intervened drifts, withdk\>0d\_\{k\}\>0for allkk\. Because the moment equations are homogeneous in\(𝚲,𝐛,𝐃\)\(\\bm\{\\Lambda\},\\mathbf\{b\},\\mathbf\{D\}\), we can impose a scale\-fixing normalization, such astr\(𝚲\)=n\\operatorname\{tr\}\(\\bm\{\\Lambda\}\)=n, to remove the global scaling ambiguity\. ## 5Related Work Herein, we mainly review methods that perform inference from stationary distributions of dynamical systems, rather than from full trajectories444There are several surveys on causal discovery from temporal data \(e\.g\.,\([Gong et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib29);[Hasan et al\., 2023](https://arxiv.org/html/2609.19955#bib.bib30)\)\)\. Identifiability and learning from steady state in SDE models\.A line of work on graphical continuous Lyapunov models treats the steady\-state covariance as the solution of a continuous Lyapunov equation\. In this setting,\([Varando and Hansen, 2020](https://arxiv.org/html/2609.19955#bib.bib14)\)proposed anl1l\_\{1\}\-regularized estimator to recover sparse drift structure from observational snapshots\.\([Dettling et al\., 2023](https://arxiv.org/html/2609.19955#bib.bib25)\)subsequently analyzed identifiability in this framework, proving that when only observational covariances are available and the diffusion matrix is known, the drift is globally identifiable from the covariance if and only if the drift graph is simple, meaning it contains no directed two\-cycles\. While these results provide valuable insights, they are limited to observational data, cannot recover models with two\-cycles in the drift, and do not handle unknown diffusion\. In contrast, our work incorporates interventional data, allows unknown diagonal diffusion, and shows that a single intervention in each SCC \(under some conditions on the DAG of SCCs\) suffices for recovery up to a global scaling\. More recently,\([Dettling et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib17)\)proposed a Lasso\-based estimator for recovering the drift structure of linear SDEs from stationary observational covariances and analyzed its identifiability properties\. However, their framework does not incorporate interventions\.Very recently,\([Zweig et al\., 2025](https://arxiv.org/html/2609.19955#bib.bib26)\)studied identifiability of linear SDEs under mean\-shift interventions, assuming a low\-rank drift matrix\. In contrast, we allow general drift matrices and base identifiability on the SCC structure of the drift graph rather than on a low rank assumption\. Other approaches aim to learn stationary SDE models without focusing on identifiability\.\([Lorch et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib15)\)proposed a kernel deviation\-from\-stationarity objective that measures how far a candidate SDE’s stationary distribution deviates from the empirical distribution\. Their framework can accommodate cycles and generalizes well to unseen interventions, but its goal is density fitting in reproducing kernel Hilbert spaces rather than moment\-based parameter identification\. Our work differs by deriving explicit graphical conditions under which the OU parameters are recoverable from first and second moments\. Recent empirical work has also combined steady\-state dynamics with interventions\.\([Rohbeck et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib16)\)introduced Bicycle, a model in which interventions alter a subset of parameters\. Bicycle achieves good performance in both structure recovery and prediction under out\-of\-distribution interventions in single\-cell datasets\. However, it is designed primarily as a predictive model and provides identifiability guarantees only when interventions are performed on all coordinates except one\. In contrast, we gave identifiability results under substantially weaker requirements, i\.e\., one intervention per SCC\. \([Boege et al\., 2025](https://arxiv.org/html/2609.19955#bib.bib19)\)have studied the conditional independence relations implied by sparsity in the drift of stationary multivariate diffusions\. Their results link graph structure to conditional independencies in the stationary state distribution\. Nonetheless, this line of work does not address the recovery of drift or diffusion parameters, nor how interventions could enable identifiability\. Finally, there exists related work for deterministic linear ODEs rather than stochastic OU models\. For example,\([Wang et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib20)\)analyzed the identifiability of linear ODE systems with hidden confounders from time series or discretely sampled trajectories and derived conditions under which latent confounding can be resolved\. Gene perturbation prediction from steady state\.In systems biology, several methods aim to predict the effects of genetic perturbations directly from stationary data\. For instance,\([Sethuraman et al\., 2023](https://arxiv.org/html/2609.19955#bib.bib21)\)proposed NODAGS\-Flow, which learns nonlinear cyclic causal structures from interventional steady\-state gene expression by fitting residual normalizing flows, producing predictions and plausible graphs\. While providing good performance in practice, NODAGS\-Flow is a likelihood\-based method without formal identifiability guarantees on recovering the parameters of the underlying system\. Other works focus on high\-performing predictive models\. For instance,\([Roohani et al\., 2022](https://arxiv.org/html/2609.19955#bib.bib22)\)proposed GEAR, which combines graph neural networks with prior biological network information to predict transcriptional responses to single or multiple gene perturbations, showing improved generalization to unseen combinations\.\([Yu et al\., 2025](https://arxiv.org/html/2609.19955#bib.bib23)\)proposed PerturbNet, which uses conditional invertible flows to model the distributional effects of unseen perturbations\. PerturBench\([Wu et al\.,](https://arxiv.org/html/2609.19955#bib.bib24)\)provides a unified benchmark suite for single\-cell perturbation modeling, facilitating comparison between predictive models\. \(a\)𝚲\\bm\{\\Lambda\} \(b\)𝑫\\bm\{D\} \(c\)𝒃\\bm\{b\} \(d\)Constraints satisfaction Figure 1:\(a–c\) Relative errors of estimated parameters of the OU process versus the number of interventions, and \(d\) percentage of instances satisfying graphical conditions in Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6)\. ## 6Experiments #### Scope of empirical evaluation\. The goal of this section is to validate the identifiability results \(e\.g\., parameter recovery\)\. Regarding comparisons, we focus on methods that assume linear SDE models and operate on stationary data\. In particular,\([Rohbeck et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib16)\)considered a Lyapunov\-based loss that corresponds to our objective without the mean terms; we include this variant explicitly as an ablation \(*Covariance only*\)\.\([Varando and Hansen, 2020](https://arxiv.org/html/2609.19955#bib.bib14)\)proposed anℓ1\\ell\_\{1\}\-regularized estimator based solely on observational covariances \(without interventions\), and is therefore not designed for the interventional setting\.\([Lorch et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib15)\)also proposed a loss for linear SDEs, but their interventions are implemented as mean shifts rather than the hard interventions assumed in our work\. We adapted their loss to our hard\-intervention setting and evaluated it on synthetic data\. However, we did not observe improvements in performance as the number of interventions increased, and hence did not report its results\. Synthetic Data\.We generated synthetic datasets by simulating steady\-state observations from a stable linear stochastic system\. The drift matrix𝚲∈ℝn×n\\bm\{\\Lambda\}\\in\\mathbb\{R\}^\{n\\times n\}was initialized as a zero matrix, then each off\-diagonal entry was set to a Gaussian random value with probabilityρ\\rhoand left at zero otherwise, whereρ\\rhois the desired density level\. For each row, the diagonal entry was set to be larger than the sum of the absolute values of off\-diagonal entries in that row by at least a positive margin, ensuring stability\. The diffusion power𝐃\\mathbf\{D\}was diagonal with strictly positive entries sampled uniformly from\[dmin=0\.2,dmax=0\.4\]\[d\_\{\\min\}=0\.2,d\_\{\\max\}=0\.4\], and the bias vector𝐛\\mathbf\{b\}was drawn uniformly from\[bmin=0\.2,bmax=1\.5\]\[b\_\{\\min\}=0\.2,b\_\{\\max\}=1\.5\]\. The observational steady\-state mean and covariance were then computed by solving the corresponding linear and Lyapunov equations for the given𝚲\\bm\{\\Lambda\},𝐛\\mathbf\{b\}, and𝐃\\mathbf\{D\}\. Interventions were simulated by zeroing all off\-diagonal entries in a selected row of𝚲\\bm\{\\Lambda\}while keeping the diagonal entry unchanged and leaving𝐛\\mathbf\{b\}unchanged\. For the observational setting and each intervention, samples were drawn from the corresponding multivariate normal distribution\. All the implementations are available in the following link:[https://github\.com/sabersalehk/OU\_ID](https://github.com/sabersalehk/OU_ID)\. More details of experiments are given in Appendix[B](https://arxiv.org/html/2609.19955#A2)\. Table 1:Evaluation of DES and PDS on the three Perturb\-seq datasets\. The interventional mean is shaded to indicate oracle access to interventional data\.In Figure[1](https://arxiv.org/html/2609.19955#S5.F1)\(a\-c\), we report the estimation error of𝚲\\bm\{\\Lambda\},𝐃\\mathbf\{D\}, and𝐛\\mathbf\{b\}as a function of the number of interventions withn=10n=10\. To evaluate recovery up to scaling, we compute the scaling factorc=⟨𝐀^,𝐀⟩/⟨𝐀,𝐀⟩c=\\langle\\hat\{\\mathbf\{A\}\},\\mathbf\{A\}\\rangle/\\langle\\mathbf\{A\},\\mathbf\{A\}\\ranglefor each parameter matrix/vector𝐀∈\{𝚲,𝐃,𝐛\}\\mathbf\{A\}\\in\\\{\\bm\{\\Lambda\},\\mathbf\{D\},\\mathbf\{b\}\\\}, where⟨𝐗,𝐘⟩=tr\(𝐗⊤𝐘\)\\langle\\mathbf\{X\},\\mathbf\{Y\}\\rangle=\\operatorname\{tr\}\(\\mathbf\{X\}^\{\\top\}\\mathbf\{Y\}\)for matrices and⟨𝐱,𝐲⟩=𝐱⊤𝐲\\langle\\mathbf\{x\},\\mathbf\{y\}\\rangle=\\mathbf\{x\}^\{\\top\}\\mathbf\{y\}for vectors\. The relative error is then measured as‖𝐀^−c𝐀‖/‖𝐀‖\\\|\\hat\{\\mathbf\{A\}\}\-c\\mathbf\{A\}\\\|/\\\|\\mathbf\{A\}\\\|\. We apply this procedure separately to𝚲\\bm\{\\Lambda\},𝐃\\mathbf\{D\}, and𝐛\\mathbf\{b\}\. The results with 90% confidence intervals are given in blue curves with the legend “Mean and Covariance \(90% CI\)”\. As shown, for𝚲\\bm\{\\Lambda\}and𝐛\\mathbf\{b\}, the relative error decreases as the number of interventions increases\. For𝐃\\mathbf\{D\}, we observe a small non\-monotone effect: the error is initially low with no interventions, increases slightly for a few interventions, and then decreases again\. One possible explanation is that𝐃\\mathbf\{D\}is already estimated accurately from observational data, so incorporating a small number of interventional data \(which may contain fewer samples than the observational data\) can worsen its estimation\. As the number of interventions increases, however, the estimation of𝚲\\bm\{\\Lambda\}improves, which in turn reduces the error in𝐃\\mathbf\{D\}\. We also report results when the loss includes only the residuals of the Lyapunov equations \(i\.e\., using covariances but not means\) with the legend “Covariance \(90% CI\),” similar to\([Dettling et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib17);[Rohbeck et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib16)\)\. This variant performs noticeably worse, highlighting that the mean terms are essential \(please note that there is no curve for𝐛\\mathbf\{b\}in this case as there is no term for estimating it\); mean estimates are typically much more accurate than covariance estimates, and incorporating them substantially improves recovery\. In Figure[1\(d\)](https://arxiv.org/html/2609.19955#S5.F1.sf4), we depict the percentage of instances of𝚲\\bm\{\\Lambda\}satisfying the graphical conditions in Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6)as a function of the number of interventions for different graph densities \(ρ\\rho\)\. For sparse graphs \(ρ=0\.3\\rho=0\.3\), less than 50% of instances satisfy the conditions with a single intervention, but the percentage increases steadily as more interventions are added\. In contrast, denser graphs \(ρ=0\.5,0\.7\\rho=0\.5,0\.7\) already satisfy the graphical conditions at a high rate with only a few interventions\. We also evaluated the SCC\-recovery step from finite samples\. Using the reconstruction procedure implied by Theorem[3\.3](https://arxiv.org/html/2609.19955#S3.Thmtheorem3)and Remark[3\.4](https://arxiv.org/html/2609.19955#S3.Thmtheorem4), we estimated the response sets from empirical mean shifts and recovered the SCC partition under both hard and soft interventions\. We measured recovery quality using the Adjusted Rand Index \(ARI\), where11indicates exact recovery\. As shown in Table[2](https://arxiv.org/html/2609.19955#S6.T2), SCC recovery improves steadily as the number of samples increases\. Hard interventions are more informative in this experiment, but soft interventions also show a clear improvement with sample size\. The details of experiments are given in Appendix[B\.5](https://arxiv.org/html/2609.19955#A2.SS5)\. Table 2:SCC recovery from estimated means\. Recovery is measured by Adjusted Rand Index \(ARI\), where11indicates exact recovery\.Our identifiability theory assumes diagonal diffusion, and this assumption is mainly used for the parameter\-identifiability result in Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6)\. The learning objective, however, can also be applied with a non\-diagonal diffusion matrix\. To evaluate this empirically, we repeated the synthetic experiment in the same setup as Figure[1](https://arxiv.org/html/2609.19955#S5.F1), but allowed the diffusion matrix to be non\-diagonal\. The results in Table[3](https://arxiv.org/html/2609.19955#S6.T3)show the same qualitative trend as in the diagonal\-diffusion setting: the relative errors for𝚲\\bm\{\\Lambda\},𝐃\\mathbf\{D\}, and𝐛\\mathbf\{b\}decrease as the number of interventions increases\. This suggests that, although the current identifiability proof uses diagonal diffusion, the proposed moment\-based estimator remains empirically useful beyond that setting\. The details of experiments are given in Appendix[B\.6](https://arxiv.org/html/2609.19955#A2.SS6)\. Table 3:Parameter recovery with non\-diagonal diffusion\. Relative errors decrease as the number of interventions increases\.We further empirically investigate how conservative the graphical conditions in Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6)are in Appendix[B\.3](https://arxiv.org/html/2609.19955#A2.SS3), and study the impact of sample size on performance in Appendix[B\.4](https://arxiv.org/html/2609.19955#A2.SS4)\. Real Data\.We assess our method on real\-world data, leveraging three published single\-cell perturbation screen datasets \(Frangieh et al\., 2021\)\. Since the true causal graph is unknown in this setting, our evaluation focuses on generalization to unseen perturbations\. Specifically, we consider a Perturb\-seq dataset containing targeted CRISPR knock\-out perturbations of 249 target genes in tumor\-infiltrating lymphocytes \(TILs\) of melanoma patients\. The perturbations were performed under three conditions, which we treat as separate datasets: a baseline culture of TILs in a neutral medium \(“Control”\), a culture of TILs with interferon\-γ\\gammaadded \(“IFN\-γ\\gamma”\), and a co\-culture of TILs with patient\-derived melanoma cells \(“Co\-Culture”\)\. Following the setup in\([Sethuraman et al\., 2023](https://arxiv.org/html/2609.19955#bib.bib21)\), we restrict our analysis to the same subset of 61 genes and adopt their reported training/test split:90%90\\%of interventions are used for training and the remaining10%10\\%are held out for evaluation, with analyses performed separately for each dataset\. Our goal is to predict the interventional mean𝝁\\bm\{\\mu\}for unseen perturbations, using the estimated parameters and the steady\-state equation for the mean\. Note that global scaling is not an issue here, as it cancels out for predicting𝝁\\bm\{\\mu\}\. For evaluation, we do not rely on Mean Absolute Error \(MAE\), as even the observational mean achieves a very close performance to the one using the interventional mean on the held\-out set\. Instead, following the recommendation in the Virtual Cell Challenge555[https://virtualcellchallenge\.org/evaluation\#scoring](https://virtualcellchallenge.org/evaluation#scoring), we report the Differential Expression Score \(DES\) and the Perturbation Differential Score \(PDS\), where higher values indicate better performance\. DES measures agreement in identifying differentially expressed genes, while PDS measures a model’s ability to distinguish between perturbations by ranking predictions according to their similarity to the true perturbational effect, regardless of their effect size\. In Table[1](https://arxiv.org/html/2609.19955#S6.T1), we compare our full method \(“Mean \+ Covariance”\) against the “Covariance” only variant \(similar to the approaches in\([Dettling et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib17);[Rohbeck et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib16)\)\), and against the one using the interventional mean of the held\-out set \(which is not available to our method\)\. As shown in Table[1](https://arxiv.org/html/2609.19955#S6.T1), our method achieves comparable and in some cases even higher scores than the oracle baseline using the interventional mean, despite not having access to held\-out interventional data\. ## 7Conclusions and Future Work We studied recovery of multivariate OU parameters from steady\-state observational and interventional data\. Our main theoretical contribution shows that one intervention per SCC of the drift graph suffices for generic recovery of\(𝚲,𝐛,𝐃\)\(\\bm\{\\Lambda\},\\mathbf\{b\},\\mathbf\{D\}\)up to a single global scaling when the DAG over SCCs is connected with a unique root\. The single\-SCC case \(Theorem[3\.1](https://arxiv.org/html/2609.19955#S3.Thmtheorem1)\) yields a one\-dimensional null space for the stacked moment equations, and the multi\-SCC result \(Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6)\) follows via a constructive, recursive decomposition that leverages mean shifts \(Theorem[3\.3](https://arxiv.org/html/2609.19955#S3.Thmtheorem3)\) to infer SCCs and a topological order\. Building on these guarantees, we considered a regularized least\-squares estimator and observed accurate parameter recovery in synthetic datasets\. While our theoretical results are developed for linear SDEs, we note that this setting is a canonical model class for studying causal inference in continuous\-time systems where causal effects and parameter recovery can be analyzed rigorously\. The identifiability results we obtain, can lead to some practical insights on when stationary data of interventions are sufficient to recover causal structure\. This might be interesting for the wider ML community working on perturbation predictions\. Our guarantees rely on some spectral/ranknondegeneracy assumptions\. Numerical results suggest these hold generically for strongly connected drift graphs\. A key avenue for future work is a formal genericity proof\. Another direction is to understand how diffusion assumptions \(e\.g\., not diagonal diffusion matrices\) change the boundary between identifiable and non\-identifiable regimes\. ## References - M\. Bansal, G\. D\. Gatta, and D\. Di BernardoInference of gene regulatory networks and compound mode of action from time course gene expression profiles\.Bioinformatics22\(7\),pp\. 815–822\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Boegeet al\.\(2025\)T\. Boege, M\. Drton, B\. Hollering, S\. Lumpp, P\. Misra, and D\. SchkodaConditional independence in stationary distributions of diffusions\.Stochastic Processes and their Applications184,pp\. 104604\.Cited by:[§5](https://arxiv.org/html/2609.19955#S5.p5.1)\. - Boeken and Mooij \(2024\)P\. Boeken and J\. M\. MooijDynamic structural causal models\.arXiv preprint arXiv:2406\.01161\.Cited by:[§2](https://arxiv.org/html/2609.19955#S2.p3.2)\. - Caoet al\.\(2019\)J\. Cao, M\. Spielmann, X\. Qiu, X\. Huang, D\. M\. Ibrahim, A\. J\. Hill, F\. Zhang, S\. Mundlos, L\. Christiansen, F\. J\. Steemers,et al\.The single\-cell transcriptional landscape of mammalian organogenesis\.Nature566\(7745\),pp\. 496–502\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Datlingeret al\.\(2017\)P\. Datlinger, A\. F\. Rendeiro, C\. Schmidl, T\. Krausgruber, P\. Traxler, J\. Klughammer, L\. C\. Schuster, A\. Kuchler, D\. Alpar, and C\. BockPooled crispr screening with single\-cell transcriptome readout\.Nature methods14\(3\),pp\. 297–301\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Dettlinget al\.\(2024\)P\. Dettling, M\. Drton, and M\. KolarOn the lasso for graphical continuous lyapunov models\.InCausal Learning and Reasoning,pp\. 514–550\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p2.1),[§5](https://arxiv.org/html/2609.19955#S5.p2.1),[§6](https://arxiv.org/html/2609.19955#S6.SS0.SSS0.Px1.p11.1),[§6](https://arxiv.org/html/2609.19955#S6.SS0.SSS0.Px1.p4.1)\. - Dettlinget al\.\(2023\)P\. Dettling, R\. Homs, C\. Améndola, M\. Drton, and N\. R\. HansenIdentifiability in continuous lyapunov models\.SIAM Journal on Matrix Analysis and Applications44\(4\),pp\. 1799–1821\.Cited by:[§5](https://arxiv.org/html/2609.19955#S5.p2.1)\. - Folland \(2009\)G\. B\. FollandA guide to advanced real analysis\.American Mathematical Soc\.\.Cited by:[§A\.4](https://arxiv.org/html/2609.19955#A1.SS4.p5.1.1)\. - Gardiner \(2009\)C\. GardinerStochastic methods\.Vol\.4,Springer Berlin Heidelberg\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Gonget al\.\(2024\)C\. Gong, C\. Zhang, D\. Yao, J\. Bi, W\. Li, and Y\. XuCausal discovery from temporal data: an overview and new perspectives\.ACM Computing Surveys57\(4\),pp\. 1–38\.Cited by:[footnote 4](https://arxiv.org/html/2609.19955#footnote4)\. - Hamilton \(2020\)J\. D\. HamiltonTime series analysis\.Princeton university press\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Hansen and Sokol \(2014\)N\. Hansen and A\. SokolCausal interpretation of stochastic differential equations\.Cited by:[§2](https://arxiv.org/html/2609.19955#S2.p3.2)\. - Hasanet al\.\(2023\)U\. Hasan, E\. Hossain, and M\. O\. GaniA survey on causal discovery methods for iid and time series data\.Transactions on Machine Learning Research\.Cited by:[footnote 4](https://arxiv.org/html/2609.19955#footnote4)\. - Hespanhaet al\.\(2009\)J\. P\. Hespanhaet al\.Linear systems theory\.Vol\.41,Princeton university press Princeton, NJ\.Cited by:[§C\.2](https://arxiv.org/html/2609.19955#A3.SS2.p2.1.1)\. - Lin \(1974\)C\. LinStructural controllability\.IEEE Transactions on Automatic Control19\(3\),pp\. 201–208\.Cited by:[§C\.2](https://arxiv.org/html/2609.19955#A3.SS2.p3.1.1)\. - Lorchet al\.\(2024\)L\. Lorch, A\. Krause, and B\. SchölkopfCausal modeling with stationary diffusions\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 1927–1935\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p2.1),[§5](https://arxiv.org/html/2609.19955#S5.p3.1),[§6](https://arxiv.org/html/2609.19955#S6.SS0.SSS0.Px1.p1.1)\. - Marbachet al\.\(2012\)D\. Marbach, J\. C\. Costello, R\. Küffner, N\. M\. Vega, R\. J\. Prill, D\. M\. Camacho, K\. R\. Allison, M\. Kellis, J\. J\. Collins,et al\.Wisdom of crowds for robust gene network inference\.Nature methods9\(8\),pp\. 796–804\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Mityagin \(2015\)B\. MityaginThe zero set of a real analytic function\.arXiv preprint arXiv:1512\.07276\.Cited by:[§A\.2](https://arxiv.org/html/2609.19955#A1.SS2.SSS0.Px1.p1.1)\. - Ohet al\.\(2024\)Y\. Oh, D\. Lim, and S\. KimStable neural stochastic differential equations in analyzing irregular time series data\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Paninskiet al\.\(2010\)L\. Paninski, Y\. Ahmadian, D\. G\. Ferreira, S\. Koyama, K\. Rahnama Rad, M\. Vidne, J\. Vogelstein, and W\. WuA new look at state\-space models for neural data\.Journal of computational neuroscience29\(1\),pp\. 107–126\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Pearl \(2009\)J\. PearlCausality\.Cambridge university press\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Peterset al\.\(2017\)J\. Peters, D\. Janzing, and B\. SchölkopfElements of causal inference: foundations and learning algorithms\.MIT press\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Rohbecket al\.\(2024\)M\. Rohbeck, B\. Clarke, K\. Mikulik, A\. Pettet, O\. Stegle, and K\. UeltzhöfferBicycle: intervention\-based causal discovery with cycles\.InCausal Learning and Reasoning,pp\. 209–242\.Cited by:[§B\.2](https://arxiv.org/html/2609.19955#A2.SS2.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.19955#S1.p2.1),[§5](https://arxiv.org/html/2609.19955#S5.p4.1),[§6](https://arxiv.org/html/2609.19955#S6.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2609.19955#S6.SS0.SSS0.Px1.p11.1),[§6](https://arxiv.org/html/2609.19955#S6.SS0.SSS0.Px1.p4.1)\. - Roohaniet al\.\(2022\)Y\. Roohani, K\. Huang, and J\. LeskovecGEARS: predicting transcriptional outcomes of novel multi\-gene perturbations\.BioRxiv,pp\. 2022–07\.Cited by:[§5](https://arxiv.org/html/2609.19955#S5.p8.1)\. - Schiebingeret al\.\(2019\)G\. Schiebinger, J\. Shu, M\. Tabaka, B\. Cleary, V\. Subramanian, A\. Solomon, J\. Gould, S\. Liu, S\. Lin, P\. Berube,et al\.Optimal\-transport analysis of single\-cell gene expression identifies developmental trajectories in reprogramming\.Cell176\(4\),pp\. 928–943\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Sethuramanet al\.\(2023\)M\. G\. Sethuraman, R\. Lopez, R\. Mohan, F\. Fekri, T\. Biancalani, and J\. HütterNodags\-flow: nonlinear cyclic causal structure learning\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 6371–6387\.Cited by:[§5](https://arxiv.org/html/2609.19955#S5.p7.1),[§6](https://arxiv.org/html/2609.19955#S6.SS0.SSS0.Px1.p10.1)\. - Varando and Hansen \(2020\)G\. Varando and N\. R\. HansenGraphical continuous lyapunov models\.InConference on Uncertainty in Artificial Intelligence,pp\. 989–998\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p2.1),[§5](https://arxiv.org/html/2609.19955#S5.p2.1),[§6](https://arxiv.org/html/2609.19955#S6.SS0.SSS0.Px1.p1.1)\. - Villaverdeet al\.\(2016\)A\. F\. Villaverde, A\. Barreiro, and A\. PapachristodoulouStructural identifiability of dynamic systems biology models\.PLoS computational biology12\(10\),pp\. e1005153\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Wanget al\.\(2024\)Y\. Wang, B\. Huang, W\. Huang, X\. Geng, and M\. GongIdentifiability analysis of linear ode systems with hidden confounders\.Advances in Neural Information Processing Systems37,pp\. 59054–59092\.Cited by:[§5](https://arxiv.org/html/2609.19955#S5.p6.1)\. - \[30\]Y\. Wu, E\. Wershof, S\. M\. Schmon, M\. Nassar, B\. Osiński, R\. Eksi, K\. Zhang, and T\. GraepelPerturBench: benchmarking machine learning models for cellular perturbation analysis\.InNeurIPS 2024 Workshop on AI for New Drug Modalities,Cited by:[§5](https://arxiv.org/html/2609.19955#S5.p8.1)\. - Yuet al\.\(2025\)H\. Yu, W\. Qian, Y\. Song, and J\. D\. WelchPerturbnet predicts single\-cell responses to unseen chemical and genetic perturbations\.Molecular Systems Biology,pp\. 1–23\.Cited by:[§5](https://arxiv.org/html/2609.19955#S5.p8.1)\. - Zhanget al\.\(2024\)X\. N\. Zhang, Y\. Pu, Y\. Kawamura, A\. Loza, Y\. Bengio, D\. Shung, and A\. TongTrajectory flow matching with applications to clinical time series modelling\.Advances in Neural Information Processing Systems37,pp\. 107198–107224\.Cited by:[§1](https://arxiv.org/html/2609.19955#S1.p1.1)\. - Zweiget al\.\(2025\)A\. Zweig, Z\. Lin, E\. Azizi, and D\. KnowlesTowards identifiability of interventional stochastic differential equations\.arXiv preprint arXiv:2505\.15987\.Cited by:[§5](https://arxiv.org/html/2609.19955#S5.p2.1.2)\. ## Appendix AProofs ### A\.1Forming the Linear System For completeness, writemt=𝔼\[𝐱t\]m\_\{t\}=\\mathbb\{E\}\[\\mathbf\{x\}\_\{t\}\]andSt=Cov\(𝐱t\)S\_\{t\}=\\operatorname\{Cov\}\(\\mathbf\{x\}\_\{t\}\)\. Taking expectations in the SDE and applying Itô’s product rule to the centered second moment gives m˙t=−𝚲mt\+𝐛,S˙t=−𝚲St−St𝚲⊤\+𝐃\.\\dot\{m\}\_\{t\}=\-\\bm\{\\Lambda\}m\_\{t\}\+\\mathbf\{b\},\\qquad\\dot\{S\}\_\{t\}=\-\\bm\{\\Lambda\}S\_\{t\}\-S\_\{t\}\\bm\{\\Lambda\}^\{\\top\}\+\\mathbf\{D\}\.The last term is the Brownian quadratic\-variation contribution\. At stationarity both derivatives vanish, yielding equation[2](https://arxiv.org/html/2609.19955#S2.E2)and equation[3](https://arxiv.org/html/2609.19955#S2.E3); the same calculation with the intervened drift gives their interventional counterparts\. Throughout,𝐈n\\mathbf\{I\}\_\{n\}denotes then×nn\\times nidentity,⊗\\otimesis the Kronecker product, and𝐞k\\mathbf\{e\}\_\{k\}is thekk\-th canonical basis vector inℝn\\mathbb\{R\}^\{n\}\.Explicitly,A⊗B=\[AjkB\]A\\otimes B=\[A\_\{jk\}B\]and\(A⊙B\)jk=AjkBjk\(A\\odot B\)\_\{jk\}=A\_\{jk\}B\_\{jk\}, where⊙\\odotdenotes the Hadamard \(entrywise\) product\. Column\-wise vectorization satisfiesvec\(AXB\)=\(B⊤⊗A\)vec\(X\)\\operatorname\{vec\}\(AXB\)=\(B^\{\\top\}\\otimes A\)\\operatorname\{vec\}\(X\); this identity produces the linear\-system blocks below\. We define the following selector matrices: - •Introduce the matrix𝐏∈ℝn2×n\\mathbf\{P\}\\in\\mathbb\{R\}^\{\\,n^\{2\}\\times n\}such thatvec\(diag\(𝐝\)\)=𝐏𝐝,\\operatorname\{vec\}\\\!\\bigl\(\\operatorname\{diag\}\(\\mathbf\{d\}\)\\bigr\)=\\mathbf\{P\}\\,\\mathbf\{d\},where thekk\-th column of𝐏\\mathbf\{P\}is given by𝐏:,k=vec\(𝐄kk\),\\mathbf\{P\}\_\{:,k\}=\\operatorname\{vec\}\(\\mathbf\{E\}\_\{kk\}\),and𝐄kk\\mathbf\{E\}\_\{kk\}denotes then×nn\\times nmatrix with a one in entry\(k,k\)\(k,k\)and zeros elsewhere\. - •For a fixed row indexii, let𝐄i:=𝐈n⊗𝐞i⊤∈ℝn×n2\\mathbf\{E\}\_\{i\}:=\\mathbf\{I\}\_\{n\}\\otimes\\mathbf\{e\}\_\{i\}^\{\\\!\\top\}\\in\\mathbb\{R\}^\{\\,n\\times n^\{2\}\}\. Therefore,𝐄ivec\(𝚲\)=\[Λi1,…,Λin\]⊤\\mathbf\{E\}\_\{i\}\\operatorname\{vec\}\(\\bm\{\\Lambda\}\)=\\bigl\[\\Lambda\_\{i1\},\\dots,\\Lambda\_\{in\}\\bigr\]^\{\\\!\\top\}extracts the entireii\-th row of𝚲\\bm\{\\Lambda\}\. - •𝐒n∈ℝn\(n\+1\)2×n2\\mathbf\{S\}\_\{n\}\\in\\mathbb\{R\}^\{\\,\\frac\{n\(n\+1\)\}\{2\}\\times n^\{2\}\}is the upper\-triangular elimination matrix\. - •Let𝐂∈ℝn2×n2\\mathbf\{C\}\\in\\mathbb\{R\}^\{\\,n^\{2\}\\times n^\{2\}\}denote the commutation matrix, so thatvec\(𝐗⊤\)=𝐂vec\(𝐗\)\\operatorname\{vec\}\(\\mathbf\{X\}^\{\\\!\\top\}\)=\\mathbf\{C\}\\,\\operatorname\{vec\}\(\\mathbf\{X\}\)for anyX∈ℝn×nX\\in\\mathbb\{R\}^\{n\\times n\}\. For the equation of observational mean, we define: 𝐌0:=\[−\(𝝁⊤⊗𝐈n\)\|𝐈n\|0\]∈ℝn×p,\\mathbf\{M\}\_\{0\}:=\\Bigl\[\-\\\!\\bigl\(\\bm\{\\mu\}^\{\\\!\\top\}\\\!\\otimes\\mathbf\{I\}\_\{n\}\\bigr\)\\;\\Big\|\\;\\mathbf\{I\}\_\{n\}\\;\\Big\|\\;\\mathbf\{0\}\\Bigr\]\\in\\mathbb\{R\}^\{\\,n\\times p\},wherep=n2\+2np=n^\{2\}\+2n\. For the equation of the interventional mean \(on coordinateii\), let𝐉i:=𝐈n−𝐞i𝐞i⊤\\mathbf\{J\}\_\{i\}:=\\mathbf\{I\}\_\{n\}\-\\mathbf\{e\}\_\{i\}\\mathbf\{e\}\_\{i\}^\{\\\!\\top\}\. We define: 𝐌1:=\[−\(\(𝝁\(i\)\)⊤⊗𝐈n\)\+𝐞i\(𝝁\(i\)\)⊤𝐉i𝐄i\|𝐈n\|0\]∈ℝn×p\.\\mathbf\{M\}\_\{1\}:=\\Bigl\[\-\\\!\\bigl\(\(\\bm\{\\mu\}^\{\(i\)\}\)^\{\\\!\\top\}\\\!\\otimes\\mathbf\{I\}\_\{n\}\\bigr\)\+\\mathbf\{e\}\_\{i\}\(\\bm\{\\mu\}^\{\(i\)\}\)^\{\\\!\\top\}\\mathbf\{J\}\_\{i\}\\mathbf\{E\}\_\{i\}\\;\\Big\|\\;\\mathbf\{I\}\_\{n\}\\;\\Big\|\\;\\mathbf\{0\}\\Bigr\]\\in\\mathbb\{R\}^\{\\,n\\times p\}\. For the observational covariance, vectorising𝚲𝚺\+𝚺𝚲⊤−diag\(𝐝\)=0\\bm\{\\Lambda\}\\bm\{\\Sigma\}\+\\bm\{\\Sigma\}\\bm\{\\Lambda\}^\{\\\!\\top\}\-\\operatorname\{diag\}\(\\mathbf\{d\}\)=0and selecting the upper\-triangular part yields 𝐊0=\[𝐒n\(𝚺⊗𝐈n\+\(𝐈n⊗𝚺\)𝐂\)\|0\|−𝐒n𝐏\]∈ℝn\(n\+1\)2×p\.\\mathbf\{K\}\_\{0\}=\\Bigl\[\\mathbf\{S\}\_\{n\}\\bigl\(\\bm\{\\Sigma\}\\otimes\\mathbf\{I\}\_\{n\}\+\(\\mathbf\{I\}\_\{n\}\\otimes\\bm\{\\Sigma\}\)\\,\\mathbf\{C\}\\bigr\)\\ \\Big\|\\ \\mathbf\{0\}\\ \\Big\|\\ \-\\mathbf\{S\}\_\{n\}\\mathbf\{P\}\\Bigr\]\\in\\mathbb\{R\}^\{\\,\\frac\{n\(n\+1\)\}\{2\}\\times p\}\. For the interventional covariance \(on coordinateii\), we define: 𝐊1=\[𝐒n\(𝚺\(i\)⊗𝐈n\+\(𝐈n⊗𝚺\(i\)\)𝐂−\[\(𝐈n⊗𝐞i\)\+\(𝐞i⊗𝐈n\)\]𝚺\(i\)𝐉i𝐄i\)\|0\|−𝐒n𝐏\]∈ℝn\(n\+1\)2×p\.\\mathbf\{K\}\_\{1\}=\\Bigl\[\\mathbf\{S\}\_\{n\}\\bigl\(\\bm\{\\Sigma\}^\{\(i\)\}\\otimes\\mathbf\{I\}\_\{n\}\+\(\\mathbf\{I\}\_\{n\}\\otimes\\bm\{\\Sigma\}^\{\(i\)\}\)\\,\\mathbf\{C\}\-\\bigl\[\(\\mathbf\{I\}\_\{n\}\\otimes\\mathbf\{e\}\_\{i\}\)\+\(\\mathbf\{e\}\_\{i\}\\otimes\\mathbf\{I\}\_\{n\}\)\\bigr\]\\bm\{\\Sigma\}^\{\(i\)\}\\mathbf\{J\}\_\{i\}\\mathbf\{E\}\_\{i\}\\bigr\)\\ \\Big\|\\ \\mathbf\{0\}\\ \\Big\|\\ \-\\mathbf\{S\}\_\{n\}\\mathbf\{P\}\\Bigr\]\\in\\mathbb\{R\}^\{\\,\\frac\{n\(n\+1\)\}\{2\}\\times p\}\. Assembling all the equations: 𝐀:=\[𝐌0𝐌1𝐊0𝐊1\]∈ℝm×p,\\mathbf\{A\}:=\\begin\{bmatrix\}\\mathbf\{M\}\_\{0\}\\\\\[2\.0pt\] \\mathbf\{M\}\_\{1\}\\\\\[2\.0pt\] \\mathbf\{K\}\_\{0\}\\\\\[2\.0pt\] \\mathbf\{K\}\_\{1\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{\\,m\\times p\},\(13\)wherem=n\+n\+n\(n\+1\)2\+n\(n\+1\)2=n2\+3nm=n\+n\+\\frac\{n\(n\+1\)\}\{2\}\+\\frac\{n\(n\+1\)\}\{2\}=n^\{2\}\+3n\. ### A\.2Notion of Genericity Fix a directed graphGG, and letΘG⊂ℝ\|E\|\+3n\\Theta\_\{G\}\\subset\\mathbb\{R\}^\{\|E\|\+3n\}denote the admissible parameter space consisting of all free parametersθ=\(𝚲,𝐛,𝐝\)\\theta=\(\\bm\{\\Lambda\},\\mathbf\{b\},\\mathbf\{d\}\)compatible withGG, with𝚲\\bm\{\\Lambda\}and the drifts for the performed interventions positive stable, and𝐃=diag\(𝐝\)≻0\\mathbf\{D\}=\\operatorname\{diag\}\(\\mathbf\{d\}\)\\succ 0\.Here, compatibility withGGmeans that the off\-diagonal support of𝚲\\bm\{\\Lambda\}is a subset ofGG\. The setΘG\\Theta\_\{G\}is open inℝ\|E\|\+3n\\mathbb\{R\}^\{\|E\|\+3n\}\. We say that a property holds*generically*onΘG\\Theta\_\{G\}if there exists a Lebesgue\-measure\-zero subsetNG⊂ΘGN\_\{G\}\\subset\\Theta\_\{G\}such that the property holds for everyθ∈ΘG∖NG\\theta\\in\\Theta\_\{G\}\\setminus N\_\{G\}\.HereΘG\\Theta\_\{G\}specifies the family of true parameters over which genericity is asserted\. The class of competing models is specified separately in each theorem\. For instance, competitors are unrestricted in graph support in Theorem[3\.1](https://arxiv.org/html/2609.19955#S3.Thmtheorem1), whereas Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6)explicitly requires compatibility withGG\. #### Rational witnesses\. We repeatedly use the fact that the zero set of a nonzero real polynomial has Lebesgue measure zero\([Mityagin, 2015](https://arxiv.org/html/2609.19955#bib.bib31)\)\. Iff=p/qf=p/qis rational, a single parameter value withq≠0q\\neq 0andf≠0f\\neq 0proves thatppis not the zero polynomial\. We call such a parameter choice a*witness*; it establishes generic nonvanishing whereffis defined\. Matrix inverses and solutions of nonsingular finite linear systems are rational by the adjugate formula\. A full\-column\-rank matrix has at least one nonzero maximal minor, so the same argument gives generic full rank from one such witness\. Finitely many rational functions that are not identically zero can be made nonzero simultaneously\. Indeed, each function has a measure\-zero zero set, and the finite union of these sets still has measure zero\. Outside this union, all the functions are nonzero at the same point\. Their individual witnesses need not coincide\. Each witness only establishes that its corresponding function is not identically zero\. In particular, every nonempty open admissible neighborhood on which the functions are defined contains such a common point, because an open neighborhood has positive Lebesgue measure and therefore cannot be contained in the exceptional union\. ### A\.3Proof of Theorem[3\.3](https://arxiv.org/html/2609.19955#S3.Thmtheorem3) Throughout this proof, genericity is understood with respect to the admissible parameter spaceΘG\\Theta\_\{G\}for any given graphGG, as defined in the genericity convention\. The statement of this theorem only depends on\(𝚲,𝐛\)\(\\bm\{\\Lambda\},\\mathbf\{b\}\), not on the diffusion parameters𝐝\\mathbf\{d\}\. Thus, any Lebesgue\-measure\-zero exceptional set in the free coordinates for\(𝚲,𝐛\)\(\\bm\{\\Lambda\},\\mathbf\{b\}\)induces a Lebesgue\-measure\-zero exceptional set in the full admissible spaceΘG\\Theta\_\{G\}\. Moreover, by the model setup, the intervened matrix𝚲~\(i\)\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}is invertible whenever the interventional mean𝝁\(i\)\\bm\{\\mu\}^\{\(i\)\}is considered\. At steady state,𝝁=𝚲−1𝐛\\bm\{\\mu\}=\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{b\}and𝝁\(i\)=\(𝚲~\(i\)\)−1𝐛\\bm\{\\mu\}^\{\(i\)\}=\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\)^\{\-1\}\\mathbf\{b\}\. Let𝐄:=𝚲~\(i\)−𝚲\\mathbf\{E\}:=\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\-\\bm\{\\Lambda\}; only theii\-th row of𝐄\\mathbf\{E\}is nonzero, withEij=−ΛijE\_\{ij\}=\-\\Lambda\_\{ij\}forj≠ij\\neq i\. By the following equation, Δ𝝁:=𝝁−𝝁\(i\)=𝚲−1𝐄\(𝚲~\(i\)\)−1𝐛=𝚲−1𝐄𝝁\(i\)\.\\Delta\\bm\{\\mu\}:=\\bm\{\\mu\}\-\\bm\{\\mu\}^\{\(i\)\}=\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{E\}\\,\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\)^\{\-1\}\\mathbf\{b\}=\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{E\}\\,\\bm\{\\mu\}^\{\(i\)\}\.\(14\)Since only rowiiof𝐄\\mathbf\{E\}is nonzero,𝐄𝐯=s\(𝐯\)𝐞i\\mathbf\{E\}\\mathbf\{v\}=s\(\\mathbf\{v\}\)\\,\\mathbf\{e\}\_\{i\}for any vector𝐯\\mathbf\{v\}, where s\(𝐯\):=−∑j≠iΛijvj\.s\(\\mathbf\{v\}\):=\-\\sum\_\{j\\neq i\}\\Lambda\_\{ij\}\\,v\_\{j\}\.Applying this to𝐯=𝝁\(i\)\\mathbf\{v\}=\\bm\{\\mu\}^\{\(i\)\}in equation[14](https://arxiv.org/html/2609.19955#A1.E14)gives Δ𝝁=s\(𝝁\(i\)\)𝚲−1𝐞i\.\\Delta\\bm\{\\mu\}=s\\\!\\big\(\\bm\{\\mu\}^\{\(i\)\}\\big\)\\,\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{e\}\_\{i\}\.\(15\) Permute coordinates by some permutation matrix that orders the SCCs topologically\. Let𝐋\\mathbf\{L\}be the block lower–triangular drift matrix after topologically ordering the SCCs, 𝐋=\[𝐋\(1\)0⋯0𝐋\(2,1\)𝐋\(2\)⋱⋱⋱0𝐋\(K,1\)⋯𝐋\(K,K−1\)𝐋\(K\)\],\\mathbf\{L\}=\\begin\{bmatrix\}\\mathbf\{L\}^\{\(1\)\}&0&\\cdots&0\\\\ \\mathbf\{L\}^\{\(2,1\)\}&\\mathbf\{L\}^\{\(2\)\}&\\ddots&\\vdots\\\\ \\vdots&\\ddots&\\ddots&0\\\\ \\mathbf\{L\}^\{\(K,1\)\}&\\cdots&\\mathbf\{L\}^\{\(K,K\-1\)\}&\\mathbf\{L\}^\{\(K\)\}\\end\{bmatrix\},with each diagonal block𝐋\(u\)\\mathbf\{L\}^\{\(u\)\}invertible\. Fix a blockrrand a coordinateiiin blockrr\. LetRRbe the union of the SCCs reachable from the component containingii, including that component\. No allowed edge leads fromRRtoRcR^\{c\}\. OrderingRcR^\{c\}beforeRRtherefore gives a block lower–triangular drift with zero\(Rc,R\)\(R^\{c\},R\)block, and its inverse has the same zero block\. Consequently\(𝐋−1𝐞i\)j=0\(\\mathbf\{L\}^\{\-1\}\\mathbf\{e\}\_\{i\}\)\_\{j\}=0wheneverj∉Rj\\notin R\. It remains to show that the entries for descendant coordinates are generically nonzero\. We show that ifttis a descendant ofrrin the DAG over SCCs andjjis a coordinate in blocktt, then there exists an admissible parameter valueθ⋆∈ΘG\\theta^\{\\star\}\\in\\Theta\_\{G\}such that\(𝐋\(θ⋆\)−1𝐞i\)j≠0\.\{\\color\[rgb\]\{0,0,0\}\\big\(\\mathbf\{L\}\(\\theta^\{\\star\}\)^\{\-1\}\\mathbf\{e\}\_\{i\}\\big\)\_\{j\}\\neq 0\.\}Indeed, by the adjugate formula, this entry can be written as\(𝐋\(θ\)−1𝐞i\)j=pj,i\(θ\)det𝐋\(θ\),\{\\color\[rgb\]\{0,0,0\}\\big\(\\mathbf\{L\}\(\\theta\)^\{\-1\}\\mathbf\{e\}\_\{i\}\\big\)\_\{j\}=\\frac\{p\_\{j,i\}\(\\theta\)\}\{\\det\\mathbf\{L\}\(\\theta\)\},\}wherepj,i\(θ\)p\_\{j,i\}\(\\theta\)is a polynomial in the free entries of𝐋\(θ\)\\mathbf\{L\}\(\\theta\)\. Sincedet𝐋\(θ\)≠0\\det\\mathbf\{L\}\(\\theta\)\\neq 0on the admissible parameter space, showing one admissible point where the entry is nonzero shows thatpj,ip\_\{j,i\}is not identically zero\. Hence the zero set of the entry is contained in the zero set of a nonzero polynomial, and therefore has Lebesgue measure zero\. Sincejjlies in a descendant block of the block containingii, there exists a coordinate\-level directed path inGGfromiitojj, sayi=v0→v1→⋯→vℓ=j\.\{\\color\[rgb\]\{0,0,0\}i=v\_\{0\}\\to v\_\{1\}\\to\\cdots\\to v\_\{\\ell\}=j\.\}If necessary, remove cycles so that the path is simple\. We construct a one\-parameter family𝐋\(ε\)\\mathbf\{L\}\(\\varepsilon\)as follows\. Chooseα\>0\\alpha\>0sufficiently large\. Set all diagonal entries of𝐋\(ε\)\\mathbf\{L\}\(\\varepsilon\)equal toα\\alpha\. For each edgevq−1→vqv\_\{q\-1\}\\to v\_\{q\}on the selected path, set the corresponding entryLvq,vq−1\(ε\)L\_\{v\_\{q\},v\_\{q\-1\}\}\(\\varepsilon\)equal to11\. Set every other allowed off\-diagonal entry of𝐋\(ε\)\\mathbf\{L\}\(\\varepsilon\)equal toε\\varepsilon, and keep all forbidden entries equal to zero\. For everyε\>0\\varepsilon\>0, every allowed off\-diagonal entry is nonzero, so𝐋\(ε\)\\mathbf\{L\}\(\\varepsilon\)has the exact support prescribed byGG\. Moreover, by choosingα\\alphalarge andε\\varepsilonsufficiently small,𝐋\(ε\)\\mathbf\{L\}\(\\varepsilon\)is strictly row\-diagonally dominant with positive diagonal entries, hence positive stable and in particular invertible\. Thus, after choosing arbitrary admissible𝐛\\mathbf\{b\}and any𝐝≻0\\mathbf\{d\}\\succ 0, this construction gives a pointθ\(ε\)∈ΘG\\theta\(\\varepsilon\)\\in\\Theta\_\{G\}for all sufficiently smallε\>0\\varepsilon\>0\. Let𝐏path\\mathbf\{P\}\_\{\\mathrm\{path\}\}be the matrix containing only the selected path entries, i\.e\.,\(𝐏path\)vq,vq−1=1\(\\mathbf\{P\}\_\{\\mathrm\{path\}\}\)\_\{v\_\{q\},v\_\{q\-1\}\}=1forq=1,…,ℓq=1,\\dots,\\elland all other entries zero\. Atε=0\\varepsilon=0,𝐋\(0\)=α𝐈\+𝐏path\.\{\\color\[rgb\]\{0,0,0\}\\mathbf\{L\}\(0\)=\\alpha\\mathbf\{I\}\+\\mathbf\{P\}\_\{\\mathrm\{path\}\}\.\}Since the selected path is simple,𝐏path\\mathbf\{P\}\_\{\\mathrm\{path\}\}is nilpotent,666A square matrix𝐀\\mathbf\{A\}is nilpotent if there exists an integerm≥1m\\geq 1such that𝐀m=𝟎\\mathbf\{A\}^\{m\}=\\mathbf\{0\}\. In this proof,𝐏path\\mathbf\{P\}\_\{\\mathrm\{path\}\}is nilpotent because it only moves forward along the finite simple pathi=v0→v1→⋯→vℓ=ji=v\_\{0\}\\to v\_\{1\}\\to\\cdots\\to v\_\{\\ell\}=j, so𝐏pathℓ\+1=𝟎\\mathbf\{P\}\_\{\\mathrm\{path\}\}^\{\\ell\+1\}=\\mathbf\{0\}\.and \(α𝐈\+𝐏path\)−1=α−1∑q≥0\(−α−1𝐏path\)q,\{\\color\[rgb\]\{0,0,0\}\(\\alpha\\mathbf\{I\}\+\\mathbf\{P\}\_\{\\mathrm\{path\}\}\)^\{\-1\}=\\alpha^\{\-1\}\\sum\_\{q\\geq 0\}\(\-\\alpha^\{\-1\}\\mathbf\{P\}\_\{\\mathrm\{path\}\}\)^\{q\},\}where the sum is finite\. Because𝐏pathℓ𝐞i=𝐞j\\mathbf\{P\}\_\{\\mathrm\{path\}\}^\{\\ell\}\\mathbf\{e\}\_\{i\}=\\mathbf\{e\}\_\{j\}, while no lower power maps𝐞i\\mathbf\{e\}\_\{i\}to𝐞j\\mathbf\{e\}\_\{j\}, we get \(𝐋\(0\)−1𝐞i\)j=\(−1\)ℓα−\(ℓ\+1\)≠0\.\{\\color\[rgb\]\{0,0,0\}\\big\(\\mathbf\{L\}\(0\)^\{\-1\}\\mathbf\{e\}\_\{i\}\\big\)\_\{j\}=\(\-1\)^\{\\ell\}\\alpha^\{\-\(\\ell\+1\)\}\\neq 0\.\}By continuity,\(𝐋\(ε\)−1𝐞i\)j≠0\\big\(\\mathbf\{L\}\(\\varepsilon\)^\{\-1\}\\mathbf\{e\}\_\{i\}\\big\)\_\{j\}\\neq 0for all sufficiently smallε\>0\\varepsilon\>0\. Hence we have exhibited admissible parameters with\(𝐋−1𝐞i\)j≠0\(\\mathbf\{L\}^\{\-1\}\\mathbf\{e\}\_\{i\}\)\_\{j\}\\neq 0\. It follows that, for each fixed descendant coordinatej∈Desc\(C\)j\\in\\mathrm\{Desc\}\(C\), the exceptional set on which\(𝐋−1𝐞i\)j=0\(\\mathbf\{L\}^\{\-1\}\\mathbf\{e\}\_\{i\}\)\_\{j\}=0has Lebesgue measure zero inΘG\\Theta\_\{G\}\. Since there are only finitely many coordinates, taking the union over all descendant coordinates still gives a Lebesgue\-measure\-zero exceptional set\. Therefore, generically, \(𝐋−1𝐞i\)j≠0for everyj∈Desc\(C\),\{\\color\[rgb\]\{0,0,0\}\\big\(\\mathbf\{L\}^\{\-1\}\\mathbf\{e\}\_\{i\}\\big\)\_\{j\}\\neq 0\\quad\\text\{for every \}j\\in\\mathrm\{Desc\}\(C\),\}whereDesc\(C\)\\mathrm\{Desc\}\(C\)includes the componentCCitself\. On the other hand, as shown above, forj∉Desc\(C\)j\\notin\\mathrm\{Desc\}\(C\), the descendant\-block argument gives\(𝐋−1𝐞i\)j=0\\big\(\\mathbf\{L\}^\{\-1\}\\mathbf\{e\}\_\{i\}\\big\)\_\{j\}=0deterministically\. If for somej≠ij\\neq i, we haveΛij≠0\\Lambda\_\{ij\}\\neq 0, thens\(𝝁\(i\)\)=−∑j≠iΛijμj\(i\)\.s\\\!\\big\(\\bm\{\\mu\}^\{\(i\)\}\\big\)=\-\\sum\_\{j\\neq i\}\\Lambda\_\{ij\}\\,\\mu^\{\(i\)\}\_\{j\}\.This is an analytic function of\(𝚲,𝐛\)\(\\bm\{\\Lambda\},\\mathbf\{b\}\)\. Moreover, if theii\-th row is non\-null, then the row vector𝐫i⊤:=−∑j≠iΛij𝐞j⊤\{\\color\[rgb\]\{0,0,0\}\\mathbf\{r\}\_\{i\}^\{\\top\}:=\-\\sum\_\{j\\neq i\}\\Lambda\_\{ij\}\\mathbf\{e\}\_\{j\}^\{\\top\}\}is nonzero\. Sinces\(𝝁\(i\)\)=𝐫i⊤\(𝚲~\(i\)\)−1𝐛,\{\\color\[rgb\]\{0,0,0\}s\\\!\\big\(\\bm\{\\mu\}^\{\(i\)\}\\big\)=\\mathbf\{r\}\_\{i\}^\{\\top\}\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\)^\{\-1\}\\mathbf\{b\},\}and\(𝚲~\(i\)\)−1\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\)^\{\-1\}is invertible, the row vector𝐫i⊤\(𝚲~\(i\)\)−1\\mathbf\{r\}\_\{i\}^\{\\top\}\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\)^\{\-1\}is nonzero\. Thus, for fixed𝚲\\bm\{\\Lambda\},s\(𝝁\(i\)\)s\(\\bm\{\\mu\}^\{\(i\)\}\)is a nonzero linear function of𝐛\\mathbf\{b\}\. Hence it is not identically zero as a function of\(𝚲,𝐛\)\(\\bm\{\\Lambda\},\\mathbf\{b\}\), and its zero set has Lebesgue measure zero in the admissible free\-coordinate space\. Therefores\(𝝁\(i\)\)≠0s\(\\bm\{\\mu\}^\{\(i\)\}\)\\neq 0generically\. Combining with equation[15](https://arxiv.org/html/2609.19955#A1.E15), the support ofΔ𝝁\\Delta\\bm\{\\mu\}matches that of𝚲−1𝐞i\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{e\}\_\{i\}whenevers\(𝝁\(i\)\)≠0s\(\\bm\{\\mu\}^\{\(i\)\}\)\\neq 0\. Thus, if for somej≠ij\\neq i,Λij≠0\\Lambda\_\{ij\}\\neq 0, generically, μk\(i\)≠μk⟺k∈Desc\(C\),\{\\color\[rgb\]\{0,0,0\}\\mu^\{\(i\)\}\_\{k\}\\neq\\mu\_\{k\}\\ \\Longleftrightarrow\\ k\\in\\mathrm\{Desc\}\(C\),\}whereDesc\(C\)\\mathrm\{Desc\}\(C\)includesCCitself\. IfΛij=0\\Lambda\_\{ij\}=0for allj≠ij\\neq i, then𝐄=𝟎\\mathbf\{E\}=\\mathbf\{0\}, sos\(𝝁\(i\)\)=0s\(\\bm\{\\mu\}^\{\(i\)\}\)=0andΔ𝝁=𝟎\\Delta\\bm\{\\mu\}=\\mathbf\{0\}\. Hence𝝁\(i\)=𝝁\\bm\{\\mu\}^\{\(i\)\}=\\bm\{\\mu\}\. ### A\.4Proof of Theorem[3\.3](https://arxiv.org/html/2609.19955#S3.Thmtheorem3)under a soft row intervention We only indicate the changes relative to the proof of Theorem[3\.3](https://arxiv.org/html/2609.19955#S3.Thmtheorem3)for the hard intervention\. The argument showing the generic support of𝚲−1𝐞i\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{e\}\_\{i\}is unchanged\. At steady state,𝝁=𝚲−1𝐛,𝝁\(i\)=\(𝚲~\(i\)\)−1𝐛\\bm\{\\mu\}=\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{b\},\\bm\{\\mu\}^\{\(i\)\}=\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\)^\{\-1\}\\mathbf\{b\}\. For a soft row intervention,𝚲~\(i\)\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}differs from𝚲\\bm\{\\Lambda\}only in rowii\. The entries in rowii, including the diagonal entry, may change\. Let𝐄:=𝚲~\(i\)−𝚲\\mathbf\{E\}:=\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\-\\bm\{\\Lambda\}\. Then only rowiiof𝐄\\mathbf\{E\}is nonzero\. Denote this row change by𝜹i⊤:=𝐞i⊤𝐄=𝐞i⊤\(𝚲~\(i\)−𝚲\)\\bm\{\\delta\}\_\{i\}^\{\\top\}:=\\mathbf\{e\}\_\{i\}^\{\\top\}\\mathbf\{E\}=\\mathbf\{e\}\_\{i\}^\{\\top\}\\big\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\-\\bm\{\\Lambda\}\\big\)\. We assume the intervention is non\-null, i\.e\.,𝜹i≠0\\bm\{\\delta\}\_\{i\}\\neq 0\.The row change is fixed independently of𝐛\\mathbf\{b\}\. As before, Δ𝝁:=𝝁−𝝁\(i\)=𝚲−1𝐄\(𝚲~\(i\)\)−1𝐛=𝚲−1𝐄𝝁\(i\)\.\\Delta\\bm\{\\mu\}:=\\bm\{\\mu\}\-\\bm\{\\mu\}^\{\(i\)\}=\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{E\}\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\)^\{\-1\}\\mathbf\{b\}=\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{E\}\\bm\{\\mu\}^\{\(i\)\}\.\(16\)Since only rowiiof𝐄\\mathbf\{E\}is nonzero, for any vector𝐯\\mathbf\{v\},𝐄𝐯=η\(𝐯\)𝐞i\\mathbf\{E\}\\mathbf\{v\}=\\eta\(\\mathbf\{v\}\)\\mathbf\{e\}\_\{i\}, where η\(𝐯\):=𝜹i⊤𝐯=∑j=1n\(Λ~ij\(i\)−Λij\)vj\.\\eta\(\\mathbf\{v\}\):=\\bm\{\\delta\}\_\{i\}^\{\\top\}\\mathbf\{v\}=\\sum\_\{j=1\}^\{n\}\\left\(\\widetilde\{\\Lambda\}^\{\(i\)\}\_\{ij\}\-\\Lambda\_\{ij\}\\right\)v\_\{j\}\.Applying this to𝐯=𝝁\(i\)\\mathbf\{v\}=\\bm\{\\mu\}^\{\(i\)\}in equation[16](https://arxiv.org/html/2609.19955#A1.E16)gives Δ𝝁=η\(𝝁\(i\)\)𝚲−1𝐞i\.\\Delta\\bm\{\\mu\}=\\eta\(\\bm\{\\mu\}^\{\(i\)\}\)\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{e\}\_\{i\}\.\(17\) The rest of the proof uses the same generic\-support claim proved in the hard intervention case: for the SCCCCcontainingii, \(𝚲−1𝐞i\)k=0deterministically fork∉Desc\(C\),\(\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{e\}\_\{i\}\)\_\{k\}=0\\quad\\text\{deterministically for \}k\\notin\\mathrm\{Desc\}\(C\),and \(𝚲−1𝐞i\)k≠0generically for everyk∈Desc\(C\),\(\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{e\}\_\{i\}\)\_\{k\}\\neq 0\\quad\\text\{generically for every \}k\\in\\mathrm\{Desc\}\(C\),whereDesc\(C\)\\mathrm\{Desc\}\(C\)includesCCitself\. It remains only to check that the scalar factor in equation[17](https://arxiv.org/html/2609.19955#A1.E17)is generically nonzero\. We haveη\(𝝁\(i\)\)=𝜹i⊤\(𝚲~\(i\)\)−1𝐛\\eta\(\\bm\{\\mu\}^\{\(i\)\}\)=\\bm\{\\delta\}\_\{i\}^\{\\top\}\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\)^\{\-1\}\\mathbf\{b\}\. Since𝜹i≠0\\bm\{\\delta\}\_\{i\}\\neq 0and𝚲~\(i\)\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}is invertible, the row vector𝜹i⊤\(𝚲~\(i\)\)−1\\bm\{\\delta\}\_\{i\}^\{\\top\}\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\)^\{\-1\}is nonzero\. Hence, for fixed\(𝚲,𝚲~\(i\)\)\(\\bm\{\\Lambda\},\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\),η\(𝝁\(i\)\)\\eta\(\\bm\{\\mu\}^\{\(i\)\}\)is a nonzero linear function of𝐛\\mathbf\{b\}\. Therefore its zero set is a Lebesgue\-measure\-zero exceptional set, andη\(𝝁\(i\)\)≠0\\eta\(\\bm\{\\mu\}^\{\(i\)\}\)\\neq 0generically\. The exceptional set is measurable, and its section in𝐛\\mathbf\{b\}at every fixed admissible drift/perturbation pair is a hyperplane of measure zero\. Fubini\-Tonelli theorem\([Folland, 2009](https://arxiv.org/html/2609.19955#bib.bib32)\)\(Theorem 2\.16 \(b\)\), applied to the indicator of this set, therefore makes it jointly null in the drift, perturbation, and input parameters\. Combining this with equation[17](https://arxiv.org/html/2609.19955#A1.E17), the support ofΔ𝝁\\Delta\\bm\{\\mu\}generically agrees with the support of𝚲−1𝐞i\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{e\}\_\{i\}\. Therefore, for a non\-null soft row intervention on coordinateii, μk\(i\)≠μk⟺k∈Desc\(C\)\\mu\_\{k\}^\{\(i\)\}\\neq\\mu\_\{k\}\\quad\\Longleftrightarrow\\quad k\\in\\mathrm\{Desc\}\(C\)generically, whereDesc\(C\)\\mathrm\{Desc\}\(C\)includesCCitself\. If the soft intervention is null, i\.e\.,𝜹i=𝟎\\bm\{\\delta\}\_\{i\}=\\mathbf\{0\}, then𝐄=𝟎\\mathbf\{E\}=\\mathbf\{0\}, soΔ𝝁=𝟎\\Delta\\bm\{\\mu\}=\\mathbf\{0\}and𝝁\(i\)=𝝁\\bm\{\\mu\}^\{\(i\)\}=\\bm\{\\mu\}\. ### A\.5Recovering SCCs and a topological order over SCCs from mean changes Assume at least one intervention is performed in every SCC\. For each intervention onii, set Ri:=\{j:μj\(i\)≠μj\}\.R\_\{i\}:=\\\{\\,j:\\ \\mu^\{\(i\)\}\_\{j\}\\neq\\mu\_\{j\}\\,\\\}\.By Theorem[3\.3](https://arxiv.org/html/2609.19955#S3.Thmtheorem3), generically, Ri=Desc\(Ci\)if the intervention oniis non\-null,Ri=∅if it is null,R\_\{i\}=\\mathrm\{Desc\}\(C\_\{i\}\)\\quad\\text\{if the intervention on $i$ is non\-null\},\\qquad R\_\{i\}=\\emptyset\\quad\\text\{if it is null\},whereCiC\_\{i\}is the SCC containingiiandDesc\(C\)\\mathrm\{Desc\}\(C\)is the set of nodes lying in SCCs reachable fromCCin the SCC–DAG, includingCCitself\. Procedure\. 1. 1\.*Singleton sources \(null interventions\)\.*IfRi=∅R\_\{i\}=\\emptyset, declare\{i\}\\\{i\\\}a singleton SCC with no incoming edges, i\.e\., a source SCC\. Do*not*merge differentiiwithRi=∅R\_\{i\}=\\emptyset\. 2. 2\.*Group non\-null response sets\.*For the remaining interventions, group by equality of response sets:i∼i′i\\sim i^\{\\prime\}iffRi=Ri′R\_\{i\}=R\_\{i^\{\\prime\}\}\. Each distinct nonempty response set corresponds to exactly one non\-null SCC, but the equivalence class of intervention targets is not necessarily the whole SCC\. We therefore use the distinct nonempty response sets as temporary SCC labels\. 3. 3\.*Recover the node set of each non\-null SCC\.*Letℛ\\mathcal\{R\}be the collection of distinct nonempty response sets obtained in step 2\. For eachR∈ℛR\\in\\mathcal\{R\}, defineC\(R\):=R∖⋃R′∈ℛR′⊊RR′C\(R\):=R\\setminus\\bigcup\_\{\\begin\{subarray\}\{c\}R^\{\\prime\}\\in\\mathcal\{R\}\\\\ R^\{\\prime\}\\subsetneq R\\end\{subarray\}\}R^\{\\prime\}\. ThenC\(R\)C\(R\)is exactly the SCC whose descendant set isRR\. 4. 4\.*Reachability order among non\-null SCCs\.*For two recovered non\-null SCCsC\(R\)C\(R\)andC\(R′\)C\(R^\{\\prime\}\),C\(R\)reachesC\(R′\)⟺R⊇R′C\(R\)\\text\{ reaches \}C\(R^\{\\prime\}\)\\Longleftrightarrow R\\supseteq R^\{\\prime\}\. Thus strict reverse inclusion of response sets recovers the SCC\-level reachability partial order\. If desired, one may draw the cover relation by adding an edge fromC\(R\)C\(R\)toC\(R′\)C\(R^\{\\prime\}\)whenR⊋R′R\\supsetneq R^\{\\prime\}and there is noR′′∈ℛR^\{\\prime\\prime\}\\in\\mathcal\{R\}such thatR⊋R′′⊋R′R\\supsetneq R^\{\\prime\\prime\}\\supsetneq R^\{\\prime\}\.Among these non\-null SCCs, reachability is a partial order\. The full condensation DAG may also have direct edges whose reachability is already implied by a longer path\. The procedure recovers the SCCs and a compatible topological order, without identifying every direct edge between them\. 5. 5\.*Topological order\.*Output all singleton sources from step 1 first, in any arbitrary order, and then output the recovered non\-null SCCs in any order consistent with reverse inclusion of their response sets, i\.e\., larger response sets before smaller response sets\. This yields a valid topological ordering of the SCCs\. ### A\.6Proof of Theorem[3\.1](https://arxiv.org/html/2609.19955#S3.Thmtheorem1) Throughout this proof, genericity is understood with respect to the admissible parameter spaceΘG\\Theta\_\{G\}for any fixed graphGG, as defined in the genericity convention in Appendix[A\.2](https://arxiv.org/html/2609.19955#A1.SS2)\. Consider the following two admissible triples satisfying the linear system equation[13](https://arxiv.org/html/2609.19955#A1.E13): - •True underlying triple:\(𝚲⋆,𝐛⋆,𝐃⋆\)\(\\bm\{\\Lambda\}\_\{\\star\},\\,\\mathbf\{b\}\_\{\\star\},\\,\\mathbf\{D\}\_\{\\star\}\)\. - •Alternative admissible triple:\(𝚲′,𝐛′,𝐃′\)\(\\bm\{\\Lambda\}^\{\\prime\},\\,\\mathbf\{b\}^\{\\prime\},\\,\\mathbf\{D\}^\{\\prime\}\)\. Define Δ𝚯:=\[vec\(Δ𝚲\)Δ𝐛Δ𝐝\],Δ𝚲:=𝚲′−𝚲⋆,Δ𝐛:=𝐛′−𝐛⋆,Δ𝐃:=𝐃′−𝐃⋆,\\Delta\\bm\{\\Theta\}:=\\begin\{bmatrix\}\\operatorname\{vec\}\(\\Delta\\bm\{\\Lambda\}\)\\\\ \\Delta\\mathbf\{b\}\\\\ \\Delta\\mathbf\{d\}\\end\{bmatrix\},\\qquad\\Delta\\bm\{\\Lambda\}:=\\bm\{\\Lambda\}^\{\\prime\}\-\\bm\{\\Lambda\}\_\{\\star\},\\quad\\Delta\\mathbf\{b\}:=\\mathbf\{b\}^\{\\prime\}\-\\mathbf\{b\}\_\{\\star\},\\quad\\Delta\\mathbf\{D\}:=\\mathbf\{D\}^\{\\prime\}\-\\mathbf\{D\}\_\{\\star\},whereΔ𝐝\\Delta\\mathbf\{d\}is the diagonal ofΔ𝐃\\Delta\\mathbf\{D\}\. Since𝐀𝚯⋆=0\\mathbf\{A\}\\bm\{\\Theta\}\_\{\\star\}=0and𝐀𝚯′=0\\mathbf\{A\}\\bm\{\\Theta\}^\{\\prime\}=0, we have𝐀Δ𝚯=0\\mathbf\{A\}\\Delta\\bm\{\\Theta\}=0\.Our goal is to show that the alternative triple is a scalar multiple of the true triple, i\.e\.,𝚲′=c𝚲⋆,𝐛′=c𝐛⋆,𝐃′=c𝐃⋆\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}^\{\\prime\}=c\\bm\{\\Lambda\}\_\{\\star\},\\mathbf\{b\}^\{\\prime\}=c\\mathbf\{b\}\_\{\\star\},\\mathbf\{D\}^\{\\prime\}=c\\mathbf\{D\}\_\{\\star\}\}for some scalarc\>0c\>0\. Fix the intervened row indexii\. The observational blocks𝐌0\\mathbf\{M\}\_\{0\}and𝐊0\\mathbf\{K\}\_\{0\}yield Δ𝚲𝝁=Δ𝐛,Δ𝚲𝚺\+𝚺Δ𝚲⊤=Δ𝐃,\\Delta\\bm\{\\Lambda\}\\bm\{\\mu\}=\\Delta\\mathbf\{b\},\\qquad\\Delta\\bm\{\\Lambda\}\\bm\{\\Sigma\}\+\\bm\{\\Sigma\}\\Delta\\bm\{\\Lambda\}^\{\\top\}=\\Delta\\mathbf\{D\},\(18\)where𝝁\\bm\{\\mu\}and𝚺\\bm\{\\Sigma\}are the true observational mean and covariance\. Since the intervention zeros out rowiioff\-diagonals, 𝚲~⋆\(i\)=𝐉i⊙𝚲⋆,𝚲~′\(i\)=𝐉i⊙𝚲′,𝐉i:=𝟏𝟏⊤−𝐞i𝟏⊤\+𝐞i𝐞i⊤\.\\widetilde\{\\bm\{\\Lambda\}\}\_\{\\star\}^\{\(i\)\}=\\mathbf\{J\}\_\{i\}\\odot\\bm\{\\Lambda\}\_\{\\star\},\\qquad\\widetilde\{\\bm\{\\Lambda\}\}^\{\\prime\(i\)\}=\\mathbf\{J\}\_\{i\}\\odot\\bm\{\\Lambda\}^\{\\prime\},\\qquad\\mathbf\{J\}\_\{i\}:=\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\-\\mathbf\{e\}\_\{i\}\\mathbf\{1\}^\{\\top\}\+\\mathbf\{e\}\_\{i\}\\mathbf\{e\}\_\{i\}^\{\\top\}\.Thus the interventional blocks𝐌1\\mathbf\{M\}\_\{1\}and𝐊1\\mathbf\{K\}\_\{1\}give \(𝐉i⊙Δ𝚲\)𝝁\(i\)=Δ𝐛,\(𝐉i⊙Δ𝚲\)𝚺\(i\)\+𝚺\(i\)\(Δ𝚲⊤⊙𝐉i⊤\)=Δ𝐃\.\(\\mathbf\{J\}\_\{i\}\\odot\\Delta\\bm\{\\Lambda\}\)\\bm\{\\mu\}^\{\(i\)\}=\\Delta\\mathbf\{b\},\\qquad\(\\mathbf\{J\}\_\{i\}\\odot\\Delta\\bm\{\\Lambda\}\)\\bm\{\\Sigma\}^\{\(i\)\}\+\\bm\{\\Sigma\}^\{\(i\)\}\(\\Delta\\bm\{\\Lambda\}^\{\\top\}\\odot\\mathbf\{J\}\_\{i\}^\{\\top\}\)=\\Delta\\mathbf\{D\}\.\(19\)Taking theii\-th row of the first equation gives \(Δ𝚲\)iiμi\(i\)=Δbi\.\(\\Delta\\bm\{\\Lambda\}\)\_\{ii\}\\mu\_\{i\}^\{\(i\)\}=\\Delta b\_\{i\}\.\(20\) Outside the measure\-zero exceptional set whereb⋆i=0b\_\{\\star i\}=0, definec:=bi′b⋆i\.\{\\color\[rgb\]\{0,0,0\}c:=\\frac\{b\_\{i\}^\{\\prime\}\}\{b\_\{\\star i\}\}\.\}Define the scaled\-difference variables 𝚲Δ:=𝚲′−c𝚲⋆,𝐛Δ:=𝐛′−c𝐛⋆,𝐃Δ:=𝐃′−c𝐃⋆\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\Delta\}:=\\bm\{\\Lambda\}^\{\\prime\}\-c\\bm\{\\Lambda\}\_\{\\star\},\\qquad\\mathbf\{b\}\_\{\\Delta\}:=\\mathbf\{b\}^\{\\prime\}\-c\\mathbf\{b\}\_\{\\star\},\\qquad\\mathbf\{D\}\_\{\\Delta\}:=\\mathbf\{D\}^\{\\prime\}\-c\\mathbf\{D\}\_\{\\star\}\.\}Then\(𝚲Δ,𝐛Δ,𝐃Δ\)\(\\bm\{\\Lambda\}\_\{\\Delta\},\\mathbf\{b\}\_\{\\Delta\},\\mathbf\{D\}\_\{\\Delta\}\)still satisfies the homogeneous linear system, and\(bΔ\)i=0\(b\_\{\\Delta\}\)\_\{i\}=0\. In what follows, all moments𝝁,𝚺,𝝁\(i\),𝚺\(i\)\\bm\{\\mu\},\\bm\{\\Sigma\},\\bm\{\\mu\}^\{\(i\)\},\\bm\{\\Sigma\}^\{\(i\)\}remain the true moments generated by\(𝚲⋆,𝐛⋆,𝐃⋆\)\(\\bm\{\\Lambda\}\_\{\\star\},\\mathbf\{b\}\_\{\\star\},\\mathbf\{D\}\_\{\\star\}\)\. For the true intervened system, theii\-th row givesΛ⋆,iiμi\(i\)=b⋆i\.\{\\color\[rgb\]\{0,0,0\}\\Lambda\_\{\\star,ii\}\\mu\_\{i\}^\{\(i\)\}=b\_\{\\star i\}\.\}Sinceb⋆i≠0b\_\{\\star i\}\\neq 0, we haveμi\(i\)≠0\\mu\_\{i\}^\{\(i\)\}\\neq 0\. Applying equation[20](https://arxiv.org/html/2609.19955#A1.E20)to the ambiguity triple and using\(bΔ\)i=0\(b\_\{\\Delta\}\)\_\{i\}=0, we get\(𝚲Δ\)ii=0\.\{\\color\[rgb\]\{0,0,0\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{ii\}=0\.\}Moreover, the\(i,i\)\(i,i\)entry of the interventional covariance equation gives2Σii\(i\)\(𝚲Δ\)ii=\(𝐃Δ\)ii2\\Sigma^\{\(i\)\}\_\{ii\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{ii\}=\(\\mathbf\{D\}\_\{\\Delta\}\)\_\{ii\}, and hence\(𝐃Δ\)ii=0\.\{\\color\[rgb\]\{0,0,0\}\(\\mathbf\{D\}\_\{\\Delta\}\)\_\{ii\}=0\.\} For the ambiguity triple, write 𝚲~Δ\(i\)=𝐉i⊙𝚲Δ=\[00\(𝚲Δ\)−i,i\(𝚲Δ\)−i,−i\]\.\\widetilde\{\\bm\{\\Lambda\}\}\_\{\\Delta\}^\{\(i\)\}=\\mathbf\{J\}\_\{i\}\\odot\\bm\{\\Lambda\}\_\{\\Delta\}=\\begin\{bmatrix\}0&0\\\\ \(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,i\}&\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}\\end\{bmatrix\}\.Using the\(−i,i\)\(\-i,i\)block of the interventional covariance equation and the fact that𝐃Δ\\mathbf\{D\}\_\{\\Delta\}is diagonal, we obtain \(𝚲Δ\)−i,i𝚺ii\(i\)\+\(𝚲Δ\)−i,−i𝚺−i,i\(i\)=0\.\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,i\}\\bm\{\\Sigma\}^\{\(i\)\}\_\{ii\}\+\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}\\bm\{\\Sigma\}^\{\(i\)\}\_\{\-i,i\}=0\.Thus \(𝚲Δ\)−i,i=−\(𝚲Δ\)−i,−i𝚺−i,i\(i\)𝚺ii\(i\)\.\{\\color\[rgb\]\{0,0,0\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,i\}=\-\\frac\{\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}\\bm\{\\Sigma\}^\{\(i\)\}\_\{\-i,i\}\}\{\\bm\{\\Sigma\}^\{\(i\)\}\_\{ii\}\}\.\}\(21\)The\(−i,−i\)\(\-i,\-i\)block then gives \(𝚲Δ\)−i,−i𝐙\+𝐙\(𝚲Δ\)−i,−i⊤=\(𝐃Δ\)−i,−i,\{\\color\[rgb\]\{0,0,0\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}\\mathbf\{Z\}\+\\mathbf\{Z\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}^\{\\\!\\top\}=\(\\mathbf\{D\}\_\{\\Delta\}\)\_\{\-i,\-i\},\}\(22\)where 𝐙:=𝚺−i,−i\(i\)−𝚺−i,i\(i\)𝚺i,−i\(i\)𝚺ii\(i\)\.\\mathbf\{Z\}:=\\bm\{\\Sigma\}^\{\(i\)\}\_\{\-i,\-i\}\-\\frac\{\\bm\{\\Sigma\}^\{\(i\)\}\_\{\-i,i\}\\bm\{\\Sigma\}^\{\(i\)\}\_\{i,\-i\}\}\{\\bm\{\\Sigma\}^\{\(i\)\}\_\{ii\}\}\.The matrix𝐙\\mathbf\{Z\}is positive definite, since it is the Schur complement of𝚺ii\(i\)\\bm\{\\Sigma\}^\{\(i\)\}\_\{ii\}in𝚺\(i\)≻0\\bm\{\\Sigma\}^\{\(i\)\}\\succ 0\. Next, the observational and interventional covariance equations for the ambiguity variables are 𝚲Δ𝚺\+𝚺𝚲Δ⊤\\displaystyle\\bm\{\\Lambda\}\_\{\\Delta\}\\bm\{\\Sigma\}\+\\bm\{\\Sigma\}\\bm\{\\Lambda\}\_\{\\Delta\}^\{\\top\}=𝐃Δ,\\displaystyle=\\mathbf\{D\}\_\{\\Delta\},\(23\)𝚲~Δ\(i\)𝚺\(i\)\+𝚺\(i\)𝚲~Δ\(i\)⊤\\displaystyle\\widetilde\{\\bm\{\\Lambda\}\}\_\{\\Delta\}^\{\(i\)\}\\bm\{\\Sigma\}^\{\(i\)\}\+\\bm\{\\Sigma\}^\{\(i\)\}\\widetilde\{\\bm\{\\Lambda\}\}\_\{\\Delta\}^\{\(i\)\\top\}=𝐃Δ\.\\displaystyle=\\mathbf\{D\}\_\{\\Delta\}\.\(24\)Define𝐰Δ∈ℝn\\mathbf\{w\}\_\{\\Delta\}\\in\\mathbb\{R\}^\{n\}by\(wΔ\)i=0\(w\_\{\\Delta\}\)\_\{i\}=0and\(𝐰Δ\)−i=\(𝚲Δ\)i,−i⊤\(\\mathbf\{w\}\_\{\\Delta\}\)\_\{\-i\}=\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{i,\-i\}^\{\\top\}, so that𝚲~Δ\(i\)=𝚲Δ−𝐞i𝐰Δ⊤\.\\widetilde\{\\bm\{\\Lambda\}\}\_\{\\Delta\}^\{\(i\)\}=\\bm\{\\Lambda\}\_\{\\Delta\}\-\\mathbf\{e\}\_\{i\}\\mathbf\{w\}\_\{\\Delta\}^\{\\top\}\.Subtracting equation[24](https://arxiv.org/html/2609.19955#A1.E24)from equation[23](https://arxiv.org/html/2609.19955#A1.E23), with𝚪:=𝚺−𝚺\(i\)\\bm\{\\Gamma\}:=\\bm\{\\Sigma\}\-\\bm\{\\Sigma\}^\{\(i\)\}, gives 𝚲Δ𝚪\+𝚪𝚲Δ⊤\+𝐞i𝐰Δ⊤𝚺\(i\)\+𝚺\(i\)𝐰Δ𝐞i⊤=0\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\Delta\}\\bm\{\\Gamma\}\+\\bm\{\\Gamma\}\\bm\{\\Lambda\}\_\{\\Delta\}^\{\\top\}\+\\mathbf\{e\}\_\{i\}\\mathbf\{w\}\_\{\\Delta\}^\{\\top\}\\bm\{\\Sigma\}^\{\(i\)\}\+\\bm\{\\Sigma\}^\{\(i\)\}\\mathbf\{w\}\_\{\\Delta\}\\mathbf\{e\}\_\{i\}^\{\\top\}=0\.\}\(25\)Substituting𝚲Δ=𝚲~Δ\(i\)\+𝐞i𝐰Δ⊤\\bm\{\\Lambda\}\_\{\\Delta\}=\\widetilde\{\\bm\{\\Lambda\}\}\_\{\\Delta\}^\{\(i\)\}\+\\mathbf\{e\}\_\{i\}\\mathbf\{w\}\_\{\\Delta\}^\{\\top\}into equation[25](https://arxiv.org/html/2609.19955#A1.E25), and then taking the\(−i,−i\)\(\-i,\-i\)block, removes all terms containing𝐞i\\mathbf\{e\}\_\{i\}and gives \(𝚲~Δ\(i\)𝚪\+𝚪𝚲~Δ\(i\)⊤\)−i,−i=0\.\\left\(\\widetilde\{\\bm\{\\Lambda\}\}\_\{\\Delta\}^\{\(i\)\}\\bm\{\\Gamma\}\+\\bm\{\\Gamma\}\\widetilde\{\\bm\{\\Lambda\}\}\_\{\\Delta\}^\{\(i\)\\top\}\\right\)\_\{\-i,\-i\}=0\.\(26\)Expanding this block, \(𝚲Δ\)−i,i𝚪i,−i\+\(𝚲Δ\)−i,−i𝚪−i,−i\+𝚪−i,i\(𝚲Δ\)−i,i⊤\+𝚪−i,−i\(𝚲Δ\)−i,−i⊤=0\.\{\\color\[rgb\]\{0,0,0\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,i\}\\bm\{\\Gamma\}\_\{i,\-i\}\+\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}\\bm\{\\Gamma\}\_\{\-i,\-i\}\+\\bm\{\\Gamma\}\_\{\-i,i\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,i\}^\{\\top\}\+\\bm\{\\Gamma\}\_\{\-i,\-i\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}^\{\\top\}=0\.\}\(27\)Using equation[21](https://arxiv.org/html/2609.19955#A1.E21), we obtain \(𝚲Δ\)−i,−i𝚵\+𝚵⊤\(𝚲Δ\)−i,−i⊤=0,\{\\color\[rgb\]\{0,0,0\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}\\bm\{\\Xi\}\+\\bm\{\\Xi\}^\{\\top\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}^\{\\top\}=0,\}\(28\)where 𝚵:=𝚪−i,−i−𝚺−i,i\(i\)𝚪i,−i𝚺ii\(i\)\.\\bm\{\\Xi\}:=\\bm\{\\Gamma\}\_\{\-i,\-i\}\-\\frac\{\\bm\{\\Sigma\}^\{\(i\)\}\_\{\-i,i\}\\bm\{\\Gamma\}\_\{i,\-i\}\}\{\\bm\{\\Sigma\}^\{\(i\)\}\_\{ii\}\}\.\(29\) This Schur\-type expression need not be symmetric or a covariance matrix\. By Lemma[A\.3](https://arxiv.org/html/2609.19955#A1.Thmtheorem3),𝚵\\bm\{\\Xi\}is invertible\. Define 𝐀:=𝚵−1𝐙,𝐒:=\(𝚲Δ\)−i,−i𝚵\.\{\\color\[rgb\]\{0,0,0\}\\mathbf\{A\}:=\\bm\{\\Xi\}^\{\-1\}\\mathbf\{Z\},\\qquad\\mathbf\{S\}:=\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}\\bm\{\\Xi\}\.\}Then equation[28](https://arxiv.org/html/2609.19955#A1.E28)implies𝐒⊤=−𝐒\\mathbf\{S\}^\{\\top\}=\-\\mathbf\{S\}\. Moreover, 𝐒𝐀−𝐀⊤𝐒=\(𝚲Δ\)−i,−i𝐙\+𝐙\(𝚲Δ\)−i,−i⊤=\(𝐃Δ\)−i,−i,\{\\color\[rgb\]\{0,0,0\}\\mathbf\{S\}\\mathbf\{A\}\-\\mathbf\{A\}^\{\\top\}\\mathbf\{S\}=\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}\\mathbf\{Z\}\+\\mathbf\{Z\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}^\{\\top\}=\(\\mathbf\{D\}\_\{\\Delta\}\)\_\{\-i,\-i\},\}where the last equality follows from equation[22](https://arxiv.org/html/2609.19955#A1.E22)\. Since\(𝐃Δ\)−i,−i\(\\mathbf\{D\}\_\{\\Delta\}\)\_\{\-i,\-i\}is diagonal, offdiag\(𝐒𝐀−𝐀⊤𝐒\)=0\.\{\\color\[rgb\]\{0,0,0\}\\operatorname\{offdiag\}\(\\mathbf\{S\}\\mathbf\{A\}\-\\mathbf\{A\}^\{\\top\}\\mathbf\{S\}\)=0\.\} ###### Assumption A\.1\. Let𝐀:=𝚵−1𝐙\\mathbf\{A\}:=\\bm\{\\Xi\}^\{\-1\}\\mathbf\{Z\}and letm=n−1m=n\-1be the reduced block dimension \(orm=\|T\|−1m=\|T\|\-1for a target SCCTT\)\. Heremmis local to this assumption\. Consider the linear mapℒ𝐀:𝒦m→ℝoffm×m\\mathcal\{L\}\_\{\\mathbf\{A\}\}:\\mathcal\{K\}\_\{m\}\\to\\mathbb\{R\}^\{m\\times m\}\_\{\\mathrm\{off\}\},where the codomain consists of matrices with zero diagonal and𝒦m:=\{𝐒∈ℝm×m:𝐒⊤=−𝐒\}\\mathcal\{K\}\_\{m\}:=\\\{\\mathbf\{S\}\\in\\mathbb\{R\}^\{m\\times m\}:\\mathbf\{S\}^\{\\top\}=\-\\mathbf\{S\}\\\}, defined by ℒ𝐀\(𝐒\)=offdiag\(𝐒𝐀−𝐀⊤𝐒\)\.\\mathcal\{L\}\_\{\\mathbf\{A\}\}\(\\mathbf\{S\}\)=\\operatorname\{offdiag\}\\\!\\big\(\\mathbf\{S\}\\mathbf\{A\}\-\\mathbf\{A\}^\{\\top\}\\mathbf\{S\}\\big\)\.Let𝐋𝐀\\mathbf\{L\}\_\{\\mathbf\{A\}\}be the matrix representation ofℒ𝐀\\mathcal\{L\}\_\{\\mathbf\{A\}\}in any basis of𝒦m\\mathcal\{K\}\_\{m\}\. Writeσ\(M\)\\sigma\(M\)for the spectrum \(set of complex eigenvalues\) ofMM\(please note that this differs from the diffusion matrix𝝈\\bm\{\\sigma\}\)\. We assume0∉σ\(𝐋𝐀⊤𝐋𝐀\)\.0\\notin\\sigma\(\\mathbf\{L\}\_\{\\mathbf\{A\}\}^\{\\top\}\\mathbf\{L\}\_\{\\mathbf\{A\}\}\)\.Equivalently, the only skew\-symmetric matrix𝐒\\mathbf\{S\}satisfyingoffdiag\(𝐒𝐀−𝐀⊤𝐒\)=0\\operatorname\{offdiag\}\\\!\\big\(\\mathbf\{S\}\\mathbf\{A\}\-\\mathbf\{A\}^\{\\top\}\\mathbf\{S\}\\big\)=0is𝐒=0\\mathbf\{S\}=0\. This is an injectivity condition: the kernel of the map, namely the set of inputs mapped to zero, is\{0\}\\\{0\\\}\. Equivalently,𝐋𝐀\\mathbf\{L\}\_\{\\mathbf\{A\}\}has full column rank and𝐋𝐀⊤𝐋𝐀\\mathbf\{L\}\_\{\\mathbf\{A\}\}^\{\\top\}\\mathbf\{L\}\_\{\\mathbf\{A\}\}is positive definite\. It therefore rules out a nonzero skew\-symmetric ambiguity in the moments\.By Assumption[A\.1](https://arxiv.org/html/2609.19955#A1.Thmtheorem1),𝐒=0\\mathbf\{S\}=0\. Since𝐒=\(𝚲Δ\)−i,−i𝚵\\mathbf\{S\}=\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}\\bm\{\\Xi\}and𝚵\\bm\{\\Xi\}is invertible,\(𝚲Δ\)−i,−i=0\.\{\\color\[rgb\]\{0,0,0\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,\-i\}=0\.\}Using equation[21](https://arxiv.org/html/2609.19955#A1.E21), we also obtain\(𝚲Δ\)−i,i=0\.\{\\color\[rgb\]\{0,0,0\}\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{\-i,i\}=0\.\}Together with\(𝚲Δ\)ii=0\(\\bm\{\\Lambda\}\_\{\\Delta\}\)\_\{ii\}=0, this shows that the only possibly nonzero entries of𝚲Δ\\bm\{\\Lambda\}\_\{\\Delta\}are the off\-diagonal entries in rowii\. Moreover, from equation[22](https://arxiv.org/html/2609.19955#A1.E22), we get\(𝐃Δ\)−i,−i=0\(\\mathbf\{D\}\_\{\\Delta\}\)\_\{\-i,\-i\}=0\. Since we already showed\(𝐃Δ\)ii=0\(\\mathbf\{D\}\_\{\\Delta\}\)\_\{ii\}=0, it follows that𝐃Δ=0\.\{\\color\[rgb\]\{0,0,0\}\\mathbf\{D\}\_\{\\Delta\}=0\.\} Now the observational covariance equation gives𝚲Δ𝚺\+𝚺𝚲Δ⊤=0\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\Delta\}\\bm\{\\Sigma\}\+\\bm\{\\Sigma\}\\bm\{\\Lambda\}\_\{\\Delta\}^\{\\top\}=0\.\}Because only rowiiof𝚲Δ\\bm\{\\Lambda\}\_\{\\Delta\}can be nonzero, write𝚲Δ=𝐞i𝐮⊤,ui=0\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\Delta\}=\\mathbf\{e\}\_\{i\}\\mathbf\{u\}^\{\\top\},u\_\{i\}=0\.\}Then𝐞i𝐮⊤𝚺\+𝚺𝐮𝐞i⊤=0\.\{\\color\[rgb\]\{0,0,0\}\\mathbf\{e\}\_\{i\}\\mathbf\{u\}^\{\\top\}\\bm\{\\Sigma\}\+\\bm\{\\Sigma\}\\mathbf\{u\}\\mathbf\{e\}\_\{i\}^\{\\top\}=0\.\}Taking theii\-th row gives𝐮⊤𝚺\+\(𝚺𝐮\)i𝐞i⊤=0\.\{\\color\[rgb\]\{0,0,0\}\\mathbf\{u\}^\{\\top\}\\bm\{\\Sigma\}\+\(\\bm\{\\Sigma\}\\mathbf\{u\}\)\_\{i\}\\mathbf\{e\}\_\{i\}^\{\\top\}=0\.\}For everyj≠ij\\neq i, this implies\(𝐮⊤𝚺\)j=0\(\\mathbf\{u\}^\{\\top\}\\bm\{\\Sigma\}\)\_\{j\}=0, while the\(i,i\)\(i,i\)entry gives2\(𝐮⊤𝚺\)i=02\(\\mathbf\{u\}^\{\\top\}\\bm\{\\Sigma\}\)\_\{i\}=0\. Therefore𝐮⊤𝚺=0\\mathbf\{u\}^\{\\top\}\\bm\{\\Sigma\}=0\. Since𝚺≻0\\bm\{\\Sigma\}\\succ 0is invertible,𝐮=0\\mathbf\{u\}=0, and hence𝚲Δ=0\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\Delta\}=0\.\}Finally, the observational mean equation gives𝚲Δ𝝁=𝐛Δ\\bm\{\\Lambda\}\_\{\\Delta\}\\bm\{\\mu\}=\\mathbf\{b\}\_\{\\Delta\}, so𝐛Δ=0\\mathbf\{b\}\_\{\\Delta\}=0\. Hence the scaled\-difference triple is zero:𝚲′=c𝚲⋆,𝐛′=c𝐛⋆,𝐃′=c𝐃⋆\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}^\{\\prime\}=c\\bm\{\\Lambda\}\_\{\\star\},\\mathbf\{b\}^\{\\prime\}=c\\mathbf\{b\}\_\{\\star\},\\mathbf\{D\}^\{\\prime\}=c\\mathbf\{D\}\_\{\\star\}\.\}Since𝐃′\\mathbf\{D\}^\{\\prime\}and𝐃⋆\\mathbf\{D\}\_\{\\star\}are positive diagonal matrices, we havec\>0c\>0\. This completes the proof\. ###### Assumption A\.2\. Consider right/left eigenbases of the true drift𝚲⋆\\bm\{\\Lambda\}\_\{\\star\}:𝚲⋆𝐫⋆,ℓ=λ⋆,ℓ𝐫⋆,ℓ,𝝆⋆,k⊤𝚲⋆=λ⋆,k𝝆⋆,k⊤,𝝆⋆,k⊤𝐫⋆,ℓ=δkℓ\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\star\}\\mathbf\{r\}\_\{\\star,\\ell\}=\\lambda\_\{\\star,\\ell\}\\mathbf\{r\}\_\{\\star,\\ell\},\\bm\{\\rho\}\_\{\\star,k\}^\{\\\!\\top\}\\bm\{\\Lambda\}\_\{\\star\}=\\lambda\_\{\\star,k\}\\bm\{\\rho\}\_\{\\star,k\}^\{\\\!\\top\},\\bm\{\\rho\}\_\{\\star,k\}^\{\\\!\\top\}\\mathbf\{r\}\_\{\\star,\\ell\}=\\delta\_\{k\\ell\}\.\}Let𝐑⋆:=\[𝐫⋆,1⋯𝐫⋆,n\],𝐏⋆:=\[𝝆⋆,1⋯𝝆⋆,n\],\{\\color\[rgb\]\{0,0,0\}\\mathbf\{R\}\_\{\\star\}:=\[\\mathbf\{r\}\_\{\\star,1\}\\,\\cdots\\,\\mathbf\{r\}\_\{\\star,n\}\],\\mathbf\{P\}\_\{\\star\}:=\[\\bm\{\\rho\}\_\{\\star,1\}\\,\\cdots\\,\\bm\{\\rho\}\_\{\\star,n\}\],\}so that𝐏⋆⊤𝐑⋆=𝐈\\mathbf\{P\}\_\{\\star\}^\{\\\!\\top\}\\mathbf\{R\}\_\{\\star\}=\\mathbf\{I\}\. Since the eigenvalues are distinct,𝐑⋆\\mathbf\{R\}\_\{\\star\}is invertible\. We choose the rows of𝐑⋆−1\\mathbf\{R\}\_\{\\star\}^\{\-1\}as the left eigenvectors\. Here⊤\\topdenotes ordinary transpose, without complex conjugation\. Let𝐰⋆∈ℝn\\mathbf\{w\}\_\{\\star\}\\in\\mathbb\{R\}^\{n\}be the true intervened\-row vector, defined by\(w⋆\)i=0\(w\_\{\\star\}\)\_\{i\}=0and\(𝐰⋆\)−i=\(𝚲⋆\)i,−i⊤\(\\mathbf\{w\}\_\{\\star\}\)\_\{\-i\}=\(\\bm\{\\Lambda\}\_\{\\star\}\)\_\{i,\-i\}^\{\\top\}, so that𝚲~⋆\(i\)=𝚲⋆−𝐞i𝐰⋆⊤\.\{\\color\[rgb\]\{0,0,0\}\\widetilde\{\\bm\{\\Lambda\}\}\_\{\\star\}^\{\(i\)\}=\\bm\{\\Lambda\}\_\{\\star\}\-\\mathbf\{e\}\_\{i\}\\mathbf\{w\}\_\{\\star\}^\{\\top\}\.\} 1. 1\.We assume that𝚲⋆\\bm\{\\Lambda\}\_\{\\star\}has simple spectrum, i\.e\.,λ⋆,k≠λ⋆,ℓ\\lambda\_\{\\star,k\}\\neq\\lambda\_\{\\star,\\ell\}fork≠ℓk\\neq\\ell\. 2. 2\.Let𝐖⋆∈ℂn×n\\mathbf\{W\}\_\{\\star\}\\in\\mathbb\{C\}^\{n\\times n\}be defined by \(𝐖⋆\)kℓ:=\(𝝆⋆,k⊤𝐞i\)\(𝐰⋆⊤𝚺\(i\)𝝆⋆,ℓ\)λ⋆,k\+λ⋆,ℓ\.\{\\color\[rgb\]\{0,0,0\}\(\\mathbf\{W\}\_\{\\star\}\)\_\{k\\ell\}:=\\frac\{\(\\bm\{\\rho\}\_\{\\star,k\}^\{\\\!\\top\}\\mathbf\{e\}\_\{i\}\)\\big\(\\mathbf\{w\}\_\{\\star\}^\{\\\!\\top\}\\bm\{\\Sigma\}^\{\(i\)\}\\bm\{\\rho\}\_\{\\star,\\ell\}\\big\)\}\{\\lambda\_\{\\star,k\}\+\\lambda\_\{\\star,\\ell\}\}\.\}We assume 0∉σ\(𝐖⋆\+𝐖⋆⊤\)\.\{\\color\[rgb\]\{0,0,0\}0\\notin\\sigma\(\\mathbf\{W\}\_\{\\star\}\+\\mathbf\{W\}\_\{\\star\}^\{\\\!\\top\}\)\.\} 3. 3\.With𝚪:=𝚺−𝚺\(i\)\\bm\{\\Gamma\}:=\\bm\{\\Sigma\}\-\\bm\{\\Sigma\}^\{\(i\)\}, we assume:𝐞i⊤𝚪−1𝚺\(i\)𝐞i≠0\.\{\\color\[rgb\]\{0,0,0\}\\mathbf\{e\}\_\{i\}^\{\\\!\\top\}\\bm\{\\Gamma\}^\{\-1\}\\bm\{\\Sigma\}^\{\(i\)\}\\mathbf\{e\}\_\{i\}\\neq 0\.\} ###### Lemma A\.3\. Under Assumption[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2), the matrix𝚵\\bm\{\\Xi\}is invertible\. ###### Proof\. Since𝚲~⋆\(i\)=𝚲⋆−𝐞i𝐰⋆⊤,\{\\color\[rgb\]\{0,0,0\}\\widetilde\{\\bm\{\\Lambda\}\}\_\{\\star\}^\{\(i\)\}=\\bm\{\\Lambda\}\_\{\\star\}\-\\mathbf\{e\}\_\{i\}\\mathbf\{w\}\_\{\\star\}^\{\\top\},\}subtracting the true observational and interventional Lyapunov equations gives 𝚲⋆𝚪\+𝚪𝚲⋆⊤\+𝐞i𝐰⋆⊤𝚺\(i\)\+𝚺\(i\)𝐰⋆𝐞i⊤=0\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\star\}\\bm\{\\Gamma\}\+\\bm\{\\Gamma\}\\bm\{\\Lambda\}\_\{\\star\}^\{\\top\}\+\\mathbf\{e\}\_\{i\}\\mathbf\{w\}\_\{\\star\}^\{\\top\}\\bm\{\\Sigma\}^\{\(i\)\}\+\\bm\{\\Sigma\}^\{\(i\)\}\\mathbf\{w\}\_\{\\star\}\\mathbf\{e\}\_\{i\}^\{\\top\}=0\.\}\(30\) Let𝐔\\mathbf\{U\}solve the Sylvester equation 𝚲⋆𝐔\+𝐔𝚲⋆⊤=−𝐞i\(𝐰⋆⊤𝚺\(i\)\)\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\star\}\\mathbf\{U\}\+\\mathbf\{U\}\\bm\{\\Lambda\}\_\{\\star\}^\{\\\!\\top\}=\-\\mathbf\{e\}\_\{i\}\\big\(\\mathbf\{w\}\_\{\\star\}^\{\\\!\\top\}\\bm\{\\Sigma\}^\{\(i\)\}\\big\)\.\}\(31\)The Sylvester operatorX↦AX\+XBX\\mapsto AX\+XBis invertible precisely whenAAand−B\-Bhave disjoint spectra\. HereA=𝚲⋆A=\\bm\{\\Lambda\}\_\{\\star\}andB=𝚲⋆⊤B=\\bm\{\\Lambda\}\_\{\\star\}^\{\\top\}; positive stability makes every eigenvalue sum nonzero\. Thus this operator is invertible\. Transposing equation[31](https://arxiv.org/html/2609.19955#A1.E31)and adding gives 𝚲⋆\(𝐔\+𝐔⊤\)\+\(𝐔\+𝐔⊤\)𝚲⋆⊤=−𝐞i𝐰⋆⊤𝚺\(i\)−𝚺\(i\)𝐰⋆𝐞i⊤\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\star\}\(\\mathbf\{U\}\+\\mathbf\{U\}^\{\\top\}\)\+\(\\mathbf\{U\}\+\\mathbf\{U\}^\{\\top\}\)\\bm\{\\Lambda\}\_\{\\star\}^\{\\top\}=\-\\mathbf\{e\}\_\{i\}\\mathbf\{w\}\_\{\\star\}^\{\\top\}\\bm\{\\Sigma\}^\{\(i\)\}\-\\bm\{\\Sigma\}^\{\(i\)\}\\mathbf\{w\}\_\{\\star\}\\mathbf\{e\}\_\{i\}^\{\\top\}\.\}Comparing with equation[30](https://arxiv.org/html/2609.19955#A1.E30)and using uniqueness gives 𝚪=𝐔\+𝐔⊤\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Gamma\}=\\mathbf\{U\}\+\\mathbf\{U\}^\{\\top\}\.\}\(32\) Expand𝐔\\mathbf\{U\}in the right\-eigenvector basis on both sides: 𝐔=∑k,ℓukℓ𝐫⋆,k𝐫⋆,ℓ⊤,\{\\color\[rgb\]\{0,0,0\}\\mathbf\{U\}=\\sum\_\{k,\\ell\}u\_\{k\\ell\}\\,\\mathbf\{r\}\_\{\\star,k\}\\mathbf\{r\}\_\{\\star,\\ell\}^\{\\\!\\top\},\}whereukℓ=𝝆⋆,k⊤𝐔𝝆⋆,ℓu\_\{k\\ell\}=\\bm\{\\rho\}\_\{\\star,k\}^\{\\\!\\top\}\\mathbf\{U\}\\bm\{\\rho\}\_\{\\star,\\ell\}\. Multiplying equation[31](https://arxiv.org/html/2609.19955#A1.E31)on the left by𝝆⋆,k⊤\\bm\{\\rho\}\_\{\\star,k\}^\{\\\!\\top\}and on the right by𝝆⋆,ℓ\\bm\{\\rho\}\_\{\\star,\\ell\}gives \(λ⋆,k\+λ⋆,ℓ\)ukℓ=−\(𝝆⋆,k⊤𝐞i\)\(𝐰⋆⊤𝚺\(i\)𝝆⋆,ℓ\)\.\{\\color\[rgb\]\{0,0,0\}\(\\lambda\_\{\\star,k\}\+\\lambda\_\{\\star,\\ell\}\)u\_\{k\\ell\}=\-\(\\bm\{\\rho\}\_\{\\star,k\}^\{\\\!\\top\}\\mathbf\{e\}\_\{i\}\)\\big\(\\mathbf\{w\}\_\{\\star\}^\{\\\!\\top\}\\bm\{\\Sigma\}^\{\(i\)\}\\bm\{\\rho\}\_\{\\star,\\ell\}\\big\)\.\}Thusukℓ=−\(𝐖⋆\)kℓu\_\{k\\ell\}=\-\(\\mathbf\{W\}\_\{\\star\}\)\_\{k\\ell\}, and 𝐔=−𝐑⋆𝐖⋆𝐑⋆⊤\.\{\\color\[rgb\]\{0,0,0\}\\mathbf\{U\}=\-\\mathbf\{R\}\_\{\\star\}\\mathbf\{W\}\_\{\\star\}\\mathbf\{R\}\_\{\\star\}^\{\\top\}\.\}\(33\)Using equation[32](https://arxiv.org/html/2609.19955#A1.E32), 𝚪=−𝐑⋆\(𝐖⋆\+𝐖⋆⊤\)𝐑⋆⊤\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Gamma\}=\-\\mathbf\{R\}\_\{\\star\}\(\\mathbf\{W\}\_\{\\star\}\+\\mathbf\{W\}\_\{\\star\}^\{\\top\}\)\\mathbf\{R\}\_\{\\star\}^\{\\top\}\.\}By Assumption[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)\(2\),𝐖⋆\+𝐖⋆⊤\\mathbf\{W\}\_\{\\star\}\+\\mathbf\{W\}\_\{\\star\}^\{\\top\}is invertible\. Since𝐑⋆\\mathbf\{R\}\_\{\\star\}is invertible,𝚪\\bm\{\\Gamma\}is invertible\. LetJ=\{1,…,n\}∖\{i\}J=\\\{1,\\dots,n\\\}\\setminus\\\{i\\\}\. Define 𝐌i:=\[𝚺ii\(i\)𝚪i,J𝚺J,i\(i\)𝚪J,J\]\.\\mathbf\{M\}\_\{i\}:=\\begin\{bmatrix\}\\bm\{\\Sigma\}^\{\(i\)\}\_\{ii\}&\\bm\{\\Gamma\}\_\{i,J\}\\\\\[2\.0pt\] \\bm\{\\Sigma\}^\{\(i\)\}\_\{J,i\}&\\bm\{\\Gamma\}\_\{J,J\}\\end\{bmatrix\}\.This matrix is obtained from𝚪\\bm\{\\Gamma\}by replacing itsii\-th column with𝚺\(i\)𝐞i\\bm\{\\Sigma\}^\{\(i\)\}\\mathbf\{e\}\_\{i\}, followed by the same row and column permutation placingiifirst; this permutation leaves the determinant unchanged\. By Cramer’s rule, det\(𝐌i\)=det\(𝚪\)𝐞i⊤𝚪−1𝚺\(i\)𝐞i\.\{\\color\[rgb\]\{0,0,0\}\\det\(\\mathbf\{M\}\_\{i\}\)=\\det\(\\bm\{\\Gamma\}\)\\,\\mathbf\{e\}\_\{i\}^\{\\top\}\\bm\{\\Gamma\}^\{\-1\}\\bm\{\\Sigma\}^\{\(i\)\}\\mathbf\{e\}\_\{i\}\.\}By Assumption[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)\(3\),det\(𝐌i\)≠0\\det\(\\mathbf\{M\}\_\{i\}\)\\neq 0\. On the other hand, the Schur determinant identity gives, det\(𝐌i\)=𝚺ii\(i\)det\(𝚵\)\.\{\\color\[rgb\]\{0,0,0\}\\det\(\\mathbf\{M\}\_\{i\}\)=\\bm\{\\Sigma\}^\{\(i\)\}\_\{ii\}\\det\(\\bm\{\\Xi\}\)\.\}Since𝚺ii\(i\)\>0\\bm\{\\Sigma\}^\{\(i\)\}\_\{ii\}\>0, we getdet\(𝚵\)≠0\\det\(\\bm\{\\Xi\}\)\\neq 0\. Thus𝚵\\bm\{\\Xi\}is invertible\.∎ ### A\.7Proof of Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6) Let𝐓\\mathbf\{T\}be the target block\. For a nonsingleton root SCC, itsobservational and interventional marginals are sub\-blocks of the observedmoments\. The isolated\-realization hypothesis makes the𝚵\\bm\{\\Xi\}invertibility and skew\-kernel conditions used in the proof of Theorem[3\.1](https://arxiv.org/html/2609.19955#S3.Thmtheorem1)generic, by the rational\-rank argument below\. That proof therefore identifies the root parameters up to a positive common scale\. If the root SCC is a singleton, its observational equationsλμ=b\\lambda\\mu=band2λΣ=d2\\lambda\\Sigma=d, withλ\>0\\lambda\>0, give the same conclusion\.Suppose the parameters of block𝐏\\mathbf\{P\}satisfy𝐛𝐏=c𝐛¯𝐏\\mathbf\{b\}\_\{\\mathbf\{P\}\}=c\\,\\bar\{\\mathbf\{b\}\}\_\{\\mathbf\{P\}\}and𝐃𝐏=c𝐃¯𝐏\\mathbf\{D\}\_\{\\mathbf\{P\}\}=c\\,\\bar\{\\mathbf\{D\}\}\_\{\\mathbf\{P\}\}for a scalarc\>0c\>0and known representatives𝐛¯𝐏,𝐃¯𝐏\\bar\{\\mathbf\{b\}\}\_\{\\mathbf\{P\}\},\\bar\{\\mathbf\{D\}\}\_\{\\mathbf\{P\}\}\. Define the residual covariances and cross terms 𝐙:=𝚺𝐓\|𝐏,𝐙\(i\):=𝚺𝐓\|𝐏\(i\),𝐁:=𝚺𝐓𝐏𝚺𝐏𝐏−1,𝐁\(i\):=𝚺𝐓𝐏\(i\)\(𝚺𝐏𝐏\(i\)\)−1\.\\mathbf\{Z\}:=\\bm\{\\Sigma\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\},\\quad\\mathbf\{Z\}^\{\(i\)\}:=\\bm\{\\Sigma\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\},\\quad\\mathbf\{B\}:=\\bm\{\\Sigma\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}^\{\-1\},\\quad\\mathbf\{B\}^\{\(i\)\}:=\\bm\{\\Sigma\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}\\big\(\\bm\{\\Sigma\}^\{\(i\)\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}\\big\)^\{\-1\}\.Define the parent mean contribution𝐡𝐏:=𝐛¯𝐏−𝐃¯𝐏𝚺𝐏𝐏−1𝝁𝐏\.\{\\color\[rgb\]\{0,0,0\}\\mathbf\{h\}\_\{\\mathbf\{P\}\}:=\\bar\{\\mathbf\{b\}\}\_\{\\mathbf\{P\}\}\-\\bar\{\\mathbf\{D\}\}\_\{\\mathbf\{P\}\}\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}^\{\-1\}\\bm\{\\mu\}\_\{\\mathbf\{P\}\}\.\}Then set 𝐯:=𝐁𝐡𝐏,𝐯\(i\):=𝐁\(i\)𝐡𝐏,𝐐:=𝐁𝐃¯𝐏𝐁⊤,𝐐\(i\):=𝐁\(i\)𝐃¯𝐏𝐁\(i\)⊤\.\{\\color\[rgb\]\{0,0,0\}\\mathbf\{v\}:=\\mathbf\{B\}\\mathbf\{h\}\_\{\\mathbf\{P\}\},\\quad\\mathbf\{v\}^\{\(i\)\}:=\\mathbf\{B\}^\{\(i\)\}\\mathbf\{h\}\_\{\\mathbf\{P\}\},\}\\qquad\\mathbf\{Q\}:=\\mathbf\{B\}\\,\\bar\{\\mathbf\{D\}\}\_\{\\mathbf\{P\}\}\\,\\mathbf\{B\}^\{\\top\},\\quad\\mathbf\{Q\}^\{\(i\)\}:=\\mathbf\{B\}^\{\(i\)\}\\bar\{\\mathbf\{D\}\}\_\{\\mathbf\{P\}\}\\,\\mathbf\{B\}^\{\(i\)\\top\}\.Here we used that the intervention is inside the target block𝐓\\mathbf\{T\}, so the upstream parent moments𝝁𝐏\\bm\{\\mu\}\_\{\\mathbf\{P\}\}and𝚺𝐏𝐏\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}are unchanged by the intervention\.Indeed, under the hard intervention, rowiiof the full drift has no off\-diagonal entries\. In particular, theii\-th row of𝚲~𝐓𝐏\(i\)\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}is zero and theii\-th row of𝚲~𝐓𝐓\(i\)\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}is\(λii,0\)\(\\lambda\_\{ii\},0\)\. Taking theii\-th row of the interventional cross\-covariance Lyapunov block gives λii𝚺i,𝐏\(i\)\+𝚺i,𝐏\(i\)𝚲𝐏𝐏⊤=0\.\\lambda\_\{ii\}\\bm\{\\Sigma\}^\{\(i\)\}\_\{i,\\mathbf\{P\}\}\+\\bm\{\\Sigma\}^\{\(i\)\}\_\{i,\\mathbf\{P\}\}\\bm\{\\Lambda\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}^\{\\top\}=0\.Sinceλii𝐈\+𝚲𝐏𝐏⊤\\lambda\_\{ii\}\\mathbf\{I\}\+\\bm\{\\Lambda\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}^\{\\top\}is invertible, we obtain𝚺i,𝐏\(i\)=0\\bm\{\\Sigma\}^\{\(i\)\}\_\{i,\\mathbf\{P\}\}=0\. Hence theii\-th row of𝐁\(i\)=𝚺𝐓𝐏\(i\)\(𝚺𝐏𝐏\(i\)\)−1\\mathbf\{B\}^\{\(i\)\}=\\bm\{\\Sigma\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}\(\\bm\{\\Sigma\}^\{\(i\)\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}\)^\{\-1\}is zero, and therefore\(𝐐\(i\)\)i,−i=\(𝐐\(i\)\)−i,i=0\(\\mathbf\{Q\}^\{\(i\)\}\)\_\{i,\-i\}=\(\\mathbf\{Q\}^\{\(i\)\}\)\_\{\-i,i\}=0\. Partition the drift on𝐓\\mathbf\{T\}as 𝚲𝐓𝐓=\[λii𝝀i,−i𝝀−i,i𝚲−i,−i\],𝚲~𝐓𝐓\(i\)=\[λii0𝝀−i,i𝚲−i,−i\],\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}=\\begin\{bmatrix\}\\lambda\_\{ii\}&\\bm\{\\lambda\}\_\{i,\-i\}\\\\ \\bm\{\\lambda\}\_\{\-i,i\}&\\bm\{\\Lambda\}\_\{\-i,\-i\}\\end\{bmatrix\},\\qquad\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}=\\begin\{bmatrix\}\\lambda\_\{ii\}&0\\\\ \\bm\{\\lambda\}\_\{\-i,i\}&\\bm\{\\Lambda\}\_\{\-i,\-i\}\\end\{bmatrix\},with unknown diagonal𝐃𝐓\\mathbf\{D\}\_\{\\mathbf\{T\}\}and vector𝐛𝐓\\mathbf\{b\}\_\{\\mathbf\{T\}\}\. We first justify the residual mean equation\. From the block mean equations 𝚲𝐏𝐏𝝁𝐏=𝐛𝐏,𝚲𝐓𝐏𝝁𝐏\+𝚲𝐓𝐓𝝁𝐓=𝐛𝐓,\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}\\bm\{\\mu\}\_\{\\mathbf\{P\}\}=\\mathbf\{b\}\_\{\\mathbf\{P\}\},\\qquad\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}\\bm\{\\mu\}\_\{\\mathbf\{P\}\}\+\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\bm\{\\mu\}\_\{\\mathbf\{T\}\}=\\mathbf\{b\}\_\{\\mathbf\{T\}\},\}and from 𝝁𝐓=𝝁𝐓\|𝐏\+𝐁𝝁𝐏,\{\\color\[rgb\]\{0,0,0\}\\bm\{\\mu\}\_\{\\mathbf\{T\}\}=\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}\+\\mathbf\{B\}\\bm\{\\mu\}\_\{\\mathbf\{P\}\},\}we obtain 𝚲𝐓𝐓𝝁𝐓\|𝐏=𝐛𝐓−\(𝚲𝐓𝐏\+𝚲𝐓𝐓𝐁\)𝝁𝐏\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}=\\mathbf\{b\}\_\{\\mathbf\{T\}\}\-\(\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}\+\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\mathbf\{B\}\)\\bm\{\\mu\}\_\{\\mathbf\{P\}\}\.\}The cross\-covariance block of the Lyapunov equation gives 𝚲𝐓𝐏𝚺𝐏𝐏\+𝚲𝐓𝐓𝚺𝐓𝐏\+𝚺𝐓𝐏𝚲𝐏𝐏⊤=0\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}\+\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\bm\{\\Sigma\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}\+\\bm\{\\Sigma\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}\\bm\{\\Lambda\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}^\{\\top\}=0\.\}Using𝐁=𝚺𝐓𝐏𝚺𝐏𝐏−1\\mathbf\{B\}=\\bm\{\\Sigma\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}^\{\-1\}, this implies 𝚲𝐓𝐏\+𝚲𝐓𝐓𝐁=−𝐁𝚺𝐏𝐏𝚲𝐏𝐏⊤𝚺𝐏𝐏−1\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}\+\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\mathbf\{B\}=\-\\mathbf\{B\}\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}\\bm\{\\Lambda\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}^\{\\top\}\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}^\{\-1\}\.\}The parent Lyapunov equation 𝚲𝐏𝐏𝚺𝐏𝐏\+𝚺𝐏𝐏𝚲𝐏𝐏⊤=𝐃𝐏\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}\+\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}\\bm\{\\Lambda\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}^\{\\top\}=\\mathbf\{D\}\_\{\\mathbf\{P\}\}\}then gives 𝚲𝐓𝐏\+𝚲𝐓𝐓𝐁=𝐁𝚲𝐏𝐏−𝐁𝐃𝐏𝚺𝐏𝐏−1\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}\+\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\mathbf\{B\}=\\mathbf\{B\}\\bm\{\\Lambda\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}\-\\mathbf\{B\}\\mathbf\{D\}\_\{\\mathbf\{P\}\}\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}^\{\-1\}\.\}Since𝐛𝐏=c𝐛¯𝐏\\mathbf\{b\}\_\{\\mathbf\{P\}\}=c\\bar\{\\mathbf\{b\}\}\_\{\\mathbf\{P\}\}and𝐃𝐏=c𝐃¯𝐏\\mathbf\{D\}\_\{\\mathbf\{P\}\}=c\\bar\{\\mathbf\{D\}\}\_\{\\mathbf\{P\}\}, we get 𝚲𝐓𝐓𝝁𝐓\|𝐏=𝐛𝐓−c𝐁\(𝐛¯𝐏−𝐃¯𝐏𝚺𝐏𝐏−1𝝁𝐏\)=𝐛𝐓−c𝐯\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}=\\mathbf\{b\}\_\{\\mathbf\{T\}\}\-c\\,\\mathbf\{B\}\\left\(\\bar\{\\mathbf\{b\}\}\_\{\\mathbf\{P\}\}\-\\bar\{\\mathbf\{D\}\}\_\{\\mathbf\{P\}\}\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}^\{\-1\}\\bm\{\\mu\}\_\{\\mathbf\{P\}\}\\right\)=\\mathbf\{b\}\_\{\\mathbf\{T\}\}\-c\\mathbf\{v\}\.\} On𝐓\\mathbf\{T\}, the stationary equations are 𝚲𝐓𝐓𝐙\+𝐙𝚲𝐓𝐓⊤=𝐃𝐓\+c𝐐,𝚲𝐓𝐓𝝁𝐓\|𝐏=𝐛𝐓−c𝐯,\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\mathbf\{Z\}\+\\mathbf\{Z\}\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}^\{\\top\}=\\mathbf\{D\}\_\{\\mathbf\{T\}\}\+c\\,\\mathbf\{Q\},\\qquad\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}=\\mathbf\{b\}\_\{\\mathbf\{T\}\}\-c\\,\\mathbf\{v\},and, under intervention, 𝚲~𝐓𝐓\(i\)𝐙\(i\)\+𝐙\(i\)𝚲~𝐓𝐓\(i\)⊤=𝐃𝐓\+c𝐐\(i\),𝚲~𝐓𝐓\(i\)𝝁𝐓\|𝐏\(i\)=𝐛𝐓−c𝐯\(i\)\.\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\mathbf\{Z\}^\{\(i\)\}\+\\mathbf\{Z\}^\{\(i\)\}\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\\top\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}=\\mathbf\{D\}\_\{\\mathbf\{T\}\}\+c\\,\\mathbf\{Q\}^\{\(i\)\},\\qquad\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\bm\{\\mu\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}=\\mathbf\{b\}\_\{\\mathbf\{T\}\}\-c\\,\\mathbf\{v\}^\{\(i\)\}\.\(34\)Subtracting gives the difference relations 𝚲𝐓𝐓𝐙\+𝐙𝚲𝐓𝐓⊤−\(𝚲~𝐓𝐓\(i\)𝐙\(i\)\+𝐙\(i\)𝚲~𝐓𝐓\(i\)⊤\)=c\(𝐐−𝐐\(i\)\),\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\mathbf\{Z\}\+\\mathbf\{Z\}\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}^\{\\top\}\-\\big\(\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\mathbf\{Z\}^\{\(i\)\}\+\\mathbf\{Z\}^\{\(i\)\}\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\\top\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\big\)=c\\,\(\\mathbf\{Q\}\-\\mathbf\{Q\}^\{\(i\)\}\),𝚲𝐓𝐓𝝁𝐓\|𝐏−𝚲~𝐓𝐓\(i\)𝝁𝐓\|𝐏\(i\)=c\(𝐯\(i\)−𝐯\)\.\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}\-\\widetilde\{\\bm\{\\Lambda\}\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\bm\{\\mu\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}=c\\,\(\\mathbf\{v\}^\{\(i\)\}\-\\mathbf\{v\}\)\. #### Interventional\(−i,i\)\(\-i,i\)block\. Taking the\(−i,i\)\(\-i,i\)block of the interventional Lyapunov equation and using that𝐃𝐓\\mathbf\{D\}\_\{\\mathbf\{T\}\}is diagonal and\(𝐐\(i\)\)−i,i=0\(\\mathbf\{Q\}^\{\(i\)\}\)\_\{\-i,i\}=0yields λii𝐙−i,i\(i\)\+𝝀−i,iZii\(i\)\+𝚲−i,−i𝐙−i,i\(i\)=0\.\\lambda\_\{ii\}\\,\\mathbf\{Z\}^\{\(i\)\}\_\{\-i,i\}\+\\bm\{\\lambda\}\_\{\-i,i\}\\,Z^\{\(i\)\}\_\{ii\}\+\\bm\{\\Lambda\}\_\{\-i,\-i\}\\,\\mathbf\{Z\}^\{\(i\)\}\_\{\-i,i\}=0\.SinceZii\(i\)\>0Z^\{\(i\)\}\_\{ii\}\>0, write 𝒖:=𝐙−i,i\(i\)Zii\(i\)\\bm\{u\}:=\\frac\{\\mathbf\{Z\}^\{\(i\)\}\_\{\-i,i\}\}\{Z^\{\(i\)\}\_\{ii\}\}to obtain 𝝀−i,i\+\(𝚲−i,−i\+λii𝐈\)𝒖=0\.\\bm\{\\lambda\}\_\{\-i,i\}\+\(\\bm\{\\Lambda\}\_\{\-i,\-i\}\+\\lambda\_\{ii\}\\mathbf\{I\}\)\\,\\bm\{u\}=0\.\(35\) #### The𝚵\\bm\{\\Xi\}\-equation on\(−i,−i\)\(\-i,\-i\)\. Let𝚪:=𝐙−𝐙\(i\)\\bm\{\\Gamma\}:=\\mathbf\{Z\}\-\\mathbf\{Z\}^\{\(i\)\}and define the data\-only matrix 𝚵:=𝚪−i,−i−𝒖𝚪i,−i\.\\bm\{\\Xi\}:=\\bm\{\\Gamma\}\_\{\-i,\-i\}\-\\bm\{u\}\\,\\bm\{\\Gamma\}\_\{i,\-i\}\.From the\(−i,−i\)\(\-i,\-i\)block of the difference Lyapunov equation and the identity above, 𝚲−i,−i𝚵\+𝚵⊤𝚲−i,−i⊤=c\(𝐐−𝐐\(i\)\)−i,−i\+λii\(𝒖𝚪i,−i\+𝚪−i,i𝒖⊤\)\.\\bm\{\\Lambda\}\_\{\-i,\-i\}\\,\\bm\{\\Xi\}\+\\bm\{\\Xi\}^\{\\top\}\\bm\{\\Lambda\}\_\{\-i,\-i\}^\{\\top\}=c\\,\(\\mathbf\{Q\}\-\\mathbf\{Q\}^\{\(i\)\}\)\_\{\-i,\-i\}\+\\lambda\_\{ii\}\\,\\Big\(\\bm\{u\}\\,\\bm\{\\Gamma\}\_\{i,\-i\}\+\\bm\{\\Gamma\}\_\{\-i,i\}\\,\\bm\{u\}^\{\\top\}\\Big\)\.\(36\) From equation[34](https://arxiv.org/html/2609.19955#A1.E34)and equation[35](https://arxiv.org/html/2609.19955#A1.E35), we derive offdiag\(𝚲−i,−i𝐒\(i\)\+𝐒\(i\)𝚲−i,−i⊤\)−λiioffdiag\(𝒖𝐙i,−i\(i\)\+𝐙−i,i\(i\)𝒖⊤\)=offdiag\(c𝐐−i,−i\(i\)\),\\operatorname\{offdiag\}\\\!\\Big\(\\bm\{\\Lambda\}\_\{\-i,\-i\}\\,\\mathbf\{S\}^\{\(i\)\}\+\\mathbf\{S\}^\{\(i\)\}\\,\\bm\{\\Lambda\}\_\{\-i,\-i\}^\{\\\!\\top\}\\Big\)\-\\lambda\_\{ii\}\\,\\operatorname\{offdiag\}\\\!\\Big\(\\bm\{u\}\\,\\mathbf\{Z\}^\{\(i\)\}\_\{i,\-i\}\+\\mathbf\{Z\}^\{\(i\)\}\_\{\-i,i\}\\,\\bm\{u\}^\{\\\!\\top\}\\Big\)=\\operatorname\{offdiag\}\\\!\\big\(c\\,\\mathbf\{Q\}^\{\(i\)\}\_\{\-i,\-i\}\\big\),\(37\)with 𝐒\(i\):=𝐙−i,−i\(i\)−𝒖𝐙i,−i\(i\)≻0\.\\mathbf\{S\}^\{\(i\)\}:=\\mathbf\{Z\}^\{\(i\)\}\_\{\-i,\-i\}\-\\bm\{u\}\\,\\mathbf\{Z\}^\{\(i\)\}\_\{i,\-i\}\\succ 0\.Define the linear maps 𝒜S\(𝐗\):=offdiag\(𝐗𝐒\(i\)\+𝐒\(i\)𝐗⊤\),𝒜Ξ\(𝐗\):=𝐗𝚵\+𝚵⊤𝐗⊤\.\\mathcal\{A\}\_\{S\}\(\\mathbf\{X\}\):=\\operatorname\{offdiag\}\(\\mathbf\{X\}\\mathbf\{S\}^\{\(i\)\}\+\\mathbf\{S\}^\{\(i\)\}\\mathbf\{X\}^\{\\\!\\top\}\),\\qquad\\mathcal\{A\}\_\{\\Xi\}\(\\mathbf\{X\}\):=\\mathbf\{X\}\\bm\{\\Xi\}\+\\bm\{\\Xi\}^\{\\\!\\top\}\\mathbf\{X\}^\{\\\!\\top\}\.Equations equation[36](https://arxiv.org/html/2609.19955#A1.E36)and equation[37](https://arxiv.org/html/2609.19955#A1.E37)form a stacked linear system in𝚲−i,−i\\bm\{\\Lambda\}\_\{\-i,\-i\}, \[𝒜S𝒜Ξ\]\(𝚲−i,−i\)=c\[offdiag\(𝐐−i,−i\(i\)\)\(𝐐−𝐐\(i\)\)−i,−i\]\+λii\[offdiag\(𝒖𝐙i,−i\(i\)\+𝐙−i,i\(i\)𝒖⊤\)𝒖𝚪i,−i\+𝚪−i,i𝒖⊤\]\.\\begin\{bmatrix\}\\mathcal\{A\}\_\{S\}\\\\\[2\.0pt\] \\mathcal\{A\}\_\{\\Xi\}\\end\{bmatrix\}\\\!\(\\bm\{\\Lambda\}\_\{\-i,\-i\}\)=c\\,\\begin\{bmatrix\}\\operatorname\{offdiag\}\(\\mathbf\{Q\}^\{\(i\)\}\_\{\-i,\-i\}\)\\\\\[2\.0pt\] \(\\mathbf\{Q\}\-\\mathbf\{Q\}^\{\(i\)\}\)\_\{\-i,\-i\}\\end\{bmatrix\}\+\\lambda\_\{ii\}\\,\\begin\{bmatrix\}\\operatorname\{offdiag\}\\\!\\big\(\\bm\{u\}\\,\\mathbf\{Z\}^\{\(i\)\}\_\{i,\-i\}\+\\mathbf\{Z\}^\{\(i\)\}\_\{\-i,i\}\\,\\bm\{u\}^\{\\\!\\top\}\\big\)\\\\\[2\.0pt\] \\bm\{u\}\\,\\bm\{\\Gamma\}\_\{i,\-i\}\+\\bm\{\\Gamma\}\_\{\-i,i\}\\,\\bm\{u\}^\{\\\!\\top\}\\end\{bmatrix\}\. Since both the true and competing drift matrices are compatible with the fixed graphGG, their relevant blocks belong to the same allowed linear space\. Let𝒱G\\mathcal\{V\}\_\{G\}be the linear space of matrices with the allowed sparsity pattern of𝚲−i,−i\\bm\{\\Lambda\}\_\{\-i,\-i\}\. Define ℒblk\(𝐗\):=\[offdiag\(𝐗𝐒\(i\)\+𝐒\(i\)𝐗⊤\)𝐗𝚵\+𝚵⊤𝐗⊤\],𝐗∈𝒱G\.\\mathcal\{L\}\_\{\\rm blk\}\(\\mathbf\{X\}\):=\\begin\{bmatrix\}\\operatorname\{offdiag\}\(\\mathbf\{X\}\\mathbf\{S\}^\{\(i\)\}\+\\mathbf\{S\}^\{\(i\)\}\\mathbf\{X\}^\{\\top\}\)\\\\\[2\.0pt\] \\mathbf\{X\}\\bm\{\\Xi\}\+\\bm\{\\Xi\}^\{\\top\}\\mathbf\{X\}^\{\\top\}\\end\{bmatrix\},\\qquad\\mathbf\{X\}\\in\\mathcal\{V\}\_\{G\}\.For a nonsingleton target SCC, choose the isolated realization supplied by the theorem’s hypothesis and set its incoming couplings to zero\. Embed it in the full graph by taking the remaining variables to be independent scalar OU processes\. The residual target moments then equal the isolated moments\. For a singleton target,𝒱G\\mathcal\{V\}\_\{G\}has dimension zero, so the stacked operator is injective automatically\. Let𝐋blk\\mathbf\{L\}\_\{\\rm blk\}be any matrix representation ofℒblk\\mathcal\{L\}\_\{\\rm blk\}\. At the point𝚲𝐓𝐏=0\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}=0, the residual target\-block equations reduce to the single\-SCC equations\. Hence, under the single\-SCC nondegeneracy assumptions, the stacked operatorℒblk\\mathcal\{L\}\_\{\\rm blk\}is injective at this point\. Since the entries of a matrix representation ofℒblk\\mathcal\{L\}\_\{\\rm blk\}depend continuously, indeed rationally, on the model parameters, this full\-column\-rank condition is not identically violated and therefore holds generically\. Thus, generically, the stacked linear system uniquely determines𝚲−i,−i\\bm\{\\Lambda\}\_\{\-i,\-i\}fromits right\-hand side\. A*rank certificate*here is one nonzero maximal minor of𝐋blk\\mathbf\{L\}\_\{\\rm blk\}, and the genericity conclusion uses the rational\-witness argument in Appendix[A\.2](https://arxiv.org/html/2609.19955#A1.SS2)\. The single\-SCC assumptions in fact give injectivity on the full reduced matrix space at the isolated witness, and hence on its subspace𝒱G\\mathcal\{V\}\_\{G\}\.Since the right\-hand side is affine inccandλii\\lambda\_\{ii\}, there exist data\-dependent matrices𝐊0,𝐊2\\mathbf\{K\}\_\{0\},\\mathbf\{K\}\_\{2\}such that 𝚲−i,−i=c𝐊0\+λii𝐊2\.\\bm\{\\Lambda\}\_\{\-i,\-i\}=c\\,\\mathbf\{K\}\_\{0\}\+\\lambda\_\{ii\}\\,\\mathbf\{K\}\_\{2\}\.\(38\)For a fixed parent scalecc, this is affine dependence onλii\\lambda\_\{ii\}\. From equation[35](https://arxiv.org/html/2609.19955#A1.E35), 𝝀−i,i=−\(c𝐊0\+λii𝐊2\+λii𝐈\)𝒖\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\lambda\}\_\{\-i,i\}=\-\\big\(c\\mathbf\{K\}\_\{0\}\+\\lambda\_\{ii\}\\mathbf\{K\}\_\{2\}\+\\lambda\_\{ii\}\\mathbf\{I\}\\big\)\\bm\{u\}\.\} Using the observational\(i,−i\)\(i,\-i\)block of the Lyapunov equation, λii𝐙i,−i\+𝝀i,−i𝐙−i,−i\+Zii𝝀−i,i⊤\+𝐙i,−i𝚲−i,−i⊤=c𝐐i,−i\.\{\\color\[rgb\]\{0,0,0\}\\lambda\_\{ii\}\\mathbf\{Z\}\_\{i,\-i\}\+\\bm\{\\lambda\}\_\{i,\-i\}\\mathbf\{Z\}\_\{\-i,\-i\}\+Z\_\{ii\}\\bm\{\\lambda\}\_\{\-i,i\}^\{\\top\}\+\\mathbf\{Z\}\_\{i,\-i\}\\bm\{\\Lambda\}\_\{\-i,\-i\}^\{\\top\}=c\\mathbf\{Q\}\_\{i,\-i\}\.\}Since𝐙−i,−i≻0\\mathbf\{Z\}\_\{\-i,\-i\}\\succ 0, we obtain 𝝀i,−i=\(c𝐐i,−i−λii𝐙i,−i−Zii𝝀−i,i⊤−𝐙i,−i𝚲−i,−i⊤\)𝐙−i,−i−1\.\{\\color\[rgb\]\{0,0,0\}\\bm\{\\lambda\}\_\{i,\-i\}=\\big\(c\\mathbf\{Q\}\_\{i,\-i\}\-\\lambda\_\{ii\}\\mathbf\{Z\}\_\{i,\-i\}\-Z\_\{ii\}\\bm\{\\lambda\}\_\{\-i,i\}^\{\\top\}\-\\mathbf\{Z\}\_\{i,\-i\}\\bm\{\\Lambda\}\_\{\-i,\-i\}^\{\\top\}\\big\)\\mathbf\{Z\}\_\{\-i,\-i\}^\{\-1\}\.\}Substituting the expressions for𝚲−i,−i\\bm\{\\Lambda\}\_\{\-i,\-i\}and𝝀−i,i\\bm\{\\lambda\}\_\{\-i,i\}gives 𝝀i,−i=𝐀0\+λii𝐀1,\{\\color\[rgb\]\{0,0,0\}\\bm\{\\lambda\}\_\{i,\-i\}=\\mathbf\{A\}\_\{0\}\+\\lambda\_\{ii\}\\mathbf\{A\}\_\{1\},\}\(39\)where 𝐀0=c\(𝐐i,−i\+Zii𝒖⊤𝐊0⊤−𝐙i,−i𝐊0⊤\)𝐙−i,−i−1,\{\\color\[rgb\]\{0,0,0\}\\mathbf\{A\}\_\{0\}=c\\Big\(\\mathbf\{Q\}\_\{i,\-i\}\+Z\_\{ii\}\\bm\{u\}^\{\\top\}\\mathbf\{K\}\_\{0\}^\{\\top\}\-\\mathbf\{Z\}\_\{i,\-i\}\\mathbf\{K\}\_\{0\}^\{\\top\}\\Big\)\\mathbf\{Z\}\_\{\-i,\-i\}^\{\-1\},\}and 𝐀1=\(−𝐙i,−i\+Zii𝒖⊤\(𝐈\+𝐊2⊤\)−𝐙i,−i𝐊2⊤\)𝐙−i,−i−1\.\{\\color\[rgb\]\{0,0,0\}\\mathbf\{A\}\_\{1\}=\\Big\(\-\\mathbf\{Z\}\_\{i,\-i\}\+Z\_\{ii\}\\bm\{u\}^\{\\top\}\(\\mathbf\{I\}\+\\mathbf\{K\}\_\{2\}^\{\\top\}\)\-\\mathbf\{Z\}\_\{i,\-i\}\\mathbf\{K\}\_\{2\}^\{\\top\}\\Big\)\\mathbf\{Z\}\_\{\-i,\-i\}^\{\-1\}\.\}Thus𝐀0\\mathbf\{A\}\_\{0\}scales linearly withcc, while𝐊2\\mathbf\{K\}\_\{2\}and𝐀1\\mathbf\{A\}\_\{1\}are data\-dependent and do not scale withcc\. From theii\-th component of the mean difference, let Δμi:=\(𝝁𝐓\|𝐏\)i−\(𝝁𝐓\|𝐏\(i\)\)i,Δvi:=vi\(i\)−vi\.\\Delta\\mu\_\{i\}:=\(\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}\)\_\{i\}\-\(\\bm\{\\mu\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}\)\_\{i\},\\qquad\\Delta v\_\{i\}:=v^\{\(i\)\}\_\{i\}\-v\_\{i\}\.Then λiiΔμi\+𝝀i,−i\(𝝁𝐓\|𝐏\)−i=cΔvi\.\\lambda\_\{ii\}\\,\\Delta\\mu\_\{i\}\+\\bm\{\\lambda\}\_\{i,\-i\}\(\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}\)\_\{\-i\}=c\\,\\Delta v\_\{i\}\.Insert𝝀i,−i=𝐀0\+λii𝐀1\\bm\{\\lambda\}\_\{i,\-i\}=\\mathbf\{A\}\_\{0\}\+\\lambda\_\{ii\}\\mathbf\{A\}\_\{1\}and group the terms to obtain λii\(Δμi\+𝐀1\(𝝁𝐓\|𝐏\)−i\)=cΔvi−𝐀0\(𝝁𝐓\|𝐏\)−i\.\\lambda\_\{ii\}\\big\(\\Delta\\mu\_\{i\}\+\\mathbf\{A\}\_\{1\}\(\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}\)\_\{\-i\}\\big\)=c\\Delta v\_\{i\}\-\\mathbf\{A\}\_\{0\}\(\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}\)\_\{\-i\}\. The denominator below is generically nonzero\. To justify the parameter choice, writeA=𝚲𝐏𝐏A=\\bm\{\\Lambda\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}andS=𝚺𝐏𝐏S=\\bm\{\\Sigma\}\_\{\\mathbf\{P\}\\mathbf\{P\}\}and normalize the parent scale to one for this witness\. The parent Lyapunov equation gives 𝐡𝐏=\(I−𝐃𝐏S−1A−1\)𝐛𝐏=−SA⊤S−1A−1𝐛𝐏\.\\mathbf\{h\}\_\{\\mathbf\{P\}\}=\\bigl\(I\-\\mathbf\{D\}\_\{\\mathbf\{P\}\}S^\{\-1\}A^\{\-1\}\\bigr\)\\mathbf\{b\}\_\{\\mathbf\{P\}\}=\-SA^\{\\top\}S^\{\-1\}A^\{\-1\}\\mathbf\{b\}\_\{\\mathbf\{P\}\}\.This map is invertible\. An allowed incoming edge followed by a simple directed path inside𝐓\\mathbf\{T\}toiigives a stable sparse realization withBij≠0B\_\{ij\}\\neq 0for somej∈𝐏j\\in\\mathbf\{P\}\.777For the sparse path realization, take the parent drift and covariance to beII, with parent diffusion2I2I, and target drift2I−N2I\-N, whereNNhas unit entries along the chosen pathk=k0→⋯→km=ik=k\_\{0\}\\to\\cdots\\to k\_\{m\}=iand is zero elsewhere\. Set𝚲𝐓𝐏=−ηekej⊤\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{P\}\}=\-\\eta e\_\{k\}e\_\{j\}^\{\\top\}withη≠0\\eta\\neq 0and use positive diagonal target diffusion\. These drifts, including hard\-intervened drifts, are triangular after ordering the path vertices and have positive diagonals\. The cross\-covariance equation givesBij=η/3m\+1≠0B\_\{ij\}=\\eta/3^\{m\+1\}\\neq 0\.Both this entry and a full\-column\-rank certificate ofℒblk\\mathcal\{L\}\_\{\\rm blk\}are rational and have nonzero witnesses, which may be different parameter choices\. By Appendix[A\.2](https://arxiv.org/html/2609.19955#A1.SS2), their nonvanishing can be made simultaneous in an arbitrarily small admissible neighborhood of the decoupled rank witness\. Fix such covariance parameters, choose𝐛𝐏\\mathbf\{b\}\_\{\\mathbf\{P\}\}so thatvi≠0v\_\{i\}\\neq 0, and take𝐛𝐓=𝐯\\mathbf\{b\}\_\{\\mathbf\{T\}\}=\\mathbf\{v\}\. Then𝝁𝐓\|𝐏=0\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}=0, whilevi\(i\)=0v\_\{i\}^\{\(i\)\}=0and the interventional mean equation giveλii\(𝝁𝐓\|𝐏\(i\)\)i=vi\\lambda\_\{ii\}\(\\bm\{\\mu\}^\{\(i\)\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}\)\_\{i\}=v\_\{i\}\. Consequently the denominator equals−vi/λii≠0\-v\_\{i\}/\\lambda\_\{ii\}\\neq 0\. The denominator is rational on the full\-rank admissible domain, so its zero set is nongeneric\. Restoring a parent scaleccgives−cvi/λii\-cv\_\{i\}/\\lambda\_\{ii\}at the corresponding witness\.Henceλii\\lambda\_\{ii\}is identified as λii=cΔvi−𝐀0\(𝝁𝐓\|𝐏\)−iΔμi\+𝐀1\(𝝁𝐓\|𝐏\)−i,\\lambda\_\{ii\}=\\frac\{c\\Delta v\_\{i\}\-\\mathbf\{A\}\_\{0\}\(\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}\)\_\{\-i\}\}\{\\Delta\\mu\_\{i\}\+\\mathbf\{A\}\_\{1\}\(\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}\)\_\{\-i\}\},where it is determined up to the same scalecc\. Withλii\\lambda\_\{ii\}identified, the previous equations give 𝚲−i,−i=c𝐊0\+λii𝐊2,𝝀−i,i=−\(c𝐊0\+λii𝐊2\+λii𝐈\)𝒖,𝝀i,−i=𝐀0\+λii𝐀1\.\\bm\{\\Lambda\}\_\{\-i,\-i\}=c\\mathbf\{K\}\_\{0\}\+\\lambda\_\{ii\}\\,\\mathbf\{K\}\_\{2\},\\qquad\\bm\{\\lambda\}\_\{\-i,i\}=\-\\Big\(c\\mathbf\{K\}\_\{0\}\+\\lambda\_\{ii\}\\,\\mathbf\{K\}\_\{2\}\+\\lambda\_\{ii\}\\mathbf\{I\}\\Big\)\\bm\{u\},\\qquad\\bm\{\\lambda\}\_\{i,\-i\}=\\mathbf\{A\}\_\{0\}\+\\lambda\_\{ii\}\\mathbf\{A\}\_\{1\}\.Thus the entire𝚲𝐓𝐓\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}is determined with the same scalingcc\. Finally, 𝐃𝐓=diag\(𝚲𝐓𝐓𝐙\+𝐙𝚲𝐓𝐓⊤−c𝐐\),\{\\color\[rgb\]\{0,0,0\}\\mathbf\{D\}\_\{\\mathbf\{T\}\}=\\mathrm\{diag\}\\big\(\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\mathbf\{Z\}\+\\mathbf\{Z\}\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}^\{\\top\}\-c\\,\\mathbf\{Q\}\\big\),\}and 𝐛𝐓=𝚲𝐓𝐓𝝁𝐓\|𝐏\+c𝐯\.\{\\color\[rgb\]\{0,0,0\}\\mathbf\{b\}\_\{\\mathbf\{T\}\}=\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\}\\bm\{\\mu\}\_\{\\mathbf\{T\}\\mid\\mathbf\{P\}\}\+c\\,\\mathbf\{v\}\.\}Thus\(𝚲𝐓𝐓,𝐃𝐓,𝐛𝐓\)\(\\bm\{\\Lambda\}\_\{\\mathbf\{T\}\\mathbf\{T\}\},\\mathbf\{D\}\_\{\\mathbf\{T\}\},\\mathbf\{b\}\_\{\\mathbf\{T\}\}\)are identified up to the same multiplicative scalecc\. ### A\.8An Example with Multiple Root SCCs Consider an OU process with state indices1,2,3,41,2,3,4, where1,2,31,2,3are root singletons and44is a child singleton\. Let 𝚲=\[λ10000λ20000λ30a1a2a3λ4\],𝐃=diag\(d1,d2,d3,d4\),\\bm\{\\Lambda\}=\\begin\{bmatrix\}\\lambda\_\{1\}&0&0&0\\\\ 0&\\lambda\_\{2\}&0&0\\\\ 0&0&\\lambda\_\{3\}&0\\\\ a\_\{1\}&a\_\{2\}&a\_\{3\}&\\lambda\_\{4\}\\end\{bmatrix\},\\qquad\\mathbf\{D\}=\\operatorname\{diag\}\(d\_\{1\},d\_\{2\},d\_\{3\},d\_\{4\}\),and assume𝐃\\mathbf\{D\}is diagonal and eachλi\>0\\lambda\_\{i\}\>0\. Suppose we observe the steady\-state mean and covariance\(𝝁,𝚺\)\(\\bm\{\\mu\},\\bm\{\\Sigma\}\)and also the interventional moments\(𝝁\(4\),𝚺\(4\)\)\(\\bm\{\\mu\}^\{\(4\)\},\\bm\{\\Sigma\}^\{\(4\)\}\)under an intervention on node44that zeros the entriesaia\_\{i\}s \(please note that interventions on root give the same mean and covariance as the observational ones\)\. For each rooti∈\{1,2,3\}i\\in\\\{1,2,3\\\}, the scalar OU gives μi=biλi,si:=Σii=di2λi\.\\mu\_\{i\}=\\frac\{b\_\{i\}\}\{\\lambda\_\{i\}\},\\qquad s\_\{i\}:=\\Sigma\_\{ii\}=\\frac\{d\_\{i\}\}\{2\\lambda\_\{i\}\}\.Letti:=Σi4t\_\{i\}:=\\Sigma\_\{i4\}denote the observed cross\-covariances with the child\. The off\-diagonal Lyapunov equations yield \(λi\+λ4\)ti\+siai=0⟺ai=−λi\+λ4siti\.\(\\lambda\_\{i\}\+\\lambda\_\{4\}\)t\_\{i\}\+s\_\{i\}a\_\{i\}=0\\quad\\Longleftrightarrow\\quad a\_\{i\}=\-\\frac\{\\lambda\_\{i\}\+\\lambda\_\{4\}\}\{s\_\{i\}\}\\,t\_\{i\}\.\(40\)The child’s observed mean and variance satisfy λ4μ4\+∑i=13aiμi\\displaystyle\\lambda\_\{4\}\\mu\_\{4\}\+\\sum\_\{i=1\}^\{3\}a\_\{i\}\\mu\_\{i\}=b4,\\displaystyle=b\_\{4\},\(41\)2λ4s4\+2∑i=13aiti\\displaystyle 2\\lambda\_\{4\}s\_\{4\}\+2\\sum\_\{i=1\}^\{3\}a\_\{i\}t\_\{i\}=d4\.\\displaystyle=d\_\{4\}\.\(42\)Under the intervention on44\(which zeros theaia\_\{i\}s\), we also observe μ4\(4\)=b4λ4,2λ4s4\(4\)=d4,ti\(4\)=0\.\\mu\_\{4\}^\{\(4\)\}=\\frac\{b\_\{4\}\}\{\\lambda\_\{4\}\},\\qquad 2\\lambda\_\{4\}s\_\{4\}^\{\(4\)\}=d\_\{4\},\\qquad t\_\{i\}^\{\(4\)\}=0\.\(43\) Now, fix any positive\(u1,u2,u3\)\(u\_\{1\},u\_\{2\},u\_\{3\}\)and a positivevv\(to be chosen\)\. Define λi′=uiλi,bi′=uibi,di′=uidi\(i=1,2,3\),λ4′=vλ4,b4′=vb4,d4′=vd4,\\lambda\_\{i\}^\{\\prime\}=u\_\{i\}\\lambda\_\{i\},\\quad b\_\{i\}^\{\\prime\}=u\_\{i\}b\_\{i\},\\quad d\_\{i\}^\{\\prime\}=u\_\{i\}d\_\{i\}\\ \\ \(i=1,2,3\),\\qquad\\lambda\_\{4\}^\{\\prime\}=v\\lambda\_\{4\},\\quad b\_\{4\}^\{\\prime\}=vb\_\{4\},\\quad d\_\{4\}^\{\\prime\}=vd\_\{4\},and chooseai′a\_\{i\}^\{\\prime\}to preserve the observedtit\_\{i\}: \(λi′\+λ4′\)ti\+siai′=0⟺ai′=−uiλi\+vλ4siti\.\(\\lambda\_\{i\}^\{\\prime\}\+\\lambda\_\{4\}^\{\\prime\}\)t\_\{i\}\+s\_\{i\}a\_\{i\}^\{\\prime\}=0\\quad\\Longleftrightarrow\\quad a\_\{i\}^\{\\prime\}=\-\\frac\{u\_\{i\}\\lambda\_\{i\}\+v\\lambda\_\{4\}\}\{s\_\{i\}\}\\,t\_\{i\}\.\(44\)Then for each root,μi′=bi′λi′=μi\\mu\_\{i\}^\{\\prime\}=\\frac\{b\_\{i\}^\{\\prime\}\}\{\\lambda\_\{i\}^\{\\prime\}\}=\\mu\_\{i\}andsi′=di′2λi′=sis\_\{i\}^\{\\prime\}=\\frac\{d\_\{i\}^\{\\prime\}\}\{2\\lambda\_\{i\}^\{\\prime\}\}=s\_\{i\}, while equation[44](https://arxiv.org/html/2609.19955#A1.E44)ensures the sametit\_\{i\}\. For the interventional moments on44, scaling\(λ4,b4,d4\)\(\\lambda\_\{4\},b\_\{4\},d\_\{4\}\)by the common factorvvpreservesμ4\(4\)=b4λ4\\mu\_\{4\}^\{\(4\)\}=\\frac\{b\_\{4\}\}\{\\lambda\_\{4\}\}ands4\(4\)=d42λ4s\_\{4\}^\{\(4\)\}=\\frac\{d\_\{4\}\}\{2\\lambda\_\{4\}\}, andti\(4\)=0t\_\{i\}^\{\(4\)\}=0still holds since theaia\_\{i\}are zeroed under intervention\. Thus all interventional moments match equation[43](https://arxiv.org/html/2609.19955#A1.E43)\. It remains to enforce that the primed parameters also satisfy the child’s*observational*equations equation[41](https://arxiv.org/html/2609.19955#A1.E41)–equation[42](https://arxiv.org/html/2609.19955#A1.E42)\. Subtractingvvtimes the original equation[42](https://arxiv.org/html/2609.19955#A1.E42)from the primed version gives 2λ4′s4−2vλ4s4\+2∑i=13\(ai′−vai\)ti=0⟺∑i=13\(ai′−vai\)ti=0\.2\\lambda\_\{4\}^\{\\prime\}s\_\{4\}\-2v\\lambda\_\{4\}s\_\{4\}\+2\\sum\_\{i=1\}^\{3\}\(a\_\{i\}^\{\\prime\}\-va\_\{i\}\)t\_\{i\}=0\\ \\Longleftrightarrow\\ \\sum\_\{i=1\}^\{3\}\(a\_\{i\}^\{\\prime\}\-va\_\{i\}\)t\_\{i\}=0\.Using equation[40](https://arxiv.org/html/2609.19955#A1.E40)and equation[44](https://arxiv.org/html/2609.19955#A1.E44), ai′−vai=−uiλi\+vλ4siti\+vλi\+λ4siti=\(v−ui\)λisiti\.a\_\{i\}^\{\\prime\}\-va\_\{i\}=\-\\frac\{u\_\{i\}\\lambda\_\{i\}\+v\\lambda\_\{4\}\}\{s\_\{i\}\}t\_\{i\}\+v\\frac\{\\lambda\_\{i\}\+\\lambda\_\{4\}\}\{s\_\{i\}\}t\_\{i\}=\\frac\{\(v\-u\_\{i\}\)\\lambda\_\{i\}\}\{s\_\{i\}\}\\,t\_\{i\}\.Hence ∑i=13\(v−ui\)λiti2si=0⟺v∑i=13αi=∑i=13uiαi,αi:=λiti2si\>0\.\\sum\_\{i=1\}^\{3\}\(v\-u\_\{i\}\)\\,\\lambda\_\{i\}\\,\\frac\{t\_\{i\}^\{2\}\}\{s\_\{i\}\}=0\\quad\\Longleftrightarrow\\quad v\\sum\_\{i=1\}^\{3\}\\alpha\_\{i\}=\\sum\_\{i=1\}^\{3\}u\_\{i\}\\alpha\_\{i\},\\qquad\\alpha\_\{i\}:=\\lambda\_\{i\}\\frac\{t\_\{i\}^\{2\}\}\{s\_\{i\}\}\>0\.\(45\)Thus v=∑iuiαi∑iαi\.v=\\frac\{\\sum\_\{i\}u\_\{i\}\\alpha\_\{i\}\}\{\\sum\_\{i\}\\alpha\_\{i\}\}\.\(46\) Similarly, subtractingvvtimes equation[41](https://arxiv.org/html/2609.19955#A1.E41)from the primed version gives λ4′μ4−vλ4μ4\+∑i=13\(ai′−vai\)μi=0⟺∑i=13\(v−ui\)λitisiμi=0,\\lambda\_\{4\}^\{\\prime\}\\mu\_\{4\}\-v\\lambda\_\{4\}\\mu\_\{4\}\+\\sum\_\{i=1\}^\{3\}\(a\_\{i\}^\{\\prime\}\-va\_\{i\}\)\\mu\_\{i\}=0\\ \\Longleftrightarrow\\ \\sum\_\{i=1\}^\{3\}\(v\-u\_\{i\}\)\\,\\lambda\_\{i\}\\,\\frac\{t\_\{i\}\}\{s\_\{i\}\}\\,\\mu\_\{i\}=0,i\.e\., v∑i=13βi=∑i=13uiβi,βi:=λitisiμi=bitisi\.v\\sum\_\{i=1\}^\{3\}\\beta\_\{i\}=\\sum\_\{i=1\}^\{3\}u\_\{i\}\\beta\_\{i\},\\qquad\\beta\_\{i\}:=\\lambda\_\{i\}\\frac\{t\_\{i\}\}\{s\_\{i\}\}\\,\\mu\_\{i\}=\\frac\{b\_\{i\}t\_\{i\}\}\{s\_\{i\}\}\.\(47\)Thus also v=∑iuiβi∑iβi\.v=\\frac\{\\sum\_\{i\}u\_\{i\}\\beta\_\{i\}\}\{\\sum\_\{i\}\\beta\_\{i\}\}\.\(48\) Equating equation[46](https://arxiv.org/html/2609.19955#A1.E46)and equation[48](https://arxiv.org/html/2609.19955#A1.E48)yields a single linear constraint on\(u1,u2,u3\)\(u\_\{1\},u\_\{2\},u\_\{3\}\): ∑iuiαi∑iαi=∑iuiβi∑iβi⟺∑i=13ui\(αi∑jβj−βi∑jαj\)=0\.\\frac\{\\sum\_\{i\}u\_\{i\}\\alpha\_\{i\}\}\{\\sum\_\{i\}\\alpha\_\{i\}\}=\\frac\{\\sum\_\{i\}u\_\{i\}\\beta\_\{i\}\}\{\\sum\_\{i\}\\beta\_\{i\}\}\\ \\Longleftrightarrow\\ \\sum\_\{i=1\}^\{3\}u\_\{i\}\\Big\(\\alpha\_\{i\}\\sum\_\{j\}\\beta\_\{j\}\-\\beta\_\{i\}\\sum\_\{j\}\\alpha\_\{j\}\\Big\)=0\.\(49\)Letγi:=αi∑jβj−βi∑jαj\\gamma\_\{i\}:=\\alpha\_\{i\}\\sum\_\{j\}\\beta\_\{j\}\-\\beta\_\{i\}\\sum\_\{j\}\\alpha\_\{j\}\. Unless\(γ1,γ2,γ3\)=\(0,0,0\)\(\\gamma\_\{1\},\\gamma\_\{2\},\\gamma\_\{3\}\)=\(0,0,0\), equation[49](https://arxiv.org/html/2609.19955#A1.E49)defines a nontrivial hyperplane inℝ3\\mathbb\{R\}^\{3\}, namely the two\-dimensional solution space of one nonzero linear equation \(Note that the conditionγi=0\\gamma\_\{i\}=0for alliiis nongeneric\)\.Indeed, it requiresβ1α1=β2α2=β3α3\\frac\{\\beta\_\{1\}\}\{\\alpha\_\{1\}\}=\\frac\{\\beta\_\{2\}\}\{\\alpha\_\{2\}\}=\\frac\{\\beta\_\{3\}\}\{\\alpha\_\{3\}\}, i\.e\.,biλiti\\frac\{b\_\{i\}\}\{\\lambda\_\{i\}t\_\{i\}\}is constant acrossii\. Moreover,∑iγi=0\\sum\_\{i\}\\gamma\_\{i\}=0, so the hyperplane contains\(1,1,1\)\(1,1,1\)\. Therefore it contains positive choices of\(u1,u2,u3\)\(u\_\{1\},u\_\{2\},u\_\{3\}\)arbitrarily close to\(1,1,1\)\(1,1,1\), including choices that are not all equal\. This imposes two independent algebraic constraints on the continuous parameters\(bi,λi,ti\)\(b\_\{i\},\\lambda\_\{i\},t\_\{i\}\), hence holds only on a measure\-zero subset of the parameter space\. Therefore, the set of positive\(u1,u2,u3\)\(u\_\{1\},u\_\{2\},u\_\{3\}\)satisfying equation[49](https://arxiv.org/html/2609.19955#A1.E49)has \(at least\) two degrees of freedom\. For any such\(u1,u2,u3\)\(u\_\{1\},u\_\{2\},u\_\{3\}\), takevvfrom equation[46](https://arxiv.org/html/2609.19955#A1.E46)\(which equals equation[48](https://arxiv.org/html/2609.19955#A1.E48)\) and defineai′a\_\{i\}^\{\\prime\}by equation[44](https://arxiv.org/html/2609.19955#A1.E44)\. Unlessu1=u2=u3=vu\_\{1\}=u\_\{2\}=u\_\{3\}=v, the transformation is not a global rescaling of\(𝚲,𝐛,𝐃\)\(\\bm\{\\Lambda\},\\mathbf\{b\},\\mathbf\{D\}\), yet all moments \(observational and interventional under node\-44intervention\) coincide\. Hence, identifiability up to a single global scale fails\. ## Appendix BDetails of Experiments ### B\.1Synthetic Data We generate synthetic datasets from a stable OU model\. For each setup\|ℐ\|∈\{0,2,4,6,8,10\}\|\\mathcal\{I\}\|\\in\\\{0,2,4,6,8,10\\\}\(number of single\-node interventions\) and4040instances, we fixn=10n=10variables and sample: - •Drift matrix𝚲\\bm\{\\Lambda\}: start from zeros; fori≠ji\\neq j, setΛij∼𝒩\(0\.2,σ2\)\\Lambda\_\{ij\}\\sim\\mathcal\{N\}\(0\.2,\\sigma^\{2\}\)with probabilityρ\\rhoand00otherwise \(defaultρ=0\.3\\rho=0\.3,σ=0\.8\\sigma=0\.8\)\. Enforce positive stability by making each row strictly diagonally dominant: Λii←max\(DIAG\_MIN,∑j\|Λij\|\+0\.2\+ui\),ui∼Unif\[0,0\.3\],DIAG\_MIN=0\.8\.\\Lambda\_\{ii\}\\leftarrow\\max\\\!\\bigl\(\\texttt\{DIAG\\\_MIN\},\\,\\sum\_\{j\}\|\\Lambda\_\{ij\}\|\+0\.2\+u\_\{i\}\\bigr\),\\quad u\_\{i\}\\sim\\mathrm\{Unif\}\[0,0\.3\],\\ \\texttt\{DIAG\\\_MIN\}=0\.8\. - •Bias vector and diffusion matrix:𝐛∼Unif\[0\.2,1\.5\]n\\mathbf\{b\}\\sim\\mathrm\{Unif\}\[0\.2,1\.5\]^\{n\};𝐃=diag\(𝐝\)\\mathbf\{D\}=\\operatorname\{diag\}\(\\mathbf\{d\}\)withdk∼Unif\[0\.2,0\.4\]d\_\{k\}\\sim\\mathrm\{Unif\}\[0\.2,0\.4\]\. Observational moments are computed via𝝁=𝚲−1𝐛\\bm\{\\mu\}=\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{b\}and𝚲𝚺\+𝚺𝚲⊤=𝐃\\bm\{\\Lambda\}\\bm\{\\Sigma\}\+\\bm\{\\Sigma\}\\bm\{\\Lambda\}^\{\\\!\\top\}=\\mathbf\{D\}\(continuous\-time Lyapunov\)\. An intervention on nodekkis implemented by zeroing rowkkof𝚲\\bm\{\\Lambda\}and restoring its diagonal entry \(self\-regulation preserved\), leaving𝐛\\mathbf\{b\}and𝐃\\mathbf\{D\}unchanged; intervened moments\(𝝁\(k\),𝚺\(k\)\)\(\\bm\{\\mu\}^\{\(k\)\},\\bm\{\\Sigma\}^\{\(k\)\}\)are computed analogously\. We then draw snapshots from the corresponding Gaussians \(by default40,00040\{,\}000observational samples and20,00020\{,\}000samples per intervention\)\. ### B\.2Real Data #### Library\-size normalization\. For each cellk∈\{1,…,Tc\}k\\in\\\{1,\\dots,T\_\{c\}\\\}in context/interventionccwith raw counts\{xki\}i=1N\\\{x\_\{k\}^\{\\,i\}\\\}\_\{i=1\}^\{N\}\(genesi=1,…,Ni=1,\\dots,N\), define the library sizesk=∑i=1Nxkis\_\{k\}=\\sum\_\{i=1\}^\{N\}x\_\{k\}^\{\\,i\}\. We normalize by fractionsfki=xki/skf\_\{k\}^\{\\,i\}=x\_\{k\}^\{\\,i\}/s\_\{k\}and then scale to a common size x~ki=104fki=104xkisk\.\\tilde\{x\}\_\{k\}^\{\\,i\}\\;=\\;10^\{4\}\\,f\_\{k\}^\{\\,i\}\\;=\\;10^\{4\}\\,\\frac\{x\_\{k\}^\{\\,i\}\}\{s\_\{k\}\}\.All downstream statistics are computed onx~ki\\tilde\{x\}\_\{k\}^\{\\,i\}\. We denote the cell vector𝐱~k=\(x~k1,…,x~kN\)⊤∈ℝN\\tilde\{\\mathbf\{x\}\}\_\{k\}=\(\\tilde\{x\}\_\{k\}^\{\\,1\},\\dots,\\tilde\{x\}\_\{k\}^\{\\,N\}\)^\{\\top\}\\in\\mathbb\{R\}^\{N\}\. #### Empirical moments with shrinkage\. Given a contextccwithTcT\_\{c\}cells\{𝐱~k\}k=1Tc\\\{\\tilde\{\\mathbf\{x\}\}\_\{k\}\\\}\_\{k=1\}^\{T\_\{c\}\}, the empirical mean and \(shrunk\) covariance are 𝝁^c=1Tc∑k=1Tc𝐱~k,𝚺^c=\(1−ηc\)𝚺^c,sample\+ηcdiag\(𝚺^c,sample\),\\hat\{\\bm\{\\mu\}\}\_\{c\}=\\frac\{1\}\{T\_\{c\}\}\\sum\_\{k=1\}^\{T\_\{c\}\}\\tilde\{\\mathbf\{x\}\}\_\{k\},\\qquad\\hat\{\\bm\{\\Sigma\}\}\_\{c\}=\(1\-\\eta\_\{c\}\)\\,\\hat\{\\bm\{\\Sigma\}\}\_\{c,\\mathrm\{sample\}\}\+\\eta\_\{c\}\\,\\operatorname\{diag\}\\\!\\big\(\\hat\{\\bm\{\\Sigma\}\}\_\{c,\\mathrm\{sample\}\}\\big\),where𝚺^c,sample=1Tc−1∑k=1Tc\(𝐱~k−𝝁^c\)\(𝐱~k−𝝁^c\)⊤\\hat\{\\bm\{\\Sigma\}\}\_\{c,\\mathrm\{sample\}\}=\\frac\{1\}\{T\_\{c\}\-1\}\\sum\_\{k=1\}^\{T\_\{c\}\}\(\\tilde\{\\mathbf\{x\}\}\_\{k\}\-\\hat\{\\bm\{\\mu\}\}\_\{c\}\)\(\\tilde\{\\mathbf\{x\}\}\_\{k\}\-\\hat\{\\bm\{\\mu\}\}\_\{c\}\)^\{\\top\}andηc=min\(0\.3,10/\(Tc−1\)\)\\eta\_\{c\}=\\min\\\!\\bigl\(0\.3,\\;10/\(T\_\{c\}\-1\)\\bigr\)\. #### Model\. We fit an OU model with parameters\(𝚲,𝐛,𝑫\)\(\\bm\{\\Lambda\},\\mathbf\{b\},\\bm\{D\}\)on wild\-type and interventional data\. #### Train/test split and metrics\. Similar to previous work\([Rohbeck et al\., 2024](https://arxiv.org/html/2609.19955#bib.bib16)\), interventions are split into train context IDs\{0,…,54\}\\\{0,\\dots,54\\\}and held\-out test context IDs\{55,…,60\}\\\{55,\\dots,60\\\}\. ### B\.3Empirical study of graphical conditions We emphasize that our identifiability result provides*sufficient*conditions\. Thus, there may exist linear SDEs whose parameters are identifiable up to a global scaling even if they do not satisfy the graphical conditions in Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6)\. To empirically study the conservativeness of the sufficient conditions in Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6), we constructed a family of OU models whose condensation graph forms a star over SCCs\. Each SCC consists of a two\-node directed cycle, with a single root SCC andmmleaf SCCs\. For each leaf SCC, we add exactly one directed edge from a randomly chosen node in the root SCC to a randomly chosen node in the leaf SCC\. Interventions are restricted to leaf SCCs and are applied only to the node within each leaf SCC that does not receive the incoming edge from the root SCC\. For a fixed number of interventions\|ℐ\|=k\|\\mathcal\{I\}\|=k, we randomly selectkkdistinct leaf SCCs and apply one hard intervention per selected SCC\. For each configuration, we evaluate whether the model parameters are identifiable up to a global scaling by analyzing the stacked linear system induced by the stationary first\- and second\-moment equations across all observational and interventional settings\. Identifiability is assessed numerically via a spectral\-gap criterion where we compute the singular values of the stacked system and declare identifiability if there is a clear separation between the smallest singular value and the remaining spectrum, corresponding to a one\-dimensional nullspace\. In all experiments, we used a threshold of101010^\{10\}for the ratio between the second\-smallest and smallest singular values\. Figure[2](https://arxiv.org/html/2609.19955#A2.F2)shows the fraction of identifiable instances as a function of the number of interventions\. We observe a sharp transition from non\-identifiability to full identifiability when the number of interventions becomes equal to the number of leaf components \(The dashed vertical lines shows the number of leaf SCCs\)\. This result indicates that the theoretical bound from Theorem[3\.6](https://arxiv.org/html/2609.19955#S3.Thmtheorem6)is close to tight in this setting\. Figure 2:Identifiability versus number of interventions in the star condensation DAG\.Figure 3:The average relative error for recovering𝚲\\mathbf\{\\Lambda\}\(up to some global scaling\) versus the sample size \(observational data\)\. ### B\.4Effect of sample size on relative error To assess the sensitivity of our estimator to the number of steady\-state samples, we conducted an experiment in which we varied the number of observational samples used to estimate\(𝝁^,𝚺^\)\(\\hat\{\\bm\{\\mu\}\},\\hat\{\\mathbf\{\\Sigma\}\}\)\. For each trial, the number of samples per intervention was set to half of the observational sample size\. We evaluated the relative error of recovering𝚲\\bm\{\\Lambda\}for observational sample sizes\{5k,10k,20k,40k\}\\\{5\\mathrm\{k\},10\\mathrm\{k\},20\\mathrm\{k\},40\\mathrm\{k\}\\\}over 30 random instances withn=10n=10\. As shown in Fig\.[3](https://arxiv.org/html/2609.19955#A2.F3), the estimation error decreases markedly when increasing the number of samples from5k5\\mathrm\{k\}to20k20\\mathrm\{k\}and then stabilizes, indicating that the estimator achieves accurate recovery with a moderate number of steady\-state samples\. This behavior is expected, since both empirical means and covariances become sufficiently accurate in this regime, after which additional samples yield diminishing returns\. ### B\.5Finite\-sample SCC recovery We evaluate the SCC\-recovery procedure implied by Theorem[3\.3](https://arxiv.org/html/2609.19955#S3.Thmtheorem3)and Remark[3\.4](https://arxiv.org/html/2609.19955#S3.Thmtheorem4)in a finite\-sample setting\. For each trial, we generate a random OU system withn=10n=10nodes\. The nodes are first partitioned into random SCCs of size between11and33\. The condensation graph is then generated as a random DAG by allowing edges only from earlier SCCs to later SCCs, with edge probability0\.30\.3\. Within each non\-singleton SCC, we include a directed cycle to ensure strong connectivity and add extra within\-SCC edges with probability0\.20\.2\. For each edge in the condensation DAG, we add one randomly selected inter\-SCC edge\. The drift matrix𝚲\\bm\{\\Lambda\}is made positive stable by strict row diagonal dominance: after assigning all off\-diagonal weights, we setΛii=∑j≠i\|Λij\|\+0\.5\\Lambda\_\{ii\}=\\sum\_\{j\\neq i\}\|\\Lambda\_\{ij\}\|\+0\.5\. The input vector is sampled asbi∼Unif\[0\.5,1\.5\]b\_\{i\}\\sim\\mathrm\{Unif\}\[0\.5,1\.5\], and the diagonal diffusion entries are sampled asdi∼Unif\[0\.2,0\.5\]d\_\{i\}\\sim\\mathrm\{Unif\}\[0\.2,0\.5\]\. The observational stationary mean and covariance are computed from𝝁=𝚲−1𝐛\\bm\{\\mu\}=\\bm\{\\Lambda\}^\{\-1\}\\mathbf\{b\}and the Lyapunov equation𝚲𝚺\+𝚺𝚲⊤=𝐃\\bm\{\\Lambda\}\\bm\{\\Sigma\}\+\\bm\{\\Sigma\}\\bm\{\\Lambda\}^\{\\top\}=\\mathbf\{D\}\. We use one intervention target per SCC, chosen as one representative node from that SCC\. For hard interventions, we zero all off\-diagonal entries in the intervened row while keeping the diagonal entry fixed\. For soft interventions, we multiply the off\-diagonal entries in the intervened row by a factorα=0\.4\\alpha=0\.4, again keeping the diagonal entry fixed\. For both observational and interventional settings, samples are drawn from the corresponding stationary Gaussian distributions\. To estimate the response setRi=\{j:μj\(i\)≠μj\}R\_\{i\}=\\\{j:\\mu\_\{j\}^\{\(i\)\}\\neq\\mu\_\{j\}\\\}, we perform coordinate\-wise two\-samplezz\-tests comparing observational and interventional samples\.The estimated response sets are used to reconstruct SCCs as follows: empty response sets are treated as singleton source SCCs; identical non\-empty response sets are grouped together; strict inclusions among response sets recover the SCC\-level partial order; and subtracting descendant response sets recovers the individual SCC node sets\. We repeat this procedure over100100independently generated OU systems\. The quality of the recovered SCC partition is measured by the Adjusted Rand Index \(ARI\), where11indicates exact recovery\. We report results for both hard and soft interventions as a function of sample size\. ### B\.6Experiment with non\-diagonal diffusion In this experiment, the data generator was modified to sample a full positive definite diffusion matrix𝛀\\bm\{\\Omega\}instead of a diagonal𝐃\\mathbf\{D\}\. Specifically, we generated a random matrix𝐀\\mathbf\{A\}, formed𝐀𝐀⊤\\mathbf\{A\}\\mathbf\{A\}^\{\\top\}, normalized its scale, and added a positive diagonal jitter\. The steady\-state covariance was then obtained by solving the full Lyapunov equation𝚲𝚺\+𝚺𝚲⊤=𝛀\\bm\{\\Lambda\}\\bm\{\\Sigma\}\+\\bm\{\\Sigma\}\\bm\{\\Lambda\}^\{\\top\}=\\bm\{\\Omega\}\. Thus, the data\-generating process no longer satisfies the diagonal diffusion assumption\. The learning model was also modified to estimate a full positive definite diffusion matrix\. We parameterized𝛀\\bm\{\\Omega\}as𝛀=𝐌𝐌⊤\+ϵ𝐈\\bm\{\\Omega\}=\\mathbf\{M\}\\mathbf\{M\}^\{\\top\}\+\\epsilon\\mathbf\{I\}, where𝐌\\mathbf\{M\}is a learned lower\-triangular matrix\. Accordingly, the covariance loss was changed from separate diagonal and off\-diagonal residuals to the full matrix residual𝚲\(c\)𝚺^\(c\)\+𝚺^\(c\)𝚲\(c\)⊤−𝛀,\\bm\{\\Lambda\}^\{\(c\)\}\\widehat\{\\bm\{\\Sigma\}\}^\{\(c\)\}\+\\widehat\{\\bm\{\\Sigma\}\}^\{\(c\)\}\\bm\{\\Lambda\}^\{\(c\)\\top\}\-\\bm\{\\Omega\},summed over the observational and interventional contextscc\. These changes ensure that both data generation and optimization use a full diffusion matrix, rather than treating diffusion as diagonal\. ## Appendix CAbout the Spectral/Rank Assumptions ### C\.1Numerical Analysis of Spectral/Rank Assumptions \(a\)Min\. eigenvalue spacing of𝚲\\bm\{\\Lambda\} \(b\)Min\. singular value of𝐋𝑨\\mathbf\{L\}\_\{\\bm\{A\}\} \(c\)Assum\.[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)\(2\)\. \(d\)Assum\.[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)\(3\)\. Figure 4:Numerical probe of spectral/rank assumptions: a\) The histogram of minimum pairwise eigenvalue spacing of𝚲\\bm\{\\Lambda\}\. b\) The histogram of minimum singular of𝐋𝑨\\mathbf\{L\}\_\{\\bm\{A\}\}\. c\) The condition in Assumption[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)\(2\) d\) The condition in Assumption[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)\(3\) \(The values on the x\-axis are inlog10\\log\_\{10\}scale\)We assessed the spectral nondegeneracy assumptions by Monte Carlo analysis on random, strongly connected OU models\. In each of250250trials, we sampled a drift matrix𝚲∈ℝ10×10\\bm\{\\Lambda\}\\in\\mathbb\{R\}^\{10\\times 10\}with off–diagonal pattern drawn i\.i\.d\. at density0\.30\.3\(Gaussian weights, standard deviation0\.80\.8\), and set each diagonal entry strictly larger than theℓ1\\ell\_\{1\}row sum \(ensuring positive stability\)\. Strong connectivity was checked on the directed graph with edgej→ij\\\!\\to\\\!iiffΛij≠0\\Lambda\_\{ij\}\\neq 0\. We drew a diagonal diffusion𝐃\\mathbf\{D\}withdk∼Unif\[0\.2,0\.4\]d\_\{k\}\\sim\\mathrm\{Unif\}\[0\.2,0\.4\], selected an intervention coordinateiiuniformly at random, and formed the intervened drift by zeroing rowii’s off–diagonals \(keepingΛii\\Lambda\_\{ii\}\)\. We then solved the continuous–time Lyapunov equations to obtain the observational and intervened covariances𝚺\\bm\{\\Sigma\}and𝚺\(i\)\\bm\{\\Sigma\}^\{\(i\)\}, formed𝚪:=𝚺−𝚺\(i\)\\bm\{\\Gamma\}:=\\bm\{\\Sigma\}\-\\bm\{\\Sigma\}^\{\(i\)\}, and built the\(−i,−i\)\(\-i,\-i\)Schur blocks𝚵\\bm\{\\Xi\}and𝐙\\mathbf\{Z\}used in the proofs, with𝐀:=𝚵−1𝐙\\mathbf\{A\}:=\\bm\{\\Xi\}^\{\-1\}\\mathbf\{Z\}\. We evaluated four quantities: \(i\) the simple\-spectrum condition for𝚲\\bm\{\\Lambda\}, measured bymink≠ℓ\|λk−λℓ\|\\min\_\{k\\neq\\ell\}\|\\lambda\_\{k\}\-\\lambda\_\{\\ell\}\|; \(ii\) the single\-SCC skew\-kernel condition in Assumption[A\.1](https://arxiv.org/html/2609.19955#A1.Thmtheorem1), measured byσmin\(𝐋𝐀\)\\sigma\_\{\\min\}\(\\mathbf\{L\}\_\{\\mathbf\{A\}\}\), where𝐋𝐀\\mathbf\{L\}\_\{\\mathbf\{A\}\}represents𝐒↦offdiag\(𝐒𝐀−𝐀⊤𝐒\)\\mathbf\{S\}\\mapsto\\operatorname\{offdiag\}\(\\mathbf\{S\}\\mathbf\{A\}\-\\mathbf\{A\}^\{\\top\}\\mathbf\{S\}\)on the skew\-symmetric subspace; \(iii\) the condition in Assumption[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)\(2\) and \(iv\) the condition in Assumption[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)\(3\), measured by\|𝐞i⊤𝚪−1𝚺\(i\)𝐞i\|\|\\mathbf\{e\}\_\{i\}^\{\\top\}\\bm\{\\Gamma\}^\{\-1\}\\bm\{\\Sigma\}^\{\(i\)\}\\mathbf\{e\}\_\{i\}\|\. We report the resulting histograms in Figure[4](https://arxiv.org/html/2609.19955#A3.F4)\. In all250250trials, the four quantities were strictly separated from zero\. This is numerical evidence for the sampled models, not a proof of genericity for every SCC pattern\. ### C\.2About the role of spectral/rank assumptions in the proofs \- Assumption[A\.1](https://arxiv.org/html/2609.19955#A1.Thmtheorem1)holds on a nonempty open subset of the ambient𝐀\\mathbf\{A\}\-space\. For example, if𝐀=diag\(a1,…,am\)\\mathbf\{A\}=\\operatorname\{diag\}\(a\_\{1\},\\dots,a\_\{m\}\)withap≠aqa\_\{p\}\\neq a\_\{q\}forp≠qp\\neq q, then for any skew\-symmetric𝐒\\mathbf\{S\},\(𝐒𝐀−𝐀⊤𝐒\)pq=\(aq−ap\)Spq\(\\mathbf\{S\}\\mathbf\{A\}\-\\mathbf\{A\}^\{\\top\}\\mathbf\{S\}\)\_\{pq\}=\(a\_\{q\}\-a\_\{p\}\)S\_\{pq\}forp≠qp\\neq q\. Henceoffdiag\(𝐒𝐀−𝐀⊤𝐒\)=0\\operatorname\{offdiag\}\(\\mathbf\{S\}\\mathbf\{A\}\-\\mathbf\{A\}^\{\\top\}\\mathbf\{S\}\)=0impliesSpq=0S\_\{pq\}=0for allp≠qp\\neq q, and since𝐒\\mathbf\{S\}is skew\-symmetric, this gives𝐒=0\\mathbf\{S\}=0\. Moreover, this injectivity is stable under sufficiently small off\-diagonal perturbations\. Indeed, write𝐀=𝐃\+𝐄\\mathbf\{A\}=\\mathbf\{D\}\+\\mathbf\{E\}, where𝐃=diag\(a1,…,am\)\\mathbf\{D\}=\\operatorname\{diag\}\(a\_\{1\},\\dots,a\_\{m\}\)andδ:=minp≠q\|ap−aq\|\>0\\delta:=\\min\_\{p\\neq q\}\|a\_\{p\}\-a\_\{q\}\|\>0\. For any skew\-symmetric𝐒\\mathbf\{S\},‖offdiag\(𝐒𝐃−𝐃𝐒\)‖F≥δ‖𝐒‖F\\\|\\operatorname\{offdiag\}\(\\mathbf\{S\}\\mathbf\{D\}\-\\mathbf\{D\}\\mathbf\{S\}\)\\\|\_\{F\}\\geq\\delta\\\|\\mathbf\{S\}\\\|\_\{F\}, whereas‖offdiag\(𝐒𝐄−𝐄⊤𝐒\)‖F≤2‖𝐄‖2‖𝐒‖F\\\|\\operatorname\{offdiag\}\(\\mathbf\{S\}\\mathbf\{E\}\-\\mathbf\{E\}^\{\\top\}\\mathbf\{S\}\)\\\|\_\{F\}\\leq 2\\\|\\mathbf\{E\}\\\|\_\{2\}\\\|\\mathbf\{S\}\\\|\_\{F\}\. Therefore, if2‖𝐄‖2<δ2\\\|\\mathbf\{E\}\\\|\_\{2\}<\\delta, thenoffdiag\(𝐒𝐀−𝐀⊤𝐒\)=0\\operatorname\{offdiag\}\(\\mathbf\{S\}\\mathbf\{A\}\-\\mathbf\{A\}^\{\\top\}\\mathbf\{S\}\)=0forces𝐒=0\\mathbf\{S\}=0\. Thus Assumption[A\.1](https://arxiv.org/html/2609.19955#A1.Thmtheorem1)holds not only for diagonal matrices with separated diagonal entries, but also for all sufficiently small off\-diagonal perturbations of such matrices\. In the model, however,𝐀\\mathbf\{A\}is determined by the induced moments, so we state the condition directly as a checkable spectral condition on the induced operatorℒ𝐀\\mathcal\{L\}\_\{\\mathbf\{A\}\}\. The perturbation argument above suggests a direction for finding a witnesses\. For instance, one may look for parameter regimes in a single SCC with suitably tuned weights, where the moment\-derived matrix𝚵−1𝐙\\bm\{\\Xi\}^\{\-1\}\\mathbf\{Z\}becomes close to a diagonal matrix with separated diagonal entries\. In such a regime, Assumption[A\.1](https://arxiv.org/html/2609.19955#A1.Thmtheorem1)holds\. For small SCCs the condition simplifies\. When\|C\|=2\|C\|=2, we havem=1m=1, so the condition is vacuous\. When\|C\|=3\|C\|=3, we havem=2m=2, and the condition reduces to the scalar requirementA22≠A11A\_\{22\}\\neq A\_\{11\}\. \- Assumption[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)guarantees invertibility of the covariance difference𝚪=𝚺−𝚺\(i\)\\bm\{\\Gamma\}=\\bm\{\\Sigma\}\-\\bm\{\\Sigma\}^\{\(i\)\}\. Its factors have a Popov–Belevitch–Hautus \(PBH\)\([Hespanha and others, 2009](https://arxiv.org/html/2609.19955#bib.bib33)\)interpretation\. The PBH criterion characterizes controllability and observability of a linear system 𝐳˙=𝐌𝐳\+𝐠u,y=𝐡⊤𝐳\.\\dot\{\\mathbf\{z\}\}=\\mathbf\{M\}\\mathbf\{z\}\+\\mathbf\{g\}u,\\qquad y=\\mathbf\{h\}^\{\\top\}\\mathbf\{z\}\.The pair\(𝐌,𝐠\)\(\\mathbf\{M\},\\mathbf\{g\}\)is controllable if and only ifℓ⊤𝐠≠0\\bm\{\\ell\}^\{\\top\}\\mathbf\{g\}\\neq 0for every nonzero left eigenvectorℓ⊤\\bm\{\\ell\}^\{\\top\}of𝐌\\mathbf\{M\}\. Similarly,\(𝐡⊤,𝐌\)\(\\mathbf\{h\}^\{\\top\},\\mathbf\{M\}\)is observable if and only if𝐡⊤𝐫≠0\\mathbf\{h\}^\{\\top\}\\mathbf\{r\}\\neq 0for every nonzero right eigenvector𝐫\\mathbf\{r\}of𝐌\\mathbf\{M\}\. These tests express whether the input can excite every mode and whether the output can detect every mode\. The structural versions ask whether these properties hold generically when the permitted entries of the matrix/input/output pair vary independently\. The graphical criterion for structural controllability requires that every state be reachable from an input and that the bipartite graph of\[𝐌𝐠\]\[\\mathbf\{M\}\\ \\mathbf\{g\}\]admit a matching of sizenn, wherennis the number of states\. The latter means selectingnnpermitted entries in distinct rows and distinct columns\. Structural observability has the dual test: every state must have a directed path to an output, and the bipartite graph of\[𝐌⊤𝐡\]⊤\[\\mathbf\{M\}^\{\\top\}\\ \\mathbf\{h\}\]^\{\\top\}must admit a matching of sizenn\([Lin, 1974](https://arxiv.org/html/2609.19955#bib.bib8)\)\. In our setting, the first factor𝝆⋆,k⊤𝐞i\\bm\{\\rho\}\_\{\\star,k\}^\{\\top\}\\mathbf\{e\}\_\{i\}in\(𝐖⋆\)kℓ\(\\mathbf\{W\}\_\{\\star\}\)\_\{k\\ell\}is the PBH controllability coefficient for\(−𝚲⋆,𝐞i\)\(\-\\bm\{\\Lambda\}\_\{\\star\},\\mathbf\{e\}\_\{i\}\): it determines whether an input acting on the intervened coordinateiican excite drift modekk\. For the second factor, set𝐡=𝚺\(i\)𝐰⋆\\mathbf\{h\}=\\bm\{\\Sigma\}^\{\(i\)\}\\mathbf\{w\}\_\{\\star\}\. Since 𝚲⋆⊤𝝆⋆,ℓ=λ⋆,ℓ𝝆⋆,ℓ,\\bm\{\\Lambda\}\_\{\\star\}^\{\\top\}\\bm\{\\rho\}\_\{\\star,\\ell\}=\\lambda\_\{\\star,\\ell\}\\bm\{\\rho\}\_\{\\star,\\ell\},the coefficient𝐰⋆⊤𝚺\(i\)𝝆⋆,ℓ=𝐡⊤𝝆⋆,ℓ\\mathbf\{w\}\_\{\\star\}^\{\\top\}\\bm\{\\Sigma\}^\{\(i\)\}\\bm\{\\rho\}\_\{\\star,\\ell\}=\\mathbf\{h\}^\{\\top\}\\bm\{\\rho\}\_\{\\star,\\ell\}is the PBH observability coefficient for\(𝐡⊤,−𝚲⋆⊤\)\(\\mathbf\{h\}^\{\\top\},\-\\bm\{\\Lambda\}\_\{\\star\}^\{\\top\}\): it determines whether this covariance\-derived output detects modeℓ\\ellof the transpose dynamics\. Under the assumed simple spectrum, nonvanishing of all the first factors is equivalent to controllability of the first pair, and nonvanishing of all the second factors is equivalent to observability of the second pair\. For these graphical tests, our independently free diagonal drift entries supply a self\-loop at every state, so the matching conditions are automatic\. Within a single SCC, every vertex is reachable from the intervened nodeii\. Hence\(−𝚲⋆,𝐞i\)\(\-\\bm\{\\Lambda\}\_\{\\star\},\\mathbf\{e\}\_\{i\}\)is structurally controllable, and the first PBH factors are generically nonzero under the assumed simple spectrum\. For the transpose dynamics, reverse the drift edges and check that every vertex can reach a vertex connected to the output\. In a single SCC this reachability condition holds for any nonempty output pattern\. It establishes structural observability when the permitted output coefficients are independently free\. These interpretations do not alone prove genericity of Assumption[A\.2](https://arxiv.org/html/2609.19955#A1.Thmtheorem2)\. In particular,𝐡=𝚺\(i\)𝐰⋆\\mathbf\{h\}=\\bm\{\\Sigma\}^\{\(i\)\}\\mathbf\{w\}\_\{\\star\}depends on the model parameters and cannot be varied independently in a structural genericity argument\. We retain the explicit condition0∉σ\(𝐖⋆\+𝐖⋆⊤\)0\\notin\\sigma\(\\mathbf\{W\}\_\{\\star\}\+\\mathbf\{W\}\_\{\\star\}^\{\\top\}\), which guarantees invertibility of𝚪=𝚺−𝚺\(i\)\\bm\{\\Gamma\}=\\bm\{\\Sigma\}\-\\bm\{\\Sigma\}^\{\(i\)\}through𝚪=−𝐑⋆\(𝐖⋆\+𝐖⋆⊤\)𝐑⋆⊤\\bm\{\\Gamma\}=\-\\mathbf\{R\}\_\{\\star\}\(\\mathbf\{W\}\_\{\\star\}\+\\mathbf\{W\}\_\{\\star\}^\{\\top\}\)\\mathbf\{R\}\_\{\\star\}^\{\\top\}\. \- The scalar condition𝐞i⊤𝚪−1𝚺\(i\)𝐞i≠0\\mathbf\{e\}\_\{i\}^\{\\top\}\\bm\{\\Gamma\}^\{\-1\}\\bm\{\\Sigma\}^\{\(i\)\}\\mathbf\{e\}\_\{i\}\\neq 0is used in the proof to pass from invertibility of𝚪:=𝚺−𝚺\(i\)\\bm\{\\Gamma\}:=\\bm\{\\Sigma\}\-\\bm\{\\Sigma\}^\{\(i\)\}to invertibility of the Schur\-type matrix𝚵\\bm\{\\Xi\}\.Consider this condition on the directed\-cycle graph1→2→⋯→n→11\\to 2\\to\\cdots\\to n\\to 1, with intervention on node11\. The intervention removes the edgen→1n\\to 1, leaving the directed chain1→2→⋯→n1\\to 2\\to\\cdots\\to n\. Thus the covariance effect generated bythe intervention can propagate through all coordinates of the block\. By Cramer’s rule, 𝐞1⊤𝚪−1𝚺\(1\)𝐞1=det\[𝚺\(1\)𝐞1,𝚪𝐞2,…,𝚪𝐞n\]det\(𝚪\)\.\\mathbf\{e\}\_\{1\}^\{\\top\}\\bm\{\\Gamma\}^\{\-1\}\\bm\{\\Sigma\}^\{\(1\)\}\\mathbf\{e\}\_\{1\}=\\frac\{\\det\\\!\\big\[\\bm\{\\Sigma\}^\{\(1\)\}\\mathbf\{e\}\_\{1\},\\,\\bm\{\\Gamma\}\\mathbf\{e\}\_\{2\},\\ldots,\\bm\{\\Gamma\}\\mathbf\{e\}\_\{n\}\\big\]\}\{\\det\(\\bm\{\\Gamma\}\)\}\.On the directed cycles, both the numerator and denominator are rational functions of the cycle weights, self\-loops, and diffusion variances\. Since the intervened graph contains a directed chain from the intervened node to all other nodes, the columns appearing in the Cramer numerator are not forced to vanish by the graph structure\. However, this does not by itself prove that the determinant is nonzero\. To prove genericity on directed cycles, or on general SCCs, one would still need to exhibit one admissible parameter value for which the condition holds\. #### Witness for directed cycle When comparing drifts supported on the same directed cycle, one can avoid the auxiliary𝐀=𝚵−1𝐙\\mathbf\{A\}=\\bm\{\\Xi\}^\{\-1\}\\mathbf\{Z\}skew\-kernel argument used in the general single\-SCC proof\. Instead, we work directly with the𝚵\\bm\{\\Xi\}\-equation𝚲Δ,−i,−i𝚵\+𝚵⊤𝚲Δ,−i,−i⊤=0\\bm\{\\Lambda\}\_\{\\Delta,\-i,\-i\}\\bm\{\\Xi\}\+\\bm\{\\Xi\}^\{\\top\}\\bm\{\\Lambda\}\_\{\\Delta,\-i,\-i\}^\{\\top\}=0, which was derived in equation[28](https://arxiv.org/html/2609.19955#A1.E28)\. When the SCC is the directed cycle1→2→⋯→n→11\\to 2\\to\\cdots\\to n\\to 1and the intervention is on node11, the hard intervention removes the edgen→1n\\to 1\. Hence, on the remaining nodes2,…,n2,\\dots,n, the support of𝐗:=𝚲Δ,−1,−1\\mathbf\{X\}:=\\bm\{\\Lambda\}\_\{\\Delta,\-1,\-1\}is lower bidiagonal, corresponding to the directed chain2→3→⋯→n2\\to 3\\to\\cdots\\to n\. For this special support pattern, the adjacent\-minor condition on𝚵\\bm\{\\Xi\}given below forces𝐗=0\\mathbf\{X\}=0\. This gives a direct injectivity argument for cycle\-supported competitors; it does not establish Assumption[A\.1](https://arxiv.org/html/2609.19955#A1.Thmtheorem1)for unrestricted competing drifts\. Consider the directed cycle1→2→⋯→n→11\\to 2\\to\\cdots\\to n\\to 1, with intervention on node11\. After the hard intervention, the edgen→1n\\to 1is removed, and the remaining support on nodes2,…,n2,\\dots,nis the directed chain2→3→⋯→n2\\to 3\\to\\cdots\\to n\. Thus the ambiguity𝐗:=𝚲Δ,−1,−1\\mathbf\{X\}:=\\bm\{\\Lambda\}\_\{\\Delta,\-1,\-1\}has lower\-bidiagonal support in the natural ordering2,…,n2,\\dots,n\. That is, ifm=n−1m=n\-1, reduced indexppcorresponds to original nodep\+1p\+1, andXpq=0unlessp=qorp=q\+1\.X\_\{pq\}=0\\text\{ unless \}p=q\\ \\text\{or\}\\ p=q\+1\. Recall that the proof gives the equation𝐗𝚵\+𝚵⊤𝐗⊤=0\\mathbf\{X\}\\bm\{\\Xi\}\+\\bm\{\\Xi\}^\{\\top\}\\mathbf\{X\}^\{\\top\}=0\. The following lemma gives a direct condition under which this equation forces𝐗=0\\mathbf\{X\}=0\. ###### Lemma C\.1\. Let𝚵∈ℝm×m\\bm\{\\Xi\}\\in\\mathbb\{R\}^\{m\\times m\}be arbitrary, not necessarily symmetric\. Suppose𝐗∈ℝm×m\\mathbf\{X\}\\in\\mathbb\{R\}^\{m\\times m\}is lower bidiagonal:Xpq=0unlessp=qorp=q\+1X\_\{pq\}=0\\qquad\\text\{unless\}\\qquad p=q\\ \\text\{or\}\\ p=q\+1\. If𝐗𝚵\+𝚵⊤𝐗⊤=0\\mathbf\{X\}\\bm\{\\Xi\}\+\\bm\{\\Xi\}^\{\\top\}\\mathbf\{X\}^\{\\top\}=0, andΞ11≠0\\Xi\_\{11\}\\neq 0, and, for everyp=1,…,m−1p=1,\\dots,m\-1,Δp\(𝚵\):=ΞppΞp\+1,p\+1−Ξp,p\+1Ξp\+1,p≠0\\Delta\_\{p\}\(\\bm\{\\Xi\}\):=\\Xi\_\{pp\}\\Xi\_\{p\+1,p\+1\}\-\\Xi\_\{p,p\+1\}\\Xi\_\{p\+1,p\}\\neq 0, then𝐗=0\\mathbf\{X\}=0\. ###### Proof\. Writexp:=Xpp,p=1,…,m,x\_\{p\}:=X\_\{pp\},p=1,\\dots,m,andℓp:=Xp\+1,p,p=1,…,m−1\.\\ell\_\{p\}:=X\_\{p\+1,p\},p=1,\\dots,m\-1\.These are the only possibly nonzero entries of𝐗\\mathbf\{X\}\. Let𝐄:=𝐗𝚵\+𝚵⊤𝐗⊤\.\\mathbf\{E\}:=\\mathbf\{X\}\\bm\{\\Xi\}\+\\bm\{\\Xi\}^\{\\top\}\\mathbf\{X\}^\{\\top\}\.By assumption,𝐄=0\\mathbf\{E\}=0\. The\(1,1\)\(1,1\)\-entry gives0=E11=2x1Ξ110=E\_\{11\}=2x\_\{1\}\\Xi\_\{11\}\. SinceΞ11≠0\\Xi\_\{11\}\\neq 0, we getx1=0x\_\{1\}=0\. Now suppose, inductively, that for somep∈\{1,…,m−1\}p\\in\\\{1,\\dots,m\-1\\\}, x1=⋯=xp=0,ℓ1=⋯=ℓp−1=0\.x\_\{1\}=\\cdots=x\_\{p\}=0,\\qquad\\ell\_\{1\}=\\cdots=\\ell\_\{p\-1\}=0\.Then rowppof𝐗\\mathbf\{X\}is zero, while rowp\+1p\+1has only two possibly nonzero entries,ℓp\\ell\_\{p\}in columnppandxp\+1x\_\{p\+1\}in columnp\+1p\+1\. Hence the\(p,p\+1\)\(p,p\+1\)\-entry of𝐄=0\\mathbf\{E\}=0gives ℓpΞpp\+xp\+1Ξp\+1,p=0,\\ell\_\{p\}\\Xi\_\{pp\}\+x\_\{p\+1\}\\Xi\_\{p\+1,p\}=0,and the\(p\+1,p\+1\)\(p\+1,p\+1\)\-entry gives ℓpΞp,p\+1\+xp\+1Ξp\+1,p\+1=0\.\\ell\_\{p\}\\Xi\_\{p,p\+1\}\+x\_\{p\+1\}\\Xi\_\{p\+1,p\+1\}=0\.Therefore, \[ΞppΞp\+1,pΞp,p\+1Ξp\+1,p\+1\]\[ℓpxp\+1\]=0\.\\begin\{bmatrix\}\\Xi\_\{pp\}&\\Xi\_\{p\+1,p\}\\\\ \\Xi\_\{p,p\+1\}&\\Xi\_\{p\+1,p\+1\}\\end\{bmatrix\}\\begin\{bmatrix\}\\ell\_\{p\}\\\\ x\_\{p\+1\}\\end\{bmatrix\}=0\.Its determinant is Δp\(𝚵\)=ΞppΞp\+1,p\+1−Ξp,p\+1Ξp\+1,p,\\Delta\_\{p\}\(\\bm\{\\Xi\}\)=\\Xi\_\{pp\}\\Xi\_\{p\+1,p\+1\}\-\\Xi\_\{p,p\+1\}\\Xi\_\{p\+1,p\},which is nonzero by assumption\. Thusℓp=0\\ell\_\{p\}=0andxp\+1=0x\_\{p\+1\}=0\. The claim follows by induction\. ∎ For a directed cycle, define κΞ:=Ξ11∏p=1m−1\(ΞppΞp\+1,p\+1−Ξp,p\+1Ξp\+1,p\)\.\\kappa\_\{\\Xi\}:=\\Xi\_\{11\}\\prod\_\{p=1\}^\{m\-1\}\\left\(\\Xi\_\{pp\}\\Xi\_\{p\+1,p\+1\}\-\\Xi\_\{p,p\+1\}\\Xi\_\{p\+1,p\}\\right\)\.By Lemma[C\.1](https://arxiv.org/html/2609.19955#A3.Thmtheorem1), ifκΞ≠0\\kappa\_\{\\Xi\}\\neq 0, then the𝚵\\bm\{\\Xi\}\-equation forces the directed\-chain ambiguity𝚲Δ,−1,−1\\bm\{\\Lambda\}\_\{\\Delta,\-1,\-1\}to vanish\. We now show thatκΞ\\kappa\_\{\\Xi\}is not identically zero on the directed\-cycle parameter space, for every cycle lengthnn\. Since𝚵\\bm\{\\Xi\}is obtained from solutions of Lyapunov equations, its entries are rational functions of the cycle weights, self\-loops, and diffusion variances\. Hence each factor inκΞ\\kappa\_\{\\Xi\}is a rational function of the model parameters\.For these rational witnesses, it suffices to work on the algebraic domain where both Lyapunov operators are nonsingular andΣ11\(1\)≠0\\Sigma^\{\(1\)\}\_\{11\}\\neq 0; positivity is not required at the intermediate witness\. A nonzero rational function cannot vanish on the nonempty open admissible parameter domain, so these witnesses imply the stated generic conclusion for stationary OU models\. Here “algebraic domain” means only that the displayed rational expressions are defined; a Lyapunov solution there need not be a positive covariance\. The witness establishes a nonzero numerator polynomial\. Appendix[A\.2](https://arxiv.org/html/2609.19955#A1.SS2)then transfers its generic nonvanishing to the open admissible OU domain\. We use induction onnn\. The base cases are direct\. Forn=2n=2, take 𝚲=\[2112\],𝐃=diag\(1,2\)\.\\bm\{\\Lambda\}=\\begin\{bmatrix\}2&1\\\\ 1&2\\end\{bmatrix\},\\qquad\\mathbf\{D\}=\\operatorname\{diag\}\(1,2\)\.A direct calculation gives 𝚵=\[364\],\\bm\{\\Xi\}=\\left\[\\frac\{3\}\{64\}\\right\],soκΞ≠0\\kappa\_\{\\Xi\}\\neq 0\. Forn=3n=3, take 𝚲=\[201120012\],𝐃=diag\(1,2,3\)\.\\bm\{\\Lambda\}=\\begin\{bmatrix\}2&0&1\\\\ 1&2&0\\\\ 0&1&2\\end\{bmatrix\},\\qquad\\mathbf\{D\}=\\operatorname\{diag\}\(1,2,3\)\.For the intervention on node11, direct calculation gives 𝚵=\[14032−1806419384−26521504\]\.\\bm\{\\Xi\}=\\begin\{bmatrix\}\\frac\{1\}\{4032\}&\-\\frac\{1\}\{8064\}\\\\\[3\.0pt\] \\frac\{19\}\{384\}&\-\\frac\{265\}\{21504\}\\end\{bmatrix\}\.Thus Ξ11=14032≠0,det\(𝚵\)=8928901376≠0\.\\Xi\_\{11\}=\\frac\{1\}\{4032\}\\neq 0,\\qquad\\det\(\\bm\{\\Xi\}\)=\\frac\{89\}\{28901376\}\\neq 0\.HenceκΞ≠0\\kappa\_\{\\Xi\}\\neq 0for the 3\-cycle\. Assume now that the claim holds for thenn\-cycle\. We argue it for the\(n\+1\)\(n\+1\)\-cycle\. Let Pp\(n\+1\)\(θ\):=Ξpp\(n\+1\)Ξp\+1,p\+1\(n\+1\)−Ξp,p\+1\(n\+1\)Ξp\+1,p\(n\+1\),p=1,…,n−1,P\_\{p\}^\{\(n\+1\)\}\(\\theta\):=\\Xi\_\{pp\}^\{\(n\+1\)\}\\Xi\_\{p\+1,p\+1\}^\{\(n\+1\)\}\-\\Xi\_\{p,p\+1\}^\{\(n\+1\)\}\\Xi\_\{p\+1,p\}^\{\(n\+1\)\},\\qquad p=1,\\dots,n\-1,and letP0\(n\+1\)\(θ\):=Ξ11\(n\+1\)P\_\{0\}^\{\(n\+1\)\}\(\\theta\):=\\Xi\_\{11\}^\{\(n\+1\)\}\. We will show that eachPp\(n\+1\)P\_\{p\}^\{\(n\+1\)\}is not identically zero as a rational function of the\(n\+1\)\(n\+1\)\-cycle parameters\. First considerp=0,…,n−2p=0,\\dots,n\-2, i\.e\., the factors involving only the oldreduced nodes2,…,n2,\\dots,n\. By the induction hypothesis and rational genericity, choose annn\-cycle in this nonsingular algebraic domain whose corresponding𝚵\(n\)\\bm\{\\Xi\}^\{\(n\)\}\-factors are allnonzero\. We now insert noden\+1n\+1as a fast relay betweennnand11\. Let𝚲old∈ℝn×n\\bm\{\\Lambda\}^\{\\rm old\}\\in\\mathbb\{R\}^\{n\\times n\}be the drift of the oldnn\-cycle, and letω:=Λ1nold\\omega:=\\Lambda^\{\\rm old\}\_\{1n\}be the closing\-edge weightn→1n\\to 1\. Define𝐀:=𝚲old−ω𝐞1𝐞n⊤\\mathbf\{A\}:=\\bm\{\\Lambda\}^\{\\rm old\}\-\\omega\\mathbf\{e\}\_\{1\}\\mathbf\{e\}\_\{n\}^\{\\top\}, so that𝐀\\mathbf\{A\}is the old drift with the closing edge removed\. Write𝐱:=\(x1,…,xn\)⊤,z:=xn\+1\\mathbf\{x\}:=\(x\_\{1\},\\dots,x\_\{n\}\)^\{\\top\},z:=x\_\{n\+1\}\.For parametersα,c\>0\\alpha,c\>0andβ∈ℝ\\beta\\in\\mathbb\{R\}satisfyingαβc=ω\\frac\{\\alpha\\beta\}\{c\}=\\omega, define the relay drift 𝚲ε=\[𝐀β𝐞1−αε−1𝐞n⊤cε−1\]\.\\bm\{\\Lambda\}\_\{\\varepsilon\}=\\begin\{bmatrix\}\\mathbf\{A\}&\\beta\\mathbf\{e\}\_\{1\}\\\\\[2\.0pt\] \-\\alpha\\varepsilon^\{\-1\}\\mathbf\{e\}\_\{n\}^\{\\top\}&c\\varepsilon^\{\-1\}\\end\{bmatrix\}\.The relay coordinate has drift−cεz\+αεxn=−cε\(z−αcxn\)\-\\frac\{c\}\{\\varepsilon\}z\+\\frac\{\\alpha\}\{\\varepsilon\}x\_\{n\}=\-\\frac\{c\}\{\\varepsilon\}\\left\(z\-\\frac\{\\alpha\}\{c\}x\_\{n\}\\right\), so asε→0\\varepsilon\\to 0, the relay enforcesz≈αcxnz\\approx\\frac\{\\alpha\}\{c\}x\_\{n\}\. Therefore the edgen\+1→1n\+1\\to 1contributes−βz≈−αβcxn=−ωxn\-\\beta z\\approx\-\\frac\{\\alpha\\beta\}\{c\}x\_\{n\}=\-\\omega x\_\{n\}, which is exactly the contribution of the old closing edgen→1n\\to 1\. We verify this at the covariance level\. Let 𝚺ε=\[𝐒ε𝐫ε𝐫ε⊤sε\]\\bm\{\\Sigma\}\_\{\\varepsilon\}=\\begin\{bmatrix\}\\mathbf\{S\}\_\{\\varepsilon\}&\\mathbf\{r\}\_\{\\varepsilon\}\\\\ \\mathbf\{r\}\_\{\\varepsilon\}^\{\\top\}&s\_\{\\varepsilon\}\\end\{bmatrix\}be the solution of the Lyapunov equation for\(𝐱,z\)\(\\mathbf\{x\},z\), with𝐒ε\\mathbf\{S\}\_\{\\varepsilon\}the old\-variable block and𝐫ε\\mathbf\{r\}\_\{\\varepsilon\}the cross block\. Keep the diagonal diffusion parameters fixed asε→0\\varepsilon\\to 0, with block form 𝐃ε=\[𝐃x00dz\]\.\\mathbf\{D\}\_\{\\varepsilon\}=\\begin\{bmatrix\}\\mathbf\{D\}\_\{x\}&0\\\\ 0&d\_\{z\}\\end\{bmatrix\}\.The\(𝐱,z\)\(\\mathbf\{x\},z\)\-block of the Lyapunov equation gives 𝐀𝐫ε\+β𝐞1sε−αε−1𝐒ε𝐞n\+cε−1𝐫ε=0\.\\mathbf\{A\}\\mathbf\{r\}\_\{\\varepsilon\}\+\\beta\\mathbf\{e\}\_\{1\}s\_\{\\varepsilon\}\-\\alpha\\varepsilon^\{\-1\}\\mathbf\{S\}\_\{\\varepsilon\}\\mathbf\{e\}\_\{n\}\+c\\varepsilon^\{\-1\}\\mathbf\{r\}\_\{\\varepsilon\}=0\.Multiplying byε\\varepsilonand rearranging, 𝐫ε=αc𝐒ε𝐞n−εc\(𝐀𝐫ε\+β𝐞1sε\)\.\\mathbf\{r\}\_\{\\varepsilon\}=\\frac\{\\alpha\}\{c\}\\mathbf\{S\}\_\{\\varepsilon\}\\mathbf\{e\}\_\{n\}\-\\frac\{\\varepsilon\}\{c\}\\left\(\\mathbf\{A\}\\mathbf\{r\}\_\{\\varepsilon\}\+\\beta\\mathbf\{e\}\_\{1\}s\_\{\\varepsilon\}\\right\)\.To justify boundedness, also multiply the\(z,z\)\(z,z\)\-equation byε\\varepsilon, obtaining 2csε−2α𝐞n⊤𝐫ε=εdz\.2cs\_\{\\varepsilon\}\-2\\alpha\\mathbf\{e\}\_\{n\}^\{\\top\}\\mathbf\{r\}\_\{\\varepsilon\}=\\varepsilon d\_\{z\}\.Atε=0\\varepsilon=0, the scaled cross and scalar equations give𝐫=\(α/c\)𝐒𝐞n\\mathbf\{r\}=\(\\alpha/c\)\\mathbf\{S\}\\mathbf\{e\}\_\{n\}ands=\(α/c\)2𝐞n⊤𝐒𝐞ns=\(\\alpha/c\)^\{2\}\\mathbf\{e\}\_\{n\}^\{\\top\}\\mathbf\{S\}\\mathbf\{e\}\_\{n\}\. Substitution into the old\-variable equation gives exactly the old Lyapunov equation\. Its nonsingularity therefore makes the full scaled linear system nonsingular at zero: its homogeneous system has only the zero solution\. In general, ifK\(ε\)K\(\\varepsilon\)is continuous andK\(0\)K\(0\)is invertible, thenK\(ε\)−1K\(\\varepsilon\)^\{\-1\}is continuous near zero, as follows from the adjugate formula and continuity of the nonzero determinant\. Applied here to the matrix of the scaled linear equations, with its continuous right\-hand side, this gives a bounded solution near zero\. The invertible matrix in this argument is the equation operator, not the limiting covariance\. Forε≠0\\varepsilon\\neq 0, scaling the equations is reversible\. Hence 𝐫ε=αc𝐒ε𝐞n\+O\(ε\)\.\\mathbf\{r\}\_\{\\varepsilon\}=\\frac\{\\alpha\}\{c\}\\mathbf\{S\}\_\{\\varepsilon\}\\mathbf\{e\}\_\{n\}\+O\(\\varepsilon\)\.HereRε=O\(ε\)R\_\{\\varepsilon\}=O\(\\varepsilon\)means‖Rε‖≤Cε\\\|R\_\{\\varepsilon\}\\\|\\leq C\\varepsilonfor sufficiently small positiveε\\varepsilon, with the other parameters fixed\. We may use any fixed norm on this finite\-dimensional space\. The\(𝐱,𝐱\)\(\\mathbf\{x\},\\mathbf\{x\}\)\-block of the Lyapunov equation is 𝐀𝐒ε\+𝐒ε𝐀⊤\+β𝐞1𝐫ε⊤\+β𝐫ε𝐞1⊤=𝐃x\.\\mathbf\{A\}\\mathbf\{S\}\_\{\\varepsilon\}\+\\mathbf\{S\}\_\{\\varepsilon\}\\mathbf\{A\}^\{\\top\}\+\\beta\\mathbf\{e\}\_\{1\}\\mathbf\{r\}\_\{\\varepsilon\}^\{\\top\}\+\\beta\\mathbf\{r\}\_\{\\varepsilon\}\\mathbf\{e\}\_\{1\}^\{\\top\}=\\mathbf\{D\}\_\{x\}\.Substituting the previous relation gives \(𝐀\+αβc𝐞1𝐞n⊤\)𝐒ε\+𝐒ε\(𝐀\+αβc𝐞1𝐞n⊤\)⊤=𝐃x\+O\(ε\)\.\\left\(\\mathbf\{A\}\+\\frac\{\\alpha\\beta\}\{c\}\\mathbf\{e\}\_\{1\}\\mathbf\{e\}\_\{n\}^\{\\top\}\\right\)\\mathbf\{S\}\_\{\\varepsilon\}\+\\mathbf\{S\}\_\{\\varepsilon\}\\left\(\\mathbf\{A\}\+\\frac\{\\alpha\\beta\}\{c\}\\mathbf\{e\}\_\{1\}\\mathbf\{e\}\_\{n\}^\{\\top\}\\right\)^\{\\top\}=\\mathbf\{D\}\_\{x\}\+O\(\\varepsilon\)\.Sinceαβ/c=ω\\alpha\\beta/c=\\omega, the drift in parentheses is 𝐀\+ω𝐞1𝐞n⊤=𝚲old\.\\mathbf\{A\}\+\\omega\\mathbf\{e\}\_\{1\}\\mathbf\{e\}\_\{n\}^\{\\top\}=\\bm\{\\Lambda\}^\{\\rm old\}\.Therefore, by uniqueness and continuity of Lyapunov solutions, 𝐒ε⟶𝚺oldasε→0\.\\mathbf\{S\}\_\{\\varepsilon\}\\longrightarrow\\bm\{\\Sigma\}^\{\\rm old\}\\qquad\\text\{as \}\\varepsilon\\to 0\. The same argument applies to the intervened covariance\. Under the intervention on node11, the edgen\+1→1n\+1\\to 1is removed, so the blockβ𝐞1\\beta\\mathbf\{e\}\_\{1\}is set to zero\. Thus the effective closing edge also disappears, matching the intervened oldnn\-cycle\. Hence the observational and interventional covariance blocks on the old variables converge to those of the oldnn\-cycle\. Consequently, the corresponding old block of𝚵ε\\bm\{\\Xi\}\_\{\\varepsilon\}converges to the old𝚵\(n\)\\bm\{\\Xi\}^\{\(n\)\}\-matrix\. Therefore, forp=0,…,n−2p=0,\\dots,n\-2, the factorsPp\(n\+1\)P\_\{p\}^\{\(n\+1\)\}converge to the nonzero factors of the oldnn\-cycle\. Hence eachPp\(n\+1\)P\_\{p\}^\{\(n\+1\)\},p=0,…,n−2p=0,\\dots,n\-2, is not identically zero as a rational function of the\(n\+1\)\(n\+1\)\-cycle parameters\. It remains to show that the last factor Pn−1\(n\+1\)=Ξn−1,n−1\(n\+1\)Ξn,n\(n\+1\)−Ξn−1,n\(n\+1\)Ξn,n−1\(n\+1\)P\_\{n\-1\}^\{\(n\+1\)\}=\\Xi\_\{n\-1,n\-1\}^\{\(n\+1\)\}\\Xi\_\{n,n\}^\{\(n\+1\)\}\-\\Xi\_\{n\-1,n\}^\{\(n\+1\)\}\\Xi\_\{n,n\-1\}^\{\(n\+1\)\}is not identically zero\. Start from the displayed 3\-cycle on the retained nodes\{1,n,n\+1\}\\\{1,n,n\+1\\\}, whose adjacent minor is nonzero\. Insert the intermediate nodes2,…,n−12,\\dots,n\-1one at a time along the edge from11tonn\. The same one\-relay calculation applies to any subdivided edge: removing the incoming edge at node11commutes with this construction\. Indeed, the inserted relays subdivide the outgoing path from11tonn, leaving the incoming edgen\+1→1n\+1\\to 1untouched\. Cutting that edge before or after insertion gives the same model, so the covariance convergence applies to both observational and interventional equations\. At each insertion, choose a sufficiently small positive relay time after fixing the preceding parameters\. Continuity preserves the nonzero minor on the two retained nodes\{n,n\+1\}\\\{n,n\+1\\\}, together with the two Lyapunov nonsingularity conditions and the nonzero interventional variance at node11\. After the finitely many insertions, the graph is the\(n\+1\)\(n\+1\)\-cycle and that retained minor is exactlyPn−1\(n\+1\)P\_\{n\-1\}^\{\(n\+1\)\}: nodes\{n,n\+1\}\\\{n,n\+1\\\}occupy reduced positions\{n−1,n\}\\\{n\-1,n\\\}\. Thus this final rational factor is not identically zero\. We have shown that every factorPp\(n\+1\)P\_\{p\}^\{\(n\+1\)\},p=0,…,n−1p=0,\\dots,n\-1, is a nonzero rational function\. Since the product of finitely many nonzero rational functions is again nonzero, we conclude that P0\(n\+1\)P1\(n\+1\)⋯Pn−1\(n\+1\)P\_\{0\}^\{\(n\+1\)\}P\_\{1\}^\{\(n\+1\)\}\\cdots P\_\{n\-1\}^\{\(n\+1\)\}is not identically zero\. Hence there exists an admissible\(n\+1\)\(n\+1\)\-cycle parameter value for which all adjacent𝚵\\bm\{\\Xi\}\-minors are nonzero\. This completes the induction\. Consequently, for every directed cycle lengthnn, the adjacent\-𝚵\\bm\{\\Xi\}condition holds generically on the directed\-cycle parameter subfamily\. By Lemma[C\.1](https://arxiv.org/html/2609.19955#A3.Thmtheorem1), the equation 𝚲Δ,−1,−1𝚵\+𝚵⊤𝚲Δ,−1,−1⊤=0\\bm\{\\Lambda\}\_\{\\Delta,\-1,\-1\}\\bm\{\\Xi\}\+\\bm\{\\Xi\}^\{\\top\}\\bm\{\\Lambda\}\_\{\\Delta,\-1,\-1\}^\{\\top\}=0has no nonzero solution with directed\-chain support, generically on this subfamily\.
Similar Articles
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
This paper introduces a novel adaptive scheduler for steering discrete diffusion language models using sparse autoencoders, demonstrating that targeting interventions based on when specific attributes commit improves control quality and strength over uniform methods.
Simultaneous inference of environmental and interaction forces in collective dynamics
This paper proposes a nonparametric variational learning framework to simultaneously infer interaction kernels and environmental forces in collective dynamics, validated on benchmark models with a model-selection procedure.
A Bayesian Filtering Approach for Learning Lagrangian Dynamics from Noisy Measurements
This paper presents a Bayesian filtering approach to learn Lagrangian dynamics from partial, noisy measurements by parameterizing kinetic and potential energies with neural networks and jointly estimating states and parameters via maximum likelihood.
State-Space NTK Collapse Near Bifurcations
This paper develops a local theory of gradient descent near bifurcations in dynamical models, showing that the state-space neural tangent kernel collapses to a rank-one operator that dominates learning dynamics, making optimization effectively low-dimensional and predictable from normal forms.
On the Identifiability of Mixed Ordinal and Exponential Family Causal DAGs under Linear Parametric Models
This paper establishes identifiability conditions for causal DAGs with mixed ordinal and exponential family nodes under linear parametric models, proving that edge orientations can be determined from the joint distribution alone with sufficient categories and support points.