Concept Modulation Models: A Unified Framework for Identifiability and Extrapolation
Summary
This paper introduces concept modulation models (CMMs), a unified framework for identifiability and extrapolation in conditional generative models. It shows that feature agreement on observed attributes induces constraints through attribute potentials, enabling algebraic extrapolation criteria that recover and generalize existing results.
View Cached Full Text
Cached at: 06/18/26, 05:43 AM
# Concept Modulation Models: A Unified Framework for Identifiability and Extrapolation
Source: [https://arxiv.org/html/2606.18509](https://arxiv.org/html/2606.18509)
Soheun YiCorrespondence: Soheun Yi,soheuny@andrew\.cmu\.edu\.Department of Statistics and Data Science, Carnegie Mellon UniversityChandler SquiresMachine Learning Department, Carnegie Mellon UniversityPradeep RavikumarMachine Learning Department, Carnegie Mellon University
###### Abstract
Reliable generalization in conditional latent variable models requires understanding both identifiability and extrapolation: how observed variation across attributes determines latent structure, and how that structure determines distributions at unseen attributes\. However, existing identifiability and extrapolation guarantees are largely model\-specific, with separate analyses in nonlinear ICA, causal representation learning, perturbation modeling, and related conditional latent variable models\. We introduce*concept modulation models \(CMMs\)*, an attribute\-indexed class of conditional generative models with structureA→Λ→C→XA\\to\{\\Lambda\}\\to C\\to X, where attributes select modulators, modulators induce latent concept laws, and concepts generate observed features\. CMMs lift transition\-based identifiability to conditional settings by showing that feature agreement on observed attributes induces a latent concept transition constrained by the CMM class\. We express these constraints through*attribute potentials*, log\-density ratios between attribute\-conditioned concept laws, separating the generic lifting step from model\-specific rigidity arguments\. The same potentials control extrapolation: agreement at unseen attributes holds exactly when the transported attribute\-potential identities extend to those attributes\. This yields algebraic extrapolation criteria, identifies the common potential\-based proof objects behind several existing identifiability and extrapolation results, and, when combined with the model\-specific rigidity arguments in those works, recovers their stated conclusions\.
## 1Introduction
Identifiability in latent\-variable representation learning asks when latent structure is not merely useful for prediction, but uniquely determined by the variation available in data, at least up to well\-characterized ambiguities\(Hyvarinenet al\.,[2019](https://arxiv.org/html/2606.18509#bib.bib4); Khemakhemet al\.,[2020a](https://arxiv.org/html/2606.18509#bib.bib2); Schölkopfet al\.,[2021](https://arxiv.org/html/2606.18509#bib.bib1)\)\. This question is central to reliable generalization because data from observed conditions alone need not determine behavior under unseen ones: two models may agree on all observed conditions while disagreeing off\-support\(D’Amouret al\.,[2022](https://arxiv.org/html/2606.18509#bib.bib3)\)\. An identifiability guarantee offers a principled route around this obstacle: if the latent structure governing how attributes affect observations is identifiable, then recovering it can help justify generalization beyond the observed training conditions\.
In this paper, we consider two problems central to this route in conditional generative models, where an observed featureXXvaries with an attributeA∈𝒜A\\in\\mathcal\{A\}, such as an environment, intervention, perturbation, or conditioning input\.*Identifiability*asks what aspects of the latent structure are determined, up to allowable ambiguities, by the conditional feature distributionsp\(x∣a\)p\(x\\mid a\)observed at attributesa∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, where𝒜o⊆𝒜\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\subseteq\\mathcal\{A\}may be a very small subset of𝒜\\mathcal\{A\}\.*Extrapolation*asks when agreement on the observed conditional distributions, together with the structural assumptions of the model class, forces agreement onp\(x∣a′\)p\(x\\mid a^\{\\prime\}\)at unseen attributesa′∈𝒜∖𝒜oa^\{\\prime\}\\in\\mathcal\{A\}\\setminus\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\.
Existing answers to these questions have largely been developed on a per\-model basis\. Nonlinear ICA and identifiable VAEs use auxiliary or conditioning variables to identify latent sources\(Hyvarinenet al\.,[2019](https://arxiv.org/html/2606.18509#bib.bib4); Khemakhemet al\.,[2020a](https://arxiv.org/html/2606.18509#bib.bib2)\)\. Causal representation learning uses environments or interventions to recover latent causal variables and causal structure\(Ahujaet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib5); Squireset al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib6); von Kügelgenet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib8); Varıcıet al\.,[2024a](https://arxiv.org/html/2606.18509#bib.bib33),[2025](https://arxiv.org/html/2606.18509#bib.bib9)\)\. Perturbation models study when learned latent responses can extrapolate to unseen perturbations\(von Kügelgenet al\.,[2025](https://arxiv.org/html/2606.18509#bib.bib13)\)\. Although these results share a common reliance on structured variation across attributes, their assumptions and conclusions are usually stated in model\-specific terms\.
Several recent works provide unifying perspectives, but their scope is different from ours\. For example,Khemakhemet al\.\([2020a](https://arxiv.org/html/2606.18509#bib.bib2)\)unify VAEs and nonlinear ICA through condition\-dependent latent priors, whileYaoet al\.\([2025](https://arxiv.org/html/2606.18509#bib.bib58)\)unify causal representation learning through an invariance principle\.Reizingeret al\.\([2025](https://arxiv.org/html/2606.18509#bib.bib14)\)study identifiability through exchangeable mechanisms or related structural assumptions\. These frameworks clarify important families of identifiability results, but they primarily address when latent representations are identifiable within particular structural regimes\. Our goal is complementary: we address a wide variety of structural regimes at once by exposing the contrastive objects that underlie several identifiability results, and we use these objects to derive conditions under for extrapolation\. We provide a more detailed survey and comparison in[App\.A](https://arxiv.org/html/2606.18509#A1)\.
To give a common formulation of these different structural regimes, we introduce a shared structural form for attribute\-conditioned generation:A→Λ→C→XA\\to\{\\Lambda\}\\to C\\to X\. Here, the attributeAAindexes a latent modulatorΛ\{\\Lambda\}, the modulator specifies a latent concept distribution, and the conceptCCgenerates the observed featureXX\. We call such models*concept modulation models*\(CMMs\)\. The modulator separates attribute\-specific indexing from the shared rule that maps modulators to concept distributions, thereby tying conditional concept laws together across attributes\.
Our analysis builds on the transition\-based identifiability perspective ofSquires and Ravikumar \([2026](https://arxiv.org/html/2606.18509#bib.bib16)\), which reduces latent identifiability to characterizing concept\-space transitions compatible with a model class\. CMMs are a conditional lift of this perspective: for each fixed attribute, a CMM induces a latent concept generative model, while the CMM class ties the family of attribute\-conditioned concept laws through a shared modulation mechanism\. This attribute\-indexed structure raises an extrapolation question absent from the unconditional setting: when does agreement of feature distributions on observed attributes force agreement at unseen attributes?
Our characterization of feature equivalence is expressed through the log\-density ratiologp\(c∣a\)−logp\(c∣a0\)\\log p\(c\\mid a\)\-\\log p\(c\\mid a\_\{0\}\), which we call the*attribute potential*\. Attribute potentials remove terms shared across attributes, isolating how the latent concept law changes withAA\. We show that feature\-equivalent CMMs are related by a latent transition that preserves the observed attribute potentials, so identifiability reduces to determining which transitions remain compatible with the model class\.
The same object also characterizes extrapolation\. Once agreement on observed attributes has been lifted to a latent transition, extrapolation is equivalent to the corresponding transported attribute\-potential identities holding not only on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, but also at the unseen attributes of interest\. These conditions recover recent guarantees for causal representation learning and perturbation modeling, and also yield new interaction\-based extrapolation guarantees for structured attribute spaces\.
Our contributions are as follows:
- •Concept modulation models:We introduce CMMs, a conditional generative framework with graphical structureA→Λ→C→XA\\to\{\\Lambda\}\\to C\\to X\([Definitions1](https://arxiv.org/html/2606.18509#Thmdefinition1)and[2](https://arxiv.org/html/2606.18509#Thmdefinition2)\)\. This separates attribute\-specific indexing from the shared modulation mechanism and captures several weakly supervised latent\-variable settings\.
- •Conditional transition\-based identifiability:Building on the transition\-based framework ofSquires and Ravikumar \([2026](https://arxiv.org/html/2606.18509#bib.bib16)\), we prove a conditional lifting theorem for CMMs: feature equivalence on observed attributes induces a latent concept transition compatible with both components of the model class \([Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)\)\. We then characterize the concept\-side compatibility condition through preservation of attribute potentials \([Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)\)\.
- •Extrapolation via attribute potentials:We prove that, once observed feature agreement has been lifted to a latent transition, agreement at unseen attributes is equivalent to extension of the transported attribute\-potential identities \([Thm\.3](https://arxiv.org/html/2606.18509#Thmtheorem3)\)\. This yields algebraic extrapolation criteria, recovers perturbation extrapolation guarantees as a special case, and gives interaction\-based extrapolation results for structured attribute spaces\.
## 2A unifying framework: Concept modulation models
### 2\.1Preliminaries and notation
Throughout, all spaces are standard Borel spaces\. For measuresμ\\mu,μ′\\mu^\{\\prime\}defined on a space𝒰\\mathcal\{U\}, we writeμ≪μ′\\mu\\ll\\mu^\{\\prime\}ifμ\\muis absolutely continuous with respect toμ′\\mu^\{\\prime\}, i\.e\.,μ′\(B\)=0⟹μ\(B\)=0\\mu^\{\\prime\}\(B\)=0\\implies\\mu\(B\)=0, and writeμ∼μ′\\mu\\sim\\mu^\{\\prime\}ifμ≪μ′\\mu\\ll\\mu^\{\\prime\}andμ′≪μ\\mu^\{\\prime\}\\ll\\mu\. We say that a condition holdsμ\\mu\-a\.e\. foru∈𝒰u\\in\\mathcal\{U\}if the set of pointsuufor which the condition fails hasμ\\mu\-measure zero\. We writeμ⊗μ′\\mu\\otimes\\mu^\{\\prime\}for the product measure ofμ\\mu,μ′\\mu^\{\\prime\}\.
For a space𝒰\\mathcal\{U\}, let𝒫\(𝒰\)\\mathcal\{P\}\(\\mathcal\{U\}\)denote the set of probability measures on𝒰\\mathcal\{U\}\. For spaces𝒰,𝒱\\mathcal\{U\},\\mathcal\{V\}, letℱ\(𝒰→𝒱\)\\mathcal\{F\}\(\\mathcal\{U\}\\to\\mathcal\{V\}\)denote the set of measurable maps from𝒰\\mathcal\{U\}to𝒱\\mathcal\{V\}, and let𝒦mk\(𝒰→𝒱\)\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{U\}\\to\\mathcal\{V\}\)denote the set of Markov kernels from𝒰\\mathcal\{U\}to𝒱\\mathcal\{V\}\. For measurable𝒰′⊆𝒰\\mathcal\{U\}^\{\\prime\}\\subseteq\\mathcal\{U\}, we denote by𝐊𝒰′\{\\mathbf\{K\}\}\_\{\\mathcal\{U\}^\{\\prime\}\}the restriction of𝐊\{\\mathbf\{K\}\}to𝒰′\\mathcal\{U\}^\{\\prime\}\. Given a measurable mapf∈ℱ\(𝒰→𝒱\)f\\in\\mathcal\{F\}\(\\mathcal\{U\}\\to\\mathcal\{V\}\), we let𝐓f:𝒫\(𝒰\)→𝒫\(𝒱\)\\mathbf\{T\}\_\{f\}\\colon\\mathcal\{P\}\(\\mathcal\{U\}\)\\to\\mathcal\{P\}\(\\mathcal\{V\}\)denote the associated pushforward operator, i\.e\.,\(𝐓fμ\)\(B\):=μ\(f−1\[B\]\)\(\\mathbf\{T\}\_\{f\}\\mu\)\(B\)\\mathrel\{:=\}\\mu\(f^\{\-1\}\[B\]\)for allμ∈𝒫\(𝒰\)\\mu\\in\\mathcal\{P\}\(\\mathcal\{U\}\)and measurableB⊆𝒱B\\subseteq\\mathcal\{V\}\. Given a Markov kernel𝐊∈𝒦mk\(𝒰→𝒱\)\{\\mathbf\{K\}\}\\in\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{U\}\\to\\mathcal\{V\}\), we define its pushforward of measureμ∈𝒫\(𝒰\)\\mu\\in\\mathcal\{P\}\(\\mathcal\{U\}\)by\(𝐊μ\)\(B\):=∫𝒰𝐊\(B∣u\)μ\(du\)\(\{\\mathbf\{K\}\}\\mu\)\(B\)\\mathrel\{:=\}\\int\_\{\\mathcal\{U\}\}\{\\mathbf\{K\}\}\(B\\mid u\)\\,\\mu\(du\)\.
As a running example throughout this paper, we consider attribute\-conditioned distributions induced by the process
p\(c∣a\)∝q\(c\)exp\{⟨f\(a\),c⟩\},X=g\(C\),p\(c\\mid a\)\\propto q\(c\)\\exp\\\{\\langle f\(a\),c\\rangle\\\},\\qquad X=g\(C\),\(1\)whereffandggare unknown\. Here, the attributeaais mapped to the modulatorλ=f\(a\)\{\\lambda\}=f\(a\), the modulatorλ\{\\lambda\}changes the latent concept law through an exponential\-family tiltq\(c\)exp\{⟨⋅,c⟩\}q\(c\)\\exp\\\{\\langle\\cdot,c\\rangle\\\}, and the concept is mapped to the observation bygg\.
### 2\.2Concept modulation models
A concept modulation model \(CMM\) factors a conditional generative model throughA→Λ→C→XA\\to\{\\Lambda\}\\to C\\to X, whereA∈𝒜A\\in\\mathcal\{A\}is an*attribute*,Λ∈ℒ\{\\Lambda\}\\in\{\\mathscr\{L\}\}is a*modulator*,C∈𝒞C\\in\\mathcal\{C\}is a latent*concept*, andX∈𝒳X\\in\\mathcal\{X\}is an observed*feature*\. Intuitively, the attribute selects a modulator, the modulator specifies a distribution over concepts, and the concept is mapped to the feature space\. Here,𝒜\\mathcal\{A\},ℒ\{\\mathscr\{L\}\},𝒞\\mathcal\{C\}, and𝒳\\mathcal\{X\}are standard Borel spaces, referred to as the attribute, modulator, concept, and feature spaces, respectively\. In the running example \([1](https://arxiv.org/html/2606.18509#S2.E1)\), theA→ΛA\\to\{\\Lambda\}stage corresponds to the mapff, theΛ→C\{\\Lambda\}\\to Cstage corresponds to the exponential tilting, and theC→XC\\to Xstage corresponds to the observation mapgg\.
We now formalize this generative structure at the level of Markov kernels\.
###### Definition 1\.
A*concept modulation model*𝖬\\mathsf\{M\}on\(𝒜,ℒ,𝒞,𝒳\)\(\\mathcal\{A\},\{\\mathscr\{L\}\},\\mathcal\{C\},\\mathcal\{X\}\)is a tuple𝖬=\(𝐐,𝐁,𝐊\)\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\), where𝐐∈𝒦mk\(𝒜→ℒ\)\{\\mathbf\{Q\}\}\\in\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{A\}\\to\{\\mathscr\{L\}\}\)is the*indexing kernel*,𝐁∈𝒦mk\(ℒ→𝒞\)\{\\mathbf\{B\}\}\\in\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\{\\mathscr\{L\}\}\\to\\mathcal\{C\}\)is the*concept modulation kernel*, and𝐊∈𝒦mk\(𝒞→𝒳\)\{\\mathbf\{K\}\}\\in\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{C\}\\to\\mathcal\{X\}\)is the*mixing kernel*\.
AAΛ\{\\Lambda\}CCXX𝐐\{\\mathbf\{Q\}\}𝐁\{\\mathbf\{B\}\}𝐊\{\\mathbf\{K\}\}\(a\) Graphical model of CMM\(b\) Valid and invalid concept transitionsτv\\tau\_\{v\}andτb\\tau\_\{b\}\(c\) Valid and invalid mixing transitionsτv\\tau\_\{v\}andτb\\tau\_\{b\}\(d\) Transition\-intersection criterion \([Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)\)Figure 1:Overview of CMMs and transition constraints\.\(a\)The observed attributeAAinfluences a modulation variableΛ\{\\Lambda\}, which modulates the latent conceptCC, which generates the observationXX\.\(b,c\)Valid concept and mixing transitions are concept\-space transformations that can be absorbed into the corresponding model class\. Arrows in \(b\) denote equivalence after restriction to𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, while arrows in \(c\) read𝐊=𝜈𝐊′𝐓τv=𝜈𝐊′′𝐓τb\{\\mathbf\{K\}\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}^\{\\prime\}\\mathbf\{T\}\_\{\\tau\_\{v\}\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}^\{\\prime\\prime\}\\mathbf\{T\}\_\{\\tau\_\{b\}\}\.\(d\)The transition\-intersection criterion selects transitions valid on both sides; if their intersection is contained in𝔊\\mathfrak\{G\}, the CMM is identifiable up to∼𝔊\\sim\_\{\\mathfrak\{G\}\}\([Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)\)\.We denote by𝐏𝖬:=𝐊𝐁𝐐∈𝒦mk\(𝒜→𝒳\)\{\\mathbf\{P\}\}^\{\\mathsf\{M\}\}\\mathrel\{:=\}\{\\mathbf\{K\}\}\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}\\in\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{A\}\\to\\mathcal\{X\}\)the*feature generation kernel*induced by𝖬\\mathsf\{M\}\. A graphical representation of[Definition1](https://arxiv.org/html/2606.18509#Thmdefinition1)is shown in[Figure1](https://arxiv.org/html/2606.18509#S2.F1): the indexing kernel𝐐\{\\mathbf\{Q\}\}specifies how attributes index modulators, while the fixed concept modulation kernel𝐁\{\\mathbf\{B\}\}specifies how each modulator induces a concept distribution\. A CMM*class*specifies candidate models with a shared concept\-modulation mechanism\. We therefore fix the kernel𝐁\{\\mathbf\{B\}\}, which maps modulators to concept distributions, while allowing the indexing kernel𝐐\{\\mathbf\{Q\}\}and the mixing kernel𝐊\{\\mathbf\{K\}\}to vary\.
###### Definition 2\.
A*concept modulation model class \(CMM class\)*is a product class
ℳ=𝒬×\{𝐁\}×𝒦,where𝒬⊆𝒦mk\(𝒜→ℒ\)and𝒦⊆𝒦mk\(𝒞→𝒳\)\\mathcal\{M\}=\\mathcal\{Q\}\\times\\\{\{\\mathbf\{B\}\}\\\}\\times\\mathcal\{K\},\\quad\\text\{where\}\\quad\\mathcal\{Q\}\\subseteq\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{A\}\\to\{\\mathscr\{L\}\}\)\\quad\\text\{and\}\\quad\\mathcal\{K\}\\subseteq\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{C\}\\to\\mathcal\{X\}\)so that all models inℳ\\mathcal\{M\}share the same concept modulation kernel𝐁\{\\mathbf\{B\}\}\. Two models𝖬=\(𝐐,𝐁,𝐊\)\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\)and𝖬′=\(𝐐′,𝐁,𝐊′\)\\mathsf\{M\}^\{\\prime\}=\(\{\\mathbf\{Q\}\}^\{\\prime\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}^\{\\prime\}\)in the CMM classℳ\\mathcal\{M\}are*feature equivalent on𝒜*o*⊆𝒜\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\subseteq\\mathcal\{A\}*if𝐏𝒜o𝖬=𝐏𝒜o𝖬′\{\\mathbf\{P\}\}^\{\\mathsf\{M\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\}=\{\\mathbf\{P\}\}^\{\\mathsf\{M\}^\{\\prime\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\}, equivalently, if𝐊𝐁𝐐\(⋅∣a\)=𝐊′𝐁𝐐′\(⋅∣a\)\{\\mathbf\{K\}\}\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\{\\mathbf\{K\}\}^\{\\prime\}\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)for alla∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\.
Givenℳ=𝒬×\{𝐁\}×𝒦\\mathcal\{M\}=\\mathcal\{Q\}\\times\\\{\{\\mathbf\{B\}\}\\\}\\times\\mathcal\{K\}, we denote the induced concept\-kernel class by𝐁𝒬:=\{𝐁𝐐\|𝐐∈𝒬\}⊆𝒦mk\(𝒜→𝒞\)\{\\mathbf\{B\}\}\\mathcal\{Q\}\\mathrel\{:=\}\\\{\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}\\;\|\\;\{\\mathbf\{Q\}\}\\in\\mathcal\{Q\}\\\}\\subseteq\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{A\}\\to\\mathcal\{C\}\)\. When𝐁\{\\mathbf\{B\}\}is clear from context, we write𝐐¯\\bar\{\\mathbf\{Q\}\}and𝒬¯\\bar\{\\mathcal\{Q\}\}for𝐁𝐐\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}and𝐁𝒬\{\\mathbf\{B\}\}\\mathcal\{Q\}, respectively; throughout, any notation𝐐¯\\bar\{\{\\mathbf\{Q\}\}\}or𝒬¯\\bar\{\\mathcal\{Q\}\}is understood relative to this fixed𝐁\{\\mathbf\{B\}\}\.
As we will see, many existing identifiability results, including nonlinear ICA\(Hyvarinenet al\.,[2019](https://arxiv.org/html/2606.18509#bib.bib4); Khemakhemet al\.,[2020a](https://arxiv.org/html/2606.18509#bib.bib2),[b](https://arxiv.org/html/2606.18509#bib.bib17)\), causal representation learning\(Squireset al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib6); Buchholzet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib18); von Kügelgenet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib8); Varıcıet al\.,[2025](https://arxiv.org/html/2606.18509#bib.bib9)\), and other works\(von Kügelgenet al\.,[2025](https://arxiv.org/html/2606.18509#bib.bib13); Schmidtet al\.,[2025](https://arxiv.org/html/2606.18509#bib.bib19)\)can be expressed in this class\-level framework\.[Tables2](https://arxiv.org/html/2606.18509#A6.T2)and[3](https://arxiv.org/html/2606.18509#A6.T3)in[App\.F](https://arxiv.org/html/2606.18509#A6)illustrate corresponding variable\-level and operator\-level dictionaries for selected results\.
We close this section by showing how the running example fits into the CMM framework\.
###### Example 1\.
For the running example in \([1](https://arxiv.org/html/2606.18509#S2.E1)\), take𝒜=ℝm\\mathcal\{A\}=\\mathbb\{R\}^\{m\},ℒ=ℝk\{\\mathscr\{L\}\}=\\mathbb\{R\}^\{k\},𝒞=ℝk\\mathcal\{C\}=\\mathbb\{R\}^\{k\}, and𝒳=ℝd\\mathcal\{X\}=\\mathbb\{R\}^\{d\}\. Then
𝐐=𝐓f,𝐁\(dc∣λ\)=q\(c\)exp\(⟨λ,c⟩\)Z\(λ\)dc,𝐊=𝐓g,\{\\mathbf\{Q\}\}=\\mathbf\{T\}\_\{f\},\\qquad\{\\mathbf\{B\}\}\(dc\\mid\{\\lambda\}\)=\\frac\{q\(c\)\\exp\(\\langle\{\\lambda\},c\\rangle\)\}\{Z\(\{\\lambda\}\)\}\\,dc,\\qquad\{\\mathbf\{K\}\}=\\mathbf\{T\}\_\{g\},whereq\(c\)\>0q\(c\)\>0for allc∈𝒞c\\in\\mathcal\{C\}andZ\(λ\)=∫q\(c\)exp\(⟨λ,c⟩\)𝑑c<∞Z\(\{\\lambda\}\)=\\int q\(c\)\\exp\(\\langle\{\\lambda\},c\\rangle\)\\,dc<\\infty\. A corresponding CMM class isℳ=𝒬×\{𝐁\}×𝒦\\mathcal\{M\}=\\mathcal\{Q\}\\times\\\{\{\\mathbf\{B\}\}\\\}\\times\\mathcal\{K\}, with𝒬=\{𝐓f\|f∈ℱ\}\\mathcal\{Q\}=\\\{\\mathbf\{T\}\_\{f\}\\;\|\\;f\\in\\mathcal\{F\}\\\}and𝒦=\{𝐓g\|g∈𝒢\}\\mathcal\{K\}=\\\{\\mathbf\{T\}\_\{g\}\\;\|\\;g\\in\\mathcal\{G\}\\\}, whereℱ⊆ℱ\(𝒜→ℒ\)\\mathcal\{F\}\\subseteq\\mathcal\{F\}\(\\mathcal\{A\}\\to\{\\mathscr\{L\}\}\)and𝒢⊆ℱ\(𝒞→𝒳\)\\mathcal\{G\}\\subseteq\\mathcal\{F\}\(\\mathcal\{C\}\\to\\mathcal\{X\}\)are classes of functions, with𝒢\\mathcal\{G\}restricted to injective smooth maps\.
## 3Conditional transition\-based identifiability and attribute potentials
Given a CMM class, we now show how identifiability can be recovered through a common transition\-intersection argument, generalizingSquires and Ravikumar \([2026](https://arxiv.org/html/2606.18509#bib.bib16)\)to the conditional generative setting\. The key idea is to view feature equivalence as evidence for a latent transitionτ\\tauon concept space\. For CMMs, the transition must not take the model outside the model class: after adjusting forτ\\tau, the resulting mixing kernel and attribute\-conditioned concept laws must still be realizable in the model class\. Identifiability up to a transition group𝔊\\mathfrak\{G\}then follows when every transition satisfying these compatibility requirements lies in𝔊\\mathfrak\{G\}, as formalized in[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)\. The compatibility condition for the attribute\-conditioned concept laws can then be characterized by log\-density ratios, as in[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)\.
### 3\.1A transition\-intersection criterion
We first fix the reference\-measure setting in which latent transitions are defined\. Letν\{\\nu\}be aσ\\sigma\-finite reference measure on𝒞\\mathcal\{C\}, and restrict attention to concept laws dominated byν\{\\nu\}\. For kernels𝐊,𝐊′∈𝒦mk\(𝒞→𝒳\)\{\\mathbf\{K\}\},\{\\mathbf\{K\}\}^\{\\prime\}\\in\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{C\}\\to\\mathcal\{X\}\), write𝐊=𝜈𝐊′\{\\mathbf\{K\}\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}^\{\\prime\}if𝐊μ=𝐊′μ\{\\mathbf\{K\}\}\\mu=\{\\mathbf\{K\}\}^\{\\prime\}\\mufor allμ≪ν\\mu\\ll\{\\nu\}\. A*latent concept transition*is a measurable automorphism of concept space, defined moduloν\{\\nu\}\-null sets:
Autν\(𝒞\):=\{τ:𝒞→𝒞\|τis invertible up toν\-null sets,𝐓τν∼ν,and𝐓τ−1ν∼ν\}\.\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\)\\mathrel\{:=\}\\\{\\tau\\colon\\mathcal\{C\}\\to\\mathcal\{C\}\\;\|\\;\\tau\\text\{ is invertible up to \}\{\\nu\}\\text\{\-null sets, \}\\mathbf\{T\}\_\{\\tau\}\{\\nu\}\\sim\{\\nu\},\\ \\text\{and \}\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\{\\nu\}\\sim\{\\nu\}\\\}\.Thus, elements ofAutν\(𝒞\)\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\)are precisely the concept\-space transformations that preserve theν\{\\nu\}\-null sets in both directions\. For example, when𝒞=ℝd\\mathcal\{C\}=\\mathbb\{R\}^\{d\}andν\{\\nu\}is Lebesgue measure, this class contains the usual diffeomorphisms\.
We next impose the structural condition that allows feature equivalence to be lifted to such a transition\. Following the transition\-based perspective ofSquires and Ravikumar \([2026](https://arxiv.org/html/2606.18509#bib.bib16)\), we require the mixing class to factor through an injective embedding of concept distributions\.
###### Definition 3\(ν\{\\nu\}\-Blackwell reducible mixing class\)\.
A mixing class𝒦⊆𝒦mk\(𝒞→𝒳\)\\mathcal\{K\}\\subseteq\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{C\}\\to\\mathcal\{X\}\)is*ν\{\\nu\}\-Blackwell reducible*if there exist a standard Borel space𝒳~\{\\widetilde\{\\mathcal\{X\}\}\}, a class𝒢\\mathcal\{G\}of bimeasurable embeddingsg:𝒞→𝒳~g\\colon\\mathcal\{C\}\\to\{\\widetilde\{\\mathcal\{X\}\}\}, and a shared kernel𝐊~∈𝒦mk\(𝒳~→𝒳\)\\widetilde\{\{\\mathbf\{K\}\}\}\\in\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\{\\widetilde\{\\mathcal\{X\}\}\}\\to\\mathcal\{X\}\)such that every𝐊∈𝒦\{\\mathbf\{K\}\}\\in\\mathcal\{K\}admits someg∈𝒢g\\in\\mathcal\{G\}with𝐊=𝜈𝐊~𝐓g\{\\mathbf\{K\}\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g\}, and the mapξ↦𝐊~ξ\\xi\\mapsto\\widetilde\{\{\\mathbf\{K\}\}\}\\xiis injective on\{𝐓gμ\|g∈𝒢,μ≪ν\}\\\{\\mathbf\{T\}\_\{g\}\\mu\\;\|\\;g\\in\\mathcal\{G\},\\ \\mu\\ll\{\\nu\}\\\}\.
The injectivity of the shared kernel𝐊~\\widetilde\{\{\\mathbf\{K\}\}\}means that if two distributions on the intermediate space𝒳~\{\\widetilde\{\\mathcal\{X\}\}\}give the same observed feature law after applying𝐊~\\widetilde\{\{\\mathbf\{K\}\}\}, then those distributions on𝒳~\{\\widetilde\{\\mathcal\{X\}\}\}must already be equal\. Thus, after writing𝐊=𝜈𝐊~𝐓g\{\\mathbf\{K\}\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g\}, the relevant concept distribution is𝐓gμ\\mathbf\{T\}\_\{g\}\\muwith the concept lawμ≪ν\\mu\\ll\{\\nu\}\. This condition is what lets equality of feature laws lift to equality ofgg\-pushed concept laws, and ultimately produces a transitionτ∈Autν\(𝒞\)\\tau\\in\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\)in[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)\.
Intuitively, a transition is valid if its effect can be absorbed without leaving the model class\. On the mixing side, this means changing only the mixing kernel, so that𝐊′=𝜈𝐊𝐓τ\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}for some𝐊′∈𝒦\{\\mathbf\{K\}\}^\{\\prime\}\\in\\mathcal\{K\}\. On the concept side, this means changing only the induced concept kernel, so that𝐐¯\(⋅∣a\)=𝐓τ𝐐¯′\(⋅∣a\)\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)for alla∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}and some𝐐¯′∈𝒬¯\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\\in\\bar\{\\mathcal\{Q\}\}\.
###### Definition 4\(Valid transition sets\)\.
For𝐐¯∈𝒬¯\\bar\{\\mathbf\{Q\}\}\\in\\bar\{\\mathcal\{Q\}\}, define the*valid concept transitions*by
𝒯𝐐¯ν\(𝒬¯;𝒜o\):=\{τ∈Autν\(𝒞\)\|∃𝐐¯′∈𝒬¯such that𝐐¯\(⋅∣a\)=𝐓τ𝐐¯′\(⋅∣a\)for alla∈𝒜o\}\.\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\\mathrel\{:=\}\\\{\\tau\\in\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\)\\;\|\\;\\exists\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\\in\\bar\{\\mathcal\{Q\}\}\\text\{ such that \}\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)\\ \\text\{ for all \}a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\\}\.For𝐊∈𝒦\{\\mathbf\{K\}\}\\in\\mathcal\{K\}, define the*valid mixing transitions*by
𝒯𝐊ν\(𝒦\):=\{τ∈Autν\(𝒞\)\|∃𝐊′∈𝒦such that𝐊′=𝜈𝐊𝐓τ\}\.\\mathcal\{T\}\_\{\{\\mathbf\{K\}\}\}^\{\{\\nu\}\}\(\\mathcal\{K\}\)\\mathrel\{:=\}\\\{\\tau\\in\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\)\\;\|\\;\\exists\{\\mathbf\{K\}\}^\{\\prime\}\\in\\mathcal\{K\}\\text\{ such that \}\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}\\\}\.
Panels \(b\) and \(c\) of[Figure1](https://arxiv.org/html/2606.18509#S2.F1)illustrate the valid transition sets\. We adopt the notion of Blackwell equivalence proposed bySquires and Ravikumar \([2026](https://arxiv.org/html/2606.18509#bib.bib16)\), extending the notion defined in the unconditional setup to the attribute\-conditional setup\.
###### Definition 5\(Blackwell equivalence, attribute\-conditional version\)\.
When two models𝖬=\(𝐐,𝐁,𝐊\)\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\)and𝖬′=\(𝐐′,𝐁,𝐊′\)\\mathsf\{M\}^\{\\prime\}=\(\{\\mathbf\{Q\}\}^\{\\prime\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}^\{\\prime\}\)satisfy𝐊′=𝜈𝐊𝐓τ\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}and𝐐¯\(⋅∣a\)=𝐓τ𝐐¯′\(⋅∣a\)\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)for alla∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}with𝐐¯=𝐁𝐐\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}and𝐐¯′=𝐁𝐐′\\bar\{\\mathbf\{Q\}\}^\{\\prime\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}^\{\\prime\}, we say𝖬\\mathsf\{M\}and𝖬′\\mathsf\{M\}^\{\\prime\}are*Blackwell equivalent viaτ\\tau*\. Whenτ∈𝔊\\tau\\in\\mathfrak\{G\}for a transition group𝔊\\mathfrak\{G\}, we write𝖬∼𝔊𝖬′\\mathsf\{M\}\\sim\_\{\\mathfrak\{G\}\}\\mathsf\{M\}^\{\\prime\}on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\.
Given that𝔊\\mathfrak\{G\}is a group,∼𝔊\\sim\_\{\\mathfrak\{G\}\}is an equivalence relation; see[Sec\.C\.1](https://arxiv.org/html/2606.18509#A3.SS1)\. We next impose a common\-support condition to keep concept transitions and density ratios well defined\. An analogous condition is assumed bySquires and Ravikumar \([2026](https://arxiv.org/html/2606.18509#bib.bib16)\), and it holds for the main continuous examples below\.
###### Definition 6\(Commonν\{\\nu\}\-support\)\.
We say that the induced concept\-kernel class𝒬¯=𝐁𝒬\\bar\{\\mathcal\{Q\}\}=\{\\mathbf\{B\}\}\\mathcal\{Q\}has*commonν\{\\nu\}\-support*if, for every𝐐¯∈𝒬¯\\bar\{\\mathbf\{Q\}\}\\in\\bar\{\\mathcal\{Q\}\}and everya∈𝒜a\\in\\mathcal\{A\}, it holds that𝐐¯\(⋅∣a\)≪ν\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)\\ll\{\\nu\}with densityp𝐐¯:=d𝐐¯\(⋅∣a\)/dνp\_\{\\bar\{\\mathbf\{Q\}\}\}\\mathrel\{:=\}d\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)/d\{\\nu\}satisfying0<p𝐐¯\(c∣a\)<∞0<p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid a\)<\\inftyforν\{\\nu\}\-almost everycc\.
The following theorem is a CMM\-level analogue of the transition\-intersection theorem ofSquires and Ravikumar \([2026](https://arxiv.org/html/2606.18509#bib.bib16)\), visualized in panel \(d\) of[Figure1](https://arxiv.org/html/2606.18509#S2.F1)\. Its CMM\-specific content is that feature equivalence induces a transition that simultaneously transports the*entire*observed set of attribute\-conditioned concept laws, rather than a single distribution\.[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)later gives the density\-level form of this concept\-side compatibility\.
###### Theorem 1\(CMM\-lift of the intersection theorem\)\.
Letℳ=𝒬×\{𝐁\}×𝒦\\mathcal\{M\}=\\mathcal\{Q\}\\times\\\{\{\\mathbf\{B\}\}\\\}\\times\\mathcal\{K\}be a CMM class, and let𝒜o≠∅\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\neq\\varnothing\. Assume that𝒦\\mathcal\{K\}isν\{\\nu\}\-Blackwell reducible and that the induced concept\-kernel class𝒬¯=𝐁𝒬\\bar\{\\mathcal\{Q\}\}=\{\\mathbf\{B\}\}\\mathcal\{Q\}has commonν\{\\nu\}\-support\. Then, for any𝖬,𝖬′∈ℳ\\mathsf\{M\},\\mathsf\{M\}^\{\\prime\}\\in\\mathcal\{M\}feature\-equivalent on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, there exists
τ∈𝒯𝐊ν\(𝒦\)∩𝒯𝐐¯ν\(𝒬¯;𝒜o\)\\tau\\in\\mathcal\{T\}\_\{\{\\mathbf\{K\}\}\}^\{\\nu\}\(\\mathcal\{K\}\)\\cap\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\\nu\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)such that𝖬\\mathsf\{M\}and𝖬′\\mathsf\{M\}^\{\\prime\}are Blackwell\-equivalent throughτ\\tauon𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\. If, moreover,𝒯𝐊ν\(𝒦\)∩𝒯𝐐¯ν\(𝒬¯;𝒜o\)⊆𝔊\\mathcal\{T\}\_\{\{\\mathbf\{K\}\}\}^\{\\nu\}\(\\mathcal\{K\}\)\\cap\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\\nu\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\\subseteq\\mathfrak\{G\}for a transition group𝔊⊆Autν\(𝒞\)\\mathfrak\{G\}\\subseteq\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\), then𝖬\\mathsf\{M\}is identifiable up to∼𝔊\\sim\_\{\\mathfrak\{G\}\}inℳ\\mathcal\{M\}on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, i\.e\., every𝖬′∈ℳ\\mathsf\{M\}^\{\\prime\}\\in\\mathcal\{M\}feature\-equivalent to𝖬\\mathsf\{M\}on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}satisfies𝖬∼𝔊𝖬′\\mathsf\{M\}\\sim\_\{\\mathfrak\{G\}\}\\mathsf\{M\}^\{\\prime\}\.
###### Proof sketch\.
We provide a proof sketch;[Sec\.C\.2](https://arxiv.org/html/2606.18509#A3.SS2)gives the full proof under a condition weaker than the one introduced in[Definition6](https://arxiv.org/html/2606.18509#Thmdefinition6)\. Let𝖬=\(𝐐,𝐁,𝐊\)\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\)and𝖬′=\(𝐐′,𝐁,𝐊′\)\\mathsf\{M\}^\{\\prime\}=\(\{\\mathbf\{Q\}\}^\{\\prime\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}^\{\\prime\}\)be feature\-equivalent on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\. Byν\{\\nu\}\-Blackwell reducibility, write𝐊=𝜈𝐊~𝐓g\{\\mathbf\{K\}\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g\}and𝐊′=𝜈𝐊~𝐓g′\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g^\{\\prime\}\}, whereggandg′g^\{\\prime\}are bimeasurable injective maps and𝐊~\\widetilde\{\{\\mathbf\{K\}\}\}is injective on the relevant pushed\-forward laws\. The common\-support condition then makesτ=g−1∘g′∈Autν\(𝒞\)\\tau=g^\{\-1\}\\circ g^\{\\prime\}\\in\{\\mathrm\{Aut\}\}\_\{\\nu\}\(\\mathcal\{C\}\)a well\-defined concept transition moduloν\{\\nu\}\-null sets\. We therefore obtain𝐊′=𝜈𝐊𝐓τ\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}and𝐁𝐐\(⋅∣a\)=𝐓τ𝐁𝐐′\(⋅∣a\)\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau\}\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)for alla∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\. Henceτ∈𝒯𝐊ν\(𝒦\)∩𝒯𝐐¯ν\(𝒬¯;𝒜o\)\\tau\\in\\mathcal\{T\}\_\{\{\\mathbf\{K\}\}\}^\{\{\\nu\}\}\(\\mathcal\{K\}\)\\cap\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\), and if this intersection lies in𝔊\\mathfrak\{G\}, the ambiguity is up to∼𝔊\\sim\_\{\\mathfrak\{G\}\}\. ∎
### 3\.2Characterizing concept\-side transitions
We now characterize the concept\-side valid transitions at the density level\. This refines the transition\-intersection criterion by isolating the common proof object behind model\-specific identifiability arguments\. Instead of working directly with equality of pushed concept laws, we use density ratios between attribute\-conditioned concept distributions; their logarithms are the*attribute potentials*\.
###### Definition 7\.
Suppose that𝐁\{\\mathbf\{B\}\}and𝐐\{\\mathbf\{Q\}\}together satisfy[Definition6](https://arxiv.org/html/2606.18509#Thmdefinition6)\. Fora,a′∈𝒜a,a^\{\\prime\}\\in\\mathcal\{A\}, define the*attribute potential*of𝐐¯=𝐁𝐐\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}byΔa,a′𝐐¯\(c\):=logp𝐐¯\(c∣a\)−logp𝐐¯\(c∣a′\)\\Delta\_\{a,a^\{\\prime\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)\\mathrel\{:=\}\\log p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid a\)\-\\log p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid a^\{\\prime\}\), which is well\-definedν\{\\nu\}\-a\.e\.
The following theorem gives an equivalent characterization of the concept\-side valid transition set𝒯𝐐¯ν\(𝒬¯;𝒜o\)\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\. This characterization will also be the main tool for extrapolation in[Sec\.4](https://arxiv.org/html/2606.18509#S4), where we ask when preservation of observed attribute potentials forces preservation at unseen attributes\.
###### Theorem 2\(Density characterization of concept transitions\)\.
Consider a fixed concept modulation kernel𝐁\{\\mathbf\{B\}\}, and assume that𝒬¯′=𝐁𝒬\\bar\{\\mathcal\{Q\}\}^\{\\prime\}=\{\\mathbf\{B\}\}\\mathcal\{Q\}has commonν\{\\nu\}\-support as defined in[Definition6](https://arxiv.org/html/2606.18509#Thmdefinition6)\. Let𝐐,𝐐′∈𝒬\{\\mathbf\{Q\}\},\{\\mathbf\{Q\}\}^\{\\prime\}\\in\\mathcal\{Q\}be two indexing kernels\. Let𝒜′⊆𝒜\\mathcal\{A\}^\{\\prime\}\\subseteq\\mathcal\{A\},τ∈Autν\(𝒞\)\\tau\\in\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\), and definerτ−1:=d\(𝐓τ−1ν\)/dνr\_\{\\tau^\{\-1\}\}\\mathrel\{:=\}d\(\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\{\\nu\}\)/d\{\\nu\}\. For any fixed anchora0∈𝒜′a\_\{0\}\\in\\mathcal\{A\}^\{\\prime\}and𝐐¯=𝐁𝐐\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\},𝐐¯′=𝐁𝐐′\\bar\{\\mathbf\{Q\}\}^\{\\prime\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}^\{\\prime\}, the following are equivalent:
1. \(i\)𝐐¯𝒜′=𝐓τ𝐐¯𝒜′′\\bar\{\\mathbf\{Q\}\}\_\{\\mathcal\{A\}^\{\\prime\}\}=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\_\{\\mathcal\{A\}^\{\\prime\}\}\.
2. \(ii\)For everya∈𝒜′a\\in\\mathcal\{A\}^\{\\prime\},p𝐐¯′\(c∣a\)=p𝐐¯\(τ\(c\)∣a\)rτ−1\(c\)p\_\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\\mid a\)=p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\)r\_\{\\tau^\{\-1\}\}\(c\)holds forν\{\\nu\}\-a\.e\.cc\.
3. \(iii\)The anchored density identityp𝐐¯′\(c∣a0\)=p𝐐¯\(τ\(c\)∣a0\)rτ−1\(c\)p\_\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\\mid a\_\{0\}\)=p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\_\{0\}\)r\_\{\\tau^\{\-1\}\}\(c\)holds forν\{\\nu\}\-a\.e\.cc, and for everya∈𝒜′a\\in\\mathcal\{A\}^\{\\prime\},Δa,a0𝐐¯′\(c\)=Δa,a0𝐐¯\(τ\(c\)\)\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)holdsν\{\\nu\}\-a\.e\.cc\.
Consequently,τ∈𝒯𝐐¯ν\(𝒬¯;𝒜o\)\\tau\\in\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\\nu\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)if and only if there exists𝐐′∈𝒬\{\\mathbf\{Q\}\}^\{\\prime\}\\in\\mathcal\{Q\}such that these equivalent conditions hold with𝒜′=𝒜o\\mathcal\{A\}^\{\\prime\}=\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\.
###### Proof sketch\.
The equivalence between \(i\) and \(ii\) is the Radon–Nikodym chain rule applied to𝐓τ−1\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\. Under the positivity assumption, \(ii\) implies \(iii\) by takinga=a0a=a\_\{0\}, then taking logs of the identities in \(ii\) foraa,a0a\_\{0\}and subtracting to cancelrτ−1\(c\)r\_\{\\tau^\{\-1\}\}\(c\)\. Conversely, if \(iii\) holds, multiplying the anchored density identity by the exponential of the attribute potential gives \(ii\)\. The final statement follows from the definition of𝒯𝐐¯ν\(𝒬¯;𝒜o\)\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\\nu\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\. ∎
Notably, statement\(iii\)of[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)characterizes the concept\-side valid transition set𝒯𝐐¯ν\(𝒬¯;𝒜o\)\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)in terms of anchored density identities and attribute\-potential preservation, avoiding the intractablerτ−1r\_\{\\tau^\{\-1\}\}term\. We now illustrate how[Thms\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)and[2](https://arxiv.org/html/2606.18509#Thmtheorem2)recover an identifiability result for the running example\.
###### Example 2\(Running example, affine recovery\)\.
Consider the running example \([Ex\.1](https://arxiv.org/html/2606.18509#Thmexample1)\)\. Taking𝒳~=𝒳\\widetilde\{\\mathcal\{X\}\}=\\mathcal\{X\}and𝐊~=𝐓id𝒳\\widetilde\{\{\\mathbf\{K\}\}\}=\\mathbf\{T\}\_\{\\mathrm\{id\}\_\{\\mathcal\{X\}\}\}, the class𝒦\\mathcal\{K\}isν\{\\nu\}\-Blackwell reducible withν\{\\nu\}taken to be Lebesgue measure onℝk\\mathbb\{R\}^\{k\}\.
Fix𝖬=\(𝐐,𝐁,𝐊\)∈ℳ\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\)\\in\\mathcal\{M\}, with𝐐=𝐓f\{\\mathbf\{Q\}\}=\\mathbf\{T\}\_\{f\}and𝐊=𝐓g\{\\mathbf\{K\}\}=\\mathbf\{T\}\_\{g\}\. Fixinga0∈𝒜oa\_\{0\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, the attribute potential is
Δa,a0𝐐¯\(c\)=⟨f\(a\)−f\(a0\),c⟩−\{logZ\(f\(a\)\)−logZ\(f\(a0\)\)\},\\Delta^\{\\bar\{\\mathbf\{Q\}\}\}\_\{a,a\_\{0\}\}\(c\)=\\langle f\(a\)\-f\(a\_\{0\}\),c\\rangle\-\\\{\\log Z\(f\(a\)\)\-\\log Z\(f\(a\_\{0\}\)\)\\\},\(2\)and[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)implies that for anyτ∈𝒯𝐐¯ν\(𝒬¯;𝒜o\)\\tau\\in\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)there must exist𝐐′∈𝒬\{\\mathbf\{Q\}\}^\{\\prime\}\\in\\mathcal\{Q\}such thatΔa,a0𝐐¯′\(c\)=Δa,a0𝐐¯\(τ\(c\)\)\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)\. In particular, ifspan\{f\(a\)−f\(a0\)\|a∈𝒜o\}=ℝk\\operatorname\{span\}\\\{f\(a\)\-f\(a\_\{0\}\)\\;\|\\;a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\\}=\\mathbb\{R\}^\{k\}, then preservation forcesτ\\tauto be affine, as we show in[Sec\.D\.1](https://arxiv.org/html/2606.18509#A4.SS1)\. Hence𝒯𝐐¯ν\(𝒬¯;𝒜o\)⊆Affine\(k\)∩Autν\(ℝk\)\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\\subseteq\\mathrm\{Affine\}\(k\)\\cap\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathbb\{R\}^\{k\}\), whereAffine\(k\)\\mathrm\{Affine\}\(k\)is the set of affine transformations onℝk\\mathbb\{R\}^\{k\}\. Therefore,
𝒯𝐊ν\(𝒦\)∩𝒯𝐐¯ν\(𝒬¯;𝒜o\)⊆𝒯𝐐¯ν\(𝒬¯;𝒜o\)⊆𝔊≔Affine\(k\)∩Autν\(ℝk\)\.\\mathcal\{T\}\_\{\{\\mathbf\{K\}\}\}^\{\{\\nu\}\}\(\\mathcal\{K\}\)\\cap\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\\subseteq\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\\subseteq\\mathfrak\{G\}\\coloneqq\\mathrm\{Affine\}\(k\)\\cap\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathbb\{R\}^\{k\}\)\.Since the exponential\-family concept laws have positive densities, the support condition in[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)holds\. Therefore,[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)implies that the running example is identifiable up to∼𝔊\\sim\_\{\\mathfrak\{G\}\}, i\.e\., invertible affine transformations of the concept space\.
Many existing identifiability results impose richness assumptions under which the observed contrasts have enough rank to identify the entire latent representation up to the target ambiguity class\. The attribute\-potential characterization also describes what remains identifiable when this rank condition fails: the observed contrast equations constrain only the component of the latent transition visible through the observed contrast span\.
## 4Extrapolation via attribute potentials
Extrapolation asks when predictive agreement on observed attributes extends beyond them: if two CMMs agree on conditional feature laws over𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, when must they also agree on𝒜ex⊇𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\\supseteq\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}? Unlike identifiability, which characterizes latent transitions compatible with feature equivalence on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, extrapolation asks whether the induced transition also enforces agreement outside𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\. For a transitionτ\\taufrom[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1), with𝐊′=𝜈𝐊𝐓τ\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}and𝐐¯𝒜o=𝐓τ𝐐¯𝒜o′\\bar\{\\mathbf\{Q\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\}=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\}, the mixing side remains fixed as the attribute set expands\. Therefore, extrapolation reduces to whether this concept\-side relation extends from𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}to𝒜ex\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\. By[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2), this is equivalent to preserving attribute potentials, as formalized next\.
###### Theorem 3\(Attribute\-potential characterization of extrapolation\)\.
Letℳ=𝒬×\{𝐁\}×𝒦\\mathcal\{M\}=\\mathcal\{Q\}\\times\\\{\{\\mathbf\{B\}\}\\\}\\times\\mathcal\{K\}be a CMM class, and let𝒜o⊆𝒜ex⊆𝒜\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\subseteq\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\\subseteq\\mathcal\{A\}\. Assume that𝒦\\mathcal\{K\}isν\{\\nu\}\-Blackwell reducible and that𝒬¯=𝐁𝒬\\bar\{\\mathcal\{Q\}\}=\{\\mathbf\{B\}\}\\mathcal\{Q\}has commonν\{\\nu\}\-support\. Fixa0∈𝒜oa\_\{0\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, and let𝖬=\(𝐐,𝐁,𝐊\)\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\)and𝖬′=\(𝐐′,𝐁,𝐊′\)\\mathsf\{M\}^\{\\prime\}=\(\{\\mathbf\{Q\}\}^\{\\prime\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}^\{\\prime\}\)be feature\-equivalent on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, and letτ\\taube a transition induced by[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1), so that𝐊′=𝜈𝐊𝐓τ\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}and𝐐¯𝒜o=𝐓τ𝐐¯𝒜o′\\bar\{\\mathbf\{Q\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\}=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\}where𝐐¯=𝐁𝐐\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\},𝐐¯′=𝐁𝐐′\\bar\{\\mathbf\{Q\}\}^\{\\prime\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}^\{\\prime\}\. Then, for this transitionτ\\tau,
𝐏𝒜ex𝖬=𝐏𝒜ex𝖬′⟺Δa,a0𝐐¯′\(c\)=Δa,a0𝐐¯\(τ\(c\)\)for everya∈𝒜ex,ν\-a\.e\.c\.\{\\mathbf\{P\}\}^\{\\mathsf\{M\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}=\{\\mathbf\{P\}\}^\{\\mathsf\{M\}^\{\\prime\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}\\quad\\Longleftrightarrow\\quad\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)\\text\{ for every \}a\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\},\\ \{\\nu\}\\text\{\-a\.e\. \}c\.Thus, if the identitiesΔa,a0𝐐¯′\(c\)=Δa,a0𝐐¯\(τ\(c\)\)\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)that hold on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}extend over𝒜ex\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}, then𝐏𝒜ex𝖬=𝐏𝒜ex𝖬′\{\\mathbf\{P\}\}^\{\\mathsf\{M\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}=\{\\mathbf\{P\}\}^\{\\mathsf\{M\}^\{\\prime\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}\.
We defer the proof of[Thm\.3](https://arxiv.org/html/2606.18509#Thmtheorem3)to[Sec\.C\.4](https://arxiv.org/html/2606.18509#A3.SS4)\. The remaining work is to provide class\-specific arguments showing that the transported attribute\-potential identities on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}force the corresponding identities on𝒜ex\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\. This implication is not automatic: in the running exponential\-family CMM with unrestricted indexing mapf:𝒜→ℝkf\\colon\\mathcal\{A\}\\to\\mathbb\{R\}^\{k\}, one may keepf\(a\)f\(a\)unchanged for alla∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}while alteringf\(aex\)f\(\{a\_\{\\textnormal\{ex\}\}\}\)arbitrarily at an unseen attributeaex∉𝒜o\{a\_\{\\textnormal\{ex\}\}\}\\notin\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\. The observed conditional feature laws are unchanged, butp\(X∣A=aex\)p\(X\\mid A=\{a\_\{\\textnormal\{ex\}\}\}\)can change\. Thus extrapolation requires structural assumptions tying unseen attribute potentials to observed ones\. The following theorem gives one such condition, requiring affine dependence through a fixed attribute representationφ\\varphi\.
###### Theorem 4\(Attribute\-potential preservation under affine attribute\-potential differences\)\.
Letℳ=𝒬×\{𝐁\}×𝒦\\mathcal\{M\}=\\mathcal\{Q\}\\times\\\{\{\\mathbf\{B\}\}\\\}\\times\\mathcal\{K\}be a CMM class satisfying the assumptions of[Thms\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)and[2](https://arxiv.org/html/2606.18509#Thmtheorem2)\. Let𝒜o⊆𝒜ex⊆𝒜\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\subseteq\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\\subseteq\\mathcal\{A\}, and fixa0∈𝒜oa\_\{0\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\. Suppose there exists a fixed attribute representationφ:𝒜→ℝm\\varphi\\colon\\mathcal\{A\}\\to\\mathbb\{R\}^\{m\}such that, for every𝐐¯∈𝒬¯=𝐁𝒬\\bar\{\\mathbf\{Q\}\}\\in\\bar\{\\mathcal\{Q\}\}=\{\\mathbf\{B\}\}\\mathcal\{Q\}, there exists a measurable mapD𝐐¯:𝒞×𝒞→ℝmD^\{\\bar\{\\mathbf\{Q\}\}\}\\colon\\mathcal\{C\}\\times\\mathcal\{C\}\\to\\mathbb\{R\}^\{m\}satisfying
Δa,a0𝐐¯\(c\)−Δa,a0𝐐¯\(c0\)=⟨φ\(a\)−φ\(a0\),D𝐐¯\(c,c0\)⟩\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\)\-\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\_\{0\}\)=\\langle\\varphi\(a\)\-\\varphi\(a\_\{0\}\),D^\{\\bar\{\\mathbf\{Q\}\}\}\(c,c\_\{0\}\)\\ranglefor everya∈𝒜exa\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}andν⊗ν\{\\nu\}\\otimes\{\\nu\}\-a\.e\.\(c,c0\)\(c,c\_\{0\}\)\. Let𝖬=\(𝐐,𝐁,𝐊\)\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\)and𝖬′=\(𝐐′,𝐁,𝐊′\)\\mathsf\{M\}^\{\\prime\}=\(\{\\mathbf\{Q\}\}^\{\\prime\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}^\{\\prime\}\)be feature\-equivalent on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, and letτ\\taube the transition induced by[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)\. Ifφ\(𝒜ex\)\\varphi\(\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\)is contained in the affine hull ofφ\(𝒜o\)\\varphi\(\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\), then the transported attribute\-potential identities extend from𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}to𝒜ex\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}: for everya∈𝒜exa\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\},Δa,a0𝐐¯′\(c\)=Δa,a0𝐐¯\(τ\(c\)\)\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)forν\{\\nu\}\-a\.e\.cc\.
###### Proof sketch\.
We defer the full proof to[Sec\.C\.5](https://arxiv.org/html/2606.18509#A3.SS5)\. For anyaex∈𝒜ex\{a\_\{\\textnormal\{ex\}\}\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}, we can writeφ\(aex\)=∑iαiφ\(ai\)\\varphi\(\{a\_\{\\textnormal\{ex\}\}\}\)=\\sum\_\{i\}\\alpha\_\{i\}\\varphi\(a\_\{i\}\), whereai∈𝒜oa\_\{i\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}and∑iαi=1\\sum\_\{i\}\\alpha\_\{i\}=1\. For any𝖬∈ℳ\\mathsf\{M\}\\in\\mathcal\{M\}, we have the centered potentialΔa,a0𝐐¯\(c\)−Δa,a0𝐐¯\(c0\)=⟨φ\(a\)−φ\(a0\),D𝐐¯\(c,c0\)⟩\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\)\-\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\_\{0\}\)=\\langle\\varphi\(a\)\-\\varphi\(a\_\{0\}\),D^\{\\bar\{\\mathbf\{Q\}\}\}\(c,c\_\{0\}\)\\rangle, so the same quantity ataex\{a\_\{\\textnormal\{ex\}\}\}is an affine combination of the centered potentials at theaia\_\{i\}’s, and the same relation holds for𝐐¯′\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\. Statement\(iii\)of[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)transports the observed identities, hence the affine expansion givesΔaex,a0𝐐¯′\(c\)−Δaex,a0𝐐¯′\(c0\)=Δaex,a0𝐐¯\(τ\(c\)\)−Δaex,a0𝐐¯\(τ\(c0\)\)\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)\-\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\_\{0\}\)=\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)\-\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\_\{0\}\)\)\. The density identity ata0a\_\{0\}and normalization ataex\{a\_\{\\textnormal\{ex\}\}\}together forceΔaex,a0𝐐¯′\(c\)=Δaex,a0𝐐¯\(τ\(c\)\)\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)\. ∎
The point of[Thm\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)is that extrapolation is driven by the linear structure of attribute potentials, not by linearity of the raw attribute space\. The representationφ\\varphimay encode nonlinear or combinatorial structure, such as interaction terms among atomic attributes, as in factorial designs\(Fisher,[1935](https://arxiv.org/html/2606.18509#bib.bib20)\)and combinatorial intervention models\(O’Donnell,[2008](https://arxiv.org/html/2606.18509#bib.bib21); Agarwalet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib22)\)\. Thus, the theorem lifts feature\-equivalence on finitely many interactions to extrapolation guarantees over a much larger combinatorial attribute space, as stated in the following corollary\.
###### Corollary 1\(Interaction\-basis extrapolation\)\.
Letℳ=𝒬×\{𝐁\}×𝒦\\mathcal\{M\}=\\mathcal\{Q\}\\times\\\{\{\\mathbf\{B\}\}\\\}\\times\\mathcal\{K\}satisfy the assumptions of[Thm\.3](https://arxiv.org/html/2606.18509#Thmtheorem3)\. Let𝒜=2\[m\]\\mathcal\{A\}=2^\{\[m\]\}, letℋ⊆2\[m\]\\mathcal\{H\}\\subseteq 2^\{\[m\]\}be a finite downward\-closed family containing∅\\varnothing, i\.e\.,H∈ℋH\\in\\mathcal\{H\}andU⊆HU\\subseteq HimplyU∈ℋU\\in\\mathcal\{H\}, and set𝒜o=ℋ\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}=\\mathcal\{H\}\. Assume every𝖬∈ℳ\\mathsf\{M\}\\in\\mathcal\{M\}induces concept densities of the formp𝖬\(c∣S\)=exp\(∑T∈ℋ,T⊆ShT𝖬\(c\)\)/Z𝖬\(S\)p^\{\\mathsf\{M\}\}\(c\\mid S\)=\\exp\\left\(\\sum\_\{T\\in\\mathcal\{H\},\\,T\\subseteq S\}h\_\{T\}^\{\\mathsf\{M\}\}\(c\)\\right\)/Z^\{\\mathsf\{M\}\}\(S\)\. Then feature equivalence on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}implies feature equivalence on allS⊆\[m\]S\\subseteq\[m\]\.
###### Proof sketch\.
We defer a full proof to[Sec\.C\.6](https://arxiv.org/html/2606.18509#A3.SS6)\. Denoting𝟏\\mathbf\{1\}by the indicator function, we haveΔS,∅𝖬\(c\)−ΔS,∅𝖬\(c0\)=⟨φℋ\(S\)−φℋ\(∅\),D𝖬\(c,c0\)⟩\\Delta\_\{S,\\varnothing\}^\{\\mathsf\{M\}\}\(c\)\-\\Delta\_\{S,\\varnothing\}^\{\\mathsf\{M\}\}\(c\_\{0\}\)=\\langle\\varphi\_\{\\mathcal\{H\}\}\(S\)\-\\varphi\_\{\\mathcal\{H\}\}\(\\varnothing\),D^\{\\mathsf\{M\}\}\(c,c\_\{0\}\)\\rangle, whereφℋ\(S\)=\(𝟏\{T⊆S\}\)T∈ℋ\\varphi\_\{\\mathcal\{H\}\}\(S\)=\(\\mathbf\{1\}\\\{T\\subseteq S\\\}\)\_\{T\\in\\mathcal\{H\}\}andDT𝖬\(c,c0\)=hT𝖬\(c\)−hT𝖬\(c0\)D\_\{T\}^\{\\mathsf\{M\}\}\(c,c\_\{0\}\)=h\_\{T\}^\{\\mathsf\{M\}\}\(c\)\-h\_\{T\}^\{\\mathsf\{M\}\}\(c\_\{0\}\)forT≠∅T\\neq\\varnothingandD∅𝖬\(c,c0\)=0D\_\{\\varnothing\}^\{\\mathsf\{M\}\}\(c,c\_\{0\}\)=0\. One can show thatφℋ\(S\)\\varphi\_\{\\mathcal\{H\}\}\(S\)is an affine combination of\{φℋ\(U\)\|U∈ℋ\}\\\{\\varphi\_\{\\mathcal\{H\}\}\(U\)\\;\|\\;U\\in\\mathcal\{H\}\\\}for anyS⊆\[m\]S\\subseteq\[m\]\.[Thms\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)and[3](https://arxiv.org/html/2606.18509#Thmtheorem3)yield desired result\. ∎
We close the section by recording an implication of[Thm\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)for the running example\.
###### Example 3\(Running example: extrapolation\)\.
For the running exponential\-family CMM, assume thatf\(a\)=Wfφ\(a\)f\(a\)=W\_\{f\}\\varphi\(a\)for a matrixWfW\_\{f\}depending onffand a fixed attribute representationφ\\varphi\. This satisfies the attribute\-potential assumption of[Thm\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)withD𝐐¯\(c,c0\)=Wf⊤\(c−c0\)D^\{\\bar\{\\mathbf\{Q\}\}\}\(c,c\_\{0\}\)=W\_\{f\}^\{\\top\}\(c\-c\_\{0\}\)\. Consequently, by[Thm\.3](https://arxiv.org/html/2606.18509#Thmtheorem3), feature equivalence on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}implies feature equivalence at every targetaex∈𝒜ex\{a\_\{\\textnormal\{ex\}\}\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}whoseφ\(aex\)\\varphi\(\{a\_\{\\textnormal\{ex\}\}\}\)is an affine combination of\{φ\(a\)\|a∈𝒜o\}\\\{\\varphi\(a\)\\;\|\\;a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\\}\.
## 5Applications: Recoveries and examples
In this section, we apply the transition\-and\-potential characterizations from[Thms\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)and[2](https://arxiv.org/html/2606.18509#Thmtheorem2)and the extrapolation criteria from[Thms\.3](https://arxiv.org/html/2606.18509#Thmtheorem3)and[4](https://arxiv.org/html/2606.18509#Thmtheorem4)to representative model classes\. Identifiability follows by combining the induced attribute potentials with class\-specific rigidity arguments\. Extrapolation follows by checking whether the transported potential identities observed on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}force the corresponding identities on target attributes in𝒜ex\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\.
#### Nonlinear ICA and iVAE\.
Nonlinear ICA with auxiliary variables\(Hyvarinenet al\.,[2019](https://arxiv.org/html/2606.18509#bib.bib4)\)and iVAE\(Khemakhemet al\.,[2020a](https://arxiv.org/html/2606.18509#bib.bib2)\)fit the CMM template by taking the attributea∈𝒜a\\in\\mathcal\{A\}to be the auxiliary or conditioning variable, the conceptc=\(c1,…,ck\)∈𝒞c=\(c\_\{1\},\\ldots,c\_\{k\}\)\\in\\mathcal\{C\}to be the latent source vector, and the featurex∈𝒳x\\in\\mathcal\{X\}to be the nonlinear observation\. In the conditionally independent source setting, the attribute potential decomposes asΔa,a0𝐐¯\(c\)=∑i=1k\[logqi\(ci∣a\)−logqi\(ci∣a0\)\]\\Delta^\{\\bar\{\\mathbf\{Q\}\}\}\_\{a,a\_\{0\}\}\(c\)=\\sum\_\{i=1\}^\{k\}\[\\log q\_\{i\}\(c\_\{i\}\\mid a\)\-\\log q\_\{i\}\(c\_\{i\}\\mid a\_\{0\}\)\], so the CMM contrast is exactly the coordinate\-wise auxiliary\-variable contrast used in nonlinear ICA and iVAE\. Under the variability assumptions ofHyvarinenet al\.\([2019](https://arxiv.org/html/2606.18509#bib.bib4)\), preservation of these potentials rules out mixing across source coordinates, forcingτ\(c\)=\(ϕ1\(cπ\(1\)\),…,ϕk\(cπ\(k\)\)\)\\tau\(c\)=\(\\phi\_\{1\}\(c\_\{\\pi\(1\)\}\),\\ldots,\\phi\_\{k\}\(c\_\{\\pi\(k\)\}\)\)up to a permutationπ\\piand invertible scalar mapsϕi\\phi\_\{i\}\. More details on recoveries are given in[Sec\.E\.1](https://arxiv.org/html/2606.18509#A5.SS1)\.
#### Causal representation learning\.
Causal representation learning fits the CMM template by takinga∈𝒜a\\in\\mathcal\{A\}to be an environment or intervention label,c=\(c1,…,ck\)∈∏i=1k𝒞ic=\(c\_\{1\},\\ldots,c\_\{k\}\)\\in\\prod\_\{i=1\}^\{k\}\\mathcal\{C\}\_\{i\}to be the latent causal representation, andx∈𝒳x\\in\\mathcal\{X\}to be the observed representation\. We work in the common\-support regime, so the induced joint concept laws have positive densities with respect to a common reference measure and the attribute potentials in[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)are well defined\. The modulator is the tuple of local mechanismsλ=\(pi\)i=1k\{\\lambda\}=\(p\_\{i\}\)\_\{i=1\}^\{k\}, indexed by a directed acyclic graph \(DAG\)GλG\_\{\{\\lambda\}\}, withℒ=⨆G∈DAG\(\[k\]\)∏i=1k𝒦mk\(𝒞paG\(i\)→𝒞i\)\{\\mathscr\{L\}\}=\\bigsqcup\_\{G\\in\\mathrm\{DAG\}\(\[k\]\)\}\\prod\_\{i=1\}^\{k\}\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{C\}\_\{\\textnormal\{pa\}\_\{G\}\(i\)\}\\to\\mathcal\{C\}\_\{i\}\)and𝐁\(dc∣λ\)=∏i=1kpi\(ci∣cpaGλ\(i\)\)dc\{\\mathbf\{B\}\}\(dc\\mid\{\\lambda\}\)=\\prod\_\{i=1\}^\{k\}p\_\{i\}\(c\_\{i\}\\mid c\_\{\\textnormal\{pa\}\_\{G\_\{\{\\lambda\}\}\}\(i\)\}\)\\,dc\. Here,DAG\(\[k\]\)\\mathrm\{DAG\}\(\[k\]\)is the set of DAGs on\[k\]\[k\],paG\(i\)\\textnormal\{pa\}\_\{G\}\(i\)is the parent set of nodeiiinGG,𝒞S≔∏i∈S𝒞i\\mathcal\{C\}\_\{S\}\\coloneqq\\prod\_\{i\\in S\}\\mathcal\{C\}\_\{i\}, and the indexing kernel𝐐=𝐓f\{\\mathbf\{Q\}\}=\\mathbf\{T\}\_\{f\}maps each environment label to its active tuple of local mechanisms\.
For an intervention labelaa, lettar\(a\)⊆\[k\]\\operatorname\{tar\}\(a\)\\subseteq\[k\]denote its target nodes, and writepiap\_\{i\}^\{a\}for the local mechanism active at nodeiiunderaa: it satisfiespia=pi0p\_\{i\}^\{a\}=p\_\{i\}^\{0\}fori∉tar\(a\)i\\notin\\operatorname\{tar\}\(a\), whilepiap\_\{i\}^\{a\}may differ frompi0p\_\{i\}^\{0\}fori∈tar\(a\)i\\in\\operatorname\{tar\}\(a\)\. Fixing an observational labela0∈𝒜oa\_\{0\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}with mechanisms\(pi0\)i=1k\(p\_\{i\}^\{0\}\)\_\{i=1\}^\{k\}, the attribute potential becomesΔa,a0𝐐¯\(c\)=∑i∈tar\(a\)\[logpia\(ci∣cpaG\(i\)\)−logpi0\(ci∣cpaG\(i\)\)\]\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\)=\\sum\_\{i\\in\\operatorname\{tar\}\(a\)\}\[\\log p\_\{i\}^\{a\}\(c\_\{i\}\\mid c\_\{\\textnormal\{pa\}\_\{G\}\(i\)\}\)\-\\log p\_\{i\}^\{0\}\(c\_\{i\}\\mid c\_\{\\textnormal\{pa\}\_\{G\}\(i\)\}\)\]\. Thus, the attribute potential records exactly the local mechanism contrasts induced by the intervention\. Across CRL variants, identifiability follows from showing that transitions preserving these contrasts must belong to the target ambiguity class: Gaussian CRL\(Buchholzet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib18)\), interventional\-to\-observational density ratios in nonparametric CRL\(von Kügelgenet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib8)\), and gradients of environment ratios in score\-based CRL\(Varıcıet al\.,[2024b](https://arxiv.org/html/2606.18509#bib.bib23),[2025](https://arxiv.org/html/2606.18509#bib.bib9)\)\.
For extrapolation, the additional structure is supplied by the independent causal mechanisms \(ICM\) principle\(Peterset al\.,[2017](https://arxiv.org/html/2606.18509#bib.bib24); Schölkopfet al\.,[2021](https://arxiv.org/html/2606.18509#bib.bib1)\): interventions may replace selected local mechanisms while leaving the remaining mechanisms unchanged\. LetJ⊆\[k\]J\\subseteq\[k\]be a set of intervention labels such that\{aj\|j∈J\}⊆𝒜o\\\{a\_\{j\}\\;\|\\;j\\in J\\\}\\subseteq\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}have pairwise disjoint targets, i\.e\.,tar\(aj\)∩tar\(aj′\)=∅\\operatorname\{tar\}\(a\_\{j\}\)\\cap\\operatorname\{tar\}\(a\_\{j^\{\\prime\}\}\)=\\varnothingfor allj≠j′j\\neq j^\{\\prime\}\. The associated composite interventionaJa\_\{J\}is defined bypiaJ=piajp\_\{i\}^\{a\_\{J\}\}=p\_\{i\}^\{a\_\{j\}\}ifi∈tar\(aj\)i\\in\\operatorname\{tar\}\(a\_\{j\}\)for the uniquej∈Jj\\in J, andpiaJ=pi0p\_\{i\}^\{a\_\{J\}\}=p\_\{i\}^\{0\}ifi∉⋃j∈Jtar\(aj\)i\\notin\\bigcup\_\{j\\in J\}\\operatorname\{tar\}\(a\_\{j\}\)\. Under the ICM principle, non\-target mechanisms cancel against the anchor and the target sets do not overlap, so the composite potential is additive:ΔaJ,a0𝐐¯\(c\)=∑j∈JΔaj,a0𝐐¯\(c\)\\Delta\_\{a\_\{J\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\)=\\sum\_\{j\\in J\}\\Delta\_\{a\_\{j\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\)\. Therefore, if the transported identitiesΔaj,a0𝐐¯′\(c\)=Δaj,a0𝐐¯\(τ\(c\)\)\\Delta\_\{a\_\{j\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a\_\{j\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)hold for allj∈Jj\\in J, additivity gives the same identity foraJa\_\{J\}\. By[Thm\.3](https://arxiv.org/html/2606.18509#Thmtheorem3), feature equivalence on the observed environments then extrapolates to the composite interventionaJa\_\{J\}\. We note that this is analogous to the intervention\-generalization strategy ofBravo\-Hermsdorffet al\.\([2023](https://arxiv.org/html/2606.18509#bib.bib25)\), which studies extrapolation under factorization assumptions on the interventional factor models\.
#### Perturbation modeling\.
Perturbation modeling invon Kügelgenet al\.\([2025](https://arxiv.org/html/2606.18509#bib.bib13)\)fits the CMM template by taking the attributea∈𝒜=ℝKa\\in\\mathcal\{A\}=\\mathbb\{R\}^\{K\}to be a perturbation label, the conceptc∈𝒞=ℝkc\\in\\mathcal\{C\}=\\mathbb\{R\}^\{k\}to be the perturbation\-relevant latent state, and the featurex∈𝒳x\\in\\mathcal\{X\}to be the observed measurement\. The deterministic indexing kernel is𝐐=𝐓f\{\\mathbf\{Q\}\}=\\mathbf\{T\}\_\{f\}, withf\(a\)=W\(a−a0\)f\(a\)=W\(a\-a\_\{0\}\), and the shared concept\-modulation kernel is𝐁\(dc∣λ\)=𝒩\(λ,I\)\{\\mathbf\{B\}\}\(dc\\mid\{\\lambda\}\)=\\mathcal\{N\}\(\{\\lambda\},I\), yielding the centered attribute potential to beΔa,a0𝐐¯\(c\)−Δa,a0𝐐¯\(c0\)=⟨a−a0,W⊤\(c−c0\)⟩\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\)\-\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\_\{0\}\)=\\langle a\-a\_\{0\},W^\{\\top\}\(c\-c\_\{0\}\)\\rangle\. Using this quantity,[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)followed by model\-specific arguments provided byvon Kügelgenet al\.\([2025](https://arxiv.org/html/2606.18509#bib.bib13)\), we recover the identifiability guarantee ofvon Kügelgenet al\.\([2025](https://arxiv.org/html/2606.18509#bib.bib13)\)under the sufficient\-diversity conditionspan\{W\(a−a0\)\|a∈𝒜o\}=ℝk\\operatorname\{span\}\\\{W\(a\-a\_\{0\}\)\\;\|\\;a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\\}=\\mathbb\{R\}^\{k\}: the perturbation effect matrix is identifiable up to an orthogonal transformation\. For extrapolation under the same condition, the centered potential satisfies the conditions of[Thm\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)withφ\(a\)=a\\varphi\(a\)=aandD𝐐¯\(c,c0\)=W⊤\(c−c0\)D^\{\\bar\{\\mathbf\{Q\}\}\}\(c,c\_\{0\}\)=W^\{\\top\}\(c\-c\_\{0\}\)\.[Thm\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)justifies extrapolation to anyaex−a0∈span\{a−a0\|a∈𝒜o\}\{a\_\{\\textnormal\{ex\}\}\}\-a\_\{0\}\\in\\operatorname\{span\}\\\{a\-a\_\{0\}\\;\|\\;a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\\}\.
## 6Discussion
CMMs provide a common language for studying identifiability and extrapolation through attribute potentials\. Their main role is to separate the generic transition\-and\-contrast step from model\-specific rigidity arguments: feature agreement first induces a latent transition, and the remaining question is whether the model class forces this transition to preserve the relevant attribute potentials\.
Our current formulation uses attribute potentials as log\-density ratios under a common\-support assumption\. This covers many exponential\-family, soft\-intervention, and stochastic hard\-intervention examples, but does not cover deterministic do\-interventions or other support\-changing mechanisms\. Extending the theory to singular or partially overlapping concept laws would broaden its applicability to causal representation learning without changing the central transition\-and\-contrast viewpoint\.
A further direction is to use CMMs as identifiability guidance for weakly supervised and constraint\-aware representation learning\. Recent neurosymbolic and concept\-based systems seek to learn intermediate symbolic or concept\-level representations from indirect supervision, such as labels, answers, captions, or consistency constraints\(Duanet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib28); Danieleet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib29); Barbieroet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib30); Oikarinenet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib31)\)\. These works show that symbolic or concept\-level structure can guide representation learning, but their results typically concern optimization, interpretability, or empirical generalization, rather than rigorous identifiability and extrapolation guarantees\. Our framework suggests one route toward such guarantees by specifying how tasks or attributes modulate latent concept laws and characterizing the induced attribute potentials\.
## Acknowledgements
This research was developed with funding from the Defense Advanced Research Projects Agency \(DARPA\) via HR0011\-25\-3\-0239, FA8750\-23\-2\-1015, ONR via N00014\-23\-1\-2368, and NSF via IIS\-1909816\.
## References
- E\. Acartürk, B\. Varıcı, K\. Shanmugam, and A\. Tajer \(2024\)Sample complexity of interventional causal representation learning\.InAdvances in Neural Information Processing Systems,pp\. 39350–39385\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1)\.
- A\. Agarwal, A\. Agarwal, and S\. Vijaykumar \(2023\)Synthetic combinations: a causal inference framework for combinatorial interventions\.Advances in Neural Information Processing Systems,pp\. 19195–19216\.Cited by:[§A\.4](https://arxiv.org/html/2606.18509#A1.SS4.p1.1),[§4](https://arxiv.org/html/2606.18509#S4.p3.1)\.
- K\. Ahuja, J\. S\. Hartford, and Y\. Bengio \(2022a\)Weakly supervised representation learning with sparse perturbations\.Advances in Neural Information Processing Systems,pp\. 15516–15528\.Cited by:[Table 2](https://arxiv.org/html/2606.18509#A6.T2.36.36.36.6),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.27.27.27.5)\.
- K\. Ahuja, J\. S\. Hartford, and Y\. Bengio \(2022b\)Properties from mechanisms: an equivariance perspective on identifiable representation learning\.InInternational Conference on Learning Representations,Cited by:[§A\.1](https://arxiv.org/html/2606.18509#A1.SS1.p1.1),[§A\.5](https://arxiv.org/html/2606.18509#A1.SS5.p1.1)\.
- K\. Ahuja, D\. Mahajan, Y\. Wang, and Y\. Bengio \(2023\)Interventional causal representation learning\.InInternational Conference on Machine Learning,pp\. 372–407\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1),[§1](https://arxiv.org/html/2606.18509#S1.p3.1)\.
- S\. Arora, R\. Ge, and A\. Moitra \(2012\)Learning topic models – going beyond SVD\.InIEEE Annual Symposium on Foundations of Computer Science,pp\. 1–10\.Cited by:[Table 2](https://arxiv.org/html/2606.18509#A6.T2.44.44.44.6),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.33.33.33.5)\.
- P\. Barbiero, G\. Ciravegna, F\. Giannini, M\. Espinosa Zarlenga, L\. C\. Magister, A\. Tonda, P\. Lio, F\. Precioso, M\. Jamnik, and G\. Marra \(2023\)Interpretable neural\-symbolic concept reasoning\.InInternational Conference on Machine Learning,pp\. 1801–1825\.Cited by:[§6](https://arxiv.org/html/2606.18509#S6.p3.1)\.
- G\. Bravo\-Hermsdorff, D\. Watson, J\. Yu, J\. Zeitler, and R\. Silva \(2023\)Intervention generalization: a view from factor graph models\.InAdvances in Neural Information Processing Systems,pp\. 43662–43675\.Cited by:[§5](https://arxiv.org/html/2606.18509#S5.SS0.SSS0.Px2.p3.15)\.
- J\. Brehmer, P\. De Haan, P\. Lippe, and T\. Cohen \(2022\)Weakly supervised causal representation learning\.InAdvances in Neural Information Processing Systems,pp\. 38319–38331\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1)\.
- T\. Bricken, A\. Templeton, J\. Batson, B\. Chen, A\. Jermyn, T\. Conerly, N\. Turner, C\. Anil, C\. Denison, A\. Askell, R\. Lasenby, Y\. Wu, S\. Kravec, N\. Schiefer, T\. Maxwell, N\. Joseph, Z\. Hatfield\-Dodds, A\. Tamkin, K\. Nguyen, B\. McLean, J\. E\. Burke, T\. Hume, S\. Carter, T\. Henighan, and C\. Olah \(2023\)Towards monosemanticity: decomposing language models with dictionary learning\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2023/monosemantic-features/index.html)Cited by:[§A\.3](https://arxiv.org/html/2606.18509#A1.SS3.p1.1)\.
- S\. Buchholz, G\. Rajendran, E\. Rosenfeld, B\. Aragam, B\. Schölkopf, and P\. K\. Ravikumar \(2023\)Learning linear causal representations from interventions under general nonlinear mixing\.InAdvances in Neural Information Processing Systems,pp\. 45419–45462\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1),[§E\.4](https://arxiv.org/html/2606.18509#A5.SS4.SSS0.Px1.p1.11),[§E\.4](https://arxiv.org/html/2606.18509#A5.SS4.SSS0.Px1.p1.8),[§E\.4](https://arxiv.org/html/2606.18509#A5.SS4.SSS0.Px2.p1.7),[Table 2](https://arxiv.org/html/2606.18509#A6.T2.24.24.24.5),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.18.18.18.4),[§2\.2](https://arxiv.org/html/2606.18509#S2.SS2.p5.1),[§5](https://arxiv.org/html/2606.18509#S5.SS0.SSS0.Px2.p2.13),[Proposition 5](https://arxiv.org/html/2606.18509#Thmproposition5),[Proposition 5](https://arxiv.org/html/2606.18509#Thmproposition5.p1.9.2)\.
- J\. Cui, Q\. Zhang, Y\. Wang, and Y\. Wang \(2026\)On the limits of sparse autoencoders: a theoretical framework and reweighted remedy\.InInternational Conference on Learning Representations,Cited by:[§A\.3](https://arxiv.org/html/2606.18509#A1.SS3.p1.1)\.
- A\. D’Amour, K\. Heller, D\. Moldovan, B\. Adlam, B\. Alipanahi, A\. Beutel, C\. Chen, J\. Deaton, J\. Eisenstein, M\. D\. Hoffman, F\. Hormozdiari, N\. Houlsby, S\. Hou, G\. Jerfel, A\. Karthikesalingam, M\. Lucic, Y\. Ma, C\. McLean, D\. Mincu, A\. Mitani, A\. Montanari, Z\. Nado, V\. Natarajan, C\. Nielson, T\. F\. Osborne, R\. Raman, K\. Ramasamy, R\. Sayres, J\. Schrouff, M\. Seneviratne, S\. Sequeira, H\. Suresh, V\. Veitch, M\. Vladymyrov, X\. Wang, K\. Webster, S\. Yadlowsky, T\. Yun, X\. Zhai, and D\. Sculley \(2022\)Underspecification presents challenges for credibility in modern machine learning\.Journal of Machine Learning Research23\(226\),pp\. 1–61\.Cited by:[§1](https://arxiv.org/html/2606.18509#S1.p1.1)\.
- A\. Daniele, T\. Campari, S\. Malhotra, and L\. Serafini \(2023\)Deep symbolic learning: discovering symbols and rules from perceptions\.InInternational Joint Conference on Artificial Intelligence,Cited by:[§6](https://arxiv.org/html/2606.18509#S6.p3.1)\.
- Y\. Du and L\. Kaelbling \(2024\)Position: compositional generative modeling: a single model is not all you need\.InInternational Conference on Machine Learning,Cited by:[§A\.4](https://arxiv.org/html/2606.18509#A1.SS4.p1.1)\.
- X\. Duan, X\. Wang, P\. Zhao, G\. Shen, and W\. Zhu \(2023\)DeepLogic: joint learning of neural perception and logical reasoning\.IEEE Transactions on Pattern Analysis and Machine Intelligence45\(4\),pp\. 4321–4334\.Cited by:[§6](https://arxiv.org/html/2606.18509#S6.p3.1)\.
- R\. A\. Fisher \(1935\)The Design of Experiments\.The Design of Experiments,Oliver & Boyd,Oxford, England\.Cited by:[§4](https://arxiv.org/html/2606.18509#S4.p3.1)\.
- L\. Gao, T\. D\. la Tour, H\. Tillman, G\. Goh, R\. Troll, A\. Radford, I\. Sutskever, J\. Leike, and J\. Wu \(2025\)Scaling and evaluating sparse autoencoders\.InInternational Conference on Learning Representations,Cited by:[§A\.3](https://arxiv.org/html/2606.18509#A1.SS3.p1.1)\.
- K\. Huang, X\. Fu, and N\. D\. Sidiropoulos \(2016\)Anchor\-free correlated topic modeling: identifiability and algorithm\.Advances in Neural Information Processing Systems\.Cited by:[Table 2](https://arxiv.org/html/2606.18509#A6.T2.48.48.48.5),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.36.36.36.4)\.
- A\. Hyvarinen, H\. Sasaki, and R\. Turner \(2019\)Nonlinear ICA using auxiliary variables and generalized contrastive learning\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 859–868\.Cited by:[§A\.1](https://arxiv.org/html/2606.18509#A1.SS1.p1.1),[§E\.1](https://arxiv.org/html/2606.18509#A5.SS1.SSS0.Px1.p1.5),[§E\.1](https://arxiv.org/html/2606.18509#A5.SS1.SSS0.Px2.p1.3),[§E\.1](https://arxiv.org/html/2606.18509#A5.SS1.SSS0.Px3.1.p1.4),[Table 2](https://arxiv.org/html/2606.18509#A6.T2.8.8.8.6),[Table 3](https://arxiv.org/html/2606.18509#A6.T3),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.6.6.6.5),[§1](https://arxiv.org/html/2606.18509#S1.p1.1),[§1](https://arxiv.org/html/2606.18509#S1.p3.1),[§2\.2](https://arxiv.org/html/2606.18509#S2.SS2.p5.1),[§5](https://arxiv.org/html/2606.18509#S5.SS0.SSS0.Px1.p1.7),[Proposition 2](https://arxiv.org/html/2606.18509#Thmproposition2),[Proposition 2](https://arxiv.org/html/2606.18509#Thmproposition2.p1.4.1)\.
- J\. Jin and V\. Syrgkanis \(2024\)Learning linear causal representations from general environments: identifiability and intrinsic ambiguity\.InAdvances in Neural Information Processing Systems,pp\. 63466–63509\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1)\.
- I\. Khemakhem, D\. Kingma, R\. Monti, and A\. Hyvarinen \(2020a\)Variational autoencoders and nonlinear ICA: a unifying framework\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 2207–2217\.Cited by:[§A\.1](https://arxiv.org/html/2606.18509#A1.SS1.p1.1),[§A\.5](https://arxiv.org/html/2606.18509#A1.SS5.p1.1),[§E\.2](https://arxiv.org/html/2606.18509#A5.SS2.SSS0.Px1.p1.12),[Table 2](https://arxiv.org/html/2606.18509#A6.T2.12.12.12.5),[Table 3](https://arxiv.org/html/2606.18509#A6.T3),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.9.9.9.4),[§1](https://arxiv.org/html/2606.18509#S1.p1.1),[§1](https://arxiv.org/html/2606.18509#S1.p3.1),[§1](https://arxiv.org/html/2606.18509#S1.p4.1),[§2\.2](https://arxiv.org/html/2606.18509#S2.SS2.p5.1),[§5](https://arxiv.org/html/2606.18509#S5.SS0.SSS0.Px1.p1.7),[Proposition 3](https://arxiv.org/html/2606.18509#Thmproposition3),[Proposition 3](https://arxiv.org/html/2606.18509#Thmproposition3.p1.10.3)\.
- I\. Khemakhem, R\. Monti, D\. Kingma, and A\. Hyvarinen \(2020b\)ICE\-BeeM: identifiable conditional energy\-based deep models based on nonlinear ICA\.Advances in Neural Information Processing Systems,pp\. 12768–12778\.Cited by:[§A\.1](https://arxiv.org/html/2606.18509#A1.SS1.p1.1),[§E\.2](https://arxiv.org/html/2606.18509#A5.SS2.SSS0.Px1.p1.12),[Table 2](https://arxiv.org/html/2606.18509#A6.T2.16.16.16.5),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.12.12.12.4),[§2\.2](https://arxiv.org/html/2606.18509#S2.SS2.p5.1)\.
- B\. Kim, M\. Wattenberg, J\. Gilmer, C\. Cai, J\. Wexler, F\. Viegas,et al\.\(2018\)Interpretability beyond feature attribution: quantitative testing with concept activation vectors \(TCAV\)\.InInternational Conference on Machine Learning,pp\. 2668–2677\.Cited by:[§A\.3](https://arxiv.org/html/2606.18509#A1.SS3.p1.1)\.
- L\. Kong, S\. Xie, W\. Yao, Y\. Zheng, G\. Chen, P\. Stojanov, V\. Akinwande, and K\. Zhang \(2022\)Partial disentanglement for domain adaptation\.InInternational Conference on Machine Learning,pp\. 11455–11472\.Cited by:[§A\.4](https://arxiv.org/html/2606.18509#A1.SS4.p1.1)\.
- S\. Lachapelle, D\. Mahajan, I\. Mitliagkas, and S\. Lacoste\-Julien \(2023\)Additive decoders for latent variables identification and cartesian\-product extrapolation\.InAdvances in Neural Information Processing Systems,pp\. 25112–25150\.Cited by:[§A\.4](https://arxiv.org/html/2606.18509#A1.SS4.p1.1)\.
- S\. Lachapelle, P\. Rodriguez, Y\. Sharma, K\. E\. Everett, R\. L\. Priol, A\. Lacoste, and S\. Lacoste\-Julien \(2022\)Disentanglement via mechanism sparsity regularization: a new principle for nonlinear ICA\.InConference on Causal Learning and Reasoning,pp\. 428–484\.Cited by:[§A\.3](https://arxiv.org/html/2606.18509#A1.SS3.p1.1)\.
- P\. Lippe, S\. Magliacane, S\. Löwe, Y\. M\. Asano, T\. Cohen, and E\. Gavves \(2022\)CITRIS: causal identifiability from temporal intervened sequences\.InInternational Conference on Machine Learning,pp\. 13557–13603\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1)\.
- P\. Lippe, S\. Magliacane, S\. Löwe, Y\. M\. Asano, T\. Cohen, and E\. Gavves \(2023\)BISCUIT: causal representation learning from binary interactions\.InConference on Uncertainty in Artificial Intelligence,pp\. 1263–1273\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1)\.
- F\. Locatello, S\. Bauer, M\. Lucic, G\. Rätsch, S\. Gelly, B\. Schölkopf, and O\. Bachem \(2019\)Challenging common assumptions in the unsupervised learning of disentangled representations\.InInternational Conference on Machine Learning,pp\. 4114–4124\.Cited by:[§A\.1](https://arxiv.org/html/2606.18509#A1.SS1.p1.1)\.
- R\. O’Donnell \(2008\)Some topics in analysis of boolean functions\.InProceedings of the Fortieth Annual ACM Symposium on Theory of Computing,pp\. 569–578\.Cited by:[§4](https://arxiv.org/html/2606.18509#S4.p3.1)\.
- T\. Oikarinen, S\. Das, L\. M\. Nguyen, and T\. Weng \(2023\)Label\-free concept bottleneck models\.InInternational Conference on Learning Representations,Cited by:[§A\.3](https://arxiv.org/html/2606.18509#A1.SS3.p1.1),[§6](https://arxiv.org/html/2606.18509#S6.p3.1)\.
- J\. Peters, D\. Janzing, and B\. Schölkopf \(2017\)Elements of Causal Inference: Foundations and Learning Algorithms\.Adaptive Computation and Machine Learning Series,MIT Press,Cambridge, MA, USA\.Cited by:[§5](https://arxiv.org/html/2606.18509#S5.SS0.SSS0.Px2.p3.15)\.
- G\. Rajendran, S\. Buchholz, B\. Aragam, B\. Schölkopf, and P\. K\. Ravikumar \(2024\)From causal to concept\-based representation learning\.InAdvances in Neural Information Processing Systems,pp\. 101250–101296\.Cited by:[§A\.3](https://arxiv.org/html/2606.18509#A1.SS3.p1.1),[Table 2](https://arxiv.org/html/2606.18509#A6.T2.52.52.52.6),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.39.39.39.5)\.
- P\. Reizinger, S\. Guo, F\. Huszár, B\. Schölkopf, and W\. Brendel \(2025\)Identifiable exchangeable mechanisms for causal structure and representation learning\.InInternational Conference on Learning Representations,Cited by:[§A\.5](https://arxiv.org/html/2606.18509#A1.SS5.p1.1),[§1](https://arxiv.org/html/2606.18509#S1.p4.1)\.
- G\. Rota \(1964\)On the foundations of combinatorial theory I\. Theory of Möbius Functions\.Zeitschrift für Wahrscheinlichkeitstheorie und Verwandte Gebiete2,pp\. 340–368\.Cited by:[§C\.6](https://arxiv.org/html/2606.18509#A3.SS6.2.p2.3)\.
- S\. Saengkyongam, E\. Rosenfeld, P\. K\. Ravikumar, N\. Pfister, and J\. Peters \(2024\)Identifying representations for intervention extrapolation\.InInternational Conference on Learning Representations,Cited by:[§A\.4](https://arxiv.org/html/2606.18509#A1.SS4.p1.1)\.
- T\. Schmidt, S\. Schneider, and M\. Bethge \(2025\)Equivariance by contrast: identifiable equivariant embeddings from unlabeled finite group actions\.InAdvances in Neural Information Processing Systems,pp\. 71760–71800\.Cited by:[Table 2](https://arxiv.org/html/2606.18509#A6.T2.56.56.56.5),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.42.42.42.4),[§2\.2](https://arxiv.org/html/2606.18509#S2.SS2.p5.1)\.
- B\. Schölkopf, F\. Locatello, S\. Bauer, N\. R\. Ke, N\. Kalchbrenner, A\. Goyal, and Y\. Bengio \(2021\)Toward causal representation learning\.Proceedings of the IEEE109\(5\),pp\. 612–634\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1),[§1](https://arxiv.org/html/2606.18509#S1.p1.1),[§5](https://arxiv.org/html/2606.18509#S5.SS0.SSS0.Px2.p3.15)\.
- C\. Squires and P\. Ravikumar \(2026\)A unifying framework for unsupervised concept extraction\.External Links:2604\.24936,[Link](https://arxiv.org/abs/2604.24936)Cited by:[§A\.3](https://arxiv.org/html/2606.18509#A1.SS3.p1.1),[§A\.5](https://arxiv.org/html/2606.18509#A1.SS5.p1.1),[2nd item](https://arxiv.org/html/2606.18509#S1.I1.i2.p1.1),[§1](https://arxiv.org/html/2606.18509#S1.p6.1),[§3\.1](https://arxiv.org/html/2606.18509#S3.SS1.p2.1),[§3\.1](https://arxiv.org/html/2606.18509#S3.SS1.p5.1),[§3\.1](https://arxiv.org/html/2606.18509#S3.SS1.p6.2),[§3\.1](https://arxiv.org/html/2606.18509#S3.SS1.p7.1),[§3](https://arxiv.org/html/2606.18509#S3.p1.4)\.
- C\. Squires, A\. Seigal, S\. S\. Bhate, and C\. Uhler \(2023\)Linear causal disentanglement via interventions\.InInternational Conference on Machine Learning,pp\. 32540–32560\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1),[Appendix E](https://arxiv.org/html/2606.18509#A5.SS0.SSS0.Px2.p2.4),[§E\.3](https://arxiv.org/html/2606.18509#A5.SS3.SSS0.Px1.p1.8),[§E\.3](https://arxiv.org/html/2606.18509#A5.SS3.SSS0.Px2.p1.4),[§E\.3](https://arxiv.org/html/2606.18509#A5.SS3.SSS0.Px3.1.p1.9),[Table 2](https://arxiv.org/html/2606.18509#A6.T2.20.20.20.6),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.15.15.15.5),[§1](https://arxiv.org/html/2606.18509#S1.p3.1),[§2\.2](https://arxiv.org/html/2606.18509#S2.SS2.p5.1),[Proposition 4](https://arxiv.org/html/2606.18509#Thmproposition4),[Proposition 4](https://arxiv.org/html/2606.18509#Thmproposition4.p1.3.1),[Proposition 5](https://arxiv.org/html/2606.18509#Thmproposition5.p1.9.2),[Remark 2](https://arxiv.org/html/2606.18509#Thmremark2.p1.1.1)\.
- B\. Varıcı, E\. Acartürk, K\. Shanmugam, A\. Kumar, and A\. Tajer \(2025\)Score\-based causal representation learning: linear and general transformations\.Journal of Machine Learning Research26\(112\),pp\. 1–90\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1),[§E\.5](https://arxiv.org/html/2606.18509#A5.SS5.SSS0.Px1.p1.2),[§E\.5](https://arxiv.org/html/2606.18509#A5.SS5.SSS0.Px3.1.p1.3),[Table 2](https://arxiv.org/html/2606.18509#A6.T2.32.32.32.5),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.24.24.24.4),[§1](https://arxiv.org/html/2606.18509#S1.p3.1),[§2\.2](https://arxiv.org/html/2606.18509#S2.SS2.p5.1),[§5](https://arxiv.org/html/2606.18509#S5.SS0.SSS0.Px2.p2.13),[Proposition 6](https://arxiv.org/html/2606.18509#Thmproposition6),[Proposition 6](https://arxiv.org/html/2606.18509#Thmproposition6.p1.11.2)\.
- B\. Varıcı, E\. Acartürk, K\. Shanmugam, and A\. Tajer \(2024a\)General identifiability and achievability for causal representation learning\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 2314–2322\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1),[§1](https://arxiv.org/html/2606.18509#S1.p3.1)\.
- B\. Varıcı, E\. Acartürk, K\. Shanmugam, and A\. Tajer \(2024b\)Linear causal representation learning from unknown multi\-node interventions\.Advances in Neural Information Processing Systems,pp\. 111614–111648\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1),[§5](https://arxiv.org/html/2606.18509#S5.SS0.SSS0.Px2.p2.13)\.
- J\. von Kügelgen, M\. Besserve, W\. Liang, L\. Gresele, A\. Kekić, E\. Bareinboim, D\. Blei, and B\. Schölkopf \(2023\)Nonparametric identifiability of causal representations from unknown interventions\.InAdvances in Neural Information Processing Systems,pp\. 48603–48638\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1),[§E\.6](https://arxiv.org/html/2606.18509#A5.SS6.SSS0.Px1.p1.16),[§E\.6](https://arxiv.org/html/2606.18509#A5.SS6.SSS0.Px2.p1.5),[§E\.6](https://arxiv.org/html/2606.18509#A5.SS6.SSS0.Px3.1.p1.3),[Table 2](https://arxiv.org/html/2606.18509#A6.T2.28.28.28.5),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.21.21.21.4),[§1](https://arxiv.org/html/2606.18509#S1.p3.1),[§2\.2](https://arxiv.org/html/2606.18509#S2.SS2.p5.1),[§5](https://arxiv.org/html/2606.18509#S5.SS0.SSS0.Px2.p2.13),[Proposition 7](https://arxiv.org/html/2606.18509#Thmproposition7),[Proposition 7](https://arxiv.org/html/2606.18509#Thmproposition7.p1.12.2),[Proposition 7](https://arxiv.org/html/2606.18509#Thmproposition7.p1.14.2)\.
- J\. von Kügelgen, X\. Shen, J\. Ketterer, N\. Meinshausen, and J\. Peters \(2025\)Representation learning for distributional perturbation extrapolation\.InLearning Meaningful Representations of Life \(LMRL\) Workshop at ICLR 2025,Cited by:[§A\.4](https://arxiv.org/html/2606.18509#A1.SS4.p1.1),[§E\.7](https://arxiv.org/html/2606.18509#A5.SS7.SSS0.Px1.p1.4),[§E\.7](https://arxiv.org/html/2606.18509#A5.SS7.SSS0.Px3.1.p1.2),[Table 2](https://arxiv.org/html/2606.18509#A6.T2.40.40.40.5),[Table 3](https://arxiv.org/html/2606.18509#A6.T3.30.30.30.4),[§1](https://arxiv.org/html/2606.18509#S1.p3.1),[§2\.2](https://arxiv.org/html/2606.18509#S2.SS2.p5.1),[§5](https://arxiv.org/html/2606.18509#S5.SS0.SSS0.Px3.p1.11),[Proposition 8](https://arxiv.org/html/2606.18509#Thmproposition8.p1.6.2)\.
- D\. Xu, D\. Yao, S\. Lachapelle, P\. Taslakian, J\. von Kügelgen, F\. Locatello, and S\. Magliacane \(2024\)A sparsity principle for partially observable causal representation learning\.InInternational Conference on Machine Learning,Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1)\.
- D\. Yao, D\. Rancati, R\. Cadei, M\. Fumero, and F\. Locatello \(2025\)Unifying causal representation learning with the invariance principle\.InInternational Conference on Learning Representations,Cited by:[§A\.5](https://arxiv.org/html/2606.18509#A1.SS5.p1.1),[§1](https://arxiv.org/html/2606.18509#S1.p4.1)\.
- D\. Yao, D\. Xu, S\. Lachapelle, S\. Magliacane, P\. Taslakian, G\. Martius, J\. von Kügelgen, and F\. Locatello \(2024\)Multi\-view causal representation learning with partial observability\.InInternational Conference on Learning Representations,Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1)\.
- J\. Zhang, K\. Greenewald, C\. Squires, A\. Srivastava, K\. Shanmugam, and C\. Uhler \(2023\)Identifiability guarantees for causal disentanglement from soft interventions\.Advances in Neural Information Processing Systems,pp\. 50254–50292\.Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1)\.
- K\. Zhang, S\. Xie, I\. Ng, and Y\. Zheng \(2024\)Causal representation learning from multiple distributions: a general setting\.InInternational Conference on Machine Learning,Cited by:[§A\.2](https://arxiv.org/html/2606.18509#A1.SS2.p1.1)\.
- Y\. Zheng, S\. Xie, and K\. Zhang \(2025\)Nonparametric identification of latent concepts\.InInternational Conference on Machine Learning,pp\. 78374–78404\.Cited by:[§A\.3](https://arxiv.org/html/2606.18509#A1.SS3.p1.1)\.
###### Contents of Appendix
1. [1Introduction](https://arxiv.org/html/2606.18509#S1)
2. [2A unifying framework: Concept modulation models](https://arxiv.org/html/2606.18509#S2)1. [2\.1Preliminaries and notation](https://arxiv.org/html/2606.18509#S2.SS1) 2. [2\.2Concept modulation models](https://arxiv.org/html/2606.18509#S2.SS2)
3. [3Conditional transition\-based identifiability and attribute potentials](https://arxiv.org/html/2606.18509#S3)1. [3\.1A transition\-intersection criterion](https://arxiv.org/html/2606.18509#S3.SS1) 2. [3\.2Characterizing concept\-side transitions](https://arxiv.org/html/2606.18509#S3.SS2)
4. [4Extrapolation via attribute potentials](https://arxiv.org/html/2606.18509#S4)
5. [5Applications: Recoveries and examples](https://arxiv.org/html/2606.18509#S5)
6. [6Discussion](https://arxiv.org/html/2606.18509#S6)
7. [References](https://arxiv.org/html/2606.18509#bib)
8. [ARelated works](https://arxiv.org/html/2606.18509#A1)1. [A\.1Identifiability in latent\-variable representation learning](https://arxiv.org/html/2606.18509#A1.SS1) 2. [A\.2Causal representation learning](https://arxiv.org/html/2606.18509#A1.SS2) 3. [A\.3Concept extraction](https://arxiv.org/html/2606.18509#A1.SS3) 4. [A\.4Extrapolation](https://arxiv.org/html/2606.18509#A1.SS4) 5. [A\.5Unifying frameworks](https://arxiv.org/html/2606.18509#A1.SS5)
9. [BPreliminaries](https://arxiv.org/html/2606.18509#A2)1. [B\.1Notation summary](https://arxiv.org/html/2606.18509#A2.SS1) 2. [B\.2Measure\-theoretic background](https://arxiv.org/html/2606.18509#A2.SS2)
10. [CProofs](https://arxiv.org/html/2606.18509#A3)1. [C\.1Proof of∼𝔊\\sim\_\{\\mathfrak\{G\}\}being equivalence relation](https://arxiv.org/html/2606.18509#A3.SS1) 2. [C\.2Proof ofThm\.1](https://arxiv.org/html/2606.18509#A3.SS2) 3. [C\.3Proof ofThm\.2](https://arxiv.org/html/2606.18509#A3.SS3) 4. [C\.4Proof ofThm\.3](https://arxiv.org/html/2606.18509#A3.SS4) 5. [C\.5Proof ofThm\.4](https://arxiv.org/html/2606.18509#A3.SS5) 6. [C\.6Proof ofCor\.1](https://arxiv.org/html/2606.18509#A3.SS6)
11. [DDeferred details](https://arxiv.org/html/2606.18509#A4)1. [D\.1Details for the running example](https://arxiv.org/html/2606.18509#A4.SS1)
12. [ERecoveries of representative prior work](https://arxiv.org/html/2606.18509#A5)1. [E\.1Nonlinear ICA](https://arxiv.org/html/2606.18509#A5.SS1) 2. [E\.2Conditionally exponential families](https://arxiv.org/html/2606.18509#A5.SS2) 3. [E\.3Linear CRL](https://arxiv.org/html/2606.18509#A5.SS3) 4. [E\.4Gaussian CRL](https://arxiv.org/html/2606.18509#A5.SS4) 5. [E\.5Score\-based CRL](https://arxiv.org/html/2606.18509#A5.SS5) 6. [E\.6Nonparametric CRL](https://arxiv.org/html/2606.18509#A5.SS6) 7. [E\.7Perturbation modeling](https://arxiv.org/html/2606.18509#A5.SS7)
13. [FCMM Translations of prior work](https://arxiv.org/html/2606.18509#A6)
## Appendix ARelated works
### A\.1Identifiability in latent\-variable representation learning
Identifiability in latent\-variable representation learning asks when latent structure is determined by observed data, up to an allowable ambiguity class such as permutation, component\-wise transformations, or affine transformations\. A canonical example is nonlinear ICA, where statistically independent latent components are observed only through a nonlinear mixing map\. Without additional structure, nonlinear ICA and unsupervised disentanglement are generally unidentifiable\[Locatelloet al\.,[2019](https://arxiv.org/html/2606.18509#bib.bib48)\]\. A central way to restore identifiability is to use auxiliary or conditioning information\.Hyvarinenet al\.\[[2019](https://arxiv.org/html/2606.18509#bib.bib4)\]show that an observed auxiliary variable can identify nonlinear ICA representations when latent components are conditionally independent and sufficiently modulated by the auxiliary variable\.Khemakhemet al\.\[[2020a](https://arxiv.org/html/2606.18509#bib.bib2)\]develop this principle in a variational latent\-variable framework, using a condition\-dependent exponential\-family prior and an injective decoder to obtain identifiability\.Khemakhemet al\.\[[2020b](https://arxiv.org/html/2606.18509#bib.bib17)\]instantiate the same idea in conditional energy\-based models, where structured conditioning of the energy function yields identifiable representations\. A complementary view derives identifiability from restrictions on latent mechanisms and their symmetries rather than from a specific exponential\-family form\[Ahujaet al\.,[2022b](https://arxiv.org/html/2606.18509#bib.bib49)\]\. Our framework follows this broader identifiability perspective, but studies conditional generative models indexed by attributes and asks both which latent concept structure is identifiable from observed attributes and when this structure extrapolates to unseen attributes\.
### A\.2Causal representation learning
Causal representation learning \(CRL\) seeks to recover latent causal variables, and sometimes their causal graph, from high\-dimensional observations\[Schölkopfet al\.,[2021](https://arxiv.org/html/2606.18509#bib.bib1)\]\. Interventional CRL uses environments or interventions as the source of variation, but the broader literature also exploits temporal structure, paired intervention data, distribution shifts, multi\-view observations, and partial observability\. In the static interventional setting,Squireset al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib6)\]study linear latent causal models under injective linear mixing and show that one perfect intervention per latent variable is sufficient, and in a worst\-case sense necessary, for identifying the latent causal model and the mixing map\.Ahujaet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib5)\]study interventional CRL under perfect and imperfect interventions, showing how intervention\-induced geometric structure can identify latent causal variables up to appropriate ambiguities\.Buchholzet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib18)\]allow general nonlinear mixing while retaining a linear\-Gaussian latent causal model, and obtain identifiability from unpaired interventional data with unknown single\-node intervention targets\.von Kügelgenet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib8)\]move to nonparametric latent causal models and diffeomorphic mixing under unknown interventions, clarifying both positive identifiability results and residual equivalence classes\. Other work relaxes or changes the intervention model, including soft interventions\[Zhanget al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib32)\], uncoupled hard interventions and general nonparametric models\[Varıcıet al\.,[2024a](https://arxiv.org/html/2606.18509#bib.bib33)\], unknown multi\-node interventions under linear transformations\[Varıcıet al\.,[2024b](https://arxiv.org/html/2606.18509#bib.bib23)\], and score\-based identification under linear and general nonlinear transformations\[Varıcıet al\.,[2025](https://arxiv.org/html/2606.18509#bib.bib9)\]\. Complementary CRL results use weaker or different supervision signals\.Brehmeret al\.\[[2022](https://arxiv.org/html/2606.18509#bib.bib50)\]identify latent causal variables and structure from paired pre\- and post\-intervention observations, generated from a shared exogenous noise, with random unknown interventions\.Lippeet al\.\[[2022](https://arxiv.org/html/2606.18509#bib.bib51)\]use temporal sequences with intervention targets to identify causal factors, whileLippeet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib52)\]replaces observed targets with binary interaction variables\. Multi\-view and multi\-distribution approaches show that explicit intervention labels are not the only route to identifiability, with guarantees under partial observability\[Yaoet al\.,[2024](https://arxiv.org/html/2606.18509#bib.bib53), Xuet al\.,[2024](https://arxiv.org/html/2606.18509#bib.bib54)\]and general distribution shifts\[Zhanget al\.,[2024](https://arxiv.org/html/2606.18509#bib.bib55)\]\. Recent work also studies intrinsic ambiguity and finite\-sample behavior in interventional CRL\[Jin and Syrgkanis,[2024](https://arxiv.org/html/2606.18509#bib.bib56), Acartürket al\.,[2024](https://arxiv.org/html/2606.18509#bib.bib57)\]\. These works provide model\-specific identifiability guarantees under particular sources of variation, whereas CMMs isolate a generic lifting step from observed feature agreement to constrained latent concept transitions and leave model\-specific rigidity to separate assumptions\.
### A\.3Concept extraction
Our use of the term “concept” is related to work on extracting latent concepts or factors from observed representations, but our goal is different\. Concept\-extraction methods in interpretability often aim to find human\-facing features in neural activations using examples, labels, sparsity, or dictionary\-learning objectives\[Kimet al\.,[2018](https://arxiv.org/html/2606.18509#bib.bib39), Oikarinenet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib31), Brickenet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib35), Gaoet al\.,[2025](https://arxiv.org/html/2606.18509#bib.bib36)\]\. Recent theoretical work instead asks when such latent concepts are identifiable\.Cuiet al\.\[[2026](https://arxiv.org/html/2606.18509#bib.bib37)\]analyze identifiability limits for sparse autoencoders,Zhenget al\.\[[2025](https://arxiv.org/html/2606.18509#bib.bib38)\]give nonparametric guarantees for identifying latent concepts, andSquires and Ravikumar \[[2026](https://arxiv.org/html/2606.18509#bib.bib16)\]provide a transition\-based framework for unsupervised concept extraction\.Rajendranet al\.\[[2024](https://arxiv.org/html/2606.18509#bib.bib7)\]connect CRL and concept\-based representation learning by relaxing recovery of true causal variables toward recovery of task\-relevant concepts represented in latent space\.Lachapelleet al\.\[[2022](https://arxiv.org/html/2606.18509#bib.bib40)\]show that mechanism sparsity can also drive disentanglement in nonlinear ICA, illustrating how sparse modulation of latent components can reduce non\-identifiability\. CMMs differ from these works by making the concept law attribute\-indexed: attributes modulate latent concept distributions, so the central questions are not only concept identifiability from observed attributes but also extrapolation of concept and feature laws to unseen attributes\.
### A\.4Extrapolation
A growing line of work studies when models learned from observed conditions can predict behavior under unseen interventions, perturbations, or attribute values\.Lachapelleet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib43)\]study additive decoders and show that additive latent structure can support Cartesian\-product extrapolation by recombining observed factors of variation\.Agarwalet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib22)\]study combinatorial interventions from a causal\-inference perspective, using low\-rank and Fourier\-sparsity structure to identify potential outcomes for unseen intervention combinations\.Du and Kaelbling \[[2024](https://arxiv.org/html/2606.18509#bib.bib41)\]discuss compositional generative modeling more broadly, emphasizing how modular structure can support generalization to unseen combinations\. In latent\-variable settings,Saengkyongamet al\.\[[2024](https://arxiv.org/html/2606.18509#bib.bib42)\]establish affine identifiability of latent representations under linear intervention effects and use this to certify extrapolation to out\-of\-support intervention values for downstream outcomes\.von Kügelgenet al\.\[[2025](https://arxiv.org/html/2606.18509#bib.bib13)\]study distributional perturbation extrapolation under latent mean\-shift models and prove representation\-identifiability and perturbation\-extrapolation guarantees\.Konget al\.\[[2022](https://arxiv.org/html/2606.18509#bib.bib44)\]study domain adaptation under partial disentanglement, where target\-domain observations are available and the goal is prediction rather than identification of a full unseen interventional distribution\. These works give extrapolation guarantees in specific structural regimes\. CMMs instead formulate extrapolation as an attribute\-indexed consistency question: agreement on observed conditional feature laws extrapolates to an unseen attribute exactly when the transported attribute\-potential identities extend to that attribute\.
### A\.5Unifying frameworks
Several works abstract away from individual identifiability theorems to identify common proof principles\.Khemakhemet al\.\[[2020a](https://arxiv.org/html/2606.18509#bib.bib2)\]unify nonlinear ICA and identifiable VAEs through condition\-dependent exponential\-family latent priors, but this unification is tied to conditional factorization and sufficient variability\.Ahujaet al\.\[[2022b](https://arxiv.org/html/2606.18509#bib.bib49)\]give a mechanism\-based and equivariance view of identifiable representation learning, characterizing residual ambiguity through symmetries shared by the latent mechanisms\.Yaoet al\.\[[2025](https://arxiv.org/html/2606.18509#bib.bib58)\]propose an invariance\-principle view of CRL, showing how several CRL guarantees can be understood through representation alignment with environment\-induced invariances\.Reizingeret al\.\[[2025](https://arxiv.org/html/2606.18509#bib.bib14)\]introduce identifiable exchangeable mechanisms as a probabilistic framework connecting ICA, CRL, and causal structure learning through exchangeable non\-i\.i\.d\. data\.Squires and Ravikumar \[[2026](https://arxiv.org/html/2606.18509#bib.bib16)\]develop a unified framework for unsupervised concept extraction, where identifiability reduces to showing that the intersection of valid transition sets lies inside the allowable ambiguity class\. These frameworks clarify identifiability within particular structural regimes\. CMMs are complementary: they provide a conditional framework in which observed feature agreement induces constrained latent concept transitions, and the same attribute\-potential constraints characterize extrapolation to unseen attributes\.
## Appendix BPreliminaries
### B\.1Notation summary
We provide in[Table1](https://arxiv.org/html/2606.18509#A2.T1)a summary of notation used throughout the appendix\.
Table 1:Notation summary\.
### B\.2Measure\-theoretic background
We use the measure\-theoretic notation introduced in the main text\. The only additional convention needed in the appendix concerns restrictions of Markov kernels\. For a measurable space𝒰\\mathcal\{U\}, writeℬ\(𝒰\)\\mathcal\{B\}\(\\mathcal\{U\}\)for itsσ\\sigma\-algebra\. If𝒰′∈ℬ\(𝒰\)\\mathcal\{U\}^\{\\prime\}\\in\\mathcal\{B\}\(\\mathcal\{U\}\), then𝒰′\\mathcal\{U\}^\{\\prime\}is equipped with the traceσ\\sigma\-algebra
ℬ\(𝒰′\):=\{E∩𝒰′\|E∈ℬ\(𝒰\)\}\.\\mathcal\{B\}\(\\mathcal\{U\}^\{\\prime\}\)\\mathrel\{:=\}\\\{E\\cap\\mathcal\{U\}^\{\\prime\}\\;\|\\;E\\in\\mathcal\{B\}\(\\mathcal\{U\}\)\\\}\.If𝐊∈𝒦mk\(𝒰→𝒱\)\{\\mathbf\{K\}\}\\in\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{U\}\\to\\mathcal\{V\}\), then𝐊𝒰′∈𝒦mk\(𝒰′→𝒱\)\{\\mathbf\{K\}\}\_\{\\mathcal\{U\}^\{\\prime\}\}\\in\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{U\}^\{\\prime\}\\to\\mathcal\{V\}\)denotes the domain restriction of𝐊\{\\mathbf\{K\}\}, defined by
𝐊𝒰′\(B∣u\):=𝐊\(B∣u\),u∈𝒰′,B∈ℬ\(𝒱\)\.\{\\mathbf\{K\}\}\_\{\\mathcal\{U\}^\{\\prime\}\}\(B\\mid u\)\\mathrel\{:=\}\{\\mathbf\{K\}\}\(B\\mid u\),\\qquad u\\in\\mathcal\{U\}^\{\\prime\},\\qquad B\\in\\mathcal\{B\}\(\\mathcal\{V\}\)\.This is well defined because, for eachB∈ℬ\(𝒱\)B\\in\\mathcal\{B\}\(\\mathcal\{V\}\), the mapu↦𝐊\(B∣u\)u\\mapsto\{\\mathbf\{K\}\}\(B\\mid u\)isℬ\(𝒰\)\\mathcal\{B\}\(\\mathcal\{U\}\)\-measurable, and its restriction to𝒰′\\mathcal\{U\}^\{\\prime\}isℬ\(𝒰′\)\\mathcal\{B\}\(\\mathcal\{U\}^\{\\prime\}\)\-measurable\. Thus𝐊𝒰′\{\\mathbf\{K\}\}\_\{\\mathcal\{U\}^\{\\prime\}\}restricts only the source measurable space; it does not restrict the target space and involves no renormalization\.
## Appendix CProofs
### C\.1Proof of∼𝔊\\sim\_\{\\mathfrak\{G\}\}being equivalence relation
###### Proposition 1\(∼𝔊\\sim\_\{\\mathfrak\{G\}\}is an equivalence relation\)\.
Fix a CMM classℳ=𝒬×\{𝐁\}×𝒦\\mathcal\{M\}=\\mathcal\{Q\}\\times\\\{\{\\mathbf\{B\}\}\\\}\\times\\mathcal\{K\}, an observed attribute set𝒜o⊆𝒜\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\subseteq\\mathcal\{A\}, and a transition group𝔊⊆Autν\(𝒞\)\\mathfrak\{G\}\\subseteq\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\)\. For𝖬=\(𝐐,𝐁,𝐊\)\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\)and𝖬′=\(𝐐′,𝐁,𝐊′\)\\mathsf\{M\}^\{\\prime\}=\(\{\\mathbf\{Q\}\}^\{\\prime\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}^\{\\prime\}\), define𝖬∼𝔊𝖬′\\mathsf\{M\}\\sim\_\{\\mathfrak\{G\}\}\\mathsf\{M\}^\{\\prime\}on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}if there existsτ∈𝔊\\tau\\in\\mathfrak\{G\}such that
𝐊′=𝜈𝐊𝐓τ,𝐐¯\(⋅∣a\)=𝐓τ𝐐¯′\(⋅∣a\)∀a∈𝒜o,\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\},\\qquad\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)\\quad\\forall a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\},where𝐐¯=𝐁𝐐\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}and𝐐¯′=𝐁𝐐′\\bar\{\\mathbf\{Q\}\}^\{\\prime\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}^\{\\prime\}\. Then∼𝔊\\sim\_\{\\mathfrak\{G\}\}is an equivalence relation onℳ\\mathcal\{M\}restricted to𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\.
###### Proof\.
Fix a CMM classℳ=𝒬×\{𝐁\}×𝒦\\mathcal\{M\}=\\mathcal\{Q\}\\times\\\{\{\\mathbf\{B\}\}\\\}\\times\\mathcal\{K\}, an observed attribute set𝒜o⊆𝒜\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\subseteq\\mathcal\{A\}, and a transition group𝔊⊆Autν\(𝒞\)\\mathfrak\{G\}\\subseteq\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\)\. Recall that𝖬=\(𝐐,𝐁,𝐊\)\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\)and𝖬′=\(𝐐′,𝐁,𝐊′\)\\mathsf\{M\}^\{\\prime\}=\(\{\\mathbf\{Q\}\}^\{\\prime\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}^\{\\prime\}\)satisfy𝖬∼𝔊𝖬′\\mathsf\{M\}\\sim\_\{\\mathfrak\{G\}\}\\mathsf\{M\}^\{\\prime\}on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}if there existsτ∈𝔊\\tau\\in\\mathfrak\{G\}such that
𝐊′=𝜈𝐊𝐓τ,𝐐¯\(⋅∣a\)=𝐓τ𝐐¯′\(⋅∣a\)∀a∈𝒜o,\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\},\\qquad\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)\\quad\\forall a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\},where𝐐¯=𝐁𝐐\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}and𝐐¯′=𝐁𝐐′\\bar\{\\mathbf\{Q\}\}^\{\\prime\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}^\{\\prime\}\.
We first prove reflexivity\. Since𝔊\\mathfrak\{G\}is a group,id𝒞∈𝔊\\operatorname\{id\}\_\{\\mathcal\{C\}\}\\in\\mathfrak\{G\}\. For every𝖬=\(𝐐,𝐁,𝐊\)\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\), we have𝐊=𝜈𝐊𝐓id𝒞\{\\mathbf\{K\}\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\operatorname\{id\}\_\{\\mathcal\{C\}\}\}and𝐐¯\(⋅∣a\)=𝐓id𝒞𝐐¯\(⋅∣a\)\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\operatorname\{id\}\_\{\\mathcal\{C\}\}\}\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)for alla∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\. Hence𝖬∼𝔊𝖬\\mathsf\{M\}\\sim\_\{\\mathfrak\{G\}\}\\mathsf\{M\}\.
We next prove symmetry\. Suppose𝖬∼𝔊𝖬′\\mathsf\{M\}\\sim\_\{\\mathfrak\{G\}\}\\mathsf\{M\}^\{\\prime\}throughτ∈𝔊\\tau\\in\\mathfrak\{G\}\. Thenτ−1∈𝔊\\tau^\{\-1\}\\in\\mathfrak\{G\}\. For anyμ≪ν\\mu\\ll\{\\nu\}, sinceτ−1∈Autν\(𝒞\)\\tau^\{\-1\}\\in\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\), we have𝐓τ−1μ≪ν\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\\mu\\ll\{\\nu\}, and therefore
𝐊′𝐓τ−1μ=𝐊𝐓τ𝐓τ−1μ=𝐊μ\.\{\\mathbf\{K\}\}^\{\\prime\}\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\\mu=\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\\mu=\{\\mathbf\{K\}\}\\mu\.Thus𝐊=𝜈𝐊′𝐓τ−1\{\\mathbf\{K\}\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}^\{\\prime\}\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\. Similarly, for everya∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\},
𝐓τ−1𝐐¯\(⋅∣a\)=𝐓τ−1𝐓τ𝐐¯′\(⋅∣a\)=𝐐¯′\(⋅∣a\),\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)=\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\),with equalities understood moduloν\{\\nu\}\-null sets\. Hence𝖬′∼𝔊𝖬\\mathsf\{M\}^\{\\prime\}\\sim\_\{\\mathfrak\{G\}\}\\mathsf\{M\}\.
Finally, we prove transitivity\. Suppose𝖬∼𝔊𝖬′\\mathsf\{M\}\\sim\_\{\\mathfrak\{G\}\}\\mathsf\{M\}^\{\\prime\}throughτ∈𝔊\\tau\\in\\mathfrak\{G\}, and𝖬′∼𝔊𝖬′′\\mathsf\{M\}^\{\\prime\}\\sim\_\{\\mathfrak\{G\}\}\\mathsf\{M\}^\{\\prime\\prime\}throughσ∈𝔊\\sigma\\in\\mathfrak\{G\}\. Write𝖬′′=\(𝐐′′,𝐁,𝐊′′\)\\mathsf\{M\}^\{\\prime\\prime\}=\(\{\\mathbf\{Q\}\}^\{\\prime\\prime\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}^\{\\prime\\prime\}\)and𝐐¯′′=𝐁𝐐′′\\bar\{\\mathbf\{Q\}\}^\{\\prime\\prime\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}^\{\\prime\\prime\}\. Since𝔊\\mathfrak\{G\}is a group,τ∘σ∈𝔊\\tau\\circ\\sigma\\in\\mathfrak\{G\}\. For everyμ≪ν\\mu\\ll\{\\nu\},𝐓σμ≪ν\\mathbf\{T\}\_\{\\sigma\}\\mu\\ll\{\\nu\}, so
𝐊′′μ=𝐊′𝐓σμ=𝐊𝐓τ𝐓σμ=𝐊𝐓τ∘σμ\.\{\\mathbf\{K\}\}^\{\\prime\\prime\}\\mu=\{\\mathbf\{K\}\}^\{\\prime\}\\mathbf\{T\}\_\{\\sigma\}\\mu=\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}\\mathbf\{T\}\_\{\\sigma\}\\mu=\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\\circ\\sigma\}\\mu\.Hence𝐊′′=𝜈𝐊𝐓τ∘σ\{\\mathbf\{K\}\}^\{\\prime\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\\circ\\sigma\}\. Moreover, for everya∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\},
𝐐¯\(⋅∣a\)=𝐓τ𝐐¯′\(⋅∣a\)=𝐓τ𝐓σ𝐐¯′′\(⋅∣a\)=𝐓τ∘σ𝐐¯′′\(⋅∣a\)\.\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau\}\\mathbf\{T\}\_\{\\sigma\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\\prime\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau\\circ\\sigma\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\\prime\}\(\\cdot\\mid a\)\.Thus𝖬∼𝔊𝖬′′\\mathsf\{M\}\\sim\_\{\\mathfrak\{G\}\}\\mathsf\{M\}^\{\\prime\\prime\}\. Therefore∼𝔊\\sim\_\{\\mathfrak\{G\}\}is reflexive, symmetric, and transitive, and hence is an equivalence relation onℳ\\mathcal\{M\}restricted to𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\. ∎
### C\.2Proof of[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)
[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)actually only requires the following condition, which is implied by[Definition6](https://arxiv.org/html/2606.18509#Thmdefinition6)when𝒜o≠∅\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\neq\\varnothing\.
###### Assumption 1\.
The conditional concept distributions induced by the CMM classℳ=𝒬×\{𝐁\}×𝒦\\mathcal\{M\}=\\mathcal\{Q\}\\times\\\{\{\\mathbf\{B\}\}\\\}\\times\\mathcal\{K\}are dominated by the reference measureν\{\\nu\}, i\.e\.,
𝐁𝐐\(⋅∣a\)≪ν∀𝐐∈𝒬,∀a∈𝒜\.\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)\\ll\{\\nu\}\\qquad\\forall\{\\mathbf\{Q\}\}\\in\\mathcal\{Q\},\\ \\forall a\\in\\mathcal\{A\}\.In addition, there exists a probability measureπ\\pion𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}such that
∫𝒜o𝐁𝐐\(⋅∣a\)π\(da\)∼ν∀𝐐∈𝒬,\\int\_\{\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\}\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)\\,\\pi\(da\)\\sim\{\\nu\}\\qquad\\forall\{\\mathbf\{Q\}\}\\in\\mathcal\{Q\},so that the observed attributes collectively cover the reference measureν\{\\nu\}up to its null sets\.
We now prove the following theorem, restated for convenience, under[1](https://arxiv.org/html/2606.18509#Thmassumption1)\. See[1](https://arxiv.org/html/2606.18509#Thmtheorem1)
###### Proof\.
Let𝖬′=\(𝐐′,𝐁,𝐊′\)∈ℳ\\mathsf\{M\}^\{\\prime\}=\(\{\\mathbf\{Q\}\}^\{\\prime\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}^\{\\prime\}\)\\in\\mathcal\{M\}be feature\-equivalent to𝖬=\(𝐐,𝐁,𝐊\)\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\)on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\. Write
μa:=𝐁𝐐\(⋅∣a\),μa′:=𝐁𝐐′\(⋅∣a\),\\mu\_\{a\}\\mathrel\{:=\}\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}\(\\cdot\\mid a\),\\qquad\\mu^\{\\prime\}\_\{a\}\\mathrel\{:=\}\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\),which satisfiesμa,μa′≪ν\\mu\_\{a\},\\mu^\{\\prime\}\_\{a\}\\ll\{\\nu\}for alla∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}by[1](https://arxiv.org/html/2606.18509#Thmassumption1)\. Byν\{\\nu\}\-Blackwell reducibility, choose bimeasurable embeddingsg,g′∈𝒢g,g^\{\\prime\}\\in\\mathcal\{G\}such that
𝐊=𝜈𝐊~𝐓g,𝐊′=𝜈𝐊~𝐓g′\.\{\\mathbf\{K\}\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g\},\\qquad\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g^\{\\prime\}\}\.Feature equivalence and carrier equality give, for everya∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\},
𝐊~𝐓gμa=𝐊~𝐓g′μa′\.\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g\}\\mu\_\{a\}=\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g^\{\\prime\}\}\\mu^\{\\prime\}\_\{a\}\.Since𝐊~\\widetilde\{\{\\mathbf\{K\}\}\}is injective on the relevant pushed\-forward carrier,
𝐓gμa=𝐓g′μa′∀a∈𝒜o\.\\mathbf\{T\}\_\{g\}\\mu\_\{a\}=\\mathbf\{T\}\_\{g^\{\\prime\}\}\\mu^\{\\prime\}\_\{a\}\\qquad\\forall a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\.Letπ\\pibe the full\-carrier anchor from[1](https://arxiv.org/html/2606.18509#Thmassumption1), and define
μ¯𝐐:=∫𝒜oμaπ\(da\),μ¯𝐐′:=∫𝒜oμa′π\(da\)\.\\overline\{\\mu\}\_\{\{\\mathbf\{Q\}\}\}\\mathrel\{:=\}\\int\_\{\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\}\\mu\_\{a\}\\,\\pi\(da\),\\qquad\\overline\{\\mu\}\_\{\{\\mathbf\{Q\}\}^\{\\prime\}\}\\mathrel\{:=\}\\int\_\{\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\}\\mu^\{\\prime\}\_\{a\}\\,\\pi\(da\)\.Integrating the previous display overa∼πa\\sim\\pigives
𝐓gμ¯𝐐=𝐓g′μ¯𝐐′\.\\mathbf\{T\}\_\{g\}\\overline\{\\mu\}\_\{\{\\mathbf\{Q\}\}\}=\\mathbf\{T\}\_\{g^\{\\prime\}\}\\overline\{\\mu\}\_\{\{\\mathbf\{Q\}\}^\{\\prime\}\}\.Sinceμ¯𝐐∼ν\\overline\{\\mu\}\_\{\{\\mathbf\{Q\}\}\}\\sim\{\\nu\}andμ¯𝐐′∼ν\\overline\{\\mu\}\_\{\{\\mathbf\{Q\}\}^\{\\prime\}\}\\sim\{\\nu\}, it follows that
𝐓gν∼𝐓g′ν\.\\mathbf\{T\}\_\{g\}\{\\nu\}\\sim\\mathbf\{T\}\_\{g^\{\\prime\}\}\{\\nu\}\.Henceg\(𝒞\)g\(\\mathcal\{C\}\)andg′\(𝒞\)g^\{\\prime\}\(\\mathcal\{C\}\)agree modulo the corresponding pushed\-forward carrier:
ν\{c:g′\(c\)∉g\(𝒞\)\}=0,ν\{c:g\(c\)∉g′\(𝒞\)\}=0\.\{\\nu\}\\\{c:g^\{\\prime\}\(c\)\\notin g\(\\mathcal\{C\}\)\\\}=0,\\qquad\{\\nu\}\\\{c:g\(c\)\\notin g^\{\\prime\}\(\\mathcal\{C\}\)\\\}=0\.Therefore the maps
τ:=g−1∘g′,τ−1:=g′−1∘g\\tau\\mathrel\{:=\}g^\{\-1\}\\circ g^\{\\prime\},\\qquad\\tau^\{\-1\}\\mathrel\{:=\}g^\{\\prime\-1\}\\circ gare well\-defined moduloν\{\\nu\}\-null sets\. Moreover, since𝐓g𝐓τν=𝐓g′ν∼𝐓gν\\mathbf\{T\}\_\{g\}\\mathbf\{T\}\_\{\\tau\}\{\\nu\}=\\mathbf\{T\}\_\{g^\{\\prime\}\}\{\\nu\}\\sim\\mathbf\{T\}\_\{g\}\{\\nu\}, injectivity ofggimplies𝐓τν∼ν\\mathbf\{T\}\_\{\\tau\}\{\\nu\}\\sim\{\\nu\}\. The same argument withggandg′g^\{\\prime\}exchanged gives𝐓τ−1ν∼ν\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\{\\nu\}\\sim\{\\nu\}\. Thus
τ∈Autν\(𝒞\),g′=g∘τν\-a\.e\.\\tau\\in\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\),\\qquad g^\{\\prime\}=g\\circ\\tau\\quad\{\\nu\}\\text\{\-a\.e\.\}
Sinceμa′≪ν\\mu^\{\\prime\}\_\{a\}\\ll\{\\nu\}, the relationg′=g∘τg^\{\\prime\}=g\\circ\\tauν\{\\nu\}\-a\.e\. implies
𝐓g′μa′=𝐓g𝐓τμa′\.\\mathbf\{T\}\_\{g^\{\\prime\}\}\\mu^\{\\prime\}\_\{a\}=\\mathbf\{T\}\_\{g\}\\mathbf\{T\}\_\{\\tau\}\\mu^\{\\prime\}\_\{a\}\.Combining this with𝐓gμa=𝐓g′μa′\\mathbf\{T\}\_\{g\}\\mu\_\{a\}=\\mathbf\{T\}\_\{g^\{\\prime\}\}\\mu^\{\\prime\}\_\{a\}, we obtain
𝐓gμa=𝐓g𝐓τμa′\.\\mathbf\{T\}\_\{g\}\\mu\_\{a\}=\\mathbf\{T\}\_\{g\}\\mathbf\{T\}\_\{\\tau\}\\mu^\{\\prime\}\_\{a\}\.Sinceggis a bimeasurable embedding,𝐓g\\mathbf\{T\}\_\{g\}is injective on measures on𝒞\\mathcal\{C\}\. Hence
μa=𝐓τμa′∀a∈𝒜o\.\\mu\_\{a\}=\\mathbf\{T\}\_\{\\tau\}\\mu^\{\\prime\}\_\{a\}\\qquad\\forall a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\.Equivalently,
𝐁𝐐\(⋅∣a\)=𝐓τ𝐁𝐐′\(⋅∣a\)∀a∈𝒜o,\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau\}\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)\\qquad\\forall a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\},soτ∈𝒯𝐐¯ν\(𝒬¯;𝒜o\)\\tau\\in\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\.
Finally, for everyμ≪ν\\mu\\ll\{\\nu\},
𝐊′μ=𝐊~𝐓g′μ=𝐊~𝐓g𝐓τμ=𝐊𝐓τμ\.\{\\mathbf\{K\}\}^\{\\prime\}\\mu=\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g^\{\\prime\}\}\\mu=\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g\}\\mathbf\{T\}\_\{\\tau\}\\mu=\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}\\mu\.Thus𝐊′=𝜈𝐊𝐓τ\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}, soτ∈𝒯𝐊ν\(𝒦\)\\tau\\in\\mathcal\{T\}\_\{\{\\mathbf\{K\}\}\}^\{\{\\nu\}\}\(\\mathcal\{K\}\)\. Therefore
τ∈𝒯𝐊ν\(𝒦\)∩𝒯𝐐¯ν\(𝒬¯;𝒜o\)⊆𝔊\.\\tau\\in\\mathcal\{T\}\_\{\{\\mathbf\{K\}\}\}^\{\{\\nu\}\}\(\\mathcal\{K\}\)\\cap\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\\subseteq\\mathfrak\{G\}\.By[Definition5](https://arxiv.org/html/2606.18509#Thmdefinition5),𝖬∼𝔊𝖬′\\mathsf\{M\}\\sim\_\{\\mathfrak\{G\}\}\\mathsf\{M\}^\{\\prime\}\. Since𝖬′\\mathsf\{M\}^\{\\prime\}was arbitrary among feature\-equivalent alternatives,𝖬\\mathsf\{M\}is identifiable up to∼𝔊\\sim\_\{\\mathfrak\{G\}\}\. ∎
### C\.3Proof of[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)
See[2](https://arxiv.org/html/2606.18509#Thmtheorem2)
###### Proof\.
We first show that \(i\) and \(ii\) are equivalent\. For anya∈𝒜′a\\in\\mathcal\{A\}^\{\\prime\}, the identity𝐐¯\(⋅∣a\)=𝐓τ𝐐¯′\(⋅∣a\)\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)is equivalent, sinceτ∈Autν\(𝒞\)\\tau\\in\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\), to
𝐐¯′\(⋅∣a\)=𝐓τ−1𝐐¯\(⋅∣a\)\.\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)\.If𝐐¯\(⋅∣a\)\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)has densityp𝐐¯\(⋅∣a\)p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\cdot\\mid a\)with respect toν\{\\nu\}, then the Radon–Nikodym chain rule for𝐓τ−1\\mathbf\{T\}\_\{\\tau^\{\-1\}\}gives
d𝐓τ−1𝐐¯\(⋅∣a\)dν\(c\)=p𝐐¯\(τ\(c\)∣a\)rτ−1\(c\),rτ−1=d\(𝐓τ−1ν\)dν\.\\frac\{d\\,\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)\}\{d\{\\nu\}\}\(c\)=p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\)r\_\{\\tau^\{\-1\}\}\(c\),\\qquad r\_\{\\tau^\{\-1\}\}=\\frac\{d\(\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\{\\nu\}\)\}\{d\{\\nu\}\}\.Thus𝐐¯′\(⋅∣a\)=𝐓τ−1𝐐¯\(⋅∣a\)\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)=\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\)holds for everya∈𝒜′a\\in\\mathcal\{A\}^\{\\prime\}if and only if
p𝐐¯′\(c∣a\)=p𝐐¯\(τ\(c\)∣a\)rτ−1\(c\)p\_\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\\mid a\)=p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\)r\_\{\\tau^\{\-1\}\}\(c\)holds for everya∈𝒜′a\\in\\mathcal\{A\}^\{\\prime\}andν\{\\nu\}\-a\.e\.cc, which proves the equivalence between \(i\) and \(ii\)\.
We next show that \(ii\) implies \(iii\)\. Takinga=a0a=a\_\{0\}in \(ii\) gives the anchored density identity
p𝐐¯′\(c∣a0\)=p𝐐¯\(τ\(c\)∣a0\)rτ−1\(c\)p\_\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\\mid a\_\{0\}\)=p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\_\{0\}\)r\_\{\\tau^\{\-1\}\}\(c\)forν\{\\nu\}\-a\.e\.cc\. Now fix anya∈𝒜′a\\in\\mathcal\{A\}^\{\\prime\}\. By the common\-support assumption,p𝐐¯′\(⋅∣a\)p\_\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(\\cdot\\mid a\),p𝐐¯′\(⋅∣a0\)p\_\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(\\cdot\\mid a\_\{0\}\),p𝐐¯\(⋅∣a\)p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\cdot\\mid a\), andp𝐐¯\(⋅∣a0\)p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\cdot\\mid a\_\{0\}\)are positive and finiteν\{\\nu\}\-a\.e\.; sinceτ∈Autν\(𝒞\)\\tau\\in\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\), the same is true ofp𝐐¯\(τ\(⋅\)∣a\)p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(\\cdot\)\\mid a\)andp𝐐¯\(τ\(⋅\)∣a0\)p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(\\cdot\)\\mid a\_\{0\}\)on aν\{\\nu\}\-full set\. Intersecting this full\-measure set with the full\-measure sets on which \(ii\) holds foraaanda0a\_\{0\}, we may take logarithms and subtract:
logp𝐐¯′\(c∣a\)−logp𝐐¯′\(c∣a0\)\\displaystyle\\log p\_\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\\mid a\)\-\\log p\_\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\\mid a\_\{0\}\)=\[logp𝐐¯\(τ\(c\)∣a\)\+logrτ−1\(c\)\]\\displaystyle=\\left\[\\log p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\)\+\\log r\_\{\\tau^\{\-1\}\}\(c\)\\right\]−\[logp𝐐¯\(τ\(c\)∣a0\)\+logrτ−1\(c\)\]\\displaystyle\\quad\-\\left\[\\log p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\_\{0\}\)\+\\log r\_\{\\tau^\{\-1\}\}\(c\)\\right\]=logp𝐐¯\(τ\(c\)∣a\)−logp𝐐¯\(τ\(c\)∣a0\)\.\\displaystyle=\\log p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\)\-\\log p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\_\{0\}\)\.Therefore
Δa,a0𝐐¯′\(c\)=Δa,a0𝐐¯\(τ\(c\)\)\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)forν\{\\nu\}\-a\.e\.cc\. Sincea∈𝒜′a\\in\\mathcal\{A\}^\{\\prime\}was arbitrary, the attribute\-potential identities hold for everya∈𝒜′a\\in\\mathcal\{A\}^\{\\prime\}, so \(iii\) follows\.
Conversely, suppose \(iii\) holds\. Then for everya∈𝒜′a\\in\\mathcal\{A\}^\{\\prime\},
p𝐐¯′\(c∣a\)\\displaystyle p\_\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\\mid a\)=p𝐐¯′\(c∣a0\)exp\{Δa,a0𝐐¯′\(c\)\}\\displaystyle=p\_\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\\mid a\_\{0\}\)\\exp\\\{\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)\\\}=p𝐐¯\(τ\(c\)∣a0\)rτ−1\(c\)exp\{Δa,a0𝐐¯\(τ\(c\)\)\}\\displaystyle=p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\_\{0\}\)r\_\{\\tau^\{\-1\}\}\(c\)\\exp\\\{\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)\\\}=p𝐐¯\(τ\(c\)∣a\)rτ−1\(c\)\\displaystyle=p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\)r\_\{\\tau^\{\-1\}\}\(c\)forν\{\\nu\}\-a\.e\.cc, so \(ii\) holds\. The final statement follows from the definition of𝒯𝐐¯ν\(𝒬¯;𝒜o\)\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\\nu\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\. ∎
### C\.4Proof of[Thm\.3](https://arxiv.org/html/2606.18509#Thmtheorem3)
See[3](https://arxiv.org/html/2606.18509#Thmtheorem3)
###### Proof\.
Write
𝐐¯=𝐁𝐐,𝐐¯′=𝐁𝐐′,μa:=𝐐¯\(⋅∣a\),μa′:=𝐐¯′\(⋅∣a\)\.\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\},\\qquad\\bar\{\\mathbf\{Q\}\}^\{\\prime\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\}^\{\\prime\},\\qquad\\mu\_\{a\}\\mathrel\{:=\}\\bar\{\\mathbf\{Q\}\}\(\\cdot\\mid a\),\\qquad\\mu^\{\\prime\}\_\{a\}\\mathrel\{:=\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\(\\cdot\\mid a\)\.By the commonν\{\\nu\}\-support assumption,μa,μa′≪ν\\mu\_\{a\},\\mu^\{\\prime\}\_\{a\}\\ll\{\\nu\}for everya∈𝒜a\\in\\mathcal\{A\}\. Sinceτ∈Autν\(𝒞\)\\tau\\in\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\), we also have𝐓τμa′≪ν\\mathbf\{T\}\_\{\\tau\}\\mu^\{\\prime\}\_\{a\}\\ll\{\\nu\}for everya∈𝒜a\\in\\mathcal\{A\}\.
The transitionτ\\tauinduced by[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)satisfies
𝐊′=𝜈𝐊𝐓τ,μa=𝐓τμa′∀a∈𝒜o\.\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\},\\qquad\\mu\_\{a\}=\\mathbf\{T\}\_\{\\tau\}\\mu^\{\\prime\}\_\{a\}\\quad\\forall a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\.In particular,
μa0=𝐓τμa0′\.\\mu\_\{a\_\{0\}\}=\\mathbf\{T\}\_\{\\tau\}\\mu^\{\\prime\}\_\{a\_\{0\}\}\.\(3\)
We first show that, for eacha∈𝒜exa\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\},
𝐏𝖬\(⋅∣a\)=𝐏𝖬′\(⋅∣a\)⟺μa=𝐓τμa′\.\{\\mathbf\{P\}\}^\{\\mathsf\{M\}\}\(\\cdot\\mid a\)=\{\\mathbf\{P\}\}^\{\\mathsf\{M\}^\{\\prime\}\}\(\\cdot\\mid a\)\\quad\\Longleftrightarrow\\quad\\mu\_\{a\}=\\mathbf\{T\}\_\{\\tau\}\\mu^\{\\prime\}\_\{a\}\.\(4\)Indeed, sinceμa′≪ν\\mu^\{\\prime\}\_\{a\}\\ll\{\\nu\}and𝐊′=𝜈𝐊𝐓τ\{\\mathbf\{K\}\}^\{\\prime\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\},
𝐏𝖬′\(⋅∣a\)=𝐊′μa′=𝐊𝐓τμa′\.\{\\mathbf\{P\}\}^\{\\mathsf\{M\}^\{\\prime\}\}\(\\cdot\\mid a\)=\{\\mathbf\{K\}\}^\{\\prime\}\\mu^\{\\prime\}\_\{a\}=\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}\\mu^\{\\prime\}\_\{a\}\.Also𝐏𝖬\(⋅∣a\)=𝐊μa\{\\mathbf\{P\}\}^\{\\mathsf\{M\}\}\(\\cdot\\mid a\)=\{\\mathbf\{K\}\}\\mu\_\{a\}\. Thus feature equality ataais equivalent to
𝐊μa=𝐊𝐓τμa′\.\{\\mathbf\{K\}\}\\mu\_\{a\}=\{\\mathbf\{K\}\}\\mathbf\{T\}\_\{\\tau\}\\mu^\{\\prime\}\_\{a\}\.It remains only to justify that𝐊\{\\mathbf\{K\}\}is injective on measures dominated byν\{\\nu\}\. Byν\{\\nu\}\-Blackwell reducibility, chooseg∈𝒢g\\in\\mathcal\{G\}and𝐊~\\widetilde\{\{\\mathbf\{K\}\}\}such that
𝐊=𝜈𝐊~𝐓g\.\{\\mathbf\{K\}\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g\}\.Ifα,β≪ν\\alpha,\\beta\\ll\{\\nu\}and𝐊α=𝐊β\{\\mathbf\{K\}\}\\alpha=\{\\mathbf\{K\}\}\\beta, then
𝐊~𝐓gα=𝐊~𝐓gβ\.\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g\}\\alpha=\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g\}\\beta\.By the injectivity condition in[Definition3](https://arxiv.org/html/2606.18509#Thmdefinition3),
𝐓gα=𝐓gβ\.\\mathbf\{T\}\_\{g\}\\alpha=\\mathbf\{T\}\_\{g\}\\beta\.Sinceggis a bimeasurable embedding,𝐓g\\mathbf\{T\}\_\{g\}is injective on probability measures on𝒞\\mathcal\{C\}, henceα=β\\alpha=\\beta\. Applying this withα=μa\\alpha=\\mu\_\{a\}andβ=𝐓τμa′\\beta=\\mathbf\{T\}\_\{\\tau\}\\mu^\{\\prime\}\_\{a\}proves \([4](https://arxiv.org/html/2606.18509#A3.E4)\)\.
Taking \([4](https://arxiv.org/html/2606.18509#A3.E4)\) for alla∈𝒜exa\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}gives
𝐏𝒜ex𝖬=𝐏𝒜ex𝖬′⟺𝐐¯𝒜ex=𝐓τ𝐐¯𝒜ex′\.\{\\mathbf\{P\}\}^\{\\mathsf\{M\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}=\{\\mathbf\{P\}\}^\{\\mathsf\{M\}^\{\\prime\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}\\quad\\Longleftrightarrow\\quad\\bar\{\\mathbf\{Q\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}\.\(5\)
We now translate the right\-hand side of \([5](https://arxiv.org/html/2606.18509#A3.E5)\) into attribute potentials\. Apply[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)with𝒜′=𝒜ex\\mathcal\{A\}^\{\\prime\}=\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\. It states that
𝐐¯𝒜ex=𝐓τ𝐐¯𝒜ex′\\bar\{\\mathbf\{Q\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}is equivalent to the conjunction of the anchored density identity
p𝐐¯′\(c∣a0\)=p𝐐¯\(τ\(c\)∣a0\)rτ−1\(c\)ν\-a\.e\.cp\_\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\\mid a\_\{0\}\)=p\_\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\\mid a\_\{0\}\)r\_\{\\tau^\{\-1\}\}\(c\)\\quad\{\\nu\}\\text\{\-a\.e\. \}c\(6\)and the transported attribute\-potential identities
Δa,a0𝐐¯′\(c\)=Δa,a0𝐐¯\(τ\(c\)\)for everya∈𝒜ex,ν\-a\.e\.c\.\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)\\quad\\text\{for every \}a\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\},\\ \{\\nu\}\\text\{\-a\.e\. \}c\.\(7\)However, \([6](https://arxiv.org/html/2606.18509#A3.E6)\) already follows from \([3](https://arxiv.org/html/2606.18509#A3.E3)\) by the density form of[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2), sincea0∈𝒜oa\_\{0\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\. Therefore, for the fixed transitionτ\\tauinduced from the observed attributes, the condition
𝐐¯𝒜ex=𝐓τ𝐐¯𝒜ex′\\bar\{\\mathbf\{Q\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}=\\mathbf\{T\}\_\{\\tau\}\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}is equivalent to \([7](https://arxiv.org/html/2606.18509#A3.E7)\) alone\.
Combining this equivalence with \([5](https://arxiv.org/html/2606.18509#A3.E5)\), we obtain
𝐏𝒜ex𝖬=𝐏𝒜ex𝖬′⟺Δa,a0𝐐¯′\(c\)=Δa,a0𝐐¯\(τ\(c\)\)for everya∈𝒜ex,ν\-a\.e\.c\.\{\\mathbf\{P\}\}^\{\\mathsf\{M\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}=\{\\mathbf\{P\}\}^\{\\mathsf\{M\}^\{\\prime\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}\\quad\\Longleftrightarrow\\quad\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)\\text\{ for every \}a\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\},\\ \{\\nu\}\\text\{\-a\.e\. \}c\.This proves the claimed equivalence\.
The final statement follows immediately: if the transported attribute\-potential identities that hold on𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}extend to alla∈𝒜exa\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}, then the right\-hand side above holds, and therefore𝐏𝒜ex𝖬=𝐏𝒜ex𝖬′\{\\mathbf\{P\}\}^\{\\mathsf\{M\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}=\{\\mathbf\{P\}\}^\{\\mathsf\{M\}^\{\\prime\}\}\_\{\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\}\. ∎
### C\.5Proof of[Thm\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)
See[4](https://arxiv.org/html/2606.18509#Thmtheorem4)
###### Proof\.
Fixaex∈𝒜ex\{a\_\{\\textnormal\{ex\}\}\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\. Sinceφ\(𝒜ex\)\\varphi\(\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\)is contained in the affine hull ofφ\(𝒜o\)\\varphi\(\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\), there exista1,…,aℓ∈𝒜oa\_\{1\},\\ldots,a\_\{\\ell\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}andα1,…,αℓ∈ℝ\\alpha\_\{1\},\\ldots,\\alpha\_\{\\ell\}\\in\\mathbb\{R\}such thatφ\(aex\)=∑i=1ℓαiφ\(ai\)\\varphi\(\{a\_\{\\textnormal\{ex\}\}\}\)=\\sum\_\{i=1\}^\{\\ell\}\\alpha\_\{i\}\\varphi\(a\_\{i\}\)and∑i=1ℓαi=1\\sum\_\{i=1\}^\{\\ell\}\\alpha\_\{i\}=1\. By the assumed affine representation of attribute\-potential differences, for𝖬=\(𝐐,𝐁,𝐊\)\\mathsf\{M\}=\(\{\\mathbf\{Q\}\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}\)we have
Δaex,a0𝐐¯\(c\)−Δaex,a0𝐐¯\(c0\)\\displaystyle\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\)\-\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\_\{0\}\)=⟨φ\(aex\)−φ\(a0\),D𝖬\(c,c0\)⟩\\displaystyle=\\langle\\varphi\(\{a\_\{\\textnormal\{ex\}\}\}\)\-\\varphi\(a\_\{0\}\),D^\{\\mathsf\{M\}\}\(c,c\_\{0\}\)\\rangle=⟨∑i=1ℓαiφ\(ai\)−∑i=1ℓαiφ\(a0\),D𝖬\(c,c0\)⟩\\displaystyle=\\langle\\sum\_\{i=1\}^\{\\ell\}\\alpha\_\{i\}\\varphi\(a\_\{i\}\)\-\\sum\_\{i=1\}^\{\\ell\}\\alpha\_\{i\}\\varphi\(a\_\{0\}\),D^\{\\mathsf\{M\}\}\(c,c\_\{0\}\)\\rangle=∑i=1ℓαi⟨φ\(ai\)−φ\(a0\),D𝖬\(c,c0\)⟩\\displaystyle=\\sum\_\{i=1\}^\{\\ell\}\\alpha\_\{i\}\\langle\\varphi\(a\_\{i\}\)\-\\varphi\(a\_\{0\}\),D^\{\\mathsf\{M\}\}\(c,c\_\{0\}\)\\rangle=∑i=1ℓαi\[Δai,a0𝐐¯\(c\)−Δai,a0𝐐¯\(c0\)\],\\displaystyle=\\sum\_\{i=1\}^\{\\ell\}\\alpha\_\{i\}\\left\[\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\)\-\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(c\_\{0\}\)\\right\],where the second equality uses∑i=1ℓαi=1\\sum\_\{i=1\}^\{\\ell\}\\alpha\_\{i\}=1\. The same argument for𝖬′=\(𝐐′,𝐁,𝐊′\)\\mathsf\{M\}^\{\\prime\}=\(\{\\mathbf\{Q\}\}^\{\\prime\},\{\\mathbf\{B\}\},\{\\mathbf\{K\}\}^\{\\prime\}\)gives
Δaex,a0𝐐¯′\(c\)−Δaex,a0𝐐¯′\(c0\)=∑i=1ℓαi\[Δai,a0𝐐¯′\(c\)−Δai,a0𝐐¯′\(c0\)\]\.\\displaystyle\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)\-\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\_\{0\}\)=\\sum\_\{i=1\}^\{\\ell\}\\alpha\_\{i\}\\left\[\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)\-\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\_\{0\}\)\\right\]\.Sinceai∈𝒜oa\_\{i\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\},[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)gives the observed transported identities
Δai,a0𝐐¯′\(c\)=Δai,a0𝐐¯\(τ\(c\)\)forν\-a\.e\.c\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)\\qquad\\text\{for $\{\\nu\}$\-a\.e\. \}cfor eachi∈\[ℓ\]i\\in\[\\ell\]\. Subtracting the identity evaluated atc0c\_\{0\}from the identity evaluated atcc, we obtain
Δai,a0𝐐¯′\(c\)−Δai,a0𝐐¯′\(c0\)=Δai,a0𝐐¯\(τ\(c\)\)−Δai,a0𝐐¯\(τ\(c0\)\)\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)\-\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\_\{0\}\)=\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)\-\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\_\{0\}\)\)forν⊗ν\{\\nu\}\\otimes\{\\nu\}\-a\.e\.\(c,c0\)\(c,c\_\{0\}\), given that𝐓τν∼ν\\mathbf\{T\}\_\{\\tau\}\{\\nu\}\\sim\{\\nu\}\. Combining the preceding two displays gives
Δaex,a0𝐐¯′\(c\)−Δaex,a0𝐐¯′\(c0\)\\displaystyle\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)\-\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\_\{0\}\)=∑i=1ℓαi\[Δai,a0𝐐¯\(τ\(c\)\)−Δai,a0𝐐¯\(τ\(c0\)\)\]\.\\displaystyle=\\sum\_\{i=1\}^\{\\ell\}\\alpha\_\{i\}\\left\[\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)\-\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\_\{0\}\)\)\\right\]\.On the other hand, applying the affine representation for𝖬\\mathsf\{M\}at the pair\(τ\(c\),τ\(c0\)\)\(\\tau\(c\),\\tau\(c\_\{0\}\)\)gives
Δaex,a0𝐐¯\(τ\(c\)\)−Δaex,a0𝐐¯\(τ\(c0\)\)\\displaystyle\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)\-\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\_\{0\}\)\)=∑i=1ℓαi\[Δai,a0𝐐¯\(τ\(c\)\)−Δai,a0𝐐¯\(τ\(c0\)\)\]\\displaystyle=\\sum\_\{i=1\}^\{\\ell\}\\alpha\_\{i\}\\left\[\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)\-\\Delta\_\{a\_\{i\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\_\{0\}\)\)\\right\]forν⊗ν\{\\nu\}\\otimes\{\\nu\}\-a\.e\.\(c,c0\)\(c,c\_\{0\}\), sinceτ∈Autν\(𝒞\)\\tau\\in\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathcal\{C\}\)preservesν\{\\nu\}\-null sets\. Therefore
Δaex,a0𝐐¯′\(c\)−Δaex,a0𝐐¯′\(c0\)=Δaex,a0𝐐¯\(τ\(c\)\)−Δaex,a0𝐐¯\(τ\(c0\)\)\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)\-\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\_\{0\}\)=\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)\-\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\_\{0\}\)\)forν⊗ν\{\\nu\}\\otimes\{\\nu\}\-a\.e\.\(c,c0\)\(c,c\_\{0\}\)\. Define
H\(c\):=Δaex,a0𝐐¯′\(c\)−Δaex,a0𝐐¯\(τ\(c\)\)\.H\(c\)\\mathrel\{:=\}\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)\-\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)\.The preceding display saysH\(c\)−H\(c0\)=0H\(c\)\-H\(c\_\{0\}\)=0forν⊗ν\{\\nu\}\\otimes\{\\nu\}\-a\.e\.\(c,c0\)\(c,c\_\{0\}\)\. This implies that, through Fubini,HHisν\{\\nu\}\-a\.e\. constant, so there exists a scalarb\(aex\)b\(\{a\_\{\\textnormal\{ex\}\}\}\)such that
Δaex,a0𝐐¯′\(c\)=Δaex,a0𝐐¯\(τ\(c\)\)\+b\(aex\)\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)\+b\(\{a\_\{\\textnormal\{ex\}\}\}\)forν\{\\nu\}\-a\.e\.cc\. Letrτ−1\(c\):=d\(𝐓τ−1ν\)/dν\(c\)r\_\{\\tau^\{\-1\}\}\(c\)\\mathrel\{:=\}d\(\\mathbf\{T\}\_\{\\tau^\{\-1\}\}\{\\nu\}\)/d\{\\nu\}\(c\)\. By the anchored density identity in[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2),
p𝐐¯′\(c∣a0\)=p𝐐¯\(τ\(c\)∣a0\)rτ−1\(c\)p\_\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\\mid a\_\{0\}\)=p\_\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\\mid a\_\{0\}\)r\_\{\\tau^\{\-1\}\}\(c\)forν\{\\nu\}\-a\.e\.cc\. Therefore,
p𝐐¯′\(c∣aex\)\\displaystyle p\_\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\\mid\{a\_\{\\textnormal\{ex\}\}\}\)=p𝐐¯′\(c∣a0\)exp\{Δaex,a0𝐐¯′\(c\)\}\\displaystyle=p\_\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\\mid a\_\{0\}\)\\exp\\\{\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)\\\}=p𝐐¯\(τ\(c\)∣a0\)rτ−1\(c\)exp\{Δaex,a0𝐐¯\(τ\(c\)\)\+b\(aex\)\}\\displaystyle=p\_\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\\mid a\_\{0\}\)r\_\{\\tau^\{\-1\}\}\(c\)\\exp\\\{\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)\+b\(\{a\_\{\\textnormal\{ex\}\}\}\)\\\}=exp\{b\(aex\)\}p𝐐¯\(τ\(c\)∣aex\)rτ−1\(c\)\.\\displaystyle=\\exp\\\{b\(\{a\_\{\\textnormal\{ex\}\}\}\)\\\}p\_\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\\mid\{a\_\{\\textnormal\{ex\}\}\}\)r\_\{\\tau^\{\-1\}\}\(c\)\.Integrating both sides with respect toν\{\\nu\}gives
1\\displaystyle 1=exp\{b\(aex\)\}∫p𝐐¯\(τ\(c\)∣aex\)rτ−1\(c\)ν\(dc\)\\displaystyle=\\exp\\\{b\(\{a\_\{\\textnormal\{ex\}\}\}\)\\\}\\int p\_\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\\mid\{a\_\{\\textnormal\{ex\}\}\}\)r\_\{\\tau^\{\-1\}\}\(c\)\\,\{\\nu\}\(dc\)=exp\{b\(aex\)\}∫p𝐐¯\(u∣aex\)ν\(du\)\\displaystyle=\\exp\\\{b\(\{a\_\{\\textnormal\{ex\}\}\}\)\\\}\\int p\_\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(u\\mid\{a\_\{\\textnormal\{ex\}\}\}\)\\,\{\\nu\}\(du\)=exp\{b\(aex\)\}\.\\displaystyle=\\exp\\\{b\(\{a\_\{\\textnormal\{ex\}\}\}\)\\\}\.Henceb\(aex\)=0b\(\{a\_\{\\textnormal\{ex\}\}\}\)=0\. ThusΔaex,a0𝐐¯′\(c\)=Δaex,a0𝐐¯\(τ\(c\)\)\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{\{a\_\{\\textnormal\{ex\}\}\},a\_\{0\}\}^\{\\bar\{\{\\mathbf\{Q\}\}\}\}\(\\tau\(c\)\)forν\{\\nu\}\-a\.e\.cc\. Sinceaex∈𝒜ex\{a\_\{\\textnormal\{ex\}\}\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}was arbitrary, the transported attribute\-potential identities extend from𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}to𝒜ex\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\. ∎
### C\.6Proof of[Cor\.1](https://arxiv.org/html/2606.18509#Thmcorollary1)
See[1](https://arxiv.org/html/2606.18509#Thmcorollary1)
###### Proof\.
Set𝒜ex=2\[m\]\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}=2^\{\[m\]\}and takea0=∅a\_\{0\}=\\varnothing\. For eachS⊆\[m\]S\\subseteq\[m\], define the interaction\-incidence vectorφℋ\(S\)∈ℝℋ\\varphi\_\{\\mathcal\{H\}\}\(S\)\\in\\mathbb\{R\}^\{\\mathcal\{H\}\}byφℋ\(S\)T=𝟏\{T⊆S\}\\varphi\_\{\\mathcal\{H\}\}\(S\)\_\{T\}=\\mathbf\{1\}\\\{T\\subseteq S\\\}forT∈ℋT\\in\\mathcal\{H\}\. We first show that the assumed density model has the affine attribute\-potential form required by[Thm\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)\. Write
ℓS𝖬\(c\)=∑T∈ℋ,T⊆ShT𝖬\(c\),\\ell\_\{S\}^\{\\mathsf\{M\}\}\(c\)=\\sum\_\{T\\in\\mathcal\{H\},T\\subseteq S\}h\_\{T\}^\{\\mathsf\{M\}\}\(c\),so thatp𝖬\(c∣S\)=exp\{ℓS𝖬\(c\)\}/ZS𝖬p^\{\\mathsf\{M\}\}\(c\\mid S\)=\\exp\\\{\\ell\_\{S\}^\{\\mathsf\{M\}\}\(c\)\\\}/Z\_\{S\}^\{\\mathsf\{M\}\}for a normalizing constantZS𝖬Z\_\{S\}^\{\\mathsf\{M\}\}\. Sincep𝖬\(c∣∅\)=exp\{h∅𝖬\(c\)\}/Z∅𝖬p^\{\\mathsf\{M\}\}\(c\\mid\\varnothing\)=\\exp\\\{h\_\{\\varnothing\}^\{\\mathsf\{M\}\}\(c\)\\\}/Z\_\{\\varnothing\}^\{\\mathsf\{M\}\}, the attribute potential relative to∅\\varnothingis
ΔS,∅𝖬\(c\)=logp𝖬\(c∣S\)p𝖬\(c∣∅\)=∑T∈ℋ,T⊆S,T≠∅hT𝖬\(c\)−logZS𝖬Z∅𝖬\.\\Delta\_\{S,\\varnothing\}^\{\\mathsf\{M\}\}\(c\)=\\log\\frac\{p^\{\\mathsf\{M\}\}\(c\\mid S\)\}\{p^\{\\mathsf\{M\}\}\(c\\mid\\varnothing\)\}=\\sum\_\{T\\in\\mathcal\{H\},T\\subseteq S,T\\neq\\varnothing\}h\_\{T\}^\{\\mathsf\{M\}\}\(c\)\-\\log\\frac\{Z\_\{S\}^\{\\mathsf\{M\}\}\}\{Z\_\{\\varnothing\}^\{\\mathsf\{M\}\}\}\.The normalizing constant does not depend oncc, so it disappears after centering at any fixedc0∈𝒞c\_\{0\}\\in\\mathcal\{C\}\. Thus
ΔS,∅𝖬\(c\)−ΔS,∅𝖬\(c0\)=∑T∈ℋ,T⊆S,T≠∅\(hT𝖬\(c\)−hT𝖬\(c0\)\)\.\\Delta\_\{S,\\varnothing\}^\{\\mathsf\{M\}\}\(c\)\-\\Delta\_\{S,\\varnothing\}^\{\\mathsf\{M\}\}\(c\_\{0\}\)=\\sum\_\{T\\in\\mathcal\{H\},T\\subseteq S,T\\neq\\varnothing\}\\bigl\(h\_\{T\}^\{\\mathsf\{M\}\}\(c\)\-h\_\{T\}^\{\\mathsf\{M\}\}\(c\_\{0\}\)\\bigr\)\.DefineD𝖬\(c,c0\)∈ℝℋD^\{\\mathsf\{M\}\}\(c,c\_\{0\}\)\\in\\mathbb\{R\}^\{\\mathcal\{H\}\}byD∅𝖬\(c,c0\)=0D\_\{\\varnothing\}^\{\\mathsf\{M\}\}\(c,c\_\{0\}\)=0andDT𝖬\(c,c0\)=hT𝖬\(c\)−hT𝖬\(c0\)D\_\{T\}^\{\\mathsf\{M\}\}\(c,c\_\{0\}\)=h\_\{T\}^\{\\mathsf\{M\}\}\(c\)\-h\_\{T\}^\{\\mathsf\{M\}\}\(c\_\{0\}\)forT≠∅T\\neq\\varnothing\. Sinceφℋ\(∅\)\\varphi\_\{\\mathcal\{H\}\}\(\\varnothing\)is equal to11in the∅\\varnothing\-coordinate and0elsewhere, the previous display can be written as
ΔS,∅𝖬\(c\)−ΔS,∅𝖬\(c0\)=⟨φℋ\(S\)−φℋ\(∅\),D𝖬\(c,c0\)⟩\.\\Delta\_\{S,\\varnothing\}^\{\\mathsf\{M\}\}\(c\)\-\\Delta\_\{S,\\varnothing\}^\{\\mathsf\{M\}\}\(c\_\{0\}\)=\\langle\\varphi\_\{\\mathcal\{H\}\}\(S\)\-\\varphi\_\{\\mathcal\{H\}\}\(\\varnothing\),D^\{\\mathsf\{M\}\}\(c,c\_\{0\}\)\\rangle\.Hence the fixed representationφ\(S\)=φℋ\(S\)\\varphi\(S\)=\\varphi\_\{\\mathcal\{H\}\}\(S\)satisfies the attribute\-potential assumption of[Thm\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)\.
It remains only to verify that each target vectorφℋ\(S\)\\varphi\_\{\\mathcal\{H\}\}\(S\), forS⊆\[m\]S\\subseteq\[m\], belongs to the affine hull of the observed vectors\{φℋ\(U\)\|U∈ℋ\}\\\{\\varphi\_\{\\mathcal\{H\}\}\(U\)\\;\|\\;U\\in\\mathcal\{H\}\\\}\. This is a standard consequence of Möbius inversion, equivalently of the invertibility of the zeta matrix of a finite poset, in the incidence\-algebra formulation ofRota \[[1964](https://arxiv.org/html/2606.18509#bib.bib47)\]\. For completeness, we recall the short argument\.
We first show linear that the observed vectors\{φℋ\(U\)\|U∈ℋ\}\\\{\\varphi\_\{\\mathcal\{H\}\}\(U\)\\;\|\\;U\\in\\mathcal\{H\}\\\}are linearly independent\. Suppose∑U∈ℋαUφℋ\(U\)=0\\sum\_\{U\\in\\mathcal\{H\}\}\\alpha\_\{U\}\\varphi\_\{\\mathcal\{H\}\}\(U\)=0\. If some coefficient is nonzero, choose an inclusion\-maximalU0∈ℋU\_\{0\}\\in\\mathcal\{H\}among the sets withαU0≠0\\alpha\_\{U\_\{0\}\}\\neq 0, which is possible becauseℋ\\mathcal\{H\}is finite\. Looking at theU0U\_\{0\}\-coordinate of the vector equality gives
0=∑U∈ℋαUφℋ\(U\)U0=∑U∈ℋU0⊆UαU\.0=\\sum\_\{U\\in\\mathcal\{H\}\}\\alpha\_\{U\}\\varphi\_\{\\mathcal\{H\}\}\(U\)\_\{U\_\{0\}\}=\\sum\_\{\\begin\{subarray\}\{c\}U\\in\\mathcal\{H\}\\\\ U\_\{0\}\\subseteq U\\end\{subarray\}\}\\alpha\_\{U\}\.By maximality ofU0U\_\{0\}, every term in the last sum exceptU=U0U=U\_\{0\}has coefficient zero\. Therefore0=αU00=\\alpha\_\{U\_\{0\}\}, contradicting the choice ofU0U\_\{0\}\. Thus all coefficients are zero, so the observed vectors are linearly independent\. There are exactly\|ℋ\|\|\\mathcal\{H\}\|such vectors in the\|ℋ\|\|\\mathcal\{H\}\|\-dimensional spaceℝℋ\\mathbb\{R\}^\{\\mathcal\{H\}\}, so they form a basis ofℝℋ\\mathbb\{R\}^\{\\mathcal\{H\}\}\.
Now fix any targetS⊆\[m\]S\\subseteq\[m\]\. Since the observed vectors form a basis, there exist coefficientsαU\(S\)\\alpha\_\{U\}\(S\),U∈ℋU\\in\\mathcal\{H\}, such that
φℋ\(S\)=∑U∈ℋαU\(S\)φℋ\(U\)\.\\varphi\_\{\\mathcal\{H\}\}\(S\)=\\sum\_\{U\\in\\mathcal\{H\}\}\\alpha\_\{U\}\(S\)\\varphi\_\{\\mathcal\{H\}\}\(U\)\.This expansion is automatically affine rather than merely linear\. Indeed, looking at the∅\\varnothing\-coordinate gives
1=φℋ\(S\)∅=∑U∈ℋαU\(S\)φℋ\(U\)∅=∑U∈ℋαU\(S\),1=\\varphi\_\{\\mathcal\{H\}\}\(S\)\_\{\\varnothing\}=\\sum\_\{U\\in\\mathcal\{H\}\}\\alpha\_\{U\}\(S\)\\varphi\_\{\\mathcal\{H\}\}\(U\)\_\{\\varnothing\}=\\sum\_\{U\\in\\mathcal\{H\}\}\\alpha\_\{U\}\(S\),because∅⊆S\\varnothing\\subseteq Sand∅⊆U\\varnothing\\subseteq Ufor everyU∈ℋU\\in\\mathcal\{H\}\. Thereforeφℋ\(S\)\\varphi\_\{\\mathcal\{H\}\}\(S\)lies in the affine hull of\{φℋ\(U\)\|U∈ℋ\}\\\{\\varphi\_\{\\mathcal\{H\}\}\(U\)\\;\|\\;U\\in\\mathcal\{H\}\\\}\. SinceS⊆\[m\]S\\subseteq\[m\]was arbitrary,φℋ\(𝒜ex\)\\varphi\_\{\\mathcal\{H\}\}\(\{\\mathcal\{A\}\_\{\\textnormal\{ex\}\}\}\)is contained in the affine hull ofφℋ\(𝒜o\)\\varphi\_\{\\mathcal\{H\}\}\(\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\.
Let𝖬,𝖬′∈ℳ\\mathsf\{M\},\\mathsf\{M\}^\{\\prime\}\\in\\mathcal\{M\}be feature\-equivalent on𝒜o=ℋ\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}=\\mathcal\{H\}\. The preceding two paragraphs verify the hypotheses of[Thm\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)withφ=φℋ\\varphi=\\varphi\_\{\\mathcal\{H\}\}\. Therefore the transported attribute\-potential identities that hold on the observed attributesℋ\\mathcal\{H\}extend to everyS⊆\[m\]S\\subseteq\[m\]\. By[Thm\.3](https://arxiv.org/html/2606.18509#Thmtheorem3), these extended attribute\-potential identities imply𝐏S𝖬=𝐏S𝖬′\{\\mathbf\{P\}\}\_\{S\}^\{\\mathsf\{M\}\}=\{\\mathbf\{P\}\}\_\{S\}^\{\\mathsf\{M\}^\{\\prime\}\}for everyS⊆\[m\]S\\subseteq\[m\]\. Hence𝖬\\mathsf\{M\}and𝖬′\\mathsf\{M\}^\{\\prime\}are feature\-equivalent on all of2\[m\]2^\{\[m\]\}\. ∎
## Appendix DDeferred details
### D\.1Details for the running example
We provide the linear\-algebra details behind[Ex\.2](https://arxiv.org/html/2606.18509#Thmexample2)\. Fixτ∈𝒯𝐐¯ν\(𝒬¯;𝒜o\)\\tau\\in\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\. By[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2), there exists𝐐′=𝐓f′∈𝒬\{\\mathbf\{Q\}\}^\{\\prime\}=\\mathbf\{T\}\_\{f^\{\\prime\}\}\\in\\mathcal\{Q\}such that, after centering at any fixedc0c\_\{0\},
⟨f′\(a\)−f′\(a0\),c−c0⟩=⟨f\(a\)−f\(a0\),τ\(c\)−τ\(c0\)⟩∀a∈𝒜o\.\\langle f^\{\\prime\}\(a\)\-f^\{\\prime\}\(a\_\{0\}\),c\-c\_\{0\}\\rangle=\\langle f\(a\)\-f\(a\_\{0\}\),\\tau\(c\)\-\\tau\(c\_\{0\}\)\\rangle\\qquad\\forall a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\.Let
V:=span\{f\(a\)−f\(a0\)\|a∈𝒜o\}\.V\\mathrel\{:=\}\\operatorname\{span\}\\\{f\(a\)\-f\(a\_\{0\}\)\\;\|\\;a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\\}\.Choosea1,…,ar∈𝒜oa\_\{1\},\\ldots,a\_\{r\}\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}such thatf\(ai\)−f\(a0\)f\(a\_\{i\}\)\-f\(a\_\{0\}\),i=1,…,ri=1,\\ldots,r, form a basis ofVV\. Define
Fo:=\(f\(a1\)−f\(a0\),…,f\(ar\)−f\(a0\)\)⊤,Fo′:=\(f′\(a1\)−f′\(a0\),…,f′\(ar\)−f′\(a0\)\)⊤\.F\_\{o\}\\mathrel\{:=\}\(f\(a\_\{1\}\)\-f\(a\_\{0\}\),\\ldots,f\(a\_\{r\}\)\-f\(a\_\{0\}\)\)^\{\\top\},\\qquad F^\{\\prime\}\_\{o\}\\mathrel\{:=\}\(f^\{\\prime\}\(a\_\{1\}\)\-f^\{\\prime\}\(a\_\{0\}\),\\ldots,f^\{\\prime\}\(a\_\{r\}\)\-f^\{\\prime\}\(a\_\{0\}\)\)^\{\\top\}\.LetV′V^\{\\prime\}be the row space ofFo′F^\{\\prime\}\_\{o\}, and letΠV\\Pi\_\{V\}andΠV′\\Pi\_\{V^\{\\prime\}\}denote the orthogonal projections ontoVVandV′V^\{\\prime\}, respectively\. Stacking the preceding identities gives
Fo′\(c−c0\)=Fo\(τ\(c\)−τ\(c0\)\)\.F^\{\\prime\}\_\{o\}\(c\-c\_\{0\}\)=F\_\{o\}\(\\tau\(c\)\-\\tau\(c\_\{0\}\)\)\.SinceFoF\_\{o\}has full row rank and row spaceVV,
ΠV=Fo⊤\(FoFo⊤\)−1Fo\.\\Pi\_\{V\}=F\_\{o\}^\{\\top\}\(F\_\{o\}F\_\{o\}^\{\\top\}\)^\{\-1\}F\_\{o\}\.Therefore,
ΠVτ\(c\)=MΠV′c\+b,M:=Fo⊤\(FoFo⊤\)−1Fo′,b:=ΠVτ\(c0\)−MΠV′c0\.\\Pi\_\{V\}\\tau\(c\)=M\\Pi\_\{V^\{\\prime\}\}c\+b,\\qquad M\\mathrel\{:=\}F\_\{o\}^\{\\top\}\(F\_\{o\}F\_\{o\}^\{\\top\}\)^\{\-1\}F^\{\\prime\}\_\{o\},\\qquad b\\mathrel\{:=\}\\Pi\_\{V\}\\tau\(c\_\{0\}\)\-M\\Pi\_\{V^\{\\prime\}\}c\_\{0\}\.This proves the projection identity used in the partial\-identifiability remark\.
IfV=ℝkV=\\mathbb\{R\}^\{k\}, thenFoF\_\{o\}has full column rank\. The stacked identity forcesFo′F^\{\\prime\}\_\{o\}to have full column rank as well; otherwiseFo\(τ\(c\)−τ\(c0\)\)F\_\{o\}\(\\tau\(c\)\-\\tau\(c\_\{0\}\)\)would lie in a proper subspace for allcc, contradictingτ∈Autν\(ℝk\)\\tau\\in\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathbb\{R\}^\{k\}\)\. HenceV′=ℝkV^\{\\prime\}=\\mathbb\{R\}^\{k\},ΠV=ΠV′=I\\Pi\_\{V\}=\\Pi\_\{V^\{\\prime\}\}=I, and the projection identity reduces to
τ\(c\)=Mc\+b\.\\tau\(c\)=Mc\+b\.Thus, under the full\-span condition,
𝒯𝐐¯ν\(𝒬¯;𝒜o\)⊆Affine\(k\)∩Autν\(ℝk\)\.\\mathcal\{T\}\_\{\\bar\{\\mathbf\{Q\}\}\}^\{\{\\nu\}\}\(\\bar\{\\mathcal\{Q\}\};\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\)\\subseteq\\mathrm\{Affine\}\(k\)\\cap\{\\mathrm\{Aut\}\}\_\{\{\\nu\}\}\(\\mathbb\{R\}^\{k\}\)\.
## Appendix ERecoveries of representative prior work
The recoveries below are not intended as new proofs of the cited identifiability theorems\. Instead, each recovery is organized as a short CMM dictionary followed by the CMM proof object that the original paper analyzes\. The generic transition step is supplied by[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)and, when a density\-level statement is needed, by[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)\. The remaining rigidity step is delegated to the corresponding theorem or proof in the cited work\.
#### Shared architecture of the recoveries\.
Each recovery follows the same template:
1. \(i\)CMM translation\.Each recovery first translates the prior model into CMM notation by specifying the attributeAA, the modulatorΛ\{\\Lambda\}, the conceptCC, the featureXX, and the kernels𝐐\{\\mathbf\{Q\}\},𝐁\{\\mathbf\{B\}\}, and𝐊\{\\mathbf\{K\}\}\.
2. \(ii\)Attribute potentialsorProof object\.The next paragraph computes the attribute potential or the equivalent object used by the cited proof, such as a centered potential, exponentiated density ratio, score contrast, covariance equation, or quadratic transition identity\.
3. \(iii\)Recovery of guarantees\.The proposition starts from feature\-equivalent CMMs, applies the CMM transition theorem, rewrites the transported identity as the cited paper’s proof object, and invokes the paper\-specific rigidity result that yields the final ambiguity class\.
Accordingly, the short proofs below only verify the CMM\-to\-paper identity displayed in the proposition; they do not reprove the paper\-specific rigidity arguments\.
#### Applicability of[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)\.
In the continuous recoveries below, we takeν\{\\nu\}to be Lebesgue measure on the corresponding concept space𝒞=ℝd\\mathcal\{C\}=\\mathbb\{R\}^\{d\}orℝn\\mathbb\{R\}^\{n\}\. Except where explicitly noted, the mixing kernels are deterministic kernels𝐓g\\mathbf\{T\}\_\{g\}, whereg:𝒞→𝒳g\\colon\\mathcal\{C\}\\to\\mathcal\{X\}is an admissible diffeomorphic embedding; in linear CRL,ggis the full\-column\-rank linear mixing matrixGG, which is a bimeasurable embedding ofℝd\\mathbb\{R\}^\{d\}onto its image\. Thus the mixing classes areν\{\\nu\}\-Blackwell reducible by taking𝒳~=𝒳\\widetilde\{\\mathcal\{X\}\}=\\mathcal\{X\},𝐊~=𝐓id𝒳\\widetilde\{\{\\mathbf\{K\}\}\}=\\mathbf\{T\}\_\{\\operatorname\{id\}\_\{\\mathcal\{X\}\}\}, and𝒢\\mathcal\{G\}to be the relevant class of admissible embeddings, sinceξ↦𝐊~ξ\\xi\\mapsto\\widetilde\{\{\\mathbf\{K\}\}\}\\xiis the identity map on the pushed\-forward measures\. For noisy iVAE\-type observation models, we instead write the mixing kernel as𝐊=𝐊~𝐓g\{\\mathbf\{K\}\}=\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g\}, whereggis the decoder embedding and𝐊~\\widetilde\{\{\\mathbf\{K\}\}\}is the shared observation\-noise channel; the required Blackwell\-reducibility condition is that this shared channel be injective on the relevant pushed\-forward latent laws, or equivalently that we work after the usual deconvolution reduction\.
On the concept side, we assume the common\-carrier condition required by[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1): the induced concept laws in the model class are dominated byν\{\\nu\}, and their densities are positive and finiteν\{\\nu\}\-a\.e\. on the attributes under consideration\. This condition holds directly for the Gaussian CRL, score\-based CRL, and nonparametric CRL recoveries under their full\-support density assumptions, and for nonlinear ICA / iVAE under the usual positivity assumptions on the conditional source or prior densities\. For linear CRL\[Squireset al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib6)\], the original theorem is covariance\-based and does not require densities; when we invoke[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1), we are considering the subclass satisfying the additional common\-carrier condition, for example when the noiseϵ\\epsilonhas a strictly positive finite Lebesgue density and allBkB\_\{k\}are invertible\. Under these standing conditions, any feature\-equivalent alternative admits the transition supplied by[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1); the individual recoveries below then identify the corresponding proof object and delegate the model\-specific rigidity step to the cited paper\.
### E\.1Nonlinear ICA
#### CMM translation\.
For nonlinear ICA\[Hyvarinenet al\.,[2019](https://arxiv.org/html/2606.18509#bib.bib4)\], takeAAto be the auxiliary variable,𝒞=ℝk\\mathcal\{C\}=\\mathbb\{R\}^\{k\},C=ZC=Z, andX=f\(Z\)X=f\(Z\)\. For each attribute valueaa, define the modulator value
λa:=\(q1\(⋅,a\),…,qk\(⋅,a\)\)\.\{\\lambda\}\_\{a\}\\mathrel\{:=\}\\bigl\(q\_\{1\}\(\\cdot,a\),\\ldots,q\_\{k\}\(\\cdot,a\)\\bigr\)\.Thus the modulator contains only the coordinate\-wise conditional source log\-potentials\. The normalizer is the derived functional
Γ\(λ\):=log∫ℝkexp\{∑i=1kλi\(ci\)\}𝑑c,\\Gamma\(\{\\lambda\}\)\\mathrel\{:=\}\\log\\int\_\{\\mathbb\{R\}^\{k\}\}\\exp\\left\\\{\\sum\_\{i=1\}^\{k\}\{\\lambda\}\_\{i\}\(c\_\{i\}\)\\right\\\}dc,and we writeΓ\(a\):=Γ\(λa\)\\Gamma\(a\)\\mathrel\{:=\}\\Gamma\(\{\\lambda\}\_\{a\}\)\. The deterministic indexing kernel is
𝐐=𝐓a↦λa\.\{\\mathbf\{Q\}\}=\\mathbf\{T\}\_\{a\\mapsto\{\\lambda\}\_\{a\}\}\.The shared concept\-modulation kernel maps a tuple of source log\-potentials to the corresponding conditionally independent source density:
𝐁\(dc∣λ\)=exp\{∑i=1kλi\(ci\)−Γ\(λ\)\}dc\.\{\\mathbf\{B\}\}\(dc\\mid\{\\lambda\}\)=\\exp\\left\\\{\\sum\_\{i=1\}^\{k\}\{\\lambda\}\_\{i\}\(c\_\{i\}\)\-\\Gamma\(\{\\lambda\}\)\\right\\\}dc\.Thus, for𝐐¯=𝐁𝐐\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\},
logp𝐐¯\(c∣a\)=∑i=1kqi\(ci,a\)−Γ\(a\)\.\\log p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid a\)=\\sum\_\{i=1\}^\{k\}q\_\{i\}\(c\_\{i\},a\)\-\\Gamma\(a\)\.The mixing kernel is deterministic,𝐊=𝐓f\{\\mathbf\{K\}\}=\\mathbf\{T\}\_\{f\}\.
#### Attribute potentials\.
Relative to an anchora0a\_\{0\}, the attribute potential is
Δa,a0𝐐¯\(c\)=∑i=1k\{qi\(ci,a\)−qi\(ci,a0\)\}−\{Γ\(a\)−Γ\(a0\)\}\.\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)=\\sum\_\{i=1\}^\{k\}\\\{q\_\{i\}\(c\_\{i\},a\)\-q\_\{i\}\(c\_\{i\},a\_\{0\}\)\\\}\-\\\{\\Gamma\(a\)\-\\Gamma\(a\_\{0\}\)\\\}\.The normalizer contrast is independent ofcc, so all mixed second derivatives of the potential are determined only by the coordinate\-wise contrasts
hi\(u;a,a0\):=qi\(u,a\)−qi\(u,a0\)\.h\_\{i\}\(u;a,a\_\{0\}\)\\mathrel\{:=\}q\_\{i\}\(u,a\)\-q\_\{i\}\(u,a\_\{0\}\)\.This separable potential is the CMM proof object whose transported mixed derivatives reproduce the nonlinear ICA equations ofHyvarinenet al\.\[[2019](https://arxiv.org/html/2606.18509#bib.bib4)\]\.
#### Recovery of guarantees\.
###### Proposition 2\(CMM recovery of Theorem 1 of\[Hyvarinenet al\.,[2019](https://arxiv.org/html/2606.18509#bib.bib4)\]\)\.
Let
𝖬=\(𝐓a↦λa,𝐁,𝐓f\),𝖬′=\(𝐓a↦λa′,𝐁,𝐓f′\)\\mathsf\{M\}=\(\\mathbf\{T\}\_\{a\\mapsto\{\\lambda\}\_\{a\}\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{f\}\),\\qquad\\mathsf\{M\}^\{\\prime\}=\(\\mathbf\{T\}\_\{a\\mapsto\{\\lambda\}^\{\\prime\}\_\{a\}\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{f^\{\\prime\}\}\)be feature\-equivalent nonlinear ICA CMMs satisfying the common\-support and smoothness assumptions needed for[Thms\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)and[2](https://arxiv.org/html/2606.18509#Thmtheorem2)\. Letτ=f−1∘f′\\tau=f^\{\-1\}\\circ f^\{\\prime\}be the transition induced by[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)\. Forhi\(u;a,a0\):=qi\(u,a\)−qi\(u,a0\)h\_\{i\}\(u;a,a\_\{0\}\)\\mathrel\{:=\}q\_\{i\}\(u,a\)\-q\_\{i\}\(u,a\_\{0\}\), the transported CMM potential identity implies, for everyj≠j′j\\neq j^\{\\prime\},
0=∑i=1k∂1hi\(τi\(c\);a,a0\)∂cjcj′2τi\(c\)\+∑i=1k∂12hi\(τi\(c\);a,a0\)∂cjτi\(c\)∂cj′τi\(c\)\.0=\\sum\_\{i=1\}^\{k\}\\partial\_\{1\}h\_\{i\}\(\\tau\_\{i\}\(c\);a,a\_\{0\}\)\\partial\_\{c\_\{j\}c\_\{j^\{\\prime\}\}\}^\{2\}\\tau\_\{i\}\(c\)\+\\sum\_\{i=1\}^\{k\}\\partial\_\{1\}^\{2\}h\_\{i\}\(\\tau\_\{i\}\(c\);a,a\_\{0\}\)\\partial\_\{c\_\{j\}\}\\tau\_\{i\}\(c\)\\partial\_\{c\_\{j^\{\\prime\}\}\}\\tau\_\{i\}\(c\)\.This is the mixed\-derivative system used in the proof of Theorem 1 ofHyvarinenet al\.\[[2019](https://arxiv.org/html/2606.18509#bib.bib4)\]\. Consequently, under their Assumption of Variability,τ\\tauis componentwise up to permutation:
τ\(c\)=\(ϕ1\(cπ\(1\)\),…,ϕk\(cπ\(k\)\)\)\.\\tau\(c\)=\\bigl\(\\phi\_\{1\}\(c\_\{\\pi\(1\)\}\),\\ldots,\\phi\_\{k\}\(c\_\{\\pi\(k\)\}\)\\bigr\)\.
###### Proof\.
By[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2), feature equivalence gives
Δa,a0𝐐¯′\(c\)=Δa,a0𝐐¯\(τ\(c\)\)\.\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)\.The candidate potentialΔa,a0𝐐¯′\(c\)\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)is coordinate\-separable inccup to an additive constant, so
∂cjcj′2Δa,a0𝐐¯′\(c\)=0for everyj≠j′\.\\partial\_\{c\_\{j\}c\_\{j^\{\\prime\}\}\}^\{2\}\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=0\\qquad\\text\{for every \}j\\neq j^\{\\prime\}\.Applying the same mixed derivative to the transported right\-hand side gives
0=∂cjcj′2\[∑i=1khi\(τi\(c\);a,a0\)\],0=\\partial\_\{c\_\{j\}c\_\{j^\{\\prime\}\}\}^\{2\}\\left\[\\sum\_\{i=1\}^\{k\}h\_\{i\}\(\\tau\_\{i\}\(c\);a,a\_\{0\}\)\\right\],becauseΓ\(a\)−Γ\(a0\)\\Gamma\(a\)\-\\Gamma\(a\_\{0\}\)is independent ofcc\. Expanding this derivative by the chain rule gives the displayed mixed\-derivative system\. The remaining step is exactly the variability\-rank and integration argument inHyvarinenet al\.\[[2019](https://arxiv.org/html/2606.18509#bib.bib4)\], which forces a monomial Jacobian and hence componentwise recovery up to permutation\. ∎
### E\.2Conditionally exponential families
#### CMM translation\.
For the conditionally exponential\-family model including iVAE\[Khemakhemet al\.,[2020a](https://arxiv.org/html/2606.18509#bib.bib2)\]and ICE\-BeeM\[Khemakhemet al\.,[2020b](https://arxiv.org/html/2606.18509#bib.bib17)\], the concept density has the form
p𝐐¯\(c∣a\)=Q\(c\)Z\(a\)exp\{⟨f\(a\),T\(c\)⟩\},X=g\(C\),p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid a\)=\\frac\{Q\(c\)\}\{Z\(a\)\}\\exp\\\{\\langle f\(a\),T\(c\)\\rangle\\\},\\qquad X=g\(C\),wheref:𝒜→ℝmf\\colon\\mathcal\{A\}\\to\\mathbb\{R\}^\{m\}is the attribute\-dependent natural parameter,T:𝒞→ℝmT\\colon\\mathcal\{C\}\\to\\mathbb\{R\}^\{m\}collects sufficient statistics,QQis a positive base density, andZZis the normalizer\. The modulator isλ\(a\)=\(f\(a\),T,Q\)\{\\lambda\}\(a\)=\(f\(a\),T,Q\), with indexing kernel𝐐=𝐓η\{\\mathbf\{Q\}\}=\\mathbf\{T\}\_\{\\eta\}forη\(a\)=λ\(a\)\\eta\(a\)=\{\\lambda\}\(a\)\. The shared concept\-modulation kernel maps\(θ,T,Q\)\(\\theta,T,Q\)to the density proportional toQ\(c\)exp\{⟨θ,T\(c\)⟩\}Q\(c\)\\exp\\\{\\langle\\theta,T\(c\)\\rangle\\\}\. The mixing kernel is𝐊=𝐓g\{\\mathbf\{K\}\}=\\mathbf\{T\}\_\{g\}, or𝐊=𝐊~𝐓g\{\\mathbf\{K\}\}=\\widetilde\{\{\\mathbf\{K\}\}\}\\mathbf\{T\}\_\{g\}in the noisy\-observation version\.
#### Attribute potentials\.
The attribute potential relative toa0a\_\{0\}is
Δa,a0𝐐¯\(c\)=⟨f\(a\)−f\(a0\),T\(c\)⟩−\{logZ\(a\)−logZ\(a0\)\}\.\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)=\\langle f\(a\)\-f\(a\_\{0\}\),T\(c\)\\rangle\-\\\{\\log Z\(a\)\-\\log Z\(a\_\{0\}\)\\\}\.After centering at anyc0c\_\{0\}, the normalizers cancel:
Δa,a0𝐐¯\(c\)−Δa,a0𝐐¯\(c0\)=⟨f\(a\)−f\(a0\),T\(c\)−T\(c0\)⟩\.\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)\-\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\_\{0\}\)=\\langle f\(a\)\-f\(a\_\{0\}\),T\(c\)\-T\(c\_\{0\}\)\\rangle\.This centered potential is the affine sufficient\-statistic equation used in the iVAE identifiability proof\.
#### Recovery of guarantees\.
###### Proposition 3\(CMM recovery of Theorem 1 of\[Khemakhemet al\.,[2020a](https://arxiv.org/html/2606.18509#bib.bib2)\]\)\.
Let𝖬=\(𝐓η,𝐁,𝐓g\)\\mathsf\{M\}=\(\\mathbf\{T\}\_\{\\eta\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{g\}\)and𝖬′=\(𝐓η′,𝐁,𝐓g′\)\\mathsf\{M\}^\{\\prime\}=\(\\mathbf\{T\}\_\{\\eta^\{\\prime\}\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{g^\{\\prime\}\}\)be feature\-equivalent conditionally exponential\-family CMMs, withη\(a\)=\(f\(a\),T,Q\)\\eta\(a\)=\(f\(a\),T,Q\)andη′\(a\)=\(f′\(a\),T′,Q′\)\\eta^\{\\prime\}\(a\)=\(f^\{\\prime\}\(a\),T^\{\\prime\},Q^\{\\prime\}\)\. Letτ=g−1∘g′\\tau=g^\{\-1\}\\circ g^\{\\prime\}be the transition induced by[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)\. Assume there are attributesa0,a1,…,ama\_\{0\},a\_\{1\},\\ldots,a\_\{m\}such that
L:=\(f\(a1\)−f\(a0\),…,f\(am\)−f\(a0\)\)L\\mathrel\{:=\}\\bigl\(f\(a\_\{1\}\)\-f\(a\_\{0\}\),\\ldots,f\(a\_\{m\}\)\-f\(a\_\{0\}\)\\bigr\)is invertible, and defineL′L^\{\\prime\}analogously\. Then the CMM potential identity gives
T\(τ\(c\)\)=MT′\(c\)\+b,M=L−⊤\(L′\)⊤,T\(\\tau\(c\)\)=MT^\{\\prime\}\(c\)\+b,\\qquad M=L^\{\-\\top\}\(L^\{\\prime\}\)^\{\\top\},for someb∈ℝmb\\in\\mathbb\{R\}^\{m\}\. This is the∼A\\sim\_\{A\}\-identifiability relation of Theorem 1 ofKhemakhemet al\.\[[2020a](https://arxiv.org/html/2606.18509#bib.bib2)\]\. The further refinement to∼P\\sim\_\{P\}is the separate rigidity step in their Theorems 2–3\.
###### Proof\.
By[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2),Δa,a0𝐐¯′\(c\)=Δa,a0𝐐¯\(τ\(c\)\)\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)\. Centering atc0c\_\{0\}cancels the normalizers and gives
⟨f′\(a\)−f′\(a0\),T′\(c\)−T′\(c0\)⟩=⟨f\(a\)−f\(a0\),T\(τ\(c\)\)−T\(τ\(c0\)\)⟩\.\\langle f^\{\\prime\}\(a\)\-f^\{\\prime\}\(a\_\{0\}\),T^\{\\prime\}\(c\)\-T^\{\\prime\}\(c\_\{0\}\)\\rangle=\\langle f\(a\)\-f\(a\_\{0\}\),T\(\\tau\(c\)\)\-T\(\\tau\(c\_\{0\}\)\)\\rangle\.Stacking this identity overa1,…,ama\_\{1\},\\ldots,a\_\{m\}and solving withLLyields
T\(τ\(c\)\)−T\(τ\(c0\)\)=L−⊤\(L′\)⊤\{T′\(c\)−T′\(c0\)\}\.T\(\\tau\(c\)\)\-T\(\\tau\(c\_\{0\}\)\)=L^\{\-\\top\}\(L^\{\\prime\}\)^\{\\top\}\\\{T^\{\\prime\}\(c\)\-T^\{\\prime\}\(c\_\{0\}\)\\\}\.Absorbing the value atc0c\_\{0\}intobbgives the affine relation\. ∎
### E\.3Linear CRL
#### CMM translation\.
For the linear CRL model ofSquireset al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib6)\], let𝒜=\{0\}∪\[K\]\\mathcal\{A\}=\\\{0\\\}\\cup\[K\],𝒞=ℝd\\mathcal\{C\}=\\mathbb\{R\}^\{d\}, and𝒳=ℝp\\mathcal\{X\}=\\mathbb\{R\}^\{p\}\. The attributek∈𝒜k\\in\\mathcal\{A\}indexes an observational or interventional environment, the concept isC=Z\(k\)C=Z^\{\(k\)\}, and the feature isX\(k\)=GZ\(k\)X^\{\(k\)\}=GZ^\{\(k\)\}withG∈ℝp×dG\\in\\mathbb\{R\}^\{p\\times d\}full column rank\. In environmentkk, the latent variables satisfy
Z\(k\)=Bk−1ϵ,𝔼\[ϵ\]=0,Cov\(ϵ\)=Id,Z^\{\(k\)\}=B\_\{k\}^\{\-1\}\\epsilon,\\qquad\\mathbb\{E\}\[\\epsilon\]=0,\\qquad\\mathrm\{Cov\}\(\\epsilon\)=I\_\{d\},whereBk=Ωk−1/2\(Id−Ak\)B\_\{k\}=\\Omega\_\{k\}^\{\-1/2\}\(I\_\{d\}\-A\_\{k\}\)is the normalized structural factor obtained through a perfect intervention: doing such on a nodeiki\_\{k\}changes only theiki\_\{k\}\-th row ofB0B\_\{0\}, replacing it byλkeik⊤\\lambda\_\{k\}e\_\{i\_\{k\}\}^\{\\top\}\. The indexing kernel is𝐐=𝐓η\{\\mathbf\{Q\}\}=\\mathbf\{T\}\_\{\\eta\}, whereη\(k\)=Bk\\eta\(k\)=B\_\{k\}\. The concept\-modulation kernel mapsBBto the law ofB−1ϵB^\{\-1\}\\epsilon, and the mixing kernel is deterministic,𝐊=𝐓G\{\\mathbf\{K\}\}=\\mathbf\{T\}\_\{G\}\.
#### Second\-moment proof object\.
WritingH=G†H=G^\{\\dagger\}, the observed pseudoprecision in environmentkkis
Θk:=Cov\(X\(k\)\)†=H⊤Bk⊤BkH\.\\Theta\_\{k\}\\mathrel\{:=\}\\mathrm\{Cov\}\(X^\{\(k\)\}\)^\{\\dagger\}=H^\{\\top\}B\_\{k\}^\{\\top\}B\_\{k\}H\.The residual relabeling set is
S\(𝒢\):=\{σ:\[d\]→\[d\]bijective\|σ\(j\)\>σ\(i\)for every edgej→iin𝒢\}\.S\(\\mathcal\{G\}\)\\mathrel\{:=\}\\\{\\sigma\\colon\[d\]\\to\[d\]\\ \\text\{bijective\}\\;\|\\;\\sigma\(j\)\>\\sigma\(i\)\\ \\text\{for every edge \}j\\to i\\text\{ in \}\\mathcal\{G\}\\\}\.Thus the proof object inSquireset al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib6)\]is not an attribute\-potential equation but the family of observed pseudoprecision equations\.
#### Recovery of guarantees\.
###### Proposition 4\(CMM recovery of Theorem 2 of\[Squireset al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib6)\]\)\.
Let𝖬=\(𝐓η,𝐁,𝐓G\)\\mathsf\{M\}=\(\\mathbf\{T\}\_\{\\eta\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{G\}\)and𝖬′=\(𝐓η~,𝐁,𝐓G~\)\\mathsf\{M\}^\{\\prime\}=\(\\mathbf\{T\}\_\{\\widetilde\{\\eta\}\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{\\widetilde\{G\}\}\)be feature\-equivalent linear CRL CMMs with full\-column\-rank mixing matrices\. Then the mixing side of[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)forces the induced transition to be linear:
G~=GT,T=G†G~∈GL\(d\)\.\\widetilde\{G\}=GT,\\qquad T=G^\{\\dagger\}\\widetilde\{G\}\\in\\mathrm\{GL\}\(d\)\.The concept\-side equality gives
Cov\(Z\(k\)\)=TCov\(Z~\(k\)\)T⊤,\\mathrm\{Cov\}\(Z^\{\(k\)\}\)=T\\,\\mathrm\{Cov\}\(\\widetilde\{Z\}^\{\(k\)\}\)\\,T^\{\\top\},and hence
Θk=H⊤Bk⊤BkH=H~⊤B~k⊤B~kH~\.\\Theta\_\{k\}=H^\{\\top\}B\_\{k\}^\{\\top\}B\_\{k\}H=\\widetilde\{H\}^\{\\top\}\\widetilde\{B\}\_\{k\}^\{\\top\}\\widetilde\{B\}\_\{k\}\\widetilde\{H\}\.Under Assumptions 1–2 ofSquireset al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib6)\]with one perfect intervention per latent node, their Theorem 2 gives identifiability up toS\(𝒢\)S\(\\mathcal\{G\}\)\.
###### Proof\.
From the mixing side relation𝐓G~=𝜈𝐓G𝐓τ\\mathbf\{T\}\_\{\\widetilde\{G\}\}\{\\,\\overset\{\{\\nu\}\}\{=\}\\,\}\\mathbf\{T\}\_\{G\}\\mathbf\{T\}\_\{\\tau\}, we haveG~z=Gτ\(z\)\\widetilde\{G\}z=G\\tau\(z\)forν\{\\nu\}\-a\.e\.zz\. Multiplying byG†G^\{\\dagger\}givesτ\(z\)=Tz\\tau\(z\)=TzwithT=G†G~T=G^\{\\dagger\}\\widetilde\{G\}, and full column rank makesTTinvertible\. The concept\-side equality transports covariances byTT, which is equivalent to the displayed pseudoprecision equations\. These are exactly the equations analyzed inSquireset al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib6)\]; their generalized\-RQ and partial\-order argument gives the stated residual ambiguity\. ∎
### E\.4Gaussian CRL
#### CMM translation\.
For the Gaussian CRL model ofBuchholzet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib18)\], let𝒜=I∪\{0\}\\mathcal\{A\}=I\\cup\\\{0\\\},𝒞=ℝd\\mathcal\{C\}=\\mathbb\{R\}^\{d\}, and𝒳=ℝd′\\mathcal\{X\}=\\mathbb\{R\}^\{d^\{\\prime\}\}\. The attributei∈Ii\\in Iindexes an interventional environment, while0is the observational environment\. The latent concept isC=Z\(i\)C=Z^\{\(i\)\}, and the observed feature isX\(i\)=f\(Z\(i\)\)X^\{\(i\)\}=f\(Z^\{\(i\)\}\)\. In environmentii, the structural factor is
B\(i\)=\(D\(i\)\)−1/2\(Id−A\(i\)\),B^\{\(i\)\}=\(D^\{\(i\)\}\)^\{\-1/2\}\(I\_\{d\}\-A^\{\(i\)\}\),and only the row indexed by the target nodetit\_\{i\}differs from the observational factorB\(0\)B^\{\(0\)\}\. The scalarη\(i\)\\eta^\{\(i\)\}denotes the intervention shift parameter fromBuchholzet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib18)\]; it is not an indexing map\.
Define the environment\-specific modulator value by
λi:=\(B\(i\),η\(i\),ti\),i∈I∪\{0\},\{\\lambda\}\_\{i\}\\mathrel\{:=\}\(B^\{\(i\)\},\\eta^\{\(i\)\},t\_\{i\}\),\\qquad i\\in I\\cup\\\{0\\\},withη\(0\)=0\\eta^\{\(0\)\}=0and arbitraryt0t\_\{0\}, since the shift term vanishes in the observational environment\. The deterministic indexing kernel is
𝐐=𝐓i↦λi\.\{\\mathbf\{Q\}\}=\\mathbf\{T\}\_\{i\\mapsto\{\\lambda\}\_\{i\}\}\.The shared concept\-modulation kernel sends a structural triple\(B,α,t\)\(B,\\alpha,t\)to the Gaussian law induced byB−1\(ϵ\+αet\)B^\{\-1\}\(\\epsilon\+\\alpha e\_\{t\}\):
𝐁\(dc∣B,α,t\)=\|detB\|\(2π\)−d/2exp\{−12∥Bc−αet∥2\}dc\.\{\\mathbf\{B\}\}\(dc\\mid B,\\alpha,t\)=\|\\det B\|\\,\(2\\pi\)^\{\-d/2\}\\exp\\left\\\{\-\\frac\{1\}\{2\}\\lVert Bc\-\\alpha e\_\{t\}\\rVert^\{2\}\\right\\\}dc\.Thus, for𝐐¯=𝐁𝐐\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\},
p𝐐¯\(c∣i\)=\|detB\(i\)\|\(2π\)−d/2exp\{−12∥B\(i\)c−η\(i\)eti∥2\}\.p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid i\)=\|\\det B^\{\(i\)\}\|\\,\(2\\pi\)^\{\-d/2\}\\exp\\left\\\{\-\\frac\{1\}\{2\}\\lVert B^\{\(i\)\}c\-\\eta^\{\(i\)\}e\_\{t\_\{i\}\}\\rVert^\{2\}\\right\\\}\.Equivalently,
C∣A=i∼Z\(i\)=\(B\(i\)\)−1\(ϵ\+η\(i\)eti\),ϵ∼𝒩\(0,Id\)\.C\\mid A=i\\sim Z^\{\(i\)\}=\(B^\{\(i\)\}\)^\{\-1\}\(\\epsilon\+\\eta^\{\(i\)\}e\_\{t\_\{i\}\}\),\\qquad\\epsilon\\sim\\mathcal\{N\}\(0,I\_\{d\}\)\.The mixing kernel is deterministic,𝐊=𝐓f\{\\mathbf\{K\}\}=\\mathbf\{T\}\_\{f\}\.
#### Attribute potentials\.
Write
r\(i\):=\(B\(i\)\)⊤eti,Θ\(i\):=\(B\(i\)\)⊤B\(i\)\.r^\{\(i\)\}\\mathrel\{:=\}\(B^\{\(i\)\}\)^\{\\top\}e\_\{t\_\{i\}\},\\qquad\\Theta^\{\(i\)\}\\mathrel\{:=\}\(B^\{\(i\)\}\)^\{\\top\}B^\{\(i\)\}\.The attribute potential relative to0is
Δi,0𝐐¯\(c\)=−12c⊤\(Θ\(i\)−Θ\(0\)\)c\+η\(i\)\(r\(i\)\)⊤c\+κ\(i\),\\Delta\_\{i,0\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)=\-\\frac\{1\}\{2\}c^\{\\top\}\(\\Theta^\{\(i\)\}\-\\Theta^\{\(0\)\}\)c\+\\eta^\{\(i\)\}\(r^\{\(i\)\}\)^\{\\top\}c\+\\kappa^\{\(i\)\},where
κ\(i\)=log\|detB\(i\)\|\|detB\(0\)\|−12\(η\(i\)\)2\\kappa^\{\(i\)\}=\\log\\frac\{\|\\det B^\{\(i\)\}\|\}\{\|\\det B^\{\(0\)\}\|\}\-\\frac\{1\}\{2\}\(\\eta^\{\(i\)\}\)^\{2\}is independent ofcc\. Equivalently, up to anii\-dependent additive constant,−Δi,0𝐐¯\-\\Delta\_\{i,0\}^\{\\bar\{\\mathbf\{Q\}\}\}is the quadratic\-linear form
12c⊤\(Θ\(i\)−Θ\(0\)\)c−η\(i\)\(r\(i\)\)⊤c\.\\frac\{1\}\{2\}c^\{\\top\}\(\\Theta^\{\(i\)\}\-\\Theta^\{\(0\)\}\)c\-\\eta^\{\(i\)\}\(r^\{\(i\)\}\)^\{\\top\}c\.This is the quadratic transition object used inBuchholzet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib18)\]\.
#### Recovery of guarantees\.
###### Proposition 5\(CMM recovery of Theorems 1 and 3 of\[Buchholzet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib18)\]\)\.
Let𝖬\\mathsf\{M\}and𝖬′\\mathsf\{M\}^\{\\prime\}be feature\-equivalent Gaussian CRL CMMs\. Write the candidate quantities with tildes, so that the candidate mixing map isf~\\widetilde\{f\}, the candidate structural factors areB~\(i\)\\widetilde\{B\}^\{\(i\)\}, the candidate shift parameters areη~\(i\)\\widetilde\{\\eta\}^\{\(i\)\}, and
r~\(i\):=\(B~\(i\)\)⊤et~i,Θ~\(i\):=\(B~\(i\)\)⊤B~\(i\)\.\\widetilde\{r\}^\{\(i\)\}\\mathrel\{:=\}\(\\widetilde\{B\}^\{\(i\)\}\)^\{\\top\}e\_\{\\widetilde\{t\}\_\{i\}\},\\qquad\\widetilde\{\\Theta\}^\{\(i\)\}\\mathrel\{:=\}\(\\widetilde\{B\}^\{\(i\)\}\)^\{\\top\}\\widetilde\{B\}^\{\(i\)\}\.Let
τ:=f~−1∘f\\tau\\mathrel\{:=\}\\widetilde\{f\}^\{\-1\}\\circ fbe the transition from ground\-truth latent coordinates to candidate latent coordinates\. Then, for every interventioni∈Ii\\in I, there existsb\(i\)∈ℝb^\{\(i\)\}\\in\\mathbb\{R\}such that
12c⊤\(Θ\(i\)−Θ\(0\)\)c−η\(i\)\(r\(i\)\)⊤c=12τ\(c\)⊤\(Θ~\(i\)−Θ~\(0\)\)τ\(c\)−η~\(i\)\(r~\(i\)\)⊤τ\(c\)\+b\(i\)\.\\frac\{1\}\{2\}c^\{\\top\}\(\\Theta^\{\(i\)\}\-\\Theta^\{\(0\)\}\)c\-\\eta^\{\(i\)\}\(r^\{\(i\)\}\)^\{\\top\}c=\\frac\{1\}\{2\}\\tau\(c\)^\{\\top\}\(\\widetilde\{\\Theta\}^\{\(i\)\}\-\\widetilde\{\\Theta\}^\{\(0\)\}\)\\tau\(c\)\-\\widetilde\{\\eta\}^\{\(i\)\}\(\\widetilde\{r\}^\{\(i\)\}\)^\{\\top\}\\tau\(c\)\+b^\{\(i\)\}\.This is the quadratic transition identity used in the proof of Theorem 3 ofBuchholzet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib18)\]\. Under the assumptions of their Theorem 3, their rigidity argument impliesτ\(c\)=Tc\\tau\(c\)=Tc, equivalentlyf~=f∘T−1\\widetilde\{f\}=f\\circ T^\{\-1\}\. Under the additional perfect\-intervention assumptions of their Theorem 1, their Appendix B combines this linearity withSquireset al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib6)\]to obtain permutation\-and\-scaling identifiability\.
###### Proof\.
Feature equivalence and[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)give a latent transition\. Using the inverse\-direction transitionτ=f~−1∘f\\tau=\\widetilde\{f\}^\{\-1\}\\circ f,[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2)gives
Δi,0𝐐¯\(c\)=Δi,0𝐐¯′\(τ\(c\)\)\.\\Delta\_\{i,0\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)=\\Delta\_\{i,0\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(\\tau\(c\)\)\.Substituting the Gaussian potential expansions for𝐐¯\\bar\{\\mathbf\{Q\}\}and𝐐¯′\\bar\{\\mathbf\{Q\}\}^\{\\prime\}gives
−12c⊤\(Θ\(i\)−Θ\(0\)\)c\+η\(i\)\(r\(i\)\)⊤c\+κ\(i\)=−12τ\(c\)⊤\(Θ~\(i\)−Θ~\(0\)\)τ\(c\)\+η~\(i\)\(r~\(i\)\)⊤τ\(c\)\+κ~\(i\)\.\-\\frac\{1\}\{2\}c^\{\\top\}\(\\Theta^\{\(i\)\}\-\\Theta^\{\(0\)\}\)c\+\\eta^\{\(i\)\}\(r^\{\(i\)\}\)^\{\\top\}c\+\\kappa^\{\(i\)\}=\-\\frac\{1\}\{2\}\\tau\(c\)^\{\\top\}\(\\widetilde\{\\Theta\}^\{\(i\)\}\-\\widetilde\{\\Theta\}^\{\(0\)\}\)\\tau\(c\)\+\\widetilde\{\\eta\}^\{\(i\)\}\(\\widetilde\{r\}^\{\(i\)\}\)^\{\\top\}\\tau\(c\)\+\\widetilde\{\\kappa\}^\{\(i\)\}\.Multiplying by−1\-1and absorbing the constantκ~\(i\)−κ\(i\)\\widetilde\{\\kappa\}^\{\(i\)\}\-\\kappa^\{\(i\)\}intob\(i\)b^\{\(i\)\}gives the displayed quadratic identity\. ∎
### E\.5Score\-based CRL
#### CMM translation\.
In the setting ofVarıcıet al\.\[[2025](https://arxiv.org/html/2606.18509#bib.bib9)\], take𝒜=ℰ\\mathcal\{A\}=\\mathcal\{E\},𝒞=ℝn\\mathcal\{C\}=\\mathbb\{R\}^\{n\}, and let the modulator space be the space of admissible local\-mechanism values
ℒ=⨆G∈DAG\(\[n\]\)∏i=1n𝒦mk\(𝒞paG\(i\)→𝒞i\)\.\{\\mathscr\{L\}\}=\\bigsqcup\_\{G\\in\\mathrm\{DAG\}\(\[n\]\)\}\\prod\_\{i=1\}^\{n\}\\mathcal\{K\}^\{\\textnormal\{mk\}\}\(\\mathcal\{C\}\_\{\\textnormal\{pa\}\_\{G\}\(i\)\}\\to\\mathcal\{C\}\_\{i\}\)\.Eachλ∈ℒ\\lambda\\in\{\\mathscr\{L\}\}specifies local conditional densities\(piλ\)i=1n\(p\_\{i\}^\{\\lambda\}\)\_\{i=1\}^\{n\}and induces a DAG𝒢λ\\mathcal\{G\}\_\{\\lambda\}on\[n\]\[n\]\. Fori∈\[n\]i\\in\[n\], writepa𝒢λ\(i\)\\textnormal\{pa\}\_\{\\mathcal\{G\}\_\{\\lambda\}\}\(i\)for the parent set induced byλ\\lambda\. Letη:ℰ→ℒ\\eta\\colon\\mathcal\{E\}\\to\{\\mathscr\{L\}\}be the deterministic indexing map and writeλa:=η\(a\)\\lambda\_\{a\}\\mathrel\{:=\}\\eta\(a\)\. The indexing kernel is𝐐=𝐓η\{\\mathbf\{Q\}\}=\\mathbf\{T\}\_\{\\eta\}\. The shared concept\-modulation kernel maps a mechanism value to the Markov density over its induced graph:
𝐁\(dc∣λ\)=∏i=1npiλ\(ci∣cpa𝒢λ\(i\)\)dc\.\{\\mathbf\{B\}\}\(dc\\mid\\lambda\)=\\prod\_\{i=1\}^\{n\}p\_\{i\}^\{\\lambda\}\(c\_\{i\}\\mid c\_\{\\textnormal\{pa\}\_\{\\mathcal\{G\}\_\{\\lambda\}\}\(i\)\}\)\\,dc\.Thus, for𝐐¯=𝐁𝐓η\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\\mathbf\{T\}\_\{\\eta\},
p𝐐¯\(c∣a\)=∏i=1npiλa\(ci∣cpa𝒢λa\(i\)\)\.p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid a\)=\\prod\_\{i=1\}^\{n\}p\_\{i\}^\{\\lambda\_\{a\}\}\(c\_\{i\}\\mid c\_\{\\textnormal\{pa\}\_\{\\mathcal\{G\}\_\{\\lambda\_\{a\}\}\}\(i\)\}\)\.The mixing kernel is deterministic,𝐊=𝐓f\{\\mathbf\{K\}\}=\\mathbf\{T\}\_\{f\}\.
#### Attribute potentials\.
Fora,a′∈ℰa,a^\{\\prime\}\\in\\mathcal\{E\},
Δa,a′𝐐¯\(c\)=∑i=1nlogpiλa\(ci∣cpa𝒢λa\(i\)\)−∑i=1nlogpiλa′\(ci∣cpa𝒢λa′\(i\)\)\.\\Delta\_\{a,a^\{\\prime\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)=\\sum\_\{i=1\}^\{n\}\\log p\_\{i\}^\{\\lambda\_\{a\}\}\(c\_\{i\}\\mid c\_\{\\textnormal\{pa\}\_\{\\mathcal\{G\}\_\{\\lambda\_\{a\}\}\}\(i\)\}\)\-\\sum\_\{i=1\}^\{n\}\\log p\_\{i\}^\{\\lambda\_\{a^\{\\prime\}\}\}\(c\_\{i\}\\mid c\_\{\\textnormal\{pa\}\_\{\\mathcal\{G\}\_\{\\lambda\_\{a^\{\\prime\}\}\}\}\(i\)\}\)\.For a coupled hard\-intervention pair\(ai,a~i\)\(a\_\{i\},\\tilde\{a\}\_\{i\}\)targeting nodeii, the two mechanism values agree at all non\-iilocal mechanisms and have parent\-freeii\-th mechanisms\. Hence
Δai,a~i𝐐¯\(c\)=logpiλai\(ci\)−logpiλa~i\(ci\),\\Delta\_\{a\_\{i\},\\tilde\{a\}\_\{i\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)=\\log p\_\{i\}^\{\\lambda\_\{a\_\{i\}\}\}\(c\_\{i\}\)\-\\log p\_\{i\}^\{\\lambda\_\{\\tilde\{a\}\_\{i\}\}\}\(c\_\{i\}\),so∇cΔai,a~i𝐐¯\(c\)\\nabla\_\{c\}\\Delta\_\{a\_\{i\},\\tilde\{a\}\_\{i\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)is one\-sparse\. The differentiated attribute potential is the latent score\-difference object used in score\-based CRL\.
#### Recovery of guarantees\.
###### Proposition 6\(CMM recovery of Theorem 25 of\[Varıcıet al\.,[2025](https://arxiv.org/html/2606.18509#bib.bib9)\]\)\.
Let𝖬=\(𝐓η,𝐁,𝐓f\)\\mathsf\{M\}=\(\\mathbf\{T\}\_\{\\eta\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{f\}\)and𝖬′=\(𝐓η′,𝐁,𝐓f′\)\\mathsf\{M\}^\{\\prime\}=\(\\mathbf\{T\}\_\{\\eta^\{\\prime\}\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{f^\{\\prime\}\}\)be feature\-equivalent score\-based CRL CMMs\. Writeλa:=η\(a\)\\lambda\_\{a\}\\mathrel\{:=\}\\eta\(a\)andλa′:=η′\(a\)\\lambda^\{\\prime\}\_\{a\}\\mathrel\{:=\}\\eta^\{\\prime\}\(a\), so that the graphs in environmentaaare𝒢λa\\mathcal\{G\}\_\{\\lambda\_\{a\}\}and𝒢λa′\\mathcal\{G\}\_\{\\lambda^\{\\prime\}\_\{a\}\}\. Letτ\\taube the transition induced by[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)\. For any coupled hard\-intervention pair\(ai,a~i\)\(a\_\{i\},\\tilde\{a\}\_\{i\}\),
∇cΔai,a~i𝐐¯′\(c\)=∂cτ\(c\)⊤∇cΔai,a~i𝐐¯\(τ\(c\)\)\.\\nabla\_\{c\}\\Delta\_\{a\_\{i\},\\tilde\{a\}\_\{i\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\partial\_\{c\}\\tau\(c\)^\{\\top\}\\nabla\_\{c\}\\Delta\_\{a\_\{i\},\\tilde\{a\}\_\{i\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)\.Moreover, the two score contrasts are one\-sparse in the corresponding candidate and ground\-truth intervention targets\. This is the transported score\-contrast and sparsity object used byVarıcıet al\.\[[2025](https://arxiv.org/html/2606.18509#bib.bib9)\]; their Theorem 25 gives componentwise recovery and recovery of the mechanism\-induced anchor graph𝒢λa0\\mathcal\{G\}\_\{\\lambda\_\{a\_\{0\}\}\}up to the corresponding permutation of𝒢λa0′\\mathcal\{G\}\_\{\\lambda^\{\\prime\}\_\{a\_\{0\}\}\}\.
###### Proof\.
By[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2), subtracting the transported identities foraia\_\{i\}anda~i\\tilde\{a\}\_\{i\}givesΔai,a~i𝐐¯′\(c\)=Δai,a~i𝐐¯\(τ\(c\)\)\\Delta\_\{a\_\{i\},\\tilde\{a\}\_\{i\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a\_\{i\},\\tilde\{a\}\_\{i\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)\. Differentiating gives the displayed score\-contrast equation\. For a coupled hard\-intervention pair, all non\-target mechanisms cancel and the target mechanisms are parent\-free, so the ground\-truth and candidate score contrasts are one\-sparse in their respective targets\. This is exactly the sparsity property formalized in Theorem 7\(iii\) ofVarıcıet al\.\[[2025](https://arxiv.org/html/2606.18509#bib.bib9)\]; the componentwise and graph\-recovery steps are their Lemmas 23–24 and Theorem 25\. ∎
### E\.6Nonparametric CRL
#### CMM translation\.
For the nonparametric CRL model ofvon Kügelgenet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib8)\], let𝒜=ℰ\\mathcal\{A\}=\\mathcal\{E\},𝒞=ℝn\\mathcal\{C\}=\\mathbb\{R\}^\{n\}, and𝒳⊆ℝd\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\}\. Fix an observational anchore0∈ℰe\_\{0\}\\in\\mathcal\{E\}\. The concept variable isC=\(C1,…,Cn\)C=\(C\_\{1\},\\ldots,C\_\{n\}\), and the observed variable isX=f\(C\)X=f\(C\), wheref:ℝn→𝒳f\\colon\\mathbb\{R\}^\{n\}\\to\\mathcal\{X\}is a diffeomorphism onto its image\. Letℒ\{\\mathscr\{L\}\}be the space of admissible local\-mechanism values\. Eachλ∈ℒ\\lambda\\in\{\\mathscr\{L\}\}specifies local conditional densities\(pjλ\)j=1n\(p\_\{j\}^\{\\lambda\}\)\_\{j=1\}^\{n\}and induces a DAG𝒢λ\\mathcal\{G\}\_\{\\lambda\}on\[n\]\[n\]\. Letη:ℰ→ℒ\\eta\\colon\\mathcal\{E\}\\to\{\\mathscr\{L\}\}be the deterministic indexing map, writeλe:=η\(e\)\\lambda\_\{e\}\\mathrel\{:=\}\\eta\(e\), and setλ0:=λe0\\lambda\_\{0\}\\mathrel\{:=\}\\lambda\_\{e\_\{0\}\}\. The indexing kernel is𝐐=𝐓η\{\\mathbf\{Q\}\}=\\mathbf\{T\}\_\{\\eta\}\. The shared concept\-modulation kernel is
𝐁\(dc∣λ\)=∏j=1npjλ\(cj∣cpa𝒢λ\(j\)\)dc\.\{\\mathbf\{B\}\}\(dc\\mid\\lambda\)=\\prod\_\{j=1\}^\{n\}p\_\{j\}^\{\\lambda\}\(c\_\{j\}\\mid c\_\{\\textnormal\{pa\}\_\{\\mathcal\{G\}\_\{\\lambda\}\}\(j\)\}\)\\,dc\.Thus, for𝐐¯=𝐁𝐓η\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\\mathbf\{T\}\_\{\\eta\},
p𝐐¯\(c∣e\)=∏j=1npjλe\(cj∣cpa𝒢λe\(j\)\)\.p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid e\)=\\prod\_\{j=1\}^\{n\}p\_\{j\}^\{\\lambda\_\{e\}\}\(c\_\{j\}\\mid c\_\{\\textnormal\{pa\}\_\{\\mathcal\{G\}\_\{\\lambda\_\{e\}\}\}\(j\)\}\)\.The mixing kernel is deterministic,𝐊=𝐓f\{\\mathbf\{K\}\}=\\mathbf\{T\}\_\{f\}, and the ground\-truth anchor graph is the mechanism\-induced graph𝒢λ0\\mathcal\{G\}\_\{\\lambda\_\{0\}\}\.
#### Attribute potentials\.
Fore,e′∈ℰe,e^\{\\prime\}\\in\\mathcal\{E\}, define
Re,e′𝐐¯\(c\):=exp\{Δe,e′𝐐¯\(c\)\}=p𝐐¯\(c∣e\)p𝐐¯\(c∣e′\)\.R\_\{e,e^\{\\prime\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)\\mathrel\{:=\}\\exp\\\{\\Delta\_\{e,e^\{\\prime\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)\\\}=\\frac\{p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid e\)\}\{p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid e^\{\\prime\}\)\}\.For a paired perfect intervention\(e,e′\)\(e,e^\{\\prime\}\)on nodeii, all non\-target mechanisms agree and cancel in the ratio\. The target mechanisms are parent\-free, so with
ρie,e′\(u\):=p~ie\(u\)p~ie′\(u\),\\rho\_\{i\}^\{e,e^\{\\prime\}\}\(u\)\\mathrel\{:=\}\\frac\{\\widetilde\{p\}\_\{i\}^\{e\}\(u\)\}\{\\widetilde\{p\}\_\{i\}^\{e^\{\\prime\}\}\(u\)\},we have
Re,e′𝐐¯\(c\)=ρie,e′\(ci\)\.R\_\{e,e^\{\\prime\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)=\\rho\_\{i\}^\{e,e^\{\\prime\}\}\(c\_\{i\}\)\.Thus the CMM proof object is the paired environment density ratio used byvon Kügelgenet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib8)\]\.
#### Recovery of guarantees\.
###### Proposition 7\(CMM recovery of Theorem 3\.4 of\[von Kügelgenet al\.,[2023](https://arxiv.org/html/2606.18509#bib.bib8)\]\)\.
Let𝖬=\(𝐓η,𝐁,𝐓f\)\\mathsf\{M\}=\(\\mathbf\{T\}\_\{\\eta\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{f\}\)and𝖬′=\(𝐓η′,𝐁,𝐓h−1\)\\mathsf\{M\}^\{\\prime\}=\(\\mathbf\{T\}\_\{\\eta^\{\\prime\}\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{h^\{\-1\}\}\)be feature\-equivalent nonparametric CRL CMMs\. Writeλe:=η\(e\)\\lambda\_\{e\}\\mathrel\{:=\}\\eta\(e\),λe′:=η′\(e\)\\lambda^\{\\prime\}\_\{e\}\\mathrel\{:=\}\\eta^\{\\prime\}\(e\),λ0:=λe0\\lambda\_\{0\}\\mathrel\{:=\}\\lambda\_\{e\_\{0\}\}, andλ0′:=λe0′\\lambda^\{\\prime\}\_\{0\}\\mathrel\{:=\}\\lambda^\{\\prime\}\_\{e\_\{0\}\}\. Leth:𝒳→ℝnh\\colon\\mathcal\{X\}\\to\\mathbb\{R\}^\{n\}be the candidate unmixing map, and define the transition from candidate to ground\-truth coordinates by
τ\(z\):=f−1\(h−1\(z\)\)\.\\tau\(z\)\\mathrel\{:=\}f^\{\-1\}\(h^\{\-1\}\(z\)\)\.For any paired perfect intervention\(e,e′\)\(e,e^\{\\prime\}\)with ground\-truth targetiiand candidate targetjj,
ρj′e,e′\(zj\)=ρie,e′\(τi\(z\)\)\.\\rho\_\{j\}^\{\\prime e,e^\{\\prime\}\}\(z\_\{j\}\)=\\rho\_\{i\}^\{e,e^\{\\prime\}\}\(\\tau\_\{i\}\(z\)\)\.For thennpaired interventions used in Theorem 3\.4 ofvon Kügelgenet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib8)\], writingj=π\(i\)j=\\pi\(i\)gives
ρπ\(i\)′e,e′\(zπ\(i\)\)=ρie,e′\(τi\(z\)\)\.\\rho\_\{\\pi\(i\)\}^\{\\prime e,e^\{\\prime\}\}\(z\_\{\\pi\(i\)\}\)=\\rho\_\{i\}^\{e,e^\{\\prime\}\}\(\\tau\_\{i\}\(z\)\)\.This is the one\-dimensional ratio equation in Appendix C\.3 ofvon Kügelgenet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib8)\]; their theorem gives coordinatewise recovery and graph isomorphism between the mechanism\-induced anchor graphs𝒢λ0\\mathcal\{G\}\_\{\\lambda\_\{0\}\}and𝒢λ0′\\mathcal\{G\}\_\{\\lambda^\{\\prime\}\_\{0\}\}\.
###### Proof\.
By[Thm\.2](https://arxiv.org/html/2606.18509#Thmtheorem2), feature equivalence givesRe,e′𝐐¯′\(z\)=Re,e′𝐐¯\(τ\(z\)\)R\_\{e,e^\{\\prime\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(z\)=R\_\{e,e^\{\\prime\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(z\)\)\. For a paired perfect intervention, all non\-target mechanisms cancel in both ratios\. Thus the ground\-truth ratio isρie,e′\(τi\(z\)\)\\rho\_\{i\}^\{e,e^\{\\prime\}\}\(\\tau\_\{i\}\(z\)\), while the candidate ratio isρj′e,e′\(zj\)\\rho\_\{j\}^\{\\prime e,e^\{\\prime\}\}\(z\_\{j\}\)\. Substitution gives the displayed equation\. The nondegeneracy, coordinatewise\-recovery, and graph\-isomorphism steps are exactly those in Appendix C\.3 and Theorem 3\.4 ofvon Kügelgenet al\.\[[2023](https://arxiv.org/html/2606.18509#bib.bib8)\]\. ∎
### E\.7Perturbation modeling
#### CMM translation\.
Perturbation modeling invon Kügelgenet al\.\[[2025](https://arxiv.org/html/2606.18509#bib.bib13)\]fits the CMM template by taking the attributea∈𝒜=ℝKa\\in\\mathcal\{A\}=\\mathbb\{R\}^\{K\}to be a perturbation label, the conceptc∈𝒞=ℝkc\\in\\mathcal\{C\}=\\mathbb\{R\}^\{k\}to be the perturbation\-relevant latent state, and the featurex∈𝒳x\\in\\mathcal\{X\}to be the observed measurement\. Fix an anchor perturbationa0a\_\{0\}\. The deterministic indexing kernel is
𝐐=𝐓a↦W\(a−a0\)\.\{\\mathbf\{Q\}\}=\\mathbf\{T\}\_\{a\\mapsto W\(a\-a\_\{0\}\)\}\.The shared concept\-modulation kernel is
𝐁\(dc∣λ\)=𝒩\(λ,Ik\)\(dc\)\{\\mathbf\{B\}\}\(dc\\mid\{\\lambda\}\)=\\mathcal\{N\}\(\{\\lambda\},I\_\{k\}\)\(dc\)where𝒩\(μ,Σ\)\\mathcal\{N\}\(\\mu,\\Sigma\)denotes a Gaussian distribution with meanμ\\muand covarianceΣ\\Sigma\. Thus, for𝐐¯=𝐁𝐐\\bar\{\\mathbf\{Q\}\}=\{\\mathbf\{B\}\}\{\\mathbf\{Q\}\},
p𝐐¯\(c∣a\)=1\(2π\)k/2exp\{−12∥c−W\(a−a0\)∥2\}\.p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid a\)=\\frac\{1\}\{\(2\\pi\)^\{k/2\}\}\\exp\\left\\\{\-\\frac\{1\}\{2\}\\lVert c\-W\(a\-a\_\{0\}\)\\rVert^\{2\}\\right\\\}\.The mixing kernel is deterministic,𝐊=𝐓g\{\\mathbf\{K\}\}=\\mathbf\{T\}\_\{g\}, whereg:ℝk→𝒳g\\colon\\mathbb\{R\}^\{k\}\\to\\mathcal\{X\}is the observation map\.
#### Attribute potentials\.
For this Gaussian mean\-shift concept law, the attribute potential relative toa0a\_\{0\}is
Δa,a0𝐐¯\(c\)\\displaystyle\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)=logp𝐐¯\(c∣a\)−logp𝐐¯\(c∣a0\)\\displaystyle=\\log p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid a\)\-\\log p\_\{\\bar\{\\mathbf\{Q\}\}\}\(c\\mid a\_\{0\}\)=⟨W\(a−a0\),c⟩−12∥W\(a−a0\)∥2\.\\displaystyle=\\langle W\(a\-a\_\{0\}\),c\\rangle\-\\frac\{1\}\{2\}\\lVert W\(a\-a\_\{0\}\)\\rVert^\{2\}\.Therefore, after centering at anyc0∈ℝkc\_\{0\}\\in\\mathbb\{R\}^\{k\},
Δa,a0𝐐¯\(c\)−Δa,a0𝐐¯\(c0\)=⟨a−a0,W⊤\(c−c0\)⟩\.\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\)\-\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(c\_\{0\}\)=\\langle a\-a\_\{0\},W^\{\\top\}\(c\-c\_\{0\}\)\\rangle\.This is the proof object used in the perturbation model: perturbations act linearly on the latent mean, and centered attribute potentials expose the perturbation effect matrixWW\.
#### Recovery of guarantees\.
###### Proposition 8\(CMM recovery of perturbation identifiability and extrapolation\)\.
Let
𝖬=\(𝐓a↦W\(a−a0\),𝐁,𝐓g\),𝖬′=\(𝐓a↦W′\(a−a0\),𝐁,𝐓g′\)\\mathsf\{M\}=\(\\mathbf\{T\}\_\{a\\mapsto W\(a\-a\_\{0\}\)\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{g\}\),\\qquad\\mathsf\{M\}^\{\\prime\}=\(\\mathbf\{T\}\_\{a\\mapsto W^\{\\prime\}\(a\-a\_\{0\}\)\},\{\\mathbf\{B\}\},\\mathbf\{T\}\_\{g^\{\\prime\}\}\)be feature\-equivalent perturbation CMMs on observed perturbations𝒜o\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}, with injective deterministic mixing maps\. Letτ=g−1∘g′\\tau=g^\{\-1\}\\circ g^\{\\prime\}be the transition induced by[Thm\.1](https://arxiv.org/html/2606.18509#Thmtheorem1)\. Then, for everya∈𝒜oa\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\},
⟨W′\(a−a0\),c⟩−12∥W′\(a−a0\)∥2=⟨W\(a−a0\),τ\(c\)⟩−12∥W\(a−a0\)∥2\.\\langle W^\{\\prime\}\(a\-a\_\{0\}\),c\\rangle\-\\frac\{1\}\{2\}\\lVert W^\{\\prime\}\(a\-a\_\{0\}\)\\rVert^\{2\}=\\langle W\(a\-a\_\{0\}\),\\tau\(c\)\\rangle\-\\frac\{1\}\{2\}\\lVert W\(a\-a\_\{0\}\)\\rVert^\{2\}\.Equivalently, after centering atc0c\_\{0\},
⟨a−a0,\(W′\)⊤\(c−c0\)⟩=⟨a−a0,W⊤\(τ\(c\)−τ\(c0\)\)⟩\.\\langle a\-a\_\{0\},\(W^\{\\prime\}\)^\{\\top\}\(c\-c\_\{0\}\)\\rangle=\\langle a\-a\_\{0\},W^\{\\top\}\(\\tau\(c\)\-\\tau\(c\_\{0\}\)\)\\rangle\.Under the sufficient\-diversity condition ofvon Kügelgenet al\.\[[2025](https://arxiv.org/html/2606.18509#bib.bib13)\], this is their perturbation\-identifiability equation, with the remaining orthogonal\-rigidity step supplied by their theorem\. The same centered identity satisfies[Thm\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)withφ\(a\)=a\\varphi\(a\)=a, so feature equivalence extrapolates to everyaex\{a\_\{\\textnormal\{ex\}\}\}satisfying
aex−a0∈span\{a−a0\|a∈𝒜o\}\.\{a\_\{\\textnormal\{ex\}\}\}\-a\_\{0\}\\in\\operatorname\{span\}\\\{a\-a\_\{0\}\\;\|\\;a\\in\{\\mathcal\{A\}\_\{\\textnormal\{o\}\}\}\\\}\.
###### Proof\.
Substituting the Gaussian mean\-shift potential into the transported identityΔa,a0𝐐¯′\(c\)=Δa,a0𝐐¯\(τ\(c\)\)\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}^\{\\prime\}\}\(c\)=\\Delta\_\{a,a\_\{0\}\}^\{\\bar\{\\mathbf\{Q\}\}\}\(\\tau\(c\)\)gives the first display\. Subtracting the same identity atc0c\_\{0\}gives the centered display\. The identifiability conclusion is the rigidity theorem ofvon Kügelgenet al\.\[[2025](https://arxiv.org/html/2606.18509#bib.bib13)\], and the extrapolation conclusion follows from[Thms\.4](https://arxiv.org/html/2606.18509#Thmtheorem4)and[3](https://arxiv.org/html/2606.18509#Thmtheorem3)\. ∎
## Appendix FCMM Translations of prior work
Table 2:Variable\-level CMM translations for representative prior work\. Each row records the objects along the chainA→Λ→C→XA\\to\{\\Lambda\}\\to C\\to X, following notations in the original papers\. In the mechanism\-based CRL rows,𝒢λ\\mathcal\{G\}\_\{\\lambda\}denotes the induced DAG from local mechanismsλ\\lambda\.Table 3:Operator\-level CMM translations for representative prior work\. Each row records the three kernels in the CMM chainA→Λ→C→XA\\to\{\\Lambda\}\\to C\\to X\. ForHyvarinenet al\.\[[2019](https://arxiv.org/html/2606.18509#bib.bib4)\],Γ\(λ\)\\Gamma\(\{\\lambda\}\)is the normalizer determined by the tupleλ=\(qi\)i=1k\{\\lambda\}=\(q\_\{i\}\)\_\{i=1\}^\{k\}, not an additional modulator component\. For\[Khemakhemet al\.,[2020a](https://arxiv.org/html/2606.18509#bib.bib2)\],𝐊~\\widetilde\{\{\\mathbf\{K\}\}\}denotes a kernel that adds noise, which are injective through deconvolution process\. In the Gaussian CRL row,η\(i\)\\eta^\{\(i\)\}is the shift parameter used in that model, not an indexing map\. In the mechanism\-based CRL rows,λ\\lambdadenotes a tuple of local mechanisms and𝒢λ\\mathcal\{G\}\_\{\\lambda\}denotes the DAG induced by those mechanisms\. For topic models,Φ∈ℝr×n\\Phi\\in\\mathbb\{R\}^\{r\\times n\}denotes the topic\-word matrix, so𝐊\(V=v∣T=t\)=Φtv\{\\mathbf\{K\}\}\(V=v\\mid T=t\)=\\Phi\_\{tv\}\.Similar Articles
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
Modus is a decoder-only model that predicts any modality from any combination of others, achieving strong performance across diverse benchmarks without modality-specific heads or losses.
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
This paper introduces C4, a cognition-inspired evaluation framework for cross-concept understanding using Chinese idioms (Chengyu), and finds that current multimodal LLMs struggle with creatively encoded meaning.
Concept-based Visual Counterfactual Explanations with Diffusion Models
Introduces C-VCE, a diffusion framework that builds an interpretable concept bottleneck layer into the generative model, enabling human-guided visual counterfactual explanations without relying on external noise-robust classifiers.
ReCBM: Uncertainty-Gated Relational Reasoning for Concept Bottleneck Models
ReCBM proposes an uncertainty-gated relational reasoning framework for Concept Bottleneck Models, introducing concept relations like co-occurrence, implication, and exclusion to recover unreliable or missing concept states and improve interpretability and downstream predictions.
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference
MACS is a training-free inference framework that mitigates the straggler effect in expert parallelism for multimodal MoE MLLMs by introducing entropy-weighted load and dynamic modality-adaptive capacity mechanisms.