Coupled Usage-Sense Processes: 时间性与可归因的词汇语义变化

arXiv cs.CL 论文

摘要

本文介绍了Coupled Usage–Sense Processes (CUSP),一个用于分析词汇语义变化的层次化模型,通过量化词语使用中随时间变化的时间、机制和归因。

arXiv:2609.30974v1 Announce Type: new Abstract: Lexical semantic change is usually summarized by a scalar distance between independently sampled period distributions. This measures how much a word changed, but does not reveal when it changed, which mechanisms and component movements carried the change, or which usages support the attribution. We introduce Coupled Usage--Sense Processes (CUSP), which derives these answers from a single marginal preserving temporal process. A hierarchical coupling relates contextual distributions through latent usage components, while Markov composition makes adjacent and longer span correspondences compatible. Displacement operators quantify change magnitude and timing, split variation exactly between movement of component centers and reorganization within components, and attribute it to transported component pairs. Word-local modes resolve distinct directions of change and their activity over time, while representative passages from attributed components ground the analysis in text. Under a Gaussian mixture specialization, we prove parametric recovery of the operators and squared distances. Synthetic experiments support the predicted rate. CUSP remains competitive on English and German DWUG and recovers controlled Janus profiles while maintaining compositionally coherent transport. A large corpus of US court opinions demonstrates transition, mode, and passage attribution in unlabeled natural text. CUSP thus makes magnitude, timing, mechanism, movement, modes, and textual evidence compatible views of one lexical history.
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:46

# Coupled Usage–Sense Processes: Temporal and Attributable Lexical Semantic Change
Source: [https://arxiv.org/html/2609.30974](https://arxiv.org/html/2609.30974)
Haruka EzoeAffiliation:Graduate School of Information Science and TechnologyAffiliation:The University of TokyoRyohei HisanoAffiliation:Graduate School of Information Science and TechnologyAffiliation:The University of TokyoAffiliation:The Canon Institute for Global Studies

###### Abstract

Lexical semantic change is usually summarized by a scalar distance between independently sampled period distributions\. This measures how much a word changed, but does not reveal when it changed, which mechanisms and component movements carried the change, or which usages support the attribution\. We introduce Coupled Usage–Sense Processes \(CUSP\), which derives these answers from a single marginal preserving temporal process\. A hierarchical coupling relates contextual distributions through latent usage components, while Markov composition makes adjacent and longer span correspondences compatible\. Displacement operators quantify change magnitude and timing, split variation exactly between movement of component centers and reorganization within components, and attribute it to transported component pairs\. Word\-local modes resolve distinct directions of change and their activity over time, while representative passages from attributed components ground the analysis in text\. Under a Gaussian mixture specialization, we prove parametric recovery of the operators and squared distances\. Synthetic experiments support the predicted rate\. CUSP remains competitive on English and German DWUG and recovers controlled Janus profiles while maintaining compositionally coherent transport\. A large corpus of US court opinions demonstrates transition, mode, and passage attribution in unlabeled natural text\. CUSP thus makes magnitude, timing, mechanism, movement, modes, and textual evidence compatible views of one lexical history\.

## 1Introduction

Computational lexical semantic change \(LSC\) studies how a word’s contextual usages evolve\. Standard evaluations condense this history to a scalar distance between periods\. A score cannot reveal when change occurred, whether it reflects movement between usage components or reorganization within them, which transitions carried it, or which usages support the attribution\. A multi\-period analysis should answer these questions as compatible views of one lexical history\.

Prior work addresses individual parts of this problem\. Aligned and dynamic embeddings track changes in a word’s distributional representation across periods\[[1](https://arxiv.org/html/2609.30974#bib.bib17),[2](https://arxiv.org/html/2609.30974#bib.bib4)\]\. Contextual approaches compare individual usages directly or cluster them into usage types, then measure change with distances, divergences, or optimal transport\[[3](https://arxiv.org/html/2609.30974#bib.bib14),[4](https://arxiv.org/html/2609.30974#bib.bib27),[5](https://arxiv.org/html/2609.30974#bib.bib23)\]\. Sense based methods track prevalence, cluster continuity, and contributing senses\[[6](https://arxiv.org/html/2609.30974#bib.bib13),[7](https://arxiv.org/html/2609.30974#bib.bib18),[8](https://arxiv.org/html/2609.30974#bib.bib37),[9](https://arxiv.org/html/2609.30974#bib.bib39),[10](https://arxiv.org/html/2609.30974#bib.bib38)\], while diachronic similarity matrices characterize temporal patterns\[[11](https://arxiv.org/html/2609.30974#bib.bib22)\]\. Unbalanced transport identifies gains and losses in usage mass\[[12](https://arxiv.org/html/2609.30974#bib.bib21)\], and embedding axes aid interpretation\[[13](https://arxiv.org/html/2609.30974#bib.bib1),[14](https://arxiv.org/html/2609.30974#bib.bib2)\]\.[Chen et al\. \[15\]](https://arxiv.org/html/2609.30974#bib.bib9)identify interpretability as a major unresolved challenge, noting that current models often cannot explain how or why meanings change and explicitly interpretable approaches remain at an early stage\. Prior methods already track senses and attribute change \(Table[3](https://arxiv.org/html/2609.30974#A1.T3)\)\. To our knowledge, no existing LSC method builds one process that preserves period marginals and makes correspondence consistent across time, exactly decomposes its change score by mechanism and transported component pair, and links its word\-local modes and attributed movements to representative passages\.

period specificusage distributions\{γt\(w\)\}t=1T\\\{\\gamma\_\{t\}^\{\(w\)\}\\\}\_\{t=1\}^\{T\}hierarchical couplingof components and usagescoherent temporal processpreserving every period marginaldisplacement operators𝐌∙\(w\)​\(t,s\)\\mathbf\{M\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)CUSP: ONE COUPLED TEMPORAL PROCESS, FOUR COMPATIBLE LSC READOUTSHow much and when?change magnitude andtemporal profileWhich mechanism?movement between centersand variation within componentsWhich component movement?transported component pairsand supporting passagesWhich directions?word\-local modesover timeFigure 1:CUSP turns independent period distributions into one marginal preserving temporal process\. Its displacement operators yield compatible LSC readouts of magnitude and timing, mechanism, component and passage attribution, and word\-local temporal modes\.Diachronic corpora provide independent samples from each period’s usage distribution, not correspondences between periods\. A directly optimizedt→t\+2t\{\\to\}t\{\+\}2transport plan can disagree with the plan composed throught\+1t\{\+\}1\. Scores and transition explanations built from those separate plans can therefore describe different histories\. What is needed is one process that preserves each period marginal and makes all readouts refer to the same temporal correspondence\.

We introduce Coupled Usage–Sense Processes \(CUSP\) to construct exactly this object \(Figure[1](https://arxiv.org/html/2609.30974#S1.F1)\)\.CUSP makes three technical contributions\. \(i\) A hierarchical coupling matches contextual distributions within component pairs, transports their prevalence mass with the prescribed marginals, and composes adjacent plans into coherent longer span correspondence\. \(ii\) Displacement operators quantify magnitude and timing and decompose change exactly by center movement, within component variation, and transported component pair\. The summed mean operator yields word\-local modes, while representative passages make the attributed movements inspectable\. \(iii\) Building on the second\-moment geometry of MENT\[[16](https://arxiv.org/html/2609.30974#bib.bib11)\], we prove marginal preservation, temporal compatibility, exact decompositions, and parametric recovery of the operators and squared distances for the trace and retained modes under identification and optimization stability conditions\. These are compatible readouts of one fitted temporal process\.

Empirically, we test each level of this construction in increasing proximity to natural LSC\. Synthetic data support the predicted recovery rate\. DWUG establishes that the operator trace retains standard scalar validity while its exact decomposition distinguishes mechanisms\. Janus tests temporal profile recovery and compositional coherence across several periods\. A large corpus of U\.S\. court opinions shows that the complete construction connects an unsupervised temporal signal to component movements, word\-local modes, and directly inspectable passages\. Together, these settings show that CUSP preserves the conventional answer to*how much*while adding compatible answers to*when*,*through which movement*, and*with what textual evidence*\.

## 2Coupled Usage–Sense Processes

A fixed encoder places independently sampled usages from each period in a common embedding space, but shared coordinates do not determine temporal correspondence\. CUSP represents each period distribution as a mixture of latent usage components and couples their prevalences and contextual distributions into a joint process whose period marginals reproduce the observed distributions\.

### 2\.1Usage distributions in each period

LetVVbe the vocabulary and\[T\]=\{1,…,T\}\[T\]=\\\{1,\\ldots,T\\\}the periods\. For wordw∈Vw\\in Vand periodt∈\[T\]t\\in\[T\], let

Xt,i\(w\)​∼i\.i\.d\.​γt\(w\),γt\(w\)=∑k=1Kt\(w\)πt,k\(w\)​νt,k\(w\),νt,k\(w\)∈𝒫2​\(ℝd\)\.X\_\{t,i\}^\{\(w\)\}\\overset\{\\mathrm\{i\.i\.d\.\}\}\{\\sim\}\\gamma\_\{t\}^\{\(w\)\},\\qquad\\gamma\_\{t\}^\{\(w\)\}=\\sum\_\{k=1\}^\{K\_\{t\}^\{\(w\)\}\}\\pi\_\{t,k\}^\{\(w\)\}\\nu\_\{t,k\}^\{\(w\)\},\\qquad\\nu\_\{t,k\}^\{\(w\)\}\\in\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)\.Here,Xt,i\(w\)X\_\{t,i\}^\{\(w\)\}is a contextual usage embedding,Kt\(w\)K\_\{t\}^\{\(w\)\}counts components,πt,k\(w\)\\pi\_\{t,k\}^\{\(w\)\}are positive prevalence weights summing to one, andνt,k\(w\)\\nu\_\{t,k\}^\{\(w\)\}is the contextual distribution of componentkk\. Components are period local, so equal indices across periods do not imply correspondence\. Mixtures are fitted independently by period, and their marginals do not identify temporal correspondence\. After coupling, prevalence changes describe relative mass redistribution, while component changes describe movement or reshaping\.

### 2\.2Hierarchical coupling and Markov composition

We use optimal transport hierarchically: contextual distributions are coupled within each candidate component pair, and their induced costs define transport between component prevalences, following the general structure of mixture and hierarchical OT\[[17](https://arxiv.org/html/2609.30974#bib.bib8),[18](https://arxiv.org/html/2609.30974#bib.bib10),[19](https://arxiv.org/html/2609.30974#bib.bib36)\]\. For each adjacent pair\(t,t\+1\)\(t,t\+1\)and possible component transitionk→ℓk\\to\\ell, CUSP first selects a contextual coupling

ηt,k,ℓ\(w\)=arg⁡min⁡∫η∈𝒜t,k,ℓ\(w\)⁡c⁡\(x,y\)​𝑑η​\(x,y\),𝒜t,k,ℓ\(w\)⊆Π⁡\(νt,k\(w\),νt\+1,ℓ\(w\)\),\\eta\_\{t,k,\\ell\}^\{\(w\)\}=\\arg\\min\_\{\\eta\\in\\mathcal\{A\}\_\{t,k,\\ell\}^\{\(w\)\}\}\\int c\(x,y\)\\,d\\eta\(x,y\),\\qquad\\mathcal\{A\}\_\{t,k,\\ell\}^\{\(w\)\}\\subseteq\\Pi\\\!\\left\(\\nu\_\{t,k\}^\{\(w\)\},\\nu\_\{t\+1,\\ell\}^\{\(w\)\}\\right\),\(1\)where𝒜t,k,ℓ\(w\)\\mathcal\{A\}\_\{t,k,\\ell\}^\{\(w\)\}is the admissible class of contextual couplings for the component pair and equalsΠ⁡\(νt,k\(w\),νt\+1,ℓ\(w\)\)\\Pi\(\\nu\_\{t,k\}^\{\(w\)\},\\nu\_\{t\+1,\\ell\}^\{\(w\)\}\)in the Gaussian specialization below\. We denote the minimum cost byCt,k,ℓ\(w\)C\_\{t,k,\\ell\}^\{\(w\)\}\. CUSP then transports the component prevalences using

Qt\(w\)=arg⁡minQ\\displaystyle Q\_\{t\}^\{\(w\)\}=\\arg\\min\_\{Q\}∑k,ℓQ⁡\(k,ℓ\)​Ct,k,ℓ\(w\)\\displaystyle\\sum\_\{k,\\ell\}Q\(k,\\ell\)C\_\{t,k,\\ell\}^\{\(w\)\}\(2\)s\.t\.\\displaystyle\\text\{s\.t\.\}∑ℓQ\(k,ℓ\)=πt,k\(w\),∑kQ\(k,ℓ\)=πt\+1,ℓ\(w\),Q\(k,ℓ\)≥0\.\\displaystyle\\sum\_\{\\ell\}Q\(k,\\ell\)=\\pi\_\{t,k\}^\{\(w\)\},\\qquad\\sum\_\{k\}Q\(k,\\ell\)=\\pi\_\{t\+1,\\ell\}^\{\(w\)\},\\qquad Q\(k,\\ell\)\\geq 0\.The two transport levels have distinct roles\.Ct,k,ℓ\(w\)C\_\{t,k,\\ell\}^\{\(w\)\}measures the contextual cost of a candidate component transition, whileQt\(w\)Q\_\{t\}^\{\(w\)\}allocates prevalence mass among those candidates subject to the two observed period marginals\.

###### Assumption 2\.1\(Uniqueness of the coupling construction\)\.

For every adjacent period pair, each problem in \([1](https://arxiv.org/html/2609.30974#S2.E1)\) has a unique finite\-cost minimizer, and \([2](https://arxiv.org/html/2609.30974#S2.E2)\) has a unique minimizer\.

The adjacent plans share the prescribed intermediate marginals and can therefore be composed into a joint coupling\[[20](https://arxiv.org/html/2609.30974#bib.bib28)\]\. CUSP uses Markov composition to induce non\-adjacent correspondences from adjacent plans\. Specifically, let\(Ω,ℱ,ℙ\)\(\\Omega,\\mathcal\{F\},\\mathbb\{P\}\)be a probability space on which the latent component processZ1:T\(w\)Z\_\{1:T\}^\{\(w\)\}and the contextual state processX1:T\(w\)X\_\{1:T\}^\{\(w\)\}are defined\. The adjacent plans define a Markov chain over usage components local to each period:

Pr⁡\(Z1\(w\)=k\)=π1,k\(w\),Pt\(w\)​\(k,ℓ\):=Pr⁡\(Zt\+1\(w\)=ℓ∣Zt\(w\)=k\)=Qt\(w\)​\(k,ℓ\)πt,k\(w\)\.\\Pr\(Z\_\{1\}^\{\(w\)\}=k\)=\\pi\_\{1,k\}^\{\(w\)\},\\qquad P\_\{t\}^\{\(w\)\}\(k,\\ell\):=\\Pr\(Z\_\{t\+1\}^\{\(w\)\}=\\ell\\mid Z\_\{t\}^\{\(w\)\}=k\)=\\frac\{Q\_\{t\}^\{\(w\)\}\(k,\\ell\)\}\{\\pi\_\{t,k\}^\{\(w\)\}\}\.Disintegratingηt,k,ℓ\(w\)​\(d​x,d​y\)=νt,k\(w\)​\(d​x\)​𝒦t,k,ℓ\(w\)​\(x,d​y\)\\eta\_\{t,k,\\ell\}^\{\(w\)\}\(dx,dy\)=\\nu\_\{t,k\}^\{\(w\)\}\(dx\)\\mathcal\{K\}\_\{t,k,\\ell\}^\{\(w\)\}\(x,dy\)defines the corresponding contextual process conditional on the entire latent component sequenceZ1:T\(w\)Z\_\{1:T\}^\{\(w\)\}:

X1\(w\)∣Z1:T\(w\)∼ν1,Z1\(w\)\(w\),Xt\+1\(w\)∣\(Xt\(w\),Z1:T\(w\)\)∼𝒦t,Zt\(w\),Zt\+1\(w\)\(w\)\(Xt\(w\),⋅\)\.X\_\{1\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}\\sim\\nu\_\{1,Z\_\{1\}^\{\(w\)\}\}^\{\(w\)\},\\qquad X\_\{t\+1\}^\{\(w\)\}\\mid\(X\_\{t\}^\{\(w\)\},Z\_\{1:T\}^\{\(w\)\}\)\\sim\\mathcal\{K\}\_\{t,Z\_\{t\}^\{\(w\)\},Z\_\{t\+1\}^\{\(w\)\}\}^\{\(w\)\}\(X\_\{t\}^\{\(w\)\},\\cdot\)\.
The construction preserves both levels of the observed mixture representation\.

###### Proposition 2\.2\(Marginal preservation and temporal compatibility\)\.

For everyw∈Vw\\in V,t∈\[T\]t\\in\[T\], andk∈\[Kt\(w\)\]k\\in\[K\_\{t\}^\{\(w\)\}\],

Pr\(Zt\(w\)=k\)=πt,k\(w\),ℒ\(Xt\(w\)∣Z1:T\(w\)\)=νt,Zt\(w\)\(w\)\.\\Pr\(Z\_\{t\}^\{\(w\)\}=k\)=\\pi\_\{t,k\}^\{\(w\)\},\\qquad\\mathcal\{L\}\(X\_\{t\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}\)=\\nu\_\{t,Z\_\{t\}^\{\(w\)\}\}^\{\(w\)\}\.Consequently,

ℒ⁡\(X1\(w\),…,XT\(w\)\)∈Π⁡\(γ1\(w\),…,γT\(w\)\)\.\\mathcal\{L\}\(X\_\{1\}^\{\(w\)\},\\ldots,X\_\{T\}^\{\(w\)\}\)\\in\\Pi\(\\gamma\_\{1\}^\{\(w\)\},\\ldots,\\gamma\_\{T\}^\{\(w\)\}\)\.\(3\)

The correspondence fromtttossis therefore induced through the intervening adjacent plans, rather than being optimized independently\. All period pairs belong to one process with the observed marginals, and there is no assertion of trajectories for individual utterances\.

### 2\.3Full, mean, and deviation processes

Letmt,k\(w\)=∫x​d​νt,k\(w\)​\(x\)m\_\{t,k\}^\{\(w\)\}=\\int x\\,d\\nu\_\{t,k\}^\{\(w\)\}\(x\)\. The coupled contextual state can be separated into a component center process and a centered deviation process within components:

ϕfull\(w\)​\(t\)=Xt\(w\),ϕmean\(w\)​\(t\)=mt,Zt\(w\)\(w\),ϕdev\(w\)​\(t\)=Xt\(w\)−mt,Zt\(w\)\(w\)\.\\phi\_\{\\mathrm\{full\}\}^\{\(w\)\}\(t\)=X\_\{t\}^\{\(w\)\},\\qquad\\phi\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t\)=m\_\{t,Z\_\{t\}^\{\(w\)\}\}^\{\(w\)\},\\qquad\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t\)=X\_\{t\}^\{\(w\)\}\-m\_\{t,Z\_\{t\}^\{\(w\)\}\}^\{\(w\)\}\.
These processes satisfy the following reconstruction and centering properties:

ϕfull\(w\)\(t\)=ϕmean\(w\)\(t\)\+ϕdev\(w\)\(t\),𝔼\[ϕdev\(w\)\(t\)∣Z1:T\(w\)\]=0\.\\phi\_\{\\mathrm\{full\}\}^\{\(w\)\}\(t\)=\\phi\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t\)\+\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t\),\\qquad\\mathbb\{E\}\[\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t\)\\mid Z\_\{1:T\}^\{\(w\)\}\]=0\.\(4\)A component may therefore move through the embedding space, or its usages may broaden, contract, or rotate around its center\. The full process retains both mechanisms\.

## 3Quantifying and attributing semantic change

The coupled process provides compatible answers to how much a word has changed, when it changed, and which mechanisms and component movements carried the change\. Second\-moment operators retain directional information, their traces yield scalar change scores and temporal profiles, and their exact decompositions attribute change to movement between components, variation within components, and transported component pairs\.

### 3\.1Displacement operators and temporal profiles

For∙∈\{full,mean,dev\}\\bullet\\in\\\{\\mathrm\{full\},\\mathrm\{mean\},\\mathrm\{dev\}\\\}, define

Δ​ϕ∙\(w\)​\(t,s\):=ϕ∙\(w\)​\(t\)−ϕ∙\(w\)​\(s\),𝐌∙\(w\)​\(t,s\):=𝔼⁡\[Δ​ϕ∙\(w\)​\(t,s\)​Δ​ϕ∙\(w\)​\(t,s\)⊤\]\.\\Delta\\phi\_\{\\bullet\}^\{\(w\)\}\(t,s\):=\\phi\_\{\\bullet\}^\{\(w\)\}\(t\)\-\\phi\_\{\\bullet\}^\{\(w\)\}\(s\),\\qquad\\mathbf\{M\}\_\{\\bullet\}^\{\(w\)\}\(t,s\):=\\mathbb\{E\}\[\\Delta\\phi\_\{\\bullet\}^\{\(w\)\}\(t,s\)\\Delta\\phi\_\{\\bullet\}^\{\(w\)\}\(t,s\)^\{\\top\}\]\.The trace distance

d∙,tr\(w\)​\(t,s\):=\{tr⁡𝐌∙\(w\)​\(t,s\)\}1/2d\_\{\\bullet,\\mathrm\{tr\}\}^\{\(w\)\}\(t,s\):=\\\{\\operatorname\{tr\}\\mathbf\{M\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)\\\}^\{1/2\}measures the total coupled displacement between two periods\. Evaluating this distance on adjacent pairs yields a temporal change profile\. The trace reflects how much change occurred, while the operator retains the embedding directions carrying that displacement\. Adjacent evaluations locate the change in time, and non\-adjacent evaluations inherit the same composed temporal correspondence\. For any fixed word\-local orthonormal basis\{𝒖r\(w\)\}r=1d\\\{\\bm\{u\}\_\{r\}^\{\(w\)\}\\\}\_\{r=1\}^\{d\}, define

d∙,r\(w\)​\(t,s\):=\{\(𝒖r\(w\)\)⊤​𝐌∙\(w\)​\(t,s\)​𝒖r\(w\)\}1/2\.d\_\{\\bullet,r\}^\{\(w\)\}\(t,s\):=\\\{\(\\bm\{u\}\_\{r\}^\{\(w\)\}\)^\{\\top\}\\mathbf\{M\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)\\bm\{u\}\_\{r\}^\{\(w\)\}\\\}^\{1/2\}\.
The trace summarizes the total semantic variation, while the mode\-wise quantities provide its orthogonal decomposition:

\(d∙,tr\(w\)​\(t,s\)\)2=∑r=1d\(d∙,r\(w\)​\(t,s\)\)2\.\\left\(d\_\{\\bullet,\\mathrm\{tr\}\}^\{\(w\)\}\(t,s\)\\right\)^\{2\}=\\sum\_\{r=1\}^\{d\}\\left\(d\_\{\\bullet,r\}^\{\(w\)\}\(t,s\)\\right\)^\{2\}\.\(5\)
Remark\.Becauseνt,k\(w\)∈𝒫2​\(ℝd\)\\nu\_\{t,k\}^\{\(w\)\}\\in\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\), allϕ∙\(w\)​\(t\)\\phi\_\{\\bullet\}^\{\(w\)\}\(t\)belong toL2​\(Ω,ℝd\)L^\{2\}\(\\Omega;\\mathbb\{R\}^\{d\}\)\. Hence𝐌∙\(w\)​\(t,s\)\\mathbf\{M\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)and the resulting distances are finite\. Moreover, the trace and mode\-wise semantic distances define finite pseudometrics on\[T\]\[T\], with the zero\-distance conditions given byϕ∙\(w\)​\(t\)=ϕ∙\(w\)​\(s\)almost surely\\phi\_\{\\bullet\}^\{\(w\)\}\(t\)=\\phi\_\{\\bullet\}^\{\(w\)\}\(s\)\\quad\\text\{almost surely\}, and\(𝒖r\(w\)\)⊤​ϕ∙\(w\)​\(t\)=\(𝒖r\(w\)\)⊤​ϕ∙\(w\)​\(s\)almost surely\(\\bm\{u\}\_\{r\}^\{\(w\)\}\)^\{\\top\}\\phi\_\{\\bullet\}^\{\(w\)\}\(t\)=\(\\bm\{u\}\_\{r\}^\{\(w\)\}\)^\{\\top\}\\phi\_\{\\bullet\}^\{\(w\)\}\(s\)\\quad\\text\{almost surely\}, respectively\. These properties are the same as those of the corresponding distances defined in\[[16](https://arxiv.org/html/2609.30974#bib.bib11)\]\. Appendix[E\.1](https://arxiv.org/html/2609.30974#A5.SS1)gives the details\.

### 3\.2Exact mechanism and component transition attribution

Conditional centering removes the mean–deviation cross term\.

###### Theorem 3\.1\(Exact mean–deviation decomposition\)\.

For everyw∈Vw\\in Vandt,s∈\[T\]t,s\\in\[T\],

𝐌full\(w\)​\(t,s\)=𝐌mean\(w\)​\(t,s\)\+𝐌dev\(w\)​\(t,s\)\.\\mathbf\{M\}\_\{\\mathrm\{full\}\}^\{\(w\)\}\(t,s\)=\\mathbf\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s\)\+\\mathbf\{M\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)\.\(6\)Consequently, forρ∈\{tr,1,…,d\}\\rho\\in\\\{\\mathrm\{tr\},1,\\ldots,d\\\},

\(dfull,ρ\(w\)​\(t,s\)\)2=\(dmean,ρ\(w\)​\(t,s\)\)2\+\(ddev,ρ\(w\)​\(t,s\)\)2\.\\left\(d\_\{\\mathrm\{full\},\\rho\}^\{\(w\)\}\(t,s\)\\right\)^\{2\}=\\left\(d\_\{\\mathrm\{mean\},\\rho\}^\{\(w\)\}\(t,s\)\\right\)^\{2\}\+\\left\(d\_\{\\mathrm\{dev\},\\rho\}^\{\(w\)\}\(t,s\)\\right\)^\{2\}\.\(7\)

Further conditioning on\(Zt\(w\),Zs\(w\)\)\(Z\_\{t\}^\{\(w\)\},Z\_\{s\}^\{\(w\)\}\)attributes each operator to transported component pairs\. Define

𝐂mean\(w\)​\(t,s,k,ℓ\)\\displaystyle\\mathbf\{C\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s;k,\\ell\):=\(mt,k\(w\)−ms,ℓ\(w\)\)​\(mt,k\(w\)−ms,ℓ\(w\)\)⊤,\\displaystyle:=\(m\_\{t,k\}^\{\(w\)\}\-m\_\{s,\\ell\}^\{\(w\)\}\)\(m\_\{t,k\}^\{\(w\)\}\-m\_\{s,\\ell\}^\{\(w\)\}\)^\{\\top\},𝐂dev\(w\)​\(t,s,k,ℓ\)\\displaystyle\\mathbf\{C\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s;k,\\ell\):=Cov⁡\(Xt\(w\)−Xs\(w\)∣Zt\(w\)=k,Zs\(w\)=ℓ\),\\displaystyle:=\\operatorname\{Cov\}\(X\_\{t\}^\{\(w\)\}\-X\_\{s\}^\{\(w\)\}\\mid Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\),q\(w\)​\(t,s,k,ℓ\)\\displaystyle q^\{\(w\)\}\(t,s;k,\\ell\):=Pr⁡\(Zt\(w\)=k,Zs\(w\)=ℓ\)\.\\displaystyle:=\\Pr\\left\(Z\_\{t\}^\{\(w\)\}=k,\\,Z\_\{s\}^\{\(w\)\}=\\ell\\right\)\.The component\-pair covariance𝐂dev\(w\)​\(t,s,k,ℓ\)\\mathbf\{C\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s;k,\\ell\)is defined only whenq\(w\)​\(t,s,k,ℓ\)\>0q^\{\(w\)\}\(t,s;k,\\ell\)\>0\. Zero\-mass pairs contribute nothing and are omitted from the weighted sums\.

###### Theorem 3\.2\(Exact component transition attribution\)\.

For∙∈\{mean,dev\}\\bullet\\in\\\{\\mathrm\{mean\},\\mathrm\{dev\}\\\},

𝐌∙\(w\)​\(t,s\)=∑k=1Kt\(w\)∑ℓ=1Ks\(w\)q\(w\)​\(t,s,k,ℓ\)​𝐂∙\(w\)​\(t,s,k,ℓ\)\.\\mathbf\{M\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)=\\sum\_\{k=1\}^\{K\_\{t\}^\{\(w\)\}\}\\sum\_\{\\ell=1\}^\{K\_\{s\}^\{\(w\)\}\}q^\{\(w\)\}\(t,s;k,\\ell\)\\mathbf\{C\}\_\{\\bullet\}^\{\(w\)\}\(t,s;k,\\ell\)\.\(8\)Consequently,

\(d∙,tr\(w\)​\(t,s\)\)2\\displaystyle\\left\(d\_\{\\bullet,\\mathrm\{tr\}\}^\{\(w\)\}\(t,s\)\\right\)^\{2\}=∑k,ℓq\(w\)​\(t,s,k,ℓ\)​tr⁡𝐂∙\(w\)​\(t,s,k,ℓ\),\\displaystyle=\\sum\_\{k,\\ell\}q^\{\(w\)\}\(t,s;k,\\ell\)\\operatorname\{tr\}\\mathbf\{C\}\_\{\\bullet\}^\{\(w\)\}\(t,s;k,\\ell\),\(9\)\(d∙,r\(w\)​\(t,s\)\)2\\displaystyle\\left\(d\_\{\\bullet,r\}^\{\(w\)\}\(t,s\)\\right\)^\{2\}=∑k,ℓq\(w\)​\(t,s,k,ℓ\)​\(𝒖r\(w\)\)⊤​𝐂∙\(w\)​\(t,s,k,ℓ\)​𝒖r\(w\)\.\\displaystyle=\\sum\_\{k,\\ell\}q^\{\(w\)\}\(t,s;k,\\ell\)\(\\bm\{u\}\_\{r\}^\{\(w\)\}\)^\{\\top\}\\mathbf\{C\}\_\{\\bullet\}^\{\(w\)\}\(t,s;k,\\ell\)\\bm\{u\}\_\{r\}^\{\(w\)\}\.\(10\)

Together, Theorems[3\.1](https://arxiv.org/html/2609.30974#S3.Thmtheorem1)and[3\.2](https://arxiv.org/html/2609.30974#S3.Thmtheorem2)form an exact accounting of change\. Every unit of the squared trace or mode\-wise variation is assigned to between center or within component change and to transported component pairs\. This attribution decomposes the score itself rather than fitting a separate explanatory model\. In the empirical analysis, passages assigned to the selected source and target components make these numerical attributions inspectable\.

### 3\.3Word\-local temporal modes

To adapt the basis to the history of wordww, we use the aggregated mean operator

ℳmean\(w\):=∑t=1T−1𝐌mean\(w\)​\(t,t\+1\)\.\\mathcal\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}:=\\sum\_\{t=1\}^\{T\-1\}\\mathbf\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,t\+1\)\.We take the orthonormal eigenvectors ofℳmean\(w\)\\mathcal\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}as the word specific basis\{𝒖r\(w\)\}r=1d\\\{\\bm\{u\}\_\{r\}^\{\(w\)\}\\\}\_\{r=1\}^\{d\}, with eigenvalues ordered decreasingly:

ℳmean\(w\)​𝒖r\(w\)=λr\(w\)​𝒖r\(w\)\.\\mathcal\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\\bm\{u\}\_\{r\}^\{\(w\)\}=\\lambda\_\{r\}^\{\(w\)\}\\bm\{u\}\_\{r\}^\{\(w\)\}\.
LetR\(w\)∈\[d\]R^\{\(w\)\}\\in\[d\]denote the number of modes retained for individual analysis\. We require their eigenvalues to be non\-degenerate111Forr\>R\(w\)r\>R^\{\(w\)\}, the eigenvectors need not be uniquely determined in the presence of eigenvalue degeneracy, but this freedom is immaterial because these modes are not used for individual analysis\.\.

###### Assumption 3\.3\(Non\-degenerate retained modes\)\.

There exists a constantδ\>0\\delta\>0such that, for everyr∈\[R\(w\)\]r\\in\[R^\{\(w\)\}\],

λr\(w\)−λr\+1\(w\)≥δ,\\lambda\_\{r\}^\{\(w\)\}\-\\lambda\_\{r\+1\}^\{\(w\)\}\\geq\\delta,where we defineλd\+1\(w\)=−∞\\lambda\_\{d\+1\}^\{\(w\)\}=\-\\infty\.

The eigenpairs admit the following spectral characterization\.

###### Proposition 3\.4\(Spectral characterization of word\-local modes\)\.

For everyr∈\[d\]r\\in\[d\],

λr\(w\)=∑t=1T−1\(dmean,r\(w\)​\(t,t\+1\)\)2,\\lambda\_\{r\}^\{\(w\)\}=\\sum\_\{t=1\}^\{T\-1\}\\left\(d\_\{\\mathrm\{mean\},r\}^\{\(w\)\}\(t,t\+1\)\\right\)^\{2\},\(11\)and

𝒖r\(w\)∈arg⁡max⁡∑t=1T−1‖𝒖‖=1𝒖⟂𝒖1\(w\),…,𝒖r−1\(w\)⁡𝒖⊤​𝐌mean\(w\)​\(t,t\+1\)​𝒖\.\\bm\{u\}\_\{r\}^\{\(w\)\}\\in\\arg\\max\_\{\\begin\{subarray\}\{c\}\\\|\\bm\{u\}\\\|=1\\\\ \\bm\{u\}\\perp\\bm\{u\}\_\{1\}^\{\(w\)\},\\ldots,\\bm\{u\}\_\{r\-1\}^\{\(w\)\}\\end\{subarray\}\}\\sum\_\{t=1\}^\{T\-1\}\\bm\{u\}^\{\\top\}\\mathbf\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,t\+1\)\\bm\{u\}\.\(12\)

The path sum identifies the directions that organize the word’s complete history, while the adjacent distances locate each direction in time\. Because the basis is word local, its modes describe distinct developments within a word rather than vocabulary wide axes\. We use the mean basis to isolate center movement while keeping the full, mean, and deviation contributions exactly comparable\.

### 3\.4Gaussian implementation and statistical recovery

For a tractable estimator, we specialize to Gaussian usage components, quadratic cost, and unrestricted component couplings using the established Gaussian and Gaussian\-mixture Wasserstein geometry\[[17](https://arxiv.org/html/2609.30974#bib.bib8),[18](https://arxiv.org/html/2609.30974#bib.bib10)\]:

νt,k\(w\)=𝒩⁡\(μt,k\(w\),Σt,k\(w\)\),𝒜t,k,ℓ\(w\)=Π⁡\(νt,k\(w\),νt\+1,ℓ\(w\)\),c⁡\(x,y\)=‖x−y‖2\.\\nu\_\{t,k\}^\{\(w\)\}=\\mathcal\{N\}\(\\mu\_\{t,k\}^\{\(w\)\},\\Sigma\_\{t,k\}^\{\(w\)\}\),\\qquad\\mathcal\{A\}\_\{t,k,\\ell\}^\{\(w\)\}=\\Pi\(\\nu\_\{t,k\}^\{\(w\)\},\\nu\_\{t\+1,\\ell\}^\{\(w\)\}\),\\qquad c\(x,y\)=\\\|x\-y\\\|^\{2\}\.
For completeness, the resulting component\-level optimal coupling has the standard Gaussian form\.

###### Proposition 3\.5\(Gaussian optimal coupling\)\.

IfΣt,k\(w\)≻0\\Sigma\_\{t,k\}^\{\(w\)\}\\succ 0, the optimizer in \([1](https://arxiv.org/html/2609.30974#S2.E1)\) is unique and equals

ηt,k,ℓ\(w\)=\(id,Tt,k,ℓ\(w\)\)\#​νt,k\(w\),Tt,k,ℓ\(w\)​\(x\)=μt\+1,ℓ\(w\)\+At,k,ℓ\(w\)​\(x−μt,k\(w\)\),\\eta\_\{t,k,\\ell\}^\{\(w\)\}=\(\\mathrm\{id\},T\_\{t,k,\\ell\}^\{\(w\)\}\)\_\{\\\#\}\\nu\_\{t,k\}^\{\(w\)\},\\qquad T\_\{t,k,\\ell\}^\{\(w\)\}\(x\)=\\mu\_\{t\+1,\\ell\}^\{\(w\)\}\+A\_\{t,k,\\ell\}^\{\(w\)\}\(x\-\\mu\_\{t,k\}^\{\(w\)\}\),\(13\)where

At,k,ℓ\(w\)=\(Σt,k\(w\)\)−1/2\[\(Σt,k\(w\)\)1/2Σt\+1,ℓ\(w\)\(Σt,k\(w\)\)1/2\]1/2\(Σt,k\(w\)\)−1/2\.A\_\{t,k,\\ell\}^\{\(w\)\}=\\left\(\\Sigma\_\{t,k\}^\{\(w\)\}\\right\)^\{\-1/2\}\\Bigl\[\\left\(\\Sigma\_\{t,k\}^\{\(w\)\}\\right\)^\{1/2\}\\Sigma\_\{t\+1,\\ell\}^\{\(w\)\}\\left\(\\Sigma\_\{t,k\}^\{\(w\)\}\\right\)^\{1/2\}\\Bigr\]^\{1/2\}\\left\(\\Sigma\_\{t,k\}^\{\(w\)\}\\right\)^\{\-1/2\}\.

Under this specialization, the component transition terms and induced distances admit closed form expressions and can be computed by dynamic programming for every period pair\(t,s\)\(t,s\)\. Details are deferred to Appendices[F](https://arxiv.org/html/2609.30974#A6)and[G](https://arxiv.org/html/2609.30974#A7)\.

Finally, we have the statistical recovery results\. Let hats denote estimated quantities\.

###### Assumption 3\.6\(Estimation and optimization conditions\)\.

The Gaussian mixtures have known, correctly specified component counts, positive weights, positive definite covariances and pairwise distinct component parameters \(see Definition[B\.2](https://arxiv.org/html/2609.30974#A2.Thmtheorem2)\)\. Their parameters are estimated by a penalized maximum likelihood estimator satisfying the conditions of[Chen and Tan \[21\]](https://arxiv.org/html/2609.30974#bib.bib7)\. Each adjacent optimal component plan is unique and has support sizeKt\(w\)\+Kt\+1\(w\)−1K\_\{t\}^\{\(w\)\}\+K\_\{t\+1\}^\{\(w\)\}\-1\.

###### Theorem 3\.7\(Statistical recovery\)\.

Under the Gaussian specialization and Assumptions[3\.3](https://arxiv.org/html/2609.30974#S3.Thmtheorem3)and[3\.6](https://arxiv.org/html/2609.30974#S3.Thmtheorem6), the following holds using the word\-local basis induced byℳmean\(w\)\\mathcal\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\. Asn\(w\):=mint⁡nt\(w\)→∞n^\{\(w\)\}:=\\min\_\{t\}n\_\{t\}^\{\(w\)\}\\to\\infty, for∙∈\{full,mean,dev\}\\bullet\\in\\\{\\mathrm\{full\},\\mathrm\{mean\},\\mathrm\{dev\}\\\}andρ∈\{tr,1,…,R\(w\)\}\\rho\\in\\\{\\mathrm\{tr\},1,\\ldots,R^\{\(w\)\}\\\},

maxt,s‖𝐌^∙\(w\)\(t,s\)−𝐌∙\(w\)\(t,s\)‖2=Op\(\(n\(w\)\)−1/2\),\\max\_\{t,s\}\\left\\\|\\widehat\{\\mathbf\{M\}\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)\-\\mathbf\{M\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)\\right\\\|\_\{2\}=O\_\{p\}\\\!\\left\(\(n^\{\(w\)\}\)^\{\-1/2\}\\right\),and

maxt,s\|d^∙,ρ\(w\)\(t,s\)2−d∙,ρ\(w\)\(t,s\)2\|=Op\(\(n\(w\)\)−1/2\)\.\\max\_\{t,s\}\\left\|\\widehat\{d\}\_\{\\bullet,\\rho\}^\{\(w\)\}\(t,s\)^\{2\}\-d\_\{\\bullet,\\rho\}^\{\(w\)\}\(t,s\)^\{2\}\\right\|=O\_\{p\}\\\!\\left\(\(n^\{\(w\)\}\)^\{\-1/2\}\\right\)\.

Thus, the operators, squared trace and mode distances used below stabilize jointly as the period sample sizes increase\. The guarantee covers the CUSP quantities derived from the fitted process, as well as the underlying Gaussian mixture parameters\.

## 4Experiments

The experiments follow Figure[1](https://arxiv.org/html/2609.30974#S1.F1): synthetic data test operator and distance recovery, DWUG\[[22](https://arxiv.org/html/2609.30974#bib.bib33)\]tests scalar validity and decomposition, Janus\[[23](https://arxiv.org/html/2609.30974#bib.bib6)\]tests multi\-period recovery and coherence, and Court opinions test passage\-grounded attribution\. The synthetic experiment uses Gaussian samples\. The lexical experiments use target\-token XL\-LEXEME embeddings\[[24](https://arxiv.org/html/2609.30974#bib.bib5)\]and Gaussian CUSP, with component couplings given by Proposition[3\.5](https://arxiv.org/html/2609.30974#S3.Thmtheorem5)\. The dominant directions shared across contextual embeddings can distort the distances\[[25](https://arxiv.org/html/2609.30974#bib.bib35)\]\. For each dataset, we therefore center its pooled usage panel, remove the leading principal direction \(k=1k=1\), and applyℓ2\\ell\_\{2\}\-normalization\[[26](https://arxiv.org/html/2609.30974#bib.bib24)\]\. The transform is fixed across words and periods within each dataset\. DWUG uses separate English and German panels, Janus uses its original released pool before schedule construction, and Court uses its balanced analysis panel\.

### 4\.1Synthetic mixture processes: finite\-sample recovery

We test Theorem[3\.7](https://arxiv.org/html/2609.30974#S3.Thmtheorem7)on a three\-period Gaussian process ind=2d=2withK=2K=2, where the component prevalences, centers, and covariances all change\. For5050repetitions atn∈\{100,200,400,800,1600\}n\\in\\\{100,200,400,800,1600\\\}usages per period, we fit independent full\-covariance Gaussian mixture models \(GMMs\) withKKfixed and withhold the planted component labels\.

Figure 2:Finite\-sample recovery of \(a\) component parameters, \(b\) full, mean, and deviation operators, and \(c\) squared full\-trace and word\-local mode distances\. Curves show means and95%95\\%confidence intervals over5050repetitions\. Gray dashed lines show then−1/2n^\{\-1/2\}rate\.Figure[2](https://arxiv.org/html/2609.30974#S4.F2)shows that the parameter errors decrease with increasing sample size\. The full operator and squared full\-trace errors both have a log–log slope of−0\.532\-0\.532, while the two squared mode distance errors have slopes of−0\.540\-0\.540and−0\.561\-0\.561\. Their confidence intervals contain−1/2\-1/2, agreeing with the rate predicted by Theorem[3\.7](https://arxiv.org/html/2609.30974#S3.Thmtheorem7)\. Appendix[I\.1](https://arxiv.org/html/2609.30974#A9.SS1)reports the mean and deviation distances and pure\-mechanism checks\.

### 4\.2DWUG: scalar validity and decomposition

The English and German DWUG datasets contain4646and5050two\-period targets with human graded\-change scores\[[22](https://arxiv.org/html/2609.30974#bib.bib33)\]\. We compare the CUSP full trace with APD, prototype distance, and AP\-JSD\[[3](https://arxiv.org/html/2609.30974#bib.bib14),[4](https://arxiv.org/html/2609.30974#bib.bib27)\], balanced usage\-level OT\[[5](https://arxiv.org/html/2609.30974#bib.bib23)\], and fSUS\[[12](https://arxiv.org/html/2609.30974#bib.bib21)\]\. We also evaluate the exact mean–deviation split \(Theorem[3\.1](https://arxiv.org/html/2609.30974#S3.Thmtheorem1)\)\. Tables[1](https://arxiv.org/html/2609.30974#S4.T1)and[6](https://arxiv.org/html/2609.30974#A9.T6)report the full/mean and complete results, respectively\.

Table 1:Spearman correlation with DWUG judgments under pooled panelk=1k=1\. Brackets give95%95\\%paired word\-bootstrap intervals\.
Figure 3:Janus atk=1k=1\. \(a\) Planted and recovered profiles\. Bands span the2\.52\.5–97\.597\.5percentiles across1616lemmas per schedule\. \(b\) Meanℓ1\\ell\_\{1\}composition error\. CUSP is consistent by construction\.
Table[1](https://arxiv.org/html/2609.30974#S4.T1)places CUSP first in English and within0\.0120\.012of the strongest German result\. The overlapping intervals support competitive scalar validity rather than uniform superiority\. Sensitivity atk=0k=0andk=3k=3preserves the same conclusion, as reported in Appendix[I\.2](https://arxiv.org/html/2609.30974#A9.SS2)\. Center movement supplies median shares of0\.8680\.868and0\.9060\.906of the full\-trace variation in English and German\. The deviation trace remains correlated with human judgments at0\.6810\.681and0\.6740\.674\. The deviation distance is the second largest in each dataset for English*bar*\(0\.3230\.323\) and German*Rezeption*\(0\.3260\.326\), both split across multiple DWUG usage clusters\. The decomposition therefore preserves a useful within\-component reorganization signal even though center displacement drives most graded change\.

### 4\.3Janus: controlled multi\-period recovery

Beyond two\-period DWUG, Janus\[[23](https://arxiv.org/html/2609.30974#bib.bib6)\]provides controlled six\-period histories:4848lemmas, each with four ten\-sentence pools resampled into4040usages per period\. Stable, gradual, and abrupt prevalence schedules each receive1616lemmas\. Released sense tags define schedules but are withheld during fitting\. CUSP receives no cross\-period usage links\. A BIC\-selected diagonal GMM is fitted independently to every lemma and period\. Figure[3](https://arxiv.org/html/2609.30974#S4.F3)\(a\) shows that CUSP achieves exact recovery of the stable profile, a TV error of0\.0440\.044for the gradual profile, and the correct boundary for all1616abrupt lemmas\. The recovered and planted adjacent changes correlate atρ=0\.954\\rho=0\.954\. The path length separates the changed and the stable schedules with an AUROC value of1\.0001\.000\. Pairwise OT fits balanced usage\-level transport independently for each period pair\. Its direct and composed plans have meanℓ1\\ell\_\{1\}disagreement0\.0840\.084, concentrated in gradual histories \(Figure[3](https://arxiv.org/html/2609.30974#S4.F3)\(b\)\)\. CUSP composes non\-adjacent plans by construction, giving numerical zero disagreement while preserving endpoint marginals \(Proposition[2\.2](https://arxiv.org/html/2609.30974#S2.Thmtheorem2)\)\. Appendix[I\.3](https://arxiv.org/html/2609.30974#A9.SS3)gives the protocol and schedule\-specific results\.

### 4\.4Court opinions: passage\-grounded attribution

Court opinions provide a natural, unlabeled setting:258,480258\{,\}480balanced usages of6060legal lemmas from100,473100\{,\}473opinions in eight decade bins \(19501950–20202020\), drawn from the May 6, 2024 CourtListener release\[[27](https://arxiv.org/html/2609.30974#bib.bib12)\]\. The vocabulary is fixed before fitting\. A fixed LLM procedure retains legal\-reasoning usages and extracts verbatim passages\. Each excerpt remains traceable through its CourtListener opinion ID\. Counts are equalized across decades within each lemma, and BIC selects period\-local diagonal GMMs withK≤5K\\leq 5\. Appendix[I\.4](https://arxiv.org/html/2609.30974#A9.SS4)details construction and fitting\.

Without gold change points or sense labels, we select transitions numerically before passage inspection\. For each lemma’s seven adjacent full\-trace distancesdtd\_\{t\}, we computezt=\(dt−d¯\)/sd⁡\(d\)z\_\{t\}=\(d\_\{t\}\-\\bar\{d\}\)/\\operatorname\{sd\}\(d\)and select the maximum\. The descriptive thresholdmaxt⁡zt≥1\.2\\max\_\{t\}z\_\{t\}\\geq 1\.2retains5353lemmas\. At the selected transition, Theorems[3\.1](https://arxiv.org/html/2609.30974#S3.Thmtheorem1)and[3\.2](https://arxiv.org/html/2609.30974#S3.Thmtheorem2)give

ak​ℓ=qk​ℓ​tr⁡\(𝐂mean​\(k,ℓ\)\+𝐂dev​\(k,ℓ\)\),dt2=∑k,ℓak​ℓ\.a\_\{k\\ell\}=q\_\{k\\ell\}\\operatorname\{tr\}\\\!\\left\(\\mathbf\{C\}\_\{\\mathrm\{mean\}\}\(k,\\ell\)\+\\mathbf\{C\}\_\{\\mathrm\{dev\}\}\(k,\\ell\)\\right\),\\qquad d\_\{t\}^\{2\}=\\sum\_\{k,\\ell\}a\_\{k\\ell\}\.Withαk​ℓ=ak​ℓ/dt2\\alpha\_\{k\\ell\}=a\_\{k\\ell\}/d\_\{t\}^\{2\}, the ratioαk​ℓ/qk​ℓ\\alpha\_\{k\\ell\}/q\_\{k\\ell\}compares the share of change in the pair with its share of transported mass\. We reportEmaxE\_\{\\max\}, the largest such ratio at the selected transition, to identify a movement that is unusually consequential for its size\. Component labels are local to each period, so their numerical indices have no semantic meaning\. Proposition[3\.4](https://arxiv.org/html/2609.30974#S3.Thmtheorem4)supplies the complementary temporal analysis\. Its word\-local mean modes distinguish directions that peak together from directions active in different decades\. Appendix[I\.5](https://arxiv.org/html/2609.30974#A9.SS5)presents the fixed thresholds, the top ten transition and temporal\-separation rankings, and the passage analyses\.

Figure 4:Court localization and word\-local modes\. \(a\) Standardized adjacent change profiles for the three highest\-enrichment lemmas and*expectation*\. Diamonds mark maxima and the line isz=1\.2z=1\.2\. \(b\)*Privacy*modes peak in different decades\. \(c\) The three leading*expectation*modes peak together\. Legends give the retained mean\-path energy shares\. Markers and line styles distinguish series\.After numerical selection, usages are assigned by maximum posterior probability\. A fixed retrieval score favors proximity to component centers and clear target\-lemma matches \(Appendix[I\.5](https://arxiv.org/html/2609.30974#A9.SS5)\)\. Source and target passages are representative endpoints, not observed sentence pairs\. Figure[4](https://arxiv.org/html/2609.30974#S4.F4)\(a\) shows*privacy*,*search*, and*standing*, the three lemmas with the highestEmaxE\_\{\\max\}scores, alongside*expectation*, whose modes peak together\. Cited decisions situate the attributed usage movements in their doctrinal context\.

Privacy\.*Privacy*ranks first in both Court rankings\. Its standardized profile peaks at1960→19701960\{\\rightarrow\}1970\(z=1\.43z=1\.43\), where center displacement supplies93\.3%93\.3\\%of the variation\. The selected pair carries2\.46%2\.46\\%of the transported mass, but37\.1%37\.1\\%of the contribution \(Emax=15\.10E\_\{\\max\}=15\.10\)\. Its source invokes Prosser’s “four categories” \(opinion1179312\) and its target says that warrantless inspection poses “only a minimal threat to justifiable expectations of privacy” \(opinion1602980\)\. At the preceding step, the leading pair moves from privacy as “a direct wrong of a personal character” \(opinion2604478\) toward the same Prosser component\. The successive pairs place Prosser’s tort classification\[[28](https://arxiv.org/html/2609.30974#bib.bib29)\]between an earlier privacy tort rationale and post\-*Katz*expectation\-of\-privacy language\[[29](https://arxiv.org/html/2609.30974#bib.bib20)\]\. The modes of this lemma separately distinguish tort, constitutional, informational, search, and family privacy across several decades, consistent with legal scholarship treating privacy as context\-dependent\[[30](https://arxiv.org/html/2609.30974#bib.bib34)\]\. The underlying passages and case anchors appear in Appendix[I\.5](https://arxiv.org/html/2609.30974#A9.SS5)\.

Expectation\.Its standardized profile peaks at1960→19701960\{\\rightarrow\}1970\(z=2\.20z=2\.20\), where center displacement supplies86\.8%86\.8\\%of the variation\. Sources concern “an expectation of pay or compensation” and contributions made “without expectation of services or direct benefits” \(opinions2226138,1636296\)\. Targets ask whether the accused showed “a reasonable expectation of privacy infringed by the search and seizure” and describe an area “protected by an expectation of privacy” \(opinions1365600,1311660\)\. The transition moves from ordinary anticipation toward the post\-*Katz*Fourth Amendment standard\[[29](https://arxiv.org/html/2609.30974#bib.bib20),[31](https://arxiv.org/html/2609.30974#bib.bib30)\]\. The three leading modes of this lemma, carrying48%48\\%,26%26\\%, and18%18\\%of retained mean\-path energy, all peak at this transition\. Unlike*privacy*, the change is synchronized rather than staggered\. The Court dataset thus demonstrates the complete empirical chain from independently sampled period distributions to a coupled process, exact contributions, temporal modes, and passages that can be checked independently\.

## 5Conclusion

Diachronic corpora provide period specific usage distributions without observed trajectories\. CUSP turns them into one marginal preserving process whose adjacent couplings determine longer span correspondence\. Its central contribution is a common process yielding magnitude and timing, exact accounting by mechanism and transported component pair, word\-local modes, and representative passages\. Gaussian specialization makes the operators and squared distances tractable and recoverable\. Synthetic, DWUG, Janus, and Court experiments support recovery, scalar validity, multi\-period coherence, and passage\-grounded attribution\. CUSP thus makes scalar measurement, coherent temporal structure, and inspectable explanation compatible views of one lexical history\.

## 6Acknowledgements

We thank Ryoma Kondo, Hiroaki Yamada, Tilmann Altwicker and Zhivko Taushanov for helpful discussions\. R\.H\. is supported by JST FOREST Program \(JPMJFR216Q\), JST PRESTO Program \(JPMJPR2469\), Grant\-in\-Aid for Scientific Research \(KAKENHI, JP24K03043\), JSPS International Joint Research Program \(JRPs with SNSF: 20251501\), and the UTEC\-UTokyo FSI Research Grant Program\.

## References

- \[1\]\(2016\)Diachronic word embeddings reveal statistical laws of semantic change\.InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Berlin, Germany,pp\. 1489–1501\.External Links:[Document](https://dx.doi.org/10.18653/v1/P16-1141),[Link](https://aclanthology.org/P16-1141/)Cited by:[§1](https://arxiv.org/html/2609.30974#S1.p2.1)\.
- \[2\]R\. Bamler and S\. Mandt\(2017\)Dynamic word embeddings\.InProceedings of the 34th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.70,pp\. 380–389\.External Links:[Link](https://proceedings.mlr.press/v70/bamler17a.html)Cited by:[§1](https://arxiv.org/html/2609.30974#S1.p2.1)\.
- \[3\]M\. Giulianelli, M\. Del Tredici, and R\. Fernández\(2020\)Analysing lexical semantic change with contextualised word representations\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,Online,pp\. 3960–3973\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.365),[Link](https://aclanthology.org/2020.acl-main.365/)Cited by:[§1](https://arxiv.org/html/2609.30974#S1.p2.1),[§4\.2](https://arxiv.org/html/2609.30974#S4.SS2.p1.1)\.
- \[4\]F\. Periti and N\. Tahmasebi\(2024\)A systematic comparison of contextualized word embeddings for lexical semantic change\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 4262–4282\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.240),[Link](https://aclanthology.org/2024.naacl-long.240/)Cited by:[§1](https://arxiv.org/html/2609.30974#S1.p2.1),[§4\.2](https://arxiv.org/html/2609.30974#S4.SS2.p1.1)\.
- \[5\]S\. Montariol and A\. Allauzen\(2021\)Transport optimal pour le changement sémantique à partir de plongements contextualisés \(optimal transport for semantic change detection using contextualised embeddings\)\.InActes de la 28e Conférence sur le Traitement Automatique des Langues Naturelles\. Volume 1 : conférence principale,Lille, France,pp\. 81–90\.External Links:[Link](https://aclanthology.org/2021.jeptalnrecital-taln.7/)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.7.1.1.1),[§I\.3](https://arxiv.org/html/2609.30974#A9.SS3.p4.1),[§1](https://arxiv.org/html/2609.30974#S1.p2.1),[§4\.2](https://arxiv.org/html/2609.30974#S4.SS2.p1.1)\.
- \[6\]L\. Frermann and M\. Lapata\(2016\)A bayesian model of diachronic meaning change\.Transactions of the Association for Computational Linguistics4,pp\. 31–45\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00081),[Link](https://aclanthology.org/Q16-1003/)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.3.1.1.1),[§1](https://arxiv.org/html/2609.30974#S1.p2.1)\.
- \[7\]R\. Hu, S\. Li, and S\. Liang\(2019\)Diachronic sense modeling with deep contextualized word embeddings: an ecological view\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,Florence, Italy,pp\. 3899–3908\.External Links:[Document](https://dx.doi.org/10.18653/v1/P19-1379),[Link](https://aclanthology.org/P19-1379/)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.3.1.1.1),[§1](https://arxiv.org/html/2609.30974#S1.p2.1)\.
- \[8\]F\. Periti, S\. Picascia, S\. Montanelli, A\. Ferrara, and N\. Tahmasebi\(2025\)Studying word meaning evolution through incremental semantic shift detection\.Language Resources and Evaluation59,pp\. 1363–1399\.External Links:[Document](https://dx.doi.org/10.1007/s10579-024-09769-1),[Link](https://link.springer.com/article/10.1007/s10579-024-09769-1)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.4.1.1.1),[§1](https://arxiv.org/html/2609.30974#S1.p2.1)\.
- \[9\]I\. Kolli, K\. Lange, J\. Rieger, and C\. Jentsch\(2026\)Word\-centered semantic graphs for interpretable diachronic sense tracking\.External Links:2601\.22410,[Link](https://arxiv.org/abs/2601.22410)Cited by:[§1](https://arxiv.org/html/2609.30974#S1.p2.1)\.
- \[10\]D\. V\. Kokosinskii, D\. Schlechtweg, and N\. V\. Arefyev\(2026\)Selecting and controlling sense granularity for lexical semantic change detection\.Program Systems: Theory and Applications17\(2\),pp\. 103–146\.External Links:[Document](https://dx.doi.org/10.25209/2079-3316-2026-17-2-103-146),[Link](https://psta.psiras.ru/en/2026/2_103-146)Cited by:[§1](https://arxiv.org/html/2609.30974#S1.p2.1)\.
- \[11\]H\. Kiyama, T\. Aida, M\. Komachi, T\. Ogiso, H\. Takamura, and D\. Mochihashi\(2025\)Analyzing continuous semantic shifts with diachronic word similarity matrices\.InProceedings of the 31st International Conference on Computational Linguistics,Abu Dhabi, UAE,pp\. 1613–1631\.External Links:[Link](https://aclanthology.org/2025.coling-main.109/)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.5.1.1.1),[§1](https://arxiv.org/html/2609.30974#S1.p2.1)\.
- \[12\]R\. Kishino, H\. Yamagiwa, R\. Nagata, S\. Yokoi, and H\. Shimodaira\(2025\)Quantifying lexical semantic shift via unbalanced optimal transport\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Vienna, Austria,pp\. 15913–15933\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.774),[Link](https://aclanthology.org/2025.acl-long.774/)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.9.1.1.1),[§I\.3](https://arxiv.org/html/2609.30974#A9.SS3.p4.1),[§1](https://arxiv.org/html/2609.30974#S1.p2.1),[§4\.2](https://arxiv.org/html/2609.30974#S4.SS2.p1.1)\.
- \[13\]T\. Aida and D\. Bollegala\(2025\)Investigating the contextualised word embedding dimensions specified for contextual and temporal semantic changes\.InProceedings of the 31st International Conference on Computational Linguistics,Abu Dhabi, UAE,pp\. 1413–1437\.External Links:[Link](https://aclanthology.org/2025.coling-main.95/)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.10.1.1.1),[§1](https://arxiv.org/html/2609.30974#S1.p2.1)\.
- \[14\]T\. Aida and D\. Bollegala\(2025\)SCDTour: embedding axis ordering and merging for interpretable semantic change detection\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 14775–14785\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.797),[Link](https://aclanthology.org/2025.findings-emnlp.797/)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.10.1.1.1),[§1](https://arxiv.org/html/2609.30974#S1.p2.1)\.
- \[15\]J\. Chen, E\. Chersoni, D\. Schlechtweg, and C\. Huang\(2026\)Lexical semantic change detection: a survey of tasks, benchmarks, models, and potential impacts in digital humanities and social sciences\.Natural Language Processing32\(4\),pp\. 391–429\.External Links:[Document](https://dx.doi.org/10.1017/nlp.2026.10032)Cited by:[§1](https://arxiv.org/html/2609.30974#S1.p2.1)\.
- \[16\]H\. Ezoe and R\. Hisano\(2026\)Multiscale euclidean network trajectories: second\-moment geometry, attribution, and change points\.arXiv preprint arXiv:2605\.04589\.External Links:2605\.04589,[Link](https://arxiv.org/abs/2605.04589)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.11.1.1.1),[§E\.2](https://arxiv.org/html/2609.30974#A5.SS2.p1.1),[Appendix H](https://arxiv.org/html/2609.30974#A8.SS0.SSS0.Px6.p1.1),[Appendix H](https://arxiv.org/html/2609.30974#A8.SS0.SSS0.Px6.p3.1.1),[Appendix H](https://arxiv.org/html/2609.30974#A8.SS0.SSS0.Px6.p4.1.1),[Appendix H](https://arxiv.org/html/2609.30974#A8.SS0.SSS0.Px6.p5.1.1),[§1](https://arxiv.org/html/2609.30974#S1.p4.1),[§3\.1](https://arxiv.org/html/2609.30974#S3.SS1.p3.1)\.
- \[17\]Y\. Chen, T\. T\. Georgiou, and A\. Tannenbaum\(2019\)Optimal transport for gaussian mixture models\.IEEE Access7,pp\. 6269–6278\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2018.2889838)Cited by:[§2\.2](https://arxiv.org/html/2609.30974#S2.SS2.p1.2),[§3\.4](https://arxiv.org/html/2609.30974#S3.SS4.p1.1)\.
- \[18\]J\. Delon and A\. Desolneux\(2020\)A wasserstein\-type distance in the space of gaussian mixture models\.SIAM Journal on Imaging Sciences13\(2\),pp\. 936–970\.External Links:[Document](https://dx.doi.org/10.1137/19M1301047)Cited by:[§2\.2](https://arxiv.org/html/2609.30974#S2.SS2.p1.2),[§3\.4](https://arxiv.org/html/2609.30974#S3.SS4.p1.1)\.
- \[19\]M\. Yurochkin, S\. Claici, E\. Chien, F\. Mirzazadeh, and J\. M\. Solomon\(2019\)Hierarchical optimal transport for document representation\.InAdvances in Neural Information Processing Systems,Vol\.32,pp\. 1601–1611\.External Links:[Link](https://proceedings.neurips.cc/paper/2019/hash/8b5040a8a5baf3e0e67386c2e3a9b903-Abstract.html)Cited by:[§2\.2](https://arxiv.org/html/2609.30974#S2.SS2.p1.2)\.
- \[20\]G\. Peyré and M\. Cuturi\(2019\)Computational optimal transport\.Foundations and Trends in Machine Learning11\(5–6\),pp\. 355–607\.External Links:[Document](https://dx.doi.org/10.1561/2200000073)Cited by:[§2\.2](https://arxiv.org/html/2609.30974#S2.SS2.p2.1)\.
- \[21\]J\. Chen and X\. Tan\(2009\)Inference for multivariate normal mixtures\.Journal of Multivariate Analysis100\(7\),pp\. 1367–1383\.External Links:[Document](https://dx.doi.org/10.1016/j.jmva.2008.12.005)Cited by:[Appendix H](https://arxiv.org/html/2609.30974#A8.SS0.SSS0.Px2.p3.1),[Appendix H](https://arxiv.org/html/2609.30974#A8.SS0.SSS0.Px2.p4.1),[Lemma H\.1](https://arxiv.org/html/2609.30974#A8.Thmtheorem1.p1.1.1),[Assumption 3\.6](https://arxiv.org/html/2609.30974#S3.Thmtheorem6.p1.1)\.
- \[22\]D\. Schlechtweg, N\. Tahmasebi, S\. Hengchen, H\. Dubossarsky, and B\. McGillivray\(2021\)DWUG: a large resource of diachronic word usage graphs in four languages\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,Online and Punta Cana, Dominican Republic,pp\. 7079–7091\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.567),[Link](https://aclanthology.org/2021.emnlp-main.567/)Cited by:[§I\.2](https://arxiv.org/html/2609.30974#A9.SS2.p1.1),[§4\.2](https://arxiv.org/html/2609.30974#S4.SS2.p1.1),[§4](https://arxiv.org/html/2609.30974#S4.p1.1)\.
- \[23\]P\. Cassotti and N\. Tahmasebi\(2025\)Sense\-specific historical word usage generation\.Transactions of the Association for Computational Linguistics13,pp\. 690–708\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00761),[Link](https://aclanthology.org/2025.tacl-1.32/)Cited by:[§I\.3](https://arxiv.org/html/2609.30974#A9.SS3.p1.1),[§4\.3](https://arxiv.org/html/2609.30974#S4.SS3.p1.1),[§4](https://arxiv.org/html/2609.30974#S4.p1.1)\.
- \[24\]P\. Cassotti, L\. Siciliani, M\. DeGemmis, G\. Semeraro, and P\. Basile\(2023\)XL\-LEXEME: WiC pretrained model for cross\-lingual LEXical sEMantic changE\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),Toronto, Canada,pp\. 1577–1585\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-short.135),[Link](https://aclanthology.org/2023.acl-short.135/)Cited by:[§I\.2](https://arxiv.org/html/2609.30974#A9.SS2.p1.1),[§4](https://arxiv.org/html/2609.30974#S4.p1.1)\.
- \[25\]W\. Timkey and M\. van Schijndel\(2021\)All bark and no bite: rogue dimensions in transformer language models obscure representational quality\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,Online and Punta Cana, Dominican Republic,pp\. 4527–4546\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.372),[Link](https://aclanthology.org/2021.emnlp-main.372/)Cited by:[§4](https://arxiv.org/html/2609.30974#S4.p1.1)\.
- \[26\]J\. Mu and P\. Viswanath\(2018\)All\-but\-the\-top: simple and effective postprocessing for word representations\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=HkuGJ3kCb)Cited by:[§4](https://arxiv.org/html/2609.30974#S4.p1.1)\.
- \[27\]Free Law Project\(2024\)CourtListener bulk legal data\.Note:Opinion\-data snapshot dated 2024\-05\-06External Links:[Link](https://wiki.free.law/c/courtlistener/help/api/bulk-data/bulk-legal-data)Cited by:[§I\.4](https://arxiv.org/html/2609.30974#A9.SS4.p1.1),[§4\.4](https://arxiv.org/html/2609.30974#S4.SS4.p1.1)\.
- \[28\]W\. L\. Prosser\(1960\)Privacy\.California Law Review48\(3\),pp\. 383–423\.External Links:[Document](https://dx.doi.org/10.2307/3478805)Cited by:[§I\.5](https://arxiv.org/html/2609.30974#A9.SS5.p5.1),[§4\.4](https://arxiv.org/html/2609.30974#S4.SS4.p4.1)\.
- \[29\]Supreme Court of the United States\(1967\)Katz v\. United States, 389 u\.s\. 347\.Cited by:[§I\.5](https://arxiv.org/html/2609.30974#A9.SS5.p5.1),[§4\.4](https://arxiv.org/html/2609.30974#S4.SS4.p4.1),[§4\.4](https://arxiv.org/html/2609.30974#S4.SS4.p5.1)\.
- \[30\]D\. J\. Solove\(2006\)A taxonomy of privacy\.University of Pennsylvania Law Review154\(3\),pp\. 477–564\.External Links:[Document](https://dx.doi.org/10.2307/40041279),[Link](https://scholarship.law.gwu.edu/faculty_publications/921/)Cited by:[§4\.4](https://arxiv.org/html/2609.30974#S4.SS4.p4.1)\.
- \[31\]Supreme Court of the United States\(1978\)Rakas v\. Illinois, 439 u\.s\. 128\.Cited by:[§4\.4](https://arxiv.org/html/2609.30974#S4.SS4.p5.1)\.
- \[32\]F\. Periti, A\. Ferrara, S\. Montanelli, and M\. Ruskov\(2022\)What is done is done: an incremental approach to semantic shift detection\.InProceedings of the 3rd Workshop on Computational Approaches to Historical Language Change,Dublin, Ireland,pp\. 33–43\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.lchange-1.4),[Link](https://aclanthology.org/2022.lchange-1.4/)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.4.1.1.1)\.
- \[33\]B\. Phan\-Tat, K\. Heylen, D\. Geeraerts, S\. De Pascale, and D\. Speelman\(2026\)SynFlow: a multidimensional diachronic semantic analysis toolkit\.External Links:2608\.19472,[Link](https://arxiv.org/abs/2608.19472)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.6.1.1.1)\.
- \[34\]C\. Arrington, M\. Gruppi, and S\. Adali\(2025\)ConShift: sense\-based language variation analysis using flexible alignment\.InFindings of the Association for Computational Linguistics: NAACL 2025,Albuquerque, New Mexico,pp\. 167–181\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.9),[Link](https://aclanthology.org/2025.findings-naacl.9/)Cited by:[Table 3](https://arxiv.org/html/2609.30974#A1.T3.5.8.1.1.1)\.
- \[35\]R\. Okano and M\. Imaizumi\(2024\)Distribution\-on\-distribution regression with wasserstein metric: multivariate gaussian case\.Journal of Multivariate Analysis203,pp\. 105334\.External Links:[Document](https://dx.doi.org/10.1016/j.jmva.2024.105334)Cited by:[Appendix F](https://arxiv.org/html/2609.30974#A6.p2.1)\.
- \[36\]F\. Santambrogio\(2017\)Euclidean, Metric, and Wasserstein gradient flows: an overview\.Bulletin of Mathematical Sciences7\(1\),pp\. 87–154\.External Links:[Document](https://dx.doi.org/10.1007/s13373-017-0101-1)Cited by:[Appendix F](https://arxiv.org/html/2609.30974#A6.p2.1)\.
- \[37\]T\. Huang, H\. Peng, and K\. Zhang\(2017\)Model selection for gaussian mixture models\.Statistica Sinica27\(1\),pp\. 147–169\.External Links:[Document](https://dx.doi.org/10.5705/ss.2014.105)Cited by:[Appendix H](https://arxiv.org/html/2609.30974#A8.SS0.SSS0.Px1.p1.1)\.
- \[38\]Supreme Court of the United States\(1967\)Warden v\. Hayden, 387 u\.s\. 294\.Cited by:[§I\.5](https://arxiv.org/html/2609.30974#A9.SS5.p12.1)\.
- \[39\]Supreme Court of the United States\(1965\)Griswold v\. Connecticut, 381 u\.s\. 479\.Cited by:[§I\.5](https://arxiv.org/html/2609.30974#A9.SS5.p12.1)\.
- \[40\]Supreme Court of the United States\(1973\)Roe v\. Wade, 410 u\.s\. 113\.Cited by:[§I\.5](https://arxiv.org/html/2609.30974#A9.SS5.p12.1)\.

## Appendix AAppendix guide and notation

The appendix follows the construction in the main text\. Appendix[B](https://arxiv.org/html/2609.30974#A2)formalizes identifiability of the period specific mixture representation\. Appendix[C](https://arxiv.org/html/2609.30974#A3)proves marginal preservation and the basic properties of the coupled process\. Appendices[D](https://arxiv.org/html/2609.30974#A4)and[E](https://arxiv.org/html/2609.30974#A5)establish the exact mean–deviation and sense\-transition decompositions and the trace and mode\-wise geometry, respectively\. Appendix[F](https://arxiv.org/html/2609.30974#A6)gives explicit Gaussian formulas, and Appendix[G](https://arxiv.org/html/2609.30974#A7)shows how to compute non\-adjacent transition quantities without enumerating latent paths\. Appendix[H](https://arxiv.org/html/2609.30974#A8)proves the finite\-sample recovery result\. Appendix[I](https://arxiv.org/html/2609.30974#A9)then provides the complete experimental constructions, robustness checks, rankings, and passage analyses\.

Throughout the paper,wwdenotes a target word,t,st,sdenote periods,k,ℓk,\\elldenote period local components, andrrdenotes a word\-local mode\. Component numbers are local to each period and do not identify the same sense across time\. Their relation across periods is determined by the fitted coupling\. We use∙∈\{full,mean,dev\}\\bullet\\in\\\{\\mathrm\{full\},\\mathrm\{mean\},\\mathrm\{dev\}\\\}for statements that apply to any of the three processes andρ∈\{tr,1,…,d\}\\rho\\in\\\{\\mathrm\{tr\},1,\\ldots,d\\\}for either the trace or a mode\-wise readout\. Table[2](https://arxiv.org/html/2609.30974#A1.T2)collects the notation used throughout the paper and appendix\.

Table 2:Notation used in the CUSP construction, theoretical results, and experiments\.### A\.1Structural comparison with related approaches

Table[3](https://arxiv.org/html/2609.30974#A1.T3)compares the structures provided by closely related approaches\. A checkmark denotes a directly provided capability,△\\trianglea related but non\-equivalent object, and a dash a capability that is not directly provided\. The comparison concerns the joint set of capabilities rather than any single column\. Compositional coupling means that adjacent and non\-adjacent usage–component correspondences arise from one multi\-period process rather than independent fits\. An exact mechanism split and transition attribution decompose the same change quantity by mechanism and transported component pair, respectively\. Temporal modes are word\-local directions with period local activity profiles\. For SynFlow,△\\triangledenotes related filler cluster continuity and value contributions rather than transported usage–component accounting\.

Table 3:Structural comparison with closely related approaches\.Temporal representationAttribution from that representationApproachMulti\-periodUsage/componentcorrespondenceCompositionalcouplingExact mechanismsplitTransitionattributionTemporalmodesDynamic and evolving sense models\[[6](https://arxiv.org/html/2609.30974#bib.bib13),[7](https://arxiv.org/html/2609.30974#bib.bib18)\]✓✓–△\\triangle––WiDiD\[[32](https://arxiv.org/html/2609.30974#bib.bib26),[8](https://arxiv.org/html/2609.30974#bib.bib37)\]✓✓––△\\triangle–Multi\-period similarity analysis\[[11](https://arxiv.org/html/2609.30974#bib.bib22)\]✓–––––SynFlow\[[33](https://arxiv.org/html/2609.30974#bib.bib40)\]✓△\\triangle–△\\triangle△\\triangle–Pairwise OT\[[5](https://arxiv.org/html/2609.30974#bib.bib23)\]△\\triangle✓––––ConShift\[[34](https://arxiv.org/html/2609.30974#bib.bib3)\]△\\triangle✓–△\\triangle△\\triangle–UOT / SUS\[[12](https://arxiv.org/html/2609.30974#bib.bib21)\]△\\triangle✓––△\\triangle–Directional methods\[[13](https://arxiv.org/html/2609.30974#bib.bib1),[14](https://arxiv.org/html/2609.30974#bib.bib2)\]–––––△\\triangleMENT\[[16](https://arxiv.org/html/2609.30974#bib.bib11)\]✓––△\\triangle–✓CUSP \(ours\)✓✓✓✓✓✓

## Appendix BIdentifiability of the sense mixture representation

The mixture representation of a period specific word distribution is not unique in general, since different parameterizations may induce the same distribution\. The following assumption formalizes the identifiability condition stated in the main text: the mixture representation is uniquely determined up to permutation of the sense labels\.

###### Assumption B\.1\(Identifiability up to sense\-label permutation\)\.

Suppose that two mixture representations

∑k=1Kt\(w\)πt,k\(w\)​νt,k\(w\)and∑k=1K~t\(w\)π~t,k\(w\)​ν~t,k\(w\)\\sum\_\{k=1\}^\{K\_\{t\}^\{\(w\)\}\}\\pi\_\{t,k\}^\{\(w\)\}\\nu\_\{t,k\}^\{\(w\)\}\\qquad\\text\{and\}\\qquad\\sum\_\{k=1\}^\{\\tilde\{K\}\_\{t\}^\{\(w\)\}\}\\tilde\{\\pi\}\_\{t,k\}^\{\(w\)\}\\tilde\{\\nu\}\_\{t,k\}^\{\(w\)\}represent the same period specific word distribution\. Then

Kt\(w\)=K~t\(w\),K\_\{t\}^\{\(w\)\}=\\tilde\{K\}\_\{t\}^\{\(w\)\},and there exists a permutationσ\\sigmaof\{1,…,Kt\(w\)\}\\\{1,\\ldots,K\_\{t\}^\{\(w\)\}\\\}such that

π~t,k\(w\)=πt,σ⁡\(k\)\(w\),ν~t,k\(w\)=νt,σ⁡\(k\)\(w\),k=1,…,Kt\(w\)\.\\tilde\{\\pi\}\_\{t,k\}^\{\(w\)\}=\\pi\_\{t,\\sigma\(k\)\}^\{\(w\)\},\\qquad\\tilde\{\\nu\}\_\{t,k\}^\{\(w\)\}=\\nu\_\{t,\\sigma\(k\)\}^\{\(w\)\},\\qquad k=1,\\ldots,K\_\{t\}^\{\(w\)\}\.

Consequently, all label\-dependent quantities are defined up to the corresponding label permutation\. In the statistical recovery analysis, whenever population and estimated quantities are compared, we implicitly choose the permutation that aligns the corresponding sense labels\.

The Gaussian\-mixture specialization used in the statistical recovery analysis provides a concrete admissible class of mixture distributions and their parameterizations satisfying the identifiability condition above\. The restrictions defining this class correspond to the first part of Assumption[3\.6](https://arxiv.org/html/2609.30974#S3.Thmtheorem6)\.

###### Definition B\.2\(Admissible Gaussian mixture class\)\.

The admissible Gaussian mixture class consists of Gaussian mixture distributions

γ=∑k=1Kπk​νk,νk=𝒩⁡\(mk,Σk\),\\gamma=\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}\\nu\_\{k\},\\qquad\\nu\_\{k\}=\\mathcal\{N\}\(m\_\{k\},\\Sigma\_\{k\}\),together with parameterizations satisfying

- •πk\>0,∑k=1Kπk=1;\\pi\_\{k\}\>0,\\qquad\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}=1;
- •Σk≻0,k=1,…,K;\\Sigma\_\{k\}\\succ 0,\\qquad k=1,\\ldots,K;
- •the component parameters are pairwise distinct: \(mk,Σk\)≠\(mℓ,Σℓ\),k≠ℓ\.\(m\_\{k\},\\Sigma\_\{k\}\)\\neq\(m\_\{\\ell\},\\Sigma\_\{\\ell\}\),\\qquad k\\neq\\ell\.

## Appendix CProofs of basic properties of the coupled usage\-sense processes

We prove Proposition[2\.2](https://arxiv.org/html/2609.30974#S2.Thmtheorem2), which establishes preservation of the prescribed sense prevalences and contextual component distributions at each period, and consequently of the observed usage distributions as the period\-wise marginals of the resulting joint process\.

###### Proof\.

Fixw∈Vw\\in V, and suppress the superscript\(w\)\(w\)throughout the proof\. We first prove the preservation of the sense marginals by induction\. By construction,

Pr⁡\(Z1=k\)=π1,k\\Pr\(Z\_\{1\}=k\)=\\pi\_\{1,k\}for everyk∈\[K1\]k\\in\[K\_\{1\}\], so the claim holds att=1t=1\. IfPr⁡\(Zt=k\)=πt,k\\Pr\(Z\_\{t\}=k\)=\\pi\_\{t,k\}for everyk∈\[Kt\]k\\in\[K\_\{t\}\], then

Pr⁡\(Zt\+1=ℓ\)\\displaystyle\\Pr\(Z\_\{t\+1\}=\\ell\)=∑k∈\[Kt\]Pr⁡\(Zt=k\)​Pt​\(k,ℓ\)\\displaystyle=\\sum\_\{k\\in\[K\_\{t\}\]\}\\Pr\(Z\_\{t\}=k\)P\_\{t\}\(k,\\ell\)=∑k∈\[Kt\]πt,k​Qt​\(k,ℓ\)πt,k=∑k∈\[Kt\]Qt​\(k,ℓ\)=πt\+1,ℓ,\\displaystyle=\\sum\_\{k\\in\[K\_\{t\}\]\}\\pi\_\{t,k\}\\frac\{Q\_\{t\}\(k,\\ell\)\}\{\\pi\_\{t,k\}\}=\\sum\_\{k\\in\[K\_\{t\}\]\}Q\_\{t\}\(k,\\ell\)=\\pi\_\{t\+1,\\ell\},wherePt​\(k,ℓ\)=Qt​\(k,ℓ\)/πt,kP\_\{t\}\(k,\\ell\)=Q\_\{t\}\(k,\\ell\)/\\pi\_\{t,k\}is the transition probability, and the last equality follows from the second marginal constraint in[2](https://arxiv.org/html/2609.30974#S2.E2), namely∑k∈\[Kt\]Qt​\(k,ℓ\)=πt\+1,ℓ\\sum\_\{k\\in\[K\_\{t\}\]\}Q\_\{t\}\(k,\\ell\)=\\pi\_\{t\+1,\\ell\}\. Thus, by induction,

Pr⁡\(Zt=k\)=πt,k\\Pr\(Z\_\{t\}=k\)=\\pi\_\{t,k\}for allt∈\[T\]t\\in\[T\]andk∈\[Kt\]k\\in\[K\_\{t\}\]\.

We next prove the preservation of the conditional contextual distributions by induction\. Fix a sense pathz1:Tz\_\{1:T\}with positive probability\. By construction,

ℒ\(X1∣Z1:T=z1:T\)=ν1,z1,\\mathcal\{L\}\(X\_\{1\}\\mid Z\_\{1:T\}=z\_\{1:T\}\)=\\nu\_\{1,z\_\{1\}\},so the claim holds att=1t=1\. Suppose that

ℒ\(Xt∣Z1:T=z1:T\)=νt,zt\.\\mathcal\{L\}\(X\_\{t\}\\mid Z\_\{1:T\}=z\_\{1:T\}\)=\\nu\_\{t,z\_\{t\}\}\.Then, for any measurable setA⊆ℝdA\\subseteq\\mathbb\{R\}^\{d\},

Pr\(Xt\+1∈A∣Z1:T=z1:T\)\\displaystyle\\Pr\(X\_\{t\+1\}\\in A\\mid Z\_\{1:T\}=z\_\{1:T\}\)=∫𝒦t,zt,zt\+1​\(x,A\)​νt,zt​\(dx\)\\displaystyle=\\int\\mathcal\{K\}\_\{t,z\_\{t\},z\_\{t\+1\}\}\(x,A\)\\,\\nu\_\{t,z\_\{t\}\}\(dx\)=∫𝟏A​\(y\)​ηt,zt,zt\+1​\(dx,dy\)\\displaystyle=\\int\\mathbf\{1\}\_\{A\}\(y\)\\,\\eta\_\{t,z\_\{t\},z\_\{t\+1\}\}\(dx,dy\)=νt\+1,zt\+1​\(A\),\\displaystyle=\\nu\_\{t\+1,z\_\{t\+1\}\}\(A\),where we used the transition kernel defined by

ηt,zt,zt\+1​\(d​x,d​y\)=νt,zt​\(d​x\)​𝒦t,zt,zt\+1​\(x,d​y\),\\eta\_\{t,z\_\{t\},z\_\{t\+1\}\}\(dx,dy\)=\\nu\_\{t,z\_\{t\}\}\(dx\)\\mathcal\{K\}\_\{t,z\_\{t\},z\_\{t\+1\}\}\(x,dy\),and the last equality follows from the marginal constraint in[1](https://arxiv.org/html/2609.30974#S2.E1), namelyηt,zt,zt\+1∈Π⁡\(νt,zt,νt\+1,zt\+1\)\\eta\_\{t,z\_\{t\},z\_\{t\+1\}\}\\in\\Pi\(\\nu\_\{t,z\_\{t\}\},\\nu\_\{t\+1,z\_\{t\+1\}\}\)\. Thus, by induction,

ℒ\(Xt∣Z1:T=z1:T\)=νt,zt\\mathcal\{L\}\(X\_\{t\}\\mid Z\_\{1:T\}=z\_\{1:T\}\)=\\nu\_\{t,z\_\{t\}\}for allt∈\[T\]t\\in\[T\], establishing the desired preservation of the conditional contextual distributions at each period\.

Finally, using the sense marginal preservation and the conditional distribution established above, the marginal law ofXtX\_\{t\}is

ℒ⁡\(Xt\)\\displaystyle\\mathcal\{L\}\(X\_\{t\}\)=∑z1:TPr\(Z1:T=z1:T\)ℒ\(Xt∣Z1:T=z1:T\)\\displaystyle=\\sum\_\{z\_\{1:T\}\}\\Pr\(Z\_\{1:T\}=z\_\{1:T\}\)\\,\\mathcal\{L\}\(X\_\{t\}\\mid Z\_\{1:T\}=z\_\{1:T\}\)=∑k\(∑z1:T:zt=kPr\(Z1:T=z1:T\)\)νt,k\\displaystyle=\\sum\_\{k\}\\left\(\\sum\_\{z\_\{1:T\}:z\_\{t\}=k\}\\Pr\(Z\_\{1:T\}=z\_\{1:T\}\)\\right\)\\nu\_\{t,k\}=∑k∈\[Kt\]πt,k​νt,k=γt\.\\displaystyle=\\sum\_\{k\\in\[K\_\{t\}\]\}\\pi\_\{t,k\}\\nu\_\{t,k\}=\\gamma\_\{t\}\.Consequently,

ℒ⁡\(X1,…,XT\)∈Π⁡\(γ1,…,γT\),\\mathcal\{L\}\(X\_\{1\},\\ldots,X\_\{T\}\)\\in\\Pi\(\\gamma\_\{1\},\\ldots,\\gamma\_\{T\}\),which proves the proposition\. ∎

We prove the reconstruction and conditional centering properties stated in \([4](https://arxiv.org/html/2609.30974#S2.E4)\)\. The reconstruction identity follows directly from the definitions, while the conditional centering property follows from the preservation of the conditional marginals established in Proposition[2\.2](https://arxiv.org/html/2609.30974#S2.Thmtheorem2)\.

###### Proof\.

By definition,

ϕfull\(w\)​\(t\)=ϕmean\(w\)​\(t\)\+ϕdev\(w\)​\(t\),\\phi\_\{\\mathrm\{full\}\}^\{\(w\)\}\(t\)=\\phi\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t\)\+\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t\),so the reconstruction identity holds\.

Moreover, by Proposition[2\.2](https://arxiv.org/html/2609.30974#S2.Thmtheorem2), for any sense pathz1:Tz\_\{1:T\}with positive probability,

𝔼\[Xt\(w\)∣Z1:T\(w\)=z1:T\]=∫xdνt,zt\(w\)\(x\)=mt,zt\(w\)\.\\displaystyle\\mathbb\{E\}\\left\[X\_\{t\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\]=\\int x\\,d\\nu\_\{t,z\_\{t\}\}^\{\(w\)\}\(x\)=m\_\{t,z\_\{t\}\}^\{\(w\)\}\.Therefore,

𝔼\[ϕdev\(w\)\(t\)∣Z1:T\(w\)\]\\displaystyle\\mathbb\{E\}\\left\[\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t\)\\mid Z\_\{1:T\}^\{\(w\)\}\\right\]=𝔼\[Xt\(w\)−mt,Zt\(w\)\(w\)∣Z1:T\(w\)\]\\displaystyle=\\mathbb\{E\}\\left\[X\_\{t\}^\{\(w\)\}\-m\_\{t,Z\_\{t\}^\{\(w\)\}\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}\\right\]=mt,Zt\(w\)\(w\)−mt,Zt\(w\)\(w\)=0\.\\displaystyle=m\_\{t,Z\_\{t\}^\{\(w\)\}\}^\{\(w\)\}\-m\_\{t,Z\_\{t\}^\{\(w\)\}\}^\{\(w\)\}=0\.Thus, both identities in \([4](https://arxiv.org/html/2609.30974#S2.E4)\) hold\. ∎

## Appendix DProofs of exact mechanistic and sense\-transition attribution

We first prove Theorem[3\.1](https://arxiv.org/html/2609.30974#S3.Thmtheorem1)\. The key step is that the deviation process is conditionally centered given the full latent sense sequence, so the cross terms between the mean and deviation processes vanish\.

###### Proof\.

Fixw∈Vw\\in Vandt,s∈\[T\]t,s\\in\[T\]\. By the reconstruction identity in \([4](https://arxiv.org/html/2609.30974#S2.E4)\),

Δ​ϕfull\(w\)​\(t,s\)\\displaystyle\\Delta\\phi\_\{\\mathrm\{full\}\}^\{\(w\)\}\(t,s\)=Δ​ϕmean\(w\)​\(t,s\)\+Δ​ϕdev\(w\)​\(t,s\)\.\\displaystyle=\\Delta\\phi\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s\)\+\\Delta\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)\.Therefore, expanding the second\-moment operator gives

𝐌full\(w\)​\(t,s\)\\displaystyle\\mathbf\{M\}\_\{\\mathrm\{full\}\}^\{\(w\)\}\(t,s\)=𝐌mean\(w\)​\(t,s\)\+𝐌dev\(w\)​\(t,s\)\\displaystyle=\\mathbf\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s\)\+\\mathbf\{M\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)\+𝔼⁡\[Δ​ϕmean\(w\)​\(t,s\)​Δ​ϕdev\(w\)​\(t,s\)⊤\]\\displaystyle\+\\mathbb\{E\}\\left\[\\Delta\\phi\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s\)\\Delta\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)^\{\\top\}\\right\]\+𝔼⁡\[Δ​ϕdev\(w\)​\(t,s\)​Δ​ϕmean\(w\)​\(t,s\)⊤\]\.\\displaystyle\+\\mathbb\{E\}\\left\[\\Delta\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)\\Delta\\phi\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s\)^\{\\top\}\\right\]\.We first consider the first cross term\.

𝔼⁡\[Δ​ϕmean\(w\)​\(t,s\)​Δ​ϕdev\(w\)​\(t,s\)⊤\]\\displaystyle\\mathbb\{E\}\\left\[\\Delta\\phi\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s\)\\Delta\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)^\{\\top\}\\right\]=𝔼\[\(mt,Zt\(w\)\(w\)−ms,Zs\(w\)\(w\)\)𝔼\[Δϕdev\(w\)\(t,s\)⊤∣Z1:T\(w\)\]\]\\displaystyle=\\mathbb\{E\}\\left\[\\left\(m\_\{t,Z\_\{t\}^\{\(w\)\}\}^\{\(w\)\}\-m\_\{s,Z\_\{s\}^\{\(w\)\}\}^\{\(w\)\}\\right\)\\mathbb\{E\}\\left\[\\Delta\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)^\{\\top\}\\mid Z\_\{1:T\}^\{\(w\)\}\\right\]\\right\]=0,\\displaystyle=0,where the last equality follows from \([4](https://arxiv.org/html/2609.30974#S2.E4)\)\. The second cross term vanishes analogously\. Consequently,

𝐌full\(w\)​\(t,s\)=𝐌mean\(w\)​\(t,s\)\+𝐌dev\(w\)​\(t,s\),\\mathbf\{M\}\_\{\\mathrm\{full\}\}^\{\(w\)\}\(t,s\)=\\mathbf\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s\)\+\\mathbf\{M\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\),which proves \([6](https://arxiv.org/html/2609.30974#S3.E6)\)\.

Applying the trace and the quadratic form𝒖r⊤​\(⋅\)​𝒖r\\bm\{u\}\_\{r\}^\{\\top\}\(\\cdot\)\\bm\{u\}\_\{r\}gives \([7](https://arxiv.org/html/2609.30974#S3.E7)\) for everyρ∈\{tr,1,…,d\}\\rho\\in\\\{\\mathrm\{tr\},1,\\ldots,d\\\}\. ∎

We next prove Theorem[3\.2](https://arxiv.org/html/2609.30974#S3.Thmtheorem2)\. We first decompose each second\-moment operator by conditioning on the latent sense path and then group the resulting terms according to the pair\(Zt\(w\),Zs\(w\)\)\(Z\_\{t\}^\{\(w\)\},Z\_\{s\}^\{\(w\)\}\)\.

###### Proof\.

Fixw∈Vw\\in V,t,s∈\[T\]t,s\\in\[T\], and∙∈\{mean,dev\}\\bullet\\in\\\{\\mathrm\{mean\},\\mathrm\{dev\}\\\}\.

For the mean process,

𝐌mean\(w\)​\(t,s\)\\displaystyle\\mathbf\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s\)=∑z1:TPr\(Z1:T\(w\)=z1:T\)\(mt,zt\(w\)−ms,zs\(w\)\)\(mt,zt\(w\)−ms,zs\(w\)\)⊤\\displaystyle=\\sum\_\{z\_\{1:T\}\}\\Pr\(Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\)\\left\(m\_\{t,z\_\{t\}\}^\{\(w\)\}\-m\_\{s,z\_\{s\}\}^\{\(w\)\}\\right\)\\left\(m\_\{t,z\_\{t\}\}^\{\(w\)\}\-m\_\{s,z\_\{s\}\}^\{\(w\)\}\\right\)^\{\\top\}=∑k=1Kt\(w\)∑ℓ=1Ks\(w\)\(∑z1:T:zt=k,zs=ℓPr\(Z1:T\(w\)=z1:T\)\)\\displaystyle=\\sum\_\{k=1\}^\{K\_\{t\}^\{\(w\)\}\}\\sum\_\{\\ell=1\}^\{K\_\{s\}^\{\(w\)\}\}\\left\(\\sum\_\{z\_\{1:T\}:z\_\{t\}=k,\\;z\_\{s\}=\\ell\}\\Pr\(Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\)\\right\)⋅\(mt,k\(w\)−ms,ℓ\(w\)\)​\(mt,k\(w\)−ms,ℓ\(w\)\)⊤\\displaystyle\\cdot\\left\(m\_\{t,k\}^\{\(w\)\}\-m\_\{s,\\ell\}^\{\(w\)\}\\right\)\\left\(m\_\{t,k\}^\{\(w\)\}\-m\_\{s,\\ell\}^\{\(w\)\}\\right\)^\{\\top\}=∑k=1Kt\(w\)∑ℓ=1Ks\(w\)q\(w\)​\(t,s,k,ℓ\)​𝐂mean\(w\)​\(t,s,k,ℓ\)\.\\displaystyle=\\sum\_\{k=1\}^\{K\_\{t\}^\{\(w\)\}\}\\sum\_\{\\ell=1\}^\{K\_\{s\}^\{\(w\)\}\}q^\{\(w\)\}\(t,s;k,\\ell\)\\mathbf\{C\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s;k,\\ell\)\.
For the deviation process, first note that, for anyk∈\[Kt\(w\)\]k\\in\[K\_\{t\}^\{\(w\)\}\]andℓ∈\[Ks\(w\)\]\\ell\\in\[K\_\{s\}^\{\(w\)\}\],

𝔼\[Xt\(w\)−Xs\(w\)∣Zt\(w\)=k,Zs\(w\)=ℓ\]\\displaystyle\\mathbb\{E\}\\left\[X\_\{t\}^\{\(w\)\}\-X\_\{s\}^\{\(w\)\}\\mid Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\]=∑z1:T:zt=k,zs=ℓPr\(Z1:T\(w\)=z1:T∣Zt\(w\)=k,Zs\(w\)=ℓ\)\\displaystyle=\\sum\_\{z\_\{1:T\}:z\_\{t\}=k,\\;z\_\{s\}=\\ell\}\\Pr\\left\(Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\mid Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\)⋅𝔼\[Xt\(w\)−Xs\(w\)∣Z1:T\(w\)=z1:T\]\\displaystyle\\cdot\\mathbb\{E\}\\left\[X\_\{t\}^\{\(w\)\}\-X\_\{s\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\]=∑z1:T:zt=k,zs=ℓPr\(Z1:T\(w\)=z1:T∣Zt\(w\)=k,Zs\(w\)=ℓ\)\\displaystyle=\\sum\_\{z\_\{1:T\}:z\_\{t\}=k,\\;z\_\{s\}=\\ell\}\\Pr\\left\(Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\mid Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\)⋅\(mt,zt\(w\)−ms,zs\(w\)\)\\displaystyle\\cdot\\left\(m\_\{t,z\_\{t\}\}^\{\(w\)\}\-m\_\{s,z\_\{s\}\}^\{\(w\)\}\\right\)=mt,k\(w\)−ms,ℓ\(w\),\\displaystyle=m\_\{t,k\}^\{\(w\)\}\-m\_\{s,\\ell\}^\{\(w\)\},where the second equality follows from the conditional marginal preservation in Proposition[2\.2](https://arxiv.org/html/2609.30974#S2.Thmtheorem2)\. Hence,

𝔼\[Δϕdev\(w\)\(t,s\)Δϕdev\(w\)\(t,s\)⊤∣Zt\(w\)=k,Zs\(w\)=ℓ\]\\displaystyle\\mathbb\{E\}\\left\[\\Delta\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)\\Delta\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)^\{\\top\}\\mid Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\]=𝔼\[\{Xt\(w\)−Xs\(w\)−\(mt,k\(w\)−ms,ℓ\(w\)\)\}\{Xt\(w\)−Xs\(w\)−\(mt,k\(w\)−ms,ℓ\(w\)\)\}⊤\|Zt\(w\)=k,Zs\(w\)=ℓ\]\\displaystyle=\\mathbb\{E\}\\left\[\\left\\\{X\_\{t\}^\{\(w\)\}\-X\_\{s\}^\{\(w\)\}\-\\left\(m\_\{t,k\}^\{\(w\)\}\-m\_\{s,\\ell\}^\{\(w\)\}\\right\)\\right\\\}\\left\\\{X\_\{t\}^\{\(w\)\}\-X\_\{s\}^\{\(w\)\}\-\\left\(m\_\{t,k\}^\{\(w\)\}\-m\_\{s,\\ell\}^\{\(w\)\}\\right\)\\right\\\}^\{\\top\}\\middle\|Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\]=Cov⁡\(Xt\(w\)−Xs\(w\)∣Zt\(w\)=k,Zs\(w\)=ℓ\)\\displaystyle=\\operatorname\{Cov\}\\left\(X\_\{t\}^\{\(w\)\}\-X\_\{s\}^\{\(w\)\}\\mid Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\)=𝐂dev\(w\)​\(t,s,k,ℓ\)\.\\displaystyle=\\mathbf\{C\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s;k,\\ell\)\.Therefore, conditioning on the pair\(Zt\(w\),Zs\(w\)\)\(Z\_\{t\}^\{\(w\)\},Z\_\{s\}^\{\(w\)\}\)gives

𝐌dev\(w\)​\(t,s\)\\displaystyle\\mathbf\{M\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)=∑k=1Kt\(w\)∑ℓ=1Ks\(w\)q\(w\)​\(t,s,k,ℓ\)\\displaystyle=\\sum\_\{k=1\}^\{K\_\{t\}^\{\(w\)\}\}\\sum\_\{\\ell=1\}^\{K\_\{s\}^\{\(w\)\}\}q^\{\(w\)\}\(t,s;k,\\ell\)⋅𝔼\[Δϕdev\(w\)\(t,s\)Δϕdev\(w\)\(t,s\)⊤∣Zt\(w\)=k,Zs\(w\)=ℓ\]\\displaystyle\\cdot\\mathbb\{E\}\\left\[\\Delta\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)\\Delta\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s\)^\{\\top\}\\mid Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\]=∑k=1Kt\(w\)∑ℓ=1Ks\(w\)q\(w\)​\(t,s,k,ℓ\)​𝐂dev\(w\)​\(t,s,k,ℓ\)\.\\displaystyle=\\sum\_\{k=1\}^\{K\_\{t\}^\{\(w\)\}\}\\sum\_\{\\ell=1\}^\{K\_\{s\}^\{\(w\)\}\}q^\{\(w\)\}\(t,s;k,\\ell\)\\mathbf\{C\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s;k,\\ell\)\.
Thus, \([8](https://arxiv.org/html/2609.30974#S3.E8)\) follows for both∙∈\{mean,dev\}\\bullet\\in\\\{\\mathrm\{mean\},\\mathrm\{dev\}\\\}\.

Applying the trace and the quadratic form𝒖r\(w\)⊤​\(⋅\)​𝒖r\(w\)\\bm\{u\}\_\{r\}^\{\(w\)\\top\}\(\\cdot\)\\bm\{u\}\_\{r\}^\{\(w\)\}gives \([9](https://arxiv.org/html/2609.30974#S3.E9)\) and \([10](https://arxiv.org/html/2609.30974#S3.E10)\)\. ∎

## Appendix EProofs of Properties of Trace and Mode\-Wise Geometry

### E\.1Validity and properties of the induced semantic distances

The semantic distances introduced in the main text define the geometry of semantic change on the period index set\. This subsection establishes the pseudometric properties of the semantic distances induced by the correspondingL2L^\{2\}random\-vector representations\.

The first property concerns the pseudo\-metric structure induced by the random\-vector representations\.

###### Proposition E\.1\(Pseudo\-metric properties of trace and mode\-wise semantic distances\)\.

For everyw∈Vw\\in Vand∙∈\{full,mean,dev\}\\bullet\\in\\\{\\mathrm\{full\},\\mathrm\{mean\},\\mathrm\{dev\}\\\}, the quantities

d∙,tr\(w\)​\(t,s\),d∙,r\(w\)​\(t,s\),r∈\[d\],d\_\{\\bullet,\\mathrm\{tr\}\}^\{\(w\)\}\(t,s\),\\qquad d\_\{\\bullet,r\}^\{\(w\)\}\(t,s\),\\quad r\\in\[d\],define finite pseudo\-metrics on\[T\]\[T\]\.

Moreover, the zero\-distance conditions are given by

d∙,tr\(w\)​\(t,s\)=0⇔ϕ∙\(w\)​\(t\)=ϕ∙\(w\)​\(s\)almost surely,d\_\{\\bullet,\\mathrm\{tr\}\}^\{\(w\)\}\(t,s\)=0\\iff\\phi\_\{\\bullet\}^\{\(w\)\}\(t\)=\\phi\_\{\\bullet\}^\{\(w\)\}\(s\)\\quad\\text\{almost surely\},and

d∙,r\(w\)​\(t,s\)=0⇔𝒖r⊤​ϕ∙\(w\)​\(t\)=𝒖r⊤​ϕ∙\(w\)​\(s\)almost surely\.d\_\{\\bullet,r\}^\{\(w\)\}\(t,s\)=0\\iff\\bm\{u\}\_\{r\}^\{\\top\}\\phi\_\{\\bullet\}^\{\(w\)\}\(t\)=\\bm\{u\}\_\{r\}^\{\\top\}\\phi\_\{\\bullet\}^\{\(w\)\}\(s\)\\quad\\text\{almost surely\}\.

Remark\.Zero distance implies equality only in the corresponding random\-vector representation and does not necessarily implyt=st=s\. For the mode\-wise distances, the zero\-distance condition only concerns the projection onto the corresponding direction𝒖r\\bm\{u\}\_\{r\}\.

###### Proof\.

Sinceνt,k\(w\)∈𝒫2​\(ℝd\)\\nu\_\{t,k\}^\{\(w\)\}\\in\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\), the mixture distribution satisfiesγt\(w\)∈𝒫2​\(ℝd\)\\gamma\_\{t\}^\{\(w\)\}\\in\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)\. By the marginal consistency of the full process in \([3](https://arxiv.org/html/2609.30974#S2.E3)\),ϕfull\(w\)​\(t\)∈L2​\(Ω,ℝd\)\\phi\_\{\\mathrm\{full\}\}^\{\(w\)\}\(t\)\\in L^\{2\}\(\\Omega;\\mathbb\{R\}^\{d\}\)\. Moreover,ϕmean\(w\)​\(t\)∈L2​\(Ω,ℝd\)\\phi\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t\)\\in L^\{2\}\(\\Omega;\\mathbb\{R\}^\{d\}\), since it takes values in the finite set of sense centers\{mt,k\(w\)\}k=1Kt\(w\)\\\{m\_\{t,k\}^\{\(w\)\}\\\}\_\{k=1\}^\{K\_\{t\}^\{\(w\)\}\}\. Therefore,

ϕdev\(w\)​\(t\)=ϕfull\(w\)​\(t\)−ϕmean\(w\)​\(t\)∈L2​\(Ω,ℝd\)\.\\phi\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t\)=\\phi\_\{\\mathrm\{full\}\}^\{\(w\)\}\(t\)\-\\phi\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t\)\\in L^\{2\}\(\\Omega;\\mathbb\{R\}^\{d\}\)\.Hence, all the followingL2L^\{2\}distances are well\-defined\.

For the trace distance, the definition gives

d∙,tr\(w\)​\(t,s\)=\(𝔼⁡\[‖ϕ∙\(w\)​\(t\)−ϕ∙\(w\)​\(s\)‖2\]\)1/2=‖ϕ∙\(w\)​\(t\)−ϕ∙\(w\)​\(s\)‖L2​\(Ω,ℝd\)\.d\_\{\\bullet,\\mathrm\{tr\}\}^\{\(w\)\}\(t,s\)=\\left\(\\mathbb\{E\}\\left\[\\left\\\|\\phi\_\{\\bullet\}^\{\(w\)\}\(t\)\-\\phi\_\{\\bullet\}^\{\(w\)\}\(s\)\\right\\\|^\{2\}\\right\]\\right\)^\{1/2\}=\\left\\\|\\phi\_\{\\bullet\}^\{\(w\)\}\(t\)\-\\phi\_\{\\bullet\}^\{\(w\)\}\(s\)\\right\\\|\_\{L^\{2\}\(\\Omega;\\mathbb\{R\}^\{d\}\)\}\.Hence,d∙,tr\(w\)d\_\{\\bullet,\\mathrm\{tr\}\}^\{\(w\)\}is theL2L^\{2\}distance between the random vectorsϕ∙\(w\)​\(t\)\\phi\_\{\\bullet\}^\{\(w\)\}\(t\)andϕ∙\(w\)​\(s\)\\phi\_\{\\bullet\}^\{\(w\)\}\(s\)\. Nonnegativity and symmetry follow directly from the properties of the norm, and the triangle inequality follows from Minkowski’s inequality\.

Similarly,

d∙,r\(w\)​\(t,s\)=‖𝒖r⊤​ϕ∙\(w\)​\(t\)−𝒖r⊤​ϕ∙\(w\)​\(s\)‖L2​\(Ω,ℝ\),d\_\{\\bullet,r\}^\{\(w\)\}\(t,s\)=\\left\\\|\\bm\{u\}\_\{r\}^\{\\top\}\\phi\_\{\\bullet\}^\{\(w\)\}\(t\)\-\\bm\{u\}\_\{r\}^\{\\top\}\\phi\_\{\\bullet\}^\{\(w\)\}\(s\)\\right\\\|\_\{L^\{2\}\(\\Omega;\\mathbb\{R\}\)\},which is theL2L^\{2\}distance between the scalar random variables obtained by projecting the process onto the direction𝒖r\\bm\{u\}\_\{r\}\. Therefore, the same arguments establish nonnegativity, symmetry, and the triangle inequality ford∙,r\(w\)d\_\{\\bullet,r\}^\{\(w\)\}\.

Finally, anL2L^\{2\}norm is zero if and only if the corresponding random elements agree almost surely\. Applying this property to the vector\-valued and projected representations gives the stated zero\-distance conditions\. ∎

### E\.2Proofs of the decomposition and characterization of the semantic geometry

This section provides proofs of the characterization and decomposition results for the semantic geometry introduced in the main text\. The proofs follow the same arguments as those in\[[16](https://arxiv.org/html/2609.30974#bib.bib11)\]\.

The following proof establishes the decomposition of the total semantic change into orthogonal mode\-wise contributions\.

###### Proof of \([5](https://arxiv.org/html/2609.30974#S3.E5)\)\.

Since\{𝒖r\}r=1d\\\{\\bm\{u\}\_\{r\}\\\}\_\{r=1\}^\{d\}is an orthonormal basis,

tr⁡𝐌∙\(w\)​\(t,s\)=∑r=1d\(𝒖r\(w\)\)⊤​𝐌∙\(w\)​\(t,s\)​𝒖r\(w\)\.\\operatorname\{tr\}\\mathbf\{M\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)=\\sum\_\{r=1\}^\{d\}\(\\bm\{u\}\_\{r\}^\{\(w\)\}\)^\{\\top\}\\mathbf\{M\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)\\bm\{u\}\_\{r\}^\{\(w\)\}\.Using the definitions ofd∙,tr\(w\)​\(t,s\)d\_\{\\bullet,\\mathrm\{tr\}\}^\{\(w\)\}\(t,s\)andd∙,r\(w\)​\(t,s\)d\_\{\\bullet,r\}^\{\(w\)\}\(t,s\)yields the claim\. ∎

The following proof establishes the spectral characterization of word\-local modes induced by each word’s aggregated mean operator\.

###### Proof of Proposition[3\.4](https://arxiv.org/html/2609.30974#S3.Thmtheorem4)\.

By definition,

ℳmean\(w\)=∑t=1T−1𝐌mean\(w\)​\(t,t\+1\)\.\\mathcal\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}=\\sum\_\{t=1\}^\{T\-1\}\\mathbf\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,t\+1\)\.Since

ℳmean\(w\)​𝒖r\(w\)=λr\(w\)​𝒖r\(w\)and‖𝒖r\(w\)‖=1,\\mathcal\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\\bm\{u\}\_\{r\}^\{\(w\)\}=\\lambda\_\{r\}^\{\(w\)\}\\bm\{u\}\_\{r\}^\{\(w\)\}\\quad\\text\{and\}\\quad\\\|\\bm\{u\}\_\{r\}^\{\(w\)\}\\\|=1,we have

λr\(w\)\\displaystyle\\lambda\_\{r\}^\{\(w\)\}=\(𝒖r\(w\)\)⊤​ℳmean\(w\)​𝒖r\(w\)\\displaystyle=\(\\bm\{u\}\_\{r\}^\{\(w\)\}\)^\{\\top\}\\mathcal\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\\bm\{u\}\_\{r\}^\{\(w\)\}=∑t=1T−1\(𝒖r\(w\)\)⊤​𝐌mean\(w\)​\(t,t\+1\)​𝒖r\(w\)\\displaystyle=\\sum\_\{t=1\}^\{T\-1\}\(\\bm\{u\}\_\{r\}^\{\(w\)\}\)^\{\\top\}\\mathbf\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,t\+1\)\\bm\{u\}\_\{r\}^\{\(w\)\}=∑t=1T−1\(dmean,r\(w\)​\(t,t\+1\)\)2,\\displaystyle=\\sum\_\{t=1\}^\{T\-1\}\\left\(d\_\{\\mathrm\{mean\},r\}^\{\(w\)\}\(t,t\+1\)\\right\)^\{2\},which proves \([11](https://arxiv.org/html/2609.30974#S3.E11)\)\.

Moreover, sinceℳmean\(w\)\\mathcal\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}is symmetric, the Courant–Fischer variational characterization gives

𝒖r\(w\)∈arg⁡max‖𝒖‖=1𝒖⟂𝒖1\(w\),…,𝒖r−1\(w\)​𝒖⊤​ℳmean\(w\)​𝒖\.\\bm\{u\}\_\{r\}^\{\(w\)\}\\in\\arg\\max\_\{\\begin\{subarray\}\{c\}\\\|\\bm\{u\}\\\|=1\\\\ \\bm\{u\}\\perp\\bm\{u\}\_\{1\}^\{\(w\)\},\\ldots,\\bm\{u\}\_\{r\-1\}^\{\(w\)\}\\end\{subarray\}\}\\bm\{u\}^\{\\top\}\\mathcal\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\\bm\{u\}\.Substituting the definition ofℳmean\(w\)\\mathcal\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}yields

𝒖r\(w\)∈arg⁡max⁡∑t=1T−1‖𝒖‖=1𝒖⟂𝒖1\(w\),…,𝒖r−1\(w\)⁡𝒖⊤​𝐌mean\(w\)​\(t,t\+1\)​𝒖,\\bm\{u\}\_\{r\}^\{\(w\)\}\\in\\arg\\max\_\{\\begin\{subarray\}\{c\}\\\|\\bm\{u\}\\\|=1\\\\ \\bm\{u\}\\perp\\bm\{u\}\_\{1\}^\{\(w\)\},\\ldots,\\bm\{u\}\_\{r\-1\}^\{\(w\)\}\\end\{subarray\}\}\\sum\_\{t=1\}^\{T\-1\}\\bm\{u\}^\{\\top\}\\mathbf\{M\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,t\+1\)\\bm\{u\},which proves \([12](https://arxiv.org/html/2609.30974#S3.E12)\)\. ∎

## Appendix FExplicit expressions under the Gaussian specialization

Under the Gaussian and quadratic\-cost specialization in Section[3](https://arxiv.org/html/2609.30974#S3), both the contextual coupling cost and the sense\-pair second\-moment components admit explicit expressions\.

First, the contextual coupling problem in \([1](https://arxiv.org/html/2609.30974#S2.E1)\) reduces to the quadratic22\-Wasserstein problem between Gaussian measures\. As established, e\.g\., in\[[35](https://arxiv.org/html/2609.30974#bib.bib25),[36](https://arxiv.org/html/2609.30974#bib.bib32)\], the optimal coupling exists and is unique under positive\-definite covariance matrices, and is induced by the affine optimal transport map given in Proposition[3\.5](https://arxiv.org/html/2609.30974#S3.Thmtheorem5)\. Its minimum cost is therefore

Ct,k,ℓ\(w\)\\displaystyle C\_\{t,k,\\ell\}^\{\(w\)\}=W22​\(νt,k\(w\),νt\+1,ℓ\(w\)\)\\displaystyle=W\_\{2\}^\{2\}\\\!\\left\(\\nu\_\{t,k\}^\{\(w\)\},\\nu\_\{t\+1,\\ell\}^\{\(w\)\}\\right\)=‖μt,k\(w\)−μt\+1,ℓ\(w\)‖2\+Tr⁡\(Σt,k\(w\)\+Σt\+1,ℓ\(w\)−2​\[\(Σt,k\(w\)\)1/2​Σt\+1,ℓ\(w\)​\(Σt,k\(w\)\)1/2\]1/2\)\.\\displaystyle=\\left\\\|\\mu\_\{t,k\}^\{\(w\)\}\-\\mu\_\{t\+1,\\ell\}^\{\(w\)\}\\right\\\|^\{2\}\+\\operatorname\{Tr\}\\\!\\left\(\\Sigma\_\{t,k\}^\{\(w\)\}\+\\Sigma\_\{t\+1,\\ell\}^\{\(w\)\}\-2\\left\[\(\\Sigma\_\{t,k\}^\{\(w\)\}\)^\{1/2\}\\Sigma\_\{t\+1,\\ell\}^\{\(w\)\}\(\\Sigma\_\{t,k\}^\{\(w\)\}\)^\{1/2\}\\right\]^\{1/2\}\\right\)\.\(14\)
We next give explicit expressions for the sense\-pair second\-moment components used in Theorem[3\.2](https://arxiv.org/html/2609.30974#S3.Thmtheorem2)\.

###### Proposition F\.1\(Explicit expressions for sense\-pair second moments\)\.

Under the Gaussian specialization, for everyw∈Vw\\in V,t<st<s, and pair\(k,ℓ\)\(k,\\ell\)withq\(w\)​\(t,s,k,ℓ\)\>0q^\{\(w\)\}\(t,s;k,\\ell\)\>0, the mean component is

𝐂mean\(w\)​\(t,s,k,ℓ\)=\(μt,k\(w\)−μs,ℓ\(w\)\)​\(μt,k\(w\)−μs,ℓ\(w\)\)⊤,\\mathbf\{C\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s;k,\\ell\)=\\left\(\\mu\_\{t,k\}^\{\(w\)\}\-\\mu\_\{s,\\ell\}^\{\(w\)\}\\right\)\\left\(\\mu\_\{t,k\}^\{\(w\)\}\-\\mu\_\{s,\\ell\}^\{\(w\)\}\\right\)^\{\\top\},\(15\)and the deviation component is

𝐂dev\(w\)​\(t,s,k,ℓ\)=Σt,k\(w\)\+Σs,ℓ\(w\)−Ξ\(w\)​\(t,s,k,ℓ\)−Ξ\(w\)​\(t,s,k,ℓ\)⊤,\\mathbf\{C\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s;k,\\ell\)=\\Sigma\_\{t,k\}^\{\(w\)\}\+\\Sigma\_\{s,\\ell\}^\{\(w\)\}\-\\Xi^\{\(w\)\}\(t,s;k,\\ell\)\-\\Xi^\{\(w\)\}\(t,s;k,\\ell\)^\{\\top\},\(16\)where

Mt,s\(w\)\\displaystyle M\_\{t,s\}^\{\(w\)\}:=\(As−1,Zs−1\(w\),Zs\(w\)\(w\)⋯At,Zt\(w\),Zt\+1\(w\)\(w\)\)⊤,\\displaystyle:=\\left\(A\_\{s\-1,Z\_\{s\-1\}^\{\(w\)\},Z\_\{s\}^\{\(w\)\}\}^\{\(w\)\}\\cdots A\_\{t,Z\_\{t\}^\{\(w\)\},Z\_\{t\+1\}^\{\(w\)\}\}^\{\(w\)\}\\right\)^\{\\top\},\(17\)Ξ\(w\)​\(t,s,k,ℓ\)\\displaystyle\\Xi^\{\(w\)\}\(t,s;k,\\ell\):=Σt,k\(w\)𝔼\[Mt,s\(w\)\|Zt\(w\)=k,Zs\(w\)=ℓ\]\\displaystyle:=\\Sigma\_\{t,k\}^\{\(w\)\}\\mathbb\{E\}\\\!\\left\[M\_\{t,s\}^\{\(w\)\}\\middle\|Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\]\(18\)

###### Proof\.

The expression for𝐂mean\(w\)​\(t,s,k,ℓ\)\\mathbf\{C\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s;k,\\ell\)follows directly from the means of the Gaussian sense components\.

For the deviation component, apply the law of total covariance by conditioning further on the complete latent sense sequence:

𝐂dev\(w\)​\(t,s,k,ℓ\)\\displaystyle\\mathbf\{C\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s;k,\\ell\)=𝔼\[Cov\(Xt\(w\)−Xs\(w\)∣Z1:T\(w\)\)\|Zt\(w\)=k,Zs\(w\)=ℓ\]\\displaystyle=\\mathbb\{E\}\\\!\\left\[\\operatorname\{Cov\}\\\!\\left\(X\_\{t\}^\{\(w\)\}\-X\_\{s\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}\\right\)\\middle\|Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\]\+Cov\(𝔼\[Xt\(w\)−Xs\(w\)∣Z1:T\(w\)\]\|Zt\(w\)=k,Zs\(w\)=ℓ\)\.\\displaystyle\\quad\+\\operatorname\{Cov\}\\\!\\left\(\\mathbb\{E\}\\\!\\left\[X\_\{t\}^\{\(w\)\}\-X\_\{s\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}\\right\]\\middle\|Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\)\.\(19\)By the conditional marginals established in Proposition[2\.2](https://arxiv.org/html/2609.30974#S2.Thmtheorem2),

Cov\(𝔼\[Xt\(w\)−Xs\(w\)∣Z1:T\(w\)\]\|Zt\(w\)=k,Zs\(w\)=ℓ\)\\displaystyle\\operatorname\{Cov\}\\\!\\left\(\\mathbb\{E\}\\\!\\left\[X\_\{t\}^\{\(w\)\}\-X\_\{s\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}\\right\]\\middle\|Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\)\(20\)=Cov\(μt,k\(w\)−μs,ℓ\(w\)\|Zt\(w\)=k,Zs\(w\)=ℓ\)=0,\\displaystyle=\\operatorname\{Cov\}\\\!\\left\(\\mu\_\{t,k\}^\{\(w\)\}\-\\mu\_\{s,\\ell\}^\{\(w\)\}\\middle\|Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\)=0,\(21\)thus the second term vanishes\.

Now fix a latent sense sequencez1:Tz\_\{1:T\}satisfyingzt=kz\_\{t\}=kandzs=ℓz\_\{s\}=\\ell\. The conditional covariance of the displacement decomposes as

Cov\(Xt\(w\)−Xs\(w\)∣Z1:T\(w\)=z1:T\)\\displaystyle\\operatorname\{Cov\}\\\!\\left\(X\_\{t\}^\{\(w\)\}\-X\_\{s\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\)=Var\(Xt\(w\)∣Z1:T\(w\)=z1:T\)\+Var\(Xs\(w\)∣Z1:T\(w\)=z1:T\)\\displaystyle=\\operatorname\{Var\}\\\!\\left\(X\_\{t\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\)\+\\operatorname\{Var\}\\\!\\left\(X\_\{s\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\)−Cov\(Xt\(w\),Xs\(w\)∣Z1:T\(w\)=z1:T\)−Cov\(Xs\(w\),Xt\(w\)∣Z1:T\(w\)=z1:T\)\.\\displaystyle\\quad\-\\operatorname\{Cov\}\\\!\\left\(X\_\{t\}^\{\(w\)\},X\_\{s\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\)\-\\operatorname\{Cov\}\\\!\\left\(X\_\{s\}^\{\(w\)\},X\_\{t\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\)\.\(22\)By Proposition[2\.2](https://arxiv.org/html/2609.30974#S2.Thmtheorem2),

Var\(Xt\(w\)∣Z1:T\(w\)=z1:T\)=Σt,k\(w\),Var\(Xs\(w\)∣Z1:T\(w\)=z1:T\)\\displaystyle\\operatorname\{Var\}\\\!\\left\(X\_\{t\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\)=\\Sigma\_\{t,k\}^\{\(w\)\},\\qquad\\operatorname\{Var\}\\\!\\left\(X\_\{s\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\)=Σs,ℓ\(w\)\.\\displaystyle=\\Sigma\_\{s,\\ell\}^\{\(w\)\}\.\(23\)
Moreover, Proposition[3\.5](https://arxiv.org/html/2609.30974#S3.Thmtheorem5)gives the affine optimal transport map for each adjacent sense transition\. Consequently, the Markov composition along the fixed sense path gives

Xs\(w\)\\displaystyle X\_\{s\}^\{\(w\)\}=μs,ℓ\(w\)\+As−1,zs−1,zs\(w\)⋯At,zt,zt\+1\(w\)\(Xt\(w\)−μt,k\(w\)\)\.\\displaystyle=\\mu\_\{s,\\ell\}^\{\(w\)\}\+A\_\{s\-1,z\_\{s\-1\},z\_\{s\}\}^\{\(w\)\}\\cdots A\_\{t,z\_\{t\},z\_\{t\+1\}\}^\{\(w\)\}\\left\(X\_\{t\}^\{\(w\)\}\-\\mu\_\{t,k\}^\{\(w\)\}\\right\)\.\(24\)Taking the conditional covariance gives

Cov\(Xt\(w\),Xs\(w\)∣Z1:T\(w\)=z1:T\)=Cov\(Xs\(w\),Xt\(w\)∣Z1:T\(w\)=z1:T\)⊤\\displaystyle\\operatorname\{Cov\}\\\!\\left\(X\_\{t\}^\{\(w\)\},X\_\{s\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\)=\\operatorname\{Cov\}\\\!\\left\(X\_\{s\}^\{\(w\)\},X\_\{t\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\)^\{\\top\}=Σt,k\(w\)\(As−1,zs−1,zs\(w\)⋯At,zt,zt\+1\(w\)\)⊤,\\displaystyle\\qquad=\\Sigma\_\{t,k\}^\{\(w\)\}\\left\(A\_\{s\-1,z\_\{s\-1\},z\_\{s\}\}^\{\(w\)\}\\cdots A\_\{t,z\_\{t\},z\_\{t\+1\}\}^\{\(w\)\}\\right\)^\{\\top\},\(25\)
Substituting these expressions into \([22](https://arxiv.org/html/2609.30974#A6.E22)\), we obtain

Cov\(Xt\(w\)−Xs\(w\)∣Z1:T\(w\)=z1:T\)\\displaystyle\\operatorname\{Cov\}\\\!\\left\(X\_\{t\}^\{\(w\)\}\-X\_\{s\}^\{\(w\)\}\\mid Z\_\{1:T\}^\{\(w\)\}=z\_\{1:T\}\\right\)=Σt,k\(w\)\+Σs,ℓ\(w\)−Σt,k\(w\)\(As−1,zs−1,zs\(w\)⋯At,zt,zt\+1\(w\)\)⊤\\displaystyle=\\Sigma\_\{t,k\}^\{\(w\)\}\+\\Sigma\_\{s,\\ell\}^\{\(w\)\}\-\\Sigma\_\{t,k\}^\{\(w\)\}\\left\(A\_\{s\-1,z\_\{s\-1\},z\_\{s\}\}^\{\(w\)\}\\cdots A\_\{t,z\_\{t\},z\_\{t\+1\}\}^\{\(w\)\}\\right\)^\{\\top\}−\[Σt,k\(w\)\(As−1,zs−1,zs\(w\)⋯At,zt,zt\+1\(w\)\)⊤\]⊤\.\\displaystyle\\quad\-\\left\[\\Sigma\_\{t,k\}^\{\(w\)\}\\left\(A\_\{s\-1,z\_\{s\-1\},z\_\{s\}\}^\{\(w\)\}\\cdots A\_\{t,z\_\{t\},z\_\{t\+1\}\}^\{\(w\)\}\\right\)^\{\\top\}\\right\]^\{\\top\}\.\(26\)
Combining the above expressions with \([19](https://arxiv.org/html/2609.30974#A6.E19)\), we obtain

𝐂dev\(w\)​\(t,s,k,ℓ\)\\displaystyle\\mathbf\{C\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s;k,\\ell\)=Σt,k\(w\)\+Σs,ℓ\(w\)−Σt,k\(w\)𝔼\[Mt,s\(w\)\|Zt\(w\)=k,Zs\(w\)=ℓ\]\\displaystyle=\\Sigma\_\{t,k\}^\{\(w\)\}\+\\Sigma\_\{s,\\ell\}^\{\(w\)\}\-\\Sigma\_\{t,k\}^\{\(w\)\}\\mathbb\{E\}\\\!\\left\[M\_\{t,s\}^\{\(w\)\}\\middle\|Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\]−𝔼\[Mt,s\(w\)⊤\|Zt\(w\)=k,Zs\(w\)=ℓ\]Σt,k\(w\)\\displaystyle\\quad\-\\mathbb\{E\}\\\!\\left\[M\_\{t,s\}^\{\(w\)\\top\}\\middle\|Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\]\\Sigma\_\{t,k\}^\{\(w\)\}\(27\)=Σt,k\(w\)\+Σs,ℓ\(w\)−Ξ\(w\)​\(t,s,k,ℓ\)−Ξ\(w\)​\(t,s,k,ℓ\)⊤,\\displaystyle=\\Sigma\_\{t,k\}^\{\(w\)\}\+\\Sigma\_\{s,\\ell\}^\{\(w\)\}\-\\Xi^\{\(w\)\}\(t,s;k,\\ell\)\-\\Xi^\{\(w\)\}\(t,s;k,\\ell\)^\{\\top\},\(28\)which proves the stated expression for𝐂dev\(w\)​\(t,s,k,ℓ\)\\mathbf\{C\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s;k,\\ell\)\. ∎

## Appendix GDynamic programming for transition quantities

We describe how to compute the sense\-transition quantities without enumerating all possible intermediate sense paths\. In particular, the sense\-pair probabilitiesq\(w\)​\(t,s,k,ℓ\)q^\{\(w\)\}\(t,s;k,\\ell\)and the cross\-period covariance quantitiesΞ\(w\)​\(t,s,k,ℓ\)\\Xi^\{\(w\)\}\(t,s;k,\\ell\), and hence the deviation components𝐂dev\(w\)​\(t,s,k,ℓ\)\\mathbf\{C\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s;k,\\ell\), can be computed by dynamic programming\.

Computational cost\.We analyze the computational cost of computing non\-adjacent quantities induced by the Markov chain after GMM fitting and adjacent period OT solving have already been completed, and hence the corresponding transport mapsAt,k,ℓ\(w\)A\_\{t,k,\\ell\}^\{\(w\)\}and transition probabilitiesPt\(w\)​\(k,ℓ\)P\_\{t\}^\{\(w\)\}\(k,\\ell\)are available\. A direct pathwise computation requires enumerating latent component paths across periods, leading to an exponential increase in computational cost with the number of periodsTT\. We therefore exploit dynamic programming to aggregate contributions from intermediate latent components without explicit path enumeration\. For one word, letK=maxt⁡Kt\(w\)K=\\max\_\{t\}K\_\{t\}^\{\(w\)\}, and letdddenote the embedding dimension\. The memory costs below refer only to additional memory required during computation and exclude storage of the precomputed mixture parameters, adjacent period transport maps and transition probabilities, as well as the final outputs\.

Considering the computation of all quantities indexed by\(t,s,k,ℓ\)\(t,s,k,\\ell\), the joint component probabilitiesq⁡\(t,s,k,ℓ\)q\(t,s;k,\\ell\)can be obtained by either explicit enumeration of latent component paths or dynamic programming\. A naive enumeration of all latent component paths yields the upper boundO⁡\(T3​KT\)O\(T^\{3\}K^\{T\}\), whereas the dynamic programming recursion requires onlyO⁡\(T2​K3\)O\(T^\{2\}K^\{3\}\)\. In terms of memory, naive path enumeration can be implemented withO⁡\(T\)O\(T\)memory by generating one path at a time, while the dynamic programming recursion requiresO⁡\(K2\)O\(K^\{2\}\)memory for the current transition\-probability matrix\.

For the computation ofΞ\(w\)​\(t,s,k,ℓ\)\\Xi^\{\(w\)\}\(t,s;k,\\ell\), naive enumeration of all latent component paths yields the upper boundO⁡\(T3​KT​d3\)O\(T^\{3\}K^\{T\}d^\{3\}\)in time andO⁡\(T\+d2\)O\(T\+d^\{2\}\)in memory\. In contrast, the dynamic programming recursion requiresO⁡\(T2​K3​d3\)O\(T^\{2\}K^\{3\}d^\{3\}\)time andO⁡\(K2​d2\)O\(K^\{2\}d^\{2\}\)memory\. The conditional weightsrt,s​\(j∣k,ℓ\)r\_\{t,s\}\(j\\mid k,\\ell\)can be computed inO⁡\(1\)O\(1\)time for each\(t,s,k,ℓ,j\)\(t,s,k,\\ell,j\)without additional memory, provided that the recursion forRt,sR\_\{t,s\}used to computeq⁡\(t,s,k,ℓ\)q\(t,s;k,\\ell\)is carried out alongside theΞ\\Xirecursion\.

OnceΞ\(w\)​\(t,s,k,ℓ\)\\Xi^\{\(w\)\}\(t,s;k,\\ell\)has been computed, the deviation component

𝐂dev\(w\)​\(t,s,k,ℓ\)=Σt,k\(w\)\+Σs,ℓ\(w\)−Ξ\(w\)​\(t,s,k,ℓ\)−Ξ\(w\)​\(t,s,k,ℓ\)⊤\\mathbf\{C\}\_\{\\mathrm\{dev\}\}^\{\(w\)\}\(t,s;k,\\ell\)=\\Sigma\_\{t,k\}^\{\(w\)\}\+\\Sigma\_\{s,\\ell\}^\{\(w\)\}\-\\Xi^\{\(w\)\}\(t,s;k,\\ell\)\-\\Xi^\{\(w\)\}\(t,s;k,\\ell\)^\{\\top\}requires onlyO⁡\(d2\)O\(d^\{2\}\)time andO⁡\(d2\)O\(d^\{2\}\)memory for each\(t,s,k,ℓ\)\(t,s,k,\\ell\), yieldingO⁡\(T2​K2​d2\)O\(T^\{2\}K^\{2\}d^\{2\}\)time overall\. This is lower\-order than theO⁡\(T2​K3​d3\)O\(T^\{2\}K^\{3\}d^\{3\}\)cost of computingΞ\\Xi, and therefore the additional computational cost of obtaining𝐂dev\\mathbf\{C\}\_\{\\mathrm\{dev\}\}is negligible in comparison\.

###### Proposition G\.1\(Dynamic programming recursions\)\.

Fort≤st\\leq s, let

Rt,s\(w\)​\(k,ℓ\):=Pr⁡\(Zs\(w\)=ℓ∣Zt\(w\)=k\)\.R\_\{t,s\}^\{\(w\)\}\(k,\\ell\):=\\Pr\\\!\\left\(Z\_\{s\}^\{\(w\)\}=\\ell\\mid Z\_\{t\}^\{\(w\)\}=k\\right\)\.Then

Rt,t\(w\)=IKt\(w\),Rt,r\+1\(w\)=Rt,r\(w\)Pr\(w\),r=t,…,s−1,R\_\{t,t\}^\{\(w\)\}=I\_\{K\_\{t\}^\{\(w\)\}\},\\qquad R\_\{t,r\+1\}^\{\(w\)\}=R\_\{t,r\}^\{\(w\)\}P\_\{r\}^\{\(w\)\},\\quad r=t,\\ldots,s\-1,\(29\)and hence

q\(w\)​\(t,s,k,ℓ\)=πt,k\(w\)​Rt,s\(w\)​\(k,ℓ\)\.q^\{\(w\)\}\(t,s;k,\\ell\)=\\pi\_\{t,k\}^\{\(w\)\}R\_\{t,s\}^\{\(w\)\}\(k,\\ell\)\.\(30\)
Fort<st<s, define the backward smoothing probability

rt,s\(w\)​\(j∣k,ℓ\):=Pr⁡\(Zs−1\(w\)=j∣Zt\(w\)=k,Zs\(w\)=ℓ\)\.r\_\{t,s\}^\{\(w\)\}\(j\\mid k,\\ell\):=\\Pr\\\!\\left\(Z\_\{s\-1\}^\{\(w\)\}=j\\mid Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\)\.Then, wheneverRt,s\(w\)​\(k,ℓ\)\>0R\_\{t,s\}^\{\(w\)\}\(k,\\ell\)\>0,

rt,s\(w\)​\(j∣k,ℓ\)=Rt,s−1\(w\)​\(k,j\)​Ps−1\(w\)​\(j,ℓ\)Rt,s\(w\)​\(k,ℓ\)\.r\_\{t,s\}^\{\(w\)\}\(j\\mid k,\\ell\)=\\frac\{R\_\{t,s\-1\}^\{\(w\)\}\(k,j\)P\_\{s\-1\}^\{\(w\)\}\(j,\\ell\)\}\{R\_\{t,s\}^\{\(w\)\}\(k,\\ell\)\}\.\(31\)The recursion forΞ\(w\)\\Xi^\{\(w\)\}is initialized by

Ξ\(w\)\(t,t;k,ℓ\)=𝟏\{k=ℓ\}Σt,k\(w\),\\Xi^\{\(w\)\}\(t,t;k,\\ell\)=\\mathbf\{1\}\_\{\\\{k=\\ell\\\}\}\\Sigma\_\{t,k\}^\{\(w\)\},\(32\)and, fors\>ts\>t,

Ξ\(w\)​\(t,s,k,ℓ\)=∑j=1Ks−1\(w\)rt,s\(w\)​\(j∣k,ℓ\)​Ξ\(w\)​\(t,s−1,k,j\)​\(As−1,j,ℓ\(w\)\)⊤\.\\Xi^\{\(w\)\}\(t,s;k,\\ell\)=\\sum\_\{j=1\}^\{K\_\{s\-1\}^\{\(w\)\}\}r\_\{t,s\}^\{\(w\)\}\(j\\mid k,\\ell\)\\,\\Xi^\{\(w\)\}\(t,s\-1;k,j\)\\left\(A\_\{s\-1,j,\\ell\}^\{\(w\)\}\\right\)^\{\\top\}\.\(33\)This recursion is evaluated for positive\-mass endpoint pairs\. Zero\-weight terms in the sum are omitted, and we setΞ\(w\)​\(t,s,k,ℓ\)=0\\Xi^\{\(w\)\}\(t,s;k,\\ell\)=0for zero\-mass endpoint pairs\.

###### Proof\.

The recursion forRt,s\(w\)R\_\{t,s\}^\{\(w\)\}follows from the Markov property\. Forr=t,…,s−1r=t,\\ldots,s\-1,

Rt,r\+1\(w\)​\(k,ℓ\)\\displaystyle R\_\{t,r\+1\}^\{\(w\)\}\(k,\\ell\)=Pr⁡\(Zr\+1\(w\)=ℓ∣Zt\(w\)=k\)\\displaystyle=\\Pr\\\!\\left\(Z\_\{r\+1\}^\{\(w\)\}=\\ell\\mid Z\_\{t\}^\{\(w\)\}=k\\right\)=∑j=1Kr\(w\)Pr⁡\(Zr\(w\)=j∣Zt\(w\)=k\)​Pr⁡\(Zr\+1\(w\)=ℓ∣Zr\(w\)=j\)\\displaystyle=\\sum\_\{j=1\}^\{K\_\{r\}^\{\(w\)\}\}\\Pr\\\!\\left\(Z\_\{r\}^\{\(w\)\}=j\\mid Z\_\{t\}^\{\(w\)\}=k\\right\)\\Pr\\\!\\left\(Z\_\{r\+1\}^\{\(w\)\}=\\ell\\mid Z\_\{r\}^\{\(w\)\}=j\\right\)=∑j=1Kr\(w\)Rt,r\(w\)​\(k,j\)​Pr\(w\)​\(j,ℓ\),\\displaystyle=\\sum\_\{j=1\}^\{K\_\{r\}^\{\(w\)\}\}R\_\{t,r\}^\{\(w\)\}\(k,j\)P\_\{r\}^\{\(w\)\}\(j,\\ell\),\(34\)which gives \([29](https://arxiv.org/html/2609.30974#A7.E29)\)\. Moreover,

q\(w\)​\(t,s,k,ℓ\)\\displaystyle q^\{\(w\)\}\(t,s;k,\\ell\)=Pr⁡\(Zt\(w\)=k\)​Pr⁡\(Zs\(w\)=ℓ∣Zt\(w\)=k\)\\displaystyle=\\Pr\\\!\\left\(Z\_\{t\}^\{\(w\)\}=k\\right\)\\Pr\\\!\\left\(Z\_\{s\}^\{\(w\)\}=\\ell\\mid Z\_\{t\}^\{\(w\)\}=k\\right\)=πt,k\(w\)​Rt,s\(w\)​\(k,ℓ\),\\displaystyle=\\pi\_\{t,k\}^\{\(w\)\}R\_\{t,s\}^\{\(w\)\}\(k,\\ell\),\(35\)which proves \([30](https://arxiv.org/html/2609.30974#A7.E30)\)\.

Next, by the definition of conditional probability,

rt,s\(w\)​\(j∣k,ℓ\)\\displaystyle r\_\{t,s\}^\{\(w\)\}\(j\\mid k,\\ell\)=Pr⁡\(Zs−1\(w\)=j∣Zt\(w\)=k\)​Pr⁡\(Zs\(w\)=ℓ∣Zs−1\(w\)=j,Zt\(w\)=k\)Pr⁡\(Zs\(w\)=ℓ∣Zt\(w\)=k\)\\displaystyle=\\frac\{\\Pr\\\!\\left\(Z\_\{s\-1\}^\{\(w\)\}=j\\mid Z\_\{t\}^\{\(w\)\}=k\\right\)\\Pr\\\!\\left\(Z\_\{s\}^\{\(w\)\}=\\ell\\mid Z\_\{s\-1\}^\{\(w\)\}=j,Z\_\{t\}^\{\(w\)\}=k\\right\)\}\{\\Pr\\\!\\left\(Z\_\{s\}^\{\(w\)\}=\\ell\\mid Z\_\{t\}^\{\(w\)\}=k\\right\)\}=Rt,s−1\(w\)​\(k,j\)​Ps−1\(w\)​\(j,ℓ\)Rt,s\(w\)​\(k,ℓ\),\\displaystyle=\\frac\{R\_\{t,s\-1\}^\{\(w\)\}\(k,j\)P\_\{s\-1\}^\{\(w\)\}\(j,\\ell\)\}\{R\_\{t,s\}^\{\(w\)\}\(k,\\ell\)\},\(36\)where the final equality follows from the Markov property\.

It remains to establish the recursion forΞ\(w\)\\Xi^\{\(w\)\}\. From the definition ofMt,s\(w\)M\_\{t,s\}^\{\(w\)\}in \([17](https://arxiv.org/html/2609.30974#A6.E17)\), fors\>ts\>t,

Mt,s\(w\)=Mt,s−1\(w\)​\(As−1,Zs−1\(w\),Zs\(w\)\(w\)\)⊤\.M\_\{t,s\}^\{\(w\)\}=M\_\{t,s\-1\}^\{\(w\)\}\\left\(A\_\{s\-1,Z\_\{s\-1\}^\{\(w\)\},Z\_\{s\}^\{\(w\)\}\}^\{\(w\)\}\\right\)^\{\\top\}\.\(37\)
For the initial case, we use the empty\-product conventionMt,t\(w\):=IdM\_\{t,t\}^\{\(w\)\}:=I\_\{d\}, so that

Ξ\(w\)​\(t,t,k,k\)=Σt,k\(w\)\.\\Xi^\{\(w\)\}\(t,t;k,k\)=\\Sigma\_\{t,k\}^\{\(w\)\}\.
Fors\>ts\>t, conditioning onZs−1\(w\)Z\_\{s\-1\}^\{\(w\)\}gives

𝔼\[Mt,s\(w\)∣Zt\(w\)=k,Zs\(w\)=ℓ\]\\displaystyle\\mathbb\{E\}\\\!\\left\[M\_\{t,s\}^\{\(w\)\}\\mid Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\}^\{\(w\)\}=\\ell\\right\]=∑j=1Ks−1\(w\)rt,s\(w\)\(j∣k,ℓ\)𝔼\[Mt,s−1\(w\)\|Zt\(w\)=k,Zs−1\(w\)=j,Zs\(w\)=ℓ\]\(As−1,j,ℓ\(w\)\)⊤\.\\displaystyle=\\sum\_\{j=1\}^\{K\_\{s\-1\}^\{\(w\)\}\}r\_\{t,s\}^\{\(w\)\}\(j\\mid k,\\ell\)\\mathbb\{E\}\\\!\\left\[M\_\{t,s\-1\}^\{\(w\)\}\\middle\|Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\-1\}^\{\(w\)\}=j,Z\_\{s\}^\{\(w\)\}=\\ell\\right\]\\left\(A\_\{s\-1,j,\\ell\}^\{\(w\)\}\\right\)^\{\\top\}\.\(38\)SinceMt,s−1\(w\)M\_\{t,s\-1\}^\{\(w\)\}depends only onZt\(w\),…,Zs−1\(w\)Z\_\{t\}^\{\(w\)\},\\ldots,Z\_\{s\-1\}^\{\(w\)\}, the Markov property gives

𝔼\[Mt,s−1\(w\)∣Zt\(w\)=k,Zs−1\(w\)=j,Zs\(w\)=ℓ\]\\displaystyle\\mathbb\{E\}\\\!\\left\[M\_\{t,s\-1\}^\{\(w\)\}\\mid Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\-1\}^\{\(w\)\}=j,Z\_\{s\}^\{\(w\)\}=\\ell\\right\]=𝔼\[Mt,s−1\(w\)∣Zt\(w\)=k,Zs−1\(w\)=j\]\.\\displaystyle\\qquad=\\mathbb\{E\}\\\!\\left\[M\_\{t,s\-1\}^\{\(w\)\}\\mid Z\_\{t\}^\{\(w\)\}=k,Z\_\{s\-1\}^\{\(w\)\}=j\\right\]\.\(39\)Multiplying \([38](https://arxiv.org/html/2609.30974#A7.E38)\) byΣt,k\(w\)\\Sigma\_\{t,k\}^\{\(w\)\}yields

Ξ\(w\)​\(t,s,k,ℓ\)\\displaystyle\\Xi^\{\(w\)\}\(t,s;k,\\ell\)=∑j=1Ks−1\(w\)rt,s\(w\)​\(j∣k,ℓ\)​Ξ\(w\)​\(t,s−1,k,j\)​\(As−1,j,ℓ\(w\)\)⊤,\\displaystyle=\\sum\_\{j=1\}^\{K\_\{s\-1\}^\{\(w\)\}\}r\_\{t,s\}^\{\(w\)\}\(j\\mid k,\\ell\)\\,\\Xi^\{\(w\)\}\(t,s\-1;k,j\)\\left\(A\_\{s\-1,j,\\ell\}^\{\(w\)\}\\right\)^\{\\top\},\(40\)which proves \([33](https://arxiv.org/html/2609.30974#A7.E33)\)\. ∎

## Appendix HStatistical recovery

This appendix studies how estimation errors propagate through the construction of the semantic geometry and establishes the error rates for the second\-moment operators and semantic distances under the Gaussian specialization given in Theorem[3\.7](https://arxiv.org/html/2609.30974#S3.Thmtheorem7)\. Figure[5](https://arxiv.org/html/2609.30974#A8.F5)illustrates the construction pipeline from estimated sense distributions and prevalences to the semantic distances\. Figure[6](https://arxiv.org/html/2609.30974#A8.F6)summarizes the error propagation from the estimated Gaussian mixture parameters to the estimated semantic geometry\. We first establish the estimation rate for the Gaussian mixture parameters and then propagate these errors through each step of the construction to obtain the convergence rate of the second\-moment operators and the induced distances\.

\(π^t,k\(w\),ν^t,k\(w\)\)\\bigl\(\\widehat\{\\pi\}\_\{t,k\}^\{\(w\)\},\\widehat\{\\nu\}\_\{t,k\}^\{\(w\)\}\\bigr\)\(η^t,k,l\(w\),C^t,k,l\(w\)\)\\bigl\(\\hat\{\\eta\}\_\{t,k,l\}^\{\(w\)\},\\hat\{C\}\_\{t,k,l\}^\{\(w\)\}\\bigr\)Q^t\(w\)\\hat\{Q\}\_\{t\}^\{\(w\)\}𝐌^∙\(w\)\\widehat\{\\mathbf\{M\}\}\_\{\\bullet\}^\{\(w\)\}d^∙,ρ\(w\)\\widehat\{d\}\_\{\\bullet,\\rho\}^\{\(w\)\}𝒖^r\\widehat\{\\bm\{u\}\}\_\{r\}Contextualcoupling \([1](https://arxiv.org/html/2609.30974#S2.E1)\)Componenttransport \([2](https://arxiv.org/html/2609.30974#S2.E2)\)when usingaggregate mean operatorFigure 5:Construction pipeline of the estimated semantic geometry\. The pipeline maps estimated sense distributions to geometric representations\.∙∈\{full,mean,dev\},ρ∈\{tr,1,…,d\}\\bullet\\in\\\{\\mathrm\{full\},\\mathrm\{mean\},\\mathrm\{dev\}\\\},\\rho\\in\\\{\\mathrm\{tr\},1,\\ldots,d\\\}\.\(π^t,k\(w\),μ^t,k\(w\),Σ^t,k\(w\)\)\\bigl\(\\widehat\{\\pi\}\_\{t,k\}^\{\(w\)\},\\widehat\{\\mu\}\_\{t,k\}^\{\(w\)\},\\widehat\{\\Sigma\}\_\{t,k\}^\{\(w\)\}\\bigr\)\(A^t,k,l\(w\),C^t,k,l\(w\)\)\\bigl\(\\widehat\{A\}\_\{t,k,l\}^\{\(w\)\},\\widehat\{C\}\_\{t,k,l\}^\{\(w\)\}\\bigr\)Q^t\(w\)\\widehat\{Q\}\_\{t\}^\{\(w\)\}𝐌^∙\(w\)\\widehat\{\\mathbf\{M\}\}\_\{\\bullet\}^\{\(w\)\}d^∙,ρ\(w\)\\widehat\{d\}\_\{\\bullet,\\rho\}^\{\(w\)\}𝒖^r\\widehat\{\\bm\{u\}\}\_\{r\}Lemma[H\.2](https://arxiv.org/html/2609.30974#A8.Thmtheorem2)Lemma[H\.3](https://arxiv.org/html/2609.30974#A8.Thmtheorem3)Lemma[H\.4](https://arxiv.org/html/2609.30974#A8.Thmtheorem4)Lemma[H\.1](https://arxiv.org/html/2609.30974#A8.Thmtheorem1)errorOp\(n−1/2\)O\_\{p\}\(n^\{\-1/2\}\)Theorem[3\.7](https://arxiv.org/html/2609.30974#S3.Thmtheorem7)errorOp\(n−1/2\)O\_\{p\}\(n^\{\-1/2\}\)Figure 6:Error propagation and recovery guarantees for the estimated semantic geometry\.∙∈\{full,mean,dev\},ρ∈\{tr,1,…,R\(w\)\}\\bullet\\in\\\{\\mathrm\{full\},\\mathrm\{mean\},\\mathrm\{dev\}\\\},\\rho\\in\\\{\\mathrm\{tr\},1,\\ldots,R^\{\(w\)\}\\\}\.#### Recovery of Gaussian mixture parameters\.

We assume throughout the theoretical analysis that the true number of mixture components is known\. This assumption does not impose a substantive restriction\. When a known finite upper boundKmaxK\_\{\\max\}on the number of components is available, an estimateK^t\(w\)\\widehat\{K\}\_\{t\}^\{\(w\)\}can be obtained by a penalized likelihood model\-selection procedure\. Under the regularity conditions in Definition[B\.2](https://arxiv.org/html/2609.30974#A2.Thmtheorem2), the model\-selection consistency established by[Huang et al\. \[37, Theorem 3\]](https://arxiv.org/html/2609.30974#bib.bib19)implies that

Pr⁡\(K^t\(w\)=Kt\(w\)\)→1\(n→∞\)\.\\Pr\\\!\\left\(\\widehat\{K\}\_\{t\}^\{\(w\)\}=K\_\{t\}^\{\(w\)\}\\right\)\\to 1\\qquad\(n\\to\\infty\)\.Thus, the true component number is recovered with probability tending to one as the sample size increases\. We therefore condition the subsequent analysis on the event of correct component\-number recovery and focus on the estimation error of the Gaussian mixture parameters and its propagation through the subsequent stages\.

#### Gaussian mixture parameter estimation\.

Taking the true component numbersKt\(w\)K\_\{t\}^\{\(w\)\}as known, we next consider the estimation of the Gaussian mixture parameters\. For a Gaussian mixture distribution

γ=∑k=1Kπk​νk,νk=𝒩⁡\(mk,Σk\),\\gamma=\\sum\_\{k=1\}^\{K\}\\pi\_\{k\}\\nu\_\{k\},\\qquad\\nu\_\{k\}=\\mathcal\{N\}\(m\_\{k\},\\Sigma\_\{k\}\),define

θk=\(mk⊤,σk\(i,j\):i=1,…,d,j=1,…,i\)⊤,\\theta\_\{k\}=\\left\(m\_\{k\}^\{\\top\},\\sigma\_\{k\}\(i,j\):i=1,\\ldots,d,\\;j=1,\\ldots,i\\right\)^\{\\top\},whereσk​\(i,j\)\\sigma\_\{k\}\(i,j\)denotes the\(i,j\)\(i,j\)\-th entry ofΣk\\Sigma\_\{k\}\. The full Gaussian mixture parameter vector is

θ=\(π1,…,πK,θ1⊤,…,θK⊤\)⊤\.\\theta=\(\\pi\_\{1\},\\ldots,\\pi\_\{K\},\\theta\_\{1\}^\{\\top\},\\ldots,\\theta\_\{K\}^\{\\top\}\)^\{\\top\}\.
The full vectorθ\\thetais retained for the componentwise parameter\-error bounds below\.

For the theoretical guarantee with fixed component counts, we use the penalized maximum likelihood estimator of[Chen and Tan \[21\]](https://arxiv.org/html/2609.30974#bib.bib7)\. To specify it, let

p​ℓ​\(θ\)=ℓn​\(θ\)\+pn​\(θ\),p\\ell\(\\theta\)=\\ell\_\{n\}\(\\theta\)\+p\_\{n\}\(\\theta\),whereℓn​\(θ\)\\ell\_\{n\}\(\\theta\)denotes the ordinary log\-likelihood based on the observed sample, and

pn\(θ\)=−1n∑k=1K\(tr\(Σk−1\)\+log\|Σk\|\)\.p\_\{n\}\(\\theta\)=\-\\frac\{1\}\{n\}\\sum\_\{k=1\}^\{K\}\\left\(\\operatorname\{tr\}\(\\Sigma\_\{k\}^\{\-1\}\)\+\\log\|\\Sigma\_\{k\}\|\\right\)\.Then, the estimator is

θ^∈arg⁡maxθ​p​ℓ​\(θ\)\.\\widehat\{\\theta\}\\in\\arg\\max\_\{\\theta\}p\\ell\(\\theta\)\.
Under the regularity conditions in Definition[B\.2](https://arxiv.org/html/2609.30974#A2.Thmtheorem2), Theorem 2 of\[[21](https://arxiv.org/html/2609.30974#bib.bib7)\]gives the asymptotic normality in a nonredundant local parameterization\. In particular, after aligning component labels, the full constrained parameter vector satisfies

‖θ^t\(w\)−θt\(w\)‖=Op\(n−1/2\)\.\\left\\\|\\widehat\{\\theta\}\_\{t\}^\{\(w\)\}\-\\theta\_\{t\}^\{\(w\)\}\\right\\\|=O\_\{p\}\(n^\{\-1/2\}\)\.
To quantify the parameter estimation errors, define

Eμ\(w\)\\displaystyle E\_\{\\mu\}^\{\(w\)\}:=maxt,k⁡‖μ^t,k\(w\)−μt,k\(w\)‖2,\\displaystyle:=\\max\_\{t,k\}\\left\\\|\\widehat\{\\mu\}\_\{t,k\}^\{\(w\)\}\-\\mu\_\{t,k\}^\{\(w\)\}\\right\\\|\_\{2\},EΣ\(w\)\\displaystyle E\_\{\\Sigma\}^\{\(w\)\}:=maxt,k⁡‖Σ^t,k\(w\)−Σt,k\(w\)‖2,\\displaystyle:=\\max\_\{t,k\}\\left\\\|\\widehat\{\\Sigma\}\_\{t,k\}^\{\(w\)\}\-\\Sigma\_\{t,k\}^\{\(w\)\}\\right\\\|\_\{2\},Eπ\(w\)\\displaystyle E\_\{\\pi\}^\{\(w\)\}:=maxt,k⁡\|π^t,k\(w\)−πt,k\(w\)\|\.\\displaystyle:=\\max\_\{t,k\}\\left\|\\widehat\{\\pi\}\_\{t,k\}^\{\(w\)\}\-\\pi\_\{t,k\}^\{\(w\)\}\\right\|\.
###### Lemma H\.1\(Gaussian mixture parameter recovery\)\.

For knownKK, under Definition[B\.2](https://arxiv.org/html/2609.30974#A2.Thmtheorem2)and the penalty and regularity conditions of[Chen and Tan \[21\]](https://arxiv.org/html/2609.30974#bib.bib7), the penalized estimator described above satisfies

Eμ\(w\),EΣ\(w\),Eπ\(w\)=Op\(n−1/2\)\.E\_\{\\mu\}^\{\(w\)\},E\_\{\\Sigma\}^\{\(w\)\},E\_\{\\pi\}^\{\(w\)\}=O\_\{p\}\(n^\{\-1/2\}\)\.

###### Proof\.

By the definition ofθt\(w\)\\theta\_\{t\}^\{\(w\)\}and the symmetry of the covariance matrices,

‖θ^t\(w\)−θt\(w\)‖2≥\\displaystyle\\left\\\|\\widehat\{\\theta\}\_\{t\}^\{\(w\)\}\-\\theta\_\{t\}^\{\(w\)\}\\right\\\|^\{2\}\\geq\{\}∑k=1Kt\(w\)\|π^t,k\(w\)−πt,k\(w\)\|2\\displaystyle\\sum\_\{k=1\}^\{K\_\{t\}^\{\(w\)\}\}\\left\|\\widehat\{\\pi\}\_\{t,k\}^\{\(w\)\}\-\\pi\_\{t,k\}^\{\(w\)\}\\right\|^\{2\}\+∑k=1Kt\(w\)‖μ^t,k\(w\)−μt,k\(w\)‖22\\displaystyle\+\\sum\_\{k=1\}^\{K\_\{t\}^\{\(w\)\}\}\\left\\\|\\widehat\{\\mu\}\_\{t,k\}^\{\(w\)\}\-\\mu\_\{t,k\}^\{\(w\)\}\\right\\\|\_\{2\}^\{2\}\+12∑k=1Kt\(w\)‖Σ^t,k\(w\)−Σt,k\(w\)‖22\.\\displaystyle\+\\frac\{1\}\{2\}\\sum\_\{k=1\}^\{K\_\{t\}^\{\(w\)\}\}\\left\\\|\\widehat\{\\Sigma\}\_\{t,k\}^\{\(w\)\}\-\\Sigma\_\{t,k\}^\{\(w\)\}\\right\\\|\_\{2\}^\{2\}\.Thus, for eachtt,

maxk\|π^t,k\(w\)−πt,k\(w\)\|,maxk‖μ^t,k\(w\)−μt,k\(w\)‖2,maxk‖Σ^t,k\(w\)−Σt,k\(w\)‖2=Op\(n−1/2\)\.\\max\_\{k\}\\left\|\\widehat\{\\pi\}\_\{t,k\}^\{\(w\)\}\-\\pi\_\{t,k\}^\{\(w\)\}\\right\|,\\max\_\{k\}\\left\\\|\\widehat\{\\mu\}\_\{t,k\}^\{\(w\)\}\-\\mu\_\{t,k\}^\{\(w\)\}\\right\\\|\_\{2\},\\max\_\{k\}\\left\\\|\\widehat\{\\Sigma\}\_\{t,k\}^\{\(w\)\}\-\\Sigma\_\{t,k\}^\{\(w\)\}\\right\\\|\_\{2\}=O\_\{p\}\(n^\{\-1/2\}\)\.Since the number of periods is fixed, taking the maximum overttgives

Eπ\(w\),Eμ\(w\),EΣ\(w\)=Op\(n−1/2\)\.E\_\{\\pi\}^\{\(w\)\},E\_\{\\mu\}^\{\(w\)\},E\_\{\\Sigma\}^\{\(w\)\}=O\_\{p\}\(n^\{\-1/2\}\)\.∎

#### Perturbation of Gaussian coupling maps and costs\.

Define the perturbation errors of the Gaussian coupling maps and transport costs by

EA\(w\):=maxt,k,ℓ⁡‖A^t,k,ℓ\(w\)−At,k,ℓ\(w\)‖2,E\_\{A\}^\{\(w\)\}:=\\max\_\{t,k,\\ell\}\\left\\\|\\widehat\{A\}\_\{t,k,\\ell\}^\{\(w\)\}\-A\_\{t,k,\\ell\}^\{\(w\)\}\\right\\\|\_\{2\},and

EC\(w\):=maxt,k,ℓ⁡\|C^t,k,ℓ\(w\)−Ct,k,ℓ\(w\)\|\.E\_\{C\}^\{\(w\)\}:=\\max\_\{t,k,\\ell\}\\left\|\\widehat\{C\}\_\{t,k,\\ell\}^\{\(w\)\}\-C\_\{t,k,\\ell\}^\{\(w\)\}\\right\|\.
###### Lemma H\.2\(Perturbation of Gaussian coupling maps and costs\)\.

Under the regularity conditions in Definition[B\.2](https://arxiv.org/html/2609.30974#A2.Thmtheorem2), asn→∞n\\to\\inftywithEμ\(w\),EΣ\(w\)→0E\_\{\\mu\}^\{\(w\)\},E\_\{\\Sigma\}^\{\(w\)\}\\to 0,

EA\(w\)=O⁡\(EΣ\(w\)\),E\_\{A\}^\{\(w\)\}=O\\\!\\left\(E\_\{\\Sigma\}^\{\(w\)\}\\right\),and

EC\(w\)=O⁡\(Eμ\(w\)\+EΣ\(w\)\)\.E\_\{C\}^\{\(w\)\}=O\\\!\\left\(E\_\{\\mu\}^\{\(w\)\}\+E\_\{\\Sigma\}^\{\(w\)\}\\right\)\.

###### Proof\.

We use the following perturbation bounds\. For matricesX,YX,Yand perturbationsΔ​X,Δ​Y\\Delta X,\\Delta Ysatisfying

‖Δ​X‖2,‖Δ​Y‖2≤E,\\\|\\Delta X\\\|\_\{2\},\\\|\\Delta Y\\\|\_\{2\}\\leq E,we have

‖\(X\+Δ​X\)​\(Y\+Δ​Y\)−X​Y‖2≤2​max⁡\{‖X‖2,‖Y‖2\}​E\+o⁡\(E\)\.\\\|\(X\+\\Delta X\)\(Y\+\\Delta Y\)\-XY\\\|\_\{2\}\\leq 2\\max\\\{\\\|X\\\|\_\{2\},\\\|Y\\\|\_\{2\}\\\}E\+o\(E\)\.ForX≻0X\\succ 0and‖Δ​X‖2≤E\\\|\\Delta X\\\|\_\{2\}\\leq E,

‖\(X\+Δ​X\)−1−X−1‖2\\displaystyle\\\|\(X\+\\Delta X\)^\{\-1\}\-X^\{\-1\}\\\|\_\{2\}≤1λmin​\(X\)2​E\+o⁡\(E\),\\displaystyle\\leq\\frac\{1\}\{\\lambda\_\{\\min\}\(X\)^\{2\}\}E\+o\(E\),‖\(X\+Δ​X\)1/2−X1/2‖2\\displaystyle\\\|\(X\+\\Delta X\)^\{1/2\}\-X^\{1/2\}\\\|\_\{2\}≤12​λmin​\(X\)1/2​E\+o⁡\(E\),\\displaystyle\\leq\\frac\{1\}\{2\\lambda\_\{\\min\}\(X\)^\{1/2\}\}E\+o\(E\),∥\(X\+ΔX\)−1/2−X−1/2∥2\\displaystyle\\\|\(X\+\\Delta X\)^\{\-1/2\}\-X^\{\-1/2\}\\\|\_\{2\}≤12​λmin​\(X\)3/2​E\+o⁡\(E\)\.\\displaystyle\\leq\\frac\{1\}\{2\\lambda\_\{\\min\}\(X\)^\{3/2\}\}E\+o\(E\)\.
Under Definition[B\.2](https://arxiv.org/html/2609.30974#A2.Thmtheorem2),

λmin​\(Σt,k\(w\)\)\>0\.\\lambda\_\{\\min\}\(\\Sigma\_\{t,k\}^\{\(w\)\}\)\>0\.Applying the above perturbation bounds successively to

At,k,ℓ\(w\)=\(Σt,k\(w\)\)−1/2\[\(Σt,k\(w\)\)1/2Σt\+1,ℓ\(w\)\(Σt,k\(w\)\)1/2\]1/2\(Σt,k\(w\)\)−1/2\.A\_\{t,k,\\ell\}^\{\(w\)\}=\\left\(\\Sigma\_\{t,k\}^\{\(w\)\}\\right\)^\{\-1/2\}\\Bigl\[\\left\(\\Sigma\_\{t,k\}^\{\(w\)\}\\right\)^\{1/2\}\\Sigma\_\{t\+1,\\ell\}^\{\(w\)\}\\left\(\\Sigma\_\{t,k\}^\{\(w\)\}\\right\)^\{1/2\}\\Bigr\]^\{1/2\}\\left\(\\Sigma\_\{t,k\}^\{\(w\)\}\\right\)^\{\-1/2\}\.gives

‖A^t,k,ℓ\(w\)−At,k,ℓ\(w\)‖2=O⁡\(‖Σ^t,k\(w\)−Σt,k\(w\)‖2\+‖Σ^t\+1,ℓ\(w\)−Σt\+1,ℓ\(w\)‖2\)\.\\left\\\|\\widehat\{A\}\_\{t,k,\\ell\}^\{\(w\)\}\-A\_\{t,k,\\ell\}^\{\(w\)\}\\right\\\|\_\{2\}=O\\\!\\left\(\\left\\\|\\widehat\{\\Sigma\}\_\{t,k\}^\{\(w\)\}\-\\Sigma\_\{t,k\}^\{\(w\)\}\\right\\\|\_\{2\}\+\\left\\\|\\widehat\{\\Sigma\}\_\{t\+1,\\ell\}^\{\(w\)\}\-\\Sigma\_\{t\+1,\\ell\}^\{\(w\)\}\\right\\\|\_\{2\}\\right\)\.SinceTTandKt\(w\)K\_\{t\}^\{\(w\)\}are fixed, taking the maximum overt,k,ℓt,k,\\ellpreserves the order, and hence

EA\(w\)=O⁡\(EΣ\(w\)\)\.E\_\{A\}^\{\(w\)\}=O\\\!\\left\(E\_\{\\Sigma\}^\{\(w\)\}\\right\)\.
Likewise, consider the transport cost

Ct,k,ℓ\(w\)=\\displaystyle C\_\{t,k,\\ell\}^\{\(w\)\}=\{\}‖μt,k\(w\)−μt\+1,ℓ\(w\)‖2\\displaystyle\\left\\\|\\mu\_\{t,k\}^\{\(w\)\}\-\\mu\_\{t\+1,\\ell\}^\{\(w\)\}\\right\\\|^\{2\}\+Tr⁡\(Σt,k\(w\)\+Σt\+1,ℓ\(w\)−2​\[\(Σt,k\(w\)\)1/2​Σt\+1,ℓ\(w\)​\(Σt,k\(w\)\)1/2\]1/2\)\.\\displaystyle\+\\operatorname\{Tr\}\\\!\\left\(\\Sigma\_\{t,k\}^\{\(w\)\}\+\\Sigma\_\{t\+1,\\ell\}^\{\(w\)\}\-2\\left\[\(\\Sigma\_\{t,k\}^\{\(w\)\}\)^\{1/2\}\\Sigma\_\{t\+1,\\ell\}^\{\(w\)\}\(\\Sigma\_\{t,k\}^\{\(w\)\}\)^\{1/2\}\\right\]^\{1/2\}\\right\)\.The squared mean term is Lipschitz in the mean parameters on the bounded parameter space, while the covariance terms are controlled by the above matrix perturbation bounds\. In particular, sinceddis fixed, the trace terms can be bounded in terms of the corresponding spectral norm perturbations\. Consequently,

\|C^t,k,ℓ\(w\)−Ct,k,ℓ\(w\)\|=O⁡\(CLOSE\\displaystyle\\left\|\\widehat\{C\}\_\{t,k,\\ell\}^\{\(w\)\}\-C\_\{t,k,\\ell\}^\{\(w\)\}\\right\|=O\\\!\\Big\(‖μ^t,k\(w\)−μt,k\(w\)‖2\+‖μ^t\+1,ℓ\(w\)−μt\+1,ℓ\(w\)‖2\\displaystyle\\left\\\|\\widehat\{\\mu\}\_\{t,k\}^\{\(w\)\}\-\\mu\_\{t,k\}^\{\(w\)\}\\right\\\|\_\{2\}\+\\left\\\|\\widehat\{\\mu\}\_\{t\+1,\\ell\}^\{\(w\)\}\-\\mu\_\{t\+1,\\ell\}^\{\(w\)\}\\right\\\|\_\{2\}OPEN\+‖Σ^t,k\(w\)−Σt,k\(w\)‖2\+‖Σ^t\+1,ℓ\(w\)−Σt\+1,ℓ\(w\)‖2\)\.\\displaystyle\+\\left\\\|\\widehat\{\\Sigma\}\_\{t,k\}^\{\(w\)\}\-\\Sigma\_\{t,k\}^\{\(w\)\}\\right\\\|\_\{2\}\+\\left\\\|\\widehat\{\\Sigma\}\_\{t\+1,\\ell\}^\{\(w\)\}\-\\Sigma\_\{t\+1,\\ell\}^\{\(w\)\}\\right\\\|\_\{2\}\\Big\)\.Again, sinceTTandKt\(w\)K\_\{t\}^\{\(w\)\}are fixed, taking the maximum overt,k,ℓt,k,\\ellpreserves the order\. Thus,

EC\(w\)=O⁡\(Eμ\(w\)\+EΣ\(w\)\)\.E\_\{C\}^\{\(w\)\}=O\\\!\\left\(E\_\{\\mu\}^\{\(w\)\}\+E\_\{\\Sigma\}^\{\(w\)\}\\right\)\.∎

#### Perturbation of the sense transport plans\.

We next consider the perturbation of the optimal sense transport plans induced by estimation errors in the sense prevalences and the transport costs\. Define

EQ\(w\):=maxt⁡‖Q^t\(w\)−Qt\(w\)‖F\.E\_\{Q\}^\{\(w\)\}:=\\max\_\{t\}\\left\\\|\\widehat\{Q\}\_\{t\}^\{\(w\)\}\-Q\_\{t\}^\{\(w\)\}\\right\\\|\_\{\\mathrm\{F\}\}\.
The perturbationEπ\(w\)E\_\{\\pi\}^\{\(w\)\}corresponds to the perturbation of the right\-hand side of the equality constraints in \([2](https://arxiv.org/html/2609.30974#S2.E2)\), whileEC\(w\)E\_\{C\}^\{\(w\)\}corresponds to the perturbation of its objective coefficients\.

The following lemma is based on a standard basis\-stability argument for linear programs\. The cost perturbationEC\(w\)E\_\{C\}^\{\(w\)\}does not appear in the leading term of the bound, but it is essential for ensuring that the support of the optimal transport planQt\(w\)Q\_\{t\}^\{\(w\)\}remains unchanged under small perturbations\. Once the support is fixed, the perturbation ofQt\(w\)Q\_\{t\}^\{\(w\)\}is determined solely by the perturbation of the marginal prevalences\.

###### Lemma H\.3\(Perturbation of the sense transport plans\)\.

Suppose that the optimal sense transport planQt\(w\)Q\_\{t\}^\{\(w\)\}is unique and satisfies

\|supp⁡\(Qt\(w\)\)\|=Kt\(w\)\+Kt\+1\(w\)−1\.\\left\|\\operatorname\{supp\}\\left\(Q\_\{t\}^\{\(w\)\}\\right\)\\right\|=K\_\{t\}^\{\(w\)\}\+K\_\{t\+1\}^\{\(w\)\}\-1\.Then, asn→∞n\\to\\inftywithEC\(w\),Eπ\(w\)→0E\_\{C\}^\{\(w\)\},E\_\{\\pi\}^\{\(w\)\}\\to 0,

EQ\(w\)=O⁡\(Eπ\(w\)\)\.E\_\{Q\}^\{\(w\)\}=O\\\!\\left\(E\_\{\\pi\}^\{\(w\)\}\\right\)\.\(41\)

###### Proof\.

For each consecutive period pair\(t,t\+1\)\(t,t\+1\), the equality constraints in \([2](https://arxiv.org/html/2609.30974#S2.E2)\) have rank

Kt\(w\)\+Kt\+1\(w\)−1,K\_\{t\}^\{\(w\)\}\+K\_\{t\+1\}^\{\(w\)\}\-1,since one of the marginal constraints is redundant\. Letbt\(w\)b\_\{t\}^\{\(w\)\}denote the right\-hand side obtained by removing one redundant marginal constraint and stacking the remaining marginal probabilities\.

LetAt\(w\)A\_\{t\}^\{\(w\)\}denote the corresponding constraint matrix, and letqt\(w\)q\_\{t\}^\{\(w\)\}denote the vector obtained by stacking all entries ofQt\(w\)Q\_\{t\}^\{\(w\)\}\. Then the equality constraints can be written as

At\(w\)​qt\(w\)=bt\(w\)\.A\_\{t\}^\{\(w\)\}q\_\{t\}^\{\(w\)\}=b\_\{t\}^\{\(w\)\}\.By the support condition,

\|supp⁡\(Qt\(w\)\)\|=Kt\(w\)\+Kt\+1\(w\)−1,\\left\|\\operatorname\{supp\}\\left\(Q\_\{t\}^\{\(w\)\}\\right\)\\right\|=K\_\{t\}^\{\(w\)\}\+K\_\{t\+1\}^\{\(w\)\}\-1,the columns ofAt\(w\)A\_\{t\}^\{\(w\)\}corresponding to the nonzero entries ofQt\(w\)Q\_\{t\}^\{\(w\)\}form a nonsingular square submatrix

Bt\(w\)∈ℝ\(Kt\(w\)\+Kt\+1\(w\)−1\)×\(Kt\(w\)\+Kt\+1\(w\)−1\)\.B\_\{t\}^\{\(w\)\}\\in\\mathbb\{R\}^\{\(K\_\{t\}^\{\(w\)\}\+K\_\{t\+1\}^\{\(w\)\}\-1\)\\times\(K\_\{t\}^\{\(w\)\}\+K\_\{t\+1\}^\{\(w\)\}\-1\)\}\.Thus, the nonzero entries ofqt\(w\)q\_\{t\}^\{\(w\)\}are given by

qt,B\(w\)=\(Bt\(w\)\)−1​bt\(w\)\.q\_\{t,B\}^\{\(w\)\}=\\left\(B\_\{t\}^\{\(w\)\}\\right\)^\{\-1\}b\_\{t\}^\{\(w\)\}\.
The uniqueness ofQt\(w\)Q\_\{t\}^\{\(w\)\}together with the support condition implies that the corresponding basic feasible solution is nondegenerate and that all nonbasic reduced costs are strictly positive\. Therefore, for sufficiently smallEC\(w\)E\_\{C\}^\{\(w\)\}andEπ\(w\)E\_\{\\pi\}^\{\(w\)\}, the same basisBt\(w\)B\_\{t\}^\{\(w\)\}remains optimal\. Consequently, the basic entries ofq^t\(w\)\\widehat\{q\}\_\{t\}^\{\(w\)\}are given by

q^t,B\(w\)=\(Bt\(w\)\)−1​b^t\(w\),\\widehat\{q\}\_\{t,B\}^\{\(w\)\}=\\left\(B\_\{t\}^\{\(w\)\}\\right\)^\{\-1\}\\widehat\{b\}\_\{t\}^\{\(w\)\},and hence

q^t,B\(w\)−qt,B\(w\)=\(Bt\(w\)\)−1​\(b^t\(w\)−bt\(w\)\)\.\\widehat\{q\}\_\{t,B\}^\{\(w\)\}\-q\_\{t,B\}^\{\(w\)\}=\\left\(B\_\{t\}^\{\(w\)\}\\right\)^\{\-1\}\\left\(\\widehat\{b\}\_\{t\}^\{\(w\)\}\-b\_\{t\}^\{\(w\)\}\\right\)\.
Since the nonbasic variables remain zero,

‖Q^t\(w\)−Qt\(w\)‖F\\displaystyle\\left\\\|\\widehat\{Q\}\_\{t\}^\{\(w\)\}\-Q\_\{t\}^\{\(w\)\}\\right\\\|\_\{\\mathrm\{F\}\}=‖q^t,B\(w\)−qt,B\(w\)‖2\\displaystyle=\\left\\\|\\widehat\{q\}\_\{t,B\}^\{\(w\)\}\-q\_\{t,B\}^\{\(w\)\}\\right\\\|\_\{2\}=‖\(Bt\(w\)\)−1​\(b^t\(w\)−bt\(w\)\)‖2\\displaystyle=\\left\\\|\\left\(B\_\{t\}^\{\(w\)\}\\right\)^\{\-1\}\\left\(\\widehat\{b\}\_\{t\}^\{\(w\)\}\-b\_\{t\}^\{\(w\)\}\\right\)\\right\\\|\_\{2\}≤‖\(Bt\(w\)\)−1‖2​‖b^t\(w\)−bt\(w\)‖2\\displaystyle\\leq\\left\\\|\\left\(B\_\{t\}^\{\(w\)\}\\right\)^\{\-1\}\\right\\\|\_\{2\}\\left\\\|\\widehat\{b\}\_\{t\}^\{\(w\)\}\-b\_\{t\}^\{\(w\)\}\\right\\\|\_\{2\}=O⁡\(Eπ\(w\)\)\.\\displaystyle=O\\\!\\left\(E\_\{\\pi\}^\{\(w\)\}\\right\)\.Taking the maximum overttyields

EQ\(w\)=O⁡\(Eπ\(w\)\)\.E\_\{Q\}^\{\(w\)\}=O\\\!\\left\(E\_\{\\pi\}^\{\(w\)\}\\right\)\.∎

#### Perturbation of the induced second\-moment operators\.

Finally, we propagate the preceding estimation errors to the induced second\-moment operators\. For∙∈\{full,mean,dev\}\\bullet\\in\\\{\\mathrm\{full\},\\mathrm\{mean\},\\mathrm\{dev\}\\\}, define

EM,∙\(w\):=maxt,s⁡‖𝐌^∙\(w\)​\(t,s\)−𝐌∙\(w\)​\(t,s\)‖2\.E\_\{M,\\bullet\}^\{\(w\)\}:=\\max\_\{t,s\}\\left\\\|\\widehat\{\\mathbf\{M\}\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)\-\\mathbf\{M\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)\\right\\\|\_\{2\}\.
###### Lemma H\.4\(Gaussian second\-moment operator perturbation\)\.

Asn→∞n\\to\\inftywith

EQ\(w\),Eμ\(w\),EΣ\(w\),EA\(w\)→0,E\_\{Q\}^\{\(w\)\},\\,E\_\{\\mu\}^\{\(w\)\},\\,E\_\{\\Sigma\}^\{\(w\)\},\\,E\_\{A\}^\{\(w\)\}\\to 0,we have

EM,mean\(w\)=O⁡\(EQ\(w\)\+Eμ\(w\)\),E\_\{M,\\mathrm\{mean\}\}^\{\(w\)\}=O\\\!\\left\(E\_\{Q\}^\{\(w\)\}\+E\_\{\\mu\}^\{\(w\)\}\\right\),EM,dev\(w\)=O⁡\(EQ\(w\)\+EΣ\(w\)\+EA\(w\)\),E\_\{M,\\mathrm\{dev\}\}^\{\(w\)\}=O\\\!\\left\(E\_\{Q\}^\{\(w\)\}\+E\_\{\\Sigma\}^\{\(w\)\}\+E\_\{A\}^\{\(w\)\}\\right\),and

EM,full\(w\)=O⁡\(EQ\(w\)\+Eμ\(w\)\+EΣ\(w\)\+EA\(w\)\)\.E\_\{M,\\mathrm\{full\}\}^\{\(w\)\}=O\\\!\\left\(E\_\{Q\}^\{\(w\)\}\+E\_\{\\mu\}^\{\(w\)\}\+E\_\{\\Sigma\}^\{\(w\)\}\+E\_\{A\}^\{\(w\)\}\\right\)\.

###### Proof\.

By \([8](https://arxiv.org/html/2609.30974#S3.E8)\), the mean and deviation operators are finite sums of terms weighted by the sense\-transition probabilities\. For the mean component,𝐂mean\(w\)​\(t,s,k,ℓ\)\\mathbf\{C\}\_\{\\mathrm\{mean\}\}^\{\(w\)\}\(t,s;k,\\ell\)is determined by the corresponding sense means\. Therefore, we obtain

EM,mean\(w\)=O⁡\(EQ\(w\)\+Eμ\(w\)\)\.E\_\{M,\\mathrm\{mean\}\}^\{\(w\)\}=O\\\!\\left\(E\_\{Q\}^\{\(w\)\}\+E\_\{\\mu\}^\{\(w\)\}\\right\)\.
For the deviation component, the covariance terms contributeO⁡\(EΣ\(w\)\)O\(E\_\{\\Sigma\}^\{\(w\)\}\)\. Moreover, the cross\-covariance terms involveΣt,k\(w\)\\Sigma\_\{t,k\}^\{\(w\)\}multiplied by conditional expectations of finite products of the coupling mapsA\(w\)A^\{\(w\)\}\. Since the number of periods is fixed, the perturbation of these products isO⁡\(EA\(w\)\)O\(E\_\{A\}^\{\(w\)\}\)\. Combining these perturbations with the perturbation of the transition probabilities gives

EM,dev\(w\)=O⁡\(EQ\(w\)\+EΣ\(w\)\+EA\(w\)\)\.E\_\{M,\\mathrm\{dev\}\}^\{\(w\)\}=O\\\!\\left\(E\_\{Q\}^\{\(w\)\}\+E\_\{\\Sigma\}^\{\(w\)\}\+E\_\{A\}^\{\(w\)\}\\right\)\.
Finally, by \([6](https://arxiv.org/html/2609.30974#S3.E6)\),

EM,full\(w\)=O⁡\(EQ\(w\)\+Eμ\(w\)\+EΣ\(w\)\+EA\(w\)\)\.E\_\{M,\\mathrm\{full\}\}^\{\(w\)\}=O\\\!\\left\(E\_\{Q\}^\{\(w\)\}\+E\_\{\\mu\}^\{\(w\)\}\+E\_\{\\Sigma\}^\{\(w\)\}\+E\_\{A\}^\{\(w\)\}\\right\)\.∎

#### Proof of distance statistical recovery\.

The preceding results establish the estimation rate for the Gaussian mixture parameters and propagate it through the coupling construction, the sense transport plans, and the induced second\-moment operators\. Combining these results with the second\-moment operator\-to\-distance perturbation bound of\[[16](https://arxiv.org/html/2609.30974#bib.bib11)\]yields the desired distance recovery rates\.

###### Proof of Theorem[3\.7](https://arxiv.org/html/2609.30974#S3.Thmtheorem7)\.

Under Assumptions[2\.1](https://arxiv.org/html/2609.30974#S2.Thmtheorem1)and[3\.6](https://arxiv.org/html/2609.30974#S3.Thmtheorem6), combining Lemma[H\.1](https://arxiv.org/html/2609.30974#A8.Thmtheorem1), Lemma[H\.2](https://arxiv.org/html/2609.30974#A8.Thmtheorem2), Lemma[H\.3](https://arxiv.org/html/2609.30974#A8.Thmtheorem3), and Lemma[H\.4](https://arxiv.org/html/2609.30974#A8.Thmtheorem4)yields

EM,∙\(w\)=Op\(\(n\(w\)\)−1/2\)\.E\_\{M,\\bullet\}^\{\(w\)\}=O\_\{p\}\\\!\\left\(\(n^\{\(w\)\}\)^\{\-1/2\}\\right\)\.
We next relate the second\-moment operator perturbation to the induced distance perturbation\. The same argument as in\[[16](https://arxiv.org/html/2609.30974#bib.bib11)\]applies, without the need for an orthogonal rotation between the population and estimated second\-moment operators\.

For the word\-local mean\-path basis, under Assumption[3\.3](https://arxiv.org/html/2609.30974#S3.Thmtheorem3), the same argument as in Proposition C\.4 of\[[16](https://arxiv.org/html/2609.30974#bib.bib11)\]yields

Eu,r\(w\):=minϵ∈\{−1,1\}⁡‖𝒖^r\(w\)−ϵ​𝒖r\(w\)‖2≤23/2​Tδ​EM,mean\(w\)E\_\{u,r\}^\{\(w\)\}:=\\min\_\{\\epsilon\\in\\\{\-1,1\\\}\}\\left\\\|\\widehat\{\\bm\{u\}\}\_\{r\}^\{\(w\)\}\-\\epsilon\\bm\{u\}\_\{r\}^\{\(w\)\}\\right\\\|\_\{2\}\\leq\\frac\{2^\{3/2\}T\}\{\\delta\}E\_\{M,\\mathrm\{mean\}\}^\{\(w\)\}for everyr∈\[R\(w\)\]r\\in\[R^\{\(w\)\}\]\. Hence,

Eu,r\(w\)=Op\(\(n\(w\)\)−1/2\),E\_\{u,r\}^\{\(w\)\}=O\_\{p\}\\\!\\left\(\(n^\{\(w\)\}\)^\{\-1/2\}\\right\),sinceT/δ=O⁡\(1\)T/\\delta=O\(1\)\.

Moreover, the same argument as in Proposition C\.7 of\[[16](https://arxiv.org/html/2609.30974#bib.bib11)\]yields

maxt,s∈\[T\]⁡\|d^∙,tr\(w\)​\(t,s\)2−d∙,tr\(w\)​\(t,s\)2\|≤d​EM,∙\(w\),\\max\_\{t,s\\in\[T\]\}\\left\|\\widehat\{d\}\_\{\\bullet,\\mathrm\{tr\}\}^\{\(w\)\}\(t,s\)^\{2\}\-d\_\{\\bullet,\\mathrm\{tr\}\}^\{\(w\)\}\(t,s\)^\{2\}\\right\|\\leq dE\_\{M,\\bullet\}^\{\(w\)\},and, for everyr∈\[R\(w\)\]r\\in\[R^\{\(w\)\}\],

maxt,s∈\[T\]⁡\|d^∙,r\(w\)​\(t,s\)2−d∙,r\(w\)​\(t,s\)2\|≤EM,∙\(w\)\+2​Eu,r\(w\)​maxt,s∈\[T\]​‖𝐌∙\(w\)​\(t,s\)‖\.\\max\_\{t,s\\in\[T\]\}\\left\|\\widehat\{d\}\_\{\\bullet,r\}^\{\(w\)\}\(t,s\)^\{2\}\-d\_\{\\bullet,r\}^\{\(w\)\}\(t,s\)^\{2\}\\right\|\\leq E\_\{M,\\bullet\}^\{\(w\)\}\+2E\_\{u,r\}^\{\(w\)\}\\max\_\{t,s\\in\[T\]\}\\\|\\mathbf\{M\}\_\{\\bullet\}^\{\(w\)\}\(t,s\)\\\|\.Since the population quantities are fixed, combining these bounds with

EM,∙\(w\)=Op\(\(n\(w\)\)−1/2\)E\_\{M,\\bullet\}^\{\(w\)\}=O\_\{p\}\\\!\\left\(\(n^\{\(w\)\}\)^\{\-1/2\}\\right\)gives

maxt,s∈\[T\]\|d^∙,ρ\(w\)\(t,s\)2−d∙,ρ\(w\)\(t,s\)2\|=Op\(\(n\(w\)\)−1/2\)\\max\_\{t,s\\in\[T\]\}\\left\|\\widehat\{d\}\_\{\\bullet,\\rho\}^\{\(w\)\}\(t,s\)^\{2\}\-d\_\{\\bullet,\\rho\}^\{\(w\)\}\(t,s\)^\{2\}\\right\|=O\_\{p\}\\\!\\left\(\(n^\{\(w\)\}\)^\{\-1/2\}\\right\)forρ∈\{tr,1,…,R\(w\)\}\\rho\\in\\\{\\mathrm\{tr\},1,\\ldots,R^\{\(w\)\}\\\}, which proves the theorem\. ∎

## Appendix ISupplementary experiments

The supplementary experiments follow the progression of Section[4](https://arxiv.org/html/2609.30974#S4)\. The synthetic study tests finite\-sample recovery of the Gaussian estimator\. DWUG provides standard LSC validation and examines the mean–deviation decomposition\. Janus isolates multi\-period profile recovery and compositional coherence\. The Court analysis documents the natural text corpus and shows how numerical transition and mode readouts lead to inspectable passages\. Each subsection supplies construction details, robustness results, or audit evidence for a claim made in the main text\.

### I\.1Synthetic mixture process and supplementary recovery

This experiment tests finite\-sample recovery under the Gaussian conditions of Theorem[3\.7](https://arxiv.org/html/2609.30974#S3.Thmtheorem7)\. It is a controlled check of the estimator rather than an additional application domain\.

Population and estimation\.We fix a three\-period process ind=2d=2with two components per period\. Prevalences and centers are

π1\\displaystyle\\pi\_\{1\}=\(\.55,\.45\),\\displaystyle=\(\.55,\.45\),μ1\\displaystyle\\mu\_\{1\}=\(−2\.402\.20\),\\displaystyle=\\begin\{pmatrix\}\-2\.4&0\\\\ 2\.2&0\\end\{pmatrix\},π2\\displaystyle\\pi\_\{2\}=\(\.40,\.60\),\\displaystyle=\(\.40,\.60\),μ2\\displaystyle\\mu\_\{2\}=\(−1\.9\.352\.7−\.25\),\\displaystyle=\\begin\{pmatrix\}\-1\.9&\.35\\\\ 2\.7&\-\.25\\end\{pmatrix\},π3\\displaystyle\\pi\_\{3\}=\(\.28,\.72\),\\displaystyle=\(\.28,\.72\),μ3\\displaystyle\\mu\_\{3\}=\(−1\.3\.703\.1\.45\)\.\\displaystyle=\\begin\{pmatrix\}\-1\.3&\.70\\\\ 3\.1&\.45\\end\{pmatrix\}\.The full covariance matrices are

tΣt,1Σt,21\(\.3000\.30\)\(\.3500\.25\)2\(1\.60\.20\.20\.40\)\(\.45−\.15−\.151\.80\)3\(3\.20\.35\.35\.50\)\(\.55\.20\.203\.50\)\.\\begin\{array\}\[\]\{c\|cc\}t&\\Sigma\_\{t,1\}&\\Sigma\_\{t,2\}\\\\ \\hline\\cr 1&\\left\(\\begin\{smallmatrix\}\.30&0\\\\ 0&\.30\\end\{smallmatrix\}\\right\)&\\left\(\\begin\{smallmatrix\}\.35&0\\\\ 0&\.25\\end\{smallmatrix\}\\right\)\\\\ 2&\\left\(\\begin\{smallmatrix\}1\.60&\.20\\\\ \.20&\.40\\end\{smallmatrix\}\\right\)&\\left\(\\begin\{smallmatrix\}\.45&\-\.15\\\\ \-\.15&1\.80\\end\{smallmatrix\}\\right\)\\\\ 3&\\left\(\\begin\{smallmatrix\}3\.20&\.35\\\\ \.35&\.50\\end\{smallmatrix\}\\right\)&\\left\(\\begin\{smallmatrix\}\.55&\.20\\\\ \.20&3\.50\\end\{smallmatrix\}\\right\)\.\\end\{array\}Thus the population contains prevalence shifts, center displacement, and anisotropic within component change\. Mean and deviation account for0\.8730\.873and0\.1270\.127of adjacent full path energy\.

For eachn∈\{100,200,400,800,1600\}n\\in\\\{100,200,400,800,1600\\\}, we drawnnindependent usages per period and repeat the experiment5050times\. We fit a full\-covariance GMM independently in every period withK=2K=2fixed, covariance regularization10−610^\{\-6\}, five initializations, and no access to sampled component labels\. FixingKKmatches Assumption[3\.6](https://arxiv.org/html/2609.30974#S3.Thmtheorem6)\. Because mixture labels are arbitrary, fitted components are aligned to the population components only when calculating errors\.

Errors and supplementary results\.We use

Eπ=maxt,k⁡\|π^t,k−πt,k\|,Eμ=maxt,k⁡‖μ^t,k−μt,k‖2,EΣ=maxt,k⁡‖Σ^t,k−Σt,k‖2\.E\_\{\\pi\}=\\max\_\{t,k\}\|\\widehat\{\\pi\}\_\{t,k\}\-\\pi\_\{t,k\}\|,\\quad E\_\{\\mu\}=\\max\_\{t,k\}\\\|\\widehat\{\\mu\}\_\{t,k\}\-\\mu\_\{t,k\}\\\|\_\{2\},\\quad E\_\{\\Sigma\}=\\max\_\{t,k\}\\\|\\widehat\{\\Sigma\}\_\{t,k\}\-\\Sigma\_\{t,k\}\\\|\_\{2\}\.For∙∈\{full,mean,dev\}\\bullet\\in\\\{\\mathrm\{full\},\\mathrm\{mean\},\\mathrm\{dev\}\\\}, operator error is

EM,∙=maxt,s⁡‖𝐌^∙​\(t,s\)−𝐌∙​\(t,s\)‖2\.E\_\{M,\\bullet\}=\\max\_\{t,s\}\\\|\\widehat\{\\mathbf\{M\}\}\_\{\\bullet\}\(t,s\)\-\\mathbf\{M\}\_\{\\bullet\}\(t,s\)\\\|\_\{2\}\.Squared\-distance error is the maximum absolute error over period pairs for either the trace or one of the two word\-local modes:

Ed,∙=maxt,s⁡\|d^∙,ρ​\(t,s\)2−d∙,ρ​\(t,s\)2\|\.E\_\{d,\\bullet\}=\\max\_\{t,s\}\\left\|\\widehat\{d\}\_\{\\bullet,\\rho\}\(t,s\)^\{2\}\-d\_\{\\bullet,\\rho\}\(t,s\)^\{2\}\\right\|\.The fitted mode distances use eigenvectors estimated from the fitted mean\-path operator\. Curves report means and normal approximation95%95\\%confidence intervals across repetitions\. Table[4](https://arxiv.org/html/2609.30974#A9.T4)estimates slopes by ordinary least squares on the five log mean errors\.

Table 4:Synthetic finite\-sample recovery\. The last column reports mean error atn=1600n=1600\.The operator, trace\-distance, and mode\-distance intervals all contain then−1/2n^\{\-1/2\}reference rate\. Parameter errors also decrease withnn, with the center and covariance curves somewhat shallower over this finite range\. Figure[7](https://arxiv.org/html/2609.30974#A9.F7)separates mean\- and deviation\-process distances\.

Figure 7:Squared\-distance recovery for the \(a\) mean and \(b\) deviation processes\. Trace and mode distances use the corresponding operator with the estimated word\-local basis\. Bands are95%95\\%confidence intervals over5050repetitions\.Pure mechanism checks\.We test the decomposition with three populations in which only centers, covariances, or prevalences change\. Center movement produces mean energy\. Changing prevalences between fixed components also transports mass between their centers and therefore produces mean energy\. The one\-component covariance process produces deviation energy\. Withn=1600n=1600usages per period and2020repetitions, Table[5](https://arxiv.org/html/2609.30974#A9.T5)shows that CUSP assigns99\.89%99\.89\\%,98\.95%98\.95\\%, and99\.95%99\.95\\%of path energy to the respective planted mechanisms\. The remaining1\.05%1\.05\\%mean allocation in the covariance\-only process is finite\-sample estimation error\.

Table 5:CUSP allocation of path energy in pure\-mechanism populations, in percent\. Each row changes only the stated population quantity\. Entries are means±\\pm95%95\\%confidence\-interval half\-widths over2020repetitions\. The theoretically correct term is bold\.
### I\.2DWUG implementation and decomposition

The English and German DWUG evaluations\[[22](https://arxiv.org/html/2609.30974#bib.bib33)\]use9,1079\{,\}107and9,1259\{,\}125XL\-LEXEME usage embeddings\[[24](https://arxiv.org/html/2609.30974#bib.bib5)\], respectively\. English contains4646targets with6161–100100usages per target–period\. German contains5050targets with3939–100100\. For every target and period, BIC selectsK∈\{1,…,5\}K\\in\\\{1,\\ldots,5\\\}and diagonal or full covariance, jointly with covariance regularization in\{10−6,10−5,10−4,10−3\}\\\{10^\{\-6\},10^\{\-5\},10^\{\-4\},10^\{\-3\}\\\}\. Components with fitted weight below0\.010\.01are excluded, and all fits use seed00\. The DWUG vectors retain1,0241\{,\}024coordinates after removal of the leading direction; no additional low\-dimensional PCA projection is applied\. Since the embedding dimension exceeds the number of usages per target–period, unregularized empirical full covariances are rank deficient\. Positive covariance regularization makes the fitted Gaussian covariances nonsingular, but does not remove the statistical uncertainty of estimating them in this regime\. The Gaussian recovery theorem assumes a correctly specified, fixed\-dimensional model and the stated estimation conditions\. In the lexical experiments, GMMs are regularized approximations: the decomposition remains exact for the fitted process, but the theorem does not guarantee recovery for normalized lexical data\.

For each language, we pool all target–period embeddings, fit one centering and principal\-direction transform, and hold it fixed across targets\. The primaryk=1k=1representation removes the leading pooled panel direction beforeℓ2\\ell\_\{2\}normalization\. Across the lexical experiments, PCA fitting uses random seed00, and preprocessing fitting and transformation use single\-threaded BLAS execution to control numerical variability\. All methods in Table[1](https://arxiv.org/html/2609.30974#S4.T1)use these same preprocessed embeddings\. CUSP selects its mixture configuration by unsupervised BIC; baseline parameters follow the fixed configurations supplied with the code\. These are different parameter\-selection rules, not matched hyperparameter\-search budgets\. Thek=0k=0andk=3k=3representations are sensitivity checks\. Confidence intervals in Table[1](https://arxiv.org/html/2609.30974#S4.T1)use3,0003\{,\}000paired bootstrap samples of target words, and Table[6](https://arxiv.org/html/2609.30974#A9.T6)reports all three preprocessing settings\.

Table 6:Spearman correlation with DWUG graded\-change judgments for the full, mean, and deviation traces\. Each transform is fitted once to the corresponding language panel\.The mean trace remains close to the full trace under every preprocessing choice\. The deviation trace is weaker but remains associated with the human score, and is comparatively large for heterogeneous words such as English*bar*and German*Rezeption*\. The decomposition therefore identifies within component reorganization that the scalar full trace would otherwise conceal\.

### I\.3Janus construction and supplementary results

[Cassotti and Tahmasebi \[23\]](https://arxiv.org/html/2609.30974#bib.bib6)generated the released Janus records with Llama 3 conditioned on a lemma, an Oxford English Dictionary sense definition, and a year\. The release contains131,811131\{,\}811sentences,13,76213\{,\}762sense–year records, and2,7682\{,\}768lemmas\. We select4848lemmas with four released sense pools and ten sentences per pool, yielding1,9201\{,\}920fixed sentences\. These populate six periods of4040usages per lemma by sampling with replacement\. The resulting11,52011\{,\}520usage slots are not independent new sentences\. Repeated contexts limit source diversity and can affect covariance estimates despite regularization\. Sampling with replacement does not itself violate independence, but the enlarged set of usage slots should not be interpreted as an equally large collection of distinct historical observations\. Resampling preserves any biases or irregularities in the released pools\. The experiment tests recovery of planted distributional changes and compositional coherence, not historical fidelity or recovery of the released sense identities\.

Stable schedules keep the four pool prevalences at\(\.25,\.25,\.25,\.25\)\(\.25,\.25,\.25,\.25\)\. In gradual schedules, the old and new pool prevalences follow\(\.90,\.72,\.54,\.36,\.18,0\)\(\.90,\.72,\.54,\.36,\.18,0\)and its reverse\. Abrupt schedules use\(\.90,\.90,\.90,0,0,0\)\(\.90,\.90,\.90,0,0,0\)and its reverse\. The two background pools have nominal prevalence\.05\.05in changed schedules, with finite counts allocated by largest remainder\. Sixteen lemmas are assigned to each schedule\. Released sense identifiers determine only these planted prevalences and are discarded before fitting\. CUSP receives neither sense identifiers nor cross\-period usage links\.

We fit one pooled panelk=1k=1transform to the original1,9201\{,\}920contexts and freeze it before constructing schedules\. A separate diagonal GMM is then fitted to every lemma–period panel, with BIC choosingK∈\{1,…,5\}K\\in\\\{1,\\ldots,5\\\}and covariance regularization in\{10−5,10−4\}\\\{10^\{\-5\},10^\{\-4\}\\\}subject to minimum component weight\.01\.01and at least three samples per component\.

To test temporal coherence, pairwise OT independently optimizes every period\-pair plan\[[5](https://arxiv.org/html/2609.30974#bib.bib23),[12](https://arxiv.org/html/2609.30974#bib.bib21)\]\. For each consecutive triple, we compare its direct planqt,t\+2q\_\{t,t\+2\}with the Markov composition ofqt,t\+1q\_\{t,t\+1\}andqt\+1,t\+2q\_\{t\+1,t\+2\}and average the resultingℓ1\\ell\_\{1\}disagreement\. CUSP defines the non\-adjacent plan by this composition, so its error is numerical zero by construction\. Table[7](https://arxiv.org/html/2609.30974#A9.T7)separates schedule recovery from the composition check\.

Table 7:Janus recovery with unsupervised period local GMMs and pooled panelk=1k=1preprocessing\. Composition error is meanℓ1\\ell\_\{1\}disagreement\.The pairwise inconsistency is concentrated in gradual replacement, where mass is redistributed over several adjacent steps\. Stable schedules and the single\-boundary abrupt schedules already compose to numerical precision\. Janus therefore does not establish uniform superiority over pairwise OT\. It shows accurate recovery of the planted temporal profiles and verifies the coherence added by CUSP when change unfolds over multiple periods\.

### I\.4Court corpus construction and fitting

The Court panel is built from the May 6, 2024 snapshot of CourtListener’s bulk opinion data, maintained by Free Law Project\[[27](https://arxiv.org/html/2609.30974#bib.bib12)\], and includes opinions filed from 1950 through 2024\. Excerpt identifiers are theidvalues of individual CourtListener opinion records in that snapshot\. The target vocabulary contains6060unigram legal lemmas selected from the tables of contents of West Academic law textbooks\.222[https://www\.westacademic\.com/](https://www.westacademic.com/)The vocabulary was fixed before model fitting and ranking, and the complete list is included with the released code\. The analysis uses eight decade bins from19501950through20202020\. The final bin contains opinions from20202020through the May 2024 snapshot, while the preceding bins cover ten years each\. Equalizing usage counts does not equalize calendar coverage\. The reported distances compare pooled period distributions rather than annualized rates of semantic change, so the shorter final bin should be interpreted with this limitation\.

Construction proceeds in five stages\. First, we locate target\-lemma matches and retain a window of approximately200200words on each side\. Up to5,0005\{,\}000candidate opinions are considered for each lemma–decade\. At most2020occurrences are scanned for an opinion–lemma pair, and at most six passages from that pair can be retained\. Second, a fixed LLM prompt receives the marked occurrence and its surrounding window\. It retains the occurrence only when the court is defining, applying, limiting, or otherwise construing the lemma as an object of legal reasoning\. Captions, repeated procedural text, proper nouns, ordinary uses, and incidental mentions are discarded\. The same prompt and decision rule are applied to every lemma and decade\. Consequently, the analysis concerns the selected legal\-reasoning usages, not changes in a word’s prevalence between ordinary discourse and legal reasoning or its unrestricted general\-language history\. The exact prompts are included with the released code\. The collection run usedgpt\-5\.6\-lunafor both this retention decision and the subsequent verbatim extraction\. It targeted up to240240distinct retained opinions per lemma–decade, with a minimum target of120120\. Third, a separate prompt copies one contiguous passage verbatim that contains the marked occurrence and completes the relevant reasoning\. It does not summarize or rewrite the opinion\. Fourth, XL\-LEXEME embeds the marked target token in that extracted passage\. Finally, ifNw,tN\_\{w,t\}is the retained count for lemmawwin decadett, we samplenw=mint⁡Nw,tn\_\{w\}=\\min\_\{t\}N\_\{w,t\}passages from every decade with seed00\. Each lemma therefore has equal counts across its eight periods, whilenwn\_\{w\}may differ between lemmas\.

The procedure yields322,222322\{,\}222passages before equalization, and the balanced panel contains258,480258\{,\}480\. The full balanced panel is used once to fit the Court centering and leading principal direction\. The same fittedk=1k=1transform is then applied to every lemma and decade\. For each lemma–decade, BIC selects a diagonal Gaussian mixture withK∈\{1,…,5\}K\\in\\\{1,\\ldots,5\\\}and minimum fitted weight\.02\.02\. Of the480480fits,471471selectK=5K=5, six selectK=4K=4, two selectK=3K=3, and one selectsK=2K=2\. Because BIC often selects the upper bound, the component cap determines the resolution of the fitted representation\. The resulting components describe recurring usage structure and are not claimed to recover a unique inventory of linguistic senses\. Alternative caps may change component assignments and the allocation between mean and deviation terms, while the decomposition remains exact for each fitted process\.

Across panel fitted ABTT settingsk=0,…,4k=0,\\ldots,4, a sensitivity analysis using a commonz≥1\.1z\\geq 1\.1screen places*privacy*among the top five lemmas by enrichment and the top two by secondary\-mode temporal separation\. Its largest change peaks in either1950→19601950\{\\rightarrow\}1960or1960→19701960\{\\rightarrow\}1970\. Thus its prominence and multi\-period character persist across these preprocessing choices, although the decade of its largest peak varies\.

The released materials include the exact retention and extraction prompts, fixed vocabulary, configuration files, frozen preprocessing and model outputs, and scripts for reproducing the reported geometry, rankings, passage retrieval, and figures\. Because the complete Court panel is large, the accompanying code package provides the frozen transform fitted on the full panel, processed inputs for the two main\-text lemmas, and numerical outputs for all 60 lemmas\. The released scripts reproduce the reported Court figures and rankings and rerun the main\-text analyses\. This subsection documents the corpus construction and fitting procedure for rebuilding the full panel from the cited CourtListener source\. The code and released results are available at[https://github\.com/hisanor013/cusp](https://github.com/hisanor013/cusp)\. The complete balanced Court text and embedding panels will be made publicly available in a separate data release\. Table[8](https://arxiv.org/html/2609.30974#A9.T8)summarizes the resulting corpus and balanced panel\.

Table 8:Court corpus and balanced analysis\-panel sizes\. Per\-decade counts are calculated after equalization within each lemma\.
### I\.5Court rankings and passage analysis

This subsection follows the Court analysis in the order in which every result is selected\. The within\-lemma standardized profile first identifies an adjacent decade that is unusually large relative to the lemma’s own history\. At that peak,EmaxE\_\{\\max\}identifies the component movement whose share of change is largest relative to its transported mass\. The word\-local analysis then asks whether the leading modes peak at different transitions\. Representative source and target passages make these numerically selected component movements inspectable\.

All numerical selections precede passage inspection\. After the analysis fixes a component movement, each usage is assigned to the fitted component with the highest posterior probability\. Representative passages within the selected source and target components are then ranked independently using a fixed score based on proximity to the component center and a surface form check favoring exact matches and forms beginning with the target lemma\. The retrieved passages provide representative source and target usages for the attributed movement, and CourtListener identifiers link every excerpt to its underlying opinion\. Named decisions situate the retrieved language historically\.

Peak transition attribution by maximum enrichment\.At a lemma’s selected peak, combine the mean and deviation contributions as

ak​ℓ=qk​ℓ​tr⁡\(𝐂mean​\(k,ℓ\)\+𝐂dev​\(k,ℓ\)\),αk​ℓ=ak​ℓ/dt2\.a\_\{k\\ell\}=q\_\{k\\ell\}\\operatorname\{tr\}\\\!\\left\(\\mathbf\{C\}\_\{\\mathrm\{mean\}\}\(k,\\ell\)\+\\mathbf\{C\}\_\{\\mathrm\{dev\}\}\(k,\\ell\)\\right\),\\qquad\\alpha\_\{k\\ell\}=a\_\{k\\ell\}/d\_\{t\}^\{2\}\.We define

Emax\(w\)=maxqk​ℓ\>0⁡αk​ℓqk​ℓ,E\_\{\\max\}^\{\(w\)\}=\\max\_\{q\_\{k\\ell\}\>0\}\\frac\{\\alpha\_\{k\\ell\}\}\{q\_\{k\\ell\}\},which identifies the component movement whose contribution share is most disproportionate to its transported\-mass share\. Sinceak​ℓ=qk​ℓ​tr⁡\(𝐂mean\+𝐂dev\)a\_\{k\\ell\}=q\_\{k\\ell\}\\operatorname\{tr\}\(\\mathbf\{C\}\_\{\\mathrm\{mean\}\}\+\\mathbf\{C\}\_\{\\mathrm\{dev\}\}\), the factorqk​ℓq\_\{k\\ell\}cancels in the ratio\. ThusEmaxE\_\{\\max\}measures displacement contribution per unit of transported mass, normalized by the total squared change; there is no uncancelled inverse\-mass factor\. Numerically, the implementation excludes pairs withqk​ℓ≤10−15q\_\{k\\ell\}\\leq 10^\{\-15\}and floors the total squared change at10−1510^\{\-15\}\. This is a per\-unit\-mass attribution statistic, not a significance measure, and rare component pairs can still have uncertain estimated displacements\. Because mixtures are fitted separately by period,kkandℓ\\ellare period local labels\. Equal or unequal indices have no intrinsic semantic interpretation\. Table[9](https://arxiv.org/html/2609.30974#A9.T9)reports the ten highest values\. The complete ranking of all5353lemmas passing this threshold and the corresponding figure are included in the supplementary material\.

Table 9:Ten highest maximum\-enrichment transitions under panel\-fittedk=1k=1\. Contributionα\\alphais the selected pair’s share of variation at the selected peak\.The following paragraphs examine the five highest ranked lemmas in Table[9](https://arxiv.org/html/2609.30974#A9.T9)\. This selection is determined by the numerical ranking rather than by the readability of the retrieved passages\. The complete ranking and component assignments are included in the supplementary material\.

Privacy\.The source says that “Professor Prosser classified these types of cases into four categories” \(opinion1179312\), while the target says that warrantless inspection poses “only a minimal threat to justifiable expectations of privacy” \(opinion1602980\)\. The pair moves from Prosser\-style tort classification\[[28](https://arxiv.org/html/2609.30974#bib.bib29)\]toward post\-*Katz*expectation\-of\-privacy language\[[29](https://arxiv.org/html/2609.30974#bib.bib20)\]\. The preceding decade’s leading pair moves from privacy as “a direct wrong of a personal character” \(opinion2604478\) toward the same Prosser component, making tort articulation, systematic classification, and expectation\-of\-privacy language successive rather than one retrospectively chosen contrast\.

Search\.The source applies a workers’ compensation rule requiring a “conscientious work search” and asks whether an unsuccessful search reflects inability to work \(opinion1139619\)\. Targets describe a “pat\-down search” and a “search of a jacket” \(opinion197788\)\. The transition is from*search*as a legally required effort to obtain work to physical inspection\. The source is retained because work search is itself part of the legal disability test, not incidental job\-seeking narration\.

Standing\.A source states that “Standing trees are a part of the realty” \(opinion1331324\)\. Targets ask whether a plaintiff has the “requisite standing to maintain the action” and describe standing as a federal\-court question under Article III \(opinions1513512,357908\)\. The pair separates the adjectival property sense from the jurisdictional doctrine\.

Warrant\.Sources state that facts “may warrant an inference of negligence” or “warrant an allowance of punitive damages” \(opinions2083592,1437656\)\. Targets discuss “probable cause for the issuance of a search warrant” and a “search warrant” \(opinions2216781,2167941\)\. Both pairs move from the verb licensing a conclusion to the noun denoting a judicial instrument\.

Trespass\.Sources plead a common\-law trespass claim or state that “a cognizable claim for trespass occurs when personal property … is injured or taken” \(opinions2421595,1557045\)\. Targets describe a “trespass\-to\-try\-title action” or an “action in trespass to try title” \(opinions4293854,1889942\)\. The retrieved language supports redistribution between distinct legal constructions of*trespass*, without assigning it to a landmark decision\.

Temporal separation of word\-local modes\.Letere\_\{r\}be moderr’s share of word\-local mean\-path energy\. Among the five modes with the largest path\-energy shares, order the two leading shares ase\(1\)≥e\(2\)e\_\{\(1\)\}\\geq e\_\{\(2\)\}and denote their peak decades byτ\(1\)\\tau\_\{\(1\)\}andτ\(2\)\\tau\_\{\(2\)\}\. For lemmas passing the standardized\-change threshold, define

S2=e\(2\)​log⁡\(1\+Emax\)S\_\{2\}=e\_\{\(2\)\}\\log\(1\+E\_\{\\max\}\)whenτ\(1\)≠τ\(2\)\\tau\_\{\(1\)\}\\neq\\tau\_\{\(2\)\}\. If the two leading modes peak together, the profile is classified as synchronized and receives noS2S\_\{2\}rank\. Multiplying the second mode’s energy share by a monotone transform ofEmaxE\_\{\\max\}favors histories with both a substantial second direction and a clearly attributable movement\. The score is fixed only to prioritize passage analyses of temporally separated histories\. It does not estimate the model, test significance, or determine whether change occurred\. Table[10](https://arxiv.org/html/2609.30974#A9.T10)reports the ten highest values\.

Table 10:Ten highest secondary\-mode temporal\-separation scores\.e\(1\)e\_\{\(1\)\}ande\(2\)e\_\{\(2\)\}are the two leading shares among the five highest\-energy mean\-path modes\.Privacy, standing, trespass, and search also appear in the enrichment ranking, where their selected transitions are grounded by the passage analyses above\. The high positions of*effect*and*taking*are consistent with broad, polysemous histories\. We retain their quantitative positions without assigning them a single landmark\-case narrative\. Complete values for all4141separated profiles are included in the supplementary material\.

Privacy modes\.Mode 2 \(30%30\\%\) peaks at1950→19601950\{\\rightarrow\}1960\. Its leading pair moves from a right of privacy “to be subordinated” to the public interest \(opinion1424064\) toward an “actionable invasion” of privacy, and the action “sounds in tort” \(opinion1741636\)\. Another target locates eavesdropping privacy in the Fourth Amendment \(opinion255989\)\. Modes 3 and 4 \(15%15\\%and11%11\\%\) peak at1960→19701960\{\\rightarrow\}1970and are led by the Prosser\-to\-expectations transition\. Mode 4 also connectsWarden v\. Hayden’s statement\[[38](https://arxiv.org/html/2609.30974#bib.bib16)\]that the Fourth Amendment protects the right of privacy, “rather than any interest in the property seized” \(opinion1982603\), to “no legitimate expectation of privacy” in a cargo hold \(opinion1378708\)\. Mode 1 \(39%39\\%\) peaks at1960→19701960\{\\rightarrow\}1970and remains strong at1970→19801970\{\\rightarrow\}1980\. At the later step, one pair moves from the “rights of privacy of practitioners of the healing arts and of their patients” \(opinion2597129\) toward statutory balancing of “personal privacy” against the “public interest” \(opinion1677943\)\. Another connects family privacy under*Griswold*and*Roe*\(opinion2142114\) to “no legitimate expectation of privacy” in plain view \(opinion2070690\)\[[39](https://arxiv.org/html/2609.30974#bib.bib15),[40](https://arxiv.org/html/2609.30974#bib.bib31)\]\. The modes therefore recover staggered and partly concurrent tort, constitutional, informational, and search\-privacy usages rather than one undifferentiated event\.

Expectation modes\.The three leading modes carry48%48\\%,26%26\\%, and18%18\\%of mean\-path energy and all peak at1960→19701960\{\\rightarrow\}1970\. Their deterministically selected passages move from “an expectation of pay or compensation” and contributions made “without expectation of services or direct benefits” \(opinions2226138,1636296\) toward “a reasonable expectation of privacy infringed by the search and seizure” and an area “protected by an expectation of privacy” \(opinions1365600,1311660\)\. The contrast with*privacy*therefore comes from a fixed temporal criterion rather than retrospective case selection\.

相似文章

基于变换与语义等价的认知过程动力学框架

arXiv cs.AI

本文提出了一种基于迭代状态变换和语义等价的认知过程建模的结构动力学框架,融合了动力系统、范畴论和反馈机制,将认知建模为朝着稳定解释演化的过程。