Foundations of Stochastic Lexical Calculus: Semantic Descent and Random Dynamics on Probability Simplices
Summary
This paper develops a stochastic lexical calculus framework for semantic updates in language models, defining conditions for when language-derived probabilities support meaningful sequential state representations. Empirical experiments validate the framework's stability and coverage under calibrated conditions.
View Cached Full Text
Cached at: 09/18/26, 09:12 AM
# Foundations of Stochastic Lexical Calculus Semantic Descent and Random Dynamics on Probability Simplices
Source: [https://arxiv.org/html/2609.20207](https://arxiv.org/html/2609.20207)
\(Working manuscript, July 2026\)
###### Abstract
Large language models produce prompt\-dependent probabilities over words, whereas scientific systems require uncertainty over meaningful states that can be updated as evidence arrives\. We develop an observable framework for determining when language\-derived probabilities support such a sequential state representation\. Theoretically, we define typed measurable transformations of contextual language, construct a minimal closed representation, and give necessary and sufficient conditions for semantic updates to exist uniquely\. We bound irreducible nonclosure and accumulated error, and under average contraction prove existence, uniqueness and stability of an external random recursion on a probability simplex\. These results define a stochastic lexical calculus without attributing an internal calculus to the language model\. Empirically, frozen experiments test the observable implications\. Raw prompt\-conditioned probabilities fail the prespecified invariance gate; after prompt\-specific calibration, a common three\-state representation passes the stability gates and covers2828of3030untouched eight\-step paths, or0\.9330\.933at nominal level0\.900\.90\. Accordingly, language probabilities support a stochastic state only conditionally on verified closure, stability and coverage within a declared operating domain\.
## 1The problem
Calculus begins by specifying objects, admissible changes, observables, and laws relating local change to composition and accumulation\. Natural language has the same need\. A word occurrence can be replaced, a qualification inserted, a clause negated, two passages concatenated, or a collection of claims reordered\. These operations act on contextual expressions rather than on isolated dictionary entries, and their order may alter meaning\.
Existing linguistic calculi primarily formalize grammatical derivability, typed composition, denotation, or symbolic rewriting\. Discrete calculus supplies general difference operators, but it does not determine which transformations preserve linguistic information, which contextual distinctions matter, or when a semantic representation remains closed under subsequent changes\. We address that missing layer\.
This gap has become operational rather than merely terminological\. A fitted language model assigns probabilities to lexical continuations, while users reason about coarser meanings such as factual alternatives, risk orientations, diagnoses, intentions, or confidence levels\. Recent work measures uncertainty by grouping semantically equivalent generations, asking models for confidence, calibrating candidate\-token probabilities, or training linguistically calibrated responses\[[3](https://arxiv.org/html/2609.20207#bib.bib3),[20](https://arxiv.org/html/2609.20207#bib.bib20),[25](https://arxiv.org/html/2609.20207#bib.bib25),[26](https://arxiv.org/html/2609.20207#bib.bib26),[27](https://arxiv.org/html/2609.20207#bib.bib27),[35](https://arxiv.org/html/2609.20207#bib.bib35)\]\. These methods establish that language\-derived uncertainty can be useful, but they leave a prior mathematical question unresolved: when does a probability\-valued summary of language retain exactly the distinctions required by later information transformations? Without such closure, successive summaries need not form a state process, even when each summary is individually well calibrated\.
The primitive object in this paper is contextual language\. Probability laws, model outputs, and declared semantic states are examples of lexical observables\. This ordering matters: the calculus is not defined by a particular language model, tokenization, ontology, or Bayesian posterior\. Those enter later as representations on which the general laws can be tested\.
The motivating question is simple\. If two contexts are treated as the same state, must every relevant linguistic intervention affect them in the same way? If not, the state has discarded transformation\-relevant lexical information\. No closed recursion on that state is justified\. This obstruction, and the canonical representation that removes it, organize the theory\.
For example, suppose two market reports are both summarized by the state probabilities\(0\.6,0\.3,0\.1\)\(0\.6,0\.3,0\.1\)for risk\-on, mixed and risk\-off\. In the first report, the favorable assessment is driven by falling inflation; in the second, it is driven by improving corporate earnings\. Appending the same new observation—“inflation unexpectedly rises”—may change the first assessment much more than the second\. The identical present state therefore does not determine the next state\. A recursion using only those three probabilities has omitted the distinction needed for the update\.
contextccevidence in one presentationtransformed contextTcTcrewording or new evidenceretained stateB\(c\)B\(c\)semantic probabilitiesnext stateB\(Tc\)B\(Tc\)semantic probabilitiestyped lexicaltransformationTTrepresentationBBrepresentationBBstate updateGTG\_\{T\}Figure 1:The central closure problem\. The upper path transforms the full context and then measures its retained semantic state\. A recursion on the lower state space is legitimate precisely when the dashed map is representative independent, so thatB\(Tc\)=GT\{B\(c\)\}B\(Tc\)=G\_\{T\}\\\{B\(c\)\\\}\. Rewordings that preserve the declared information should leave the state stable; genuine new evidence may move it, but must do so through the retained state if that state is to support autonomous sequential modelling\.Figure[1](https://arxiv.org/html/2609.20207#S1.F1)separates two operations that are often conflated\. A presentation change alters the words while preserving the declared information; an evidence update changes the information itself\. A useful semantic state should be stable to the former and responsive to the latter\. More strongly, if it is to be updated recursively, its response to every admitted transformation must be determined by the retained state rather than by lexical detail that the representation has already discarded\.
Why does this lead to a*stochastic lexical calculus*? In a deployed system, neither evidence nor the linguistic transformations through which it is presented arrive as a fixed deterministic sequence\. Reports, questions, qualifications and corrections arrive over time; their order and content are uncertain; and each may alter a probability distribution over the declared semantic states\. For three states, for example, the evolving quantity is a point\(p1,p2,p3\)\(p\_\{1\},p\_\{2\},p\_\{3\}\)on the probability simplex, not a single response\. Successive lexical transformations therefore generate a random path on that simplex\.
Calling this path a stochastic state process requires more than observing that its coordinates move\. The same retained state must imply the same successor state under a declared transformation; otherwise the apparent recursion still depends on lexical information that has been discarded\. Once this closure condition holds, finite lexical differences describe one\-step movement, cocycle laws describe composition, and contraction and perturbation bounds determine whether approximation error dissipates or accumulates\. These are the components of the stochastic lexical calculus developed here\.
The terminology is intentionally discrete and representation based\. We do not begin by imposing a continuous\-time diffusion, stochastic integral or Itô formula on language\. Rather, we derive a random dynamical system from typed contextual transformations and prove when it descends to a stable simplex\-valued recursion\. Continuous\-time limits, when scientifically appropriate, would come only after this elementary closure problem has been solved\.
### 1\.1Contributions and boundary of the claim
The term*calculus*has long been used in linguistics and logic\. Lambek calculus formalizes grammatical composition; lambda calculi support formal semantics; differential lambda calculus differentiates computational terms; Brzozowski derivatives act on formal languages; and discrete calculus supplies finite\-difference laws\. Earlier works have also used the phrases*calculus of words*,*lexical calculi*, and*calculus of language*\. We therefore make no claim to have introduced mathematical reasoning about words\.
The focal contribution is structural and stochastic, not merely lexical\.The problem lies at the intersection of several established mathematical structures in computer science: typed semantics specifies legal language transformations, partial\-map categories represent operations that are not defined on every context, behavioral equivalence determines which contexts may share a state, and probability kernels describe random evolution\. The deterministic transformation algebra is therefore a necessary foundation, but it is not the paper’s endpoint\. Our main question is when random arrivals of contextual evidence induce a well\-defined stochastic process on semantic probability simplices\. This requires a bridge that existing lexical calculi and generic stochastic systems do not generally supply: a measurable quotient must preserve enough future behavior for random transformations to descend, compose and remain stable through time\. The existence, causal uniqueness, synchronization and perturbation results for that descended process turn the lexical foundation into a stochastic lexical calculus and place language\-derived uncertainty within a compositional theory of state\-based computation\.
The closest modern lines of work are coalgebraic behavioral metrics and fibrational proof techniques\[[2](https://arxiv.org/html/2609.20207#bib.bib2),[7](https://arxiv.org/html/2609.20207#bib.bib7)\], restriction categories for partial computation\[[11](https://arxiv.org/html/2609.20207#bib.bib11),[12](https://arxiv.org/html/2609.20207#bib.bib12)\], categorical probability and Markov categories\[[21](https://arxiv.org/html/2609.20207#bib.bib21),[24](https://arxiv.org/html/2609.20207#bib.bib24)\], and renewed categorical treatments of natural\-language semantics\[[9](https://arxiv.org/html/2609.20207#bib.bib9),[13](https://arxiv.org/html/2609.20207#bib.bib13)\]\. These theories explain behavior, partiality, stochastic composition and linguistic structure at a high level of generality\. They do not by themselves specify which transformations of a contextual record preserve statistical information, how a probability law over verbal continuations is coarsened to an application state, or how failure of that coarsening is exposed by a commutation defect\. The present paper occupies that interface rather than claiming an alternative foundation to those literatures\.
The contribution claimed here is therefore more specific\. We formulate a stochastic lexical calculus in which:
1. 1\.the primitive domain consists of contextual lexical expressions;
2. 2\.typed transformations are the directions of change;
3. 3\.equivalence is operational and relative to declared observables and interventions;
4. 4\.the canonical state is constructed from future transform–observation behavior;
5. 5\.a proposed semantic representation is valid only when transformations descend through it;
6. 6\.random descended transformations define a simplex\-valued recursion with explicit existence and stability conditions; and
7. 7\.violations of the calculus laws are measurable rather than assumed away\.
To the best of our literature search, this combination—particularly terminal minimality for typed partial measurable language transformations and testable semantic descent—has not been developed for contextual natural language\. The constituent quotient, factorization, random\-iteration and finite\-difference arguments are standard\. Our contribution is their lexical construction and the resulting separation of three questions: what information a representation retains, whether admitted transformations descend through it, and whether the descended recursion is stable\.
The theoretical development is organized around three principal results\. First, the behavioral signature gives the smallest closed observable representation in the declared category\. Second, exact descent is equivalent to saturation of transformation domains and preservation of representation fibres; approximate descent has an irreducible fibre\-diameter error\. Third, once descent has been established, standard average\-contraction arguments yield a unique causal stochastic recursion and an explicit propagation bound for descent error\. Statistical identification and finite\-sample recovery are included only to make the observable implications testable; they are not presented as a new general theory of semiparametric estimation\. Nor do the results imply that a fitted language model literally contains a Bayesian belief state\.
The exposition follows the same order as the certification problem\. Section 2 defines contextual transformations, information equivalence and the canonical behavioral representation\. Section 3 develops semantic descent, error propagation and finite examples\. Section 4 constructs the stochastic lexical calculus on probability simplices\. Section 5 reports the completed bounded validation, including the gates that failed, and Section 6 concludes\. The appendices contain the categorical foundations, statistical recovery results and extended comparison with related mathematical structures\.
## 2Contextual transformation systems and information equivalence
### 2\.1Contextual lexical spaces and transformations
Before asking whether a semantic state is stable, we must say what is being changed and what can be observed\. This is less trivial for language than for a vector in Euclidean space\. Replacing a number in a structured record, reordering two sentences, and inserting a negation are all string operations, but they have different scientific roles\. The first may change evidence, the second may preserve it, and the third may reverse meaning\. We therefore make the transformation type and its legal domain part of the mathematical object\.
For example, consider the record “temperature is 39 C; oxygen saturation is falling\.” Replacing 39 by 37 changes the measured evidence\. Reversing the order of the two clauses can preserve the same evidence\. Inserting “not” before “falling” changes the second observation to its negation\. Although each operation edits a string, only the reordering is naturally treated as information preserving for this example\.
With these distinctions in place, we now formalize the contextual language space, its observable outputs, and the typed transformations that act upon it\. LetΣ\\Sigmabe a finite or countable vocabulary and letΣ∗\\Sigma^\{\*\}denote the free monoid of finite strings under concatenation\. The admissible context space𝒞⊆Σ∗\\mathcal\{C\}\\subseteq\\Sigma^\{\*\}is equipped with a sigma\-algebra𝒜\\mathcal\{A\}\. The continuation space𝒲⊆Σ∗\\mathcal\{W\}\\subseteq\\Sigma^\{\*\}has sigma\-algebra𝒢\\mathcal\{G\}\.
An isolated word type is generally not a sufficient unit: the occurrence of*strong*in “strong earnings” differs from its occurrence in “strong inflation\.” We therefore take a lexical occurrence to be a triple\(c,i,wi\)\(c,i,w\_\{i\}\)consisting of a context, a location, and the expression occupying that location\. Phrase spans and structured evidence records are included by allowingiito index a finite interval or a typed field\.
###### Definition 2\.1\(Contextual lexical system\)\.
A contextual lexical system is a tuple
𝔏=\(Σ,𝒞,𝔗,𝒪,≃\),\\mathfrak\{L\}=\(\\Sigma,\\mathcal\{C\},\\mathfrak\{T\},\\mathcal\{O\},\\simeq\),where𝔗\\mathfrak\{T\}is a category of typed partial transformations of contexts,𝒪\\mathcal\{O\}is a family of measurable lexical observables, and≃\\simeqis a declared information or semantic equivalence relation\. An observable is a mapF:𝒞→VFF:\\mathcal\{C\}\\to V\_\{F\}into a specified measurable, metric, algebraic, or normed codomain\.
Examples of observables include the presence of a phrase, the provenance record recovered from a passage, a formal denotation, a distribution over possible continuations, a semantic orientation, or a probability composition over declared states\. The theory does not require every observable to be probabilistic\.
###### Definition 2\.2\(Primitive lexical edits\)\.
For a contextual occurrence or span, primitive edits include typed substitutionRu→vR\_\{u\\to v\}, insertionIvI\_\{v\}, deletionDuD\_\{u\}, permutationPπP\_\{\\pi\}, qualificationQvQ\_\{v\}, and negationNN\. Their legal domains are part of their definitions\. Composite transformations are finite well\-typed paths of primitive edits\.
The category formulation prevents arbitrary string manipulation from being treated as meaningful calculus\. For example, a provenance\-preserving reordering and a chronology\-reversing reordering have different types even when both are permutations of the same words\.
### 2\.2Probability\-valued lexical observables
The following important representation is a special case, not the definition of the calculus\.
###### Definition 2\.3\(Observable lexical kernel\)\.
An observable lexical kernel is a Markov kernel
𝖪:𝒞×𝒢⟶\[0,1\],\\mathsf\{K\}:\\mathcal\{C\}\\times\\mathcal\{G\}\\longrightarrow\[0,1\],such thatA↦𝖪\(c,A\)A\\mapsto\\mathsf\{K\}\(c,A\)is a probability measure for everyc∈𝒞c\\in\\mathcal\{C\}andc↦𝖪\(c,A\)c\\mapsto\\mathsf\{K\}\(c,A\)is measurable for everyA∈𝒢A\\in\\mathcal\{G\}\. We write𝖪c:=𝖪\(c,⋅\)∈𝒫\(𝒲\)\\mathsf\{K\}\_\{c\}:=\\mathsf\{K\}\(c,\\cdot\)\\in\\mathcal\{P\}\(\\mathcal\{W\}\)\.
The kernel is the externally observable conditional distribution over verbal continuations\. An observation mechanism may expose only a coarsening or truncation of this law; that case is represented by a separate observation operator and must not be silently identified with𝖪c\\mathsf\{K\}\_\{c\}\.
###### Definition 2\.4\(Semantic map\)\.
ForK≥2K\\geq 2, a semantic map is a measurable function
Π:𝒫\(𝒲\)⟶ΔK−1,ΔK−1:=\{𝒑∈\[0,1\]K:𝟏⊤𝒑=1\}\.\\Pi:\\mathcal\{P\}\(\\mathcal\{W\}\)\\longrightarrow\\Delta^\{K\-1\},\\qquad\\Delta^\{K\-1\}:=\\\{\\bm\{p\}\\in\[0,1\]^\{K\}:\\bm\{1\}^\{\\top\}\\bm\{p\}=1\\\}\.The observable semantic state is
𝐁:=Π∘𝖪:𝒞⟶ΔK−1\.\\mathbf\{B\}:=\\Pi\\circ\\mathsf\{K\}:\\mathcal\{C\}\\longrightarrow\\Delta^\{K\-1\}\.
For a measurable partitionA1,…,AKA\_\{1\},\\ldots,A\_\{K\}of𝒲\\mathcal\{W\}, the simplest pushforward is
Π\(μ\)=\(μ\(A1\),…,μ\(AK\)\)\.\\Pi\(\\mu\)=\\bigl\(\\mu\(A\_\{1\}\),\\ldots,\\mu\(A\_\{K\}\)\\bigr\)\.Calibrated maps may be nonlinear; none of the descent results below requires linearity\.
###### Definition 2\.5\(Typed lexical transformation\)\.
A lexical transformation is a measurable partial mapT:dom\(T\)⊆𝒞→𝒞T:\\operatorname\{dom\}\(T\)\\subseteq\\mathcal\{C\}\\to\\mathcal\{C\}supplied with a declared type, such as insertion, deletion, substitution, negation, permutation, paraphrase, or evidence update\. A collection𝔗\\mathfrak\{T\}is composition\-closed whenever domains are compatible and contains the identity transformation\.
The type is part of the experimental intervention\. The theory does not infer from the model alone whether a transformation is a paraphrase or a contradiction\.
### 2\.3A concrete foundation for information equivalence
The word “equivalent” is dangerous unless the experiment relative to which it is asserted has been fixed\. For example, “temperature 39 C; oxygen falling” and “oxygen falling; temperature 39 C” may encode the same clinical record, although a fitted language system can assign different continuation probabilities to the two strings\. By contrast, replacing “falling” with “stable” changes the record\. Similarity of embeddings or agreement of modal responses cannot distinguish these cases by itself\.
We therefore begin with Blackwell’s comparison of statistical experiments\[[5](https://arxiv.org/html/2609.20207#bib.bib5)\]and then move from experiment\-level equivalence to a realized\-record criterion that can be constructed in data\. This places prompt perturbations within a familiar decision\-theoretic framework: two presentations count as information equivalent only when one contains no state\-relevant information absent from the other\.
Information equivalence cannot mean that two passages have similar embeddings, share vocabulary, or receive the same label from a language model\. It must be defined relative to a statistical question\. Let\(𝒳,ℋ\)\(\\mathcal\{X\},\\mathcal\{H\}\)be a measurable state space and let an information experiment be a Markov kernel
𝖤:𝒳×𝒴⟶\[0,1\],\\mathsf\{E\}:\\mathcal\{X\}\\times\\mathcal\{Y\}\\longrightarrow\[0,1\],where𝖤\(x,⋅\)\\mathsf\{E\}\(x,\\cdot\)is the distribution of an observable recordYYunder statexx\. A rendering mapr:𝒴→𝒞r:\\mathcal\{Y\}\\to\\mathcal\{C\}converts a structured record into language while retaining its source identifiers and values\.
###### Definition 2\.6\(Blackwell information equivalence\[[5](https://arxiv.org/html/2609.20207#bib.bib5)\]\)\.
Two experiments𝖤1:𝒳↝𝒴1\\mathsf\{E\}\_\{1\}:\\mathcal\{X\}\\rightsquigarrow\\mathcal\{Y\}\_\{1\}and𝖤2:𝒳↝𝒴2\\mathsf\{E\}\_\{2\}:\\mathcal\{X\}\\rightsquigarrow\\mathcal\{Y\}\_\{2\}are information equivalent, written𝖤1≃B𝖤2\\mathsf\{E\}\_\{1\}\\simeq\_\{\\mathrm\{B\}\}\\mathsf\{E\}\_\{2\}, if there exist state\-independent Markov kernelsQ12Q\_\{12\}andQ21Q\_\{21\}such that
𝖤2=𝖤1Q12,𝖤1=𝖤2Q21\.\\mathsf\{E\}\_\{2\}=\\mathsf\{E\}\_\{1\}Q\_\{12\},\\qquad\\mathsf\{E\}\_\{1\}=\\mathsf\{E\}\_\{2\}Q\_\{21\}\.
This is the standard equivalence induced by Blackwell’s comparison of statistical experiments: each observation can be simulated from the other without access to the state\. Consequently the two experiments have the same attainable Bayes risk for every bounded decision problem\.
For a fixed realized record, a more operational criterion is available\. Suppose𝒳=\{1,…,K\}\\mathcal\{X\}=\\\{1,\\ldots,K\\\}and the experiment is dominated by a measureν\\nu, with likelihoodex\(y\)e\_\{x\}\(y\)\. Define the likelihood ray
\[𝒆\(y\)\]:=\{a\(e1\(y\),…,eK\(y\)\):a\>0\}\.\[\\bm\{e\}\(y\)\]:=\\\{a\(e\_\{1\}\(y\),\\ldots,e\_\{K\}\(y\)\):a\>0\\\}\.
###### Definition 2\.7\(Realized information equivalence\)\.
Two recordsy,y′y,y^\{\\prime\}are likelihood equivalent, writteny≃Ly′y\\simeq\_\{\\mathrm\{L\}\}y^\{\\prime\}, when
ex\(y′\)=a\(y,y′\)ex\(y\)for everyx∈𝒳e\_\{x\}\(y^\{\\prime\}\)=a\(y,y^\{\\prime\}\)e\_\{x\}\(y\)\\quad\\text\{for every \}x\\in\\mathcal\{X\}for some scalara\(y,y′\)\>0a\(y,y^\{\\prime\}\)\>0independent ofxx\.
###### Proposition 2\.8\(Posterior characterization; cf\. Blackwell equivalence\[[5](https://arxiv.org/html/2609.20207#bib.bib5)\]\)\.
For strictly positive likelihood vectors,y≃Ly′y\\simeq\_\{\\mathrm\{L\}\}y^\{\\prime\}if and only if the Bayesian posterior afteryyequals the posterior aftery′y^\{\\prime\}for every full\-support prior on𝒳\\mathcal\{X\}\.
###### Proof\.
Proportional likelihoods cancel to the same normalized posterior\. Conversely, equality under every full\-support prior implies equality of every pairwise posterior odds; henceex\(y\)/ej\(y\)=ex\(y′\)/ej\(y′\)e\_\{x\}\(y\)/e\_\{j\}\(y\)=e\_\{x\}\(y^\{\\prime\}\)/e\_\{j\}\(y^\{\\prime\}\)for allx,jx,j, which is equivalent to proportionality\. ∎
###### Definition 2\.9\(Provenance\-preserving rendering equivalence\)\.
Letyybe a structured evidence record carrying values, units, times, sources, and target definitions\. Two contextsc=r\(y\)c=r\(y\)andc′=r′\(y\)c^\{\\prime\}=r^\{\\prime\}\(y\)are provenance\-preserving renderings of the same record when both renderings are injective on the declared information fields and a deterministic audit map recoversyyfrom either context\. We writec≃Pc′c\\simeq\_\{\\mathrm\{P\}\}c^\{\\prime\}\.
Provenance\-preserving equivalence is deliberately stronger than ordinary paraphrase\. It gives an experimentally constructible subset of Blackwell\-equivalent presentations: both strings deterministically encode the same statistical record\. Evidence order, bullet versus prose format, and equivalent numerical notation can be tested this way without claiming that unrestricted natural\-language paraphrases are automatically equivalent\.
Let≃ℐ\\simeq\_\{\\mathcal\{I\}\}denote the equivalence relation selected for an experiment, whereℐ\\mathcal\{I\}records the target state, decision class, horizon, and provenance rules\. The quotient
𝒞ℐ:=𝒞/≃ℐ\\mathcal\{C\}\_\{\\mathcal\{I\}\}:=\\mathcal\{C\}/\{\\simeq\_\{\\mathcal\{I\}\}\}is the information\-context space\. The quotient is therefore task relative, not a universal partition of language\.
###### Definition 2\.10\(Information\-preserving transformation\)\.
A transformationTTis information preserving on𝒞0\\mathcal\{C\}\_\{0\}when
c≃ℐTcfor everyc∈𝒞0\.c\\simeq\_\{\\mathcal\{I\}\}Tc\\quad\\text\{for every \}c\\in\\mathcal\{C\}\_\{0\}\.It is provenance preserving when the equivalence can be certified by recovery of the same structured record\.
This separates two empirical questions\. WhetherTTpreserves information is certified from the data\-generating experiment and provenance; whether the language service preserves the corresponding semantic state is tested through𝐁\(Tc\)≈𝐁\(c\)\\mathbf\{B\}\(Tc\)\\approx\\mathbf\{B\}\(c\)\.
#### 2\.3\.1Lexical observational equivalence
Blackwell and likelihood equivalence require a specified statistical experiment\. A lexical calculus also needs an internal analogue that applies to nonprobabilistic observables\.
###### Definition 2\.11\(Observational lexical equivalence\)\.
For observables𝒪0⊆𝒪\\mathcal\{O\}\_\{0\}\\subseteq\\mathcal\{O\}and transformations𝔗0⊆𝔗\\mathfrak\{T\}\_\{0\}\\subseteq\\mathfrak\{T\}, define
c≃𝒪0,𝔗0c′⟺F\(Tc\)=F\(Tc′\)c\\simeq\_\{\\mathcal\{O\}\_\{0\},\\mathfrak\{T\}\_\{0\}\}c^\{\\prime\}\\quad\\Longleftrightarrow\\quad F\(Tc\)=F\(Tc^\{\\prime\}\)for everyF∈𝒪0F\\in\\mathcal\{O\}\_\{0\}and everyT∈𝔗0T\\in\\mathfrak\{T\}\_\{0\}for which both sides are defined\.
Two contexts are lexically equivalent when no admitted observable, either now or after an admitted intervention, distinguishes them\. This is relative to a declared experimental vocabulary\. If𝒪0\\mathcal\{O\}\_\{0\}contains all bounded decision losses generated by an experiment, it recovers decision\-theoretic equivalence\. If it contains all posterior\-odds observables, it recovers likelihood\-ray equivalence\. If it contains an auditable record\-recovery map, it respects provenance equivalence\.
###### Proposition 2\.12\(Compatibility hierarchy\)\.
Letc=r\(y\)c=r\(y\)andc′=r′\(y′\)c^\{\\prime\}=r^\{\\prime\}\(y^\{\\prime\}\)be renderings of observations from a dominated finite\-state experiment\.
1. 1\.Provenance equivalence of renderings of the same record implies likelihood equivalence\.
2. 2\.Likelihood equivalence implies equality of every Bayesian posterior observable for every full\-support prior\.
3. 3\.Blackwell\-equivalent experiments induce the same attainable risk for every bounded decision problem\.
4. 4\.None of the converses holds without additional injectivity or richness assumptions\.
###### Proof\.
The first statement follows because the recovered record and hence its likelihood vector are identical\. The second is Proposition[2\.8](https://arxiv.org/html/2609.20207#S2.Thmdefinition8)\. The third is Blackwell’s comparison theorem\. Coarsened renderings and restricted decision or observable classes give counterexamples to the converses\. ∎
Observational lexical equivalence therefore does not compete with Blackwell equivalence\. It extends the same operational principle to a chosen family of contextual transformations and lexical observables\.
### 2\.4The canonical lexical representation
The foundational problem is to construct the least informative representation that nevertheless supports all selected lexical observations after all selected transformations\.
This is best understood as a memory question\. Suppose two contexts produce the same response today\. They should be merged only if no admissible sequence of later interventions can make the selected observations distinguish them\. A state that remembers less is not closed; a state that remembers more carries unnecessary lexical detail\. This is the future\-behavior principle behind Myhill–Nerode minimization and coalgebraic behavior\[[31](https://arxiv.org/html/2609.20207#bib.bib31),[32](https://arxiv.org/html/2609.20207#bib.bib32)\]\. We formulate that principle for typed, partial and measurable contextual transformations with heterogeneous readouts\.
Assume that𝔗0\\mathfrak\{T\}\_\{0\}contains the identity and is closed under right composition by every transformation whose dynamics are modelled\. Define
𝒵∗:=∏\(F,T\)∈𝒪0×𝔗0VF\\mathcal\{Z\}\_\{\*\}:=\\prod\_\{\(F,T\)\\in\\mathcal\{O\}\_\{0\}\\times\\mathfrak\{T\}\_\{0\}\}V\_\{F\}with its product sigma\-algebra\.
###### Definition 2\.13\(Canonical lexical signature\)\.
The canonical lexical signature is
𝖲∗:𝒞⟶𝒵∗,𝖲∗\(c\):=\(F\(Tc\)\)\(F,T\)∈𝒪0×𝔗0\.\\mathsf\{S\}\_\{\*\}:\\mathcal\{C\}\\longrightarrow\\mathcal\{Z\}\_\{\*\},\\qquad\\mathsf\{S\}\_\{\*\}\(c\):=\\bigl\(F\(Tc\)\\bigr\)\_\{\(F,T\)\\in\\mathcal\{O\}\_\{0\}\\times\\mathfrak\{T\}\_\{0\}\}\.
Its fibres are exactly the equivalence classes in Definition[2\.11](https://arxiv.org/html/2609.20207#S2.Thmdefinition11)\.
###### Theorem 2\.14\(Canonical minimal closed representation; extension of future\-behavior quotients\[[31](https://arxiv.org/html/2609.20207#bib.bib31),[32](https://arxiv.org/html/2609.20207#bib.bib32)\]\)\.
The signature𝖲∗\\mathsf\{S\}\_\{\*\}has the following properties\.
1. 1\.For each modelledUU, there is a unique induced transformationU¯∗\\overline\{U\}\_\{\*\}on𝖲∗\(𝒞\)\\mathsf\{S\}\_\{\*\}\(\\mathcal\{C\}\)satisfying 𝖲∗∘U=U¯∗∘𝖲∗\.\\mathsf\{S\}\_\{\*\}\\circ U=\\overline\{U\}\_\{\*\}\\circ\\mathsf\{S\}\_\{\*\}\.
2. 2\.EveryF∈𝒪0F\\in\\mathcal\{O\}\_\{0\}factors through𝖲∗\\mathsf\{S\}\_\{\*\}\.
3. 3\.LetR:𝒞→𝒵R:\\mathcal\{C\}\\to\\mathcal\{Z\}be any representation such that, for everyF∈𝒪0F\\in\\mathcal\{O\}\_\{0\}andT∈𝔗0T\\in\\mathfrak\{T\}\_\{0\}, there is a mapgF,Tg\_\{F,T\}with F∘T=gF,T∘R\.F\\circ T=g\_\{F,T\}\\circ R\.Then𝖲∗\\mathsf\{S\}\_\{\*\}factors uniquely throughRRonR\(𝒞\)R\(\\mathcal\{C\}\): 𝖲∗=h∘R\.\\mathsf\{S\}\_\{\*\}=h\\circ R\.Consequently,RRcannot identify two contexts distinguished by𝖲∗\\mathsf\{S\}\_\{\*\}\. Up to a bijection of its image,𝖲∗\\mathsf\{S\}\_\{\*\}is the coarsest representation sufficient for the declared observables and transformations\.
###### Proof\.
If𝖲∗\(c\)=𝖲∗\(c′\)\\mathsf\{S\}\_\{\*\}\(c\)=\\mathsf\{S\}\_\{\*\}\(c^\{\\prime\}\), thenF\(Tc\)=F\(Tc′\)F\(Tc\)=F\(Tc^\{\\prime\}\)for every\(F,T\)\(F,T\)\. Closure under right composition givesF\(TUc\)=F\(TUc′\)F\(TUc\)=F\(TUc^\{\\prime\}\), hence𝖲∗\(Uc\)=𝖲∗\(Uc′\)\\mathsf\{S\}\_\{\*\}\(Uc\)=\\mathsf\{S\}\_\{\*\}\(Uc^\{\\prime\}\)\. The descent theorem therefore defines a uniqueU¯∗\\overline\{U\}\_\{\*\}\. Since the identity lies in𝔗0\\mathfrak\{T\}\_\{0\},F\(c\)F\(c\)is a coordinate projection of𝖲∗\(c\)\\mathsf\{S\}\_\{\*\}\(c\), proving the second claim\. For the third, define
h\(R\(c\)\):=\(gF,T\(R\(c\)\)\)F,T\.h\(R\(c\)\):=\\bigl\(g\_\{F,T\}\(R\(c\)\)\\bigr\)\_\{F,T\}\.The assumed factorizations make this definition representative independent and equal to𝖲∗\(c\)\\mathsf\{S\}\_\{\*\}\(c\)\. Uniqueness follows onR\(𝒞\)R\(\\mathcal\{C\}\)\. ∎
### 2\.5Axioms of lexical calculus
The representation theory above determines when a closed state exists\. We next record the elementary laws obeyed by finite changes of an observable\. These laws act as a ledger: they track change under one edit, accumulation along a sequence, and order effects between noncommuting edits\. They do not postulate a smooth manifold of sentences or an infinitesimal derivative of a word\.
For example, ifTTappends an evidence item andSSchanges its numerical format, the total change fromcctoS\(Tc\)S\(Tc\)is the change caused byTTplus the change caused bySSat the already transformed contextTcTc\. This is an exact cocycle identity, not a Taylor approximation\.
We now state the laws required of the deterministic calculus\. They define the target mathematical structure\. An observed language service need not satisfy them exactly; its departures are reported as defects\.
Let𝖮𝖻𝗌\(𝒞,V\)\\mathsf\{Obs\}\(\\mathcal\{C\},V\)be a vector space of measurable observablesF:𝒞→VF:\\mathcal\{C\}\\to V, whereVVis a normed vector space\. Let𝔗\\mathfrak\{T\}be a small category whose objects are admissible context domains and whose arrows are typed lexical transformations\. Composition is writtenS∘TS\\circ T, meaning firstTT, thenSS\.
LC1: Typed closure\.The identity belongs to𝔗\\mathfrak\{T\}, and compatible transformations have an associative composition\. Transformations that violate the declared grammar, provenance, or domain constraints are not composable arrows\.
LC2: Linearity in observables\.For scalarsa,ba,b,
𝖣T\(aF\+bG\)=a𝖣TF\+b𝖣TG\.\\mathsf\{D\}\_\{T\}\(aF\+bG\)=a\\mathsf\{D\}\_\{T\}F\+b\\mathsf\{D\}\_\{T\}G\.
LC3: Null and identity laws\.𝖣IF=0,𝖣TF=0wheneverF∘T=F\.\\mathsf\{D\}\_\{\\mathrm\{I\}\}F=0,\\qquad\\mathsf\{D\}\_\{T\}F=0\\ \\text\{whenever \}F\\circ T=F\.
LC4: Cocycle law\.𝖣S∘TF\(c\)=𝖣TF\(c\)\+𝖣SF\(Tc\)\.\\mathsf\{D\}\_\{S\\circ T\}F\(c\)=\\mathsf\{D\}\_\{T\}F\(c\)\+\\mathsf\{D\}\_\{S\}F\(Tc\)\.
LC5: Product law\.For scalar observablesF,GF,G,
𝖣T\(FG\)\(c\)=F\(c\)𝖣TG\(c\)\+G\(c\)𝖣TF\(c\)\+𝖣TF\(c\)𝖣TG\(c\)\.\\mathsf\{D\}\_\{T\}\(FG\)\(c\)=F\(c\)\\mathsf\{D\}\_\{T\}G\(c\)\+G\(c\)\\mathsf\{D\}\_\{T\}F\(c\)\+\\mathsf\{D\}\_\{T\}F\(c\)\\mathsf\{D\}\_\{T\}G\(c\)\.Equivalently,𝖣T\(FG\)=F𝖣TG\+\(G∘T\)𝖣TF\\mathsf\{D\}\_\{T\}\(FG\)=F\\mathsf\{D\}\_\{T\}G\+\(G\\circ T\)\\mathsf\{D\}\_\{T\}F\.
LC6: Information congruence\.Ifc≃ℐc′c\\simeq\_\{\\mathcal\{I\}\}c^\{\\prime\}, then an information\-invariant observable satisfiesF\(c\)=F\(c′\)F\(c\)=F\(c^\{\\prime\}\)\. Moreover, every admissible transformation declared to act on information classes obeys
c≃ℐc′⟹Tc≃ℐTc′\.c\\simeq\_\{\\mathcal\{I\}\}c^\{\\prime\}\\Longrightarrow Tc\\simeq\_\{\\mathcal\{I\}\}Tc^\{\\prime\}\.Thus it descends to𝒞ℐ\\mathcal\{C\}\_\{\\mathcal\{I\}\}\.
LC7: Semantic naturality\.A dynamically adequate semantic map satisfies
𝐁∘T=T¯∘𝐁\\mathbf\{B\}\\circ T=\\overline\{T\}\\circ\\mathbf\{B\}for every transformation in the modelled family\. Approximate calculi replace equality by a uniform metric defect\.
LC8: Pushforward compatibility\.For a measurable semantic mapΠ\\Piand lexical kernel𝖪\\mathsf\{K\},
𝖣T\(Π∘𝖪\)\(c\)=Π\(𝖪Tc\)−Π\(𝖪c\)\.\\mathsf\{D\}\_\{T\}\(\\Pi\\circ\\mathsf\{K\}\)\(c\)=\\Pi\(\\mathsf\{K\}\_\{Tc\}\)\-\\Pi\(\\mathsf\{K\}\_\{c\}\)\.IfΠ\\Piis bounded linear on signed measures, then
𝖣T\(Π∘𝖪\)=Π\(𝖣T𝖪\)\.\\mathsf\{D\}\_\{T\}\(\\Pi\\circ\\mathsf\{K\}\)=\\Pi\(\\mathsf\{D\}\_\{T\}\\mathsf\{K\}\)\.
LC9: Fundamental path law\.Along every finite admissible path, total change equals the sum of transported lexical differentials\. Path independence holds exactly when circulation vanishes on all cycles\.
LC10: Continuity\.For specified metricsd𝒞d\_\{\\mathcal\{C\}\}anddVd\_\{V\}, admissible observables and induced state maps have declared moduli of continuity\. This converts approximate information equivalence and measurement error into explicit output\-error bounds\.
LC1–LC5 and LC9 are algebraic laws inherited from transformation and finite\-difference calculus\. LC6–LC8 are the language\-and\-semantics laws that distinguish the present construction\. LC10 is required for statistical falsifiability\.
#### 2\.5\.1Empirical defects
For each axiom, define a defect rather than treating approximate equality as success by inspection\. Important examples are
δinfo\(T\)\\displaystyle\\delta\_\{\\mathrm\{info\}\}\(T\):=supc∈𝒞0:c≃ℐTcd\(𝐁\(c\),𝐁\(Tc\)\),\\displaystyle:=\\sup\_\{c\\in\\mathcal\{C\}\_\{0\}:\\,c\\simeq\_\{\\mathcal\{I\}\}Tc\}d\\bigl\(\\mathbf\{B\}\(c\),\\mathbf\{B\}\(Tc\)\\bigr\),δnat\(T,T¯\)\\displaystyle\\delta\_\{\\mathrm\{nat\}\}\(T,\\overline\{T\}\):=supc∈𝒞0d\(𝐁\(Tc\),T¯\(𝐁\(c\)\)\),\\displaystyle:=\\sup\_\{c\\in\\mathcal\{C\}\_\{0\}\}d\\bigl\(\\mathbf\{B\}\(Tc\),\\overline\{T\}\(\\mathbf\{B\}\(c\)\)\\bigr\),δcomp\(S,T\)\\displaystyle\\delta\_\{\\mathrm\{comp\}\}\(S,T\):=sup𝒃d\(S∘T¯\(𝒃\),S¯\(T¯\(𝒃\)\)\)\.\\displaystyle:=\\sup\_\{\\bm\{b\}\}d\\bigl\(\\overline\{S\\circ T\}\(\\bm\{b\}\),\\overline\{S\}\(\\overline\{T\}\(\\bm\{b\}\)\)\\bigr\)\.The calculus is\(εinfo,εnat,εcomp\)\(\\varepsilon\_\{\\mathrm\{info\}\},\\varepsilon\_\{\\mathrm\{nat\}\},\\varepsilon\_\{\\mathrm\{comp\}\}\)\-valid on an operating domain when simultaneous confidence bounds for these defects lie below thresholds frozen before testing\.
## 3Minimal representations and semantic descent
### 3\.1The target commuting diagram
The canonical signature may be much larger than the state an application wants to retain\. A clinician may keep three triage probabilities; a control system may keep normal, warning and critical probabilities\. The practical question is therefore not whether some closed representation exists, but whether this particular coarsening supports the proposed transformation\.
The test is a commuting diagram\. Transforming the full context and then measuring its state must agree with first measuring the state and then applying an update defined solely on that state\. If two contexts have the same retained state but the same transformation sends them to different successor states, no state\-only update exists\. This is a sufficiency failure, not a finite\-sample inconvenience\.
ForT∈𝔗T\\in\\mathfrak\{T\}, we seek an induced state transformation
T¯:𝐁\(domT\)⟶ΔK−1\\overline\{T\}:\\mathbf\{B\}\(\\operatorname\{dom\}T\)\\longrightarrow\\Delta^\{K\-1\}such that
𝐁∘T=T¯∘𝐁\.\\mathbf\{B\}\\circ T=\\overline\{T\}\\circ\\mathbf\{B\}\.\(1\)
c∈𝒞c\\in\\mathcal\{C\}Tc∈𝒞Tc\\in\\mathcal\{C\}𝐁\(c\)∈ΔK−1\\mathbf\{B\}\(c\)\\in\\Delta^\{K\-1\}𝐁\(Tc\)∈ΔK−1\\mathbf\{B\}\(Tc\)\\in\\Delta^\{K\-1\}TTΠ∘𝖪\\Pi\\circ\\mathsf\{K\}Π∘𝖪\\Pi\\circ\\mathsf\{K\}T¯\\overline\{T\}
Equation \([1](https://arxiv.org/html/2609.20207#S3.E1)\) is semantic closure: the measured state retains exactly the information needed to determine the effect ofTT\.
#### 3\.1\.1Exact descent
Define the semantic equivalence relation
c∼𝐁c′⟺𝐁\(c\)=𝐁\(c′\)\.c\\sim\_\{\\mathbf\{B\}\}c^\{\\prime\}\\quad\\Longleftrightarrow\\quad\\mathbf\{B\}\(c\)=\\mathbf\{B\}\(c^\{\\prime\}\)\.
###### Theorem 3\.1\(Descent, existence, and uniqueness—quotient factorization\)\.
For a transformationTT, the following statements are equivalent\.
1. 1\.There existsT¯\\overline\{T\}satisfying \([1](https://arxiv.org/html/2609.20207#S3.E1)\)\.
2. 2\.TTpreserves semantic fibres: 𝐁\(c\)=𝐁\(c′\)⟹𝐁\(Tc\)=𝐁\(Tc′\)\\mathbf\{B\}\(c\)=\\mathbf\{B\}\(c^\{\\prime\}\)\\Longrightarrow\\mathbf\{B\}\(Tc\)=\\mathbf\{B\}\(Tc^\{\\prime\}\)for everyc,c′∈dom\(T\)c,c^\{\\prime\}\\in\\operatorname\{dom\}\(T\)\.
3. 3\.TTinduces a well\-defined map on the quotientdom\(T\)/∼𝐁\\operatorname\{dom\}\(T\)/\{\\sim\_\{\\mathbf\{B\}\}\}\.
When these conditions hold,T¯\\overline\{T\}is unique on𝐁\(domT\)\\mathbf\{B\}\(\\operatorname\{dom\}T\)\.
###### Proof\.
If \(1\) holds and𝐁\(c\)=𝐁\(c′\)\\mathbf\{B\}\(c\)=\\mathbf\{B\}\(c^\{\\prime\}\), then𝐁\(Tc\)=T¯\(𝐁\(c\)\)=T¯\(𝐁\(c′\)\)=𝐁\(Tc′\)\\mathbf\{B\}\(Tc\)=\\overline\{T\}\(\\mathbf\{B\}\(c\)\)=\\overline\{T\}\(\\mathbf\{B\}\(c^\{\\prime\}\)\)=\\mathbf\{B\}\(Tc^\{\\prime\}\), proving \(2\)\. Condition \(2\) makes the definitionT¯\(𝐁\(c\)\):=𝐁\(Tc\)\\overline\{T\}\(\\mathbf\{B\}\(c\)\):=\\mathbf\{B\}\(Tc\)independent of the representativecc, proving \(1\) and \(3\)\. Any map satisfying \([1](https://arxiv.org/html/2609.20207#S3.E1)\) must take𝐁\(c\)\\mathbf\{B\}\(c\)to𝐁\(Tc\)\\mathbf\{B\}\(Tc\), which proves uniqueness on the image\. ∎
#### 3\.1\.2Approximate descent
Let\(ΔK−1,d\)\(\\Delta^\{K\-1\},d\)be a metric space\. Define the diameter of the transformed fibre at𝒃\\bm\{b\}by
ωT\(𝒃\):=supc,c′:𝐁\(c\)=𝐁\(c′\)=𝒃d\(𝐁\(Tc\),𝐁\(Tc′\)\),\\omega\_\{T\}\(\\bm\{b\}\):=\\sup\_\{c,c^\{\\prime\}:\\,\\mathbf\{B\}\(c\)=\\mathbf\{B\}\(c^\{\\prime\}\)=\\bm\{b\}\}d\\bigl\(\\mathbf\{B\}\(Tc\),\\mathbf\{B\}\(Tc^\{\\prime\}\)\\bigr\),andωT∗:=sup𝒃ωT\(𝒃\)\\omega\_\{T\}^\{\*\}:=\\sup\_\{\\bm\{b\}\}\\omega\_\{T\}\(\\bm\{b\}\)\.
###### Theorem 3\.3\(Sharp representation\-relative approximate descent\)\.
Letddbe induced by a norm and assume that each transformed fibre has nonempty image\. For any state mapGG,
supcd\(𝐁\(Tc\),G\(𝐁\(c\)\)\)≥12ωT∗\.\\sup\_\{c\}d\\bigl\(\\mathbf\{B\}\(Tc\),G\(\\mathbf\{B\}\(c\)\)\\bigr\)\\geq\\frac\{1\}\{2\}\\omega\_\{T\}^\{\*\}\.If transformed fibre images admit Chebyshev centres of radius at mostrTr\_\{T\}, there existsT¯\\overline\{T\}satisfying
supcd\(𝐁\(Tc\),T¯\(𝐁\(c\)\)\)≤rT,\\sup\_\{c\}d\\bigl\(\\mathbf\{B\}\(Tc\),\\overline\{T\}\(\\mathbf\{B\}\(c\)\)\\bigr\)\\leq r\_\{T\},whereωT∗/2≤rT≤ωT∗\\omega\_\{T\}^\{\*\}/2\\leq r\_\{T\}\\leq\\omega\_\{T\}^\{\*\}\. In one\-dimensional state coordinates,rT=ωT∗/2r\_\{T\}=\\omega\_\{T\}^\{\*\}/2\.
###### Proof\.
For any two points in the same transformed fibre, the triangle inequality implies that at least one lies at distance at least half their separation fromG\(𝒃\)G\(\\bm\{b\}\)\. Taking suprema proves the lower bound\. Selecting a Chebyshev centre in each transformed fibre proves the upper bound\. On the real line the midpoint of the extrema has radius one\-half the diameter\. ∎
This theorem identifies the irreducible state\-compression error\. It cannot be repaired by collecting more observations while retaining the same semantic state\.
For example, suppose two records are both summarized as\(0\.6,0\.3,0\.1\)\(0\.6,0\.3,0\.1\), but the same new evidence sends one to\(0\.8,0\.15,0\.05\)\(0\.8,0\.15,0\.05\)and the other to\(0\.3,0\.4,0\.3\)\(0\.3,0\.4,0\.3\)\. The present three numbers have omitted information needed for the update\. Theorem[3\.3](https://arxiv.org/html/2609.20207#S3.Thmdefinition3)turns that objection into a lower bound: every state\-only update must err on at least one record\.
#### 3\.1\.3Composition of approximate descents
One\-step closure is useful only if its error can be controlled under composition\. LetTj:Cj−1⇢CjT\_\{j\}:C\_\{j\-1\}\\dashrightarrow C\_\{j\}be a compatible path, letBj:Cj→\(Mj,dj\)B\_\{j\}:C\_\{j\}\\to\(M\_\{j\},d\_\{j\}\)be semantic representations, and letGj:Mj−1→MjG\_\{j\}:M\_\{j\-1\}\\to M\_\{j\}be proposed induced updates\. Define the exact represented path and its closed approximation by
y0=B0\(c\),yj=Bj\(Tj⋯T1c\),y\_\{0\}=B\_\{0\}\(c\),\\qquad y\_\{j\}=B\_\{j\}\(T\_\{j\}\\cdots T\_\{1\}c\),and
y^0=B0\(c\),y^j=Gj\(y^j−1\)\.\\widehat\{y\}\_\{0\}=B\_\{0\}\(c\),\\qquad\\widehat\{y\}\_\{j\}=G\_\{j\}\(\\widehat\{y\}\_\{j\-1\}\)\.
###### Theorem 3\.4\(Multi\-step approximate descent; discrete Gronwall bound\)\.
SupposeGjG\_\{j\}isLjL\_\{j\}\-Lipschitz and, on every exact state reached by the path,
dj\(Bj\(Tjx\),Gj\(Bj−1\(x\)\)\)≤εj\.d\_\{j\}\\bigl\(B\_\{j\}\(T\_\{j\}x\),G\_\{j\}\(B\_\{j\-1\}\(x\)\)\\bigr\)\\leq\\varepsilon\_\{j\}\.Then for every admissible initial context,
dn\(yn,y^n\)≤∑j=1n\(∏k=j\+1nLk\)εj,d\_\{n\}\(y\_\{n\},\\widehat\{y\}\_\{n\}\)\\leq\\sum\_\{j=1\}^\{n\}\\left\(\\prod\_\{k=j\+1\}^\{n\}L\_\{k\}\\right\)\\varepsilon\_\{j\},where an empty product equals one\. In particular:
1. 1\.ifLj≤1L\_\{j\}\\leq 1andεj≤ε\\varepsilon\_\{j\}\\leq\\varepsilon, the error is at mostnεn\\varepsilon;
2. 2\.ifLj≤ρ<1L\_\{j\}\\leq\\rho<1andεj≤ε\\varepsilon\_\{j\}\\leq\\varepsilon, the error is at mostε\(1−ρn\)/\(1−ρ\)\\varepsilon\(1\-\\rho^\{n\}\)/\(1\-\\rho\);
3. 3\.without a bound on the products of Lipschitz constants, uniformly small one\-step defects need not yield a stable pathwise representation\.
###### Proof\.
Letej=dj\(yj,y^j\)e\_\{j\}=d\_\{j\}\(y\_\{j\},\\widehat\{y\}\_\{j\}\)\. InsertingGj\(yj−1\)G\_\{j\}\(y\_\{j\-1\}\)and applying the triangle inequality and Lipschitz property gives
ej≤dj\(yj,Gj\(yj−1\)\)\+dj\(Gj\(yj−1\),Gj\(y^j−1\)\)≤εj\+Ljej−1\.e\_\{j\}\\leq d\_\{j\}\(y\_\{j\},G\_\{j\}\(y\_\{j\-1\}\)\)\+d\_\{j\}\(G\_\{j\}\(y\_\{j\-1\}\),G\_\{j\}\(\\widehat\{y\}\_\{j\-1\}\)\)\\leq\\varepsilon\_\{j\}\+L\_\{j\}e\_\{j\-1\}\.Sincee0=0e\_\{0\}=0, iteration proves the displayed sum\. The first two conclusions follow by bounding the products; the third follows because those products are the amplification factors in the exact bound\. ∎
### 3\.2Ontology adequacy
The preceding results make ontology choice part of the mathematics rather than a preliminary naming exercise\. A semantic representationB:𝒞→SB:\\mathcal\{C\}\\to Sis*transformation sufficient*for𝔗0⊆𝔗\\mathfrak\{T\}\_\{0\}\\subseteq\\mathfrak\{T\}when everyT∈𝔗0T\\in\\mathfrak\{T\}\_\{0\}descends throughBB\. Equivalently, by Theorem[3\.1](https://arxiv.org/html/2609.20207#S3.Thmdefinition1),
B\(c\)=B\(c′\)⟹B\(Tc\)=B\(Tc′\)\(T∈𝔗0\)\.B\(c\)=B\(c^\{\\prime\}\)\\quad\\Longrightarrow\\quad B\(Tc\)=B\(Tc^\{\\prime\}\)\\qquad\(T\\in\\mathfrak\{T\}\_\{0\}\)\.This is the precise adequacy condition required by stochastic lexical calculus: a retained state must preserve every distinction needed to determine its admissible future updates\. Approximate adequacy is measured by the transformed\-fibre diameter in Theorem[3\.3](https://arxiv.org/html/2609.20207#S3.Thmdefinition3), and its sequential cost is bounded by Theorem[3\.4](https://arxiv.org/html/2609.20207#S3.Thmdefinition4)\.
Failure has a clear interpretation\. The proposed ontology has merged contexts that respond differently to a relevant transformation, so no closed state recursion exists at that resolution\. Refining the state may restore closure; adding a more elaborate dynamic model to the same insufficient state cannot remove the representation\-relative lower bound\. The appendix records brief additional remarks on balancing closure against statistical complexity\.
### 3\.3Exact and approximate finite examples
#### 3\.3\.1An exact closed quotient
LetC=\{00,01,10,11\}C=\\\{00,01,10,11\\\}with the discrete sigma\-algebra\. LetT\(x1x2\)=\(1−x1\)x2T\(x\_\{1\}x\_\{2\}\)=\(1\-x\_\{1\}\)x\_\{2\}flip the first bit, and let the observable beF\(x1x2\)=x1F\(x\_\{1\}x\_\{2\}\)=x\_\{1\}\. Take the transformation category generated byTT, soT2=1CT^\{2\}=1\_\{C\}\. The behavioral signature is
S\(x1x2\)=\(F\(x1x2\),F\(T\(x1x2\)\)\)=\(x1,1−x1\)\.S\(x\_\{1\}x\_\{2\}\)=\\bigl\(F\(x\_\{1\}x\_\{2\}\),F\(T\(x\_\{1\}x\_\{2\}\)\)\\bigr\)=\(x\_\{1\},1\-x\_\{1\}\)\.It has exactly two fibres, and the quotientB\(x1x2\)=x1B\(x\_\{1\}x\_\{2\}\)=x\_\{1\}is measurably isomorphic toS\(C\)S\(C\)\. The induced update is
T¯B\(b\)=1−b\.\\overline\{T\}\_\{B\}\(b\)=1\-b\.The second bit is discarded because neither the observable nor any admitted future transformation makes it visible\. Thus the calculus removes irrelevant detail while retaining an exact closed recursion\.
A linguistic reading is immediate but not required by the mathematics\. The two bits could encode a semantic orientation and a stylistic feature; if the declared transformations alter only orientation and the readout observes only orientation, style is correctly omitted from the minimal state\.
#### 3\.3\.2A sharp approximate obstruction
LetC=\{a,b,x0,x1\}C=\\\{a,b,x\_\{0\},x\_\{1\}\\\}with the discrete sigma\-algebra and define
B\(a\)=B\(b\)=12,B\(x0\)=0,B\(x1\)=1\.B\(a\)=B\(b\)=\\tfrac\{1\}\{2\},\\qquad B\(x\_\{0\}\)=0,\\qquad B\(x\_\{1\}\)=1\.Let the partial transformationTThave domain\{a,b\}\\\{a,b\\\}and satisfyT\(a\)=x0T\(a\)=x\_\{0\},T\(b\)=x1T\(b\)=x\_\{1\}\. The two source contexts are indistinguishable underBB, but their transformed states are maximally separated:
ωB,T=\|B\(Ta\)−B\(Tb\)\|=1\.\\omega\_\{B,T\}=\|B\(Ta\)\-B\(Tb\)\|=1\.No function of the retained state alone can reproduce both successors\. Indeed, for everyG:\[0,1\]→\[0,1\]G:\[0,1\]\\to\[0,1\],
max\{\|G\(1/2\)−0\|,\|G\(1/2\)−1\|\}≥12,\\max\\\{\|G\(1/2\)\-0\|,\|G\(1/2\)\-1\|\\\}\\geq\\tfrac\{1\}\{2\},and equality is achieved byG\(1/2\)=1/2G\(1/2\)=1/2\. The lower bound in Theorem[3\.3](https://arxiv.org/html/2609.20207#S3.Thmdefinition3)is therefore sharp\. In a language application,aaandbbmay receive the same present semantic score while responding differently to a subsequent qualification; the example shows why present calibration alone cannot justify a state recursion\.
## 4Stochastic lexical calculus on probability simplices
### 4\.1Differentials and transformation algebra
Once closure is settled, it becomes meaningful to discuss change\. The appropriate primitive is a finite difference along a typed transformation, because language edits are discrete and often noncommutative\. The term “differential” is used in this finite sense\. It supplies composition and path identities for sequential analysis without importing unsupported smoothness\.
LetVVbe a normed vector space andF:𝒞→VF:\\mathcal\{C\}\\to V\.
###### Definition 4\.1\(Lexical differential\)\.
The finite lexical differential ofFFin directionTTis
𝖣TF\(c\):=F\(Tc\)−F\(c\)\.\\mathsf\{D\}\_\{T\}F\(c\):=F\(Tc\)\-F\(c\)\.For probability measures, subtraction is interpreted as a finite signed measure; for semantic compositions it is an element of the simplex tangent space\{𝐯:𝟏⊤𝐯=0\}\\\{\\bm\{v\}:\\bm\{1\}^\{\\top\}\\bm\{v\}=0\\\}\.
This is a difference operator\. Its specifically lexical content comes from the typed transformation action, equivalence structure, and observable kernel—not from claiming an infinitesimal geometry on isolated words\.
For compatibleS,T∈𝔗S,T\\in\\mathfrak\{T\}, adding and subtractingF\(Tc\)F\(Tc\)gives the elementary cocycle identity
𝖣S∘TF\(c\)=𝖣TF\(c\)\+𝖣SF\(Tc\)\.\\mathsf\{D\}\_\{S\\circ T\}F\(c\)=\\mathsf\{D\}\_\{T\}F\(c\)\+\\mathsf\{D\}\_\{S\}F\(Tc\)\.
###### Definition 4\.2\(Lexical commutator\)\.
The order effect ofSSandTTonFFis
\[S,T\]F\(c\):=F\(S∘T\(c\)\)−F\(T∘S\(c\)\)\.\[S,T\]\_\{F\}\(c\):=F\(S\\circ T\(c\)\)\-F\(T\\circ S\(c\)\)\.
This definition remains meaningful when the transformation semigroup is noncommutative\. It directly measures whether two evidence operations have order\-dependent semantic effects\.
Likewise, for a pathγ=\(T1,…,Tm\)\\gamma=\(T\_\{1\},\\ldots,T\_\{m\}\)withcj=Tjcj−1c\_\{j\}=T\_\{j\}c\_\{j\-1\}, telescoping gives
F\(cm\)−F\(c0\)=∑j=1m𝖣TjF\(cj−1\)\.F\(c\_\{m\}\)\-F\(c\_\{0\}\)=\\sum\_\{j=1\}^\{m\}\\mathsf\{D\}\_\{T\_\{j\}\}F\(c\_\{j\-1\}\)\.Thus accumulated lexical differentials are path independent on a connected transformation graph exactly when their circulation vanishes around every directed cycle\. These identities require no empirical claim: they follow from the definition of a finite difference\. The substantive statistical question begins with whether the semantic state through which they are computed is closed and stable\.
### 4\.2Semantic pushforwards
An observable language law may live on a vast continuation space, while the retained state lies in a finite simplex\. A semantic map pushes the former law to the latter\. Its Lipschitz modulus measures how much an upstream perturbation can be amplified in semantic probability space\. The next bound therefore connects an observable language discrepancy to a downstream state\-error budget\.
Let∥⋅∥TV\\\|\\cdot\\\|\_\{\\mathrm\{TV\}\}be total variation distance on𝒫\(𝒲\)\\mathcal\{P\}\(\\mathcal\{W\}\)and supposeΠ\\PiisLΠL\_\{\\Pi\}\-Lipschitz from total variation to\(ΔK−1,d\)\(\\Delta^\{K\-1\},d\)\.
###### Theorem 4\.3\(Pushforward stability\)\.
For everyTTandcc,
d\(𝐁\(Tc\),𝐁\(c\)\)≤LΠ‖𝖪Tc−𝖪c‖TV\.d\\bigl\(\\mathbf\{B\}\(Tc\),\\mathbf\{B\}\(c\)\\bigr\)\\leq L\_\{\\Pi\}\\,\\\|\\mathsf\{K\}\_\{Tc\}\-\\mathsf\{K\}\_\{c\}\\\|\_\{\\mathrm\{TV\}\}\.If𝖪^\\widehat\{\\mathsf\{K\}\}andΠ^\\widehat\{\\Pi\}satisfy uniform errorsεK\\varepsilon\_\{K\}andεΠ\\varepsilon\_\{\\Pi\}, then the estimated semantic differential obeys
d\(𝖣T𝐁^\(c\),𝖣T𝐁\(c\)\)≤2εΠ\+2LΠεK,d\\bigl\(\\widehat\{\\mathsf\{D\}\_\{T\}\\mathbf\{B\}\}\(c\),\\mathsf\{D\}\_\{T\}\\mathbf\{B\}\(c\)\\bigr\)\\leq 2\\varepsilon\_\{\\Pi\}\+2L\_\{\\Pi\}\\varepsilon\_\{K\},under the product metric induced bydd\.
###### Proof\.
The first claim is the Lipschitz property\. For the second, insert the true kernel and map at bothccandTcTc, then apply the triangle inequality twice\. ∎
For a linear partition pushforward,LΠ≤1L\_\{\\Pi\}\\leq 1in total variation\. A calibrated nonlinear map requires an estimated or theoretically bounded modulus of continuity\.
### 4\.3Semantic probability maps as a special representation
The abstract theory becomes measurable when the observable is a conditional law over possible verbal continuations\. If phrases such as “urgent review” and “immediate escalation” belong to the same declared state, their probabilities may be grouped and calibrated into a state probability\. That operation is useful but not automatic: phrase probabilities remain prompt dependent, and different groupings can discard different future\-relevant distinctions\.
The general calculus becomes a statistical measurement theory when one selected observable is a continuation kernel𝖪c\\mathsf\{K\}\_\{c\}and a semantic mapΠ\\Piaggregates or calibrates lexical outcomes:
𝒞→𝖪𝒫\(𝒲\)→ΠΔK−1\.\\mathcal\{C\}\\xrightarrow\{\\mathsf\{K\}\}\\mathcal\{P\}\(\\mathcal\{W\}\)\\xrightarrow\{\\Pi\}\\Delta^\{K\-1\}\.This construction includes recent semantic\-uncertainty and confidence\-measurement schemes as special cases: a continuation law is grouped or calibrated into a finite collection of meanings\[[20](https://arxiv.org/html/2609.20207#bib.bib20),[25](https://arxiv.org/html/2609.20207#bib.bib25),[26](https://arxiv.org/html/2609.20207#bib.bib26),[35](https://arxiv.org/html/2609.20207#bib.bib35)\]\. A companion measurement paper develops and empirically certifies this semantic map\[[17](https://arxiv.org/html/2609.20207#bib.bib17)\], while a companion statistical paper establishes identification, recovery rates, and weak\-identification limits for the resulting observation kernel\[[18](https://arxiv.org/html/2609.20207#bib.bib18)\]\. The present construction remains distinct because neither𝖪\\mathsf\{K\}norΠ\\Pialone defines the lexical transformation category, observational equivalence, or canonical signature\.
The special case is nevertheless important\. It makes lexical differentials observable through externally supplied probabilities and turns semantic closure into a testable condition\. If
Π\(𝖪Tc\)=T¯\(Π\(𝖪c\)\),\\Pi\(\\mathsf\{K\}\_\{Tc\}\)=\\overline\{T\}\\bigl\(\\Pi\(\\mathsf\{K\}\_\{c\}\)\\bigr\),then the semantic probability composition is sufficient for the transformationTT\. If this equality fails on a fibre, the semantic map has coarsened distinctions needed by the subsequent dynamics\.
For an interior composition𝒃∈intΔK−1\\bm\{b\}\\in\\operatorname\{int\}\\Delta^\{K\-1\}, let
ilr:intΔK−1⟶ℝK−1\\operatorname\{ilr\}:\\operatorname\{int\}\\Delta^\{K\-1\}\\longrightarrow\\mathbb\{R\}^\{K\-1\}be any fixed isometric log\-ratio chart and define𝒛\(c\)=ilr\(𝐁\(c\)\)\\bm\{z\}\(c\)=\\operatorname\{ilr\}\(\\mathbf\{B\}\(c\)\)\. The coordinate lexical differential is
𝖣T𝒛\(c\)=ilr\(𝐁\(Tc\)\)−ilr\(𝐁\(c\)\)\.\\mathsf\{D\}\_\{T\}\\bm\{z\}\(c\)=\\operatorname\{ilr\}\(\\mathbf\{B\}\(Tc\)\)\-\\operatorname\{ilr\}\(\\mathbf\{B\}\(c\)\)\.Changing the orthonormal ILR basis rotates coordinates but does not change Aitchison distances or intrinsic closure\. ILR therefore provides Euclidean coordinates for estimation and dynamics without defining the lexical state itself\.
The relation between the theories is consequently:
lexical transformations⟶lexical observables⟶semantic map⟶log\-ratio representation\.\\text\{lexical transformations\}\\longrightarrow\\text\{lexical observables\}\\longrightarrow\\text\{semantic map\}\\longrightarrow\\text\{log\-ratio representation\}\.Lexical calculus supplies the first two arrows and the conditions under which the third supports a closed state\. An applied semantic\-measurement system estimates and tests a particular third arrow\. A stochastic lexical calculus would model random compositions only after these closure conditions have been checked\.
#### 4\.3\.1Prompt\-calibrated semantic descent
Letuuindex a prompt construction and letcu\(e\)c\_\{u\}\(e\)be the resulting context for evidenceee\. Two contexts may be information equivalent even when their continuation kernels differ\. The raw assignmentcu\(e\)↦𝖪cu\(e\)c\_\{u\}\(e\)\\mapsto\\mathsf\{K\}\_\{c\_\{u\}\(e\)\}therefore need not descend to the evidence quotient\. Introduce instead a prompt\-indexed semantic pushforwardΠu\\Pi\_\{u\}and calibration mapψu\\psi\_\{u\}, and define
Bu\(cu\(e\)\)=ψu\{Πu\(𝖪cu\(e\)\)\}\.B\_\{u\}\(c\_\{u\}\(e\)\)=\\psi\_\{u\}\\\!\\left\\\{\\Pi\_\{u\}\(\\mathsf\{K\}\_\{c\_\{u\}\(e\)\}\)\\right\\\}\.This family descends to evidence space whenBu\(cu\(e\)\)=Bu′\(cu′\(e\)\)B\_\{u\}\(c\_\{u\}\(e\)\)=B\_\{u^\{\\prime\}\}\(c\_\{u^\{\\prime\}\}\(e\)\)for every admissible pairu,u′u,u^\{\\prime\}\. Thus calibration is not merely a numerical correction; it can be the morphism that removes prompt\-specific coordinates before semantic state is formed\.
###### Proposition 4\.4\(Bayesian update descends through likelihood rays\)\.
Let𝛑∈intΔK−1\\bm\{\\pi\}\\in\\operatorname\{int\}\\Delta^\{K\-1\}, letPPbe a positive transition matrix, and let a calibrated semantic map return a positive likelihood vector𝐠u\(e\)\\bm\{g\}\_\{u\}\(e\)\. Define
ℬ𝒈\(𝝅\)=𝒈⊙P⊤𝝅𝟏⊤\(𝒈⊙P⊤𝝅\)\.\\mathcal\{B\}\_\{\\bm\{g\}\}\(\\bm\{\\pi\}\)=\\frac\{\\bm\{g\}\\odot P^\{\\top\}\\bm\{\\pi\}\}\{\\mathbf\{1\}^\{\\top\}\(\\bm\{g\}\\odot P^\{\\top\}\\bm\{\\pi\}\)\}\.The updateℬ𝐠u\(e\)\\mathcal\{B\}\_\{\\bm\{g\}\_\{u\}\(e\)\}descends through information\-equivalent prompts for every positive prior if and only if𝐠u\(e\)=a𝐠u′\(e\)\\bm\{g\}\_\{u\}\(e\)=a\\,\\bm\{g\}\_\{u^\{\\prime\}\}\(e\)for somea\>0a\>0\. Exact equality of the raw continuation kernels is unnecessary\.
###### Proof\.
Proportional likelihoods give the same normalized update because the common scalar cancels\. Conversely, equality of the updates for every positive prior implies equality of every pairwise posterior odds ratio\. The common prior odds cancel, so all pairwise likelihood ratios agree\. The two positive vectors are therefore proportional\. ∎
The proposition expresses the correct quotient\. Language probabilities live before semantic descent and may retain prompt\-specific distinctions\. Bayesian updating lives on projective likelihood space, where common scale is irrelevant\. Approximate descent is assessed by the distance between calibrated likelihood rays or resulting posteriors, not by equality of word distributions\.
### 4\.4Random dynamics after semantic descent
Static calibration is not enough for sequential use\. When evidence arrives repeatedly, yesterday’s retained state is fed into today’s update\. A small one\-step closure error can then be damped, accumulated or amplified depending on the update dynamics\. This section asks the limited question justified by the preceding theory: once semantic descent has been verified, under what conditions does the resulting external random recursion exist uniquely and remain stable?
The word “external” is essential\. The recursion is constructed from observable language\-derived measurements and declared updates\. Nothing here requires, or purports to reveal, hidden activations, model weights or an internal Bayesian computation\.
Let\(Ξn\)n≥1\(\\Xi\_\{n\}\)\_\{n\\geq 1\}be random marks selecting transformations and let
Cn\+1=TΞn\+1Cn\.C\_\{n\+1\}=T\_\{\\Xi\_\{n\+1\}\}C\_\{n\}\.If every relevant transformation descends, then
𝐁n\+1=T¯Ξn\+1\(𝐁n\)\.\\mathbf\{B\}\_\{n\+1\}=\\overline\{T\}\_\{\\Xi\_\{n\+1\}\}\(\\mathbf\{B\}\_\{n\}\)\.
###### Corollary 4\.5\(Closed lexical\-state recursion\)\.
Under exact descent, the semantic state process is Markov whenever the transformation marks are conditionally Markov given the current semantic state\. Underε\\varepsilon\-descent, it admits a closed recursion with one\-step deterministic misspecification at mostε\\varepsilon\.
This is the required foundation for stochastic lexical dynamics\. The next subsection develops random compositions, existence, uniqueness, contraction, and perturbation only after the closure defect has been defined\. Diffusion and jump limits, statistical estimation, and stopping rules remain separate extensions rather than prerequisites for the present calculus\.
#### 4\.4\.1Random semantic descent and average contraction
We now develop the part of that sequel needed to turn closed semantic descent into a mathematically well\-defined stochastic recursion\. Let\(Ω,ℱ,ℙ,θ\)\(\\Omega,\\mathcal\{F\},\\mathbb\{P\},\\theta\)be an invertible measure\-preserving dynamical system and letΞn=Ξ0∘θn\\Xi\_\{n\}=\\Xi\_\{0\}\\circ\\theta^\{n\}be a stationary sequence of transformation marks\. On a complete separable metric semantic space\(S,d\)\(S,d\), write
Xn\+1=FΞn\+1\(Xn\),Fξ:S→S,X\_\{n\+1\}=F\_\{\\Xi\_\{n\+1\}\}\(X\_\{n\}\),\\qquad F\_\{\\xi\}:S\\to S,whereFξF\_\{\\xi\}is the descended action of the corresponding lexical transformation\. Define its random Lipschitz coefficient by
L\(ξ\)=supx≠yd\{Fξ\(x\),Fξ\(y\)\}d\(x,y\)\.L\(\\xi\)=\\sup\_\{x\\neq y\}\\frac\{d\\\{F\_\{\\xi\}\(x\),F\_\{\\xi\}\(y\)\\\}\}\{d\(x,y\)\}\.The exact\-descent theorem above establishes that this recursion is a well\-defined representation of the contextual process\. The following result states when it has a unique causal state rather than merely a finite sequence of updates\.
###### Assumption 4\.6\(Average contraction\)\.
The driving system is ergodic,𝔼log\+L\(Ξ0\)<∞\\mathbb\{E\}\\log^\{\+\}L\(\\Xi\_\{0\}\)<\\infty, and, for somex0∈Sx\_\{0\}\\in S,
𝔼log\+d\{FΞ0\(x0\),x0\}<∞,χ:=𝔼logL\(Ξ0\)<0\.\\mathbb\{E\}\\log^\{\+\}d\\\{F\_\{\\Xi\_\{0\}\}\(x\_\{0\}\),x\_\{0\}\\\}<\\infty,\\qquad\\chi:=\\mathbb\{E\}\\log L\(\\Xi\_\{0\}\)<0\.We setlog0=−∞\\log 0=\-\\infty; a zero Lipschitz coefficient gives immediate coalescence and can be handled separately\.
###### Theorem 4\.7\(Causal existence, uniqueness, and synchronization\)\.
Under Assumption[4\.6](https://arxiv.org/html/2609.20207#S4.Thmdefinition6), the backward iterates
X0\(m\)=FΞ0∘FΞ−1∘⋯∘FΞ−m\+1\(x0\)X\_\{0\}^\{\(m\)\}=F\_\{\\Xi\_\{0\}\}\\circ F\_\{\\Xi\_\{\-1\}\}\\circ\\cdots\\circ F\_\{\\Xi\_\{\-m\+1\}\}\(x\_\{0\}\)converge almost surely to a random variableX0⋆X\_\{0\}^\{\\star\}independent of the choice ofx0x\_\{0\}\. The processXn⋆=X0⋆∘θnX\_\{n\}^\{\\star\}=X\_\{0\}^\{\\star\}\\circ\\theta^\{n\}is stationary, causal, and solves the recursion\. Any other causal stationary solution is equal to it almost surely\. Moreover, two forward trajectories driven by the same marks synchronize at exponential rate:
lim supn→∞1nlogd\(Xn,X~n\)≤χ<0almost surely\.\\limsup\_\{n\\to\\infty\}\\frac\{1\}\{n\}\\log d\(X\_\{n\},\\widetilde\{X\}\_\{n\}\)\\leq\\chi<0\\quad\\text\{almost surely\}\.
###### Proof\.
PutAj=L\(Ξ−j\)A\_\{j\}=L\(\\Xi\_\{\-j\}\)andDj=d\{FΞ−j\(x0\),x0\}D\_\{j\}=d\\\{F\_\{\\Xi\_\{\-j\}\}\(x\_\{0\}\),x\_\{0\}\\\}\. By the ergodic theorem,
1n∑j=0n−1logAj⟶χ<0almost surely\.\\frac\{1\}\{n\}\\sum\_\{j=0\}^\{n\-1\}\\log A\_\{j\}\\longrightarrow\\chi<0\\quad\\text\{almost surely\}\.Thus, for eachδ∈\(0,−χ\)\\delta\\in\(0,\-\\chi\), there is an almost surely finiteNδN\_\{\\delta\}such that∏j=0n−1Aj≤exp\{\(χ\+δ\)n\}\\prod\_\{j=0\}^\{n\-1\}A\_\{j\}\\leq\\exp\\\{\(\\chi\+\\delta\)n\\\}whenevern≥Nδn\\geq N\_\{\\delta\}\. The logarithmic moment condition onD0D\_\{0\}, stationarity, the tail\-sum formula and Borel–Cantelli implyDn≤exp\(δn\)D\_\{n\}\\leq\\exp\(\\delta n\)eventually\. Choosingδ<−χ/2\\delta<\-\\chi/2therefore gives
∑n≥0Dn∏j=0n−1Aj<∞almost surely\.\\sum\_\{n\\geq 0\}D\_\{n\}\\prod\_\{j=0\}^\{n\-1\}A\_\{j\}<\\infty\\quad\\text\{almost surely\}\.
Form\>ℓm\>\\ell, insert one additional map at a time in the backward compositions\. The triangle inequality and the Lipschitz bounds give
d\(X0\(m\),X0\(ℓ\)\)≤∑j=ℓm−1Dj∏h=0j−1Ah\.d\(X\_\{0\}^\{\(m\)\},X\_\{0\}^\{\(\\ell\)\}\)\\leq\\sum\_\{j=\\ell\}^\{m\-1\}D\_\{j\}\\prod\_\{h=0\}^\{j\-1\}A\_\{h\}\.The summability above makes the backward sequence Cauchy\. Completeness givesX0⋆X\_\{0\}^\{\\star\}, and the same product estimate applied to two starting points removes dependence onx0x\_\{0\}\. The limit is measurable with respect to the past marks, so it is causal\. Shifting the construction proves stationarity and the recursion\.
For two causal stationary solutions, pull both solutions backmmsteps and apply the same product estimate\. The Lipschitz product tends to zero almost surely, whereas stationarity makes the family of pulled\-back distances tight\. Their product therefore converges to zero in probability\. Since the time\-zero distance is bounded by that product for everymm, it is zero almost surely, proving uniqueness\. Finally,
d\(Xn,X~n\)≤d\(X0,X~0\)∏j=1nL\(Ξj\)d\(X\_\{n\},\\widetilde\{X\}\_\{n\}\)\\leq d\(X\_\{0\},\\widetilde\{X\}\_\{0\}\)\\prod\_\{j=1\}^\{n\}L\(\\Xi\_\{j\}\)and the ergodic limit of the logarithmic product proves synchronization\. ∎
Uniform contraction is therefore sufficient but unnecessary\. Individual lexical updates may expand semantic distance; the closed stochastic state still exists when contraction dominates on the logarithmic average\. This distinction is important for language, where a contradiction or negation may create a large local movement without making the entire recursion unstable\.
###### Theorem 4\.8\(Perturbation of an approximately descending recursion\)\.
Let
Xn\+1=FΞn\+1\(Xn\),X^n\+1=F^Ξn\+1\(X^n\),X\_\{n\+1\}=F\_\{\\Xi\_\{n\+1\}\}\(X\_\{n\}\),\\qquad\\widehat\{X\}\_\{n\+1\}=\\widehat\{F\}\_\{\\Xi\_\{n\+1\}\}\(\\widehat\{X\}\_\{n\}\),and supposed\{F^Ξn\(x\),FΞn\(x\)\}≤εnd\\\{\\widehat\{F\}\_\{\\Xi\_\{n\}\}\(x\),F\_\{\\Xi\_\{n\}\}\(x\)\\\}\\leq\\varepsilon\_\{n\}on the supported domain\. Then, pathwise,
d\(X^n,Xn\)≤\(∏j=1nL\(Ξj\)\)d\(X^0,X0\)\+∑s=1n\(∏j=s\+1nL\(Ξj\)\)εs\.d\(\\widehat\{X\}\_\{n\},X\_\{n\}\)\\leq\\left\(\\prod\_\{j=1\}^\{n\}L\(\\Xi\_\{j\}\)\\right\)d\(\\widehat\{X\}\_\{0\},X\_\{0\}\)\+\\sum\_\{s=1\}^\{n\}\\left\(\\prod\_\{j=s\+1\}^\{n\}L\(\\Xi\_\{j\}\)\\right\)\\varepsilon\_\{s\}\.Under Assumption[4\.6](https://arxiv.org/html/2609.20207#S4.Thmdefinition6), a uniformly bounded defectεs≤ε¯\\varepsilon\_\{s\}\\leq\\bar\{\\varepsilon\}produces an almost surely finite geometric resolvent\. If insteadL\(Ξs\)≤ρ<1L\(\\Xi\_\{s\}\)\\leq\\rho<1, the explicit bound is
d\(X^n,Xn\)≤ρnd\(X^0,X0\)\+ε¯1−ρn1−ρ\.d\(\\widehat\{X\}\_\{n\},X\_\{n\}\)\\leq\\rho^\{n\}d\(\\widehat\{X\}\_\{0\},X\_\{0\}\)\+\\bar\{\\varepsilon\}\\frac\{1\-\\rho^\{n\}\}\{1\-\\rho\}\.
###### Proof\.
Insert and subtractFΞn\+1\(X^n\)F\_\{\\Xi\_\{n\+1\}\}\(\\widehat\{X\}\_\{n\}\)and apply the triangle inequality:
d\(X^n\+1,Xn\+1\)≤εn\+1\+L\(Ξn\+1\)d\(X^n,Xn\)\.d\(\\widehat\{X\}\_\{n\+1\},X\_\{n\+1\}\)\\leq\\varepsilon\_\{n\+1\}\+L\(\\Xi\_\{n\+1\}\)d\(\\widehat\{X\}\_\{n\},X\_\{n\}\)\.Iteration proves the convolution formula\. Average contraction makes the backward products exponentially summable almost surely\. The uniform case is the finite geometric series\. ∎
#### 4\.4\.2Simplex\-valued lexical states
For the semantic probability representationS=intΔK−1S=\\operatorname\{int\}\\Delta^\{K\-1\}, choose any orthonormal contrast matrixVVand use isometric log\-ratio coordinates
ilrV\(𝒑\)=V⊤log𝒑\.\\operatorname\{ilr\}\_\{V\}\(\\bm\{p\}\)=V^\{\\top\}\\log\\bm\{p\}\.Two choices ofVVdiffer by an orthogonal matrix\. Consequently distances, singular values, contraction exponents, and the perturbation bounds above do not depend on the arbitrary coordinate basis\. The stochastic calculus is therefore simplex\-valued in its interpretation and Euclidean only in its local representation\.
###### Corollary 4\.9\(Basis\-invariant stochastic semantic filter\)\.
Suppose a prompt\-calibrated semantic map descends to positive likelihood rays as in Proposition[4\.4](https://arxiv.org/html/2609.20207#S4.Thmdefinition4), and suppose the resulting Bayesian update maps satisfy Assumption[4\.6](https://arxiv.org/html/2609.20207#S4.Thmdefinition6)in one ILR basis\. Then the causal filter exists, is unique, and has the same synchronization exponent in every ILR basis\. Approximate prompt descent enters only through the perturbation defectsεn\\varepsilon\_\{n\}in Theorem[4\.8](https://arxiv.org/html/2609.20207#S4.Thmdefinition8)\.
###### Proof\.
An ILR basis change is an orthogonal conjugacy\. Orthogonal maps preserve the metric and operator norms, hence preserve each Lipschitz coefficient and the Lyapunov exponentχ\\chi\. Apply Theorems[4\.7](https://arxiv.org/html/2609.20207#S4.Thmdefinition7)and[4\.8](https://arxiv.org/html/2609.20207#S4.Thmdefinition8)\. ∎
These results define the proper mathematical boundary of stochastic lexical calculus\. The language model is not asserted to perform Bayesian filtering internally\. Rather, a validated semantic representation supports an external random dynamical system whose existence, uniqueness, coordinate invariance, and approximation error can be proved\.
#### 4\.4\.3Controlled illustration of semantic descent
A frozen three\-state hidden Markov experiment tested Proposition[4\.4](https://arxiv.org/html/2609.20207#S4.Thmdefinition4)with two information\-equivalent prompt constructions, six lexical evidence symbols and eight updates per path\. The raw prompt kernels failed the maximum Jensen–Shannon invariance gate: their mean discrepancy was0\.008910\.00891, but their maximum was0\.133970\.13397against a frozen limit of0\.100\.10\. The failure is the expected obstruction at the wrong representation level\.
Prompt\-specific affine calibration maps were then fitted on a disjoint partition and frozen\. In calibrated posterior space, mean and maximum cross\-prompt Jensen–Shannon discrepancies fell to0\.0011220\.001122and0\.0107200\.010720\. Both passed their frozen limits of0\.020\.02and0\.100\.10\. A pathwise conformal radius of0\.1906690\.190669in total variation was calibrated on separate paths; 28 of 30 untouched eight\-update paths were covered, giving empirical coverage0\.9330\.933at nominal level0\.900\.90\. The minimum retained candidate mass was0\.9999998380\.999999838\.
This experiment does not prove a universal lexical calculus or an internal Bayesian mechanism\. It supplies a finite example in which a raw probability\-valued observable does not descend through prompt equivalence, whereas a prompt\-indexed calibrated semantic representation approximately does and supports an externally specified Bayesian recursion\. It thereby illustrates the distinction between closure failure of one representation and successful descent of a better one\.
## 5Completed bounded validation
### 5\.1Questions and design
The theory is population\-level, but its central conditions have observable consequences\. The validation therefore asks three deliberately ordered questions\. First, is the proposed lexical observation complete enough to be used at all? Second, do information\-equivalent prompts induce sufficiently similar measurements at the representation being tested? Third, after calibration and recursive updating, do held\-out path errors obey their frozen coverage bound? A positive answer to the third question is meaningful only if the first two are answered at the same representation level\.
This ordering is important\. It would be easy to report only the calibrated result and conclude that language probabilities are stable\. The raw\-kernel experiment shows why that conclusion would be false: the same evidence can produce a large prompt\-specific discrepancy before semantic calibration\. Conversely, failure of the raw kernel need not invalidate every coarser representation\. The descent criterion asks whether the*chosen*state, not every upstream lexical quantity, is closed under the declared equivalence\.
All three analyses used archived observations and frozen design files\. No threshold, partition or prompt was changed after inspection of the untouched test results\. The state space was\{risk\-on,mixed,risk\-off\}\\\{\\text\{risk\-on\},\\text\{mixed\},\\text\{risk\-off\}\\\}; the sequential experiments used six lexical evidence symbols, two information\-equivalent prompt constructions and eight updates per path\. “Unsure” was retained as an open\-world measurement component and was not silently conditioned away\. Table[1](https://arxiv.org/html/2609.20207#S5.T1)records the role of each experiment\.
Table 1:Completed validation design\. Each row tests a different logical level of the theory; failure at one level is not relabelled as success at another\.
### 5\.2Results
Table[2](https://arxiv.org/html/2609.20207#S5.T2)reports every principal frozen gate, including failures\. The complete lexical\-process claim was not supported\. Although the observed candidate mass, average contraction, uniform path coverage and termwise perturbation bound passed, the active design was nearly singular and the convergence\-slope interval included zero\. This is direct evidence against claiming a generally identified lexical process from that design\.
The raw sequential filter gave a more focused diagnosis\. Its emission system had full rank, all test normalizers were positive, and empirical path coverage was exactly 0\.90 at nominal level 0\.90\. Nevertheless, the maximum cross\-prompt Jensen–Shannon discrepancy was 0\.13397, above the frozen 0\.10 gate\. The conjunctive claim therefore failed\. The failure says that raw prompt\-conditioned language probabilities cannot be treated as an information\-invariant state merely because their average discrepancy is small\.
Table 2:Frozen validation results\. “Pass” applies only to the stated gate and operating domain\. The final column is the conclusion permitted by the conjunction of gates, not a post hoc interpretation of individual favorable statistics\.GateStatisticResultInterpretationObserved candidate massminimum0\.9999998080\.999999808PassNegligible missing mass in the sampled candidate system\.Lexical\-process identificationminimum singular value2\.96×10−162\.96\\times 10^\{\-16\}FailThe active design does not identify the proposed process\.Average contractionmean log ratio−0\.1686\-0\.1686; interval\[−0\.2033,−0\.1366\]\[\-0\.2033,\-0\.1366\]PassSupported on the sampled support\.Convergence trendlog–log slope−0\.0357\-0\.0357; interval\[−0\.0741,0\.0042\]\[\-0\.0741,0\.0042\]FailNo strictly negative convergence slope established\.Uniform path recovery0\.910\.91over 100 paths; 95% Wilson interval\[0\.838,0\.952\]\[0\.838,0\.952\]PassCovers the nominal 0\.90 level in this experiment\.Termwise perturbation bound0 violations in 100 pathsPassObserved transformed increments obeyed the computed bound\.Raw prompt invariancemax JS0\.133970\.13397versus limit0\.100\.10FailRaw lexical kernels do not descend through prompt equivalence\.Raw sequential coverage36/40=0\.9036/40=0\.90pathsPassCoverage alone cannot rescue the failed conjunctive claim\.Calibrated prompt invariancemean/max JS0\.001122/0\.0107200\.001122/0\.010720PassCommon posterior meaning is stable for the two frozen prompts\.Calibrated sequential coverage28/30=0\.93328/30=0\.933paths at nominal0\.900\.90PassFrozen pathwise radius0\.1906690\.190669covers held\-out paths\.Prompt\-specific affine calibration was then fitted on a disjoint partition and frozen\. In the common posterior representation, the maximum cross\-prompt discrepancy fell to 0\.010720 and 28 of 30 untouched paths lay within the frozen bound, giving coverage 0\.933\. The calibrated experiment passed all of its declared gates\. In the language of Theorem[3\.1](https://arxiv.org/html/2609.20207#S3.Thmdefinition1), the result supports approximate descent for this representation, transformation family and horizon\. It does not establish raw\-kernel invariance, global identification of the broader lexical process, or validity outside the frozen prompt and evidence classes\.
### 5\.3What has and has not been validated
The completed experiments support one positive, bounded statement: within the frozen three\-state system, prompt\-specific calibration produced a common semantic representation whose information\-equivalent prompt variants were stable and whose eight\-step posterior errors achieved the prescribed held\-out coverage\. They also support two negative statements: the raw prompt kernel is not invariant at the frozen maximum\-discrepancy threshold, and the broader lexical\-process design is not identified\.
These outcomes are scientifically useful because they distinguish a repairable representation failure from a universal impossibility\. Calibration changes the representation at which descent is assessed; it does not prove that the upstream language law was invariant\. Replication across transformation families, fitted models and service epochs remains necessary before any broad empirical claim about lexical calculus\. Those replications are future work, not part of the evidence reported here\.
## 6Discussion and conclusion
### 6\.1Discussion
The construction answers a practical question: when may a probability\-valued interpretation of language be treated as a state rather than merely a score? The answer is not calibration alone\. A state must also be closed, at least approximately, under the transformations that drive its proposed dynamics\.
This criterion prevents a common modelling error\. A smooth time series can always be fitted to successive semantic scores, but such a fit does not establish that the present score contains enough information to determine the next update\. Fibre preservation supplies the missing sufficiency condition\. Approximate fibre diameter quantifies the cost of violating it\.
The results separate four requirements that are easily conflated in applications\. Calibration asks whether the reported semantic coordinates agree with a declared reference task\. Behavioral separation asks whether scientifically distinct mechanisms induce distinguishable observable signatures\. Closure asks whether a declared language transformation has a well\-defined action on the retained semantic coordinates\. Minimality asks whether any strictly coarser interpretable representation retains these properties\. None of the first three implies the others, and a representation that passes them need not be minimal\. The terminal representation theorem explains the ideal population target; finite\-sample recovery theory is reserved for the supplementary extensions\.
This end\-to\-end distinction matters for stochastic modelling\. Calibration alone permits two contexts with the same reported state to respond differently to the next admissible information transformation\. Identifiability alone permits a representation that is unnecessarily large\. Closure alone permits a stable but scientifically meaningless quotient\. Only their conjunction supports the interpretation of the retained coordinates as an empirically recoverable state on which a stochastic recursion can be defined without silently reintroducing discarded language information\.
Several limitations are deliberate\. Transformations are declared rather than discovered; informational equivalence is task relative; the semantic map may depend on the observation mechanism; and finite access yields partial observation of a continuation law\. These are not reasons to abandon the calculus\. They identify the statistical errors that an empirical implementation must report\.
The framework also does not guarantee that a smallest admissible representation exists in every user\-chosen model class\. Existence follows for the canonical behavioral quotient under the stated measurability construction, while a restricted interpretable family requires non\-emptiness and a unique minimal element\. Failure of either condition is a substantive diagnostic: the transformation family may demand additional semantic coordinates, the observations may not separate competing representations, or the proposed ontology may not be stable enough to support autonomous dynamics\.
### 6\.2Conclusion
Lexical calculus is defined here narrowly: it studies observable contextual transformation systems and the conditions under which their actions descend to a retained semantic state\. The population theory constructs the canonical closed behavioral representation, characterizes exact descent by fibre preservation, bounds the unavoidable error of approximate descent, and shows how finite defects propagate\. Once this gate has been passed, average contraction supplies a unique causal external recursion and synchronization; simplex coordinates do not alter those conclusions\.
The empirical evidence respects the same order\. It does not support the broad claim that a raw lexical process is identified, nor that raw prompt kernels are information invariant\. It does support a smaller claim: within one frozen three\-state experiment, prompt\-specific calibration produced a common semantic representation with stable prompt meaning and 0\.933 finite\-horizon coverage at nominal level 0\.90\. The failed gates are as important as the passed ones because they show that calibration at one time point is not enough to justify a state recursion\.
The practical implication is a falsifiable workflow rather than a metaphor\. Declare the information equivalence and transformation family; construct or estimate a semantic map; test fibre preservation, identification and coverage; and only then use the resulting coordinates as a stochastic state\. This is the precise sense in which the framework is foundational\. It does not attribute a hidden calculus to a language model\. It states the mathematical and empirical conditions under which an observable language\-derived representation can legitimately support one\.
## Ancillary theoretical material
The following appendix records the typed measurable and coalgebraic foundations needed to support the paper’s structural claims\. It then gives compact statements of how ontology choice and statistical recovery enter the construction\.
## Appendix ATyped partial measurable systems and coalgebraic recovery
The one\-space construction hides two features that matter in applications\. Not every operation is legal everywhere, and operations may move between different sorts, such as a structured record, a rendered prompt and a probability\-valued readout\. Category and restriction\-category language provides bookkeeping for these domains and sorts\. It prevents ill\-typed compositions rather than adding abstraction for its own sake\.
The construction is not intrinsically linguistic\. We first define it for a general observable transformation system and later specialize contexts to natural language\. LetIIbe a countable set of sorts and let𝐂=\(\(Ci,𝒜i\)\)i∈I\\mathbf\{C\}=\(\(C\_\{i\},\\mathcal\{A\}\_\{i\}\)\)\_\{i\\in I\}be measurable spaces\. Fori,j∈Ii,j\\in I, a partial measurable arrowT:i⇢jT:i\\dashrightarrow jis an equivalence class of pairs\(DT,t\)\(D\_\{T\},t\), whereDT∈𝒜iD\_\{T\}\\in\\mathcal\{A\}\_\{i\}andt:\(DT,𝒜i\|DT\)→\(Cj,𝒜j\)t:\(D\_\{T\},\\mathcal\{A\}\_\{i\}\|\_\{D\_\{T\}\}\)\\to\(C\_\{j\},\\mathcal\{A\}\_\{j\}\)is measurable; pairs are identified when they have the same domain and agree pointwise\. ForT:i⇢jT:i\\dashrightarrow jandU:j⇢kU:j\\dashrightarrow k, set
DU∘T:=DT∩t−1\(DU\),\(U∘T\)\(c\):=u\(t\(c\)\)\.D\_\{U\\circ T\}:=D\_\{T\}\\cap t^\{\-1\}\(D\_\{U\}\),\\qquad\(U\\circ T\)\(c\):=u\(t\(c\)\)\.The identity1i:i⇢i1\_\{i\}:i\\dashrightarrow iis\(Ci,idCi\)\(C\_\{i\},\\operatorname\{id\}\_\{C\_\{i\}\}\)\. Associativity follows from associativity of ordinary composition and equality of the displayed domains\. These data form the many\-sorted partial\-map category𝐏𝐚𝐫𝐌𝐞𝐚𝐬\(𝐂\)\\mathbf\{ParMeas\}\(\\mathbf\{C\}\)\. An*admissible transformation category*𝒯\\mathcal\{T\}is a small wide subcategory of𝐏𝐚𝐫𝐌𝐞𝐚𝐬\(𝐂\)\\mathbf\{ParMeas\}\(\\mathbf\{C\}\)\.
For every sortjj, let𝒪j\\mathcal\{O\}\_\{j\}be a countable family of measurable mapsF:Cj→VFF:C\_\{j\}\\to V\_\{F\}, where eachVFV\_\{F\}is standard Borel\. ForT:i⇢jT:i\\dashrightarrow j, introduce the pointed spaceVF∂:=VF⊔\{∂\}V\_\{F\}^\{\\partial\}:=V\_\{F\}\\sqcup\\\{\\partial\\\}and define
\[F,T\]\(c\):=\{F\(Tc\),c∈DT,∂,c∉DT\.\[F,T\]\(c\):=\\begin\{cases\}F\(Tc\),&c\\in D\_\{T\},\\\\ \\partial,&c\\notin D\_\{T\}\.\\end\{cases\}The cemetery value records inadmissibility and therefore distinguishes an undefined operation from a defined operation whose observable value happens to be zero\.
###### Definition A\.1\(Behavioral signature and quotient sigma\-algebra\)\.
For sortii, let
Pi:=∏j∈I∏T∈𝒯\(i,j\)∏F∈𝒪jVF∂P\_\{i\}:=\\prod\_\{j\\in I\}\\prod\_\{T\\in\\mathcal\{T\}\(i,j\)\}\\prod\_\{F\\in\\mathcal\{O\}\_\{j\}\}V\_\{F\}^\{\\partial\}with its product sigma\-algebra, and define
Si:Ci⟶Pi,Si\(c\):=\(\[F,T\]\(c\)\)j,T,F\.S\_\{i\}:C\_\{i\}\\longrightarrow P\_\{i\},\\qquad S\_\{i\}\(c\):=\\bigl\(\[F,T\]\(c\)\\bigr\)\_\{j,T,F\}\.WriteZi∗:=Si\(Ci\)Z\_\{i\}^\{\*\}:=S\_\{i\}\(C\_\{i\}\)and equip it with the final sigma\-algebra
𝒵i∗:=\{A⊆Zi∗:Si−1\(A\)∈𝒜i\}\.\\mathcal\{Z\}\_\{i\}^\{\*\}:=\\\{A\\subseteq Z\_\{i\}^\{\*\}:S\_\{i\}^\{\-1\}\(A\)\\in\\mathcal\{A\}\_\{i\}\\\}\.ThusSi:Ci↠Zi∗S\_\{i\}:C\_\{i\}\\twoheadrightarrow Z\_\{i\}^\{\*\}is a measurable quotient map by construction\.
###### Definition A\.2\(Closed observable quotient representation\)\.
A closed representation of\(𝐂,𝒯,𝒪\)\(\\mathbf\{C\},\\mathcal\{T\},\\mathcal\{O\}\)is a family of measurable quotient mapsRi:Ci↠ZiR\_\{i\}:C\_\{i\}\\twoheadrightarrow Z\_\{i\}, whereZiZ\_\{i\}carries the final sigma\-algebra induced byRiR\_\{i\}, together with:
1. 1\.for everyT:i⇢jT:i\\dashrightarrow j, a partial measurable mapT¯:Zi⇢Zj\\overline\{T\}:Z\_\{i\}\\dashrightarrow Z\_\{j\}satisfying DT=Ri−1\(DT¯\),Rj∘T=T¯∘RionDT;D\_\{T\}=R\_\{i\}^\{\-1\}\(D\_\{\\overline\{T\}\}\),\\qquad R\_\{j\}\\circ T=\\overline\{T\}\\circ R\_\{i\}\\quad\\text\{on \}D\_\{T\};
2. 2\.for everyF∈𝒪iF\\in\\mathcal\{O\}\_\{i\}, a measurable readoutfF:Zi→VFf\_\{F\}:Z\_\{i\}\\to V\_\{F\}satisfyingF=fF∘RiF=f\_\{F\}\\circ R\_\{i\}\.
A morphismh:R→R′h:R\\to R^\{\\prime\}is a family of measurable mapshi:Zi→Zi′h\_\{i\}:Z\_\{i\}\\to Z\_\{i\}^\{\\prime\}such thatRi′=hi∘RiR\_\{i\}^\{\\prime\}=h\_\{i\}\\circ R\_\{i\}\. This equality forces preservation of readouts, domains and induced transitions\. Closed representations and these morphisms form a category𝐑𝐞𝐩\(𝐂,𝒯,𝒪\)\\mathbf\{Rep\}\(\\mathbf\{C\},\\mathcal\{T\},\\mathcal\{O\}\)\.
The orientation of morphisms is deliberate: an arrowR→R′R\\to R^\{\\prime\}means thatR′R^\{\\prime\}is a measurable coarsening ofRR\.
###### Theorem A\.3\(Terminal behavioral representation—typed measurable extension\)\.
AssumeII, the hom\-sets of𝒯\\mathcal\{T\}, and the observable families are countable\. ThenS=\(Si\)i∈IS=\(S\_\{i\}\)\_\{i\\in I\}is a closed representation and is terminal in𝐑𝐞𝐩\(𝐂,𝒯,𝒪\)\\mathbf\{Rep\}\(\\mathbf\{C\},\\mathcal\{T\},\\mathcal\{O\}\)\. Consequently:
1. 1\.everyT:i⇢jT:i\\dashrightarrow jhas a unique induced partial measurable mapT¯∗:Zi∗⇢Zj∗\\overline\{T\}\_\{\*\}:Z\_\{i\}^\{\*\}\\dashrightarrow Z\_\{j\}^\{\*\};
2. 2\.every selected observable factors uniquely throughSS;
3. 3\.every closed representationRRadmits a unique measurable morphismR→SR\\to S;
4. 4\.the terminal representation is unique up to unique measurable isomorphism\.
HenceSSis the coarsest measurable quotient that preserves admissibility, selected observations and all future behavior under𝒯\\mathcal\{T\}\.
###### Proof\.
Countability and componentwise measurability make everySiS\_\{i\}measurable\. SupposeSi\(c\)=Si\(c′\)S\_\{i\}\(c\)=S\_\{i\}\(c^\{\\prime\}\)\. The cemetery coordinates implyc∈DTc\\in D\_\{T\}if and only ifc′∈DTc^\{\\prime\}\\in D\_\{T\}\. IfT:i⇢jT:i\\dashrightarrow jis admissible at both points, then for everyU:j⇢kU:j\\dashrightarrow kandF∈𝒪kF\\in\\mathcal\{O\}\_\{k\},
\[F,U\]\(Tc\)=\[F,U∘T\]\(c\)=\[F,U∘T\]\(c′\)=\[F,U\]\(Tc′\)\.\[F,U\]\(Tc\)=\[F,U\\circ T\]\(c\)=\[F,U\\circ T\]\(c^\{\\prime\}\)=\[F,U\]\(Tc^\{\\prime\}\)\.ThusSj\(Tc\)=Sj\(Tc′\)S\_\{j\}\(Tc\)=S\_\{j\}\(Tc^\{\\prime\}\), and
T¯∗\(Si\(c\)\):=Sj\(Tc\),DT¯∗:=Si\(DT\),\\overline\{T\}\_\{\*\}\(S\_\{i\}\(c\)\):=S\_\{j\}\(Tc\),\\qquad D\_\{\\overline\{T\}\_\{\*\}\}:=S\_\{i\}\(D\_\{T\}\),is representative independent\. Domain saturation follows from the same cemetery coordinate\. To prove measurability, letA∈𝒵j∗A\\in\\mathcal\{Z\}\_\{j\}^\{\*\}\. By the definition of the final sigma\-algebra,
Si−1\(T¯∗−1\(A\)\)=DT∩T−1\(Sj−1\(A\)\)∈𝒜i;S\_\{i\}^\{\-1\}\\\!\\left\(\\overline\{T\}\_\{\*\}^\{\-1\}\(A\)\\right\)=D\_\{T\}\\cap T^\{\-1\}\(S\_\{j\}^\{\-1\}\(A\)\)\\in\\mathcal\{A\}\_\{i\};henceT¯∗−1\(A\)∈𝒵i∗\\overline\{T\}\_\{\*\}^\{\-1\}\(A\)\\in\\mathcal\{Z\}\_\{i\}^\{\*\}\. The identity coordinate gives the readout factorization, soSSis closed\.
LetRRbe any closed representation\. Definehi:Zi→Zi∗h\_\{i\}:Z\_\{i\}\\to Z\_\{i\}^\{\*\}byhi\(Ri\(c\)\)=Si\(c\)h\_\{i\}\(R\_\{i\}\(c\)\)=S\_\{i\}\(c\)\. IfRi\(c\)=Ri\(c′\)R\_\{i\}\(c\)=R\_\{i\}\(c^\{\\prime\}\), domain preservation, repeated transition closure and readout factorization imply\[F,T\]\(c\)=\[F,T\]\(c′\)\[F,T\]\(c\)=\[F,T\]\(c^\{\\prime\}\)for every coordinate; hencehih\_\{i\}is well defined\. SinceSi=hi∘RiS\_\{i\}=h\_\{i\}\\circ R\_\{i\}andZiZ\_\{i\}has the final sigma\-algebra ofRiR\_\{i\},hih\_\{i\}is measurable\. Surjectivity ofRiR\_\{i\}gives uniqueness\. ThusSSis terminal\. The final assertion is the usual uniqueness of terminal objects\. ∎
### A\.1Restriction\-category structure
Categories of partial maps carry a canonical restriction operation\[[11](https://arxiv.org/html/2609.20207#bib.bib11)\]\. ForT=\(DT,t\):i⇢jT=\(D\_\{T\},t\):i\\dashrightarrow j, define
T¯:=\(DT,idDT\):i⇢i\.\\overline\{T\}:=\(D\_\{T\},\\operatorname\{id\}\_\{D\_\{T\}\}\):i\\dashrightarrow i\.This partial identity records exactly whereTTis defined\.
###### Proposition A\.5\(Restriction structure\[[11](https://arxiv.org/html/2609.20207#bib.bib11)\]\)\.
The category𝐏𝐚𝐫𝐌𝐞𝐚𝐬\(𝐂\)\\mathbf\{ParMeas\}\(\\mathbf\{C\}\)is a restriction category\. With ordinary right\-to\-left composition, its restriction operation satisfies
T∘T¯\\displaystyle T\\circ\\overline\{T\}=T,\\displaystyle=T,T¯∘U¯\\displaystyle\\overline\{T\}\\circ\\overline\{U\}=U¯∘T¯\\displaystyle=\\overline\{U\}\\circ\\overline\{T\}whenT,Uhave the same source,\\displaystyle\\text\{when $T,U$ have the same source\},T∘U¯¯\\displaystyle\\overline\{T\\circ\\overline\{U\}\}=T¯∘U¯\\displaystyle=\\overline\{T\}\\circ\\overline\{U\}whenT,Uhave the same source,\\displaystyle\\text\{when $T,U$ have the same source\},U¯∘T\\displaystyle\\overline\{U\}\\circ T=T∘U∘T¯\\displaystyle=T\\circ\\overline\{U\\circ T\}whenU∘Tis typed\.\\displaystyle\\text\{when $U\\circ T$ is typed\}\.Every admissible transformation category𝒯\\mathcal\{T\}that contains the restrictions of its arrows is therefore a restriction subcategory\. Moreover, the induced partial maps in every closed representation preserve restrictions:
T¯∗¯=\(T¯\)∗¯\.\\overline\{\\,\\overline\{T\}\_\{\*\}\\,\}=\\overline\{\(\\overline\{T\}\)\_\{\*\}\}\.
###### Proof\.
The first identity holds because restricting toDTD\_\{T\}before applyingTTchanges neither its domain nor its values\. The second follows because both composites are the partial identity onDT∩DUD\_\{T\}\\cap D\_\{U\}\. The third has domainDT∩DUD\_\{T\}\\cap D\_\{U\}and is the identity there\. For the fourth, both sides have domain
DT∩T−1\(DU\)D\_\{T\}\\cap T^\{\-1\}\(D\_\{U\}\)and agree withTTon that domain\. These are the restriction axioms\. A closed representation satisfiesDT=Ri−1\(DT¯∗\)D\_\{T\}=R\_\{i\}^\{\-1\}\(D\_\{\\overline\{T\}\_\{\*\}\}\); hence the restriction of the induced arrow is the partial identity on the image of the saturated domain, which is precisely the arrow induced byT¯\\overline\{T\}\. ∎
The restriction formulation is not an alternative theory\. It identifies the established categorical structure underlying admissibility and clarifies the new step: constructing a terminal observable quotient inside a measurable, many\-sorted restriction system\.
### A\.2Recovery of the classical final\-coalgebra behavior map
Consider the total, one\-sorted case\. LetAAbe a countable action alphabet, let\(O,𝒪\)\(O,\\mathcal\{O\}\)be a measurable output space, and let
H\(X\):=O×XAH\(X\):=O\\times X^\{A\}on measurable spaces and measurable maps, with countable\-product sigma\-algebras\. A measurable Moore system is anHH\-coalgebra
γ:C⟶O×CA,γ\(c\)=\(o\(c\),\(Tac\)a∈A\)\.\\gamma:C\\longrightarrow O\\times C^\{A\},\\qquad\\gamma\(c\)=\\bigl\(o\(c\),\(T\_\{a\}c\)\_\{a\\in A\}\\bigr\)\.Forw=a1⋯an∈A∗w=a\_\{1\}\\cdots a\_\{n\}\\in A^\{\*\}, writeTw=Tan∘⋯∘Ta1T\_\{w\}=T\_\{a\_\{n\}\}\\circ\\cdots\\circ T\_\{a\_\{1\}\}andTϵ=1CT\_\{\\epsilon\}=1\_\{C\}\.
###### Theorem A\.6\(Coalgebraic recovery; standard Moore behavior theorem\[[29](https://arxiv.org/html/2609.20207#bib.bib29),[32](https://arxiv.org/html/2609.20207#bib.bib32)\]\)\.
The measurable spaceOA∗O^\{A^\{\*\}\}, equipped with
ζ\(q\)=\(q\(ϵ\),\(qa\)a∈A\),qa\(w\):=q\(aw\),\\zeta\(q\)=\\bigl\(q\(\\epsilon\),\(q\_\{a\}\)\_\{a\\in A\}\\bigr\),\\qquad q\_\{a\}\(w\):=q\(aw\),is a finalHH\-coalgebra\. Its unique coalgebra morphism from\(C,γ\)\(C,\\gamma\)is
β:C⟶OA∗,β\(c\)\(w\)=o\(Twc\)\.\\beta:C\\longrightarrow O^\{A^\{\*\}\},\\qquad\\beta\(c\)\(w\)=o\(T\_\{w\}c\)\.If the selected observable family is\{o\}\\\{o\\\}and the transformation category is the action ofA∗A^\{\*\}, then the lexical behavioral signatureSSis exactlyβ\\beta\. Consequently its terminal quotient agrees with the ordinary Moore behavioral quotient and, for Boolean language acceptance, with the Myhill–Nerode quotient\.
###### Proof\.
Countability ofA∗A^\{\*\}makesOA∗O^\{A^\{\*\}\}measurable with the product sigma\-algebra, and every coordinate ofζ\\zetais a coordinate projection, soζ\\zetais measurable\. The displayedβ\\betais measurable coordinatewise\. It is a coalgebra morphism because
β\(c\)\(ϵ\)=o\(c\),β\(Tac\)\(w\)=o\(TwTac\)=β\(c\)\(aw\)\.\\beta\(c\)\(\\epsilon\)=o\(c\),\\qquad\\beta\(T\_\{a\}c\)\(w\)=o\(T\_\{w\}T\_\{a\}c\)=\\beta\(c\)\(aw\)\.Iff:C→OA∗f:C\\to O^\{A^\{\*\}\}is any coalgebra morphism, its output equation givesf\(c\)\(ϵ\)=o\(c\)f\(c\)\(\\epsilon\)=o\(c\)and its transition equation givesf\(c\)\(aw\)=f\(Tac\)\(w\)f\(c\)\(aw\)=f\(T\_\{a\}c\)\(w\)\. Induction on word length yieldsf\(c\)\(w\)=o\(Twc\)f\(c\)\(w\)=o\(T\_\{w\}c\), sof=βf=\\beta\. Finally, the coordinates ofS\(c\)S\(c\)are precisely\(o\(Twc\)\)w∈A∗\(o\(T\_\{w\}c\)\)\_\{w\\in A^\{\*\}\}; henceS=βS=\\beta, and equality of signatures is classical future\-output equivalence\. ∎
###### Proposition A\.7\(Strictness beyond total Moore systems\)\.
The framework contains countable\-action measurable Moore systems behaviorally faithfully, by Theorem[A\.6](https://arxiv.org/html/2609.20207#A1.Thmdefinition6), but it is strictly more expressive when transformation admissibility is observable\. In particular, letC=\{c0,c1\}C=\\\{c\_\{0\},c\_\{1\}\\\}, let the only readout be constant, and letTTbe defined only atc0c\_\{0\}\. The behavioral signature separatesc0c\_\{0\}andc1c\_\{1\}through the cemetery coordinate\. No total Moore system on the same carrier, with the same constant readout and action label, can represent this distinction\. It can do so only by changing the model—for example by adjoining a failure state or an admissibility readout\.
###### Proof\.
The inclusion of total Moore systems and preservation of their behavioral equivalence follow from Theorem[A\.6](https://arxiv.org/html/2609.20207#A1.Thmdefinition6)\. In the partial system,\[F,T\]\(c0\)=F\(Tc0\)\[F,T\]\(c\_\{0\}\)=F\(Tc\_\{0\}\)while\[F,T\]\(c1\)=∂\[F,T\]\(c\_\{1\}\)=\\partial, soS\(c0\)≠S\(c1\)S\(c\_\{0\}\)\\neq S\(c\_\{1\}\)\. In a total Moore system on the same two points,TTmust be defined at both\. Since the present readout is constant and no admissibility coordinate exists, the undefined\-versus\-defined distinction cannot be expressed\. Adding a failure state or readout changes the carrier or observable signature, proving strictness relative to the stated total theory\. ∎
## Appendix BFurther remarks on ontology adequacy
Section 3 states the adequacy condition required by the main theorem chain\. This appendix records its limited model\-selection consequence without developing a separate theory of ontology learning\.
LetB:𝒞→SB:\\mathcal\{C\}\\to Sbe a candidate semantic representation and let𝔗0⊆𝔗\\mathfrak\{T\}\_\{0\}\\subseteq\\mathfrak\{T\}be the transformations relevant to the intended dynamics\. CallBB*transformation sufficient*for𝔗0\\mathfrak\{T\}\_\{0\}when everyT∈𝔗0T\\in\\mathfrak\{T\}\_\{0\}descends throughBB\. By Theorem[3\.1](https://arxiv.org/html/2609.20207#S3.Thmdefinition1), this is equivalent to
B\(c\)=B\(c′\)⟹B\(Tc\)=B\(Tc′\)\(T∈𝔗0\)\.B\(c\)=B\(c^\{\\prime\}\)\\quad\\Longrightarrow\\quad B\(Tc\)=B\(Tc^\{\\prime\}\)\\qquad\(T\\in\\mathfrak\{T\}\_\{0\}\)\.The approximate version replaces equality by the fibre\-diameter bound of Theorem[3\.3](https://arxiv.org/html/2609.20207#S3.Thmdefinition3)\. Thus ontology adequacy has a direct observable implication: contexts assigned the same current state should have sufficiently similar transformed states\.
If a candidate state fails this condition, there are two principled responses\. One may refine the state so that the offending contexts are no longer merged, or restrict the transformation family to the operations for which closure is actually required\. Refinement can remove a particular closure obstruction by retaining the distinction that caused it\. Coarsening cannot remove that obstruction without also discarding the transformed distinction\. Conversely, unnecessary refinement raises dimension and may make calibration unstable\. The practical target is therefore the least complex scientifically interpretable representation that passes calibration, identification, and approximate\-closure gates on a declared operating domain\. Existence or uniqueness of such a minimal representation is not asserted without specifying and testing the candidate family\.
## Appendix CFrom an observable language law to an estimable state
The abstract calculus begins with a state mapBB, whereas an application observes a conditional probability law over verbal continuations\. The statistical bridge is the composition
𝒞→𝖪𝒫\(𝒲\)→ΠΔK−1→𝜓ΔK−1\.\\mathcal\{C\}\\xrightarrow\{\\;\\mathsf\{K\}\\;\}\\mathcal\{P\}\(\\mathcal\{W\}\)\\xrightarrow\{\\;\\Pi\\;\}\\Delta^\{K\-1\}\\xrightarrow\{\\;\\psi\\;\}\\Delta^\{K\-1\}\.Here𝖪c\\mathsf\{K\}\_\{c\}is the observable continuation law,Π\\Pigroups prespecified meaning\-equivalent continuations, andψ\\psiis a finite\-dimensional calibration map estimated on data disjoint from evaluation\. The continuation law is otherwise unrestricted, so the construction is semiparametric: the language component is an unknown probability kernel, while only the low\-dimensional inverse calibration is parametrized\. The resultingB\(c\)=ψ\{Π\(𝖪c\)\}B\(c\)=\\psi\\\{\\Pi\(\\mathsf\{K\}\_\{c\}\)\\\}is the state to which the descent and stochastic stability results apply\.
This construction separates three questions that should not be conflated\. First,*measurement completeness*asks whether the selected continuations retain enough probability mass\. Second,*statistical identification*asks whether distinct reference states induce distinguishable values of the grouped language law on the operating domain\. Third,*dynamic closure*asks whether the calibrated state retains enough information to predict the effect of subsequent transformations\. Accurate one\-time calibration does not imply the third property\.
For clarity, suppose the grouped observable isq\(c\)q\(c\)and the calibration model isψθ\(q\)\\psi\_\{\\theta\}\(q\)\. Conditional identification on a compact operating set𝒬\\mathcal\{Q\}requires
ψθ1\(q\)=ψθ2\(q\)for allq∈𝒬⟹θ1=θ2,\\psi\_\{\\theta\_\{1\}\}\(q\)=\\psi\_\{\\theta\_\{2\}\}\(q\)\\quad\\text\{for all \}q\\in\\mathcal\{Q\}\\quad\\Longrightarrow\\quad\\theta\_\{1\}=\\theta\_\{2\},or the corresponding local full\-rank condition for a differentiable model\. Stable recovery additionally requires the inverse modulus not to approach zero\. These are assumptions about the declared measurement design, not properties granted by fluency or by a model’s printed numerical confidence\. The completed experiments accordingly test retained mass, a lower identification modulus, perturbation error, and held\-out sequential coverage\.
The main paper does not require a universal estimator forψ\\psi\. It requires only a frozen estimator whose error can enter Theorem[4\.3](https://arxiv.org/html/2609.20207#S4.Thmdefinition3)and whose induced closure defects enter Theorem[4\.8](https://arxiv.org/html/2609.20207#S4.Thmdefinition8)\. Full asymptotic theory for nonlinear semiparametric recovery is therefore supporting material rather than part of the foundational theorem chain; it is developed in\[[18](https://arxiv.org/html/2609.20207#bib.bib18)\]\. The measurement design and its held\-out certification are developed separately in\[[17](https://arxiv.org/html/2609.20207#bib.bib17)\]\.
## Appendix DBoundary with adjacent theories
The construction uses established ideas at three levels but addresses a different junction between them\. Future\-behaviour quotients, automata, and coalgebra explain how a minimal state can be constructed from observable responses\[[24](https://arxiv.org/html/2609.20207#bib.bib24),[32](https://arxiv.org/html/2609.20207#bib.bib32)\]\. Random dynamical systems and iterated random functions establish existence and synchronization once a state and its random update maps are given\[[1](https://arxiv.org/html/2609.20207#bib.bib1),[16](https://arxiv.org/html/2609.20207#bib.bib16)\]\. Semantic uncertainty methods group or calibrate language probabilities into application\-relevant meanings\[[20](https://arxiv.org/html/2609.20207#bib.bib20),[26](https://arxiv.org/html/2609.20207#bib.bib26)\]\.
The missing link is whether a prompt\-dependent language law descends to a state on which stochastic evolution is well defined\. The contribution here is to make that link explicit: semantic descent supplies the admissible state; its fibre defect quantifies irreducible nonclosure; and random\-iteration theory supplies the stochastic recursion only after that defect has been controlled\. The term*stochastic lexical calculus*names this combined construction, not a replacement for probability theory, stochastic calculus, or existing linguistic calculi\.
## References
- \[1\]L\. Arnold\.*Random Dynamical Systems*\. Springer, Berlin, 1998\.
- \[2\]P\. Baldan, F\. Bonchi, H\. Kerstan, and B\. König\. Coalgebraic behavioral metrics\.*Logical Methods in Computer Science*, 14\(3\), 2018\.
- \[3\]N\. Band, X\. Li, T\. Ma, and T\. Hashimoto\. Linguistic calibration of long\-form generations\.*Proceedings of the 41st International Conference on Machine Learning*, 235:2732–2778, 2024\.
- \[4\]F\. Bianchi, D\. Nozza, and D\. Hovy\. Language invariant properties in natural language processing\.*Proceedings of NLP Power\!*, pages 84–92, 2022\.
- \[5\]D\. Blackwell\. Equivalent comparisons of experiments\.*Annals of Mathematical Statistics*, 24\(2\):265–272, 1953\.
- \[6\]R\. F\. Blute, J\. R\. B\. Cockett, and R\. A\. G\. Seely\. Differential categories\.*Mathematical Structures in Computer Science*, 16\(6\):1049–1083, 2006\.
- \[7\]F\. Bonchi, B\. König, and D\. Petrişan\. Up\-to techniques for behavioural metrics via fibrations\.*Mathematical Structures in Computer Science*, 33\(4–5\):182–221, 2023\.
- \[8\]J\. A\. Brzozowski\. Derivatives of regular expressions\.*Journal of the ACM*, 11\(4\):481–494, 1964\.
- \[9\]S\. Chatzikyriakidis and R\. Cooper\. Mathematical structures in natural language semantics: Zawadowski’s contribution to linguistics\.*Mathematical Structures in Computer Science*, 36:e16, 1–12, 2026\.
- \[10\]S\. Clark\. Vector space models of lexical meaning\. In S\. Lappin and C\. Fox, editors,*The Handbook of Contemporary Semantic Theory*\. Wiley, 2016\.
- \[11\]J\. R\. B\. Cockett and S\. Lack\. Restriction categories I: Categories of partial maps\.*Theoretical Computer Science*, 270\(1–2\):223–259, 2002\.
- \[12\]R\. Cockett, G\. Cruttwell, M\. Kerjean, and J\.\-S\. Pacaud Lemay\. Foreword for the special issue “Differential Structures in Computer Science and Mathematics\.”*Mathematical Structures in Computer Science*, 35:e27, 2025\.
- \[13\]M\. Cousin\. Meaning–Text Theory within abstract categorial grammars: Toward paraphrase and lexical function modeling for text generation\.*Proceedings of IWCS*, pages 134–143, 2023\.
- \[14\]Data representation and lexical calculi\.*Information Processing & Management*, 20\(1–2\):151–174, 1984\.
- \[15\]J\. Desharnais, A\. Edalat, and P\. Panangaden\. Bisimulation for labelled Markov processes\.*Information and Computation*, 179\(2\):163–193, 2002\.
- \[16\]P\. Diaconis and D\. Freedman\. Iterated random functions\.*SIAM Review*, 41\(1\):45–76, 1999\.
- \[17\]M\. F\. Dixon\. Calibrating semantic uncertainty from observable language\-model probabilities\.*arXiv:2607\.17447*, 2026\.
- \[18\]M\. Dixon\. Identification and learning of semantic observation kernels: Partial observation, uniform recovery, and minimax limits\. Companion manuscript, 2026\.
- \[19\]T\. Ehrhard and L\. Régnier\. The differential lambda\-calculus\.*Theoretical Computer Science*, 309:1–41, 2003\.
- \[20\]S\. Farquhar, J\. Kossen, L\. Kuhn, and Y\. Gal\. Detecting hallucinations in large language models using semantic entropy\.*Nature*, 630:625–630, 2024\.
- \[21\]T\. Fritz, P\. Perrone, and S\. Rezagholi\. Probability, valuations, hyperspace: Three monads on top and the support as a morphism\.*Mathematical Structures in Computer Science*, 31\(8\):850–897, 2022\.
- \[22\]R\. Fuchs\. The calculus of language: Explicit representation of emergent linguistic structure through type\-theoretical paradigms\.*Interdisciplinary Science Reviews*, 2021\.
- \[23\]W\. Hoeffding\. Probability inequalities for sums of bounded random variables\.*Journal of the American Statistical Association*, 58\(301\):13–30, 1963\.
- \[24\]J\. Jacobs and T\. Wißmann\. Fast coalgebraic bisimilarity minimization\.*Proceedings of the ACM on Programming Languages*, 7\(POPL\):1514–1541, 2023\.
- \[25\]S\. Kadavath et al\. Language models \(mostly\) know what they know\.*arXiv:2207\.05221*, 2022\.
- \[26\]L\. Kuhn, Y\. Gal, and S\. Farquhar\. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation\.*arXiv:2302\.09664*, 2023\.
- \[27\]A\. Kumar, R\. Morabito, S\. Umbet, J\. Kabbara, and A\. Emami\. Confidence under the hood: An investigation into confidence–probability alignment in large language models\.*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics*, 2024\.
- \[28\]J\. Lambek\. The mathematics of sentence structure\.*American Mathematical Monthly*, 65\(3\):154–170, 1958\.
- \[29\]F\. Loregian\. Automata and coalgebras in categories of species\.*Mathematical Structures in Computer Science*, 35:e36, 1–48, 2025\.
- \[30\]M\. Moortgat\. Categorial grammar and formal semantics\. In*Encyclopedia of Language and Linguistics*\. Wiley, 2011\.
- \[31\]A\. Nerode\. Linear automaton transformations\.*Proceedings of the American Mathematical Society*, 9\(4\):541–544, 1958\.
- \[32\]J\. J\. M\. M\. Rutten\. Universal coalgebra: A theory of systems\.*Theoretical Computer Science*, 249\(1\):3–80, 2000\.
- \[33\]Shilpika, C\. Graziani, B\. Lusch, V\. Vishwanath, and M\. E\. Papka\. Probabilistic attribution for large language models\.*arXiv:2605\.21726*, 2026\.
- \[34\]A\. Silva, F\. Bonchi, M\. Bonsangue, and J\. Rutten\. Generalizing determinization from automata to coalgebras\.*Logical Methods in Computer Science*, 9\(1\), 2013\.
- \[35\]K\. Tian, E\. Mitchell, A\. Zhou, A\. Sharma, R\. Rafailov, H\. Yao, C\. Finn, and C\. Manning\. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine\-tuned with human feedback\.*Proceedings of EMNLP*, pages 5433–5442, 2023\.
- \[36\]F\. van Breugel and J\. Worrell\. A behavioural pseudometric for probabilistic transition systems\.*Theoretical Computer Science*, 331\(1\):115–142, 2005\.
- \[37\]P\. A\. Verburg\. Hobbes’ calculus of words\.*COLING 1969*, Preprint 39, 1969\.
- \[38\]S\. Wein, Z\. Wang, and N\. Schneider\. Measuring fine\-grained semantic equivalence with Abstract Meaning Representation\.*Proceedings of IWCS*, pages 144–154, 2023\.Similar Articles
From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models
The paper proposes ModelLog, a declarative probabilistic framework for evaluating language models by defining semantic constraints over token predictions, linking evaluation to learning through shared semantics.
HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation
This paper introduces HawkesLLM, a framework that models semantic uncertainty propagation in multi-step agentic text simulations by combining a multivariate Hawkes process for temporal influence and memory selection with a language model for text generation. Evaluation on a GDELT news-cascade case study shows improved late-stage semantic alignment under compact prompt-memory constraints.
Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models
This paper proposes Semantic Lenia, a framework that transforms LLM inference into a continuous dynamical system in logit space, demonstrating the emergence of autonomous semantic solitons and homeostatic limit cycles through nonlinear feedback.
Sequential statistical inference for Large Language Models: Representation, validity, and monitoring
This paper argues for a sequential inference framework to enhance LLM trustworthiness by modeling interactions as dependent stochastic processes, ensuring validity under repeated use, and enabling online monitoring for behavioral shifts.
Semantic DLM+: Improving Diffusion Language Models through Bias-variance Trade-off in Transition Kernel Design
This paper theoretically analyzes diffusion language models through a bias-variance lens, identifying trade-offs between masking and uniform diffusion kernels. It proposes SemDLM+, which adds a global transition and semantic-frequency penalty to overcome the semantic basin problem, achieving competitive generation quality on LM1B and OpenWebText benchmarks.