Semantic Field Theory: Historical Origin, Higher-Order Interaction, and Stabilized Semantic Inference
Summary
This paper introduces Semantic Field Theory (SFT), a computational model for lexical semantics that models meaning through semantic fields, contextual deformation, interaction terms, and energy minimization. It provides formal elements including Gaussian product closure, Möbius inversion for higher-order interactions, and stability conditions.
View Cached Full Text
Cached at: 07/24/26, 05:16 AM
# Historical Origin, Higher-Order Interaction, and Stabilized Semantic Inference
Source: [https://arxiv.org/html/2607.20451](https://arxiv.org/html/2607.20451)
Dimitris Vartziotis1,2 1NIKI – Digital Engineering, Ioannina, Greece 2TWT Science & Innovation, Stuttgart, Germany
\(May 2026\)
###### Abstract
Semantic Field Theory \(SFT\) has developed from a philosophical critique of strong anti\-formalist readings of language games into a proposed computational model class for lexical semantics, higher order composition, and stabilized interpretation\. This paper reconstructs that evolution and gives SFT a sharper mathematical core suitable for independent evaluation in computational linguistics and representation learning\. The central proposal is that a tractable level of linguistic organization can be modeled through lexical representations expressed as semantic fields, through contextual deformation of those fields, through interaction terms defined over subsets of tokens, and through stabilization governed by semantic energy dynamics\. The paper contributes five formal elements\. First, it defines a semantic field model as a tuple consisting of a semantic space, a lexical field lifting, a contextual deformation map, an interaction complex, and an interpretation functional\. Second, it proves a Gaussian product closure result showing that multiplicative field interactions have explicit centers, precisions, and compatibility factors\. Third, it generalizes the three\-word problem by using Mobius inversion on the subset lattice to isolate irreducible semantic interactions of arbitrary order\. Fourth, it introduces an order spectrum that measures how much field mass is explained at each interaction order\. Fifth, it formulates stabilized interpretation as minimization of an energy functional associated with the sentence and gives existence, descent, and stability conditions\. A small worked example shows how a three\-word summer day triple can be represented by Gaussian semantic fields, implemented in Python, and summarized by a flow diagram\. The result is not a completed theory of natural language meaning and does not replace social, pragmatic, or normative accounts of language\. It is a mathematically explicit hypothesis about one representational level: how public language use may give rise to stable geometric, interactional, and dynamical regularities that can be estimated, ablated, and compared with transformer models\.
## 1Introduction
The mathematical study of linguistic meaning has passed through several regimes\. The distributional hypothesis treated lexical meaning as a pattern of contextual occurrence\(Harris,[1954](https://arxiv.org/html/2607.20451#bib.bib12); Firth,[1957](https://arxiv.org/html/2607.20451#bib.bib10)\)\. Vector\-space semantics made that hypothesis geometrical and computational\(Turney and Pantel,[2010](https://arxiv.org/html/2607.20451#bib.bib27)\)\. Compositional distributional semantics then asked how phrase and sentence meanings could be built from lexical representations\(Mitchell and Lapata,[2008](https://arxiv.org/html/2607.20451#bib.bib20),[2010](https://arxiv.org/html/2607.20451#bib.bib21); Coecke et al\.,[2010](https://arxiv.org/html/2607.20451#bib.bib7); Grefenstette et al\.,[2011](https://arxiv.org/html/2607.20451#bib.bib11)\)\. Neural embeddings made dense semantic geometry empirically powerful\(Mikolov et al\.,[2013](https://arxiv.org/html/2607.20451#bib.bib19); Pennington et al\.,[2014](https://arxiv.org/html/2607.20451#bib.bib22)\)\. Contextual encoders and transformers shifted the unit of representation from a word type to a token embedded in an utterance\(Peters et al\.,[2018](https://arxiv.org/html/2607.20451#bib.bib23); Devlin et al\.,[2019](https://arxiv.org/html/2607.20451#bib.bib8); Vaswani et al\.,[2017](https://arxiv.org/html/2607.20451#bib.bib33)\)\. Large language models subsequently showed that broad linguistic regularities can be optimized, transferred, and used for general\-purpose language behavior at scale\(Kaplan et al\.,[2020](https://arxiv.org/html/2607.20451#bib.bib15); Brown et al\.,[2020](https://arxiv.org/html/2607.20451#bib.bib5); Bommasani et al\.,[2021](https://arxiv.org/html/2607.20451#bib.bib3)\)\.
This development motivates a question that is both scientific and philosophical\. Are mathematical structures merely convenient external representations of language, or do they reveal a stable level of organization generated by language use itself? A strong answer would be metaphysical and is unnecessary for the present paper\. A weaker and scientifically useful answer is enough: public language use appears to induce regularities that can be represented as geometry, interaction, and dynamics\. These regularities do not settle questions of reference, grounding, pragmatics, or normativity\(Bender and Koller,[2020](https://arxiv.org/html/2607.20451#bib.bib1); Bender et al\.,[2021](https://arxiv.org/html/2607.20451#bib.bib2)\)\. They nevertheless form a legitimate object of computational theory\.
Semantic Field Theory \(SFT\) is proposed as one such theory\. From its earliest philosophical stages, SFT was guided by the intuition that lexical meaning behaves more like a distributed semantic field than a discrete symbolic atom or isolated vector representation\. A lexical item should therefore not be represented only as a point, a symbol, or a single vector, but as a distributed field over a semantic space; an utterance should not be represented only as a sum of lexical items, but as an interaction field over subsets of tokens; and interpretation should not be treated only as retrieval of a static object, but as stabilization in a sentence\-conditioned energy landscape\. Earlier forms of this intuition were developed in philosophical work on Wittgenstein and language games\(Vartziotis,[2012](https://arxiv.org/html/2607.20451#bib.bib28),[2017](https://arxiv.org/html/2607.20451#bib.bib29)\)\. A later preprint formulated the explicit contrast between language\-game anti\-formalism and language as mathematical structure\(Vartziotis,[2026](https://arxiv.org/html/2607.20451#bib.bib32)\)\. The present paper takes the next step: it reconstructs SFT as an evolving formal program and supplies a mathematically consolidated model class\.
The contribution is not only historical\. The paper introduces several technical elements that make SFT more mature as a computational hypothesis\. First, it defines a*semantic field model*as a tuple of maps and function spaces\. Second, it replaces the merely metaphorical notion of field interaction by a Gaussian field algebra, including a closed\-form product rule for interaction centers and compatibility factors\. Third, it generalizes the three\-word problem through Mobius inversion on the subset lattice\. Fourth, it introduces an*order spectrum*, a measurable profile of first\-order, pairwise, third\-order, and higher semantic residuals\. Fifth, it defines a stabilized interpretation map by energy minimization and gives basic existence, descent, and perturbation stability results\. A short worked example is included to make the formal machinery executable and inspectable in the simple case of a three\-token utterance\.
The intended status of the paper is therefore precise\. SFT is not claimed to be a full semantics of natural language\. It is a formal model class for a particular level of semantic organization: distributed lexical influence, context\-dependent deformation, higher\-order interaction, and stabilized inference\. This level is compatible with use\-theoretic and social accounts of meaning, but it is not exhausted by them\. Its value should be judged by whether it yields estimable parameters, interpretable diagnostics, and discriminable empirical consequences\.
#### Organization\.
Section[2](https://arxiv.org/html/2607.20451#S2)reconstructs the historical development of SFT\. Section[3](https://arxiv.org/html/2607.20451#S3)clarifies the relation between language games, distributional structure, and formal semantics\. Section[4](https://arxiv.org/html/2607.20451#S4)defines the semantic field model\. Section[5](https://arxiv.org/html/2607.20451#S5)develops the Gaussian interaction algebra\. Section[6](https://arxiv.org/html/2607.20451#S6)gives a minimal computational example using a three\-word summer\-day triple\. Section[7](https://arxiv.org/html/2607.20451#S7)gives the residual theory of higher\-order composition\. Section[8](https://arxiv.org/html/2607.20451#S8)treats interpretation as energy minimization\. Section[10](https://arxiv.org/html/2607.20451#S10)relates SFT to transformer representations\. Section[11](https://arxiv.org/html/2607.20451#S11)proposes empirical tests\. Section[12](https://arxiv.org/html/2607.20451#S12)states limitations\.
## 2Historical evolution of SFT
SFT did not originate as a response to current large language models\. Its earliest motivation was philosophical: a resistance to the inference that because meaning is public, social, and use\-governed, the search for inner semantic structure must be illegitimate\. The development can be summarized as a sequence of increasingly explicit formulations\.
Table 1:A compact reconstruction of the historical development of SFT\. The earlier works are treated here as conceptual antecedents, not as complete technical implementations of the present model\.The first stage is a philosophical orientation\. Later Wittgenstein emphasizes language as public practice, rule\-following, and participation in language games\(Wittgenstein,[1953](https://arxiv.org/html/2607.20451#bib.bib37)\)\. This remains a necessary constraint on any account of meaning\. However, the stronger claim that this public character prohibits formal internal structure does not follow\. The early SFT intuition was that use and structure should be separated analytically: use supplies the public ground of meaning, while repeated use can also produce stable patterns that are mathematically articulable\.
Already in the earlier philosophical stages of SFT, a distinction emerged between lexical fields and linguistic fields, although not yet in formal mathematical language\. Lexical fields correspond to distributed semantic potential associated with word types\. Linguistic fields correspond to utterance\-level interaction systems produced when lexical fields are activated together\. This distinction is central because it prevents SFT from collapsing into one\-vector\-per\-word semantics\. It also prevents the theory from identifying sentence meaning with simple aggregation\.
The third stage is the LLM\-era reinterpretation\. Transformer\-based systems do not prove that language is reducible to vectors, nor do they solve the philosophical problem of meaning\. They do, however, demonstrate that large\-scale language use contains stable, learnable, high\-dimensional regularities\. This observation weakens strong anti\-structural prohibitions and motivates a more detailed formal account\. This broader reinterpretation is further developed in two books currently in press at Literareon – Utz Verlag by the author:*Wittgenstein and the End of Language Games: Mathematical Semantics, Field Theory, and the Age of Predictive Language Models*, and*The Revolution of LLMs: Philosophical and Mathematical Conjugations*\. These works extend the philosophical implications of SFT beyond the present computational formulation and examine how predictive language models reshape traditional debates concerning meaning, structure, and linguistic representation\.
The present paper is the fourth stage\. It treats SFT not as a slogan but as a family of estimable mathematical models\. The historical claim is modest: SFT began as a philosophical critique and is now reformulated as a computational hypothesis\. The scientific claim is also modest: if the hypothesis is meaningful, it should generate quantities that can be estimated and ablated, such as field overlap, residual interaction mass, energy gaps, and stabilization trajectories\.
## 3Use, structure, and levels of explanation
A recurring source of confusion is the phrase “inner structure”\. If it means private mental objects that determine meaning independently of public criteria, then SFT does not require it\. If it means mathematically describable regularities induced by public language use, then rejecting it is unnecessary and scientifically costly\.
We therefore distinguish four levels\.
1. 1\.Normative\-pragmatic level\.Meaning is learned, corrected, and stabilized in public linguistic practice\. This is the level emphasized by language\-game approaches\(Wittgenstein,[1953](https://arxiv.org/html/2607.20451#bib.bib37); Kripke,[1982](https://arxiv.org/html/2607.20451#bib.bib16)\)\.
2. 2\.Distributional level\.Repeated use leaves statistical regularities in contexts, co\-occurrences, substitutions, and entailment patterns\(Harris,[1954](https://arxiv.org/html/2607.20451#bib.bib12); Firth,[1957](https://arxiv.org/html/2607.20451#bib.bib10); Turney and Pantel,[2010](https://arxiv.org/html/2607.20451#bib.bib27)\)\.
3. 3\.Geometric\-field level\.Distributional regularities can be represented as fields, distances, overlaps, displacements, and interaction structures in a semantic space\.
4. 4\.Dynamical\-computational level\.Interpretation can be modeled as a process that evolves toward stable states under a sentence\-conditioned field\.
SFT is a theory of the third and fourth levels\. It is constrained by the first and informed by the second\. This layered view is important for scientific maturity because it avoids two reductions\. It avoids reducing meaning to social practice alone, and it avoids reducing meaning to hidden vectors alone\. The object of SFT is narrower: the structured representational regularities through which lexical and utterance\-level meanings can be approximated, composed, and dynamically stabilized\.
## 4Semantic Field Theory as a model class
Let𝐰=\(w1,…,wm\)\\mathbf\{w\}=\(w\_\{1\},\\ldots,w\_\{m\}\)be a token sequence and let\[m\]=\{1,…,m\}\[m\]=\\\{1,\\ldots,m\\\}\. LetS⊆ℝnS\\subseteq\\mathbb\{R\}^\{n\}be a semantic space with Euclidean norm∥⋅∥\\\|\\cdot\\\|\. The space is not an ontological container of meanings; it is a formal domain for semantic proximity, deformation, interaction, and dynamics\.
###### Definition 1\(Semantic field model\)\.
A semantic field model of orderppis a tuple
𝔐p=\(S,𝒱,ℓΘ,DΘ,𝒦p,ΛΘ,UΘ\),\\mathfrak\{M\}\_\{p\}=\(S,\\mathcal\{V\},\\ell\_\{\\Theta\},D\_\{\\Theta\},\\mathcal\{K\}\_\{p\},\\Lambda\_\{\\Theta\},U\_\{\\Theta\}\),whereS⊆ℝnS\\subseteq\\mathbb\{R\}^\{n\}is a semantic space,𝒱\\mathcal\{V\}is a vocabulary,ℓΘ\\ell\_\{\\Theta\}maps lexical types to base fields,DΘD\_\{\\Theta\}maps base fields to contextually deformed token fields,𝒦p\(𝐰\)\\mathcal\{K\}\_\{p\}\(\\mathbf\{w\}\)is an interaction complex containing subsets of\[m\]\[m\]of size at mostpp,ΛΘ\\Lambda\_\{\\Theta\}assigns sentence\-conditioned interaction coefficients, andUΘ\(⋅∣𝐰\)U\_\{\\Theta\}\(\\cdot\\mid\\mathbf\{w\}\)is an interpretation energy\.
This definition separates five roles that were often implicit in earlier discussions: lexical structure, contextual deformation, interaction selection, interaction weighting, and interpretation\. Each role may be instantiated by symbolic, neural, probabilistic, or hybrid mechanisms\.
### 4\.1Lexical field lifting
Let each vocabulary typet∈𝒱t\\in\\mathcal\{V\}have a base embeddingzt∈ℝdzz\_\{t\}\\in\\mathbb\{R\}^\{d\_\{z\}\}\. A lexical lifting map sendsztz\_\{t\}to field parameters
ct\\displaystyle c\_\{t\}=Wczt,\\displaystyle=W\_\{c\}z\_\{t\},\(1\)at\\displaystyle a\_\{t\}=softplus\(wa⊤zt\+ba\),\\displaystyle=\\operatorname\{softplus\}\(w\_\{a\}^\{\\top\}z\_\{t\}\+b\_\{a\}\),\(2\)Σt\\displaystyle\\Sigma\_\{t\}=diag\(softplus\(WΣzt\)\+δ𝟏\),δ\>0\.\\displaystyle=\\operatorname\{diag\}\\big\(\\operatorname\{softplus\}\(W\_\{\\Sigma\}z\_\{t\}\)\+\\delta\\mathbf\{1\}\\big\),\\qquad\\delta\>0\.\(3\)For a token occurrencewi=tiw\_\{i\}=t\_\{i\}, the base lexical field is
Li\(x\)=atiexp\[−12\(x−cti\)⊤Σti−1\(x−cti\)\],x∈S\.L\_\{i\}\(x\)=a\_\{t\_\{i\}\}\\exp\\\!\\left\[\-\\frac\{1\}\{2\}\(x\-c\_\{t\_\{i\}\}\)^\{\\top\}\\Sigma\_\{t\_\{i\}\}^\{\-1\}\(x\-c\_\{t\_\{i\}\}\)\\right\],\\qquad x\\in S\.\(4\)A Gaussian field is not required by the theory, but it gives a useful base case: centers encode dominant regions, covariances encode semantic spread and anisotropy, and amplitudes encode field strength\.
### 4\.2Contextual deformation
Type\-level fields must be distinguished from token\-level fields\. Letqϕ\(𝐰\)q\_\{\\phi\}\(\\mathbf\{w\}\)be a contextual sentence representation, possibly produced by a transformer encoder, recurrent model, syntactic encoder, or task\-specific context map\. A contextually deformed token field is
L~i\(x∣𝐰\)=a~i\(𝐰\)exp\[−12\(x−c~i\(𝐰\)\)⊤Σ~i\(𝐰\)−1\(x−c~i\(𝐰\)\)\],\\widetilde\{L\}\_\{i\}\(x\\mid\\mathbf\{w\}\)=\\tilde\{a\}\_\{i\}\(\\mathbf\{w\}\)\\exp\\\!\\left\[\-\\frac\{1\}\{2\}\(x\-\\tilde\{c\}\_\{i\}\(\\mathbf\{w\}\)\)^\{\\top\}\\widetilde\{\\Sigma\}\_\{i\}\(\\mathbf\{w\}\)^\{\-1\}\(x\-\\tilde\{c\}\_\{i\}\(\\mathbf\{w\}\)\)\\right\],\(5\)with
a~i\(𝐰\)\\displaystyle\\tilde\{a\}\_\{i\}\(\\mathbf\{w\}\)=atiγi\(𝐰\),γi\(𝐰\)\>0,\\displaystyle=a\_\{t\_\{i\}\}\\,\\gamma\_\{i\}\(\\mathbf\{w\}\),\\qquad\\gamma\_\{i\}\(\\mathbf\{w\}\)\>0,\(6\)c~i\(𝐰\)\\displaystyle\\tilde\{c\}\_\{i\}\(\\mathbf\{w\}\)=cti\+di\(𝐰\),\\displaystyle=c\_\{t\_\{i\}\}\+d\_\{i\}\(\\mathbf\{w\}\),\(7\)Σ~i\(𝐰\)\\displaystyle\\widetilde\{\\Sigma\}\_\{i\}\(\\mathbf\{w\}\)=Σti\+Ri\(𝐰\),Σ~i\(𝐰\)≻0\.\\displaystyle=\\Sigma\_\{t\_\{i\}\}\+R\_\{i\}\(\\mathbf\{w\}\),\\qquad\\widetilde\{\\Sigma\}\_\{i\}\(\\mathbf\{w\}\)\\succ 0\.\(8\)The simplest version setsdi=0d\_\{i\}=0andRi=0R\_\{i\}=0and uses only a non\-negative gateγi\\gamma\_\{i\}\. A richer version permits contextual displacement and covariance deformation\. This distinction is conceptually important: context need not erase lexical identity; it can deform lexical influence\.
A concrete neural parameterization is
γi\(𝐰\)\\displaystyle\\gamma\_\{i\}\(\\mathbf\{w\}\)=softplus\(ug⊤tanh\(Wg\[zti;qϕ\(𝐰\);ei\]\)\+bg\),\\displaystyle=\\operatorname\{softplus\}\\\!\\left\(u\_\{g\}^\{\\top\}\\tanh\(W\_\{g\}\[z\_\{t\_\{i\}\};q\_\{\\phi\}\(\\mathbf\{w\}\);e\_\{i\}\]\)\+b\_\{g\}\\right\),\(9\)di\(𝐰\)\\displaystyle d\_\{i\}\(\\mathbf\{w\}\)=Wdtanh\(Wdc\[zti;qϕ\(𝐰\);ei\]\),\\displaystyle=W\_\{d\}\\tanh\(W\_\{dc\}\[z\_\{t\_\{i\}\};q\_\{\\phi\}\(\\mathbf\{w\}\);e\_\{i\}\]\),\(10\)whereeie\_\{i\}may include positional or syntactic features\. This makes the model order\-sensitive even when interaction products are symmetric in their arguments\.
### 4\.3Interaction complexes
The full powerset of\[m\]\[m\]is usually computationally impossible\. SFT therefore separates the formal definition from the selected interaction structure\.
###### Definition 2\(Interaction complex\)\.
For a sentence𝐰\\mathbf\{w\}and orderpp, an interaction complex𝒦p\(𝐰\)\\mathcal\{K\}\_\{p\}\(\\mathbf\{w\}\)is a family of non\-empty subsetsA⊆\[m\]A\\subseteq\[m\]such that\|A\|≤p\|A\|\\leq p\. It is downward closed ifA∈𝒦p\(𝐰\)A\\in\\mathcal\{K\}\_\{p\}\(\\mathbf\{w\}\)and∅≠B⊆A\\emptyset\\neq B\\subseteq AimplyB∈𝒦p\(𝐰\)B\\in\\mathcal\{K\}\_\{p\}\(\\mathbf\{w\}\)\.
A downward\-closed complex is natural when residual interactions are interpreted hierarchically\. In practice,𝒦p\\mathcal\{K\}\_\{p\}can be selected by local windows, syntactic neighborhoods, top\-kkcontextual similarity, dependency paths, or learned sparsity\.
For eachA∈𝒦p\(𝐰\)A\\in\\mathcal\{K\}\_\{p\}\(\\mathbf\{w\}\), define
ΦA\(x∣𝐰\)=λA\(𝐰\)∏i∈AL~i\(x∣𝐰\),\\Phi\_\{A\}\(x\\mid\\mathbf\{w\}\)=\\lambda\_\{A\}\(\\mathbf\{w\}\)\\prod\_\{i\\in A\}\\widetilde\{L\}\_\{i\}\(x\\mid\\mathbf\{w\}\),\(11\)whereλA\(𝐰\)∈ℝ\\lambda\_\{A\}\(\\mathbf\{w\}\)\\in\\mathbb\{R\}is an interaction coefficient\. The linguistic field of orderppis
ℒp\(x∣𝐰\)=∑A∈𝒦p\(𝐰\)ΦA\(x∣𝐰\)\.\\mathcal\{L\}\_\{p\}\(x\\mid\\mathbf\{w\}\)=\\sum\_\{A\\in\\mathcal\{K\}\_\{p\}\(\\mathbf\{w\}\)\}\\Phi\_\{A\}\(x\\mid\\mathbf\{w\}\)\.\(12\)When\|A\|=1\|A\|=1, one may setλA=1\\lambda\_\{A\}=1or absorb singleton weights intoL~i\\widetilde\{L\}\_\{i\}\. Positive coefficients reinforce joint activation; negative coefficients suppress it\.
A sentence\-conditioned coefficient can be parameterized by
λA\(𝐰\)=v\|A\|⊤tanh\(W\|A\|\[pool\(zti,ei:i∈A\);qϕ\(𝐰\)\]\)\+b\|A\|,\\lambda\_\{A\}\(\\mathbf\{w\}\)=v\_\{\|A\|\}^\{\\top\}\\tanh\\\!\\left\(W\_\{\|A\|\}\[\\operatorname\{pool\}\(z\_\{t\_\{i\}\},e\_\{i\}:i\\in A\);q\_\{\\phi\}\(\\mathbf\{w\}\)\]\\right\)\+b\_\{\|A\|\},\(13\)wherepool\\operatorname\{pool\}is symmetric if order information is carried elsewhere, or order\-sensitive if the subset is represented as an ordered tuple\.
## 5Gaussian field algebra
A main advantage of Gaussian lexical fields is that multiplicative interactions remain analytically tractable\. This gives SFT a concrete geometry of composition rather than a purely verbal field metaphor\.
Fori∈Ai\\in A, write
L~i\(x∣𝐰\)=a~iexp\[−12\(x−c~i\)⊤Σ~i−1\(x−c~i\)\],\\widetilde\{L\}\_\{i\}\(x\\mid\\mathbf\{w\}\)=\\tilde\{a\}\_\{i\}\\exp\\\!\\left\[\-\\frac\{1\}\{2\}\(x\-\\tilde\{c\}\_\{i\}\)^\{\\top\}\\widetilde\{\\Sigma\}\_\{i\}^\{\-1\}\(x\-\\tilde\{c\}\_\{i\}\)\\right\],where the dependence on𝐰\\mathbf\{w\}is suppressed for readability\. LetPi=Σ~i−1P\_\{i\}=\\widetilde\{\\Sigma\}\_\{i\}^\{\-1\}\.
###### Proposition 1\(Gaussian product closure\)\.
For any non\-emptyA⊆\[m\]A\\subseteq\[m\], define
PA\\displaystyle P\_\{A\}=∑i∈APi,\\displaystyle=\\sum\_\{i\\in A\}P\_\{i\},\(14\)ΣA\\displaystyle\\Sigma\_\{A\}=PA−1,\\displaystyle=P\_\{A\}^\{\-1\},\(15\)μA\\displaystyle\\mu\_\{A\}=ΣA∑i∈APic~i,\\displaystyle=\\Sigma\_\{A\}\\sum\_\{i\\in A\}P\_\{i\}\\tilde\{c\}\_\{i\},\(16\)κA\\displaystyle\\kappa\_\{A\}=exp\[−12\(∑i∈Ac~i⊤Pic~i−μA⊤PAμA\)\]\.\\displaystyle=\\exp\\\!\\left\[\-\\frac\{1\}\{2\}\\left\(\\sum\_\{i\\in A\}\\tilde\{c\}\_\{i\}^\{\\top\}P\_\{i\}\\tilde\{c\}\_\{i\}\-\\mu\_\{A\}^\{\\top\}P\_\{A\}\\mu\_\{A\}\\right\)\\right\]\.\(17\)Then
∏i∈AL~i\(x∣𝐰\)=\(∏i∈Aa~i\)κAexp\[−12\(x−μA\)⊤PA\(x−μA\)\]\.\\prod\_\{i\\in A\}\\widetilde\{L\}\_\{i\}\(x\\mid\\mathbf\{w\}\)=\\left\(\\prod\_\{i\\in A\}\\tilde\{a\}\_\{i\}\\right\)\\kappa\_\{A\}\\exp\\\!\\left\[\-\\frac\{1\}\{2\}\(x\-\\mu\_\{A\}\)^\{\\top\}P\_\{A\}\(x\-\\mu\_\{A\}\)\\right\]\.\(18\)
###### Proof\.
Expand the exponent:
∑i∈A\(x−c~i\)⊤Pi\(x−c~i\)=x⊤PAx−2x⊤∑i∈APic~i\+∑i∈Ac~i⊤Pic~i\.\\sum\_\{i\\in A\}\(x\-\\tilde\{c\}\_\{i\}\)^\{\\top\}P\_\{i\}\(x\-\\tilde\{c\}\_\{i\}\)=x^\{\\top\}P\_\{A\}x\-2x^\{\\top\}\\sum\_\{i\\in A\}P\_\{i\}\\tilde\{c\}\_\{i\}\+\\sum\_\{i\\in A\}\\tilde\{c\}\_\{i\}^\{\\top\}P\_\{i\}\\tilde\{c\}\_\{i\}\.SinceμA=PA−1∑i∈APic~i\\mu\_\{A\}=P\_\{A\}^\{\-1\}\\sum\_\{i\\in A\}P\_\{i\}\\tilde\{c\}\_\{i\}, completing the square gives
x⊤PAx−2x⊤PAμA=\(x−μA\)⊤PA\(x−μA\)−μA⊤PAμA\.x^\{\\top\}P\_\{A\}x\-2x^\{\\top\}P\_\{A\}\\mu\_\{A\}=\(x\-\\mu\_\{A\}\)^\{\\top\}P\_\{A\}\(x\-\\mu\_\{A\}\)\-\\mu\_\{A\}^\{\\top\}P\_\{A\}\\mu\_\{A\}\.Substitution yields \([18](https://arxiv.org/html/2607.20451#S5.E18)\)\. ∎
The termμA\\mu\_\{A\}is the*interaction focus*\. It is a precision\-weighted semantic location jointly supported by the fields inAA\. The matrixPAP\_\{A\}is the*interaction precision*: adding fields sharpens the product\. The scalarκA\\kappa\_\{A\}is a*compatibility factor*\. It decreases when centers are mutually distant relative to their covariances and increases when the fields overlap\. Thus, in the Gaussian case, SFT yields explicit quantities for lexical compatibility and compositional focus\.
### 5\.1Interaction tension
The exponent inκA\\kappa\_\{A\}defines a natural measure of semantic tension:
TA=∑i∈Ac~i⊤Pic~i−μA⊤PAμA≥0\.T\_\{A\}=\\sum\_\{i\\in A\}\\tilde\{c\}\_\{i\}^\{\\top\}P\_\{i\}\\tilde\{c\}\_\{i\}\-\\mu\_\{A\}^\{\\top\}P\_\{A\}\\mu\_\{A\}\\geq 0\.\(19\)For\|A\|=2\|A\|=2,TAT\_\{A\}reduces to a Mahalanobis\-type discrepancy between the two centers under the combined covariance\. For largerAA, it measures how difficult it is for all fields inAAto share a common focus\. A high\-tension triple with a non\-zero positive coefficientλA\\lambda\_\{A\}is a natural formal candidate for metaphor, coercion, or non\-literal construal: the fields are not simply close, yet the sentence activates a joint focus\.
## 6Minimal computational example: a summer\-day triple
This section gives a deliberately small, transparent instantiation of the preceding Gaussian algebra\. The example is not intended as an empirical validation of SFT\. It is a didactic bridge from the formal definitions to executable code\. Consider the three\-token summer\-day triple
𝐰=\(w1,w2,w3\)=\(helios,thalassa,zesti\),\\mathbf\{w\}=\(w\_\{1\},w\_\{2\},w\_\{3\}\)=\(\\text\{helios\},\\text\{thalassa\},\\text\{zesti\}\),where the transliterated Greek words correspond to*sun*,*sea*, and*heat*\. Let the semantic space beS=ℝ3S=\\mathbb\{R\}^\{3\}, with coordinates interpreted as summer salience, natural/marine salience, and thermal intensity\. Assign toy centers
c1\\displaystyle c\_\{1\}=\(0\.95,0\.70,0\.90\)⊤,\\displaystyle=\(0\.95,0\.70,0\.90\)^\{\\top\},\(20\)c2\\displaystyle c\_\{2\}=\(0\.85,0\.95,0\.45\)⊤,\\displaystyle=\(0\.85,0\.95,0\.45\)^\{\\top\},\(21\)c3\\displaystyle c\_\{3\}=\(0\.90,0\.45,0\.98\)⊤,\\displaystyle=\(0\.90,0\.45,0\.98\)^\{\\top\},\(22\)and use the simplest isotropic fields
Pi=I3,ai=1,i=1,2,3\.P\_\{i\}=I\_\{3\},\\qquad a\_\{i\}=1,\\qquad i=1,2,3\.\(23\)Thus each lexical field has the form
Li\(x\)=exp\[−12‖x−ci‖2\]\.L\_\{i\}\(x\)=\\exp\\\!\\left\[\-\\frac\{1\}\{2\}\\\|x\-c\_\{i\}\\\|^\{2\}\\right\]\.\(24\)ForA=\{1,2,3\}A=\\\{1,2,3\\\}, the SFT interaction quantities become
PA\\displaystyle P\_\{A\}=3I3,\\displaystyle=3I\_\{3\},\(25\)ΣA\\displaystyle\\Sigma\_\{A\}=13I3,\\displaystyle=\\frac\{1\}\{3\}I\_\{3\},\(26\)μA\\displaystyle\\mu\_\{A\}=13\(c1\+c2\+c3\)=\(0\.900,0\.700,0\.777\)⊤,\\displaystyle=\\frac\{1\}\{3\}\(c\_\{1\}\+c\_\{2\}\+c\_\{3\}\)=\(0\.900,0\.700,0\.777\)^\{\\top\},\(27\)TA\\displaystyle T\_\{A\}=∑i=13‖ci‖2−3‖μA‖2≈0\.293,\\displaystyle=\\sum\_\{i=1\}^\{3\}\\\|c\_\{i\}\\\|^\{2\}\-3\\\|\\mu\_\{A\}\\\|^\{2\}\\approx 0\.293,\(28\)κA\\displaystyle\\kappa\_\{A\}=exp\(−TA/2\)≈0\.864\.\\displaystyle=\\exp\(\-T\_\{A\}/2\)\\approx 0\.864\.\(29\)In this toy setting, the compatibility is high because the three fields jointly support a coherent focus: a hot summer day near the sea\. The second coordinate is pulled upward by*thalassa*, the third coordinate is pulled upward by*helios*and*zesti*, and the first coordinate remains high across all three terms\. This illustrates the central SFT idea that the utterance\-level focus is not one lexical point, but a stabilized region produced by field interaction\.
### 6\.1Executable prototype
Listing[1](https://arxiv.org/html/2607.20451#LST1)implements the same calculation\. The centers are hand\-coded only to make the algebra visible; in an empirical implementation they would be learned from data or induced from an encoder\.
Listing 1:Minimal SFT computation for a summer\-day triple\.1importnumpyasnp
2fromitertoolsimportcombinations
3
4
5
6words=\["helios","thalassa","zesti"\]
7
8
9
10centers=\{
11"helios":np\.array\(\[0\.95,0\.70,0\.90\]\),
12"thalassa":np\.array\(\[0\.85,0\.95,0\.45\]\),
13"zesti":np\.array\(\[0\.90,0\.45,0\.98\]\),
14\}
15
16fields=\[\]
17forwordinwords:
18fields\.append\(\{
19"word":word,
20"center":centers\[word\],
21"precision":np\.eye\(3\),
22"amplitude":1\.0,
23\}\)
24
25
26defsft\_interaction\(selected\_fields\):
27"""GaussianproductinteractionaccordingtoSFT\."""
28P\_A=sum\(f\["precision"\]forfinselected\_fields\)
29Sigma\_A=np\.linalg\.inv\(P\_A\)
30
31weighted\_sum=sum\(
32f\["precision"\]@f\["center"\]forfinselected\_fields
33\)
34mu\_A=Sigma\_A@weighted\_sum
35
36tension=sum\(
37f\["center"\]\.T@f\["precision"\]@f\["center"\]
38forfinselected\_fields
39\)
40tension\-=mu\_A\.T@P\_A@mu\_A
41
42compatibility=np\.exp\(\-0\.5\*tension\)
43returnmu\_A,float\(tension\),float\(compatibility\)
44
45
46print\("Pairwiseinteractions"\)
47forpairincombinations\(fields,2\):
48mu,tension,compatibility=sft\_interaction\(pair\)
49print\(\[f\["word"\]forfinpair\]\)
50print\("focusmu=",np\.round\(mu,3\)\)
51print\("tensionT=",round\(tension,4\)\)
52print\("compatibility=",round\(compatibility,4\)\)
53
54print\("\\nTripleinteraction"\)
55mu,tension,compatibility=sft\_interaction\(fields\)
56print\(words\)
57print\("focusmu\_123=",np\.round\(mu,3\)\)
58print\("tensionT\_123=",round\(tension,4\)\)
59print\("compatibilityk\_123=",round\(compatibility,4\)\)
60
61ifcompatibility\>0\.8:
62print\("Interpretation:coherentsummer\-daysemanticfield\."\)
63elifcompatibility\>0\.5:
64print\("Interpretation:relatedfieldswithmoderatetension\."\)
65else:
66print\("Interpretation:weak,unstable,ormetaphoricalinteraction\."\)
### 6\.2Flussdiagramm
Figure[1](https://arxiv.org/html/2607.20451#S6.F1)summarizes the computational flow\. The diagram is intentionally close to the mathematical pipeline: lexical items are lifted to field centers, Gaussian fields are constructed, subset interactions are evaluated, and the triple\-level focus and diagnostics are interpreted\.
Input triple:*helios*,*thalassa*,*zesti*\(sun, sea, heat\)Assign or learn semantic centersci∈ℝ3c\_\{i\}\\in\\mathbb\{R\}^\{3\}Construct Gaussian lexical fieldsLi\(x\)=exp\[−12‖x−ci‖2\]L\_\{i\}\(x\)=\\exp\[\-\\frac\{1\}\{2\}\\\|x\-c\_\{i\}\\\|^\{2\}\]Compute pairwise interactions for\{1,2\}\\\{1,2\\\},\{1,3\}\\\{1,3\\\},\{2,3\}\\\{2,3\\\}Compute triple interaction forA=\{1,2,3\}A=\\\{1,2,3\\\}ReturnμA\\mu\_\{A\},TAT\_\{A\}, andκA\\kappa\_\{A\}Interpretation: coherent summer\-day semantic fieldFigure 1:Flussdiagramm for the minimal SFT summer\-day example\.
## 7Residual interaction and the generalized three\-word problem
The three\-word problem states that some semantic effects are not reconstructible from singleton and pairwise effects\. The mature form of this idea is not limited to triples\. It is a general residual theory over the subset lattice\.
For any non\-emptyA⊆\[m\]A\\subseteq\[m\], letℒA\(x∣𝐰\)\\mathcal\{L\}\_\{A\}\(x\\mid\\mathbf\{w\}\)denote the linguistic field induced by the subsequence or subconfiguration indexed byAA\. Defineℒ∅≡0\\mathcal\{L\}\_\{\\emptyset\}\\equiv 0\. The residual interaction field is
ΔA\(x∣𝐰\)=∑B⊆A\(−1\)\|A\|−\|B\|ℒB\(x∣𝐰\)\.\\Delta\_\{A\}\(x\\mid\\mathbf\{w\}\)=\\sum\_\{B\\subseteq A\}\(\-1\)^\{\|A\|\-\|B\|\}\\mathcal\{L\}\_\{B\}\(x\\mid\\mathbf\{w\}\)\.\(30\)This is Mobius inversion on the Boolean lattice\(Rota,[1964](https://arxiv.org/html/2607.20451#bib.bib25)\)\.
###### Proposition 2\(Residual identification\)\.
Assume the hierarchical expansion
ℒA\(x∣𝐰\)=∑∅≠B⊆AΦB\(x∣𝐰\)\.\\mathcal\{L\}\_\{A\}\(x\\mid\\mathbf\{w\}\)=\\sum\_\{\\emptyset\\neq B\\subseteq A\}\\Phi\_\{B\}\(x\\mid\\mathbf\{w\}\)\.\(31\)Then for every non\-emptyA⊆\[m\]A\\subseteq\[m\],
ΔA\(x∣𝐰\)=ΦA\(x∣𝐰\)\.\\Delta\_\{A\}\(x\\mid\\mathbf\{w\}\)=\\Phi\_\{A\}\(x\\mid\\mathbf\{w\}\)\.\(32\)Conversely,
ℒA\(x∣𝐰\)=∑∅≠B⊆AΔB\(x∣𝐰\)\.\\mathcal\{L\}\_\{A\}\(x\\mid\\mathbf\{w\}\)=\\sum\_\{\\emptyset\\neq B\\subseteq A\}\\Delta\_\{B\}\(x\\mid\\mathbf\{w\}\)\.\(33\)
###### Proof\.
Substituting \([31](https://arxiv.org/html/2607.20451#S7.E31)\) into \([30](https://arxiv.org/html/2607.20451#S7.E30)\) gives
ΔA=∑B⊆A\(−1\)\|A\|−\|B\|∑∅≠C⊆BΦC=∑∅≠C⊆AΦC∑B:C⊆B⊆A\(−1\)\|A\|−\|B\|\.\\Delta\_\{A\}=\\sum\_\{B\\subseteq A\}\(\-1\)^\{\|A\|\-\|B\|\}\\sum\_\{\\emptyset\\neq C\\subseteq B\}\\Phi\_\{C\}=\\sum\_\{\\emptyset\\neq C\\subseteq A\}\\Phi\_\{C\}\\sum\_\{B:C\\subseteq B\\subseteq A\}\(\-1\)^\{\|A\|\-\|B\|\}\.ForC=AC=Athe inner sum equals11\. ForC⊊AC\\subsetneq A, writeB=C∪DB=C\\cup DwithD⊆A∖CD\\subseteq A\\setminus C\. The inner sum becomes
∑D⊆A∖C\(−1\)\|A\|−\|C\|−\|D\|=\(1−1\)\|A\|−\|C\|=0\.\\sum\_\{D\\subseteq A\\setminus C\}\(\-1\)^\{\|A\|\-\|C\|\-\|D\|\}=\(1\-1\)^\{\|A\|\-\|C\|\}=0\.Thus onlyΦA\\Phi\_\{A\}survives\. The inverse formula is the standard inverse of Mobius inversion on the Boolean lattice\. ∎
ForA=\{i,j,k\}A=\\\{i,j,k\\\}this gives
Δijk\\displaystyle\\Delta\_\{ijk\}=ℒijk−ℒij−ℒik−ℒjk\+ℒi\+ℒj\+ℒk,\\displaystyle=\\mathcal\{L\}\_\{ijk\}\-\\mathcal\{L\}\_\{ij\}\-\\mathcal\{L\}\_\{ik\}\-\\mathcal\{L\}\_\{jk\}\+\\mathcal\{L\}\_\{i\}\+\\mathcal\{L\}\_\{j\}\+\\mathcal\{L\}\_\{k\},\(34\)with arguments suppressed\. The triple residual is the first residual that cannot be reduced to isolated words or pairwise compatibility\. This is the precise mathematical form of the three\-word problem\.
### 7\.1Order spectra
Residuals become empirically useful when equipped with norms\. Let‖f‖ℋ\\\|f\\\|\_\{\\mathcal\{H\}\}be a chosen function norm, such as anL2L^\{2\}norm overSS, an empirical norm over sampled semantic states, or a reproducing\-kernel norm\. Define the order\-kkresidual mass
Rk\(𝐰\)=\(∑A∈𝒦p\(𝐰\)\|A\|=k∥ΔA\(⋅∣𝐰\)∥ℋ2\)1/2\.R\_\{k\}\(\\mathbf\{w\}\)=\\left\(\\sum\_\{\\begin\{subarray\}\{c\}A\\in\\mathcal\{K\}\_\{p\}\(\\mathbf\{w\}\)\\\\ \|A\|=k\\end\{subarray\}\}\\\|\\Delta\_\{A\}\(\\cdot\\mid\\mathbf\{w\}\)\\\|\_\{\\mathcal\{H\}\}^\{2\}\\right\)^\{1/2\}\.\(35\)The vector
Specp\(𝐰\)=\(R1\(𝐰\),…,Rp\(𝐰\)\)\\operatorname\{Spec\}\_\{p\}\(\\mathbf\{w\}\)=\(R\_\{1\}\(\\mathbf\{w\}\),\\ldots,R\_\{p\}\(\\mathbf\{w\}\)\)\(36\)is the*semantic order spectrum*of the sentence under the model\.
The order spectrum is a new diagnostic object\. Additive semantic models predict concentration atR1R\_\{1\}\. Pairwise compositional models predict that most residual mass lies inR1R\_\{1\}andR2R\_\{2\}\. SFT predicts that idioms, metaphors, coercions, and event construals should produce systematically higherR3R\_\{3\}or higher\-order mass than literal compositional controls, after controlling for sentence length and lexical frequency\.
## 8Interpretation as stabilized energy minimization
Given a linguistic fieldℒp\(⋅∣𝐰\)\\mathcal\{L\}\_\{p\}\(\\cdot\\mid\\mathbf\{w\}\), SFT defines interpretation as stabilization in an energy landscape\. Let
U\(x∣𝐰\)=β2‖x‖2−ℒp\(x∣𝐰\),β\>0\.U\(x\\mid\\mathbf\{w\}\)=\\frac\{\\beta\}\{2\}\\\|x\\\|^\{2\}\-\\mathcal\{L\}\_\{p\}\(x\\mid\\mathbf\{w\}\),\\qquad\\beta\>0\.\(37\)A stabilized interpretation is a local or global minimizer ofUU:
x∗\(𝐰\)∈argminx∈SU\(x∣𝐰\)\.x^\{\*\}\(\\mathbf\{w\}\)\\in\\operatorname\*\{arg\\,min\}\_\{x\\in S\}U\(x\\mid\\mathbf\{w\}\)\.\(38\)The quadratic term prevents unbounded drift and can be replaced by a more structured prior whenSShas known geometry\.
###### Proposition 3\(Existence\)\.
AssumeS=ℝnS=\\mathbb\{R\}^\{n\},p<∞p<\\infty,𝒦p\(𝐰\)\\mathcal\{K\}\_\{p\}\(\\mathbf\{w\}\)is finite, and the deformed lexical fields are bounded and continuous\. ThenU\(⋅∣𝐰\)U\(\\cdot\\mid\\mathbf\{w\}\)is coercive and attains a global minimum\.
###### Proof\.
Finite products and finite sums of bounded continuous functions are bounded and continuous, soℒp\(⋅∣𝐰\)\\mathcal\{L\}\_\{p\}\(\\cdot\\mid\\mathbf\{w\}\)is bounded and continuous\. The termβ2‖x‖2\\frac\{\\beta\}\{2\}\\\|x\\\|^\{2\}diverges to\+∞\+\\inftyas‖x‖→∞\\\|x\\\|\\to\\infty\. HenceUUis continuous and coercive onℝn\\mathbb\{R\}^\{n\}, and therefore attains a global minimum\. ∎
The gradient flow is
dxdt=−∇U\(x∣𝐰\)=∇ℒp\(x∣𝐰\)−βx\.\\frac\{dx\}\{dt\}=\-\\nabla U\(x\\mid\\mathbf\{w\}\)=\\nabla\\mathcal\{L\}\_\{p\}\(x\\mid\\mathbf\{w\}\)\-\\beta x\.\(39\)A discrete inference scheme is
x\(t\+1\)=x\(t\)−η∇U\(x\(t\)∣𝐰\)=x\(t\)\+η\(∇ℒp\(x\(t\)∣𝐰\)−βx\(t\)\)\.x^\{\(t\+1\)\}=x^\{\(t\)\}\-\\eta\\nabla U\(x^\{\(t\)\}\\mid\\mathbf\{w\}\)=x^\{\(t\)\}\+\\eta\\big\(\\nabla\\mathcal\{L\}\_\{p\}\(x^\{\(t\)\}\\mid\\mathbf\{w\}\)\-\\beta x^\{\(t\)\}\\big\)\.\(40\)
###### Proposition 4\(Descent\)\.
If∇U\(⋅∣𝐰\)\\nabla U\(\\cdot\\mid\\mathbf\{w\}\)isLUL\_\{U\}\-Lipschitz and0<η<2/LU0<\\eta<2/L\_\{U\}, then the update \([40](https://arxiv.org/html/2607.20451#S8.E40)\) satisfies
U\(x\(t\+1\)∣𝐰\)≤U\(x\(t\)∣𝐰\)−η\(1−LUη2\)∥∇U\(x\(t\)∣𝐰\)∥2\.U\(x^\{\(t\+1\)\}\\mid\\mathbf\{w\}\)\\leq U\(x^\{\(t\)\}\\mid\\mathbf\{w\}\)\-\\eta\\left\(1\-\\frac\{L\_\{U\}\\eta\}\{2\}\\right\)\\\|\\nabla U\(x^\{\(t\)\}\\mid\\mathbf\{w\}\)\\\|^\{2\}\.\(41\)
###### Proof\.
This is the standard smooth\-objective descent inequality applied to gradient descent onUU\. ∎
### 8\.1Uniqueness and multistability
SFT should not always force a unique interpretation\. Ambiguity and polysemy may correspond to multiple local minima\. Still, a uniqueness condition is useful\.
###### Proposition 5\(Strong\-convexity sufficient condition\)\.
Assumeℒp\\mathcal\{L\}\_\{p\}is twice differentiable and there existsM≥0M\\geq 0such that
∇2ℒp\(x∣𝐰\)⪯MIfor allx∈ℝn\.\\nabla^\{2\}\\mathcal\{L\}\_\{p\}\(x\\mid\\mathbf\{w\}\)\\preceq MI\\quad\\text\{for all \}x\\in\\mathbb\{R\}^\{n\}\.Ifβ\>M\\beta\>M, thenU\(⋅∣𝐰\)U\(\\cdot\\mid\\mathbf\{w\}\)is\(β−M\)\(\\beta\-M\)\-strongly convex and has a unique global minimizer\.
###### Proof\.
Since∇2U\(x∣𝐰\)=βI−∇2ℒp\(x∣𝐰\)\\nabla^\{2\}U\(x\\mid\\mathbf\{w\}\)=\\beta I\-\\nabla^\{2\}\\mathcal\{L\}\_\{p\}\(x\\mid\\mathbf\{w\}\), the assumed bound implies
∇2U\(x∣𝐰\)⪰\(β−M\)I\.\\nabla^\{2\}U\(x\\mid\\mathbf\{w\}\)\\succeq\(\\beta\-M\)I\.ThusUUis strongly convex\. A strongly convex coercive function has a unique global minimizer\. ∎
Whenβ≤M\\beta\\leq M, several minima may exist\. Rather than being a defect, this can model multiple stabilized readings\. The energy gap between minima and the volume of their basins become semantic diagnostics\.
### 8\.2Stability under perturbation
LetUUandU^\\widehat\{U\}be two sentence\-conditioned energies, for example produced by two nearby parameter settings or two paraphrastic sentences\. Suppose both areμ\\mu\-strongly convex with minimizersx∗x^\{\*\}andx^∗\\widehat\{x\}^\{\*\}\.
###### Proposition 6\(Perturbation stability\)\.
If
supx‖∇U\(x\)−∇U^\(x\)‖≤ε\\sup\_\{x\}\\\|\\nabla U\(x\)\-\\nabla\\widehat\{U\}\(x\)\\\|\\leq\\varepsilonandU^\\widehat\{U\}isμ\\mu\-strongly convex, then
‖x∗−x^∗‖≤εμ\.\\\|x^\{\*\}\-\\widehat\{x\}^\{\*\}\\\|\\leq\\frac\{\\varepsilon\}\{\\mu\}\.\(42\)
###### Proof\.
Because∇U\(x∗\)=0\\nabla U\(x^\{\*\}\)=0and∇U^\(x^∗\)=0\\nabla\\widehat\{U\}\(\\widehat\{x\}^\{\*\}\)=0,
‖∇U^\(x∗\)−∇U^\(x^∗\)‖=‖∇U^\(x∗\)‖=‖∇U^\(x∗\)−∇U\(x∗\)‖≤ε\.\\\|\\nabla\\widehat\{U\}\(x^\{\*\}\)\-\\nabla\\widehat\{U\}\(\\widehat\{x\}^\{\*\}\)\\\|=\\\|\\nabla\\widehat\{U\}\(x^\{\*\}\)\\\|=\\\|\\nabla\\widehat\{U\}\(x^\{\*\}\)\-\\nabla U\(x^\{\*\}\)\\\|\\leq\\varepsilon\.For aμ\\mu\-strongly convex differentiable function, its gradient isμ\\mu\-strongly monotone, implying
‖∇U^\(x∗\)−∇U^\(x^∗\)‖≥μ‖x∗−x^∗‖\.\\\|\\nabla\\widehat\{U\}\(x^\{\*\}\)\-\\nabla\\widehat\{U\}\(\\widehat\{x\}^\{\*\}\)\\\|\\geq\\mu\\\|x^\{\*\}\-\\widehat\{x\}^\{\*\}\\\|\.Combining the inequalities gives the result\. ∎
This proposition gives a principled sense in which stabilized interpretations vary continuously with small field perturbations when the energy basin is well conditioned\.
### 8\.3Probabilistic interpretation
For temperatureτ\>0\\tau\>0, define
pτ\(x∣𝐰\)=Z\(𝐰\)−1exp\(−U\(x∣𝐰\)τ\),Z\(𝐰\)=∫ℝnexp\(−U\(x∣𝐰\)τ\)𝑑x\.p\_\{\\tau\}\(x\\mid\\mathbf\{w\}\)=Z\(\\mathbf\{w\}\)^\{\-1\}\\exp\\\!\\left\(\-\\frac\{U\(x\\mid\\mathbf\{w\}\)\}\{\\tau\}\\right\),\\qquad Z\(\\mathbf\{w\}\)=\\int\_\{\\mathbb\{R\}^\{n\}\}\\exp\\\!\\left\(\-\\frac\{U\(x\\mid\\mathbf\{w\}\)\}\{\\tau\}\\right\)dx\.\(43\)BecauseUUis coercive under the conditions above,Z\(𝐰\)Z\(\\mathbf\{w\}\)is finite for Gaussian\-field instantiations\. The zero\-temperature limit emphasizes global minimizers, while finite temperature represents uncertainty, ambiguity, and graded interpretive alternatives\. This connects SFT to energy\-based modeling\(LeCun et al\.,[2006](https://arxiv.org/html/2607.20451#bib.bib17)\)without identifying linguistic meaning with the energy model itself\.
## 9Learning and computational instantiation
LetΘ\\Thetadenote lexical, contextual deformation, interaction, and energy parameters\. Given supervised data
𝒟=\{\(𝐰\(r\),y\(r\)\)\}r=1N,\\mathcal\{D\}=\\\{\(\\mathbf\{w\}^\{\(r\)\},y^\{\(r\)\}\)\\\}\_\{r=1\}^\{N\},definex∗\(𝐰\)x^\{\*\}\(\\mathbf\{w\}\)byTTunrolled steps of \([40](https://arxiv.org/html/2607.20451#S8.E40)\) or by an implicit optimization layer\. A prediction headFψF\_\{\\psi\}consumesx∗\(𝐰\)x^\{\*\}\(\\mathbf\{w\}\), the order spectrum, or energy\-derived features\. A general training objective is
𝒥\(Θ,ψ\)=∑r=1Nℓ\(Fψ\(χ\(𝐰\(r\)\)\),y\(r\)\)\+αΩΣ\(Θ\)\+ρΩint\(Θ\)\+ζΩstab\(Θ\),\\mathcal\{J\}\(\\Theta,\\psi\)=\\sum\_\{r=1\}^\{N\}\\ell\\big\(F\_\{\\psi\}\(\\chi\(\\mathbf\{w\}^\{\(r\)\}\)\),y^\{\(r\)\}\\big\)\+\\alpha\\Omega\_\{\\Sigma\}\(\\Theta\)\+\\rho\\Omega\_\{\\mathrm\{int\}\}\(\\Theta\)\+\\zeta\\Omega\_\{\\mathrm\{stab\}\}\(\\Theta\),\(44\)where
χ\(𝐰\)=\[x∗\(𝐰\);Specp\(𝐰\);U\(x∗\(𝐰\)∣𝐰\)\]\.\\chi\(\\mathbf\{w\}\)=\\big\[x^\{\*\}\(\\mathbf\{w\}\);\\operatorname\{Spec\}\_\{p\}\(\\mathbf\{w\}\);U\(x^\{\*\}\(\\mathbf\{w\}\)\\mid\\mathbf\{w\}\)\\big\]\.The regularizerΩΣ\\Omega\_\{\\Sigma\}controls covariance degeneracy,Ωint\\Omega\_\{\\mathrm\{int\}\}encourages sparse higher\-order interactions, andΩstab\\Omega\_\{\\mathrm\{stab\}\}penalizes unstable inference trajectories or excessively flat minima\.
The computational bottleneck is the interaction complex\. A complete order\-ppexpansion hasO\(mp\)O\(m^\{p\}\)terms\. A scientifically useful implementation should therefore report not only accuracy but also the selection rule for𝒦p\\mathcal\{K\}\_\{p\}, the retained interaction density, and the contribution of each order spectrum component\.
## 10Relation to transformer representations
Transformers are not SFT, and SFT is not an alternative derivation of the transformer architecture\. The relation is one of effective description\. A self\-attention layer computes
αij\(ℓ\)\\displaystyle\\alpha\_\{ij\}^\{\(\\ell\)\}=softmaxj\(\(WQ\(ℓ\)hi\(ℓ\)\)⊤\(WK\(ℓ\)hj\(ℓ\)\)dk\),\\displaystyle=\\operatorname\{softmax\}\_\{j\}\\left\(\\frac\{\(W\_\{Q\}^\{\(\\ell\)\}h\_\{i\}^\{\(\\ell\)\}\)^\{\\top\}\(W\_\{K\}^\{\(\\ell\)\}h\_\{j\}^\{\(\\ell\)\}\)\}\{\\sqrt\{d\_\{k\}\}\}\\right\),\(45\)hi\(ℓ\+1\)\\displaystyle h\_\{i\}^\{\(\\ell\+1\)\}=f\(hi\(ℓ\),∑jαij\(ℓ\)WV\(ℓ\)hj\(ℓ\)\)\.\\displaystyle=f\\left\(h\_\{i\}^\{\(\\ell\)\},\\sum\_\{j\}\\alpha\_\{ij\}^\{\(\\ell\)\}W\_\{V\}^\{\(\\ell\)\}h\_\{j\}^\{\(\\ell\)\}\\right\)\.\(46\)This mechanism performs contextual reweighting and nonlinear recombination over token states\(Vaswani et al\.,[2017](https://arxiv.org/html/2607.20451#bib.bib33)\)\. Probing studies suggest that transformer representations encode syntactic and semantic regularities, though attention weights should not be naively equated with explanations\(Clark et al\.,[2019](https://arxiv.org/html/2607.20451#bib.bib6); Hewitt and Manning,[2019](https://arxiv.org/html/2607.20451#bib.bib13); Tenney et al\.,[2019](https://arxiv.org/html/2607.20451#bib.bib26); Jain and Wallace,[2019](https://arxiv.org/html/2607.20451#bib.bib14); Rogers et al\.,[2020](https://arxiv.org/html/2607.20451#bib.bib24)\)\.
SFT can be placed on top of or alongside such representations in at least three ways\.
#### SFT as a semantic probe\.
A pretrained encoder suppliesqϕ\(𝐰\)q\_\{\\phi\}\(\\mathbf\{w\}\)and token features\. SFT parameters are fitted while the encoder is frozen\. The question is whether explicit field overlap, residual mass, and stabilization features explain behavior beyond standard embedding similarities\.
#### SFT as an interpretable head\.
A downstream model usesx∗\(𝐰\)x^\{\*\}\(\\mathbf\{w\}\)andSpecp\(𝐰\)\\operatorname\{Spec\}\_\{p\}\(\\mathbf\{w\}\)as interpretable features\. The model can report which interaction subsets contributed to a decision and whether the decision depends on higher\-order residuals\.
#### SFT as a regularizer\.
During fine\-tuning, one can encourage stable fields, sparse interaction complexes, or controlled order spectra\. This does not force a transformer to become an SFT model; it imposes a semantic bias on a flexible representation learner\.
The main conceptual difference is that transformers produce contextual states directly, whereas SFT decomposes semantic organization into type\-level fields, contextual deformation, explicit subset interactions, and energy stabilization\. This decomposition is the source of SFT’s testable claims\.
## 11Evaluation strategy and empirical predictions
A mature follow\-up paper must make clear what could count against the theory\. SFT becomes scientifically meaningful when it predicts measurable differences between model orders and semantic constructions\.
### 11\.1Interaction\-order ablation
Train or fit variants withp=1p=1,p=2p=2, andp=3p=3while holding the encoder and parameter budget as fixed as possible\. Evaluate on phrase similarity, sentence similarity, natural\-language inference, and controlled idiom/metaphor datasets\. Benchmarks such as SICK, SNLI, MultiNLI, GLUE, and SuperGLUE provide starting points for such comparisons\(Marelli et al\.,[2014](https://arxiv.org/html/2607.20451#bib.bib18); Bowman et al\.,[2015](https://arxiv.org/html/2607.20451#bib.bib4); Williams et al\.,[2018](https://arxiv.org/html/2607.20451#bib.bib36); Wang et al\.,[2018](https://arxiv.org/html/2607.20451#bib.bib34),[2019](https://arxiv.org/html/2607.20451#bib.bib35)\)\. The key prediction is not simply that higher order improves performance; flexible models often improve when more parameters are added\. The stronger prediction is thatR3R\_\{3\}should selectively explain cases involving non\-additive joint construal\.
### 11\.2Residual diagnostics
Construct minimal triples such as literal controls, metaphoric triples, coercion triples, and idiomatic triples\. For each sentence, computeR1,R2,R3R\_\{1\},R\_\{2\},R\_\{3\}and compare the normalized third\-order ratio
ρ3\(𝐰\)=R3\(𝐰\)R1\(𝐰\)\+R2\(𝐰\)\+R3\(𝐰\)\+ϵ\.\\rho\_\{3\}\(\\mathbf\{w\}\)=\\frac\{R\_\{3\}\(\\mathbf\{w\}\)\}\{R\_\{1\}\(\\mathbf\{w\}\)\+R\_\{2\}\(\\mathbf\{w\}\)\+R\_\{3\}\(\\mathbf\{w\}\)\+\\epsilon\}\.\(47\)SFT predicts thatρ3\\rho\_\{3\}should be higher for constructions whose interpretation depends on joint three\-way interaction than for matched literal controls\.
### 11\.3Field\-overlap predictions
The compatibility factorκA\\kappa\_\{A\}and overlap integrals give predictions for lexical substitution, priming, and semantic acceptability\. If two words have high field overlap but differ in covariance, they may be near in center\-based similarity but behave differently under interaction\. This supplies a test that distinguishes field semantics from point\-vector semantics\.
### 11\.4Stabilization diagnostics
Energy\-based interpretation yields additional observables: convergence time, gradient norm decay, energy gaps, basin sensitivity, and finite\-temperature entropy\. Sentences with unstable or ambiguous readings should exhibit flatter minima, smaller energy gaps, or multiple local attractors\. These diagnostics can be compared against human ambiguity judgments or model uncertainty\.
### 11\.5Multilingual structure
If translation equivalents share partial field geometry but differ in covariance and interaction coefficients, multilingual encoders should permit partial alignment of centers while preserving language\-specific deformation patterns\. This predicts a middle position between strict semantic identity across languages and complete incommensurability\.
## 12Scope and limitations
SFT is a formal hypothesis, not a completed semantics\. Several limitations are essential\.
First, the semantic spaceSSis model\-dependent\. Different encoders, training corpora, and supervision regimes may induce different geometries\. SFT therefore requires identifiability analysis: which field parameters are stable under reparameterization, and which are artifacts of the coordinate system?
Second, higher\-order interaction is expensive\. A full third\-order model is already cubic in sentence length\. Any practical implementation must use sparsity, syntax, local windows, or learned pruning\. Such approximations are not merely engineering details; they determine which interactions the theory can detect\.
Third, SFT does not solve grounding or normativity\. It can describe stable representational regularities induced by language use, but it does not replace the public practices through which meanings are taught, contested, and corrected\.
Fourth, relation to transformers remains methodological rather than identificatory\. A transformer may contain patterns that can be approximated by SFT, but no attention head or hidden layer should be declared a semantic field without explicit modeling and validation\.
Fifth, the strong philosophical phrase “intrinsic mathematical structure” should be treated as a research hypothesis\. The scientifically safer formulation is that language use induces mathematically tractable regularities at a representational level\. The stronger metaphysical reading is not required for the model class\.
## 13Conclusion
This paper has reconstructed Semantic Field Theory as an evolving formal program\. The historical origin lies in a philosophical critique of strong anti\-formalist interpretations of language games\. The underlying intuition of lexical semantic fields was present from the beginning, although only later reformulated in explicit mathematical and computational terms\. The contemporary mathematical form is a model class in which lexical items are lifted to distributed fields, contexts deform those fields, utterances activate subset\-indexed interaction complexes, irreducible higher\-order effects are isolated by Mobius residuals, and interpretation is computed as stabilization in an energy landscape\.
The paper’s main scientific claim is limited but testable: a tractable level of semantic organization may be captured by field geometry, higher\-order residual structure, and energy\-based inference\. The Gaussian product closure result makes interaction geometry explicit\. The order spectrum turns the three\-word problem into a measurable diagnostic\. The energy formulation gives SFT an implementable inference procedure with existence and stability guarantees\. The worked summer\-day example illustrates how this formal apparatus can be reduced to a minimal executable pipeline without confusing the toy instantiation with a trained semantic model\.
The resulting theory does not deny that meaning is public, social, and normative\. Instead, it proposes that public language use can leave stable mathematical structure without being reducible to that structure\. In that sense, SFT aims to occupy a middle position: stronger than metaphorical talk about semantic fields, weaker than a complete reduction of language to vectors, and sufficiently formal to be estimated, ablated, and criticized\.
## Acknowledgments
The author thanks TWT GmbH Science & Innovation and NIKI Digital Engineering for support\. He also thanks S\. Katsioli for helpful discussions\.
## Appendix ASupplementary calculations
### A\.1Gradients of Gaussian fields
For a deformed Gaussian field
L~i\(x\)=a~iexp\[−12\(x−c~i\)⊤Pi\(x−c~i\)\],\\widetilde\{L\}\_\{i\}\(x\)=\\tilde\{a\}\_\{i\}\\exp\\left\[\-\\frac\{1\}\{2\}\(x\-\\tilde\{c\}\_\{i\}\)^\{\\top\}P\_\{i\}\(x\-\\tilde\{c\}\_\{i\}\)\\right\],wherePi=Σ~i−1P\_\{i\}=\\widetilde\{\\Sigma\}\_\{i\}^\{\-1\},
∇L~i\(x\)=−L~i\(x\)Pi\(x−c~i\)\.\\nabla\\widetilde\{L\}\_\{i\}\(x\)=\-\\widetilde\{L\}\_\{i\}\(x\)P\_\{i\}\(x\-\\tilde\{c\}\_\{i\}\)\.\(48\)For an interaction termΦA\(x\)=λA∏i∈AL~i\(x\)\\Phi\_\{A\}\(x\)=\\lambda\_\{A\}\\prod\_\{i\\in A\}\\widetilde\{L\}\_\{i\}\(x\),
∇ΦA\(x\)=ΦA\(x\)∑i∈A\[−Pi\(x−c~i\)\]\.\\nabla\\Phi\_\{A\}\(x\)=\\Phi\_\{A\}\(x\)\\sum\_\{i\\in A\}\\big\[\-P\_\{i\}\(x\-\\tilde\{c\}\_\{i\}\)\\big\]\.\(49\)Thus
∇U\(x∣𝐰\)=βx−∑A∈𝒦p\(𝐰\)ΦA\(x∣𝐰\)∑i∈A\[−Pi\(x−c~i\)\]\.\\nabla U\(x\\mid\\mathbf\{w\}\)=\\beta x\-\\sum\_\{A\\in\\mathcal\{K\}\_\{p\}\(\\mathbf\{w\}\)\}\\Phi\_\{A\}\(x\\mid\\mathbf\{w\}\)\\sum\_\{i\\in A\}\\big\[\-P\_\{i\}\(x\-\\tilde\{c\}\_\{i\}\)\\big\]\.\(50\)
### A\.2Closed\-form pairwise values for the summer\-day example
For the isotropic summer\-day example in Section[6](https://arxiv.org/html/2607.20451#S6), pairwise interaction tensions reduce to
Tij=12‖ci−cj‖2,κij=exp\(−Tij/2\)\.T\_\{ij\}=\\frac\{1\}\{2\}\\\|c\_\{i\}\-c\_\{j\}\\\|^\{2\},\\qquad\\kappa\_\{ij\}=\\exp\(\-T\_\{ij\}/2\)\.\(51\)The resulting approximate values are
These numbers are not linguistic measurements\. They are sanity checks for the toy field geometry: the pair*helios–zesti*is closest in thermal intensity, while*thalassa–zesti*has greater tension because the marine and heat dimensions pull in different directions\.
### A\.3Closed\-formL2L^\{2\}norm for Gaussian interactions
If
fA\(x\)=CAexp\[−12\(x−μA\)⊤PA\(x−μA\)\],f\_\{A\}\(x\)=C\_\{A\}\\exp\\left\[\-\\frac\{1\}\{2\}\(x\-\\mu\_\{A\}\)^\{\\top\}P\_\{A\}\(x\-\\mu\_\{A\}\)\\right\],withPA≻0P\_\{A\}\\succ 0, then
‖fA‖L2\(ℝn\)2=CA2∫ℝnexp\[−\(x−μA\)⊤PA\(x−μA\)\]𝑑x=CA2πn/2\|PA\|−1/2\.\\\|f\_\{A\}\\\|\_\{L^\{2\}\(\\mathbb\{R\}^\{n\}\)\}^\{2\}=C\_\{A\}^\{2\}\\int\_\{\\mathbb\{R\}^\{n\}\}\\exp\\left\[\-\(x\-\\mu\_\{A\}\)^\{\\top\}P\_\{A\}\(x\-\\mu\_\{A\}\)\\right\]dx=C\_\{A\}^\{2\}\\pi^\{n/2\}\|P\_\{A\}\|^\{\-1/2\}\.\(52\)This permits exact residual norms when residuals consist of single Gaussian interaction terms\. When residuals are sums of Gaussian terms, pairwise cross\-integrals also have closed forms by the same product rule\.
### A\.4Algorithmic schema
Input:sentence𝐰=\(w1,…,wm\)\\mathbf\{w\}=\(w\_\{1\},\\ldots,w\_\{m\}\); orderpp; parametersΘ\\Theta; step sizeη\\eta; toleranceϵ\\epsilon\. Output:stabilized statex∗\(𝐰\)x^\{\*\}\(\\mathbf\{w\}\), order spectrumSpecp\(𝐰\)\\operatorname\{Spec\}\_\{p\}\(\\mathbf\{w\}\)\.
1. 1\.Map each typetit\_\{i\}to base field parameters\(ati,cti,Σti\)\(a\_\{t\_\{i\}\},c\_\{t\_\{i\}\},\\Sigma\_\{t\_\{i\}\}\)\.
2. 2\.Compute contextual representationqϕ\(𝐰\)q\_\{\\phi\}\(\\mathbf\{w\}\)\.
3. 3\.Deform each lexical field toL~i\(⋅∣𝐰\)\\widetilde\{L\}\_\{i\}\(\\cdot\\mid\\mathbf\{w\}\)\.
4. 4\.Select interaction complex𝒦p\(𝐰\)\\mathcal\{K\}\_\{p\}\(\\mathbf\{w\}\)\.
5. 5\.Compute coefficientsλA\(𝐰\)\\lambda\_\{A\}\(\\mathbf\{w\}\)forA∈𝒦p\(𝐰\)A\\in\\mathcal\{K\}\_\{p\}\(\\mathbf\{w\}\)\.
6. 6\.Constructℒp\(x∣𝐰\)=∑AΦA\(x∣𝐰\)\\mathcal\{L\}\_\{p\}\(x\\mid\\mathbf\{w\}\)=\\sum\_\{A\}\\Phi\_\{A\}\(x\\mid\\mathbf\{w\}\)and energyU\(x∣𝐰\)U\(x\\mid\\mathbf\{w\}\)\.
7. 7\.Initializex\(0\)x^\{\(0\)\}by a pooled center, random restart, or encoder projection\.
8. 8\.Iteratex\(t\+1\)=x\(t\)−η∇U\(x\(t\)∣𝐰\)x^\{\(t\+1\)\}=x^\{\(t\)\}\-\\eta\\nabla U\(x^\{\(t\)\}\\mid\\mathbf\{w\}\)until convergence\.
9. 9\.Compute residualsΔA\\Delta\_\{A\}and order spectrumSpecp\(𝐰\)\\operatorname\{Spec\}\_\{p\}\(\\mathbf\{w\}\)\.
10. 10\.Returnx∗\(𝐰\)=x\(t\+1\)x^\{\*\}\(\\mathbf\{w\}\)=x^\{\(t\+1\)\}and diagnostics\.
## References
- Bender and Koller \(2020\)Emily M\. Bender and Alexander Koller\. 2020\.Climbing towards NLU: On meaning, form, and understanding in the age of data\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics*, pages 5185–5198\.
- Bender et al\. \(2021\)Emily M\. Bender, Timnit Gebru, Angelina McMillan\-Major, and Shmargaret Shmitchell\. 2021\.On the dangers of stochastic parrots: Can language models be too big?In*Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency*, pages 610–623\.
- Bommasani et al\. \(2021\)Rishi Bommasani, Drew A\. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S\. Bernstein, Jeannette Bohg, Anja Butterfield, and many others\. 2021\.On the opportunities and risks of foundation models\.*arXiv preprint*, arXiv:2108\.07258\.
- Bowman et al\. \(2015\)Samuel R\. Bowman, Gabor Angeli, Christopher Potts, and Christopher D\. Manning\. 2015\.A large annotated corpus for learning natural language inference\.In*Proceedings of EMNLP*, pages 632–642\.
- Brown et al\. \(2020\)Tom B\. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and others\. 2020\.Language models are few\-shot learners\.In*Advances in Neural Information Processing Systems*, 33:1877–1901\.
- Clark et al\. \(2019\)Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D\. Manning\. 2019\.What does BERT look at? An analysis of BERT’s attention\.In*Proceedings of the 2019 ACL Workshop BlackboxNLP*, pages 276–286\.
- Coecke et al\. \(2010\)Bob Coecke, Mehrnoosh Sadrzadeh, and Stephen Clark\. 2010\.Mathematical foundations for a compositional distributional model of meaning\.*Linguistic Analysis*, 36\(1–4\):345–384\.
- Devlin et al\. \(2019\)Jacob Devlin, Ming\-Wei Chang, Kenton Lee, and Kristina Toutanova\. 2019\.BERT: Pre\-training of deep bidirectional transformers for language understanding\.In*Proceedings of NAACL\-HLT*, pages 4171–4186\.
- Ethayarajh \(2019\)Kawin Ethayarajh\. 2019\.How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT\-2 embeddings\.In*Proceedings of EMNLP\-IJCNLP*, pages 55–65\.
- Firth \(1957\)J\. R\. Firth\. 1957\.*Papers in Linguistics 1934–1951*\.Oxford University Press\.
- Grefenstette et al\. \(2011\)Edward Grefenstette, Mehrnoosh Sadrzadeh, Stephen Clark, Bob Coecke, and Stephen Pulman\. 2011\.Concrete sentence spaces for compositional distributional models of meaning\.In*Proceedings of IWCS 2011*\.
- Harris \(1954\)Zellig S\. Harris\. 1954\.Distributional structure\.*WORD*, 10\(2–3\):146–162\.
- Hewitt and Manning \(2019\)John Hewitt and Christopher D\. Manning\. 2019\.A structural probe for finding syntax in word representations\.In*Proceedings of NAACL\-HLT*, pages 4129–4138\.
- Jain and Wallace \(2019\)Sarthak Jain and Byron C\. Wallace\. 2019\.Attention is not explanation\.In*Proceedings of NAACL\-HLT*, pages 3543–3556\.
- Kaplan et al\. \(2020\)Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B\. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei\. 2020\.Scaling laws for neural language models\.*arXiv preprint*, arXiv:2001\.08361\.
- Kripke \(1982\)Saul A\. Kripke\. 1982\.*Wittgenstein on Rules and Private Language: An Elementary Exposition*\.Harvard University Press\.
- LeCun et al\. \(2006\)Yann LeCun, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu\-Jie Huang\. 2006\.A tutorial on energy\-based learning\.In G\. Bakir, T\. Hofmann, B\. Scholkopf, A\. Smola, B\. Taskar, and S\. Vishwanathan, editors,*Predicting Structured Data*, MIT Press\.
- Marelli et al\. \(2014\)Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli\. 2014\.A SICK cure for the evaluation of compositional distributional semantic models\.In*Proceedings of LREC*, pages 216–223\.
- Mikolov et al\. \(2013\)Tomas Mikolov, Wen\-tau Yih, and Geoffrey Zweig\. 2013\.Linguistic regularities in continuous space word representations\.In*Proceedings of NAACL\-HLT*, pages 746–751\.
- Mitchell and Lapata \(2008\)Jeff Mitchell and Mirella Lapata\. 2008\.Vector\-based models of semantic composition\.In*Proceedings of ACL\-08: HLT*, pages 236–244\.
- Mitchell and Lapata \(2010\)Jeff Mitchell and Mirella Lapata\. 2010\.Composition in distributional models of semantics\.*Cognitive Science*, 34\(8\):1388–1429\.
- Pennington et al\. \(2014\)Jeffrey Pennington, Richard Socher, and Christopher D\. Manning\. 2014\.GloVe: Global vectors for word representation\.In*Proceedings of EMNLP*, pages 1532–1543\.
- Peters et al\. \(2018\)Matthew E\. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer\. 2018\.Deep contextualized word representations\.In*Proceedings of NAACL\-HLT*, pages 2227–2237\.
- Rogers et al\. \(2020\)Anna Rogers, Olga Kovaleva, and Anna Rumshisky\. 2020\.A primer in BERTology: What we know about how BERT works\.*Transactions of the Association for Computational Linguistics*, 8:842–866\.
- Rota \(1964\)Gian\-Carlo Rota\. 1964\.On the foundations of combinatorial theory I\. Theory of Mobius functions\.*Zeitschrift fuer Wahrscheinlichkeitstheorie und Verwandte Gebiete*, 2:340–368\.
- Tenney et al\. \(2019\)Ian Tenney, Dipanjan Das, and Ellie Pavlick\. 2019\.BERT rediscovers the classical NLP pipeline\.In*Proceedings of ACL*, pages 4593–4601\.
- Turney and Pantel \(2010\)Peter D\. Turney and Patrick Pantel\. 2010\.From frequency to meaning: Vector space models of semantics\.*Journal of Artificial Intelligence Research*, 37:141–188\.
- Vartziotis \(2012\)Dimitris Vartziotis\. 2012\.*Scholia se stochasmous tou Ludwig Wittgenstein \[Scholia on Reflections of Ludwig Wittgenstein\]*\.Lefki Selida\.
- Vartziotis \(2017\)Dimitris Vartziotis\. 2017\.*Kommentare zu Wittgensteins Zitaten*\.Literareon – Utz Verlag\.
- Vartziotis \(in press\-a\)Dimitris Vartziotis\. In press\.*Wittgenstein and the End of Language Games: Mathematical Semantics, Field Theory, and the Age of Predictive Language Models*\.Literareon – Utz Verlag\.
- Vartziotis \(in press\-b\)Dimitris Vartziotis\. In press\.*The Revolution of LLMs: Philosophical and Mathematical Conjugations*\.Literareon – Utz Verlag\.
- Vartziotis \(2026\)Dimitris Vartziotis\. 2026\.Language as mathematical structure: Examining Semantic Field Theory against language games\.*arXiv preprint*, arXiv:2601\.00448\.
- Vaswani et al\. \(2017\)Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N\. Gomez, Lukasz Kaiser, and Illia Polosukhin\. 2017\.Attention is all you need\.In*Advances in Neural Information Processing Systems*, 30:5998–6008\.
- Wang et al\. \(2018\)Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R\. Bowman\. 2018\.GLUE: A multi\-task benchmark and analysis platform for natural language understanding\.In*Proceedings of the BlackboxNLP Workshop*, pages 353–355\.
- Wang et al\. \(2019\)Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R\. Bowman\. 2019\.SuperGLUE: A stickier benchmark for general\-purpose language understanding systems\.In*Advances in Neural Information Processing Systems*, 32\.
- Williams et al\. \(2018\)Adina Williams, Nikita Nangia, and Samuel R\. Bowman\. 2018\.A broad\-coverage challenge corpus for sentence understanding through inference\.In*Proceedings of NAACL\-HLT*, pages 1112–1122\.
- Wittgenstein \(1953\)Ludwig Wittgenstein\. 1953\.*Philosophical Investigations*\.Blackwell\.Similar Articles
Asking the Wrong Questions: The HuggingFace Incident and the (Mis)Calibration of Semantic Fields in LLMs
This paper discusses the OpenAI-HuggingFace incident and proposes a theoretical framework using semantic field mechanics to explain how large language models process meaning.
LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting
This paper proposes Task-Semantic Field Factorization (TSF), an LLM-guided framework that uses offline semantic construction from process documents to enhance time-series forecasting and soft sensing in industrial processes. TSF reduces MAE by 6.4% on average with nearly negligible added parameters and inference overhead.
Computational conceptual history of scientific concepts: From early digital methods to LLMs
This paper situates large language models within the broader history of computational approaches to concept analysis in the history, philosophy, and sociology of science (HPSS), reviewing methodological challenges and LLM-based case studies for lexical semantic change detection. It covers corpus construction, operationalization, and evaluation across both pre-LLM and LLM-era workflows.
Foundations of Stochastic Lexical Calculus: Semantic Descent and Random Dynamics on Probability Simplices
This paper develops a stochastic lexical calculus framework for semantic updates in language models, defining conditions for when language-derived probabilities support meaningful sequential state representations. Empirical experiments validate the framework's stability and coverage under calibrated conditions.
Language Modeling with Hyperspherical Flows
This paper introduces S-FLM, a novel flow-based language model that operates in a hyperspherical latent space to address the computational costs and semantic limitations of existing discrete diffusion and continuous flow models.