Exact semantic readout from compressed vector representations

arXiv cs.CL Papers

Summary

This paper develops a vector logic for formal semantics, characterizing when compressed vector representations allow exact linear or affine readouts of truth conditions, with experiments on GloVe and word2vec embeddings.

arXiv:2609.18047v1 Announce Type: new Abstract: We characterize when compressed vector representations admit exact linear or affine readouts of a finite lexicon's truth conditions: one fixed map per predicate, sending each entity vector to the corresponding truth vector. A necessary and sufficient row-space condition determines existence; the augmented truth matrix has rank r, giving minimum dimension r in the linear case, and r-1 in the affine. Exact readouts return values in a shared truth basis on which Boolean connectives act unchanged; separability alone requires an intervening threshold. For binary relations, exact bilinear readout of identity or strict total order requires linearly independent entity vectors. Experiments with GloVe and word2vec distinguish exact affine recovery, linear separability, and held-out prediction: most predicates are strictly separable, but none admits an exact affine readout from the pretrained embeddings. Supervised transductive training attains exact affine recovery to numerical precision at every tested dimension meeting the bound. At the embeddings' original dimension, geometries constrained to exact linear recovery retain 98-99 percent of the pretrained variance on the feature norms, and 80-83 percent on the WordNet lexicon.
Original Article
View Cached Full Text

Cached at: 09/17/26, 09:05 AM

# Exact semantic readout from compressed vector representations
Source: [https://arxiv.org/html/2609.18047](https://arxiv.org/html/2609.18047)
## Exact semantic readout from compressed vector representationsThanks:Third in a series developing a vector logic for formal semantics\[[25](https://arxiv.org/html/2609.18047#bib.bib1),[26](https://arxiv.org/html/2609.18047#bib.bib2)\]\.

[Daniel Quigley](https://orcid.org/0009-0004-7957-1806)Affiliation:Center for Possible MindsAffiliation:Indiana University BloomingtonAffiliation:Bloomington, IN 47408Email:[dgquigle@iu\.edu](mailto:)

###### Abstract

We characterize when compressed vector representations admit exact linear or affine readouts of a finite lexicon’s truth conditions: one fixed map per predicate, sending each entity vector to the corresponding truth vector\. A necessary and sufficient row\-space condition determines existence; the augmented truth matrix has rankrr, giving minimum dimensionrrin the linear case, andr−1r\-1in the affine\. Exact readouts return values in a shared truth basis on which Boolean connectives act unchanged; separability alone requires an intervening threshold\. For binary relations, exact bilinear readout of identity or strict total order requires linearly independent entity vectors\. Experiments with GloVe and word2vec distinguish exact affine recovery, linear separability, and held\-out prediction: most predicates are strictly separable, but none admits an exact affine readout from the pretrained embeddings\. Supervised transductive training attains exact affine recovery to numerical precision at every tested dimension meeting the bound\. At the embeddings’ original dimension, geometries constrained to exact linear recovery retain 98–99 percent of the pretrained variance on the feature norms, and 80–83 percent on the WordNet lexicon\.

*Keywords*formal semantics⋅\\cdotvector logic⋅\\cdotsemantic space⋅\\cdotencoding⋅\\cdotembedding⋅\\cdotmathematical linguistics

## 1Introduction

Montagovian semantics represents entities and predicates in typed domains, with composition governed by the model’s functions\[[20](https://arxiv.org/html/2609.18047#bib.bib6),[9](https://arxiv.org/html/2609.18047#bib.bib7)\]\. Distributional embeddings represent lexical items as vectors, whose geometry reflects patterns of use\[[4](https://arxiv.org/html/2609.18047#bib.bib8)\]\. When these vectors represent the entities of a finite semantic model, which predicates can be recovered by linear readouts, and what does compression prevent?

A vector logic for formal semantics\[[25](https://arxiv.org/html/2609.18047#bib.bib1),[26](https://arxiv.org/html/2609.18047#bib.bib2)\]supplies an exact construction: each primitive domain element receives its own basis vector, and semantic functions extend linearly from free carriers, preserving composition along primitive intermediate types\. Linear independence makes these lifts possible\. Learned embeddings generally have fewer dimensions than entities, and, therefore, introduce linear dependences that may obstruct the lifts\.

An embedding lookup with matrix𝐄∈ℝV×d\\mathbf\{E\}\\in\\mathbb\{R\}^\{V\\times d\}is equivalent, on one\-hot inputs, to a bias\-free linear map\[[27](https://arxiv.org/html/2609.18047#bib.bib3)\]\. It, therefore, factors through the free entity carrier as𝐄⊤:ℝV→ℝd\\mathbf\{E\}^\{\\top\}:\\mathbb\{R\}^\{V\}\\to\\mathbb\{R\}^\{d\}\. This identifies two regimes: a*free*geometry has linearly independent entity vectors, and admits every semantic lift; a*compressed*geometry has dependent vectors, and admits only those lifts compatible with its dependences\. Compression need not erase entity identity: distinct columns still permit arbitrary predicate lookup on a finite domain; we are interested, then, in which predicates remain linearly recoverable\.

For a monadic lexicon with truth matrix𝐓\\mathbf\{T\}, we show that exact readouts into a shared truth basis exist precisely when the geometry’s row space containsrow⁡\[𝐓;𝟏\]\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]\. The rank of this augmented truth matrix gives the minimum dimension, while its kernel specifies the admissible dependences among entity vectors\. A restricted lexicon can, therefore, admit exact compression, although requiring every predicate forces the free regime\. Once atomic readouts return the truth basis, Boolean connectives act unchanged\.

Exact readout is stronger than linear separability: a score may distinguish true instances from false ones without, itself, returning their truth values\. This distinction connects linear probing\[[3](https://arxiv.org/html/2609.18047#bib.bib17)\]to the vector logic\. Whether such a readout exists, and whether it can be learned from a subset of entities, are separate questions\. We test exact recovery, separability, and held\-out prediction on distributional embeddings, then examine how much pretrained variance survives supervised training toward exact recovery\. Figure[1](https://arxiv.org/html/2609.18047#S1.F1)summarizes the program and the dependencies among these questions\.

Free carrierEmbeddingH=E⊤H=E^\{\\top\}Exact readoutrow⁡\[T;𝟏\]⊆row⁡\(H\)\\operatorname\{row\}\[T;\\mathbf\{1\}\]\\subseteq\\operatorname\{row\}\(H\)Semantic computationCompression boundsDefect and separabilitySupervised geometriesPretrained geometriesFigure 1:Dependencies among the formal results and empirical analyses\.
## 2Vector logic

We recall from\[[25](https://arxiv.org/html/2609.18047#bib.bib1)\]the definitions and the homomorphism theorem, and from\[[26](https://arxiv.org/html/2609.18047#bib.bib2)\]the index\-sort material\. We refrain from re\-proving the relevant content; see, instead, papers proper therein\.

### 2\.1Extensional models and free carrier

Types are generated fromeeandttby⟨σ,τ⟩\\langle\\sigma,\\tau\\rangle\. A typed extensional modelℳe​x​t=⟨\(𝒟τ\)τ,ℐ⟩\\mathcal\{M\}\_\{ext\}=\\langle\(\\mathcal\{D\}\_\{\\tau\}\)\_\{\\tau\},\\mathcal\{I\}\\ranglehas entity domain𝒟e\\mathcal\{D\}\_\{e\}, truth domain𝒟t=\{1,0\}\\mathcal\{D\}\_\{t\}=\\\{1,0\\\}, function domains𝒟⟨σ,τ⟩=𝒟τ𝒟σ\\mathcal\{D\}\_\{\\langle\\sigma,\\tau\\rangle\}=\\mathcal\{D\}\_\{\\tau\}^\{\\mathcal\{D\}\_\{\\sigma\}\}, and an interpretationℐ\\mathcal\{I\}; denotations⟦⋅⟧\\llbracket\\cdot\\rrbracketfollow the recursion of\[[9](https://arxiv.org/html/2609.18047#bib.bib7)\]\. Throughout, let𝒟e\\mathcal\{D\}\_\{e\}be finite111The finiteness restriction is what makes the rank statements below meaningful; the extensional theorem itself holds for domains of any cardinality\., with\|𝒟e\|=V\|\\mathcal\{D\}\_\{e\}\|=Vand elementsd1,…,dVd\_\{1\},\\dots,d\_\{V\}\.

The vector space modelℳ𝒮\\mathcal\{M\}\_\{\\mathcal\{S\}\}assigns to each domain a real vector space𝒮𝒟τ\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{\\tau\}\}and an injectionhτ:𝒟τ→𝒮𝒟τh\_\{\\tau\}:\\mathcal\{D\}\_\{\\tau\}\\to\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{\\tau\}\}\. The construction takes the form

he​\(di\)=𝐞i∈ℝV,ht​\(1\)=𝐛1=\[10\],ht​\(0\)=𝐛0=\[01\],h\_\{e\}\(d\_\{i\}\)=\\mathbf\{e\}\_\{i\}\\in\\mathbb\{R\}^\{V\},\\qquad h\_\{t\}\(1\)=\\mathbf\{b\}\_\{1\}=\\begin\{bmatrix\}1\\\\ 0\\end\{bmatrix\},\\qquad h\_\{t\}\(0\)=\\mathbf\{b\}\_\{0\}=\\begin\{bmatrix\}0\\\\ 1\\end\{bmatrix\},where𝐞i\\mathbf\{e\}\_\{i\}is theii\-th standard basis vector, and for a function type sendsf∈𝒟⟨σ,τ⟩f\\in\\mathcal\{D\}\_\{\\langle\\sigma,\\tau\\rangle\}to the linear mapLfL\_\{f\}, determined on the basis\{𝐞a\}a∈𝒟σ\\\{\\mathbf\{e\}\_\{a\}\\\}\_\{a\\in\\mathcal\{D\}\_\{\\sigma\}\}of the free carrier ofσ\\sigmabyLf​𝐞a=hτ​\(f⁡\(a\)\)L\_\{f\}\\,\\mathbf\{e\}\_\{a\}=h\_\{\\tau\}\(f\(a\)\); a linear map is determined by its values on a basis, so this fixesLfL\_\{f\}uniquely\. Note that\[[25](https://arxiv.org/html/2609.18047#bib.bib1)\]embeds every domain, function\-type domains included, by basis vectors, and defines the lifthfh\_\{f\}pointwise on the image ofhσh\_\{\\sigma\}through the left inversehσ−1h\_\{\\sigma\}^\{\-1\};\[[26](https://arxiv.org/html/2609.18047#bib.bib2)\]attaches to each type a free carrierℱτ\\mathcal\{F\}\_\{\\tau\}, with basis indexed by𝒟τ\\mathcal\{D\}\_\{\\tau\}, and an operator carrier𝒮τ\\mathcal\{S\}\_\{\\tau\}, with𝒮⟨σ,τ⟩=Hom⁡\(ℱσ,𝒮τ\)\\mathcal\{S\}\_\{\\langle\\sigma,\\tau\\rangle\}=\\operatorname\{Hom\}\(\\mathcal\{F\}\_\{\\sigma\},\\mathcal\{S\}\_\{\\tau\}\), so that the free carrier stands in argument position, and the two carriers coincide on primitive types\. We adopt the second convention, since the results below concern linear readouts on the entity space, whose type is primitive, and we call the family\{ℱτ\}\\\{\\mathcal\{F\}\_\{\\tau\}\\\}the free carrier: elements of primitive domains go to basis vectors, and the primitive spaces are free on their domains\. In particular, a monadic predicateP∈𝒟⟨e,t⟩P\\in\\mathcal\{D\}\_\{\\langle e,t\\rangle\}becomes the2×V2\\times Vmatrix

𝐌P=\[ht\(P\(d1\)\)⋯ht\(P\(dV\)\)\],\\mathbf\{M\}\_\{P\}=\\bigl\[\\,h\_\{t\}\(P\(d\_\{1\}\)\)\\;\\;\\cdots\\;\\;h\_\{t\}\(P\(d\_\{V\}\)\)\\,\\bigr\],whoseii\-th column is the truth vector thatPPassigns todid\_\{i\}, and𝐌P​𝐞i=ht​\(P⁡\(di\)\)\\mathbf\{M\}\_\{P\}\\mathbf\{e\}\_\{i\}=h\_\{t\}\(P\(d\_\{i\}\)\)is functional application\. Truth\-functional connectives become fixed matrices on tensor powers of the truth space, annn\-ary connectiveccacting as a2×2n2\\times 2^\{n\}matrix𝐌c\\mathbf\{M\}\_\{c\}on⨂iht​\(ti\)\\bigotimes\_\{i\}h\_\{t\}\(t\_\{i\}\); for instance

𝐌¬=\[0110\],𝐌∧=\[10000111\],\\mathbf\{M\}\_\{\\neg\}=\\begin\{bmatrix\}0&1\\\\ 1&0\\end\{bmatrix\},\\qquad\\mathbf\{M\}\_\{\\wedge\}=\\begin\{bmatrix\}1&0&0&0\\\\ 0&1&1&1\\end\{bmatrix\},in the ordering𝐛1⊗𝐛1,𝐛1⊗𝐛0,𝐛0⊗𝐛1,𝐛0⊗𝐛0\\mathbf\{b\}\_\{1\}\\otimes\\mathbf\{b\}\_\{1\},\\mathbf\{b\}\_\{1\}\\otimes\\mathbf\{b\}\_\{0\},\\mathbf\{b\}\_\{0\}\\otimes\\mathbf\{b\}\_\{1\},\\mathbf\{b\}\_\{0\}\\otimes\\mathbf\{b\}\_\{0\}of the tensor basis, after\[[19](https://arxiv.org/html/2609.18047#bib.bib10),[31](https://arxiv.org/html/2609.18047#bib.bib11)\]\.

### 2\.2Homomorphism theorem

###### Theorem 2\.1\.

For every extensional modelℳe​x​t\\mathcal\{M\}\_\{ext\}, there exist injections\{hτ\}\\\{h\_\{\\tau\}\\\}into vector spaces\{𝒮𝒟τ\}\\\{\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{\\tau\}\}\\\}such that every semantic functionf:𝒟σ→𝒟τf:\\mathcal\{D\}\_\{\\sigma\}\\to\\mathcal\{D\}\_\{\\tau\}has a unique linear liftLf:ℱσ→𝒮𝒟τL\_\{f\}:\\mathcal\{F\}\_\{\\sigma\}\\to\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{\\tau\}\}from the free carrier ofσ\\sigma, withLf​\(𝐛a\)=hτ​\(f⁡\(a\)\)L\_\{f\}\(\\mathbf\{b\}\_\{a\}\)=h\_\{\\tau\}\(f\(a\)\)for everya∈𝒟σa\\in\\mathcal\{D\}\_\{\\sigma\}, where𝐛a\\mathbf\{b\}\_\{a\}is the free encoding ofaa, which coincides withhσ​\(a\)h\_\{\\sigma\}\(a\)whenσ\\sigmais primitive;nn\-ary functions lift to multilinear maps on the free carriers, and composition of semantic functions corresponds to composition of lifts along primitive intermediate types\.

The square

𝒟σ\{\\lx@inpgf@ignorespaces\\mathcal\{D\}\_\{\\sigma\}\}𝒟τ\{\\lx@inpgf@ignorespaces\\mathcal\{D\}\_\{\\tau\}\}𝒮𝒟σ\{\\lx@inpgf@ignorespaces\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{\\sigma\}\}\}𝒮𝒟τ\{\\lx@inpgf@ignorespaces\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{\\tau\}\}\}f\\scriptstyle\{\\lx@inpgf@ignorespaces f\}hσ\\scriptstyle\{\\lx@inpgf@ignorespaces h\_\{\\sigma\}\}hτ\\scriptstyle\{\\lx@inpgf@ignorespaces h\_\{\\tau\}\}Lf\\scriptstyle\{\\lx@inpgf@ignorespaces L\_\{f\}\}

commutes for everyffwith primitive argument type, which covers every function below, so the recursive evaluation of any expression inℳe​x​t\\mathcal\{M\}\_\{ext\}has a step\-by\-step counterpart inℳ𝒮\\mathcal\{M\}\_\{\\mathcal\{S\}\}once arguments are carried in their free encodings\. The injections are non\-surjective by design:im⁡\(ht\)=\{𝐛1,𝐛0\}\\operatorname\{im\}\(h\_\{t\}\)=\\\{\\mathbf\{b\}\_\{1\},\\mathbf\{b\}\_\{0\}\\\}is a proper subset ofℝ2\\mathbb\{R\}^\{2\}, and denotation is confined to the images\.

Theorem[2\.1](https://arxiv.org/html/2609.18047#S2.Thmtheorem1)is existential: the free carrier is one family of injections for which the lifts exist, and other families are the subject of Section[4](https://arxiv.org/html/2609.18047#S4)\. The commuting square for a single function is\[[25](https://arxiv.org/html/2609.18047#bib.bib1)\], where the lift is defined on the image alone; linearity on the span for primitive argument types, the multilinear lift ofnn\-ary functions, and the operator carriers of function types follow\[[26](https://arxiv.org/html/2609.18047#bib.bib2)\], whose descent theorems determine when a functional of a function domain acts linearly on operator encodings as well: on a power set, exactly the constants, the ultrafilter indicators, and their complements, which on a finite domain are the constants, Montague’s individualsλ​P\.P⁡\(d\)\\lambda P\.\\,P\(d\), and their negations\.

### 2\.3Index sorts

The intensional layer adjoins index sorts \(worlds and times, among others\), collected in a compound index spaceS=∏σ𝒟σS=\\prod\_\{\\sigma\}\\mathcal\{D\}\_\{\\sigma\}, with its own free carrierhS​\(s\)=𝐞sh\_\{S\}\(s\)=\\mathbf\{e\}\_\{s\}, the carrierℱs\\mathcal\{F\}\_\{s\}of the compound index typess\. An intensiong:S→𝒟τg:S\\to\\mathcal\{D\}\_\{\\tau\}becomes the linear operator𝒮S→𝒮𝒟τ\\mathcal\{S\}\_\{S\}\\to\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{\\tau\}\}, with𝐞s↦hτ​\(g⁡\(s\)\)\\mathbf\{e\}\_\{s\}\\mapsto h\_\{\\tau\}\(g\(s\)\); a propositionφ\\varphibecomes𝐏φ∈ℝ2×\|S\|\\mathbf\{P\}\_\{\\varphi\}\\in\\mathbb\{R\}^\{2\\times\|S\|\}, with truth profile𝐯⁡\(φ\)∈\{0,1\}\|S\|\\mathbf\{v\}\(\\varphi\)\\in\\\{0,1\\\}^\{\|S\|\}its top row\. Over a single discrete sort of worldsW=\{w1,…,wn\}W=\\\{w\_\{1\},\\dots,w\_\{n\}\\\}with accessibility matrix𝐀∈\{0,1\}n×n\\mathbf\{A\}\\in\\\{0,1\\\}^\{n\\times n\}, the modal operators of\[[26](https://arxiv.org/html/2609.18047#bib.bib2)\]read

\(□​φ\)​\(wi\)=1⇔\(𝐀⁡\(𝟏−𝐯⁡\(φ\)\)\)i=0,\(◇​φ\)​\(wi\)=1⇔\(𝐀​𝐯​\(φ\)\)i\>0,\(\\Box\\varphi\)\(w\_\{i\}\)=1\\iff\(\\mathbf\{A\}\\,\(\\mathbf\{1\}\-\\mathbf\{v\}\(\\varphi\)\)\)\_\{i\}=0,\\qquad\(\\Diamond\\varphi\)\(w\_\{i\}\)=1\\iff\(\\mathbf\{A\}\\,\\mathbf\{v\}\(\\varphi\)\)\_\{i\}\>0,\(1\)a linear accumulation followed by a decision; when the out\-degreedeg𝐀⁡\(wi\)=∑j𝐀i​j\\operatorname\{deg\}\_\{\\mathbf\{A\}\}\(w\_\{i\}\)=\\sum\_\{j\}\\mathbf\{A\}\_\{ij\}is finite, as it is here, the first condition is equivalent to\(𝐀​𝐯​\(φ\)\)i=deg𝐀⁡\(wi\)\(\\mathbf\{A\}\\,\\mathbf\{v\}\(\\varphi\)\)\_\{i\}=\\operatorname\{deg\}\_\{\\mathbf\{A\}\}\(w\_\{i\}\)\. Further, operators are defined over measure frames, in which the counting measure of the discrete case is but one choice among others; a finite geometry matrix carries the discrete case with finitely many indices, the case to which we restrict ourselves in[Section7](https://arxiv.org/html/2609.18047#S7)\.

## 3Embedding lookup on free carrier

The observation that an embedding layer is a linear map on one\-hot inputs is due to\[[27](https://arxiv.org/html/2609.18047#bib.bib3),[28](https://arxiv.org/html/2609.18047#bib.bib4)\], who also explains why frameworks introduce a transpose and why implementations use the lookup form at all\. We develop that here in detail, and apply it to our vector logic\.

### 3\.1Setup

LetV≥1V\\geq 1be the vocabulary size, andd≥1d\\geq 1the embedding dimension\. An embedding matrix is any𝐄∈ℝV×d\\mathbf\{E\}\\in\\mathbb\{R\}^\{V\\times d\}, with rows𝐄i∈ℝ1×d\\mathbf\{E\}\_\{i\}\\in\\mathbb\{R\}^\{1\\times d\}and entriesEi​jE\_\{ij\}\. The lookup isℓ𝐄​\(i\)=𝐄i\\ell\_\{\\mathbf\{E\}\}\(i\)=\\mathbf\{E\}\_\{i\}fori∈\{1,…,V\}i\\in\\\{1,\\dots,V\\\}, and the bias\-free linear layer with weight𝐄\\mathbf\{E\}isL𝐄​\(x\)=x⊤​𝐄L\_\{\\mathbf\{E\}\}\(x\)=x^\{\\top\}\\mathbf\{E\}forx∈ℝVx\\in\\mathbb\{R\}^\{V\}\. Outputs are row vectors, such that batches stack vertically\. Under the reading of the vocabulary as the entity domain,𝐞i=he​\(di\)\\mathbf\{e\}\_\{i\}=h\_\{e\}\(d\_\{i\}\)andL𝐄​\(𝐞i\)=\(𝐄⊤​𝐞i\)⊤L\_\{\\mathbf\{E\}\}\(\\mathbf\{e\}\_\{i\}\)=\(\\mathbf\{E\}^\{\\top\}\\mathbf\{e\}\_\{i\}\)^\{\\top\}; the composite𝐄⊤∘he\\mathbf\{E\}^\{\\top\}\\circ h\_\{e\}is a second injection of𝒟e\\mathcal\{D\}\_\{e\}into a vector space, whenever the rows of𝐄\\mathbf\{E\}are distinct\.

### 3\.2Forward pass

###### Proposition 3\.1\(single token\)\.

For everyi∈\{1,…,V\}i\\in\\\{1,\\dots,V\\\},L𝐄​\(𝐞i\)=ℓ𝐄​\(i\)L\_\{\\mathbf\{E\}\}\(\\mathbf\{e\}\_\{i\}\)=\\ell\_\{\\mathbf\{E\}\}\(i\)\.

###### Proof\.

Fix a columnjj\. By the definition of the matrix product,

\(𝐞i⊤​𝐄\)j=∑k=1V\(𝐞i\)k​Ek​j=∑k=1Vδi​k​Ek​j=Ei​j,\(\\mathbf\{e\}\_\{i\}^\{\\top\}\\mathbf\{E\}\)\_\{j\}=\\sum\_\{k=1\}^\{V\}\(\\mathbf\{e\}\_\{i\}\)\_\{k\}E\_\{kj\}=\\sum\_\{k=1\}^\{V\}\\delta\_\{ik\}E\_\{kj\}=E\_\{ij\},thejjth entry of𝐄i\\mathbf\{E\}\_\{i\}\. Sincejjwas arbitrary,𝐞i⊤​𝐄=𝐄i\\mathbf\{e\}\_\{i\}^\{\\top\}\\mathbf\{E\}=\\mathbf\{E\}\_\{i\}\. ∎

###### Corollary 3\.1\(batch\)\.

Leti1,…,in∈\{1,…,V\}i\_\{1\},\\dots,i\_\{n\}\\in\\\{1,\\dots,V\\\}and let𝐗∈ℝn×V\\mathbf\{X\}\\in\\mathbb\{R\}^\{n\\times V\}have rows𝐞i1⊤,…,𝐞in⊤\\mathbf\{e\}\_\{i\_\{1\}\}^\{\\top\},\\dots,\\mathbf\{e\}\_\{i\_\{n\}\}^\{\\top\}\. Then rowttof𝐗𝐄\\mathbf\{X\}\\mathbf\{E\}equals𝐄it\\mathbf\{E\}\_\{i\_\{t\}\}for everytt, so𝐗𝐄\\mathbf\{X\}\\mathbf\{E\}is the vertical stack ofℓ𝐄​\(i1\),…,ℓ𝐄​\(in\)\\ell\_\{\\mathbf\{E\}\}\(i\_\{1\}\),\\dots,\\ell\_\{\\mathbf\{E\}\}\(i\_\{n\}\)\.

###### Proof\.

Rowttof𝐗𝐄\\mathbf\{X\}\\mathbf\{E\}is𝐞it⊤​𝐄=𝐄it\\mathbf\{e\}\_\{i\_\{t\}\}^\{\\top\}\\mathbf\{E\}=\\mathbf\{E\}\_\{i\_\{t\}\}by Proposition[3\.1](https://arxiv.org/html/2609.18047#S3.Thmproposition1), applied to each row independently; repeated identifiers amongi1,…,ini\_\{1\},\\dots,i\_\{n\}are, likewise, covered\. ∎

###### Corollary 3\.2\(orientation\)\.

With𝐖=𝐄⊤\\mathbf\{W\}=\\mathbf\{E\}^\{\\top\}, the mapx↦x​𝐖⊤x\\mapsto x\\mathbf\{W\}^\{\\top\}agrees withL𝐄L\_\{\\mathbf\{E\}\}onℝV\\mathbb\{R\}^\{V\}and, hence, withℓ𝐄\\ell\_\{\\mathbf\{E\}\}on one\-hot inputs\.

###### Proof\.

𝐖⊤=𝐄\\mathbf\{W\}^\{\\top\}=\\mathbf\{E\}, sox​𝐖⊤=x​𝐄x\\mathbf\{W\}^\{\\top\}=x\\mathbf\{E\}for every row vectorxx; apply Corollary[3\.1](https://arxiv.org/html/2609.18047#S3.Thmcorollary1)\. ∎

𝐖=𝐄⊤\\mathbf\{W\}=\\mathbf\{E\}^\{\\top\}is the geometry matrix: its columns are the entity vectors𝐄i⊤\\mathbf\{E\}\_\{i\}^\{\\top\}, and the stored parameter of the framework’s linear layer is the geometry itself\.

### 3\.3Backward pass

When considering the backward pass, we are, essentially, encountering gradients\. The identity extends to gradients with respect to the parameters, which makes the implementations interchangeable during training, and covers accumulation on repeated identifiers\. Letℒ\\mathcal\{L\}be a scalar loss depending on the parameters only by𝐘=𝐗𝐄∈ℝn×d\\mathbf\{Y\}=\\mathbf\{X\}\\mathbf\{E\}\\in\\mathbb\{R\}^\{n\\times d\}, and write𝐆=∂ℒ/∂𝐘∈ℝn×d\\mathbf\{G\}=\\partial\\mathcal\{L\}/\\partial\\mathbf\{Y\}\\in\\mathbb\{R\}^\{n\\times d\}with rows𝐆1,…,𝐆n\\mathbf\{G\}\_\{1\},\\dots,\\mathbf\{G\}\_\{n\}\. In the lookup, the backward pass is defined as a scatter\-add: the gradient with respect to𝐄\\mathbf\{E\}is initialized to zero, and𝐆t\\mathbf\{G\}\_\{t\}is added to rowiti\_\{t\}for eachtt\.

###### Proposition 3\.2\(gradient\)\.

Under the linear\-layer parameterization,∂ℒ/∂𝐄=𝐗⊤​𝐆\\partial\\mathcal\{L\}/\\partial\\mathbf\{E\}=\\mathbf\{X\}^\{\\top\}\\mathbf\{G\}, whose rowkkequals∑t:it=k𝐆t\\sum\_\{t:\\,i\_\{t\}=k\}\\mathbf\{G\}\_\{t\}, with the empty sum being zero\.

Proposition[3\.2](https://arxiv.org/html/2609.18047#S3.Thmproposition2)coincides with the scatter\-add gradient of the lookup implementation\.

###### Proof\.

SinceYt​j=∑kXt​k​Ek​jY\_\{tj\}=\\sum\_\{k\}X\_\{tk\}E\_\{kj\}, the chain rule gives

∂ℒ∂Ek​j=∑t=1n∂ℒ∂Yt​j​∂Yt​j∂Ek​j=∑t=1nGt​j​Xt​k=\(𝐗⊤​𝐆\)k​j\.\\frac\{\\partial\\mathcal\{L\}\}\{\\partial E\_\{kj\}\}=\\sum\_\{t=1\}^\{n\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial Y\_\{tj\}\}\\frac\{\\partial Y\_\{tj\}\}\{\\partial E\_\{kj\}\}=\\sum\_\{t=1\}^\{n\}G\_\{tj\}X\_\{tk\}=\(\\mathbf\{X\}^\{\\top\}\\mathbf\{G\}\)\_\{kj\}\.BecauseXt​k=δit​kX\_\{tk\}=\\delta\_\{i\_\{t\}k\}, the sum overttretains exactly the indices withit=ki\_\{t\}=k, so rowkkof𝐗⊤​𝐆\\mathbf\{X\}^\{\\top\}\\mathbf\{G\}is∑t:it=k𝐆t\\sum\_\{t:i\_\{t\}=k\}\\mathbf\{G\}\_\{t\}\. The scatter\-add produces the same row by construction: it deposits𝐆t\\mathbf\{G\}\_\{t\}into rowiti\_\{t\}for eachtt, and leaves the remaining rows at zero\. ∎

Under Corollary[3\.2](https://arxiv.org/html/2609.18047#S3.Thmcorollary2), the gradient with respect to𝐖=𝐄⊤\\mathbf\{W\}=\\mathbf\{E\}^\{\\top\}is\(𝐗⊤​𝐆\)⊤=𝐆⊤​𝐗\(\\mathbf\{X\}^\{\\top\}\\mathbf\{G\}\)^\{\\top\}=\\mathbf\{G\}^\{\\top\}\\mathbf\{X\}; the identification is a transpose, and the row\-selection structure is unchanged\. Training moves only the columns of the geometry, indexed by tokens observed in the batch, so any property of the trained geometry is a property of the training distribution and the objective\. The linear\-layer backward materializes𝐗⊤​𝐆∈ℝV×d\\mathbf\{X\}^\{\\top\}\\mathbf\{G\}\\in\\mathbb\{R\}^\{V\\times d\}densely, with, at most,nnnonzero rows, whereas the scatter\-add sees only those rows; the free carrier is a mathematical object that implementations never materialize; Table[1](https://arxiv.org/html/2609.18047#S3.T1)shows\.

Table 1:Cost comparison for a batch ofnntokens\.
### 3\.4Identifier bias and linearity

Now,\[[27](https://arxiv.org/html/2609.18047#bib.bib3)\]states the identity for a bias\-free layer\. The restriction concerns the identification of parameters, and leaves the function class untouched\.

Letb∈ℝdb\\in\\mathbb\{R\}^\{d\}, and considerx↦x⊤​𝐄\+b⊤x\\mapsto x^\{\\top\}\\mathbf\{E\}\+b^\{\\top\}\. On the input𝐞i\\mathbf\{e\}\_\{i\}, this equals𝐄i\+b⊤\\mathbf\{E\}\_\{i\}\+b^\{\\top\}, which is rowiiof𝐄′=𝐄\+𝟏​b⊤\\mathbf\{E\}^\{\\prime\}=\\mathbf\{E\}\+\\mathbf\{1\}b^\{\\top\}with𝟏∈ℝV\\mathbf\{1\}\\in\\mathbb\{R\}^\{V\}the all\-ones vector\. A linear layer with bias, restricted to one\-hot inputs, is, again, a lookup, with table𝐄′\\mathbf\{E\}^\{\\prime\}; the set of functions on one\-hot inputs realized with bias equals the set realized without, and the parameterization with bias has add\-dimensional redundancy, since\(𝐄,b\)\(\\mathbf\{E\},b\)and\(𝐄\+𝟏​c⊤,b−c\)\(\\mathbf\{E\}\+\\mathbf\{1\}c^\{\\top\},b\-c\)realize the same function for everyc∈ℝdc\\in\\mathbb\{R\}^\{d\}\. We see, in Section[5\.3](https://arxiv.org/html/2609.18047#S5.SS3), the same redundancy, only from the side of the readout instead, where it becomes the all\-ones row of the rank criterion\.

L𝐄L\_\{\\mathbf\{E\}\}is linear onℝV\\mathbb\{R\}^\{V\}by construction\. The compositei↦ℓ𝐄​\(i\)i\\mapsto\\ell\_\{\\mathbf\{E\}\}\(i\)on the integers admits a linear extension only when𝐄i=i​c\\mathbf\{E\}\_\{i\}=icfor a fixedc∈ℝ1×dc\\in\\mathbb\{R\}^\{1\\times d\}and allii, which fails for generic𝐄\\mathbf\{E\}, as any𝐄\\mathbf\{E\}with𝐄2≠2​𝐄1\\mathbf\{E\}\_\{2\}\\neq 2\\mathbf\{E\}\_\{1\}shows; and the encodingi↦𝐞ii\\mapsto\\mathbf\{e\}\_\{i\}is, itself, incompatible with integer addition, since𝐞1\+𝐞2\\mathbf\{e\}\_\{1\}\+\\mathbf\{e\}\_\{2\}lies outside the set of one\-hot vectors\. Token identifiers are labels; in the vector logic, this is the statement that the free carrier encodes entities as atoms: every relation among is via the geometry; we see in Section[5](https://arxiv.org/html/2609.18047#S5)which relations a geometry must preserve\.

## 4Regimes

The extensional theorem is proved for the free carrier\. We see in Section[3](https://arxiv.org/html/2609.18047#S3)that the free carrier is the input to every embedding matrix; any other injection of𝒟e\\mathcal\{D\}\_\{e\}into a real vector space is a candidate carrier, and which of them admit the lifts that Theorem[2\.1](https://arxiv.org/html/2609.18047#S2.Thmtheorem1)guarantees for the free one is decided by the geometry matrix of Definition[4\.1](https://arxiv.org/html/2609.18047#S4.Thmdefinition1)\.

###### Definition 4\.1\(geometry matrix\)\.

Leth:𝒟e→ℝdh:\\mathcal\{D\}\_\{e\}\\to\\mathbb\{R\}^\{d\}be any map; its geometry matrix is𝐇∈ℝd×V\\mathbf\{H\}\\in\\mathbb\{R\}^\{d\\times V\}, with columns𝐇𝐞i=h⁡\(di\)\\mathbf\{H\}\\mathbf\{e\}\_\{i\}=h\(d\_\{i\}\), such thath=𝐇∘heh=\\mathbf\{H\}\\circ h\_\{e\}, whereheh\_\{e\}is the free carrier\.

The free carrier itself has𝐇=IV\\mathbf\{H\}=I\_\{V\}; a trained embedding has𝐇=𝐄⊤\\mathbf\{H\}=\\mathbf\{E\}^\{\\top\}\. We callhhfree when the columns of𝐇\\mathbf\{H\}are linearly independent, and compressed otherwise\.

Every map out of𝒟e\\mathcal\{D\}\_\{e\}into a vector space factors through the free carrier in this way, which is the universal property that justifies the name in the first place; the geometry matrix is the linear map through which it factors\. Injectivity ofhhis, technically, a weaker condition than freeness: distinct columns suffice for injectivity; a compressed geometry may well have distinct columns\.

ddrank⁡\[𝐓;𝟏\]\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]VVno exact liftof whole lexiconcompressed regimeexact⇔\\Leftrightarrowrow⁡\(𝐇\)⊇row⁡\[𝐓;𝟏\]\\operatorname\{row\}\(\\mathbf\{H\}\)\\supseteq\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]free regimeevery function liftstrained embeddings:d≪Vd\\ll VFigure 2:The dimension axis for a geometry𝐇∈ℝd×V\\mathbf\{H\}\\in\\mathbb\{R\}^\{d\\times V\}over a lexicon with truth matrix𝐓\\mathbf\{T\}\. Belowrank⁡\[𝐓;𝟏\]\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\], no geometry carries the whole lexicon; at and aboveVVwith independent columns, every semantic function lifts; between them, exact lift is the row\-space condition of Theorem[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1), and is where the geometries of Section[8\.4](https://arxiv.org/html/2609.18047#S8.SS4)are\. One instance of each regime is drawn in Figure[4](https://arxiv.org/html/2609.18047#S5.F4)\.### 4\.1Free regime

###### Proposition 4\.1\(reparameterization\)\.

Let\{hτ\}\\\{h\_\{\\tau\}\\\}be the free carrier with lifts\{Lf\}\\\{L\_\{f\}\\\}, and, for each type, letTτT\_\{\\tau\}be an injective linear map on𝒮𝒟τ\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{\\tau\}\}\. Sethτ′=Tτ∘hτh^\{\\prime\}\_\{\\tau\}=T\_\{\\tau\}\\circ h\_\{\\tau\}\. Then, for every semantic functionf:𝒟σ→𝒟τf:\\mathcal\{D\}\_\{\\sigma\}\\to\\mathcal\{D\}\_\{\\tau\}withσ\\sigmaprimitive, there is a linearLf′L^\{\\prime\}\_\{f\}, withLf′∘hσ′=hτ′∘fL^\{\\prime\}\_\{f\}\\circ h^\{\\prime\}\_\{\\sigma\}=h^\{\\prime\}\_\{\\tau\}\\circ f, and composition along primitive types is preserved\.

In particular, every free geometry of Definition[4\.1](https://arxiv.org/html/2609.18047#S4.Thmdefinition1)admits lifts for every semantic function\.

###### Proof\.

SinceTσT\_\{\\sigma\}is injective, it has a linear left inverseTσ−T\_\{\\sigma\}^\{\-\}withTσ−​Tσ=IT\_\{\\sigma\}^\{\-\}T\_\{\\sigma\}=Ion𝒮𝒟σ\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{\\sigma\}\}\. DefineLf′=Tτ​Lf​Tσ−L^\{\\prime\}\_\{f\}=T\_\{\\tau\}L\_\{f\}T\_\{\\sigma\}^\{\-\}\. Fora∈𝒟σa\\in\\mathcal\{D\}\_\{\\sigma\},

Lf′​hσ′​\(a\)=Tτ​Lf​Tσ−​Tσ​hσ​\(a\)=Tτ​Lf​hσ​\(a\)=Tτ​hτ​\(f⁡\(a\)\)=hτ′​\(f⁡\(a\)\)\.L^\{\\prime\}\_\{f\}\\,h^\{\\prime\}\_\{\\sigma\}\(a\)=T\_\{\\tau\}L\_\{f\}T\_\{\\sigma\}^\{\-\}T\_\{\\sigma\}h\_\{\\sigma\}\(a\)=T\_\{\\tau\}L\_\{f\}h\_\{\\sigma\}\(a\)=T\_\{\\tau\}h\_\{\\tau\}\(f\(a\)\)=h^\{\\prime\}\_\{\\tau\}\(f\(a\)\)\.For composition, ifg:𝒟τ→𝒟ρg:\\mathcal\{D\}\_\{\\tau\}\\to\\mathcal\{D\}\_\{\\rho\}withτ\\tauprimitive, then

Lg′​Lf′=Tρ​Lg​Tτ−​Tτ​Lf​Tσ−=Tρ​Lg​Lf​Tσ−=Tρ​Lg∘f​Tσ−=Lg∘f′,L^\{\\prime\}\_\{g\}L^\{\\prime\}\_\{f\}=T\_\{\\rho\}L\_\{g\}T\_\{\\tau\}^\{\-\}T\_\{\\tau\}L\_\{f\}T\_\{\\sigma\}^\{\-\}=T\_\{\\rho\}L\_\{g\}L\_\{f\}T\_\{\\sigma\}^\{\-\}=T\_\{\\rho\}L\_\{g\\circ f\}T\_\{\\sigma\}^\{\-\}=L^\{\\prime\}\_\{g\\circ f\},using Theorem[2\.1](https://arxiv.org/html/2609.18047#S2.Thmtheorem1)forLg​Lf=Lg∘fL\_\{g\}L\_\{f\}=L\_\{g\\circ f\}\. Finally, a free geometry𝐇∈ℝd×V\\mathbf\{H\}\\in\\mathbb\{R\}^\{d\\times V\}has linearly independent columns, hence is an injective linear mapℝV→ℝd\\mathbb\{R\}^\{V\}\\to\\mathbb\{R\}^\{d\}; takeTe=𝐇T\_\{e\}=\\mathbf\{H\}andTτ=IT\_\{\\tau\}=Ifor the other types\. ∎

Here, we see that the orthonormal basis of the free carrier is a convenience: any linearly independent family of entity vectors will do\. Every function lifts here, so every measurement on a free geometry passes the homomorphism conditions\. The function\-type spaces are left as they were, so a predicate still lives inHom⁡\(ℝV,ℝ2\)\\operatorname\{Hom\}\(\\mathbb\{R\}^\{V\},\\mathbb\{R\}^\{2\}\), while entities live inℝd\\mathbb\{R\}^\{d\}; the readout picture of Section[5](https://arxiv.org/html/2609.18047#S5), in which a predicate over a geometry𝐇\\mathbf\{H\}is a linear mapLP:ℝd→ℝ2L\_\{P\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{2\}, is the specializationLP=𝐌P​𝐇−L\_\{P\}=\\mathbf\{M\}\_\{P\}\\mathbf\{H\}^\{\-\}, with𝐇−\\mathbf\{H\}^\{\-\}a left inverse of𝐇\\mathbf\{H\}\.

### 4\.2Compressed regime

A trained embedding matrix has𝐄⊤∈ℝd×V\\mathbf\{E\}^\{\\top\}\\in\\mathbb\{R\}^\{d\\times V\}withddin the hundreds \(or low thousands\), andVVin the tens of thousands, sorank⁡𝐄⊤≤d<V\\operatorname\{rank\}\\mathbf\{E\}^\{\\top\}\\leq d<V, and the geometry is compressed by a wide margin\. Proposition[4\.1](https://arxiv.org/html/2609.18047#S4.Thmproposition1)requires injectiveTT, and, therefore, leaves it be\. A criterion onker⁡𝐇\\ker\\mathbf\{H\}replaces the proposition, which is the space of linear dependences among the entity vectors: a vector𝐜∈ℝV\\mathbf\{c\}\\in\\mathbb\{R\}^\{V\}lies inker⁡𝐇\\ker\\mathbf\{H\}exactly when∑ici​h​\(di\)=0\\sum\_\{i\}c\_\{i\}\\,h\(d\_\{i\}\)=0\. We must show that lifts exist exactly when these dependences lie in the kernel of the lexicon’s augmented truth matrix\.

## 5Rank criterion for monadic predicates

Fix a geometry𝐇∈ℝd×V\\mathbf\{H\}\\in\\mathbb\{R\}^\{d\\times V\}and a finite lexicon𝒫⊆𝒟⟨e,t⟩\\mathcal\{P\}\\subseteq\\mathcal\{D\}\_\{\\langle e,t\\rangle\}of monadic predicates\. ForP∈𝒫P\\in\\mathcal\{P\}, write𝐭P∈\{0,1\}1×V\\mathbf\{t\}\_\{P\}\\in\\\{0,1\\\}^\{1\\times V\}for its truth row,\(𝐭P\)i=P⁡\(di\)\(\\mathbf\{t\}\_\{P\}\)\_\{i\}=P\(d\_\{i\}\), and let𝐓∈\{0,1\}\|𝒫\|×V\\mathbf\{T\}\\in\\\{0,1\\\}^\{\|\\mathcal\{P\}\|\\times V\}be the truth matrix with rows𝐭P\\mathbf\{t\}\_\{P\}\. Write\[𝐓;𝟏\]\[\\mathbf\{T\};\\mathbf\{1\}\]for𝐓\\mathbf\{T\}augmented by the all\-ones row\. Row spacesrow⁡\(⋅\)\\operatorname\{row\}\(\\cdot\)are subspaces ofℝ1×V\\mathbb\{R\}^\{1\\times V\}\.

### 5\.1Exact lifts

An exact lift ofPPover𝐇\\mathbf\{H\}is a linear mapLP:ℝd→ℝ2L\_\{P\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{2\}withLP​h​\(di\)=ht​\(P⁡\(di\)\)L\_\{P\}\\,h\(d\_\{i\}\)=h\_\{t\}\(P\(d\_\{i\}\)\)for everyii, the truth basis𝐛1,𝐛0\\mathbf\{b\}\_\{1\},\\mathbf\{b\}\_\{0\}being shared across the lexicon, such that the connective matrices of Section[2](https://arxiv.org/html/2609.18047#S2)act on outputs unchanged\. In matrix form,LP∈ℝ2×dL\_\{P\}\\in\\mathbb\{R\}^\{2\\times d\}, and the condition is

LP​𝐇=𝐌P,L\_\{P\}\\,\\mathbf\{H\}=\\mathbf\{M\}\_\{P\},\(2\)with𝐌P\\mathbf\{M\}\_\{P\}the predicate matrix of the free carrier\. Over the free carrier,𝐇=IV\\mathbf\{H\}=I\_\{V\}andLP=𝐌PL\_\{P\}=\\mathbf\{M\}\_\{P\}solves \([2](https://arxiv.org/html/2609.18047#S5.E2)\); over a compressed geometry, \([2](https://arxiv.org/html/2609.18047#S5.E2)\) is a linear system in the unknownLPL\_\{P\}that may well fail to be solvable\.

𝒟e\{\\lx@inpgf@ignorespaces\\mathcal\{D\}\_\{e\}\}ℝV\{\\lx@inpgf@ignorespaces\\mathbb\{R\}^\{V\}\}ℝd\{\\lx@inpgf@ignorespaces\\mathbb\{R\}^\{d\}\}\{0,1\}\{\\lx@inpgf@ignorespaces\\\{0,1\\\}\}ℝ2\{\\lx@inpgf@ignorespaces\\mathbb\{R\}^\{2\}\}he\\scriptstyle\{\\lx@inpgf@ignorespaces h\_\{e\}\}P\\scriptstyle\{\\lx@inpgf@ignorespaces P\}𝐇\\scriptstyle\{\\lx@inpgf@ignorespaces\\mathbf\{H\}\}𝐌P\\scriptstyle\{\\lx@inpgf@ignorespaces\\mathbf\{M\}\_\{P\}\}LP\\scriptstyle\{\\lx@inpgf@ignorespaces L\_\{P\}\}ht\\scriptstyle\{\\lx@inpgf@ignorespaces h\_\{t\}\}Figure 3:The left square commutes for every predicate by Theorem[2\.1](https://arxiv.org/html/2609.18047#S2.Thmtheorem1), with𝐌P=𝐛1​𝐭P\+𝐛0​\(𝟏−𝐭P\)\\mathbf\{M\}\_\{P\}=\\mathbf\{b\}\_\{1\}\\mathbf\{t\}\_\{P\}\+\\mathbf\{b\}\_\{0\}\(\\mathbf\{1\}\-\\mathbf\{t\}\_\{P\}\)the free readout; the dashed map exists when𝐌P\\mathbf\{M\}\_\{P\}factors through𝐇\\mathbf\{H\}, which is condition \([2](https://arxiv.org/html/2609.18047#S5.E2)\)\. When𝐇\\mathbf\{H\}is free \(injective\), the factorization isLP=𝐌P​𝐇−L\_\{P\}=\\mathbf\{M\}\_\{P\}\\mathbf\{H\}^\{\-\}; when𝐇\\mathbf\{H\}is compressed, it exists only ifker⁡𝐇⊆ker⁡𝐌P\\ker\\mathbf\{H\}\\subseteq\\ker\\mathbf\{M\}\_\{P\}, which is a rank condition as Theorem[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1)\.###### Lemma 5\.1\.

For everyPP,row⁡\(𝐌P\)=span⁡\{𝐭P,𝟏\}\\operatorname\{row\}\(\\mathbf\{M\}\_\{P\}\)=\\operatorname\{span\}\\\{\\mathbf\{t\}\_\{P\},\\mathbf\{1\}\\\}\.

###### Proof\.

The rows of𝐌P\\mathbf\{M\}\_\{P\}are𝐭P\\mathbf\{t\}\_\{P\}and𝟏−𝐭P\\mathbf\{1\}\-\\mathbf\{t\}\_\{P\}, since columniiis𝐛1\\mathbf\{b\}\_\{1\}whenP⁡\(di\)=1P\(d\_\{i\}\)=1and𝐛0\\mathbf\{b\}\_\{0\}otherwise\. The spans of\{𝐭P,𝟏−𝐭P\}\\\{\\mathbf\{t\}\_\{P\},\\mathbf\{1\}\-\\mathbf\{t\}\_\{P\}\\\}and\{𝐭P,𝟏\}\\\{\\mathbf\{t\}\_\{P\},\\mathbf\{1\}\\\}coincide\. ∎

###### Theorem 5\.1\(rank criterion\)\.

The following are equivalent\.

1. 1\.EveryP∈𝒫P\\in\\mathcal\{P\}has an exact lift over𝐇\\mathbf\{H\}\.
2. 2\.row⁡\(𝐇\)⊇row⁡\[𝐓;𝟏\]\\operatorname\{row\}\(\\mathbf\{H\}\)\\supseteq\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]\.
3. 3\.ker⁡𝐇⊆ker⁡\[𝐓;𝟏\]\\ker\\mathbf\{H\}\\subseteq\\ker\[\\mathbf\{T\};\\mathbf\{1\}\]; that is, every linear dependence∑ici​h​\(di\)=0\\sum\_\{i\}c\_\{i\}\\,h\(d\_\{i\}\)=0among the entity vectors satisfies∑ici​P​\(di\)=0\\sum\_\{i\}c\_\{i\}P\(d\_\{i\}\)=0for everyP∈𝒫P\\in\\mathcal\{P\}and∑ici=0\\sum\_\{i\}c\_\{i\}=0\.

###### Proof sketch\.

The equationL​𝐇=𝐌L\\mathbf\{H\}=\\mathbf\{M\}has a solutionLLif and only if every row of𝐌\\mathbf\{M\}is a linear combination of the rows of𝐇\\mathbf\{H\}, that is,row⁡\(𝐌\)⊆row⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{M\}\)\\subseteq\\operatorname\{row\}\(\\mathbf\{H\}\)\. By Lemma[5\.1](https://arxiv.org/html/2609.18047#S5.Thmlemma1), \([2](https://arxiv.org/html/2609.18047#S5.E2)\) is solvable forPPif and only if𝐭P,𝟏∈row⁡\(𝐇\)\\mathbf\{t\}\_\{P\},\\mathbf\{1\}\\in\\operatorname\{row\}\(\\mathbf\{H\}\), and solvability for all of𝒫\\mathcal\{P\}isrow⁡\[𝐓;𝟏\]⊆row⁡\(𝐇\)\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]\\subseteq\\operatorname\{row\}\(\\mathbf\{H\}\), which is the equivalence of \(1\) and \(2\)\.

For \(2\) and \(3\), the row space and kernel of a matrix are orthogonal complements inℝV\\mathbb\{R\}^\{V\}, sorow⁡\(𝐇\)⊇row⁡\[𝐓;𝟏\]\\operatorname\{row\}\(\\mathbf\{H\}\)\\supseteq\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]if and only ifker⁡𝐇⊆ker⁡\[𝐓;𝟏\]\\ker\\mathbf\{H\}\\subseteq\\ker\[\\mathbf\{T\};\\mathbf\{1\}\]; and𝐜∈ker⁡\[𝐓;𝟏\]\\mathbf\{c\}\\in\\ker\[\\mathbf\{T\};\\mathbf\{1\}\]to𝐭P​𝐜=0\\mathbf\{t\}\_\{P\}\\mathbf\{c\}=0for eachPPand𝟏​𝐜=0\\mathbf\{1\}\\mathbf\{c\}=0\.

From \(2\) to \(1\) also admits a direct construction, and the direction from \(1\) to \(3\) a direct computation\. Given \(2\), choose row vectors𝐰P,𝐰P′∈ℝ1×d\\mathbf\{w\}\_\{P\},\\mathbf\{w\}^\{\\prime\}\_\{P\}\\in\\mathbb\{R\}^\{1\\times d\}with𝐰P​𝐇=𝐭P\\mathbf\{w\}\_\{P\}\\mathbf\{H\}=\\mathbf\{t\}\_\{P\}and𝐰P′​𝐇=𝟏−𝐭P\\mathbf\{w\}^\{\\prime\}\_\{P\}\\mathbf\{H\}=\\mathbf\{1\}\-\\mathbf\{t\}\_\{P\}, and setLP=𝐛1​𝐰P\+𝐛0​𝐰P′L\_\{P\}=\\mathbf\{b\}\_\{1\}\\mathbf\{w\}\_\{P\}\+\\mathbf\{b\}\_\{0\}\\mathbf\{w\}^\{\\prime\}\_\{P\}, so thatLP​x=\(𝐰P​x\)​𝐛1\+\(𝐰P′​x\)​𝐛0L\_\{P\}x=\(\\mathbf\{w\}\_\{P\}x\)\\,\\mathbf\{b\}\_\{1\}\+\(\\mathbf\{w\}^\{\\prime\}\_\{P\}x\)\\,\\mathbf\{b\}\_\{0\}\. ThenLP​h​\(di\)=P⁡\(di\)​𝐛1\+\(1−P⁡\(di\)\)​𝐛0=ht​\(P⁡\(di\)\)L\_\{P\}h\(d\_\{i\}\)=P\(d\_\{i\}\)\\,\\mathbf\{b\}\_\{1\}\+\(1\-P\(d\_\{i\}\)\)\\,\\mathbf\{b\}\_\{0\}=h\_\{t\}\(P\(d\_\{i\}\)\)\. Conversely, given a liftLPL\_\{P\}and𝐜∈ker⁡𝐇\\mathbf\{c\}\\in\\ker\\mathbf\{H\}, applyLPL\_\{P\}to∑ici​h​\(di\)=0\\sum\_\{i\}c\_\{i\}\\,h\(d\_\{i\}\)=0and expand:∑ici​\(P⁡\(di\)​𝐛1\+\(1−P⁡\(di\)\)​𝐛0\)=0\\sum\_\{i\}c\_\{i\}\\bigl\(P\(d\_\{i\}\)\\,\\mathbf\{b\}\_\{1\}\+\(1\-P\(d\_\{i\}\)\)\\,\\mathbf\{b\}\_\{0\}\\bigr\)=0, and independence of𝐛1,𝐛0\\mathbf\{b\}\_\{1\},\\mathbf\{b\}\_\{0\}forces∑ici​P​\(di\)=0\\sum\_\{i\}c\_\{i\}P\(d\_\{i\}\)=0and∑ici​\(1−P⁡\(di\)\)=0\\sum\_\{i\}c\_\{i\}\(1\-P\(d\_\{i\}\)\)=0, whose sum is∑ici=0\\sum\_\{i\}c\_\{i\}=0\. ∎

### 5\.2Corollaries

Here we take a brief tour of some nice properties of the lifts\.

###### Corollary 5\.1\(minimal dimension\)\.

The leastddfor which some𝐇∈ℝd×V\\mathbf\{H\}\\in\\mathbb\{R\}^\{d\\times V\}carries exact lifts of all of𝒫\\mathcal\{P\}isrank⁡\[𝐓;𝟏\]\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\], attained by any𝐇\\mathbf\{H\}, whose rows form a basis ofrow⁡\[𝐓;𝟏\]\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]\.

###### Proof sketch\.

By \(2\) in Theorem[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1),row⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\)must contain a subspace of dimensionrank⁡\[𝐓;𝟏\]\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\], sod≥rank⁡𝐇≥rank⁡\[𝐓;𝟏\]d\\geq\\operatorname\{rank\}\\mathbf\{H\}\\geq\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]; taking the rows of𝐇\\mathbf\{H\}to be a basis ofrow⁡\[𝐓;𝟏\]\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]gives equality\. ∎

###### Corollary 5\.2\(closure forces the free regime\)\.

If𝒫=𝒟⟨e,t⟩\\mathcal\{P\}=\\mathcal\{D\}\_\{\\langle e,t\\rangle\}, then exact lifts of all of𝒫\\mathcal\{P\}exist only over free geometries; in particular,d≥Vd\\geq V\.

###### Proof sketch\.

𝒟⟨e,t⟩\\mathcal\{D\}\_\{\\langle e,t\\rangle\}contains the singleton predicatesPiP\_\{i\}with𝐭Pi=𝐞i⊤\\mathbf\{t\}\_\{P\_\{i\}\}=\\mathbf\{e\}\_\{i\}^\{\\top\}, sorow⁡\[𝐓;𝟏\]=ℝ1×V\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]=\\mathbb\{R\}^\{1\\times V\}, and \(2\) in[Theorem5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1)forcesrank⁡𝐇=V\\operatorname\{rank\}\\mathbf\{H\}=V, which is linear independence of theVVcolumns\. ∎

The extensional theorem embeds every domain𝒟⟨e,t⟩\\mathcal\{D\}\_\{\\langle e,t\\rangle\}in full; Corollary[5\.2](https://arxiv.org/html/2609.18047#S5.Thmcorollary2)requires a free geometry for it\. Linear independence of the entity vectors is, thereby, derived from predicate closure, and the one\-hot construction is one coordinate choice among the free geometries, which are its injective linear images by Definition[4\.1](https://arxiv.org/html/2609.18047#S4.Thmdefinition1), each admitting every lift by Proposition[4\.1](https://arxiv.org/html/2609.18047#S4.Thmproposition1); the equal pairwise distances of the one\-hot basis are a feature of that choice alone\. The theorem embeds an object, the full function space𝒟⟨e,t⟩\\mathcal\{D\}\_\{\\langle e,t\\rangle\}, that lies beyond what any learned system represents; the results that follow concern the compressed regime, where closure fails by construction, and the conditions have content\.

###### Corollary 5\.3\(compressibility and the lexicon\)\.

A dependence𝐜∈ℝV\\mathbf\{c\}\\in\\mathbb\{R\}^\{V\}among entity vectors is admissible, in the sense that some geometry with𝐜∈ker⁡𝐇\\mathbf\{c\}\\in\\ker\\mathbf\{H\}carries exact lifts of𝒫\\mathcal\{P\}, if and only if𝐜⟂row⁡\[𝐓;𝟏\]\\mathbf\{c\}\\perp\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]; the admissible dependences form the orthogonal complement of the augmented row space of the lexicon, of dimensionV−rank⁡\[𝐓;𝟏\]V\-\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]\.

###### Proof\.

Immediate from \(3\) in Theorem[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1), taking𝐇\\mathbf\{H\}with row space exactlyrow⁡\[𝐓;𝟏\]\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]for the converse\. ∎

A geometry may well identify entity vectors up to a dependence exactly when every predicate in the lexicon, and the constant predicate, assign that dependence weight zero\. The dimensionddof a learned embedding bounds its rank; whetherddis compatible with exactness depends on the lexicon, which lexicalizes a minuscule fraction of the2V2^\{V\}available predicates, and Corollary[5\.1](https://arxiv.org/html/2609.18047#S5.Thmcorollary1)computes the exchange rate between rank and lexicon\.

###### Corollary 5\.4\(indiscernibility\)\.

Let𝐇\\mathbf\{H\}haverow⁡\(𝐇\)=row⁡\[𝐓;𝟏\]\\operatorname\{row\}\(\\mathbf\{H\}\)=\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]\. Thenh⁡\(di\)=h⁡\(dj\)h\(d\_\{i\}\)=h\(d\_\{j\}\)if and only ifP⁡\(di\)=P⁡\(dj\)P\(d\_\{i\}\)=P\(d\_\{j\}\)for everyP∈𝒫P\\in\\mathcal\{P\}\. Hence, a minimal exact geometry is injective on𝒟e\\mathcal\{D\}\_\{e\}if and only if𝒫\\mathcal\{P\}separates entities\.

###### Proof sketch\.

h⁡\(di\)=h⁡\(dj\)h\(d\_\{i\}\)=h\(d\_\{j\}\)if and only if𝐞i−𝐞j∈ker⁡𝐇=row⁡\[𝐓;𝟏\]⟂\\mathbf\{e\}\_\{i\}\-\\mathbf\{e\}\_\{j\}\\in\\ker\\mathbf\{H\}=\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]^\{\\perp\}, if and only if𝐭P​\(𝐞i−𝐞j\)=0\\mathbf\{t\}\_\{P\}\(\\mathbf\{e\}\_\{i\}\-\\mathbf\{e\}\_\{j\}\)=0for allPP\(the condition𝟏​\(𝐞i−𝐞j\)=0\\mathbf\{1\}\(\\mathbf\{e\}\_\{i\}\-\\mathbf\{e\}\_\{j\}\)=0holding automatically\), which isP⁡\(di\)=P⁡\(dj\)P\(d\_\{i\}\)=P\(d\_\{j\}\)for allPP\. ∎

### 5\.3Affine readout

We see the all\-ones row in Theorem[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1)through the shared truth basis: the second row of𝐌P\\mathbf\{M\}\_\{P\}is𝟏−𝐭P\\mathbf\{1\}\-\\mathbf\{t\}\_\{P\}; a linearLPL\_\{P\}must produce it\. IfLPL\_\{P\}is permitted to be affine of the formLP​x=A​x\+𝐜PL\_\{P\}x=Ax\+\\mathbf\{c\}\_\{P\}, then the requirement changes\.

###### Proposition 5\.1\(affine lifts\)\.

Affine exact lifts of all of𝒫\\mathcal\{P\}over𝐇\\mathbf\{H\}exist if and only ifrow⁡\(𝐓\)⊆row⁡\(𝐇\)\+span⁡\{𝟏\}\\operatorname\{row\}\(\\mathbf\{T\}\)\\subseteq\\operatorname\{row\}\(\\mathbf\{H\}\)\+\\operatorname\{span\}\\\{\\mathbf\{1\}\\\}; equivalently, if and only if linear exact lifts exist over the homogenized geometry𝐇\+=\[𝐇;𝟏\]∈ℝ\(d\+1\)×V\\mathbf\{H\}^\{\+\}=\[\\mathbf\{H\};\\mathbf\{1\}\]\\in\\mathbb\{R\}^\{\(d\+1\)\\times V\}\.

###### Proof sketch\.

An affine map onℝd\\mathbb\{R\}^\{d\}is a linear map onℝd\+1\\mathbb\{R\}^\{d\+1\}, restricted to the affine hyperplane of vectors with last coordinate11, andh\+​\(di\)=\(h⁡\(di\),1\)h^\{\+\}\(d\_\{i\}\)=\(h\(d\_\{i\}\),1\)has geometry matrix𝐇\+\\mathbf\{H\}^\{\+\}; the equivalence of affine lifts over𝐇\\mathbf\{H\}and linear lifts over𝐇\+\\mathbf\{H\}^\{\+\}is this identification\. Theorem[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1), applied to𝐇\+\\mathbf\{H\}^\{\+\}, requiresrow⁡\[𝐓;𝟏\]⊆row⁡\(𝐇\+\)=row⁡\(𝐇\)\+span⁡\{𝟏\}\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]\\subseteq\\operatorname\{row\}\(\\mathbf\{H\}^\{\+\}\)=\\operatorname\{row\}\(\\mathbf\{H\}\)\+\\operatorname\{span\}\\\{\\mathbf\{1\}\\\}, in which the𝟏\\mathbf\{1\}requirement is automatic, leavingrow⁡\(𝐓\)⊆row⁡\(𝐇\)\+span⁡\{𝟏\}\\operatorname\{row\}\(\\mathbf\{T\}\)\\subseteq\\operatorname\{row\}\(\\mathbf\{H\}\)\+\\operatorname\{span\}\\\{\\mathbf\{1\}\\\}\. ∎

###### Corollary 5\.5\(minimal affine dimension\)\.

Write𝐓c=𝐓−𝐭¯​𝟏⊤\\mathbf\{T\}\_\{c\}=\\mathbf\{T\}\-\\bar\{\\mathbf\{t\}\}\\mathbf\{1\}^\{\\top\}for the truth matrix with its row means removed\. Thenrank⁡𝐓c=rank⁡\[𝐓;𝟏\]−1\\operatorname\{rank\}\\mathbf\{T\}\_\{c\}=\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]\-1, and the leastddfor which some𝐇∈ℝd×V\\mathbf\{H\}\\in\\mathbb\{R\}^\{d\\times V\}carries affine exact lifts of all of𝒫\\mathcal\{P\}isrank⁡\[𝐓;𝟏\]−1\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]\-1, attained by any𝐇\\mathbf\{H\}whose rows form a basis ofrow⁡\(𝐓c\)\\operatorname\{row\}\(\\mathbf\{T\}\_\{c\}\)\.

###### Proof\.

Each row of𝐓c\\mathbf\{T\}\_\{c\}sums to zero, sorow⁡\(𝐓c\)⊆𝟏⟂\\operatorname\{row\}\(\\mathbf\{T\}\_\{c\}\)\\subseteq\\mathbf\{1\}^\{\\perp\}and𝟏∉row⁡\(𝐓c\)\\mathbf\{1\}\\notin\\operatorname\{row\}\(\\mathbf\{T\}\_\{c\}\), whilerow⁡\[𝐓;𝟏\]=row⁡\(𝐓c\)\+span⁡\{𝟏\}\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]=\\operatorname\{row\}\(\\mathbf\{T\}\_\{c\}\)\+\\operatorname\{span\}\\\{\\mathbf\{1\}\\\}, giving the rank identity\. By Proposition[5\.1](https://arxiv.org/html/2609.18047#S5.Thmproposition1), affine exact lifts exist if and only ifrow⁡\(𝐓\)⊆row⁡\(𝐇\)\+span⁡\{𝟏\}\\operatorname\{row\}\(\\mathbf\{T\}\)\\subseteq\\operatorname\{row\}\(\\mathbf\{H\}\)\+\\operatorname\{span\}\\\{\\mathbf\{1\}\\\}, equivalentlyrow⁡\(𝐓c\)⊆row⁡\(𝐇\)\+span⁡\{𝟏\}\\operatorname\{row\}\(\\mathbf\{T\}\_\{c\}\)\\subseteq\\operatorname\{row\}\(\\mathbf\{H\}\)\+\\operatorname\{span\}\\\{\\mathbf\{1\}\\\}\. The right side has dimension at mostd\+1d\+1and must containrow⁡\(𝐓c\)\+span⁡\{𝟏\}\\operatorname\{row\}\(\\mathbf\{T\}\_\{c\}\)\+\\operatorname\{span\}\\\{\\mathbf\{1\}\\\}of dimensionrank⁡\[𝐓;𝟏\]\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\], whenced≥rank⁡\[𝐓;𝟏\]−1d\\geq\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]\-1, with equality whenrow⁡\(𝐇\)=row⁡\(𝐓c\)\\operatorname\{row\}\(\\mathbf\{H\}\)=\\operatorname\{row\}\(\\mathbf\{T\}\_\{c\}\)\. ∎

Appending a constant coordinate converts affine readouts into linear ones; the bias supplies the constant row required by the shared truth basis\. Over the free carrier, linear and affine readouts realize the same predicate extensions, since𝟏∈row⁡\(IV\)\\mathbf\{1\}\\in\\operatorname\{row\}\(I\_\{V\}\)\. We retain the linear formulation as primary, and obtain affine readouts by homogenization\. Once either readout returns the truth basis, the Boolean connective matrices act unchanged\.

### 5\.4Connectives and sentences

Once every predicate in𝒫\\mathcal\{P\}lifts exactly, so does every sentence built from atomic predications by truth\-functional connectives, with the same connective matrices as over the free carrier\.

###### Proposition 5\.2\.

Suppose everyP∈𝒫P\\in\\mathcal\{P\}has an exact liftLPL\_\{P\}over𝐇\\mathbf\{H\}\. Then, for every Boolean combinationΦ\\Phiof atomic predicationsP⁡\(di\)P\(d\_\{i\}\), the vector computed by applying𝐌c\\mathbf\{M\}\_\{c\}to tensor products of lifted outputs equalsht​\(⟦Φ⟧\)h\_\{t\}\(\\llbracket\\Phi\\rrbracket\)\.

###### Proof\.

We proceed by induction onΦ\\Phi\. For the simple case:LP​h​\(di\)=ht​\(P⁡\(di\)\)L\_\{P\}h\(d\_\{i\}\)=h\_\{t\}\(P\(d\_\{i\}\)\)by hypothesis\. Inductive case: if the immediate subformulas evaluate toht​\(t1\),…,ht​\(tn\)h\_\{t\}\(t\_\{1\}\),\\dots,h\_\{t\}\(t\_\{n\}\), then𝐌c\(ht\(t1\)⊗ht\(t2\)⊗⋯⊗ht\(tn\)\)=ht\(c\(t1,…,tn\)\)\\mathbf\{M\}\_\{c\}\(h\_\{t\}\(t\_\{1\}\)\\otimes h\_\{t\}\(t\_\{2\}\)\\otimes\\cdots\\otimes h\_\{t\}\(t\_\{n\}\)\)=h\_\{t\}\(c\(t\_\{1\},\\dots,t\_\{n\}\)\)by the definition of𝐌c\\mathbf\{M\}\_\{c\}over the truth space, which is unchanged by compression of the entity space\. ∎

Compression of the entity carrier is, therefore, confined to the leaves of a derivation; the connective level of the vector logic is insensitive to it\. This is a first indication of where the compressed regime places its constraints: on the geometry of atoms and their readouts; everything from the truth space upward is unchanged\. The compound predicates themselves need no lifts of their own, and in general have none: given the𝟏\\mathbf\{1\}row, the exact lexicon is closed under negation, since𝟏−𝐭P∈row⁡\(𝐇\)\\mathbf\{1\}\-\\mathbf\{t\}\_\{P\}\\in\\operatorname\{row\}\(\\mathbf\{H\}\), while𝐭P⊙𝐭Q\\mathbf\{t\}\_\{P\}\\odot\\mathbf\{t\}\_\{Q\}, the truth row ofP∧QP\\wedge Qas a monadic predicate, may lie outsiderow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\), as it does for the two predicates of Example[5\.1](https://arxiv.org/html/2609.18047#S5.Thmexample1)atV=4V=4, where the conjunction is the singleton\{d1\}\\\{d\_\{1\}\\\}that Figure[4](https://arxiv.org/html/2609.18047#S5.F4)excludes\. Proposition[5\.2](https://arxiv.org/html/2609.18047#S5.Thmproposition2)evaluatesP⁡\(di\)∧Q⁡\(di\)P\(d\_\{i\}\)\\wedge Q\(d\_\{i\}\)exactly all the same, because the tensor product of the two readouts is quadratic inh⁡\(di\)h\(d\_\{i\}\), and the connective matrix acts on that product\.

###### Example 5\.1\.

Take𝒟e=\{Daniel,Thomas\}\\mathcal\{D\}\_\{e\}=\\\{\\textsc\{Daniel\},\\textsc\{Thomas\}\\\}, soV=2V=2, and the lexicon𝒫=\{write,published\}\\mathcal\{P\}=\\\{\\mathrm\{write\},\\mathrm\{published\}\\\}with𝐭write=\(1,0\)\\mathbf\{t\}\_\{\\mathrm\{write\}\}=\(1,0\)and𝐭published=\(0,0\)\\mathbf\{t\}\_\{\\mathrm\{published\}\}=\(0,0\), so that at the world of evaluation Daniel writes and neither is published\. Then\[𝐓;𝟏\]\[\\mathbf\{T\};\\mathbf\{1\}\]has rows\(1,0\),\(0,0\),\(1,1\)\(1,0\),\(0,0\),\(1,1\)and rank2=V2=V; by Corollary[5\.1](https://arxiv.org/html/2609.18047#S5.Thmcorollary1), the minimal exact dimension is22, so this lexicon already forces the free regime on two entities\. Extend toV=4V=4, with𝐭write=\(1,1,0,0\)\\mathbf\{t\}\_\{\\mathrm\{write\}\}=\(1,1,0,0\)and𝐭published=\(1,0,1,0\)\\mathbf\{t\}\_\{\\mathrm\{published\}\}=\(1,0,1,0\): nowrank⁡\[𝐓;𝟏\]=3<4\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]=3<4, exact compression tod=3d=3exists, and[Section8\.3](https://arxiv.org/html/2609.18047#S8.SS3)identifies the one admissible dependence\.

height11writeh⁡\(Daniel\)h\(\\textsc\{Daniel\}\)h⁡\(Thomas\)h\(\\textsc\{Thomas\}\)free,V=2V=2

writepublishedh⁡\(d1\)h\(d\_\{1\}\)h⁡\(d2\)h\(d\_\{2\}\)h⁡\(d3\)h\(d\_\{3\}\)h⁡\(d4\)h\(d\_\{4\}\)compressed,V=4V=4,d=3d=3

Figure 4:Both regimes on the lexicon of Example[5\.1](https://arxiv.org/html/2609.18047#S5.Thmexample1), in minimal exact geometries, whose rows are truth rows and𝟏\\mathbf\{1\}, the zero row ofpublished\\mathrm\{published\}atV=2V=2dropped; each coordinate reads off a predicate \(or the constant\); every entity vector reaches the affine slice at height one\. Left:V=2V=2:rank⁡\[𝐓;𝟏\]=2=V\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]=2=V, the lexicon forces the free regime, and the two entity vectors are independent; the sole dependence∑ici​h​\(di\)=0\\sum\_\{i\}c\_\{i\}\\,h\(d\_\{i\}\)=0isc=0c=0\. Right:V=4V=4:rank⁡\[𝐓;𝟏\]=3\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]=3, and the four entity vectors satisfy the single admissible dependenceh⁡\(d1\)−h⁡\(d2\)−h⁡\(d3\)\+h⁡\(d4\)=0h\(d\_\{1\}\)\-h\(d\_\{2\}\)\-h\(d\_\{3\}\)\+h\(d\_\{4\}\)=0, a parallelogram as in Section[8\.3](https://arxiv.org/html/2609.18047#S8.SS3)\. Every truth row of the lexicon annihilates it, so𝐰write=\(1,0,0\)\\mathbf\{w\}\_\{\\mathrm\{write\}\}=\(1,0,0\)and𝐰published=\(0,1,0\)\\mathbf\{w\}\_\{\\mathrm\{published\}\}=\(0,1,0\)are exact; the singleton\{d1\}\\\{d\_\{1\}\\\}assigns it weight11, and \(3\) of Theorem[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1)excludes its exact lift over this geometry\.
### 5\.5Determiners

Once the atomic lexicon lifts exactly, quantification over entities computes in the compressed space as well\. A determiner meaning that is conservative, extension\-invariant, and isomorphism\-invariant depends on its restrictorPPand scopeQQonly through the pair\(\|P∖Q\|,\|P∩Q\|\)\(\|P\\setminus Q\|,\|P\\cap Q\|\)\[[30](https://arxiv.org/html/2609.18047#bib.bib27),[11](https://arxiv.org/html/2609.18047#bib.bib28)\], the tree of numbers, so thateveryis\|P∖Q\|=0\|P\\setminus Q\|=0,someis\|P∩Q\|≥1\|P\\cap Q\|\\geq 1,mostis\|P∩Q\|\>\|P∖Q\|\|P\\cap Q\|\>\|P\\setminus Q\|, andat leastnnis\|P∩Q\|≥n\|P\\cap Q\|\\geq n\.

###### Proposition 5\.3\(compressed determiners\)\.

Suppose everyP∈𝒫P\\in\\mathcal\{P\}has an exact lift over𝐇\\mathbf\{H\}, with readouts𝐰P​𝐇=𝐭P\\mathbf\{w\}\_\{P\}\\mathbf\{H\}=\\mathbf\{t\}\_\{P\}and𝐰𝟏​𝐇=𝟏\\mathbf\{w\}\_\{\\mathbf\{1\}\}\\mathbf\{H\}=\\mathbf\{1\}, and let𝐆=𝐇𝐇⊤∈ℝd×d\\mathbf\{G\}=\\mathbf\{H\}\\mathbf\{H\}^\{\\top\}\\in\\mathbb\{R\}^\{d\\times d\}\. Then, for allP,Q∈𝒫P,Q\\in\\mathcal\{P\},

\|P∩Q\|=𝐰P​𝐆​𝐰Q⊤,\|P∖Q\|=𝐰P​𝐆​\(𝐰𝟏−𝐰Q\)⊤,\|P\\cap Q\|=\\mathbf\{w\}\_\{P\}\\,\\mathbf\{G\}\\,\\mathbf\{w\}\_\{Q\}^\{\\top\},\\qquad\|P\\setminus Q\|=\\mathbf\{w\}\_\{P\}\\,\\mathbf\{G\}\\,\(\\mathbf\{w\}\_\{\\mathbf\{1\}\}\-\\mathbf\{w\}\_\{Q\}\)^\{\\top\},so every conservative, extension\-invariant, isomorphism\-invariant determiner evaluates onPPandQQby twod×dd\\times dbilinear accumulations followed by its decision on the tree of numbers\.

###### Proof\.

\|P∩Q\|=𝐭P​𝐭Q⊤=𝐰P​𝐇𝐇⊤​𝐰Q⊤\|P\\cap Q\|=\\mathbf\{t\}\_\{P\}\\mathbf\{t\}\_\{Q\}^\{\\top\}=\\mathbf\{w\}\_\{P\}\\mathbf\{H\}\\mathbf\{H\}^\{\\top\}\\mathbf\{w\}\_\{Q\}^\{\\top\}and\|P∖Q\|=𝐭P​\(𝟏−𝐭Q\)⊤=𝐰P​𝐇𝐇⊤​\(𝐰𝟏−𝐰Q\)⊤\|P\\setminus Q\|=\\mathbf\{t\}\_\{P\}\(\\mathbf\{1\}\-\\mathbf\{t\}\_\{Q\}\)^\{\\top\}=\\mathbf\{w\}\_\{P\}\\mathbf\{H\}\\mathbf\{H\}^\{\\top\}\(\\mathbf\{w\}\_\{\\mathbf\{1\}\}\-\\mathbf\{w\}\_\{Q\}\)^\{\\top\}; the rest is the cited classification\. ∎

The Gram matrix𝐆\\mathbf\{G\}is formed once from the geometry and shared across the lexicon\. The quantified sentence is a decision on bilinear forms in the readouts and, by the descent theorem of\[[26](https://arxiv.org/html/2609.18047#bib.bib2)\], a linear readout of no single vector\. The two accumulations have the form of the modal accumulation \([1](https://arxiv.org/html/2609.18047#S2.E1)\) with the restrictor in place of the accessible set, which is the sense in which modals quantify over worlds\[[13](https://arxiv.org/html/2609.18047#bib.bib29)\]; the compressed regime places one constraint on both, and leaves the connective and quantifier levels as they are\. Beyond the exact regime, the counts inherit the defect of Section[8](https://arxiv.org/html/2609.18047#S8)through𝐭P−𝐰P​𝐇\\mathbf\{t\}\_\{P\}\-\\mathbf\{w\}\_\{P\}\\mathbf\{H\}, and the decisions then act on approximate counts, which we leave to future work\.

## 6Relations

A binary relationR⊆𝒟e×𝒟eR\\subseteq\\mathcal\{D\}\_\{e\}\\times\\mathcal\{D\}\_\{e\}is, at type⟨e,⟨e,t⟩⟩\\langle e,\\langle e,t\\rangle\\rangle, a curried function, and theHom\\operatorname\{Hom\}construction of\[[26](https://arxiv.org/html/2609.18047#bib.bib2)\]lifts it to a linear map𝒮𝒟e→Hom⁡\(𝒮𝒟e,𝒮𝒟t\)\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{e\}\}\\to\\operatorname\{Hom\}\(\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{e\}\},\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{t\}\}\), equivalently, a bilinear map𝒮𝒟e×𝒮𝒟e→𝒮𝒟t\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{e\}\}\\times\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{e\}\}\\to\\mathcal\{S\}\_\{\\mathcal\{D\}\_\{t\}\}, equivalently, a linear map on the tensor square\. Over a compressed geometry, the corresponding object is a linearLR:ℝd⊗ℝd→ℝ2L\_\{R\}:\\mathbb\{R\}^\{d\}\\otimes\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{2\}with

LR​\(h⁡\(di\)⊗h⁡\(dj\)\)=ht​\(R⁡\(di,dj\)\)for all​i,j\.L\_\{R\}\\,\\bigl\(h\(d\_\{i\}\)\\otimes h\(d\_\{j\}\)\\bigr\)=h\_\{t\}\(R\(d\_\{i\},d\_\{j\}\)\)\\quad\\text\{for all \}i,j\.\(3\)Write𝐓R∈\{0,1\}V×V\\mathbf\{T\}\_\{R\}\\in\\\{0,1\\\}^\{V\\times V\}for the truth matrix ofRR,\(𝐓R\)i​j=R⁡\(di,dj\)\(\\mathbf\{T\}\_\{R\}\)\_\{ij\}=R\(d\_\{i\},d\_\{j\}\)\.

### 6\.1Bilinear criterion

###### Theorem 6\.1\(relations\)\.

An exact lift \([3](https://arxiv.org/html/2609.18047#S6.E3)\) ofRRover𝐇\\mathbf\{H\}exists if and only if𝟏∈row⁡\(𝐇\)\\mathbf\{1\}\\in\\operatorname\{row\}\(\\mathbf\{H\}\), and there is𝐀R∈ℝd×d\\mathbf\{A\}\_\{R\}\\in\\mathbb\{R\}^\{d\\times d\}with

𝐓R=𝐇⊤​𝐀R​𝐇\.\\mathbf\{T\}\_\{R\}=\\mathbf\{H\}^\{\\top\}\\mathbf\{A\}\_\{R\}\\,\\mathbf\{H\}\.

###### Proof sketch\.

Bookkeeping is an irritant and laborious here\. We proceed cautiously\.

The tensor squareh⁡\(di\)⊗h⁡\(dj\)h\(d\_\{i\}\)\\otimes h\(d\_\{j\}\)is column\(i,j\)\(i,j\)of𝐇⊗𝐇∈ℝd2×V2\\mathbf\{H\}\\otimes\\mathbf\{H\}\\in\\mathbb\{R\}^\{d^\{2\}\\times V^\{2\}\}, so \([3](https://arxiv.org/html/2609.18047#S6.E3)\) isLR​\(𝐇⊗𝐇\)=𝐌RL\_\{R\}\(\\mathbf\{H\}\\otimes\\mathbf\{H\}\)=\\mathbf\{M\}\_\{R\}, with𝐌R∈ℝ2×V2\\mathbf\{M\}\_\{R\}\\in\\mathbb\{R\}^\{2\\times V^\{2\}\}having columnsht​\(R⁡\(di,dj\)\)h\_\{t\}\(R\(d\_\{i\},d\_\{j\}\)\)in the same ordering of pairs\(i,j\)\(i,j\)as the Kronecker product; as in Lemma[5\.1](https://arxiv.org/html/2609.18047#S5.Thmlemma1),row⁡\(𝐌R\)=span⁡\{vec⁡\(𝐓R\)⊤,𝟏V2⊤\}\\operatorname\{row\}\(\\mathbf\{M\}\_\{R\}\)=\\operatorname\{span\}\\\{\\operatorname\{vec\}\(\\mathbf\{T\}\_\{R\}\)^\{\\top\},\\mathbf\{1\}\_\{V^\{2\}\}^\{\\top\}\\\}withvec\\operatorname\{vec\}taken in that ordering, and solvability isrow⁡\(𝐌R\)⊆row⁡\(𝐇⊗𝐇\)\\operatorname\{row\}\(\\mathbf\{M\}\_\{R\}\)\\subseteq\\operatorname\{row\}\(\\mathbf\{H\}\\otimes\\mathbf\{H\}\)\. The rows of𝐇⊗𝐇\\mathbf\{H\}\\otimes\\mathbf\{H\}are the Kronecker products𝐇k⊗𝐇l\\mathbf\{H\}\_\{k\}\\otimes\\mathbf\{H\}\_\{l\}of pairs of rows of𝐇\\mathbf\{H\}, sorow⁡\(𝐇⊗𝐇\)=row⁡\(𝐇\)⊗row⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\\otimes\\mathbf\{H\}\)=\\operatorname\{row\}\(\\mathbf\{H\}\)\\otimes\\operatorname\{row\}\(\\mathbf\{H\}\); under the identification ofℝ1×V⊗ℝ1×V\\mathbb\{R\}^\{1\\times V\}\\otimes\\mathbb\{R\}^\{1\\times V\}withV×VV\\times Vmatrices,row⁡\(𝐇\)⊗row⁡\(𝐇\)=\{𝐇⊤​𝐀𝐇:𝐀∈ℝd×d\}\\operatorname\{row\}\(\\mathbf\{H\}\)\\otimes\\operatorname\{row\}\(\\mathbf\{H\}\)=\\\{\\mathbf\{H\}^\{\\top\}\\mathbf\{A\}\\mathbf\{H\}:\\mathbf\{A\}\\in\\mathbb\{R\}^\{d\\times d\}\\\}, sincerow⁡\(𝐇\)=\{𝐚⊤​𝐇\}\\operatorname\{row\}\(\\mathbf\{H\}\)=\\\{\\mathbf\{a\}^\{\\top\}\\mathbf\{H\}\\\}and\(𝐚⊤​𝐇\)⊤​\(𝐚′⁣⊤​𝐇\)=𝐇⊤​\(𝐚𝐚′⁣⊤\)​𝐇\(\\mathbf\{a\}^\{\\top\}\\mathbf\{H\}\)^\{\\top\}\(\\mathbf\{a\}^\{\\prime\\top\}\\mathbf\{H\}\)=\\mathbf\{H\}^\{\\top\}\(\\mathbf\{a\}\\mathbf\{a\}^\{\\prime\\top\}\)\\mathbf\{H\}, with sums of such rank\-one terms filling out all𝐀\\mathbf\{A\}\. Sovec⁡\(𝐓R\)⊤∈row⁡\(𝐇⊗𝐇\)\\operatorname\{vec\}\(\\mathbf\{T\}\_\{R\}\)^\{\\top\}\\in\\operatorname\{row\}\(\\mathbf\{H\}\\otimes\\mathbf\{H\}\)is𝐓R=𝐇⊤​𝐀R​𝐇\\mathbf\{T\}\_\{R\}=\\mathbf\{H\}^\{\\top\}\\mathbf\{A\}\_\{R\}\\mathbf\{H\}for some𝐀R\\mathbf\{A\}\_\{R\}, and𝟏V2⊤=𝟏V⊤⊗𝟏V⊤∈row⁡\(𝐇\)⊗row⁡\(𝐇\)\\mathbf\{1\}\_\{V^\{2\}\}^\{\\top\}=\\mathbf\{1\}\_\{V\}^\{\\top\}\\otimes\\mathbf\{1\}\_\{V\}^\{\\top\}\\in\\operatorname\{row\}\(\\mathbf\{H\}\)\\otimes\\operatorname\{row\}\(\\mathbf\{H\}\)if and only if𝟏V⊤∈row⁡\(𝐇\)\\mathbf\{1\}\_\{V\}^\{\\top\}\\in\\operatorname\{row\}\(\\mathbf\{H\}\), since𝟏𝟏⊤=𝐇⊤​𝐀𝐇\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}=\\mathbf\{H\}^\{\\top\}\\mathbf\{A\}\\mathbf\{H\}requires the rank\-one matrix𝟏𝟏⊤\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}to have column space insidecol⁡\(𝐇⊤\)=row⁡\(𝐇\)⊤\\operatorname\{col\}\(\\mathbf\{H\}^\{\\top\}\)=\\operatorname\{row\}\(\\mathbf\{H\}\)^\{\\top\}, and conversely𝟏V⊤=𝐚⊤​𝐇\\mathbf\{1\}\_\{V\}^\{\\top\}=\\mathbf\{a\}^\{\\top\}\\mathbf\{H\}gives𝟏𝟏⊤=𝐇⊤​𝐚​𝐚⊤​𝐇\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}=\\mathbf\{H\}^\{\\top\}\\mathbf\{a\}\\,\\mathbf\{a\}^\{\\top\}\\mathbf\{H\}\. ∎

𝐓R\\mathbf\{T\}\_\{R\}==𝐇⊤\\mathbf\{H\}^\{\\top\}𝐀R\\mathbf\{A\}\_\{R\}𝐇\\mathbf\{H\}V×VV\\times VV×dV\\times dd×dd\\times dd×Vd\\times VR⁡\(di,dj\)R\(d\_\{i\},d\_\{j\}\)h​\(di\)⊤h\(d\_\{i\}\)^\{\\top\}h⁡\(dj\)h\(d\_\{j\}\)Figure 5:Factorization𝐓R=𝐇⊤​𝐀R​𝐇\\mathbf\{T\}\_\{R\}=\\mathbf\{H\}^\{\\top\}\\mathbf\{A\}\_\{R\}\\,\\mathbf\{H\}, withd≪Vd\\ll V\. The shaded row of𝐇⊤\\mathbf\{H\}^\{\\top\}ish​\(di\)⊤h\(d\_\{i\}\)^\{\\top\}and the shaded column of𝐇\\mathbf\{H\}ish⁡\(dj\)h\(d\_\{j\}\); the light bands in𝐓R\\mathbf\{T\}\_\{R\}are rowiiand columnjj, and their crossing is the entry\(𝐓R\)i​j=R⁡\(di,dj\)\(\\mathbf\{T\}\_\{R\}\)\_\{ij\}=R\(d\_\{i\},d\_\{j\}\), the readouth​\(di\)⊤​𝐀R​h​\(dj\)h\(d\_\{i\}\)^\{\\top\}\\mathbf\{A\}\_\{R\}\\,h\(d\_\{j\}\)\. The inner dimension isdd, givingrank⁡𝐓R≤d\\operatorname\{rank\}\\mathbf\{T\}\_\{R\}\\leq d, the obstruction of Corollary[6\.1](https://arxiv.org/html/2609.18047#S6.Thmcorollary1); the second condition of the theorem,𝟏∈row⁡\(𝐇\)\\mathbf\{1\}\\in\\operatorname\{row\}\(\\mathbf\{H\}\), constrains𝐇\\mathbf\{H\}alone\.The lifted relation222This is the bilinear scoring function of relational embedding models such as RESCAL\[[21](https://arxiv.org/html/2609.18047#bib.bib20)\], which Theorem[6\.1](https://arxiv.org/html/2609.18047#S6.Thmtheorem1)recovers as the exact form of a lifted binary relation; under the vector logic, the score has a truth\-conditional reading\.is a bilinear form𝐀R\\mathbf\{A\}\_\{R\}on the compressed space, andR⁡\(di,dj\)R\(d\_\{i\},d\_\{j\}\)is read off ash​\(di\)⊤​𝐀R​h​\(dj\)h\(d\_\{i\}\)^\{\\top\}\\mathbf\{A\}\_\{R\}\\,h\(d\_\{j\}\)\.

### 6\.2Obstructions

###### Corollary 6\.1\(rank obstruction\)\.

IfRRlifts exactly over𝐇\\mathbf\{H\}, thenrank⁡𝐓R≤rank⁡𝐇≤d\\operatorname\{rank\}\\mathbf\{T\}\_\{R\}\\leq\\operatorname\{rank\}\\mathbf\{H\}\\leq d\. In particular, the identity relation\{\(di,di\)\}\\\{\(d\_\{i\},d\_\{i\}\)\\\}, with𝐓==IV\\mathbf\{T\}\_\{=\}=I\_\{V\}, lifts exactly only over free geometries\.

###### Proof\.

rank⁡\(𝐇⊤​𝐀R​𝐇\)≤rank⁡𝐇\\operatorname\{rank\}\(\\mathbf\{H\}^\{\\top\}\\mathbf\{A\}\_\{R\}\\mathbf\{H\}\)\\leq\\operatorname\{rank\}\\mathbf\{H\}; andrank⁡IV=V\\operatorname\{rank\}I\_\{V\}=Vforcesrank⁡𝐇=V\\operatorname\{rank\}\\mathbf\{H\}=V\. ∎

The rank bound is a lower bound only; the factorization constrains both argument positions through the same geometry, and the exact value follows from the proof of Theorem[6\.1](https://arxiv.org/html/2609.18047#S6.Thmtheorem1)\.

###### Corollary 6\.2\(minimal dimension for a relation\)\.

The leastddfor which some𝐇∈ℝd×V\\mathbf\{H\}\\in\\mathbb\{R\}^\{d\\times V\}carries an exact lift ofRRisdimspan⁡\(\{𝟏\}∪row⁡\(𝐓R\)∪row⁡\(𝐓R⊤\)\)\\dim\\operatorname\{span\}\\bigl\(\\\{\\mathbf\{1\}\\\}\\cup\\operatorname\{row\}\(\\mathbf\{T\}\_\{R\}\)\\cup\\operatorname\{row\}\(\\mathbf\{T\}\_\{R\}^\{\\top\}\)\\bigr\), attained by any𝐇\\mathbf\{H\}whose rows form a basis of that span\.

###### Proof\.

By the identification in the proof of Theorem[6\.1](https://arxiv.org/html/2609.18047#S6.Thmtheorem1),\{𝐇⊤​𝐀𝐇:𝐀∈ℝd×d\}\\\{\\mathbf\{H\}^\{\\top\}\\mathbf\{A\}\\mathbf\{H\}:\\mathbf\{A\}\\in\\mathbb\{R\}^\{d\\times d\}\\\}is the set ofV×VV\\times Vmatrices whose rows lie inrow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\)and whose columns lie inrow⁡\(𝐇\)⊤\\operatorname\{row\}\(\\mathbf\{H\}\)^\{\\top\}, so𝐓R=𝐇⊤​𝐀R​𝐇\\mathbf\{T\}\_\{R\}=\\mathbf\{H\}^\{\\top\}\\mathbf\{A\}\_\{R\}\\mathbf\{H\}is solvable if and only ifrow⁡\(𝐓R\)⊆row⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{T\}\_\{R\}\)\\subseteq\\operatorname\{row\}\(\\mathbf\{H\}\)androw⁡\(𝐓R⊤\)⊆row⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{T\}\_\{R\}^\{\\top\}\)\\subseteq\\operatorname\{row\}\(\\mathbf\{H\}\)\. With the requirement𝟏∈row⁡\(𝐇\)\\mathbf\{1\}\\in\\operatorname\{row\}\(\\mathbf\{H\}\), exact lift is containment of the stated span inrow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\), whenced≥rank⁡𝐇≥dimspan⁡\(\{𝟏\}∪row⁡\(𝐓R\)∪row⁡\(𝐓R⊤\)\)d\\geq\\operatorname\{rank\}\\mathbf\{H\}\\geq\\dim\\operatorname\{span\}\\bigl\(\\\{\\mathbf\{1\}\\\}\\cup\\operatorname\{row\}\(\\mathbf\{T\}\_\{R\}\)\\cup\\operatorname\{row\}\(\\mathbf\{T\}\_\{R\}^\{\\top\}\)\\bigr\), with equality when the rows of𝐇\\mathbf\{H\}are a basis of the span\. ∎

A single monadic predicate hasrank⁡\[𝐭P;𝟏\]≤2\\operatorname\{rank\}\[\\mathbf\{t\}\_\{P\};\\mathbf\{1\}\]\\leq 2, and lifts exactly in dimension two, so forcing high dimension at type⟨e,t⟩\\langle e,t\\ranglerequires a lexicon, and Corollary[5\.2](https://arxiv.org/html/2609.18047#S5.Thmcorollary2)uses the whole closed such one\. A single binary relation can force the free regime by itself, and the relation that does so is identity, the denotation of the copula inDaniel is Daniel\. A strict total order does so as well, one dimension above its rank:𝐓<\\mathbf\{T\}\_\{<\}is strictly upper triangular with ones above the diagonal, of rankV−1V\-1, whilerow⁡\(𝐓<\)=span⁡\{𝐞2⊤,…,𝐞V⊤\}\\operatorname\{row\}\(\\mathbf\{T\}\_\{<\}\)=\\operatorname\{span\}\\\{\\mathbf\{e\}\_\{2\}^\{\\top\},\\dots,\\mathbf\{e\}\_\{V\}^\{\\top\}\\\}androw⁡\(𝐓<⊤\)=span⁡\{𝐞1⊤,…,𝐞V−1⊤\}\\operatorname\{row\}\(\\mathbf\{T\}\_\{<\}^\{\\top\}\)=\\operatorname\{span\}\\\{\\mathbf\{e\}\_\{1\}^\{\\top\},\\dots,\\mathbf\{e\}\_\{V\-1\}^\{\\top\}\\\}together spanℝ1×V\\mathbb\{R\}^\{1\\times V\}forV≥2V\\geq 2, so Corollary[6\.2](https://arxiv.org/html/2609.18047#S6.Thmcorollary2)returnsVV\. Equivalence relations withkkclasses haverank⁡𝐓R=k\\operatorname\{rank\}\\mathbf\{T\}\_\{R\}=k, and, since𝐓R\\mathbf\{T\}\_\{R\}is symmetric with𝟏\\mathbf\{1\}in its row space, compress to dimensionkkexactly\. The rank of a relation’s truth matrix bounds the dimension of any exact carrier from below, in the same way thatrank⁡\[𝐓;𝟏\]\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]bounds the dimension for a monadic lexicon, and Corollary[6\.2](https://arxiv.org/html/2609.18047#S6.Thmcorollary2)gives the value\.

## 7Index sorts

The intensional layer places its free carrier on the index space,hS​\(s\)=𝐞sh\_\{S\}\(s\)=\\mathbf\{e\}\_\{s\}; for a discrete sort with finitely many indices, the results of Sections[5](https://arxiv.org/html/2609.18047#S5)and[6](https://arxiv.org/html/2609.18047#S6)apply to any compressed geometry𝐇W∈ℝd×n\\mathbf\{H\}\_\{W\}\\in\\mathbb\{R\}^\{d\\times n\}of the world sort with the same proofs333We defer them here, and encourage the ambitious reader to try them\., once the objects are identified\. Here, propositions play the role of monadic predicates:φ\\varphihas truth profile𝐯​\(φ\)⊤\\mathbf\{v\}\(\\varphi\)^\{\\top\}as its truth row overWW, and a lexicon of propositions𝒫W\\mathcal\{P\}\_\{W\}has truth matrix𝐓W\\mathbf\{T\}\_\{W\}\. Accessibility plays the role of a binary relation, with truth matrix𝐀\\mathbf\{A\}\.

###### Proposition 7\.1\(compressed worlds\)\.

Let𝐇W∈ℝd×n\\mathbf\{H\}\_\{W\}\\in\\mathbb\{R\}^\{d\\times n\}be a geometry ofWW\.

1. 1\.Everyφ∈𝒫W\\varphi\\in\\mathcal\{P\}\_\{W\}has an exact liftLφ​𝐇W=𝐏φL\_\{\\varphi\}\\mathbf\{H\}\_\{W\}=\\mathbf\{P\}\_\{\\varphi\}if and only ifrow⁡\(𝐇W\)⊇row⁡\[𝐓W;𝟏\]\\operatorname\{row\}\(\\mathbf\{H\}\_\{W\}\)\\supseteq\\operatorname\{row\}\[\\mathbf\{T\}\_\{W\};\\mathbf\{1\}\]; the minimal exact dimension isrank⁡\[𝐓W;𝟏\]\\operatorname\{rank\}\[\\mathbf\{T\}\_\{W\};\\mathbf\{1\}\]\.
2. 2\.Accessibility lifts exactly as a bilinear form if and only if𝟏∈row⁡\(𝐇W\)\\mathbf\{1\}\\in\\operatorname\{row\}\(\\mathbf\{H\}\_\{W\}\)and𝐀=𝐇W⊤​𝐀^​𝐇W\\mathbf\{A\}=\\mathbf\{H\}\_\{W\}^\{\\top\}\\widehat\{\\mathbf\{A\}\}\\mathbf\{H\}\_\{W\}for some𝐀^∈ℝd×d\\widehat\{\\mathbf\{A\}\}\\in\\mathbb\{R\}^\{d\\times d\}; in particular,rank⁡𝐀≤d\\operatorname\{rank\}\\mathbf\{A\}\\leq d\.
3. 3\.Under \(1\) and \(2\) the modal accumulation of \([1](https://arxiv.org/html/2609.18047#S2.E1)\) computes in the compressed space: with𝐯⁡\(φ\)=\(𝐰φ​𝐇W\)⊤\\mathbf\{v\}\(\\varphi\)=\(\\mathbf\{w\}\_\{\\varphi\}\\mathbf\{H\}\_\{W\}\)^\{\\top\}and𝟏=\(𝐰𝟏​𝐇W\)⊤\\mathbf\{1\}=\(\\mathbf\{w\}\_\{\\mathbf\{1\}\}\\mathbf\{H\}\_\{W\}\)^\{\\top\}for the readouts of \(1\), 𝐀​𝐯​\(φ\)=𝐇W⊤​𝐀^​𝐇W​𝐇W⊤​𝐰φ⊤,𝐀⁡\(𝟏−𝐯⁡\(φ\)\)=𝐇W⊤​𝐀^​𝐇W​𝐇W⊤​\(𝐰𝟏−𝐰φ\)⊤,\\mathbf\{A\}\\,\\mathbf\{v\}\(\\varphi\)=\\mathbf\{H\}\_\{W\}^\{\\top\}\\,\\widehat\{\\mathbf\{A\}\}\\,\\mathbf\{H\}\_\{W\}\\mathbf\{H\}\_\{W\}^\{\\top\}\\,\\mathbf\{w\}\_\{\\varphi\}^\{\\top\},\\qquad\\mathbf\{A\}\\,\(\\mathbf\{1\}\-\\mathbf\{v\}\(\\varphi\)\)=\\mathbf\{H\}\_\{W\}^\{\\top\}\\,\\widehat\{\\mathbf\{A\}\}\\,\\mathbf\{H\}\_\{W\}\\mathbf\{H\}\_\{W\}^\{\\top\}\\,\(\\mathbf\{w\}\_\{\\mathbf\{1\}\}\-\\mathbf\{w\}\_\{\\varphi\}\)^\{\\top\},each ad×dd\\times dcomputation followed by one expansion through𝐇W⊤\\mathbf\{H\}\_\{W\}^\{\\top\}, after which the decisions of \([1](https://arxiv.org/html/2609.18047#S2.E1)\) apply unchanged\.

###### Proof\.

\(1\) and \(2\) are simply Theorems[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1)and[6\.1](https://arxiv.org/html/2609.18047#S6.Thmtheorem1), withWWfor𝒟e\\mathcal\{D\}\_\{e\}\. \(3\) is substitution\. ∎

𝐀\\mathbf\{A\}𝐯⁡\(φ\)\\mathbf\{v\}\(\\varphi\)==𝐇W⊤\\mathbf\{H\}\_\{W\}^\{\\top\}𝐀^\\widehat\{\\mathbf\{A\}\}𝐇W​𝐇W⊤\\mathbf\{H\}\_\{W\}\\mathbf\{H\}\_\{W\}^\{\\top\}𝐰φ⊤\\mathbf\{w\}\_\{\\varphi\}^\{\\top\}n×nn\\times nn×1n\\times 1n×dn\\times dd×dd\\times dd×dd\\times dd×1d\\times 1oneexpansiond×dd\\times dcomputationFigure 6:The compressed modal accumulation of \(3\),𝐀​𝐯​\(φ\)=𝐇W⊤​𝐀^​\(𝐇W​𝐇W⊤\)​𝐰φ⊤\\mathbf\{A\}\\,\\mathbf\{v\}\(\\varphi\)=\\mathbf\{H\}\_\{W\}^\{\\top\}\\,\\widehat\{\\mathbf\{A\}\}\\,\(\\mathbf\{H\}\_\{W\}\\mathbf\{H\}\_\{W\}^\{\\top\}\)\\,\\mathbf\{w\}\_\{\\varphi\}^\{\\top\}, withd≪nd\\ll n\. To the right of𝐇W⊤\\mathbf\{H\}\_\{W\}^\{\\top\}, every block isd×dd\\times dord×1d\\times 1, so the accumulation is ad×dd\\times dcomputation, and one expansion through𝐇W⊤\\mathbf\{H\}\_\{W\}^\{\\top\}; the threshold checks of \([1](https://arxiv.org/html/2609.18047#S2.E1)\) then apply to the expanded vector unchanged\. The𝐇W​𝐇W⊤\\mathbf\{H\}\_\{W\}\\mathbf\{H\}\_\{W\}^\{\\top\}is as a singled×dd\\times dblock, because it is formed once from the geometry and shared across propositionsφ\\varphi\.The reflexive and transitive frames of the modal logics that concern the intensional layer have accessibility matrices of varying rank; the empty relation has rank zero, and imposes only the𝟏\\mathbf\{1\}requirement, a universal relation has rank one and compresses to a single dimension, a strict linear order, of rankn−1n\-1, forces the free regime by Corollary[6\.2](https://arxiv.org/html/2609.18047#S6.Thmcorollary2), its row and column spaces together spanningℝ1×n\\mathbb\{R\}^\{1\\times n\}, as does a reflexive linear order, whose matrix is upper triangular with unit diagonal, and the identity relation \(the frame of the trivial modality\), by Corollary[6\.1](https://arxiv.org/html/2609.18047#S6.Thmcorollary1); the strict future accessibility of a discrete time sort, therefore, admits exact lift only over the free carrier of that sort\. Since the operators of \([1](https://arxiv.org/html/2609.18047#S2.E1)\) read𝐀\\mathbf\{A\}only through its Boolean support, an approximate carrier need only reproduce the support of𝐀\\mathbf\{A\}for the modal verdicts to survive, a weaker requirement than exact bilinear factorization444The product\-modality is future work, which concerns the support factorization in its own right; we do not pursue the approximate modal case here\.\.

## 8Relaxation

Exact lift is a subspace condition that either holds or fails outright; learned geometries will fail it for most predicates, so we might wonder, then how far is a geometry from carrying a predicate at all? The direct construction in the proof of Theorem[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1)suggests that exact lift places𝐭P\\mathbf\{t\}\_\{P\}inrow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\), and the distance from𝐭P\\mathbf\{t\}\_\{P\}torow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\)is a least\-squares quantity\.

### 8\.1Defect

###### Definition 8\.1\(defect\)\.

For a predicatePPwith truth row𝐭P\\mathbf\{t\}\_\{P\}and a geometry𝐇\\mathbf\{H\}, the defect is

δ\(P;𝐇\)=min𝐰∈ℝ1×d∥𝐰𝐇−𝐭P∥2\.\\delta\(P;\\mathbf\{H\}\)=\\min\_\{\\mathbf\{w\}\\in\\mathbb\{R\}^\{1\\times d\}\}\\bigl\\lVert\\mathbf\{w\}\\mathbf\{H\}\-\\mathbf\{t\}\_\{P\}\\bigr\\rVert\_\{2\}\.

row⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\)𝐭P\\mathbf\{t\}\_\{P\}𝐭P​𝐇†​𝐇\\mathbf\{t\}\_\{P\}\\mathbf\{H\}^\{\\dagger\}\\mathbf\{H\}δ⁡\(P,𝐇\)\\delta\(P;\\mathbf\{H\}\)θ\\theta𝟏\\mathbf\{1\}Figure 7:The defect of Definition[8\.1](https://arxiv.org/html/2609.18047#S8.Thmdefinition1), in which the truth row𝐭P\\mathbf\{t\}\_\{P\}is projected ontorow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\); the length isδ⁡\(P,𝐇\)\\delta\(P;\\mathbf\{H\}\)and the angleθ\\thetais the principal angle between𝐭P\\mathbf\{t\}\_\{P\}and the row space, withδ=∥𝐭P∥​sin⁡θ\\delta=\\lVert\\mathbf\{t\}\_\{P\}\\rVert\\sin\\theta\. Exact lift isθ=0\\theta=0\. The𝟏\\mathbf\{1\}row is drawn nearly in the plane, as Section[8\.4](https://arxiv.org/html/2609.18047#S8.SS4)finds it for trained geometries; the principal angles of Table[2](https://arxiv.org/html/2609.18047#S8.T2)are the anglesθ\\thetataken jointly overrow⁡\[𝐓;𝟏\]\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]\.###### Proposition 8\.1\.

δ⁡\(P,𝐇\)=∥𝐭P​\(IV−𝐇†​𝐇\)∥2\\delta\(P;\\mathbf\{H\}\)=\\lVert\\mathbf\{t\}\_\{P\}\(I\_\{V\}\-\\mathbf\{H\}^\{\\dagger\}\\mathbf\{H\}\)\\rVert\_\{2\}, where𝐇†\\mathbf\{H\}^\{\\dagger\}is the Moore–Penrose pseudoinverse, and𝐇†​𝐇\\mathbf\{H\}^\{\\dagger\}\\mathbf\{H\}is the orthogonal projector ontorow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\); the minimizer is𝐰=𝐭P​𝐇†\\mathbf\{w\}=\\mathbf\{t\}\_\{P\}\\mathbf\{H\}^\{\\dagger\}\. Moreover,δ⁡\(P,𝐇\)=0\\delta\(P;\\mathbf\{H\}\)=0if and only if𝐭P∈row⁡\(𝐇\)\\mathbf\{t\}\_\{P\}\\in\\operatorname\{row\}\(\\mathbf\{H\}\), so that if𝟏∈row⁡\(𝐇\)\\mathbf\{1\}\\in\\operatorname\{row\}\(\\mathbf\{H\}\), exact lift ofPPisδ⁡\(P,𝐇\)=0\\delta\(P;\\mathbf\{H\}\)=0\.

###### Proof sketch\.

\{𝐰𝐇\}\\\{\\mathbf\{w\}\\mathbf\{H\}\\\}isrow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\), and the closest point of a subspace to𝐭P\\mathbf\{t\}\_\{P\}is its orthogonal projection𝐭P​𝐇†​𝐇\\mathbf\{t\}\_\{P\}\\mathbf\{H\}^\{\\dagger\}\\mathbf\{H\}, with𝐭P​\(I−𝐇†​𝐇\)\\mathbf\{t\}\_\{P\}\(I\-\\mathbf\{H\}^\{\\dagger\}\\mathbf\{H\}\)vanishes exactly on the subspace\. The final clause is Theorem[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1)\. ∎

The defect is computable on any trained embedding by one least\-squares solve per predicate, with predicate extensions supplied by lexical resources or feature norms\[[15](https://arxiv.org/html/2609.18047#bib.bib22)\]; it is a statistic of the trained geometry and a property of the training data and objective, which is unconstrained; the condition is, therefore, a specification, and the defect is a measurement\.

### 8\.2Threshold and separability

Conceptual spaces\[[8](https://arxiv.org/html/2609.18047#bib.bib9)\]and degree semantics\[[12](https://arxiv.org/html/2609.18047#bib.bib23)\]treat graded predicates as regions in a structured space, with the classical predicate recovered by a threshold,⟦P⟧\(x\)=𝟏\[d\(x,RP\)≤θP\]\\llbracket P\\rrbracket\(x\)=\\mathbf\{1\}\[\\,d\(x,R\_\{P\}\)\\leq\\theta\_\{P\}\\,\]; its simplest instance is a half\-space, what the exact lift becomes when the equality in \([2](https://arxiv.org/html/2609.18047#S5.E2)\) is weakened to a sign condition\.

###### Definition 8\.2\(thresholded lift\)\.

PPis linearly separable over𝐇\\mathbf\{H\}if there are𝐰∈ℝ1×d\\mathbf\{w\}\\in\\mathbb\{R\}^\{1\\times d\}andθ∈ℝ\\theta\\in\\mathbb\{R\}with𝐰​h​\(di\)\>θ\\mathbf\{w\}h\(d\_\{i\}\)\>\\thetawhenP⁡\(di\)=1P\(d\_\{i\}\)=1and𝐰​h​\(di\)<θ\\mathbf\{w\}h\(d\_\{i\}\)<\\thetawhenP⁡\(di\)=0P\(d\_\{i\}\)=0\.

###### Proposition 8\.2\.

IfPPhas an exact lift over𝐇\\mathbf\{H\}, thenPPis linearly separable over𝐇\\mathbf\{H\}; the converse fails\.

###### Proof\.

With𝐰=𝐰P\\mathbf\{w\}=\\mathbf\{w\}\_\{P\}from the direct construction,𝐰​h​\(di\)=P⁡\(di\)∈\{0,1\}\\mathbf\{w\}h\(d\_\{i\}\)=P\(d\_\{i\}\)\\in\\\{0,1\\\}, andθ=12\\theta=\\tfrac\{1\}\{2\}separates\. For the converse, taked=1d=1,𝐇=\[1 2 3 4\]\\mathbf\{H\}=\[\\,1\\;2\\;3\\;4\\,\], and𝐭P=\(0,0,1,1\)\\mathbf\{t\}\_\{P\}=\(0,0,1,1\): the thresholdθ=52\\theta=\\tfrac\{5\}\{2\}separates, while𝐭P∈span⁡\{𝐇,𝟏\}\\mathbf\{t\}\_\{P\}\\in\\operatorname\{span\}\\\{\\mathbf\{H\},\\mathbf\{1\}\\\}would requirea\+c=0a\+c=0and2​a\+c=02a\+c=0from the first two coordinates, forcinga=c=0a=c=0, which contradicts the third coordinate; exact lift, therefore, fails, in the affine sense of Proposition[5\.1](https://arxiv.org/html/2609.18047#S5.Thmproposition1), as well as the linear one\. ∎

h⁡\(di\)h\(d\_\{i\}\)PP11223344best affine readoutthresholdθ=5/2\\theta=5/2Figure 8:The example of[Proposition8\.2](https://arxiv.org/html/2609.18047#S8.Thmproposition2): four entities on a line,𝐭P=\(0,0,1,1\)\\mathbf\{t\}\_\{P\}=\(0,0,1,1\)\. No affine function of position takes the values0,0,1,10,0,1,1, so the affine defect is positive, while the threshold at5/25/2separates the extension exactly\. The measurements of Section[8\.4](https://arxiv.org/html/2609.18047#S8.SS4)find learned geometries in this position for nearly every predicate: high separability, positive defect\.A linear probe tests the linear separability of a predicate over a learned geometry:\[[1](https://arxiv.org/html/2609.18047#bib.bib15)\]keep the probe linear, so that accuracy tracks the representation, and\[[10](https://arxiv.org/html/2609.18047#bib.bib16)\]caution that it tracks the capacity of the probe as well; we must account for this, so we do so with held\-out scoring throughout, and a Gaussian null on the held\-out error of Section[8\.4](https://arxiv.org/html/2609.18047#S8.SS4)\. Proposition[8\.2](https://arxiv.org/html/2609.18047#S8.Thmproposition2)is the probe in the vector logic: thresholded relaxation of the homomorphism condition, strictly weaker than the exact condition, and the defectδ⁡\(P,𝐇\)\\delta\(P;\\mathbf\{H\}\)is the condition’s own measure\. Probing practice fits each predicate independently; the rank criterion adds that a shared readout basis across the lexicon requires the𝟏\\mathbf\{1\}\-row, or, equivalently, a bias, and that the ranks of Corollaries[5\.1](https://arxiv.org/html/2609.18047#S5.Thmcorollary1)and[6\.1](https://arxiv.org/html/2609.18047#S6.Thmcorollary1)bound what any geometry of a given dimension can carry, before any probe is fit\.

### 8\.3Parallelograms

Let us now return to the four\-entity lexicon of Example[5\.1](https://arxiv.org/html/2609.18047#S5.Thmexample1):𝐭write=\(1,1,0,0\)\\mathbf\{t\}\_\{\\mathrm\{write\}\}=\(1,1,0,0\),𝐭published=\(1,0,1,0\)\\mathbf\{t\}\_\{\\mathrm\{published\}\}=\(1,0,1,0\),𝟏=\(1,1,1,1\)\\mathbf\{1\}=\(1,1,1,1\)\. The three rows are independent, sorank⁡\[𝐓;𝟏\]=3\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]=3, and, by Corollary[5\.3](https://arxiv.org/html/2609.18047#S5.Thmcorollary3), the admissible dependences form a one\-dimensional space \(the orthogonal complement of the row space\), spanned by

𝐜=\(1,−1,−1,1\)⊤:𝐭write​𝐜=1−1=0,𝐭published​𝐜=1−1=0,𝟏​𝐜=0\.\\mathbf\{c\}=\(1,\-1,\-1,1\)^\{\\top\}:\\qquad\\mathbf\{t\}\_\{\\mathrm\{write\}\}\\,\\mathbf\{c\}=1\-1=0,\\quad\\mathbf\{t\}\_\{\\mathrm\{published\}\}\\,\\mathbf\{c\}=1\-1=0,\\quad\\mathbf\{1\}\\,\\mathbf\{c\}=0\.Every minimal exact geometry for this lexicon, therefore, satisfies exactly one dependence,

h⁡\(d1\)−h⁡\(d2\)=h⁡\(d3\)−h⁡\(d4\),h\(d\_\{1\}\)\-h\(d\_\{2\}\)=h\(d\_\{3\}\)\-h\(d\_\{4\}\),which form a parallelogram, with the interpretation: the difference between a writer who is published and one who is unpublished equals the difference between a nonwriter who is published and one who is unpublished\. The simplest nontrivial solutions of the constraints the rank criterion imposes are analogy structures of the kind reported for word embeddings since the classic\[[18](https://arxiv.org/html/2609.18047#bib.bib12)\]\.

h⁡\(d1\)h\(d\_\{1\}\): writer, publishedh⁡\(d2\)h\(d\_\{2\}\): writer, unpublishedh⁡\(d3\)h\(d\_\{3\}\): nonwriter, publishedh⁡\(d4\)h\(d\_\{4\}\): nonwriter, unpublished\+published\+\\mathrm\{published\}\+published\+\\mathrm\{published\}\+write\+\\mathrm\{write\}\+write\+\\mathrm\{write\}Figure 9:The parallelogram forced on every minimal exact geometry for the four\-entity lexicon \(simplified from Figure[4](https://arxiv.org/html/2609.18047#S5.F4)\):h⁡\(d1\)−h⁡\(d2\)=h⁡\(d3\)−h⁡\(d4\)h\(d\_\{1\}\)\-h\(d\_\{2\}\)=h\(d\_\{3\}\)\-h\(d\_\{4\}\), so the displacement forpublishedis the same, whether taken from a writer or from a nonwriter, and likewise forwrite\. The single admissible dependence𝐜=\(1,−1,−1,1\)\\mathbf\{c\}=\(1,\-1,\-1,1\)is this figure\.In general, the admissible dependences arerow⁡\[𝐓;𝟏\]⟂\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]^\{\\perp\}, and integer vectors in that complement with two entries\+1\+1and two entries−1\-1are exactly the parallelograms the lexicon permits at all\. This is not deep, but follows from pairs of entity pairs that agree on feature difference\.

###### Proposition 8\.3\(Boolean cube\)\.

Let𝒫\\mathcal\{P\}consist ofmmbinary features on theV=2mV=2^\{m\}entities of\{0,1\}m\\\{0,1\\\}^\{m\}, featurekkhaving truth rowx↦xkx\\mapsto x\_\{k\}\. Thenrank⁡\[𝐓;𝟏\]=m\+1\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]=m\+1, and the space of admissible dependencesrow⁡\[𝐓;𝟏\]⟂\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]^\{\\perp\}, of dimension2m−m−12^\{m\}\-m\-1, is spanned by the parallelogram vectors𝐞x−𝐞x\+𝐞k−𝐞z\+𝐞z\+𝐞k\\mathbf\{e\}\_\{x\}\-\\mathbf\{e\}\_\{x\+\\mathbf\{e\}\_\{k\}\}\-\\mathbf\{e\}\_\{z\}\+\\mathbf\{e\}\_\{z\+\\mathbf\{e\}\_\{k\}\}over coordinateskkand pointsx,zx,zwithxk=zk=0x\_\{k\}=z\_\{k\}=0\.

###### Proof\.

The rowsx↦xkx\\mapsto x\_\{k\}andx↦1x\\mapsto 1are the affine functions’ basis on the cube, and are independent, giving the rank\. Each parallelogram vector is orthogonal to every affine functionf⁡\(x\)=a0\+∑kak​xkf\(x\)=a\_\{0\}\+\\sum\_\{k\}a\_\{k\}x\_\{k\}, sincef⁡\(x\)−f⁡\(x\+𝐞k\)−f⁡\(z\)\+f⁡\(z\+𝐞k\)=−ak\+ak=0f\(x\)\-f\(x\+\\mathbf\{e\}\_\{k\}\)\-f\(z\)\+f\(z\+\\mathbf\{e\}\_\{k\}\)=\-a\_\{k\}\+a\_\{k\}=0, so the parallelogram span lies inrow⁡\[𝐓;𝟏\]⟂\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]^\{\\perp\}\.

Conversely, letf∈ℝVf\\in\\mathbb\{R\}^\{V\}be orthogonal to every parallelogram vector; thenf⁡\(x\+𝐞k\)−f⁡\(x\)=f⁡\(z\+𝐞k\)−f⁡\(z\)f\(x\+\\mathbf\{e\}\_\{k\}\)\-f\(x\)=f\(z\+\\mathbf\{e\}\_\{k\}\)\-f\(z\)for allx,zx,zwithxk=zk=0x\_\{k\}=z\_\{k\}=0, so the increment along coordinatekkis a constantaka\_\{k\}, and induction on the number of nonzero coordinates givesf⁡\(x\)=f⁡\(0\)\+∑kak​xkf\(x\)=f\(0\)\+\\sum\_\{k\}a\_\{k\}x\_\{k\}, an affine function\. The orthogonal complement of the parallelogram span is, therefore, the space of affine functions, and the parallelogram span is its complement\. ∎

000000100100010010110110001001101101011011111111a1a\_\{1\}a2a\_\{2\}a3a\_\{3\}h⁡\(x\)=a0\+∑kak​xkh\(x\)=a\_\{0\}\+\\sum\_\{k\}a\_\{k\}x\_\{k\}shaded face:𝐞010−𝐞110−𝐞011\+𝐞111\\mathbf\{e\}\_\{010\}\-\\mathbf\{e\}\_\{110\}\-\\mathbf\{e\}\_\{011\}\+\\mathbf\{e\}\_\{111\}rank⁡\[𝐓;𝟏\]=4\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]=4,dimrow⁡\[𝐓;𝟏\]⟂=4\\;\\dim\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]^\{\\perp\}=4Figure 10:The Boolean cube of Proposition[8\.3](https://arxiv.org/html/2609.18047#S8.Thmproposition3)form=3m=3: eight entities, three features, and a minimal exact geometry inℝ4\\mathbb\{R\}^\{4\}\(drawn in three dimensions, with the affine offset suppressed\)\. Every minimal exact geometry is an affine image of the cube, so each face is a parallelogram; the four independent faces span the admissible dependences, of dimension23−3−1=42^\{3\}\-3\-1=4, and the affine functionsa0\+∑kak​xka\_\{0\}\+\\sum\_\{k\}a\_\{k\}x\_\{k\}are licensed by the lexicon\.In a minimal exact geometry for a full factorial feature lexicon, every coordinate difference is, thus, a constant vector, which is the setting in which analogy by vector arithmetic is exact; an exact geometry of higher rank satisfies a subset of these equalities, and the free geometry keeps the2m2^\{m\}entity vectors affinely independent\.

### 8\.4Measurement

We turn, now, to implementation and experiment555Code is available at the project repo\. Minimal working verifiable code was written by the author; LLM assistance for this paper was used in the writing and expanding of the experiment software at larger scale, building from the author’s minimal working code\.\. We computed a held\-out errorδ^\\hat\{\\delta\}, written so as to keep it apart from the defectδ\\deltaof Definition[8\.1](https://arxiv.org/html/2609.18047#S8.Thmdefinition1), the probe, the principal angles, and the dimension sweep of Corollary[5\.1](https://arxiv.org/html/2609.18047#S5.Thmcorollary1), on two classic lexicons against two likewise classic embeddings\. The geometries are a 300\-dimensional GloVe embedding\[[23](https://arxiv.org/html/2609.18047#bib.bib13)\]and the 300\-dimensional word2vec embedding of\[[17](https://arxiv.org/html/2609.18047#bib.bib26)\]; each is used as its geometry matrix𝐇\\mathbf\{H\}without centering, whitening, normalization, or truncation\. The projection residual of Proposition[8\.1](https://arxiv.org/html/2609.18047#S8.Thmproposition1)depends onrow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\)alone; the held\-out error below depends on the coordinates through its penalty, and is reported on the coordinates as published\.

The first lexicon uses the McRae feature norms\[[15](https://arxiv.org/html/2609.18047#bib.bib22)\]\. Following the norms’ own inclusion threshold, a feature holds of a concept when at least five participants produced it\. Entries below the threshold enter𝐓\\mathbf\{T\}as false, so the truth\-conditional reading treats nonproduction as a negative judgment, which is an assumption about the norms rather than a datum in them\.

Concepts the norms distinguish by sense, such asbatin its animal and baseball senses, share one word vector and receive the union of their features\. The 541 concepts collapse to 532 words, all present in GloVe and 531 in word2vec\. Retaining features assigned to at least fifteen words, and to at most fifteen fewer than all of them, gives 76 predicates andrank⁡\[𝐓;𝟏\]=77\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]=77in both embeddings\.

The second lexicon uses WordNet hypernyms\[[7](https://arxiv.org/html/2609.18047#bib.bib25)\]\. We select monosemous nouns among each embedding’s twenty thousand most frequent tokens\. Hypernyms of depth at least four with 30–500 members become predicates\. GloVe supplies 6271 nouns and 171 predicates, with 165 distinct extensions andrank⁡\[𝐓;𝟏\]=166\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]=166; for word2vec the figures are 5257 nouns and 143 predicates, with 139 distinct extensions andrank⁡\[𝐓;𝟏\]=139\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]=139\. Predicates with identical extensions remain separate rows\.

In every case,rank⁡\[𝐓;𝟏\]<d=300<V\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]<d=300<V, which places all four cells in the compressed regime\. The available dimension therefore permits exact monadic compression, although whether a pretrained geometry realizes it remains to be tested\.

The full\-domain quantity is the affine defect of Proposition[5\.1](https://arxiv.org/html/2609.18047#S5.Thmproposition1), normalized by the centered norm of the truth row,

δ\+​\(P,𝐇\)=∥𝐭P\(IV−\(𝐇\+\)†𝐇\+\)∥2∥𝐭P−t¯P​𝟏∥2,t¯P=\|P\|/V,\\delta^\{\+\}\(P;\\mathbf\{H\}\)=\\frac\{\\bigl\\lVert\\mathbf\{t\}\_\{P\}\\bigl\(I\_\{V\}\-\(\\mathbf\{H\}^\{\+\}\)^\{\\dagger\}\\mathbf\{H\}^\{\+\}\\bigr\)\\bigr\\rVert\_\{2\}\}\{\\lVert\\mathbf\{t\}\_\{P\}\-\\bar\{t\}\_\{P\}\\mathbf\{1\}\\rVert\_\{2\}\},\\qquad\\bar\{t\}\_\{P\}=\|P\|/V,so that a constant readout scores11and an exact affine lift scores00\. The held\-out errorδ^​\(P,𝐇\)\\hat\{\\delta\}\(P;\\mathbf\{H\}\)is a predictive quantity, and differs from it in target, in regularization, and in normalization\. The entities are split into five foldsF1,…,F5F\_\{1\},\\dots,F\_\{5\}; for eachkk, a readout\(𝐰\(k\),b\(k\)\)\(\\mathbf\{w\}^\{\(k\)\},b^\{\(k\)\}\)is fit by ridge regression over𝐇\+\\mathbf\{H\}^\{\+\}on the entities outsideFkF\_\{k\}, with the intercept unpenalized, and

δ^​\(P,𝐇\)=\(∑k∑i∈Fk\(𝐰\(k\)​h​\(di\)\+b\(k\)−𝐭P​\(i\)\)2∑k∑i∈Fk\(𝐭P​\(i\)−t¯P\(k\)\)2\)1/2,\\hat\{\\delta\}\(P;\\mathbf\{H\}\)=\\left\(\\frac\{\\sum\_\{k\}\\sum\_\{i\\in F\_\{k\}\}\\bigl\(\\mathbf\{w\}^\{\(k\)\}h\(d\_\{i\}\)\+b^\{\(k\)\}\-\\mathbf\{t\}\_\{P\}\(i\)\\bigr\)^\{2\}\}\{\\sum\_\{k\}\\sum\_\{i\\in F\_\{k\}\}\\bigl\(\\mathbf\{t\}\_\{P\}\(i\)\-\\bar\{t\}^\{\(k\)\}\_\{P\}\\bigr\)^\{2\}\}\\right\)^\{1/2\},witht¯P\(k\)\\bar\{t\}^\{\(k\)\}\_\{P\}the mean of𝐭P\\mathbf\{t\}\_\{P\}overFkF\_\{k\}, so predicting the fold mean scores11\. The penalty is chosen per predicate from\{10−2,10−1,1,10,102,103\}\\\{10^\{\-2\},10^\{\-1\},1,10,10^\{2\},10^\{3\}\\\}by the minimum of this same held\-out error, a selection that biasesδ^\\hat\{\\delta\}downward; choosing it on inner folds of the training data instead moves the medians below by at most0\.010\.01\. The null is the same statistic on a Gaussian matrix of the same shape\. A positiveδ^\\hat\{\\delta\}is, therefore, consistent with an exact affine lift on the full domain, which a penalized fit to a subset need not recover, and the two quantities are reported side by side\. The probe is a logistic regression with unit inverse regularization and balanced class weights, its probabilities cross\-fitted over stratified five\-fold splits and scored by the area under the curve\. Ranks are taken at the default tolerance of the numerical library, the largest singular value times the larger matrix dimension times machine precision; principal angles are the arccosines of the singular values ofQU⊤​QWQ\_\{U\}^\{\\top\}Q\_\{W\}, withQUQ\_\{U\}andQWQ\_\{W\}orthonormal bases ofrow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\)androw⁡\[𝐓;𝟏\]\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]from singular value decompositions at a relative cutoff of10−1010^\{\-10\}\. Table[2](https://arxiv.org/html/2609.18047#S8.T2)gives medians over predicates\.

Table 2:Median full\-domain affine defectδ\+\\delta^\{\+\}, with its minimum over predicates in parentheses; median held\-out errorδ^\\hat\{\\delta\}, with Gaussian null; median probe AUC; and the second through tenth principal angles in degrees betweenrow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\)androw⁡\[𝐓;𝟏\]\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]\(null given in parentheses\)\. The first principal angle is1\.11\.1,2\.12\.1,2\.12\.1, and5\.25\.2degrees; the angle between the𝟏\\mathbf\{1\}row itself androw⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\)is1\.31\.3,2\.62\.6,2\.12\.1, and5\.45\.4degrees, so the first principal direction lies near the𝟏\\mathbf\{1\}row\.No predicate admits an exact affine lift in any of the four cells: the median full\-domain defect ranges from0\.510\.51to0\.890\.89\(see Table[2](https://arxiv.org/html/2609.18047#S8.T2)\), and the smallest over all predicates is0\.250\.25\. The embeddings, nevertheless, have lower defects than Gaussian controls\. For a centered truth row and a randomdd\-dimensional subspace of𝟏⟂\\mathbf\{1\}^\{\\perp\}, the expected squared defect is1−d/\(V−1\)1\-d/\(V\-1\), and the observed control medians match its square root to two decimals\.

Jointly over the lexicon, the principal angles give the same account of the obstruction: the lexical row space lies closer torow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\)than the Gaussian controls do, and the constant row is nearly contained in it, while the smallest principal angle stays positive in every cell, beyond numerical tolerance\. Writing

U=row⁡\(𝐇\),S=row⁡\[𝐓;𝟏\],U=\\operatorname\{row\}\(\\mathbf\{H\}\),\\qquad S=\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\],we, therefore, have

U∩S=\{0\},\(U\+span⁡\{𝟏\}\)∩S=span⁡\{𝟏\}\.U\\cap S=\\\{0\\\},\\qquad\\bigl\(U\+\\operatorname\{span\}\\\{\\mathbf\{1\}\\\}\\bigr\)\\cap S=\\operatorname\{span\}\\\{\\mathbf\{1\}\\\}\.The second equality follows because𝟏∈S\\mathbf\{1\}\\in S: ifu\+c​𝟏∈Su\+c\\mathbf\{1\}\\in Swithu∈Uu\\in U, thenu∈U∩Su\\in U\\cap S\. Thus no predicate of the lexicon admits an exact linear readout, and none other than a constant one an exact affine readout, agreeing with the individual defects\.

Under the held\-out error, recovery is imperfect as well: only one McRae predicate,musical instrument, hasδ^<0\.5\\hat\{\\delta\}<0\.5, and none does on WordNet\. Held\-out error and probe AUC are strongly negatively correlated on McRae \(ρ=−0\.86\\rho=\-0\.86and−0\.82\-0\.82\), so predicates with better discrimination generally have lower prediction error\. High AUC, however, establishes neither exact affine recovery nor strict separability\.

We test separability directly by seeking𝐰\\mathbf\{w\}andbbsuch that

𝐰​h​\(di\)\+b≥1if​P​\(di\)=1,𝐰​h​\(di\)\+b≤−1otherwise\.\\mathbf\{w\}h\(d\_\{i\}\)\+b\\geq 1\\quad\\text\{if \}P\(d\_\{i\}\)=1,\\qquad\\mathbf\{w\}h\(d\_\{i\}\)\+b\\leq\-1\\quad\\text\{otherwise\}\.Every McRae predicate is separable in both embeddings and their Gaussian controls, so this test does not distinguish the geometries there\. On WordNet, GloVe separates159159of171171predicates and word2vec136136of143143, compared with8787and8080in the controls\.

Size accounts for the controls, which separate no predicate above6363members in the GloVe cell or6161in the word2vec cell\. The geometries separate every predicate up to100100members,1919of2222and2121of2424between100100and200200, and33of1212and33of77above; among the predicates larger than any sampled control managed,6262of7474over GloVe and5353of6060over word2vec are half\-spaces\. Failures includeaction,activity, andcontent, whereascity,municipality, andurban areaare separable in both embeddings\.Urban areagives a direct instance of Proposition[8\.2](https://arxiv.org/html/2609.18047#S8.Thmproposition2), strictly separable at affine defects of0\.680\.68and0\.540\.54;districtandadministrative districtgive the converse, the two predicates with the lowest held\-out error over GloVe being among its failures\.

Replacing first\-sense WordNet labels with monosemous assignments raises median AUC from0\.890\.89to0\.950\.95, and lowers held\-out error only from0\.940\.94to0\.930\.93\. Retaining progressively more principal directions of𝐇\\mathbf\{H\}lowers held\-out error gradually, without a pronounced transition at the lexicon’s rank, which Corollary[5\.1](https://arxiv.org/html/2609.18047#S5.Thmcorollary1)permits: the bound guarantees that some geometry of sufficient dimension carries the lexicon, not that a truncation of a pretrained geometry does\.

Recovery varies by predicate type as well: on McRae, taxonomic predicates have median held\-out errors of0\.650\.65and0\.710\.71, against0\.870\.87and0\.900\.90for attributive predicates, with color and size worst; on WordNet, administrative and geographic categories and substance nouns are predicted best, and abstract nouns such asideaandinformationworst\.666This agrees with\[[29](https://arxiv.org/html/2609.18047#bib.bib24)\], who found better distributional recovery of taxonomic than attributive properties\.

Figure 11:Per\-predicate held\-out errorδ^\\hat\{\\delta\}under the isotonic readout against cross\-validated probe AUC, drawn from the per\-predicate tables\. Predicates concentrate at high AUC and high error: discriminable and poorly recovered\. The lower right corner, low error at high AUC, is nearly empty\.Monotone transformations of the affine score lower the held\-out error to between0\.770\.77and0\.860\.86, against Gaussian controls near1\.001\.00, a gain over the ridge readout of0\.030\.03to0\.060\.06on McRae and0\.100\.10to0\.150\.15on WordNet\. Both are fit within training folds, isotonic regression on cross\-fitted training scores\.

Recovery is much better in a small tail: formusical instrument, on eighteen concepts, isotonic error falls to0\.160\.16and0\.050\.05, and birds and their parts, fruit, cities, and countries follow between0\.400\.40and0\.600\.60, at AUC above0\.980\.98\. The tenth percentile of the isotonic error lies between0\.540\.54and0\.680\.68, so substantial error remains for most predicates\.

A two\-layer readout with 64 hidden units \(validated on synthetic data requiring nonlinear recovery\) raises median error on McRae by0\.030\.03and0\.040\.04relative to isotonic regression, and lowers it on WordNet by0\.060\.06and0\.070\.07, to0\.770\.77and0\.750\.75, while its AUC does not exceed the logistic probe’s\. These are results about held\-out entities: on the finite domain itself, distinct entity vectors permit recovery of every predicate by an unrestricted decoder\.

No predicate of either lexicon admits an exact linear or affine lift, so the geometries lie outside the exact regime of Theorem[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1)\. Many predicates are, nevertheless, strictly separable, placing the WordNet cells inside the thresholded regime of Proposition[8\.2](https://arxiv.org/html/2609.18047#S8.Thmproposition2)well beyond the sampled controls, and the McRae cells inside it at a rate the controls match\. Held\-out recovery varies across predicates and remains imperfect under every readout family examined\. Exact representability, separability, and predictive performance, therefore, give distinct assessments of one geometry\.

### 8\.5Training toward exact regime

Whether the exact regime is reachable at the dimensions current embeddings use, and at what distributional cost, is the question of this section\.

We train a geometry𝐇∈ℝd×V\\mathbf\{H\}\\in\\mathbb\{R\}^\{d\\times V\}under a mixed objective\. Let𝐄∈ℝ300×V\\mathbf\{E\}\\in\\mathbb\{R\}^\{300\\times V\}be the pretrained geometry with its row means removed, and𝐓c\\mathbf\{T\}\_\{c\}the truth matrix with its row means removed\. For readouts𝐖∈ℝ300×d\\mathbf\{W\}\\in\\mathbb\{R\}^\{300\\times d\}and𝐖T∈ℝ\|𝒫\|×d\\mathbf\{W\}\_\{T\}\\in\\mathbb\{R\}^\{\|\\mathcal\{P\}\|\\times d\}and an intercept𝐰0∈ℝ\|𝒫\|\\mathbf\{w\}\_\{0\}\\in\\mathbb\{R\}^\{\|\\mathcal\{P\}\|\},

ℒ⁡\(𝐇\)=\(1−λ\)​min𝐖​∥𝐖𝐇−𝐄∥F2∥𝐄∥F2\+λ​min𝐖T,𝐰0​∥𝐖T​𝐇\+𝐰0​𝟏⊤−𝐓∥F2∥𝐓c∥F2,\\mathcal\{L\}\(\\mathbf\{H\}\)=\(1\-\\lambda\)\\,\\min\_\{\\mathbf\{W\}\}\\frac\{\\lVert\\mathbf\{W\}\\mathbf\{H\}\-\\mathbf\{E\}\\rVert\_\{F\}^\{2\}\}\{\\lVert\\mathbf\{E\}\\rVert\_\{F\}^\{2\}\}\+\\lambda\\,\\min\_\{\\mathbf\{W\}\_\{T\},\\mathbf\{w\}\_\{0\}\}\\frac\{\\lVert\\mathbf\{W\}\_\{T\}\\mathbf\{H\}\+\\mathbf\{w\}\_\{0\}\\mathbf\{1\}^\{\\top\}\-\\mathbf\{T\}\\rVert\_\{F\}^\{2\}\}\{\\lVert\\mathbf\{T\}\_\{c\}\\rVert\_\{F\}^\{2\}\},\(4\)so the distributional term is the squared relative error of reconstructing the pretrained vectors linearly from𝐇\\mathbf\{H\}, and the truth term is the squared relative residual of the affine lift of Proposition[5\.1](https://arxiv.org/html/2609.18047#S5.Thmproposition1), taken jointly over the lexicon\. The objective is minimized in these squared terms; the numbers reported below, and plotted in Figures[12](https://arxiv.org/html/2609.18047#S8.F12)and[13](https://arxiv.org/html/2609.18047#S8.F13), are their square roots, writtenϵ⁡\(𝐇\)\\epsilon\(\\mathbf\{H\}\)for the distributional reconstruction error andδ𝒫​\(𝐇\)\\delta\_\{\\mathcal\{P\}\}\(\\mathbf\{H\}\)for the joint truth defect, so that both are relative norms on the scale ofδ\+\\delta^\{\+\}\. Both inner minima are least\-squares problems with closed\-form solutions, andℒ\\mathcal\{L\}is minimized by alternating least squares between𝐇\\mathbf\{H\}and the readouts, with a ridge of10−810^\{\-8\}on the𝐇\\mathbf\{H\}solve and the rows of𝐇\\mathbf\{H\}renormalized after each step, sinceℒ\\mathcal\{L\}is invariant underG​L​\(d\)GL\(d\)acting on𝐇\\mathbf\{H\}; each run starts from a Gaussian𝐇\\mathbf\{H\}with a fixed seed and takes forty iterations\. Every entity is a column of𝐇\\mathbf\{H\}, so the experiment is transductive, as an embedding layer is\. Atλ=0\\lambda=0, training reconstructs the pretrained geometry; atλ=1\\lambda=1, the truth term alone is minimized, the surplus directions of𝐇\\mathbf\{H\}above the rank are untrained, and the distributional error atλ=1\\lambda=1carries no information, so the frontier is read atλ∈\(0,1\)\\lambda\\in\(0,1\)\.

Two quantities are available in closed form\.

1. 1\.The minimum of the truth term over alldd\-dimensional geometries is\(∑k\>dσk2\)1/2/∥𝐓c∥F\\bigl\(\\sum\_\{k\>d\}\\sigma\_\{k\}^\{2\}\\bigr\)^\{1/2\}/\\lVert\\mathbf\{T\}\_\{c\}\\rVert\_\{F\}for the singular valuesσk\\sigma\_\{k\}of𝐓c\\mathbf\{T\}\_\{c\}, since the intercept absorbs the row means; this floor is positive ford<rank⁡𝐓c=rank⁡\[𝐓;𝟏\]−1d<\\operatorname\{rank\}\\mathbf\{T\}\_\{c\}=\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]\-1by Corollary[5\.5](https://arxiv.org/html/2609.18047#S5.Thmcorollary5), and zero from there on, which is Corollary[5\.1](https://arxiv.org/html/2609.18047#S5.Thmcorollary1)in the affine form of Proposition[5\.1](https://arxiv.org/html/2609.18047#S5.Thmproposition1), and it fixes the location of the exactness transition in advance\.
2. 2\.The second is a linear\-exact benchmark: the variance retained at dimensionddunder the constraintrow⁡\(𝐇\)⊇row⁡\[𝐓;𝟏\]\\operatorname\{row\}\(\\mathbf\{H\}\)\\supseteq\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]of Theorem[5\.1](https://arxiv.org/html/2609.18047#S5.Thmtheorem1)is the fraction of∥𝐄∥F2\\lVert\\mathbf\{E\}\\rVert\_\{F\}^\{2\}captured by the bestdd\-dimensional row space containingrow⁡\[𝐓;𝟏\]\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\], namelyrow⁡\[𝐓;𝟏\]\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\]together with the topd−rank⁡\[𝐓;𝟏\]d\-\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]principal directions of𝐄\\mathbf\{E\}projected off it, and at a least\-squares optimum the retained variance is1−ϵ21\-\\epsilon^\{2\}\.

The benchmark is one dimension more constrained than the training criterion, which is affine, and needsrow⁡\(𝐇\)\\operatorname\{row\}\(\\mathbf\{H\}\)to contain a complement of𝟏\\mathbf\{1\}inrow⁡\[𝐓;𝟏\]\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\], such asrow⁡\(𝐓c\)\\operatorname\{row\}\(\\mathbf\{T\}\_\{c\}\); the affine benchmark, withrow⁡\(𝐓c\)\\operatorname\{row\}\(\\mathbf\{T\}\_\{c\}\)in place ofrow⁡\[𝐓;𝟏\]\\operatorname\{row\}\[\\mathbf\{T\};\\mathbf\{1\}\], retains at least as much, and the two differ by at most the variance of one direction\.

The grid runs overd∈\{10,20,40,80,120,160,200,300\}d\\in\\\{10,20,40,80,120,160,200,300\\\}andλ∈\{0,0\.1,0\.5,0\.9,1\}\\lambda\\in\\\{0,0\.1,0\.5,0\.9,1\\\}\. A predicate counts as recovered within tolerance when its affine defect on the trained geometry,∥𝐰P​𝐇\+bP​𝟏⊤−𝐭P∥2/∥𝐭P−t¯P​𝟏∥2\\lVert\\mathbf\{w\}\_\{P\}\\mathbf\{H\}\+b\_\{P\}\\mathbf\{1\}^\{\\top\}\-\\mathbf\{t\}\_\{P\}\\rVert\_\{2\}/\\lVert\\mathbf\{t\}\_\{P\}\-\\bar\{t\}\_\{P\}\\mathbf\{1\}\\rVert\_\{2\}with the readout refit by least squares, is below0\.050\.05; Figure[12](https://arxiv.org/html/2609.18047#S8.F12)plots this fraction\. The McRae cell with word2vec in this was matched to the embedding by the training loader, which looks tokens up as given, whereas the diagnostics loader of Section[8\.4](https://arxiv.org/html/2609.18047#S8.SS4)adds a case fallback; that holdsV=529V=529, one predicate falls below fifteen positives, and it trains on7575predicates withrank⁡\[𝐓;𝟏\]=76\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]=76, while Table[2](https://arxiv.org/html/2609.18047#S8.T2)reports the diagnostics cell with7676predicates and rank7777; the dotted line of Figure[12](https://arxiv.org/html/2609.18047#S8.F12)for that cell sits at7575accordingly\.

Atλ=1\\lambda=1, the optimizer attains the closed\-form floor ofδ𝒫\\delta\_\{\\mathcal\{P\}\}at every grid point: to four decimals where the floor is positive \(on McRae0\.7020\.702,0\.5540\.554, and0\.3500\.350atd=10,20,40d=10,20,40; on WordNet with GloVe0\.0140\.014atd=160d=160; with word2vec0\.0580\.058atd=120d=120\), and to10−1310^\{\-13\}where the floor is zero, which is every grid point at or aboverank⁡\[𝐓;𝟏\]−1\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]\-1\. The affine exact regime is, therefore, reached at numerical tolerance atd=80d=80and above on McRae, atd=200d=200and above on WordNet with GloVe, and atd=160d=160and above with word2vec, and the location of the transition is the floor’s,rank⁡\[𝐓;𝟏\]−1\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]\-1, with the grid serving to check that the optimizer finds it\. The per\-predicate fraction within tolerance atλ=1\\lambda=1is11at those points,0\.9470\.947atd=160<165d=160<165on WordNet with GloVe, and0\.7060\.706atd=120<138d=120<138with word2vec, which is the approximate carriage a positive floor admits: below the rank no geometry carries the whole lexicon, and most of it can still lie within tolerance\.

Atd=300d=300, the constraint is relatively cheap on McRae and relatively affordable on WordNet: the linear\-exact benchmark retains99\.299\.2and98\.598\.5percent of the pretrained variance on McRae and82\.582\.5and79\.579\.5percent on WordNet, the constraint having rank166166and139139over a spectral tail atV≈6000V\\approx 6000\. On the frontier atλ=0\.9\\lambda=0\.9,δ𝒫\\delta\_\{\\mathcal\{P\}\}is0\.0030\.003and0\.0040\.004on McRae atϵ=0\.085\\epsilon=0\.085and0\.1210\.121\(retention99\.399\.3and98\.598\.5percent\), with every predicate within tolerance; on WordNet it is0\.0460\.046and0\.0440\.044atϵ=0\.365\\epsilon=0\.365and0\.4150\.415\(retention86\.786\.7and82\.882\.8percent\), with0\.750\.75of the predicates within tolerance after forty iterations\.

The affine exact regime is, therefore, reached at numerical tolerance atd=80d=80and above on McRae, atd=200d=200and above on WordNet with GloVe, and atd=160d=160and above with word2vec, which are the grid points at or above the minimal affine dimension of Corollary[5\.5](https://arxiv.org/html/2609.18047#S5.Thmcorollary5)\.

Figure 12:Fraction of predicates with affine defect below0\.050\.05on the trained geometry \(recovered within tolerance\), after alternating least squares atλ=1\\lambda=1, against the trained dimensiond∈\{10,20,40,80,120,160,200,300\}d\\in\\\{10,20,40,80,120,160,200,300\\\}, drawn from the run’s output\. Dotted lines mark the minimal affine dimensionrank⁡\[𝐓;𝟏\]−1\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]\-1of Corollary[5\.5](https://arxiv.org/html/2609.18047#S5.Thmcorollary5):7676and7575\(McRae; the word2vec training cell holds7575predicates, see the text\),165165\(WordNet, GloVe\),138138\(WordNet, word2vec\)\. Wherever the fraction reads11, the joint residual is below10−1310^\{\-13\}; at the two grid points just below that dimension it equals the closed\-form minimum,0\.0140\.014and0\.0580\.058\.Figure 13:The frontier atd=300d=300: joint truth defectδ𝒫\\delta\_\{\\mathcal\{P\}\}against distributional reconstruction errorϵ\\epsilon, both relative, forλ∈\{0\.1,0\.5,0\.9\}\\lambda\\in\\\{0\.1,0\.5,0\.9\\\}after forty iterations, all four cells\. On McRae,δ𝒫\\delta\_\{\\mathcal\{P\}\}reaches0\.0030\.003and0\.0040\.004atϵ=0\.085\\epsilon=0\.085and0\.1210\.121; on WordNet,0\.0460\.046and0\.0440\.044at0\.3650\.365and0\.4150\.415, where the optimizer is still descending atλ=0\.9\\lambda=0\.9\.Reaching the floor on the training entities shows that an embedding layer attains the affine exact regime under the truth objective; the frontier shows that most of the distributional variance survives within tolerance of it; exact realizability atd≥rank⁡\[𝐓;𝟏\]d\\geq\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]is Corollary[5\.1](https://arxiv.org/html/2609.18047#S5.Thmcorollary1), and the experiment shows that the optimizer finds it\. The result is transductive: generalization of the truth\-conditional structure to unseen entities, and acquisition of the same geometry from distributional training alone, remain open\.

## 9Discussion

Consider the distinction between*representability*and*representation*\. The free construction establishes when a truth\-conditional lexicon can be carried exactly by a finite\-dimensional vector space; what follows is the empirical analysis, in which we ask how closely existing representations approach that construction\. The defect introduced above makes this quantitative: zero defect means exact, while positive defect measures difference from it\.

We showed that ordinary distributional embeddings already lie substantially closer to the truth\-conditional geometry than a random subspace of the same dimension; this is consistent with the fact that semantic features are often linearly represented in learned vector spaces\[[22](https://arxiv.org/html/2609.18047#bib.bib19)\]; it also gives a geometric interpretation of superposition\[[5](https://arxiv.org/html/2609.18047#bib.bib18)\]: when the available dimension is below the minimum required for exactness, several truth conditions must share dimensions, and the resulting dependencies appear as nonzero defect\. The principal angles \(as reported in Table[2](https://arxiv.org/html/2609.18047#S8.T2)\), therefore, quantify the extent to which a distributional geometry already contains the structure required by the lexicon, as opposed to mere generic similarity\.

Distributional learning, by itself, leaves every predicate of both lexicons outside the exact regime, by the affine defects and the principal angles, and the held\-out error orders the predicates: concrete category predicates are recovered far better than attributive and abstract predicates, and monotone \(or nonlinear\) readouts recover much of the same ordering\. This is compatible with the literature on conceptual spaces\[[8](https://arxiv.org/html/2609.18047#bib.bib9)\], while also placing a limit on a purely geometric account: a region in a conceptual space need not constitute an exact extension\. Human categorization provides an independent reason not to expect exact linear separability as a universal property\[[16](https://arxiv.org/html/2609.18047#bib.bib21)\]\.

Atd=300d=300, the truth\-conditional constraint can be imposed, while preserving most of the variance of the original embeddings\. On McRae, the lexicon is brought within tolerance with very little distributional loss; on WordNet, the constraint is more expensive, and remains compatible with substantial retention\. Under the truth term alone, the optimizer reaches the affine exact regime at numerical tolerance at every trained dimension fromrank⁡\[𝐓;𝟏\]−1\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]\-1upward, where the closed\-form floor is zero, and matches the positive floor below it\. Dimension is thereby eliminated as the reason the pretrained embeddings fail to carry the lexicon; the objective, the corpus, and the labels remain as candidates, which the present experiments leave apart from one another\.

This has a useful consequence for the relation between distributional and truth\-conditional semantics: the two need not compete for representation space\. A single geometry can retain substantial distributional structure, while also carrying a truth\-conditional lexicon\. The experiments leave open whether language exposure alone produces such a geometry\. Skip\-gram’s relation to shifted PMI factorization\[[14](https://arxiv.org/html/2609.18047#bib.bib14)\]gives the distributional geometry a corpus\-level interpretation, but there is nothing in that objective that requires the resulting space to satisfy the truth\-conditional constraints\. The present results, therefore, support a weaker \(and more precise\) claim:*distributional learning supplies information from which truth\-conditional structure may be \(partially\) recovered; exactness requires either an additional constraint or some mechanism that supplies equivalent information\.*

The framework also clarifies the status of compositionality\. Once the leaves of a derivation are represented exactly, the homomorphism conditions determine the corresponding Boolean composition without further learning\. Outside the exact regime, each leaf is only approximately represented, so compositional error can accumulate\. A system may, therefore, perform well on individual semantic probes, while failing a composed entailment\. Good distributional similarity alone, consequently, leaves exact logical behavior undetermined\.

The empirical scope of the present study is deliberately narrower than the formal framework we are likewise developing; here, the measurements concern one\-place predicates over static entity geometries\. Relations, represented bilinearly on𝐇⊗𝐇\\mathbf\{H\}\\otimes\\mathbf\{H\}, and higher arities follow the same geometric strategy, but were not evaluated here, which we leave for future work\. Quantifiers are treated in Proposition[5\.3](https://arxiv.org/html/2609.18047#S5.Thmproposition3)for the exact regime, where they compute by bilinear accumulation on the readouts, and their behavior under positive defect remains to be measured\. Likewise, the formal conditions governing computation after the embedding layer raise a separate question\. Proposition[5\.2](https://arxiv.org/html/2609.18047#S5.Thmproposition2)confines the present compression result to the leaves of a derivation; whether attention and feed\-forward computation can, themselves, realize the required multilinear maps remains open\.

The vector logic supplies a specification of what a representation carrying a truth\-conditional structure has to satisfy, and the defect turns the specification into a measurable property of an empirical geometry, so that the question whether a learned representation carries such a structure has a computable answer\. That is the principal role of the framework; the origin of the gradient observed here, and the emergence of truth\-conditional structure from language exposure alone, lie outside what it decides\.

## 10Conclusion

A truth\-conditional lexicon imposes linear constraints on the representation space; the rank of those constraints gives the minimum dimension for exactness, while the defect measures how closely a lower\-dimensional \(or otherwise unconstrained\) representation approaches that ideal\. The empirical results show that existing distributional embeddings contain substantial structure relevant to the lexicon, while every predicate of both lexicons fails the exact criterion, linear and affine, over both geometries\. Once the truth\-conditional constraint enters the objective, the affine exact regime is attained at numerical tolerance from the predicted dimensionrank⁡\[𝐓;𝟏\]−1\\operatorname\{rank\}\[\\mathbf\{T\};\\mathbf\{1\}\]\-1upward, and, atd=300d=300, the lexicon comes within tolerance of it while most of the original distributional variance is retained; distributional training, by itself, leaves exactness to a further constraint\.

The formal framework extends beyond the monadic case studied here\. Relations and higher\-arity predicates can be treated on tensor\-product spaces, quantifiers require corresponding operators on compressed readouts, and contextual representations can be evaluated layer by layer\. The most direct empirical continuation is to replace the pretrained distributional matrix in the mixed objective of Section[8\.5](https://arxiv.org/html/2609.18047#S8.SS5)with corpus statistics to test whether exact truth\-conditional structure can emerge from co\-occurrence information alone, and at what distributional cost; we expect not\.

What does it mean for a vector geometry to carry a truth\-conditional semantics compositionally is answered here for finite monadic lexicons and binary relations; how far does an empirical geometry fall short of that condition is measured for two lexicons and two embeddings; how a learning system might acquire such a representation is another problem\.

## Acknowledgments

The observation that an embedding layer is a linear map on one\-hot inputs, and its pedagogical framing, are due entirely to Sebastian Raschka, which set this paper in motion\.

## References

- \[1\]\(2018\)Understanding intermediate layers using linear classifier probes\.External Links:1610\.01644,[Link](https://arxiv.org/abs/1610.01644)Cited by:[§8\.2](https://arxiv.org/html/2609.18047#S8.SS2.p3.1)\.
- \[2\]C\. Allen and T\. Hospedales\(2019\)Analogies explained: towards understanding word embeddings\.InProceedings of the 36th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.97,pp\. 223–231\.Cited by:[Remark 8\.1](https://arxiv.org/html/2609.18047#S8.Thmremark1.p1.1.1)\.
- \[3\]Y\. Belinkov\(2022\)Probing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.External Links:[Link](https://aclanthology.org/2022.cl-1.7/),[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00422)Cited by:[§1](https://arxiv.org/html/2609.18047#S1.p5.1)\.
- \[4\]G\. Boleda\(2020\)Distributional semantics and linguistic theory\.Annual Review of Linguistics6,pp\. 213–234\.Cited by:[§1](https://arxiv.org/html/2609.18047#S1.p1.1)\.
- \[5\]N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen, R\. Grosse, S\. McCandlish, J\. Kaplan, D\. Amodei, M\. Wattenberg, and C\. Olah\(2022\)Toy models of superposition\.External Links:2209\.10652,[Link](https://arxiv.org/abs/2209.10652)Cited by:[§9](https://arxiv.org/html/2609.18047#S9.p2.1)\.
- \[6\]K\. Ethayarajh, D\. Duvenaud, and G\. Hirst\(2019\)Towards understanding linear word analogies\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,Florence,pp\. 3253–3262\.External Links:[Document](https://dx.doi.org/10.18653/v1/P19-1315)Cited by:[Remark 8\.1](https://arxiv.org/html/2609.18047#S8.Thmremark1.p1.1.1)\.
- \[7\]C\. Fellbaum \(Ed\.\)\(1998\)WordNet: an electronic lexical database\.MIT Press,Cambridge, MA\.Cited by:[§8\.4](https://arxiv.org/html/2609.18047#S8.SS4.p4.1)\.
- \[8\]P\. Gärdenfors\(2000\)Conceptual spaces: the geometry of thought\.The MIT Press\.Cited by:[§8\.2](https://arxiv.org/html/2609.18047#S8.SS2.p1.1),[§9](https://arxiv.org/html/2609.18047#S9.p3.1)\.
- \[9\]I\. Heim and A\. Kratzer\(1998\)Semantics in generative grammar\.Blackwell\.Cited by:[§1](https://arxiv.org/html/2609.18047#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.18047#S2.SS1.p1.1)\.
- \[10\]J\. Hewitt and P\. Liang\(2019\)Designing and interpreting probes with control tasks\.External Links:1909\.03368,[Link](https://arxiv.org/abs/1909.03368)Cited by:[§8\.2](https://arxiv.org/html/2609.18047#S8.SS2.p3.1)\.
- \[11\]E\. L\. Keenan and J\. Stavi\(1986\)A semantic characterization of natural language determiners\.Linguistics and Philosophy9\(3\),pp\. 253–326\.External Links:[Document](https://dx.doi.org/10.1007/BF00630273)Cited by:[§5\.5](https://arxiv.org/html/2609.18047#S5.SS5.p1.1)\.
- \[12\]C\. Kennedy\(2007\)Vagueness and grammar: the semantics of relative and absolute gradable adjectives\.Linguistics and Philosophy30\(1\),pp\. 1–45\.External Links:[Document](https://dx.doi.org/10.1007/s10988-006-9008-0)Cited by:[§8\.2](https://arxiv.org/html/2609.18047#S8.SS2.p1.1)\.
- \[13\]A\. Kratzer\(1991\)Modality\.InSemantics: An International Handbook of Contemporary Research,A\. von Stechow and D\. Wunderlich \(Eds\.\),pp\. 639–650\.Cited by:[§5\.5](https://arxiv.org/html/2609.18047#S5.SS5.p3.1)\.
- \[14\]O\. Levy and Y\. Goldberg\(2014\)Neural word embedding as implicit matrix factorization\.InAdvances in Neural Information Processing Systems,Z\. Ghahramani, M\. Welling, C\. Cortes, N\. Lawrence, and K\. Weinberger \(Eds\.\),Vol\.27\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2014/file/b78666971ceae55a8e87efb7cbfd9ad4-Paper.pdf)Cited by:[§9](https://arxiv.org/html/2609.18047#S9.p5.1)\.
- \[15\]K\. McRae, G\. S\. Cree, M\. S\. Seidenberg, and C\. Mcnorgan\(2005\)Semantic feature production norms for a large set of living and nonliving things\.Behavior Research Methods37\(4\),pp\. 547–559\.External Links:[Document](https://dx.doi.org/10.3758/BF03192726)Cited by:[§8\.1](https://arxiv.org/html/2609.18047#S8.SS1.p2.1),[§8\.4](https://arxiv.org/html/2609.18047#S8.SS4.p2.1)\.
- \[16\]D\. Medin and P\. Schwanenflugel\(1981\)Linear separability in classification learning\.Journal of Experimental Psychology: Human Learning and Memory7\(5\),pp\. 355–368\.Cited by:[§9](https://arxiv.org/html/2609.18047#S9.p3.1)\.
- \[17\]T\. Mikolov, I\. Sutskever, K\. Chen, G\. Corrado, and J\. Dean\(2013\)Distributed representations of words and phrases and their compositionality\.InAdvances in Neural Information Processing Systems 26,pp\. 3111–3119\.Cited by:[§8\.4](https://arxiv.org/html/2609.18047#S8.SS4.p1.1)\.
- \[18\]T\. Mikolov, W\. Yih, and G\. Zweig\(2013\)Linguistic regularities in continuous space word representations\.InProceedings of NAACL\-HLT 2013,Atlanta, GA,pp\. 746–751\.Cited by:[§8\.3](https://arxiv.org/html/2609.18047#S8.SS3.p1.3)\.
- \[19\]E\. Mizraji\(1992\)Vector logics: the matrix\-vector representation of logical calculus\.Fuzzy Sets and Systems50\(2\),pp\. 179–185\.Cited by:[§2\.1](https://arxiv.org/html/2609.18047#S2.SS1.p2.4)\.
- \[20\]R\. Montague\(1974\)English as a formal language\.InFormal Philosophy: Selected Papers of Richard Montague,R\. Thomason \(Ed\.\),pp\. 188–221\.Cited by:[§1](https://arxiv.org/html/2609.18047#S1.p1.1)\.
- \[21\]M\. Nickel, V\. Tresp, and H\. Kriegel\(2011\)A three\-way model for collective learning on multi\-relational data\.InProceedings of the 28th International Conference on International Conference on Machine Learning,ICML’11,Madison, WI, USA,pp\. 809–816\.External Links:ISBN 9781450306195Cited by:[footnote 2](https://arxiv.org/html/2609.18047#footnote2)\.
- \[22\]K\. Park, Y\. J\. Choe, and V\. Veitch\(2024\)The linear representation hypothesis and the geometry of large language models\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 39643–39666\.External Links:[Link](https://proceedings.mlr.press/v235/park24c.html)Cited by:[§9](https://arxiv.org/html/2609.18047#S9.p2.1)\.
- \[23\]J\. Pennington, R\. Socher, and C\. D\. Manning\(2014\)GloVe: global vectors for word representation\.InEmpirical Methods in Natural Language Processing \(EMNLP\),pp\. 1532–1543\.Cited by:[§8\.4](https://arxiv.org/html/2609.18047#S8.SS4.p1.1)\.
- \[24\]PyTorch ContributorsTorch\.nn\.linear\.Note:[https://pytorch\.org/docs/stable/generated/torch\.nn\.Linear\.html](https://pytorch.org/docs/stable/generated/torch.nn.Linear.html)Accessed 2026\-08\-15Cited by:[Remark 3\.2](https://arxiv.org/html/2609.18047#S3.Thmremark2.p1.1.1)\.
- \[25\]D\. Quigley\(2025\)A vector logic for extensional formal semantics\.Journal of Logic, Language and Information34\(5\),pp\. 557–599\.External Links:[Document](https://dx.doi.org/10.1007/s10849-025-09443-x)Cited by:Exact semantic readout from compressed vector representations,[§1](https://arxiv.org/html/2609.18047#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.18047#S2.SS1.p2.2),[§2\.2](https://arxiv.org/html/2609.18047#S2.SS2.p2.1),[§2](https://arxiv.org/html/2609.18047#S2.p1.1)\.
- \[26\]D\. Quigley\(2026\)A vector logic for intensional formal semantics\.Note:Under review, Journal of Logic, Language and Information; arXiv:2602\.02940External Links:2602\.02940,[Link](https://arxiv.org/abs/2602.02940)Cited by:Exact semantic readout from compressed vector representations,[§1](https://arxiv.org/html/2609.18047#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.18047#S2.SS1.p2.2),[§2\.2](https://arxiv.org/html/2609.18047#S2.SS2.p2.1),[§2\.3](https://arxiv.org/html/2609.18047#S2.SS3.p1.1),[§2](https://arxiv.org/html/2609.18047#S2.p1.1),[§5\.5](https://arxiv.org/html/2609.18047#S5.SS5.p3.1),[Remark 5\.1](https://arxiv.org/html/2609.18047#S5.Thmremark1.p1.1.1),[§6](https://arxiv.org/html/2609.18047#S6.p1.1)\.
- \[27\]S\. RaschkaWhy can an embedding layer be interpreted as a linear layer applied to one\-hot encoded tokens?\.Note:[https://sebastianraschka\.com/faq/docs/embedding\-linear\-onehot\.html](https://sebastianraschka.com/faq/docs/embedding-linear-onehot.html)Accessed 2026\-08\-15Cited by:[§1](https://arxiv.org/html/2609.18047#S1.p3.1),[§3\.4](https://arxiv.org/html/2609.18047#S3.SS4.p1.1),[Remark 3\.2](https://arxiv.org/html/2609.18047#S3.Thmremark2.p1.1.1),[§3](https://arxiv.org/html/2609.18047#S3.p1.1)\.
- \[28\]S\. Raschka\(2024\)Build a large language model \(from scratch\)\.Manning\.Cited by:[§3](https://arxiv.org/html/2609.18047#S3.p1.1)\.
- \[29\]D\. Rubinstein, E\. Levi, R\. Schwartz, and A\. Rappoport\(2015\)How well do distributional models capture different types of semantic knowledge?\.InProceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing \(Volume 2: Short Papers\),pp\. 726–730\.Cited by:[footnote 6](https://arxiv.org/html/2609.18047#footnote6)\.
- \[30\]J\. van Benthem\(1986\)Essays in logical semantics\.Studies in Linguistics and Philosophy, Vol\.29,Reidel,Dordrecht\.Cited by:[§5\.5](https://arxiv.org/html/2609.18047#S5.SS5.p1.1)\.
- \[31\]J\. Westphal and J\. Hardy\(2005\)Logic as a vector system\.Journal of Logic and Computation15\(5\),pp\. 751–765\.External Links:[Document](https://dx.doi.org/10.1093/logcom/exi040)Cited by:[§2\.1](https://arxiv.org/html/2609.18047#S2.SS1.p2.4)\.

Similar Articles