Geometry of Semantic Space: Comparative Study of Discrete and Continuous Models
Summary
This paper compares the geometric structures induced by deep learning vector embeddings (CamemBERT) and lexical co-occurrence graph models on the French 'Great National Debate' corpus, finding similar local topology but distinct global organization, highlighting complementarity between the two approaches.
View Cached Full Text
Cached at: 06/08/26, 09:22 AM
# Comparative Study of Discrete and Continuous Models
Source: [https://arxiv.org/html/2606.07183](https://arxiv.org/html/2606.07183)
## Geometry of Semantic Space: Comparative Study of Discrete and Continuous Models
Gabriel Bounias1,2Sabine Ploux2 1ISC\-PIF \(Institut des Systemes Complexes de Paris IdF\), CNRS, France 2CAMS \(Centre d’analyse et de mathématique sociales\), CNRS & EHESS, Paris, France gabriel\.bounias@ens\-lyon\.frsabine\.ploux@ehess\.fr
###### Abstract
This work examines the semantic geometry underlying NLP models\. We compare supervised vector embeddings, such as CamemBERT, with lexical co\-occurrence graphs that encode semantic relations more directly\. While transformer\-based embeddings achieve strong performance, their induced geometries often display unsatisfactory distributions\. In contrast, graph\-based models reveal a clearer and more human\-readable organization of meaning\.
We have implemented a methodology that allows us to perform a comparative analysis either based on the structure of the graphs or based on the topology of the embeddings induced by these two approaches\.
The results of the comparison – applied to the French "Great National Debate" corpus a collection of citizen contributions to the public debate– show a similar local topology but a very different overall structure and topology\. Theses findings suggest complementary perspectives between deep supervised models and graph\-based models, considering a new pathway to guide neural architectures toward more stable and interpretable convergence with graphs structures\.
Geometry of Semantic Space: Comparative Study of Discrete and Continuous Models
Gabriel Bounias1,2Sabine Ploux21ISC\-PIF \(Institut des Systemes Complexes de Paris IdF\), CNRS, France2CAMS \(Centre d’analyse et de mathématique sociales\), CNRS & EHESS, Paris, Francegabriel\.bounias@ens\-lyon\.frsabine\.ploux@ehess\.fr
## 1Introduction
Natural languages structure a space of meaning where each unit finds its place based on semantic proximity\. This emerging geometry is an essential marker of how an NLP model reflects the semantic functioning of the language it deals with\. This production of internal geometry reflects a more or less effective understanding of the use of linguistic units\.
In the meaning space representation word, there are two main approaches\. Vector models derived from deep learning \(BERT, GPT, etc\.\) map each word into a large\-dimensional continuous space\. Conversely, graph models rely on discrete structures, where linguistic entities are linked by co\-occurrence or lexical proximity\. This construction makes the emergence of neighborhoods more explicit and allows for direct geometric analysis of meaning relationships\.
We therefore hypothesize that the quality of an NLP model can be evaluated through the geometric consistency of the structures it induces\. Comparing the geometries derived from vector and graph models, applied to the same corpus, highlights their respective limitations, but also their complementarity\. Our goal is to analyze how these structures reflect, or conversely distort, the organization of meaning in language, as conceived by humans\.
## 2Modelling
### 2\.1Corpus and linguistic unit
In this comparative study, we must first agree on a common constructive basis\. To do this, we need to base the construction of models on the analysis of the same corpus\. The analysis is based on the Grand Débat National corpus, derived from the 2019 public consultation that followed the Gilets Jaunes movement\. This corpus, comprising around 10 million sentences, provides sufficient lexical and thematic diversity to explore the dynamics of meaning within a well\-defined discursive context\. This choice helps limit interpretative complexity while ensuring comparability between the resulting semantic geometries\. Both models — the vector\-based and the graph\-based — are applied to the same textual foundation, following a lemmatization step \(using*Simplemma*,Barbaresi \([2023](https://arxiv.org/html/2606.07183#bib.bib2)\)\) to neutralize morphological variations and ensure linguistic consistency throughout the analysis\.
The comparative analysis also requires defining a common semantic unit to ensure the coherence of comparisons\. Rather than working with isolated words prone to contextual polysemy, this study adopts regular co\-occurrence cliques as the linguistic unitJiet al\.\([2003](https://arxiv.org/html/2606.07183#bib.bib4)\)\. A co\-occurrence clique is defined as a set of words that appear together with a higher\-than\-expected probability within the same context, like a sentence skeleton\. This approach captures polysemy more precisely, as each clique represents a contextualized and recurrent configuration of meaning within the corpus\. Co\-occurrences are measured using Pointwise Mutual Information \(PMI\), which quantifies the strength of association between two words beyond chance\. By applying thresholds on frequency and PMI, a graph connecting highly correlated words is constructed, from which maximal cliques are extracted to serve as reference semantic units\. This methodological choice provides strong semantic precision while ensuring a consistent foundation for comparing the geometry of meaning emerging from both paradigms\. We can now begin building the two models on this basis\.
### 2\.2Graph model
In constructing the graph\-based model, the first step is to define a simple yet robust measure to quantify the proximity between cliques of co\-occurrences, i\.e\., the semantic units extracted from the corpus\. We introduce a one\-parameter similarity function, notedss, which combines two complementary dimensions of linguistic proximity: on the one hand, contextual similarity, captured by the regular co\-occurrence of words within similar contexts, and on the other, lexical overlap, captured by shared words between cliques\. Formally, the similarity between two cliquesC1C\_\{1\}andC2C\_\{2\}is defined as:
s\(C1,C2\)=\(1\+λ\|C1∩C2\|max\(\|C1\|,\|C2\|\)\)×\\displaystyle s\(C\_\{1\},C\_\{2\}\)=\\Bigg\(1\+\\lambda\\frac\{\|C\_\{1\}\\cap C\_\{2\}\|\}\{\\max\(\|C\_\{1\}\|,\|C\_\{2\}\|\)\}\\Bigg\)\\times∑\(w1,w2\)∈C1×C2PPMI\(w1,w2\)\|C1\|\|C2\|\\displaystyle\\sum\_\{\(w\_\{1\},w\_\{2\}\)\\in C\_\{1\}\\times C\_\{2\}\}\\frac\{\\operatorname\{PPMI\}\(w\_\{1\},w\_\{2\}\)\}\{\|C\_\{1\}\|\|C\_\{2\}\|\}
This equation simultaneously captures thematic proximity through the cross\-average of Positive PMI values and lexical overlap, via the left\-hand coefficient weighted byλ\\lambda\. The parameterλ\\lambdacontrols the balance between the contextual and lexical components, ensuring that both contribute comparably to the final similarity measure\.
Once this similarity is computed between all pairs of cliques, we construct a graphGcG\_\{c\}whose nodes correspond to the co\-occurrence cliques, and whose edges connect cliques that exceed a similarity threshold onss\.
### 2\.3Supervised vector model
### 2\.4Construction of the continuous model
To model meaning within a continuous semantic space, we rely on CamemBERTMartinet al\.\([2020](https://arxiv.org/html/2606.07183#bib.bib7)\), a RoBERTa\-basedLiuet al\.\([2019](https://arxiv.org/html/2606.07183#bib.bib6)\)Transformer model pre\-trained on 138 GB of French text from the OSCAR corpus using the Masked Language Modeling \(MLM\) objective\. CamemBERT provides a bidirectional encoder architecture, where each token attends to all others within the same sequence\. The model operates with1212attention layers and1212heads per layer, producing contextualized embeddings of dimension768768\. Its tokenizer, SentencePiece, segments words into subword units to better capture morphological and lexical variations specific to French\. This setup ensures that the model handles polysemy and inflectional morphology in a robust and linguistically coherent way\.
We use this model to embed the co\-occurrence cliques extracted from theGrand Débat Nationalcorpus111[https://www\.data\.gouv\.fr/datasets/grand\-debat\-national\-propositions](https://www.data.gouv.fr/datasets/grand-debat-national-propositions)\. Each cliqueCCis represented by averaging the contextual embeddings of its constituent words, computed acrossnpn\_\{p\}sentences in which the clique occurs most prominently\. And each resulting vector𝐯C\\mathbf\{v\}\_\{C\}take place in an high\-dimensional continuous semantic space, where each point corresponds to a meaning\.
An useful remark is that this representation enables the derivation of a graph structureGbG\_\{b\}based on cosine or Euclidean similarity between clique embeddings\. Finally we made procedure to construct 2 types of graphsGc=\(V,Ec,Wc\)G\_\{c\}=\(V,E\_\{c\},W\_\{c\}\)andGb=\(V,Eb,Wb\)G\_\{b\}=\(V,E\_\{b\},W\_\{b\}\):
V=\{C\|\\displaystyle V=\\\{C\\,\|\\,∀\(wi,wj\)∈C2,PMI\(wi,wj\)\>thPMI\},\\displaystyle\\forall\(w\_\{i\},w\_\{j\}\)\\\!\\in\\\!C^\{2\},\\ \\text\{PMI\}\(w\_\{i\},w\_\{j\}\)\\\!\>\\\!th\_\{\\text\{PMI\}\}\\\},Ec\\displaystyle E\_\{c\}=\{\(Ci,Cj\)\|s\(Ci,Cj\)\>thc\},\\displaystyle=\\\{\(C\_\{i\},C\_\{j\}\)\\,\|\\,s\(C\_\{i\},C\_\{j\}\)\\\!\>\\\!th\_\{c\}\\\},\\quadEb\\displaystyle E\_\{b\}=\{\(Ci,Cj\)\|sv\(𝐯𝐂𝐢,𝐯𝐂𝐣\)\>thb\},\\displaystyle=\\\{\(C\_\{i\},C\_\{j\}\)\\,\|\\,s\_\{v\}\(\\mathbf\{v\_\{C\_\{i\}\}\},\\mathbf\{v\_\{C\_\{j\}\}\}\)\\\!\>\\\!th\_\{b\}\\\},Wc\(Ci,Cj\)=norm\(s\(Ci,Cj\)\),\\displaystyle W\_\{c\}\(C\_\{i\},C\_\{j\}\)=\\text\{norm\}\\\!\\big\(s\(C\_\{i\},C\_\{j\}\)\\big\),\\quadWb\(Ci,Cj\)=norm\(sv\(𝐯𝐂𝐢,𝐯𝐂𝐣\)\)\\displaystyle W\_\{b\}\(C\_\{i\},C\_\{j\}\)=\\text\{norm\}\\\!\\big\(s\_\{v\}\(\\mathbf\{v\_\{C\_\{i\}\}\},\\mathbf\{v\_\{C\_\{j\}\}\}\)\\big\)
with :
sv\(𝐯,𝐰\)=\{1−‖𝐯−𝐰‖2dmax,\(euclidean\)𝐯⋅𝐰‖𝐯‖‖𝐰‖,\(cosine\)s\_\{v\}\(\\mathbf\{v\},\\mathbf\{w\}\)=\\begin\{cases\}1\-\\dfrac\{\\\|\\mathbf\{v\}\-\\mathbf\{w\}\\\|\_\{2\}\}\{d\_\{\\max\}\},&\\text\{\(euclidean\)\}\\\\\[4\.0pt\] \\dfrac\{\\mathbf\{v\}\\cdot\\mathbf\{w\}\}\{\\\|\\mathbf\{v\}\\\|\\\|\\mathbf\{w\}\\\|\},&\\text\{\(cosine\)\}\\end\{cases\}\(1\)
## 3Comparison methods
### 3\.1Graph embedding
A first idea to compare semantic content conveyed by those two models is to start from the graph structure and embed it in a vector space\. These procedures are well documented principally because it allows graph structures to be visualized while retaining relational information\. In practice a graph embedding is a application that assigns each node of the graph a vector in ann\-dimensional space\. This type of application gives a set of vectors, each one corresponding to a node ofGcG\_\{c\}so a clique inVV\. The resulting set and its geometry can be directly compared to the emerging geometry in the vector space generated by CamemBERT\. I this study we use 4 kinds on graph embeddings to enforce the comparison :
- •Force directed methodEades \([1984](https://arxiv.org/html/2606.07183#bib.bib14)\), an usual method to visualize graph structure in 2D\. It’s based on energy minimization problem in a physical system by modeling spring links\.
- •Spectral method, based on the eigen\-decomposition of the graph Laplacian\. It embeds nodes by minimizing distances according to the spectrum of the Laplacian, revealing low\-dimensional structures that preserve global connectivity patterns\.
- •Isomap methodTenenbaumet al\.\([2000](https://arxiv.org/html/2606.07183#bib.bib13)\), a technique that preserves geodesic distances between nodes on the manifold induced by the graph\. It use spectral decomposition to recover the intrinsic geometry of the structure\.
- •Node2VecGrover and Leskovec \([2016](https://arxiv.org/html/2606.07183#bib.bib12)\), a random walk method that encode proximity in the emergent metric by analysing relation between nodes in random walks within the graph\. It is a learning method that usually need many dimensions\.
### 3\.2Comparison between graphs structure
Another way to compare the two models is to start from the vector space generated by the CamemBERT\-based procedure and to construct a graphGbG\_\{b\}whose edges reflect the similarity between vectors\. The comparison method applied to the structures of the two graphsGcG\_\{c\}andGbG\_\{b\}uses graph theory tools as clustering or Breath\-first Search \(BFS\) subgraphs for example\.
Clustering enables to highlight lexical fields in the corpus using precise, controllable structural criteria\. This is because cliques are grouped together by construction between cliques used globally in the same contexts\. Thus, by structurally analyzing the groupings within the emerging geometries, we are able to identify areas of meaning to a greater or lesser extent, which will also serve as a criterion for the quality of the models\. Here, we rely in particular on the Infomap methodRosvall and Bergstrom \([2011](https://arxiv.org/html/2606.07183#bib.bib11)\)\. This algorithm detects communities by modeling the flow of information along random walks on the graph\. The principle is to minimize the description length of a random walker’s trajectory using information theory\. In other words, the algorithm seeks a partition of the graph that allows an optimal compression of movements, grouping together nodes that are frequently visited together\. This approach is particularly well suited to our context, as it emphasizes the structural coherence of semantic regions and thus form zone of meaning within the global topology of structures\.
The advantage of these comparison methods is that they allow to switch between well\-documented two ways of representing the semantic geometry\.
## 4Results
In order to carry out these comparisons between the two types of models, we distinguish between the issues related to the scales highlighted by each type of comparison\. We begin by applying the concept of graph embedding to examine the arrangements of meaning at the local level\. Next, we attempt to compare the geometries that emerge at the global level\.
### 4\.1Local structures
#### 4\.1\.1BFS subtrees
Let’s begin by analyzing what emerges from comparing the relevance of local structures generated by both models\. Each model defines, for every clique, a specific semantic neighborhood — discrete in the case ofGcG\_\{c\}, and continuous in the embedding space\. To evaluate how these two geometries align locally, we extract subgraphs fromGcG\_\{c\}using BFS trees of limited depth around a given root cliqueCiC\_\{i\}\. Each BFS subtree thus represents the local semantic environment of that clique, as defined by the discrete co\-occurrence structure\. So we compare in a first place spatial distribution in the space generated by theGcG\_\{c\}embedding and spatial distribution in the vector space generated by CamemBERT using BFS subtrees ofGcG\_\{c\}and vector spaces similarities\.
Figure 1:BFS subgraph \(depth=3\) ofGcG\_\{c\}starting from the cliqueC=C=\(danse, opéra, théâtre\), \(dance, opera, theater\)\) embedded using the force\-directed method\. Each node is colored according to the cosine similarity between the clique of that node andCCin the space generated by the CamemBERT model\. The size of the nodes is inversely proportional to their distance from the cliqueCCin the BFS path\. Note that labels are originally in French but translate here\.Figure[1](https://arxiv.org/html/2606.07183#S4.F1)shows us an example of how the local geometry of the two models is organized\. The graph model shows a decrease in semantic coherence \(due to the metric distance within the embedding\) over a distance consistent with the meaning we give to the semantic proximity between two cliques\. We can highlight, for example, the gradation of semantic proximity that emerges between the cliques \(danse, opéra, théâtre\) \(dance, opera, theater\) and \(danse, musique, peinture\), \(dance, music, painting\) and between \(danse, opéra, théatre\), \(dance, opera, theater\) and \(bibliothèque, musée\), \(library, museum\)\. Indeed, we would expect this decrease in proximity\. For the geometry within the continuous model, we can see from the cosine similarities that the semantic distinction via distance is very quickly stifled\. After about ten cliques, the meaning is averaged out\.
Thus, very locally, the structures generated by the models coincide\. However, the models diverge in terms of their semantic description after a dozen cliques on this example\. The discrete model predicts a longer progression than the graph model, which averages proximities more quickly\.
It is important to note that in this example, we are able to embed the structure ofGcG\_\{c\}in two dimensions because it is an example taken from a BFS with few nodes\. Thus, even if the embedding comes from the complex embedding ofGcG\_\{c\}\(and is most likely impossible to embed “correctly” in two dimensions\), we do not see any overlap with other lexical fields\.
#### 4\.1\.2Local metric correlation analysis
To move beyond the purely visual or two\-dimensional comparison of neighborhoods, we propose a quantitative assessment of the coherence between the two semantic spaces\. For each node within a BFS\-extracted subgraph ofGcG\_\{c\}, we compute two complementary measures relative to the root cliqueCC: \(1\) the metric distance toCCin the graph embedding space, and \(2\) the cosine similarity toCCin the CamemBERT space\. Each node is then represented as a point in a two\-dimensional scatter plot, with its graph\-based distance on thexx\-axis and its BERT\-based similarity on theyy\-axis\. This representation allows us to evaluate the correlation between the two metrics — that is, whether the semantic proximities encoded by the discrete co\-occurrence structure are preserved in the continuous embedding space\.
Figure 2:Correlation between metric distances in the graph embedding space \(x\-axis\) and cosine similarity in the space generated by CamemBERT \(y\-axis\)\. The starting clique is \(banquise , fondre\), \(ice floe, melt\) and the embedding is of type Node2Vec in dimension 50 ofGcG\_\{c\}constructed atthPMI=9th\_\{\\text\{PMI\}\}=9andp=0\.1%p=0\.1\\%\. For the readability of the local aspect of the comparison, only nodes from the BFS path of depth 6 \(∼1100\\sim 1100cliques\) are represented\. The correlation coefficient approximately \-0\.137\. Note that labels are originally in French but translate here\.In this framework, a negative correlation would indicate that the two geometries are globally consistent: nodes that are close in the graph embedding \(smallxx\) should correspond to high similarity in the BERT space \(largeyy\), while nodes farther apart in the graph should gradually exhibit lower similarity\. Conversely, the absence of such correlation would reveal a breakdown of structural coherence between the two models\.
The results exhibit a clear decay of correlation as one moves away from the root cliqueCC\. In the immediate neighborhood \(first BFS shell\), CamemBERT tends to agree with the graph structure: the closest co\-occurrence cliques remain semantically close in the embedding space\. However, this correspondence quickly vanishes beyond a small radius — typically after ten to fifteen nodes as describe the correlation coefficientr=−0\.137r=\-0\.137by it’s small size\. The scatter plots become diffuse, and the correlation between graph\-based distance and embedding\-based similarity collapses toward zero\. This pattern is illustrated in Figure[2](https://arxiv.org/html/2606.07183#S4.F2), where the local consistency around the root clique \(banquise, fondre\) \(ice floe, melt\) is visible for only a limited set of related cliques, such as \(Alpes, glacier\), \(Alpes, glacier\)\. Beyond this local range, however, semantic coherence fades: distant cliques like \(enneigement, montagne\), \(snowfall, mountain\), although structurally connected inGcG\_\{c\}, show no particular similarity in CamemBERT’s representation\.
Those observations reveals a structural limitation of large language model embeddings such as CamemBERT, while they capture local semantic regularities efficiently, they fail to reproduce the gradual and progressive organization of meaning that emerges from discrete co\-occurrence graphs\. The continuous embedding space tends to smooth and flatten semantic distinctions, yielding dense local neighborhoods but no coherent sense of large\-scale organization by this averaging effect\. In contrast,GcG\_\{c\}induces a layered and gradated geometry of lexical proximity, where transitions between fields of meaning occur progressively through overlapping cliques\. This gradual property is largely absent from the continuous model, which instead favors isotropic distributions driven by statistical regularities rather than explicit relational structures\.
### 4\.2Global structure comparison
We now aim to compare the global organization of meaning\. A key finding from the previous section suggests that, for the graph model, the semantic structure can extend gradually beyond local and specific neighborhoods\. Here, we attempt to shed further light on this effect of overall architecture\.
#### 4\.2\.1Embeddings repartition
A first step in gaining an understanding of the overall organization is to examine the distribution of vectors in the embedding spaces\. The occupation of space by the embeddings provides information about the precision and convergence of the algorithms involved\. As for the embedding ofGcG\_\{c\}, analyzing the vector distribution allow us to use dimension as a convergence parameter\. Conceptually, we expect there to be fewer and fewer cliques as proximity increases \(in the sense that there are fewer “paraphrastic” cliques than semantically uncorrelated cliques\)\. This observation is reflected in the embedding space as a decreasing cosine similarity distribution as the cosine tends towards11\. We also expect this distribution to be centered at 0, which would indicate that all directions are being exploited\. Thus, the expected bell curve in this distribution would be a marker of convergence\. This can be seen in Figure[3](https://arxiv.org/html/2606.07183#S4.F3), which shows the distributions for the different types of embeddings ofGcG\_\{c\}in3030dimensional space\. One can see a bell curve centered at0for the Isomap and Force\-Directed algorithms \(referred to as FD, in the following text\), indicating that they have already converged\. One can see that the Node2Vec algorithm does not cover the entire space, and that the spectral algorithm has not converged when space dimension equals3030\. It is important to note that these distributions do not provide information about the quality of the semantic information carried by the graphGcG\_\{c\}, but only give an initial measure of the quality of the embedding, i\.e\., whether the graph embedding, at a given dimension, can reveal its geometry or whether the implementation is correct\. These distributions are a simple verification of the embeddings rather than a measure of semantic content\. One example is an FD\-type embedding which, due to its implementation, have a reasonable curve even in 2D but unsatisfactory semantic content\.
Figure 3:Distribution of cosine similarities between vectors of the four types of embeddings ofGcG\_\{c\}\(constructed atthPMI=9,p=0\.001th\_\{\\text\{PMI\}\}=9,\\;p=0\.001\) in3030\-dimensional space\.Figure 4:Distribution of cosine similarities between pairs of embeddings of cliques from the CamemBERT model\.If we now look at the distribution of cosine similarities across all vectors generated by CamemBERT in dimension768768, shown in Figure[4](https://arxiv.org/html/2606.07183#S4.F4), one can observe the phenomenon described above\. The curve is “centered” around0\.90\.9, which implies the existence of a cone in the vector space where all the vectors of the embedding are concentrated\. This problem, documented in particular for learning models, is known as the Curse of High DimensionalityZhanget al\.\([2025](https://arxiv.org/html/2606.07183#bib.bib10)\)\. The addition of dimensions makes the space so vast that the data is highly dispersed, and so the distances between points lose their discriminating power and exponentially more data is needed to learn effectively\. It then becomes difficult for the loss function to effectively enforce the use of all dimensions\. One can also note a fairly short tail to the right of the peak, which seems to reflect a rapid loss of semantic coherence\.
These spatial organization measures thus provide us with some clues for understanding the convergence of semantic spaces constructions, even if it is difficult to draw any conclusions about semantic content from these analyses\. To do this, we now attempt to compare the graph structures in order to characterize the global distribution from a semantic point of view\.
Figure 5:GcG\_\{c\}\(blue\) andGbG\_\{b\}\(orange\) degree \(top\) and betweenness centrality \(bottom\) distributions\.
#### 4\.2\.2Graph structures comparison
In order to compare the structure of graphsGbG\_\{b\}etGcG\_\{c\}, degree and betweenness centrality distributions were calculated for all nodes \(Figures[5](https://arxiv.org/html/2606.07183#S4.F5)\)\. These distributions reveal two characteristics of the structure of the vector\-based graphGbG\_\{b\}\(in orange\): a few high\-degree nodes connected to a significant portion of the graph’s nodes, and many low\-degree \(<4<4\) nodes that are therefore relatively isolated from the rest of the graph\. The same distributions for clique graphGcG\_\{c\}\(in blue\) show less variation: the peak of the degree histogram, located in the8−128\-12range, is much lower than the peak of the degree histogram forGbG\_\{b\}\. This indicates a smaller number of isolated or sparsely connected nodes\. Furthermore, the maximum degree, which is2525, is much lower than that ofGbG\_\{b\}, which is100100\. Therefore, unlikeGbG\_\{b\},GcG\_\{c\}does not contain a subset of hyper\-connected nodes\. Betweenness centrality distributions, meanwhile, reveals a structural divergence regarding role nodes within these two graphs\. On the one hand,GbG\_\{b\}shows an aggregate of betweenness centralities around low values, indicating that no node is structurally significant for paths within this graph\. On the other hand,GcG\_\{c\}exhibits a wider range of betweenness centrality values, highlighting the existence of cliques that are significant for paths through the graph\. These cliques with high betweenness values serve as bridges between lexical fields\. For example, the clique \(huile,colza,soja\) \(oil, rapeseed, soy\) has a betweenness centrality of0\.080\.08and links, on the one hand, issues related to ecology and, on the other hand, those related to agriculture\. These same centrality values are absent withinGbG\_\{b\}\.
Based on the values of these two indicators, the graphGcG\_\{c\}has a more homogeneous structure in terms of the distribution of links between nodes than the graphGbG\_\{b\}, which contains a few central nodes with high transition power and many peripheral nodes\. This disparity should have an impact on the classification and semantic consistency of the clusters\. We will demonstrate this\.
#### 4\.2\.3Clustering on semantic graphs
A community detection and graph partitioning was applied using the Infomap algorithm\. Figure[6](https://arxiv.org/html/2606.07183#S4.F6)shows the results for each of the graphsGcG\_\{c\}andGbG\_\{b\}, with the clusters colored\. Overlaid on each partition is a histogram of cluster sizes\. We observe strong heterogeneity inGbG\_\{b\}: few large clusters and many small ones\. The histogram forGcG\_\{c\}is flatter, with a maximum size \(181181\) much smaller than that ofGbG\_\{b\}\(480480\), a peak that is not at zero, and a steeper slope\. In short, the distribution of cluster sizes inGcG\_\{c\}is more homogeneous than that ofGbG\_\{b\}\.


Figure 6:Upper,GbG\_\{b\}\(left\) andGcG\_\{c\}\(right\) colored according to the Infomap partition of each graph\. Graphs are constructed atthPMI=9th\_\{\\text\{PMI\}\}=9,p=0\.05%p=0\.05\\%andsv=seucls\_\{\\textbf\{v\}\}=s\_\{\\text\{eucl\}\}, FD\-type embedding\. Lower, the respective distribution of cluster sizes in the associated Infomap partition\.Beyond these measures, which highlight significant structural differences, it is important to gain an understanding of the semantic coherence within the clusters for each of the two graphs\. Table[1](https://arxiv.org/html/2606.07183#S4.T1)illustrates this aspect\. It presents cliques randomly selected from a few of the first clusters in theGbG\_\{b\}andGcG\_\{c\}partitions\. We observe stronger semantic coherence between cliques within the same cluster forGcG\_\{c\}than forGbG\_\{b\}\. For example, in the largest cluster ofGbG\_\{b\}, the pair of cliques \[\(wind power, energy sector, nuclear power\), \(noisy, conversation, music\)\] have little thematic connection\. In contrast, based on this same random selection, one can observe good semantic coherence starting within the largest clusters inGcG\_\{c\}\.
Table 1:Randomly selected cliques extracted from Infomap\-derived clusters \(graphsGcG\_\{c\}andGbG\_\{b\}\) constructed atp=0\.1%p=0\.1\\%\. The first column lists the cluster indices in descending order of size\.
## 5Graph structure and embedding dimensionality
Language models are known for their high dimensionality\. By varying the dimension of the embedding space using the methods described in Section[3\.1](https://arxiv.org/html/2606.07183#S3.SS1), we sought to determine the number of dimensions required for a “good” representation of the calculated clusters\. Figure[7](https://arxiv.org/html/2606.07183#S5.F7)illustrates trustworthiness222The trustworthiness measureTwTwis the average of the number of thekknearest neighbors ofeein the embedding space that belong to the same cluster asee\. It ranges from0to11; a value close to11indicates that the local structure of the graph predicted by the clustering is well preserved in the embedded space\.TwTwdepending on the dimension of the embedding\.
Twk\(E,P\)=1n∑i=1n1k∑j∈𝒩k\(ei\)𝟏\[P\(ei\)=P\(ej\)\]Tw\_\{k\}\(E,P\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\frac\{1\}\{k\}\\sum\_\{j\\in\\mathcal\{N\}\_\{k\}\(e\_\{i\}\)\}\\mathbf\{1\}\_\{\[P\(e\_\{i\}\)=P\(e\_\{j\}\)\]\}whereEEis the set of thennvectors of the embedding,PPis the partition, and𝒩k\(ei\)\\mathcal\{N\}\_\{k\}\(e\_\{i\}\)denotes thekknearest neighbors ofeie\_\{i\}in the vector space\. We observe that the curves converge rapidly to a maximum value at\>0\.7\>0\.7for a dimension<100<100for three of the methods\. This shows that the embedding space of the graph is consistent with the classification while maintaining a reasonable dimensionality\.
Figure 7:Evolution of trustworthiness atk=10k=10on the Infomap partitioning as a function of the dimension of the embedding ofGcG\_\{c\}atp=0\.1%p=0\.1\\%across the four types of embeddings\.
## 6Conclusion
We have developed a method for comparing the geometry of semantic spaces and the structure of the graphs induced by a language model and a regular co\-occurrence clique\-based model\. This comparison was conducted using the same corpus of texts\. The results highlight a consistency in the local semantic similarities between the two types of models\. They also highlight significant differences in geometry and global structure\. Based on this comparison, the graph structure and the geometry of the embeddings generated by the co\-occurrence cliques appear to demonstrate a better gradation of semantic distances, a more evenly distributed use of embedding space, and a more homogeneous graph structure in terms of degree, betweenness, and cluster size\. An assessment of the semantic coherence of the clusters also supports the discrete model\. Finally, the various embeddings tested show good classification performance starting at a few dozen dimensions\.
These findings call for further research\. While language model training and architectures provide a highly effective solution for all generation tasks that primarily require context\-local performance, could they benefit from a better representation of the overall structure of semantic spaces? Would this approach make it possible to reduce the size of the effective dimensions? If these assumptions were to prove true, they would result not only in quantitative and qualitative gains but also in an increase in the amount of required resources\.
## 7Limits
This study was conducted for a single language \(French\), a unique corpus of texts, and a sole language model \(CamenBERT\)\. Its extension to other languages, other corpora, and different models remains to be done\.
## References
- Simplemma\.Zenodo\.Cited by:[§2\.1](https://arxiv.org/html/2606.07183#S2.SS1.p1.1)\.
- P\. Eades \(1984\)A heuristic for graph drawing\.Congressus numerantium42\(11\),pp\. 149–160\.Cited by:[1st item](https://arxiv.org/html/2606.07183#S3.I1.i1.p1.1)\.
- A\. Grover and J\. Leskovec \(2016\)Node2vec: scalable feature learning for networks\.InProceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining,pp\. 855–864\.Cited by:[4th item](https://arxiv.org/html/2606.07183#S3.I1.i4.p1.1)\.
- H\. Ji, S\. Ploux, and E\. Wehrli \(2003\)Lexical knowledge representation with contextonyms\.InProceedings of Machine Translation Summit IX: Papers,Cited by:[§2\.1](https://arxiv.org/html/2606.07183#S2.SS1.p2.1)\.
- Y\. Liu, M\. Ott, N\. Goyal, J\. Du, M\. Joshi, D\. Chen, O\. Levy, M\. Lewis, L\. Zettlemoyer, and V\. Stoyanov \(2019\)Roberta: a robustly optimized bert pretraining approach\.arXiv preprint arXiv:1907\.11692\.Cited by:[§2\.4](https://arxiv.org/html/2606.07183#S2.SS4.p1.3)\.
- L\. Martin, B\. Muller, P\. O\. Suarez, Y\. Dupont, L\. Romary, É\. V\. de La Clergerie, D\. Seddah, and B\. Sagot \(2020\)CamemBERT: a tasty french language model\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 7203–7219\.Cited by:[§2\.4](https://arxiv.org/html/2606.07183#S2.SS4.p1.3)\.
- M\. Rosvall and C\. T\. Bergstrom \(2011\)Multilevel compression of random walks on networks reveals hierarchical organization in large integrated systems\.PloS one6\(4\),pp\. e18209\.Cited by:[§3\.2](https://arxiv.org/html/2606.07183#S3.SS2.p2.1)\.
- J\. B\. Tenenbaum, V\. d\. Silva, and J\. C\. Langford \(2000\)A global geometric framework for nonlinear dimensionality reduction\.science290\(5500\),pp\. 2319–2323\.Cited by:[3rd item](https://arxiv.org/html/2606.07183#S3.I1.i3.p1.1)\.
- S\. Zhang, Z\. You, Y\. Chen, Z\. Wen, Q\. Wang, Z\. Qiu, Y\. Li, and M\. Tan \(2025\)Curse of high dimensionality issue in transformer for long\-context modeling\.arXiv preprint arXiv:2505\.22107\.Cited by:[§4\.2\.1](https://arxiv.org/html/2606.07183#S4.SS2.SSS1.p2.2)\.Similar Articles
Relation Geometry in Semantic Space of Language Models
This paper explores how semantic relations are encoded in the geometry of language model semantic spaces, finding that asymmetric relations occupy distinct regions and that lexical information matters more for causal models while contextual information matters more for masked and diffusion models.
The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture
Presents a continuous geometric framework modeling Transformer operations as integro-differential equations on a semantic fiber bundle, validated across multiple architectures.
Analysing the Linearity of Linguistic Relations in Language Model Embedding Spaces
This paper proposes a framework to analyze how linguistic relations are linearly encoded in language model embeddings, revealing differences across models like GloVe, RoBERTa, and ModernBERT and relation types.
Psychological Constructs in Shared Semantic Space
This paper proposes a framework using Supervised Semantic Differential to represent psychological constructs as directions in a shared word-embedding space, enabling comparison across different measurement instruments and research traditions.
Semantic Space of Parts of Speech
This paper uses word2vec embeddings and neural networks to analyze the inherent fuzziness in parts of speech categorization, creating a 3D semantic space to visualize prototypical words and boundaries between linguistic categories.