Relation Geometry in Semantic Space of Language Models

arXiv cs.CL Papers

Summary

This paper explores how semantic relations are encoded in the geometry of language model semantic spaces, finding that asymmetric relations occupy distinct regions and that lexical information matters more for causal models while contextual information matters more for masked and diffusion models.

arXiv:2607.26762v1 Announce Type: new Abstract: When it comes to generating vector representations of words, current language models are achieving high-quality results. However, what is not known is the extent to which knowledge about semantic relations is represented in the geometry of the semantic spaces created in this way. In order to answer this question, we study the relation geometry of such semantic spaces from three perspectives. We first examine whether words standing in a particular relation to a target word~(called relata) occupy the same region in semantic space, and whether the regions corresponding to different relations are distinct from each other. We then verify to what extent semantic spaces reflect certain well-known properties of relations, such as symmetry, asymmetry, and transitivity. Finally, we consider which information about the target words and relata is more important for relation geometry: their surface forms, or their contexts. We conduct experiments on six semantic relations using causal, masked, and diffusion language models. The results show that relata in asymmetric relations relatively clearly occupy a distinct region in semantic space. Asymmetric relations' properties are only moderately well encoded in the semantic space, yet better than those of symmetric ones. Furthermore, when considering the question which information source has the strongest impact on results amongst the models we evaluated, we find that lexical information tends to be more important for the causal language model, whereas contextual information is more important for the masked and diffusion language models. Our results empirically show that relation geometry is not equally well-represented for all relations in semantic space, suggesting that there is a difference in how well semantic relations might be learned from distributional information alone.
Original Article
View Cached Full Text

Cached at: 07/30/26, 09:59 AM

# Relation Geometry in Semantic Space of Language Models
Source: [https://arxiv.org/abs/2607.26762](https://arxiv.org/abs/2607.26762)
[View PDF](https://arxiv.org/pdf/2607.26762)

> Abstract:When it comes to generating vector representations of words, current language models are achieving high\-quality results\. However, what is not known is the extent to which knowledge about semantic relations is represented in the geometry of the semantic spaces created in this way\. In order to answer this question, we study the relation geometry of such semantic spaces from three perspectives\. We first examine whether words standing in a particular relation to a target word~\(called relata\) occupy the same region in semantic space, and whether the regions corresponding to different relations are distinct from each other\. We then verify to what extent semantic spaces reflect certain well\-known properties of relations, such as symmetry, asymmetry, and transitivity\. Finally, we consider which information about the target words and relata is more important for relation geometry: their surface forms, or their contexts\. We conduct experiments on six semantic relations using causal, masked, and diffusion language models\. The results show that relata in asymmetric relations relatively clearly occupy a distinct region in semantic space\. Asymmetric relations' properties are only moderately well encoded in the semantic space, yet better than those of symmetric ones\. Furthermore, when considering the question which information source has the strongest impact on results amongst the models we evaluated, we find that lexical information tends to be more important for the causal language model, whereas contextual information is more important for the masked and diffusion language models\. Our results empirically show that relation geometry is not equally well\-represented for all relations in semantic space, suggesting that there is a difference in how well semantic relations might be learned from distributional information alone\.

## Submission history

From: Zhihan Cao \[[view email](https://arxiv.org/show-email/f5ba6cde/2607.26762)\] **\[v1\]**Wed, 29 Jul 2026 11:02:47 UTC \(140 KB\)

Similar Articles

Geometry of Semantic Space: Comparative Study of Discrete and Continuous Models

arXiv cs.CL

This paper compares the geometric structures induced by deep learning vector embeddings (CamemBERT) and lexical co-occurrence graph models on the French 'Great National Debate' corpus, finding similar local topology but distinct global organization, highlighting complementarity between the two approaches.

Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders

arXiv cs.LG

This research investigates whether multilingual large language models develop shared internal representations for mathematical reasoning across languages, introducing a novel Geometry-Invariant Sparse Autoencoder (GI-SAE) method. It finds that cross-language feature sharing is model-dependent and that geometric similarity does not consistently imply functional interchangeability.