Do General NLP Embeddings Capture Ontological Reasoning?

arXiv cs.CL Papers

Summary

The paper introduces AVA, a framework to evaluate NLP embeddings' ability to capture ontological reasoning, finding significant limitations and challenging the assumption that strong NLP performance translates to Semantic Web competence.

arXiv:2609.00177v1 Announce Type: new Abstract: General-purpose NLP embedding models perform well on linguistic tasks, but their ability to capture symbolic ontological structure remains unclear. We introduce AVA, a systematic framework for evaluating whether embeddings distinguish logic-sensitive relational semantics in ontologies and knowledge graphs. AVA comprises 171,007 contrastive triplets derived from 163 heterogeneous ontologies using hierarchy inversion, relation substitution, and disjointness injection. Each triplet contains an ontology statement, a semantically equivalent paraphrase, and a logic-sensitive hard negative with contradictory relational meaning. We evaluate more than 25 state-of-the-art embedding models and find substantial limitations: the best model achieves only 0.739 triplet accuracy, while hard negative accuracy falls to 0.135. Fine-tuning improves discrimination by a large margin but transfers poorly to downstream Semantic Web tasks, including taxonomy discovery and ontology alignment. Further analysis suggests that improvements stem partly from perturbation-specific pattern recognition rather than robust ontological understanding. These findings reveal a persistent gap between linguistic representation learning and ontology-level discrimination, challenging the assumption that strong NLP benchmark performance translates to Semantic Web competence.
Original Article
View Cached Full Text

Cached at: 09/02/26, 05:46 AM

# Do General NLP Embeddings Capture Ontological Reasoning?
Source: [https://arxiv.org/html/2609.00177](https://arxiv.org/html/2609.00177)
Conference:Proceedings of the 35th ACM International Conference on Information and Knowledge Management; November 07–11, 2026; Rome, ItalyProceedings of the 35th ACM International Conference on Information and Knowledge Management \(CIKM ’26\), November 07–11, 2026, Rome, ItalyDOI:[10\.1145/3799682\.3840036](https://doi.org/10.1145/3799682.3840036)ISBN:979\-8\-4007\-2539\-5/2026/11CCS:Information systems Language modelsCCS:Computing methodologies Information extractionCCS:Computing methodologies Lexical semanticsCCS:Computing methodologies Ontology engineeringCCS:Computing methodologies Supervised learningCCS:Information systems Similarity measuresCCS:Information systems Retrieval effectivenessHamed Babaei Giglou[https://orcid.org/0000-0003-3758-1454](https://orcid.org/0000-0003-3758-1454)email:[hamed\.babaei@tib\.eu](mailto:[email protected])Affiliation:TIB Leibniz Information Centre for Science and Technology,Hannover,Lower Saxony,GermanyJennifer D’Souza[https://orcid.org/0000-0002-6616-9509](https://orcid.org/0000-0002-6616-9509)email:[jennifer\.dsouza@tib\.eu](mailto:[email protected])Affiliation:TIB Leibniz Information Centre for Science and Technology,Hannover,Lower Saxony,GermanyandSören Auer[https://orcid.org/0000-0002-0698-2864](https://orcid.org/0000-0002-0698-2864)email:[auer@tib\.eu](mailto:[email protected])Affiliation:TIB Leibniz Information Centre for Science and Technology,L3S Research Center, Leibniz University of Hannover,Hannover,Lower Saxony,Germany

Received 5 June 2009

###### Abstract\.

General\-purpose NLP embedding models perform well on linguistic tasks, but their ability to capture symbolic ontological structure remains unclear\. We introduce AVA, a systematic framework for evaluating whether embeddings distinguish logic\-sensitive relational semantics in ontologies and knowledge graphs\. AVA comprises 171,007 contrastive triplets derived from 163 heterogeneous ontologies using hierarchy inversion, relation substitution, and disjointness injection\. Each triplet contains an ontology statement, a semantically equivalent paraphrase, and a logic\-sensitive hard negative with contradictory relational meaning\. We evaluate more than 25 state\-of\-the\-art embedding models and find substantial limitations: the best model achieves only 0\.739 triplet accuracy, while hard negative accuracy falls to 0\.135\. Fine\-tuning improves discrimination by a large margin but transfers poorly to downstream Semantic Web tasks, including taxonomy discovery and ontology alignment\. Further analysis suggests that improvements stem partly from perturbation\-specific pattern recognition rather than robust ontological understanding\. These findings reveal a persistent gap between linguistic representation learning and ontology\-level discrimination, challenging the assumption that strong NLP benchmark performance translates to Semantic Web competence\.

###### Keywords:

Embedding, Large Language Models, Ontology, Semantic Textual Similarity, Ontology Engineering

††cc\-license:by## 1\.Introduction

Recent advances in representation learning have transformed both Natural Language Processing \(NLP\) and the Semantic Web\. General\-purpose embedding models, ranging from Word2Vec\([Mikolov et al\., 2013](https://arxiv.org/html/2609.00177#bib.bib25)\)and GloVe\([Pennington et al\., 2014](https://arxiv.org/html/2609.00177#bib.bib28)\)to transformer\-based architectures such as BERT\([Devlin et al\., 2019](https://arxiv.org/html/2609.00177#bib.bib12)\), MPNet\([Song et al\., 2020](https://arxiv.org/html/2609.00177#bib.bib33)\), E5\([Wang et al\., 2022](https://arxiv.org/html/2609.00177#bib.bib37)\), GTE\([Zhang et al\., 2024](https://arxiv.org/html/2609.00177#bib.bib41);[Li et al\., 2023](https://arxiv.org/html/2609.00177#bib.bib23)\), and BGE\([Xiao et al\., 2024](https://arxiv.org/html/2609.00177#bib.bib39)\), achieve strong performance on retrieval and Semantic Textual Similarity \(STS\) benchmarks\. In parallel, the Semantic Web community has developed ontology and knowledge graph embedding approaches, including RDF2Vec\([Ristoski and Paulheim, 2016](https://arxiv.org/html/2609.00177#bib.bib32)\), OWL2Vec\*\([Chen et al\., 2021b](https://arxiv.org/html/2609.00177#bib.bib8)\), DL2Vec\([Chen et al\., 2021a](https://arxiv.org/html/2609.00177#bib.bib7)\), and knowledge graph embedding \(KGE\) methods such as TransE\([Bordes et al\., 2013](https://arxiv.org/html/2609.00177#bib.bib6)\), ComplEx\([Trouillon et al\., 2016](https://arxiv.org/html/2609.00177#bib.bib36)\), and RotatE\([Sun et al\., 2019](https://arxiv.org/html/2609.00177#bib.bib35)\)\. More recent approaches, including KG\-BERT\([Yao et al\., 2019](https://arxiv.org/html/2609.00177#bib.bib40)\)and KEPLER\([Wang et al\., 2021](https://arxiv.org/html/2609.00177#bib.bib38)\), attempt to combine linguistic representations with structured knowledge\. Despite these advances, a fundamental challenge remains unresolved:*generalization in Semantic Web representation learning*\. Traditional KGE methods operate on fixed entity and relation vocabularies and often require retraining when applied to new ontologies\([Chen et al\., 2023](https://arxiv.org/html/2609.00177#bib.bib9);[Liang et al\., 2023](https://arxiv.org/html/2609.00177#bib.bib24)\)\. Ontology\-aware methods capture rich structural semantics within a particular ontology but may exhibit limited transferability across heterogeneous schemas\([Qiang, 2023](https://arxiv.org/html/2609.00177#bib.bib29);[Bian, 2025](https://arxiv.org/html/2609.00177#bib.bib5);[Li et al\., 2025](https://arxiv.org/html/2609.00177#bib.bib22)\)\. Conversely, transformer\-based sentence embeddings generalize well across linguistic tasks yet are not explicitly optimized for symbolic reasoning over OWL/RDFS structures\. Hyperbolic representations provide a promising alternative for hierarchical knowledge\([Nickel and Kiela, 2017](https://arxiv.org/html/2609.00177#bib.bib27);[Dhingra et al\., 2018](https://arxiv.org/html/2609.00177#bib.bib13)\), but their effectiveness for transferable ontology relation discrimination remains unclear\.

This challenge is closely related to STS, a primary evaluation paradigm for sentence embeddings\. Modern embedding models achieve strong STS performance\([Kumar et al\., 2025](https://arxiv.org/html/2609.00177#bib.bib20)\), suggesting they can capture semantic equivalence between textual expressions\. However, standard STS benchmarks rarely require distinguishing between statements that are lexically similar yet differ in ontology\-level relational semantics\([Gatto et al\., 2023](https://arxiv.org/html/2609.00177#bib.bib16);[Sun et al\., 2025](https://arxiv.org/html/2609.00177#bib.bib34)\)\. As a result, it remains unclear whether strong STS performance reflects genuine sensitivity to subclass relations, domain/range constraints, disjointness axioms, and other forms of symbolic knowledge\. To investigate this question, we introduceAVA, a logic\-sensitive ontology similarity benchmark based on structured ontology perturbations\. AVA generates ontology\-aware hard negatives through hierarchy inversion, relation substitution, and disjointness injection, producing sentence pairs that preserve substantial lexical overlap while expressing contradictory ontological meaning\. Using 163 heterogeneous ontologies, we construct a dataset of 171,007 contrastive triplets and evaluate both pre\-trained and fine\-tuned embedding models under cross\-ontology generalization settings\.

Our experiments reveal three key findings\. First, even the strongest general\-purpose embeddings achieve only moderate performance on ontology\-sensitive similarity judgments, with substantial degradation on hard negatives\. Second, contrastive fine\-tuning dramatically improves triplet discrimination, with hyperbolic objectives achieving near\-perfect ranking accuracy\. Third, these improvements transfer only weakly to downstream ontology engineering tasks such as taxonomy discovery\([Babaei Giglou et al\., 2023](https://arxiv.org/html/2609.00177#bib.bib2)\)and ontology alignment\([Hertling and Paulheim, 2023](https://arxiv.org/html/2609.00177#bib.bib18)\)\. Together, these results reveal an optimization–generalization gap:*embeddings can learn to discriminate AVA perturbations without acquiring robust, transferable representations of ontological structure and pattern\-specific discrimination rather than general ontological reasoning*\.

The contributions of this work are threefold: \(1\) we introduce AVA, a large\-scale benchmark for evaluating ontology\-aware semantic similarity using structured logic\-sensitive perturbations; \(2\) we provide a comprehensive evaluation of modern embedding models and contrastive learning objectives, including Euclidean and hyperbolic formulations; and \(3\) we demonstrate that high contrastive discrimination accuracy does not necessarily imply transferable ontology understanding, highlighting important limitations of current embedding\-based approaches for Semantic Web applications\. Furthermore, we make the implementation publicly available to the research community at[https://github\.com/sciknoworg/AVA](https://github.com/sciknoworg/AVA)\.

Figure 1\.Dataset analysis: semantic similarity \(cosine similarity computed using MPNET\-base\) and lexical similarity \(token\-set ratio\), distributions for anchor\-positive \(A\-P\), anchor\-negative \(A\-N\), and positive\-negative \(P\-N\) pairs\.Table 1\.Examples of logic\-sensitive hard negatives\.Anchor: An online gaming account is a subclass of anonline account\.Hard Negative: An online gaming account is a subclass of anagent\.Anchor: Online chat accounts are defined as a subclass ofonline accounts\.Hard Negative: Online chat accounts are defined as a subclass ofonline e\-commerce accounts\.Anchor: An online gaming account is a specific type ofonline accounts\.Hard Negative: An online gaming account is a specific type ofonline chat account\.

Table 2\.Dataset statistics\.![Refer to caption](https://arxiv.org/html/2609.00177v1/images/example_bfs_triplet.png)Figure 2\.Overview of triplet generation\. \(A\) A two\-hop BFS subgraph extracted from an ontology\. \(B\) Example contrastive triplet derived from the subgraph\.
## 2\.AVA

AVA is an evaluation framework for assessing whether embedding models capture ontology\-level relational semantics beyond surface lexical similarity\. It combines \(1\) ontology perturbations that generate logic\-sensitive contrastive triplets and \(2\) contrastive objectives that evaluate ontology\-aware discrimination and cross\-ontology transfer using triplet loss, hyperbolic loss, and reinforcement learning techniques\([Christiano et al\., 2017](https://arxiv.org/html/2609.00177#bib.bib10);[Lambert et al\., 2022](https://arxiv.org/html/2609.00177#bib.bib21)\)\.

### 2\.1\.Structured Ontology Perturbations

Ontology Graph Extraction\.We used 163 ontologies from diverse domains, including biomedical \(i\.e\., GO\([Consortium, 2026](https://arxiv.org/html/2609.00177#bib.bib11)\), OBI\([Bandrowski et al\., 2016](https://arxiv.org/html/2609.00177#bib.bib4)\)\), geospatial, social, engineering, and schema ontologies\. We accessed these collections via the OntoLearner library\([Giglou et al\., 2026](https://arxiv.org/html/2609.00177#bib.bib17)\)\. For each ontology, we construct an undirected graph in which nodes represent OWL classes and object/datatype/annotation properties, and edges encode structural OWL/RDFS relations \(i\.e\.,rdfs:subClassOf,owl:domain,owl:range, etc\)\. Existential restriction axioms of the formC⊑∃R\.DC\\sqsubseteq\\exists R\.Dare reified as direct labeled edges betweenCCandDD\. To construct plausible hard negatives—such as sibling swaps—without exceeding the LLM context window, we extract subgraphs using a two\-hop breadth\-first search \(BFS\) seeded at every OWL class\. Single\-hop neighborhoods frequently lack sibling classes, whereas a two\-hop radius reliably captures the necessary sibling and grandparent relationships\. We retain subgraphs bounded between 3 and 20 nodes to exclude trivial or unwieldy structures, and deduplicate them by the MD5 fingerprint of their sorted node sets, resulting in 50,548 unique subgraphs\.

Contrastive Triplet Generation Prompt<task\>
The following content represents a structural sub\-graph extracted from a formal knowledge ontology, detailing entities, their definitions, and their semantic relationships \(triples\)\.
</task\><objective\>
Synthesize a high\-quality dataset of natural language sentence triplets \(Anchor, Positive, Negative\)\. This dataset will be utilized for contrastive learning to fine\-tune an embedding model \(e\.g\., a Sentence\-Transformer\) for semantic representation and ontology alignment\.
</objective\><guidelines\>
Guidelines for Triplet Generation:
1\.Anchor: Formulate a clear, unambiguous declarative sentence that reflects a true fact, definition, or exact relation directly asserted explicitly in the provided sub\-graph\.
2\.Positive: Construct a semantically equivalent paraphrase of the Anchor\. This sentence must preserve the exact factual meaning but utilize morphological variations, synonyms, or alternative syntactic structures \(e\.g\., active versus passive voice\) to encourage robust semantic representation\.
3\.Negative: Construct a ’hard negative’ statement\. This sentence must exhibit high lexical overlap with the Anchor \(employing similar entities, relations, or vocabulary from the ontology domain\) but assert a factually incorrect relationship, a contradictory definition, or pair disjoint entities\. It must serve as a difficult distractor for the embedding model\.
</guidelines\><output\-format\>
\- Generate exactly 5 unique triplets\. Output strictly as a valid JSON array
\- Output the result strictly as a valid JSON array of objects, with the exact keys: "anchor", "positive", and "negative"\.
Example Output:```
[{"anchor": "...", "positive": "...", "negative": "..."}]
```

</output\-format\><input\-subgraph\>
\{subgraph\}
</input\-subgraph\>Figure 3\.LLM prompt used for generating contrastive triplets\.\{subgraph\}is a placeholder for structured presentation of subgraphs\.Logic\-Sensitive Hard Negative Synthesis\.Each subgraph is converted into a structured natural\-language prompt presenting its entities with labels and definitions alongside their RDF triples\. We instruct aQwen3\.5\-35B\-A3BLLM\([Qwen Team, 2026](https://arxiv.org/html/2609.00177#bib.bib30)\)to generate five contrastive triplets per subgraph under the following constraints \(see prompt at[Figure 3](https://arxiv.org/html/2609.00177#S2.F3)\):1\) Anchor, a declarative sentence expressing a fact directly asserted in the subgraph \(hierarchy assertion, domain/range constraint, equivalence, or disjointness\);2\) Positive, a semantically equivalent paraphrase using morphological variation, synonymy, or syntactic restructuring \(e\.g\., active to passive voice\)\. The factual content must be preserved exactly;3\) Hard negative, a statement exhibiting high lexical overlap with the anchor but asserting an ontologically incorrect relationship, for instance, swapping a subclass for a sibling, inverting a property domain, or replacing an entity with a disjoint concept\. This synthesis procedure is specifically designed to produce*logic\-sensitive*negatives rather than randomly sampled distractors\. After post\-processing the raw outputs, we obtained a total of 197,326 candidate \(Anchor, Positive, Negative\) triplets\. Representative examples are shown in[Table 1](https://arxiv.org/html/2609.00177#S1.T1), and[Figure 2](https://arxiv.org/html/2609.00177#S1.F2)shows how, using BFS subgraphs, triplets are generated\.

Post\-Processing and Dataset Statistics\.We apply two filtering passes using token\-set ratio similarity scores computed withRapidFuzz\. First, we flag samples where the anchor–negative similarity exceeds 90 as hard negatives and move them to a dedicated evaluation split, as they represent the most challenging cases\. Second, samples where the positive–negative similarity exceeds 90 are discarded outright, as the two cannot be meaningfully distinguished \(w\.r\.t[Figure 1](https://arxiv.org/html/2609.00177#S1.F1), the A\-P distributions\)\. After de\-duplication at the anchor level, the dataset contains 171,007 samples\. We perform the train/test split at the ontology level: all samples from a given ontology are assigned exclusively to train or test, with the test set constructed from ontologies whose label sets share fewer than 100 labels with the training ontologies\. Source and subgraph disjointness is verified programmatically to ensure that there is no leakage between test and train sets, except for a small number of subgraphs \(less than 0\.1% of triplets, which is≈\\approx148 triplets\) rooted in upper\-ontology classes shared across multiple imported ontologies\. Dataset statistics are summarized in[Table 2](https://arxiv.org/html/2609.00177#S1.T2)\. General NLP embeddings \(specifically MPNET\-base\) cannot separate anchor\-positive from anchor\-negatives in semantic space, yet hard negatives are deliberately more lexically similar to the anchor than the positives \(see[Figure 1](https://arxiv.org/html/2609.00177#S1.F1)\)\. A model that learns to discriminate these triplets, therefore, cannot rely on lexical overlap and must learn something about relational structure\.

### 2\.2\.Contrastive Learning

Triplet Loss\.We fine\-tune using the standard cosine triplet loss, which optimizes the margin between anchor\-positive and anchor\-negative similarity scores\. Given embeddings𝒜\\mathcal\{A\},𝒫\\mathcal\{P\}, and𝒩\\mathcal\{N\}, the loss isℒtri=max⁡\(0,sim⁡\(𝒜,𝒩\)−sim⁡\(𝒜,𝒫\)\+m\)\\mathcal\{L\}\_\{\\text\{tri\}\}=\\max\(0,\\,\\mathrm\{sim\}\(\\mathcal\{A\},\\mathcal\{N\}\)\-\\mathrm\{sim\}\(\\mathcal\{A\},\\mathcal\{P\}\)\+m\), wheremmis a margin hyperparameter set to0\.30\.3\.

Hyperbolic Triplet Loss\.Ontological class hierarchies are inherently tree\-structured, which Euclidean space represents poorly\. We replace the Euclidean margin with a Poincaré\-ball distance\([Nickel and Kiela, 2017](https://arxiv.org/html/2609.00177#bib.bib27)\), computingℒhyp=max⁡\(0,dc​\(𝒜,𝒫\)−dc​\(𝒜,𝒩\)\+m\)\\mathcal\{L\}\_\{\\text\{hyp\}\}=\\max\(0,\\,d\_\{c\}\(\\mathcal\{A\},\\mathcal\{P\}\)\-d\_\{c\}\(\\mathcal\{A\},\\mathcal\{N\}\)\+m\), wheredcd\_\{c\}is the hyperbolic distance at curvaturec=0\.3c=0\.3\. This encourages the model to reflect hierarchical structure in its embedding geometry\.

Embedding\-Adapted DPO\.We adapt Direct Preference Optimization\([Rafailov et al\., 2023](https://arxiv.org/html/2609.00177#bib.bib31)\)to the embedding setting by treating cosine similarity as an implicit reward\. A frozen reference encoderfreff\_\{\\text\{ref\}\}acts as a regularizer, and the loss penalizes the policy encoderfθf\_\{\\theta\}whenever its similarity marginΔθ=simθ​\(𝒜,𝒫\)−simθ​\(𝒜,𝒩\)\\Delta\_\{\\theta\}=\\mathrm\{sim\}\_\{\\theta\}\(\\mathcal\{A\},\\mathcal\{P\}\)\-\\mathrm\{sim\}\_\{\\theta\}\(\\mathcal\{A\},\\mathcal\{N\}\)falls behind the reference marginΔref\\Delta\_\{\\text\{ref\}\}, with temperatureβ=0\.5\\beta=0\.5\.

Table 3\.Evaluation results\.Table 4\.Taxonomy discovery results on MPNET\-Base/MiniLM\-L6 models before/after fine\-tuning using the perpetuated dataset\.Table 5\.Ontology alignment results on MPNET\-Base/MiniLM\-L6 models before/after fine\-tuning using the perpetuated dataset\.

## 3\.Results

Evaluation is performed under a cross\-ontology setting where evaluation ontological samples are fully excluded from training\. We use L2\-normalized embeddings and cosine similaritysim⁡\(x,y\)=x⋅y\|x\|​\|y\|\\mathrm\{sim\}\(x,y\)=\\frac\{x\\cdot y\}\{\|x\|\|y\|\}\. Performance is measured viatriplet accuracy, defined as the proportion of cases wheresim⁡\(𝒜,𝒫\)\>sim⁡\(𝒜,𝒩\)\\mathrm\{sim\}\(\\mathcal\{A\},\\mathcal\{P\}\)\>\\mathrm\{sim\}\(\\mathcal\{A\},\\mathcal\{N\}\)\(ties counted as incorrect\)\. We also reportHard Negative Accuracyon samples where anchor–negative lexical similarity exceeds 90 \(token\-set similarity\), measuring how often the model correctly ranks the positive above the hard negative \(s​i​m​\(𝒜,𝒫\)\>s​i​m​\(𝒜,𝒩\)sim\(\\mathcal\{A\},\\mathcal\{P\}\)\>sim\(\\mathcal\{A\},\\mathcal\{N\}\)\)\.

Performance of Pre\-trained Embeddings\.The[Table 3](https://arxiv.org/html/2609.00177#S2.T3)summarizes the performance of more than 25 pre\-trained embedding models\. Overall, results reveal substantial limitations in current general\-purpose embeddings when confronted with ontology\-sensitive semantic distinctions\. Among all evaluated models,Qwen3\-Embedding\-0\.6Bachieves the strongest performance, reaching a triplet accuracy of0\.7390\.739and a hard negative accuracy of0\.5720\.572\. Larger variants of the same family show comparable results, suggesting that model scale alone does not guarantee improved ontological discrimination\. Sentence embedding models such asMiniLM\-L6,MPNET\-base,GTE, andmultilingual\-E5achieve moderate performance, with triplet accuracies generally between0\.630\.63and0\.660\.66\. In contrast, several embedding models perform substantially worse; specifically, the OpenAIText\-Embedding\-3\-largereaches only0\.3880\.388triplet accuracy, whileLlama\-Embed\-Nemotron\-8Bobtains0\.2170\.217triplet accuracy and only 0\.135 hard negative accuracy\.

A consistent trend across all models is the large performance drop on hard negatives\. Although some embeddings achieve reasonable overall triplet accuracy, they still struggle to distinguish ontology\-consistent statements from highly similar contradictory statements\. This suggests that many embeddings rely primarily on lexical and distributional similarity rather than explicit sensitivity to relational semantics such as subclass structure, domain/range constraints, or disjointness relations\. These findings indicate that strong performance on standard semantic similarity or retrieval benchmarks does not necessarily translate into competence on ontology\-level understanding or discrimination\.

Effect of Contrastive Fine\-Tuning\.Fine\-tuning dramatically improves triplet discrimination performance\. Across bothMiniLM\-L6andMPNET\-base\(a widely used model in semantic web engineering tasks\), all three optimization objectives substantially outperform their corresponding pre\-trained baselines\. ForMPNET\-base, standard triplet loss increases triplet accuracy from0\.6360\.636to0\.9850\.985, while hyperbolic triplet loss further improves performance to0\.9890\.989\. Hard negative accuracy exhibits a similar pattern, increasing from0\.4270\.427to0\.9720\.972and0\.9800\.980, respectively\. Comparable improvements are observed forMiniLM\-L6, where hyperbolic loss achieves the strongest overall results with0\.9810\.981triplet accuracy and0\.9660\.966hard negative accuracy\. The hyperbolic objective consistently outperforms Euclidean triplet loss by a small but measurable margin\. This result is consistent with prior works suggesting that hyperbolic geometry provides a more natural representation space for hierarchical structures\([Nickel and Kiela, 2017](https://arxiv.org/html/2609.00177#bib.bib27);[Dhingra et al\., 2018](https://arxiv.org/html/2609.00177#bib.bib13)\)\. In contrast, the embedding\-adapted DPO objective performs substantially worse than both triplet\-based approaches, although it still improves considerably over the corresponding pre\-trained models\. This suggests the primary bottleneck may be reliance on a reference module whose representations do not encode fine\-grained ontological distinctions\. Consequently, preference optimization is likely constrained by the semantic limitations of the reference model, reducing its ability to learn transferable ontology\-level representations\.

Taken together, these results demonstrate that ontology\-aware contrastive supervision enables embeddings to separate semantically valid statements from logic\-sensitive negatives with near\-perfect accuracy\. A similar scenario is also observed in KEPLER\([Wang et al\., 2021](https://arxiv.org/html/2609.00177#bib.bib38)\), where the fine\-tuned model on the generated dataset performed very well\. However, high performance alone might not establish that models have acquired transferable ontological understanding\. Instead, the models likely learn to recognize perturbation\-specific structural patterns rather than genuine symbolic semantics\.

Transfer to Downstream Ontology Engineering Tasks\.Despite the dramatic gains, downstream evaluation reveals a substantial optimization–generalization gap\. The[Table 4](https://arxiv.org/html/2609.00177#S2.T4)reports taxonomy discovery task performance, where the aim is that, for a given set of ontological classes, the models should construct taxonomic pairs \(parent, child\) that form a subclass \(or is\-a\) relation \(experiments were performed with the OntoLearner\([Giglou et al\., 2026](https://arxiv.org/html/2609.00177#bib.bib17)\)library\)\. Contrary to expectations, fine\-tuning often degrades performance relative to the original models\. Standard triplet loss consistently reduces recall across most ontology benchmarks\. Hyperbolic loss partially mitigates this decline and occasionally yields modest improvements, particularly on SchemaOrg, OBI, SWEET, and Plant Ontology \(PO\), but gains remain relatively small compared to the near\-perfect improvements on earlier stages\. DPO fine\-tuning produces severe degradation across all datasets\. A similar pattern appears in ontology alignment \(see[Table 5](https://arxiv.org/html/2609.00177#S2.T5)\), the task of finding equivalent classes between two different ontologies \(experiments were conducted using the OntoAligner\([Babaei Giglou et al\., 2025](https://arxiv.org/html/2609.00177#bib.bib3)\)library\)\. Triplet and hyperbolic losses provide only modest improvements on several benchmarks, including ENVO–SWEET and MaterialInformation–MatOnto, while performance on other datasets remains largely unchanged\. Hyperbolic loss achieves the strongest overall transfer performance, but improvements are typically measured in only a few percentage points\. In contrast, DPO again substantially reduces alignment accuracy across all evaluation datasets\.

The discrepancy between near\-perfect triplet ranking accuracy and comparatively small downstream gains suggests that optimization does not necessarily induce robust ontology\-level understanding\. Instead, models may learn perturbation\-specific decision boundaries that effectively distinguish the generated negatives without acquiring transferable representations of hierarchical and logical structure\. These findings indicate that contrastive discrimination and ontology generalization should be treated as distinct evaluation objectives rather than interchangeable measures of semantic understanding\.

## 4\.Discussion

Because AVA’s triplets are synthesized by a single LLM \(Qwen3\.5\-35B\-A3B\) and the strongest pretrained encoder in our evaluation belongs to the same model family \(Qwen3\-Embedding\), we cannot rule out that part of its advantage reflects shared stylistic or distributional artifacts between generator and encoder rather than genuinely superior ontological sensitivity\. We treat this as a limitation of single\-generator benchmark construction; incorporating multiple LLMs for triplet synthesis would increase lexical and stylistic diversity and further mitigate generator\-specific artifacts, and we leave this, together with broader cross\-generator validation, to future work\.

We also note that, as commonly observed for retrieval modules in RAG pipelines, taxonomy discovery performance degrades when candidates exhibit high semantic overlap\. Analysis in\([Giglou et al\., 2026](https://arxiv.org/html/2609.00177#bib.bib17)\)indicates substantial semantic overlap among candidate classes across the ontologies used here, suggesting that the persistently low absolute recall observed for all models partly reflects this intrinsic task difficulty, which ontology\-aware fine\-tuning only partially mitigates\.

Furthermore, the AVA specifically evaluates ontology discrimination rather than logical reasoning\. It measures whether a model can distinguish between two lexically similar statements according to their ontological consistency, but it does not assess inference or entailment\. Consequently, performance on the AVA should not be interpreted as evidence of logical reasoning ability\. A separate evaluation would be required to test reasoning, for example, by assessing whether a model can inferA⊑CA\\sqsubseteq CfromA⊑BA\\sqsubseteq BandB⊑CB\\sqsubseteq CwhenA⊑CA\\sqsubseteq Cis not explicitly asserted\. Such a setting would directly evaluate the model’s ability to derive new knowledge from formal ontology axioms rather than merely discriminate between ontology\-consistent and ontology\-violating statements\.

## 5\.Conclusion

In conclusion, while embedding\-based models can perform well on similarity and retrieval tasks, this does not necessarily imply that they can reliably discriminate between ontology\-consistent and ontology\-violating statements\. In ontology engineering, such models may capture statistical patterns in the data while failing to reflect formal logical constraints\. It therefore remains unclear whether embedding optimization alone is sufficient for reliable ontology\-level discrimination, or whether explicit logical mechanisms are needed to adequately capture such constraints\.

## GenAI Disclosure

In preparing this manuscript, generative AI tools, specifically ChatGPT and Gemini, were used solely for grammar checking, spelling checks, and readability of some sentences\. All suggested changes were carefully reviewed and adapted by the authors to ensure accuracy and appropriateness\. The scientific content, research design, analysis, and conclusions were developed and verified exclusively by the authors without AI involvement\.

## References

- Babaei Giglou et al\.\(2023\)Hamed Babaei Giglou, Jennifer D’Souza, and Sören Auer\. 2023\.LLMs4OL: Large language models for ontology learning\. In*International semantic web conference*\. Springer, 408–427\.
- Babaei Giglou et al\.\(2025\)Hamed Babaei Giglou, Jennifer D’Souza, Oliver Karras, and Sören Auer\. 2025\.Ontoaligner: A comprehensive modular and robust python toolkit for ontology alignment\. In*European Semantic Web Conference*\. Springer, 174–191\.
- Bandrowski et al\.\(2016\)Anita Bandrowski, Ryan Brinkman, Mathias Brochhausen, Matthew H Brush, Bill Bug, Marcus C Chibucos, Kevin Clancy, Mélanie Courtot, Dirk Derom, Michel Dumontier, et al\.2016\.The ontology for biomedical investigations\.*PloS one*11, 4 \(2016\), e0154556\.
- Bian \(2025\)Haonan Bian\. 2025\.LLM\-empowered knowledge graph construction: A survey\.*arXiv preprint arXiv:2510\.20345*\(2025\)\.
- Bordes et al\.\(2013\)Antoine Bordes, Nicolas Usunier, Alberto Garcia\-Duran, Jason Weston, and Oksana Yakhnenko\. 2013\.Translating embeddings for modeling multi\-relational data\.*Advances in neural information processing systems*26 \(2013\)\.
- Chen et al\.\(2021a\)Jun Chen, Azza Althagafi, and Robert Hoehndorf\. 2021a\.Predicting candidate genes from phenotypes, functions and anatomical site of expression\.*Bioinformatics*37, 6 \(2021\), 853–860\.
- Chen et al\.\(2021b\)Jiaoyan Chen, Pan Hu, Ernesto Jimenez\-Ruiz, Ole Magnus Holter, Denvar Antonyrajah, and Ian Horrocks\. 2021b\.OWL2Vec\*: embedding of OWL ontologies\.*Machine Learning*110, 7 \(2021\), 1813–1845\.
- Chen et al\.\(2023\)Mingyang Chen, Wen Zhang, Yuxia Geng, Zezhong Xu, Jeff Z\. Pan, and Huajun Chen\. 2023\.Generalizing to unseen elements: a survey on knowledge extrapolation for knowledge graphs\. In*Proceedings of the Thirty\-Second International Joint Conference on Artificial Intelligence*\(Macao, P\.R\.China\)*\(IJCAI ’23\)*\. Article 737, 9 pages\.[doi:10\.24963/ijcai\.2023/737](https://doi.org/10.24963/ijcai.2023/737)
- Christiano et al\.\(2017\)Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei\. 2017\.Deep reinforcement learning from human preferences\.*Advances in neural information processing systems*30 \(2017\)\.
- Consortium \(2026\)The Gene Ontology Consortium\. 2026\.The Gene Ontology knowledgebase in 2026\.*Nucleic Acids Research*54, D1 \(01 2026\), D1779–D1792\.arXiv:https://academic\.oup\.com/nar/article\-pdf/54/D1/D1779/66009250/gkaf1292\.pdf[doi:10\.1093/nar/gkaf1292](https://doi.org/10.1093/nar/gkaf1292)
- Devlin et al\.\(2019\)Jacob Devlin, Ming\-Wei Chang, Kenton Lee, and Kristina Toutanova\. 2019\.Bert: Pre\-training of deep bidirectional transformers for language understanding\. In*Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 \(long and short papers\)*\. 4171–4186\.
- Dhingra et al\.\(2018\)Bhuwan Dhingra, Christopher Shallue, Mohammad Norouzi, Andrew Dai, and George Dahl\. 2018\.Embedding text in hyperbolic spaces\. In*Proceedings of the Twelfth Workshop on Graph\-Based Methods for Natural Language Processing \(TextGraphs\-12\)*\. 59–69\.
- Dragisic et al\.\(2017\)Zlatan Dragisic, Valentina Ivanova, Huanyu Li, and Patrick Lambrix\. 2017\.Experiences from the anatomy track in the ontology alignment evaluation initiative\.*Journal of biomedical semantics*8 \(2017\), 1–28\.
- Fallatah et al\.\(2020\)Omaima Fallatah, Ziqi Zhang, and Frank Hopfgartner\. 2020\.A gold standard dataset for large knowledge graphs matching\. In*Ontology Matching 2020: Proceedings of the 15th International Workshop on Ontology Matching co\-located with the 19th International Semantic Web Conference \(ISWC 2020\)*, Vol\. 2788\. CEUR Workshop Proceedings, 24–35\.
- Gatto et al\.\(2023\)Joseph Gatto, Omar Sharif, Parker Seegmiller, Philip Bohlman, and Sarah M\. Preum\. 2023\.Text Encoders Lack Knowledge: Leveraging Generative LLMs for Domain\-Specific Semantic Textual Similarity\. In*Proceedings of the Third Workshop on Natural Language Generation, Evaluation, and Metrics \(GEM\)*, Sebastian Gehrmann, Alex Wang, João Sedoc, Elizabeth Clark, Kaustubh Dhole, Khyathi Raghavi Chandu, Enrico Santus, and Hooman Sedghamiz \(Eds\.\)\. Association for Computational Linguistics, Singapore, 277–288\.[https://aclanthology\.org/2023\.gem\-1\.23/](https://aclanthology.org/2023.gem-1.23/)
- Giglou et al\.\(2026\)Hamed Babaei Giglou, Jennifer D’Souza, Andrei Aioanei, Nandana Mihindukulasooriya, and Sören Auer\. 2026\.OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models\.*arXiv preprint arXiv:2607\.01977*\(2026\)\.
- Hertling and Paulheim \(2023\)Sven Hertling and Heiko Paulheim\. 2023\.Olala: Ontology matching with large language models\. In*Proceedings of the 12th knowledge capture conference 2023*\. 131–139\.
- Karam et al\.\(2020\)Naouel Karam, Abderrahmane Khiat, Alsayed Algergawy, Melanie Sattler, Claus Weiland, and Marco Schmidt\. 2020\.Matching biodiversity and ecology ontologies: challenges and evaluation results\.*The Knowledge Engineering Review*35 \(2020\), e9\.
- Kumar et al\.\(2025\)Lokendra Kumar, Neelesh S Upadhye, and Kannan Piedy\. 2025\.Advances and Challenges in Semantic Textual Similarity: A Comprehensive Survey\.*arXiv preprint arXiv:2601\.03270*\(2025\)\.
- Lambert et al\.\(2022\)Nathan Lambert, Louis Castricato, Leandro von Werra, and Alex Havrilla\. 2022\.Illustrating reinforcement learning from human feedback \(rlhf\)\.*Hugging Face Blog*9 \(2022\)\.
- Li et al\.\(2025\)Wenda Li, Tongya Zheng, Shunyu Liu, Yu Wang, Kaixuan Chen, Hanyang Yuan, Bingde Hu, Zujie Ren, Mingli Song, and Gang Chen\. 2025\.Towards Efficient LLM\-aware Heterogeneous Graph Learning\.*arXiv preprint arXiv:2511\.17923*\(2025\)\.
- Li et al\.\(2023\)Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang\. 2023\.Towards general text embeddings with multi\-stage contrastive learning\.*arXiv preprint arXiv:2308\.03281*\(2023\)\.
- Liang et al\.\(2023\)Xinyu Liang, Guannan Si, Jianxin Li, Pengxin Tian, Zhaoliang An, and Fengyu Zhou\. 2023\.A survey of inductive knowledge graph completion\.*Neural Comput\. Appl\.*36, 8 \(Dec\. 2023\), 3837–3858\.[doi:10\.1007/s00521\-023\-09286\-2](https://doi.org/10.1007/s00521-023-09286-2)
- Mikolov et al\.\(2013\)Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean\. 2013\.Efficient estimation of word representations in vector space\.*arXiv preprint arXiv:1301\.3781*\(2013\)\.
- Nas and Huschka \(2023\)E\. Nas and M\. Huschka\. 2023\.MSE Benchmark\.[https://github\.com/EngyNasr/MSE\-Benchmark](https://github.com/EngyNasr/MSE-Benchmark)\.
- Nickel and Kiela \(2017\)Maximillian Nickel and Douwe Kiela\. 2017\.Poincaré embeddings for learning hierarchical representations\.*Advances in neural information processing systems*30 \(2017\)\.
- Pennington et al\.\(2014\)Jeffrey Pennington, Richard Socher, and Christopher D Manning\. 2014\.Glove: Global vectors for word representation\. In*Proceedings of the 2014 conference on empirical methods in natural language processing \(EMNLP\)*\. 1532–1543\.
- Qiang \(2023\)Zhangcheng Qiang\. 2023\.Ontology\-compliant knowledge graphs\. In*European Semantic Web Conference*\. Springer, 298–309\.
- Qwen Team \(2026\)Qwen Team\. 2026\.Qwen3\.5: Towards Native Multimodal Agents\.[https://qwen\.ai/blog?id=qwen3\.5](https://qwen.ai/blog?id=qwen3.5)
- Rafailov et al\.\(2023\)Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn\. 2023\.Direct preference optimization: Your language model is secretly a reward model\.*Advances in neural information processing systems*36 \(2023\), 53728–53741\.
- Ristoski and Paulheim \(2016\)Petar Ristoski and Heiko Paulheim\. 2016\.Rdf2vec: Rdf graph embeddings for data mining\. In*International semantic web conference*\. Springer, 498–514\.
- Song et al\.\(2020\)Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie\-Yan Liu\. 2020\.Mpnet: Masked and permuted pre\-training for language understanding\.*Advances in neural information processing systems*33 \(2020\), 16857–16867\.
- Sun et al\.\(2025\)Yiqun Sun, Qiang Huang, Anthony KH Tung, and Jun Yu\. 2025\.Text embeddings should capture implicit semantics, not just surface meaning\.*arXiv preprint arXiv:2506\.08354*\(2025\)\.
- Sun et al\.\(2019\)Zhiqing Sun, Zhi\-Hong Deng, Jian\-Yun Nie, and Jian Tang\. 2019\.RotatE: Knowledge Graph Embedding by Relational Rotation in Complex Space\.*CoRR*abs/1902\.10197 \(2019\)\.[http://arxiv\.org/abs/1902\.10197](http://arxiv.org/abs/1902.10197)
- Trouillon et al\.\(2016\)Théo Trouillon, Johannes Welbl, Sebastian Riedel, Éric Gaussier, and Guillaume Bouchard\. 2016\.Complex embeddings for simple link prediction\. In*International conference on machine learning*\. PMLR, 2071–2080\.
- Wang et al\.\(2022\)Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei\. 2022\.Text embeddings by weakly\-supervised contrastive pre\-training\.*arXiv preprint arXiv:2212\.03533*\(2022\)\.
- Wang et al\.\(2021\)Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang\. 2021\.KEPLER: A unified model for knowledge embedding and pre\-trained language representation\.*Transactions of the Association for Computational Linguistics*9 \(2021\), 176–194\.
- Xiao et al\.\(2024\)Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian\-Yun Nie\. 2024\.C\-Pack: Packed Resources For General Chinese Embeddings\. In*Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval*\(Washington DC, USA\)*\(SIGIR ’24\)*\. Association for Computing Machinery, New York, NY, USA, 641–649\.[doi:10\.1145/3626772\.3657878](https://doi.org/10.1145/3626772.3657878)
- Yao et al\.\(2019\)Liang Yao, Chengsheng Mao, and Yuan Luo\. 2019\.KG\-BERT: BERT for knowledge graph completion\.*arXiv preprint arXiv:1909\.03193*\(2019\)\.
- Zhang et al\.\(2024\)Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li, and Min Zhang\. 2024\.mGTE: Generalized Long\-Context Text Representation and Reranking Models for Multilingual Text Retrieval\. In*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track*, Franck Dernoncourt, Daniel Preoţiuc\-Pietro, and Anastasia Shimorina \(Eds\.\)\. Association for Computational Linguistics, Miami, Florida, US, 1393–1412\.[doi:10\.18653/v1/2024\.emnlp\-industry\.103](https://doi.org/10.18653/v1/2024.emnlp-industry.103)

Similar Articles