AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching
摘要
This paper introduces Hybrid Ontology Matching (HOM), unifying equivalence and subsumption discovery, and proposes AgentMap, an LLM-based multi-agent framework for joint ontology matching. Experiments show promising results on hybrid, equivalence-only, and subsumption-only settings.
查看缓存全文
缓存时间: 2026/07/31 04:01
# AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching
Source: [https://arxiv.org/html/2607.27130](https://arxiv.org/html/2607.27130)
Yiping Song1Jiaoyan Chen1Renate A\. Schmidt1Hui Yang1Wen Zhang2 1Department of Computer Science, The University of Manchester, UK \{jiaoyan\.chen, renate\.schmidt, hui\.yang\-2\}@manchester\.ac\.uk, yiping\.song@postgrad\.manchester\.ac\.uk 2School of Software Technology, Zhejiang University, China zhang\.wen@zju\.edu\.cn
###### Abstract
Ontology matching \(OM\) has traditionally been formulated as either equivalence discovery or subsumption matching\. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously discover equivalence and subsumption mappings\. In this paper, we introduceHybrid Ontology Matching \(HOM\), a new OM task that unifies equivalence and subsumption discovery, and accordingly propose a Large Language Model \(LLM\)\-based multi\-agent OM frameworkAgentMapthat is implemented by a series of interdependent semantic decisions\. Given a concept in the source ontology, AgentMap integrates semantic retrieval, hierarchical search, and collaborative multi\-agent LLM reasoning to progressively explore the target ontology, identifying either the equivalent concept, if one exists, or the most fine\-grained subsumer\. We further extend four OM datasets for a HOM benchmark and evaluate AgentMap under hybrid, equivalence\-only, and subsumption\-only settings\. Experimental results show that AgentMap achieves promising performance on the hybrid setting, and at the same time outperforms equivalence matching and subsumption matching baselines on the equivalence\-only and subsumption\-only settings, respectively\.
Keywords:Ontology Matching, Large Language Models, Subsumption Matching, Multi\-Agent Systems
## 1Introduction
Ontologies have been widely used for knowledge representation in many domains, such as SNOMED CT for healthcare and FoodOn for food and agriculture\[[2](https://arxiv.org/html/2607.27130#bib.bib43),[4](https://arxiv.org/html/2607.27130#bib.bib49),[23](https://arxiv.org/html/2607.27130#bib.bib48)\]\. For knowledge integration, reuse and interpretation, ontology matching \(OM\), which is to discover pairs of concepts across ontologies with a specific relationship like equivalence and subsumption, has been widely investigated\[[20](https://arxiv.org/html/2607.27130#bib.bib44)\]\. However, different ontologies may use different terminologies and adopt conceptualisations with different levels of granularity, and thus accurate OM is challenging, requiring semantic interpretation beyond lexical and structural matching\.
OM systems have evolved from traditional lexical matching and rule\-based approaches, such as LogMap\[[14](https://arxiv.org/html/2607.27130#bib.bib25)\]and AgreementMakerLight \(AML\)\[[6](https://arxiv.org/html/2607.27130#bib.bib9)\], to encoder\-based pre\-trained language model \(PLM\)\-based approaches, such as BERTMap\[[8](https://arxiv.org/html/2607.27130#bib.bib2)\], and more recently to generative LLM\-based frameworks, such as GenOM\[[22](https://arxiv.org/html/2607.27130#bib.bib30)\]\. Most of the current OM systems focus on equivalence matching, although there are also some OM systems like BERTSub\[[3](https://arxiv.org/html/2607.27130#bib.bib3)\]that aim to discover concept pairs with subsumption relationships\. More importantly, to the best of our knowledge, there is a shortage of OM systems that simultaneously discover equivalence and subsumption mappings\.
To close this gap, we introduceHybrid Ontology Matching \(HOM\), a new OM task that requires jointly discovering equivalence and subsumption correspondences for each source concept \(formalised in the Problem Statement\)\.
Accordingly, we propose a multi\-agent OM framework named AgentMap, which decomposes HOM into a series of interdependent semantic reasoning and decision\-making steps\. This design is inspired by LLM agent paradigms that interleave reasoning with action\[[27](https://arxiv.org/html/2607.27130#bib.bib37)\], refine an initial decision using feedback from subsequent verification steps\[[17](https://arxiv.org/html/2607.27130#bib.bib50)\], and route control between specialised agents via handoffs\[[19](https://arxiv.org/html/2607.27130#bib.bib46)\]\. AgentMap combines embedding\-based candidate retrieval, hierarchy\-aware LLM reasoning, and lexical matching with conflict resolution, decomposing reasoning into specialised agents for equivalence screening, equivalence verification, and subsumption discovery\.
We further extend four existing OM datasets involving ontologies of the food and biomedicine domains for a benchmark of the new HOM task\. The main contributions of this work can be summarised as follows:
- •We introduce the new task of Hybrid Ontology Matching \(HOM\), which requires simultaneously considering equivalence and subsumption mappings\. It better reflects the requirements of real\-world knowledge integration and reuse than conventional OM tasks\.
- •We develop the framework AgentMap that decomposes the complex HOM task into a set of sub\-tasks orchestrated by an agentic workflow, effectively integrating established ontology matching techniques with LLM\-based reasoning\.
- •We evaluate AgentMap on four HOM datasets, demonstrating consistent improvements over alternative LLM\-based strategies\. We further compare AgentMap with both classic and recent OM systems under the conventional equivalence\-only and subsumption\-only settings, where it also achieves superior performance\.
## 2Problem Statement
Existing OM research is aimed at either equivalence matching or subsumption matching\. In this work, we formulate a new task termedHybrid Ontology Matching \(HOM\), where the OM system is expected to identify both equivalence and subsumption mappings simultaneously\.
In particular, given a source named conceptcsc\_\{s\}from a source ontology𝒪s\\mathcal\{O\}\_\{s\}and a target ontology𝒪t\\mathcal\{O\}\_\{t\}, HOM aims to identify a mapping represented as a triplem=\(cs,ct,r\)m=\(c\_\{s\},c\_\{t\},r\), wherectc\_\{t\}is a named concept from𝒪t\\mathcal\{O\}\_\{t\}andr∈\{equivalence,subsumption\}r\\in\\\{\\texttt\{equivalence\},\\texttt\{subsumption\}\\\}is the semantic relation betweencsc\_\{s\}andctc\_\{t\}\. The equivalence mapping is prioritised, as it represents more fine\-grained semantic association\. Namely, the system is expected to identifyctc\_\{t\}as the equivalent concept ofcsc\_\{s\}in𝒪t\\mathcal\{O\}\_\{t\}if it exists; otherwise, the system is expected to identifyctc\_\{t\}as the most specific subsumer ofcsc\_\{s\}in𝒪t\\mathcal\{O\}\_\{t\}\(i\.e\.,cs⊑ctc\_\{s\}\\sqsubseteq c\_\{t\}and there exists no named conceptct′≠ctc\_\{t\}^\{\\prime\}\\neq c\_\{t\}in𝒪t\\mathcal\{O\}\_\{t\}such thatcs⊑ct′⊑ctc\_\{s\}\\sqsubseteq c\_\{t\}^\{\\prime\}\\sqsubseteq c\_\{t\}\)\.
## 3AgentMap
AgentMap addresses HOM by restricting LLM reasoning to a progressively refined candidate space, combining embedding\-based retrieval with iterative exploration of the ontology hierarchy for efficient mapping discovery\. Figure[1](https://arxiv.org/html/2607.27130#S3.F1)presents the overall architecture of AgentMap, which consists of three modules:
Data Preprocessing and Candidate Retrievalextracts textual information from the source concept and the target ontology, encodes concepts into semantic embeddings, performs cosine\-similarity\-based retrieval, and constructs two candidate sets for lexical matching and agent\-based reasoning, respectively\.
Agent\-Based Reasoningperforms collaborative reasoning through multiple LLM\-based agents and ontology\-aware search to progressively identify semantic correspondences between concepts\.
Lexical Matching and Conflict Resolutioncomplements agent\-based semantic reasoning with lexical evidence and produces a unified correspondence by resolving potential inconsistencies between the two matching strategies\.
Figure 1:Overview of the proposed AgentMap framework\. Semantic retrieval first constructs task\-specific candidate sets for the agent\-based reasoning and lexical matching modules\. Their predictions are subsequently reconciled through LLM\-based conflict resolution to jointly predict the target concept and semantic relation\.### 3\.1Data Preprocessing and Candidate Retrieval
Given a source conceptcsc\_\{s\}and a target ontology𝒪t\\mathcal\{O\}\_\{t\}, this module constructs two task\-specific candidate sets for the agent and lexical matching modules\.
AgentMap extracts the textual information ofcsc\_\{s\}and the concepts in𝒪t\\mathcal\{O\}\_\{t\}, specifically the label and synonyms of each concept, encodes them into semantic embeddings, and ranks the target concepts according to their cosine similarity tocsc\_\{s\}\. Two candidate sets are then constructed from this ranking using different top\-k thresholds: \(1\) A smaller candidate setC0C\_\{0\}is constructed for the agent module, limiting costly LLM reasoning to the most semantically relevant concepts\. \(2\) A larger candidate setC\+C^\{\+\}is constructed for lexical matching, whose low computational cost permits broader candidate coverage\. This dual candidate\-set design provides each downstream module with a candidate scope suited to its computational characteristics and matching objective\.
### 3\.2Agent\-Based Reasoning
The agent\-based reasoning module consists of three specialised agents:AgentESAgent\_\{ES\}for initial equivalence screening,AgentEVAgent\_\{EV\}for equivalence verification, andAgentSDAgent\_\{SD\}for subsumption discovery\. The module follows an equivalence\-first strategy\.AgentESAgent\_\{ES\}first examines the initial candidate setC0C\_\{0\}for a potential equivalent concept\. If an equivalence candidate is identified,AgentEVAgent\_\{EV\}verifies the prediction through an ontology\-structure\-guided candidate update\. Otherwise,AgentSDAgent\_\{SD\}performs iterative subsumption discovery by progressively expanding the candidate set along the target ontology hierarchy\.
#### Equivalence Discovery\.
AgentESAgent\_\{ES\}receives the source conceptcsc\_\{s\}and the initial candidate setC0C\_\{0\}, and performs LLM reasoning to determine whetherC0C\_\{0\}contains a valid equivalent target concept\. If an equivalence candidate is identified,AgentESAgent\_\{ES\}outputs the most likely target conceptc^t\\hat\{c\}\_\{t\}and passes it toAgentEVAgent\_\{EV\}for further verification\.
AgentEVAgent\_\{EV\}receives the source conceptcsc\_\{s\}together with a candidate setCrefC\_\{\\mathrm\{ref\}\}, constructed from the direct parent and child concepts ofc^t\\hat\{c\}\_\{t\}in the target ontology\. Specifically, AgentMap retrieves the direct parent and child concepts ofc^t\\hat\{c\}\_\{t\}, denotedP\(c^t\)P\(\\hat\{c\}\_\{t\}\)andCh\(c^t\)Ch\(\\hat\{c\}\_\{t\}\), respectively, and, to avoid introducing every parent and child concept into the verification process, selects only the parent concept and the child concept that are most semantically similar tocsc\_\{s\}:
p∗=argmaxp∈P\(c^t\)sim\(cs,p\),p^\{\*\}=\\arg\\max\_\{p\\in P\(\\hat\{c\}\_\{t\}\)\}sim\(c\_\{s\},p\),ch∗=argmaxch∈Ch\(c^t\)sim\(cs,ch\),ch^\{\*\}=\\arg\\max\_\{ch\\in Ch\(\\hat\{c\}\_\{t\}\)\}sim\(c\_\{s\},ch\),wheresim\(⋅,⋅\)sim\(\\cdot,\\cdot\)denotes cosine similarity of the embeddings\.
The resulting candidate set is
Cref=\{c^t,p∗,ch∗\}\.C\_\{\\mathrm\{ref\}\}=\\\{\\hat\{c\}\_\{t\},p^\{\*\},ch^\{\*\}\\\}\.Ifc^t\\hat\{c\}\_\{t\}has no direct parent concept \(respectively, no direct child concept\),p∗p^\{\*\}\(respectively,ch∗ch^\{\*\}\) is undefined and is simply omitted fromCrefC\_\{\\mathrm\{ref\}\}; in this caseCrefC\_\{\\mathrm\{ref\}\}contains onlyc^t\\hat\{c\}\_\{t\}together with whichever ofp∗p^\{\*\}andch∗ch^\{\*\}exists\.
AgentEVAgent\_\{EV\}then performs LLM reasoning overCrefC\_\{\\mathrm\{ref\}\}to verify the equivalence correspondence, comparingc^t\\hat\{c\}\_\{t\}against its most semantically relevant parent and child concepts\. Based on this comparison,AgentEVAgent\_\{EV\}outputs a single final candidate fromCrefC\_\{\\mathrm\{ref\}\}: it either retainsc^t\\hat\{c\}\_\{t\}, or replaces it withp∗p^\{\*\}orch∗ch^\{\*\}, whichever candidate’s granularity better matchescsc\_\{s\}\.
#### Subsumption Discovery\.
IfAgentESAgent\_\{ES\}does not identify a valid equivalent concept, AgentMap invokesAgentSDAgent\_\{SD\}to iteratively search for the closest valid subsumer\.AgentSDAgent\_\{SD\}initially receives the source conceptcsc\_\{s\}together with the initial candidate setC0C\_\{0\}\. Atithi^\{th\}iteration, it performs LLM reasoning over the current candidate setCiC\_\{i\}to determine whether any candidate is a valid subsumer ofcsc\_\{s\}\. If a valid subsumerctc\_\{t\}is identified, AgentMap returns\(cs,ct,subsumption\)\.\(c\_\{s\},c\_\{t\},\\texttt\{subsumption\}\)\.Otherwise, the candidate set is updated using the structure of the target ontology\. Specifically, all direct parents of each concept inCiC\_\{i\}are collected to form the candidate set for the next iteration:
Ci\+1=Parents\(Ci\)=⋃c∈CiP\(c\)\.C\_\{i\+1\}=Parents\(C\_\{i\}\)=\\bigcup\_\{c\\in C\_\{i\}\}P\(c\)\.AgentSDAgent\_\{SD\}then performs the same reasoning process overCi\+1C\_\{i\+1\}\. Therefore, each iteration examines the direct parents of the complete candidate set from the preceding iteration, allowing the search to move upward through the ontology hierarchy level by level\.
This ontology\-structure\-guided update continues until a valid subsumer is identified or the maximum iterationsdmaxd\_\{\\max\}is reached\. If no valid subsumer is found afterdmaxd\_\{\\max\}upward expansions,AgentSDAgent\_\{SD\}performs a final LLM reasoning step over all candidates visited throughout the search process:
C0∪C1∪⋯∪Cdmax\.C\_\{0\}\\cup C\_\{1\}\\cup\\cdots\\cup C\_\{d\_\{\\max\}\}\.The agent then selects the most appropriate subsumer from these visited concepts\.
### 3\.3Lexical Matching and Conflict Resolution
In parallel with the agent\-based reasoning module, AgentMap performs lexical matching to identify equivalence correspondences supported by lexical evidence\. The lexical matching module operates on the candidate setC\+C^\{\+\}generated during candidate retrieval\. For each candidate concept inC\+C^\{\+\}, the module compares its representations with the source concept, including labels and synonyms\. A candidate is regarded as a lexical match if one of its labels or synonyms matches with some label or synonym of the source concept\. Because lexical matching alone cannot determine hierarchical relations, this module predicts onlyequivalencecorrespondences\.
In the end, the lexical matching result and the agent reasoning result are reconciled through LLM\-based conflict resolution if they are not identical: \(i\) if no candidate inC\+C^\{\+\}satisfies the lexical matching criterion above \(i\.e\., the lexical matching module identifies no equivalence candidate\), AgentMap returns the agent reasoning result; \(ii\) if lexical matching outputs an equivalent concept while the agents infer a subsumer, AgentMap returns the lexical matching result, following the equivalence\-first problem setting; \(iii\) if lexical matching and agent reasoning output two different equivalent target concepts, AgentMap performs an additional LLM reasoning step to determine the final output\.
## 4Evaluation
### 4\.1Benchmark and Metrics
#### Benchmark Construction
Existing benchmarks, such as Bio\-ML used by Ontology Alignment Evaluation Initiative \(OAEI\)\[[10](https://arxiv.org/html/2607.27130#bib.bib4)\], provide equivalence and subsumption reference mappings independently\. They can only evaluate OM systems for either equivalence or subsumption matching\. We therefore construct a dedicated benchmark for the new task of HOM\.
The proposed benchmark is built upon four equivalence OM datasets \(tasks\), including three from OAEI Bio\-ML \(SNOMED–FMA–Body, SNOMED–NCIT–Pharm, and NCIT–DOID–Disease\) which match medical ontologies, and the HeLiS\-FoodOn dataset\[[3](https://arxiv.org/html/2607.27130#bib.bib3)\]which mathches a health lifestyle ontology with the food ontology\. Since HOM requires jointly evaluating equivalence and subsumption matching against a single target ontology, we reorganize the source concepts of the original equivalence benchmark into two disjoint subsets: one for equivalence evaluation and one for subsumption evaluation\. Subsumption ground truth, however, is not independently annotated; instead, following\[[10](https://arxiv.org/html/2607.27130#bib.bib4)\], it is derived from the equivalence mappings by taking the direct parent of each equivalence target as the corresponding subsumer\.
This construction creates an ambiguity: for concepts assigned to the subsumption subset, their true equivalence target is still present in the original target ontology, even though the intended ground truth is now its parent\. Since the equivalence target is, by definition, a more specific and semantically closer match than its parent, any system able to identify it would report it instead of the coarser subsumption target used as ground truth, rendering subsumption evaluation ill\-defined on the unmodified ontology\. To resolve this, we remove, from the original target ontology, only the equivalence target concepts of the source concepts assigned to the subsumption subset, and reattach their child concepts to the corresponding parent to preserve the original taxonomy\. The resulting ontology, denoted𝒪t\\mathcal\{O\}\_\{t\}, is used consistently by all methods \(AgentMap and baselines\) throughout the paper, and jointly supports equivalence\-only, subsumption\-only, and HOM evaluation on a single, shared target ontology\.
Each resulting HOM dataset consists of a source ontology, a target ontology and a test set of source conceptsTtestT\_\{test\}, each of which is annotated by exactly one ground\-truth target concept and the corresponding matching type \(equivalenceorsubsumption\)\. The testing subset with equivalence \(resp\. subsumption\) target concept is denoted asTeqT\_\{eq\}\(resp\.TsubT\_\{sub\}\)\. See Table[1](https://arxiv.org/html/2607.27130#S4.T1)for more dataset statistics\.
Table 1:Statistics of the constructed HOM benchmark\.
#### Evaluation Metrics
Three evaluation settings are considered: HOM, equivalence\-alone, and subsumption\-alone\. TheHOM settinguses the whole test setTtestT\_\{test\}\. A source concept is regarded as correctly processed only if both the target concept and the matching type are correctly identified\. The overall performance is measured byOverallAcc\.\{\}\_\{\\text\{Acc\.\}\}which is the ratio of the corrected predicted source concepts among all the source conceptsTtestT\_\{test\}\. To further analyse performance on different matching types, we additionally report the accuraciesEqvAcc\.\{\}\_\{\\text\{Acc\.\}\}andSubAcc\.\{\}\_\{\\text\{Acc\.\}\}on the equivalence and subsumption subsets of the whole testing set, i\.e\.,TeqT\_\{eq\}andTsubT\_\{sub\}, respectively\.
In the HOM setting, the system does not know the matching type, while in theequivalence or subsumption\-alone setting, we let the system know the matching type in advance, so as to fairly comparing AgentMap with existing OM systems for either equivalence or subsumption matching\. A source concept is regarded as correctly processed if its ground\-truth target concept is identified\. The equivalence matching performance is measured byAccuracyeqAccuracy\_\{eq\}, which is the ratio of correctly processed source concepts amongTeqT\_\{eq\}, and the subsumption matching performance is measured byAccuracysubAccuracy\_\{sub\}which is the ratio of correctly processed source concepts amongTsubT\_\{sub\}\.
### 4\.2Baselines
#### HOM Setting\.
We construct four LLM\-based baselines by varying the prompting strategy and the available candidate set\. Two prompting strategies are considered: direct prompting and Chain\-of\-Thought \(CoT\) prompting\[[24](https://arxiv.org/html/2607.27130#bib.bib47)\]\. Two candidate configurations are evaluated\. The first uses the same reasoning candidate setC0C\_\{0\}as AgentMap, allowing the LLM to directly predict both the target concept and the semantic relation\. The second expandsC0C\_\{0\}by including the parent concepts \(up to two ontology levels\) and direct child concepts of each retrieved candidate, forming an expanded neighbourhood candidate set that approximates the maximum ontology neighbourhood explored by AgentMap\.
Combining the two prompting strategies with the two candidate configurations yields four baselines:
- •LLM \+C0C\_\{0\}: Direct prompting over the reasoning candidate set\.
- •LLM \+C0C\_\{0\}\+ CoT: Chain\-of\-Thought prompting over the reasoning candidate set\.
- •LLM \+ Neighbour: Direct prompting over the expanded neighbourhood candidate set\.
- •LLM \+ Neighbour \+ CoT: Chain\-of\-Thought prompting over the expanded neighbourhood candidate set\.
These baselines control both the prompting strategy and the available candidate coverage, helping isolate the contribution of AgentMap’s staged multi\-agent reasoning process\.
#### Equivalence\-alone Setting\.
For equivalence OM, we compare AgentMap with representative methods from three categories\. Traditional OM systems, LogMap\[[14](https://arxiv.org/html/2607.27130#bib.bib25)\]and AML\[[6](https://arxiv.org/html/2607.27130#bib.bib9)\], take two complete ontologies as input and produce a single global alignment for the whole ontology pair rather than a per\-concept prediction\. To adapt them toTeqT\_\{eq\}, we look up each source concept in this global alignment and use its mapped target, if any, as the prediction\. In contrast, pre\-trained language model methods, represented by BERTMap\[[8](https://arxiv.org/html/2607.27130#bib.bib2)\], and LLM methods, represented by GenOM\[[22](https://arxiv.org/html/2607.27130#bib.bib30)\], natively predict a target concept for each source concept individually, and are therefore evaluated directly onTeqT\_\{eq\}following the same protocol as AgentMap\.
#### Subsumption\-alone Setting\.
For subsumption OM, we compare AgentMap with five representative methods that score each candidate subsumer in an embedding space and rank them accordingly:\(1\)General sentence embedding methods for semantic similarity, such as SBERT and OpenAItext\-embedding\-3\-smallEmbedding;\(2\)Fine\-tuned hirerachy embeddings methods, such as HiT\[[11](https://arxiv.org/html/2607.27130#bib.bib32)\]and OnT\[[26](https://arxiv.org/html/2607.27130#bib.bib33)\]; and\(3\)Fine\-tuned PLM\-based classification model, such as BERTSub\[[3](https://arxiv.org/html/2607.27130#bib.bib3)\], which is a representative subsumption OM system\. For all baselines, accuracy is computed as Hit@1, i\.e\., whether the top\-scored candidate matches the ground truth\. Among these baselines, embedding\-based methods encode each concept once and reuse the encoding for every comparison, and can therefore consider all named concepts as candidates; BERTSub instead reruns the model for every source–candidate pair, making exhaustive scoring over the full ontology infeasible, and therefore requires a predefined candidate list as input\. To ensure a fair comparison, we provide BERTSub with a candidate set consisting of the initial retrieved candidatesC0C\_\{0\}together with their parent concepts expanded up to two ontology levels, matching the maximum search range explored byAgentSDAgent\_\{SD\}during subsumption discovery\. This controls candidate coverage while allowing different ranking and reasoning strategies to be compared\.
### 4\.3Experimental Setup
Unless otherwise specified, AgentMap uses GPT\-4\.1\-mini as the backbone LLM throughout all experiments\. All LLM\-based methods are evaluated with a decoding temperature of 0\. Ontology concepts are encoded using the OpenAItext\-embedding\-3\-smallembedding model, and cosine similarity is used for candidate retrieval\. The agent\-based reasoning module operates on the top\-5 retrieved candidates \(C0C\_\{0\}\), while the lexical matching module uses the top\-20 retrieved candidates \(C\+C^\{\+\}\)\. During ontology\-guided subsumption search, the maximum upward traversaldmaxd\_\{\\max\}is set to 2 \(i\.e\., up to the grandparent level of the initially retrieved target candidates in the ontology hierarchy\)\. The same configuration is adopted throughout all experiments unless explicitly stated otherwise\.
To evaluate the robustness of AgentMap, we additionally replace GPT\-4\.1\-mini with several representative open\-source LLMs\.111See Appendix for LLM prompt templates, SBERT embedding results,C0C\_\{0\}size sensitivity, and further details; code and data will be released\.
### 4\.4Experimental Results
#### HOM Results\.
Table 2:HOM Results\. All the methods \(AgentMap and the baselines\) use GPT\-4\.1\-mini as the backbone LLM\. Bold indicates the best result and underline indicates the second\-best in each column\.Table[2](https://arxiv.org/html/2607.27130#S4.T2)reports results for the HOM setting\. AgentMap achieves the highest overall accuracy on the three biomedical benchmarks \(SNOMED\-FMA\-Body, SNOMED\-NCIT\-Pharm, and NCIT\-DOID\-Disease\), and is only slightly behind the best baseline method on HeLiS–FoodOn\.
Specifically, AgentMap achieved consistently best subsumption accuracy across all four datasets\. For example, on SNOMED\-NCIT\-Pharm, AgentMap’s subsumption accuracy is up to 30\.8% higher than the best baseline \(0\.378 vs\. 0\.289\)\. These results show that decomposing HOM into staged semantic decisions, combined with iterative ontology\-guided search, is substantially more effective than single\-step reasoning over a fixed candidate set\.
For equivalence accuracy, the gap between AgentMap and the other baselines is much smaller\. In some cases, such as HeLiS–FoodOn, LLM\+Neighbourhood achieves higher equivalence accuracy than AgentMap\. This is due to the fact that the baseline’s expanded neighbourhood decides equivalence and subsumption together in a single LLM call, so the candidate set it sees when making the equivalence decision already includes the parent, grandparent, and child concepts needed for subsumption\. AgentMap’sAgentESAgent\_\{ES\}, in contrast, judges equivalence using onlyC0C\_\{0\}, which is a much smaller set\. With only 174 equivalence instances in HeLiS–FoodOn, this gap in candidate coverage is further amplified, noticeably lowering AgentMap’s equivalence accuracy relative to the baseline\.
#### Equivalence\-alone Results\.
SNOMED\-FMA
SNOMED\-NCIT
NCIT\-DOID
HeLiS\-FoodOn
Table 3:Comparison with existing ontology matching system on equivalence alignment\. Bold indicates the best result and underline the second\-best in each column\.Table[3](https://arxiv.org/html/2607.27130#S4.T3)compares AgentMap with representative OM systems on equivalence\-alone matching\. As GenOM relies on next\-token probabilities for candidate selection, it is only evaluated with open\-source LLMs; we thus additionally report AgentMap under the same backbone \(Qwen2\.5\-32B\-Instruct\[[25](https://arxiv.org/html/2607.27130#bib.bib40)\]\) for a fair comparison, retaining GPT\-4\.1\-mini as the default elsewhere\.
AgentMap achieves the best performance on all four benchmarks under both backbones\. LogMap fluctuates sharply \(0\.470 to 0\.901\): as an earlier\-generation system, it relies heavily on external lexicons for matching, whose coverage varies across domains, unlike BERT\- and LLM\-based methods that draw on learned semantic representations instead\. BERTMap is competitive on the three medical benchmarks but collapses to 0\.391 on HeLiS–FoodOn: it is built on BioClinicalBERT, a BERT variant pre\-trained specifically on clinical and biomedical text, whose vocabulary and representations are tailored to medical terminology and therefore transfer poorly to the food and lifestyle domain of HeLiS–FoodOn\. GenOM stays comparatively stable across all four, consistent with LLMs carrying broader, less domain\-specific knowledge\. Switching AgentMap’s backbone from Qwen2\.5\-32B to GPT\-4\.1\-mini yields only modest further gains, indicating the improvement stems mainly from the reasoning framework rather than the backbone LLM\.
#### Subsumption\-alone Results\.
SNOMED–FMA–Body
SNOMED–NCIT–Pharm
NCIT–DOID–Disease
HeLiS–FoodOn
Table 4:Comparison with existing subsumption matching methods\. Bold indicates the best result and underline the second\-best in each column\.Table[4](https://arxiv.org/html/2607.27130#S4.T4)presents results that compare AgentMap with the subsumption OM baselines\. AgentMap achieves the highest accuracy on all four benchmarks, with gains most pronounced on the three biomedical benchmarks; it improves the accuracy over the best baseline BERTSub from 0\.191 to 0\.401 on SNOMED\-FMA\-Body, from 0\.046 to 0\.398 on SNOMED\-NCIT\-Pharm, and from 0\.336 to 0\.564 on NCIT\-DOID\-Disease\. There is also a small gain \(from 0\.269 to 0\.278\) on HeLiS\-FoodOn\.Note that BERTSub is given the the initial retrieved candidatesC0C\_\{0\}as AgentMap \(see Baselines Section\)\.
The gain is especially large on SNOMED\-NCIT\-Pharm, where the top\-ranked candidates are often near\-synonymous, making it hard for embedding\- or BERT\-based scoring to separate the correct subsumer from its closest competitors; AgentMap’s LLM\-based reasoning is better able to resolve such fine\-grained distinctions\. More broadly, these results show that purely ranking is insufficient for subsumption matching and it requires progressively exploring the hierarchy through iterative agent\-based reasoning\.
### 4\.5Ablation Study
Table[5](https://arxiv.org/html/2607.27130#S4.T5)reports ablation results on SNOMED\-FMA\-Body, the largest of the four benchmarks, examining two components: the hierarchical search mechanism and the Lexical Matching and Conflict Resolution \(LM&CR\) module\.
The hierarchical search appear to be the primary contributor to AgentMap’s subsumption performance\. Removing hierarchical search can reduces overall accuracy by 15\.7% and subsumption accuracy by more than half \(58\.2%\), while equivalence accuracy is unaffected, Removing LM&CR, in contrast, causes only a modest drop in overall \(1\.4%\) and equivalence accuracy \(1\.7%\), with no effect on subsumption, indicating that LM&CR mainly refines equivalence prediction\.
Table 5:Ablation study of AgentMap on the SNOMED\-FMA\-Body benchmark \(LM&CR denotes the Lexical Matching and Conflict Resolution module\)\.
### 4\.6Effect of Backbone LLMs
Table 6:Effect of different backbone LLMs on HOM, Bold indicates the best result\.Table[6](https://arxiv.org/html/2607.27130#S4.T6)evaluates the effect of backbone LLMs on AgentMap\. The framework performs stably across LLMs such as GPT\-4\.1\-mini, Qwen2\.5\-32B\-Instruct, Qwen2\.5\-72B\-Instruct\[[25](https://arxiv.org/html/2607.27130#bib.bib40)\], and Llama3\.1\-70B\-Instruct\[[5](https://arxiv.org/html/2607.27130#bib.bib41)\], indicating that its effectiveness is largely independent of the backbone\. GPT\-4\.1\-mini and the Qwen2\.5 variants are the strongest overall, each best on two of the four benchmarks, with Llama3\.1\-70B\-Instruct remaining competitive; Superisingly, in SNOMED\-NCIT\-Pharm task, Qwen2\.5\-32B even archives better performance than bigger Qwen2\.5\-72B model\. Mixtral\-8x7B\-Instruct\[[13](https://arxiv.org/html/2607.27130#bib.bib42)\], in contrast, consistently underperforms across all four, which may be due to its weaker capabilities\.
## 5Related Work
### 5\.1Equivalence Ontology Matching
Equivalence OM has evolved from rule\-based systems to neural representation learning and, more recently, LLM\-based reasoning\. Early systems, such as LogMap\[[14](https://arxiv.org/html/2607.27130#bib.bib25)\]and AML\[[6](https://arxiv.org/html/2607.27130#bib.bib9)\], combine lexical similarity, ontology structures, and logical reasoning to construct high\-quality ontology alignments, and remain strong baselines in the OAEI benchmark\.
With the development of deep learning, OM has increasingly relied on learned semantic representations\[[8](https://arxiv.org/html/2607.27130#bib.bib2)\]\. Early neural approaches employed convolutional neural networks \(CNNs\) to encode ontology concepts from textual descriptions\[[1](https://arxiv.org/html/2607.27130#bib.bib29)\]\. More recent methods adopt transformer\-based language models, including BERTMap\[[8](https://arxiv.org/html/2607.27130#bib.bib2)\], BioSTransformer\[[18](https://arxiv.org/html/2607.27130#bib.bib19)\], BioGITOM\[[21](https://arxiv.org/html/2607.27130#bib.bib23)\], and Magneto\[[15](https://arxiv.org/html/2607.27130#bib.bib27)\], which leverage contextual language representations, graph neural architectures, or hybrid small\-large language models to improve semantic matching accuracy\.
Recent advances in LLMs have further shifted OM from representation learning toward semantic reasoning\[[9](https://arxiv.org/html/2607.27130#bib.bib5)\]\. Representative systems, including GenOM\[[22](https://arxiv.org/html/2607.27130#bib.bib30)\], LogMap\-LLM\[[16](https://arxiv.org/html/2607.27130#bib.bib31)\], Olala\[[12](https://arxiv.org/html/2607.27130#bib.bib7)\], and LLM4OM\[[7](https://arxiv.org/html/2607.27130#bib.bib1)\], exploit the reasoning capability of LLMs to perform ontology alignment through semantic understanding and multi\-step inference\.
Although these methods differ substantially in architecture, they are all designed for equivalence OM, where the objective is to identify concepts referring to the same real\-world entity across heterogeneous ontologies\.
### 5\.2Subsumption Ontology Matching
Compared with equivalence OM, subsumption OM has received considerably less attention\. Existing studies commonly formulate the task as a candidate ranking problem\[[10](https://arxiv.org/html/2607.27130#bib.bib4)\], where the objective is to identify the correct subsumer from a predefined candidate list\.
BERTSub\[[3](https://arxiv.org/html/2607.27130#bib.bib3)\]is one of the few approaches specifically designed for subsumption OM\. It uses contextual language models to rank candidate subsumers under existing benchmark settings\. However, BERTSub assumes that the benchmark\-supported subsumer is already contained in the candidate list and therefore evaluates candidate ranking rather than ontology\-wide subsumption discovery\.
Ontology representation learning methods such as OnT\[[26](https://arxiv.org/html/2607.27130#bib.bib33)\]and HiT\[[11](https://arxiv.org/html/2607.27130#bib.bib32)\]have also been used as concept encoders in related ranking settings\. However, they are not designed specifically for subsumption OM\. In this work, we include them only as encoder\-based baselines to examine how ontology\-aware representations perform under our evaluation protocol\.
HOM instead requires searching the target ontology to jointly identify the target concept and determine whether the relation isequivalenceorsubsumption— a joint formulation that, to the best of our knowledge, no existing OM framework supports\.
## 6Conclusion
We introduced Hybrid Ontology Matching \(HOM\), reframing equivalence and subsumption discovery as a single task evaluated jointly, reflecting the fact that, in practice, whether a source concept has an exact match or only a broader one is not known in advance\. AgentMap addresses HOM by decomposing this joint decision into staged, interdependent agent reasoning steps rather than a single LLM judgment\. Our ablation and cross\-backbone experiments show that this staged decomposition, rather than candidate coverage or backbone strength, drives AgentMap’s gains, suggesting iterative, structure\-aware reasoning as a general principle for LLM agents over hierarchical structures\. AgentMap also sets a new state of the art on subsumption matching, though its absolute accuracy remains below 0\.5 on three of four benchmarks, showing this sub\-task is still considerably harder than equivalence matching\. Future work includes closing this gap through more targeted hierarchical search and extending HOM to richer semantic relations\.
## References
- \[1\]\(2020\-05\)Ontology matching using convolutional neural networks\.InProceedings of the Twelfth Language Resources and Evaluation Conference,N\. Calzolari, F\. Béchet, P\. Blache, K\. Choukri, C\. Cieri, T\. Declerck, S\. Goggi, H\. Isahara, B\. Maegaard, J\. Mariani, H\. Mazo, A\. Moreno, J\. Odijk, and S\. Piperidis \(Eds\.\),Marseille, France,pp\. 5648–5653\(eng\)\.External Links:[Link](https://aclanthology.org/2020.lrec-1.693/),ISBN 979\-10\-95546\-34\-4Cited by:[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p2.1)\.
- \[2\]L\. Bos and K\. Donnelly\(2006\)SNOMED\-ct: the advanced terminology and coding system for ehealth\.Stud Health Technol Inform121,pp\. 279–290\.Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p1.1)\.
- \[3\]J\. Chen, Y\. He, Y\. Geng, E\. Jiménez\-Ruiz, H\. Dong, and I\. Horrocks\(2023\-09\)Contextual semantic embeddings for ontology subsumption prediction\.World Wide Web26\(5\),pp\. 2569–2591\(en\)\.External Links:ISSN 1573\-1413,[Link](https://doi.org/10.1007/s11280-023-01169-9),[Document](https://dx.doi.org/10.1007/s11280-023-01169-9)Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p2.1),[§4\.1](https://arxiv.org/html/2607.27130#S4.SS1.SSS0.Px1.p2.1),[§4\.2](https://arxiv.org/html/2607.27130#S4.SS2.SSS0.Px3.p1.2),[§5\.2](https://arxiv.org/html/2607.27130#S5.SS2.p2.1)\.
- \[4\]D\. M\. Dooley, E\. J\. Griffiths, G\. S\. Gosal, P\. L\. Buttigieg, R\. Hoehndorf, M\. C\. Lange, L\. M\. Schriml, F\. S\. Brinkman, and W\. W\. Hsiao\(2018\)FoodOn: a harmonized food ontology to increase global food traceability, quality control and data integration\.npj Science of Food2\(1\),pp\. 23\.Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p1.1)\.
- \[5\]A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Yang, A\. Fan,et al\.\(2024\)The llama 3 herd of models\.External Links:2407\.21783Cited by:[§4\.6](https://arxiv.org/html/2607.27130#S4.SS6.p1.1)\.
- \[6\]D\. Faria, C\. Pesquita, E\. Santos, M\. Palmonari, I\. F\. Cruz, and F\. M\. Couto\(2013\)The AgreementMakerLight Ontology Matching System\.InOn the Move to Meaningful Internet Systems: OTM 2013 Conferences,R\. Meersman, H\. Panetto, T\. Dillon, J\. Eder, Z\. Bellahsene, N\. Ritter, P\. De Leenheer, and D\. Dou \(Eds\.\),Berlin, Heidelberg,pp\. 527–541\(en\)\.External Links:ISBN 978\-3\-642\-41030\-7,[Document](https://dx.doi.org/10.1007/978-3-642-41030-7%5F38)Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p2.1),[§4\.2](https://arxiv.org/html/2607.27130#S4.SS2.SSS0.Px2.p1.2),[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p1.1)\.
- \[7\]H\. B\. Giglou, J\. D’Souza, F\. Engel, and S\. Auer\(2024\-04\)LLMs4OM: Matching Ontologies with Large Language Models\.arXiv\.Note:arXiv:2404\.10317External Links:[Link](http://arxiv.org/abs/2404.10317),[Document](https://dx.doi.org/10.48550/arXiv.2404.10317)Cited by:[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p3.1)\.
- \[8\]Y\. He, J\. Chen, D\. Antonyrajah, and I\. Horrocks\(2022\-06\)BERTMap: A BERT\-Based Ontology Alignment System\.Proceedings of the AAAI Conference on Artificial Intelligence36\(5\),pp\. 5684–5691\(en\)\.Note:Number: 5External Links:ISSN 2374\-3468,[Link](https://ojs.aaai.org/index.php/AAAI/article/view/20510),[Document](https://dx.doi.org/10.1609/aaai.v36i5.20510)Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p2.1),[§4\.2](https://arxiv.org/html/2607.27130#S4.SS2.SSS0.Px2.p1.2),[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p2.1)\.
- \[9\]Y\. He, J\. Chen, H\. Dong, and I\. Horrocks\(2023\-09\)Exploring Large Language Models for Ontology Alignment\.arXiv\.Note:arXiv:2309\.07172External Links:[Link](http://arxiv.org/abs/2309.07172),[Document](https://dx.doi.org/10.48550/arXiv.2309.07172)Cited by:[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p3.1)\.
- \[10\]Y\. He, J\. Chen, H\. Dong, E\. Jiménez\-Ruiz, A\. Hadian, and I\. Horrocks\(2022\)Machine Learning\-Friendly Biomedical Datasets for Equivalence and Subsumption Ontology Matching\.InThe Semantic Web – ISWC 2022,U\. Sattler, A\. Hogan, M\. Keet, V\. Presutti, J\. P\. A\. Almeida, H\. Takeda, P\. Monnin, G\. Pirrò, and C\. d’Amato \(Eds\.\),Cham,pp\. 575–591\(en\)\.External Links:ISBN 978\-3\-031\-19433\-7,[Document](https://dx.doi.org/10.1007/978-3-031-19433-7%5F33)Cited by:[§4\.1](https://arxiv.org/html/2607.27130#S4.SS1.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2607.27130#S4.SS1.SSS0.Px1.p2.1),[§5\.2](https://arxiv.org/html/2607.27130#S5.SS2.p1.1)\.
- \[11\]Y\. He, Z\. Yuan, J\. Chen, and I\. Horrocks\(2024\)Language models as hierarchy encoders\.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA\.External Links:ISBN 9798331314385Cited by:[§4\.2](https://arxiv.org/html/2607.27130#S4.SS2.SSS0.Px3.p1.2),[§5\.2](https://arxiv.org/html/2607.27130#S5.SS2.p3.1)\.
- \[12\]S\. Hertling and H\. Paulheim\(2023\-12\)OLaLa: Ontology Matching with Large Language Models\.InProceedings of the 12th Knowledge Capture Conference 2023,K\-CAP ’23,New York, NY, USA,pp\. 131–139\.External Links:ISBN 979\-8\-4007\-0141\-2,[Link](https://dl.acm.org/doi/10.1145/3587259.3627571),[Document](https://dx.doi.org/10.1145/3587259.3627571)Cited by:[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p3.1)\.
- \[13\]A\. Q\. Jiang, A\. Sablayrolles, A\. Roux, A\. Mensch, B\. Savary, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, E\. B\. Hanna, F\. Bressand,et al\.\(2024\)Mixtral of experts\.External Links:2401\.04088,[Document](https://dx.doi.org/10.48550/arXiv.2401.04088)Cited by:[§4\.6](https://arxiv.org/html/2607.27130#S4.SS6.p1.1)\.
- \[14\]E\. Jiménez\-Ruiz and B\. Cuenca Grau\(2011\)LogMap: logic\-based and scalable ontology matching\.InThe Semantic Web–ISWC 2011,Lecture Notes in Computer Science, Vol\.7031,pp\. 273–288\.Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p2.1),[§4\.2](https://arxiv.org/html/2607.27130#S4.SS2.SSS0.Px2.p1.2),[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p1.1)\.
- \[15\]Y\. Liu, E\. Pena, A\. Santos, E\. Wu, and J\. Freire\(2025\-06\)Magneto: Combining Small and Large Language Models for Schema Matching\.arXiv\.Note:arXiv:2412\.08194 \[cs\]External Links:[Link](http://arxiv.org/abs/2412.08194),[Document](https://dx.doi.org/10.48550/arXiv.2412.08194)Cited by:[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p2.1)\.
- \[16\]S\. Lushnei, D\. Shumskyi, S\. Shykula, E\. Jiménez\-Ruiz, and A\. d\. Garcez\(2026\-03\)Large language models as oracles for ontology alignment\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 2435–2449\.External Links:[Link](https://aclanthology.org/2026.eacl-long.110/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.110),ISBN 979\-8\-89176\-380\-7Cited by:[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p3.1)\.
- \[17\]A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang,et al\.\(2023\)Self\-refine: iterative refinement with self\-feedback\.Advances in neural information processing systems36,pp\. 46534–46594\.Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p4.1)\.
- \[18\]S\. Menad, W\. Laddada, S\. Abdeddaïm, and L\. Soualmia\(2023\)BioSTransformers for Biomedical Ontologies Alignment:\.InProceedings of the 15th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management,Rome, Italy,pp\. 73–84\(en\)\.External Links:ISBN 978\-989\-758\-671\-2,[Link](https://www.scitepress.org/DigitalLibrary/Link.aspx?doi=10.5220/0012188600003598),[Document](https://dx.doi.org/10.5220/0012188600003598)Cited by:[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p2.1)\.
- \[19\]OpenAI\(2025\)OpenAI agents sdk: handoffs\.Note:[https://openai\.github\.io/openai\-agents\-python/handoffs/](https://openai.github.io/openai-agents-python/handoffs/)Accessed: 2026\-07\-21Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p4.1)\.
- \[20\]L\. Otero\-Cerdeira, F\. J\. Rodríguez\-Martínez, and A\. Gómez\-Rodríguez\(2015\)Ontology matching: a literature review\.Expert Systems with Applications42\(2\),pp\. 949–971\.Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p1.1)\.
- \[21\]S\. Oulefki, L\. Berkani, N\. Boudjenah, L\. Bellatreche, and A\. Mokhtari\(2025\-09\)BioGITOM: Matching Biomedical Ontologies with Graph Isomorphism Transformer\.The VLDB Journal34\(6\),pp\. 65\(en\)\.External Links:ISSN 0949\-877X,[Link](https://doi.org/10.1007/s00778-025-00943-7),[Document](https://dx.doi.org/10.1007/s00778-025-00943-7)Cited by:[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p2.1)\.
- \[22\]Y\. Song, J\. Chen, and R\. A\. Schmidt\(2026\)GenOM: ontology matching with description generation and large language models\.World Wide Web29\(3\),pp\. 29\.Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p2.1),[§4\.2](https://arxiv.org/html/2607.27130#S4.SS2.SSS0.Px2.p1.2),[§5\.1](https://arxiv.org/html/2607.27130#S5.SS1.p3.1)\.
- \[23\]S\. Staab and R\. Studer\(2013\)Handbook on ontologies\.Springer Science & Business Media\.Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p1.1)\.
- \[24\]J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, E\. H\. Chi, Q\. Le, and D\. Zhou\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.35,pp\. 24824–24837\.External Links:[Link](https://neurips.cc/)Cited by:[§4\.2](https://arxiv.org/html/2607.27130#S4.SS2.SSS0.Px1.p1.2)\.
- \[25\]A\. Yanget al\.\(2024\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[§4\.4](https://arxiv.org/html/2607.27130#S4.SS4.SSS0.Px2.p1.1),[§4\.6](https://arxiv.org/html/2607.27130#S4.SS6.p1.1)\.
- \[26\]H\. Yang, J\. Chen, Y\. He, Y\. Gao, and I\. Horrocks\(2025\)Language models as ontology encoders\.InThe Semantic Web – ISWC 2025: 24th International Semantic Web Conference, Nara, Japan, November 2–6, 2025, Proceedings, Part I,Berlin, Heidelberg,pp\. 443–461\.External Links:ISBN 978\-3\-032\-09526\-8,[Link](https://doi.org/10.1007/978-3-032-09527-5_24),[Document](https://dx.doi.org/10.1007/978-3-032-09527-5%5F24)Cited by:[§4\.2](https://arxiv.org/html/2607.27130#S4.SS2.SSS0.Px3.p1.2),[§5\.2](https://arxiv.org/html/2607.27130#S5.SS2.p3.1)\.
- \[27\]S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. Cao\(2022\)React: synergizing reasoning and acting in language models\.arXiv preprint arXiv:2210\.03629\.Cited by:[§1](https://arxiv.org/html/2607.27130#S1.p4.1)\.
## Appendix AImplementation Details
The benchmark is split into equivalence and subsumption evaluation subsets using a fixed random seed of 42 to ensure reproducibility\. All open\-source models, including the open\-source backbone LLMs \(Qwen2\.5\-32B\-Instruct, Qwen2\.5\-72B\-Instruct, Llama3\.1\-70B\-Instruct, Mixtral\-8x7B\-Instruct\) and the local embedding, equivalence and subsumption baselines are run on two NVIDIA A100 GPUs\.
## Appendix BPrompt Templates
### B\.1Agent Initial Equivalence Screening Prompt
Agent A Prompt: Initial Equivalence Screening \(System Message\)Given one source concept and a list of candidate concepts\.Task: Determine whether any candidate is the correct equivalent concept of the source\. Equivalent means:•refers to the same real\-world entity or concept•same meaning•same level of specificityIf no exact equivalent exists, selectNONE\.Rules:•Select exactly one candidate ID, orNONE\.•Do not invent candidate IDs\.Return exactly this format:``` Reasoning: <short reasoning> Selected: <candidate_id or NONE> ``` The source concept and candidate list are supplied via the user message at inference time\.
### B\.2Agent Equivalence Verification Prompt
Agent Prompt: Equivalence Verification \(System Message\)Given one source concept and a small list of candidate concepts\.Task: Select the single best candidate as the final equivalence result\. The correct result should:•have the same core meaning as the source•match the source granularity bestRules:•Always select exactly one candidate ID from the given list\.•Do not invent candidate IDs\.Return exactly this format and no extra text:``` Reasoning: <short reasoning> Selected: <candidate_id> ``` The source concept and the updated candidate setCref=\{c^t,p∗,ch∗\}C\_\{\\text\{ref\}\}=\\\{\\hat\{c\}\_\{t\},p^\{\*\},ch^\{\*\}\\\}are supplied via the user message at inference time\.
### B\.3Agent Subsumption Discovery Prompt
Agent Prompt \(Iterative Search\): Subsumption Discovery \(System Message\)Given one source concept and a list of candidate concepts\.Task: Choose the nearest broader candidate for the source\. If none is clearly broader, chooseNONE\.Rules:•A valid choice must be broader than the source\.•It must be the closest available parent\-level concept\.•Do not choose candidates that are only related, part\-based, sibling\-level, or more specific\.•Use only the given labels and synonyms\.•Choose exactly one candidate ID, orNONE\.Return exactly this format:``` Reasoning: <short reasoning> Selected: <candidate_id or NONE> ``` The source concept and the current candidate setCiC\_\{i\}are supplied via the user message at inference time\. IfNONEis returned, the candidate set is expanded toCi\+1=Parents\(Ci\)C\_\{i\+1\}=Parents\(C\_\{i\}\)and this prompt is reapplied\.
Agent Prompt \(Fallback\): Final Subsumer Selection \(System Message\)Given one source concept and a list of candidate concepts\.Task: Select the candidate that is the closest broader concept of the source concept\.Definitions:•A broader concept is a parent\-level or ancestor\-level concept that can subsume the source\.•The selected candidate should be semantically broader than the source, not equivalent to it, not narrower than it, and not merely related to it\.•If multiple candidates are broader, select the most specific and closest broader candidate\.•You must select exactly one candidate from the list\.Reasoning procedure:1\.Identify whether each plausible candidate is broader than the source\.2\.Exclude candidates that are equivalent, narrower, sibling concepts, parts, attributes, or merely related concepts\.3\.Among the remaining broader candidates, choose the closest and most specific one\.4\.If no candidate is a perfect direct parent, choose the best available broader candidate\.Return exactly this format:``` Reasoning: <short reasoning> Selected: <candidate_id> ``` This prompt is applied once, afterdmaxd\_\{\\max\}upward expansions, over the complete search trajectoryC0∪C1∪⋯∪CdmaxC\_\{0\}\\cup C\_\{1\}\\cup\\cdots\\cup C\_\{d\_\{\\max\}\}\.
### B\.4Conflict Resolution Prompt: Equivalence Arbitration
Conflict Resolution Prompt: Equivalence Arbitration \(System Message\)You are an ontology alignment arbitration agent\. Your task is to choose which target concept is more appropriate as the equivalence match for the source concept\. Choose exactly one option:AorB\.Definitions:•equivalencemeans that the source concept and target concept refer to the same or nearly identical concept at the same conceptual granularity\.Rules:•Use only labels and synonyms\.•Do not use or infer from IRIs\.•Do not introduce a new target concept\.•Do not output subsumption\.•Briefly explain the conflict\.•Then output the final choice\.Return exactly this format:``` Reason: <short reasoning> Choice: A or B ``` This prompt is invoked when the lexical matching module and the agent\-based reasoning module output two different equivalence target concepts \(options A and B, supplied via the user message\), following the third case of the conflict resolution rules in the Lexical Matching and Conflict Resolution section\.
### B\.5HOM Baseline Prompts
Direct Prompting Baseline \(System Message\)Given one source concept and a list of target candidate concepts\.Task: Select exactly one candidate and assign exactly one relation:equivalenceorsubsumption\.Definitions:•equivalencemeans the candidate refers to the same real world concept as the source\.•subsumptionmeans the candidate is broader than the source and can act as a direct or near direct parent level concept\.Rules:•Chooseequivalenceif one candidate has the same meaning as the source\.•If no equivalent candidate exists, choose the nearest broader candidate\.•Do not choose siblings, children, parts, or merely related concepts\.•Prefer the most specific broader candidate when choosingsubsumption\.•You must choose exactly one candidate from the given list\.•Use only the provided labels and synonyms\.•Candidate IDs are only placeholders\. They do not contain semantic information\.Output JSON only:``` { "candidate_id": "...", "relation": "equivalence" or "subsumption" } ``` The source concept and the candidate set are supplied via the user message at inference time\. This prompt is used for both the LLM \+C0C\_\{0\}and LLM \+ Neighbourhood baselines, differing only in the candidate set provided\.
Chain\-of\-Thought Prompting Baseline \(System Message\)Given one source concept and a list of target candidate concepts\.Task: Select exactly one candidate and assign exactly one relation:equivalenceorsubsumption\.Definitions:•equivalencemeans the candidate refers to the same real world concept as the source\.•subsumptionmeans the candidate is broader than the source and can act as a direct or near direct parent level concept\.Rules:•Chooseequivalenceif one candidate has the same meaning as the source\.•If no equivalent candidate exists, choose the nearest broader candidate\.•Do not choose siblings, children, parts, or merely related concepts\.•Prefer the most specific broader candidate when choosingsubsumption\.•You must choose exactly one candidate from the given list\.•Use only the provided labels and synonyms\.•Candidate IDs are only placeholders\. They do not contain semantic information\.•Keep the reasoning concise\.•The reasoning should be natural language\.•The final answer must be a JSON object\.Output format:``` Reasoning: <concise natural-language reasoning> Final answer: { "candidate_id": "...", "relation": "equivalence" or "subsumption" } ``` The source concept and the candidate set are supplied via the user message at inference time\. This prompt is used for both the LLM \+C0C\_\{0\}\+ CoT and LLM \+ Neighbourhood \+ CoT baselines, differing only in the candidate set provided\.
## Appendix CEffect of Open\-Source Resources
### C\.1Closed\-Source vs\. Open\-Source Configurations
Table[7](https://arxiv.org/html/2607.27130#A3.T7)reports results on SNOMED–FMA–Body under four configurations, combining OpenAI or SBERT embeddings with GPT\-4\.1\-mini or Qwen2\.5\-32B\-Instruct as the backbone\. The fully closed\-source configuration achieves the best performance \(0\.657 overall\), while the fully open\-source configuration \(SBERT \+ Qwen2\.5\-32B\-Instruct\) trails by 13\.3 points \(0\.524\), a moderate gap given the removal of all commercial components\. Comparing the two hybrid configurations suggests that the embedding model contributes more to this gap than the backbone LLM: replacing the backbone alone \(OpenAI \+ Qwen2\.5\-32B\-Instruct\) costs only 2\.2 points relative to the closed\-source setting, whereas replacing the embedding model alone \(SBERT \+ GPT\-4\.1\-mini\) costs 11\.2 points, driven primarily by a large drop in equivalence accuracy \(0\.957 to 0\.777\)\. We examine this effect of embedding choice in more detail across all four benchmarks in the next subsection\.
Table 7:Performance of AgentMap under closed\-source, hybrid, and open\-source settings on the SNOMED–FMA Body task\.
### C\.2Effect of Embedding Model Across Benchmarks
While Table[7](https://arxiv.org/html/2607.27130#A3.T7)controls for both factors on a single benchmark, we further isolate the effect of the embedding model alone by fixing the backbone to Qwen2\.5\-32B\-Instruct across all four benchmarks\. Table[8](https://arxiv.org/html/2607.27130#A3.T8)reports the results\.
OpenAI embeddings outperform SBERT on both equivalence and subsumption accuracy on three of the four benchmarks \(Body, Pharm, and Disease\), indicating that retrieval quality generally affects both stages of AgentMap\. However, which stage is more affected varies substantially across datasets: on SNOMED–FMA–Body, the embedding choice affects equivalence more \(a 21\.0% relative drop with SBERT\) than subsumption \(6\.9 %\), whereas on SNOMED–NCIT–Pharm the pattern reverses sharply, with subsumption accuracy dropping by over half \(52\.3%\) while equivalence remains largely unaffected \(0\.6%\)\. NCIT–DOID–Disease shows a more balanced effect on both metrics\. This inconsistency suggests that the relative sensitivity of each stage to retrieval quality is dataset\-dependent, though identifying the specific underlying factors is beyond the scope of this analysis\.
On HeLiS–FoodOn, the trend reverses entirely: SBERT achieves higher overall and subsumption accuracy than OpenAI embeddings, while OpenAI retains an advantage on equivalence\. Given the small size of this dataset \(174 equivalence and 190 subsumption instances\), this reversal likely reflects sampling variance rather than a systematic advantage of SBERT\. Overall, these results indicate that while AgentMap remains functional with a fully open\-source embedding model, retrieval quality has a measurable and dataset\-dependent impact on both equivalence and subsumption performance\.
Table 8:Effect of embedding model \(OpenAItext\-embedding\-3\-smallvs\. SBERTall\-MiniLM\-L6\-v2\) on HOM performance across the four benchmarks, with Qwen2\.5\-32B\-Instruct fixed as the backbone LLM\. Bold indicates the better result within each dataset\.
## Appendix DEffect of Candidate Set Size \(Top\-kk\)
We further examine the sensitivity of AgentMap to the size of the agent reasoning candidate setC0C\_\{0\}, using Qwen2\.5\-32B\-Instruct as the backbone, withk∈\{5,7,10,15\}k\\in\\\{5,7,10,15\\\}\. Figure[2](https://arxiv.org/html/2607.27130#A4.F2)reports the results across all four benchmarks\.
Askkincreases, Overall and Sub accuracy decrease monotonically, while Eqv accuracy rises only marginally\. This indicates that a larger candidate set does not meaningfully improve equivalence discovery, since the correct equivalent concept is typically already covered by a smallkk, but it introduces more distractor candidates into Agent C’s subsumption search, increasing the likelihood of selecting an incorrect broader concept\. This trend is consistent across the three biomedical benchmarks; HeLiS–FoodOn shows a similar pattern but with more fluctuation, reflecting its much smaller test set size\. These results support our choice ofk=5k=5as the default configuration, which achieves the best overall and subsumption accuracy while minimising the LLM reasoning cost associated with a larger candidate set\.

\(a\) SNOMED–FMA–Body

\(b\) SNOMED–NCIT–Pharm

\(c\) NCIT–DOID–Disease

\(d\) HeLiS–FoodOn
Figure 2:Effect of candidate set sizekkon HOM results across the four benchmarks, using Qwen2\.5\-32B\-Instruct as the backbone\.相似文章
Org-Agent:超越个人助理,走向组织级智能体
论文提出 Org-Agent,这是一个统一的以约束为中心的推理框架,能够使 LLM 智能体协调来自多个用户的请求,并利用分布在交互过程中的知识,通过依赖图和拓扑调度对任务进行分解。在 MUSES-Bench 和 GroupMemBench 上的实验表明,该框架在跨用户决策和记忆能力方面均有所提升。
HMACE:面向组合优化的异构多智能体协同进化
本文介绍了 HMACE,这是一种异构多智能体协同进化框架,利用大型语言模型(LLM)自动化设计启发式算法,以解决 NP 难组合优化问题。实验表明,在旅行商问题(TSP)和装箱问题(BPP)等任务上,该方法在质量与效率的权衡方面优于单智能体和基准多智能体方法。
CoWeaver:面向人与智能体混合科学协作的双向、可学习且可解释的匹配引擎
CoWeaver是一种双向、可学习且可解释的匹配引擎,旨在通过填补能力差距、两阶段排序和不确定性感知探索,在科学网络中促成人类与基于LLM的智能体之间的强大协作。
OPINE-World: 使用本体错误优先的交互式探索进行程序化世界建模
OPINE-World 引入了一个 LLM 智能体,通过交互在线学习以对象为中心的程序化世界模型,采用本体错误优先的探索和协作的假设-测试智能体,在 ARC-AGI-3 上取得了强劲的结果。
HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents
HyperAgent is a research framework that models tool relations via a Tool-Schema Hypergraph to improve planning and execution for LLM agents, reducing API calls and token usage on the AppWorld benchmark.