On Measuring Semantic Preservation in Legal Ontology Learning
Summary
This paper proposes a task-based evaluation methodology for measuring semantic preservation in ontology learning, comparing LLM performance on source documents versus transformed representations in the legal domain.
View Cached Full Text
Cached at: 08/14/26, 09:24 AM
# On Measuring Semantic Preservation in Legal Ontology Learning Source: [https://arxiv.org/html/2608.12326](https://arxiv.org/html/2608.12326) Jarosław A\. Chudziak Warsaw University of Technology Warsaw, Poland jaroslaw\.chudziak@pw\.edu\.pl ###### Abstract Ontology learning transforms unstructured text into structured representations for automated reasoning\. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on structural correctness while failing to measure whether meaning survives transformation\. We propose an evaluation methodology that addresses this: comparing LLM task performance on source documents against performance on transformed representations, with the difference quantifying semantic loss\. We demonstrate this approach on legal merger agreement analysis, a domain chosen for its complex language and precise semantic requirements, comparing direct LLM application against three ontology learning methods across six language models\. The results reveal systematic semantic loss with significant variation based on reasoning complexity and model\-method interactions\. Our contributions are: \(1\) an evaluation framework for measuring semantic preservation in ontology learning, and \(2\) empirical evidence that semantic loss varies dramatically with model\-method pairing, providing guidance for selecting optimal configurations in legal knowledge systems\. *K*eywordsOntology Learning⋅\\cdotSemantic Preservation⋅\\cdotLegal Ontology⋅\\cdotLLM Evaluation⋅\\cdotKnowledge Representation⋅\\cdotTask\-Based Evaluation ## 1Introduction Ontology learning seeks to automatically transform unstructured text into structured, machine\-readable representations that can support reasoning and integration across systems\[[4](https://arxiv.org/html/2608.12326#bib.bib4),[18](https://arxiv.org/html/2608.12326#bib.bib20)\]\. This approach has attracted renewed interest in specialized domains, particularly in law\[[10](https://arxiv.org/html/2608.12326#bib.bib24),[5](https://arxiv.org/html/2608.12326#bib.bib5)\], where the volume and complexity of textual information creates challenges for manual knowledge organization efforts\. Ontologies play a complementary role alongside large language models \(LLMs\) in modern knowledge systems\. While natural language serves humans well, it functions poorly as a communication protocol between computational systems due to inherent ambiguity and inconsistency\[[20](https://arxiv.org/html/2608.12326#bib.bib22)\]\. Ontologies provide controlled vocabularies essential for data interoperability, transparent knowledge representations, and logical consistency requirements that LLMs cannot guarantee\[[20](https://arxiv.org/html/2608.12326#bib.bib22)\]\. This creates opportunities for hybrid architectures where LLMs interface with symbolic reasoning systems through structured ontological representations\[[24](https://arxiv.org/html/2608.12326#bib.bib28),[17](https://arxiv.org/html/2608.12326#bib.bib37)\]\. However, a critical question remains underexplored: when we transform natural language text into ontological representations, do we lose the semantic content needed for downstream tasks? Current evaluation methodologies focus on structural correctness\[[2](https://arxiv.org/html/2608.12326#bib.bib2),[3](https://arxiv.org/html/2608.12326#bib.bib3)\], assessing whether ontologies conform to formal specifications and exhibit logical consistency\. These measures ensure technical validity but cannot detect semantic loss\. An ontology may pass all structural tests yet fail to support the reasoning tasks for which it was intended\. Without methods to measure semantic preservation, practitioners cannot make evidence\-based decisions about when ontological transformation helps or hinders their applications\. Figure 1:Overview of our evaluation methodology\. Left: source documents \(baseline 100% accessible content\) are transformed via three ontology learning methods, each introducing semantic loss \(average accuracy, LLMs4OL: \-21pp, NeOn\-GPT: \-13pp, NeOn\-CoT: \-16pp\)\. Right: traditional evaluation validates structure but cannot detect information loss; our proposed task\-based approach measures semantic preservation by comparing LLM accuracy on identical tasks before and after transformation\.We propose an evaluation methodology that addresses this gap by using LLM performance as a proxy for semantic accessibility\. Our approach treats baseline LLM performance on source documents as a model\-specific upper bound on accessible semantic content, then measures how much of that content remains accessible after ontological transformation\. This methodology builds on task\-based evaluation paradigms from NLP\[[16](https://arxiv.org/html/2608.12326#bib.bib18),[22](https://arxiv.org/html/2608.12326#bib.bib26)\]and recent behavioral testing frameworks\[[23](https://arxiv.org/html/2608.12326#bib.bib27)\]\. Rather than claiming to measure absolute semantic content, we quantify the relative preservation of information that a given model can extract, enabling direct comparison between source documents and their ontological transformations under consistent conditions\. We demonstrate this approach using the MAUD dataset from LegalBench\[[13](https://arxiv.org/html/2608.12326#bib.bib13),[29](https://arxiv.org/html/2608.12326#bib.bib33)\], comparing direct LLM application against three ontology learning methods: LLMs4OL\[[1](https://arxiv.org/html/2608.12326#bib.bib1)\], NeOn\-GPT\[[11](https://arxiv.org/html/2608.12326#bib.bib11)\], and a simplified NeOn\-CoT method we introduce\. Our evaluation spans six language models to ensure robust findings across different architectures\. Results reveal systematic semantic loss across all approaches, with particularly severe degradation for complex legal reasoning tasks\. The contributions of this work are: \(1\) an evaluation framework for measuring semantic preservation in ontology learning, and \(2\) empirical evidence that semantic loss in legal ontology learning varies dramatically with model\-method pairing, providing guidance for selecting optimal configurations\. The remainder of this paper is organized as follows\. Section[2](https://arxiv.org/html/2608.12326#S2)reviews related work\. Section[3](https://arxiv.org/html/2608.12326#S3)presents the proposed methodology, the three ontology learning approaches, and the experimental setup\. Section[4](https://arxiv.org/html/2608.12326#S4)reports results, including a task\-level analysis\. Sections[5](https://arxiv.org/html/2608.12326#S5)and[6](https://arxiv.org/html/2608.12326#S6)discuss implications, limitations, and future work\. Section[7](https://arxiv.org/html/2608.12326#S7)concludes\. ## 2Related Work Ontology engineering has evolved from traditional manual approaches\[[5](https://arxiv.org/html/2608.12326#bib.bib5),[10](https://arxiv.org/html/2608.12326#bib.bib24)\]to recent LLM\-based automated methods\[[1](https://arxiv.org/html/2608.12326#bib.bib1),[11](https://arxiv.org/html/2608.12326#bib.bib11)\]\. Traditional approaches used lexico\-syntactic pattern mining and clustering\[[14](https://arxiv.org/html/2608.12326#bib.bib15)\], frameworks like Text2Onto\[[7](https://arxiv.org/html/2608.12326#bib.bib8)\], and the NeOn methodology for structured ontology development\[[25](https://arxiv.org/html/2608.12326#bib.bib29)\]\. Early work explored semantic nets as meta\-level representations for maintaining semantic accessibility across heterogeneous data types\[[6](https://arxiv.org/html/2608.12326#bib.bib7)\]\. Legal ontology projects drew on jurisprudential foundations and employed bottom\-up, top\-down, and hybrid development strategies\[[10](https://arxiv.org/html/2608.12326#bib.bib24)\]\. Construction methodologies included Methontology\[[9](https://arxiv.org/html/2608.12326#bib.bib10)\], Ontology Development 101\[[21](https://arxiv.org/html/2608.12326#bib.bib23)\], and CommonKADS\[[26](https://arxiv.org/html/2608.12326#bib.bib30)\]\. Systematic analysis showed 93\.5% of legal ontology development was manual, with 41% lacking validation activity\[[10](https://arxiv.org/html/2608.12326#bib.bib24)\]\. Evaluation followed Gold Standard, Case\-Based, Data\-Driven, and User\-Based approaches\[[2](https://arxiv.org/html/2608.12326#bib.bib2)\]\. The legal domain posed challenges including complex writing styles, heterogeneous sources, and evolving documents that contributed to syntactic and semantic anomalies\[[10](https://arxiv.org/html/2608.12326#bib.bib24)\]\. LLMs introduced new paradigms based on automated knowledge extraction\[[1](https://arxiv.org/html/2608.12326#bib.bib1)\]\. LLMs4OL addresses term typing \(MAP@1\), taxonomy discovery, and relation extraction \(F1\-scores\), reporting term typing as the most challenging task \(60\-80% baseline\) with fine\-tuning gains of 25%, 18%, and 3% respectively\[[1](https://arxiv.org/html/2608.12326#bib.bib1)\]\. NeOn\-GPT combines structured methodology with LLM capabilities, using structural metrics such as axiom counts and hierarchical depth\[[11](https://arxiv.org/html/2608.12326#bib.bib11)\]\. Both approaches focus on structural correctness rather than knowledge representation quality or semantic preservation\[[1](https://arxiv.org/html/2608.12326#bib.bib1),[11](https://arxiv.org/html/2608.12326#bib.bib11)\]\. Structural metrics fail to capture semantic richness or semantic loss during transformation, particularly in specialized domains like law\[[10](https://arxiv.org/html/2608.12326#bib.bib24)\]\. Task\-based evaluation offers an alternative perspective\. Sparck Jones and Galliers established the distinction between intrinsic and extrinsic evaluation\[[16](https://arxiv.org/html/2608.12326#bib.bib18)\], and Mollá and Hutchinson showed disconnects between structural accuracy and downstream performance\[[19](https://arxiv.org/html/2608.12326#bib.bib21)\]\. For knowledge systems, Porzel and Malaka pioneered task\-based ontology evaluation\[[22](https://arxiv.org/html/2608.12326#bib.bib26)\], extended by Gangemi et al\.\[[12](https://arxiv.org/html/2608.12326#bib.bib12)\]and Brewster et al\.\[[3](https://arxiv.org/html/2608.12326#bib.bib3)\]\. Similar to behavioral testing frameworks like CheckList\[[23](https://arxiv.org/html/2608.12326#bib.bib27)\], we treat LLM performance as a signal for evaluating representation quality\. Information\-theoretic approaches provide quantitative frameworks for measuring preservation during representation transformations\. Knowledge distillation research introduced loyalty measures using Jensen\-Shannon divergence\[[30](https://arxiv.org/html/2608.12326#bib.bib34)\]and methods for measuring semantic retention across representational levels\[[15](https://arxiv.org/html/2608.12326#bib.bib17)\]\. Multi\-scale mutual information methods show that preservation must be measured across granularities to capture local and global coherence\[[31](https://arxiv.org/html/2608.12326#bib.bib35)\]\. LLM\-as\-evaluator studies show strong LLMs reach over 80% agreement with human preferences\[[32](https://arxiv.org/html/2608.12326#bib.bib36)\], though bias studies note limitations that require careful calibration\[[28](https://arxiv.org/html/2608.12326#bib.bib32)\]\. Ensemble approaches with diverse LLM evaluators show promise for reliable assessment\[[27](https://arxiv.org/html/2608.12326#bib.bib31)\]\. These results support our use of baseline LLM performance as a reference point for measuring semantic preservation during ontological transformation\. ## 3Methodology Current ontology learning evaluation methods focus on structural correctness rather than knowledge transformation quality\[[2](https://arxiv.org/html/2608.12326#bib.bib2),[1](https://arxiv.org/html/2608.12326#bib.bib1),[11](https://arxiv.org/html/2608.12326#bib.bib11)\]\. Building on the task\-based evaluation paradigm from NLP \(Section[2](https://arxiv.org/html/2608.12326#S2)\), we propose an evaluation framework that quantifies semantic preservation by using LLM performance on downstream tasks as a proxy for semantic accessibility\[[22](https://arxiv.org/html/2608.12326#bib.bib26),[23](https://arxiv.org/html/2608.12326#bib.bib27)\]\. Our approach treats baseline LLM performance on source documents as the total accessible content extractable by that LLM, defined asLLM\(text,question\)→answerLLM\(text,question\)\\rightarrow answerperformance on original legal text\. This baseline establishes an upper bound for semantic accessibility from that model’s perspective, following the principle that extrinsic task performance provides a more meaningful signal than intrinsic structural measures\[[16](https://arxiv.org/html/2608.12326#bib.bib18),[19](https://arxiv.org/html/2608.12326#bib.bib21)\]\. While ontological transformation may intentionally prioritize structural organization and formal reasoning capabilities over complete semantic preservation, measuring accessibility trade\-offs enables informed decisions about when ontological transformation enhances versus diminishes practical utility for specific applications\. We quantify this as:SemanticLoss=BaselinePerf−TransformedPerfSemanticLoss=BaselinePerf\-TransformedPerf\. This framework provides domain\-agnostic assessment applicable across legal specializations without requiring domain\-specific gold standards, measures semantic preservation from a user\-centric perspective focusing on practical utility for downstream applications, and enables systematic comparison of multiple ontology learning approaches\. Importantly, our methodology measures semantic preservation relative to each LLM’s capabilities, accounting for model\-specific baseline variations\. ### 3\.1Ontology Learning Approaches We evaluate three ontology learning approaches against the direct LLM application baseline established in the previous section\. Each approach transforms source legal documents into structured ontological representations using different computational strategies, representing different levels of methodological structure and computational complexity, as well as varying degrees of domain awareness\. The approaches vary in complexity from parallel task decomposition to sequential methodology\-guided construction, and differ in their degree of domain targeting \- from completely domain\-agnostic \(LLMs4OL\) to explicitly domain\-aware \(NeOn\-GPT and NeOn\-CoT\)\. This variation enables systematic analysis of how both structural complexity and domain specificity affect semantic preservation during ontological transformation\. #### LLMs4OL Based on the framework proposed by Babaei Giglou et al\.\[[1](https://arxiv.org/html/2608.12326#bib.bib1)\], this approach decomposes ontology construction into three parallel tasks: term typing \(classifying significant terms\), taxonomy discovery \(identifying hierarchical relationships\), and non\-taxonomic relation extraction \(discovering semantic relationships between entities\)\. The approach is designed to be domain\-agnostic, making no assumptions about the specific domain or downstream evaluation tasks\. Since the original LLMs4OL framework focuses on individual ontology learning tasks without integration, we extend the approach with an additional integration step that synthesizes results from all three tasks into a complete OWL ontology in Turtle format\. This integration step uses zero\-shot prompting to combine term classifications, hierarchical relationships, and semantic relationships into a coherent ontological representation without domain\-specific guidance\. #### NeOn\-GPT Following the methodology described by Fathallah et al\.\[[11](https://arxiv.org/html/2608.12326#bib.bib11)\], this approach implements a structured five\-stage pipeline: requirements specification, competency question generation, conceptual model creation, ontology draft implementation, and enrichment with instances and annotations\. Unlike LLMs4OL, this approach explicitly incorporates domain awareness by identifying the text as mergers and acquisitions \(M&A\) contract fragments and steering the ontology construction process toward legal document representation\. While the original NeOn\-GPT includes validation mechanisms \(syntax checking, consistency verification, and pitfall resolution\), we implement the core methodology stages without the validation loops, as syntactic correctness is less critical for our semantic preservation evaluation than substantive knowledge representation\. #### NeOn\-CoT This approach consolidates the NeOn methodology principles into a single chain\-of\-thought prompting strategy\. Similar to NeOn\-GPT, it maintains domain awareness by explicitly referencing the legal context and M&A domain\. The prompt guides the LLM through the conceptual phases of ontology development \(domain analysis, competency questions, entity extraction, conceptual modeling, and formal implementation\) within a unified processing step, reducing computational overhead while maintaining the structured reasoning approach\. ### 3\.2Representation Challenges and Constraints Our approach adopts two methodological constraints\. First, LLMs process ontological representations as textual strings rather than formal logical structures\. This aligns with our goal of measuring semantic accessibility from a user\-centric perspective, reflecting how practitioners typically interact with these representations\. Second, the framework operates under information conservation principles: ontology learning approaches reorganize existing content without external knowledge augmentation\. This enables attribution of performance differences to transformation effects rather than knowledge enrichment\. Suboptimal baseline LLM performance reflects model\-specific interpretation capabilities rather than insufficient source information\. These choices complement rather than replace evaluations of ontologies’ formal computational capabilities\. ### 3\.3Dataset and Experimental Setup Our evaluation employs the Merger Agreement Understanding Dataset \(MAUD\) from LegalBench\[[13](https://arxiv.org/html/2608.12326#bib.bib13),[29](https://arxiv.org/html/2608.12326#bib.bib33)\]\. MAUD contains 34 multiple\-choice tasks focusing on merger agreement analysis, including material adverse effect definitions, representations and warranties, and fiduciary obligations \(see Appendix[A](https://arxiv.org/html/2608.12326#A1)for complete task listing\)\. We use all 34 tasks\. To ensure balanced evaluation, we selected 69 examples from each task \(the minimum available across all tasks\), yielding 2,346 total examples\. For each model, we measure accuracy on the multiple\-choice tasks under two conditions: \(1\) baseline, where the model answers directly from source text using zero\-shot prompting, and \(2\) ontology\-transformed, where the model first generates an ontology, then answers from that ontology\. Each model generates its own ontologies, ensuring consistent capabilities between transformation and evaluation\. Temperature is set to 0 where supported\. Answers are extracted via structured outputs\. We report single\-run results; statistical significance is established across the 2,346 test examples\. The implementation uses Python 3\.13 with[LangChain 0\.3\.25](https://github.com/langchain-ai/langchain)and[LangGraph 0\.4\.8](https://www.langchain.com/langgraph)\. Models were accessed through the respective provider APIs: OpenAI, Anthropic, Google Gemini API, and Fireworks AI \(DeepSeek v3 and Llama 4 Maverick\)\. Complete prompts and the code are available in an online code repository111[https://github\.com/albsadowski/ontology\-learning\-eval](https://github.com/albsadowski/ontology-learning-eval)\. ## 4Results Our experimental results reveal substantial semantic loss across all ontology learning approaches, with performance degradation ranging from 8\.4 percentage points \(pp\) to 27\.4pp compared to baseline accessibility levels\. Table[1](https://arxiv.org/html/2608.12326#S4.T1)presents semantic loss measurements calculated as the percentage reduction from baseline performance for each model\-method combination\. Table 1:Information Preservation Analysis Across Models and Ontology Learning MethodsNote:Loss = information loss \(pp\)\. Success = success case accuracy\. Failure = failure case accuracy\. Bold = best per model; underlined = overall top performers\. LLMs4OL exhibits the most severe semantic loss \(13\.9\-27\.4pp\), likely due to information fragmentation during parallel task decomposition\. NeOn\-GPT demonstrates the best preservation characteristics \(8\.8\-18\.1pp\.\), suggesting that explicit domain awareness and sequential processing better preserve legal document semantics\. NeOn\-CoT shows intermediate performance with notable variability \(8\.4\-25\.3pp\.\), trading some preservation benefits for computational efficiency\. When baseline models successfully extract information, ontology learning methods retain 45\-83% of accessible content, with NeOn\-GPT achieving the highest success preservation rates \(60\-83%\)\. Analysis of failure cases reveals limited compensatory benefits from ontological transformation, with most approaches showing only 9\-36% accuracy on baseline failures, indicating that ontology learning primarily reorganizes existing accessible content rather than making previously inaccessible information available\. #### Cross\-Model Consistency The semantic loss patterns demonstrate notable consistency across diverse model architectures, providing evidence that observed degradation reflects systematic limitations of current ontology learning approaches rather than model\-specific artifacts\. The performance hierarchy NeOn\-GPT\>\>NeOn\-CoT\>\>LLMs4OL holds across all models tested, despite architectural differences between transformer variants \(GPT 4\.1 mini, Claude Sonnet 4\), mixture\-of\-experts systems \(DeepSeek v3\), and reasoning\-enhanced models \(o4 mini\)\. This consistency suggests structured, domain\-aware approaches systematically outperform parallel task decomposition methods for legal ontology learning\. However, absolute sensitivity varies considerably across models\. Claude Sonnet 4 shows higher sensitivity \(45\-61% preservation rates\), while Gemini 2\.5 Flash demonstrates more robust preservation \(57\-83% rates\), suggesting model selection represents a critical factor for practical applications\. Despite baseline performance differences \(59\.8% to 75\.14%\), all models exhibit comparable semantic loss magnitudes, with standard deviation below 0\.09 for each method, indicating the transformation process imposes consistent accessibility constraints regardless of initial model capabilities\. All performance differences between baseline and ontology learning methods were statistically significant \(Wilcoxon signed\-rank tests,p<0\.001p<0\.001\), with effect sizes ranging from medium \(d=0\.53d=0\.53for NeOn\-GPT\) to large \(d=0\.85d=0\.85for LLMs4OL\)\. Pairwise comparisons confirmed the observed performance hierarchy, with all method differences significant \(p<0\.001p<0\.001\)\. #### Semantic Loss Analysis To explore patterns of semantic loss in legal ontology learning, we analyzed MAUD tasks with sufficient baseline performance \(≥50%\\geq 50\\%\) to ensure meaningful statistical analysis\. Figure[2](https://arxiv.org/html/2608.12326#S4.F2)reveals systematic patterns varying by linguistic and logical complexity, ranging from minimal degradation for straightforward categorical determinations to severe semantic loss \(up to 65% average loss\) for complex multi\-standard legal reasoning tasks\. Critically, significant model\-method interactions emerge, with Gemini demonstrating exceptional performance when paired with NeOn methods, particularly NeOn\-CoT, suggesting optimal ontology learning requires careful model\-method matching\. Figure 2:Performance loss across the subset of MAUD tasks with baseline accuracy≥50%\\geq 50\\%; tasks below this threshold \(e\.g\., t11\) are excluded from the heatmap but remain part of the aggregate results in Table[1](https://arxiv.org/html/2608.12326#S4.T1)\. The full task list is given in Appendix[A](https://arxiv.org/html/2608.12326#A1)\.The most severe semantic loss occurs with nuanced legal standards and modal language defining procedural requirements\. Taskt13\(fiduciary duty standards for superior offer recommendations\) shows the highest average loss at 65%, followed byt12\(fiduciary duty standards for intervening event recommendations\) at 53%\. This illustrates a challenge: while source documents specify that boards may act when continuation would more likely than not result in a violation of fiduciary duties, ontological representations abstract this to generic concepts likeLikelyViolationOfFiduciaryDutiesLikelyViolationOfFiduciaryDuties, losing precise probabilistic language \(more likely than not vs\. reasonably likely vs\. could reasonably be expected\) that distinguishes legal standards\. However, model\-method interactions reveal striking differences\. Whilet13shows devastating loss with most combinations, Gemini paired with NeOn\-CoT achieves 88% semantic retention\. Similarly,t12shows 53% average loss, but Gemini with NeOn\-CoT maintains 86% baseline performance compared to just 13% with LLMs4OL, illustrating how modal expressions respond dramatically differently to various ontological approaches\. Conversely, tasks involving straightforward binary determinations demonstrate near\-perfect preservation across all combinations\. Taskt19\(whether financial point of view is sole consideration for superior offers\) achieved best preservation with only 1% average loss, followed byt34\(consideration type classification, 6% loss\) andt28\(efforts standard determination, 6% loss\)\. These leverage ontologies’ strength in capturing explicit relationships while avoiding linguistic nuance causing degradation in complex reasoning tasks\. The pronounced model\-method interaction effects reveal that effectiveness depends critically on appropriate pairing\. Tasks involving complex multi\-standard determinations show 45\-65% average losses, but Gemini consistently outperforms other models with NeOn methods, winning 9 out of 10 challenging tasks analyzed\. This suggests Gemini’s architecture is particularly suited to NeOn’s structured reasoning patterns, potentially due to superior handling of chain\-of\-thought and step\-by\-step ontological reasoning\. Method\-specific analysis reveals the sophistication hierarchy \(NeOn\-GPT\>\>NeOn\-CoT\>\>LLMs4OL\) applies primarily to certain model types, with Gemini representing a clear outlier\. While NeOn\-GPT generally achieves best average preservation across all models, Gemini specifically excels with NeOn\-CoT, suggesting synergistic effects for legal ontology learning\. This indicates future research should focus on identifying optimal model\-method pairings rather than pursuing universal approaches, with Gemini \+ NeOn\-CoT representing a particularly promising direction for maintaining semantic fidelity in legal ontology transformations\. ## 5Discussion The consistency of semantic loss across diverse models and methods suggests this may reflect characteristics of current transformation approaches rather than implementation\-specific artifacts\. These findings suggest the need for more evidence\-based approaches to ontological transformation, moving from assumptions about structural benefits to empirical validation of when ontologies are appropriate for specific applications\. Semantic loss appears intrinsic to the abstraction process \- when ontological representations transform “more likely than not” into generic concepts likeLikelyViolationOfFiduciaryDutiesLikelyViolationOfFiduciaryDuties, critical legal distinctions vanish irreversibly\. This is not a technical problem to be solved through better algorithms, but a fundamental characteristic of structural abstraction that practitioners must acknowledge\. The discovery that optimal performance requires specific model\-method pairings \(e\.g\., Gemini \+ NeOn\-CoT achieving 86\-88% semantic retention while other combinations preserve only 12\-45%\) suggests ontology learning is not universally applicable\. Rather than pursuing one\-size\-fits\-all approaches, the field should focus on identifying optimal configurations for specific legal reasoning tasks and content types\. Our LLM\-based evaluation methodology could represent a useful complement to existing ontology assessment approaches\. While structural correctness metrics ensure technical validity, they may not fully capture the semantic richness essential for downstream applications\. Semantic preservation metrics could supplement structural measures, enabling more evidence\-based decisions about when ontological transformation enhances versus diminishes practical utility\. Despite systematic semantic loss during transformation, ontologies remain important in the LLM era for bridging neural and symbolic systems\. Natural language functions poorly as a communication protocol between computational systems due to inherent ambiguity and inconsistency\[[20](https://arxiv.org/html/2608.12326#bib.bib22)\]\. Ontologies provide controlled vocabularies essential for data interoperability, transparent knowledge representations, and logical consistency requirements that LLMs cannot guarantee\[[20](https://arxiv.org/html/2608.12326#bib.bib22)\]\. This complementary relationship suggests ontologies’ role as intermediaries in hybrid architectures, enabling LLMs to interface with symbolic reasoning systems, databases, and automated decision\-making processes that require formal guarantees\. ## 6Limitations and Future Work Several limitations should temper the interpretation of these results\. The evaluation covers a single legal subdomain, merger and acquisition contract analysis through MAUD, so the magnitude of semantic loss and the observed NeOn\-GPT\>\>NeOn\-CoT\>\>LLMs4OL hierarchy may not transfer to other legal subdomains or to non\-legal domains\. The task format is multiple\-choice, which is convenient for automated scoring but bounds the semantic distinctions the framework can surface; open\-ended legal tasks may exhibit different patterns\. Coverage of ontology learning approaches is limited to three methods\. The framework itself relies on baseline LLM performance as a proxy for accessible content, so semantic loss is measured relative to each model’s capabilities rather than against ground\-truth semantic content; results should be read as comparative rather than absolute\. Ontologies are consumed by LLMs as text, which means the framework does not credit benefits that would materialize only through formal reasoning such as SPARQL queries or automated inference\. We also assume information conservation, that is, ontology generation introduces no external knowledge; approaches that intentionally enrich ontologies fall outside the framework as defined here\. Given the systematic semantic loss observed across all transformation approaches, future research should investigate how ontological representations might better preserve complex legal standards\. Tasks like fiduciary duty determinations showed the highest degradation, suggesting that current methods struggle with the precise formulations that distinguish legal standards in practice\. Domain generalization studies should test whether observed patterns hold across legal specializations and other knowledge domains where precise language matters\. Computational context evaluation should also assess whether semantic loss is offset by formal reasoning capabilities through SPARQL queries and automated inference within ontologies’ intended computational environment\[[2](https://arxiv.org/html/2608.12326#bib.bib2)\]\. Selective transformation approaches could optimize the accessibility\-structure trade\-off by applying ontological methods only to content types that benefit from formal representation while preserving linguistically nuanced passages in natural language\. Finally, systematic model\-method optimization should explore architecture\-methodology interactions to identify optimal pairings for specific legal reasoning requirements\[[8](https://arxiv.org/html/2608.12326#bib.bib9)\]\. ## 7Summary Ontology learning transforms complex documents into structured representations for automated reasoning\. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on structural correctness while failing to measure whether meaning survives transformation\. We propose an evaluation methodology that addresses this gap: comparing LLM task performance on source documents against performance on transformed representations, with the difference quantifying semantic loss\. This approach treats baseline performance as accessible content and enables evidence\-based assessment of when ontological transformation helps or hinders downstream applications\. Evaluating this methodology on legal merger agreement analysis using the MAUD dataset, we compare direct LLM application against three ontology learning approaches \(LLMs4OL, NeOn\-GPT, and NeOn\-CoT\) across six state\-of\-the\-art language models\. Our results reveal systematic semantic loss across all approaches, with severe losses for complex legal reasoning tasks involving nuanced modal language, while simple categorical determinations show near\-perfect preservation\. We discover significant model\-method interactions, with optimal pairings \(e\.g\., Gemini \+ NeOn\-CoT\) achieving strong semantic retention while suboptimal combinations showed severe degradation\. Our contributions are: \(1\) an evaluation framework for measuring semantic preservation in ontology learning, and \(2\) empirical evidence that semantic loss varies dramatically with model\-method pairing, providing guidance for selecting optimal configurations in legal knowledge systems\. These findings suggest that practitioners should empirically validate semantic preservation rather than assuming structural correctness guarantees downstream utility\. ## References - \[1\]G\. Babaeiet al\.\(2023\)LLMs4OL: large language models for ontology learning\.InThe Semantic Web – ISWC 2023,pp\. 408–427\.External Links:ISBN 978\-3\-031\-47240\-4Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p5.1),[§2](https://arxiv.org/html/2608.12326#S2.p1.1),[§2](https://arxiv.org/html/2608.12326#S2.p2.1),[§3\.1](https://arxiv.org/html/2608.12326#S3.SS1.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2608.12326#S3.p1.1)\. - \[2\]J\. Branket al\.\(2005\)A survey of ontology evaluation techniques\.InProceedings of the Conference on Data Mining and Data Warehouses \(SiKDD 2005\),Ljubljana, Slovenia,pp\. 166–170\.Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p3.1),[§2](https://arxiv.org/html/2608.12326#S2.p1.1),[§3](https://arxiv.org/html/2608.12326#S3.p1.1),[§6](https://arxiv.org/html/2608.12326#S6.p4.1)\. - \[3\]C\. Brewsteret al\.\(2004\)Data driven ontology evaluation\.InProceedings of the Fourth International Conference on Language Resources and Evaluation \(LREC’04\),Lisbon, Portugal\.Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p3.1),[§2](https://arxiv.org/html/2608.12326#S2.p3.1)\. - \[4\]P\. Buitelaaret al\.\(2005\)Ontology learning from text: an overview\.pp\. 3–12\.Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p1.1)\. - \[5\]N\. Casellas\(2011\)Legal ontology engineering: methodologies, modelling trends, and the ontology of professional judicial knowledge\.Law, Governance and Technology Series,Springer,Netherlands\.Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p1.1),[§2](https://arxiv.org/html/2608.12326#S2.p1.1)\. - \[6\]J\. A\. Chudziak and M\. Piotrowski\(1995\)Semantic support for multimedia information system\.In1995 IEEE International Conference on Systems, Man and Cybernetics\. Intelligent Systems for the 21st Century,Vol\.5,pp\. 3914–3919 vol\.5\.External Links:[Document](https://dx.doi.org/10.1109/ICSMC.1995.538400)Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p1.1)\. - \[7\]P\. Cimiano and J\. Völker\(2005\)Text2Onto\.InNatural Language Processing and Information Systems,pp\. 227–238\.External Links:ISBN 978\-3\-540\-32110\-1Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p1.1)\. - \[8\]B\. C\. Colelough and W\. Regli\(2025\)Neuro\-symbolic ai in 2024: a systematic review\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2501.05435)Cited by:[§6](https://arxiv.org/html/2608.12326#S6.p5.1)\. - \[9\]O\. Corchoet al\.\(2005\)Building legal ontologies with methontology and webode\.InLaw and the Semantic Web: Legal Ontologies, Methodologies, Legal Information Retrieval, and Applications,pp\. 142–157\.External Links:ISBN 978\-3\-540\-32253\-5,[Document](https://dx.doi.org/10.1007/978-3-540-32253-5%5F9)Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p1.1)\. - \[10\]C\. M\. de Oliveira Rodrigueset al\.\(2019\)Legal ontologies over time: a systematic mapping study\.Expert Systems with Applications130,pp\. 12–30\.External Links:[Document](https://dx.doi.org/10.1016/j.eswa.2019.04.009)Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p1.1),[§2](https://arxiv.org/html/2608.12326#S2.p1.1),[§2](https://arxiv.org/html/2608.12326#S2.p2.1)\. - \[11\]N\. Fathallahet al\.\(2025\)NeOn\-gpt: a large language model\-powered pipeline for ontology learning\.InThe Semantic Web: ESWC 2024 Satellite Events,pp\. 36–50\.External Links:ISBN 978\-3\-031\-78952\-6Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p5.1),[§2](https://arxiv.org/html/2608.12326#S2.p1.1),[§2](https://arxiv.org/html/2608.12326#S2.p2.1),[§3\.1](https://arxiv.org/html/2608.12326#S3.SS1.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2608.12326#S3.p1.1)\. - \[12\]A\. Gangemiet al\.\(2005\)A theoretical framework for ontology evaluation and validation\.InCEUR Workshop Proceedings,Vol\.166\.Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p3.1)\. - \[13\]N\. Guhaet al\.\(2023\)LEGALBENCH: a collaboratively built benchmark for measuring legal reasoning in large language models\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2308.11462)Cited by:[Table A\.1](https://arxiv.org/html/2608.12326#A1.T1.2),[Appendix A](https://arxiv.org/html/2608.12326#A1.p1.1),[§1](https://arxiv.org/html/2608.12326#S1.p5.1),[§3\.3](https://arxiv.org/html/2608.12326#S3.SS3.p1.1)\. - \[14\]M\. A\. Hearst\(1998\)Automated discovery of wordnet relations\.WordNet: An Electronic Lexical Database and Some of its Applications2\.Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p1.1)\. - \[15\]X\. Jiaoet al\.\(2020\)TinyBERT: distilling BERT for natural language understanding\.InFindings of the Association for Computational Linguistics: EMNLP 2020,pp\. 4163–4174\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.372)Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p4.1)\. - \[16\]K\. S\. Jones and J\. R\. Galliers\(1996\)Evaluating natural language processing systems: an analysis and review\.Springer\-Verlag\.External Links:ISBN 3540613099Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p4.1),[§2](https://arxiv.org/html/2608.12326#S2.p3.1),[§3](https://arxiv.org/html/2608.12326#S3.p2.1)\. - \[17\]A\. Kostka and J\. A\. Chudziak\(2024\)Synergizing logical reasoning, knowledge management and collaboration in multi\-agent LLM system\.InProceedings of the 38th Pacific Asia Conference on Language, Information and Computation,Tokyo, Japan,pp\. 203–212\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2507.02170)Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p2.1)\. - \[18\]A\. Maedche\(2002\)Ontology learning for the semantic web\.The Springer International Series in Engineering and Computer Science, Vol\.665,Springer Science\+Business Media,New York, NY\.External Links:[Document](https://dx.doi.org/10.1007/978-1-4615-0925-7)Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p1.1)\. - \[19\]D\. Mollá and B\. Hutchinson\(2003\)Intrinsic versus extrinsic evaluations of parsing systems\.InProceedings of the EACL 2003 Workshop on Evaluation Initiatives in Natural Language Processing: Are Evaluation Methods, Metrics and Resources Reusable?,USA,pp\. 43–50\.External Links:[Document](https://dx.doi.org/10.3115/1641396.1641403)Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p3.1),[§3](https://arxiv.org/html/2608.12326#S3.p2.1)\. - \[20\]F\. Neuhaus\(2023\)Ontologies in the era of large language models – a perspective\.Applied Ontology18,pp\. 399–407\.External Links:[Document](https://dx.doi.org/10.3233/AO-230072)Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p2.1),[§5](https://arxiv.org/html/2608.12326#S5.p5.1)\. - \[21\]N\. Noy and D\. Mcguinness\(2001\)Ontology development 101: a guide to creating your first ontology\.Knowledge Systems Laboratory32\.Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p1.1)\. - \[22\]R\. Porzel and R\. Malaka\(2004\)A task\-based approach for ontology evaluation\.ECAI Workshop on Ontology Learning and Population\.Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p4.1),[§2](https://arxiv.org/html/2608.12326#S2.p3.1),[§3](https://arxiv.org/html/2608.12326#S3.p1.1)\. - \[23\]M\. T\. Ribeiroet al\.\(2020\)Beyond accuracy: behavioral testing of NLP models with CheckList\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,pp\. 4902–4912\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.442)Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p4.1),[§2](https://arxiv.org/html/2608.12326#S2.p3.1),[§3](https://arxiv.org/html/2608.12326#S3.p1.1)\. - \[24\]A\. Sadowski and J\. A\. Chudziak\(2025\)On verifiable legal reasoning: a multi\-agent framework with formalized knowledge representations\.InProceedings of the 34th ACM International Conference on Information and Knowledge Management,New York, NY, USA,pp\. 2535–2545\.External Links:ISBN 9798400720406,[Document](https://dx.doi.org/10.1145/3746252.3761057)Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p2.1)\. - \[25\]M\. C\. Suárez\-Figueroaet al\.\(2015\)The neon methodology framework: a scenario\-based methodology for ontology development\.Applied Ontology10\(2\),pp\. 107–145\.External Links:[Document](https://dx.doi.org/10.3233/AO-150145)Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p1.1)\. - \[26\]A\. Valenteet al\.\(1999\)Legal modeling and automated reasoning with on\-line\.International Journal of Human\-Computer Studies51\(6\),pp\. 1079–1125\.External Links:[Document](https://dx.doi.org/10.1006/ijhc.1999.0298)Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p1.1)\. - \[27\]P\. Vergaet al\.\(2024\)Replacing judges with juries: evaluating llm generations with a panel of diverse models\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2404.18796)Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p4.1)\. - \[28\]P\. Wanget al\.\(2024\)Large language models are not fair evaluators\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 9440–9450\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.511)Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p4.1)\. - \[29\]S\. Wanget al\.\(2023\)MAUD: an expert\-annotated legal NLP dataset for merger agreement understanding\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 16369–16382\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.1019)Cited by:[§1](https://arxiv.org/html/2608.12326#S1.p5.1),[§3\.3](https://arxiv.org/html/2608.12326#S3.SS3.p1.1)\. - \[30\]C\. Xuet al\.\(2021\)Beyond preserved accuracy: evaluating loyalty and robustness of BERT compression\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 10653–10659\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.832)Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p4.1)\. - \[31\]S\. Zhaoet al\.\(2019\)Region mutual information loss for semantic segmentation\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.1910.12037)Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p4.1)\. - \[32\]L\. Zhenget al\.\(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.InProceedings of the 37th International Conference on Neural Information Processing Systems,Red Hook, NY, USA\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2306.05685)Cited by:[§2](https://arxiv.org/html/2608.12326#S2.p4.1)\. ## Appendix AMAUD Tasks Reference This appendix provides the mapping between the task identifiers used throughout this paper and the corresponding questions from the MAUD task as implemented in LegalBench\[[13](https://arxiv.org/html/2608.12326#bib.bib13)\]\. Table[A\.1](https://arxiv.org/html/2608.12326#A1.T1)lists all 34 tasks with their full text\. Table A\.1:MAUD Task ReferenceNote:All MAUD tasks originate from LegalBanch\[[13](https://arxiv.org/html/2608.12326#bib.bib13)\]\. Complete task definitions, including answer options and examples, are available in the[LegalBench repository](https://github.com/HazyResearch/legalbench)\.
Similar Articles
Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs
This paper proposes a semantic verification framework using Natural Language Inference (NLI) to evaluate the sensitivity of clinical LLMs to meaning-preserving prompt variations, introducing metrics such as MVS, ΔC, and WCI. Results show that domain specialization does not consistently improve robustness, with both domain-specific and general-purpose models showing mixed performance.
LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation
This paper introduces LP-Eval, a rubric and dataset for evaluating legal proposition generation by large language models, with annotations by legal experts. Results show that rubric-guided LLM evaluations align more closely with expert assessments than direct scoring.
LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs
PropMe is a propensity-aware framework for evaluating LLM memorization, distinguishing between forced reproduction capabilities and natural propensity using SimpleTrace for deterministic attribution across open models and datasets.
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
A comprehensive dual-aspect evaluation framework for large language models on Vietnamese legal text simplification, combining quantitative benchmarking (Accuracy, Readability, Consistency) with qualitative error analysis across GPT-4o, Claude 3 Opus, Gemini 1.5 Pro, and Grok-1.
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
Researchers from PNNL and Washington University introduce a systematic framework to test how five LLMs detect subtle semantic changes in documents, revealing positional bias, context coherence effects, and model-specific scoring fingerprints.