Semantic Layer Induction from Raw Telemetry via Hierarchical LLM and RAG Abstraction
Summary
This paper presents an end-to-end framework for automating the construction of a business semantic layer from raw telemetry data using hierarchical LLM inference and RAG abstraction, significantly improving semantic quality and reducing maintenance effort.
View Cached Full Text
Cached at: 09/18/26, 08:59 AM
# Semantic Layer Induction from Raw Telemetry via Hierarchical LLM and RAG Abstraction
Source: [https://arxiv.org/html/2609.19615](https://arxiv.org/html/2609.19615)
###### Abstract
Modern applications generate massive volumes of raw telemetry data, but translating those noisy, heterogeneous event streams into actionable business insights remains a fundamental challenge\. Data engineers and analysts expend substantial effort reconciling semantic discrepancies, hand\-crafting parsing logics, and maintaining fragile mappings between raw data and business KPIs\. In this paper, we present an end\-to\-end framework that fully automates the construction of a business semantic layer from application raw logs\. Our approach introduces a two\-stage semantic abstraction: first, high\-level business features are identified via LLM inference augmented with domain\-specific industry knowledge; second, fine\-grained business nodes are derived through a structured pipeline comprising data refinement, hybrid retrieval, multi\-stage filtering, semantic clustering, and canonical naming\. Evaluation on production\-scale telemetry demonstrates that our system improves human\-assessed semantic quality from 50 to 80\+ on a 100\-point scale, reduces maintenance effort by 80%, filters out 74% of noise, and achieves 0\.87 Cohen’s kappa via an integrated LLM\-as\-Judge evaluation, enabling continuous, scalable quality assurance\. Overall, our work distinguishes itself from prior work by addressing the novel problem of business semantic layer induction from raw telemetry, operating without labeled training data or manual rule engineering\. The relevant code is publicly available on GitHub111[https://github\.com/yuanzhe\-jia/semantic\-layer](https://github.com/yuanzhe-jia/semantic-layer)\.
###### Keywords:
Semantic Layer, Hierarchical Abstraction, Retrieval Augmented Generation, LLM\-as\-Judge
## 1Introduction
Modern software platforms generate petabytes of raw telemetry data every day—user interactions, backend events, and API calls—capturing the full spectrum of system activity\. It is the lifeblood of data\-driven decision\-making, powering conversion funnels, product analytics, and KPI monitoring\. However, organizations consistently struggle to extract reliable business insights from this kind of data\. The challenge lies not only in data volume but also in semantic heterogeneity\. The same action—say, "users search for a product"—may be tracked across platforms and versions assearch\_click\(Android\),search\_submit\(IOS\) andbutton\_click\(Web\), each with different parameters and schema\. This fragmentation forces teams into an endless cycle of manual mapping, custom SQL logic per dashboard, and cross\-functional debate about "what the telemetry data actually means"\. The operational cost is substantial: a typical enterprise data platform may maintain thousands of custom parsers, consuming thousands of engineer\-hours monthly\.
Existing solutions fall short\. Manual parser construction, while precise, is brittle and does not scale across heterogeneous log formats\. Syntax\-based parsing methods can extract templates but lack business semantic understanding\. Supervised learning approaches require extensive labeled data for each tracking point, making them impractical for rapidly evolving applications\. Even recent LLM\-based approaches to log parsing focus on syntax\-level template extraction rather than semantic mapping to business concepts\. We argue that what enterprises require is not merely structured logs, but a business semantic layer—a stable, canonical mapping from raw telemetry to human\-interpretable insights that directly align with customer needs and business objectives\. In this paper, we present an LLM\-powered framework for business semantic layer induction\. The rest of the paper is structured as follows: Section[2](https://arxiv.org/html/2609.19615#S2)reviews and critiques existing approaches in log parsing, semantic data management, and RAG\-based structured extraction; Section[3](https://arxiv.org/html/2609.19615#S3)details our hierarchical abstraction framework and its algorithmic implementation; Section[4](https://arxiv.org/html/2609.19615#S4)describes the dataset, experimental setup, evaluation metrics, and empirical results; and Section[5](https://arxiv.org/html/2609.19615#S5)concludes the paper with a summary of contributions and directions for future work\.
## 2Related Work
### 2\.1Log Parsing and Telemetry Analysis
Automated log parsing has been extensively studied in the systems community\. Traditional approaches use clustering or frequent pattern mining to extract log templates\. Methods like Drain\[[3](https://arxiv.org/html/2609.19615#bib.bib1)\]and LogParser\[[11](https://arxiv.org/html/2609.19615#bib.bib2)\]achieve high template extraction accuracy but focus on syntax\-level patterns—they identify what changes across log lines but do not understand what those changes semantically represent\. More recent LLM\-based parsers, such as LogParser\-LLM\[[13](https://arxiv.org/html/2609.19615#bib.bib3)\], demonstrate superior performance by seamlessly blending semantic insights with statistical nuances, obviating the need for hyper\-parameter tuning and labeled training data while ensuring rapid adaptability through online parsing\. However, these approaches still focus on template extraction and field naming rather than mapping logs to semantic concepts\. Similarly, Matryoshka et al\.\[[7](https://arxiv.org/html/2609.19615#bib.bib4)\]use LLMs to generate semantically\-aware log parsers by inferring log syntax, variable naming, and schema normalization\. Although impressive, they focus on mapping log fields to standardized security schema for threat detection\. In process mining, researchers have explored semantics\-aware event log analysis using LLMs\[[8](https://arxiv.org/html/2609.19615#bib.bib5)\], but these approaches typically analyze existing logs rather than constructing them from raw telemetry\.
### 2\.2Semantic Data Management
Beyond log parsing and telemetry analysis, the data management community has long investigated semantic enrichment of enterprise data assets\. Hoseini et al\.\[[4](https://arxiv.org/html/2609.19615#bib.bib6)\]provide a comprehensive survey on semantic data management in data lakes, covering ontology\-based data access and semantic modeling approaches that link metadata to knowledge graphs\. In parallel, significant research has focused on knowledge graph construction as a means of structuring and organizing business semantics\. Bian et al\.\[[1](https://arxiv.org/html/2609.19615#bib.bib7)\]survey LLM\-empowered knowledge graph construction, analyzing how LLMs reshape ontology engineering and knowledge extraction pipelines, while Zhao et al\.\[[12](https://arxiv.org/html/2609.19615#bib.bib8)\]review machine learning approaches for entity and ontology learning\. From a metadata management perspective, recent surveys on data catalog tools by Kropshofer et al\.\[[5](https://arxiv.org/html/2609.19615#bib.bib9)\]and Tonnarelli et al\.\[[9](https://arxiv.org/html/2609.19615#bib.bib10)\]examine how technical metadata annotated with domain knowledge improves data accessibility and interoperability\. However, these approaches rely on pre\-defined ontologies and struggle with the heterogeneity of multi\-platform naming convention, requiring additional translation to aggregate fine\-grained graph relations into business KPIs\. None address the unique challenge of inducing hierarchical business semantics directly from noisy, high\-volume telemetry\.
### 2\.3RAG for Structured Data Extraction
RAG has been increasingly applied to structured knowledge extraction and domain\-specific reasoning tasks\. EventRAG\[[10](https://arxiv.org/html/2609.19615#bib.bib11)\]introduces an event\-centric RAG framework that constructs event knowledge graphs from narrative documents to enhance LLM generation with structured event semantics and temporal reasoning\. TM\-RAG\[[14](https://arxiv.org/html/2609.19615#bib.bib12)\]employs ontology\-guided graph retrieval with a timeline ontology for automated construction claim report generation, demonstrating the effectiveness of structured knowledge organization in domain\-specific RAG systems\. GenDFIR\[[6](https://arxiv.org/html/2609.19615#bib.bib13)\]applies RAG to cyber incident timeline analysis, retrieving relevant forensic events from a structured knowledge base to support investigation\. While these approaches share insights that event\-centric organization and structured retrieval improve reasoning, they are designed for task\-specific, one\-off query answering\. In contrast, our method uses RAG to retrieve candidate mapping rules for each business capability with the distinct objective of constructing a reusable, generalizable semantic layer that can serve diverse downstream analytics without task\-specific re\-engineering\.
## 3Methodology
### 3\.1Problem Formulation
We formalize the semantic layer induction problem as follows:
- •Given: A raw telemetry event corpusℰ=\{e1,…,eN\}\\mathcal\{E\}=\\\{e\_\{1\},\\dots,e\_\{N\}\\\}, where eacheie\_\{i\}is a tracking event that consists of a name and a set of key\-value conditions; and an optional industry taxonomy𝒯\\mathcal\{T\}\.
- •Find: A semantic layer𝒮=\{\(fj,𝒩j\)\}\\mathcal\{S\}=\\\{\(f\_\{j\},\\mathcal\{N\}\_\{j\}\)\\\}, wherefjf\_\{j\}is a business feature, and𝒩j=\{nj,1,…,nj,K\}\\mathcal\{N\}\_\{j\}=\\\{n\_\{j,1\},\\dots,n\_\{j,K\}\\\}are business nodes, such that𝒮\\mathcal\{S\}maximizes a semantic coherence objective while minimizing feature sparsity\.
Algorithm 1The proposed framework0:raw telemetry
ℰ\\mathcal\{E\}, industry prior
𝒯\\mathcal\{T\}, top\-
kkthreshold
0:semantic layer
𝒮\\mathcal\{S\}
1:
ℰrefined←RefineEvents\(ℰ,𝒯\)\\mathcal\{E\}\_\{refined\}\\leftarrow\\textsc\{RefineEvents\}\(\\mathcal\{E\},\\mathcal\{T\}\)\{Stage 1\}
2:
ℱ←IdentifyFeatures\(ℰrefined,𝒯\)\\mathcal\{F\}\\leftarrow\\textsc\{IdentifyFeatures\}\(\\mathcal\{E\}\_\{refined\},\\mathcal\{T\}\)\{Stage 2\}
3:foreach feature
f∈ℱf\\in\\mathcal\{F\}do
4:
ℛf←RuleRetrieve\(f,ℰrefined\)\\mathcal\{R\}\_\{f\}\\leftarrow\\textsc\{RuleRetrieve\}\(f,\\mathcal\{E\}\_\{refined\}\)\{Stage 3\}
5:
ℛfclean←FilterCandidateRules\(ℛf\)\\mathcal\{R\}\_\{f\}^\{clean\}\\leftarrow\\textsc\{FilterCandidateRules\}\(\\mathcal\{R\}\_\{f\}\)\{Stage 4\}
6:
𝒩f←ClusterAndNameNodes\(ℛfclean\)\\mathcal\{N\}\_\{f\}\\leftarrow\\textsc\{ClusterAndNameNodes\}\(\\mathcal\{R\}\_\{f\}^\{clean\}\)\{Stage 5\}
7:endfor
8:
𝒮←⋃f\(f,𝒩f\)\\mathcal\{S\}\\leftarrow\\bigcup\_\{f\}\(f,\\mathcal\{N\}\_\{f\}\)
9:return
𝒮\\mathcal\{S\}
A straightforward approach to this problem is to learn a flat mapping directly from raw telemetry data to business labels\. However, such a strategy suffers from several fundamental limitations: the mapping space is large; semantically similar events may be scattered across unrelated labels; and the resulting mappings are brittle to platform\-specific naming conventions\. Our proposed model solves this problem via two\-level hierarchy \(see Algorithm 1\)\. This hierarchy \(Feature → Node\) serves as an inductive bias that constrains the search space: the framework first reasons about high\-level business capabilities, and then refines each capability into its constituent actions\. By decoupling the problem into two nested subproblems, the hierarchy reduces the effective complexity of the mapping task and enables the system to leverage industry priors at the feature level before committing to fine\-grained node assignments\. Additionally, unlike a naive pipeline, the model framework maintains a shared latent semantic space across stages: representations learned in data refinement directly influence feature identification, which in turn constrains the retrieval space via a feedback loop, ensuring that downstream errors do not cascade catastrophically\.
### 3\.2Stage 1: Data Preparation
Raw telemetry data often includes high\-cardinality fields and semantically weak columns that hinder effective retrieval\. To address this, we perform a two\-step data refinement process\.
#### Column Importance Scoring:
We first identify columns that carry significant business context from application telemetry, which is typically stored as tracking event logs, where each event contains a name and numerous metadata columns \(e\.g\.,url\_path,page\_title,element\_id\)\. For each metadata column, we prompt an LLM with statistical profiles \(e\.g\., null ratio, distinct count\) and a predefined business taxonomy to assign an importance score ranging from 1 to 5\. Columns falling below a calibrated threshold are excluded from subsequent processing, effectively pruning irrelevant attributes that would otherwise introduce noise into semantic matching\. This step constitutes a lightweight but effective schema\-level filter that prioritizes business\-relevant dimensions over purely technical or transient fields\.
#### Enumeration Normalization:
For each column retained after the importance scoring, we tokenize its enumeration values using common delimiters \(e\.g\., "/", whitespace\) to obtain fine\-grained tokens\. We then apply a random string detector—a entropy\-based classifier augmented with LLM\-based pattern recognition—to identify transient identifiers such as UUIDs, session tokens, and request IDs\. These detected strings are replaced with a uniform mask symbol \("\*"\), effectively stripping away instance\-specific noise\. Subsequently, we consolidate enumerations that share identical masked patterns via a set of regular expression rules, merging semantically equivalent variants into a canonical representation\. The entire process yields a clean, low\-cardinality vector space\. Finally, we aggregate identical cleaned records to produce a compact set of raw telemetry data, and the refined data will serve as the input for subsequent feature identification and mapping retrieval stages\.
### 3\.3Stage 2: Business Feature Identification
We leverage an LLM to generate a set of high\-level business features from the refined data\. Given that the data volume after Stage 1 remains prohibitively large for direct LLM consumption, we first perform a stratified sampling to obtain a representative subset that preserves the diversity of data patterns and their frequency distribution\. The sampling strategy ensures the LLM operates within its context window while maintaining sufficient coverage of the application’s behavioral landscape\. To ensure the generated features are both comprehensive and industry\-relevant, we augment the LLM context with domain\-specific reference materials\. Specifically, for a given vertical \(e\.g\., e\-commerce\), we supply the LLM with a curated list of canonical business features that are typical for that industry, such asSignin,Search,Cart,Checkout, andOrder\. This external knowledge acts as a strong prior, anchoring the generation to established business taxonomies and preventing the LLM from producing overly granular, UI\-focused, or platform\-specific labels\. This dual\-input strategy—combining sampled data with external industry knowledge—enables the system to produce features that are both empirically grounded in the observed telemetry and semantically aligned with real\-world logic\.
#### Constraints:
The LLM is prompted to produce feature names that:
- •Use one or two nouns\.
- •Prefer general names \(e\.g\.,Searchrather thanKeyword Search\)\.
- •Avoid UI\-specific nouns \(e\.g\.,Button,Form\)\.
- •Reuse industry\-standard feature names when semantically matching\.
#### Quality Check:
After generation, we apply a gate that verifies:
- •All supplied industry\-standard features are covered\.
- •Each feature name contains fewer than three words\.
- •The semantic similarity among features remains below a threshold\.
### 3\.4Stage 3: Candidate Rule Retrieval
We retrieve candidate SQL\-like mapping rules for each business feature identified in Stage 2\. We first encode the refined data \(produced in Stage 1\) into dense embeddings using a pre\-trained sentence transformer\. For a given business feature, we generate a query embedding from its name \(optionally augmented with a brief description\) and perform a similarity search over the vector index\. Critically, each indexed unit represents not an entire tracking event, but a specific key\-value condition\. The retrieval mechanism is designed to identify the top\-kkmost semantically aligned condition subsets with respect to the target business feature\. This design ensures that the search focuses on attribute\-level semantics, rather than being biased by surface\-level naming conventions\. To improve recall, we complement dense retrieval with BM25 keyword matching and combine both scores via a weighted linear fusion\.
For each business feature, the retrieval process returns a diverse set of tracking event conditions that are semantically related but span different facets of the feature\. For instance, for the business featureSearch, the retrieved event conditions may include abutton\_clickevent indicating the initiation of a search action \(e\.g\.,element\_id = "search\_init"\) as well as apage\_viewevent reflecting the subsequent display of search results \(e\.g\.,page\_title LIKE "%Search%"\)\. Each retrieved event condition is translated into a SQL\-like predicate based on its structural conditions\. The union of these predicates for a given business feature forms an initial candidate rule set, which encapsulates multiple semantic variants underlying the same business feature\. This broad coverage ensures semantic comprehensiveness, while the subsequent filtering and clustering stages are responsible for disentangling these variants and assigning them to distinct business nodes\.
### 3\.5Stage 4: Candidate Rule Filtering
Retrieved rules contain significant noise, thus we apply a two\-phase filter:
#### Hard Rule Filtering
: Blocked events \(e\.g\.,ad\_click,api\_error\)\.
#### LLM Semantic Filtering
: For each candidate rule, an LLM independently judges whether it belongs to the corresponding business feature\. The LLM returnsYes\(retain\) orNo\(reject\)\. Rules are rejected if they are:
- •Not semantically relevant to the business feature\.
- •Logs that do not reflect user interactions\.
- •Pure input without results\.
- •No meaningful API calls\.
### 3\.6Stage 5: Business Node Clustering and Naming
Rules retained after filtering often correspond to multiple distinct user actions or system processes falling under the same business feature\. To derive semantically coherent business nodes, we first perform a clustering step over the retained rules for each business feature\. Specifically, we use the LLM to group rules that share the same underlying business semantics and represent exactly the same stage of a business module \(e\.g\., for theSearchfeature, rules indicating the initiation of a search are clustered together, while those indicating the viewing of results form a separate cluster\)\. This ensures that all mapping rules within a given cluster are semantically equivalent\.
Subsequently, for each cluster, the LLM assigns a canonical name that concisely captures the common semantics of its member rules\. The naming format follows a structured pattern:\[noun\]\+\[verb\], where the noun part consists of one or two nouns that denote the parent business node \(e\.g\.,Search Result,Cart Item\), and the verb part is a single base\-form word that describes the precise user action or state transition \(e\.g\.,View,Add\)\. Additional constraints enforce name consistency: multiple rules reflecting the same behavior must share an identical name, and when semantically similar candidates appear, the most general name is preferred\. This hierarchical clustering\-then\-naming strategy ensures that each business node corresponds to a unique business action/status, maintaining a clean, interpretable, and platform\-agnostic semantic layer\.
### 3\.7Deterministic Output
To ensure reproducibility and avoid hallucination, we:
- •Use fixed random seeds for LLM inference\.
- •Settemperature = 0\.0for deterministic sampling\.
- •Constrain the output format to valid JSON objects with concrete examples\.
- •Parse LLM responses withjson\.loads\(\)\.
The final output is a JSON object keyed bynode\_id, each containing a canonical business node name, matching conditions, and relevant metadata for downstream usage\.
```
{
"n0": {
"node_name": "Search Result View",
"node_rule": [
{
"event_name": "page_view",
"conditions": [
{
"key": "page_title",
"value": "%search%",
"operator": "LIKE"
}
]
}
]
"feature_name": "Search",
"industry": "e-commerce"
}
}
```
## 4Experiments
### 4\.1Experimental Setup
#### Dataset:
We evaluated our system on a large\-scale e\-commerce dataset\[[2](https://arxiv.org/html/2609.19615#bib.bib14)\]derived from real\-world user interaction logs collected from an online retailer’s website over a six\-month period\. With millions of user sessions spanning the full e\-commerce user journey—from browsing and searching to carting and purchasing—this dataset directly mirrors the heterogeneous, multi\-platform telemetry scenarios targeted by our business semantic layer induction framework\.
#### Evaluation Metrics:
- •Human Assessment: Semantic correctness rated by human experts\.
- •Noise Reduction: Percentage of data filtered out by the pipeline\.
- •Manual Effort Reduction: Hours per week saved by human experts previously maintaining custom semantic mappings\.
- •LLM\-as\-Judge Agreement: Cohen’s kappa between LLM\-as\-Judge and human experts on annotated data\.
#### Baseline:
Initial prompt engineering iteration \(no business feature identification, no RAG retrieval, and no LLM semantic filtering\) versus the full pipeline\.
#### Implementation:
The framework is implemented in Python 3\.10\+ as a modular CLI tool, with Milvus for vector indexing, OpenAI API for LLM inference and evaluation, Apache Airflow for batch orchestration, and Docker for containerized deployment\.
### 4\.2Human Assessment
We conducted a blind data review \(see Table[1](https://arxiv.org/html/2609.19615#S4.T1)\)\. For each of 100 sampled semantic mappings \(business features↔\\leftrightarrowbusiness nodes↔\\leftrightarrowSQL\-like rules\), 5 human experts rated semantic correctness on a 100\-point scale\. The full pipeline achieved a mean score of 82\.3 \(σ=9\.7\\sigma=9\.7\), compared to the baseline of 51\.6 \(σ=14\.2\\sigma=14\.2\)—a statistically significant improvement \(p<0\.01p<0\.01, two\-tailed t\-test\)\.
Table 1:Semantic quality comparison
### 4\.3Noise Reduction
Hard\-coded and LLM\-based filtering collectively eliminated 74% of candidate rules as noise \(see Table[2](https://arxiv.org/html/2609.19615#S4.T2)\)\. The retained 26% of rules account for\>\>90% of semantic coverage, confirming that a small number of critical semantics drive the most business insights\.
Table 2:Noise filtering effectiveness
### 4\.4Manual Effort Reduction
To quantify the efficiency gains of our system, we also conducted a controlled experiment comparing a data science team with 5 human experts against the automated pipeline on equivalent semantic mapping tasks\. Human experts required an average of 10 hours per week to generate and correct semantic mappings, whereas our system reduced this effort to approximately 2 hours—an 80% reduction\. Moreover, for querying newly defined metrics, the manual workflow consumed 4 hours per query \(including SQL construction and data validation\), while the semantic layer enabled self\-service answers in under 25 minutes\. These results demonstrate that automation substantially reduces the operational burden on domain experts, allowing them to focus on higher\-level analytical work\.
### 4\.5LLM\-as\-Judge Agreement
We constructed a golden set of 500 manually annotated semantic mappings \(business features↔\\leftrightarrowbusiness nodes↔\\leftrightarrowSQL\-like rules\) and used an LLM\-as\-Judge with a structured prompt to rate the quality of those mappings on a 1–5 scale\. Comparing the LLM\-as\-Judge ratings against human expert annotations on the same set yielded a Cohen’s kappa of 0\.87, indicating near\-perfect agreement and validating the framework’s reliability for automated quality assessment\. This result is particularly significant because it demonstrates that LLMs can detect quality degradation early and trigger targeted refinement, ensuring long\-term stability in production environments\. This closed\-loop evaluation mechanism fundamentally shifts the maintenance paradigm from reactive, expert\-dependent corrections to proactive, automated quality stewardship\.
### 4\.6Ablation Study
We finally conducted an ablation study to isolate the contribution of each major component \(see Table[3](https://arxiv.org/html/2609.19615#S4.T3)\)\. The experiment confirms that every component contributes positively, with enumeration normalization, business feature identification and embedding search retrieval being the most critical\.
Table 3:Ablation study with human assessment
## 5Conclusion
In this paper, we formalize and address the problem of business semantic layer induction from raw application telemetry\. Unlike prior work on log parsing and schema matching, our framework tackles the unique challenge of deriving hierarchical, business\-aligned semantics without manual curation or labeled training data\. Our contributions are fourfold: \(1\) a systematic data refinement pipeline combining LLM\-driven column importance scoring and enumeration normalization to substantially reduce feature sparsity; \(2\) a hierarchical abstraction framework that decomposes the problem into coarse\-grained business feature identification and fine\-grained business node classification, mirroring the natural reasoning structure of business analytics; \(3\) a principled induction pipeline integrating hybrid retrieval, multi\-stage filtering, and contrastive clustering with canonical naming; and \(4\) an integrated LLM\-as\-Judge evaluation achieving 0\.87 Cohen’s kappa with human experts, enabling scalable and continuous quality monitoring\. Extensive experiments on production\-scale e\-commerce telemetry demonstrate that our approach reduces manual maintenance effort by 80%, improves semantic quality from 50 to 80\+ on a 100\-point scale, and filters 74% of noisy candidates\. By shifting the burden from ad\-hoc parser maintenance to principled semantic induction, our framework empowers organizations to focus on extracting actionable business insights rather than debating customized logics\. Future work includes incorporating causal reasoning and cross\-domain transfer learning to further enhance the generalization and analytical depth of the induced semantic layer\.
## References
- \[1\]H\. Bian\(2025\)LLM\-empowered knowledge graph construction: a survey\.arXiv preprint arXiv:2510\.20345\.Cited by:[§2\.2](https://arxiv.org/html/2609.19615#S2.SS2.p1.1)\.
- \[2\]J\. Dąbrowski, M\. Janicka, Ł\. Sienkiewicz, G\. Stomfai, J\. Dietmar, F\. Barile, M\. Polignano, C\. Pomo, and A\. Srivastava\(2025\)The synerise dataset: an e\-commerce dataset for sequential recommendation, universal behavior modeling and deep relational learning\.InProceedings of the Recommender Systems Challenge 2025,pp\. 1–6\.Cited by:[§4\.1](https://arxiv.org/html/2609.19615#S4.SS1.SSS0.Px1.p1.1)\.
- \[3\]P\. He, J\. Zhu, Z\. Zheng, and M\. R\. Lyu\(2017\)Drain: an online log parsing approach with fixed depth tree\.In2017 IEEE international conference on web services \(ICWS\),pp\. 33–40\.Cited by:[§2\.1](https://arxiv.org/html/2609.19615#S2.SS1.p1.1)\.
- \[4\]S\. Hoseini, J\. Theissen\-Lipp, and C\. Quix\(2024\)A survey on semantic data management as intersection of ontology\-based data access, semantic modeling and data lakes\.Journal of Web Semantics81,pp\. 100819\.Cited by:[§2\.2](https://arxiv.org/html/2609.19615#S2.SS2.p1.1)\.
- \[5\]J\. Kropshofer, J\. Schrott, W\. Wöß, and L\. Ehrlinger\(2025\)A survey on the functionalities of data catalog tools\.IEEE Access\.Cited by:[§2\.2](https://arxiv.org/html/2609.19615#S2.SS2.p1.1)\.
- \[6\]F\. Y\. Loumachi, M\. C\. Ghanem, and M\. A\. Ferrag\(2025\)Advancing cyber incident timeline analysis through retrieval\-augmented generation and large language models\.Computers14\(2\),pp\. 67\.Cited by:[§2\.3](https://arxiv.org/html/2609.19615#S2.SS3.p1.1)\.
- \[7\]J\. Piet, V\. Fang, R\. Khare, S\. Coull, V\. Paxson, R\. A\. Popa, and D\. Wagner\(2025\)Semantic\-aware parsing for security logs\.arXiv preprint arXiv:2506\.17512\.Cited by:[§2\.1](https://arxiv.org/html/2609.19615#S2.SS1.p1.1)\.
- \[8\]V\. Pyrih, A\. Rebmann, and H\. van der Aa\(2025\)LLMs that understand processes: instruction\-tuning for semantics\-aware process mining\.In2025 7th International Conference on Process Mining \(ICPM\),pp\. 1–8\.Cited by:[§2\.1](https://arxiv.org/html/2609.19615#S2.SS1.p1.1)\.
- \[9\]M\. Tonnarelli, I\. Kumara, S\. Driessen, D\. A\. Tamburri, W\. Van Den Heuvel, and P\. Oor\(2025\)Data catalog tools: a systematic multivocal literature review\.Journal of Systems and Software,pp\. 112584\.Cited by:[§2\.2](https://arxiv.org/html/2609.19615#S2.SS2.p1.1)\.
- \[10\]Z\. Yang, Y\. Wang, Z\. Shi, Y\. Yao, L\. Liang, K\. Ding, E\. Yilmaz, H\. Chen, and Q\. Zhang\(2025\)Eventrag: enhancing llm generation with event knowledge graphs\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics,pp\. 16967–16979\.Cited by:[§2\.3](https://arxiv.org/html/2609.19615#S2.SS3.p1.1)\.
- \[11\]C\. Zhang, W\. Xu, J\. Liu, L\. Zhang, G\. Liu, J\. Guan, Q\. Zhou, and S\. Zhou\(2025\)SemanticLog: towards effective and efficient large\-scale semantic log parsing\.IEEE Transactions on Software Engineering\.Cited by:[§2\.1](https://arxiv.org/html/2609.19615#S2.SS1.p1.1)\.
- \[12\]Z\. Zhao, X\. Luo, M\. Chen, and L\. Ma\(2024\)A survey of knowledge graph construction using machine learning\.Computer Modeling in Engineering & Sciences139\(1\),pp\. 225\.Cited by:[§2\.2](https://arxiv.org/html/2609.19615#S2.SS2.p1.1)\.
- \[13\]A\. Zhong, D\. Mo, G\. Liu, J\. Liu, Q\. Lu, Q\. Zhou, J\. Wu, Q\. Li, and Q\. Wen\(2024\)Logparser\-llm: advancing efficient log parsing with large language models\.InProceedings of the 30th ACM SIGKDD,pp\. 4559–4570\.Cited by:[§2\.1](https://arxiv.org/html/2609.19615#S2.SS1.p1.1)\.
- \[14\]W\. Zhu, X\. Li, L\. Wang, J\. Wang, and Y\. Wei\(2026\)TM\-rag: a tree\-mapped retrieval\-augmented generation framework for construction claim report generation\.Advanced Engineering Informatics69,pp\. 104092\.Cited by:[§2\.3](https://arxiv.org/html/2609.19615#S2.SS3.p1.1)\.Similar Articles
GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions
The paper introduces GROUND, a governed semantic-retrieval framework that constrains LLM-generated analytics to approved business definitions, effectively reducing hallucinations in enterprise data analysis while ensuring compliance with security and data policies.
Self-Describing Structured Data with Dual-Layer Guidance: A Lightweight Alternative to RAG for Precision Retrieval in Large-Scale LLM Knowledge Navigation
SDSR proposes lightweight self-describing structured data with dual-layer guidance to exploit LLM primacy bias, achieving 100% routing accuracy without vector DBs.
A Semantic-Layer-Mediated Agent for Natural Language to SQL over Heterogeneous Enterprise Databases
This paper presents a semantic-layer-mediated NL2SQL agent that decouples intent from physical execution by reasoning over a curated semantic model, achieving 94.15% execution accuracy on the Spider2-snow benchmark.
RSF-GLLM: Bridging the Semantic Gap in Multi-Hop Knowledge Graph QA via Recurrent Soft-Flow and Decoupled LLM Generation
This paper introduces RSF-GLLM, a framework that decouples differentiable graph reasoning from LLM generation to address the semantic gap in multi-hop knowledge graph question answering, achieving competitive performance with superior inference efficiency.
LLMs on Tabular Data with Limited Semantics: Evidence from Industrial Car Retrofit Prediction
This paper evaluates LLM-based strategies (embedding, prompt, hybrid) against classical tabular models on an industrial car retrofit prediction dataset with hashed categorical features. It finds that tree ensembles outperform LLMs overall, but embeddings and hybrid approaches remain useful, while direct prompting fails without semantic cues.