REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles
Summary
REALMS is a conversational AI system for real-time exact audience sizing using LLMs and embeddings, enabling marketers to query high-dimensional profile data with natural language and receive precise counts quickly in production.
View Cached Full Text
Cached at: 09/28/26, 09:39 AM
# REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles Source: [https://arxiv.org/html/2609.30547](https://arxiv.org/html/2609.30547) Haixu Ma\*††thanks:\*All authors contributed equally\.Aditya Bansal\*Affiliation:Adobe Inc\. San Jose, USA adibansal@adobe\.comAffiliation:Affiliation:Shubham Lohiya\*Affiliation:Adobe Inc\. San Jose, USA slohiya@adobe\.comAffiliation:Affiliation:Sumit Ranjan\*Affiliation:Adobe Inc\. San Jose, USA sumit\.nitt@gmail\.com ###### Abstract Audience sizing is a critical component of digital marketing\. It enables precise resource allocation, campaign planning, and performance optimization\. Traditional approaches using skeleton audiences, sampling, or predictive modeling suffer from significant delays, estimation errors, and poor scalability over high\-dimensional profile data\. We present REALMS \(Real\-timeExactAudience sizing viaLLM\-basedMulti\-attributeSearch\), a conversational system for exact audience sizing deployed in production on an enterprise customer data platform\. REALMS enables marketers to query massive profile stores with millions of profiles and thousands of attributes using natural language and receive precise counts in seconds\. The system introduces three key components: \(1\) a categorical attribute retrieval mechanism using embedding\-based vector search to dynamically identify relevant schema attributes without manual configuration; \(2\) an LLM\-powered NL2SQL pipeline with template\-based in\-context learning for accurate query generation over complex nested schemas; and \(3\) schema standardization enabling industry\-agnostic deployment across diverse enterprise environments\. Evaluation on real enterprise data demonstrates strong recall for attribute retrieval, high SQL execution accuracy, and low latency, which enables real\-time interactive audience insights where prior methods required hours\. ###### Index Terms: Audience Sizing, Conversational Systems, Embedding\-based Retrieval, High\-Dimensional Data, Large Language Models, NL2SQL, Real\-time Analytics ## IIntroduction Digital marketing platforms rely on audience sizing to estimate how many users satisfy behavioral, demographic, and transactional constraints prior to campaign deployment\. Accurate estimates support campaign planning, budget allocation, and personalization, whereas inaccurate estimates can lead to inefficient spending and degraded outcomes\. The increasing availability of large\-scale user data has further strengthened the role of analytics in enterprise decision making\[[20](https://arxiv.org/html/2609.30547#bib.bib2),[2](https://arxiv.org/html/2609.30547#bib.bib3)\]\. In enterprise customer data platforms, audience sizing is often implemented via approximate query processing, such as sampling, sketches, or predictive modeling\. Although these techniques offer established efficiency–accuracy trade\-offs\[[1](https://arxiv.org/html/2609.30547#bib.bib4),[3](https://arxiv.org/html/2609.30547#bib.bib5)\], they can be brittle in practice: estimation error is difficult to control for correlated, high\-dimensional, and nested attributes, and batch\-oriented pre\-computation can inflate end\-to\-end latency, limiting iterative and interactive use\. In parallel, natural language interfaces have become an important access modality for non\-technical data exploration\. Conversational information seeking systems reduce the need to manually author structured queries\[[14](https://arxiv.org/html/2609.30547#bib.bib6),[4](https://arxiv.org/html/2609.30547#bib.bib7)\]\. Audience sizing, however, differs from conventional conversational information retrieval because it requires executing aggregation queries over structured profile stores rather than retrieving ranked documents, necessitating robust intent understanding coupled with reliable structured execution\. Recent advances in Large Language Models \(LLMs\) have improved Natural\-Language\-to\-SQL \(NL2SQL\) and semantic parsing\[[23](https://arxiv.org/html/2609.30547#bib.bib8),[21](https://arxiv.org/html/2609.30547#bib.bib9),[18](https://arxiv.org/html/2609.30547#bib.bib10)\], but deploying these methods for enterprise audience sizing raises additional requirements beyond standard benchmark settings\. We summarize the key requirements that shape our system design as follows: \(1\)Adaptability to high\-dimensional schemas and heterogeneous data sources: Enterprise profile schemas are highly diverse and often contain thousands of attributes organized in nested hierarchies\. In addition, customer data platforms integrate multiple sources, such as web interactions, transactions, and offline records, with differing representations and semantics\. A practical solution should therefore perform robust schema grounding and schema alignment, mapping natural\-language constraints to the appropriate structured attributes without manual rules or per\-customer customization; \(2\)Scalability over large profile datasets: Audience sizing operates over very large profile stores, where high dimensionality and nested structures make exact counting computationally expensive\. The system should employ efficient retrieval and execution strategies to support exact computation while controlling resource consumption; \(3\)Real\-time conversational interaction: In an AI\-assistant setting, end\-to\-end latency directly affects usability\. The system should reliably produce executable queries and return precise results within seconds to preserve interactive dialogue\. Fig\. 1:Comparative Analysis of Conventional and REALMS \(Our Approach\) for Real\-time Exact Audience Sizing Over Large\-Scale Enterprise Profiles\.To address these challenges, we presentREALMS\(Real\-time Exact Audience sizing via LLM\-based Multi\-attribute Search\), a conversational system for natural\-language audience sizing over large\-scale enterprise profile datasets\. REALMS integrates information retrieval, semantic parsing, and database execution within an AI assistant\. The system comprises three components: \(1\)Schema\-aware attribute retrieval, which uses embedding\-based dense retrieval to map user utterances to relevant schema attributes\[[17](https://arxiv.org/html/2609.30547#bib.bib13),[9](https://arxiv.org/html/2609.30547#bib.bib14)\]; \(2\)LLM\-driven structured query generation, which produces executable NL2SQL queries over deeply nested schemas via in\-context generation; \(3\)Schema standardization and execution, which normalizes heterogeneous data sources and supports exact counting at scale\. TABLE I:Comparison of audience sizing approaches\.Our experiments on real enterprise datasets show high attribute\-retrieval recall, strong SQL execution accuracy, and low end\-to\-end latency, enabling interactive audience insights in seconds where conventional pipelines may require minutes to hours\. More broadly, REALMS exemplifies a conversational analytics paradigm in which an LLM\-based assistant functions as a semantic query planner that bridges information retrieval and database execution: it retrieves relevant schema elements and composes structured aggregation queries to return exact audience counts\. We further demonstrate practicality through a production deployment in an enterprise conversational AI assistant\. ## IIRelated Work Conversational Information Retrieval\.Foundational work on conversational search\[[14](https://arxiv.org/html/2609.30547#bib.bib6)\]and conversational information seeking\[[22](https://arxiv.org/html/2609.30547#bib.bib15)\]investigates multi\-turn dialogue for complex information needs\. Liu et al\.\[[12](https://arxiv.org/html/2609.30547#bib.bib20)\]proposed SUQL, augmenting SQL with free\-text primitives for conversational search over hybrid data\. While prior conversational systems primarily retrieve ranked documents\[[4](https://arxiv.org/html/2609.30547#bib.bib7)\], our work retrieves structured schema elements and computes exact aggregation queries, extending conversational IR to structured enterprise analytics\. Natural Language Interfaces to Databases \(NL2SQL\)\.Semantic parsing into SQL has advanced through benchmarks such as WikiSQL\[[23](https://arxiv.org/html/2609.30547#bib.bib8)\], Spider\[[21](https://arxiv.org/html/2609.30547#bib.bib9)\], and BIRD\[[11](https://arxiv.org/html/2609.30547#bib.bib19)\], which evaluates text\-to\-SQL over large\-scale databases with noisy data\. LLMs with in\-context learning achieve competitive accuracy without fine\-tuning\[[16](https://arxiv.org/html/2609.30547#bib.bib16)\], with DIN\-SQL\[[13](https://arxiv.org/html/2609.30547#bib.bib17)\]and DAIL\-SQL\[[7](https://arxiv.org/html/2609.30547#bib.bib18)\]further advancing prompt\-based generation through decomposed reasoning and optimized prompt design\. However, existing work predominantly assumes curated schemas of moderate size, whereas enterprise profile stores contain thousands of heterogeneous, deeply nested attributes requiring real\-time interactive querying\. LLM\-based Enterprise Analytics Systems\.SiriusBI\[[8](https://arxiv.org/html/2609.30547#bib.bib21)\]deploys an LLM\-based business intelligence system with multi\-round dialogue and schema\-aware prompting\. Floratou et al\.\[[6](https://arxiv.org/html/2609.30547#bib.bib1)\]identify persistent production challenges including schema complexity and execution reliability\. Sharma et al\.\[[18](https://arxiv.org/html/2609.30547#bib.bib10)\]propose tree\-guided token decoding for schema\-aware SQL generation\. Our system addresses these challenges through template\-based in\-context learning, categorical attribute retrieval, and schema standardization\. Schema Matching and Data Integration\.Classical schema matching\[[15](https://arxiv.org/html/2609.30547#bib.bib11)\]and data integration\[[5](https://arxiv.org/html/2609.30547#bib.bib12)\]address aligning heterogeneous sources into unified queryable representations\. Recent work applies LLMs: ReMatch\[[19](https://arxiv.org/html/2609.30547#bib.bib22)\]employs retrieval\-enhanced LLMs for large\-scale schema matching, and CRUSH4SQL\[[10](https://arxiv.org/html/2609.30547#bib.bib23)\]uses LLM\-generated schema hallucinations with composite dense retrieval to identify relevant schema subsets\. Our system similarly leverages embedding\-based retrieval but introduces categorical bucketing for diverse and precise attribute coverage\. ## IIISystem Design To address the challenges of schema heterogeneity, large\-scale profile data, and interactive latency outlined in Section 1, REALMS is designed around three core principles that jointly enable accurate and efficient audience sizing from natural\-language requests\. First, REALMS adopts*schema\-agnostic attribute retrieval*to bridge the gap between user language and highly diverse enterprise schemas\. Enterprise profile stores often contain thousands of attributes organized in deeply nested structures, with naming conventions and semantics that vary substantially across customers and data sources\. To address this challenge, REALMS leverages dense embedding representations combined with categorical attribute bucketing to encode schema semantics into a shared embedding space\[[17](https://arxiv.org/html/2609.30547#bib.bib13)\]\. At query time, natural\-language constraints are matched against candidate attributes through semantic retrieval rather than hand\-crafted rules or schema\-specific mappings\. This design allows the system to generalize across heterogeneous schemas while minimizing customer\-specific engineering effort\. Second, REALMS employs*template\-based in\-context learning*for robust NL2SQL generation\[[13](https://arxiv.org/html/2609.30547#bib.bib17),[7](https://arxiv.org/html/2609.30547#bib.bib18)\]\. Instead of relying on task\-specific fine\-tuning for each customer schema, the system constructs prompts using retrieved schema context and reusable query templates\. Large language models then translate natural\-language audience definitions into executable SQL queries over complex nested profile structures\. This approach improves portability across domains while maintaining the flexibility needed to support diverse audience construction tasks\. Third, REALMS introduces a*schema standardization layer*that provides a unified execution interface over heterogeneous enterprise data sources\. Customer data platforms typically integrate information from web interactions, transactions, and offline records, each with different schemas and storage formats\. REALMS materializes these data into a standardized columnar representation that abstracts away source\-specific differences and exposes a consistent query surface\. By decoupling query generation from physical data organization, the standardization layer enables efficient execution and exact audience counting at scale while preserving compatibility with existing customer data infrastructures\. Together, these design principles establish the foundation of REALMS, enabling semantic understanding of natural\-language requests, portable query generation across heterogeneous schemas, and scalable exact computation suitable for real\-time conversational interaction\. Below is a detailed introduction for the architexture of our system\. The high\-level architecture is illustrated in Figure[2](https://arxiv.org/html/2609.30547#S3.F2)\. The system comprises two stages: anoffline pre\-processingstage that prepares and indexes enterprise profile data, and anAI Assistant runtimethat handles real\-time query processing\. This separation ensures that computationally expensive operations, such as data materialization, embedding generation, and index construction are performed asynchronously, while the runtime path remains lightweight and latency\-optimized\. Fig\. 2:High\-level system architecture of REALMS: offline pre\-processing and real\-time AI Assistant runtime\.### III\-AOffline Pre\-processing The offline stage transforms raw enterprise profile data into a queryable, indexed representation optimized for real\-time retrieval and SQL execution\. Figure[3](https://arxiv.org/html/2609.30547#S3.F3)illustrates the pre\-processing pipeline\. Fig\. 3:Offline pre\-processing and hydration flow\.Scheduled compute jobs export and flatten hierarchical profile records from the source data platform into a columnar representation stored in a scalable analytical database\. This materialization step is necessary because enterprise profiles are typically stored in deeply nested schemas, where direct analytical queries over the raw structures incur prohibitive latency\. The columnar representation enables efficient predicate evaluation and aggregation while preserving the full attribute space\. Concurrently, the system generates attribute\-level metadata for each schema element, including natural\-language descriptions derived from attribute paths, data types, and sample values\. Dense vector embeddings are computed for each attribute description using a sentence embedding model\[[17](https://arxiv.org/html/2609.30547#bib.bib13)\]\. Attributes are then organized into*categorical buckets*based on their semantic type, for example, demographic attributes \(age, gender, location\), behavioral attributes \(page views, app visits\), and transactional attributes \(purchase history, order value\)\. Each bucket is independently indexed in a vector search service to support parallel, category\-aware retrieval\. This categorical organization addresses a key limitation of naïve dense retrieval over large schemas: enterprise schemas often contain hundreds of semantically similar attributes \(e\.g\., multiple date fields or numeric counters\), and without categorical separation, retrieval results tend to concentrate within a single semantic cluster, reducing coverage of the attribute space and degrading downstream SQL generation accuracy\. Additionally, data governance labels from the platform schema are ingested during pre\-processing to annotate attributes with access control and privacy policies, enabling compliance enforcement at query time\. ### III\-BAI Assistant Runtime Query Pipeline At inference time, REALMS transforms a natural\-language audience request into an executable analytical query through a five\-stage runtime pipeline and end\-to\-end runtime flow, illustrated in Figures[4](https://arxiv.org/html/2609.30547#S3.F4)and[5](https://arxiv.org/html/2609.30547#S3.F5)\. The pipeline combines schema\-aware retrieval, graph\-based retrieval\-augmented generation, LLM\-powered semantic parsing, scalable query execution, and result interpretation to support accurate and interactive audience sizing over large\-scale enterprise profile datasets\. Fig\. 4:Five stages demonstration of the REALMS runtime query processing pipeline\.Fig\. 5:End\-to\-end REALMS runtime flow\.Schema\-Aware Attribute Retrieval\.Given a user query, REALMS first identifies the subset of schema attributes most relevant to the requested audience definition\. The query is encoded into a dense embedding using the same model employed during offline indexing\[[9](https://arxiv.org/html/2609.30547#bib.bib14)\]\. The embedding is then used to perform parallel approximate nearest\-neighbor searches across all attribute buckets\. Unlike global retrieval approaches that may overemphasize a single attribute category, REALMS retrieves the top\-kkcandidate attributes from each semantic bucket, ensuring balanced coverage across heterogeneous profile dimensions\. For example, given the query*“How many users in California made a purchase in the last 30 days?”*, the retrieval stage identifies geographic attributes from demographic schemas and purchase\-related attributes from behavioral schemas\. The retrieved attributes, together with their schema paths, descriptions, and data types, are assembled into a structured schema context that grounds downstream query generation\. Knowledge\-Graph Retrieval\-Augmented Generation \(KG\-RAG\)\.After identifying the relevant schema attributes, REALMS retrieves semantically and structurally similar examples from a repository of historical audience requests and their validated SQL implementations\. Rather than maintaining a flat collection of examples, REALMS organizes this repository as a query knowledge graph, where nodes represent natural\-language audience definitions, schema attributes, SQL query templates, operators, and aggregation patterns, while edges capture semantic, structural, and execution\-level relationships among them\. Given a new request, REALMS performs retrieval over the knowledge graph using both the user query and the retrieved schema context\. The retrieval process identifies the top\-kkquestion\-SQL pairs that are most relevant to the target analytical intent and query structure\. This graph\-based retrieval strategy differs from conventional RAG systems that rely solely on semantic similarity in a vector index\. By explicitly modeling relationships among user intents, schema elements, and SQL patterns, the knowledge graph enables retrieval of examples that are not only semantically related but also structurally aligned with the target query\. The retrieved examples serve as retrieval\-augmented context for downstream SQL generation\. Because enterprise audience\-sizing requests frequently exhibit recurring analytical patterns, such as demographic filtering, behavioral segmentation, temporal constraints, cohort construction, and aggregation operations, grounding generation in previously validated SQL examples significantly improves generation accuracy, reduces hallucinations, and promotes consistent query construction across heterogeneous schemas\. LLM\-Powered NL2SQL Generation\.Using both the schema context and the retrieved examples, REALMS translates the user request into an executable SQL query\. The generation prompt consists of three components: \(1\) the original natural\-language request, \(2\) the retrieved schema attributes and metadata, and \(3\) the top\-kkretrieved question\-SQL demonstrations obtained from the KG\-RAG stage\. In addition to retrieval\-augmented examples, REALMS employs a library of template\-based prompt structures that capture common analytical patterns such as filtering, aggregation, ranking, counting, percentage computation, and cohort analysis\. These templates remain independent of any organization\-specific schema and operate through placeholder\-based attribute references\. During inference, retrieved schema attributes are injected into the prompt, enabling the same prompt framework to generalize across heterogeneous enterprise datasets without customer\-specific fine\-tuning\[[16](https://arxiv.org/html/2609.30547#bib.bib16),[13](https://arxiv.org/html/2609.30547#bib.bib17),[7](https://arxiv.org/html/2609.30547#bib.bib18)\]\. Guided by both retrieved examples and schema context, the language model generates SQL over the standardized profile representation, resolving attribute references and constructing the appropriate filtering, grouping, aggregation, and temporal operators required by the user request\. SQL Validation and Execution\.Since generated queries directly access enterprise data, REALMS performs a validation stage prior to execution\. The validator checks schema consistency, attribute existence, type compatibility, aggregation correctness, and governance constraints\. In particular, governance enforcement ensures that restricted or personally identifiable information \(PII\) attributes cannot be exposed through generated queries or query outputs\. Queries that fail validation are not executed\. Instead, structured validation feedback is returned to the language model, enabling iterative query refinement in a manner similar to self\-correction frameworks proposed in recent NL2SQL systems\[[13](https://arxiv.org/html/2609.30547#bib.bib17)\]\. Once validated, the query is executed on a columnar analytical database engine optimized for large\-scale aggregation workloads\. This execution layer enables exact audience counting and analytical computation over hundreds of millions of profile records while maintaining interactive response latency\. Result Interpretation and Explanation\.Finally, REALMS converts execution results into a user\-facing response\. In addition to returning the requested audience size or aggregate statistic, the system generates a concise natural\-language explanation describing the interpreted query logic, selected attributes, filtering criteria, and aggregation operations\. This explanation provides transparency into the system’s reasoning process and allows users to verify that the generated query accurately reflects their intent\. Furthermore, the explanation serves as a foundation for iterative refinement, enabling users to modify constraints, add additional conditions, or explore related audience segments through subsequent conversational turns\. By combining exact execution with interpretable explanations, REALMS delivers both analytical accuracy and usability for non\-technical business users\. ## IVEvaluation We evaluate REALMS along three dimensions: \(1\) end\-to\-end query correctness, \(2\) attribute retrieval component effectiveness, and \(3\) system latency and scalability under realistic workloads\. ### IV\-AEvaluation Dataset We construct an evaluation dataset consisting of 600 natural\-language audience\-sizing requests designed to reflect realistic enterprise analytics workloads\. The dataset combines two complementary sources\. First, we collect real customer questions observed in production environments\. These queries capture authentic business terminology, natural language variations, and practical information needs encountered by marketing and customer\-engagement teams\. Second, we supplement the real queries with LLM\-generated questions to increase coverage and systematically explore a broader space of linguistic expressions and logical compositions\. The generated queries are designed to preserve realistic business semantics while varying attribute combinations, aggregation structures, and constraint formulations\. To evaluate robustness across common analytical tasks, the dataset covers three representative query intents: \(1\) Count queries, which return the exact cardinality of profiles satisfying a predicate \(e\.g\., “How many profiles have visited the mobile application within the last six months and are older than 30 years?”\); \(2\) Top\-kkaggregation queries, which identify the highest\-ranked groups according to an aggregate metric \(e\.g\., “What are the top five cities with the highest average number of e\-commerce purchases under the specified constraints?”\); \(3\) Percentage queries, which compute proportions relative to a reference population \(e\.g\., “What percentage of customers in Segment A reside in California?”\)\. To further evaluate compositional reasoning, we vary query complexity by controlling the number of referenced attributes \(one, two, or three attributes per query\)\. Increasing the number of attributes simultaneously increases the difficulty of schema grounding, logical composition, and SQL generation\. ### IV\-BSystem Latency and Scalability To evaluate the operational characteristics of REALMS, we conduct load\-testing experiments in a staging environment under varying request rates and concurrency levels\. For each configuration, we report the total number of requests, failure count, and latency statistics including median, mean, 95th percentile, and 99th percentile response times\. Table[II](https://arxiv.org/html/2609.30547#S4.T2)summarizes the results\. REALMS maintains stable performance under moderate load levels of 10–30 requests per minute \(RPM\), exhibiting no observed failures and median response times below nine seconds\. As workload intensity increases to 60 RPM with ten concurrent users, the system continues to provide interactive response times, achieving a median latency of 10\.0 seconds and a 95th\-percentile latency of 14\.0 seconds\. Under this configuration, only two failures are observed among 105 requests, indicating that the system remains operational while approaching the capacity limits of the current deployment\. These results demonstrate that the retrieval, generation, and execution pipeline can support interactive conversational workloads while maintaining predictable latency characteristics across a range of operating conditions\. TABLE II:Load test results under varying RPM and conc\. users\. ### IV\-CRetrieval and Query Correctness REALMS employs schema\-aware attribute retrieval to identify a compact set of candidate attributes for downstream SQL generation\. This retrieval stage serves two purposes: \(1\) reducing prompt size to remain within the context constraints of the language model, and \(2\) improving generation accuracy by grounding the model on the most relevant schema elements\. We evaluate retrieval quality using two metrics\. Recall@kkmeasures the fraction of ground\-truth attributes that appear among the top\-kkretrieved candidates\. Exact\-match rate@kkmeasures the percentage of queries for which all required attributes are successfully retrieved within the top\-kkresults\. For end\-to\-end query correctness, we compare generated SQL against reference SQL annotations\. Because multiple SQL expressions may be semantically equivalent despite syntactic differences, we prioritize execution match, which executes both queries against the underlying analytical database and compares their outputs\. When execution\-based comparison is unavailable, we additionally report an LLM\-as\-a\-judge metric that assesses semantic equivalence between generated and reference SQL queries\. TABLE III:Performance on the evaluation dataset \(accuracy %\)\.Table[III](https://arxiv.org/html/2609.30547#S4.T3)reports performance whenk=5k=5\. REALMS achieves 94% Recall@kkand 93% Exact\-match@kk, indicating that the retrieval module consistently identifies the attributes required for query construction\. This strong retrieval performance translates into high downstream query accuracy, with 90% execution match and 95% semantic equivalence according to the LLM\-based evaluator\. Together, these results demonstrate that schema\-aware retrieval and KG\-RAG effectively ground SQL generation over large heterogeneous enterprise schemas\. To better understand the behavior of REALMS, we also conduct a series of additional analyses\. First, we perform an ablation study to quantify the contribution of schema\-aware retrieval, and KG\-RAG\. Second, we evaluate robustness under increasing query complexity by varying the number of referenced attributes\. Finally, we analyze performance across different analytical intents, including count, top\-kk, and percentage queries\. Together, these experiments provide a more comprehensive understanding of the factors driving end\-to\-end query correctness\. #### IV\-C1Ablation Study TABLE IV:Ablation study of REALMS for structured query correctness\.Table[IV](https://arxiv.org/html/2609.30547#S4.T4)evaluates the contribution of individual components in REALMS\. Removing either schema\-aware retrieval or KG\-RAG substantially degrades execution accuracy, indicating that both schema grounding and retrieval\-augmented examples are critical for robust NL2SQL generation\. Randomly retrieved examples provide limited benefit, demonstrating the importance of knowledge\-graph\-guided retrieval\. The validation\-and\-refinement loop further improves correctness by correcting schema and syntax errors prior to execution\. #### IV\-C2Query Complexity Breakdown TABLE V:Performance under different query complexities\.Table[V](https://arxiv.org/html/2609.30547#S4.T5)reports performance as a function of query complexity\. As the number of referenced attributes increases, the difficulty of schema grounding and logical composition grows\. Nevertheless, REALMS maintains strong retrieval and execution accuracy, demonstrating robustness to increasingly complex audience definitions\. #### IV\-C3Query Intent Breakdown TABLE VI:Performance across query intents\.Table[VI](https://arxiv.org/html/2609.30547#S4.T6)breaks down performance by analytical intent\. REALMS achieves consistently strong results across count, top\-kk, and percentage queries, indicating that the proposed retrieval and generation pipeline generalizes beyond simple counting tasks and supports a diverse range of audience analytics workloads\. ## VConclusion We present REALMS, a real\-time conversational system for*exact*audience sizing over large\-scale enterprise profile stores\. REALMS unifies schema\-aware attribute retrieval, LLM\-driven NL2SQL generation, and a standardized execution layer to handle heterogeneous high\-dimensional schemas, nested data, and strict interactive latency constraints\. Experiments on real enterprise datasets demonstrate strong retrieval quality, high SQL execution accuracy, and low end\-to\-end latency, enabling audience insights in seconds\. More broadly, REALMS exemplifies conversational analytics in which an LLM acts as a semantic query planner that links information retrieval with database execution for structured aggregation\. ## References - \[1\]S\. Chaudhuri, R\. Motwani, and V\. Narasayya\(1998\)Random sampling for histogram construction: how much is enough?\.ACM SIGMOD Record27\(2\),pp\. 436–447\.Cited by:[§I](https://arxiv.org/html/2609.30547#S1.p2.1)\. - \[2\]H\. Chen, R\. H\. Chiang, and V\. C\. Storey\(2012\)Business intelligence and analytics: from big data to big impact\.MIS quarterly36\(4\),pp\. 1165–1188\.Cited by:[§I](https://arxiv.org/html/2609.30547#S1.p1.1)\. - \[3\]G\. Cormode, M\. Garofalakis, P\. J\. Haas, and C\. Jermaine\(2011\)Synopses for massive data: samples, histograms, wavelets, sketches\.Foundations and Trends in Databases4\(1\-3\),pp\. 1–294\.Cited by:[§I](https://arxiv.org/html/2609.30547#S1.p2.1)\. - \[4\]J\. Dalton, S\. Fischer, P\. Owoicho, F\. Radlinski, F\. Rossetto, J\. R\. Trippas, and H\. Zamani\(2022\)Conversational information seeking: theory and application\.InProceedings of the 45th international ACM SIGIR conference on research and development in information retrieval,pp\. 3455–3458\.Cited by:[§I](https://arxiv.org/html/2609.30547#S1.p3.1),[§II](https://arxiv.org/html/2609.30547#S2.p1.1)\. - \[5\]A\. Doan, A\. Halevy, and Z\. Ives\(2012\)Principles of data integration\.Elsevier\.Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p4.1)\. - \[6\]A\. Floratou, F\. Psallidas, F\. Zhao, S\. Deep, G\. Hagleither, W\. Tan, J\. Cahoon, R\. Alotaibi, J\. Henkel, A\. Singla,et al\.\(2024\)Nl2sql is a solved problem… not\!\.InCIDR,Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p3.1)\. - \[7\]D\. Gao, H\. Wang, Y\. Li, X\. Sun, Y\. Qian, B\. Ding, and J\. Zhou\(2024\)Text\-to\-sql empowered by large language models: a benchmark evaluation\.Proceedings of the VLDB Endowment17\(7\),pp\. 1132–1145\.Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p2.1),[§III\-B](https://arxiv.org/html/2609.30547#S3.SS2.p8.1),[§III](https://arxiv.org/html/2609.30547#S3.p3.1)\. - \[8\]J\. Jiang, H\. Xie, Y\. Shen, Z\. Zhang, M\. Lei, Y\. Zheng, Y\. Fang, C\. Li, D\. Huang, W\. Zhang,et al\.\(2024\)Siriusbi: building end\-to\-end business intelligence enhanced by large language models\.arXiv preprint arXiv:2411\.06102\.Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p3.1)\. - \[9\]V\. Karpukhin, B\. Oguz, S\. Min, P\. Lewis, L\. Wu, S\. Edunov, D\. Chen, and W\. Yih\(2020\)Dense passage retrieval for open\-domain question answering\.InProceedings of the 2020 conference on empirical methods in natural language processing \(EMNLP\),pp\. 6769–6781\.Cited by:[§I](https://arxiv.org/html/2609.30547#S1.p9.1),[§III\-B](https://arxiv.org/html/2609.30547#S3.SS2.p2.1)\. - \[10\]M\. Kothyari, D\. Dhingra, S\. Sarawagi, and S\. Chakrabarti\(2023\)CRUSH4SQL: collective retrieval using schema hallucination for text2sql\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 14054–14066\.Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p4.1)\. - \[11\]J\. Li, B\. Hui, G\. Qu, J\. Yang, B\. Li, B\. Li, B\. Wang, B\. Qin, R\. Geng, N\. Huo,et al\.\(2024\)Can LLM already serve as a database interface? A BIg Bench for Large\-Scale Database Grounded Text\-to\-SQL\.InAdvances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track,Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p2.1)\. - \[12\]S\. Liu, J\. Xu, W\. Tjangnaka, S\. Semnani, C\. Yu, and M\. Lam\(2024\)SUQL: conversational search over structured and unstructured data with large language models\.InFindings of the Association for Computational Linguistics: NAACL 2024,pp\. 4535–4555\.Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p1.1)\. - \[13\]M\. Pourreza and D\. Rafiei\(2023\)DIN\-SQL: decomposed in\-context learning of text\-to\-sql with self\-correction\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p2.1),[§III\-B](https://arxiv.org/html/2609.30547#S3.SS2.p11.1),[§III\-B](https://arxiv.org/html/2609.30547#S3.SS2.p8.1),[§III](https://arxiv.org/html/2609.30547#S3.p3.1)\. - \[14\]F\. Radlinski and N\. Craswell\(2017\)A theoretical framework for conversational search\.InProceedings of the 2017 conference on conference human information interaction and retrieval,pp\. 117–126\.Cited by:[§I](https://arxiv.org/html/2609.30547#S1.p3.1),[§II](https://arxiv.org/html/2609.30547#S2.p1.1)\. - \[15\]E\. Rahm and P\. A\. Bernstein\(2001\)A survey of approaches to automatic schema matching\.the VLDB Journal10\(4\),pp\. 334–350\.Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p4.1)\. - \[16\]N\. Rajkumar, R\. Li, and D\. Bahdanau\(2022\)Evaluating the text\-to\-sql capabilities of large language models\.arXiv preprint arXiv:2204\.00498\.Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p2.1),[§III\-B](https://arxiv.org/html/2609.30547#S3.SS2.p8.1)\. - \[17\]N\. Reimers and I\. Gurevych\(2019\)Sentence\-bert: sentence embeddings using siamese bert\-networks\.InProceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing \(EMNLP\-IJCNLP\),pp\. 3982–3992\.Cited by:[§I](https://arxiv.org/html/2609.30547#S1.p9.1),[§III\-A](https://arxiv.org/html/2609.30547#S3.SS1.p3.1),[§III](https://arxiv.org/html/2609.30547#S3.p2.1)\. - \[18\]C\. Sharma, R\. Narayanam, S\. Pal, K\. Yeturu, S\. K\. Saini, and K\. Mukherjee\(2025\)TTD\-sql: tree\-guided token decoding for efficient and schema\-aware sql generation\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track,pp\. 1287–1298\.Cited by:[§I](https://arxiv.org/html/2609.30547#S1.p3.1),[§II](https://arxiv.org/html/2609.30547#S2.p3.1)\. - \[19\]E\. Sheetrit, M\. Brief, M\. Mishaeli, and O\. Elisha\(2024\)ReMatch: retrieval enhanced schema matching with LLMs\.arXiv preprint arXiv:2403\.01567\.Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p4.1)\. - \[20\]M\. Wedel and P\. Kannan\(2016\)Marketing analytics for data\-rich environments\.Journal of marketing80\(6\),pp\. 97–121\.Cited by:[§I](https://arxiv.org/html/2609.30547#S1.p1.1)\. - \[21\]T\. Yu, R\. Zhang, K\. Yang, M\. Yasunaga, D\. Wang, Z\. Li, J\. Ma, I\. Li, Q\. Yao, S\. Roman,et al\.\(2018\)Spider: a large\-scale human\-labeled dataset for complex and cross\-domain semantic parsing and text\-to\-sql task\.InProceedings of the 2018 conference on empirical methods in natural language processing,pp\. 3911–3921\.Cited by:[§I](https://arxiv.org/html/2609.30547#S1.p3.1),[§II](https://arxiv.org/html/2609.30547#S2.p2.1)\. - \[22\]H\. Zamani, J\. R\. Trippas, J\. Dalton, and F\. Radlinski\(2022\)Conversational information seeking\.arXiv preprint arXiv:2201\.08808\.Cited by:[§II](https://arxiv.org/html/2609.30547#S2.p1.1)\. - \[23\]V\. Zhong, C\. Xiong, and R\. Socher\(2017\)Seq2sql: generating structured queries from natural language using reinforcement learning\.arXiv preprint arXiv:1709\.00103\.Cited by:[§I](https://arxiv.org/html/2609.30547#S1.p3.1),[§II](https://arxiv.org/html/2609.30547#S2.p2.1)\.
Similar Articles
REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening
This paper presents REALM, a coarse-to-fine generative framework for audio-driven reactive listening in embodied AI, addressing challenges in timing and facial motion with evaluations and robot deployment showing improvements.
Too Good to Be Real? Diagnosing and Reducing the Gap Between AI Preference and Real User Engagement
The paper diagnoses a systematic gap between AI preferences and real user engagement in content generation, finding that LLMs overemphasize logical structure while real engagement favors affective and expressive elements, and proposes OMRA to reduce this gap by 54.4%.
To Memories and Beyond: From Remembering to Knowing You across Long-Term Multimodal Personal Archives
This paper introduces ReaLMem, the first benchmark for evaluating long-term multimodal memory in AI using authentic personal archives, and proposes ChronoProfiler for temporal weighting to improve personalization in AI companions.
SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context
Introduces SocialPersona, a benchmark for evaluating multimodal large language models on their ability to recover revealed preferences from longitudinal social-media timelines and use them in personalized dialogue.
OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants
OmniInteract introduces a streaming benchmark for real-time omnimodal LLMs, evaluating online audio-visual processing with temporal grounding and interactive response requirements. Experiments show that current models perform poorly, with the best overall IA-QTF1 score reaching only 0.368.