Data Quality Rule Generation with LLMs
Summary
This paper introduces LeDQeR, a framework that uses large language models to automatically generate data quality rules for enterprise tools, improving rule coverage and reducing manual maintenance effort.
View Cached Full Text
Cached at: 09/10/26, 08:31 AM
# Data Quality Rule Generation with LLMs Source: [https://arxiv.org/html/2609.06053](https://arxiv.org/html/2609.06053) CCS:Information systems Data cleaningCCS:Applied computing Business rulesCCS:Applied computing Enterprise data managementAnna\-Christina Glock,Thomas Hütter[https://orcid.org/0000-0002-7190-6825](https://orcid.org/0000-0002-7190-6825)Affiliation:Software Competence Center,Hagenberg,Austriaemail:[Thomas\.Huetter@scch\.at](mailto:[email protected]),Johannes Fürnkranz[https://orcid.org/0000-0002-1207-0159](https://orcid.org/0000-0002-1207-0159)Affiliation:Johannes Kepler University,Linz,Austriaemail:[juffi@faw\.jku\.at](mailto:[email protected]),Wolfram Wöß[https://orcid.org/0009-0006-8352-0205](https://orcid.org/0009-0006-8352-0205)Affiliation:Johannes Kepler University,Linz,Austriaemail:[wolfram\.woess@jku\.at](mailto:[email protected]),Christine Dominka\-KissAffiliation:Austrian Post,Vienna,Austriaemail:[christine\.dominka\-kiss@post\.at](mailto:[email protected])andLisa Ehrlinger[https://orcid.org/0000-0002-1825-0097](https://orcid.org/0000-0002-1825-0097)Affiliation:Hasso Plattner Institute, University of Potsdam,Potsdam,Germanyemail:[lisa\.ehrlinger@hpi\.de](mailto:[email protected]) ###### Abstract\. The validation of data, such as customer and employee data, is an important task in many organizations\. Errors in data can have severe consequences\. For example, a wrong drug unit in a patient record can lead to life\-threatening medication errors, and a missing street number in an address to failed deliveries\. Companies often employ rule\-based enterprise data quality \(DQ\) tools, which allow domain experts to specify rules to validate the data over time\. While rule\-based DQ tools are computationally efficient and provide explainable reports, maintaining a comprehensive rule set manually is challenging, as domain experts often overlook essential rules, especially in complex domains and large data volumes\. Hence, closing these gaps remains an open problem in practice\. In this paper, we address the challenge of automated DQ rule generation\. For this, we formalize a generalizable*generate\-filter framework*and introduceLeDQeR, anLLM\-basedDQrule generation approach\. First, a large language model \(LLM\) generates candidates rules from an observed dirty data tuple for a given rule\-based DQ tool syntax\. Second, we apply four filter techniques that ensure the \(i\) executability, \(ii\) correctness, and \(iii\) generalizability, and avoid \(iv\) redundancy of the generated rules\. An extensive experimental evaluation suggests thatLeDQeRis able to produce effective and compact rule sets for various datasets and error types\. Figure 1\.LeDQeRextends the rule set of a rule\-based enterprise DQ tool\. Incoming data is validated against an existing rule setR=\{r1,…,rm\}R=\\\{r\_\{1\},\\ldots,r\_\{m\}\\\}\. Errors covered by a rule are detected, while errors not covered by any rule pass the validation unnoticed and cause failures in downstream processes\.LeDQeRcloses these gaps: undetected dirty tuples are passed to the generate\-filter pipeline, which produces new rules\{rm\+1,rm\+2\}\\\{r\_\{m\+1\},r\_\{m\+2\}\\\}that extend the existing rule set such that the same error type is detected in the future\. ###### Keywords: Data quality, rule generation, rule\-based tools, large language models ## 1\.Introduction Data quality \(DQ\) has long been recognized as a critical prerequisite for effective decision\-making in organizations\([Wang and Strong, 1996](https://arxiv.org/html/2609.06053#bib.bib33);[Redman, 1998](https://arxiv.org/html/2609.06053#bib.bib34);[Pipino et al\., 2002](https://arxiv.org/html/2609.06053#bib.bib35)\)\.[Haug et al\.](https://arxiv.org/html/2609.06053#bib.bib26)\([Haug et al\., 2011](https://arxiv.org/html/2609.06053#bib.bib26)\)show the impact of poor DQ on organizations that can range from incorrect decision\-making and decreased customer satisfaction to life\-threatening medication errors in healthcare\([de Andrade et al\., 2026](https://arxiv.org/html/2609.06053#bib.bib27)\)\. As organizations increasingly deploy artificial intelligence \(AI\) in practice, the impact of poor DQ becomes even more severe: AI models trained on low\-quality data show reduced performance\([Mori et al\., 2023](https://arxiv.org/html/2609.06053#bib.bib28);[Mohammed et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib29)\)\. Beyond these technical considerations, DQ has also become a regulatory obligation: the EU AI Act mandates that training, validation, and test data of high\-risk AI systems must be*relevant*,*representative*,*free of errors*, and*complete*\([European Parliament and Council of the European Union, 2024](https://arxiv.org/html/2609.06053#bib.bib39)\)\. Consequently, even in the era of AI and large language models \(LLMs\), detecting and correcting data errors in enterprise settings remains an essential requirement\. Machine learning \(ML\)\-based approaches to automate DQ tasks, such as error detection and correction, with minimal expert input have been proposed in research\([Mahdavi et al\., 2019](https://arxiv.org/html/2609.06053#bib.bib19);[Rekatsinas et al\., 2017](https://arxiv.org/html/2609.06053#bib.bib8), e\.g\.,\)\. However, these tools lack vendor support and system integration capabilities that are often required by large organizations, and are hence rarely deployed in enterprise settings\. More recently, also LLMs have been applied directly to detect errors in data\([Glock et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib20);[Chandru et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib40);[Narayan et al\., 2022](https://arxiv.org/html/2609.06053#bib.bib42);[Bodensohn et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib41);[Yang et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib44)\)\. While effective especially for semantic and context\-dependent error types, these approaches require passing the data through the LLM, which is computationally expensive at the scale of enterprise data and typically conflicts with strict data protection requirements, such as the European General Data Protection Regulation \(GDPR\)\([European Parliament and Council of the European Union, 2016](https://arxiv.org/html/2609.06053#bib.bib43)\), when the data contains personal information\. Moreover,[Bodensohn et al\.](https://arxiv.org/html/2609.06053#bib.bib41)\([Bodensohn et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib41)\)show that the accuracy of LLMs declines sharply on real\-world enterprise data, e\.g\., due to large table sizes and missing internal knowledge\. Consequently, organizations predominantly rely on enterprise DQ tools\([Ehrlinger and Wöß, 2022](https://arxiv.org/html/2609.06053#bib.bib18);[Chien et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib30)\), which are largely rule\-based solutions with some ML and AI support\. These tools are computationally efficient, produce deterministic and interpretable results, and are well\-established in production\. While some enterprise DQ tools recently integrate LLMs to translate natural\-language descriptions into rules\([Rehberger et al\., 2026](https://arxiv.org/html/2609.06053#bib.bib38)\), the rule creation itself remains a manual process: a data steward still needs to know and specify which rules are required\. ##### Problem statement\. However, a rule\-based DQ tool is only as good as its rule set\.[Figure1](https://arxiv.org/html/2609.06053#S0.F1)illustrates this setting: incoming data is validated against an existing rule setR=\{r1,…,rm\}R=\\\{r\_\{1\},\\ldots,r\_\{m\}\\\}\. Data errors that are covered by a rule are detected and can be handled accordingly\. Data errors that are not covered by any rule, however, pass the validation unnoticed\. As a consequence, the data is classified as clean and propagates into downstream processes, such as ETL pipelines or master data management, where the undetected errors cause processing failures\. Consider our industry use case at the Austrian Post: a missing street number or a semantically hard error, such as the confusion of the Austrian cities Lienz and Linz, can lead to failed deliveries\. Such failures are particularly costly because they surface late and reoccur until the rule set is extended accordingly\. Since manually defining and maintaining DQ rules is complex and time\-consuming\([Chiang and Miller, 2008](https://arxiv.org/html/2609.06053#bib.bib10);[Ehrlinger and Wöß, 2022](https://arxiv.org/html/2609.06053#bib.bib18);[Taleb et al\., 2021](https://arxiv.org/html/2609.06053#bib.bib17)\), closing these gaps in the rule set remains an open problem in practice\. ##### Our approach\. In this paper, we solve this problem withLeDQeR, a LLM\-based framework that iteratively extends the rule set of an enterprise DQ tool whenever an undetected error is identified \(cf\.[Figure1](https://arxiv.org/html/2609.06053#S0.F1)\): the dirty tuple that caused a downstream failure is passed toLeDQeR, which generates DQ rules that detect the error such that the same error type cannot pass the validation unnoticed again\.LeDQeRfollows a generate\-filter approach\. In the*generate*step, a LLM automatically generates candidate DQ rules from a dirty tuple and a given DQ rule format\. In contrast to traditional rule\-learning methods\([Chiang and Miller, 2008](https://arxiv.org/html/2609.06053#bib.bib10);[Yeh and Puri, 2010](https://arxiv.org/html/2609.06053#bib.bib9)\), which struggle to capture semantic context and frequently require labeled data, LLMs leverage pre\-trained knowledge and semantic reasoning capabilities to generate DQ rules in code\-like syntax \(e\.g\., SQL or Python\) from minimal input\([Yadav and Mondal, 2025](https://arxiv.org/html/2609.06053#bib.bib21);[R et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib22)\)\. In the*filter*step, the candidate rules are progressively refined with respect to \(i\) executability, \(ii\) correctness, \(iii\) generalizability, and \(iv\) redundancy, producing a minimal and precise rule set that is ready for operational deployment\. ##### Contributions In summary, this paper makes the following contributions to automate DQ rule generation in enterprise settings: - •We formalize the problem of DQ rule generation and introduce a generalizable generate\-filter framework\. - •We proposeLeDQeR, a LLM\-based implementation of the framework that generates DQ rules from observed dirty tuples and incrementally extends an existing rule set\. - •We introduce a suite of four filter techniques that ensure the \(i\) executability, \(ii\) correctness, \(iii\) generalizability, and \(iv\) redundancy of the new rules\. - •We comprehensively evaluateLeDQeRon five datasets from diverse domains with up to nine different error types\. ##### Results Our evaluation shows thatLeDQeRenables the generation of executable rules that generalize well to unseen data without producing false positives\. In particular, our experiments suggest that LLMs are promising for generating DQ rules and show the effectiveness of our filters\. We also observed that the ability of the generated rules to detect a dirty tuple is influenced by the error type and the amount of error context provided in the prompt\. ##### Outline We first introduce necessary background on data errors and rule\-based DQ tools in[Section2](https://arxiv.org/html/2609.06053#S2), where we also present the running example used throughout this paper\. In[Section3](https://arxiv.org/html/2609.06053#S3), we first formalize the problem of DQ rule generation and subsequently proposeLeDQeR, an LLM\-based framework for DQ rule generation that follows our generate\-filter formalization\. We evaluateLeDQeRin[Section4](https://arxiv.org/html/2609.06053#S4)and discuss related work in[Section5](https://arxiv.org/html/2609.06053#S5)\.[Section6](https://arxiv.org/html/2609.06053#S6)concludes this paper with an outlook on future work\. ## 2\.Data Errors and Rule\-based DQ Measurement In this section, we introduce the necessary background on data error types in[Section2\.1](https://arxiv.org/html/2609.06053#S2.SS1)and rule\-based DQ tools in[Section2\.2](https://arxiv.org/html/2609.06053#S2.SS2)\. Throughout this section, we will also introduce a running example that is derived from our industry use case at the Austrian Post: a personal contact information \(PCI\) dataset \(cf\.[Table1](https://arxiv.org/html/2609.06053#S2.T1)\)\. This example serves as a walkthrough to explain our generate\-filter framework in[Section3](https://arxiv.org/html/2609.06053#S3)and its experimental evaluation in[Section4](https://arxiv.org/html/2609.06053#S4)\. ### 2\.1\.Data Quality and Error Types [Table1](https://arxiv.org/html/2609.06053#S2.T1)shows our running example: a subset of a personal contact information \(PCI\) dataset from our industry use case at the Austrian Post\. Each tuple describes a person and their address with the columnsStreet,FirstName,PostalCode, andCity\. The example is a reduced version of the fullPCIdataset with1414columns, which we use in our experimental evaluation \(cf\.[Section4](https://arxiv.org/html/2609.06053#S4)\)\. Data quality is per definition context\-dependent, as it is commonly defined as “fitness for use”\([Wang and Strong, 1996](https://arxiv.org/html/2609.06053#bib.bib33)\), which means that the data must be fit for the specific purpose it is used for\. Consequently, DQ measurement must reflect domain\-specific semantics and usage expectations\. ###### Example 2\.1\. The use case of our running example introduces the following requirements on the data: \(1\) street names must contain actual street names and no placeholder values such as “Unknown”, \(2\) a four\-digit numerical representation like “1220” is required for the postal code, even though a leading “A” is commonly added in Austria \(e\.g\., “A\-1220”\), and \(3\) the city solely consists of the city name, with no subareas being allowed\. Values that violate these requirements are data errors and are highlighted in color in[Table1](https://arxiv.org/html/2609.06053#S2.T1)\. Table 1\.Subset of the PCI dataset as running example\.IDStreetFirstNamePostalCodeCity1UnknownIna\-MariaA\-1220Wien2BrahmsplatzHans1040Wien, Wieden3BerggasseAnnaA\-4020Linz4FeldgasseAnna\-MariaA6020Innsbruck5UnknownMaria8010GrazType of error:■\\blacksquareEmbedded Value■\\blacksquareIncorrect Format■\\blacksquareDisguised Missing ValueAs illustrated in Example[2\.1](https://arxiv.org/html/2609.06053#S2.Thmtheorem1), data errors are observable effects of poor data quality and are inherently context\- and usage\-specific\. This has led to the development of numerous data\-error taxonomies\([Rahm and Do, 2000](https://arxiv.org/html/2609.06053#bib.bib11);[Müller and Freytag, 2005](https://arxiv.org/html/2609.06053#bib.bib12);[Kim et al\., 2003](https://arxiv.org/html/2609.06053#bib.bib13);[Oliveira et al\., 2005](https://arxiv.org/html/2609.06053#bib.bib14);[Josko et al\., 2016](https://arxiv.org/html/2609.06053#bib.bib15);[Ilyas and Naumann, 2022](https://arxiv.org/html/2609.06053#bib.bib16);[Bhadauria et al\., 2026](https://arxiv.org/html/2609.06053#bib.bib36)\)\. Based on these taxonomies and feedback from domain experts, we selected nine error types that cover a broad range of DQ issues that appear in single values or tuples for evaluating our framework: Explicit missing values \(emv\):are values explicitly set to null\. This is a common error type\([Pearson, 2006](https://arxiv.org/html/2609.06053#bib.bib25)\)and well\-known in literature\([Rahm and Do, 2000](https://arxiv.org/html/2609.06053#bib.bib11);[Oliveira et al\., 2005](https://arxiv.org/html/2609.06053#bib.bib14);[Kim et al\., 2003](https://arxiv.org/html/2609.06053#bib.bib13);[Müller and Freytag, 2005](https://arxiv.org/html/2609.06053#bib.bib12);[Ilyas and Naumann, 2022](https://arxiv.org/html/2609.06053#bib.bib16);[Jung et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib37);[Bhadauria et al\., 2026](https://arxiv.org/html/2609.06053#bib.bib36)\)\. Disguised missing values \(dmv\):are semantically missing \(like emv\) but contain placeholder values\. For example, John Doe asFirstNameandLastNameor 01\.01\.1900 asDateOfBirth\. Detecting this error type is challenging because they might be syntactically valid but semantically incorrect\([Pearson, 2006](https://arxiv.org/html/2609.06053#bib.bib25);[Bhadauria et al\., 2026](https://arxiv.org/html/2609.06053#bib.bib36)\)\. Contradictions \(con\):are mutually inconsistent values among different values of a tuple\([Müller and Freytag, 2005](https://arxiv.org/html/2609.06053#bib.bib12)\)\. For example, both the postal code and the city may be syntactically correct, but do not match the same entity\. Even though literature sometimes describes such errors under the umbrella of “functional dependency violations”\([Oliveira et al\., 2005](https://arxiv.org/html/2609.06053#bib.bib14);[Josko et al\., 2016](https://arxiv.org/html/2609.06053#bib.bib15)\), practical scenarios can be more complex\. For example, no functional dependency holds between city and postal code\. In other words, one postal code can refer to multiple cities \(e\.g\., 4232 is the postal code for99municipalities in Austria\), whereas one city can have multiple postal codes \(e\.g\., Vienna has2323postal codes\)\. Therefore, we use the more general term “contradiction” introduced by[Müller and Freytag](https://arxiv.org/html/2609.06053#bib.bib12)\([Müller and Freytag, 2005](https://arxiv.org/html/2609.06053#bib.bib12)\)in this paper\. Misfielded values \(mfv\):are correct values that are switched and therefore end up in the wrong field/column\. For example, a value of aFirstNameis stored in the respectiveLastNameof the same tuple and vice versa\. This error type has already been described by\([Kim et al\., 2003](https://arxiv.org/html/2609.06053#bib.bib13);[Rahm and Do, 2000](https://arxiv.org/html/2609.06053#bib.bib11);[Jung et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib37);[Bhadauria et al\., 2026](https://arxiv.org/html/2609.06053#bib.bib36)\)\. Embedded values \(ebv\):are values that contain additional, unwanted information\([Oliveira et al\., 2005](https://arxiv.org/html/2609.06053#bib.bib14);[Rahm and Do, 2000](https://arxiv.org/html/2609.06053#bib.bib11);[Bhadauria et al\., 2026](https://arxiv.org/html/2609.06053#bib.bib36)\)\. For example, the postal code “1220, Donaustadt” also contains the district name\. Spelling mistakes \(spm\):are incorrect values resulting from added, deleted, or exchanged characters/numbers\([Rahm and Do, 2000](https://arxiv.org/html/2609.06053#bib.bib11);[Oliveira et al\., 2005](https://arxiv.org/html/2609.06053#bib.bib14);[Kim et al\., 2003](https://arxiv.org/html/2609.06053#bib.bib13);[Josko et al\., 2016](https://arxiv.org/html/2609.06053#bib.bib15);[Jung et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib37);[Bhadauria et al\., 2026](https://arxiv.org/html/2609.06053#bib.bib36)\)\. Common examples include typos, such as missing an ”n” in ”Viena”\. Domain violations \(dov\):refer to values that are outside of an expected range or violates business constraints\([Oliveira et al\., 2005](https://arxiv.org/html/2609.06053#bib.bib14);[Josko et al\., 2016](https://arxiv.org/html/2609.06053#bib.bib15);[Müller and Freytag, 2005](https://arxiv.org/html/2609.06053#bib.bib12)\)\. For example, a negativePhoneNumberor the violation of a business rule that enforces the country column to contain the full name of the country\. Incorrect formats \(ifo\):may appear for various columns and data formats\. For example, the date might follow different standards \(’MM/DD/YYYY’ vs\. ’DD/MM/YYYY’\)\. This error type is listed as a syntax violation in\([Oliveira et al\., 2005](https://arxiv.org/html/2609.06053#bib.bib14);[Bhadauria et al\., 2026](https://arxiv.org/html/2609.06053#bib.bib36)\), as a domain format error in\([Müller and Freytag, 2005](https://arxiv.org/html/2609.06053#bib.bib12)\), and as a format error in\([Jung et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib37)\)\. Incorrect character sets \(ics\):are used in the data\. For example, the Cyrillic character set was used to write the address instead of the Latin character set\. This error type was also introduced by[Kim et al\.](https://arxiv.org/html/2609.06053#bib.bib13)\([Kim et al\., 2003](https://arxiv.org/html/2609.06053#bib.bib13)\), who focus on wrong encodings \(e\.g\., ASCII vs\. EBCDIC\) and also mentioned as incorrect encoding in\([Jung et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib37);[Bhadauria et al\., 2026](https://arxiv.org/html/2609.06053#bib.bib36)\)\. While not exhaustive, the selected error types capture the major classes of DQ issues described in literature and encountered by practitioners\. ###### Example 2\.2\. Each tuple in[Table1](https://arxiv.org/html/2609.06053#S2.T1)is identifiable via column ID and contains at least one dirty value, which can be classified as one of the error types: the errors inStreetare of type dmv, as two tuples contain the valueUnknown, which indicates a missing street name without actually being a missing value\.FirstNamehas no errors\. Three values inPostalCodeare of error type ifo, since the correct format is the four\-digit postal code as shown in tuples 2 and 5\. Tuple 2 violates the requirement that the city is limited to the city name and is hence of type ebv\. ### 2\.2\.Rule\-based Data Quality Measurement Rule\-based DQ measurement\([Loshin, 2002](https://arxiv.org/html/2609.06053#bib.bib32)\)is a foundational approach to measuring data quality, relying on an explicitly defined rule set that captures the expected data properties\. These rules formalize domain knowledge and business constraints \(e\.g\., allowable value ranges, referential integrity conditions, or pattern\-based formats\) and are systematically evaluated against datasets to detect violations\. Several practical data quality tools and frameworks support the rule\-based approach\([Ehrlinger and Wöß, 2022](https://arxiv.org/html/2609.06053#bib.bib18)\)\. Prominent open\-source systems include Great Expectations111[https://github\.com/great\-expectations/great\_expectations](https://github.com/great-expectations/great_expectations)\(GX\), Amazon Deequ222[https://github\.com/awslabs/deequ](https://github.com/awslabs/deequ), and Apache Griffin333[https://griffin\.apache\.org/](https://griffin.apache.org/)\. In the commercial domain, platforms such as Informatica444[https://www\.informatica\.com/products/data\-quality\.html](https://www.informatica.com/products/data-quality.html), Talend Data Quality555[https://www\.talend\.com/products/data\-quality/](https://www.talend.com/products/data-quality/), and Ataccama666[https://www\.ataccama\.com/platform/data\-quality](https://www.ataccama.com/platform/data-quality)offer support for rule\-based data quality evaluation\. Despite differences in implementation details, these tools share a common architecture in which explicitly defined rules are evaluated against data to identify violations that subsequently provide data quality indicators\. In the remainder of this paper, we focus on GX as a representative rule\-based data quality framework due to its widespread adoption and open\-source nature\. A list of expectations \(i\.e\., rules\) in GX syntax is shown in Example[2\.3](https://arxiv.org/html/2609.06053#S2.Thmtheorem3)\. ###### Example 2\.3\. Based on the data in Table[1](https://arxiv.org/html/2609.06053#S2.T1), a set of hypothetically generated, possibly incorrect rules is given below: 1. \(1\)gx\.expectations\.ExpectColumnValuesToNotMatchRegex\( column=’PostalCode’, regex=`r'A\-\[0\-9\]\*'`\) 2. \(2\)gx\.expectations\.ExpectColumnValuesToMatchRegex\( column=’Street’, regex=`r'^Unknown$'`\) 3. \(3\)gx\.expectations\.ExpectColumnValuesToNotMatchRegex\( column=’FirstName’, regex=`r'^\[A\-Z\]\[a\-z\\\-\]\*$'`\) 4. \(4\)gx\.expectations\.ColumnNotMatchRegex\( column=’Street’, regex=`r'\[A\-Z\]\[a\-z\]\*'`\) Each rule specifies a pattern with a regular expression, where the operation defines whether matching values are considered correct or incorrect\. Rule \(1\) states that postal codes starting with ’A\-’ are incorrect, which correctly identifies the entries marked in red in Table[1](https://arxiv.org/html/2609.06053#S2.T1)\. Rule \(2\) is an example for an incorrect rule because it expects street names to match the value “Unknown”, thereby erroneously flagging the valid street names in tuples 2–4 while accepting the actual errors\. Rule \(3\) states that first names that consist of a capitalized letter followed by lower\-case letters, possibly separated by hyphens, are incorrect, which would erroneously mark entries 2, 3, and 4 of the first name column as errors\. Rule \(4\) syntactically violates the GX format and is not executable at all\. Note that these hypothetical rules are intentionally imperfect and illustrate typical problems that occur during rule generation\. In[Section3](https://arxiv.org/html/2609.06053#S3), we show how our filters remove such rules\. While rule\-based DQ frameworks provide deterministic and transparent data validation, there are some notable limitations: First, data stewards often lack complete insight into newly emerging or fast evolving data domains, which makes it difficult to define comprehensive rule sets upfront and maintaining hundreds of rules over time\. Second, defining rules for semantic error types \(e\.g\., “empty” refers to a missing value\) is more difficult than for syntactic error types \(e\.g\., Austrian postal codes must consists of four digits\)\. In summary, the manual generation of suitable DQ rules remains a huge and practical challenge in enterprise settings\. ## 3\.LLM\-based Data Quality Rule Generation We introduce a LLM\-based approach for data quality rule generation that addresses the shortcomings of manual approaches\. In this section, we formalize the problem, introduce a general generate\-filter framework for data quality rule generation, and then present our LLM\-based approach to instantiate the framework\. ### 3\.1\.Problem Formalization Data quality rule generation aims at automatically deriving a set of DQ rulesℛ=\{r1,r2,…,rm\}\\mathcal\{R\}=\\\{r\_\{1\},r\_\{2\},\\ldots,r\_\{m\}\\\}for a given dataset𝒟=\{t1,t2,…,tn\}\\mathcal\{D\}=\\\{t\_\{1\},t\_\{2\},\\ldots,t\_\{n\}\\\}ofnntuples and a rule\-based DQ tool supporting rule syntax𝒮\\mathcal\{S\}\. The quality of the generated rules depends on their executability \(i\.e\., syntactic correctness\) and their ability to identify the corresponding error types in the data\. To foster the discussion of existing\([Xie et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib24);[Schneider, 2025](https://arxiv.org/html/2609.06053#bib.bib23)\), our proposed, and future approaches, we introduce a generalized framework for data quality rule generation, called*generate\-filter framework*\(cf\.[Figure2](https://arxiv.org/html/2609.06053#S3.F2)\)\. The framework produces a rule setℛ\\mathcal\{R\}in two steps: 1. \(1\)*Rule Generation:*Given an input dataset𝒟\\mathcal\{D\}, the objective is to generate an initial set of candidate DQ rules𝒞\\mathcal\{C\}\. 2. \(2\)*Rule Filtering:*Since the generated rule candidates𝒞\\mathcal\{C\}may include erroneous, redundant, or overlapping rules, a subsequent filtering phase is applied to identify and remove low\-quality rules\. Thereby, the overall correctness and coverage of the resulting rule setℛ\\mathcal\{R\}is improved\. GenerateFilter𝒟\\mathcal\{D\}𝒞\\mathcal\{C\}ℛ\\mathcal\{R\}Figure 2\.Generate\-filter framework for DQ rule generation\.###### Example 3\.1\. Consider dataset𝒟\\mathcal\{D\}from Example[2\.1](https://arxiv.org/html/2609.06053#S2.Thmtheorem1)\. In the generate step, a set of\|𝒞\|=4\|\\mathcal\{C\}\|=4candidate rules for the DQ tool GX is generated \(cf\. Example[2\.3](https://arxiv.org/html/2609.06053#S2.Thmtheorem3)\)\. Assume that in the filtering step, the syntactical correctness of the rules is verified\. Consequently,ℛ\\mathcal\{R\}contains three rules \(1\-3\), while rule \(4\) is discarded because it does not conform to GX’s syntax\. Figure 3\.Architecture ofLeDQeR: an LLM generates candidate rules𝒞\\mathcal\{C\}, which are progressively refined through four filters𝒞⊇𝒞′⊇𝒞′′⊇𝒞′′′\\mathcal\{C\}\\supseteq\\mathcal\{C\}^\{\\prime\}\\supseteq\\mathcal\{C\}^\{\\prime\\prime\}\\supseteq\\mathcal\{C\}^\{\\prime\\prime\\prime\}until a final rule setℛ\\mathcal\{R\}is returned\. ### 3\.2\.LeDQeR: An LLM\-based Generate\-Filter Framework to Generate DQ Rules We introduceLeDQeR, a LLM\-based instantiation of the generate\-filter framework\.[Figure3](https://arxiv.org/html/2609.06053#S3.F3)shows the architecture ofLeDQeR: First, we generate a candidate DQ rule set using a LLM\-based technique\. Second, we propose a set of four complementary filters that progressively refine the candidate rule set𝒞⊇𝒞′⊇𝒞′′⊇𝒞′′′⊇ℛ\\mathcal\{C\}\\supseteq\\mathcal\{C\}^\{\\prime\}\\supseteq\\mathcal\{C\}^\{\\prime\\prime\}\\supseteq\\mathcal\{C\}^\{\\prime\\prime\\prime\}\\supseteq\\mathcal\{R\}\. Algorithm 1DQ Rule Generation1: 𝒟\\mathcal\{D\}\(Dirty Dataset\), ℳ\\mathcal\{M\}\(LLM Model\), 𝒫\\mathcal\{P\}\(Prompt Template\), 𝒮\\mathcal\{S\}\(Rule Syntax\), DescDesc\(Dataset Description\) 2: 𝒞\\mathcal\{C\}\(Candidate set of DQ Rules\) 3:proceduregenerateRules\( 𝒟,ℳ,𝒫,𝒮,Desc\\mathcal\{D\},\\mathcal\{M\},\\mathcal\{P\},\\mathcal\{S\},Desc\) 4: 𝒞←∅\\mathcal\{C\}\\leftarrow\\emptyset 5:for 𝐭i∈𝒟\\mathbf\{t\}\_\{i\}\\in\\mathcal\{D\}do 6: prompt←constructPrompt\(𝒫,𝐭i,𝒮,Desc\)prompt\\leftarrow\\text\{constructPrompt\}\(\\mathcal\{P\},\\mathbf\{t\}\_\{i\},\\mathcal\{S\},Desc\) 7: response←queryLLM\(prompt,ℳ\)response\\leftarrow\\text\{queryLLM\}\(prompt,\\mathcal\{M\}\) 8: 𝒞i←parseResponse\(response\)\\mathcal\{C\}\_\{i\}\\leftarrow\\text\{parseResponse\}\(response\) 9: 𝒞←𝒞∪𝒞i\\mathcal\{C\}\\leftarrow\\mathcal\{C\}\\cup\\mathcal\{C\}\_\{i\} 10:endfor 11:return 𝒞\\mathcal\{C\} 12:endprocedure #### 3\.2\.1\.LLM\-based DQ Rule Generation Instead of defining the set of DQ rules by hand, as it is typically expected by state\-of\-the\-art DQ tools \(cf\.[Section 2](https://arxiv.org/html/2609.06053#S2)\), we present an automated LLM\-based approach to generate DQ rules\. Algorithm[1](https://arxiv.org/html/2609.06053#alg1)describes the generation process, which requires the following inputs: a \(dirty\) dataset𝒟\\mathcal\{D\}, a LLM modelℳ\\mathcal\{M\}, a prompt template𝒫\\mathcal\{P\}, a DQ rule syntax𝒮\\mathcal\{S\}, and a textual descriptionDescDescof the semantic context of the dataset \(e\.g\., “The data contain information about the scheduled and actual departure and arrival times of airplanes for a flight dataset\.”\)\. Each tuple𝐭i∈𝒟\\mathbf\{t\}\_\{i\}\\in\\mathcal\{D\}contains at least one error\. For each tuple𝐭i\\mathbf\{t\}\_\{i\}, a prompt template𝒫\\mathcal\{P\}is populated with the dirty record, a guideline for the DQ rule syntax𝒮\\mathcal\{S\}, and an overall dataset description\. For this prompt, the LLM generates a set of candidate rules𝒞i\\mathcal\{C\}\_\{i\}for tuple𝐭i\\mathbf\{t\}\_\{i\}\. The number of generated rules\|𝒞i\|\|\\mathcal\{C\}\_\{i\}\|may differ from the actual error count of the dirty record, as the LLM is unaware of the exact number of errors and may generate multiple rules or none for a specific error\. The global candidate set𝒞\\mathcal\{C\}is defined as the union of all individual rule sets𝒞i\\mathcal\{C\}\_\{i\}generated for each tuple𝐭i\\mathbf\{t\}\_\{i\}with1≤i≤n1\\leq i\\leq n, such that𝒞=⋃i=1n𝒞i\\mathcal\{C\}=\\bigcup\_\{i=1\}^\{n\}\\mathcal\{C\}\_\{i\}\. ##### Prompt construction\. The prompt is a central part of the generation process and consists of the following four sections: the prompt template𝒫\\mathcal\{P\}\(containing the initial framing and tasks\), the input data𝐭i\\mathbf\{t\}\_\{i\}, the rule syntax𝒮\\mathcal\{S\}, and the dataset descriptionDescDesc\. For each tuple𝐭i\\mathbf\{t\}\_\{i\}, the functionconstructPrompt\(𝒫,𝐭i,𝒮,Desc\)\\text\{constructPrompt\}\(\\mathcal\{P\},\\mathbf\{t\}\_\{i\},\\mathcal\{S\},Desc\)instantiates the prompt template into a concrete prompt \(cf\. Algorithm[1](https://arxiv.org/html/2609.06053#alg1), line 4\)\. The*initial framing*of the prompt contains overall information such as the LLM’s role \(e\.g\., “You are a GX rule expert\.”\) and the overall task \(e\.g\., “Generate GX rules to prevent an error in the data sample\)\. We support and evaluate three variants of providing the*input data*𝐭i\\mathbf\{t\}\_\{i\}with increasing difficulty: 1. \(1\)*Dirty & Clean:*The LLM is provided with clean and dirty data\. While this scenario is uncommon in practice, it provides the most information to the LLM and serves as an upper bound for our evaluation\. 2. \(2\)*Dirty & Type:*In addition to dirty data, the LLM also receives information about the error types present in the columns of a record\. This scenario is applicable in case the applied error detection approach returns additional error type information rather than a binary assessment\. 3. \(3\)*Just Dirty:*Only dirty data is provided to the LLM\. This variant is the most realistic scenario in practice, where no additional data and ground truth are available\. Compared to the other input data variants, the LLM needs to detect the error without additional information before generating a rule\. The*rule*section specifies the types of rules the LLM should generate, including information about the DQ tool: \(1\) A clear description of the application scenario, e\.g\., focus on generating a single rule set without additional information\. \(2\) Details on the targeted DQ tool, e\.g\., GX version numbers and rule syntax𝒮\\mathcal\{S\}examples\. This includes information on generating custom rule functions as well as a list of existing GX rule functions with \(e\.g\., ”ExpectColumnValuesToNotMatchRegex\(column: str, regex: str\)”\) and without parameters \(e\.g\., ”ExpectColumnValuesToNotMatchRegex”\)\. Otherwise, there is a high chance to generate incorrect rules\. In the last prompt section, we guide the LLM by defining four consecutive sub\-tasks to achieve the overall goal: 1. \(1\)detect the error in the provided data, 2. \(2\)generate rules to detect the identified errors, 3. \(3\)check if the rules are general and fit the given syntax, and 4. \(4\)return the rules in a specific format \(in our case JSON\)\. #### 3\.2\.2\.Filtering pipeline After generating candidate rules𝒞\\mathcal\{C\}in the first step of the framework, filters are applied to improve the overall quality of𝒞\\mathcal\{C\}\. InLeDQeR, we implement a set of four filters that are sequently applied to progressively refine the rules and produce a final rule set𝒞⊇𝒞′⊇𝒞′′⊇𝒞′′′⊇ℛ\\mathcal\{C\}\\supseteq\\mathcal\{C\}^\{\\prime\}\\supseteq\\mathcal\{C\}^\{\\prime\\prime\}\\supseteq\\mathcal\{C\}^\{\\prime\\prime\\prime\}\\supseteq\\mathcal\{R\}\. ##### Syntactic filter The syntactic filter verifies whether the generated rules𝒞\\mathcal\{C\}are executable and match the given DQ rule syntax𝒮\\mathcal\{S\}\. The result is a reduced candidate rule set𝒞′⊆𝒞\\mathcal\{C\}^\{\\prime\}\\subseteq\\mathcal\{C\}that contains only syntactically correct \(i\.e\., executable\) rules\. No additional information, such as labeled data, is required for this filter\. ###### Example 3\.2\. For Example[2\.3](https://arxiv.org/html/2609.06053#S2.Thmtheorem3), this means that the syntactical correctness rate is0\.750\.75since the fourth rule contains an operation that is not supported in GX \(ColumnNotMatchRegex\) and hence only33out of44rules are syntactically correct\. ##### Correctness filter The correctness filter removes rules that fail to filter the tuple from which they were generated\. Therefore, we re\-apply and verify each generated rule𝐫∈𝒞′\\mathbf\{r\}\\in\\mathcal\{C\}^\{\\prime\}on the tuple that has triggered its generation\. A rule either recognizes the error \(*true positive TP*\), fails to recognize the error \(*false negative FN*\), or mistakenly detects an error in a clean value \(*false positive FP*\)\. The filter removes all rules that do not lead to a TP from𝒞′\\mathcal\{C\}^\{\\prime\}leading to a reduced set of rules𝒞′′⊆𝒞′\\mathcal\{C\}^\{\\prime\\prime\}\\subseteq\\mathcal\{C\}^\{\\prime\}\. ###### Example 3\.3\. Applying this filter to the rules in our Example[2\.3](https://arxiv.org/html/2609.06053#S2.Thmtheorem3)with the data from[Table1](https://arxiv.org/html/2609.06053#S2.T1)yields an F1 score of 0\.5\. Specifically, the first rule correctly detects that thePostalCodeis wrong \(TP\)\. TheStreetis wrong but not detected by any rule \(FN\)\. Rule three incorrectly flags the cleanFirstNameas wrong, resulting in a false positive \(FP\)\. ##### Coverage filter To identify rules𝐫∈𝒞′′\\mathbf\{r\}\\in\\mathcal\{C\}^\{\\prime\\prime\}that do not generalize well beyond the original error instance within the observed input tuple𝐭i\\mathbf\{t\}\_\{i\}, the coverage filter tests each rule against all tuples that share the same error type in the training dataset𝒟\\mathcal\{D\}\. Rules are removed if they achieve a recall above a predefined threshold across these variants, resulting in𝒞′′′⊆𝒞′′\\mathcal\{C\}^\{\\prime\\prime\\prime\}\\subseteq\\mathcal\{C\}^\{\\prime\\prime\}\. While precision was addressed in correctness filter, the coverage filter focuses exclusively on recall\. A high recall indicates strong generalizability across various variants of the same error type\. However, a low recall does not necessarily imply that a rule is ineffective, as some errors may require highly specific rules to avoid confusion with correct values\. Unfortunately, such rules may exhibit low recall and could be filtered out depending on the chosen threshold\. In the current implementation ofLeDQeR, we do not set a minimum recall threshold to prioritize maximum coverage and avoid the risk of filtering highly specific rules\. ###### Example 3\.4\. The remaining rule inC′′C^\{\\prime\\prime\}is applied on tuples 1, 3 and 4 \([Table1](https://arxiv.org/html/2609.06053#S2.T1)\)\. Tuple 2 has a different error and thus will not be used\. Tuple 5 shares one error column and error type with tuple 1\. However, the LLM did not generate a rule to detect this error successfully in tuple 1, thus the tuple is not selected\. As only rule 1 passed the two previous filters, we consider only this rule for the coverage filter\. The recall of 0\.66 is calculated from - •TP \(2\): tuple 1 \(PostalCode\), tuple 3 \(PostalCode\) - •FP \(0\): - •FN \(1\): tuple 4 \(PostalCode\) This means the generalizability of the rule is not optimal, as it captures only a subset of error variants\. While it successfully detects the dirty tuple 3 \(with thePostalCodeformat "A\-<4 digits\>"\), it fails to identify another dirty tuple \(with thePostalCodeformat "A<4 digits\>"\)\. As we do not set a minimum recall threshold,C′′′C^\{\\prime\\prime\\prime\}consists of rule 1\. ##### Redundancy filter The redundancy filter removes duplicated and overlapping rules by using a greedy covering algorithm to reduce the learned rules to a minimal rule set that collectively covers all errors of an error type\([Fürnkranz, 1999](https://arxiv.org/html/2609.06053#bib.bib31)\)\. We therefore use a separate test dataset𝒟test\\mathcal\{D\}\_\{test\}to estimate precision and recall\. In practice,𝒟test\\mathcal\{D\}\_\{test\}can be a curated dataset used to validate rule performance\. While our evaluation employs large\-scale datasets to demonstrate scalability, smaller curated sets are sufficient for practical application\. We iteratively select rules with the highest recall among those achieving precision11, mark their covered tuples, and repeat until all tuples are covered, resulting in the final rule setℛ⊆𝒞′′′\\mathcal\{R\}\\subseteq\\mathcal\{C\}^\{\\prime\\prime\\prime\}\. Since false positives are costly in practice, precision is the primary filter metric\. Unlike the previous filters, the redundancy filter removes overlapping rules in the𝒞′′′\\mathcal\{C\}^\{\\prime\\prime\\prime\}across an unseen𝒟test\\mathcal\{D\}\_\{test\}\. Consequently, its behavior cannot be meaningfully demonstrated using the small\-scale[Table1](https://arxiv.org/html/2609.06053#S2.T1)\. We therefore defer its illustration to the experimental evaluation in[Section4\.5](https://arxiv.org/html/2609.06053#S4.SS5)\. ## 4\.Experiments In this section, we experimentally evaluateLeDQeRby analyzing each step of the generate\-filter framework: the rule generation and the syntactic filter in[Section4\.2](https://arxiv.org/html/2609.06053#S4.SS2), the correctness filter in[Section4\.3](https://arxiv.org/html/2609.06053#S4.SS3), the coverage filter in[Section4\.4](https://arxiv.org/html/2609.06053#S4.SS4), and the redundancy filter in[Section4\.5](https://arxiv.org/html/2609.06053#S4.SS5)\. The experimental setup is described in[Section4\.1](https://arxiv.org/html/2609.06053#S4.SS1)\. As each step targets a different aspect of rule quality, the datasets, evaluation metrics, and number of rules vary by experiment\. ### 4\.1\.Experimental Setup All experiments were implemented in Python and are available on GitHub777https://github\.com/Anna\-Christina\-Glock/dq\_rule\_generation\. We ran the experiments on our high\-performance computing platform, which features a single AMD EPYC 7643 \(48\-core, 2\.3 GHz\) CPU, four NVIDIA A100 GPUs \(each with 160GB high\-bandwidth memory\), and 1TB RAM\. #### 4\.1\.1\.Data quality tool In our experiments, we create rules for the open source DQ tool Great Expectations\(GX\), which enables users to create rules \(“expectations” in GX terminology\) to check whether the validated data contains errors \(see Example[2\.3](https://arxiv.org/html/2609.06053#S2.Thmtheorem3)for a syntax example\)\. GX provides multiple predefined rules that can be customized\. We have chosen GX in version 1\.9\.1 \(i\.e\., the latest stable version\) for our experiments, since it is freely available, based on Python, has a large and active open\-source community, as well as active partnerships with major industry companies such as Snowflake and Databricks\([Great Expectations, 2026](https://arxiv.org/html/2609.06053#bib.bib1)\)\. #### 4\.1\.2\.Datasets We used five datasets for our evaluation, summarized in[Table2](https://arxiv.org/html/2609.06053#S4.T2)\. Thebeers,flights,hospital, andtaxdatasets are adapted from Mahdavi et al\.\([Mahdavi et al\., 2019](https://arxiv.org/html/2609.06053#bib.bib19)\)\. ThePCI\(personal contact information\) dataset is synthetically generated from real Austrian addresses\([Glock et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib20)\)\. Compared to the original version in\([Glock et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib20)\), we polluted thePCIdataset with additional error types for this study888https://github\.com/Anna\-Christina\-Glock/pci\-llm\-toolkit/tree/main/data\-generator\. Table 2\.Dataset characteristics used in our experiments\.Dataset\# Columns\# Dirty Tuples\# Error TypesBeers102,4101Flights71,9043Hospital194071Tax156,0002PCI146,0009Table 3\.Syntactic filter: Number of generated rules and the corresponding percentage of executable rules across three LLMs, five𝒟\\mathcal\{D\}, and three input variants \(Dirty & Clean, Dirty & Type, Just Dirty\), and two𝒮\\mathcal\{S\}variations \(with vs\. without parameter information\)\.GLM\-4\.7Gemma\-4Qwen3\-Coderno ParameterParameterno ParameterParameterno ParameterParameterDirty & CleanBeersNo\. Rules2,5342,5692,9843,1193,2233,137% Executable0\.280\.840\.991\.000\.980\.94FlightsNo\. Rules2,0022,7483,1549027,8607,804% Executable0\.450\.770\.950\.930\.980\.94HospitalNo\. Rules468502476484530621% Executable0\.500\.800\.930\.950\.930\.82PCINo\. Rules6,5396,3216,4676,83511,39611,379% Executable0\.370\.780\.860\.930\.900\.85TaxNo\. Rules14,81815,0351932,02714,25113,307% Executable0\.480\.850\.880\.970\.940\.80Dirty & TypeBeersNo\. Rules1,9772,0712,2672,3982,4842,475% Executable0\.320\.851\.001\.000\.980\.95FlightsNo\. Rules1,6371,6941,9652,0152,5482,836% Executable0\.430\.820\.991\.000\.970\.90HospitalNo\. Rules492543571558481429% Executable0\.430\.810\.980\.990\.920\.90PCINo\. Rules6,8357,0334,0764,24116,78516,691% Executable0\.350\.800\.910\.940\.980\.91TaxNo\. Rules4,6604,8583,4281,6445,6265,472% Executable0\.390\.811\.001\.000\.990\.98Just DirtyBeersNo\. Rules2,4512,4664,0524,1152,6862,602% Executable0\.340\.760\.680\.830\.890\.58FlightsNo\. Rules2,5032,4944,8194,8928,5818,033% Executable0\.460\.780\.980\.990\.970\.92HospitalNo\. Rules7467461,4891,4581,0711,048% Executable0\.480\.870\.981\.000\.900\.84PCINo\. Rules9,2938,8878,9049,73927,74526,222% Executable0\.330\.780\.810\.850\.990\.88TaxNo\. Rules8,5359,2558,50010,05916,01016,470% Executable0\.440\.850\.790\.930\.930\.84 #### 4\.1\.3\.LLMs To interface with the three selected local LLMs, we utilized the OpenAI Python client999https://github\.com/openai/openai\-python\. These models were selected based on their rankings in the Chatbot Arena LLM Leaderboard101010https://lmarena\.ai/?leaderboardand their compatibility with our computing infrastructure: - •GLM\-4\.7\-Flash \(cyankiwi/GLM\-4\.7\-Flash\-AWQ\-4bit\)111111https://huggingface\.co/cyankiwi/GLM\-4\.7\-Flash\-AWQ\-4bit:a lightweight general\-purpose model chosen for its ability to run on weaker hardware, serving as a lightweight baseline\. - •Qwen3\-Coder \(cpatonn/Qwen3\-Coder\-30B\-A3B\-Instruct\-AWQ\-4bit\)121212https://huggingface\.co/cyankiwi/Qwen3\-Coder\-30B\-A3B\-Instruct\-AWQ\-4bit:a code\-specialized model that may benefit the generation of syntactically correct rules\. - •Gemma\-4 \(RedHatAI/gemma\-4\-31B\-it\-FP8\-Dynamic\)131313https://huggingface\.co/RedHatAI/gemma\-4\-31B\-it\-FP8\-Dynamic:a high\-performance general\-purpose model serving as a strong baseline for comparison\. This selection allows us to compare a lightweight model, a code\-specialized model, and a high\-capacity general\-purpose model\. ### 4\.2\.Evaluation of the Syntactic Filter [Table3](https://arxiv.org/html/2609.06053#S4.T3)shows the number of rules generated by the LLMs for each dataset, LLM, input variant, and𝒮\\mathcal\{S\}variant \(i\.e\. with/without parameter\)\. Interestingly, the models typically generated more than one rule per tuple\. This contradicts our initial expectation of a single rule per tuple, since most tuples contain only one error affecting a single value\. This suggests that the LLM generates additional rules to detect other potential errors, which may improve the generalizability ofℛ\\mathcal\{R\}\. However, this also increases the likelihood of overlapping and duplicated rules, further justifying the importance of the filtering pipeline\. To evaluate the syntactic filter, we use the fraction\|C′\|\|C\|\\frac\{\|C^\{\\prime\}\|\}\{\|C\|\}of syntactically correct rules as the metric\. The findings in[Table3](https://arxiv.org/html/2609.06053#S4.T3)indicate that the LLMs are reliable in generating executable rules, especially when𝒮\\mathcal\{S\}with parameter information is provided\. Depending on the model and the dataset, 87\.13% to 100% of the rules are executable in that case\. This reliability was consistent across different datasets and various model, although the coding\-specialized Qwen3\-Coder \(mean executability = 91%\) and the general\-purpose Gemma\-4 \(mean executability = 93%\) outperformed the lightweight GLM\-4\.7\-Flash \(mean executability = 61%\)\. ### 4\.3\.Evaluation of the Correctness Filter Figure 4\.Correctness filter: Aggregated F1\-Score per error type across the three LLMs, three input variants, and five datasets\.The correctness filter is applied to all rules that remain after the syntactic filter\. As mentioned in[Section3\.2\.2](https://arxiv.org/html/2609.06053#S3.SS2.SSS2.Px2), we use F1 to evaluate whether the generated rules correctly identify the error in each tuple\. [Figure4](https://arxiv.org/html/2609.06053#S4.F4)presents the mean F1 per error type and per types of error information \(Just Dirty, Dirty & Type and Dirty & Clean\)\. Because the results of the rules generated with and without parameter information are very similar \(max std: 0\.087\), we will present the mean results for both\. Just Dirty performs worst overall, which is expected, as, in addition to creating rules, the dirty values have to be detected\. Which is especially difficult cases where business knowledge is needed, for example for the dov or the ifo where the right format might depend of organizational preferences\. The results for Dirty & Type and Dirty & Clean demonstrate that it is possible to create successful rules for these error information types as well\. The results for Dirty & Type indicate that providing information about the error type and column is enough to increase the overall performance by 0\.33 compared to Just Dirty\. Consistent with the observations from the syntactic filter, GLM\-4\.7\-Flash exhibited the lowest performance among the three models\. Interestingly, Qwen3\-Coder also underperformed relative to Gemma\-4\. This is likely due to Gemma\-4’s better general reasoning capabilities, whereas the specialized coding capabilities that benefited Qwen3\-Coder in the syntactic filter were less critical for ensuring the correctness of the rules\. Overall, the rule generation works well across all error types, with a mean F1 of 0\.66, Performance varied by error type, with the LLM struggling most with con \(F1 = 0\.4\) and achieving the highest success with ebv \(F1 = 0\.87\)\. The success with ebv, is likely due to these errors being easily detectable via regular expressions and only affecting a single value, given how they were generated\. In contrast, con involves inconsistencies across multiple columns, necessitating more complex rules\. ### 4\.4\.Evaluation of the Coverage Filter Following[Section3\.2\.2](https://arxiv.org/html/2609.06053#S3.SS2.SSS2.Px3), we use recall as the metric since false positives have been eliminated in previous stages\. A high recall indicates good generalizability across different variants of the same error type\. To prioritize maximum coverage and avoid the risk of filtering out highly specific rules \(as cautioned in[Section3\.2\.2](https://arxiv.org/html/2609.06053#S3.SS2.SSS2.Px3)\), no minimum recall threshold was enforced during this stage\. Consequently, all candidate rules were retained for the final redundancy filtering process\. Figure 5\.Coverage filter: Aggregated recall per error type across the three LLMs, three input variants, and five datasets\.[Figure5](https://arxiv.org/html/2609.06053#S4.F5)shows that regardless of the specific type of error information provided, the rules demonstrate similar generalizability over other instances of the same type, achieving a mean recall of 0\.88 \(std = 0\.16\)\. Contrary to the observations in the syntactic filter and the correctness filter, GLM\-4\.7\-Flash is no longer the lowest\-performing model in this stage\. The rules generalize similarly across all LLM models \(recall std = 0\.16\), although those generated by Qwen3\-Coder often perform the worst of the three\. An exception is the con error type, in which Gemma\-4 rules exhibit the lowest generalization\. The reasons for this may be linked to the semantic complexity of con errors, leading Gemma\-4 to produce overly specific rules to avoid specific contradictions\. Even though the generalization works for all error types, performance is notably lower for those with multiple variants, such as dov \(recall = 0\.58\), con \(recall = 0\.67\), and spm \(recall = 0\.84\)\. For example, for the column, the LLM created a rule with a regex that correctly detects values starting with a lowercase letter\. However, the rule does not generalize to all Austrian city names: cities containing umlauts \(e\.g\., “ö”, “ü”, “ä”\), hyphens, or multiple words \(e\.g\., “Hagenberg im Mühlkreis”\) are not detected, making the rule overly general but not optimal\. Figure 6\.Precision and recall of each rule measured on an unseen dataset𝒟test\\mathcal\{D\}\_\{test\}for each of the three LLMs, three data input variants, and five datasets\. The results have been ordered by F1 score\.Figure 7\.Redundancy filter: Percentage of covered clean and dirty tuples found by the rules inℛ\\mathcal\{R\}, calculated for each of the five𝒟test\\mathcal\{D\}\_\{test\}, three input variants and three LLMs with different precision settings for the rules inC′′′C^\{\\prime\\prime\\prime\}\. ### 4\.5\.Evaluation of the Redundancy Filter Before applying the redundancy filter as described in[Section3\.2\.2](https://arxiv.org/html/2609.06053#S3.SS2.SSS2.Px4), we evaluate how well each rule inC′′′C^\{\\prime\\prime\\prime\}generalizes to a new, unseen dataset𝒟test\\mathcal\{D\}\_\{test\}\. For this, recall and precision are computed per rule; see[Figure6](https://arxiv.org/html/2609.06053#S4.F6)for the results\. The figure shows that the rules inC′′′C^\{\\prime\\prime\\prime\}generalize well to unseen data, as most of the rules either have a high precision or high recall and a subset exhibits both high precision and high recall\. These results are consistent with those reported in the coverage filter, which already demonstrated that the LLM is capable of generating rules that generalize to unseen data but may miss some dirty tuples\. Taken together, the individual rule quality suggests that the rule set as a whole should perform well in combination\. We then apply the redundancy filter toC′′′C^\{\\prime\\prime\\prime\}and evaluate the resulting final rule setℛ\\mathcal\{R\}against𝒟test\\mathcal\{D\}\_\{test\}\. For thebeers,flights,hospital, andtaxdataset, we aggregate the results over 5 folds, as only one dataset is available\. For thePCIwe generated a𝒟test\\mathcal\{D\}\_\{test\}with 1,000,000 tuples 25% of them dirty ones\.[Figure7](https://arxiv.org/html/2609.06053#S4.F7)shows the total number of detected tuples split into TP \(dirty tuple\) and FP \(clean tuple\) across different minimum precision thresholds\. At a precision of 1 \(i\.e\., no false positives\), the rules achieve at least 50% coverage \(except for thetaxdataset\), depending on the error type existing in each dataset\. For thetaxdataset, a slight increase in false positives \(up to 6 FP or 0\.007%\) results in near\-complete error coverage \(99\.58%\), indicating a favorable trade\-off\. In contrast, for theflightsdataset, relaxing the precision threshold yields comparatively smaller gains of 9% in coverage while introducing a larger number of false positives \(\+6\.7%\), suggesting a more conservative threshold selection\. This supports our design choice to filter and rank rules by precision, allowing data stewards to adapt the rule set to different error characteristics and tolerance levels\. Furthermore, the redundancy filter significantly reduces the number of rules, resulting in the final rule setℛ\\mathcal\{R\}only containing between 2 and 136 rules\. Lowering the precision threshold increases coverage at the cost of introducing false positives, exposing a controllable precision–coverage trade\-off\. This trade\-off remains consistent across the different LLMs for thebeers,flights, andtaxdatasets\. However, for thehospitalandPCIdatasets, Gemma\-4 exhibits the most favorable trade\-off, followed closely by GLM\-4\.7\-Flash and Qwen3\-Coder\. ### 4\.6\.Summary of Results Our experiments show that an LLM is able to generate syntactically correct rules with executability rates ranging from 87\.13% to 100% when provided with both rule names and valid parameters\. With purposeful filtering, those rules can be reduced to a small rule set that generalizes well to unseen data\. The results indicate that even providing only the error type and column increases the overall performance \(F1 score\) by an average of 0\.33 compared to providing only dirty values\. Generally speaking, the rule generation works well across most error types, yielding a mean F1 score of 0\.66, although complex error types such as contradictions showed lower performance \(F1 = 0\.4\)\. The generalizability of the generated rules is lower for those error types with multiple variants, such as dmv \(recall=0\.58\) or con \(recall=0\.67\), compared to the mean recall of 0\.87 across all types\. The evaluation of the redundancy filter demonstrates that at a precision of 1 the reduced rule set achieves at least 50% coverage of the dirty tuples across four of the five unseen datasets\. Decreasing this precision threshold can, in some cases, improve coverage significantly\. For example, in thetaxdataset, introducing only 6 false positives increased the percentage of found dirty values by approximately 70 percentage points to reach near\-complete coverage\. This highlights a controllable trade\-off between precision and recall that a data steward can adjust based on the tolerance for false positives in the respective organization and use case\. ## 5\.Related Work We discuss related work along three lines: rule learning in the machine learning \(ML\) community, DQ rule discovery in the database community, and LLM\-based approaches to DQ rule generation\. ##### Rule learning in the ML community Learning interpretable error\-correcting rules has a long tradition in machine learning in general, and in inductive rule learning in particular\. Generalized additive models\([Hastie and Tibshirani, 1986](https://arxiv.org/html/2609.06053#bib.bib5)\)may be viewed as a family of algorithms that learn multiple layers classifiers where each layer incrementally improves upon the previous ones\. In particular boosting\-based approaches\([Freund and Schapire, 1997](https://arxiv.org/html/2609.06053#bib.bib6)\)have a clear focus on learning models that correct classification errors\. Patching\([Kauschke and Fürnkranz, 2018](https://arxiv.org/html/2609.06053#bib.bib4)\)takes a somewhat different angle in that it investigates the potential of learning interpretable rules that characterize regions of errors in an underlying immutable black\-box system\.[Gao et al\. \(2022\)](https://arxiv.org/html/2609.06053#bib.bib7)propose a system for learning rules to correct an underlying NLP system\. Collectively, these approaches provide important foundations for learning human\-understandable rule sets from examples and focus on improving predictive systems by characterizing or correcting model errors\. In contrast,LeDQeRfocuses on generating DQ rules that are executable by a DQ tool to detect errors in structured datasets\. ##### DQ rule discovery in the database community Within the database community, automatic discovery of DQ rules has been studied primarily in the context of integrity constraints and functional dependencies\.[Chiang and Miller \(2008\)](https://arxiv.org/html/2609.06053#bib.bib10)and[Yeh and Puri \(2010\)](https://arxiv.org/html/2609.06053#bib.bib9)explored the discovery of DQ rules from data, focusing on conditional functional dependencies and related dependency patterns\. These traditional approaches are largely combinatorial, and rely on discovering statistical regularities across large datasets\. As a result, they are well suited for structural constraints, but less effective to learn semantic validation rules that depend on contextual information or domain knowledge\. To overcome these limitations, recent work explores LLMs for generating higher\-level semantic constraints\. ##### LLM\-based DQ rule generation The use of LLMs to generate interpretable rules and features has received increasing attention\.[Balek et al\. \(2025\)](https://arxiv.org/html/2609.06053#bib.bib3)propose LLMs for generating tabulated features for natural language documents, and[Gottlob \(2024\)](https://arxiv.org/html/2609.06053#bib.bib2)propose the use of LLMs for automatic database curation via the Chat2Data framework\. Table 4\.Comparison ofLeDQeRwith related LLM\-based rule generation frameworks\. Xie et al\.\([Xie et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib24)\)and Schneider\([Schneider, 2025](https://arxiv.org/html/2609.06053#bib.bib23)\)\.StepXie et al\.SchneiderLeDQeRInputRequirements \+MetadataTemplate \+MetadataDirty tuplesGenerationOne\-timeOne\-timeIncrementalFilteringSyntactic✓✓a✓Correctness✗✓a✓Coverage✗✗✓Redundancy✓✓a✓OutputSQLrulesProlog\-likerulesGXrules - aRequires / involves a LLM\-as\-a\-judge\. Most closely related to our work are two approaches that also use LLMs to generate DQ rules, which we compare along the lines of our generate\-filter formalization in[Table4](https://arxiv.org/html/2609.06053#S5.T4)\.[Xie et al\. \(2025\)](https://arxiv.org/html/2609.06053#bib.bib24)follow a requirement\-driven approach in which rules are derived from natural\-language business requirements and database metadata, tailored to the medical domain of electronic health records\.[Schneider \(2025\)](https://arxiv.org/html/2609.06053#bib.bib23)proposes a template\-driven approach where templates specify the tables and operations an LLM should use to generate the rules\. Both approaches generate a complete rule set from scratch in a one\-time manner, based on what errors the LLM expects rather than on errors observed in the data\. In contrast,LeDQeRgenerates rules from observed \(i\.e\., real\) errors in the data: for each dirty tuple,LeDQeR\(1\) detects the respective error, \(2\) suggests a rule to detect this error, and \(3\) applies the filtering steps outlined in[Section3\.2](https://arxiv.org/html/2609.06053#S3.SS2)\. Rather than generating rules from scratch,LeDQeRiteratively refines an existing rule base from a small number of examples, relieving data stewards from manually specifying rules or requirements\([Schneider, 2025](https://arxiv.org/html/2609.06053#bib.bib23)\)\. Consequently, a direct quantitative comparison with[Xie et al\. \(2025\)](https://arxiv.org/html/2609.06053#bib.bib24)and[Schneider \(2025\)](https://arxiv.org/html/2609.06053#bib.bib23)is not meaningful, since the three approaches differ in their required input as well as their output \(i\.e\., rule syntax\)\. While[Xie et al\. \(2025\)](https://arxiv.org/html/2609.06053#bib.bib24)require domain\-specific business requirements and[Schneider \(2025\)](https://arxiv.org/html/2609.06053#bib.bib23)require manually specified requirements,LeDQeRcan be used solely on dirty tuples as input, which reflects a more practical scenario\. ## 6\.Conclusion In this work, we introducedLeDQeR, an LLM\-based generate\-filter framework that generates DQ rules for the enterprise DQ tool GX\. Unlike previous works\([Xie et al\., 2025](https://arxiv.org/html/2609.06053#bib.bib24);[Schneider, 2025](https://arxiv.org/html/2609.06053#bib.bib23)\), which rely on the manual specification of domain\-specific requirements or templates,LeDQeRtargets a practical setting: a data steward has identified a dirty tuple that passed the validation unnoticed and needs to extend the rule set of GX accordingly\. In the*generate*step, an LLM generates candidate rules from the dirty tuple\. In the*filter*step, four filters remove low\-quality rules\. They check the executability and correctness of individual rules, their generalizability to other errors of the same type, and remove redundant rules from the final rule set\. Our evaluation across five datasets demonstrates thatLeDQeRgenerates correct and generalizable rules from dirty tuples alone\. Providing additional information about the error, such as its type and location, or the clean tuple, further improves the rule quality\. Importantly for real\-world scenarios, the error type and location alone are sufficient to significantly improve performance\. This reduces the need to find clean data for the successful generation of DQ rules\. For less complex error types, the dirty tuple alone can still yield sufficient rule sets\. Furthermore,LeDQeRperforms consistently across different LLMs\. Interestingly, higher\-capacity general\-purpose models outperform lightweight or coding\-specialized models as error complexity increases\. This suggests that general reasoning is more critical for rule correctness than specialized coding proficiency\. Finally, we observed an occasional inverse relationship between initial rule correctness and generalizability\. This highlights the importance of the generate\-filter approach, which filters a wide range of candidate rules\. Future work\.Based on our industry partner’s feedback, we plan to systematically evaluate the ability of LLMs to generate natural\-language descriptions of the generated rules\. Such descriptions also allow non\-technical users to understand and use these rules and automate documentation processes\. We consider this direction promising since our manual reviews revealed that the LLM occasionally produced useful rule descriptions without explicit prompting\. Furthermore, we aim to explore methods for automatically detecting the error type of a given dirty tuple\. Our results already demonstrate its potential: providing the error type significantly improves the rule generation performance\. ###### Acknowledgements\. The research was supported by the Austrian ministries BMIMI, BMWET and the State of Upper Austria in the frame of the SCCH COMET competence center INTEGRATE \(FFG 892418\) and by the “ICT of the Future” project QuanTD \(no\. 898626\), as well as by the Austrian Science Fund FWF \(10\.55776/COE12\)\. RedHatAI/gemma\-4\-31B\-it\-FP8\-Dynamic and Claude Pro were used to generate code and code documentation for the experimental evaluation, and to revise and proofread parts of the paper\. All ideas and the scientific content of this paper originate from the authors\. Where text was AI\-assisted, it was rewritten and verified by the authors, who take full responsibility for the content of this work\. ## Artifacts The code implementingLeDQeRis available in a public GitHub repository141414https://github\.com/Anna\-Christina\-Glock/dq\_rule\_generation\. The code to generate thePCIdataset is also available in a public GitHub repository151515https://github\.com/Anna\-Christina\-Glock/pci\-llm\-toolkit/tree/main/data\-generator, and the other four \(polluted and labeled\) datasets are also publicly available161616https://github\.com/Anna\-Christina\-Glock/dq\_rule\_generation\. The LLMs used during the experiments are available via HuggingFace171717https://huggingface\.co/models\. ## References - Baleket al\.\(2025\)V\. Balek, L\. Sýkora, V\. Sklenák, and T\. KliegrLLM\-based feature generation from text for interpretable machine learning\.Machine Learning114\(11\),pp\. 241\.External Links:[Document](https://dx.doi.org/10.1007/S10994-025-06867-1)Cited by:[§5](https://arxiv.org/html/2609.06053#S5.SS0.SSS0.Px3.p1.1)\. - Bhadauriaet al\.\(2026\)D\. Bhadauria, H\. Harmouch, F\. Naumann, D\. Srivastava, and L\. EhrlingerA catalog of data errors\.External Links:2604\.09277,[Link](https://arxiv.org/abs/2604.09277)Cited by:[item Explicit missing values \(emv\)](https://arxiv.org/html/2609.06053#S2.I1.ix1.p1.1),[item Disguised missing values \(dmv\)](https://arxiv.org/html/2609.06053#S2.I1.ix2.p1.1),[item Misfielded values \(mfv\)](https://arxiv.org/html/2609.06053#S2.I1.ix4.p1.1),[item Embedded values \(ebv\)](https://arxiv.org/html/2609.06053#S2.I1.ix5.p1.1),[item Spelling mistakes \(spm\)](https://arxiv.org/html/2609.06053#S2.I1.ix6.p1.1),[item Incorrect formats \(ifo\)](https://arxiv.org/html/2609.06053#S2.I1.ix8.p1.1),[item Incorrect character sets \(ics\)](https://arxiv.org/html/2609.06053#S2.I1.ix9.p1.1),[§2\.1](https://arxiv.org/html/2609.06053#S2.SS1.p3.1)\. - Bodensohnet al\.\(2025\)J\. Bodensohn, U\. Brackmann, L\. Vogel, A\. Sanghi, and C\. BinnigUnveiling challenges for llms in enterprise data engineering\.Proc\. VLDB Endow\.19\(2\),pp\. 196–209\.Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p2.1)\. - Chandruet al\.\(2025\)R\. Chandru, P\. Kadusabe, A\. M\. Qureshi, S\. Sharma, and A\. KaushikLarge language models \(llms\) in medical error detection and correction: a comprehensive review\.Next\-Gen Healthcare: AI\-Powered Medical Innovations,pp\. 149–173\.Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p2.1)\. - Chiang and Miller \(2008\)F\. Chiang and R\. J\. MillerDiscovering data quality rules\.Proceedings of the VLDB Endowment1\(1\),pp\. 1166–1177\.Cited by:[§1](https://arxiv.org/html/2609.06053#S1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.06053#S1.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.06053#S5.SS0.SSS0.Px2.p1.1)\. - Chienet al\.\(2025\)M\. Chien, D\. Radhakrishnan, and S\. WaiteMagic quadrant for augmented data quality solutions\.Technical reportGartner\.External Links:[Link](https://www.gartner.com/en/documents/6246519)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p3.1)\. - de Andradeet al\.\(2026\)J\. B\. C\. de Andrade, M\. A\. O\. de Medeiros Cavalcante, T\. L\. M\. Lopes, J\. M\. S\. Treigher, M\. D\. Balsells, J\. L\. Vasconcelos, L\. C\. Monteiro, and D\. D\. da Silveira MotaDiscovery of data quality issues in electronic health records: profound consequences for critical care medicine applications – a systematized review\.Critical Care30\(1\)\.External Links:ISSN 1364\-8535,[Link](http://dx.doi.org/10.1186/s13054-025-05677-0),[Document](https://dx.doi.org/10.1186/s13054-025-05677-0)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p1.1)\. - Ehrlinger and Wöß \(2022\)L\. Ehrlinger and W\. WößA survey of data quality measurement and monitoring tools\.Frontiers Big Data5,pp\. 850611\.Cited by:[§1](https://arxiv.org/html/2609.06053#S1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.06053#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.06053#S2.SS2.p2.1)\. - European Parliament and Council of the European Union \(2016\)European Parliament and Council of the European UnionGeneral data protection regulation \(gdpr\)\.External Links:[Link](https://eur-lex.europa.eu/legal-content/EN/TXT/PDF/?uri=CELEX:02016R0679-20160504)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p2.1)\. - European Parliament and Council of the European Union \(2024\)European Parliament and Council of the European UnionRegulation \(eu\) 2024/1689 laying down harmonised rules on artificial intelligence \(artificial intelligence act\)\.Note:Official Journal of the European Union, L seriesExternal Links:[Link](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p1.1)\. - Freund and Schapire \(1997\)Y\. Freund and R\. E\. SchapireA decision\-theoretic generalization of on\-line learning and an application to boosting\.Journal of Computer and System Sciences55\(1\),pp\. 119–139\.Cited by:[§5](https://arxiv.org/html/2609.06053#S5.SS0.SSS0.Px1.p1.1)\. - Fürnkranz \(1999\)J\. FürnkranzSeparate\-and\-conquer rule learning\.Artificial Intelligence Review13\(1\),pp\. 3–54\.Cited by:[§3\.2\.2](https://arxiv.org/html/2609.06053#S3.SS2.SSS2.Px4.p1.1)\. - Gaoet al\.\(2022\)T\. Gao, S\. Singh, and R\. J\. MooneyTowards automated error analysis: learning to characterize errors\.InProceedings of the 35th International Florida Artificial Intelligence Research Society Conference \(FLAIRS\),Hutchinson Island, Jensen Beach, Florida\.External Links:[Document](https://dx.doi.org/10.32473/FLAIRS.V35I.130632)Cited by:[§5](https://arxiv.org/html/2609.06053#S5.SS0.SSS0.Px1.p1.1)\. - Glocket al\.\(2025\)A\. Glock, C\. Dominka\-Kiss, P\. Korom, and L\. EhrlingerDetecting and cleaning errors in personal contact information with large language models\.InProceedings of the 2nd International Workshop on Data Acquisition and Analytics \(DATAI@VLDB\),External Links:[Link](https://www.vldb.org/2025/Workshops/VLDB-Workshops-2025/DATAI/DATAI25_12.pdf)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p2.1),[§4\.1\.2](https://arxiv.org/html/2609.06053#S4.SS1.SSS2.p1.1)\. - Gottlob \(2024\)G\. GottlobEnhancing data precision with large language models: analyzing failures and innovating database curation\.InProceedings of the 32nd Symposium of Advanced Database Systems,CEUR Workshop Proceedings, Vol\.3741,Villasimius, Italy,pp\. 1–2\.Cited by:[§5](https://arxiv.org/html/2609.06053#S5.SS0.SSS0.Px3.p1.1)\. - Great Expectations \(2026\)Great ExpectationsGreat expectations’ official partners\.Great Expectations\.External Links:[Link](https://greatexpectations.io/partners/)Cited by:[§4\.1\.1](https://arxiv.org/html/2609.06053#S4.SS1.SSS1.p1.1)\. - Hastie and Tibshirani \(1986\)T\. Hastie and R\. TibshiraniGeneralized additive models\.\.Statistical Science1\(3\),pp\. 297–310\.External Links:[Document](https://dx.doi.org/10.1214/ss/1177013604)Cited by:[§5](https://arxiv.org/html/2609.06053#S5.SS0.SSS0.Px1.p1.1)\. - Hauget al\.\(2011\)A\. Haug, F\. Zachariassen, and D\. van LiempdThe costs of poor data quality\.J\. Ind\. Eng\. Manag\.4\(2\) \(en\)\.External Links:[Document](https://dx.doi.org/10.3926/jiem.2011.v4n2.p168-193)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p1.1)\. - Ilyas and Naumann \(2022\)I\. F\. Ilyas and F\. NaumannData errors: symptoms, causes and origins\.IEEE Data Eng\. Bull\.45\(1\),pp\. 4–9\.Cited by:[item Explicit missing values \(emv\)](https://arxiv.org/html/2609.06053#S2.I1.ix1.p1.1),[§2\.1](https://arxiv.org/html/2609.06053#S2.SS1.p3.1)\. - Joskoet al\.\(2016\)J\. M\. B\. Josko, M\. K\. Oikawa, and J\. E\. FerreiraA formal taxonomy to improve data defect description\.InDASFAA,Lecture Notes in Computer Science, Vol\.9645,pp\. 307–320\.External Links:[Link](https://doi.org/10.1007/978-3-319-32055-7/_25),[Document](https://dx.doi.org/10.1007/978-3-319-32055-7%5F25)Cited by:[item Contradictions \(con\)](https://arxiv.org/html/2609.06053#S2.I1.ix3.p1.1),[item Spelling mistakes \(spm\)](https://arxiv.org/html/2609.06053#S2.I1.ix6.p1.1),[item Domain violations \(dov\)](https://arxiv.org/html/2609.06053#S2.I1.ix7.p1.1),[§2\.1](https://arxiv.org/html/2609.06053#S2.SS1.p3.1)\. - Junget al\.\(2025\)P\. Jung, S\. Jäger, N\. Chandler, and F\. BiessmannTowards realistic error models for tabular data\.ACM J\. Data Inf\. Qual\.17\(4\),pp\. 28:1–28:27\.Cited by:[item Explicit missing values \(emv\)](https://arxiv.org/html/2609.06053#S2.I1.ix1.p1.1),[item Misfielded values \(mfv\)](https://arxiv.org/html/2609.06053#S2.I1.ix4.p1.1),[item Spelling mistakes \(spm\)](https://arxiv.org/html/2609.06053#S2.I1.ix6.p1.1),[item Incorrect formats \(ifo\)](https://arxiv.org/html/2609.06053#S2.I1.ix8.p1.1),[item Incorrect character sets \(ics\)](https://arxiv.org/html/2609.06053#S2.I1.ix9.p1.1)\. - Kauschke and Fürnkranz \(2018\)S\. Kauschke and J\. FürnkranzBatchwise patching of classifiers\.InProceedings of the 32nd AAAI Conference on Artificial Intelligence \(AAAI\-18\),pp\. 3374–3381\.External Links:[Link](https://www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/16099)Cited by:[§5](https://arxiv.org/html/2609.06053#S5.SS0.SSS0.Px1.p1.1)\. - Kimet al\.\(2003\)W\. Kim, B\. Choi, E\. Hong, S\. Kim, and D\. LeeA taxonomy of dirty data\.Data Mining and Knowledge Discovery7\(1\),pp\. 81–99\.External Links:ISSN 1384\-5810,[Link](https://doi.org/10.1023/A:1021564703268),[Document](https://dx.doi.org/10.1023/A%3A1021564703268)Cited by:[item Explicit missing values \(emv\)](https://arxiv.org/html/2609.06053#S2.I1.ix1.p1.1),[item Misfielded values \(mfv\)](https://arxiv.org/html/2609.06053#S2.I1.ix4.p1.1),[item Spelling mistakes \(spm\)](https://arxiv.org/html/2609.06053#S2.I1.ix6.p1.1),[item Incorrect character sets \(ics\)](https://arxiv.org/html/2609.06053#S2.I1.ix9.p1.1),[§2\.1](https://arxiv.org/html/2609.06053#S2.SS1.p3.1)\. - Loshin \(2002\)D\. LoshinRule\-based data quality\.InProceedings of the Eleventh International Conference on Information and Knowledge Management,CIKM ’02,New York, NY, USA,pp\. 614–616\.External Links:ISBN 1581134924,[Link](https://doi.org/10.1145/584792.584894),[Document](https://dx.doi.org/10.1145/584792.584894)Cited by:[§2\.2](https://arxiv.org/html/2609.06053#S2.SS2.p1.1)\. - Mahdaviet al\.\(2019\)M\. Mahdavi, Z\. Abedjan, R\. Castro Fernandez, S\. Madden, M\. Ouzzani, M\. Stonebraker, and N\. TangRaha: a configuration\-free error detection system\.InProceedings of the 2019 International Conference on Management of Data,SIGMOD ’19,New York, NY, USA,pp\. 865–882\.External Links:ISBN 9781450356435,[Link](https://doi.org/10.1145/3299869.3324956),[Document](https://dx.doi.org/10.1145/3299869.3324956)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p2.1),[§4\.1\.2](https://arxiv.org/html/2609.06053#S4.SS1.SSS2.p1.1)\. - Mohammedet al\.\(2025\)S\. Mohammed, L\. Budach, M\. Feuerpfeil, N\. Ihde, A\. Nathansen, N\. Noack, H\. Patzlaff, F\. Naumann, and H\. HarmouchThe effects of data quality on machine learning performance on tabular data\.Information Systems132,pp\. 102549\.Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p1.1)\. - Moriet al\.\(2023\)L\. Mori, B\. Richardson, T\. Saleh, R\. Sellschop, and I\. WellsClearing data\-quality roadblocks: Unlocking AI in manufacturing\.Note:[https://www\.mckinsey\.com/capabilities/tech\-and\-ai/our\-insights/clearing\-data\-quality\-roadblocks\-unlocking\-ai\-in\-manufacturing\#/](https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/clearing-data-quality-roadblocks-unlocking-ai-in-manufacturing#/)\[Accessed 28\-01\-2026\]Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p1.1)\. - Müller and Freytag \(2005\)H\. Müller and J\.C\. FreytagProblems, methods, and challenges in comprehensive data cleansing\.Informatik\-Berichte,Humboldt\-Univ\. zu Berlin\.Cited by:[item Explicit missing values \(emv\)](https://arxiv.org/html/2609.06053#S2.I1.ix1.p1.1),[item Contradictions \(con\)](https://arxiv.org/html/2609.06053#S2.I1.ix3.p1.1),[item Domain violations \(dov\)](https://arxiv.org/html/2609.06053#S2.I1.ix7.p1.1),[item Incorrect formats \(ifo\)](https://arxiv.org/html/2609.06053#S2.I1.ix8.p1.1),[§2\.1](https://arxiv.org/html/2609.06053#S2.SS1.p3.1)\. - Narayanet al\.\(2022\)A\. Narayan, I\. Chami, L\. J\. Orr, and C\. RéCan foundation models wrangle your data?\.Proc\. VLDB Endow\.16\(4\),pp\. 738–746\.Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p2.1)\. - Oliveiraet al\.\(2005\)P\. Oliveira, F\. Rodrigues, and P\. R\. HenriquesA formal definition of data quality problems\.InProceedings of the 2005 International Conference on Information Quality \(MIT ICIQ Conference\),Cited by:[item Explicit missing values \(emv\)](https://arxiv.org/html/2609.06053#S2.I1.ix1.p1.1),[item Contradictions \(con\)](https://arxiv.org/html/2609.06053#S2.I1.ix3.p1.1),[item Embedded values \(ebv\)](https://arxiv.org/html/2609.06053#S2.I1.ix5.p1.1),[item Spelling mistakes \(spm\)](https://arxiv.org/html/2609.06053#S2.I1.ix6.p1.1),[item Domain violations \(dov\)](https://arxiv.org/html/2609.06053#S2.I1.ix7.p1.1),[item Incorrect formats \(ifo\)](https://arxiv.org/html/2609.06053#S2.I1.ix8.p1.1),[§2\.1](https://arxiv.org/html/2609.06053#S2.SS1.p3.1)\. - Pearson \(2006\)R\. PearsonThe problem of disguised missing data\.InSIGKDD Explorations Newsletter,Vol\.8,pp\. 83–92\.Cited by:[item Explicit missing values \(emv\)](https://arxiv.org/html/2609.06053#S2.I1.ix1.p1.1),[item Disguised missing values \(dmv\)](https://arxiv.org/html/2609.06053#S2.I1.ix2.p1.1)\. - Pipinoet al\.\(2002\)L\. L\. Pipino, Y\. W\. Lee, and R\. Y\. WangData quality assessment\.Communications of the ACM45\(4\),pp\. 211–218\.External Links:[Document](https://dx.doi.org/10.1145/505248.506010)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p1.1)\. - Ret al\.\(2025\)R\. R, A\. P\. Vemali, D\. Aswal, R\. Ramesh, and A\. BhupalScottyPoseidon at SemEval\-2025 task 8: LLM\-driven code generation for zero\-shot question answering on tabular data\.InProceedings of the 19th International Workshop on Semantic Evaluation \(SemEval\-2025\),Vienna, Austria,pp\. 2197–2204\.External Links:[Link](https://aclanthology.org/2025.semeval-1.285/),ISBN 979\-8\-89176\-273\-2Cited by:[§1](https://arxiv.org/html/2609.06053#S1.SS0.SSS0.Px2.p1.1)\. - Rahm and Do \(2000\)E\. Rahm and H\. H\. DoData cleaning: problems and current approaches\.IEEE Data Eng\. Bull\.23\(4\),pp\. 3–13\.Cited by:[item Explicit missing values \(emv\)](https://arxiv.org/html/2609.06053#S2.I1.ix1.p1.1),[item Misfielded values \(mfv\)](https://arxiv.org/html/2609.06053#S2.I1.ix4.p1.1),[item Embedded values \(ebv\)](https://arxiv.org/html/2609.06053#S2.I1.ix5.p1.1),[item Spelling mistakes \(spm\)](https://arxiv.org/html/2609.06053#S2.I1.ix6.p1.1),[§2\.1](https://arxiv.org/html/2609.06053#S2.SS1.p3.1)\. - Redman \(1998\)T\. C\. RedmanThe impact of poor data quality on the typical enterprise\.Communications of the ACM41\(2\),pp\. 79–82\.External Links:[Document](https://dx.doi.org/10.1145/269012.269034)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p1.1)\. - Rehbergeret al\.\(2026\)T\. Rehberger, T\. Hütter, L\. Ehrlinger, and W\. WößEvaluating data quality tools: measurement capabilities and llm integration\.arXiv preprint arXiv:2604\.09163\.External Links:[Link](https://arxiv.org/abs/2604.09163)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p3.1)\. - Rekatsinaset al\.\(2017\)T\. Rekatsinas, X\. Chu, I\. F\. Ilyas, and C\. RéHoloClean: holistic data repairs with probabilistic inference\.Proceedings of the VLDB Endowment10\(11\),pp\. 1190–1201\.External Links:[Link](http://www.vldb.org/pvldb/vol10/p1190-rekatsinas.pdf),[Document](https://dx.doi.org/10.14778/3137628.3137631)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p2.1)\. - Schneider \(2025\)A\. SchneiderTemplate\-guided rule generation and evaluation for data quality using large language models\.Master’s Thesis,TU Wien\.Cited by:[§3\.1](https://arxiv.org/html/2609.06053#S3.SS1.p2.1),[§5](https://arxiv.org/html/2609.06053#S5.SS0.SSS0.Px3.p2.1),[Table 4](https://arxiv.org/html/2609.06053#S5.T4),[§6](https://arxiv.org/html/2609.06053#S6.p1.1)\. - Talebet al\.\(2021\)I\. Taleb, M\. A\. Serhani, C\. Bouhaddioui, and R\. DssouliBig data quality framework: a holistic approach to continuous quality management\.J\. Big Data8\(1\),pp\. 76\.External Links:[Document](https://dx.doi.org/10.1186/S40537-021-00468-0)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.SS0.SSS0.Px1.p1.1)\. - Wang and Strong \(1996\)R\. Y\. Wang and D\. M\. StrongBeyond accuracy: what data quality means to data consumers\.Journal of Management Information Systems12\(4\),pp\. 5–33\.External Links:[Document](https://dx.doi.org/10.1080/07421222.1996.11518099)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.06053#S2.SS1.p2.1)\. - Xieet al\.\(2025\)S\. Xie, H\. Cai, Y\. Sun, and X\. LvLLM\-DQR: large language model\-based automated generation of data quality rules for electronic health records\.J\. Biomed\. Inform\.172\(104951\),pp\. 104951\(en\)\.Cited by:[§3\.1](https://arxiv.org/html/2609.06053#S3.SS1.p2.1),[§5](https://arxiv.org/html/2609.06053#S5.SS0.SSS0.Px3.p2.1),[Table 4](https://arxiv.org/html/2609.06053#S5.T4),[§6](https://arxiv.org/html/2609.06053#S6.p1.1)\. - Yadav and Mondal \(2025\)D\. Yadav and S\. MondalEvaluating pre\-trained large language models on zero shot prompts for parallelization of source code\.J\. Syst\. Softw\.230\(112543\),pp\. 112543\(en\)\.External Links:[Document](https://dx.doi.org/10.1016/j.jss.2025.112543)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.SS0.SSS0.Px2.p1.1)\. - Yanget al\.\(2025\)Q\. Yang, Z\. Hong, D\. Cao, H\. Wang, Z\. Xie, T\. He, Y\. Liu, Y\. Yang, and D\. ZhangAddrLLM: address rewriting via large language model on nationwide logistics data\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.1,KDD ’25,New York, NY, USA,pp\. 2756–2767\.External Links:ISBN 9798400712456Cited by:[§1](https://arxiv.org/html/2609.06053#S1.p2.1)\. - Yeh and Puri \(2010\)P\. Z\. Yeh and C\. A\. PuriAn efficient and robust approach for discovering data quality rules\.In22nd IEEE International Conference on Tools with Artificial Intelligence, ICTAI,pp\. 248–255\.External Links:[Document](https://dx.doi.org/10.1109/ICTAI.2010.43)Cited by:[§1](https://arxiv.org/html/2609.06053#S1.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.06053#S5.SS0.SSS0.Px2.p1.1)\.
Similar Articles
Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation
This paper introduces DoLQ, a multi-agent framework that uses Large Language Models to perform both qualitative and quantitative evaluations for discovering ordinary differential equations from observational data.
RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning
RimRule proposes a neuro-symbolic method that distills compact, interpretable rules from failure traces using the Minimum Description Length principle, improving LLM tool-use performance without modifying weights, and demonstrating rule portability across models.
RuleChef: Grounding LLM Task Knowledge in Human-Editable Rules
RuleChef is a framework that uses LLMs to generate human-editable, executable rules for NLP tasks, iteratively improving them based on examples and human feedback, resulting in fast, deterministic, and inspectable rule systems.
When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering
This paper investigates quality issues in LLM-generated answers for hardware description language questions, finding over-answering tendencies like redundancy (65.7%) and verbosity (69.1%), and proposes a multi-agent framework that reduces core answers by 37% and non-core content length by 31% while improving quality scores.
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
This paper introduces RLearner-LLM, a framework using Hybrid-DPO to balance logical correctness and fluency in LLM-generated explanations, achieving significant NLI entailment improvements across multiple domains and base models while mitigating the verbosity bias of standard preference signals.