无需训练的合成:面向表格、时序与关系型数据的纯推理合成数据流水线
摘要
论文提出 GenScript,一种纯推理的合成数据流水线。它不再针对每个数据集训练模型,而是将源数据的确定性统计画像输入 LLM,由其推断字段语义与完整性约束,再由编码代理将其编译为可审计的采样器。该方法统一了表格、时序与关系型数据的合成,可在几分钟内构建生成器,同时在保真度上接近领先方法,并保留关键的完整性约束。
arXiv:2609.38414v1 Announce Type: new
Abstract: Synthetic data generation is dominated by the fit-then-sample paradigm: a generative model is trained on a private dataset and then sampled from. Despite its widespread adoption, this paradigm faces three challenges: (1) a new training run is required for every dataset; (2) different data modalities, such as single tables, time series, and relational databases, require task-specific models and feature engineering; and (3) the resulting model is opaque, making its behavior under data constraints difficult to inspect. We propose GENSCRIPT, an inference-only pipeline that eliminates model training. GENSCRIPT computes a deterministic statistical profile of the source data (column types, ranges, missingness, categories, correlations, etc.) and passes it--rather than raw rows--to a language model to infer field semantics and cross-column integrity constraints. A coding agent then compiles the profile and constraints into an executable, auditable sampler. This unified approach supports single-table, temporal, and relational data without task-specific modeling. Across four single-table benchmarks, GENSCRIPT builds generators in 2 minutes and samples 50k rows within 6 seconds, while remaining within a few points of leading methods in marginal fidelity. Notably, it is the only method that perfectly preserves a 1-to-1 mapping between columns in the Adult dataset. On a smart-building dataset, it produces conditional time series that more closely match the real distribution than two baselines and perfectly preserves primary- and foreign-key relationships in the corresponding relational database.
查看缓存全文
缓存时间: 2026/10/02 09:49
# Synthesis Without Training:An Inference-Only Pipeline for Tabular,Temporal, and Relational Synthetic Data
Source: [https://arxiv.org/html/2609.38414](https://arxiv.org/html/2609.38414)
\\workshoptitle
Beyond Private Training: The New Landscape of AI Privacy
Abdul RaheemAffiliation:Betterdata AIJiayu LiAffiliation:University of Illinois Urbana\-ChampaignSohei ArisakaAffiliation:KAJIMA Technical Research Institute SingaporeDarius Lim Hong YiAffiliation:KAJIMA Technical Research Institute SingaporeMilad AbdollahzadehAffiliation:Betterdata AIUzair JavaidAffiliation:Betterdata AIBiplab SikdarAffiliation:National University of Singapore
###### Abstract
Synthetic data generation is dominated by the*fit\-then\-sample*paradigm: a generative model is trained on a private dataset and then sampled from\. Despite its widespread adoption, this paradigm faces three challenges: \(1\) a new training run is required for every dataset; \(2\) different data modalities, such as single tables, time series, and relational databases, require task\-specific models and feature engineering; and \(3\) the resulting model is opaque, making its behavior under data constraints difficult to inspect\. We proposeGenScript, an*inference\-only*pipeline that eliminates model training\.GenScriptcomputes a deterministic statistical profile of the source data \(column types, ranges, missingness, categories, correlations, etc\.\) and passes it—rather than raw rows—to a language model to infer field semantics and cross\-column integrity constraints\. A coding agent then compiles the profile and constraints into an executable, auditable sampler\. This unified approach supports single\-table, temporal, and relational data without task\-specific modeling\. Across four single\-table benchmarks,GenScriptbuilds generators in 2 minutes and samples 50k rows within 6 seconds, while remaining within a few points of leading methods in marginal fidelity\. Notably, it is the only method that perfectly preserves a 1\-to\-1 mapping between columns in the Adult dataset\. On a smart\-building dataset, it produces conditional time series that more closely match the real distribution than two baselines and perfectly preserves primary\- and foreign\-key relationships in the corresponding relational database\.
## 1Introduction
Organisations holding sensitive tabular data increasingly rely on synthetic surrogates for sharing, testing, and downstream model development\. The standard recipe is to fit a deep generative model to the private data and sample from it\. Conditional GANs and VAEs for tables\[[18](https://arxiv.org/html/2609.38414#bib.bib4)\], diffusion models\[[6](https://arxiv.org/html/2609.38414#bib.bib5),[14](https://arxiv.org/html/2609.38414#bib.bib11)\], and large\-language\-model\-based generators\[[1](https://arxiv.org/html/2609.38414#bib.bib6)\]all follow this recipe, as do the specialised architectures used for sequential\[[21](https://arxiv.org/html/2609.38414#bib.bib7),[9](https://arxiv.org/html/2609.38414#bib.bib8)\]and relational data\[[13](https://arxiv.org/html/2609.38414#bib.bib1),[15](https://arxiv.org/html/2609.38414#bib.bib9)\]\.
This paradigm imposes a structural tax\.Per\-dataset trainingmeans every new dataset needs its own run, hyperparameter search, and GPU budget, so cost scales linearly rather than amortising\.Modality fragmentationmeans a single table, a time series, and a relational database are served by three model families with three training procedures and three sets of feature engineering, making model*selection*the practitioner’s first task\.Opacitymeans that when a generator violates a business rule—a refund on an unpaid order—there is no direct place to inspect or repair it\.
#### Our position\.
For a large and practically important class of datasets, the generative model is unnecessary\. What a practitioner needs is an accurate description of the data’s marginal and joint structure, and of the rules it obeys\. Both can be obtained*without*gradient descent: the first from deterministic profiling code, the second from a language model reasoning over that profile\. A coding agent compiles both into an ordinary program that emits rows\. We call this*inference\-only*generation: nothing is trained at any point, and the only learned component is a frozen, general\-purpose LLM used at inference time\. We instantiate this asGenScript\(Section[2](https://arxiv.org/html/2609.38414#S2)\), which contributes atraining\-free pipelinereplacing the trained generator with an LLM\-authored sampling program;one pipeline across modalities, since the data–generator interface is a profile rather than a tensor, so modality enters only as extra profile fields \(Section[2\.4](https://arxiv.org/html/2609.38414#S2.SS4)\); andexplicit constraints, materialised as readable predicates that can be audited, edited, or supplied by an expert\. We claim no advantage on data whose value lies in high\-order structure a profile cannot summarise; our claim is that for schema\-driven operational data, an explicit program offers substantial reductions in computational cost and improved auditability, at a measurable cost in fidelity\.
#### Related work\.
Tabular generators fit a model to the private data\[[18](https://arxiv.org/html/2609.38414#bib.bib4),[25](https://arxiv.org/html/2609.38414#bib.bib23),[6](https://arxiv.org/html/2609.38414#bib.bib5),[26](https://arxiv.org/html/2609.38414#bib.bib24),[14](https://arxiv.org/html/2609.38414#bib.bib11),[7](https://arxiv.org/html/2609.38414#bib.bib12),[1](https://arxiv.org/html/2609.38414#bib.bib6),[24](https://arxiv.org/html/2609.38414#bib.bib25),[15](https://arxiv.org/html/2609.38414#bib.bib9),[23](https://arxiv.org/html/2609.38414#bib.bib26)\], and sequential and relational data each bring their own families again\[[21](https://arxiv.org/html/2609.38414#bib.bib7),[9](https://arxiv.org/html/2609.38414#bib.bib8),[17](https://arxiv.org/html/2609.38414#bib.bib13),[16](https://arxiv.org/html/2609.38414#bib.bib14),[12](https://arxiv.org/html/2609.38414#bib.bib15),[8](https://arxiv.org/html/2609.38414#bib.bib16)\]\. Closest to us are methods that use an LLM to read a table’s meaning and then hand generation to a fitted model: LLM\-TabFlow\[[10](https://arxiv.org/html/2609.38414#bib.bib19)\]recovers inter\-column logical relations for a diffusion model, and SPADA\[[19](https://arxiv.org/html/2609.38414#bib.bib18)\]induces a sparse dependency graph and samples by kernel density estimation\. We differ in handing generation to a*program*, so no density is estimated at any point\. More discussions are provided in Appendix[A](https://arxiv.org/html/2609.38414#A1)\.
## 2Method
real dataDDIngestnormaliseProfilemeasureAnalyzeinterpretAssemblecompileΦ\\PhiΣ\\SigmaGeneraterunGGsyntheticD^\\hat\{D\}EvaluatescoreGGverifyΣ\\SigmaagainstDDLLM \+ coding agentFigure 1:TheGenScriptworking pipeline\. The implemented UI is provided in Appendix[D](https://arxiv.org/html/2609.38414#A4)\.LetD=\{T1,…,TK\}D=\\\{T\_\{1\},\\dots,T\_\{K\}\\\}be a source dataset ofKKtables \(K=1K=1for the single\-table case\)\.GenScriptproducesD^\\hat\{D\}through the six phases of Figure[1](https://arxiv.org/html/2609.38414#S2.F1):*Ingest*,*Profile*,*Analyze*and*Assemble*construct the generator,*Generate*runs it, and*Evaluate*scores the output\.
### 2\.1Ingest and Profile: deterministic statistical profiling
A fixed, hand\-written script computes a profileΦ\(D\)\\Phi\(D\); no language model is involved, so this phase is exact, cheap, and reproducible\. Per column we record the inferred type \(continuous, integer, categorical, datetime, text, identifier\); the observed range, or for categorical columns the category inventory with empirical frequencies; the missing\-value rate and any conditional missingness; and summary statistics\. Per table we compute a correlation matrix; for relational inputs we record keys and each parent–child cardinality distribution; for time series the profile adds the timestamp granularity and the distribution of inter\-arrival gaps, per\-entity sequence\-length and coverage statistics\.
### 2\.2Analyze: semantic analysis and constraint induction
We passΦ\(D\)\\Phi\(D\), including the correlation matrix, to a language model and ask for a semantic specificationΣ\\Sigma\. This recovers information present in the data but invisible to the profiler:field semantics, such as that a five\-digit integer column is a postal code rather than a quantity;cross\-column constraintssuch as “ifA\>0\>0thenB=0=0”, for which strong correlation entries and structured missingness are evidence that a rule exists while column names supply its form; andconditional structure, indicating which columns to sample conditioned on which others\.
FormallyΣ=\{\(cj,wj\)\}\\Sigma=\\\{\(c\_\{j\},w\_\{j\}\)\\\}is a set of predicates over a row—or a row and its parent, in the relational case—with confidence weights, short enough to be reviewed and corrected by a domain expert before generation\. Every candidate predicate is additionally*verified against the real data*: those holding on fewer thanτ\\tauof real rows are demoted or discarded, guarding against rules hallucinated from suggestive column names alone\.
### 2\.3Assemble, Generate and Evaluate: synthesizing and running the sampler
A coding agent receivesΦ\\PhiandΣ\\Sigmaand writes an executable generatorGG, given a target interface and a sandbox in which to run its own output\. Generation proceeds by constrained rejection sampling: the program draws a candidate row from a factorized proposal built from the profile’s marginals and conditional structure, accepting it only if the hard predicates hold,
x^∼qΦ\(⋅\),accept if∏j∈ℋcj\(x^\)=1,\\hat\{x\}\\sim q\_\{\\Phi\}\(\\cdot\),\\qquad\\text\{accept if \}\\textstyle\\prod\_\{j\\in\\mathcal\{H\}\}c\_\{j\}\(\\hat\{x\}\)=1,\(1\)whereℋ\\mathcal\{H\}indexes the hard constraints\. Where a constraint is deterministic—an arithmetic identity, a derived field—the agent*imputes*rather than rejects, computing the dependent value directly; rejection is reserved for inequality\-shaped constraints, which keeps acceptance rates high\. Eventually, synthetic dataD^\\hat\{D\}are evaluated as shown in Section[3](https://arxiv.org/html/2609.38414#S3)\.
### 2\.4Extension to time series and relational data
The modalities differ only in what*Profile*records and what*Generate*emits\. For time series the profile adds the sampling interval, per\-entity sequence\-length distribution, trend and seasonality estimates, and lag\-kkautocorrelations\. For relational databases it adds the schema graph and the distribution of child\-row counts per parent, and the agent walks the schema in topological order with foreign keys assigned by construction, so referential integrity holds by design rather than by filtering\. The practitioner selects no model, because there is none to select\.
## 3Experiments
#### Setup\.
For single tables we use CoverType, Credit, Intrusion, and Adult, covering heavy categorical structure, class imbalance, and large row counts; per\-dataset statistics for these and for the two relational tables are given in Appendix[B](https://arxiv.org/html/2609.38414#A2)\. Baselines are TabTreeFormer\[[7](https://arxiv.org/html/2609.38414#bib.bib12)\], TabDiff\[[14](https://arxiv.org/html/2609.38414#bib.bib11)\], and REaLTabFormer\[[15](https://arxiv.org/html/2609.38414#bib.bib9)\], one from each of the tree\-based, diffusion, and auto\-regressive language\-model families\. Each method generates as many rows as the source table contains: 50 000 for the first three, 32 561 for Adult\.τ\\tauis set to 100 rows for the semantic analysis phase\.GenScriptuses gpt\-5\.6\-sol as its LLM backend; the baselines were trained on a server with 32 CPU cores, 200 GB RAM, and two RTX 4090 Ti GPUs\.
#### Metrics\.
Fidelity uses*Shape*and*Trend*\[[2](https://arxiv.org/html/2609.38414#bib.bib17),[14](https://arxiv.org/html/2609.38414#bib.bib11)\]: Shape is the similarity of each column’s marginal density, Trend the fidelity of correlations between column pairs; higher is better for both\. Utility uses machine\-learning efficacy under train\-on\-synthetic, test\-on\-real: we split the real data:28\\\!:\\\!2, fit an XGBoost classifier on the synthetic table, and report AUC on the held\-out real20%20\\%\. We also report generator\-construction \(“Train”\) and sampling time; forGenScript, “Train” covers profiling, constraint induction, and program synthesis, with no gradient step in it\.
Table 1:Fidelity and utility on the four single\-table benchmarks\. Best per column in bold\.Table 2:Generator\-construction time \(“Train”\) and sampling time, in seconds\. ForGenScript, “Train” covers profiling, constraint induction, and program synthesis\.
#### Fidelity and utility\.
On ShapeGenScriptis close to the leaders \(0\.984, 0\.936, 0\.930 and 0\.874 on CoverType, Adult, Credit and Intrusion\), second of four on CoverType and third elsewhere, which is a good outcome for a sampler built from per\-column marginals alone\. On Trend it is last on all four \(0\.748 to 0\.897\), the direct cost of reproducing only the dependencies named inΣ\\Sigma\. MLE is the weakest axis and degrades with the number of target classes: 0\.894 on Credit and 0\.828 on Adult, both binary, against 0\.572 on Intrusion and 0\.467 on CoverType, which have twenty and seven classes respectively\. Further discussions are provided in Appendix[C](https://arxiv.org/html/2609.38414#A3)\.
#### Constraint preservation\.
The aggregate metrics are blind to hard integrity rules\. In Adult,educationandeducation\-numencode the same fact twice, so the real table admits exactly 16 diploma/year pairs\. Every trained baseline emits rows violating this dependency, whereasGenScriptmaterialises it as a hard predicate and emits none\. Such rows are impossible rather than improbable, and no metric in Table[1](https://arxiv.org/html/2609.38414#S3.T1)registers them; Appendix[F](https://arxiv.org/html/2609.38414#A6)gives per\-method rates\.
#### Cost\.
Table[2](https://arxiv.org/html/2609.38414#S3.T2)is where the inference\-only design pays\.GenScriptis the fastest method among all the baselines\. It builds its generator in 73 to 115 seconds, against 175 to 6760 seconds for the baselines, and samples in 0\.9 to 5\.4 seconds, against 8\.6 to 534 seconds\. Measured against whichever baseline is fastest on each dataset, construction is 1\.9 to 3\.6 times quicker and sampling 2\.4 to 9\.6 times quicker; End\-to\-end the totals are 74 to 117 seconds against 213 to 566 seconds for the best baseline on each dataset, a margin of 2\.3 to 4\.9 times achieved\.
#### Time series and relational structure\.
For the remaining modalities we use a mock\-up dataset from an industrial partner in the construction sector\. An access\-control table \(acc\) records everyone granted entry to a building, keyed byUserID; a camera\-detection table \(aicamera\) records their sightings as a timestamped time series whoseUserIDreferencesacc, so the pair is both a conditional time series and a two\-table relational database\. On the relational axis nothing separates the methods: all preserve primary\-key uniqueness and referential integrity\. What differs is that the baselines must be told the key, whereasGenScriptidentifies it from the profile alone\. On the temporal axis they do separate\. The t\-SNE result \( Figure[3](https://arxiv.org/html/2609.38414#A5.F3)in Appendix[E](https://arxiv.org/html/2609.38414#A5)\) reverses the single\-table picture:GenScript’s points interleave with the real ones throughout the main manifold, whereas both trained baselines form clusters entirely disjoint from it\. Two properties ofaicameramake it well suited to an induced rule set\. The building is an office, so presence follows a strong daily routine\. Andsource\.id\(i\.e\., the time series column inaicamera\) is categorical, so its per\-category frequencies are recorded exactly in the profile and reproduced by construction\.
## 4Conclusion and future work
GenScriptshows that a useful synthetic\-data generator can be obtained without training one\. Across four single\-table benchmarks it is the fastest method on both construction and sampling while staying close to the leaders on marginal fidelity; on a real building\-telemetry time series it matches the real distribution more closely than two specialised temporal models; and it alone emits no rows violating a functional dependency, because the rule is named and enforced rather than fitted\. What it produces is a program a practitioner can read, edit and re\-run\. The limits are equally clear: it is last on Trend on all four datasets, so label\-conditional structure is not yet reproduced\.
Three directions follow\. The constraints we induce are essentially pairwise, since the correlation matrix is the evidence driving them; dependencies over three or more columns are invisible to it and are exactly what wide tables contain, so extending induction beyond pairwise structure is the most direct route to closing the Trend gap\. The profile is also the only input to the semantic phase: a short description of the dataset would let the model propose candidate relationships from domain knowledge, which the verification step of Section[2\.2](https://arxiv.org/html/2609.38414#S2.SS2)can confirm or reject against the real data\. Finally, our relational evaluation covers one two\-table schema; deeper hierarchies, composite keys and cycles remain untested\.
## References
- \[1\]V\. Borisov, K\. Seßler, T\. Leemann, M\. Pawelczyk, and G\. Kasneci\(2023\)Language models are realistic tabular data generators\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.p1.1)\.
- \[2\]DataCebo, Inc\.\(2023\)SDMetrics: metrics for synthetic data evaluation\.Note:Version 0\.9\.0External Links:[Link](https://docs.sdv.dev/sdmetrics/)Cited by:[§3](https://arxiv.org/html/2609.38414#S3.SS0.SSS0.Px2.p1.1)\.
- \[3\]B\. Feuer, Y\. Liu, C\. Hegde, and J\. Freire\(2024\)ArcheType: a novel framework for open\-source column type annotation using large language models\.Proceedings of the VLDB Endowment17\(9\),pp\. 2279–2292\.External Links:[Document](https://dx.doi.org/10.14778/3665844.3665857)Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px3.p1.1)\.
- \[4\]L\. Gao, A\. Madaan, S\. Zhou, U\. Alon, P\. Liu, Y\. Yang, J\. Callan, and G\. Neubig\(2023\)PAL: program\-aided language models\.InInternational Conference on Machine Learning \(ICML\),pp\. 10764–10799\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px2.p1.1)\.
- \[5\]S\. Hegselmann, A\. Buendia, H\. Lang, M\. Agrawal, X\. Jiang, and D\. Sontag\(2023\)TabLLM: few\-shot classification of tabular data with large language models\.InInternational Conference on Artificial Intelligence and Statistics \(AISTATS\),pp\. 5549–5581\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px3.p1.1)\.
- \[6\]A\. Kotelnikov, D\. Baranchuk, I\. Rubachev, and A\. Babenko\(2023\)TabDDPM: modelling tabular data with diffusion models\.InInternational Conference on Machine Learning \(ICML\),pp\. 17564–17579\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.p1.1)\.
- \[7\]J\. Li, B\. Zhao, Z\. Zhao, U\. Javaid, K\. Yee, and B\. Sikdar\(2025\)TabTreeFormer: tabular data generation using hybrid tree\-transformer\.arXiv preprint arXiv:2501\.01216\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.38414#S3.SS0.SSS0.Px1.p1.1)\.
- \[8\]J\. Li, Z\. Zhao, M\. Abdollahzadeh, B\. Sikdar, and Y\.C\. Tay\(2026\)IRG: modular synthetic relational database generation with complex relational schemas\.InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining \(KDD\),External Links:[Document](https://dx.doi.org/10.1145/3770854.3780313)Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1)\.
- \[9\]Z\. Lin, A\. Jain, C\. Wang, G\. Fanti, and V\. Sekar\(2020\)Using GANs for sharing networked time series data: challenges, initial promise, and open questions\.InACM Internet Measurement Conference \(IMC\),pp\. 464–483\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.p1.1)\.
- \[10\]Y\. Long, L\. Xu, and A\. Brintrup\(2025\)LLM\-TabFlow: synthetic tabular data generation with inter\-column logical relationship preservation\.arXiv preprint arXiv:2503\.02161\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1)\.
- \[11\]A\. Narayan, I\. Chami, L\. Orr, and C\. Ré\(2022\)Can foundation models wrangle your data?\.Proceedings of the VLDB Endowment16\(4\),pp\. 738–746\.External Links:[Document](https://dx.doi.org/10.14778/3574245.3574258)Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px3.p1.1)\.
- \[12\]W\. Pang, M\. Shafieinejad, L\. Liu, S\. Hazlewood, and X\. He\(2024\)ClavaDDPM: multi\-relational data synthesis with cluster\-guided diffusion models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1)\.
- \[13\]N\. Patki, R\. Wedge, and K\. Veeramachaneni\(2016\)The synthetic data vault\.InIEEE International Conference on Data Science and Advanced Analytics \(DSAA\),pp\. 399–410\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.p1.1)\.
- \[14\]J\. Shi, M\. Xu, H\. Hua, H\. Zhang, S\. Ermon, and J\. Leskovec\(2025\)TabDiff: a mixed\-type diffusion model for tabular data generation\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.p1.1),[§3](https://arxiv.org/html/2609.38414#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2609.38414#S3.SS0.SSS0.Px2.p1.1)\.
- \[15\]A\. V\. Solatorio and O\. Dupriez\(2023\)REaLTabFormer: generating realistic relational and tabular data using transformers\.arXiv preprint arXiv:2302\.02041\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.p1.1),[§3](https://arxiv.org/html/2609.38414#S3.SS0.SSS0.Px1.p1.1)\.
- \[16\]N\. Suh, Y\. Yang, D\. Hsieh, Q\. Luan, S\. Xu, S\. Zhu, and G\. Cheng\(2025\)TimeAutoDiff: a unified framework for generation, imputation, forecasting, and time\-varying metadata conditioning of heterogeneous time series tabular data\.Transactions on Machine Learning Research \(TMLR\)\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1)\.
- \[17\]P\. Tiwald, I\. Krchova, A\. Sidorenko, M\. Vargas Vieyra, M\. Scriminaci, and M\. Platzer\(2025\)TabularARGN: a flexible and efficient auto\-regressive framework for generating high\-fidelity synthetic data\.arXiv preprint arXiv:2501\.12012\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1)\.
- \[18\]L\. Xu, M\. Skoularidou, A\. Cuesta\-Infante, and K\. Veeramachaneni\(2019\)Modeling tabular data using conditional GAN\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.32\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.p1.1)\.
- \[19\]S\. Yang, Z\. Zhang, B\. Prenkaj, and G\. Kasneci\(2025\)Doubling your data in minutes: ultra\-fast tabular data generation via LLM\-induced dependency graphs\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 10337–10358\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.525)Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1)\.
- \[20\]S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. Cao\(2023\)ReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px2.p1.1)\.
- \[21\]J\. Yoon, D\. Jarrett, and M\. van der Schaar\(2019\)Time\-series generative adversarial networks\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.32\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38414#S1.p1.1)\.
- \[22\]J\. Zhang, G\. Cormode, C\. M\. Procopiuc, D\. Srivastava, and X\. Xiao\(2017\)PrivBayes: private data release via bayesian networks\.ACM Transactions on Database Systems42\(4\),pp\. 1–41\.Cited by:[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2609.38414#A1.SS0.SSS0.Px3.p1.1)\.
- \[23\]Z\. Zhao, R\. Birke, and L\. Y\. Chen\(2023\)FCT\-gan: enhancing global correlation of table synthesis via fourier transform\.InProceedings of the 32nd ACM international conference on information and knowledge management,pp\. 4450–4454\.Cited by:[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1)\.
- \[24\]Z\. Zhao, R\. Birke, and L\. Y\. Chen\(2025\)Tabula: harnessing language models for tabular data synthesis\.InPacific\-Asia Conference on Knowledge Discovery and Data Mining,pp\. 247–259\.Cited by:[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1)\.
- \[25\]Z\. Zhao, A\. Kunar, R\. Birke, and L\. Y\. Chen\(2021\)Ctab\-gan: effective table data synthesizing\.InAsian conference on machine learning,pp\. 97–112\.Cited by:[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1)\.
- \[26\]Z\. Zhao, A\. Kunar, R\. Birke, H\. Van der Scheer, and L\. Y\. Chen\(2024\)Ctab\-gan\+: enhancing tabular data synthesis\.Frontiers in big Data6,pp\. 1296508\.Cited by:[§1](https://arxiv.org/html/2609.38414#S1.SS0.SSS0.Px2.p1.1)\.
## Appendix AExtended related work
#### Related work: single tables\.
CTGAN and TVAE\[[18](https://arxiv.org/html/2609.38414#bib.bib4)\], the diffusion models TabDDPM\[[6](https://arxiv.org/html/2609.38414#bib.bib5)\]and TabDiff\[[14](https://arxiv.org/html/2609.38414#bib.bib11)\], the hybrid tree\-transformer TabTreeFormer\[[7](https://arxiv.org/html/2609.38414#bib.bib12)\], and the language\-model generators GReaT\[[1](https://arxiv.org/html/2609.38414#bib.bib6)\]and REaLTabFormer\[[15](https://arxiv.org/html/2609.38414#bib.bib9)\]all train on the private data\. TabDiff and TabTreeFormer are the closest analogues to our aim, but both unify*within*the single\-table modality and stay inside weight space, whereas we unify*across*modalities by leaving weight space altogether\. PrivBayes\[[22](https://arxiv.org/html/2609.38414#bib.bib3)\]and SDV\[[13](https://arxiv.org/html/2609.38414#bib.bib1)\]share our commitment to an explicit model, but fit their dependency structure and cannot express logical constraints\.
#### Related work: sequential and relational\.
For sequential data, TimeGAN\[[21](https://arxiv.org/html/2609.38414#bib.bib7)\]and DoppelGANger\[[9](https://arxiv.org/html/2609.38414#bib.bib8)\]are adversarial; TabularARGN\[[17](https://arxiv.org/html/2609.38414#bib.bib13)\]trains auto\-regressive conditionals across the column, time, and table axes; and TimeAutoDiff\[[16](https://arxiv.org/html/2609.38414#bib.bib14)\]pairs a VAE with latent diffusion for heterogeneous time\-series tables\. For relational schemas, ClavaDDPM\[[12](https://arxiv.org/html/2609.38414#bib.bib15)\]propagates cluster latent variables across foreign keys, and IRG\[[8](https://arxiv.org/html/2609.38414#bib.bib16)\]generates tables incrementally along a depth\-first traversal to handle composite and overlapping keys\. Each is strong on its home ground, and collectively they make our point: three modalities, three literatures, three training procedures\. Our aim is not to beat any one in its own setting, but to cover all three with one pipeline and no training at all\. Our third stage separately builds on the observation that language models are more reliable emitting programs than answers\[[4](https://arxiv.org/html/2609.38414#bib.bib2),[20](https://arxiv.org/html/2609.38414#bib.bib10)\]: ours is never asked to produce rows, only the sampler that produces them, which bounds its contribution to a short verifiable artifact and makes generation deterministic given a seed\.
#### Related work: LLMs making sense of tabular data\.
A separate line of work uses language models not as density estimators but as readers of a table’s meaning\. TabLLM\[[5](https://arxiv.org/html/2609.38414#bib.bib20)\]shows that serialising rows into text lets a pretrained model classify from a handful of examples, evidence that column names and value formats carry usable semantics on their own\. The data\-management literature has pushed this further:[Narayan et al\. \[11\]](https://arxiv.org/html/2609.38414#bib.bib21)cast entity matching, error detection, and imputation as prompting tasks and find that foundation models reach state\-of\-the\-art without task\-specific training, and ArcheType\[[3](https://arxiv.org/html/2609.38414#bib.bib22)\]performs zero\-shot semantic column\-type annotation, assigning meaning to a column from its name and a sample of its values rather than from a fixed type vocabulary learned in advance\. Our*Analyze*phase is the same operation put to a different end: where that work annotates a column to clean, match, or integrate it, we annotate it in order to*generate*it, and we ask additionally for the constraints that hold*between*columns\. LLM\-TabFlow\[[10](https://arxiv.org/html/2609.38414#bib.bib19)\]uses LLM reasoning to recover inter\-column logical relationships and then delegates density modelling to a score\-based diffusion model, and SPADA\[[19](https://arxiv.org/html/2609.38414#bib.bib18)\]induces a sparse dependency graph with an LLM and synthesises by traversing it with kernel density estimation or a normalising flow, avoiding LLM calls at sampling time entirely\. These are the closest precedents for what we do, and they bracket our position: both extract structure with an LLM and then hand generation to a fitted statistical model, whereas we hand it to a program the LLM writes, so no density is estimated at any point and the extracted structure stays legible in the artifact that generates the data\. Purely statistical generators such as PrivBayes\[[22](https://arxiv.org/html/2609.38414#bib.bib3)\]and SDV\[[13](https://arxiv.org/html/2609.38414#bib.bib1)\]sit at the other end: inspectable and cheap, but with the dependency structure fitted rather than named, and no mechanism for the semantics that column names carry\.
## Appendix BDataset statistics
Table 3:The six source tables\. Columns counts include the target where one exists\. Following common practice for these benchmarks, CoverType, Credit and Intrusion are subsampled to 50 000 rows stratified on the target; Adult is used at its standard training\-split size\. Rows generated equals rows in the source table for every method\.TableDomainRowsCols\.ClassesSourceCoverTypeforest cover50 000557UCI Covertype \(581 012 rows\)Creditcard transactions50 000312Kaggle creditcardfraud \(284 807\)Intrusionnetwork traffic50 0004220UCI KDD Cup 1999Adultcensus income32 561152UCI Adult \(48 842 in full\)accbuilding access1006—industrial partneraicameracamera detections138 3414—industrial partnerThe four public tables are the benchmark suite used by the CTAB\-GAN line of work and its successors, which is why the subsampling protocol follows theirs: 50 000 rows drawn stratified on the target for the three large tables, Adult taken as\-is\. Intrusion’s target is the raw KDD Cup 1999 label, not the coarse five\-way grouping into normal traffic plus four attack families that much of the intrusion\-detection literature uses\. The 10% KDD file carries 23 distinct labels, and stratified subsampling to 50 000 rows drops the rarest of them \(spy,perlandphfeach occur fewer than five times\), leaving the 20 classes we synthesise\. This matters for the MLE numbers in Table[1](https://arxiv.org/html/2609.38414#S3.T1): Intrusion is a 20\-way problem, so its AUC is macro\-averaged over far more classes than CoverType’s seven\. The two partner tables have no target column, so no class count is reported for them\. They are also very differently shaped:accis a small reference table of 100 rows, whileaicameraholds 138 341 detections, an average of roughly 1 400 per registered user\. Itssource\.idcolumn takes 242 distinct values, which is the high\-cardinality categorical whose frequencies the profile records exactly and the generator reproduces by construction \(Section[3](https://arxiv.org/html/2609.38414#S3)\)\.
## Appendix CPer\-dataset fidelity and utility
#### Fidelity\.
On Shape,GenScriptis competitive without being best: 0\.984 on CoverType places it second of four, ahead of both TabDiff and REaLTabFormer and within 0\.7 points of TabTreeFormer, and 0\.936 on Adult and 0\.930 on Credit sit within 5 and 6 points of the leader\. This is a good outcome for a sampler built directly from per\-column marginals\. Intrusion is the hardest case at 0\.874, though TabTreeFormer fares far worse there \(0\.557\), suggesting its high\-cardinality categorical columns are difficult for any method that discretises them\. On Trend,GenScriptis last on all four datasets \(0\.748–0\.897\)\. This is the cost of the design: the sampler reproduces only the dependencies named inΣ\\Sigmaor carried by the proposal’s conditional structure, so pairwise correlations that no induced rule captured are not modelled\. Trend is where a trained generator has the clearest advantage, and our results do not dispute it\.
#### Utility\.
MLE is whereGenScripttrails furthest, and the pattern follows the target’s cardinality\. On the binary\-target datasets it is usable if behind: 0\.894 AUC on Credit and 0\.828 on Adult, against 0\.999 and 0\.925 for the best baseline\. On the multi\-class datasets it degrades sharply—0\.572 on Intrusion and 0\.467 on CoverType, the latter still below chance, so a classifier trained on that output is worse than useless on real data\. No baseline drops below chance on any dataset, so this points at our handling of multi\-class labels rather than at a limit of the approach: the sampler draws the target from its marginal without conditioning on the features that determine it, which costs little when the target is binary and much more when it is not\.
## Appendix DImplementation
Figure 2:TheGenScriptapplication user interface, on a completed Adult run\. The phase bar tracks the pipeline of Section[2](https://arxiv.org/html/2609.38414#S2)\. The cards report what this run recovered: nine induced constraints, two of them enforced structurally, and 30 of 33 validation checks passed\. The generator is exportable, and so auditable and re\-runnable independently of the pipeline that wrote it\.
## Appendix ETime\-series visualisation
Figure 3:t\-SNE ofaicamerasequences embedded bysource\.id\.GenScript\(orange\) lies inside the real manifold \(blue\); TabularARGN \(red\) and TimeAutoDiff \(green\) occupy disjoint regions\.
## Appendix FConstraint preservation on Adult
Table 4:Consistency of theeducation↔\\leftrightarroweducation\-numfunctional dependency on Adult\. The real table maps each of its 16 diploma labels to exactly one year count; a row whose year count does not match its label is a violation\. Percentages are of the 32 561 generated rows\.In Adult,educationandeducation\-numencode the same fact twice—a diploma label and the year count it corresponds to—so the real table contains exactly 16 distinct pairs and the mapping is a functional dependency\. Table[4](https://arxiv.org/html/2609.38414#A6.T4)shows that no trained baseline reproduces it exactly, and that the distinct\-pair count alone is not a sufficient diagnostic: a method can emit the right*number*of pairs and still pair a diploma with the wrong year count, so the row\-level violation rate is the measure that matters\.
The violation rates are small for some baselines and large for others, but the distinction that matters is between zero and non\-zero rather than between rates\. A functional dependency admits no exceptions: a row pairing a diploma with an impossible year count is invalid however rare it is, and rare violations are in some ways worse than frequent ones, because they survive spot\-checks and surface later in whatever downstream job joins or filters on the column\. None of these violations is visible to Shape, Trend, or MLE\.
GenScriptrecovers the dependency in*Analyze*\(Section[2\.2](https://arxiv.org/html/2609.38414#S2.SS2)\)—the correlation is the evidence that a rule exists, the column names supply its form—and materialises it as a hard predicate, so the violation count is zero by construction rather than by filtering afterwards\. This is the concrete payoff of naming constraints instead of fitting them\.相似文章
制作用于微调的合成数据集
作者提出了一种使用形式求解器为LLM生成多样化合成推理训练数据的流程,并询问现有工作以及关于避免重复模板的建议。
想要更好的合成数据?引导它:用于低资源语言生成的激活引导
本文研究了激活引导作为替代少样本提示的方法,用于生成低资源语言的合成数据。作者提出了LanguageSteering和QualitySteering策略,表明在早期层进行引导可以提高数据多样性并改善下游模型性能。
高效数据合成的一个信息论准则
本文从信息论角度解释了合成数据何时会改善或降低LLM训练效果,区分了信息开放和信息封闭的生成循环,并通过数据处理不等式解释了模型崩溃的原因。
GraphGen:利用知识驱动的合成数据生成增强大语言模型的监督微调
GraphGen是一个知识图谱引导的框架,用于生成合成问答数据,以改进大语言模型的监督微调,通过多跳采样和风格控制生成来针对知识缺口。实验表明,它优于传统的合成数据方法。
生成更好训练数据的智能体(25分钟阅读)
Autodata 引入了一种智能体数据科学家,它能够迭代生成并优化合成训练数据,并通过元优化进一步提升数据质量,在计算机科学和法律推理任务上取得了更好的效果。