Tag
The paper introduces a dependency-aware fidelity diagnostic to measure inter-column dependency in synthetic tabular data, revealing that standard metrics are blind to dependency and that current generators have a residual gap not closed by capacity increases or common fixes.
Introduces K-IPO, a generate-then-select oversampling framework that preserves the original data's feature importance ranking (measured by Kendall's tau) during augmentation for imbalanced tabular data, showing improved preservation, explanation consistency, and predictive performance across 20 datasets.