Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models
摘要
This paper shows that matching a marginal Gaussian prior in factorized generative models does not prevent conditional style leakage, where style latents carry class information. Multiple remedies are explored, but the authors conclude that marginal statistics alone cannot certify class-invariance.
查看缓存全文
缓存时间: 2026/08/07 07:49
# Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models
Source: [https://arxiv.org/html/2608.05243](https://arxiv.org/html/2608.05243)
###### Abstract
Factorized generative models regularize a style latentzsz\_\{\\mathrm\{s\}\}toward a fixed prior with a marginal statistic,q\(zs\)≈𝒩\(0,I\)q\(z\_\{\\mathrm\{s\}\}\)\\approx\\mathcal\{N\}\(0,I\), and treat the result as a certificate thatzsz\_\{\\mathrm\{s\}\}carries no class information\. The certificate does not hold\. Matching the marginal constrains nothing about the class conditionalsq\(zs∣y\)q\(z\_\{\\mathrm\{s\}\}\\mid y\), sozsz\_\{\\mathrm\{s\}\}can be exactly Gaussian in aggregate while remaining maximally informative about the label\. We show this gap is one of four terms in an exact decomposition of what a factorized sampler must match, and that eliminating it is necessary but not sufficient for the sampling procedure the factorization exists to support\. Empirically, a case\-study model and four reference latent baselines with near\-zero global MMD all permit a linear probe to recover the label fromzsz\_\{\\mathrm\{s\}\}at7474–100%100\\%\(chance10%10\\%\); our case\-study model reaches99\.15%99\.15\\%clustering accuracy while externally evaluated class\-conditional generation succeeds16%16\\%of the time\. The leakage survives six one\-at\-a\-time perturbations of capacity, curriculum, prior geometry, and supervision on two datasets\. Four remedies attacking the leakage through different mechanisms span2121–46%46\\%probe recovery while leaving the within\-class dependence proxy essentially unchanged\. A post\-hoc conditional prior raises external generated\-class accuracy to0\.970\.97on MNIST without retraining but reaches only0\.410\.41on CIFAR\-10; an empirical style bank reaches0\.880\.88on CIFAR\-10\. The methodological point: no divergence computed onq\(zs\)q\(z\_\{\\mathrm\{s\}\}\)alone can certifyzs⟂yz\_\{\\mathrm\{s\}\}\\perp y, and reporting marginal statistics does not verify the property practitioners claim from them\.
## Introduction
Matching an aggregate posterior to a Gaussian is a statement about a marginal\. Class\-invariant style is a statement about a family of conditionals\. The two are different claims, and the first does not imply the second:q\(zs\)q\(z\_\{\\mathrm\{s\}\}\)can be exactly𝒩\(0,I\)\\mathcal\{N\}\(0,I\)while everyq\(zs∣y=k\)q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\)occupies its own region of latent space, so long as the regions average out\. A regularizer that only ever seesq\(zs\)q\(z\_\{\\mathrm\{s\}\}\)cannot detect this, and therefore cannot prevent it\.
The gap has teeth because of how factorized models generate\. During training the decoder only ever receives\(zc,zs\)\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\)drawn from the same image\. Ifq\(zs∣y=k\)q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\)differs across classes, and nothing in the objective stops it, the decoder is free to learn class\-specific interactions between the two codes\. At generation time a class prior supplieszcz\_\{\\mathrm\{c\}\}for classkkwhilezsz\_\{\\mathrm\{s\}\}comes from the marginal, which averages over all classes\. The resulting pair is one the decoder never saw\. We call the failure*conditional style leakage*\.
What makes it dangerous is that the metrics a practitioner would check are all blind to it\. Reconstruction is measured on the posterior manifold, wherezsz\_\{\\mathrm\{s\}\}comes from a real image and the learned interaction is exactly the one being exercised, so it looks fine\. Global MMD is a marginal statistic by construction\. Clustering accuracy reads onlyzcz\_\{\\mathrm\{c\}\}\. A model can therefore pass every standard diagnostic and still fail completely at the one capability the factorization was built to provide\.
Our starting point is that this is not a defect of any particular loss but a structural property of factorized sampling\. We give an exact decomposition of the mismatch such a sampler incurs \(Theorem[1](https://arxiv.org/html/2608.05243#Thmtheorem1)\) into four non\-negative terms: global style\-prior mismatch, conditional style leakage, semantic\-prior mismatch, and within\-class dependence between the two codes\. Valid sampling requires all four to vanish, so removing the leakage is necessary but not sufficient, and the latent mismatch upper\-bounds the image\-space mismatch but not conversely\. That asymmetry is what makes a decoder\-level intervention necessary rather than redundant\.
We then audit a case\-study model and four reference latent baselines with finite\-sample proxies for the four terms, and use F\-CS\-WAE \(Factorized Class\-Structured Spherical Cauchy WAE\) as a concentrated case study\. We use it to exhibit and repair the failure, not to argue it beats the baselines it is compared against\. The four reference families share one diagnostic architecture and budget, while F\-CS\-WAE is evaluated in its native configuration, so the cross\-model table supports within\-model auditing rather than a ranking\. F\-CS\-WAE is label\-guided by design; the proposed per\-class style MMD reuses the labels already required by its semantic objective and therefore adds no annotation requirement\.
#### Contributions\.
\(1\) An auditing framework for factorized generative models that separates marginal prior fit, conditional leakage, semantic\-prior mismatch, and within\-class dependence, supported by a structural KL decomposition\. \(2\) A conditional audit protocol built from finite\-sample proxies for those terms, applied to a case\-study model and reference latent baselines\. \(3\) A decoder\-level validation study showing when information detected by the audit is actually used during recombination, together with a six\-way robustness check on two datasets, showing no single design choice accounts for the leakage while its magnitude is strongly capacity\-dependent on CIFAR\-10\. \(4\) An evaluation of representation\- and sampling\-level repairs, including a negative transfer result\.
## Related Work
We organize prior work by a single question: does the method certify conditional invariancezs⟂yz\_\{\\mathrm\{s\}\}\\perp y, or only a statistic compatible withzs⟂̸yz\_\{\\mathrm\{s\}\}\\not\\perp y?
VAEs\(Kingma and Welling[2013](https://arxiv.org/html/2608.05243#bib.bib1)\)impose a per\-example KL toward a fixed prior; WAEs\(Tolstikhinet al\.[2017](https://arxiv.org/html/2608.05243#bib.bib2)\)relax this to a divergence between the aggregated posterior and the prior, computed with kernel two\-sample tests\(Grettonet al\.[2012](https://arxiv.org/html/2608.05243#bib.bib9)\)or projected distances\(Kolouriet al\.[2018](https://arxiv.org/html/2608.05243#bib.bib8)\)\. The VAE penalty is not simply an aggregate regularizer in disguise: averaged over the data it splits exactly into𝔼q\(x\)\[KL\(q\(z∣x\)∥p\(z\)\)\]=Iq\(X;Z\)\+KL\(q\(z\)∥p\(z\)\)\\mathbb\{E\}\_\{q\(x\)\}\[\\mathrm\{KL\}\(q\(z\\mid x\)\\\|p\(z\)\)\]=I\_\{q\}\(X;Z\)\+\\mathrm\{KL\}\(q\(z\)\\\|p\(z\)\)\(Hoffman and Johnson[2016](https://arxiv.org/html/2608.05243#bib.bib25)\), an information\-capacity term plus an aggregate\-mismatch term of the kind WAE\-MMD targets directly\. Neither certifiesI\(zs;y\)=0I\(z\_\{\\mathrm\{s\}\};y\)=0\. The capacity term bounds whatZZencodes aboutXXin general, not about the label, and the aggregate term is precisely the marginal statistic Proposition[1](https://arxiv.org/html/2608.05243#Thmproposition1)shows is insufficient\.
β\\beta\-VAE\(Higginset al\.[2017](https://arxiv.org/html/2608.05243#bib.bib19)\),β\\beta\-TCVAE\(Chenet al\.[2018](https://arxiv.org/html/2608.05243#bib.bib18)\), FactorVAE\(Kim and Mnih[2018](https://arxiv.org/html/2608.05243#bib.bib17)\), and DIP\-VAE\(Kumaret al\.[2018](https://arxiv.org/html/2608.05243#bib.bib39)\)penalize total correlation or aggregate moments\. These constrain dependence among latent*dimensions*, not between a designated style code and the label\. More generally, unsupervised disentanglement is not identifiable without inductive biases in the model or data\(Locatelloet al\.[2019](https://arxiv.org/html/2608.05243#bib.bib40)\); weakly supervised pairs can supply such a bias\(Locatelloet al\.[2020](https://arxiv.org/html/2608.05243#bib.bib41)\), but a bias toward factorization is not itself a certificate of the particular invariancezs⟂yz\_\{\\mathrm\{s\}\}\\perp y\. Semi\-supervised VAEs\(Kingmaet al\.[2014](https://arxiv.org/html/2608.05243#bib.bib20)\)and CVAE\(Sohnet al\.[2015](https://arxiv.org/html/2608.05243#bib.bib21)\)condition generation on the label, sidestepping the need forzcz\_\{\\mathrm\{c\}\}to encode identity, but do not by themselves test whether a residual latent remains label\-informative\.
Content–style separation predates deep generative models\(Tenenbaum and Freeman[2000](https://arxiv.org/html/2608.05243#bib.bib30)\)\. Modern split\-latent models use class labels, grouped observations, pairwise similarity, mutual\-information penalties, or latent optimization to divide specified from residual variation\(Mathieuet al\.[2016](https://arxiv.org/html/2608.05243#bib.bib31); Bouchacourtet al\.[2018](https://arxiv.org/html/2608.05243#bib.bib32); Jhaet al\.[2018](https://arxiv.org/html/2608.05243#bib.bib33); Klyset al\.[2018](https://arxiv.org/html/2608.05243#bib.bib34); Zheng and Sun[2019](https://arxiv.org/html/2608.05243#bib.bib35); Ilseet al\.[2020](https://arxiv.org/html/2608.05243#bib.bib36); Gabbay and Hoshen[2020](https://arxiv.org/html/2608.05243#bib.bib37)\)\. These works establish that explicit inductive biases can improve recombination; they also make a plain VAE latent an inappropriate surrogate for a declared class\-invariant style code\. Most directly,Ridgeway and Mozer \([2018](https://arxiv.org/html/2608.05243#bib.bib38)\)introduce leakage filtering, probe class from a style posterior, and evaluate content–style recombination on MNIST and more complex domains\. Our novelty is therefore neither the first observation that content can leak into style nor the algebraic chain rule in isolation\. It is the auditing framework that uses the decomposition to separate four distinct failure sources, connects finite\-sample latent diagnostics to the claim practitioners make from marginal matching, and validates at decoder level whether detected leakage changes recombination and sampling\.
Image\-translation systems operationalize the same recombination requirement: MUNIT\(Huanget al\.[2018](https://arxiv.org/html/2608.05243#bib.bib42)\)and DRIT\(Leeet al\.[2018](https://arxiv.org/html/2608.05243#bib.bib43)\)combine content with sampled style, while DMIT\(Yuet al\.[2019](https://arxiv.org/html/2608.05243#bib.bib44)\)exposes the generator to random cross\-domain combinations during training\. Their notion of domain\-specific style is task\-dependent—the domain may itself be the conditioning attribute—but their training strategies highlight the same train–sample support issue\. Recent methods add other inductive biases: V3 exploits variance–invariance patterns across domains\(Wuet al\.[2025](https://arxiv.org/html/2608.05243#bib.bib45)\), while SCFlow learns invertible merging from combinatorial style–content coverage\(Maet al\.[2025](https://arxiv.org/html/2608.05243#bib.bib46)\)\. These methods address how to learn a separation\. Reference\-guided diffusion work also calls unwanted reference content carried by style features “content leakage” and suppresses it by masking\(Zhuet al\.[2025](https://arxiv.org/html/2608.05243#bib.bib47)\); our leakage is instead label information in a learned style latent\. We ask when such a latent can be sampled independently\.
Several invariance methods directly targetzs⟂̸yz\_\{\\mathrm\{s\}\}\\not\\perp y\. The variational fair autoencoder\(Louizoset al\.[2015](https://arxiv.org/html/2608.05243#bib.bib26)\)matches nuisance\-conditional posteriors to a common target; the per\-class style MMD in Section[Method and Case Study](https://arxiv.org/html/2608.05243#Sx4)is a direct instantiation of that idea, not a new mechanism\. Gradient reversal\(Ganin and Lempitsky[2015](https://arxiv.org/html/2608.05243#bib.bib27)\), Fader Networks\(Lampleet al\.[2017](https://arxiv.org/html/2608.05243#bib.bib28)\), and information\-theoretic invariance\(Moyeret al\.[2018](https://arxiv.org/html/2608.05243#bib.bib29)\)remove specified information through different objectives\. They primarily optimize representation invariance; our decomposition asks the additional question of whether removingzs⟂̸yz\_\{\\mathrm\{s\}\}\\not\\perp ysuffices for the independent sampler\. It does not when the other three terms remain\.
Deep generative clustering \(DEC\(Xieet al\.[2015](https://arxiv.org/html/2608.05243#bib.bib4)\), IDEC\(Guoet al\.[2017](https://arxiv.org/html/2608.05243#bib.bib23)\), VaDE\(Jianget al\.[2016](https://arxiv.org/html/2608.05243#bib.bib3)\)\) evaluates via ACC/NMI/ARI on the cluster variable, a protocol that cannot see this failure because it never inspectszsz\_\{\\mathrm\{s\}\}\. Our case study makes the point concrete:99\.15%99\.15\\%clustering accuracy on the very checkpoint whose class\-conditional generation reaches16%16\\%external Gen\-ACC\. For the semantic variable we use a Spherical Cauchy prior with a Möbius reparameterization\(Sablica and Hornik[2025](https://arxiv.org/html/2608.05243#bib.bib6)\), a heavier\-tailed alternative to the von Mises\-Fisher posterior of S\-VAE\(Davidsonet al\.[2018](https://arxiv.org/html/2608.05243#bib.bib5)\); this is an implementation choice, not a claim about the leakage, which is defined independently of howzcz\_\{\\mathrm\{c\}\}is parameterized\.
## Factorized Sampling and Its Mismatch
Let𝒟=\{\(xi,yi\)\}\\mathcal\{D\}=\\\{\(x\_\{i\},y\_\{i\}\)\\\}haveKKclasses\. A factorized model encodes each input into a semantic variablezcz\_\{\\mathrm\{c\}\}and a style variablezsz\_\{\\mathrm\{s\}\}, and regularizes the latter by
ℒstyle=D\(q\(zs\),𝒩\(0,I\)\),\\mathcal\{L\}\_\{\\mathrm\{style\}\}=D\\bigl\(q\(z\_\{\\mathrm\{s\}\}\),\\,\\mathcal\{N\}\(0,I\)\\bigr\),\(1\)for some aggregate divergenceDD, withq\(zs\)=𝔼y\[q\(zs∣y\)\]q\(z\_\{\\mathrm\{s\}\}\)=\\mathbb\{E\}\_\{y\}\[q\(z\_\{\\mathrm\{s\}\}\\mid y\)\]\. Writep\(zs\):=𝒩\(0,I\)p\(z\_\{\\mathrm\{s\}\}\):=\\mathcal\{N\}\(0,I\)for the style prior\. Together with a class\-conditional semantic priorp\(zc∣y=k\)p\(z\_\{\\mathrm\{c\}\}\\mid y\{=\}k\)these define the*factorized sampler*: drawzc∼p\(zc∣y=k\)z\_\{\\mathrm\{c\}\}\\sim p\(z\_\{\\mathrm\{c\}\}\\mid y\{=\}k\)andzs∼p\(zs\)z\_\{\\mathrm\{s\}\}\\sim p\(z\_\{\\mathrm\{s\}\}\)independently, which is the naive class\-conditional sampling procedure in standard use\.
### Marginal matching certifies nothing
###### Proposition 1\(No certificate\)\.
Letzs∼𝒩\(0,I\)z\_\{\\mathrm\{s\}\}\\sim\\mathcal\{N\}\(0,I\)onℝds\\mathbb\{R\}^\{d\_\{s\}\}and let\{Ak\}k=1K\\\{A\_\{k\}\\\}\_\{k=1\}^\{K\}be any measurable partition withP\(zs∈Ak\)=1/KP\(z\_\{\\mathrm\{s\}\}\\in A\_\{k\}\)=1/K\. Definey=k⇔zs∈Aky=k\\iff z\_\{\\mathrm\{s\}\}\\in A\_\{k\}\. Thenq\(zs\)=𝒩\(0,I\)q\(z\_\{\\mathrm\{s\}\}\)=\\mathcal\{N\}\(0,I\)exactly, soD\(q\(zs\),𝒩\(0,I\)\)=0D\(q\(z\_\{\\mathrm\{s\}\}\),\\mathcal\{N\}\(0,I\)\)=0for*every*divergenceDD; andI\(zs;y\)=H\(y\)=logKI\(z\_\{\\mathrm\{s\}\};y\)=H\(y\)=\\log K, the maximum possible\.
The proof is immediate:yyis defined post hoc as a measurable function ofzsz\_\{\\mathrm\{s\}\}, so the law ofzsz\_\{\\mathrm\{s\}\}is untouched, whileH\(y∣zs\)=0H\(y\\mid z\_\{\\mathrm\{s\}\}\)=0makes the mutual information maximal\. The construction is deliberately adversarial and we do not claim trained decoders produce anything like it\. The point is narrow and worst\-case: no divergence computed onq\(zs\)q\(z\_\{\\mathrm\{s\}\}\)alone, MMD or otherwise, can rule outzsz\_\{\\mathrm\{s\}\}being maximally class\-informative, however the dependence happens to be shaped\.
### What valid factorized sampling requires
Proposition[1](https://arxiv.org/html/2608.05243#Thmproposition1)is the extremal case of term \(2\) below\. The following decomposition covers any encoder, not just that construction\.
###### Theorem 1\(Factorized\-sampling mismatch\)\.
Letyybe uniform on\{1,…,K\}\\\{1,\\dots,K\\\}and defineMfact:=𝔼y\[KL\(q\(zc,zs∣y\)∥p\(zc∣y\)p\(zs\)\)\]M\_\{\\mathrm\{fact\}\}:=\\mathbb\{E\}\_\{y\}\[\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\\mid y\)\\,\\\|\\,p\(z\_\{\\mathrm\{c\}\}\\mid y\)\\,p\(z\_\{\\mathrm\{s\}\}\)\)\]\. Then
Mfact=\\displaystyle M\_\{\\mathrm\{fact\}\}=\\;Iq\(zc;zs∣y\)⏟\(4\) within\-class dep\.\+𝔼y\[KL\(q\(zc∣y\)∥p\(zc∣y\)\)\]⏟\(3\) semantic\-prior mismatch\\displaystyle\\underbrace\{I\_\{q\}\(z\_\{\\mathrm\{c\}\};z\_\{\\mathrm\{s\}\}\\mid y\)\}\_\{\\text\{\(4\) within\-class dep\.\}\}\+\\underbrace\{\\mathbb\{E\}\_\{y\}\[\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{c\}\}\\mid y\)\\\|p\(z\_\{\\mathrm\{c\}\}\\mid y\)\)\]\}\_\{\\text\{\(3\) semantic\-prior mismatch\}\}\+Iq\(zs;y\)⏟\(2\) style leakage\+KL\(q\(zs\)∥p\(zs\)\)⏟\(1\) style\-prior mismatch,\\displaystyle\+\\underbrace\{I\_\{q\}\(z\_\{\\mathrm\{s\}\};y\)\}\_\{\\text\{\(2\) style leakage\}\}\+\\underbrace\{\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{s\}\}\)\\\|p\(z\_\{\\mathrm\{s\}\}\)\)\}\_\{\\text\{\(1\) style\-prior mismatch\}\},and all four terms are non\-negative\.
The proof is a chain\-rule expansion applied twice, given in the supplementary material\. Two consequences matter\.
###### Corollary 1\(Sampling validity\)\.
q\(zc,zs∣y=k\)=p\(zc∣y=k\)p\(zs\)q\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\)=p\(z\_\{\\mathrm\{c\}\}\\mid y\{=\}k\)\\,p\(z\_\{\\mathrm\{s\}\}\)for a\.e\.kkif and only ifMfact=0M\_\{\\mathrm\{fact\}\}=0, i\.e\. iff all four terms vanish\. Class\-invariant style \(term 2\) is thus necessary but not sufficient; it must be accompanied by a matched style prior, a matched semantic prior, and within\-class independence ofzcz\_\{\\mathrm\{c\}\}andzsz\_\{\\mathrm\{s\}\}\.
###### Corollary 2\(Decoder pushforward\)\.
For a decoderDec\\mathrm\{Dec\}with pushforwardDec\#\\mathrm\{Dec\}\_\{\\\#\},𝔼y\[KL\(Dec\#q\(⋅∣y\)∥Dec\#\[p\(zc∣y\)p\(zs\)\]\)\]≤Mfact\\mathbb\{E\}\_\{y\}\[\\mathrm\{KL\}\(\\mathrm\{Dec\}\_\{\\\#\}q\(\\cdot\\mid y\)\\\|\\mathrm\{Dec\}\_\{\\\#\}\[p\(z\_\{\\mathrm\{c\}\}\\mid y\)p\(z\_\{\\mathrm\{s\}\}\)\]\)\]\\leq M\_\{\\mathrm\{fact\}\}by the data\-processing inequality\.
Corollary[2](https://arxiv.org/html/2608.05243#Thmcorollary2)runs one way\. A smallMfactM\_\{\\mathrm\{fact\}\}suffices for the fixed decoder to make posterior\-decoded and sampler\-decoded distributions close, but it is not by itself a certificate that decoded samples match the data distribution\. Conversely, a large value does not force a visible failure, since a decoder can be insensitive to the mismatched directions and map distinct latent distributions to nearly the same images\. Whether a trained decoder actually conditions on the leaked structure is an empirical question that the decomposition leaves open, which is why the intervention in Section[The decoder reads identity off the style code](https://arxiv.org/html/2608.05243#Sx5.SSx2)is needed rather than redundant\.
### Diagnostics
Each term gets a finite\-sample proxy\. We measure the proxies, not the information quantities, and never estimateMfactM\_\{\\mathrm\{fact\}\}as a whole\.Term 1is proxied by global MMD,MMD2\(q\(zs\),𝒩\(0,I\)\)\\mathrm\{MMD\}^\{2\}\(q\(z\_\{\\mathrm\{s\}\}\),\\mathcal\{N\}\(0,I\)\)\.Term 2by inter\-class style separationΔinter=\(K2\)−1∑j<k‖μ¯s\(j\)−μ¯s\(k\)‖2\\Delta\_\{\\mathrm\{inter\}\}=\\binom\{K\}\{2\}^\{\-1\}\\sum\_\{j<k\}\\\|\\bar\{\\mu\}\_\{\\mathrm\{s\}\}^\{\(j\)\}\-\\bar\{\\mu\}\_\{\\mathrm\{s\}\}^\{\(k\)\}\\\|\_\{2\}and by linear\-probe accuracyLP\(zs→y\)\\mathrm\{LP\}\(z\_\{\\mathrm\{s\}\}\\to y\)\.Δinter\\Delta\_\{\\mathrm\{inter\}\}is a mean\-separation statistic and thus a convenient symptom, not what Proposition[1](https://arxiv.org/html/2608.05243#Thmproposition1)is about: that proposition concerns dependence in general and holds even when class means coincide\. An HSIC\-based proxy would also apply here, but our fixed\-bandwidth estimator saturates atzsz\_\{\\mathrm\{s\}\}’s empirical scale, returning a near\-constant value across checkpoints whoseΔinter\\Delta\_\{\\mathrm\{inter\}\}and LP vary widely, so we do not report it\.Term 3is already computed each run asℒclass\\mathcal\{L\}\_\{\\mathrm\{class\}\}\.Term 4has no existing proxy; we compare real within\-class pairs against independently recombined ones,
JointMMD=1K∑kMMD2\(\{\(zci,zsi\)\}yi=k,\{\(zci,zsπ\(i\)\)\}yi=k\),\\begin\{split\}\\mathrm\{JointMMD\}=\\tfrac\{1\}\{K\}\\textstyle\\sum\_\{k\}\\mathrm\{MMD\}^\{2\}\\bigl\(&\\\{\(z\_\{\\mathrm\{c\}\}^\{i\},z\_\{\\mathrm\{s\}\}^\{i\}\)\\\}\_\{y\_\{i\}=k\},\\\\ &\\\{\(z\_\{\\mathrm\{c\}\}^\{i\},z\_\{\\mathrm\{s\}\}^\{\\pi\(i\)\}\)\\\}\_\{y\_\{i\}=k\}\\bigr\),\\end\{split\}\(2\)whereπ\\pipermutes style codes within classkk, breaking the within\-class coupling while leaving both per\-class marginals, and hence terms \(1\)–\(3\), untouched\.
## Method and Case Study
The standard style regularizer, Eq\.[1](https://arxiv.org/html/2608.05243#Sx3.E1), matches only the marginal\. The remedy we study instead matches each class conditional:
ℒstyle\-cls=1K∑k=1KMMDRBF\(\{zsi:yi=k\},\{εj\}\),\\mathcal\{L\}\_\{\\text\{style\-cls\}\}=\\tfrac\{1\}\{K\}\\textstyle\\sum\_\{k=1\}^\{K\}\\mathrm\{MMD\}\_\{\\mathrm\{RBF\}\}\\bigl\(\\\{z\_\{\\mathrm\{s\}\}^\{i\}:y\_\{i\}\{=\}k\\\},\\\{\\varepsilon\_\{j\}\\\}\\bigr\),\(3\)εj∼𝒩\(0,I\)\\varepsilon\_\{j\}\\sim\\mathcal\{N\}\(0,I\)\. Eq\.[1](https://arxiv.org/html/2608.05243#Sx3.E1)has no gradient incentive to reduce inter\-class separation, since the mixture can match𝒩\(0,I\)\\mathcal\{N\}\(0,I\)while every conditional stays class\-structured; Eq\.[3](https://arxiv.org/html/2608.05243#Sx4.E3)penalizes exactly that\. Nothing about it is specific to our architecture: it needs a labeled batch and a style latent\. In F\-CS\-WAE, the semantic MMD and auxiliary classifier already consume the same labels, so adding Eq\.[3](https://arxiv.org/html/2608.05243#Sx4.E3)does not change the supervision regime or annotation budget\. Applied during training to an otherwise unsupervised model, it would require labels or pseudo\-labels; this is distinct from the diagnostics of Section[Diagnostics](https://arxiv.org/html/2608.05243#Sx3.SSx3), which need labels only on the evaluation set and apply post hoc to any trained model\.
F\-CS\-WAE instantiates the setup concretely: a hyperspherical semantic variablezc∈𝕊dc−1z\_\{\\mathrm\{c\}\}\\in\\mathbb\{S\}^\{d\_\{c\}\-1\}aligned to class\-conditional Spherical Cauchy priors, and a Euclidean style variablezs∈ℝdsz\_\{\\mathrm\{s\}\}\\in\\mathbb\{R\}^\{d\_\{s\}\}regularized by Eqs\.[1](https://arxiv.org/html/2608.05243#Sx3.E1)–[3](https://arxiv.org/html/2608.05243#Sx4.E3), with a shared ResNet\-18 trunk feeding two heads and a residual upsampling decoder \(dc=64d\_\{c\}\{=\}64,ds=128d\_\{s\}\{=\}128\)\. The objective adds a supervised semantic MMD, an aggregated semantic MMD, and an auxiliary classifier onμc\\mu\_\{\\mathrm\{c\}\}that prevents semantic collapse; these shapezcz\_\{\\mathrm\{c\}\}and are not part of the leakage remedy\. Coefficients ramp over a four\-phase 300\-epoch curriculum in which per\-class style regularization begins in phase B, after a 50\-epoch reconstruction\-only phase A during which the decoder does see real\(zc,zs\)\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\)pairs unregularized\. Section[Transfer and robustness](https://arxiv.org/html/2608.05243#Sx5.SSx4)tests whether removing that warmup matters; it does not\. Full details are in the supplementary material\.
At generation we either sample naively,zc∼p\(zc∣y=k\)z\_\{\\mathrm\{c\}\}\\sim p\(z\_\{\\mathrm\{c\}\}\\mid y\{=\}k\)andzs∼𝒩\(0,I\)z\_\{\\mathrm\{s\}\}\\sim\\mathcal\{N\}\(0,I\), or apply a post\-hoc class\-conditional style prior: estimateμ¯s\(k\)\\bar\{\\mu\}\_\{\\mathrm\{s\}\}^\{\(k\)\}andσ¯s2\(k\)\\bar\{\\sigma\}\_\{s\}^\{2\(k\)\}from encoded training data and drawzs∼𝒩\(μ¯s\(k\),τ2diag\(σ¯s2\(k\)\)\)z\_\{\\mathrm\{s\}\}\\sim\\mathcal\{N\}\(\\bar\{\\mu\}\_\{\\mathrm\{s\}\}^\{\(k\)\},\\tau^\{2\}\\mathrm\{diag\}\(\\bar\{\\sigma\}\_\{s\}^\{2\(k\)\}\)\)withτ=0\.25\\tau\{=\}0\.25\. The second needs no retraining, only a forward pass\.
## Experiments
CIFAR\-10\(Krizhevsky[2009](https://arxiv.org/html/2608.05243#bib.bib15)\)is the primary clustering/generation benchmark; MNIST\(LeCunet al\.[1998](https://arxiv.org/html/2608.05243#bib.bib14)\)and Fashion\-MNIST\(Xiaoet al\.[2017](https://arxiv.org/html/2608.05243#bib.bib13)\)are used for diagnostics\. We report clustering \(ACC/NMI/ARI\), reconstruction \(SSIM/LPIPS\(Zhanget al\.[2018](https://arxiv.org/html/2608.05243#bib.bib12)\)\), sample quality \(FID\(Heuselet al\.[2017](https://arxiv.org/html/2608.05243#bib.bib11)\), 10k vs 10k\), and generated\-class accuracy \(Gen\-ACC\), the fraction of generated images assigned to the sampled class by an independently trained external classifier\.
Table 1:CIFAR\-10 against baselines matched in backbone \(ResNet\-18\) and latent dimension \(d=192d\{=\}192\), grouped by how the label is used\. F\-CS\-WAE is 3\-seed mean\. The fair clustering comparison is against the label\-guided group, which like F\-CS\-WAE pushes the label into the latent\.Before using F\-CS\-WAE to study leakage we check it is not degenerate, since a degenerate model could leak trivially\. Table[1](https://arxiv.org/html/2608.05243#Sx5.T1)compares it against seven baselines matched in backbone and latent dimension\. Against the label\-guided autoencoders, the fair comparison, F\-CS\-WAE \(80\.9%80\.9\\%\) exceeds the best \(66\.5%66\.5\\%\) and has the lowest FID in the table\. The conditional\-generative baselines cluster near chance because their label enters only at the decoder, leaving the encoder latent unshaped; they are included as the closest generative analogues, not as a clustering comparison\. This is evidence of competitiveness, not a controlled superiority claim: the baselines are single\-seed, F\-CS\-WAE uses both latent\-shaping and a generative prior where each baseline uses one, and we did not equalize per\-method tuning\.
#### Evaluation independence\.
All generation\-side metrics in the main paper are scored directly from generated pixels by a classifier trained only on real images and independent of the model under test: a 4\-layer CNN for MNIST \(99\.4%99\.4\\%real\-test accuracy\) and a WideResNet\-28\-10\(Zagoruyko and Komodakis[2016](https://arxiv.org/html/2608.05243#bib.bib49)\)for CIFAR\-10 \(94\.1%94\.1\\%\)\. Under this fixed protocol, MNIST naive Gen\-ACC is0\.160\.16, the class\-conditional prior reaches0\.970\.97, and latent\-swap style\-following is96\.5%96\.5\\%\. On CIFAR\-10, the class\-conditional diagonal prior atτ=0\.25\\tau\{=\}0\.25reaches0\.410\.41and the empirical style bank reaches0\.880\.88\. Internal\-classifier scores are reported only in the supplementary material as a robustness comparison; they are not mixed with the primary results below\.
### Marginal metrics hide the leakage
Table 2:Conditional style leakage across a case\-study model and four reference latent baselines on MNIST\. Global MMD is near zero everywhere;Δinter\\Delta\_\{\\mathrm\{inter\}\}and LP are not\. Gen\-ACC is scored by the external CNN and uses naivezs∼𝒩\(0,I\)z\_\{\\mathrm\{s\}\}\\sim\\mathcal\{N\}\(0,I\)\. The four reference families share one architecture and budget; F\-CS\-WAE uses its native configuration, so conclusions are within\-row rather than ranked across rows\.Table[2](https://arxiv.org/html/2608.05243#Sx5.T2)applies the diagnostic to four unsupervised baselines and F\-CS\-WAE on MNIST\. The baselines are purpose\-built for this comparison and share the same encoder family, latent dimension, optimizer, and 100\-epoch budget with one another\. F\-CS\-WAE is evaluated in its native configuration\. Accordingly, the table is an audit rather than a leaderboard: for every row individually, global MMD is decoupled fromΔinter\\Delta\_\{\\mathrm\{inter\}\}and LP, and the argument does not depend on ranking leakage severity across model families\.
Every baseline shows LP far above chance despite near\-zero global MMD\. WAE\-MMD has the largest global MMD \(0\.05760\.0576, some3030–300×300\\timesthe others\), suggesting its own regularizer had not converged, and correspondingly the largest baselineΔinter\\Delta\_\{\\mathrm\{inter\}\}and LP\. More telling areβ\\beta\-TCVAE and FactorVAE: their objectives explicitly penalize latent dependence, and they do achieve the smallestΔinter\\Delta\_\{\\mathrm\{inter\}\}of the five, yet LP stays at7474–76%76\\%\. Penalizing total correlation among latent*dimensions*reduces, but does not remove, dependence between the latent and the*label*\. These are different quantities\.
Figure 1:Real image, posterior reconstruction, naive\-prior sample, and class\-conditional\-prior sample, for five digits\. Reconstructions are indistinguishable from the real image \(SSIM0\.9830\.983\); naive\-prior samples are legible digits but frequently the*wrong*one; the class\-conditional prior restores the intended identity\.F\-CS\-WAE without the remedy is the most extreme case \(Δinter=6\.27\\Delta\_\{\\mathrm\{inter\}\}=6\.27, LP=100%=100\\%\), and the consequence is concrete rather than statistical\. On the same checkpoint it reaches99\.15%99\.15\\%clustering accuracy and SSIM0\.9830\.983, yet class\-conditional generation under naive sampling reaches only16%16\\%external Gen\-ACC\. The failure is also structured rather than random: probability mass collapses onto a handful of attractor digits regardless of the class requested, which is what one expects if the decoder resolves an unfamiliar\(zc,zs\)\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\)pair by trusting whichever code it learned to read identity from\. A 2D t\-SNE ofzsz\_\{\\mathrm\{s\}\}makes the leakage visually obvious only for the two largest\-Δinter\\Delta\_\{\\mathrm\{inter\}\}rows; for VAE,β\\beta\-TCVAE, and FactorVAE it is statistically real but invisible in projection, which is a reason to prefer the quantitative diagnostic over a plot\.
### The decoder reads identity off the style code
A probe showszsz\_\{\\mathrm\{s\}\}*contains*class information; it does not show the decoder*uses*it\. For every ordered class pair\(a,b\)\(a,b\)we takeμc\\mu\_\{\\mathrm\{c\}\}from a real test image of classaaandμs\\mu\_\{\\mathrm\{s\}\}from an independent real image of classbb, decode, and classify\. This uses only real, individually\-encoded outputs rather than prior draws, so a failure to recover classaacannot be blamed on an out\-of\-support prior sample, though the pairing itself may lie off the joint support the decoder trained on\.
Figure 2:MNIST latent swap atδ=0\\delta\{=\}0: rows donatezcz\_\{\\mathrm\{c\}\}\(content\) and columns donatezsz\_\{\\mathrm\{s\}\}\(style\)\. The external CNN assigns the output to the style\-donor class in96\.5%96\.5\\%of off\-diagonal swaps\. Internal\-classifier heatmaps for bothδ\\deltasettings are retained only as a supplementary robustness comparison\.Atδ=0\\delta\{=\}0on MNIST, external style\-following is96\.5%96\.5\\%\(Figure[2](https://arxiv.org/html/2608.05243#Sx5.F2)\): the decoder reads identity predominantly fromzsz\_\{\\mathrm\{s\}\}\. Every row of the swap grid \(fixedzcz\_\{\\mathrm\{c\}\}\) looks nearly identical; the column, thezsz\_\{\\mathrm\{s\}\}donor, determines the digit\. The per\-class remedy shifts the qualitative dependence toward content on both datasets\. Exact rates under the model’s internal classifier are provided only in the supplementary robustness comparison and are not used as primary evidence here\.
Per\-class conditional MMD confirms the same structure behindΔinter\\Delta\_\{\\mathrm\{inter\}\}’s single scalar: atδ=0\\delta\{=\}0, mean per\-classMMD2\(q\(zs∣y=k\),𝒩\(0,I\)\)\\mathrm\{MMD\}^\{2\}\(q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\),\\mathcal\{N\}\(0,I\)\)is21×21\\timesthe global MMD on MNIST \(0\.01860\.0186vs0\.000890\.00089\) and15×15\\timeson CIFAR\-10\. Every one of the ten per\-class values sits far above the global figure the marginal regularizer actually optimizes\. Atδ=1\\delta\{=\}1the gap narrows but does not close \(4\.4×4\.4\\times;5\.3×5\.3\\times\)\.
#### Is this an off\-support artifact?
The obvious objection is that cross\-class pairs simply leave the training joint’s support\. Grading the intervention by support answers it\. Same\-class cross\-image pairs sit at1\.2×1\.2\\timesthe same\-image kNN distance to the training joint and preserve identity, whereas cross\-class pairs sit at2\.1×2\.1\\timesand largely lose the content\-donor identity\. Interpolatingzsz\_\{\\mathrm\{s\}\}between a same\-class and a cross\-class donor hands identity over monotonically, crossing atα≈0\.38\\alpha\\approx 0\.38, well before the pair reaches cross\-class support distance\. Style\-dominance turns on as soon aszsz\_\{\\mathrm\{s\}\}carries a competing class signal, not only once the pair leaves the support\. Cross\-class pairs are measurably farther out, so the objection is not fully dissolved, but the effect does not reduce to that distance\.
### What the remedies do and do not fix
Per\-class style MMD works in the intended direction without finishing the job\. It cutsΔinter\\Delta\_\{\\mathrm\{inter\}\}by81%81\\%\(6\.27→1\.226\.27\\to 1\.22\) and LP by57\.457\.4points \(100%→42\.6%100\\%\\to 42\.6\\%\), still far above the10%10\\%floor and higher than three of the four unsupervised baselines\. Theorem[1](https://arxiv.org/html/2608.05243#Thmtheorem1)explains why this term\-2 remedy is necessarily partial: Eq\.[3](https://arxiv.org/html/2608.05243#Sx4.E3)drives terms \(1\)–\(2\) down by construction but places no pressure on term \(4\), the within\-class dependence a decoder can still exploit once everyq\(zs∣y=k\)q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\)individually matches the prior\.
Table 3:Representation\-level remedies on MNIST at matched backbone, dimensions, and training budget\. Generation scores from the model’s internal classifier are moved to the supplementary robustness comparison\. Single\-seed point estimates are reported\.Table 4:Six\-point sweep of the per\-class style MMD weightδ\\delta\. Single\-seed point estimates are reported\.Δinter\\Delta\_\{\\mathrm\{inter\}\}decreases steadily on both datasets, while ACC is non\-monotonic and has one pronounced dip per dataset\.Table[4](https://arxiv.org/html/2608.05243#Sx5.T4)sweeps six values\.Δinter\\Delta\_\{\\mathrm\{inter\}\}decreases steadily on both datasets, so the leakage\-reduction mechanism behaves as expected across the range rather than only at the endpoints\. Naive FID broadly improves but is not monotonic, and clustering ACC has a pronounced single\-seed dip atδ=1\\delta\{=\}1on MNIST andδ=0\.3\\delta\{=\}0\.3on CIFAR\-10\. With one seed per point, we cannot attribute those dips toδ\\deltarather than initialization and make no significance claim across sweep settings\.
Because per\-class MMD only reduces term 2 to42\.6%42\.6\\%LP, it cannot separate “term 2 is still too large” from “term 2 is not the whole story”\. Table[3](https://arxiv.org/html/2608.05243#Sx5.T3)therefore compares four remedies that attack term 2 through different mechanisms, plus one that also targets term 4\. Gradient reversal is much the strongest term\-2 remedy, driving LP to21%21\\%, within1111points of chance and less than half what per\-class MMD achieves\. Yet every term\-2\-only intervention leaves JointMMD in the narrow0\.00410\.0041–0\.00440\.0044range\. Targeting term 4 directly cuts JointMMD by2\.7×2\.7\\times, the only intervention that moves this proxy substantially, while its LP remains higher than gradient reversal’s\. Thus reducingI\(zs;y\)I\(z\_\{\\mathrm\{s\}\};y\)does not by itself reduce within\-classzcz\_\{\\mathrm\{c\}\}–zsz\_\{\\mathrm\{s\}\}dependence, exactly the distinction made by Corollary[1](https://arxiv.org/html/2608.05243#Thmcorollary1)\. Internal\-evaluator generation scores for these checkpoints are retained only as a robustness comparison in the supplementary material\.
We calibrate the term\-4 proxy against a within\-class double\-permutation null using 200 permutations\. Every checkpoint sits44–5×5\\timesabove its null \(p=0\.005p\{=\}0\.005\), so term 4 is genuinely non\-zero, while theδ=0→1\\delta\{=\}0\\to 1change is not significant \(p=0\.31p\{=\}0\.31\): per\-class style MMD leaves term 4 statistically untouched even as it cutsΔinter\\Delta\_\{\\mathrm\{inter\}\}and LP sharply\. A conditional\-HSIC cross\-check agrees\. One comparison we avoid: MNIST and CIFAR\-10 JointMMD values are close, but even after standardization, equal MMD on two different latent distributions does not imply equal dependence, so we draw no cross\-dataset conclusion from it\.
### Transfer and robustness
Table 5:Single\-variable perturbations from theδ=0\\delta\{=\}0checkpoint, each removing one candidate confound, both datasets\. Single\-seed point estimates are reported\.Figure 3:Linear\-probe recovery of the label fromzsz\_\{\\mathrm\{s\}\}for the baseline and six perturbations\. Leakage survives all of them on both datasets, but the two capacity perturbations move CIFAR\-10 far more than they move MNIST, while no\-warmup and the vMF prior barely move either\.Several design choices in F\-CS\-WAE could plausibly be the real explanation for whyzsz\_\{\\mathrm\{s\}\}absorbs so much class information: the style latent being larger than the semantic one, the reconstruction\-only warmup, the prior geometry, or the auxiliary classifier\. Table[5](https://arxiv.org/html/2608.05243#Sx5.T5)and Figure[3](https://arxiv.org/html/2608.05243#Sx5.F3)perturb each independently on both datasets\. LP never drops below92\.6%92\.6\\%on MNIST or44\.0%44\.0\\%on CIFAR\-10, and neither approaches chance\. On MNIST every perturbation is mild, costing at most7\.47\.4points; on CIFAR\-10 the same six span a4444\-point range\. Smaller style capacity is the largest effect on both \(7\.47\.4and42\.942\.9points\), followed by balanced dims, but capacity is not uniquely special: removing the classifier costs19\.819\.8points on CIFAR\-10 and the Gaussian prior12\.412\.4, while no\-warmup and the vMF prior barely move it\.
Two readings follow\. Capacity is a real, dataset\-dependent contributor, consistent with severity tracking how much class structure a dataset forces intozsz\_\{\\mathrm\{s\}\}, and a paper that checked only MNIST would have badly understated how much one hyperparameter can move this number\. We also flag a confound we do not disentangle: the perturbations that reduce LP most also reduce representational capacity, and removing the classifier collapses clustering from99%99\\%to55\.4%55\.4\\%, so part of each drop may reflect a weaker representation rather than genuine invariance\. Even so, no single tested perturbation drives LP near chance, so none of these choices*eliminates*the leakage, even though capacity clearly modulates it\.
Leakage severity also follows the expected dataset ordering\. Measuring intra\-class diversity independently of our model, as mean within\-class pairwise cosine distance in a frozen DINOv2 space, confirms the ordering with clear separation: MNIST0\.330\.33, Fashion\-MNIST0\.470\.47, CIFAR\-100\.660\.66\. BothΔinter\\Delta\_\{\\mathrm\{inter\}\}and LP decrease along it \(MNIST6\.276\.27/100%100\\%, Fashion\-MNIST4\.864\.86/97\.7%97\.7\\%, CIFAR\-103\.803\.80/86\.9%86\.9\\%\)\. We do not use Gen\-ACC as a cross\-dataset proxy for leakage severity: unlikeΔinter\\Delta\_\{\\mathrm\{inter\}\}and LP, it is measured after decoding and therefore also inherits dataset difficulty, sample quality, and external\-classifier error\.
#### The repair does not transfer\.
Table 6:Generation underzsz\_\{\\mathrm\{s\}\}sampling strategies on theδ=0\\delta\{=\}0checkpoints \(zcz\_\{\\mathrm\{c\}\}always from the class prior\)\. Gen\-ACC uses the fixed external CNN on MNIST and WideResNet\-28\-10 on CIFAR\-10\.On MNIST, the external CNN raises Gen\-ACC from0\.160\.16under global Gaussian sampling to0\.970\.97under the class\-conditional diagonal prior atτ=0\.25\\tau\{=\}0\.25\(Table[6](https://arxiv.org/html/2608.05243#Sx5.T6)\)\. On CIFAR\-10, the same parametric recipe reaches only0\.410\.41under the external WideResNet\. The empirical style bank reaches0\.880\.88and also has higher measured diversity \(0\.2410\.241versus0\.1010\.101\), indicating that matching only a unimodal diagonal Gaussian misses important shape inq\(zs∣y=k\)q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\)\. Controls for mean shift, variance shrinkage, and shuffled labels show the same qualitative mechanism under the internal evaluator and are reported only as a supplementary robustness comparison\. The diagnostic therefore transfers across datasets more cleanly than the parametric repair does\.
## Discussion
CIFAR\-10’s higher intra\-class diversity weakens conditional style dependence, andΔinter\\Delta\_\{\\mathrm\{inter\}\}/LP confirm this cleanly, but the practical remedy does not inherit that cleanliness\. CIFAR\-10 samples stay visibly blurry \(FID∼83\\sim 83–8484\), and we avoid calling naive output “visually plausible”; at this FID the accurate description is recognizable but not sharp\. Global aggregate matching is standard in WAEs and disentanglement models, and our results show it is insufficient whenever the style variable must be class\-invariant for generation\. Per\-class style MMD is best read as a supervised analogue of total\-correlation penalties, aimed at the single dependencyzs⟂̸yz\_\{\\mathrm\{s\}\}\\not\\perp yrather than at dependence among latent dimensions in general\. Finally, the MNIST checkpoint reaches SSIM0\.9830\.983yet FID73\.9973\.99under naive sampling: reconstruction measures fidelity on the posterior manifold, not the prior, and evaluation of generative clustering should report both, since the two can diverge substantially\.
#### Limitations\.
We report no component\-wise ablation beyond theδ\\deltaaxis\. The per\-class remedy assumes labels already available to the label\-guided training objective and is not, by itself, an unsupervised disentanglement loss\. Gen\-ACC is confounded by dataset\-intrinsic difficulty\. The CIFAR\-10 transfer failure is measured on one checkpoint per dataset, and we have not tested whether a richer conditional \(a per\-class mixture\) closes it\. The robustness check perturbs one variable at a time, so joint perturbations, which could compound given how much larger each effect is on CIFAR\-10, remain untested\. Table[3](https://arxiv.org/html/2608.05243#Sx5.T3)is MNIST\-only; whether gradient reversal’s term\-2 advantage and the joint remedy’s term\-4 advantage replicate on CIFAR\-10 is untested\. Three datasets remain too few to treat the diversity–leakage relationship as a measured correlation rather than a consistent ordering\.
## Conclusion
Conditional style leakage is a real failure mode in factorized generative models: aggregate regularization of a style variable does not prevent it from becoming class\-dependent, because marginal matching is compatible withzsz\_\{\\mathrm\{s\}\}being maximally informative about the label rather than merely correlated with it\. We placed it inside an exact decomposition of what factorized sampling requires, measured it across a case\-study model and four reference latent baselines, and showed that a decoder trained this way is measurably more sensitive to the leaked code than to the semantic one\. No single perturbation of capacity, curriculum, prior geometry, or supervision eliminates the leakage on either dataset, though capacity strongly modulates its magnitude on CIFAR\-10\. Comparing four mechanistically different term\-2 remedies settles what one remedy could not: they span2121–46%46\\%probe recovery yet leave the term\-4 proxy unchanged, whereas the joint remedy cuts that proxy by2\.7×2\.7\\timesdespite a higher probe score than gradient reversal\. Eliminating conditional style leakage is necessary but not sufficient, exactly as Corollary[1](https://arxiv.org/html/2608.05243#Thmcorollary1)states\. Under the fixed external protocol, the conditional prior reaches0\.970\.97Gen\-ACC on MNIST but only0\.410\.41on CIFAR\-10; an empirical bank reaches0\.880\.88\. Thus fixing a diagnostic finding on one dataset does not fix the underlying generative gap on another\. Factorized generative models should verify class\-invariance of style variables with conditional diagnostics rather than marginal statistics, and should report where their remedies do and do not generalize\.
## References
- D\. Bouchacourt, R\. Tomioka, and S\. Nowozin \(2018\)Multi\-level variational autoencoder: learning disentangled representations from grouped observations\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.32,pp\. 2095–2102\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v32i1.11867)Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p4.1)\.
- R\. T\. Q\. Chen, X\. Li, R\. Grosse, and D\. Duvenaud \(2018\)Isolating sources of disentanglement in variational autoencoders\.arXiv preprint arXiv:1802\.04942\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p3.4)\.
- T\. R\. Davidson, L\. Falorsi, N\. De Cao, T\. Kipf, and J\. M\. Tomczak \(2018\)Hyperspherical variational auto\-encoders\.arXiv preprint arXiv:1804\.00891\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p7.4)\.
- A\. Gabbay and Y\. Hoshen \(2020\)Demystifying inter\-class disentanglement\.InInternational Conference on Learning Representations,Cited by:[J\.3 Explicit split\-latent baselines](https://arxiv.org/html/2608.05243#Sx17.SSx3.p1.2),[Related Work](https://arxiv.org/html/2608.05243#Sx2.p4.1)\.
- Y\. Ganin and V\. Lempitsky \(2015\)Unsupervised domain adaptation by backpropagation\.InInternational Conference on Machine Learning,pp\. 1180–1189\.Cited by:[J\.3 Explicit split\-latent baselines](https://arxiv.org/html/2608.05243#Sx17.SSx3.p1.2),[Related Work](https://arxiv.org/html/2608.05243#Sx2.p6.2)\.
- A\. Gretton, K\. M\. Borgwardt, M\. J\. Rasch, B\. Schölkopf, and A\. Smola \(2012\)A kernel two\-sample test\.Journal of Machine Learning Research13,pp\. 723–773\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p2.4)\.
- X\. Guo, L\. Gao, X\. Liu, and J\. Yin \(2017\)Improved deep embedded clustering with local structure preservation\.InProceedings of the International Joint Conference on Artificial Intelligence,pp\. 1753–1759\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p7.4)\.
- M\. Heusel, H\. Ramsauer, T\. Unterthiner, B\. Nessler, and S\. Hochreiter \(2017\)GANs trained by a two time\-scale update rule converge to a local nash equilibrium\.InAdvances in Neural Information Processing Systems,Cited by:[Experiments](https://arxiv.org/html/2608.05243#Sx5.p1.1)\.
- I\. Higgins, L\. Matthey, A\. Pal, C\. Burgess, X\. Glorot, M\. Botvinick, S\. Mohamed, and A\. Lerchner \(2017\)beta\-VAE: learning basic visual concepts with a constrained variational framework\.International Conference on Learning Representations\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p3.4)\.
- M\. D\. Hoffman and M\. J\. Johnson \(2016\)ELBO surgery: yet another way to carve up the variational evidence lower bound\.InNIPS Workshop on Advances in Approximate Bayesian Inference,Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p2.4)\.
- X\. Huang, M\. Liu, S\. Belongie, and J\. Kautz \(2018\)Multimodal unsupervised image\-to\-image translation\.InProceedings of the European Conference on Computer Vision,pp\. 172–189\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p5.1)\.
- M\. Ilse, J\. M\. Tomczak, C\. Louizos, and M\. Welling \(2020\)DIVA: domain invariant variational autoencoders\.InProceedings of the Third Conference on Medical Imaging with Deep Learning,Proceedings of Machine Learning Research, Vol\.121,pp\. 322–348\.Cited by:[J\.3 Explicit split\-latent baselines](https://arxiv.org/html/2608.05243#Sx17.SSx3.p1.2),[Related Work](https://arxiv.org/html/2608.05243#Sx2.p4.1)\.
- A\. H\. Jha, S\. Anand, M\. Singh, and V\. S\. R\. Veeravasarapu \(2018\)Disentangling factors of variation with cycle\-consistent variational auto\-encoders\.InProceedings of the European Conference on Computer Vision,pp\. 805–820\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p4.1)\.
- Z\. Jiang, Y\. Zheng, H\. Tan, B\. Tang, and H\. Zhou \(2016\)Variational deep embedding: an unsupervised and generative approach to clustering\.arXiv preprint arXiv:1611\.05148\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p7.4)\.
- H\. Kim and A\. Mnih \(2018\)Disentangling by factorising\.arXiv preprint arXiv:1802\.05983\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p3.4)\.
- D\. P\. Kingma, S\. Mohamed, D\. J\. Rezende, and M\. Welling \(2014\)Semi\-supervised learning with deep generative models\.Advances in Neural Information Processing Systems\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p3.4)\.
- D\. P\. Kingma and M\. Welling \(2013\)Auto\-encoding variational bayes\.arXiv preprint arXiv:1312\.6114\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p2.4)\.
- J\. Klys, J\. Snell, and R\. Zemel \(2018\)Learning latent subspaces in variational autoencoders\.InAdvances in Neural Information Processing Systems,Vol\.31\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p4.1)\.
- S\. Kolouri, P\. E\. Pope, C\. E\. Martin, and G\. K\. Rohde \(2018\)Sliced\-wasserstein autoencoder: an embarrassingly simple generative model\.arXiv preprint arXiv:1804\.01947\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p2.4)\.
- A\. Krizhevsky \(2009\)Learning multiple layers of features from tiny images\.Technical reportUniversity of Toronto\.Cited by:[Experiments](https://arxiv.org/html/2608.05243#Sx5.p1.1)\.
- A\. Kumar, P\. Sattigeri, and A\. Balakrishnan \(2018\)Variational inference of disentangled latent concepts from unlabeled observations\.InInternational Conference on Learning Representations,Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p3.4)\.
- G\. Lample, N\. Zeghidour, N\. Usunier, A\. Bordes, L\. Denoyer, and M\. Ranzato \(2017\)Fader networks: manipulating images by sliding attributes\.InAdvances in Neural Information Processing Systems,Cited by:[J\.3 Explicit split\-latent baselines](https://arxiv.org/html/2608.05243#Sx17.SSx3.p1.2),[Related Work](https://arxiv.org/html/2608.05243#Sx2.p6.2)\.
- Y\. LeCun, L\. Bottou, Y\. Bengio, and P\. Haffner \(1998\)Gradient\-based learning applied to document recognition\.Proceedings of the IEEE86\(11\),pp\. 2278–2324\.Cited by:[Experiments](https://arxiv.org/html/2608.05243#Sx5.p1.1)\.
- H\. Lee, H\. Tseng, J\. Huang, M\. Singh, and M\. Yang \(2018\)Diverse image\-to\-image translation via disentangled representations\.InProceedings of the European Conference on Computer Vision,pp\. 35–51\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p5.1)\.
- F\. Locatello, S\. Bauer, M\. Lucic, G\. Raetsch, S\. Gelly, B\. Schölkopf, and O\. Bachem \(2019\)Challenging common assumptions in the unsupervised learning of disentangled representations\.InProceedings of the 36th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.97,pp\. 4114–4124\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p3.4)\.
- F\. Locatello, B\. Poole, G\. Raetsch, B\. Schölkopf, O\. Bachem, and M\. Tschannen \(2020\)Weakly\-supervised disentanglement without compromises\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 6348–6359\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p3.4)\.
- C\. Louizos, K\. Swersky, Y\. Li, M\. Welling, and R\. Zemel \(2015\)The variational fair autoencoder\.arXiv preprint arXiv:1511\.00830\.Cited by:[J\.3 Explicit split\-latent baselines](https://arxiv.org/html/2608.05243#Sx17.SSx3.p1.2),[Related Work](https://arxiv.org/html/2608.05243#Sx2.p6.2)\.
- P\. Ma, X\. Yang, Y\. Li, M\. Gui, F\. Krause, J\. Schusterbauer, and B\. Ommer \(2025\)SCFlow: implicitly learning style and content disentanglement with flow models\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 14919–14929\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p5.1)\.
- M\. F\. Mathieu, J\. J\. Zhao, A\. Ramesh, P\. Sprechmann, and Y\. LeCun \(2016\)Disentangling factors of variation in deep representation using adversarial training\.InAdvances in Neural Information Processing Systems,Vol\.29\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p4.1)\.
- D\. Moyer, S\. Gao, R\. Brekelmans, A\. Galstyan, and G\. Ver Steeg \(2018\)Invariant representations without adversarial training\.InAdvances in Neural Information Processing Systems,pp\. 9102–9111\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p6.2)\.
- K\. Ridgeway and M\. C\. Mozer \(2018\)Open\-ended content\-style recombination via leakage filtering\.arXiv preprint arXiv:1810\.00110\.Cited by:[J\.3 Explicit split\-latent baselines](https://arxiv.org/html/2608.05243#Sx17.SSx3.p1.2),[Related Work](https://arxiv.org/html/2608.05243#Sx2.p4.1)\.
- L\. Sablica and K\. Hornik \(2025\)Hyperspherical variational autoencoders using efficient spherical cauchy distribution\.arXiv preprint arXiv:2506\.21278\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p7.4)\.
- K\. Sohn, H\. Lee, and X\. Yan \(2015\)Learning structured output representation using deep conditional generative models\.Advances in Neural Information Processing Systems\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p3.4)\.
- C\. Szegedy, V\. Vanhoucke, S\. Ioffe, J\. Shlens, and Z\. Wojna \(2016\)Rethinking the inception architecture for computer vision\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,pp\. 2818–2826\.Cited by:[K\.4 FID: feature extractor, resizing, normalization](https://arxiv.org/html/2608.05243#Sx18.SSx4.p1.3)\.
- J\. B\. Tenenbaum and W\. T\. Freeman \(2000\)Separating style and content with bilinear models\.Neural Computation12\(6\),pp\. 1247–1283\.External Links:[Document](https://dx.doi.org/10.1162/089976600300015349)Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p4.1)\.
- I\. Tolstikhin, O\. Bousquet, S\. Gelly, and B\. Schölkopf \(2017\)Wasserstein auto\-encoders\.arXiv preprint arXiv:1711\.01558\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p2.4)\.
- Y\. Wu, Z\. Wang, B\. Raj, and G\. Xia \(2025\)Unsupervised disentanglement of content and style via variance\-invariance constraints\.InInternational Conference on Learning Representations,Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p5.1)\.
- H\. Xiao, K\. Rasul, and R\. Vollgraf \(2017\)Fashion\-MNIST: a novel image dataset for benchmarking machine learning algorithms\.arXiv preprint arXiv:1708\.07747\.Cited by:[Experiments](https://arxiv.org/html/2608.05243#Sx5.p1.1)\.
- J\. Xie, R\. Girshick, and A\. Farhadi \(2015\)Unsupervised deep embedding for clustering analysis\.arXiv preprint arXiv:1511\.06335\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p7.4)\.
- X\. Yu, Y\. Chen, T\. Li, S\. Liu, and G\. Li \(2019\)Multi\-mapping image\-to\-image translation via learning disentanglement\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p5.1)\.
- S\. Zagoruyko and N\. Komodakis \(2016\)Wide residual networks\.InProceedings of the British Machine Vision Conference,Cited by:[K\.6 External classifier: architecture and training protocol](https://arxiv.org/html/2608.05243#Sx18.SSx6.p3.4),[Evaluation independence\.](https://arxiv.org/html/2608.05243#Sx5.SSx3.SSS0.Px1.p1.8)\.
- R\. Zhang, P\. Isola, A\. A\. Efros, E\. Shechtman, and O\. Wang \(2018\)The unreasonable effectiveness of deep features as a perceptual metric\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,Cited by:[Experiments](https://arxiv.org/html/2608.05243#Sx5.p1.1)\.
- Z\. Zheng and L\. Sun \(2019\)Disentangling latent space for VAE by label relevant/irrelevant dimensions\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 12192–12201\.Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p4.1)\.
- L\. Zhu, X\. Wang, C\. Zhou, Q\. Gu, and N\. Ye \(2025\)Less is more: masking elements in image condition features avoids content leakages in style transfer diffusion models\.InInternational Conference on Learning Representations,Cited by:[Related Work](https://arxiv.org/html/2608.05243#Sx2.p5.1)\.
Supplementary Material Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models
## A\. Proofs
### A\.1 Proposition 1 \(No certificate\)
###### Proposition 2\.
Letzs∼𝒩\(0,I\)z\_\{\\mathrm\{s\}\}\\sim\\mathcal\{N\}\(0,I\)onℝds\\mathbb\{R\}^\{d\_\{s\}\}, and let\{Ak\}k=1K\\\{A\_\{k\}\\\}\_\{k=1\}^\{K\}be any measurable partition ofℝds\\mathbb\{R\}^\{d\_\{s\}\}withP\(zs∈Ak\)=1/KP\(z\_\{\\mathrm\{s\}\}\\in A\_\{k\}\)=1/Kfor eachkk\(such a partition always exists, e\.g\. via level sets of a continuous statistic ofzsz\_\{\\mathrm\{s\}\}\)\. Definey=k⇔zs∈Aky=k\\iff z\_\{\\mathrm\{s\}\}\\in A\_\{k\}\. Then \(i\)q\(zs\)=𝒩\(0,I\)q\(z\_\{\\mathrm\{s\}\}\)=\\mathcal\{N\}\(0,I\)exactly, soD\(q\(zs\),𝒩\(0,I\)\)=0D\(q\(z\_\{\\mathrm\{s\}\}\),\\mathcal\{N\}\(0,I\)\)=0for every divergenceDD; and \(ii\)I\(zs;y\)=H\(y\)=logKI\(z\_\{\\mathrm\{s\}\};y\)=H\(y\)=\\log K\.
###### Proof\.
\(i\)yyis defined post hoc as a measurable function ofzsz\_\{\\mathrm\{s\}\}; the distribution ofzsz\_\{\\mathrm\{s\}\}itself is never altered, soq\(zs\)=𝒩\(0,I\)q\(z\_\{\\mathrm\{s\}\}\)=\\mathcal\{N\}\(0,I\)holds exactly, and every divergence to𝒩\(0,I\)\\mathcal\{N\}\(0,I\), aggregate MMD included, is identically zero\. \(ii\) Sinceyyis a deterministic function ofzsz\_\{\\mathrm\{s\}\},H\(y∣zs\)=0H\(y\\mid z\_\{\\mathrm\{s\}\}\)=0, soI\(zs;y\)=H\(y\)−H\(y∣zs\)=H\(y\)I\(z\_\{\\mathrm\{s\}\};y\)=H\(y\)\-H\(y\\mid z\_\{\\mathrm\{s\}\}\)=H\(y\)\. By constructionP\(y=k\)=1/KP\(y=k\)=1/Kfor allkk, soH\(y\)=logKH\(y\)=\\log K, which is also the maximum possible entropy, and hence the maximum possible mutual information with any other variable, for aKK\-ary label\. ∎
The construction is deliberately adversarial:\{Ak\}\\\{A\_\{k\}\\\}need not resemble anything a trained decoder produces, andq\(zs∣y=k\)q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\)here iszsz\_\{\\mathrm\{s\}\}restricted toAkA\_\{k\}and renormalized, not a smooth family with class\-specific means\. That is the point\. It shows that no divergence computed onq\(zs\)q\(z\_\{\\mathrm\{s\}\}\)alone can rule outzsz\_\{\\mathrm\{s\}\}being maximally class\-informative, regardless of how the dependence is structured geometrically\. It is a worst\-case non\-identifiability result, not a claim about typical behavior\.
### A\.2 Theorem 1 \(Factorized\-sampling mismatch\)
###### Theorem 2\.
Letyybe uniform on\{1,…,K\}\\\{1,\\dots,K\\\}, letq\(zc,zs∣y=k\)q\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\)be the encoder’s joint class conditional, letp\(zc∣y=k\)p\(z\_\{\\mathrm\{c\}\}\\mid y\{=\}k\)andp\(zs\)p\(z\_\{\\mathrm\{s\}\}\)be the semantic and style priors, and letq\(zs\)=𝔼y\[q\(zs∣y\)\]q\(z\_\{\\mathrm\{s\}\}\)=\\mathbb\{E\}\_\{y\}\[q\(z\_\{\\mathrm\{s\}\}\\mid y\)\]\. Define
Mfact:=𝔼y\[KL\(q\(zc,zs∣y\)∥p\(zc∣y\)p\(zs\)\)\]\.M\_\{\\mathrm\{fact\}\}:=\\mathbb\{E\}\_\{y\}\\bigl\[\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\\mid y\)\\,\\\|\\,p\(z\_\{\\mathrm\{c\}\}\\mid y\)\\,p\(z\_\{\\mathrm\{s\}\}\)\)\\bigr\]\.Then
Mfact=\\displaystyle M\_\{\\mathrm\{fact\}\}=\\;Iq\(zc;zs∣y\)\+𝔼y\[KL\(q\(zc∣y\)∥p\(zc∣y\)\)\]\\displaystyle I\_\{q\}\(z\_\{\\mathrm\{c\}\};z\_\{\\mathrm\{s\}\}\\mid y\)\+\\mathbb\{E\}\_\{y\}\[\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{c\}\}\\mid y\)\\\|p\(z\_\{\\mathrm\{c\}\}\\mid y\)\)\]\+Iq\(zs;y\)\+KL\(q\(zs\)∥p\(zs\)\),\\displaystyle\+I\_\{q\}\(z\_\{\\mathrm\{s\}\};y\)\+\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{s\}\}\)\\\|p\(z\_\{\\mathrm\{s\}\}\)\),and all four terms are non\-negative\.
###### Proof\.
Fixkkand expand the class\-kkterm by inserting±logq\(zc∣k\)±logq\(zs∣k\)\\pm\\log q\(z\_\{\\mathrm\{c\}\}\\mid k\)\\pm\\log q\(z\_\{\\mathrm\{s\}\}\\mid k\)into the integrand:
KL\(q\(zc,zs∣k\)∥p\(zc∣k\)p\(zs\)\)\\displaystyle\\mathrm\{KL\}\\bigl\(q\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\\mid k\)\\,\\\|\\,p\(z\_\{\\mathrm\{c\}\}\\mid k\)\\,p\(z\_\{\\mathrm\{s\}\}\)\\bigr\)=𝔼q\(zc,zs∣k\)\[logq\(zc,zs∣k\)q\(zc∣k\)q\(zs∣k\)\]\\displaystyle=\\mathbb\{E\}\_\{q\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\\mid k\)\}\\Bigl\[\\log\\tfrac\{q\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\\mid k\)\}\{q\(z\_\{\\mathrm\{c\}\}\\mid k\)\\,q\(z\_\{\\mathrm\{s\}\}\\mid k\)\}\\Bigr\]\+𝔼q\(zc∣k\)\[logq\(zc∣k\)p\(zc∣k\)\]\+𝔼q\(zs∣k\)\[logq\(zs∣k\)p\(zs\)\]\\displaystyle\\quad\+\\mathbb\{E\}\_\{q\(z\_\{\\mathrm\{c\}\}\\mid k\)\}\\Bigl\[\\log\\tfrac\{q\(z\_\{\\mathrm\{c\}\}\\mid k\)\}\{p\(z\_\{\\mathrm\{c\}\}\\mid k\)\}\\Bigr\]\+\\mathbb\{E\}\_\{q\(z\_\{\\mathrm\{s\}\}\\mid k\)\}\\Bigl\[\\log\\tfrac\{q\(z\_\{\\mathrm\{s\}\}\\mid k\)\}\{p\(z\_\{\\mathrm\{s\}\}\)\}\\Bigr\]=Iq\(zc;zs∣y=k\)\+KL\(q\(zc∣k\)∥p\(zc∣k\)\)\\displaystyle=I\_\{q\}\(z\_\{\\mathrm\{c\}\};z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\)\+\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{c\}\}\\mid k\)\\\|p\(z\_\{\\mathrm\{c\}\}\\mid k\)\)\+KL\(q\(zs∣k\)∥p\(zs\)\),\\displaystyle\\quad\+\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{s\}\}\\mid k\)\\\|p\(z\_\{\\mathrm\{s\}\}\)\),where the first term is, by definition, the mutual information betweenzcz\_\{\\mathrm\{c\}\}andzsz\_\{\\mathrm\{s\}\}underq\(⋅,⋅∣y=k\)q\(\\cdot,\\cdot\\mid y\{=\}k\)\. Taking𝔼y\[⋅\]\\mathbb\{E\}\_\{y\}\[\\cdot\]gives the first two terms of the statement directly, since𝔼y\[Iq\(zc;zs∣y=k\)\]=Iq\(zc;zs∣y\)\\mathbb\{E\}\_\{y\}\[I\_\{q\}\(z\_\{\\mathrm\{c\}\};z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\)\]=I\_\{q\}\(z\_\{\\mathrm\{c\}\};z\_\{\\mathrm\{s\}\}\\mid y\)by definition of conditional mutual information\. For the remaining piece,𝔼y\[KL\(q\(zs∣y\)∥p\(zs\)\)\]\\mathbb\{E\}\_\{y\}\[\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{s\}\}\\mid y\)\\\|p\(z\_\{\\mathrm\{s\}\}\)\)\], insert±logq\(zs\)\\pm\\log q\(z\_\{\\mathrm\{s\}\}\)and useq\(zs\)=𝔼y\[q\(zs∣y\)\]q\(z\_\{\\mathrm\{s\}\}\)=\\mathbb\{E\}\_\{y\}\[q\(z\_\{\\mathrm\{s\}\}\\mid y\)\]:
𝔼y\[KL\(q\(zs∣y\)∥p\(zs\)\)\]=\\displaystyle\\mathbb\{E\}\_\{y\}\\bigl\[\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{s\}\}\\mid y\)\\\|p\(z\_\{\\mathrm\{s\}\}\)\)\\bigr\]=\{\}𝔼y\[KL\(q\(zs∣y\)∥q\(zs\)\)\]⏟=Iq\(zs;y\)\\displaystyle\\underbrace\{\\mathbb\{E\}\_\{y\}\\bigl\[\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{s\}\}\\mid y\)\\\|q\(z\_\{\\mathrm\{s\}\}\)\)\\bigr\]\}\_\{=\\,I\_\{q\}\(z\_\{\\mathrm\{s\}\};y\)\}\+KL\(q\(zs\)∥p\(zs\)\)\.\\displaystyle\+\\mathrm\{KL\}\(q\(z\_\{\\mathrm\{s\}\}\)\\\|p\(z\_\{\\mathrm\{s\}\}\)\)\.where the first bracket is exactlyIq\(zs;y\)I\_\{q\}\(z\_\{\\mathrm\{s\}\};y\): the average KL of a conditional from its own marginal is, by definition, the mutual information\. Non\-negativity of every term follows from non\-negativity of KL divergence and of mutual information \(itself a KL divergence\)\. ∎
### A\.3 Corollaries
###### Corollary 3\(Exact sampling validity\)\.
q\(zc,zs∣y=k\)=p\(zc∣y=k\)p\(zs\)q\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\)=p\(z\_\{\\mathrm\{c\}\}\\mid y\{=\}k\)\\,p\(z\_\{\\mathrm\{s\}\}\)for almost everykkif and only ifMfact=0M\_\{\\mathrm\{fact\}\}=0, i\.e\. iff all four terms vanish simultaneously\. In particular, class\-invariant style \(Iq\(zs;y\)=0I\_\{q\}\(z\_\{\\mathrm\{s\}\};y\)=0\) is necessary but not sufficient\.
###### Proof\.
Every term of Theorem 1 is a non\-negative expectation of a KL divergence \(terms 2 and 4 are mutual informations, themselves KL divergences\), so their sumMfactM\_\{\\mathrm\{fact\}\}is zero iff each term is zero; and𝔼y\[KL\(⋅∥⋅\)\]=0\\mathbb\{E\}\_\{y\}\[\\mathrm\{KL\}\(\\cdot\\\|\\cdot\)\]=0iff the integrand vanishes for a\.e\.yy, iff the two per\-class distributions coincide a\.e\. ∎
###### Corollary 4\(Decoder pushforward bound\)\.
LetDec\\mathrm\{Dec\}be the decoder andDec\#μ\\mathrm\{Dec\}\_\{\\\#\}\\muthe law ofDec\(zc,zs\)\\mathrm\{Dec\}\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\)when\(zc,zs\)∼μ\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\)\\sim\\mu\. Then
𝔼y\[KL\(Dec\#q\(⋅∣y\)∥Dec\#\[p\(zc∣y\)p\(zs\)\]\)\]≤Mfact\.\\mathbb\{E\}\_\{y\}\\bigl\[\\mathrm\{KL\}\(\\mathrm\{Dec\}\_\{\\\#\}q\(\\cdot\\mid y\)\\,\\\|\\,\\mathrm\{Dec\}\_\{\\\#\}\[p\(z\_\{\\mathrm\{c\}\}\\mid y\)\\,p\(z\_\{\\mathrm\{s\}\}\)\]\)\\bigr\]\\leq M\_\{\\mathrm\{fact\}\}\.
###### Proof\.
The data\-processing inequality states that pushing two distributions through the same Markov kernel \(hereDec\\mathrm\{Dec\}\) cannot increase their KL divergence; applying it inside𝔼y\[⋅\]\\mathbb\{E\}\_\{y\}\[\\cdot\]and invoking Theorem 1 gives the bound\. ∎
This runs one way only\. A smallMfactM\_\{\\mathrm\{fact\}\}suffices for the fixed decoder to make posterior\-decoded and sampler\-decoded class\-conditional distributions close, but it is not by itself a certificate that decoded samples match the data distribution\. Conversely, a large value does not force a visible failure, because a decoder can be insensitive to the mismatched latent directions and map distinct latent distributions to nearly identical image distributions\. Whether the trained decoder actually conditions on the leaked structure that terms \(2\) and \(4\) quantify is an empirical question the decomposition does not settle, which is why the latent\-swap intervention in the main paper is necessary rather than redundant\.
## B\. Architecture Details
The semantic head produces\(μ~c,rc\)\(\\tilde\{\\mu\}\_\{\\mathrm\{c\}\},r\_\{c\}\), givingμc=μ~c/‖μ~c‖2\\mu\_\{\\mathrm\{c\}\}=\\tilde\{\\mu\}\_\{\\mathrm\{c\}\}/\\\|\\tilde\{\\mu\}\_\{\\mathrm\{c\}\}\\\|\_\{2\}and concentrationρc=σ\(rc\)\(1−ε\)\\rho\_\{c\}=\\sigma\(r\_\{c\}\)\(1\-\\varepsilon\); the semantic latent is sampled by the Spherical Cauchy Möbius reparameterizationηc∼Unif\(𝕊dc−1\)\\eta\_\{c\}\\sim\\mathrm\{Unif\}\(\\mathbb\{S\}^\{d\_\{c\}\-1\}\),zc=Tμc,ρc\(ηc\)z\_\{\\mathrm\{c\}\}=T\_\{\\mu\_\{\\mathrm\{c\}\},\\rho\_\{c\}\}\(\\eta\_\{c\}\)\. The style head produces\(μs,logσs2\)\(\\mu\_\{\\mathrm\{s\}\},\\log\\sigma\_\{s\}^\{2\}\)andzs=μs\+σs⊙εsz\_\{\\mathrm\{s\}\}=\\mu\_\{\\mathrm\{s\}\}\+\\sigma\_\{s\}\\odot\\varepsilon\_\{s\},εs∼𝒩\(0,I\)\\varepsilon\_\{s\}\\sim\\mathcal\{N\}\(0,I\)\.
The ResNet\-18 backbone operates on32×3232\\times 32inputs\. It replaces the first7×77\\times 7convolution \(stride 2\) with a3×33\\times 3convolution \(stride 1, padding 1\) and removes the max\-pool layer, yielding a4×44\\times 4spatial output with 512 channels; a shared FC layer maps512×4×4=8,192512\\times 4\\times 4\{=\}8\{,\}192to 256 dimensions with SiLU, from which separate linear heads produce the semantic and style parameters\. The decoder concatenates\[zc;zs\]\[z\_\{\\mathrm\{c\}\};z\_\{\\mathrm\{s\}\}\]\(dc=64d\_\{c\}\{=\}64,ds=128d\_\{s\}\{=\}128in all experiments unless stated\) and maps it through three ResBlockUp stages \(bilinear2×2\\timesupsample,1×11\\times 1projection, GroupNorm\+SiLU\+Conv residual block;512→256→128→64512\{\\to\}256\{\\to\}128\{\\to\}64channels,4×4→32×324\\times 4\\to 32\\times 32\), followed by a3×33\\times 3output convolution and sigmoid\.
For each classkk, F\-CS\-WAE maintains a centermk∈𝕊dc−1m\_\{k\}\\in\\mathbb\{S\}^\{d\_\{c\}\-1\}with priorp\(zc∣y=k\)=SCauchy\(mk,ρp\)p\(z\_\{\\mathrm\{c\}\}\\mid y\{=\}k\)=\\mathrm\{SCauchy\}\(m\_\{k\},\\rho\_\{p\}\),ρp=0\.7\\rho\_\{p\}\{=\}0\.7, updated by EMA from batch class\-mean directions once per epoch,mk←normalize\(τmk\+\(1−τ\)μ¯c\(k\)\)m\_\{k\}\\leftarrow\\mathrm\{normalize\}\(\\tau m\_\{k\}\+\(1\-\\tau\)\\bar\{\\mu\}\_\{\\mathrm\{c\}\}^\{\(k\)\}\),τ=0\.95\\tau\{=\}0\.95\. The full training objective is
ℒ\\displaystyle\\mathcal\{L\}=ℒrec\+α\(t\)ℒclass\+β\(t\)ℒagg\+γ\(t\)ℒstyle\\displaystyle=\\mathcal\{L\}\_\{\\mathrm\{rec\}\}\+\\alpha\(t\)\\mathcal\{L\}\_\{\\mathrm\{class\}\}\+\\beta\(t\)\\mathcal\{L\}\_\{\\mathrm\{agg\}\}\+\\gamma\(t\)\\mathcal\{L\}\_\{\\mathrm\{style\}\}\+δ\(t\)ℒstyle\-cls\+η\(t\)ℒcls,\\displaystyle\\quad\+\\delta\(t\)\\mathcal\{L\}\_\{\\mathrm\{style\\text\{\-\}cls\}\}\+\\eta\(t\)\\mathcal\{L\}\_\{\\mathrm\{cls\}\},withℒrec=λ1‖x−x^‖1\+λLPIPSLPIPS\(x,x^\)\\mathcal\{L\}\_\{\\mathrm\{rec\}\}=\\lambda\_\{1\}\\\|x\-\\hat\{x\}\\\|\_\{1\}\+\\lambda\_\{\\mathrm\{LPIPS\}\}\\mathrm\{LPIPS\}\(x,\\hat\{x\}\),ℒclass=1K∑kMMD2\(Qk,Pk\)\\mathcal\{L\}\_\{\\mathrm\{class\}\}=\\frac\{1\}\{K\}\\sum\_\{k\}\\mathrm\{MMD\}^\{2\}\(Q\_\{k\},P\_\{k\}\)withQk=\{zci:yi=k\}Q\_\{k\}=\\\{z\_\{\\mathrm\{c\}\}^\{i\}\{:\}y\_\{i\}\{=\}k\\\}andPk∼p\(zc∣y=k\)P\_\{k\}\\sim p\(z\_\{\\mathrm\{c\}\}\\mid y\{=\}k\),ℒagg=MMD2\(\{zci\},\{zc,rjp\}\)\\mathcal\{L\}\_\{\\mathrm\{agg\}\}=\\mathrm\{MMD\}^\{2\}\(\\\{z\_\{\\mathrm\{c\}\}^\{i\}\\\},\\\{z\_\{c,r\_\{j\}\}^\{p\}\\\}\)against a random\-class prior mixture, andℒcls=CE\(h\(μc\),y\)\\mathcal\{L\}\_\{\\mathrm\{cls\}\}=\\mathrm\{CE\}\(h\(\\mu\_\{\\mathrm\{c\}\}\),y\)a linear auxiliary classifier preventing semantic collapse\.
## C\. Training Hyperparameters and Schedule
Table 7:Hyperparameters used in all F\-CS\-WAE experiments\.Table 8:Training phase boundaries and active loss terms; ramps are linear\.Per\-class style regularization ramps up in Phase B, beforeℒclass\\mathcal\{L\}\_\{\\mathrm\{class\}\}andℒagg\\mathcal\{L\}\_\{\\mathrm\{agg\}\}are introduced in Phase C, so that it begins as early as possible relative to the semantic MMD terms, limiting rather than eliminating the epochs during which class\-specific\(zc,zs\)\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\)interactions can form without style\-conditional counter\-pressure\. Phase A itself \(epochs 0–49\) is reconstruction\-only, so the decoder does see real\(zc,zs\)\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\)pairs before any style or class regularization is active\. The no\-warmup perturbation reported in the main paper tests whether removing this phase changes the leakage and finds that it does not\.
The eight baselines in the main paper’s CIFAR\-10 competence table—ResNetAE plus seven label\-using methods—use the same ResNet\-18 backbone, total latent dimension \(d=192d\{=\}192\), 300\-epoch budget, and optimizer settings as F\-CS\-WAE\. The four MNIST cross\-model diagnostic controls \(VAE, WAE\-MMD,β\\beta\-TCVAE, and FactorVAE\) instead share a lightweight 3\-layer CNN encoder, latent dimension 64, and a 100\-epoch budget\. The two groups serve different comparisons and are not ranked against one another\.
## D\. Diagnostic Protocol
All leakage diagnostics usen=2048n\{=\}2048held\-out test samples with a fixed seed, so repeated invocations across checkpoints evaluate the same subset rather than adding sampling noise on top of single\-seed training runs\. The Euclidean MMD estimator uses the average of a multi\-scale RBF bandwidth ladderσ∈\{0\.5,1,2,5,10,20,50\}\\sigma\\in\\\{0\.5,1,2,5,10,20,50\\\}; a single median\-heuristic bandwidth saturates atzsz\_\{\\mathrm\{s\}\}’s empirical scale and returns near\-zero regardless of distributional mismatch\. The same saturation is why we do not report the HSIC proxy: with a single fixed bandwidth it returns a near\-constant value across checkpoints whoseΔinter\\Delta\_\{\\mathrm\{inter\}\}and LP vary by factors of five\.
JointMMD, the within\-class dependence diagnostic defined in the main paper, uses a product kernel: the multi\-scale spherical RBF above forzcz\_\{\\mathrm\{c\}\}and the multi\-scale Euclidean RBF forzsz\_\{\\mathrm\{s\}\}, with per\-block standardization\. The permutation null independently permutes both sample sets within each class, so any residual value reflects estimator bias rather than dependence; we report200200permutations per checkpoint across 3 seeds\. With the standard plus\-one correction, the minimum attainable value is1/\(200\+1\)≈0\.0051/\(200\+1\)\\approx 0\.005, so null\-exceedance cases are reported asp=0\.005p\{=\}0\.005\.
The linear probe is a single linear layer trained on the encodedμs\\mu\_\{\\mathrm\{s\}\}of the evaluation split with an 80/20 train/validation split, Adam at10−210^\{\-2\}for 100 epochs; we report validation accuracy\. Because the probe is linear, its validation accuracy is a conservative proxy for predictability within the chosen classifier family; it is not presented as a direct numerical lower bound on mutual information\.
The latent\-swap intervention drawsn=100n\{=\}100pairs per ordered class pair \(9090off\-diagonal pairs forK=10K\{=\}10\), decodesDec\(μc,μs\)\\mathrm\{Dec\}\(\\mu\_\{\\mathrm\{c\}\},\\mu\_\{\\mathrm\{s\}\}\)using posterior means rather than samples to remove reparameterization noise, and re\-encodes the result\. We report the diagonal\-excluded means ofP\(pred=a\)P\(\\mathrm\{pred\}\{=\}a\)\(content\-following\) andP\(pred=b\)P\(\\mathrm\{pred\}\{=\}b\)\(style\-following\); these do not sum to one, and the residual is the mass assigned to the other eight classes\.
## E\. Extended Results
### E\.1 Fullδ\\deltasweep
Table[9](https://arxiv.org/html/2608.05243#Sx12.T9)gives the complete six\-point sweep including reconstruction and sample\-quality columns omitted from the main paper\. Naive FID improves overall but not monotonically withδ\\deltaon MNIST, from74\.074\.0atδ=0\\delta\{=\}0to41\.541\.5atδ=3\\delta\{=\}3, with a temporary rise from52\.652\.6to59\.459\.4atδ=0\.3\\delta\{=\}0\.3\. The remedy makes naive prior samples*better*even though it does not make them correct\. The two quantities are not the same thing, and conflating them is precisely the error the paper is about\. Reconstruction is essentially flat across the sweep \(SSIM0\.9770\.977–0\.9830\.983on MNIST\), so the remedy is not purchasing leakage reduction with reconstruction quality\.
Table 9:Fullδ\\deltasweep, single seed per point \(measured\)\.Note that theΔinter\\Delta\_\{\\mathrm\{inter\}\}values here are computed on theδ\\delta\-sweep checkpoints and differ in the third significant figure from the values in the main paper’s cross\-model diagnostic table, which come from separately trainedδ=0\\delta\{=\}0/δ=1\\delta\{=\}1checkpoints used for the cross\-model comparison \(6\.216\.21vs6\.276\.27atδ=0\\delta\{=\}0; both round to1\.221\.22atδ=1\\delta\{=\}1\)\. This is ordinary run\-to\-run variation between independently trained models at the same setting and illustrates why the single\-seed caveat matters\.
Figure 4:Clustering ACC against linear\-probe leakage across the six\-pointδ\\deltasweep, both datasets \(single seed per point; the connecting line is a visual aid, not a claim of a smooth underlying curve\)\. Leakage \(thexx\-axis\) moves smoothly withδ\\delta; ACC does not\.Figure 5:Naive\-sampling FID againstΔinter\\Delta\_\{\\mathrm\{inter\}\}across the same sweep\.Δinter\\Delta\_\{\\mathrm\{inter\}\}falls steadily, whereas FID improves overall but non\-monotonically; lower leakage tends to accompany better naive sample quality without making sampling*correct*\.
### E\.2 Cross\-dataset leakage
Table 10:Leakage across three datasets atδ=0\\delta\{=\}0\(measured\), ordered by measured intra\-class diversity\.Figure 6:Left:Δinter\\Delta\_\{\\mathrm\{inter\}\}and linear\-probe accuracy across the three datasets, ordered by intra\-class visual diversity; both decrease monotonically\. Right: naive\-Gaussian generation self\-accuracy does*not*follow the same order \(CIFAR\-10 is lowest, not highest\), because it additionally inherits dataset\-intrinsic classification difficulty; see the main paper’s discussion\.Fashion\-MNIST atδ=0\\delta\{=\}0reaches ACC0\.93410\.9341, NMI0\.86720\.8672, ARI0\.86330\.8633, and FID72\.3772\.37; its leakage metrics lie between MNIST and CIFAR\-10, while its FID is slightly lower than MNIST’s \(Figure[6](https://arxiv.org/html/2608.05243#Sx12.F6)\)\. On MNIST the single\-latent CS\-WAE variant \(one Spherical Cauchy semantic latent, no style variable\) reaches ACC0\.87350\.8735, NMI0\.88630\.8863, ARI0\.84110\.8411, FID26\.9526\.95\. It is worth noting explicitly that this ablation has*no style latent and therefore cannot leak by construction*, and it also achieves the best FID of any model we trained\. That is not an argument for removing the style variable, since the factorization exists to provide independent style control, but it does bound how much the factorization is buying on this dataset\.
### E\.3 Per\-class conditional MMD
The single scalarΔinter\\Delta\_\{\\mathrm\{inter\}\}compresses aKK\-way structure\. ComputingMMD2\(q\(zs∣y=k\),𝒩\(0,I\)\)\\mathrm\{MMD\}^\{2\}\(q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\),\\mathcal\{N\}\(0,I\)\)per class atδ=0\\delta\{=\}0gives a mean of0\.01860\.0186on MNIST against a global MMD of0\.000890\.00089, a factor of2121; on CIFAR\-10,0\.00850\.0085against0\.000560\.00056, a factor of1515\. Every one of the ten per\-class values on both datasets sits above the global figure, so this is not a few confusable classes dragging an average: every class contributes \(Figure[7](https://arxiv.org/html/2608.05243#Sx12.F7)\)\. Atδ=1\\delta\{=\}1the ratios fall to4\.4×4\.4\\times\(MNIST\) and5\.3×5\.3\\times\(CIFAR\-10\) without reaching one\. The pairwise structureMMD2\(q\(zs∣y=i\),q\(zs∣y=j\)\)\\mathrm\{MMD\}^\{2\}\(q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}i\),q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}j\)\)likewise shows broad off\-diagonal separation on both datasets rather than a few isolated pairs \(Figure[8](https://arxiv.org/html/2608.05243#Sx12.F8)\)\.


Figure 7:Per\-classMMD2\(q\(zs∣y=k\),𝒩\(0,I\)\)\\mathrm\{MMD\}^\{2\}\(q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}k\),\\mathcal\{N\}\(0,I\)\)atδ=0\\delta\{=\}0\(bars\) against the global MMD the marginal regularizer actually optimizes \(dashed line\)\. Left: MNIST; right: CIFAR\-10\. Every class sits far above the line the model was trained to minimize\.

Figure 8:PairwiseMMD2\(q\(zs∣y=i\),q\(zs∣y=j\)\)\\mathrm\{MMD\}^\{2\}\(q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}i\),q\(z\_\{\\mathrm\{s\}\}\\mid y\{=\}j\)\)between all class pairs atδ=0\\delta\{=\}0\(left MNIST, right CIFAR\-10\)\. Broad off\-diagonal structure, not a few isolated confusable pairs\.Figure 9:Per\-class histograms ofzsz\_\{\\mathrm\{s\}\}projected onto the top principal component of the class\-mean matrix \(MNIST,δ=0\\delta\{=\}0\)\. Some classes separate cleanly along this single direction even where a 2D t\-SNE does not make the separation visible, illustrating that a fixed low\-dimensional projection can either overstate or understate leakage\.
## F\. Failure\-Mode Analysis
The naive\-sampling failure is systematic rather than random\. Confusion between the sampled classy=ky\{=\}kand the class the classifier assigns to the generated image \(Figure[10](https://arxiv.org/html/2608.05243#Sx13.F10)\) shows probability mass collapsing onto a small set of attractor digits, notably88, largely independent of which class was requested\. This is what one expects if the decoder resolves an out\-of\-training\-distribution\(zc,zs\)\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\)pair by falling back on whichever class its style input most resembles, and it explains why model\-internal self\-accuracy sits near0\.150\.15rather than near the0\.100\.10that uniform random errors would produce: the errors are concentrated, not spread\. The main paper reports the corresponding external Gen\-ACC\.
Figure 10:Model\-internal robustness check: confusion between the sampled class and the class assigned after re\-encoding the generated image\. Naive Gaussian sampling \(left\) collapses onto attractor digits \(notably 8\), while the class\-conditional prior \(right,τ=0\.25\\tau\{=\}0\.25\) is nearly diagonal\. The main paper uses an independent external classifier for primary Gen\-ACC\.The latent\-swap grid makes the same point at the pixel level \(Figure[11](https://arxiv.org/html/2608.05243#Sx13.F11)\)\. Atδ=0\\delta\{=\}0every row \(fixedzcz\_\{\\mathrm\{c\}\}donor\) produces visually identical output; only the column \(zsz\_\{\\mathrm\{s\}\}donor\) determines the digit\. On CIFAR\-10 the rows are more similar to each other than in the MNIST grid but still show column\-driven drift in object identity for several classes, consistent with a decoder that was never as exclusively dependent onzsz\_\{\\mathrm\{s\}\}\. Under the model\-internal robustness evaluator, the CIFAR\-10 swap heatmaps \(Figure[12](https://arxiv.org/html/2608.05243#Sx13.F12)\) show the sameδ=0→δ=1\\delta\{=\}0\\to\\delta\{=\}1flip as MNIST but weaker in magnitude \(content\-following4\.2%→20\.0%4\.2\\%\\to 20\.0\\%, style\-following68\.8%→38\.5%68\.8\\%\\to 38\.5\\%\)\.


Figure 11:Latent\-swap grids atδ=0\\delta\{=\}0: row=zc=z\_\{\\mathrm\{c\}\}donor class, column=zs=z\_\{\\mathrm\{s\}\}donor class\. Left, MNIST: every row is visually identical; identity is set entirely by the column\. Right, CIFAR\-10: weaker but still visible column\-driven drift\.



Figure 12:Model\-internal robustness check: CIFAR\-10 latent\-swap heatmaps \(left content\-following, right style\-following; topδ=0\\delta\{=\}0, bottomδ=1\\delta\{=\}1\)\. Primary swap rates in the main paper use the external classifier\.A 2D t\-SNE ofzsz\_\{\\mathrm\{s\}\}colored by class, across all six rows of the cross\-model table \(Figure[13](https://arxiv.org/html/2608.05243#Sx13.F13)\), is worth reading with a caution attached: unlike a typical clustering figure,*color\-separated clusters here mean worse*\. Separation is visually obvious only for WAE\-MMD and F\-CS\-WAE without the remedy, the two largest\-Δinter\\Delta\_\{\\mathrm\{inter\}\}rows, and visibly reduced after adding per\-class style MMD\. For VAE,β\\beta\-TCVAE, and FactorVAE the leakage is statistically real \(LP7474–87%87\\%\) but not visually apparent in the projection\. Projectingzsz\_\{\\mathrm\{s\}\}onto the top principal component of theK×dsK\\times d\_\{s\}class\-mean matrix, the same signalΔinter\\Delta\_\{\\mathrm\{inter\}\}summarizes, separates some classes cleanly even where t\-SNE does not \(Figure[9](https://arxiv.org/html/2608.05243#Sx12.F9)\)\. A fixed low\-dimensional projection can therefore both overstate and understate leakage depending on which directions it happens to preserve, which is the argument for the quantitative diagnostics over any single plot\.
Figure 13:t\-SNE of the style latentzsz\_\{\\mathrm\{s\}\}, colored by MNIST digit class, for all six rows of the cross\-model table\. Color\-separated clusters here mean*worse*\(class information leaking intozsz\_\{\\mathrm\{s\}\}\): separation is visually obvious only for WAE\-MMD and F\-CS\-WAE without the remedy, and visibly reduced after adding per\-class style MMD; for the other three families the leakage is statistically real but invisible in this projection\.
## G\. Sampling Strategies: Qualitative Results
Figure 14:Model\-internal robustness check for MNIST sampling strategies\. Internal self\-accuracy jumps from near\-chance to100%100\\%whenzsz\_\{\\mathrm\{s\}\}is conditioned on class, while diversity is lowest for the class\-mean point and highest for the empirical style bank\. External Gen\-ACC is reported in the main paper\.



Figure 15:MNIST generation under four style\-sampling strategies \(left to right\): global Gaussian, class mean, class\-conditional diagonalτ=0\.25\\tau\{=\}0\.25, and empirical bank\. Under the model\-internal robustness evaluator their self\-accuracies are15%15\\%,100%100\\%,100%100\\%, and100%100\\%; these are not the external Gen\-ACC values used in the main paper\.Figure 16:CIFAR\-10 naive \(odd rows\) vs\. class\-conditionalτ=0\.25\\tau\{=\}0\.25\(even rows\) samples, four classes\. Diversity within a row collapses underτ=0\.25\\tau\{=\}0\.25, and identity is often unclear by eye even where the model\-internal evaluator scores it correctly \(0\.6810\.681\)\. The corresponding primary external Gen\-ACC in the main paper is0\.410\.41\.Figure 17:Model\-internal robustness check, the CIFAR\-10 analogue of Figure[14](https://arxiv.org/html/2608.05243#Sx14.F14): class\-conditional Gaussian strategies plateau near70%70\\%, while the empirical bank is higher\. Primary external Gen\-ACC is reported separately in the main paper\.Figures[14](https://arxiv.org/html/2608.05243#Sx14.F14)–[17](https://arxiv.org/html/2608.05243#Sx14.F17)are model\-internal robustness counterparts to the main paper’s externally evaluated sampling table\. Two details are visible in the grids that the numbers compress away\. First, the MNIST global\-Gaussian samples are not noise: they are clean, legible digits of the*wrong*class, which is exactly the signature of a decoder resolving an unfamiliar\(zc,zs\)\(z\_\{\\mathrm\{c\}\},z\_\{\\mathrm\{s\}\}\)pair by trusting the style code\. Second, the CIFAR\-10τ=0\.25\\tau\{=\}0\.25rows show within\-class collapse: the variance shrinkage that is harmless on MNIST \(whose per\-class style distributions are effectively unimodal\) visibly suppresses diversity on CIFAR\-10, and the empirical bank restores it, which is the qualitative counterpart of its0\.2410\.241diversity score\.
## H\. Case\-Study Reference Results
Figure 18:CIFAR\-10 clustering and FID across three F\-CS\-WAE seeds\. Variance is low \(ACC std0\.520\.52points\), which is why the main paper reports the 3\-seed mean for F\-CS\-WAE while the baselines remain single\-seed\.Figure 19:CIFAR\-10 class\-conditional prior samples \(one row per class\)\. FID8383places these in the recognizable\-but\-blurry regime typical of L1/LPIPS\-trained decoders without an adversarial loss, well short of photorealistic\.Figures[18](https://arxiv.org/html/2608.05243#Sx15.F18)–[19](https://arxiv.org/html/2608.05243#Sx15.F19)document the case\-study model’s baseline competence: stable clustering across seeds and recognizable class\-conditional samples\. These support the main paper’s non\-degeneracy claim and nothing stronger\.
## I\. Baseline Implementation
The four cross\-model diagnostic baselines \(VAE, WAE\-MMD,β\\beta\-TCVAE, FactorVAE\) share a lightweight 3\-layer convolutional encoder with latent dimension 64, trained 100 epochs with Adam at10−310^\{\-3\}on MNIST\.β\\beta\-TCVAE uses a minibatch\-weighted\-sampling estimate of the total correlation term; FactorVAE uses a density\-ratio discriminator on permuted latent dimensions\. These are purpose\-built for the diagnostic and are not tuned to compete on generation quality, which is why we draw only within\-model conclusions from them\.
The eight baselines in the main paper’s CIFAR\-10 competence comparison share F\-CS\-WAE’s ResNet\-18 backbone and total latent dimension \(d=192d\{=\}192\), trained 300 epochs under the same optimizer settings\. Seven use labels; ResNetAE is the unsupervised reference\. The label\-using methods divide by how the label enters\.*Label\-guided*models put the label on the latent: AEWithCE adds a cross\-entropy head, AEWithSupCon a supervised contrastive loss on L2\-normalized projections, AEWithCenterLoss a learnable per\-class center pull, AEWithTriplet an online hard\-mined triplet margin loss\.*Conditional\-generative*models put the label only on the decoder: ConditionalVAE concatenates a label embedding to the decoder input, ConditionalWAE\-MMD does the same with an MMD rather than a KL penalty, and GaussianClassPriorWAE matches per\-class Gaussian priors in a Euclidean latent\. The clustering gap between the two groups \(up to66\.5%66\.5\\%versus near\-chance\) is a direct consequence of that difference: a model whose label never touches the encoder has no reason to produce a class\-structured latent, and does not\.
## J\. Experiments
### J\.1 Five\-seed remedy replication
Table 11:Measured mean±\\pmstandard deviation over five seeds\.
### J\.2 Five\-seed sampling replication
Table 12:Measured external Gen\-ACC \(mean±\\pmstandard deviation\) over five seeds\.
### J\.3 Explicit split\-latent baselines
Table 13:Measured ranges for explicit content–style baselines on MNIST and CIFAR\-10\. Each range is the minimum–maximum across the named methods and datasets, not an uncertainty interval\.The rows instantiate GRL/Fader\(Ganin and Lempitsky[2015](https://arxiv.org/html/2608.05243#bib.bib27); Lampleet al\.[2017](https://arxiv.org/html/2608.05243#bib.bib28)\), VFAE/DIVA\(Louizoset al\.[2015](https://arxiv.org/html/2608.05243#bib.bib26); Ilseet al\.[2020](https://arxiv.org/html/2608.05243#bib.bib36)\), and LORD/leakage filtering\(Gabbay and Hoshen[2020](https://arxiv.org/html/2608.05243#bib.bib37); Ridgeway and Mozer[2018](https://arxiv.org/html/2608.05243#bib.bib38)\)\. Each implementation uses the common ResNet\-18 backbone, total latent dimensiond=192d\{=\}192, 300\-epoch optimizer schedule, and seed 0 used for the matched CIFAR\-10 comparison, while retaining its method\-specific objective\. LP is evaluated on the fixedn=2048n\{=\}2048subset and naive Gen\-ACC with the external classifiers of Section K\.6\. The range in each row pools the two named implementations across MNIST and CIFAR\-10; it is included as a compact robustness summary and is not used for cross\-family ranking\.
### J\.4 Global\-MMD calibration
For each checkpoint, pool encoded style samples with an equal\-size Gaussian reference, recompute the unbiased MMD after 1,000 random label permutations, and bootstrap the observed MMD over evaluation examples\. The resulting estimate is MNIST MMD0\.00130\.0013with a bootstrap 95% interval\[0\.0008,0\.0019\]\[0\.0008,0\.0019\]and permutationp≈0\.10p\\approx 0\.10, and CIFAR\-10 MMD0\.00160\.0016with interval\[0\.0010,0\.0024\]\[0\.0010,0\.0024\]andp≈0\.07p\\approx 0\.07\. The substantive claim does not depend on nonsignificance: even if the marginal mismatch is statistically detectable, its scale and its test do not certifyzs⟂yz\_\{s\}\\perp y\. We report the null quantiles, effect size, interval, sample size, kernel, and bandwidth ladder together\.
## K\. Compute and Reproducibility
### K\.1 Compute environment
A single 300\-epoch F\-CS\-WAE run takes roughly4\.54\.5–1313hours wall\-clock on one NVIDIA A30 \(24GB\), the spread driven by contention from other jobs on the same node rather than by the configuration; PyTorch 2\.1, CUDA 12\.1, mixed precision disabled \(all reported numbers use full FP32, since AMP changedΔinter\\Delta\_\{\\mathrm\{inter\}\}by more than run\-to\-run noise in preliminary testing on MNIST\)\. Diagnostics run in seconds to low minutes on a saved checkpoint \(FID and the JointMMD permutation null are the slowest, at roughly one and three minutes respectively\)\. The core single\-seed and case\-study results correspond to approximately3030training runs \(sixδ\\delta\-sweep points×\\timestwo datasets, three datasets atδ=0\\delta\{=\}0in Table[10](https://arxiv.org/html/2608.05243#Sx12.T10), one single\-latent ablation, four cross\-model diagnostic baselines, seven matched\-supervision baselines, and three additional F\-CS\-WAE seeds on CIFAR\-10 for the case study\)\. The five\-seed replications and explicit split\-latent baseline evaluations in Section J are additional to this core count\. The external classifiers of Section K\.6 are each trained once per dataset\.
### K\.2 Seed count and seed IDs, table by table
Unless stated otherwise, all training uses PyTorch’s global seed \(torch\.manual\_seed\), which also seeds NumPy and Python’srandomvia the standard PyTorch Lightning/utility seeding hook, and all reported single\-seed runs use seed0\. Table[14](https://arxiv.org/html/2608.05243#Sx18.T14)states, for every checkpoint\-based table and figure with a quantitative claim, how many independent training seeds it aggregates and which seed IDs\. The ranges in Table[13](https://arxiv.org/html/2608.05243#Sx17.T13)aggregate distinct configurations and datasets, each run with seed 0, rather than repeated seeds\.*No table in this paper computes error bars from fewer than the seed count listed*; single\-seed entries report the point estimate from that one training run and no variance, which is the honest reading of them\.
Table 14:Seed count and seed IDs by table/figure\. “1 \(checkpoint\-specific\)” means the point comes from a specific, separately\-trained checkpoint \(e\.g\. theδ=0\\delta\{=\}0/δ=1\\delta\{=\}1pair used for cross\-model comparison\), so distinct rows at the same nominalδ\\deltaacross tables are not the same weights and can differ in the third significant figure \(Section E\.1\)\.Table / FigureSeedsSeed IDsMain matched\-supervision baselines10Main F\-CS\-WAE competence rows30, 1, 2Main cross\-model diagnostic \(δ=0/1\\delta\{=\}0/1\)1 \(checkpoint\-specific\)0Main paper invariance remedies10Main paperδ\\deltasweep10Main paper robustness perturbations10Main paper sampling strategies10Supp\. Table[11](https://arxiv.org/html/2608.05243#Sx17.T11)\(remedies\)50, 1, 2, 3, 4Supp\. Table[12](https://arxiv.org/html/2608.05243#Sx17.T12)\(sampling\)50, 1, 2, 3, 4Supp\. Table[13](https://arxiv.org/html/2608.05243#Sx17.T13)\(each configuration\)10Supp\. Table[9](https://arxiv.org/html/2608.05243#Sx12.T9)\(δ\\deltasweep\)10Supp\. Table[10](https://arxiv.org/html/2608.05243#Sx12.T10)\(cross\-dataset\)10Fig\.[18](https://arxiv.org/html/2608.05243#Sx15.F18)\(case study\)30, 1, 2JointMMD permutation null \(Sec\. D\)30, 1, 2Per\-class/pairwise conditional MMD \(Sec\. E\.3\)10Latent\-swap grids and heatmaps \(Sec\. F\)10t\-SNE \(Fig\.[13](https://arxiv.org/html/2608.05243#Sx13.F13)\)10The 3\- and 5\-seed entries report mean±\\pmsample standard deviation across their respective runs; Figure[18](https://arxiv.org/html/2608.05243#Sx15.F18)’s caption states the resulting ACC standard deviation \(0\.520\.52points\) explicitly as the justification for treating F\-CS\-WAE as the only row with multi\-seed variance reported in the main text\. We flag this asymmetry rather than resolve it: the baselines, invariance remedies,δ\\deltasweep, and robustness perturbations are single\-seed, so a between\-configuration comparison where one side has three points and the other has one should be read as indicative, not as a statistically tested difference\. We do not run significance tests across single\-seed points for that reason\.
### K\.3 Dataset preprocessing and augmentation
MNIST and Fashion\-MNIST images are resized from their native28×2828\\times 28resolution to32×3232\\times 32by bilinear interpolation before entering the model; they remain single\-channel because the modified ResNet\-18 first convolution accepts one channel\. This matches the decoder’s32×3232\\times 32output and makes the pixelwise reconstruction loss well defined\. CIFAR\-10 images remain at their native32×32×332\\times 32\\times 3resolution\. All three datasets use the standard train/test split as distributed bytorchvision\.datasets\(MNIST: 60,000/10,000; Fashion\-MNIST: 60,000/10,000; CIFAR\-10: 50,000/10,000\); then=2048n\{=\}2048evaluation subset used throughout Section D is drawn once, with a fixed seed, from the test split of each dataset, and reused across checkpoints so that every diagnostic is evaluated on the same 2048 held\-out images regardless of which model produced them\.
Pixel values are scaled to\[0,1\]\[0,1\]\(min\-max, not per\-channel standardization\), matching the sigmoid output of the decoder and theL1L\_\{1\}reconstruction loss\.*No data augmentation is applied when training F\-CS\-WAE or any of the matched\-supervision/cross\-model baselines*: the reconstruction loss requires an exact pixel\-wise correspondence between the encoder input and reconstruction target\. Although applying a shared transformation to both would preserve that correspondence, we omit augmentation to isolate the leakage phenomenon from any confound introduced by augmentation\-induced invariances\. The CIFAR\-10 external classifier of Section K\.6 uses augmentation; the MNIST and Fashion\-MNIST external classifiers do not\.
### K\.4 FID: feature extractor, resizing, normalization
FID is computed with the standard Inception\-v3 network\(Szegedyet al\.[2016](https://arxiv.org/html/2608.05243#bib.bib48)\)pretrained on ImageNet, using activations from the final average\-pooling layer \(2048\-dimensionalpool3features\), via thepytorch\-fidreference implementation and its bundled Inception weights, which is the same feature extractor used by the great majority of FID numbers reported in the generative\-modeling literature and is necessary for our numbers to be comparable to others’\. Images are converted to 3\-channel \(MNIST/Fashion\-MNIST grayscale is replicated across channels for this step only, not for training\), resized from the model resolution to299×299299\\times 299with bilinear interpolation, and normalized to the Inception\-v3 input range the reference implementation expects \(\[−1,1\]\[\-1,1\], per\-channel\)\. Thus both real and generated MNIST/Fashion\-MNIST images pass through the same32×3232\\times 32model resolution before the FID resize\. The real\-image reference distribution is the full test split \(10,000 images per dataset\); the generated distribution is 10,000 samples drawn under the sampling strategy named in each table \(naive prior, class\-conditional, empirical bank, etc\.\)\. FID is the squared Fréchet distance between Gaussians fit to the two 2048\-dimensional feature sets,
FID=‖μr−μg‖22\+Tr\(Σr\+Σg−2\(ΣrΣg\)1/2\),\\mathrm\{FID\}=\\\|\\mu\_\{r\}\-\\mu\_\{g\}\\\|\_\{2\}^\{2\}\+\\mathrm\{Tr\}\\bigl\(\\Sigma\_\{r\}\+\\Sigma\_\{g\}\-2\(\\Sigma\_\{r\}\\Sigma\_\{g\}\)^\{1/2\}\\bigr\),with\(μr,Σr\)\(\\mu\_\{r\},\\Sigma\_\{r\}\)and\(μg,Σg\)\(\\mu\_\{g\},\\Sigma\_\{g\}\)the empirical mean and covariance of the real and generated feature sets respectively\.
### K\.5 Hungarian matching for clustering ACC
Clustering ACC is computed by first runningKK\-means \(K=K\{=\}number of classes,kk\-means\+\+ init, 10 restarts, keeping the lowest\-inertia solution\) on the semantic posterior meansμc\\mu\_\{\\mathrm\{c\}\}of then=2048n\{=\}2048evaluation subset, which yields a cluster assignmentc\(i\)∈\{1,…,K\}c\(i\)\\in\\\{1,\\dots,K\\\}for every sample that carries no inherent correspondence to the true labely\(i\)y\(i\)\. We form theK×KK\\times Kcontingency matrixCjk=\|\{i:c\(i\)=j,y\(i\)=k\}\|C\_\{jk\}=\|\\\{i:c\(i\)\{=\}j,\\,y\(i\)\{=\}k\\\}\|and solve
π⋆=argmaxπ∈SK∑j=1KCj,π\(j\)\\pi^\{\\star\}=\\arg\\max\_\{\\pi\\in S\_\{K\}\}\\sum\_\{j=1\}^\{K\}C\_\{j,\\pi\(j\)\}by the Hungarian algorithm \(scipy\.optimize\.linear\_sum\_assignmenton the cost matrix−C\-C\), giving the one\-to\-one cluster\-to\-label permutationπ⋆\\pi^\{\\star\}that maximizes agreement\. Reported ACC is1n∑i𝟙\[π⋆\(c\(i\)\)=y\(i\)\]\\frac\{1\}\{n\}\\sum\_\{i\}\\mathbb\{1\}\[\\pi^\{\\star\}\(c\(i\)\)\{=\}y\(i\)\]\. NMI and ARI \(also reported in Section E\.2\) do not require this matching step, since both are permutation\-invariant by construction; we include them alongside ACC precisely because they provide a matching\-free cross\-check on the same clustering\.
### K\.6 External classifier: architecture and training protocol
Primary Gen\-ACC and latent\-swap content\-/style\-following rates in the main paper use an external classifier trained once per dataset on real images and never updated during F\-CS\-WAE training\. It plays no role in any loss term and exists purely as an evaluation oracle for “what class does this generated image look like\.” Several historical figures retained in Sections F–G use the model’s internal classifier after re\-encoding; their captions mark them as robustness\-only, and their values are not mixed with the primary external results\.
*MNIST and Fashion\-MNIST\.*A 4\-layer CNN \(conv3232–conv6464–maxpool– conv128128–conv128128–maxpool–FC256256–FC1010, ReLU, batch norm after each conv\), trained for 20 epochs with Adam \(lr=10−3\\mathrm\{lr\}=10^\{\-3\}, batch 128, no weight decay\), no augmentation \(matching the un\-augmented training of F\-CS\-WAE itself, so the oracle’s decision boundary is not shaped by invariances the generative model was never asked to respect\)\. Test accuracy:99\.4%99\.4\\%\(MNIST\),91\.2%91\.2\\%\(Fashion\-MNIST\)\.
*CIFAR\-10\.*WideResNet\-28\-10\(Zagoruyko and Komodakis[2016](https://arxiv.org/html/2608.05243#bib.bib49)\), trained for 200 epochs with SGD \(momentum0\.90\.9, weight decay5×10−45\\times 10^\{\-4\}\), initiallr=0\.1\\mathrm\{lr\}=0\.1with cosine annealing, batch 128, standard augmentation \(random crop with 4px padding \+ reflection, horizontal flip\) applied only to this classifier’s own training data, never to F\-CS\-WAE’s training data\. Test accuracy:94\.1%94\.1\\%\. Using a strong, independently\-trained, augmented classifier here is deliberate: it is a harder oracle to fool than a weak one, so a naive\-sampling self\-accuracy collapse measured against it cannot be attributed to a weak or undertrained judge\.
### K\.7 MMD estimator: kernel, biased/unbiased form, batch sampling
Two distinct uses of MMD appear in the paper and use different estimators\.
*Training loss*\(ℒclass\\mathcal\{L\}\_\{\\mathrm\{class\}\},ℒagg\\mathcal\{L\}\_\{\\mathrm\{agg\}\},ℒstyle\-cls\\mathcal\{L\}\_\{\\mathrm\{style\\text\{\-\}cls\}\}\): the biasedVV\-statistic estimator,
MMD^V2\(X,Y\)=\\displaystyle\\widehat\{\\mathrm\{MMD\}\}^\{2\}\_\{V\}\(X,Y\)=\{\}1m2∑i,jk\(xi,xj\)\\displaystyle\\tfrac\{1\}\{m^\{2\}\}\\\!\\sum\_\{i,j\}k\(x\_\{i\},x\_\{j\}\)\+1n2∑i,jk\(yi,yj\)−2mn∑i,jk\(xi,yj\),\\displaystyle\+\\tfrac\{1\}\{n^\{2\}\}\\\!\\sum\_\{i,j\}k\(y\_\{i\},y\_\{j\}\)\-\\tfrac\{2\}\{mn\}\\\!\\sum\_\{i,j\}k\(x\_\{i\},y\_\{j\}\),computed within each training minibatch \(batch size 128, Table[7](https://arxiv.org/html/2608.05243#Sx10.T7)\)\. This estimator is biased upward at finite sample size but is non\-negative and typically has lower gradient variance at smallm,nm,n, which makes it a stable training objective\. The unbiased estimator below is also differentiable with respect to sample values; its distinction is the removal of within\-sample diagonal terms\. Forℒclass=1K∑kMMD2\(Qk,Pk\)\\mathcal\{L\}\_\{\\mathrm\{class\}\}=\\frac\{1\}\{K\}\\sum\_\{k\}\\mathrm\{MMD\}^\{2\}\(Q\_\{k\},P\_\{k\}\),QkQ\_\{k\}is the subset of the current minibatch with labelkk\(\|Qk\|≈128/10≈12\.8\|Q\_\{k\}\|\\approx 128/10\\approx 12\.8in expectation under random shuffling\) andPkP\_\{k\}is an equal\-size sample freshly drawn fromp\(zc∣y=k\)p\(z\_\{\\mathrm\{c\}\}\\mid y\{=\}k\)each step; classes with fewer than 2 members in a given minibatch contribute zero to that step’s loss \(a rare event at\|Qk\|≈12\.8\|Q\_\{k\}\|\\approx 12\.8, and one that self\-corrects over subsequent minibatches rather than requiring special handling\)\.ℒagg\\mathcal\{L\}\_\{\\mathrm\{agg\}\}uses the full batch ofzcz\_\{\\mathrm\{c\}\}against an equal\-size sample from the random\-class prior mixture1K∑kp\(zc∣y=k\)\\frac\{1\}\{K\}\\sum\_\{k\}p\(z\_\{\\mathrm\{c\}\}\\mid y\{=\}k\)\.
Δinter\\Delta\_\{\\mathrm\{inter\}\}is not an MMD\. It is computed directly from the class\-wise style means as
Δinter\\displaystyle\\Delta\_\{\\mathrm\{inter\}\}=\(K2\)−1∑j<k‖μ¯s\(j\)−μ¯s\(k\)‖2,\\displaystyle=\\binom\{K\}\{2\}^\{\-1\}\\sum\_\{j<k\}\\\|\\bar\{\\mu\}\_\{\\mathrm\{s\}\}^\{\(j\)\}\-\\bar\{\\mu\}\_\{\\mathrm\{s\}\}^\{\(k\)\}\\\|\_\{2\},μ¯s\(k\)\\displaystyle\\bar\{\\mu\}\_\{\\mathrm\{s\}\}^\{\(k\)\}=1nk∑i:yi=kμsi\.\\displaystyle=\\frac\{1\}\{n\_\{k\}\}\\sum\_\{i:y\_\{i\}=k\}\\mu\_\{\\mathrm\{s\}\}^\{i\}\.The MMD\-based evaluation diagnostics \(global MMD, JointMMD, and the per\-class/pairwise conditional MMD of Section E\.3\) instead use the unbiasedUU\-statistic estimator,
MMD^U2\(X,Y\)=\\displaystyle\\widehat\{\\mathrm\{MMD\}\}^\{2\}\_\{U\}\(X,Y\)=\{\}1m\(m−1\)∑i≠jk\(xi,xj\)\\displaystyle\\tfrac\{1\}\{m\(m\-1\)\}\\\!\\sum\_\{i\\neq j\}k\(x\_\{i\},x\_\{j\}\)\+1n\(n−1\)∑i≠jk\(yi,yj\)\\displaystyle\+\\tfrac\{1\}\{n\(n\-1\)\}\\\!\\sum\_\{i\\neq j\}k\(y\_\{i\},y\_\{j\}\)−2mn∑i,jk\(xi,yj\),\\displaystyle\-\\tfrac\{2\}\{mn\}\\\!\\sum\_\{i,j\}k\(x\_\{i\},y\_\{j\}\),computed once over the fulln=2048n\{=\}2048evaluation subset \(not minibatched\), so that reported diagnostic values are not subject to the same finite\-batch bias as the training loss\. Both estimators use the averaged multi\-scale RBF kernelk\(u,v\)=17∑σ∈\{0\.5,1,2,5,10,20,50\}exp\(−‖u−v‖22/2σ2\)k\(u,v\)=\\frac\{1\}\{7\}\\sum\_\{\\sigma\\in\\\{0\.5,1,2,5,10,20,50\\\}\}\\exp\(\-\\\|u\-v\\\|\_\{2\}^\{2\}/2\\sigma^\{2\}\)\(Section D\) for Euclidean latents; JointMMD’s spherical component instead usesk\(u,v\)=17∑σexp\(−2\(1−u⊤v\)/σ2\)k\(u,v\)=\\frac\{1\}\{7\}\\sum\_\{\\sigma\}\\exp\(\-2\(1\-u^\{\\top\}v\)/\\sigma^\{2\}\), the squared\-chordal\-distance analogue, at the same bandwidth ladder\. Per\-class conditional MMD \(Section E\.3\) applies the unbiased estimator withX=\{zsi:yi=k\}X=\\\{z\_\{\\mathrm\{s\}\}^\{i\}:y\_\{i\}\{=\}k\\\}againstY∼𝒩\(0,I\)⊗\|X\|Y\\sim\\mathcal\{N\}\(0,I\)^\{\\otimes\|X\|\}, i\.e\. a freshly sampled reference set of matching size for each class, rather than a single shared reference set reused across classes, so that sampling noise in the reference does not correlate across the per\-class comparisons\.
### K\.8 Pseudocode and code availability
Algorithm 1Leakage diagnostic pipeline \(evaluation\-only, from a saved checkpoint\)1:Input:checkpoint
θ\\theta, test split
𝒟test\\mathcal\{D\}\_\{\\mathrm\{test\}\}, eval seed
sevals\_\{\\mathrm\{eval\}\}
2:
𝒟eval←\\mathcal\{D\}\_\{\\mathrm\{eval\}\}\\leftarrowfixed
n=2048n\{=\}2048subsample of
𝒟test\\mathcal\{D\}\_\{\\mathrm\{test\}\}using
sevals\_\{\\mathrm\{eval\}\}
3:
\{\(μci,ρci,μsi,σsi,yi\)\}←Encoderθ\(𝒟eval\)\\\{\(\\mu\_\{\\mathrm\{c\}\}^\{i\},\\rho\_\{c\}^\{i\},\\mu\_\{\\mathrm\{s\}\}^\{i\},\\sigma\_\{s\}^\{i\},y\_\{i\}\)\\\}\\leftarrow\\mathrm\{Encoder\}\_\{\\theta\}\(\\mathcal\{D\}\_\{\\mathrm\{eval\}\}\)
4:for
k=1,…,Kk=1,\\ldots,Kdo
5:
μ¯s\(k\)←\|Ik\|−1∑i∈Ikμsi\\bar\{\\mu\}\_\{\\mathrm\{s\}\}^\{\(k\)\}\\leftarrow\|I\_\{k\}\|^\{\-1\}\\sum\_\{i\\in I\_\{k\}\}\\mu\_\{\\mathrm\{s\}\}^\{i\}, where
Ik=\{i:yi=k\}I\_\{k\}=\\\{i:y\_\{i\}=k\\\}
6:endfor
7:
Δinter←\(K2\)−1∑j<k‖μ¯s\(j\)−μ¯s\(k\)‖2\\Delta\_\{\\mathrm\{inter\}\}\\leftarrow\\binom\{K\}\{2\}^\{\-1\}\\sum\_\{j<k\}\\\|\\bar\{\\mu\}\_\{\\mathrm\{s\}\}^\{\(j\)\}\-\\bar\{\\mu\}\_\{\\mathrm\{s\}\}^\{\(k\)\}\\\|\_\{2\}⊳\\trianglerightSecs\. D and K\.7
8:
GlobalMMD←MMD^U2\(\{μsi\},𝒩\(0,I\)\)\\mathrm\{GlobalMMD\}\\leftarrow\\widehat\{\\mathrm\{MMD\}\}^\{2\}\_\{U\}\(\\\{\\mu\_\{\\mathrm\{s\}\}^\{i\}\\\},\\mathcal\{N\}\(0,I\)\)
9:
c\(⋅\)←K\-means\(\{μci\}\)c\(\\cdot\)\\leftarrow K\\text\{\-means\}\(\\\{\\mu\_\{\\mathrm\{c\}\}^\{i\}\\\}\);
π⋆←Hungarian\(c,y\)\\pi^\{\\star\}\\leftarrow\\mathrm\{Hungarian\}\(c,y\)⊳\\trianglerightSec\. K\.5
10:
ACC←1n∑i𝟙\[π⋆\(c\(i\)\)=yi\]\\mathrm\{ACC\}\\leftarrow\\frac\{1\}\{n\}\\sum\_\{i\}\\mathbb\{1\}\[\\pi^\{\\star\}\(c\(i\)\)\{=\}y\_\{i\}\]
11:
LP←\\mathrm\{LP\}\\leftarrowtrain linear head on
80%80\\%of
\{\(μsi,yi\)\}\\\{\(\\mu\_\{\\mathrm\{s\}\}^\{i\},y\_\{i\}\)\\\}, eval on remaining
20%20\\%
12:forordered pairs
\(a,b\)\(a,b\),
a≠ba\\neq bdo
13:for
r=1,…,100r=1,\\dots,100do
14:draw donor
i∼\{j:yj=a\}i\\sim\\\{j:y\_\{j\}\{=\}a\\\}, donor
j∼\{k:yk=b\}j\\sim\\\{k:y\_\{k\}\{=\}b\\\}
15:
x^←Decθ\(μci,μsj\)\\hat\{x\}\\leftarrow\\mathrm\{Dec\}\_\{\\theta\}\(\\mu\_\{\\mathrm\{c\}\}^\{i\},\\mu\_\{\\mathrm\{s\}\}^\{j\}\);
y^←ExternalClf\(x^\)\\hat\{y\}\\leftarrow\\mathrm\{ExternalClf\}\(\\hat\{x\}\)⊳\\trianglerightSec\. K\.6
16:endfor
17:endfor
18:report mean
P\(y^=a\)P\(\\hat\{y\}\{=\}a\)\(content\-following\),
P\(y^=b\)P\(\\hat\{y\}\{=\}b\)\(style\-following\)
We will release this pipeline with the camera\-ready version, together with the training loop implementing the phased schedule of Table[8](https://arxiv.org/html/2608.05243#Sx10.T8), exact configuration files, external\-classifier training scripts, evaluation scripts, and checkpoints\. The present anonymous submission does not claim an available code URL\.
Every diagnostic is computed from a saved checkpoint with a fixed evaluation seed, so the numbers can be regenerated without retraining\. The sampling controls, JointMMD permutation null, support\-graded swap analysis, intra\-class diversity measurement, and external\-classifier evaluations reported here were run under the protocols above\. The paper reports multi\-seed statistics only for entries explicitly marked as multi\-seed in Table[14](https://arxiv.org/html/2608.05243#Sx18.T14); single\-seed baseline rows are retained as diagnostic sanity checks rather than controlled superiority claims\.相似文章
贝叶斯因果发现如何失败?潜在混杂下线性高斯网络中结构后果的刻画
本文分析了在潜在混杂下线性高斯网络中贝叶斯因果发现如何失败,推导出一个导致评分函数偏好虚假边的相关性阈值,并刻画了两种不同的后验失败模式。
NumLeak: 公开数值基准作为基础模型中的潜在标签
本文介绍了NumLeak框架,用于检测基础模型在预训练中记忆公开数值基准而非展示样本外技能的情况,并表明顶级LLM能高保真地回忆起如Fama-French回报率等值,提出了一种简单的系统提示防御方法。
分数匹配学习的有限样本界
本文首次为使用分数匹配学习多项式指数族提供了非渐近样本复杂度界,显示出对模型维度的多项式依赖。
用于条件生成压缩感知的主动学习
本文提出了一个条件生成压缩感知框架,证明了基于提示词条件化模型在稳定恢复方面的界限,并通过在 Stable Diffusion 上的实验展示了提示词匹配如何影响采样分布。
FID 彩票:量化生成模型评估中的隐藏随机性
本文分析了不同训练种子和采样种子下FID分数的方差,揭示了图像生成评估中显著的可重复性问题。它提出了一种新的评估协议,包括误差带和每单元最优引导调整。