Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
Summary
This paper identifies a failure mode in classifier-free guidance distillation called Negative Branch Asymmetry, where errors in the positive and negative CFG branches cancel out, and proposes Positive-Direction Matching to supervise branches separately for more robust distilled models.
View Cached Full Text
Cached at: 07/28/26, 06:33 AM
Paper page - Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
Source: https://huggingface.co/papers/2607.24731 What happens when on-policy diffusion distillation meets classifier-free guidance?
A natural approach is to match the teacher’s and student’s guided predictions. But this seemingly reasonable objective hides a subtle ambiguity: errors from the positive and negative CFG branches can cancel each other, making the guided output look correct even when the individual branches are not.
We find that this ambiguity is usually harmless when teacher and student share the same negative conditioning. However, when the teacher retains privileged information in its negative branch, the two branches can start moving in opposite directions during training. We call this failure modeNegative Branch Asymmetry (NBA).
Our solution,Positive–Direction Matching (PDM), supervises the positive prediction and the CFG conditional direction separately. This removes cross-branch error compensation and makes the distilled model substantially more robust to changes in inference guidance scale.
We study this behavior across image-domain diagnostics and dense-to-sparse video control, including pose, depth, and scribble conditions.
We hope this work offers a useful perspective: when distilling CFG-based models, matching the final guided prediction may not be enough—the branches themselves matter.
Similar Articles
Commitment Before Realization: When Classifier-Free Guidance Becomes Unnecessary in Masked Diffusion Language Models
This paper investigates when classifier-free guidance (CFG) is actually necessary in masked diffusion language models, showing that guidance dependence is prompt-specific and can often be removed without losing constraint satisfaction, leading to a defined 'commitment horizon'.
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
This paper introduces a training-free diagnostic framework to analyze per-token distillation signals for reasoning models, revealing that guidance is more beneficial on incorrect rollouts and depends on student capacity and task context.
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance
This paper introduces RSTG, a method that selectively applies on-policy distillation to recover learning signals from zero-variance GRPO groups during LLM post-training, achieving substantial gains on math and code reasoning benchmarks.
Persistent Negatives for Adversarial Black-Box On-Policy Distillation
This paper introduces persistent-negative adversarial distillation to address the moving-target problem in black-box on-policy distillation, improving discriminator stability and performance across benchmarks.
Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance
This paper identifies limitations in token-level supervision for on-policy distillation of LLMs and proposes TOPD, which uses near-future trajectory information to better identify divergent reasoning states and distribute guidance across multiple tokens, achieving gains on AIME benchmarks.