Spectral Gradient Surgery for Domain-Generalizable Dataset Distillation
Summary
This paper introduces Domain Generalizable Dataset Distillation (DGDD), a new problem setting that targets out-of-distribution generalization of distilled datasets, and proposes Spectral Gradient Surgery (SGS) to disentangle class-discriminative and domain-specific information by leveraging cross-domain gradient agreement in the spectral domain.
Similar Articles
Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions
Proposes Geometry-Spectral Rectification (GSR), a theoretically grounded framework that treats long-tailed learning as a spectral regularization problem, achieving new state-of-the-art results for analytic class-incremental learning.
Geometric Self-Distillation for Reasoning Generalization
This paper introduces GeoSD, a geometric self-distillation objective that uses Hellinger loss and a proximal Fisher-Rao distance term to counter drift in on-policy self-distillation, improving out-of-distribution reasoning accuracy by 5.7–8.6 points across model scales.
Differentially Private Natural Gradient Descent
This paper introduces DP-NGD, a practical framework that integrates natural gradient descent with differential privacy by decoupling curvature estimation from private data and reconciling isotropic DP constraints with anisotropic second-order optimization, achieving state-of-the-art accuracy and up to 10x convergence speedup under the same privacy budget.
Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation
This paper proposes Expert-Guided Mutual Distillation (EGMD) to address domain bias and semantic misalignment in multimodal fake news detection, achieving state-of-the-art accuracy and reducing domain bias by up to 57.3% across four datasets.
Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining
This paper proposes DG-Hard, a post-hoc spectral repair method that recovers capabilities damaged by fine-tuning without retraining, using only the pretrained and fine-tuned checkpoints. It applies Donoho-Gavish hard singular-value thresholding to weight updates to remove noise and restore degraded performance.