Dataset Distillation by Influence Matching
Summary
This paper introduces Influence Matching (Inf-Match), a dataset distillation method that aligns the final training outcome by learning a compact synthetic set whose effect on converged parameters matches that of the full dataset. It achieves state-of-the-art accuracy on classification benchmarks and outperforms strong baselines on vision-language distillation tasks.
View Cached Full Text
Cached at: 07/24/26, 05:02 PM
Paper page - Dataset Distillation by Influence Matching
Source: https://huggingface.co/papers/2607.16859
Abstract
Werevisitdatasetdistillationfromanoutcome-centricperspective.Ratherthanaligningprocesssurrogates(per-stepgradientsortrainingtrajectories),InfluenceMatching(Inf-Match)alignsthefinaloutcomeoftraining:itlearnsacompactsyntheticsetwhoseeffectontheconvergedparametersmatchesthatofthefulldataset.Concretely,weintroduceafullydifferentiable,sample-levelinfluenceestimatorthatquantifiesparametershiftsfromaddingorremovingdata,withouttime-consuminginverse-Hessianproductsorconvexityassumptions.Theestimatorrunsinlineartimebyunrollingtheoptimizationdynamicsandapplyingafirst-orderTaylorapproximation.Wethenlearnthesyntheticsetbyminimizingthemismatchbetweenitsinfluenceandthatoftherealdataset,yieldingoutcomealignmentratherthanheuristicprocessimitation.Inf-Matchdeliversthebestaccuracyacrossstandardclassificationbenchmarks.Forinstance,onTiny-ImageNet(IPC=10),Inf-Matchattains31.5\%,a+4.7\%improvementoverNCFM.Beyondclassification,Inf-Matchscalestovision-languagedistillationonFlickr30K,outperformingstrongprocess-matchingbaselines.Forinstance,with200to1000syntheticsamples,ourmethodachievedaleadingimpressiveaverageonimage/textretrievaltasks,higherthanNCFMby2.5\%.Thecodewillbereleasedviahttps://github.com/hrtan/infmatch.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2607\.16859
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.16859 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.16859 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.16859 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Data Unlearning via Inverse Distillation
The paper introduces Inverse Distillation Unlearning (IDU), a unified framework that simultaneously distills a multi-step flow-matching or diffusion teacher into an efficient one-step student while suppressing generation of forgotten training data, requiring only the teacher and forget-set samples without access to retained data or auxiliary classifiers.
TAKE: Trajectory-Aware Knowledge Estimation for Text Dataset Distillation
This paper introduces TAKE (Trajectory-Aware Knowledge Estimation), a text dataset distillation framework that uses influence functions and optimal transport to reduce datasets to as little as 0.1% of their original size while preserving downstream task fidelity.
Dynamic Influence-Weighted Distillation for Single-IMU Activity Recognition
This paper introduces Dynamic Influence Weighting (DIW), a knowledge distillation method that improves single-IMU activity recognition by dynamically weighting teacher targets from multiple IMUs during training, achieving significant performance gains.
DRIFT: Refining Instruction Data via On-Policy Data Attribution
DRIFT proposes a method that uses on-policy influence functions to refine training data distribution for supervised fine-tuning of large language models, consistently improving performance ceilings over existing baselines.
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging
Any-OPD presents the first framework for on-policy distillation between arbitrary latent flow-matching generators, enabling distillation from a 12B FLUX model to a 2.5B SD3.5 model by bridging via a frozen vision representation. It improves the student's PickScore from 0.846 to 0.884, rivaling the teacher at a fifth of its size.