Tag
Proposes a unified post-hoc detection framework for copyright infringement in AI models, using conditional sensitivity and differential privacy to measure memorization across modalities.
This paper introduces TAKE (Trajectory-Aware Knowledge Estimation), a text dataset distillation framework that uses influence functions and optimal transport to reduce datasets to as little as 0.1% of their original size while preserving downstream task fidelity.
DRIFT proposes a method that uses on-policy influence functions to refine training data distribution for supervised fine-tuning of large language models, consistently improving performance ceilings over existing baselines.
DeMix is a novel framework that detects erroneous training samples and identifies their specific error types (label errors, feature errors, spurious correlations) by analyzing influence vectors, achieving a 22.61% improvement in debugging F1-score and 9.32% gain in task performance after data repair.
This paper proposes CLIF, a method using influence functions to interpret NLP models at both sample and concept levels within Concept Bottleneck Models, enabling transparent debugging and concept-level analysis.
This paper introduces a framework for token-level influence attribution in large language models by learning orthogonal latent spaces with sparse autoencoders, enabling precise identification of training data tokens that jointly influence predictions, with applications in high-stakes domains like healthcare.