Tag
This paper demonstrates that a small coordinated Wikipedia editing campaign can measurably shape how language models handle topics, using animal welfare as a case study.
DRIFT proposes a method that uses on-policy influence functions to refine training data distribution for supervised fine-tuning of large language models, consistently improving performance ceilings over existing baselines.
The paper proposes NULLs (Natively Unlearnable LLMs), a model class that isolates source-specific contributions in sparsely activated sinks while sharing backbone neurons, enabling clean unlearning of individual data sources without retraining and preserving general language capabilities.
GRASP introduces a geometry-aware, interaction-based method for scalable pretraining data attribution that models subset dynamics, outperforming existing additive approaches by over double the task-level rank correlation while reducing computation costs.
This paper provides the first systematic analysis of error sources in trajectory-based data attribution methods, identifies optimizer mismatch as the dominant error, proposes AdamW-influence to address it, and offers practical guidelines for data selection via a K-step look-ahead framework.