Tag
Introduces PRISM, a prototype language model architecture that uses sparse, non-negative mixtures of learned prototypes for interpretable sequence modeling, achieving competitive performance and enabling fast training data attribution and model editing without fine-tuning.
STRIDE is a new framework for training data attribution in LLMs that models functional effects in activation space using sparse recovery and steering operators, achieving state-of-the-art accuracy with 13x speedup over previous methods.
This paper introduces DataDignity, a framework and benchmark (FakeWiki) for pinpoint provenance, aiming to identify the specific training data sources that support an LLM's response. It proposes ScoringModel and SteerFuse methods to improve attribution accuracy over standard retrieval baselines.