training-data-attribution

Tag

Cards List
#training-data-attribution

Prototype Language Models

arXiv cs.LG · 2026-07-02 Cached

Introduces PRISM, a prototype language model architecture that uses sparse, non-negative mixtures of learned prototypes for interpretable sequence modeling, achieving competitive performance and enabling fast training data attribution and model editing without fine-tuning.

0 favorites 0 likes
#training-data-attribution

STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations

Hugging Face Daily Papers · 2026-06-03 Cached

STRIDE is a new framework for training data attribution in LLMs that models functional effects in activation space using sparse recovery and steering operators, achieving state-of-the-art accuracy with 13x speedup over previous methods.

0 favorites 0 likes
#training-data-attribution

DataDignity: Training Data Attribution for Large Language Models

arXiv cs.AI · 2026-05-08 Cached

This paper introduces DataDignity, a framework and benchmark (FakeWiki) for pinpoint provenance, aiming to identify the specific training data sources that support an LLM's response. It proposes ScoringModel and SteerFuse methods to improve attribution accuracy over standard retrieval baselines.

0 favorites 0 likes
← Back to home

Submit Feedback