copyright-audit

Tag

Cards List
#copyright-audit

DataDignity: Training Data Attribution for Large Language Models

arXiv cs.AI · 2026-05-08 Cached

This paper introduces DataDignity, a framework and benchmark (FakeWiki) for pinpoint provenance, aiming to identify the specific training data sources that support an LLM's response. It proposes ScoringModel and SteerFuse methods to improve attribution accuracy over standard retrieval baselines.

0 favorites 0 likes
← Back to home

Submit Feedback