data-attribution

Tag

Cards List
#data-attribution

Form Over Content In Gradient-Based Data Attribution Methods

arXiv cs.CL · 16h ago Cached

This paper resolves a debate on gradient-based data attribution methods for large language models by demonstrating that they primarily track answer format rather than task semantics, challenging their reliability in targeted instruction tuning.

0 favorites 0 likes
#data-attribution

@RampLabs: AI spend is measured in tokens, model calls, and dollars. However, none of these fields describe the actual work being …

X AI KOLs Following · 2026-09-01 Cached

Ramp developed a semantic layer to attribute AI agent spend to objectives and outcomes, moving from monitoring spend to understanding AI ROI.

0 favorites 0 likes
#data-attribution

Small edits, large models: How Wikipedia advocacy shapes LLM values

arXiv cs.CL · 2026-06-25 Cached

This paper demonstrates that a small coordinated Wikipedia editing campaign can measurably shape how language models handle topics, using animal welfare as a case study.

0 favorites 0 likes
#data-attribution

DRIFT: Refining Instruction Data via On-Policy Data Attribution

arXiv cs.LG · 2026-06-18 Cached

DRIFT proposes a method that uses on-policy influence functions to refine training data distribution for supervised fine-tuning of large language models, consistently improving performance ceilings over existing baselines.

0 favorites 0 likes
#data-attribution

Natively Unlearnable Large Language Models

arXiv cs.LG · 2026-06-15 Cached

The paper proposes NULLs (Natively Unlearnable LLMs), a model class that isolates source-specific contributions in sparsely activated sinks while sharing backbone neurons, enabling clean unlearning of individual data sources without retraining and preserving general language capabilities.

0 favorites 0 likes
#data-attribution

GRASP: Geometry-aware Residual Alignment for Scalable Pretraining Data Attribution

arXiv cs.LG · 2026-06-08 Cached

GRASP introduces a geometry-aware, interaction-based method for scalable pretraining data attribution that models subset dynamics, outperforming existing additive approaches by over double the task-level rank correlation while reducing computation costs.

0 favorites 0 likes
#data-attribution

How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines

arXiv cs.LG · 2026-05-20

This paper provides the first systematic analysis of error sources in trajectory-based data attribution methods, identifies optimizer mismatch as the dominant error, proposes AdamW-influence to address it, and offers practical guidelines for data selection via a K-step look-ahead framework.

0 favorites 0 likes
← Back to home

Submit Feedback