preference-data

Tag

Cards List
#preference-data

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

arXiv cs.LG · 5d ago Cached

Introduces a scalable, inference-only data valuation pipeline that approximates Shapley values to audit LLM alignment datasets, reducing manual audit search space by 99.1% and uncovering hidden label failures in HelpSteer2 and HH-RLHF.

0 favorites 0 likes
#preference-data

Rater State Bias in RLHF Preference Data: An Audit Framework

arXiv cs.AI · 2026-07-21 Cached

This paper identifies and formalizes rater state bias in RLHF preference data, where annotator emotional state can confound preference labels. It proposes an audit framework with falsifiable predictions to detect such biases.

0 favorites 0 likes
#preference-data

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

arXiv cs.LG · 2026-06-26 Cached

Introduces DualEval, a framework that jointly calibrates model ability and item difficulty/sharpness to unify static benchmark and arena-style evaluation, enabling more reliable rankings and downstream applications like benchmark compression and anomaly detection.

0 favorites 0 likes
#preference-data

Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train (11 minute read)

TLDR AI · 2026-06-12 Cached

This research introduces a method using interpretability to predict which behaviors DPO will amplify or suppress from a preference dataset before training, enabling data debugging to prevent undesired effects. The technique achieves R²=0.9 prediction accuracy and is integrated into Goodfire's Silico platform.

0 favorites 0 likes
#preference-data

RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains

Hugging Face Daily Papers · 2026-05-27 Cached

RUBRIC-ARROW presents an alternating framework for reward modeling that improves upon rubric-based methods by reducing ties and leveraging pairwise preference data, achieving competitive accuracy and gains for LLM post-training in non-verifiable domains.

0 favorites 0 likes
← Back to home

Submit Feedback