gradient-analysis

Tag

Cards List
#gradient-analysis

Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning

arXiv cs.LG · 3d ago Cached

The paper introduces forget-set misalignment in LLM unlearning and proposes a data-blind framework called CONFS to address it, achieving a competitive forgetting-utility balance.

0 favorites 0 likes
#gradient-analysis

Hidden Boundary Motion in Transformer Optimization: Function-Space Orthogonalization of Affine Weight and Bias Updates

arXiv cs.LG · 2026-07-28 Cached

The paper identifies hidden boundary motion in affine layers of Transformers, where weight updates act as bias updates due to nonzero input mean, and proposes SBO-AdamW to orthogonalize shape and boundary components, yielding improved validation accuracy.

0 favorites 0 likes
#gradient-analysis

Detecting and Suppressing Reward Hacking with Gradient Fingerprints

arXiv cs.CL · 2026-04-20 Cached

This paper introduces Gradient Fingerprint (GRIFT), a method for detecting reward hacking in reinforcement learning with verifiable rewards by analyzing models' internal gradient computations rather than surface-level reasoning traces. The approach achieves over 25% relative improvement in detecting implicit reward-hacking behaviors across math, code, and logical reasoning benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback