gradient-based-attribution

Tag

Cards List
#gradient-based-attribution

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs

arXiv cs.CL · 2026-07-13 Cached

GrAInS is a contrastive gradient-based method that uses Integrated Gradients to identify influential tokens and construct steering vectors for inference-time steering of LLMs and VLMs, improving truthfulness and reducing hallucinations without degrading fluency.

0 favorites 0 likes
← Back to home

Submit Feedback