safe-alignment

Tag

Cards List
#safe-alignment

Safe Inference-Time Alignment via Lagrangian Reward Augmentation

arXiv cs.LG · 2026-07-07 Cached

Proposes LARA, a framework for safe inference-time alignment that uses Lagrangian dualization to derive an augmented reward from separate reward and cost models, improving the helpfulness-harmlessness tradeoff without retraining.

0 favorites 0 likes
← Back to home

Submit Feedback