Tag
IDEEA proposes a training-free, input-dependent steering method for large language models that clusters activations and uses optimal matching to improve truthfulness in TruthfulQA by up to 23.5% over baselines.
An unsupervised method called activation-matched finetuning is proposed to detect hidden behaviors in large language models by comparing activations with a reference model, reliably identifying triggers without prior knowledge.