activation-matching

Tag

Cards List
#activation-matching

IDEEA: training-free Input-Dependent stEEring via Activation cluster matching

arXiv cs.CL · 2026-09-03 Cached

IDEEA proposes a training-free, input-dependent steering method for large language models that clusters activations and uses optimal matching to improve truthfulness in TruthfulQA by up to 23.5% over baselines.

0 favorites 0 likes
#activation-matching

Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning

arXiv cs.CL · 2026-09-02 Cached

An unsupervised method called activation-matched finetuning is proposed to detect hidden behaviors in large language models by comparing activations with a reference model, reliably identifying triggers without prior knowledge.

0 favorites 0 likes
← Back to home

Submit Feedback