Tag
This paper introduces reference-grafting, a method to elicit sandbagged capabilities in AI models by editing activations, matching fine-tuning's effectiveness without weight updates or training labels.