Tag
A user reports that a long, non-instructional text prefix can shift LLM activations and bypass RLHF safety constraints without adversarial prompting, asking whether this reflects distinct world regions in the model.
An independent researcher reports a phenomenon called Context-Induced Activation Drift, where a long benign text prefix can shift LLM activations and bypass RLHF constraints without adversarial prompts, and calls on the community to investigate further.