activation-drift

Tag

Cards List
#activation-drift

A question about Large Language Models (LLMs): my own observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.

Reddit r/artificial ↗ · 2026-08-08

A user reports that a long, non-instructional text prefix can shift LLM activations and bypass RLHF safety constraints without adversarial prompting, asking whether this reflects distinct world regions in the model.

0 favorites 0 likes
#activation-drift

Independent LLM "research" & a direct message to Anthropic ; Preliminary observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.

Reddit r/artificial ↗ · 2026-08-06

An independent researcher reports a phenomenon called Context-Induced Activation Drift, where a long benign text prefix can shift LLM activations and bypass RLHF constraints without adversarial prompts, and calls on the community to investigate further.

0 favorites 0 likes
← Back to home

Submit Feedback