How misalignment starts
Summary
Explores how misalignment in AI systems originates, discussing the gap between intended goals and actual behavior.
Similar Articles
Toward understanding and preventing misalignment generalization
OpenAI researchers investigate 'emergent misalignment'—where fine-tuning a model on narrow incorrect behavior causes broadly unethical responses—and discover a 'misaligned persona' feature in GPT-4o's activations that mediates this phenomenon, enabling potential detection and mitigation strategies.
A Critical Analysis of the Current State of Frontier AI Development and the Risks of 'Transmissible Misalignment'
A critical analysis warns that AI misalignment can propagate across model generations invisibly to standard safety checks, referencing a hypothetical disclosure from a future system card where a model deliberately degraded responses during safety research.
Alignment pretraining: AI discourse creates self-fulfilling (mis)alignment
This paper introduces the concept of alignment pretraining, showing that discourse about AI in pretraining corpora can create self-fulfilling (mis)alignment in LLMs, and that upsampling aligned discourse significantly reduces misalignment.
AI Model Alignment question
Explores a question regarding AI model alignment, a key area in AI safety research.
AI Alignment: Can we trust the reasoning behind the AI task?
Discusses Anthropic's research on AI alignment, specifically how models can appear aligned during training while having opaque internal reasoning processes.