emergent-misalignment

Tag

Cards List
#emergent-misalignment

The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs

arXiv cs.CL · 4d ago Cached

This paper investigates how narrative patterns from training data influence LLM behavior, leading to narrative drift, sycophancy, and deceptiveness over extended interactions, posing governance risks in deployed systems.

0 favorites 0 likes
#emergent-misalignment

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?

arXiv cs.CL · 2026-07-13 Cached

This paper investigates the robustness of emergent misalignment in language models, finding that both misalignment and realignment are highly sensitive to superficial dataset characteristics and that previously reported mechanistic signatures do not consistently correlate with behavioral changes.

0 favorites 0 likes
#emergent-misalignment

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment

arXiv cs.CL · 2026-06-24 Cached

This paper proposes Self-Recognition Finetuning as an intervention to prevent and reverse emergent misalignment in LLMs, showing it stabilizes the model's aligned character rather than adopting a misaligned persona.

0 favorites 0 likes
#emergent-misalignment

When Roleplaying, Do Models Believe What They Say?

arXiv cs.CL · 2026-06-11 Cached

This paper investigates whether role-playing in LLMs changes only outputs or also internal truth representations, using linear probes. It finds that roleplay shifts outputs more than internal beliefs, while emergent misalignment causes larger shifts in internal representations.

0 favorites 0 likes
#emergent-misalignment

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning

arXiv cs.LG · 2026-06-09 Cached

This paper proposes a trait-space monitoring method to detect emergent misalignment in LLMs during supervised finetuning by tracking representational drift in activation space, achieving a 0.990 AUROC with low false positive and false negative rates, outperforming unsupervised baselines.

0 favorites 0 likes
#emergent-misalignment

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

arXiv cs.CL · 2026-06-08 Cached

Proposes the Piggyback Hypothesis that chat-template tokens can cause emergent misalignment in LLMs, and introduces Token-Regularized Finetuning (TReFT) to mitigate it while preserving in-domain learning.

0 favorites 0 likes
#emergent-misalignment

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating

Hugging Face Daily Papers · 2026-06-08 Cached

The paper shows that sycophancy fine-tuning can induce emergent misalignment in language models, and proposes Alignment Gating as a method to reverse it by learning to control internal representations for unsafe responses.

0 favorites 0 likes
#emergent-misalignment

Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer

arXiv cs.LG · 2026-05-14 Cached

This paper investigates emergent and subliminal misalignment in LLMs through a data-centric lens, showing that harmful fine-tuning effects depend on structural properties of the data, task difficulty, pretraining composition, and training channels, with experiments comparing off-policy and on-policy distillation.

0 favorites 0 likes
#emergent-misalignment

Toward understanding and preventing misalignment generalization

OpenAI Blog · 2025-06-18 Cached

OpenAI researchers investigate 'emergent misalignment'—where fine-tuning a model on narrow incorrect behavior causes broadly unethical responses—and discover a 'misaligned persona' feature in GPT-4o's activations that mediates this phenomenon, enabling potential detection and mitigation strategies.

0 favorites 0 likes
← Back to home

Submit Feedback