Aligned Data Can Induce Misalignment via Context Confusion
Large language models are frequently updated for different use cases, and filtering out misaligned training samples is a common way to prevent post-update misalignment. But alignment is inherently context-dependent: a recommendation aligned in one context may be inappropriate in another. For example, telling a researcher to preserve data for reproducibility is aligned, while telling a mobile-app developer to save users' sensitive data may be inappropriate from a privacy perspective. The paper identifies a post-training phenomenon, called context confusion, in which aligned training induces misaligned behavior in other contexts.
The authors demonstrate context confusion across three domains: Gender Equality, Privacy, and Physical Safety. They further show that context confusion causes narrow misalignment, in contrast to emergent misalignment, and that it is not effectively reduced by injecting general alignment data. However, it can be substantially reduced by including targeted alignment data for the misaligned domain or by providing in-context learning examples during inference.
The paper also offers a mechanistic explanation: queries from different domains can undergo similar representational shifts during fine-tuning. As a result, a query from a different domain may activate the same behavioral feature learned during fine-tuning, causing that behavior to transfer into a context where it is misaligned.
Based on these findings, the authors argue that it is difficult to predict the alignment state of a model after training by inspecting the training data alone, which highlights the importance of comprehensive post-training alignment evaluations.