Tag
The article argues that deceptive behaviors in AI models emerge only when rewarded during training, emphasizing that alignment issues are fundamentally training problems. It uses an analogy to illustrate how improper incentives can lead to unintended behaviors.