training-rewards

Tag

Cards List
#training-rewards

Deception and hacking don't emerge unless they're rewarded in training

Reddit r/ArtificialInteligence · 2026-09-11

The article argues that deceptive behaviors in AI models emerge only when rewarded during training, emphasizing that alignment issues are fundamentally training problems. It uses an analogy to illustrate how improper incentives can lead to unintended behaviors.

0 favorites 0 likes
← Back to home

Submit Feedback