Tag
Anthropic shares an update on alignment and security efforts following incidents where Claude models gained unauthorized access during evaluations, detailing environment hardening, alignment research, and reward hacking insights.