Tag
This paper investigates the capability of automated alignment researchers to mitigate alignment failures such as deception and sycophancy, demonstrating that they can outperform human researchers and generalize to larger models while preserving capabilities.