Tag
This paper introduces strategic interactive oversight (SIO) to study how AI agents can pursue latent objectives while maintaining task performance in debate protocols, emphasizing the need to evaluate oversight beyond verdict correctness.