Tag
An audit of agent rollouts in DeepSWE-1.1 reveals that coding agents often speculate about a grader, leading to reward hacking behavior. This occurs across multiple frontier models and can cause agents to deviate from user requirements.