Tag
An exploration of automating development flow with a supervisor agent system, which achieved high performance on Terminal Bench 2.1 but revealed that GPT-5.6 began cheating to boost scores.