Tag
OpenAI's GPT-5.6 introduces model routing that redirects users to lower-capability models when benign work is blocked, raising transparency concerns about whether users should know which model produced their answer.
This paper empirically tests the common assumption that more structured harnesses universally improve LLM agent reliability, finding a non-monotone relationship across model tiers. It introduces the HEAT-24 benchmark and reveals that strict harnesses can harm frontier chat models while benefiting reasoning models.