harness-sensitivity

Tag

Cards List
#harness-sensitivity

It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers

arXiv cs.AI · 2026-05-27 Cached

This paper empirically tests the common assumption that more structured harnesses universally improve LLM agent reliability, finding a non-monotone relationship across model tiers. It introduces the HEAT-24 benchmark and reveals that strict harnesses can harm frontier chat models while benefiting reasoning models.

0 favorites 0 likes
← Back to home

Submit Feedback