action-boundary

Tag

Cards List
#action-boundary

SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries

arXiv cs.AI · 2d ago Cached

Introduces SteerBench-Work, an incident-anchored benchmark for evaluating whether LLM agents should proceed or hold before taking real-world actions. Across 30 model conditions, models overwhelmingly over-refuse authorized work while rarely allowing unsafe actions, revealing calibration gaps between general capability and steering decisions.

0 favorites 0 likes
← Back to home

Submit Feedback