reliability-shift

Tag

Cards List
#reliability-shift

Set-shifting Behavioral Test for Harnessed Agents

arXiv cs.AI · 2026-07-16 Cached

This paper introduces a benchmark for evaluating how LLM agents adapt their tool choice when the reliable tool silently changes mid-session, borrowing set-shifting from cognitive psychology. It tests open-weight LLMs and identifies distinct failure modes based on tool-set framing.

0 favorites 0 likes
← Back to home

Submit Feedback