Tag
UserToolBench is a new benchmark for evaluating personalized decision-making in tool-use LLMs, testing whether models can infer latent user preferences, decide when to clarify, and produce user-aligned tool-call trajectories under incomplete information. Experiments show current models struggle with multi-tool coordination and long-horizon consistency.