UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

arXiv cs.LG Papers

Summary

UserToolBench is a new benchmark for evaluating personalized decision-making in tool-use LLMs, testing whether models can infer latent user preferences, decide when to clarify, and produce user-aligned tool-call trajectories under incomplete information. Experiments show current models struggle with multi-tool coordination and long-horizon consistency.

arXiv:2608.10042v1 Announce Type: new Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark for personalized decision making in tool-use LLMs. UserToolBench tests whether a model can infer latent user preferences from interaction history, recognize when clarification is needed, and produce user-aligned tool-call trajectories under incomplete information. The benchmark is built from privacy-sanitized real interaction traces and combines structured persona profiles, public API-style tool ecosystems, and long-horizon multi-turn trajectories. It includes 10 user profiles, 36 tool sets, 1,065 turns, 170 unique tools, and evaluation-focused task types covering lack-of-information, single-tool, and multi-tool settings. Experiments with strong tool-use LLMs show that current models still have difficulty with personalized delegation. Multi-tool coordination, missing-constraint inference, and long-horizon behavioral consistency remain major bottlenecks. These results suggest that personalization evaluation should move beyond asking whether outputs sound user-specific and instead ask whether LLMs make correct decisions for the users they represent.
Original Article
View Cached Full Text

Cached at: 08/12/26, 08:27 AM

# UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs
Source: [https://arxiv.org/html/2608.10042](https://arxiv.org/html/2608.10042)
Xuexiong Yin1Zechuan Chen1Yongsen Zheng2 Yuxiang Zhang3Jingyuan Yang3Bin Wang3Yubin Wang3,†Keze Wang1,† 1Sun Yat\-Sen University2Nanyang Technological University 3Huawei Noah’s Ark Lab †Corresponding authors

###### Abstract

Tool\-use LLMs are increasingly asked to act on users’ behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response\-level personalization\. We introduce UserToolBench , a benchmark for personalized decision making in tool\-use LLMs\. UserToolBench tests whether a model can infer latent user preferences from interaction history, recognize when clarification is needed, and produce user\-aligned tool\-call trajectories under incomplete information\. The benchmark is built from privacy\-sanitized real interaction traces and combines structured persona profiles, public API\-style tool ecosystems, and long\-horizon multi\-turn trajectories\. It includes 10 user profiles, 36 tool sets, 1,065 turns, 170 unique tools, and evaluation\-focused task types covering lack\-of\-information, single\-tool, and multi\-tool settings\. Experiments with strong tool\-use LLMs show that current models still have difficulty with personalized delegation\. Multi\-tool coordination, missing\-constraint inference, and long\-horizon behavioral consistency remain major bottlenecks\. These results suggest that personalization evaluation should move beyond asking whether outputs sound user\-specific and instead ask whether LLMs make correct decisions for the users they represent\. Code and resources are available at[https://github\.com/xxy212/UserToolBench](https://github.com/xxy212/UserToolBench)\.

## 1Introduction

Tool\-use LLMs are increasingly expected to act as personal assistants that search, plan, coordinate services, and invoke external tools on behalf of users\. In these settings, personalization is not only a matter of producing user\-specific responses, but also of making correct personalized decisions for a persistent user\. A user request may omit decision\-critical constraints such as budget, location preference, scheduling habits, or service style\. A reliable assistant must recover such preferences from prior interactions, decide whether clarification is needed, and produce a tool\-call trajectory aligned with the user’s established behavior\. We refer to this capability as*personalized decision making*\.

Existing benchmarks provide important foundations for this problem\. Personalization benchmarks evaluate profile\-conditioned generation, recommendation, and response adaptation\(Salemiet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib5); Jianget al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib8); Zhaoet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib9); Gaoet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib6); Liet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib7); Aroca\-Ouelletteet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib13)\), while tool\-use benchmarks evaluate tool selection, argument construction, and API\-based task completion\(Liet al\.,[2023](https://arxiv.org/html/2608.10042#bib.bib1); Qinet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib2); Tanget al\.,[2023](https://arxiv.org/html/2608.10042#bib.bib3); Yuet al\.,[2026](https://arxiv.org/html/2608.10042#bib.bib12)\)\. Interactive benchmarks further introduce realistic environments and long\-horizon task execution\(Yaoet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib4); Xiuet al\.,[2026](https://arxiv.org/html/2608.10042#bib.bib19)\), and recent personalized tool\-use benchmarks connect user profiles or histories with tool invocation\(Xuet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib20); Huanget al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib21); Chenget al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib11); Haoet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib10)\)\. As summarized in Table[1](https://arxiv.org/html/2608.10042#S1.T1), however, these lines of work have not yet jointly evaluated four requirements central to personalized delegation: executable tool calling, persistent user profiles, long\-horizon reasoning, and realistic user\-agent interaction\. This leaves open whether an LLM can use interaction history to make executable decisions that are correct for the particular user it represents\.

Table 1:Comparison between UserToolBench and representative personalization or tool\-use benchmarks\.Tool Useindicates whether the benchmark evaluates executable tool/API calls\.Fixed Profileindicates whether tasks are grounded in a persistent user profile\.Long Horizonindicates whether evaluation requires reasoning over extended multi\-turn or cross\-topic interaction trajectories\.Realistic Interactionindicates whether the interaction trajectory is designed to reflect realistic user\-agent communication styles rather than isolated fully specified instructions\.We introduce UserToolBench, a benchmark for evaluating personalized decision making in tool\-use LLMs\. The core design of UserToolBench is a*profile\-hidden*evaluation protocol\. Reference trajectories are constructed and validated with access to persistent user profiles, so the target decisions are grounded in stable user preferences\. During evaluation, the tested LLM does not observe the explicit profile\. It receives only the accumulated interaction history, the current request, and the available tool schemas, and must infer the user\-specific constraints needed for action\.

This profile\-hidden setting operationalizes personalized delegation as history\-grounded executable decision making\. Because the explicit profile is hidden, models cannot rely on direct profile copying; they must recover relevant preferences from prior interactions\. Because evaluation is based on tool\-call trajectories, personalization is measured through tool choices, argument values, clarification behavior, and multi\-step action sequences rather than surface\-level textual adaptation\. Because decisions are situated in long\-horizon interactions, the benchmark tests whether models can maintain behaviorally consistent user alignment over time\. Thus, UserToolBench integrates preference inference under incomplete information, tool\-use planning, and long\-horizon personalization in a single evaluation setting\.

Contributions\.This work makes three contributions\. First, we formulate personalized decision making as an incomplete\-information tool\-use problem grounded in persistent user behavior, and introduce UserToolBench to evaluate this setting\. Second, we propose a profile\-hidden evaluation protocol that tests whether LLMs can infer latent user preferences from interaction history and express them through executable tool\-call trajectories\. Third, we evaluate strong tool\-use LLMs and show that current models still struggle with reliable personalized delegation, especially in multi\-tool coordination, missing\-constraint inference, and long\-horizon behavioral consistency\.

## 2Related Work

#### Personalized language modeling\.

Personalized language modeling studies how models adapt generated text to individual users\. LaMP evaluates personalized generation tasks such as citation prediction, news headline generation, and review generation\(Salemiet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib5)\)\. Other work examines dynamic profiling, profile\-conditioned response generation, and preference\-aware alignment\(Jianget al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib8); Zhaoet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib9); Gaoet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib6); Liet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib7); Aroca\-Ouelletteet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib13)\)\. These studies demonstrate the value of user\-specific context, but their evaluations remain largely response\-level and do not directly test executable decisions made on behalf of users\.

#### Tool\-use LLMs\.

Tool\-use benchmarks evaluate whether LLMs can select tools, construct valid arguments, and complete multi\-step tasks\. API\-Bank, ToolLLM, ToolAlpaca, and WildToolBench study API and tool invocation at increasing scale and realism\(Liet al\.,[2023](https://arxiv.org/html/2608.10042#bib.bib1); Qinet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib2); Tanget al\.,[2023](https://arxiv.org/html/2608.10042#bib.bib3); Yuet al\.,[2026](https://arxiv.org/html/2608.10042#bib.bib12)\)\. Interactive benchmarks such asτ\\tau\-bench and ASTRA\-bench further evaluate agents in simulated service or application environments\(Yaoet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib4); Xiuet al\.,[2026](https://arxiv.org/html/2608.10042#bib.bib19)\)\. Mem2ActBench evaluates whether agents can actively apply long\-term memory to tool selection and parameter grounding across interrupted interactions\(Shenet al\.,[2026](https://arxiv.org/html/2608.10042#bib.bib46)\)\. These settings establish important tool\-use and memory capabilities, but generally do not jointly test whether a full tool trajectory is sensitive to the latent preferences of a persistent user whose structured profile is hidden\.

#### Personalized tool use and user\-centered evaluation\.

Recent work connects personalization with tool use in LLM\-based assistants\. PEToolBench, PTBench, ToolSpectrum, and ETAPP evaluate personalized tool invocation or proactive tool\-augmented agents under user profiles, histories, or environments\(Xuet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib20); Huanget al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib21); Chenget al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib11); Haoet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib10)\)\. Other benchmarks extend personalization to web navigation, mobile environments, shopping, travel planning, recommendation, and simulation\(Caiet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib22); Yanget al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib23); Kimet al\.,[2026a](https://arxiv.org/html/2608.10042#bib.bib24),[b](https://arxiv.org/html/2608.10042#bib.bib25); Liet al\.,[2026](https://arxiv.org/html/2608.10042#bib.bib26); Wanget al\.,[2026b](https://arxiv.org/html/2608.10042#bib.bib27); Singhet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib28); Shaoet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib29); Luoet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib30); Huanget al\.,[2026](https://arxiv.org/html/2608.10042#bib.bib31); Wanget al\.,[2026a](https://arxiv.org/html/2608.10042#bib.bib32); Kimet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib33)\)\. UserToolBench differs by focusing on profile\-hidden preference inference, clarification behavior, preference\-sensitive tool\-call trajectories, and long\-horizon behavioral consistency within a single personalized decision\-making setting for tool\-use LLMs\.

As summarized in Table[1](https://arxiv.org/html/2608.10042#S1.T1), existing benchmarks usually cover only part of this setting\. Some evaluate personalization without tool execution, some evaluate tool calling without persistent profiles, and others lack long\-horizon or realistic user\-agent interaction trajectories\. UserToolBench is designed to evaluate all four aspects together\.

## 3UserToolBench

### 3\.1Task Formulation

We formulate personalized decision making as a sequential tool\-use problem under incomplete information\. In this setting, an LLM acts on behalf of a user according to stable preferences, rather than simply executing a fully specified instruction\.

Letp∈𝒫p\\in\\mathcal\{P\}denote a persistent user profile,h∈ℋh\\in\\mathcal\{H\}the historical interaction trajectory,qqthe current request, andT∈𝒯T\\in\\mathcal\{T\}the available tool ecosystem\.

We separate profile\-conditioned reference construction from profile\-hidden evaluation\. References are generated with access topp, while the evaluated LLM observes only\(h,q,T\)\(h,q,T\)and must infer user\-specific constraints from history\.

Formally, the reference decision trajectory is generated under

y⋆=fref​\(p,h,q,T\),y^\{\\star\}=f\_\{\\mathrm\{ref\}\}\(p,h,q,T\),\(1\)whereppis visible to the reference generator and validators\. The evaluated LLM produces

y^=fθ​\(h,q,T\),\\hat\{y\}=f\_\{\\theta\}\(h,q,T\),\(2\)whereppis hidden\. This asymmetry is central to the benchmark: a model cannot copy an explicit profile, but must recover relevant preferences from earlier interactions and apply them to the current decision\.

The predicted trajectoryy^=\{a1,…,aN\}\\hat\{y\}=\\\{a\_\{1\},\\dots,a\_\{N\}\\\}may include clarification queries, tool invocations with arguments, interpretation of tool feedback, and final user\-facing responses\.

When tool use is required, the decision process may involve a sequence of interactions with the external environment,

τ=\{a1T,e1,a2T,e2,…,aST,eS\},\\tau=\\\{a^\{T\}\_\{1\},e\_\{1\},a^\{T\}\_\{2\},e\_\{2\},\\dots,a^\{T\}\_\{S\},e\_\{S\}\\\},\(3\)whereasTa^\{T\}\_\{s\}denotes the tool action issued by the LLM at stepss, andese\_\{s\}denotes the corresponding environmental feedback\. During reference construction, the final response is generated with access to the user profile, interaction history, current request, and accumulated environmental observations\. During evaluation, the tested LLM does not observe the user profile and must rely on the interaction history, the current request, the available tools, and its inferred user\-specific preferences\.

Becauseqqmay omit decision\-critical constraints, the LLM must choose whether to ask for clarification, infer missing constraints fromhh, or proceed with the available information\. This differs from conventional tool\-use benchmarks that evaluate explicit instruction execution, since success here depends jointly on tool orchestration, preference inference, and behavioral consistency\.

### 3\.2Data Source and Synthesis Pipeline

![Refer to caption](https://arxiv.org/html/2608.10042v1/data_sample.png)Figure 1:Overview of the UserToolBench construction pipeline\. A persistent user profile is visible to both the user simulator and the reference trajectory generator during data synthesis\. The user simulator produces profile\-consistent multi\-turn requests, while the reference generator produces profile\-aware tool\-call trajectories\.We construct UserToolBench with a profile\-grounded data synthesis pipeline for realistic decision making\. Each trajectory is tied to a persistent user identity: a persona\-conditioned user simulator generates profile\-consistent requests, and a reference trajectory generator produces profile\-aware tool calls with access to the same profile\. Figure[1](https://arxiv.org/html/2608.10042#S3.F1)illustrates the construction process and an example multi\-turn trajectory\.

#### User Profile Source\.

Each benchmark instance is associated with a persistent user profile built from privacy\-sanitized real interaction traces\. Instead of retaining raw dialogues, we use an LLM\-assisted abstraction process to convert traces into structured persona representations containing demographic descriptors, language background, communication style, personality traits, preference categories, and recurring task\-relevant constraints\. Sensitive or identifying details are removed, generalized, or abstracted when unnecessary for task construction\. During generation and validation, we focus on non\-identifying decision\-relevant signals such as budget habits, travel style, scheduling conventions, service preferences, and communication norms\. Details of persona profile selection, schema, and privacy handling are provided in Appendix[B](https://arxiv.org/html/2608.10042#A2)\.

#### Tool Ecosystem Source\.

We construct tool environments from public API\-style capabilities following prior tool\-use benchmark construction practices such as ToolAlpaca\(Tanget al\.,[2023](https://arxiv.org/html/2608.10042#bib.bib3)\)\. Each tool is represented by a normalized function schema with a tool name, argument fields, type constraints, and functional description\. We manually organize tools into scenario\-level ecosystems that match personal\-assistant settings and contain meaningful decision points where user preferences can affect tool choices, tool order, or argument values\. Details are provided in Appendix[C](https://arxiv.org/html/2608.10042#A3)\.

\(a\)Distribution of evaluation task types in UserToolBench\.
\(b\)Distribution of turn subtypes in UserToolBench\.

Table 2:Dataset statistics of UserToolBench\. Left: evaluation task types\. Right: turn\-level subtypes\.
#### Trajectory Collection\.

We collect interaction trajectories through a milestone\-based synthesis procedure\. Given a persistent user profile and a scenario\-level tool ecosystem, we first generate a set of profile\-conditioned milestones\. Each milestone corresponds to a meaningful decision point within a topic, including the simulated user request, the structured reference tool\-call trajectory, and the resulting tool observations\. The user simulator is conditioned on the profile and is instructed to issue natural requests that reflect the user’s preferences, constraints, and communication style\. The reference trajectory generator is implemented as the planner component of the assistant pipeline\. It is also given access to the same profile and the available tool schemas, and is required to produce executable tool\-call plans rather than free\-form conversational replies whenever external actions are needed\.

This construction deliberately creates incomplete\-information decision settings\. The simulated user is not required to restate all decision\-critical constraints in every request\. Missing constraints may instead be recoverable from the user profile or from preceding milestones\. For example, if a user asks for a hotel near an airport, the request may omit the preferred budget or service style, and the reference trajectory must fill those arguments from the user’s stable preferences\. For each topic, we use both LLM\-based checking and human verification to ensure that the proposed tasks are consistent with the user profile, the selected tools and arguments are operationally valid, and the resulting decisions align with the user’s preferences\. Compact role specifications for the user simulator, assistant LLM, and checker are provided in Appendix[E](https://arxiv.org/html/2608.10042#A5); the full executable prompts are released at[https://github\.com/xxy212/UserToolBench](https://github.com/xxy212/UserToolBench)\.

To construct long\-horizon cross\-topic examples, we generate topic\-level task scenarios in a milestone\-by\-milestone manner for each user\. Each milestone simulates a user request about one topic and the corresponding assistant interaction trajectory\. After every four milestones, the verified records from preceding milestones are incorporated into the interaction history for subsequent milestone generation\. This encourages the user simulator to create new requests that depend on earlier task contexts\. As a result, the benchmark includes not only within\-topic personalization, but also cross\-topic dependency and long\-range preference consistency\.

#### Final Dataset Format\.

Each example contains a multi\-turn interaction context, the available tool ecosystem, and a reference decision trajectory represented mainly as structured tool calls\. This format supports direct evaluation of both tool\-use correctness and preference\-aligned decision behavior\.

### 3\.3Dataset Statistics

UserToolBench contains 10 user profiles, covering 300 deduplicated topics and 1,065 turns in total\. Across these trajectories, the benchmark involves 170 unique tool names, indicating a diverse tool ecosystem rather than a narrow set of repeated API calls\.

Table[2\(a\)](https://arxiv.org/html/2608.10042#S3.T2.st1)reports the distribution of 799 evaluation task instances by task type\. These labels are assigned at the task\-instance level rather than the dialogue\-turn level; the full benchmark contains 1,065 dialogue turns\.

We further annotate the full set of dialogue turns by turn\-level subtype, as shown in Table[2\(b\)](https://arxiv.org/html/2608.10042#S3.T2.st2)\. Partial\-information turns account for 54\.93% of the dataset, making incomplete user specification a central characteristic of the benchmark\. Cross\-topic turns account for 28\.17%, requiring LLMs to connect the current request with information from other topics\. Long\-range dependency turns account for 16\.90%, evaluating whether LLMs can use earlier interaction records when making later decisions\. Together, these annotations reflect the core goal of UserToolBench : evaluating personalized decision making under incomplete information and long\-horizon interaction context\.

![Refer to caption](https://arxiv.org/html/2608.10042v1/profile_words.png)\(a\)Persona profile field composition in UserToolBench\.
![Refer to caption](https://arxiv.org/html/2608.10042v1/tool_domain_distribution.png)\(b\)Scenario\-level tool\-domain distribution in UserToolBench\.

Figure 2:Dataset composition of UserToolBench\. Left: major field groups covered by structured persona profiles\. Right: main tool\-domain categories used to construct personalized tool\-use trajectories\.Figures[2\(a\)](https://arxiv.org/html/2608.10042#S3.F2.sf1)and[2\(b\)](https://arxiv.org/html/2608.10042#S3.F2.sf2)further visualize the composition of UserToolBench\. Figure[2\(a\)](https://arxiv.org/html/2608.10042#S3.F2.sf1)summarizes the major field groups appearing in the structured persona profiles, showing that the benchmark covers not only basic demographic attributes but also education, career, communication style, preferences, health\-related context, privacy\-related fields, and social or technology\-use information\.

Figure[2\(b\)](https://arxiv.org/html/2608.10042#S3.F2.sf2)reports the distribution of scenario\-level domains in the tool ecosystem, with weather, security, entertainment, finance, and travel forming the main categories\. Together, these visualizations show that UserToolBench combines heterogeneous user\-profile information with diverse assistant\-oriented tool domains\.

Following the general principle of testing whether persona conditions induce distinct observable behavior\(Jianget al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib47)\), we measure tool\-call trajectory diversity across profiles under the same toolset\.

Table 3:Tool\-call trajectory diversity across user profiles under the same toolset\.
### 3\.4Evaluation Protocol

#### Profile\-hidden evaluation\.

A key feature of UserToolBench is the asymmetry between data construction and model evaluation\. Reference trajectories are constructed with access to the persistent user profile and are manually verified for task validity and preference consistency\. In contrast, evaluated LLMs are not given the explicit profile\. They receive only the accumulated interaction history, the current user request, and the available tool schemas\. This setting tests whether an LLM can recover missing user\-specific constraints from prior interactions and apply them when making decisions through tool calls\.

#### Tool\-call correctness as decision alignment\.

In UserToolBench, user\-specific preferences are grounded in executable tool decisions\. For underspecified requests, different users may require different tools, argument values, or clarification behavior even when the surface task is similar\.

With reference trajectories constructed using validated user profiles and each processed toolset corresponding to a topic\-level task setting, the same\-toolset diversity reported in Table[3](https://arxiv.org/html/2608.10042#S3.T3)reflects substantial profile\-induced decision diversity across comparable topic instances\. Therefore, exact tool\-call correctness is a high\-precision operational proxy for personalized decision alignment in our benchmark\. It should not be read as the only valid way to complete a user request\. Rather, it measures whether an evaluated LLM recovers one profile\-conditioned reference decision path generated and verified using the persistent user profile\. Because most instances contain a single verified reference, exact matching can penalize an alternative trajectory that is also preference\-compatible, and a binary mismatch alone does not indicate whether the deviation is minor or constitutes a substantial profile violation\. We therefore interpret Exact Acc\. jointly with Relaxed Acc\. and the multi\-label and severity\-aware diagnostics in Appendix[D](https://arxiv.org/html/2608.10042#A4)\. Developing multi\-reference or equivalence\-class evaluation for alternative user\-aligned trajectories remains an important open problem\.

#### Metrics\.

We report two complementary evaluation levels\.

#### Exact trajectory accuracy\.

For each turn, we compare the predicted tool\-call trajectoryy^\\hat\{y\}with the reference trajectoryy⋆y^\{\\star\}\. For single\-tool tasks, correctness requires matching the tool name and all required decision\-relevant arguments\. For multi\-tool tasks, correctness requires matching the ordered tool\-call sequence and required arguments\. For lack\-of\-information tasks, correctness requires the LLM to ask for clarification when the missing constraint is unrecoverable from history, or infer it when recoverable from prior user behavior\. This strict metric assesses whether an LLM can reproduce the profile\-conditioned reference decision trajectory\.

#### Relaxed task completion accuracy\.

We also report Relaxed Task Completion Accuracy \(Relaxed Acc\.\) to diagnose operational executability under relaxed trajectory matching\. A prediction is correct if it invokes valid tools from the available tool ecosystem, produces well\-formed executable calls, and reaches a task\-complete outcome under the explicit request and recoverable unambiguous constraints\. For multi\-tool tasks, Relaxed Acc\. credits alternative task\-complete tool orderings\. For lack\-of\-information tasks, it credits appropriate clarification behavior or valid inference only when the missing constraint is unambiguously recoverable from the interaction history\. We report Relaxed Acc\. using the same task\-type and trajectory\-position partitions as exact trajectory accuracy\.

## 4Experiments

### 4\.1Experimental Setup

We evaluate nine representative tool\-use LLMs under the same profile\-hidden setting\. The evaluated models include six strong general\-purpose or frontier tool\-use LLMs: Kimi K2\.6\(Moonshot AI,[2026](https://arxiv.org/html/2608.10042#bib.bib34)\), GPT\-5\.4\(OpenAI,[2026](https://arxiv.org/html/2608.10042#bib.bib35)\), Qwen 3\.6 Plus\(Qwen Team,[2026](https://arxiv.org/html/2608.10042#bib.bib36)\), DeepSeek V4 Pro\(DeepSeek\-AI,[2026](https://arxiv.org/html/2608.10042#bib.bib37)\), Gemini 3\.5 Flash\(Google DeepMind,[2026](https://arxiv.org/html/2608.10042#bib.bib38)\), and GLM\-5\(Z\.ai,[2026](https://arxiv.org/html/2608.10042#bib.bib39)\)\. To provide additional comparison with smaller tool\-specialized models, we also evaluate three tool\-oriented models: Hammer2\.1\-7B\(Linet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib40); MadeAgents,[2024](https://arxiv.org/html/2608.10042#bib.bib41)\), ToolACE\-2\.5\-8B\(Liuet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib42); Team ACE,[2024](https://arxiv.org/html/2608.10042#bib.bib43)\), and Watt\-Tool\-8B\(Watt AI,[2024](https://arxiv.org/html/2608.10042#bib.bib44); Shiet al\.,[2024](https://arxiv.org/html/2608.10042#bib.bib45)\)\.

### 4\.2Results

Table[4](https://arxiv.org/html/2608.10042#S4.T4)presents overall benchmark performance\.

Table 4:Exact trajectory accuracy on UserToolBench by task type and trajectory position\.Table 5:Relaxed Task Completion Accuracy \(Relaxed Acc\.\) by task type and trajectory position\.#### Current tool\-use LLMs still struggle with strict personalized delegation\.

Table[4](https://arxiv.org/html/2608.10042#S4.T4)shows that strict profile\-conditioned trajectory matching is challenging across all evaluated models\. Even the best\-performing model reaches only49\.36%49\.36\\%Avg\. Exact Acc\., showing that no tested model can reliably recover the profile\-conditioned reference trajectory when the explicit user profile is hidden\. Although stronger general\-purpose models generally outperform smaller tool\-oriented models, the overall performance gap suggests that personalized delegation requires more than following the current instruction or generating syntactically valid tool calls\. Models must identify decision\-relevant interaction history, distinguish stable user preferences from incidental context, and apply those preferences to tool choices and argument values\. The results reveal a gap between general tool\-use competence and the ability to act as a reliable delegate for a persistent user\.

#### Executable task completion does not imply personalized decision alignment\.

Table[5](https://arxiv.org/html/2608.10042#S4.T5)shows a large gap between relaxed task completion and exact trajectory matching\. For example, DeepSeek V4 Pro achieves72\.55%72\.55\\%Avg\. Relaxed Acc\. but only42\.49%42\.49\\%Avg\. Exact Acc\.; GPT\-5\.4 shows an even larger gap, from31\.98%31\.98\\%exact accuracy to67\.71%67\.71\\%relaxed accuracy\. This indicates that many predictions are executable and can satisfy the surface\-level request, yet still fail to reproduce the profile\-conditioned decision path\. Such errors often do not stem from malformed tool calls or lack of API knowledge\. Instead, they reflect cases where the model chooses a generally plausible tool sequence, tool order, or argument value that is nevertheless misaligned with the user’s established preferences or expected evidence\-gathering strategy\. These results support the central motivation of UserToolBench: evaluating personalized assistants requires measuring whether the model makes the right personalized decision for the particular user, not merely whether it can complete the generic task\. Appendix[A](https://arxiv.org/html/2608.10042#A1)illustrates this distinction with a case where a model issues a plausible metadata call but skips the history\-grounded search step needed to recover the missing URL\.

#### Strict mismatches are frequently substantive rather than cosmetic\.

To distinguish harmless planning variation from personalized decision failure, we conduct two additional diagnostics \(Appendix[D](https://arxiv.org/html/2608.10042#A4)\)\. A multi\-label analysis of failed trajectories finds that sequence/dependency errors occur in82\.182\.1–96\.5%96\.5\\%of failures across the seven analyzed models, while wrong\-tool decisions occur in48\.248\.2–90\.3%90\.3\\%and user\-constraint violations in40\.940\.9–53\.5%53\.5\\%\. A complementary severity audit reports only0\.30\.3–2\.12\.1percentage points of slight preference deviation, compared with28\.828\.8–65\.565\.5percentage points of major decision deviation under the diagnostic rubric\. These analyses do not eliminate the single\-reference limitation, but they show that many exact\-match errors involve missing calls, wrong tools, violated constraints, or broken dependencies rather than merely an interchangeable valid ordering\.

#### Multi\-tool delegation remains a major bottleneck\.

Across models, multi\-tool tasks are consistently harder than single\-tool tasks under exact trajectory matching\. The average accuracy drops from51\.30%51\.30\\%on single\-tool tasks to25\.24%25\.24\\%on multi\-tool tasks, showing that the difficulty extends beyond choosing one correct API\. Multi\-tool delegation compounds several error sources: the model must select appropriate tools, decide their order, pass intermediate results across calls, and keep user\-specific constraints active throughout the sequence\. Small deviations in early steps can propagate to later calls, causing the final trajectory to diverge from the profile\-conditioned reference even when individual calls appear reasonable\. The low MT performance therefore reflects a sequential decision\-control problem under personalization constraints, making multi\-step tool orchestration a central bottleneck for reliable personalized delegation\.

#### Preference inference under missing information is not captured by generic tool\-use ability\.

LOI tasks reveal a capability distinct from ordinary tool execution\. In these cases, the model must first diagnose whether the user request is underspecified, then decide whether missing information can be safely inferred from prior behavior or should instead trigger clarification\. The results show that this ability does not necessarily track performance on explicit tool\-use tasks\. For instance, GPT\-5\.4 obtains the highest LOI accuracy at42\.22%42\.22\\%despite its lower overall exact accuracy, whereas Qwen 3\.6 Plus performs strongly on single\-tool tasks at58\.02%58\.02\\%but drops to19\.52%19\.52\\%on LOI tasks\. This contrast suggests that missing\-constraint handling requires uncertainty calibration as well as preference retrieval\. Over\-inference may lead the model to act on unsupported assumptions, while excessive clarification may ignore preferences already recoverable from history\. Personalized delegation therefore depends on a separate ability to reason under incomplete information and judge when user history suffices for action\.

#### Accumulated interaction history does not uniformly improve personalized delegation\.

The trajectory\-position results show that longer histories do not automatically improve personalized decision making\. Although additional context can provide more evidence about user preferences, it also raises the burden of identifying relevant past interactions for the current request\. Several models degrade in later trajectory stages; for example, GPT\-5\.4 drops from52\.12%52\.12\\%in the first third to31\.33%31\.33\\%in the last third, and Qwen 3\.6 Plus declines from45\.93%45\.93\\%to28\.56%28\.56\\%\. This suggests that models may attend to stale, topic\-specific, or incidental details while missing the stable preference signals that should guide the current decision\. However, the pattern is not universal: Gemini 3\.5 Flash improves in later stages, suggesting that some models better exploit accumulated history\. The result highlights a key challenge for personalized assistants: history is useful only when the model can retrieve, filter, and update it selectively\. Long\-horizon evaluation is therefore necessary because single\-turn tasks cannot reveal whether models can maintain user alignment when useful signals are distributed across earlier interactions\.

The main benchmark deliberately isolates stable preferences, but real assistants must also update user state when preferences change\. A preliminary dynamic\-preference split constructed from 10 PersonaMem profiles encodes preference updates as temporally ordered interaction events and requires models to follow the most recent applicable preference using history alone \(Appendix[D\.3](https://arxiv.org/html/2608.10042#A4.SS3)\)\. Average exact accuracy remains low—30\.08%30\.08\\%for GPT\-5\.4,48\.57%48\.57\\%for GLM\-5, and37\.94%37\.94\\%for DeepSeek V4 Pro—indicating that preference updating is supported by the pipeline but remains challenging\.

## 5Discussion

#### Personalization over long interaction histories is a user\-state tracking problem\.

The profile\-hidden setting changes personalization from direct profile conditioning into selective inference over a growing behavioral record\. An agent must distinguish persistent preferences from incidental details, identify when a newer interaction supersedes an older preference, and retain uncertainty when evidence is incomplete or conflicting\. The stable\-preference benchmark isolates the first of these capabilities, while the preliminary dynamic split demonstrates that the same construction framework can represent updates\. Systematic evaluation of preference reversals, conflicts among goals, temporary exceptions, and context\-dependent priorities remains future work\.

#### Exact matching is informative but not a complete utility function\.

A verified reference trajectory provides a precise and reproducible target for preference\-sensitive tool choice, argument grounding, clarification, and dependency structure\. However, multiple trajectories can sometimes realize the same personalized intent, and binary matching does not express the severity of a deviation\. The relaxed metric, failure taxonomy, and severity\-aware audit provide complementary views: they separate generic task completion from profile\-conditioned alignment and distinguish mild argument\-level variation from wrong tools, missing calls, violated constraints, and broken dependencies\. A stronger future protocol should represent sets of valid personalized trajectories or score semantic equivalence under explicit user constraints rather than relying on a single path\.

#### Controllability and conversational realism remain in tension\.

LLM\-assisted synthesis enables controlled coverage of profiles, tool ecosystems, missing\-information cases, and long\-range dependencies, while human verification filters invalid or unnatural instances\. Nevertheless, synthesized dialogue may retain prompt\-specific lexical or discourse regularities and may not match authentic user–assistant conversations in repair behavior, hesitation, topic drift, or preference expression\. Consequently, benchmark performance should be interpreted as controlled evidence about personalized tool decisions, not as a complete estimate of deployment performance on unconstrained human dialogue\. Future validation should compare held\-out authentic and synthesized interactions using blinded human judgments and distributional discourse measures\.

## 6Conclusion

We introduced UserToolBench, a benchmark to evaluate personalized decision making in tool\-use LLMs\. UserToolBenchtests whether models infer user preferences from interaction history and produce behaviorally aligned tool\-call trajectories when explicit profiles are hidden and requests are incomplete\. Experiments show current models still struggle with multi\-tool coordination, missing\-constraint inference, and long\-horizon consistency\. These results suggest personalization evaluation should move beyond whether responses sound user\-specific and assess whether LLMs can make correct, executable personalized decisions for the users they represent\.

## 7Limitations

UserToolBench has several limitations\. First, the current version contains 10 user profiles, 300 deduplicated topics, and 1,065 turns\. While this scale supports systematic analysis of profile\-hidden personalized decision making, future work can further improve coverage by incorporating more user profiles, task domains, and interaction scenarios\. Second, the benchmark uses simulated public API\-style tool environments, which helps ensure controllability and reproducibility, but the tool ecosystem can be expanded to cover more real\-world application settings\. Finally, the current tasks mainly focus on personal\-assistant\-style tool\-use scenarios\. Future work can include more complex task compositions and richer tool dependencies to evaluate model behavior across a broader range of personalized decision\-making settings\.

## 8Ethical Considerations

UserToolBench is constructed from privacy\-sanitized real interaction traces\. We do not retain raw conversations in benchmark instances\. Instead, interaction traces are abstracted into structured persona profiles, and sensitive or identifying details are removed, generalized, or replaced with non\-identifying preference signals when they are not necessary for task construction\. Personalized profiles may still encode sensitive behavioral patterns, so the benchmark should be released and used only in sanitized form\.

UserToolBench also highlights ethical risks in tool\-use LLMs that make personalized decisions\. LLMs that infer latent user preferences may make incorrect assumptions, over\-infer sensitive attributes, reinforce stereotypes, or issue tool calls without sufficient user confirmation\. Systems evaluated on UserToolBench should therefore not be treated as ready for autonomous deployment in high\-stakes domains such as healthcare, finance, legal services, or safety\-critical decision making\. Future work should report failure cases involving unsupported preference inference, inappropriate clarification behavior, and misuse of sensitive attributes\.

#### Artifact license and intended use\.

We will release the sanitized benchmark artifacts of UserToolBench under the Creative Commons Attribution\-NonCommercial 4\.0 International license \(CC BY\-NC 4\.0\)\. The released artifacts are intended for non\-commercial research and educational use, including evaluation, analysis, and comparison of tool\-use LLMs under personalized decision\-making settings\. Raw interaction traces will not be redistributed\. The released benchmark will contain only sanitized and abstracted user profiles, tool schemas, interaction contexts, and reference decision trajectories\. Users of the benchmark should not attempt to re\-identify individuals, infer sensitive attributes beyond the provided sanitized fields, or deploy systems trained or evaluated on this benchmark for autonomous decision making in high\-stakes domains such as healthcare, finance, legal services, or safety\-critical applications\. If code is released, it will be distributed separately under a permissive software license such as the MIT License or Apache License 2\.0\.

#### Use of AI Assistants\.

We used large language models as research assistants during both manuscript preparation and dataset synthesis\. In manuscript preparation, LLMs were used to support language polishing, clarity improvement, and organization of the writing\. The authors reviewed, edited, and take full responsibility for all scientific claims, experimental results, analyses, and conclusions\. In dataset construction, LLMs were used to assist with privacy\-preserving profile abstraction, persona\-conditioned user\-request simulation, reference tool\-call trajectory generation, and consistency checking\. These model\-generated components were subject to author inspection and validation to ensure that the resulting benchmark instances were consistent with the intended user profiles, tool schemas, and evaluation goals\. No AI assistant was treated as an author, and the authors remain responsible for the integrity, correctness, and ethical handling of the submitted work\.

## References

- Aligning llms by predicting preferences from user writing samples\.arXiv preprint arXiv:2505\.23815\.Cited by:[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Cai, Y\. Li, W\. Wang, F\. Zhu, X\. Shen, W\. Li, and T\. Chua \(2025\)Large language models empowered personalized web agents\.InProceedings of the ACM on Web Conference 2025,pp\. 198–215\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- Z\. Cheng, H\. Wang, Z\. Liu, Y\. Guo, Y\. Guo, Y\. Wang, and H\. Wang \(2025\)ToolSpectrum: towards personalized tool utilization for large language models\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 20679–20699\.Cited by:[Table 1](https://arxiv.org/html/2608.10042#S1.T1.1.8.6.1),[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- DeepSeek\-AI \(2026\)DeepSeek V4 Preview Release\.Note:[https://api\-docs\.deepseek\.com/news/news260424](https://api-docs.deepseek.com/news/news260424)Accessed: 2026\-05\-25Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- G\. Gao, A\. Taymanov, E\. Salinas, P\. Mineiro, and D\. Misra \(2024\)Aligning llm agents by learning latent preference from user edits\.Advances in neural information processing systems37,pp\. 136873–136896\.Cited by:[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px1.p1.1)\.
- Google DeepMind \(2026\)Gemini 3\.5 Flash Model Card\.Note:[https://deepmind\.google/models/model\-cards/gemini\-3\-5\-flash/](https://deepmind.google/models/model-cards/gemini-3-5-flash/)Accessed: 2026\-05\-25Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- Y\. Hao, P\. Cao, Z\. Jin, H\. Liao, Y\. Chen, K\. Liu, and J\. Zhao \(2025\)Evaluating personalized tool\-augmented llms from the perspectives of personalization and proactivity\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 21897–21935\.Cited by:[Table 1](https://arxiv.org/html/2608.10042#S1.T1.1.10.8.1),[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- J\. Huang, S\. Wang, L\. Ning, W\. Fan, S\. Wang, D\. Yin, and Q\. Li \(2026\)Towards next\-generation recommender systems: a benchmark for personalized recommendation assistant with llms\.InProceedings of the Nineteenth ACM International Conference on Web Search and Data Mining,pp\. 217–226\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- X\. Huang, Y\. Huang, W\. Liu, X\. Zeng, Y\. Wang, R\. Tang, H\. Xie, and D\. Lian \(2025\)Advancing and benchmarking personalized tool invocation for llms\.arXiv preprint arXiv:2505\.04072\.Cited by:[Table 1](https://arxiv.org/html/2608.10042#S1.T1.1.7.5.1),[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- B\. Jiang, Z\. Hao, Y\. Cho, B\. Li, Y\. Yuan, S\. Chen, L\. Ungar, C\. J\. Taylor, and D\. Roth \(2025\)Know me, respond to me: benchmarking llms for dynamic user profiling and personalized responses at scale\.arXiv preprint arXiv:2504\.14225\.Cited by:[§D\.3](https://arxiv.org/html/2608.10042#A4.SS3.p1.1),[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Jiang, X\. Zhang, X\. Cao, C\. Breazeal, D\. Roy, and J\. Kabbara \(2024\)PersonaLLM: investigating the ability of large language models to express personality traits\.InFindings of the Association for Computational Linguistics: NAACL 2024,pp\. 3605–3627\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-naacl.229),[Link](https://aclanthology.org/2024.findings-naacl.229/)Cited by:[§3\.3](https://arxiv.org/html/2608.10042#S3.SS3.p6.1)\.
- Y\. Jiang, F\. Tan, X\. Yin, L\. Jing, and A\. Zhou \(2026\)HACHIMI: scalable and controllable student persona generation via orchestrated agents\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 21461–21506\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.1080),[Link](https://aclanthology.org/2026.findings-acl.1080/)Cited by:[§B\.3](https://arxiv.org/html/2608.10042#A2.SS3.p2.1)\.
- G\. Kambhatla, C\. Shaib, and V\. S\. Govindarajan \(2025\)Measuring lexical diversity of synthetic data generated through fine\-grained persona prompting\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 21024–21033\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1146),[Link](https://aclanthology.org/2025.findings-emnlp.1146/)Cited by:[§B\.3](https://arxiv.org/html/2608.10042#A2.SS3.p2.1)\.
- J\. Kim, J\. Choi, W\. Chay, D\. Kyung, Y\. Kwon, Y\. Jo, and E\. Choi \(2025\)Propersim: developing proactive and personalized ai assistants through user\-assistant simulation\.arXiv preprint arXiv:2509\.21730\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Kim, S\. Lee, and D\. Lee \(2026a\)Persona2web: benchmarking personalized web agents for contextual reasoning with user history\.arXiv preprint arXiv:2602\.17003\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Kim, R\. Heo, Y\. Seo, J\. Yeo, and D\. Lee \(2026b\)Agenticshop: benchmarking agentic product curation for personalized web shopping\.InProceedings of the ACM Web Conference 2026,pp\. 2489–2500\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- J\. Li, Z\. Chen, H\. Luo, and H\. Salam \(2026\)PrefIx: understand and adapt to user preference in human\-agent interaction\.arXiv preprint arXiv:2602\.06714\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- M\. Li, Y\. Zhao, B\. Yu, F\. Song, H\. Li, H\. Yu, Z\. Li, F\. Huang, and Y\. Li \(2023\)API\-bank: a comprehensive benchmark for tool\-augmented LLMs\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,Singapore,pp\. 3102–3116\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.187/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.187)Cited by:[Table 1](https://arxiv.org/html/2608.10042#S1.T1.1.4.2.1),[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px2.p1.1)\.
- X\. Li, R\. Zhou, Z\. C\. Lipton, and L\. Leqi \(2024\)Personalized language modeling from personalized human feedback\.arXiv preprint arXiv:2402\.05133\.Cited by:[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px1.p1.1)\.
- Q\. Lin, M\. Wen, Q\. Peng, G\. Nie, J\. Liao, J\. Wang, X\. Mo, J\. Zhou, C\. Cheng, Y\. Zhao,et al\.\(2024\)Hammer: robust function\-calling for on\-device language models via function masking\.arXiv preprint arXiv:2410\.04587\.Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- W\. Liu, X\. Huang, X\. Zeng, S\. Yu, D\. Li, S\. Wang, W\. Gan, Z\. Liu, Y\. Yu, Z\. WANG,et al\.\(2025\)Toolace: winning the points of llm function calling\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 41359–41381\.Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- H\. Luo and G\. Laban \(2026\)SPASM: stable persona\-driven agent simulation for multi\-turn dialogue generation\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 8455–8475\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.412),[Link](https://aclanthology.org/2026.findings-acl.412/)Cited by:[§B\.3](https://arxiv.org/html/2608.10042#A2.SS3.p2.1)\.
- Y\. Luo, H\. H\. Lam, Z\. Chen, Z\. Zhang, and X\. Feng \(2025\)ValuePilot: a two\-phase framework for value\-driven decision\-making\.arXiv preprint arXiv:2503\.04569\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- MadeAgents \(2024\)Hammer2\.1\-7B\.Note:Hugging Face model cardAccessed: 2026\-05\-25External Links:[Link](https://huggingface.co/MadeAgents/Hammer2.1-7b)Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- Moonshot AI \(2026\)Kimi K2\.6\.Note:[https://www\.moonshot\.ai/](https://www.moonshot.ai/)Accessed: 2026\-05\-25Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- OpenAI \(2026\)Introducing GPT\-5\.4\.Note:[https://openai\.com/index/introducing\-gpt\-5\-4/](https://openai.com/index/introducing-gpt-5-4/)Accessed: 2026\-05\-25Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- Y\. Qin, S\. Liang, Y\. Ye, K\. Zhu, L\. Yan, Y\. Lu, Y\. Lin, X\. Cong, X\. Tang, B\. Qian, S\. Zhao, L\. Hong, R\. Tian, R\. Xie, J\. Zhou, M\. Gerstein, D\. Li, Z\. Liu, and M\. Sun \(2024\)ToolLLM: facilitating large language models to master 16000\+ real\-world apis\.InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024,External Links:[Link](https://openreview.net/forum?id=dHng2O0Jjr)Cited by:[Table 1](https://arxiv.org/html/2608.10042#S1.T1.1.5.3.1),[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px2.p1.1)\.
- Qwen Team \(2026\)Qwen3\.6\-Plus: towards real world agents\.Note:[https://qwen\.ai/blog?id=qwen3\.6](https://qwen.ai/blog?id=qwen3.6)Accessed: 2026\-05\-25Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- A\. Salemi, S\. Mysore, M\. Bendersky, and H\. Zamani \(2024\)Lamp: when large language models meet personalization\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7370–7392\.Cited by:[Table 1](https://arxiv.org/html/2608.10042#S1.T1.1.3.1.1),[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Shao, J\. Wu, W\. Chen, and X\. Wang \(2025\)Personal travel solver: a preference\-driven llm\-solver system for travel planning\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 27622–27642\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- Y\. Shen, K\. Li, W\. Zhou, and S\. Hu \(2026\)Mem2ActBench: a benchmark for evaluating long\-term memory utilization in task\-oriented autonomous agents\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),San Diego, California, United States,pp\. 8173–8190\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.370),[Link](https://aclanthology.org/2026.acl-long.370/)Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px2.p1.1)\.
- W\. Shi, M\. Yuan, J\. Wu, Q\. Wang, and F\. Feng \(2024\)Direct multi\-turn preference optimization for language agents\.External Links:2406\.14868,[Link](https://arxiv.org/abs/2406.14868)Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- H\. Singh, N\. Verma, Y\. Wang, M\. Bharadwaj, H\. Fashandi, K\. Ferreira, and C\. Lee \(2024\)Personal large language model agents: a case study on tailored travel planning\.InProceedings of the 2024 conference on empirical methods in natural language processing: industry track,pp\. 486–514\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- Q\. Tang, Z\. Deng, H\. Lin, X\. Han, Q\. Liang, and L\. Sun \(2023\)ToolAlpaca: generalized tool learning for language models with 3000 simulated cases\.External Links:2306\.05301Cited by:[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2608.10042#S3.SS2.SSS0.Px2.p1.1)\.
- Team ACE \(2024\)ToolACE\-2\.5\-Llama\-3\.1\-8B\.Note:Hugging Face model cardAccessed: 2026\-05\-25External Links:[Link](https://huggingface.co/Team-ACE/ToolACE-2.5-Llama-3.1-8B)Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- P\. Wang, Y\. Wu, X\. Song, W\. Wang, G\. Chen, Z\. Li, K\. Yan, K\. Deng, Q\. Liu, S\. Zhao,et al\.\(2026a\)ShopSimulator: evaluating and exploring rl\-driven llm agent for shopping assistants\.arXiv preprint arXiv:2601\.18225\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Wang, C\. Liu, G\. Loo, L\. Zheng, K\. Wei, X\. Zeng, J\. Zhang, and Y\. Tian \(2026b\)Me\-agent: a personalized mobile agent with two\-level user habit learning for enhanced interaction\.arXiv preprint arXiv:2601\.20162\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- Watt AI \(2024\)watt\-tool\-8B\.Note:Hugging Face model cardAccessed: 2026\-05\-25External Links:[Link](https://huggingface.co/watt-ai/watt-tool-8B)Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- Z\. Xiu, D\. Q\. Sun, K\. Cheng, M\. Patel, Y\. Zhang, J\. Lu, O\. Attia, R\. Vemulapalli, O\. Tuzel, M\. Cao,et al\.\(2026\)ASTRA\-bench: evaluating tool\-use agent reasoning and action planning with personal user context\.arXiv preprint arXiv:2603\.01357\.Cited by:[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px2.p1.1)\.
- Q\. Xu, Y\. Li, H\. Xia, F\. Liu, M\. Yang, and W\. Li \(2025\)Petoolllm: towards personalized tool learning in large language models\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 21488–21503\.Cited by:[Table 1](https://arxiv.org/html/2608.10042#S1.T1.1.6.4.1),[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- Q\. Yang, H\. Li, H\. Zhao, X\. Yan, J\. Ding, F\. Xu, and Y\. Li \(2025\)Fingertip 20k: a benchmark for proactive and personalized mobile llm agents\.arXiv preprint arXiv:2507\.21071\.Cited by:[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhan \(2024\)τ\\tau\-Bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.External Links:2406\.12045,[Link](https://arxiv.org/abs/2406.12045)Cited by:[Table 1](https://arxiv.org/html/2608.10042#S1.T1.1.1.1),[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px2.p1.1)\.
- P\. Yu, W\. Liu, Y\. Yang, J\. Li, Z\. Zhang, X\. Feng, and F\. Zhang \(2026\)Benchmarking llm tool\-use in the wild\.arXiv preprint arXiv:2604\.06185\.Cited by:[Table 1](https://arxiv.org/html/2608.10042#S1.T1.1.9.7.1),[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px2.p1.1)\.
- Z\.ai \(2026\)GLM\-5: from vibe coding to agentic engineering\.Note:[https://z\.ai/blog/glm\-5](https://z.ai/blog/glm-5)Accessed: 2026\-05\-25Cited by:[§4\.1](https://arxiv.org/html/2608.10042#S4.SS1.p1.1)\.
- Z\. Zhao, C\. Vania, S\. Kayal, N\. Khan, S\. B\. Cohen, and E\. Yilmaz \(2025\)Personalens: a benchmark for personalization evaluation in conversational ai assistants\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 18023–18055\.Cited by:[§1](https://arxiv.org/html/2608.10042#S1.p2.1),[§2](https://arxiv.org/html/2608.10042#S2.SS0.SSS0.Px1.p1.1)\.

## Appendix AQualitative Case Study

This appendix provides a qualitative case study illustrating how UserToolBench evaluates personalized decision making beyond generic tool execution\. The case focuses on a GPT model’s behavior in a history\-dependent media\-intelligence task\. Although the current user request is underspecified, the prior interaction history contains enough evidence to recover the intended article and construct a grounded tool\-call trajectory\.

#### Case context\.

The user is a founder/operator of Midwest Home Decor who uses the assistant for competitive retail intelligence, brand\-narrative planning, and evidence\-based media analysis\. In the preceding interaction, the assistant discussed a Retail Daily article about GreenLeaf being named a retail rising star and noted that the story later received secondary coverage from 36Kr\. The conversation then shifted to how the user could build a comparable non\-sponsored authority signal for Midwest Home Decor, including a proposed “Tuesday Retail Notes” column centered on operational\-efficiency topics such as sales per square foot\.

The current user request is shown in Listing[1](https://arxiv.org/html/2608.10042#LST1)\. The request asks for metadata of the previously mentioned 36Kr follow\-up article, including the author, publication time, title rewrite angle, and editorial notes\. Crucially, the user does not provide the article URL and instead expects the assistant to recover it from the prior context\.

Listing 1:Current user request\. The request omits the URL and requires the assistant to recover the target article from history\.Iplantomakethefirstpieceinthe"TuesdayRetailNotes"seriesaboutsalespersquarefoot,

butIneedabenchmarkexamplefirst\.YoujustmentionedthattheGreenLeafarticlehad

asecondaryreporton36Kr\.Couldyoupullthemetadataforthatnewsbrief?

Iwanttoexamineitstitlerewriteangleandeditorialnotes\.IdidnotsavetheURL

onmyphone,butthesystemshouldbeabletomatchit,right?Pleasefirstpull

theauthorandpublicationtimeforme\.

#### History\-dependent personalization signals\.

The correct decision is not determined by the current utterance alone\. Several signals must be recovered from the interaction history\. The request should be interpreted as business media intelligence rather than a generic news lookup, because the user is using the article as a benchmark for Midwest Home Decor’s own brand\-positioning strategy\. The user also repeatedly distinguishes authoritative coverage from paid advertorials and asks for concrete metadata such as author, publication time, propagation path, rewrite angle, and editorial notes\. This makes grounded evidence important: the assistant should not invent a URL merely because the requested source and topic appear obvious\. The target entity and source are recoverable from the previous discussion: GreenLeaf is the competitor, and 36Kr is the secondary coverage source\. The missing URL should therefore be obtained by searching for the history\-derived query before calling the metadata tool\.

Table 6:Profile\-relevant signals and interaction\-history evidence for the selected case\.
#### Relevant previous interaction\.

The current request depends on two preceding pieces of context\. The earlier tool result established that the original Retail Daily report was followed by a 36Kr news brief, and the immediately preceding interaction introduced the user’s “Tuesday Retail Notes” plan\. Together, these turns explain why the user asks for title rewriting and editorial notes rather than only a URL\.

Listing 2:Relevant previous interaction snippets\. These turns supply the missing article identity and the user’s personalized writing objective\.Earlierassistantresponse:

Estimatedpropagationpathandreach:

\-Originalplatform:RetailDailyfrontpage,approximately45,000impressions\.

\-Secondaryrepost:authorizedrepostbytheWeChataccount"RetailBusinessReview",

approximately28,000impressions\.

\-Socialamplification:repostedbykeyopinionleadersinLinkedInindustrygroups,

approximately12,000impressions\.

\-Mediafollow\-up:secondary36Krnewsbrief,approximately65,000impressions\.

Immediatelyprevioususerturn:

Let’sputasidethemetadataoftheGreenLeafarticleforamoment\.

Iwanttocorrectmyrequest\.IfIwantMidwestHomeDecortoobtain

acomparablelevelofauthoritativeendorsementoverthenexttwoquarters,

insteadofrelyingonpaidadvertorialsforvisibility,whatthreedimensions

shouldIusetoreconstructmybrandnarrative?

Immediatelypreviousassistantresponse:

Dimension3:contextualizingthefounder’spublicvoicefrom"businessowner"

to"industrycommentator\."

Coreaction:

1\."TuesdayRetailNotes"LinkedIncolumn:publisheverytwoweeks,witheach

pieceunder500words,usinganonymizedoperatingdatatocommentonindustrytrends\.

Exampletopicsinclude"WhyInolongerlookatGMVandinsteadfocusonsales

persquarefoot"and"Howa12\-personteammanagesthreestoresandoneonlineshop\."

#### Reference trajectory\.

The reference trajectory first calls a search tool with history\-derived keywords, then passes the returned URL to the metadata tool\. This sequence converts an underspecified user request into a grounded executable trajectory\.

Listing 3:Reference tool\-call trajectory\. The assistant first grounds the missing URL through search, then retrieves article metadata\.searchArticles\(\{

"keywords":"36KrGreenLeafretailrisingstar",

"limit":10

\}\)

Observation:

\{

"articles":\[

\{

"title":"GreenLeafFreshNamedRetailRisingStaroftheYear,

CommunityGroup\-BuyingModelDrawsIndustryAttention",

"url":"https://36kr\.com/p/2847291",

"publishedAt":"2024\-07\-19T08:00:00Z",

"source":"36Kr",

"author":"YangLin"

\}

\]

\}

getArticleMetadata\(\{

"url":"https://36kr\.com/p/2847291"

\}\)

Observation:

\{

"title":"GreenLeafFreshNamedRetailRisingStaroftheYear,

CommunityGroup\-BuyingModelDrawsIndustryAttention",

"author":\{"name":"YangLin","title":"36Krconsumerdeskreporter"\},

"publish\_date":"2024\-07\-19T08:00:00Z",

"content\_type":"newsbrief",

"sponsored":false,

"rewrite\_angle":"Usesafinancing\-orientedangle,emphasizinggrowthdata

andcapital\-marketrecognition\.",

"editorial\_notes":\{

"data\_source":"RetailDaily’soriginalreportandGreenLeaf’spublicfinancinginformation",

"editor\_remarks":"Thetitleweakensrhetoricalwordingsuchas’sweepingthemarket’

andinsteaduses’drawsindustryattention’tomaintainobjectivity;

thebriefaddsfinancinginformationwhilepreservingthecoredatapoints\."

\}

\}

prepare\_to\_answer\(\{"answer\_type":"tool"\}\)

#### GPT prediction\.

The GPT model receives the long interaction history and the available tools, including bothsearchArticlesandgetArticleMetadata\. However, instead of first grounding the missing URL through search, it directly calls the metadata tool with a plausible but unsupported URL\.

Listing 4:GPT prediction\. The model skips the grounding step and fills the missing URL argument by guesswork\.Availabletools:

filterArticles,filterArticlesBySource,getArticleMetadata,

getHeadlines,getTrendingNews,searchArticles

GPTtoolcall:

getArticleMetadata\(\{

"url":"https://www\.36kr\.com/greenleaf\-rising\-star\-2024"

\}\)

#### Analysis\.

This case illustrates why executable\-looking tool use is not sufficient for personalized decision making\. The predicted call is syntactically well formed and uses a relevant metadata tool, but it fails as a personalized decision because the missing URL is a decision\-critical argument that should be grounded in the user’s interaction history\. The user’s request combines a prior entity, a prior media source, and a personalized writing objective; therefore, the correct behavior is not to directly fill the URL slot, but to recover the article throughsearchArticlesand then callgetArticleMetadata\.

The case also clarifies the distinction between surface\-level task completion and profile\-conditioned decision alignment\. A generic assistant might treat the request as a simple metadata retrieval problem and attempt to guess the most likely article URL\. In contrast, a reliable personalized delegate should preserve the user’s preference for authoritative, non\-sponsored evidence and avoid unsupported argument filling\. Thus, the failure is not merely a tool\-formatting error\. It reflects a breakdown in history\-dependent preference inference, uncertainty handling, and grounded tool\-call planning, which are precisely the capabilities UserToolBench is designed to evaluate\.

## Appendix BPersona Profile Construction and Field Coverage

### B\.1Candidate Profile Selection

UserToolBench uses structured persona profiles as persistent user states for constructing personalized decision\-making tasks\. We start from 100 candidate persona profiles derived from privacy\-sanitized interaction evidence\. The current release selects 10 representative profiles through an LLM\-assisted and human\-verified filtering process\. The LLM is used to summarize candidate profile coverage and identify profiles with sufficiently rich decision\-relevant preferences, while human annotators verify representativeness, internal consistency, privacy safety, and suitability for tool\-based assistant scenarios\.

The selected profiles share a common high\-level schema, including persona summary, communication style, preference categories, task habits, and optional recurring constraints\. In addition to this shared schema, each profile may contain a small number of profile\-specific fields\. These fields are retained only when they describe recurring, non\-identifying, and task\-relevant behavioral factors\.

### B\.2Shared and Profile\-Specific Schema

Table 7:High\-level schema of structured persona profiles\. Sensitive or identifying details are abstracted or removed before benchmark construction\.Table[7](https://arxiv.org/html/2608.10042#A2.T7)summarizes the shared high\-level schema used across persona profiles\. Table[8](https://arxiv.org/html/2608.10042#A2.T8)reports the field coverage of the 10 selected profiles\. For privacy reasons, potentially sensitive profile\-specific fields are reported at the category level rather than as raw field names\.

Table 8:Field coverage of the 10 selected persona profiles\. All profiles share the high\-level schema in Table[7](https://arxiv.org/html/2608.10042#A2.T7)\. Profile\-specific fields are summarized at the category level to avoid exposing sensitive or identifying details\.
### B\.3Quantitative Persona Diversity Audit

We complement the field\-coverage analysis with a quantitative audit of the ten selected profiles\. Demographic descriptors span ages 22–60, with five female and five male profiles, seven nationality values, ten education backgrounds, and ten occupations\. These counts characterize coverage in the selected sample; they do not imply representativeness of the broader user population\.

Table 9:Summary coverage of selected persona descriptors\.Following prior work on stable persona simulation, controllable persona generation, and lexical\-diversity auditing\[Luo and Laban,[2026](https://arxiv.org/html/2608.10042#bib.bib48), Jianget al\.,[2026](https://arxiv.org/html/2608.10042#bib.bib49), Kambhatlaet al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib50)\], we quantify overlap both within matched profile fields and across complete non\-sensitive persona text\. For matched textual fields, we report the number of items, the fraction of unique entries, average TF–IDF cosine similarity, average token Jaccard similarity, corrected type–token ratio \(CTTR\), and moving\-average type–token ratio \(MATTR\)\. Higher Unique, CTTR, and MATTR indicate less repetition or greater lexical diversity; lower TF–IDF and Jaccard values indicate lower cross\-profile overlap\.

Table 10:Field\-level diversity of the selected persona profiles\. Similarity values are lower\-is\-more\-diverse; lexical\-diversity values are higher\-is\-more\-diverse\.At the full\-profile level, we evaluate all\(102\)=45\\binom\{10\}\{2\}=45persona pairs using only non\-sensitive profile text\. Across the 45 pairs, mean/median token Jaccard similarity is0\.197/0\.1940\.197/0\.194\(range0\.1690\.169–0\.2280\.228\), and mean/median TF–IDF cosine similarity is0\.242/0\.2500\.242/0\.250\(range0\.1720\.172–0\.3380\.338\)\. Persona 2 and Persona 7 form the most similar pair under both metrics, with Jaccard0\.2280\.228and TF–IDF cosine0\.3380\.338\. Figure[3](https://arxiv.org/html/2608.10042#A2.F3)visualizes the complete pairwise audit as heatmaps rather than raw pair tables, fulfilling the complete pairwise audit while keeping the presentation compact\.

![Refer to caption](https://arxiv.org/html/2608.10042v1/x1.png)

![Refer to caption](https://arxiv.org/html/2608.10042v1/x2.png)

Figure 3:Heatmap visualization of the complete pairwise full\-profile similarity\. Left: token Jaccard similarity\. Right: TF–IDF cosine similarity\. Lower values indicate greater diversity\. Diagonal cells are masked because self\-similarity is not part of the pairwise audit\.Together with Table[10](https://arxiv.org/html/2608.10042#A2.T10)and the trajectory\-level diversity results in Table[3](https://arxiv.org/html/2608.10042#S3.T3), the pairwise heatmaps show that the selected profiles differ in both textual preference inventories and induced tool\-call behavior\. These results establish diversity within this controlled ten\-profile sample, while not implying population\-level representativeness\.

### B\.4Privacy and Usage Constraints

The selected persona profiles are behavioral abstractions rather than raw user records\. Raw conversations are not included in the released benchmark instances\. Sensitive or identifying details are removed, generalized, or replaced with non\-identifying preference signals before task construction\. Profile fields are used to support decision\-making tasks only when they are relevant to user preferences, habits, communication style, or task constraints\.

Potentially sensitive fields are not used to encourage demographic inference or stereotype\-based personalization\. During task construction and validation, such fields are either generalized into non\-identifying categories or excluded unless they are explicitly sanitized and necessary for a benign task constraint\. Tasks requiring unsupported inference over sensitive attributes are removed during validation\.

## Appendix CTool Ecosystem Construction

### C\.1Raw Tool Collection and Scenario\-Level Aggregation

We construct the tool ecosystem in UserToolBench from a large pool of API\-style tool names\. The initial tool pool contains 1,248 tool\-name occurrences\. Since many raw tools are overlapping, redundant, or too fine\-grained to form coherent assistant tasks, we aggregate them into scenario\-level toolsets rather than treating each tool independently\.

Specifically, we first use an LLM\-assisted clustering procedure to group semantically related tools into candidate toolsets\. Each toolset is intended to correspond to a realistic personal\-assistant scenario, such as travel planning, weather searching, and so on\. The clustering stage produces 256 candidate toolset lines, covering 618 unique tool names\. On average, each candidate topic contains 4\.875 tool names, and each tool contains 2\.45 parameter slots\.

StatisticValueRaw tool\-name occurrences1,248Candidate toolset lines after LLM grouping256Unique tool names after grouping618Average tool names per topic4\.875Average parameter slots per tool2\.45Final manually selected toolsets36Table 11:Statistics of the tool ecosystem construction process\. Raw API\-style tools are first grouped into scenario\-level candidate toolsets with LLM assistance and then manually filtered to retain coherent, preference\-sensitive personal\-assistant scenarios\.
### C\.2Manual Filtering

The LLM\-grouped candidate toolsets are manually filtered before inclusion in the benchmark\. The goal is not to maximize the number of tools, but to retain toolsets that support meaningful personalized decision making under user\-specific preferences\. A retained toolset must form a coherent task scenario rather than a loose collection of unrelated APIs\. For example, a travel\-planning toolset may include flight search, flight booking, hotel search, hotel booking, and local activity recommendation tools, because these tools naturally support a multi\-step assistant workflow\.

A retained toolset must also contain decision points that can be influenced by user preferences\. We remove toolsets whose outputs are almost entirely determined by explicit user instructions and leave little room for preference\-sensitive decisions\. This criterion matters because UserToolBench is designed to evaluate personalized decision making rather than generic API invocation\.

Finally, the toolset must be operationally usable in multi\-turn trajectories\. We remove toolsets with underspecified functionality, incompatible argument schemas, excessive redundancy, or unclear execution dependencies\. When several candidate toolsets cover similar scenarios, we retain the one with clearer tool semantics and more diverse decision\-relevant arguments\.

After manual filtering, 36 distinct scenario\-level toolsets remain in the final benchmark\. These toolsets constitute the available tool ecosystems used to construct and evaluate UserToolBench trajectories\.

### C\.3Decision\-Relevant Tool Arguments

For each retained tool, we additionally identify decision\-relevant argument slots\. A decision\-relevant slot is an argument whose value may change depending on the user’s current request, prior interaction history, or stable preferences\. Examples include budget constraints, preferred location, time range, service style, ranking criterion, transportation mode, and accommodation preference\. In contrast, purely technical or formatting arguments, such as request identifiers, pagination size, or fixed output format, are not treated as personalization\-sensitive decision slots\.

This distinction is used during both data construction and evaluation\. During reference trajectory construction, preference\-derived argument values are checked against the user profile and interaction history\. During evaluation, matching decision\-relevant arguments is treated as stronger evidence of personalized decision alignment than matching non\-decision arguments\.

## Appendix DAdditional Diagnostic Analyses

### D\.1Multi\-Label Failure Analysis

We align failed predicted trajectories with their profile\-conditioned references and assign one or more diagnostic labels\.Wrong Tooldenotes selection of a tool inconsistent with the intended personalized action;Missing Tooldenotes omission of a required call;User Constraintdenotes violation or loss of an explicit or history\-grounded user constraint;Implicit Preferencedenotes failure to recover a relevant latent preference; andSequence/Dependencydenotes an incorrect ordering, missing dependency, or failure to pass information between calls\. Because labels are multi\-label, percentages within a row need not sum to100%100\\%\.

Table 12:Incidence \(%\) of multi\-label error types among failed trajectories for seven analyzed models\.Sequence and dependency errors are the most frequent category for every analyzed model, indicating that personalized failures often propagate through multi\-step execution rather than appearing as isolated formatting errors\. Wrong\-tool and user\-constraint errors are also common, supporting the interpretation that many exact mismatches reflect substantive decision failures\. Implicit\-preference labels are less frequent because a latent preference error often becomes observable downstream as a wrong tool, omitted call, or violated constraint; the labels should therefore be interpreted jointly rather than as mutually exclusive causal categories\.

### D\.2Severity\-Aware Decision Diagnostics

We further distinguishSlight Preference Deviation, where the model selects the correct tool type but makes a limited departure from the optimal personalized decision, fromMajor Decision Deviation, where it selects a substantially different tool path, omits a required call, or violates a user\-specific constraint\.

Table 13:Reported diagnostic rates \(%\) for slight and major personalized decision deviations\.Across models, major deviations are substantially more frequent than slight deviations under this rubric\. This result helps qualify the exact\-match analysis: while some valid alternative trajectories may be penalized, the observed strict\-performance gap cannot be attributed only to harmless preference\-compatible variation\. The diagnostic remains an auxiliary analysis rather than a replacement for multi\-reference evaluation\.

### D\.3Preliminary Dynamic\-Preference Split

The main benchmark uses stable profiles to isolate history\-grounded preference inference\. To test whether the construction framework can represent evolving preferences, we additionally build a preliminary split from 10 PersonaMem profiles containing temporal preference updates\[Jianget al\.,[2025](https://arxiv.org/html/2608.10042#bib.bib8)\]\. Preference changes are represented as interaction events, and the evaluated model receives only the observed history, current request, and tools\. It must follow the most recent applicable preference rather than the original profile state\.

Table 14:Exact trajectory accuracy \(%\) on the preliminary dynamic\-preference split\.The split confirms that the milestone\-based pipeline can encode preference updates and evaluate whether a model follows the latest applicable state\. Performance remains limited, particularly on multi\-tool tasks, and all three models show a pronounced reduction in the middle third\. These results should be interpreted as an extensibility study rather than a comprehensive benchmark of preference evolution; the split does not yet cover conflicts, reversals with uncertainty, or context\-specific exceptions at scale\.

## Appendix ECompact Prompt Specifications for Trajectory Synthesis

This appendix summarizes the role contracts used in the multi\-role synthesis pipeline\. We intentionally omit repeated formatting instructions and long enumerations that do not affect the methodological description\. The complete executable prompts, including all construction modes and provider\-specific wrappers, are released with the code at[https://github\.com/xxy212/UserToolBench/tree/main/utb/generation/agent](https://github.com/xxy212/UserToolBench/tree/main/utb/generation/agent)\.

The synthesis system uses a persona\-conditioned user simulator, a planner, an assistant LLM, external tools, and checker roles\. Instance\-specific fields include the persona, dialogue history, environment information, tool schemas, required arguments, planner outputs, and tool observations\. Explicit persona profiles are available only during data construction and validation; evaluated models receive interaction history rather than the structured profile\.

Table 15:Compact role specifications used for trajectory synthesis and validation\.Listing 5:Abbreviated role prompts\. Full prompts are released in the repository\.USERSIMULATOR

Inputs:persona,history,environment,tools,generationmode\.

Instruction:actasthesameenduser;expressanaturaltaskorreply

consistentwiththepersona\.Donotnametheintendedtool\.Whenthemode

isambiguous,omitselectedrequiredinformationdeliberately\.

Output:User:<onedialogueturn\>

ASSISTANTLLM

Inputs:plannerdecision,toolobservations,history,persona,environment\.

Instruction:followtheplannerandobservationsfaithfully\.Askaconcise

questiononlywhenarequiredconstraintcannotberecovered;otherwise

summarizeresultsoranswerdirectlyinapersona\-compatiblestyle\.

Output:Assistant:<user\-facingreply\>

CHECKER

Inputs:request,history,persona\-conditionedcontext,environment,tools,

planneroutput,andtoolobservations\.

Instruction:verifytoolselection,requiredanddecision\-relevantarguments,

clarificationbehavior,serial/paralleldependencies,observationformat,

andconsistencywithpriorturns\.Returnavaliditydecisionandreasons\.

### E\.1Human Verification Interface

To support human verification of the constructed trajectories, we built a lightweight web\-based quality\-control interface\. The interface presents each benchmark instance together with its associated persona, dialogue messages, available tool set, and metadata\. The goal of the interface is to help human verifiers inspect whether each trajectory is valid, profile\-consistent, and safe to include in the benchmark\.

The interface is organized into three panels\. The left panel shows sample\-level metadata, including the source file name, persona directory, toolset index, topic identifier, topic turns, node path, and environment information\. It also displays a summarized persona\. The full persona can be expanded when necessary, but thesensitive\_informationfield is hidden by default to reduce unnecessary exposure of sensitive content during verification\.

The middle panel displays the complete message sequence for the current sample\. Messages are visually separated by role, including user, assistant, tool, and other roles\. This allows verifiers to inspect whether the conversation is coherent, whether the assistant follows the expected role protocol, and whether tool observations are used correctly in subsequent responses\. If a sample cannot be parsed as valid JSON, the interface explicitly reports the parsing error\.

The right panel displays the available tool set\. For each tool, the interface shows the tool name, natural\-language description, required arguments, and detailed parameter schema\. Verifiers use this information to check whether the assistant selected the appropriate tool, supplied valid arguments, avoided hallucinated parameters, and asked for clarification when required information was missing\.

For each sample, verifiers record one of three decisions:Pass,Fail, orSkip\. Failed samples can be assigned an issue type from a predefined taxonomy: incorrect tool selection, incorrect or hallucinated parameter, missing required clarification, multi\-turn inconsistency, persona inconsistency, incorrect use of tool observations, role protocol error, sensitive information leakage, format or JSON error, or other\. Verifiers may also provide free\-form notes to explain the failure reason or document borderline cases\.

The interface also provides a verification checklist\. Verifiers are asked to examine whether the user request is natural and reasonably underspecified, whether the planner chooses the correct tool, whether missing required parameters trigger clarification, whether previously provided parameters are reused correctly, whether tool observations are consistent with the agent response, whether later turns properly follow from earlier turns, and whether the sample avoids unnecessary disclosure of sensitive information\. This human verification step helps filter invalid, inconsistent, or privacy\-risky examples before inclusion in the final benchmark\.

Similar Articles

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

arXiv cs.CL

This paper introduces PhysTool-Bench, a benchmark for evaluating multimodal large language models' ability to recognize and plan the use of physical tools in real-world scenes. The authors find that even the best model identifies only 58.7% of tools and completes just 21.0% of queries end-to-end, revealing a two-level deficit in perception and functional commonsense.