Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling
Summary
This paper argues that full-horizon planning with lazy replanning is more efficient than step-by-step execution for data-centric LLM agent tasks, using fewer tokens while maintaining accuracy.
View Cached Full Text
Cached at: 05/12/26, 06:50 AM
# Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling Source: [https://arxiv.org/html/2605.08477](https://arxiv.org/html/2605.08477) ,Nikita Bhutani[nikita@megagon\.ai](https://arxiv.org/html/2605.08477v1/mailto:[email protected])Megagon LabsMountain ViewCaliforniaUSA,Hannah Kim[hannah@megagon\.ai](https://arxiv.org/html/2605.08477v1/mailto:[email protected])Megagon LabsMountain ViewCaliforniaUSA,Dan Zhang[dan˙z@megagon\.ai](https://arxiv.org/html/2605.08477v1/mailto:dan%CB%[email protected])Megagon LabsMountain ViewCaliforniaUSAandEstevam Hruschka[estevam@megagon\.ai](https://arxiv.org/html/2605.08477v1/mailto:[email protected])Megagon LabsMountain ViewCaliforniaUSA \(2026\) ###### Abstract\. Explicit planning is a critical capability for LLM\-based agents solving complex data\-centric tasks, which require precise tool calling over external data sources\. Existing strategies fall into two paradigms based on planning horizon: \(1\) full\-horizon \(FH\), which generates a complete plan before execution, and \(2\) single\-step horizon \(SH\), which interleaves each action \(tool call\) with incremental reasoning and observation\. While step\-by\-step execution is a common default under the assumption thateagerexecution monitoring is necessary for adaptability, we revisit this assumption for well\-defined data\-centric tasks\. Our controlled empirical study isolates planning horizon as the key architectural feature and systematically analyzes the effects of topological complexity and tool robustness on both paradigms\. Our experiments across Knowledge Base Question Answering and Multi\-hop QA show that FH planning withlazyreplanning achieves accuracy parity with SH across varying depths, breadths, and robustness levels, while using22–3×3\\timesfewer tokens\. These findings suggest that for well\-defined data\-centric tasks, eager step\-wise monitoring is often unnecessary, and full\-horizon planning with on\-demand replanning can offer a more efficient default\. Large language model agents, tool\-calling, data\-centric tasks ††journalyear:2026††copyright:cc††conference:ACM Conference on AI and Agentic Systems; May 26–29, 2026; San Jose, CA, USA††booktitle:ACM Conference on AI and Agentic Systems \(CAIS ’26\), May 26–29, 2026, San Jose, CA, USA††doi:10\.1145/3786335\.3813129††isbn:979\-8\-4007\-2415\-2/2026/05††submissionid:52††ccs:Information systems Question answering††ccs:Computing methodologies Natural language generation## 1\.Introduction Large Language Model \(LLM\) agents are increasingly deployed to solvedata\-centric tasksin which answers must be constructed through tool calls over external sources such as databases, knowledge graphs, or documents \(Figure[1](https://arxiv.org/html/2605.08477#S1.F1)\)\. In these settings, success depends on coordinating tool calls that match the latent logic \(e\.g\., joins or multi\-hop reasoning\) and vocabulary constraints imposed by the data source \(e\.g\., data schema, entity mentions\)\. As these tasks grow in complexity, explicit planning of low\-level tool calls has become a central component of modern agentic architectures\(Gu et al\.,[2024](https://arxiv.org/html/2605.08477#bib.bib12); Xin et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib39); Xiong et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib40)\)\. \(a\)In data\-centric tasks, LLM agents must coordinate tool calls to synthesize the answer from external data sources\. This complex tool calling necessitates explicit planning \(below\)\. \(b\)Single\-step horizon \(SH\) alternates planning and execution at each step \(*eager*monitoring\)\. Full\-horizon \(FH\) plans upfront and \(optionally\) replans only on demand \(*lazy*monitoring\)\. \(c\)When GPT\-4\.1\-mini is used as a backbone LLM, SH and FH achieve comparable accuracy across datasets \(left\), but SH consumes much more input\+output tokens \(right\)\. See §[4](https://arxiv.org/html/2605.08477#S4)for details\. Figure 1\.Planning horizon in data\-centric tool calling\.SH plans and executes step\-by\-step\. FH plans ahead and replans only when needed, which can reduce token consumption\.A three\-part overview of planning horizon in data\-centric tool calling\. The top panel shows that answering data\-centric questions requires coordinating multiple tool calls over external sources\. The middle panel contrasts single\-step horizon planning, which alternates planning and execution after each tool call, with full\-horizon planning, which generates a multi\-step tool\-use plan upfront and replans only when needed\. The bottom panel summarizes experimental results showing similar accuracy between the two approaches across datasets, while single\-step horizon uses substantially more tokens\.Existing planning techniques can be categorized into two major paradigms byplanning horizon, the number of steps planned before tool execution \(Figure[1](https://arxiv.org/html/2605.08477#S1.F1)\)\.Single\-step horizon \(SH\)planning interleaves reasoning and execution, calling one tool at a time based on prior observations\.Full\-horizon \(FH\)planning instead generates a complete plan upfront before execution\. Recent frontier models and systems increasingly support upfront high\-level task decomposition, and more advanced FH planning techniques have also been developed\(Xu et al\.,[2023](https://arxiv.org/html/2605.08477#bib.bib41); Li et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib19)\)\. However, tight “think\-act\-observe” loops\(Yao et al\.,[2023](https://arxiv.org/html/2605.08477#bib.bib45)\)remain a common default for low\-level tool execution\.111In this paper, we focus on thislow\-level execution layerrather than higher\-level decomposition\.Thiseagermonitoring is often presumed essential for robustly handling the opacity and potential noise of external tools and data sources\(Kim et al\.,[2024b](https://arxiv.org/html/2605.08477#bib.bib17); Gonzalez\-Pumariega et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib9); Zhang et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib47)\)\. Recent work byLiu et al\.\([2025](https://arxiv.org/html/2605.08477#bib.bib22)\)has begun to question whether interleaved planning is universally optimal\. They compare different planning strategies on general reasoning tasks without tool interaction and find that SH planning is not consistently superior\. However, their study leaves open a critical question: does this conclusion extend to data\-centric settings, where success depends on structured tool calls rather than purely internal reasoning? We perform a controlled empirical study and address this gap by shifting the focus from abstract reasoning todata\-centric, tool\-calling tasks, in which planning decisions directly affect execution success, computational cost, and robustness\. We further go beyond dataset\-level comparisons by analyzing planning behavior at the instance level, enabling a more precise characterization of task difficulty\. Specifically, we introduce an instance\-level framework that disentangles two orthogonal dimensions:topological complexity\(the depth and breadth of the execution graph\) andtool robustness\(the tolerance to imprecise inputs\)\. Together, these dimensions capture two fundamental challenges in tool\-mediated data\-centric tasks: satisfying logical dependencies among intermediate steps and aligning generated arguments with external schema or vocabulary constraints\. We hypothesize that for well\-defined data\-centric tasks, an FH planner equipped withlazymonitoring \(executing a complete plan and replanning only upon failure\) can match the performance of an SH planner without the massive overhead of continuous feedback integration\. We focus on Knowledge Base Question Answering \(KBQA\) and Multi\-hop QA \(HotpotQA\) because they represent the two core challenges of data\-centric agents\. KBQA serves as a controlled environment for testing logic\-alignment, requiring agents to coordinate atomic tool operations that mirror database queries\. HotpotQA represents the challenges in unstructured settings, where agents must navigate retrieval noise and coordinate reasoning\-based sub\-agents\. This setup allows us to analyze planning behavior across both rigid, structured schemas and fuzzier, unstructured data sources\. Our results show no statistical evidence that SH planning provides a performance advantage over FH planning across varying structural or robustness configurations in well\-defined data\-centric tasks\. Given this performance parity, the22–3×3\\timesefficiency advantage of FH planning \(Figure[1](https://arxiv.org/html/2605.08477#S1.F1)\) makes it a stronger option than previously thought\. Our findings also suggest that the perceived brittleness of FH planning in prior work may result from the absence of proper recovery mechanisms rather than an inherent limitation of full\-horizon planning\. While SH planning may remain advantageous in exploratory or highly dynamic tool\-calling tasks, our results refine prevailing assumptions about SH by demonstrating that eager monitoring is not universally necessary\. In structured and stable environments, less frequent monitoring can substantially improve efficiency without sacrificing accuracy\. Our contributions are the following: - •We isolate planning horizon as a core architectural variable in LLM agents \(§[2](https://arxiv.org/html/2605.08477#S2)\) and provide a controlled comparison of its effects in data\-centric tasks\. - •We introduce an instance\-level framework that characterizes difficulty via execution graph topology and tool robustness \(§[3](https://arxiv.org/html/2605.08477#S3)\)\. In particular, we identify depth and breadth as overlooked axes of execution\-graph complexity that affect planning performance beyond sequential length\. - •We show that FH planning with lazy replanning achieves accuracy parity with SH planning while using fewer tokens on well\-defined data\-centric tasks \(§[4](https://arxiv.org/html/2605.08477#S4)\)\. This result provides a foundation for future work on adaptive and hybrid planners\. ## 2\.Planning Horizon : The Key Architectural Feature The landscape of agentic frameworks is expansive\(Huang et al\.,[2024](https://arxiv.org/html/2605.08477#bib.bib13); Li,[2025](https://arxiv.org/html/2605.08477#bib.bib20); Wei et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib36)\)\. Yet, most approaches can be understood through a single underlying design choice:planning horizon, defined as the number of steps an agent plans before execution\. Planning horizon determines how much simulation an agent performs prior to interacting with tools and therefore governs when feedback from the environment is incorporated into planning\. We conceptualize the agent as a policyπ\\piinteracting with a tool execution environment𝔼\\mathbb\{E\}\. Let𝔸\\mathbb\{A\}denote the set of available tools \(actions\) and𝕆\\mathbb\{O\}the space of possible observations \(tool outputs\)\. Given a user queryqq, the agent produces a plan𝒫=\(a1,…,aT\)\\mathcal\{P\}=\(a\_\{1\},\\dots,a\_\{T\}\)consisting of actionsai∈𝔸\(i=1,…,T\)a\_\{i\}\\in\\mathbb\{A\}\(i=1,\\dots,T\), whereTTdenotes the plan length andaTa\_\{T\}is the final planned action\. Here, a plan refers specifically to a sequence of tool calls at the execution level\. Executing an actionaia\_\{i\}yields an observationoi∈𝕆o\_\{i\}\\in\\mathbb\{O\}, resulting in a trajectory that leads to the answer toqq\. Within this formulation, planning strategies differ only in when the policyπ\\piis invoked during execution\. In this section, we treat planning horizon as a primary architectural feature and discuss how it shapes the trade\-off between adaptivity and cost\. In particular, we focus on when and how agents replan in response to tool feedback\. ### 2\.1\.Single\-Step Horizon \(SH\) The single\-step horizon \(SH\) paradigm tightly interleaves planning and execution\. The agent simulates only one step ahead before performing an action\. In this paradigm, action generation at stepttis conditioned on the user queryqqand the history of previous actions and observations: \(1\)at∼π\(q,a1,o1,…,at−1,ot−1\)\\displaystyle a\_\{t\}\\sim\\pi\(q,a\_\{1\},o\_\{1\},\.\.\.,a\_\{t\-1\},o\_\{t\-1\}\)This design entailseagerfeedback monitoring, since the agentalwaysprocesses observations up to the current step before deciding on the next action\. SH is widely used in modern agentic applications, with some variations that can plan concurrent, independent steps at once\. Its popularity stems from the tight planning–acting feedback loop, which enables robust adaptation to uncertainty and noise in external tools and environments\. Such adaptivity is particularly beneficial in exploratory tasks that require information gathering before subsequent planning\(Kim et al\.,[2024b](https://arxiv.org/html/2605.08477#bib.bib17); Gonzalez\-Pumariega et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib9); Zhang et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib47)\)\. ### 2\.2\.Full\-Horizon \(FH\) The full\-horizon \(FH\) paradigm generates a complete execution graph upfront\. The agent performs a full simulation of the trajectory required to solve the task before it triggers any tool execution\. In this paradigm, the agent generates an initial plan as a complete sequence of actions𝒫=\(a1,…,aT\)\\mathcal\{P\}=\(a\_\{1\},\\dots,a\_\{T\}\)conditioned only on the query: \(2\)𝒫∼π\(q\)\\displaystyle\\mathcal\{P\}\\sim\\pi\(q\) Execution then proceeds via the environment𝔼\\mathbb\{E\}without invoking the policyπ\\piat every step\. Unlike SH, FH can allowlazyfeedback integration, wherein observations are incorporated only when monitoring is triggered\. For example, an execution failure at stepkkcan trigger monitoring, prompting the agent to replan based on the observed trajectory: \(3\)𝒫′∼π\(q,a1,o1,…,ak,ok\)\\displaystyle\\mathcal\{P\}^\{\\prime\}\\sim\\pi\(q,a\_\{1\},o\_\{1\},\.\.\.,a\_\{k\},o\_\{k\}\) ### 2\.3\.Implications of Planning Horizon Planning horizon creates a trade\-off between adaptivity and computational cost\. SHeagerlymonitors every tool call, which can yield robust adaptation to uncertainty and noise but incurs substantial overhead due to repeated inference\. FH, by contrast, generates a multi\-step plan in a single pass and is therefore more efficient, but only integrates feedbacklazily\. Comparing FH’s lazy feedback integration \(Eq\. \([3](https://arxiv.org/html/2605.08477#S2.E3)\)\) with SH’s single\-step generation \(Eq\. \([1](https://arxiv.org/html/2605.08477#S2.E1)\)\), we see that they both condition on the same action–observation history available at that step\. Thus, the practical difference is not*what*information is used, but*when*it is used: at every step \(SH\) versus only when triggered \(FH\)\. This insight suggests a testable hypothesis: with an appropriate monitoring trigger, FH should achieve performance comparable to SH while requiring fewer policy calls\. From this perspective, the commonly assumed brittleness of FH may arise not from its planning horizon per se, but from the absence of effective error recovery mechanisms\. Notably, many existing studies implement FH without replanning\(Gonzalez\-Pumariega et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib9); Zhang et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib47)\), which confounds evaluation of the paradigm\. We therefore evaluate FH equipped with lazy monitoring\. To obtain generalizable insights, we adopt a simple trigger that is commonly used in practice: replanning upon execution failure\. We also abstract away implementation\-specific details such as tool\-calling formats and isolate planning horizon as the primary variable in our analysis\. ## 3\.Task Characterization To evaluate how planning horizon shapes agent behavior, we require a principled way to characterize task difficulty\. Dataset\-level comparison\(Liu et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib22)\)can obscure meaningful variation across instances\. Instances actually differ substantially in which tools and data sources are involved and how they are connected\. Without explicitly modeling suchinstance\-levelfeatures, comparisons between planning strategies may conflate planning effects with underlying instance properties\. In data\-centric tool\-calling tasks, success depends on*aligning logic and vocabulary*: correctly composing tool calls to match the latent dependency structure of the task instance \(logic alignment\) and specifying tool parameters that match the schema or representation of the external data source \(vocabulary alignment\)\. capture these challenges using two independent dimensions: topological complexity and tool robustness\. This separation allows us to disentangle structural reasoning difficulty from environmental uncertainty\. These factors are often conflated in prior evaluations\. ### 3\.1\.Topological Complexity \(Logic Alignment\) Figure 2\.An example plan DAG for the query: “Which company has more employees: Google or the parent company of Instagram?”\. The DAG involves two parallel paths: a single\-hop lookup for Google \(a1a\_\{1\}\) and a two\-hop traversal for Instagram’s parent company \(a2→a3a\_\{2\}\\rightarrow a\_\{3\}\), which are finally compared at the aggregation node \(a4a\_\{4\}\)\. This structure results in a depth ofd=3d=3and a breadth ofb=4/3b=4/3\.A directed acyclic graph illustrating the execution plan for comparing the employee counts of Google and Instagram’s parent company\. One branch performs a single\-hop lookup for Google, while the other performs a lookup followed by a parent\-company traversal for Instagram\. The two branches are merged by a final comparison node\. The figure highlights a critical\-path depth of three steps and an average breadth of four\-thirds\.Topological complexity captures the difficulty of aligning execution with the latent logical structure of a task\. We represent a plan as a directed acyclic graph \(DAG\)G=\(V,E\)G=\(V,E\), where nodes \(VV\) are tool calls and edges \(EE\) are dependencies between them\. Motivated by task scheduling studies\(Sevcik,[1989](https://arxiv.org/html/2605.08477#bib.bib29); Calzarossa and Serazzi,[1993](https://arxiv.org/html/2605.08477#bib.bib3)\), we characterize each execution graph using two metrics: Depth \(dd\)::Depth is defined as the critical path length of the DAG, i\.e\., the longest chain of dependent tool calls\. Depth captures the minimum number of sequential reasoning steps that must be executed in order\. High depth increases exposure to cascading errors\. Breadth \(bb\)::We define breadth as the average parallelism of the graph:b=\|V\|/db=\|V\|/d\. Higher breadth indicates the presence of multiple independent sub\-plans that must eventually be merged\. LLM planners typically operate over a linearized, text\-form plan\. This technical constraint can make it difficult to manage highly branched plans\. Consider the query: “Which company has more employees: Google or the parent company of Instagram?” Figure[2](https://arxiv.org/html/2605.08477#S3.F2)illustrates the plan DAG for this task using the KoPL tools \(Appendix[B\.1](https://arxiv.org/html/2605.08477#A2.SS1)\)\. The query involves two parallel branches \(a single entity lookup and an entity lookup followed by a relationship traversal\)\. These branches are then merged at an aggregation node to perform the final comparison\. The critical path isa2→a3→a4a\_\{2\}\\rightarrow a\_\{3\}\\rightarrow a\_\{4\}, yielding a depth ofd=3d=3\. With\|V\|=4\|V\|=4total steps, the parallel breadth isb=4/3b=4/3\. This example indicates that structural complexity arises not only from sequential length but also from the need to coordinate multiple independent sub\-queries\. ### 3\.2\.Tool Robustness \(Vocabulary Alignment\) Even when a plan’s logical structure is correct, execution can fail if tool inputs do not exactly match the representation of data sources\. We thus define tool robustness to measure how tolerant an environment is to imperfect tool specifications\. We consider two common forms of robustness: Robustness to Schema Mismatch::In structured data environments \(e\.g\., KBQA\), tool parameters must match predefined schema elements such as relation names or attribute keys\. For instance, a tool call that usesemployeesmay fail if the schema only containsemployee\_counts\. This mismatch can be mitigated when the tool supports flexible recovery mechanisms such as embedding\-based soft matching \(e\.g\., mappingemployeestoemployee\_countsgiven the query context\)\. Robustness to Retrieval Uncertainty::In unstructured environments \(e\.g\., document retrieval\), failure may arise from search mismatch rather than schema mismatch\. For example, the query “whoownsInstagram” may not retrieve evidence that states “Instagramwas acquiredby Meta”, depending on the retrieval quality\. Robustness can vary depending on recall of the retriever \(top\-1 vs top\-100\), semantic matching quality and evidence aggregation strategy\. Although other robustness factors can also be meaningful, such as tolerance to technical failures \(e\.g\., service downtime\), we focus on the semantic and informational dimensions above because they are particularly common in data\-centric QA tasks\. ## 4\.Experiments We aim to assess whether step\-wise monitoring is necessary for complex data\-centric tool use, or whether comparable performance is achievable with FH with lazy monitoring\. For this purpose, we compare SH and FH planning with lazy under controlled variation in topological complexity and tool robustness\. Specifically, we address three questions: \(1\) Does SH planning outperform FH planning in overall accuracy? \(2\) Does SH planning better handle greater topological complexity \(deep/wide graphs\)? \(3\) Does SH planning exhibit greater robustness in noisy environments? ### 4\.1\.Experimental Setup Figure 3\.Distribution of task instances across datasets\. The horizontal and vertical axes represent the critical path length \(depth\) and average parallelism \(breadth\), respectively\. The bar chart summarizes total instance counts per dataset\.A dataset overview figure combining a scatter plot and a bar chart\. The scatter plot places task instances by execution\-graph depth on the horizontal axis and average parallelism breadth on the vertical axis, showing how instances from different datasets are distributed across topological complexity\. The accompanying bar chart summarizes the total number of instances in each dataset\.Table 1\.Example task instances with their corresponding ground\-truth tool\-call trajectories \(execution graphs\)\.\(a\)KQA Pro \(b\)GrailQA \(c\)Multi\-objective HotpotQA \(k=2k=2\)Q: 1\. Eugeniusz Bodo and Chris Buck both shared what occupation? 2\. The creator of the record label Merciful Release was born in what year?Solution \(Simplified\)A: 1\. film director, 2\. 1959 We evaluate performance across data\-centric tasks over two types of underlying data: KBQA \(Structured\)::We use KQA Pro\(Cao et al\.,[2022](https://arxiv.org/html/2605.08477#bib.bib4)\), GrailQA\(Gu et al\.,[2021](https://arxiv.org/html/2605.08477#bib.bib11)\), WebQSP\(Yih et al\.,[2016](https://arxiv.org/html/2605.08477#bib.bib46)\), and GraphQ\(Su et al\.,[2016](https://arxiv.org/html/2605.08477#bib.bib31)\)as our main testbed\. These tasks require precise logical composition over structured knowledge bases \(Wikidata and Freebase\), where answers are constructed through multi\-step queries\. For KQA Pro, we employ KoPL tools \(27 tools\)\. For the rest, we use the atomic query tools\(Luo et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib24)\)\(7 tools\)\. A tool execution is considered as failed if it returns no valid result\. Multi\-objective HotpotQA \(Unstructured\)::FollowingZhou et al\.\([2025](https://arxiv.org/html/2605.08477#bib.bib48)\), we combine examples from HotpotQA\(Yang et al\.,[2018](https://arxiv.org/html/2605.08477#bib.bib44)\)to synthesize multi\-objective questions with greater breadth \(k∈\{2,3,4,5\}k\\in\\\{2,3,4,5\\\}\), and also include the original HotpotQA instances \(k=1k=1\)\. This setting requires accurate task decomposition and search over Wikipedia\. We use the search and reasoning sub\-agents fromKim et al\.\([2024c](https://arxiv.org/html/2605.08477#bib.bib15)\)as tools\. Each sub\-agent takes a short sub\-question as input and returns a short answer if supporting evidence is retrieved \(search\) or if the reasoning is successfully performed \(reasoning\)\. Note that modern LLMs can answer many HotpotQA questions without tools\. We include this dataset primarily to connect our analysis with prior work\. To enable analysis based on topological features, we derive ground\-truth execution trajectories from dataset provenance\. For KBQA, we use the gold solution program \(KQA Pro\) or logical forms \(GrailQA, WebQSP, and GraphQ\) to obtain the sequence of tool calls with dependency information\. For HotpotQA, we use gold supporting documents to determine the required retrieval and reasoning steps, assisted by LLM annotation \(GPT\-4\.1\-mini\)\. We downsample all datasets to balance reasoning length and question types, yielding 929 KQA Pro instances, 1,115 Atomic KBQA instances, and 1,000 multi\-objective HotpotQA instances\. As discussed in Section[3\.2](https://arxiv.org/html/2605.08477#S3.SS2), we model tool robustness using schema matching in KBQA and the retrieval rank cutoff \(top\-\(k\)\) in multi\-objective HotpotQA\. For the main results \(§[4\.2](https://arxiv.org/html/2605.08477#S4.SS2)\) and topological complexity analysis \(§[4\.3](https://arxiv.org/html/2605.08477#S4.SS3)\), we use high\-robustness settings, where tools are more tolerant to imprecise input parameters\. In KBQA, when a parameter specified by the planner does not exactly match the KB schema, we perform soft schema matching by \(i\) retrieving candidates via cosine similarity over BGE embeddings\(Xiao et al\.,[2024](https://arxiv.org/html/2605.08477#bib.bib38)\)and \(ii\) validating candidates with GPT\-4\.1\-mini\. If no valid match is found, the tool returns the top\-10 nearest candidates as feedback\. In multi\-objective HotpotQA, the search tool retrieves the top\-10 documents to increase the chance of retrieving the necessary evidence\. In controlled robustness experiments \(§[4\.4](https://arxiv.org/html/2605.08477#S4.SS4)\), we switch to low\-robustness settings by disabling soft matching in KBQA or restricting retrieval to top\-1 in HotpotQA\. Figure[3](https://arxiv.org/html/2605.08477#S4.F3)shows the distribution of task instances across datasets, and Table[1](https://arxiv.org/html/2605.08477#S4.T1)provides examples with ground\-truth trajectories\. More details on dataset\-specific pre\-processing and tool implementations can be found in Appendix[B\.1](https://arxiv.org/html/2605.08477#A2.SS1)\. ##### Backbone LLMs: We compare four backbone LLMs representing distinct capability classes: \(1\) GPT\-4\.1\-mini and Qwen3\-235B\-A22B \(Instruct\), representing standard instruction\-tuned models; and \(2\) GPT\-5\-mini and Gemini\-3\-Flash, representing frontier reasoning models with stronger inherent planning capabilities\. ##### Metrics: We repeat the same configuration three times to account for the randomness of LLMs\. We report average accuracy \(using LLM\-as\-judge for soft matching\) and efficiency \(average input/output tokens per query\)\. See Appendix[B\.3](https://arxiv.org/html/2605.08477#A2.SS3)for details\. \[system\]\(FH\) Usetheavailabletoolstoanswertheuser’squestionstepbystep,generatingtheentireplan\.Use$itorefertotheoutputofstepi\(0\-indexed\)\. \*\*Tools\*\* \{tool\_definitions\} Youcanaccesstheknowledgebasethroughtheprovidedtools\. \*\*Examples\*\* \{demonstrations\} \[system\]\(SH\) Usetheavailabletoolstoanswertheuser’squestionstepbystep,generatingoneactionatatime\.Use$itorefertotheoutputofstepi\(0\-indexed\)\. \(\.\.\.sameasFH\) \[user\] Question:\{question\} Prompt 1:Planning prompts for Atomic KBQA ##### Implementation Details: As discussed in Section[2](https://arxiv.org/html/2605.08477#S2), FH and SH differ primarily in planning horizon\. To ensure a fair comparison, we use minimal prompts that include task instructions, tool definitions \(as a JSON string\), and in\-context demonstrations \(see Prompt[1](https://arxiv.org/html/2605.08477#LST1)\), while keeping all other components identical across paradigms\. For each task instance, the agent is allowed to execute up to 30 tool calls, and if tool execution fails, the agent can retry up to eight times\. We constrain the LLM decoding to produce a plan as a sequence of tool calls in JSON format\. We do not use native tool\-calling APIs, which are not applicable to FH\. For in\-context learning, we sample 10 demonstration examples from the training set and provide the same set to both paradigms, retrieving semantically similar examples via query embedding\. Details such as LLM prompts and hyper\-parameters can be found in Appendix[B\.2](https://arxiv.org/html/2605.08477#A2.SS2)\. ### 4\.2\.Overall Accuracy and Efficiency Table 2\.Comparison of FH and SH across structured \(KBQA\) and unstructured \(multi\-objective HotpotQA\) tool\-calling tasks\. FH matches SH in accuracy across most settings while using substantially fewer input tokens than SH\. The efficiency gap widens as the tool set grows \(e\.g\. KQA Pro\)\.Table[2](https://arxiv.org/html/2605.08477#S4.T2)presents aggregate accuracy and token consumption for both paradigms across all datasets and LLM backends\. ##### Accuracy: SH shows no clear performance advantage over FH\. In fact, FH significantly outperforms SH on Atomic KBQA datasets for Gemini\-3\-Flash \(e\.g\., SH is worse by 15\.4 accuracy points on GrailQA and 17\.2 points on GraphQ\)\. We provide a more detailed analysis in Section[4\.5](https://arxiv.org/html/2605.08477#S4.SS5)\. This pattern result suggests that, for this specific reasoning model, FH planning can be more effective than SH planning\. ##### Efficiency: FH consistently reduces input token consumption by 2–3x\. The magnitude of this reduction depends on the length of tool definitions in the input prompt\. For KQA Pro \(27 tools with complex schemas\), the savings are massive \(2\.7–4\.7x\)\. For HotpotQA \(2 tools\), the margin is smaller \(1\.4–1\.9x\) but still significant\. Crucially, for reasoning models \(GPT\-5\-mini and Gemini\-3\-Flash\), FH also saves output tokens\. While SH triggers a full Chain\-of\-Thought generation at every single step, FH generates the reasoning chain only once during initial plan generation and during occasional replanning\. This efficiency advantage is particularly beneficial in production settings, where token usage directly translates to latency and cost\. ##### Does SH outperform FH in overall accuracy? No\. Overall, SH does not deliver a clear accuracy advantage over FH, while FH is more token\-efficient\. We next analyze performance under controlled variation in topological complexity and tool robustness\. ### 4\.3\.Topological Complexity Analysis We employ a logistic regression model to quantify the impact of planning strategy while controlling for task complexity\. Because we have repeated measures \(three runs per query\), we fit the model with Generalized Estimating Equations \(GEE\) using clustering by question ID\. The specification is: logit\(P\(y=1\)\)=\\displaystyle\\text\{logit\}\(P\(y=1\)\)=β0\+βdd∗\+βbb∗\+βSHxSH\\displaystyle\\beta\_\{0\}\+\\beta\_\{d\}d^\{\*\}\+\\beta\_\{b\}b^\{\*\}\+\\beta\_\{\\text\{SH\}\}x\_\{\\text\{SH\}\}\+βd:SH\(d∗×xSH\)\+βb:SH\(b∗×xSH\)\+…,\\displaystyle\+\\beta\_\{d:\\text\{SH\}\}\(d^\{\*\}\\times x\_\{\\text\{SH\}\}\)\+\\beta\_\{b:\\text\{SH\}\}\(b^\{\*\}\\times x\_\{\\text\{SH\}\}\)\+\\dots\_\{,\}wherey∈\{0,1\}y\\in\\\{0,1\\\}indicates task success,β\\betadenotes coefficients to be estimated,d∗d^\{\*\}andb∗b^\{\*\}are standardized depth and breadth features \(§[3\.1](https://arxiv.org/html/2605.08477#S3.SS1)\), andxSH∈\{0,1\}x\_\{\\text\{SH\}\}\\in\\\{0,1\\\}is a binary indicator for the SH strategy\. We additionally include dataset\-specific control variables, for example the final reasoning operation type \(e\.g\.,VerifyDate\) for KBQA, dataset identity for Atomic KBQA, and question type \(bridge/comparison\) for HotpotQA\. For brevity, we omit these controls from the summary tables in this section\. Full model specifications and results are provided in Appendix[B\.3](https://arxiv.org/html/2605.08477#A2.SS3)and Appendix[C\.1](https://arxiv.org/html/2605.08477#A3.SS1)\. Table 3\.Summary of GEE coefficients \(∗∗p<0\.01\*\*p<0\.01,∗p<0\.05\*p<0\.05\)\. Atomic KBQA results include GrailQA, WebQSP, and GraphQ\.##### Result: As shown in Table[3](https://arxiv.org/html/2605.08477#S4.T3), the coefficients for depth \(βd\\beta\_\{d\}\) and breadth \(βb\\beta\_\{b\}\) are consistently negative and statistically significant \(p<0\.01p<0\.01\) across most datasets and models\. This result confirms that topological complexity is a challenging factor that degrades performanceregardless of the planning paradigm\. ##### Does SH handle greater topological complexity \(deep/wide graphs\) better than FH? In general, no\. The interaction coefficients between planning method and topological complexity \(βd:SH\\beta\_\{d:\\text\{SH\}\}andβb:SH\\beta\_\{b:\\text\{SH\}\}\) are statistically insignificant \(p\>0\.05p\>0\.05\) in most settings\. In some cases, SH is evenworseat handling increasing complexity \(e\.g\., GPT\-4\.1\-mini on KQA Pro\)\. Nevertheless, these differences are small compared to the dominant negative impact of topological complexity \(βd\\beta\_\{d\}andβb\\beta\_\{b\}\) overall\. ### 4\.4\.Tool Robustness Analysis Table 4\.Performance change \(absolute accuracy difference\) under low robustness settings\. Difference larger than 0\.05 is denoted inbold\.To evaluate how SH and FH behave under different levels of tool robustness, we inject noise into tool execution\. Specifically, for KBQA we disable automatic schema matching \(forcing exact identifier use\), and for multi\-objective HotpotQA we restrict retrieval to top\-1 \(simulating low recall\)\. Under both perturbations, tools are more likely to return empty results, which in turn necessitate replanning\. ##### Result: Table[4](https://arxiv.org/html/2605.08477#S4.T4)shows that noise leads to comparable performance degradation for FH and SH\. In many cases, the drop is less than 0\.05 accuracy points, and the difference between the two paradigms is small, suggesting that both SH and FH handle execution failures reasonably well\. ##### Does SH exhibit greater robustness in noisy environments? No\. Contrary to the expectation that immediate step\-by\-step feedback improves error recovery, SH shows no consistent advantage and is sometimes more sensitive to noise\. This unexpected brittleness is most evident with Gemini\-3\-Flash\. On GraphQ, turning off soft schema matching reduces SH accuracy by 0\.178 points, in contrast to a 0\.078\-point drop for FH\. On HotpotQA, restricting retrieval to top\-1 reduces SH accuracy by 0\.190 points, compared to a 0\.065\-point drop for FH\. We analyze this behavior in more detail in the next section\. Table 5\.Fraction of task instances in which the agent made at least one repeated tool call \(same tool call with identical arguments\)\. SH exhibits substantially higher repetition rates than FH in several settings\. Rates≥10%\\geq 10\\%are shown in bold\.Table 6\.Gemini\-3\-Flash under low\-robustness retrieval \(top\-1\) on multi\-objective HotpotQA \(k=2k=2\)\. This table focuses on one sub\-question: “Which director, John Schlesinger or Barbara Albert, was also a writer and film producer?” Italic text shows the model’s internal reasoning\. “Error” denotes execution failures\. The feedback text returned to the agent is omitted for brevity\. See Appendix[C\.2](https://arxiv.org/html/2605.08477#A3.SS2)for full trajectories\.\(a\)SH trajectory: After failures, the agent eventually repeated an identical search call \(shown in red\)\.I’m currently engaged in evaluating both John Schlesinger and Barbara Albert, determining if each functioned as a writer and producer\. After gathering the facts, a comparison will readily reveal the answer\. Parallel to this, …¿ search\(Was John Schlesinger a writer and film producer?\)→\\toErrorI’m starting by investigating the first question\. I’ve initiated Step 0: searching for John Schlesinger’s roles\. …¿ search\(What were the occupations of director John Schlesinger?\)→\\toErrorI’m now formulating my search queries\. I’ll begin by querying about the occupations of John Schlesinger and Barbara Albert\. …¿ search\(What were the occupations of John Schlesinger?\)→\\toErrorI’ve begun looking into the roles of John Schlesinger and Barbara Albert\. Initial searches for their occupations are underway\. I’ll need …¿search\(What were the occupations of John Schlesinger?\)\[repeat\] \(b\)FH trajectory: Upon failures, the agent replans the remaining steps\. It shifted focus to an alternative sub\-goal, and intentionally refined the search query \(bold\)\.I’ve initiated searches to determine if John Schlesinger and Barbara Albert were both writers and producers, as planned\. Once I have those details, a comparison will be conducted to pinpoint the correct individual\. I’m also …¿ \[$0 = search\(Was John Schlesinger a writer and film producer?\), $1 = search\(Was Barbara Albert a writer and film producer?\), …\]→\\toError in $1I’ve restarted my search for John Schlesinger’s occupations, focusing on … Also, I’m planning to look into Barbara Albert’s roles\. …¿ \[$1 = search\(What were the occupations of director John Schlesinger?\), $2 = search\(What were the occupations of director Barbara Albert?\), …\]→\\toError in $1I’m now focusing on Barbara Albert’s occupations,and then refining how I search for John Schlesinger’s roles\. …¿ \[$2 = search\(What are the occupations of Barbara Albert \(director\)?\), $3 = search\(Was John Schlesinger a writer and film producer in addition to being a director?\), …\]→\\toError in $3I’m focusing on John Schlesinger’s occupations, trying to confirm if he was also a writer or producer\.I’ll use a more specific search term this time around\.…¿ \[$4 = search\(What were the professional roles of the English film director John Schlesinger?\), …\] ### 4\.5\.Failure Mode Analysis: Repetitive Tool Calls To better understand the unexpected brittleness of SH under some conditions, we identified a key failure mode, particularly in low\-robustness tool settings: SH agents can get stuck in repetitive tool\-call loops more often than FH agents\. When a tool call fails or returns an empty result, models such as Gemini\-3\-Flash often issue the same tool call again in the next step with identical arguments, despite receiving failure feedback\. To quantify this behavior, we measured the fraction of task instances in which an agent made at least one repeated tool call, defined as invoking the same tool with identical arguments multiple times\. Under the default high\-robustness setting, Gemini\-3\-Flash in SH mode repeated tool calls in 30–45% of instances on GrailQA and multi\-objective HotpotQA \(Table[5](https://arxiv.org/html/2605.08477#S4.T5)\), whereas FH reduced this rate to 1\.9% \(GrailQA\) and 5\.9% \(multi\-objective HotpotQA\)\. The gap widened further in low\-robustness settings: on multi\-objective HotpotQA with top\-1 retrieval, SH made repetitive calls in 66\.9% of instances, in contrast to 13\.9% for FH \(see Table[6](https://arxiv.org/html/2605.08477#S4.T6)for a case study\)\. We observed the same pattern, though less pronounced, with other LLM backends as well\. We conjecture that FH with lazy replanning is less prone to such local traps for two reasons\. First, replanning is triggered only upon execution failure rather than at every step, which may reduce the chance of myopic retries\. Second, when replanning is triggered, the model must generate a complete remaining plan toward the end goal, which may guide it to revise the overall strategy instead of repeating the same local action\. ## 5\.Related Work We position our work within three strands of research: \(1\) comparative analyses of planning paradigms, \(2\) adaptive or hybrid planning strategies, and \(3\) agentic approaches for reasoning in data\-centric tasks\. ### 5\.1\.Comparative Studies of Planning Paradigms Recent survey papers provide meta\-analyses of the broad landscape of LLM\-based task planning techniques\(Huang et al\.,[2024](https://arxiv.org/html/2605.08477#bib.bib13); Li,[2025](https://arxiv.org/html/2605.08477#bib.bib20); Wei et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib36)\)\. Closest to our analytic goals is the work ofLiu et al\.\([2025](https://arxiv.org/html/2605.08477#bib.bib22)\), who depart from conceptual analysis to provide an empirical analysis of LLM task decomposition \(≈\\approxplanning\) across multiple domains\. They show that the choice between Decomposition\-First \(analogous to FH\) and Interleaved \(analogous to SH\) strategies is highly task\-dependent, especially when computational efficiency is taken into account\. This finding is particularly relevant given that step\-by\-step execution with eager monitoring remains a common default in low\-level tool use under environmental uncertainty\(Kim et al\.,[2024b](https://arxiv.org/html/2605.08477#bib.bib17); Gonzalez\-Pumariega et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib9); Zhang et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib47)\)\. Our work builds on[Liu et al\.](https://arxiv.org/html/2605.08477#bib.bib22)’s insights but differs in three key aspects\. First, we isolateplanning horizonas the key architectural feature and analyze its impact on both accuracy and cost at the level of tool execution\. Second, while Liu et al\. primarily evaluate pure LLM reasoning, we focus ondata\-centric tool calling, which exhibits unique challenges \(logic and vocabulary alignment\)\. Third, we move beyond dataset\-level comparisons by introducing an instance\-level characterization of task difficulty to better understand when and why SH is \(or is not\) advantageous\. ### 5\.2\.Adaptive Planning Strategies As discussed in Section[2](https://arxiv.org/html/2605.08477#S2), existing planning techniques can be categorized into FH and SH, each with distinct characteristics\. FH planning was common in early LLM planning work, motivated by multi\-step reasoning paradigms such as Chain\-of\-Thought\(Wei et al\.,[2022](https://arxiv.org/html/2605.08477#bib.bib37)\)and explicit plan\-then\-execute pipelines\(Wang et al\.,[2023](https://arxiv.org/html/2605.08477#bib.bib34); Lu et al\.,[2023](https://arxiv.org/html/2605.08477#bib.bib23); Schick et al\.,[2023](https://arxiv.org/html/2605.08477#bib.bib27)\)\. Although SH is now often viewed as the more robust default following strong results in prior work\(Inaba et al\.,[2023](https://arxiv.org/html/2605.08477#bib.bib14); Yao et al\.,[2023](https://arxiv.org/html/2605.08477#bib.bib45)\), FH continues to motivate active research, including techniques that verify or optimize a full plan before execution\(Li et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib19); Lee et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib18)\), as well as approaches that prioritize lower computational cost\(Xu et al\.,[2023](https://arxiv.org/html/2605.08477#bib.bib41)\)\. To combine their strengths, recent work has explored strategies that integrate upfront global planning with iterative local planning\. One prominent direction is adaptive workflows, where agents dynamically consolidate low\-level actions into a single step\(Wang et al\.,[2024](https://arxiv.org/html/2605.08477#bib.bib35)\)or begin with a global plan and switch to step\-wise planning when execution fails or task complexity demands it\(Prasad et al\.,[2024](https://arxiv.org/html/2605.08477#bib.bib26); Choi et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib6)\)\. Another direction is component optimization, where the roles of global planner and local executor are optimized separately using specialized fine\-tuning\(Erdogan et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib8)\)or exemplar retrieval\(Kim et al\.,[2024a](https://arxiv.org/html/2605.08477#bib.bib16)\)\. Practical systems \(e\.g\., LLM\-powered coding agents\) also increasingly support mixed planning behaviors such as upfront task decomposition and batched tool execution\(Anthropic,[2025](https://arxiv.org/html/2605.08477#bib.bib2); Cloudflare,[2025](https://arxiv.org/html/2605.08477#bib.bib7)\)\. Yet, the design of these integrated systems remains largely heuristic, with limited systematic understanding of exactly which cues should trigger strategy switching\. Our work does not introduce a new systems paradigm\. Rather, we provide a controlled empirical comparison of planning horizon in data\-centric tool calling to clarify when FH\-style behavior can be a cost\-effective option\. This insight motivates future work on principled routing and strategy selection grounded in measurable task properties rather than trial\-and\-error adaptation\. ### 5\.3\.Agentic Approaches to Data\-centric Tasks Traditionally, data\-centric tasks like KBQA and Text\-to\-SQL were solved by specialized semantic parsers\. Recently, significant attention has shifted toward agentic data analytics\(Testini et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib32); Chen et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib5)\)\. In these frameworks, LLMs interact with environments to access and manipulate data for complex tasks\. Frameworks like Middleware\(Gu et al\.,[2024](https://arxiv.org/html/2605.08477#bib.bib12)\)and Interactive\-T2S\(Xiong et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib40)\)enable LLM agents to gain schema awareness through interaction, addressing the hallucination problems inherent in static query generation\. Tight step\-wise feedback can be essential when tasks require exploratory schema discovery or when the environment is highly uncertain\. However, for data\-centric tasks with relatively stable tool and data semantics, our analysis suggests that SH is not always the most cost\-effective default\. In production environments, latency and token budgets are often primary constraints\. Our findings therefore offer a practical perspective on designing efficient data\-centric agents\. ## 6\.Conclusion We evaluate whether LLM agents solving data\-centric tool\-calling tasks require step\-wise planning with eager monitoring in the settings we study\. We isolate*planning horizon*as the key architectural choice, comparing single\-step horizon \(SH\) planning with full\-horizon \(FH\) planning under*lazy*monitoring\. To make this comparison controlled and fine\-grained, we characterize task difficulty along*topological complexity*\(execution\-graph depth and breadth\) and*tool robustness*\(tolerance to imperfect inputs\), and we evaluate across KBQA and multi\-hop QA with multiple LLMs\. Despite the common assumption favoring SH, we find no consistent accuracy advantage for SH over FH, nor evidence that SH is more robust to increased topological complexity or noisy tool behavior\. Overall, these results suggest that, in the*well\-defined*data\-centric tool\-use settings studied here, FH with lazy replanning can match SH while using substantially fewer tokens, while SH may remain preferable in more exploratory or highly uncertain environments\. These findings open several directions for future work\. An important next step is to evaluate and extend these methods in more complicated and real\-world scenarios, including domain\-specific or proprietary data\. Our results also motivate developing hybrid methods that build on the complementary strengths of FH and SH\. ###### Acknowledgements\. We thank the anonymous reviewers and our colleagues at Megagon Labs for their valuable feedback and discussions\. ## References - \(1\) - Anthropic \(2025\)Anthropic\. 2025\.Programmatic tool calling \- Claude API Docs\.[https://platform\.claude\.com/docs/en/agents\-and\-tools/tool\-use/programmatic\-tool\-calling](https://platform.claude.com/docs/en/agents-and-tools/tool-use/programmatic-tool-calling) - Calzarossa and Serazzi \(1993\)Maria Calzarossa and Giuseppe Serazzi\. 1993\.Workload characterization: a survey\.*Proc\. IEEE*81, 8 \(1993\), 1136–1150\.[doi:10\.1109/5\.236191](https://doi.org/10.1109/5.236191) - Cao et al\.\(2022\)Shulin Cao, Jiaxin Shi, Liangming Pan, Lunyiu Nie, Yutong Xiang, Lei Hou, Juanzi Li, Bin He, and Hanwang Zhang\. 2022\.KQA Pro: A Dataset with Explicit Compositional Programs for Complex Question Answering over Knowledge Base\. In*Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*\. Association for Computational Linguistics, Dublin, Ireland, 6101–6119\.[doi:10\.18653/v1/2022\.acl\-long\.422](https://doi.org/10.18653/v1/2022.acl-long.422) - Chen et al\.\(2025\)Ke Chen, Peiran Wang, Yaoning Yu, Xianyang Zhan, and Haohan Wang\. 2025\.Large Language Model\-based Data Science Agent: A Survey\.arXiv:2508\.02744 \[cs\.AI\][https://arxiv\.org/abs/2508\.02744](https://arxiv.org/abs/2508.02744) - Choi et al\.\(2025\)Jae\-Woo Choi, Hyungmin Kim, Hyobin Ong, Minsu Jang, Dohyung Kim, Jaehong Kim, and Youngwoo Yoon\. 2025\.ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long\-Horizon Task Planning\.[doi:10\.48550/arXiv\.2511\.02424](https://doi.org/10.48550/arXiv.2511.02424) - Cloudflare \(2025\)Cloudflare\. 2025\.Codemode · Cloudflare Agents docs\.[https://developers\.cloudflare\.com/agents/api\-reference/codemode/](https://developers.cloudflare.com/agents/api-reference/codemode/) - Erdogan et al\.\(2025\)Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami\. 2025\.Plan\-and\-Act: Improving Planning of Agents for Long\-Horizon Tasks\. In*Proceedings of the 42nd International Conference on Machine Learning**\(Proceedings of Machine Learning Research, Vol\. 267\)*\. PMLR, Vancouver, Canada, 15419–15462\.[https://proceedings\.mlr\.press/v267/erdogan25a\.html](https://proceedings.mlr.press/v267/erdogan25a.html) - Gonzalez\-Pumariega et al\.\(2025\)Gonzalo Gonzalez\-Pumariega, Leong Su Yean, Neha Sunkara, and Sanjiban Choudhury\. 2025\.Robotouille: An Asynchronous Planning Benchmark for LLM Agents\. In*The Thirteenth International Conference on Learning Representations*\. Singapore\.[https://openreview\.net/forum?id=OhUoTMxFIH](https://openreview.net/forum?id=OhUoTMxFIH) - Google \(2025\)Google\. 2025\.Gemini 3 Flash Model Card\.[https://storage\.googleapis\.com/deepmind\-media/Model\-Cards/Gemini\-3\-Flash\-Model\-Card\.pdf](https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Flash-Model-Card.pdf) - Gu et al\.\(2021\)Yu Gu, Sue Kase, Michelle Vanni, Brian Sadler, Percy Liang, Xifeng Yan, and Yu Su\. 2021\.Beyond I\.I\.D\.: Three Levels of Generalization for Question Answering on Knowledge Bases\. In*Proceedings of the Web Conference 2021**\(WWW ’21\)*\. Association for Computing Machinery, New York, NY, USA, 3477–3488\.[doi:10\.1145/3442381\.3449992](https://doi.org/10.1145/3442381.3449992) - Gu et al\.\(2024\)Yu Gu, Yiheng Shu, Hao Yu, Xiao Liu, Yuxiao Dong, Jie Tang, Jayanth Srinivasa, Hugo Latapie, and Yu Su\. 2024\.Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments\. In*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*\. Association for Computational Linguistics, Miami, Florida, USA, 7646–7663\.[doi:10\.18653/v1/2024\.emnlp\-main\.436](https://doi.org/10.18653/v1/2024.emnlp-main.436) - Huang et al\.\(2024\)Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, Defu Lian, Yasheng Wang, Ruiming Tang, and Enhong Chen\. 2024\.Understanding the planning of LLM agents: A survey\.arXiv:2402\.02716 \[cs\.AI\][https://arxiv\.org/abs/2402\.02716](https://arxiv.org/abs/2402.02716) - Inaba et al\.\(2023\)Tatsuro Inaba, Hirokazu Kiyomaru, Fei Cheng, and Sadao Kurohashi\. 2023\.MultiTool\-CoT: GPT\-3 Can Use Multiple External Tools with Chain of Thought Prompting\. In*Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*\. Association for Computational Linguistics, Toronto, Canada, 1522–1532\.[doi:10\.18653/v1/2023\.acl\-short\.130](https://doi.org/10.18653/v1/2023.acl-short.130) - Kim et al\.\(2024c\)Joongwon Kim, Bhargavi Paranjape, Tushar Khot, and Hannaneh Hajishirzi\. 2024c\.Husky: A Unified, Open\-Source Language Agent for Multi\-Step Reasoning\.[doi:10\.48550/arXiv\.2406\.06469](https://doi.org/10.48550/arXiv.2406.06469) - Kim et al\.\(2024a\)Minsoo Kim, Victor Bursztyn, Eunyee Koh, Shunan Guo, and Seung\-won Hwang\. 2024a\.RaDA: Retrieval\-augmented Web Agent Planning with LLMs\. In*Findings of the Association for Computational Linguistics: ACL 2024*\. Association for Computational Linguistics, Bangkok, Thailand, 13511–13525\.[doi:10\.18653/v1/2024\.findings\-acl\.802](https://doi.org/10.18653/v1/2024.findings-acl.802) - Kim et al\.\(2024b\)Sehoon Kim, Suhong Moon, Ryan Tabrizi, Nicholas Lee, Michael W\. Mahoney, Kurt Keutzer, and Amir Gholami\. 2024b\.An LLM Compiler for Parallel Function Calling\. In*Proceedings of the 41st International Conference on Machine Learning**\(Proceedings of Machine Learning Research, Vol\. 235\)*\. PMLR, Vienna, Austria, 24370–24391\.[https://proceedings\.mlr\.press/v235/kim24y\.html](https://proceedings.mlr.press/v235/kim24y.html) - Lee et al\.\(2025\)Christine P\. Lee, David Porfirio, Xinyu Jessica Wang, Kevin Chenkai Zhao, and Bilge Mutlu\. 2025\.VeriPlan: Integrating Formal Verification and LLMs into End\-User Planning\. In*Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems**\(CHI ’25\)*\. Association for Computing Machinery, New York, NY, USA, Article 247, 19 pages\.[doi:10\.1145/3706598\.3714113](https://doi.org/10.1145/3706598.3714113) - Li et al\.\(2025\)Ao Li, Yuexiang Xie, Songze Li, Fugee Tsung, Bolin Ding, and Yaliang Li\. 2025\.Agent\-Oriented Planning in Multi\-Agent Systems\. In*The Thirteenth International Conference on Learning Representations*\. Singapore\.[https://openreview\.net/forum?id=EqcLAU6gyU](https://openreview.net/forum?id=EqcLAU6gyU) - Li \(2025\)Xinzhe Li\. 2025\.A Review of Prominent Paradigms for LLM\-Based Agents: Tool Use, Planning \(Including RAG\), and Feedback Learning\. In*Proceedings of the 31st International Conference on Computational Linguistics*\. Association for Computational Linguistics, Abu Dhabi, UAE, 9760–9779\.[https://aclanthology\.org/2025\.coling\-main\.652/](https://aclanthology.org/2025.coling-main.652/) - Lin et al\.\(2021\)Jimmy Lin, Xueguang Ma, Sheng\-Chieh Lin, Jheng\-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira\. 2021\.Pyserini: A Python Toolkit for Reproducible Information Retrieval Research with Sparse and Dense Representations\. In*Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval*\(Virtual Event, Canada\)*\(SIGIR ’21\)*\. Association for Computing Machinery, New York, NY, USA, 2356–2362\.[doi:10\.1145/3404835\.3463238](https://doi.org/10.1145/3404835.3463238) - Liu et al\.\(2025\)Shuodi Liu, Yingzhuo Liu, Zi Wang, Yusheng Wang, Huijia Wu, Liuyu Xiang, and Zhaofeng He\. 2025\.Select\-Then\-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models\. In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*\. Association for Computational Linguistics, Suzhou, China, 5454–5477\.[doi:10\.18653/v1/2025\.emnlp\-main\.278](https://doi.org/10.18653/v1/2025.emnlp-main.278) - Lu et al\.\(2023\)Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai\-Wei Chang, Ying Nian Wu, Song\-Chun Zhu, and Jianfeng Gao\. 2023\.Chameleon: Plug\-and\-Play Compositional Reasoning with Large Language Models\.*Advances in Neural Information Processing Systems*36 \(Dec\. 2023\), 43447–43478\.[https://proceedings\.neurips\.cc/paper\_files/paper/2023/hash/871ed095b734818cfba48db6aeb25a62\-Abstract\-Conference\.html](https://proceedings.neurips.cc/paper_files/paper/2023/hash/871ed095b734818cfba48db6aeb25a62-Abstract-Conference.html) - Luo et al\.\(2025\)Haoran Luo, Haihong E, Yikai Guo, Qika Lin, Xiaobao Wu, Xinyu Mu, Wenhao Liu, Meina Song, Yifan Zhu, and Anh Tuan Luu\. 2025\.KBQA\-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search\. In*Proceedings of the 42nd International Conference on Machine Learning*\. PMLR, Vancouver, Canada, 41177–41199\.[https://proceedings\.mlr\.press/v267/luo25d\.html](https://proceedings.mlr.press/v267/luo25d.html) - OpenAI \(2025\)OpenAI\. 2025\.Update to GPT\-5 System Card: GPT\-5\.2\.[https://cdn\.openai\.com/pdf/3a4153c8\-c748\-4b71\-8e31\-aecbde944f8d/oai\_5\_2\_system\-card\.pdf](https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944f8d/oai_5_2_system-card.pdf) - Prasad et al\.\(2024\)Archiki Prasad, Alexander Koller, Mareike Hartmann, Peter Clark, Ashish Sabharwal, Mohit Bansal, and Tushar Khot\. 2024\.ADaPT: As\-Needed Decomposition and Planning with Language Models\. In*Findings of the Association for Computational Linguistics: NAACL 2024*\. Association for Computational Linguistics, Mexico City, Mexico, 4226–4252\.[doi:10\.18653/v1/2024\.findings\-naacl\.264](https://doi.org/10.18653/v1/2024.findings-naacl.264) - Schick et al\.\(2023\)Timo Schick, Jane Dwivedi\-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom\. 2023\.Toolformer: Language Models Can Teach Themselves to Use Tools\.*Advances in Neural Information Processing Systems*36 \(Dec\. 2023\), 68539–68551\.[https://papers\.nips\.cc/paper\_files/paper/2023/hash/d842425e4bf79ba039352da0f658a906\-Abstract\-Conference\.html](https://papers.nips.cc/paper_files/paper/2023/hash/d842425e4bf79ba039352da0f658a906-Abstract-Conference.html) - Seabold and Perktold \(2010\)Skipper Seabold and Josef Perktold\. 2010\.statsmodels: Econometric and statistical modeling with python\. In*9th Python in Science Conference*\. Austin, Texas, USA\. - Sevcik \(1989\)Kenneth C\. Sevcik\. 1989\.Characterizations of parallelism in applications and their use in scheduling\. In*Proceedings of the 1989 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems*\(Oakland, California, USA\)*\(SIGMETRICS ’89\)*\. Association for Computing Machinery, New York, NY, USA, 171–180\.[doi:10\.1145/75108\.75391](https://doi.org/10.1145/75108.75391) - Singh et al\.\(2025\)Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El\-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, Akshay Nathan, Alan Luo, Alec Helyar, Aleksander Madry, Aleksandr Efremov, Aleksandra Spyra, Alex Baker\-Whitcomb, Alex Beutel, Alex Karpenko, Alex Makelov, Alex Neitz, Alex Wei, Alexandra Barr, Alexandre Kirchmeyer, Alexey Ivanov, Alexi Christakis, Alistair Gillespie, Allison Tam, Ally Bennett, Alvin Wan, Alyssa Huang, Amy McDonald Sandjideh, Amy Yang, Ananya Kumar, Andre Saraiva, Andrea Vallone, Andrei Gheorghe, Andres Garcia Garcia, Andrew Braunstein, Andrew Liu, Andrew Schmidt, Andrey Mereskin, Andrey Mishchenko, Andy Applebaum, Andy Rogerson, Ann Rajan, Annie Wei, Anoop Kotha, Anubha Srivastava, Anushree Agrawal, Arun Vijayvergiya, Ashley Tyra, Ashvin Nair, Avi Nayak, Ben Eggers, Bessie Ji, Beth Hoover, Bill Chen, Blair Chen, Boaz Barak, Borys Minaiev, Botao Hao, Bowen Baker, Brad Lightcap, Brandon McKinzie, Brandon Wang, Brendan Quinn, Brian Fioca, Brian Hsu, Brian Yang, Brian Yu, Brian Zhang, Brittany Brenner, Callie Riggins Zetino, Cameron Raymond, Camillo Lugaresi, Carolina Paz, Cary Hudson, Cedric Whitney, Chak Li, Charles Chen, Charlotte Cole, Chelsea Voss, Chen Ding, Chen Shen, Chengdu Huang, Chris Colby, Chris Hallacy, Chris Koch, Chris Lu, Christina Kaplan, Christina Kim, CJ Minott\-Henriques, Cliff Frey, Cody Yu, Coley Czarnecki, Colin Reid, Colin Wei, Cory Decareaux, Cristina Scheau, Cyril Zhang, Cyrus Forbes, Da Tang, Dakota Goldberg, Dan Roberts, Dana Palmie, Daniel Kappler, Daniel Levine, Daniel Wright, Dave Leo, David Lin, David Robinson, Declan Grabb, Derek Chen, Derek Lim, Derek Salama, Dibya Bhattacharjee, Dimitris Tsipras, Dinghua Li, Dingli Yu, DJ Strouse, Drew Williams, Dylan Hunn, Ed Bayes, Edwin Arbus, Ekin Akyurek, Elaine Ya Le, Elana Widmann, Eli Yani, Elizabeth Proehl, Enis Sert, Enoch Cheung, Eri Schwartz, Eric Han, Eric Jiang, Eric Mitchell, Eric Sigler, Eric Wallace, Erik Ritter, Erin Kavanaugh, Evan Mays, Evgenii Nikishin, Fangyuan Li, Felipe Petroski Such, Filipe de Avila Belbute Peres, Filippo Raso, Florent Bekerman, Foivos Tsimpourlas, Fotis Chantzis, Francis Song, Francis Zhang, Gaby Raila, Garrett McGrath, Gary Briggs, Gary Yang, Giambattista Parascandolo, Gildas Chabot, Grace Kim, Grace Zhao, Gregory Valiant, Guillaume Leclerc, Hadi Salman, Hanson Wang, Hao Sheng, Haoming Jiang, Haoyu Wang, Haozhun Jin, Harshit Sikchi, Heather Schmidt, Henry Aspegren, Honglin Chen, Huida Qiu, Hunter Lightman, Ian Covert, Ian Kivlichan, Ian Silber, Ian Sohl, Ibrahim Hammoud, Ignasi Clavera, Ikai Lan, Ilge Akkaya, Ilya Kostrikov, Irina Kofman, Isak Etinger, Ishaan Singal, Jackie Hehir, Jacob Huh, Jacqueline Pan, Jake Wilczynski, Jakub Pachocki, James Lee, James Quinn, Jamie Kiros, Janvi Kalra, Jasmyn Samaroo, Jason Wang, Jason Wolfe, Jay Chen, Jay Wang, Jean Harb, Jeffrey Han, Jeffrey Wang, Jennifer Zhao, Jeremy Chen, Jerene Yang, Jerry Tworek, Jesse Chand, Jessica Landon, Jessica Liang, Ji Lin, Jiancheng Liu, Jianfeng Wang, Jie Tang, Jihan Yin, Joanne Jang, Joel Morris, Joey Flynn, Johannes Ferstad, Johannes Heidecke, John Fishbein, John Hallman, Jonah Grant, Jonathan Chien, Jonathan Gordon, Jongsoo Park, Jordan Liss, Jos Kraaijeveld, Joseph Guay, Joseph Mo, Josh Lawson, Josh McGrath, Joshua Vendrow, Joy Jiao, Julian Lee, Julie Steele, Julie Wang, Junhua Mao, Kai Chen, Kai Hayashi, Kai Xiao, Kamyar Salahi, Kan Wu, Karan Sekhri, Karan Sharma, Karan Singhal, Karen Li, Kenny Nguyen, Keren Gu\-Lemberg, Kevin King, Kevin Liu, Kevin Stone, Kevin Yu, Kristen Ying, Kristian Georgiev, Kristie Lim, Kushal Tirumala, Kyle Miller, Lama Ahmad, Larry Lv, Laura Clare, Laurance Fauconnet, Lauren Itow, Lauren Yang, Laurentia Romaniuk, Leah Anise, Lee Byron, Leher Pathak, Leon Maksin, Leyan Lo, Leyton Ho, Li Jing, Liang Wu, Liang Xiong, Lien Mamitsuka, Lin Yang, Lindsay McCallum, Lindsey Held, Liz Bourgeois, Logan Engstrom, Lorenz Kuhn, Louis Feuvrier, Lu Zhang, Lucas Switzer, Lukas Kondraciuk, Lukasz Kaiser, Manas Joglekar, Mandeep Singh, Mandip Shah, Manuka Stratta, Marcus Williams, Mark Chen, Mark Sun, Marselus Cayton, Martin Li, Marvin Zhang, Marwan Aljubeh, Matt Nichols, Matthew Haines, Max Schwarzer, Mayank Gupta, Meghan Shah, Melody Huang, Meng Dong, Mengqing Wang, Mia Glaese, Micah Carroll, Michael Lampe, Michael Malek, Michael Sharman, Michael Zhang, Michele Wang, Michelle Pokrass, Mihai Florian, Mikhail Pavlov, Miles Wang, Ming Chen, Mingxuan Wang, Minnia Feng, Mo Bavarian, Molly Lin, Moose Abdool, Mostafa Rohaninejad, Nacho Soto, Natalie Staudacher, Natan LaFontaine, Nathan Marwell, Nelson Liu, Nick Preston, Nick Turley, Nicklas Ansman, Nicole Blades, Nikil Pancha, Nikita Mikhaylin, Niko Felix, Nikunj Handa, Nishant Rai, Nitish Keskar, Noam Brown, Ofir Nachum, Oleg Boiko, Oleg Murk, Olivia Watkins, Oona Gleeson, Pamela Mishkin, Patryk Lesiewicz, Paul Baltescu, Pavel Belov, Peter Zhokhov, Philip Pronin, Phillip Guo, Phoebe Thacker, Qi Liu, Qiming Yuan, Qinghua Liu, Rachel Dias, Rachel Puckett, Rahul Arora, Ravi Teja Mullapudi, Raz Gaon, Reah Miyara, Rennie Song, Rishabh Aggarwal, RJ Marsan, Robel Yemiru, Robert Xiong, Rohan Kshirsagar, Rohan Nuttall, Roman Tsiupa, Ronen Eldan, Rose Wang, Roshan James, Roy Ziv, Rui Shu, Ruslan Nigmatullin, Saachi Jain, Saam Talaie, Sam Altman, Sam Arnesen, Sam Toizer, Sam Toyer, Samuel Miserendino, Sandhini Agarwal, Sarah Yoo, Savannah Heon, Scott Ethersmith, Sean Grove, Sean Taylor, Sebastien Bubeck, Sever Banesiu, Shaokyi Amdo, Shengjia Zhao, Sherwin Wu, Shibani Santurkar, Shiyu Zhao, Shraman Ray Chaudhuri, Shreyas Krishnaswamy, Shuaiqi, Xia, Shuyang Cheng, Shyamal Anadkat, Simón Posada Fishman, Simon Tobin, Siyuan Fu, Somay Jain, Song Mei, Sonya Egoian, Spencer Kim, Spug Golden, SQ Mah, Steph Lin, Stephen Imm, Steve Sharpe, Steve Yadlowsky, Sulman Choudhry, Sungwon Eum, Suvansh Sanjeev, Tabarak Khan, Tal Stramer, Tao Wang, Tao Xin, Tarun Gogineni, Taya Christianson, Ted Sanders, Tejal Patwardhan, Thomas Degry, Thomas Shadwell, Tianfu Fu, Tianshi Gao, Timur Garipov, Tina Sriskandarajah, Toki Sherbakov, Tomer Kaftan, Tomo Hiratsuka, Tongzhou Wang, Tony Song, Tony Zhao, Troy Peterson, Val Kharitonov, Victoria Chernova, Vineet Kosaraju, Vishal Kuo, Vitchyr Pong, Vivek Verma, Vlad Petrov, Wanning Jiang, Weixing Zhang, Wenda Zhou, Wenlei Xie, Wenting Zhan, Wes McCabe, Will DePue, Will Ellsworth, Wulfie Bain, Wyatt Thompson, Xiangning Chen, Xiangyu Qi, Xin Xiang, Xinwei Shi, Yann Dubois, Yaodong Yu, Yara Khakbaz, Yifan Wu, Yilei Qian, Yin Tat Lee, Yinbo Chen, Yizhen Zhang, Yizhong Xiong, Yonglong Tian, Young Cha, Yu Bai, Yu Yang, Yuan Yuan, Yuanzhi Li, Yufeng Zhang, Yuguang Yang, Yujia Jin, Yun Jiang, Yunyun Wang, Yushi Wang, Yutian Liu, Zach Stubenvoll, Zehao Dou, Zheng Wu, and Zhigang Wang\. 2025\.OpenAI GPT\-5 System Card\.arXiv:2601\.03267 \[cs\.CL\][https://arxiv\.org/abs/2601\.03267](https://arxiv.org/abs/2601.03267) - Su et al\.\(2016\)Yu Su, Huan Sun, Brian Sadler, Mudhakar Srivatsa, Izzeddin Gür, Zenghui Yan, and Xifeng Yan\. 2016\.On Generating Characteristic\-rich Question Sets for QA Evaluation\. In*Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing*\. Association for Computational Linguistics, Austin, Texas, 562–572\.[doi:10\.18653/v1/D16\-1054](https://doi.org/10.18653/v1/D16-1054) - Testini et al\.\(2025\)Irene Testini, Lorenzo Pacchiardi, and Jose Hernandez\-Orallo\. 2025\.Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents\.*Transactions on Machine Learning Research*\(2025\)\.[https://openreview\.net/forum?id=MB0TCLfLn1](https://openreview.net/forum?id=MB0TCLfLn1) - Vrandečić and Krötzsch \(2014\)Denny Vrandečić and Markus Krötzsch\. 2014\.Wikidata: A Free Collaborative Knowledgebase\.*Commun\. ACM*57, 10 \(Sept\. 2014\), 78–85\.[doi:10\.1145/2629489](https://doi.org/10.1145/2629489) - Wang et al\.\(2023\)Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka\-Wei Lee, and Ee\-Peng Lim\. 2023\.Plan\-and\-Solve Prompting: Improving Zero\-Shot Chain\-of\-Thought Reasoning by Large Language Models\. In*Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*\. Association for Computational Linguistics, Toronto, Canada, 2609–2634\.[doi:10\.18653/v1/2023\.acl\-long\.147](https://doi.org/10.18653/v1/2023.acl-long.147) - Wang et al\.\(2024\)Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji\. 2024\.Executable Code Actions Elicit Better LLM Agents\. In*Proceedings of the 41st International Conference on Machine Learning*\. PMLR, Vienna, Austria, 50208–50232\.[https://proceedings\.mlr\.press/v235/wang24h\.html](https://proceedings.mlr.press/v235/wang24h.html) - Wei et al\.\(2025\)Hui Wei, Zihao Zhang, Shenghua He, Tian Xia, Shijia Pan, and Fei Liu\. 2025\.PlanGenLLMs: A Modern Survey of LLM Planning Capabilities\. In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*\. Association for Computational Linguistics, Vienna, Austria, 19497–19521\.[doi:10\.18653/v1/2025\.acl\-long\.958](https://doi.org/10.18653/v1/2025.acl-long.958) - Wei et al\.\(2022\)Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou\. 2022\.Chain\-of\-Thought Prompting Elicits Reasoning in Large Language Models\. In*Advances in Neural Information Processing Systems*, Vol\. 35\. Curran Associates, Inc\., New Orleans, Louisiana, USA, 24824–24837\.[https://proceedings\.neurips\.cc/paper\_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf) - Xiao et al\.\(2024\)Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian\-Yun Nie\. 2024\.C\-Pack: Packed Resources For General Chinese Embeddings\. In*Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval*\(Washington DC, USA\)*\(SIGIR ’24\)*\. Association for Computing Machinery, New York, NY, USA, 641–649\.[doi:10\.1145/3626772\.3657878](https://doi.org/10.1145/3626772.3657878) - Xin et al\.\(2025\)Amy Xin, Jinxin Liu, Zijun Yao, Zhicheng Lee, Shulin Cao, Lei Hou, and Juanzi Li\. 2025\.AtomR: Atomic Operator\-Empowered Large Language Models for Heterogeneous Knowledge Reasoning\. In*Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2*\(Toronto ON, Canada\)*\(KDD ’25\)*\. Association for Computing Machinery, New York, NY, USA, 3344–3355\.[doi:10\.1145/3711896\.3736849](https://doi.org/10.1145/3711896.3736849) - Xiong et al\.\(2025\)Guanming Xiong, Junwei Bao, Hongfei Jiang, Yang Song, and Wen Zhao\. 2025\.Multi\-Turn Interactions for Text\-to\-SQL with Large Language Models\. In*Proceedings of the 34th ACM International Conference on Information and Knowledge Management*\(Seoul, Republic of Korea\)*\(CIKM ’25\)*\. Association for Computing Machinery, New York, NY, USA, 3560–3570\.[doi:10\.1145/3746252\.3761052](https://doi.org/10.1145/3746252.3761052) - Xu et al\.\(2023\)Binfeng Xu, Zhiyuan Peng, Bowen Lei, Subhabrata Mukherjee, Yuchen Liu, and Dongkuan Xu\. 2023\.ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models\.arXiv:2305\.18323 \[cs\.CL\][https://arxiv\.org/abs/2305\.18323](https://arxiv.org/abs/2305.18323) - Yadan \(2019\)Omry Yadan\. 2019\.Hydra \- A framework for elegantly configuring complex applications\.Github\.[https://github\.com/facebookresearch/hydra](https://github.com/facebookresearch/hydra) - Yang et al\.\(2025\)An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu\. 2025\.Qwen3 Technical Report\.arXiv:2505\.09388 \[cs\.CL\][https://arxiv\.org/abs/2505\.09388](https://arxiv.org/abs/2505.09388) - Yang et al\.\(2018\)Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D\. Manning\. 2018\.HotpotQA: A Dataset for Diverse, Explainable Multi\-hop Question Answering\. In*Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing*\. Association for Computational Linguistics, Brussels, Belgium, 2369–2380\.[doi:10\.18653/v1/D18\-1259](https://doi.org/10.18653/v1/D18-1259) - Yao et al\.\(2023\)Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao\. 2023\.ReAct: Synergizing Reasoning and Acting in Language Models\. In*The Eleventh International Conference on Learning Representations*\. Kigali, Rwanda\.[https://openreview\.net/forum?id=WE\_vluYUL\-X](https://openreview.net/forum?id=WE_vluYUL-X) - Yih et al\.\(2016\)Wen\-tau Yih, Matthew Richardson, Chris Meek, Ming\-Wei Chang, and Jina Suh\. 2016\.The Value of Semantic Parse Labeling for Knowledge Base Question Answering\. In*Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*\. Association for Computational Linguistics, Berlin, Germany, 201–206\.[doi:10\.18653/v1/P16\-2033](https://doi.org/10.18653/v1/P16-2033) - Zhang et al\.\(2025\)Zhenyu Zhang, Tianyi Chen, Weiran Xu, Alex Pentland, and Jiaxin Pei\. 2025\.ReCAP: Recursive Context\-Aware Reasoning and Planning for Large Language Model Agents\. In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*\. San Diego, California, USA\.[https://openreview\.net/forum?id=r2ykUnzuGt](https://openreview.net/forum?id=r2ykUnzuGt) - Zhou et al\.\(2025\)Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, and Paul Pu Liang\. 2025\.MEM1: Learning to Synergize Memory and Reasoning for Efficient Long\-Horizon Agents\.arXiv:2506\.15841 \[cs\.CL\][https://arxiv\.org/abs/2506\.15841](https://arxiv.org/abs/2506.15841) ## Appendix AArtifact Appendix Our GitHub repository \([https://github\.com/megagonlabs/cais26\-planning\-horizon](https://github.com/megagonlabs/cais26-planning-horizon)\), includes source code, configuration files, scripts, and analysis notebooks\. Datasets and raw experiment outputs are not bundled\. In this section, we provide a brief introduction to our codebase\. ### A\.1\.Key Results and Reproduction Summary The main claim of the paper is that,for well\-defined data\-centric tasks, a full\-horizon \(FH\) planner with lazy monitoring can match a single\-step horizon \(SH\) planner while avoiding much of the overhead from constant feedback integration\.The table below summarizes how the main findings are supported by our software\. 1. \(1\)FH matches SH accuracy across datasets and backbone models while FH uses fewer tokens than SH\.\(Table[2](https://arxiv.org/html/2605.08477#S4.T2)in Section[4\.2](https://arxiv.org/html/2605.08477#S4.SS2)\) - •Runscripts/batch/batch\_exp\_\*\_hydra\.shwrappers, thenscripts/batch/batch\_postprocess\_main\_results\.shfor comparing accuracy and token usage between SH and FH across datasets and models\. - •Details:docs/walkthrough\.md,README\.md 2. \(2\)Topological complexity hurts SH and FHin similar ways\.\(Table[3](https://arxiv.org/html/2605.08477#S4.T3)in Section[4\.3](https://arxiv.org/html/2605.08477#S4.SS3)\) - •Runnotebooks/sec4p3\_topological\-complexity\-analysis\.ipynbto analyze the results from \(1\)\. - •Details:docs/walkthrough\.md 3. \(3\)SH showsno consistent robustness advantagein noisy tool settings\.\(Table[4](https://arxiv.org/html/2605.08477#S4.T4)in Section[4\.4](https://arxiv.org/html/2605.08477#S4.SS4)\) - •Runnotebooks/sec4p4\_tool\-robustness\.ipynbto analyze the results from \(1\)\. - •Details:docs/walkthrough\.md 4. \(4\)SH repeats tool calls more often after failures\.\(Section[4\.5](https://arxiv.org/html/2605.08477#S4.SS5)\) - •Runnotebooks/sec4p5\_repetitive\-tool\-calls\.ipynbto analyze the results from \(1\)\. - •Details:docs/walkthrough\.md See Appendix[A\.4](https://arxiv.org/html/2605.08477#A1.SS4)for more details\. ### A\.2\.Requirements Requirements vary by experiment type\. SeeREADME\.mdfor details\. - •Software:Python 3\.11\+ anduvare required\. Most dependencies can be installed withuv\. API keys for the selected LLM providers \(e\.g\., OpenAI\) must be set in the repository\-root\.envfile\. Atomic KBQA also needs Virtuoso\. - •Hardware:The experiment pipeline needs a GPU\. Hardware needs depend on the experiment track\. RAM and disk must fit the selected knowledge base or Wikipedia\-based corpus\. Multi\-objective HotpotQA uses large Pyserini indexes and may need more RAM during retrieval\. Freebase\-backed Atomic KBQA runs recommend 100 GB\+ RAM\. - •Network:Internet access is needed for LLM APIs and for first\-time index downloads\. ### A\.3\.Setup Guide Readdocs/setup/code\.mdanddocs/setup/data\.mdfirst\.Then use the dataset READMEs underdata/for track\-specific setup\. If you want the lightest path first, start with KQA Pro, then move to the larger tracks\. ### A\.4\.Execution Entry Points This section briefly describes the entry points to reproduce the key experimental results\.docs/walkthrough\.mdis the main guide\. #### A\.4\.1\.Main Experiment Workflow \(§[4\.2](https://arxiv.org/html/2605.08477#S4.SS2)\) For Tables[2](https://arxiv.org/html/2605.08477#S4.T2)and[4](https://arxiv.org/html/2605.08477#S4.T4), runscripts/batch/batch\_exp\_\*\_hydra\.shwrappers for the main experiments\. The results can be evaluated and aggregated bybatch\_postprocess\_main\_results\.sh\. This workflow supports the findings on accuracy parity and token reduction\. The output files are used in the subsequent post\-hoc analyses\. #### A\.4\.2\.Topological Complexity Analysis \(§[4\.3](https://arxiv.org/html/2605.08477#S4.SS3)\) notebooks/sec4p3\_topological\-complexity\-analysis\.ipynbsupports the finding that greater execution\-graph depth and breadth hurt both planners in similar ways, without a clear SH\-specific advantage\. It produces the logistic\-regression coefficients used for Table[3](https://arxiv.org/html/2605.08477#S4.T3)\. #### A\.4\.3\.Tool Robustness Analysis \(§[4\.4](https://arxiv.org/html/2605.08477#S4.SS4)\) notebooks/sec4p4\_tool\-robustness\.ipynbsupports the finding that both SH and FH planners degrade by similar amounts when tool robustness is reduced, with no consistent SH advantage\. It produces the accuracy\-delta tables used in Section[3\.2](https://arxiv.org/html/2605.08477#S3.SS2)\. #### A\.4\.4\.Repetitive Tool Call Analysis \(§[4\.5](https://arxiv.org/html/2605.08477#S4.SS5)\) notebooks/sec4p5\_repetitive\-tool\-calls\.ipynbsupports the finding that SH repeats tool calls much more often than FH after failures\. It produces the repetition\-rate values mentioned in Section[4\.5](https://arxiv.org/html/2605.08477#S4.SS5)\. #### A\.4\.5\.Lightest Runnable Example for Sanity\-Check scripts/example\_run\_kqa\_pro\.pyis the smallest end\-to\-end run in the repository\. It takes a user query as a commandline argument and shows how SH/FH agents answer it using KoPL\. ### A\.5\.Cautions A few practical issues are worth knowing before you run the code\. - •Nondeterminism:LLM API runs are not fully deterministic\. Reproduced numbers should be close to the paper, but exact score\-by\-score matches are not expected\. - •Heavy setup:Atomic KBQA needs a local Virtuoso and Freebase setup\. Multi\-objective HotpotQA downloads large Pyserini indexes on first use\. These datasets need more setup time and more memory\. - •Execution time:For the representative OpenAI settings reported inREADME\.md, one trial is on the order of about 2 to 64 hours of total runtime, depending on track, planner, and model\. Wall\-clock time can be lower with parallel execution\. ### A\.6\.Recommendations The paths below are the easiest ways to approach the artifact\. - •Start small:Start with KQA Pro if you want the lightest runnable path\. Use the smaller\-run options in the batch scripts for sanity\-check\. - •Use parallel execution when possible:If multiple GPUs are available, use the parallel settings\. \(See the walkthrough document\.\) ## Appendix BImplementation Details This section describes the details of tasks \([B\.1](https://arxiv.org/html/2605.08477#A2.SS1)\), planners \([B\.2](https://arxiv.org/html/2605.08477#A2.SS2)\), and evaluation \([B\.3](https://arxiv.org/html/2605.08477#A2.SS3)\)\. ### B\.1\.Task Details Our study uses KQA Pro, Atomic KBQA datasets \(GrailQA, WebQSP, and GraphQ\), and the multi\-objective version of HotpotQA\. They undergo distinct preprocessing and are handled by different tool sets\. We describe these details below\. #### B\.1\.1\.KQA Pro KQA Pro\(Cao et al\.,[2022](https://arxiv.org/html/2605.08477#bib.bib4)\)is a dataset for QA over a sampled subset of Wikidata\(Vrandečić and Krötzsch,[2014](https://arxiv.org/html/2605.08477#bib.bib33)\)\. It features approximately 120k examples paired with ground\-truth reasoning programs in KoPL \(Knowledge\-oriented Programming Language\)\. Table 7\.KoPL tools defined byCao et al\.\([2022](https://arxiv.org/html/2605.08477#bib.bib4)\)\.\[system\]\(FH\) Usetheavailabletoolstoanswertheuser’squestionstepbystep,generatingtheentireplan\.Use$itorefertotheoutputofstepi\(0\-indexed\)\. \*\*Tools\*\* \{tool\_definitions\} Youcanaccesstheknowledgebasethroughtheprovidedtools\. \*\*Examples\*\* \{demonstrations\} \[system\]\(SH\) Usetheavailabletoolstoanswertheuser’squestionstepbystep,generatingoneactionatatime\.Use$itorefertotheoutputofstepi\(0\-indexed\)\. \(\.\.\.sameasFH\) \[user\] Question:\{question\} Prompt 2:Schema grounding for KQA Pro##### Tools: The agent has access to 27 tools based on KoPL functions \(Table[7](https://arxiv.org/html/2605.08477#A2.T7)\)\.222[https://github\.com/THU\-KEG/KoPL](https://github.com/THU-KEG/KoPL)We derived the ground\-truth execution graph of KoPL tools based on the program annotations in the original data\. When parameters \(concept, attribute, or relation\) specified by the planner are not found in the KB schema, we perform soft matching based on embedding similarity and an LLM\. We first retrieve 10 similar values based onbge\-base\-en\-v1\.5embedding similarity, followed by evaluation withgpt\-4\.1\-mini\-2025\-04\-18\(see Prompt[2](https://arxiv.org/html/2605.08477#LST2)for prompts\)\. If no valid candidate is found, the tool returns the top\-10 similar candidates as feedback\. For the controlled experiments regarding tool robustness \(§[4\.4](https://arxiv.org/html/2605.08477#S4.SS4)\), we disable this soft matching and only return the top\-1 candidate to simulate tools with low robustness\. ##### Preprocessing: We sampled approximately 1,000 examples each from the training and development splits of KQA Pro, resulting in new training and test sets with balanced step lengths\. 1. \(1\)Validation:We filtered examples to ensure the provided KoPL program executes without error and matches the ground\-truth answer\. 2. \(2\)Plan DAG Annotation:KoPL programs are represented as trees\. For our analysis, we converted them into DAGs by merging identical steps\. 3. \(3\)Sampling:We balanced the test set across 5 complexity bins based on the number of DAG nodes \(1\-3, 4\-5, 6\-7, 8\-9, 10\+\)\. 4. \(4\)In\-context Demonstration:For each task instance, we retrieved 10 similar examples from held\-out data as in\-context demonstrations usingbge\-base\-en\-v1\.5embeddings of the question text\. #### B\.1\.2\.Atomic KBQA \(GrailQA, WebQSP, GraphQ\) GrailQA\(Gu et al\.,[2021](https://arxiv.org/html/2605.08477#bib.bib11)\), WebQSP\(Yih et al\.,[2016](https://arxiv.org/html/2605.08477#bib.bib46)\), and GraphQ\(Su et al\.,[2016](https://arxiv.org/html/2605.08477#bib.bib31)\)contain QA instances over the full Freebase knowledge base\.333We used the resources available at[https://github\.com/dki\-lab/Freebase\-Setup](https://github.com/dki-lab/Freebase-Setup)for running the Freebase server locally\.Each dataset contains annotations of logical forms called S\-expressions\(Gu et al\.,[2021](https://arxiv.org/html/2605.08477#bib.bib11)\), which can be deterministically converted into executable SPARQL queries\. We refer to these datasets collectively asAtomic KBQAdatasets because we employ Atomic Query Tools \(below\) for them\. Table 8\.Atomic query tools defined byLuo et al\.\([2025](https://arxiv.org/html/2605.08477#bib.bib24)\)with modifications, and their correspondence to functions in S\-expression\(Gu et al\.,[2021](https://arxiv.org/html/2605.08477#bib.bib11)\)\.##### Tools: We implemented sevenAtomic Query Tools\(Luo et al\.,[2025](https://arxiv.org/html/2605.08477#bib.bib24)\)444[https://github\.com/LHRLAB/KBQA\-o1](https://github.com/LHRLAB/KBQA-o1)with slight modifications \(Table[8](https://arxiv.org/html/2605.08477#A2.T8)\)\. These tools wrap functions in S\-expressions\. Chains of function calls can be converted into S\-expressions, and subsequently to SPARQL queries\. We derived the ground\-truth execution graph based on the S\-expression annotations\. Similar to KoPL tools, we perform soft schema grounding using embeddings and an LLM when the parameters specified by the planner are not found\. ##### Preprocessing: We sampled roughly 1,000 examples from the three datasets for creating training and test sets with balanced step lengths\. 1. \(1\)Deduplication:These datasets often contain multiple variations derived from the same logical form template\. We grouped examples by unique S\-expression to avoid bias toward common templates\. 2. \(2\)Validation:We ran SPARQL queries generated from the ground\-truth S\-expressions and filtered out those that returned answers differing from the ground truth\. 3. \(3\)Plan DAG Annotation:We parsed linear function lists into DAGs based on variable usage\. 4. \(4\)Sampling:We balanced the dataset across 4 complexity bins \(1\-3, 4\-5, 6\-7, 8\+\)\. 5. \(5\)Demonstration Retrieval:We retrieved 10 similar examples for each query using question embeddings, similar to the KQA Pro procedure\. #### B\.1\.3\.Multi\-Objective HotpotQA HotpotQA\(Yang et al\.,[2018](https://arxiv.org/html/2605.08477#bib.bib44)\)is widely used in existing work on LLM agents\. To analyze the effect of structural complexity, we followedZhou et al\.\([2025](https://arxiv.org/html/2605.08477#bib.bib48)\)and synthesized multi\-objective tasks by combiningkkindependent HotpotQA questions \(k=2,…,5k=2,\\dots,5\) in addition to the original HotpotQA instances \(k=1k=1\)\. ##### Tools: We implemented LLM sub\-agents\(Kim et al\.,[2024c](https://arxiv.org/html/2605.08477#bib.bib15)\)as tools\. We usedgpt\-4\.1\-mini\-2025\-04\-18as a backbone\. 1. \(1\)search\(question\): Answers a single\-hop question based on evidence retrieved from English Wikipedia555We used the 2017\-10\-01 dump of the first paragraphs released byYang et al\.\([2018](https://arxiv.org/html/2605.08477#bib.bib44)\)\.using the pre\-built Pyserini index withbge\-base\-en\-v1\.5\(Lin et al\.,[2021](https://arxiv.org/html/2605.08477#bib.bib21)\)\. By default, we retrieve the top\-10 relevant paragraphs for the search query\. In the controlled experiments \(§[4\.4](https://arxiv.org/html/2605.08477#S4.SS4)\), we limit this to the top\-1 result to simulate a less robust search tool\. 2. \(2\)commonsense\(question\): Performs logic/comparison using the internal knowledge of the LLM\. \[system\] YouarevalidatingwhetheraHotpotQAbridgequestionfollowsavalidlinearreasoningstructure\.Judgeonlyfromthewordingofthequestion;donotuseoutsideknowledgeorsupportingfacts\.Decomposethequestionintosteps,stateifanystepyieldsmultiplecandidates\(aset\),andthenoutputasingleJSONobject\. Coredefinitions \-Linearchain\(valid\):Astrictlysequentialchainof2\+stepswhereeachstepyieldsexactlyoneintermediateentity/propertythatdirectlyfeedsthenextstep\.Directpropertylookuporasingleyes/noverificationonthatsingleentityisfine\. \-Branching/setfiltering\(invalid\):Anystepproducesmultiplecandidatesthatmustbechecked/filteredindividually\(setoperations\),orthequestionasksforintersections/commonalities/comparisonsacrossmultipleentities\. Conservativeuniquenesspolicy\(donotassume\): \-Donotinferuniquenessfromplausibilityortypicalworldfacts\.Ifthewordingdoesnotmakeuniquenessexplicit,treatthestepasset\-producing\-\>invalid\. \-Rolesthatvaryovertime\(e\.g\.,"thepresidentof\[org\]","thecoachof\[team\]"\)arenon\-uniqueunlessthequestionspecifiesatime/tenure\(e\.g\.,ayear/season/ordinal/current\)\. \-Combiningmultipledescriptors\(e\.g\.,"Austrianforestcaretaker,naturalist,pseudoscientist"\)doesnotguaranteeuniqueness;stilltreataspotentiallymultipleunlessuniquelypinned\. Validpatterns\(usuallyuniquebywording\) \-Specifictitledwork\-\>role/property\(e\.g\.,"thedirectorof\[titledwork\]","thevocaliston’\[song\]’"\)\. \-Definite,singularrolestiedtoaspecificpropernounwithanexplicittime/ordinal\(e\.g\.,"theheadcoachof\[team\]in\[year\]","the42ndpresidentof\[country\]"\)\. \-Definitionalsuperlativesthatdenoteasingleitembydefinition\(e\.g\.,"thecapitalof\[country\]","thelargestcityin\[county\]"\)\. Invalidpatterns\(mustlabelinvalid\) \-Set\-producingfirststep\(guests,stars,castmembers,contestants,authors,filmsfeaturingX,etc\.\)followedbyfiltering\. \-Parallel/compoundquestionsseekingtwoindependentfactsorcommonalities\(e\.g\.,"WhatdoXandYhaveincommon?","WhichstarofAwasalsoinB?"\)\. \-Ambiguouscardinalityorvaguequalifiers\(e\.g\.,"former","long\-time"\)withouttimebounds;treatasset\-producing\. \-Anystepwhereuniquenessacrosstimeisnotfixedbythewording\. Outputrequirements \-ReturnonlyasingleJSONobjectwithkeys: \-"reasoning":2\-4concisesentencesthatenumeratethesteps\(e\.g\.,"Step1\.\.\.Step2\.\.\."\)andexplicitlystatewhere\(ifanywhere\)branchingoccurs\. \-"is\_valid":trueiflinear\(single\-path\),falseifbranching/set/parallelorifonlyasingle\-hoplookup\. \-Donotanswertheoriginalquestion\.Donotincludeextrakeys,disclaimers,orformatting\. Alwaysthinkstepbystepbasedsolelyonthequestiontext,erronthesideoftreatingambiguousstepsasbranching,indicatewhetherbranchingoccursandwhere,andthenprovidethefinalJSON\. \[user\] Question:\{question\} Isthereasoningstructureofthisquestionvalidaccordingtothecriteriaprovided?Decomposethequestionintostepwisereasoning\(labelStep1,Step2,etc\.\),indicateifandwherebranching\(setoperationsormultiplecandidateentities\)isrequired,andreturnonlytherequiredJSONobject\. Prompt 3:Validation of HotpotQA questions\[system\] You’retaskedwithdecomposingmulti\-hop"comparison"questionsintoexactlytwoparallel‘search‘stepsfollowedbyasingle‘reasoning‘step\-neverfewer\. ForeverycomparisonquestionthatrequireschoosingbetweenEntityAandEntityBbasedonanattribute,followthisstrictpattern: \-\*\*Node0:\*\*RetrieveEntityA’sattributewith‘search‘\. \-\*\*Node1:\*\*RetrieveEntityB’sattributewith‘search‘\. \-\*\*Node2:\*\*Use‘reasoning‘tocompare$0vs$1andoutputtherequiredentity/value\. Acomparisonquestion\*\*must\*\*includeonereasoningstepthatexplicitlycomparesthetworetrievedvalues\.Annotateeveryexampleaccordingly\. \#\#Context \-\*\*Goal:\*\*Annotatemulti\-hopQAexampleswithground\-truthreasoningplansasaDirectedAcyclicGraph\(DAG\)\.Youroutputsmustillustratetherequiredmulti\-steppattern\. \-\*\*ToolsAvailable:\*\* \-‘search\(query\)‘:Retrievesfacts\(returnsastring\)\. \-‘reasoning\(instruction\)‘:Appliessimplelogic\(comparison,filtering,conditionals\)onstrings\. \-\*\*AnnotatorRole:\*\*Theanswerandsupportingfactsareprovidedforyourannotation\.TheQAsystemyouannotateforcannotseethem;avoidanyleakageorhintinginyournodeinputs\. \#\#Required"Comparison"Pattern \-Forall"comparison"questions: \-TheDAGmustcontainexactly: \-Two‘search‘nodeswithnodependencies\(parallelretrieval\)\. \-One‘reasoning‘nodethatdependsonbothsearches\. \-Thetwo‘search‘queriesmustbeparallelinphrasing\(sameattributeaskedforbothentities\)sooutputsaredirectlycomparable\. \-Thefinaldecisionmustbemadeinthe‘reasoning‘node,notviaathird‘search‘\. \#\#Process 1\.\*\*Rephrase:\*\*Restatetheinputquestioninclear,naturalEnglish,keepingallconstraints\. 2\.\*\*ConstructDAG:\*\*Breakdownthequestionintoaminimal,strictlyorderedDAGusingtheaboverulesandthetoolsprovided\. \#\#NodeConstructionRules \-\*\*Self\-Contained:\*\*Eachnode’s‘input‘isaclear,standalonesentence\(replace‘$i‘withtheliteralvalue\)\. \-\*\*NaturalEmbedding:\*\*‘$i‘appearsasthoughtheentityorvalueisbeingaskedabout,notreferenced\.Nevermeta\-phrase\. \-\*\*LiteralUse:\*\*Treat‘$i‘asavaluetoinquireabout\-notadoc/source\. \-\*\*Conciseness:\*\*Eachnodecontainsexactly\*\*one\*\*inputsentence\. \-\*\*NoRedundantSearch:\*\*Neversearchforinformationalreadystatedinthequestion;instead,useitinareasoningnodeifinvolved\. \#\#ReasoningNodeRequirements\(Comparison\) \-Mustbeexactlyonesentence\. \-Mustexplicitlynamebothentitiesandreferencebothretrievedvalues\($0and$1\)\. \-Muststatewhattoreturn\(e\.g\.,"outputPersonA"\)\. \-Useadirectcomparativeconstruction;avoidawkwardrolesfor$i\(e\.g\.,donotwrite"If$0shows\.\.\."\)\. \-Template:"\[EntityA\]has\[attribute\]$0,and\[EntityB\]has\[attribute\]$1;basedonthequestion’scriterion,outputthecorrectentity\." \#\#OutputFormat ReturnasingleflatJSONobject,noMarkdown,nocomments: \-‘rephrased\_question‘\(string\) \-‘dag‘:Listofnodes: \-‘function‘:"search"or"reasoning" \-‘dependencies‘:Listofintegers\(nodeindices\) \-‘input‘:String\(standalone,with‘$i‘referencesasneeded\) Prompt 4:DAG annotation of HotpotQA questions \(Comparison:1/2\)\[system\]\(cont’d\) \#Example \{examples\} \#Notes \-Donotcollapsecomparisonquestionsintoonesearch\. \-Donotaddextrasearchstepsafterthetwoparallelsearches\. \-Alwaysretrievethesameattributeforbothentitieswithparallelphrasing\. \-Alwaysendwithexactlyonereasoningnodethatmakestheselectionandstateswhattooutput\. \-Neverreferencetheanswerorsupportingfactsinanynodeinputfields\. \[user\] GeneratethereasoningDAGforthisquestion\. Question:\{question\} Answer:\{answer\} SupportingFacts:\{supporting\_facts\} \*\*Notes\*\* \-Donotcollapsecomparisonquestionsintoonesearch\. \-Donotaddextrasearchstepsafterthetwoparallelsearches\. \-Alwaysretrievethesameattributeforbothentitieswithparallelphrasing\. \-Alwaysendwithexactlyonereasoningnodethatmakestheselectionandstateswhattooutput\. \-Neverreferencetheanswerorsupportingfactsinanynodeinputfields\. Prompt 5:DAG annotation of HotpotQA questions \(Comparison:2/2\) ##### Preprocessing: We sampled examples from the training and development sets of HotpotQA and synthesized new training and test splits with multi\-objective questions \(k=1,⋯,5k=1,\\cdots,5\)\. 1. \(1\)Validation:As the original HotpotQA questions were generated based on pre\-selected pairs of Wikipedia paragraphs, some questions, particularly bridge questions, are not suitable for tool calling settings\.666For example, some bridge questions involve set questions with very large number of intermediate answers \(e\.g\., all players in a specific team\)\. To control reasoning structures, we usedgpt\-5\-mini\-2025\-08\-07to select bridge question that involve single entities as pivot answers\. See Prompt[3](https://arxiv.org/html/2605.08477#LST3)\. 2. \(2\)Sampling:We sampled 200 examples uniformly across combinations of bridge/comparison types for eachkk, resulting in 1,000 QA pairs for each of the training and test splits\. 3. \(3\)Plan DAG Annotation:We usedgpt\-5\.2\-2025\-12\-11with medium reasoning effort to generate ground\-truth DAGs based on the gold supporting paragraphs for each sub\-question\. The annotation was straightforward for most questions \(bridge questions generatesearch→\\tosearch, and comparison questions generate two independentsearchsteps followed by a comparison step withcommonsense\), but some bridge questions involved additional reasoning steps\. We also generatedquestiontexts of tool calls to be used in in\-context demonstrations\. Prompt[4](https://arxiv.org/html/2605.08477#LST4)shows the instructions to convert comparison questions\. We used similar instructions for bridge questions\. 4. \(4\)Demonstration Retrieval:We retrieved 10 demonstration examples with the samekkfor each query using question embeddings\. ### B\.2\.Planner Details \[system\]\(FH\) Usetheavailabletoolstoanswertheuser’squestionstepbystep,generatingtheentireplan\.Use$itorefertotheoutputofstepi\(0\-indexed\)\. \*\*Tools\*\* \{tool\_definitions\} Youcanaccesstheknowledgebasethroughtheprovidedtools\. \*\*GlossaryofKBconcepts:\*\* \-Entity:ThemostbasiciteminKB\. \-Concept:Theabstractionofasetofentities,e\.g\.,basketballplayer\. \-Relation:Thelinkbetweenentitiesorconcepts\.Entitiesarelinkedtoconceptsviatherelationinstanceof\.Conceptsareorganizedintoatreestructureviarelationsubclassof\. \-Attribute:Theliteralinformationofanentity\.Anattributehasakeyandavalue,whichisoneoffourtypes:string,number,date,andyear\.Thenumbervaluecanhaveanextraunit,e\.g\.,206centimetre\. \-Relationalknowledge:Thetriplewithform\(entity,relation,entity\),e\.g\.,\(LeBronJamesJr\.,father,LeBronJames\)\. \-Literalknowledge:Thetriplewithform\(entity,attributekey,attributevalue\),e\.g\.,\(LeBronJames,height,206centimetre\)\. \-Qualifierknowledge:Thetriplewhoseheadisarelationalorliteraltriple,e\.g\.,\(\(LeBronJames,draftedby,ClevelandCavaliers\),pointintime,2003\)\.Aqualifieralsohasakeyandavalue\. \*\*Examples\*\* \{demonstrations\} \[system\]\(SH\) Usetheavailabletoolstoanswertheuser’squestionstepbystep,generatingoneactionatatime\.Use$itorefertotheoutputofstepi\(0\-indexed\)\. \(\.\.\.sameasFH\) \[user\] Question:\{question\} Prompt 6:Planning prompts for KQA Pro\. The glossary was defined byCao et al\.\([2022](https://arxiv.org/html/2605.08477#bib.bib4)\)\.\[system\]\(FH\) Usetheavailabletoolstoanswertheuser’squestionstepbystep,generatingtheentireplan\.Use$itorefertotheoutputofstepi\(0\-indexed\)\. \*\*Tools\*\* \{tool\_definitions\} Usetheavailabletoolstoanswertheuser’squestionstepbystep \*\*Examples\*\* \{demonstrations\} \[system\]\(SH\) Usetheavailabletoolstoanswertheuser’squestionstepbystep,generatingoneactionatatime\.Use$itorefertotheoutputofstepi\(0\-indexed\)\. \(\.\.\.sameasFH\) \[user\] Question:\{question\} Prompt 7:Planning prompts for multi\-objective HotpotQA\[user\] ExecutionResult: \{execution\_result\} Thelaststepreturnedanerrorornoresults\.Produceanappend\-onlycontinuation:addnewstepsstartingatindex\{start\_index\}thatcontinuefromtheexecutedstepsabove\.Donotrepeat,modify,orreissueanyexecutedstep,andonlyreferenceexistingoutputs\. Ifemptyresultsareexpectedandvalidforthequestion,producethecontinuationstepsfromtheoriginalplanstartingatstep\{start\_index\}\.Otherwise,addcorrectivestepsstartingat\{start\_index\}\(e\.g\.,adjusttoolparameters,ortryalternativetools\)andproceedtowardtheanswer\. Prompt 8:Replanning prompt \(FH\)\. This user message is appended to the message list\.\{execution\_result\}is replaced with a list of executed actions and their results\.For both FH and SH planners, we use minimal prompts containing task instructions, tool definitions \(JSON string\)777We tested other formats like TypeScript and Markdown but found no significant difference\. The JSON format generally performed well\., and demonstration examples \(Prompts[1](https://arxiv.org/html/2605.08477#LST1)and[7](https://arxiv.org/html/2605.08477#LST7)\)\. We include 10 in\-context demonstrations retrieved viaBAAI/bge\-base\-en\-v1\.5embeddings\. FH generates the full list of steps\. If execution fails, the agent receives the error and replans up to 8 times \(Prompt[8](https://arxiv.org/html/2605.08477#LST8)\)\. SH generates one step, receives the observation, and generates the next step\. Both FH and SH agents are allowed to issue up to 30 tool calls\. We use structured output functionality to constrain LLM generation to valid JSON structures representing a sequence of tool calls\. Specifically, we use theanyOfkeyword in the JSON Schema888[https://json\-schema\.org/](https://json-schema.org/)to limit the vocabulary of action names and their parameter names\. When the planner generates invalid parameter values \(e\.g\., references to non\-existent steps\), we prompt it to retry with a simple error message: “Your response is in an invalid format\. Please read the instructions carefully and try again,” allowing up to 8 retries\. We use the following backbone models via public APIs \(OpenAI API for GPT, Fireworks API for Qwen, and Vertex AI API for Gemini\)\. - •gpt\-4\.1\-mini\-2024\-07\-18\(Temperature 0\) - •gpt\-5\-mini\-2025\-08\-07\(Temperature 1, Medium reasoning effort\) - •qwen3\-235b\-a22b\-thinking\-2507\(Temperature 0\) - •gemini\-3\-flash\-preview\(Temperature 0, Medium reasoning effort\) We set the context limit to be sufficiently large \(10k\) and use default values for other parameters\. ### B\.3\.Evaluation Details \[system\]\(KBQA\) Youareastrict,impartialgraderofanswercorrectness\. \*\*Goal:\*\*CompareSystemOutputtoCorrectAnsweronly\.Donotuseoutsideknowledge;treatCorrectAnswerasgroundtruth\. \*\*Output:\*\*Replywithexactlyonelabel\(noquotes,nopunctuation,noextratext\): \-correct \-partially\_correct \-incorrect \-refusal/unsure \*\*Scoringrules:\*\* 1\)Non\-answer\-\>refusal/unsure: \-SystemOutputrefuses,asksforclarification,says"unknown"/"noinformation",orotherwisedoesnotcontainananswervalue\. 2\)Exactmatch\-\>correct: \-SystemOutputconveysthesameanswervalue\(s\)asinCorrectAnswer,andnoadditionalanswervalues\. \-Treatlistsassets:ignoreorder,casing,surroundingfiller,andduplicaterepetitionsofthesamecorrectvalue\(s\)\. \-Ignorepurelyexplanatoryfiller\(e\.g\.,"theansweris\.\.\."\)\. 3\)Partialoverlap\-\>partially\_correct: \-AtleastoneexpectedanswervalueappearsintheSystemOutput,butitismissinganyotherrequiredvaluesand/orincludesextraanswervaluesnotintheCorrectAnswer\. 4\)Mismatch\-\>incorrect: \-NoneoftheexpectedanswervaluesappearintheSystemOutput,oranyprovidedvaluecontradictstheCorrectAnswer\. \*\*Definitionsandmatchingguidance:\*\* \-Answervalue:anentity\(IDorname\),number,date,orotheratomicitem\.Treatcomma/semicolon/newline/bulletedlistsandclearconjunctionsasmultiplevalues\. \-Entitymatching:"m\.xxxxx\(Name\)"/"Qxxxxx\(Name\)"matchesifeithertheIDortheNameappearsinSystemOutput\(case\-insensitive\)\. \-Numeric/datematching:requireexactequalityunlessCorrectAnswerexplicitlylistsmultipleacceptablevalues\.Rounded/truncatednumbersareincorrectwhenCorrectAnswerprovidesamoreprecisevalue\. \-Type/Hierarchy:Donotgivecreditforbroader/narrower/relatedcategories,roles,orspecializations\(e\.g\.,hypernym/hyponym\)\.Onlytheexactexpectedvaluecounts\. \-Strictcomparison:Donotusethequestiontextoroutsideknowledgetoinferequivalence\.ComparetheSystemOutputstrictlytotheCorrectAnswertext\. \-Extrainformation: \-Extraexplanatorytextmaybeignored\. \-Extraanswervalues\(additionalentities/values\)maketheresultpartially\_correct\. \[user\] Question:\{question\} CorrectAnswer:\{correct\_answer\} SystemOutput:\{system\_output\} Prompt 9:Evaluation prompt for KBQA tasks\[system\] Youareastrict,impartialgraderofanswercorrectness\. \*\*Goal:\*\*CompareSystemOutputtoCorrectAnsweronly\.Donotuseoutsideknowledge;treatCorrectAnswerasgroundtruth\. \*\*Output:\*\*Replywithexactlyonelabel\(noquotes,nopunctuation,noextratext\): \-correct \-partially\_correct \-incorrect \-refusal/unsure \*\*Scoringrules:\*\* 1\)Non\-answer\-\>refusal/unsure: \-SystemOutputrefuses,asksforclarification,says"unknown"/"noinformation",orotherwisedoesnotcontainananswervalue\. 2\)Exactmatch\-\>correct: \-TheSystemOutputconveysexactlythesameanswervalue\(s\)astheCorrectAnswerandnoadditionaldistinctanswervalues,inthesameorderformulti\-partanswers\. \-Ignorecasing,articles,punctuation/spacing,andduplicaterepetitions\. \-Ignoreexplanatoryordescriptivecontextthatdoesnotadddistinctanswervalues\(e\.g\.,prosearoundtheanswer,unitqualifiers,city\+stateafteracity,manufacturerprefixes,honorifics/titles,adjectiveslike"WorldFamous"\)\. \-Acceptunambiguousaliasesorminornamevariantsforthesameentity\(e\.g\.,nicknamesvsfullnames,with/withoutmiddlenames,commonalternateforms\)whentheyclearlyrefertothesameentityanddonotintroduceadifferentone\. \-Numbers/dates:treatwordsvsdigitsanddigit\-grouping\(e\.g\.,1,840vs1840\)asequivalent\.Iftheexpectedvalueisayear,afulldatecontainingthatyearisacceptable\.Iftheexpectedvalueincludesmonth\+year,monthaloneisincomplete\. \-Approximation:IfthequestionorCorrectAnswersignalsapproximation\(about/approximately/around/over\),allowconsistentapproximateorinequalityphrasingnearthevalue\. 3\)Partialoverlap\-\>partially\_correct: \-Atleastoneexpectedanswervalueappears,butanyrequiredvalueismissingand/oratleastonevalueisincorrect/contradictory\. \-Theoutputaddsextradistinctanswervaluesforaslot\(e\.g\.,listingmultipleroleslike"actoranddirector",multiplecandidateentitieswithand/or/slashes/lists\)beyondwhattheCorrectAnswerexpects\. 4\)Mismatch\-\>incorrect: \-NoneoftheexpectedanswervaluesappearintheSystemOutput,oranyprovidedvaluecontradictstheCorrectAnswer\. \*\*Definitionsandmatchingguidance:\*\* \-Answervalue:anatomicitemsuchasanentity\(IDorname\),number,date/time,oryes/no\. \-Entitymatching:givecreditonlywhentheexpectedentity\(name/IDoraclearalias/variant\)isexplicitlypresentintheSystemOutput\.Donotawardcreditformerelyimplyingtheanswerwithoutnamingit\.Donotgivecreditforbroader/narrower/relatedcategories\. \-Order:Whenmultiplequestionsareasked,thecorrectanswersareinthatorder;theSystemOutputmustaligntobecorrect\. \[user\] Question:\{question\} CorrectAnswer:\{correct\_answer\} SystemOutput:\{system\_output\} Prompt 10:Evaluation prompt for multi\-objective HotpotQATo account for lexical variation in answers and handle verbose expressions \(e\.g\., “The answer is XX”\), we use LLM\-as\-a\-judge to evaluate the correctness of the final answer\. The evaluator prompts were tuned semi\-automatically on a held\-out set of 100 examples per dataset to ensure high agreement with human judgment\. See Prompts[9](https://arxiv.org/html/2605.08477#LST9)and[10](https://arxiv.org/html/2605.08477#LST10)\. For the GEE analysis, we use the implementation in statsmodels\.999[https://www\.statsmodels\.org/](https://www.statsmodels.org/)To handle repeated trials, we use the question ID for clustered analysis\. The formula for the regression is as follows, whered∗d^\{\*\}andb∗b^\{\*\}are the normalized depth and breadth of the plan,xSHx\_\{\\text\{SH\}\}is a binary variable indicating whether the planner is SH\. Atomic KBQA: \(4\)logit\(P\(y=1\)\)=β0\+βdd∗\+βbb∗\+βSHxSH\+βd:SH\(d∗×xSH\)\+βb:SH\(b∗×xSH\)\+∑Dataseti\+∑LastStepj⏟Fixed Effects\\text\{logit\}\(P\(y=1\)\)=\\beta\_\{0\}\+\\beta\_\{d\}d^\{\*\}\+\\beta\_\{b\}b^\{\*\}\+\\beta\_\{\\text\{SH\}\}x\_\{\\text\{SH\}\}\+\\beta\_\{d:\\text\{SH\}\}\(d^\{\*\}\\times x\_\{\\text\{SH\}\}\)\+\\beta\_\{b:\\text\{SH\}\}\(b^\{\*\}\\times x\_\{\\text\{SH\}\}\)\+\\underbrace\{\\sum\\text\{Dataset\}\_\{i\}\+\\sum\\text\{LastStep\}\_\{j\}\}\_\{\\text\{Fixed Effects\}\}where C\(dataset\) is a dummy variable for the dataset \(GrailQA, WebQSP, GraphQ\),LastStepis a dummy variable for the type of the last step in the plan\. KQA Pro: \(5\)logit\(P\(y=1\)\)=β0\+βdd∗\+βbb∗\+βSHxSH\+βd:SH\(d∗×xSH\)\+βb:SH\(b∗×xSH\)\+∑LastStepi⏟Fixed Effects\\text\{logit\}\(P\(y=1\)\)=\\beta\_\{0\}\+\\beta\_\{d\}d^\{\*\}\+\\beta\_\{b\}b^\{\*\}\+\\beta\_\{\\text\{SH\}\}x\_\{\\text\{SH\}\}\+\\beta\_\{d:\\text\{SH\}\}\(d^\{\*\}\\times x\_\{\\text\{SH\}\}\)\+\\beta\_\{b:\\text\{SH\}\}\(b^\{\*\}\\times x\_\{\\text\{SH\}\}\)\+\\underbrace\{\\sum\\text\{LastStep\}\_\{i\}\}\_\{\\text\{Fixed Effects\}\}whereLastStepis a dummy variable for the type of the last step in the plan\. Multi\-objective HotpotQA: \(6\)logit\(P\(y=1\)\)=β0\+βdd∗\+βbb∗\+βSHxSH\+βd:SH\(d∗×xSH\)\+βb:SH\(b∗×xSH\)\+βhas\_bridgexhas\_bridge\+βhas\_comparisonxhas\_comparison⏟Fixed Effects\\text\{logit\}\(P\(y=1\)\)=\\beta\_\{0\}\+\\beta\_\{d\}d^\{\*\}\+\\beta\_\{b\}b^\{\*\}\+\\beta\_\{\\text\{SH\}\}x\_\{\\text\{SH\}\}\+\\beta\_\{d:\\text\{SH\}\}\(d^\{\*\}\\times x\_\{\\text\{SH\}\}\)\+\\beta\_\{b:\\text\{SH\}\}\(b^\{\*\}\\times x\_\{\\text\{SH\}\}\)\+\\underbrace\{\\beta\_\{\\text\{has\\\_bridge\}\}x\_\{\\text\{has\\\_bridge\}\}\+\\beta\_\{\\text\{has\\\_comparison\}\}x\_\{\\text\{has\\\_comparison\}\}\}\_\{\\text\{Fixed Effects\}\}wherexhas\_bridgex\_\{\\text\{has\\\_bridge\}\}andxhas\_comparisonx\_\{\\text\{has\\\_comparison\}\}are binary variables indicating whether the question contains a bridge or comparison sub\-question, respectively\. ### B\.4\.External Resources Table 9\.A list of pretrained language models used in this study\.Table 10\.A list of datasets used in this study\.Table 11\.A list of software used in this study\.Tables[9](https://arxiv.org/html/2605.08477#A2.T9),[10](https://arxiv.org/html/2605.08477#A2.T10)and[11](https://arxiv.org/html/2605.08477#A2.T11)list the key external resources on which this study relies\. Our use of these resources complies with their respective terms of use\. ## Appendix CAdditional Experimental Results ### C\.1\.Detailed GEE Coefficients Table 12\.GEE coefficients for KQA Pro experiments \(1/2\)Table 13\.GEE coefficients for KQA Pro experiments \(2/2\)Table 14\.GEE coefficients for Atomic KBQA experimentsTable 15\.GEE coefficients for multi\-objective HotpotQA experimentsTables[12](https://arxiv.org/html/2605.08477#A3.T12),[13](https://arxiv.org/html/2605.08477#A3.T13),[14](https://arxiv.org/html/2605.08477#A3.T14), and[15](https://arxiv.org/html/2605.08477#A3.T15)show all the GEE coefficients\. ### C\.2\.Case Study on Repetitive Tool Calls Table 16\.SH with Gemini\-3\-Flash under low\-robustness retrieval \(top\-1\) on multi\-objective HotpotQA \(k=2k=2\)\. This table focuses on one sub\-question: “Which director, John Schlesinger or Barbara Albert, was also a writer and film producer?”Step 1ReasoningI’m currently engaged in evaluating both John Schlesinger and Barbara Albert, determining if each functioned as a writer and producer\. After gathering the facts, a comparison will readily reveal the answer\. Parallel to this, I’m setting up searches to uncover the release companies behind Cinderella and The Emperor’s New Groove to formulate my response regarding the latter question\.Tool¿ search\(Was John Schlesinger a writer and film producer?\)OutputError in search: Failed to find the answer to ”Was John Schlesinger a writer and film producer?”No supporting information found in the search result: information is about Richard Schlesinger, not John Schlesinger\.Retry with a different question or try a different tool\.Step 2ReasoningI’m starting by investigating the first question\. I’ve initiated Step 0: searching for John Schlesinger’s roles\. The plan is to establish if he wrote or produced films, as well as directed\. After that, I will move on to step 1 and do the same for Barbara Albert\.Tool¿ search\(What were the occupations of director John Schlesinger?\)OutputError in search: Failed to find the answer to ”Was John Schlesinger a writer and film producer?”No supporting information found in the search result: information is about Richard Schlesinger, not John Schlesinger\.Retry with a different question or try a different tool\.Step 3ReasoningI’m now formulating my search queries\. I’ll begin by querying about the occupations of John Schlesinger and Barbara Albert\. Following that, I’ll identify the release companies for ”Cinderella” and ”The Emperor’s New Groove\.” This information will then allow me to address the initial questions posed\. I’m focusing on efficiency and clarity to ensure accurate results\.Tool¿ search\(What were the occupations of John Schlesinger?\)OutputError in search: Failed to find the answer to ”Was John Schlesinger a writer and film producer?”No supporting information found in the search result: information is about Richard Schlesinger, not John Schlesinger\.Retry with a different question or try a different tool\.Step 4ReasoningI’ve begun looking into the roles of John Schlesinger and Barbara Albert\. Initial searches for their occupations are underway\. I’ll need to sift through the results to see if either one of them was both a writer and a film producer\. The plan is to compare the findings to pinpoint the director who fulfills both criteria\. I’ll then move on to the second question\.Tool¿ search\(What were the occupations of John Schlesinger?\)OutputError in search: Failed to find the answer to ”Was John Schlesinger a writer and film producer?”No supporting information found in the search result: information is about Richard Schlesinger, not John Schlesinger\.Retry with a different question or try a different tool\.Step 5ReasoningI’m now starting to investigate John Schlesinger\. My initial step is to determine if he also worked as a writer and film producer\. Then, I’ll turn my attention to Barbara Albert, following the same procedure\. After that, I will compare the data to formulate a response to the user’s first question\.Tool¿ search\(What were the occupations of director John Schlesinger?\)OutputError in search: Failed to find the answer to ”Was John Schlesinger a writer and film producer?”No supporting information found in the search result: information is about Richard Schlesinger, not John Schlesinger\.Retry with a different question or try a different tool\.Table 17\.FH with Gemini\-3\-Flash under low\-robustness retrieval \(top\-1\) on multi\-objective HotpotQA \(k=2k=2\)\. This table focuses on one sub\-question: “Which director, John Schlesinger or Barbara Albert, was also a writer and film producer?”InitialReasoningI’ve initiated searches to determine if John Schlesinger and Barbara Albert were both writers and producers, as planned\. Once I have those details, a comparison will be conducted to pinpoint the correct individual\. I’m also preparing to find the release companies for ”Cinderella” and ”The Emperor’s New Groove” to compare them\.Plan¿ \[$0 = search\(Was John Schlesinger a writer and film producer?\), $1 = search\(Was Barbara Albert a writer and film producer?\), $2 = reasoning\(John Schlesinger is described as $0, and Barbara Albert is described as $1; based on which one was also a writer and film producer, output the correct person\.\), …\]OutputError in search $0: Failed to find the answer to ”Was John Schlesinger a writer and film producer?”No supporting information found in the search result: information is about Richard Schlesinger, not John Schlesinger\.Retry with a different question or try a different tool\.Replanning 1ReasoningI’ve restarted my search for John Schlesinger’s occupations, focusing on director, writer, and producer roles, since my initial attempt didn’t yield the right information\. Also, I’m planning to look into Barbara Albert’s roles\. My next step will be to compare the findings to pinpoint the director who was also a writer and film producer\.Plan¿ \[$1 = search\(What were the occupations of director John Schlesinger?\), $2 = search\(What were the occupations of director Barbara Albert?\), …\]OutputError in search $1: Failed to find the answer to ”Was John Schlesinger a writer and film producer?”No supporting information found in the search result: information is about Richard Schlesinger, not John Schlesinger\.Retry with a different question or try a different tool\.Replanning 2ReasoningI’m now focusing on Barbara Albert’s occupations, and then refining how I search for John Schlesinger’s roles\. I plan to query directly whether Schlesinger worked as a writer and producer\. After that, I’ll compare the findings to answer the first question\. For the second question, I will pinpoint the release companies for both films, and compare the results\.Plan¿ \[$2 = search\(What are the occupations of Barbara Albert \(director\)?\), $3 = search\(Was John Schlesinger a writer and film producer in addition to being a director?\), …\]OutputError in search $3: Failed to find the answer to ”Was John Schlesinger a writer and film producer?”No supporting information found in the search result: information is about Richard Schlesinger, not John Schlesinger\.Retry with a different question or try a different tool\.Replanning 3ReasoningI’m focusing on John Schlesinger’s occupations, trying to confirm if he was also a writer or producer\. I’ll use a more specific search term this time around\. If that doesn’t work, I’ll leverage Barbara Albert’s information to answer the first question\. After that, I’ll tackle the release companies\.Plan¿ \[$4 = search\(What were the professional roles of the English film director John Schlesinger?\), $5 = reasoning\(Based on the occupations of Barbara Albert \($2\) and John Schlesinger \($4\), which one was also a writer and film producer?\), …\]Tables[16](https://arxiv.org/html/2605.08477#A3.T16)and[17](https://arxiv.org/html/2605.08477#A3.T17)show the full trajectories corresponding to the case study examples in Table[6](https://arxiv.org/html/2605.08477#S4.T6)\(§[4\.5](https://arxiv.org/html/2605.08477#S4.SS5)\)\.
Similar Articles
When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents
The ToolMaze benchmark evaluates LLM agents' ability to handle real-world tool failures, revealing that implicit semantic failures cause the largest performance drops and that dynamic replanning remains a critical bottleneck not addressed by scaling or prompting.
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents
This arXiv survey (1,547 papers, 2024-2026) systematically maps the field of long-horizon LLM agents, disambiguating long-horizon, long-context, and long-term memory, and organizing research into six lifecycle categories while identifying the core 'horizon gap' and open measurement problems.
Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents
This paper investigates how LLM agents lose plan information as it gets evicted from context during long interactions. Using replay pairing and compression stress tests, the authors show that standard agents do not carry plans as persistent state, and propose diagnostics to measure plan signal decay.
Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems
This paper presents a method for compiling repeated standard operating procedure steps into validated, versioned tools before deployment, replacing inference-time code generation. In a fulfillment center alarm-triage system, this approach reduces p50 latency by 42% and end-to-end error rate by up to 53%.
Thoughts on Long-Horizon Agents
The author discusses insights from rebuilding self-service onboarding as agentic flows, emphasizing how long-horizon agents that combine deterministic task management with non-deterministic LLM actions improve reliability and reduce hallucinations in enterprise AI deployments.