FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows
Summary
FlowScout is a framework that automatically generates tool-integrated agentic workflows from historical task-solving records, using Monte Carlo tree search guided by execution feedback. Experiments show it improves tool invocation correctness and execution score over baselines.
View Cached Full Text
Cached at: 08/12/26, 08:27 AM
# FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows
Source: [https://arxiv.org/html/2608.10039](https://arxiv.org/html/2608.10039)
###### Abstract
Agentic workflows have become an important abstraction for building reliable LLM\-based automation systems by organizing large language models \(LLMs\), tools, and control logic into explicit execution structures\. However, constructing high\-quality agentic workflows remains largely manual and requires substantial domain expertise\. Recent studies have explored automatic agentic workflow generation from historical task\-solving records, but they mainly produce LLM\-centric workflows, where real tool executions are abstracted and simulated by LLM nodes, limiting theusability and stabilityof generated workflows\. To address these limitations, we proposeFlowScout, an execution\-guided framework for generating tool\-integrated agentic workflows from historical task\-solving records\. Specifically,FlowScoutrepresents an agentic workflow as a directed graph composed of LLM nodes, tool\-calling nodes, and dependency edges\. It first minesa commontool coordination skeleton from historical records to construct an initial workflow,and then refines the workflow topology through Monte Carlo tree search guided by execution feedback\.We evaluateFlowScouton four representative task domains, and compare it with three baselines,i\.e\.,PM4Py, ReAct and AFlow\. Experimental results show that agentic workflows generated byFlowScoutimprove the tool invocation correctness by at least92\.69%and theexecution scoreby at least17\.66%over the baselines, while achieving lower performance variation across repeated runs\.
## IIntroduction
The rapid development of large language models \(LLMs\)\[[24](https://arxiv.org/html/2608.10039#bib.bib2),[9](https://arxiv.org/html/2608.10039#bib.bib3)\]has introduced LLM agents as a new automation paradigm for complex tasks\. By combining LLM reasoning with external tools, LLM agents can interpret user queries, decompose tasks, invoke tools, and adapt subsequent actions according to intermediate results\[[42](https://arxiv.org/html/2608.10039#bib.bib19),[32](https://arxiv.org/html/2608.10039#bib.bib20),[41](https://arxiv.org/html/2608.10039#bib.bib21),[46](https://arxiv.org/html/2608.10039#bib.bib22),[3](https://arxiv.org/html/2608.10039#bib.bib23),[37](https://arxiv.org/html/2608.10039#bib.bib24),[13](https://arxiv.org/html/2608.10039#bib.bib25),[30](https://arxiv.org/html/2608.10039#bib.bib26),[28](https://arxiv.org/html/2608.10039#bib.bib27),[14](https://arxiv.org/html/2608.10039#bib.bib29)\]\. However, the flexibility and the black\-box nature of LLM reasoning make agent executions less stable and harder to inspect\. For similar user queries, LLM agents may make inconsistent decisions, follow different action orders, and select different tools, leading to unpredictable execution results\[[1](https://arxiv.org/html/2608.10039#bib.bib40)\]\.
To improve the usability and stability of agent executions, agentic workflows provide a promising way to combine the strengths of traditional workflows and LLMs\[[43](https://arxiv.org/html/2608.10039#bib.bib7)\]\. Instead of allowing an LLM agent to freely decide every action at runtime, an agentic workflow organizes LLMs, tools, and control logic into an explicit execution structure\[[19](https://arxiv.org/html/2608.10039#bib.bib39)\]\. Such agentic workflows preserve the semantic reasoning capability of LLMs, while constraining their execution with reusable and observable structures\. Therefore, agentic workflows have become an increasingly important abstraction for building reliable LLM\-based automation systems\[[45](https://arxiv.org/html/2608.10039#bib.bib4),[10](https://arxiv.org/html/2608.10039#bib.bib5),[27](https://arxiv.org/html/2608.10039#bib.bib6),[48](https://arxiv.org/html/2608.10039#bib.bib9),[34](https://arxiv.org/html/2608.10039#bib.bib8)\]\.
Prior studies have explored workflow generation, including synthesizing web\-service workflows from formal service specifications\[[33](https://arxiv.org/html/2608.10039#bib.bib11),[2](https://arxiv.org/html/2608.10039#bib.bib12)\], extracting process models from execution logs\[[7](https://arxiv.org/html/2608.10039#bib.bib13),[35](https://arxiv.org/html/2608.10039#bib.bib17)\], and generating software workflows from natural language descriptions\[[23](https://arxiv.org/html/2608.10039#bib.bib14),[39](https://arxiv.org/html/2608.10039#bib.bib15)\]\. These studies provide important foundations for deriving workflows from specifications or observed task executions\. However, they are mainly designed for settings where task activities have explicit semantics, stable execution patterns, and well\-defined interfaces\. As a result, they are difficult to directly apply to open\-ended tasks that require semantic understanding, dynamic decision\-making, and flexible coordination among external tools\. Currently, the construction of high\-quality agentic workflows remains largely manual\. Existing agent frameworks,e\.g\.,LangChain\[[15](https://arxiv.org/html/2608.10039#bib.bib41)\], LangGraph\[[16](https://arxiv.org/html/2608.10039#bib.bib42)\], and AutoGen\[[36](https://arxiv.org/html/2608.10039#bib.bib43)\], and low\-code platforms,e\.g\.,Dify\[[17](https://arxiv.org/html/2608.10039#bib.bib44)\]and Coze\[[6](https://arxiv.org/html/2608.10039#bib.bib45)\], provide useful abstractions for implementing agentic workflows, but developers still need to decide how the agentic workflow should be built for a target task domain, requiring substantial domain expertise and iterative debugging effort\.
Recently, a few studies\[[44](https://arxiv.org/html/2608.10039#bib.bib50),[47](https://arxiv.org/html/2608.10039#bib.bib51)\]have started to explore the automatic generation of agentic workflows\. These approaches employ historical task\-solving records,e\.g\.,tool orchestration logs collected from human demonstrations, or LLM agent traces where tool invocation orders have been validated by successful task completion, to generate domain\-specific agentic workflows\. However, they can only generate LLM\-centric workflows, where nodes are all LLM operations such as planning, programming, formatting, revision, and context generation, failing to explicitly preserve the real tool invocations embedded in historical task\-solving records\. Thus, the generated agentic workflows still rely heavily on LLMs to simulate tool execution, which may lead to unstable behaviors, hallucinated outputs, and limited interpretability of domain\-specific tool coordination\.
To address these limitations, we proposeFlowScout, an execution\-guided framework for automatically generating tool\-integrated agentic workflows from historical task\-solving records\. Specifically, we model an agentic workflow as a directed graph composed of LLM nodes, tool\-calling nodes, and dependency edges that encode data flow and control flow, and formulate its generation as a search problem over candidate workflow graphs\.Given available tools and a set of historical task\-solving records consisting of domain\-specific user queries, reference execution results, and validated tool orchestration logs,FlowScoutfirst mines the most common tool coordination skeleton from historical records, and constructs an initial agentic workflow by augmenting the skeleton with necessary LLM nodes\. Then,FlowScoutperforms execution\-guided graph search, where candidate workflows are executed on historical user queries and their outputs are compared against reference results\. The feedback guides Monte Carlo tree search\[[4](https://arxiv.org/html/2608.10039#bib.bib30)\]to refine the workflow graph by adding, removing or refining LLM and tool\-calling nodes, as well as modifying dependency edges\. Finally,FlowScoutproduces a reusable and stable agentic workflow that explicitly coordinates real tools for solving new user queries in the target domain\.
We have conducted extensive experiments to evaluate the effectiveness and efficiency ofFlowScout\. First, we evaluateFlowScouton four representative task domains in ToolBench\[[28](https://arxiv.org/html/2608.10039#bib.bib27)\],i\.e\.,finance, sports, travel, and weather, and compare it with three baselines,i\.e\.,PM4Py\[[35](https://arxiv.org/html/2608.10039#bib.bib17)\], ReAct\[[42](https://arxiv.org/html/2608.10039#bib.bib19)\], and AFlow\[[44](https://arxiv.org/html/2608.10039#bib.bib50)\]\. The results show that agentic workflows generated byFlowScoutimprove the tool invocation correctness by at least92\.69%and theexecution scoreby at least17\.66%over the baselines,while reducing the coefficient of variation across repeated runs by 62\.09% and 27\.90% compared with ReAct and AFlow, respectively\.Second,the workflows generated byFlowScoutincur at least 24\.12% higher runtime cost than ReAct and AFlow for improved effectiveness\. Besides, ablation studies show the contribution of our workflow miner and Monte Carlo tree search inFlowScout\. Finally, we demonstrate the generalization capability of agentic workflows generated byFlowScoutby deploying them at runtime with tools and LLMs different from those used during workflow generation\.
The main contributions of this work are as follows:
- •We represent an agentic workflow as a structured graph consisting of LLM nodes, tool\-calling nodes,and edges that encode data dependency or control dependency, and cast the generation of agentic workflows as a graph search problem\.
- •We proposeFlowScout, an execution\-guided generation framework, that automatically searches for tool\-integrated agentic workflows based on historical task\-solving records, improving the usability and stability of generated workflows\.
- •We implement a prototype ofFlowScout, and conduct experiments to demonstrate its effectiveness and efficiency\.
## IIProblem Formulation
Given available tools and historical task\-solving records, including user queries, reference execution results, and validated tool orchestration logs, we aim to automatically generate tool\-integrated agentic workflows for autonomous task completion\. We first define such workflows and then formulate their generation as a search problem over valid workflow graphs\.
### II\-AWorkflow Definition
Given a task domain with a set of available tools𝒯\\mathcal\{T\}, we define a tool\-integrated agentic workflow as a directed graph\.
###### Definition 1\(Tool\-Integrated Agentic Workflow\)\.
A tool\-integrated agentic workflow is a directed graph𝒢=\(𝒱,ℰ\)\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\}\), where𝒱\\mathcal\{V\}is a set of workflow nodes consisting of LLM nodes and tool\-calling nodes, andℰ\\mathcal\{E\}is a set of directed dependency edges that define the control flow and data flow among nodes\.
###### Definition 2\(LLM Node\)\.
An LLM nodevl∈𝒱v\_\{l\}\\in\\mathcal\{V\}, which performs semantic reasoning on textual context, is a tuplevl=⟨δvl,ρvl⟩v\_\{l\}=\\langle\\delta\_\{v\_\{l\}\},\\rho\_\{v\_\{l\}\}\\rangle, whereδvl\\delta\_\{v\_\{l\}\}is an LLM andρvl\\rho\_\{v\_\{l\}\}is its prompt\.
###### Definition 3\(Tool\-Calling Node\)\.
A tool\-calling nodevt∈𝒱v\_\{t\}\\in\\mathcal\{V\}, which invokes real tools, is a tuple⟨t,ϑt⟩\\langle t,\\vartheta\_\{t\}\\rangle, wheret∈𝒯t\\in\\mathcal\{T\}is the tool andϑt\\vartheta\_\{t\}is the arguments for tool invocation\.
###### Definition 4\(Dependency Edge\)\.
An edgee⟨vi,vj⟩∈ℰe\_\{\\langle v^\{i\},v^\{j\}\\rangle\}\\in\\mathcal\{E\}is a tuple⟨vi,vj,μ⟩\\langle v^\{i\},v^\{j\},\\mu\\rangle, denoting a directed dependency from nodeviv^\{i\}to nodevjv^\{j\}, whereμ\\muspecifies the dependency semantics,i\.e\.,data dependency or control dependency\.
Workflow execution starts from an LLM node for user query parsing, and proceeds along dependency edges in the graph until a final response is produced by another LLM node forresponse synthesis\.During execution, LLM nodes reason over available context and intermediate results, whereas tool\-calling nodes provide external capabilities through real tool invocations\.
### II\-BWorkflow Generation as Graph Search
The input consists of a set of available tools𝒯\\mathcal\{T\}and a collection of historical task\-solving records𝒟=\{\(qi,ri,πi\)\}i=1N\\mathcal\{D\}=\\\{\(q\_\{i\},r\_\{i\},\\pi\_\{i\}\)\\\}\_\{i=1\}^\{N\}, whereNNis the number of records,qiq\_\{i\}is a user query,rir\_\{i\}is the reference execution result, andπi\\pi\_\{i\}is the validated tool orchestration log\. These records support both the automatic construction and evaluation of tool\-integrated agentic workflows\.
Based on the definition in Sec\.[II\-A](https://arxiv.org/html/2608.10039#S2.SS1), we define the search space of candidate workflows over available tools𝒯\\mathcal\{T\}by Eq\.[1](https://arxiv.org/html/2608.10039#S2.E1),
Ω\(𝒯\)=\{𝒢=\(𝒱,ℰ\)∣Valid\(𝒢,𝒯\)=True\}\\small\\Omega\(\\mathcal\{T\}\)=\\left\\\{\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\}\)\\mid\\mathrm\{Valid\}\(\\mathcal\{G\},\\mathcal\{T\}\)=\\text\{True\}\\right\\\}\(1\)whereValid\(𝒢,𝒯\)\\mathrm\{Valid\}\(\\mathcal\{G\},\\mathcal\{T\}\)is a validity predicate indicating whethera candidate workflow𝒢\\mathcal\{G\}satisfies the basic constraints of a tool\-integrated workflow\. Specifically,𝒢\\mathcal\{G\}is valid only if its tool\-calling nodes can invoke real tools in𝒯\\mathcal\{T\}, its dependency edges form an executable graph, and it contains at least one complete execution path from the input query to the final response\.
To estimate the quality of a candidate workflow𝒢∈Ω\(𝒯\)\\mathcal\{G\}\\in\\Omega\(\\mathcal\{T\}\), we execute𝒢\\mathcal\{G\}on each historical queryqiq\_\{i\}in𝒟\\mathcal\{D\}, producing a tool orchestration traceπ^i\\hat\{\\pi\}\_\{i\}and an execution resultr^i\\hat\{r\}\_\{i\}\.Since complex tool\-use tasks require both correct tool selection and execution scheduling\[[38](https://arxiv.org/html/2608.10039#bib.bib28)\], we evaluate candidate workflows at both the tool\-invocation level and the final\-result level\.Specifically, we compareπ^i\\hat\{\\pi\}\_\{i\}with the validated orchestration logπi\\pi\_\{i\}to evaluatetool invocation correctness, and comparer^i\\hat\{r\}\_\{i\}with the reference execution resultrir\_\{i\}to obtainexecution score\.
Tool Invocation Correctness \(TC\)measures whether the agentic workflow invokes the correct tools with correct arguments while avoiding redundant invocations\. Given a historical queryqiq\_\{i\}, let its validated tool orchestration log beπi=\{\(tik,θik\)\}k=1Ki\\pi\_\{i\}=\\\{\(t\_\{i\}^\{k\},\\theta\_\{i\}^\{k\}\)\\\}\_\{k=1\}^\{K\_\{i\}\},whereKiK\_\{i\}is the number of tool invocations, andtikt\_\{i\}^\{k\}andθik\\theta\_\{i\}^\{k\}denote the tool name and arguments of thekk\-th reference invocation\.Similarly, let the tool orchestration trace produced by a candidate workflow𝒢\\mathcal\{G\}beπ^i=\{\(t^ik,θ^ik\)\}k=1K^i\\hat\{\\pi\}\_\{i\}=\\\{\(\\hat\{t\}\_\{i\}^\{k\},\\hat\{\\theta\}\_\{i\}^\{k\}\)\\\}\_\{k=1\}^\{\\hat\{K\}\_\{i\}\}, whereK^i\\hat\{K\}\_\{i\}is the number of tool invocations, andt^ik\\hat\{t\}\_\{i\}^\{k\}andθ^ik\\hat\{\\theta\}\_\{i\}^\{k\}denote the tool name and arguments of thekk\-th invocation\. We define the tool\-invocation correctness of𝒢\\mathcal\{G\}on queryqiq\_\{i\}by Eq\.[2](https://arxiv.org/html/2608.10039#S2.E2),
TCi\(𝒢\)=\|LCSt=t^∧θ=θ^\(πi,π^i\)\|max\(Ki,K^i\)\\small TC\_\{i\}\(\\mathcal\{G\}\)=\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}\\pgfsys@color@gray@stroke\{0\}\\pgfsys@color@gray@fill\{0\}\\frac\{\\left\|\\operatorname\{LCS\}\_\{t=\\hat\{t\}\\land\\theta=\\hat\{\\theta\}\}\(\\pi\_\{i\},\\hat\{\\pi\}\_\{i\}\)\\right\|\}\{\\max\(K\_\{i\},\\hat\{K\}\_\{i\}\)\}\}\(2\)whereLCSt=t^∧θ=θ^\(πi,π^i\)\\operatorname\{LCS\}\_\{t=\\hat\{t\}\\land\\theta=\\hat\{\\theta\}\}\(\\pi\_\{i\},\\hat\{\\pi\}\_\{i\}\)denotes the longest common subsequence between the reference tool invocations and the workflow\-produced tool invocations, where matched invocations must have identical tool names and arguments\. The numerator counts the number of correctly matched tool invocations\.Compared with normalizing byKiK\_\{i\}, the denominatormax\(Ki,K^i\)\\max\(K\_\{i\},\\hat\{K\}\_\{i\}\)penalizes both missing and redundant tool invocations\. The overall tool invocation correctness of𝒢\\mathcal\{G\}on historical records𝒟\\mathcal\{D\}is then defined by Eq\.[3](https://arxiv.org/html/2608.10039#S2.E3),
TC\(𝒢\)=1N∑i=1NTCi\(𝒢\)\\small TC\(\\mathcal\{G\}\)=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathrm\{TC\}\_\{i\}\(\\mathcal\{G\}\)\(3\)whereNNis the number of historical user queries in𝒟\\mathcal\{D\}\.
Execution Score \(ES\)measures whether the final execution result produced by the agentic workflow satisfies the user query and matches the reference execution result\.Since the execution result may be expressed as natural language, structured results, or task\-completion outcomes,rule\-based evaluation is often insufficient\. Therefore, we adopt an LLM\-as\-a\-judge evaluator,i\.e\.,G\-eval\[[21](https://arxiv.org/html/2608.10039#bib.bib32)\], to assess the final execution result\.
Given a historical queryqiq\_\{i\}with reference execution resultrir\_\{i\}andthe final execution resultr^i\\hat\{r\}\_\{i\}produced by candidate workflow𝒢\\mathcal\{G\}, the evaluator scoresr^i\\hat\{r\}\_\{i\}againstrir\_\{i\}from dimensions𝒳=\{correctness,completeness,relevance,clarity\}\\mathcal\{X\}=\\\{\\textit\{correctness\},\\textit\{completeness\},\\textit\{relevance\},\\textit\{clarity\}\\\}\. Letsix∈\[0,1\]s\_\{i\}^\{x\}\\in\[0,1\]denote the score ofr^i\\hat\{r\}\_\{i\}on dimensionx∈𝒳x\\in\\mathcal\{X\}\. We define the execution score of𝒢\\mathcal\{G\}on historical records𝒟\\mathcal\{D\}by Eq\.[4](https://arxiv.org/html/2608.10039#S2.E4),
ES\(𝒢\)=1N∑i=1N1\|𝒳\|∑x∈𝒳six\\small ES\(\\mathcal\{G\}\)=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\frac\{1\}\{\|\\mathcal\{X\}\|\}\\sum\_\{x\\in\\mathcal\{X\}\}s\_\{i\}^\{x\}\(4\)whereNNis the number of historical user queries in𝒟\\mathcal\{D\}\.
The objective of tool\-integrated agentic workflow generation is to identify a valid workflow graph that achieves hightool invocationcorrectness and execution score on historical task\-solving records\. Given the search spaceΩ\(𝒯\)\\Omega\(\\mathcal\{T\}\), we define the optimal tool\-integrated agentic workflow𝒢∗\\mathcal\{G\}^\{\*\}on tools𝒯\\mathcal\{T\}by Eq\.[5](https://arxiv.org/html/2608.10039#S2.E5)\.
𝒢∗=argmax𝒢∈Ω\(𝒯\)\(TC\(𝒢\)\+ES\(𝒢\)\)\\mathcal\{G\}^\{\*\}=\\arg\\max\_\{\\mathcal\{G\}\\in\\Omega\(\\mathcal\{T\}\)\}\\left\(TC\(\\mathcal\{G\}\)\+ES\(\\mathcal\{G\}\)\\right\)\(5\)This formulation casts workflow generation as a graph search problem over the constrained workflow spaceΩ\(𝒯\)\\Omega\(\\mathcal\{T\}\)\.
## IIIMethodology
Figure 1:Approach Overview ofFlowScoutAs formulated in Sec\.[II\-B](https://arxiv.org/html/2608.10039#S2.SS2), tool\-integrated agentic workflow generation can be viewed as a graph search problem over the constrained workflow space\. However, directly searching this space is inefficient because many graph edits may lead to invalid or ineffective workflows\. Therefore,FlowScoutcombines skeleton mining with execution\-guided graph search to generate stable and reusable tool\-integrated agentic workflows\.
Fig\.[1](https://arxiv.org/html/2608.10039#S3.F1)shows the overview ofFlowScout, which consists of three modules,i\.e\.,Workflow Miner,Workflow Executor, andWorkflow Optimizer\. Given historical task\-solving records and available tools,FlowScoutfirst splits the records into optimization and validation sets, mines the frequent tool coordination skeleton from validated orchestration logsin the optimization set, and constructs an initial tool\-integrated candidate workflow by augmenting the mined skeleton with necessary LLM nodes \(see Sec\.[III\-A](https://arxiv.org/html/2608.10039#S3.SS1)\)\. Then,FlowScoutcompiles the candidate workflow into an executable form, and runs it on the optimization set and validation set to collectoptimization feedback and validation feedback\(see Sec\.[III\-B](https://arxiv.org/html/2608.10039#S3.SS2)\)\. Based on thisexecution feedback,FlowScoutperforms Monte Carlo tree search to refine the topology of the agentic workflow through selection, expansion, simulation, and backpropagation \(see Sec\.[III\-C](https://arxiv.org/html/2608.10039#S3.SS3)\)\. After iterative optimization,FlowScoutoutputs the candidate workflow with the best validation results as the final tool\-integrated agentic workflow\.
### III\-AWorkflow Miner
Directly searching for agentic workflows from an empty graph is inefficient, because most random graph edits either produce invalid workflows or fail to capture the common tool invocation structure in the target domain\. Therefore,FlowScoutfirst splits historical task\-solving records, mines the tool coordination skeleton, and initializes a workflow𝒢0\\mathcal\{G\}^\{0\}, providing a structured warm start for subsequent optimization\.
Task\-Solving Record Splitting\.Given the historical task\-solving records𝒟=\{\(qi,ri,πi\)\}i=1N\\mathcal\{D\}=\\\{\(q\_\{i\},r\_\{i\},\\pi\_\{i\}\)\\\}\_\{i=1\}^\{N\},FlowScoutsplits them into an optimization set𝒟opt\\mathcal\{D\}\_\{opt\}and a validation set𝒟val\\mathcal\{D\}\_\{val\}\.𝒟opt\\mathcal\{D\}\_\{opt\}is used to mine the tool coordination skeleton and provide optimization feedback for workflow generation, while𝒟val\\mathcal\{D\}\_\{val\}is reserved for selecting the best candidate workflow\.
Coordination Skeleton Mining\.After obtaining the optimization set𝒟opt\\mathcal\{D\}\_\{opt\}, we mine a process model from validated orchestration logs to obtain a reusable tool coordination skeleton that can serve as the template for initializing the agentic workflow\. For each validated orchestration logπi=\(tik,θik\)k=1Ki\\pi\_\{i\}=\{\(t\_\{i\}^\{k\},\\theta\_\{i\}^\{k\}\)\}\_\{k=1\}^\{K\_\{i\}\}from historical task\-solving recordsin𝒟opt\\mathcal\{D\}\_\{opt\}, we derive the ordered tool sequenceτi=\(ti1,ti2,…,tiKi\)\\tau\_\{i\}=\(t\_\{i\}^\{1\},t\_\{i\}^\{2\},\\ldots,t\_\{i\}^\{K\_\{i\}\}\), which preserves tool dependencies such as execution order and repeated invocations\.
Based on the ordered tool sequences, we construct an event log for process mining\.In the event log,τi\\tau\_\{i\}is treated as a case, and each tool invocationtikt\_\{i\}^\{k\}inτi\\tau\_\{i\}is treated as one activity in the corresponding case\.Then, we apply PM4Py\[[35](https://arxiv.org/html/2608.10039#bib.bib17)\]to discover the most common process model from the constructed event log\. Theminedmodel summarizes how tools are commonly coordinated in most historical executions,i\.e\.,sequential relations, alternative paths, parallel relations, and repeated invocations\.
Then, we convert the discovered process model into a reusable tool coordination skeleton\. Each activity in the mined model is mapped to a generic tool\-calling position, rather than being bound to a concrete tool\. This abstraction indicates that a tool invocation should be performed at this position, while leaving the concrete tool selection and argument generation to the later workflow execution\. Each process relation is mapped to a workflow dependency\. Specifically, sequential relations are converted into ordered dependencies between tool\-calling positions, alternative paths are converted into conditional branches, parallel relations are converted into parallel workflow branches, and repeated invocations are converted into retry loops\.
Agentic Workflow Initialization\.The resulting coordination skeleton captures the common tool\-invocation structure in historical executions, but it only describes the abstract positions of tool invocations and their coordination relations\. To make the skeleton executable for new user queries, we further instantiate it as an initial tool\-integrated agentic workflow\. Specifically, each generic tool\-calling position in the skeleton is instantiated as a tool\-calling node, which will select a concrete tool from the available tool set and generate arguments during execution\. We further insert an LLM node for planning before the tool\-calling structure to parse the user query, and an LLM node after the tool\-calling structure forsynthesizingthe final response to the user\. Fig\.[2](https://arxiv.org/html/2608.10039#S3.F2)shows an example of the workflow mining process\.
Figure 2:Example of Initial Agentic Workflow Mining
### III\-BWorkflow Executor
The workflow executor turns a candidate tool\-integrated agentic workflow into an executable program, and collects execution results\. Since candidate workflows produced by miner or optimizer may be incomplete or ineffective,FlowScoutexecutes them under an observable runtime, and uses the observedfeedbackto guide subsequent workflow optimization\.
Workflow Compilation\.Given a candidate workflow𝒢′=\(𝒱′,ℰ′\)\\mathcal\{G\}^\{\\prime\}=\(\\mathcal\{V\}^\{\\prime\},\\mathcal\{E\}^\{\\prime\}\),FlowScoutcompiles it into an executable Python script𝒢′^\\hat\{\\mathcal\{G\}^\{\\prime\}\}\. Specifically,FlowScouttraverses the workflow graph, and generates the corresponding Python code for nodes𝒱′\\mathcal\{V\}^\{\\prime\}and edgesℰ′\\mathcal\{E\}^\{\\prime\}\. Each LLM nodevl∈𝒱′v\_\{l\}\\in\\mathcal\{V\}^\{\\prime\}is assigned a function that calls the backbone LLMδvl\\delta\_\{v\_\{l\}\}with certain promptsρvl\\rho\_\{v\_\{l\}\}to obtain the response\.Each tool\-calling nodevt∈𝒱′v\_\{t\}\\in\\mathcal\{V\}^\{\\prime\}is compiled into a function that performs tool selection and argument generation when the node is executed\. Specifically, the function selects a concrete toolt∈𝒯t\\in\\mathcal\{T\}based on the intermediate context, and fills the argumentsϑt\\vartheta\_\{t\}to invoke the tool, recording the returned observation\.Dependency edges inℰ′\\mathcal\{E\}^\{\\prime\}are compiled into the data and control flow among these functions\.
During compilation, we instrument eachvlv\_\{l\}andvtv\_\{t\}with observation hooks\.For a queryqiq\_\{i\}, the runtime records the input, output, and possible error message of each executed node as node\-level observations𝒪i\\mathcal\{O\}\_\{i\}\. For an LLM node, the input consists of the intermediate context generated by preceding nodes and the node’s system prompt, while the output is the model response passed to downstream nodes or returned to the user\. For a tool\-calling node, the input consists of the tool\-oriented subtask decomposed by preceding nodes and the set of available tools𝒯\\mathcal\{T\}\. Its output includes the selected tool, the generated arguments, and the tool response passed to downstream nodes\.These observations make candidate execution traceable, allowingFlowScoutto identify failures such as wrong tool selection, incorrect argument generation, skipped branches, failed tool calls, or poor response synthesis\.
Workflow Execution\.After compilation, we execute𝒢′^\\hat\{\\mathcal\{G\}^\{\\prime\}\}on both𝒟opt\\mathcal\{D\}\_\{opt\}and𝒟val\\mathcal\{D\}\_\{val\}\. For𝒟opt\\mathcal\{D\}\_\{opt\}, we collect optimization feedbackℱopt\(𝒢′\)=\{\(π^i,r^i,𝒪i\)∣\(qi,ri,πi\)∈𝒟opt\}\\mathcal\{F\}\_\{opt\}\(\\mathcal\{G\}^\{\\prime\}\)=\\\{\(\\hat\{\\pi\}\_\{i\},\\hat\{r\}\_\{i\},\\mathcal\{O\}\_\{i\}\)\\mid\(q\_\{i\},r\_\{i\},\\pi\_\{i\}\)\\in\\mathcal\{D\}\_\{opt\}\\\}, whereπ^i\\hat\{\\pi\}\_\{i\}is the produced tool orchestration trace,r^i\\hat\{r\}\_\{i\}is the execution result, and𝒪i\\mathcal\{O\}\_\{i\}denotes the node\-level observations collected during execution\. This feedback is passed to the optimizer for graph search\.
For𝒟val\\mathcal\{D\}\_\{val\}, we evaluate the candidate workflow by comparing eachπ^i\\hat\{\\pi\}\_\{i\}with the validated orchestration logπi\\pi\_\{i\}and comparing eachr^i\\hat\{r\}\_\{i\}with the reference resultrir\_\{i\}to obtain the validation feedbackTC\(𝒢′\)𝒟val\\textit\{TC\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{val\}\}andES\(𝒢′\)𝒟val\\textit\{ES\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{val\}\}, respectively\. This feedback is used to guide the final workflow selection\.
### III\-CWorkflow Optimizer
Starting from the initial workflow𝒢0\\mathcal\{G\}^\{0\}, the workflow optimizer searches for a better tool\-integrated agentic workflow through execution\-guided graph search\. We adopt Monte Carlo tree search \(MCTS\)\[[4](https://arxiv.org/html/2608.10039#bib.bib30)\]because it naturally fits execution\-guided workflow optimization, where each iteration expands one candidate graph, evaluates it through real execution, and uses the observed feedback to guide later search\. In the search tree, each nodemmstores a tool\-integrated agentic workflow𝒢m\\mathcal\{G\}^\{m\}, its visit countCmC^\{m\}that records how many times the workflow𝒢m\\mathcal\{G\}^\{m\}has been evaluated or passed through during backpropagation, accumulated rewardWmW^\{m\}that sums the scalar rewards propagated tomm, mean rewardQm=WmCmQ^\{m\}=\\frac\{W^\{m\}\}\{C^\{m\}\}that estimates the historical quality of𝒢m\\mathcal\{G\}^\{m\}, and optimization feedbackℱopt\(𝒢m\)\\mathcal\{F\}\_\{opt\}\(\\mathcal\{G\}^\{m\}\)collected by the workflow executor\. Each tree edge corresponds to applying one graph mutation to the parent agentic workflow\.Our MCTS for workflow optimization consists of four key stages,i\.e\.,selection,expansion,simulation, andbackpropagation,
Selection\.In each iteration,FlowScoutselects an MCTS node to expand from the current search tree\. Starting from the root node corresponding to𝒢0\\mathcal\{G\}^\{0\}, we recursively descend the search tree untilreaching an expandable nodem⋆m^\{\\star\}\. To keep the branching factor manageable, each tree node is allowed to generate at most three children in our implementation\.WhenFlowScoutreaches a nodemmthat has generated fewer than three children, the selection process stops, andmmis selected asm⋆m^\{\\star\}for expansion\.Otherwise,FlowScoutchooses one of its children according to a selection score that balances historical execution quality and exploration\.For each child nodemcm^\{c\}ofmm,the selection scoreScore\(mc\)\\mathrm\{Score\}\(m^\{c\}\)is defined by Eq\.[6](https://arxiv.org/html/2608.10039#S3.E6),
Score\(mc\)=\{\+∞,Cmc=0,Qmc\+λln\(Cm\+1\)Cmc\+1,Cmc\>0\.\\scriptsize\\mathrm\{Score\}\(m^\{c\}\)=\\begin\{cases\}\+\\infty,&C^\{m^\{c\}\}=0,\\\\ Q^\{m^\{c\}\}\+\\lambda\\sqrt\{\\dfrac\{\\ln\(C^\{m\}\+1\)\}\{C^\{m^\{c\}\}\+1\}\},&C^\{m^\{c\}\}\>0\.\\end\{cases\}\(6\)whereQmcQ^\{m^\{c\}\}is the mean reward ofmcm^\{c\},CmC^\{m\}andCmcC^\{m^\{c\}\}denote the visit counts of the parent and child nodes, andλ\\lambdais the current exploration coefficient\. The second term gives higher scores to less visited children, preventing the search from prematurely focusing on a small set of candidates\. The coefficientλ\\lambdais adjusted by a cosine annealing schedule\[[22](https://arxiv.org/html/2608.10039#bib.bib33)\], encouraging broader exploration in early iterations and gradually shifting the search toward candidates with better scalar rewards\.We select the child node with the highest score asm⋆m^\{\\star\}for expansion\.
Expansion\.After selecting an expandable nodem⋆m^\{\\star\}, we expand it by generating a new child workflow from𝒢m⋆\\mathcal\{G\}^\{m^\{\\star\}\}\. The expansion is guided by an LLM\-based optimizer that analyzes the optimization feedbackℱopt\(𝒢m⋆\)\\mathcal\{F\}\_\{opt\}\(\\mathcal\{G\}^\{m^\{\\star\}\}\)collected by the workflow executor\. Recall that each feedback item\(π^i,r^i,𝒪i\)\(\\hat\{\\pi\}\_\{i\},\\hat\{r\}\_\{i\},\\mathcal\{O\}\_\{i\}\)contains the produced orchestration trace, the execution result, and node\-level observations for each queryqi∈𝒟optq\_\{i\}\\in\\mathcal\{D\}\_\{opt\}\. Thus, we can compareπ^i\\hat\{\\pi\}\_\{i\}with the validated orchestration logπi\\pi\_\{i\}to identify tool\-invocation mismatches, comparer^i\\hat\{r\}\_\{i\}withrir\_\{i\}to identify execution failures, and inspect𝒪i\\mathcal\{O\}\_\{i\}to find problematic workflow nodes\.
Based on these feedback signals, the LLM\-based optimizer proposes a workflow refinement plan\. The plan may contain one or multiple graph edits denoted asΔm⋆=\{o1,o2,…,oL\}\\Delta^\{m^\{\\star\}\}=\\\{o\_\{1\},o\_\{2\},\\ldots,o\_\{L\}\\\}, where each operationolo\_\{l\}is a graph edit applied to𝒢m⋆\\mathcal\{G\}^\{m^\{\\star\}\}\. Applying the refinement plan produces a new candidate workflow𝒢′=Δm⋆\(𝒢m⋆\)\\mathcal\{G\}^\{\\prime\}=\\Delta^\{m^\{\\star\}\}\(\\mathcal\{G\}^\{m^\{\\star\}\}\)\. The operation set includes four types of edits\.Node\-level editsadd, remove, or replace LLM nodes and tool\-calling nodes\.Edge\-level editsadd, remove, or reconnect data flow dependencies\.Control\-flow editsintroduce or refine conditional branches, parallel branches, and retry loops\.Prompt\-level editsrefine the system prompts of LLM nodes\.
For example, tool\-invocation mismatches may lead to inserting or refining tool\-calling nodes, while invalid tool arguments may lead to revising the argument generation prompt\. Wrong branch activation may lead to refining conditional branches or retry loops\. When the tool trace is mostly correct but the final execution score remains low, the optimizer may refine the prompt for the LLM node for response synthesis\.These edits can be combined in one refinement plan when the feedback indicates coupled problems\. To reduce the cost and instability of generating prompts for newly added LLM nodes,FlowScoutmaintains a small library of predefined LLM\-node prompt templates\. These templates cover common node roles in agentic workflows, including planning, extracting, reasoning, verifying, and synthesizing\. When a refinement plan inserts or replaces an LLM node,FlowScoutselects the corresponding prompt template according to the intended node role, and fills it with the current workflow context and execution feedback\.All prompt templates for LLM nodes along with the prompts for the LLM\-based optimizer are provided on our replication website\[[8](https://arxiv.org/html/2608.10039#bib.bib1)\]\.
Finally,FlowScoutapplies the refinement plan to𝒢m⋆\\mathcal\{G\}^\{m^\{\\star\}\}, and the resulting candidate workflow𝒢′\\mathcal\{G\}^\{\\prime\}is added as a new childm′m^\{\\prime\}ofm⋆m^\{\\star\}and passed to the workflow executor for execution\.
Simulation\.After expansion, we invoke the workflow executor to run the new candidate workflow𝒢′\\mathcal\{G\}^\{\\prime\}, obtaining the execution feedbackℱopt\(𝒢′\)\\mathcal\{F\}\_\{opt\}\(\\mathcal\{G\}^\{\\prime\}\)\. Since MCTS requires a scalar reward for backpropagation, we calculate the tool invocation correctnessTC\(𝒢′\)𝒟opt\\textit\{TC\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{opt\}\}and the execution scoreES\(𝒢′\)𝒟opt\\textit\{ES\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{opt\}\}on𝒟opt\\mathcal\{D\}\_\{opt\}, and combine them into a scalar reward by Eq\.[7](https://arxiv.org/html/2608.10039#S3.E7)\.
R\(𝒢′\)𝒟opt=12\(TC\(𝒢′\)𝒟opt\+ES\(𝒢′\)𝒟opt\)\\small\\mathrm\{R\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{opt\}\}=\\frac\{1\}\{2\}\\left\(\\textit\{TC\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{opt\}\}\+\\textit\{ES\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{opt\}\}\\right\)\(7\)The rewardR\(𝒢′\)𝒟opt\\mathrm\{R\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{opt\}\}is used for backpropagation, whileℱopt\(𝒢′\)\\mathcal\{F\}\_\{opt\}\(\\mathcal\{G\}^\{\\prime\}\)is cached in the new child node for later expansion\. During optimization,FlowScoutalso records theTC\(𝒢′\)𝒟val\\textit\{TC\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{val\}\}andES\(𝒢′\)𝒟val\\textit\{ES\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{val\}\}, and computes the validation scoreR\(𝒢′\)𝒟val\\mathrm\{R\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{val\}\}in the same form as Eq\.[7](https://arxiv.org/html/2608.10039#S3.E7)\. TheR\(𝒢′\)𝒟val\\mathrm\{R\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{val\}\}is used only for candidate selection and is not used to update MCTS statistics\.
Backpropagation\.After simulation,FlowScoutpropagates the scalar rewardR\(𝒢′\)𝒟opt\\mathrm\{R\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{opt\}\}from the expanded child node back to the root node along the selected path\.For each nodemim^\{i\}on this path, we updateCmiC^\{m^\{i\}\},WmiW^\{m^\{i\}\}, andQmiQ^\{m^\{i\}\}by Eq\.[8](https://arxiv.org/html/2608.10039#S3.E8)\.
Cmi←Cmi\+1,Wmi←Wmi\+R\(𝒢′\)𝒟opt,Qmi←WmiCmi\.\\small C^\{m^\{i\}\}\\leftarrow C^\{m^\{i\}\}\+1,\\\\ W^\{m^\{i\}\}\\leftarrow W^\{m^\{i\}\}\+\\mathrm\{R\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{opt\}\},\\\\ Q^\{m^\{i\}\}\\leftarrow\\frac\{W^\{m^\{i\}\}\}\{C^\{m^\{i\}\}\}\.\(8\)These updated statistics affect the selection in later iterations\.
Input:Initial workflow𝒢0\\mathcal\{G\}^\{0\}, optimization set𝒟opt\\mathcal\{D\}\_\{opt\}, validation set𝒟val\\mathcal\{D\}\_\{val\}, maximum iterationsJJ, patiencePP
Output:Optimized workflow
𝒢∗\\mathcal\{G\}^\{\*\}
1
2Initialize root node of the tree
𝒮\\mathcal\{S\}with
𝒢0\\mathcal\{G\}^\{0\}and
ℱopt\(𝒢0\)\\mathcal\{F\}\_\{opt\}\(\\mathcal\{G\}^\{0\}\);
3
𝒢∗←𝒢0\\mathcal\{G\}^\{\*\}\\leftarrow\\mathcal\{G\}^\{0\},
Rval∗←R\(𝒢0\)𝒟val\\mathrm\{R\}^\{\*\}\_\{val\}\\leftarrow\\mathrm\{R\}\(\\mathcal\{G\}^\{0\}\)\_\{\\mathcal\{D\}\_\{val\}\},
p←0p\\leftarrow 0;
4
5for*j←1j\\leftarrow 1toJJ*do
6
m⋆←Select\(𝒮\)m^\{\\star\}\\leftarrow\\mathrm\{Select\}\(\\mathcal\{S\}\);
7
𝒢′←Expand\(m⋆,𝒮\.m⋆\.ℱopt\(𝒢m⋆\)\)\\mathcal\{G\}^\{\\prime\}\\leftarrow\\mathrm\{Expand\}\(m^\{\\star\},\\mathcal\{S\}\.m^\{\\star\}\.\\mathcal\{F\}\_\{opt\}\(\\mathcal\{G\}^\{m^\{\\star\}\}\)\);
8
9
R\(𝒢′\)𝒟opt\\mathrm\{R\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{opt\}\},
𝒮\.m′\.ℱopt\(𝒢′\)\\mathcal\{S\}\.m^\{\\prime\}\.\\mathcal\{F\}\_\{opt\}\(\\mathcal\{G\}^\{\\prime\}\)←Simulate\(𝒢′\\leftarrow\\mathrm\{Simulate\}\(\\mathcal\{G\}^\{\\prime\},
𝒟opt\)\\mathcal\{D\}\_\{opt\}\);
10
R\(𝒢′\)𝒟val←Simulate\(𝒢′\\mathrm\{R\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{val\}\}\\leftarrow\\mathrm\{Simulate\}\(\\mathcal\{G\}^\{\\prime\},
𝒟val\)\\mathcal\{D\}\_\{val\}\);
11
12
Backpropagate\(𝒮\.m′,R\(𝒢′\)𝒟opt\)\\mathrm\{Backpropagate\}\(\\mathcal\{S\}\.m^\{\\prime\},\\mathrm\{R\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{opt\}\}\);
13
14if*R\(𝒢′\)𝒟val\>Rval∗\\mathrm\{R\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{val\}\}\>\\mathrm\{R\}^\{\*\}\_\{val\}*then
15
𝒢∗←𝒢′\\mathcal\{G\}^\{\*\}\\leftarrow\\mathcal\{G\}^\{\\prime\};
16
Rval∗←R\(𝒢′\)𝒟val\\mathrm\{R\}^\{\*\}\_\{val\}\\leftarrow\\mathrm\{R\}\(\\mathcal\{G\}^\{\\prime\}\)\_\{\\mathcal\{D\}\_\{val\}\};
17
p←0p\\leftarrow 0;
18
19end if
20else
21
p←p\+1p\\leftarrow p\+1;
22
23end if
24
25if*p≥Pp\\geq P*then
26break;
27
28end if
29
30end for
31
32return*𝒢∗\\mathcal\{G\}^\{\*\}*;
Algorithm 1FlowScoutWorkflow OptimizerOptimization Procedure\.Algorithm[1](https://arxiv.org/html/2608.10039#algorithm1)summarizes the optimization procedure\.FlowScoutfirst initializes the search tree𝒮\\mathcal\{S\}with the initial workflow𝒢0\\mathcal\{G\}^\{0\}and its execution feedbackℱopt\(𝒢0\)\\mathcal\{F\}\_\{opt\}\(\\mathcal\{G\}^\{0\}\), and sets the initial workflow as the current bestcandidate workflow\(Lines[1](https://arxiv.org/html/2608.10039#algorithm1)\-[1](https://arxiv.org/html/2608.10039#algorithm1)\)\. In each iteration,FlowScoutselects one node from the current search tree, and expands it into a new child workflow using the cached execution feedback \(Lines[1](https://arxiv.org/html/2608.10039#algorithm1)\-[1](https://arxiv.org/html/2608.10039#algorithm1)\)\. The new workflow is then simulated on the optimization set to obtain the scalar reward for MCTS update and the execution feedback for later expansion, and is also evaluated on the validation set to obtain its validation score \(Lines[1](https://arxiv.org/html/2608.10039#algorithm1)\-[1](https://arxiv.org/html/2608.10039#algorithm1)\)\. The scalar reward is propagated along the selected path to update the MCTS statistics \(Line[1](https://arxiv.org/html/2608.10039#algorithm1)\)\. If the validation score of the new workflow is higher than the best score observed so far,FlowScoutupdates the best candidate and resets the early\-stopping counterpp\(Lines[1](https://arxiv.org/html/2608.10039#algorithm1)\-[1](https://arxiv.org/html/2608.10039#algorithm1)\); otherwise, the counter is increased \(Line[1](https://arxiv.org/html/2608.10039#algorithm1)\)\. The search terminates when the counter reaches the patience thresholdPP\(Line[1](https://arxiv.org/html/2608.10039#algorithm1)\) or the maximum number of iterationsJJis reached, and finally returns the best validation candidate as𝒢∗\\mathcal\{G\}^\{\*\}\(Line[1](https://arxiv.org/html/2608.10039#algorithm1)\)\. Following\[[44](https://arxiv.org/html/2608.10039#bib.bib50)\], we set the maximum number of iterations to 30\. To reduce unnecessary search cost, we further apply early stopping when the validation score does not improve for 5 iterations\.
## IVEvaluation
We implement a prototype ofFlowScoutwith7,768lines of Python code\. To evaluate the effectiveness and efficiency ofFlowScout, we design the following four research questions\.
- •RQ1 Effectiveness Evaluation\.How effective isFlowScoutin generating tool\-integrated agentic workflows?
- •RQ2 Efficiency Evaluation\.How efficient isFlowScoutin generating tool\-integrated agentic workflows?
- •RQ3 Ablation Study\.How do the workflow miner and MCTS contribute to the final effectiveness and efficiency?
- •RQ4 Generalization Evaluation\.Can the tool\-integrated agentic workflows generated byFlowScoutgeneralize when deployed at runtime with tools and LLM backbones different from those used during workflow generation?
### IV\-AEvaluation Setup
Dataset\.We select four task domains,i\.e\.,finance, sports, travel, and weather from ToolBench\[[28](https://arxiv.org/html/2608.10039#bib.bib27)\], which is a widely used benchmark including user queries, available tools \(i\.e\.,remote APIs\), reference execution results, and validated tool orchestration logs for task completion\. We first filter out the tools yielding empty responses due to unavailable services\. Then, for each task domain, we divide the remaining tools into two disjoint sets, denoted asSeen\-ToolsandUnseen\-Tools, where each seen tool is matched with an unseen tool that provides similar functionality\. The seen tools are used for agentic workflow generation, while the unseen tools are reserved for evaluating whether the generated workflows can generalize to functionally similar but previously unseen tools\. Based on the validated tool orchestration logs, we associate each tool with the user queries that invoke it during task completion\. For the seen tools, we collect the associated user queries along with their reference execution results and validated tool orchestration logs in each task domain, and split them into optimization, validation, and test sets with a ratio of 5:2:3, denoted as𝒬opt\\mathcal\{Q\}\_\{opt\},𝒬val\\mathcal\{Q\}\_\{val\}, and𝒬test\\mathcal\{Q\}\_\{test\}\. The optimization set is used to guide workflow generation, the validation set is used to evaluate candidate workflows during the search, and the test set is held out for the final effectiveness evaluation\. For the unseen tools, we collect the same number of associated user queries as𝒬gene\\mathcal\{Q\}\_\{gene\}along with reference execution results separately, and use them only for the generalization evaluation, so that the seen and unseen test sets are of equal size for a fair comparison\. The final dataset statistics are shown in Table[I](https://arxiv.org/html/2608.10039#S4.T1)\.
TABLE I:Statistics of the Evaluation DatasetDomainSeen\-ToolsUnseen\-Tools\#Tools\#𝒬opt\\mathcal\{Q\}\_\{opt\}\#𝒬val\\mathcal\{Q\}\_\{val\}\#𝒬test\\mathcal\{Q\}\_\{test\}\#Tools\#𝒬gene\\mathcal\{Q\}\_\{gene\}Finance341467187280341280Sports288253101152288152Travel10121787130101130Weather488333504850
Baselines\.We compareFlowScoutwith three representative baselines,i\.e\.,PM4Py\[[35](https://arxiv.org/html/2608.10039#bib.bib17)\], ReAct\[[42](https://arxiv.org/html/2608.10039#bib.bib19)\], and AFlow\[[44](https://arxiv.org/html/2608.10039#bib.bib50)\]\. PM4Py is a representative workflow mining framework that discovers process models from event logs\. We include PM4Py to examine whether traditional workflow mining techniques can derive effective workflows from validated tool orchestration logs\. Specifically, we adapt each tool invocation in the logs as an activity and use the discovered process model as the workflow structure for task execution\. ReAct is a widely used LLM agent paradigm that interleaves reasoning and acting\. It serves as a standard agent baseline that solves tool\-use tasks without explicitly constructing reusable workflow structures\. AFlow is a recent automatic agentic workflow generation approach that searches for workflows from historical task\-solving records\. We include AFlow as the most related baseline, since it also automates workflow generation but mainly constructs LLM\-centric workflows\. For fair comparison, all approaches are evaluated on the same task domains, available tools, and test queries\. Approaches requiring historical records are provided with the same seen\-tool data, including user queries, reference execution results, and validated tool orchestration logs\.Both AFlow andFlowScoutuse GPT\-4o\[[26](https://arxiv.org/html/2608.10039#bib.bib34)\]as the optimization LLM for workflow construction\. The workflows generated by AFlow andFlowScout, together with the ReAct agent, use GPT\-4o\-mini\[[25](https://arxiv.org/html/2608.10039#bib.bib35)\]as the execution LLM backbone and have access to the same tool documentation\.For all baselines, we only align the input and output interfaces of their official replication artifacts to fit our evaluation dataset\.
RQ Setup\.ForRQ1, we use the optimization set𝒬opt\\mathcal\{Q\}\_\{opt\}and validation set𝒬val\\mathcal\{Q\}\_\{val\}to generate workflows with PM4Py, AFlow, andFlowScout\. We then evaluate the generated workflows, together with the ReAct agent, on the test set𝒬test\\mathcal\{Q\}\_\{test\}\. We report the averageTCandESacross different task domains, and compareFlowScoutwith the baselines\. We further run the workflows generated by AFlow andFlowScoutfor 10 times, as well as the ReAct agent, on 50 queries selected randomly in𝒬test\\mathcal\{Q\}\_\{test\}for each task domain, and compare the coefficient of variation ofESto evaluate their execution stability\.
ForRQ2, we evaluate the efficiency ofFlowScoutfrom both offline workflow generation and runtime execution\. For offline generation, we plot the validation score of candidate workflows in 30 iterations to show how the quality of workflows evolves during the search process\. For runtime execution, we measure the average execution time of the workflow generated byFlowScouton𝒬test\\mathcal\{Q\}\_\{test\}, and compare it with that of the ReAct agent and the workflow generated by AFlow\.
ForRQ3, we conduct ablation studies on the finance domain to examine the contribution of the workflow miner and the MCTS\. We construct three variants ofFlowScout, denoted asFlowScout\-NoMCTS,FlowScout\-NoMiner, andFlowScout\-Beam, respectively\.FlowScout\-NoMCTSremoves the MCTS\-based workflow optimizer, and directly uses the workflow mined from the historical task\-solving records without further optimization\.FlowScout\-NoMinerremoves the workflow miner fromFlowScout, and searches for workflows from an empty skeleton using MCTS\-based workflow optimizer\.FlowScout\-Beamreplaces MCTS with beam search\[[29](https://arxiv.org/html/2608.10039#bib.bib38)\], which keeps the top\-3 workflow candidates at each search step\. For fair comparison,FlowScout\-NoMiner,FlowScout\-BeamandFlowScoutare allowed to generate 30 candidate workflows in total\. We compare theTCandESof their best candidate workflows to evaluate effectiveness, and report the total generation time cost to evaluate efficiency\.
ForRQ4, we evaluate the workflow generated byFlowScoutfrom two aspects,i\.e\.,tool\-set replacement and LLM\-backbone replacement, on the finance domain compared with the ReAct agent and the workflow generated by AFlow\. For tool\-set generalization, the tool\-integrated agentic workflow is generated usingSeen\-Tools, and then evaluated after replacing the available tools withUnseen\-Tools\. For LLM backbone generalization, we keep the generated workflow topology unchanged, and replace the execution LLM backbone with DeepSeek\-V3\.2\[[20](https://arxiv.org/html/2608.10039#bib.bib36)\]and Qwen3\-32B\[[40](https://arxiv.org/html/2608.10039#bib.bib37)\], respectively\. We reportTCandESto measure whether each approach can still support effective tool invocation and task completion\.
Environment\.We conduct all the experiments on Ubuntu 20\.04\.4 LTS servers with 4 NVIDIA GeForce RTX 3090 GPUs, Intel\(R\) Xeon\(R\) Silver 4310 @ 2\.10GHz and 128GB memory\.
### IV\-BEffectiveness Evaluation \(RQ1\)
TABLE II:Results of the General EffectivenessApproachMetricTask DomainFinanceSportsTravelWeatherAvg\.PM4PyTC0\.05500\.03530\.03580\.29260\.1047ES0\.15300\.17150\.06200\.26800\.1636ReActTC0\.51150\.35010\.50460\.19390\.3900ES0\.61130\.64180\.61840\.63760\.6273AFlowTC–––––ES0\.61820\.60070\.68480\.56510\.6172FlowScoutTC0\.65830\.71600\.80280\.82890\.7515ES0\.73490\.69440\.75280\.77010\.7381
Overall Effectiveness\.Table[II](https://arxiv.org/html/2608.10039#S4.T2)reports the effectiveness results ofFlowScoutand the baselines on the test set\. We do not reportTCfor AFlow because its generated workflow only contains LLM nodes and does not produce real\-tool invocations\. Overall,FlowScoutconsistently achieves the best performance across all four domains in terms of both tool\-invocation correctness \(TC\) and execution score \(ES\)\.
PM4Py performs markedly worse than all LLM\-based approaches\. Although process mining can extract tool orchestration patterns from historical task\-solving records, directly translating these patterns into executable workflows provides limited support for query understanding, tool selection, argument construction, and result synthesis\. For fair comparison, we mainly compareFlowScoutwith LLM\-based approaches\.
With respect toTC,FlowScoutimproves the averageTCover ReAct by92\.69%\. The improvement is observed in every domain, ranging from28\.70%in the finance domain to327\.49%in the weather domain\. In particular,FlowScoutmore than doubles theTCof ReAct in the sports domain and achieves over four times itsTCin the weather domain, demonstrating that the generated tool\-integrated agentic workflows substantially improve the accuracy of tool selection and orchestration\.
With respect toES,FlowScoutimproves the averageESby17\.66%over ReAct and19\.59%over AFlow\.FlowScoutalso consistently surpasses the best\-performing baseline in each domain, with relative improvements of18\.88%,8\.20%,9\.93%, and20\.78%in the finance, sports, travel, and weather domains, respectively\.FlowScoutoutperforms both baselines in all cases, indicating that it can generate workflows that are both more effective and less sensitive to domain\-specific tasks\.
Overall, these results demonstrate thatFlowScoutgenerates workflows with more accurate tool invocation and better end\-to\-end task\-solving performance than both general\-purpose LLM agents and existing workflow\-generation approaches\.
Figure 3:Results of the Execution StabilityStability Analysis\.We further evaluate execution stability by repeatedly running each approach 10 times on the same test queries and computing the coefficient of variation \(CV\) ofESfor each query\. A lower CV indicates that the approach has more consistent execution quality across repeated runs\. As shown in Fig\.[3](https://arxiv.org/html/2608.10039#S4.F3),FlowScoutachieves the lowest average CV across domains\. Overall, it reduces the average CV by62\.09%and27\.90%compared with ReAct and AFlow, respectively\.
Compared with ReAct, the improvement is particularly substantial in the finance and sports domains, whereFlowScoutreduces the average CV by68\.14%and68\.03%, respectively\. These domains involve relatively large tool spaces and more possible tool combinations, making unconstrained planning more susceptible to different tool selections and invocation orders across runs\. This consistent reduction in CV across all domains indicates that explicitly structuring workflow benefits in complex scenarios\. Compared with AFlow,FlowScoutreduces the average CV by20\.75%,33\.22%,8\.31%, and40\.04%in the finance, sports, travel, and weather domains, respectively\. The box\-plot distributions show thatFlowScoutgenerally exhibits narrower interquartile ranges and fewer high\-variation cases\. Although AFlow generates reusable workflow structures, its workflows are primarily composed of LLM\-centric operators without actual tool invocations\. In contrast,FlowScoutintegrates tool\-calling nodes and their dependencies explicitly, and refines them using execution feedback\. This reduces the number of runtime decisions left to unstable LLM planning which may lead to hallucinations and consequently improves the reproducibility of execution outcomes\.
Breakdown Analysis\.The observed ineffectiveness and instability can be attributed to uncertainty introduced at three stages of execution,i\.e\.,LLM planning, tool invocation, and response synthesis\. For ReAct, all three stages are dynamically determined at runtime\. Variations in LLM generation may lead the agent to decompose the same query differently, select different tools, alter the invocation order, or terminate at different points, producing substantially different final execution scores\. AFlow reduces part of this variation by providing reusable agentic workflows\. However, because tool invocation is not represented and the agentic workflows generated by AFlow rely heavily on LLM nodes, AFlow is generally more stable than ReAct but still exhibits noticeable variation\.
FlowScoutreduces both ineffectiveness and instability by constructing the workflow topology and explicitly modeling tool invocations\. However, the generated tool\-integrated agentic workflows are not fully deterministic or universally optimal\. Their effectiveness remains limited by insufficient coverage of real world queries, and errors in LLM\-based query interpretation, argument generation, and response synthesis\. Meanwhile, stochastic decisions within LLM nodes, varying intermediate outputs, and occasional tool failures can still lead to inconsistent execution results across repeated runs\.
Figure 4:Examples of the Workflow Search Process and the Tool\-Integrated Agentic Workflow in the Travel DomainCase Study\.Fig\.[4](https://arxiv.org/html/2608.10039#S4.F4)illustrates an abstract workflow search process in the travel domain and the execution of the resulting tool\-integrated workflow\. The left part shows the selected path in the MCTS search tree, where each node represents a candidate workflow and each edge represents a workflow mutation\. Through operations such as adding nodes, inserting conditional branches, and modifying prompts,FlowScoutprogressively improves the validation score from0\.3420to0\.7164\.
The right part shows the execution of the tool\-integrated agentic workflow on a query that requests both the distance from Birmingham to Sacramento and hotel recommendations in Sacramento\. The workflow first decomposes the query into distance calculation and hotel recommendation subtasks\. It then invokes\[GreatCircleDistance\-KM\]and\[GreatCircleDistance\-Mile\]to obtain the distance in kilometers and miles, respectively\. When the intermediate context is found to omit the requested hotel information by the reasoning node, the judging node activates the optional branch and invokes\[HotelSearch\]to retrieve hotels in Sacramento\. Finally, the synthesizing node integrates the outputs of the three tools into a final response\.
Summary\.FlowScoutimproves the tool invocation correctness by at least92\.69%and execution score by at least17\.66%with better stability, compared with baselines\.
### IV\-CEfficiency Evaluation \(RQ2\)
Figure 5:Validation Scores during the Search ProcessOffline Workflow Generation\.Fig\.[5](https://arxiv.org/html/2608.10039#S4.F5)shows the validation scores of candidate workflows over 30 optimization iterations\. Across all four domains,FlowScoutfinds its best workflow within 20 iterations, indicating that execution\-guided search can obtain effective workflows under a bounded optimization budget\. Compared with the initial workflows, the best candidates improve the validation scores by92\.20%,98\.13%,109\.47%, and105\.58%in the finance, sports, travel, and weather domains, respectively, demonstrating the benefit of iterative workflow optimization\. Besides, these best candidates appear at iterations 20, 17, 14, and 12, respectively, where the ordering is consistent with the tool\-space sizes of these four task domains\. This result suggests that larger tool spaces may require more iterations to explore the optimal tool\-integrated agentic workflow\.
Runtime Execution\.At runtime, the agentic workflow generated byFlowScoutrequires 22\.80 seconds per query, on average, compared with 11\.44 seconds for the ReAct agent and 18\.37 seconds for the agentic workflow generated by AFlow\. The overhead is more pronounced relative to the ReAct agent, which performs planning and tool invocation within a comparatively compact interaction loop, whileFlowScoutmay involve multiple specialized LLM nodes for planning, reasoning, judging, and synthesizing intermediate results\. These nodes reduce unconstrained runtime decisions, but also introduce additional runtime cost\. In contrast, the smaller gap betweenFlowScoutand AFlow indicates that part of the runtime cost is common to workflow\-based execution, whileFlowScoutintroduces time overhead for tool invocation compared with AFlow\. Considering thatFlowScoutimproves both general effectiveness and execution stability over the baselines, the observed runtime overhead represents a trade\-off between execution efficiency and workflow reliability\.
Summary\.FlowScoutfinds effective workflows within a limited offline search budget, with at least24\.12%higher runtime cost than baselines for improved effectiveness\.
### IV\-DAblation Study \(RQ3\)
TABLE III:Results of the Ablation StudyMinerMCTSBeamTCESTime \(h\)FlowScout\-NoMCTS✓0\.45980\.21180\.01FlowScout\-NoMiner✓0\.49850\.562110\.9FlowScout\-Beam✓✓0\.67680\.532738\.88FlowScout✓✓0\.65830\.734915\.39
Table[III](https://arxiv.org/html/2608.10039#S4.T3)reports the ablation results on the finance domain\. Among all variants,FlowScout\-NoMCTSshows the largest performance degradation compared withFlowScout, reducingTCandESby30\.15%and71\.18%, respectively\. This indicates that the mined workflow provides a useful tool coordination skeleton, but still requires further optimization to become an effective tool\-integrated agentic workflow\.
Removing the workflow miner also reduces performance\. Compared withFlowScout,FlowScout\-NoMinerreducesTCby24\.27%andESby23\.51%\. Without the mined skeleton, MCTS cannot discover an optimal candidate workflow within the same candidate budget\. The miner therefore provides a warm start that improves both search quality and efficiency\.
Replacing MCTS with beam search yields a2\.81%higherTC, but itsESis27\.51%lower than that ofFlowScout\. Beam search retains only the top\-performing candidates at each step and may prematurely discard workflows that can lead to better optimization results\. In contrast, MCTS models each candidate agentic workflow as a node in the search tree and explores the entire search tree for better optimization results\.
Regarding the time cost,FlowScout\-NoMCTSis the fastest because it performs no iterative optimization, but its effectiveness is substantially lower\. Under the same candidate budget,FlowScouttakes15\.39hours, which is41\.19%more thanFlowScout\-NoMinerand60\.42%less thanFlowScout\-Beam\. The high cost of beam search arises from maintaining and expanding multiple candidates at each search step\. These results show that both workflow mining and MCTS contribute to the final effectiveness and efficiency\.
Summary\.The workflow miner provides an effective initialization, while MCTS further improves TC and ES over the mined\-only workflow by 43\.17% and 246\.98%\. Compared with beam search, MCTS achieves substantially higher agentic workflow quality with approximately half the time for search\.
### IV\-EGeneralization Evaluation \(RQ4\)
TABLE IV:Results of Tool GeneralizationMetricReActAFlowFlowScoutSeen\-ToolsTC0\.5115–0\.6583ES0\.61130\.61820\.7349Unseen\-ToolsTC0\.4574–0\.5660ES0\.60500\.58740\.6751Tool\-Set Generalization\.Table[IV](https://arxiv.org/html/2608.10039#S4.T4)reports the results after replacing the tools used during workflow generation withUnseen\-Tools\. Although all approaches show some degradation,FlowScoutremains the best\-performing approach on both metrics\. Compared with ReAct,FlowScoutimprovesTCby23\.74%andESby11\.59%under the unseen tool set\. It also outperforms AFlow by14\.93%inES\.
After tool\-set replacement,FlowScoutretains85\.98%of its originalTCand91\.86%of its originalES\. In particular, the relatively small decrease inESshows that the generated agentic workflow can preserve most of its task\-solving capability when deployed with different available tools\. Together with its continued advantage over the baselines, this result demonstrates that the generated agentic workflow is not restricted to the tools used during workflow generation\. This generalization mainly stems from the separation between workflow\-level orchestration and runtime tool binding\.FlowScoutpreserves reusable task decomposition, control\-flow, and dependency structures, while its tool\-calling nodes dynamically select tools and construct arguments according to the tools available at runtime, rather than binding the workflow to specific tools\.
TABLE V:Results of LLM GeneralizationMetricReActAFlowFlowScoutGPT\-4o\-miniTC0\.5115–0\.6583ES0\.61130\.61820\.7349DeepSeek\-V3\.2TC0\.5463–0\.5548ES0\.56760\.51580\.7221Qwen3\-32BTC0\.3908–0\.5885ES0\.39020\.46380\.7239LLM\-Backbone Generalization\.Table[V](https://arxiv.org/html/2608.10039#S4.T5)reports the results after replacing GPT\-4o\-mini with DeepSeek\-V3\.2 and Qwen3\-32B while keeping the generated workflow topology unchanged\.FlowScoutconsistently achieves the highestTCandESacross both replacement LLMs\. With DeepSeek\-V3\.2,FlowScoutoutperforms ReAct by1\.56%inTCand27\.22%inES, and exceeds AFlow by40\.00%inES\. With Qwen3\-32B, the improvements over ReAct reach50\.59%and85\.52%, respectively, while itsESis56\.08%higher than that of AFlow\.
After replacing the execution LLM,FlowScoutretains84\.28%and89\.39%of its originalTCwith DeepSeek\-V3\.2 and Qwen3\-32B, respectively\. More importantly, it retains98\.26%and98\.50%of its originalES, showing that the generated workflow preserves nearly all of its capability across different LLM backbones\. This is because our generated workflow explicitly specifies the roles of different LLM nodes, tool\-calling positions, and their dependencies, while the execution LLM is only responsible for completing the local function of each node\. Therefore, replacing the LLM backbone may affect individual decisions to some extent, but the overall task\-solving capability encoded by the workflow remains reusable\.
Summary\.The workflows generated byFlowScoutretain91\.37%of their originalTCandESon average when deployed with different tools and LLM backbones\.
## VThreats to Validity
First, the selection of evaluation metrics poses a threat to validity\.TCmay penalize alternative but valid tool invocations, whileESmay be affected by the bias of the LLM\-based evaluator\. To mitigate this threat, we evaluate all approaches using the same reference execution results, employing a widely used LLM\-based evaluator with the same prompts, and jointly report both tool\-level and result\-level metrics\.
Second, the selection and configuration of baselines pose a threat to validity, since the compared approaches adopt different workflow abstractions\. To reduce this threat, we use their official replication artifacts, modify only the input and output interfaces, and align the available tools, tool documentation, execution LLM, and optimization budget whenever applicable\.
Third, domain coverage, workflow complexity, and historical\-record quality pose threats to validity\. The four evaluated domains differ in tool\-set size and query distribution, but mainly contain tool\-oriented tasks with relatively short tool invocation chains with at most 6 tool invocations\. Thus, the findings may not fully generalize to workflows with longer tool invocation chains\. Noisy or incomplete records may affect both the quality of workflow mining and optimization\. We mitigate this threat by conducting experiments on manually validated records\.
Finally, randomness in LLM generation, workflow search, and LLM\-based evaluation poses a threat to result reliability\. To mitigate this threat, we independently generate and evaluate workflows in four task domains with different tool sets and query distributions, whereFlowScoutconsistently outperforms the baselines\. Query\-level paired tests further show that these improvements are statistically significant across domains after Holm correction\[[11](https://arxiv.org/html/2608.10039#bib.bib68)\]\(padj<0\.05\\text\{p\}\_\{\\mathrm\{adj\}\}<0\.05\), with 95% bootstrap confidence intervals excluding zero and Cohen’s d\>0\.63\>0\.63\[[5](https://arxiv.org/html/2608.10039#bib.bib69)\]\. We also repeat execution 10 times on 50 randomly selected queries per domain, and report the coefficient of variation ofES\. These cross\-domain, statistical, and repeated\-execution results reduce the likelihood that the observed improvements are caused by a particular domain or an individual execution run\.
## VIRelated Work
Recent studies and platforms\[[15](https://arxiv.org/html/2608.10039#bib.bib41),[16](https://arxiv.org/html/2608.10039#bib.bib42),[36](https://arxiv.org/html/2608.10039#bib.bib43),[17](https://arxiv.org/html/2608.10039#bib.bib44),[6](https://arxiv.org/html/2608.10039#bib.bib45)\]have facilitated the construction of agentic workflows\. Developer\-oriented frameworks such as LangChain\[[15](https://arxiv.org/html/2608.10039#bib.bib41)\], LangGraph\[[16](https://arxiv.org/html/2608.10039#bib.bib42)\], and AutoGen\[[36](https://arxiv.org/html/2608.10039#bib.bib43)\]provide programming abstractions for composing LLMs, tools, memory modules and control logic\. These frameworks make it easier for developers to implement agentic workflows, but they still require manual programming and careful design of workflow structures\. Low\-code platforms such as Dify\[[17](https://arxiv.org/html/2608.10039#bib.bib44)\]and Coze\[[6](https://arxiv.org/html/2608.10039#bib.bib45)\]offer visual interfaces for connecting workflow components\. Although these platforms reduce the engineering effort of building agentic workflows, they still rely on users to manually specify workflow structures, tool connections, and control logic based on domain knowledge\.
To further reduce the manual effort, some studies have started to explore automatic generation of LLM agents\[[12](https://arxiv.org/html/2608.10039#bib.bib48),[31](https://arxiv.org/html/2608.10039#bib.bib49)\]or agentic workflows\[[18](https://arxiv.org/html/2608.10039#bib.bib46),[44](https://arxiv.org/html/2608.10039#bib.bib50),[47](https://arxiv.org/html/2608.10039#bib.bib51)\]\. Hu et al\.\[[12](https://arxiv.org/html/2608.10039#bib.bib48)\]and Shang et al\.\[[31](https://arxiv.org/html/2608.10039#bib.bib49)\]search over modular agent designs by composing LLMs, memory modules, external tools and prompts\. AutoFlow\[[18](https://arxiv.org/html/2608.10039#bib.bib46)\]parses user queries and uses LLMs to decompose them into executable workflows at the query level, relying on the task decomposition ability of LLMs for each individual query\. The most related works to ours are AFlow\[[44](https://arxiv.org/html/2608.10039#bib.bib50)\]and A2Flow\[[47](https://arxiv.org/html/2608.10039#bib.bib51)\], which also use historical task\-solving records, including user queries and execution results, to search for agentic workflows\. However, their generated agentic workflows only consist of LLM nodes, where LLMs act as general\-purpose modules for solving different subtasks\. In contrast,FlowScoutmodels real tool invocations as workflow nodes, and constructs the topology between LLM reasoning and external tools\. Guided by execution feedback on historical task\-solving records, it uses Monte Carlo tree search to generate agentic workflows that are more usable and stable\.We do not compareFlowScoutwith A2Flow, as it follows the general idea of AFlow and is not publicly available\.
## VIIConclusion
We have proposed and implementedFlowScout, a framework for automatic domain\-specific agentic workflow generation\.FlowScoutformulates the agentic workflow as a structured graph composed of LLM nodes, tool\-calling nodes and dependency edges that encode control and data dependencies, and searches for the optimal workflow topology guided by historical task\-solving records\. Large\-scale experiments have been conducted to demonstrate its effectiveness and efficiency\.
## VIIIData Availability
The source code and data of our work are available at\[[8](https://arxiv.org/html/2608.10039#bib.bib1)\]\.
## References
- \[1\]Anthropic\(2024\)Building effective agents\.Note:[https://www\.anthropic\.com/engineering/building\-effective\-agents](https://www.anthropic.com/engineering/building-effective-agents)Accessed: 2026\-05\-28Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1)\.
- \[2\]P\. Bertoli, M\. Pistore, and P\. Traverso\(2010\)Automated composition of web services via planning in asynchronous domains\.Artificial Intelligence174\(3\-4\),pp\. 316–361\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p3.1.1)\.
- \[3\]I\. Bouzenia, P\. Devanbu, and M\. Pradel\(2025\)RepairAgent: an autonomous, llm\-based agent for program repair\.InProceedings of the 47th IEEE/ACM International Conference on Software Engineering,pp\. 2188–2200\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1)\.
- \[4\]C\. B\. Browne, E\. Powley, D\. Whitehouse, S\. M\. Lucas, P\. I\. Cowling, P\. Rohlfshagen, S\. Tavener, D\. Perez, S\. Samothrakis, and S\. Colton\(2012\)A survey of monte carlo tree search methods\.IEEE Transactions on Computational Intelligence and AI in games4\(1\),pp\. 1–43\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p5.1.2),[§III\-C](https://arxiv.org/html/2608.10039#S3.SS3.p1.10)\.
- \[5\]J\. Cohen\(1988\)Statistical power analysis for the behavioral sciences\.2 edition,Lawrence Erlbaum Associates\.Cited by:[§V](https://arxiv.org/html/2608.10039#S5.p4.2)\.
- \[6\]Coze\(2025\)Coze studio: an ai agent development platform\.Note:[https://github\.com/coze\-dev/coze\-studio](https://github.com/coze-dev/coze-studio)Accessed: 2026\-05\-19Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p3.1),[§VI](https://arxiv.org/html/2608.10039#S6.p1.1)\.
- \[7\]F\. Friedrich, J\. Mendling, and F\. Puhlmann\(2011\)Process model generation from natural language text\.InProceedings of the 23rd international conference on Advanced information systems engineering,pp\. 482–496\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p3.1.1)\.
- \[8\]A\. Github\(2026\)FlowScout\(Website\)Note:Accessed: 2026\-05\-28External Links:[Link](https://anonymous.4open.science/r/xxxx/)Cited by:[§III\-C](https://arxiv.org/html/2608.10039#S3.SS3.p5.1.4),[§VIII](https://arxiv.org/html/2608.10039#S8.p1.1)\.
- \[9\]Google\(2023\)Google gemini\.Note:[https://gemini\.google\.com/](https://gemini.google.com/)Accessed: 2026\-05\-28Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1)\.
- \[10\]H\. He, W\. Yao, K\. Ma, W\. Yu, Y\. Dai, H\. Zhang, Z\. Lan, and D\. Yu\(2024\)Webvoyager: building an end\-to\-end web agent with large multimodal models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 6864–6890\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p2.1)\.
- \[11\]S\. Holm\(1979\)A simple sequentially rejective multiple test procedure\.Scandinavian Journal of Statistics6\(2\),pp\. 65–70\.Cited by:[§V](https://arxiv.org/html/2608.10039#S5.p4.2)\.
- \[12\]S\. Hu, C\. Lu, and J\. Clune\(2025\)Automated design of agentic systems\.InProceedings of the 13th International Conference on Learning Representations,Cited by:[§VI](https://arxiv.org/html/2608.10039#S6.p2.2)\.
- \[13\]M\. Kim, T\. Stennett, S\. Sinha, and A\. Orso\(2025\)A multi\-agent approach for rest api testing with semantic graphs and llm\-driven inputs\.InProceedings of the 47th IEEE/ACM International Conference on Software Engineering,pp\. 1409–1421\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1)\.
- \[14\]S\. Kim, S\. Moon, R\. Tabrizi, N\. Lee, M\. W\. Mahoney, K\. Keutzer, and A\. Gholami\(2024\)An llm compiler for parallel function calling\.InProceedings of the 41st International Conference on Machine Learning,pp\. 24370–24391\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1)\.
- \[15\]LangChain\(2025\)LangChain: the agent engineering platform\.Note:[https://github\.com/langchain\-ai/langchain](https://github.com/langchain-ai/langchain)Accessed: 2026\-05\-19Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p3.1),[§VI](https://arxiv.org/html/2608.10039#S6.p1.1)\.
- \[16\]LangChain\(2025\)LangGraph: low\-level orchestration framework for building stateful agents\.Note:[https://github\.com/langchain\-ai/langgraph](https://github.com/langchain-ai/langgraph)Accessed: 2026\-05\-19Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p3.1),[§VI](https://arxiv.org/html/2608.10039#S6.p1.1)\.
- \[17\]LangGenius\(2025\)Dify: production\-ready platform for agentic workflow\.Note:[https://github\.com/langgenius/dify](https://github.com/langgenius/dify)Accessed: 2026\-05\-19Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p3.1),[§VI](https://arxiv.org/html/2608.10039#S6.p1.1)\.
- \[18\]Z\. Li, S\. Xu, K\. Mei, W\. Hua, B\. Rama, O\. Raheja, H\. Wang, H\. Zhu, and Y\. Zhang\(2024\)AutoFlow: automated workflow generation for large language model agents\.arXiv preprint arXiv:2407\.12821\.Cited by:[§VI](https://arxiv.org/html/2608.10039#S6.p2.2)\.
- \[19\]F\. Lin, D\. J\. Kim, and T\. Chen\(2025\)Soen\-101: code generation by emulating software process models using large language model agents\.In2025 IEEE/ACM 47th International Conference on Software Engineering,pp\. 1527–1539\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p2.1)\.
- \[20\]A\. Liu, A\. Mei, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong,et al\.\(2025\)Deepseek\-v3\. 2: pushing the frontier of open large language models\.arXiv preprint arXiv:2512\.02556\.External Links:[Link](https://arxiv.org/abs/2512.02556)Cited by:[§IV\-A](https://arxiv.org/html/2608.10039#S4.SS1.p6.2.2)\.
- \[21\]Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. Zhu\(2023\)G\-eval: nlg evaluation using gpt\-4 with better human alignment\.InProceedings of the 2023 conference on empirical methods in natural language processing,pp\. 2511–2522\.Cited by:[§II\-B](https://arxiv.org/html/2608.10039#S2.SS2.p5.1)\.
- \[22\]I\. Loshchilov and F\. HutterSGDR: stochastic gradient descent with warm restarts\.Learning10,pp\. 3\.Cited by:[§III\-C](https://arxiv.org/html/2608.10039#S3.SS3.p2.15)\.
- \[23\]A\. Mastropaolo, F\. Zampetti, G\. Bavota, and M\. Di Penta\(2024\)Toward automatically completing github workflows\.InProceedings of the 46th IEEE/ACM International Conference on Software Engineering,pp\. 1–12\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p3.1.1)\.
- \[24\]OpenAI\(2023\)Introducing chatgpt\.Note:[https://openai\.com/blog/chatgpt](https://openai.com/blog/chatgpt)Accessed: 2026\-05\-28Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1)\.
- \[25\]OpenAI\(2024\)GPT\-4o\-mini\.Note:[hhttps://developers\.openai\.com/api/docs/models/gpt\-4o\-mini](hhttps://developers.openai.com/api/docs/models/gpt-4o-mini)Accessed: 2026\-05\-28Cited by:[§IV\-A](https://arxiv.org/html/2608.10039#S4.SS1.p2.1.4)\.
- \[26\]OpenAI\(2024\)GPT\-4o\.Note:[https://developers\.openai\.com/api/docs/models/gpt\-4o/](https://developers.openai.com/api/docs/models/gpt-4o/)Accessed: 2026\-05\-28Cited by:[§IV\-A](https://arxiv.org/html/2608.10039#S4.SS1.p2.1.4)\.
- \[27\]OpenAI\(2025\)Introducing deep research\.Note:[https://openai\.com/index/introducing\-deep\-research/](https://openai.com/index/introducing-deep-research/)Accessed: 2026\-06\-03Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p2.1)\.
- \[28\]Y\. Qin, S\. Liang, Y\. Ye, K\. Zhu, L\. Yan, Y\. Lu, Y\. Lin, X\. Cong, X\. Tang, B\. Qian,et al\.\(2024\)ToolLLM: facilitating large language models to master 16000\+ real\-world apis\.InProceedings of the 12th International Conference on Learning Representations,Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1),[§I](https://arxiv.org/html/2608.10039#S1.p6.1),[§IV\-A](https://arxiv.org/html/2608.10039#S4.SS1.p1.4)\.
- \[29\]S\. J\. Russell and P\. Norvig\(2021\)Artificial intelligence: a modern approach\.4 edition,Pearson\.Note:Global EditionCited by:[§IV\-A](https://arxiv.org/html/2608.10039#S4.SS1.p5.1.1)\.
- \[30\]T\. Schick, J\. Dwivedi\-Yu, R\. Dessí, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. Scialom\(2023\)Toolformer: language models can teach themselves to use tools\.InProceedings of the 37th International Conference on Neural Information Processing Systems,pp\. 68539–68551\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1)\.
- \[31\]Y\. Shang, Y\. Li, K\. Zhao, L\. Ma, J\. Liu, F\. Xu, and Y\. Li\(2025\)AgentSquare: automatic llm agent search in modular design space\.InProceedings of the 13th International Conference on Learning Representations,Cited by:[§VI](https://arxiv.org/html/2608.10039#S6.p2.2)\.
- \[32\]N\. Shinn, F\. Cassano, E\. Berman, A\. Gopinath, K\. Narasimhan, and S\. Yao\(2023\)Reflexion: language agents with verbal reinforcement learning\.InProceedings of the 37th Conference on Neural Information Processing Systems,Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1)\.
- \[33\]E\. Sirin, B\. Parsia, D\. Wu, J\. Hendler, and D\. Nau\(2004\)HTN planning for web service composition using shop2\.Journal of Web Semantics1\(4\),pp\. 377–396\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p3.1.1)\.
- \[34\]J\. Su, Q\. Lan, Y\. Xia, L\. Sun, W\. Tian, T\. Shi, and L\. He\(2026\)Difficulty\-aware agentic orchestration for query\-specific multi\-agent workflows\.InProceedings of the ACM Web Conference 2026,pp\. 2060–2070\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p2.1)\.
- \[35\]W\. Van der Aalst, T\. Weijters, and L\. Maruster\(2004\)Workflow mining: discovering process models from event logs\.IEEE Transactions on Knowledge and Data Engineering16\(9\),pp\. 1128–1142\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p3.1.1),[§I](https://arxiv.org/html/2608.10039#S1.p6.1),[§III\-A](https://arxiv.org/html/2608.10039#S3.SS1.p4.3),[§IV\-A](https://arxiv.org/html/2608.10039#S4.SS1.p2.1)\.
- \[36\]Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu, A\. H\. Awadallah, A\. White, D\. Burger, and C\. Wang\(2023\)AutoGen: enabling next\-gen LLM applications via multi\-agent conversation\.arXiv preprint arXiv:2308\.08155\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p3.1),[§VI](https://arxiv.org/html/2608.10039#S6.p1.1)\.
- \[37\]C\. S\. Xia, Y\. Deng, S\. Dunn, and L\. Zhang\(2025\)Demystifying llm\-based software engineering agents\.Proceedings of the ACM on Software Engineering2\(FSE\),pp\. 801–824\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1)\.
- \[38\]H\. Xu, X\. Huang, Y\. Liu, and Z\. Deng\(2026\)TPS\-bench: evaluating ai agents’ tool planning & scheduling abilities in compounding tasks\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 34949–34961\.Cited by:[§II\-B](https://arxiv.org/html/2608.10039#S2.SS2.p3.10.1)\.
- \[39\]J\. Xu, W\. Du, X\. Liu, and X\. Li\(2024\)LLM4Workflow: an llm\-based automated workflow model generation tool\.InProceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering,pp\. 2394–2398\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p3.1.1)\.
- \[40\]A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.External Links:[Link](https://arxiv.org/abs/2505.09388)Cited by:[§IV\-A](https://arxiv.org/html/2608.10039#S4.SS1.p6.2.2)\.
- \[41\]J\. Yang, C\. E\. Jimenez, A\. Wettig, K\. Lieret, S\. Yao, K\. Narasimhan, and O\. Press\(2024\)SWE\-agent: agent\-computer interfaces enable automated software engineering\.InProceedings of the 38th International Conference on Neural Information Processing Systems,Advances in Neural Information Processing Systems, Vol\.37,pp\. 50528–50652\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1)\.
- \[42\]S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. Cao\(2023\)ReAct: synergizing reasoning and acting in language models\.InProceedings of the 11th International Conference on Learning Representations,Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1),[§I](https://arxiv.org/html/2608.10039#S1.p6.1),[§IV\-A](https://arxiv.org/html/2608.10039#S4.SS1.p2.1)\.
- \[43\]L\. Yue, K\. R\. Bhandari, C\. Ko, D\. Patel, S\. Lin, N\. Zhou, J\. Gao, P\. Chen, and S\. Pan\(2026\)From static templates to dynamic runtime graphs: a survey of workflow optimization for llm agents\.arXiv preprint arXiv:2603\.22386\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p2.1)\.
- \[44\]J\. Zhang, J\. Xiang, Z\. Yu, F\. Teng, X\. Chen, J\. Chen, M\. Zhuge, X\. Cheng, S\. Hong, J\. Wang, B\. Zheng, B\. Liu, Y\. Luo, and C\. Wu\(2025\)AFlow: automating agentic workflow generation\.InProceedings of the 13th International Conference on Learning Representations,Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p4.1),[§I](https://arxiv.org/html/2608.10039#S1.p6.1),[§III\-C](https://arxiv.org/html/2608.10039#S3.SS3.p9.7),[§IV\-A](https://arxiv.org/html/2608.10039#S4.SS1.p2.1),[§VI](https://arxiv.org/html/2608.10039#S6.p2.2)\.
- \[45\]W\. Zhang, Y\. Li, Y\. Bei, J\. Luo, G\. Wan, L\. Yang, C\. Xie, Y\. Yang, W\. Huang, C\. Miao,et al\.\(2025\)From web search towards agentic deep research: incentivizing search with reasoning agents\.arXiv e\-prints,pp\. arXiv–2506\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p2.1)\.
- \[46\]Y\. Zhang, H\. Ruan, Z\. Fan, and A\. Roychoudhury\(2024\)AutoCodeRover: autonomous program improvement\.InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis,pp\. 1592–1604\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p1.1)\.
- \[47\]M\. Zhao, X\. Wei, Y\. Shao, K\. Zhou, L\. Yang, S\. Rao, J\. Zhan, and Z\. Chen\(2026\)A2A^\{2\}Flow: automating agentic workflow generation via self\-adaptive abstraction operators\.InProceedings of the 40th AAAI Conference on Artificial Intelligence,Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p4.1),[§VI](https://arxiv.org/html/2608.10039#S6.p2.2)\.
- \[48\]C\. Zheng, J\. Chen, Y\. Lyu, W\. Z\. T\. Ng, H\. Zhang, Y\. Ong, I\. Tsang, and H\. Yin\(2025\)Mermaidflow: redefining agentic workflow generation via safety\-constrained evolutionary programming\.arXiv preprint arXiv:2505\.22967\.Cited by:[§I](https://arxiv.org/html/2608.10039#S1.p2.1)\.Similar Articles
MoFlow: Multi-Objective Agentic Workflow Generation
MoFlow formulates agentic workflow generation as a multi-objective MDP and solves it with Convex-Hull Monte Carlo Tree Search with set-valued backups, covering the Pareto front so workflows can be retrieved for any preference without retraining. It outperforms six baselines on math, code, and QA benchmarks in terms of hypervolume.
FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
FlowEvo is a training-free framework that enables large language model agents to co-evolve reusable skills and workflows at inference time, achieving state-of-the-art accuracy and efficiency across benchmarks like ALFWorld, HumanEval, and GSM8K.
@ycombinator: flowscope deploys agents that learn and document how businesses operate. From there, their agents redesign and automate…
Flowscope is a Y Combinator-backed AI-native consulting firm that deploys agents to map, redesign, and automate business processes within days by integrating directly into existing enterprise systems.
FlowBank: Query-Adaptive Agentic Workflows Optimization through Precompute-and-Reuse
FlowBank introduces a three-stage framework for optimizing agentic workflows in LLM multi-agent systems by precomputing a diverse set of reusable workflows and adaptively selecting the best one per query, achieving higher scores while maintaining cost competitiveness.
SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
SkillFlow introduces a benchmark of 166 tasks across 20 families for evaluating autonomous agents' ability to discover, repair, and maintain skills over time through a lifelong learning protocol. Experiments reveal a substantial capability gap among leading models, with Claude Opus 4.6 improving significantly while others show limited or negative gains from skill evolution.