It Takes Workflows to Evolve Better Workflows
Summary
FloWright proposes a hierarchical, structure-aware reward paradigm enabling LLM agents in multi-agent workflows to self-evolve and co-evolve without extra models or labels, while DataWright adaptively hardens datasets into workflow-level tasks. Small open models trained with FloWright gain up to +7.41% performance, with co-evolving multiple roles (+5.03%) outperforming optimizing a single role (+2.83%).
View Cached Full Text
Cached at: 10/02/26, 04:28 AM
Paper page - It Takes Workflows to Evolve Better Workflows
Source: https://huggingface.co/papers/2610.01026
Abstract
Tacklingcomplexreal-worldtaskscanexceedthecapabilitiesofasinglelargelanguagemodel(LLM),motivatingtheuseofmulti-agentworkflowsthatcoordinatespecializedagentstoworktogetheronthesetasks.RecentmethodstrainLLMstoconstructbetterworkflowsfromexecutionoutcomes,buttheyoptimizeonlytheworkflowgenerator,whiletheotheragentsthatbuildorexecuteeachworkflowremainfixedeventhougheveryoutcomedependsonallofthem.However,extendingtrainingbeyondthegeneratorischallenging:theagentsarecoupled,andaworkflow’soutcomeisasinglesparsescorethatcannottellwhichagentcausesafailure.WeproposeFloWright,whichleveragestheworkflowasaharnesstooptimizeworkflows.Byintroducingahierarchical,structure-awarerewardparadigm,FloWrightenablesoneroletoself-evolveandtwoormorerolestoco-evolve,withnoadditionalmodels,labels,orexecutions.Consideringthelimitationthatworkflowsarecommonlytrainedandevaluatedondatathatasingleagentcanalreadyhandle,wefurtherproposeDataWright,anadaptivedatahardeningapproachthatconvertsexistingdatasetsintoworkflow-leveltaskswithincreaseddifficulty.Acrossdocument,slide,chart,code,math,andfinancetasks,smallopenmodelstrainedwithFloWrightachieveimprovedperformancebyupto+7.41%,withco-evolving(+5.03%)morerolesgainingmorethanoptimizingoneofthemalone(+2.83%).Ourprojectpage:https://xhguo7.github.io/FloWright/.
View arXiv pageView PDFAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2610.01026 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2610.01026 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2610.01026 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
FlowEvo is a training-free framework that enables large language model agents to co-evolve reusable skills and workflows at inference time, achieving state-of-the-art accuracy and efficiency across benchmarks like ALFWorld, HumanEval, and GSM8K.
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
CoEvolve proposes an agent-data mutual evolution framework for training LLM agents through closed-loop, interaction-driven learning that adapts both the agent and its training data distribution. The method extracts feedback signals from rollout trajectories to guide LLM-based task synthesis, demonstrating significant improvements (15-19% absolute gains) across multiple Qwen models on AppWorld and BFCL benchmarks.
Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
This paper proposes a reward-driven LLM agent workflow that integrates POMDP routing and self-correcting reward models, achieving a 24.5% improvement in task success rate on benchmarks like ALFWorld and WebShop.
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
This paper studies when end-to-end reinforcement learning training improves multi-agent LLM workflows, comparing shared-policy and isolated-policy training across different workflows, tasks, and model scales, revealing conditional tradeoffs.
@rohanpaul_ai: Better self-improving agents need better solvers, not bigger update-writing models. This challenges the common habit of…
This paper disentangles the roles of evolver and agent in self-improving LLM agents, showing that a small evolver can write sufficiently good updates, while a mid-tier agent benefits most from using them. It recommends using the strongest model as the task executor, not the update writer.