It Takes Workflows to Evolve Better Workflows

Hugging Face Daily Papers Papers

Summary

FloWright proposes a hierarchical, structure-aware reward paradigm enabling LLM agents in multi-agent workflows to self-evolve and co-evolve without extra models or labels, while DataWright adaptively hardens datasets into workflow-level tasks. Small open models trained with FloWright gain up to +7.41% performance, with co-evolving multiple roles (+5.03%) outperforming optimizing a single role (+2.83%).

Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows that coordinate specialized agents to work together on these tasks. Recent methods train LLMs to construct better workflows from execution outcomes, but they optimize only the workflow generator, while the other agents that build or execute each workflow remain fixed even though every outcome depends on all of them. However, extending training beyond the generator is challenging: the agents are coupled, and a workflow's outcome is a single sparse score that cannot tell which agent causes a failure. We propose FloWright, which leverages the workflow as a harness to optimize workflows. By introducing a hierarchical, structure-aware reward paradigm, FloWright enables one role to self-evolve and two or more roles to co-evolve, with no additional models, labels, or executions. Considering the limitation that workflows are commonly trained and evaluated on data that a single agent can already handle, we further propose DataWright, an adaptive data hardening approach that converts existing datasets into workflow-level tasks with increased difficulty. Across document, slide, chart, code, math, and finance tasks, small open models trained with FloWright achieve improved performance by up to +7.41%, with co-evolving (+5.03%) more roles gaining more than optimizing one of them alone (+2.83%). Our project page: https://xhguo7.github.io/FloWright/.
Original Article
View Cached Full Text

Cached at: 10/02/26, 04:28 AM

Paper page - It Takes Workflows to Evolve Better Workflows

Source: https://huggingface.co/papers/2610.01026

Abstract

Tacklingcomplexreal-worldtaskscanexceedthecapabilitiesofasinglelargelanguagemodel(LLM),motivatingtheuseofmulti-agentworkflowsthatcoordinatespecializedagentstoworktogetheronthesetasks.RecentmethodstrainLLMstoconstructbetterworkflowsfromexecutionoutcomes,buttheyoptimizeonlytheworkflowgenerator,whiletheotheragentsthatbuildorexecuteeachworkflowremainfixedeventhougheveryoutcomedependsonallofthem.However,extendingtrainingbeyondthegeneratorischallenging:theagentsarecoupled,andaworkflow’soutcomeisasinglesparsescorethatcannottellwhichagentcausesafailure.WeproposeFloWright,whichleveragestheworkflowasaharnesstooptimizeworkflows.Byintroducingahierarchical,structure-awarerewardparadigm,FloWrightenablesoneroletoself-evolveandtwoormorerolestoco-evolve,withnoadditionalmodels,labels,orexecutions.Consideringthelimitationthatworkflowsarecommonlytrainedandevaluatedondatathatasingleagentcanalreadyhandle,wefurtherproposeDataWright,anadaptivedatahardeningapproachthatconvertsexistingdatasetsintoworkflow-leveltaskswithincreaseddifficulty.Acrossdocument,slide,chart,code,math,andfinancetasks,smallopenmodelstrainedwithFloWrightachieveimprovedperformancebyupto+7.41%,withco-evolving(+5.03%)morerolesgainingmorethanoptimizingoneofthemalone(+2.83%).Ourprojectpage:https://xhguo7.github.io/FloWright/.

View arXiv pageView PDFAdd to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2610.01026 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2610.01026 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2610.01026 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution

arXiv cs.CL

CoEvolve proposes an agent-data mutual evolution framework for training LLM agents through closed-loop, interaction-driven learning that adapts both the agent and its training data distribution. The method extracts feedback signals from rollout trajectories to guide LLM-based task synthesis, demonstrating significant improvements (15-19% absolute gains) across multiple Qwen models on AppWorld and BFCL benchmarks.