The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows

Hugging Face Daily Papers Papers

Summary

ReASearch is a unified framework where a single tool-using agent internalizes the search policy for optimizing prompts, programs, and ML workflows, outperforming specialized baselines across 14 tasks.

Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search policy can be internalized by a single tool-using agent? We present ReASearch, a unified framework for reasoning-driven optimization in which the agent autonomously decides what to evaluate, how to diagnose failures, which edits to make, and when to verify or restart. Rather than serving only as a proposal generator guided by hand-designed heuristics, the agent actively analyzes outcomes, allocates budget, and refines its strategy over long horizons through persistent memory. With a shared agent loop and domain-specific tools, ReASearch instantiates the exact same scaffold to optimize prompts, programs, and ML workflows. Across 14 diverse tasks, it is competitive with and mostly better than specialized optimization systems, achieving gains of 2% to 40% over strong domain-specific baselines, and in some cases discovering solutions that improve on prior human best-known results. Crucially, we observe that complex search behaviors, which are typically implemented by explicit controllers, emerge naturally from the agent's reasoning process.
Original Article
View Cached Full Text

Cached at: 08/10/26, 06:14 AM

Paper page - The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows

Source: https://huggingface.co/papers/2608.06714

Abstract

Recentsystemsforoptimizingprompts,programs,andMLworkflowstypicallyrelyonexplicitouter-loopcontrollerssuchasevolutionarysearch,bandits,ortextual-gradientmethods.Weaskafundamentallydifferentquestion:howmuchofthissearchpolicycanbeinternalizedbyasingletool-usingagent?WepresentReASearch,aunifiedframeworkforreasoning-drivenoptimizationinwhichtheagentautonomouslydecideswhattoevaluate,howtodiagnosefailures,whicheditstomake,andwhentoverifyorrestart.Ratherthanservingonlyasaproposalgeneratorguidedbyhand-designedheuristics,theagentactivelyanalyzesoutcomes,allocatesbudget,andrefinesitsstrategyoverlonghorizonsthroughpersistentmemory.Withasharedagentloopanddomain-specifictools,ReASearchinstantiatestheexactsamescaffoldtooptimizeprompts,programs,andMLworkflows.Across14diversetasks,itiscompetitivewithandmostlybetterthanspecializedoptimizationsystems,achievinggainsof2%to40%overstrongdomain-specificbaselines,andinsomecasesdiscoveringsolutionsthatimproveonpriorhumanbest-knownresults.Crucially,weobservethatcomplexsearchbehaviors,whicharetypicallyimplementedbyexplicitcontrollers,emergenaturallyfromtheagent’sreasoningprocess.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2608\.06714

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.06714 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.06714 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.06714 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

SAGE: Stochastic Prompt Optimization via Agent-Guided Exploration

arXiv cs.CL

Introduces SPO, a stochastic search framework for automatic prompt optimization, with three strategies including SAGE, an agent-guided multi-agent pipeline. Evaluated on benchmarks and deployed on a mental-health chatbot, showing improvements in retention through continuous optimization.

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Hugging Face Daily Papers

AREX introduces a family of recursively self-improving agents for deep research, alternating between an inner research loop and an outer self-improvement loop, trained with long-horizon reinforcement learning. It substantially outperforms comparable-scale baselines on benchmarks like BrowseComp and Humanity's Last Exam.