CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Hugging Face Daily Papers Papers

Summary

CAFE is a framework that couples a search agent and critic via shared parameters to learn in-trajectory corrective feedback, improving search performance and reducing hallucinations across benchmarks.

Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound. Treating corrective feedback as a learned in-trajectory intervention couples the two roles: the agent must decide when to request and use feedback, while the critic must infer useful corrections from outcome-confounded rollouts whose failure patterns shift as the agent improves. We introduce CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles. CAFE initializes feedback-conditioned recovery from trajectories built around the base agent's own failures, then couples online and offline optimization. During online RL, a comparative feedback estimate uses a prompt-level call--skip success gap to shape request returns, while feedback-aware advantage shaping reweights token advantages before and after feedback. Offline, rollout-derived preference optimization learns feedback from matched successful and unsuccessful trajectories. On seven agentic search benchmarks, CAFE outperforms the evaluated RL-based search agents on average, retains its gains across all six out-of-domain benchmarks, and reduces answer-level hallucinations. One-sided ablations show that improving only the agent or only the critic eventually plateaus, whereas alternating the two updates continues to improve performance. These findings suggest that a self-improving search agent needs feedback that co-evolves with the policy it guides.
Original Article
View Cached Full Text

Cached at: 08/26/26, 03:17 AM

Paper page - CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Source: https://huggingface.co/papers/2608.24794 Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

CAFE couples a search agent and critic via shared parameters to learn in-trajectory corrective feedback, improving search performance and reducing hallucinations across benchmarks.

Outcome-supervised search agentslearn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound. Treatingcorrective feedbackas a learnedin-trajectory interventioncouples the two roles: the agent must decide when to request and use feedback, while the critic must infer useful corrections from outcome-confounded rollouts whose failure patterns shift as the agent improves. We introduceCAFE(Coupled Agent--Feedback Evolution), a framework in which ashared-parameter modelalternates between search-agent and critic roles.CAFEinitializesfeedback-conditioned recoveryfrom trajectories built around the base agent’s own failures, then couples online and offline optimization. Duringonline RL, acomparative feedback estimateuses a prompt-level call--skip success gap to shape request returns, whilefeedback-aware advantage shapingreweightstoken advantagesbefore and after feedback. Offline,rollout-derived preference optimizationlearns feedback from matched successful and unsuccessful trajectories. On sevenagentic search benchmarks,CAFEoutperforms the evaluatedRL-based search agentson average, retains its gains across all sixout-of-domain benchmarks, and reducesanswer-level hallucinations. One-sided ablations show that improving only the agent or only the critic eventually plateaus, whereas alternating the two updates continues to improve performance. These findings suggest that a self-improving search agent needs feedback that co-evolves with the policy it guides.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2608\.24794

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.24794 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.24794 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.24794 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

CAS: Conformalized Agentic Search via Adaptive Retrieval and Policy Weighting

arXiv cs.AI

This paper introduces Conformalized Agentic Search (CAS), a framework that uses Conformal Prediction to enhance the reliability of search agents by adaptively retrieving documents and weighting policies during reinforcement learning, thereby improving accuracy and reducing redundant tool calls.