Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

Hugging Face Daily Papers 05/26/26, 12:00 AM Papers

agentic-rl llm-agents tool-use knowledge-boundary on-policy reward-hacking accuracy-efficiency

Summary

This paper proposes AKBE, an on-policy method for LLM agent reinforcement learning that dynamically identifies when tool use is needed versus when internal knowledge suffices, improving accuracy by +1.85 on average and reducing tool calls by 18% over standard agentic RL.

Agentic reinforcement learning (RL) has proven effective for training LLM-based agents with external tool-use capabilities. However, we identify that agentic RL training induces increasing redundant tool calls and blurs the model's intrinsic knowledge boundary, where the model fails to distinguish when tools are needed versus when parametric knowledge suffices. Existing solutions based on reward shaping create coarse-grained optimization targets that tend to incentivize indiscriminate tool-call suppression, leading to reward hacking. In this paper, we propose AKBE (Agentic Knowledge Boundary Enhancement), an on-policy method that dynamically probes the model's intrinsic knowledge boundary through dual-path (with-tool and no-tool) rollouts during training. We define the knowledge boundary as the per-instance determination of whether tools are required and the minimum tool calls necessary. By comparing correctness across paths, AKBE categorizes trajectories and constructs targeted supervisory signals that guide efficient tool-use patterns for each question. These signals are integrated seamlessly into the agentic RL training loop. Experiments on seven QA benchmarks demonstrate that AKBE improves task accuracy by +1.85 on average and reduces tool calls by 18% over standard agentic RL, yielding 25% higher tool productivity without any accuracy-efficiency trade-off. Further analysis suggests its plug-and-play compatibility across different RL algorithms and the mechanism of each signal category. Our code is available at https://github.com/CuSO4-Chen/AKBE.

Original Article

View Cached Full Text

Cached at: 05/27/26, 02:47 AM

Paper page - Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

Source: https://huggingface.co/papers/2605.26952

Abstract

AKBE enhances LLM agent training by dynamically identifying when tools are needed versus when internal knowledge suffices, improving accuracy and reducing unnecessary tool usage through targeted supervisory signals.

Agentic reinforcement learning(RL) has proven effective for training LLM-based agents with externaltool-use capabilities. However, we identify that agentic RL training induces increasing redundant tool calls and blurs the model’s intrinsicknowledge boundary, where the model fails to distinguish when tools are needed versus when parametric knowledge suffices. Existing solutions based onreward shapingcreate coarse-grained optimization targets that tend to incentivize indiscriminate tool-call suppression, leading toreward hacking. In this paper, we propose AKBE (AgenticKnowledge BoundaryEnhancement), anon-policy methodthat dynamically probes the model’s intrinsicknowledge boundarythrough dual-path (with-tool and no-tool) rollouts during training. We define theknowledge boundaryas the per-instance determination of whether tools are required and the minimum tool calls necessary. By comparing correctness across paths, AKBE categorizes trajectories and constructs targetedsupervisory signalsthat guide efficient tool-use patterns for each question. These signals are integrated seamlessly into the agentic RL training loop. Experiments on seven QA benchmarks demonstrate that AKBE improvestask accuracyby +1.85 on average and reduces tool calls by 18% over standard agentic RL, yielding 25% highertool productivitywithout any accuracy-efficiency trade-off. Further analysis suggests its plug-and-play compatibility across different RL algorithms and the mechanism of each signal category. Our code is available at https://github.com/CuSO4-Chen/AKBE.

View arXiv page View PDF GitHub Add to collection

Get this paper in your agent:

hf papers read 2605\.26952

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2605.26952 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2605.26952 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2605.26952 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

Paper page - Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

Abstract

Models citing this paper0

Datasets citing this paper0

Spaces citing this paper0

Collections including this paper0

Similar Articles

Learning Agentic Policy from Action Guidance

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents

Milestone-Guided Policy Learning for Long-Horizon Language Agents

PolicyBank: Evolving Policy Understanding for LLM Agents

Submit Feedback

Similar Articles

Learning Agentic Policy from Action Guidance

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents

Milestone-Guided Policy Learning for Long-Horizon Language Agents

PolicyBank: Evolving Policy Understanding for LLM Agents