Beacon: Knowing When and How to Perform Agentic Visual Reasoning
Summary
This paper introduces Beacon, an agentic visual reasoning model that improves multimodal LLMs' ability to decide when to use tools and benefit from tool use, using Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion in reinforcement learning.
View Cached Full Text
Cached at: 07/31/26, 05:53 AM
Paper page - Beacon: Knowing When and How to Perform Agentic Visual Reasoning
Source: https://huggingface.co/papers/2607.28595 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Thefundamentalgoalofagenticvisualreasoningistoimprovethesuccessrateofmultimodallargelanguagemodels(MLLMs)oncomplextasks,ratherthanmerelyequippingthemwithasophisticatedyetinefficientreasoningparadigm.Inthiswork,werethinkagenticvisualreasoningthroughtwokeydimensionsoftooluse:ModeAdaptiveness(MA)andToolEffect(TE).ModeAdaptivenesscharacterizeswhetheranMLLMcanrecognizewhentoolsaretrulynecessaryandinvokethemaccordingly,therebyavoidingunnecessarycomputationaloverheadwhileimprovingperformanceonchallengingproblemsthatrequiretoolassistance.ToolEffectcharacterizestheactualimpactoftooluse:toolsshouldextendthemodel’scapabilitiesonproblemsunsolvablethroughtext-onlyreasoning,whileavoidingadditionalerrorsonproblemsthatthemodelcanalreadysolvewithouttools.WeconductacomprehensiveanalysistoquantifythesetwopropertiesandempiricallyrevealthatexistingagenticvisualreasoningmodelsexhibitlimitedModeAdaptiveness,whilethegainsproducedbytooluseonhardexamplesarelargelyoffsetbytheharmintroducedoneasyexamplesthatthemodelscanalreadysolve.Motivatedbytheseobservations,weproposeBeacon,anovelagenticvisualreasoningmodelthatachievesstrongeroverallperformance,improvedModeAdaptiveness,andgenuinetool-inducedperformancegains.AtthecoreofBeaconaretheNecessity-AwareAdaptiveRewardandtheHint-GuidedCapabilityExpansionmechanisminthereinforcementlearningstage,whichrespectivelyencourageadaptivetoolinvocationbasedontasknecessityandstrengthenthemodel’stool-usecapabilityonthemostchallengingproblems.ExtensiveexperimentsacrossdiversebenchmarksdemonstratethestrongoverallperformanceofBeaconanditssubstantialimprovementsinbothModeAdaptivenessandToolEffect.
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2607\.28595
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.28595 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.28595 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.28595 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
ATLAS presents a visual reasoning framework that combines agentic operations and latent representations using functional tokens, enabling efficient training via next-token prediction and reinforcement learning while avoiding intermediate image generation.
Bad Seeing or Bad Thinking? Rewarding Perception for Vision-Language Reasoning
This paper introduces a reinforcement learning framework that improves perception-reasoning synergy in vision-language models by explicitly rewarding perceptual fidelity, using a 'blindfolded reasoning' proxy and structured verbal verification to address ambiguity in modality credit assignment.
Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning
Introduces AutoTool, a model that adaptively decides whether to invoke tools for multimodal LLM reasoning, achieving significant accuracy and efficiency gains through reinforcement learning and dual-mode reasoning.
Adaptive Latent Agentic Reasoning
This paper introduces Adaptive Latent Agentic Reasoning (ALAR), a dual-mode framework for LLM agents that uses compact latent reasoning for routine turns and selectively escalates to explicit chain-of-thought for harder decisions, achieving up to 84.6% token reduction while maintaining task accuracy.
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
This paper introduces LedgerMind, a provenance-constrained multimodal agentic reasoning framework that uses a Structured Evidence Ledger to ensure grounded, faithful reasoning in visual question answering, addressing failure patterns like hallucination and over-reasoning.