CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding

Hugging Face Daily Papers Papers

Summary

CoEvoWhen proposes a policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from VLM agentic reasoning trajectories to improve ultra-long video temporal grounding without updating model parameters.

Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters. During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition. Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model. Extensive experiments spanning five benchmarks and three VLMs show that policy-tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.
Original Article
View Cached Full Text

Cached at: 10/01/26, 04:20 AM

Paper page - CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding

Source: https://huggingface.co/papers/2609.40048

Abstract

Ultra-longvideotemporalgroundingrequiresbalancinglong-rangeevidencesearchwithfine-grainedeventunderstandingunderalimitedvisualbudget,yetexistingagenticmethodsstillrelylargelyonpredefinedpoliciesandtoolcapabilities.Motivatedbythis,weproposeanovelpolicy-toolcoevolutionframeworkthatjointlyevolveshigh-levelpoliciesandexecutablemediatoolsfromtheagenticreasoningtrajectoriesofaVLM,formingareusableskillwithoutupdatingmodelparameters.Duringevolution,anexternalskillupdaterdistillstransferabletaskexperienceinlong-videotemporalgrounding,accordinglyrefiningtheorchestrationoflong-rangeimage-basedandfine-grainedvideo-basedobservations.Alongsidethesepolicyupdates,theupdateremploysitscodingcapabilitiestoupgradeexistingtoolsorcreatenewones,adaptingthetoolstolong-videoevidenceacquisition.Equippedwiththeevolvedskill,theVLMautonomouslyorchestratestoolsundertheguidanceoftheevolvedpolicy,coordinatingimageandvideoobservationsforagenticinferencewithoutrelyingonaseparate,strongerplanningmodel.ExtensiveexperimentsspanningfivebenchmarksandthreeVLMsshowthatpolicy-toolcoevolutionconsistentlyimprovestemporalgroundingaccuracyinultra-longvideoswhilereducingvisualtokencostatinference,andthattheevolvedskillyieldssubstantialperformancegainsongenerallong-videoQAwithoutadditionaltask-specificevolution,demonstratingtheeffectivenessandgeneralizabilityofourframeworkforlong-videounderstanding.

View arXiv pageView PDFProject pageGitHub0Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.40048 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.40048 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.40048 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution

arXiv cs.CL

CoEvolve proposes an agent-data mutual evolution framework for training LLM agents through closed-loop, interaction-driven learning that adapts both the agent and its training data distribution. The method extracts feedback signals from rollout trajectories to guide LLM-based task synthesis, demonstrating significant improvements (15-19% absolute gains) across multiple Qwen models on AppWorld and BFCL benchmarks.