Tag
This paper introduces VHop, a data generation framework and benchmark for multi-step visual retrieval, and VHop-Router, an end-to-end trained autoregressive multi-step retriever that operates in visual latent space, boosting retrieval from under 5% to 76.3% and improving agentic search success by 52.7% while drastically cutting tokens and API payload.
WorldAuditBench is a benchmark of 213 anomaly-auditing tasks across 13 interactive 3D environments (Unreal Engine 5 / Three.js) that evaluates whether multimodal agents can couple action with visual reasoning. Frontier models achieve only 6.6–42.3% success versus 83.4% human performance, highlighting current limitations in evidence gathering and interpretation.
The paper proposes VideoLoop, a multimodal agent with looped working memory to address semantic thrashing in long-form video understanding, demonstrating performance gains on benchmarks like VideoMME.
The paper introduces NavHarness, a training-free embodied navigation harness that integrates memory processing into the navigation loop, enabling agents to carry experience across sessions for lifelong navigation. On GOAT-Bench, it improves success rate by up to 22.6 points over context-only sessions and achieves state-of-the-art results with GPT-6/Astra.
The paper introduces CoCA, an on-policy learning framework for dynamic capability allocation in long-horizon multimodal agents, addressing computational overhead and stage-dependent demands through conditional comparisons and reinforcement learning.
Qwen3.8-Omni-Flash is a natively multimodal agentic model designed for real-world productivity, improving multimodal understanding and reasoning with a MoE architecture and extended context window, along with releasing open-source frameworks for multimodal applications.
UI-Venus-2 is an open-source multimodal GUI agent designed for digital automation across mobile, web, and desktop environments, using unified reasoning-action loops and robust verification to enable reliable real-world applications.
This paper presents a case study on skill discovery and routing in a multimodal agent harness, showing that partial in-prompt exposure of skills can create lexical competition that hinders correct selection, linking small-scale in-context retrieval to large-scale approaches.
This survey analyzes the impact of multimodality on agentic frameworks, examining core modules like perception, reasoning, and planning, and reviews applications across robotics, GUI navigation, and other domains.
A new benchmark for evaluating multimodal agents in building 3D open worlds from natural language shows GPT-5.5 and Qwen3.8-Max scoring below 60%, while an open 30B model achieves the best performance through reinforcement learning.
This paper proposes VibeWorlding, a framework for benchmarking and training multimodal agents to construct 3D open worlds from user queries, showing that reinforcement learning improves open-source models to compete with closed-source frontiers.
This paper presents OpenVisTool, an open framework for synthesizing instructive visual tool-use trajectories, along with a dataset (OpenVisTool-42K) and benchmark. It shows that fine-tuning on causally grounded supervision improves visual tool-use performance across multiple model backbones.
The paper argues that simply scaling multimodal environments does not always improve agent training, and proposes Ability-aware Environment Selection (AES) and Hierarchical Difficulty Curriculum (HDC) to better structure environment distributions along diversity and difficulty dimensions.
DeepVoyager-VL proposes a long-horizon multimodal deep-search framework that integrates visual evidence into intermediate reasoning, using a multimodal event graph for data synthesis and fine-tuning without reinforcement learning, achieving strong performance across ten benchmarks.
CaM-Wolf is the first multimodal AI agent for social deduction games like Werewolf, integrating video perception and generation with a causal-aware reasoner trained via reinforcement learning to handle hidden roles and social reasoning. Experiments and user studies show improved gameplay and human-AI interaction quality.
The paper introduces GPTNT, a benchmark built on Keep Talking and Nobody Explodes that requires two multimodal agents to collaborate in real-time under time pressure and information asymmetry, revealing critical weaknesses in state-of-the-art systems.
Introduces DMV-Bench, an interactive benchmark for evaluating visual memory in multimodal agents using incidental visual cues from product images, and proposes DualMem, a dual-coding memory architecture that outperforms text-only and other multimodal baselines across various chain lengths.
This paper introduces AgentViSS, a benchmark evaluating visual social intelligence in multimodal social simulation, containing 240 scenarios with aligned visual-textual evidence. Evaluating seven recent MLLMs reveals a gap between local role enactment and visually grounded interaction management.
SpatialWorld is a unified benchmark for evaluating interactive spatial reasoning in multimodal agents across diverse real-world tasks, revealing that even the strongest models achieve low task success rates.
Introduces TaskMem, a reinforcement-learning-based framework for dynamic memorization in multimodal agents, achieving accuracy improvements of 6.3%, 7.0%, and 5.3% on streaming video benchmarks.