multimodal-agents

Tag

Cards List
#multimodal-agents

Learning to Route in Visual Space via Multi-Step Embedding Retrieval

arXiv cs.AI ↗ · yesterday Cached

This paper introduces VHop, a data generation framework and benchmark for multi-step visual retrieval, and VHop-Router, an end-to-end trained autoregressive multi-step retriever that operates in visual latent space, boosting retrieval from under 5% to 76.3% and improving agentic search success by 52.7% while drastically cutting tokens and API payload.

0 favorites 0 likes
#multimodal-agents

WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

Hugging Face Daily Papers ↗ · 2d ago Cached

WorldAuditBench is a benchmark of 213 anomaly-auditing tasks across 13 interactive 3D environments (Unreal Engine 5 / Three.js) that evaluates whether multimodal agents can couple action with visual reasoning. Frontier models achieve only 6.6–42.3% success versus 83.4% human performance, highlighting current limitations in evidence gathering and interpretation.

0 favorites 0 likes
#multimodal-agents

VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents

Hugging Face Daily Papers ↗ · 3d ago Cached

The paper proposes VideoLoop, a multimodal agent with looped working memory to address semantic thrashing in long-form video understanding, demonstrating performance gains on benchmarks like VideoMME.

0 favorites 0 likes
#multimodal-agents

NavHarness: Towards Lifelong Embodied Navigation

Hugging Face Daily Papers ↗ · 4d ago Cached

The paper introduces NavHarness, a training-free embodied navigation harness that integrates memory processing into the navigation loop, enabling agents to carry experience across sessions for lifelong navigation. On GOAT-Bench, it improves success rate by up to 22.6 points over context-only sessions and achieves state-of-the-art results with GPT-6/Astra.

0 favorites 0 likes
#multimodal-agents

Learning What to Activate: Combinatorial Capability Allocation for Long-Horizon Multimodal Agents

arXiv cs.AI ↗ · 2026-09-24 Cached

The paper introduces CoCA, an on-policy learning framework for dynamic capability allocation in long-horizon multimodal agents, addressing computational overhead and stage-dependent demands through conditional comparisons and reinforcement learning.

0 favorites 0 likes
#multimodal-agents

Qwen3.8-Omni: Towards Native Omni-Modal Agents

arXiv cs.CL ↗ · 2026-09-23 Cached

Qwen3.8-Omni-Flash is a natively multimodal agentic model designed for real-world productivity, improving multimodal understanding and reasoning with a MoE architecture and extended context window, along with releasing open-source frameworks for multimodal applications.

0 favorites 0 likes
#multimodal-agents

UI-Venus-2 Technical Report

Hugging Face Daily Papers ↗ · 2026-08-27 Cached

UI-Venus-2 is an open-source multimodal GUI agent designed for digital automation across mobile, web, and desktop environments, using unified reasoning-action loops and robust verification to enable reliable real-world applications.

0 favorites 0 likes
#multimodal-agents

Representation Affects Retrieval: A Case Study of Skill Discovery and Routing in a Multimodal Agent Harness

arXiv cs.AI ↗ · 2026-08-24 Cached

This paper presents a case study on skill discovery and routing in a multimodal agent harness, showing that partial in-prompt exposure of skills can create lexical competition that hinders correct selection, linking small-scale in-context retrieval to large-scale approaches.

0 favorites 0 likes
#multimodal-agents

A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications

arXiv cs.AI ↗ · 2026-08-24 Cached

This survey analyzes the impact of multimodality on agentic frameworks, examining core modules like perception, reasoning, and planning, and reviews applications across robotics, GUI navigation, and other domains.

0 favorites 0 likes
#multimodal-agents

VibeWorlding benchmark: open 30B model tops GPT-5.5 on 3D worlds

Reddit r/ArtificialInteligence ↗ · 2026-08-18

A new benchmark for evaluating multimodal agents in building 3D open worlds from natural language shows GPT-5.5 and Qwen3.8-Max scoring below 60%, while an open 30B model achieves the best performance through reinforcement learning.

0 favorites 0 likes
#multimodal-agents

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

Hugging Face Daily Papers ↗ · 2026-08-15 Cached

This paper proposes VibeWorlding, a framework for benchmarking and training multimodal agents to construct 3D open worlds from user queries, showing that reinforcement learning improves open-source models to compete with closed-source frontiers.

0 favorites 0 likes
#multimodal-agents

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper presents OpenVisTool, an open framework for synthesizing instructive visual tool-use trajectories, along with a dataset (OpenVisTool-42K) and benchmark. It shows that fine-tuning on causally grounded supervision improves visual tool-use performance across multiple model backbones.

0 favorites 0 likes
#multimodal-agents

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

The paper argues that simply scaling multimodal environments does not always improve agent training, and proposes Ability-aware Environment Selection (AES) and Hierarchical Difficulty Curriculum (HDC) to better structure environment distributions along diversity and difficulty dimensions.

0 favorites 0 likes
#multimodal-agents

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

Hugging Face Daily Papers ↗ · 2026-08-03 Cached

DeepVoyager-VL proposes a long-horizon multimodal deep-search framework that integrates visual evidence into intermediate reasoning, using a multimodal event graph for data synthesis and fine-tuning without reinforcement learning, achieving strong performance across ten benchmarks.

0 favorites 0 likes
#multimodal-agents

CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

arXiv cs.AI ↗ · 2026-07-31 Cached

CaM-Wolf is the first multimodal AI agent for social deduction games like Werewolf, integrating video perception and generation with a causal-aware reasoner trained via reinforcement learning to handle hidden roles and social reasoning. Experiments and user studies show improved gameplay and human-AI interaction quality.

0 favorites 0 likes
#multimodal-agents

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

arXiv cs.AI ↗ · 2026-06-30 Cached

The paper introduces GPTNT, a benchmark built on Keep Talking and Nobody Explodes that requires two multimodal agents to collaborate in real-time under time pressure and information asymmetry, revealing critical weaknesses in state-of-the-art systems.

0 favorites 0 likes
#multimodal-agents

DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection

arXiv cs.CL ↗ · 2026-06-29 Cached

Introduces DMV-Bench, an interactive benchmark for evaluating visual memory in multimodal agents using incidental visual cues from product images, and proposes DualMem, a dual-coding memory architecture that outperforms text-only and other multimodal baselines across various chain lengths.

0 favorites 0 likes
#multimodal-agents

Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation

arXiv cs.CL ↗ · 2026-06-16 Cached

This paper introduces AgentViSS, a benchmark evaluating visual social intelligence in multimodal social simulation, containing 240 scenarios with aligned visual-textual evidence. Evaluating seven recent MLLMs reveals a gap between local role enactment and visually grounded interaction management.

0 favorites 0 likes
#multimodal-agents

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

Hugging Face Daily Papers ↗ · 2026-06-08 Cached

SpatialWorld is a unified benchmark for evaluating interactive spatial reasoning in multimodal agents across diverse real-world tasks, revealing that even the strongest models achieve low task success rates.

0 favorites 0 likes
#multimodal-agents

Task-Focused Memorization for Multimodal Agents

Hugging Face Daily Papers ↗ · 2026-05-29 Cached

Introduces TaskMem, a reinforcement-learning-based framework for dynamic memorization in multimodal agents, achieving accuracy improvements of 6.3%, 7.0%, and 5.3% on streaming video benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback