multimodal-agent

Tag

Cards List
#multimodal-agent

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery

Hugging Face Daily Papers · 2026-07-07 Cached

This paper proposes VaseMuseum, a lightweight multimodal agent framework that combines 3D digitization with vision-language models to create an interactive digital museum for ancient Greek pottery, addressing challenges of evidence grounding and hallucination through source-level and response-level reliability control.

0 favorites 0 likes
#multimodal-agent

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration

Hugging Face Daily Papers · 2026-07-06 Cached

This paper introduces CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and CanvasAgent, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interactions and hybrid reward optimization.

0 favorites 0 likes
#multimodal-agent

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory

Hugging Face Daily Papers · 2026-07-06 Cached

Light-Omni is a multimodal agent framework for efficient video understanding that uses dual contextual states (global state and parametric latent state) to avoid iterative reasoning, achieving faster and more accurate processing with significant speedup and memory savings.

0 favorites 0 likes
#multimodal-agent

Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

arXiv cs.AI · 2026-06-16 Cached

Visual-Seeker proposes a visual-native multimodal deep search agent that actively reasons over fine-grained visual details and synthesizes multimodal evidence, achieving state-of-the-art performance on five challenging multimodal search benchmarks.

0 favorites 0 likes
#multimodal-agent

VisualClaw: A Real-Time, Personalized Agent for the Physical World

Hugging Face Daily Papers · 2026-06-15 Cached

VisualClaw is a self-evolving multimodal agent that reduces deployment costs through hybrid encoding and skill evolution, while improving video-QA accuracy across multiple benchmarks.

0 favorites 0 likes
#multimodal-agent

@MaxForAI: Yesterday, ByteDance Seed open-sourced a very interesting checkpoint, TaskMem. It is trained on Qwen3-VL-30B-A3B, with the goal not being to directly answer questions, but to enable multimodal Agents to learn to generate more useful long-term memory from video/environment streams. The key is to let the Agent learn in continuous video…

X AI KOLs Timeline · 2026-06-03 Cached

ByteDance Seed has open-sourced the TaskMem checkpoint, trained on Qwen3-VL-30B-A3B. It uses two-stage reinforcement learning to enable multimodal Agents to learn to generate long-term memory from video streams, achieving significant improvements on benchmarks such as VideoMME and EgoLife.

0 favorites 0 likes
← Back to home

Submit Feedback