Tag
PenEcho integrates DeepSeek Harness to create an AI agent that can visually create and edit content on its canvas, enhancing user interaction through chat and providing professional visual output tools.
This paper introduces ToolSciVer, the first tool-augmented framework for multimodal scientific claim verification (MSCV), which equips a VLM with type-aware visual tools and trains the policy using GRPO to achieve superior performance on SciVer and MuSciClaims datasets across multiple model families.
This paper introduces CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and CanvasAgent, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interactions and hybrid reward optimization.
Stanford professor Judy Fan discusses at MIT how humans make the invisible visible through visual tools, contrasting with AI's limitations in visual understanding, and presents research on drawing, sketch recognition, and graph reading performance gaps between humans and AI models like GPT-4V.
DAIR Academy announces a free live session on building visual LLM artifacts to make LLM knowledge bases more actionable, with updates on new tools and releases for Pro members.