Tag
This paper argues that agent skills should incorporate visual information, not just text, and proposes a multimodal skill paradigm combining textual logic with visual support. Experiments show visual skills outperform text-only approaches in visual-centric tasks.