Tag
Researchers introduce CaptureGuide-Bench, a benchmark for capture-time photography guidance, and ShutterMuse, a unified multimodal LLM trained to provide composition and pose recommendations, demonstrating improved performance over general-purpose models.
This paper identifies and addresses the 'editing decoupling failure' in Multimodal LLMs, where knowledge updates via multimodal inputs fail to generalize to unimodal queries. The authors propose DECODE, a method to disentangle and localize modality-specific neurons for more effective knowledge editing.
This paper studies the ability of multimodal large language models (MLLMs) to detect when the correct answer is absent in video understanding tasks, finding that models systematically fail by selecting plausible distractors instead of recognizing no valid option exists. The failure worsens in temporal reasoning and dense frame sampling, and chain-of-thought prompting only partially mitigates the issue.
Introduces SynCred-Bench, a benchmark of 600 AI-generated misinformation images across six credible-form categories, showing that existing detectors (including MLLMs, open-source AIGC detectors, and commercial APIs) perform poorly, with human annotators also struggling.
Introduces ESI-BENCH, a comprehensive benchmark for embodied spatial intelligence built on OmniGibson, covering 10 task categories and 29 subcategories. Experiments show active exploration substantially outperforms passive approaches, with failures mainly due to action blindness rather than perception, revealing a metacognitive gap in models compared to humans.