Tag
The paper proposes δ-Vision, a method to reduce computational overhead in multimodal language models by using lightweight low-rank adapters to reconstruct visual states efficiently without discarding visual tokens.
This paper explores calibrated ambiguity as a generative resource in human communication versus multimodal language models, using the Dixit game to show that AI exhibits ambiguity collapse and lacks cultural references compared to humans.
This paper introduces SPACE, the first source-free unlearning framework for multimodal large language models (MLLMs), which uses text-guided proxy anchor selection and dual-constraint semantic isolation to erase target concepts without requiring access to original training data, achieving performance comparable to data-dependent methods.
Researchers introduce the MM-OCEAN dataset and a three-tier evaluation framework for grounded personality reasoning in multimodal LLMs, revealing a 'Prejudice Gap' where models often make correct predictions without proper grounding.
SpaceDG is a large-scale dataset and benchmark that evaluates multimodal language models' spatial reasoning robustness under visual degradations like motion blur and low light, revealing significant performance gaps and showing that fine-tuning on SpaceDG improves robustness without degrading clean image performance.