Tag
This paper introduces a contrastive modeling framework for multimodal in-context learning to improve reasoning path alignment, enhancing performance on tasks like visual question answering.