Tag
FLAT is a representation pre-training framework that jointly optimizes a shared multimodal encoder with downstream decoders for text-to-image and image-to-text tasks, using flexible-length aligned 1D sequences to enable cross-modal retrieval and generation with state-of-the-art results.
MKGR is a multimodal framework that combines protein sequence encoding with four biomedical knowledge graphs to improve cold-start protein-protein interaction prediction, outperforming baselines on benchmark datasets.
Alibaba researchers propose AFMRL, a two-stage framework that uses MLLMs to extract product attributes and enhance fine-grained multimodal representation learning for e-commerce retrieval tasks.