MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

arXiv cs.LG Papers

Summary

This paper introduces MAG, a manifold-guided framework for semi-supervised multi-modal in-context demonstration selection, leveraging unlabeled data to improve few-shot ICL for MLLMs. Experiments on eight benchmarks show consistent gains in label-scarce regimes.

arXiv:2608.12724v1 Announce Type: new Abstract: Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them for ICL. We propose MAG (MAnifold-Guided semi-supervised in-context demonstra- tion selection), an efficient framework that leverages unlabeled data to improve multi-modal ICL. MAG formulates demonstration selection as a semi-supervised propagation problem on a multi-modal graph and adopts a two-stage strategy: (i) relevance score propagation identifies a compact set of high-impact unlabeled samples for pseudo-labeling, reducing MLLM inference cost; (ii) multi-modal relevance is used to select the final demonstrations. We show that textual represen- tations are more effective for relevance propagation, while both visual and textual modalities are crucial for high-quality demonstration selection. Experiments on eight multi-modal benchmarks demonstrate that MAG consistently outperforms strong baselines in label-scarce regimes, achieving significant gains with a limited pseudo-labeling budget.
Original Article
View Cached Full Text

Cached at: 08/14/26, 09:31 AM

# MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning
Source: [https://arxiv.org/html/2608.12724](https://arxiv.org/html/2608.12724)
Xun XuAffiliation:Institute for Infocomm Research, A⁢STAR, Singaporezirui\_c@u\.nus\.eduTiankai ChenAffiliation:Institute for Infocomm Research, A⁢STAR, Singaporezirui\_c@u\.nus\.eduFady RezkAffiliation:Institute for Infocomm Research, A⁢STAR, Singaporezirui\_c@u\.nus\.eduBowen ZhengXiaodong ShiShijie LiAffiliation:Institute for Infocomm Research, A⁢STAR, Singaporezirui\_c@u\.nus\.eduKangkang LuAffiliation:Institute for Infocomm Research, A⁢STAR, Singaporezirui\_c@u\.nus\.eduBharadwaj VeeravalliNancy F\. ChenAffiliation:Institute for Infocomm Research, A⁢STAR, Singaporezirui\_c@u\.nus\.edu\[3pt\] College of DesignEngineeringNational University of Singapore

###### Abstract

Few\-shot in\-context learning \(ICL\) with multi\-modal large language models \(MLLMs\) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations\. While unlabeled multi\-modal data is abundant, it remains elusive how to exploit them for ICL\. We proposeMAG\(MAnifold\-Guided semi\-supervised in\-context demonstration selection\), an efficient framework that leverages unlabeled data to improve multi\-modal ICL\. MAG formulates demonstration selection as a semi\-supervised propagation problem on a multi\-modal graph and adopts a two\-stage strategy: \(i\) relevance score propagation identifies a compact set of high\-impact unlabeled samples for pseudo\-labeling, reducing MLLM inference cost; \(ii\) multi\-modal relevance is used to select the final demonstrations\. We show that textual representations are more effective for relevance propagation, while both visual and textual modalities are crucial for high\-quality demonstration selection\. Experiments on eight multi\-modal benchmarks demonstrate that MAG consistently outperforms strong baselines in label\-scarce regimes, achieving significant gains with a limited pseudo\-labeling budget\.

## 1Introduction

Large language models \(LLMs\) and their multi\-modal extensions \(MLLMs\) have shown strong in\-context learning \(ICL\) abilities, allowing adaptation to new tasks by conditioning on a small set of input–output demonstrations, e\.g\., visual samples, textual queries, and answer tuples, without parameter updates\([4](https://arxiv.org/html/2608.12724#bib.bib19);[1](https://arxiv.org/html/2608.12724#bib.bib30);[26](https://arxiv.org/html/2608.12724#bib.bib32);[29](https://arxiv.org/html/2608.12724#bib.bib31)\)\. In multi\-modal settings, ICL is especially attractive, as it enables MLLMs to address diverse vision–language tasks through prompt construction alone\. However, the effectiveness of few\-shot multi\-modal ICL critically depends on the quality, relevance, and coverage of the demonstrations, making their construction both central and challenging\([46](https://arxiv.org/html/2608.12724#bib.bib9);[38](https://arxiv.org/html/2608.12724#bib.bib2);[8](https://arxiv.org/html/2608.12724#bib.bib8);[5](https://arxiv.org/html/2608.12724#bib.bib38)\)\.

In practice, such high\-quality input–output demonstrations, i\.e\., labeled data, are often scarce\. For instance, new users may have limited historical interactions, or insufficient labeled data may be available to bootstrap a new downstream task\. This inherent label scarcity substantially limits the effectiveness of ICL\([19](https://arxiv.org/html/2608.12724#bib.bib3);[34](https://arxiv.org/html/2608.12724#bib.bib12)\)\. In contrast, it is typically much easier to acquire large amounts of unlabeled multi\-modal data, such as visual samples paired with textual queries but without answers\. Yet, in the absence of ground\-truth outputs, these unlabeled samples are difficult to exploit using standard retrieval\-based ICL methods\([38](https://arxiv.org/html/2608.12724#bib.bib2);[35](https://arxiv.org/html/2608.12724#bib.bib24)\)\. This tension naturally motivates semi\-supervised learning which has long shown that unlabeled data can significantly improve generalization by capturing intrinsic data structure that is inaccessible from limited labeled examples\([47](https://arxiv.org/html/2608.12724#bib.bib39);[37](https://arxiv.org/html/2608.12724#bib.bib40);[15](https://arxiv.org/html/2608.12724#bib.bib35);[36](https://arxiv.org/html/2608.12724#bib.bib41)\)\. These insights suggest that abundant unlabeled multi\-modal data, if properly utilized, could serve as a powerful resource for constructing more informative and robust in\-context demonstrations\.

Figure 1:Comparison of our proposed semi\-supervised in\-context learning \(ICL\) against vanilla ICL\.A central challenge is how to effectively propagate information from labeled data to unlabeled data in a manner compatible with ICL\. We address this challenge through a two\-stage framework designed to balance scalability with demonstration quality\.

First of all, given the pool of unlabeled multi\-modal samples and the high inference cost of MLLMs, directly pseudo\-labeling all unlabeled data is too expensive\. Instead, we employ a graph\-based relevance score propagation mechanism\([48](https://arxiv.org/html/2608.12724#bib.bib33);[44](https://arxiv.org/html/2608.12724#bib.bib34)\)to identify a compact candidate set of unlabeled samples that are most relevant to the labeled data\. This step propagates supervision along the intrinsicdata manifold, enabling efficient filtering of unlabeled samples that are likely to provide informative and transferable context\. Importantly, the procedure is formulated as label propagation on the graph which is both highly computational efficient with closed\-form solution and offers principled mathematical explanation\. The approach is also highly scalable to very large unlabeled data pool due to the sparsity of the manifold graph\. Eventually, only selected unlabeled candidates are then forwarded to the MLLM for pseudo\-labeling, reducing inference cost while preserving semantic relevance\. In the second stage, we further select the most effective in\-context demonstrations from the combined labeled and pseudo\-labeled candidate pool\. Specifically, we apply score propagation again by treating the query sample as positive node and propagate relevance score to all candidates for demonstration selection\.

Critically, we reveal that relying on textual descriptions of visual content is a stable option for relevance propagation in the first stage for coarse filtering\. In contrast, incorporating both visual and textual modalities is essential for high\-quality demonstration selection in the second stage to ensure fine\-grained alignment between the demonstrations and the query\. In particular, late fusion between visual and textual modalities proves more effective in the second stage, as it better synergizes complementary modalities while decoupling modality\-specific features\.

To reflect theMAnifold\-Guided nature of our semi\-supervised in\-context demonstration selection, we term the proposed methodMAG\. Unlike learning\-based approaches\([12](https://arxiv.org/html/2608.12724#bib.bib13);[23](https://arxiv.org/html/2608.12724#bib.bib14)\)that rely on fine\-tuning or training auxiliary models, MAG is learning\-free, enabling rapid deployment across diverse tasks\. Across eight benchmarks spanning visual emotion recognition\([40](https://arxiv.org/html/2608.12724#bib.bib42);[30](https://arxiv.org/html/2608.12724#bib.bib43)\), scene text understanding\([49](https://arxiv.org/html/2608.12724#bib.bib44)\), visual reasoning\([7](https://arxiv.org/html/2608.12724#bib.bib45)\), and visual question answering\([14](https://arxiv.org/html/2608.12724#bib.bib46);[28](https://arxiv.org/html/2608.12724#bib.bib47)\), MAG consistently outperforms methods retrieving from labeled data alone\. These results demonstrate that principled incorporation of unlabeled data through relevance\-guided pseudo\-labeling and multi\-modal graph propagation substantially improves ICL performance under label\-scarce conditions, with particularly pronounced gains on tasks requiring complex cross\-modal reasoning\.

We summarize the contributions as follows\.

- •We identify label scarcity as a key bottleneck in few\-shot multi\-modal ICL and propose to leverage abundant unlabeled multi\-modal data via a semi\-supervised formulation, enabling more robust and informative demonstration construction under limited labeled data\.
- •We introduce a two\-stage, training\-free framework that employs graph\-based relevance score propagation to efficiently select a compact, query\-relevant subset of unlabeled samples for pseudo\-labeling and ICL, substantially reducing MLLM inference cost while preserving semantic relevance\.
- •Across 8 diverse benchmarks, we show consistent and substantial gains with particularly large improvements on reasoning\-intensive tasks, demonstrating that unlabeled data can be effectively leveraged for MAG\.

## 2Related Work

In\-Context Learning \(ICL\): ICL enables large language models to solve new tasks by conditioning on demonstrations without parameter updates\([4](https://arxiv.org/html/2608.12724#bib.bib19)\)\. Research has explored ICL mechanisms as implicit gradient descent\([31](https://arxiv.org/html/2608.12724#bib.bib20);[13](https://arxiv.org/html/2608.12724#bib.bib21)\)or Bayesian inference\([39](https://arxiv.org/html/2608.12724#bib.bib22)\), while other work examines how demonstration selection, ordering, and formatting affect performance\([27](https://arxiv.org/html/2608.12724#bib.bib23);[35](https://arxiv.org/html/2608.12724#bib.bib24);[2](https://arxiv.org/html/2608.12724#bib.bib1)\)\. Methods for improving ICL fall into learning\-free and learning\-based categories\. Learning\-based approaches optimize prompts or selection strategies through training\. Examples include automatic prompt engineering\([45](https://arxiv.org/html/2608.12724#bib.bib25)\), gradient\-based prompt optimization\([32](https://arxiv.org/html/2608.12724#bib.bib26)\), and parameter\-efficient adaptation\([25](https://arxiv.org/html/2608.12724#bib.bib5)\)\. CEIL\([41](https://arxiv.org/html/2608.12724#bib.bib6)\)learns selection policies using determinant point processes\. While effective, these methods require additional training or task\-specific supervision, unlike our learning\-free approach\. Learning\-free methods use heuristics for demonstration selection: TopK\+MDL uses content similarity and minimum description length\([38](https://arxiv.org/html/2608.12724#bib.bib2)\), while Cover\-LS emphasizes structural similarity and diversity\([19](https://arxiv.org/html/2608.12724#bib.bib3)\)\. MAPLE\([10](https://arxiv.org/html/2608.12724#bib.bib4)\)introduces graph\-based influence propagation to select demonstrations and incorporates unlabeled data through adaptive pseudo\-labeling for text\-only ICL\. Unlike MAPLE, which addresses many\-shot pseudo\-labeling for in\-context learning with large context windows, our work targets few\-shot multi\-modal ICL by introducing a two\-stage, modality\-aware score propagation that efficiently filters unlabeled data and selects high\-quality demonstrations under strict context length and inference cost constraints\. The introduced score propagation presents a more computational efficient and mathematically principled perspective\. We also identify effective multi\-modal fusion practices in the context of ICL\.

Multi\-Modal In\-Context Learning: Vision\-language models\([33](https://arxiv.org/html/2608.12724#bib.bib27);[18](https://arxiv.org/html/2608.12724#bib.bib28);[21](https://arxiv.org/html/2608.12724#bib.bib29)\)established shared representation spaces, enabling multi\-modal ICL with models like Flamingo\([1](https://arxiv.org/html/2608.12724#bib.bib30)\), GPT\-4V\([29](https://arxiv.org/html/2608.12724#bib.bib31)\), and LLaVA\([26](https://arxiv.org/html/2608.12724#bib.bib32)\)\. However, MM\-ICL performance remains highly sensitive to demonstration selection\([22](https://arxiv.org/html/2608.12724#bib.bib7);[8](https://arxiv.org/html/2608.12724#bib.bib8)\)\. Learning\-free MM\-ICL methods address this through improved retrieval\. MMICES\([8](https://arxiv.org/html/2608.12724#bib.bib8)\)and VICL\([46](https://arxiv.org/html/2608.12724#bib.bib9)\)combine visual concept consistency with textual similarity, while ICCG\([20](https://arxiv.org/html/2608.12724#bib.bib10)\)uses diversity\-coverage matching\. MPCAR\([34](https://arxiv.org/html/2608.12724#bib.bib12)\)incorporates multi\-perspective demonstrations, and Cola\([6](https://arxiv.org/html/2608.12724#bib.bib11)\)coordinates multiple VLMs\. However, these methods operate only on labeled demonstrations and cannot leverage abundant unlabeled data, limiting their effectiveness in label\-scarce settings\. Learning\-based methods adapt MLLMs through training\. STIC\([12](https://arxiv.org/html/2608.12724#bib.bib13)\)uses self\-training for image comprehension, TACO\([23](https://arxiv.org/html/2608.12724#bib.bib14)\)employs task\-aware attention, CVR\-LLM\([24](https://arxiv.org/html/2608.12724#bib.bib15)\)generates context\-aware textual descriptions, and CONTEXTNA\([24](https://arxiv.org/html/2608.12724#bib.bib15)\)proposes an agentic retrieval framework\. These methods require training or complex pipelines, while our approach is learning\-free and directly applicable across tasks\.

Semi\-Supervised Learning and Label Propagation: Label propagation diffuses label information from labeled to unlabeled samples over similarity graphs\([48](https://arxiv.org/html/2608.12724#bib.bib33);[44](https://arxiv.org/html/2608.12724#bib.bib34)\), assuming nearby samples share labels\. Modern variants leverage deep representations: Iscen et al\.\([16](https://arxiv.org/html/2608.12724#bib.bib16)\)proposed efficient diffusion on region manifolds, while subsequent work integrated pseudo\-labeling into deep learning\([15](https://arxiv.org/html/2608.12724#bib.bib35);[3](https://arxiv.org/html/2608.12724#bib.bib36);[43](https://arxiv.org/html/2608.12724#bib.bib37)\)\. More expressive approaches use hypergraph structures\([42](https://arxiv.org/html/2608.12724#bib.bib17)\)to model higher\-order relations\. Chen et al\.\([9](https://arxiv.org/html/2608.12724#bib.bib18)\)propose graph\-based pseudo\-labeling for text\-only ICL, propagating labels over example\-query relations\. To the best of our knowledge, no prior work has applied label propagation paradigms to multi\-modal ICL\. We build on these advances by extending influence\-guided pseudo\-labeling to multi\-modal ICL and demonstrating that visual and textual graphs capture complementary structures requiring late fusion\.

## 3Methodology

![Refer to caption](https://arxiv.org/html/2608.12724v1/framework.png)Figure 2:Overview of MAG\. Stage 1 \(top\): We generate textual descriptions of all images using the MLLM, construct a textual relationship graph, identify high\-relevance unlabeled samples, and pseudo\-label them\. Stage 2 \(bottom\): We construct visual and textual graphs over the expanded candidate pool and perform query\-conditioned relevance propagation in both modalities to select demonstrations for each query\.#### Problem Setup\.

We are given a labeled dataset𝒟L=\{\(xi,qi,ai\)\}i=1NL\\mathcal\{D\}\_\{L\}=\\\{\(x\_\{i\},q\_\{i\},a\_\{i\}\)\\\}\_\{i=1\}^\{N\_\{L\}\}and an unlabeled dataset𝒟U=\{\(xj,qj\)\}j=1NU\\mathcal\{D\}\_\{U\}=\\\{\(x\_\{j\},q\_\{j\}\)\\\}\_\{j=1\}^\{N\_\{U\}\}, wherexix\_\{i\},qiq\_\{i\}, andaia\_\{i\}denote the image, text \(e\.g\., a question in VQA\), and answer, respectively\. In typical semi\-supervised settings,NL≪NUN\_\{L\}\\ll N\_\{U\}\. Our goal is to select effective in\-context demonstrations for each queryqqunder label\-scarce conditions\.

MAG addresses this problem through two components: \(1\)manifold\-guided pseudo\-labeling, which identifies and pseudo\-labels high\-relevance unlabeled samples to expand the candidate pool while controlling cost; and \(2\)query\-conditioned multi\-modal graph\-based demonstration selection, which performs relevance propagation on complementary visual and textual graphs to select most relevant demonstrations from the candidate pool\.

### 3\.1Manifold\-Guided Pseudo\-Labeling

In the first stage, we identify unlabeled samples that are strongly connected to labeled data and pseudo\-label them to augment the limited labeled pool\. To this end, we construct a textual relationship graph that captures semantic structure shared across labeled and unlabeled samples\.

Textual Description Generation\.Contrary to the practice in MAPLE[10](https://arxiv.org/html/2608.12724#bib.bib4)which directly uses questionqiq\_\{i\}to compute affinity, we generate textual descriptions of image using the MLLM,di=ℳdesc​\(xi\)d\_\{i\}=\\mathcal\{M\}\_\{\\mathrm\{desc\}\}\(x\_\{i\}\), whereℳdesc\\mathcal\{M\}\_\{\\mathrm\{desc\}\}produces a detailed description of the visual content\. These descriptions, combined with any existing textual information \(e\.g\., task\-specific questions\), form the textual representationti=\[di,qi\]t\_\{i\}=\[d\_\{i\},q\_\{i\}\]\. This design ensures more visual information is captured in the construction of graph affinity and enables handling the tasks where questions are identical across samples \(e\.g\. many VQA tasks have an identical question for different images\)\.

Manifold Graph Construction\.We construct a textual relationship graph𝒢\(t\)=\(𝒱\(t\),ℰ\(t\)\)\\mathcal\{G\}^\{\(t\)\}=\(\\mathcal\{V\}^\{\(t\)\},\\mathcal\{E\}^\{\(t\)\}\)over𝒟L∪𝒟U\\mathcal\{D\}\_\{L\}\\cup\\mathcal\{D\}\_\{U\}\. Each node corresponds to a text embeddinghi\(t\)=ft​\(ti\)h\_\{i\}^\{\(t\)\}=f\_\{t\}\(t\_\{i\}\), whereftf\_\{t\}is a pre\-trained text encoder \(e\.g\., Contriver\([17](https://arxiv.org/html/2608.12724#bib.bib49)\)\)\. We build akk\-NN graph and symmetrize it to capture the underlying manifold, with edge weights defined as

wi​j\(t\)=\{si​j\(t\),if​j∈NNk​\(i\),0,otherwise,si​j\(t\)=⟨hi\(t\),hj\(t\)⟩\.w\_\{ij\}^\{\(t\)\}=\\begin\{cases\}s\_\{ij\}^\{\(t\)\},&\\text\{if \}j\\in\\mathrm\{NN\}\_\{k\}\(i\),\\\\ 0,&\\text\{otherwise\},\\end\{cases\}\\quad s\_\{ij\}^\{\(t\)\}=\\langle h\_\{i\}^\{\(t\)\},h\_\{j\}^\{\(t\)\}\\rangle\.\(1\)
Relevance Score Propagation\.To identify candidate unlabeled samples for pseudo\-labeling, we propagate relevance scores from labeled nodes to unlabeled ones\. We initialize an relevance score vectorI\(0\)\(t\)∈\{0,1\}NL\+NUI^\{\(t\)\}\_\{\(0\)\}\\in\\\{0,1\\\}^\{N\_\{L\}\+N\_\{U\}\}by assigning value11to labeled samples and00to unlabeled samples\. Relevance is iteratively propagated over𝒢\(t\)\\mathcal\{G\}^\{\(t\)\}via

I\(k\)\(t\)=α​W^\(t\)​I\(k−1\)\(t\)\+\(1−α\)​I\(0\)\(t\),I^\{\(t\)\}\_\{\(k\)\}=\\alpha\\hat\{W\}^\{\(t\)\}I^\{\(t\)\}\_\{\(k\-1\)\}\+\(1\-\\alpha\)I^\{\(t\)\}\_\{\(0\)\},\(2\)wherekkindicates the propagation step,α∈\[0,1\]\\alpha\\in\[0,1\]controls the propagation strength andW^\(t\)=\(D\(t\)\)−1/2W\(t\)\(D\(t\)\)−1/2\\hat\{W\}^\{\(t\)\}=\(D^\{\(t\)\}\)^\{\-1/2\}W^\{\(t\)\}\(D^\{\(t\)\}\)^\{\-1/2\}is the normalized adjacency matrix withD\(t\)D^\{\(t\)\}being the degree matrix\. The iteration converges to the closed\-form solution

I\(∞\)\(t\)=\(1−α\)​\(𝐈−α​W^\(t\)\)−1​I\(0\)\(t\)\.I^\{\(t\)\}\_\{\(\\infty\)\}=\(1\-\\alpha\)\(\\mathbf\{I\}\-\\alpha\\hat\{W\}^\{\(t\)\}\)^\{\-1\}I^\{\(t\)\}\_\{\(0\)\}\.\(3\)
We select the top\-KKunlabeled samples with the highest relevance scores and pseudo\-label them using the MLLM conditioned on labeled demonstrations:

a^j=ℳ⁡\(xj,qj,𝒟L\)\.\\hat\{a\}\_\{j\}=\\mathcal\{M\}\(x\_\{j\},q\_\{j\};\\mathcal\{D\}\_\{L\}\)\.\(4\)The pseudo\-labeled set𝒟P\\mathcal\{D\}\_\{P\}and labeled set𝒟L\\mathcal\{D\}\_\{L\}together form the expanded candidate pool𝒟E=𝒟L∪𝒟P\\mathcal\{D\}\_\{E\}=\\mathcal\{D\}\_\{L\}\\cup\\mathcal\{D\}\_\{P\}\. The above score propagation is easily scalable to very large unlabeled data pool sinceWWis sparse which allows efficiency solution to Eq\.[3](https://arxiv.org/html/2608.12724#S3.E3), e\.g\. through conjugate gradient or Cholesky decomposition\.

### 3\.2Multi\-Modal Graph Construction & Query\-Conditioned Propagation

In the second stage, we select the most relevant demonstrations for ICL from the expanded candidate pool𝒟E\\mathcal\{D\}\_\{E\}\. We construct separate relationship graphs in the visual and textual embedding spaces\. These practices ensure most relevant demonstrations are selected for ICL while the inference cost is manageable\.

#### Multi\-modal graph construction & Score Propagation\.

For each sample in𝒟E\\mathcal\{D\}\_\{E\}, we compute both visual and textual embeddings,

hi\(v\)=fv​\(xi\),hi\(t\)=ft​\(ti\),h\_\{i\}^\{\(v\)\}=f\_\{v\}\(x\_\{i\}\),\\quad h\_\{i\}^\{\(t\)\}=f\_\{t\}\(t\_\{i\}\),\(5\)wherefvf\_\{v\}is a visual encoder \(e\.g\., CLIP ViT\-L/14\([33](https://arxiv.org/html/2608.12724#bib.bib27)\)\)\. We then construct modality\-specific graphs𝒢\(v\)\\mathcal\{G\}^\{\(v\)\}and𝒢\(t\)\\mathcal\{G\}^\{\(t\)\}usingkk\-NN with cosine similarity following a similar fashion in Eq\.[1](https://arxiv.org/html/2608.12724#S3.E1)\. This results in two graphs featuring both visual𝒢\(v\)\\mathcal\{G\}^\{\(v\)\}and textual𝒢\(t\)\\mathcal\{G\}^\{\(t\)\}modalities\.

Given a queryqq, we initialize modality\-specific relevance vectorsI0\(m\)​\(q\)I\_\{0\}^\{\(m\)\}\(q\)by assigning value11to thekknearest neighbors ofqqin modalitym∈\{v,t\}m\\in\\\{v,t\\\}and00to all other nodes\. Relevance is then propagated following the closed\-form solution:

I\(∞\)\(m\)=\(1−α\)​\(𝐈−α​W^\(m\)\)−1​I\(0\)\(m\)I^\{\(m\)\}\_\{\(\\infty\)\}=\(1\-\\alpha\)\(\\mathbf\{I\}\-\\alpha\\hat\{W\}^\{\(m\)\}\)^\{\-1\}I^\{\(m\)\}\_\{\(0\)\}\(6\)
Demonstration Selection and Inference\. After convergence, we apply late fusion to combine the modality\-specific relevance scores via,

I=β​I\(∞\)\(v\)\+\(1−β\)​I\(∞\)\(t\),I=\\beta I^\{\(v\)\}\_\{\(\\infty\)\}\+\(1\-\\beta\)I^\{\(t\)\}\_\{\(\\infty\)\},\(7\)whereβ∈\[0,1\]\\beta\\in\[0,1\]balances the visual and textual modalities\. For each queryqq, we rank all candidates in𝒟E\\mathcal\{D\}\_\{E\}by the fused relevance scoreIIand select the top\-kkdemonstrations:

D​S​\(q\)=Top​\-​k​\(\{Ii\}xi∈𝒟E\)\.DS\(q\)=\\mathrm\{Top\}\\text\{\-\}k\\left\(\\\{I\_\{i\}\\\}\_\{x\_\{i\}\\in\\mathcal\{D\}\_\{E\}\}\\right\)\.\(8\)The selected demonstrations are then provided to the MLLM for inference:

a^q=ℳ⁡\(xq,qq,D​S​\(q\)\)\.\\hat\{a\}\_\{q\}=\\mathcal\{M\}\(x\_\{q\},q\_\{q\};DS\(q\)\)\.\(9\)
Efficient Incremental Inference\.The static graph construction described above requires attaching query samples to the graph\. Without seeing all query samples, a naive way to incrementally do inference for a new query sample may need to rebuild the graph from scratch\. However, for a graph withN=\|𝒟E\|\+1N=\|\\mathcal\{D\}\_\{E\}\|\+1nodes, this incurs additional computational cost\. To enable efficient inference while still exploiting the manifold structure, we introduce a simple graph update mechanism\. Given an existing modality\-specific graph𝒢\(m\)=\(𝒱\(m\),ℰ\(m\)\)\\mathcal\{G\}^\{\(m\)\}=\(\\mathcal\{V\}^\{\(m\)\},\\mathcal\{E\}^\{\(m\)\}\)with adjacency matrixW\(m\)∈ℝN×NW^\{\(m\)\}\\in\\mathbb\{R\}^\{N\\times N\}, we insert a new query samplexqx\_\{q\}by measuring its similarity to all existing nodes:

sq​i\(m\)=⟨hq\(m\),hi\(m\)⟩,∀i∈\{1,…,\|𝒟E\|\}\.s\_\{qi\}^\{\(m\)\}=\\langle h\_\{q\}^\{\(m\)\},h\_\{i\}^\{\(m\)\}\\rangle,\\quad\\forall i\\in\\\{1,\\dots,\|\\mathcal\{D\}\_\{E\}\|\\\}\.\(10\)
Rather than recomputingkk\-nearest neighbors for all nodes, we only remove the previous query sample from the graph and create edges for the top K nodes and connect them to the new query sample\. This operation only requires calculating the inner product between query and\|𝒟E\|\|\\mathcal\{D\}\_\{E\}\|samples when each new query arrives\.

## 4Experiments

### 4\.1Experimental Setup

Datasets\. We evaluate on eight benchmarks spanning visual emotion recognition, scene text understanding, visual reasoning, and visual question answering\. These tasks require diverse visual\-textual understanding capabilities, from recognizing subtle emotional cues to complex spatial reasoning\.Visual Emotion Recognition:EmoSet\([40](https://arxiv.org/html/2608.12724#bib.bib42)\)contains images annotated with eight emotion categories for emotion classification, while Emotion6\([30](https://arxiv.org/html/2608.12724#bib.bib43)\)provides six basic emotion labels collected via controlled user studies\.Scene Text Understanding:TextOCR\([49](https://arxiv.org/html/2608.12724#bib.bib44)\)evaluates text detection and recognition in natural scene images\.Visual Reasoning:MMStar\([7](https://arxiv.org/html/2608.12724#bib.bib45)\)contains 1,500 challenge samples testing multi\-modal reasoning with reduced language bias\. MatchingMI\([49](https://arxiv.org/html/2608.12724#bib.bib44)\)requires cross\-image matching and correspondence\. CLEVR\([49](https://arxiv.org/html/2608.12724#bib.bib44)\)is a synthetic diagnostic benchmark for compositional reasoning including counting, comparison, and logical operations\.Visual Question Answering:GQA\([14](https://arxiv.org/html/2608.12724#bib.bib46)\)features compositional questions grounded in scene graphs\. OK\-VQA\([28](https://arxiv.org/html/2608.12724#bib.bib47)\)evaluates open\-domain VQA requiring external commonsense knowledge\. For all datasets, we report accuracy as an evaluation metric\.

Competing Methods\. We compare against seven baselines spanning both zero\-shot approaches and few\-shot retrieval methods\.Few\-shot methods:Few\-shot[4](https://arxiv.org/html/2608.12724#bib.bib19)randomly samples labeled demonstrations\.VICL\([46](https://arxiv.org/html/2608.12724#bib.bib9)\)uses retrieval and reranking with language\-based demonstrations\.Top\-K\+MDL\([38](https://arxiv.org/html/2608.12724#bib.bib2)\)ranks candidates by content similarity and minimum description length\.MMICES\([8](https://arxiv.org/html/2608.12724#bib.bib8)\)performs two\-stage selection via visual concept consistency and textual similarity\.CVR\-LLM\([24](https://arxiv.org/html/2608.12724#bib.bib15)\)generates context\-aware visual descriptions through iterative refinement and conducts demonstration selection using multi\-modal features\.MAPLE\([10](https://arxiv.org/html/2608.12724#bib.bib4)\)performs pseudo\-labeling on unlabeled data and selects demonstrations from a graph constructed using textual questions only\.Zero\-shot methods:Zero\-shotuses task instructions without demonstrations\.Cola\-Zero\([6](https://arxiv.org/html/2608.12724#bib.bib11)\)coordinates multiple VLMs via an LLM\.

### 4\.2Implementation Details

ICL Configurations\.For each dataset, we randomly sample 15 labeled examples, 585 unlabeled examples as the candidate pool, and 300 test samples\. From the unlabeled pool, we pseudo\-label the top 45 high\-relevance samples identified in Stage 1, forming an expanded candidate pool of 60 samples \(\|𝒟E\|=15\+45\|\\mathcal\{D\}\_\{E\}\|=15\+45\)\. For few\-shot methods, both MAPLE and MAG select 10 demonstrations from 15 labeled \+ 585 unlabeled examples while all others select 10 demonstrations from 15 labeled ones\.

Text Description Generation\.For all samples, we generate image descriptions using Gemini\-2\.0\-Flash\([11](https://arxiv.org/html/2608.12724#bib.bib48)\)with the following prompt:“Generate a description for the image, taking the question as guidance, elucidating both the visual content and the underlying purpose or intention depicted\. Craft a clear and concise description that integrates details from the image, highlighting visual cues and semantic meaning which may be useful to get the answer\.”These descriptions are used to construct textual representations in both Stage 1 and Stage 2\.

Models and hyperparameters\.We adopt Gemini\-2\.0\-Flash as the base MLLM\. Contriever\([17](https://arxiv.org/html/2608.12724#bib.bib49)\)is used for text encoding \(ftf\_\{t\}\), and CLIP ViT\-L/14\([33](https://arxiv.org/html/2608.12724#bib.bib27)\)for visual encoding \(fvf\_\{v\}\)\. To investigate the generalization of method, we also evaluate against alternative MLLMs, including Gemini\-3\.0\-Flash, GPT\-4o, GPT\-4\.1, GPT\-5 and Qwen models in the appendix\. For graph construction and propagation, we set the number of nearest neighbors tokn​n=5k\_\{nn\}=5, the propagation strength toα=0\.9\\alpha=0\.9, and the modality fusion weight toβ=0\.5\\beta=0\.5\. These hyperparameters are fixed across all experiments for consistency\.

### 4\.3Quantitave Evaluations

Table[1](https://arxiv.org/html/2608.12724#S4.T1)compares MAG against baselines under label\-scarce conditions, reported with average and standard deviation over 3 random runs\. We make the following observations from the results\.

Table 1:Main results on multi\-modal benchmarks under label\-scarce settings with Gemini\-2\.0\-Flash \(mean accuracy±\\pmstandard dev % across 3 runs\)\.Unlabeled data provides substantial gains under label scarcity\.MAG consistently outperforms all baselines across eight benchmarks using only 15 labeled samples\. Compared with the strongest retrieval baseline \(MMICES\), MAG achieves notable improvements on challenging tasks, including \+30\.5% on CLEVR, \+7\.1% on MMStar, and \+25\.1% on TextOCR\. Relative to standard few\-shot prompting, MAG further improves TextOCR \(\+23\.8%\), CLEVR \(\+30\.8%\), and EmoSet \(\+15\.3%\)\. These results demonstrate the effectiveness of leveraging unlabeled data through relevance\-guided pseudo\-labeling and graph propagation\.

Graph\-based propagation outperforms similarity\-based retrieval\.Compared with similarity\-based retrieval methods such as VICL and Top\-K\+MDL, MAG consistently achieves stronger performance by modeling global relational structure among samples\. For instance, MAG surpasses Top\-K\+MDL by \+34\.3% on CLEVR and \+27\.1% on TextOCR, validating that graph connectivity provides more reliable supervision than local similarity alone under limited labeled data\.

Our method complements description\-based approaches\.Although CVR\-LLM achieves competitive performance through iterative description refinement, MAG still improves results on most benchmarks, including \+3\.5% on EmoSet, \+3\.5% on TextOCR, \+6\.7% on GQA, and \+10\.9% on OKVQA\. Unlike CVR\-LLM, which shows unstable cross\-task generalization, MAG maintains consistently strong performance, suggesting that pseudo\-label expansion and graph propagation provide a more robust way to exploit unlabeled data\.

Stronger performance than single modality manifold method\.Compared with the single modality manifold based method \(MAPLE\), we observe consistent improvement across all datasets\. Despite both methods selecting demonstrations from the same labeled \+ unlabeled examples, our method excels by exploiting multiple modalities and principled selection strategy\.

Overall, these results confirm that MAG enables effective multi\-modal ICL by incorporating unlabeled data, achieving strong and stable performance\. Qualitative examples demonstrating improved demonstration selection are provided in Appendix[C](https://arxiv.org/html/2608.12724#A3)\.

### 4\.4Ablation Study

We conduct ablation studies to validate the key components of MAG\. Table[2](https://arxiv.org/html/2608.12724#S4.T2)evaluates three design choices: pseudo\-labeling strategy, demonstration selection method, and modality usage\. We report mean and standard deviation across 3 runs and draw the following observations\.

Manifold\-guided pseudo\-labeling is critical, and naive pseudo\-labeling can be harmful\.Comparing three strategies, we observe a clear trend: \(1\) removing pseudo\-labeling yields 72\.0% on EmoSet, \(2\) adding randomly selected pseudo\-labels provides only marginal gains on EmoSet \(73\.1%\) and even degrades performance on some tasks \(e\.g\., GQA: 54\.7% vs\. 56\.4%\), while \(3\) our manifold\-guided pseudo\-labeling significantly improves performance across all benchmarks \(76\.3% on EmoSet, 91\.4% on TextOCR, 67\.9% on MMStar, 58\.0% on GQA\)\. This indicates that pseudo\-label quality, rather than quantity, is the determining factor\. Naively introducing pseudo\-labels can inject noise and offset potential gains, whereas our relevance\-guided selection effectively identifies high\-value samples, leading to consistent improvements\.

Graph\-based demonstration selection consistently outperforms heuristic strategies\.Compared to random selection, graph\-based selection yields substantial gains \(\+3\.5% on EmoSet, \+2\.5% on TextOCR, \+1\.8% on GQA\)\. Replacing graph propagation with TopK similarity leads to smaller improvements and consistently underperforms the full method\. This suggests that selecting demonstrations purely based on query\-example similarity is insufficient\. Instead, graph propagation captures higher\-order relationships among samples, enabling retrieval of globally relevant and representative demonstrations\. This advantage becomes more pronounced in low\-resource settings where labeled examples are sparse and potentially biased\.

Both modalities are necessary and complementary\.Removing either modality results in significant performance degradation\. Text\-only graphs \(–img\) particularly hurt visually grounded reasoning \(GQA: 44\.2% vs\. 58\.0%\), while visual\-only graphs \(–desc\) fail on semantically intensive tasks \(EmoSet: 60\.2% vs\. 76\.3%, TextOCR: 67\.0% vs\. 91\.4%\)\. Notably, visual\-only selection introduces misleading context due to superficial visual similarity, leading to large drops on language\-heavy tasks\. These results highlight that visual and textual modalities encode complementary structures, and effective demonstration selection requires jointly modeling both\.

Table 2:Ablation study\. We evaluate the contribution of \(1\) relevance\-guided pseudo\-labeling \(MAPLE vs\. Random vs\. None\), \(2\) graph\-based propagation \(Graph vs\. TopK vs\. Random\), and \(3\) multi\-modal representations \(with/without images and descriptions\)\.
### 4\.5Modality Analysis for Stage 1

We further evaluate MAG using different modality representations for Stage 1 pseudo\-label selection\. Table[3](https://arxiv.org/html/2608.12724#S4.T3)shows results for textual\-only, visual\-only, and combined representations\.

Textual representations provide robust cross\-task performance\.Textual\-only achieves best performance on 6 of 8 datasets, particularly excelling on emotion recognition \(EmoSet: 77\.0%, Emotion6: 65\.0%\), scene text understanding \(TextOCR: 91\.7%\), and visual reasoning \(MMStar: 70\.3%, GQA: 62\.3%\)\. This suggests that MLLM\-generated image descriptions capture semantic concepts that are well\-aligned with these task requirements\.

Visual representations excel slightly on spatial reasoning\.Visual\-only achieves marginally higher accuracy on spatially\-intensive tasks \(CLEVR: 92\.3% vs\. 92\.0%, OKVQA: 58\.0% vs\. 56\.7%\), where geometric relationships are critical\. However, the performance gap is small \(\+0\.3% to \+1\.3%\)\.

Choose textual representation only for Stage 1\.Given that textual representations provide strong and stable performance across diverse task types, we use textual\-only for Stage 1 in all experiments\. Importantly, even on datasets where visual representations provide small gains in Stage 1 isolation, our full method with textual Stage 1 still substantially outperforms all baselines \(e\.g\., 92\.0% vs\. 90\.3% baseline on CLEVR, 56\.7% vs\. 54\.0% baseline on OKVQA in Table[1](https://arxiv.org/html/2608.12724#S4.T1)\)\. This demonstrates that Stage 2’s multi\-modal propagation effectively compensates for Stage 1 modality choice, as it incorporates both visual and textual information through complementary graphs\.

Table 3:Performance using different Stage 1 modalities for pseudo\-label selection \(accuracy %\)\. Textual representations provide robust performance across tasks\. Visual representations achieve slightly higher accuracy on spatially\-intensive tasks \(CLEVR, OKVQA\) but the gap is small \(\+0\.3% to \+1\.3%\)\. We use textual representations as default for consistency\.
### 4\.6Multi\-Modal Fusion Strategy

We compare three modality fusion strategies at Stage 2 for computing relevance score: early fusion \(concatenate the textual embeddings and visual embeddings before graph construction\), mid fusion \(utilize the weighted sum of adjacency matrix for relevance propagation\), and late fusion \(do relevance propagation of two modalities separately, then average these two relevance scores\)\. Table[4](https://arxiv.org/html/2608.12724#S4.T4)shows that late fusion consistently achieves the best performance across all benchmarks, with particularly large gains on reasoning\-intensive tasks \(MMStar: \+6\.6% over early fusion, GQA: \+2\.3%\)\. Early fusion performs worst, suggesting that directly mixing heterogeneous visual and textual features before relational modeling weakens modality\-specific structure learning\. Late fusion preserves complementary relational signals by allowing each modality’s graph to capture its own semantic structure before combining scores, which is particularly beneficial for tasks requiring cross\-modal reasoning\.

Table 4:Performance using different fusion strategies for combining visual and textual relevance scores \(accuracy %\)\. Late fusion consistently outperforms early and mid fusion by preserving modality\-specific graph structures during propagation\.
### 4\.7Hyperparameter Sensitivity

Figure 3:Sensitivity analysis across three datasets \(EmoSet, Emotion6, MMStar\)\. Performance is stable across wide ranges ofα\\alpha,β\\beta,kn​nk\_\{nn\}, pool size, and label rate, with clear optimal configurations:α≈0\.9\\alpha\\approx 0\.9,kn​n=5k\_\{nn\}=5,β=0\.5\\beta=0\.5, pool size of 60, label rate of 25%, and 10 demonstrations\.We evaluate MAG’s sensitivity to key hyperparameters\. Figure[3](https://arxiv.org/html/2608.12724#S4.F3)shows results across EmoSet, Emotion6, and MMStar for \(a\) propagation strength \(α\\alpha\), \(b\) modality fusion weight \(β\\beta\), \(c\) nearest neighbors \(kn​nk\_\{nn\}\), \(d\) candidate pool size \(\|𝒟E\|\|\\mathcal\{D\}\_\{E\}\|\), \(e\) label rate within candidate pool size \(\|𝒟L\|/\|𝒟E\|\|\\mathcal\{D\}\_\{L\}\|/\|\\mathcal\{D\}\_\{E\}\|\), and \(f\) number of demonstrations \(\|D​S​\(q\)\|\|DS\(q\)\|\)\. Performance remains stable across wide ranges for most parameters, with clear optimal configurations emerging:α≈0\.9\\alpha\\approx 0\.9\(emphasizing propagation over initialization\),kn​n=5k\_\{nn\}=5\(moderate neighborhood size\),β=0\.5\\beta=0\.5\(balanced modality fusion\), pool size of 60 \(15 labeled \+ 45 pseudo\-labeled\), label rate of 25% \(demonstrating label efficiency\), and 10 demonstrations \(beyond which performance plateaus or degrades\)\. Notably, the label rate analysis shows that performance peaks at only 25% labeled data, validating that our relevance\-guided pseudo\-labeling effectively leverages unlabeled samples\. The framework demonstrates strong robustness to hyperparameter choices while maintaining consistent optimal configurations across diverse tasks\.

## 5Conclusion

We introduced MAG, a semi\-supervised framework for multi\-modal in\-context learning that effectively leverages unlabeled data under label\-scarce settings\. By modeling demonstration selection as relevance propagation on multi\-modal graphs, MAG identifies a compact set of high\-impact unlabeled samples for pseudo\-labeling and selects query\-conditioned demonstrations that are globally relevant across visual and textual modalities\. This design enables efficient use of unlabeled data while respecting inference\-time and context\-length constraints\. Extensive experiments across eight diverse multi\-modal benchmarks show that MAG consistently outperforms strong few\-shot and semi\-supervised baselines, achieving substantial gains with a limited pseudo\-labeling budget\. Our analysis highlights the complementary roles of textual and visual manifolds in ICL and demonstrates that principled graph\-based selection is key to scalable multi\-modal prompting\. We hope MAG provides a practical foundation for exploiting unlabeled data in future multi\-modal ICL systems\.

## References

- Alayracet al\.\(2022\)J\. Alayrac, J\. Donahue, P\. Luc, A\. Miech, I\. Barr, Y\. Hasson, K\. Lenc, A\. Mensch, K\. Millican, M\. Reynolds,et al\.Flamingo: a visual language model for few\-shot learning\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p1.1),[§2](https://arxiv.org/html/2608.12724#S2.p2.1)\.
- Anet al\.\(2023\)S\. An, Z\. Lin, Q\. Fu, B\. Chen, N\. Zheng, J\. Lou, and D\. ZhangHow do in\-context examples affect compositional generalization?\.InAnnual Meeting of the Association for Computational Linguistics,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p1.1)\.
- Basak and Yin \(2023\)H\. Basak and Z\. YinPseudo\-label guided contrastive learning for semi\-supervised medical image segmentation\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p3.1)\.
- Brownet al\.\(2020\)T\. Brown, B\. Mann, N\. Ryder,et al\.Language models are few\-shot learners\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p1.1),[§2](https://arxiv.org/html/2608.12724#S2.p1.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.12724#S4.T1.2.1.6.1.1)\.
- Chenet al\.\(2025a\)C\. Chen, Y\. Zhai, Y\. Zhao, J\. Gao, B\. Ding, and J\. LiProvoking multi\-modal few\-shot lvlm via exploration\-exploitation in\-context learning\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p1.1)\.
- Chenet al\.\(2023\)L\. Chen, B\. Li, S\. Shen, J\. Yang, C\. Li, K\. Keutzer, T\. Darrell, and Z\. LiuLarge language models are visual reasoning coordinators\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p2.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.12724#S4.T1.2.1.4.1.1)\.
- Chenet al\.\(2024\)L\. Chen, J\. Li, X\. Dong, P\. Zhang, Y\. Zang, Z\. Chen, H\. Duan, J\. Wang, Y\. Qiao, D\. Lin,et al\.Are we on the right way for evaluating large vision\-language models?\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p6.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p1.1)\.
- Chenet al\.\(2025b\)S\. Chen, Z\. Han, B\. He, J\. Liu, M\. Buckley, Y\. Qin, P\. Torr, V\. Tresp, and J\. GuCan multimodal large language models truly perform multimodal in\-context learning?\.InIEEE/CVF Winter Conference on Applications of Computer Vision,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p1.1),[§2](https://arxiv.org/html/2608.12724#S2.p2.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.12724#S4.T1.2.1.9.1.1)\.
- Chenet al\.\(2025c\)Z\. Chen, S\. Wang, X\. Fu, C\. Shi, Z\. Lei, C\. Shen, and J\. LiFrom cross\-task examples to in\-task prompts: a graph\-based pseudo\-labeling framework for in\-context learning\.InConference on Empirical Methods in Natural Language Processing,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p3.1)\.
- Chenet al\.\(2025d\)Z\. Chen, S\. Wang, Z\. Tan, J\. Li, and C\. ShenMAPLE: many\-shot adaptive pseudo\-labeling for in\-context learning\.InInternational Conference on Machine Learning,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p1.1),[§3\.1](https://arxiv.org/html/2608.12724#S3.SS1.p2.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.12724#S4.T1.2.1.11.1.1)\.
- Comaniciet al\.\(2025\)G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein, O\. Ram, D\. Zhang, E\. Rosen,et al\.Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.arXiv preprint arXiv:2507\.06261\.Cited by:[§4\.2](https://arxiv.org/html/2608.12724#S4.SS2.p2.1)\.
- Denget al\.\(2024\)Y\. Deng, P\. Lu, F\. Yin, Z\. Hu, S\. Shen, Q\. Gu, J\. Y\. Zou, K\. Chang, and W\. WangEnhancing large vision language models with self\-training on image comprehension\.Advances in Neural Information Processing Systems\.Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p6.1),[§2](https://arxiv.org/html/2608.12724#S2.p2.1)\.
- Garget al\.\(2022\)S\. Garg, D\. Tsipras, P\. S\. Liang, and G\. ValiantWhat can transformers learn in\-context? a case study of simple function classes\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p1.1)\.
- Hudson and Manning \(2019\)D\. A\. Hudson and C\. D\. ManningGqa: a new dataset for real\-world visual reasoning and compositional question answering\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p6.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p1.1)\.
- Iscenet al\.\(2019\)A\. Iscen, G\. Tolias, Y\. Avrithis, and O\. ChumLabel propagation for deep semi\-supervised learning\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p2.1),[§2](https://arxiv.org/html/2608.12724#S2.p3.1)\.
- Iscenet al\.\(2017\)A\. Iscen, G\. Tolias, Y\. Avrithis, T\. Furon, and O\. ChumEfficient diffusion on region manifolds: recovering small objects with compact cnn representations\.InIEEE conference on computer vision and pattern recognition,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p3.1)\.
- Izacardet al\.\(2021\)G\. Izacard, M\. Caron, L\. Hosseini, S\. Riedel, P\. Bojanowski, A\. Joulin, and E\. GraveUnsupervised dense information retrieval with contrastive learning\.arXiv preprint arXiv:2112\.09118\.Cited by:[§3\.1](https://arxiv.org/html/2608.12724#S3.SS1.p3.1),[§4\.2](https://arxiv.org/html/2608.12724#S4.SS2.p3.1)\.
- Jiaet al\.\(2021\)C\. Jia, Y\. Yang, Y\. Xia, Y\. Chen, Z\. Parekh, H\. Pham, Q\. Le, Y\. Sung, Z\. Li, and T\. DuerigScaling up visual and vision\-language representation learning with noisy text supervision\.InInternational Conference on Machine Learning,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p2.1)\.
- Levyet al\.\(2023\)I\. Levy, B\. Bogin, and J\. BerantDiverse demonstrations improve in\-context compositional generalization\.InAnnual Meeting of the Association for Computational Linguistics,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p2.1),[§2](https://arxiv.org/html/2608.12724#S2.p1.1)\.
- Liet al\.\(2024a\)C\. Li, C\. Jing, Z\. Li, M\. Zhai, Y\. Wu, and Y\. JiaIn\-context compositional generalization for large vision\-language models\.InConference on Empirical Methods in Natural Language Processing,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p2.1)\.
- Li and et al\. \(2022\)J\. Li and et al\.BLIP: bootstrapped language\-image pre\-training for unified vision\-language understanding and generation\.InInternational Conference on Machine Learning,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p2.1)\.
- Liet al\.\(2024b\)L\. Li, J\. Peng, H\. Chen, C\. Gao, and X\. YangHow to configure good in\-context sequence for visual question answering\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p2.1)\.
- Liet al\.\(2025\)Y\. Li, J\. Yang, T\. Yun, P\. Feng, J\. Huang, and R\. TangTaco: enhancing multimodal in\-context learning via task mapping\-guided sequence configuration\.InConference on Empirical Methods in Natural Language Processing,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p6.1),[§2](https://arxiv.org/html/2608.12724#S2.p2.1)\.
- Liet al\.\(2024c\)Z\. Li, D\. Liu, C\. Zhang, H\. Wang, T\. Xue, and W\. CaiEnhancing advanced visual reasoning ability of large language models\.InConference on Empirical Methods in Natural Language Processing,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p2.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.12724#S4.T1.2.1.10.1.1)\.
- Liuet al\.\(2022\)H\. Liu, D\. Tam, M\. Muqeeth, J\. Mohta, T\. Huang, M\. Bansal, and C\. A\. RaffelFew\-shot parameter\-efficient fine\-tuning is better and cheaper than in\-context learning\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p1.1)\.
- Liuet al\.\(2023\)H\. Liu, C\. Li, Q\. Wu, and Y\. J\. LeeVisual instruction tuning\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p1.1),[§2](https://arxiv.org/html/2608.12724#S2.p2.1)\.
- Luet al\.\(2022\)Y\. Lu, M\. Bartolo, A\. Moore, S\. Riedel, and P\. StenetorpFantastically ordered prompts and where to find them: overcoming few\-shot prompt order sensitivity\.InAnnual Meeting of the Association for Computational Linguistics,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p1.1)\.
- Marinoet al\.\(2019\)K\. Marino, M\. Rastegari, A\. Farhadi, and R\. MottaghiOK\-vqa: a visual question answering benchmark requiring external knowledge\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p6.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p1.1)\.
- OpenAI \(2023\)OpenAIGPT\-4 technical report\.arXiv preprint arXiv:2303\.08774\.Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p1.1),[§2](https://arxiv.org/html/2608.12724#S2.p2.1)\.
- Pandaet al\.\(2018\)R\. Panda, J\. Zhang, H\. Li, J\. Lee, X\. Lu, and A\. K\. Roy\-ChowdhuryContemplating visual emotions: understanding and overcoming dataset bias\.InEuropean Conference on Computer Vision,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p6.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p1.1)\.
- Peebleset al\.\(2022\)W\. Peebles, I\. Radosavovic, T\. Brooks, A\. A\. Efros, and J\. MalikLearning to learn with generative models of neural network checkpoints\.arXiv preprint arXiv:2209\.12892\.Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p1.1)\.
- Pryzantet al\.\(2023\)R\. Pryzant, D\. Iter, J\. Li, Y\. T\. Lee, C\. Zhu, and M\. ZengAutomatic prompt optimization with" gradient descent" and beam search\.InAnnual Meeting of the Association for Computational Linguistics,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p1.1)\.
- Radfordet al\.\(2021\)A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.Learning transferable visual models from natural language supervision\.InInternational Conference on Machine Learning,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p2.1),[§3\.2](https://arxiv.org/html/2608.12724#S3.SS2.SSS0.Px1.p1.2),[§4\.2](https://arxiv.org/html/2608.12724#S4.SS2.p3.1)\.
- Rahmanet al\.\(2025\)A\. Rahman, Q\. Xu, and X\. HuangMPCAR: multi\-perspective contextual augmentation for enhanced visual reasoning in large vision\-language models\.arXiv preprint arXiv:2508\.12400\.Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p2.1),[§2](https://arxiv.org/html/2608.12724#S2.p2.1)\.
- Rubinet al\.\(2022\)O\. Rubin, J\. Herzig, and J\. BerantLearning to retrieve prompts for in\-context learning\.InAnnual Conference of the North American Chapter of the Association for Computational Linguistics,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p2.1),[§2](https://arxiv.org/html/2608.12724#S2.p1.1)\.
- Sohnet al\.\(2020\)K\. Sohn, D\. Berthelot, N\. Carlini, Z\. Zhang, H\. Zhang, C\. A\. Raffel, E\. D\. Cubuk, A\. Kurakin, and C\. LiFixmatch: simplifying semi\-supervised learning with consistency and confidence\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p2.1)\.
- Tarvainen and Valpola \(2017\)A\. Tarvainen and H\. ValpolaMean teachers are better role models: weight\-averaged consistency targets improve semi\-supervised deep learning results\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p2.1)\.
- Wuet al\.\(2023\)Z\. Wu, Y\. Wang, J\. Ye, and L\. KongSelf\-adaptive in\-context learning: an information compression perspective for in\-context example selection and ordering\.InAnnual Meeting of the Association for Computational Linguistics,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p1.1),[§1](https://arxiv.org/html/2608.12724#S1.p2.1),[§2](https://arxiv.org/html/2608.12724#S2.p1.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.12724#S4.T1.2.1.8.1.1)\.
- Xieet al\.\(2021\)S\. M\. Xie, A\. Raghunathan, P\. Liang, and T\. MaAn explanation of in\-context learning as implicit bayesian inference\.arXiv preprint arXiv:2111\.02080\.Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p1.1)\.
- Yanget al\.\(2023\)J\. Yang, Q\. Huang, T\. Ding, D\. Lischinski, D\. Cohen\-Or, and H\. HuangA large\-scale visual emotion dataset with rich attributes\.InIEEE/CVF International Conference on Computer Vision,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p6.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p1.1)\.
- Yeet al\.\(2023\)J\. Ye, Z\. Wu, J\. Feng, T\. Yu, and L\. KongCompositional exemplars for in\-context learning\.InInternational Conference on Machine Learning,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p1.1)\.
- Zhanget al\.\(2020\)Y\. Zhang, N\. Wang, Y\. Chen, C\. Zou, H\. Wan, X\. Zhao, and Y\. GaoHypergraph label propagation network\.InAAAI Conference on Artificial Intelligence,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p3.1)\.
- Zhenget al\.\(2023\)M\. Zheng, S\. You, L\. Huang, C\. Luo, F\. Wang, C\. Qian, and C\. XuSimmatchv2: semi\-supervised learning with graph consistency\.InIEEE/CVF International Conference on Computer Vision,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p3.1)\.
- Zhouet al\.\(2004\)D\. Zhou, O\. Bousquet, T\. Lal, J\. Weston, and B\. SchölkopfLearning with local and global consistency\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p4.1),[§2](https://arxiv.org/html/2608.12724#S2.p3.1)\.
- Zhouet al\.\(2022\)Y\. Zhou, A\. I\. Muresanu, Z\. Han, K\. Paster, S\. Pitis, H\. Chan, and J\. BaLarge language models are human\-level prompt engineers\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2608.12724#S2.p1.1)\.
- Zhouet al\.\(2024\)Y\. Zhou, X\. Li, Q\. Wang, and J\. ShenVisual in\-context learning for large vision\-language models\.arXiv preprint arXiv:2402\.11574\.Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p1.1),[§2](https://arxiv.org/html/2608.12724#S2.p2.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.12724#S4.T1.2.1.7.1.1)\.
- Zhuet al\.\(2003\)X\. Zhu, Z\. Ghahramani, and J\. D\. LaffertySemi\-supervised learning using gaussian fields and harmonic functions\.InInternational Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p2.1)\.
- Zhuet al\.\(2002\)X\. Zhu, Z\. Ghahramani, and J\. LaffertyLearning from labeled and unlabeled data with label propagation\.InInternational Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p4.1),[§2](https://arxiv.org/html/2608.12724#S2.p3.1)\.
- Zonget al\.\(2024\)Y\. Zong, O\. Bohdal, and T\. HospedalesVL\-icl bench: the devil in the details of multimodal in\-context learning\.arXiv preprint arXiv:2403\.13164\.Cited by:[§1](https://arxiv.org/html/2608.12724#S1.p6.1),[§4\.1](https://arxiv.org/html/2608.12724#S4.SS1.p1.1)\.

## Appendix

## Appendix AAlgorithm Psuedo Code

Algorithm[1](https://arxiv.org/html/2608.12724#alg1)summarizes the overall procedure of MAG, including the generation of the construction of the expanded pool, the final query\-conditioned selection and prediction process\.

Algorithm 1MAG1:Input:Labeled data

𝒟L\\mathcal\{D\}\_\{L\}, unlabeled data

𝒟U\\mathcal\{D\}\_\{U\}, query

xqx\_\{q\},

qqq\_\{q\}
2:Output:Predicted answer

a^q\\hat\{a\}\_\{q\}
3:

4:// Generate textual descriptions for all images

5:for

xi∈𝒟L∪𝒟Ux\_\{i\}\\in\\mathcal\{D\}\_\{L\}\\cup\\mathcal\{D\}\_\{U\}do

6:

di←ℳdesc​\(xi\)d\_\{i\}\\leftarrow\\mathcal\{M\}\_\{\\text\{desc\}\}\(x\_\{i\}\)// Generate image description

7:endfor

8:

9:// Stage 1: Relevance\-guided pseudo\-labeling

10:Build textual graph

𝒢\(t\)\\mathcal\{G\}^\{\(t\)\}over

𝒟L∪𝒟U\\mathcal\{D\}\_\{L\}\\cup\\mathcal\{D\}\_\{U\}using text embeddings

11:Compute relevance scores

IjI\_\{j\}for

xj∈𝒟Ux\_\{j\}\\in\\mathcal\{D\}\_\{U\}via Eq\. \(4\)

12:for

xj∈Top\-​K​\(\{Ij\}xj∈𝒟U\)x\_\{j\}\\in\\text\{Top\-\}K\\left\(\\\{I\_\{j\}\\\}\_\{x\_\{j\}\\in\\mathcal\{D\}\_\{U\}\}\\right\)do

13:Pseudo\-label:

a^j←ℳ⁡\(xj,qj,𝒟L\)\\hat\{a\}\_\{j\}\\leftarrow\\mathcal\{M\}\(x\_\{j\},q\_\{j\};\\mathcal\{D\}\_\{L\}\)
14:Add

\(xj,a^j\)\(x\_\{j\},\\hat\{a\}\_\{j\}\)to

𝒟P\\mathcal\{D\}\_\{P\}
15:endfor

16:Expand pool:

𝒟E←𝒟L∪𝒟P\\mathcal\{D\}\_\{E\}\\leftarrow\\mathcal\{D\}\_\{L\}\\cup\\mathcal\{D\}\_\{P\}
17:

18:// Stage 2: Multi\-modal query\-conditioned selection

19:Build visual graph

𝒢\(v\)\\mathcal\{G\}^\{\(v\)\}\(image embeddings\) and textual graph

𝒢\(t\)\\mathcal\{G\}^\{\(t\)\}\(text embeddings\) over

𝒟E\\mathcal\{D\}\_\{E\}
20:Propagate score

I\(∞\)\(m\)I\_\{\(\\infty\)\}^\{\(m\)\}via Eq \(7\)

21:Late Fusion:

I←β​I\(v\)\+\(1−β\)​I\(t\)I\\leftarrow\\beta I^\{\(v\)\}\+\(1\-\\beta\)I^\{\(t\)\}
22:Select:

D​S​\(q\)←Top\-​k​\(\{Ii\}i=1\|𝒟E\|\)DS\(q\)\\leftarrow\\text\{Top\-\}k\\left\(\\\{I\_\{i\}\\\}\_\{i=1\}^\{\|\\mathcal\{D\}\_\{E\}\|\}\\right\)
23:Predict:

a^q←ℳ⁡\(xq,qq,D​S​\(q\)\)\\hat\{a\}\_\{q\}\\leftarrow\\mathcal\{M\}\(x\_\{q\},q\_\{q\};DS\(q\)\)
24:return

a^q\\hat\{a\}\_\{q\}

## Appendix BAdditional Experimental Analysis

### B\.1Multi\-Modal Backbone Analysis

We evaluate MAG across seven state\-of\-the\-art multi\-modal foundation models to assess performance variation with backbone choice\. Table[5](https://arxiv.org/html/2608.12724#A2.T5)reports results across all eight benchmarks\.

Table 5:Performance of MAG with different multi\-modal backbones \(accuracy %\)\. Gemini\-2\.0\-Flash provides the most balanced performance across task types\.Overall trends\.We observe substantial performance variance across backbones, indicating that MAG is sensitive to the underlying multi\-modal reasoning and perception capabilities\. While all models benefit from MAG, their relative strengths differ significantly across task types\.

Balanced vs\. specialized performance\.Gemini\-2\.0\-Flash achieves the most balanced performance, maintaining consistently strong results across both perception\-heavy tasks \(e\.g\., TextOCR, CLEVR\) and reasoning\-intensive benchmarks \(e\.g\., GQA, OKVQA\)\. In contrast, Gemini\-3\.0\-Flash exhibits more polarized behavior, achieving state\-of\-the\-art results on MMStar and CLEVR, but suffering a notable drop on TextOCR, suggesting potential trade\-offs between visual reasoning and text recognition\.

Effect of backbone capability\.Stronger general\-purpose models such as GPT\-5 demonstrate clear advantages on high\-level reasoning tasks \(e\.g\., GQA, OKVQA\), while also maintaining competitive performance on structured reasoning benchmarks like CLEVR and MatchingMI\. However, earlier generations \(e\.g\., GPT\-4o, GPT\-4\.1\) show clear limitations on multi\-modal understanding tasks such as MMStar and TextOCR, highlighting the importance of improved visual\-text alignment in newer models\.

Open\-source vs\. proprietary models\.Qwen\-VL variants exhibit competitive performance on several benchmarks, particularly MatchingMI and GQA, but show instability on tasks requiring fine\-grained semantic understanding \(e\.g\., Emotion6\)\. Notably, Qwen\-VL\-Plus suffers a significant drop on Emotion6, indicating that emotion recognition remains challenging for some open\-source models\.

Key takeaway\.These results suggest that MAG is robust across a wide range of backbones, but its effectiveness is amplified by models with strong cross\-modal alignment and reasoning capabilities\. In particular, backbones that achieve a better balance between perception and reasoning \(e\.g\., Gemini\-2\.0\-Flash, GPT\-5\) tend to yield the most consistent gains across diverse benchmarks\.

### B\.2Scalability Analysis

We further analyze the scalability of our method in terms of both computational and memory complexity\.

As shown in Eq\. \(4\) of the main paper, the steady\-state solution of graph propagation is:

I\(∞\)\(m\)=\(1−α\)​\(𝐈−α​W^\(m\)\)−1​I\(0\)\(m\)\.I^\{\(m\)\}\_\{\(\\infty\)\}=\(1\-\\alpha\)\(\\mathbf\{I\}\-\\alpha\\hat\{W\}^\{\(m\)\}\)^\{\-1\}I^\{\(m\)\}\_\{\(0\)\}\.
The main computational cost arises from solving the linear system involving\(𝐈−α​W^\(m\)\)\(\\mathbf\{I\}\-\\alpha\\hat\{W\}^\{\(m\)\}\)\. In practice,W^\(m\)\\hat\{W\}^\{\(m\)\}is a sparse affinity matrix constructed viakk\-nearest neighbors\. Therefore, this reduces to solving asparse linear system, rather than performing dense matrix inversion\.

Such systems can be efficiently solved using iterative methods \(e\.g\., Conjugate Gradient, Lanczos\) or sparse Cholesky decomposition, with complexity typically ranging fromO⁡\(n​log⁡n\)O\(n\\log n\)toO⁡\(n1\.5\)O\(n^\{1\.5\}\), significantly lower than theO⁡\(n3\)O\(n^\{3\}\)complexity of dense inversion\. Memory consumption scales linearly asO⁡\(n​k\)O\(nk\), since only non\-zero edges are stored\.

#### Empirical Validation\.

We validate the above analysis with empirical measurements shown in Table[6](https://arxiv.org/html/2608.12724#A2.T6)\.

Table 6:Latency and memory usage of graph propagation\.As the number of samples increases, the number of edges grows linearly, consistent with theO⁡\(n​k\)O\(nk\)complexity\. The latency increases approximately linearly \(doubling the samples roughly doubles runtime\), confirming that our method avoids cubic complexity\. Memory usage also exhibits near\-linear growth, with minor deviations due to implementation overhead\.

Overall, both theoretical and empirical results demonstrate that our approach remains efficient and scalable for larger unlabeled datasets\.

### B\.3Pseudo\-Label Accuracy Analysis

We present a quantitative comparison between different pseudo\-label selection strategies in Table[7](https://arxiv.org/html/2608.12724#A2.T7)\.

Table 7:Pseudo\-label accuracy \(p\-label acc\) and final performance\.We observe a strong positive correlation between pseudo\-label accuracy and final performance\. In 4 out of 5 datasets, our method outperforms random selection in both pseudo\-label quality and downstream performance, validating the effectiveness of our selection strategy\.

#### Robustness to Real\-World Noise\.

To address real\-world challenges such as cross\-modal inconsistencies, our method adopts modality\-specific disentanglement\. Each modality is processed independently, preventing noise in one modality \(e\.g\., low\-quality images\) from contaminating the other during propagation\.

As demonstrated in Section 4\.6, the proposed late\-fusion strategy consistently outperforms alternative designs across seven datasets\. By selectively integrating reliable steady\-state representations from each modality, our method exhibits strong robustness to semantic gaps and noisy inputs\.

## Appendix CQualitative Analysis

### C\.1Comparison with Baselines

Figure[4](https://arxiv.org/html/2608.12724#A3.F4)compares MAG with strong baselines across emotion recognition \(EmoSet\), visual QA \(GQA\), reasoning \(MMStar\), and OCR \(TextOCR\)\. Our method consistently produces more accurate predictions\. In emotion recognition, we correctly identify subtle affective cues \(*disgust*\) while baselines produce generic interpretations\. For reasoning tasks, graph\-based selection helps distinguish linguistically plausible but visually inconsistent options\. In OCR, we accurately recognize scene text \(*REINICIO*\) where baselines hallucinate\. These examples demonstrate that multi\-modal graph propagation enables selection of globally relevant demonstrations beyond local similarity\.

![Refer to caption](https://arxiv.org/html/2608.12724v1/qualitative_baseline.png)Figure 4:Qualitative comparison between MAG and baselines \(VICL, MMICES, CVR\-LLM\) across four tasks\. Our method consistently produces more accurate predictions by selecting semantically relevant demonstrations via graph\-based propagation\.
### C\.2Ablation Visualizations

![Refer to caption](https://arxiv.org/html/2608.12724v1/qualitative_ablation_w.png)Figure 5:Visualization of the full MAG method across EmoSet, GQA, MMStar, and OCR tasks\.![Refer to caption](https://arxiv.org/html/2608.12724v1/qualitative_ablation_1.png)Figure 6:MAG with random Stage 1 pseudo\-label selection \(vs\. relevance\-guided\)\. Random selection leads to weaker semantic alignment and less accurate predictions\.![Refer to caption](https://arxiv.org/html/2608.12724v1/qualitative_ablation_2.png)Figure 7:MAG with random Stage 2 demonstration selection \(vs\. graph\-based\)\. Random selection introduces noisy demonstrations, harming reasoning and OCR tasks\.Figures[5](https://arxiv.org/html/2608.12724#A3.F5)–[7](https://arxiv.org/html/2608.12724#A3.F7)present qualitative ablations comparing the full method with randomized Stage 1 \(pseudo\-label selection\) and randomized Stage 2 \(demonstration selection\)\.

When Stage 1 selection is randomized \(Figure[6](https://arxiv.org/html/2608.12724#A3.F6)\), the retrieved demonstrations are less semantically aligned with the queries, which in turn yields less stable emotion recognition and weaker object\-centric reasoning\. Moreover, random selection increases the incidence of corrupted OCR text outputs, suggesting that the model is less likely to incorporate informative unlabeled samples under this setting\.

Replacing Stage 2 with random demonstration selection \(Figure[7](https://arxiv.org/html/2608.12724#A3.F7)\) further degrades performance\. Irrelevant demonstrations introduce noisy visual\-text associations, harming multi\-choice reasoning and scene\-text understanding\.

In contrast, the full model \(Figure[5](https://arxiv.org/html/2608.12724#A3.F5)\) consistently selects semantically coherent demonstrations, producing accurate predictions across all task types\. These qualitative results verify that both relevance\-guided pseudo\-labeling \(Stage 1\) and graph\-based demonstration selection \(Stage 2\) are essential for robust multi\-modal reasoning\.

## Appendix DEvaluation Protocol

Following prior multi\-modal in\-context learning works, we adopt an exact containment\-based evaluation criterion for all benchmarks\. Specifically, given the model responserrand the ground\-truth answeryy, a prediction is considered correct if the ground\-truth text appears in the generated response:

Correct⁡\(r,y\)=\{1,if​y⊆r,0,otherwise\.\\mathrm\{Correct\}\(r,y\)=\\begin\{cases\}1,&\\text\{if \}y\\subseteq r,\\\\ 0,&\\text\{otherwise\}\.\\end\{cases\}
That is, if the response contains the ground\-truth answer as a substring, the prediction is counted as correct; otherwise, it is treated as incorrect\. We report the average accuracy over all evaluation samples\. This evaluation protocol is particularly suitable for open\-ended multi\-modal generation settings, where models may produce additional explanatory text while still containing the correct answer\.

## Appendix EPrompts for ICL

The prompts we use for in\-context learning on different tasks are as follows\.

- •For visual emotion recognition task\(EmoSet and Emotion6\), the prompt we use in our method is "What is the emotion category of this image? Give only the emotion category, and no extra commentary, formatting, or chattiness\. You can only make prediction from the following categories: \[Label List\]\. Given an image and its description, please predict the emotion category of this image\. Here are several examples\. Image: image\-1, Description: description\-1, Question: question\-1, Emotion category: label\-1\. Image: image\-2, Description: description\-2, Question: question\-2, Emotion category: label\-2… Image: image\-N, Description: description\-N, Question: question\-N, Emotion category: label\-N\. Image: image\-q, Description: description\-q, Question: question\-q, Emotion category:"\.
- •For TextOCR, MMStar, CLEVR, GQA and OKVQA, the prompt is "Given an image and its description, please answer a question about the image\. Here are several examples\. Image: image\-1, Description: description\-1, Question: question\-1, Answer: label\-1\. Image: image\-2, Description: description\-2, Question: question\-2, Answer: label\-2… Image: image\-N, Description: description\-N, Question: question\-N, Answer: label\-N\. Image: image\-q, Description: description\-q, Question: question\-q, Answer:"\.
- •For MatchingMI, the prompt is "Given two images and their description, please answer a question about the two images\. Here are several examples\. Images: image\-1\-Aimage\-1\-B, Description: description\-1, Question: question\-1, Answer: label\-1\. Image: image\-2\-Aimage\-2\-B, Description: description\-2, Question: question\-2, Answer: label\-2… Image: image\-N\-Aimage\-N\-B, Description: description\-N, Question: question\-N, Answer: label\-N\. Image: image\-q\-Aimage\-q\-B, Description: description\-q, Question: question\-q, Answer:"\.

Similar Articles

Context-Aware RL for Agentic and Multimodal LLMs

Hugging Face Daily Papers

Introduces ContextRL, a reinforcement learning approach that teaches LLMs to identify which context supports an answer, achieving gains on agentic and multimodal benchmarks.

Many-Shot CoT-ICL: Making In-Context Learning Truly Learn

Hugging Face Daily Papers

This paper investigates many-shot chain-of-thought in-context learning for reasoning tasks, revealing that standard scaling rules do not transfer and proposing Curvilinear Demonstration Selection (CDS) for improved ordering, achieving up to 5.42 percentage-point gain.