Automatic Combination of Sample Selection Strategies for Few-Shot Learning
Summary
This paper proposes ACSESS, a method for automatically combining multiple sample selection strategies to improve few-shot learning across both in-context learning and gradient-based approaches. The work demonstrates that combining strategies consistently outperforms individual selection methods across 14 datasets with both text and image modalities.
View Cached Full Text
Cached at: 04/20/26, 08:32 AM
# Automatic Combination of Sample Selection Strategies for Few-Shot Learning
Source: https://arxiv.org/html/2402.03038
Branislav Pecher♠†, Ivan Srba†, Maria Bielikova†, Joaquin Vanschoren‡
♠Faculty of Information Technology, Brno University of Technology, Brno, Czechia
†Kempelen Institute of Intelligent Technologies, Bratislava, Slovakia
‡Eindhoven University of Technology, Eindhoven, Netherlands
{branislav.pecher, ivan.srba, maria.bielikova}@kinit.sk, [email protected]
###### Abstract
In few-shot learning, the selection of samples has a significant impact on the performance of the model. While effective sample selection strategies are well-established in supervised settings, research on large language models largely overlooks them, favouring strategies specifically tailored to individual in-context learning settings. In this paper, we propose a new method for Automatic Combination of Sample Selection Strategies (ACSESS) to leverage the strengths and complementarity of various well-established selection objectives. We investigate and compare the impact of 23 sample selection strategies on the performance of 5 in-context learning models and 3 few-shot learning approaches (meta-learning, few-shot fine-tuning) over 6 text and 8 image datasets. The experimental results show that the combination of strategies through the ACSESS method consistently outperforms all individual selection strategies and performs on par or exceeds the in-context learning specific baselines. Lastly, we demonstrate that sample selection remains effective even on smaller datasets, yielding the greatest benefits when only a few shots are selected, while its advantage diminishes as the number of shots increases.
## 1 Introduction
Many domains are characterised by a labelled data scarcity due to data collection/annotation costs or privacy considerations, making the training of typical deep learning models unfeasible. Few-shot learning addresses the challenge of adapting models to new tasks when only a handful of labelled samples are available (Song et al., 2023). Two main approaches exist: 1) in-context learning, where a pretrained large language model is conditioned on few examples without any parameter updates (Dong et al., 2022; Liu et al., 2023); and 2) gradient-based few-shot learning, which includes meta-learning and fine-tuning, where the knowledge from relevant datasets is transferred to a new problem using gradient-based training (Song et al., 2023; Chen et al., 2019; Vanschoren, 2018; Hospedales et al., 2021). More details regarding their differences are included in Appendix J.
The choice of which sample to use is critical in both approaches, as performance can vary drastically depending on sample quality. Existing studies in different settings, such as in LLM alignment (Zhou et al., 2024), have shown that curating a smaller subset of high-quality samples can lead to better performance than training on the full set of available samples. This holds for few-shot learning as well, especially for noisy or unbalanced datasets or for in-context learning, which is known to be sensitive to sample choice (Pecher et al., 2024; Agarwal et al., 2021; Köksal et al., 2022).
Many works have already proposed diverse sample selection strategies to identify an optimal set of samples. A sample selection strategy follows its specific objective by considering one or few properties of available samples, e.g., to maximise similarity, diversity, informativeness or quality of selected samples (Li and Qiu, 2023; Zhang et al., 2022; Chang and Jia, 2023). With the popularisation of large language models (LLMs), many new selection strategies are proposed for in-context learning. However, they are often tailored only for specific settings, limiting their applicability. At the same time, the well-established and long-standing selection strategies that were shown to perform well in supervised settings, such as active learning, are mostly ignored as part of LLM sample selection (Pecher et al., 2024; Albalak et al., 2024).
In this work, we demonstrate that well-established strategies can match or exceed the performance of in-context learning specific strategies. To achieve this, we automatically and optimally combine the sample selection objectives used by well-established sample selection strategies to identify samples with complementary properties. In order to leverage the strengths of different sample selection strategies, we propose ACSESS, an effective method for Automatic Combination of Sample Selection Strategies. First, a subset of relevant strategies that can improve the overall success of few-shot learning is identified. Afterwards, the selection objectives from the identified strategies are combined together, using weighting based on their expected contribution, in order to identify the most informative and high-quality samples that can provide the most benefit.
Our contributions and findings are as follows[^1]:
- We provide the first large-scale study of 23 sample selection strategies across 5 in-context learning models and 3 few-shot learning approaches, including meta-learning and few-shot fine-tuning, across 6 text and 8 image datasets. We show that sample selection can yield consistent gains (up to 55 and 33 percentage points for in-context learning and gradient few-shot learning, respectively), though the effects depend heavily on dataset and approach.
- We propose ACSESS, a method that automatically combines and weights sample selection objectives to leverage their strengths and identify samples with complementary properties (i.e., informativeness, representativeness and learnability). Our results show that the combination of well-established strategies identified by the ACSESS method consistently leads to performance on par or better than the existing in-context learning specific baselines.
- Through ablation studies, we find the following key insights: 1) sample selection has higher impact when the number of shots is low or when using noisy datasets; 2) at higher number of shots (30-40 on average), the impact of sample selection is negligible as all strategies regress to random selection; 3) after a certain point, the boost in performance from using more shots becomes negligible (30 shots for in-context learning; 50-shot for gradient few-shot learning); 4) sample selection is beneficial even for small dataset sizes (achieving similar performance when selecting from only 25% for in-context learning or 10% of the dataset for gradient few-shot learning).
## 2 Related Work
A large body of work is dedicated to selecting high-quality samples for in-context learning, where the overall performance was found to be sensitive to this choice, leading to large variability in results (Pecher et al., 2024; Köksal et al., 2022; Zhang et al., 2022). In the supervised setting, the selection is mostly focused on distilling datasets to a lower number of samples (Yu et al., 2023), selecting a core set representative of the full set (Guo et al., 2022), or reducing annotation costs using active learning (Ren et al., 2021). However, the impact of these strategies remains underexplored for other few-shot learning approaches.
Within the majority of the sample selection strategies, the sample selection objective takes into consideration a single property of available samples. To this end, various heuristics and unsupervised metrics serve as an estimate of their potential to increase the model's performance. For in-context learning, the most popular approach is selecting samples based on the similarity to the test sample, differing only in employed similarity measure and representation (Liu et al., 2022; An et al., 2023; Zemlyanskiy et al., 2022; Gupta et al., 2022; Pasupat et al., 2021; Wang et al., 2022; Nashid et al., 2023; Agrawal et al., 2023; Gao et al., 2021). Besides similarity, the samples are often selected based on their informativeness or the uncertainty, often using active learning strategies (Köksal et al., 2022; Schröder et al., 2022; Margatina et al., 2021; Park et al., 2022; Mavromatis et al., 2023), bias of the samples (Ma et al., 2023) or adversarial training (Agarwal et al., 2021). Other methods define a notion of quality for each sample, either by prompting a LLM to rate the samples (Shin et al., 2021) or observing how the inclusion or removal of the sample affects the performance (Maronikolis et al., 2023; Nguyen and Wong, 2023).
When considering representativeness, the samples are selected based on how well they are representative of the full dataset (Guo et al., 2022; Killamsettty et al., 2021; Mirzasoleimanetal et al., 2020; Paul et al., 2021; Coleman et al., 2020). Learnability metrics either determine how easy it is to learn the specific sample (Swayamdipta et al., 2020; Zhang and Plank, 2021), or how often the sample is forgotten after being learned (Toneva et al., 2018). Finally, some approaches balance multiple sample properties at the same time. Often, the similarity of samples is combined with their diversity (Wu et al., 2023; Qin et al., 2023) or the informativeness of the samples is balanced with their representativeness (Su et al., 2022; Levy et al., 2023; Ye et al., 2023; Liu and Wang, 2023). In specific cases, a two-step sample search is done, such as finding a set of informative samples and then using diversity-guided search to improve this set (Li and Qiu, 2023), or using different uncertainty measures to select the top samples (Diao et al., 2024).
Newer strategies leverage optimisation to select a single representative set of samples. One possibility is to train a sample retriever, such as a separate scoring model (Rubin et al., 2022; Luo et al., 2023; Li et al., 2023; Wang et al., 2023; Aimen et al., 2023) or define a surrogate scorer for modelling the sample's goodness and use multi-armed bandit to select the best exemplars (Purohit et al., 2024, 2025). Another option is to leverage reinforcement learning (Zhang et al., 2022; Shum et al., 2023; Scarlatos and Lan, 2023). Finally, the Datamodels approach trains a linear model to predict the performance gain of a set of samples to select the subset that would lead to the highest possible performance increase (Ilyas et al., 2022; Chang and Jia, 2023; Jundi and Lapesa, 2022; Vilar et al., 2023).
In this work, we investigate whether the combination of objectives from well-established single-property strategies can rival the selection from the strategies tailored specifically to in-context learning. Similar to the close works of Li and Qiu (2023); Purohit et al. (2024, 2025), we select a single set of samples, but do not focus solely on in-context learning. In essence, we complement work by Agarwal et al. (2021) by exploring and effectively combining sample properties important for few-shot learning performance. Finally, our proposed method is inspired by the Datamodels approach (Ilyas et al., 2022), but performs the selection at the level of strategies instead of samples.
## 3 ACSESS: Automatic Combination of Selection Strategies
Despite the large number of existing sample selection strategies, most remain constrained to a single objective, typically considering only one sample property. This narrow focus often leads to suboptimal selection, whereas greater performance gains can be achieved by appropriately combining multiple objectives that capture complementary sample properties. For example, the most informative sample that is hard to learn or is often forgotten may not contribute as much for few-shot methods that utilise gradient-based training. On the other hand, for in-context learning, the samples on the decision boundary, which are often hard to learn, may provide more benefit. Instead of focusing on a single property, we propose to select examples that are characterised by
[^1]: To support replicability and extension of our results, we openly publish the source code of our experiments at https://github.com/kinit-sk/ACSESSSimilar Articles
ClusterFewshot: Improving Few-shot Optimization for LLMs workflow
ClusterFewshot is a novel method for improving few-shot demonstration selection in LLM workflows by integrating semantic clustering and utility scoring, which reduces optimization costs and enhances accuracy in DSPy-based pipelines.
Selective Synergistic Learning for Video Object-Centric Learning
Selective Synergistic Learning (SSync) improves video object-centric learning by selectively distilling reliable cues via pseudo-labeling and transitive merging, avoiding error propagation from indiscriminate dense alignment.
Stage-adaptive Token Selection for Efficient Omni-modal LLMs
SEATS is a training-free, stage-adaptive token selection method that reduces computational overhead in omni-modal LLMs by progressively pruning redundant visual and audio tokens, achieving a 9.3x FLOPs reduction and 4.8x prefill speedup while preserving 96.3% performance.
ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning
ALPINE introduces an ultra-lightweight spatial-relational architecture for few-shot image classification that achieves accuracy gains with fewer parameters, faster convergence, and better robustness compared to baselines like Prototypical Networks and MAML.
Many-Shot CoT-ICL: Making In-Context Learning Truly Learn
This paper investigates many-shot chain-of-thought in-context learning for reasoning tasks, revealing that standard scaling rules do not transfer and proposing Curvilinear Demonstration Selection (CDS) for improved ordering, achieving up to 5.42 percentage-point gain.