Tag
CRISP is a scalable coreset method for imbalanced tabular learning that efficiently reduces dataset size while maintaining high accuracy in tasks like fraud detection.
HERALD introduces a gradient-free graph condensation framework that adapts to heterophily in graphs by selecting nodes and features based on measured heterophily, achieving competitive performance on benchmark datasets.
This paper introduces GLOBE, a trajectory-aligned coreset selection framework that uses gradient trajectories across multiple checkpoints and multi-order matching with structured sparse optimization to select compact, representative training subsets, outperforming existing methods on six benchmarks.
This paper proposes F2CTO, the first distributed first-order constrained trilevel optimization method for robust coreset selection over distributed networks, with a non-asymptotic convergence guarantee of O(ε^(-3/2)).
This paper introduces TAKE (Trajectory-Aware Knowledge Estimation), a text dataset distillation framework that uses influence functions and optimal transport to reduce datasets to as little as 0.1% of their original size while preserving downstream task fidelity.
This paper proposes a submodular coreset selection method for LLM benchmarks that selects a subset of prompts without using model evaluation outcomes, achieving score preservation across 35 benchmarks and 18 LLMs.
The paper proposes a method to mitigate spurious correlations by disentangling learning dynamics of core and spurious features using a two-stage sample scoring function, achieving state-of-the-art debiasing performance with only 10% of training data.
SemiPrune is a label-efficient dataset pruning framework that uses semi-supervised learning to generate pseudo-labels from a small labeled subset, enabling existing supervised pruning methods to work with unlabeled data. It achieves state-of-the-art performance on domain-specific, image-corrupted, and long-tailed datasets.