Tag
The paper explores using code-switched text to align representations in small language models under data constraints, demonstrating that a curriculum involving code-switching improves multilingual performance.
This paper presents a synthetic training pipeline that uses perturbed public documents to generate context-dependent training data, significantly improving large language models' performance on context learning tasks like CL-bench without human annotation.
The paper introduces MFP, a method using role-aware Morgan fingerprints for predicting reaction yields in chemistry, achieving high accuracy and faster training compared to existing methods like YieldBERT and GNAN.
REPAIR is a self-evolving data augmentation framework for scientific dense retrievers that resolves long-tail confusion via fact-verified iterative refinement, demonstrating significant performance improvements on materials science and biomedical benchmarks.
This paper proposes a coverage-aware virtual IMU augmentation framework to address data scarcity in human activity recognition, generating and selecting synthetic sensor data to improve model performance with limited labeled data.
The paper proposes CoMA-DiT, a bidirectional cross-modal Diffusion Transformer for latent augmentation in multimodal brain state decoding, which enhances performance in tasks like auditory attention decoding and emotion recognition by using paired modalities as mutual supervisory signals.
KItCAT introduces a lightweight training strategy for auto-regressive LLMs that uses input corruption to generate diverse training inputs, improving knowledge injection from niche documents without costly paraphrasing.
This paper investigates scenarios where adding correctly labeled data can harm model performance, introducing insertion-stability and examining the limits of dimension-based theory in machine learning generalization.
This paper proposes a geometric filtering framework that selects high-quality LLM-generated samples by evaluating their Euclidean distance to real class examples in an embedding space, improving few-shot text classification performance by +2.61 percentage points over SMOTE.
This paper proposes a simulation-based methodology to generate augmented traffic datasets by replacing physical sensors with virtual ones, extending sensor coverage in urban networks while preserving traffic patterns.
This paper introduces techniques for the second MLC-SLM Challenge, including random leading-silence cropping and synthetic data generation to enhance multilingual conversational speech tasks, achieving improved accuracy and reduced error rates.
The paper introduces S^2VOPD, a self-supervised method that improves vision-language models by distilling from original images into strongly augmented student views, achieving performance surpassing GPT-5.4 on benchmarks.
TailBooster is a dual-layer generative framework that synthesizes operationally valid extreme air-transport events using statistical tail extraction and autoencoder-based cleaning, significantly improving extreme-event prediction accuracy.
This paper systematically evaluates 15 machine learning models, including the TabPFN foundation model, for post-wildfire debris-flow prediction using USGS basin-scale data, finding TabPFN achieves the best performance (threat score 0.637) and that synthetic data augmentation improves most models.
This paper presents a novel unsupervised data augmentation method combining Gaussian Mixture Models and Large Language Models to improve clustering on imbalanced text datasets by generating synthetic documents for underrepresented clusters.
Proposes a GAN-based framework (GT-GAN) for synthesizing high-fidelity data from incomplete LEO satellite Internet observations, showing robustness even with 40% missing data.
This paper presents a three-stage neural-symbolic pipeline for gaming toxicity detection, combining transformer ensembles with rule-based mediation, achieving top accuracy in the EEUCA 2026 shared task.
This paper proposes a cross-lingual data augmentation strategy that uses machine translation to transfer expert-annotated difficulty labels from high-resource languages to low-resource languages. Experiments with BERT-based regression models show that augmenting scarce native data with translated corpora significantly improves the accuracy of text difficulty assessment.
Introduces K-IPO, a generate-then-select oversampling framework that preserves the original data's feature importance ranking (measured by Kendall's tau) during augmentation for imbalanced tabular data, showing improved preservation, explanation consistency, and predictive performance across 20 datasets.
Researchers from MIT and others developed GIFT, a system that teaches vision-language models to automatically convert 2D designs into accurate CAD programs for rapid prototyping, using model-generated data to correct mistakes and improve performance with less computation.