data-augmentation

Tag

Cards List
#data-augmentation

Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence

arXiv cs.CL ↗ · 2d ago Cached

The paper explores using code-switched text to align representations in small language models under data constraints, demonstrating that a curriculum involving code-switching improves multilingual performance.

0 favorites 0 likes
#data-augmentation

Learning to Learn from Context: Synthetic Training from Perturbed Public Documents

Hugging Face Daily Papers ↗ · 3d ago Cached

This paper presents a synthetic training pipeline that uses perturbed public documents to generate context-dependent training data, significantly improving large language models' performance on context learning tasks like CL-bench without human annotation.

0 favorites 0 likes
#data-augmentation

Role-Aware Morgan Fingerprints for Reaction Yield Prediction

arXiv cs.LG ↗ · 2026-09-22 Cached

The paper introduces MFP, a method using role-aware Morgan fingerprints for predicting reaction yields in chemistry, achieving high accuracy and faster training compared to existing methods like YieldBERT and GNAN.

0 favorites 0 likes
#data-augmentation

REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement

arXiv cs.AI ↗ · 2026-09-17 Cached

REPAIR is a self-evolving data augmentation framework for scientific dense retrievers that resolves long-tail confusion via fact-verified iterative refinement, demonstrating significant performance improvements on materials science and biomedical benchmarks.

0 favorites 0 likes
#data-augmentation

Coverage-Aware Virtual IMU Augmentation for Low-Resource Human Activity Recognition

arXiv cs.AI ↗ · 2026-09-16 Cached

This paper proposes a coverage-aware virtual IMU augmentation framework to address data scarcity in human activity recognition, generating and selecting synthetic sensor data to improve model performance with limited labeled data.

0 favorites 0 likes
#data-augmentation

Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding

arXiv cs.AI ↗ · 2026-09-12 Cached

The paper proposes CoMA-DiT, a bidirectional cross-modal Diffusion Transformer for latent augmentation in multimodal brain state decoding, which enhances performance in tasks like auditory attention decoding and emotion recognition by using paired modalities as mutual supervisory signals.

0 favorites 0 likes
#data-augmentation

KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training

arXiv cs.CL ↗ · 2026-09-02 Cached

KItCAT introduces a lightweight training strategy for auto-regressive LLMs that uses input corruption to generate diverse training inputs, improving knowledge injection from niche documents without costly paraphrasing.

0 favorites 0 likes
#data-augmentation

When Does More Correct Data Hurt? Insertion-Stability and the Limits of Dimension-Based Theory

arXiv cs.LG ↗ · 2026-08-17 Cached

This paper investigates scenarios where adding correctly labeled data can harm model performance, introducing insertion-stability and examining the limits of dimension-based theory in machine learning generalization.

0 favorites 0 likes
#data-augmentation

Geometric Filtering of LLM-Generated Samples for Few-Shot Text Classification

arXiv cs.CL ↗ · 2026-08-17 Cached

This paper proposes a geometric filtering framework that selects high-quality LLM-generated samples by evaluating their Euclidean distance to real class examples in an embedding space, improving few-shot text classification performance by +2.61 percentage points over SMOTE.

0 favorites 0 likes
#data-augmentation

Simulation-Driven Vehicular Traffic Data Augmentation: Extending Sensor Coverage Through Virtual Sensing

arXiv cs.AI ↗ · 2026-08-17 Cached

This paper proposes a simulation-based methodology to generate augmented traffic datasets by replacing physical sensors with virtual ones, extending sensor coverage in urban networks while preserving traffic patterns.

0 favorites 0 likes
#data-augmentation

Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge

arXiv cs.CL ↗ · 2026-08-17 Cached

This paper introduces techniques for the second MLC-SLM Challenge, including random leading-silence cropping and synthetic data generation to enhance multilingual conversational speech tasks, achieving improved accuracy and reduced error rates.

0 favorites 0 likes
#data-augmentation

Self-Supervised Visual On-Policy Distillation

Hugging Face Daily Papers ↗ · 2026-08-14 Cached

The paper introduces S^2VOPD, a self-supervised method that improves vision-language models by distilling from original images into strongly augmented student views, achieving performance surpassing GPT-5.4 on benchmarks.

0 favorites 0 likes
#data-augmentation

TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement

Hugging Face Daily Papers ↗ · 2026-08-12 Cached

TailBooster is a dual-layer generative framework that synthesizes operationally valid extreme air-transport events using statistical tail extraction and autoencoder-based cleaning, significantly improving extreme-event prediction accuracy.

0 favorites 0 likes
#data-augmentation

Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction

arXiv cs.LG ↗ · 2026-08-07 Cached

This paper systematically evaluates 15 machine learning models, including the TabPFN foundation model, for post-wildfire debris-flow prediction using USGS basin-scale data, finding TabPFN achieves the best performance (threat score 0.637) and that synthetic data augmentation improves most models.

0 favorites 0 likes
#data-augmentation

Imbalanced Data Clustering via Targeted Data Augmentation Using GMM and LLM

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper presents a novel unsupervised data augmentation method combining Gaussian Mixture Models and Large Language Models to improve clustering on imbalanced text datasets by generating synthetic documents for underrepresented clusters.

0 favorites 0 likes
#data-augmentation

A GAN-Based Framework for Robust Data Synthesis in Satellite Internet Observations

arXiv cs.AI ↗ · 2026-07-29 Cached

Proposes a GAN-based framework (GT-GAN) for synthesizing high-fidelity data from incomplete LEO satellite Internet observations, showing robustness even with 40% missing data.

0 favorites 0 likes
#data-augmentation

thaulab@EEUCA 2026: Who Said What to Whom? A Targeting-Aware Neural-Symbolic Pipeline for Gaming Toxicity Detection

arXiv cs.CL ↗ · 2026-07-24 Cached

This paper presents a three-stage neural-symbolic pipeline for gaming toxicity detection, combining transformer ensembles with rule-based mediation, achieving top accuracy in the EEUCA 2026 shared task.

0 favorites 0 likes
#data-augmentation

Translation as Augmentation: Effect of Translated Data on Assessment of Difficulty

arXiv cs.CL ↗ · 2026-07-22 Cached

This paper proposes a cross-lingual data augmentation strategy that uses machine translation to transfer expert-annotated difficulty labels from high-resource languages to low-resource languages. Experiments with BERT-based regression models show that augmenting scarce native data with translated corpora significantly improves the accuracy of text difficulty assessment.

0 favorites 0 likes
#data-augmentation

K-IPO: Kendall-constrained Importance Preserving Oversampling for Imbalanced Tabular Data

arXiv cs.LG ↗ · 2026-07-21 Cached

Introduces K-IPO, a generate-then-select oversampling framework that preserves the original data's feature importance ranking (measured by Kendall's tau) during augmentation for imbalanced tabular data, showing improved preservation, explanation consistency, and predictive performance across 20 datasets.

0 favorites 0 likes
#data-augmentation

A better way to turn 2D designs into 3D models for rapid prototyping

MIT News — Artificial Intelligence ↗ · 2026-07-16 Cached

Researchers from MIT and others developed GIFT, a system that teaches vision-language models to automatically convert 2D designs into accurate CAD programs for rapid prototyping, using model-generated data to correct mistakes and improve performance with less computation.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback