data-augmentation

Tag

Cards List
#data-augmentation

Evaluating the Effect of Frame Rate in Sequence-Based Classification of Autism-Related Self-Stimulatory Hand Idiosyncrasies

arXiv cs.AI ↗ · 2026-07-10 Cached

This paper evaluates the effect of frame sampling rate on sequence-based classification of autism-related self-stimulatory hand behaviors using LSTM and GRU models, achieving up to 98.75% accuracy at a 15-frame interval, and analyzes data augmentation strategies for small behavioral video datasets.

0 favorites 0 likes
#data-augmentation

When Does Generating More Help? Disentangling Fixed-Source Synthesis from Source Expansion in Synthetic Data Scaling

arXiv cs.CL ↗ · 2026-07-03 Cached

This paper isolates Fixed-Source Synthesis (FSS) from Source Expansion in synthetic data scaling, proposing a rectified scaling law that predicts performance at high budgets from low-budget fits. Empirical results show FSS is bounded and that adding seed questions outperforms increasing response budgets at large scales.

0 favorites 0 likes
#data-augmentation

@neural_avb: This is very close to how the text-albumentations library works. Inputs a passage source, and generates task-oriented d…

X AI KOLs Timeline ↗ · 2026-07-02 Cached

The article compares the text-albumentations library to the new Autodata paper, which adds review mechanics with a weak/strong resolver and an external judge to maintain synthetic dataset quality.

0 favorites 0 likes
#data-augmentation

AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition

Hugging Face Daily Papers ↗ · 2026-07-02 Cached

AGVBench is a reliability-oriented benchmark for data augmentation in vein recognition, evaluating 30 augmentation strategies across multiple datasets and backbones, revealing decoupling between accuracy and security.

0 favorites 0 likes
#data-augmentation

From SRA to Self-Flow: Data Augmentation or Self-Supervision?

Hugging Face Daily Papers ↗ · 2026-07-02 Cached

This paper investigates the mechanisms behind self-alignment methods in diffusion transformers, revealing that performance improvements from methods like Self-Flow primarily come from data augmentation along the noise dimension rather than token interactions between noise levels. The authors introduce Attention Separation to demonstrate this and propose an effective design combining self-representation alignment with dual-timestep augmentation.

0 favorites 0 likes
#data-augmentation

Data and Evaluation Closed-Loop for Model Capability Enhancement

arXiv cs.AI ↗ · 2026-06-30 Cached

Introduces the capability slice, a unit for linking evaluation failures to data interventions in LLMs, enabling a closed-loop process that diagnoses and fixes model weaknesses. Demonstrated on two case studies, showing recovery from training regression and significant math reasoning improvements.

0 favorites 0 likes
#data-augmentation

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

arXiv cs.CL ↗ · 2026-06-30 Cached

This paper investigates the distributional gap between synthetic and real speech in LLM-based ASR systems, identifies where the LLM separates them, and proposes using layer-selection and RIR augmentation to match real-data baselines with less real data.

0 favorites 0 likes
#data-augmentation

Counterfactual Residual Data Augmentation for Regression

arXiv cs.LG ↗ · 2026-06-30 Cached

Proposes Counterfactual Residual Data Augmentation (CRDA) for tabular regression, leveraging residual invariance under feature perturbations to generate realistic training samples, achieving significant MSE reduction on benchmarks.

0 favorites 0 likes
#data-augmentation

MirrorPPR: Exemplar-Based Portrait Photo Retouching

Hugging Face Daily Papers ↗ · 2026-06-28 Cached

MirrorPPR introduces an exemplar-based portrait retouching framework using Diffusion Transformer with LoRA adaptation and self-augmented training data, achieving superior quality and identity preservation.

0 favorites 0 likes
#data-augmentation

\textsc{DiARC}: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

arXiv cs.CL ↗ · 2026-06-26 Cached

DiARC is a method that improves the reasoning ability of large language models on ARC-like tasks by constructing preference pairs from positive and negative samples, outperforming baselines across multiple benchmarks.

0 favorites 0 likes
#data-augmentation

Equivariance and Augmentation for Bayesian Neural Networks

arXiv cs.LG ↗ · 2026-06-26 Cached

This paper studies data augmentation for Bayesian neural networks trained with variational inference, deriving conditions for exact equivariance and introducing novel symmetrization techniques like orbit expansion to improve symmetry and performance.

0 favorites 0 likes
#data-augmentation

Extracting Problem and Method Sentence from Scientific Papers: A Context-enhanced Transformer Using Formulaic Expression Desensitization

arXiv cs.CL ↗ · 2026-06-26 Cached

This paper proposes a context-enhanced transformer using formulaic expression desensitization for extracting problem and method sentences from scientific papers, achieving improvements of 3.71% and 2.67% in macro F1 score on two datasets.

0 favorites 0 likes
#data-augmentation

Data Augmentation: A Fourier Analysis Perspective

arXiv cs.LG ↗ · 2026-06-24 Cached

This paper develops a Fourier analysis framework to study data augmentation under group invariances, showing that partial augmentation can achieve the same minimax rates as full augmentation up to a vanishing approximation error, while also proving that exact invariance requires full group averaging.

0 favorites 0 likes
#data-augmentation

Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining

Hugging Face Daily Papers ↗ · 2026-06-19 Cached

This paper investigates training-time data augmentation techniques to mitigate overfitting in autoregressive language model pretraining under data-constrained, compute-abundant regimes, finding that combining token-level noise, sequence permutations, and target offset prediction improves validation loss.

0 favorites 0 likes
#data-augmentation

Low-resource Language Discrimination Towards Chinese Dialects with Transfer learning and Data Augmentation

arXiv cs.CL ↗ · 2026-06-18 Cached

The paper proposes a novel framework (CDDTLDA) using transfer learning and data augmentation to improve Chinese dialects discrimination under low-resource conditions, achieving state-of-the-art results on two benchmark corpora.

0 favorites 0 likes
#data-augmentation

Want Better Synthetic Data? Steer It: Activation Steering for Low-Resource Language Generation

arXiv cs.CL ↗ · 2026-06-18 Cached

This paper investigates activation steering as an alternative to few-shot prompting for generating synthetic data in low-resource languages. The authors propose LanguageSteering and QualitySteering strategies, showing that steering on early layers improves diversity and downstream model performance.

0 favorites 0 likes
#data-augmentation

REVES: REvision and VErification--Augmented Training for Test-Time Scaling

Hugging Face Daily Papers ↗ · 2026-06-17 Cached

Proposes REVES, a two-stage iterative framework that alternates between data augmentation and policy optimization to improve LLM reasoning by leveraging intermediate correction steps, achieving superior performance on coding benchmarks and constraint satisfaction problems.

0 favorites 0 likes
#data-augmentation

CoCoGEC: Counterfactual Generation for Robust Grammatical Error Correction

arXiv cs.CL ↗ · 2026-06-16 Cached

Proposes CoCoGEC, a counterfactual generation framework that alters error-irrelevant contexts in GEC training data to improve model robustness, achieving significant F0.5 gains on perturbed benchmarks.

0 favorites 0 likes
#data-augmentation

PiDA: Phonetically-Informed Data Augmentation for Robust Vietnamese Speech Translation

arXiv cs.CL ↗ · 2026-06-12 Cached

This paper presents PiDA, a phonetically-informed data augmentation method for Vietnamese speech translation that improves robustness by generating ASR-like corruptions using phonetic word embeddings, achieving up to +2.04 BLEU on noisy outputs.

0 favorites 0 likes
#data-augmentation

One Jailbreak, Many Tongues: Learning Language-Insensitive Intention Representations for Multilingual Jailbreak Detection

arXiv cs.CL ↗ · 2026-06-11 Cached

This paper proposes MLJailDe, a multilingual jailbreak detection framework that uses back-translation data augmentation and relative-distance constraints to improve cross-lingual generalization and robustness, achieving 98.5% F1 score across 11 languages.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback