Tag
This paper evaluates the effect of frame sampling rate on sequence-based classification of autism-related self-stimulatory hand behaviors using LSTM and GRU models, achieving up to 98.75% accuracy at a 15-frame interval, and analyzes data augmentation strategies for small behavioral video datasets.
This paper isolates Fixed-Source Synthesis (FSS) from Source Expansion in synthetic data scaling, proposing a rectified scaling law that predicts performance at high budgets from low-budget fits. Empirical results show FSS is bounded and that adding seed questions outperforms increasing response budgets at large scales.
The article compares the text-albumentations library to the new Autodata paper, which adds review mechanics with a weak/strong resolver and an external judge to maintain synthetic dataset quality.
AGVBench is a reliability-oriented benchmark for data augmentation in vein recognition, evaluating 30 augmentation strategies across multiple datasets and backbones, revealing decoupling between accuracy and security.
This paper investigates the mechanisms behind self-alignment methods in diffusion transformers, revealing that performance improvements from methods like Self-Flow primarily come from data augmentation along the noise dimension rather than token interactions between noise levels. The authors introduce Attention Separation to demonstrate this and propose an effective design combining self-representation alignment with dual-timestep augmentation.
Introduces the capability slice, a unit for linking evaluation failures to data interventions in LLMs, enabling a closed-loop process that diagnoses and fixes model weaknesses. Demonstrated on two case studies, showing recovery from training regression and significant math reasoning improvements.
This paper investigates the distributional gap between synthetic and real speech in LLM-based ASR systems, identifies where the LLM separates them, and proposes using layer-selection and RIR augmentation to match real-data baselines with less real data.
Proposes Counterfactual Residual Data Augmentation (CRDA) for tabular regression, leveraging residual invariance under feature perturbations to generate realistic training samples, achieving significant MSE reduction on benchmarks.
MirrorPPR introduces an exemplar-based portrait retouching framework using Diffusion Transformer with LoRA adaptation and self-augmented training data, achieving superior quality and identity preservation.
DiARC is a method that improves the reasoning ability of large language models on ARC-like tasks by constructing preference pairs from positive and negative samples, outperforming baselines across multiple benchmarks.
This paper studies data augmentation for Bayesian neural networks trained with variational inference, deriving conditions for exact equivariance and introducing novel symmetrization techniques like orbit expansion to improve symmetry and performance.
This paper proposes a context-enhanced transformer using formulaic expression desensitization for extracting problem and method sentences from scientific papers, achieving improvements of 3.71% and 2.67% in macro F1 score on two datasets.
This paper develops a Fourier analysis framework to study data augmentation under group invariances, showing that partial augmentation can achieve the same minimax rates as full augmentation up to a vanishing approximation error, while also proving that exact invariance requires full group averaging.
This paper investigates training-time data augmentation techniques to mitigate overfitting in autoregressive language model pretraining under data-constrained, compute-abundant regimes, finding that combining token-level noise, sequence permutations, and target offset prediction improves validation loss.
The paper proposes a novel framework (CDDTLDA) using transfer learning and data augmentation to improve Chinese dialects discrimination under low-resource conditions, achieving state-of-the-art results on two benchmark corpora.
This paper investigates activation steering as an alternative to few-shot prompting for generating synthetic data in low-resource languages. The authors propose LanguageSteering and QualitySteering strategies, showing that steering on early layers improves diversity and downstream model performance.
Proposes REVES, a two-stage iterative framework that alternates between data augmentation and policy optimization to improve LLM reasoning by leveraging intermediate correction steps, achieving superior performance on coding benchmarks and constraint satisfaction problems.
Proposes CoCoGEC, a counterfactual generation framework that alters error-irrelevant contexts in GEC training data to improve model robustness, achieving significant F0.5 gains on perturbed benchmarks.
This paper presents PiDA, a phonetically-informed data augmentation method for Vietnamese speech translation that improves robustness by generating ASR-like corruptions using phonetic word embeddings, achieving up to +2.04 BLEU on noisy outputs.
This paper proposes MLJailDe, a multilingual jailbreak detection framework that uses back-translation data augmentation and relative-distance constraints to improve cross-lingual generalization and robustness, achieving 98.5% F1 score across 11 languages.