data-augmentation

Tag

Cards List
#data-augmentation

FlowLet: Conditional 3D Brain MRI Synthesis using Wavelet Flow Matching

Hugging Face Daily Papers ↗ · 2026-06-08 Cached

FlowLet is a conditional generative framework that synthesizes age-conditioned 3D brain MRIs using flow matching in an invertible wavelet domain, improving brain age prediction accuracy for underrepresented age groups with high efficiency.

0 favorites 0 likes
#data-augmentation

WISE-HAR: A Generalizable Ensemble Deep Learning Framework for WiFi-Based Human Activity Recognition

arXiv cs.AI ↗ · 2026-06-03 Cached

This paper presents WISE-HAR, an ensemble deep learning framework for WiFi-based human activity recognition, achieving robust performance and generalization across scenarios with minimal accuracy drops.

0 favorites 0 likes
#data-augmentation

GLENS: Global Search via Learning from Solver Iterates with Diffusion Models

arXiv cs.LG ↗ · 2026-06-02 Cached

GLENS is a data-efficient global search method that uses diffusion models to generate diverse, high-quality initial guesses for local minima in non-convex optimization problems by leveraging intermediate solver iterates as free data augmentation.

0 favorites 0 likes
#data-augmentation

Measuring the Symmetry--Data Exchange Rate

Hugging Face Daily Papers ↗ · 2026-05-31

This exploratory study empirically measures the symmetry–data exchange rate predicted by equivariance theory on controlled C_n-symmetric tasks, finding that wrong-group constraints are actively harmful, augmentation with test-time orbit averaging matches equivariant models exactly, and the empirical exchange rate is broadly consistent with theory but statistically inconclusive. The authors emphasize the study's exploratory nature and call for registered replications.

0 favorites 0 likes
#data-augmentation

SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction

arXiv cs.CL ↗ · 2026-05-25 Cached

Proposes SSDAU, a structured semantic data augmentation method for joint entity and relation extraction that preserves semantic structure by segmenting text based on entity labels and using BERTTopic for topic consistency, significantly outperforming existing augmentation methods.

0 favorites 0 likes
#data-augmentation

TERGAD: Structure-Aware Text-Enhanced Representations for Graph Anomaly Detection

arXiv cs.CL ↗ · 2026-05-20 Cached

TERGAD is a novel data augmentation framework that uses large language models to translate node-level topological properties into semantic narratives, then fuses these with original node attributes via a gated dual-branch autoencoder for graph anomaly detection, achieving state-of-the-art results on six datasets.

0 favorites 0 likes
#data-augmentation

How Data Augmentation Shapes Neural Representations

arXiv cs.LG ↗ · 2026-05-18 Cached

This paper uses shape analysis tools to characterize how different data augmentation strategies reshape the geometry of neural network representations, finding that augmentation strength and type lead to distinct, well-behaved trajectories in shape space.

0 favorites 0 likes
#data-augmentation

Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction

arXiv cs.CL ↗ · 2026-05-18 Cached

This paper proposes a large language model-driven data augmentation framework using GPT-5 to generate synthetic oral monologues from written anchors for cognitive score prediction from speech. A similarity-guided selection strategy consistently reduces prediction error, particularly for minority low-score participants.

0 favorites 0 likes
#data-augmentation

Mitigating Data Scarcity in Psychological Defense Classification with Context-Aware Synthetic Augmentation

arXiv cs.CL ↗ · 2026-05-15 Cached

This paper proposes a context-aware synthetic augmentation framework combined with a hybrid classification model to address data scarcity and class imbalance in classifying psychological defense mechanisms from text. The method achieves significant improvements on the PsyDefDetect shared task benchmark.

0 favorites 0 likes
#data-augmentation

DataArc-SynData-Toolkit: A Unified Closed-Loop Framework for Multi-Path, Multimodal, and Multilingual Data Synthesis

arXiv cs.LG ↗ · 2026-05-12 Cached

The article introduces DataArc-SynData-Toolkit, an open-source framework designed to simplify multi-path, multimodal, and multilingual synthetic data generation. It aims to lower technical barriers and improve usability for training large language models through a unified, configuration-driven pipeline.

0 favorites 0 likes
#data-augmentation

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations

arXiv cs.CL ↗ · 2026-05-11 Cached

This paper introduces GSM-SEM, a framework for generating semantically diverse benchmark variants to mitigate memorization in mathematical reasoning evaluations. The authors demonstrate that this approach reveals significant performance drops in current SOTA LLMs compared to static benchmarks.

0 favorites 0 likes
#data-augmentation

Active Tabular Augmentation via Policy-Guided Diffusion Inpainting

Hugging Face Daily Papers ↗ · 2026-05-11 Cached

Proposes TAP, a tabular augmentation policy that couples diffusion inpainting with a learner-conditioned policy to improve downstream model performance under data scarcity, outperforming strong baselines on real-world datasets.

0 favorites 0 likes
#data-augmentation

When Informal Text Breaks NLI: Tokenization Failure, Distribution Shift, and Targeted Mitigations

arXiv cs.CL ↗ · 2026-04-21 Cached

This paper investigates how informal text (slang, emoji, Gen-Z filler tokens) degrades NLI accuracy in ELECTRA-small and RoBERTa-large models, identifying two distinct failure mechanisms—tokenization failure (emoji mapped to [UNK]) and distribution shift (out-of-domain noise tokens)—and proposes targeted mitigations that recover accuracy without harming clean-text performance.

0 favorites 0 likes
#data-augmentation

Efficient training of language models to fill in the middle

OpenAI Blog ↗ · 2022-07-28 Cached

OpenAI presents a simple data augmentation technique that enables autoregressive language models to perform fill-in-the-middle (FIM) text generation without harming left-to-right performance, with extensive ablations and best practices provided for training such models.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback