data-curation

Tag

Cards List
#data-curation

The Next Evolution of AI Is Learning From Your Dodgy Gaming Skills

Wired ↗ · 3d ago Cached

A British startup, Worldmodeldata, is using video game data to train AI world models, aiming to overcome data shortages for teaching AI to understand and interact with the physical world.

0 favorites 0 likes
#data-curation

SALSA: Semi-Autonomous Literature Summarization Assistant

arXiv cs.CL ↗ · 2026-09-22 Cached

SALSA is an open-source human-in-the-loop platform that combines AI and user intervention to extract structured datasets from multimodal scientific literature, focusing on materials research.

0 favorites 0 likes
#data-curation

SynthSentry: Detecting Synthetic Data Contamination in Language Model Training Data

arXiv cs.CL ↗ · 2026-09-14 Cached

SynthSentry introduces a corpus-level, model-agnostic method to detect synthetic data contamination in language model training data without access to generating models, using distributional divergence over lexical, n-gram, and perplexity statistics.

0 favorites 0 likes
#data-curation

@vanstriendaniel: The real tokenmaxxing move: use frontier agents to build small classifiers for large-scale data curation. Small classif…

X AI KOLs Following ↗ · 2026-09-10 Cached

Using frontier agents to build small classifiers for data curation can reduce costs significantly compared to LLM labeling, as demonstrated in a case study on the FinePDFs-Edu dataset.

0 favorites 0 likes
#data-curation

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

arXiv cs.LG ↗ · 2026-08-28 Cached

This paper demonstrates that scaling up off-task model-generated distillation data can amplify latent teacher traits in students, even when the data appears benign, suggesting the need for trait-aware curation in AI training.

0 favorites 0 likes
#data-curation

Bart- A vintage llm [R]

Reddit r/MachineLearning ↗ · 2026-08-24

Unbounded Labs introduces Bart, a vintage LLM trained from scratch on pre-1931 English text, with open-sourced datasets, benchmarks, and methodology for advancing vintage LLM research.

0 favorites 0 likes
#data-curation

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

arXiv cs.CL ↗ · 2026-08-20 Cached

The paper introduces TranslatePsy-AfriSLM, an open-source collection of machine translation resources for 19 Sub-Saharan African languages, demonstrating that fine-tuned small language models with filtered synthetic data outperform much larger models like TranslateGemma-27B and Qwen3.5-122B-A10B.

0 favorites 0 likes
#data-curation

From volunteers to data miners

Reddit r/ArtificialInteligence ↗ · 2026-08-19 Cached

Reddit's volunteer moderators inadvertently create a structured dataset vital for AI training, but their biases and arbitrary rules can embed errors and biases into AI systems, leading to real-world issues like hallucinations and manipulation.

0 favorites 0 likes
#data-curation

Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data

arXiv cs.CL ↗ · 2026-08-18 Cached

This paper assesses the prevalence of extremist speech in the training data of large language models, focusing on the Dolma corpus, and discusses implications for data curation and model safety.

0 favorites 0 likes
#data-curation

RL Environments Are All You Need (6 minute read)

TLDR AI ↗ · 2026-08-06 Cached

The author argues that RL environments serve as the essential data for building AI agents, enabling systematic training, prompt optimization, and evaluation rather than manual iteration.

0 favorites 0 likes
#data-curation

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

Hugging Face Daily Papers ↗ · 2026-08-01 Cached

Introduces Poplar, a scalable Specify-Render-Inspect pipeline for synthesizing human-centric image datasets, and releases Poplar-9K, a curated dataset of 9,401 image-text pairs with auditable inspection records.

0 favorites 0 likes
#data-curation

DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes

Hugging Face Daily Papers ↗ · 2026-07-27 Cached

DecoupleMix introduces a systematic framework for optimizing pretraining data mixtures for Vision-Language Models by decoupling inter-class and intra-class ratio search, using convex optimization to improve scalability and performance over heuristic baselines.

0 favorites 0 likes
#data-curation

From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin

arXiv cs.CL ↗ · 2026-07-22 Cached

This paper presents an engineering study adapting NVIDIA Nemotron 3.5 ASR Streaming 0.6B to Kikuyu, Dholuo, and Kalenjin, achieving 42.97% and 33.98% WER on internal sets for Kikuyu and Dholuo, respectively, through data-centric techniques including corpus auditing, normalization, and streaming evaluation.

0 favorites 0 likes
#data-curation

BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking

arXiv cs.AI ↗ · 2026-07-14 Cached

This paper introduces BatteryLake, a governed data lakehouse that uses LLM agents for evidence-grounded metadata extraction and schema mapping, with human-in-the-loop verification, to curate heterogeneous battery aging datasets and release an open benchmark.

0 favorites 0 likes
#data-curation

Physics-Informed Machine Learning Under Small-Data Constraints: Lessons from Abrasive Waterjet Milling

arXiv cs.LG ↗ · 2026-07-10 Cached

This paper presents methodological contributions for physics-informed machine learning under small-data constraints, using an abrasive waterjet milling dataset of 155 points. It shows that data curation choices, evaluation design, and physics integration form matter significantly, with Gaussian Process variants outperforming other models.

0 favorites 0 likes
#data-curation

Data for Agents

Hugging Face Blog ↗ · 2026-07-08 Cached

NVIDIA discusses the importance of open and synthetic data for building robust AI agents, highlighting their Nemotron open datasets for training, reasoning, and tool-use.

0 favorites 0 likes
#data-curation

CurateEvo: Data-Curation Evolving for Agentic Post-Training

arXiv cs.CL ↗ · 2026-07-08 Cached

CurateEvo is a failure-driven dynamic evolution framework for agentic post-training data curation. It iteratively rewrites curation strategies using failed trajectories, improving effectiveness and efficiency on benchmarks like ACEBench-Agent and BFCL-V4.

0 favorites 0 likes
#data-curation

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

Hugging Face Daily Papers ↗ · 2026-07-08 Cached

MedPMC is an automated framework that transforms medical literature into high-fidelity multimodal data for foundation models, achieving significant improvements across multiple benchmarks and clinical settings.

0 favorites 0 likes
#data-curation

Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting

arXiv cs.LG ↗ · 2026-07-07 Cached

This paper proposes a generator-agnostic post-generation curation method that selects informative subsets of synthetic images by splitting real classes into canonical homogeneous and non-redundant heterogeneous subsets, and scoring synthetic images via a fidelity-diversity criterion. It consistently outperforms existing data-selection baselines and matches real-data performance with up to 40% fewer synthetic samples.

0 favorites 0 likes
#data-curation

DataComp-VLM: Improved Open Datasets for Vision-Language Models

Hugging Face Daily Papers ↗ · 2026-06-26 Cached

This paper introduces DataComp-VLM (DCVLM), a comprehensive benchmark for evaluating data curation strategies for vision-language models. The authors find that data mixing, rather than filtering, significantly improves performance, and their resulting DCVLM-Baseline dataset achieves state-of-the-art results on 33 downstream tasks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback