Tag
A British startup, Worldmodeldata, is using video game data to train AI world models, aiming to overcome data shortages for teaching AI to understand and interact with the physical world.
SALSA is an open-source human-in-the-loop platform that combines AI and user intervention to extract structured datasets from multimodal scientific literature, focusing on materials research.
SynthSentry introduces a corpus-level, model-agnostic method to detect synthetic data contamination in language model training data without access to generating models, using distributional divergence over lexical, n-gram, and perplexity statistics.
Using frontier agents to build small classifiers for data curation can reduce costs significantly compared to LLM labeling, as demonstrated in a case study on the FinePDFs-Edu dataset.
This paper demonstrates that scaling up off-task model-generated distillation data can amplify latent teacher traits in students, even when the data appears benign, suggesting the need for trait-aware curation in AI training.
Unbounded Labs introduces Bart, a vintage LLM trained from scratch on pre-1931 English text, with open-sourced datasets, benchmarks, and methodology for advancing vintage LLM research.
The paper introduces TranslatePsy-AfriSLM, an open-source collection of machine translation resources for 19 Sub-Saharan African languages, demonstrating that fine-tuned small language models with filtered synthetic data outperform much larger models like TranslateGemma-27B and Qwen3.5-122B-A10B.
Reddit's volunteer moderators inadvertently create a structured dataset vital for AI training, but their biases and arbitrary rules can embed errors and biases into AI systems, leading to real-world issues like hallucinations and manipulation.
This paper assesses the prevalence of extremist speech in the training data of large language models, focusing on the Dolma corpus, and discusses implications for data curation and model safety.
The author argues that RL environments serve as the essential data for building AI agents, enabling systematic training, prompt optimization, and evaluation rather than manual iteration.
Introduces Poplar, a scalable Specify-Render-Inspect pipeline for synthesizing human-centric image datasets, and releases Poplar-9K, a curated dataset of 9,401 image-text pairs with auditable inspection records.
DecoupleMix introduces a systematic framework for optimizing pretraining data mixtures for Vision-Language Models by decoupling inter-class and intra-class ratio search, using convex optimization to improve scalability and performance over heuristic baselines.
This paper presents an engineering study adapting NVIDIA Nemotron 3.5 ASR Streaming 0.6B to Kikuyu, Dholuo, and Kalenjin, achieving 42.97% and 33.98% WER on internal sets for Kikuyu and Dholuo, respectively, through data-centric techniques including corpus auditing, normalization, and streaming evaluation.
This paper introduces BatteryLake, a governed data lakehouse that uses LLM agents for evidence-grounded metadata extraction and schema mapping, with human-in-the-loop verification, to curate heterogeneous battery aging datasets and release an open benchmark.
This paper presents methodological contributions for physics-informed machine learning under small-data constraints, using an abrasive waterjet milling dataset of 155 points. It shows that data curation choices, evaluation design, and physics integration form matter significantly, with Gaussian Process variants outperforming other models.
NVIDIA discusses the importance of open and synthetic data for building robust AI agents, highlighting their Nemotron open datasets for training, reasoning, and tool-use.
CurateEvo is a failure-driven dynamic evolution framework for agentic post-training data curation. It iteratively rewrites curation strategies using failed trajectories, improving effectiveness and efficiency on benchmarks like ACEBench-Agent and BFCL-V4.
MedPMC is an automated framework that transforms medical literature into high-fidelity multimodal data for foundation models, achieving significant improvements across multiple benchmarks and clinical settings.
This paper proposes a generator-agnostic post-generation curation method that selects informative subsets of synthetic images by splitting real classes into canonical homogeneous and non-redundant heterogeneous subsets, and scoring synthetic images via a fidelity-diversity criterion. It consistently outperforms existing data-selection baselines and matches real-data performance with up to 40% fewer synthetic samples.
This paper introduces DataComp-VLM (DCVLM), a comprehensive benchmark for evaluating data curation strategies for vision-language models. The authors find that data mixing, rather than filtering, significantly improves performance, and their resulting DCVLM-Baseline dataset achieves state-of-the-art results on 33 downstream tasks.