data-curation

Tag

Cards List
#data-curation

@harold_matmul: dspy.GEPA used in pretraining data curation in the new Microsoft AI effort :-)

X AI KOLs Timeline ↗ · 2026-06-24 Cached

The article explains how GEPA (Genetic-Pareto Optimization) within DSPy is used for efficient prompt tuning, specifically applied to pretraining data curation at Microsoft AI, allowing researchers to replace manual prompt engineering with automated compute-driven optimization.

0 favorites 0 likes
#data-curation

OpenThoughts-Agent: Data Recipes for Agentic Models

Hugging Face Daily Papers ↗ · 2026-06-23 Cached

This paper introduces OpenThoughts-Agent, an open-source data curation pipeline for training agentic language models, achieving a 44.8% average accuracy across seven benchmarks and outperforming prior open datasets through systematic experiments.

0 favorites 0 likes
#data-curation

DRIFT: Refining Instruction Data via On-Policy Data Attribution

arXiv cs.LG ↗ · 2026-06-18 Cached

DRIFT proposes a method that uses on-policy influence functions to refine training data distribution for supervised fine-tuning of large language models, consistently improving performance ceilings over existing baselines.

0 favorites 0 likes
#data-curation

Characterizing Narrative Content in Web-scale LLM Pretraining Data

Hugging Face Daily Papers ↗ · 2026-06-17 Cached

A fine-grained study of narrative features in web-scale LLM pretraining data, introducing NarraBERT and NarraDolma to measure narrative patterns and their distribution across sources.

0 favorites 0 likes
#data-curation

Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation

arXiv cs.AI ↗ · 2026-06-12 Cached

This paper presents STG, a structured testbench generation framework for LLM-driven hardware design workflows that reduces token cost and improves verification reliability compared to existing prompt-based approaches.

0 favorites 0 likes
#data-curation

Can Generalist Agents Automate Data Curation?

arXiv cs.AI ↗ · 2026-06-04 Cached

Researchers introduce Curation-Bench, a benchmark to evaluate whether generalist coding agents can automate the iterative data curation loop in AI development. Results show agents can match strong baselines within ten iterations, but reliable data research requires scaffolded method adaptation rather than open-ended prompting alone.

0 favorites 0 likes
#data-curation

Can Generalist Agents Automate Data Curation?

Hugging Face Daily Papers ↗ · 2026-06-02 Cached

This paper explores whether generalist coding agents (Claude Code, Codex, etc.) can automate data curation loops, achieving published baselines within 10 iterations but revealing a gap in exploring new methods. A scaffold that forces agents to adapt prior research yields policies that beat baselines using 10x less data.

0 favorites 0 likes
#data-curation

Exploring Autonomous Agentic Data Engineering for Model Specialization

Hugging Face Daily Papers ↗ · 2026-05-28 Cached

This paper introduces Autonomous Agentic Data Engineering, a task where LLMs autonomously execute end-to-end data curation pipelines for model specialization, showing significant performance gains (e.g., GPT-5.2 improves a student model by 57.29%).

0 favorites 0 likes
#data-curation

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion

Hugging Face Daily Papers ↗ · 2026-05-28 Cached

LoMo proposes a data curation method that reformulates single-modality prompts into interleaved multimodal sequences to improve cross-modal representation alignment in vision-language models, achieving consistent gains on multiple benchmarks.

0 favorites 0 likes
#data-curation

GEM: Geometric Entropy Mixing for Optimal LLM Data Curation

arXiv cs.LG ↗ · 2026-05-27 Cached

GEM reformulates LLM data curation as a variational problem on the hypersphere, using geometric entropy mixing and a minorize-maximize algorithm to discover balanced semantic clusters, achieving state-of-the-art improvements in data mixing strategies by up to 1.2% average downstream accuracy.

0 favorites 0 likes
#data-curation

GRACE: Gradient-aligned Reasoning Data Curation for Efficient Post-training

arXiv cs.AI ↗ · 2026-05-14 Cached

GRACE proposes a gradient-aligned method that scores individual reasoning steps to select the most valuable data for post-training, achieving 108.8% of full-data performance with only 20% of the data.

0 favorites 0 likes
#data-curation

Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods

arXiv cs.LG ↗ · 2026-05-08 Cached

This paper introduces ADAPT, an online reweighting framework for LLM data curation that dynamically adjusts sample importance during training via loss weighting, outperforming offline selection and mixing methods in cross-benchmark generalization.

0 favorites 0 likes
#data-curation

Target-Oriented Pretraining Data Selection via Neuron-Activated Graph

arXiv cs.CL ↗ · 2026-04-20 Cached

This paper introduces Neuron-Activated Graph (NAG) Ranking, a training-free framework for selecting pretraining data aligned with target tasks by identifying and ranking candidate data based on similarity in neuron activation patterns. The approach achieves 4.9% average improvement over random sampling and demonstrates that sparse neuron patterns capture functional capabilities for target learning.

0 favorites 0 likes
#data-curation

marin-community/marin

GitHub Trending (daily) ↗ · 2026-08-25 Cached

Marin is an open-source research program and software platform dedicated to the transparent development of foundation models, covering data curation to model training and evaluation.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback