training-data

Tag

Cards List
#training-data

Are you in the Weights?

Product Hunt ↗ · 2026-06-20

A tool that lets you check if your data is included in the training sets of large language models, exploring the concept of digital immortality in AI.

0 favorites 0 likes
#training-data

Why don't frontier labs say how much data they are training on?

Reddit r/ArtificialInteligence ↗ · 2026-06-17

Article questions why frontier AI labs like OpenAI and Anthropic do not disclose the size of their training data, suggesting that improvements may come from data volume rather than genuine intelligence.

0 favorites 0 likes
#training-data

Donate your coding sessions to an open CC-BY-4.0 dataset to help train open-weight and open source models

Reddit r/LocalLLaMA ↗ · 2026-06-16

A new initiative called Trace Commons aims to collect coding agent traces into an open CC-BY-4.0 dataset to help train open-weight and open-source models, countering the data advantage of proprietary models from Anthropic and OpenAI.

0 favorites 0 likes
#training-data

DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence Vectors

arXiv cs.LG ↗ · 2026-06-11 Cached

DeMix is a novel framework that detects erroneous training samples and identifies their specific error types (label errors, feature errors, spurious correlations) by analyzing influence vectors, achieving a 22.61% improvement in debugging F1-score and 9.32% gain in task performance after data repair.

0 favorites 0 likes
#training-data

ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories

arXiv cs.CL ↗ · 2026-06-11 Cached

This paper introduces ISE, a three-stage synthesis paradigm for generating multi-turn OS-agent trajectories with grounded execution, demonstrating that fine-tuning on the resulting ISE-Trace dataset significantly improves agent performance on ClawEval.

0 favorites 0 likes
#training-data

@rohanpaul_ai: A Primer paper about how reasoning models improve after training Shows that better reasoning models depend less on raw …

X AI KOLs Following ↗ · 2026-06-07 Cached

This primer paper explores how reasoning models improve after training, arguing that effective reasoning data relies more on checkable training evidence than raw data size. It categorizes reasoning data by verification methods and emphasizes preserving messy agent data for learning signals.

0 favorites 0 likes
#training-data

LIMMT: Less is More for Motion Tracking

Hugging Face Daily Papers ↗ · 2026-06-05 Cached

This paper introduces LIMMT, a data-centric study showing that training with high-quality, minimal subsets of motion data (under 3% of AMASS) outperforms using the full dataset for physics-based humanoid motion tracking, defining motion data quality through physics feasibility, diversity, and complexity.

0 favorites 0 likes
#training-data

@anyscalecompute: GPUs in Mumbai, training data in Iowa? Cross-region reads tax every epoch. We put @Alluxio NVMe caching in front of the…

X AI KOLs Following ↗ · 2026-06-04 Cached

Anyscale demonstrates a 20x speedup in cross-region training data reads by using Alluxio NVMe caching with Ray Data, showing warm cache reads drop from 4,241 to 208 seconds for 1TB.

0 favorites 0 likes
#training-data

The interesting part of model collapse isn't technical, it's epistemic

Reddit r/AI_Agents ↗ · 2026-06-01

This article explores model collapse not as a technical bug but as an epistemic problem: when an AI model's outputs become its own inputs, the model's representation of reality gradually flattens into a self-referential average, raising questions about how we distinguish a model that models the world from one that models only itself.

0 favorites 0 likes
#training-data

Data Isn't Scarce. Your Imagination Is (8 minute read)

TLDR AI ↗ · 2026-05-29 Cached

Asuka Zheng argues that the 'running out of training data' panic is misplaced; the real scarcity is a lack of imagination in collecting diverse, long-horizon data, illustrated by her SRE replacement project and broader research trends.

0 favorites 0 likes
#training-data

Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces

Hugging Face Daily Papers ↗ · 2026-05-28 Cached

This paper identifies harmful continuations in answer-correct long chain-of-thought training traces for LLM SFT, characterized by uncertainty-geometry mismatches, and proposes a lightweight boundary proxy method to remove them.

0 favorites 0 likes
#training-data

Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories

arXiv cs.CL ↗ · 2026-05-27 Cached

This paper diagnoses the low diversity in LLM-generated stories, finding that 88.3% of sampled stories contain one of 11 common words (e.g., Elias, lighthouse) across models, and traces this homogeneity to post-training data and alignment rather than prevalence in pre-training data.

0 favorites 0 likes
#training-data

Is Position Bias in Dense Retrievers Built In-or Learned from Data?

Hugging Face Daily Papers ↗ · 2026-05-26 Cached

This paper investigates whether positional bias in dense retrievers originates from architecture or training data, finding that training data distribution strongly influences bias and that balanced training can reduce sensitivity by up to 87% while maintaining retrieval performance.

0 favorites 0 likes
#training-data

AI Generated Code Quality

Reddit r/AI_Agents ↗ · 2026-05-25

The article discusses concerns that as AI tools generate increasing amounts of code, future models trained on this synthetic code may suffer from reduced quality and originality, and asks how major AI labs like OpenAI, Anthropic, and GitHub plan to address this issue.

0 favorites 0 likes
#training-data

Brain-LLM Alignment Tracks Training Data, Not Typology

arXiv cs.CL ↗ · 2026-05-25 Cached

This paper investigates brain-LLM alignment across English, Chinese, and French using fMRI data and multiple LLMs, finding that training-language dominance and typological distance, not an inherent English advantage, drive alignment patterns.

0 favorites 0 likes
#training-data

More and more workers in India are collecting video data to train humanoid robots using head-mounted cameras

Reddit r/singularity ↗ · 2026-05-23

Workers in India are increasingly using head-mounted cameras to collect video data for training humanoid robots, highlighting a growing trend in AI data collection.

0 favorites 0 likes
#training-data

I fine-tuned an LLM to be C-3PO to test which training data format works best for persona injection [P]

Reddit r/MachineLearning ↗ · 2026-05-23 Cached

An experiment comparing three Supervised Fine-Tuning data formats (demonstrations, first-person statements, synthetic documents) for injecting a C-3PO persona into Qwen3-4B, finding first-person statements best for generalization and synthetic documents best for factual knowledge.

0 favorites 0 likes
#training-data

Training Data - AI Microgames

Product Hunt ↗ · 2026-05-22

Training Data - AI Microgames is a product where users play microgames to help gather training data for AI.

0 favorites 0 likes
#training-data

@kothasuhas: really really cool work. TLDR: it probably does not make sense to filter _any_ data in the infinite compute regime

X AI KOLs Following ↗ · 2026-05-21 Cached

New research suggests that with sufficient compute, filtering training data for language models may be unnecessary, and models can benefit from low-quality data.

0 favorites 0 likes
#training-data

Terminal-World: Scaling Terminal-Agent Environments via Agent Skills

arXiv cs.CL ↗ · 2026-05-21 Cached

Terminal-World introduces a fully automated pipeline that uses agent skills to synthesize high-quality training data for terminal agents, enabling models to outperform baselines with only 1.2% of the training data. The method co-derives task instructions, environments, and teacher trajectories from skill primitives.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback