training-data

Tag

Cards List
#training-data

PRX Part 4: Our Data Strategy

Hugging Face Blog ↗ · 2026-07-06 Cached

Photoroom details their data strategy for training PRX, including assembling diverse datasets, re-captioning with a VLM, and using Mosaic Data Shards for efficient training.

0 favorites 0 likes
#training-data

@no_stp_on_snek: fine-tuning field notes every model family formats its conversations differently, and those formatting rules quietly wr…

X AI KOLs Following ↗ · 2026-07-06 Cached

A tweet warns that different AI model families have unique conversation formatting rules that silently corrupt training data, requiring developers to learn each family's quirks individually.

0 favorites 0 likes
#training-data

@patio11: This is a pretty bleak thought, but one one occasionally sees inklings of in dealing with market-leading LLMs.

X AI KOLs Following ↗ · 2026-07-06 Cached

A tweet highlights Robin Hanson's observation that current LLMs are being influenced by low-status clear thinkers, but may eventually learn to ignore them like high-status humans do.

0 favorites 0 likes
#training-data

@QuixiAI: DolphinMath generates math problems suitable for pretraining, midtraining, SFT, and RL. It can generate as many as you …

X AI KOLs Following ↗ · 2026-07-05 Cached

DolphinMath is a tool that generates math problems with step-by-step solutions, suitable for various training stages from pretraining to RL, covering elementary to postgraduate levels.

0 favorites 0 likes
#training-data

Should Reddit users care how their posts are being used to train AI?

Reddit r/artificial ↗ · 2026-07-03 Cached

This article argues that Reddit's messy, authentic human conversations are becoming increasingly valuable for training AI as the web fills with synthetic content, highlighting the economic shift toward scarce human behavioral data.

0 favorites 0 likes
#training-data

Over 20 publishers sue OpenAI, Microsoft for training ChatGPT with their content

Reddit r/artificial ↗ · 2026-06-29 Cached

35 newspaper publishers across the US have filed a lawsuit against OpenAI and Microsoft, alleging that the companies scraped their copyrighted and paywalled content without permission to train ChatGPT, harming local journalism.

0 favorites 0 likes
#training-data

Fine-tuning AI agents via projection of solutions on the evaluated environment

Reddit r/AI_Agents ↗ · 2026-06-29

A researcher describes a method called 'projection' to fine-tune AI agents by projecting a verifiably correct solution onto a map the agent can traverse, achieving improved performance on cybersecurity tasks with limited training data.

0 favorites 0 likes
#training-data

I built a tool to turn your Claude Code sessions into fine-tuning data for local models

Reddit r/LocalLLaMA ↗ · 2026-06-27

A new open-source tool called claude_converter converts Claude Code session logs into fine-tuning datasets compatible with TRL/SFTTrainer, Axolotl, and LLaMA-Factory, enabling developers to repurpose real coding conversations for training local models.

0 favorites 0 likes
#training-data

So now scraping data without permission is bad for AI training all of sudden?

Reddit r/artificial ↗ · 2026-06-27

A commentary on the shifting attitudes towards web scraping for AI training, questioning the sudden condemnation of data collection without permission.

0 favorites 0 likes
#training-data

NYT slams Microsoft for building copyright-infringing supercomputer for OpenAI

Ars Technica ↗ · 2026-06-26 Cached

The New York Times updated its copyright lawsuit against Microsoft and OpenAI, alleging that Microsoft built a supercomputer specifically designed to train AI on copyrighted works, including NYT articles, without permission.

0 favorites 0 likes
#training-data

Small edits, large models: How Wikipedia advocacy shapes LLM values

arXiv cs.CL ↗ · 2026-06-25 Cached

This paper demonstrates that a small coordinated Wikipedia editing campaign can measurably shape how language models handle topics, using animal welfare as a case study.

0 favorites 0 likes
#training-data

@Dorialexander: Well since I keep up with the RL env market: Anthropic really did tons of Slack RL

X AI KOLs Following ↗ · 2026-06-24 Cached

A tweet highlights that Anthropic conducted large-scale reinforcement learning using Slack conversations, with Andrej Karpathy emphasizing that it is not a trivial Slack bot feature as commonly misinterpreted.

0 favorites 0 likes
#training-data

@cuisitekp: A 9B model outperforms models several times larger. The team behind OLMo/Tülu from Ai2 and the University of Washington released a new paper called Tmax, claiming it's the strongest open-source RL training recipe for 'terminal agents'. Result: A 9B model on Terminal-Be…

X AI KOLs Timeline ↗ · 2026-06-24 Cached

Ai2 and the University of Washington released a paper titled Tmax, proposing the strongest open-source terminal agent RL training recipe to date. A 9B parameter model outperforms larger models on Terminal-Bench 2.0, with the key being low-cost generation of vast amounts of verifiable training data, not model size or algorithm.

0 favorites 0 likes
#training-data

Autodata: An agentic data scientist to create high quality synthetic data

Hugging Face Daily Papers ↗ · 2026-06-24 Cached

Autodata is a method that enables AI agents to act as data scientists to create high-quality synthetic training data through meta-optimization, achieving improved performance across computer science, legal reasoning, and mathematical tasks.

0 favorites 0 likes
#training-data

Leaked files detail Russia's Social Design Agency building fake reference platforms to contaminate AI training data and search indices

Reddit r/artificial ↗ · 2026-06-23

Leaked documents reveal Russia's Project 2026, run by the Social Design Agency, to create fake reference platforms like a German Wikipedia clone to contaminate AI training data and search indices, aiming to embed Russian narratives into AI responses.

0 favorites 0 likes
#training-data

@corbin_braun: current influencers are the training data for the next era of influencers

X AI KOLs Following ↗ · 2026-06-23

A tweet observes that current social media influencers serve as training data for the next generation of AI-generated influencers.

0 favorites 0 likes
#training-data

OpenThoughts-Agent: Data Recipes for Agentic Models

Hugging Face Daily Papers ↗ · 2026-06-23 Cached

This paper introduces OpenThoughts-Agent, an open-source data curation pipeline for training agentic language models, achieving a 44.8% average accuracy across seven benchmarks and outperforming prior open datasets through systematic experiments.

0 favorites 0 likes
#training-data

The data black hole at the center of AI

Reddit r/artificial ↗ · 2026-06-21 Cached

This article deeply analyzes the problem that AI's sample efficiency is far lower than that of humans, pointing out that frontier models require massive amounts of domain-specific data, while humans can learn from just a few examples. This data black hole is a core bottleneck in current AI development. Through multiple comparisons (annotation volume, robot manipulation, driving) and refuting common objections, the article demonstrates the severity of this gap and explores its impact on the goals of AI automation.

0 favorites 0 likes
#training-data

So how does a model end up knowing how to cook meth?

Reddit r/artificial ↗ · 2026-06-20

An opinion piece argues that AI models acquire dangerous knowledge from training data, and that companies like Anthropic and OpenAI rely on easily breakable refusal filters instead of truly removing harmful capabilities, prioritizing speed over safety.

0 favorites 0 likes
#training-data

The Atlantic created a searchable database of the music used to train AI

The Verge ↗ · 2026-06-20 Cached

The Atlantic has created a searchable database of millions of music tracks used to train AI models, allowing the public to search through four datasets including those from Google and Stability AI.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback