training-data

Tag

Cards List
#training-data

Do Large Language Models Capture the Diversity in their Training Data?

arXiv cs.CL ↗ · 2026-09-03 Cached

The paper investigates the conditional diversity gap in large language models by comparing the entropy of generated outputs with training data and proposes an information-theoretic framework to measure and mitigate this gap.

0 favorites 0 likes
#training-data

Some people said the Minecraft clone I fully vibecoded with Qwen3.8-27B Q4 is not that impressive because Minecraft is in the training data, so I had the model add 4 things that are probably not.

Reddit r/LocalLLaMA ↗ · 2026-08-30

An individual built a Minecraft clone using the Qwen3.8-27B Q4 model and added custom features to address criticism that Minecraft data influenced the training.

0 favorites 0 likes
#training-data

Creepy crawlies

Lobsters Hottest ↗ · 2026-08-29 Cached

AI crawlers are causing significant CPU load on kernel.org by inefficiently scraping git commits for training data, consuming more resources than all legitimate access combined.

0 favorites 0 likes
#training-data

Who’s Training on Your AI Chats? The Big Players, Audited

Reddit r/ArtificialInteligence ↗ · 2026-08-27 Cached

The article audits major AI providers like OpenAI and Anthropic, revealing that they train on user chats by default on consumer plans, with opt-out options often hidden in settings.

0 favorites 0 likes
#training-data

Google Reportedly in Advanced Talks for $1.5 Billion Deal With AI Coding Startup Mechanize (4 minute read)

TLDR AI ↗ · 2026-08-27

Google is reportedly in advanced talks for a $1.5 billion deal with AI coding startup Mechanize, which develops virtual environments, benchmarks, and training data to help AI agents handle complex tasks.

0 favorites 0 likes
#training-data

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training

Hugging Face Daily Papers ↗ · 2026-08-25 Cached

The paper introduces Game2World Engine, a framework for removing UI from gameplay videos to create cleaner training data for video world models, with GameCleaner model achieving state-of-the-art results in UI removal.

0 favorites 0 likes
#training-data

Napster's homepage is now entirely AI agents. It's a clean test case for how fast training data goes stale.

Reddit r/artificial ↗ · 2026-08-23

Napster has pivoted to an AI agent platform, illustrating how AI training data quickly becomes outdated, leading models to provide confident but incorrect information about company identities.

0 favorites 0 likes
#training-data

@rohanpaul_ai: New Stanford + Carnegie paper. Screen recordings of real work look like free training data for agents. They aren't, bec…

X AI KOLs Timeline ↗ · 2026-08-23 Cached

The paper presents Task Model Induction, a method that decomposes screen recordings of real computer work into separate tasks and goal trees to better train AI agents, improving performance by 30% on unseen tasks compared to current summarizers.

0 favorites 0 likes
#training-data

Is (or will) AI learn backwards? (Since most of its training data is now AI-generated data)

Reddit r/ArtificialInteligence ↗ · 2026-08-21

The article raises concerns about AI models potentially training on data generated by AI itself, which could lead to issues like model collapse, citing examples such as Deezer removing millions of AI-generated songs and the high proportion of AI-created blog articles.

0 favorites 0 likes
#training-data

AI data startup Micro1 reaches $500M gross run rate amid AI training boom

TechCrunch AI ↗ · 2026-08-21 Cached

Micro1, an AI data startup, has reached a $500M gross annual run rate in eight months due to high demand for AI training data, though it remains smaller than competitors like Mercor and Handshake.

0 favorites 0 likes
#training-data

@garrytan: YC is the YC for AI Researchers

X AI KOLs Timeline ↗ · 2026-08-20 Cached

Y Combinator's Paper Club focuses on data challenges in AI research, with domain experts discussing training data and benchmarks.

0 favorites 0 likes
#training-data

@rohanpaul_ai: LLMs got the internet. Robots have to build their own internet. That is the data problem Humyn Labs is going after. @hu…

X AI KOLs Timeline ↗ · 2026-08-20 Cached

Humyn Labs is tackling the robotics data problem by converting human experience into synchronized training data with multiple sensors, such as IMU, depth cameras, and hand tracking, to make it useful for robot training.

0 favorites 0 likes
#training-data

@rohanpaul_ai: Brilliant piece by Zhipu Founder Tang Jie. AI scaling is moving past parameter growth. “How many parameters?” is becomi…

X AI KOLs Following ↗ · 2026-08-20 Cached

Zhipu Founder Tang Jie discusses how AI scaling is evolving beyond parameter count to include factors like training data, compute per forward pass, and post-training, with GLM-5.3 as an example.

0 favorites 0 likes
#training-data

@rohanpaul_ai: The Meta/Oxford study finds, a multimodal model may need surprisingly little image-generation data if language and visu…

X AI KOLs Following ↗ · 2026-08-14 Cached

A Meta/Oxford study finds that multimodal models need surprisingly little image-generation data if language and visual understanding are trained together, suggesting an optimal 70/25/5 split of language, image understanding, and image generation tokens. It also warns that delaying vision training causes 'vision laziness' where models ignore images.

0 favorites 0 likes
#training-data

Does pre-generative-AI data become more valuable as the internet fills with synthetic material?

Reddit r/artificial ↗ · 2026-08-12

A blog post and accompanying tweet explore whether pre-generative-AI data becomes more valuable as the internet fills with synthetic content, discussing provenance, model collapse, and Anthropic's book scanning.

0 favorites 0 likes
#training-data

Why isn't there more talk of AI being stopped by copyright laws?

Reddit r/ArtificialInteligence ↗ · 2026-08-09

Discussion of whether AI companies should face copyright lawsuits to force ethical training data sourcing, citing Suno's loss in Germany and arguing laws should require consent for training on people's data.

0 favorites 0 likes
#training-data

@macrodata_labs: Everyone is betting on Egocentric data to scale robotics But turning that footage into training data requires recoverin…

X AI KOLs Following ↗ · 2026-08-06 Cached

Macrodata Labs releases a research blog on scaling robotics with egocentric video data by recovering 3D hand motion signals using only open-source models.

0 favorites 0 likes
#training-data

TIL AI can draw a watch showing an actual time

Reddit r/singularity ↗ · 2026-08-04

The author observes that ChatGPT can now draw a watch with a specified time, noting this was a classic AI failure a year ago due to training data over-representing 10:10, and asks what other '10:10 problems' remain.

0 favorites 0 likes
#training-data

@turingbook: The greatest potential value of Harness (Agent) may be going into the real work scenarios of seasoned experts across industries, observing how they work and think, and recording and "distilling" their valuable experience—data that was previously largely tacit.

X AI KOLs Timeline ↗ · 2026-08-02 Cached

The tweet highlights the potential of agent harnesses to capture and distill experts' tacit knowledge as new training data, referencing a Berkeley AI Summit talk by Jianfeng Gao on agentic modeling as an emerging AI paradigm.

0 favorites 0 likes
#training-data

@julien_c: Happy to partner with @trufflesec to help them perform the largest secret scan of AI training data ever

X AI KOLs Following ↗ · 2026-07-31 Cached

Truffle Security, in partnership with Julien Chaumond, conducted the largest secret scan of AI training data on HuggingFace, finding 221,303 live unique credentials across 6,003 public datasets.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback