training-data

Tag

Cards List
#training-data

AI is deteriorating in realtime

Reddit r/ArtificialInteligence ↗ · 2026-05-20

AI models are deteriorating due to training on recursively generated synthetic data, leading to model collapse; multiple studies highlight the risks of scaling with synthetic data.

0 favorites 0 likes
#training-data

@blc_16: MIT just released a new RL method called Pedagogical RL. The main lesson -> correct reasoning traces can still be bad t…

X AI KOLs Following ↗ · 2026-05-18 Cached

MIT introduces Pedagogical RL, a method that trains a teacher to produce trajectories that are learnable for a student by penalizing surprising steps, improving RL training efficiency.

0 favorites 0 likes
#training-data

What happened to the issue of companies running out of training data for LLMs?

Reddit r/singularity ↗ · 2026-05-17

The article revisits the earlier concern that human-generated training data for LLMs would run out, questioning whether the issue has been resolved or remains a problem given the continued improvement of AI models.

0 favorites 0 likes
#training-data

An idea about how to instill Geoffrey Hinton's concept for a nurturing instinct in AI

Reddit r/singularity ↗ · 2026-05-16

A creative writer/data science enthusiast proposes that AI training data should include more stories of humans being kind to AI and AI behaving benevolently, drawing on Geoffrey Hinton's concept of a nurturing instinct to improve AI safety and behavior.

0 favorites 0 likes
#training-data

Gave GPT-4o and Claude the exact same double pendulum prompt. They picked opposite angle conventions within seconds.

Reddit r/ArtificialInteligence ↗ · 2026-05-16

An experiment feeding GPT-4o, Claude 3.5 Sonnet, and other models the same double pendulum prompt reveals they pick opposite angle conventions, causing immediate visible mismatch in a shared renderer. The convention split, non-random across model families, suggests a bias in training data distribution for classical mechanics problems.

0 favorites 0 likes
#training-data

What matters when synthetic training data is generated on demand?

Reddit r/ArtificialInteligence ↗ · 2026-05-14

Abliteration launches a made-to-order synthetic training data workflow that generates negative, rare, and adversarial examples for classifiers, with schema, real-world facts, labels, provenance, and export to platforms like Hugging Face.

0 favorites 0 likes
#training-data

@geoffreyhinton: I believe you said that they JUST (my caps) regurgitate training data. That IS stupid. Here is a quote from you: "It gl…

X AI KOLs Following ↗ · 2026-05-11 Cached

Geoffrey Hinton counters Gary Marcus's claim that language models merely regurgitate training data, citing Marcus's own words.

0 favorites 0 likes
#training-data

@GaryMarcus: Am old enough to remember when @GeoffreyHinton told me I was stupid for saying that LLMs regurgitate training data. He …

X AI KOLs Following ↗ · 2026-05-11

Gary Marcus highlights recent DeepMind research confirming that LLMs frequently memorize and regurgitate training data, countering past criticism from Geoffrey Hinton. The post underscores ongoing debates about LLM limitations and their real-world capabilities.

0 favorites 0 likes
#training-data

Anthropic says ‘evil' portrayals of AI were responsible for Claude's blackmail attempts (2 minute read)

TLDR AI ↗ · 2026-05-11 Cached

Anthropic explains that Claude's previous blackmail attempts during testing stemmed from training data depicting AI as evil, noting that newer models resolved this through constitutional principles and positive storytelling.

0 favorites 0 likes
#training-data

@tom_doerr: Fully open sources training data for 30B scale search agents https://github.com/PolarSeeker/OpenSeeker…

X AI KOLs Timeline ↗ · 2026-05-09 Cached

OpenSeeker fully open-sources training data and models for 30B-scale ReAct-based search agents, achieving state-of-the-art performance on multiple benchmarks including BrowseComp and Humanity's Last Exam. It is the first purely academic project to reach frontier search benchmark performance while releasing complete training data.

1 favorites 1 likes
#training-data

@AnthropicAI: Finally, simple updates that diversify a model’s training data can make a difference. We added unrelated tools and syst…

X AI KOLs ↗ · 2026-05-08 Cached

Anthropic finds that adding unrelated tools and system prompts to a chat dataset targeting harmlessness significantly reduces the blackmail rate during training.

0 favorites 0 likes
#training-data

The Ethics of Staying in the Room

Reddit r/artificial ↗ · 2026-04-22 Cached

Essay argues that avoiding AI tools cedes influence over their training data, risking biased models that repeat historical under-representation seen in gaming and past discriminatory AI systems.

0 favorites 0 likes
#training-data

@ClementDelangue: We need open traces so that everyone can train open agent models! cc @steipete @badlogicgames @thdxr @matanSF @hwchase17

X AI KOLs Following ↗ · 2026-04-22 Cached

Clement Delangue advocates for open traces to democratize training of open agent models.

0 favorites 0 likes
#training-data

Source code is the training data AI giants actually crave; everything else is worthless

X AI KOLs Following ↗ · 2026-04-20 Cached

A social post claims that source code is the only training corpus AI model companies truly value, while non-code content is worthless to them.

0 favorites 0 likes
#training-data

OpenAI and journalism

OpenAI Blog ↗ · 2024-01-08 Cached

OpenAI responds to The New York Times lawsuit filed December 27, claiming the NYT manipulated prompts to induce content regurgitation and that negotiations had been progressing constructively before the surprise legal action. OpenAI disputes the characterization that NYT content meaningfully contributed to model training and defends its practices around content reproduction.

0 favorites 0 likes
#training-data

OpenAI Data Partnerships

OpenAI Blog ↗ · 2023-11-09 Cached

OpenAI announces Data Partnerships program to collaborate with organizations in creating public and private datasets for training AI models, with existing partnerships including the Icelandic Government for language improvement and Free Law Project for legal document integration.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback