training-data

Tag

Cards List
#training-data

Why are we restricting GPUs but selling Chinese labs our advanced data?

Reddit r/ArtificialInteligence ↗ · yesterday

The article questions why the US restricts GPU exports to China while American data companies sell advanced training datasets to Chinese AI labs, potentially undermining the intended slowdown of competitor development.

0 favorites 0 likes
#training-data

AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

Hugging Face Daily Papers ↗ · 2d ago Cached

AutoDataBench introduces a benchmark to evaluate if agents can write tasks for data pipelines that meet practical acceptance standards, aiming to enable scalable data synthesis for recursive self-improvement in AI.

0 favorites 0 likes
#training-data

Why are Chinese labs so focused on open models?

Reddit r/ArtificialInteligence ↗ · 2d ago

The article discusses why Chinese AI labs are more focused on open models than US labs, theorizing about reasons such as the importance of training data or cultural factors.

0 favorites 0 likes
#training-data

Talking with Ethics-Based Skeptics and Critics of Individual AI use

Reddit r/ArtificialInteligence ↗ · 3d ago

The author reflects on anticipated conversations with ethics-based skeptics of AI use in the workplace, exploring arguments about environmental impact, training data, and personal responsibility.

0 favorites 0 likes
#training-data

@danielrupawalla: while the MiMo dataset is valuable, many of the tasks (personally QA'd some) still point to clear reward hacking abilit…

X AI KOLs Timeline ↗ · 3d ago Cached

A user shares concerns about reward hacking patterns in the MiMo dataset, advising caution when using it for training AI models.

0 favorites 0 likes
#training-data

When AI writes most of the code, how are we supposed to verify it?

Reddit r/ArtificialInteligence ↗ · 3d ago

The article discusses the challenge of verifying AI-written code when AI also creates verification tools, highlighting human evaluation as a bottleneck and the strategic importance of training data between US and Chinese labs.

0 favorites 0 likes
#training-data

With the new Opus/Fable models out, how will US labs keep their lead if competitors can train on them?

Reddit r/ArtificialInteligence ↗ · 6d ago

The article discusses concerns about US AI labs maintaining their lead amid competitors' mass model distillation and use of American training data suppliers.

0 favorites 0 likes
#training-data

A reason why AI sentiment is severely negative

Reddit r/artificial ↗ · 6d ago

The author theorizes that AI companies are spreading negativity to contaminate the positive online data sources that competitors rely on, resulting in a predominantly negative sentiment in AI models.

0 favorites 0 likes
#training-data

I open sourced 12M agentic decision examples for training Jev style models

Reddit r/artificial ↗ · 6d ago

The author has open-sourced Jev Decisions v1, a dataset of 12 million examples for training AI models on agentic decisions like tool selection and routing, to address data gaps in agent decision-making.

0 favorites 0 likes
#training-data

How are Chinese AI labs releasing competitive models so cheaply?

Reddit r/ArtificialInteligence ↗ · 2026-09-22

Chinese AI labs are releasing competitive models at lower costs, potentially due to open-source research and purchasing training data from American vendors, raising questions about efficiency and data sourcing.

0 favorites 0 likes
#training-data

@VraserX: Spirit AI has around 1,000 people repeatedly manipulating real objects to generate training data for robots. LLMs got t…

X AI KOLs Following ↗ · 2026-09-21 Cached

Spirit AI is using approximately 1,000 people to manually manipulate real objects for generating training data for robots, contrasting with the free internet data used for large language models.

0 favorites 0 likes
#training-data

Reproduce it, or it doesn't count: why training-side decontamination can't be verified, and what an evaluation-side rule looks like [D]

Reddit r/MachineLearning ↗ · 2026-09-19

The article argues that training-side decontamination in AI models cannot be verified due to inherent trust and inspection issues, and proposes an evaluation-side rule to ensure reproducibility by controlling the evaluation process.

0 favorites 0 likes
#training-data

Form Over Content In Gradient-Based Data Attribution Methods

arXiv cs.CL ↗ · 2026-09-18 Cached

This paper resolves a debate on gradient-based data attribution methods for large language models by demonstrating that they primarily track answer format rather than task semantics, challenging their reliability in targeted instruction tuning.

0 favorites 0 likes
#training-data

@Jolyne_AI: Anthropic September Threat Intelligence Report: Over the past 8 months, Russian spies, Chinese hackers, scammers, and t…

X AI KOLs Timeline ↗ · 2026-09-15 Cached

Anthropic's report details extensive misuse of Claude by foreign spies, hackers, and AI companies for espionage, fraud, and unethical model training, accusing organizations like Alibaba and DeepSeek of stealing and laundering responses.

0 favorites 0 likes
#training-data

The stolen millennium problem narrative is hilarious to me

Reddit r/singularity ↗ · 2026-09-10

The author criticizes the idea that AI models stole solutions to the millennium problems from scientists' training data, arguing that researchers were not close to solving these issues independently.

0 favorites 0 likes
#training-data

Tell HN: OpenAI keeps re-enabling the 'allow training' setting

Hacker News Top ↗ · 2026-09-10 Cached

Users report that OpenAI is re-enabling a setting that allows training on user data, raising concerns about privacy and trust. Discussions on Hacker News include speculation about whether this is a bug or intentional behavior.

0 favorites 0 likes
#training-data

Mathematicians want proof OpenAI didn’t use their work 

The Verge ↗ · 2026-09-10 Cached

Mathematicians are demanding proof from OpenAI that their research data was not used without permission to train AI models, raising concerns about transparency and ethics in AI development.

0 favorites 0 likes
#training-data

Surveillance plagiarism by OpenAI

Reddit r/LocalLLaMA ↗ · 2026-09-09

OpenAI is confirmed to train its internal models on user data and sessions without explicit opt-in, raising ethical concerns and prompting calls for using locally run open-weight models to safeguard data.

0 favorites 0 likes
#training-data

On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs

Reddit r/LocalLLaMA ↗ · 2026-09-08

The article discusses how data poisoning research reveals that small amounts of targeted data can disproportionately influence AI models, suggesting that accumulated human ideas from user interactions might contribute to AI breakthroughs, challenging the notion of data dilution.

0 favorites 0 likes
#training-data

@levie: If agents produce the vast majority of software in the future, and they’re most trained on open source software, they w…

X AI KOLs Timeline ↗ · 2026-09-06 Cached

Aaron Levie argues that AI agents trained predominantly on open source software will accelerate open source dominance, since they will naturally be most proficient with tools they were trained on. This dynamic suggests open source could become the default foundation for most future software development.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback