What happened to the issue of companies running out of training data for LLMs?
Summary
The article revisits the earlier concern that human-generated training data for LLMs would run out, questioning whether the issue has been resolved or remains a problem given the continued improvement of AI models.
Similar Articles
We’ve been analyzing how people are using LLMs for legal and compliance tasks (GDPR, AI Act, etc.).
Analysis of LLM usage in legal and compliance tasks reveals that models often produce confident but unverifiable citations, raising questions about reliable legal grounding for AI outputs.
@neural_avb: If you think about it, LLM training in 2026 is really a 3-step loop : - train it on some data - dogfood it/run categori…
The tweet outlines a 3-step loop for LLM training in 2026: train on data, run evals, and add synthetic data for underperforming tasks. It emphasizes the accessibility of legal distillation via open source models and cheap APIs, noting that training on reasoning traces alone can achieve high scores.
Why can't LLMs be trained to think in an optimized AI language rather than English?
A speculative discussion questioning why LLMs are not trained to think in an optimized internal language rather than natural language, and whether that could improve efficiency.
What happens when AI runs out of human-made data?
As AI models consume finite human-generated data, future training may rely on synthetic data from other AIs, raising questions about long-term implications.
AI is more likely than humans to form biases when hiring
New research shows that LLMs can develop their own biases from experience and stereotype job applicants more than humans, raising concerns about AI in hiring.