Tag
As AI models consume finite human-generated data, future training may rely on synthetic data from other AIs, raising questions about long-term implications.
This article analyzes the capital landscape of the AI for Science field, pointing out that data exhaustion has driven capital's attention to scientific experiment data, and summarizes representative financing cases.
This essay applies the Moneyball philosophy to Physical AI, arguing that the industry overvalues raw data volume and teleoperation hours while undervaluing data novelty and marginal utility. It provides a framework for pricing data and recommends strategies for capital efficiency in robotics.
A significant portion of remaining AI training data is on undigitized magnetic tapes stored in warehouses, highlighting a potential data source as internet-based data runs out.
Introduces TimeMoDE, a framework combining Diffusion Transformers with Mixture-of-Experts for generating realistic time series under data scarcity, using pre-training on multi-domain datasets and domain prompts to handle domain-specific features and diffusion timestep signals for adaptive denoising.
This paper introduces a method that trains a population of models instead of a single model to achieve significantly lower loss when scaling multi-epoch pretraining, especially under data scarcity.
Asuka Zheng argues that the 'running out of training data' panic is misplaced; the real scarcity is a lack of imagination in collecting diverse, long-horizon data, illustrated by her SRE replacement project and broader research trends.
Scientists claim to have found a solution to prevent AI models from cannibalizing themselves when human-generated data runs out, addressing the problem of model collapse where LLMs trained on synthetic data produce gibberish and hallucinations.
The article revisits the earlier concern that human-generated training data for LLMs would run out, questioning whether the issue has been resolved or remains a problem given the continued improvement of AI models.
Proposes TAP, a tabular augmentation policy that couples diffusion inpainting with a learner-conditioned policy to improve downstream model performance under data scarcity, outperforming strong baselines on real-world datasets.
This paper proposes a self-supervised physics-informed neural network (PINN) framework with a learnable blending neuron to adaptively balance physics-based and data-driven losses, and integrates transfer learning to improve efficiency under data scarcity. It is validated on liquid-metal miniature heat sink CFD data with only 87 datapoints, achieving under 8% error.