Tag
This paper analyzes the scaling limits of constant-stepsize SGD near flat minima, showing that the invariant law concentrates at scale α^(1/m) for objectives with flatness exponent m ≥ 2, and converges to non-Gaussian stationary distributions for m > 2.
Chamath revealed that his company found AI token costs double every 45 days while downstream productivity improves by at most 5%. He believes AI development is approaching a bottleneck and suggests companies reconsider their strategies or even consider exiting.
A blog post comparing hallucination rates of major AI models reveals that smaller open-source models like GLM-5.2 hallucinate significantly less than larger proprietary models like GPT-5.5, suggesting diminishing returns from scaling model size.
A Citadel Securities report argues that frontier AI is facing real economic limits due to compute and inference costs, leading to a shift toward cost discipline and model substitution. The note validates recent experiences of high token bills and predicts a bifurcation in AI usage.
A debate on whether AGI is inevitable or facing a wall, weighing AI self-improvement and reasoning against issues like lack of understanding, power constraints, and shifting goalposts.
This paper shows that layer-local training methods like Forward-Forward (FF) do not scale to realistic image sizes and datasets, and that synthetic benchmarks overstate their performance. The authors introduce a strong FF variant (DTG-FF) and demonstrate that on real data (e.g., ImageNet-100 at 224x224) FF achieves only 49.4% versus typical BP above 75%, while on synthetic tasks the gap narrows or reverses.
A deep-dive analysis exploring why AI companies continue to scale systems despite prominent researchers declaring the end of the scaling era and widespread acknowledgment of diminishing returns, examining the structural and financial incentives driving the industry.
AI models are deteriorating due to training on recursively generated synthetic data, leading to model collapse; multiple studies highlight the risks of scaling with synthetic data.