Puzzling Success of Overparameterization: Lottery Tickets or Escape Dimensions?
Summary
A paper investigating the reasons behind the success of overparameterization in neural networks, comparing the lottery ticket hypothesis with escape dimensions.
Similar Articles
Double-Scoring: Reliable Extraction of Strong Lottery Tickets
This paper introduces double-scoring, an augmented score-space parameterization for extracting strong lottery tickets from neural networks. It improves upon edge-popup and pruning-at-initialization baselines and reduces sensitivity to sparsity hyperparameters.
@che_shr_cat: 1/ A 5M-parameter model just beat frontier LLMs on hard logical puzzles at less than 1/100,000th of the inference cost.…
A 5M-parameter model outperforms frontier LLMs on hard logical puzzles at a fraction of the inference cost by using continuous latent space test-time compute.
Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks
This paper analyzes when random low-dimensional reparameterizations can train neural networks, deriving an orientation-resolved master formula for the random-slice residual and introducing RaMaN, a scalable framework that predicts required latent dimensions while dramatically reducing memory costs.
Feature Lottery? A Bifurcation Theory of Concept Emergence
This paper introduces a bifurcation theory of representation dynamics to detect when neural networks acquire structured representations during training, using a Hessian analysis of a GMM probe. The resulting ratio β/β_c serves as a label-free phase coordinate that predicts the onset of usable structure and can forecast feature interpretability in sparse autoencoders early in training.
Better exploration with parameter noise
OpenAI presents parameter noise, a technique that adds adaptive noise to neural network policy parameters rather than action spaces, enabling agents to learn tasks significantly faster than traditional action noise approaches. The method achieves 2x faster learning on HalfCheetah and represents a middle ground between evolution strategies and deep RL approaches like TRPO and DDPG.