@linghaokong76: Can networks perform better without adding non-zero weights? Our ICML 2026 paper says yes: spreading the same active we…
Summary
A new ICML 2026 paper shows that spreading the same active weights across more neurons reduces collisions and improves accuracy in neural networks, suggesting networks can perform better without adding non-zero weights.
View Cached Full Text
Cached at: 07/07/26, 01:22 AM
Can networks perform better without adding non-zero weights?
Our ICML 2026 paper says yes: spreading the same active weights across more neurons reduces collisions and improves accuracy.
Presenting today at ICML, 2pm!
W/ @inimai_s @yonashav @micahadler @DAlistarh & Nir Shavit https://t.co/hUxBOoIIrS
Similar Articles
Bug or Feature^2: Weight Drift, Activation Sparsity, and Spikes
This paper formally proves that training neural networks with asymmetric activation functions like ReLU, GELU, or SiLU causes weights to drift negative, leading to up to 90% activation sparsity. It also shows that squared activations like ReLU² improve performance but cause activation spikes, which can be fixed by clipping, with GELU² achieving the best validation loss.
@nicholasturner0: Even if we have ways to break up neural networks into interpretable parts, the largest weights between those parts can …
The research note investigates how weight superposition in neural networks leads to interference weights, causing confusion even when the networks are broken into interpretable parts.
Understanding neural networks through sparse circuits
OpenAI researchers present methods for training sparse neural networks that are easier to interpret by forcing most weights to zero, enabling the discovery of small, disentangled circuits that can explain model behavior while maintaining performance. This work aims to advance mechanistic interpretability as a complement to post-hoc analysis of dense networks and support AI safety goals.
@oneill_c: 1/ Can you actually get new facts into an LLM's weights without breaking the model? This question decides how we approa…
This thread presents research on whether new facts can be added to an LLM's weights without breaking the model, and finds that it breaks unexpectedly, making compressed KV caches and in-context learning more promising for continual learning.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
OpenAI presents weight normalization, a reparameterization technique that decouples weight vector length from direction to improve neural network training convergence and computational efficiency without introducing minibatch dependencies, making it suitable for RNNs and noise-sensitive applications.