Tag
This paper shows that being a valid curvature upper bound does not automatically justify using a last-layer relative-flatness proxy as an adversarial-robustness certificate or as a differentiable training regularizer. The authors derive a gauge-invariant repair, demonstrate the proxy's unboundedness under symmetry-preserving softmax shifts, and show via CIFAR-10 experiments that quotient-regularized predictors stay aligned while raw-regularized ones can have their generalization suppressed.
UniSpace introduces a unified visual representation that unifies semantic understanding, high-fidelity reconstruction, and image generation in a single space using a reparameterized ViT, eliminating the need for a separate VAE.
Identifies two coupled causes of posterior collapse in VAEs and introduces λ-VAE, a modification to the reparameterization step that equalizes variance across latent dimensions, reducing collapse and improving information capacity.
A Python implementation of the Gumbel-Sinkhorn neural network for sorting lists of numbers, based on the 2018 paper by Mena et al.
CODA reparameterizes memory-bound operations in LLM training to fuse them into the matmul epilogue, achieving near state-of-the-art performance with LLM-generated kernels.
This paper challenges the common belief that flat minima cause better generalization in neural networks, arguing that 'weakness'—a reparameterization-invariant measure of function simplicity—is the true driver. Empirical results on MNIST and Fashion-MNIST show that weakness predicts generalization while sharpness anticorrelates, and the large-batch generalization advantage vanishes as training data increases.