reparameterization

Tag

Cards List
#reparameterization

When a Flatness Proxy Is Not a Function: Robustness Certificates and Training Interventions

arXiv cs.LG ↗ · 2d ago Cached

This paper shows that being a valid curvature upper bound does not automatically justify using a last-layer relative-flatness proxy as an adversarial-robustness certificate or as a differentiable training regularizer. The authors derive a gauge-invariant repair, demonstrate the proxy's unboundedness under symmetry-preserving softmax shifts, and show via CIFAR-10 experiments that quotient-regularized predictors stay aligned while raw-regularized ones can have their generalization suppressed.

0 favorites 0 likes
#reparameterization

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

Hugging Face Daily Papers ↗ · 2026-08-09 Cached

UniSpace introduces a unified visual representation that unifies semantic understanding, high-fidelity reconstruction, and image generation in a single space using a reparameterized ViT, eliminating the need for a separate VAE.

0 favorites 0 likes
#reparameterization

$\mathbf{\lambda}$-VAE: Variance Equalization for Posterior Collapse

arXiv cs.LG ↗ · 2026-07-08 Cached

Identifies two coupled causes of posterior collapse in VAEs and introduces λ-VAE, a modification to the reparameterization step that equalizes variance across latent dimensions, reducing collapse and improving information capacity.

0 favorites 0 likes
#reparameterization

@murage_kibicho: I added a neural sorting algorithm. It builds on the reparametarization trick from Stable diffusion! It's called a Gumb…

X AI KOLs Timeline ↗ · 2026-06-28 Cached

A Python implementation of the Gumbel-Sinkhorn neural network for sorting lists of numbers, based on the 2018 paper by Mena et al.

0 favorites 0 likes
#reparameterization

@HanGuo97: LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes…

X AI KOLs Following ↗ · 2026-05-21 Cached

CODA reparameterizes memory-bound operations in LLM training to fuse them into the matmul epilogue, achieving near state-of-the-art performance with LLM-generated kernels.

0 favorites 0 likes
#reparameterization

Are Flat Minima an Illusion?

arXiv cs.LG ↗ · 2026-05-08 Cached

This paper challenges the common belief that flat minima cause better generalization in neural networks, arguing that 'weakness'—a reparameterization-invariant measure of function simplicity—is the true driver. Empirical results on MNIST and Fashion-MNIST show that weakness predicts generalization while sharpness anticorrelates, and the large-batch generalization advantage vanishes as training data increases.

0 favorites 0 likes
← Back to home

Submit Feedback