Tag
提出了 DC-SAE(解耦紧凑语义自编码器),在保持高保真重建的同时实现 32 倍空间压缩并加速扩散模型训练收敛,显著超越此前 SOTA 高压缩分词器 DC-AE。
The paper proposes an attention-free approach to masked language modeling using autoencoders and iterative refinement, achieving performance comparable to BERT with fewer FLOPs.
FuseReg is a regularization method for Representation Autoencoders that mitigates the reconstruction-generation gap by using random layer-subset sampling during training, improving generation performance and decoder robustness across different layer inputs.
This paper shows that fine-tuning autoencoders for reconstruction reduces effective dimensionality, making standard velocity prediction inefficient in diffusion models, and proposes using x0-prediction to focus on the signal manifold, consistently improving text-to-image generation.
This paper introduces Sparse Koopman Autoencoders (SKAEs) to identify local dynamical regimes in multibasin nonlinear systems, demonstrating superior forecasting performance and interpretable latent supports.
This paper investigates fixed points and stability in randomly initialized autoencoders, introducing local and global edge-of-chaos concepts using random matrix theory and Gaussian processes.
Introduces JUGAAD, a deep learning framework that uses autoencoders and census/geospatial data to downscale socioeconomic indicators in India from coarse survey resolution to fine-grained village-cluster scale, validated against district-level NSSO data.
This paper investigates the failure of boundary-seeking knowledge distillation (CAKE) when applied to bottlenecked generative autoencoders, showing that the shared latent manifold creates gradient conflicts that prevent effective synthesis of contrastive samples. A simple noise forward pass baseline is proposed instead.
This paper introduces attention-free latent memory and dynamic re-encoding to improve long-horizon predictions in Koopman autoencoders, reducing error accumulation on benchmark dynamical systems.
This paper provides a mathematical analysis of superposition in neural networks, deriving upper and lower bounds on L2 reconstruction loss for simple autoencoders with power activation functions, corroborating empirical findings by Elhage et al.
Physics-conforming Latent Twins is a framework for learning latent surrogate solution operators that enforce physical principles such as conservation laws and dissipative inequalities by design, using a constraint-transfer approach and structure-preserving latent dynamics.
Introduces Rational Sparse Autoencoder (RSAE), which replaces fixed encoder activations with trainable rational functions, improving reconstruction and sparsity trade-offs on residual-stream activations of open-weight language models across multiple baseline families.
The author shares their work on reducing the cost of multi-vector retrieval by using k-means as top-1 sparse coding. Omar Khattab adds that late-interaction sparse retrieval with neuron-level inverted indexing on unsupervised sparse autoencoders works well.
This paper proposes Single-stage Sparse Retrieval (SSR), which replaces K-means clustering with sparse autoencoders and inverted indexing, achieving 15x faster indexing and halved retrieval latency while improving accuracy on the BEIR benchmark.
This article introduces Prior-Aligned Autoencoders (PAE), a new method for creating diffusion-friendly latent manifolds that achieves state-of-the-art image generation quality while enabling 13x faster training convergence.
An educational blog post explaining the Vector Quantized Variational Autoencoder (VQ-VAE) architecture, a key component of OpenAI's DALL-E image generation model.