Tag
Presents a machine learning framework using an autoencoder for efficient modeling of FinFET devices, achieving high accuracy with minimal training data.
IDEAL proposes an in-depth alignment framework for discrete representation autoencoding, jointly aligning quantized tokens with shallow and deep VFM features to achieve superior reconstruction and generation performance.
Proposes AE-YOLO, an attention-guided autoencoder-enhanced YOLO framework for robust insulator defect detection in UAV transmission-line imagery, achieving 95.10% [email protected] and outperforming YOLO baselines by 5 points.
The paper introduces MilliVid, a method for improving long-range consistency in video generation by using a multi-scale autoencoder to compress frames into hierarchical tokens and then generating them with a coarse-to-fine diffusion model, outperforming baselines on Minecraft videos.
SwiftVR is a real-time one-step generative video restoration framework that achieves high frame rates on consumer GPUs using efficient attention mechanisms and a lightweight restoration-aware autoencoder.
Introduces SelfBootTok, a self-bootstrapped tokenization method that separates global and local information, reducing generator computation by ~40% and achieving a new state-of-the-art gFID of 1.56 with only 64 tokens.
This paper proposes a Cycle-Space Detector (CSD) for detecting blind false data injection attacks on power systems, where an autoencoder generates stealthy perturbations aligned with the measurement Jacobian null space. The CSD uses topology-derived cycle constraints to improve detection without requiring precise line parameters.
Proposes CALAD, a channel-aware contrastive learning framework for multivariate time series anomaly detection that uses estimated channel relevance to construct contrastive samples, achieving state-of-the-art performance.
Tadpole introduces a foundation model for 3D PDEs, pre-trained as an autoencoder via efficient online data generation, enabling large-scale diverse training without storage overhead. It demonstrates strong fine-tuning performance for dynamics learning and generative modeling across heterogeneous physical systems.
This paper introduces DRoRAE, a method that improves visual tokenization by fusing multi-layer features from pretrained vision encoders rather than relying solely on the last layer. It demonstrates significant improvements in reconstruction and generation quality on ImageNet and establishes a scaling law between fusion capacity and performance.
This paper addresses the issue of dimensional collapse in VQ-VAEs, showing that representations often occupy a low-dimensional subspace. It proposes an 'AE Warm-Up' strategy that trains the model as an unquantized autoencoder first, which improves reconstruction quality and increases effective latent dimensionality.
This article introduces a polynomial autoencoder that improves upon PCA for compressing transformer embeddings by using a quadratic decoder to capture nonlinear variance. Benchmarks on BEIR show it significantly outperforms standard PCA and Matryoshka embeddings in retrieval quality while maintaining high compression ratios.
This article explains the architecture of DALL-E, focusing on its transformer component that correlates language with discrete image representations to generate high-quality images from text prompts.