Tag
GeoRVQ introduces a decoder-aware masked token model for physiological signals that improves accuracy and reduces decoded distortion by accounting for local response and residual dependencies in residual vector quantization.
This paper proposes a method to integrate full-reference image quality metrics into rate-distortion optimization for video codecs by approximating them with input-dependent quadratic distortions using stochastic Hessian estimates, achieving BD-rate savings in VVC.
This paper introduces a rate-distortion theoretical framework for factual hallucination in closed-book question answering, distinguishing between errors from missing coverage and compression distortion under finite memory.
This paper proposes a unified rate-distortion perspective on discrete visual tokenization, resolving key questions about quantization objectives and comparisons, and shows that vector quantization achieves the lowest distortion under controlled conditions.
HeadWiseKV is a training-free framework that compresses KV caches in hybrid long-context language models, reducing GPU memory usage and extending context lengths while maintaining quality.
This paper unifies memory compaction techniques across LLMs and agents under a rate-distortion framework, proposing a taxonomy and benchmark for evaluating compression across different layers.
FRAPPE is a novel autoencoding framework that uses a projection pursuit encoder to predict residuals from full input, enabling efficient variable-rate image compression with fast CPU-based encoding. At high compression ratios, FRAPPE-Image achieves higher perceptual quality than AVIF with 47x faster encoding, making real-time 1080p 30fps CPU-only encoding possible.
This paper introduces LiVeAction, a lightweight neural codec designed for real-time operation on resource-constrained devices. It utilizes an FFT-like structure and variance-based rate penalty to achieve superior rate-distortion performance while remaining practical for low-power sensors.