Tag
This paper proposes a unified rate-distortion perspective on discrete visual tokenization, resolving key questions about quantization objectives and comparisons, and shows that vector quantization achieves the lowest distortion under controlled conditions.
UniSVQ proposes a unified 2-bit quantization framework that bridges scalar and vector quantization by parameterizing codewords as an affine transform of integer lattices, achieving state-of-the-art performance among scalar methods and matching vector methods with higher throughput.