Tag
This paper presents a statistically-lossless quantization method for large language models, aiming to reduce model size without information loss.
This paper introduces ALP, a state-of-the-art lossless compression algorithm for IEEE 754 floating-point data, exploiting decimal and high-precision patterns. It achieved superior performance in decoding speed, compression ratio, and compression speed, earning the SIGMOD Best Artifact Award.
A tool for lossless compression of prompt caches designed specifically for AI agents.