Tag
This paper introduces BITCOS, a distribution-adaptive layout for storing ternary LLM weights more efficiently, achieving up to 1.28× speedup in matrix-vector multiplication and 1.27× in inference throughput on GPUs.
A new open-source repo compresses 60 million text chunks from 201 GB to 6 GB with zero loss in accuracy, making vector databases potentially obsolete for many use cases.