storage-efficiency

Tag

Cards List
#storage-efficiency

Breaking the 1.58-bit Barrier for Ternary LLMs

arXiv cs.AI ↗ · 2026-09-16 Cached

This paper introduces BITCOS, a distribution-adaptive layout for storing ternary LLM weights more efficiently, achieving up to 1.28× speedup in matrix-vector multiplication and 1.27× in inference throughput on GPUs.

0 favorites 0 likes
#storage-efficiency

@HowToPrompt__: Vector databases are officially cooked This repo shrinks 60 million text chunks from 201 GB to just 6 GB without any lo…

X AI KOLs Timeline ↗ · 2026-06-17 Cached

A new open-source repo compresses 60 million text chunks from 201 GB to 6 GB with zero loss in accuracy, making vector databases potentially obsolete for many use cases.

0 favorites 0 likes
← Back to home

Submit Feedback