binary-quantization

Tag

Cards List
#binary-quantization

@sudoingX: this lab took qwen 3.6 27b, the model i've been calling king of the 24gb tier all month, and crushed it down to 3.9gb. …

X AI KOLs Timeline · 2026-07-16 Cached

PrismML announces Bonsai 27B, a binary-quantized version of Qwen3.6 27B that runs on a phone using only 1.125 bits per weight, claiming 89.5% intelligence retention. The model is being independently tested by @sudoingX to verify performance.

0 favorites 0 likes
#binary-quantization

Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction

Hacker News Top · 2026-06-29 Cached

Mixedbread Search introduces asymmetric quantization for late interaction retrieval, achieving near-lossless quality with 97% storage reduction by storing document vectors as binary signs while keeping query vectors at higher precision.

0 favorites 0 likes
#binary-quantization

An Implementation of NanoQuant: A flexible binary quantization method

Reddit r/LocalLLaMA · 2026-06-08

NanoQuant is a flexible binary quantization method that compresses dense transformers to sub-1-bit per weight. This repository provides a PyTorch implementation, still a work in progress, capable of quantizing models like Qwen3-0.6B and Qwen3-4B.

0 favorites 0 likes
← Back to home

Submit Feedback