Tag
User reports successfully running a 27B parameter model quantized to 1-bit on a Jetson Orin NX 16GB edge device, expressing amazement at the feasibility.
A 1-bit quantized version of the Hy3 295B model achieves 2.2x faster inference speed compared to the cloud API with no quality loss.
Bonsai 27B achieves 93% size reduction via 1-bit quantization while retaining 90% intelligence, enabling local browser inference with custom WebGPU kernels.