Our 1-bit quant of Hy3 295B runs 2.2x faster than the cloud API with no quality loss

Reddit r/LocalLLaMA Models

Summary

A 1-bit quantized version of the Hy3 295B model achieves 2.2x faster inference speed compared to the cloud API with no quality loss.

No content available
Original Article

Similar Articles

An official 1-bit quant for Hy4??? 👀

Reddit r/LocalLLaMA

This article presents official 1-bit quantization builds for the Hy4 preview model, offering GGUF files with reduced sizes and instructions for running on patched llama.cpp.

We quantized the new Ornith 1.5 9B and 35B-A3B

Reddit r/LocalLLaMA

The post details the quantization of Ornith 1.5 9B and 35B-A3B AI models using Atomic Dynamic methods, providing benchmarks against stock quantizations and sharing Hugging Face collections.