Our 1-bit quant of Hy3 295B runs 2.2x faster than the cloud API with no quality loss
Summary
A 1-bit quantized version of the Hy3 295B model achieves 2.2x faster inference speed compared to the cloud API with no quality loss.
Similar Articles
@atomic_chat_hq: 1-bit Hy3 running locally is 2.2x faster than its API at the same quality! We gave both models the same task and compar…
Tencent's Hy3 295B model now available in 1-bit and 4-bit GGUF formats, achieving 2.2x faster local inference compared to cloud API while maintaining quality, as demonstrated by running on 4x RTX 5090 with 128GB VRAM.
An official 1-bit quant for Hy4??? 👀
This article presents official 1-bit quantization builds for the Hy4 preview model, offering GGUF files with reduced sizes and instructions for running on patched llama.cpp.
We quantized the new Ornith 1.5 9B and 35B-A3B
The post details the quantization of Ornith 1.5 9B and 35B-A3B AI models using Atomic Dynamic methods, providing benchmarks against stock quantizations and sharing Hugging Face collections.
2.5x faster Qwen3.6 NVFP4 Unsloth quants
Unsloth releases quantized Qwen3.6 models using NVFP4 format, achieving 2.5x faster inference speeds.
@anirudhbv_ce: Introducing SpectralQuant.. here to save your KV cache :)
SpectralQuant is a new KV cache quantization technique achieving 5.95× compression on Mistral 7B with only 7.5% perplexity overhead, significantly outperforming TurboQuant while requiring only 15 seconds of calibration per model.