Hy3 1Bit 89-93 GB
Summary
Announcement of Hy3 1-bit quantized model with 89-93 GB memory footprint.
Similar Articles
An official 1-bit quant for Hy4??? 👀
This article presents official 1-bit quantization builds for the Hy4 preview model, offering GGUF files with reduced sizes and instructions for running on patched llama.cpp.
@aisearchio: Minimax H3 GGUFs are here! Q2 is only 8.49 GB, so you could fit it in lower-end GPUs. https://huggingface.co/realrebela…
Announcement that GGUF quantizations of MiniMax H3 are available, with the Q2 version being only 8.49 GB for lower-end GPUs.
@Xudong07452910: A flagship large model with 295B parameters can now run on a single 96GB inference GPU, with 50% faster decoding. Tencent Hunyuan team releases quantized versions for Hy3 (295B parameters). The 1-bit version (IQ1_M) compresses weights from 598GB to 85.5GB, a 6…
Tencent Hunyuan team releases quantized versions for the 295B-parameter Hy3 large model. The 1-bit version compresses weights to 85.5GB, enabling deployment on a single 96GB inference GPU with ~50% faster decoding. The open-source GGUF format is compatible with the llama.cpp ecosystem.
@atomic_chat_hq: 1-bit Hy3 running locally is 2.2x faster than its API at the same quality! We gave both models the same task and compar…
Tencent's Hy3 295B model now available in 1-bit and 4-bit GGUF formats, achieving 2.2x faster local inference compared to cloud API while maintaining quality, as demonstrated by running on 4x RTX 5090 with 128GB VRAM.
Our 1-bit quant of Hy3 295B runs 2.2x faster than the cloud API with no quality loss
A 1-bit quantized version of the Hy3 295B model achieves 2.2x faster inference speed compared to the cloud API with no quality loss.