@aisearchio: Minimax H3 GGUFs are here! Q2 is only 8.49 GB, so you could fit it in lower-end GPUs. https://huggingface.co/realrebela…
Summary
Announcement that GGUF quantizations of MiniMax H3 are available, with the Q2 version being only 8.49 GB for lower-end GPUs.
View Cached Full Text
Cached at: 08/03/26, 11:55 PM
Minimax H3 GGUFs are here! Q2 is only 8.49 GB, so you could fit it in lower-end GPUs. https://huggingface.co/realrebelai/MiniMax-H3_GGUFs/tree/main…
realrebelai/MiniMax-H3_GGUFs at main
Source: https://huggingface.co/realrebelai/MiniMax-H3_GGUFs/tree/main
![]()
Upload MiniMax-H3-FL2VA-Q2_K-(Mixed_Precision).gguf
verified
about 9 hours ago
Similar Articles
realrebelai/MiniMax-H3_GGUFs
Hugging Face repository providing GGUF quantizations of MiniMax-H3 models for use with ComfyUI, including directory structure and links to required VAEs.
@aisearchio: GLM 5.2 GGUF is already here! 8-bit is ~half the size of the full model. Smaller versions coming soon https://huggingfa…
GLM 5.2 GGUF quantized model is released, with 8-bit version half the size of the full model; smaller versions are coming soon.
Unsloth Minimax M3 GGUF
Unsloth is uploading a GGUF quantized version of the MiniMax M3 model to Hugging Face.
sakamakismile/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4
Release of an NVFP4-quantized uncensored MiniMax-H3 text encoder (Qwen3-VL-32B Heretic) that fits on a single 16GB GPU and serves as a drop-in replacement in ComfyUI workflows.
@TeksEdge: Unsloth released the fastest Qwen3.6-27B MTP GGUF I've tested. Time to upgrade. Compared to the previous GGUF, Q4/Q6 XL…
Unsloth has released an optimized GGUF version of the Qwen3.6-27B MTP model, achieving significantly faster inference speeds (up to 114 tok/s on an RTX 5090) compared to previous quantizations.