@0xSero: Just added 2 new model compressions: Hy3-FP8 & NVFP4 I recommend trying this model it's very strong and fits on 256gb o…
Summary
0xSero has released new FP8 and NVFP4 quantized versions of the Tencent Hy3-preview model, enabling it to run on 256GB VRAM with full context.
View Cached Full Text
Cached at: 05/10/26, 08:23 AM
Just added 2 new model compressions:
Hy3-FP8 & NVFP4
I recommend trying this model it’s very strong and fits on 256gb of vram with full context
https://t.co/UQI63BCFiJ
0xSero/Hy3-preview-NVFP4 · Hugging Face
Source: https://huggingface.co/0xSero/Hy3-preview-NVFP4
https://huggingface.co/0xSero/Hy3-preview-NVFP4#hy3-preview-nvfp4a16Hy3-preview NVFP4A16
This is a checkpoint-onlyNVFP4A16quantization oftencent/Hy3\-preview, produced withllmcompressor\.entrypoints\.model\_free\.model\_free\_ptq.
- Base model:
tencent/Hy3\-preview - Quantization scheme:
NVFP4A16 - Ignored modules/patterns:
lm\_head, model\.embed\_tokens, re:\.\*router\.gate$, re:\.\*expert\_bias$ - Source snapshot: recorded in
QUANTIZATION\_MANIFEST\.json - License: inherits Tencent Hy Community License Agreement from the base model; original
LICENSEis included.
https://huggingface.co/0xSero/Hy3-preview-NVFP4#notesNotes
This release quantizes safetensors weights without importing the custom HYV3 model class. Router gates, expert bias tensors, embeddings, and lm_head are preserved unquantized for compatibility/conservatism.
Similar Articles
sakamakismile/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4
Release of an NVFP4-quantized uncensored MiniMax-H3 text encoder (Qwen3-VL-32B Heretic) that fits on a single 16GB GPU and serves as a drop-in replacement in ComfyUI workflows.
NVFP4 kv cache quantization on sm120 will make 32GB VRAM systems very capable
NVFP4 KV cache quantization on sm120 significantly improves memory efficiency for large language models, enabling 32GB VRAM systems to achieve ~60 tok/sec inference at 196k context size with Qwen3.6-27B.
@0xSero: Best models for your hardware this week. 8-12GB - https://huggingface.co/LiquidAI/LFM2.5-8B-A1B… incredible model, so f…
A curated weekly roundup of the best AI models for different hardware configurations, from 8GB to 768GB VRAM, highlighting performance and benchmarks.
@HuggingPapers: NVIDIA just released the NVFP4 quantized Kimi-K2.7-Code on Hugging Face A 1T-parameter Moonshot AI model quantized to F…
NVIDIA released the NVFP4 quantized Kimi-K2.7-Code, a 1 trillion-parameter Moonshot AI model quantized to FP4 for Blackwell GPUs, preserving accuracy with reduced memory usage.
@0xSero: Best models for your hardware - 4gb to 12gb vram - VibeThinker-3B - smokes everything remotely close to its weight clas…
This thread recommends AI models optimized for different VRAM levels, highlighting VibeThinker-3B for its strong reasoning performance at 3B parameters, along with other models for coding and general use.