4-bit-quantization

Tag

Cards List
#4-bit-quantization

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

Hugging Face Daily Papers · 2026-08-21 Cached

Quantization-Aware Healing is a method that recovers compressed 4-bit language models by distilling directly from the original uncompressed model, offering faster and more stable performance than Quantization-Aware Training.

0 favorites 0 likes
#4-bit-quantization

Bringing Nunchaku 4-bit Diffusion Inference to Diffusers

Hugging Face Blog · 2026-07-23 Cached

Nunchaku, a 4-bit diffusion inference engine based on SVDQuant, is now natively integrated into Hugging Face Diffusers, enabling fast and memory-efficient loading of quantized diffusion models with a simple from_pretrained() call.

0 favorites 0 likes
#4-bit-quantization

nota-ai/Solar-Open2-250B-Nota-NVFP4

Hugging Face Models Trending · 2026-07-21 Cached

Nota AI releases a 4-bit quantized version of Upstage's Solar Open2 250B MoE model, using proprietary NVFP4 quantization that requires NVIDIA Blackwell GPUs.

0 favorites 0 likes
#4-bit-quantization

New wave of miniboss models you can run on dual DGX Spark

Reddit r/LocalLLaMA · 2026-07-15

A new wave of large language models including GLM 4.5, Qwen 3.5, MiniMax M2.7, Deepseek V4 Flash, Xiaomi MiMo 2.5, StepFun 3.7 Flash, and Tencent Hy3 can now be run locally on a dual DGX Spark setup with 250GB usable memory at 4-bit quantization, costing approximately $7,000–$8,000.

0 favorites 0 likes
#4-bit-quantization

Updates on North Mini Code: 4 bit quant + Ollama + OpenRouter

Reddit r/LocalLLaMA · 2026-06-18 Cached

Cohere releases North Mini Code, a 30B-A3B open-weights model with 4-bit quantization for code generation and agentic coding tasks, supporting 256K context.

0 favorites 0 likes
#4-bit-quantization

Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit

arXiv cs.LG · 2026-05-29 Cached

This paper adapts paired binary sample-size calculations to 4-bit quantization benchmarks, providing a conservative minimum detectable effect (MDE) bound that helps benchmark designers determine reliability before running experiments. A pilot audit shows that much of the observed variance across small subsamples is binomial sampling noise, not true model unreliability.

0 favorites 0 likes
#4-bit-quantization

@HuggingModels: Gemma 4 is here, and it's optimized for Apple Silicon. This 4-bit quantized model runs fast on your Mac, not just in th…

X AI KOLs Timeline · 2026-05-24 Cached

Gemma 4 is a 4-bit quantized model optimized for Apple Silicon, enabling fast local inference on Mac devices, reducing reliance on cloud computing.

0 favorites 0 likes
← Back to home

Submit Feedback