quantization-aware-training

Tag

Cards List
#quantization-aware-training

PETITION FOR QUANTIZATION AWARE TRAINING TO BE A NORM!!!

Reddit r/AI_Agents · 2026-09-02

The post questions why quantization-aware training is not a standard practice for open-weight AI models and asks about potential barriers like compute costs or performance impacts.

0 favorites 0 likes
#quantization-aware-training

Capability-Stratified Degradation in Ternary Language Models

arXiv cs.AI · 2026-09-01 Cached

This research paper analyzes capability-stratified degradation in ternary quantized language models, showing that while factual knowledge deteriorates significantly, commonsense reasoning and downstream task adaptability are retained, making the models viable for efficient edge deployment.

0 favorites 0 likes
#quantization-aware-training

QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

arXiv cs.CL · 2026-08-17 Cached

QUASAR is a quantization-aware training method that uses loss-aware reconstruction to lower the loss floor, improving low-bit model performance in large language models with significant accuracy gains at 2-4 bits.

0 favorites 0 likes
#quantization-aware-training

@TeksEdge: The most surprising thing to me about Google’s Gemma 4 technical report is how aggressively they optimized through heav…

X AI KOLs Timeline · 2026-07-08 Cached

Commentary on Google's Gemma 4 technical report, noting aggressive optimization via Quantization-Aware Training and Multi-Token Prediction for faster inference.

0 favorites 0 likes
#quantization-aware-training

Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models

arXiv cs.LG · 2026-07-07 Cached

Introduces Variable Bit-width Quantization (VBQ), a training-time method where each group of 64 weights learns its own bit-width (1,2,4,8) via Gumbel-Softmax relaxation. VBQ discovers a heterogeneous allocation that yields a 'bigger-but-smaller' regime, e.g., a 131M parameter model at 1.82 mean bits beats a 55M FP16 model while using less storage, and a 1.46B model matches a 593M FP16 with ~3.7x less storage.

0 favorites 0 likes
#quantization-aware-training

Gemma4-12B-QAT Uncensored Balanced is out with MTP (~60% speed boost)!

Reddit r/LocalLLaMA · 2026-06-22

Release of Gemma4-12B-QAT Uncensored Balanced, a fine-tuned uncensored model with a multi-token-prediction draft head for ~60% faster speculative decoding, optimized for llama.cpp and offering vision support.

0 favorites 0 likes
#quantization-aware-training

@_philschmid: Weights: https://huggingface.co/collections/google/gemma-4-qat-q4-0… Blog: https://blog.google/innovation-and-ai/techno…

X AI KOLs Following · 2026-06-08 Cached

Google released Gemma 4 models with quantization-aware training (QAT) at Q4_0 precision on Hugging Face, offering efficient variants from 5B to 33B parameters.

0 favorites 0 likes
#quantization-aware-training

@_philschmid: More Gemma 4! New QAT Gemma 4 checkpoints with similar performance while using ~4x less memory! It comes with a new mob…

X AI KOLs Following · 2026-06-08 Cached

New QAT Gemma 4 checkpoints offer similar performance with ~4x less memory, enabling a 1GB memory footprint for Gemma 4 E2B via a new mobile quantization format.

0 favorites 0 likes
#quantization-aware-training

Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

Hacker News Top · 2026-06-05 Cached

Google releases Gemma 4 models optimized with Quantization-Aware Training (QAT) to improve efficiency for mobile and laptop deployment, reducing memory footprint to 1GB for the E2B model while preserving quality.

0 favorites 0 likes
#quantization-aware-training

google/gemma-4-12B-it-qat-q4_0-gguf

Hugging Face Models Trending · 2026-06-05 Cached

Google DeepMind releases Gemma 4 models optimized with Quantization-Aware Training (QAT) in multiple formats including GGUF, enabling high quality with reduced memory requirements.

0 favorites 0 likes
#quantization-aware-training

Max-Window Scale Estimation for Near-Lossless HiF8 W8A8 Quantization-Aware Training

arXiv cs.LG · 2026-05-27 Cached

This paper systematically studies HiF8 W8A8 quantization-aware training for OpenPangu-Embedded-1B, identifying and addressing failure modes such as amax saturation and catastrophic forgetting, achieving near-lossless performance with a 64-step max-algorithm DTS strategy and a 500-step BF16 warmup.

0 favorites 0 likes
← Back to home

Submit Feedback