Tag
NVIDIA demonstrates how quantization-aware distillation (QAD) using NVIDIA Model Optimizer improves the Nemotron 3.5 Lightning model, reducing memory usage and increasing throughput while preserving accuracy for agentic benchmarks.
LiquidAI has released updated 4-bit Q4_0 checkpoints for LFM2.5 models using Quantization-Aware Distillation, recovering 97% of accuracy lost to quantization while maintaining low memory and high speed.