Softmax Reparameterization for Output-Head Quantization

Hugging Face Daily Papers Papers

Summary

This paper proposes softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization to reduce inference cost in small language models. The method demonstrates significant improvements across various quantization techniques and datasets.

Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. This one-dimensional search includes the original head and fixed mean-centering, preserves the full-precision softmax distribution, and leaves the trained decoder unchanged; a rank-one correction handles nonlinear logit paths such as soft-capping. Across seven heads, W4 gains concentrate where baseline quantization substantially distorts predictions: on Phi-4-mini, AW-MSE KL falls from 0.936 to 0.256. The gains survive stronger GPTQ calibration and remain complementary to exact per-channel scaling and affine quantization. Across four heads and three W4 quantizers, frozen WikiText-selected coefficients also transfer to C4 and OpenWebMath, outperforming mean-centering in all 18 comparisons where the frozen coefficient differs from 1 and matching it in the remaining six. At W2, used as a compression stress test, benefits broaden across nearly the full model--quantizer matrix. Matched residual analysis shows that improved fidelity can accompany greater logit reconstruction error while reducing the residual's Fisher-weighted cost. For shift-compatible heads, reparameterization adds no inference operation and preserves packed W4 execution: with the decoder held in BF16, quantizing the Phi output head reduces batch-one generation latency by 10.8% relative to the BF16-head baseline.
Original Article
View Cached Full Text

Cached at: 09/28/26, 04:05 PM

Paper page - Softmax Reparameterization for Output-Head Quantization

Source: https://huggingface.co/papers/2609.31291

Abstract

Largevocabulariesmakeoutputheadsasubstantialinferencecostinsmalllanguagemodels.Weproposesoftmaxreparameterization,apost-trainingmethodthatselectsafunctionallyequivalentoutputheadbeforequantization.Themethodsubtractsascalarmultipleofthevocabulary-rowmeanfromeveryoutputrowandselectsthecoefficientbyvalidationKLseparatelyforRTN,activation-weightedMSE,andfull-HessianGPTQ.Thisone-dimensionalsearchincludestheoriginalheadandfixedmean-centering,preservesthefull-precisionsoftmaxdistribution,andleavesthetraineddecoderunchanged;arank-onecorrectionhandlesnonlinearlogitpathssuchassoft-capping.Acrosssevenheads,W4gainsconcentratewherebaselinequantizationsubstantiallydistortspredictions:onPhi-4-mini,AW-MSEKLfallsfrom0.936to0.256.ThegainssurvivestrongerGPTQcalibrationandremaincomplementarytoexactper-channelscalingandaffinequantization.AcrossfourheadsandthreeW4quantizers,frozenWikiText-selectedcoefficientsalsotransfertoC4andOpenWebMath,outperformingmean-centeringinall18comparisonswherethefrozencoefficientdiffersfrom1andmatchingitintheremainingsix.AtW2,usedasacompressionstresstest,benefitsbroadenacrossnearlythefullmodel--quantizermatrix.Matchedresidualanalysisshowsthatimprovedfidelitycanaccompanygreaterlogitreconstructionerrorwhilereducingtheresidual’sFisher-weightedcost.Forshift-compatibleheads,reparameterizationaddsnoinferenceoperationandpreservespackedW4execution:withthedecoderheldinBF16,quantizingthePhioutputheadreducesbatch-onegenerationlatencyby10.8%relativetotheBF16-headbaseline.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.31291

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.31291 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.31291 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.31291 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization

arXiv cs.AI

ReQuant introduces a backpropagation-free, fixed-grid discrete refinement stage for post-training quantization (PTQ) that iteratively improves initial quantized models while preserving the quantized format, showing consistent gains across various LLMs and bit-widths.

Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models

arXiv cs.LG

Introduces Variable Bit-width Quantization (VBQ), a training-time method where each group of 64 weights learns its own bit-width (1,2,4,8) via Gumbel-Softmax relaxation. VBQ discovers a heterogeneous allocation that yields a 'bigger-but-smaller' regime, e.g., a 131M parameter model at 1.82 mean bits beats a 55M FP16 model while using less storage, and a 1.46B model matches a 593M FP16 with ~3.7x less storage.