Softmax Reparameterization for Output-Head Quantization
Summary
This paper proposes softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization to reduce inference cost in small language models. The method demonstrates significant improvements across various quantization techniques and datasets.
View Cached Full Text
Cached at: 09/28/26, 04:05 PM
Paper page - Softmax Reparameterization for Output-Head Quantization
Source: https://huggingface.co/papers/2609.31291
Abstract
Largevocabulariesmakeoutputheadsasubstantialinferencecostinsmalllanguagemodels.Weproposesoftmaxreparameterization,apost-trainingmethodthatselectsafunctionallyequivalentoutputheadbeforequantization.Themethodsubtractsascalarmultipleofthevocabulary-rowmeanfromeveryoutputrowandselectsthecoefficientbyvalidationKLseparatelyforRTN,activation-weightedMSE,andfull-HessianGPTQ.Thisone-dimensionalsearchincludestheoriginalheadandfixedmean-centering,preservesthefull-precisionsoftmaxdistribution,andleavesthetraineddecoderunchanged;arank-onecorrectionhandlesnonlinearlogitpathssuchassoft-capping.Acrosssevenheads,W4gainsconcentratewherebaselinequantizationsubstantiallydistortspredictions:onPhi-4-mini,AW-MSEKLfallsfrom0.936to0.256.ThegainssurvivestrongerGPTQcalibrationandremaincomplementarytoexactper-channelscalingandaffinequantization.AcrossfourheadsandthreeW4quantizers,frozenWikiText-selectedcoefficientsalsotransfertoC4andOpenWebMath,outperformingmean-centeringinall18comparisonswherethefrozencoefficientdiffersfrom1andmatchingitintheremainingsix.AtW2,usedasacompressionstresstest,benefitsbroadenacrossnearlythefullmodel--quantizermatrix.Matchedresidualanalysisshowsthatimprovedfidelitycanaccompanygreaterlogitreconstructionerrorwhilereducingtheresidual’sFisher-weightedcost.Forshift-compatibleheads,reparameterizationaddsnoinferenceoperationandpreservespackedW4execution:withthedecoderheldinBF16,quantizingthePhioutputheadreducesbatch-onegenerationlatencyby10.8%relativetotheBF16-headbaseline.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.31291
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.31291 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.31291 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.31291 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
PRQuant: Permutation Residual Quantization for Low-Overhead Inference
PRQuant is a training-free and low-overhead framework for quantizing linear layers in large language models, using permutation and residual compensation to reduce inference latency while improving accuracy over baselines like MXFP4.
MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models
MixQuant proposes an adaptive mixed-precision quantization framework for LLMs that handles variable memory budgets by marginalizing layer distortion over random upstream configurations, outperforming existing methods across multiple models and budgets.
ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization
ReQuant introduces a backpropagation-free, fixed-grid discrete refinement stage for post-training quantization (PTQ) that iteratively improves initial quantized models while preserving the quantized format, showing consistent gains across various LLMs and bit-widths.
Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models
Introduces Variable Bit-width Quantization (VBQ), a training-time method where each group of 64 weights learns its own bit-width (1,2,4,8) via Gumbel-Softmax relaxation. VBQ discovers a heterogeneous allocation that yields a 'bigger-but-smaller' regime, e.g., a 131M parameter model at 1.82 mean bits beats a 55M FP16 model while using less storage, and a 1.46B model matches a 593M FP16 with ~3.7x less storage.
Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal Residuals
This paper presents a decomposition framework for quantization error in large language models, separating it into activation-guided weight compensation and orthogonal residual, and derives practical guidelines for improving W4A4 quantization through techniques like Hadamard rotation and sign selection.