A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models

arXiv cs.LG Papers

Summary

This paper introduces a composite metric for quantization in small large language models that balances speed and quality by combining information retention and throughput gains, evaluated on models like Gemma 3 1B.

arXiv:2608.26926v1 Announce Type: new Abstract: Small language models (sLLMs) are nowadays hosted on devices with limited memory and computational budget. In an autoregressive setup, inference is memory-bandwidth bound: uniform quantization is often detrimental to such models, since their architecture has limited redundancies and only a few layers are not very sensitive to lower precision. We propose a composite metric that combines two orthogonal criteria: information retention (measured in terms of a normalized SQNR-based coefficient) and throughput gains (modeled using a roofline-based latency analysis). By profiling Gemma 3 1B, we find that Feed-Forward Network blocks and the embedding matrix are the most promising targets for acceleration. For each candidate, we estimate a normalized quality score based on simulated quantization and a normalized speed score based on roofline modeling with no actual execution needed. We combine the two scores in a composite priority coefficient, allowing us to tune the trade-off between speed and quality as needed. Our metric is general and can be used to prioritize individual blocks, their projection sublayers, or transformer layers as a whole. We evaluate our approach on several model architectures, showing that our estimates have at around 4% prediction error for the accelerated speedup. We find that our method generally allocates more resources to the most expressive layers compared to evolutionary search, specialized accelerators, or Shapley-value-based approaches that require expensive approximate inference. Our analytical approach makes sLLM quantization a predictable engineering task.
Original Article
View Cached Full Text

Cached at: 08/28/26, 09:45 AM

# A Layer Importance Metric for Quantization Accounting for the Speed–Quality Trade-off in Autoregressive Models
Source: [https://arxiv.org/html/2608.26926](https://arxiv.org/html/2608.26926)
###### Abstract

Small large language models \(sLLMs\) are now deployed on devices with limited memory and compute budget\. Within the autoregressive setting, inference is memory\-bandwidth bound: typical homogeneous quantization is typically detrimental to such models due to their compactness, as their architectures do not retain much redundancy, and in particular a handful of components display extremely poor tolerance to lower precision\.

This paper introduces a composite metric that combines two orthogonal assessment criteria for such components: the information retention after quantization, captured by a normalized coefficient, and the gains in throughput that can be expected from accelerating individual layers, analytically estimated using a roofline\-based latency model\. Using an architectural profiling benchmark on Gemma 3 1B, we identify the Feed\-Forward Network \(FFN\) blocks and the embedding matrix as the most promising candidates for accelerated inference, and we estimate for each candidate a normalized quality score, computed via simulated quantization, and a normalized speed score, computed without any model execution\. The metric uses a single prioritization coefficient to combine the two scores into a composite value, which can be tuned to prioritize speed vs quality according to a specific application’s needs\. The metric is general and can be applied to blocks of layers \(such as FFNs in the transformer\), their internal projections \(such as the feed forward network within a transformer block\), or single transformer layers, with no modifications to the computation\.

We evaluated our method on multiple model architectures, measuring around 4% error in predicted speedups on average, and demonstrated that it enables significantly better resource allocation decisions than existing approaches, which usually rely on extensive evolutionary search strategies, domain\-specific inference accelerators, or require access to expensive Shapley\-value approximations\. In short, our analytical approach makes sLLM quantization both faster, easier, and more intuitive, turning a trial\-and\-error process into a mathematical exercise\.

Keywords:quantization, small language models, sLLM, roofline model, SQNR, inference acceleration, model compression

## 1Introduction

The current practice in large language model \(LLM\) design appears to be contradictory, in that while model architectures become increasingly complex, their physical realization is subject to rather strict limitations imposed by memory bus widths, battery capacity, or the available silicon area\. AI assistants are now omnipresent in smartphones and other local machines where they perform a variety of non\-critical tasks such as content creation and web browsing\[[23](https://arxiv.org/html/2608.26926#bib.bib23)\]\. Yet even in this setting, when an LLM is used for generic\-purpose assistance, the emphasis is placed on the model’s ability to operate in a memory\-constrained environment rather than perform hundreds of billions of operations per second; put otherwise, a compactly represented LLM is significantly more valuable than a large one at the same level of functional complexity\. Now, the critical question is: how to achieve a high degree of model reduction while preserving the response quality, as witnessed by the recent advances in knowledge distillation and parameter sharing\.

It is not accidental that currently available small LLMs \(Llama 3 1B/3B, Gemma 3 1B, Qwen 3 1\.7B, etc\.\) are the result of aggressive model distillation: a large\-scale LLM possesses ample resources to withstand the pruning of its parameter count, which has little effect on its accuracy, while an sLLM \(small LLM\) does not have these advantages\. If post\-training quantization \(PTQ\) is applied to an sLLM, it may not produce acceptable results: the internal structure of the model remains intact, but its behavioral patterns are altered, so that the response quality decreases substantially\.

Another issue is that modern accelerators are often memory\-bandwidth\-constrained\[[3](https://arxiv.org/html/2608.26926#bib.bib3)\]: the majority of the time spent in autoregressive decoding is occupied by offloading weights from video memory, not performing arithmetic operations, which also appears to be the case when naively implementing the attention mechanism\[[16](https://arxiv.org/html/2608.26926#bib.bib16)\]\.

Thus, the utilization of a GPU in this scenario is significantly lower than could be the case, and the user must wait for the next token to be generated\.

A uniform reduction in the precision of weights and activations, applicable to each layer of the model, is not the only possibility; indeed, different network layers may respond differently to such a reduction\. Some layers are relatively inert and can safely undergo aggressive quantization without noticeable effects on the quality of generation, while others store crucial information that should be left untouched\. Several methods have been developed to date that allow one to quantify different elements of the model with varying degrees of precision\. However, each of them is limited by certain design decisions: CARVQ\[[4](https://arxiv.org/html/2608.26926#bib.bib4)\]and IMPQ\[[6](https://arxiv.org/html/2608.26926#bib.bib6)\]utilize multi\-level vector quantization or require extensive computation to estimate each component’s contribution to the total model performance using Shapley\-value estimation\[[21](https://arxiv.org/html/2608.26926#bib.bib21)\], while HAPM\[[5](https://arxiv.org/html/2608.26926#bib.bib5)\]requires specialized inference software to demonstrate its effectiveness\. Thus, none of these approaches provides a practical solution for easily identifying the set of layers within an sLLM that could benefit from higher precision\.

The paper suggests that in order to achieve the best possible trade\-off between performance and quality, one must move away from uniform quantization and instead attempt to find groups of operations within the model that contribute the most to its quality\. In other words, it is more useful to spend computational resources \(measured in FLOPS\) on operations that preserve information rather than discard it\.

We propose a method by which one can assess the quality of an sLLM quantization configuration using a single metric that estimates the cost\-quality trade\-off\. The paper reports the following results:

- •profiling the Gemma 3 1B architecture to identify the blocks that contribute most to inference latency, and justifying the choice of the FFN and embedding matrix as priority quantization targets;
- •obtaining a normalized quality score for the FFN and Embedding blocks via the SQNR coefficient using the SA\-PTQ operator;
- •obtaining a normalized speed score for each block based on an analytical roofline model, without re\-running the model;
- •combining both scores into an integrated metric with a priority parameterα\\alpha, allowing the balance between speed and quality to be shifted depending on the requirements of a specific task\.

The analysis presented confirms the applicability of the metric on standard hardware without quantizing the model, and substantiates the possibility of extending it to layer\-wise heterogeneous quantization\.

Goal of the present work\.The goal is to develop an analytical metric for estimating layer importance that accounts for the trade\-off between inference speed and generation quality in autoregressive models\. This tool is intended to identify structural blocks and layers whose quantization will provide the maximum hardware gain at minimal information loss\.

### 1\.1Domain analysis

To justify the necessity of the methodology for the choice of blocks presented in the paper, it is essential to explain the technological context of optimization\. Namely, the common industry practice which allows to reduce VRAM usage and thus accelerate inference without retraining is model quantization \- a process of reducing the bit width of weights to 4 or 8 bits\. It goes without saying that quantization is never the sole optimization step \- it rather represents the final stage when most of the optimizations are already performed\. And the issue with optimization is not quantitative \- the tooling for quantitative analysis is lacking\. In other words, a significant amount of blocks can be subjected to quantization without affecting the performance of the model, while some blocks cannot\. This is why quantization as a concept is not the ultimate goal but only the stage that requires analysis\. And this is the analysis that the given paper provides with, which is why quantization is framed as a task to solve and not a goal to achieve\.

It is vital to understand that ultimately, the objective of the proposed metric is to prioritize model blocks and layers in terms of information robustness\. This way, all blocks that require high precision to maintain their functional capabilities are separated from those that allow for aggressive optimization with little loss in terms of accuracy\. And this prioritization is ultimately the basis for heterogeneous models which can provide the most significant speed gains with the least losses in performance\.

### 1\.2Related work

Small language models have enabled to run LLMs on the end\-user device\. However, due to the nature of the Transformer architecture, generating tokens is a slow process, and an unquantized model is often too slow to be useful\[[22](https://arxiv.org/html/2608.26926#bib.bib22)\]\.

Quantization, which reduces the bit\-width of weights, is the standard method to tackle the problem\. However, it is not trivial to apply existing methods to smaller models\.

Compared to large models, small ones are more information\-dense: they do not have the luxury of discarding some weights as not useful\. As a result, applying the same compression to every layer harms the usefulness of the model more than its size\.

Several papers addressed this issue from different angles\. CARVQ aims to tackle the problem of embedding matrix size by compressing it to around 1\.6 bits using a corrective adaptor and group residual vector quantization\[[4](https://arxiv.org/html/2608.26926#bib.bib4)\]\. The authors note that, after PTQ of the transformer itself, the embedding matrix comprises the largest portion of the model \(52% for LLaMA\-3\.2\-1B\)\[[4](https://arxiv.org/html/2608.26926#bib.bib4)\], so it should be compressed aggressively\. CARVQ requires lookup tables and nonlinear transformations to be applied at inference time, which complicates the integration with other code\. The metric proposed in this paper, in contrast, utilizes conventional scalar quantization \(SA\-PTQ\) and measures the SQNR of the compressed weights, without requiring any additional structures\. As a result, it can be used in existing tools such as llama\.cpp\.

HAPM paper\[[5](https://arxiv.org/html/2608.26926#bib.bib5)\]suggests another approach, prunning the weights of the language models on mobile phones with hardware\-aware latency modeling\. They introduce the concept of real latency sensitivity, which is linked directly to the accuracy drop of a pruned model\. The structure of the resulting model becomes sparse, requiring specialized inference machinery to attain the reported speed gains\. TheSnS\_\{n\}metric from the current paper does not depend on model structure and can be applied to any model\. It measures the information contained in the weights, normalized by the time spent on processing them \(Δ​L\\Delta L\)\. It achieves the desired speed gains by eliminating lower\-ranked layers outright rather than changing their structure, and can utilize existing accelerators\.

IMPQ paper\[[6](https://arxiv.org/html/2608.26926#bib.bib6)\]analyzes interactions between layers using Shapley values from game theory to determine the importance of different weight tensors\. While it provides theoretical justification for the layer ranking, the computation of the Shapley values is expensive to apply to models with tens of billions of weights, as in the LLaMA series\.SnS\_\{n\}, on the other hand, only needs to perform a relatively quick analytical probing of the model to approximate the importance of each layer\. This allows to spend less time on model analysis and focus on finding the best compression option\.

QRazor paper\[[7](https://arxiv.org/html/2608.26926#bib.bib7)\]combines two\-step weight and activation compression with Significant Data Razoring to compress the LLM by removing the least\-significant bits\. They modify the Attention blocks, but the authors claim that their changes to the performance are minimal\. Unlike QRazor, which modifies the Attention blocks directly, this work focuses primarily on FFN and Embedding: efficient attention variants such as GQA\[[20](https://arxiv.org/html/2608.26926#bib.bib20)\]already reduce its memory footprint architecturally to some extent, and our tests \(Section 2\) showed that Attention contributes comparatively little to per\-token latency in the configurations we profiled, making it a lower priority target here\. QRazor uses a heuristic based on amplitudes of the weights to rank the importance of bits, whileSnS\_\{n\}uses an energy\-based SQNR metric to approximate the value of each bit more accurately\. This allows to select the most valuable bits to keep, providing the flexibility to choose the desired generation quality/bit\-width balance\.

### 1\.3Problem statement

The goal of the present study is to develop a metric for evaluating neural network components that allows ranking them by significance for inference acceleration while preserving generation quality\. Formally, the task can be represented as a search for a heterogeneous quantized model configuration that provides maximum performance under a given quality\-loss threshold\.

Formal description of parameters\.Let the original modelMMbe represented as an ordered set of components\{C1,C2,…,Cn\}\\\{C\_\{1\},C\_\{2\},\\ldots,C\_\{n\}\\\}, where a component may be an entire structural block \(FFN, Embedding\), an individual projection within a block \(Wg​a​t​e,Wu​p,Wd​o​w​nW\_\{gate\},W\_\{up\},W\_\{down\}\), or a separate transformer layer\. The quantization process consists in mapping each layer to the space of admissible bit widthsB=\{b4,b8,b16\}B=\\\{b\_\{4\},b\_\{8\},b\_\{16\}\\\}\. The state of the system is characterized by the following functions:

- •L⁡\(Li,b\)L\(L\_\{i\},b\)– the latency function, defining the inference execution time of blockLiL\_\{i\}at bit widthbb;
- •Q⁡\(M,\{b1,…,bn\}\)Q\(M,\\\{b\_\{1\},\\ldots,b\_\{n\}\\\}\)– the integral quality indicator \(functional robustness\), characterizing the deviation of the quantized model’s output signal from the reference\.

To find the optimal bit\-width distribution\{b1,b2,…,bn\}\\\{b\_\{1\},b\_\{2\},\\ldots,b\_\{n\}\\\}, we formulate an optimization problem\. The objective function, aimed at minimizing total inference time \(the efficiency criterion\), is given by \(1\)\.

∑i=1n\(LM​L​P\(i\)​\(bi\)\+LE​m​b​\(b\)\)⟶min\\sum\_\{i=1\}^\{n\}\\left\(L\_\{MLP\}^\{\(i\)\}\(b\_\{i\}\)\+L\_\{Emb\}\(b\)\\right\)\\longrightarrow\\min\(1\)
At the same time, the system is subject to a boundary condition \(2\), defining the maximum admissible degradation of the information signal \(the quality constraint\)\.

S​Q​N​Rt​o​t​a​l≥S​Q​N​Rb​a​s​e​l​i​n​e−ϵSQNR\_\{total\}\\geq SQNR\_\{baseline\}\-\\epsilon\(2\)
where expression \(1\) defines the efficiency criterion, and inequality \(2\) fixes the maximum admissible signal\-energy lossϵ\\epsilon, expressed in decibels \(dB\)\.

The result of this formalization is the development of an integral selectivity criterions​c​o​r​e​\(α\)score\(\\alpha\)\(Fig\.[1](https://arxiv.org/html/2608.26926#S1.F1)\)\.

![Refer to caption](https://arxiv.org/html/2608.26926v1/figures/image1.png)Figure 1:Algorithmic scheme for constructing the selectivity metrics​c​o​r​e​\(α\)score\(\\alpha\)\.*Source: compiled by the author\.*This criterion combines two normalized components – the information preservation of a componentQkQ\_\{k\}and its speed potentialSkS\_\{k\}– through the priority parameterα\\alpha, allowing a heterogeneous architecture to be formed that is balanced with respect to speed and generation quality\.

## 2Formalizing the Latency Parameters of FFN and Embedding

For an effective compression\-evaluation metric, one has to move away from an abstract consideration of weight counts to the concrete impact each component’s size has on the resulting inference speed\. And indeed, when speaking of sLLMs, the relevant metric is not the raw power, but the latency introduced by reading data from memory\. A cost metric, discussed here, aims to provide a means of precise prediction by answering the question of what “gain” in performance, measured in seconds saved, is projected from reducing the bit\-width of a particular layer\. This metric turns the vague question of how many layers to quantize into an exacting calculation, the parameters of which are defined by the desired acceleration and the quality of the signal preserved\.

Beforehand, however, a profiling operation was performed on the Gemma 1B model to identify the major contributors to token\-generation latency\. The metric of choice, for this particular study, was simply the token generation speed, as it defines the applicability of a quantized model in practice\. The blocks under review were 8\-bit quantized to identify the primary sources of overhead, which were then isolated for further analysis\.

The operations of profiling and identifying computational dominants were performed in Google Colab with T4 GPU using the resources of the llama\.cpp toolkit\. In particular, thellama\-quantizeutility, which allows for changing the bit\-width of individual blocks by means of typecasting particular tensors, was used\. Gemma 3 1B was quantized to Q8\_0 for the relevant parameter groups: FFN, Attention mechanism layers \(Q, K, V, Out\) and Embedding matrix, with F16\-weights model serving as the reference point\. The metrics for each configuration were gathered automatically with the help of thellama\-cpp\-pythonlibrary with at least three runs for each set of parameters to eliminate hardware variation noise\.

The characteristics of architectural performance were defined with the consideration of three factors: token generation speed \(tokens/s\), context processing latency and dynamic VRAM consumption\. The latter, in particular, was evaluated by means of direct reading of the VRAM utilization percentages from thenvidia\-smiutility with 0\.05s sampling frequency throughout the duration of calculations\. Diversified input sequences were used for the calculations to observe the behavior of the model under both light \(short queries\) and heavy \(8192 tokens, typical for gaming sessions\) loads\.

The results of the profiling runs \(see Table[1](https://arxiv.org/html/2608.26926#S2.T1)\) confirm the nonlinear nature of the relationship between the bit\-width of model sections and their impact on the inference speed\. In particular, the analysis demonstrates that the contribution of the FFN blocks and the Embedding matrix to the overall performance is disproportionately high compared to the input dimensions: thus, the speed gain from reducing their bit\-depth from F16 to Q8 is substantially larger for both than for Attention layers\.

On the other hand, the comparison of the configurations with reduced precision of FFN and Embedding \(“Q8 Q8 Q8”\) and the baseline \(“Q8 ffn\_all=F16”\) show that the omission of these blocks from quantization results in significantly decreased performance of the 8\-bit model\. In particular, for the short context length, the token generation speed decreases to 72\.3 tok/s for the latter configuration, compared to the fully quantized alternative\. And while the contribution of the embedding layer is less obvious, the drop in throughput for completely unquantized Embedding is still substantial\.

By comparison, the impact of the Attention layers is significantly lower: even with a similar reduction in precision \(from F16 to Q8\), the loss of performance from omitting these blocks is smaller for both context lengths\. Altogether, the observations confirm the hypothesis that in the Gemma 3 1B architecture, the FFN blocks and the Embedding matrix are the critical performance factors\. Their contribution to the overall latency dominates over other sources of overhead and, therefore, should be prioritized when selecting layers for quantization in a local deployment\.

Table 1:Performance measurements for different quantization variants of Gemma 3 1B blocks \(FFN/Attn/Emb\) at context lengths of 10, 50, and 100 tokens\.The results obtained justify the selection of these blocks as priority optimization targets\. Thus, the search for a balance between speed gain and accuracy preservation should be focused on these segments, which served as the direct premise for applying modification metrics to them\.

Analytical latency metrics\.This work uses a simulation of uniform quantization – an activation\-quantization metric that also accounts for sensitivity, i\.e\., Sensitivity\-Aware PTQ \(SA\-PTQ\) – which moves the optimization process from the realm of guesswork into the domain of precise computation and energy\-based signal analysis via the SQNR coefficient\. This also forgoes labor\-intensive latency profiling in favor of an analytical hardware\-cost metric of the layers \(LF​F​NL\_\{FFN\},LE​m​bL\_\{Emb\}\)\. The integrated selectivity metricSnS\_\{n\}formed in the course of the study reveals the “information endurance” of each block, providing maximum speed gain under strict quality control\.

### 2\.1Theoretical Basis for Layer Sensitivity Estimation

The optimization problem for language models \(sLLMs\) is posed by their parameter set’s information density: unlike huge models, small ones have little or no redundancy, and even minor distortions produce significant accuracy drops\. The traditional Quantization\-Aware Training \(QAT\) is prohibitively expensive for sLLMs due to requiring substantial resources for quantization\[[8](https://arxiv.org/html/2608.26926#bib.bib8),[12](https://arxiv.org/html/2608.26926#bib.bib12)\], while the canonical Post\-Training Quantization \(PTQ\) methods\[[1](https://arxiv.org/html/2608.26926#bib.bib1),[9](https://arxiv.org/html/2608.26926#bib.bib9),[10](https://arxiv.org/html/2608.26926#bib.bib10),[13](https://arxiv.org/html/2608.26926#bib.bib13),[17](https://arxiv.org/html/2608.26926#bib.bib17),[19](https://arxiv.org/html/2608.26926#bib.bib19)\], reviewed in detail in\[[14](https://arxiv.org/html/2608.26926#bib.bib14)\], apply “blind” uniform compression, destroying the connectivity structure\. In this work, we propose the Sensitivity\-Aware PTQ metric that evaluates a layer’s information “endurance” by quantizing its parameters and measuring the output vector’s norm deviation\. By avoiding the calculation of the Hessian matrix, required for Hessian\-aware mixed\-precision methods\[[15](https://arxiv.org/html/2608.26926#bib.bib15)\]and too expensive for realistic settings, the sensitivity estimation step transforms the model into a heterogeneous structure in terms of block\-wise bit\-width, providing excellent speed gains while maintaining the quality within the given constraints\.

The choice of PTQ as the basis for our method is motivated by the need for quick adaptation of the model to the target hardware without additional training\[[11](https://arxiv.org/html/2608.26926#bib.bib11)\]\. Unlike traditional PTQ, which assumes uniform weight compression and is optimized to bound the parameter error, our SA\-PTQ focuses on the output signal, enabling the analytical treatment of the degradation, which is crucial for the sensitivity analysis of sLLMs\.

The SA\-PTQ algorithm is a mathematical function that converts a smoothly varying number into a stepped approximation and vice versa\. This procedure involves two operations: Quantization and Dequantization\.

The step width calculation: the distance between grid points is computed based on the number of bits used to represent the numbers\. Forbbbits \(e\.g\.,b=4b=4\), the grid has2b−12^\{b\}\-1intervals: ForΔ\\Delta\-bit precision \(e\.g\.,bb\), the value range ofb=4b=4is divided into2b−12^\{b\}\-1equal intervals \(i\.e\., for 4 bits, the range would be divided into 15 equal segments\):

Δ=max⁡\(x\)−min⁡\(x\)2b−1\\Delta=\\frac\{\\max\(x\)\-\\min\(x\)\}\{2^\{b\}\-1\}\(3\)
The step \(4\) serves as the unit of measurement for the entire subsequent process\. Knowing it, we can derive the complete SA\-PTQ mapping, which describes the number\-transformation cycle:

SA\-PTQ​\(x,b\)=round​\(x−min⁡\(x\)Δ\)⋅Δ\+min⁡\(x\)\\text\{SA\-PTQ\}\(x,b\)=\\text\{round\}\\left\(\\frac\{x\-\\min\(x\)\}\{\\Delta\}\\right\)\\cdot\\Delta\+\\min\(x\)\(4\)
Here, in the expression, theroundfunction discards insignificant signal components, leaving only the informational skeleton corresponding to the chosen bit width\. Depending on the data type inside the transformer architecture, this base logic is adapted for specific tasks\.

#### 2\.1\.1Formula variant for FFN \(weight quantization\)

The FFN block \(also referred to as MLP in some architectures, hence theM​L​PMLPsubscript used in the formulas below\) in the transformer architecture is traditionally viewed as the model’s “memory”: it is where the factual information learned by the model during training is stored\. In the Gemma 3 architecture, the fully connected block is implemented as a SwiGLU variant\[[18](https://arxiv.org/html/2608.26926#bib.bib18)\]: instead of a standard two\-layer perceptron, the output of the block is given by the element\-wise product of two parallel projections, one of which is passed through a nonlinearity\. While such a construction is more expressive than a standard perceptron, it becomes more challenging to analyze its sensitivity to individual matrix transformations: modifying any of the three matrices \(Wg​a​t​e,Wu​p,Wd​o​w​nW\_\{gate\},W\_\{up\},W\_\{down\}\) involved introduces non\-linear distortions to the output of the FFN block\.

From the point of view of quantization, the FFN block is also of particular interest: the weights of this block comprise the majority of the model’s weights, and their compression directly impacts the speed of model inference\. At the same time, the hypothesis about the uniformity of sensitivity is by no means obvious: the behavior of early and late layers in the model can be significantly different\.

The weights of the FFN layer are quantized using the SA\-PTQ operator \(4\)\. For the per\-channel scheme, the parametersΔ\\Deltaandmin⁡\(W\)\\min\(W\)are computed for each row of the matrix independently, thus minimizing the expected quantization error \(5\)\.

W^=round​\(Wi−min⁡\(Wi\)Δi\)⋅Δi\+min⁡\(Wi\)\\widehat\{W\}=\\text\{round\}\\left\(\\frac\{W\_\{i\}\-\\min\(W\_\{i\}\)\}\{\\Delta\_\{i\}\}\\right\)\\cdot\\Delta\_\{i\}\+\\min\(W\_\{i\}\)\(5\)
whereΔi=max⁡\(Wi\)−min⁡\(Wi\)2b−1\\Delta\_\{i\}=\\dfrac\{\\max\(W\_\{i\}\)\-\\min\(W\_\{i\}\)\}\{2^\{b\}\-1\}is the quantization step of rowii\. As a pessimistic baseline variant, a per\-tensor scheme is considered, in which a singleΔ\\Deltais computed from the global maximum of the entire matrix\.

To evaluate degradation at eachnn\-th layer, the reference and distorted outputs of the FFN block are computed\. The Gemma 3 architecture uses SiLU gating, so the output is formed as in \(6\)\.

yclean=Wdown⋅\(σ⁡\(Wgate​h\)⋅Wup​h\),ydirty=W^down⋅\(σ⁡\(W^gate​h\)⋅W^up​h\)y\_\{\\text\{clean\}\}=W\_\{\\text\{down\}\}\\cdot\\left\(\\sigma\(W\_\{\\text\{gate\}\}h\)\\cdot W\_\{\\text\{up\}\}h\\right\),\\qquad y\_\{\\text\{dirty\}\}=\\widehat\{W\}\_\{\\text\{down\}\}\\cdot\\left\(\\sigma\(\\widehat\{W\}\_\{\\text\{gate\}\}h\)\\cdot\\widehat\{W\}\_\{\\text\{up\}\}h\\right\)\(6\)
whereh∈ℝdh\\in\\mathbb\{R\}^\{d\}is the hidden vector of the last token, andσ⁡\(x\)=x1\+e−x\\sigma\(x\)=\\dfrac\{x\}\{1\+e^\{\-x\}\}is the SiLU activation function\. The resulting vectors are fed into the SQNR formula \(7\), forming a layer\-wise sensitivity profile of the model with respect to FFN weight quantization\. Becauseyc​l​e​a​n\(n\)y\_\{clean\}^\{\(n\)\}andyd​i​r​t​y\(n\)y\_\{dirty\}^\{\(n\)\}are computed independently per layer, the result is a full layer\-wise sensitivity profile, from which the quality drop for any subset of quantized layers can be predicted\.

#### 2\.1\.2Formula variant for Embedding

The question of quantizing language models always involves identifying the blocks that allow for aggressive compression and, therefore, are not sensitive to precision and blocks that are extremely sensitive to precision loss and, therefore, should not be compressed\. The Attention block traditionally falls into the second category since the attention mechanism demonstrates fragile behavior at low bit\-width and insufficient compression efficiency\. As a result, this block is typically kept at a higher precision\. By contrast, the Embedding matrix is a good candidate for compression since its weights are fixed and independent from the input context\. Moreover, given the large vocabulary size, the embedding matrix in Gemma 3 1B constitutes a considerable portion of the model’s weights\. Thus, reducing its size would lead to significant memory savings and acceleration of the model\. However, this hypothesis requires empirical verification: up to which bit\-width can the embedding matrix be safely compressed without affecting the quality of generated tokens? This value can be identified by quantizing the embedding matrix and measuring the tokenization quality either by direct evaluation or some proxy metric that does not require full model quantization\.

The output\-layer \(LM Head\) weights are responsible for converting the hidden state of the last token into a logit distribution over the vocabulary, from which the next token is selected\. Accordingly, LM Head weights directly impact the quality of generated text: their incorrect quantization would result in a significant degradation of the logit distribution\. The current work evaluates the sensitivity of the LM Head weights to low\-bitwidth quantization\. Notably, this analysis is performed after model training, assuming that the LM Head weights are not fine\-tuned during quantization\. Accordingly, the impact of reduced precision on the output\-layer performance can be estimated using two proxy metrics: SQNR, which characterizes the distortion of the logit vector, and the gap metricΔg​a​p\\Delta\_\{gap\}\.

The output\-layer weightsWWare quantized using the SA\-PTQ operator \(4\), which is applied to the tensor directly, as in the case of fully connected layers \(5\)\. In the per\-channel variant, the parametersΔ\\Deltaandmin⁡\(W\)\\min\(W\)are computed for each row of the weight matrix, thus allowing each row to have its own dynamic range and minimizing the quantization error \(7\)\.

Wq=round​\(Wi−min⁡\(Wi\)Δi\)⋅Δi\+min⁡\(Wi\)W\_\{q\}=\\text\{round\}\\left\(\\frac\{W\_\{i\}\-\\min\(W\_\{i\}\)\}\{\\Delta\_\{i\}\}\\right\)\\cdot\\Delta\_\{i\}\+\\min\(W\_\{i\}\)\(7\)
whereΔi=max⁡\(Wi\)−min⁡\(Wi\)2b−1\\Delta\_\{i\}=\\dfrac\{\\max\(W\_\{i\}\)\-\\min\(W\_\{i\}\)\}\{2^\{b\}\-1\}is the quantization step of rowii\. Since the weights of the Embedding matrix are static, the parametersΔi\\Delta\_\{i\}andmin⁡\(W\)\\min\(W\)are computed once and fixed for the entire inference session\.

To evaluate the degradation of the output layer at each step of autoregressive generation \(8\), two logit vectors are computed – the reference and the distorted one:

lc​l​e​a​n=WF​16⋅ht,ld​i​r​t​y=Wq⋅htl\_\{clean\}=W\_\{F16\}\\cdot h\_\{t\},\\qquad l\_\{dirty\}=W\_\{q\}\\cdot h\_\{t\}\(8\)
whereWqW\_\{q\}is the weight matrix after applying SA\-PTQ, andht∈ℝdh\_\{t\}\\in\\mathbb\{R\}^\{d\}is the hidden vector of the last token at steptt\. Feedinglc​l​e​a​nl\_\{clean\}andld​i​r​t​yl\_\{dirty\}into the SQNR formula then gives a quantitative estimate of how much the entire logit distribution is distorted by quantizing the output\-layer weights\.

### 2\.2SQNR as a Universal Degradation Indicator

Within the proposed framework of activation\-quantization metrics and SA\-PTQ, we decided to adopt the SQNR coefficient\[[2](https://arxiv.org/html/2608.26926#bib.bib2)\], which regards the compression process as the addition of noise to the information transmission channel\. The main advantage of using SQNR is that the metric has a logarithmic scale and is based on energy considerations\. Decibels allow to estimate distortions in both the large FFN weights and the small Embedding vectors with equal precision, while also providing a simple way to evaluate gradient of the compression ratio when reducing the bit\-width \(eg, from 16 to 8 bits\)\. At the same time, SQNR is cheap to calculate for each layer, and its correlation with the final quality of generation is high, which makes it suitable for use within heterogeneous sLLMs, since these two properties are critical when selecting a metric\.

The need for such a metric arises precisely from the fact that we are comparing two types of layers: FFNs and Embeddings\. To do this, one needs a dimensionless coefficient that would indicate to what extent the information contained in these layers is distorted\.

To do this, we use the SQNR metric \(signal\-to\-quantization\-noise ratio\), which is conventionally measured in decibels\. In this case, if the value of SQNR is high \(eg, 40–50 dB\), then this means that the layer is "not noticeable" after quantization, or, in other words, is not distorted\. On the other hand, a value of SQNR below 20 dB indicates a significant loss of information during quantization\. This property is very convenient for comparing different types of layers, since SQNR is a universal metric that quantifies the amount of information loss during compression, regardless of whether it is stored in vectors or tensors\. The value of SQNR makes sense for any data, since it estimates the percentage of useful information in the compressed data, regardless of the initial size \(whether it is millions of tokens or a hundredth of a token\)\.

Thus, we use the SQNR \(9\) metric to estimate the informativeness of each layer\. It calculates the ratio between the energy of information in a clean signal and the energy of information in a noisy signal\. This metric helps to find the most informative layers of the neural network, which differ depending on the type of layer, within the limits of one common metric\.

S​Q​N​RM​L​P\(n\)=10⋅log10⁡\(‖yc​l​e​a​n\(n\)‖22‖yc​l​e​a​n\(n\)−yd​i​r​t​y\(n\)‖22\)SQNR\_\{MLP\}^\{\(n\)\}=10\\cdot\\log\_\{10\}\\left\(\\frac\{\\left\\\|y\_\{clean\}^\{\(n\)\}\\right\\\|\_\{2\}^\{2\}\}\{\\left\\\|y\_\{clean\}^\{\(n\)\}\-y\_\{dirty\}^\{\(n\)\}\\right\\\|\_\{2\}^\{2\}\}\\right\)\(9\)
The SQNR calculation for FFN measures the impact of weight quantization on the transformation of the input vector, comparing the reference valueyc​l​e​a​n=x⋅Wy\_\{clean\}=x\\cdot Wwith the distorted resultyd​i​r​t​y=x⋅SA\-PTQ​\(W,b\)y\_\{dirty\}=x\\cdot\\text\{SA\-PTQ\}\(W,b\)to assess the layer’s robustness to compression, whereas the KV\-cache analysis determines the direct degradation of vectors when written to memory, by comparing the original vectoryc​l​e​a​n=vy\_\{clean\}=vwith its quantized versionyd​i​r​t​y=SA\-PTQ​\(v,b\)y\_\{dirty\}=\\text\{SA\-PTQ\}\(v,b\), in order to confirm context preservation and the uniformity of the data distribution\.

For a more granular view, we computeS​Q​N​R\(n\)SQNR^\{\(n\)\}separately for eachnn\-th transformer block, which pinpoints exactly which layers are most sensitive to precision loss\.

## 3Testing: Embedding Data

Two experiments were conducted \(Figs\.[2](https://arxiv.org/html/2608.26926#S3.F2),[3](https://arxiv.org/html/2608.26926#S3.F3)\)\. The first \(top\-5 tokens for a single prompt\) allows a specific moment – a single token – to be observed, giving a detailed picture of which tokens compete and how their ranking changes after quantization\. The second \(difference \+ SQNR over 10 tokens\) allows the dynamics of generation to be observed, averaging over several steps – a more statistically robust difference, showing how quantization affects the model’s confidence during generation\.

As part of an initial analysis of the degradation of the output\-layer weight matrix, we perform an ablation experiment comparing model predictions at different degrees of quantization against the reference precision F16 \(Fig\.[2](https://arxiv.org/html/2608.26926#S3.F2)\)\. The “Top\-1 match with F16” plot shows the ratio of test prompts where the winning token matched the reference model\. The metric is the clearest indication of prediction degradation, as its value drops when further weight matrix reductions distort the model’s final prediction\.

The weight\-matrix MAE is the mean absolute error between the elements of the weight matrix of the reduced precision and the reference F16 matrix\. This metric characterizes the deviation of the distribution of the weights at the matrix level before they are projected to logits, enabling a direct comparison between different schemes’ impact\. The deviation of the value of the winning token for each prompt at each degree of reduction serves as another diagnostic metric\. By plotting these values, we see “Top\-1 token logit by prompt”, the logit of the winning token for each prompt, reduced to a certain precision, versus the F16 reference\. The lower value of the logit is an indication of reduced potency of this particular token in the output distribution after reduction, which could affect its rank\. Similarly, delta logit of the winning token is a useful metric for inspecting individual prompts\. It is calculated as the value of the logit of the winning token for the reference minus the value of the logit of the winning token for the reduced precision model, for each prompt\. A positive value signifies a reduction in the winning token’s potency after reduction compared to the reference, which can be a sign of prediction degradation\. We analyze this value for each prompt individually, which allows us to rank prompts by their sensitivity to the reduction\.

![Refer to caption](https://arxiv.org/html/2608.26926v1/figures/image2.png)Figure 2:Token ranking for a single prompt\.To visually assess the degradation of the output block at different quantization levels, three plots were constructed \(Fig\.[3](https://arxiv.org/html/2608.26926#S3.F3)\), each reflecting a separate aspect of compression’s effect on prediction quality\.

The first plot shows the average gap between the top\-1 and top\-2 token logits, averaged over several autoregressive generation steps\. This metric characterizes the model’s confidence in choosing the next token: the larger the gap, the more robust the prediction is to external perturbations, including those introduced by quantization\. The plot allows visual assessment of how well different bit widths preserve this confidence relative to the reference F16 model\.

The second plot shows the error vectorΔg​a​p\\Delta\_\{gap\}\(10\), the difference of the average gap between the reference and the quantized model for each test prompt\. A positive value indicates a reduction in the gap between the winning token’s logits and its closest competitor, which indicates degradation of the model’s confidence in choosing the next token\. A negative value means an increase in this difference due to the stochastic nature of quantization error and cannot be interpreted as an improvement in prediction quality\. The plot makes it possible to identify prompts most vulnerable to compression and to assess the stability of degradation across different input contexts\.

Δg​a​p=g​a​pF​16−g​a​pQ\\Delta\_\{gap\}=gap\_\{F16\}\-gap\_\{Q\}\(10\)
The third plot shows the SQNR of the logits for each prompt\. Unlike the gap metric, which focuses exclusively on the two extreme values of the distribution, SQNR evaluates the distortion of the entire logit vector\. Jointly examining SQNR andΔg​a​p\\Delta\_\{gap\}makes it possible to distinguish two degradation scenarios: uniform distortion of the whole distribution versus a local shift in the region of candidate tokens that directly affects the final choice \(Fig\.[3](https://arxiv.org/html/2608.26926#S3.F3)\)\.

![Refer to caption](https://arxiv.org/html/2608.26926v1/figures/image3.png)Figure 3:Difference and SQNR over 10 tokens\.For the subsequent use of SQNR in the generalized metric, in order to obtain a quantization\-efficiency score, SQNR is normalized according to formula \(11\)\.

S​Q​N​Rn​o​r​m=1−log10⁡\(2\)log10⁡\(S​Q​N​RQ​8\)SQNR\_\{norm\}=1\-\\frac\{\\log\_\{10\}\(2\)\}\{\\log\_\{10\}\(SQNR\_\{Q8\}\)\}\(11\)
The mean SQNR value for Q8\-ch was 45\.61 dB, giving a normalized score of 0\.819\. The value for Q16 \(93\.82 dB\) serves as the reference and confirms that at full precision there is no distortion\. The normalized valueQe​m​b=0\.819Q\_\{emb\}=\\mathbf\{0\.819\}reflects the information preservation of the output layer after quantization and is subsequently combined with the speed component into the unified metric\.

The same normalization formula can also be applied not only to the SQNR value averaged over the entire block but also to each layer independently\. Since SA\-PTQ probing is performed for each layer separately, the natural output is not a single number but a vector\{Q\(1\),Q\(2\),…,Q\(L\)\}\\\{Q^\{\(1\)\},Q^\{\(2\)\},\\ldots,Q^\{\(L\)\}\\\}, where each element characterizes the information preservation of a specific layer\. This opens up the possibility of layer\-wise ranking, in which layers with highQ\(n\)Q^\{\(n\)\}can tolerate aggressive compression, while layers with lowQ\(n\)Q^\{\(n\)\}require preserving higher precision\. Thus, the proposed metric scales from a block\-level estimate to a full layer\-wise model profile without changing the formula itself, simply by adjusting the granularity of the input data\.

## 4Testing: FFN Data

Several tests were carried out, first, to obtain data for each FFN block \(gate, up, and down\) for analysis\. Testing predicted the quality loss of the FFN at 8\-4 bits, obtaining SQNR for each of its components and layers\.

Two experiments were carried out \(Figs\.[4](https://arxiv.org/html/2608.26926#S4.F4),[7](https://arxiv.org/html/2608.26926#S4.F7)\-[11](https://arxiv.org/html/2608.26926#S4.F11)\): one is a static weight analysis, yielding SQNR by projection and layer, SQNR by prompt, and weight MAE, which characterizes the general degradation picture for the given quantization scheme; the others visualize the dynamics of the autoregressive generation process by means of a token×\\timeslayer heatmap, which makes it possible to see how the quantization affects each step of the generation in all layers of the model\.

As part of the analysis of the degradation of the FFN block, a comparison of SQNR values was made for different levels of quantization relative to the baseline F16 \(Fig\.[4](https://arxiv.org/html/2608.26926#S4.F4)\)\. The “SQNR comparison by layer \- all configs and projections” plot characterizes the dynamics of SQNR for all three projections \(Wg​a​t​e,Wu​p,Wd​o​w​nW\_\{gate\},W\_\{up\},W\_\{down\}\) for each layer of the FFN block for all four configurations\. This allows one to estimate roughly the vulnerability of each projection to quantization\. The drop in SQNR values in the middle layers with a tendency to rise at the output layer is characteristic of all configurations\. For the Q8\-ch configuration, theWd​o​w​nW\_\{down\}projection demonstrated significantly lower SQNR values compared toWg​a​t​eW\_\{gate\}andWu​pW\_\{up\}, while for Q4\-ch, all three projections showed SQNR values below 20 dB, which indicates the destructive effect of quantization with 4 bits per channel on the FFN block\.

![Refer to caption](https://arxiv.org/html/2608.26926v1/figures/image4.png)Figure 4:Predicted degradation across bit widths for each component of Gemma 3 1B\-it, relative to the mean\.The plot in Fig\.[5](https://arxiv.org/html/2608.26926#S4.F5)shows the mean SQNR value of the block’s final output, averaged over all 26 layers of the model, for each test prompt\. The results demonstrate high homogeneity of degradation: SQNR values are practically independent of the semantic content of the input context and are determined solely by the quantization scheme\. Q8\-ch shows a mean value of 29\.0 dB – a zone of noticeable but non\-destructive degradation\. Q4\-ch, Q8\-tensor, and Q4\-tensor drop below 20 dB, with Q4\-tensor showing values around 0 dB, corresponding to complete signal destruction\.

![Refer to caption](https://arxiv.org/html/2608.26926v1/figures/image5.png)Figure 5:FFN output SQNR by prompt\.The plot in Fig\.[6](https://arxiv.org/html/2608.26926#S4.F6)characterizes the degree of weight distortion at the level of the matrices themselves, prior to projection into activation space\. Q8\-ch shows the smallest error, whereas Q4\-tensor gives an order\-of\-magnitude larger deviation\. Notably, theWd​o​w​nW\_\{down\}projection consistently shows lower MAE relative toWg​a​t​eW\_\{gate\}andWu​pW\_\{up\}across all configurations, which is explained by differences in matrix dimensionality and weight distribution\.

![Refer to caption](https://arxiv.org/html/2608.26926v1/figures/image6.png)Figure 6:FFN weight MAE by configuration and projection\.To visually assess the dynamics of degradation during autoregressive generation, SQNR heatmaps were constructed in the token×\\timeslayer space \(Figs\.[7](https://arxiv.org/html/2608.26926#S4.F7)–[11](https://arxiv.org/html/2608.26926#S4.F11)\)\. Each row corresponds to one generation step, each column to a layer number; color reflects the mean SQNR value across the three FFN projections\.

![Refer to caption](https://arxiv.org/html/2608.26926v1/figures/image7.png)Figure 7:Per\-layer/per\-token prediction heatmap, geography\-knowledge prompt\.![Refer to caption](https://arxiv.org/html/2608.26926v1/figures/image8.png)Figure 8:Per\-layer/per\-token prediction heatmap, function\-knowledge prompt\.![Refer to caption](https://arxiv.org/html/2608.26926v1/figures/image9.png)Figure 9:Per\-layer/per\-token prediction heatmap, historical\-knowledge prompt\.![Refer to caption](https://arxiv.org/html/2608.26926v1/figures/image10.png)Figure 10:Per\-layer/per\-token prediction heatmap, composition prompt\.![Refer to caption](https://arxiv.org/html/2608.26926v1/figures/image11.png)Figure 11:Per\-layer/per\-token prediction heatmap, code\-reading prompt\.Across all five prompts \(Figs\.[7](https://arxiv.org/html/2608.26926#S4.F7)–[11](https://arxiv.org/html/2608.26926#S4.F11)\), a consistent pattern is observed: Q8\-ch forms a gradient transition from yellow\-green tones in the early layers to warmer tones in the middle layers, with partial recovery toward the last layer, which correlates with the layer\-wise dynamics observed in the SQNR\-by\-layer plot\. Q4\-ch, in contrast, colors nearly the entire space red, showing SQNR values below 10 dB as early as the second or third layer\. The dominant degradation factor is the layer index, and this pattern is reproduced across all prompts\. At the same time, context type introduces local deviations: in technical and structured prompts \(“In Python, to open a file you use”, “The theory of relativity was developed by”\), individual generation tokens show anomalous SQNR values not observed in narrative contexts\. This indicates that FFN block degradation is predominantly structural in nature, although the semantics of the input context can locally amplify its manifestations\.

To quantitatively summarize the observed degradation picture and for subsequent use, the normalized SQNR score of the FFN block is computed using the same scheme as for the output layer \(12\)\.

S​Q​N​Rn​o​r​m=1−log10⁡\(2\)log10⁡\(S​Q​N​RQ​8\)SQNR\_\{norm\}=1\-\\frac\{\\log\_\{10\}\(2\)\}\{\\log\_\{10\}\(SQNR\_\{Q8\}\)\}\(12\)
The mean SQNR value at Q8\-ch for theWd​o​w​nW\_\{down\}projection was 29\.70 dB – noticeably lower than the corresponding value for the embedding matrix \(45\.61 dB\), which is consistent with the visual picture from the heatmaps: the FFN block as a whole is more sensitive \(since the gate, up, and down blocks are all quantized\) to quantization than the output layer\. The reference value Q16 \(77\.95 dB\) confirms that at the original precision degradation is minimal\. The normalized scoreQf​f​n=0\.796Q\_\{ffn\}=\\mathbf\{0\.796\}reflects the higher sensitivity of the FFN to compression losses, subsequently allowing a score to be obtained based on the ratio between information loss and inference speedup\.

## 5Formalizing the Latency Parameters of FFN and Embedding \(Speed Component\)

To construct a general metric of quality\-vs\-speed dependence, the speed component must be taken into account\. Since SQNR measures the ratio of signal energy to quantization\-noise energy – i\.e\., the degree of distortion of the output vector after weight quantization – it is a purely informational metric that contains no information about speed\. Speed is determined differently: by the number of bytes the GPU must read from memory per generation step\.

To combine speed and quality in a single metric, both indicators need to be expressed on a logarithmic \(decibel\) scale\. However, direct comparison is impossible because of the difference in scale: the speed score and SQNR lie in fundamentally different value ranges\. Normalization via the logarithm – which represents the theoretical maximum speedup for Q8 and is derived from the ratio of F16 and Q8 byte sizes – brings both metrics to a common range\. After normalization, both indicators reflect a fraction of their physical limit and can be combined into a single quantization\-efficiency metric\.

### 5\.1Embedding

To answer the question of how quickly the GPU can read the weight matrix from memory, only the roofline model can help – it describes the performance of an operation through the ratio of computation to data volume\. For memory\-bound operations, this is sufficient, while other, more complex profiling methods or empirical measurements are redundant\.

Roofline, as an analytical tool, predicts the execution time of an operation based on two physical constraints: memory bandwidth and GPU compute power\. Since multiplying the embedding matrix by the hidden\-state vector is memory\-bound, execution time comes down entirely to memory bandwidth – which means token time for any weight bit width can be computed analytically, from architectural parameters and hardware specs alone, without running the model at all\. Applying this principle to the transformer architecture, the execution time of any layer is expressed through the size of its weights and the GPU’s memory bandwidth \(13\)\.

tl​a​y​e​r=Wb​y​t​e​sB​Wt\_\{layer\}=\\frac\{W\_\{bytes\}\}\{BW\}\(13\)
whereWb​y​t​e​sW\_\{bytes\}is the weight size in bytes andB​WBWis the GPU’s memory bandwidth \(320 GB/s for the T4\)\. The formula holds for memory\-bound operations: the arithmetic intensityA​I=1AI=1FLOP/byte is significantly below the ridge point of 203 FLOP/byte, hence execution time is determined exclusively by the speed of reading weights from memory\. Applying this dependency to the lm\_head matrix gives its execution time for F16 and Q8 separately \(14\)\.

tl​mF​16=V⋅H⋅2B​W,tl​mQ​8=V⋅H⋅1B​Wt\_\{lm\}^\{F16\}=\\frac\{V\\cdot H\\cdot 2\}\{BW\},\\qquad t\_\{lm\}^\{Q8\}=\\frac\{V\\cdot H\\cdot 1\}\{BW\}\(14\)
whereVVis the vocabulary size andHHis the hidden\-layer dimension\. Under Q8 quantization, each matrix element occupies 1 byte instead of 2, which halves the volume of data read and proportionally speeds up the operation\. Knowing the lm\_head time for both bit widths, the total time to generate one token is the sum of the time for all transformer blocks and the time for the output layer \(15\)\.

tt​o​kF​16=L⋅tl​a​y​e​r\+tl​mF​16,tt​o​kQ​8=L⋅tl​a​y​e​r\+tl​mQ​8t\_\{tok\}^\{F16\}=L\\cdot t\_\{layer\}\+t\_\{lm\}^\{F16\},\\qquad t\_\{tok\}^\{Q8\}=L\\cdot t\_\{layer\}\+t\_\{lm\}^\{Q8\}\(15\)
whereLLis the number of transformer blocks andtl​a​y​e​rt\_\{layer\}is the total time of all operations of one block except lm\_head\. When only the embedding matrix is quantized, block time remains unchanged and only the lm\_head time changes\. The ratio of times for the two configurations gives the speedup coefficient \(16\)\.

ηe​m​b=tt​o​kF​16tt​o​kQ​8=L⋅tl​a​y​e​r\+2​V​HB​WL⋅tl​a​y​e​r\+V​HB​W\\eta\_\{emb\}=\\frac\{t\_\{tok\}^\{F16\}\}\{t\_\{tok\}^\{Q8\}\}=\\frac\{L\\cdot t\_\{layer\}\+\\frac\{2VH\}\{BW\}\}\{L\\cdot t\_\{layer\}\+\\frac\{VH\}\{BW\}\}\(16\)
whereη\\etais the speedup coefficient; the memory bandwidthB​WBWcancels out, and the speedup is expressed purely through the model’s architectural parameters, independent of specific hardware characteristics\. When tested on the Qwen 0\.5B, Qwen 1\.5B, and Gemma 3 1B models, the mean error between predicted and actual results was approximately 4%\.

To verify the proposed prediction method, a series of experiments was conducted on the Gemma 3 1B\-it model using Google Colab \(NVIDIA T4 GPU\) as the runtime environment\. Embedding\-matrix quantization was performed via LLM\-Viewer and llama\.cpp with the Q8\_0 format, while all other model components were kept in F16\. The actual speedup was measured as the ratio of the decoding speed of the quantized configuration to that of the baseline F16 model\.

The results are presented in Table[2](https://arxiv.org/html/2608.26926#S5.T2)\. The analytically predicted token time for Q8 was 5\.37 ms versus an actual 5\.51 ms, corresponding to a prediction error of 2\.6%\. The predicted speedup of×\\times1\.178 matched the actual×\\times1\.148 within the acceptable margin of error\. After normalization vialog10⁡\(2\)\\log\_\{10\}\(2\), the speed score wasSe​m​b=0\.236S\_\{emb\}=\\mathbf\{0\.236\}, reflecting the fraction of the theoretical Q8 speedup potential that was actually realized\.

Table 2:Prediction error relative to actual speedup for Gemma 3 1B\-it \(Embedding\)\.The normalized speed score corresponds to the fraction of the theoretical Q8 speedup potential that was realized\. Together with the quality component, this value forms the basis for building the unified metric\.

### 5\.2FFN

For the FFN, an anchor F16 measurement is used – a single real F16 measurement via llama\-bench, which serves as the reference point for the prediction\. The reasoning is that the roofline model theoretically predicts the difference between F16 and Q8 \(ΔF​F​N\\Delta\_\{FFN\}\), but not the absolute token time – there is GPU load, the scheduler, CUDA graphs, and other factors that roofline does not account for\. Since the FFN block is a heavy operation, its execution time is fully determined by the volume of weights read and by memory bandwidth \(17, 18\)\.

tF​F​NF​16=WF​F​NB​W,tF​F​NQ​8=WF​F​N/2B​Wt\_\{FFN\}^\{F16\}=\\frac\{W\_\{FFN\}\}\{BW\},\\qquad t\_\{FFN\}^\{Q8\}=\\frac\{W\_\{FFN\}/2\}\{BW\}\(17\)
whereWF​F​N=3⋅H⋅I⋅2W\_\{FFN\}=3\\cdot H\\cdot I\\cdot 2bytes \(gate, up, down projections\)\. The coefficient 3 reflects the three matrices of the SwiGLU block, and the multiplier 2 is the size of one element in F16 format, in bytes\. When moving to Q8, each element occupies 1 byte, so the volume of data read is halved and operation time drops proportionally\. Knowing the times for both bit widths, we can compute the theoretical saving for a single block and extend it to allLLtransformer blocks of the model \(19\)\.

ΔF​F​N=\(tF​F​NF​16−tF​F​NQ​8\)⋅L\\Delta\_\{FFN\}=\\left\(t\_\{FFN\}^\{F16\}\-t\_\{FFN\}^\{Q8\}\\right\)\\cdot L\(18\)
The remaining components of each block – the attention mechanism, normalizations, and residual connections – remain in F16 and are not included in the saving computation\. To move from theoretical savings to an absolute prediction, a real anchor point is needed\. It is obtained as the reciprocal of the measured generation speed of the F16 model \(20\)\.

Tr​e​a​lF​16=1tok/sF​16T\_\{real\}^\{F16\}=\\frac\{1\}\{\\text\{tok/s\}\_\{F16\}\}\(19\)
This measurement absorbs all the system overheads that roofline does not model\. Given the real token time and the theoretically predicted saving, the predicted Q8 token time is obtained by subtraction \(21\):

Tp​r​e​dQ​8=Tr​e​a​lF​16−ΔF​F​NT\_\{pred\}^\{Q8\}=T\_\{real\}^\{F16\}\-\\Delta\_\{FFN\}\(20\)
This approach preserves the accuracy of the absolute measurement while not requiring the model to be re\-run for each configuration\. The ratio of the original time to the predicted time gives a dimensionless speedup coefficient \(22\):

ηf​f​n=Tr​e​a​lF​16Tp​r​e​dQ​8\\eta\_\{ffn\}=\\frac\{T\_\{real\}^\{F16\}\}\{T\_\{pred\}^\{Q8\}\}\(21\)
To evaluate the proposed prediction method, experiments were carried out on three models: Gemma 3 1B\-it, LLaMA 3\.2 1B, and Qwen 2\.5 1\.5B, using Google Colab \(NVIDIA T4 GPU\) as the runtime environment\. For each model, the savingΔF​F​N\\Delta\_\{FFN\}was calculated analytically from the LLM\-Viewer data, and the predicted speedup was compared to the actual value obtained by benchmarking with llama\-bench, with only the FFN blocks being quantized to int8 and the rest of the model left in float16\.

The architectural parameters and roofline\-computation results for each of the three models are given in Table[3](https://arxiv.org/html/2608.26926#S5.T3)\. For Gemma 3 1B\-it, the analytical saving is 1\.94 ms, which would reduce the 12\.53 ms per\-token time to 10\.59 ms\. On the other hand, while LLaMA 3\.2 1B has fewer blocks \(16 vs\. 26\), its per\-block saving is larger due to the higher FFN matrix width, 157\.5μ\\mus vs\. 74\.7μ\\mus for Gemma 3 1B\-it, for a total of 2\.52 ms\. Finally, with 28 blocks and the largest per\-block weight volume, Qwen 2\.5 1\.5B has the absolute maximum saving of 3\.61 ms, reducing the 17\.53 ms per\-token time to 13\.92 ms\.

Table 3:Architectural parameters and roofline results at test time for three models\.The comparison between predicted and actual speedup coefficients is given in Table[4](https://arxiv.org/html/2608.26926#S5.T4)\. The best match was achieved for Gemma 3 1B\-it, with an error of 0\.7%; for LLaMA 3\.2 1B the error was 3\.3%\. The largest deviation was observed for Qwen 2\.5 1\.5B, at 6\.3%, which may be explained by FFN architectural specifics and CUDA\-scheduler behavior for this model\. The mean error across the three models was 3\.4%, confirming the applicability of the hybrid approach for predicting speedup without re\-running the model\. After normalization, the FFN speed score was 0\.261\.

Table 4:Prediction error relative to actual speedup \(FFN\)\.For subsequent use, the speedup coefficient is converted into a normalized scale vialog10⁡\(2\)\\log\_\{10\}\(2\), where a value of one corresponds to the theoretical maximum speedup at a two\-fold reduction in weight volume\. For Gemma 3 1B\-it, the predicted speedup of×\\times1\.199 gives a normalized FFN speed score ofSf​f​n=0\.261S\_\{ffn\}=\\mathbf\{0\.261\}\.

The presented approach admits a natural generalization to the case of layer\-wise heterogeneous quantization, where each of theLLlayers receives its own bit widthbnb\_\{n\}\. Since all Gemma 3 layers are architecturally identical, the execution time of one layer at bit widthbnb\_\{n\}is computed analytically via roofline \(23\)\.

tF​F​N,nbn=WF​F​N⋅16bnB​Wt\_\{FFN,n\}^\{b\_\{n\}\}=\\frac\{W\_\{FFN\}\\cdot\\frac\{16\}\{b\_\{n\}\}\}\{BW\}\(22\)
where the factor16bn\\frac\{16\}\{b\_\{n\}\}reflects the reduction in the volume of data read relative to F16\. The total saving over all layers is defined as \(24\)

ΔF​F​N=∑n=1L\(tF​F​N,nF​16−tF​F​N,nbn\)\\Delta\_\{FFN\}=\\sum\_\{n=1\}^\{L\}\\left\(t\_\{FFN,n\}^\{F16\}\-t\_\{FFN,n\}^\{b\_\{n\}\}\\right\)\(23\)
The predicted token time is computed by the same scheme as before\. Since the architectural parameters of all 26 layers of Gemma 3 are known, the theoretical calculation for any bit\-width configuration can likewise be performed analytically without additional measurements\.

Considering a configuration in which the middle 10 layers – the most sensitive according to the SQNR heatmap data – are kept at Q8, while the first 8 and last 8 layers, being more robust, are quantized to Q4 \(25, 26\):

ΔF​F​N=8⋅\(tF​F​NF​16−tF​F​NQ​4\)\+10⋅\(tF​F​NF​16−tF​F​NQ​8\)\+8⋅\(tF​F​NF​16−tF​F​NQ​4\)\\Delta\_\{FFN\}=8\\cdot\\left\(t\_\{FFN\}^\{F16\}\-t\_\{FFN\}^\{Q4\}\\right\)\+10\\cdot\\left\(t\_\{FFN\}^\{F16\}\-t\_\{FFN\}^\{Q8\}\\right\)\+8\\cdot\\left\(t\_\{FFN\}^\{F16\}\-t\_\{FFN\}^\{Q4\}\\right\)\(24\)
At a baseline token time oftF​F​NF​16=149\.4​μ​st\_\{FFN\}^\{F16\}=149\.4\\,\\mu s,tF​F​NQ​8=74\.7​μ​st\_\{FFN\}^\{Q8\}=74\.7\\,\\mu s,tF​F​NQ​4≈37\.4​μ​st\_\{FFN\}^\{Q4\}\\approx 37\.4\\,\\mu s:

ΔF​F​N=16⋅112\.0\+10⋅74\.7=1792\.0\+747\.0=2539\.0​μ​s≈2\.54​ms\\Delta\_\{FFN\}=16\\cdot 112\.0\+10\\cdot 74\.7=1792\.0\+747\.0=2539\.0\\,\\mu s\\approx 2\.54\\text\{ ms\}\(25\)
At a baseline token time ofTr​e​a​lF​16=12\.53T\_\{real\}^\{F16\}=12\.53ms, the predicted time is12\.53−2\.54=9\.9912\.53\-2\.54=9\.99ms, corresponding to a speedup ofη≈1\.254\\eta\\approx 1\.254versusη=1\.183\\eta=1\.183for uniform Q8\. Layer\-wise heterogeneity theoretically gives an additional speedup of about 6% relative to block\-level Q8, with the sensitive middle layers retaining Q8 precision while the robust outer layers are more aggressively quantized\. However, it is possible that under layer\-wise heterogeneous quantization the overhead of dequantizing weights when loading them into compute cores may differ for Q4 and Q8, potentially introducing an additional prediction error\. Practical verification of this result requires tools with fine\-grained control over the computation graph, such as ExLlamaV3 or TensorRT\-LLM, which were not available for the Gemma 3 architecture within the scope of this study\.

## 6The Unified Metric

The experiments conducted for measuring quality degradation and predicting inference speed yielded four normalized values:Qe​m​b=0\.819Q\_\{emb\}=0\.819,Qf​f​n=0\.796Q\_\{ffn\}=0\.796,Se​m​b=0\.236S\_\{emb\}=0\.236,Sf​f​n=0\.261S\_\{ffn\}=0\.261\. Each of them independently characterizes one aspect of quantizing a specific block and is brought to a common scale from zero to one\. The task of the final metric is to combine these four numbers into a single estimate of a quantization configuration, while preserving the ability to control the balance between speed and quality depending on the requirements of a specific task\. For this, a weighted average of the speed and quality components is introduced, with a priority parameterα\\alpha\(27\)\.

score\(α\)=\(1−α\)⋅1N∑kSk\+α⋅1N∑kQk\\text\{score\}\(\\alpha\)=\(1\-\\alpha\)\\cdot\\frac\{1\}\{N\}\\sum\_\{k\}S\_\{k\}\+\\alpha\\cdot\\frac\{1\}\{N\}\\sum\_\{k\}Q\_\{k\}\(26\)
where the summation runs over all evaluated blockskk\. In the block\-level variant,k∈\{e​m​b,f​f​n\}k\\in\\\{emb,ffn\\\}andN=2N=2; in the layer\-wise variant,k∈\{1,2,…,L\}k\\in\\\{1,2,\\ldots,L\\\}andN=LN=L\. HereSkS\_\{k\}is the normalized speed score of blockkk, reflecting the fraction of the theoretical speedup potential realized at a two\-fold reduction in weight volume, andQkQ\_\{k\}is the normalized SQNR score of blockkk, characterizing the information preservation of the output representations after quantization\. The parameterα∈\[0,1\]\\alpha\\in\[0,1\]sets the priority between speed and quality: atα=0\\alpha=0the metric reflects purely the speed potential, atα=1\\alpha=1purely the information preservation, and atα=0\.5\\alpha=0\.5both aspects contribute equally\. All components are normalized on a unified scale, which allows quantization configurations to be directly compared both at the block level and at the level of individual layers, without reference to absolute latency values or decibels\.

Atα=0\.5\\alpha=0\.5, for the Q8 configuration on the Gemma 3 1B\-it model, the final score wass​c​o​r​e=0\.497score=0\.497, where the speed component equals 0\.187 and the quality component equals 0\.808\.

## 7Conclusion

This paper aims to produce a metric that quantifies the efficiency of the sLLM quantization, combining two criteria usually studied in isolation, namely, the information preserved in a given block, captured by the normalized SQNR, and the gain in the quantized model’s performance, captured by an analytical roofline latency model\. Unlike prior work, where a significant computational budget was spent on either quantization\-aware training or lengthy evolutionary searches over design space, the proposed metric is purely analytical, requiring no training for any of the candidate configurations\.

Profiling revealed that FFN blocks and embedding contribute the most to the overall latency, with attention being surprisingly less costly, informing the choice of target blocks to optimize over\. These blocks were scored according to the metrics, achieving a value ofQe​m​b=0\.819Q\_\{emb\}=0\.819

andQf​f​n=0\.796Q\_\{ffn\}=0\.796for their quality retention score, and a value ofSe​m​b=0\.236S\_\{emb\}=0\.236andSf​f​n=0\.261S\_\{ffn\}=0\.261for their performance gain score, respectively\. A sanity check on the performance gain scores across three model architectures resulted in an average error of approximately 3\.4%, demonstrating the metric’s practical applicability\.

Together with theα=0\.5\\alpha=0\.5, this allows us to score a particular configuration, achieving a value of 0\.497 for the Q8 configuration for Gemma 3 1B\-it, balancing the information preservation and performance efficiency\. As demonstrated by theα\\alpha, this metric can be adjusted to reflect the priority of either quality or performance, depending on the specific use case and system constraints\. Finally, due to the metric’s independence from the granularity of input blocks, it can be applied to blocks, layers, or projections interchangeably, enabling fine\-grained optimization at various levels\. This property is particularly useful, as it allows one to estimate the quality drop and performance gain for a particular configuration ahead of time\.

## References

- \[1\]Hasan J\. Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques // arXiv preprint\. 2411\.06084 – 2024\. –[https://arxiv\.org/abs/2411\.06084](https://arxiv.org/abs/2411.06084)\.
- \[2\]Castagnetti A\., Pegatoquet A\., Miramond B\. Trainable quantization for Speedy Spiking Neural Networks // Frontiers in Neuroscience\. – 2023\. – Vol\. 17\. – Art\. 1154241\. –[https://doi\.org/10\.3389/fnins\.2023\.1154241](https://doi.org/10.3389/fnins.2023.1154241)\.
- \[3\]Davies M\. Efficient LLM Inference: Bandwidth, Compute, Synchronization, and Capacity are all you need // arXiv preprint\. 2507\.14397 – 2025\. –[https://arxiv\.org/html/2507\.14397v1](https://arxiv.org/html/2507.14397v1)\.
- \[4\]Gou D\., et al\. CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression // arXiv preprint\. 2510\.12721 – 2025\. –[https://arxiv\.org/abs/2510\.12721](https://arxiv.org/abs/2510.12721)\.
- \[5\]Murshed M\.G\.S\., et al\. HAPM – Hardware Aware Pruning Method for CNN hardware accelerators in resource constrained devices // arXiv preprint\. 2408\.14055 – 2024\. –[https://arxiv\.org/abs/2408\.14055](https://arxiv.org/abs/2408.14055)\.
- \[6\]Zhang X\., et al\. IMPQ: Interaction\-Aware Layerwise Mixed Precision Quantization for LLMs // arXiv preprint\. 2509\.15455 – 2025\. –[https://arxiv\.org/html/2509\.15455v1](https://arxiv.org/html/2509.15455v1)\.
- \[7\]Liu J\., et al\. 4\-bit Reliable Quantization // arXiv preprint\. 2501\.13331 – 2025\. –[https://arxiv\.org/html/2501\.13331v1](https://arxiv.org/html/2501.13331v1)\.
- \[8\]Liu Z\., et al\. EfficientQAT: Efficient Quantization\-Aware Training for Large Language Models // Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(ACL\)\. – 2024\. –[https://arxiv\.org/abs/2407\.11062](https://arxiv.org/abs/2407.11062)\.
- \[9\]Frantar E\., Ashkboos S\., Hoefler T\., Alistarh D\. GPTQ: Accurate Post\-Training Quantization for Generative Pre\-trained Transformers // International Conference on Learning Representations \(ICLR\)\. – 2023\. –[https://arxiv\.org/abs/2210\.17323](https://arxiv.org/abs/2210.17323)\.
- \[10\]Lin J\., et al\. AWQ: Activation\-aware Weight Quantization for LLM Compression and Acceleration // Proceedings of Machine Learning and Systems \(MLSys\)\. – 2023\. –[https://arxiv\.org/abs/2306\.00978](https://arxiv.org/abs/2306.00978)\.
- \[11\]Liu Z\., et al\. LLM\-QAT: Data\-Free Quantization Aware Training for Large Language Models // arXiv preprint\. 2305\.17888 – 2023\. –[https://arxiv\.org/abs/2305\.17888](https://arxiv.org/abs/2305.17888)\.
- \[12\]Han S\., Mao H\., Dally W\.J\. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding // International Conference on Learning Representations \(ICLR\)\. – 2016\. –[https://arxiv\.org/abs/1510\.00149](https://arxiv.org/abs/1510.00149)\.
- \[13\]Jacob B\., et al\. Quantization and Training of Neural Networks for Efficient Integer\-Arithmetic\-Only Inference // IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)\. – 2018\. – P\. 2704–2713\. –[https://arxiv\.org/abs/1712\.05877](https://arxiv.org/abs/1712.05877)\.
- \[14\]Gholami A\., et al\. A White Paper on Neural Network Quantization // arXiv preprint\. 2106\.08295 – 2021 \(updated 2024\)\. –[https://arxiv\.org/abs/2106\.08295](https://arxiv.org/abs/2106.08295)\.
- \[15\]Yao Z\., et al\. HAWQV3: Dyadic Neural Network Quantization // Proceedings of the 38th International Conference on Machine Learning \(ICML\)\. – 2021\. – P\. 11875–11886\. –[https://arxiv\.org/abs/2011\.10680](https://arxiv.org/abs/2011.10680)\.
- \[16\]Dao T\., Fu D\.Y\., Ermon S\., Rudra A\., Ré C\. FlashAttention: Fast and Memory\-Efficient Exact Attention with IO\-Awareness // Advances in Neural Information Processing Systems \(NeurIPS\)\. – 2022\. – Vol\. 35\. – P\. 16344–16359\. –[https://arxiv\.org/abs/2205\.14135](https://arxiv.org/abs/2205.14135)\.
- \[17\]Xiao G\., et al\. SmoothQuant: Accurate and Efficient Post\-Training Quantization for Large Language Models // Proceedings of the 40th International Conference on Machine Learning \(ICML\)\. – 2023\. –[https://arxiv\.org/abs/2211\.10438](https://arxiv.org/abs/2211.10438)\.
- \[18\]Shazeer N\. GLU Variants Improve Transformer // arXiv preprint\. 2002\.05202 – 2020\. –[https://arxiv\.org/abs/2002\.05202](https://arxiv.org/abs/2002.05202)\.
- \[19\]Dettmers T\., Lewis M\., Belkada Y\., Zettlemoyer L\. LLM\.int8\(\): 8\-bit Matrix Multiplication for Transformers at Scale // Advances in Neural Information Processing Systems \(NeurIPS\)\. – 2022\. – Vol\. 35\. – P\. 22128–22141\. –[https://arxiv\.org/abs/2208\.07339](https://arxiv.org/abs/2208.07339)\.
- \[20\]Ainslie J\., Lee\-Thorp J\., de Jong M\., Zemlyanskiy Y\., Lebrón F\., Sanghai S\. GQA: Training Generalized Multi\-Query Transformer Models from Multi\-Head Checkpoints // arXiv preprint\. 2305\.13245 – 2023\. –[https://arxiv\.org/abs/2305\.13245](https://arxiv.org/abs/2305.13245)\.
- \[21\]Shapley L\.S\. A Value for n\-person Games // Contributions to the Theory of Games\. – 1953\. – Vol\. 2, No\. 28\. – P\. 307–317\.
- \[22\]Ganesh P\., Chen Y\., Lou X\. et al\. Compressing Large\-Scale Transformer\-Based Models: A Case Study on BERT // Transactions of the Association for Computational Linguistics\. – 2021\. – Vol\. 9\. – P\. 1061–1080\. –[https://doi\.org/10\.1162/tacl\_a\_00413](https://doi.org/10.1162/tacl_a_00413)\.
- \[23\]Chu J\., Leng Y\., Li M\., Shen Y\., Shen X\., Zhang Y\. GEO\-Flag: Detecting and Measuring GEO\-Optimized Web Content // arXiv preprint\. 2608\.16824 – 2026\. –[https://arxiv\.org/abs/2608\.16824](https://arxiv.org/abs/2608.16824)\.

Similar Articles