From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs

arXiv cs.LG Papers

Summary

This paper presents an analytically structured, empirically calibrated methodology for estimating LLM inference energy on NVIDIA H100 GPUs without direct measurement, separating prefill and decoding phases and decomposing energy into compute, parameter-access, KV-cache write, and attention-read components.

arXiv:2607.26571v1 Announce Type: new Abstract: The operational energy consumption of large language model (LLM) inference is becoming an increasingly important component of the environmental footprint of deployed AI systems. However, direct measurement of inference energy often requires hardware telemetry, power instrumentation, or infrastructure-specific monitoring, limiting its applicability in comparative studies, early-stage system design, and sustainability reporting. This report presents an analytically structured, empirically calibrated, GPU-level methodology for estimating LLM inference energy on NVIDIA H100-class accelerators without direct runtime measurement. The proposed estimator combines parameter-scaled transformer FLOP accounting, calibrated memory-traffic factors, and hardware-specific energy coefficients for FP16/BF16 tensor-core computation and high-bandwidth-memory movement. It explicitly separates prompt prefill from autoregressive decoding, enabling energy estimates for input tokens, output tokens, and complete inference requests. The methodology further decomposes total energy into compute, parameter-access, key-value-cache write, and attention-read components, allowing the scaling behavior with model size, context length, and generated-token count to be analyzed. The resulting estimates are not intended to replace physical power measurements; rather, they provide transparent, reproducible, and assumption-explicit approximations suitable for model comparison, green-coding analysis, and design-time evaluation of LLM inference workloads.
Original Article
View Cached Full Text

Cached at: 07/30/26, 09:59 AM

# From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs
Source: [https://arxiv.org/html/2607.26571](https://arxiv.org/html/2607.26571)
\\copyrightclause

Copyright for this paper by its authors\. Use permitted under Creative Commons License Attribution 4\.0 International \(CC BY 4\.0\)\.

\\conference

\[orcid=0000\-0002\-0877\-7063, email=tina\.vartziotis@twt\-gmbh\.de, \]\\cormark\[1\]\\fnmark\[1\]

\[orcid=0000\-0001\-7116\-9338, email=rodopi\.kosteli@nikitec\.gr, \]\\fnmark\[1\]\\fnmark\[1\]

\[orcid=0009\-0006\-1309\-8598, email=elli\.vartziotis@twt\-gmbh\.de, \]

\[orcid=0000\-0002\-0562\-5136, email=gdasoulas@hsph\.harvard\.edu, \] \[ email=michael\.keckeisen@twt\-gmbh\.de, \] \[orcid=0000\-0002\-9421\-8566, email=kskianis@uoi\.gr, \]

\[orcid=0000\-0002\-9421\-8566, email=skots@mit\.edu, \]

\[orcid=0000\-0002\-9421\-8566, email=fdominic@hsph\.harvard\.edu, \]

\\cortext

\[1\]Corresponding author\.\\fntext\[1\]These authors contributed equally\.

Tina VartziotisNational Technical University of Athens, Patission Complex 42, 10682 Athens, GreeceHarvard University, 1350 Massachusetts Avenue, 02138 Cambridge, MA, USATWT GmbH Science & Innovation, IndustriestraSSe 6, 70565 Stuttgart, DERodopi KosteliNIKI Ltd Digital Engineering, 205 National Resistance Street, 45500 Ioannina, GreeceUniversity of Ioannina, Campus, 451 10 Ioannina, GreeceElli Danae VartziotisGeorge DasoulasMichael KeckeisenKonstantinos SkianisSotirios KotsopoulosMassachusetts Institute of Technology, 77 Massachusetts Avenue, 02139 Cambridge, MA, USAFrancesca Dominici

\(2022\)

###### Abstract

The operational energy consumption of large language model \(LLM\) inference is becoming an increasingly important component of the environmental footprint of deployed AI systems\. However, direct measurement of inference energy often requires hardware telemetry, power instrumentation, or infrastructure\-specific monitoring, limiting its applicability in comparative studies, early\-stage system design, and sustainability reporting\. This report presents an analytically structured, empirically calibrated, GPU\-level methodology for estimating LLM inference energy on NVIDIA H100\-class accelerators without direct runtime measurement\. The proposed estimator combines parameter\-scaled transformer FLOP accounting, calibrated memory\-traffic factors, and hardware\-specific energy coefficients for FP16/BF16 tensor\-core computation and high\-bandwidth\-memory movement\. It explicitly separates prompt prefill from autoregressive decoding, enabling energy estimates for input tokens, output tokens, and complete inference requests\. The methodology further decomposes total energy into compute, parameter\-access, key–value\-cache write, and attention\-read components, allowing the scaling behavior with model size, context length, and generated\-token count to be analyzed\. The resulting estimates are not intended to replace physical power measurements; rather, they provide transparent, reproducible, and assumption\-explicit approximations suitable for model comparison, green\-coding analysis, and design\-time evaluation of LLM inference workloads\.

###### keywords:

Large Language Models\\sepGPU Inference\\sepInference Energy Modeling\\sepTransformer Systems\\sepComputing Emissions\\sepTensor\-Core Computing\\sepEnergy\-Efficient AI\\sepGreen AI\\sepSustainable Computing\\sepHigh\-Performance Computing

## 1Introduction

As large language models \(LLMs\) scale to serve millions of users across cloud, edge, and on\-premise deployments, their cumulative inference energy has emerged as a first\-order sustainability concern, driven by the substantial computational and environmental costs of large\-scale AI inference\. Prior work in Green AI has shown that advances in model capability are often accompanied by increases in energy consumption and carbon emissions, motivating energy\-aware approaches to machine learning research and deployment\[strubell2019energy,schwartz2020green\]\.

While early work focused primarily on training costs\[patterson2021carbon\], broader Green AI research highlighted the importance of computational efficiency, energy\-aware reporting, and environmental impact across the ML lifecycle\[schwartz2020green,henderson2020systematic,lacoste2019quantifying,lannelongue2021green\], including recent efforts to extend energy\-aware optimization to model selection\[betello2025one\]\. Attention has since shifted toward inference, which is becoming an increasingly important operational component of deployed LLM systems\. Unlike training, which is performed occasionally, inference workloads run continuously and scale with user demand through autoregressive token generation, accumulating substantial energy costs over a model’s deployment lifetime\[lim2024carbon\]\. Their energy footprint depends strongly on generated\-token count, sequence length, model size, hardware configuration, and batching behavior\[fernandez2025energy,fu2024llmco2,Vartziotis\_LLMasaService\]\.

This makes inference energy a practical concern not only for large cloud providers, but also for applied AI teams, research groups, and small\-to\-medium organizations deploying LLMs on local or on\-premise GPU infrastructure\. For such deployments, accelerator\-side energy does not capture the full service\-level cost, since system and facility overheads, including cooling, networking, and power delivery, also contribute\[uptime2024survey,mlperfpower2024\]\. Nevertheless, GPU\-side energy is often a major controllable component of inference cost in GPU\-dense systems\[nvidia\_h100\_datasheet\], making GPU\-level estimation useful for comparing model and workload choices, assessing green\-coding interventions\[Vartziotis\_GreenCode\], and making deployment\-time decisions when full datacenter\-level instrumentation is not available\.

Existing approaches for estimating or reporting AI\-related energy consumption often rely on hardware telemetry, external power measurement, infrastructure\-specific monitoring, or coarse\-grained carbon\-accounting tools\[lannelongue2021green,lacoste2019quantifying\]\. Measurement\-based studies provide valuable empirical evidence across hardware configurations and task types\[samsi\_words\_2023,luccioni\_power\_2024\], but their results are tied to specific hardware, serving systems, and runtime configurations\. Carbon\-aware inference work such as SPROUT further demonstrates that autoregressive generation can itself be optimized for sustainability\[li2024sprout\]\. Conversely, transformer scaling\-law studies and GPU microarchitectural benchmarks provide the basis for analytical FLOP and hardware\-energy modeling\[kaplan2020scaling,hoffmann2022training,antepara2025benchmark\], but do not by themselves provide request\-level or token\-level inference energy estimates\. This leaves a gap for transparent, reproducible estimators that combine transformer compute modeling, memory\-traffic attribution, and hardware\-specific energy coefficients to estimate request\-level and normalized token\-level GPU inference energy without requiring runtime instrumentation\.

We address this gap by presenting a semi\-analytical methodology for estimating accelerator\-side GPU compute and memory\-movement energy for LLM inference\. The estimator combines parameter\-scaled transformer FLOP modeling, calibrated HBM memory\-traffic estimation, and hardware\-specific energy coefficients\. It separates prompt prefill from autoregressive decode, reports normalized input\-token, output\-token, and request\-level energy estimates, and decomposes total energy into compute, parameter\-access, KV\-cache\-write, and attention\-read components\. Beyond comparative estimation, the framework also highlights actionable mechanisms for reducing GPU inference energy, including smaller models, shorter generations, KV\-cache quantization, prompt compression, and improved batching efficiency\.

## 2Methodology

This section defines the analytically structured, empirically calibrated methodology used to estimate GPU\-level energy consumption for LLM inference\. The estimator is designed for settings in which direct runtime power measurement is unavailable, impractical, or not comparable across systems\. It therefore provides assumption\-explicit energy estimates rather than replacing hardware\-level measurement\. The scope of the estimator is accelerator\-side operational energy\. We model only the energy associated with GPU\-side computation and GPU memory movement during inference\. System\-level and datacenter\-level contributions, including CPU execution, host memory, networking, storage, power\-supply losses, cooling, and power usage effectiveness \(PUE\), are outside the scope of this work and are not included in the reported estimates\. Similarly, this paper does not estimate carbon emissions; all reported quantities are expressed as GPU\-level energy\.

Each inference request is separated into two phases: prompt prefill and autoregressive decode\. During prefill, the model processes the input prompt and constructs the initial key–value \(KV\) cache\. During decode, the model generates output tokens sequentially, with each step attending to the accumulated context through the KV cache\. This phase separation is required because input\-token and output\-token costs have different compute and memory\-access patterns\.

### 2\.1Request\-Level Energy Decomposition

LetTinT\_\{\\mathrm\{in\}\}denote the number of input tokens, i\.e\., the tokenized prompt provided to the model, and letToutT\_\{\\mathrm\{out\}\}denote the number of generated output tokens\. The total GPU energy of one request, denotedEGPUE\_\{\\mathrm\{GPU\}\}, is decomposed into prefill and decode energy:

EGPU=Epre\+Edec\.E\_\{\\mathrm\{GPU\}\}=E\_\{\\mathrm\{pre\}\}\+E\_\{\\mathrm\{dec\}\}\.\(1\)whereEpreE\_\{\\mathrm\{pre\}\}is the energy consumed during the prefill phase andEdecE\_\{\\mathrm\{dec\}\}is the energy consumed during the autoregressive decode phase\. The reported input\-token and output\-token energy values are defined as:

Ein/token=EpreTin,Eout/token=EdecTout\.E\_\{\\mathrm\{in/token\}\}=\\frac\{E\_\{\\mathrm\{pre\}\}\}\{T\_\{\\mathrm\{in\}\}\},\\qquad E\_\{\\mathrm\{out/token\}\}=\\frac\{E\_\{\\mathrm\{dec\}\}\}\{T\_\{\\mathrm\{out\}\}\}\.\(2\)whereEin/tokenE\_\{\\mathrm\{in/token\}\}andEout/tokenE\_\{\\mathrm\{out/token\}\}denote the average energy per input token and per generated output token, respectively\. These quantities are not intrinsic constants of a model\. They depend on prompt length, generated\-token count, inference precision, batching, cache reuse, hardware characteristics, and serving implementation\. At the GPU level, request energy is modeled as the sum of tensor\-core compute energy and high\-bandwidth\-memory \(HBM\) movement energy:

EGPU=Ecompute\+Ememory\.E\_\{\\mathrm\{GPU\}\}=E\_\{\\mathrm\{compute\}\}\+E\_\{\\mathrm\{memory\}\}\.\(3\)whereEcomputeE\_\{\\mathrm\{compute\}\}denotes energy consumed by tensor\-core computation, andEmemoryE\_\{\\mathrm\{memory\}\}denotes energy consumed by memory movement\. This decomposition follows common GPU energy modeling approaches, which treat computation and data movement as the dominant contributors to energy consumption\[antepara2025benchmark\]\. The compute term is proportional to the number of tensor\-core floating\-point operations:

Ecompute=αTC​CTC,E\_\{\\mathrm\{compute\}\}=\\alpha\_\{\\mathrm\{TC\}\}C\_\{\\mathrm\{TC\}\},\(4\)whereCTCC\_\{\\mathrm\{TC\}\}denotes tensor\-core FLOPs andαTC\\alpha\_\{\\mathrm\{TC\}\}is the hardware\-specific energy per tensor\-core FLOP\. The memory term is proportional to the number of bits transferred through HBM:

Ememory=eHBM​QHBM,E\_\{\\mathrm\{memory\}\}=e\_\{\\mathrm\{HBM\}\}Q\_\{\\mathrm\{HBM\}\},\(5\)whereQHBMQ\_\{\\mathrm\{HBM\}\}denotes HBM traffic in bits andeHBMe\_\{\\mathrm\{HBM\}\}is the energy per transferred HBM bit\. Combining these terms gives:

EGPU=αTC​CTC\+eHBM​QHBM\.E\_\{\\mathrm\{GPU\}\}=\\alpha\_\{\\mathrm\{TC\}\}C\_\{\\mathrm\{TC\}\}\+e\_\{\\mathrm\{HBM\}\}Q\_\{\\mathrm\{HBM\}\}\.\(6\)
This formulation separates workload\-dependent quantities\(CTC,QHBM\)\(C\_\{\\mathrm\{TC\}\},Q\_\{\\mathrm\{HBM\}\}\)from hardware\-specific coefficients\(αTC,eHBM\)\(\\alpha\_\{\\mathrm\{TC\}\},e\_\{\\mathrm\{HBM\}\}\)\[antepara2025benchmark\]\.

### 2\.2Compute Model

For dense decoder\-only transformers, we approximate the dominant tensor\-core FLOPs using the standard parameter\-scaled transformer estimate:

Cpre,dense\\displaystyle C\_\{\\mathrm\{pre,dense\}\}=K​N​Tin,\\displaystyle=KNT\_\{\\mathrm\{in\}\},Cdec,dense\\displaystyle C\_\{\\mathrm\{dec,dense\}\}=K​N​Tout,\\displaystyle=KNT\_\{\\mathrm\{out\}\},\(7\)whereNNis the number of model parameters andK=6K=6is the FLOP coefficient per parameter per token\. For long contexts, architecture\-aware attention corrections are added through

Cpre=Cpre,dense\+Cpre,attn,Cdec=Cdec,dense\+Cdec,attn,C\_\{\\mathrm\{pre\}\}=C\_\{\\mathrm\{pre,dense\}\}\+C\_\{\\mathrm\{pre,attn\}\},\\qquad C\_\{\\mathrm\{dec\}\}=C\_\{\\mathrm\{dec,dense\}\}\+C\_\{\\mathrm\{dec,attn\}\},\(8\)so that

CTC=Cpre\+Cdec\.C\_\{\\mathrm\{TC\}\}=C\_\{\\mathrm\{pre\}\}\+C\_\{\\mathrm\{dec\}\}\.\(9\)
The full attention correction terms are provided in Supplementary Section[A](https://arxiv.org/html/2607.26571#A1)\.

### 2\.3Memory\-Movement Model

In addition to tensor\-core computation, LLM inference requires substantial data movement through the GPU memory hierarchy\. We approximate the dominant off\-chip memory traffic as high\-bandwidth\-memory \(HBM\) traffic and decompose it into parameter\-access traffic, KV\-cache write traffic, and attention\-related KV\-cache read traffic:

Bitstotal=Bitsparams\+BitsKV\+Bitsattn′\.\\mathrm\{Bits\}\_\{\\mathrm\{total\}\}=\\mathrm\{Bits\}\_\{\\mathrm\{params\}\}\+\\mathrm\{Bits\}\_\{\\mathrm\{KV\}\}\+\\mathrm\{Bits\}\_\{\\mathrm\{attn\}\}^\{\\prime\}\.\(10\)This decomposition captures the primary sources of memory traffic in autoregressive transformer inference, including parameter access, KV\-cache storage, and attention\-related reads, as observed in modern LLM serving systems\[kwon2023efficient\]\. The corresponding memory energy is modeled as

Ememory=Bitstotal⋅eHBM⋅η​\(N\),E\_\{\\mathrm\{memory\}\}=\\mathrm\{Bits\}\_\{\\mathrm\{total\}\}\\cdot e\_\{\\mathrm\{HBM\}\}\\cdot\\eta\(N\),\(11\)whereeHBMe\_\{\\mathrm\{HBM\}\}is the energy per bit transferred from HBM andη​\(N\)\\eta\(N\)is a calibrated memory\-inefficiency factor\. The individual memory\-traffic terms are defined in Section[3\.4](https://arxiv.org/html/2607.26571#S3.SS4), with full derivations provided in Supplementary Section[B](https://arxiv.org/html/2607.26571#A2)\.

### 2\.4Total Request and Token\-Level Energy

The total GPU energy for a request, denotedErequestE\_\{\\mathrm\{request\}\}, is obtained by combining the tensor\-core compute term and the calibrated HBM memory\-movement term:

Erequest=αTC​\(Cpre\+Cdec\)\+Bitstotal⋅eHBM⋅η​\(N\)\.E\_\{\\mathrm\{request\}\}=\\alpha\_\{\\mathrm\{TC\}\}\\left\(C\_\{\\mathrm\{pre\}\}\+C\_\{\\mathrm\{dec\}\}\\right\)\+\\mathrm\{Bits\}\_\{\\mathrm\{total\}\}\\cdot e\_\{\\mathrm\{HBM\}\}\\cdot\\eta\(N\)\.\(12\)whereCpreC\_\{\\mathrm\{pre\}\}andCdecC\_\{\\mathrm\{dec\}\}denote the total prefill and decode tensor\-core FLOPs, obtained from the dense terms in Section[2\.2](https://arxiv.org/html/2607.26571#S2.SS2)and, when applicable, the attention correction terms in Supplementary Section[A](https://arxiv.org/html/2607.26571#A1)\. The termBitstotal\\mathrm\{Bits\}\_\{\\mathrm\{total\}\}denotes total HBM traffic \(Equation[10](https://arxiv.org/html/2607.26571#S2.E10)\)\. The factorη​\(N\)\\eta\(N\)is a dimensionless inefficiency multiplier applied to the memory term, representing non\-ideal HBM behavior\. The average energy per processed token is:

Eavg/token=ErequestTin\+Tout\.E\_\{\\mathrm\{avg/token\}\}=\\frac\{E\_\{\\mathrm\{request\}\}\}\{T\_\{\\mathrm\{in\}\}\+T\_\{\\mathrm\{out\}\}\}\.\(13\)whereTinT\_\{\\mathrm\{in\}\}andToutT\_\{\\mathrm\{out\}\}are the numbers of input and generated output tokens, respectively\. For phase\-level reporting, we compute:

Epre=Ecompute,pre\+Ememory,pre,Edec=Ecompute,dec\+Ememory,dec\.E\_\{\\mathrm\{pre\}\}=E\_\{\\mathrm\{compute,pre\}\}\+E\_\{\\mathrm\{memory,pre\}\},\\qquad E\_\{\\mathrm\{dec\}\}=E\_\{\\mathrm\{compute,dec\}\}\+E\_\{\\mathrm\{memory,dec\}\}\.\(14\)whereEcompute,preE\_\{\\mathrm\{compute,pre\}\}andEcompute,decE\_\{\\mathrm\{compute,dec\}\}denote the tensor\-core compute energy consumed during the prefill and decode phases, respectively, andEmemory,preE\_\{\\mathrm\{memory,pre\}\}andEmemory,decE\_\{\\mathrm\{memory,dec\}\}denote the corresponding memory\-movement energy\. The input\-token and output\-token metrics are then obtained using Equation[2](https://arxiv.org/html/2607.26571#S2.E2)\.

### 2\.5Hardware Coefficients

The hardware coefficients convert tensor\-core FLOPs and HBM traffic into energy\. In the main evaluation, we instantiate the estimator for H100\-class FP16/BF16 inference using the fixed coefficients reported in Table[1](https://arxiv.org/html/2607.26571#S3.T1):αTC=0\.52\\alpha\_\{\\mathrm\{TC\}\}=0\.52pJ/FLOP andeHBM=11\.68e\_\{\\mathrm\{HBM\}\}=11\.68pJ/bit\. These values are based on microarchitectural accelerator\-energy measurements of GPU tensor\-core computation and HBM access reported by Antepara et al\.\[antepara2025benchmark\]\. These coefficients represent microarchitectural accelerator energy rather than wall\-plug or datacenter\-level energy\. They are GPU\-level compute and HBM\-transfer coefficients\. Additional coefficients for other accelerator classes, such as A100, are provided in the Supplementary Table[6](https://arxiv.org/html/2607.26571#A5.T6)\.

### 2\.6Simplified Parameter\-Only Estimator

For model inventories where detailed architecture information is unavailable, we also define a simplified estimator\. In this case, output\-token energy is approximated as:

E^out/token=αeff​K​N,\\widehat\{E\}\_\{\\mathrm\{out/token\}\}=\\alpha\_\{\\mathrm\{eff\}\}KN,\(15\)whereαeff\\alpha\_\{\\mathrm\{eff\}\}is an effective energy\-per\-FLOP coefficient that absorbs compute, memory, and utilization effects\. This parameter\-scaled approximation is consistent with prior analyses showing that transformer compute scales linearly with model size\[kaplan2020scaling,hoffmann2022training\]\. Input\-token energy is modeled as a prompt\-length\-dependent multiple of output\-token energy:

E^in/token=M​\(Tin\)​αeff​K​N\.\\widehat\{E\}\_\{\\mathrm\{in/token\}\}=M\(T\_\{\\mathrm\{in\}\}\)\\alpha\_\{\\mathrm\{eff\}\}KN\.\(16\)
The request\-level estimate is:

E^request=Tin​E^in/token\+Tout​E^out/token\.\\widehat\{E\}\_\{\\mathrm\{request\}\}=T\_\{\\mathrm\{in\}\}\\widehat\{E\}\_\{\\mathrm\{in/token\}\}\+T\_\{\\mathrm\{out\}\}\\widehat\{E\}\_\{\\mathrm\{out/token\}\}\.\(17\)
whereM​\(Tin\)M\(T\_\{\\mathrm\{in\}\}\)is a prompt\-length\-dependent prefill multiplier\. Its functional form is defined in Section[3\.4](https://arxiv.org/html/2607.26571#S3.SS4)\.

The simplified estimator is useful when only parameter counts and token counts are known\. The architecture\-aware estimator in equation[12](https://arxiv.org/html/2607.26571#S2.E12)should be preferred whenever layer count, hidden dimension, KV\-cache dimension, precision, and workload assumptions are available\.

## 3Evaluation Setup and Assumptions

This section specifies the hardware configuration, model set, workload assumptions, and fixed estimator parameters used in the analytical evaluation\. The purpose is to make the reported energy estimates reproducible and to separate methodological assumptions from the numerical results\.

### 3\.1Target Hardware and Precision

The evaluation is instantiated for NVIDIA H100\-class accelerators, corresponding to the inference hardware considered in this work\. Unless otherwise stated, all models are assumed to run using FP16 or BF16 tensor\-core inference\. Model weights and KV\-cache entries are therefore assumed to use 16\-bit precision:

bw=bkv=16\.b\_\{w\}=b\_\{\\mathrm\{kv\}\}=16\.\(18\)For the H100\-class configuration, we use the accelerator\-level energy coefficients summarized in Table[1](https://arxiv.org/html/2607.26571#S3.T1)\. Here,αTC\\alpha\_\{\\mathrm\{TC\}\}denotes the energy per FP16/BF16 tensor\-core FLOP andeHBMe\_\{\\mathrm\{HBM\}\}denotes the energy per bit transferred through high\-bandwidth memory\. These coefficients represent GPU\-level microarchitectural energy and do not include CPU, cooling, networking, or datacenter\-level overheads\.

### 3\.2Model Set

We evaluate a representative set of transformer\-based LLMs that cover small \(<3​B<3B\), medium \(3​B−30​B3B\-30B\), and large \(\>30​B\>30B\) parameter regimes as summarized in Supplementary Table[5](https://arxiv.org/html/2607.26571#A3.T5)\. The model set includes embedding models, general\-purpose decoder\-only LLMs, code models, reasoning models, and vision\-language models\. For each model, the estimator requires at minimum the parameter countNNand the input/output token counts\. When available, additional architecture\-specific quantities such as the number of layersnℓn\_\{\\ell\}, hidden dimensiondmodeld\_\{\\mathrm\{model\}\}, Kv and effective KV\-cache dimensiondkvd\_\{\\mathrm\{kv\}\}are used by the architecture\-aware estimator\. If full architecture metadata is unavailable, the simplified parameter\-only estimator is used\. In this case, the model is represented by its parameter count and token workload only\. This enables consistent comparison across heterogeneous model inventories while preserving explicit assumptions\.

### 3\.3Workload Scenarios

The analytical evaluation considers request\-level inference workloads defined by the number of input tokensTinT\_\{\\mathrm\{in\}\}and generated output tokensToutT\_\{\\mathrm\{out\}\}\. The default workload used for model\-size comparisons is:

Tin=500,Tout=500\.T\_\{\\mathrm\{in\}\}=500,\\qquad T\_\{\\mathrm\{out\}\}=500\.\(19\)Additional experiments varyToutT\_\{\\mathrm\{out\}\}while holding model size and input length fixed in order to study how inference energy scales with generated sequence length\. This isolates the effect of autoregressive decoding and attention\-related KV\-cache reads\. Unless otherwise stated, the evaluation assumes single\-request execution and does not explicitly model continuous batching\. Effects of cache reuse, locality, and batching on parameter\-access traffic are represented through the parameter\-access factorγ\\gamma\.

### 3\.4Fixed Model Parameters

We adopt standard constants for transformer inference, including the dense transformer FLOP coefficientKK, along with hardware\-related energy parameters\. All fixed modeling constants and hardware coefficients are summarised in Table[1](https://arxiv.org/html/2607.26571#S3.T1)\.

Table 1:Fixed constants and hardware coefficients used in the evaluation\.ParameterValueDescriptionKK6FLOPs per parameter per token \(transformer constant\)αTC\\alpha\_\{\\mathrm\{TC\}\}0\.52 pJ/FLOPH100 tensor\-core energy coefficienteHBMe\_\{\\mathrm\{HBM\}\}11\.68 pJ/bitHBM energy per bit transferredbwb\_\{w\}16Bits per model weight \(FP16/BF16\)bkvb\_\{\\mathrm\{kv\}\}16Bits per KV\-cache elementThe individual memory\-traffic terms in Equation[10](https://arxiv.org/html/2607.26571#S2.E10)are specified by the parameter\-access, KV\-cache\-write, and attention\-read terms below:

Bitsparams\\displaystyle\\mathrm\{Bits\}\_\{\\mathrm\{params\}\}=bw​N​γ​\(N\),γ​\(N\)=γ0​\(NN0\)β,\\displaystyle=b\_\{w\}N\\gamma\(N\),\\hskip 15\.00002pt\\gamma\(N\)=\\gamma\_\{0\}\\left\(\\frac\{N\}\{N\_\{0\}\}\\right\)^\{\\beta\},\(20\)BitsKV\\displaystyle\\mathrm\{Bits\}\_\{\\mathrm\{KV\}\}=2​bkv​dmodel​nℓ​Tout,\\displaystyle=2b\_\{\\mathrm\{kv\}\}d\_\{\\mathrm\{model\}\}n\_\{\\ell\}T\_\{\\mathrm\{out\}\},\(21\)Bitsattn′\\displaystyle\\mathrm\{Bits\}\_\{\\mathrm\{attn\}\}^\{\\prime\}=\[2​bkv​dmodel​nℓ​\(Tout​Tin\+Tout​\(Tout−1\)2\)\]​sattn​\(N\)\.\\displaystyle=\\left\[2b\_\{\\mathrm\{kv\}\}d\_\{\\mathrm\{model\}\}n\_\{\\ell\}\\left\(T\_\{\\mathrm\{out\}\}T\_\{\\mathrm\{in\}\}\+\\frac\{T\_\{\\mathrm\{out\}\}\(T\_\{\\mathrm\{out\}\}\-1\)\}\{2\}\\right\)\\right\]s\_\{\\mathrm\{attn\}\}\(N\)\.\(22\)Here,γ​\(N\)\\gamma\(N\)captures effective parameter reuse, whilesattn​\(N\)s\_\{\\mathrm\{attn\}\}\(N\)captures non\-ideal KV\-cache read overhead; their calibrated forms are given in Table[2](https://arxiv.org/html/2607.26571#S3.T2)\. Full derivations and interpretation are provided in Supplementary Section[B](https://arxiv.org/html/2607.26571#A2)\.

For the simplified parameter\-only estimator, the prefill multiplierM​\(Tin\)M\(T\_\{\\mathrm\{in\}\}\)is selected according to the input length bucket:

M​\(Tin\)=\{1\.2,Tin≤2048,1\.8,2048<Tin≤5120,3\.0,5120<Tin≤10240,4\.0,Tin\>10240\.M\(T\_\{\\mathrm\{in\}\}\)=\\begin\{cases\}1\.2,&T\_\{\\mathrm\{in\}\}\\leq 2048,\\\\ 1\.8,&2048<T\_\{\\mathrm\{in\}\}\\leq 5120,\\\\ 3\.0,&5120<T\_\{\\mathrm\{in\}\}\\leq 10240,\\\\ 4\.0,&T\_\{\\mathrm\{in\}\}\>10240\.\\end\{cases\}\(23\)
This multiplier is used only in the simplified parameter\-only estimator\.

### 3\.5Calibration Procedure

We estimate the model parameters using a data\-driven calibration procedure based on reported energy measurements for LLM inference\[caravaca2025prompts\]\. For each model size, we first decompose the total energy into compute and memory components, using the analytical expressions derived in the previous sections\. The remaining memory contribution is then further decomposed into parameter, KV\-cache, and attention terms\. Using this decomposition, we obtain approximate estimates of the effective parameter\-access factorγ​\(N\)\\gamma\(N\)for each model:

γ​\(N\)=Bitsparamsbw​N,\\gamma\(N\)=\\frac\{\\mathrm\{Bits\}\_\{\\mathrm\{params\}\}\}\{b\_\{w\}N\},\(24\)whereBitsparams\\mathrm\{Bits\}\_\{\\mathrm\{params\}\}is inferred from measured energy after subtracting compute and other memory contributions\. The parametersγ0\\gamma\_\{0\}andβ\\betaare then determined by fitting the power\-law model to these inferred values, while also ensuring consistency with the overall energy estimates\. The parameters ofsattn​\(N\)s\_\{\\mathrm\{attn\}\}\(N\)\(Equation[22](https://arxiv.org/html/2607.26571#S3.E22)\) andη​\(N\)\\eta\(N\)\(Equation[11](https://arxiv.org/html/2607.26571#S2.E11)\) are calibrated by minimizing the deviation between model predictions and measurement\-based reported energy values across the evaluated models\. In practice, we perform a low\-dimensional parameter search over the coefficients of the scaling functions, selecting values that minimize the relative error while preserving the expected scaling trends with model size\.

This procedure is a constrained least\-squares fitting over a small number of parameters, where the objective is to match both the magnitude and growth behavior of measured energy consumption rather than exactly fitting individual data points\. Due to the small number of calibration parameters and limited data points, the calibration is intentionally kept simple to avoid overfitting and to retain generalizability across models and workloads\. After calibration, these factors are kept fixed across the scaling and decomposition experiments\.

This calibration makes the estimator an analytically structured, empirically calibrated model: transformer FLOP counts and memory\-traffic terms provide the analytical structure, while the calibrated factors represent non\-ideal reuse, attention\-access overhead, and global HBM inefficiency\. The factors should therefore be recalibrated when applying the methodology to a different GPU generation, inference engine, batching regime, or serving configuration\.

Table 2:Calibrated model parameters used in the analytical estimator\.ParameterValueDescriptionγ0\\gamma\_\{0\}0\.10Baseline parameter\-access factor at reference model sizeN0N\_\{0\}N0N\_\{0\}24BReference model size for scaling relationshipsβ\\beta0\.8Exponent controlling degradation of parameter reuse with model sizesattn​\(N\)s\_\{\\mathrm\{attn\}\}\(N\)1\+1\.5​\(NN0\)0\.9\\displaystyle 1\+1\.5\\left\(\\frac\{N\}\{N\_\{0\}\}\\right\)^\{0\.9\}Attention scaling factor capturing KV\-cache read overheadη​\(N\)\\eta\(N\)1\+0\.8​\(NN0\)0\.8\\displaystyle 1\+0\.8\\left\(\\frac\{N\}\{N\_\{0\}\}\\right\)^\{0\.8\}Global memory\-inefficiency factor for HBM traffic
### 3\.6Reported Metrics

For each model–workload pair, the estimator reports prefill energyEpreE\_\{\\mathrm\{pre\}\}, decode energyEdecE\_\{\\mathrm\{dec\}\}, total request energyErequestE\_\{\\mathrm\{request\}\}, input\-token energy, output\-token energy, and average energy per processed token\. The full description of these metrics is provided in Supplementary Section[D](https://arxiv.org/html/2607.26571#A4)\.

### 3\.7Interpretation of Estimates

All reported values are analytical GPU\-level estimates and are scenario\-dependent approximations rather than direct measurements or intrinsic properties of the models\. Differences between measured and estimated energy may arise from batching, tensor\-parallel communication, framework overheads, kernel fusion, quantization, cache behavior, and runtime GPU utilization\. These effects are discussed further in the limitations section\.

## 4Results and Evaluation

This section presents the analytical energy estimates obtained using the setup defined in Section[3](https://arxiv.org/html/2607.26571#S3)\. We evaluate the estimator along four dimensions: scaling with generated sequence length, scaling with model size, decomposition of energy into compute and memory components, and comparison against measurement\-based results from prior work\. We also report token\-level energy estimates for the model inventory considered in this study\. Unless otherwise stated, all results assume FP16/BF16 inference on H100\-class accelerators, a dense transformer FLOP constantK=6K=6, the calibrated parameter\-access factorγ​\(N\)\\gamma\(N\), the attention\-access factorsattn​\(N\)s\_\{\\mathrm\{attn\}\}\(N\), the memory\-inefficiency factorη​\(N\)\\eta\(N\), and the hardware coefficientsαTC=0\.52\\alpha\_\{\\mathrm\{TC\}\}=0\.52pJ/FLOP andeHBM=11\.68e\_\{\\mathrm\{HBM\}\}=11\.68pJ/bit\. The set of models considered in this study and their key architectural characteristics are summarized in Table S[5](https://arxiv.org/html/2607.26571#A3.T5), which can be found in the Supplementary material\.

### 4\.1Scaling with Generated Output Length

We first evaluate how inference energy scales with the number of generated tokens\. For this experiment, we fix the model size to 32B parameters and use an input length ofTin=100T\_\{\\mathrm\{in\}\}=100tokens\. We then vary the number of generated output tokensToutT\_\{\\mathrm\{out\}\}\.

![Refer to caption](https://arxiv.org/html/2607.26571v1/figures/scaling_with_output_sequence_length.png)Figure 1:Energy breakdown as a function of generated output lengthToutT\_\{\\mathrm\{out\}\}for a 32B\-parameter model withTin=100T\_\{\\mathrm\{in\}\}=100input tokens\. Compute and KV\-cache write costs scale almost linearly withToutT\_\{\\mathrm\{out\}\}, while attention\-related KV\-cache reads exhibit quadratic growth and become increasingly significant at larger output lengths\.Figure[1](https://arxiv.org/html/2607.26571#S4.F1)shows that compute energy grows approximately linearly withToutT\_\{\\mathrm\{out\}\}, consistent with the autoregressive decoding process in which each generated token requires one forward pass through the model\. KV\-cache write traffic also grows linearly, because each generated token contributes one new key and value entry per layer\. In contrast, scaled attention\-related memory traffic grows super\-linearly\. This is due both to the quadratic growth of KV\-cache reads with output length and to the attention\-access factorsattn​\(N\)s\_\{\\mathrm\{attn\}\}\(N\)used in the calibrated memory model\. During decode, each newly generated token attends over the accumulated context, so the total number of KV\-cache reads increases as:

Tout​Tin\+Tout​\(Tout−1\)2\.T\_\{\\mathrm\{out\}\}T\_\{\\mathrm\{in\}\}\+\\frac\{T\_\{\\mathrm\{out\}\}\(T\_\{\\mathrm\{out\}\}\-1\)\}\{2\}\.\(25\)

### 4\.2Scaling with Model Size

We next evaluate the dependence of request\-level energy on model size\. The workload is fixed to:

Tin=500,Tout=500\.T\_\{\\mathrm\{in\}\}=500,\\qquad T\_\{\\mathrm\{out\}\}=500\.\(26\)
This isolates the effect of model parameter count while holding the token workload constant\.

![Refer to caption](https://arxiv.org/html/2607.26571v1/figures/energy_vs_model_size_loglog.png)Figure 2:Energy per request as a function of model size for a fixed workload of 500 input tokens and 500 output tokens\. Total energy, compute energy, and memory energy are shown on a log–log scale\.Figure[2](https://arxiv.org/html/2607.26571#S4.F2)shows that total request energy increases approximately linearly with model size on a log–log scale\. This behavior is expected from the parameter\-scaled compute model:

Cdec,dense=K​N​Tout,Cpre,dense=K​N​Tin\.C\_\{\\mathrm\{dec,dense\}\}=KNT\_\{\\mathrm\{out\}\},\\qquad C\_\{\\mathrm\{pre,dense\}\}=KNT\_\{\\mathrm\{in\}\}\.\(27\)
For fixed input and output lengths, the dominant dense matrix operations scale linearly with the number of parametersNN\. Consequently, the compute component follows:

Ecompute∝N\.E\_\{\\mathrm\{compute\}\}\\propto N\.\(28\)
The memory component increases with model size through the calibrated factorγ​\(N\)\\gamma\(N\), the attention\-access factorsattn​\(N\)s\_\{\\mathrm\{attn\}\}\(N\), and the global memory\-inefficiency factorη​\(N\)\\eta\(N\)\. Under the fixed 500\-token/500\-token workload, compute remains the dominant component for most evaluated model sizes, although memory energy increases nonlinearly with the calibrated memory factors\. While the overall scaling trend is approximately linear in model size, the memory\-energy component exhibits non\-monotonic behavior for certain models\. This is due to differences in architectural configurations, rather than parameter count alone\. In particular, key contributors to memory traffic, such as the hidden dimensiondmodeld\_\{\\mathrm\{model\}\}and the number of layersnℓn\_\{\\ell\}, do not scale uniformly withNNacross different models\. As a result, some medium\-sized models may exhibit lower memory traffic than nearby models with slightly different parameter counts, leading to localized deviations \(e\.g\., a reduction in the memory\-energy curve\)\. This highlights that memory costs are sensitive to architectural design choices, not solely parameter count\.

### 4\.3Energy Decomposition

To understand which terms dominate request energy, we decompose total energy into compute, parameter movement, KV\-cache writes, and attention\-related KV\-cache reads\. The decomposition is evaluated for a 32B\-parameter model while varyingToutT\_\{\\mathrm\{out\}\}\. Figure[3](https://arxiv.org/html/2607.26571#S4.F3)shows the absolute energy decomposition as a function of generated output length for a 32B\-parameter model\. The stacked representation highlights the additive structure of the energy model, where total energy is the sum of compute and memory components\.

![Refer to caption](https://arxiv.org/html/2607.26571v1/figures/energy_decomposition_vs_tokens.png)Figure 3:Fractional energy contribution as a function of generated output length for a 32B\-parameter model\. Compute dominates at short and moderate output lengths, while attention\-related memory traffic increases with sequence length due to quadratic KV\-cache read scaling\.Compute energy increases approximately linearly withToutT\_\{\\mathrm\{out\}\}, reflecting the constant per\-token cost of autoregressive decoding\. In contrast, attention\-related memory energy exhibits super\-linear growth, due to the quadratic scaling of KV\-cache reads with sequence length\. At small output lengths, total energy is dominated by compute\. AsToutT\_\{\\mathrm\{out\}\}increases, the attention\-related component grows rapidly and becomes a substantial contributor to the overall energy\. This results in an upward curvature of the total energy trend, indicating the increasing impact of memory movement at longer sequence lengths\.

The contributions from parameter access and KV\-cache writes are negligible relative to compute and attention\-related memory traffic across the evaluated range\. In particular, the KV\-cache write component, although included in the model, is not visually distinguishable in the stacked representation\. This is because KV writes scale linearly withToutT\_\{\\mathrm\{out\}\}, while attention\-related KV\-cache reads scale quadratically and dominate memory traffic at larger sequence lengths\.

### 4\.4Token\-Level Energy Estimates for the Model Inventory

Table[3](https://arxiv.org/html/2607.26571#S4.T3)reports simplified token\-level estimates for the model inventory using the parameter\-only compute\-dominated estimator:

E^out/token=αTC​K​N,E^in/token≈1\.2​E^out/token,\\widehat\{E\}\_\{\\mathrm\{out/token\}\}=\\alpha\_\{\\mathrm\{TC\}\}KN,\\qquad\\widehat\{E\}\_\{\\mathrm\{in/token\}\}\\approx 1\.2\\,\\widehat\{E\}\_\{\\mathrm\{out/token\}\},\(29\)where the input\-token multiplier corresponds to the short\-prompt setting\. These values are simplified, compute\-dominated estimates; further interpretation is provided in Supplementary Section[D](https://arxiv.org/html/2607.26571#A4)\.

Table 3:Estimated energy consumption per token and per request across evaluated models\.ModelParams \(B\)Eout/tokenE\_\{\\mathrm\{out/token\}\}Ein/tokenE\_\{\\mathrm\{in/token\}\}ErequestE\_\{\\mathrm\{request\}\}\(mJ/token\)\(mJ/token\)\(Wh\)EmbeddingGemma0\.3080\.9611\.1530\.000599MXBAI Embed Large0\.3341\.0421\.2500\.000696Qwen3 Embedding0\.6001\.8722\.2460\.001027Qwen3 \(1\.7B\)1\.7005\.3046\.3650\.002641Granite 3\.2 Vision2\.5307\.8949\.4720\.004923Qwen3 \(8B\)8\.00024\.96029\.9520\.011795Granite 3\.38\.17025\.49030\.5880\.012465Ministral 3 \(14B\)14\.00043\.68052\.4160\.021313DeepSeek\-Coder V216\.00049\.92059\.9040\.017425GPT\-OSS \(20B\)20\.00062\.40074\.8800\.022281Qwen3 \(32B\)32\.00099\.840119\.8080\.052721Qwen2\.5\-Coder \(32B\)32\.00099\.840119\.8080\.052721Qwen3\-VL \(32B\)32\.00099\.840119\.8080\.052721DeepSeek\-R132\.00099\.840119\.8080\.052721Llama 3\.3 \(70B\)70\.000218\.400262\.0800\.170747GPT\-OSS \(120B\)120\.000374\.400449\.2800\.155204Table[3](https://arxiv.org/html/2607.26571#S4.T3)shows the expected linear dependence of token\-level energy on model size\. Sub\-billion\-parameter embedding models have estimated token\-level costs below approximately 2\.3 mJ/input token under the simplified estimator, whereas 70B\- and 120B\-parameter models require substantially higher per\-token energy\. For example, the simplified estimate for LLaMA 3\.3 70B is 218\.4 mJ/output token and 262\.1 mJ/input token under the short\-prompt assumption\. The 120B model reaches 374\.4 mJ/output token and 449\.3 mJ/input token\. These values are useful for comparative model selection because they expose the energy scaling of energy consumption with parameter count\. The analysis isolates the cost of token processing and does not account for task performance\.

Models with identical parameter counts yield identical token\-level estimates under the parameter\-scaled compute model, since per\-token FLOPs depend only on the number of parameters\. Architectural differences influence memory\-related energy and full request\-level costs, but are not reflected in these simplified token\-level estimates\. Finally, modality\-specific models such as vision\-language architectures may incur additional costs that are not captured by the parameter\-only formulation\.

### 4\.5Comparison with Measurement\-Based Results

Finally, we compare the analytical estimates against the measurement\-based study of Caravaca et al\.\[caravaca2025prompts\]\. The comparison uses the same nominal workload:

Tin=500,Tout=500\.T\_\{\\mathrm\{in\}\}=500,\\qquad T\_\{\\mathrm\{out\}\}=500\.\(30\)
Table 4:Analytical versus measured energy per prompt for a 500\-token input and 500\-token output workload\.Model sizeAnalytical energyMeasured energyError\(B parameters\)\(Wh/request\)\(Wh/request\)\(%\)80\.0117950\.00927027\.23240\.0324270\.02628023\.39700\.1707470\.1792304\.73720\.1707470\.23326026\.80Table[4](https://arxiv.org/html/2607.26571#S4.T4)shows that the analytical estimator matches measurement\-based values within approximately 5–27% for the compared cases\. The lowest error is observed for the 70B model, where the analytical estimate differs from the measured value by 4\.73%\. For the 8B, 24B, and 72B cases, the relative error remains below 30%\. The remaining discrepancies are expected\. The analytical model estimates GPU\-level compute and memory movement under controlled assumptions, whereas measurement\-based studies include additional effects from inference engines, batching policies, runtime scheduling, tensor\-parallel execution, kernel fusion, GPU utilization, and system\-level overheads\. Differences may also arise from comparing models of similar size but different architecture\. Therefore, the comparison shows that the estimator captures the correct order of magnitude and scaling behavior, rather than as exact per\-deployment energy accounting\.

Overall, the results indicate that the proposed analytical estimator provides a transparent and reproducible approximation of LLM inference energy\. It captures the expected linear scaling with model size, the super\-linear effect of attention\-related memory traffic in long generations, and the distinction between input\-token, output\-token, and request\-level energy\.

## 5Limitations and Conclusion

The proposed estimator is limited to accelerator\-side operational energy\. It excludes CPU execution, host memory, networking, storage, power\-supply losses, cooling, and datacenter\-level power usage effectiveness\. Consequently, the reported values are GPU\-level operational energy estimates, not end\-to\-end service energy, datacenter energy, carbon emissions, or lifecycle emissions\. Nevertheless, accelerator\-side energy is a major controllable component in GPU\-dense on\-premise inference deployments: H100 SXM\-class accelerators have a thermal design power of approximately 700 W, and multi\-GPU servers can therefore be dominated by accelerator power under high utilization\[nvidia\_h100\_datasheet\]\. Facility\-level energy remains larger because of cooling and power\-delivery overheads; for example, Uptime Institute reports an average data\-center PUE of approximately 1\.56, and MLPerf Power treats inference energy as a system\-level measurement problem\[uptime2024survey,mlperfpower2024,mlperfpowerdocs\]\.

The estimator also abstracts away several architecture\- and system\-specific effects\. The parameter\-scaled FLOP model does not fully capture feed\-forward expansion ratios, grouped\-query attention, multi\-query attention, mixture\-of\-experts routing, or modality\-specific processing in vision\-language models\. The memory model approximates weight access, KV\-cache writes, and attention reads using calibrated factorsγ​\(N\)\\gamma\(N\),sattn​\(N\)s\_\{\\mathrm\{attn\}\}\(N\), andη​\(N\)\\eta\(N\), which summarize parameter reuse, attention\-access overhead, and HBM inefficiency\. These factors do not explicitly model the complete GPU memory hierarchy, kernel scheduling, cache residency, batching behavior, or inference engine implementation, and should be recalibrated for other hardware platforms, serving engines, or batching regimes\.

Despite these limitations, the methodology provides a reproducible framework for comparing GPU\-level inference energy across models and workloads when direct instrumentation is unavailable\. It also identifies actionable inference\-stage levers for reducing energy: smaller or task\-specialized models reduce parameter\-scaled compute; shorter outputs reduce autoregressive decoding cost; prompt compression and retrieval filtering reduce context length; quantized weights and KV caches reduce memory traffic; and batching, prefix caching, efficient attention kernels, speculative decoding, and model routing can improve serving efficiency\. The method is therefore best understood as a complementary tool for green\-coding analysis, comparative model selection, and design\-time evaluation, while precise deployment accounting still requires system\-level power measurement\. Future work should integrate measured serving traces, extend the estimator to quantized and mixture\-of\-experts models, and incorporate tensor\-parallel communication and batching effects\.

## Declaration on Generative AI

During the preparation of this work, the authors used OpenAI ChatGPT for language editing, restructuring, consistency checking, and drafting assistance\. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the publication’s content\.

## Acknowledgments

The authors gratefully acknowledge the ITEA GreenCode research project for fostering discussions on energy\-aware AI systems and for funding parts of this work\. The authors also thank their colleagues at TWT Science & Innovation, Stefanos Papanikolaou and Michael Herrnberger, for their valuable discussions and insights\.

## References

## Supplementary Material

## Appendix ADerivation of the Compute Model

This section provides the architecture\-aware compute terms used by the estimator\. The main paper reports the compact parameter\-scaled form, while the full attention correction terms are given here for reproducibility\.

For dense decoder\-only transformers, the dominant computation arises from matrix multiplications in attention projections and feed\-forward layers\. When only the model parameter count is available, the dense decode\-phase FLOPs per generated token, denotedCdec/tokenC\_\{\\mathrm\{dec/token\}\}, are approximated as:

Cdec/token≈K​N,C\_\{\\mathrm\{dec/token\}\}\\approx KN,\(31\)whereNNis the number of model parameters andKKis a transformer FLOP constant\. Following standard transformer FLOP accounting, we use:

which corresponds to a commonly used approximation that dense transformer inference requires on the order of six FLOPs per parameter per token for the forward pass\[kaplan2020scaling,hoffmann2022training\]\.

The dense compute cost of the decode phase is therefore:

Cdec,dense=K​N​Tout\.C\_\{\\mathrm\{dec,dense\}\}=KNT\_\{\\mathrm\{out\}\}\.\(33\)
Similarly, the dense compute cost of the prefill phase is:

Cpre,dense=K​N​Tin\.C\_\{\\mathrm\{pre,dense\}\}=KNT\_\{\\mathrm\{in\}\}\.\(34\)
For long contexts, additional attention terms can be included\. If architecture\-level parameters are available, the prefill attention correction is approximated as:

Cpre,attn≈2​nℓ​dmodel​Tin2,C\_\{\\mathrm\{pre,attn\}\}\\approx 2n\_\{\\ell\}d\_\{\\mathrm\{model\}\}T\_\{\\mathrm\{in\}\}^\{2\},\(35\)wherenℓn\_\{\\ell\}is the number of transformer layers anddmodeld\_\{\\mathrm\{model\}\}is the hidden dimension\.

During decode, the token generated at stepttattends to a context of lengthTin\+t−1T\_\{\\mathrm\{in\}\}\+t\-1\. The attention\-related decode cost is approximated as:

Cdec,attn≈4​nℓ​dmodel​\(Tout​Tin\+Tout​\(Tout−1\)2\)\.C\_\{\\mathrm\{dec,attn\}\}\\approx 4n\_\{\\ell\}d\_\{\\mathrm\{model\}\}\\left\(T\_\{\\mathrm\{out\}\}T\_\{\\mathrm\{in\}\}\+\\frac\{T\_\{\\mathrm\{out\}\}\(T\_\{\\mathrm\{out\}\}\-1\)\}\{2\}\\right\)\.\(36\)
The total tensor\-core compute estimate is:

CTC=Cpre,dense\+Cdec,dense\+Cpre,attn\+Cdec,attn\.C\_\{\\mathrm\{TC\}\}=C\_\{\\mathrm\{pre,dense\}\}\+C\_\{\\mathrm\{dec,dense\}\}\+C\_\{\\mathrm\{pre,attn\}\}\+C\_\{\\mathrm\{dec,attn\}\}\.\(37\)
When detailed architectural parameters are unavailable, the attention correction terms are omitted and the parameter\-scaled approximation is used\.

## Appendix BDerivation of the Memory\-Movement Model

This section provides the full memory\-traffic decomposition used by the architecture\-aware estimator\. The total HBM traffic is decomposed into parameter access, KV\-cache writes, and attention\-related KV\-cache reads:

Bitstotal=Bitsparams\+BitsKV\+Bitsattn′,\\mathrm\{Bits\}\_\{\\mathrm\{total\}\}=\\mathrm\{Bits\}\_\{\\mathrm\{params\}\}\+\\mathrm\{Bits\}\_\{\\mathrm\{KV\}\}\+\\mathrm\{Bits\}\_\{\\mathrm\{attn\}\}^\{\\prime\},\(38\)whereBitsparams\\mathrm\{Bits\}\_\{\\mathrm\{params\}\}denotes parameter\-access traffic,BitsKV\\mathrm\{Bits\}\_\{\\mathrm\{KV\}\}denotes KV\-cache write traffic, andBitsattn′\\mathrm\{Bits\}\_\{\\mathrm\{attn\}\}^\{\\prime\}denotes scaled attention\-related KV\-cache read traffic\.

### B\.1Parameter\-Access Traffic

Model weights are stored in GPU memory and accessed during the forward passes required for inference\. A naive upper\-bound formulation would assume that the full set of model weights is transferred from HBM at every decoding step, leading to memory traffic proportional to the number of generated tokens\. However, modern GPU implementations reduce this cost through on\-chip caching, kernel fusion, memory locality, and overlap between memory access and computation\.

To account for imperfect parameter reuse, we introduce a model\-size\-dependent parameter\-access factorγ​\(N\)\\gamma\(N\):

Bitsparams=bw​N​γ​\(N\),\\mathrm\{Bits\}\_\{\\mathrm\{params\}\}=b\_\{w\}N\\gamma\(N\),\(39\)wherebwb\_\{w\}is the number of bits per model weight andNNis the number of model parameters\.

The parameter\-access factorγ​\(N\)∈\(0,1\]\\gamma\(N\)\\in\(0,1\]represents the effective fraction of model parameters retrieved from HBM throughout the inference request\. We model its dependence on model size as:

γ​\(N\)=γ0​\(NN0\)β,\\gamma\(N\)=\\gamma\_\{0\}\\left\(\\frac\{N\}\{N\_\{0\}\}\\right\)^\{\\beta\},\(40\)whereN0N\_\{0\}is a reference model size,γ0\\gamma\_\{0\}is the baseline reuse factor atN0N\_\{0\}, andβ\\betaapproximates the degradation of effective reuse as the model working set exceeds on\-chip memory capacity\.

### B\.2KV\-Cache Write Traffic

For each generated token, the model stores key and value vectors in each transformer layer\. The KV\-cache write traffic is approximated as:

BitsKV=2​bkv​dmodel​nℓ​Tout,\\mathrm\{Bits\}\_\{\\mathrm\{KV\}\}=2b\_\{\\mathrm\{kv\}\}d\_\{\\mathrm\{model\}\}n\_\{\\ell\}T\_\{\\mathrm\{out\}\},\(41\)wherebkvb\_\{\\mathrm\{kv\}\}is the number of bits per KV\-cache element,dmodeld\_\{\\mathrm\{model\}\}is the hidden dimension,nℓn\_\{\\ell\}is the number of transformer layers, andToutT\_\{\\mathrm\{out\}\}is the number of generated output tokens\. The factor of two accounts for storing both keys and values\.

### B\.3Attention\-Related KV\-Cache Read Traffic

During autoregressive decoding, each newly generated token attends over the input prompt and the previously generated tokens through the KV cache\. For an input lengthTinT\_\{\\mathrm\{in\}\}and an output sequence of lengthToutT\_\{\\mathrm\{out\}\}, the baseline attention\-related memory traffic is:

Bitsattn=2​bkv​dmodel​nℓ​\(Tout​Tin\+Tout​\(Tout−1\)2\)\.\\mathrm\{Bits\}\_\{\\mathrm\{attn\}\}=2b\_\{\\mathrm\{kv\}\}d\_\{\\mathrm\{model\}\}n\_\{\\ell\}\\left\(T\_\{\\mathrm\{out\}\}T\_\{\\mathrm\{in\}\}\+\\frac\{T\_\{\\mathrm\{out\}\}\(T\_\{\\mathrm\{out\}\}\-1\)\}\{2\}\\right\)\.\(42\)
This term includes a linear prompt\-attention component and a quadratic generated\-token component\. To account for non\-ideal memory behavior, including irregular access patterns, limited locality, cache contention, and increasing memory pressure at larger model sizes, we apply an attention\-specific scaling factor:

Bitsattn′=Bitsattn⋅sattn​\(N\),\\mathrm\{Bits\}\_\{\\mathrm\{attn\}\}^\{\\prime\}=\\mathrm\{Bits\}\_\{\\mathrm\{attn\}\}\\cdot s\_\{\\mathrm\{attn\}\}\(N\),\(43\)
wheresattn​\(N\)≥1s\_\{\\mathrm\{attn\}\}\(N\)\\geq 1is a dimensionless scaling factor\.

The attention\-related terms follow the quadratic dependence on sequence length characteristic of standard self\-attention, along with the associated memory\-access costs on GPU memory hierarchies\[dao2022flashattention\]\.

The importance of KV\-cache memory traffic in LLM serving has also been emphasized in prior system\-level work, including PagedAttention and vLLM\[kwon2023efficient\]\.

### B\.4HBM\-Dominated Memory Energy

The memory\-energy term is computed as:

Ememory=Bitstotal⋅eHBM⋅η​\(N\),E\_\{\\mathrm\{memory\}\}=\\mathrm\{Bits\}\_\{\\mathrm\{total\}\}\\cdot e\_\{\\mathrm\{HBM\}\}\\cdot\\eta\(N\),\(44\)
whereeHBMe\_\{\\mathrm\{HBM\}\}is the energy per bit transferred from HBM andη​\(N\)≥1\\eta\(N\)\\geq 1is a global memory\-inefficiency factor\. The factorη​\(N\)\\eta\(N\)approximates bandwidth saturation, memory\-controller overhead, cache contention, and pipeline stalls under high memory pressure\.

The factorsγ​\(N\)\\gamma\(N\),sattn​\(N\)s\_\{\\mathrm\{attn\}\}\(N\), andη​\(N\)\\eta\(N\)are empirical calibration terms fitted against measurement\-based LLM inference energy results reported by Caravaca et al\.\[caravaca2025prompts\]\.

## Appendix CModel Architecture Details

In this section we present key architectural details of the models that have been used in this study\. Table[5](https://arxiv.org/html/2607.26571#A3.T5)shows the parameter, layer, and hidden dimension \(dmodeld\_\{\\mathrm\{model\}\}\) of the model architectures\.

Table 5:Overview of Model Architectural Configurations\.ModelParamsLayersdm​o​d​e​ld\_\{model\}Ministral 3 \(14B\)\[ministral3\]14B405120EmbeddingGemma\[embeddinggemma\]308M26768MXBAI Embed Large\[mxbai\]334M241024Qwen3 Embedding \(0\.6B\)\[qwen3embed\]0\.6B281024Qwen3\-VL \(32B\)\[qwen3vl\]32B645120Granite 3\.3 \(8B\)\[granite33\]8\.17B404096Granite 3\.2 Vision\[granite32\]2\.53B324096DeepSeek\-Coder V2 \(16B\)\[deepseekcoder\]16B272048DeepSeek\-R1 \(32B\)\[deepseekr1\]32B645120Qwen3 \(8B\)\[qwen38b\]8B364096Qwen3 \(32B\)\[qwen332b\]32B645120Qwen3 \(1\.7B\)\[qwen31p7b\]1\.7B282048Qwen2\.5\-Coder \(32B\)\[qwen25coder\]32B645120Llama 3\.3 \(70B\)\[llama33\]70B808192GPT\-OSS \(120B\)\[gptoss120\]120B362880GPT\-OSS \(20B\)\[gptoss20\]20B242880
## Appendix DReported Metrics

For each model–workload pair, the estimator reports energy at three levels of granularity: phase\-level energy, token\-normalized energy, and aggregate request\-level energy\. This separation is necessary because LLM inference is not a homogeneous operation: prompt prefill and autoregressive decoding differ in their compute structure, memory\-access pattern, and dependence on sequence length\.

The phase\-level quantities are the prefill energy,EpreE\_\{\\mathrm\{pre\}\}, and the decode energy,EdecE\_\{\\mathrm\{dec\}\}\. These terms represent the estimated GPU energy required to process the input prompt and to generate the output sequence, respectively\. The total request energy is defined as

Erequest=Epre\+Edec\.E\_\{\\mathrm\{request\}\}=E\_\{\\mathrm\{pre\}\}\+E\_\{\\mathrm\{dec\}\}\.\(45\)
To compare requests with different prompt and generation lengths, we normalize phase\-level energy by the corresponding token counts\. The input\-token energy is defined as

Ein/token=EpreTin,E\_\{\\mathrm\{in/token\}\}=\\frac\{E\_\{\\mathrm\{pre\}\}\}\{T\_\{\\mathrm\{in\}\}\},\(46\)whereTinT\_\{\\mathrm\{in\}\}is the number of prompt tokens\. Analogously, the output\-token energy is defined as

Eout/token=EdecTout,E\_\{\\mathrm\{out/token\}\}=\\frac\{E\_\{\\mathrm\{dec\}\}\}\{T\_\{\\mathrm\{out\}\}\},\(47\)whereToutT\_\{\\mathrm\{out\}\}is the number of generated tokens\.

We additionally report the average energy per processed token:

Eavg/token=ErequestTin\+Tout\.E\_\{\\mathrm\{avg/token\}\}=\\frac\{E\_\{\\mathrm\{request\}\}\}\{T\_\{\\mathrm\{in\}\}\+T\_\{\\mathrm\{out\}\}\}\.\(48\)
This aggregate metric mixes prefill and decode costs and is therefore sensitive to the input/output token ratio\.

Finally, to analyze the dominant sources of energy consumption, the request energy is decomposed into compute and memory contributions:

Erequest=Ecompute\+Ememory\.E\_\{\\mathrm\{request\}\}=E\_\{\\mathrm\{compute\}\}\+E\_\{\\mathrm\{memory\}\}\.\(49\)
The memory component is further attributed to parameter access, KV\-cache writes, and KV\-cache reads\. Reporting these components enables the evaluation to distinguish compute\-dominated regimes from memory\-influenced or long\-context regimes\.

### D\.1Interpretation of Simplified Token\-Level Estimates

The model\-inventory table in the main paper reports simplified token\-level energy estimates based on a parameter\-only, compute\-dominated approximation:

E^out/token=αTC​K​N,E^in/token≈1\.2​E^out/token\.\\widehat\{E\}\_\{\\mathrm\{out/token\}\}=\\alpha\_\{\\mathrm\{TC\}\}KN,\\qquad\\widehat\{E\}\_\{\\mathrm\{in/token\}\}\\approx 1\.2\\,\\widehat\{E\}\_\{\\mathrm\{out/token\}\}\.\(50\)
These values do not include the calibrated memory factorsγ​\(N\)\\gamma\(N\),sattn​\(N\)s\_\{\\mathrm\{attn\}\}\(N\), orη​\(N\)\\eta\(N\)used in the architecture\-aware evaluation\. They are first\-order comparative estimates across the model inventory, rather than full compute\-plus\-memory request\-level estimates\. In contrast, the scaling and decomposition figures in the main paper use the full calibrated compute\-plus\-memory model\.

## Appendix EImplementation and Execution Assumptions

This section describes the implementation details and system assumptions underlying the analytical energy estimator\.

#### Estimator implementation\.

The estimator is implemented as a Python\-based analytical tool that computes energy consumption from the model and the workload parameters\. Given the number of model parametersNN, number of layersnℓn\_\{\\ell\}, hidden dimensiondmodeld\_\{\\mathrm\{model\}\}, and token counts, the tool calculates FLOPs and memory traffic using the analytical formulas described in the main text\.

The implementation does not execute neural networks or perform runtime profiling\. Instead, it deterministically estimates energy based on compute and memory abstractions\. No direct GPU telemetry, wall\-plug power instrumentation, or runtime energy measurement APIs \(e\.g\., NVML or nvidia\-smi\) are used during estimation\.

#### Hardware assumptions\.

The main evaluation uses H100\-class coefficients, while additional accelerator coefficients are reported here for completeness\. Energy conversion coefficients for tensor\-core operations \(αTC\\alpha\_\{\\mathrm\{TC\}\}\) and HBM traffic \(eHBMe\_\{\\mathrm\{HBM\}\}\) are taken from measurement\-based prior work\[antepara2025benchmark\]\.

Table 6:Accelerator\-level energy coefficients for FP16/BF16 tensor\-core inference\.HardwareαTC\\alpha\_\{\\mathrm\{TC\}\}\(pJ/FLOP\)eHBMe\_\{\\mathrm\{HBM\}\}\(pJ/bit\)H100 / GH2000\.5211\.68A1000\.7013\.11
#### Software stack assumptions\.

The estimator assumes inference execution on optimized tensor\-core GPU kernels using FP16/BF16 arithmetic and HBM\-resident model weights\. It is intended to approximate inference behavior of optimized GPU\-based implementations built on CUDA and high\-performance libraries such as cuBLAS and cuDNN, as well as modern LLM inference frameworks \(e\.g\., TensorRT\-LLM, Megatron\-LM, and vLLM\)\.

The model does not explicitly simulate kernel\-level execution or software\-specific optimizations such as kernel fusion, scheduling, or memory tiling\. Instead, these system\-level effects are treated implicitly and are approximated through the empirical scaling factorsγ​\(N\)\\gamma\(N\),sattn​\(N\)s\_\{\\mathrm\{attn\}\}\(N\), andη​\(N\)\\eta\(N\), which collectively capture deviations from ideal compute and memory behavior observed in optimized inference systems\.

#### Calibration reference\.

Model parameters are calibrated using reported energy measurements for optimized LLM inference from prior work\[caravaca2025prompts\]\. These measurements serve as reference values for matching the magnitude and scaling behavior of energy consumption\.

#### Limitations\.

The estimator does not model GPU execution at the kernel or instruction level\. In particular, it does not simulate thread\-level parallelism, CUDA scheduling, or detailed memory hierarchy behavior\. Instead, such effects are approximated through calibrated scaling factors\.

As a result, the model is an analytical approximation rather than a cycle\-accurate simulation or direct hardware measurement\.

Similar Articles