Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum

arXiv cs.AI Papers

Summary

This paper introduces agentic-eCAL, a metric for evaluating energy costs in multi-agent AI workflows across edge-cloud networks, demonstrating that transmission costs are minimal but inference costs are significant.

arXiv:2609.18283v1 Announce Type: new Abstract: As telecommunication networks evolve toward autonomous 5G-Advanced and 6G operations, agentic artificial intelligence (AI) workflows, where large language models (LLMs) execute multi-step reasoning, invoke diagnostic tools, retrieve domain knowledge, and coordinate across agent teams, are increasingly embedded across the edge-cloud continuum. While the biological brain accomplishes complex cognition on an exceptionally modest metabolic power budget of approximately 20W contemporary LLMs are profoundly energy- and memory-intensive, making sustainable lifecycle orchestration a critical operational priority. However, existing AI lifecycle metrics evaluate only isolated, single-model inferences or overlook multi-agent execution graphs entirely. Consequently, network operators lack foundational models to determine whether distributed agent communication incurs meaningful energy costs and where across edge-cloud tiers agent teams should physically reside. To address this gap, we introduce agentic-eCAL, generalizing the Energy Cost of AI Lifecycle (eCAL) metric to directed multi-agent workflows by coupling a closed-form two-rate single-call energy model (compute-bound prefill and memory-bound decode) with 7-layer OSI data transport. Grounded in hundreds of GPU benchmark configurations on NVIDIA A100 and H100, 16 open-weight models and 8 orchestration topologies, we validate components of the metric and study workflow placement implications. Our findings demonstrate that inter-agent text transport incurs 0.25% of workflow energy across 5G RAN, metro, and optical links. Therefore in edge-cloud agent placement the dominant energy cost of distribution is often not the transmission of inter-agent text itself, but the additional inference and context processing induced by that communication.
Original Article
View Cached Full Text

Cached at: 09/17/26, 09:37 AM

# Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum
Source: [https://arxiv.org/html/2609.18283](https://arxiv.org/html/2609.18283)
Vid HanželTim Strnad and Blaž Bertalani膆thanks:The authors are with the Department of Communication Systems, Jožef Stefan Institute, Jamova cesta 39, 1000 Ljubljana, Slovenia \(e\-mail: \{carolina\.fortuna, blaz\.bertalanic\}@ijs\.si\)\. Manuscript submitted to the IEEE Journal on Selected Areas in Communications \(JSAC\) Special Issue on Agentic AI for Intelligent Networks\.

###### Abstract

As telecommunication networks evolve toward autonomous 5G\-Advanced and 6G operations, agentic artificial intelligence \(AI\) workflows, where large language models \(LLMs\) execute multi\-step reasoning, invoke diagnostic tools, retrieve domain knowledge, and coordinate across agent teams, are increasingly embedded across the edge\-cloud continuum\. While the biological brain accomplishes complex cognition on an exceptionally modest metabolic power budget of approximately 20 W, contemporary LLMs are profoundly energy\- and memory\-intensive, making sustainable lifecycle orchestration a critical operational priority\. However, existing AI lifecycle metrics evaluate only isolated, single\-model inferences or overlook multi\-agent execution graphs entirely\. Consequently, network operators lack foundational models to determine whether distributed agent communication incurs meaningful energy costs and where across edge\-cloud tiers agent teams should physically reside\. To address this gap, we introduceagentic\-eCAL, generalizing the Energy Cost of AI Lifecycle \(eCAL\) metric to directed multi\-agent workflows by coupling a closed\-form two\-rate single\-call energy model \(compute\-bound prefill and memory\-bound decode\) with 7\-layer OSI data transport\. Grounded in hundreds of GPU benchmark configurations on NVIDIA A100 and H100, 16 open\-weight models and 8 orchestration topologies, we validate components of the metric and study workflow placement implications\. Our findings demonstrate that inter\-agent text transport incurs<0\.25%<0\.25\\%of workflow energy across 5G RAN, metro, and optical links\. Therefore in edge\-cloud agent placement the dominant energy cost of distribution is often not the transmission of inter\-agent text itself, but the additional inference and context processing induced by that communication\. Furthermore, on an ETSI ZSM\-aligned telco edge infrastructure benchmark, the evaluated multi\-agent workflows increase energy by up to23\.9×23\.9\\timesfor Qwen2\.5\-7B and4\.0×4\.0\\timesfor Qwen3\.5\-9B, without producing a consistent improvement in incident remediation success\.

###### Index Terms:

Agentic AI, energy consumption, AI lifecycle, eCAL, sustainable networks, edge computing, service placement, large language models, open\-weight models\.

## IIntroduction

The biological brain accomplishes multi\-step reasoning, contextual memory retrieval, and coordinated action on an exceptionally modest metabolic power budget of approximately 20 W\[[1](https://arxiv.org/html/2609.18283#bib.bib29)\]\. In stark contrast, contemporary artificial intelligence \(AI\) based on large language models \(LLMs\) is profoundly energy\- and memory\-inefficient: state\-of\-the\-art serving nodes draw hundreds of watts per accelerator, while per\-token generation and key\-value \(KV\) context retention rapidly saturate memory bandwidth and capacity\[[2](https://arxiv.org/html/2609.18283#bib.bib30),[3](https://arxiv.org/html/2609.18283#bib.bib8)\]\. Despite these physical costs, agentic AI, where LLMs reason across multi\-step loops, call external tools, retrieve domain knowledge, and coordinate as distributed teams, is increasingly embedded across the edge\-cloud continuum of telecommunication networks\. Because operators increasingly deploy*open\-weight*models distributed between user devices, edge sites, and the core, establishing a rigorous energy\-memory characterization of these workloads is both fundamental and urgent\.

Fig\. 1:Agentic eCAL framework and example multi\-agent workflow placement across the edge\-cloud continuum\.As conceptualized in Fig\.[1](https://arxiv.org/html/2609.18283#S1.F1), an agentic workflowW=\(𝒜,ℰ\)W=\(\\mathcal\{A\},\\mathcal\{E\}\)coordinates heterogeneous operations \(LLM inference withEcallE\_\{\\mathrm\{call\}\}energy, diagnostic tool invocations withEtoolE\_\{\\mathrm\{tool\}\}energy, and vector database retrievalsEretE\_\{\\mathrm\{ret\}\}energy\) across candidate network tiers\. In classical distributed edge computing, placement solves an offloading trade\-off between local computation and transmission energy \(EtxE\_\{\\mathrm\{tx\}\}\) across wireless \(5G RAN\) and optical bearers\. Yet, as we will show in this paper and anticipated in Fig\.[1](https://arxiv.org/html/2609.18283#S1.F1), agentic AI fundamentally overturns this premise: inter\-agent transmission is insignificant, accounts for less than0\.25%0\.25\\%of workflow energy \(Etx/Eprefill≈1/496E\_\{\\mathrm\{tx\}\}/E\_\{\\mathrm\{prefill\}\}\\approx 1/496for a 2,700\-token hand\-off over 5G\)\. Consequently, model size \(β​Nparams\\beta N\_\{\\mathrm\{params\}\}\) and dynamic resident context \(γ​ca\\gamma c\_\{a\}\) can dominate the energy implications of edge\-cloud placement\.

A placement problem is conventionally treated as an offloading trade between compute and communication energy\. Yet the energy cost of such AI workflows is poorly understood: existing metrics either stop at a single model inference or ignore the compute entirely\. The originaleCALmetric\[[4](https://arxiv.org/html/2609.18283#bib.bib1)\]captured the end\-to\-end energy of a*single*AI model’s lifecycle \(i\.e\. data collection, preprocessing, training, evaluation, and inference\), as a per\-bit quantity\[J/b\]\[\\mathrm\{J/b\}\], and showed that the more a trained model is used, the more energy\-efficient each inference becomes as the fixed development cost is amortised over a growing number of inferences\. However, this no longer describes how a growing class of network AI services are built\. Instead,*agentic*workflows take a large,*pre\-trained, open\-weight*LLM and orchestrate it typically as directed graph of: model reasoning over multiple steps, external tools calls, document retrieval, and cooperation with other agents as per Fig\.[1](https://arxiv.org/html/2609.18283#S1.F1)\. Because such workflows are natural candidates for distribution \(an agent close to the user for latency and privacy, an agent in the core for large data corpus access\), they inherit the placement question above, and this paper asks whether the classical answer still holds:*does distributing the agents of an LLM workflow across the network cost meaningful energy?*Three properties make the energy accounting fundamentally different from a single inference:

#### No bespoke training

The model is taken off the shelf; the analogue ofeCAL’s\[[4](https://arxiv.org/html/2609.18283#bib.bib1)\]development energy is the model’s*embodied*\(pre\-training\) energy, which is the most significant but shared across all of the model’s users\.

#### The unit of work is a workflow, not an inference

A single task triggers a graph of many LLM calls, whose prompts*carry the accumulating transcript*, so token volume, and hence energy, grows super\-linearly in the number of steps\.

#### Heterogeneous components

Beyond LLM inference, agentic execution incorporates external tool execution \(EtoolE\_\{\\mathrm\{tool\}\}\), vector database retrieval \(EretE\_\{\\mathrm\{ret\}\}\), and 7\-layer OSI transmission \(EtxE\_\{\\mathrm\{tx\}\}\), as depicted in Fig\.[1](https://arxiv.org/html/2609.18283#S1.F1)\.

Our contributions are:

- •We introduceagentic\-eCAL, a rigorous per\-bit energy framework for multi\-agent workflows executing across the edge\-cloud continuum, summing LLM inference, tool execution, vector retrieval, 7\-layer OSI transmission, and amortized embodied costs \(Sec\.[III](https://arxiv.org/html/2609.18283#S3)\)\.agentic\-eCALgeneralizeseCAL\[[4](https://arxiv.org/html/2609.18283#bib.bib1)\], reducing to its inference metric for one\-shot calls\.
- •We establish a physics\- and hardware\-grounded two\-rate closed\-form energy model for individual LLM calls \(Secs\.[III\-A](https://arxiv.org/html/2609.18283#S3.SS1)and[III\-B](https://arxiv.org/html/2609.18283#S3.SS2)\), decomposing inference into compute\-bound prefill and memory\-bandwidth\-bound decode regimes, parameterized by serving batch amortization and context\-dependent KV\-cache traffic\.
- •We validate the framework against extensive GPU measurements on NVIDIA A100 and H100 with vLLM continuous batching \(R2≥0\.99R^\{2\}\\geq 0\.99, MAPE≈10%\\approx 10\\%\)\. Systemic characterization reveals that: \(i\) history\-carrying loops exhibit super\-linear prompt\-prefill energy scaling; \(ii\) inter\-agent transmission across 5G RAN, metro, and optical links accounts for<0\.25%<0\.25\\%of workflow energy; and \(iii\) on an ETSI ZSM\-aligned telco edge infrastructure benchmark, compared with the single\-agent baseline, the evaluated multi\-agent topologies increase energy consumption by up to23\.9×23\.9\\timesfor Qwen2\.5\-7B and4\.0×4\.0\\timesfor Qwen3\.5\-9B, without producing a consistent improvement in incident remediation success rates111The code will be open sourced upon acceptance\. https://github\.com/sensorlab/a\-eCAL\.

Paper Organization\.The remainder of this paper is structured as follows\. Sec\.[II](https://arxiv.org/html/2609.18283#S2)surveys related work across AI lifecycle metrics, LLM serving energy, and multi\-agent systems\. Sec\.[III](https://arxiv.org/html/2609.18283#S3)formalizes theagentic\-eCALmetric and system model, deriving the closed\-form two\-rate single\-call energy model, super\-linear context accumulation, and the 7\-layer OSI transmission formulation\. Sec\.[IV](https://arxiv.org/html/2609.18283#S4)presents empirical GPU validation and scaling laws across reasoning depth, agent count, and batching\. Sec\.[V](https://arxiv.org/html/2609.18283#S5)evaluates multi\-agent placement feasibility, accelerator memory boundaries, and transmission energy across network tiers\. Sec\.[VI](https://arxiv.org/html/2609.18283#S6)evaluates the framework on an ETSI ZSM\-aligned telco edge infrastructure benchmark\. Sec\.[VII](https://arxiv.org/html/2609.18283#S7)discusses limitations and future extensions, and Sec\.[VIII](https://arxiv.org/html/2609.18283#S8)concludes the paper\.

## IIRelated Work

Lifecycle energy metrics\.eCAL\[[4](https://arxiv.org/html/2609.18283#bib.bib1)\]introduced a per\-bit, component\-resolved view of AI energy spanning the OSI and cloud\-computing reference architectures\. We adopt the same per\-bit framing and the same transmission/virtualization machinery, and extend the*development*and*inference*components to agentic LLM workflows\.

LLM inference compute and energy\.The compute of a decoder\-only transformer is governed by a well\-established FLOP accounting: a forward \(inference\) pass costs≈2​Nparams\\approx 2N\_\{\\mathrm\{params\}\}FLOPs per token and a full training step≈6​Nparams\\approx 6N\_\{\\mathrm\{params\}\}\(the backward pass being roughly twice the forward\), with a context\-dependent attention term that is negligible until the context length exceeds≈8​d\\approx 8d\[[5](https://arxiv.org/html/2609.18283#bib.bib4),[6](https://arxiv.org/html/2609.18283#bib.bib5)\]\. On the measurement side, Samsi*et al\.*\[[7](https://arxiv.org/html/2609.18283#bib.bib6)\]report33–44J per output token for LLaMA 65B on V100/A100, the ML\.ENERGY benchmark\[[8](https://arxiv.org/html/2609.18283#bib.bib7)\]shows per\-generation energy falling∼3\.5×\\sim\\\!3\.5\\timesfrom batch 4 to batch 64 on H100 and warns that estimating energy from nameplate TDP overestimates measured energy by up to4\.1×4\.1\\times, and frontier\-scale serving reaches a median0\.310\.31Wh per query\[[3](https://arxiv.org/html/2609.18283#bib.bib8)\]\. We adopt theβ​Nparams\\beta N\_\{\\mathrm\{params\}\}withβ=2\\beta=2accounting and calibrate a utilization factor to these measurements rather than to peak FLOP/s\.

Embodied cost of open\-weight models\.Pre\-training energy is documented in primary model cards and peer\-reviewed audits: Llama 2 required3\.33\.3M A100\-hours \(539539tCO2eq\) and Llama 37\.77\.7M H100\-hours \(22902290tCO2eq, of which1\.31\.3M h for the 8B and6\.46\.4M h for the 70B\)\[[9](https://arxiv.org/html/2609.18283#bib.bib11),[10](https://arxiv.org/html/2609.18283#bib.bib12)\]; Luccioni*et al\.*\[[11](https://arxiv.org/html/2609.18283#bib.bib9)\]decompose BLOOM\-176B’s50\.550\.5tCO2eq lifecycle into∼22%\\sim\\\!22\\%embodied and∼78%\\sim\\\!78\\%operational\. Crucially, inference only overtakes training at fleet scale, with parity after∼205\\sim\\\!205M \(BLOOMz\-560M\) to∼593\\sim\\\!593M \(BLOOMz\-7B\) inferences\[[12](https://arxiv.org/html/2609.18283#bib.bib10)\], which motivates amortizing embodied energy over deployment volume\.

Agentic, reasoning, and RAG systems\.The defining energy property of agentic AI is*token amplification*: reasoning models emit∼18×\\sim\\\!18\\timesmore tokens than standard models on the same task \(e\.g\. Phi\-4\-reasoning\-plus6,7806\{,\}780vs Phi\-4378378tokens\), sometimes at lower accuracy\[[13](https://arxiv.org/html/2609.18283#bib.bib13)\], and test\-time scaling with∼15×\\sim\\\!15\\timestokens raises per\-query energy∼13×\\sim\\\!13\\times\[[3](https://arxiv.org/html/2609.18283#bib.bib8)\]\. Because energy scales near\-linearly with generated tokens, this directly multiplies the per\-token term\. Agent composition compounds it: single agents use∼4×\\sim\\\!4\\timesand multi\-agent systems∼15×\\sim\\\!15\\timesmore tokens than a chat call\[[14](https://arxiv.org/html/2609.18283#bib.bib14)\], and multi\-agent debate can reach3232–158×158\\timesa single call before sparse\-communication methods recover much of it\[[15](https://arxiv.org/html/2609.18283#bib.bib15)\]\. In retrieval\-augmented generation the injected context dominates: a∼100\\sim\\\!100\-token request plus∼1000\\sim\\\!1000retrieved tokens makes the augmented request\>10×\>\\\!10\\timescostlier\[[16](https://arxiv.org/html/2609.18283#bib.bib16)\], while prefix/KV\-cache reuse of that context cuts time\-to\-first\-token up to4×4\\times\[[16](https://arxiv.org/html/2609.18283#bib.bib16)\]and tuned RAG pipelines save up to60%60\\%\[[17](https://arxiv.org/html/2609.18283#bib.bib17)\]\. Model choice is a complementary lever: measured per\-token energy rises∼18×\\sim\\\!18\\timesfrom Llama\-3 1B to 70B, so a 1B model costs roughly5%5\\%of a 70B model per generated token\[[18](https://arxiv.org/html/2609.18283#bib.bib19)\], as is sparsity: Mixture\-of\-Experts energy tracks*active*not total parameters, e\.g\. a 30B\-A3B model uses3\.56×3\.56\\timesless energy/token than a dense 32B\[[19](https://arxiv.org/html/2609.18283#bib.bib18)\]\. Crucially, all agentic/tool evidence is in*token*units, and no published study measures the joules of a full open\-weight agentic run, so we bridge tokens to energy via the per\-token model and flag the tool/orchestration terms as analytical\.

Agentic AI in network management\.Modern agentic AI has the potential to deliver the autonomy that the zero\-touch literature specifies: ETSI ZSM defines closed management loops over network domains\[[20](https://arxiv.org/html/2609.18283#bib.bib31)\], and O\-RAN places those loops on a virtualized cloud substrate spanning far\-edge, edge and regional sites\[[21](https://arxiv.org/html/2609.18283#bib.bib32)\]\. As LLM\-based agents are adopted as the reasoning element of such loops, their energy becomes an operational quantity rather than an academic one\. Energy accounting for the AI/ML workflow inside O\-RAN has so far addressed*single\-model*pipelines\[[22](https://arxiv.org/html/2609.18283#bib.bib2),[23](https://arxiv.org/html/2609.18283#bib.bib3)\]; we extend it to the multi\-call, multi\-agent case\.

Placement and orchestration\.Distributing computation across network tiers is classically an offloading trade, in which compute energy saved by moving work outward is repaid in transmission\[[24](https://arxiv.org/html/2609.18283#bib.bib33)\]\. Our measurements indicate that this trade does not arise for text hand\-offs between agents: what a hand\-off carries is small in bits but large in the prefill it triggers on arrival, an asymmetry consistent with the load\-independence of router power\[[25](https://arxiv.org/html/2609.18283#bib.bib25)\]\. Orchestration architectures for teams of agents have been analyzed from the perspective of collaboration modes, roles, models, or workflows per query\[[26](https://arxiv.org/html/2609.18283#bib.bib34),[27](https://arxiv.org/html/2609.18283#bib.bib35),[28](https://arxiv.org/html/2609.18283#bib.bib36)\], and communication\-efficient multi\-agent reasoning has been pursued by sparsifying which agents address which\[[15](https://arxiv.org/html/2609.18283#bib.bib15)\]; we evaluate the graph instead for the traffic its edge count implies and for the memory its retained context occupies\.

TABLE I:Summary of mathematical notation and parameters\.
## IIIAgentic Workflows and the Agentic\-eCAL Metric

We define an agentic workflow as a directed graphW=\(𝒜,ℰ\)W=\(\\mathcal\{A\},\\mathcal\{E\}\)ofNNagents joined byE=\|ℰ\|E=\|\\mathcal\{E\}\|hand\-offs and executed inKKsteps; stepkkis run by agenta⁡\(k\)∈𝒜a\(k\)\\in\\mathcal\{A\}and is one LLM call ofpin\(k\)p\_\{\\mathrm\{in\}\}^\{\(k\)\}prompt andpout\(k\)p\_\{\\mathrm\{out\}\}^\{\(k\)\}generated tokens, optionally with a retrieval before it and a tool invocation after\. As per the conceptual illustration in Fig\.[1](https://arxiv.org/html/2609.18283#S1.F1), it is hosted on a set𝒟\\mathcal\{D\}of candidate deployment sites spanning various network segments from the user device to the core: siteuuoffers memoryMuM\_\{u\}for weights and cache and sustains a serving batchbub\_\{u\}, and sitesu,vu,vare joined by a bearer of energy intensityεu​v\\varepsilon\_\{uv\}\[J/b\], withεu​u=0\\varepsilon\_\{uu\}=0for co\-located agents\. A deployment of a workflow is then the triple𝒞=\(x,π,ℰ\)\\mathcal\{C\}=\(x,\\pi,\\mathcal\{E\}\): the*placement*xa​u∈\{0,1\}x\_\{au\}\\in\\\{0,1\\\}, one if agentaaruns at siteuu, with∑uxa​u=1\\sum\_\{u\}x\_\{au\}=1; the*hand\-off protocol*π∈𝒫\\pi\\in\\mathcal\{P\}, which fixes the payloadθπ​\(a,a′\)\\theta\_\{\\pi\}\(a,a^\{\\prime\}\)in bits each edge carries, whether the accumulated transcript, an increment, a bounded summary or the key\-value cache itself; and the*architecture*ℰ\\mathcal\{E\}, which fixes how many edges there are\.

Agent role types vary across the sources this paper draws on, so we keep each source’s names and classify agents by function instead in four classes as follows\.*Generate*agents produce a candidate or a plan, and are charged principally forpoutp\_\{\\mathrm\{out\}\}\.*Transform*agents refine a single line of work\.*Ground*agents acquire state external toWW, through retrieval or a tool, and so import intopinp\_\{\\mathrm\{in\}\}a segment no agent generated\.*Aggregate*agents reduce a set of competing candidates and read all of it, so theirpinp\_\{\\mathrm\{in\}\}grows with the size of that set\.

We define a placement configuration as admissible only if each siteu∈𝒟u\\in\\mathcal\{D\}can hold the weights of the model it runs together with the resident context of every agent it hosts,

β​Nparams⏟static model weights\+γ​∑a∈𝒜xa​u​ca⏟dynamic resident context≤Mu,∀u∈𝒟,\\underbrace\{\\beta N\_\{\\mathrm\{params\}\}\}\_\{\\text\{static model weights\}\}\\;\+\\;\\underbrace\{\\gamma\\sum\_\{a\\in\\mathcal\{A\}\}x\_\{au\}c\_\{a\}\}\_\{\\text\{dynamic resident context\}\}\\;\\leq\\;M\_\{u\},\\qquad\\forall u\\in\\mathcal\{D\},\(1\)where the first term is the static memory footprint of model weights loaded once per site atβ\\betabytes per parameter, and the second term is the dynamic footprint of the key\-value \(KV\) cache, requiringγ\\gammabytes per context tokencac\_\{a\}for each agentaahosted at siteuu\(xa​u=1x\_\{au\}=1\)\. While model weights load once, KV caches scale with agent concurrency and retained context length\.

To quantify the energy of an agentic workflow, a new metric is required\. For a workflowWW, we generalize theeCALmetric\[[4](https://arxiv.org/html/2609.18283#bib.bib1)\], by definingagentic\-eCALas the ratio of total energy consumed to the useful application\-level data produced:

agentic\-eCAL=EW\+γe​\(Eemb\+Eemb,ret\)Buseful\[J/b\],\\textit\{agentic\-eCAL\}=\\frac\{E\_\{W\}\+\\gamma\_\{\\mathrm\{e\}\}\(E\_\{\\mathrm\{emb\}\}\+E\_\{\\mathrm\{emb,ret\}\}\)\}\{B\_\{\\mathrm\{useful\}\}\}\\quad\[\\mathrm\{J/b\}\],\(2\)whereEWE\_\{W\}is the operational energy of the workflow,EembE\_\{\\mathrm\{emb\}\}is the embodied \(pre\-training\) energy of the open\-weight model,Eemb,retE\_\{\\mathrm\{emb,ret\}\}is the embodied \(pre\-training\) energy of the open\-weight model used when the optional retrieval step is included,γe∈\[0,1\]\\gamma\_\{\\mathrm\{e\}\}\\in\[0,1\]is the fraction of that embodied energy attributed to one workflow invocation, andBuseful=f​ToutB\_\{\\mathrm\{useful\}\}=f\\,T\_\{\\mathrm\{out\}\}is the useful output measured in bits \(ToutT\_\{\\mathrm\{out\}\}useful output tokens atffbits/token\)\. Table[I](https://arxiv.org/html/2609.18283#S2.T1)summarizes all the notations in the paper\.

#### The operational energy of a workflowEWE\_\{W\}

sums the energy of every component, scaled by a virtualization and system orchestration overheadγv\\gamma\_\{\\mathrm\{v\}\}\(the same factoreCALuses for cloud platform overheads, Eq\. \(30\) of\[[4](https://arxiv.org/html/2609.18283#bib.bib1)\]\):

EW=\(1\+γv\)​\[∑k=1KEcall\(k\)\+∑kEtool\(k\)\+∑kEret\(k\)\+Etx\]E\_\{W\}=\(1\+\\gamma\_\{\\mathrm\{v\}\}\)\\\!\\left\[\\sum\_\{k=1\}^\{K\}E\_\{\\mathrm\{call\}\}^\{\(k\)\}\+\\sum\_\{k\}E\_\{\\mathrm\{tool\}\}^\{\(k\)\}\+\\sum\_\{k\}E\_\{\\mathrm\{ret\}\}^\{\(k\)\}\+E\_\{\\mathrm\{tx\}\}\\right\]\(3\)whereEcall\(k\)E\_\{\\mathrm\{call\}\}^\{\(k\)\}is the energy of the LLM call at step k,Etool\(k\)E\_\{\\mathrm\{tool\}\}^\{\(k\)\}the execution energy of any tool that step invokes,Eret\(k\)E\_\{\\mathrm\{ret\}\}^\{\(k\)\}the cost of any retrieval preceding it \(query embedding plus index search\), andEtxE\_\{\\mathrm\{tx\}\}the transmission energy of the messages received from other agents\. AtK=1\\\!K\\\!=\\\!1with no tools or retrieval and with co\-located agents,EW→\(1\+γv\)​EcallE\_\{W\}\\\!\\to\\\!\(1\{\+\}\\gamma\_\{\\mathrm\{v\}\}\)E\_\{\\mathrm\{call\}\}and Eq\.[2](https://arxiv.org/html/2609.18283#S3.E2)reduces toeCAL’s energy\-per\-bit of a single inference, with the denominator specialisingeCAL’s*manipulated*application level data bits to the workflow’s*useful output*bits\.agentic\-eCALis thus a strict generalization ofeCAL, adding the multi\-call graph on top of the single\-inference term\.

#### The embodied energy of workflow

If a model servesGGagentic workflow invocations over its deployed lifetime, the embodied energy attributable to one invocation isγe​\(Eemb\+Eemb,ret\)\\gamma\_\{\\mathrm\{e\}\}\(E\_\{\\mathrm\{emb\}\}\+E\_\{\\mathrm\{emb,ret\}\}\)from Eq\.[2](https://arxiv.org/html/2609.18283#S3.E2)whereγe=1/G\\gamma\_\{\\mathrm\{e\}\}=1/Grefers to the fraction of the energy from training the LLM models used for answering the callsEembE\_\{\\mathrm\{emb\}\}and for retrieval when this is presentEemb,retE\_\{\\mathrm\{emb,ret\}\}\. The more workflows an LLM powers, the lower the fixed training energy attributable to that workflow\.

The existingeCAL\[[4](https://arxiv.org/html/2609.18283#bib.bib1)\]amortizes a model’s one\-time development energyEDE\_\{\\mathrm\{D\}\}over the number of inferences served\. For an open\-weight model the development cost is the*pre\-training*energyEembE\_\{\\mathrm\{emb\}\}, which is incurred once by the model producer but shared across*all*downstream deployments\. As witheCAL’s amortization\-over\-inferences result, this share vanishes at scale, so for widely served open\-weight models the operational termEWE\_\{W\}dominatesagentic\-eCAL\. The crossover point \(how many invocations are needed before operation overtakes embodied cost\) depends mostly onEembE\_\{\\mathrm\{emb\}\}\.

### III\-AEnergy of a Single LLM Call:EcallE\_\{\\mathrm\{call\}\}

An LLM call executes in two distinct phases, yielding a*two\-rate*energy model:

Ecall​\(pin,pout,b\)≈cpre​pin\+cdec​\(b\)​pout,E\_\{\\mathrm\{call\}\}\(p\_\{\\mathrm\{in\}\},p\_\{\\mathrm\{out\}\};b\)\\;\\approx\\;c\_\{\\mathrm\{pre\}\}\\,p\_\{\\mathrm\{in\}\}\\;\+\\;c\_\{\\mathrm\{dec\}\}\(b\)\\,p\_\{\\mathrm\{out\}\},\(4\)wherecprec\_\{\\mathrm\{pre\}\}is the per\-prefill\-token energy andcdec​\(b\)c\_\{\\mathrm\{dec\}\}\(b\)is the per\-decode\-token energy at serving batch sizebb\.

Prefillprocesses allpinp\_\{\\mathrm\{in\}\}prompt tokens in parallel\. It is compute\-bound at near\-peak utilization \(ηpre\\eta\_\{\\mathrm\{pre\}\}\) and, being a single batched matrix multiplication, is essentially independent ofbb:

cpre≈Nparams​β​P/\(ηpre​Π\)c\_\{\\mathrm\{pre\}\}\\approx N\_\{\\mathrm\{params\}\}\\beta\\,P/\(\\eta\_\{\\mathrm\{pre\}\}\\Pi\)\(5\)Decodeemitspoutp\_\{\\mathrm\{out\}\}tokens sequentially\. It is memory bandwidth bound accounting for7777–91%91\\%of inference time\[[29](https://arxiv.org/html/2609.18283#bib.bib21)\]as weights and the growing KV\-cache stream from High Bandwidth Memory \(HBM\) while arithmetic units sit largely idle\. Because weights are read once per batch, this cost amortizes over thebbconcurrently\-decoding sequences:

cdec​\(b\)≈PBHBM​\(Nparams​βb⏟weights/batch\+γ​c¯⏟KV\-cache\),c\_\{\\mathrm\{dec\}\}\(b\)\\;\\approx\\;\\frac\{P\}\{B\_\{\\mathrm\{HBM\}\}\}\\\!\\left\(\\underbrace\{\\frac\{N\_\{\\mathrm\{params\}\}\\beta\}\{b\}\}\_\{\\text\{weights\}/\\text\{batch\}\}\\;\+\\;\\underbrace\{\\gamma\\,\\bar\{c\}\}\_\{\\text\{KV\-cache\}\}\\right\),\(6\)whereBHBMB\_\{\\mathrm\{HBM\}\}is HBM bandwidth,222In practice a serving stack sustains only about70%70\\%of the theoretical peak, so Eq\.[6](https://arxiv.org/html/2609.18283#S3.E6)at the nameplateBHBMB\_\{\\mathrm\{HBM\}\}is a lower bound oncdecc\_\{\\mathrm\{dec\}\}\. Model\-bandwidth utilization,\(achieved bandwidth\)/\(peak bandwidth\)\(\\text\{achieved bandwidth\}\)/\(\\text\{peak bandwidth\}\), is measured at72%72\\%for Llama\-7B in fp16 on an A100\-80GB, where the ceiling is itself under85%85\\%because “even just copying memory struggles to break” it\[[30](https://arxiv.org/html/2609.18283#bib.bib27)\], and at60%60\\%and55%55\\%on2×2\{\\times\}H100\-80GB and4×4\{\\times\}A100\-40GB\[[31](https://arxiv.org/html/2609.18283#bib.bib28)\]\. All three are single\-stream figures; utilization falls further asbbgrows and attention over the accumulated cache displaces the weight read\.β\\betais bytes per parameter, andγ​c¯\\gamma\\bar\{c\}is the per\-token KV\-cache traffic at mean contextc¯\\bar\{c\}\.

This two\-rate formulation reveals an important crossover\. Mathematically, a call becomes prefill\-dominated whencpre​pin\>cdec​\(b\)​poutc\_\{\\mathrm\{pre\}\}\\,p\_\{\\mathrm\{in\}\}\>c\_\{\\mathrm\{dec\}\}\(b\)\\,p\_\{\\mathrm\{out\}\}, which occurs when the token ratiopin/pout\>cdec​\(b\)/cprep\_\{\\mathrm\{in\}\}/p\_\{\\mathrm\{out\}\}\>c\_\{\\mathrm\{dec\}\}\(b\)/c\_\{\\mathrm\{pre\}\}\. At small batches, weight\-loading dominates \(cdec≫cprec\_\{\\mathrm\{dec\}\}\\gg c\_\{\\mathrm\{pre\}\}\) and energy tracks the linear decode workload\. However, asbbscales, the weight term in Eq\.[6](https://arxiv.org/html/2609.18283#S3.E6)vanishes andcdec​\(b\)c\_\{\\mathrm\{dec\}\}\(b\)decreases toward the KV\-cache floor, rapidly driving down the right side of the inequality\. This dynamic fundamentally shapes agentic AI: because transcripts accumulate across steps, production\-batched workflows inevitably cross this threshold into the prefill\-dominated regime\. Consequently, the workflow’s energy scales super\-linearly, whereas the exact same workflow under single\-stream serving would remain linear\.

### III\-BEnergy of the Workflow LLM Calls:Ecall,wE\_\{\\mathrm\{call,w\}\}

In a history\-carrying loop the prompt of stepkkcarries the system/tool\-schema promptpsysp\_\{\\mathrm\{sys\}\}, any newly introduced tokens, retrieved context, and the accumulated transcripthk=∑j<k\(pout\(j\)\+o\(j\)\)h\_\{k\}=\\sum\_\{j<k\}\\\!\\big\(p\_\{\\mathrm\{out\}\}^\{\(j\)\}\+o^\{\(j\)\}\\big\)\(o\(j\)o^\{\(j\)\}: tool observations\)\. With a fixed emissionssper step,pin\(k\)≈psys\+s​kp\_\{\\mathrm\{in\}\}^\{\(k\)\}\\\!\\approx\\\!p\_\{\\mathrm\{sys\}\}\+s\\,k, so the*token*volume is exactly quadratic while the generated output is linear:

∑k=1Kpin\(k\)=s2​K2\+O⁡\(K\),∑k=1Kpout\(k\)=s​K\.\\sum\_\{k=1\}^\{K\}p\_\{\\mathrm\{in\}\}^\{\(k\)\}=\\tfrac\{s\}\{2\}K^\{2\}\+O\(K\),\\qquad\\sum\_\{k=1\}^\{K\}p\_\{\\mathrm\{out\}\}^\{\(k\)\}=sK\.\(7\)Substituting into Eq\.[4](https://arxiv.org/html/2609.18283#S3.E4)through the two\-rate model, the operational energy is

Ec​a​l​l,w∝cpre​\(s2​K2\)\+cdec​\(b\)​s​K,E\_\{call,w\}\\;\\propto\\;c\_\{\\mathrm\{pre\}\}\\,\\big\(\\tfrac\{s\}\{2\}K^\{2\}\\big\)\\;\+\\;c\_\{\\mathrm\{dec\}\}\(b\)\\,sK,\(8\)soagentic\-eCALscales asKaK^\{a\}with a*batch\-set*exponent≤a≤21\\\!\\leq\\\!a\\\!\\leq\\\!2: at small batchcdec​\(b\)≫cprec\_\{\\mathrm\{dec\}\}\(b\)\\\!\\gg\\\!c\_\{\\mathrm\{pre\}\}and the linear decode term dominates \(a→1a\\\!\\to\\\!1\); asbbgrows,cdec​\(b\)→0c\_\{\\mathrm\{dec\}\}\(b\)\\\!\\to\\\!0and the quadratic prefill term takes over \(a→2a\\\!\\to\\\!2\)\. Eq\.[7](https://arxiv.org/html/2609.18283#S3.E7)shows that the*context*grows quadratically unconditionally, but as Eq\. \([8](https://arxiv.org/html/2609.18283#S3.E8)\) shows, the*energy of a sequence of calls*inherits that growth only when prefill\-dominated\. This quadratic scaling governs stateless agent invocations across distributed network hosts or microservices without shared memory, where each step must re\-prefill the accumulated prompt history\. While stateful single\-host serving engines leverage Automatic Prefix Caching \(APC\) to reduce redundant prompt prefill to linear scaling O\(K\) for co\-located calls, distributed edge\-cloud agent pipelines that dispatch steps across distinct edge servers or microservices may incur full transcript re\-prefill penalties\.

### III\-CEnergy of Tool Calls:EtoolE\_\{\\mathrm\{tool\}\}

A tool invocation contributes its own execution energyEtool\(k\)E\_\{\\mathrm\{tool\}\}^\{\(k\)\}and injectso\(k\)o^\{\(k\)\}observation tokens into the transcript\. Non\-LLM tools that fetch data over the network are priced by the OSI data\-collection modelED​CE\_\{DC\}ofeCAL\[[4](https://arxiv.org/html/2609.18283#bib.bib1)\]\); local compute tools by their FLOPs and the same energy mapping as Eq\.[5](https://arxiv.org/html/2609.18283#S3.E5)\.

Fig\. 2:Empirical validation of the two\-rate model \(Qwen2\.5\-7B and Llama\-3\.1\-8B; A100, vLLM;270270configurations per model,22–3030agents, serving batchb∈\[2,256\]b\\\!\\in\\\!\[2,256\]\)\. \(a\) Predicted vs\. measured per\-query GPU energy \(R2\>0\.99R^\{2\}\\\!\>\\\!0\.99, mean absolute error∼10%\{\\sim\}10\\%\); \(b\) The two rates recovered by regressing energy on\(pin,pout\)\(p\_\{\\mathrm\{in\}\},p\_\{\\mathrm\{out\}\}\): the prefill ratecprec\_\{\\mathrm\{pre\}\}is batch\-independent, while the decode rate collapses asadec/ba\_\{\\mathrm\{dec\}\}/btoward a small floor \(Eq\.[6](https://arxiv.org/html/2609.18283#S3.E6)\)\. Llama’s higher floor keeps it partly decode\-bound at large batch\.
### III\-DEnergy of Retrieval:EretE\_\{\\mathrm\{ret\}\}

A retrieval step embeds the query with an embedding model \(≈β​Nemb​pq\\approx\\beta N\_\{\\mathrm\{emb\}\}\\,p\_\{q\}FLOPs forpqp\_\{q\}query tokens\), searches a vector index ofMvecM\_\{\\mathrm\{vec\}\}vectors of dimensiondembd\_\{\\mathrm\{emb\}\}\(cost∝demb​Mvec\\propto d\_\{\\mathrm\{emb\}\}M\_\{\\mathrm\{vec\}\}for exact search, or∝demb​log⁡Mvec\\propto d\_\{\\mathrm\{emb\}\}\\log M\_\{\\mathrm\{vec\}\}for graph\-based approximate search\)\[[32](https://arxiv.org/html/2609.18283#bib.bib22)\], and injectskrk\_\{r\}chunks ofLrL\_\{r\}tokens that enlarge the next call’s prompt\. By mapping the embedding energy via Eq\.[5](https://arxiv.org/html/2609.18283#S3.E5)and index search operations via memory bandwidth yields:

Eret\(k\)=β​Nemb​pq​Pembηemb​Πemb\+αidx​demb​log⁡\(Mvec\),E\_\{\\mathrm\{ret\}\}^\{\(k\)\}=\\frac\{\\beta N\_\{\\mathrm\{emb\}\}\\,p\_\{q\}\\,P\_\{\\mathrm\{emb\}\}\}\{\\eta\_\{\\mathrm\{emb\}\}\\,\\Pi\_\{\\mathrm\{emb\}\}\}\\;\+\\;\\alpha\_\{\\mathrm\{idx\}\}\\,d\_\{\\mathrm\{emb\}\}\\log\(M\_\{\\mathrm\{vec\}\}\),\(9\)whereαidx\\alpha\_\{\\mathrm\{idx\}\}characterizes index distance calculation energy \[J/op\]\. The added context tokens are captured by Eqs\.[7](https://arxiv.org/html/2609.18283#S3.E7)\. Empirically, growing the context from22K to1010K tokens raises measured energy per token∼3×\\sim\\\!3\\timesfor a7070B model\[[18](https://arxiv.org/html/2609.18283#bib.bib19),[33](https://arxiv.org/html/2609.18283#bib.bib20)\]\.

### III\-EEnergy of Transmission for Multi\-agent Communication:EtxE\_\{\\mathrm\{tx\}\}

When agents are co\-located on the same physical host, the hand\-off cost is strictly confined to local memory transfers and additional prompt tokens\. When agents are distributed across network tiers \(e\.g\., edge devices, base station MEC servers, metro aggregation sites, or central cloud datacenters\), inter\-agent messages incur data communication energy\. In Eq\.[3](https://arxiv.org/html/2609.18283#S3.E3), the total transmission energy isEtx=∑\(u,v\)∈ℰBpayload,u​v⋅εu​vE\_\{\\mathrm\{tx\}\}=\\sum\_\{\(u,v\)\\in\\mathcal\{E\}\}B\_\{\\mathrm\{payload\},uv\}\\cdot\\varepsilon\_\{uv\}, whereBpayload,u​vB\_\{\\mathrm\{payload\},uv\}is the application payload volume in bits along directed edge\(u,v\)\(u,v\), andεu​v\\varepsilon\_\{uv\}is the effective end\-to\-end transport energy intensity \[J/bit\]\.

Following the model in the eCAL\[[4](https://arxiv.org/html/2609.18283#bib.bib1)\]framework, the actual bits transmitted over the network forBpayload,u​vB\_\{\\mathrm\{payload\},uv\}can be computed with:

BT,p\\displaystyle B\_\{\\mathrm\{T\},p\}\[b\]=⌈BT,LOSI,p∏l=1LOSI\(R​RDP,l​\(1\+O​HDP,l\)CLOSE⏟contribution from DP\\displaystyle~\[b\]=\\lceil B\_\{\\mathrm\{T\},L\_\{\\mathrm\{OSI\}\},p\}\\prod\_\{l=1\}^\{L\_\{\\mathrm\{OSI\}\}\}\\underbrace\{\\left\(RR\_\{\\mathrm\{DP\},l\}\(1\+OH\_\{\\mathrm\{DP\},l\}\)\\right\.\}\_\{\\text\{contribution from DP\}\}\+OPENR​RCP,l​γl​O​HCP,l\)⏟contribution from CP⌉,whereLOSI=7,\\displaystyle\+\\underbrace\{\\left\.RR\_\{\\mathrm\{CP\},l\}\\gamma\_\{l\}OH\_\{\\mathrm\{CP\},l\}\\right\)\}\_\{\\text\{contribution from CP\}\}\\rceil,~\\mathrm\{where\}~L\_\{\\mathrm\{OSI\}\}=7,\(10\)whereO​HDP,lOH\_\{\\mathrm\{DP\},l\}is the protocol overhead ratio at layerll\(encompassing L2 MAC framing, L3 IP routing, L4 TCP headers, L6 TLS 1\.3 encryption, and L7 HTTP/2 or gRPC protobuf serialization, typically adding100%100\\%to200%200\\%overhead\), andR​RDP,lRR\_\{\\mathrm\{DP\},l\}is the retransmission ratio\. For wireless cellular links \(e\.g\., 5G NR\), channel fading and interference induce Hybrid Automatic Repeat reQuest \(HARQ\) retransmissions, yieldingR​RDP≈1\.5RR\_\{\\mathrm\{DP\}\}\\approx 1\.5to2\.02\.0under cell\-edge or high\-mobility channel conditions\.

The same framework establishes the total energyEDC,pE\_\{\\mathrm\{DC\},p\}required to transport payloadBpayloadB\_\{\\mathrm\{payload\}\}across physical linkppsums RF transmission/reception with protocol execution:

EDC,p=\\displaystyle E\_\{\\mathrm\{DC\},p\}=PT,pRT,p​BT,p⏟PHY transmit​ET,p\+PR,pRR,p​BT,p⏟PHY receive​ER,p\\displaystyle\\underbrace\{\\frac\{P\_\{\\mathrm\{T\},p\}\}\{R\_\{\\mathrm\{T\},p\}\}B\_\{\\mathrm\{T\},p\}\}\_\{\\text\{PHY transmit \}E\_\{\\mathrm\{T\},p\}\}\+\\underbrace\{\\frac\{P\_\{\\mathrm\{R\},p\}\}\{R\_\{\\mathrm\{R\},p\}\}B\_\{\\mathrm\{T\},p\}\}\_\{\\text\{PHY receive \}E\_\{\\mathrm\{R\},p\}\}\+∑l=2LOSI\(BT,l,p⋅Ndev,l,p⋅Pdev,p⏟L2 to L7 protocol processing at end device\\displaystyle\+\\sum\_\{l=2\}^\{L\_\{\\mathrm\{OSI\}\}\}\\Big\(\\underbrace\{B\_\{\\mathrm\{T\},l,p\}\\cdot N\_\{\\mathrm\{dev\},l,p\}\\cdot P\_\{\\mathrm\{dev\},p\}\}\_\{\\text\{L2 to L7 protocol processing at end device\}\}OPEN\+BT,l,p⋅Ngw,l,p⋅Pgw,p⏟L2 to L7 protocol processing at gateway\),\\displaystyle\\qquad\\quad\+\\underbrace\{B\_\{\\mathrm\{T\},l,p\}\\cdot N\_\{\\mathrm\{gw\},l,p\}\\cdot P\_\{\\mathrm\{gw\},p\}\}\_\{\\text\{L2 to L7 protocol processing at gateway\}\}\\Big\),\(11\)wherePT,pP\_\{\\mathrm\{T\},p\}andPR,pP\_\{\\mathrm\{R\},p\}are the physical transceiver powers \[W\],RT,pR\_\{\\mathrm\{T\},p\}andRR,pR\_\{\\mathrm\{R\},p\}the operational data rates \[b/s\],Ndev,l,pN\_\{\\mathrm\{dev\},l,p\}andNgw,l,pN\_\{\\mathrm\{gw\},l,p\}the CPU clock cycles per bit for protocol processing at device and gateway, andPdev,p,Pgw,pP\_\{\\mathrm\{dev\},p\},P\_\{\\mathrm\{gw\},p\}the computational energy efficiencies \[J/cycle\]\.

TABLE II:Model vs\. measured per\-output\-token energy, each row evaluated at the serving batch its source reports\.ModelHWbbEcallE\_\{\\mathrm\{call\}\}measuredLlama\-3 70BH100 NVL1281280\.450\.39\[[18](https://arxiv.org/html/2609.18283#bib.bib19)\]Llama\-3 8BH100 NVL1281280\.0250\.07\[[18](https://arxiv.org/html/2609.18283#bib.bib19)\]Qwen2\.5\-7BH100 NVL1281280\.0210\.07\[[18](https://arxiv.org/html/2609.18283#bib.bib19)\]Units: J per output token; model and measurement both GPU\-only\.Measured: GPU\-only,00–22K context band, of\[[18](https://arxiv.org/html/2609.18283#bib.bib19)\]\(4×4\{\\times\}H100 node\)\.The 7B and 8B runs occupy one GPU, so∼38%\{\\sim\}38\\%of theirvalue is idle sibling draw \(≈0\.07\\approx\\\!0\.07\); the 70B is sharded over all four\.Their published0\.110\.11is therefore quoted here at the active device,0\.11×0\.620\.11\{\\times\}0\.62, the share of node GPU power that device drawsacross the7474released runs; the7070B needs no such correction\.

## IVEnergy and Memory Characterization of Agentic Workflows

This section aims at experimentally and numerically validatingagentic\-eCALcomponent by component, as introduced in Sec\.[III](https://arxiv.org/html/2609.18283#S3)\. We first calibrate and test the single\-call model of Eq\.[4](https://arxiv.org/html/2609.18283#S3.E4)against direct GPU energy measurement, then examine the scaling law it implies for a sequence of calls Eq\.[8](https://arxiv.org/html/2609.18283#S3.E8), then apply the composed metric to a complete workflow Eq\.[3](https://arxiv.org/html/2609.18283#S3.E3), and finally consider how the embodied term amortizes Eq\.[2](https://arxiv.org/html/2609.18283#S3.E2)\.

### IV\-AExperimental Validation of the LLM Call Energy Model

Eq\.[3](https://arxiv.org/html/2609.18283#S3.E3)breaks down the energy of a workflowEWE\_\{W\}per component, while Sec\.[III\-A](https://arxiv.org/html/2609.18283#S3.SS1)focuses on the first termEc​a​l​lE\_\{call\}and introduces a two rate energy model of a single LLM call in Eq\.[4](https://arxiv.org/html/2609.18283#S3.E4)\. To validate the model in Eq\.[4](https://arxiv.org/html/2609.18283#S3.E4), we vary the number of agentsNN, stepsKKand batch sizebb\. We select two open weight 7\-8bn LLMs, namely Qwen2\.5\-7B, and Llama\-3\.1\-8B \(NVIDIA A100, vLLM continuous batching\), across team sizes of22to3030agents\. Energy is metered with NVML\. As depicted in Fig\.[2](https://arxiv.org/html/2609.18283#S3.F2)a, the energy provided by Eq\.[4](https://arxiv.org/html/2609.18283#S3.E4)\(y axis\) closely matches the measurements \(x axis\) withR2\>0\.99R^\{2\}\\\!\>\\\!0\.99and mean absolute error∼10%\{\\sim\}10\\%\.

We notice that the two terms of Eq\.[6](https://arxiv.org/html/2609.18283#S3.E6)can also be expressed ascdec​\(b\)=adec/b\+c0c\_\{\\mathrm\{dec\}\}\(b\)=a\_\{\\mathrm\{dec\}\}/b\+c\_\{0\}, in whichadeca\_\{\\mathrm\{dec\}\}is the weight\-read coefficient, the part that amortises over the batch, andc0c\_\{0\}the KV\-cache floor, which does not\. By regressing the measured energies at each batch we findadec=3\.2a\_\{\\mathrm\{dec\}\}=3\.2and3\.33\.3andc0=0\.02c\_\{0\}=0\.02and0\.060\.06J/token for Qwen2\.5\-7B and Llama\-3\.1\-8B respectively\. The fittedadeca\_\{\\mathrm\{dec\}\}lies a factor of about2\.72\.7below the datasheet value of the terms of Eq\.[6](https://arxiv.org/html/2609.18283#S3.E6)\. We therefore read that equation as establishing the form of the two terms rather than as predicting their magnitude to better than a third\. We plot the curves with these coefficients together with the measurements in Fig\.[2](https://arxiv.org/html/2609.18283#S3.F2)b and used throughout\.

It can be seen from Fig\.[2](https://arxiv.org/html/2609.18283#S3.F2)b that the measurements validate the two rate model from Eqs\.[5](https://arxiv.org/html/2609.18283#S3.E5)and[6](https://arxiv.org/html/2609.18283#S3.E6): a near\-constantcpre≈0\.02c\_\{\\mathrm\{pre\}\}\\\!\\approx\\\!0\.02–0\.030\.03J/token \(the compute\-bound prefill floor\) and a decode rate that decreases with∼1/b\\sim\\\!1/b\(p<0\.001p\\\!<\\\!0\.001\)\. Both rates match first principles:cprec\_\{\\mathrm\{pre\}\}implies a prefill utilizationηpre≈0\.74\\eta\_\{\\mathrm\{pre\}\}\\\!\\approx\\\!0\.74–0\.850\.85\(compute\-bound, as expected\), and the decode coefficient agrees with the memory\-bandwidth predictionP​N​β/BHBMPN\\beta/B\_\{\\mathrm\{HBM\}\}to within10%10\\%\.

As an independent cross\-hardware check, Table[II](https://arxiv.org/html/2609.18283#S3.T2)evaluates Eq\.[4](https://arxiv.org/html/2609.18283#S3.E4)on hardware and at serving batches we did not measure ourselves, from datasheet quantities alone\. Considering input from the benchmark, including its Standard Load profile \(b=128b\\\!=\\\!128,pout=500p\_\{\\mathrm\{out\}\}\\\!=\\\!500\), its Alpaca prompts and its4×4\{\\times\}H100 NVL node, the model returns 0\.45 J/token for Llama\-3 70B, 0\.025 for Llama\-3 8B and 0\.021 for Qwen2\.5\-7B, against measured 0\.39, 0\.07 and 0\.07\[[18](https://arxiv.org/html/2609.18283#bib.bib19)\]\. The two smaller models occupy one device of four, so their published figures are corrected to that device before comparison; while the7070B shards over all four and needs no correction\. The check shows that the model underestimatesEc​a​l​lE\_\{call\}of the smaller models and overestimates for the larger one\. This discrepancy is consistent across a tenfold range of model size and lies in the magnitude of Eq\.[6](https://arxiv.org/html/2609.18283#S3.E6)and not its form: evaluated at the datasheetBHBMB\_\{\\mathrm\{HBM\}\}the decode rate is a lower bound on served energy, because a serving stack sustains only∼70%\{\\sim\}70\\%of peak bandwidth even in the favourable single\-stream case \(footnote to Eq\.[6](https://arxiv.org/html/2609.18283#S3.E6)\) and the equation omits the empirical decode floorc0c\_\{0\}, KV\-cache fragmentation, idle draw and the cost of attention over the accumulating cache\. This shortfall widens with batch, which is why the coefficients calibrated at our serving batches \(Fig\.[2](https://arxiv.org/html/2609.18283#S3.F2)\) sit above the datasheet curve and are the ones we use throughout\.

Fig\. 3:Measured scaling \(A100, vLLM\) and model overlays\. \(a\) Reasoning depth on Qwen2\.5\-7B and Llama\-3\.1\-8B: per\-query energy vs\. rounds, history\-carrying vs\. history\-free\. \(b\) Agent count on Qwen2\.5\-7B: energy vs\.NNat a small and a large serving batch\. \(c\) The energy\-vs\-agent\-count exponentaarises with batch on both Qwen2\.5\-7B and Llama\-3\.1\-8B: the prefill/decode crossover\.We rely on the decode from memory bandwidth rather than peak FLOP/s because nameplate thermal design power overestimates measured energy by up to4\.1×4\.1\\times: decode is memory\-bandwidth\-bound, occupies7777–91%91\\%of inference time and is nearly insensitive to compute clock\[[8](https://arxiv.org/html/2609.18283#bib.bib7),[29](https://arxiv.org/html/2609.18283#bib.bib21)\]\. Eq\.[6](https://arxiv.org/html/2609.18283#S3.E6)therefore carries a bandwidth term and a batch divisor in place of a utilization factor, and serving choices \(batch size, quantization, KV\-cache reuse\) act on energy by influencingcdec​\(b\)c\_\{\\mathrm\{dec\}\}\(b\): batching alone a33–5×5\\timesreduction\[[19](https://arxiv.org/html/2609.18283#bib.bib18)\], latency\-constrained Pareto configurations a further44%44\\%\[[8](https://arxiv.org/html/2609.18283#bib.bib7)\]and hence the crossover the two\-rate model makes explicit\. The bound presumes saturated batched serving, which early multi\-GPU deployments that ran each device far below its power budget did not reach: the33–44J per output token reported for a 65B model in 2023\[[7](https://arxiv.org/html/2609.18283#bib.bib6)\]splits the model across88–3232GPUs that each draw well under a third of board TDP, and Eq\.[6](https://arxiv.org/html/2609.18283#S3.E6)estimates such a run several\-fold low\.

Fig\. 4:Four\-agent RAG workflow case study \(Llama\-3 8B, A100\)\. \(a\) Sequential execution graph with per\-step token counts, tinted by functional class \(Sec\.[III](https://arxiv.org/html/2609.18283#S3)\)\. \(b\) Prompt token composition: transcript accumulation supplies56%56\\%of the writer’s prompt\. \(c\)agentic\-eCALper agent across serving batchesb∈\{1,64,256\}b\\in\\\{1,64,256\\\}\. \(d\) Inter\-agent transmission energy across physical bearers \(Table[III](https://arxiv.org/html/2609.18283#S4.T3)\) vs\. workflow compute energyEWE\_\{W\}, demonstrating that transport energy is orders of magnitude below compute\.
### IV\-BExperimental Validation of Scaling the Call Energy Model

Due to the relatively large size of the models corroborated with the context,Ec​a​l​lE\_\{call\}dominatesEWE\_\{W\}\. Sec\.[III\-B](https://arxiv.org/html/2609.18283#S3.SS2)models this process through Eqs\.[7](https://arxiv.org/html/2609.18283#S3.E7)and[8](https://arxiv.org/html/2609.18283#S3.E8)\. Under the same experimental conditions from Sec\.[III\-A](https://arxiv.org/html/2609.18283#S3.SS1), we validate the scaling\. Fig\.[3](https://arxiv.org/html/2609.18283#S4.F3)shows the measured scaling against reasoning depth \(carry vs\. free\), agent count, and serving batch\. We also provide model overlays in Fig\.[3](https://arxiv.org/html/2609.18283#S4.F3)\(a\) and \(b\): Eq\.[4](https://arxiv.org/html/2609.18283#S3.E4)evaluated on the same measured token workload at the same serving batch, with i\) the calibrated coefficients of Fig\.[2](https://arxiv.org/html/2609.18283#S3.F2)and ii\) the datasheet numbers\.

In Fig\.[3](https://arxiv.org/html/2609.18283#S4.F3)a, we isolate the cost of conversational memory by comparing history\-carrying against history\-free loops across reasoning depthsK∈\[1,6\]K\\in\[1,6\]\. In a history\-carrying loop, each sequential round re\-reads the accumulated transcript of prior steps, driving quadratic prompt token growth in Eq\.[7](https://arxiv.org/html/2609.18283#S3.E7)\. Consequently, measured energy compounds super\-linearly with depth \(filled markers\), whereas history\-free execution scales strictly linearly \(open markers\)\. AtK=6K=6, history carry inflates energy by32\.8%32\.8\\%on Qwen2\.5\-7B \(8989J vs\.6767J\) and21\.7%21\.7\\%on Llama\-3\.1\-8B \(185185J vs\.152152J\)\. Llama consumes roughly2×2\\timesthe energy of Qwen across all depths due to its higher decode floor \(c0=0\.06c\_\{0\}=0\.06vs\.0\.020\.02J/token\) and larger per\-token key\-value context footprint \(γ=131\\gamma=131vs\.5656kB/token\)\.

It can be seen from Fig\.[3](https://arxiv.org/html/2609.18283#S4.F3)that the two\-rate model follows the curves given by the measured points: the fitted version closer than the datasheet version\. In Fig\.[3](https://arxiv.org/html/2609.18283#S4.F3)a, the hardware datasheet derivation falls34%34\\%below Llama\-3\.1\-8B atK=6K=6because Eq\.[6](https://arxiv.org/html/2609.18283#S3.E6)models pure memory traffic and omits the empirical decode floor \(c0=0\.060c\_\{0\}=0\.060J/token, driven by PagedAttention fragmentation and idle draw\), an omission magnified by Llama’s more decode\-weighted workload \(pin:pout≈2\.7:1p\_\{\\mathrm\{in\}\}\\\!:\\\!p\_\{\\mathrm\{out\}\}\\approx 2\.7:1vs\.4\.5:14\.5:1\)\. This is the same observation discussed in Sec\.[IV\-A](https://arxiv.org/html/2609.18283#S4.SS1)related to Table[II](https://arxiv.org/html/2609.18283#S3.T2)\.

Fig\.[3](https://arxiv.org/html/2609.18283#S4.F3)b evaluates team width scaling \(N∈\[2,30\]N\\in\[2,30\]\) on Qwen2\.5\-7B across serving batchesb=16b=16andb=256b=256\. Under all\-to\-all communication, inter\-agent messages scale asΘ⁡\(N2\)\\Theta\(N^\{2\}\)\. At small batch \(b=16b=16\), memory\-bandwidth\-bound decode is poorly amortized \(cdec≈0\.22c\_\{\\mathrm\{dec\}\}\\approx 0\.22J/token\), dampening the scaling slope \(a≈1\.25a\\approx 1\.25\)\. Conversely, at production batch \(b=256b=256\), weight\-streaming overhead collapses toward the asymptotic key\-value floor \(cdec→c0≈0\.02c\_\{\\mathrm\{dec\}\}\\to c\_\{0\}\\approx 0\.02J/token\), exposing the compute\-bound quadratic prefill term and steepening the scaling exponent toa≈1\.42a\\approx 1\.42\. Crucially, while batching reduces energy by38\.1%38\.1\\%atN=2N=2\(130130J vs\.210210J\), this efficiency margin shrinks to just15\.4%15\.4\\%atN=30N=30\(4,4004\{,\}400J vs\.5,2005\{,\}200J\), demonstrating that batch amortization cannot overcome the quadratic prefill explosion of large, densely coupled multi\- agent teams\.

Fig\.[3](https://arxiv.org/html/2609.18283#S4.F3)c shows the energy\-vs\-agent\-count exponent rising with batch accordingly, froma≈1\.10a\\\!\\approx\\\!1\.10to1\.421\.42for Qwen \(b:→256b\\\!:\\\!2\\\!\\to\\\!256\) and from1\.021\.02to1\.171\.17for Llama, a direct GPU\-energy measurement that confirms the predicted super\-linear, batch\-dependent crossover\. The milder Llama slope reflects the crossover directly: a higher decode floor \(cdec≈0\.06c\_\{\\mathrm\{dec\}\}\\\!\\approx\\\!0\.06vs∼0\.02\{\\sim\}0\.02atb=256b\\\!=\\\!256\) keeps it partly decode\-bound, so architecture sets how far the workflow crosses\. The measured exponent reflects the partial prefill dominance at the tested batches\. The multi\-agent*width*axis is identical: under all\-to\-all communication each ofNNagents reads the otherN−1N\\\!\-\\\!1, so per\-call input isΘ⁡\(N\)\\Theta\(N\)and∑kpin\(k\)=Θ⁡\(N2\)\\sum\_\{k\}p\_\{\\mathrm\{in\}\}^\{\(k\)\}=\\Theta\(N^\{2\}\), giving the same≤a≤21\\\!\\leq\\\!a\\\!\\leq\\\!2law inNN\. Sparser coupling lowers the exponent, and*independent*agents \(no peer context\) collapse it to a linearΘ⁡\(N\)\\Theta\(N\)\(the efficient fan\-out case\)\. A single crossover thus governs both reasoning depth \(KK\) and team width \(NN\)\. In an agentic workflow, the embodied energy still atomizes with workflow invocations as shown witheCAL’s however history*accumulates*within one workflow\.

### IV\-CEnergy of a RAG Workflow

Assuming an agentic workflow including tool call and retrieval, this section aims to study an example 4 agent workflow\. Fig\.[4](https://arxiv.org/html/2609.18283#S4.F4)a defines the workflow used as the case study throughout this section\. A planner, a retriever, an analyst and a writer execute in sequence over a shared transcript, which each step re\-reads in full; the retriever adds1,2801\{,\}280tokens of retrieved context and the analyst one tool observation of150150tokens\. Under these assumptions, the four steps consume4,9904\{,\}990prompt tokens against1,0201\{,\}020generated, a ratio of4\.94\.9\. The imbalance is structural rather than incidental: prompts accumulate across steps while generated output does not, so the transcript supplies56%56\\%of the final step’s prompt against none of the first\.

TABLE III:Inter\-agent transmission as a share of workflow energy, by placement and bearer\.The case\-study workflow is as follows\. Four agents run in sequence, the retriever drawing55chunks of256256tokens from the index and the analyst invoking one tool that returns150150observation tokens; each box gives the step’s prompt and generated token counts\. Token budgets are stipulated; the values in Fig\.[4](https://arxiv.org/html/2609.18283#S4.F4)\(b\) to \(d\) follows from these through the model of Sec\.[III](https://arxiv.org/html/2609.18283#S3)\. As can be seen in Fig\.[4](https://arxiv.org/html/2609.18283#S4.F4)b, all agents ingest a system prompt \(with grey\) and local prompts\. The retriever’s prompt includes retrieval and history tokens while being dominated by retrieved context\. The accumulated transcript grows monotonically across planner, retriever and analyst, and supplies56%56\\%of the writer’s prompt, which is what Eq\.[7](https://arxiv.org/html/2609.18283#S3.E7)characterizes\. With red dots on the figure, the token outputs of the agents are depicted\. It can be seen that they are relatively small compared to the prompts and reach the size of the system prompt only for the writer in this case study\.

Fig\.[4](https://arxiv.org/html/2609.18283#S4.F4)c evaluates each agent of the workflow with the proposedagentic\-eCALfrom Eq\.[2](https://arxiv.org/html/2609.18283#S3.E2)\. It can be seen that the per\-agent figures are nearly uniform atb=1b=1, spanning0\.2030\.203to0\.2250\.225J/bit, because decode dominates and decode energy is proportional to the tokens an agent generates\. Batching removes that proportionality and the agents separate: atb=64b=64the planner, analyst and writer sit between0\.01160\.0116and0\.01230\.0123J/bit whereas the retriever reaches0\.03400\.0340, a factor of2\.82\.8, since it prefills2,0002\{,\}000tokens to emit120120\. The separation persists atb=256b=256, the largest measured batch, where the same three sit between0\.00930\.0093and0\.01010\.0101against the retriever’s0\.03170\.0317\. The agent that is least efficient per useful bit is therefore identifiable only in the regime in which the workflow would actually be served\. The workflow as a whole costs0\.01620\.0162J/bit against0\.01460\.0146for its calls alone, the difference being the tool, the retrieval and the10%10\\%orchestration overhead of Eq\.[3](https://arxiv.org/html/2609.18283#S3.E3)\.

Fig\.[4](https://arxiv.org/html/2609.18283#S4.F4)\(d\) evaluates the energy that the three hand\-offs required by the four\-agent workflow of this case study if they were distributed and the token data were transmitted through the bearers of Table[III](https://arxiv.org/html/2609.18283#S4.T3)\. The1,2901\{,\}290tokens crossing the agent boundaries cost between8\.98\.9mJ on metro fibre and0\.520\.52J on a loaded edge cell, that is between0\.003%0\.003\\%and0\.186%0\.186\\%ofEWE\_\{W\}atb=64b=64, and a further13\.9×13\.9\\timessmaller in proportional terms atb=1b=1, whereEWE\_\{W\}is correspondingly larger\. Distributing this workflow over any bearer is consequently free to within the precision of the energy model itself\.

Fig\. 5:Embodied vs\. operationalagentic\-eCALas a function of served invocationsGGfor Llama\-3 8B and Llama\-3 70B, on the A100 atb=64b\\\!=\\\!64; the crossover atG≈1\.2×1010G\\approx 1\.2\\times 10^\{10\}for the 8B marks where operation overtakes amortized pre\-training\. Qwen2\.5\-7B is omitted as it publishes no pre\-training energy data\.Fig\. 6:The eight orchestration architectures motivated by\[[26](https://arxiv.org/html/2609.18283#bib.bib34),[27](https://arxiv.org/html/2609.18283#bib.bib35),[28](https://arxiv.org/html/2609.18283#bib.bib36)\]nodes colored by functional class \(Sec\.[III](https://arxiv.org/html/2609.18283#S3)\): generate \(theLLproposals\), transform \(critics, refiners, synthesizers, duelists\), the manager, and Diamond’s non\-answering planner\.Tournamentselects through pairwise duels,Treemerges via fan\-in\-three synthesis, andPersona\-Starshares Star’s graph but differs in prompts\. Counts are illustrative\.![Refer to caption](https://arxiv.org/html/2609.18283v1/fig_capacity.png)Fig\. 7:Where an agent team fits across accelerator tiers \(Eq\.[1](https://arxiv.org/html/2609.18283#S3.E1)\)\. \(a\) Model loading across 16 open\-weight models by weight class \(β​Nparams≤Mu\\beta N\_\{\\mathrm\{params\}\}\\leq M\_\{u\}\)\. \(b\) Concurrent serving capacity \(sessions of a 10\-agent team\) under resident KV\-cache context \(γ​∑axa​u​ca≤Mu−β​Nparams\\gamma\\sum\_\{a\}x\_\{au\}c\_\{a\}\\leq M\_\{u\}\-\\beta N\_\{\\mathrm\{params\}\}\), showing density is governed by KV widthγ\\gamma\. \(c\) Inter\-agent edge countEEacross team sizesN∈\{5,…​30\}N\\in\\\{5,\.\.\.30\\\}for the eight orchestration architectures\.
### IV\-DAmortization

Having validated the operational components ofagentic\-eCALfrom Eq\.[3](https://arxiv.org/html/2609.18283#S3.E3), we now evaluate how the embodied \(pre\-training\) energy amortizes over the total number of workflow invocations,GGwith results depicted in Fig\.[5](https://arxiv.org/html/2609.18283#S4.F5)\. The embodied energy of Llama\-3 8B \(0\.910\.91GWh, i\.e\.1\.31\.3M H100\-hours\) dominates the per\-invocationagentic\-eCALuntilG≈1\.2×1010G\\approx 1\.2\\times 10^\{10\}\. Beyond this point, the fixed pre\-training cost atomizes and the operational workflow energyEWE\_\{W\}dictates the efficiency floor\.

This workflow invocation amortization is the agentic equivalent ofeCAL’s single\-inference amortization result\[[4](https://arxiv.org/html/2609.18283#bib.bib1)\]\. Notably, it exceeds the22–6×1086\\times 10^\{8\}inference parity threshold reported empirically for the earlier BLOOMz model by a factor of2020–6060\[[12](https://arxiv.org/html/2609.18283#bib.bib10)\]\. Furthermore, larger models cross this threshold*sooner*: Llama\-3 70B reaches operational parity at just0\.66×0\.66\\timesthe invocations of the 8B model\. This occurs because the pre\-training energy grew by only4\.9×4\.9\\timesbetween the 8B and 70B variants, whereas the operational energy per invocation scales directly with parameter count\. Consequently, the operational cost per call overtakes the amortized pre\-training cost significantly earlier in a massive model’s lifecycle\.

## VPlacement and Communication Characterization of Agentic Workflows

This section evaluates the placement and communication characteristics of agentic workflows across the edge\-cloud continuum, experimentally and numerically grounding the accelerator memory admissibility constraints of Eq\.[1](https://arxiv.org/html/2609.18283#S3.E1)and the network tier substrate illustrated in Fig\.[1](https://arxiv.org/html/2609.18283#S1.F1)\. In classical distributed computing, service placement solves an offloading trade\-off between local computation and transmission energy\. In agentic AI, because inter\-agent text transmission is small \(<0\.25%<0\.25\\%of workflow energy\), the dominant energy cost of distribution is often not the transmission of inter\-agent text itself, but the additional inference and context processing induced by that communication\. Consequently, deployment feasibility and serving density across the network are governed principally by the memory capacity boundaries of Eq\.[1](https://arxiv.org/html/2609.18283#S3.E1)\.

### V\-AWhere a Team Fits: Accelerators, Models and Transports

Fig\. 8:Empirical breakdown of hand\-off across network links\. \(a\) Data volume transmitted per session under four protocols \(K=3K=3\): transcript re\-transmission \(0\.330\.33Mb\), incremental \(0\.0510\.051Mb\), summary \(0\.0790\.079Mb\), and KV\-cache migration \(18\.918\.9Gb\)\. \(b\) Transmitted volume vs\. loop depthKK\(quadratic for transcript and cache; linear for incremental/summary\)\. \(c\) Parity bearer intensityε⋆\\varepsilon^\{\\star\}where transport energy equals compute energy: text hand\-offs remain orders of magnitude above deployed bearers, while KV\-cache migration exceeds compute energy on all cellular bearers\.An agentic workflow is characterised by two properties that a single\-model analysis does not expose\. The first is heterogeneity of scale: the members of a team need not be, and in practice frequently are not, of the same size, so a site must accommodate whichever class of model the workflow employs\. The second is the communication topology, which determines how many messages the members exchange and hence how much traffic a distributed deployment must carry\. This subsection treats the two together\.

We consider1616open\-weight models spanning5\.75\.7to50\.350\.3GB of weights, grouped into four classes by the memory those weights occupy\. The lightest class, below88GB, contains SmolLM3\-3B, Qwen2\.5\-3B and Qwen3\-4B\. The88–1616GB class is the most populous with seven members, comprising the77–88B dense models Mistral\-7B, OLMo\-2\-7B, Qwen2\.5\-7B, Marin\-8B, Llama\-3\.1\-8B and Qwen3\-8B together with the mixture\-of\-experts GPT\-OSS\-20B, whose routed weights occupy11\.711\.7GB\. The1616–3232GB class holds the1414B models Phi\-4, Qwen3\-14B and R1\-Distill\-Qwen\-14B, and the heaviest class,3232to6464GB, holds Llama\-3\.3\-70B and Qwen2\.5\-72B at four\-bit weights together with Qwen3\.6\-27B\. Weight class is not, however, a proxy for what a device can host: quantisation places two7070B\-class models in the same class as a2727B model at native precision, and the key\-value width within a single class varies by an order of magnitude, from3636to512512kB per token of context\. The experiments evaluated1616models in a full mesh communication topology on H100 GPU on 6 tasks overK=3K=3communication rounds and teams ofN∈\{1,\.\.,30\}N\\in\\\{1,\.\.,30\\\}\. We sub\-sample one round and form the communication topologies from Fig\.[6](https://arxiv.org/html/2609.18283#S4.F6)assuming one round of reasoning andN=10N=10agents\.

The teams themselves are organised by one of the eight orchestration architectures of Fig\.[6](https://arxiv.org/html/2609.18283#S4.F6), which differ in how many first\-tier proposals they generate and in how those proposals are subsequently transformed\. Star and Persona\-Star aggregateL=N−1L=N\-1proposals at a single manager and differ only in the prompts assigned to the workers\. Proposer\-Critic, Tournament, Tree and Diamond interpose a hierarchical or filtering tier, pairing proposals with critics, resolving them through a pairwise bracket, merging them by fan\-in\-three synthesis, or conditioning solvers on separately generated plans, so thatLLfalls to between55and77atN=10N=10calls\. Chain and Cascading\-Chain reduceLLto unity and advance a single evolving solution through successive refiners, the latter permitting each refiner to read up to five preceding reports\. Because every edge of these graphs carries one agent’s output into another’s context, the architecture fixes the number of inter\-agent messages, and with it the traffic a distributed deployment must transport\.

Fig\.[7](https://arxiv.org/html/2609.18283#S4.F7)evaluates the consequences of both properties, takingN=10N=10as a mid\-range team size\. Fig\.[7](https://arxiv.org/html/2609.18283#S4.F7)a evaluates the static weight\-loading term \(β​Np​a​r​a​m​s≤Mu\\beta N\_\{params\}\\leq M\_\{u\}\): an88GB module admits two models of the lightest class, a2424GB accelerator admits that class and the88–1616GB class entire, and6464GB or more admits all four classes, including the quantised7070B models\. Above the far edge, therefore, the weight class of a model does not determine whether it can be deployed\.

Fig\.[7](https://arxiv.org/html/2609.18283#S4.F7)b reports what the same accelerators sustain once each agent’s resident context is included, expressed in the unit an operator provisions: concurrent sessions of a ten\-agent team, each device holding one copy of the weights together with one context per agent\. Capacity ranges from one or two sessions on the88GB module to5959on the A100, and its ordering does not follow the weight classes of panel \(a\)\. Within the88–1616GB class alone it varies16×16\\times, because capacity is governed by the key\-value widthγ\\gammarather than by the weights: GPT\-OSS\-20B and Qwen2\.5\-7B, at4848and5656kB per token, sustain112112and8787sessions, whereas OLMo\-2\-7B, a lighter model that retains multi\-head attention at512512kB per token, sustains seven\. The binding constraint on serving density is accordingly the context an agent must retain, not the model it runs\.

Fig\.[7](https://arxiv.org/html/2609.18283#S4.F7)c reports how the architectures differ in the traffic they generate\. Since each edge of the orchestration graph contributes one peer message to some agent’s prompt, the summed prefill of a workflow isN​p0\+E​τNp\_\{0\}\+E\\tauforEEinter\-agent edges: the shape of the graph does not affect cost, and only its edge count does\. The consequence is that most of the architectures are indistinguishable\. Star, Persona\-Star, Proposer\-Critic, Chain and, to within one edge, Tournament and Tree all requireE≈N−1E\\approx N\-1at every team size, so atN=10N=10they lie between88and1010edges\. Only two separate: Diamond atE≈1\.5​NE\\approx 1\.5N, because each solver is conditioned on a separately generated plan, and Cascading\-Chain atE≈5​NE\\approx 5N, because each of its refiners re\-reads up to five preceding reports\. The dispersion is moreover bounded, widening from2\.5×2\.5\\timesatN=5N=5to4\.8×4\.8\\timesatN=30N=30rather than growing with team size\.

Summarizing Fig\.[7](https://arxiv.org/html/2609.18283#S4.F7), we find that the feasibility of deploying an agentic workflow is governed by memory and, within a weight class, by key\-value width\. The orchestration architecture is a traffic lever only through its edge count, and for six of the eight architectures that count is the same\. What distinguishes the two remaining designs is the re\-reading of context, which is the same mechanism that makes prefill the dominant term of Sec\.[III\-B](https://arxiv.org/html/2609.18283#S3.SS2), and it is therefore in the retention of context rather than in the choice of graph that the design freedom lies\.

Fig\. 9:Energy and task completion on an infrastructure\-repair benchmark for six agent topologies on Qwen2\.5\-7B and Qwen3\.5\-9B\. Energies are obtained from the measured tokens using the two rate model calibrated for the respective models\. \(a\) Fraction of the twenty\-four tasks solved against energy per run\. \(b\) Energy cost per solved task\. \(c\) Energy per solved task against proposer count, annotated with the score attained\.
### V\-BInter\-agent Message Payload Assessment

Sec\.[V\-A](https://arxiv.org/html/2609.18283#S5.SS1)establishes how many messages an agent team exchanges, while this section studies what the messages contain subsampling from the same agent trace dataset\. We evaluate four candidate hand\-off mechanisms in Fig\.[8](https://arxiv.org/html/2609.18283#S5.F8)a for the four\-agent workflow atK=3K=3rounds\. Re\-shipping the accumulated transcript, as a stateless endpoint must, transmits0\.330\.33Mbit per session\. Carrying only the increment produced since the recipient last observed the conversation transmits0\.0510\.051Mbit, a factor of6\.56\.5less, and a bounded running summary lies between the two at0\.0790\.079Mbit\. Migrating the key\-value cache rather than reconstructing it transmits18\.918\.9Gbit, five orders of magnitude above the transcript, because the cache entry for a single token occupies131131kB against the1717bits of the token itself\. The protocol choice therefore spans a dynamic range of3\.7×1053\.7\\times 10^\{5\}, compared to4\.8×4\.8\\timesfor the choice of architecture \(Fig\.[7](https://arxiv.org/html/2609.18283#S4.F7)c\) and roughly10310^\{3\}across physical network bearers \(Table[III](https://arxiv.org/html/2609.18283#S4.T3)\): what a hand\-off contains dominates both where it is sent and how it is carried\.

These hand\-off mechanisms also differ fundamentally in their asymptotic scaling with reasoning depth, as evaluated in Fig\.[8](https://arxiv.org/html/2609.18283#S5.F8)b\. Re\-shipping the transcript grows quadratically inKK, since each round re\-transmits everything produced before it, rising30×30\\timesforK=\{1,…​6\}K=\\\{1,\.\.\.6\\\}\. In contrast, incremental hand\-off grows linearly and rises6×6\\timesover the same range, while the bounded summary rises6\.5×6\.5\\times\. The gap between the stateless and incremental strategies widens with the depth of the loop, reflecting the same history accumulation that makes prompt prefill the dominant term in Sec\.[III\-B](https://arxiv.org/html/2609.18283#S3.SS2)\. Cache migration inherits both this quadratic growth and the five\-order magnitude offset simultaneously\.

To evaluate under what conditions transmission energy could challenge compute dominance, Fig\.[8](https://arxiv.org/html/2609.18283#S5.F8)c inverts the comparison and calculates the parity bearer intensityε⋆\\varepsilon^\{\\star\}required for transmission energy to equal the accompanying compute energy\. For every text\-carrying protocol,ε⋆\\varepsilon^\{\\star\}exceeds3×10−33\\times 10^\{\-3\}J/b across all reasoning depths, which represents more than two orders of magnitude above even a loaded edge cell and five orders of magnitude above an optical core: no deployed bearer brings a transcript hand\-off within reach of the compute it serves\. Migrating the key\-value cache reverses this conclusion entirely: its parity thresholdε⋆\\varepsilon^\{\\star\}falls to5\.1×10−85\.1\\times 10^\{\-8\}J/b byK=6K=6, leaving a margin of only5×5\\timesover an optical backbone and falling below mobile bearers, where cache transport exceeds local compute energy by a factor of twenty\. Attention state is therefore effectively non\-transportable across network links, binding an agent’s context to the local accelerator that built it and confirming that the memory capacity boundaries of Sec\.[V\-A](https://arxiv.org/html/2609.18283#S5.SS1)cannot be relieved by offloading raw cache tensors over the network\.

In this section we showed that communication efficiency is obtained by transmitting increments rather than transcripts, keeping the cache local\. The bearer and the orchestration graph are, by comparison, second\-order choices\.

## VIEnergy per Solved Task on an Infrastructure Benchmark

To evaluate agentic efficiency on delivered network utility, this section measures task completion on an infrastructure incident benchmark representative of modern Cloud\-Native Network Functions \(CNFs\) at the telco edge\.

In 5G\-Advanced and emerging 6G architectures, edge computing platforms and O\-RAN cloud infrastructures \(O\-Cloud\) rely on containerized network microservices orchestrated via Kubernetes \(e\.g\., k3s for far\-edge appliances\)\. Following the ETSI Zero\-touch network and Service Management \(ZSM\) framework, autonomous agent teams serve as closed\-loop diagnostic and remediation engines\. We evaluate five multi\-agent topologies and a single\-agent baseline on a balanced 24\-incident subset of the*kubernetes\-core*suite of Kubeply’s Infra\-Bench333Kubeply, “infra\-bench: Open benchmark tasks for evaluating AI agents on real infrastructure work,” Apr\. 2026\. https://github\.com/kubeply/infra\-benchover ephemeral k3s edge clusters, with eight tasks per difficulty tier\. The configurations comprise a single ReAct agent \(Single\), a worker with critic feedback \(W\+C\), a worker with constraint verification \(W\+V\), parallel proposers followed by a synthesiser \(PS\), PS with verification \(PS\+V\), and a planner–reviewer–worker–verifier pipeline \(P\+R\+W\+V\)\. The benchmark spans incidents like service routing, access control faults, and gives agents 60 minutes and 50 topology turns to diagnose and repair each live cluster before deterministic state verification\.

The results depicted in Fig\.[9](https://arxiv.org/html/2609.18283#S5.F9)are averaged over at least three runs per configuration, with energy obtained from measured token counts assuming an effective decode batch ofb=64b=64using model\-specific two rate coefficients:cpre=0\.023c\_\{\\mathrm\{pre\}\}=0\.023J/token andcdec​\(64\)=0\.07c\_\{\\mathrm\{dec\}\}\(64\)=0\.07J/token for Qwen2\.5\-7B, andcpre=0\.024c\_\{\\mathrm\{pre\}\}=0\.024J/token andcdec​\(64\)=0\.15c\_\{\\mathrm\{dec\}\}\(64\)=0\.15J/token for Qwen3\.5\-9B\. From Fig\.[9](https://arxiv.org/html/2609.18283#S5.F9)a it can be seen that for Qwen3\.5\-9B, the topologies consume between257\.0257\.0and1029\.21029\.2kJ per run, a factor of4\.04\.0, while performsnce scores range from50\.0%50\.0\\%to64\.6%64\.6\\%\. Single solves61\.7%61\.7\\%of tasks; only P\+R\+W\+V exceeds this score, reaching64\.6%64\.6\\%while consuming2\.9×2\.9\\timesmore energy\. For Qwen2\.5\-7B, Single solves9\.2%9\.2\\%of tasks, while the multi\-agent topologies score between0\.4%0\.4\\%and4\.6%4\.6\\%despite consuming up to2611\.12611\.1kJ per run\. Fig\.[9](https://arxiv.org/html/2609.18283#S5.F9)b reports aggregate energy consumption divided by successful task completions\. For Qwen3\.5\-9B, Single costs17\.417\.4kJ per successful repair, compared with29\.029\.0to79\.279\.2kJ for the multi\-agent topologies\. For Qwen2\.5\-7B, the corresponding costs are49\.749\.7kJ for Single and320\.0320\.0to5222\.35222\.3kJ for the multi\-agent topologies, as several configurations complete fewer than one task per run on average\. Fig\.[9](https://arxiv.org/html/2609.18283#S5.F9)c examines scaling from one to five proposers\. For Qwen3\.5\-9B, PS cost rises from31\.231\.2to79\.279\.2kJ per solved task while its score remains approximately constant\. For PS\+V, the score rises from45\.8%45\.8\\%to59\.7%59\.7\\%as the cost increases from51\.551\.5to68\.468\.4kJ\. Qwen2\.5\-7B remains below5%5\\%throughout\. Replication increases energy in every series, while the largest observed score gain occurs for Qwen3\.5\-9B PS\+V\.

## VIIDiscussion and Limitations

Whileagentic\-eCALestablishes a closed\-form foundation for multi\-agent lifecycle accounting, three operational dimensions warrant further refinement\. First, our formulation treats prompt prefill as computing over the full prefix, which represents a conservative upper bound\. In production serving stacks with prefix caching \(e\.g\., vLLM automatic prefix caching and RadixAttention\), reusable system prompts and shared conversational branches avoid redundant tensor evaluations, cutting time\-to\-first\-token by up to4×4\\times\[[16](https://arxiv.org/html/2609.18283#bib.bib16)\]\. Incorporating cache hit\-rate statistics directly into the prefill summation is a natural extension\. Second, for Mixture\-of\-Experts architectures, the parameter countNparamsN\_\{\\mathrm\{params\}\}must be replaced by active routed parametersNactN\_\{\\mathrm\{act\}\}, as unrouted expert weights reside in memory but do not incur per\-token matmul FLOPs\[[19](https://arxiv.org/html/2609.18283#bib.bib18)\]\. Third, our empirical calibration captures PagedAttention memory fragmentation on A100; newer microarchitectures featuring hardware\-accelerated FP8 or FP4 tensor cores and unified memory will shift the decode roofline slopeadec/ba\_\{\\mathrm\{dec\}\}/band reduce the asymptotic floorc0c\_\{0\}\. Finally, embodied pre\-training energy is amortized across an estimated global invocation volumeGG, which remains challenging to track for open\-weight models redistributed across autonomous private network operators\.

## VIIIConclusion and Future Work

This paper extended eCAL to agentic AI workflows, introducing agentic\-eCAL to account for LLM inference, context accumulation, tool execution, retrieval, inter\-agent communication, and embodied energy\. The results show that context processing, rather than text transmission, is often the dominant energy cost of multi\-agent execution\. Under the evaluated network conditions, text\-based inter\-agent communication contributes only a small fraction of workflow energy, while accumulated context can drive super\-linear energy growth\. In contrast, KV\-cache migration introduces orders\-of\-magnitude larger payloads and can make network transfer energetically impractical\. The infrastructure benchmark further shows that increasing agent count does not necessarily improve task success and can substantially increase energy per solved task\. These findings suggest that energy\-aware agentic systems should prioritize context management, incremental hand\-offs, local KV state, and agent selection based on measurable task benefit\. Future work will incorporate prefix\-cache reuse, additional serving stacks and accelerators, speculative decoding, and direct energy measurements of complete agentic workloads\.

## References

- \[1\]A\. Mehonic and A\. J\. Kenyon\(2022\)Brain\-inspired computing needs a master plan\.Nature604\(7905\),pp\. 255–260\.Cited by:[§I](https://arxiv.org/html/2609.18283#S1.p1.1)\.
- \[2\]K\. Crawford\(2024\)Generative ai’s environmental costs are soaring—and mostly secret\.Nature626\(8000\),pp\. 693–693\.Cited by:[§I](https://arxiv.org/html/2609.18283#S1.p1.1)\.
- \[3\]F\. Oviedo, F\. Kazhamiaka, E\. Choukse, A\. Kim, A\. Luers, M\. Nakagawa, R\. Bianchini, and J\. M\. Lavista Ferres\(2026\)Energy use of ai inference, efficiency pathways, and test\-time scaling\.Joule,pp\. 102430\.External Links:ISSN 2542\-4351,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.joule.2026.102430),[Link](https://www.sciencedirect.com/science/article/pii/S2542435126001145)Cited by:[§I](https://arxiv.org/html/2609.18283#S1.p1.1),[§II](https://arxiv.org/html/2609.18283#S2.p2.1),[§II](https://arxiv.org/html/2609.18283#S2.p4.1)\.
- \[4\]S\. Chou, J\. Hribar, V\. Hanžel, M\. Mohorčič, and C\. Fortuna\(2026\)The energy cost of artificial intelligence lifecycle in communication networks\.IEEE Journal on Selected Areas in Communications44,pp\. 2427–2443\.External Links:[Document](https://dx.doi.org/10.1109/JSAC.2025.3642835)Cited by:[1st item](https://arxiv.org/html/2609.18283#S1.I1.i1.p1.1),[§I](https://arxiv.org/html/2609.18283#S1.SS0.SSS0.Px1.p1.1),[§I](https://arxiv.org/html/2609.18283#S1.p3.1),[§II](https://arxiv.org/html/2609.18283#S2.p1.1),[§III](https://arxiv.org/html/2609.18283#S3.SS0.SSS0.Px1.p1.1),[§III](https://arxiv.org/html/2609.18283#S3.SS0.SSS0.Px2.p2.1),[§III\-C](https://arxiv.org/html/2609.18283#S3.SS3.p1.1),[§III\-E](https://arxiv.org/html/2609.18283#S3.SS5.p2.1),[§III](https://arxiv.org/html/2609.18283#S3.p4.1),[§IV\-D](https://arxiv.org/html/2609.18283#S4.SS4.p2.1)\.
- \[5\]J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. Amodei\(2020\)Scaling laws for neural language models\.arXiv preprint arXiv:2001\.08361\.Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p2.1)\.
- \[6\]J\. Austin, S\. Douglas, R\. Frostig, A\. Levskaya, C\. Chen, S\. Vikram, F\. Lebron, P\. Choy, V\. Ramasesh, A\. Webson, and R\. Pope\(2025\)How to scale your model\.Note:[https://jax\-ml\.github\.io/scaling\-book/transformers/](https://jax-ml.github.io/scaling-book/transformers/)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p2.1)\.
- \[7\]S\. Samsi, D\. Zhao, J\. McDonald, B\. Li, A\. Michaleas, M\. Jones, W\. Bergeron, J\. Kepner, D\. Tiwari, and V\. Gadepally\(2023\)From words to watts: benchmarking the energy costs of large language model inference\.In2023 IEEE High Performance Extreme Computing Conference \(HPEC\),Note:arXiv:2310\.03003Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p2.1),[§IV\-A](https://arxiv.org/html/2609.18283#S4.SS1.p5.1)\.
- \[8\]J\. Chung, J\. J\. Ma, R\. Wu, J\. Liu, O\. J\. Kweon, Y\. Xia, Z\. Wu, and M\. ChowdhuryD\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\)\(2025\)The ml\.energy benchmark: toward automated inference energy measurement and optimization\.Vol\.38, Main Conference,Curran Associates, Inc\.\.External Links:[Document](https://dx.doi.org/10.52202/085713-3652),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/9dc510e3d7b0b3b2a58ffed7a3ad6b0f-Paper-Datasets_and_Benchmarks_Track.pdf)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p2.1),[§IV\-A](https://arxiv.org/html/2609.18283#S4.SS1.p5.1)\.
- \[9\]H\. Touvron, L\. Martin, K\. Stone,et al\.\(2023\)Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p3.1)\.
- \[10\]Meta AI\(2024\)Llama 3 model card\.Note:[https://github\.com/meta\-llama/llama3/blob/main/MODEL\_CARD\.md](https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p3.1)\.
- \[11\]A\. S\. Luccioni, S\. Viguier, and A\. Ligozat\(2023\)Estimating the carbon footprint of BLOOM, a 176B parameter language model\.J\. Mach\. Learn\. Res\.24\(253\),pp\. 1–15\.Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p3.1)\.
- \[12\]S\. Luccioni, Y\. Jernite, and E\. Strubell\(2024\)Power hungry processing: Watts driving the cost of AI deployment?\.InProc\. ACM Conf\. on Fairness, Accountability, and Transparency \(FAccT\),FAccT ’24,pp\. 85–99\.External Links:[Document](https://dx.doi.org/10.1145/3630106.3658542)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p3.1),[§IV\-D](https://arxiv.org/html/2609.18283#S4.SS4.p2.1)\.
- \[13\]G\. Srivastava, A\. S\. Hussain, S\. Srinivasan, and X\. Wang\(2026\)Do LLMs overthink basic math reasoning? benchmarking the accuracy\-efficiency tradeoff in language models\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 25784–25826\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.1285)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p4.1)\.
- \[14\]J\. Hadfield, B\. Zhang, K\. Lien, F\. Scholz, J\. Fox, and D\. Ford\(2025\)How we built our multi\-agent research system\.Note:[https://www\.anthropic\.com/engineering/multi\-agent\-research\-system](https://www.anthropic.com/engineering/multi-agent-research-system)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p4.1)\.
- \[15\]Y\. Zeng, W\. Huang, L\. Jiang, T\. Liu, X\. Jin, C\. T\. Tiana, J\. Li, and X\. Xu\(2025\)S2\{\}^\{2\}\-MAD: breaking the token barrier to enhance multi\-agent debate efficiency\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 9393–9408\.Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p4.1),[§II](https://arxiv.org/html/2609.18283#S2.p6.1)\.
- \[16\]C\. Jin, Z\. Zhang, X\. Jiang, F\. Liu, S\. Liu, X\. Liu, and X\. Jin\(2025\)RAGCache: efficient knowledge caching for retrieval\-augmented generation\.ACM Transactions on Computer Systems44\(1\),pp\. 1–27\.External Links:[Document](https://dx.doi.org/10.1145/3768628)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p4.1),[§VII](https://arxiv.org/html/2609.18283#S7.p1.1)\.
- \[17\]Z\. Guo, C\. Gao, and J\. Bogner\(2026\)On the effectiveness of proposed techniques to reduce energy consumption in RAG systems: a controlled experiment\.InProceedings of the IEEE/ACM 48th International Conference on Software Engineering,pp\. 101–112\.Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p4.1)\.
- \[18\]C\. Niu, W\. Zhang, J\. Li, Y\. Zhao, T\. Wang, X\. Wang, and Y\. Chen\(2026\)TokenPowerBench: benchmarking the power consumption of LLM inference\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 32582–32590\.Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p4.1),[§III\-D](https://arxiv.org/html/2609.18283#S3.SS4.p1.2),[TABLE II](https://arxiv.org/html/2609.18283#S3.T2.2.2.5.1),[TABLE II](https://arxiv.org/html/2609.18283#S3.T2.2.3.5.1),[TABLE II](https://arxiv.org/html/2609.18283#S3.T2.2.4.5.1),[TABLE II](https://arxiv.org/html/2609.18283#S3.T2.2.6.1.1),[§IV\-A](https://arxiv.org/html/2609.18283#S4.SS1.p4.1)\.
- \[19\]J\. Chung, R\. Wu, J\. J\. Ma, and M\. Chowdhury\(2026\)Where do the joules go? diagnosing inference energy consumption\.External Links:2601\.22076,[Link](https://arxiv.org/abs/2601.22076)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p4.1),[§IV\-A](https://arxiv.org/html/2609.18283#S4.SS1.p5.1),[§VII](https://arxiv.org/html/2609.18283#S7.p1.1)\.
- \[20\]ETSI\(2026\)Zero\-touch network and service management \(ZSM\); study on the utilization of agents in autonomous networks\.Group ReportTechnical ReportETSI GR ZSM 020 V1\.1\.1,European Telecommunications Standards Institute \(ETSI\)\.External Links:[Link](https://www.etsi.org/deliver/etsi_gr/ZSM/001_099/020/01.01.01_60/gr_zsm020v010101p.pdf)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p5.1)\.
- \[21\]O\-RAN Alliance Working Group 6\(2025\)Cloud architecture and deployment scenarios for O\-RAN virtualized RAN\.Technical ReportTechnical ReportO\-RAN\.WG6\.CADS\-v08\.01,O\-RAN Alliance\.Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p5.1)\.
- \[22\]S\. Chou, J\. Hribar, B\. Bertalanič, M\. Mohorčič, T\. Lagkas, P\. Sarigiannidis, and C\. Fortuna\(2025\)Energy cost of the AI/ML workflow in O\-RAN\.In2025 IEEE Conf\. on Network Function Virtualization and Software\-Defined Networking \(NFV\-SDN\),pp\. 1–6\.External Links:[Document](https://dx.doi.org/10.1109/NFV-SDN66355.2025.11349371)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p5.1)\.
- \[23\]S\. Chou, J\. Hribar, M\. Mohorčič, and C\. Fortuna\(2024\)Towards the standardization of energy efficiency metrics of the AI lifecycle in 6G and beyond\.In2024 IEEE Conf\. on Standards for Communications and Networking \(CSCN\),pp\. 187–190\.External Links:[Document](https://dx.doi.org/10.1109/CSCN63874.2024.10849732)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p5.1)\.
- \[24\]P\. Mach and Z\. Becvar\(2017\)Mobile edge computing: a survey on architecture and computation offloading\.IEEE communications surveys & tutorials19\(3\),pp\. 1628–1656\.External Links:[Document](https://dx.doi.org/10.1109/COMST.2017.2682318)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p6.1)\.
- \[25\]R\. Jacob, L\. Röllin, J\. Lim, J\. Chung, M\. Béhanzin, W\. Wang, A\. Hunziker, T\. Moroianu, S\. Tabaeiaghdaei, A\. Perrig, and L\. Vanbever\(2025\)Fantastic joules and where to find them\. modeling and optimizing router energy demand\.InProc\. ACM Internet Measurement Conference \(IMC\),pp\. 48–62\.External Links:[Document](https://dx.doi.org/10.1145/3730567.3732920)Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p6.1)\.
- \[26\]Y\. Yue, G\. Zhang, B\. Liu, G\. Wan, K\. Wang, D\. Cheng, and Y\. Qi\(2025\)MasRouter: learning to route LLMs for multi\-agent systems\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15549–15572\.Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p6.1),[Fig\. 6](https://arxiv.org/html/2609.18283#S4.F6)\.
- \[27\]G\. Zhang, L\. Niu, J\. Fang, K\. Wang, L\. Bai, and X\. Wang\(2025\)Multi\-agent architecture search via agentic supernet\.InProceedings of the 42nd ICML,Vol\.267,pp\. 75834–75852\.Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p6.1),[Fig\. 6](https://arxiv.org/html/2609.18283#S4.F6)\.
- \[28\]G\. Yu\(2026\)AdaptOrch: task\-adaptive multi\-agent orchestration in the era of LLM performance convergence\.arXiv preprint arXiv:2602\.16873\.Cited by:[§II](https://arxiv.org/html/2609.18283#S2.p6.1),[Fig\. 6](https://arxiv.org/html/2609.18283#S4.F6)\.
- \[29\]P\. J\. Maliakel, S\. Ilager, and I\. Brandic\(2026\)Characterizing LLM inference energy\-performance tradeoffs across workloads and GPU scaling\.InProceedings of the 26th IEEE International Symposium on Cluster, Cloud and Internet Computing \(CCGrid\),pp\. 33–43\.External Links:[Document](https://dx.doi.org/10.1109/CCGrid68966.2026.00013)Cited by:[§III\-A](https://arxiv.org/html/2609.18283#S3.SS1.p2.2),[§IV\-A](https://arxiv.org/html/2609.18283#S4.SS1.p5.1)\.
- \[30\]PyTorch Foundation\(2023\)Accelerating generative AI with PyTorch II: GPT, fast\.Note:PyTorch Blog[https://pytorch\.org/blog/accelerating\-generative\-ai\-2/](https://pytorch.org/blog/accelerating-generative-ai-2/)Cited by:[footnote 2](https://arxiv.org/html/2609.18283#footnote2)\.
- \[31\]M\. Agarwal, A\. Qureshi, N\. Sardana, L\. Li, J\. Quevedo, and D\. Khudia\(2023\)LLM inference performance engineering: best practices\.Note:Databricks Blog[https://www\.databricks\.com/blog/llm\-inference\-performance\-engineering\-best\-practices](https://www.databricks.com/blog/llm-inference-performance-engineering-best-practices)Cited by:[footnote 2](https://arxiv.org/html/2609.18283#footnote2)\.
- \[32\]Y\. A\. Malkov and D\. A\. Yashunin\(2020\)Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs\.IEEE Transactions on Pattern Analysis and Machine Intelligence42\(4\),pp\. 824–836\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2018.2889473)Cited by:[§III\-D](https://arxiv.org/html/2609.18283#S3.SS4.p1.1)\.
- \[33\]H\. Pizzini Cavagna, A\. Proia, G\. Madella, G\. B\. Esposito, F\. Antici, D\. Cesarini, Z\. Kiziltan, and A\. Bartolini\(2026\)SweetSpot: an analytical model for predicting energy efficiency of LLM inference\.InProceedings of the 17th ACM/SPEC International Conference on Performance Engineering,pp\. 83–95\.External Links:[Document](https://dx.doi.org/10.1145/3777884.3797011)Cited by:[§III\-D](https://arxiv.org/html/2609.18283#S3.SS4.p1.2)\.
- \[34\]W\. Van Heddeghem, F\. Idzikowski, W\. Vereecken, D\. Colle, M\. Pickavet, and P\. Demeester\(2012\)Power consumption modeling in optical multilayer networks\.Photonic Network Communications24\(2\),pp\. 86–102\.External Links:[Document](https://dx.doi.org/10.1007/s11107-011-0370-7)Cited by:[TABLE III](https://arxiv.org/html/2609.18283#S4.T3.2.2.1.1)\.
- \[35\]J\. Lorincz, E\. Čusto, and D\. Begušić\(2025\)A comprehensive analysis of methods for improving and estimating energy efficiency of passive and active fiber\-to\-the\-home optical access networks\.Sensors25\(19\),pp\. 6012\.External Links:[Document](https://dx.doi.org/10.3390/s25196012)Cited by:[TABLE III](https://arxiv.org/html/2609.18283#S4.T3.2.3.1.1)\.
- \[36\]D\. Xu, A\. Zhou, X\. Zhang, G\. Wang, X\. Liu, C\. An, Y\. Shi, L\. Liu, and H\. Ma\(2020\)Understanding operational 5G: a first measurement study on its coverage, performance and energy consumption\.InProc\. ACM SIGCOMM,pp\. 479–494\.External Links:[Document](https://dx.doi.org/10.1145/3387514.3405882)Cited by:[TABLE III](https://arxiv.org/html/2609.18283#S4.T3.2.4.1.1),[TABLE III](https://arxiv.org/html/2609.18283#S4.T3.2.5.1.1),[TABLE III](https://arxiv.org/html/2609.18283#S4.T3.2.7.1.1)\.

Similar Articles

AI agents are changing how people think about compute costs

Reddit r/AI_Agents

The article discusses how AI agent workflows are shifting optimization focus from pure inference costs to broader challenges like latency, orchestration overhead, and reliability. It highlights a trend toward hybrid architectures and dynamic model routing to address these multi-step workflow complexities.

Improving the speed and energy-efficiency of AI agents

MIT News — Artificial Intelligence

Researchers from MIT and Microsoft developed an intelligent system that automatically optimizes agentic workflows, reducing computational resources and energy usage while maintaining performance.