energy-estimation

Tag

Cards List
#energy-estimation

Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors

arXiv cs.LG · 2026-08-10 Cached

This paper introduces HYMELL, a hybrid analytical-machine-learning framework for estimating LLM inference latency and energy across prefill and decode phases, validated on NVIDIA H100 with under 5% error for LLaMA 3 8B.

0 favorites 0 likes
#energy-estimation

From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs

arXiv cs.LG · 2026-07-30 Cached

This paper presents an analytically structured, empirically calibrated methodology for estimating LLM inference energy on NVIDIA H100 GPUs without direct measurement, separating prefill and decoding phases and decomposing energy into compute, parameter-access, KV-cache write, and attention-read components.

0 favorites 0 likes
#energy-estimation

WattLayer: Get Layers Right to Estimate Inference Energy of Neural Networks

arXiv cs.LG · 2026-06-29 Cached

This paper introduces WattLayer, a task-independent layer-wise energy estimation model for neural networks, evaluated on over 100,000 layers across 295 architectures, achieving a median error of 19.6% and outperforming state-of-the-art methods.

0 favorites 0 likes
← Back to home

Submit Feedback