Tag
This paper presents an analytically structured, empirically calibrated methodology for estimating LLM inference energy on NVIDIA H100 GPUs without direct measurement, separating prefill and decoding phases and decomposing energy into compute, parameter-access, KV-cache write, and attention-read components.
The Low Resource Computing 2026 workshop challenges the trend of infinite growth in computing, advocating for resource-efficient computation inspired by historical achievements, aiming to put processing power back in individuals' hands for a sustainable future.
This survey comprehensively reviews resource-efficient architectures and hardware-software co-design for green AI, covering efficient model construction, training/deployment strategies, and sustainable hardware, aiming to guide sustainable large model development.