Tag
PALS is a power-aware runtime for LLM serving that treats GPU power caps as a controllable knob, jointly optimizing them with batch size to maximize energy efficiency while meeting throughput targets. The system improves energy efficiency by up to 26.3% and reduces QoS violations by 4x-7x under power constraints.