power-aware

Tag

Cards List
#power-aware

PALS: Power-Aware LLM Serving for Mixture-of-Experts Models

arXiv cs.AI ↗ · 2026-05-22 Cached

PALS is a power-aware runtime for LLM serving that treats GPU power caps as a controllable knob, jointly optimizing them with batch size to maximize energy efficiency while meeting throughput targets. The system improves energy efficiency by up to 26.3% and reduces QoS violations by 4x-7x under power constraints.

0 favorites 0 likes
← Back to home

Submit Feedback