Tag
This paper proposes a phase-decoupled, model-calibrated power control method for disaggregated LLM serving, demonstrating improved energy efficiency with minimal latency impact compared to NVIDIA's Max-Q profiles on MoE models.