Tag
This paper presents a tail-risk-aware scheduling method for agentic LLM workflows that reduces tail latency by optimizing turn release decisions, achieving up to a 3.50x speedup in P95 workflow flow time under contention.