head-of-line-blocking

Tag

Cards List
#head-of-line-blocking

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

arXiv cs.CL · 2026-07-13 Cached

AugServe introduces a state-aware request scheduling framework with dynamic batch-level token budgets to mitigate head-of-line blocking and improve effective throughput for augmented LLM inference serving, achieving up to 6.5x higher throughput than vLLM.

0 favorites 0 likes
#head-of-line-blocking

Duration Aware Scheduling for ASR Serving Under Workload Drift

Hugging Face Daily Papers · 2026-03-11 Cached

This paper presents a duration-aware scheduling approach for automatic speech recognition (ASR) serving that uses audio duration as a proxy for job time to reduce head-of-line blocking latency under workload drift.

0 favorites 0 likes
← Back to home

Submit Feedback