Tag
AugServe introduces a state-aware request scheduling framework with dynamic batch-level token budgets to mitigate head-of-line blocking and improve effective throughput for augmented LLM inference serving, achieving up to 6.5x higher throughput than vLLM.
Charlie Marsh reports that Codex improved request scheduling, making uncached complex resolutions (e.g., Transformers) in uv over 40% faster.