Tag
AugServe introduces a state-aware request scheduling framework with dynamic batch-level token budgets to mitigate head-of-line blocking and improve effective throughput for augmented LLM inference serving, achieving up to 6.5x higher throughput than vLLM.
This research paper introduces adaptive correction scheduling for enforcing hard constraints in generative sampling, demonstrating that it improves the cost-accuracy frontier compared to terminal or stepwise projection methods.