Tag
AOSpec is a lossless framework that co-speculates actions and observations across the LLM agent-environment loop to reduce latency, achieving notable end-to-end latency reductions across various serving settings.
A computer science paper presents a graph-based workflow serving engine that unifies agent operations into a global wGraph, using dynamic graph synthesis and differential KV-cache reuse to boost agent accuracy by 4.95% while cutting GPU memory usage by 4x.