Tag
A paper introducing predictive speculative KV replication to handle bursty LLM inference workloads, with code available on GitHub.