reusable-computation

Tag

Cards List
#reusable-computation

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference

arXiv cs.AI · 4d ago Cached

MiniCache is a program caching framework that reuses computation across similar requests by parameterizing Program-of-Thought programs, using small models for semantic variable extraction and speculative drafting to improve LLM inference efficiency.

0 favorites 0 likes
← Back to home

Submit Feedback