answer-caching

标签

Cards List
#answer-caching

FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents

arXiv cs.AI · 昨天 缓存

FinCacheServe is a system for dependency-consistent answer reuse in RAG serving over mutable enterprise documents, using document versions, evidence fingerprints, and tool fingerprints to invalidate caches. Evaluations show it skips over 53% of LLM calls with zero stale outputs, reducing GPU cost compared to versioned semantic caching.

0 人收藏 0 人点赞
← 返回首页

提交意见反馈