dependency-consistency

Tag

Cards List
#dependency-consistency

FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents

arXiv cs.AI · yesterday Cached

FinCacheServe is a system for dependency-consistent answer reuse in RAG serving over mutable enterprise documents, using document versions, evidence fingerprints, and tool fingerprints to invalidate caches. Evaluations show it skips over 53% of LLM calls with zero stale outputs, reducing GPU cost compared to versioned semantic caching.

0 favorites 0 likes
← Back to home

Submit Feedback