cache-reuse

Tag

Cards List
#cache-reuse

I tested DeepSeek Harness with GLM, Kimi, Opus, and GPT to see if prompt caching still works with other models

Reddit r/AI_Agents · 5d ago

The article tests whether DeepSeek Harness maintains high prompt caching rates when using alternative AI models, finding that GLM and Kimi achieve 97-99% cache reuse, while Opus shows no cache activity and GPT test failed.

0 favorites 0 likes
#cache-reuse

From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

arXiv cs.AI · 2026-08-14 Cached

This paper proposes Global-ImpactCache (GCache), a bilevel optimization framework that learns cache reuse policies for diffusion models by aligning error weighting with final generation quality, instead of relying on local similarity heuristics. It achieves significant speedups and quality improvements on image and video generation tasks, including a 2.17x speedup on Wan2.1 with lower LPIPS.

0 favorites 0 likes
#cache-reuse

@m_sirovatka: KV Cache re-use is the most important thing for agentic rollouts. We've integrated Mooncake Store into prime-rl with vL…

X AI KOLs Following · 2026-06-02 Cached

vLLM integrates Mooncake Store for distributed KV cache reuse, enabling cross-node prefix caching to efficiently serve agentic workloads with high token reuse.

0 favorites 0 likes
#cache-reuse

KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs

Hugging Face Daily Papers · 2026-04-14 Cached

KV Packet proposes a recomputation-free cache reuse framework for LLMs that uses trainable soft-token adapters to bridge context discontinuities, eliminating overhead while maintaining performance comparable to full recomputation baselines on Llama-3.1 and Qwen2.5.

0 favorites 0 likes
← Back to home

Submit Feedback