prompt-cache

Tag

Cards List
#prompt-cache

@DeRonin_: follow these three rules and Claude Code gets radically better: - use /subtask instead of spawning a fresh subagent, it…

X AI KOLs Following · 2026-08-28 Cached

The article advises using /subtask, background execution, and worktree isolation in Claude Code to enhance efficiency by sharing prompt cache and managing context effectively, reducing costs.

0 favorites 0 likes
#prompt-cache

Your Agentic Workflow's Cache Keepalive Costs 8x Too Much

Lobsters Hottest · 2026-07-21 Cached

A detailed measurement study across Anthropic, OpenAI, Gemini, and DeepSeek finds that the conventional 30-second prompt cache keepalive is 8x too frequent; a 4-minute interval is optimal, and only Anthropic's cache saves money at long idle gaps.

0 favorites 0 likes
#prompt-cache

Self-hosting with llama-server? Fix for prompt cache reloading on every turn

Reddit r/openclaw · 2026-06-29

A troubleshooting guide showing how changing OpenClaw's contextInjection setting from 'always' to 'continuation-skip' fixes prompt cache reloading on every turn when using llama-server, resulting in a 100x speed improvement for long sessions.

0 favorites 0 likes
#prompt-cache

@rohanpaul_ai: TokenPilot reduces LLM agent costs via ingestion-aware compaction and lifecycle-aware eviction. Achieves 61–87% cost re…

X AI KOLs Following · 2026-06-16 Cached

TokenPilot reduces LLM agent costs via ingestion-aware compaction and lifecycle-aware eviction, achieving 61–87% cost reduction on PinchBench and Claw-Eval with competitive scores.

0 favorites 0 likes
#prompt-cache

TokenPilot: Cache-Efficient Context Management for LLM Agents

Hugging Face Daily Papers · 2026-06-15 Cached

TokenPilot is a dual-granularity context management framework that reduces inference costs in long-horizon LLM sessions by stabilizing prompt prefixes and conservatively managing context segments, achieving 61-87% cost reduction on benchmarks while maintaining competitive performance.

0 favorites 0 likes
#prompt-cache

@ClaudeDevs: With Opus 4.8, you can add system instructions mid-conversation without breaking the prompt cache. More cache hits mean…

X AI KOLs Following · 2026-05-29 Cached

Claude Opus 4.8 allows adding system instructions mid-conversation without breaking the prompt cache, reducing cost and latency for API requests.

0 favorites 0 likes
#prompt-cache

@Michaelzsguo: So you bought the 128GB MacBook Pro. Now the question is not, “Which local model gets the highest TPS?” It is: which se…

X AI KOLs Timeline · 2026-05-17 Cached

This thread recommends a local AI coding stack for the 128GB MacBook Pro, using Qwen 3.6 model with MLX server and specific configurations for reliable coding assistance.

0 favorites 0 likes
← Back to home

Submit Feedback