Taught Claude to talk like a caveman to use 75% less tokens.
Summary
A user experimented with prompting Claude to communicate concisely, resulting in a 75% reduction in token usage while monitoring potential impacts on model intelligence.
Similar Articles
Companies Are Making Claude and Codex Talk Like Cavemen to Stop AI’s Soaring Costs
Companies are adopting a plugin called 'Caveman' that forces AI models like Claude and Codex to speak in terse, caveman-like language to reduce token consumption and curb soaring AI costs. The tool can cut output tokens by up to 75%, and is being used by employees at OpenAI, Nvidia, GitHub, and Legrand.
@_avichawla: Claude Code used 3x fewer tokens with one change: - Before: 10.4M tokens · 10 errors · $9.21 - After: 3.7M tokens · 0 e…
By swapping to Insforge Skills + CLI as the backend context layer, a user cut Claude Code token usage by 64 %, eliminated all errors and reduced cost from $9.21 to $2.81.
@_avichawla: A smarter Claude model burns more tokens, not fewer! And it's not a minor 3-5% difference. But 54% higher token usage. …
The article analyzes why smarter AI agents like Claude consume more tokens when interacting with human-centric backends like Supabase due to inefficient context discovery. It introduces InsForge, an open-source backend tool designed for agents that provides structured context to significantly reduce token usage and manual interventions.
Does Speaking to Agents Like Cavemen Save 65% of Tokens? We Test
JetBrains benchmarked the Caveman token-compression skill on Claude Code across 86 tasks, finding real output-token savings of about 8.5% (not the advertised 65%) with no detectable degradation in task quality.
Cut my Claude Code token burn by 30-40% — the stack that's actually real (2026)
Sharing a practical stack to reduce Claude Code token usage by 30-40%, focusing on real-world efficiency gains for AI coding.