Tag
This paper explores techniques for token optimization and context window management in multi-agent AI workflows to improve efficiency and performance.
Anthropic published a practical guide for Claude Code, emphasizing context management through session commands and discussing token optimization and prompt caching to improve efficiency in AI-assisted coding.
Anthropic has published instructions on how to run Claude Code efficiently by using the /compact command before the prompt cache expires to save tokens.
mcptoon is a zero-dependency Python CLI client for MCP servers that uses a compact TOON format to reduce tool discovery tokens by up to 97%, saving context window space for AI agents.
Addy Osmani shares a tip to periodically run /doctor in Claude Code or ask Codex to audit for unused skills, MCPs, and context use to reduce token bloat.
A developer shares a detailed case study on reducing an AI agent's token consumption by 72% through system prompt reduction, tighter retrieval, tool output pruning, and other techniques, with minimal impact on success rate.
The author argues that AI coding harnesses and model routing are as important as the model itself, sharing tests with Oh-My-Pi and OpenCode that cut token usage and errors, and recommending tiered model subscriptions for high-volume lightweight tasks.
Book-to-skill is an open-source tool that converts technical books into structured skills for AI coding agents, reducing token usage by 24x-51x by loading only the needed chapters on demand.
AGENTS.md summarized by Vercel engineers can save you 80% of tokens in project development. The core is to reduce AI's ineffective iterations and token consumption through engineering decisions and solution design.
A developer shares a curated installer that packages 11 token-optimization tools to help AI agents manage context and reduce token usage, asking the community for additional recommendations especially for multi-agent setups.
Webcmd is an open-source tool that teaches browser agents to learn site navigation once and reuse it, dramatically reducing token waste on repetitive tasks. It supports 99 sites and runs locally.
The author describes building a custom AI research pipeline using multiple subscriptions and cheaper models to reduce token costs, learning firsthand how to optimize token usage while researching tokenomics.
An engineer tracked his coding agent's token usage over a week, finding that only 0.67% of tokens were spent on actual tasks, with 99% consumed by tool directories, skill descriptions, and system prompts. He provides optimization strategies, including shell output filtering which saved 46.9% of tokens.
pxpipe is a local open-source proxy tool that reduces Claude Code bills by approximately 70% by rendering large amounts of text (such as system prompts, code, and logs) into PNG images and feeding them through the large model's vision channel. It exploits the fact that images are billed by pixel rather than by word count. The tool is perfectly compatible with the Fable 5 model, operates with clever automation, but uses lossy compression and is unsuitable for sensitive data.
Boost is a lazy-loaded local CLI layer that truncates verbose terminal output and replaces it with semantic markers to reduce token consumption for coding agents, caching logs locally for on-demand inspection. The tool, incubated within JFrog, runs entirely locally and seeks community feedback on its heuristic.
A developer claims a system saved 95% of tokens in Claude Code by providing a single Markdown file with secret Anthropic documentation, making the agent precise and boosting productivity 10x.
ClaudeDevs shares the pattern of using Fable 5 as a consultant, called by the executor Sonnet 5 to leverage lower billing rates and save on token costs.
A free GitHub plugin called Context Mode reduces token waste in Claude Code by up to 98% by sandboxing tool calls and only sending back essential output, saving session state and allowing Claude to resume where it left off.
pxpipe is a local proxy that reduces Claude Code's token usage by converting bulky context (system prompts, tool docs, history) into compact images, achieving 59-70% cost savings with minimal accuracy loss. It exploits the token efficiency of images over text for dense content.
Superpowers 6.0 is released, achieving a 50% speed increase and 60% token consumption reduction through Fable 5 optimization, while maintaining the same high-quality output.