token-optimization

Tag

Cards List
#token-optimization

Token Optimization and Context Window Management in Multi-Agent AI Workflows

arXiv cs.CL ↗ · 2026-08-19 Cached

This paper explores techniques for token optimization and context window management in multi-agent AI workflows to improve efficiency and performance.

0 favorites 0 likes
#token-optimization

@Xudong07452910: Anthropic just wrote a very practical Claude Code usage guide. There's a detail I hadn't thought much about before: files read by Claude, tests run, terminal outputs—once they enter the current Session, they're basically carried forward in every subsequent round. So a Claude Code…

X AI KOLs Timeline ↗ · 2026-08-16 Cached

Anthropic published a practical guide for Claude Code, emphasizing context management through session commands and discussing token optimization and prompt caching to improve efficiency in AI-assisted coding.

0 favorites 0 likes
#token-optimization

@Saboo_Shubham_: Anthropic just published HOW to run Claude Code without burning tokens. Your prompt cache expires after an hour. Run /c…

X AI KOLs Timeline ↗ · 2026-08-15 Cached

Anthropic has published instructions on how to run Claude Code efficiently by using the /compact command before the prompt cache expires to save tokens.

0 favorites 0 likes
#token-optimization

Show HN: Mcptoon – MCP CLI client that cuts tool discovery tokens by 97%

Hacker News Top ↗ · 2026-08-11 Cached

mcptoon is a zero-dependency Python CLI client for MCP servers that uses a compact TOON format to reduce tool discovery tokens by up to 97%, saving context window space for AI agents.

0 favorites 0 likes
#token-optimization

@addyosmani: Tip: Periodically run /doctor in Claude Code or ask Codex to audit for unused skills, MCPs, context use. You've probabl…

X AI KOLs Timeline ↗ · 2026-08-10 Cached

Addy Osmani shares a tip to periodically run /doctor in Claude Code or ask Codex to audit for unused skills, MCPs, and context use to reduce token bloat.

0 favorites 0 likes
#token-optimization

Cut my agent’s tokens by 72% (11.9k ➝ 3.3k per task). Here’s exactly what I changed, with numbers

Reddit r/AI_Agents ↗ · 2026-08-10

A developer shares a detailed case study on reducing an AI agent's token consumption by 72% through system prompt reduction, tighter retrieval, tool output pruning, and other techniques, with minimal impact on success rate.

0 favorites 0 likes
#token-optimization

Same AI model. Better results. Lower cost.

Reddit r/AI_Agents ↗ · 2026-08-07

The author argues that AI coding harnesses and model routing are as important as the model itself, sharing tests with Oh-My-Pi and OpenCode that cut token usage and errors, and recommending tiered model subscriptions for high-volume lightweight tasks.

0 favorites 0 likes
#token-optimization

@alex_verem: You're burning 200,000 tokens every time you ask your agent about a book you already own. Dump a 400-page technical boo…

X AI KOLs Timeline ↗ · 2026-08-05 Cached

Book-to-skill is an open-source tool that converts technical books into structured skills for AI coding agents, reducing token usage by 24x-51x by loading only the needed chapters on demand.

0 favorites 0 likes
#token-optimization

@Saccc_c: AGENTS.md, summarized by Vercel engineers, can save you 80% of tokens in project development. The core is to reduce AI's ineffective iterations and token consumption through engineering decisions and solution design. Full version below, recommended to place in your project folder: # AGENTS.md - Not aimed at maintaining backward...

X AI KOLs Following ↗ · 2026-08-03 Cached

AGENTS.md summarized by Vercel engineers can save you 80% of tokens in project development. The core is to reduce AI's ineffective iterations and token consumption through engineering decisions and solution design.

0 favorites 0 likes
#token-optimization

Curated 11 token-optimization tools into one installer — what am I missing?

Reddit r/AI_Agents ↗ · 2026-07-24

A developer shares a curated installer that packages 11 token-optimization tools to help AI agents manage context and reduce token usage, asking the community for additional recommendations especially for multi-agent setups.

0 favorites 0 likes
#token-optimization

@heyshrutimishra: The most expensive part of running browser agents is not the task. It is everything before the task. Navigating, Findin…

X AI KOLs Timeline ↗ · 2026-07-24 Cached

Webcmd is an open-source tool that teaches browser agents to learn site navigation once and reuse it, dramatically reducing token waste on repetitive tasks. It supports 99 sites and runs locally.

0 favorites 0 likes
#token-optimization

I burned all my tokens researching how to save tokens

Hacker News Top ↗ · 2026-07-19 Cached

The author describes building a custom AI research pipeline using multiple subscriptions and cheaper models to reduce token costs, learning firsthand how to optimize token usage while researching tokenomics.

0 favorites 0 likes
#token-optimization

@yibie: Recommends this hardcore real-world test. An engineer tracked his coding agent session for a week and found that only 0.67% of tokens were spent on actual tasks—the remaining 99% all went to moving tool directories, skill descriptions, and system prompts. Work-to-overhead ratio 1:1…

X AI KOLs Timeline ↗ · 2026-07-15 Cached

An engineer tracked his coding agent's token usage over a week, finding that only 0.67% of tokens were spent on actual tasks, with 99% consumed by tool directories, skill descriptions, and system prompts. He provides optimization strategies, including shell output filtering which saved 46.9% of tokens.

0 favorites 0 likes
#token-optimization

@CycleDecoded: Ridiculous, guys—a wild trick just stormed GitHub trending, directly exploiting a bug in the large model billing system. This thing is called pxpipe (MIT license), a local open-source proxy tool specifically designed to counter Claude Code's billing shock. The principle is absolutely genius: large models charge text tokens by word count, but images are billed by fixed pixels. So it simply takes your long, bloated system prompts, code, and history logs, snaps them into a dense PNG image, and feeds it through the model's vision channel. This is a brutal move—a direct bypass of expensive text billing!

X AI KOLs Timeline ↗ · 2026-07-10 Cached

pxpipe is a local open-source proxy tool that reduces Claude Code bills by approximately 70% by rendering large amounts of text (such as system prompts, code, and logs) into PNG images and feeding them through the large model's vision channel. It exploits the fact that images are billed by pixel rather than by word count. The tool is perfectly compatible with the Fable 5 model, operates with clever automation, but uses lossy compression and is unsuitable for sensitive data.

0 favorites 0 likes
#token-optimization

Stripping terminal noise from agent context via a lazy-loaded local CLI layer. Looking for brutal feedback on this heuristic.

Reddit r/LocalLLaMA ↗ · 2026-07-09

Boost is a lazy-loaded local CLI layer that truncates verbose terminal output and replaces it with semantic markers to reduce token consumption for coding agents, caching logs locally for on-demand inspection. The tool, incubated within JFrog, runs entirely locally and seeks community feedback on its heuristic.

0 favorites 0 likes
#token-optimization

@noisyb0y1: CLAUDE CODE DEVELOPER MAKING $1.4M/YEAR SHOWED THE SYSTEM THAT SAVES 95% OF TOKENS one MD file into Claude Code - and t…

X AI KOLs Timeline ↗ · 2026-07-09 Cached

A developer claims a system saved 95% of tokens in Claude Code by providing a single Markdown file with secret Anthropic documentation, making the agent precise and boosting productivity 10x.

0 favorites 0 likes
#token-optimization

@FinanceYF5: ClaudeDevs shares a few patterns they often use with Fable 5: Treat Fable 5 as a “consultant.” Have the executor Sonnet 5 call Fable 5 for guidance. This way, most tokens are billed at the lower executor rate.

X AI KOLs Following ↗ · 2026-07-09 Cached

ClaudeDevs shares the pattern of using Fable 5 as a consultant, called by the executor Sonnet 5 to leverage lower billing rates and save on token costs.

0 favorites 0 likes
#token-optimization

@PrajwalTomar_: You're losing thousands of tokens every Claude Code session. The people who installed ONE free plugin aren't. The reaso…

X AI KOLs Following ↗ · 2026-07-07 Cached

A free GitHub plugin called Context Mode reduces token waste in Claude Code by up to 98% by sandboxing tool calls and only sending back essential output, saving session state and allowing Claude to resume where it left off.

0 favorites 0 likes
#token-optimization

@paul_cal: They actually tried some stuff themselves. Haven't audited but would be easy to point an agent at the repo & extend wit…

X AI KOLs Following ↗ · 2026-07-05 Cached

pxpipe is a local proxy that reduces Claude Code's token usage by converting bulky context (system prompts, tool docs, history) into compact images, achieving 59-70% cost savings with minimal accuracy loss. It exploits the token efficiency of images over text for dense content.

0 favorites 0 likes
#token-optimization

@Gorden_Sun: Superpowers 6.0 Release: Enhanced by Fable 5. Superpowers is a set of skills and instruction combinations that follow a complete software development methodology, suitable for various agents like CC and Codex. Version 6.0 is optimized with Fable 5, achieving a 50% increase in running speed and consuming…

X AI KOLs Timeline ↗ · 2026-07-04 Cached

Superpowers 6.0 is released, achieving a 50% speed increase and 60% token consumption reduction through Fable 5 optimization, while maintaining the same high-quality output.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback