token-efficiency

Tag

Cards List
#token-efficiency

@ericzakariasson: here's a prompt to improve your agent harness based on what we've learned at cursor. enjoy # Improve this agent harness…

X AI KOLs Timeline · 16h ago Cached

This article shares a prompt and practical guidelines for improving the token efficiency of LLM agent harnesses, based on lessons learned at Cursor, aiming to reduce costs without sacrificing task quality.

0 favorites 0 likes
#token-efficiency

@PrajwalTomar_: Most browser agents are burning your tokens on mistakes they already made. Your browser agent has ZERO memory so every …

X AI KOLs Following · 23h ago Cached

WebCMD provides memory for browser agents like Chrome, saving site paths to avoid repeating mistakes and reducing token waste. It was tested on Reddit and ranked as the most accurate and cheapest per task in BU Bench V1.

0 favorites 0 likes
#token-efficiency

@trq212: Opus 5.5 is the result of your feedback. It communicates clearly, it's cheaper per token than Opus 5.0 with the intelli…

X AI KOLs Timeline · yesterday Cached

Introducing Claude Opus 5.5, the first model in the Claude 5.5 family, which performs at the level of Claude Fable 5.1 but costs 40% less to run with increased rate limits.

0 favorites 0 likes
#token-efficiency

Bigger context windows just give you a bigger dead zone in the middle

Reddit r/AI_Agents · yesterday

An analysis of 847 AI agent runs reveals that larger context windows cause performance drops due to attention cliffs, and Synap is presented as a tool to efficiently manage context and reduce token usage.

0 favorites 0 likes
#token-efficiency

@aigclink: NVIDIA's open-source agent harness focused on token efficiency: SoL-Pi, 35–64% fewer tokens than the native harness, 50…

X AI KOLs Timeline · 5d ago Cached

NVIDIA's SoL-Pi is an open-source agent harness that reduces token usage by 35-64% and API costs by 50-54% by automating the inspection and refactoring of AI workflows for efficiency.

0 favorites 0 likes
#token-efficiency

@omarsar0: Build your own harness, folks. This is absolute banger paper from NVIDIA on self-evolving agent harnesses. (bookmark it…

X AI KOLs Timeline · 5d ago Cached

The paper introduces SoL-Pi, an automated harness discovery method that reduces token usage by nearly half and cuts API costs by about a third while maintaining performance on benchmarks using models like GPT-5.6 Sol and Opus 5.

0 favorites 0 likes
#token-efficiency

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

Hugging Face Daily Papers · 2026-09-17 Cached

SoL-Pi introduces a method for recursively scaling auto-research loops in coding agents, achieving significant token and cost reductions while maintaining performance on benchmarks.

0 favorites 0 likes
#token-efficiency

@composio: We ran GPT-6 Astra across 6 agent harnesses (Codex, Claude Code, OpenCode, Hermes Agent, Pi Agent, Command Code) on 29 …

X AI KOLs Timeline · 2026-09-16 Cached

The post describes an evaluation of GPT-6 Astra across six agent harnesses on 29 agentic tasks, showing similar success rates but token usage varying by 3–5x on failures.

0 favorites 0 likes
#token-efficiency

Token Efficient Task Execution via Application Behavior Modeling for Web Agents

arXiv cs.AI · 2026-09-15 Cached

OdoBot is a novel web-agent architecture that uses application behavior modeling to reduce token consumption and improve task success rates, outperforming agents like Agent-E and WebVoyager on the Canvas LMS.

0 favorites 0 likes
#token-efficiency

@Greptime: Evaluate the query interface alongside the RCA model. In Agent RCA Bench, six models investigated the same 14 incidents…

X AI KOLs Timeline · 2026-09-14 Cached

Agent RCA Bench evaluation found that GreptimeDB's SQL/PromQL interface produced 40% fewer wrong diagnoses and used ~48% fewer input tokens than Prometheus/Loki/Tempo wrappers across six AI models.

0 favorites 0 likes
#token-efficiency

Internal Documentation: Written for humans, agents, or both?

Reddit r/AI_Agents · 2026-09-14

The article discusses whether internal documentation should be written for human readability or optimized for AI agent efficiency, exploring trade-offs and advocating for context-specific formatting.

0 favorites 0 likes
#token-efficiency

@cjzafir: In the last 17 days, I burned $6k worth of tokens on my Claude 20x and Codex 20x subs. But I got so much work done by j…

X AI KOLs Following · 2026-09-12 Cached

The author open-sources a 'lean-thinking' skill for AI coding agents that reduces token usage by up to 35% while maintaining quality, based on personal testing with Claude and Codex models.

0 favorites 0 likes
#token-efficiency

Eggshell

Product Hunt · 2026-09-12 Cached

Eggshell is a local memory tool for AI agents that stores work results and evidence to reduce repeated investigation and save tokens, without requiring LLM calls for memory organization.

0 favorites 0 likes
#token-efficiency

What’s the most token efficient web search API in 2026? I measured token counts across 4 tools

Reddit r/AI_Agents · 2026-09-10

This article measures and compares the token efficiency of four web search APIs—Brave Search, Tavily, Exa, and Firecrawl—for AI agent contexts, concluding that Firecrawl is the most efficient for reducing context window bloat.

0 favorites 0 likes
#token-efficiency

@Leechael: The new generation of models should no longer use subagents. Subagents aren't as effective as having sessions exchange …

X AI KOLs Following · 2026-09-09 Cached

The article discusses the inefficiency of using subagents in AI models, suggesting that direct message exchange between sessions might be more effective, based on observations about GPT-6 Astra's token usage in Codex.

0 favorites 0 likes
#token-efficiency

@j_dekoninck: GPT-6-Astra takes first place on MathArena, with a massive 90% expected performance! It is also very token efficient: o…

X AI KOLs Timeline · 2026-09-05 Cached

GPT-6-Astra takes first place on MathArena with a 90% expected performance and high token efficiency.

0 favorites 0 likes
#token-efficiency

Gemini's 88% video token cut landed on 3.7 Flash, not 3.8

Reddit r/ArtificialInteligence · 2026-09-03

Google's update to Gemini 3.7 Flash introduces video token reduction by up to 88%, lowering costs significantly, though accuracy improvements are questionable.

0 favorites 0 likes
#token-efficiency

@danshipper: BREAKING: Anthropic just dropped Fable 5.1—and CLAUDE IS SO BACK. We’ve spent the last week testing it at @every across…

X AI KOLs Following · 2026-09-01 Cached

Anthropic has released Fable 5.1, an updated Claude model that excels in coding, writing, and knowledge work, with improved speed, token efficiency, and business readiness.

0 favorites 0 likes
#token-efficiency

@_philschmid: Gemini video understanding is now agentic. Gemini can now iteratively navigate video timelines, decide watch what, pick…

X AI KOLs Following · 2026-09-01 Cached

Gemini now offers agentic video understanding, enabling iterative navigation and adaptive processing to reduce token usage and costs while boosting accuracy in video analysis.

0 favorites 0 likes
#token-efficiency

Introducing agentic video understanding with Gemini

Google DeepMind Blog · 2026-09-01 Cached

Google DeepMind introduces agentic video understanding for Gemini models, reducing token consumption by up to 88% and improving accuracy in video analysis.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback