token-cost

Tag

Cards List
#token-cost

DeepSeek prefix caching hacks to cut token costs 90% and enable ads-supported agents

Reddit r/AI_Agents · 2026-08-12

A developer shares how they cut DeepSeek API costs by roughly 90% through prefix caching optimizations, reducing browser agent task costs below $0.005 and enabling a free, ads-supported browser agent.

0 favorites 0 likes
#token-cost

@hanakoxbt: your agent has thirty tools. it calls two of them. the other twenty eight are not sitting idle somewhere. they are in t…

X AI KOLs Following · 2026-08-08 Cached

Unused tools in an AI agent's toolset still consume tokens and add noise to tool selection, so agents should load only the tools required for the current task.

0 favorites 0 likes
#token-cost

Tested whether my coding CLI actually reads AGENTS.md.

Reddit r/AI_Agents · 2026-08-07

An experiment testing whether a coding CLI actually reads AGENTS.md found the file was silently ignored, and even when read, bloated instruction files increased token costs and failed to improve performance. The author recommends writing only what the model cannot infer from code.

0 favorites 0 likes
#token-cost

I measured what 13 search APIs actually cost to run inside an agent. The pricing page is the smaller half of the bill

Reddit r/AI_Agents · 2026-08-06

A developer benchmarks 13 search API configurations inside an AI agent, revealing that hidden token costs from reading payloads can dominate the total bill and vary by up to 67x across providers.

0 favorites 0 likes
#token-cost

@PrajwalTomar_: Wait this is actually INSANE. Most of your Claude Code bill is you paying to re-read the same code. Every edit, it read…

X AI KOLs Timeline · 2026-08-05 Cached

A tweet points out that most Claude Code token costs come from re-reading the entire codebase on each edit, and highlights a new category of tools that use codebase mapping, context pruning, and project memory to drastically cut token usage, with a list of 10 tools in the linked article.

0 favorites 0 likes
#token-cost

@omarsar0: Finally a good paper testing whether self-reflection loops are worth it. Setup: Seven methods, open models at 1.5B, 3B …

X AI KOLs Timeline · 2026-08-04 Cached

A paper finds that self-reflection loops like Self-Refine and Reflexion do not beat repeated sampling at equal token cost across open models from 1.5B to 7B, with all 18 self-inspection comparisons negative.

0 favorites 0 likes
#token-cost

@Saccc_c: Someone has open-sourced the Codex infinite ammo method. The core remains that Sol is responsible for orchestration and planning, while Luna max handles the actual execution. This method saves about 59% of token cost on simple tasks, and reaches 65% on complex tasks.

X AI KOLs Following · 2026-08-03 Cached

Someone has open-sourced the Codex infinite ammo method, using Sol for planning and Luna Max for execution, saving about 59% of token costs on simple tasks and 65% on complex tasks.

0 favorites 0 likes
#token-cost

Would you use a circuit breaker for AI agents?

Reddit r/LocalLLaMA · 2026-07-30

The author discusses the problem of AI agents getting stuck in tool loops and proposes building 'Moven AI', an open-source SDK that acts as a circuit breaker to abort runaway agents before they incur excessive costs.

0 favorites 0 likes
#token-cost

Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

Hacker News Top · 2026-07-29 Cached

An updated benchmark shows self-hosting Kimi K3 on 8×B300 nodes achieves 86.4% task resolution at roughly 20% higher hardware cost compared to GLM-5.2 on B200 nodes, though with lower throughput.

0 favorites 0 likes
#token-cost

My Codex run burned 100,000+ tokens over 8 hours and still handed me an unfinished mess... so I built something that just finishes

Reddit r/AI_Agents · 2026-07-28

After a Codex run wasted over 100,000 tokens without completing a task, the author built a tool that reliably finishes code generation jobs.

0 favorites 0 likes
#token-cost

@PyTorch: Open Source Amplifies the Full-Stack Advantage to Power the Lowest Token Cost. PyTorch is a leading example: Launched i…

X AI KOLs Timeline · 2026-07-21 Cached

NVIDIA details how its full-stack inference software, co-developed with open-source ecosystems like PyTorch, reduces token costs by up to 5x on Blackwell GPUs, with real-world deployments from Baseten, Cognition, Deep Infra, and others demonstrating performance gains.

0 favorites 0 likes
#token-cost

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

NVIDIA Blog · 2026-07-21 Cached

NVIDIA announces Vera Rubin platform, delivering 10x more throughput per megawatt than Blackwell and claiming lowest token cost for AI factories through extreme co-design across seven chips and five rack trays.

0 favorites 0 likes
#token-cost

@FinanceYF5: Chamath said he asked the CTO about the company's AI token spending today and got a shocking answer: token costs double every 45 days, but downstream productivity improves by at most 5%. His exact words were that costs double, benefits stay roughly flat, and the team found that to iterate to the next generation of capabilities, the amount of tokens required grows exponentially, because...

X AI KOLs Timeline · 2026-07-14 Cached

Chamath revealed that his company found AI token costs double every 45 days while downstream productivity improves by at most 5%. He believes AI development is approaching a bottleneck and suggests companies reconsider their strategies or even consider exiting.

0 favorites 0 likes
#token-cost

@0xCheshire: Chamath just revealed a deeply unsettling truth for the AI industry. He asked his own CTO to review the company's spending, and the result was staggering: "Our token costs are doubling every 45 days," yet the downstream productivity gains are at most about 5%. Costs are skyrocketing exponentially, while returns remain basically flat...

X AI KOLs Timeline · 2026-07-13 Cached

Chamath reveals the harsh reality of AI costs vs. returns: token costs double every 45 days, but downstream productivity gains are at most 5%. Large model capability improvement has hit an asymptote, and within the next 3-4 years, every company will face an ultimate reckoning between cost and benefit.

0 favorites 0 likes
#token-cost

Will Meta start a token price war and drive down API pricing across the AI industry?

Reddit r/ArtificialInteligence · 2026-07-11

Meta's aggressive pricing for its Muse Spark 1.1 model undercuts competitors like OpenAI and Anthropic, potentially sparking a token price war that could lower costs, prevent a duopoly, and spur innovation in the application layer.

0 favorites 0 likes
#token-cost

@FinanceYF5: Oh my god... Fable 5 is back, and it's insanely powerful. Someone asked Fable to make a game called 'Super Smart Racing'... With just 4 prompts and $173 worth of tokens, Fable 5 created this game. (Prompts below)

X AI KOLs Timeline · 2026-07-02 Cached

Fable 5 model only used 4 prompts and $173 worth of tokens to create a game called 'Super Smart Racing', demonstrating its extremely strong generative capabilities.

0 favorites 0 likes
#token-cost

How NVIDIA’s Inference Software Stack Powers the Lowest Token Cost

NVIDIA Blog · 2026-06-30 Cached

NVIDIA's full-stack inference software, codesigned with hardware, has reduced token costs by up to 5x on the Blackwell platform in just one month, enabling lower cost per token for AI factories. Companies like Baseten, Cognition, Deep Infra, and Together AI are using the stack to optimize inference performance.

0 favorites 0 likes
#token-cost

Orbital AI datacenter token costs x8-x12 of Earth one

Reddit r/ArtificialInteligence · 2026-06-17

A report indicates that operating an AI datacenter in orbit costs 8 to 12 times more per token than a terrestrial datacenter, highlighting significant cost barriers for space-based AI computation.

0 favorites 0 likes
#token-cost

What is the real cost of a token and token futures market

Reddit r/ArtificialInteligence · 2026-06-17 Cached

Bellwethr is developing an open methodology for tracking the real USD cost of a single inference token from capable models, with a draft benchmark suite and community contributions underway.

0 favorites 0 likes
#token-cost

@mattpocockuk: Announcing mattpocock/skills v1 - Achieved a 63% reduction in token cost for skill descriptions - Split skills into mod…

X AI KOLs Following · 2026-06-17 Cached

Announcing version 1 of mattpocock/skills, a collection of AI skill definitions that reduces token costs by 63% and introduces new skills for codebase design, domain modeling, and more.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback