Tag
A video visualizes the top 20 AI models by weekly token usage on OpenRouter from December 2024 to August 2026, showing a shift from US to Chinese labs dominating the leaderboard.
This article introduces a cost-effective tool for accessing AI large models, ideal for emergency use by small companies, at a daily cost of only 80-100 USD or 20 RMB.
The author criticizes the common practice of including cached input tokens in discussions about token usage in AI models.
The article discusses the finding that AI agents are currently utilizing five times more tokens than humans, indicating a notable trend in AI capabilities and efficiency.
The article describes a scenario where an LLM agent exhausts its weekly token quota on excessive testing without producing actual code, highlighting the concept of 'vibe tax' on software developers due to AI overengineering.
The post argues that open models' token share will rise due to cost-effectiveness in multi-agent systems, referencing data showing open-source AI gaining market share from OpenAI and Anthropic.
Open weight model usage on Vercel AI Gateway reached a record 62% share of tokens, up from 28.4% two months ago, signaling growing adoption as enterprise use develops.
A developer describes turning Grok Bot into a virtual CTO to automate GitHub repo management with cloud agents, noting high token consumption and limitations.
A firsthand report detailing the costs incurred from running AI agents in production for one week, emphasizing that token usage dominates expenses and sharing insights on cost management like integer math for billing.
The author has created an open-source CLI tool that live monitors token usage and API costs per prompt for AI agents, and is seeking interest to share it.
Shruti Mishra highlights the explosive growth in AI token usage—from an OpenAI employee using 100k tokens/month six years ago to a worldwide average of 100k/month today—and argues that demand for cheap, high-quality AI intelligence is effectively uncapped.
The user found that Codex's auto-review mode calls an independent agent for approval, causing fast token consumption, and is planning to switch to full-access mode.
A technical deep-dive that records and analyzes the exact HTTP requests Codex CLI sends to the model, measuring token counts and how instructions, tools, and context are bundled.
GPT-5.6 Sol uses more than twice the tokens per session compared to GPT-5.5 in Codex workflows, leading to higher costs and faster depletion of subscription quotas.
Initial tests of DeepSeek v4 Flash show notable gains in UI/UX design capabilities, though the model remains token-hungry.
ccusage is an open-source command-line tool that helps developers inspect token usage and costs from local coding-agent CLI data, offering daily/weekly/monthly reports, model breakdowns, and JSON export.
The author criticizes Claude Code's increasingly large system prompt (32k tokens) for degrading cost, latency, and performance, and praises Pi's minimalist 1k-token approach with plugins as a better philosophy for coding agents.
The US Army has exhausted its annual AI token allocation from Ask Sage within a month after encouraging widespread use, forcing it to reinstate limits and raising questions about the sustainability of generative AI adoption in the Department of Defense.
El CEO de Anthropic observa que una Skill gratuita de GitHub reduce el uso de tokens de Claude Code en un 90%, mientras los usuarios aún pagan por Max. La herramienta Ponytail optimiza el código generado para reducir costos.
A study comparing Claude Code and OpenCode reveals that Claude Code sends 33k tokens before reading the prompt while OpenCode sends only 7k, highlighting significant inefficiency in Claude Code's cache strategy and token usage.