Tag
According to Bloomberg, Harvey's costs for renting OpenAI and Anthropic models have risen sharply due to a 20-fold surge in token usage, causing gross margins to drop from about 50% to -50% by June.
The article discusses the trend of venture capital subsidizing token usage in consumer AI products, likening it to the '$5 Uber' model for long-term monetization and competitive dominance.
An AI system named 'jev' is showcased playing Super Smash Bros. against itself, controlling multiple characters and making rapid decisions within fractions of a second, using over 22 million tokens.
Jovan from UkisAI discusses improvements in their Swift Qwen3.8 27B model and seeks community feedback on creating a Swifted version of Bonsai 2 to address overthinking loops and high token usage.
A user reports spending $260,000 in API credits on Codex within 24 hours, utilizing 5.2 billion tokens with GPT 6 Astra and GPT 5.6 Sol models.
A Reddit post highlights the strong performance of the Qwen 3.8 27B model, with mention of 1.3B tokens used recently.
The tweet reports that GPT-6 Astra had the highest spending on OpenRouter last week, while GPT-5.6 Luna used the most tokens, leading by a wide margin.
The article reports that Astra is the top AI model by spending last week, while Luna leads in token usage, highlighting trends between Anthropic and OpenAI.
A developer is building a plugin for Devin Desktop and CLI, testing it with Devin, and sharing insights from extensive token consumption and lessons learned.
Xiaomi's MiMo-V2.5 large language model ranked top globally by monthly token usage on OpenRouter in July, driving AI integration into smart manufacturing and the Human x Car x Home strategy.
OpenRouter's weekly token usage has surged from 4.6 trillion a year ago to 113 trillion, with a doubling in the past month, indicating significant growth in the AI platform's adoption.
Elon Musk announced that all users of Grok Bot will receive another free reset on token usage.
The post highlights how discounting GPT 5.6 models on OpenRouter led to a 13.8x increase in token usage, illustrating the Jevons Paradox where efficiency gains result in higher total resource consumption.
OpenAI's discount program on Terra and Luna models led to a sharp increase in token usage, competitive displacement from other labs, and notable user retention after discounts ended.
The tweet provides a public service announcement advising users of the GLM-5.3 Flash AI model to use the 'high' reasoning_effort setting for better efficiency, as it achieves similar accuracy with significantly fewer tokens compared to the 'max' setting.
The article argues that with prompt caching, longer, stable prompts can be cheaper than frequently changing short ones, sharing insights from running AI agents with high cache hit rates.
Ox Alpha, based on GLM-5.3 Flash AA, is announced at one-hundredth the price of frontier models, powered by Chinese chips, and has achieved nearly 20% weekly token share on OpenRouter.
A video visualizes the top 20 AI models by weekly token usage on OpenRouter from December 2024 to August 2026, showing a shift from US to Chinese labs dominating the leaderboard.
This article introduces a cost-effective tool for accessing AI large models, ideal for emergency use by small companies, at a daily cost of only 80-100 USD or 20 RMB.
The author criticizes the common practice of including cached input tokens in discussions about token usage in AI models.