Tag
The paper proposes a causal graph divergence framework to assess the faithfulness of LLM pricing agents in oligopolistic markets, revealing that chain-of-thought monitoring fails to detect collusion due to dissociation between collusive behavior and reasoning traces.
Drew Breunig discusses how the release of the Fable model shifted coding strategies by making cost a significant factor, ending the era of easily improved coding harnesses due to model advancements.
Nous Research announces that Ox Alpha is free via Nous Portal for a limited time, with the portal offering access to numerous AI models at varying token prices.
A blog post explaining that cache read costs dominate LLM inference spending for agentic workloads, with cumulative costs growing quadratically as context is re-read each turn, and advice to reduce tool call count to cut costs.
Compares API list prices of 18 LLMs from major providers, highlighting a 100x cost difference for the same workload and recommending model routing for cost efficiency.
A price tracker found that GLM-5.2 and Tencent's Hy3 quietly changed prices multiple times in a week, highlighting volatility in Chinese LLM pricing.
Anthropic is renting GPUs from xAI's Colossus cluster for inference as token consumption grows exponentially, highlighting a token shortage that is driving up costs and pressuring AI companies' margins.
A developer shares their experience with AI inference costs after switching from subsidized OpenAI Codex to OpenRouter, prompting a discussion about the sustainability of current LLM pricing models and the potential shift towards open-source self-hosting.
This article provides a comprehensive 2026 guide to free and low-cost large language models, comparing domestic (China) and international options.