I think “use fewer tokens” is too shallow as LLM cost advice
Summary
This article argues that common LLM cost advice focusing on token reduction is too shallow, and that the more impactful strategy in production is to route different workflow steps to different models rather than using a single default model.
Similar Articles
LLM Routing is not the problem to solve; token efficiency is
The article argues that model routing isn't the real problem to solve—token efficiency is. It advocates for a holistic closed-loop approach combining cheaper defaults, preference-aware routing, and better caching (e.g., Coinbase's 5%→60% cache hit improvement) to maximize useful intelligence per dollar.
We optimize LLM costs before we ask what the AI is for
An AI consultant reflects on how teams optimize LLM costs without questioning whether the task needs a model at all, and advocates measuring cost per successful outcome rather than per token.
What I'm Finding About LLM Code Style and Token Costs
The article discusses how LLM code style choices affect token consumption and costs, offering optimizations such as using Web API standards and simpler indentation to reduce output tokens.
Using LLMs to measure what LLMs cost and why smaller models aren't always cheaper
Using smaller, cheaper LLMs can increase total workflow costs due to hidden expenses like review time and error correction, highlighting the need for comprehensive cost tracking.
Could better human–LLM coordination reduce token costs without changing the model?
The paper explores whether improved coordination between humans and large language models can reduce token costs without modifying the model itself.