I think “use fewer tokens” is too shallow as LLM cost advice

Reddit r/AI_Agents News

Summary

This article argues that common LLM cost advice focusing on token reduction is too shallow, and that the more impactful strategy in production is to route different workflow steps to different models rather than using a single default model.

A lot of LLM cost advice seems to stop at prompt compression, caching, or token limits. But in production workflows, I suspect the bigger issue is model choice. Example: - classify ticket intent - summarize context - retrieve docs - draft reply - final high-risk response Those steps probably should not all use the same model. For teams running AI agents or RAG in production: are you routing different steps to different models, or still using one default model everywhere?
Original Article

Similar Articles

LLM Routing is not the problem to solve; token efficiency is

Reddit r/AI_Agents

The article argues that model routing isn't the real problem to solve—token efficiency is. It advocates for a holistic closed-loop approach combining cheaper defaults, preference-aware routing, and better caching (e.g., Coinbase's 5%→60% cache hit improvement) to maximize useful intelligence per dollar.