Save Money Automagically by Auto Routing LLM Choice

Reddit r/AI_Agents Tools

Summary

The author built an automated system to route LLM choices based on cost and performance data, aiming to optimize AI spending, with plans to open-source the tool.

AI is both plummeting in cost (when intelligence is held steady) and exploding in cost because we use AI more and more. Cost control and quality control are essential. A management truism is: You can't manage what you don't measure. And while I've had my gripes about this in the past (just because something can be measured doesn't mean it's the right metric) - it's long past time for effective management of AI costs. First is to measure the right things. Cost per token is a bad measurement. Cost per successful task is a much more beneficial measurement. I created this measurement two ways. I built a bespoke benchmark that uses my real coding data and ran it based on roles: planner, supervisor, coder and code review. I measure against models and harnesses. Reading reviews online is helpful, but nothing is as good as metrics coming from how YOU and your organization develop code. One thing I've learned is at the current moment, running the SAME model in Pi verses Codex, cuts the cost and time more than 1/3rd. That may change over time, which is why I built my own benchmark to keep track as new models and updated harnesses come out. The second measurement is ongoing. Every time I run code, I now keep track of the model, the harness, the time elapsed, success or failure. Now I have data to make future decisions. Getting to this point was a huge upgrade. The next is even better - live model routing optimized for cost and performance. I minimize expected cost of an accepted outcome subject to quality, independence, capacity, and time constraints. I don't have to manually set or guess at the model or harness. It's now all data driven and live. All optimized with actual data that considers costs per successful task. I have subscriptions to OpenAI, Claude, Gemini, OpenCode Go -- and I try others out over time. My router understands that the cost of a model using a subscription is a lot cheaper than that of the public API cost. I have a formula that is "playing horse shoes not darts" accurate. And I make sure not to route to plans that have 10% or less available. If all my subs are outstripped, then I call the most cost effective model by pay as you go API. The cost of models like DeepSeek V4.1 Flash are so cheap, they are practically free. You don't need subscriptions. But can they do everything? No. But why not start with them, and escalate up to a more powerful model if they fail? Make that automagic and you reap great cost savings. And do so based on YOUR data and actual usage. And if I have plenty of OpenAi subscription available, then Luna xHigh is actually cheaper for me than DeeepSeek. Right now my system is highly integrated to my AI Employee Factory platform, but I intend to break it out and open-source it. I had $100/mo subs to OpenAI, Anthropic and Gemini. At that, I was running out of both OpenAI and Anthropic. Plenty of Gemini but they are lacking a Fable or Astra/Sol quality model. The cheap flash models NEED quality planning, supervision and code review in order to be effective. Moved to $200/mo OpenAI as it gives 4x the $100/mo plan. Downgraded both Anthropic and Gemini to the $20/mo plans. I've more than doubled my usage as OpenAI gives you more dollar per dollar than Anthropic. And if I need even more, I can use retail DeepSeek V4.1 Flash or GLM 5.3 Flash which are more cost effective and better than Gemini even on subscription. I will miss having as much Fable 5.1 but Astra is even better at most (but not all) things. These real life data driven decisions will change as the landscape changes. But no longer am I going by feel or latest x/Twitter post.
Original Article

Similar Articles

A client paid me to rip the AI out of the tool I built them.

Reddit r/AI_Agents

A developer built an LLM-powered ticket routing tool, but the support team distrusted the black-box decisions. The client paid to replace the LLM with a simple rules engine, resulting in higher accuracy, lower costs, and greater user trust.