@omarsar0: Switching models mid-session is one of the fastest ways to inflate an agent bill. Every switch resets your prompt cache…

X AI KOLs Timeline Tools

Summary

An open-weight auto-routing model called @daridotdev is released for coding agents, offering cache-aware model switching to reduce costs by up to 70% while maintaining performance.

Switching models mid-session is one of the fastest ways to inflate an agent bill. Every switch resets your prompt cache, and a cheaper model can end up costing you more. @daridotdev built their router to be cache-aware. It only switches when the move actually saves money, and it plugs into Claude Code, Codex, or whatever harness you already use. The router itself is an open-weight SLM.
Original Article
View Cached Full Text

Cached at: 07/25/26, 03:59 AM

Switching models mid-session is one of the fastest ways to inflate an agent bill.

Every switch resets your prompt cache, and a cheaper model can end up costing you more.

@daridotdev built their router to be cache-aware. It only switches when the move actually saves money, and it plugs into Claude Code, Codex, or whatever harness you already use.

The router itself is an open-weight SLM.

Avyay Varadarajan (@avyvar): Today, we’re releasing our open-weight, auto-routing model @daridotdev, built for coding agents.

We’re state-of-the-art on the Pareto Frontier, w/ 70% cost reduction + comparable coding performance to Fable.

Bring your own evals, choose your models, or use our defaults.

Similar Articles