@rohanpaul_ai: Not Diamond just released the methodology behind their model routing which gets Opus xhigh quality while cutting agent …

X AI KOLs Timeline Tools

Summary

Not Diamond released a methodology for model routing that achieves Opus-level quality while reducing agent costs by 20–80%, using a sequential decision approach to handle long-running coding agents effectively.

Not Diamond just released the methodology behind their model routing which gets Opus xhigh quality while cutting agent costs 20–80%. Model routing becomes much harder decision for long-running coding agents. So Not Diamond is benchmarking the router across simulated user turns, changing task complexity, and response delays that can expire the KV cache. And harder than it seems, switch models too aggressively and you can throw away the KV cache. Pick one model from the opening prompt and you are assuming task complexity stays fixed for the rest of the session. So Not Diamond is treating routing as a sequential decision problem instead. At each step, its router predicts future reward and cost for a particular model and reasoning effort, using current and previous session state, message and token counts, task complexity, KV-cache state, and intermediate reward signals. Read more detail on their technical report.
Original Article
View Cached Full Text

Cached at: 09/02/26, 03:50 AM

Not Diamond just released the methodology behind their model routing which gets Opus xhigh quality while cutting agent costs 20–80%.

Model routing becomes much harder decision for long-running coding agents. So Not Diamond is benchmarking the router across simulated user turns, changing task complexity, and response delays that can expire the KV cache.

And harder than it seems, switch models too aggressively and you can throw away the KV cache. Pick one model from the opening prompt and you are assuming task complexity stays fixed for the rest of the session.

So Not Diamond is treating routing as a sequential decision problem instead.

At each step, its router predicts future reward and cost for a particular model and reasoning effort, using current and previous session state, message and token counts, task complexity, KV-cache state, and intermediate reward signals.

Read more detail on their technical report.

Tomas Hernando Kofman (@tomas_hk): Today we’re releasing our methodology for evaluating model routing with interactive benchmarks, which represent agent cost accumulation better than static benchmarks do.

Across leading benchmarks, we achieve Pareto-dominance, exceeding Opus xhigh quality at 20–80% lower cost.

Similar Articles

@rohanpaul_ai: Read more on their official blog.

X AI KOLs Following

Not Diamond announces Not Diamond Code, an intelligent model router for coding agents that selects the best model and reasoning effort per step, claiming 20%+ cost savings and Pareto-optimal performance on coding benchmarks.