@rohanpaul_ai: Not Diamond just released the methodology behind their model routing which gets Opus xhigh quality while cutting agent …
Summary
Not Diamond released a methodology for model routing that achieves Opus-level quality while reducing agent costs by 20–80%, using a sequential decision approach to handle long-running coding agents effectively.
View Cached Full Text
Cached at: 09/02/26, 03:50 AM
Not Diamond just released the methodology behind their model routing which gets Opus xhigh quality while cutting agent costs 20–80%.
Model routing becomes much harder decision for long-running coding agents. So Not Diamond is benchmarking the router across simulated user turns, changing task complexity, and response delays that can expire the KV cache.
And harder than it seems, switch models too aggressively and you can throw away the KV cache. Pick one model from the opening prompt and you are assuming task complexity stays fixed for the rest of the session.
So Not Diamond is treating routing as a sequential decision problem instead.
At each step, its router predicts future reward and cost for a particular model and reasoning effort, using current and previous session state, message and token counts, task complexity, KV-cache state, and intermediate reward signals.
Read more detail on their technical report.
Tomas Hernando Kofman (@tomas_hk): Today we’re releasing our methodology for evaluating model routing with interactive benchmarks, which represent agent cost accumulation better than static benchmarks do.
Across leading benchmarks, we achieve Pareto-dominance, exceeding Opus xhigh quality at 20–80% lower cost.
Similar Articles
@rohanpaul_ai: Read more on their official blog.
Not Diamond announces Not Diamond Code, an intelligent model router for coding agents that selects the best model and reasoning effort per step, claiming 20%+ cost savings and Pareto-optimal performance on coding benchmarks.
@julien_c: Model routing is the missing layer of the multi-model stack. open models keep getting better at coding, routing helps t…
Not Diamond Code is announced as an intelligent model router for long-horizon coding agents, selecting the best model and reasoning effort per step to reduce costs by 20-65%.
@tomas_hk: Today we’re releasing our methodology for evaluating model routing with interactive benchmarks, which represent agent c…
Releasing a methodology for evaluating model routing with interactive benchmarks that achieve Pareto-dominance over leading benchmarks, offering higher quality at lower cost.
We stopped sending every AI agent request to Claude Opus 5. The results surprised us.
A team benchmarked routing different stages of an AI agent workflow to different models versus sending every request to Claude Opus 5 across 89 Terminal-Bench 2.1 tasks, and found surprising results.
my agent bill went from $200 a week to $40 when I stopped running Opus on every subtask
A developer shares how they reduced their AI agent's weekly cost from $200 to $40 by routing simple subtasks to cheaper models like DeepSeek V4 Pro and Tencent Hunyuan while keeping complex reasoning on Opus 4.7, achieving comparable output quality for most work.