@mckaywrigley: bullish model routers. same perf at lower cost is obvious. but there are massive gains to be had by creating "smoother"…
Summary
A tweet highlights the promise of model routers and 'model melding', referencing the launch of Not Diamond Code, an intelligent model router for coding agents that cuts costs by 20-65%.
View Cached Full Text
Cached at: 08/04/26, 10:17 PM
bullish model routers.
same perf at lower cost is obvious.
but there are massive gains to be had by creating “smoother” intelligence via blending multiple jagged models together.
the era of model melding begins.
Tomas Hernando Kofman (@tomas_hk): Today we’re announcing Not Diamond Code, the world’s most powerful intelligent model router for long-horizon coding agents.
Not Diamond works with any gateway or harness, including Claude Code, to select the best model and reasoning effort for each step, reducing costs by 20-65%
Similar Articles
@julien_c: Model routing is the missing layer of the multi-model stack. open models keep getting better at coding, routing helps t…
Not Diamond Code is announced as an intelligent model router for long-horizon coding agents, selecting the best model and reasoning effort per step to reduce costs by 20-65%.
@cryptopunk7213: this is pretty genius. in a world of increasingly expensive and abundant ai models products like this are a dream AI mo…
Factory Router automatically selects the best AI model for each task, claiming to cut costs by 25% while maintaining frontier performance, a promising tool for large enterprises.
@alexatallah: If you're a researcher looking to: → conduct rigorous studies on how multiple models can outperform the frontier → leve…
OpenRouter launches Fusion API, a compound model that achieves high intelligence at half the price, leveraging the largest LLM marketplace.
@rohanpaul_ai: Read more on their official blog.
Not Diamond announces Not Diamond Code, an intelligent model router for coding agents that selects the best model and reasoning effort per step, claiming 20%+ cost savings and Pareto-optimal performance on coding benchmarks.
Split my agent into a cheap router model and a premium synthesis model, bill dropped about 75%
A developer splits their AI agent's LLM calls into a cheap router model (GPT-OSS 120B) for tool-picking and a premium model (gpt-5.4) for synthesis, cutting costs by ~78% while maintaining output quality.