@ClementDelangue: Routing and post-training open-source models won't only give you more accurate systems but also meaningfully faster and…
Summary
Discussion on how routing and post-training open-source models can outperform frontier models in accuracy, speed, and cost, with Harvey's partnership with Fireworks AI demonstrating hybrid legal agents beating frontier models on quality and cost.
View Cached Full Text
Cached at: 06/03/26, 07:54 PM
Routing and post-training open-source models won’t only give you more accurate systems but also meaningfully faster and cheaper systems as most companies are currently learning (in addition to giving you more control and privacy).
The idea that a “frontier” model (by frontier we mean is slightly more accurate on a few very limited benchmarks) will be better for all domains, all tasks, all setups just doesn’t hold up! It’s marketing for making you pay more!
Harvey (@harvey): We partnered with @FireworksAI_HQ to train open-source models for legal. Here’s what we found:
- Hybrid legal agents can beat frontier models on quality and cost by routing selectively to a frontier advisor.
We tested a hybrid setup where GLM 5.1 served as the primary worker,
Similar Articles
@alexatallah: If you're a researcher looking to: → conduct rigorous studies on how multiple models can outperform the frontier → leve…
OpenRouter launches Fusion API, a compound model that achieves high intelligence at half the price, leveraging the largest LLM marketplace.
@harvey: Inference-time model routing based on legal practice area dramatically improves agent performance. So do other types of…
Harvey is experimenting with inference-time model routing and blended intelligence techniques to improve agent performance in legal tasks, highlighting both promise and risks.
@aiDotEngineer: Your Agent Can Now Train Models The argument from @mervenoyann: open source models have caught up. GLM 5.1 is leading t…
The talk by @mervenoyann demonstrates that open source models like GLM 5.1 have caught up to closed models, and shows how Hugging Face's ecosystem enables agents to train models, run inference, and build workflows.
@julien_c: Model routing is the missing layer of the multi-model stack. open models keep getting better at coding, routing helps t…
Not Diamond Code is announced as an intelligent model router for long-horizon coding agents, selecting the best model and reasoning effort per step to reduce costs by 20-65%.
@gabepereyra: Harvey partnered with @appliedcompute to train a legal agent. We optimized each part of the agent stack, including the …
Harvey partnered with Applied Compute to train a legal agent, optimizing the agent stack and post-training the GLM-5.1 model using reward signals from their Legal Agent Benchmark.