Tag
Nvidia released Nemotron 3.5 Lightning, a 30B open mixture-of-experts model, and NeMo Switchyard, an open-source routing library that dynamically assigns each step of an AI agent workflow to the most suitable model. Nvidia claims the combination can cut agent task costs to about a third while maintaining frontier-level performance.
LangChain tested NVIDIA's open-source router Switchyard with Deep Agents, showing that routing between models can cut costs by ~70% while retaining ~90% accuracy, and introduced new middleware integration for NVIDIA models.