NVIDIA announced Nemotron 3.5 Lightning, a 30B mixture-of-experts open model optimized for high-volume agentic AI workloads, alongside NeMo Switchyard, an open-source library for intelligent model routing across heterogeneous model ecosystems.
<div id="bsf_rt_marker"></div><p><span style="font-weight: 400;">As AI shifts from chatbots to autonomous agents, open models are serving market demands for full control over where AI runs and how it’s deployed and evolves.</span></p>
<p><span style="font-weight: 400;">Today, NVIDIA is expanding its Nemotron 3 model family with Nemotron 3.5 Lightning, the highest-efficiency model in its class for long-running agentic AI workloads. This release follows Nemotron 3 Nano and reflects NVIDIA’s commitment to continually improving open models for greater accuracy and speed. </span></p>
<p><span style="font-weight: 400;">Built for specialized tasks within larger multi-agent systems, Nemotron 3.5 Lightning, a 30-billion-parameter </span><a target="_blank" href="https://www.nvidia.com/en-us/glossary/mixture-of-experts/"><span style="font-weight: 400;">mixture-of-experts</span></a><span style="font-weight: 400;"> model, helps create smarter and more efficient agentic applications.</span></p>
<p><span style="font-weight: 400;">Also, NVIDIA is releasing NeMo Switchyard, an open source library for smart routing inside popular agent tools. Enterprises can use it to build a router based on their specific needs. When deployed, NeMo Switchyard can intelligently direct each request to the most capable and suitable model for the job, across developers’ own mix of open, proprietary and NVIDIA models, without requiring developers to rewrite their applications.</span></p>
<p><span style="font-weight: 400;">Together, Nemotron 3.5 Lightning and NeMo Switchyard deliver greater control over how AI is deployed, where it runs and how efficiently it operates — across PCs, workstations, data centers and the cloud.</span></p>
<figure id="attachment_97408" aria-describedby="caption-attachment-97408" style="width: 1200px" class="wp-caption alignnone"><img loading="lazy" decoding="async" class="size-large wp-image-97408" src="https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3-5-lightening-launch-1680x945.png" alt="" width="1200" height="675" srcset="https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3-5-lightening-launch-1680x945.png 1680w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3-5-lightening-launch-960x540.png 960w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3-5-lightening-launch-1280x720.png 1280w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3-5-lightening-launch-1536x864.png 1536w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3-5-lightening-launch-1290x725.png 1290w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3-5-lightening-launch-630x354.png 630w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3-5-lightening-launch-300x169.png 300w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3-5-lightening-launch-400x225.png 400w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3-5-lightening-launch.png 1920w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /><figcaption id="caption-attachment-97408" class="wp-caption-text">Nemotron 3.5 Lightning delivers frontier-level intelligence in a small, customizable open model built for high-volume agentic workflows.</figcaption></figure>
<h2><b>Always-On Agents Need a System of Models </b></h2>
<p><span style="font-weight: 400;">Modern agentic systems — always-on agents — increasingly operate as </span><a target="_blank" href="https://youtu.be/Np0afRWtdp8?si=om8lMihXpHOTXc3e"><span style="font-weight: 400;">systems of models</span></a><span style="font-weight: 400;">, or model ensembles, with different models specialized for different tasks. </span></p>
<p><a target="_blank" href="https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/"><span style="font-weight: 400;">NVIDIA Nemotron</span></a><span style="font-weight: 400;"> open models are designed for this architecture. A frontier reasoning model such as Nemotron 3 Ultra or GPT-5.6 may plan and orchestrate a workflow, while smaller specialized models like Nemotron 3.5 Lightning can perform targeted tasks such as code review, tool use, security alert monitoring and answering billing questions.</span></p>
<h2><b>Powering High-Volume Specialized Tasks With Nemotron 3.5 Lightning</b></h2>
<p><a target="_blank" href="https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/"><span style="font-weight: 400;">NVIDIA Nemotron 3.5 Lightning</span></a><span style="font-weight: 400;"> is a fully customizable open model built for high-volume tasks powering always-on agents. It was developed with contributions from the Nemotron Coalition, whose members provided evaluation methodologies, inference software and datasets to help advance the model.</span></p>
<p><span style="font-weight: 400;">The model delivers up to 4x faster output speed, leading to 30% faster agentic task completion compared with other models in its class. And because it’s open and customizable, Nemotron 3.5 Lightning can be easily post-trained with NVIDIA NeMo on an organization’s own domain data, tools and workflows to improve accuracy for specialized tasks.</span></p>
<figure id="attachment_97394" aria-describedby="caption-attachment-97394" style="width: 1200px" class="wp-caption alignnone"><img loading="lazy" decoding="async" class="size-large wp-image-97394" src="https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart3-1680x945.png" alt="" width="1200" height="675" srcset="https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart3-1680x945.png 1680w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart3-960x540.png 960w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart3-1280x720.png 1280w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart3-1536x864.png 1536w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart3-1290x725.png 1290w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart3-630x354.png 630w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart3-300x169.png 300w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart3-400x225.png 400w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart3.png 1920w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /><figcaption id="caption-attachment-97394" class="wp-caption-text">PinchBench benchmarks demonstrate that Nemotron 3.5 Lightning delivers faster agentic task completion with frontier-level accuracy compared to other models in its class.</figcaption></figure>
<p><span style="font-weight: 400;">AI leaders across industries are customizing Nemotron 3.5 Lightning for their workloads, including </span><span style="font-weight: 400;">CrowdStrike </span><span style="font-weight: 400;">for cybersecurity, </span><span style="font-weight: 400;">Harvey</span><span style="font-weight: 400;"> with </span><span style="font-weight: 400;">Trajectory</span><span style="font-weight: 400;"> for legal services and </span><a target="_blank" href="https://www.coderabbit.ai/blog/teaching-nvidia-nemotron-3-5-lightning-to-route-code-reviews"><span style="font-weight: 400;">CodeRabbit</span></a> <span style="font-weight: 400;">with </span><span style="font-weight: 400;">Baseten</span> <span style="font-weight: 400;">for code review, helping improve accuracy for domain-specific agentic tasks. Additionally</span><span style="font-weight: 400;">, </span><a target="_blank" href="https://www.lila.ai/news/building-the-agent-driven-era-of-science-with-nvidia-bionemo-agent-toolkit"><span style="font-weight: 400;">Lila Sciences</span></a> <span style="font-weight: 400;">i</span><span style="font-weight: 400;">s helping to improve reasoning capabilities for agentic tasks across physical and life sciences, and </span><a target="_blank" href="http://fastino.ai/blog/fastino-nemotron-3-5-lightning-finance-and-healthcare"><span style="font-weight: 400;">Fastino Labs</span></a> <span style="font-weight: 400;">customized the model and is seeing leading accuracies for software development, finance and healthcare workloads. </span></p>
<figure id="attachment_97417" aria-describedby="caption-attachment-97417" style="width: 1200px" class="wp-caption alignnone"><img loading="lazy" decoding="async" class="size-large wp-image-97417" src="https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart2-1680x945.png" alt="" width="1200" height="675" srcset="https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart2-1680x945.png 1680w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart2-960x540.png 960w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart2-1280x720.png 1280w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart2-1536x864.png 1536w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart2-1290x725.png 1290w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart2-630x354.png 630w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart2-300x169.png 300w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart2-400x225.png 400w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart2.png 1920w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /><figcaption id="caption-attachment-97417" class="wp-caption-text">Enterprises have customized Nemotron 3.5 Lightning to achieve leading accuracy for their specialized task in their agentic workflows.</figcaption></figure>
<p><span style="font-weight: 400;">Nemotron 3.5 Lightning also gives organizations control over privacy and deployment. It can run on local AI systems — including </span><a target="_blank" href="https://www.nvidia.com/en-us/ai-on-rtx/"><span style="font-weight: 400;">NVIDIA RTX PCs</span></a><span style="font-weight: 400;">, </span><a target="_blank" href="https://www.nvidia.com/en-us/products/workstations/dgx-spark/"><span style="font-weight: 400;">NVIDIA DGX Spark</span></a><span style="font-weight: 400;">, </span><a target="_blank" href="https://www.nvidia.com/en-us/products/workstations/dgx-station/"><span style="font-weight: 400;">NVIDIA DGX Station</span></a><span style="font-weight: 400;"> and </span><a target="_blank" href="https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/"><span style="font-weight: 400;">NVIDIA Jetson</span></a><span style="font-weight: 400;"> — to help users maximize existing infrastructure investments, or scale across edge AI devices, </span><a target="_blank" href="https://www.nvidia.com/en-us/products/workstations/"><span style="font-weight: 400;">NVIDIA RTX PRO</span></a><span style="font-weight: 400;"> workstations, data centers and cloud environments for enterprise use cases. And Nemotron 3.5 Lightning can run locally or on premises for high-volume, specialized tasks that require fast responses.</span></p>
<p><span style="font-weight: 400;">Also, as with every Nemotron launch, NVIDIA publishes as much of the training data and techniques as licensing permits, which allows for traceability, auditing and training of other models. Alongside Lightning, NVIDIA is releasing </span><a target="_blank" href="https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Terminal-Pivot-v1-nano35-release"><span style="font-weight: 400;">Nemotron-RL-Agentic-Terminal-Pivot</span></a><span style="font-weight: 400;">, an agentic reinforcement learning dataset used to post-train it for coding agent capabilities.</span></p>
<h2><b>More Efficient AI Apps With Model Routing </b></h2>
<p><span style="font-weight: 400;">Some models are better for coding, some for reasoning, some for lightweight tasks and some are optimized to run locally for greater privacy and efficiency. If customers rely on one default model, they might either overspend or lose quality; if they manage routing manually, it becomes integration work that can slow down a deployment.</span></p>
<p><a target="_blank" href="https://developer.nvidia.com/blog/route-ai-agent-workloads-across-models-with-nvidia-nemo-switchyard"><span style="font-weight: 400;">NVIDIA NeMo Switchyard</span></a><span style="font-weight: 400;"> is an open source model routing library for AI agents. The technology routes prompts to the most capable and efficient model for each step of an agent workflow automatically, based on specific needs. Agent application developers can tune or modify the router with different routing algorithms to match their priorities, such as quality, latency and cost requirements. In a system of models, enterprises can create powerful AI agents with improved tokenomics. </span></p>
<p><span style="font-weight: 400;">Internal benchmarks show that </span><span style="font-weight: 400;">NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone.</span></p>
<figure id="attachment_97397" aria-describedby="caption-attachment-97397" style="width: 1200px" class="wp-caption alignnone"><img loading="lazy" decoding="async" class="size-large wp-image-97397" src="https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart4-1680x945.png" alt="" width="1200" height="675" srcset="https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart4-1680x945.png 1680w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart4-960x540.png 960w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart4-1280x720.png 1280w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart4-1536x864.png 1536w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart4-1290x725.png 1290w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart4-630x354.png 630w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart4-300x169.png 300w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart4-400x225.png 400w, https://blogs.nvidia.com/wp-content/uploads/2026/08/nemotron-3.5-lightning-positioning-chart4.png 1920w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /><figcaption id="caption-attachment-97397" class="wp-caption-text">NVIDIA internal benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone.</figcaption></figure>
<p><span style="font-weight: 400;">NVIDIA is working with partners across the AI ecosystem to bring intelligent model routing into the tools and platforms developers already use. </span></p>
<ul>
<li><a target="_blank" href="https://boomi.com/blog/why-open-model-routing-matters"><b>Boomi</b></a><b>:</b> <span style="font-weight: 400;">Evaluated Switchyard across five routing capabilities, achieving 100% domain-routing accuracy, sending 59% of traffic to a 5x faster fine-tuned model and reducing later-turn latency by 21%.</span></li>
<li><b>Cadence</b><span style="font-weight: 400;">: Improved efficiency by 9.9% by using the </span><a target="_blank" href="https://www.cadence.com/en_US/home/company/newsroom/press-releases/pr/2026/cadence-unveils-industrys-first-fully-autonomous-virtual.html"><span style="font-weight: 400;">ChipStack AI Super Agent</span></a><span style="font-weight: 400;"> for a formal verification use case.</span></li>
<li><a target="_blank" href="https://dev.classmethod.jp/en/articles/nvidia-nemo-switchyard-first-touch/"><b>Classmethod</b></a><span style="font-weight: 400;">: Is running opencode and Fireworks workloads using NeMo Switchyard internally, with initial testing showing a 27% cost reduction while maintaining quality.</span></li>
<li><b>Cognition</b><span style="font-weight: 400;">: </span><span style="font-weight: 400;">Integrated the NVIDIA NeMo Switchyard staged router into Devin Desktop for NVIDIA internal use, achieving near-frontier performance on </span><a target="_blank" href="https://cognition.com/frontiercode"><span style="font-weight: 400;">FrontierCode Main </span></a><span style="font-weight: 400;">while reducing mean cost by 28% relative to routing all requests to a single underlying frontier model.</span></li>
<li><a target="_blank" href="https://konghq.com/blog/preview/6a6cd1d40364e90001c019a5"><b>Kong</b></a><span style="font-weight: 400;">:</span><span style="font-weight: 400;"> Delivers routing with NeMo Switchyard natively through Kong AI Gateway.</span></li>
<li><a target="_blank" href="http://www.langchain.com/blog/switchyard-agent-routing-benchmark"><b>LangChain</b></a><span style="font-weight: 400;">: With NeMo Switchyard, achieved 74% lower cost in 145 multi-turn Deep Agents tasks by routing only 7% of calls to a frontier model, at a 6% accuracy tradeoff.</span></li>
<li><b>LiteLLM</b><span style="font-weight: 400;">: Is adding NeMo Switchyard as a plug-in into its proxy layer so developers can access these benefits without changing their existing stack.</span></li>
<li><b>Nous Research</b><span style="font-weight: 400;">: Integrated NeMo Switchyard into Hermes to provide developers with an easy-to-configure routing system to improve agent efficiency.</span></li>
<li><a target="_blank" href="https://x.com/RampLabs/status/2087163448513765449?s=20"><b>Ramp</b></a><span style="font-weight: 400;">: Used NeMo Switchyard to match a frontier model’s performance while cutting costs by 58% and runtime by 33% in </span><a target="_blank" href="https://labs.ramp.com/swebench"><span style="font-weight: 400;">Ramp SWE-Bench</span></a><span style="font-weight: 400;">.</span></li>
<li><b>Siemens</b><span style="font-weight: 400;">: Is benchmarking to improve efficiency in its </span><a target="_blank" href="https://www.siemens.com/en-us/products/fuse-eda-ai-system/agent/"><span style="font-weight: 400;">Fuse EDA AI Agent</span></a><span style="font-weight: 400;">.</span></li>
</ul>
<p><i><span style="font-weight: 400;">Nemotron 3.5 Lightning is available on </span></i><a target="_blank" href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4"><i><span style="font-weight: 400;">Hugging Face</span></i></a><i><span style="font-weight: 400;">, </span></i><a target="_blank" href="https://modelscope.ai/collections/nv-community/Nemotron-35-Lightning"><i><span style="font-weight: 400;">ModelScope</span></i></a><i><span style="font-weight: 400;">, </span></i><a target="_blank" href="https://openrouter.ai/nvidia/nemotron-3.5-lightning:free"><i><span style="font-weight: 400;">OpenRouter</span></i></a><i><span style="font-weight: 400;"> and </span></i><a target="_blank" href="https://build.nvidia.com/nvidia/nemotron-3.5-lightning-30b-a3b"><i><span style="font-weight: 400;">build.nvidia.com</span></i></a><i><span style="font-weight: 400;"> as an NVIDIA NIM microservice as well as through a broad ecosystem of NVIDIA Cloud Partners, post-training platforms, inference platforms and cloud service providers. NeMo Switchyard is available on </span></i><a target="_blank" href="https://github.com/NVIDIA-NeMo/Switchyard"><i><span style="font-weight: 400;">GitHub</span></i></a><i><span style="font-weight: 400;"> and coming to partner platforms soon. </span></i></p>
# NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI
Source: [https://blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx/](https://blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx/)
As AI shifts from chatbots to autonomous agents, open models are serving market demands for full control over where AI runs and how it’s deployed and evolves\.
Today, NVIDIA is expanding its Nemotron 3 model family with Nemotron 3\.5 Lightning, the highest\-efficiency model in its class for long\-running agentic AI workloads\. This release follows Nemotron 3 Nano and reflects NVIDIA’s commitment to continually improving open models for greater accuracy and speed\.
Built for specialized tasks within larger multi\-agent systems, Nemotron 3\.5 Lightning, a 30\-billion\-parameter[mixture\-of\-experts](https://www.nvidia.com/en-us/glossary/mixture-of-experts/)model, helps create smarter and more efficient agentic applications\.
Also, NVIDIA is releasing NeMo Switchyard, an open source library for smart routing inside popular agent tools\. Enterprises can use it to build a router based on their specific needs\. When deployed, NeMo Switchyard can intelligently direct each request to the most capable and suitable model for the job, across developers’ own mix of open, proprietary and NVIDIA models, without requiring developers to rewrite their applications\.
Together, Nemotron 3\.5 Lightning and NeMo Switchyard deliver greater control over how AI is deployed, where it runs and how efficiently it operates — across PCs, workstations, data centers and the cloud\.
Nemotron 3\.5 Lightning delivers frontier\-level intelligence in a small, customizable open model built for high\-volume agentic workflows\.## **Always\-On Agents Need a System of Models**
Modern agentic systems — always\-on agents — increasingly operate as[systems of models](https://youtu.be/Np0afRWtdp8?si=om8lMihXpHOTXc3e), or model ensembles, with different models specialized for different tasks\.
[NVIDIA Nemotron](https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/)open models are designed for this architecture\. A frontier reasoning model such as Nemotron 3 Ultra or GPT\-5\.6 may plan and orchestrate a workflow, while smaller specialized models like Nemotron 3\.5 Lightning can perform targeted tasks such as code review, tool use, security alert monitoring and answering billing questions\.
## **Powering High\-Volume Specialized Tasks With Nemotron 3\.5 Lightning**
[NVIDIA Nemotron 3\.5 Lightning](https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/)is a fully customizable open model built for high\-volume tasks powering always\-on agents\. It was developed with contributions from the Nemotron Coalition, whose members provided evaluation methodologies, inference software and datasets to help advance the model\.
The model delivers up to 4x faster output speed, leading to 30% faster agentic task completion compared with other models in its class\. And because it’s open and customizable, Nemotron 3\.5 Lightning can be easily post\-trained with NVIDIA NeMo on an organization’s own domain data, tools and workflows to improve accuracy for specialized tasks\.
PinchBench benchmarks demonstrate that Nemotron 3\.5 Lightning delivers faster agentic task completion with frontier\-level accuracy compared to other models in its class\.AI leaders across industries are customizing Nemotron 3\.5 Lightning for their workloads, includingCrowdStrikefor cybersecurity,HarveywithTrajectoryfor legal services and[CodeRabbit](https://www.coderabbit.ai/blog/teaching-nvidia-nemotron-3-5-lightning-to-route-code-reviews)withBasetenfor code review, helping improve accuracy for domain\-specific agentic tasks\. Additionally,[Lila Sciences](https://www.lila.ai/news/building-the-agent-driven-era-of-science-with-nvidia-bionemo-agent-toolkit)is helping to improve reasoning capabilities for agentic tasks across physical and life sciences, and[Fastino Labs](http://fastino.ai/blog/fastino-nemotron-3-5-lightning-finance-and-healthcare)customized the model and is seeing leading accuracies for software development, finance and healthcare workloads\.
Enterprises have customized Nemotron 3\.5 Lightning to achieve leading accuracy for their specialized task in their agentic workflows\.Nemotron 3\.5 Lightning also gives organizations control over privacy and deployment\. It can run on local AI systems — including[NVIDIA RTX PCs](https://www.nvidia.com/en-us/ai-on-rtx/),[NVIDIA DGX Spark](https://www.nvidia.com/en-us/products/workstations/dgx-spark/),[NVIDIA DGX Station](https://www.nvidia.com/en-us/products/workstations/dgx-station/)and[NVIDIA Jetson](https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/)— to help users maximize existing infrastructure investments, or scale across edge AI devices,[NVIDIA RTX PRO](https://www.nvidia.com/en-us/products/workstations/)workstations, data centers and cloud environments for enterprise use cases\. And Nemotron 3\.5 Lightning can run locally or on premises for high\-volume, specialized tasks that require fast responses\.
Also, as with every Nemotron launch, NVIDIA publishes as much of the training data and techniques as licensing permits, which allows for traceability, auditing and training of other models\. Alongside Lightning, NVIDIA is releasing[Nemotron\-RL\-Agentic\-Terminal\-Pivot](https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Terminal-Pivot-v1-nano35-release), an agentic reinforcement learning dataset used to post\-train it for coding agent capabilities\.
## **More Efficient AI Apps With Model Routing**
Some models are better for coding, some for reasoning, some for lightweight tasks and some are optimized to run locally for greater privacy and efficiency\. If customers rely on one default model, they might either overspend or lose quality; if they manage routing manually, it becomes integration work that can slow down a deployment\.
[NVIDIA NeMo Switchyard](https://developer.nvidia.com/blog/route-ai-agent-workloads-across-models-with-nvidia-nemo-switchyard)is an open source model routing library for AI agents\. The technology routes prompts to the most capable and efficient model for each step of an agent workflow automatically, based on specific needs\. Agent application developers can tune or modify the router with different routing algorithms to match their priorities, such as quality, latency and cost requirements\. In a system of models, enterprises can create powerful AI agents with improved tokenomics\.
Internal benchmarks show thatNeMo Switchyard maintains frontier\-level accuracy while reducing task completion cost to nearly one\-third of Opus 4\.8 alone\.
NVIDIA internal benchmarks show that NeMo Switchyard maintains frontier\-level accuracy while reducing task completion cost to nearly one\-third of Opus 4\.8 alone\.NVIDIA is working with partners across the AI ecosystem to bring intelligent model routing into the tools and platforms developers already use\.
- [**Boomi**](https://boomi.com/blog/why-open-model-routing-matters)**:**Evaluated Switchyard across five routing capabilities, achieving 100% domain\-routing accuracy, sending 59% of traffic to a 5x faster fine\-tuned model and reducing later\-turn latency by 21%\.
- **Cadence**: Improved efficiency by 9\.9% by using the[ChipStack AI Super Agent](https://www.cadence.com/en_US/home/company/newsroom/press-releases/pr/2026/cadence-unveils-industrys-first-fully-autonomous-virtual.html)for a formal verification use case\.
- [**Classmethod**](https://dev.classmethod.jp/en/articles/nvidia-nemo-switchyard-first-touch/): Is running opencode and Fireworks workloads using NeMo Switchyard internally, with initial testing showing a 27% cost reduction while maintaining quality\.
- **Cognition**:Integrated the NVIDIA NeMo Switchyard staged router into Devin Desktop for NVIDIA internal use, achieving near\-frontier performance on[FrontierCode Main](https://cognition.com/frontiercode)while reducing mean cost by 28% relative to routing all requests to a single underlying frontier model\.
- [**Kong**](https://konghq.com/blog/preview/6a6cd1d40364e90001c019a5):Delivers routing with NeMo Switchyard natively through Kong AI Gateway\.
- [**LangChain**](http://www.langchain.com/blog/switchyard-agent-routing-benchmark): With NeMo Switchyard, achieved 74% lower cost in 145 multi\-turn Deep Agents tasks by routing only 7% of calls to a frontier model, at a 6% accuracy tradeoff\.
- **LiteLLM**: Is adding NeMo Switchyard as a plug\-in into its proxy layer so developers can access these benefits without changing their existing stack\.
- **Nous Research**: Integrated NeMo Switchyard into Hermes to provide developers with an easy\-to\-configure routing system to improve agent efficiency\.
- [**Ramp**](https://x.com/RampLabs/status/2087163448513765449?s=20): Used NeMo Switchyard to match a frontier model’s performance while cutting costs by 58% and runtime by 33% in[Ramp SWE\-Bench](https://labs.ramp.com/swebench)\.
- **Siemens**: Is benchmarking to improve efficiency in its[Fuse EDA AI Agent](https://www.siemens.com/en-us/products/fuse-eda-ai-system/agent/)\.
*Nemotron 3\.5 Lightning is available on*[*Hugging Face*](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4)*,*[*ModelScope*](https://modelscope.ai/collections/nv-community/Nemotron-35-Lightning)*,*[*OpenRouter*](https://openrouter.ai/nvidia/nemotron-3.5-lightning:free)*and*[*build\.nvidia\.com*](https://build.nvidia.com/nvidia/nemotron-3.5-lightning-30b-a3b)*as an NVIDIA NIM microservice as well as through a broad ecosystem of NVIDIA Cloud Partners, post\-training platforms, inference platforms and cloud service providers\. NeMo Switchyard is available on*[*GitHub*](https://github.com/NVIDIA-NeMo/Switchyard)*and coming to partner platforms soon\.*
Nvidia released Nemotron 3.5 Lightning, a 30B open mixture-of-experts model, and NeMo Switchyard, an open-source routing library that dynamically assigns each step of an AI agent workflow to the most suitable model. Nvidia claims the combination can cut agent task costs to about a third while maintaining frontier-level performance.
NVIDIA released Nemotron 3.5 Lightning, a 30B-parameter MoE model with only 3B active parameters, optimized for agent execution tasks. It claims faster, cheaper tool calls and agent execution while staying fully open-source under OpenMDW-1.1.
NVIDIA announces Nemotron 3 Nano Omni, an open multimodal model that unifies vision, audio, and language processing to enable faster and more efficient AI agents, achieving up to 9x higher throughput compared to other open omni models.