NVIDIA and CoreWeave are bringing Vera Rubin NVL72 systems into production on CoreWeave Cloud, with Cognition (maker of Devin AI) as the first production customer reporting up to 4.8x higher token throughput on software engineering inference workloads. CoreWeave also launched CoreWeave Forge and will offer the NVIDIA Vera CPU, marketed as the first CPU built for AI agents.
<div id="bsf_rt_marker"></div><p><span style="font-weight: 400;">Building on nearly a decade of co-engineering, CoreWeave has built NVIDIA compute, networking and software into a cloud purpose-built for AI that’s still returning on investment across multiple generations of deployment. Now, CoreWeave is bringing the next generation of NVIDIA infrastructure to production.</span></p>
<p><span style="font-weight: 400;">At CoreWeave Fully Connected, running this week in San Francisco, CoreWeave announced availability of </span><a target="_blank" href="https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/"><span style="font-weight: 400;">NVIDIA Vera Rubin NVL72</span></a><span style="font-weight: 400;"> systems with Spectrum-X 102.4T Ethernet networking. Cognition, the applied AI lab behind the Devin AI software engineer, is the first customer running production workloads on Vera Rubin. </span></p>
<p><span style="font-weight: 400;">CoreWeave will also offer </span><a target="_blank" href="https://www.nvidia.com/en-us/data-center/vera-cpu/"><span style="font-weight: 400;">NVIDIA Vera</span></a><span style="font-weight: 400;">, the first CPU built for AI agents. In addition, CoreWeave launched CoreWeave Forge, a connected environment for training, evaluating and improving models and agents on NVIDIA accelerated computing.</span></p>
<p><span style="font-weight: 400;">“NVIDIA accelerated computing delivers value across generations,” said Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA. “CoreWeave’s NVIDIA V100 GPUs are still running customer workloads nearly a decade after Volta launched, even as CoreWeave brings Vera Rubin NVL72 into production. That’s the strength of the NVIDIA platform: infrastructure that keeps earning for years, and the flexibility to put the right GPU on the right workload.</span><span style="font-weight: 400;">”</span></p>
<h2><b>Cognition Runs on Vera Rubin NVL72 With 4.8x Higher Token Throughput</b></h2>
<p><span style="font-weight: 400;">Cognition runs training, reinforcement learning and production inference for Devin on CoreWeave. The company scaled to thousands of GPUs on CoreWeave in nine months, powering Cognition inference workloads.</span></p>
<p><span style="font-weight: 400;">Earlier this month, CoreWeave received its first Vera Rubin NVL72 production racks. Shortly after, Cognition benchmarked Vera Rubin’s inference performance against a GB200 NVL72 baseline using a real-world software engineering workload. To generate this workload, it sampled a subset of tasks from FrontierCode and deployed AI agents to solve them. </span></p>
<p><span style="font-weight: 400;">In its early tests, Cognition saw Vera Rubin NVL72 deliver up to a 4.8x increase in total token throughput for SWE-2 inference workloads over GB200 NVL72. For Devin, those gains mean faster real-time code generation and more responsive multistep reasoning.</span></p>
<p><span style="font-weight: 400;">“Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship,” said Silas Alberti, founding team at Cognition. “Having all of it on one platform, with NVIDIA and CoreWeave engineers who work the hard problems alongside ours, matters more to us than any single spec.”</span></p>
<h2><b>CoreWeave Announces NVIDIA Vera Rubin Availability on CoreWeave Cloud</b></h2>
<p><span style="font-weight: 400;">CoreWeave announced availability of NVIDIA Vera Rubin NVL72 on CoreWeave Cloud, making it one of the first cloud providers to deliver the platform in customers’ hands.</span></p>
<p><span style="font-weight: 400;">Early-access customers can put the performance of NVIDIA’s full-stack AI factory platform to work quickly on CoreWeave Cloud. In days, CoreWeave stood up a production Vera Rubin cluster for Cognition, achieved as a result of the codesign and collaboration between NVIDIA and CoreWeave up and down the stack, from infrastructure to tokens served. </span></p>
<p><span style="font-weight: 400;">Capacity can be operated through CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes and CoreWeave Inference. </span></p>
<h2><b>NVIDIA Vera CPU to Come to CoreWeave Cloud, Tests Show More Than 3x Faster Agentic Sandbox Startups</b></h2>
<p><span style="font-weight: 400;">Agentic AI puts pressure on infrastructure from two directions: serving agents demands low-latency compute at scale, while improving them through post-training requires thousands of isolated environments running at once.</span></p>
<p><span style="font-weight: 400;">NVIDIA Vera CPU is purpose-built for agentic workloads. For agentic AI, a key performance measure is how many isolated agent environments can run at once and how consistent and performant each one stays as that number grows. </span></p>
<p><span style="font-weight: 400;">CoreWeave’s deployment of Vera puts 128 CPUs and 11,264 cores in a single rack, enough for more than 11,000 concurrent environments at one core each. With CoreWeave Sandboxes, these environments are hardware-isolated and run alongside the training jobs they support, with Spectrum-X Ethernet switches and BlueField-4 DPUs ensuring secure, high-performance, secure agent communication at low latency. </span></p>
<p><span style="font-weight: 400;">In testing, CoreWeave achieved more than 3x faster agent sandbox startup times on NVIDIA Vera CPUs, accelerating and scaling its sandboxes, an execution layer for reinforcement learning (RL), agent tool use and model evaluation that let AI teams run code in isolated environments on CoreWeave. </span><span style="font-weight: 400;">On Terminal-Bench, CoreWeave saw a 1.7x performance gain on Vera CPU across all passing tasks.</span></p>
<h2><b>CoreWeave Forge: Closing the AI Loop From Production Back to Training</b></h2>
<p><span style="font-weight: 400;">Models and agents improve by running a loop: production behavior informs the next training run, and each evaluation sharpens the next version. This loop has historically been split across tools from different vendors, with signal lost at every handoff.</span></p>
<p><span style="font-weight: 400;">CoreWeave Forge unifies Weights & Biases, post-training expertise from OpenPipe and the open source marimo notebook project in one connected environment built for continuous model and agent improvement. It stays open across models, frameworks and clouds.</span></p>
<p><span style="font-weight: 400;">New and expanded capabilities available include:</span></p>
<ul>
<li><b>CoreWeave ARIA </b><span style="font-weight: 400;">— now generally available — helps users learn, research, code and iterate across the AI loop, analyzing runs, proposing experiments, recommending code changes and storing them in GitHub, and bringing back actionable evidence that analyzes experiment data, surfaces what drove a change and proposes the next experiments to run.</span></li>
<li><b>CoreWeave Agent Lens </b><span style="font-weight: 400;">— a new service — turns production agent observability into continuous improvement and understandable insights. It improves failure detection by 20% and fixes issues at half of the cost, which turns tens of millions of production agent traces into insights that drive fixes.</span></li>
<li><b>CoreWeave Sandboxes</b><span style="font-weight: 400;"> — now generally available — let users run agents, tool calls, RL and evaluations in isolated CPU or GPU execution environments, on serverless infrastructure or on the infrastructure they already train on, providing a fresh, isolated environment for every agent tool call, RL run or evaluation.</span></li>
<li><b>Post-training improves model quality and cuts latency and costs</b><span style="font-weight: 400;"> harnessing users’ own production signals, with no training cluster required.</span> <a target="_blank" href="https://www.coreweave.com/products/coreweave-forge/serverless-sft"><span style="font-weight: 400;">Serverless supervised fine-tuning</span></a><span style="font-weight: 400;"> and</span> <a target="_blank" href="https://www.coreweave.com/products/coreweave-forge/serverless-rl"><span style="font-weight: 400;">serverless RL</span></a><span style="font-weight: 400;"> let users experiment with their own training recipes. Serverless RL trains 1.4x faster at 40% lower cost than a self-managed setup. </span></li>
</ul>
<p><a target="_blank" href="https://www.nvidia.com/en-us/ai/dynamo/"><span style="font-weight: 400;">NVIDIA Dynamo</span></a><span style="font-weight: 400;">, an open source inference framework for AI factories, powers CoreWeave’s managed inference service as well as RL Rollouts, now in private preview. RL Rollouts loads new checkpoints into a live deployment while it’s running, so reinforcement learning continues without redeploys, and post-training gets the same inference efficiency as production.</span></p>
<p><span style="font-weight: 400;">Canva, Capital One and MasterClass are among the first companies building on Forge.</span></p>
<p><a target="_blank" href="https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/"><span style="font-weight: 400;">NVIDIA Nemotron</span></a><span style="font-weight: 400;"> open models give teams on Forge a direct path to customizing and deploying reasoning and multimodal models for agentic workflows.</span></p>
<h2><b>Proven Impact From Startups to Global Enterprises</b></h2>
<p><span style="font-weight: 400;">AI labs, AI-natives and global enterprises are using the co-engineered NVIDIA and CoreWeave platform to move from prototype to production faster. </span></p>
<p><span style="font-weight: 400;">In healthcare, Ennoble Care, a home-based care provider serving about 50,000 high-need Medicare patients across 15 states, selected CoreWeave to run clinical AI inference. It will use reserved NVIDIA RTX PRO 6000 GPU capacity on CoreWeave Kubernetes Service to scale AI agents for clinical documentation, decision support and back-office automation.</span></p>
<p><span style="font-weight: 400;">CoreWeave has delivered record </span><a target="_blank" href="https://coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra"><span style="font-weight: 400;">MLPerf results</span></a><span style="font-weight: 400;"> in every round of training and inference. It’s the only cloud provider who holds the Platinum ranking in SemiAnalysis ClusterMAX 1.0, 2.0 and 3.0, and serves nine of the 10 leading AI labs.</span></p>
<p><span style="font-weight: 400;">Together, NVIDIA and CoreWeave are giving customers a platform to turn experimental agents into production systems that write software, support clinicians and do useful work in the real world.</span></p>
<p><i><span style="font-weight: 400;">Learn more by attending </span></i><a target="_blank" href="https://www.nvidia.com/en-us/events/coreweave-fully-connected/"><i><span style="font-weight: 400;">NVIDIA sessions, demos and workshops at CoreWeave Fully Connected</span></i></a><i><span style="font-weight: 400;">.</span></i></p>
# From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI
Source: [https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/)
Building on nearly a decade of co\-engineering, CoreWeave has built NVIDIA compute, networking and software into a cloud purpose\-built for AI that’s still returning on investment across multiple generations of deployment\. Now, CoreWeave is bringing the next generation of NVIDIA infrastructure to production\.
At CoreWeave Fully Connected, running this week in San Francisco, CoreWeave announced availability of[NVIDIA Vera Rubin NVL72](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/)systems with Spectrum\-X 102\.4T Ethernet networking\. Cognition, the applied AI lab behind the Devin AI software engineer, is the first customer running production workloads on Vera Rubin\.
CoreWeave will also offer[NVIDIA Vera](https://www.nvidia.com/en-us/data-center/vera-cpu/), the first CPU built for AI agents\. In addition, CoreWeave launched CoreWeave Forge, a connected environment for training, evaluating and improving models and agents on NVIDIA accelerated computing\.
“NVIDIA accelerated computing delivers value across generations,” said Ian Buck, vice president of hyperscale and high\-performance computing at NVIDIA\. “CoreWeave’s NVIDIA V100 GPUs are still running customer workloads nearly a decade after Volta launched, even as CoreWeave brings Vera Rubin NVL72 into production\. That’s the strength of the NVIDIA platform: infrastructure that keeps earning for years, and the flexibility to put the right GPU on the right workload\.”
## **Cognition Runs on Vera Rubin NVL72 With 4\.8x Higher Token Throughput**
Cognition runs training, reinforcement learning and production inference for Devin on CoreWeave\. The company scaled to thousands of GPUs on CoreWeave in nine months, powering Cognition inference workloads\.
Earlier this month, CoreWeave received its first Vera Rubin NVL72 production racks\. Shortly after, Cognition benchmarked Vera Rubin’s inference performance against a GB200 NVL72 baseline using a real\-world software engineering workload\. To generate this workload, it sampled a subset of tasks from FrontierCode and deployed AI agents to solve them\.
In its early tests, Cognition saw Vera Rubin NVL72 deliver up to a 4\.8x increase in total token throughput for SWE\-2 inference workloads over GB200 NVL72\. For Devin, those gains mean faster real\-time code generation and more responsive multistep reasoning\.
“Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship,” said Silas Alberti, founding team at Cognition\. “Having all of it on one platform, with NVIDIA and CoreWeave engineers who work the hard problems alongside ours, matters more to us than any single spec\.”
## **CoreWeave Announces NVIDIA Vera Rubin Availability on CoreWeave Cloud**
CoreWeave announced availability of NVIDIA Vera Rubin NVL72 on CoreWeave Cloud, making it one of the first cloud providers to deliver the platform in customers’ hands\.
Early\-access customers can put the performance of NVIDIA’s full\-stack AI factory platform to work quickly on CoreWeave Cloud\. In days, CoreWeave stood up a production Vera Rubin cluster for Cognition, achieved as a result of the codesign and collaboration between NVIDIA and CoreWeave up and down the stack, from infrastructure to tokens served\.
Capacity can be operated through CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes and CoreWeave Inference\.
## **NVIDIA Vera CPU to Come to CoreWeave Cloud, Tests Show More Than 3x Faster Agentic Sandbox Startups**
Agentic AI puts pressure on infrastructure from two directions: serving agents demands low\-latency compute at scale, while improving them through post\-training requires thousands of isolated environments running at once\.
NVIDIA Vera CPU is purpose\-built for agentic workloads\. For agentic AI, a key performance measure is how many isolated agent environments can run at once and how consistent and performant each one stays as that number grows\.
CoreWeave’s deployment of Vera puts 128 CPUs and 11,264 cores in a single rack, enough for more than 11,000 concurrent environments at one core each\. With CoreWeave Sandboxes, these environments are hardware\-isolated and run alongside the training jobs they support, with Spectrum\-X Ethernet switches and BlueField\-4 DPUs ensuring secure, high\-performance, secure agent communication at low latency\.
In testing, CoreWeave achieved more than 3x faster agent sandbox startup times on NVIDIA Vera CPUs, accelerating and scaling its sandboxes, an execution layer for reinforcement learning \(RL\), agent tool use and model evaluation that let AI teams run code in isolated environments on CoreWeave\.On Terminal\-Bench, CoreWeave saw a 1\.7x performance gain on Vera CPU across all passing tasks\.
## **CoreWeave Forge: Closing the AI Loop From Production Back to Training**
Models and agents improve by running a loop: production behavior informs the next training run, and each evaluation sharpens the next version\. This loop has historically been split across tools from different vendors, with signal lost at every handoff\.
CoreWeave Forge unifies Weights & Biases, post\-training expertise from OpenPipe and the open source marimo notebook project in one connected environment built for continuous model and agent improvement\. It stays open across models, frameworks and clouds\.
New and expanded capabilities available include:
- **CoreWeave ARIA**— now generally available — helps users learn, research, code and iterate across the AI loop, analyzing runs, proposing experiments, recommending code changes and storing them in GitHub, and bringing back actionable evidence that analyzes experiment data, surfaces what drove a change and proposes the next experiments to run\.
- **CoreWeave Agent Lens**— a new service — turns production agent observability into continuous improvement and understandable insights\. It improves failure detection by 20% and fixes issues at half of the cost, which turns tens of millions of production agent traces into insights that drive fixes\.
- **CoreWeave Sandboxes**— now generally available — let users run agents, tool calls, RL and evaluations in isolated CPU or GPU execution environments, on serverless infrastructure or on the infrastructure they already train on, providing a fresh, isolated environment for every agent tool call, RL run or evaluation\.
- **Post\-training improves model quality and cuts latency and costs**harnessing users’ own production signals, with no training cluster required\.[Serverless supervised fine\-tuning](https://www.coreweave.com/products/coreweave-forge/serverless-sft)and[serverless RL](https://www.coreweave.com/products/coreweave-forge/serverless-rl)let users experiment with their own training recipes\. Serverless RL trains 1\.4x faster at 40% lower cost than a self\-managed setup\.
[NVIDIA Dynamo](https://www.nvidia.com/en-us/ai/dynamo/), an open source inference framework for AI factories, powers CoreWeave’s managed inference service as well as RL Rollouts, now in private preview\. RL Rollouts loads new checkpoints into a live deployment while it’s running, so reinforcement learning continues without redeploys, and post\-training gets the same inference efficiency as production\.
Canva, Capital One and MasterClass are among the first companies building on Forge\.
[NVIDIA Nemotron](https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/)open models give teams on Forge a direct path to customizing and deploying reasoning and multimodal models for agentic workflows\.
## **Proven Impact From Startups to Global Enterprises**
AI labs, AI\-natives and global enterprises are using the co\-engineered NVIDIA and CoreWeave platform to move from prototype to production faster\.
In healthcare, Ennoble Care, a home\-based care provider serving about 50,000 high\-need Medicare patients across 15 states, selected CoreWeave to run clinical AI inference\. It will use reserved NVIDIA RTX PRO 6000 GPU capacity on CoreWeave Kubernetes Service to scale AI agents for clinical documentation, decision support and back\-office automation\.
CoreWeave has delivered record[MLPerf results](https://coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra)in every round of training and inference\. It’s the only cloud provider who holds the Platinum ranking in SemiAnalysis ClusterMAX 1\.0, 2\.0 and 3\.0, and serves nine of the 10 leading AI labs\.
Together, NVIDIA and CoreWeave are giving customers a platform to turn experimental agents into production systems that write software, support clinicians and do useful work in the real world\.
*Learn more by attending*[*NVIDIA sessions, demos and workshops at CoreWeave Fully Connected*](https://www.nvidia.com/en-us/events/coreweave-fully-connected/)*\.*
NVIDIA introduces Vera, a new max single-threaded CPU designed for the agentic AI era, optimizing per-core performance to accelerate AI agent loops and maximize AI factory revenue.
NVIDIA hand-delivers the first Vera CPUs to Anthropic, OpenAI, Oracle Cloud Infrastructure, and SpaceXAI, marking the arrival of a CPU purpose-built for agentic AI workloads.
NVIDIA announced its new Vera architecture CPUs, built from the ground up for agentic AI and reinforcement learning, claiming 2x performance over x86 alternatives. The Vera Rubin NVL72 platform integrates 72 GPUs and 36 CPUs, with major customers including Meta, Oracle, and Alibaba.
NVIDIA's Vera Rubin NVL72 system demonstrates up to 30x higher throughput per megawatt for agentic AI workloads, setting a new efficiency standard for AI infrastructure.
NVIDIA announces its Vera CPU will power new supercomputers at Los Alamos National Laboratory, delivering significant performance improvements for agentic AI simulations and scientific workloads.