Chips

Nvidia's Vera CPU and CoreWeave tackle the processing bottleneck in agentic AI systems

As AI agents move beyond answering questions to executing tasks, infrastructure must balance GPU-powered reasoning with CPU-intensive execution work. Nvidia and CoreWeave are deploying the Vera CPU to eliminate this emerging constraint.

4 min read
Nvidia and CoreWeave tackle the CPU bottleneck in agentic AI infrastructure

Building infrastructure for agentic AI requires systems that can handle both the computational work between a model's decisions and the execution phase as agents move from responding to queries to performing actions.

Graphics processing units drive the reasoning capabilities of AI models, while central processing units manage the bulk of execution tasks. This division of labor informs Nvidia Corp.'s strategy around its Vera CPU, which the company plans to deploy in partnership with CoreWeave Inc. Hannah Coutand, director of product marketing for Vera CPU at Nvidia, outlined the reasoning behind the effort.

The reason why Vera is so important is because we recognize that extreme co-design means that we look across the AI factory and we want to make sure we solve for any inefficient bottlenecks. We recognized the CPU was becoming one of them. We didn't set out to say, 'Let's go and build CPUs.' We set out to solve this bottleneck.

Hannah Coutand, Nvidia

Coutand and Harsh Banwait, senior director of product at CoreWeave, discussed the technical challenges and solutions during an exclusive broadcast at the Fully Connected event, speaking with theCUBE Research's Dave Vellante and John Furrier. Their conversation covered CPU performance requirements, sandbox capabilities and security considerations for AI infrastructure.

CPU performance becomes critical as agentic workloads expand

The convergence of training and inference operations means infrastructure must support both model computation and the execution environments where agents operate. Tool invocations, API calls and SQL queries all generate CPU-bound work that grows with agentic deployments, according to Coutand.

I think the volume certainly plays a role. That's why having a CPU that handles those types of calls extremely well, with fast, beefy cores, high memory bandwidth and low latency, it handles both operating in this new world … and serves as a great foundation for simply agentic use cases [and] reinforcement learning.

Hannah Coutand, Nvidia

CoreWeave intends to make Vera CPU capacity available as a standalone offering alongside its existing infrastructure. The company's Sandboxes product creates isolated execution environments where agents can run tasks during reinforcement learning and inference phases. Testing with Vera CPUs yielded measurable improvements, Banwait reported.

Quite recently, we also tested that with Vera CPUs, and we were proud to share that we saw about a 3x improvement in performance in terms of Sandbox startup times. For us, it's a very important combination of getting the performance that we need from the silicon and getting our customers the experience that they expect from a product.

Harsh Banwait, CoreWeave

Security and performance must advance together across the infrastructure

Protecting agentic AI systems requires security mechanisms integrated into the hardware and software stack. Nvidia's Open Agent Safety Platform combines OpenShell runtime controls with Nvidia Sentry, which operates on BlueField-4 data processing units to provide independent monitoring and enforcement capabilities.

They're all in a single Vera Rubin tray, which includes Vera CPU as well as the DPU; that's all included in the same hardware infrastructure. So, Open Agent Safety Platform, the secure runtime, which is OpenShell, runs on the Vera CPU in that tray, and DOCA Sentry runs on the DPU part.

Hannah Coutand, Nvidia

Deploying this infrastructure at scale demands careful attention to power consumption, thermal management and operational automation. CoreWeave's engineering roadmap prioritizes networking, CPUs and storage as critical components that must scale in concert, Banwait explained.

https://www.youtube.com/embed/qTBDZljPmqQ?feature=oembed

All of that needs to be able to keep up. We're going to continue to focus on wherever the bottleneck is so that when the entire system is kind of advancing, it's doing that in one cohesive way. Otherwise, the weakest link in the chain kind of holds it all back.

Harsh Banwait, CoreWeave

Source: SiliconANGLE · Reporting supplemented by The Silicon Ledger staff.