Nvidia Reframes AI Infrastructure Around Token Economics and Power Efficiency
As AI systems grow more complex, the economics of data centers are shifting from individual chip performance to system-wide efficiency measured in tokens per watt, according to Nvidia executives.

The financial viability of artificial intelligence infrastructure now extends well beyond simply acquiring powerful graphics processors. When agentic systems leverage multiple models, databases and tools in concert, the data center itself must function as a unified computing platform. This evolution is redirecting focus away from isolated chip performance toward the broader infrastructure ecosystem that converts raw computing power into actionable intelligence. Networking, storage, processors and software must work together at scale while maximizing the intelligence produced from each unit of electrical power, according to Ian Buck, vice president and general manager of hyperscale and HPC at Nvidia Corp.
Instead of cars or devices or PCs, it's tokens. These assets are not IT; they're not cost. They're actually appreciating, revenue-generating, fungible, durable, productive parts of an economy.
Ian Buck, Nvidia
Buck discussed these shifts in AI infrastructure economics during the Fully Connected event, speaking with theCUBE Research's Dave Vellante and John Furrier in an exclusive broadcast on theCUBE, SiliconANGLE Media's livestreaming studio.
Inference reshapes AI factory economics
The commercial value generated by an AI factory stems from inference, where deployed models handle incoming requests and generate tokens. However, inference does not eliminate the need for training, since organizations must continuously refresh their deployed models as new data emerges and market conditions shift, according to Buck.
It's not just fire and forget on all these services. As companies are using these models, they're refining them, they're aligning them, they're adding more data to them. Having them up to date and aware — that actually is a little bit of training. We're seeing the work in reinforcement learning and online alignment.
Ian Buck, Nvidia
Latency considerations introduce a separate economic dimension for workloads where faster response times command premium value. Nvidia's Groq 3 LPX inference accelerator pairs with its Vera Rubin platform to boost per-user token throughput for time-critical applications.
If there's value in those tokens to have the fastest possible thinking, LPX can be boosted on top of Vera Rubin to make that possible. We're seeing a lot of interest in areas like fintech and other areas where things are happening in real time.
Ian Buck, Nvidia
Power makes efficiency a system-level priority
The electrical capacity available to a data center ultimately determines the maximum volume of computing infrastructure that facility can support. This physical constraint elevates tokens per watt to a critical metric for AI factory performance and compels hardware manufacturers to deliver substantial efficiency gains with each new generation.
Data centers have a natural cap, and that cap is actually their power. With every generation of GPU, we make sure that our tokens per watt is upwards of 10 times more efficient. In fact, we saw that with Blackwell — we got, in the end, a 30x improvement in tokens per watt.
Ian Buck, Nvidia
This reorientation also affects the architectural level at which infrastructure must be designed and managed. CoreWeave enables customers to select specific configurations or leverage higher-level inference services that fine-tune the relationship between throughput and token latency, according to Buck.
https://www.youtube.com/embed/64vTNUuQaeQ?feature=oembed
CoreWeave can do that for customers. They don't have to feel overwhelmed by all the choices. That's where our partner ecosystem is so important.
Ian Buck, Nvidia