Nvidia’s AI advantage is moving beyond the GPU

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane August 29, 2026 3 min read
Nvidia’s AI advantage is moving beyond the GPU

Nvidia’s market cap growth slowed after a decade of expansion, with shares showing a more modest trajectory in the year following a tenfold increase between early 2023 and mid-2025.

Investors worry about the durability of Nvidia’s position as hyperscalers like Amazon and Google build their own chips. However, the company’s advantage now extends well beyond its graphics processing units.

As AI compute requirements reach gigawatt scales, managing the systems that surround the GPU has become a complex task. Nvidia has built the hardware needed to handle this orchestration, securing a significant lead even as competition for the chips themselves intensifies.

Operating a megascale data centre at peak efficiency remains extremely difficult. This challenge grows as deployments become larger and faster.

Rack by Rack

Nvidia is currently rolling out its Vera Rubin architecture. This system pairs the Rubin GPU with other units, including the Vera CPU, the Groq 3 LPX inference accelerator, and similar racks for storage and networking.

Discussions with Nvidia staff reveal that these systems are highly specialised. They do not process tokens like the Rubin GPU. Instead, they ensure everything outside the GPU operates as efficiently as possible. If the GPU is the engine, these are the rest of the car.

The Vera CPU focuses on orchestrating data. “Vera is important because there’s only so much memory that you can put in a single server or any sort of compute platform,” Jason Hardy, Nvidia’s VP of storage technology, told me.

As data centres scale computing power, memory capacity has also increased. Companies like Micron have profited from this second wave of the infrastructure boom. Moving that data to the GPU at the right time is not straightforward. Companies driving tokens-per-watt lower are realising how crucial traffic direction is.

“We saw upwards of 3x improvement in these operations, where the Vera CPU is allowing for acceleration,” Hardy said. “So now we can use our flash to its fullest potential, because we can get all that performance out of it without bottlenecking.”

Similar problems exist outside Nvidia. When OpenAI developed its Jalapeño chip, a major focus was avoiding these challenges by minimising the amount of data that needs to be moved around.

“We designed Jalapeño to minimize data movement and communication delays,” the company said in a blog post earlier this month. “Its large domain allows the entire workload to remain within one connected system, minimizing data movement and helping the complete request stay fast and efficient from beginning to end.”

This is a different approach. It avoids data movement entirely by conducting a workload within one integrated chip. The overall logic remains the same: increasing efficiency with smarter traffic control instead of just more processor cycles. This opens a new layer of infrastructure for companies to compete over.

This new focus on data orchestration is not automatically a win for Nvidia. The company must compete with rival chipmakers and hyperscalers just as it has with GPUs. But the competition has moved to a new layer, where building a rival GPU matters less than being able to make the entire system work efficiently.

At least in the early stages, Nvidia looks to have a commanding lead.

What it means

For people making things with AI, the bottleneck is shifting. Previously, the focus was on raw processing power. Now, the constraint is how quickly data can reach the processor. Tools that manage this flow, like Nvidia’s Vera CPU or OpenAI’s integrated Jalapeño design, will determine how fast and efficient creative workflows remain as models grow larger.

Scroll to Top