Nvidia used its GTC 2026 keynote in San Jose to unveil Vera Rubin, a five-rack AI platform that pairs its own GPUs with Groq’s low-latency processors, and to raise its revenue projection to $1 trillion through 2027. The Data Center Knowledge report on the March keynote captures the headline numbers: a 350x token-generation jump on a hypothetical 1 GW factory, a $300 billion annual revenue opportunity, and a roadmap that extends to orbital data centers. But the real news is quieter and more consequential. Vera Rubin is Nvidia’s first public admission that its GPUs are not the answer for every AI workload.
That admission arrives in the form of a licensing deal. Nvidia signed a deal with Groq in December 2025, and Vera Rubin is the first product to integrate Groq’s LPU processors. The platform’s third rack is a Groq 3 LPX inference accelerator rack with 256 LPU processors, designed for low-latency, large-context agentic systems. Nvidia could have kept this integration quiet. Instead, Huang put Groq on stage, in the rack, and in the revenue model. The $300 billion annual opportunity he cited is explicitly tied to combining Vera Rubin racks with Groq LPX racks. That is not a partnership of convenience. It is a strategic surrender on the claim that one architecture serves all of AI.
The five-rack design is a disaggregation play, and analysts noticed. Matt Kimball of Moor Insights & Strategy told Data Center Knowledge that Nvidia is “quietly acknowledging that their GPUs are not the answer for every single workload out there, especially with agentic AI.” Karl Freund of Cambrian AI Research called the Groq integration significant but framed the broader shift as positioning for agentic AI leadership. The architecture itself tells the story: NVL72 GPU racks for pretraining, Vera CPU racks for tool calling and reinforcement learning, Groq LPX racks for low-latency decode, BlueField-4 DPU racks for storage, and Spectrum-6 SPX for rack-to-rack networking. Each rack is a bet that a different kind of silicon handles a different phase of AI better than a monolithic GPU.
The phase model matters here. Ian Buck, Nvidia’s vice president of hyperscale and HPC, described four phases of AI that Vera Rubin accelerates: pretraining, post-training fine-tuning, test-time scaling, and a new phase he called “agentic scaling.” The last one is new, and it is where the Groq racks earn their place. Agentic scaling means AI systems interacting with other AI systems and tools, which requires low latency and large context windows. Nvidia’s GPUs are excellent at massive parallel compute. They are less excellent at the sequential, latency-sensitive token generation that agentic workflows demand. Groq’s LPUs, built for exactly that, fill the gap.
Huang’s keynote rhetoric leaned heavily on inference economics. “Finally, AI is able to do productive work, and therefore the inflection point of inference has arrived,” he said. “AI now has to think. In order to think, it has to inference.” He claimed computing demand has jumped 10,000 times over two years, with usage up roughly 100 times. He cited OpenAI and Anthropic by name, saying both would generate more tokens and more revenue if they had more capacity. The framing is consistent with Nvidia’s broader narrative: training was phase one, inference is phase two, and the economics of inference are what justify the $1 trillion projection.
The numbers deserve scrutiny. Huang said a 1 GW AI factory running Hopper generates roughly 2 million tokens per second, while a Vera Rubin system reaches about 700 million tokens per second. That is a 350x increase. He also claimed the NVL72 GPU racks deliver 10x higher inference throughput per watt at one-tenth the cost per token compared with Blackwell, and can train models with one-quarter the GPUs. The Vera CPUs, successors to Grace, offer twice the energy efficiency and three times the memory bandwidth per core versus x86. The BlueField-4 STX storage racks deliver four times the performance per watt. These are vendor claims, presented without independent verification, and they should be read as targets rather than measured results.
The revenue story is where the skepticism should land hardest. Huang’s $300 billion annual opportunity assumes buyers will pay for a new ultra-tier of AI service. Futurum Group analyst Brendan Burke did the math for Data Center Knowledge: the implied ultra-tier price of $150 per million tokens is a 50x jump over the $3 per million tokens medium tier. “The company remains dependent on application-layer customers to justify the 50x token price increase that Jensen set out as the goal of the combined Rubin and Groq systems,” Burke said. That is a polite way of saying the entire financial thesis rests on AI companies finding customers willing to pay 50 times more for premium tokens. Nvidia can build the racks. It cannot build the demand.
The orbital data center vision is the most speculative piece. Huang disclosed early work on a Vera Rubin Space Module powered by the Rubin GPU, and acknowledged the cooling problem: no conduction, no convection, only radiation in space. Freund called the ambitions “aspirational.” That is charitable. Terrestrial data centers struggle with power and cooling today; space adds launch costs, radiation hardening, and latency that physics will not negotiate. The satellite presence Nvidia already has is real, but a full orbital data center is a decade away at best, and the energy argument cuts both ways. Solar power in space is abundant, but beaming that compute value back to Earth is not a solved problem.
What Vera Rubin means for AI builders is more immediate. The platform signals that heterogeneous infrastructure is becoming the default for serious AI work. A builder planning agentic systems in 2027 will not buy one architecture. They will match workloads to processors: GPUs for pretraining and dense compute, CPUs for tool calling and code compilation, LPUs for low-latency inference, DPUs for storage, and Ethernet or InfiniBand for the fabric. Nvidia is selling the whole stack, but the stack itself is a recognition that no single chip wins every phase.
The Groq deal is the tell. Nvidia spent years building CUDA into a moat, and it still licenses a competitor’s processor for its flagship platform. That is not a hedge. It is an acknowledgment that the agentic AI wave has different compute requirements, and that Nvidia would rather own the platform than lose the customer. The $1 trillion target depends on that platform story holding together, which depends on application-layer customers accepting the 50x token premium. Huang’s keynote was confident. The architecture underneath it is a bet that Nvidia’s customers can sell that premium to their own customers.
The most telling line in the keynote was not about tokens or racks. It was Huang describing the feeling every startup has, the feeling OpenAI has, the feeling Anthropic has: if they could just get more capacity, they could generate more tokens and their revenues would go up. That is the bet Vera Rubin is making. Nvidia is assuming the capacity constraint is the binding one, and that the market for tokens is elastic enough to absorb a 50x price increase at the top tier. If the constraint is actually demand, not capacity, then the $300 billion opportunity shrinks to a rounding error. The racks will ship either way. The question is whether anyone will pay for the ultra-tier.