NVIDIA’s GTC 2026 Bet Moves Beyond GPUs as Vera Reaches Four Partners

Jensen Huang’s March 16 GTC 2026 keynote presented NVIDIA’s next phase as an infrastructure business organized around producing AI tokens, not merely selling faster GPUs. The clearest subsequent change came in May: first Vera CPU systems reached four partners, moving one component of the strategy from an onstage announcement into customer evaluation.
That progress does not mean the entire “AI factory” stack is broadly deployed. The keynote’s lasting significance is instead the operating model it proposed: judge infrastructure by useful token throughput, response speed and cost under a fixed power budget, then coordinate CPUs, GPUs, networking, storage and software as one production system.
The keynote made the system—not the GPU—the product
At GTC 2026 in San Jose, Huang described a platform whose components are designed together around inference and agentic workloads. NVIDIA’s official keynote record dates the address to March 16 and identifies Vera Rubin as a full-stack platform comprising seven chips, five rack-scale systems and one supercomputer, alongside the Vera CPU, Rubin GPUs, BlueField-4 infrastructure and associated networking and software.
This is more consequential than a conventional generation-to-generation accelerator launch. NVIDIA is asking buyers to evaluate an integrated production line: the processor running a model, the host CPU coordinating tools, the interconnect moving data, the storage layer supplying context and the serving software scheduling inference. Performance at the level of one chip matters, but it no longer represents the whole commercial proposition.
The strategy also tightens NVIDIA’s control over system design. A customer can still encounter an open ecosystem of cloud providers, server manufacturers and software frameworks, yet more of the reference architecture originates with NVIDIA. For buyers, the relevant question becomes whether gains from integration outweigh reduced freedom to substitute individual components.
“Token factory” is an economic model, not a new name for every data center
Huang treated tokens as the output of AI infrastructure and separated three measurements: throughput per watt, interactive speed per user and cost per token. EE Times’ technical account of the keynote confirms that framing and reports NVIDIA’s claim of up to 50 times higher performance per watt for GB300 than Hopper under the comparison presented onstage.
The qualification matters. A vendor benchmark is evidence for a specified configuration and workload, not proof that every model or application will obtain the same multiplier. Model architecture, precision, batch size, latency target, memory pressure, networking and software optimization can all change realized throughput. “Tokens per watt” is therefore useful only when the tokens perform comparable work under disclosed service-level conditions.
The factory metaphor is most applicable to organizations operating inference continuously and at scale. A cloud provider serving many models has reason to optimize utilization across racks; an enterprise running a smaller, intermittent workload may care more about deployment simplicity, data governance or total ownership cost. Calling both environments factories does not make their purchasing criteria identical.
Vera gives agentic workloads a dedicated CPU layer
One of the keynote’s most concrete departures from a GPU-only narrative was Vera, NVIDIA’s first custom CPU. Agentic systems do not spend all their time inside neural-network calculations: they retrieve context, call tools, execute code, manage sandboxes and coordinate multiple tasks. Those operations create CPU, memory and data-movement demands even when GPUs perform the model inference.
On May 18, NVIDIA documented the first Vera system handoffs to Anthropic, OpenAI, SpaceXAI and Oracle Cloud Infrastructure. The company lists 88 custom Olympus cores and 1.2 TB/s of memory bandwidth, while describing SpaceXAI’s system as being evaluated for reinforcement-learning and agent-based simulation workloads.
This is meaningful post-keynote progress, but its scope should remain precise. Delivery of early systems to four organizations demonstrates working hardware outside NVIDIA’s labs; it does not establish general availability, fleet-scale adoption or independently measured application performance. The next evidence to watch is repeatable workload data and the transition from evaluation systems to production deployments.
Physical AI extends the same stack into machines
The keynote carried the factory idea beyond data-center inference into robotics, autonomous vehicles and industrial systems. NVIDIA’s pitch links model development with simulation, training, deployment at the edge and real-world control. In that formulation, Omniverse and physics simulation help create and test environments, while robotics models and Jetson-class computers support machines operating outside the data center.
This part of the strategy expands NVIDIA’s addressable workload, but it also introduces constraints that token throughput alone cannot capture. A physical system must meet requirements for sensing, control latency, reliability and safety in an unpredictable environment. A faster model response can be valuable without being sufficient for a safe robot or vehicle.
The live Olaf demonstration illustrated the breadth of the proposed stack, combining simulation and onboard computing in a moving character. It was an effective demonstration of integration, not evidence that unrelated industrial robots are ready for unsupervised deployment. Production claims still need to be evaluated in the context of a specific machine, task and operating environment.
What infrastructure buyers should take from GTC 2026
The keynote’s strongest idea is that AI capacity should be evaluated as an operational system. Buyers comparing platforms need measurements from their own models and service targets, including sustained throughput, latency distribution, power draw, utilization and software effort. A headline accelerator figure cannot reveal whether CPUs, memory, networking or orchestration become the bottleneck after deployment.
Procurement teams should also distinguish the stages behind NVIDIA’s announcements. A platform description, an early partner delivery, a cloud preview and a generally purchasable production service are different statuses. As of August 14, 2026, Vera’s documented partner handoffs provide a tangible update, while broader conclusions about adoption require later deployment evidence.
GTC 2026 ultimately recast NVIDIA’s competitive unit from a processor to an engineered production line for inference, agents and physical AI. Vera’s arrival at four partners gives that thesis its first post-keynote checkpoint. The unanswered question is no longer whether NVIDIA can describe an AI factory, but whether customers can reproduce its promised economics across real workloads without surrendering too much flexibility.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.