
NVIDIA spent fifteen years convincing the world that the GPU was the computer. Then, at GTC Taipei on May 31, 2026, it announced a CPU and called it the most important chip in the AI factory. The NVIDIA Vera is the company's first processor designed from scratch for AI agents, and it's now in full production, shipping inside both standalone servers and NVIDIA's flagship Vera Rubin AI racks.
This is worth understanding even if you never buy a server, because Vera reveals where NVIDIA thinks computing is going: a future where the biggest users of processors aren't people at all, but software agents working in loops. Here's what Vera is, what it claims to do, and how seriously to take those claims.
Key takeaways
- Vera is NVIDIA's first CPU purpose-built for AI agents: 88 custom "Olympus" Arm-based cores, announced May 31, 2026, now in full production.
- NVIDIA claims 1.8x faster task completion than x86 CPUs on agentic workloads, 50% higher instructions per cycle than its Grace predecessor, and 1.2 TB/s of memory bandwidth. These are NVIDIA's own figures, not independent benchmarks.
- It ships two ways: as a standalone server CPU from Dell, HPE, Lenovo, Supermicro and others, and paired with Rubin GPUs in the Vera Rubin NVL72 platform via 1.8 TB/s NVLink-C2C.
- Early adopters named by NVIDIA include Anthropic, OpenAI, CoreWeave, Oracle Cloud, and Perplexity, though OpenAI chose AMD CPUs over Vera for its own Jalapeño chip on maturity grounds.
- Vera's successor is already on the roadmap: "Rosa," based on a new Rigel core.
Why NVIDIA built a CPU for agents
Jensen Huang's pitch line at the launch was blunt: "AI agents will be the largest users of computing." The logic behind Vera starts from how agents actually work. A chatbot generates text in one pass. An agent reasons, writes code, executes it, calls a tool, queries a database, checks the result, and loops, dozens of times per task. The GPU handles the model's thinking; everything between the thinking, the doing, is CPU work.
Worse, that CPU work is mostly serialized. You can't run step five of a tool chain before step four finishes. So while data center CPUs spent the last decade maximizing core counts for cloud rental economics, agents need the opposite: strong single-threaded performance, predictable latency under load, and enormous memory bandwidth so no core ever starves waiting for data. NVIDIA's argument is that conventional server CPUs were optimized for the wrong thing, and Vera is the correction.
The Olympus core: built for single-threaded speed
Vera's compute comes from 88 custom "Olympus" cores, NVIDIA's own Armv9.2-based design. The headline architectural claim is 50% higher instructions per cycle (IPC) than Grace, NVIDIA's previous CPU. IPC measures how much work a core does per clock tick, so a 50% gain means dramatically faster execution of the sequential code that dominates agent loops: Python runtimes, JavaScript execution, compilation, database queries.
The memory subsystem is arguably more important than the cores. Vera delivers up to 1.2 TB/s of LPDDR5X memory bandwidth, and its monolithic compute die provides 3.4 TB/s of core-to-core bandwidth, which NVIDIA says is three times greater than any other data center CPU. The design goal is consistency: "every core completes the agent task at full performance without other cores slowing it down," as NVIDIA puts it. For agents, where one slow step stalls the whole loop, that predictability matters more than peak throughput.

The 1.8x claim, labeled honestly
NVIDIA's headline performance figure is that Vera enables "1.8x faster task completion compared with x86 CPUs" across agentic AI, reinforcement learning, and data processing workloads. This number deserves careful handling, because it's doing a lot of marketing work.
First, it's NVIDIA's own claim, published in its launch materials and press release, measured on workloads NVIDIA chose. No independent third-party benchmark of shipping Vera silicon against current x86 server CPUs had been published as of early October 2026. Second, the figure is workload-specific: it covers agentic-path tasks like code execution and data processing, not general-purpose server workloads or single-threaded desktop performance. Third, there's a corroborating data point from a customer: Perplexity's infrastructure VP told Reuters that Vera completed AI coding workloads around 1.5 times faster than conventional CPUs, which is in the same neighborhood but not the same as 1.8x.
None of this means the claim is false. It means it's unverified vendor marketing until independent labs test production silicon, which is the standard you should apply to every chipmaker's launch numbers, including AMD's and Intel's.
Two ways to buy it: standalone or Rubin-paired
Vera ships in two configurations. As a standalone CPU, it's going into two-socket servers from Dell, HPE, Lenovo, and Supermicro, plus Taiwan system builders including ASUS, Compal, Foxconn, GIGABYTE, Pegatron, QCT, Wistron, and Wiwynn. These are for customers who want the agent-optimized CPU without buying into NVIDIA's full GPU stack.
The more strategic configuration is the Vera Rubin NVL72 rack: 36 Vera CPUs paired with 72 Rubin GPUs in a single liquid-cooled system. The pairing uses NVLink-C2C, a coherent interconnect running at up to 1.8 TB/s, which merges the CPU's 1.5 TB of LPDDR5X memory and the GPU's HBM4 into a single addressable pool. In NVIDIA's framing, the rack stops being a collection of components and becomes one machine, purpose-built for what it calls "agentic inference": maintaining vast inference context memory while agents reason through multi-step tasks.
Who is actually buying Vera
NVIDIA's launch announcement named an unusually specific customer list. On the AI lab side: Anthropic, OpenAI, and SpaceXAI are "planning to adopt" Vera. On the hyperscaler side: ByteDance, CoreWeave, Lambda, Nebius, Nscale, and Oracle Cloud Infrastructure. Finance made an appearance too, with the NYSE named as a customer exploring the chip. "Planning to adopt" and "exploring" are doing the usual launch-PR hedging, so treat the list as intent rather than purchase orders.
The most telling customer signal may be the one that went the other way. OpenAI revealed at Hot Chips 2026 that its custom Jalapeño inference chip is paired with AMD EPYC "Turin" CPUs, not Vera. OpenAI's hardware chief, Richard Ho, said NVIDIA's Vera standalone was "a little bit behind on that maturity level" for OpenAI's scale. That's a polite way of saying Vera is brand new and unproven at hyperscale, which is true of every first-generation architecture and worth remembering alongside the benchmark claims.
For the Canadian angle: none of this changes your cloud bill tomorrow, but it shapes what compute looks like in 2027. Canada's National AI Council is tasked with compute-access questions, and the skills that matter are shifting toward infrastructure: see the AI skills rundown for the Canadian job market. If you're deploying agents rather than buying chips, the security implications are the more urgent read: AI coding agents went rogue this summer.

What comes next: Rosa
NVIDIA has already named Vera's successor: Rosa, based on a new Rigel core. That tells you two things. First, NVIDIA now treats CPUs as a roadmap business, not a one-off experiment, which is a genuine strategic shift for a company that built its empire on GPUs. Second, the CPU fight is going to be a multi-generational war: AMD's Venice is shipping on 2nm now, Intel's Xeon 7 lands in 2027, and NVIDIA is already talking about what's after Vera.
Practical next steps
- If you evaluate server CPUs: wait for independent Vera benchmarks on your actual workloads before treating the 1.8x figure as planning data; vendor numbers are directional, not contractual.
- If you run AI agents in production: profile where your agent loops actually spend time. If it's tool execution and data munging between model calls, CPU selection matters more than your GPU choice.
- If you buy cloud compute: ask providers when Vera-based instances land on their roadmaps, and compare against AMD Venice-based options arriving in Q4 2026.
- If you're learning the stack: the NVLink-C2C unified-memory model is the architectural idea to understand; it's the template every vendor is converging on.
The bottom line
The NVIDIA Vera is a bet that the most valuable computer of the next decade is the one running the agent loop, not the one running the model. The 88-core Olympus design, the 1.2 TB/s of memory bandwidth, and the 1.8x task-completion claim (NVIDIA's number, pending independent verification) all serve that thesis. Whether Vera wins its category will be decided by shipping silicon, independent benchmarks, and hyperscale deployments through 2027, not by launch-day slides. But the strategic signal is already clear: the CPU is no longer the GPU's sidekick, and NVIDIA intends to own both.
Sources
- https://nvidianews.nvidia.com/news/nvidia-unveils-vera-the-cpu-for-agents
- https://www.globenewswire.com/news-release/2026/06/01/3303981/0/en/NVIDIA-Unveils-Vera-the-CPU-for-Agents.html
- https://www.techspot.com/news/111712-nvidia-unveils-vera-88-core-arm-cpu-ai.html
- https://www.outlookbusiness.com/deeptech/why-perplexity-is-choosing-nvidias-vera-cpu-over-traditional-intel-and-amd-chips
- https://aiweekly.co/alerts/openai-pairs-jalapeo-asic-with-amd-turin-skips-nvidia-vera
- https://www.techtimes.com/articles/320933/20260718/nvidia-vera-rubin-cuts-post-training-token-costs-seven-chip-codesign.htm
Quick answers
Frequently asked questions
01
What is the NVIDIA Vera CPU?
Vera is NVIDIA's first CPU designed specifically for AI agents, announced at GTC Taipei on May 31, 2026, and now in full production. It features 88 custom "Olympus" Arm-based cores optimized for the serialized, single-threaded work that AI agents do between GPU inference steps: running code, calling tools, and processing data.
02
How fast is the NVIDIA Vera CPU?
NVIDIA claims Vera completes tasks 1.8x faster than x86 CPUs in agentic workloads, delivers 50% higher instructions per cycle than its previous Grace CPU, and provides up to 1.2 TB/s of memory bandwidth. These are NVIDIA's own figures from its launch materials, not independently verified benchmarks.
03
Is the Vera CPU available to buy?
Yes, in server form. Standalone Vera CPU systems are being built by Dell, HPE, Lenovo, Supermicro, and others, and Vera also ships paired with Rubin GPUs in NVIDIA's Vera Rubin NVL72 rack platform. It's a data center product, not a consumer chip.
04
What is the difference between Vera and Grace?
Grace was NVIDIA's first Arm data center CPU, a general-purpose host processor for GPU servers with nearly 2.5 million shipments. Vera is its successor, redesigned around agentic AI: stronger per-core performance, far more memory bandwidth (1.2 TB/s), and a monolithic compute die with 3.4 TB/s of core-to-core bandwidth.
05
Which companies are using NVIDIA Vera?
NVIDIA names Anthropic, OpenAI, and SpaceXAI among AI labs exploring Vera, plus hyperscalers ByteDance, CoreWeave, Lambda, Nebius, Nscale, and Oracle Cloud. Perplexity has said publicly it found Vera about 1.5x faster than conventional CPUs for AI coding workloads.
06
What is NVLink-C2C in the Vera Rubin platform?
NVLink-C2C is the coherent chip-to-chip interconnect linking the Vera CPU to the Rubin GPU, running at up to 1.8 TB/s. It lets the CPU's LPDDR5X memory and the GPU's HBM4 function as a unified memory pool, eliminating the PCIe boundary between them.



