“Olympus Cores Delivering The Best Performance Ever Seen On ARM.”
That was the Phoronix headline. It is accurate. It is also the least important sentence on the page.
Read what the benchmark actually settles, and what it does not. Lot’s of technical specs, sorry in advance: Phoronix measured NVIDIA’s new Vera CPU at 1.63x its own Grace predecessor on the geometric mean, 1.55x Intel’s 128-core Xeon 6980P, and roughly 10% ahead of AMD’s EPYC 9575F. The silicon underneath: 88 custom Armv9.2 Olympus cores, 176 threads via spatial multi-threading, FP8 support, 2MB of L2 per core and a 164MB unified L3 (double the L2 of Grace and a far larger L3), LPDDR5X delivering up to 1.2TB/s of bandwidth, PCIe Gen 6 and CXL 3.1, in a 450W socket that lands near 500W with memory, broadly comparable to the x86 parts it is built to displace.
Those numbers are real. They are not the point.
The point is which socket just changed hands. For years NVIDIA has been moving the unit of competition from the chip to the rack: it builds the GPU, the networking through NVLink and Spectrum, and the coherent CPU-GPU link through NVLink-C2C. One socket was still a merchant part, bought from Intel or AMD: the head-node CPU. Vera is that socket. With Vera, the last major merchant socket moves in-house.
The benchmark is the distraction. The ownership map is the story.
The CPU Stopped Being a Supporting Actor
The head-node CPU only became worth winning because agentic AI made it worth winning.
The AI narrative was GPUs: more FLOPS, more HBM, more NVLink. The CPU booted the server, coordinated the system, and stayed out of the narrative. Agentic workloads broke that. The orchestration between GPU calls now lands on the CPU, and at GTC Jensen put a number on it: 12,000 GPUs require 400,000 CPU cores for agentic AI and reinforcement learning, a 33-to-1 core ratio.
The clearest statement of why the socket matters came from the person whose company licenses the cores underneath it. In The Agentic CPU, Arm’s Mohamed Awad told me after the keynote: “The GPUs are gonna generate the tokens, the CPUs are gonna figure out what to do with them. As more and more tokens get generated, you need more and more CPUs to handle it.” That orchestration socket used to be rented from x86. Vera ends the lease.
Vera Closes the Loop
The reframe paid readers carry from here is The Last x86 Island: the head-node CPU was the final major merchant socket in NVIDIA’s AI rack, and Vera takes it in-house.
This is the layer NVIDIA spent a decade assembling everywhere else. In NVIDIA’s Inference Stack Depth Strategy: “They sell chips; NVIDIA sells stack depth.” That line was written about the GPU. It now applies one socket lower. A merchant vendor sells a CPU. NVIDIA sells a CPU designed knowing the exact GPU it feeds, the exact fabric it talks across, and the exact runtime it schedules for. Vera exists to feed the Rubin GPU over NVLink-C2C. That is not a spec-sheet comparison. That is a different product category.
The stack-depth tell is in the boring infrastructure. NVIDIA upstreamed Olympus compiler support into GCC and LLVM Clang more than a year before the chip ships. Kernel support is already upstream. Vera boots on stock ARM64 distributions like Ubuntu and Fedora through ACPI, with no Device Tree headaches. That is what owning the stack looks like two layers below the headline.

The table is the whole argument in one frame. Four CPUs, one power envelope, one benchmark sweep. Vera leads each measured geomean. But the column that matters is the last one: NVIDIA Vera and Grace are in-house silicon; AMD EPYC and Intel Xeon are the merchant parts being displaced.
The benchmark column will be stale in a quarter. The ownership column will not.
The benchmark tells you Vera is fast. The ownership map tells you why AMD and Intel should care. Below the divider: why winning the socket beats buying it, the perf-per-watt math that turns a CPU spec into a cost-per-token line, the three bears ranked by what actually changes the thesis (including the one chip that could retake the raw lead this year), and the read across NVDA, AMD, INTC, and ARM.

Leave a Reply