Nvidia has begun shipping its Vera Rubin NVL72 rack-scale AI systems to major cloud providers, delivering a dramatic leap in how much artificial intelligence work each watt of electricity can accomplish. CoreWeave, the first cloud operator to publicly validate the platform, reports that a single Vera Rubin rack can churn out ten times more reasoning tokens per second for every megawatt of power compared to the previous Grace Blackwell generation.

The milestone, confirmed on July 21, moves the conversation beyond chip speed alone. For the first time, Nvidia is explicitly pitching an integrated rack—combining custom CPUs, GPUs, networking, and liquid cooling—as the atomic unit of AI deployment. Systems are now running at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius, with a manufacturing footprint spanning more than 350 factories across 30 countries.

What Exactly Is Vera Rubin NVL72?

Vera Rubin NVL72 is not a graphics card. It’s a pre-assembled rack-scale platform that packs 72 Rubin GPUs and 36 Vera CPUs into a single cabinet, all linked by a sixth-generation NVLink fabric capable of moving data at 260 terabytes per second. The rack also includes ConnectX-9 SuperNICs for 1.6 Tb/s of per-GPU network throughput, BlueField-4 data processing units, and Spectrum-6 Ethernet switches—all designed to work together from the start.

Nvidia calls this “extreme co-design.” The Vera CPU handles agentic reasoning and data coordination while the Rubin GPU tackles heavy matrix math with its HBM4 memory and 50 petaflop NVFP4 Transformer Engine. By engineering every piece as a system, Nvidia can squeeze out efficiencies that a pieced-together cluster would miss. The result: a single Vera Rubin NVL72 rack does the work that would have required multiple racks of the previous generation—at a fraction of the electricity.

The Efficiency Breakthrough: Tokens per Megawatt

The headline benchmark comes from CoreWeave, which tested Vera Rubin against Grace Blackwell NVL72 on the DeepSeek-R1 reasoning model. Under matched interactivity targets—ensuring both systems delivered tokens to users at the same responsive speed—Vera Rubin generated approximately 10 times more tokens per second per megawatt. That’s not just a theoretical number; it’s a real-world demonstration using Nvidia’s TensorRT-LLM software, NVFP4 precision, multi-token prediction, and disaggregated prefill and decode.

Nvidia also claims that for deep-reasoning workloads, the cost per million tokens can drop to one-tenth what it was on Grace Blackwell. These numbers are workload-dependent—they rely on a specific model, software stack, and precision format—but they illustrate a clear trend. When electricity is the limiting resource, an AI platform that can produce more useful output from the same power draw wins.

Why This Matters for Cloud AI Users

If you’re a developer or an IT decision-maker using cloud AI services, you may soon notice two changes: lower inference costs and the ability to handle more complex, interactive workloads.

Chatbots that simply answer one-sentence queries are giving way to agentic AI—reasoning agents that plan, call tools, inspect results, and revise their work over many steps. Each step generates a flood of intermediate tokens. A system that can produce those tokens faster and more efficiently means lower latency for users and lower bills for developers. Cloud providers, in turn, can pack more revenue-generating inference work into data centers that are already capped by available electricity.

For enterprises that run their own AI infrastructure, the implications are even sharper. Transitioning to rack-scale architectures can boost the value of every megawatt they’ve secured from a utility. But it also requires new data center designs: high-density power delivery, direct liquid cooling, and networking that can keep up with the 1.6 Tb/s data flows between racks. A Vera Rubin deployment is not a drop-in replacement for a few GPU servers; it’s a platform that demands holistic planning.

How We Got Here: From GPUs to AI Factories

Two years ago, an AI cluster was essentially a bunch of servers stuffed with GPUs. Performance was measured in FLOPS, and the primary bottleneck was how many accelerators you could buy. But as models grew larger and inference workloads became the main source of demand, the bottlenecks multiplied: memory bandwidth, network latency, power density, and cooling all started to throttle real-world performance.

The International Energy Agency projects that global data center electricity consumption will more than double by 2030, driven primarily by AI. In many regions, building a new data center can take years—not because of server supply, but because of grid interconnection, transformer lead times, and local permitting. That’s why “tokens per megawatt” has become an economic necessity. Vera Rubin is Nvidia’s answer to that constraint: a rack that radically improves the output of every watt before anyone needs to pour a new foundation.

What to Do Now: Preparing for the Vera Rubin Era

For cloud consumers: Keep an eye on pricing announcements from Azure, Google Cloud, and Oracle. As Vera Rubin instances become generally available, compare the cost per token for your most demanding agentic AI workloads. You may be able to process more reasoning cycles within the same budget—or scale down the number of instances you need.

For data center operators: Start scoping the requirements for dense AI racks. Even if you’re not deploying Vera Rubin tomorrow, power infrastructure with 100+ kW per cabinet and direct-to-chip liquid cooling are quickly becoming table stakes. Work with facilities teams to audit available power and cooling capacity, and begin piloting rack-scale integrations with your preferred hardware vendors.

For AI developers: Familiarize yourself with Nvidia’s inference stack. The impressive efficiency numbers come from a tight coupling of hardware, networking, and software. Using TensorRT-LLM, Dynamo, and NVFP4 precision lets you exploit the platform’s full capabilities. Test your models on early Vera Rubin instances through your cloud provider to see how they behave under the new architecture.

Outlook: The AI Efficiency Race Is On

Vera Rubin NVL72 is not a singular event; it’s the first unambiguous statement that AI infrastructure economics are no longer about who has the fastest chip. They’re about who can build the most efficient—and most deployable—rack-scale system. Nvidia will likely extend this design philosophy to future platforms, further tightening the integration between computation, communication, and cooling.

Competitors are racing to offer their own tightly coupled solutions, but for now, Nvidia’s head start in system-level co-design and manufacturing scale is formidable. The real test will come as Vera Rubin reaches broader general availability later this year, giving thousands of customers real-world experience with the platform. The 10x efficiency number will be validated—or nuanced—by a diversity of models and workloads. For anyone building or buying AI, the era of judging a data center by its GPU count alone is over.