Artificial Intelligence

NVIDIA Vera Rubin Ushers in the Gigascale Era for AI Infrastructure

NVIDIA’s latest generation of AI supercomputing, the Vera Rubin platform, is now entering full production, marking a significant leap forward in the scale and efficiency of artificial intelligence infrastructure. This revolutionary system, built from chip to grid, is designed to meet the escalating demands of AI development and deployment, particularly for the burgeoning field of agentic systems. Production is actively ramping up across key partners, including CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure, leveraging a vast, globally distributed supply chain spanning over 350 factory sites in 30 countries. This extensive network represents the largest and most mature rack-scale supply chain ever assembled to address the burgeoning compute needs of the AI industry.

The Vera Rubin platform is engineered to deliver unprecedented performance per watt and a dramatically reduced cost per token, a critical metric for the economic viability of large-scale AI operations. Early benchmarks are already showcasing its transformative capabilities. CoreWeave, a prominent AI cloud provider, reported a groundbreaking 10x increase in throughput per megawatt compared to the previous generation Grace Blackwell NVL72 system when running the DeepSeek-R1 benchmark. This dramatic improvement directly addresses the primary constraint for many "AI factories": power limitations.

Advancing Performance Through Extreme Codesign

The remarkable performance gains of the Vera Rubin platform are attributed to a philosophy of "extreme codesign," where seven distinct chips and five rack trays are engineered as a singular, cohesive system rather than being assembled from disparate components. This integrated approach encompasses the Vera Rubin NVL72 compute nodes, the Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX networking, and the Vera BlueField-4 STX network interface cards. This holistic design process ensures seamless interoperability and optimized performance across the entire system.

At the heart of this innovation lies the NVIDIA Vera CPU. Redefining the role of a CPU in AI factories, the Vera CPU is specifically architected for the demands of the "agent era." Its custom Olympus core delivers a substantial 2x increase in single-threaded performance, 3x improvement in core-to-core bandwidth, and a 40% reduction in memory latency compared to competing chiplet designs. This makes it exceptionally efficient for agentic workloads, which rely heavily on rapid, single-threaded operations and efficient data handling.

Accelerating AI Factories with Purpose-Built Networking

Networking is a crucial component of any large-scale computing infrastructure, and NVIDIA has equipped the Vera Rubin platform with cutting-edge networking solutions. The platform features the sixth generation of NVLink scale-up technology, which provides more than double the throughput for complex workloads, a threefold reduction in latency, and a tenfold increase in packet rates over standard Ethernet. For scale-out connectivity, the Spectrum-X Ethernet solution integrates 102.4 terabit per second (Tb/s) Spectrum-6 switch systems with 1.6 Tb/s ConnectX-9 SuperNICs. This advanced networking incorporates adaptive routing, sophisticated congestion control mechanisms, comprehensive telemetry, and support for open operating systems, enabling a 1.6x higher Remote Direct Memory Access (RDMA) bandwidth compared to off-the-shelf Ethernet solutions.

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

Leading AI infrastructure builders, including CoreWeave, Microsoft, SpaceXAI, and Tesla, are among the first to integrate Spectrum-6 switches into their AI factories. Further enhancing scale-out capabilities, NVIDIA’s Photonics technology, featuring co-packaged optics, represents the industry’s first volume-manufactured switch with integrated optics. This innovation slashes power consumption by fivefold and significantly increases Mean Time Between Failures (MTBF) compared to pluggable transceivers. Early adopters like CoreWeave, Lambda, and Oracle Cloud Infrastructure (OCI) are already leveraging this technology. Spectrum-XGS Ethernet extends performance across multiple sites, achieving 1.9x higher multi-site throughput, acknowledging that gigascale AI is a distributed challenge, not confined to a single facility.

NVLink Fusion further expands the NVIDIA infrastructure ecosystem by opening the platform to third-party XPUs (eXtreme Processing Units), providing partners with a streamlined path to market on the robust NVLink scale-up stack and its extensive ecosystem.

Operational Efficiency and Sustainability Gains

NVIDIA’s commitment to rack-scale codesign over three generations has resulted in the Vera Rubin NVL72 system being designed without any cables, fans, or hoses within the compute tray. This streamlined design drastically reduces compute tray assembly time from hours to a mere minute.

Moreover, the Vera Rubin platform incorporates a liquid cooling inlet temperature design of 45 degrees Celsius, enabling chiller-free dry-cooler operation. For new AI factories, this higher-temperature dry cooling, coupled with a closed-loop liquid cooling system, is projected to save millions of gallons of water per megawatt annually. This focus on efficiency and sustainability is increasingly critical as AI infrastructure expands globally.

NVIDIA Vera Rubin Powers Europe’s Open Model Era

In a significant development for the European AI landscape, the Vera Rubin platform is powering a newly expanded partnership between Microsoft and Mistral AI. This collaboration aims to bring frontier AI capabilities to the region, combining open European models with secure cloud and customer-controlled environments, thereby enabling governments and regulated industries to adopt advanced AI on their own terms.

A multibillion-dollar agreement underpins this partnership, focusing on expanding AI infrastructure within Europe. Mistral AI is augmenting its GPU capacity by deploying thousands of the latest NVIDIA Vera Rubin GPUs, significantly increasing AI compute availability for customers and providing a unified platform for training, inference, and large-scale AI deployments. The Vera Rubin NVL72, a rack-scale AI supercomputer that unifies seven newly codesigned chips into a single system, will be the backbone of this initiative, powering Mistral Compute and Microsoft’s European AI infrastructure with tens of thousands of GPUs.

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

The demand for sophisticated AI capabilities within Europe, governed by regional laws and data control requirements, is a driving force behind this expansion. The increasing complexity of agentic systems, which can consume up to 15 times more tokens than traditional AI applications, underscores the strategic importance of efficient infrastructure. Meeting these demands requires not only substantial computing power but also infrastructure that scales efficiently while adhering to regional mandates for data sovereignty, governance, resilience, and strategic autonomy.

The Vera Rubin platform provides the foundational computing power for Europe’s burgeoning open-model ecosystem. Its full-stack architecture integrates accelerated computing, advanced networking, and sophisticated software, ensuring that models and agents can operate efficiently from the initial training phase through to production deployment.

Open Models Built for Enterprise AI

The collaborative effort between Microsoft, Mistral, and NVIDIA is set to deliver sovereign-ready AI solutions accessible across public cloud, cloud-connected, and fully disconnected private cloud environments. This initiative pairs the flexibility of open models with the robust infrastructure necessary for their scaled operation. Mistral Medium 3.5 and OCR 4 models are now available within Microsoft Foundry, and Mistral models are integrated into Microsoft Copilot Studio. Through Azure Local and Foundry Local offerings, customers will have the ability to utilize the same models, tools, and operational patterns across both cloud and customer-controlled environments.

AI on Europe’s Terms

This infrastructure empowers government agencies to apply AI to sensitive workflows, allows healthcare and financial services organizations to deploy within strict regional compliance frameworks, and enables manufacturers to process data on-site for accelerated decision-making. When compared to NVIDIA’s GB200 NVL72, the Vera Rubin NVL72 delivers up to 10 times more tokens per megawatt and a tenfold reduction in cost per million tokens, offering enhanced intelligence within the same power footprint. The vision is clear: sovereign AI should not necessitate trade-offs between innovation, economic efficiency, and control. With Vera Rubin as the computing foundation, complemented by Microsoft and Mistral’s sovereign cloud infrastructure and open European models, Europe is poised to achieve advancements across all three critical areas.

NVIDIA Vera Rubin NVL72 on CoreWeave Demonstrates 10x More Tokens Per Megawatt Than Blackwell in Benchmark

CoreWeave, a leading AI cloud provider, has embraced NVIDIA’s extreme codesign philosophy, achieving a significant performance leap with the Vera Rubin NVL72 system. This close collaboration with partners like CoreWeave is pivotal in integrating new generations of NVIDIA accelerated computing into AI factories. Following extensive co-engineering efforts, CoreWeave has become the first AI cloud provider to successfully bring up and validate the Vera Rubin NVL72, and has now released its initial performance metrics from live hardware.

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

A DeepSeek-R1 benchmark conducted on the Vera Rubin NVL72 at CoreWeave revealed a remarkable 10x improvement in tokens per second per megawatt when compared to the Grace Blackwell NVL72. Tokens per megawatt is a crucial metric for determining the scalability and profitability of AI infrastructure. A higher tokens-per-megawatt ratio translates to more computational intelligence derived from a given power budget or the ability to execute the same workload using substantially less power. Enterprises and AI research labs, such as Jane Street, are set to leverage the Vera Rubin platform on the CoreWeave cloud to scale their AI operations.

Beating Bottlenecks with NVIDIA Spectrum-X

The DeepSeek R1 benchmark’s mixture-of-experts architecture necessitates extensive all-to-all GPU communication, as each token must be routed across distributed expert sub-networks. The Vera Rubin NVL72’s 260 TB/s all-to-all NVLink 6 fabric effectively eliminates this bottleneck, enabling the entire rack to function as a unified accelerator.

CoreWeave is among the pioneering adopters of the NVIDIA Spectrum-X Ethernet SN6600-LD as the switching fabric for its Vera Rubin NVL72 deployments. Built upon the 102.4 Tb/s Spectrum-6 switch chip and featuring a liquid-cooled design, CoreWeave is deploying dense switching racks that deliver 1.64 petabits per second (Pb/s) per rack, offering 100% more capacity than previous-generation air-cooled switches. Furthermore, CoreWeave provides a fully non-blocking, multi-plane, multi-rail spine and leaf fabric that connects Vera Rubin NVL72 GPUs without any oversubscription.

NVIDIA Vera Rubin NVL72 Drives New Google Cloud A5X Instance for Ineffable Intelligence

Google Cloud’s new A5X instance, powered by the NVIDIA Vera Rubin NVL72, is now operational and supporting Ineffable Intelligence, a London-based startup developing a new generation of "superlearner" systems. These systems are designed to learn continuously through experience, accelerating breakthroughs across diverse fields.

Unlike traditional AI approaches that rely on large language models trained on static datasets, Ineffable Intelligence’s agents learn directly from interaction with their environments. They leverage reinforcement learning, generating experience within continuously simulated, massively parallel environments and rapidly translating this experience into policy updates and evaluations. These tightly coupled learning loops place extreme demands on compute power, memory bandwidth, and interconnectivity, necessitating infrastructure capable of operating at immense scale with exceptionally low latency.

"The next era of research requires the next era of hardware," stated Lasse Espeholt, co-founder of Ineffable Intelligence. "We feel privileged to work with the teams at NVIDIA and Google Cloud, who were able to grant us early access to Vera Rubin. The support across both teams has been unmatched; we were up and running almost immediately and are already testing infra for our superlearners."

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

The NVIDIA Vera Rubin NVL72 is purpose-built for this type of agentic training, offering predictable latency, high utilization, and significantly greater intelligence per dollar compared to previous-generation systems, making it an ideal platform for large-scale reinforcement learning.

Google Cloud A5X instances, announced at Google Cloud Next, are bare-metal configurations built on NVIDIA Vera Rubin NVL72 rack-scale systems. These instances promise up to a tenfold reduction in inference cost per token and a tenfold increase in token throughput per megawatt compared to the prior generation. The A5X instances leverage NVIDIA ConnectX-9 SuperNICs in conjunction with next-generation Google Virgo networking. This combination enables clusters that can scale to tens of thousands of NVIDIA Rubin GPUs within a single site and extend to nearly a million GPUs across multisite configurations. This provides customers with a unified, AI-optimized stack for training, tuning, and serving frontier, open, agentic, and physical AI models, while simultaneously optimizing for performance, cost, and sustainability. This advanced infrastructure is poised to unlock the next generation of reinforcement learning systems, driving breakthroughs in superlearning and superintelligence.

NVIDIA Vera CPU Doubles Orchestration Speed for DeepInfra AI Cloud

Independent benchmark results from DeepInfra highlight the performance capabilities of the NVIDIA Vera CPU, demonstrating it to be more than twice as fast and capable of supporting a greater number of concurrent AI agents compared to other CPUs. DeepInfra, an early participant in the NVIDIA open AI ecosystem, conducted these benchmarks using its production AI agent infrastructure. The cloud platform processes nearly 5 trillion tokens weekly, with approximately 30% of this volume driven by agentic systems, underscoring its focus on high-throughput AI inference.

The benchmarks indicate that the Vera CPU supports up to 1.6 times more concurrent AI agents while maintaining the same quality of service and achieving up to 2.2 times faster orchestration compared to alternative CPUs. This is achieved while simultaneously improving infrastructure utilization and cost efficiency. These findings underscore the Vera CPU’s ability to deliver the cost efficiency, low latency, and high throughput essential for production-grade agentic AI.

As AI agents undertake increasingly complex reasoning, planning, tool utilization, and data manipulation tasks, CPU performance has become a critical factor in orchestrating workflows around each model call. The Vera CPU, a key component of NVIDIA’s extreme codesign approach for AI factories, is specifically engineered for agentic workloads. DeepInfra’s benchmark findings illustrate how the NVIDIA Vera CPU empowers cloud providers to enhance infrastructure utilization, boost cost efficiency, and support a larger volume of concurrent AI agents without compromising service quality.

The introduction of the Vera Rubin platform and its associated technologies signifies a pivotal moment in the evolution of AI infrastructure. By focusing on extreme codesign, advanced networking, and operational efficiency, NVIDIA is enabling a new era of AI development and deployment, capable of tackling the most complex computational challenges of the future.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button