NVIDIA Unveils Breakthrough AI Factory Efficiency and Vera Rubin Infrastructure at Packed AI Infra Summit

The Santa Clara Convention Center transformed into a bustling hub of technological innovation on Tuesday, September 15, as the annual AI Infra Summit drew a record-breaking crowd of over 8,000 attendees—more than double the 3,500 participants recorded the previous year. Dubbed by industry insiders as the "Coachella of infrastructure tech," the event served as the backdrop for major announcements regarding the future of artificial intelligence hardware, energy efficiency, and data center scaling. Taking center stage during the morning keynote session was Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA. Addressing a packed auditorium, Buck detailed how modern AI workloads are evolving rapidly, shifting the baseline metrics of performance from raw compute speed to sustainable, highly optimized energy consumption.
The urgency for these architectural shifts is driven primarily by the explosive growth of agentic AI. Unlike traditional, straightforward request-response models, agentic workflows require systems to execute complex chains of reasoning, perform continuous tool calls, and dynamically spawn sub-agents. These processes drastically inflate context lengths and input-output token volumes, placing unprecedented demands on underlying data center infrastructure. To meet this paradigm shift, NVIDIA showcased its comprehensive, full-stack AI factory platform. This end-to-end ecosystem integrates upcoming Vera Rubin systems, Dynamo inference software, NeMo libraries, and advanced networking frameworks—including NVLink for high-bandwidth scale-up computing, Spectrum-X Ethernet, ConnectX SuperNICs, BlueField-powered context-memory storage, and BlueField DPUs dedicated to robust infrastructure security.
At the heart of NVIDIA’s latest strategy is a fundamental recalibration of data center performance metrics. The industry is rapidly moving away from evaluating infrastructure solely on peak performance toward a more practical standard: validated agentic tokens per megawatt. Because modern AI factories operate under severe power constraints imposed by regional electrical grids, these massive computing facilities must be codesigned from the foundational silicon all the way up to the electrical utility grid. NVIDIA’s newly introduced DSX MaxLPS system addresses this challenge directly, utilizing factory-wide power optimization to deliver up to 1.4 times more tokens per megawatt. Combined with NVLink technology, which unites massive clusters of accelerated computing hardware into a cohesive, high-performance system, these advancements ensure that data center operators can extract maximum productivity and financial return from every single unit of electricity consumed.

Unlocking Grid Capacity Through Automated Power Management
The intersection of artificial intelligence infrastructure and local power grid management took center stage with the unveiling of a strategic collaboration involving Silicon Valley Power and NVIDIA ecosystem partner Emerald AI. Operating within Silicon Valley Power’s territory, the initiative introduces an advanced flexible-load interconnection program designed to help AI data centers actively support regional grid stability rather than merely consuming resources. During initial testing, Emerald AI successfully integrated NVIDIA DSX Flex technology into its Conductor grid-responsive power management software, demonstrating automated load reduction across hundreds of live demand-response events without compromising the performance of underlying AI workloads.
Under traditional circumstances, data centers maintain a static power draw, creating severe bottlenecks for municipal utilities struggling to balance energy supply and demand. The integration of DSX Flex changes this dynamic entirely by enabling AI factories to act as responsive grid resources. The software continuously monitors real-time electricity grid signals, pricing fluctuations, and hybrid energy inputs, dynamically adjusting power consumption based on a predefined workload hierarchy. When the utility grid requires immediate relief, the system automatically sheds load by temporarily pausing lower-priority tasks, allowing critical computational workloads to proceed uninterrupted. Once grid conditions stabilize, paused jobs seamlessly resume. This automated load-balancing capability proves that controllable AI loads can help utility providers unlock critical electrical capacity, paving the way for sustainable infrastructure expansion without overwhelming local power grids.
Maximizing Performance per Watt with DSX MaxLPS

Cloud providers and hyperscalers are already capitalizing on these efficiency breakthroughs, with Lambda emerging as an early validator of NVIDIA DSX MaxLPS technology on Blackwell servers. Lambda shared compelling benchmark results at the summit, illustrating how continuous, dynamic power monitoring across GPUs and server racks can eliminate the inefficiencies inherent in static provisioning. Because artificial intelligence training and inference workloads exhibit vastly different electrical power profiles, static power allocation often leaves massive amounts of compute capacity stranded and unused.
MaxLPS solves this issue by continuously shifting available electrical power to the exact components where demand is highest. The operational impact demonstrated by Lambda was substantial: the cloud provider successfully operated 19 active nodes within the exact power budget typically reserved for just 16 full-power nodes. This optimization yielded a 24% increase in cluster-wide token throughput—scaling from approximately 4 million to 5 million tokens per second—while simultaneously boosting performance per watt by 23%. Looking ahead toward next-generation NVIDIA Vera Rubin NVL72 AI factories, DSX MaxLPS is projected to enable up to 40% more GPU capacity within a static megawatt budget, redefining the economic viability of power-constrained data center deployments.
Vera Rubin Architecture and Groq 3 LPX Integration
Addressing the relentless demand for higher token generation under strict power limitations, NVIDIA also highlighted the synergistic capabilities of its Vera Rubin NVL72 architecture paired with NVIDIA Groq 3 LPX. At the individual rack level, Intelligent Power Smoothing software and expanded energy buffering mechanisms absorb sudden electrical spikes, enabling hardware to operate much closer to sustained peak demand and turning previously wasted headroom into productive compute power.

For complex agentic AI applications where latency and context scaling compound rapidly, Groq 3 LPX provides deterministic, ultralow-latency inference capabilities. When integrated with Vera Rubin systems and managed via DSX MaxLPS, the combined platform delivers up to 35 times higher token throughput per megawatt compared to previous-generation GB200 NVL72 setups when handling massive models exceeding two trillion parameters at extended context lengths. In rigorous evaluations utilizing a 100,000-context Qwen 3.8 27B workload, Groq 3 LPX achieved an impressive 2,529 output tokens per second per user. This computational headroom empowers AI agents to execute significantly more reasoning steps and API tool calls within a standardized response window, even as operational scale increases.
Real-World Validation via SemiAnalysis AgentX
To accurately measure these performance gains, the industry is increasingly relying on sophisticated evaluation frameworks that mirror real-world usage rather than synthetic benchmarks. SemiAnalysis introduced its AgentX dashboard, which evaluates inference capabilities by analyzing recorded, real-world agentic coding sessions complete with authentic context growth, tool call delays, and multi-tier sub-agent spawning. According to live data on the SemiAnalysis dashboard, the NVIDIA Vera Rubin NVL72 running the DeepSeek V4 Pro model delivers up to 30 times higher throughput per megawatt than previous-generation NVIDIA GB300 NVL72 systems.
Because agentic workflows generate input token volumes roughly 15 times larger than conventional chat prompts, capturing their performance requires robust end-to-end testing methodologies. The extreme efficiency documented by AgentX translates directly into profound economic advantages for data center operators. Achieving up to 30 times higher throughput per megawatt means facilities can generate drastically more commercial value from an identical energy footprint, while achieving up to 45 times lower cost per million tokens. This leap in productivity is the direct result of deep hardware-software codesign, encompassing the NVL72 scale-up domain, sixth-generation NVLink interconnects, fifth-generation Tensor Cores utilizing NVFP4 precision, and optimized inference software stacks like TensorRT-LLM and NVIDIA Dynamo.

Startup Ecosystem Success and the NVIDIA Vera CPU
Beyond massive hyperscale deployments, the broader technology ecosystem is rapidly adopting NVIDIA’s latest hardware innovations, with emerging startups reporting exceptional performance metrics across a diverse array of workloads utilizing the NVIDIA Vera CPU. Perplexity integrated the Vera CPU into its SPACE secure sandbox platform for agentic AI, noting a 1.9-fold increase in sandbox startup speeds. Similarly, Daytona subjected the processor to intensive agentic execution tests, reporting substantial performance gains for complex multi-step workflows.
In analytical database testing, ClickHouse benchmarked the Vera CPU on ClickBench, identifying it as the fastest machine measured to date and calling it a strong indicator of future CPU capabilities for data-intensive tasks. DeepInfra reported sweeping victories across multiple benchmarks, highlighting a 2.2-fold improvement in orchestration step latency. Furthermore, Prime Intellect confirmed that the Vera CPU maintains exceptional memory bandwidth and consistently low latency even under heavy parallel workloads—a critical requirement for predictable agentic execution. Additional performance validations came from Redpanda, which demonstrated 5.5 times lower latency and 73% higher throughput; Starburst, which achieved triple the query throughput of competing processors; and Kinetica, which recorded a 2.7-fold improvement in analytical query speeds.
Multilayer Resiliency with NVIDIA NVLink 6

As artificial intelligence factories expand to encompass hundreds of thousands of interconnected GPUs, maintaining operational reliability becomes just as important as raw computational speed. Training and inference workloads require uninterrupted continuity, meaning that transient electrical errors, signal degradation, and hardware anomalies pose constant operational risks. To safeguard massive compute clusters against these disruptions, NVIDIA engineered NVLink 6 with a comprehensive, multilayer resiliency architecture designed to detect, contain, and autonomously recover from faults before applications experience downtime.
At the physical layer, the architecture incorporates custom forward error correction, physical layer retries, and universal recovery protocols to preserve a lossless communication fabric while minimizing latency overhead. At the network layer, advanced credit-based flow control, dynamic routing mechanisms, and automated link rebalancing isolate faults locally, preventing cascading network stalls that could otherwise severely degrade overall AI factory throughput. By combining unmatched energy efficiency, breakthrough token generation capabilities, and enterprise-grade reliability, NVIDIA’s latest infrastructure ecosystem establishes a definitive blueprint for the future of scalable artificial intelligence production.







