Create
NVIDIA's Vera Rubin Architecture Promises Tenfold Reduction in AI Inference Costs

NVIDIA's Vera Rubin Architecture Promises Tenfold Reduction in AI Inference Costs

Arkadiy Andrienko

At CES 2026, NVIDIA CEO Jensen Huang unveiled the next generation of computing architecture for artificial intelligence — Vera Rubin — which may later form the basis for new gaming graphics cards. The company has already begun full-scale production of the new chips, with shipments to customers starting in the second half of this year.

The new platform is positioned as the successor to the current Blackwell generation. According to NVIDIA, the Rubin GPU delivers a fivefold performance increase in inference tasks compared to its predecessor. For training large language models, efficiency improves by 3.5 times, which will reduce the number of required GPUs by a factor of four for similar tasks.

The Vera Rubin architecture is based on six specialized components:

  • The Vera Central Processing Unit with 88 cores based on the Armv9.2 architecture;
  • The Rubin Graphics Processing Unit delivering 50 petaflops of performance when working with the NVFP4 format;
  • A 6th-generation NVLink switch with a bandwidth of 3.6 TB/s per GPU;
  • A ConnectX-9 SuperNIC network adapter and a BlueField-4 data processing unit;
  • A Spectrum 6 Ethernet switch.

Special attention is paid to energy efficiency and reliability. According to the company, the new Spectrum-X Ethernet Photonics technology with integrated optics reduces power consumption fivefold while increasing connection reliability tenfold. Jensen Huang emphasized that the transition to the new manufacturing process enabled this leap with only a 1.6-fold increase in the number of transistors on the chip. Among the first to receive the new platform will be cloud providers CoreWeave and Microsoft Azure.

NVIDIA's Vera Rubin Architecture Promises Tenfold Reduction in AI Inference Costs

The announcement comes amid increasing competition in the AI accelerator market. In addition to its traditional rival AMD, NVIDIA also faces competition from major clients like Google, who are developing their own chips. In response, the company is betting on a substantial increase in performance and cost efficiency: the cost per token for AI inference is expected to drop by a factor of ten.

    About the author
    Comments0