Rubin GPU is a new-generation artificial intelligence (AI) accelerator introduced by NVIDIA. It was officially unveiled at CES 2026 in January 2026 and is scheduled to enter mass production in the third or fourth quarter of 2026. The chip completed key design milestones and tape-out in June 2025, with customer samples provided in September of the same year. In response to competition from AMD’s MI450 accelerator, NVIDIA subsequently increased the GPU’s power consumption from 1,800W to 2,000W and conducted a new tape-out.
Rubin GPU is manufactured using TSMC’s third-generation 3nm process technology (N3P) and CoWoS-L advanced packaging technology. It contains approximately 336 billion transistors. The compute dies are fabricated using the N3P process and integrated with one N5B-process I/O die through SoIC 3D vertical stacking technology, combining two compute dies with one I/O die. Rubin is also the first NVIDIA GPU platform to support eight stacks of HBM4 high-bandwidth memory, providing up to 288GB of HBM4 memory per GPU and memory bandwidth of up to 22 TB/s.
The Rubin GPU delivers up to 50 PFLOPS of NVFP4 inference performance, approximately five times that of the previous-generation Blackwell architecture. A single Rubin GPU can run AI models with hundreds of billions of parameters. For large-scale AI model training, the number of GPUs required can be reduced to approximately one-quarter of that needed by previous-generation systems, while overall throughput can increase by up to 10×. Its total interconnect bandwidth reaches 260 TB/s.
The accompanying Vera CPU features 88 customized Arm architecture cores and is interconnected through sixth-generation NVLink technology. The platform’s motherboard design integrates one Vera CPU and two high-performance Rubin GPUs.
The NVL144 rack-scale system supports up to 144 GPUs and can deliver 3.6 exaflops of FP4 computing performance. A derivative product, the Rubin CPX GPU, is specifically designed for long-context AI inference. The NVL144 CPX rack-scale system combines 144 Rubin GPUs, 144 CPX GPUs, and 36 Vera CPUs, delivering up to 8 exaflops of FP4 performance.
NVIDIA plans to introduce the Rubin Ultra NVL576 platform in the second half of 2027. The system is expected to feature four reticle-scale compute dies, 1TB of HBM4e memory, and up to 15 exaflops of FP4 inference performance. It will also introduce microchannel cold-plate cooling technology to address the substantially higher thermal and power requirements of next-generation AI computing systems.
AI Smart Device Home
