AI Smart Device Home

News

News

NVIDIA Reportedly Restarts Rubin CPX Inference Prefill Chip Project with Major Design Overhaul

According to a report from ITHome, analyst Ming-Chi Kuo stated yesterday that his latest industry research indicates that NVIDIA has restarted its “Rubin CPX” AI inference prefill accelerator GPU project. However, the design has undergone substantial revisions, with production expected to begin in Q1 2027.


At the chip level, the new Rubin CPX delivers computing performance close to that of the standard Rubin GPU, with further improved prefill performance. Its maximum power consumption per unit is also 2,300W, and it is equipped with 168GB of HBM4 memory. (ITHome note: the previous Rubin CPX was equipped with 128GB of GDDR7 memory; the 168GB capacity may indicate a 7 × 24GB configuration.)

At the tray and rack levels, eight new Rubin CPX GPUs form one compute tray, while eight compute trays form one rack module. Internal scale-out connectivity is provided through Spectrum-6 all-copper Ethernet. The new Rubin CPX uses a dedicated MGX ETL rack, with each rack containing one to four rack modules. Scale-out connectivity between rack modules uses a Spectrum-6 + OSFP optical solution.

At the system level, NVIDIA recommends deploying Rubin CPX and standard Rubin GPUs in a 1:1 ratio. Connectivity between the Rubin CPX ETL rack and the Vera Rubin NVL72 rack is based on Ethernet RDMA.

The analyst commented that more than 50% of current AI inference workloads are devoted to processing input data and building the corresponding KV cache. The new Rubin CPX can handle prefill workloads with greater deployment flexibility and lower cost. A new Rubin CPX compute tray provides 1.34TB of HBM memory, which is sufficient for most long-context prefill workloads and KV cache generation requirements.

PREVIOUS:NVIDIA Takes a $400 Million Charge in H1 Due to Excess H200 Inventory

NEXT:RUBIN

Leave a Reply

+852 4748 9818

84798770582

yba960270@gmail.com

Leave a message