IT Times reported on September 1 that the latest industry survey by an analyst indicates that NVIDIA has restarted its “Rubin CPX” GPU project for AI inference prefill acceleration. However, the design has undergone substantial revisions, with production expected to begin in Q1 2027.
At the chip level, the new Rubin CPX delivers compute performance close to that of the standard Rubin GPU, while offering further improvements in prefill performance. Its maximum power consumption is also 2,300W, and it is equipped with 168GB of HBM4 memory.
At the compute-tray and rack levels, eight new Rubin CPX GPUs form one compute tray, while eight compute trays form one rack module. Scale-out connectivity within the rack module is provided by Spectrum-6 all-copper Ethernet. The new Rubin CPX uses a dedicated MGX ETL rack, with each rack containing one to four rack modules. Scale-out connectivity between the modules is implemented through a Spectrum-6 and OSFP optical interconnect solution.
At the system level, NVIDIA recommends pairing Rubin CPX with standard Rubin GPUs in a 1:1 configuration. The Rubin CPX ETL rack and the Vera Rubin NVL72 rack are interconnected using Ethernet-based RDMA.
The analyst noted that more than 50% of current AI inference workloads involve processing input data and building the corresponding KV cache. The new Rubin CPX can handle these prefill workloads with greater deployment flexibility and lower costs. Each Rubin CPX compute tray provides 1.34TB of HBM memory, which is sufficient to meet the requirements of most long-context prefill workloads and KV cache construction.
AI Smart Device Home
