India, July 27 -- Nvidia's Vera Rubin production ramp represents more than the beginning of another accelerator cycle. The platform changes how investors should evaluate AI infrastructure because its principal performance claim is no longer based solely on GPU speed.

Nvidia is now emphasizing how many tokens an integrated rack-scale system can generate from a fixed amount of electrical power. In an industry increasingly constrained by grid capacity, power density and cooling requirements, tokens per megawatt may become a more economically important measure than peak processor performance.

On July 21, Nvidia confirmed that Vera Rubin NVL72 production was ramping, with systems already operating at CoreWeave, Google Cloud, Microsoft Azure ...