India, Sept. 16 -- Nvidia's NVLink Fusion is emerging as a way for hyperscalers and AI infrastructure providers to combine specialised AI accelerators with Nvidia's rack-scale architecture, potentially changing how inference systems are designed and deployed.

The technology is gaining relevance as AI infrastructure moves beyond a predominantly GPU-centric model. While GPUs remain central to large-scale AI training and inference, the growing demand for real-time AI services is creating a market for application-specific processors designed around particular inference characteristics, including low latency, memory efficiency and performance per watt.

The latest developments involving AI chipmaker d-Matrix and MediaTek illustrate how Nvidia...