It’s easier for the CSPs to move into hardware than it is for Nvidia to move into cloud hosting.
Although as a middle ground I’ve been quite happy with Nvidia Brev for on-demand GPU instances from a select marketplace of CSP offerings. It’s a well kept secret IMO — great product (from an acquisition iirc).
Also, not sure how well CSPs inference stack is compared with vllm + nvidia. A lot of open weight models uses MoE, making the inference stack more complex.