upvote
Problem is that might go away or get nerfed.
reply
then you switch provider, it's not a monopoly
reply
If that happens you can still buy hardware later with almost certainly more (tok/s)/$ and better capabilities to run newer models more efficiently (remember native MXFP4?). Right now basically every generation of accelerator is adding new capabilities. These aren't yearly DirectX 9.0c-compatible GPU performance bumps.

As an individual, for average privacy needs (e.g. open source or at-home coding and automation), it's pretty much complete nonsense financially to self-host LLMs currently or select hardware now based on the capability to do so, and pay thousands of bucks extra.

reply
If you don't mind exfiltrating all your IP to the API provider
reply
Haha wow. I’m trying to even imagine the AI landscape in 15 years and I can’t.
reply
Instead of saying "I have a MBP with 64gb of RAM" you'll hear people say: "I'm subscribed to Model 9.x11B" and others will comment: "Oh dang, that's a nice model!"
reply