upvote
Yeah, they all do it though. And now that models are being natively trained at FP8/NVFP4, I’m not sure it matters.

Only way you can really know you’re getting the full model is to host it yourself. Every inference provider has every reason to lie and it’s impossible to find out the degree to which they are.

reply