upvote
Sure, but that's for "personal" serving. I meant for 3rd party providers. Usually we get a good indication on what it costs to host this, as the prices settle on open router. That's why I said it's tougher to serve than kimi k3 on launch. As a provider you'd do fp8 if the model creator didn't do QAT on q4, or until someone does a good calibrated nvfp4. And that's usually nvda :)
reply
That makes sense, but your specific phrasing precluded the possibility of non-QAT quantization.
reply
Should have worded that better, my bad.
reply
QAT is an optimizing quantization algorithm, not naive quant.
reply
Right, but the way they phrased it suggested that without QAT it could not be quanted at all.
reply