Even now on the $200 plan I use up my Fable credits in a single day and had to start using codex and openrouter for more usage because Fable burns $100s an hour when billed on usage.
It became an easy decision, even the $200/month by Anthropic sounds like a bad deal.
It's also already the case that in larger companies people do not have access to these plans as employees and must use API rates.
There's also the political angle. When Anthropic and OpenAI held back their top of the line models because of Bessent & Trump's bullshit, and threatened to deny unwashed foreigners like me access... I dropped my Codex plan and made do purely with GLM 5.2 for three weeks before OpenAI finally released 5.6 Sol. Feels inevitable that this will happen again.
Or, somebody will come up with a way to serve e.g. Kimi K3 or the new Qwen model in an extremely cheap way. Or DeepSeek releases a competitive model at their cut-throat rates. And then the cost argument just wins.
we'll get new quants, dspark speculators, distills and optimized kernels
as long as there are near frontier models available there will be inference providers selling them at or below cost of inference in attempt to get market share.
This is a very large model. Much larger (3x) than GLM. The resources to run it are very expensive.
I just looked at DeepInfra -- I've got an account there already etc -- and it's at FP4 quant. How much that effects the quality of inference for GLM, I can't say. I could see using it as a backup when other things run out but don't think I'd trust it.