upvote
Between GLM 5.3 Flash on my legacy Z.ai coding plan, and Qwen 3.8 Flash locally on my DGX Spark-like, I'm barely using my Anthropic/OpenAI subscriptions, likely to cancel them soon.
reply
It's been slow like molasses on the coding plan. I ended up using it more on fireworks. But DeepSeek is so much cheaper because the cache cost is much better.
reply
I’m kind of lucky that most of the time I don’t have to use it through peak times so it’s not that bad speed wise

The legacy plan I have is so good as to be basically unlimited usage for my workloads, so I’m kind of stuck with it til they stop renewing it haha

reply
GLM-5.3-flash is my implementation model after GLM-5.3 writes the plan.

It's an excellent workhorse. When I am running out of my GLM quota I switch GLM-5.3-flash to DS-4.1-flash.

reply
Same. It's barely even touching the credit I have on openrouter, it's great.
reply
Do you switch model mid session, or do you use subagent to do the switch after planning?
reply
Neither, I use different sessions for each step.
reply