upvote
Yes, I was being facetious. Chinese labs are doing the most interesting research and publishing it. But Anthropic beats the "distillation attack" drum every time doubts surface about the depth of their technical moat. Of course, when they need a scary bogeyman, then the story shifts to how China is recklessly racing ahead building dangerously powerful models. That's the thing with Anthropic: they speak out of all 13 sides of their mouths.
reply
> China is recklessly racing ahead building dangerously powerful models

Their fear-mongering about GLM 5.3 got me to try it out. Its very good, I'll only go back to Claude if GLM isn't available (it forgot how to do tool calls yesterday).

Interestingly enough, it seems to compact at about 10% of the 1mn context, which makes sense if they're trying to run profitably.

reply
I like to do some discussions and back and forth with GLM. It doesn't hallucinate as much or gaslight like DeepSeek does. DeepSeek is great if you have a plan and you let it go. Also can plan a bit. But it has this sparse attention and especially the index_topk is too small in common deployments so it doesn't "see" all the context. And it then imagines things which is annoying if you chat with it about stuff you develop in the session.

I recommend trying Coralbricks with GLM due to their cheap input cache prices. GLM can get quite expensive elsewhere.

reply