upvote
GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or...

Also a lot of questions to benchmark because opus 5 is completely useless model right now.

I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?

reply
Opus 5 generates really good code and terminal commands though. It's just bad at the accompanying text it tells you. These benchmarks don't grade the text generation of the response I don't think, only the task outcome.
reply
Terminal Bench 4.0 did not introduce new questions. All tasks were public for a while. If you look at GLM 5.2, which is using the same base model as 5.3, but was released prior to most tasks, it does extremely horribly on terminal Bench 3.0 (4 to 8 times worse than every other model) - source: https://benchlm.ai/benchmarks/terminal-bench-3
reply
Astra is 58%. The current title says it's "rivaling Astra"
reply
It is rivaling Astra, on their own benchmark that they made (FrontierCode), that they ran themselves in their own closed-source ecosystem that isn’t reproducible by anyone.
reply
reply
I have been using it today the whole day and it is definitely better than Luna. Great that bench agrees.
reply
This explains a lot about Sonnet 5.
reply
So, better than Sonnet and Luna? lol
reply