upvote
hence the "former deepseek v4 pro". I tried it out this morning and have had no complaints. I already liked glm 5.2
reply
> Deepseek v4 pro 0813

Which itself released yesterday? You're writing, reading, and evaluating enough software in a ~36 hour period to form, reject, and form another opinion about which model makes better _architectural_ choices?

reply
The sentence still doesn't make sense, because "settled on" implies a long testing phase with a verdict eventually emerging out of that.

What you're currently doing is "testing out"

reply
Honest question, how do you assess models this quickly? What metrics are you using? Would love to get my suite from multiple days and hundreds of prompts down to minutes. Got a few first pass tasks I run upon release for an initial experience, but those only work because even Fable and Sol fail despite objectively correct solutions existing, so it works because most models fail, but then, those are consciously not enough for coding, tool use, adherence or task specific inference and assessment…
reply
What are you working on? That can dictate which models are best.
reply
Anything from ML pipelines for language specific pruning over a Rust/JS/CSS mix codebase to assistance in motorcycle maintenance and different canvas coatings. Most of my evals build on those requirements, just takes a while to get any serious opinion on a model. Doubt anyone can do that in such short time, unless their tasks are so simple that most modern models not only succeed but could themselves accurately rate output. If even Fable still confuses PU or wax coated cotton canvas with a nylon shell, that needs experience for human assessment and the time that comes with it.
reply