upvote
Same problem with every chinese model currently, they overthink way too much and take too much tokens and time.
reply
More or less, yeah. I've found mild success with deepseek-v4-flash though, and also Qwen3.5-122B-A10B-NVFP4 running locally, especially in terms of "doesn't overthink every single prompt" and somewhat reasonable quality. Really wishing for a 3.8 update of the 122B variant, that'd be really competitive (for local usage) :)
reply
A consequence of aggressive distillation?
reply
[dead]
reply