upvote
To be fair a lot of models have quirks, I've never found a model swap that was transparent.
reply
Yes, if you set reasoning to none you can force the granularity of the thinking.

It will actually adhere to your request for e.g. 3 sentences max.

Thinking mode will override any instructions in the prompt (at least for other models in my experience).

Of course this will probably hurt performance, but works great for easy tasks that you know are trivial. Tons of pipeline, image recognition etc use cases where this works well.

I'd be curious to see Qwen 3.8 27B low thinking benchmarks though.

reply