upvote
I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

reply
This is interesting, thanks. - https://github.com/Neroued/ninfer
reply
I prefer https://ornith.ai/ornith_1_5.html to Qwen 3.8 not only because it is much faster on my hardware but better responses.

But this Qwen 3.8 Flash next coder is amazing running with Strata.

reply
Is this true for 27b Q4_K_XL vs flash next IQ3_S? I thought under Q4 models start quickly degrading?
reply
While this is generally true, it's _a little_ less true the larger the model is.

Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.

reply
To add to the other comment, there's also Ridge quantisation - the majority of weights are indeed Q3_x, but the most sensitive layers are FP8.
reply
Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.
reply
I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.
reply
do LLMs tend to be homesick when not used in the same harness they sat in during some training phase?
reply
iirc there was a sectionin Qwen’s paper where they talked anout how they post-trained flash or 3.8 to work just as well regardless of the harness or eval used. I think that used to be true but not sure if it is any longer
reply