upvote
> Qwen 3.6 27b locally a few months ago

I've been doing local inference for a couple of years on the side, and I'm astonished at the number of variables you need to have control over to get a reliable result. Inference engine, model parameters (top_k, temp, MTP-enabled/not), quant level, and harness all have a big impact on the results.

DS4-0731 at 2bit on llama.cpp (ROCm) and 250k context with omp.sh has been consistently reliable for me, just a bit slow (10 t/s) compared to what I'd prefer. Trying out DwarfStar today (benching it right now) to see if I can get better speed, but otherwise I've found it to be great on my side projects that are smaller (up to 10ksloc).

There could also be a domain issue - I tend to do lots of web programming and sysadmin work in these projects; if your work is more esoteric, it might not be nearly as good. I haven't tested much outside of my narrow domain.

reply
I started using Claude right before 4.5 came out, and 4.6 is where it turned a corner for my use.

Qwen 3.8 Flash Next is there. 3.8 27b is fairly close.

I'm excited to see what Qwen 4 will bring.

I'm running on a 128GB Strix Halo for Flash Next and an Intel Arc Pro B70 (32GB) for 27b.

reply
I just used 5.5 xhigh reasoning to make a massive implementation spec (for a vibey throwaway project/exploration, not anything important, burned 80% of the 5h window), now my Strix is in the process of implementing it.

I think there's probably low-hanging fruit to outsource reasoning from rote read/writes.... Just speculation though, I'm not a token optimization expert.

reply