I've been doing local inference for a couple of years on the side, and I'm astonished at the number of variables you need to have control over to get a reliable result. Inference engine, model parameters (top_k, temp, MTP-enabled/not), quant level, and harness all have a big impact on the results.
DS4-0731 at 2bit on llama.cpp (ROCm) and 250k context with omp.sh has been consistently reliable for me, just a bit slow (10 t/s) compared to what I'd prefer. Trying out DwarfStar today (benching it right now) to see if I can get better speed, but otherwise I've found it to be great on my side projects that are smaller (up to 10ksloc).
There could also be a domain issue - I tend to do lots of web programming and sysadmin work in these projects; if your work is more esoteric, it might not be nearly as good. I haven't tested much outside of my narrow domain.
Qwen 3.8 Flash Next is there. 3.8 27b is fairly close.
I'm excited to see what Qwen 4 will bring.
I'm running on a 128GB Strix Halo for Flash Next and an Intel Arc Pro B70 (32GB) for 27b.
I think there's probably low-hanging fruit to outsource reasoning from rote read/writes.... Just speculation though, I'm not a token optimization expert.