The tradeoff is time (especially on RDMA4 hardware) - it does take a long time and spend a lot of tokens to get to the result, but I've found I can trust the results enough that I can queue a lot of work, essentially have it running all the time and achieve a decent velocity.
It's the first small local model I've felt like I can do real work with.
unsloth/Qwen3.8-27B-GGUF UD-Q3_K_XL DSH (pi)
Any tips?
we went from 62% completion to 92% using a claude code harness