upvote
Everyone's definition of usable is different, but I disagree with your estimation of the specs required to be useful. I am able to do useful coding on Qwen3.8-27b, with 100k context and 16gb VRAM, q8 cache. I feel I'm living right on the cusp... my GPU is old (2016, pascal), so to get usable speeds I have to drop to a Q2 quant - which still gets stuff done, but the difference with q4 is noticeable. Q3 is close enough I don't really notice the difference between it and Q4, but it's too slow on my system. More context would be nice, but it's not that hard to work within ~100k.
reply
I’m using qwen3.8 quants (q3) effectively on a 5070 ti. A lot of it is about guardrails, but you also have to figure out how far (and in what ways) you can push a given model.
reply
27b runs perfectly fine on 2x 24GB at ~100 t/s (4090) with speculative decoding on 8 bit quants
reply
I bet! Just 2x24GB is not super basic hardware imho.
reply
What about Bonsai 2? You can fit Qwen3.8 27B on an 8GB GPU with it, and upstream llama.cpp support is already being worked on (they just got System1 support too).
reply
The Bonsai models are really bad when you actually use them for more than short responses.

Their marketing made it look like a breakthrough, but in my experience it’s just the next step down from the Q2 quants in both size and quality.

Q2 quants are already not very useful in my experience. The Bonsai models are even worse.

If you only need 80% plausible outputs that don’t need to reference a lot of context they can be useful. If you try to use them for real tasks it feels like time warping back to 2023 when you LLMs were barely useful if you babysat every word of the output.

reply
Have you actually used Bonsai 2 though and not just the original Bonsai? The experience is vastly improved but still requires a custom llama.cpp fork to use as of right now.
reply
The ternary model? Hopefully those are worth a damn in a few years, but currently just an interesting toy from what I understand.
reply
> Flash attention with who knows how much draft ensures it makes the same mistakes all the time and cannot call tools, format output or follow basic guidelines reliably.

> Draft MTP <= 2 so it doesn't trip

I am not sure you understand what either of these things do.

Do you think that FA or MTP are lossy?

reply
I mixed FA and MTP wrong in my original post, thanks for pointing it out.

My experience is that draft-mtp=2 gives 25% improvement in tokens/s but the model is unable to call tools with the right arguments reliably. I have since then gone for smaller quants at draft-mtp=1 and that problem is at least gone so far.

An agent needs to repeat tool calls in the right way, so cannot penalize repetition. This however leads to the agent retrying the wrong calls constantly.

reply
Can you please make your substantive points without snark or swipes? This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.

There's a lot of good information here but the comment spoils itself by coming across as aggressive in this way.

This is particularly a problem when responding to someone else's work. We need commenters to point out problems respectfully, not put down what other people have been making.

reply
Sorry, edited
reply
Appreciated!
reply