Their marketing made it look like a breakthrough, but in my experience it’s just the next step down from the Q2 quants in both size and quality.
Q2 quants are already not very useful in my experience. The Bonsai models are even worse.
If you only need 80% plausible outputs that don’t need to reference a lot of context they can be useful. If you try to use them for real tasks it feels like time warping back to 2023 when you LLMs were barely useful if you babysat every word of the output.
> Draft MTP <= 2 so it doesn't trip
I am not sure you understand what either of these things do.
Do you think that FA or MTP are lossy?
My experience is that draft-mtp=2 gives 25% improvement in tokens/s but the model is unable to call tools with the right arguments reliably. I have since then gone for smaller quants at draft-mtp=1 and that problem is at least gone so far.
An agent needs to repeat tool calls in the right way, so cannot penalize repetition. This however leads to the agent retrying the wrong calls constantly.
There's a lot of good information here but the comment spoils itself by coming across as aggressive in this way.
This is particularly a problem when responding to someone else's work. We need commenters to point out problems respectfully, not put down what other people have been making.