upvote
Just run /goal to optimise it and you should be good in less than an hour. Also best to use models that support speculative decoding.
reply
Optimize llama.cpp? Hmm.

Wrt speculative decode, basically zero finetunes keep it. I'm testing with some ridiculous abliterated amalgamation so spec decode has been gone for most of its ancestry.

reply