What size context are you able to squeeze in with less than 2gb of headroom? I have had some luck using a quantized kv cache but i fear that also decreases overall quality.
replyTry the llama.cpp fork by thetom. It's called turboquant after the technique
reply