upvote
probably the best experience would be deepseek v4 flash 0731 (it takes about 170GB RAM on the server side for the full thing and RAM reserved for 1M context) via opencode's $10 a month plan until you use that up, it's either Q8 or full precision. Assuming you're ok with doing things with external inference.
reply
Why would they have listed how much vram they had if they were looking to rent gpu time on someone else's machine?
reply
A casual review of my comment history would show that I've been nothing but the biggest proponent of running models locally, and I do so myself a great deal. But one also has to be realistic about the capabilities of what you can do in a 16GB GPU these days. I already said an extra small Q2 quantization was effectively lobotomized so I didn't want to repeat myself.

This person has basically run into the limit of state of the art for even a modestly sized local model (this isn't deepseek v4 flash 0731 Q8 which I am running myself locally on a great deal more hardware), this is a 27B dense, but they're just not going to have a good time if they expect good quality results out of a Q2. The choices are either upgrade hardware or pay for external inference.

reply
Fine, but they already said they are using a Q3, so Q2 being unusable (disagree, but whatever) isn’t helpful new info.
reply