upvote
Sure, but shouldn’t the programs to run the LLMs go “the user has this much vram and the model is this size, so I’ll start with sensible defaults based on that”?

You could override, obviously.

reply
Yes, llama.cpp does that.
reply