Hacker News
new
past
comments
ask
show
jobs
points
by
LeBit
15 hours ago
|
comments
by
chorizo
13 hours ago
|
[-]
Had a similar experience. Llama.cpp compiled natively; parameter sweep to find best options fitting my use case for the qwen models with 16GB VRAM. The whole thing packaged into a portable container.
reply