Hacker News
new
past
comments
ask
show
jobs
points
by
KeplerBoy
3 hours ago
|
comments
by
boredatoms
1 hours ago
|
next
[-]
It also depends on the runtime, vllm is unbelievably slow at model loading compared to llama.cpp
reply
by
teaearlgraycold
3 hours ago
|
prev
|
[-]
A lot of cloud platforms have terrible slow network storage. They also might need to compile the GPU kernels fresh as they might not have a persistent CUDA cache.
reply