upvote
It also depends on the runtime, vllm is unbelievably slow at model loading compared to llama.cpp
reply
A lot of cloud platforms have terrible slow network storage. They also might need to compile the GPU kernels fresh as they might not have a persistent CUDA cache.
reply