upvote
What model has to trade between those? I have both. You just have different, independently-optimized forward passes for each.
reply
I'm running on Strix Halo so memory bandwidth is my constraint. In that example I'm describing the choice between using ROCm or Vulkan. I have a llama-swap config that can call different instances of llama-server running a toolbox with either runtime.
reply