upvote
Not sure I have a strong opinion but I am sort of okay with current state of affairs. I am optimizing for V100. EOL cards on EOL CUDA. Llama is a good enough base for this. A couple weeks of grunting at Claude has gotten the inference /fast/ for my uses. 150-160 t/s on 2 GPU for 27B and 125 t/s on Flash Next. Asking them to upstream every random feature does not make sense. They sacrifice a lot of speed to maintain stability and a reasonable feature set that works across a diverse range of models and systems. They could maybe merge some features like this and gate them on flags a little faster, but you can cobble together what you need and the big models can figure out how to make it fast.
reply
Does anything else support Pascal gpu's though?
reply