upvote
I have a 36gb M3 Max. I tested it across quite a few different options: llama.cpp, oLMX, ollama with different options.

So far ollama managed to be the most performant of them all. I will get 30 to 40 tokes/sec with it when using the -mlx version of Qwen3.8.

Whatever the sauce the ollama folks baked into the mlx + MTP mix is currently working the best out of the box.

reply