Have you tried it with MTPLX? I get around 30 tok/s with it, also on an M1 Max with 64GB.
I tried it with the author’s 4-bit quant of Qwen 3.8 27B: https://huggingface.co/Youssofal/Qwen3.8-27B-MTPLX-Optimized... (but no need to download it manually; MTPLX will ask which one you want).