Thank you, found some benchmarks and they look really promising. To my understanding those Neural Accelerators are like AMX but for GPUs. With those accelerators the GPU performance on a M5 Max in LLM inferencing would totally be on par with a 5090, that's quite impressive!
replyWhich inference benchmark are you looking at? The prefill speeds should be comparable on some LLMs, but the 5090 has much a higher theoretical max decode speed.
reply