How so? In tokens per second when running major open-weights models, or something else?
Nvidia has CUDA, AMD has CDNA, and Apple has... compute shaders, I guess?
How much it matters in inference? Most GPUs have enough computing for that and the bottleneck is the RAM speed and size. And M5 Ultra is becoming to challenge this.
It's kinda why memory bandwidth is an enormous red herring, even for datacenter applications. Nvidia's huge advantage is a compute-optimized GPU architecture and their Infiniband networking, their memory controllers aren't really the star of the show.