upvote
It's likely that a stack of Macs will draw more power for slower prefill/decode than equivalently priced Nvidia GPUs. If power efficient inference is the goal, Macs are a non-starter.
reply
So if it isn't a comparative ability, now it's a power cost issue? This reads like goal post moving.
reply
Oh, it's absolutely both. The power you waste waiting for TFTT on prefill will absolutely compound at the "medium sized labs and businesses" scale.
reply