upvote
Meanwhile the GB300 used by hosted llms:

GPU Memory Bandwidth: 7.1 TB/s Interconnect Bandwidth: 900 GB/s bidirectional

https://pi3g.com/nvidia-gb300-specifications-including-memor...

If you think M7 will hit even 15% of these speeds you're very optimistic.

reply
A hosted instance serves multiple customers at a time. A local model only one.
reply
How many though? At 1m context you quickly fill a full gb300's 280gb of memory
reply