upvote
Tokens per second is almost entirely memory bandwidth at inference time, training obviously needs more compute but you can add more chips for that.
reply
According to SemiAnalysis, both inference and post-training (RLVR) is mostly memory bandwidth bound. Only pre-training is compute bound, but it now only takes a small share of overall data center capacity.

https://x.com/EugeneNg/status/2099315982959616369

reply
Huawei uses their own non-standard HBM called HiZQ probably not produced by CXMT.
reply
China should invest in an analog inference chip. It's a hail mary but why not.
reply
They can probably afford to do both.
reply
I read HBM yields are 25-30% (vs 80-90%) making them 3 to 5 times as expensive. They are 4-5 years behind, that probably means 1-2 in Chinese time.
reply
Does that include yield from packaging?
reply