upvote
The methods for splitting weights across multiple chips are well established. Groq/Cerebras can't hold a model on one chip either.
reply
Umm I have an extra 35, do you have layer 6?
reply
They pipeline-parallelize across multiple chips. DeepSeek v4 Pro will be 30 chips.
reply
I think it's enough that a single layer fits on each chip if you can daisy-chain them with good interconnects.
reply