Hacker News
new
past
comments
ask
show
jobs
points
by
cubefox
7 hours ago
|
comments
by
wmf
6 hours ago
|
next
[-]
The methods for splitting weights across multiple chips are well established. Groq/Cerebras can't hold a model on one chip either.
reply
by
pyrolistical
4 hours ago
|
parent
|
[-]
Umm I have an extra 35, do you have layer 6?
reply
by
octoberfranklin
2 hours ago
|
prev
|
next
[-]
They pipeline-parallelize across multiple chips. DeepSeek v4 Pro will be 30 chips.
reply
by
IsTom
6 hours ago
|
prev
|
[-]
I think it's enough that a single layer fits on each chip if you can daisy-chain them with good interconnects.
reply