upvote
Imagine we did that, split up a model layers as A->B->C. C will need to wait for B to compute a forward pass, which is waiting for A to compute its forward pass. To compute the forward pass, B needs all of the outputs from A, which is an upload and a download (maybe these can be done concurrently).

Then A waits for B to compute its backwards pass, which is waiting for C to do the same thing. Again you are sending around potentially gigabytes of data.

This is in contrast to mining bitcoins for example which doesn’t require any coordination from miners because their work is completely independent, and the answer is very small compared to the work needed to get it.

reply
Yep. You need to transport all the weights at the boundary regardless of Backprop/NPC.

But the cool thing is that if your NN is split into mostly self contained chunks then you can go widthwise parallel.

An architecture like MOE exploits this fact so that the active weights during pre-training you're backproping only through active experts

reply
The problem (and contrast with other approaches) is that mat muls requires synchronization. Arranging your networking and training structure to maximize compute and minimize communication is the main craft of ML training infra folks. In your example, yes you can compute layers on different machines (i.e. Tensor Parallelism), but you must be very careful in how you arrange it.
reply