But the cool thing is that if your NN is split into mostly self contained chunks then you can go widthwise parallel.
An architecture like MOE exploits this fact so that the active weights during pre-training you're backproping only through active experts