upvote
They say you can cluster up to four with a shared memory pool, and get three times the inference performance of a single machine.
reply
RDMA is buggy and Thunderbolt only delivers 1/10th the throughput of native connectivity. 1TB of Unified Memory w/ 1.2TB/s of bandwidth with marginally ~$30k cost is a different story than 1TB of sorta Unified Memory w/ an effective 120GB/s of bandwidth with a marginally ~$40k cost + all the RDMA bugs.
reply
You need latency for token parallelism, not bandwidth. Hence actual RDMA that bypasses the software TCP stack (ROCe or whatever).
reply