System Design · Amazon · Hard
1. Main parallelism strategies Data parallelism How it works: Every device keeps a complete copy of the model. The mini-batch is split across devices, each device runs forward and backward propagation on its own batch slice, and then the gradients are averaged across all devices before the optimizer updates the weights. Pros: Very simple to implement. Scales computational throughput well when the model fits on each device. Minimal communication relative to model parallelism…
Checking your access…