Algorithm · OpenAI · Medium
Sharded Matrix Multiplication with Gradients Hard · Matrix Multiplication, Distributed Computing, Backpropagation, Debugging · General · Hints: Chain rule, required transposes, all‑reduce sum Implement the forward and backward passes for the matrix product $$C = A B$$ when $$A$$ is split row‑wise across multiple devices and $$B$$ is replicated on every device. You are given: A list A_shards of length n_dev, where each element is a local slice of $$A$$ with shape (m_local,…
Checking your access…