Back to problems

PyTorch Multi-Head Self-Attention

Algorithm · Uber · Hard

PyTorch Multi-Head Self-Attention Problem Overview Implement a multi-head self-attention module in PyTorch. The exercise evaluates whether you can correctly handle: separate query, key, and value projections reshaping tensors so each attention head is processed independently scaled dot-product attention combining the head outputs and applying the final projection Use the following conventions: The input tensor x has shape (batch_size, seq_len, d_model). num_heads specifies…

Checking your access…