Algorithm · Google · Medium
Task Write the forward pass for a simplified Transformer block. The inputs are a sequence tensor and the block's weight parameters; calculate the resulting tensor Y using the operations below. The block receives: A sequence representation X with length T and hidden width D Single-head self-attention weights Wq, Wk, Wv, and Wo Feed-forward network parameters W1, b1, W2, and b2 An optional attention mask that can prevent attention to selected positions Use these computations,…
Checking your access…