Back to problems

Hand-Code Self-Attention and Cross-Entropy

Algorithm · ByteDance · Hard

Requirements During the coding segment of an MLE internship phone interview, complete these two tasks in parallel: Write scaled dot-product self-attention using pseudocode, NumPy, or PyTorch. A one-head implementation is sufficient, although you may add multiple heads if time remains. Starting from the underlying probability model, derive and write the binary cross-entropy loss. Include the sigmoid probability, the loss for one observation, and the batch-average loss. As you…

Checking your access…