Algorithm · ByteDance · Hard
Requirements During the coding segment of an MLE internship phone interview, complete these two tasks in parallel: Write scaled dot-product self-attention using pseudocode, NumPy, or PyTorch. A one-head implementation is sufficient, although you may add multiple heads if time remains. Starting from the underlying probability model, derive and write the binary cross-entropy loss. Include the sigmoid probability, the loss for one observation, and the batch-average loss. As you…
Checking your access…