Back to problems

Attention Implementation + Flash / Linear Attention Follow-ups

Algorithm · Meta · Hard

Requirements Implement the conventional attention layer used by Transformer models. Use the following function interface: Clearly document the tensor dimensions of the query, key, and value inputs. Form the attention logits, normalize them, and use the resulting weights to aggregate the values. Give the algorithm's runtime and space costs. Follow-ups: describe how FlashAttention differs, the reason for its name, and the computation that linear attention aims to approximate…

Checking your access…