Algorithm · Meta · Medium
Question: Build Scaled Dot-Product Attention Create an implementation of single-head scaled dot-product attention. The input consists of three matrices: Query matrix Q, whose dimensions are $$n \times d$$ Key matrix K, whose dimensions are $$n \times d$$ Value matrix V, whose dimensions are $$n \times d$$ Return the matrix defined by: \[ Attention(Q,K,V)=softmax\left(\frac{QK^T}{\sqrt d}\right)V \] Apply softmax separately to every row. Input Format Output Format Produce the…
Checking your access…