Algorithm · Oracle · Hard
This assessment contains two independent parts. Implement both. Part 1 — Multi-Head Attention from scratch Given three matrices Q (queries), K (keys), and V (values), all of shape (batch, seq_len, d_model), together with a causal mask flag causal, implement a standard multi-head scaled dot-product attention function. Split d_model into num_heads equal-sized slices of size d_k = d_model / num_heads. Assume d_model is divisible by num_heads. For each head: Transform Q, K, V…
Checking your access…