ByteDance · ML & AI Fundamentals
Define QKV for recommender cross-attention
TrueInterview
October 7, 2026 · 1 min read
You are building a deep-learning recommendation system that uses a Transformer-style cross-attention block to capture the interaction between a user and a candidate item. The model typically receives the following inputs:
- A user behavior sequence: the items the user has interacted with previously, each already converted to an embedding vector of size .
- A candidate item whose relevance score is to be predicted, also represented by an embedding vector of size .
- Optional context features (time, device, location, etc.) that may also be embedded. You choose to place a cross-attention layer somewhere in the model instead of relying only on self-attention.
- Describe one concrete way to set the Query (Q), Key (K), and Value (V) tensors in this cross-attention block from the inputs above. Explain the semantic role of each of Q, K, and V.
- Provide at least two different reasonable design choices for assigning Q, K, and V (for instance, one where the candidate item acts as the query and one where the user history acts as the query). For each design, explain:
- Which inputs are used as Q, K, and V.
- What interaction the attention mechanism is modeling.
- Pros and cons, or the situations where that design is preferable.
- Briefly explain how cross-attention here differs from self-attention within the user behavior sequence, and why cross-attention can be useful in recommendation systems. Overview: This question tests understanding of Transformer-style cross-attention and the concrete construction of Query, Key, and Value tensors for deep-learning recommender systems, including representation semantics, embedding alignment, and interaction modeling among user history, candidate items, and context.
Loading comments…