Back to problems

Debug a GPT-Style Transformer with KV-Cache Generation

AI Coding · OpenAI · Hard

Requirements Follow-up task (one of the following): Convert to a classifier: swap the last head for a classification head and revise prediction plus loss logic; some interviewers want mean pooling before producing the final logits. Add a KV cache: a starter class is supplied, and you must connect caching to attention, update positional-embedding treatment, and thread through the required parameters. This is fairly direct for candidates who have practiced it. Expected…

Checking your access…