Algorithm · OpenAI · Hard
Work with an existing decoder-only Transformer language model implementation that contains four or five intentionally introduced bugs. You must first locate and correct those bugs, then build key-value (KV) caching from scratch so that autoregressive decoding reuses attention keys and values from earlier positions instead of recomputing the entire prefix for every new token. The buggy source file is not included here. Unless stated otherwise, assume a standard decoder-only…
Checking your access…