System Design · Waymo · Hard
Imagine you are part of the team responsible for the runtime that executes machine learning models on an accelerator. A particular model shows elevated end-to-end latency, and at times the runtime runs short on memory. Your job is to describe how you would investigate and improve this situation. Start by explaining how you would profile the workload and pinpoint where time and memory are going. Then discuss choices for tensor-level computation, including how data is laid out…
Checking your access…